logo

Annotation, training and models

We train our own models.
None of them diagnose anything.

Every radiology product says AI now, and almost none of them say which kind they mean. Ours is trained here, on radiology images, and it works on the clerical half of the job: what was scanned, what is written on the film, and what the radiologist just said. This page is the whole of it, including the parts that are unglamorous.

Where it stands today
Models in service
One
What it does
Recognises the body part
Models that diagnose
None
Regulatory clearance
None
Who approves a correction
A person
Measured on the live system on 12 September 2026 and rounded.
82%
Top answer correct, body part
On about 5,700 held-back images. An engineering measure, not a clinical one
140,000
Predictions since late June
At 0.89 average confidence
5 min
Until a new model reaches every site
Swapped in without a restart
0
Models that diagnose anything
And none in training

What runs today

Four jobs, all of them clerical, all of them constant.

None of these is a diagnosis. Each is something a person does by hand many times a day, and each one is measured on the live system rather than estimated.

  • Recognises what was scanned

    A model we trained reads the image and says which body part and view it is, so the study files itself against the right procedure. It is the one doing the volume: about 140,000 predictions since late June at an average confidence of 0.89.

  • Reads the text burnt into the image

    The letters a radiographer puts on a film, marking left or right and the projection. About 176,000 images have been read this way, and each signal carries the anatomy, view and side it found.

  • Types while a radiologist talks

    Dictation is transcribed on our server rather than in the browser, so it behaves the same on every machine and does not depend on what speech engine a device happens to have.

  • Segments where you click

    In the viewer you click inside a structure and it outlines it, Ctrl-click to exclude, then accept or reject. It runs on our server so the browser downloads nothing. It finds nothing on its own: it outlines what you point at.

How it gets better

The system argues with itself, and a person settles it.

A model that trains on its own output drifts, and one that waits for people to label everything never gets anywhere. So this sits in between, and the rule is strict. Two independent signals must agree with each other and disagree with what the record says before anything is proposed. One signal on its own is never enough to raise a case.

When the image model and the text reader both say the study is a knee and the paperwork says something else, that becomes a proposed correction carrying its evidence: the prediction, its confidence and what each signal saw. A person opens it next to the actual image, and approves or rejects it. Only an approved correction changes the next training set.

Nothing is ever applied because the machine was confident. A person approves every correction.

Being straight about the scale: this is careful work done by a small number of people, not a labelling floor. What makes it work is that the system finds the cases worth looking at.

Corrections on the live system
Decisions recorded
About 16,000
Relabelled
About 12,000
Confirmed as correct
About 3,500
Thrown out of training
About 300
Applied without a person
None
Every decision carries a written reason. Roughly three in four were raised by the system, the rest by hand.

Before a single image trains anything

Most of the work is throwing data away.

The taxonomy is versioned, and each training run records the exact revision it used. Editing it changes the next run and can never rewrite what a published model was trained on.

The annotation workspace

Drawing on images, inside the system that already holds them.

Labelling radiology images usually means exporting them to a separate tool, which means copying patient data somewhere else and reconciling it afterwards. Here the workspace is the same viewer radiologists already use, reading from the same archive, so nothing is exported anywhere.

Reviewers work a queue of studies that have already been reported, so the labelling effort never delays a patient. What they draw is stored against the exact image and frame it belongs to.

An honest limit: there is no second-reader workflow here. Annotations are not independently adjudicated and no agreement between annotators is measured. If you need that, it is a conversation to have before you plan a project, not after.

What a reviewer can draw
A box
Yes
An outline or freehand region
Yes
Points and arrows
Not as training data
Where it is stored
Against the exact image and frame
Saving
Automatic, as you draw
A separate queue of already-reported studies, so annotation never sits in front of clinical work.

How a model reaches a site

Trained in one place, serving everywhere, switched on by a person.

A trained model is stored with everything needed to judge it: accuracy per class, what each class was most often confused with, the training curves, and how long it takes on one image. It is created switched off, and stays off until somebody looks at those numbers and promotes it.

Once promoted, sites collect it themselves. Each installation checks for a newer version every few minutes and swaps it in memory, so nothing restarts and nobody visits. If a model arrives whose output does not line up with its own list of labels, it is refused outright and the previous model keeps serving, because a mislabelled answer is worse than an old one.

The path a model takes
After training it is
Inactive
What is kept with it
Metrics, curves, confusion matrix
Who turns it on
An administrator
Sites pick it up in
About 5 minutes
If its output looks wrong
Refused, previous one kept

What we will not say

The list matters more than the claims.

The accuracy figure on this page describes a filing task and was measured the way engineers measure such things. It is not evidence of how a model behaves on a population it has never seen, and we will not offer it as though it were.

The segmentation tool is a well-known general-purpose model, not one we trained and not one specialised for medicine. The transcription and the text reading are third-party models too. We build the pipeline around them and we say which parts are ours.

And nothing here writes a report. There was an experiment in that direction and it was deleted rather than shipped, which we think was the right call and would rather tell you than have you read it somewhere.

Say it plainly
Detects disease
No
Clinically validated
No
Cleared by a regulator
No
Writes your report
No
Segmentation finds things itself
No, you click
Every one of these is something other products in this market imply. We would rather be the ones who wrote the list down.

Straight answers

The questions people actually ask about AI

Does your AI detect abnormalities?
No. There is no model in this platform that detects a disease or a finding, and we have not trained one that is in service. Everything running today is about filing the study correctly and saving typing. If someone tells you our software flags pathology, they are mistaken.
Then what is the AI actually for?
Removing clerical work from radiology. Knowing which body part was scanned, reading the markers burnt into an image, and turning speech into report text are all jobs that people currently do by hand hundreds of times a day. That is where the time is.
How accurate is the model that is running?
The one in service gets the body part right about 82 in 100 times as its first answer, and about 97 in 100 within its top five, measured on roughly 5,700 images it had not trained on. That is an engineering measurement of a filing task, not a clinical result, and we will not dress it up as one.
Is it validated clinically, or approved by a regulator?
Neither. No sensitivity, specificity or area-under-curve figure is computed anywhere in this system, there is no reader study, and there is no clearance from any authority. Anyone who needs those things needs a different product, and we would rather say so in the first meeting.
Does it learn from our data automatically?
Not automatically, and that is deliberate. The system proposes corrections when its own signals disagree with the recorded label, a person approves or rejects each one, and training is started by hand afterwards. A new model then sits inactive until an administrator promotes it.
Can we label our own data for our own models?
There is an annotation workspace in the viewer where reviewers draw boxes and outlines on images, and those annotations can train a detection model. Be clear-eyed about what that is: it is a labelling tool and a training pipeline, not a validated clinical product at the end of it.
Do the models run on our hardware or yours?
Ours, on the server. The same build runs on a graphics card or on an ordinary processor, so a site without a card still gets the same answers, more slowly. No model weights are sent to a browser.
What happens when a new model is trained?
It is stored with its metrics, its confusion matrix and its training curves, and it does nothing until a person switches it on. When they do, every site picks it up within about five minutes without anyone restarting anything. If a model's output does not match its own label list, it is refused and the previous one keeps running.

Next step

Ask us the hard version of the question

Bring the claim another vendor made and we will tell you whether our system does that, and if it does not, what it does instead.