There are no negatives in the cohort

Microsoft Research's post on CARE-X reports a second study in which measurement-dependent findings — aortic enlargement, hilar mass, pulmonary artery enlargement — were evaluated on an outpatient cohort of 122 positive cases with CT-confirmed ground truth. The tool-augmented variant reached 94.26 percent recall, 10.65 percentage points above the best perception-only baseline. The post sets the limit itself: these numbers are all recall, a model that flags everything achieves perfect recall and is useless in practice, and an extended study including CT-confirmed negative cohorts is under way.[1]

A cohort built only of positives can show that a method finds what is there; it cannot show how often the same method calls something enlarged when it is not. Without that side, the gain of 10.65 percentage points is compatible both with a genuine improvement and with an effectively lowered threshold, and the published figures do not separate the two. The competing reading deserves saying out loud: the tool path computes clinically defined quantities such as the cardiothoracic ratio instead of estimating them by eye, so the gain may well survive a negative cohort — but that would be an expectation about the next study rather than a result from this one.[1]

Which model do the numbers belong to?

The post describes the measurement work as a research experiment separate from CARE-X: it pairs Qwen3-VL-4B-Instruct with deterministic measurement tools in an inference-time loop and involves no task-specific training. CARE-X is a different system: a SigLIP2-so400M vision encoder, a Phi-4-mini-instruct language model of 3.8 billion parameters, auxiliary classification and grounding heads co-trained with the language objective, and a reinforcement-learning stage using DAPO. The F1 figures in the measurement table — cardiomegaly from 74.56 to 96.00, mediastinal widening from 72.63 to 97.47 — therefore describe the tool pipeline.[1]

That distinction changes what a reader can carry away. The numbers to weigh against a baseline are CARE-X's own: on 1,047 de-identified radiographs from Narayana Health, covering five conditions with prevalence between 2.6 percent and 5.2 percent, it reports the highest sensitivity on three of five, while on fracture it reads 0.62 sensitivity with 0.64 specificity against CheXOne's 0.41 and 0.90. Read as a pair, those two lines describe a different operating point rather than a uniformly better model, and which one is preferable depends on what a missed fracture and a false alarm each cost in the ward where it runs.[1]

Which number would change the view?

The next informative figure is already named in the post: the extended study with CT-confirmed negative cohorts. If it is published by the end of February 2027 and reports specificity or positive predictive value for the tool-augmented path at the same operating point that produced 94.26 percent recall, the screening claim about aortic dilation moves from plausible to testable. The signal to watch is narrow and concrete: a false-positive rate on CT-confirmed negatives. Without that number, the recall figure stays where it stands today.[1]

There is a smaller companion result the post reports from a related study accepted at the EACTS 2026 conference: for mild aortic dilation, measurement-driven reasoning detected 40 of 43 CT-confirmed cases, against 5 of 43 on the initial radiology reads. Read carefully, that comparison is about what a chest X-ray report is normally asked to do, since aortic enlargement is usually not the reason the film was taken; and here too what is measured is sensitivity on positives. Not a miracle, a method: the method here is written down clearly enough to let a reader name the missing half of the measurement, which is the most useful thing a research post can offer.[1]