
Table of Contents
We recently ran a set of mock radiology examinations, modelled on the UK Royal College of Radiologists FRCR 2B Short Cases. Against a pass threshold of 73.2, our foundation model Harrison.Rad 1.5 scored a median of 86.5; GPT-5.4, the best general-purpose model, scored 44. Other models all scored below 38. This gap illustrates a core argument about general-purpose AI radiology benchmarks: intelligence is multidimensional, and applying the right benchmarks and tests is critical to picking the best model for the task.
Why the Headlines Around General-Purpose AI Radiology Benchmarks Can Mislead
You may have seen a recent Nature Medicine paper reporting that general-purpose language models outperform specialized clinical AI tools across several medical benchmarks. The convenient reading is that generalist models beat specialized medical AI, but this is too broad, particularly in radiology.
What the Nature Medicine Paper Actually Tested
Look at what was tested: licensing-style questions, clinician preference alignment, and de-identified clinical queries. These are text in, text out. On those tasks, frontier models did better, which is meaningful. But it doesn’t tell you how these models read a scan.
Why the Signal Is in the Image, Not the Text
Radiology is different because the information that matters is not conveyed in text, but in images. The clinical notes may reveal an extensive smoking history, but the cancerous lung nodule itself can only be found in the pixels. A guideline may tell you what to do with a finding, but it does not contain the visual signal you need to detect it.
Why Frontier Models Struggle With Image Interpretation
Frontier models can reason their way into a treatment plan, but reasoning cannot surface a finding it never saw in the image. To confidently detect these abnormalities requires training upon hospital data, PACS data, radiologist reports and addendums, with the variation of real scanners and protocols. A frontier model can be extraordinarily capable and still have seen very little of the data that matter most for reading a scan.
Making the Test Match the Task in General-Purpose AI Radiology Benchmarks
The deeper issue is one every health system already applies to its own quality programs: does the benchmark actually measure what you care about? A text-based medical benchmark tests knowledge and reasoning in language. A radiology short case is closer to the real work of a radiologist: look at the images, find what is relevant, ignore what is not.
Why This Distinction Matters for Health Systems
Neither test is a perfect stand-in for clinical practice, but they measure very different things, and only one of them tells you about image interpretation. Be wary of impressive scores on the wrong test. A model can ace a medical-text leaderboard and still be the wrong tool for chest X-rays.
Holding Harrison.ai’s Own Models to This Same Standard
The 86.5 result is an internal one: the examination set has not been released, and Harrison.Rad 1.5 has not yet faced external scrutiny, so it should be treated as a useful stress test, not the final word. That scrutiny is exactly what external researchers are welcome to partner on.
Independent Evidence on the Earlier Harrison.Rad 1 Model
In a 2026 American Journal of Roentgenology study, radiologists at Stanford and Mass General Brigham compared four AI systems on 212 chest radiographs. Harrison.Rad 1, the earlier version of the model, produced the reports radiologists most often accepted, up to 75.5%, versus 16% to 57% for the others, and rated highest for quality and preference. Separately, at the 2025 American College of Radiology Annual Meeting, a challenge run with Mass General Brigham had 113 radiologists blind-rate Harrison.Rad 1 reports acceptable 65.4% of the time, against 79.6% for radiologist-written reports.
What This Means for How Health Systems Should Evaluate Radiology AI
The Nature Medicine paper may be a useful result for medical question-answering and sitting examinations, but the practice of radiology is a different task, and if health systems are choosing AI to read scans, they should insist it be evaluated as one. Harrison.Rad 1 and Harrison.Rad 1.5 are research-only foundation models and are not medical devices regulatory-cleared for clinical use; Harrison.ai is seeking regulatory clearance for medical devices powered by these models in various markets.
Why Vendor Selection Criteria Should Reflect This Distinction
Given this argument that text-based benchmarks and image-interpretation exams measure fundamentally different capabilities, health system leaders evaluating radiology AI vendors may want to specifically request image-based performance data, such as blind radiologist acceptance rates on real chest radiographs, rather than relying solely on general medical-knowledge benchmark scores when comparing competing tools.
What This General-Purpose AI Radiology Benchmarks Debate Means Going Forward
As health systems continue navigating a growing field of both general-purpose and specialized clinical AI tools, this piece underscores a broader lesson applicable well beyond radiology: benchmark selection itself shapes which tools appear to “win,” making it essential for health system leaders to match evaluation criteria to the actual clinical task at hand rather than relying on headline-grabbing aggregate scores. Given that Harrison.Rad 1.5’s cited 86.5 score remains an unreleased internal result pending external validation, health systems should treat vendor-reported benchmarks generally with appropriate scrutiny until independently verified, regardless of which company is making the claim.
What to Watch Going Forward
As Harrison.ai pursues regulatory clearance for its Harrison.Rad models and invites external researchers to validate its internal exam results, industry observers will likely watch whether independent studies confirm the performance gap described here between specialized radiology AI and general-purpose frontier models. Given the broader debate this piece raises about benchmark selection in general-purpose AI radiology benchmarks evaluation, health systems purchasing radiology AI tools may increasingly demand image-specific validation data, rather than general medical-knowledge scores, as a standard part of their vendor selection process going forward.
For more healthcare industry updates, insights and news, visit DistilINFO. Click here to subscribe to stay informed.
