DAMO RADAR: The Radiology AI That Beat Most Radiologists
6 min read
Photo by National Cancer Institute on Unsplash
DAMO RADAR Just Became the Radiology AI to Watch
Most radiology AI stories follow a familiar script: a lab claims impressive numbers on a narrow task, the paper gets some attention, and then it quietly disappears into the pile of research that never touches a real hospital. Last week's release from Alibaba's DAMO Academy is a bit different, and it's worth slowing down for.
On September 18, 2026, DAMO Academy and Zhejiang University School of Medicine published a study in Science describing a model called DAMO RADAR. It reads contrast-enhanced CT scans of the abdomen, covering 18 organs, and flags nearly 150 distinct conditions, from fatty liver disease and acute appendicitis to liver, pancreatic, gastric, and colorectal cancers. The researchers tested it against almost 40,000 real-world exams and ran a reader study pitting it against 26 practicing radiologists. On average, the model outperformed 23 of them. Alibaba then open-sourced the whole thing: weights, code, and training framework, all released alongside the paper.
That combination, a peer-reviewed study with a large real-world evaluation, plus a genuinely open release, is rare enough that it deserves a closer look at what the model actually does, what the numbers mean, and where the obvious limits are.
What DAMO RADAR Actually Does
DAMO RADAR is a vision-language model, the same broad architecture family behind consumer multimodal chatbots, but trained specifically on CT scans paired with the clinical reports radiologists write about them. Instead of producing a single yes-or-no answer, it generates structured findings across a long list of possible conditions, similar to how a radiologist dictates a report organ by organ.
The headline number from the paper is a mean area under the curve (AUC) of 0.913 across 146 clinical findings, evaluated on that pool of roughly 40,000 exams. AUC measures how well a model separates true cases from false ones on a scale where 1.0 is perfect and 0.5 is a coin flip, so 0.913 puts it solidly in the range clinicians consider strong for a screening tool, though not flawless.
The more attention-grabbing figure is the head-to-head comparison: in a multi-hospital reader study, DAMO RADAR beat 23 of the 26 radiologists it was tested against, on average, across the full set of conditions. Alibaba also reported that using the model alongside radiologists in a workflow setting reduced missed diagnoses by about 10% and cut diagnosis time by more than 30%. Those workflow numbers matter more for near-term adoption than the raw AUC score does, since most hospitals won't hand imaging over to software unsupervised. What they're actually buying is a second reader that catches things a tired human might miss on the fortieth scan of a shift.
Why the Open-Source Release Matters as Much as the Score
A lot of medical AI research stays locked behind a paywall or a commercial product, which makes independent verification slow and expensive. DAMO Academy released DAMO RADAR's weights, code, and technical framework in full alongside the Science publication. That means outside researchers, and other hospital systems, can run the exact model on their own data rather than taking the benchmark numbers on faith.
This matters because AI models trained on one hospital's scanners and patient population sometimes stumble when they hit a different scanner brand, a different demographic mix, or imaging protocols the training data didn't include. An open release invites exactly the kind of scrutiny that catches those gaps early, instead of after a model is already deployed somewhere.
The researchers describe DAMO RADAR as "the world's first expert-level generalist medical imaging model," and say the underlying method could extend to imaging types beyond abdominal CT. That's a notable claim, and it's the kind of thing the open release should let other labs actually test rather than just repeat.
The Limits Nobody Should Skip Over
None of this means DAMO RADAR is ready to read scans on its own in a clinic tomorrow. A few caveats came with the release itself, not from outside skeptics:
Third-party clinical validation is still pending. A single institution's reader study, even a well-designed one at 40,000 exams, is not the same as independent labs reproducing the result on their own patient populations.
Regulatory clearance is a separate, slower process. Publishing in Science and open-sourcing weights doesn't clear the FDA or equivalent bodies elsewhere, and clinical deployment in most countries requires that approval regardless of how strong the paper's numbers look.
Generalization outside the training distribution is unproven. Performance on scanners, imaging protocols, or patient populations that differ from what the model trained on needs its own testing, and abdominal CT is a fairly specific slice of what radiology departments actually see day to day.
As the research team itself put it, open weights don't equal clinical safety. An assistant model that flags 146 conditions is not an autonomous doctor, and treating it as one, or racing to deploy it that way before validation catches up, would be the wrong lesson to take from a genuinely strong result.
How This Fits the Broader Radiology AI Landscape
Radiology has been one of the more active corners of applied computer vision for years now, with tools from companies like Aidoc and platforms built around workflow triage and report generation already in clinical use at many hospitals. What's shifted lately is the scope of what a single model attempts. Earlier tools tended to specialize narrowly, catching pulmonary embolisms or flagging a single type of lesion. A generalist model that handles 146 findings across 18 organs from one CT protocol is a meaningfully bigger ask, and DAMO RADAR's numbers suggest that scope isn't automatically a tradeoff against accuracy.
For developers and researchers working in medical imaging, the open weights are the practical story here. Instead of waiting for a commercial vendor to license access, teams can pull DAMO RADAR down, test it against their own datasets, and see where it holds up and where it doesn't. That's a different kind of contribution to the field than another closed benchmark result, and it's likely to shape how quickly (and how carefully) generalist imaging models move from research papers into actual hospital workflows.
Key Takeaways
DAMO RADAR is a vision-language model from Alibaba's DAMO Academy and Zhejiang University School of Medicine that reads abdominal CT scans for nearly 150 conditions, published in Science on September 18, 2026. In a reader study, it outperformed 23 of 26 radiologists on average, with a mean AUC of 0.913 across 146 findings on almost 40,000 real exams. Alibaba open-sourced the model's weights, code, and framework alongside the paper, which is unusual for research at this scale and lets independent teams verify the results directly. The model still needs broader third-party validation and regulatory clearance before it can be used clinically, and its own creators are clear that it's an assistant, not a replacement for a radiologist's judgment.
FAQ
Is DAMO RADAR available to use right now? The weights and code are open-sourced for research use, but it hasn't cleared regulatory approval for clinical deployment, so hospitals can't yet use it to make diagnostic decisions on patients.
What does it mean that the model "beat" radiologists? In a controlled reader study, its average performance across 146 conditions exceeded that of 23 out of 26 participating radiologists. It doesn't mean the model is right every time or that the three radiologists it didn't outperform were doing anything wrong; reader studies compare averages across many cases, not single decisions.
Does this mean AI will replace radiologists? Nothing in the paper suggests that, and the research team explicitly frames it as an assistant tool. The workflow benefit Alibaba reported, a 10% drop in missed diagnoses and faster turnaround, points toward augmenting radiologists rather than removing them from the loop.
Why does the open-source release matter so much? Because it lets other researchers and hospitals test the model against their own data instead of trusting a single institution's benchmark, which is exactly the kind of scrutiny generalist medical AI models need before anyone should trust them in practice.
- radiology AI
- medical imaging AI
- computer vision
- vision-language models
- healthcare AI
- Alibaba DAMO Academy