← Back to Home

Generative AI Still Trails in Radiological Diagnosis: Meta-Analysis of 48 Studies Finds Experts 13 Percentage Points More Accurate

How far is generative AI from being a reliable clinical assistant when making diagnoses from images and medical histories? A new study finds that response format and input information affect performance, while average results across studies are insufficient to establish whether it is safe to use in clinical practice.

By SURL BioNews

An abnormality on an image often needs to be considered alongside medical history and other clues before it can be translated into a diagnosis. Generative artificial intelligence has been tested on this ability, but producing a fluent answer does not mean its judgment is reliable. A systematic review and meta-analysis published on October 3 in the *Japanese Journal of Radiology* found that generative AI remained less accurate than expert physicians in the included studies of radiological diagnosis.

The analysis pooled 48 studies identified through searches conducted through March 2025 and was registered in advance with PROSPERO; the abstract of the same paper indexed in PubMed also confirmed the study scope and main findings. The evaluations covered text, images, and combined text and image inputs, testing whether models could provide the correct diagnosis. The research therefore measured more than the ability to directly interpret images; it also assessed disease inference based on textual descriptions.

Response format affected performance: pooled accuracy was 42.9% when models were required to write their own diagnoses, compared with 58.1% when answer options were provided. A statistical model comparing results across studies estimated that expert physicians were approximately 13.0 percentage points more accurate than AI, with a 95% confidence interval of 0.9 to 25.2 percentage points and a P value of 0.038. This is an absolute difference in accuracy and cannot be interpreted as AI consistently trailing by the same margin for every type of examination or disease.

The basis for comparison also has limits. Only 12 of the 48 studies included physician comparators; the difference above came from a model estimate across studies, rather than all AI systems and physicians being tested on the same cases under the same conditions. Differences in case mix and difficulty may affect the results, and these test scores cannot be directly converted into error rates for routine hospital diagnosis.

Another finding that can easily be misinterpreted is that text-only inputs performed better than image-only or combined text and image inputs. The authors cautioned that text sometimes already contained image findings written and organized by people, effectively providing important clues in advance; task difficulty may also have differed between groups. This association supports further research into text-assisted applications, but it does not prove that removing images makes diagnosis more accurate.

The quality of the evidence also limits how far the conclusions can be extended: only 1 study was rated as having a low risk of bias, while most others had an unclear or high risk. The test materials were also almost entirely in English, leaving insufficient evidence of applicability to Chinese-language medical settings. With the literature search ending in March 2025, this analysis reflects the evidence on models and evaluation methods at that time and cannot represent the capabilities of all subsequent versions.

For clinical practice, the next step is to establish which parts of the workflow AI can reliably assist with. The authors suggested that uses such as differential diagnosis suggestions for physician review could continue to be evaluated, but deployment should be preceded by standardized testing, adequate sample sizes, and transparent reporting of methods. Determining whether these tools can improve care will also require testing physicians’ performance and safety when using AI under conditions that closely reflect actual practice.

References

  1. Japanese Journal of Radiology
  2. PubMed