← Back to Home

AI Reading of Cancer Papers Nears Expert Level, but Spotting Research Flaws Remains the Hardest Task

Four large language models underwent a structured critical appraisal test involving 24 studies on breast cancer and viruses. The overall agreement distribution of the GPT-5 series was close to that of domain experts, but methodological assessment, contradictions within full texts, and hallucinations by smaller models delineated the limits of automated evidence review.

By SURL BioNews

As the medical literature grows rapidly, the truly scarce resource is no longer just reading time, but the expert attention needed to determine whether evidence supports a conclusion. A peer-reviewed study also published on arXiv shows that large language models can, in some respects, approach domain experts in highly structured critical appraisal of cancer literature. However, once the task delves into study design and internal contradictions, the machines’ weaknesses become more apparent.

Using the potential association between an “MMTV-like virus” and breast cancer as a case study, the researchers selected 24 papers and designed 77 evidence-extraction and critical-appraisal questions covering multiple-choice questions, rating scales, multiple-response questions, and free-text answers. The models tested included Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, and GPT-5 Nano. The research team then compared the degree of answer agreement among experts and between experts and models.

The results showed that the distributions of agreement scores achieved by GPT-5 and GPT-5 Nano were statistically indistinguishable from the distribution of agreement among experts. This does not mean that the models matched experts on every question, and it certainly cannot be interpreted as evidence that they can independently adjudicate medical evidence. It means that, within this appraisal framework—limited to a specific topic and structured in advance—the overall response patterns of the two models fell within the range formed by differences among experts.

The models also displayed different biases. The two Gemini models were more permissive when applying criteria for microbial causation of cancer, potentially making it easier for associations supported by insufficient evidence to clear the threshold. Fabricated content was rare overall but occurred more often in smaller models. Even when overall agreement appears favorable, a small number of errors could still distort the direction of a systematic review if they concern pivotal studies or core criteria.

The most persistent difficulties centered on methodological assessment and identifying contradictory statements within the same full text. These two tasks require an understanding of experimental design, whether controls are appropriate, whether inferences overreach, and how to connect information dispersed across the methods, results, and discussion sections. These are precisely the aspects of biomedical evidence review that are hardest to reduce to data extraction.

The study supports using large language models for initial screening, field extraction, or as a second reviewer to help manage a volume of literature that is difficult for human reviewers to track comprehensively. However, it has not demonstrated that models can replace experts. The test covered only a single microbe–cancer question and 24 papers, and it did not establish whether models can maintain their performance across other diseases, literature of varying quality, or real-world review workflows that are continuously updated. The researchers raised the possibility of a multi-model strategy, but broader external validation, error tracking, and human review mechanisms are still needed before deployment in biomedical decision-making.

References

  1. arXiv
  2. Frontiers in Cellular and Infection Microbiology