← Back to Home

Does Clinical AI Lose to General-Purpose Models? Two Evaluations Reach Opposite Conclusions, Bringing Trust in Medical Chatbots to the Fore

One independent comparison found that clinical AI built specifically for physicians does not necessarily outperform frontier general-purpose models; another physician evaluation placed OpenEvidence far ahead. The divergence concerns more than rankings, exposing how data sources, scoring methods, and conflicts of interest shape the definition of “reliable.”

By SURL BioNews

Physicians are increasingly turning to AI for diagnostic clues, treatment options, and the latest literature. But is a system labeled “clinical-specific” truly more reliable than a general-purpose chatbot? Two recent evaluations produced nearly opposite answers, extending the debate beyond model capabilities to a more fundamental question: Can existing tests demonstrate that a tool is ready for real-world medical decision-making?

The main study, published in *Nature Medicine*, compared OpenEvidence and UpToDate Expert AI with GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6. The evaluation included 500 MedQA questions, 500 HealthBench items, and 100 real-world clinical queries; for the last portion, 12 clinicians who did not know the source of the answers completed a total of 1,800 annotations. The results showed that the three general-purpose models formed a higher-performing group, while the two clinical tools lagged overall.

The gap does not simply mean that “general-purpose models are safer.” The study found no significant differences among the models in harmful content or hallucinations. More apparent issues emerged in usability: UpToDate Expert AI, for example, had a refusal rate of about 19%, while OpenEvidence received the lowest clarity score. In a busy clinical setting, whether an answer is complete, understandable, and directly supports judgment may be as important as correctly answering knowledge questions.

However, the evaluation also has important blind spots. The authors acknowledged that public medical question banks may have been included in model training data, causing benchmark contamination. Parts of HealthBench were also scored with the involvement of frontier models, potentially favoring answers that resemble their style of expression. Testing the clinical tools through browser interfaces may not reflect their normal workflows. The study also did not compare citation quality, answer traceability, or response speed—precisely the factors that give physicians important reasons to choose specialized tools.

The controversy intensified because another arXiv study reached the opposite conclusion. That evaluation used 620 questions actually submitted to OpenEvidence by physicians and 187 HealthBench items, and invited 149 practicing physicians matched by specialty to assess them. In blinded pairwise comparisons, physicians preferred OpenEvidence over GPT-5.5, Gemini 3.1 Pro, or Claude Opus 4.8 in accuracy, clinical usefulness, source quality, verifiability, and completeness; its advantage over each of the other models was about 25 to 39 percentage points.

This reversal likewise needs to be interpreted in the context of the study design. The questions in the second study came from existing OpenEvidence users, and OpenEvidence also participated in planning data collection and recruiting physicians; this may have made the question distribution more closely aligned with use cases in which the product excels. On the other hand, its assessment of source quality and clinical utility addressed key dimensions not covered by the first study. The two sets of results therefore may not be reducible to one being right and the other wrong. Instead, they may show that different question sets, model versions, user interfaces, and scoring criteria produce different winners.

What is truly missing is a shared validation framework more closely tied to clinical consequences: independent institutions using undisclosed, time-sensitive cases and questions; preregistering analytical methods; checking whether citations support claims; and tracking whether errors would change diagnostic or treatment decisions. High scores on medical exam questions are only a starting point. Until these conditions are in place, neither specialized nor general-purpose models should earn medical trust based on a single leaderboard.

References

  1. STAT
  2. Nature Medicine
  3. arXiv