Medical Artificial Intelligence · eu
Medical AI enters multidisciplinary consultations: A 100-case study finds it useful, but recommendation agreement remains limited
For complex inflammatory diseases, OpenEvidence can provide information physicians find useful; however, full agreement with consultation conclusions was approximately 40–50%, and answers to the same question also changed over time.
When inflammatory diseases involve different organs, treatment choices often require physicians from multiple specialties to weigh the options together. Can artificial intelligence that synthesizes medical literature in real time help clarify these difficult questions? A study published in Scientific Reports on October 2 brought the medical large language model OpenEvidence into real-world consultation workflows and reached a cautious conclusion: the information was useful, but agreement between recommendations remained limited.
This single-center, prospective observational study included 100 consecutively enrolled cases of complex inflammatory diseases. Physicians first discussed the cases independently and recorded preliminary recommendations, then reviewed the AI responses and confirmed their final recommendations after further discussion. Twenty-one physicians from seven specialties completed a total of 759 case ratings; these ratings reflected experts’ experiences using the tool, rather than treatment outcomes for 759 patients.
Physicians generally endorsed the AI’s usefulness. On a five-point scale, “providing clinical information relevant to the case” received a mean score of 4.06, while recommending continued use in the future scored 3.95. This suggests that it can be incorporated into discussions, but finding information helpful and accepting its diagnostic and treatment recommendations remain two different judgments.
Two physicians separately assessed 97 evaluable cases. The proportions in which each judged the AI response to be fully consistent with the final consultation recommendations were 40.2% and 50.5%, respectively. Their ratings also differed. These figures should therefore be understood as the degree of agreement with expert conclusions, rather than directly interpreted as the accuracy of the AI’s diagnosis and treatment recommendations. In particular, the final recommendations were confirmed after physicians had reviewed the AI responses, so they were not a fully independent benchmark for comparison.
The research team then adjusted the prompting approach, asking the model to first check whether the information was sufficient before making recommendations. Agreement showed a trend toward improvement, but this did not reach statistical significance. Because the two rounds of testing were several months apart and the platform continued to be updated during that period, the study could not clearly separate the effects of improved prompting from changes to the system.
Differences over time also warrant attention: when the researchers entered the same questions again, both the responses and the cited literature changed. For clinical use, a response that performs well on one occasion cannot be guaranteed to be reproduced later; how much the model knows about a patient’s condition also depends on the information physicians provide beforehand.
This study supports the feasibility of AI as a consultation support tool, but has not demonstrated that it can improve patient outcomes. The study did not assess clinical harm or systematically examine “hallucinations”—generated content that appears plausible but is actually incorrect. In addition, because the cases came from a single center, the generalizability of the results remains limited. How to continually validate updated systems and enable experts to identify omissions in responses are questions that still need to be answered as AI moves toward routine medical use.