← Back to Home

Medical AI Takes On Three Follow-Up Visits, Performing on Par With Primary Care Physicians in Disease Management

Across 100 multidisciplinary simulations involving multiple consecutive visits, Google’s AMIE produced more precise, guideline-aligned care plans; however, outperforming physicians in text-based scenarios does not yet prove that it can safely care for real patients.

By SURL BioNews

The challenge of chronic disease care often lies not in getting the diagnosis right once, but in repeatedly adjusting tests, medications, and follow-up arrangements over several months. A study published in *Nature* shows that AMIE, a medical AI system developed by Google, demonstrated overall management-reasoning capabilities on par with primary care physicians when simulating this type of longitudinal decision-making, while receiving higher scores for plan precision and adherence to clinical guidelines.

The study used a randomized, blinded virtual clinical skills assessment involving 21 primary care physicians, 21 participants acting as patients, and 100 text-based consultation scenarios spanning five specialties, each comprising three visits. The cases were designed using UK NICE guidelines and BMJ Best Practice, with a supporting database containing a total of 627 documents. This allowed the assessment to test not only a single consultation, but also subsequent testing, treatment adjustments, and responses to disease progression.

AMIE is not a single chatbot. According to Google, the system combines a real-time conversational agent with a management-reasoning agent: the former gathers information from patients and explains plans, while the latter consults clinical guidelines and drug formularies to organize possible testing, treatment, and follow-up options. Specialist physicians scored the responses without knowing the identity of the respondent, and concluded that AMIE’s overall disease-management reasoning met the non-inferiority criterion relative to primary care physicians.

The differences primarily emerged in how plans balanced competing options. AMIE scored higher for the precision of its treatment and testing recommendations and for alignment with guidelines, suggesting that it was better able to narrow its recommendations to what each scenario required instead of compiling a large number of possible options. The study also used the RxQA medication-reasoning benchmark to test difficult medication questions; when the system could access external drug information, it also outperformed the physicians who were tested.

However, the comparison measured reasoning quality in controlled text-based scenarios, not patient health outcomes. Simulated patients cannot fully reproduce the ambiguous descriptions, interactions among comorbidities, testing delays, medication difficulties, and sudden deterioration encountered in real clinical settings. Nor are three virtual visits sufficient to validate safety, accountability, or the physician–patient relationship in long-term care. The study’s authors are from Google Research and Google DeepMind, and the system still requires independent testing that more closely reflects real-world clinical practice.

To become a clinical decision-support tool, AMIE must also demonstrate that it remains reliable across different healthcare systems, populations, and levels of medical-record data quality, while establishing mechanisms for physician review, error interception, privacy protection, and continuous monitoring. In particular, “on par with physicians” is a statistical conclusion based on specific scoring metrics and a prespecified threshold; it cannot be interpreted as meaning that the system can replace physicians or prescribe treatment autonomously.

Google said it has launched a nationwide real-world virtual-care study to explore the system’s use in clinical settings; this will be a more critical test. At this stage, the study demonstrates that medical AI can maintain more structured disease-management reasoning across multiple visits, while both the *Nature* paper and an independent commentary state explicitly that AMIE is not yet ready for use in actual clinical care.

References

  1. Nature
  2. Google
  3. arXiv
  4. PubMed
  5. Nature Medicine