← Back to Home

Nearly 14,000 People Test AI Symptom Assessment: Asking Follow-Up Questions More Closely Resembles Clinical Reasoning Than Answering Directly

Google recruited nearly 14,000 people through Fitbit to test five conversational agents. Systems that actively asked follow-up questions outperformed a general chat mode in differential diagnosis, but the key findings were based on only a small number of participants who self-reported diagnoses received during medical visits and are not yet sufficient to replace physicians.

By SURL BioNews

As generative AI begins answering health questions about headaches, coughs, chest tightness, and other symptoms, the real test is not merely how much it “knows,” but whether it can conduct an assessment by drawing out the course of illness, associated symptoms, and risk signals from fragmented descriptions. A large field study by a Google research team found that purpose-designed conversational agents that actively collect medical histories can indeed produce differential diagnoses that more closely match subsequent clinical judgments than general chat models guided by users themselves. However, these results remain far from proving that AI can safely conduct consultations independently.

The study recruited 13,917 consenting participants through the Fitbit app and randomly assigned them to five SymptomAI agents based on Gemini Flash 2.0. Four agents led interviews using fixed-question, flexible-question, or dynamic-reasoning approaches, while another group used a user-led mode resembling a general chatbot. At the end of each conversation, the system listed several possible causes and recommended next steps. All content was provided solely for research analysis and did not constitute a formal medical diagnosis.

The sample that could actually be used to verify the results was far smaller than the total number of participants. Two weeks later, only 1,228 people reported diagnoses provided by healthcare professionals, of which 517 cases entered more than 250 hours of review by clinical experts. Board-certified physicians reviewed the same conversation transcripts under blinded conditions and generated their own differential diagnoses, after which clinical reviewers compared the AI and physician lists. The latest preprint reports that the odds of SymptomAI’s differential diagnosis matching a subsequently self-reported diagnosis were approximately 2.56 times those of the comparator physicians, a statistically significant difference. This figure differs slightly from the 2.47-fold result reported in an earlier version.

The clearest message from the study may not be that “AI outperformed physicians,” but rather the importance of active questioning. The four agent-led strategies, which asked more follow-up questions about medical history, performed similarly and all substantially outperformed the basic user-led chat mode. In other words, the value of a medical conversational system depends not only on the model that generates the final answer, but also on whether it knows what to ask, when to follow up, and how to organize incomplete everyday language into clinically interpretable information.

The team also matched the conversation results with more than 500,000 days of wearable-device data covering nearly 400 conditions. Among people classified by the AI as having acute respiratory infections, group-level changes in cardiovascular, respiratory, skin-temperature, sleep, and other measures appeared in the period leading up to the day of the symptom conversation. This temporal alignment may serve as evidence of physiological association, but it cannot conversely prove that the AI’s diagnosis was correct for an individual participant. Wearable signals themselves also mostly lack disease specificity.

The study’s limitations center on the “ground truth” and the method of comparison. Subsequent diagnoses were self-reported by participants and may have been affected by memory, willingness to seek medical care, and selection bias. Of nearly 14,000 participants, only 517 cases were included in the primary expert review. The independent physicians could only read transcripts shaped by the AI’s questions; they could not ask their own follow-up questions, examine patients, or review medical records. The comparison therefore cannot be equated with AI being tested against a complete clinical consultation. An independent analysis also noted that the top-five diagnostic hit rate was approximately 70%, leaving a non-negligible margin for error.

SymptomAI is currently an experimental research prototype, not an approved clinical diagnostic product. If such systems enter care workflows in the future, they will require prospective clinical trials and validation in external populations, as well as answers to questions involving incorrect triage, missed emergency symptoms, allocation of responsibility, and health-data governance. The study’s more credible contribution is that it moves the evaluation of medical AI beyond curated case questions and toward the ambiguous, fragmented conversations people actually have. It demonstrates a better way to ask questions, but has not yet delivered a diagnostic tool that can safely be entrusted with the task.

References

  1. Google Research
  2. arXiv
  3. Evidence Triage