Health Technology · global
Clinical AI Is Moving Into the Exam Room, and Adoption Is Outpacing Validation Systems
From real-time prompts providing the basis for diagnosis and treatment during medical-record conversations to literature-based answers for physicians’ queries, specialized clinical AI systems are competing to become the gateway to decision-making. Yet usage rates, internal safety scores, and actual improvements in patient outcomes remain three different things.
As physicians speak with patients, AI may no longer merely organize medical records in the background. Based on the conversation and patient data, it may also suggest medications, differential diagnoses, or treatment directions on the spot. This kind of clinical decision support is evolving from a search tool into part of the care workflow, making one question increasingly urgent: Are products entering the exam room faster than evidence-generation and oversight systems can keep pace?
*Nature Medicine* reported that specialized systems including Abridge, Atropos Health, and OpenEvidence are already competing in the US market, while physicians are also using general-purpose chatbots. The competition is not only about who can answer questions, but who can occupy the interface physicians use most often. Once recommendations appear directly alongside the medical record or consultation screen, every search eliminated may translate into greater product stickiness and clinical influence.
Abridge exemplifies this shift. Originally known for its ambient clinical documentation tool, the company later added decision-support features that present contextualized evidence within its existing interface, based on consultation conversations, patient records, and information sources including UpToDate. Physicians can select questions suggested by the system or enter their own. Answers are not automatically added to the medical record, but they are positioned as close as possible to the point of decision-making.
To demonstrate that its product is ready for use in healthcare institutions, Abridge says it has used more than 1,000 physician-developed scoring criteria and incorporated adversarial testing involving harmful prompts, requests for unauthorized actions, discriminatory scenarios, and attempts to bypass safety restrictions. The company also says it scored 99.5% in a clinical safety evaluation covering more than 1,000 cases and is expanding use progressively through internal testing, small-scale trials with partner physicians, and deployment in designated departments. However, these figures come from the company itself and cannot yet replace independent research. Nor do they answer whether AI reduces misdiagnosis, improves treatment outcomes, or decreases patient harm.
OpenEvidence, meanwhile, entered the market through medical-literature question answering. Data released by the company from the NOHARM study cover 101 US board-certified physicians, 45 large language models, 4 clinical AI systems, 12,747 expert annotations, and 4,249 potential clinical actions. In the study group allowed to choose tools freely, physicians used OpenEvidence for 22.3% of their answers, compared with a combined 19.8% for other external AI chatbots. But this benchmark study has not yet undergone peer review, and the data were released by the company through a press release. Frequency of use may reflect preference, but it cannot directly prove better quality of care.
The two companies also represent different paths to trust. Abridge emphasizes multilayered testing and post-deployment feedback, while OpenEvidence has introduced EvidenceGrade, which attempts to label the quality of the evidence underlying each answer according to the GRADE framework. The former focuses on whether answers pass safety checks; the latter emphasizes making the strength of the evidence visible to physicians. But even when a system cites reliable literature, it may still mismatch the patient’s circumstances, while strong offline scores may not reproduce the time pressure, missing data, and automation bias found in busy exam rooms.
The key to the next stage, therefore, is not another impressive accuracy figure, but the establishment of comparable clinical evidence that can be continuously tracked: which patients, specialties, and workflows use the system; how errors are detected and reported; whether revalidation is needed after model updates; and whether patient outcomes ultimately change. Clinical AI has already begun participating in the selection and presentation of evidence. If evaluation methods remain confined to scales defined separately by each vendor, the market may decide what physicians see first, while science follows behind to ask whether those recommendations are truly reliable.