Biomedical AI · us
Rerunning Old Trials With AI: EmulatRx Tests Clinical Research Designs Using Medical Records
Five AI agents with clearly defined roles worked together—from trial protocols and medical-record fields to causal analysis—to reconstruct 20 historical studies. The results demonstrate the potential to support study design while also highlighting the limits that prevent real-world data from replacing randomized trials.
The costliest mistakes in clinical trials are often embedded before the first patient is enrolled: eligibility criteria may be too narrow, endpoints may be difficult to obtain from medical records, or key confounding factors may not be handled properly. EmulatRx, developed by a team at Weill Cornell Medicine, attempts to have multiple AI agents first “rerun” historical trials using existing electronic health records, thereby assessing whether a study design is feasible. The peer-reviewed study of the system has been published in *Nature Communications*, but it remains positioned as a research tool rather than a source of evidence that can replace prospective clinical trials.
The researchers evaluated 20 trial emulations. Ten acute-disease studies used the critical-care database MIMIC-IV, while 10 chronic-disease studies used data from the INSIGHT Network. The latter integrates medical records from five healthcare systems in New York, enabling the team to test whether the system could reconstruct existing treatment comparisons from real-world data across different disease timescales and clinical settings.
EmulatRx does not rely on a single chatbot to perform the entire analysis. Instead, it assigns five role-based nodes: the Supervisor coordinates the workflow, the Trialist develops the trial framework, the Clinician handles clinical meaning, the Informatician maps eligibility criteria and variables to electronic health records, and the Statistician estimates causal effects and interprets the results. This division of labor seeks to transform work that ordinarily requires repeated coordination in clinical research into a traceable and revisable computational workflow.
Across multiple emulations, the system reproduced treatment effects previously reported in historical studies and identified differences in certain patient subgroups. Technical materials also presented a corticosteroid trial in intensive-care patients with sepsis as a proof of concept: after the system adjusted eligibility criteria, handled missing variables, and revised covariates, the resulting average treatment effect and risk ratio were directionally consistent with the original randomized trial. Such results suggest that it may help researchers identify data gaps, estimate feasible populations, or compare the potential effects of different protocol configurations before formal enrollment begins.
However, obtaining similar answers from historical data does not mean the system can predict a new trial. The way electronic health records are documented, treatment selection, and patient composition can all introduce bias; unmeasured confounding factors also do not disappear as the number of agents increases. The authors therefore noted that validation is currently limited to MIMIC-IV and INSIGHT. Applicability across databases, healthcare systems, and additional disease areas has not yet been established, and causal estimates also require expert review.
Version 1.0.0 of the EmulatRx software has been archived through Zenodo, providing a citable snapshot of the implementation. Weill Cornell’s technical materials indicate that a U.S. patent application has been filed, and future uses may extend to post-marketing effectiveness studies, drug safety monitoring, indication expansion, and drug repurposing. What will ultimately determine its value is not merely whether it can reproduce old trials, but whether it can consistently generate transparent and verifiable analyses across different hospitals, varying levels of data quality, and new questions for which no standard answer yet exists.