Biomedical Artificial Intelligence · global
AI Can Select Molecules, but Has Yet to Prove It Can Create Better Drugs: Pharmaceutical Evaluation Must Move Beyond Leaderboards
A perspective article in Nature Reviews Drug Discovery notes that improvements in model accuracy have yet to translate into clear clinical benefits; the next generation of evaluations should ask whether AI truly improves R&D decision-making.
Artificial intelligence can now predict protein structures, screen compounds, design new molecules, and is often expected to shorten the lengthy and costly drug development process. Yet the question that truly matters to patients is not whether a model can score a few more points on a test set, but whether it enables research teams to eliminate unsuccessful candidates earlier, select safer and more effective drugs, and accelerate their path to the clinic. A recent perspective article in Nature Reviews Drug Discovery argues that this crucial evidence remains quite limited.
The article reviews more than a decade of developments in AI drug discovery. It notes that although model development, applications, and benchmarking have increased substantially, there is still insufficient clinically relevant evidence to demonstrate that patients are consequently receiving safer or more effective drugs sooner. This does not mean AI has no value. Rather, current evidence largely remains focused on predictive performance, computational speed, or individual stages of R&D, with multiple hurdles still separating these measures from a new drug’s ultimate success.
One challenge stems from the nature of life sciences data. Results for compound activity, toxicity, and pharmacokinetics often vary with experimental conditions, cell lines, doses, and data sources. Even if a model is adept at fitting existing data, it may not be able to handle new chemical spaces or clinical contexts. If training tasks are not defined precisely enough, several models with similar leaderboard performance may reach entirely different judgments when deployed in real-world R&D workflows.
The authors also distinguish between “technology push,” driven by technical capabilities, and “science pull,” guided by biomedical questions. The former tends to begin with a new algorithm and then look for data to which it can be applied; the latter first defines the decisions an R&D team must make and the costs of incorrect judgments. The article argues that for AI tools to have a practical impact, problem formulation, data generation, and validation methods should all be designed backward from specific use cases.
This also means evaluation standards must change. Rather than comparing only model accuracy on fixed datasets, researchers should assess whether AI increases the hit rate for candidate compounds, reduces unnecessary experiments, improves the subsequent ranking of candidates, or identifies safety and translational risks earlier. Such prospective validation embedded in real-world workflows is slower and more expensive, but it comes closer to answering the questions that pharmaceutical companies, regulators, and patients actually need answered.
The article is an interdisciplinary perspective review, not a new study of a single AI-developed drug, clinical trial, or patient dataset, and it does not present a standardized figure capable of quantifying benefits across the industry. It therefore cannot prove the success or failure of individual programs. A more precise conclusion is that current evaluations remain insufficient to support the industry’s most ambitious clinical promises. The next test for AI in drug development is not merely whether it can design seemingly novel molecules, but whether it can leave behind traceable, comparable decision records demonstrating that it has genuinely changed the path drugs take to reach patients.