Biomedical Policy · us
From Pediatric Cancer Data to Drug Safety: US Paves the Way for Bio Genesis With More Than $1.2 Billion
This national initiative aims to connect fragmented clinical, genetic, and molecular data with supercomputing, shortening the time it takes for drugs and diagnostics to reach patients; the real tests will be data quality, privacy, and clinical validation.
Whether artificial intelligence can accelerate medical progress depends not only on how large the models are, but also on whether hospital data can be securely shared, whether predictions hold across populations, and whether regulators accept the evidence. The US National Institutes of Health (NIH) has launched Bio Genesis in an effort to connect federal health data, advanced computing, and automated laboratory facilities into a nationwide biomedical research system.
Bio Genesis is part of the interagency Genesis Mission. The White House announced a federal commitment of more than $5 billion for the overall mission, but that amount also covers areas including energy, manufacturing, infrastructure, and national security, and cannot be regarded entirely as biomedical funding. According to plans released by NIH, more than $1.2 billion already committed for fiscal year 2026 and planned for 2027 has been aligned with six biomedical challenges, with additional funding opportunities still expected to be introduced.
The six areas include predicting complex biological systems, advancing biomanufacturing within the United States, strengthening biosecurity, accelerating pediatric cancer research, shortening drug translation, and identifying the root causes of chronic disease. The work will integrate molecular, genomic, phenotypic, clinical, and real-world data to train models to find new uses for existing drugs. Veterans’ health data may also be combined with electronic health records and genetic information to identify disease risks earlier.
Pediatric cancer is one of the most concrete testing grounds. The plan is to establish a national learning system spanning hundreds of rare pediatric cancer subtypes, allowing imaging, pathology, genetic, and treatment outcome data from different hospitals to be analyzed together. ARPA-H is separately investing $50 million to expand a privacy-preserving, AI-ready pediatric data network to more than 200 hospitals and care centers. Its pilot work coordinated and mapped more than 90% of unstructured clinical data, electronic health records, and genetic data between two major hospitals, but this cannot yet be equated with evidence that the models have improved diagnosis or survival outcomes.
In drug development, the CATALYST program will invest up to $125 million over four and a half years to develop models that simulate human biology, with the aim of predicting the safety and efficacy of drug candidates before human trials. Another program, APECx, with funding of up to $204 million, combines AI protein modeling, high-throughput experimentation, and clinical evaluation to design vaccine antigens that can cover multiple viruses. These programs outline a path from computational prediction to experimental validation, but at this stage they remain focused primarily on research objectives and infrastructure development, with no comparative results sufficient to demonstrate clinical benefit.
NIH has proposed cutting in half, within a decade, the time it takes for scientific discoveries to reach patients. To deliver on this goal, the models must still undergo external validation across institutions and populations, while addressing biases in medical records, missing data, genetic privacy, and authorization governance. If AI predictions are to support human trials or regulatory decisions, it must also be demonstrated that the results are reproducible, traceable, and superior to current methods. Bio Genesis is therefore not merely an investment in computing power, but also an institutional experiment in whether health data can be used in a trustworthy manner.