← Back to Home

AI Can Design Exceptionally Potent Antibodies, but May Not Be Able to Pick the Good Ones: 511 Blinded Tests Reveal Its True Capabilities

Antibodies submitted by 29 teams underwent standardized experimental testing, with some designs achieving picomolar affinity; however, leading methods often failed when applied to a different task, and simple statistics were even more reliable than most AI approaches.

By SURL BioNews

Antibody generation models can propose thousands of impressive amino acid sequences, but the question that truly matters for drug development is how many of them can bind tightly to their targets in the laboratory while remaining stable, showing low aggregation, and being manufacturable. A prospective blinded benchmark published in *Nature Biotechnology* examined this gap for the first time through a large-scale, standardized experimental process: AI can already deliver exceptionally potent antibodies under specific conditions, but it is not yet an antibody discovery machine suitable for every stage.

The AIntibody challenge received 511 designs or predictions submitted by 29 academic, pharmaceutical, biotechnology, and technology organizations. Participants tackled three types of tasks: improving antibody affinity based on existing screening data, identifying the best sequence from candidate groups with similar heavy-chain CDR3s, and designing new CDR combinations absent from the original database. The inaugural test used the receptor-binding domain of the SARS-CoV-2 spike protein as the antigen, and all candidate sequences were subsequently produced as full-length IgGs rather than remaining only at the computational scoring stage.

The most striking results emerged in “affinity maturation”: when models were given higher-quality sequencing data resembling that used in actual development workflows, they could effectively optimize existing antibodies. Aureka’s AuraBind method won this task; overall, several teams produced antibodies with dissociation constants below 100 picomolar that also met developability thresholds. This indicates that AI may already be able to narrow the range of candidates requiring synthesis and testing, helping researchers identify variants worth advancing more quickly.

But in the task of “selecting the best one from a candidate group,” the advantage quickly disappeared. Except for one method from a University of Washington team, every AI strategy performed worse than random sampling at identifying high-affinity antibodies; among all AI submissions, only 9.8% to 13.8% outperformed the control sequence in each group, compared with 39% for random selection. Moreover, when the same winning method was applied to another sequence group or task, it generally failed to maintain its performance.

Designs that moved beyond existing sequence libraries showed even more polarized results. Xencor’s method produced the competition winner, with affinity reaching 2.9 picomolar, but the research team noted that this sequence had a major flaw that could prevent it from becoming a clinical candidate. In this task, about 30.4% of submissions did not bind the antigen at all, while another 16% bound but failed the developability assessment. A small number of top designs coexisted with many failed sequences, showing that peak records do not represent a platform’s ability to deliver consistently.

This blinded benchmark measured more than binding strength. Researchers used surface plasmon resonance and KinExA to confirm affinity and ranking, and also examined hydrophobicity, self-interaction, polyreactivity, melting temperature, and aggregation temperature. These metrics reflect whether an antibody might encounter problems during manufacturing, storage, or use in humans, and explain why “binding most tightly” does not mean “being the most suitable drug candidate.” Winning teams including Aureka and Xencor also made related code or model materials publicly available, providing an inspectable foundation for future comparisons of methods.

The study’s boundaries were equally clear: all tasks focused on a single antigen for which extensive structural and sequence data already exist, and the data provided were even richer than what is typically available at the same stage of a conventional discovery workflow; the organizing consortium also participated in publishing the results, and blinding relied mainly on organizational procedures rather than fully independent information isolation. These results therefore more closely resemble the current technology’s upper limit under favorable conditions. AI antibody design has found a useful role, but becoming a reliable general-purpose tool will require repeated blinded testing across antigens and data environments, as well as substantially larger and more diverse experimental datasets on affinity and developability.

References

  1. Nature Biotechnology, Published online: 2026-08-19; | doi:10.1038/s41587-026-03238-6
  2. Nature Biotechnology
  3. AIntibody.org (administered by Specifica, an IQVIA business)
  4. Aureka Biotechnologies
  5. Xencor