← Back to Home

AI Can Understand Proteins, but Still Cannot Pinpoint Where Antibodies Bind: EpiBench Reveals a Gap in Epitope Reasoning

A new benchmark encompassing 1,609 samples backed by experimental and structural evidence shows that general-purpose large language models can capture broad epitope signals but struggle to reliably connect specific antibodies, amino acid positions, and immune escape.

By SURL BioNews

Whether an antibody can become a good drug depends not only on whether it binds to an antigen, but also on exactly where it binds. This region, known as an “epitope,” may influence blocking efficacy, therapeutic specificity, and immune escape following mutation. However, a new benchmark shows that even though large language models can now answer complex biomedical questions, their ability to identify this critical binding site directly from protein sequences remains quite limited.

The research team’s EpiBench includes 1,609 samples, with labels derived from structural contacts in antibody–antigen complexes, curated B-cell functional experiments, and escape effects measured through deep mutational scanning. The benchmark deliberately removes antibody names, antigen names, Protein Data Bank identifiers, and paper information, retaining only antigen and antibody sequences, residue positions, or mutation data to reduce the opportunity for models to answer by memorizing the questions.

The questions are divided into five categories spanning the antibody development workflow: identifying regions in an antigen sequence that may be targeted, determining the epitope corresponding to a specific antibody, grouping antibodies whose binding sites overlap, identifying epitopes with functional effects, and predicting whether a single-point mutation will render an antibody ineffective. These tasks correspond respectively to candidate-region selection, clarification of mechanisms of action, antibody combination design, and assessment of drug-resistance risk, rather than merely general biological knowledge question answering.

After nine general-purpose models underwent a standardized zero-shot evaluation, none led across the board. Some models could place their predicted ranges roughly near possible epitopes, but their performance remained close to random when distinguishing epitope residues from non-epitope residues one by one. Questions requiring the binding site or escape effect to be determined in conjunction with a specific antibody also exposed deficiencies in sequence alignment and structural reasoning. In tasks that allowed direct comparison, specialized epitope models still outperformed general-purpose large language models overall.

Long sequences were particularly likely to cause models to lose track of positions. In the region-exploration task, in which no specific antibody was provided, the nine models’ average top-50 candidate-residue recall score fell from 81.3 for antigens shorter than 200 amino acids to 12.8 for antigens longer than 800 amino acids. Requiring models to provide an explicit reasoning process also did not consistently improve the results. Incorrect answers often relied on seemingly plausible sequence motifs, general physicochemical properties, or antibody-sequence similarity without truly addressing the antibody–epitope relationship being asked about.

These findings are better viewed as a diagnostic tool than as a final verdict on AI’s capabilities in antibody design. EpiBench uses a closed, sequence-only, automatically scored setting and does not directly simulate real-world development workflows in which structural modeling, experimental feedback, and human interpretation all play a role. Some known epitopes may also be incomplete. The research is currently available only as a preprint and has not yet undergone peer review, but it draws a clear boundary: general-purpose models can provide broad clues, yet they still lack a reliable foundation in residue localization, antibody specificity, and biological mechanisms needed to independently support antibody-candidate selection and escape-risk assessment.

References

  1. arXiv