AI and Protein Engineering · global
Claude Autonomously Designs Protein Binders, with Wet-Lab Validation for 14 of 15 Targets
Two independent laboratories confirmed that 354 of 1,320 AI-designed candidates could bind their targets, a hit rate above the reference range for typical protein design programs; but being able to “grab” a protein is still far from becoming a safe and effective drug.
One of the most time-consuming steps in designing protein therapeutics is identifying, from a large pool of candidate sequences, the molecules that can genuinely grab a disease target. Anthropic has released the results of an experiment in which Claude autonomously designed protein binders: among 15 targets with interpretable measurement results, the model found wet-lab-confirmed binders for 14, indicating that general-purpose AI may be starting to participate in early-stage R&D workflows that previously required multiple rounds of computation and human judgment.
The technical report shows that Claude Opus 4.8 and Mythos Preview conducted multiple autonomous design tasks lasting 24 to 48 hours for 16 protein targets, using existing protein design tools to generate, evaluate, and iteratively improve sequences. Researchers submitted 1,320 designs for testing and confirmed that 354 could bind their targets, for an overall hit rate of about 27%; depending on the experimental setup, the success rate for an individual design ranged from 22% to 35%. Anthropic compared this with the 10% to 15% commonly seen in previous protein design programs, but the targets, screening thresholds, and assay methods were not entirely consistent, so the difference cannot be interpreted directly as a standardized performance ranking.
The significance of these results lies not only in computer simulations. Adaptyv Bio and Twist Bioscience separately synthesized proteins from the sequences delivered by the model and measured binding performance using different experimental formats, reducing potential bias from a single laboratory or assay platform. Looking only at the designs Claude ranked first for each target in each task, 49% were experimentally confirmed, suggesting that the model’s ranking capability may be as important as its generation capability.
Individual cases also showed more concrete signals of potency. For RBX1, which is involved in protein ubiquitination, 28 of the 90 sequences proposed by Claude successfully bound the target; the strongest measured affinity was 3.9 nanomolar, while a competition-winning design tested on the same experimental plate measured 45 nanomolar. However, this comparison involved a single target under specific testing conditions and cannot be used to infer that the model has the same advantage across all protein or drug-development tasks.
Anthropic also released on Hugging Face a dataset containing summaries of 1,440 designs across 16 targets, including protein sequences, generation tools, optimization rounds, model rankings, predicted structures, and binding and affinity data obtained by the two experimental organizations. The dataset contains more records than the 1,320 principal experimental designs described above, reflecting that the public release covers a more complete set of task records; this release may help other researchers reproduce the analysis and examine whether the model favors particular protein structures or scoring methods.
The real barriers to drug development begin to emerge one by one only after binding has been confirmed. Binders must still be shown to alter disease-related functions and to have sufficient selectivity, stability, manufacturability, and in vivo exposure; immune responses, toxicity, efficacy in animals, and human trials also remain unanswered by this study. The results are therefore better regarded as validation of the efficiency of early-stage candidate generation and ranking, rather than evidence that AI can already complete new-drug R&D independently.