Anthropic says its Claude AI models successfully designed protein binders for 14 of 15 targets in a laboratory-backed experiment, marking a significant expansion of the technology’s potential role in biological research. The company said Claude generated designs from scratch and, depending on the experimental setup, between 22% and 35% of individual designs successfully bound to their intended targets.
The experiment goes beyond a purely computational demonstration. Anthropic worked with Adaptyv Bio and Twist Bioscience to physically produce and test the proteins designed by Claude. Across 1,320 designs, 354 were confirmed to bind successfully, according to Anthropic. The company said the results compare with a typical 10% to 15% success rate for protein-design campaigns, although the findings come from Anthropic’s own research and represent an early demonstration rather than proof that Claude can independently develop medicines.
Claude Designed Binders For 14 Of 15 Targets
The experiment focused on de novo protein binder design, a process in which scientists attempt to create a new protein that attaches to a specific target protein.
Protein binders are important in biomedical research because many biological interventions depend on a molecule attaching to a particular protein and changing or blocking its activity. Designing a successful binder can therefore be an important early step in drug-discovery research, although a working binder is still far from becoming an approved medicine.
Anthropic gave Claude access to scientific literature, computing resources and specialist protein-design tools. The AI was then tasked with designing binders against 15 different protein targets.
It succeeded in producing measurable binding against 14 of those 15 targets.
Claude Protein Design Experiment At A Glance
| Metric | Result |
|---|---|
| Protein targets tested | 15 |
| Targets with successful binders | 14 |
| Targets without successful binder | 1 |
| Total designs generated | 1,320 |
| Confirmed successful binders | 354 |
| Overall confirmed-binder ratio | ~26.8% |
| Reported hit-rate range | 22%–35% |
| Typical protein-design campaign | 10%–15% |
The 354 successful binders divided by 1,320 total designs gives an overall ratio of approximately 26.8%. Anthropic’s reported 22% to 35% range reflects different experimental configurations rather than a single universal success rate.
Claude’s Hit Rate Was Above The Typical Range
One of the most important findings is the reported percentage of designs that actually bound to their intended targets.
Anthropic said Claude achieved binding rates between 22% and 35%, depending on the model and setup. The company compared those results with a typical 10% to 15% success rate in protein-design campaigns today.
Protein Binder Hit-Rate Comparison
The comparison suggests Claude’s reported hit rate was roughly 1.5 to 3.5 times the typical range, depending on which points are compared.
However, these figures should be interpreted carefully. The typical 10% to 15% benchmark is a broad industry comparison, while Claude’s results came from a specific experimental setup selected by Anthropic.
1,320 Designs Produced 354 Confirmed Binders
The experiment generated a large number of candidate proteins rather than relying on a handful of designs.
Claude produced 1,320 designs, of which 354 were confirmed as successful binders in the reported testing. That works out to approximately 26.8% across the combined set.
From AI Design To Laboratory Result
15 TARGETS
↓
Claude designs proteins
↓
1,320 candidate designs
↓
Physical production
↓
Laboratory testing
↓
354 confirmed binders
↓
14 of 15 targets achieved
This workflow is important because computational predictions alone can produce false positives. A protein can appear promising in a computer simulation but fail when physically produced and tested.
Anthropic’s experiment therefore included an experimental validation stage rather than stopping at AI-generated predictions.
Two Companies Tested Claude’s Designs
Anthropic said Adaptyv Bio and Twist Bioscience independently produced and tested the proteins generated during the experiment.
The involvement of external laboratory companies provides an additional layer of validation compared with an experiment evaluated solely by the AI developer.
However, the available results should still be viewed as an early research demonstration. Independent testing of the physical proteins establishes that at least some of the designs worked under laboratory conditions, but it does not establish that the system is ready for therapeutic development.
Validation Process
| Stage | Organisation / System | Role |
|---|---|---|
| AI design | Claude | Generated candidate protein binders |
| Computational workflow | Specialist protein tools | Supported design and evaluation |
| Protein production | Adaptyv Bio / Twist Bioscience | Physically produced candidates |
| Experimental testing | Laboratory partners | Tested binding |
| Final result | Anthropic analysis | Reported successful binders |
The distinction between designing a binder and developing a drug is crucial. A successful binder is an early research result, not evidence of a completed therapeutic.
Claude Used Existing Scientific Tools
The experiment did not involve Claude independently inventing every underlying scientific algorithm.
Anthropic’s research describes Claude as orchestrating specialist tools and models, including existing protein-design and protein-folding systems. This means the AI’s contribution was partly in determining how to use available scientific resources, iterate through designs and select candidates.
AI-Driven Scientific Workflow
Scientific Literature
+
Protein-Design Models
+
Protein-Folding Tools
+
GPU Computing
↓
CLAUDE
↓
Selects / orchestrates workflow
↓
Generates candidate binders
↓
Filters candidates
↓
Laboratory validation
That distinction matters when assessing the significance of the result. Claude is demonstrating the ability to coordinate complex scientific workflows rather than replacing every specialised computational tool involved in protein design.
Some Designs Were Stronger Than Previous Results
Anthropic said some of Claude’s strongest designs bound several times more tightly than the best previously published result for the relevant targets.
Binding strength is an important characteristic because a stronger interaction can potentially make a binder more effective at engaging its target.
However, stronger binding by itself does not determine whether a protein is suitable as a medicine.
A potential therapeutic candidate must satisfy many additional requirements, including stability, specificity, biological activity, safety and manufacturability.
From Binder To Medicine
AI-Designed Binder
↓
Laboratory Binding
↓
Binding Strength
↓
Specificity & Stability
↓
Biological Testing
↓
Preclinical Development
↓
Clinical Trials
↓
Regulatory Review
↓
Potential Medicine
Claude’s current achievement sits near the beginning of this chain rather than the end.
The Experiment Could Accelerate Early Drug Discovery
The potential commercial significance lies in the time required to generate and test protein designs.
Anthropic said designing protein binders for a target has historically required specialist work lasting weeks or months per target. The company’s experiment suggests AI agents could compress some of that computational work into a much shorter period.
If the approach scales, researchers could potentially evaluate many more candidate designs before committing significant laboratory resources.
Traditional Vs AI-Assisted Workflow
| Traditional Approach | AI-Assisted Approach |
|---|---|
| Specialist researchers design candidates | AI generates and evaluates candidates |
| Weeks or months per target can be required | Anthropic reports much faster design cycles |
| Manual tool selection | AI orchestrates multiple tools |
| Smaller number of candidates | Potentially hundreds of candidates |
| Laboratory testing follows design | Laboratory validation remains essential |
The biggest potential advantage is not necessarily that AI eliminates scientists. Instead, AI could allow researchers to explore a much larger design space before spending money and laboratory time on physical testing.
Claude Also Analysed Chemistry Lab Data
Anthropic’s research announcement included a second experiment involving analytical chemistry.
Claude Opus 5 was given raw nuclear magnetic resonance (NMR) and liquid chromatography-mass spectrometry (LC-MS) data from a contract laboratory. Anthropic said Claude produced completed analyses in 23 and 19 minutes and matched the laboratory’s results on hydrogen counts and purity.
For one purity measurement, Claude reportedly produced a result of 96.4%, compared with 96.33% from the laboratory’s own analysis.
Anthropic’s Chemistry Experiment
| Task | Claude Result |
|---|---|
| NMR analysis | 23 minutes |
| LC-MS analysis | 19 minutes |
| Reported purity | 96.4% |
| Lab purity | 96.33% |
| Input | Raw lab files + two-sentence prompt |
The chemistry experiment reinforces Anthropic’s broader argument that AI can help automate routine but time-consuming parts of scientific research.
Anthropic Is Expanding Claude Beyond Coding
The protein experiment represents another step in Anthropic’s attempt to position Claude as a general-purpose scientific agent rather than simply a chatbot or coding assistant.
Claude’s ability to work with scientific literature, specialised tools and laboratory data could eventually make it useful across multiple stages of research.
The potential applications include protein design, chemistry, materials science and biological data analysis.
Claude’s Emerging Scientific Role
CHATBOT
↓
CODING ASSISTANT
↓
RESEARCH ASSISTANT
↓
SCIENTIFIC AGENT
↓
AI + COMPUTATIONAL TOOLS
↓
AI + LABORATORY WORKFLOWS
The final stage is particularly important because scientific discoveries ultimately require physical-world validation.
Why Laboratory Validation Matters
AI models can generate plausible scientific outputs without those outputs necessarily being physically correct.
Protein design provides a clear example. A computationally attractive protein sequence still has to be produced, folded correctly and tested to determine whether it actually binds to the intended target.
That makes laboratory testing a critical bottleneck.
Anthropic’s experiment addresses part of this bottleneck by demonstrating that a meaningful fraction of AI-generated designs can survive physical testing. But it does not eliminate the need for experiments.
Computational Prediction Vs Physical Validation
| Computational Stage | Physical Stage |
|---|---|
| Generate protein sequence | Manufacture protein |
| Predict structure | Verify physical structure |
| Estimate binding | Test actual binding |
| Rank candidates | Confirm experimental performance |
| Fast and scalable | Slower and more expensive |
The long-term value of AI in drug discovery could therefore come from reducing the number of failed candidates that reach expensive laboratory stages.
Biosafety Remains A Major Consideration
Anthropic has not made the protein-design capability generally available, citing dual-use biosafety concerns.
This is an important distinction from conventional AI features. A system capable of designing proteins could have beneficial applications in medicine and biotechnology while also creating potential risks if used to design harmful biological agents.
As a result, Anthropic is treating access to advanced protein-design capabilities differently from ordinary consumer-facing Claude features.
Opportunity Vs Risk
| Potential Benefit | Potential Risk |
|---|---|
| Faster drug discovery | Dual-use biological applications |
| More protein candidates | Misuse of design capabilities |
| Lower computational expertise barrier | Biosafety concerns |
| Faster scientific analysis | Incorrect AI-generated results |
| Better research productivity | Need for laboratory verification |
The challenge for AI companies will be to provide legitimate researchers with useful capabilities while limiting misuse.
The Bigger Picture
Anthropic’s protein-design experiment represents a significant shift in the role AI could play in scientific research. Claude did not simply answer a biology question; it coordinated specialised computational tools to generate physical protein candidates that were subsequently tested in laboratories. The reported 14-of-15 target result and 354 confirmed binders show that AI-generated designs can produce measurable biological results under experimental conditions.
At the same time, the achievement should not be confused with autonomous drug discovery. Protein binders are an early research component, and the path from a successful binder to a safe and effective medicine involves many additional stages. Anthropic’s decision not to make the protein-design capability generally available also underscores the dual-use risks associated with increasingly capable biological AI.
Looking Ahead
The next test for Anthropic will be whether the approach can reproduce these results across larger and more diverse sets of targets, with independent researchers and broader laboratory validation. A 14-of-15 result is promising, but a larger body of experiments will be needed to determine how consistently Claude can outperform conventional workflows and whether the reported hit-rate advantage survives outside Anthropic’s selected testing environment.
If the results continue to hold, AI agents could become an important layer between computational biology and physical experimentation, allowing researchers to generate, rank and refine far more candidates before committing laboratory resources. The immediate opportunity is therefore not to replace scientists, but to accelerate the earliest and most computationally intensive parts of scientific discovery while keeping experimental validation and human oversight firmly in the loop.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.

