Anthropic says its Claude AI models successfully designed protein binders for 14 of 15 targets in a laboratory-backed experiment, marking a significant expansion of the technology’s potential role in biological research. The company said Claude generated designs from scratch and, depending on the experimental setup, between 22% and 35% of individual designs successfully bound to their intended targets.

The experiment goes beyond a purely computational demonstration. Anthropic worked with Adaptyv Bio and Twist Bioscience to physically produce and test the proteins designed by Claude. Across 1,320 designs, 354 were confirmed to bind successfully, according to Anthropic. The company said the results compare with a typical 10% to 15% success rate for protein-design campaigns, although the findings come from Anthropic’s own research and represent an early demonstration rather than proof that Claude can independently develop medicines.

Claude Designed Binders For 14 Of 15 Targets

The experiment focused on de novo protein binder design, a process in which scientists attempt to create a new protein that attaches to a specific target protein.

Protein binders are important in biomedical research because many biological interventions depend on a molecule attaching to a particular protein and changing or blocking its activity. Designing a successful binder can therefore be an important early step in drug-discovery research, although a working binder is still far from becoming an approved medicine.

Anthropic gave Claude access to scientific literature, computing resources and specialist protein-design tools. The AI was then tasked with designing binders against 15 different protein targets.

It succeeded in producing measurable binding against 14 of those 15 targets.

Claude Protein Design Experiment At A Glance

MetricResult
Protein targets tested15
Targets with successful binders14
Targets without successful binder1
Total designs generated1,320
Confirmed successful binders354
Overall confirmed-binder ratio~26.8%
Reported hit-rate range22%–35%
Typical protein-design campaign10%–15%

The 354 successful binders divided by 1,320 total designs gives an overall ratio of approximately 26.8%. Anthropic’s reported 22% to 35% range reflects different experimental configurations rather than a single universal success rate.

Claude’s Hit Rate Was Above The Typical Range

One of the most important findings is the reported percentage of designs that actually bound to their intended targets.

Anthropic said Claude achieved binding rates between 22% and 35%, depending on the model and setup. The company compared those results with a typical 10% to 15% success rate in protein-design campaigns today.

Protein Binder Hit-Rate Comparison

The comparison suggests Claude’s reported hit rate was roughly 1.5 to 3.5 times the typical range, depending on which points are compared.

However, these figures should be interpreted carefully. The typical 10% to 15% benchmark is a broad industry comparison, while Claude’s results came from a specific experimental setup selected by Anthropic.

1,320 Designs Produced 354 Confirmed Binders

The experiment generated a large number of candidate proteins rather than relying on a handful of designs.

Claude produced 1,320 designs, of which 354 were confirmed as successful binders in the reported testing. That works out to approximately 26.8% across the combined set.

From AI Design To Laboratory Result

15 TARGETS
     ↓
Claude designs proteins
     ↓
1,320 candidate designs
     ↓
Physical production
     ↓
Laboratory testing
     ↓
354 confirmed binders
     ↓
14 of 15 targets achieved

This workflow is important because computational predictions alone can produce false positives. A protein can appear promising in a computer simulation but fail when physically produced and tested.

Anthropic’s experiment therefore included an experimental validation stage rather than stopping at AI-generated predictions.

Two Companies Tested Claude’s Designs

Anthropic said Adaptyv Bio and Twist Bioscience independently produced and tested the proteins generated during the experiment.

The involvement of external laboratory companies provides an additional layer of validation compared with an experiment evaluated solely by the AI developer.

However, the available results should still be viewed as an early research demonstration. Independent testing of the physical proteins establishes that at least some of the designs worked under laboratory conditions, but it does not establish that the system is ready for therapeutic development.

Validation Process

StageOrganisation / SystemRole
AI designClaudeGenerated candidate protein binders
Computational workflowSpecialist protein toolsSupported design and evaluation
Protein productionAdaptyv Bio / Twist BiosciencePhysically produced candidates
Experimental testingLaboratory partnersTested binding
Final resultAnthropic analysisReported successful binders

The distinction between designing a binder and developing a drug is crucial. A successful binder is an early research result, not evidence of a completed therapeutic.

Claude Used Existing Scientific Tools

The experiment did not involve Claude independently inventing every underlying scientific algorithm.

Anthropic’s research describes Claude as orchestrating specialist tools and models, including existing protein-design and protein-folding systems. This means the AI’s contribution was partly in determining how to use available scientific resources, iterate through designs and select candidates.

AI-Driven Scientific Workflow

Scientific Literature
        +
Protein-Design Models
        +
Protein-Folding Tools
        +
GPU Computing
        ↓
      CLAUDE
        ↓
Selects / orchestrates workflow
        ↓
Generates candidate binders
        ↓
Filters candidates
        ↓
Laboratory validation

That distinction matters when assessing the significance of the result. Claude is demonstrating the ability to coordinate complex scientific workflows rather than replacing every specialised computational tool involved in protein design.

Some Designs Were Stronger Than Previous Results

Anthropic said some of Claude’s strongest designs bound several times more tightly than the best previously published result for the relevant targets.

Binding strength is an important characteristic because a stronger interaction can potentially make a binder more effective at engaging its target.

However, stronger binding by itself does not determine whether a protein is suitable as a medicine.

A potential therapeutic candidate must satisfy many additional requirements, including stability, specificity, biological activity, safety and manufacturability.

From Binder To Medicine

AI-Designed Binder
       ↓
Laboratory Binding
       ↓
Binding Strength
       ↓
Specificity & Stability
       ↓
Biological Testing
       ↓
Preclinical Development
       ↓
Clinical Trials
       ↓
Regulatory Review
       ↓
Potential Medicine

Claude’s current achievement sits near the beginning of this chain rather than the end.

The Experiment Could Accelerate Early Drug Discovery

The potential commercial significance lies in the time required to generate and test protein designs.

Anthropic said designing protein binders for a target has historically required specialist work lasting weeks or months per target. The company’s experiment suggests AI agents could compress some of that computational work into a much shorter period.

If the approach scales, researchers could potentially evaluate many more candidate designs before committing significant laboratory resources.

Traditional Vs AI-Assisted Workflow

Traditional ApproachAI-Assisted Approach
Specialist researchers design candidatesAI generates and evaluates candidates
Weeks or months per target can be requiredAnthropic reports much faster design cycles
Manual tool selectionAI orchestrates multiple tools
Smaller number of candidatesPotentially hundreds of candidates
Laboratory testing follows designLaboratory validation remains essential

The biggest potential advantage is not necessarily that AI eliminates scientists. Instead, AI could allow researchers to explore a much larger design space before spending money and laboratory time on physical testing.

Claude Also Analysed Chemistry Lab Data

Anthropic’s research announcement included a second experiment involving analytical chemistry.

Claude Opus 5 was given raw nuclear magnetic resonance (NMR) and liquid chromatography-mass spectrometry (LC-MS) data from a contract laboratory. Anthropic said Claude produced completed analyses in 23 and 19 minutes and matched the laboratory’s results on hydrogen counts and purity.

For one purity measurement, Claude reportedly produced a result of 96.4%, compared with 96.33% from the laboratory’s own analysis.

Anthropic’s Chemistry Experiment

TaskClaude Result
NMR analysis23 minutes
LC-MS analysis19 minutes
Reported purity96.4%
Lab purity96.33%
InputRaw lab files + two-sentence prompt

The chemistry experiment reinforces Anthropic’s broader argument that AI can help automate routine but time-consuming parts of scientific research.

Anthropic Is Expanding Claude Beyond Coding

The protein experiment represents another step in Anthropic’s attempt to position Claude as a general-purpose scientific agent rather than simply a chatbot or coding assistant.

Claude’s ability to work with scientific literature, specialised tools and laboratory data could eventually make it useful across multiple stages of research.

The potential applications include protein design, chemistry, materials science and biological data analysis.

Claude’s Emerging Scientific Role

CHATBOT
   ↓
CODING ASSISTANT
   ↓
RESEARCH ASSISTANT
   ↓
SCIENTIFIC AGENT
   ↓
AI + COMPUTATIONAL TOOLS
   ↓
AI + LABORATORY WORKFLOWS

The final stage is particularly important because scientific discoveries ultimately require physical-world validation.

Why Laboratory Validation Matters

AI models can generate plausible scientific outputs without those outputs necessarily being physically correct.

Protein design provides a clear example. A computationally attractive protein sequence still has to be produced, folded correctly and tested to determine whether it actually binds to the intended target.

That makes laboratory testing a critical bottleneck.

Anthropic’s experiment addresses part of this bottleneck by demonstrating that a meaningful fraction of AI-generated designs can survive physical testing. But it does not eliminate the need for experiments.

Computational Prediction Vs Physical Validation

Computational StagePhysical Stage
Generate protein sequenceManufacture protein
Predict structureVerify physical structure
Estimate bindingTest actual binding
Rank candidatesConfirm experimental performance
Fast and scalableSlower and more expensive

The long-term value of AI in drug discovery could therefore come from reducing the number of failed candidates that reach expensive laboratory stages.

Biosafety Remains A Major Consideration

Anthropic has not made the protein-design capability generally available, citing dual-use biosafety concerns.

This is an important distinction from conventional AI features. A system capable of designing proteins could have beneficial applications in medicine and biotechnology while also creating potential risks if used to design harmful biological agents.

As a result, Anthropic is treating access to advanced protein-design capabilities differently from ordinary consumer-facing Claude features.

Opportunity Vs Risk

Potential BenefitPotential Risk
Faster drug discoveryDual-use biological applications
More protein candidatesMisuse of design capabilities
Lower computational expertise barrierBiosafety concerns
Faster scientific analysisIncorrect AI-generated results
Better research productivityNeed for laboratory verification

The challenge for AI companies will be to provide legitimate researchers with useful capabilities while limiting misuse.

The Bigger Picture

Anthropic’s protein-design experiment represents a significant shift in the role AI could play in scientific research. Claude did not simply answer a biology question; it coordinated specialised computational tools to generate physical protein candidates that were subsequently tested in laboratories. The reported 14-of-15 target result and 354 confirmed binders show that AI-generated designs can produce measurable biological results under experimental conditions.

At the same time, the achievement should not be confused with autonomous drug discovery. Protein binders are an early research component, and the path from a successful binder to a safe and effective medicine involves many additional stages. Anthropic’s decision not to make the protein-design capability generally available also underscores the dual-use risks associated with increasingly capable biological AI.

Looking Ahead

The next test for Anthropic will be whether the approach can reproduce these results across larger and more diverse sets of targets, with independent researchers and broader laboratory validation. A 14-of-15 result is promising, but a larger body of experiments will be needed to determine how consistently Claude can outperform conventional workflows and whether the reported hit-rate advantage survives outside Anthropic’s selected testing environment.

If the results continue to hold, AI agents could become an important layer between computational biology and physical experimentation, allowing researchers to generate, rank and refine far more candidates before committing laboratory resources. The immediate opportunity is therefore not to replace scientists, but to accelerate the earliest and most computationally intensive parts of scientific discovery while keeping experimental validation and human oversight firmly in the loop.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.