AI-Designed Protein Binders

Anthropic × Adaptyv Bio × Twist Bioscience | August 2026

AI-Designed Protein Binders
AI Designed Protein Binders

The Anthropic–Adaptyv Bio results are an important inflection point for AI-enabled biotechnology, but not because Claude has suddenly learned to invent drugs from first principles.

The deeper development is that a frontier general-purpose AI agent can now act as a computational protein engineer and research orchestrator: research a biological target, select an epitope, choose among specialist protein-generation models, install and run those models, allocate compute, generate and filter candidates, rank them, and deliver sequences for physical testing—with very little human intervention.

Anthropic's Claude Opus 4.8 and Mythos Preview generated experimentally confirmed binders against 14 of 15 evaluable targets. Across 1,320 experimentally measured designs, 354 bound their intended targets, an overall hit rate of approximately 27%. Multi-target campaigns produced hit rates of 22.6–26.7%; Mythos Preview reached 35.1% when working on targets separately. Two external laboratories—Adaptyv Bio and Twist Bioscience—physically synthesized and tested the designs.

The investment thesis is therefore not simply:

“AI will design drugs.”

It is:

Drug discovery is beginning to acquire a software-like execution layer in which autonomous agents operate models, compute infrastructure and increasingly automated laboratories in a continuous design–build–test–learn loop.

That transition could materially change where value accumulates across the biotech stack.


What actually happened

Claude was given an approximately 30,000-token expert protocol rather than individual protein designs. It then operated publicly available specialist protein-design and structure-prediction models including tools derived from RFdiffusion, PXDesign, Genie, BoltzGen, ProteinMPNN-family methods, ESMFold and Protenix. Humans selected the overall benchmark set, provided infrastructure and approved access requests, but did not make individual design decisions during the campaigns.

The campaign used substantial computational resources: up to 12,500 NVIDIA H100 GPU-hours for the multi-target runs. The independent analysis by The Decoder estimates the stated budgets at approximately $50,000 for a multi-target campaign and $10,000 per single-target campaign. Importantly, the improvement to 35.1% in single-target mode came with roughly 2.8× more compute per target, meaning improved performance cannot yet be separated cleanly from increased computational expenditure.

The strongest example was RBX1. In a previous Adaptyv competition, only 9 of 245 submitted de novo proteins bound RBX1. Claude generated 28 binders from 90 designs, and its strongest binder measured approximately 3.9 nM KD, versus approximately 45 nM for the competition winner when the two were evaluated on the same assay plate.

Claude also generated cross-species binders: 130 of 233 tested binders recognized the mouse orthologue of their target. Opus 4.8 generated binders against therapeutically important TNFα, including designs reactive across human, cynomolgus monkey and mouse proteins.

That is legitimate experimental evidence rather than an in-silico benchmark.


The bull case

1. The experiment crossed the computational-to-physical boundary

AI-biotech announcements routinely stop at predicted structures, docking scores or computational benchmarks.

This experiment did not.

The molecules were synthesized and experimentally tested by external organizations, and the submitted designs were tested essentially as delivered. That greatly increases the credibility of the result.

2. Agentic orchestration may be as important as another protein foundation model

Claude did not invent RFdiffusion, BoltzGen or ProteinMPNN.

It learned which tools to use, when to use them and how to combine them.

That represents an important transition:

Model → workflow → autonomous research agent.

Much as coding agents now combine compilers, repositories, terminals, tests and cloud infrastructure, scientific agents can combine literature, biological databases, specialist ML models, HPC clusters and laboratory services.

The scientific literature was already pointing in this direction before Anthropic's announcement. A May 2026 Communications Biology commentary argued that routine de novo protein-binder design may be undergoing an “AlphaFold moment,” while a major Nature review concluded that several foundational protein-design problems are approaching the point where the question shifts from how to design toward what should be designed.

3. Parallelization fundamentally changes research economics

A human protein engineer can intensively operate only a limited number of campaigns.

An agent can potentially operate:

100 targets
→ 1,000 targets
→ tens of thousands of design hypotheses

in parallel.

Even if the AI is only comparable to a good scientist on each individual campaign, enormous parallelism can transform total R&D throughput.

The scarce resource potentially moves from scientific labor toward:

experimental capacity + high-quality biological data + capital + downstream development expertise.

4. Adaptyv may represent the more interesting structural change

Adaptyv describes itself as a cloud laboratory for protein designers. Its API allows software—and explicitly AI agents—to query targets, create experiments, obtain cost estimates, track experiments and retrieve structured wet-lab results programmatically.

That makes the following loop possible:

AI agent → protein design → API → automated wet lab → experimental measurement → structured data → AI agent → next design

This is substantially more consequential than simply improving a protein-generation model.

It creates the beginnings of a biological equivalent of CI/CD in software.

Adaptyv also operates ProteinBase, which aggregates standardized positive and negative experimental protein-design data. Negative experimental data is particularly valuable because much of it never appears in scientific publications.

That combination—automation + standardized assays + API access + accumulated experimental data—can create a genuine data flywheel.

Adaptyv's last publicly disclosed major financing was an $8 million seed round, led by ACE Ventures, after reporting more than 10,000 proteins tested during 2025 and customers spanning pharma, AI laboratories and biotech startups.

From an investor perspective, this is a remarkably strategically positioned asset for a company at that stage.


The bear case — and why the headlines overstate the result

“27% versus 10–15%” should not be interpreted as “Claude is twice as good as humans.”

This is perhaps the most important qualification.

Anthropic's 10–15% comparison is derived from protein-design campaigns represented in ProteinBase rather than a contemporaneous, randomized human-expert control running exactly the same targets, models and compute budget.

And success rates across modern protein-design systems vary enormously.

A 2026 review of experimentally validated AI protein-design systems reports individual campaigns ranging from zero hits to very high success rates. Chai-2, for example, reported an overall 68% rate across its own validation set, while individual AlphaProteo targets ranged from zero to very high hit rates. Direct comparison is difficult because targets, assays, candidate-selection rules and experimental conditions differ.

Therefore the breakthrough is not necessarily state-of-the-art molecular generation.

The breakthrough is autonomous orchestration of a state-of-the-art toolchain.

That distinction is crucial for investors.


Binder ≠ therapeutic

Anthropic explicitly acknowledges this.

The experiments primarily demonstrate binding.

They do not yet demonstrate, across these candidates:

functional antagonism or agonism,
cellular efficacy,
target selectivity across the proteome,
pharmacokinetics,
tissue penetration,
immunogenicity,
toxicity,
formulation stability,
manufacturability,
in-vivo therapeutic efficacy, or
clinical benefit.

This limitation has precedent. Stanford's 2026 AI Index notes that Adaptyv's previous Nipah competition produced 99 experimentally confirmed binders—including very high-affinity molecules—but none neutralized the targeted protein.

That single observation captures the gap between:

binding

and

medicine.


The system still fails unpredictably

Against maltose-binding protein, none of 90 designs were confirmed as binders.

Claude also struggled with BBF-14.

More concerningly, computational confidence metrics did not clearly distinguish these failed campaigns from successful ones.

This means the AI cannot yet reliably recognize:

“I am working in a region where my design stack does not work.”

That is one of the central challenges for autonomous science.


Human expertise has not disappeared; it has been compiled into the system

The protocol itself was approximately 16,000 words / 30,000 tokens and encoded extensive expert knowledge.

Humans also selected the benchmark targets and constructed the experimental environment.

So the correct analogy is not “scientist eliminated.”

It is closer to:

expert scientific knowledge → encoded workflow → reusable agentic infrastructure.

That is still enormously valuable—but economically different from replacing scientific expertise entirely.


The key investor insight: protein generation itself may commoditize

Because Claude used predominantly open-source protein-design models, this experiment contains a somewhat uncomfortable message for companies whose moat is simply:

“We have an AI model that generates proteins.”

That layer may commoditize rapidly.

As general AI agents become good at selecting and operating RFdiffusion-like, Boltz-like and ProteinMPNN-like systems, access to competent in-silico protein design will become increasingly widespread.

The durable value is more likely to migrate toward four layers:

LayerLikely defensibility
General scientific agent/orchestrationScale, reasoning quality, integrations, enterprise adoption
Automated experimental infrastructurePhysical automation, assay throughput, standardization
Proprietary experimental dataVery high; especially standardized negative data
Therapeutic developmentDisease biology, IP, translational expertise, clinical execution

This is why the Anthropic result may actually strengthen the investment case for companies like Adaptyv and other automated experimental infrastructure providers, rather than only for AI-model companies.


What the future probably looks like

Phase I — 2026–2028: “AI protein engineer”

Protein scientists increasingly stop manually operating every computational model.

An AI agent reads the literature, constructs the target dossier, identifies epitopes, chooses design algorithms, submits GPU jobs, evaluates structures and proposes candidates.

Human scientists supervise the scientific objective and make portfolio decisions.

Design becomes dramatically cheaper and more abundant.

The new bottleneck becomes experimental validation.


Phase II — 2027–2030: “Design campaigns as an API”

The real transformation arrives when Adaptyv-like laboratories become machine-callable.

The workflow becomes:

Target → Agent → Design → Build → Assay → Analyze → Redesign

without humans manually transferring information between stages.

Adaptyv has already built much of the infrastructure required for this architecture through its API.

The broader trend is not unique to Adaptyv. Nature Reviews Chemistry now describes self-driving laboratories as moving from narrow automation toward multipurpose infrastructure where algorithms propose, execute and interpret experiments with limited intervention. Northwestern received $20 million from the NSF in July to construct an openly accessible AI-directed protein-engineering cloud lab.

This suggests a new infrastructure category is forming.


Phase III — 2028–2032: multi-objective therapeutic optimization

Binding affinity becomes only one objective.

Agents simultaneously optimize:

affinity + selectivity + solubility + stability + expression + immunogenicity + half-life + tissue penetration + manufacturability + species cross-reactivity.

Instead of generating thousands of molecules and manually selecting a winner, the system iteratively moves through a multidimensional therapeutic fitness landscape.

This is where AI protein design starts affecting actual pharmaceutical economics.


Phase IV — 2030s: autonomous therapeutic R&D organizations

The long-term architecture could resemble:

Scientific foundation agent
↓
specialized biology models
↓
cloud compute
↓
DNA/protein synthesis
↓
automated cloud laboratory
↓
cellular assays / organoids / in-vivo studies
↓
structured experimental data
↓
active-learning agent
↓
next-generation candidates

Drug discovery begins to behave less like a sequence of manually operated departments and more like a continuously optimizing computational–physical system.

At that point the valuable corporate asset will not be “an AI model.”

It will be the closed loop itself.


Investment implications

Anthropic

The strategic opportunity extends beyond selling tokens.

Claude Science could become an operating system for pharmaceutical R&D, sitting above BioRxiv, PubMed, HPC infrastructure, specialist biology models, ELNs, cloud laboratories and internal pharma systems.

Anthropic has already positioned Claude Science as an integrated scientific workbench and has publicly said it intends to gain direct experience developing medicines itself. Earlier 2026 reporting also showed Bristol Myers Squibb deploying Claude across more than 30,000 employees, including R&D workflows.

The biggest risk is that orchestration is not unique to Claude. Other frontier models can potentially operate the same open-source protein stack.


Adaptyv Bio

Potentially one of the most strategically interesting private infrastructure companies in this emerging ecosystem.

Its moat is not merely automation.

It is potentially:

lab automation + assay standardization + APIs + customer workflows + ProteinBase + accumulated positive/negative experimental data.

If AI-generated candidate volume increases by orders of magnitude, experimental validation becomes the scarce commodity.

That is classic picks-and-shovels exposure.


Twist Bioscience

Synthetic DNA and protein-production infrastructure gains from increasing candidate throughput regardless of which AI model wins.

The Anthropic experiment itself demonstrates this role: Twist acted as one of the independent physical validation layers.


AI-first protein-design companies

The outlook is mixed.

Companies with only an incremental generative model face commoditization pressure.

Companies combining proprietary datasets, differentiated modalities, high-value therapeutic programs, experimental infrastructure or superior multi-objective optimization retain stronger defensibility.

The winning question changes from:

“Whose model makes proteins?”

to:

“Whose system repeatedly turns designs into drug-quality molecules?”


Bottom line for investors

This announcement should not be interpreted as evidence that autonomous AI drug discovery has been solved.

It should be interpreted as evidence that one of the largest remaining pieces of manual scientific workflow—orchestrating sophisticated computational biology tools—can increasingly be delegated to general-purpose AI agents.

That matters enormously.

Protein design is moving from:

artisan science → computational science → automated engineering.

And the investment opportunity will increasingly sit at the interfaces between digital intelligence and physical biology.

The most valuable companies may therefore not be the ones producing the most impressive protein-generation demo.

They may be the companies controlling:

experimental throughput, proprietary data, biological feedback loops, translational validation and ultimately the clinical path to patients.

My base-case expectation is that de novo binder generation becomes increasingly commoditized over the next several years, while the value of automated experimental validation increases substantially.

The winner of the AI-biotech era is unlikely to be a single “AlphaFold for drugs.”

It is much more likely to be a self-improving network of AI agents, specialist molecular models and automated laboratories that can repeatedly convert hypotheses into experimentally verified biology.

That is what the Anthropic–Adaptyv result is beginning to show.


News and research coverage reviewed

SourceWhy it matters
Anthropic, Aug. 18 — “How Claude is accelerating protein design and analytical chemistry”Primary announcement and experimental numbers.
Anthropic technical report — “Autonomous de novo protein binder design with Claude”Full methods, datasets and limitations.
The Decoder, Aug. 19Particularly useful independent analysis of compute budgets, experimental controls and limitations.
Times of India, Aug. 21Latest mainstream coverage of the 14/15-target result.
IBTimes UK, Aug. 20Focuses on therapeutic limitations and biosecurity implications.
AIDB, Aug. 20Detailed breakdown of campaign design and reported hit rates.
NewsBytes, Aug. 19Mainstream technology/science coverage of the results.
Gadgets Now, Aug. 19Places binder work alongside Claude's broader scientific workflow automation.
Techmeme discussion clusterUseful aggregation of reactions from protein scientists, AI researchers and skeptics.
Adaptyv — “Can LLMs design proteins?” May 2026Earlier controlled TREM2 experiment: autonomous agents broadly matched human participants.
Adaptyv API announcementShows how agents can programmatically send designs into a physical wet lab.
Adaptyv ProteinBaseImportant standardized experimental-data layer and potential data flywheel.
Nature, 2026 — “The past, present and future of de novo protein design”Independent scientific context suggesting fundamental protein-design problems are approaching engineering maturity.
Communications Biology, May 2026Describes binder design as approaching an “AlphaFold moment.”
Current Opinion in Structural Biology, June 2026Shows why cross-study hit-rate comparisons require considerable caution.
Stanford AI Index 2026Critical counterpoint: high-affinity Nipah binders did not translate into target neutralization.