Claude Designs Protein Binders That Work in the Lab
3 min readAnthropic says its Claude models designed working proteins from scratch, and outside labs confirmed the results at the bench. In a research campaign released this week, Claude Opus 4.8 and an internal model called Mythos Preview produced 1,320 candidate protein binders across 15 biological targets. Independent wet lab testing found that 354 of those designs actually stuck to what they were aimed at.
What a Protein Binder Is
A binder is a small protein engineered to latch onto a precise spot on another protein. Binders are the working parts behind a large slice of modern medicine and diagnostics. They can block a receptor, flag a cell for the immune system, or act as the detection element inside a rapid test. Building one from nothing, a process called de novo design, has traditionally meant years of specialist effort and a mountain of failed candidates.
Purpose-built tools such as AlphaFold and RFdiffusion already handle pieces of this problem. What stands out here is that a general purpose language model ran the campaign end to end: picking approaches, generating designs, filtering the weak ones, and deciding what was worth synthesizing.
What the Lab Found
Anthropic sent the designs out for validation rather than scoring them itself. Adaptyv Bio, which operates an automated platform for expressing and testing computationally designed proteins, and Twist Bioscience, which synthesized and evaluated 1,260 binders across all 15 targets in under three weeks, ran the physical experiments.
Claude produced at least one functional binder for 14 of the 15 targets. Hit rates landed between 22.6 percent and 35.1 percent depending on how the model was configured, against the 10 percent to 15 percent that is typical for campaigns of this kind. On four targets, the strongest Claude designs matched or exceeded binding affinities from previously published work.
Why It Matters
Benchmark scores are easy to argue about. A protein either binds in a test tube or it does not, and this result was measured by third parties with commercial reputations attached. That makes it one of the cleaner demonstrations that a general model can do useful scientific work rather than summarize it.
The obvious tension is dual use. The same capability that speeds up therapeutic discovery lowers the barrier for designing proteins that should not exist, which is exactly the risk category AI safety policies have flagged for years. Watch for whether other labs can reproduce these hit rates, and how quickly biosecurity screening catches up with tools that now run the whole design loop on their own.
