Anthropic let Claude run protein design alone for two days; outside labs built every design, the 27% hit rate came with one clean miss.
Over one to two days, with a written protocol and a hard time limit on a cloud account, Claude read up on each of sixteen protein targets, picked which surface to aim at, and ran the design software itself. Two outside labs, hired by Anthropic, then built and measured all 1,320 candidate proteins it produced, exactly as delivered. Anthropic's technical report lays out the work, the dataset is public on Hugging Face, and a preprint mirrors the report on alphaxiv. The autonomy went further than expected. The cost did not move.
The judgment that used to take months of specialist protein design got compressed into a roughly 16,000-word protocol. Humans only chose targets, paid for compute, placed the lab orders, and read the final data. Every design decision in between was the model's call. A binder, in this context, is a small protein built to latch onto one specific spot on a larger protein. Finding useful binders against disease targets is the early step in most modern drug discovery, and it is where the human bottleneck has lived for two decades.
Across all 1,320 designs, 354 bound their target, a 27% overall rate. Among the top-ranked design per target, the rate jumped to 49%. On fifteen of the sixteen targets, the labs got readable measurements; fourteen of those fifteen yielded at least one working design. Anthropic's research page and Forbes's synthesis report the same numbers.
The head-to-head on RBX1, a target from a public human design competition, was sharper. Human competitors entered 245 designs; 9 bound. Claude produced 90 designs; 28 bound. When the winning entry from each side was re-made and measured on the same plate, Claude's best held roughly 10x tighter than the human competition winner. On TNFα, a target where published efforts had reported zero working binders, the model produced twelve.
One target did not work. On MBP, the labs reported 0 of 90 designs bound. A clean miss is what makes the rest of the numbers usable. The 14-of-15 headline has one genuine design failure, and the report does not bury it. That miss is the comparison point for the 27% rate.
The cost picture is the second half of the news. Every piece of design software in the pipeline is open source and free, and Anthropic paid no licensing fees. The team explicitly excluded AlphaFold 3's model weights and two other well-known packages because their licenses were restrictive, a constraint Anthropic flagged in the technical report and that India Today covered in its synthesis. The real money sat with the two external labs that physically built all 1,320 proteins and ran the binding assays, not with compute or software.
Months of specialist judgment got absorbed into a written protocol run on a cloud account. The wet-lab cost of producing and testing every design did not move. The bottleneck for autonomous protein design has moved to the bench work, not the design step.