Logo
FrontierNews.ai

Claude Just Solved Protein Design Problems That Stumped Human Experts

Anthropic's Claude models have achieved a breakthrough in protein design, hitting 14 out of 15 targets with success rates that dramatically exceed what human experts typically achieve. In a new experiment released by the company, Claude's Mythos and Opus models designed 1,320 protein candidates from scratch, with 354 successfully binding to their targets. The highest hit rate reached 35.1%, compared to the typical 10 to 15 percent success rate in current protein design projects.

What Makes This Protein Design Achievement Different From AlphaFold?

Many people associate protein work with Google DeepMind's AlphaFold project, but Claude's task is fundamentally different. AlphaFold predicts the three-dimensional shape a protein will fold into when given its amino acid sequence. Claude, by contrast, faced the inverse problem: given a target protein, design an entirely new protein that doesn't exist in nature and can bind to it with extreme precision.

The proteins Claude designed are called minibinders, miniature binding proteins typically only 50 to 70 amino acids long. Their job is to attach to a specific site on a target protein with high accuracy and intensity. This binding action is crucial in drug development; many medications work by binding to harmful proteins to deactivate them, binding to beneficial proteins to activate them, or delivering drug molecules directly to disease sites.

Traditionally, this work is extraordinarily labor-intensive. Protein designers must run multiple professional models, repeatedly generate and optimize candidates, and screen them through a process that relies heavily on experience. As Anthropic noted, each target typically requires experts to spend weeks or even months screening through numerous candidate molecules to find a few truly effective ones.

How Did Claude Achieve These Results?

Anthropic gave Claude access to a comprehensive toolkit: research papers, network resources, multiple professional protein models, and substantial computing power. The company tested two different approaches to see which worked better.

  • Multi-target mode: Claude processed multiple targets simultaneously in a single 48-hour task, with access to up to 12,500 H100 GPU hours of computing power.
  • Single-target mode: Claude focused on one target at a time, running all tasks in parallel for 24 hours, with each target able to use up to 2,500 H100 GPU hours.
  • Performance difference: Focusing on one task proved significantly better than multitasking, with Mythos Preview's hit rate jumping from 26.7 percent in multi-target mode to 35.1 percent in single-target mode.

The results were striking. On the RBX1 target, which regulates protein degradation, Mythos Preview achieved a 40 percent hit rate. This far exceeded the 3.7 percent overall hit rate from all participants in a previous design competition held by Adaptyv Bio. Claude's top-ranked design also produced a high-affinity binding protein whose binding intensity exceeded the championship design selected from 245 competing designs.

On another difficult target called TNFα, an inflammatory signaling protein, Opus 4.8 succeeded where Mythos Preview failed. Opus 4.8 designed multiple effective binding proteins, some of which can bind to TNFα across humans, cynomolgus monkeys, and mice simultaneously. This cross-species binding ability is critical for future animal testing, eliminating the need to redesign proteins for different species.

Can Claude Analyze Lab Data as Well as It Designs Proteins?

Anthropic's experiment didn't stop at design. The company wanted to test whether Claude could independently read experimental data and determine what compounds were created and how pure they were. Claude Opus 5 tackled two labor-intensive analytical chemistry tasks: nuclear magnetic resonance spectroscopy (NMR) and liquid chromatography-mass spectrometry (LC-MS).

These tasks normally require chemists to manually process raw data in proprietary formats, performing calibration, peak picking, integration, and verification step by step. The work is tedious and relies heavily on experience; small errors can lead to wrong conclusions. Anthropic gave Claude only the raw files from the laboratory and a two-sentence task instruction, with no vendor software or operator guidance.

Claude completed the NMR analysis in 23 minutes and the LC-MS analysis in 19 minutes. For NMR, Claude identified 18 signal peaks and calculated the number of hydrogen atoms corresponding to each peak, with an error of no more than 0.08 hydrogen atoms compared to the laboratory's results. The final measured sample purity was 96.4 percent, while the laboratory reported 96.33 percent, demonstrating remarkable consistency.

What Are the Limitations of Claude's Protein Design Capabilities?

Despite these impressive results, Claude's abilities remain uneven. When challenged to design beta-sheet structures, Claude successfully produced 15 effective binding proteins on 6 targets. However, when switching to the maltose-binding protein (MBP), Claude failed completely; none of the 90 designs were confirmed successful, with only one showing a weak signal.

Anthropic acknowledged that it has not yet determined why Mythos Preview succeeded on some targets while Opus 4.8 succeeded on others. The company stated that the scientific research capabilities of large language models remain uneven across different problem types. This suggests that while Claude has demonstrated remarkable potential in protein design, the technology is still in early stages and cannot yet reliably handle all protein design challenges.

Steps to Understanding Claude's Impact on Drug Development

  • Current limitations: Claude cannot yet automatically convert a target protein into a finished drug; it still fails on certain protein types and requires human oversight and validation of results.
  • Efficiency gains: Claude can autonomously schedule professional models and complete design and analysis processes that previously required experts to spend weeks or months organizing and screening candidates.
  • Practical applications: The technology could accelerate early-stage drug discovery by rapidly generating and analyzing protein candidates, reducing the time and cost of initial screening phases.

The breakthrough demonstrates that large language models like Claude can move beyond text-based tasks into specialized scientific domains. However, the uneven performance across different protein types suggests that the field is still learning how to best apply these models to complex biological problems. Anthropic's open-sourcing of the relevant data and prompts may help other researchers understand and improve upon these results.