Science's next AI leap will come from reasoning agents, not data-hungry models like AlphaFold

AlphaFold was a landmark, but the conditions that made it possible took 53 years and $21 billion to build. For most of science, a different kind of AI, one that reasons and experiments step by step, is already showing what comes next.

AI2Day NewsdeskUpdated Editor: Lee Brown4 min read
Photoreal news-editorial style, 16:9 framing, full-frame edge-to-edge
Share

Key points

  • Google DeepMind's AlphaFold won Demis Hassabis and John Jumper the 2024 Nobel Prize in chemistry for predicting protein shapes.
  • The protein database AlphaFold trained on took 53 years and an estimated $21 billion in experimental work to assemble.
  • Most scientific fields cannot produce datasets of that quality, making AlphaFold-style AI a poor model for wider science.
  • AI agents, software that reasons through problems and calls on multiple tools in sequence, are now matching decade-long human research findings in hours.
  • A side benefit: agents automatically log every step, which could help fix science's long-running reproducibility crisis.

Two years ago, the idea that a computer program could solve one of biology's hardest puzzles seemed extraordinary. Then AlphaFold did it.

The program, built by Google DeepMind, predicts the three-dimensional shape of a protein, a molecule whose shape determines what it does inside the body, from its genetic sequence alone. Scientists had chased this problem for fifty years. Demis Hassabis and John Jumper won the 2024 Nobel Prize in chemistry for it. Startups flooded into biology and chemistry, raising billions on the assumption that AI plus data would now open discovery after discovery.

That assumption, according to an analysis first reported by MIT Technology Review, deserves scrutiny.

Why can't we just repeat what AlphaFold did?

The data required almost never exists. AlphaFold trained on the Protein Data Bank, roughly 170,000 experimentally confirmed protein shapes built over 53 years of international cooperation at an estimated cost of $21 billion. That kind of dataset is extraordinarily rare. We reported in July on Google's $40 million AI-credits push into national science labs, where one lab cut a 90-minute task to 13 minutes, precisely because useful data there already existed at scale.

Most experimental science is messier. Cell cultures change subtly over time. Chemicals carry trace impurities. Lab humidity shifts between readings. The clean, standardised, large-scale data that a modern neural network, the statistical pattern-matcher at the heart of tools like AlphaFold, needs to learn from simply doesn't exist in most fields, and building it would take decades.

A handful of areas, weather forecasting and genomics among them, do have the right kind of data and may see AlphaFold-style breakthroughs soon. They're the exception, not the rule.

So what does work?

AI agents are showing real promise, and they work very differently. An agent is software that reasons through a multi-step problem by calling on different tools, keeping what works and discarding what fails. It behaves less like a lookup table and more like a careful scientist.

Google's AI Co-Scientist, announced in May, is one early example. Researchers gave it a single-page brief and one goal: work out how antibiotic resistance, the ability of bacteria to survive drugs designed to kill them, spreads between bacterial species. The system divided the task across several sub-programs. One generated hypotheses from the scientific literature. A second challenged each hypothesis the way a peer reviewer would. A third ranked the survivors, and a fourth refined the winner.

The agent concluded that resistance genes were hitching rides inside bacterial viruses, borrowing whichever virus could carry them into a new host. Researchers at Imperial College London had spent a decade reaching the same conclusion through painstaking lab work. Their paper was still in peer review when Co-Scientist produced its answer.

What does this mean for patients and the public?

Faster, more reproducible science means faster routes from laboratory finding to clinical trial. It also matters for trust. Science's reproducibility crisis, where researchers repeatedly find they can't replicate each other's results, has quietly undermined confidence in medical research for years. Agents log every step automatically, creating an exact record of method. That record makes replication straightforward in a way that hand-written, often incomplete lab notebooks rarely are.

Agents aren't perfect. They still occasionally produce confident-sounding errors, and their reasoning isn't always consistent. Those are real constraints. But they're engineering problems, not fundamental barriers.

The bigger shift is quieter than a Nobel Prize, and it may matter more. My read: the reproducibility benefit is underplayed in most coverage. A tool that speeds up science and makes it auditable at the same time is rarer than headlines suggest, and that's the part worth watching.

Common questions

Does this mean AI is about to replace scientists?

No. Agents work best as a tireless assistant that tests ideas in sequence and keeps precise records. Human judgement and ethical oversight remain central to deciding which questions are worth asking.

Will this speed up drug discovery for patients?

Potentially yes, though timelines are hard to predict. If agents can compress the hypothesis-testing stage of research from years to weeks, clinical trials, the lengthy human-safety studies required before any drug reaches patients, would begin sooner. The trials themselves still take the same time.

© 2026 AI2Day