Science's next AI leap will come from reasoning agents, not data-hungry models like AlphaFold

AlphaFold was a landmark, but the conditions that made it possible took 53 years and $21 billion to build. For most of science, a different kind of AI, one that reasons and experiments step by step, is already showing what comes next.

AI2Day Newsdesk4 min read
Photoreal news-editorial style, 16:9 framing, full-frame edge-to-edge
Share

Key points

  • Google DeepMind's AlphaFold won Demis Hassabis and John Jumper the 2024 Nobel Prize in chemistry for predicting the shapes of proteins.
  • The protein database AlphaFold trained on took 53 years and an estimated $21 billion in experimental work to assemble.
  • Most scientific fields cannot produce datasets of that quality, making AlphaFold-style AI a poor model for wider science.
  • AI agents, software that reasons through problems and uses multiple tools in sequence, are now matching decade-long human research findings in hours.
  • A key side benefit: agents automatically log every step they take, which could help fix science's long-running reproducibility crisis.

Two years ago, the idea that a computer program could solve one of biology's hardest puzzles seemed extraordinary. Then AlphaFold did it.

The program, built by Google DeepMind, can predict the three-dimensional shape of a protein, a molecule whose shape determines what it does inside the body, from its genetic sequence alone. Scientists had chased this problem for fifty years. In 2024, Demis Hassabis and John Jumper won the Nobel Prize in chemistry for it. Startups flooded into biology and chemistry, raising billions on the assumption that AI plus data would now unlock discovery after discovery.

That assumption, according to a detailed analysis first reported by MIT Technology Review, deserves a closer look.

Why can't we just repeat what AlphaFold did?

Simply put: the data required to do so almost never exists. AlphaFold trained on the Protein Data Bank, a collection of roughly 170,000 experimentally confirmed protein shapes built over 53 years of international cooperation, at an estimated cost of $21 billion. That kind of dataset is extraordinarily rare.

Most experimental science is messier. Cell cultures change subtly over time. Chemicals carry trace impurities. Lab humidity shifts between readings. Results that one team can replicate on a Tuesday may not hold for another team on a Thursday. The clean, standardised, large-scale data that a modern neural network, the statistical pattern-matcher at the heart of tools like AlphaFold, needs to learn from simply does not exist in most fields, and building it would take decades.

A handful of areas, weather forecasting and parts of genomics among them, do have the right kind of data and may see AlphaFold-style breakthroughs soon. But they are the exception.

So what does work?

AI agents are showing real promise, and they work very differently. An agent is software that reasons through a multi-step problem by calling on different tools, picking the best result, discarding what fails, and trying again. It behaves less like a lookup table and more like a careful scientist.

Google's AI Co-Scientist, announced in May, is one early example. Researchers gave it a single-page brief and one goal: work out how antibiotic resistance, the ability of bacteria to survive drugs designed to kill them, spreads between bacterial species. The system divided the task across several sub-programs. One generated hypotheses from the scientific literature. A second challenged each hypothesis the way a peer reviewer would. A third ranked the survivors. A fourth refined the winner.

The agent concluded that resistance genes were hitching rides inside bacterial viruses, borrowing whichever virus could carry them into a new host. Researchers at Imperial College London had spent a decade reaching the same conclusion through painstaking lab work. Their paper was still in peer review when Co-Scientist produced its answer.

What does this mean for patients and the public?

Faster and more reproducible science means faster routes from laboratory finding to clinical trial. It also matters for trust. Science's reproducibility crisis, where researchers repeatedly find they cannot replicate each other's results, has quietly undermined confidence in medical research for years. Agents log every step they take automatically, creating an exact record of their method. That record makes replication straightforward in a way that human lab notebooks, often hand-written and incomplete, rarely are.

Agents are not perfect. They still occasionally produce confident-sounding errors, their reasoning is not always consistent, and they have limits on how long they can work without human oversight. Those are real constraints. But they are engineering problems, not fundamental barriers.

The bigger shift is quieter than a Nobel Prize, and it may matter more.

Common questions

Does this mean AI is about to replace scientists?

No. Agents work best as a tireless assistant that tests ideas in sequence and keeps precise records. Human judgment, creativity and ethical oversight remain central to deciding which questions are worth asking.

Will this speed up drug discovery for patients?

Potentially yes, though timelines are hard to predict. If agents can compress the hypothesis-testing stage of research from years to weeks, clinical trials, which are the lengthy human-safety studies required before any drug reaches patients, would begin sooner. The trials themselves still take the same time.

© 2026 AI2Day