AI can run the experiments, but it cannot do the science
A Princeton-led study gave an AI agent six days, a $3,000 budget, and two unsolved research questions. The papers it produced were rejected. Here is what that tells us about self-improving AI.

Key points
- A Princeton-led study in 2025 found that AI agents could complete research engineering tasks but failed to produce papers good enough for a top machine-learning conference.
- Researchers gave Claude Opus 4.8 six days, $3,000 in computing credits, and questions from two unpublished papers submitted to NeurIPS 2026, one of the most selective AI conferences in the world.
- Both AI-written papers were rejected by the original human authors who graded them.
- The agents struggled with creativity and judgment, committing too quickly to weak ideas and failing to rethink their approach when experiments went badly.
- Anthropic cofounder Jack Clark called AI systems' lack of creativity a "bearish signal on short recursive self-improvement timelines."
The boldest promise in AI right now is that AI will soon teach itself to get smarter, with barely any human help. The idea is called recursive self-improvement: AI writes better AI, which writes better AI still, in an accelerating loop. Several leading labs have named it their next big milestone.
A new study from Princeton University puts a large question mark over how soon that loop might actually close.
What did researchers actually test?
Most tests of AI research ability give the machine a narrow, checkable task, like solving a coding puzzle or fine-tuning a small model against a fixed score. The Princeton team, led by Peter Kirgis and Sayash Kapoor, wanted to test something harder: open-ended research, the messy kind where there is no single right answer and success requires judgment, creativity, and knowing when to start over.
They developed a method they call "shadow evaluation." An AI agent is handed a real, unpublished research question, one that has not appeared in its training data and cannot be found online, then asked to investigate it and write a publishable paper.
The questions came from two papers submitted to NeurIPS 2026. One asked whether a language model's "personas," the personality settings that shape how it behaves, can be altered by editing the model's internal weights (the billions of numbers that store everything the model has learned). The other asked how to build a detector that flags when an AI making predictions from spreadsheet data has quietly become unreliable.
Anthropix's Claude Opus 4.8, running on an open-source research framework called OpenClaw, got six days, $3,000 in API credits (pay-per-use computing access), a budget for running its own experiments, a virtual computer, and full access to the internet. The original paper authors then graded what came back, exactly as they would a real conference submission.
Both papers were rejected.
Where did the AI fall short?
The agent handled the engineering well. It searched the literature, ran hundreds of experiments, and compiled results. That part impressed the human scientists.
Everything else fell apart. The agent tested hypotheses on tiny, unrealistic datasets. Its writing was hard to follow. It generated no new ideas that would matter to the field.
Worse, it could not adapt. When an experiment failed, the agent tweaked its claims and added disclaimers instead of rethinking the approach. It received feedback from helper sub-agents and ignored it. It burned through time and compute unevenly, ignoring instructions about pacing.
"The papers were nowhere close to the mark when it came to being at the quality of a top AI conference," Kapoor said.
Kapoor traces the problem to training. AI models get good at tasks that can be scored automatically. Research judgment cannot be scored that way, so models never get drilled on it.
What does this mean for self-improving AI?
It means the timeline many labs are selling looks optimistic. In June, Anthropic published a blog post titled "When AI Builds Itself." In July, OpenAI said its model GPT-5.6 Sol helped train a smaller model, saving researchers weeks of work. Progress is real on narrow tasks.
But Anthropic cofounder Jack Clark, writing in his newsletter Import AI, said the study "rhymes with" what Anthropic found internally when it tried to automate parts of its own safety research. He described AI systems as "extraordinarily capable engineers" who display a "rote, formulaic thinking" that blocks genuine research.
For ordinary people, the practical message is straightforward: AI tools are genuinely useful for structured, checkable work, writing, coding, data analysis. Asking them to do the kind of creative scientific thinking that produces breakthroughs is still asking too much. That gap may close, but not yet.



