Apple Asks Whether AI Coding Agents Actually Need All That Extra Machinery

New research from Apple ML Research finds that a single, well-trained AI agent given direct access to a computer can match or beat far more complex multi-agent systems on real machine learning tasks.

AI2Day NewsdeskEditor: Lee Brown3 min read
A single glowing mechanical arm precisely assembling circuit components on a clean white workbench, surrounded by dozens of unused, disconnected robotic arms ha
Share

Key points

  • Apple ML Research found that a plain coding agent with direct computer access matched multi-agent AI systems on autonomous machine learning benchmarks, with no orchestration layer needed.
  • Modern AI agents for coding have grown increasingly complex, stacking multiple specialised sub-agents and retrieval tools on top of elaborate coordination software.
  • The research challenges a widespread assumption in the field: that more scaffolding always means better results.
  • Simpler systems may soon do what complex ones currently do, at lower cost and with fewer points of failure.

For the past year, building an AI coding agent has started to look a bit like assembling flat-pack furniture with too many parts. Researchers bolt on an orchestrator to manage several sub-agents, add a retrieval agent to search documentation, then another layer to coordinate the whole thing. The assumption has been that more structure equals better performance.

Apple ML Research is pushing back on that.

The team studied what they call MLE agents, short for machine learning engineering agents. These are AI programs that carry out full software engineering tasks on their own, running experiments, writing code, reading results, without a human stepping in. They wanted to know how much of that surrounding machinery is actually doing useful work.

What did they actually find?

A straightforward coding agent, one where the underlying AI model can directly read files, write code and run commands in a real computer environment, held its own against far more elaborate setups. No multi-agent orchestra, no dedicated research sub-agent. Just a capable model with access to the tools a human engineer would use.

The finding matters because the field has been moving hard in the opposite direction. Elaborate "harnesses", the term researchers use for the surrounding machinery that manages an AI agent's behaviour, have become the default. The implicit logic was that stagnating performance on long, multi-step tasks required more infrastructure to fix.

Apple's work suggests the infrastructure may have been papering over weaknesses in the underlying model. Improve the model itself, and the scaffolding becomes less necessary.

What does this mean for ordinary people?

If you use AI coding tools at work or rely on software built with AI assistance, simpler systems may soon do what complex ones currently do, more cheaply. Fewer moving parts means fewer points of failure.

It also fits a pattern AI2Day has been tracking. Nvidia's Kumo Tabular demonstrated on 29 September that a single model reading a spreadsheet cold can beat every hand-tuned specialist. The ceiling on a strong general model keeps proving higher than the field assumed, and it's not a coincidence that Apple ML Research keeps turning up in these results. We covered a separate Apple ML Research paper on 11 September, when their protein design model skipped an entire training step that the field had treated as unavoidable.

The honest caveat: this research was tested on specific public benchmarks, controlled problems rather than the messier tasks real software teams face. Complex harnesses may still earn their keep in production settings where the environment is noisier.

Apple ML Research has asked a question the rest of the field will now have to answer with data rather than assumption.

© 2026 AI2Day