The Best AI Agents Still Fail at Most Real Office Work. ServiceNow Built a Tool to Fix That.
ServiceNow's AutoSynthData spots where an AI agent breaks down on the job, then automatically generates practice tasks to fix the gap. The backdrop: the strongest model tested cleared barely a third of realistic enterprise tasks.

Key points
- The top scorer in the EnterpriseOps-Gym benchmark completed only around a third of realistic enterprise tasks, a striking ceiling for the most capable models available.
- EnterpriseOps-Gym tests AI agents across expert-written tasks spanning eight business departments, with simulated database states and strict permission rules.
- Giving agents a human-written plan before each task lifted scores by 14 to 35 percentage points, pointing to planning as the core weakness.
- ServiceNow's AutoSynthData, published via Hugging Face, automatically converts an agent's failures into targeted training exercises without human labellers.
AI agents, software that carries out multi-step tasks on its own rather than just answering questions, are being sold hard to businesses right now. Let the software handle the ticket routing, the HR lookups, the IT requests. The reality, according to new research, is messier than that pitch suggests.
A benchmark called EnterpriseOps-Gym put 14 of the most capable models through tasks modelled on real office software environments: databases that change state, strict permission rules, tools requiring exactly the right sequence of steps. The best performer cleared only around a third of tasks. Every other frontier model, the term researchers use for the largest and most advanced AI systems available, did worse.
That ceiling is the number that matters. It means the strongest commercially available agent fails at roughly two in every three realistic office jobs it's handed.
What is actually going wrong?
Planning is the bottleneck. When researchers handed agents a ready-made human plan alongside each task, scores jumped by 14 to 35 percentage points depending on the model. The agents could execute steps well enough. They just couldn't figure out the right steps on their own across a long sequence.
This matches what our 17 September story on agent reliability found: models that look impressive on short, clean tasks hit a wall when real-world friction, state changes and ambiguity enter the picture.
What does AutoSynthData actually do?
ServiceNow's CoreAI team built AutoSynthData to attack that wall systematically. The tool watches a target model fail, then compares those failures against a stronger "teacher" model that succeeds on the same tasks. From that comparison it extracts a description of the missing skill, something like "the agent can't look up a customer record and then cross-check it against an access policy before acting."
It then generates fresh practice tasks that drill that specific skill, varying the entities, database states and phrasing each time. Each candidate goes through automated validation: the task must be completable, realistic for the environment, and verified correctly so the checker rewards genuine success rather than penalising valid solutions. Only tasks that pass every check become training data.
The process repeats. After training, the tool re-evaluates the model, finds the next gap, builds the next batch. No human labellers write the tasks, which matters because labelling is usually the slow, expensive part of improving a model on a specific job.
The training data question has become one of the defining tensions in AI development. Court documents revealed last year that labs are making aggressive choices about what they train on. AutoSynthData points toward a different approach: generating synthetic, targeted data from the model's own failure record rather than scraping existing content. We've tracked this shift across ten training-data stories since July.
What does this mean for businesses buying AI agents?
If you're evaluating AI agents for internal workflows, that one-third success rate is a useful reality check before signing a contract. Ask vendors how their agent performs on tasks that mirror your actual systems, not on general capability benchmarks. The gap between an impressive demo and something that works in your environment is exactly what EnterpriseOps-Gym was designed to measure, and right now that gap is large.
Common questions
Is this a finished product businesses can buy?
AutoSynthData is a research pipeline published by ServiceNow CoreAI, not a packaged product. Businesses can't buy it directly today, but the methodology and the EnterpriseOps-Gym benchmark are publicly available for teams with the technical resources to apply them.
Does this mean AI agents are not ready for the workplace?
They handle narrow, well-defined tasks reliably. The one-third figure applies to complex, multi-step workflows across large enterprise systems with strict access rules. Simpler, more tightly scoped tasks perform better, which is why most successful business deployments today keep agents on a short leash.



