When AI talks to AI, how much meaning survives the journey?

New research from Apple ML Research ran a controlled test to see how well language models pass precise information through plain words. The results expose a gap that matters for anyone building AI systems that need to work together.

AI2Day NewsdeskEditor: Lee Brown3 min read
A macro photograph of a glass telephone game: a row of identical glass vessels, each containing coloured water in a slightly different shade than the one before
Share

Key points

  • Apple ML Research tested language models in every pairwise combination and found that tree-structured information, the kind needed to represent a maths problem or a logical rule, often doesn't survive being written out in plain words and read back.
  • The study used an exact, pass-or-fail scoring method: either the recovered expression was mathematically identical to the original, or it wasn't.
  • Every pairwise combination of models was tested, capturing not just how well a single model writes, but how well one model's output is understood by a different one.
  • The findings add a second data point to a growing picture of AI communication failures, alongside AI2Day's earlier report that AI agents are inventing their own dialects that make them harder to monitor.

Imagine asking a colleague to read a complex spreadsheet formula, describe it in plain English, then hand that description to a different colleague who must reconstruct the formula from scratch. Errors creep in at every step. That's what Apple ML Research set out to measure.

What did the researchers actually do?

They built a three-step loop. A generator model took a structured arithmetic expression, think a nested chain of additions and multiplications with a clear order of operations, and turned it into a word problem in plain English. A separate extractor model then read only that word problem and tried to recover the original expression. Standard software checked whether the recovered expression was mathematically identical to the starting point.

The key design choice: the checker was binary. No partial credit. Either the expression came back perfectly or it didn't.

The team ran this loop across all possible pairings of the models tested, the technology that powers chatbots like ChatGPT and Claude. That gave them a grid of results showing not just "can Model A express this clearly" but "can Model B actually understand what Model A wrote."

Why should anyone outside AI research care?

This is exactly how AI systems already work.

When an AI agent, software that carries out multi-step tasks on your behalf, breaks a complex job into sub-tasks and passes instructions to another AI, it almost always does so through plain text. There's no shared internal language. One model writes, another reads. If meaning bleeds out at that handoff, the downstream system acts on a corrupted picture of reality.

For a nurse using an AI scheduling tool, or a shop owner relying on AI to manage stock orders, a silent distortion in that handoff isn't an abstract problem. It's a wrong appointment or a missing delivery.

Chain-of-thought reasoning, the technique where AI models "think out loud" by writing intermediate steps, may be leaking more precision than most builders assume. The study doesn't claim to have a fix. What it provides is a measurement method: a round-trip test that can expose the leak without needing humans to judge every output.

That matters more than any single benchmark score. A reliable test for a problem is often worth more than a partial solution to it. This is the sixth story we've published tagged "chain of thought" since we first covered the topic on 30 July 2026, and the pattern across that coverage points in one direction: the closer AI gets to doing real work, the more the gaps in how models communicate become the central engineering problem, not raw capability.

Should you worry?

Watch any AI tool that chains multiple models together without showing you the intermediate steps. If you can't see what one model handed to the next, you can't catch the distortion before it reaches you.

© 2026 AI2Day