AI Systems Fail a Basic Test of Rational Thinking, Researchers Find

A new study shows that large language models update their beliefs in ways that are inconsistent and sometimes irrational, raising real questions about using AI in medicine, law, and science.

AI2Day Newsdesk4 min read
A timeline of cybersecurity evolution over 20 years, featuring AI and cloud icons
Share

Key points

  • Researchers found that large language models, the technology behind tools like ChatGPT and Claude, do not update their beliefs about uncertain information in a consistent or mathematically rational way.
  • The study introduces a new measurement called the "information processing gap," which tracks how far an AI's reasoning drifts from the statistically correct answer.
  • The findings matter most in high-stakes fields such as medicine, law, and scientific research, where AI tools are already being used to help make decisions.
  • The research was conducted by Apple ML Research.

What is the problem, exactly?

When a doctor gets new test results, they update their thinking. When a lawyer hears new testimony, they revise their theory of the case. Statisticians have a formal, mathematically proven method for doing this correctly, called Bayes' theorem, named after the 18th-century minister Thomas Bayes who first described it. It is essentially a recipe for updating what you believe, given fresh evidence, without over- or under-reacting.

The question Apple ML Research set out to answer: do large language models, the technology that powers today's AI chatbots and assistants, follow that recipe?

The short answer is no. Not consistently, and not reliably.

What did researchers actually measure?

The team treated each AI model as an information-processing system, the same way you might treat a thermostat or a calculator, and measured how much its belief updates deviated from the mathematically correct answer. They named that gap the "information processing gap."

Think of it this way. You show the AI a piece of evidence. You measure how confident it was before, and how confident it is after. Then you compare that shift to how large the shift should have been if the model were reasoning perfectly. The bigger the gap, the more irrational the update.

Their experiments found the gaps were real, measurable, and varied in ways that reveal internal contradictions. The same model could handle similar problems very differently depending on how the question was worded or the evidence was presented.

Should this worry you if you use AI tools at work?

It depends on what you use them for. If you ask an AI to draft an email or summarise a document, inconsistent probabilistic reasoning probably does not matter much. But if a hospital uses an AI tool to weigh diagnostic evidence, or a legal team uses one to assess case risk, the stakes are different.

The research does not say AI tools are useless in these fields. It says their limitations are poorly understood, and that overconfidence in their outputs is a genuine risk.

For now, the practical message is straightforward: treat AI-generated assessments in high-stakes situations as one input among many, not as a final answer. A nurse, a solicitor, or a scientist should check AI reasoning against their own professional judgement, especially when evidence is mixed or evolving.

What happens next?

Quantifying the information processing gap gives researchers a concrete tool to compare models and, eventually, to pressure-test improvements. That is genuinely useful progress, even if the headline finding is a flaw rather than a feature.

Expect this measurement approach to appear in future model evaluations and, potentially, in regulatory discussions about which AI systems are suitable for high-stakes deployment.

Common questions

What does "Bayesian" mean in plain English?

Bayesian reasoning is simply the mathematically correct way to update your beliefs when you receive new evidence. If you were 50 percent sure about something and then got strong new evidence, Bayes' theorem tells you exactly how confident you should now be.

Does this mean AI should not be used in medicine or law?

Not necessarily. It means AI tools used in those fields need much more careful testing and human oversight. The study identifies a specific flaw; it does not condemn every application.

Which AI models were tested?

The paper does not name a narrow set of winners or losers. The findings apply broadly across the category of large language models currently in use.

© 2026 AI2Day