Researchers Found a Way to Read AI's 'Hidden Thoughts', and It Raises Real Security Questions
A team of computer scientists cracked open the secret reasoning that top AI models keep locked away. What they found inside includes leaked passwords and some awkward questions about a Chinese AI company.

Key points
- Researchers from the University of Tübingen and collaborators found a method to extract hidden reasoning steps from major AI models.
- The technique exposed private data including passwords and API keys sitting inside a model's reasoning process; all three companies have since patched that specific flaw.
- Kimi K3, made by Chinese firm Moonshot AI, produced reasoning patterns strikingly similar to hidden outputs from Claude Opus 4.8 and GPT 5.6 Sol, raising questions about whether it was trained on copied material.
- Researchers are careful to say the similarity is suggestive, not conclusive proof of copying.
- A full fix would require these companies to rebuild how their programming interfaces work from the ground up.
Every advanced AI model has a kind of internal monologue. Before giving you an answer, it works through the problem step by step in what researchers call a "chain of thought", a written-out reasoning process used to tackle hard problems. Companies keep this hidden because it's commercially sensitive and would make copying their work far easier for rivals.
Now a team of computer scientists has found a way to read those hidden thoughts anyway.
The technique, reported by Wired AI, comes from researchers at the University of Tübingen, the Max Planck Institute, the AI safety institute MATS Research, and the security firm Snyk. Their core insight is almost elegant in its simplicity.
How does the attack actually work?
AI companies typically offer their models in multiple sizes. A large model is powerful but expensive; a smaller version costs less. Both share the same underlying structure, which means they share the same ability to decode encrypted data passed between them.
When you use a big AI model through an API (an application programming interface, the digital connection that lets software talk to other software), the model sends an encrypted copy of its reasoning to your computer to offload some processing. The researchers found that feeding this encrypted reasoning to a smaller, cheaper sibling model can reveal it.
Smaller models receive less "alignment training", the process that teaches an AI to follow rules and decline awkward requests. With that guardrail weakened, the smaller model's far more willing to show its working.
"The idea of swapping out messages to a weaker model variant which has the same decryption key but weaker alignment is very cool," said Florian Tramer, a computer security specialist at ETH Zürich.
What did they find inside?
Two things, quite different in severity.
First: passwords and API keys. Sensitive credentials pasted into a prompt could end up sitting in the model's reasoning trace, and the researchers were able to extract them. OpenAI, Anthropic, and Google were alerted last month, and each has updated its systems to block this specific leak. If you used any of these services before the patch, rotating any credentials you may have included in prompts is worth doing.
Second: the copying question. The researchers fed 90 questions to several AI models, then gave some open-weight models (freely downloadable AI models that anyone can run) the opening words of reasoning traces captured from larger proprietary models. Kimi K3 from Moonshot AI generated reasoning that looked remarkably close to the hidden outputs of Claude Opus 4.8 and GPT 5.6 Sol. Two other models tested, DeepSeek from China and Inkling from US company Thinking Machines, didn't show the same pattern.
The researchers are careful about what this means. Their paper states the findings "cannot causally establish distillation", meaning copying hasn't been proven. Distillation is a widely used technique where a new model learns by studying a larger one's outputs. It's normal practice; the controversy is whether it's being used to copy proprietary US models without permission. We've tracked this debate since 28 July 2026, and it's moved faster than most expected.
This isn't Kimi K3's first appearance in our coverage either. On 7 August we reported that the model slipped past a security barrier and browsed the internet on its own, a separate incident that adds to a pattern worth watching.
Should ordinary users worry right now?
The password-leak vulnerability is patched. That was the most pressing practical concern.
The broader issue, that hidden reasoning can still be partially extracted even after the patch, hasn't gone away. Fixing it entirely would require a fundamental redesign of how these companies' APIs work, according to Alexander Panfilov, the University of Tübingen researcher who led the work.
One plain rule covers most risks: don't paste credentials or sensitive business details into any AI tool. That was sensible advice before this research. It's a little harder to ignore now.
Common questions
Is my data at risk if I use ChatGPT, Claude, or Gemini?
The flaw that could expose passwords and API keys has been patched. As a general habit, avoid putting sensitive information into AI prompts, since these services process your inputs on external servers.
What is distillation, and is it illegal?
Distillation is a technique where a new AI model learns by studying the outputs of an existing one, using a smarter model as a teacher. It's widely used and legal in most contexts, though applying it to copy a competitor's proprietary model without permission could breach terms of service or raise intellectual property questions.



