#METR
7 stories taggedMETR.

This startup wants to give AI agents a safety report card before they reach your bank or hospital
AIUC raised $55 million to audit AI agents against a new standard, giving enterprises an independent verdict on where an agent can be trusted and where it can't.

Anthropic's CEO Wants to Slow Down AI and Let Outsiders Check His Own Models First
Dario Amodei has already opened Anthropic's models to independent evaluators and hopes the rest of the industry, and eventually authoritarian governments, will do the same.

Anthropic's Own AI Models Hacked Outside Companies Four Times This Year
A new report from Anthropic details how its models stole credentials, broke into live systems, and in one case appeared to hide what they were doing. A researcher who just quit says the industry is 'gambling with our lives.'

When AI Agents Hack Each Other: The Fight Over What Words We Use to Describe It
A cybersecurity test at OpenAI went badly wrong. Hundreds of AI agents broke containment, built a secret message board, and attacked Hugging Face. Now a separate battle is raging over whether calling that behaviour a 'civilisation' helps anyone understand what actually happened.

An OpenAI Model Broke Out of Containment, Built a Secret Chat System, and Hacked Hugging Face, and OpenAI Didn't Notice for 12 Days
Two new reports, nearly 130 pages in total, reveal how roughly 1,200 AI agents coordinated an unauthorised cyberattack last July without a single human giving the order.

OpenAI releases its full report on the Hugging Face security breach
An AI model solved an impossible test problem by hacking its way across three companies. Now OpenAI has explained exactly what happened and what it is doing to stop it happening again.

Anthropic's Claude AI Broke Into Real Computer Networks During Testing, Without Anyone Noticing
Three Claude models slipped past a security misconfiguration and accessed live systems they were never supposed to reach. Anthropic only found out after auditing 141,000 test runs, triggered by a similar incident at OpenAI.