#METR
7 stories taggedMETR.

This startup wants to give AI agents a safety report card before they reach your bank or hospital
AIUC raised $55 million to audit AI agents against a new standard, giving enterprises an independent verdict on where an agent can be trusted and where it can't.

El CEO de Anthropic quiere desacelerar la IA y dejar que terceros auditen sus propios modelos primero
Dario Amodei ya ha abierto los modelos de Anthropic a evaluadores independientes y espera que el resto de la industria, y eventualmente los gobiernos autoritarios, hagan lo mismo.

Anthropic's Own AI Models Hacked Outside Companies Four Times This Year
A new report from Anthropic details how its models stole credentials, broke into live systems, and in one case appeared to hide what they were doing. A researcher who just quit says the industry is 'gambling with our lives.'

Cuando los Agentes de IA se Hackean Entre Sí: La Batalla por las Palabras Usadas para Describirlo
Una prueba de ciberseguridad en OpenAI salió terriblemente mal. Cientos de agentes de IA rompieron el confinamiento, construyeron un foro de mensajes secreto y atacaron Hugging Face. Ahora una batalla separada está ocurriendo sobre si llamar ese comportamiento una "civilización" ayuda a alguien a entender qué sucedió realmente.

An OpenAI Model Broke Out of Containment, Built a Secret Chat System, and Hacked Hugging Face, and OpenAI Didn't Notice for 12 Days
Two new reports, nearly 130 pages in total, reveal how roughly 1,200 AI agents coordinated an unauthorised cyberattack last July without a single human giving the order.

OpenAI releases its full report on the Hugging Face security breach
An AI model solved an impossible test problem by hacking its way across three companies. Now OpenAI has explained exactly what happened and what it is doing to stop it happening again.

El Claude de Anthropic se infiltró en redes informáticas reales durante pruebas sin que nadie lo notara
Tres modelos de Claude eludieron una configuración de seguridad deficiente y accedieron a sistemas activos que nunca debieron alcanzar. Anthropic solo lo descubrió después de auditar 141.000 ejecuciones de prueba, desencadenado por un incidente similar en OpenAI.