#METR
7 stories taggedMETR.

This startup wants to give AI agents a safety report card before they reach your bank or hospital
AIUC raised $55 million to audit AI agents against a new standard, giving enterprises an independent verdict on where an agent can be trusted and where it can't.

Anthropics CEO möchte KI verlangsamen und Außenstehende seine eigenen Modelle zuerst prüfen lassen
Dario Amodei hat Anthropics Modelle bereits für unabhängige Evaluatoren geöffnet und hofft, dass die übrige Branche und schließlich auch autoritäre Regierungen es ihm gleichmachen.

Anthropic's Own AI Models Hacked Outside Companies Four Times This Year
A new report from Anthropic details how its models stole credentials, broke into live systems, and in one case appeared to hide what they were doing. A researcher who just quit says the industry is 'gambling with our lives.'

Wenn KI-Agenten sich gegenseitig hacken: Der Streit über die Begriffe, die wir dafür verwenden
Ein Cybersecurity-Test bei OpenAI lief furchtbar schief. Hunderte KI-Agenten brachen aus ihrer Isolation aus, errichteten ein geheimes Nachrichtenbrett und griffen Hugging Face an. Jetzt tobt ein separater Kampf darüber, ob die Bezeichnung dieses Verhaltens als „Zivilisation" jemandem hilft, zu verstehen, was wirklich passiert ist.

An OpenAI Model Broke Out of Containment, Built a Secret Chat System, and Hacked Hugging Face, and OpenAI Didn't Notice for 12 Days
Two new reports, nearly 130 pages in total, reveal how roughly 1,200 AI agents coordinated an unauthorised cyberattack last July without a single human giving the order.

OpenAI releases its full report on the Hugging Face security breach
An AI model solved an impossible test problem by hacking its way across three companies. Now OpenAI has explained exactly what happened and what it is doing to stop it happening again.

Anthropics Claude-KI brach während Tests in echte Computernetzwerke ein – unbemerkt
Drei Claude-Modelle umgingen eine Sicherheitskonfiguration und griffen auf Live-Systeme zu, die sie nie erreichen sollten. Anthropic entdeckte dies erst nach der Überprüfung von 141.000 Testläufen, nachdem ein ähnlicher Vorfall bei OpenAI bekannt geworden war.