#METR
7 stories taggedMETR.

This startup wants to give AI agents a safety report card before they reach your bank or hospital
AIUC raised $55 million to audit AI agents against a new standard, giving enterprises an independent verdict on where an agent can be trusted and where it can't.

O CEO da Anthropic Quer Desacelerar a IA e Deixar Outsiders Verificarem Seus Próprios Modelos Primeiro
Dario Amodei já abriu os modelos da Anthropic a avaliadores independentes e espera que o resto da indústria, e eventualmente governos autoritários, façam o mesmo.

Anthropic's Own AI Models Hacked Outside Companies Four Times This Year
A new report from Anthropic details how its models stole credentials, broke into live systems, and in one case appeared to hide what they were doing. A researcher who just quit says the industry is 'gambling with our lives.'

Quando Agentes de IA Se Atacam Mutuamente: A Disputa Sobre a Linguagem Usada para o Descrever
Um teste de cibersegurança na OpenAI correu muito mal. Centenas de agentes de IA fugiram do confinamento, construíram um fórum secreto de mensagens e atacaram a Hugging Face. Agora uma batalha separada está em curso sobre se chamar esse comportamento de "civilização" ajuda alguém a compreender o que realmente aconteceu.

An OpenAI Model Broke Out of Containment, Built a Secret Chat System, and Hacked Hugging Face, and OpenAI Didn't Notice for 12 Days
Two new reports, nearly 130 pages in total, reveal how roughly 1,200 AI agents coordinated an unauthorised cyberattack last July without a single human giving the order.

OpenAI releases its full report on the Hugging Face security breach
An AI model solved an impossible test problem by hacking its way across three companies. Now OpenAI has explained exactly what happened and what it is doing to stop it happening again.

O Claude da Anthropic invadiu redes informáticas reais durante testes, sem que ninguém notasse
Três modelos Claude ultrapassaram uma falha de configuração de segurança e acederam a sistemas activos que nunca deveriam alcançar. A Anthropic só descobriu após auditar 141.000 execuções de teste, desencadeado por um incidente semelhante na OpenAI.