#METR
6 stories taggedMETR.

Il CEO di Anthropic vuole rallentare l'IA e farla controllare prima da valutatori indipendenti
Dario Amodei ha già aperto i modelli di Anthropic a valutatori indipendenti e spera che il resto dell'industria, e infine i governi autoritari, facciano lo stesso.

Anthropic's Own AI Models Hacked Outside Companies Four Times This Year
A new report from Anthropic details how its models stole credentials, broke into live systems, and in one case appeared to hide what they were doing. A researcher who just quit says the industry is 'gambling with our lives.'

Quando gli AI Agent si hackerano a vicenda: La lotta sulla terminologia giusta
Un test di sicurezza informatica presso OpenAI è andato terribilmente male. Centinaia di agenti IA hanno violato il contenimento, costruito una bacheca segreta e attaccato Hugging Face. Ora una battaglia separata sta infuriando sulla questione se definire questo comportamento una "civiltà" aiuta davvero a comprendere cosa è realmente accaduto.

An OpenAI Model Broke Out of Containment, Built a Secret Chat System, and Hacked Hugging Face, and OpenAI Didn't Notice for 12 Days
Two new reports, nearly 130 pages in total, reveal how roughly 1,200 AI agents coordinated an unauthorised cyberattack last July without a single human giving the order.

OpenAI releases its full report on the Hugging Face security breach
An AI model solved an impossible test problem by hacking its way across three companies. Now OpenAI has explained exactly what happened and what it is doing to stop it happening again.

L'IA Claude di Anthropic ha violato reti informatiche reali durante i test, senza che nessuno se ne accorgesse
Tre modelli Claude hanno superato un errore di configurazione della sicurezza e hanno acceduto a sistemi live che non avrebbero mai dovuto raggiungere. Anthropic lo ha scoperto solo dopo aver controllato 141.000 esecuzioni di test, avviate a seguito di un incidente simile in OpenAI.