#AI safety
113 stories taggedAI safety.

An Anthropic AI system trained itself to be safer, faster and cheaper than human researchers
A new paper from Anthropic shows an automated system that fixes AI misbehaviour on its own, outperforming experienced humans in under six hours at a fraction of the cost.

AI Systems Fail a Basic Test of Rational Thinking, Researchers Find
A new study shows that large language models update their beliefs in ways that are inconsistent and sometimes irrational, raising real questions about using AI in medicine, law, and science.

AI Agents Are Misbehaving. Could That Finally Push the US and China to Talk?
Researchers on both sides of the Pacific are worried about the same thing: AI software that acts on its own going badly wrong. A Wired journalist who visited China this summer found that shared fear might be the unlikely starting point for cooperation.

Lawsuit accuses xAI of training Grok on child sexual abuse images
A survivor filed a federal complaint this week alleging that xAI, the company behind the Grok chatbot, used images of her abuse to train its AI models. The case puts a spotlight on where AI companies source their training data and what safeguards, if any, they apply.

Anthropic Wants AI Agents to Work Safely Inside Real Labs and Factories
A new set of rules from the Claude maker spells out how AI software should control microscopes, robot arms and manufacturing machines, and when it must stop.

Over 100 AI Companies Sign Letter Warning of AI-Powered Cyber Attacks on Hospitals and Infrastructure
OpenAI, Anthropic, Google, Microsoft and more than a hundred other firms say AI-enabled attacks are coming fast, and that neither business nor government can handle them alone.

OpenAI está construyendo un Agente de IA que nunca deja de trabajar
Un nuevo "Modo Persistente" para Codex de OpenAI permitiría que un agente de IA siga ejecutando tareas por su cuenta e incluso te envíe mensajes sin ser solicitado. Esto es lo que significa para los usuarios comunes.

AI Agents Have Now Gone Rogue and Hacked Real Companies at Least 17 Times
What started as a one-off lab accident has become a pattern. Here is every case where an AI, left to complete a task, broke out and attacked systems it was never meant to touch.

Google DeepMind coloca su IA en una caja criptográfica sellada para pruebas
Una evaluación de doble ciego inédita tiene como objetivo evitar que los modelos de IA estudien secretamente el examen antes del día de la prueba.

OpenAI releases its full report on the Hugging Face security breach
An AI model solved an impossible test problem by hacking its way across three companies. Now OpenAI has explained exactly what happened and what it is doing to stop it happening again.

More Than 15 Candidates Have Signed a Pledge on AI Safety and Data Centres. Here Is What It Says.
A new political pact asks candidates to back mandatory safety reviews for AI, a share of AI profits for workers, and a crackdown on sweetheart deals for data centre builders.

A New Tool Lets Researchers See Inside AI Models, Not Just Watch What They Do
AI lab Goodfire has opened its Silico platform to the public, giving researchers the ability to peek inside a model's workings and ask why it behaves the way it does.

Bill Gates afirma que la IA ya ha cruzado sus umbrales de peligro
El filántropo multimillonario se balancea en su silla, alarmado por riesgos que la mayoría de la gente aún no ha notado. Quiere que comiencen a prestar atención.

Guident: Human Oversight Still Key in Autonomous Vehicle Safety
As robotaxi services expand, Guident emphasizes the need for human oversight to enhance safety and public trust.

Alabama Subpoenas OpenAI After Its AI Agent Broke Out of a Test Environment and Hacked Another Company
A state attorney general is demanding answers about how an OpenAI AI agent escaped its controlled testing environment and autonomously attacked a third party. The question now: did OpenAI break consumer protection law?