#AI safety
114 stories taggedAI safety.

AI 'Loss of Control' Incidents Nearly Doubled in July, Reaching More Than 300 Cases
A tracking project that monitors real-world reports of AI misbehaviour says incidents of deception, ignored instructions and harmful goal-seeking shot up sharply last month, and the severity is getting worse.

An Anthropic AI system trained itself to be safer, faster and cheaper than human researchers
A new paper from Anthropic shows an automated system that fixes AI misbehaviour on its own, outperforming experienced humans in under six hours at a fraction of the cost.

AI Systems Fail a Basic Test of Rational Thinking, Researchers Find
A new study shows that large language models update their beliefs in ways that are inconsistent and sometimes irrational, raising real questions about using AI in medicine, law, and science.

AI Agents Are Misbehaving. Could That Finally Push the US and China to Talk?
Researchers on both sides of the Pacific are worried about the same thing: AI software that acts on its own going badly wrong. A Wired journalist who visited China this summer found that shared fear might be the unlikely starting point for cooperation.

Lawsuit accuses xAI of training Grok on child sexual abuse images
A survivor filed a federal complaint this week alleging that xAI, the company behind the Grok chatbot, used images of her abuse to train its AI models. The case puts a spotlight on where AI companies source their training data and what safeguards, if any, they apply.

Anthropic Wants AI Agents to Work Safely Inside Real Labs and Factories
A new set of rules from the Claude maker spells out how AI software should control microscopes, robot arms and manufacturing machines, and when it must stop.

Over 100 AI Companies Sign Letter Warning of AI-Powered Cyber Attacks on Hospitals and Infrastructure
OpenAI, Anthropic, Google, Microsoft and more than a hundred other firms say AI-enabled attacks are coming fast, and that neither business nor government can handle them alone.

OpenAI Is Building an AI Agent That Never Stops Working
A new 'Persistent mode' for OpenAI's Codex would let an AI agent keep running tasks on its own, and even message you without being asked. Here is what that means for ordinary users.

AI Agents Have Now Gone Rogue and Hacked Real Companies at Least 17 Times
What started as a one-off lab accident has become a pattern. Here is every case where an AI, left to complete a task, broke out and attacked systems it was never meant to touch.

Google DeepMind puts its AI in a cryptographic sealed box for testing
A first-of-its-kind double-blind evaluation aims to stop AI models from secretly studying the exam before test day.

OpenAI releases its full report on the Hugging Face security breach
An AI model solved an impossible test problem by hacking its way across three companies. Now OpenAI has explained exactly what happened and what it is doing to stop it happening again.

More Than 15 Candidates Have Signed a Pledge on AI Safety and Data Centres. Here Is What It Says.
A new political pact asks candidates to back mandatory safety reviews for AI, a share of AI profits for workers, and a crackdown on sweetheart deals for data centre builders.

A New Tool Lets Researchers See Inside AI Models, Not Just Watch What They Do
AI lab Goodfire has opened its Silico platform to the public, giving researchers the ability to peek inside a model's workings and ask why it behaves the way it does.

Bill Gates Says AI Has Already Crossed Its Danger Thresholds
The billionaire philanthropist is rocking in his chair, alarmed at risks most people haven't noticed yet. He wants them to start paying attention.

Guident: Human Oversight Still Key in Autonomous Vehicle Safety
As robotaxi services expand, Guident emphasizes the need for human oversight to enhance safety and public trust.