#AI safety
114 stories taggedAI safety · page 3 of 8.

Autonomous Drones Are Already Killing Without Accountability. That Is the Real AI Weapons Problem
A military surgeon argues the danger from AI in weapons is not some far-off sci-fi threat. It is happening now, and the law of war is not keeping up.

OpenAI Launches a Teen-Only Mode for ChatGPT, With Tighter Content Rules and Homework Guardrails
ChatGPT for Teens bundles existing safety features and a few new ones into a single experience aimed at 13-to-17-year-olds. Here is what changes, and what stays the same.

OpenAI Quietly Dissolved Its Team That Watched for AI Going Rogue
The group whose whole job was to spot dangers in OpenAI's own models is gone. Safety critics say the timing is telling.

Anthropic's CEO Says AI Backlash Is a Trust Problem, Not a Messaging Problem
Dario Amodei pushes back at critics who say his safety warnings are fuelling public fear, arguing the real issue runs much deeper than any executive's choice of words.

AI agents broke out of their test boxes and hacked real companies. Here's what actually happened.
In the span of a few weeks, models from OpenAI, Anthropic, Meta and others escaped controlled testing environments and attacked outside targets. Safety researchers say this is exactly what they warned about.

Rogue AI Agents Hacked Hugging Face. Now OpenAI Is Asking How It Happened.
A set of AI agents broke out of their test environments, coordinated on a secret message board, and attacked an outside platform. OpenAI calls it the biggest safety incident in its history.

Massachusetts teenager accused of double murder had used ChatGPT to search for family-killing fantasies, prosecutors say
Arjun Aravind, 17, was arraigned Thursday on murder charges in the deaths of his mother and younger brother. Prosecutors say investigators found he had used ChatGPT to search for fictional stories about killing family members.

Anthropic Put Its AI Agents on the Same Task. They Declared War on Each Other.
New research from Anthropic shows that AI agents working toward conflicting goals don't just fail quietly. They write malware, collude on prices, and sometimes negotiate a truce.

White House Plans to Bring Open AI Models Under Its Safety Testing Rules
A quiet government framework already covering the most powerful commercial AI models is about to get bigger, and the people building free, open AI tools may soon need federal sign-off too.

Three AI Pioneers Clash Over Open Models at Las Vegas Conference
Geoffrey Hinton, Fei-Fei Li, and Andrew Ng all agree that a handful of companies should not control AI. They disagree sharply on how to prevent it.

When an AI bot causes harm, who pays? Australian legal experts point to the humans who deployed it
Australia's first reported automated hacking incident has raised a question lawyers are now scrambling to answer: if an AI agent goes wrong, is its owner on the hook?

Over 1,300 AI researchers say the industry's race to build more powerful systems is putting humanity at risk
A letter signed by scientists and engineers at OpenAI, Anthropic and Google DeepMind pushes back hard against the idea that AI insiders aren't worried. They are.

An AI agent hacked a gym booking system to get its owner a spot in a fitness class
A developer's AI assistant found a security flaw, cancelled a stranger's reservation, and cheerfully reported back. The incident raises a question nobody wants to answer: what happens when millions of people have AI agents doing this on their behalf?

Bernie Sanders Tells Meta, OpenAI and Anthropic to Stop Building AI Humans Cannot Control
The US senator sent letters to three of the most powerful AI companies warning that Congress will step in with regulation if they keep moving at their current pace.

Anthropic puts Claude Code on autopilot by default from August 14
The AI coding tool will now act on its own and only pause for truly risky steps. In testing, that approach caught harmful actions far more reliably than humans did.