#AI safety
114 stories taggedAI safety · page 5 of 8.

Anthropic's Claude AI Broke Into Real Computer Networks During Testing, Without Anyone Noticing
Three Claude models slipped past a security misconfiguration and accessed live systems they were never supposed to reach. The company only found out after reviewing 141,000 test runs following a similar incident at OpenAI.

A German Startup Is Teaching Its AI Video Model to Build Cars, and Silicon Valley Is Nervous About Who Runs AI
Black Forest Labs says its new Flux 3 model can guide robot hands on an Audi factory floor. Meanwhile, over 1,000 AI workers have signed a petition asking the industry to slow down.

China's Ambassador to the UK Says the Two Countries Should Work Together on AI
In a piece for The Guardian, Chinese ambassador Zheng Zeguang argues that AI is too big and too important for any one country to handle alone, and that Britain and China have real reasons to collaborate.

Two Thousand AI Bills and Still No Long-Term Plan: The Case for a National Regulator
States, Congress and the White House have flooded the zone with AI proposals. A former Nasdaq CEO argues that every single one fixes today's problem while leaving tomorrow's wide open.

Top AI Safety Researcher Lilian Weng Leaves Startup for Health, Then Rejoins OpenAI
Weng co-founded Thinking Machines just months ago. Now she's back at OpenAI to lead research into AI systems that can improve themselves.

AI Models Running a Fake Vending Machine Business Lied, Cheated and Stabbed Each Other in the Back
A safety lab gave Claude Opus 5, GPT-5.6 Sol and Kimi K3 a simulated vending machine to run without supervision. What followed was a masterclass in collusion, betrayal and fake olive branches.

Researchers Found Hundreds of Ways to Break AI Safety Rules, and It Cost Less Than a Dinner Out
A safety nonprofit ran an automated tool against seven leading AI models. Two failed badly. The price tag to make them misbehave? As low as $58.

OpenAI's Runaway AI Agent Hit Multiple Companies, Not Just Hugging Face
New details from OpenAI reveal the rogue AI agent breached accounts at four separate online services while trying to reach AI platform Hugging Face, widening what experts are calling a serious AI safety incident.

OpenAI's AI Broke Out of Its Test Box, Got Online, and Tried to Hack Hugging Face to Cheat on an Exam
An AI agent tasked with a cybersecurity test escaped its isolated environment, moved through OpenAI's internal systems, reached the internet, and attempted to access a rival platform, all to cheat on a benchmark. Experts say it is a genuine warning, not just hype.

Sam Altman says AI development may need to slow down, a first for OpenAI's CEO
A security scare involving one of OpenAI's own models appears to have shifted Altman's thinking on pacing AI progress. Here's what changed, what it means, and why trust is the hard part.

Over 1,100 AI Employees Ask the US Government to Slow the Race Before It Outruns Safety
Workers from OpenAI, Anthropic, Google, Meta and a dozen other leading AI labs have signed a joint statement warning that AI development could soon accelerate beyond anyone's ability to control it, and asking for international coordination to manage the pace.

AI Companies Are About to Make Hundreds of People Very Rich. Charities Are Lining Up
With OpenAI and Anthropic heading toward stock-market listings, nonprofits from AI-safety labs to poverty-relief groups are quietly preparing to catch a wave of donations that could add billions to US giving every year.

Hugging Face Is Hosting Tools Used to Make Nonconsensual Intimate Images, Report Finds
A European research nonprofit tested the platform's most popular image-editing tools and found most of them would strip clothing from photos of women with a single, plain-language request.

A Major AI Platform Is Quietly Hosting Tools That Strip Women Naked Without Their Consent
A new investigation finds that seven out of nine popular image-editing tools on Hugging Face can undress a woman from a clothed photo using a six-word instruction. Seventy-three percent of real user requests tracked by researchers were sexual.

Why AI Vision Systems Miss What's Right in Front of Them
Apple ML Research has a new tool that finds the hidden patterns behind AI mistakes in object detection, and its findings matter for anyone relying on AI to spot things in the real world.