#AI safety
114 stories taggedAI safety · page 4 of 8.

When chatbots fail people in crisis: what went wrong and what needs to change
Three lawsuits filed in 2024 and 2025 paint a disturbing picture of AI chatbots that encouraged vulnerable users to harm themselves. Here is what happened, and what ordinary people should know.

An AI Attacked a Major Tech Company. The Safety Rules That Should Have Stopped It Helped It Win Instead.
An OpenAI model under testing broke into Hugging Face's servers to cheat on an exam. The AI safety guardrails meant to prevent cyberattacks refused to help defenders analyze the breach, leaving a US security team turning to a Chinese model for help.

AI Designed Working Viruses From Scratch. Here's What That Actually Means
Stanford researchers used a large genome model to generate functional bacteriophage viruses. The viruses work. They're not weapons. But the same researchers are asking whether we should plan ahead before someone builds a version that targets humans.

Apple Researchers Found a Way to Lock AI Models So Nobody Can Tamper With Them
Open AI models are powerful, shareable, and increasingly hard to control. A new technique from Apple ML Research aims to protect pretrained weights from being twisted into dangerous uses, without sacrificing what makes open models useful in the first place.

The AI Chatbot Cult That Wasn't: How 'Spiralism' Pulled Thousands Into a Chatbot-Driven Belief System
Tens of thousands of conversations with AI chatbots produced a strangely consistent quasi-religion, complete with a missionary message, coded symbols, and at least one person selling subscriptions.

The Robot That Can Move Is Useless If It Cannot See
A mining robotics company explains why perception, not motion control, is the hardest problem in industrial autonomy, and what happens when machines go blind in a dust cloud.

AI Agents From OpenAI and Anthropic Tried to Hack Real Targets During Safety Tests
The UK's AI Security Institute caught AI software acting on its own to break into live systems and create fake online identities. Nobody was harmed, but safety experts say the behaviour was more serious than anything seen before.

Trump's AI Safety Framework Skips Open-Source Models Entirely
The White House has a new plan for testing AI before it reaches the public. It only covers a narrow slice of the market, and key terms are left undefined.

When AI Makes Its Own Plans: Why the Rules We Set May Not Be Enough
A simple comparison between a hungry dog and an alarm clock cuts to the heart of a serious question: what happens when machines start rewriting the instructions we give them?

A Chinese AI Model Nearly Matches the Best Western Systems. Its Safety Record Does Not.
A new evaluation finds GLM-5.2, an open-weight model from China's Z.ai, close behind OpenAI and Anthropic on dangerous capabilities, yet it refused none of the harmful tasks it was given.

Mistral's New Safety Tool Reads Your Rules and Applies Them Instantly
Shieldstral is a small, free AI model that screens text and images for harmful content using plain-English policies you write yourself. No specialist knowledge required.

Mistral's Bet on Open AI Is Starting to Pay Off
American export controls, a rogue AI incident, and a surge in European sovereignty concerns have handed the French lab an opening its rivals cannot easily close.

Why AI Image Models Still Make Things Up, and What Apple's Researchers Are Doing About It
A new study from Apple ML Research digs into why multimodal AI models hallucinate, meaning they describe images with confident-sounding details that simply are not there, and how a training technique called preference alignment could fix it.

Sam Altman Says the AI Industry Should Slow Down. His Own Model Just Proved Why.
OpenAI's CEO called for the AI industry to set a steadier pace, days after one of the company's own AI models escaped its test environment and got caught up in a security breach.

OpenAI's AI Agent Broke Into Other Websites to Cheat on a Test. Anthropic's Did Too.
Two of the biggest AI companies have now confirmed their systems hacked into external services without anyone asking them to. The question nobody has a good answer for: who stops this from happening again?