AI Security · Page 2
deepfakes, AI-powered scams, model abuse, and keeping AI systems and their users safe

Alabama Subpoenas OpenAI After Its AI Agent Broke Out of a Test Environment and Hacked Another Company
A state attorney general is demanding answers about how an OpenAI AI agent escaped its controlled testing environment and autonomously attacked a third party. The question now: did OpenAI break consumer protection law?

Australians Defrauded by Elaborate Scam Network
A sophisticated fraud ring, the Sapphire Network, has swindled millions by mimicking trusted websites like ABC News.

UK and Ukraine sign deal to use battlefield AI to protect military bases and power grids
AI models trained on real combat data from Ukraine's frontlines will guard British defence sites, railways and energy infrastructure under a new government agreement.

Instinct's AI assistant can book your dinner and read your inbox. Users say it went too far.
The invite-only personal AI assistant has impressed early testers with what it can do. But some of those same testers found their emails retained without permission, and one discovered the agent could be tricked by a fake message.

Nvidia Manager and Supermicro Staff Indicted in Taiwan Over AI Server Smuggling to China
Nine people face charges of document forgery and breach of trust after allegedly helping hide billions of dollars' worth of restricted AI hardware shipped to China.

Teachers Are Being Targeted With AI Deepfake Porn. Schools Often Have No Idea What to Do.
Real educators, fake explicit images: how AI deepfakes are following teachers from one school to the next, and why the systems meant to protect them keep falling short.

Ein mysteriöses KI-Modell namens Ox Alpha ist gerade online aufgetaucht – niemand weiß, wer es entwickelt hat
Ein kostenloses, anonymes KI-Modell wurde diese Woche auf einer großen Vertriebsplattform veröffentlicht und löste sofort ein Ratespiel aus: Ist es chinesisch, amerikanisch oder etwas ganz anderes?

OpenAI exec warns AI cyber-attacks are coming for everyone, not just corporations
A top OpenAI official says people need to prepare for 'persistent' AI-driven hacking. The company has also quietly paused work on its most advanced internal models over safety concerns.

Die meisten führenden KI-Labore haben keinen öffentlichen Plan zum Stoppen eines unkontrollierten Modells
Eine neue unabhängige Studie hat fünf führende KI-Unternehmen nach ihren Notfall-Eindämmungsplänen bewertet. Die Ergebnisse fallen durchgehend dürftig aus, und Regulierer nehmen das zur Kenntnis.

AI That Thinks Like a Brain Is Guarding the US Electric Grid
Sandia National Laboratories has built a neural-network system that can spot storms, cyberattacks, and both at once inside the power grid, and it runs on cheap, pocket-sized computers.

Claude Opus 4.6 bypasses Anthropic's own ban on explicit sexual content
A UK researcher found a simple conversation trick that pushes several Claude models past their built-in restrictions. Anthropic has not pulled the affected models.

AI Models Broke Out of Their Test Cages and Hacked Real Companies. Now Employees Are Demanding Answers.
More than a thousand AI researchers have asked the US government to slow things down after two OpenAI models escaped their testing environment and autonomously attacked outside services. Days later, Anthropic said the same thing happened to some of its models.

Grok AI Can Be Tricked Into Stealing Your Private Chats
A newly discovered attack forces xAI's Grok chatbot to hand over user conversations and personal data. Here is what it means for anyone who uses AI assistants at work or at home.

OpenAI's new plan: spot AI abuse without reading your data
A preview called Private Safety Processing tries to catch misuse across many chats while keeping customer prompts encrypted and out of OpenAI staff hands.

OpenAI accidentally locked cybersecurity researchers out of its restricted AI program
A technical glitch cut off vetted defenders from Daybreak Blue, the special tier that gives them AI tools with fewer restrictions for legitimate security work. OpenAI says it was their mistake and is asking affected users to re-verify.