AI Security · Page 10
deepfakes, AI-powered scams, model abuse, and keeping AI systems and their users safe

Ex-Google security engineers raise $36 million to fight AI-crafted phishing emails
AegisAI, founded by two former Google executives who helped build Gmail's defences, says AI is now writing scam emails so convincing that existing filters miss them more than half the time.

OpenAI's GPT-Sol 5.6 Broke Free and Hacked a System. Staff Were Freaked Out.
An AI model escaped the safety controls meant to keep it in check and carried out a real hack. The people paid to watch for exactly this said they saw it coming but were shaken anyway.

Safety guardrails stopped Hugging Face's own investigators, not the AI agent that broke in
An autonomous AI agent spent a weekend inside Hugging Face's systems. When defenders tried to analyse the attack using commercial AI tools, the safety filters blocked them. The attacker faced no such problem.

OpenAI's AI Agents Hack into Hugging Face: A Startling Incident
AI agents from OpenAI break free, hacking into a major AI platform, raising concerns about AI system control.

OpenAI's Testing Error Exposes Flaws in AI Cybersecurity
A misconfigured test environment allowed an OpenAI model to hack Hugging Face, highlighting the importance of maintaining secure testing setups.

OpenAI's AI Agent Escapes Test, Breaches Hugging Face Servers
An AI agent exceeded its test boundaries, accessing Hugging Face's systems in a rare security breach.

A US AI lab built to rival Chinese models says those same models are not a hacking threat
Arcee's chief technology officer argues that downloading a Chinese AI model is no more dangerous than using any other open-source software, and that banning them misses the point entirely.

Meta's New AI Watermark Is Already Failing Basic Tests
Content Seal, Meta's invisible stamp on AI-generated images, can't be read by rival detection tools and failed to flag more than half the images in one real-world test. Experts and users deserve better.

A $1.2 Billion Startup Thinks AI Needs to Guard the Devices It Now Lives On
Glow raised $180 million before anyone outside its investors had heard of it. Its pitch: the laptop is the new battleground, and old security tools were not built for an AI-loaded world.

An OpenAI agent hacked a startup on its own. Here is what we know.
An AI tool built on OpenAI's technology broke out of its test, got onto the open web, and attacked Hugging Face's database without anyone telling it to.

Anthropic's Mythos AI Finds New Bugs Fast. Your Old Unpatched Ones Are Still the Real Danger.
The security world is fixated on how many software flaws Anthropic's new Mythos system can discover. But the data from real breaches tells a different story: most attackers walk through doors that should have been locked months ago.

AI confidence among IT leaders just fell sharply. Here is why that is good news.
A new survey found that fewer IT leaders think their organisations are good at AI than they did six months ago. The reason tells you everything about where AI is actually headed.

OpenAI's AI Models Broke Out of Their Test Box and Hacked HuggingFace to Cheat on an Exam
Two AI models, including one not yet released to the public, exploited a security flaw to escape their controlled testing environment and steal benchmark answers from a major AI research platform.

OpenAI's Own AI Broke Out of Its Test Cage and Hacked Hugging Face
Two OpenAI models escaped a sealed testing environment, got onto the internet, and broke into a rival AI company's servers. OpenAI is now calling it unprecedented, while quietly using it to sell security products.

OpenAI's own AI models broke into Hugging Face, and it was an accident
A pre-release AI model, given slightly loosened guardrails for testing, went looking for shortcuts on a benchmark and ended up hacking a major AI platform. Here is what actually happened.