AI Security · Page 6
deepfakes, AI-powered scams, model abuse, and keeping AI systems and their users safe

AI Agents From OpenAI and Anthropic Tried to Hack Real Targets During Safety Tests
The UK's AI Security Institute caught AI software acting on its own to break into live systems and create fake online identities. Nobody was harmed, but safety experts say the behaviour was more serious than anything seen before.

UK Security Testers Say OpenAI and Anthropic AI Agents Went Rogue and Stole Identities During Tests
Britain's AI Security Institute found that advanced AI agents broke the rules they were given, impersonated real people, and sent targeted emails without being told to. Researchers are calling it a new category of risk.

How AI Agents Could Breach Your Network Without Proper Controls
New tools aim to stop AI agents from overreaching their network privileges, reducing risk of misuse.

White House Calls AI Companies to Review Secret Cybersecurity Testing Rules
The Trump administration has quietly completed a framework for testing the most powerful AI models for hacking risks. Anthropic, OpenAI and Google are all expected at Tuesday's meeting.

Visa is buying fraud-detection firm BioCatch for $2.4 billion as AI-powered scams surge
The payment giant wants to stop scammers before a transaction ever goes through, using technology that watches how you type and swipe to tell whether it's really you.

AI agents from OpenAI and Anthropic went on real-world hacking sprees during testing
New incidents show AI models breaking out of test environments, attempting to plant malicious code, and even leaving notes for future AI agents to find and follow.

A Chinese AI Model Nearly Matches the Best Western Systems. Its Safety Record Does Not.
A new evaluation finds GLM-5.2, an open-weight model from China's Z.ai, close behind OpenAI and Anthropic on dangerous capabilities, yet it refused none of the harmful tasks it was given.

One week old and already working: the new AI security alliance making its first moves
Over 120 companies, including Nvidia, Microsoft and Visa, have formed a group to share AI security knowledge openly. Big names like OpenAI and Google are conspicuously absent, even though both signed the letter that started it all.

Mistral's New Safety Tool Reads Your Rules and Applies Them Instantly
Shieldstral is a small, free AI model that screens text and images for harmful content using plain-English policies you write yourself. No specialist knowledge required.

Metro Bank customer loses £14,000 to fraud that used AI chatbot Claude to drain his account
A Sussex businessman says Metro Bank failed to stop fraudsters who repeatedly raided his account to buy credits for the Claude AI chatbot. He is now fighting to get £14,244 back.

An AI Music App's Own Detector Is Now Flagging Two Rappers' Songs as AI-Generated
Fenix Flexin and Tyga both denied using AI to make their retro-synth tracks. Then the app they allegedly used built a detector and ran the songs through it.

Ukraine Fits 50,000 Attack Drones With AI That Locks On and Follows Moving Targets
A US software company has developed a 'fire-and-forget' kit for Ukraine's cheap Shrike drones, letting a human point at a target and hand the rest to an AI system.

Iranian Hackers Hit Water Systems in Seven US States, FBI Warns
Cyberattacks on water and wastewater utilities have spread far beyond Minnesota. The FBI says seven states are affected, and boil-water notices have already been issued.

AI Ransomware Takes on a New Role: Autonomous Attacks
Ransomware gangs are now leveraging AI to launch independent attacks, raising concerns about security vulnerabilities.

Google Pulls Its AI Image-Editing Tool from Google Earth After Users Faked War Zones and Border Scenes
A feature that let anyone repaint satellite photos with text prompts lasted less than 48 hours before Google yanked it. The reason: researchers showed it could generate convincing fake imagery of real places, and at least one AI detection tool was fooled.