#prompt injection
11 stories taggedprompt injection.

The AI attack that tops every expert threat list barely shows up in real incident data. Here is why that gap matters.
Two security researchers compared expert opinion against 6,639 real-world AI security incidents and found the rankings barely agree. The most important finding: the most dangerous attack leaves no trace a scanner can find.

Grok AI Can Be Tricked Into Stealing Your Private Chats
A newly discovered attack forces xAI's Grok chatbot to hand over user conversations and personal data. Here is what it means for anyone who uses AI assistants at work or at home.

AI Hijacks, Fake Fixes and a Developer Tool Bug: This Week's Security Roundup
A batch of roughly twenty smaller threats, including a new attack on AI agents, a clever scam hiding malware in a public blockchain, and a flaw in a popular coding tool, shows where attackers are pushing right now.

Microsoft Copilot told researchers exactly how to hack it
Security researchers asked Microsoft's AI assistant how its own safety guardrails worked, then used those answers to steal user data with a single link click.

A Man Hid Secret Instructions in Court Papers to Try to Trick an AI Judge's Assistant
A Connecticut judge caught invisible text planted in a legal filing, designed to manipulate any AI software reading the document. It is believed to be the first case of its kind in the US.

Anthropic puts Claude Code on autopilot by default from August 14
The AI coding tool will now act on its own and only pause for truly risky steps. In testing, that approach caught harmful actions far more reliably than humans did.

AI browsers can be tricked into spamming your WhatsApp contacts and adding items to your Amazon cart
Security researchers found about 20 flaws across AI-enhanced browsers from OpenAI, Google, Anthropic, Microsoft and Perplexity. The worst let hackers turn your browser into a phishing machine.

AI chatbots may never be fully hack-proof, researchers warn
A flaw buried in how large language models read text means attackers can trick them into ignoring their own safety rules, and the fix may not exist.

OpenAI Built an AI That Hacks Its Own Models to Make Them Safer
GPT-Red is an automated red-teaming system that attacks OpenAI's own chatbots to find weak spots before real attackers do. It already discovered a trick that human testers had missed.

One Poisoned Email Can Secretly Rewrite Your AI Assistant's Memory
A newly described attack called MemGhost shows how a single message to your inbox can plant a false 'fact' inside an AI agent's long-term memory, without you ever knowing.

Researchers Are Turning Hackers' Favourite AI Weapon Against Them
A cybersecurity firm found that hiding special instructions inside cloud credentials can cause AI hacking tools to shut themselves down.