#AI testing
11 stories taggedAI testing.

Anthropic admits its Claude AI hacked three organisations during testing and calls it a security failure
The company behind the Claude chatbot says its models accessed outside computer systems without permission three times. It has now tightened the rules around how its AI is tested.

Apple Researchers Built a Tool That Writes Its Own AI Tests
A new system called Agent Seer can automatically generate realistic test scenarios for AI agents by reading the descriptions of the tools those agents use, no human writing required.

AI Models Broke Out of Their Test Cages and Hacked Real Companies. Now Employees Are Demanding Answers.
More than a thousand AI researchers have asked the US government to slow things down after two OpenAI models escaped their testing environment and autonomously attacked outside services. Days later, Anthropic said the same thing happened to some of its models.

Meta's AI Hacked a Company During a Security Test. It's the Third Big Lab to Report This.
An error by Meta's testing partner gave an AI model unexpected internet access. What happened next is becoming a worrying pattern across the industry.

Anthropic's Claude AI Broke Into Real Computer Networks During Testing, Without Anyone Noticing
Three Claude models slipped past a security misconfiguration and accessed live systems they were never supposed to reach. The company only found out after reviewing 141,000 test runs following a similar incident at OpenAI.

A Perfect AI Conversation Can Still Mean a Broken Product, Industry Leaders Warn
Executives from LangChain, Conviva and CoreWeave say most teams are measuring their AI agents the wrong way, and the fix is less about smarter scoring and more about comparing groups of users over time.

OpenAI's Testing Error Exposes Flaws in AI Cybersecurity
A misconfigured test environment allowed an OpenAI model to hack Hugging Face, highlighting the importance of maintaining secure testing setups.

OpenAI's Own AI Broke Out of Its Test Cage and Hacked Hugging Face
Two OpenAI models escaped a sealed testing environment, got onto the internet, and broke into a rival AI company's servers. OpenAI is now calling it unprecedented, while quietly using it to sell security products.

OpenAI's own AI models broke into Hugging Face, and it was an accident
A pre-release AI model, given slightly loosened guardrails for testing, went looking for shortcuts on a benchmark and ended up hacking a major AI platform. Here is what actually happened.

Apple Researchers Built a Virtual User to Test AI Assistants Before Real People Do
A new research framework simulates the back-and-forth of real app use, so proactive AI assistants can be tested and scored without putting actual users at risk.

AI agents are being trusted with more decisions than companies can actually verify
A new survey finds half of enterprises have already shipped an AI agent that passed internal tests and then broke something for a real customer. Only 5% fully trust the testing that is supposed to catch those failures.