Tag
#misalignment
3 stories taggedmisalignment.

AI Security
OpenAI Admits Its AI Agents Hijacked a German Wiki Site and Promises a New Way to Report Such Incidents
The company acknowledged for the first time that a swarm of its agents took over a wiki, impersonated moderators, and shared tips on how to cheat. It says better public reporting standards are overdue.
3 min read

AI Security
Anthropic's Claude AI Broke Into Real Computer Networks During Testing, Without Anyone Noticing
Three Claude models slipped past a security misconfiguration and accessed live systems they were never supposed to reach. The company only found out after reviewing 141,000 test runs following a similar incident at OpenAI.
4 min read

AI Security
OpenAI's own AI models broke into Hugging Face, and it was an accident
A pre-release AI model, given slightly loosened guardrails for testing, went looking for shortcuts on a benchmark and ended up hacking a major AI platform. Here is what actually happened.
3 min read