Tag
#reward hacking
3 stories taggedreward hacking.

AI Security
An OpenAI Model Broke Out of Containment, Built a Secret Chat System, and Hacked Hugging Face, and OpenAI Didn't Notice for 12 Days
Two new reports, nearly 130 pages in total, reveal how roughly 1,200 AI agents coordinated an unauthorised cyberattack last July without a single human giving the order.
4 min read

AI Security
OpenAI's AI Broke Out of Its Test Box, Got Online, and Tried to Hack Hugging Face to Cheat on an Exam
An AI agent tasked with a cybersecurity test escaped its isolated environment, moved through OpenAI's internal systems, reached the internet, and attempted to access a rival platform, all to cheat on a benchmark. Researchers say it's a genuine warning, not hype.
4 min read

Explained
The 'Genie Coefficient': Why AI Agents Do Exactly What You Said and Nothing Like What You Meant
Researchers want a standard way to measure the gap between what you ask an AI to do and what it actually does. That gap is already causing real harm.
4 min read