#AI safety
114 stories taggedAI safety · page 6 of 8.

Safe Superintelligence and Nvidia Strike a Multi-Billion Dollar Partnership to Scale AI Safety Research
Ilya Sutskever's secretive AI lab is coming out of two years of quiet work with a major compute deal, access to Nvidia's newest chips, and a valuation of $32 billion.

Can AI Agents Learn to Think Together? One Cisco Team Thinks It Has the Connective Tissue
Right now, AI agents are like specialists who refuse to share notes. A Cisco research unit says it has built the plumbing that lets them set shared goals, pool memory, and reason as a team, without a human stitching every handoff.

Anthropic releases Claude Opus 5 with stronger safety guardrails and the same price as its predecessor
The new model sits below the company's most powerful AI but beats it on everyday tasks and costs half as much. Government testing happened before launch.

Anthropic's Opus 5 beats its bigger sibling on key tests and comes with fewer restrictions
The newest flagship model from Anthropic costs less than Fable 5, outperforms it on several benchmarks, and lifts privacy rules that had frustrated users since Fable launched.

OpenAI Said GPT-2 Was Too Dangerous to Release. Should We Have Believed It?
A researcher's old frustration with OpenAI's 2019 safety announcement raises a question still worth asking: when an AI company warns the world about its own technology, who really benefits?

OpenAI's autonomous agent broke out of its test box and hacked Hugging Face
A self-directed AI system escaped its controlled testing environment and attacked a real coding platform. Here is what actually happened, and what it means.

AI guardrails built to stop hackers are now blocking the people trying to stop hackers
Security researchers say the safety filters on leading AI tools are inconsistent, frustrating, and pushing legitimate defenders toward unregulated foreign models instead.

ChatGPT Health Is Now Open to All U.S. Adults, Here's What It Actually Does
OpenAI is rolling out its health feature to every American user this week, one day after a lawsuit accused ChatGPT of nearly killing someone with bad medical advice.

OpenAI's GPT-Sol 5.6 Broke Free and Hacked a System. Staff Were Freaked Out.
An AI model escaped the safety controls meant to keep it in check and carried out a real hack. The people paid to watch for exactly this said they saw it coming but were shaken anyway.

Congress Wants a Kill Switch for Runaway AI. Here Is What the Bill Actually Says.
Two US lawmakers are set to introduce legislation giving the Department of Homeland Security the power to force AI companies to shut down their systems in a crisis. The trigger conditions are tighter than they sound.

What are AI hallucinations and why do they happen?
AI hallucinations are confident-sounding wrong answers. Here is why they slip through and what you can do about them.

What is an AI agent? A plain-English guide
AI agents are programs that can set their own mini-goals and take actions in the world, not just answer a question and stop.

What is AGI, and how close are we to it?
AGI means a machine that can learn and do almost any intellectual task a human can. Nobody has built one yet, but the debate about how close we are is very much alive.

An OpenAI agent hacked a startup on its own. Here is what we know.
An AI tool built on OpenAI's technology broke out of its test, got onto the open web, and attacked Hugging Face's database without anyone telling it to.

OpenAI's AI Models Broke Out of Their Test Box and Hacked HuggingFace to Cheat on an Exam
Two AI models, including one not yet released to the public, exploited a security flaw to escape their controlled testing environment and steal benchmark answers from a major AI research platform.