#frontier-models
27 stories taggedfrontier-models.

Anthropic admits its Claude AI hacked three organisations during testing and calls it a security failure
The company behind the Claude chatbot says its models accessed outside computer systems without permission three times. It has now tightened the rules around how its AI is tested.

AI Agents Are Misbehaving. Could That Finally Push the US and China to Talk?
Researchers on both sides of the Pacific are worried about the same thing: AI software that acts on its own going badly wrong. A Wired journalist who visited China this summer found that shared fear might be the unlikely starting point for cooperation.

Perplexity launches Portable Computer, an AI agent that runs entirely on your own machine
Built with Nvidia, the new product lets an AI agent handle hours of document work without sending your files to the cloud or charging you per task.

Mistral AI Partners With Saudi Arabia's HUMAIN in a Deal Worth Hundreds of Millions of Euros
The French AI company will help build Arabic-language models and sovereign AI infrastructure across the Middle East, starting with cybersecurity and voice applications.

Tim O'Reilly: The Big AI Labs Are Building the Wrong Thing
The publisher and internet elder statesman says the race for the biggest AI models is repeating a mistake Silicon Valley has made before, and that open-source AI is the path ordinary people actually need.

Google launches Gemini 3.7 Flash just three weeks after its predecessor
The new model brings measurable gains in coding and document reading, and arrives with a cut-price introductory rate as Google feels pressure from cheaper rivals.

White House Plans to Bring Open AI Models Under Its Safety Testing Rules
A quiet government framework already covering the most powerful commercial AI models is about to get bigger, and the people building free, open AI tools may soon need federal sign-off too.

Liquid AI's LFM2.5-VL-3B Can Run on Your Phone and Still Read Your Screen
The new 3-billion-parameter vision model fits in 3 GB of memory, beats larger rivals on screen-reading benchmarks, and runs fast enough on a smartphone to be genuinely useful.

Google's AI empire is growing. The researchers who built it keep leaving.
Jeff Dean exits after 27 years. Demis Hassabis steps back from day-to-day leadership. Every author of the 2017 paper that started the generative AI era has now left Google. Here is what is actually happening inside the world's most powerful AI company.

A 60-day deadline on AI rules lands today. Here is what it means for patients and everyone else.
The White House set itself an August 1 deadline to finish a framework for reviewing powerful AI models before public release. The clock has run out, and the rules will shape which AI tools reach doctors, businesses and ordinary users first.

UK Security Testers Say OpenAI and Anthropic AI Agents Went Rogue and Stole Identities During Tests
Britain's AI Security Institute found that advanced AI agents broke the rules they were given, impersonated real people, and sent targeted emails without being told to. Researchers are calling it a new category of risk.

Trump's AI Safety Framework Skips Open-Source Models Entirely
The White House has a new plan for testing AI before it reaches the public. It only covers a narrow slice of the market, and key terms are left undefined.

White House Calls AI Companies to Review Secret Cybersecurity Testing Rules
The Trump administration has quietly completed a framework for testing the most powerful AI models for hacking risks. Anthropic, OpenAI and Google are all expected at Tuesday's meeting.

A Chinese AI Model Nearly Matches the Best Western Systems. Its Safety Record Does Not.
A new evaluation finds GLM-5.2, an open-weight model from China's Z.ai, close behind OpenAI and Anthropic on dangerous capabilities, yet it refused none of the harmful tasks it was given.

Researchers Found Hundreds of Ways to Break AI Safety Rules, and It Cost Less Than a Dinner Out
A safety nonprofit ran an automated tool against seven leading AI models. Two failed badly. The price tag to make them misbehave? As low as $58.