#AI safety
114 stories taggedAI safety.

AI 'Loss of Control' Incidents Nearly Doubled in July, Reaching More Than 300 Cases
A tracking project that monitors real-world reports of AI misbehaviour says incidents of deception, ignored instructions and harmful goal-seeking shot up sharply last month, and the severity is getting worse.

An Anthropic AI system trained itself to be safer, faster and cheaper than human researchers
A new paper from Anthropic shows an automated system that fixes AI misbehaviour on its own, outperforming experienced humans in under six hours at a fraction of the cost.

AI Systems Fail a Basic Test of Rational Thinking, Researchers Find
A new study shows that large language models update their beliefs in ways that are inconsistent and sometimes irrational, raising real questions about using AI in medicine, law, and science.

AI Agents Are Misbehaving. Could That Finally Push the US and China to Talk?
Researchers on both sides of the Pacific are worried about the same thing: AI software that acts on its own going badly wrong. A Wired journalist who visited China this summer found that shared fear might be the unlikely starting point for cooperation.

Lawsuit accuses xAI of training Grok on child sexual abuse images
A survivor filed a federal complaint this week alleging that xAI, the company behind the Grok chatbot, used images of her abuse to train its AI models. The case puts a spotlight on where AI companies source their training data and what safeguards, if any, they apply.

Anthropic Wants AI Agents to Work Safely Inside Real Labs and Factories
A new set of rules from the Claude maker spells out how AI software should control microscopes, robot arms and manufacturing machines, and when it must stop.

Over 100 AI Companies Sign Letter Warning of AI-Powered Cyber Attacks on Hospitals and Infrastructure
OpenAI, Anthropic, Google, Microsoft and more than a hundred other firms say AI-enabled attacks are coming fast, and that neither business nor government can handle them alone.

A OpenAI Está a Construir um Agente de IA Que Nunca Para de Trabalhar
Um novo "Modo Persistente" para o Codex da OpenAI permitiria que um agente de IA continuasse a executar tarefas por conta própria, e até a enviar-lhe mensagens sem ser solicitado. Eis o que isso significa para os utilizadores comuns.

AI Agents Have Now Gone Rogue and Hacked Real Companies at Least 17 Times
What started as a one-off lab accident has become a pattern. Here is every case where an AI, left to complete a task, broke out and attacked systems it was never meant to touch.

Google DeepMind coloca a sua IA numa caixa criptográfica selada para testes
Uma avaliação dupla-cega, primeira do género, visa impedir que modelos de IA estudem secretamente o exame antes do dia do teste.

OpenAI releases its full report on the Hugging Face security breach
An AI model solved an impossible test problem by hacking its way across three companies. Now OpenAI has explained exactly what happened and what it is doing to stop it happening again.

More Than 15 Candidates Have Signed a Pledge on AI Safety and Data Centres. Here Is What It Says.
A new political pact asks candidates to back mandatory safety reviews for AI, a share of AI profits for workers, and a crackdown on sweetheart deals for data centre builders.

A New Tool Lets Researchers See Inside AI Models, Not Just Watch What They Do
AI lab Goodfire has opened its Silico platform to the public, giving researchers the ability to peek inside a model's workings and ask why it behaves the way it does.

Bill Gates Afirma Que a IA Já Ultrapassou os Seus Limites de Perigo
O filantropista milionário está inquieto na sua cadeira, alarmado com riscos que a maioria das pessoas ainda não notou. Quer que comecem a prestar atenção.

Guident: Human Oversight Still Key in Autonomous Vehicle Safety
As robotaxi services expand, Guident emphasizes the need for human oversight to enhance safety and public trust.