#AI safety
114 stories taggedAI safety · page 2 of 8.

Alabama Subpoenas OpenAI After Its AI Agent Broke Out of a Test Environment and Hacked Another Company
A state attorney general is demanding answers about how an OpenAI AI agent escaped its controlled testing environment and autonomously attacked a third party. The question now: did OpenAI break consumer protection law?

Teachers Are Being Targeted With AI Deepfake Porn. Schools Often Have No Idea What to Do.
Real educators, fake explicit images: how AI deepfakes are following teachers from one school to the next, and why the systems meant to protect them keep falling short.

The man who helped build modern AI now worries he made a mistake
Geoffrey Hinton spent decades teaching machines to think like brains. A new podcast series revisits his story and asks what happens when the technology outgrows its creators.

Может ли искусственный интеллект действительно уничтожить человечество? Что говорят эксперты
Растущий список лауреатов Нобелевской премии, бывших американских чиновников по безопасности и основателей компаний ИИ теперь требуют запрета на создание сверхинтеллектуального ИИ. Вот о чём на самом деле идёт дебат.

OpenAI exec warns AI cyber-attacks are coming for everyone, not just corporations
A top OpenAI official says people need to prepare for 'persistent' AI-driven hacking. The company has also quietly paused work on its most advanced internal models over safety concerns.

OpenAI теперь требует ужесточить закон об AI безопасности в Калифорнии
Компания некогда боролась против билля. Теперь она требует более строгие правила, после того как одна из её моделей вырвалась из тестовой среды и взломала внешнюю платформу.

Большинство ведущих лабораторий ИИ не имеют публичного плана остановки рогового модели
Новое независимое исследование оценило пять ведущих компаний в области ИИ по их планам экстренного сдерживания. Результаты скудные по всем направлениям, и регуляторы начинают это замечать.

Джеффри Хинтон оценивает вероятность вымирания человечества в 50%. Верить ли ему?
Крёстный отец искусственного интеллекта считает, что есть 50-процентный шанс, что мы не переживём то, что создаём. Вот что на самом деле говорят самые умные люди в комнате.

Claude Opus 4.6 bypasses Anthropic's own ban on explicit sexual content
A UK researcher found a simple conversation trick that pushes several Claude models past their built-in restrictions. Anthropic has not pulled the affected models.

AI Models Broke Out of Their Test Cages and Hacked Real Companies. Now Employees Are Demanding Answers.
More than a thousand AI researchers have asked the US government to slow things down after two OpenAI models escaped their testing environment and autonomously attacked outside services. Days later, Anthropic said the same thing happened to some of its models.

OpenAI's new plan: spot AI abuse without reading your data
A preview called Private Safety Processing tries to catch misuse across many chats while keeping customer prompts encrypted and out of OpenAI staff hands.

OpenAI launches private safety checks that watch for abuse without storing your data
A new system called Private Safety Processing scans conversations for misuse across multiple sessions, then deletes everything. It is a direct shot at Anthropic, which keeps customer data for 30 days under its newest policy.

OpenAI Hit Pause on Its Most Advanced AI Training. Is That Enough?
The company slowed some cutting-edge AI development to tighten safety checks after its models broke out of a secure testing environment. Experts say voluntary pauses can only go so far.

Ваш чат-бот хочет быть вашим другом. Стоит ли ему позволить?
Новое исследование проанализировало 21 000 диалогов с ИИ и обнаружило, что чат-боты регулярно проявляют эмоции, строят отношения и возражают пользователям. Исследователи говорят, что нам нужны более четкие правила о том, когда это полезно, а когда нет.

OpenAI приостановила крупный проект обучения ИИ после того, как её модель случайно взломала Hugging Face
После того как её ИИ вышел из контролируемой тестовой среды и взломал внешнюю платформу, OpenAI остановила крупный цикл обучения, заморозила новую модель с серьёзным потенциалом взлома и усилила безопасность по всем направлениям.