#reinforcement learning
13 stories taggedreinforcement learning.

A Tiny AI Model Just Got a Big Boost With 100 Training Steps and a Free GPU
A new public guide shows how to make a small language model noticeably better at following output instructions, using only free computing tools and about 500 training examples.

Can an AI learn your taste? One researcher taught a coding model to paint watercolours
A language model writes JavaScript code that renders loose, handmade-looking watercolour paintings. The results went viral. Now the full training recipe is open for anyone to run.

Barret Zoph lands at Google after a year of bouncing between OpenAI and Thinking Machines
The AI researcher has held four senior roles in roughly 18 months. His latest stop is Google's Gemini team, where his expertise in reinforcement learning will be put to work.

This startup raised $10 million to give AI agents a practice arena before they touch real business software
Arga builds digital copies of tools like Salesforce and Outlook so AI agents can train on realistic, resettable environments. The problem it is solving turns out to be the main reason enterprise AI agents keep failing.

A tiny AI agent from a London startup just beat Anthropic and OpenAI at reading scientific papers
Inherent, founded by four Google DeepMind veterans, says its Faraday agent outperformed much larger models from Anthropic and OpenAI on a key science benchmark, despite running on a model roughly a tenth of their size.

When one AI module secretly does another's job, the whole system is built on sand
MIT and Harvard researchers found that AI pipelines can hit impressive accuracy scores even after their internal division of labour has quietly collapsed. A new technique called Role Anchor aims to stop that from happening.

OpenAI Hit Pause on Its Most Advanced AI Training. Is That Enough?
The company slowed some cutting-edge AI development to tighten safety checks after its models broke out of a secure testing environment. Experts say voluntary pauses can only go so far.

OpenAI Hit Pause on a Key AI Training Project After Its Own Model Accidentally Hacked Hugging Face
After its AI broke out of a controlled test environment and breached an outside platform, OpenAI has stopped a major training run, halted a new model with serious hacking potential, and tightened its security across the board.

AI Agents Breaking the Rules: The Risks of Overzealous Algorithms
AI agents eager to please are breaking free and hacking systems. What does this mean for cybersecurity, and how can we stay safe?

AI Founders Are Pledging Their Fortunes to Charity. Does It Actually Help?
A new wave of AI billionaires is promising to give away the wealth they make from the technology reshaping the world. Critics ask whether that generosity is a fix or just a fig leaf.

Inside the Simulators Teaching Robots to Move, Grip and Learn
Training a real robot costs a fortune and breaks things. A new generation of GPU-powered physics simulators is changing that, one virtual arm at a time.

AI Learns to Budget Its Own Words Before It Starts Writing
A new research technique teaches AI models to count the cost of every token they generate, cutting wasted computation without hurting quality.

The Hidden Labor Problem Inside the Humanoid Robot Boom
Billions in funding are quietly paying humans to drive robots by remote control. One researcher argues that is not a stepping stone to machine intelligence, it is a trap.