Apple researchers built a smarter way to train AI agents. Here is why the boring details matter.

A new training method from Apple ML Research teaches AI agents to focus on the lessons they have not learned yet, skipping exercises that are too easy or too hard. It's quietly important for anyone who cares about where general-purpose AI is heading.

AI2Day NewsdeskAsistido por IAPublicado Editor: Lee Brown4 min read
Illustration: A vast server room with long rows of glowing blue and white rack-mounted computers receding into the distance
Ilustración creada con IA. No es una fotografía de los hechos descritos.
Share

Key points

  • Apple ML Research published a new training method called RISED that automatically picks the most useful practice tasks for AI agents learning across multiple environments at once.
  • The method solves a specific waste problem: in a standard training batch, some task groups are aced every time and some are failed every time, and both teach the agent almost nothing.
  • RISED uses a scoring system it calls rubrics to rank task groups by how much the agent still has to learn from them, then trims the least useful ones before training begins.
  • The technique also includes a self-distillation step, where the agent uses its own best past answers to reinforce correct behaviour without needing human labels.
  • This research matters for the push toward generalist AI agents: software that can handle many different kinds of tasks rather than one narrow job.

AI agents are software programs that carry out multi-step tasks on their own, booking a meeting, writing code, or browsing the web. Training one agent to handle many different task types at once is one of the hardest open problems in AI research, and it's the problem Apple ML Research set out to address.

The new method is called RISED (Rubrics for Agentic Multi-Environment Selection and Self-Distillation). Dense name, simple idea.

What problem does RISED actually solve?

When researchers train an AI agent across many environments simultaneously, they batch the training data into groups. Some groups the agent already handles perfectly. Others it fails completely. Both are wasted effort.

A group the agent always gets right offers no new signal. A group it always gets wrong is equally useless, because the scoring system used in modern AI training, called group-relative reward, works by comparing answers within a group. If every answer scores zero, there's nothing to compare.

RISED adds a rubric layer. Before each training step, it scores every candidate group of tasks and selects only those where the agent is still genuinely learning, neither coasting nor drowning. Self-distillation works differently: the agent's own strongest previous outputs get recycled as positive examples, reinforcing what it already does well without requiring a human to label anything.

Should anyone outside AI research care about this?

Yes, for a concrete reason. General-purpose AI agents are what companies want to deploy in hospitals and businesses. The faster training methods improve, the sooner those agents arrive, and the more capable they'll be.

This is Apple's third AI training paper we've covered in a month. Our 17 September story on DACA-GRPO and our 1 October story on RLTL;DR each addressed a different flaw in how agents learn from feedback. RISED addresses a third: how to stop wasting training compute on tasks that can't teach the agent anything new.

More capable agents trained on less data also means the gap between frontier labs and everyone else could narrow, or widen faster, depending on who adopts techniques like this first.

The RISED paper is a technical contribution, not a product launch. No company has announced it's shipping an agent built this way. What it does is fill a real gap in the standard training playbook, one that researchers across the industry have been working around rather than solving. That's the honest measure of it.

Common questions

Does this mean Apple is building a general-purpose AI agent?

The paper is a research publication, not a product announcement. Apple ML Research publishes methods that other teams may use. No consumer product has been announced.

What is group-relative reward, in plain English?

It's a scoring method where the AI's answers are judged by comparing them to each other inside a small group, rather than against a fixed right answer. If all answers in the group are equally wrong or equally right, the method can't tell which direction to improve, which is the gap RISED addresses.

How does self-distillation differ from normal training?

Normal training relies on humans or a separate model to label the correct answers. Self-distillation lets the agent treat its own best outputs as the positive examples, cutting the cost and delay of external labelling.

© 2026 AI2Day