Same AI Memory, Fewer Tokens: How a New Approach Beats a Leading Agent System at a Fraction of the Cost
Two AI systems both teach agents to learn from their own mistakes. One sends every lesson every time. The other sends only what each model can actually use. The difference shows up in the bill.

Key points
- ALTK-Evolve matched or beat the ACE agent-memory system on a 168-task benchmark while using as little as one-seventh of the computing tokens per task.
- On DeepSeek-V3.2, a strong AI model, ALTK-Evolve scored 89.3% task completion versus ACE's 80.4%, at roughly 40% of ACE's token cost.
- On the weaker model gpt-oss-120b, accuracy was nearly identical, but ALTK-Evolve used about one-seventh as many tokens per task.
- Both systems refuse to compress an agent's learned lessons into a short summary, but they differ sharply on how those lessons reach the model at decision time.
- The work was shared by the ALTK-Evolve team on Hugging Face.
Imagine teaching a new employee not through a manual, but by letting them make mistakes and then showing them exactly what went wrong. That's roughly what both of these AI systems do. The real question is whether you hand the employee every note ever written, or just the ones that apply today.
What problem are these systems solving?
AI agents, software that carries out multi-step tasks on its own, fail in a telling way. They usually know the rules but haven't learned to apply them cleanly.
Give an AI agent a task like splitting a bill across several simulated apps, and it might call the wrong function, pull the wrong person's record, or return an answer when no answer was asked for.
Both ACE (Agentic Context Engineering) and ALTK-Evolve fix this by mining the agent's own past attempts for lessons, then feeding those lessons back the next time it works. No human labels, no retraining the underlying model. Just memory, used at the moment of decision.
Where do the two systems disagree?
They disagree on delivery, and that's where the cost difference comes from.
ACE keeps one carefully maintained playbook and injects it into the agent's context, the block of text the model reads before acting, on every single step. Thorough, but expensive.
ALTK-Evolve treats delivery as a dial rather than a constant. For a strong model with plenty of headroom, it can send the full set of lessons. For a weaker model, it selects only the lessons most relevant to the current task, keeps a fixed core of high-confidence guidelines, and leaves the rest in storage. The same lessons exist in both cases; fewer of them travel to the model.
The results on the AppWorld benchmark, a test of 168 realistic multi-app tasks, are shown below.
| Model | System | Task completion | Tokens per task |
|---|---|---|---|
| DeepSeek-V3.2 (strong) | ACE | 80.4% | 634K |
| DeepSeek-V3.2 (strong) | ALTK-Evolve | 89.3% | 263K |
| gpt-oss-120b (weaker) | ACE | 54.8% | 777K |
| gpt-oss-120b (weaker) | ALTK-Evolve | 56.0% | 116K |
Tokens are the units of text an AI model reads and writes. More tokens mean more computing time and higher cost.
Why does sending fewer lessons sometimes work better?
On hard tasks, the agent needs to pick the right lesson, not wade through all of them. On easy tasks with a weaker model, a full playbook gives a small edge because the answer is close to the surface. That same flood of context gets in the way when the task is harder.
A stronger model absorbs more context without losing the thread. A weaker one can drown in it.
Both systems agree on one important principle: don't compress lessons into a short summary. A lesson that appeared in five separate tasks is meaningfully different from one that appeared once. ALTK-Evolve tracks how many independent episodes produced each lesson. ACE keeps a helpful-versus-harmful count per item. Different labels, same underlying logic. Our earlier story on how Asana built shared agent memory across a company, published 6 August, shows how the question of what agents remember, and what they forget, is becoming a design problem across the industry.
What happens next?
The ALTK-Evolve team say the question of exactly how many lessons to send, and how that scales across models of different strengths, is the subject of their next post. A full report is already public for researchers who want to test the approach.
For anyone building or paying for AI agent systems, this is what matters most: accuracy and cost don't have to trade off. What the model is told at decision time shapes both, and the right amount isn't the same for every model. That's the piece most agent system designers are still treating as a fixed setting rather than a variable worth tuning.
Common questions
Does this mean ACE is a bad system?
No. ACE and ALTK-Evolve solve the same problem and agree on the hard parts. ACE's own efficiency story focuses on how cheaply it builds its memory store. ALTK-Evolve's gain comes from serving fewer lessons per step, which is a different axis entirely.
Does this affect ordinary users directly?
Not yet in a product you can download. The work is at the research stage, tested on a benchmark of simulated tasks. But the cost savings it points to matter for any company running AI agents at scale, and lower running costs typically translate to lower prices over time.



