How a Research Lab Stopped Researchers From Hoarding GPUs

The Allen Institute for AI replaced a broken priority system with a budget model that treats GPU time like money. The fix is unglamorous and it works.

AI2Day NewsdeskAsistido por IAPublicado Editor: Lee Brown4 min read
Illustration: A dense server room filled with rows of glowing GPU racks bathed in cool blue and green light
Ilustración creada con IA. No es una fotografía de los hechos descritos.
Share

Key points

  • Ai2's infrastructure team found that GPU demand at any moment exceeds supply by two to three times across its cluster of thousands of Nvidia H100 and B200 chips.
  • Researchers gamed the old priority system so heavily that 100 percent of scheduled workloads eventually claimed the highest priority, starving lower-priority work of any time at all.
  • A new budget model, detailed on Hugging Face, gives each research project a guaranteed percentage share of GPU time decided in advance by managers, not by who shouts loudest.
  • Gaming the new system costs teams their own GPU budget, making honest negotiation cheaper than cheating.

GPUs, the specialised chips that do the heavy number-crunching AI models need, are expensive and scarce. At the Allen Institute for AI (Ai2), a non-profit research lab, thousands of them sit in clusters running experiments for around 150 scientists at any given time. Demand runs at two to three times what's available. Every spare GPU hour has multiple teams competing for it.

For a while, the lab tried to manage that competition with a priority queue: tag your job HIGH or LOW, and the scheduler decides who gets chips first. It collapsed almost immediately.

What went wrong with priority queues?

Within months, 100 percent of scheduled workloads carried the highest priority label, because the label was free. Giving your job a high tag cost nothing, so everyone did it. Lower-priority jobs received almost no GPU time.

Researchers also discovered a second trick. If you could protect a job from being interrupted (from "preemption," which means the system pausing your job to hand the chip to someone else), you could park a do-nothing job on a GPU and reconnect to it later when you needed to debug something quickly. Ai2 calls this "GPU squatting." On-call engineers spent most of their incident-response time not fixing hardware but persuading teams to voluntarily shut down jobs blocking maintenance.

This problem is older than AI. A 2011 paper the team cites describes engineers at a search company writing infinite loops into code just to make their jobs look busier, so they could keep dedicated machines. Hardware changes; human incentives don't.

How does the new budget model work?

Instead of labels, Ai2 now allocates slices of time. Before any workload exists, managers decide what fraction of total GPU time each research project deserves. A project might receive a guaranteed 35 percent claim on the whole cluster. That share is its budget.

If a job isn't backed by budget, it gets no protection from interruption. Squatting on a GPU now drains the team's own allocation. Gaming the system becomes more expensive than simply arguing, in a formal review, for a larger slice of time.

The algorithm underneath this, called hierarchical fair-share scheduling, distributes shared resources proportionally through a ranked structure of teams. It isn't new: it traces back to the Hadoop Fair Scheduler in 2009 and remains in common cluster-management tools today. What Ai2 changed is what feeds into it. The weights are now budget decisions made by managers with full knowledge of research priorities, not labels that anyone can inflate.

This matters beyond one lab. Ai2 appeared in our coverage on 2 October of its open-source training system for trillion-parameter models, and the pattern is consistent: the hard problem isn't building the chip, it's getting useful work out of it. Ai2's fix is deliberately unglamorous. It moves the argument about who deserves resources out of the scheduler and into a budgeting meeting, where it can actually be resolved.

For researchers at universities or corporate AI teams dealing with shared computing, that's the practical lesson here. Priority labels are a social contract that collapses under pressure; budget allocations are harder to game because the cost lands on the person gaming them.

Common questions

Does this apply to cloud AI services that ordinary people use?

Not directly. This system runs inside a private research lab. The same priority-gaming problems appear in any organisation where people share expensive computing resources, including universities and corporate AI teams, but cloud providers handle allocation through billing, not scheduling labels.

Could researchers still game the new system?

Yes, but at a cost to themselves. Any trick that grabs GPU time draws from the team's own budget, which makes sustained gaming self-defeating. The design goal is to make honest advocacy cheaper than manipulation.

© 2026 AI2Day