The Allen Institute Just Open-Sourced a Training System for Trillion-Parameter AI Models
Olmo-core 3 lets academic labs and smaller teams build huge AI models without the usual sky-high compute bill. Here's what changed and why it matters.

Key points
- Olmo-core 3, released by the Allen Institute for AI, has benchmarked a 1.2-trillion-parameter AI model running across 512 specialised chips, reaching 858 trillion floating-point operations per second per chip.
- A redesigned training method delivered 2.7 times more throughput on eight Nvidia B300 GPUs compared with the previous system, in preliminary tests.
- A lower-precision number format called MXFP8 cut peak memory use from 103 gigabytes to 95 gigabytes and lifted training speed by roughly 21 percent in controlled tests.
- The framework is fully open-source, published via Hugging Face; universities and smaller labs can build on it without paying for proprietary tooling.
- The Allen Institute says the next generation of its Olmo AI model will be built on this stack and will carry its largest training dataset yet.
Most AI training infrastructure sits behind closed doors at the big labs. The Allen Institute for AI just opened one.
Olmo-core 3 landed this week: a technical framework for training the Olmo family of large language models, the technology behind chatbots like ChatGPT and Claude. It's built specifically for a class of AI called mixture-of-experts models, or MoEs.
What is a mixture-of-experts model, and why does it matter?
An MoE model is a smarter way to pack more capability into an AI without running up the electricity bill every time it processes a sentence.
Ordinary "dense" AI models activate every component for every word they process. An MoE contains specialised sub-networks called experts and routes each word to only a small handful of them. The full model is enormous; the active slice stays manageable.
Coordination is the hard part. Spreading experts across many GPUs, the specialised chips that do AI's heavy number-crunching, creates communication costs that can eat away at the efficiency gains. Olmo-core 3 is built to tackle that at serious scale.
In one benchmark, the team expanded the expert pool from 8 to 128 specialists. Each token, a small chunk of text roughly the size of a syllable, still only consulted four of them. Total model capacity grew from 4.6 billion to 47 billion parameters (the learned values that give an AI its abilities) while training throughput dropped by less than 5 percent. The same infrastructure has been tested at over one trillion total parameters.
What does this mean for researchers on a tight budget?
A university team or a small lab can now reach training scales that were essentially locked away at Google or Meta a year ago.
We first covered the Allen Institute for AI on 28 July 2026, and four stories later the pattern is consistent: the institute keeps releasing open tools while its peers keep theirs proprietary. Our 1 September story on Allen AI's benchmark-auditing tool caught the same instinct at work.
The new system drops a method called fully sharded data parallelism, which gathers and reshards model weights repeatedly across chips, in favour of one that keeps experts sitting on their assigned GPUs and routes data to them instead. Combined with MXFP8, a lower-precision number format that fits more computation into less memory, that's where most of the speed gains come from.
Practical example: a researcher writing a grant proposal could run a 47-billion-parameter MoE on a university GPU cluster without needing a corporate cloud account.
Privacy note: Olmo-core 3 is a training framework, not a consumer app. Your training data stays on your own hardware; nothing is sent to the Allen Institute.
Price: Free and open-source. You still pay for the GPUs.
The Allen Institute says the next Olmo model, trained on its largest dataset ever, will be the first to run on this stack. That's the thing worth watching: if open-source MoE training matures here before the proprietary labs decide to care, the gap between well-funded AI and everyone else gets a little less comfortable for the well-funded side.



