← News

Ai2 Olmo-core 3 official release graphic naming the open MoE training framework

1 Oct 2026

Ai2

Ai2 releases Olmo-core 3 to train trillion-parameter MoE models more efficiently

On Oct. 1, 2026, Ai2 said Olmo-core 3 switches MoE training to a DDP-plus-expert-parallelism design meant to keep large expert pools efficient — including benchmarks past one trillion total parameters.

Open labs win or lose on whether outsiders can actually train the same class of sparse models the closed labs run. Shipping the MoE training stack — not just weights — is how Ai2 tries to keep trillion-parameter sparse training from becoming a closed-shop skill.

On Thursday, 1 October 2026, the Allen Institute for AI, which shortens its name to Ai2, released Olmo-core 3. The post is titled “Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs.” The page dates it October 1, 2026. It does not print an hour. Ai2 calls the release a significant upgrade to its open framework for developing large language models, and it says the upgrade includes a redesigned open mixture-of-experts training system. A large language model is software that reads and writes text. Mixture of experts, shortened to MoE, means the model is split into many specialist pieces, called experts. For each small piece of text, a token, only some of those experts run. A parameter is one learned number inside the model. Those lines are Ai2’s.

What Ai2 says the stack is for. Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency. A trillion parameters is a thousand billion learned numbers. Efficiency, here, means getting that work done without wasting as much of the computers’ time. Ai2 says this stack is one of the core systems behind the next generation of Olmo, and part of the work of opening the tools and training setup behind each new model. Training, in that sentence, is the long job of setting the model’s numbers. The post says the next Olmo will use an MoE design. It says Ai2 is aiming for that model to be its most capable Olmo yet, trained on its largest dataset and with its longest context window. A context window is how much text the model can hold in view at once. The post describes that model as what Ai2 is building next. It does not say a finished next Olmo shipped on October 1. Those lines are Ai2’s.

Why a bigger expert pool is hard, in Ai2’s account. An MoE can hold many more parameters than a dense model, because each token uses only part of the model. A dense model runs nearly all of itself for every token. The full MoE still has to sit in the memory of the graphics chips, the GPUs, and those numbers still have to be updated while training runs. Sending each token to the right experts across a cluster of chips adds its own traffic. Ai2 says that as MoEs grow, that traffic can eat the savings that came from using only part of the model. Olmo-core 3 is built to close that gap. Those lines are Ai2’s.

A smaller test of that gap. Ai2 says that in one benchmark it grew the expert pool from 8 experts to 128, and still picked only four experts for each token. The number of parameters that actually run for each token stayed about the same, roughly 3.2 billion. The whole model grew from 4.6 billion parameters to 47 billion. Training throughput fell by less than 5 percent. Throughput is how much text the stack processes in a given time. That comparison is Ai2’s. The page does not name an outside lab that repeated it.

The change in how the chips share the work. Ai2 says its earlier MoE path in Olmo-core used fully sharded data parallelism, shortened to FSDP. That setup gathered the model’s weights for each small batch of training data and then split them up again. Weights are the learned numbers. Olmo-core 3 switches to distributed data parallelism, shortened to DDP, plus expert parallelism. Expert parallelism spreads the experts across GPUs, so each chip stores only part of the pool. Under the new design, the experts stay on the GPUs and the relevant data is routed to them, instead of gathering the weights over and over. Ai2 also names NVIDIA’s Megatron-Core as an established option for training large MoEs. It says Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, and that the redesign improves throughput over Ai2’s earlier FSDP path. Those lines are Ai2’s.

The speed comparison Ai2 reports. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack. The earlier implementation processed 19,400 tokens per second per GPU. Ai2 says that is about 2.7 times the throughput. A B300 is an NVIDIA data-center chip. Per GPU means the count is for each chip, not a total across the eight. Preliminary means Ai2 is reporting an early measurement, not a finished product score. Those figures are Ai2’s. The 19,400 figure on this page is the earlier Olmo-core implementation. The page does not print a Megatron-Core token rate for this test.

The large runs. Ai2 says it has benchmarked a 1.2-trillion-parameter configuration, with 58.36 billion parameters active per token, across 512 GPUs. Active per token means the slice that actually runs for each piece of text. The highest throughput Ai2 reports for that run is 858 TFLOP/s/GPU. A teraflop is a trillion simple math operations. The figure is how many of those operations each GPU was doing per second, on the model’s work. Ai2 says these tests used random routing, a way of sending tokens to experts without regard to which expert would have been best, so the measurement is of the system, not of how good a trained model is. Ai2 also says it tried DeepEP v2, a different way of moving data between experts across GPUs, and reached a configuration of about 2.38 trillion total parameters. That run was a short-capacity test, not a full training run. Ai2 says it shows the scale the stack can reach, not a sustained training speed. Those lines are Ai2’s.

Who can use it. Ai2 says Olmo-core 3 is fully open. Researchers and developers can use it to train their own MoEs, adapt it to different hardware, and try different ways of routing tokens and splitting the model across chips. The post points to a technical report and to the Olmo-core code on GitHub. The post does not print a price, a customer count, or a funding figure. Those lines are Ai2’s.

The picture is Ai2’s Olmo-core 3 title graphic. A dark teal field fills the frame. A pink mark made of small squares sits beside the words “Olmo-core 3,” set in the same pink. The frame does not print a calendar date. It is the release graphic. It is not a photograph of a lab, a chip, or a person.

In plain terms, Ai2 said on Thursday that Olmo-core 3 is an open training stack for mixture-of-experts models, the kind that keep most of their numbers idle on any one piece of text. Ai2 says the new path leaves the expert pieces sitting on the GPUs and sends the data to them, instead of gathering the weights for every small batch. On eight B300 chips, Ai2 reports about 52,000 tokens a second per chip for a 47-billion-parameter model, against about 19,400 on its earlier setup, roughly 2.7 times the throughput. Ai2 also reports a 1.2-trillion-parameter setup across 512 chips, with about 58.36 billion parameters active per token, and a short test that reached about 2.38 trillion parameters. The next Olmo, Ai2 says, will be an MoE. The post does not say that model is out.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources