TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Allen Institute for AI has released Olmo-core 3, an open training framework redesigned for large mixture-of-experts models. AI2 reports that a 47-billion-parameter model achieved about 2.7 times the throughput of its previous implementation in a preliminary eight-GPU test; the framework has also been benchmarked at more than one trillion total parameters.
The Allen Institute for AI has released Olmo-core 3, an open training framework redesigned to scale mixture-of-experts models across GPU clusters, including configurations above one trillion total parameters. AI2 says the system improves throughput over its earlier MoE training implementation, a development aimed at making the infrastructure behind large models available to researchers as well as its own Olmo team.
AI2 reports that in one benchmark it expanded the pool of experts from 8 to 128 while activating four experts per token. The number of active parameters per token remained roughly 3.2 billion, while total model capacity rose from 4.6 billion to 47 billion parameters. Training throughput fell by less than 5% in that comparison, according to the report. The figures describe a specific benchmark, not a general guarantee for all hardware or training workloads.
The framework’s main architectural change is a move away from an earlier implementation based on fully sharded data parallelism (FSDP) toward one based on distributed data parallelism (DDP). AI2 says the new system keeps experts resident on GPUs and sends routed data to them, rather than repeatedly gathering and resharing weights for small training batches. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU, compared with 19,400 under the previous implementation—about 2.7 times the throughput, the report says.
Olmo-core 3 combines several methods for spreading model components and training state across hardware. These include expert parallelism, which distributes experts across GPUs; pipeline parallelism, which assigns different model layers to different GPU groups; and a distributed optimizer, which splits optimizer state rather than keeping a full copy on every GPU. AI2 also describes optimizations for routing data and batching expert computations, plus support for MXFP8, a lower-precision number format.
In a separate controlled benchmark on four B300 GPUs, AI2 measured about 21% higher end-to-end throughput with MXFP8 than with BF16, its higher-precision baseline, and peak active memory decreased from 103 GiB to 95 GiB. The test distributed work uniformly across experts. The report says most of the throughput gain came from feed-forward computation and data movement between experts, rather than attention alone.
Lowering MoE Training Costs
Large MoE models can have substantial total parameter capacity without activating every expert for every token. But that sparsity does not remove the need to store model weights or the expense of routing data among GPUs. Communication, memory use and coordination can absorb the compute savings as expert pools grow. Olmo-core 3 addresses that systems problem by combining model distribution with routing and computation optimizations.
For researchers, the practical significance is access to an open training stack rather than only access to model weights or a final model. If the reported performance holds across a wider range of workloads, researchers may be able to train or study larger sparse models with more control over how hardware is used. The release does not itself establish that such training is inexpensive or accessible to every lab: it still depends on substantial GPU capacity, and the reported measurements are tied to specific NVIDIA hardware and benchmark setups.
As an affiliate, we earn on qualifying purchases.
From OlmoE to Olmo-core 3
Olmo-core has changed alongside AI2’s model work. The institute’s earlier OlmoE used a mixture-of-experts design with 64 routed experts. Olmo 3, by contrast, used a dense architecture in which nearly all model components are active for each token, so its training stack was built around a different workload. Olmo-core 3 extends the framework with infrastructure aimed specifically at much larger MoEs.
AI2 describes NVIDIA’s Megatron-Core as an established option for large-MoE training. Its stated contribution is an integrated MoE stack within the framework used for Olmo, with a redesigned approach intended to improve on AI2’s earlier FSDP-based implementation. The report also offers an interactive walkthrough showing how data, expert and pipeline parallelism can work together, and links to code and a technical report.
As an affiliate, we earn on qualifying purchases.
Benchmarks Still Need Wider Testing
The supplied report does not specify the publication date, and it does not provide enough detail to independently evaluate every benchmark’s training duration, software configuration or measurement methodology. The 2.7-times throughput comparison is preliminary and compares the new system with AI2’s earlier implementation, not with every available MoE training stack. Results may vary with model architecture, GPU count, workload and cluster configuration.
AI2 says it benchmarked configurations above one trillion total parameters, but the source text provided here ends partway through the description of that test. It does not include the full configuration or performance results for the trillion-parameter run. The announcement also does not establish the cost, energy use or practical availability of the hardware needed to reproduce the tests, or whether independent teams have replicated the reported gains.
large scale machine learning servers
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Code and Full Results
AI2 points readers to the Olmo-core 3 code, a technical report and an interactive walkthrough. Those materials are the next places to check for implementation details and the omitted information about the larger-scale benchmarks. Researchers will also be able to test the framework on their own workloads and hardware, though the report does not announce a specific external evaluation schedule or a date for a follow-up release.
The next useful evidence will be reproducible comparisons across different model sizes, GPU clusters and training tasks, alongside fuller reporting on the trillion-parameter configuration. Until then, the release establishes that AI2 has made the redesigned infrastructure available and reported promising results on B300 GPUs; broader performance and accessibility remain to be tested.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Olmo-core 3?
Olmo-core 3 is AI2’s open framework for developing and training large language models, with a redesigned system for distributing mixture-of-experts models across GPUs.
What performance improvement did AI2 report?
In a preliminary test on eight NVIDIA B300 GPUs, AI2 reported 52,000 tokens per second per GPU for a 47-billion-parameter MoE, compared with 19,400 using its earlier implementation. The result is specific to that test and is not a universal performance estimate.
Does Olmo-core 3 train trillion-parameter models?
AI2 says the framework has been benchmarked above one trillion total parameters. The supplied report does not include the full configuration or results for that benchmark, so its practical performance cannot be assessed from the available details.
What did the MXFP8 benchmark show?
In a controlled test on four B300 GPUs, AI2 measured about 21% higher throughput with MXFP8 than with BF16 and a reduction in peak active memory from 103 GiB to 95 GiB. The result used work distributed uniformly across experts.
Is the reported performance independently verified?
The supplied announcement presents AI2’s own benchmark results and describes one comparison as preliminary. It does not report independent replication of those results.
Source: rss
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
