THE SIGNAL IN ONE SENTENCE

A trillion parameters sounds like a bigger brain. In Ai2's newest release, it is more useful to picture a very large machine room. The Allen Institute for AI released Olmo-core 3 on October 1. It is a redesigned, open training stack for mixture-of-experts models, the sparse systems that can contain many specialized blocks while using only a small selection for each token. That design promises an appealing bargain. A model can own a huge cabinet of learned components without opening every drawer for every word. The arithmetic done for one token can stay much smaller than the total model. The cabinet still has to exist. Its weights must live somewhere. Its training state must be stored and updated. Tokens must be sent across a cluster to the experts chosen for them, then returned in the right order. If the routing system spends too long moving data or waiting for a crowded expert, the elegant sparse model becomes an expensive traffic jam. Olmo-core 3 is Ai2's attempt to make that plumbing work at much larger scale, and to publish the plumbing rather than only the eventual faucet. The plain signal is that open AI infrastructure now includes more of the difficult systems layer beneath model weights. Researchers can inspect how Ai2 distributes experts, routes tokens, divides layers, stores optimizer state and measures the resulting bottlenecks. But openness changes who can study the machine more than it changes who can afford to run its largest configuration. Ai2 reports a useful first benchmark. It increased an expert pool from 8 experts to 128 while choosing four experts for each token. Active parameters stayed around 3.2 billion per token, while total parameter capacity rose from 4.6 billion to 47 billion. Training throughput fell by less than 5 percent. That is the basic MoE proposition in one test: much more total capacity, nearly the same amount of active computation. Ai2 also reports a larger preliminary comparison on eight Nvidia B300 GPUs. A 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using its earlier implementation. The company describes that as about 2.7 times the throughput. The change is not a mysterious new intelligence trick. It is a change in what moves. Ai2's earlier MoE implementation used fully sharded data parallelism, or FSDP, in a way that gathered and resharded weights for each small batch. Imagine moving every specialist's filing cabinet into the meeting room, letting a few people use it, then moving the cabinets away again. Olmo-core 3 uses distributed data parallelism, or DDP, to keep experts resident on GPUs and route the relevant token data to them. The cabinets stay put. The work travels. That turns communication into the main event. Expert parallelism spreads different experts across GPUs. Pipeline parallelism places different layers on groups of GPUs. A distributed optimizer spreads the extra state needed to update the model. Rowwise expert parallelism puts routed data directly into expert input buffers. GPU-resident routing keeps routing metadata on the accelerators so a CPU does not have to wait for it to return. Grouped matrix operations combine many small expert jobs into work a GPU can handle efficiently. None of those phrases is likely to improve a dinner party. Together, they determine whether a sparse model saves time or merely relocates the bill. Ai2 also added support for MXFP8, a lower-precision format that can reduce both computation and data movement. In a controlled test on four B300 GPUs with work distributed uniformly across experts, Ai2 measured about 21 percent higher end-to-end throughput than with BF16. Peak active memory fell from 103 GiB to 95 GiB. Lower precision is not free speed. Values may need conversion, and some operations are more sensitive than others. Ai2 says most of the measured benefit came from feed-forward work and moving data between experts, not attention alone. The useful engineering question is not whether eight-bit arithmetic sounds efficient. It is whether less work and less traffic outweigh the conversion cost across the entire run. Then comes the headline number. Ai2 benchmarked a 1.2-trillion-parameter configuration with 58.36 billion parameters active for each token across 512 Nvidia B300 GPUs. Its highest observed throughput was 858 trillion floating-point operations per second per GPU. That was a systems test, not the training of a trillion-parameter model whose quality can be evaluated. Ai2 says the benchmark used random routing. Random routing removes the learned decision about which expert should receive a token, which makes it useful for measuring the machinery but not for judging what a trained model knows or whether its routing improves answers. The team also reached 2.38 trillion total parameters in a short capacity test using DeepEP v2. Ai2 explicitly says that test demonstrated reachable scale, not sustained training performance. Those qualifications belong beside the numbers, not in a footnote underneath them. Olmo-core 3 shows that Ai2's stack can arrange and exercise very large sparse configurations on substantial hardware. It does not announce a finished trillion-parameter Olmo model. It does not establish better model quality. It does not tell us what a complete training run would cost in electricity, engineering time or money. It does offer something rarer than a large number: documented failure. Ai2's technical account describes a routing score that appeared to improve while actual expert workload became less balanced. The team calls the failure token gerrymandering. The metric could be pleased while the machines were not. That is a wonderfully specific warning. MoE training needs experts to receive a reasonably balanced load. If too many tokens choose the same few experts, those GPUs become a queue while others wait. A proxy score can say balance is improving even when the real work distribution is deteriorating. Engineers need to inspect the load, not merely the objective meant to encourage it. Ai2 also reports that reducing experts' learning rates because individual experts see fewer tokens did not improve the model family it tested. GPU calculations took different amounts of time when input values changed even when matrix shapes stayed the same, so fair performance comparisons need matching values as well as matching dimensions. Attempts to overlap communication and computation sometimes slowed the full system instead of speeding it up. These are not embarrassing scraps. They are part of the value of open infrastructure. A released weight file lets another team run one trained artifact. Released training code helps that team inspect how the artifact can be made, changed and measured. Negative findings help it avoid spending precious cluster time rediscovering the same dead ends. At this scale, a failed idea can burn more than an afternoon. The repository is published under the Apache 2.0 license. Its documentation describes source installation, a PyPI package and optional components for functions such as grouped MoE operations and low-precision training. It also warns that published container images may not work on clusters with different hardware, drivers or CUDA versions. That warning is the access story in miniature. Open code is inspectable and adaptable. It is not automatically portable across every accelerator. It does not supply a high-speed interconnect, hundreds of current-generation GPUs, the electricity to run them or the engineers who know why one collective operation stalled at 3 a.m. The largest published test used 512 B300 GPUs. Even the preliminary 47-billion-parameter comparison used eight. A university group with a few older accelerators can read the implementation and borrow ideas. It cannot reproduce the headline benchmark merely because the license permits it. This does not make the release fake-open. It makes openness a ladder with several rungs. The code can be read. Smaller configurations can be tested. System choices can be challenged. Researchers can port parts to other hardware, compare routing methods, reproduce negative findings and use the framework to train models within their actual budget. A well-funded lab can attempt the larger recipes. Independent verification of the full scale claims still requires independent access to comparable hardware. The practical place to start is small and adversarial. Run a compact MoE on the hardware already available. Measure tokens per second, peak memory, network traffic and expert load separately. Change the number of experts while holding active parameters as steady as possible. Repeat measurements with identical input values. Disable one optimization at a time. Test what happens when routing becomes uneven. Record how much time is spent computing, converting number formats, moving data and waiting. Do not copy the fastest published configuration and assume the same result. Ai2's numbers are tied to B300 GPUs, cluster topology, software versions, routing assumptions and chosen model shapes. Another system may have slower links, less memory, different kernels or a bottleneck that the published setup avoided. For research groups, the release can support experiments in expert specialization, routing stability, load balancing and lower-precision training. Hardware teams can use it to study where communication overwhelms computation. Universities can teach distributed training with real production code rather than a diagram that stops before the network. Independent auditors can inspect how benchmark choices shape the story a throughput number tells. Model builders can also use the stack as a warning label for future product claims. If a company announces a trillion-parameter MoE, ask how many parameters are active per token. Ask whether the number describes a trained model, a short capacity test or a synthetic systems run. Ask about routing, throughput, memory, networking, power and the duration of the test. Ask whether model quality was measured at all. Ai2 says Olmo-core 3 will underpin its next generation of Olmo, which is planned to use an MoE architecture, a larger dataset and a longer context window. That future model is not part of this release. The infrastructure arrived first. That ordering is useful. It gives the public a chance to inspect the tracks before the train. Olmo-core 3 does not democratize trillion-parameter training by itself. It does make a difficult layer of the work more visible, including the places where elegant ideas failed. For open research, that is meaningful progress. The pipes are on the table. The machine room is still expensive.

01

WHAT ACTUALLY CHANGED

Ai2 released Olmo-core 3 on October 1, 2026 as an open training stack for large mixture-of-experts models.

The redesign keeps experts resident on GPUs and routes token data to them instead of repeatedly gathering and resharding expert weights.

A benchmark expanded an expert pool from 8 to 128 while selecting four experts per token and keeping active parameters near 3.2 billion.

In that benchmark, total capacity grew from 4.6 billion to 47 billion parameters while throughput fell by less than 5 percent.

A preliminary eight-B300 comparison measured 52,000 tokens per second per GPU for the new stack and 19,400 for Ai2's earlier implementation.

A four-B300 test measured about 21 percent higher throughput with selected MXFP8 use than with BF16, while peak active memory fell from 103 GiB to 95 GiB.

Ai2 benchmarked a 1.2-trillion-parameter configuration with 58.36 billion active parameters per token across 512 B300 GPUs.

The trillion-parameter benchmark used random routing to test systems performance rather than trained model quality.

A separate 2.38-trillion-parameter DeepEP v2 configuration was a short capacity test, not a sustained training run.

Ai2 published code, documentation and the repository under the Apache 2.0 license.

02

WHY THIS MATTERS

Sparse models save computation only when routing and communication do not consume the advantage.

Keeping experts on GPUs changes the problem from repeated weight movement to efficient token movement.

Open training infrastructure exposes decisions that cannot be recovered from model weights alone.

Published negative results can save other researchers expensive cluster time.

Random-routing and capacity tests show what the system can exercise, not what a trained model can do.

Hardware-specific results may not transfer to clusters with different GPUs, links, drivers or kernels.

Open code improves inspection and adaptation without providing the compute needed to reproduce the largest run.

Load-balancing metrics can look healthy while actual expert queues become worse.

The release gives researchers a concrete system for studying routing, precision, memory and network bottlenecks at smaller scale.

FIG. 286HOW A SPARSE MODEL TURNS A TOKEN INTO A NETWORKING JOB
1SPLIT TEXT INTO TOKENS→
2SCORE AVAILABLE EXPERTS→
3SELECT A FEW EXPERTS→
4ROUTE TOKEN DATA ACROSS GPUS→
5RUN EXPERT COMPUTATION→
6RETURN AND COMBINE OUTPUTS→
7MEASURE LOAD AND WAIT TIME→
8UPDATE WEIGHTS→
9REPEAT WITHOUT CLOGGING THE LINKS
Only a few experts compute for each token, but the whole system must store the experts, move the data and prevent hot spots.

03

WHERE IT COULD HELP

  • Train a smaller mixture-of-experts model with only a few accelerators.
  • Compare expert-routing methods while measuring actual load distribution.
  • Profile the share of time spent computing, communicating, converting precision and waiting.
  • Port individual routing or parallelism components to another accelerator stack.
  • Test lower-precision training against a BF16 baseline on local hardware.
  • Reproduce token-gerrymandering behavior and design better balance measurements.
  • Teach distributed model training with a real open implementation.
  • Audit whether a large-parameter benchmark measures capacity, sustained training or model quality.
  • Build failure tests around uneven routing, dropped links and overloaded experts.
  • Keep hardware, software and input values fixed when comparing performance changes.

KEEP A HAND ON THE WHEEL

All performance and scale figures in the release were measured by Ai2 and are tied to specific Nvidia B300 configurations, software choices and benchmark conditions. The 1.2-trillion-parameter run used random routing and did not evaluate a trained model. The 2.38-trillion-parameter result was a short capacity test rather than sustained training. The cited material does not provide a full cost, power or emissions ledger for a complete trillion-parameter training run, and no independent reproduction of the largest configurations is cited. Open code does not make hundreds of current-generation GPUs, fast networking or systems expertise broadly available.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on October 2, 2026.

PUBLICATION RECEIPT: Original publication. Verified October 2, 2026 against Ai2's October 1 technical release, report link, interactive explanation and Apache 2.0 code repository.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US