THE SIGNAL IN ONE SENTENCE

A small language model can fit on a laptop and still become sluggish when the conversation gets long. The reason is not mysterious. In a conventional transformer, the model keeps a key-value cache for earlier tokens so later tokens can attend to them. As the context grows, that cache grows too. Each new token can require more memory traffic, and memory traffic is often the practical speed limit for generation on a personal device. Sevren, a research lab in Copenhagen, has published an experimental model that attacks that bottleneck directly. Heimr 570M has 2.7 billion parameters in total, but the lab says only about 570 million are active for each token. It was trained for a context window of 64,000 tokens and combines sixteen Mamba-3 state-space layers with four attention layers. The model's feed-forward blocks use a mixture of experts, routing each token through four of 32 experts plus one shared expert. That is a lot of architecture language for one fairly plain idea. Most layers carry a fixed-size running state instead of rereading an ever-growing history. A small number of attention layers still look back over the context, but they add much less cache than a full-attention model with attention in every layer. The model also activates only part of its expert network for each token. The plain signal is that a local model can keep a long memory without paying the full transformer tax on every new word. Sevren reports that Heimr decodes two to eight times faster than comparable open models on long contexts. In one specific test at 64K context, using 4-bit weights on an Apple M5 Pro, the lab reports a 14 times advantage over Qwen3-0.6B, 2.4 times over Qwen3.5-0.8B and 2.2 times over Gemma 3 1B. Those are developer measurements, not independent results. They also describe one part of performance: single-stream decode speed on stated hardware with quantized weights. They do not tell us how quickly the system ingests a new document, how much memory the complete application consumes, how it behaves on a different processor or whether its answers are useful for a particular task. Speed is not quality. A model that produces a wrong answer quickly remains a fast wrong-answer machine. Sevren did publish a broad base-model evaluation. It ran Heimr and four comparison models through the same harness on the full 91,037 examples in the DataComp-LM CORE suite. The lab reports a CORE score of 0.411 for Heimr, ahead of scores from 0.345 to 0.375 for the comparison models. The per-task table is more revealing than the headline average. Heimr leads several tasks and trails on others. Its reported BoolQ, CommonsenseQA, SQuAD and repeat-copy results are below at least one comparison model. The technical note also discloses that part of its pretraining mix was built against three tasks in the evaluation. When those tasks are removed, the reported lead over the two Qwen models narrows. That disclosure is useful. It is also a reminder that a benchmark score is not a transferable guarantee. The model was trained on 200 billion tokens in two phases. Sevren says the first 160 billion used a 4K context and distillation from Mistral 7B. The final 40 billion used a 64K context on a long-document mix without the teacher. The company reports using 16 H100 GPUs throughout training. Distillation lets a smaller model learn from the output distribution of a larger one. It can improve a compact model without recreating the larger model's full size. It can also carry forward the teacher's blind spots, data assumptions and uneven capabilities. The total and active parameter counts need care too. Calling Heimr a 570M model describes the approximate number of parameters used per token, not the full storage footprint. The model has 2.7 billion total parameters because the expert network contains many components that are not all active at once. Sparse activation can reduce computation per token, but a device still needs access to the full set of weights or an implementation that manages them efficiently. That distinction matters for the word local. Local inference can keep documents on a person's device, reduce network dependence and avoid per-request cloud latency. It can help a laptop summarize a private folder, classify records, extract fields, search notes or support an offline workflow. It can also make useful language tools possible where connectivity is expensive or unreliable. But local does not automatically mean private. An application can run the model locally while still sending telemetry, prompts or results elsewhere. It can store unencrypted transcripts. It can download unverified model files. It can give another application broad access to personal documents. Privacy depends on the complete product, not only the location of matrix multiplication. Local also does not mean energy-free. Moving fewer bytes can reduce delay and power use for one operation, but the real energy cost depends on hardware, context length, token count, quantization, cooling and how often the model runs. Sevren's technical note does not publish a full energy comparison. The strongest part of the Heimr release is that the performance claim is connected to an architectural reason. Sixteen of twenty layers carry a fixed-size state. Only four use attention that grows with context. On a memory-bandwidth-bound device, keeping the bytes read per token closer to constant should protect decode speed as the history expands. That is a testable claim, not magic dust. The next useful tests should separate the pieces. Measure prompt ingestion and generation independently. Report time to first token, tokens per second at several context lengths, peak memory and energy per completed task. Use multiple consumer chips, not only one high-end configuration. Compare equal quantization, batch size and output length. Publish the exact commands and model files. Then ask independent teams to reproduce the curve. Long-context quality needs its own evaluation. A model may accept 64,000 tokens without reliably using information from the beginning. It may retrieve one detail while losing relationships across a document. It may handle prose but fail on tables, code or multilingual material. Needle-in-a-haystack tests can check retrieval, but real workflows need questions that combine evidence across the context and expose unsupported answers. For European developers, the broader attraction is control. Sevren presents Heimr as part of a programme for models that can run on phones and laptops, and frames the work around European technical capacity. A locally runnable model with documented architecture could let companies and public institutions keep more of their workflow under their own operational rules. That possibility is larger than this experiment, but the experiment does not prove it by itself. The public technical note documents training, architecture, speed tests and benchmark results. The lab's website says its methods, weights and reference code are intended to ship openly. The Heimr article, however, does not clearly provide a downloadable Heimr 570M checkpoint, license or reproducible release package. Readers should distinguish an open research explanation from a model they can independently run today. The useful move now is boring in the best way. Release the exact weights, tokenizer, code, license, evaluation harness, quantization recipe and hardware settings. Let other people reproduce the speed curve, test different chips and find the tasks where quality breaks. Measure privacy and energy at the product level. Compare against strong simple baselines. Heimr's architecture suggests that small local models do not have to choose between a long context and tolerable generation speed. That is an interesting signal from Copenhagen. It becomes a dependable result when the rest of the world can run the same experiment and get the same answer.

01

WHAT ACTUALLY CHANGED

Copenhagen research lab Sevren published its experimental Heimr 570M model architecture and measurements on September 30.

The model has 2.7 billion total parameters, with about 570 million active for each token through sparse expert routing.

Its twenty-layer architecture combines sixteen Mamba-3 state-space layers with four Heimr attention layers.

Sevren trained the model for a 64K context on 200 billion tokens using 16 H100 GPUs across two training phases.

The lab reports two to eight times faster long-context decoding than comparable open models in its tests.

At 64K context on an Apple M5 Pro with 4-bit weights, Sevren reports advantages ranging from 2.2 times to 14 times against the selected comparison models.

Using one common harness on 91,037 CORE examples, the company reports an aggregate score of 0.411, with uneven results across individual tasks.

The release is a technical research note for an experimental model, not independent validation or a documented production deployment.

02

WHY THIS MATTERS

Long contexts can slow conventional transformers because the key-value cache and memory traffic grow with the history.

A mostly fixed-state architecture can keep generation speed steadier as the context expands.

Sparse expert routing uses only part of a larger model for each token, separating active computation from total model size.

Faster local inference could support private-document workflows, offline tools and lower-latency assistance on personal devices.

The distinction between 570 million active parameters and 2.7 billion total parameters matters for storage, memory and honest comparisons.

Company benchmarks are useful early evidence, but hardware, quantization, kernels and task selection can materially change the result.

Accepting a 64K input does not prove reliable reasoning across all 64,000 tokens.

Open methods help European technical capacity only when the exact artifacts and license allow other teams to reproduce and extend the work.

FIG. 301KEEP THE LONG HISTORY, LIMIT THE GROWING CACHE
1READ THE PROMPT→
2UPDATE A FIXED STATE→
3OPEN AN ATTENTION GATE→
4ROUTE TO SELECTED EXPERTS→
5GENERATE ONE TOKEN→
6REPEAT AS CONTEXT GROWS→
7MEASURE SPEED AND QUALITY→
8REPRODUCE ON ANOTHER DEVICE
Most Heimr layers carry a fixed-size state, while four attention layers and selected experts preserve some look-back and capacity. The architecture can reduce memory traffic, but the measured advantage still needs reproduction.

03

WHERE IT COULD HELP

  • Summarize or search long private documents on a laptop after confirming the full application keeps data local.
  • Extract fields from contracts, reports and manuals without sending every page to a cloud service.
  • Support offline or low-connectivity assistants where network latency and availability are practical constraints.
  • Prototype a local coding or research tool that needs a long working context but modest per-token computation.
  • Compare state-space and attention layers under identical hardware, quantization and context settings.
  • Measure time to first token, decode speed, peak memory and energy separately instead of collapsing them into one speed claim.
  • Test retrieval and multi-step reasoning at many positions inside a 64K context.
  • Audit whether a supposedly local product sends telemetry, prompts, documents or outputs to another service.
  • Reproduce the CORE results with the published harness and report task-level variation, not only the aggregate score.
  • Evaluate smaller models against the exact business task before treating model size or context length as a purchasing shortcut.

KEEP A HAND ON THE WHEEL

Heimr 570M is explicitly experimental. The speed and quality figures were produced by Sevren, the model's developer, on selected hardware, quantization settings, kernels and evaluation tasks. They have not been independently reproduced here. The 570M name refers to active parameters per token, while the model contains 2.7 billion parameters in total. A 64K trained context does not establish reliable use of every token, and the public note does not report production reliability, prompt-ingestion speed, complete memory use, energy per task, safety behavior or performance across a broad range of consumer hardware. The article also does not clearly link a downloadable checkpoint, license and end-to-end reproduction package. Watch for exact weights and code, license terms, independent speed and CORE replications, long-context retrieval and reasoning tests, multilingual results, memory and energy measurements and evidence from real local applications.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on Verified against Sevren primary sources on October 3, 2026.

PUBLICATION RECEIPT: Original reporting and analysis. Performance and benchmark results are attributed to Sevren and are not presented as independently reproduced.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US