THE SIGNAL IN ONE SENTENCE
LG AI Research has built a forecasting model for the series that make normal forecasting systems sweat: the spare part ordered twice a year, the medicine that suddenly disappears from shelves, the train platform that fills after a concert, and the winter power spike that was not visible last Tuesday. EXAONE Demand 1.0 is described in a September 25 preprint from the South Korean laboratory. The paper assembles 11,306,740 demand series containing 48,355,437,247 observations from 73 converted sources, adds synthetic data for rare demand patterns, and attaches a demand-specific adapter to a frozen general forecasting model. On 22 held-out datasets, both the full version and a synthetic-only version rank ahead of the 36 time-series foundation models the authors tested. That is a serious research result. It is also the beginning of the useful conversation, not the end. Demand is not simply a line waiting to be extended. A promotion can move the line. A holiday can bend it. A product launch can break it. A stockout can hide the very demand a retailer needs to estimate because sales fall to zero when the shelf is empty, not when customers stop wanting the item. The model only sees those outside events through the traces they leave in the history. LG's authors name that limitation directly. Their router reads the context window, not a calendar of promotions, prices and holidays. Stockout censoring is represented in the synthetic generator, but the system does not recover the hidden demand lost behind real out-of-stock sales. The missing promotion still wins because no amount of clever pattern recognition can read a fact that never entered the room. The paper starts by treating demand as its own species of time series. General forecasting models train across weather, finance, traffic, energy, health and many other domains. Actual demand can be a small part of that mixture. It is often short, full of zeros and divided across thousands of individual products, stores or services. A national electricity curve can be long and smooth. A replacement component for one machine can be mostly silence followed by one urgent order. Training one generic behavior across both can make the rare, commercially painful cases disappear inside the average. LG's corpus reaches across retail and e-commerce, energy, transport, healthcare, hospitality, telecommunications, economics, web traffic and computing. The team collected 93 sources and converted 73 into one schema. It excluded sources that were non-tabular, duplicated material already collected or required unavailable credentials. The converted collection has a heavily skewed shape. The lower quartile contains only 128 observations per series, while the median has 1,373. By the paper's classification, about 6.34 million series are smooth, 4.14 million are erratic, 727,598 are intermittent and 97,558 are lumpy. The difficult intermittent and lumpy cases remain a minority, so the synthetic generator deliberately produces more of them. The four labels come from established demand-forecasting practice. Smooth demand appears often and varies relatively little. Erratic demand appears often but jumps in size. Intermittent demand arrives infrequently with more stable sizes. Lumpy demand is both infrequent and highly variable, which is forecasting's polite term for a series that likes to throw furniture. EXAONE Demand does not force every series into one hard box. It computes soft membership across the four behaviors, then teaches a small router to produce weights from eight scale-free statistics. A series near the boundary can be partly intermittent and partly lumpy instead of receiving a brittle label. The adapter has one shared low-rank branch that is always active and four routed branches tied to the demand classes. Those branches modify a frozen general-domain backbone without retraining all of its weights. Each layer can mix the branches differently for the same series. The result is specialization without building four separate forecasting models and a fifth one to choose among them. The training data is also handled with more care than a giant-number headline suggests. Published time-series collections often recycle the same underlying series under different names. A model can look excellent if a renamed copy of its test set appeared during training. LG compares raw values, not merely dataset titles, and removes matching training candidates. The paper reports 358,008 series removed from the training side through that leakage check. The held-out side contains 899,618 series from evaluation suites. The authors also state an important asymmetry: competing models were pretrained before these benchmarks and their original training data was not necessarily filtered with the same procedure. That does not invalidate the comparison, but it belongs in the receipt. LG reran 36 foundation-model baselines and the frozen EXAONE backbone under one protocol instead of copying favorable numbers from other papers. The list includes versions of Chronos, TimesFM, Moirai, TiRex, Sundial, Toto, TabPFN-TS, Granite, YingLong and others. No baseline was selected according to its performance on the 22 datasets, and the paper says no system was tuned on this evaluation suite. EXAONE Demand ranks first in the main table with an average rank of 4.09 across datasets, a 91.9 percent pairwise win rate and a mean absolute scaled error, or MASE, of 1.0667. The synthetic-only version ranks second with an average rank of 5.55, an 88.0 percent win rate and MASE of 1.0742. TiRex-1.1 ranks third in the table. The full version also reports lower normalized deviation, weighted quantile loss, percentage error and absolute error than the synthetic-only version, while the synthetic-only version has a slightly lower mean scaled interval score. In other words, one neat winner depends on which part of the forecast a decision actually needs. Then comes the number that should be taped above every procurement meeting. The paper explains that MASE compares a model with a seasonal naive forecast. A value above one means repeating the last season would have been the better option. Every row in the reported comparison, including EXAONE Demand, sits above one on the present evaluation suite. LG's model is the best-performing foundation model in that table while still losing, on the aggregate MASE interpretation supplied by the authors, to an extremely simple baseline. That is not a gotcha. It is excellent scientific context. A new model can improve the frontier among its peers and still fail the cheapest test a planner should run. The useful deployment question is not whether EXAONE beat TimesFM. It is whether it beats last season, a store-specific statistical model and the current human process after every cost is counted. The model trained with real and synthetic demand performs better than the synthetic-only version on 16 of the 22 evaluation datasets and worse on six. The authors report that the gains on helped datasets outweigh the losses on the others, which supports the aggregate advantage of real data. It also warns against buying one universal score. A retailer, grid operator, transit agency and pharmacy do not share the same error costs. Underestimating insulin demand is not the mirror image of ordering too many seasonal decorations. A forecast can have lower average error while making the expensive mistake more often. EXAONE Demand emits nine quantiles, which can express a distribution instead of one point. That is useful because a planner can see a likely range and choose a service level. The paper evaluates weighted quantile loss and the quality of an 80 percent interval alongside point metrics. Production teams still have to connect those distributions to consequences. What is the cost of a stockout? What expires on the shelf? How quickly can supply respond? Which customers are harmed first? What level of reserve is legally required? A probability band becomes a decision only after someone writes the loss function in human terms. The paper includes six selected forecast plots where EXAONE Demand follows collapses to zero or sharp peaks more closely than two strong baselines. The authors explicitly say those panels were chosen to show the adaptation when it works and that selected curves cannot carry the aggregate claim. That sentence should be standard equipment in every AI release. A pretty chart is an example. A held-out protocol is evidence. A live operating result is something else again. Nothing in the preprint establishes that EXAONE Demand has reduced waste, prevented a stockout, improved staffing, cut reserve margins or saved money in an LG business or another South Korean organization. No production deployment, intervention study, service-level comparison, planner workflow, operating cost or independent replication is presented. The public evidence is a company-authored preprint and benchmark package. That is enough to take the method seriously. It is not enough to let the model order the warehouse. The first practical deployment should run in shadow mode. Feed the model the same historical series available to planners, freeze each forecast before the outcome arrives and compare it with the seasonal naive baseline, the current production system and human adjustments. Keep promotion, holiday, price, stockout and lifecycle data separate at first so the team can see what the univariate model misses. Score by product, location, horizon and demand class. Report the mean and the tail. A model that trims average error but creates catastrophic misses on scarce medicines, transformer parts or peak electricity days is not an improvement. Next, record every override. If a planner raises the forecast because a campaign begins Friday, capture the reason and the source of the information. If the model beats the override, keep that too. Over time, the override log becomes the covariate roadmap the paper already points toward. Promotions, holidays and prices can be added as explicit features only after the organization proves their definitions, availability and timing. A promotion entered after sales begin is not a forecast input. It is hindsight wearing an employee badge. Stockouts need their own treatment. Observed sales are censored when inventory reaches zero. Training on those zeros can teach a model that customers wanted nothing precisely when demand exceeded supply. Teams need inventory availability, unfilled orders, substitutions, wait lists and replenishment timing to estimate lost demand. The paper leaves this recovery problem open. A production pilot should label it openly instead of replacing every zero with a guess. Sensitive domains also need access and purpose controls. Pharmacy demand, transportation usage and household energy can reveal patterns about communities and individuals even when the forecasting target looks operational. The paper combines open datasets across domains, but an adopter still has to inventory its own data, minimize identifiers, set retention periods and restrict who can drill from a forecast into the underlying series. More data is not automatically a better reason to collect it. South Korea has a strong practical reason to care about this work. Manufacturing, logistics, retail, transport, energy and healthcare all depend on forecasts made across enormous catalogs and networks. A domestic research lab building a reusable demand model can lower the effort required to test specialized forecasting across those sectors. The significance beyond Korea is the architecture: build a domain corpus, verify source counts, remove value-level leakage, preserve a frozen backbone and route small adapters according to interpretable characteristics. That recipe can travel. The benchmark victory should travel with its caveats. The plain signal is that EXAONE Demand makes demand look less like one generic line and more like the family of awkward behaviors planners actually meet. Its corpus is large, its routing idea is readable and its evaluation is unusually explicit. The paper also tells us the model cannot see a promotion that is not recorded, does not recover hidden demand from real stockouts and remains above the seasonal naive benchmark on aggregate MASE. That honesty makes the work more useful, not less. Put it beside the planner, not over the planner. Make it beat last season before asking it to beat the world.
01
WHAT ACTUALLY CHANGED
LG AI Research published the EXAONE Demand 1.0 preprint on September 25.
The project collected 93 sources and converted 73 into a common demand-series schema.
The resulting corpus contains 11,306,740 series and 48,355,437,247 observations.
The lower quartile has 128 observations per series and the median has 1,373.
The corpus includes retail, energy, transport, healthcare, hospitality, telecommunications, economics, web and computing demand.
The authors classify demand behavior as smooth, intermittent, erratic or lumpy.
A synthetic generator adds more of the rare intermittent and lumpy patterns.
A router reads eight scale-free statistics and produces soft weights across four demand-specific branches.
One shared low-rank branch remains active while the four specialized branches are mixed by the router.
The general-domain backbone remains frozen during adapter training.
The team removed 358,008 training series after checking for value-level overlap with evaluation data.
The evaluation covers 22 held-out datasets and 36 time-series foundation models plus the frozen backbone.
EXAONE Demand ranks first with a 4.09 average rank, a 91.9 percent pairwise win rate and MASE of 1.0667.
The synthetic-only version ranks second with a 5.55 average rank, an 88.0 percent win rate and MASE of 1.0742.
The version using real and synthetic demand improves on the synthetic-only version on 16 datasets and worsens on six.
The paper says every compared row has MASE above one, meaning a seasonal naive forecast performs better on that aggregate interpretation.
Promotions, holidays and price changes are not supplied as explicit external variables.
Recovery of hidden demand from real stockout-censored sales remains future work.
02
WHY THIS MATTERS
Demand series are often shorter, sparser and more irregular than the data dominating general forecasting corpora.
A domain-specific corpus can preserve rare patterns that disappear inside a general average.
Soft routing avoids pretending that every demand series belongs cleanly to one behavior class.
Low-rank adapters can specialize a frozen model without retraining every backbone parameter.
Raw-value leakage checks are stronger than comparing dataset names when public collections redistribute the same series.
A first-place model comparison does not prove superiority to a simple seasonal baseline.
Average forecast error is not the same as inventory cost, service level, waste or public benefit.
Under-forecasting and over-forecasting carry different costs in retail, healthcare, energy and transport.
A forecast distribution is more useful than one point only when planners connect uncertainty to a decision rule.
A missing promotion or holiday cannot be recovered reliably from the historical series before its effect appears.
Sales falling to zero during a stockout can hide strong customer demand.
Synthetic censoring does not establish recovery of latent demand in a live supply chain.
Results that help 16 datasets and hurt six should be examined by domain rather than sold as one universal win.
Selected forecast plots can illustrate behavior but cannot substitute for the complete held-out evaluation.
Company-run benchmarks need independent replication and shadow deployment before operational authority expands.
Human overrides can supply valuable event context when their reasons and timing are logged.
Operational demand data can expose sensitive community or customer patterns even when names are removed.
The architecture offers a reusable recipe for specialization beyond South Korea, but the decision evidence must travel with it.
03
WHERE IT COULD HELP
- Run EXAONE Demand in shadow mode beside the current forecasting process before it influences orders or staffing.
- Compare it with a seasonal naive forecast, a simple local statistical model and the existing human process.
- Freeze every forecast before the outcome and preserve the model, data version, horizon and quantiles.
- Score results separately by product, location, horizon and demand class.
- Measure service levels, shortages, waste, expiration, reserve margin and cost instead of relying only on average error.
- Choose asymmetric loss functions that reflect the real difference between too much and too little supply.
- Record every planner override with its reason, source and timestamp.
- Use the override ledger to identify promotions, holidays, prices and lifecycle events worth adding as covariates.
- Keep late-entered event data out of historical backtests to prevent hindsight leakage.
- Track inventory availability, unfilled orders, substitutions and replenishment timing to estimate stockout-censored demand.
- Do not replace every zero with an inferred value unless the inference is labeled and evaluated.
- Evaluate quantile calibration so stated demand ranges match observed frequencies.
- Stress-test abrupt launches, closures, disasters, policy changes and supplier interruptions.
- Create domain-specific safety rules for medicines, grid reserves and other high-consequence demand.
- Require human approval before a forecast changes a consequential order, schedule or public-service allocation.
- Audit training licenses, source provenance, retention and access for every local dataset.
- Publish the failure slices and operating cost alongside the headline rank.
- Invite independent teams to reproduce the 22-dataset protocol and the seasonal-naive comparison.
KEEP A HAND ON THE WHEEL
EXAONE Demand 1.0 is described in a company-authored preprint, not an independently reviewed production study. LG AI Research ran the 36 comparison models and the frozen backbone under its own protocol. The 22 held-out datasets support a broad benchmark comparison but do not establish fewer stockouts, lower waste, better staffing, safer reserves or financial return in a live organization. The paper reports a first-place average rank and 91.9 percent pairwise win rate, while also reporting MASE of 1.0667. Its own metric explanation says a value above one means a seasonal naive forecast would have been better, and every model row is above one on this suite. Real and synthetic training improves 16 datasets and worsens six relative to synthetic-only training. Promotions, holidays and prices reach the model only through their historical trace. Real lost demand behind stockouts is not recovered, and the four demand classes and routed branches are fixed rather than learned. The selected forecast plots are explicitly illustrative. Watch for independent replication, released weights or code, model and serving cost, country and sector deployments, calibration by demand class, comparison with simple local baselines, covariate-aware routing, stockout recovery, planner override evidence and actual changes in service, waste and cost.
04
TERMS WORTH KEEPING
OPEN GLOSSARY CARD
Demand forecasting
Estimating how much of a product or service will be wanted during a future period.
OPEN GLOSSARY CARD
Time-series foundation model
A forecasting model pretrained on many sequences so it can make predictions for new series with little or no task-specific training.
OPEN GLOSSARY CARD
Intermittent demand
Demand that occurs infrequently, leaving many zero periods between nonzero orders or uses.
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on September 28, 2026.
PUBLICATION RECEIPT: Revision 1. Published September 28, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US