THE SIGNAL IN ONE SENTENCE
London startup Mantic has raised $25 million in seed funding after its system finished ahead of every human contestant in the Summer 2026 Metaculus Cup. Reuters reported the financing on September 18 and said Mantic placed behind only one other bot, called laertes. Metaculus records the tournament as closed, with 58 questions, a $5,000 prize pool and winners announced on September 5. A separate official question asking whether a bot would beat all humans resolved Yes. This is a real result, and it deserves more than a shrug. It is also narrower than the phrase superhuman forecasting can make it sound. Mantic did not predict the future with certainty. It assigned probabilities to clearly written questions, updated them before deadlines and was scored after outcomes resolved. The company says it improves frontier models from other laboratories by testing them on historical data, grading the forecasts and iterating. Reuters reported examples in music and Colombian politics where Mantic stayed closer to outcomes than the human consensus, and said the company has drawn interest from businesses and government agencies. Customers were not named, and no evidence of customer outcomes was published. The plain signal is that AI systems can now compete seriously in a structured craft once dominated by motivated human forecasters. The next test is transfer: whether the same discipline survives private data, vague executive questions, changing incentives, rare events and decisions where the customer can move the outcome it is trying to predict.
01
WHAT ACTUALLY CHANGED
Reuters reported on September 18 that Mantic raised $25 million in a seed round led by Radical Ventures. Microsoft venture fund M12, Thinking Machines Lab and Balderton Capital were among the other backers. The valuation was not disclosed.
Mantic was co-founded in 2024 by Toby Shevlane and Ben Day. Shevlane previously worked as a research scientist at Google DeepMind and told Reuters that the company specializes frontier models from other laboratories for forecasting rather than training a general frontier model from scratch.
The company says it tests its system on historical data, grades the resulting forecasts and improves the system through repeated evaluation. That describes a development loop. The public material does not provide the full training set, model versions, prompts, retrieval sources, compute budget or ablation studies needed to reproduce the result.
Metaculus lists the Summer 2026 Cup as closed, with 58 questions, a $5,000 prize pool and winners announced September 5. The platform describes itself as an online forecasting service focused on questions of global importance.
Reuters reported that Mantic outperformed every human contestant and all but one bot, called laertes. An official Metaculus question asking whether a bot would beat all humans resolved Yes and shows 129 forecasters. Together, those sources support bot superiority within this tournament, not across every forecasting task.
The contest covered political, economic and cultural developments. Reuters reported that Mantic did not lean as heavily as the human consensus against a Shakira song overtaking another track on the Billboard Hot 100, and assigned a higher early probability to Abelardo de la Espriella winning Colombia's presidential election. Both outcomes favored Mantic's positions.
Those examples illustrate a possible advantage without proving its cause. Shevlane attributed part of the performance to avoiding herd behavior. The published evidence does not isolate whether the gain came from model reasoning, broader search, faster updating, different priors, prompt design, question selection or simple variance across a finite set of forecasts.
A Radical Ventures partner told Reuters that companies and government agencies around the world had shown interest and that some had integrated Mantic. Hedge funds and trading firms were described as especially interested. Mantic declined to name customers, and no customer-level accuracy, decision or financial result was published.
The funding arrived after the contest result, but the result alone does not establish a durable commercial advantage. Investors are making a forward-looking bet on transfer, distribution and future performance. The tournament is historical evidence, not a warranty for a government briefing or trading desk.
02
WHY THIS MATTERS
Forecasting is an unusually useful AI test because it produces an answer before the world supplies the label. A model cannot quietly absorb the resolved outcome if the question closes first and the evaluation preserves the forecast history. The calendar creates a natural audit trail.
A probability is more useful than a confident paragraph. Saying an event has a 35 percent chance lets a decision-maker compare options, set a threshold and later check calibration. It also exposes uncertainty that fluent language can otherwise hide.
Proper scoring matters because a lucky call is not enough. A forecasting system should be rewarded for putting higher probability on events that happen and penalized for misplaced confidence. The exact rule and timing can change rankings, so the scoreboard is meaningful only with its evaluation design attached.
Resolution criteria do part of the intellectual work. A public tournament can define the event, deadline and deciding source in advance. Real organizations often ask foggier questions such as whether a supplier is risky or a policy will work. Turning those into resolvable forecasts may improve the decision before any model answers.
Beating individual humans does not necessarily beat a well-designed human institution. A strong workflow can combine model forecasts, domain experts, dissenting views and explicit decision thresholds. The relevant comparison is not machine versus person as a species contest. It is which process stays calibrated and catches its own blind spots.
Private deployments introduce information that the public contest did not test. Customer data may be incomplete, strategically manipulated, legally restricted or unlike the historical examples used for tuning. A model that shines on public questions can still fail when the evidence arrives in a messy spreadsheet five minutes before a meeting.
Decision-makers can also change the outcome. A company may cancel a launch because a forecast looks bad, making the forecast appear wrong even though it caused the intervention. Evaluation must distinguish prediction quality from the effects of acting on the prediction.
Rare, high-impact events are especially treacherous. Calibration requires many comparable cases, while wars, financial crises and technological discontinuities provide small, shifting samples. A system can be well calibrated on routine questions and still miss the event a government cares about most.
The venture funding matters because it moves probabilistic forecasting from a research and enthusiast niche toward a product category. That can spread disciplined uncertainty. It can also create incentives to market a bounded tournament win as general foresight before transfer evidence exists.
03
WHERE IT COULD HELP
- Rewrite strategic questions with a specific outcome, deadline, resolution source and update schedule before asking either people or models to forecast them
- Run the AI forecaster beside an existing human team for a fixed evaluation period without allowing either side to see the other's initial estimate
- Score forecasts with a published proper scoring rule and preserve every timestamped update instead of judging only the final headline call
- Measure calibration by probability band, topic, forecast horizon, geography, data source and model version rather than reporting one blended accuracy number
- Compare the system with strong human baselines, a simple historical base-rate model and the crowd consensus, not with a convenient novice
- Ask for a rationale and evidence map while keeping the probability as the scored output, since persuasive prose is not proof of predictive quality
- Set decision thresholds before seeing the model output and record which action each probability range is supposed to trigger
- Track cases where a forecast changes behavior so evaluators can separate predictive error from an outcome altered by intervention
- Create an abstention and escalation route for questions with weak data, unstable definitions, legal sensitivity or consequences too large for an automated recommendation
- Publish customer case studies with denominators, baselines, losses, decision changes and failures before claiming that a tournament advantage transfers to operations
KEEP A HAND ON THE WHEEL
The verified result is bounded. Mantic finished ahead of every human in one short-horizon Metaculus competition and behind one other bot, according to Reuters. Metaculus documents 58 questions and a resolved Yes to the proposition that a bot would beat all humans. That does not establish that Mantic is the best forecasting system, that bots will win every future cup or that performance transfers to private government, trading or corporate decisions. The company uses frontier models from other laboratories and has not published the full system, reproducible evaluation package or customer results. The reported integrations and investor enthusiasm are evidence of interest, not evidence of value created. Watch for the complete leaderboard and scoring methodology, per-question forecasts, calibration curves, time-stamped updates, comparisons with simple baselines, independent replication, model-version stability, performance on longer horizons, named customer studies, disclosure of failures and proof that decisions improved rather than merely receiving more precise-looking numbers.
04
TERMS WORTH KEEPING
OPEN GLOSSARY CARD
Probabilistic forecast
A prediction expressed as a probability instead of a certain yes, no, or single-point answer.
OPEN GLOSSARY CARD
Forecast calibration
The degree to which stated probabilities match observed frequencies across many resolved forecasts.
OPEN GLOSSARY CARD
Proper scoring rule
A scoring method designed to reward honest probability estimates by considering both the prediction and the resolved outcome.
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on September 20, 2026.
PUBLICATION RECEIPT: Revision 1. Published September 20, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US