THE SIGNAL IN ONE SENTENCE

Mantic, a London startup founded in 2024, has raised $25 million after its forecasting system finished ahead of every human participant in the Summer 2026 Metaculus Cup. It also finished behind one bot, called laertes, so the honest headline contains both facts. Mantic builds a system around frontier models from other laboratories, tests it on historical questions, grades the predictions, and tunes the process for medium-term forecasts about politics, economics, business, technology, conflict, and culture. That is a real performance result on a public tournament, not a crystal ball. A forecast is a probability attached to a precisely defined event before it resolves. Its quality depends on the question set, scoring rule, update timing, resolution source, participation requirements, model version, comparison group, and whether the probabilities remain calibrated across many outcomes. The plain signal is that AI forecasting has moved from a parlor trick toward a measurable technical competition. The next test is whether the same system improves decisions outside the tournament, where organizations must act before the outcome arrives and pay for being wrong in unequal ways.

01

WHAT ACTUALLY CHANGED

Reuters reported on September 18 that Mantic raised a $25 million seed round led by Radical Ventures, with participation from Microsoft venture fund M12, Thinking Machines Lab, Balderton Capital, and other investors. The valuation was not disclosed. Reuters says some companies and government agencies have integrated Mantic, citing Radical Ventures partner Aaron Rosenberg, who joined the startup's board. Mantic declined to identify customers, so there is no public customer list, contract value, deployment record, or measured decision outcome to inspect.

The funding follows the Summer 2026 Metaculus Cup, an online forecasting tournament that concluded in September. The published result, confirmed by Reuters and reflected in Mantic's tournament materials, placed Mantic ahead of every human participant and behind one bot named laertes. This means the system performed better on the tournament's scored question set under its rules. It does not mean Mantic predicted every event correctly or outperforms every person and system in every domain.

Mantic says its product specializes frontier AI models from other laboratories for judgmental forecasting. CEO and co-founder Toby Shevlane told Reuters that the company tests the system on historical data, grades its performance, and improves it. Mantic's site says the intended horizon is roughly one week to one year and the subject areas include geopolitics, business, policy, technology, and culture. The company has not publicly disclosed the complete model stack, training corpus, tournament prompts, inference budget, update schedule, or ablation tests showing which component produced the result.

Reuters highlighted two examples supplied by Shevlane. Mantic placed less weight than the human consensus on a claim about a Shakira song, and it assigned roughly 40 percent to Abelardo De La Espriella winning Colombia's presidential election when the consensus was near 30 percent. Those resolved examples illustrate useful disagreement. They do not establish the average quality of the system because any forecaster can look brilliant in a handpicked win and foolish in a handpicked miss.

Mantic markets probabilistic forecasts with reasoning and references, along with monitoring as new information arrives. This is a better shape than a one-word prediction because users can inspect how uncertain the system was and how its probability changed. Still, a rationale can sound excellent while the probability is poorly calibrated, the evidence is stale, or the event is defined in a way that does not match the decision a customer must make.

The money changes the company's capacity, not the meaning of the tournament. A seed round can fund engineering, data, evaluation, sales, and deployments. It cannot convert one leaderboard into proof of universal accuracy, business value, market profits, government competence, or safe automated decisions. Investors and the company use the word superhuman. The public evidence supports a narrower claim: Mantic beat the human participants in this tournament and nearly all of the bots.

02

WHY THIS MATTERS

Forecasting is one of the rare AI claims that can be scored against reality. A system states 70 percent before an event closes, the event later resolves, and the scoring rule rewards both accuracy and appropriately expressed uncertainty. That is healthier than a benchmark where the answer can be massaged after the model sees the test. It also creates new ways to game the comparison through question selection, update timing, hidden ensembles, participation thresholds, or favorable resolution rules.

Beating a field is not the same as being calibrated. If a forecaster assigns 70 percent to many events, roughly seven in ten should occur over time. A system can win a small tournament by making several bold calls that land and still be unreliable at 70 percent in a broader set. Calibration needs enough questions, several probability ranges, stable methods, and results reported over repeated tournaments, not one dramatic finish.

The baseline matters. Comparing Mantic with every registered human can produce a different story from comparing it with the best experienced forecasters, the crowd aggregate, a simple base-rate model, a prediction market, the strongest available bot, or a blended human-machine team. A useful report should publish all of those comparisons and show how performance changes by topic, forecast horizon, question type, and amount of research time.

Resolution is part of the model test. A question must define what counts, which source decides, when the window closes, and what happens if the event becomes ambiguous or the source changes. A forecaster can be penalized for misunderstanding the contract rather than misunderstanding the world. That is not a flaw if the contract was clear in advance. It is a serious flaw if resolution judgment drifts after predictions are locked.

A probability becomes valuable only when it changes an action. A government may prepare for an unlikely crisis because the damage would be enormous. A trader may ignore a likely event if the price already reflects it. A manufacturer may order extra parts only above a threshold set by delay costs and storage costs. The same 40 percent can sensibly produce three different decisions. Mantic needs evidence that its forecasts improve choices after costs, alternatives, and accountability are included.

Scale could be genuinely useful. A machine can maintain thousands of forecasts, update them as evidence changes, preserve version history, and expose disagreement that a committee quietly smoothed away. The risk is false precision arriving faster than anyone can inspect it. Organizations should treat the system as a disciplined probability generator with an audit trail, not an oracle that owns the decision.

FIG. 166TURN A LEADERBOARD WIN INTO DECISION EVIDENCE
1DEFINE THE QUESTION, DEADLINE AND RESOLUTION SOURCE→
2LOCK THE MODEL, INFORMATION WINDOW AND BASELINES→
3RECORD EACH PROBABILITY AND EVERY UPDATE→
4SCORE CALIBRATION, ACCURACY, COST AND STABILITY→
5TEST WHETHER THE FORECAST CHANGED A REAL DECISION
A tournament can prove who scored best on its board. A deployment must also prove that the probability was timely, reproducible, affordable, and useful.

03

WHERE IT COULD HELP

  • Publish a versioned forecast card for every question with its exact wording, open and close times, resolution source, initial probability, every update, cited evidence, model and tool versions, compute and research budget, final outcome, score, and any adjudication
  • Evaluate the system against several baselines at once, including the Metaculus community, top human forecasters, a simple base-rate model, a prediction market where available, a frontier-model prompt, the best competing bot, and a human-machine team
  • Report calibration by probability band, topic, horizon, question format, update frequency, and tournament, while preserving misses, withdrawals, ambiguous resolutions, inactive questions, and participation thresholds rather than publishing only overall rank
  • Connect probabilities to decision rules before the outcome is known by naming the action threshold, expected cost of each error, responsible decision-maker, override process, review date, and result of the action that was actually taken
  • Run prospective customer trials with frozen evaluation plans and independent scoring, then publish whether the forecasts improved timing, inventory, risk controls, policy preparation, research priorities, or returns after fees and operational costs

KEEP A HAND ON THE WHEEL

Mantic beat every human participant in the Summer 2026 Metaculus Cup and finished behind one bot named laertes. That is a strong tournament result, not proof of superior forecasting across all people, domains, time horizons, question sets, or deployment conditions. The public materials reviewed here do not disclose a complete model stack, inference budget, prompts, training and test separation, update policy, per-question results, error distribution, participation requirements, or customer outcome study. Mantic's example wins are selected illustrations rather than a complete error analysis. A leaderboard rank depends on the tournament's exact questions, scoring formula, timing, resolution rules, entrants, and missing-prediction treatment. The $25 million financing is a seed round at an undisclosed valuation. Investor interest, unnamed integrations, and demand from trading firms do not prove customer value or profitable forecasts. Superhuman is company and investor language. Watch for a complete Summer 2026 result export, repeated performance across future tournaments, calibration curves, strong and simple baselines, model-version history, compute and research costs, independent replication, named customer evaluations, and evidence that better probabilities caused better decisions.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on September 18, 2026.

PUBLICATION RECEIPT: Revision 1. Published September 18, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US