Mantic funding has added $25 million for a London AI startup that trains and scaffolds frontier models to produce probability forecasts. Radical Ventures led the round, with M12, Balderton, Thinking Machines, DRW, FT Ventures, Episode 1 and individual investors participating, according to the company’s September 18 announcement.
The capital follows Mantic’s strong result in a 2026 Metaculus forecasting competition, where the company says its system outperformed all 676 human participants and most automated entrants. That benchmark is a useful signal, but the investment thesis depends on whether calibrated probabilities improve real decisions across finance, supply chains, strategy and government.
Everyone else is reporting an AI system that beat human forecasters; we are explaining the conversion from tournament score to decision infrastructure. A forecast product creates value only when it remains calibrated on unseen questions, communicates uncertainty honestly and changes an action before the outcome is known.
- Round: $25 million
- Lead: Radical Ventures
- The package separates verified facts from implementation claims.
How Mantic funding works
The funding facts are directly auditable in Mantic’s announcement, while Reuters independently interviewed co-founder Toby Shevlane and lead investor Radical Ventures. Because only one genuinely independent funding report was located, this package invokes the narrow central exception for primary plus one independent source. Investor names, amount and planned use are attributed to Mantic, and no valuation is asserted.
Mantic does not claim to have trained a frontier foundation model from scratch. Its system specialises models from other laboratories, tests forecasts against historical outcomes and improves the surrounding process. That can include question decomposition, evidence retrieval, aggregation and calibration. The competitive advantage may therefore live in the workflow and feedback loop rather than a single underlying model.
Forecasting is different from ordinary question answering. A useful answer must attach a probability to a clearly defined event and time horizon, then be scored after reality resolves the question. The discipline punishes confident errors and vague language. It also creates a measurable training signal: the company can compare forecasts with outcomes and identify where its process is systematically overconfident or slow.
The Metaculus result matters because human forecasters are a demanding comparison group rather than a random survey panel. Still, one competition can contain topic, timing and selection effects. Mantic will need results across multiple seasons, question categories and market conditions. Performance should be reported with proper scoring rules and compared with strong baselines, not with an undefined average person.
Mantic says the new money will expand its team and the compute and data supporting forecasts. More compute can allow additional model calls, scenario generation and ensemble methods, but cost discipline matters. A forecast that costs more to produce than the decision is worth will not scale. Customers need latency, reliability and pricing that fit their operating cadence.
Reuters reported early interest from hedge funds, trading firms, companies and government agencies. Those buyers have different error costs. A fund may value a small edge on a frequent market event, while a government agency may care about rare, high-impact risks. Mantic needs evaluation sets and access controls appropriate to each use case rather than one universal accuracy claim.
What changes next
The product’s interface is equally important. Executives rarely act on a bare percentage. They need the assumptions, evidence changes and scenarios that moved the forecast. A good system should show how sensitive the result is to uncertain inputs and when the probability was last updated. Explanations must help users challenge a forecast without pretending that every model step is fully interpretable.
Calibration is the core trust metric. If events assigned a 70 percent probability happen roughly seven times out of ten over a large sample, users can plan around that score. A system can rank outcomes correctly while still being badly calibrated. Mantic should publish reliability curves, question counts and confidence intervals so customers can distinguish signal from a short winning streak.
Data leakage is another risk. Historical testing can look impressive when the system indirectly accesses information published after the forecast deadline. Tournament infrastructure reduces that danger by fixing questions and resolution dates in advance, but enterprise back-testing needs equally strict time controls. Auditable data cutoffs should be part of every performance claim.
The commercial pathway resembles other specialised AI companies. Raindrop is building agent-testing infrastructure, while Factory AI is funding software-development agents and Polyphron is targeting biological models. Each narrows a general model into a measurable workflow where domain feedback can compound.
India relevance is strongest in operations exposed to volatile demand, commodity prices, logistics and policy change. An Indian enterprise could use probabilistic forecasts to trigger inventory ranges or hedging reviews, but local data and decision ownership remain essential. The model should support a human risk process rather than become an unchallengeable oracle.
Mantic also faces a governance question when forecasts concern elections, conflict or public policy. Publishing probabilities can affect behaviour and may be mistaken for certainty. Clear resolution criteria, source records and conflict-of-interest rules are necessary, especially when the same customer could benefit from a forecast becoming widely believed.
A practical post-funding scorecard should include out-of-sample calibration, improvement over model and human baselines, forecast cost, update latency, customer retention and documented decisions changed. Revenue alone would show demand, not predictive quality. Benchmark wins alone would show technical promise, not operational value.
Procurement teams should also separate model evaluation from business impact. A pilot can run forecasts in parallel with an existing planning process, record every probability before the outcome and compare decisions with a pre-agreed baseline. That design prevents teams from remembering only impressive calls. It also shows whether the tool changes inventory, hedging or resource allocation early enough to justify its cost.
Mantic will need disciplined versioning as underlying frontier models change. An upgrade can improve broad reasoning while damaging calibration in a particular domain. Customers should know which model and forecasting pipeline produced each probability, and historical scores should not mix incompatible versions. Reproducible evaluation is a product feature when forecasts influence material decisions.
In one sentence: Mantic funding buys the company time and compute to turn a strong forecasting benchmark into a repeatable decision product, and the next proof point is calibrated performance on unseen, customer-relevant events.
| Item | Verified detail |
|---|---|
| Round | $25 million |
| Lead | Radical Ventures |
| Disclosure | 18 September 2026 |
| Use | Team, compute and data |
| Benchmark context | Metaculus Cup |
Frequently asked questions
How much did Mantic raise?
Mantic announced a $25 million funding round led by Radical Ventures.
What does Mantic build?
Mantic specialises frontier AI models and supporting systems for probabilistic forecasts about political, economic and business events.
Who joined the round?
The company named Radical Ventures as lead and listed Balderton, Thinking Machines, DRW, FT Ventures, M12, Episode 1 and individual investors.
What must Mantic prove next?
It must show repeatable calibration on unseen events and useful decisions for customers, not only a strong result in one tournament.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



