Are prediction markets accurate? What the data says
By insiderz6 min read

Prediction markets are accurate where they are deep and short dated, and unreliable where they are not. On a cross platform sample of resolved markets measured to 14 October 2025, Polymarket scored a Brier loss of 0.1652 and Kalshi 0.1982, against 0.25 for always guessing 50 percent. On the 2024 US election specifically, one study found that only 67 percent of Polymarket markets priced the eventual winner above 50 percent.
What do the published accuracy numbers look like?
Brier score is the standard measure. You square the gap between the probability you gave and the outcome, coded 1 or 0, and average. Lower is better. Saying 50 percent to everything gives 0.25. The arithmetic is in Brier score explained.
The one public cross platform comparison that publishes its sample sizes is brier.fyi, which scores resolved binary and multiple choice markets across four venues.
| Platform | Money required | Overall Brier score | Markets in sample | Data as of |
|---|---|---|---|---|
| Polymarket | Yes | 0.1652 | 259 | 14 October 2025 |
| Metaculus | No | 0.1664 | 153 | 14 October 2025 |
| Kalshi | Yes | 0.1982 | 183 | 14 October 2025 |
| Manifold | No | 0.2118 | 376 | 14 October 2025 |
| insiderz | No | No resolved record yet | 0 | 4 September 2026 |
Figures from brier.fyi. The samples are small and not matched question for question, so treat the ordering as suggestive rather than settled. The honest summary is that all four sit in a band between roughly 0.16 and 0.21, and that the play money platform Metaculus is inside that band, not outside it.
insiderz launched in September 2026 and has no resolved history of its own. Any accuracy claim about insiderz today would be invented, so there is none on this page. The counts that do exist are below.
insiderz right now
Live from insiderz.- calls
- 0
- insiders
- 0
- resolved events
- 0
- open events
- 1,904
How is accuracy actually measured?
Three different things get called accuracy and they are not the same.
- Hit rate. What share of markets had the eventual winner priced above 50 percent. Easy to state, but it ignores confidence entirely.
- Brier score. Average squared error of the probability. Rewards confidence when it is justified and punishes it when it is not.
- Calibration. Whether claims made at 70 percent come true about 70 percent of the time. A market can have a good hit rate and still be badly calibrated.
Most headline claims about prediction markets use hit rate, because it produces the biggest number. When you see "94 percent accurate", check whether that means a Brier score or a count of favorites that won.
What did the 2024 election studies actually find?
Two findings pull in opposite directions and both are real.
The critical one comes from Joshua D. Clinton and TzuFeng Huang at Vanderbilt, who analysed more than 2,500 political markets across the Iowa Electronic Markets, Kalshi, PredictIt and Polymarket in the final five weeks of the 2024 campaign. They found 93 percent of PredictIt markets, 78 percent of Kalshi markets and 67 percent of Polymarket markets priced the eventual outcome better than chance, and reported little evidence of efficiency: prices for identical contracts diverged across exchanges, daily price changes were weakly correlated or negatively autocorrelated, and arbitrage opportunities peaked in the final two weeks (Clinton and Huang, SocArXiv, 2025). The largest venue scored worst. Reporting on the study noted that Kalshi disputed the methodology, arguing calibration rather than hit rate is the right measure (DL News, 5 December 2025).
The favorable one comes from a study of Polymarket's own microstructure over 5 January to 6 November 2024. Price sensitivity to order flow fell from about 0.53 at the end of July 2024 to about 0.01 by October, and the half life of yes plus no pricing deviations fell from several hours in early 2024 to well under a minute by October and November (Tsang and Yang, arXiv:2603.03136, August 2026). By the end, the headline market was hard to push around and quick to close its own gaps.
Both can be true. A deep flagship market can behave well while the two thousand markets around it behave badly.
Where are prediction markets weakest?
Four conditions predict a bad price, and they compound.
- Low volume. Thin markets are one person's opinion with a price attached.
- Long horizon. The further from resolution, the more prices sit near 50.
- Ambiguous wording. The market settles on the written rule, not the headline.
- Contested resolution. Where the rule can be argued, the price includes dispute risk. See how prediction markets resolve.
The horizon effect has been measured. Across 353 million trades and 429,000 binary contracts on Kalshi and Polymarket, political markets showed persistent underconfidence, with prices compressed toward 50 percent: a contract at 70 cents a month out corresponded to a true probability closer to 75 percent. The model explained 87.3 percent of calibration variance in sample on Kalshi and 71.5 percent out of sample (Le, arXiv:2602.19520, February 2026). The important word is structured. The error is not noise, it has a direction, and a direction is something a forecaster can exploit.
How often does an individual beat the market?
Often enough to be worth measuring, rarely enough that one example proves nothing.
The best public benchmark for individuals is ForecastBench, which put the same unresolved questions to superforecasters, the general public and large language models. On the human question set, superforecasters reached a mean Brier score of 0.096, against 0.121 for the general public and 0.122 for the best model tested, with both gaps significant at p below 0.001 (Karger et al., arXiv:2409.19839, revised February 2025). Skilled individuals really do separate from the crowd, and they separate by a measurable amount.
That is a different claim from beating a specific market price on a specific event. To show that, you need your probability, the market's probability at the same instant, and a resolved outcome, repeated across dozens of events. Anything less is a story.
Method: how this page was built, and when it updates
- Scope. Published accuracy studies of Polymarket, Kalshi, PredictIt, Metaculus and Manifold, plus benchmark results for human forecasters. No affiliate sources, no platform marketing pages.
- Metric preference. Brier score first, calibration second, hit rate last and always labelled as such.
- Sample sizes. Every number on this page carries its sample size and its date, or it is not on this page.
- Excluded. Claims of the form "X percent accurate" with no sample, no date and no metric definition. There are many of them.
- Cadence. This page is reviewed at least every 90 days and after any major election resolves. Last checked 4 September 2026.
- What insiderz adds later. Calls on insiderz are locked with the time and the Polymarket price at that moment, so once events resolve there is a clean comparison of a person's probability against the market's probability at the same instant. That data does not exist yet in useful quantity, so it is not published here yet.
Open events are on Events, and the people ranked on Beats market, Events, Edge and Early are on the leaderboard.
Questions people ask
- How accurate is Polymarket?
- On a cross platform sample of resolved markets measured to 14 October 2025, Polymarket's overall Brier score was 0.1652 across 259 markets, the best of the four platforms compared. Accuracy drops sharply on low volume and long horizon markets.
- Are prediction markets better than polls?
- On short horizon elections, usually yes, because they aggregate polls plus everything else. On long horizon or technical questions the edge shrinks.
- What is a good Brier score?
- Around 0.25 is a coin flip. Superforecasters reached a mean of 0.096 on the ForecastBench human question set. Below 0.10 is elite.
- Do prediction market prices efficiently aggregate information?
- Not always. A study of more than 2,500 markets in the final five weeks of the 2024 US campaign found prices for identical contracts diverging across exchanges and arbitrage opportunities persisting to election day.
Sources
- Prediction Markets? The Accuracy and Efficiency of $2.4 Billion in the 2024 Presidential Election, Joshua D. Clinton and TzuFeng Huang, SocArXiv, 2025
- Are Polymarket and Kalshi as reliable as they say? Not quite, study warns, DL News, 5 December 2025
- Prediction market accuracy by platform, brier.fyi, data updated 14 October 2025
- ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities, Karger et al., arXiv:2409.19839, revised 28 February 2025
- Decomposing Crowd Wisdom: Domain-Specific Calibration Dynamics in Prediction Markets, Nam Anh Le, arXiv:2602.19520, February 2026
- The Anatomy of a Blockchain Prediction Market: Polymarket in the 2024 U.S. Presidential Election, Kwok Ping Tsang and Zichao Yang, arXiv:2603.03136, August 2026



