Connect wallet

Are prediction markets accurate? What the data says

By insiderz6 min read

Abstract flat illustration on a dark background of a calibration curve, a diagonal reference line with a slightly bowed plotted line and scattered points around it

Prediction markets are accurate where they are deep and short dated, and unreliable where they are not. On a cross platform sample of resolved markets measured to 14 October 2025, Polymarket scored a Brier loss of 0.1652 and Kalshi 0.1982, against 0.25 for always guessing 50 percent. On the 2024 US election specifically, one study found that only 67 percent of Polymarket markets priced the eventual winner above 50 percent.

What do the published accuracy numbers look like?

Brier score is the standard measure. You square the gap between the probability you gave and the outcome, coded 1 or 0, and average. Lower is better. Saying 50 percent to everything gives 0.25. The arithmetic is in Brier score explained.

The one public cross platform comparison that publishes its sample sizes is brier.fyi, which scores resolved binary and multiple choice markets across four venues.

Platform Money required Overall Brier score Markets in sample Data as of
Polymarket Yes 0.1652 259 14 October 2025
Metaculus No 0.1664 153 14 October 2025
Kalshi Yes 0.1982 183 14 October 2025
Manifold No 0.2118 376 14 October 2025
insiderz No No resolved record yet 0 4 September 2026

Figures from brier.fyi. The samples are small and not matched question for question, so treat the ordering as suggestive rather than settled. The honest summary is that all four sit in a band between roughly 0.16 and 0.21, and that the play money platform Metaculus is inside that band, not outside it.

insiderz launched in September 2026 and has no resolved history of its own. Any accuracy claim about insiderz today would be invented, so there is none on this page. The counts that do exist are below.

insiderz right now

Live from insiderz.
calls
0
insiders
0
resolved events
0
open events
1,904

How is accuracy actually measured?

Three different things get called accuracy and they are not the same.

  • Hit rate. What share of markets had the eventual winner priced above 50 percent. Easy to state, but it ignores confidence entirely.
  • Brier score. Average squared error of the probability. Rewards confidence when it is justified and punishes it when it is not.
  • Calibration. Whether claims made at 70 percent come true about 70 percent of the time. A market can have a good hit rate and still be badly calibrated.

Most headline claims about prediction markets use hit rate, because it produces the biggest number. When you see "94 percent accurate", check whether that means a Brier score or a count of favorites that won.

What did the 2024 election studies actually find?

Two findings pull in opposite directions and both are real.

The critical one comes from Joshua D. Clinton and TzuFeng Huang at Vanderbilt, who analysed more than 2,500 political markets across the Iowa Electronic Markets, Kalshi, PredictIt and Polymarket in the final five weeks of the 2024 campaign. They found 93 percent of PredictIt markets, 78 percent of Kalshi markets and 67 percent of Polymarket markets priced the eventual outcome better than chance, and reported little evidence of efficiency: prices for identical contracts diverged across exchanges, daily price changes were weakly correlated or negatively autocorrelated, and arbitrage opportunities peaked in the final two weeks (Clinton and Huang, SocArXiv, 2025). The largest venue scored worst. Reporting on the study noted that Kalshi disputed the methodology, arguing calibration rather than hit rate is the right measure (DL News, 5 December 2025).

The favorable one comes from a study of Polymarket's own microstructure over 5 January to 6 November 2024. Price sensitivity to order flow fell from about 0.53 at the end of July 2024 to about 0.01 by October, and the half life of yes plus no pricing deviations fell from several hours in early 2024 to well under a minute by October and November (Tsang and Yang, arXiv:2603.03136, August 2026). By the end, the headline market was hard to push around and quick to close its own gaps.

Both can be true. A deep flagship market can behave well while the two thousand markets around it behave badly.

Where are prediction markets weakest?

Four conditions predict a bad price, and they compound.

  • Low volume. Thin markets are one person's opinion with a price attached.
  • Long horizon. The further from resolution, the more prices sit near 50.
  • Ambiguous wording. The market settles on the written rule, not the headline.
  • Contested resolution. Where the rule can be argued, the price includes dispute risk. See how prediction markets resolve.

The horizon effect has been measured. Across 353 million trades and 429,000 binary contracts on Kalshi and Polymarket, political markets showed persistent underconfidence, with prices compressed toward 50 percent: a contract at 70 cents a month out corresponded to a true probability closer to 75 percent. The model explained 87.3 percent of calibration variance in sample on Kalshi and 71.5 percent out of sample (Le, arXiv:2602.19520, February 2026). The important word is structured. The error is not noise, it has a direction, and a direction is something a forecaster can exploit.

How often does an individual beat the market?

Often enough to be worth measuring, rarely enough that one example proves nothing.

The best public benchmark for individuals is ForecastBench, which put the same unresolved questions to superforecasters, the general public and large language models. On the human question set, superforecasters reached a mean Brier score of 0.096, against 0.121 for the general public and 0.122 for the best model tested, with both gaps significant at p below 0.001 (Karger et al., arXiv:2409.19839, revised February 2025). Skilled individuals really do separate from the crowd, and they separate by a measurable amount.

That is a different claim from beating a specific market price on a specific event. To show that, you need your probability, the market's probability at the same instant, and a resolved outcome, repeated across dozens of events. Anything less is a story.

Method: how this page was built, and when it updates

  • Scope. Published accuracy studies of Polymarket, Kalshi, PredictIt, Metaculus and Manifold, plus benchmark results for human forecasters. No affiliate sources, no platform marketing pages.
  • Metric preference. Brier score first, calibration second, hit rate last and always labelled as such.
  • Sample sizes. Every number on this page carries its sample size and its date, or it is not on this page.
  • Excluded. Claims of the form "X percent accurate" with no sample, no date and no metric definition. There are many of them.
  • Cadence. This page is reviewed at least every 90 days and after any major election resolves. Last checked 4 September 2026.
  • What insiderz adds later. Calls on insiderz are locked with the time and the Polymarket price at that moment, so once events resolve there is a clean comparison of a person's probability against the market's probability at the same instant. That data does not exist yet in useful quantity, so it is not published here yet.

Open events are on Events, and the people ranked on Beats market, Events, Edge and Early are on the leaderboard.

Questions people ask

How accurate is Polymarket?
On a cross platform sample of resolved markets measured to 14 October 2025, Polymarket's overall Brier score was 0.1652 across 259 markets, the best of the four platforms compared. Accuracy drops sharply on low volume and long horizon markets.
Are prediction markets better than polls?
On short horizon elections, usually yes, because they aggregate polls plus everything else. On long horizon or technical questions the edge shrinks.
What is a good Brier score?
Around 0.25 is a coin flip. Superforecasters reached a mean of 0.096 on the ForecastBench human question set. Below 0.10 is elite.
Do prediction market prices efficiently aggregate information?
Not always. A study of more than 2,500 markets in the final five weeks of the 2024 US campaign found prices for identical contracts diverging across exchanges and arbitrage opportunities persisting to election day.

Sources

  1. Prediction Markets? The Accuracy and Efficiency of $2.4 Billion in the 2024 Presidential Election, Joshua D. Clinton and TzuFeng Huang, SocArXiv, 2025
  2. Are Polymarket and Kalshi as reliable as they say? Not quite, study warns, DL News, 5 December 2025
  3. Prediction market accuracy by platform, brier.fyi, data updated 14 October 2025
  4. ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities, Karger et al., arXiv:2409.19839, revised 28 February 2025
  5. Decomposing Crowd Wisdom: Domain-Specific Calibration Dynamics in Prediction Markets, Nam Anh Le, arXiv:2602.19520, February 2026
  6. The Anatomy of a Blockchain Prediction Market: Polymarket in the 2024 U.S. Presidential Election, Kwok Ping Tsang and Zichao Yang, arXiv:2603.03136, August 2026

Keep reading

Prediction markets: how a price becomes a probability

A prediction market is a market where people trade contracts that pay 1 if an event happens and 0 if it does not. The last traded price sits between 0 and 1, so a contract at 34 cents reads as a 34 percent chance. The price is a claim about the future made by everyone trading at once. Like any claim, it can be beaten.

7 min read

Market, poll or pundit: who was actually right?

On the 2024 US presidential race the market leaned toward the eventual winner and the leading poll model did not. Polymarket had Trump at 58 percent on 4 November 2024, while the 538 model's final forecast gave Harris 50 in 100 and Trump 49 in 100. Pundits produced no scoreable number at all. One election does not settle the general question, and the same market data set looks much worse when you widen the sample.

6 min read

Play money vs real money: does money help forecasts?

Money does not appear to buy accuracy. The head to head test that settled the question ran a real money exchange against a play money exchange across 208 NFL games in 2003 and found no statistically significant difference on four separate scoring rules. Cross platform Brier scores measured to October 2025 put the money and no money platforms in the same band. What money reliably buys is attention, liquidity and a way for large holders to push a price.

6 min read

Brier score explained in plain words

A Brier score measures how far your probabilities were from reality. For each forecast, take the probability you gave, subtract the outcome written as 1 for happened and 0 for did not, and square the result. Average that over all your forecasts. Zero is perfect, 0.25 is what you get by saying 50 percent every time, and 1 is as wrong as it is possible to be.

8 min read