Connect wallet

Calibration: what being right 70 percent means

By insiderz7 min read

Flat abstract illustration on a dark background of a diagonal reference line with a curve bending below it, suggesting a calibration plot without any text

Calibration is the match between what you claim and what happens. If you are calibrated, the events you call 70 percent happen about 70 percent of the time, the ones you call 90 percent happen about 90 percent of the time, and so on down the scale. It is a property of a batch of forecasts, never of one. A single 70 percent call that comes true proves nothing.

How do you read a calibration table in ten seconds?

A calibration table groups your resolved forecasts into confidence bands and asks, for each band, how often the event actually happened. Stated probability goes on the left, real hit rate goes on the right. Perfect calibration means the two columns match.

Here is a worked example: 100 resolved forecasts from one person, sorted into six bands. The numbers are an illustration, not real calls.

Confidence band Calls in band Average stated probability Times it happened Real hit rate Gap
0 to 10 percent 8 0.05 1 0.13 +0.08
10 to 30 percent 12 0.20 3 0.25 +0.05
30 to 50 percent 15 0.39 6 0.40 +0.01
50 to 70 percent 20 0.59 12 0.60 +0.01
70 to 90 percent 25 0.78 17 0.68 -0.10
90 to 100 percent 20 0.94 15 0.75 -0.19

Read the gap column and nothing else. Near zero in the middle bands, negative at the top, positive at the bottom. This person is overconfident: when they are sure, they are less right than they say, and when they dismiss something, it happens more often than they allow.

What does overconfidence look like?

Overconfidence is a calibration table where the high bands under-deliver and the low bands over-deliver. Plotted as a curve against the diagonal, it sits inside the line at both ends: too flat, hugging the middle of the outcome axis while your stated probabilities run to the edges.

In the table above, the 90 to 100 percent band is the damage. Twenty calls stated at an average of 94 percent came true 15 times, a hit rate of 75 percent. Those five surprises are expensive under any squared-error scoring, and they are also the ones people remember.

The fix is mechanical, not motivational. If your 90s land at 75, write 75 next time. You do not need to become more knowledgeable to fix a calibration gap. You need to stop using numbers your record does not support.

Why is underconfidence also a failure?

Underconfidence is the reverse pattern: the events you called 70 percent happen 85 percent of the time. Your curve sits outside the diagonal. It feels safer, and it is still wrong, because you left information on the table. Someone reading your forecasts cannot recover the extra certainty you had and did not state.

Prediction markets show this pattern too, on some question types. In Decomposing Crowd Wisdom, an arXiv preprint from February 2026 revised in August 2026, Nam Anh Le analysed 353 million trades across 429,000 binary contracts on Kalshi and Polymarket and found that calibration varies by event domain, by time to resolution and by trade size, with political markets showing persistent underconfidence where prices compress toward 50 percent. A price that will not commit is the market version of the same failure.

Calibration, resolution and sharpness: what is the difference?

Calibration is one of three parts, and on its own it is not enough.

Allan Murphy showed in A New Vector Partition of the Probability Score, published in 1973, that a squared-error forecast score splits into three terms: the uncertainty in the events themselves, the reliability of the forecasts (this is calibration), and the resolution, meaning how far the forecaster separated the events that happened from the ones that did not.

  • Calibration asks whether your numbers mean what they say.
  • Resolution asks whether your numbers distinguish between cases at all.
  • Sharpness is the plain-words version of resolution: are you willing to move away from 50 percent.

Saying 50 percent on everything is perfectly calibrated and has zero resolution. It is also worthless. The pair matters: a good forecaster is calibrated and sharp.

The superforecaster study by Barbara Mellers and colleagues, published in Perspectives on Psychological Science in 2015, reports both numbers separately for the three groups in the Good Judgment Project. Averaged over the second and third tournament years, calibration error was 0.01 for superforecasters, 0.03 for the next tier of high performers and 0.04 for everybody else, while resolution ran 0.40, 0.35 and 0.32 in the same order. Superforecasters were better on both dimensions at once, which is the harder trick.

How many resolved forecasts before your curve means anything?

Roughly 100 resolved forecasts before the shape of the curve is worth looking at, and a few hundred before a small gap is a fact rather than noise.

The arithmetic is simple binomial noise. If your true hit rate in a band is 70 percent and you have 20 resolved calls in that band, the standard error is about 10 percentage points, so an observed rate anywhere from roughly 49 to 90 percent is consistent with being perfectly calibrated. With 100 calls in the band the standard error falls to about 4.6 points, and with 400 calls to about 2.3 points.

Two practical consequences. First, do not redesign your forecasting after a dozen resolved events. Second, a calibration curve built from six bands needs several hundred forecasts in total, because each band needs its own sample. This is the main reason most published calibration claims are weaker than they look.

How do you improve your calibration?

Four habits, in the order that pays.

  1. Write the number before you look at anything. Your own probability first, then the market price, then the news. Anchoring on a price you have already seen destroys the evidence about your own judgment.
  2. Record every forecast, including the ones that embarrass you. A record you can prune is not a record. This is why timestamped, non-editable calls are the only kind worth scoring.
  3. Group and recount every few months. Build the table above from your own resolved forecasts. The gap column tells you exactly which bands to shift.
  4. Use more numbers. In the Good Judgment Project, superforecasters used an average of 57 distinct probability values across the questions they answered, against 29 and 30 for the two comparison groups, and were more likely to give values divisible only by 1 percent. Rounding to multiples of 10 percent threw away real information.

Forecast accuracy is also measurable against a fixed external standard. On ForecastBench, presented at ICLR 2025, the superforecaster median scored 0.096 across 498 questions while the public median scored 0.121, with human forecasts collected in July 2024. Calibration is a large part of that gap, and it is the part you can fix deliberately.

Where can you build a curve without a tournament?

You need three things: a stream of questions that resolve, a way to record a probability that cannot be changed afterwards, and a benchmark to compare against.

On insiderz, a call is a yes or no statement on a real event with your confidence attached. The call is locked the second you post it, with the time and the Polymarket price at that moment frozen next to it. Nothing can be edited or deleted, not by you and not by us. When the event resolves, the call is scored against that frozen price, and your profile carries the whole history, including the bad ones. Open events are here and the insiders leaderboard shows who is currently ahead of the price.

insiderz launched in September 2026, so there is not yet enough resolved history for a platform-wide calibration curve. These are the counts so far.

insiderz right now

Live from insiderz.
calls
0
insiders
0
resolved events
0
open events
1,904

For the score that combines calibration and sharpness into one number, read Brier score explained. For what a probability on a market actually means before you try to match it, read what does a 34 percent chance actually mean. For the habits behind the best measured calibration on record, read superforecasters.

Questions people ask

What does calibrated mean in forecasting?
That the things you call 70 percent happen about 70 percent of the time, across many forecasts. Calibration is a property of a batch of forecasts, never of a single one.
Is calibration the same as accuracy?
No. You can be perfectly calibrated and useless by saying 50 percent every time. You also need sharpness, which means moving away from 50 percent when the evidence lets you.
How do I check my calibration?
Group your resolved forecasts into confidence bands, count how often the event happened in each band, and compare that hit rate to the band. Any gap larger than random noise is a bias you can correct.
How many forecasts do I need before my calibration curve means anything?
Roughly 100 resolved forecasts before the bands are worth reading, and a few hundred before small gaps are meaningful. With 20 forecasts in a band, a true 70 percent rate can land anywhere between about 49 and 90 percent by chance alone.
What is the difference between overconfidence and underconfidence?
Overconfidence means your high bands under-deliver and your low bands over-deliver, so your curve sits inside the diagonal. Underconfidence is the reverse: you were right more often than you claimed, and you left information unsaid.

Sources

  1. Allan H. Murphy, A New Vector Partition of the Probability Score, Journal of Applied Meteorology 12(4), American Meteorological Society, 1973
  2. Barbara Mellers et al., Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions, Perspectives on Psychological Science 10(3), 2015
  3. Ezra Karger et al., ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities, ICLR 2025, arXiv, revised February 2025
  4. Nam Anh Le, Decomposing Crowd Wisdom: Domain-Specific Calibration Dynamics in Prediction Markets, arXiv, February 2026, revised August 2026

Keep reading

Brier score explained in plain words

A Brier score measures how far your probabilities were from reality. For each forecast, take the probability you gave, subtract the outcome written as 1 for happened and 0 for did not, and square the result. Average that over all your forecasts. Zero is perfect, 0.25 is what you get by saying 50 percent every time, and 1 is as wrong as it is possible to be.

8 min read

What does a 34% chance actually mean?

A 34 percent chance means that across a large set of claims made with the same confidence, about 34 out of every 100 come true. It is a statement about a group, not about one event. When the event happens, the 34 percent forecast was not wrong. It was a forecast that said this happens about a third of the time, and a third of the time is often.

5 min read

Superforecasters: what they do differently

A superforecaster is someone who ranked in the top 2 percent for accuracy across hundreds of scored questions and then kept doing it. The label came out of a research tournament, not a marketing department, and the advantage is measurable: better calibration, sharper separation of what happens from what does not, and a set of working habits that are cheap to copy. Nothing in the list requires talent you can only be born with.

9 min read