Connect wallet

Superforecasters: what they do differently

By insiderz9 min read

Flat abstract illustration on a dark background of one bright dot far ahead of a dense cluster of dimmer dots along a horizontal accuracy scale

A superforecaster is someone who ranked in the top 2 percent for accuracy across hundreds of scored questions and then kept doing it. The label came out of a research tournament, not a marketing department, and the advantage is measurable: better calibration, sharper separation of what happens from what does not, and a set of working habits that are cheap to copy. Nothing in the list requires talent you can only be born with.

Where does the term come from?

Between 2011 and 2014 the US Intelligence Advanced Research Projects Activity ran three geopolitical forecasting tournaments in which five university research programmes competed to produce accurate probability estimates. The winning programme was the Good Judgment Project, led by Barbara Mellers and Philip Tetlock at the University of Pennsylvania. Every group was scored on the same metric, the Brier score that Glenn Brier introduced for weather forecasts in 1950.

The details matter, because they define the sample. Each tournament ran about nine months. The first two posed slightly more than 100 questions, the third about 150, and questions stayed open for an average of 102 days. Each year began with between 2,200 and 3,900 forecasters. At the end of the first year the project ranked everyone and took the top five forecasters from each of twelve experimental conditions, 60 people, and put them into five elite teams of twelve. Those teams were given the name superforecasters. All of this is documented in Identifying and Cultivating Superforecasters, published by Mellers and colleagues in Perspectives on Psychological Science in 2015.

The expectation was regression to the mean. It did not happen. Superforecasters stayed ahead in the second year and again in the third.

What does the measured advantage look like?

Measure Superforecasters or elite crowd Comparison group Study and date
Standardized Brier score, average of years 2 and 3 -0.34 -0.14 next tier, +0.04 all others Mellers et al., 2015
Calibration error 0.01 0.03 next tier, 0.04 all others Mellers et al., 2015
Resolution 0.40 0.35 next tier, 0.32 all others Mellers et al., 2015
Area under the ROC curve 96 percent 84 percent next tier, 75 percent all others Mellers et al., 2015
Distinct probability values used 57 29 next tier, 30 all others Mellers et al., 2015
Forecasts submitted per question, year 1 2.77 1.47 comparison groups Mellers et al., 2015
News stories opened, years 2 and 3 255 55 and 58 Mellers et al., 2015
Aggregate Brier score, elite poll versus 300+ non-elite crowd 0.166 0.228 Atanasov et al., 2024
Median Brier score, 498 benchmark questions 0.096 0.121 public median Karger et al., ICLR 2025

Two notes on reading this table. The Good Judgment tournament used the original two-outcome Brier definition, which runs from 0 to 2, so those numbers are not directly comparable to the 0 to 1 scores in the last row. And the standardized scores are relative to the pool on the same questions, which is the only fair way to compare forecasters who did not all answer the same things.

What are the habits, stated as steps?

Seven, in the order you would actually run them.

  1. Take the outside view first. Ask how often things in this reference class happen, before you look at the specifics of this case.
  2. Break the question into parts you can estimate. Decomposition is the single most transferable skill on this list.
  3. Look for the resolution criteria before you look for an answer. Half of forecasting disputes are definition disputes.
  4. State a specific number. Superforecasters used an average of 57 distinct probability values across their questions, against 29 and 30 for the two comparison groups, and were the group most likely to use numbers divisible by 1 percent rather than rounding to multiples of 10.
  5. Update often, in small steps. In the first tournament year superforecasters submitted 2.77 forecasts per question against 1.47 for the comparison groups. Frequency of belief updating turned out to be the strongest single behavioural predictor of accuracy in the project.
  6. Do the reading. Across the second and third years superforecasters opened an average of 255 news stories through the project's reader, against 55 and 58 for the comparison groups.
  7. Keep score, and look at the score. The project measured learning as the reduction in error over a tournament year, and superforecasters improved fastest.

The habits also came with real cognitive differences. Superforecasters scored at least one standard deviation above the general population on tests of fluid intelligence and on tests of political knowledge. The honest summary from the paper is that they were partly discovered and partly created. The created part is the part in the list above.

How does decomposition work in practice?

Decomposition means replacing one question you cannot estimate with several you can.

"Will this bill pass before the end of the year" is unestimable as a single lump. Broken up, it becomes: does it get out of committee, is there floor time in the remaining sessions, does the leadership want a vote, has any comparable bill passed in this configuration before. Each part has a rough base rate. Multiply the ones that must all happen, and you have a starting number that came from something rather than from a feeling.

The mechanical benefit is that it exposes the weak link. When you write out four sub-questions and find that one carries all the uncertainty, you know exactly what to research and what to ignore. It also gives you something to update: when news arrives, it usually bears on one sub-question, not on the whole.

Why update in small increments?

Because most news moves a probability a little, and the discipline of moving it a little is what separates a forecast from a reaction.

Frequent small updates were the strongest behavioural predictor of accuracy in the Good Judgment data. The mechanism is not mysterious. A forecaster who revisits a question weekly incorporates ten pieces of small evidence over ten weeks. A forecaster who sets a number and forgets it incorporates none, and a forecaster who swings from 20 percent to 80 percent on a single headline has substituted one piece of evidence for all the others.

Small increments also protect the granularity that makes the numbers useful. If your only moves are 10 percentage points, you cannot express the difference between mild and moderate evidence.

Why does keeping score matter so much?

Because without a score there is no feedback, and without feedback the habits above are just preferences.

This is where most public forecasting fails. Predictions get made in places where they can be deleted, ignored, or reinterpreted after the fact. A record you can edit tells you nothing about the person, and worse, it tells the person nothing about themselves. The Good Judgment result depends entirely on the fact that every forecast was timestamped, scored against a fixed rule, and impossible to withdraw.

A second requirement is a benchmark. A raw accuracy number depends on how hard the questions were. Scoring against something that answered the same questions at the same moments, another forecaster, a base rate, or a market price, is what turns a number into evidence about you.

How do superforecasters compare with markets and with AI in 2026?

Against markets, the useful finding is that the crowd matters more than the mechanism. In Crowd prediction systems: Markets, polls, and elite forecasters, published in the International Journal of Forecasting in 2024, Pavel Atanasov, Jens Witkowski, Barbara Mellers and Philip Tetlock compared prediction markets against team prediction polls using Good Judgment tournament data, with each system run by either a large non-elite crowd or a small elite one. Small elite crowds outperformed larger ones, while the two systems were statistically tied with each other. The aggregated elite poll scored a Brier score of 0.166 against 0.228 for a non-elite crowd of more than 300 forecasters on the same questions, and a decomposition showed the gap came entirely from better discrimination, not better calibration. The elite cutoff in that data was the top 2 percent. The same paper notes earlier published work finding that elite forecasters in team polls were more accurate than professional intelligence analysts with access to classified information.

Against AI, the picture is dated the moment you write it down. On ForecastBench, presented at ICLR 2025, the superforecaster median Brier score was 0.096 across 498 questions, against 0.121 for the public median and 0.122 for the best language model tested, Claude 3.5 Sonnet. The human forecasts were collected in July 2024. The same paper fits a trend between model quality and forecasting accuracy and projects that models would reach superforecaster level at an Arena score of about 1406, with a wide confidence interval. As of September 2026, any claim about the current gap needs a current benchmark run behind it.

Can you practise without a tournament?

Yes, and the requirements are short: questions that actually resolve, a way to record a probability that cannot be changed afterwards, and a benchmark to be scored against.

Metaculus and Good Judgment Open both run scored questions. insiderz gives you the third ingredient by default. A call on insiderz is a yes or no statement on a real event with your confidence attached, on the same events Polymarket lists. The call is locked the second you post it, with the time and the Polymarket price at that moment frozen beside it. Nothing can be edited or deleted, not by you and not by us. When the event resolves, the call is scored against that frozen price, which means the benchmark is built in rather than chosen after the fact.

There is no money involved and nothing to bet. The insiders leaderboard ranks people on how often they beat the market price, how many events they have resolved, their average edge against the price, and how early they were. Open events are here.

For the score itself, in plain words, read Brier score explained. For the failure mode these habits are designed to fix, read calibration. For the practical version of all of this applied to a live price, read how to beat the market.

Questions people ask

What is a superforecaster?
In the Good Judgment Project the label was operational, not honorary: the top 2 percent of forecasters by accuracy at the end of a tournament year, promoted into small elite teams. It means a measured record on scored questions, not a job title.
Can you learn to forecast better?
Yes. In the Good Judgment Project the strongest single behavioural predictor of accuracy was how often people updated their forecasts, and a short probability training module improved accuracy. Both are things anyone can copy, and both only work if the forecasts are scored.
Do superforecasters beat prediction markets?
Small elite crowds beat large non-elite crowds in both formats, while prediction markets and team prediction polls were statistically tied with each other in the Good Judgment data. Who is in the crowd mattered more than which mechanism was used.
Do superforecasters still beat AI in 2026?
On ForecastBench, whose human forecasts were collected in July 2024, the superforecaster median Brier score was 0.096 against 0.122 for the best language model. The paper projects that gap closing as models improve, so treat any claim about it as dated.
How do I become one?
There is no shortcut around the record. Make forecasts on questions that resolve, state a number rather than an opinion, keep every forecast including the bad ones, and score them against a benchmark on the same events.

Sources

  1. Barbara Mellers et al., Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions, Perspectives on Psychological Science 10(3), 2015
  2. Pavel Atanasov, Jens Witkowski, Barbara Mellers and Philip Tetlock, Crowd prediction systems: Markets, polls, and elite forecasters, International Journal of Forecasting, 2024
  3. Ezra Karger et al., ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities, ICLR 2025, arXiv, revised February 2025
  4. Glenn W. Brier, Verification of Forecasts Expressed in Terms of Probability, Monthly Weather Review 78(1), American Meteorological Society, January 1950

Keep reading

How to beat the market: a practical guide

Beating a prediction market means being right where the price was wrong, often enough and on enough events that luck stops being the explanation. There is no trick that works everywhere. What works is a method: start from the price rather than your opinion, disagree only in the few situations where prices are known to be soft, put a specific number on it, and score every call against the price at the moment you made it.

10 min read

Brier score explained in plain words

A Brier score measures how far your probabilities were from reality. For each forecast, take the probability you gave, subtract the outcome written as 1 for happened and 0 for did not, and square the result. Average that over all your forecasts. Zero is perfect, 0.25 is what you get by saying 50 percent every time, and 1 is as wrong as it is possible to be.

8 min read

Calibration: what being right 70 percent means

Calibration is the match between what you claim and what happens. If you are calibrated, the events you call 70 percent happen about 70 percent of the time, the ones you call 90 percent happen about 90 percent of the time, and so on down the scale. It is a property of a batch of forecasts, never of one. A single 70 percent call that comes true proves nothing.

7 min read