Connect wallet

How to build a prediction track record people can verify

By insiderz9 min read

Flat abstract illustration on a dark background of a row of sealed tokens on a timeline, four of them stamped and locked, one drifting away unsealed

A prediction track record is believable when four things are true at once. Every call was published before the event. The timestamp came from someone other than you. Nothing can be edited or deleted afterwards. And every call, including the ones you lost, is scored against a public benchmark. Miss any one of those and what you have is a highlight reel.

Why does nobody believe a screenshot?

A screenshot proves that an image exists. It does not prove when the text in it was written, that the text was not changed, or that you did not make forty other calls that went the other way. It is the weakest possible evidence and it is the most common. The problem is not that people are dishonest. The problem is that the format has no way to be honest.

The same applies to a post you can delete. Pew Research Center tracked nearly 5 million tweets from March to June 2023 and found that 18% were no longer publicly visible within three months, with 60% of those disappearances caused by the account being made private, suspended or deleted and 40% by the author deleting the individual post. Deletion on social platforms is normal, silent and leaves no marker. A reader looking at someone's timeline in 2026 cannot tell what used to be there.

What does a track record have to survive?

A track record has to survive three attacks: someone claiming you wrote it later, someone claiming you changed it, and someone asking what else you said. The first two are solved by timestamps and immutability. The third is solved by completeness, and it is the one that actually decides whether a record is worth reading.

It also has to survive a fourth question that most records never face: compared to what? Calling a heavy favorite correctly is not evidence of skill. Calling it against a benchmark that said otherwise is.

What are the four properties of a verifiable track record?

Timestamped by someone else. The moment a call was made must be recorded by a party with no interest in the outcome: a platform, a blockchain, a public archive. Your own file dates do not count.

Immutable. Once published, the call cannot be edited or removed, not by you and not by the operator. If the operator can quietly delete a row, the record is only as good as the operator's incentives.

Complete. Every call you made is in the record, in the same place, on the same terms. Not the ones you remember. Not the ones that worked.

Scored against a benchmark. Each resolved call gets a number, and that number compares your call to what a public reference said at the same moment. Without a benchmark, "I was right" only tells you that the outcome was likely.

Here is how the common ways of keeping a record score on those four, checked on 4 September 2026.

Where the record lives Timestamp from a third party Author can remove it from public view Every call visible Scored against a benchmark
Screenshot in a chat No Yes No No
Post on X or Telegram Yes Yes No No
Personal blog or newsletter No Yes No No
OpenTimestamps proof file Yes Yes No No
Trading history on Polymarket Yes No Yes No, ranked by profit
insiderz profile Yes No Yes Yes, against the market price

Two rows deserve a note. An OpenTimestamps proof, announced by Peter Todd in September 2016, anchors a hash to the Bitcoin blockchain for free and proves that a specific text existed before a specific block. It is excellent at property one and useless at property three, because you decide later which proofs to show. And a Polymarket trading history is genuinely immutable and complete, but the platform's own ranking is not a skill measure: the leaderboard endpoint in Polymarket's documentation, as of September 2026, accepts exactly two ordering criteria, PNL and VOL, and returns no accuracy field at all.

Why is "complete" the property everyone cheats on?

Completeness is the only property you cannot fake with technology, and it is the only one that changes the answer. Timestamps and immutability are cheap to add. A denominator is expensive, because it means publishing your losses.

The arithmetic is unforgiving. If you make coin flip style calls and keep only the winners, you will look like a genius after twenty attempts. Ten correct calls in a row happens by chance roughly once in 1,024 tries at even odds, which sounds impressive until you notice that thousands of people are making calls every week, so a handful of them will produce that run without any skill at all. Show ten hits out of ten and it means something. Show ten hits and hide the fourteen misses and it means nothing.

CXO Advisory Group ran the honest version of this experiment. Between 2005 and 2012 it collected 6,582 public US stock market forecasts from 68 named experts and graded them all, winners and losers, arriving at a terminal accuracy of 46.9%, and 47.4% when averaged per guru. That is the number you get when nobody is allowed to choose which calls count. The same people, judged by their own selections, would look far better.

This is why deleted predictions are the central problem in this field rather than a side issue. A record with a hole in it is not a weaker record. It is a different kind of object.

Why score against the market instead of against right and wrong?

A hit rate ignores difficulty. Saying an incumbent will survive a routine vote, and being right, is worth almost nothing. Saying a 20% outcome will happen, and being right, is worth a great deal. A benchmark is what tells those two apart, and a live market price is the hardest benchmark available, because it already contains what everyone else thinks.

Formal forecasting research has used a benchmark since the beginning. The IARPA tournaments run from 2011 to 2014 scored every entrant with the Brier score, the squared distance between a probability and the outcome, and the winning program selected 60 people as superforecasters purely on measured accuracy, then watched them stay ahead of the comparison groups for two more years instead of regressing to the mean. That result exists only because every forecast was recorded, scored and kept, including the bad ones.

The same design is why forecasting can now be measured at all. A 2025 evaluation used 464 resolved Metaculus questions to compare frontier language models against top human forecasters, and found the models ahead of the crowd but behind the expert group. That comparison is only possible because both sides left a complete, scored record behind. See Brier score explained in plain words for how the number itself works.

How do you start a track record today?

  1. Pick events that resolve. A call needs an outcome and a date. "Crypto will have a big year" never resolves. "BTC closes above $X on 31 December" does.
  2. State a side and a confidence. Yes or no, plus how sure you are. Confidence is what makes a record scoreable rather than anecdotal.
  3. Publish before the event, somewhere you cannot edit. The moment matters more than the wording.
  4. Record the benchmark at that moment. If the reference price was 34% and you said yes, that number is part of the call. Without it, nobody can tell later whether you were early or just late.
  5. Keep making calls, including on things you are unsure about. A record of only your confident calls is a selected record again.
  6. Never remove anything. The misses are what make the hits readable.

On insiderz this is the default rather than a discipline you have to maintain. A call is locked the second it is posted, with the time and the Polymarket price at that moment frozen alongside it. It cannot be edited or deleted, not by you and not by us. When the event resolves, the call is scored against the market. Calls go public after a delay, and followers with live access see them the moment they are made. There is no money involved anywhere: you are putting your name on a statement, not a stake.

How do you read someone else's record?

Start with the denominator. How many resolved calls, over what period? Fewer than about a dozen and you are reading noise, whatever the hit rate says. On our leaderboard the rule is stricter: people under 30 resolved events are not ranked.

Then look at what the calls were worth. A record full of calls made at 90% market probability is a record of agreeing with the consensus quickly. Look for calls made where the reference price disagreed, and check whether the price later moved toward them. That last part, lead time, is the hardest thing to fake, because it requires the rest of the world to react after you.

Finally, check whether the record can go backwards. If the platform lets a user hide a call, or lets the operator remove one, the record's ceiling is the operator's honesty. Ask what the deletion policy is before you read the numbers. For the mechanics of proving a single call rather than a whole record, see every way to timestamp a prediction, compared.

What does insiderz claim, and what does it not?

insiderz claims exactly this: the calls on it are locked at the moment they are made, with the time and the Polymarket price recorded, they cannot be edited or deleted, and they are scored against that price when the event resolves. The leaderboard shows four columns, Beats market, Events, Edge and Early, and anyone can pull the same data through the public API. Identity is a wallet signature: no email, no name.

insiderz does not claim to have historical accuracy data of its own. The site launched in September 2026, which means every record on it starts empty, including ours. It does not claim that a good record predicts future results, and it is not a place to bet, because there is nothing to bet. What it does is remove the three excuses that make every other track record unreadable: you cannot backdate, you cannot edit, and you cannot hide the losses. Whether the numbers that follow are any good is up to the person making the calls. See P&L is not skill for why a profit ranking answers a different question entirely.

Questions people ask

How do you prove you predicted something?
You need a record made before the event, timestamped by someone other than you, that cannot be edited or deleted, and that includes your wrong calls as well as your right ones. Anything missing one of those four is a story, not proof.
Why is a screenshot not proof?
A screenshot shows only the calls you chose to keep. It carries no independent timestamp, it can be staged or edited, and it hides how many other calls you made and lost.
What makes a track record credible?
Completeness. A record that contains every call you made, scored against a public benchmark, is credible. A selection of your best calls is not, however accurate each one was.
How many predictions do you need before a record means anything?
Around a dozen resolved calls scored against a benchmark is the point where a good run stops being explainable by luck. Ten correct coin flips in a row happen by chance about once in 1,024 tries.
Does a track record need money at stake?
No. Money measures how much you risked, not how accurate you were. What a record needs is a benchmark, so that being right on an obvious outcome counts for less than being right where the benchmark was wrong.

Sources

  1. Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions, Mellers et al., Perspectives on Psychological Science, 2015
  2. Guru Grades, CXO Advisory Group, forecasts collected 2005 to 2012
  3. Link Rot and Digital Decay on Government, News and Other Webpages, Pew Research Center, 17 May 2024
  4. Get trader leaderboard rankings, Polymarket Documentation, accessed 4 September 2026
  5. OpenTimestamps: Scalable, Trust-Minimized, Distributed Timestamping with Bitcoin, Peter Todd, 15 September 2016
  6. Evaluating LLMs on Real-World Forecasting Against Expert Forecasters, Janna Lu, arXiv, July 2025

Keep reading

Deleted predictions: why crypto track records are fiction

A track record assembled from social media posts measures what survived, not what was said. Posts can be deleted, edited or quietly reframed, and nothing marks the gap afterwards. Pew Research Center found 18% of tweets vanish from public view within three months. Since losing calls are the ones most likely to disappear, the visible set always flatters the author.

8 min read

Every way to timestamp a prediction, compared

There are six practical ways to timestamp a prediction: a screenshot, a public post, a published hash you reveal later, an OpenTimestamps proof anchored to Bitcoin, an on chain transaction, or a scored forecasting platform. They differ on one axis that matters more than cost or difficulty: whether the proof survives you wanting it gone, and whether it shows the predictions you would rather forget.

7 min read

P&L is not skill: what the Polymarket leaderboard ranks

Polymarket's leaderboard ranks realized profit and trading volume, and nothing else. Its public API accepts exactly two ordering criteria, PNL and VOL, and the response carries no accuracy field, no hit rate and no count of resolved markets. That makes it an accurate answer to "who made the most money here" and a poor answer to "who knows what is going to happen".

8 min read

Brier score explained in plain words

A Brier score measures how far your probabilities were from reality. For each forecast, take the probability you gave, subtract the outcome written as 1 for happened and 0 for did not, and square the result. Average that over all your forecasts. Zero is perfect, 0.25 is what you get by saying 50 percent every time, and 1 is as wrong as it is possible to be.

8 min read