Prediction markets, graded

Who actually makes money in prediction markets, measured.

We started as a forecast desk that tried to beat Kalshi's weather prices and lost, in public, with the receipts. Now we do the measuring: the famous trades decoded from their own fills, the exchanges' tapes read from the kill to the close, the price scored week by week, and our own calls graded when their dates arrive. Sources and timestamps on every figure. Never a recommendation.

Illustration · the desk is closed NYC daily high ≥ 90°F · today
0%
Market
0%
NWS data
0%50%100%
GAP +23 pts 12Z update raised the forecast high to 91°F; market hasn't caught up.
The score, kept in public

The desk that started this, graded against what happened.

After each market settles, the desk's published probability is scored against the official NWS climate report — and, because the market's own price is captured at the same read, head-to-head against the price on the same contracts. The pipeline writes this section, not the marketing department; it can get worse as easily as better.

Head-to-head Brier · same settled markets · lower is better
The desk
Market price
settled outcomes hit rate calibrating — first settlements landing
Calibration · what we said vs what happened
0% predicted probability 100% 0 100
On the dashed line = perfectly calibrated · dot size = sample count
See every settled market and the full graphs

Brier score is the mean squared error of probability forecasts — 0 is perfect, and always guessing 50% scores 0.25. Scored on settled Kalshi weather markets against official NWS climate reports over a rolling 90-day window, one prediction of record per contract. Past calibration is not a guarantee of future results, and none of this is a recommendation.


The opening

Everyone repeats who wins in these markets. Almost nobody reads the tape.

The problem

The stories that circulate about prediction markets, the quiet whale, the boring strategy that prints, the edge that anyone could farm, are told from leaderboards and screenshots. The fills, the order books and the trade tapes that would confirm or kill them are public, and they go unread.

What the desk does

We pull the fills, resolve every market, and score them. We read the exchange's own tape from the moment a contract is decided to the close. We log the ladders every hour with the size that was there, score the price week by week, and post dated calls that get graded when their dates arrive, with the misses left up.


How it is done

Built like a desk, not a tip line.

Sources on every claim

fills · tape · order book · settlement · time

Every figure names the endpoint it came from and the moment it was read. The code that computed it is in the open, with tests.

We never tell you what to bet

descriptive, not prescriptive

You get the measurement and the receipts. Most of what we find is that an edge belongs to someone else, and we say so.

We grade ourselves first

dated call · rule · probability · result

The desk this site began as lost to the market's price and published the loss. The calls page keeps that habit.

The pieces

Decoded from the fills. Read off the tape.

Each piece takes a claim people repeat about these markets and tests it against the venue's own record, with the method and the code at the end.

One email per piece, nothing else. No brief, no alerts, no seats.

For information only. Not financial, investment, or trading advice. Figures are computed from public exchange data and shown with sources and timestamps; verify independently. Prediction-market trading carries risk and may be restricted in your jurisdiction.