Sports Edge Model — every game called, every call graded, in public

The Lab · The math · updated 2026-08-15

How many graded games until a win rate means anything?

We just passed 500. The honest answer is that a win rate still tells you almost nothing — and that the number worth checking was never the win rate.

This week our board graded its 519th game. The favorite has won 286 of them, a 55.1% hit rate, every call published before the game and scored afterwards.

Five hundred sounds like a lot. It is the point where most people would say the record has proven something. So here is what 519 games actually buys, which is less than you would think in one direction and more than it sounds in another.

What a win rate is worth at 519 games

A hit rate is an estimate, and every estimate has a width. For a percentage the standard error is sqrt(p(1-p)/n), which at 55.1% over 519 games comes to 2.18 points.

Two standard errors either side is the usual bar for "we're fairly sure". That gives:

After 519 graded games, the true hit rate is somewhere between 50.8% and 59.4%.

Read that again, because it is the whole article. After more than five hundred published calls, our own data cannot distinguish "barely better than a coin flip" from "wins three out of five". Both are comfortably inside the range. Anyone quoting 55.1% as a fact is quoting the middle of an interval eight and a half points wide.

And it narrows slowly, because the n is under a square root. Halving that interval takes four times the games — around 2,000 more. Cutting it to a single point takes about 9,500 games, which at fifteen a day is roughly a decade and a half of baseball.

Here is the same arithmetic as a table. How many graded games before a calibration error of a given size becomes detectable at all:

error to detect games needed
3 points ~1,056
2 points ~2,376
1 point ~9,504

We are at 519. We are not close to the top row.

It is worse for profit

Win rate at least has the decency to be a single number. Return on investment mixes in price, and prices vary, so the variance is larger still.

To establish that a strategy really returns +3% — a genuinely good edge, the kind a professional would take all day — from win/loss records alone takes roughly 4,265 bets. Not 500. Not 1,000.

This is why we say, repeatedly and to our own cost, that our published picks record proves nothing. It is 3 graded, 2–1, +97.9% ROI. That ROI is arithmetically true and completely meaningless: three bets at plus money, two of which landed. A fourth result would move it by tens of points. We report it because hiding it would be worse, not because it is evidence.

If a service shows you a 60% win rate over 50 picks, you now know exactly what to do with that number. The standard error on 50 picks is about 7 points. Their 60% is somewhere between 46% and 74%. They have not shown you anything, and if they are charging for it, the sample size is the tell.

So what does 500 games buy?

A different question, and a better one.

Instead of "is the model right?" — which needs thousands of games — ask "do the model's numbers mean what they say?" When it says 62%, does that group win about 62% of the time? That is calibration, and it is answerable much sooner, because every band is its own test.

model said games actual inside its own error bar?
50–55% 226 51.8% yes
55–60% 178 54.5% yes
60–65% 76 61.8% yes
65–70% 27 63.0% yes
70–75% 10 60.0% yes

Every band lands within one standard error of what it claimed. The ordering holds too: higher confidence really does win more often, monotonically, across the 480 games in the first three rows.

Across the whole book the model claimed an average of 56.6% and delivered 55.1% — a miss of 1.5 points, which is 0.68 standard errors. Statistically invisible. The numbers mean what they say.

That is a real result and we are pleased with it. It is also not a claim that we beat anything, and the difference between those two sentences is the difference between this site and an advertisement.

The part that stings

Being calibrated is not the same as being profitable, and we know precisely why ours isn't.

Our published probability is the de-vigged consensus of the sportsbooks. No signal has ever been promoted past shadow testing, so the model agrees with the market by construction. Being well calibrated to the market's own opinion is a fair description of "having no edge", stated politely.

We have spent this season trying to find one, and the trying is mostly a list of things that did not work:

Four hypotheses, four nulls. That is what the search for an edge actually looks like, and almost nobody publishes this part.

What to do with any record, including ours

Ask for the sample size before the percentage. It is the number that decides whether the percentage means anything, and it is the number least often volunteered.

Prefer calibration to win rate. "When they say 60%, does it happen 60% of the time?" is answerable in hundreds of games. "Do they beat the market?" needs thousands.

Treat closing line value as the faster instrument. Comparing the price you got against the market's final number is far less noisy than waiting for games to land — it is roughly 400 times more efficient than scoring against outcomes, which is why we grade every pick that way.

Be suspicious of a good number on a small sample, most of all when it is ours.

We publish the record because you should be able to check it, not because 519 games has settled anything. At the current rate the win rate becomes genuinely informative somewhere around 2030. The calibration table is already telling you the truth today.

Model output is informational and entertainment content, not betting or financial advice. If you bet, bet what you can afford to lose.

Track your own bets against this board

Log what you actually bet — units and dollars — and see it graded beside a model that publishes every call it makes, wins and losses alike. Free, no card.

Start a free tracker

← Back to The Lab · How the model works · The record