Model statistics

Model accuracy: why 80% can be worth nothing

TL;DR. Hit rate is the figure everybody asks for and the easiest to inflate: just forecast the obvious. A model that gets 80% of its forecasts right can be worse than one getting 45% right, and understanding why is the difference between reading a statistic and understanding it.

Note. Informational and statistical content. It is not betting advice nor a promise of profit. No statistic predicts a single match: they are there to estimate probabilities over the long run. 18+ only. Play responsibly.


Why a high hit rate is trivially achievable

Forecast only matches with an overwhelming favourite and you will be right the vast majority of the time. You have demonstrated nothing: you selected the easy cases. Anyone with access to the table would have done the same.

We have verified it in our own data: our high-probability section posts a very high hit rate, and that is exactly what is expected of it. A high hit rate on high-probability forecasts is not a model achievement, it is the definition of the category.

Hit rate only says something when compared with what you expected to get right given the kind of matches forecast. 80% on matches with a clear favourite is normal; 55% on evenly matched games would be remarkable.

Accuracy is not profitability

They are independent, and this is the most expensive confusion in the industry. You can be right often and gain nothing, or right rarely and gain, because what matters is not how often you are right but whether the probability you estimated was better than the one the price reflected.

In our own data there are categories with a very high hit rate and essentially flat long-run performance. That is not a model defect: it is what happens when the market already had that same match well estimated.

MetricWhat it measuresWhen it misleads
Hit rateHow often you are rightWhenever the type of match is not stated
CalibrationWhether a 70% happens 70% of the timeRarely: it is the honest metric of a probabilistic model
SampleHow many forecasts sit behind itAny figure based on fewer than a hundred cases

The honest metric is calibration

A probabilistic model is well calibrated if, of all the matches it gave 70%, roughly 70% are won. That is what makes a model useful: not that it is right often, but that its numbers mean what they say.

A calibrated model saying 55% and being right 55% of those times is more useful than one saying 90% and being right 70%, even though the second has a better raw hit rate. The first gives you a probability you can rely on; the second gives you a decorative figure.

How we publish it

On every team page we show the model's accuracy on that team, but only when at least five matches have been judged. With three matches, a «67%» is noise presented as data, and we would rather show nothing than show a figure that means nothing.

And we always say what it is: the percentage is the accuracy of the forecast, not a return. You can see aggregate performance on the performance page and the per-team detail on the team pages.

Frequently asked questions

Is a model with 80% accuracy a good model?

There is no way to tell without knowing which matches. Forecasting only games with an overwhelming favourite reaches that figure trivially. Hit rate only says something compared with what you expected given the matches chosen.

Does being right often mean making money?

No. They are independent: what matters is not how often you are right but whether your estimated probability beat the price's. There are categories with very high accuracy and essentially flat performance.

Which metric really measures a model's quality?

Calibration: that of the matches it gave 70%, roughly 70% are won. A calibrated model saying 55% and being right 55% of the time is more useful than one saying 90% and being right 70%.