Prediction accuracy & track record
Every prediction we publish, graded against the result — including the ones we got wrong.
Covering 13 November 2025 to 28 August 2026. Every figure below carries the number of predictions behind it, because a percentage without a sample size is not something you can check.
By market
| Market | Correct | Graded | Hit rate |
|---|---|---|---|
| Match result | 2,424 | 4,312 | 56.2% |
| Exact score | 419 | 4,312 | 9.7% |
| Goals line | 2,332 | 4,307 | 54.1% |
| Both teams to score | 2,379 | 4,305 | 55.3% |
The exact scoreline is listed alongside the rest even though it lands around one time in ten — it is genuinely hard, and dropping it would quietly lift every other number on this page. Markets that neither won nor lost, such as a goals line the match landed exactly on, are excluded from the graded count rather than counted as wins.
Are our confidence scores honest?
When we say a call is 70% likely, it ought to come in about seven times in ten. Here is what actually happened, grouped by the confidence we stated at the time:
| We said | Predictions | Actually right | Gap |
|---|---|---|---|
| 45–49% | 7 | 71.4% | +23.4 |
| 50–54% | 141 | 43.3% | -8.1 |
| 55–59% | 118 | 50.0% | -7.7 |
| 60–64% | 1,280 | 54.0% | -9.2 |
| 65–69% | 1,034 | 49.1% | -18.4 |
| 70–74% | 1,129 | 56.2% | -15.9 |
| 75–79% | 419 | 74.9% | -2.5 |
| 80–84% | 142 | 83.8% | +1.5 |
| 85–89% | 41 | 78.0% | -8.4 |
| 90–94% | 1 | 100.0% | +10.0 |
Our confidence scores are currently overstated.
Across 2,163 predictions — a substantial share of this record — we claimed materially more confidence than the results justified. The extremes of the scale hold up well; the middle, where most of our output sits, does not. We are publishing this rather than quietly rescaling the numbers first, and correcting it is the current priority. Until it is fixed, treat a mid-range confidence score as weaker than it reads.
By competition
Competitions with at least 50 graded predictions. Anything thinner is listed as a total rather than a rate — a hit rate from a dozen matches tells you nothing.
| Competition | Graded | Result | Score | Goals | BTTS |
|---|---|---|---|---|---|
| Major League Soccer | 306 | 52.0% | 9.8% | 61.1% | 61.8% |
| Serie A(Brazil) | 277 | 56.7% | 10.8% | 52.2% | 56.5% |
| Serie A(Italy) | 269 | 58.7% | 11.2% | 47.2% | 45.4% |
| Premier League | 269 | 51.7% | 7.8% | 55.0% | 58.0% |
| La Liga | 267 | 56.9% | 10.1% | 51.7% | 56.9% |
| Pro League | 249 | 62.2% | 10.8% | 54.2% | 56.5% |
| J1 League | 242 | 54.1% | 10.7% | 47.9% | 54.1% |
| Jupiler Pro League | 226 | 50.4% | 12.4% | 53.5% | 54.9% |
| Primeira Liga | 225 | 60.0% | 8.9% | 53.6% | 49.1% |
| Liga MX | 221 | 54.3% | 12.2% | 55.0% | 58.4% |
| Eredivisie | 219 | 57.1% | 5.5% | 59.4% | 62.6% |
| Bundesliga | 216 | 58.3% | 10.2% | 63.4% | 64.4% |
| Süper Lig | 206 | 58.3% | 14.6% | 57.8% | 55.6% |
| UEFA Champions League | 200 | 55.0% | 6.0% | 59.5% | 54.5% |
| Ligue 1 | 199 | 53.3% | 8.0% | 52.8% | 51.3% |
| UEFA Europa League | 185 | 59.5% | 9.7% | 51.1% | 49.5% |
| Premiership | 177 | 63.3% | 7.3% | 51.4% | 50.8% |
| K League 1 | 159 | 41.5% | 5.7% | 41.8% | 49.0% |
| World Cup | 104 | 64.4% | 8.7% | 52.9% | 54.8% |
| World Cup - Qualification Europe | 54 | 66.7% | 9.3% | 57.4% | 59.3% |
| 11 smaller competitions | 42 | too few to report a rate | |||
This record is what improves the model
The point of grading every prediction is not the scoreboard — it is the correction. This record is the measurement the model gets corrected against, so it is judged on what happened rather than on its own assumptions.
The overconfidence in the table above is the first correction we have made. It was visible only because we grade ourselves; nothing else would have surfaced it. Predictions are now published with their confidence mapped onto what confidences like them have actually achieved, which is why the figures on the site are lower than they used to be.
The reliability table still shows the gap, because it covers every prediction we have ever graded — including the thousands made before the correction existed. As predictions made under it accumulate, this is where you will see whether it worked. If the gap does not close, that will be visible here too.
More results make each correction sharper and more specific. Competitions where we have hundreds of graded predictions can be adjusted on their own evidence; thinner ones will lean on the overall picture until they have accumulated enough to stand alone, rather than being tuned on a handful of matches and mistaking noise for a pattern.
Every prediction records which version of the model produced it. That is deliberate, and it cuts against us: it means a change either shows up in these numbers or it did not work, and you can check which. A track record is only worth anything if it can prove us wrong as readily as right — so expect these figures to move, and to move on the evidence.
How these are counted
- Markets settle on ninety minutes. A tie level at full time and won in extra time is graded as the draw it was. Extra time and penalties never turn a wrong call into a right one.
- Every published prediction is included. There is no selection — a fixture we priced appears here whatever happened, and postponed or abandoned matches are excluded because they have no result, not because of how they went.
- What we said is frozen when the match is graded. A prediction cannot be revised after the fact and re-scored.
- Nothing here is edited for presentation. The preview on a played fixture is exactly what was published before kick-off.
Football is high variance, and a percentage over a few dozen matches is mostly noise. Judge any prediction service — including this one — over hundreds of results, which is why the sample size sits beside every figure above.