How we price a match
Where our numbers come from, what the AI is allowed to change, and how we mark our own work.
It starts with arithmetic, not a guess
Every prediction begins with one number: how many goals we expect each side to score. That figure is calculated from league data before any AI is involved — each team's scoring and conceding rate, measured against the average for their own division, and combined to estimate the goals they should produce against this particular opponent.
We measure each side in the venue they are about to play in. Home and away records are kept apart the whole way through, because collapsing them loses real information: in the 2025 Eredivisie, home teams averaged 1.80 goals a game and away teams 1.38. A model that treats those as one number and adds a generic home bonus throws away a figure the league table gives it for free.
That calculation needs a league table with enough football behind it, and in August there is not one. Rather than give up on league data for the opening weeks of a season, we start from last season's final table and let this season take over as it accumulates — a side's first home game counts for a little, its tenth for a great deal more. The handover is gradual rather than a switch thrown on some particular matchday.
Newly promoted teams are the honest gap in that. They have no record in the division they have just entered, so we price them at exactly league average and mark the prediction accordingly. That is a placeholder, not a measurement, and those fixtures deserve more scepticism than the rest. Where neither season offers a usable table — a cup with no group stage, a competition in its first weeks — we fall back to recent scoring and conceding rates, and where even that is thin the model estimates the figure directly. Predictions are not all built on the same footing, and we would rather say so than imply a uniformity that does not exist.
Being concrete about it: for much of this site's history the league-table calculation was not running at all, and the estimate came from the model alone. That is visible in the record — it produced roughly the same home advantage in every competition, which is not how football works. Korea's K League 1, where away teams have outscored home teams across the matches we have graded, was priced with the same home bias as the Eredivisie, where the advantage is real and large. It is the single clearest reason our worst-performing competition performs as badly as it does.
What the AI actually does
Where a baseline exists, the language model does not invent the numbers. It receives that calculated figure and is asked for one thing: a bounded adjustment for what the arithmetic cannot see — injuries, suspensions, rotation, a side already through in a cup tie, a manager resting players before a bigger fixture. Where no baseline could be calculated, the model estimates expected goals itself, and that is the weaker path described above.
That adjustment is capped in both directions. The model can nudge an estimate; it cannot replace it. If it returns nothing usable, we use the baseline unchanged — a statistical estimate is a perfectly serviceable prediction on its own.
This is deliberate. Asked to produce expected goals unaided, a language model gives a different answer each time you ask, and that variance carries through to every number on the page. Anchoring it to a calculation removes the wobble and means a bad AI judgement degrades to a reasonable statistical estimate rather than to nonsense.
Why our confidence numbers went down
Grading thousands of our own predictions showed something uncomfortable: the confidence we stated was consistently higher than what we delivered. Calls we rated in the mid-to-high sixties were landing closer to one time in two. The extremes were fine — it was the middle of the range, where most of our output sits, that was overstated.
So confidence is no longer reported as the model states it. It is mapped onto what confidences like it have actually achieved across our graded record. In practice that means the numbers you see now are lower than the ones this site showed before, sometimes by fifteen points or more.
That is not the model getting worse. It is the same model, reporting honestly. A published 70% that wins half the time is worse than useless to anyone acting on it — it inverts the decision. We would rather show you a 52% that means 52%.
Every prediction records which version of this correction produced it, so the track record can show whether the correction actually worked rather than asking you to assume it did.
One estimate, every market
From that single expected-goals figure we calculate the probability of every possible scoreline, then read the markets off it: the win/draw/win split, over and under goals, both teams to score, and the most likely exact score.
This matters more than it sounds. When those markets are estimated separately they contradict each other — we measured our own earlier version stating an over/under probability that its own win/draw/win split did not support, in the same direction, in roughly four cases out of five. Deriving everything from one distribution makes agreement automatic rather than something to check. The scoreline list on a prediction page and the goals tip beside it are the same calculation viewed two ways.
Why we sometimes say “or draw”
Naming the single most likely result sounds decisive and is often misleading. Because the three outcomes share 100% between them, a draw can only ever be the outright favourite when it clears roughly a third on its own — which is rarer than draws actually are. A model that always names one winner will almost never call a draw, however even the match is.
It also flattens the difference between a coin-flip and a near-certainty: both come out as “away win”. So when one outcome is genuinely ahead we say so, and when the match is tight we give a double chance that includes the draw, and flag it as close. Giving up a little apparent confidence for a call that is actually informative is the right trade.
How we mark our own work
Once a match is played, the prediction page shows the result and grades what we said against it, market by market. A few points about how that is scored:
- Markets settle on ninety minutes. A tie level at full time and won in extra time is graded as the draw it was. Extra time and penalties do not turn a correct call into a wrong one, or the reverse.
- The exact scoreline is counted like every other market, even though it lands roughly one time in ten. Leaving it out would flatter the record.
- Wrong calls are shown as prominently as right ones. A scorecard that only displayed the hits would be worth less than no scorecard.
- The analysis is never edited after the result. What you read on a played fixture is the preview exactly as it was published before kick-off.
What we do not model
Some of this is football being football, and some of it is work we have not done yet. Both are worth knowing:
- We treat the two teams' scoring as independent. In reality low-scoring draws happen slightly more often than that assumption implies, so those scorelines are marginally under-weighted.
- We have no live injury or suspension feed. The model works from whatever team news is available at the time, which varies by league and by fixture.
- Confidence scores are corrected against our own record rather than taken at face value — see below. The correction is only as good as the data behind it, and competitions where we have graded fewer matches are corrected less confidently than the ones where we have hundreds.
- Weather, referee decisions, and in-game incidents are not modelled at all.
The model is not finished
Every prediction we publish is graded against the result, and that record is what the method above gets corrected against. It is the current version, not a fixed formula — where the record shows a systematic bias, the correction is measured from the data rather than guessed at. The confidence recalibration above is the first of those, and it will not be the last.
More results make that sharper. A competition with hundreds of graded predictions can be adjusted on its own evidence; one with a few dozen leans on the broader picture until it has enough to stand alone. Each prediction records which version of the model produced it, so a change either shows up in the track record or it did not work — and you can check which rather than take our word for it.
Using this sensibly
These are probabilities, not forecasts. A 70% call losing is not a broken model — it is the 30% arriving, and it will keep arriving about three times in ten. The only honest way to judge a prediction service is over hundreds of results, which is why we publish ours rather than a highlight reel.
Treat our numbers as one input alongside your own judgement, never as a substitute for it. Please gamble responsibly. If it stops being fun, visit BeGambleAware.