Handbook

How the model works, what it can't do, and how to read what it gives you.

10 sections

Start here

What BankerAI is

BankerAI is a subscription information service. We publish statistical estimates of how likely each outcome in a football match is, and we publish how those estimates performed afterwards.

We are not a bookmaker. We do not accept, hold or process bets, and we are not a licensed gambling operator. Nothing here is advice, a tip, or a guarantee — every number on the site is the output of a model, and models are wrong regularly.

You place your own bets with your own bookmaker, and the outcome is yours.

What we cover

The model is trained on 143,411 completed matches across 35 leagues, with 12 to 16 seasons of history per league.

Coverage is not uniform. Some leagues carry richer source data than others, and a few are between seasons at any given time. If a league has no upcoming fixtures in the next 30 days, it will not appear in your picks — that is the season gate working, not a fault.

The model

How a prediction is made

For each fixture the model reads pre-match features only — nothing that happened during or after the game — and produces a probability for every outcome it supports.

Two families of model sit behind that. A gradient-boosted classifier (XGBoost and LightGBM, blended) estimates the match result directly. A Poisson goals model estimates how many goals each side is likely to score, which is what makes markets like Over/Under, BTTS and correct score possible from the same run.

The output you see is a probability, not a prediction of certainty. A 70% call is expected to be wrong three times in ten. That is the model working correctly, not failing.

What goes into it

Team strength as an Elo rating, updated after every match. Recent form over the last 5 and 10 games — points, goals scored, goals conceded. Head-to-head history between the two clubs. Home advantage, fitted per league rather than assumed. Rest days and fixture congestion. And the bookmaker's own price, where our feed carries one.

On expected goals, be precise about what we have: we compute an xG PROXY from shot counts, not true xG from a shot-quality model. In 21 leagues that proxy is shots-based and carries real information. In the other 13 the source data has no shot counts, so the proxy falls back to goals — meaning it adds nothing there. We would rather say that than imply a shot-quality model we do not have.

Elo ratings are published for 12 of the 35 leagues. Where a league is not covered, the reasoning panel says so instead of showing a number.

Odds, edge and value

Fair odds are simply 1 divided by our probability. They are a way of expressing the same number, not a price you can bet at — no bookmaker offers a margin-free price.

Where our feed carries a bookmaker price we de-vig it (strip the built-in margin) to get the market's implied probability, then compare. The difference is the edge.

We deliberately shrink our own probability toward the market rather than trusting it outright. The market is sharp, and an earlier version of this product that chased disagreement with the closing line performed markedly worse than one anchored to it. Being confident and being right are different things.

Using it

Reading a confidence score

The percentage on a pick is the model's probability that the selection lands. Bands are labelled Very high, High, Strong, Moderate and Long-shot, and they describe probability only.

None of them mean certain. An 85% pick loses roughly one time in seven; a 60% pick loses two times in five. If a run of high-confidence picks loses, that is within normal variance rather than evidence something has broken.

Accumulators and compounding

An accumulator combines several predictions into one bet. The odds multiply, so the return grows quickly — but every leg has to be right, and one wrong leg loses the whole slip.

The compounding is harsher than it looks. Five legs at 85% each land together only 44% of the time. In our own 30-day backtest, individual picks won 57% of the time while all five of a day's picks landing together happened on 1 day in 10.

Accumulators are the highest-variance way to use these predictions. Longer slips pay more precisely because they land less often.

How we report performance

The Track Record page leads with the aggregate: every published prediction, the overall win rate, the overall return, and the profit or loss — with the sample size beside every percentage. It is not filtered, and a negative return is shown as negative.

A win rate on its own is the most flattering and least useful number we could show you. A high hit rate on short-odds favourites still loses money, so we always show the return next to the rate.

Per-league and per-market breakdowns are available below the aggregate. They are there for detail, never as a substitute for the headline number.

What this cannot do

It cannot tell you what will happen. It estimates how often each outcome would occur across many repetitions of the same fixture.

It does not know about a manager sacked this morning, a squad rotation for a midweek cup tie, or a player ruled out an hour before kick-off. Injury data exists for Premier League clubs only, sourced from the Fantasy Premier League API — there is no injury feed for the other 33 leagues.

It has no expected line-ups, no weather, no squad valuations and no live in-play adjustment. Where a screen has nothing real to show, it says so rather than filling the space.

Responsible use

18+ only. Gambling can be addictive and gambling losses are real losses.

Only stake money you can afford to lose. Never chase a losing run by increasing stakes — the maths does not reward it, and it is the most common route from a hobby to a problem.

If betting has stopped feeling like entertainment, free and confidential support is available at BeGambleAware.org.

Prefer something to keep? The same material as a PDF.