Football analytics · Essay · July 31, 2026

Introducing Surprisal#

Adding predictability into value-based conventions of football analysis

Does predictable decision-making just mean good decision-making?

One Useful Analogy#

Imagine this for a change: You are a teacher in a class and you asked one of the students a question. Nevermind the question but there are two completely separate things you can judge from the student’s response.

1. Was the answer correct?; and

2. Was that the answer you expected from her?

As you, the teacher, stand there pondering on all of the possibilities while your student rambles on, you unfailingly come to a single set of outcomes.

Those two questions were independent and brought to you to the following situations where the answer was,

  • Correct and exactly what you expected — she delivered it with textbook accuracy as she always does.

  • Correct but entirely unexpected — she ended up at the right answer somehow that you could not see coming

  • Incorrect and exactly what you expected — she makes that mistake often and you know it well

  • Incorrect but entirely unexpected — concerning since the answer was so bizarre. Unusual of her.

Whether pedagogy is your cup of tea is immaterial here unfortunately. Football, like any sport, is a results-first business. Coloured by this exact candor, instead of asking both questions of the set, football analytics has historically only measured one half of it.

In-complete world of xG#

Existing value-based metrics endeavour to ask only that first question in the analogy above: Was the outcome good? If you are here even as a casual fan of football, you must have heard about xG and probably much to most casual fans’ chagrin, the stat economy using terms like these in football analysis is only growing more prevalent. xG or Expected Goals attempts to give a value to an attempted shot. It says “how likely is the player to score from that location in that state”.

There are other, more sophisticated metrics that attempt to do something similar, they all ask: “Was the outcome good”?

xT (or Expected Threat), an advanced metric introduced by Sarah Rudd and repurposed by Karun Singh, provides more context to xG by assigning a “threat-value” to a location on the football pitch and not just for the terminal action (the shot). Many recent complex metrics like VAEP or OBV greatly widen the context window to recent actions but essentially ask the same question “how good”. There can be another question to be asked: “Was the choice typical”. It is an easy mistake to make since it makes no sense to ask the second question in the absence of the first. It matters little whether a choice is typical if the outcome doesn’t even exist.

Introducing Surprisal#

The contrast between the concepts of overcoached teams and game-breakers, midfield monotones and risk-takers pulled me in the direction of the primary thesis of this work. Functionally, the following contrasts also need to be discussed.

Value vs Predictability#

Now, the differences stated between the two cannot be purely analogical. So, Surprisal, borrowing a concept from Claude Shannon's original information theory (1948), effectively gives the probability of an action given its state. Footballingly, it puts a number on how rare it is that my winger shoots from the byline instead of doing anything else like dribble, pass or cross. How is this different to value-based metrics, you ask? Surprisal has absolutely nothing to do with what happens in the future. It has no bearing on the outcome which is all that xG, xT or most of the advanced metrics are concerned about.

Mathematically, surprisal assigns a cost to an action becoming less likely in a given state:

Surprisal(actionstate)=logP(actionstate)\operatorname{Surprisal}(action \mid state) = -\log P(action \mid state)

Value-based metrics instead estimate the future outcome of a decision:

P(outcomestate,action)P(\operatorname{outcome} \mid state, action)

Spatial vs Sequential#

The most common lens in football analytics views football through spatial value models (xG/xT) but when we attempt to view football viewed how it’s played: every possession is a sequence of discrete tokens – (event_type, pitch_zone) pairs, one per action: (Pass, zone_7) → (Carry, zone_9) → (Pass, zone_12) → (Shot, zone_14). This is structurally identical to a sentence being a sequence of words. Modern language models are trained to do exactly one thing that is minimise the average surprisal of the next word given everything before it. So in direct modeling comparison, a possession is a sentence and a word that is added at the end is an action. “Surprisal” calculates how unexpected it is for that word to appear there. This is the exact technique I have replicated for the application of football analytics.

Full derivation of Surprisal and its possession-long derivative Perplexity is available on the project page.

I should strongly emphasise that Surprisal, a predictability-based metric, is meant to complement and not compete with value-based metrics like xG or xT. Recalling the student analogy from the beginning, Surprisal, being orthogonal to Value, adds to the one-dimensionality in the footballing context which on the pitch translates to a four outcome set that we saw earlier.

Let’s repeat the four combinations from before but this time applied to football:

A quadrant chart comparing surprisal and value

The upper-right quadrant holds high-value, high-surprise actions. #

For the uninitiated, the four combinations result in

  • Low surprise, high value — a well-drilled possession pattern that works because it's rehearsed

  • High surprise, high value — an unexpected, creative moment that pays off (an unscripted piece of individual brilliance).

  • Low surprise, low value — mundane, expected sideways/backward recycling that goes nowhere.

  • High surprise, low value — chaotic, undisciplined play that doesn't lead anywhere.

Surprise! It’s real — Validating the metric#

During the process of ideation of a predictability metric for the value dominant, I was wondering whether it will show on data and whether or not it is completely absorbed by the existing spatial metrics which do use the previous states and actions as information to predict the next outcome. So, to work out if Surprisal is really orthogonal to Value and independent of the already existing metrics, validation is necessary.

Since the idea of this metric came from the natural language models, validating it would undoubtedly be influenced by how language models are validated. To reiterate from the previous section, every possession is a series of discrete tokens, say, (Carry, zone_9) → (Pass, zone_12) → (Shot, zone_14) like words in a sentence. If the sequence was jumbled up here randomly to (Pass, zone_12) → (Shot, zone_14) → (Carry, zone_9), the surprisal metric should be through the roof as this sequence is too bizarre to actually happen in a game. Think of jumbling the word order in a sentence and see if you can make it make sense. Following this logic, randomised sequences should reliably score higher on Surprisal (therefore less predictable) than real possession sequences. The results exactly followed this through multiple datasets thereby proving this hypothesis and thus the viability of the metric.

You can play with demo below which showcases the argument above. Play with it and see if you can shuffle your way to a lower Surprisal score.

Real possession
Own thirdBuild-upMidfieldWide final thirdBox edgeInside box
Avg surprisal
1.29 bits
Shuffled version
Own thirdBuild-upMidfieldWide final thirdBox edgeInside box
Avg surprisal
1.29 bits
Click “Shuffle again” a few times to see how consistently the real order beats random ones.

You can't.

In language as in football, coherent sentences and real possessions consistently score lower average surprisal than their shuffled counterparts. A direct confirmation the model has learned genuine structure in how football is actually played, not just noise, before any of its output is trusted for analysis. The chart below illustrates the same with real data points collected over 8 seasons of football across 5 leagues.

Findings and Usability#

Forewarned in the previous sections, Surprisal on its own sits alone and fairly useless. Alongside a value-based metric, it has a friend and becomes really interesting. For this purpose, I used an xT model I had built from the first-principles so I wouldn’t make any mistakes due to faulty or misunderstood assumptions modelling further experiments — which you can find here. Unlike vanilla xT models, the model I built is turnover-aware which essentially means the threat at every location is not only determined by shots or moves (passes and carries) but also by the possibility of a turnover in possession. I believe this maps more accurately to players’ and coaches’ mindsets than just forward-facing models. Fear of failure is just as big a motivator as desire for success.

When plotted against xT delta per action – the difference of threat values from one action to the next, Surprisal per action resulted in the following funnel-shaped graph

Note: The following interactive plot is a down-sampled version of the fully-populated plot to make the load times faster. You can access it here.

What we can see and infer from this graph are close to real-world footballing observations:

  • Low-surprisal actions are densely populated towards the zero value of xT delta. Predictable actions cluster near zero precisely because they're common so they tend to be low-risk by nature.

  • High-surprisal actions are, almost definitionally, attempts that deviate from the well-worn pattern. It’s important to note that deviation can fail in either direction. It can fail because the attempt itself was low-percentage and gets cut out (a turnover in a promising position, dragging xT delta negative) or it can succeed precisely because the defence wasn't organised for something outside the normal pattern (a genuine breakthrough, dragging xT delta strongly positive).

What are some possible concrete, practical uses of this? I’ve got a few:

  1. Game-state-dependent strategy. This is a well-established idea in decision theory and other sports (chess players deliberately play sharper, riskier lines when losing on the clock; underdogs in many sports statistically benefit from higher-variance strategies): when behind, low expected-value-difference strategies don't help — you need variance with quality, to escape a losing position. When protecting a lead, the opposite is true — you want to minimize variance even at some cost to average value, since a single high-upside failure (losing the ball in your own third chasing a hopeful pass) is disproportionately costly when the game state punishes risk. A coach could use a live version of this metric to say something like "we're 1-0 down with 15 minutes left — our current pattern is running at low average surprisal; we need to deliberately break our own patterns" — a genuinely quantified, real-time signal for exactly the kind of adjustment coaches already try to make by feel.

  2. Player profiling. A player's average surprisal, combined with their own value-variance, gives you a real, separable trait: some players are "control" players (low surprisal, tight value distribution — reliable, low-risk contributors) and others are "variance generators" (high surprisal, wide value distribution — capable of the moment of magic, but also more likely to give the ball away). This is genuinely different information than xG or xT totals alone, and it's exactly the kind of individual signal a recruitment analyst would want when trying to answer "do we need control in this position, or do we need someone who can unlock a low block."

  3. Opponent scoutability. If an upcoming opponent's danger is mostly generated through low-surprisal patterns — repeated, drilled sequences — that's genuinely more preparable in advance, since a pattern predictable enough for your model to learn is plausibly predictable enough for a defensive coach to prepare a specific plan against. A team whose threat comes disproportionately from high-surprisal moments is structurally harder to scout, because there's no stable repeated shape to drill against — the danger comes from individual moments, not system.

Objections and Limitations#

The big one#

Why does this metric exist at all? Isn't it just a fancier, over-engineered way to express xT?"

If you weren't convinced by my conceptual and epistemical ramblings why Surprisal is something different, I came prepared. The weak correlation (0.20) is itself proof that surprisal and value are answering genuinely different questions. If they were strongly correlated, surprisal would just be a noisier, more roundabout way of measuring what xT already measures, and wouldn't be worth building. The fact that they are only weakly linked is the actual evidence this is a legitimate second dimension of analysis, not a redundant one.

For more technical readers, the rest of this section might be of interest. If you don’t find yourself in that camp, feel free to skip ahead.

Let’s snap back to what was found and verified during the experiments, the core claim being real football sequences are more predictable than random ones, and predictability correlates with narrower value spread. This survived a shuffled-baseline check, event-type stratification, and a formal bootstrap test against sample-size artifacts.

For the scientific rigour valuing folks, the rest of this is about what the above scrutiny didn't cover.

  1. The causal horror. There is still a question of a lurking "phase-of-play chaos" variable driving both surprisal and value-variance independently that was never tested, because it can't be with event data alone. Everything said about why the funnel exists is a plausible story, not a demonstrated one. The honest claim is still "surprisal and value-variance are robustly associated," not "predictability causes narrower variance."

  2. Surprisal is relative to the training corpus, not an objective property of the action. This is worth stating explicitly because it's easy to misread: a given action's surprisal score isn't "how surprising a knowledgeable coach would find it" — it's "how rare this exact (context, token) pair was in this specific pooled dataset of five leagues, 2018-2024." Train the same model on a different set of leagues, a different era, or a different competition level, and the identical real-world action could score completely differently. Any claim phrased as if surprisal measures something intrinsic to the action, rather than something relative to the model's training distribution, overstates what's actually been built.

  3. The man that is Markov. The model only ever conditions on the single preceding action — never on the score, the minute, the opponent, or anything more than one step back. Two actions that are contextually worlds apart (a give-and-go in minute 3 of a 0-0 versus the identical zone-transition in minute 93 of a game chasing an equaliser) get identical surprisal, because the model literally cannot see the difference.

  4. The Three Stooges – Errors compound across models. The xT-delta cross-reference isn't just "surprisal versus ground truth" — it's surprisal versus xT, which itself depends on xG, and both carry their own already-documented limitations. The funnel finding is only as trustworthy as the weakest link in that three-model chain, not just the surprisal model on its own.

  5. Statistical significance isn't the same as practical significance. The bootstrap test proved the variance increase is real, at p < 0.0005. With a sample in the hundreds of thousands, almost any genuine effect will clear that bar — the test says the pattern isn't noise, it says nothing about whether the size of the effect is large enough to change a real coaching decision. That's a distinct, unaddressed question. Also the most potent.

This is all you can extract from me for now. The rest of the scathing criticism of my work and time are listed and analysed in the project page. I wonder if I have mentioned it enough times?

Conclusion#

So, Professor [insert your name], was the answer correct, or was it expected? The student who gives the correct answer gets the same marks as the student with the most expected one. Most measurement systems generally don’t separate the two. It’s something that happens when a field thinks in a single axis because then the second-axis becomes harder to see.

We’ve established how most of the analytics work goes into a single axis. Just like any other system, this myopic view has a real cost in how the sport gets evaluated. Scouting and player development lean, structurally, toward rewarding the predictable. A player or an approach whose value comes from doing the unexpected thing is much harder to build a report around, precisely because the same trait that makes them dangerous also makes them error-prone, and a single-axis view of "how much did this help" can't tell those two costs apart. Quantifying the second-axis is not only an analytics curiosity but rather it's a real lever on the question of who a club decides is worth developing versus worth flagging as inconsistent, based on identical numbers on the stat sheet.

So walk with me,

Doesn’t predictable decision-making just mean good decision-making?

Not necessarily.

I and a lot of football watchers know this intimately but it’s hard to put a number next to it. That’s the difference between Manchester City’s centurions and Real Madrid’s three-peat UCL winning dark-magic xG breakers, to make a poor analogy. Top-level football analytics, scouting and data-driven coaching is such a tightly-knit community that it’s hard to see from the outside through their opaque windows to ascertain which approaches are being currently used or favoured internally. There are still a lot of questions regarding the specificity and applicability of Surprisal and its big sister Perplexity (also a language model term). With the dearth of publicly available data however, those questions are even harder to answer for now. We mourn the passing of FBRef and yet we persevere.