<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://seanelvidge.github.io/feed.xml" rel="self" type="application/atom+xml"/><link href="https://seanelvidge.github.io/" rel="alternate" type="text/html" hreflang="en"/><updated>2026-08-16T03:52:21+00:00</updated><id>https://seanelvidge.github.io/feed.xml</id><title type="html">Sean Elvidge</title><entry><title type="html">Space Weather - The Musical</title><link href="https://seanelvidge.github.io/articles/2026/Space_Weather_Sonification/" rel="alternate" type="text/html" title="Space Weather - The Musical"/><published>2026-06-17T22:00:00+00:00</published><updated>2026-06-17T22:00:00+00:00</updated><id>https://seanelvidge.github.io/articles/2026/Space_Weather_Sonification</id><content type="html" xml:base="https://seanelvidge.github.io/articles/2026/Space_Weather_Sonification/"><![CDATA[<p>What does a geomagnetic storm sound like?</p> <p>Not metaphorically. Literally.</p> <p>Imagine taking the indices we use to drive our models of near-Earth space (solar radio flux, sunspot number, geomagnetic activity, ring-current disturbance) and translating it into a piece of music. This isn’t a “space-inspired” soundtrack, but a deterministic piano score where every pitch, rhythm, chord, tempo change and accent is driven by real space weather indices.</p> <p>That is what this blog post is all about. About how we turn space weather events into piano music. The result is a structured musical translation of the space environment, built so that the data can be heard.</p> <h2 id="the-may-2024-superstorm">The May 2024 Superstorm</h2> <figure> <iframe src="/assets/sonification/May_Storm_20240505-20240515.pdf" class="rounded z-depth-1" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="" width="100%" height="700" title="May 2024 storm score"/> </figure> <figure> <audio src="/assets/sonification/May_Storm_20240505-20240515.mp3" controls=""/> </figure> <p>(Click play and listen along whilst you read the rest of the post explaining where the music comes from)</p> <h2 id="a-storm-compressed-into-the-hands-of-a-pianist">A storm, compressed into the hands of a pianist</h2> <p>The example above uses data from May 5th to 15th 2024. This 10-day window captures the period around the major May 2024 geomagnetic storm — an event of intense aurora, severe geomagnetic activity and unusually (at least in recent times) dynamic near-Earth conditions.</p> <p>Musically, the piece is placed onto a fixed daily grid:</p> <ul> <li>one day becomes four bars;</li> <li>each bar is in 6/4;</li> <li>one day is therefore 24 quarter-note beats;</li> <li>one hour corresponds to one quarter-note beat;</li> <li>thirty minutes corresponds to half a beat.</li> </ul> <p>So however expressive the notes become (more details below), the structure remains anchored in time. Every layer of space weather indices realigns at midnight. Each day takes up the same “musical space”, allowing the changing behaviour of the Sun-Earth system to become clear through changes in texture, pitch, rhythm and intensity.</p> <h2 id="five-indices-five-jobs">Five indices, five jobs</h2> <p>Here we use five space weather indices, and each are given a different job to do in defining the music:</p> <ul> <li>F10.7: the adjusted solar radio flux, controls the harmonic root in the left hand.</li> <li>Sunspot number: adds ‘density’ to the left-hand. When the sunspot number is sufficiently high, we add an octave root.</li> <li>Hp30: drives the right-hand melody. Because Hp30 is available every 30 minutes, it provides a natural melodic line.</li> <li>Kp: controls the broad rhythmic regime. Quiet geomagnetic periods lead to slower figures; storm periods create denser patterns.</li> <li>Dst: as it becomes more negative, the music becomes more forceful through dynamics, accents and tempo.</li> </ul> <p>I was keen to use as many indices as possible to give the music as much ‘texture’ as I could do - entirely deterministically. I have tried to make solar activity shape harmony, geomagnetic activity shape motion and storm intensity shape tension.</p> <h2 id="the-right-hand-melody-hp30">The right hand melody: Hp30</h2> <p>The right hand is the most active. It uses a pitch drawn from F# minor (my favourite), ranging from F#4 to E6. Each Hp30 value is normalised and mapped to one of fourteen pitches.</p> <p>This means that higher geomagnetic activity tends to push the melody upward. But the mapping is not absolute. To create the score we do some local normalisation, i.e. we care both about how large Hp30 is on a physical scale and how large it is relative to the other values in the selected date range. This was needed so that even quite periods (geomagnetically) stay interesting.</p> <p>The Kp (a logarithmic scale from 0 to 9) is used to define the rhythm:</p> <ul> <li>below Kp 5, the rhythm is relatively slow;</li> <li>from Kp 5 to 7, the rhythm becomes more active;</li> <li>above Kp 7, each hour becomes a four-note semiquaver figure.</li> </ul> <p>But the exact rhythm inside each hour also responds to local movement in Hp30 and Dst. If Hp30 jumps within the hour, or if Dst changes sharply from one hour to the next, the right hand becomes more animated. A quiet Kp period is therefore not forced to be dull if the other data are still moving.</p> <h2 id="the-left-hand-harmony-and-tension">The left hand: harmony and tension</h2> <p>The left hand provides the harmonic ‘frame’ of the piece. It chooses roots from an ordered set:</p> <p>F#, A, B, C#, D, E, G#.</p> <p>These are not arranged chromatically but arranged to give useful harmonies inside the F# minor world we’re working in.</p> <p>Each day has four bars, and each bar receives one harmonic root. F10.7 selects the base root. Then the day’s pattern of F10.7 and sunspot-number change determines how the four roots move. If solar activity and sunspots are both rising, the harmony tends to climb. If both are falling, it descends. If they disagree, the progression takes a mixed path.</p> <p>The chords themselves are simple diatonic triads: F# minor, A major, B minor, C# minor, D major, E major and G# diminished. This keeps the harmonic language coherent while still allowing the data to move the music through different regions of the scale.</p> <p>Dst and Kp then decide how much weight the left hand carries. In low-tension periods, a bar may simply be one long chord. In moderate tension, a bass note is added before the chord. In high tension, the left hand pulses with repeated bass-plus-chord figures. A storm is therefore not only heard in the treble melody but also changes the ‘weight’ of the piano texture.</p> <h2 id="tempo-dynamics-and-accents">Tempo, dynamics and accents</h2> <p>The score also has some performance notes in it. Whilst the overall tempo of the piece remains constant throughout, average note length is (slightly) determined by solar activity, Hp30, Dst, and Kp can all push the tempo upward (having an impact of making the piece feel like the speed is changing, but only really ranges from about 60 to 120 bpm so remains playable (at least to my limits!).</p> <p>We also use four dynamic levels: p, mp, mf and f. This allows for clear changes in intensity without going over the top. Accents appear when the normalised activity is high enough. Staccato is added to the shortest right-hand notes. During intense storm intervals, the music becomes not just higher or faster, but sharper and more articulated.</p> <h2 id="why-this-is-more-than-sonification">Why this is more than sonification</h2> <p>Many data-to-music projects work by assigning one variable to pitch and another to volume. That can be effective, but it often produces a thin musical result. This piano mapping is a little more ambitious. It treats the space weather system as a set of interacting musical lines:</p> <ul> <li>long-timescale solar conditions become harmony;</li> <li>short-timescale geomagnetic changes become melody;</li> <li>storm thresholds become rhythmic regimes;</li> <li>Dst depression becomes tension;</li> <li>sunspot number becomes chord density;</li> <li>daily evolution becomes harmonic progression.</li> </ul> <p>This is also all deterministic. Given the same date range then the same notes, rhythms, dynamics and accents are produced. There are no random choices. This is important because it means the musical output is reproducible. The piece can be discussed, analysed and regenerated.</p> <h2 id="hearing-the-may-2024-storm">Hearing the May 2024 storm</h2> <figure> <audio src="/assets/sonification/May_Storm_20240505-20240515.mp3" controls=""/> </figure> <p>In calm moments of the space environment, the music has ‘space’. The left hand can hold long chords while the right hand moves slowly through the scale. As activity increases, the melody becomes more agile. Rhythms subdivide. The tempo rises. Accents appear. The left hand gains pulse. The storm becomes audible not as noise, but as structure.</p> <p>That is the most compelling aspect of the approach. Space weather is often communicated through plots, maps, alerts and indices. Those are essential. But music offers a different route into the same system. It lets us hear change, pressure, release, escalation and recovery.</p> <p>The May 2024 storm was a scientific event, an operational challenge and, for many people, a spectacular visual experience. In this piano version, it becomes something else as well: a short, reproducible musical portrait of a disturbed geospace environment.</p>]]></content><author><name></name></author><category term="spaceWeather"/><category term="mathematics"/><summary type="html"><![CDATA[What does a geomagnetic storm sound like? Not metaphorically. Literally.]]></summary></entry><entry><title type="html">Who wins in a 48 team World Cup - Panini</title><link href="https://seanelvidge.github.io/articles/2026/A_48_Team_World_Cup/" rel="alternate" type="text/html" title="Who wins in a 48 team World Cup - Panini"/><published>2026-06-11T19:30:00+00:00</published><updated>2026-06-11T19:30:00+00:00</updated><id>https://seanelvidge.github.io/articles/2026/A_48_Team_World_Cup</id><content type="html" xml:base="https://seanelvidge.github.io/articles/2026/A_48_Team_World_Cup/"><![CDATA[<p>Since 1970 in Mexico <a href="https://paninistore.com/">Panini</a> have published the official World Cup sticker book. It has become synonymous with the tournament, with a huge following and has become a multi-billion dollar business. It is estimated that for this World Cup between 500 and 800 million sticker packs will be sold <a href="https://www.marca.com/en/world-cup/2026/06/05/world-cup-stickers-1-4-billion-euro-business-with-an-underlying-controversy.html">[1]</a>.</p> <p>So what does it take to complete the sticker book? And how has this changed from the last 32-team World Cup (Qatar 2022) to this, the first 48-team World Cup? Well, the 50% increase in teams, “only” leads to a 10% increase in required stickers (but a 42% increase in costs…).</p> <p>For this analysis we assume that stickers are randomly distributed in packs and are each equally likely (which has been claimed to be the case by the Panini CEO, however <a href="https://academic.oup.com/jrssig/article/18/3/5/7038156?searchresult=1&amp;login=false">research has shown</a> that ‘shiny’ stickers are systematically rarer than others by about 2x).</p> <p>To see where these numbers come from, let \(N\) by the total number of stickers in the album and \(k\) be the number of stickers in one pack. So each pack contains \(k\) <strong>distinct</strong> stickers which are sampled uniformly at random from the \(N\) possible stickers. So each pack is equally likely to be any one of</p> \[{N\choose k}\] <p>which means “\(N\) choose \(k\)” and, simply, is used to represent the number of ways to choose \(k\) items from a set of \(N\) items where order does not matter. You can calculate its numerical value using:</p> \[{N\choose k} = \frac{N!}{(N-k)!k!}\] <p>where “!” means <a href="https://en.wikipedia.org/wiki/Factorial">factorial</a>.</p> <p>So the number we are trying to find is:</p> \[E_m = \text{ expected number of additional packs needed when you already have } m \text{ distinct stickers}\] <p>and in our specific case the goal is to compute \(E_0\) because we start with zero (0) stickers. The boundary condition is \(E_N = 0\) because once you already have all \(N\) stickers, you need to more packs. Now suppose you currently have \(m\) distinct stickers, then</p> \[N-m\] <p>stickers are still missing. In the next pack of \(k\) stickers, suppose exactly \(r\) of them are new stickers. To get exactly \(r\) new stickers:</p> <ul> <li>choose \(r\) stickers from the \(N-m\) missing stickers,</li> <li>choose the remaining \(k-r\) stickers from the \(m\) stickers you already have.</li> </ul> <p>So the number of packs that contain exactly \(r\) new stickers is:</p> \[\binom{N-m}{r}\binom{m}{k-r}.\] <p>Since all possible packs are equally likely, the probability of getting exactly \(r\) new stickers is:</p> \[p_{m,r} = \frac{\binom{N-m}{r}\binom{m}{k-r}}{\binom{N}{k}},\] <p>where invalid binomial terms are treated as zero. Then, the possible values of \(r\) are:</p> \[\max(0,k-m) \le r \le \min(k,N-m)\] <p>because you cannot get more new stickers than are missing, and you cannot get more than \(k\) new stickers from one pack. So now we need fewer stickers to complete our collection, specifically we move from ‘state’ \(m\) to state \(m+r\) with probability \(p_{m,r}\). Therefore:</p> \[E_m = 1 + \sum p_{m,r}E_{m+r},\] <p>where the “1” represents the pack you just bought. However, one possible outcome is \(r=0\), meaning the pack we bought contains no new stickers. In that case we remain at state \(m\), so \(E_m\) appears on both sides:</p> \[E_m = 1 + p_{m,0}E_m+\sum_{r\ge 1}p_{m,r}E_{m+r},\] <p>which we can rearrange to get:</p> \[\begin{eqnarray*} E_m - p_{m,0}E_m &amp;=&amp; 1+\sum_{r\ge 1}p_{m,r}E_{m+r},\\ E_m(1-p_{m,0}) &amp;=&amp; 1 + \sum_{r\ge 1}p_{m,r}E_{m+r},\\ E_m &amp;=&amp; \frac{1+\sum_{r\ge 1}p_{m,r}E_{m+r}}{1-p_{m,0}}. \end{eqnarray*}\] <p>This gives an exact recursive calculation. Since \(E_m\) depends only on \(E_{m+1}, E_{m+2},\ldots,E_N\) we compute backwards from:</p> \[E_N=0\] <p>down to</p> \[E_0.\] <p>In 2022 each Panini pack contained 5 stickers and the whole album required 670 stickers. So that gives \(N=670\) and \(k=5\), using the recurrence relation described above gives:</p> \[E\approx 946.9837488\] <p>so the expected number of packs required to complete the album is 947 packs.</p> <p>In 2026, each pack contains 7 stickers and an album needs 980 stickers. So with \(N=980\) and \(k=7\), and using the defined relation the expected number of packs is:</p> \[E_0 \approx 1042.36\] <p>or about 1042 packs. The increase from 947 packs to 1042 packs is an increase of about 10% (10.031…% to be more precise).</p> <p>A multi-pack of stickers at my local supermarket costs £7.50 for 6 packs. To get the expected number of 1042 packs requires 174 multi-packs at a cost of £1,305.</p> <p>For comparison, in 2022, a 6 pack multi-pack (remembering each contained 5 stickers per pack), cost £4.99, so to get the expected number of 947 packs, 158 packs were needed, at a total cost of £788.42. Adjusting for inflation that comes to £918.89.</p> <p>So whilst we only need 10% more stickers, we should expect to spend 42% more to complete the album.</p>]]></content><author><name></name></author><category term="football"/><category term="mathematics"/><summary type="html"><![CDATA[Who is the biggest winner of the new, larger, 48 team World Cup? Panini, with the increase in the numbers of stickers to complete their album.]]></summary></entry><entry><title type="html">When Good Ideas Don’t Survive Contact With Data - A December Football Myth Tested</title><link href="https://seanelvidge.github.io/articles/2025/A_December_Football_Myth/" rel="alternate" type="text/html" title="When Good Ideas Don’t Survive Contact With Data - A December Football Myth Tested"/><published>2025-12-31T19:30:00+00:00</published><updated>2025-12-31T19:30:00+00:00</updated><id>https://seanelvidge.github.io/articles/2025/A_December_Football_Myth</id><content type="html" xml:base="https://seanelvidge.github.io/articles/2025/A_December_Football_Myth/"><![CDATA[<p>Every so often I come up with what feels like a great idea for a blog post. A clever angle, an interesting hypothesis, something that surely must be hiding in the data just waiting to be revealed. And then, after hours of analysis, the result is… nothing. Completely flat. No effect.</p> <p>This is one of those stories.</p> <p>I wanted to look at whether the famously congested English football December schedule gives an advantage to teams (either weaker or stronger teams). With matches every few days, tired legs, rotation, winter weather, and general chaos, it seemed plausible that perhaps squad rotation from the stronger sides might be key, or perhaps stronger sides might stumble more than usual and that underdogs might pick up a few unexpected points.</p> <p>It’s a nice idea. Unfortunately, as we’ll see, a nice idea does not guarantee a nice result.</p> <h2 id="why-december-might-matter">Why December Might Matter</h2> <p>Consider what happens in December in English football:</p> <ul> <li>Fixture congestion increases dramatically.</li> <li>Elite teams face multiple competitions and may rotate heavily.</li> <li>Injuries accumulate and squad depth becomes critical.</li> <li>Weather worsens, pitches soften, and match conditions become less predictable.</li> </ul> <p>Perhaps if stronger teams depend more on structure and precision, and weaker teams rely more on disruption and variance, you can build a reasonable argument that December’s chaos environment might narrow the gap.</p> <p>In other words, if \(S\) is team strength and \(\epsilon\) is “football randomness”, maybe the effective strength looks more like:</p> \[S_{\text{effective}} = S + \epsilon,\] <p>and perhaps December has a larger \(\epsilon\).</p> <p>If so, underdogs might perform better than expected.</p> <h2 id="how-to-measure-underdog-performance">How to Measure “Underdog Performance”?</h2> <p>To test this we need a <a href="https://seanelvidge.com/articles/2024/All_England_football_league_results/">very large dataset of football results</a> and a way to <a href="https://seanelvidge.com/articles/2025/Football_team_rankings/">quantify which team was stronger <em>before</em> the match</a> was played.</p> <p>In the team strength database a higher values means a stronger team.</p> <p>Thus:</p> <ul> <li>The stronger team is whichever side has the higher rank.</li> <li>The weaker team (the “underdog”) is the side with the lower rank.</li> </ul> <p>And for each match, we compute:</p> <ul> <li>Underdog points</li> <li>Underdog goal difference</li> <li>Whether the match took place in December</li> </ul> <p>Everything you’d need to test whether December helps the weaker (or stronger) team.</p> <h2 id="december-vs-non-december-the-raw-numbers">December vs Non-December: The Raw Numbers</h2> <p>Aggregating across every division, every season, and every match where a clear favourite exists, we get:</p> <ul> <li>Outside December, underdogs earn: <strong>1.178 points per match</strong></li> <li>In December, underdogs earn: <strong>1.181 points per match</strong></li> </ul> <p>The difference, <strong>0.0026 points</strong>, is essentially zero.</p> <div class="row mt-3"> <div class="col-sm mt-3 mt-md-0"> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/ppm_dec_nonDec-480.webp 480w,/assets/img/ppm_dec_nonDec-800.webp 800w,/assets/img/ppm_dec_nonDec-1400.webp 1400w," sizes="95vw" type="image/webp"/> <img src="/assets/img/ppm_dec_nonDec.png" class="img-fluid rounded z-depth-1" width="100%" height="auto" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> </figure> </div> </div> <p>The underlying distribution of goal differences tells the same story.</p> <div class="row mt-3"> <div class="col-sm mt-3 mt-md-0"> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/Underdog_Goal_Differences-480.webp 480w,/assets/img/Underdog_Goal_Differences-800.webp 800w,/assets/img/Underdog_Goal_Differences-1400.webp 1400w," sizes="95vw" type="image/webp"/> <img src="/assets/img/Underdog_Goal_Differences.png" class="img-fluid rounded z-depth-1" width="100%" height="auto" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> </figure> </div> </div> <p>I will admit, this was mildly disappointing. But maybe aggregate data hides the truth? Perhaps the Premier League behaves differently? Or lower leagues, with smaller squads, show something?</p> <p>Nope.</p> <p>However you slice the dataset, Tier 1, Tier 2, Tier 3, Tier 4, the same conclusion emerges:</p> <ul> <li>The difference between underdog performance in December vs other months is tiny.</li> <li>In every tier, the effect is statistically insignificant.</li> <li>Some tiers lean slightly positive, others slightly negative, all within noise.</li> </ul> <p>So weaker teams, as a group, do <strong>not</strong> seem to benefit from the December schedule.</p> <p>You might reasonably ask whether the <strong>favourites</strong> show any seasonal change. If weaker teams don’t improve, perhaps stronger teams deteriorate?</p> <p>I ran the same comparison for the stronger side in each match.</p> <p>The result? Exactly the same story.</p> <p>Stronger teams’ points and goal differences in December are statistically indistinguishable from the rest of the season. No noticeable dip, no seasonal weakness, no hidden December curse.</p> <p>In short:</p> <blockquote> <p>Neither underdogs nor favourites show any meaningful change in performance during December.</p> </blockquote> <h2 id="a-more-formal-statistical-model">A More Formal Statistical Model</h2> <p>To be thorough, I fitted a regression model of underdog points:</p> \[\text{pts} = \beta_0 + \beta_1\,\text{December} + \beta_2\,\text{Home} + C(\text{Tier}) + C(\text{Season}) + \varepsilon .\] <p>This controls for:</p> <ul> <li>Home advantage</li> <li>League tier</li> <li>Season-by-season variation</li> </ul> <p>The coefficient for the December effect came out as:</p> \[\beta_1 = -0.0018 \pm 0.013,\] <p>which is, again, indistinguishable from zero.</p> <h2 id="when-a-good-hypothesis-fails">When a Good Hypothesis Fails</h2> <p>This leads to what I think is the most important lesson from the whole exercise.</p> <p>In science many ideas do not survive contact with real data. They’re plausible, they’re elegant, and they would make for a great story, but the universe simply refuses to cooperate.</p> <p>That does not make the work wasted.</p> <p>It makes it honest.</p> <h2 id="the-academic-problem-with-negative-results">The Academic Problem With Negative Results</h2> <p>In academia, “negative” or null results are notoriously hard to publish. Journals often prefer dramatic or surprising findings, which unintentionally encourages selective reporting:</p> <ul> <li>Effects that <strong>don’t</strong> exist are quietly forgotten.</li> <li>Effects that <strong>appear by chance</strong> get published.</li> <li>The literature accumulates exciting stories but not necessarily accurate ones.</li> </ul> <p>This December analysis is a textbook example: an interesting idea, diligently tested, clearly unsupported.</p> <p>Yet these non-effects matter. They give us a clearer picture of reality, and they remind us that the absence of a pattern is itself information, and sometimes, important information.</p>]]></content><author><name></name></author><category term="football"/><category term="mathematics"/><summary type="html"><![CDATA[Every so often I come up with what feels like a good idea for a blog post. But they don't always work out, this is the story of my investigations into the congested English football December schedule and whether it gives certain teams an advantage.]]></summary></entry><entry><title type="html">Christmas Day Football - The Lost Tradition</title><link href="https://seanelvidge.github.io/articles/2025/Christmas_Day_Football/" rel="alternate" type="text/html" title="Christmas Day Football - The Lost Tradition"/><published>2025-12-25T09:00:00+00:00</published><updated>2025-12-25T09:00:00+00:00</updated><id>https://seanelvidge.github.io/articles/2025/Christmas_Day_Football</id><content type="html" xml:base="https://seanelvidge.github.io/articles/2025/Christmas_Day_Football/"><![CDATA[<p>It’s been exactly 60 years since the last English league match was played on Christmas Day. On 25 December 1965, Blackpool hosted Blackburn Rovers in a festive Lancashire derby that would become the final flicker of a long-standing tradition. While Boxing Day remains a staple of the football calendar, the idea of heading to the ground after unwrapping presents has quietly faded into the past. Here’s a look back at how Christmas Day football began, what made it special, and why it came to an end.</p> <h2 id="victorian-beginnings">Victorian Beginnings</h2> <p>The Football League didn’t include Christmas Day matches in its inaugural season (1888), but by the very next year, Preston North End hosted Aston Villa on 25 December 1889. That 3-2 win for Preston drew around 9,000 fans and marked the beginning of Christmas football as a regular fixture.</p> <p>By the early 20th century, the tradition was well-established. Local derbies were often scheduled to reduce travel, and many clubs played the same opponent home and away over Christmas and Boxing Day. It wasn’t unusual for fans to watch a game in the morning and be home in time for turkey in the afternoon.</p> <h2 id="when-christmas-meant-football">When Christmas Meant Football</h2> <p>For much of the 20th century, football on Christmas Day was as normal as carols and crackers. The interwar years and post-WWII period were the high points, with full fixture lists and packed crowds. In 1949, over 3 million fans attended league matches during the Christmas week. Players would sometimes play three matches in four days. Fans shared flasks and sang carols on the terraces.</p> <p>Holiday scheduling quirks added to the charm. Tranmere once lost 4–1 on Christmas Day, only to win the return fixture 13–4 the next day. Matches were sometimes played in snow, fog, or biting cold, but still the crowds came.</p> <p>Christmas Day 1914 was particularly notable. Despite the shadow of World War I, nine First Division matches went ahead, drawing a combined crowd of around 173,000. Football was still seen as a morale-booster, and a newspaper at the time declared that to millions:</p> <blockquote> <p>Christmas without football would not be Christmas at all. (<a href="https://www.britishnewspaperarchive.co.uk/viewer/bl/0000530/19451221/019/0003">Essex Newsman, 21 December, 1945</a>)</p> </blockquote> <p>And across the Channel that same day, one of football’s most poignant and enduring moments occurred. British and German troops along the Western Front took part in the famous 1914 Christmas Truce. In the frozen fields of Flanders, soldiers from both sides briefly set down their arms and came together to sing carols, exchange gifts, and even play informal matches in No Man’s Land.</p> <h2 id="christmas-day-classics">Christmas Day Classics</h2> <p>There were high-scoring thrillers (Chelsea 7–4 Portsmouth in 1957), and even a match in 1940 where Norwich beat a scratch Brighton team 18–0. The largest Christmas Day crowds topped 60,000 at places like Newcastle and Arsenal.</p> <p>Some stories veer into folklore: players turning up tipsy after too much Christmas cheer, supporters pelting the pitch with orange peel, and squads sharing a turkey dinner on the train between back-to-back games.</p> <p>And for Coventry fans: our own Ken Satchwell scored the last ever Football League hat-trick on Christmas Day, in a 5–3 win over Wrexham in 1959. One of the last festive hurrahs before the tradition faded.</p> <h2 id="the-final-whistle-25-december-1965">The Final Whistle: 25 December 1965</h2> <p>60 years ago today, in 1965, Blackpool beat Blackburn Rovers 4–2 in front of just over 20,000 fans. Young Alan Ball, not yet a World Cup winner, scored one of the goals. The weather was chilly but the game was lively, and the matchday programme was wrapped in festive green and tangerine.</p> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/Blackpool-Xmas-Day-programme-480.webp 480w,/assets/img/Blackpool-Xmas-Day-programme-800.webp 800w,/assets/img/Blackpool-Xmas-Day-programme-1400.webp 1400w," sizes="95vw" type="image/webp"/> <img src="/assets/img/Blackpool-Xmas-Day-programme.jpg" class="img-fluid rounded z-depth-1" width="100%" height="auto" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> </figure> <p>That game became, quietly, the last of its kind.</p> <h2 id="why-it-ended">Why It Ended</h2> <p>The reasons are a mix of culture, logistics, and changing times. More families preferred to stay home on Christmas morning. Public transport no longer ran on Christmas Day. Policing and stewarding became harder to organise. Players were increasingly reluctant to spend Christmas away from their families.</p> <p>The spread of floodlights and TV also meant fixtures could be moved to evenings or rescheduled around the holidays. Boxing Day, a public holiday since the 19th century, proved more popular for fans and clubs alike. By the early 60s, most clubs had already opted out. No ban was needed. The tradition just slipped away.</p> <h2 id="remembering-christmas-football">Remembering Christmas Football</h2> <p>Christmas Day football was sometimes chaotic, often charming, and always memorable. It belongs to another era, one of packed terraces, handwritten programmes, and roaring fires waiting at home. Brentford tried to make it come back in 1983 when the tried to schedule a game for Christmas morning. But fans rebelled and the club backed down.</p> <p>These days, players can enjoy their mince pies in peace, and fans save their voices for Boxing Day.</p> <p>But 60 years ago today, Blackpool beat Blackburn in the last of 1,243 Christmas Day fixtures, as English football took its final bow on December 25th.</p>]]></content><author><name></name></author><category term="football"/><summary type="html"><![CDATA[It's been exactly 60 years since the last English league match was played on Christmas Day. Here’s a look back at how Christmas Day football began, what made it special, and why it came to an end.]]></summary></entry><entry><title type="html">How to Calculate League Position Probabilities</title><link href="https://seanelvidge.github.io/articles/2025/League_table_prediction_probabilities/" rel="alternate" type="text/html" title="How to Calculate League Position Probabilities"/><published>2025-12-22T19:30:00+00:00</published><updated>2025-12-22T19:30:00+00:00</updated><id>https://seanelvidge.github.io/articles/2025/League_table_prediction_probabilities</id><content type="html" xml:base="https://seanelvidge.github.io/articles/2025/League_table_prediction_probabilities/"><![CDATA[<p>By mid-season, football fans instinctively start doing probability in their heads.</p> <blockquote> <p>If we win our next two, and they drop points away at… actually, hang on.</p> </blockquote> <p>This usually ends with a napkin full of scribbles and a strong emotional commitment to a particular set of results. What follows in this post an attempt to do the same thing — but with statistics.</p> <p>On the <a href="https://seanelvidge.com/tableProbs">predicted league tables</a> page of this site you’ll find the probabilities for every team finishing in every possible league position. Not via vast simulations. Not using “expected points” (xPts). Not vibes. Actual probabilities.</p> <p>This post explains how those tables are built, what assumptions sit underneath them, and, importantly, what they don’t claim to do.</p> <h2 id="estimating-the-future">Estimating the future</h2> <p>For each remaining fixture, we estimate the probability of:</p> <ul> <li>a home win,</li> <li>a draw,</li> <li>an away win.</li> </ul> <p>These probabilities come from my <a href="https://seanelvidge.com/articles/2025/Football_team_rankings/">rating model</a> that combines:</p> <ul> <li>team strength,</li> <li>home advantage,</li> <li>historical era-dependent draw rates,</li> </ul> <p>(you can try the <a href="https://seanelvidge.com/matchProbs">match probability calculator</a> yourself).</p> <p>For these forecasts we assume that team strength remains constant for the rest of the season. This is not because teams don’t change, obviously they do, but because this assumption lets us ask a very precise question:</p> <p>“Given what we know right now, what does the future look like?”</p> <p>No form, no momentum, no injuries, no managerial bounce. Just the present frozen in amber.</p> <h2 id="turning-matches-into-points">Turning matches into points</h2> <p>Once we know the probability that a team finishes on, say, 64 points, we can ask the next question:</p> <p>“What position will that correspond to?”</p> <p>For each possible final points total:</p> <ul> <li>we calculate the probability that other teams finish above that total,</li> <li>the probability they finish below it,</li> <li>and the probability they finish on exactly the same points.</li> </ul> <p>When teams are tied on points, we assume a random tie-break between them. This is not how the real leagues work, goal difference matters, but it is a deliberately neutral assumption that avoids injecting additional modelling choices.</p> <p>By aggregating over all possible point totals, we obtain the probability that a given team finishes in position 1, 2, 3, …, N.</p> <p>That’s what fills the table.</p> <p>For example:</p> <div class="row mt-3"> <div class="col-sm mt-3 mt-md-0"> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/position-odds-2025-2026-EFL_Championship-480.webp 480w,/assets/img/position-odds-2025-2026-EFL_Championship-800.webp 800w,/assets/img/position-odds-2025-2026-EFL_Championship-1400.webp 1400w," sizes="95vw" type="image/webp"/> <img src="/assets/img/position-odds-2025-2026-EFL_Championship.png" class="img-fluid rounded z-depth-1" width="100%" height="auto" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> </figure> </div> </div> <h2 id="why-this-isnt-a-simulation">Why this isn’t a simulation</h2> <p>Many league predictors simulate the rest of the season thousands or millions of times (e.g. <a href="https://theanalyst.com/articles/opta-football-predictions">Opta</a>). Those approaches can be flexible and intuitive. However this approach is different.</p> <p>Here every probability is computed exactly so the results are fully reproducible and even small probabilities (like a 0.3% chance of winning the league) are real, not artefacts of random noise. However the cost of this precision is a stronger set of assumptions, here we assume:</p> <ul> <li>match outcomes are independent,</li> <li>team strengths do not change,</li> <li>tie-breaks are random.</li> </ul> <h2 id="reading-the-table">Reading the table</h2> <p>Each row shows a team; each column shows a finishing position.</p> <ul> <li>A cell marked “12%” means exactly that: a 12% chance of finishing in that position.</li> <li>“&lt;1%” means possible, but very unlikely.</li> <li>“–” means impossible — the team cannot mathematically finish there.</li> </ul> <p>The strongest probability in each row is highlighted, and other positions are shaded relative to it. This gives a quick visual sense of where a team’s likely finishing range lies.</p> <p>This can be hard to see on a mobile scream, so the full table is replaced by a compact summary:</p> <ul> <li>chance of finishing 1st,</li> <li>chance of finishing Top 6,</li> <li>chance of finishing in the Bottom 3.</li> </ul> <p>The full table is always available via a toggle, and the entire table can be downloaded as an image for closer inspection.</p> <h2 id="what-this-doesnt-say">What this doesn’t say</h2> <p>These tables do not say:</p> <ul> <li>who will win the league,</li> <li>that form and injuries don’t matter.</li> </ul> <p>They say something subtler, and, I think, more interesting:</p> <p>“If the football world stays roughly as it is today, how wide is the space of possible futures?”</p> <p>Sometimes that space is narrow. Sometimes it can be surprisingly large.</p> <p>Check out the <a href="https://seanelvidge.com/tableProbs">predicted league tables</a> for yourself. Updated everytime the <a href="https://seanelvidge.com/articles/2024/All_England_football_league_results/">football database</a> is updated.</p>]]></content><author><name></name></author><category term="football"/><category term="mathematics"/><summary type="html"><![CDATA[The methodology behind how I calculate the probabilities of where the different teams in the English football league will end up.]]></summary></entry><entry><title type="html">Tracking Football Team Strengths with a Bayesian Kalman Model</title><link href="https://seanelvidge.github.io/articles/2025/Football_team_rankings/" rel="alternate" type="text/html" title="Tracking Football Team Strengths with a Bayesian Kalman Model"/><published>2025-12-15T09:20:00+00:00</published><updated>2025-12-15T09:20:00+00:00</updated><id>https://seanelvidge.github.io/articles/2025/Football_team_rankings</id><content type="html" xml:base="https://seanelvidge.github.io/articles/2025/Football_team_rankings/"><![CDATA[<p>Not all football-rating systems are the same. Many of the public ones, like the excellent <a href="http://clubelo.com/System">ClubElo</a>, do a fine job of ranking teams using elegant updates. Win and your rating rises, lose and it falls; the amount you move depends on how “surprised” the model was.</p> <p>This model begins in the same spirit but replaces those heuristic updates with a probabilistic engine: a <em>Bayesian Extended Kalman Filter</em> coupled to a modern version of the Bradley–Terry model (with added draws). The core idea is simple: treat each team’s strength as something hidden that we estimate and track through time, with uncertainty that expands between matches and contracts when new results arrive.</p> <p>Across the <a href="https://seanelvidge.com/articles/2024/All_England_football_league_results/">full historical dataset</a> (every English league match since 1888) the system achieves a mean Brier score of 0.2035. These values are significantly better than many published computer models (e.g. <a href="https://www.stat.cmu.edu/cmsac/sure/2023/showcase/soccer/report.html">Nyamdorj et al. 2014</a>, <a href="https://bsic.it/odds-at-play-testing-efficiency-in-the-premier-league-and-serie-a/">BSIC, 2024</a> and <a href="https://harvardsportsanalysis.org/2015/07/5988/">Harvard Sports Analysis Collective, 2015</a>) which indicates a stable, well-calibrated predictive performance. Crucially, because the filter quantifies its own uncertainty, it can tell us not only who is strongest, but how confident we should be in that judgement.</p> <p>The rest of this post goes into the mathematical details of the ranking algorithm, but if you want to access the underlying data it is <a href="https://github.com/seanelvidge/England-football-results">available here</a> (specifically the file <a href="https://raw.githubusercontent.com/seanelvidge/England-football-results/refs/heads/main/EnglandLeagueResults_wRanks.csv"><code class="language-plaintext highlighter-rouge">EnglandLeagueResults_wRanks.csv</code></a>).</p> <h2 id="the-big-picture">The Big Picture</h2> <p>Imagine every team has a hidden “true strength” \(s_i\). Before a match, the home and away teams carry beliefs about their current strengths, each with an associated uncertainty. When they play, the result provides new information that updates those beliefs.</p> <p>In the ClubElo framework this update is written directly as</p> \[R_{\text{new}} = R_{\text{old}} + K(S - E),\] <p>where \(S\) is the score (1 for a win, 0.5 for a draw, 0 for a loss), \(E\) is the expected probability, and \(K\) is a fixed responsiveness parameter.</p> <p>Here, the same logic is embedded in a <a href="https://en.wikipedia.org/wiki/Kalman_filter">Kalman filter</a>, which means the effective \(K\) is learned automatically. The update size depends on two things: how uncertain we are about the teams, and how surprising the result was. Upsets between uncertain teams lead to large updates; shocks between well-understood teams barely move the needle.</p> <p>Between matches, team strengths are not frozen. Instead, they evolve according to a <em>mean-reverting stochastic process</em>, allowing form to drift while preventing runaway behaviour. Newly promoted teams begin with large uncertainty and adapt quickly; established teams gravitate toward long-term baselines rather than rising indefinitely.</p> <h2 id="modelling-match-outcomes-wins-draws-losses">Modelling Match Outcomes (Wins, Draws, Losses)</h2> <p>Match outcomes are modelled using the <a href="https://www.jstor.org/stable/2283595">Davidson extension</a> of the <a href="https://en.wikipedia.org/wiki/Bradley%E2%80%93Terry_model">Bradley–Terry model</a>, which naturally incorporates draws. The probability of each outcome is</p> \[\Pr(\text{Home}) = \frac{e^{\Delta}}{e^{\Delta} + e^{-\Delta} + \kappa},\] \[\Pr(\text{Draw}) = \frac{\kappa}{e^{\Delta} + e^{-\Delta} + \kappa},\] \[\Pr(\text{Away}) = \frac{e^{-\Delta}}{e^{\Delta} + e^{-\Delta} + \kappa},\] <p>where</p> \[\Delta = \beta \bigl(s_H - s_A + h\bigr).\] <p>Here, \(s_H\) and \(s_A\) are the latent strengths of the home and away teams, \(\beta\) is a scaling parameter, \(\kappa\) controls the draw rate, and \(h\) is the home-advantage term.</p> <h2 id="home-advantage---explicit-and-time-varying">Home Advantage - Explicit and Time-Varying</h2> <p>Home advantage is not treated as a fixed constant. Instead, it is explicitly modelled and allowed to vary over time. Historical analysis shows that home-win rates in English football have declined dramatically since the late 19th century, falling from well above 60% to closer to 40–45% in the modern era (see <a href="https://seanelvidge.com/articles/2025/Home_advantage_in_English_football/">this analysis</a>).</p> <p>By allowing the home-advantage parameter \(h\) to evolve slowly with time, the model correctly distinguishes between a home match in 1890 and one in 2025. This prevents systematic bias when comparing teams across eras and is a key improvement over static-advantage rating systems.</p> <h2 id="validation-and-rating-scale">Validation and Rating Scale</h2> <p>Predictive accuracy is measured using the <a href="https://en.wikipedia.org/wiki/Brier_score">Brier score</a>, defined as the mean squared error between predicted probabilities and observed outcomes. Over the full dataset the score is 0.2035 (for the 2024/25 season it is 0.2085), indicating robust calibration both historically and in the present day.</p> <p>Internally, the filter operates on a latent “skill” scale roughly spanning \(-3\) to \(+3\). However for presentation, these values are mapped linearly onto an Elo-style scale, centred on 1000 points, with elite teams reaching 1800+ and lower-league teams clustering about a thousand points below. This transformation is purely cosmetic; all inference happens on the latent scale.</p> <h1 id="part-ii--if-you-dare-read-on-the-mathematics">Part II — If You Dare Read On: The Mathematics</h1> <h2 id="1-state-evolution-ornsteinuhlenbeck-dynamics">1. State Evolution (<a href="https://en.wikipedia.org/wiki/Ornstein%E2%80%93Uhlenbeck_process">Ornstein–Uhlenbeck</a> Dynamics)</h2> <p>Each team’s latent strength evolves as</p> \[s_{i,t+1} = \rho\, s_{i,t} + (1 - \rho)\, \mu_{\text{tier}(i,t)} + \varepsilon_{i,t},\] <p>with</p> \[\varepsilon_{i,t} \sim \mathcal{N}(0, q_i).\] <p>Here, \(\rho\) controls persistence, \(\mu_{\text{tier}}\) is the long-term baseline for the team’s division tier, and \(q_i\) is the process variance. This formulation allows ratings to drift while remaining anchored to realistic league-level expectations.</p> <h2 id="2-observation-model">2. Observation Model</h2> <p>Given predicted strengths \(s_H\) and \(s_A\), the model produces a probability vector</p> \[\mathbf{p} = \begin{bmatrix} p_H \\ p_D \\ p_A \end{bmatrix},\] <p>using the Davidson–Bradley–Terry equations above. The Jacobian matrix \(H\) is computed by differentiating \(\mathbf{p}\) with respect to the state vector \([s_H, s_A]^T\), enabling linearisation of the nonlinear observation model.</p> <h2 id="3-extended-kalman-filter-update">3. Extended Kalman Filter Update</h2> <p>Prediction step:</p> \[x^- = F x_t,\] \[P^- = F P_t F^\top + Q,\] <p>where \(F = \rho I\) and \(Q\) is the process-noise covariance.</p> <p>Observation noise:</p> \[R = \operatorname{diag}(\mathbf{p}) - \mathbf{p}\mathbf{p}^\top.\] <p>Update step:</p> \[K = P^- H^\top \bigl(H P^- H^\top + R\bigr)^{-1},\] \[x_{t+1} = x^- + K (y - \mathbf{p}),\] \[P_{t+1} = (I - K H) P^-.\] <p>The Kalman gain \(K\) replaces the fixed \(K\)-factor of traditional Elo systems, adapting automatically to uncertainty and surprise. A secondary control loop monitors the normalised innovation squared to ensure statistical consistency over time.</p> <h2 id="why-this-matters">Why This Matters</h2> <p>This framework enforces mathematical honesty. Uncertainty is explicit, calibration is measurable, and every parameter (\(\beta\), \(\rho\), \(q\), \(\kappa\), tier baselines, and historical home advantage) has a clear interpretation. Instead of a single magic constant, the model becomes a living system that adapts across seasons, divisions, and eras.</p>]]></content><author><name></name></author><category term="football"/><category term="mathematics"/><summary type="html"><![CDATA[The methodology behind my football team strength model.]]></summary></entry><entry><title type="html">Football Bogey Grounds and How Statistics Can Prove Them</title><link href="https://seanelvidge.github.io/articles/2025/Football_bogey_grounds/" rel="alternate" type="text/html" title="Football Bogey Grounds and How Statistics Can Prove Them"/><published>2025-12-12T19:04:00+00:00</published><updated>2025-12-12T19:04:00+00:00</updated><id>https://seanelvidge.github.io/articles/2025/Football_bogey_grounds</id><content type="html" xml:base="https://seanelvidge.github.io/articles/2025/Football_bogey_grounds/"><![CDATA[<p>Football supporters are never short of folklore. Some of it heroic, some of it tragic, and some of it mathematically suspicious. Among the more enduring tales is the <em>bogey ground</em>: that one venue where your club never quite manages to win, no matter how many times the fixture computer sends you there with fresh optimism, a new manager and a good run of form.</p> <p>But one question always pops up in the back of my mind when I hear this: When is a bogey ground actually a bogey ground, and when is it just a trick of small numbers?</p> <p>Most fans instinctively understand the point. Losing your only ever visit to Carlisle does not make Brunton Park a cursed ground. A couple of failed trips to Luton are not evidence of eldritch forces at work. Yet, scattered across the long history of English league football, a few pairs of teams have met often enough, and still produced zero away wins, that bogey grounds become statistically real.</p> <p>In the Premier League an example of this is Fulham, and their trips to Arsenal.</p> <p>Fulham have played 32 league matches away at Arsenal. They have never won.</p> <p>Not once. Not in 1914 (lost 2-0). Not in 1964 (2-2). Not in 2014 (lost 2-0).</p> <p>Zero wins from thirty-two attempts. Seven draws. Twenty-five defeats.</p> <p>Even more striking, this isn’t just a bad example, this is, among every club in the entire Football League dataset, over 135 years of football, no team has a larger, more emphatic, more mathematically convincing record of away futility than Fulham at Arsenal.</p> <p>So let’s take a look, not just at the record itself, but at how we can quantify bogeyness in a way that recognises your intuition (“two games aren’t enough!”) but takes advantage of the huge dataset football provides.</p> <h2 id="why-never-won-isnt-enough-the-curse-of-small-n">Why “never won” isn’t enough (the curse of small \(n\))</h2> <p>Imagine you flip a coin twice and get two tails. Does that mean the coin is biased?</p> <p>Of course not. You need <em>more</em> flips before you start thinking that someone is up.</p> <p>The same goes for football. A team losing its only away visit to Manchester City tells you nothing. Losing twice? Still nothing except a faint sense of déjà vu. Losing five times? Now you’re paying attention. Losing ten times? Stop buying tickets to the away game.</p> <p>So what we need is a way of answering this question:</p> <p>Given a team has played \(n\) away matches at a ground and won none, what is a reasonable upper limit on how good their <em>true</em> chance of winning there might be?</p> <p>Fortunately statistics, as it always does, gives us a handy tool for working this out, the <a href="https://en.wikipedia.org/wiki/Binomial_proportion_confidence_interval#Wilson_score_interval">Wilson confidence interval.</a></p> <h2 id="a-gentle-introduction-to-the-wilson-interval">A (gentle) introduction to the Wilson interval</h2> <p>Don’t panic; no equations are necessary. Just an idea. A Wilson interval takes two bits of information:</p> <ul> <li>the number of games played (\(n\)), and</li> <li>the number of games won (0, in our bogey-ground case)</li> </ul> <p>and asks:</p> <p>“If their true underlying chance of winning were \(p\), how big could \(p\) realistically be before the observed data (zero wins) would start to look implausible?”</p> <p>For small \(n\) (not many games played) the answer is: “\(p\) could still be quite large.”</p> <p>For large \(n\) (a large number of games played) the answer is: “\(p\) must be quite small.”</p> <p>This lines up with our intuition.</p> <p>A team with 0 wins from 2 tries? Their true win rate might easily be 30%, 40%, even 50%.</p> <p>A team with 0 wins from 32 tries? Now the interval collapses. It tells you the true away win probability is almost certainly very small, low enough that even by bad luck alone, you’d be surprised <em>not</em> to have picked up a single win after that many attempts.</p> <p>The Wilson interval is good for this particular problem (compared to other approaches like the Wald Interval) because it behaves properly at the boundaries (like at 0% or 100%), where ordinary methods can breakdown. Football fans should love it because it gives scientific legitimacy to something they’ve always known: <strong>some bogey grounds are imaginary, but a few are real</strong>.</p> <h2 id="the-most-convincing-bogey-of-them-all-fulham-at-arsenal">The most convincing bogey of them all: Fulham at Arsenal</h2> <p>So, what does Wilson say about Fulham’s 0 wins from 32 away league games at Arsenal?</p> <p>It says this:</p> <blockquote> <p>In each of Fulham’s away matches against Arsenal their probability of winning must have been (at the very most) 10% for 0 wins out of 32 to not be statistically suspicious.</p> </blockquote> <p>Mathematically this comes from plugging \(n\) (number of games) and \(z\) (which is equal to 1.96 for 95% ‘confidence’) into the equation for calculating the upper bound of the Wilson interval:</p> \[\mbox{Wilson Upper Bound} = \frac{z^2}{n+z^2} = \frac{1.96^2}{32+1.96^2} = 0.107 = 10.7\%\] <p>So the only way we could avoid saying Arsenal is Fulham’s away team nemesis is if we believe Fulham, across the whole history of the fixture (starting in 1913; <a href="https://seanelvidge.com/h2h?team1=Arsenal&amp;team2=Fulham">see my head-to-head tool here</a>), have not had better than a 10% chance of winning those games. Whilst we can’t be (mathematically) certain that it is true, it is incredibly unlikely. Typical away win rates in the top division over history are roughly 25-30% range. Even clear cut underdogs often have pre-match win probabilities around 15-20%.</p> <p>So, therefore, I think it is safe to say to conclude:</p> <blockquote> <p>Fulham away at Arsenal is not just a bogey ground, it is the bogey ground.</p> </blockquote> <p>The Premier League’s most mathematically defensible curse.</p> <h2 id="other-bogey-teams">Other bogey teams</h2> <p>Fulham–Arsenal is the only pair in the <a href="https://seanelvidge.com/articles/2024/All_England_football_league_results/">entire database</a> whose Wilson upper bound hovers at just above the 10% threshold. But several others get close, here are the best (or worst) of the rest:</p> <table style="border-collapse: collapse; width: 75%;"> <thead> <tr> <th style="border: 1px solid black; padding: 8px;">Away Team</th> <th style="border: 1px solid black; padding: 8px;">Home Team</th> <th style="border: 1px solid black; padding: 8px;">Played</th> <th style="border: 1px solid black; padding: 8px;">Record</th> <th style="border: 1px solid black; padding: 8px;">Wilson Upper</th> </tr> </thead> <tbody> <tr> <td style="border: 1px solid black; padding: 8px;">Fulham</td> <td style="border: 1px solid black; padding: 8px;">Arsenal</td> <td style="border: 1px solid black; padding: 8px;">32</td> <td style="border: 1px solid black; padding: 8px;">0W–7D–25L</td> <td style="border: 1px solid black; padding: 8px;">10.7%</td> </tr> <tr> <td style="border: 1px solid black; padding: 8px;">Grimsby Town</td> <td style="border: 1px solid black; padding: 8px;">Blackburn Rovers</td> <td style="border: 1px solid black; padding: 8px;">28</td> <td style="border: 1px solid black; padding: 8px;">0W–9D–19L</td> <td style="border: 1px solid black; padding: 8px;">12.1%</td> </tr> <tr> <td style="border: 1px solid black; padding: 8px;">Mansfield Town</td> <td style="border: 1px solid black; padding: 8px;">Reading</td> <td style="border: 1px solid black; padding: 8px;">26</td> <td style="border: 1px solid black; padding: 8px;">0W–5D–21L</td> <td style="border: 1px solid black; padding: 8px;">12.9%</td> </tr> <tr> <td style="border: 1px solid black; padding: 8px;">Tranmere Rovers</td> <td style="border: 1px solid black; padding: 8px;">Barnsley</td> <td style="border: 1px solid black; padding: 8px;">26</td> <td style="border: 1px solid black; padding: 8px;">0W–11D–15L</td> <td style="border: 1px solid black; padding: 8px;">12.9%</td> </tr> <tr> <td style="border: 1px solid black; padding: 8px;">Oldham Athletic</td> <td style="border: 1px solid black; padding: 8px;">Charlton Athletic</td> <td style="border: 1px solid black; padding: 8px;">25</td> <td style="border: 1px solid black; padding: 8px;">0W–11D–14L</td> <td style="border: 1px solid black; padding: 8px;">13.3%</td> </tr> <tr> <td style="border: 1px solid black; padding: 8px;">Newport County</td> <td style="border: 1px solid black; padding: 8px;">Luton Town</td> <td style="border: 1px solid black; padding: 8px;">24</td> <td style="border: 1px solid black; padding: 8px;">0W–7D–17L</td> <td style="border: 1px solid black; padding: 8px;">13.8%</td> </tr> <tr> <td style="border: 1px solid black; padding: 8px;">Coventry City</td> <td style="border: 1px solid black; padding: 8px;">Preston North End</td> <td style="border: 1px solid black; padding: 8px;">24</td> <td style="border: 1px solid black; padding: 8px;">0W–9D–15L</td> <td style="border: 1px solid black; padding: 8px;">13.8%</td> </tr> </tbody> </table> <p><br/></p> <p>These are not small samples. These are not casual coincidences. When you are approaching 20 or 30 visits with no victory, it is time to stop thinking you are unlucky and start accepting you have a bogey ground.</p>]]></content><author><name></name></author><category term="football"/><category term="mathematics"/><summary type="html"><![CDATA[Can a football club actually have a statistically verifiable football bogey ground or is it just bad luck?]]></summary></entry><entry><title type="html">Why Our Next Big Solar Storm Warning Could Still Be a Guess</title><link href="https://seanelvidge.github.io/articles/2025/Guessing_at_Space_Weather/" rel="alternate" type="text/html" title="Why Our Next Big Solar Storm Warning Could Still Be a Guess"/><published>2025-11-30T19:30:00+00:00</published><updated>2025-11-30T19:30:00+00:00</updated><id>https://seanelvidge.github.io/articles/2025/Guessing_at_Space_Weather</id><content type="html" xml:base="https://seanelvidge.github.io/articles/2025/Guessing_at_Space_Weather/"><![CDATA[<p>Eighteen months ago, a beautiful auroral display swept across the world, accompanying the first extreme geomagnetic storm in more than two decades. It was a landmark event not just for the spectacle but for what it revealed about our growing global dependence on reliable space-weather forecasting. When another storm drew headlines this month (November 2025), with an early “G5 watch” issued by forecasters, governments and operators activated contingency plans in anticipation of another major impact (the geomagnetic storm scale, G-scale, ranges from one to five, with five denoting the most extreme events). Yet the storm ultimately peaked at only G3, a level we expect hundreds of times in a solar cycle. That mismatch exposed a persistent weakness in how we observe and interpret the Sun.</p> <p>The November events began with three coronal mass ejections (CMEs) erupting from the same active region over consecutive days. Using direct, head-on observations from the Sun–Earth line, scientists estimated their speeds and trajectories and issued a G4 storm watch for 12 November. When the first CME struck earlier than expected and with stronger parameters than the initial estimate, the risk escalated: successive CMEs often travel faster along the newly cleared path, and the potential for a combined impact grows. A G5 watch was released, triggering international preparedness actions. Power-grid operators revisited protection settings, communications providers reviewed continuity plans, and governments braced for potential disruption.</p> <p>Then nothing happened, at least not at the predicted intensity or time. The CME did arrive, but just before midnight, and with a magnetic field that was only briefly and weakly aligned in the southward direction needed to strongly couple with Earth’s magnetic system. Without that sustained southward field, even an otherwise fast and dense CME cannot deliver a major geomagnetic storm. The result was a modest G3. This forecasting gap did not arise from poor judgement but from a fundamental observational void: after we watch a CME leave the Sun, we lose sight of its evolution for roughly 99% of the journey to Earth.</p> <p>The limitations of this “single viewpoint” forecasting have been known for years. We can estimate initial CME speed, but we cannot track how it deforms, interacts with the solar wind or merges with other CMEs en route. Crucially, we cannot determine its magnetic structure, the key factor in storm severity, until the CME reaches the L1 point (Lagrange point 1), only about 30 minutes before Earth impact. Even this narrow window is presently fragile: the DSCOVR spacecraft, one of only two satellites providing L1 measurements, is offline, and ACE is approaching 30 years in space with ageing instruments.</p> <p>Given the rising global reliance on vulnerable infrastructure: power networks, navigation systems, satellite constellations, living with this degree of uncertainty is no longer tenable. Forecasts that overshoot create costly false alarms; forecasts that undershoot risk real damage. The November mis-forecast was harmless, but it underscored how close we are to the limits of what current assets can deliver.</p> <p>New capabilities are on the horizon, though they require sustained commitment. The NOAA–NASA SWFO-L1 mission is expected to restore robust solar-wind monitoring from mid-2026, stabilising the 30-minute warning baseline. Far more transformative would be the European Space Agency’s Vigil mission, planned for launch to the L5 Lagrange point in 2031. From that vantage point, Vigil would view CMEs from the side, enabling reliable speed and trajectory estimates and dramatically reducing arrival-time uncertainty. But its path to launch depends on political support through multiple rounds of ministerial scrutiny.</p> <p>Beyond L5, deeper innovation is needed. A constellation of sensors placed in distant retrograde orbits (DRO) could offer early in-situ magnetic-field measurements, potentially hours, not minutes, before a CME reaches Earth. ESA’s planned demonstration of a DRO spacecraft in 2026, HENON, is a first step, but only that. These observational advances must be matched by improved modelling of CME structure, solar precursors and the complex interactions that determine whether an incoming CME will actually be geoeffective.</p> <p>Accurate space-weather forecasting will always involve uncertainty, but it need not remain dominated by blind spots. The November storm was a reminder that our current system can predict the possibility of major events yet cannot confirm their true character until it is almost too late. If we want forecasts that are genuinely actionable, capable of guiding grid operators, satellite controllers and governments with confidence, then investment in new vantage points, new missions and new science is essential. Without it, our next “G5 watch” may leave us waiting once again, unsure whether a major storm is coming or whether the sky will stay quiet.</p>]]></content><author><name></name></author><category term="spaceWeather"/><summary type="html"><![CDATA[Accurate space-weather forecasting will always involve uncertainty, but it need not remain dominated by blind spots. The November storm was a reminder that our current system can predict the possibility of major events yet cannot confirm their true character until it is almost too late.]]></summary></entry><entry><title type="html">40 points to avoid relegation?</title><link href="https://seanelvidge.github.io/articles/2025/40_points_to_avoid_relegation/" rel="alternate" type="text/html" title="40 points to avoid relegation?"/><published>2025-11-08T23:04:00+00:00</published><updated>2025-11-08T23:04:00+00:00</updated><id>https://seanelvidge.github.io/articles/2025/40_points_to_avoid_relegation</id><content type="html" xml:base="https://seanelvidge.github.io/articles/2025/40_points_to_avoid_relegation/"><![CDATA[<html> <head> <style>.chart-figure{width:100%;margin:1rem 0}.chart-container{position:relative;width:100%;height:55vh;min-height:420px;max-height:80vh}.chart-container canvas{width:100%!important;height:100%!important;display:block}@media(max-width:640px){.chart-container{height:65vh;min-height:480px}}</style> </head> </html> <p>The number has become part of Premier League folk law. Forty points. Reach 40 and you can relax, the trapdoor to the Championship won’t open.</p> <p>At the start of the 2015/16 season (the season Leicester City won the League!) their manager Claudio Ranieri set them the target of reaching 40 points to avoid relegation. When they hit the target:</p> <blockquote> <p>“We have 40 points which was the goal. It’s champagne for my players!”</p> </blockquote> <blockquote> <p>Claudio Ranieri, 2 Jan, 2016.</p> </blockquote> <p>But where did this number come from, and it is the right target?</p> <p>Many people have noted that the 40 point mark is a “myth” (e.g. <a href="https://www.premierleague.com/en/news/3932287">The Premier League</a>, <a href="https://www.bbc.co.uk/sport/football/43049564">BBC</a>, <a href="https://www.nytimes.com/athletic/6126560/2025/02/12/leicester-city-fixtures-premier-league-relegation/">The Athletic</a>). But those articles are really just saying that, on average, you need less than 40 points to survive (and the number of points you need seems to be decreasing).</p> <figure class="chart-figure"> <div class="chart-container"> <canvas id="pointsChart"></canvas> </div> <figcaption></figcaption> </figure> <p>The plot above shows the number of points required to stay in the top division of English football since 3 points for a win was introduced in 1985 (this includes before the Premier League started in 1992). The league tables were created using my <a href="https://seanelvidge.com/leaguetable">arbitrary league table generator</a>. It is also worth noting that during the time range there have been a varying number of teams in the top division (between 20 and 22) which obviously impacts the points required. We have normalized this to a 38-game season for comparison.</p> <p>A few things are immediatly obvious:</p> <ol> <li>Most of the time you do not need 40 points to avoid relegation (34/44; 77%),</li> <li>The average number of points needed to avoid relegation over the last 40 years is 37 (in the Premier League era, it is 36 points),</li> <li>There is a clear decreasing trend of the number of points required.</li> </ol> <p>So where does the idea that you need 40 points come from? There is no obvious origin of the phrase. My best guess is that in the mid 90s, after Southampton survied (just, on goal difference) with 38 points in 1995/96, followed by Coventry City with 41 the following year, then Everton with 40 and then Southampton (again) with 41 points in 1998/99, that that run of four was enough for the (rough) number to stick.</p> <p>But despite the previous analysis (by me and many others) I think 40 is still the correct, pre-season, target. Moving to (for example) a 37 point target would only give you slightly better than a 50-50 chance of staying up (54.5%). Perhaps you would like a little more certainty…</p> <p>One way to look at this is to calculate the cumulative distribution function (CDF). The CDF is a function that shows the probability that a random variable is less than or equal to a specific value. It “accumulates” or “adds up” the probabilities for all outcomes up to a certain point.</p> <figure class="chart-figure"> <div class="chart-container"> <canvas id="cdfChart"></canvas> </div> <figcaption></figcaption> </figure> <p>The above figure shows two, slightly different, CDFs, in blue across the whole dataset (since 1985) and in red just in the Premier League era. To read the plot look at the number of points on the x-axis and then read off the corresponding probability of avoiding relegation on the y-axis for that number of points.</p> <p>Now the 40 point value makes a lot more sense, as it gives a team over an 80% chance of staying in the division, now I prefer those odds!</p> <p>So, whilst there are plenty of posts online telling you that the 40 point target is a myth, I think, as pre-season target for teams, it is still a good one.</p> <script src="https://cdn.jsdelivr.net/npm/chart.js@4"></script> <canvas id="pointsChart"></canvas> <script>function formatSeason(t){const e=parseInt(t,10);return`${e-1}/${String(e).slice(-2).padStart(2,"0")}`}function linearFit(t,e){const a=t.length,r=t.reduce((t,e)=>t+e,0),o=e.reduce((t,e)=>t+e,0),i=(a*t.reduce((t,a,r)=>t+a*e[r],0)-r*o)/(a*t.reduce((t,e)=>t+e*e,0)-r*r);return{m:i,c:(o-i*r)/a}}const chartData={datasets:[{label:"(Normalized) Points Needed to Avoid Relegation",data:[{x:"1982",y:38.9},{x:"1983",y:43.4},{x:"1984",y:44.3},{x:"1985",y:45.2},{x:"1986",y:38},{x:"1987",y:38.9},{x:"1988",y:34.2},{x:"1989",y:40},{x:"1990",y:44},{x:"1991",y:38},{x:"1992",y:38.9},{x:"1993",y:45.2},{x:"1994",y:38.9},{x:"1995",y:39.8},{x:"1996",y:39},{x:"1997",y:41},{x:"1998",y:41},{x:"1999",y:37},{x:"2000",y:34},{x:"2001",y:35},{x:"2002",y:37},{x:"2003",y:43},{x:"2004",y:34},{x:"2005",y:34},{x:"2006",y:35},{x:"2007",y:39},{x:"2008",y:37},{x:"2009",y:35},{x:"2010",y:31},{x:"2011",y:40},{x:"2012",y:37},{x:"2013",y:37},{x:"2014",y:34},{x:"2015",y:36},{x:"2016",y:38},{x:"2017",y:35},{x:"2018",y:34},{x:"2019",y:35},{x:"2020",y:35},{x:"2021",y:29},{x:"2022",y:36},{x:"2023",y:35},{x:"2024",y:27},{x:"2025",y:26}],errorBars:{"(Normalized) Points Needed to Avoid Relegation":[]}}]},years=chartData.datasets[0].data.map(t=>t.x),yPoints=chartData.datasets[0].data.map(t=>t.y),xNums=years.map(t=>parseInt(t,10)),{m:m,c:c}=linearFit(xNums,yPoints),fitY=xNums.map(t=>m*t+c),primaryDataset={label:"(Normalized) Points Needed to Avoid Relegation",data:yPoints,borderColor:"blue",borderWidth:2,pointRadius:2,tension:0},fortyLine={label:"40-point benchmark",data:years.map(()=>40),borderColor:"red",borderWidth:2,pointRadius:0,tension:0,fill:!1,order:0},fitLine={label:"Linear fit",data:fitY,borderColor:"grey",borderWidth:2,borderDash:[6,4],pointRadius:0,tension:0,fill:!1,order:2},vertical1993={label:"Start of Premier League",data:[{x:"1993",y:25},{x:"1993",y:50}],borderColor:"purple",borderWidth:.5,pointRadius:0,tension:0,fill:!1,borderDash:[10,10],order:3},ctx=document.getElementById("pointsChart").getContext("2d");new Chart(ctx,{type:"line",data:{labels:years,datasets:[fortyLine,primaryDataset,fitLine,vertical1993]},options:{responsive:!0,maintainAspectRatio:!1,plugins:{legend:{position:"top"},tooltip:{callbacks:{title:t=>formatSeason(t[0].label)}}},scales:{x:{type:"category",title:{display:!0,text:"Season"},ticks:{callback:function(t){const e=this.getLabelForValue(t);return parseInt(e,10)%5==0?formatSeason(e):""},maxRotation:0,autoSkip:!1}},y:{title:{display:!0,text:"Points"},beginAtZero:!1}},interaction:{mode:"nearest",intersect:!1}}});</script> <script>const cdfPercents=[0,0,0,0,0,0,0,0,0,0,0,2.27,4.55,4.55,6.82,6.82,9.09,9.09,9.09,20.45,38.64,43.18,54.55,61.36,75,81.82,86.36,86.36,88.64,93.18,95.45,100,100,100,100,100,100,100,100,100,100],cdfPremier=[0,0,0,0,0,0,0,0,0,0,0,3.03,6.06,6.06,9.09,9.09,12.12,12.12,12.12,27.27,48.48,54.55,69.7,72.73,81.82,87.88,93.94,93.94,96.97,96.97,96.97,100,100,100,100,100,100,100,100,100,100],cdfData1=cdfPercents.map((t,e)=>({x:e+15,y:t})),cdfData2=cdfPremier.map((t,e)=>({x:e+15,y:t})),cdfCtx=document.getElementById("cdfChart").getContext("2d");new Chart(cdfCtx,{type:"line",data:{datasets:[{label:"Historical CDF",data:cdfData1,stepped:!0,borderWidth:2,pointRadius:0},{label:"Premier League Era CDF",data:cdfData2,stepped:!0,borderWidth:2,borderDash:[6,4],pointRadius:0}]},options:{parsing:!1,maintainAspectRatio:!1,scales:{x:{type:"linear",min:15,max:55,ticks:{stepSize:5},title:{display:!0,text:"Points"}},y:{min:0,max:100,ticks:{callback:t=>t+"%"},title:{display:!0,text:"Cumulative probability (%)"}}},plugins:{legend:{display:!0},tooltip:{callbacks:{label:t=>`P(X \u2264 ${t.parsed.x}) = ${t.parsed.y.toFixed(2)}%`}}},elements:{line:{tension:0}}}});</script>]]></content><author><name></name></author><category term="football"/><category term="mathematics"/><summary type="html"><![CDATA[40 points is the classic benchmark to avoid relegation from the Premier League, is this the right value?]]></summary></entry><entry><title type="html">New space weather modelling suite enables upper atmosphere forecasting</title><link href="https://seanelvidge.github.io/articles/2025/New_space_weather_modelling_suite/" rel="alternate" type="text/html" title="New space weather modelling suite enables upper atmosphere forecasting"/><published>2025-10-08T14:31:00+00:00</published><updated>2025-10-08T14:31:00+00:00</updated><id>https://seanelvidge.github.io/articles/2025/New_space_weather_modelling_suite</id><content type="html" xml:base="https://seanelvidge.github.io/articles/2025/New_space_weather_modelling_suite/"><![CDATA[<p>A pioneering new space weather forecasting modelling suite will enable operational modelling of the upper atmosphere at the Met Office for the first time in a major breakthrough for UK atmospheric science.</p> <p>The Advanced Ensemble Networked Assimilation System (AENeAS) is a new suite of space weather forecasting models available to the Met Office that focuses on how space weather can influence the thermosphere and ionosphere here on Earth.</p> <p>The suite, built at the University of Birmingham, and developed in collaboration with Lancaster University, the Universities of Leeds, Bath and Leicester and the British Antarctic Survey is now running on the Met Office’s new supercomputer.</p> <blockquote> <p>The deployment of this suite at the UK Met Office is the realization of a 10-year vision of SERENE, to build and deliver a state-of-the-art upper atmosphere modelling capability into operational use. The tools will be able to support a wide range of users and ultimately allow people to make informed decisions earlier, being proactive rather than reactive in their response to space weather.</p> </blockquote> <p><strong>Professor Sean Elvidge, Head of Space Environment and Radio Engineering (SERENE) - University of Birmingham</strong></p> <p>Complementing the Met Office’s existing space weather forecasting models, which include predicting the arrival of events from the surface of the Sun, this system introduces new forecasting capability for modelling impacts from space weather on satellites, aviation, communications and services which rely on GNSS.</p> <p>Professor Sean Elvidge, Head of Space Environment and Radio Engineering (SERENE) at the University of Birmingham, and the lead developer of the system said: “The deployment of this suite at the UK Met Office is the realization of a 10-year vision of SERENE, to build and deliver a state-of-the-art upper atmosphere modelling capability into operational use.</p> <div class="row mt-3"> <div class="col-sm mt-3 mt-md-0"> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/aeneas_day_group-480.webp 480w,/assets/img/aeneas_day_group-800.webp 800w,/assets/img/aeneas_day_group-1400.webp 1400w," sizes="95vw" type="image/webp"/> <img src="/assets/img/aeneas_day_group.jpg" class="img-fluid rounded z-depth-1" width="100%" height="auto" data-zoomable="" loading="eager" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> </figure> </div> </div> <p>“The tools will be able to support a wide range of users and ultimately allow people to make informed decisions earlier, being proactive rather than reactive in their response to space weather.”</p> <p>The new suite means that, for the first time, forecasters at the Met Office Space Weather Operations Centre (MOSWOC) will have access to forecast model output on the impacts of space weather on the ionosphere, as well as enhanced modelling of the thermosphere.</p> <blockquote> <p>Once again, cutting-edge British innovation is making a remarkable difference to our daily lives - this time from way up in the atmosphere. This is a really exciting example of how better understanding of what’s happening in space can protect the tech we all rely on, from GPS on our phones to keeping the power grid working.</p> </blockquote> <p><strong>UK Science Minister Lord Vallance</strong></p> <p>Met Office Space Weather Manager Simon Machin said: “This delivers a world-leading capability that provides greater confidence and forecasting skill than any models currently in operation anywhere else in the world.</p> <p>“This isn’t just about science - it’s about protecting the systems we rely on every day. From aircraft communications to GPS in your phone, space weather can affect us all.”</p> <p>Science Minister Lord Vallance said: “Once again, cutting-edge British innovation is making a remarkable difference to our daily lives - this time from way up in the atmosphere.</p> <p>“This is a really exciting example of how better understanding of what’s happening in space can protect the tech we all rely on, from GPS on our phones to keeping the power grid working.”</p> <p>The new modelling capability will be able to assimilate near real-time data about the current state of the ionosphere and thermosphere, combine with forecasts of solar activity from the Sun to produce accurate and actionable forecasts of the upper atmosphere. This will help enable service providers to take mitigating actions to prevent impacts from space weather where possible.</p> <p>Professor Farideh Honary from Lancaster University said: “We are happy to see our research being translated into a useful product to be used by industry. The research and modelling led by Lancaster is relevant to the aviation industry and in particular to flights using polar routes which are dependent on high frequency communications.</p> <p>“These flights have significantly increased since their initial opening in the 1990s due to their operational advantages such as reduced flight times and fuel consumption, which translate to cost savings and environmental benefits like lower carbon emissions.”</p> <p>More accurate and precise forecast information will help enable service providers to take mitigating actions to prevent impacts from space weather where possible. One example of this is with users of Global Navigation Satellite Systems (GNSS), such as GPS, for positioning or navigational purposes. If they understand that there is likely to be a loss of accuracy in GNSS, they can switch to using other systems.</p> <p>Together, the new modelling suite has been delivered as part of Space Weather Instrumentation, Measurement, Modelling and Risk (SWIMMR) programme, which was funded through the UKRI Strategic Priorities Fund and designed to enhance the UK’s capability for monitoring, modelling and forecasting space weather.</p> <p>Professor Ian McCrea, SWIMMR Programme Lead at STFC RAL Space, said: “These new models represent a significant step forward for the UK’s capacity to model, forecast, and understand key components of our upper atmosphere. By coupling advances in physical modelling with global scale observations, they will enable unparalleled awareness of Earth’s geospatial environment.</p> <p>“The models promise to address a wide range of use cases, ranging from radio communications to the prediction of satellite orbits, and we expect that they will be of great importance to a huge variety of stakeholders. They also provide a clear demonstration of how the collaboration between academics and end users, which SWIMMR has enabled, can benefit members of both communities.”</p>]]></content><author><name></name></author><category term="spaceWeather"/><summary type="html"><![CDATA[New suite of space weather forecasting models focuses on how space weather can influence the thermosphere and ionosphere here on Earth.]]></summary></entry></feed>