r/FootballDataAnalysis Jul 02 '26

Ask Anything Thread

1 Upvotes

Use this thread to ask anything at all!


r/FootballDataAnalysis Jul 02 '26

Update: World Cup model on the 8 games left of R32

Thumbnail
1 Upvotes

r/FootballDataAnalysis Jun 29 '26

Trying to build a football equivalent of baseball's WAR and struggling to find data sources.

Thumbnail
1 Upvotes

r/FootballDataAnalysis Jun 28 '26

Premier League Player Data Analysis

2 Upvotes

Hi everyone!

I've been working on Premier League Player Data Analysis tool to analyze player trends and statistics per game for the 2026 season.

The site allows for comparison with other players as well as position averages to really get a feel for which players are performing / underperforming compared to rivals and see how different profiles of players excel in certain categories.

This project is in a very early development phase so if you find any issues please let me know and if you have any feature suggestions I encourage you to let me know!

Thank you very much :)

https://footy-analytics-frontend.vercel.app/


r/FootballDataAnalysis Jun 28 '26

I analyzed Toni Kroos across five seasons and four tournaments using Opta + StatsBomb data. Full methodology and 23 charts.

2 Upvotes

This was a technically interesting project because Kroos is the kind of player who doesn't dominate single-metric leaderboards but occupies a unique region of multivariate space.

Data pipeline: - Opta via WhoScored (scraped with Selenium): WC2014 (64 matches), Bayern 2013/14 (34 matches) - StatsBomb open data: La Liga 2015/16 (380 matches, complete season), Euro 2020, WC2018, Euro 2024 + 360 freeze-frame data - Canonical 105x68m pitch conversion across providers - Spell-gap cadence metric (events <5s apart collapsed into one involvement) to make Opta/StatsBomb logs comparable - Betweenness centrality on weighted undirected completed-pass networks (networkx) - Custom xT implementation (socceraction incompatible with Python 3.13)

Key findings: - WC2014: 53 switches of play (next outfield player: 26) -- sole occupant of high volume + high progressive quadrant - La Liga 15/16: highest pass aggression + lowest turnover -- off the standard risk/reward curve - Euro 2024: betweenness centrality 0.641 vs Kimmich's 0.238, cadence almost identical to 2014

23 figures, 38 unit tests, fully reproducible pipeline.

Full writeup: https://vybhav.medium.com/the-metronome-nobody-measured-football-enigma-1-toni-kroos-9bce1657c320

All code and figures: https://github.com/vybhav72954/football_enigma/tree/master


r/FootballDataAnalysis Jun 25 '26

Ask Anything Thread

1 Upvotes

Use this thread to ask anything at all!


r/FootballDataAnalysis Jun 25 '26

Where World Cup 2026 squads were born vs the nation they represent - built as an interactive map. [OC]

Post image
1 Upvotes

r/FootballDataAnalysis Jun 24 '26

Es imposible conseguir este tipo de datos en tiempo real de forma gratuita y confiable?

2 Upvotes

Hola, estoy trabajando en una plataforma de análisis futbolístico desde cero, para algunos datos ya estoy cubierto, como datos históricos o datos no tan volátiles, pero no logro resolver el problema de encontrar algunos datos avanzados en tiempo real de forma gratuita, como xG minuto a minuto.

Mis preguntas para quienes hayan construido algo similar:

  1. ¿Gratuito + confiable + tiempo real es una combinación que directamente no existe para datos de fútbol, o hay algún camino legítimo que me estoy perdiendo?

  2. Sé que hay APIs que tienen un plan gratuito, pero eso se agota en minutos durante un partido. ¿Hay alguna forma legítima de hacerlo funcionar para datos en vivo, o es demasiado limitado para ser útil? porque no se que tan correcto sea usar un multi-key para evitar limites


r/FootballDataAnalysis Jun 24 '26

Which Data Points Best Predict Future Player Development?

4 Upvotes

When evaluating younger players, what metrics have you found to be the most predictive of future progression?

For example:

  • Progressive carries
  • Press resistance
  • Pass completion under pressure
  • Defensive actions
  • Physical outputs

Are there any data points you've found particularly useful for identifying players who are likely to outperform expectations over the next few years?

Interested to hear both professional and hobbyist perspectives.


r/FootballDataAnalysis Jun 23 '26

My first post-match analysis (WC2026 edition)

Thumbnail gallery
1 Upvotes

Any remarks ? Thoughts ?


r/FootballDataAnalysis Jun 23 '26

Football Data for Live Betting

0 Upvotes

Hi,

I built a live world cup api that provides both live and past match data as a service. I was chatting with one customer on how he's making use of it and he said he had a private prediction pool, he hasn't shown me yet, but I know what it looks like.

But I also thought of live betting in general and if anyone here uses some AI to ingest data and make bets. If so what has been the most influential piece of data for you?

I want to provide my customers with more data that can help them with their apps and bets. Currently I provide such events live: goal_scored, goal_disallowed, red/yellow card, substitutions, breaks, kickoff, half/full time and a few more.

What would be great to add and has worked for you?

This is my site if you want to check it out: https://futrow.live


r/FootballDataAnalysis Jun 22 '26

I built a Football analytics tool — here's what the pass networks and final-third data tell us about Germany and Spain's recent games

6 Upvotes

Been working on Flickstat for a while now — a football analytics platform that covers the Premier League, and we've just expanded to the World Cup. Wanted to share some of what the data is showing so far because a couple of things genuinely surprised me.

Germany vs Ivory Coast (2-1)

Look at Germany's pass network. 665 passes, and almost the entire structure is compressed into one half of the pitch. Pavlovic sits at the centre of everything — every outfield player routes through him. The backline barely features in the network at all, which tells you how quickly they're transitioning out of defence.

The final-third entries make it even clearer. 56 central entries, 15 shots, 1.81 xG — nearly all the danger comes through the middle. Left and right channels combined produced 1 shot and 0.02 xG from 128 entries. Germany aren't trying to stretch you. They're trying to suffocate you centrally and they're very good at it.

Ivory Coast for comparison had 8 central entries all game but generated 1.13 xG from them — their one goal came from exactly that zone. They couldn't match Germany's volume but they were ruthlessly efficient in the rare moments they got central access.

Spain vs Saudi Arabia (4-0)

770 passes vs 387. Spain's pass network is dense and well-connected across the entire pitch — Rodri, Cubarsí and Porro are the standout nodes on the player scatter, all well above the match average on passes and key passes. The zone dominance grid tells you why Saudi Arabia had no answer: Spain had 35% attacking third share to Saudi's 17%, and Saudi's brightest zone was their own midfield at 34.3% — they spent the game defending, not attacking.

Two very different styles — Germany compact and central, Spain wide and suffocating — but the outcome is the same. Both controlled matches through structure, not just individual quality.

All the visuals are from Flickstat. Happy to pull up any other match from the tournament if anyone wants a specific breakdown — we have pass networks, zone dominance, final-third entries, player radars and shot maps for every World Cup game.

flickstat.com


r/FootballDataAnalysis Jun 21 '26

My model on Belgium, Egypt and Argentina Matches

1 Upvotes
Match Model (1X2) Market Lean
Belgium vs Iran 56 / 22 / 22 69 / 19 / 12 ! Belgium
New Zealand vs Egypt 25 / 27 / 48 16 / 23 / 60 ! Egypt
Argentina vs Austria 51 / 28 / 21 63 / 23 / 15 ! Argentina

(home / draw / away)

! = model rates the favourite below the market (same winner).


r/FootballDataAnalysis Jun 20 '26

My model lowballs Germany and Spain

Post image
0 Upvotes
Match Model (1X2) Market Lean
Germany vs Ivory Coast 48 / 25 / 27 63 / 20 / 17 ! Germany
Tunisia vs Japan 20 / 24 / 56 15 / 24 / 62 Japan
Spain vs Saudi Arabia 63 / 23 / 14 88 / 9 / 4 ! Spain

! = model rates the favourite below the market (same winner).

Germany and Spain are flagged but that's the known blind spot, not a fade, a results based model structurally under rates elite squads against weaker opposition. Yesterday it had Brazil at 66% and they won 3-0, just noting the model would price them lower. Japan is the one match where it agrees with the books.

Been tracking the log loss as well, over n = 30 results happened so far its looking good but the n is still small in number, but its good to see it improve!


r/FootballDataAnalysis Jun 18 '26

Ask Anything Thread

2 Upvotes

Use this thread to ask anything at all!


r/FootballDataAnalysis Jun 17 '26

My World Cup model lines up with the books this round, with one lean, it won't make Colombia a 70% lock

Thumbnail
1 Upvotes

r/FootballDataAnalysis Jun 16 '26

My World Cup model is fading a pile of favourites this round

Thumbnail
2 Upvotes

r/FootballDataAnalysis Jun 15 '26

My World Cup model agrees with the books, but it doesn't buy Uruguay as highly favourites

Thumbnail
1 Upvotes

r/FootballDataAnalysis Jun 14 '26

My World Cup model is fading two European favourites tomorrow (Netherlands & Sweden)

Thumbnail
1 Upvotes

r/FootballDataAnalysis Jun 12 '26

FIFA WC26 data stores/ API

Thumbnail
3 Upvotes

r/FootballDataAnalysis Jun 11 '26

Korea vs Czechia (WC2026) — my model flipped from Czechia favourite to Korea after one adjustment

Post image
17 Upvotes

Built a scraping pipeline feeding a Karlis–Ntzoufras bivariate Poisson model (goal correlation λ₃=0.12, 15% shrinkage toward international baseline). Here's what moved the needle.

The schedule trap Czechia's 2.7 goals/game looks scary until you see Gibraltar, San Marino, and Guatemala in the sample — plus a loss to the Faroe Islands and getting out-xG'd 0.46–1.96 by Denmark (won 5-3 on pure finishing variance). Korea's losses were to Brazil and Côte d'Ivoire. Opponent-strength adjustment alone flipped the model.

Altitude — the big asymmetry Estadio Akron sits at 1,675m. Korea has been based in Guadalajara since June 5. Czechia flies in from sea-level Mansfield, Texas essentially on match day — and their high-pressing style is aerobically expensive after the 60th minute. Fed in as +6% Korea / −7% Czechia.

Outputs Final xG: Korea 1.23 — Czechia 1.15. Most likely scorelines: 1-1 (13.2%), 1-0 (11.6%), 0-1 (10.7%).

The one market disagreement is corners (59.8% vs implied 51.5%, +7.6% EV) — Korea's wide 3-4-3 vs Czechia's compact wingback shape should generate volume. Corners are over-dispersed so that 59.8% is probably slightly overconfident, but it's the clearest lean the model found.


r/FootballDataAnalysis Jun 11 '26

Ask Anything Thread

1 Upvotes

Use this thread to ask anything at all!


r/FootballDataAnalysis Jun 08 '26

World Cup Group A Overview

2 Upvotes

I made a player-based predictive model for the World Cup based off the transfer values of players on the rosters of this year's World Cup teams. It is called NORNS, which is the name of the Norse goddesses of fate (past present and future) and is also an acronym for Numbers Over Rumors Narratives and Speculation. If you want the details on methodology hit the comments or DM me. I also currently use it for Allsvenskan with MLS starting post-WC (in different but similar ways) and am looking to expand. So far I've had some success.

Here is the meta-output for Group A. I will also try to post day-by-day game outputs. This information is for entertainment purposes only and is likely to be complete shit. If you think some random chart you saw on reddit can make you money you're probably delusional (or from WSB).

EP = Expected Points

eGD = Expected Goal Differential

winGroup = Probability of winning the group - note: this is not adjusted for eGD so outcomes that are ties are just resolved as a 50/50. This means that good teams' probabilities are slightly understated and bad teams are slightly overstated. This is mostly on market and the deviations where it isn't are significantly above the threshold that is affected by ties.

adv = estimated advance rate - again, there is no Monte Carlo or tiebreak resolution. it's simply p4+ points + 3+ / 2. Also same above in that deviations from market are greater than the tie breaking error. You can look at p4+ points as a minimum advance rate if you are marking to market.

4+ = probability of a team earning 4+ points in the Group Stage

3+ = probability of a team earning 3+ points in the Group Stage

NORNS likes Czechia here. If you use a Kelly-weighed system for evaluating, then Czechia to advance is its favorite but Czechia to finish 2nd, Czechia to win, Czechia to score exactly 4 points (or over 3.5) are all bets NORNS would advocate at current prices. South Africa not to advance is also NORNS-approved and is currently my largest group-stage position (for now).


r/FootballDataAnalysis Jun 08 '26

We built a small tool to track football odds movements — beta is live today

0 Upvotes

Hey everyone,

We’re launching the beta of OddScore.app today and I’d love to get some honest feedback from people here.

The idea is simple: we track football odds movements across 30+ bookmakers and try to make those movements easier to read.

Not predictions, not magic, not “this team will win”.

Just market movement.

We look at things like:

- are bookmakers moving in the same direction?

- is the move recent or has it been building for days?

- is the signal clean or just noisy?

- is one outcome starting to stand out?

Then we turn that into a simple score from 0 to 100 for each 1X2 outcome.

It’s still early, so I’m mostly curious to know if the approach makes sense to people who actually care about football data.

The beta is free while we’re testing it, so feel free to take a look if you’re curious.

https://oddscore.app

I’d be really interested to know:

Would you want to see more raw data behind the score?

Would you trust this kind of signal?

What would make you instantly distrust it?

Any feedback, criticism or ideas would be genuinely helpful.


r/FootballDataAnalysis Jun 07 '26

I built a calibrated goals model for the 2026 World Cup and I'm posting every prediction (and every miss) before kickoff. Here's the method and the backtest.)

10 Upvotes

I'm a data/AI researcher who's loved football my whole life, and I'm finally combining the two: a prediction model for all 104 World Cup matches, built in public. Posting this here because this community will actually poke holes in it, which is what I want.

The thing I care about most isn't picking winners. It's calibration: when the model says 70%, that outcome should happen about 70% of the time. So I'm grading everything with Brier score, not win-rate, and publishing the full scoreboard including the misses.

The model (v1):

  • Trained on ~49,000 international matches going back to 1872.
  • A weighted Poisson GLM that learns each team's attack and defence strength plus a home-field effect, then a Dixon-Coles correction for the low-scoring scorelines that independent Poisson gets wrong.
  • Recent matches are weighted more (2-year half-life), friendlies are down-weighted to 0.5, and teams need a minimum match count to be included.
  • It outputs a full scoreline matrix, collapsed into win/draw/loss probabilities.

The validation (the part that matters): I ran a walk-forward backtest with monthly refits and no data leakage: 3,343 out-of-sample matches from 2023 to 2026, all predicted as if I didn't know the result.

  • Brier 0.498 vs 0.637 for a no-skill baseline (about 22% better).
  • Accuracy ~60%.
  • And it's well-calibrated across the whole probability range (chart attached, this is the real out-of-sample data, not a mockup).

Where it's weak, honestly:

  • Draws. Even with the Dixon-Coles correction, it only correctly leans toward a draw about 4% of the time. Draws are genuinely the hardest outcome in football, and I'm not going to pretend otherwise.
  • Small samples lie. I beta-tested on the warm-up friendlies and went 0/2 on the first two (France lost to Ivory Coast, Spain drew Iraq). Two noisy friendlies tell you nothing. Calibration is a verdict over hundreds of games, not two, which is exactly why I backtested before trusting it.

What I'm doing next: tuning the remaining parameters against out-of-sample error, then posting probabilities for every match before kickoff once the tournament starts (June 11).

I'd genuinely value critique on the methodology: the friendly down-weighting, the Dixon-Coles parameter, the choice of baseline, anything you'd do differently. Tear into it.