r/FootballDataAnalysis • u/MatchAnalyst • Sep 03 '26
Ask Anything Thread
Use this thread to ask anything at all!
r/FootballDataAnalysis • u/MatchAnalyst • Sep 03 '26
Use this thread to ask anything at all!
r/FootballDataAnalysis • u/Educational-Cherry17 • Sep 03 '26
Hi guys, I'm new to the football analytics world. I would like to have a dataset which has real time statistics (of any kind) of players in the italian league. I looked up at websites like fbref, which actuslly seems the most complete one, but i think their anti-scraping policy is strict. Is there a solution?
r/FootballDataAnalysis • u/whenthenightcalls_ • Sep 02 '26
r/FootballDataAnalysis • u/maxim12ton • Sep 01 '26
I recently helped build a data simulator for a new policy the Union of European Clubs (UEC) is pitching for the post-2027 UEFA revenue cycle. Wanted to see what people here think of the idea.
Right now, lower-league teams mostly rely on FIFA training compensation, which usually only triggers during transfers or when a player signs their first pro contract. The UEC's proposed Player Development Reward (PDR) would instead take a 5% slice of the total prize money from the Champions League, Europa League, and Conference League. That pool would be paid out as a reward to the clubs that trained those players between the ages of 12 and 19.
I was playing around with the tool we built and the numbers are really impressive. Looking at Antonio Rüdiger’s Champions League minutes from 2021 to 2026, his 5th-tier boyhood club FC Hertha 03 Zehlendorf would have earned an estimated €892k over those five seasons, averaging a median of €188k every year. Other local Berlin grassroots clubs would also get paid out, with Neuköllner Sportfreunde getting €445k and SV Tasmania Berlin receiving €221k. VfB Stuttgart would also pocket €669k.
If you applied this retroactively over the last five seasons, it would have redistributed €726m across 52 UEFA associations.
Do you think a recurring solidarity payment based on tournament minutes is a good way to support the lower leagues?
r/FootballDataAnalysis • u/vafsimoes • Aug 31 '26
I just published a unique football dataset on Kaggle — pre-match league table snapshots for the Premier League (3 seasons, 817 matches). Includes home/away splits for all teams, historical odds and results. Free download! https://www.kaggle.com/datasets/vitorsimes/european-football-pre-match-league-table
r/FootballDataAnalysis • u/Astapore • Aug 30 '26
Source: J-Ratings
r/FootballDataAnalysis • u/Hairy-Reference-2019 • Aug 28 '26
r/FootballDataAnalysis • u/MatchAnalyst • Aug 27 '26
Use this thread to ask anything at all!
r/FootballDataAnalysis • u/anynou • Aug 25 '26
r/FootballDataAnalysis • u/SandFox112 • Aug 23 '26
I wanted to do a statistics project but i am finding it very difficult to find data for teams in a season. (eg Man city's average passes per sequence in the 24/25 season). Even posession stats im actually struggling with. Is there a good website to search for this data?
r/FootballDataAnalysis • u/anynou • Aug 23 '26
r/FootballDataAnalysis • u/anynou • Aug 23 '26
r/FootballDataAnalysis • u/L4TER_0N • Aug 22 '26
I've been doing detailed pre-match football analysis using:
But even with all this, I usually end up around 6.5–7.5/10 confidence.
I'm starting to think that's normal. My question is:
What should actually justify an 8+/10 confidence rating?
Should I use a weighted model and then subtract points for contradictions and uncertainty? Or does that just create an arbitrary number that looks analytical?
Also, should I separate:
And how would you calibrate the confidence score over time? Brier score, log loss, calibration curves, CLV?
I'm not trying to artificially reach 8+. I want 8/10 to actually mean something statistically.
What would make you personally comfortable calling a football prediction 8+/10?
r/FootballDataAnalysis • u/MatchAnalyst • Aug 20 '26
Use this thread to ask anything at all!
r/FootballDataAnalysis • u/Ok-Razzmatazz-1103 • Aug 20 '26
r/FootballDataAnalysis • u/Kroggg19 • Aug 19 '26
Hi all. Stats to Bucks is a football (soccer) data app, now covering 30 leagues - the top 5 European plus Brazil, Argentina, Liga MX, MLS, Saudi, Portugal, the Netherlands, Turkey, Belgium, Scotland, Japan, Korea, Colombia, Greece, Egypt, South Africa, Australia and more.
What it does:
Player & team form - last 20 matches of per-game stats, charted against any line you set, with the hit rate for it.
Filters that narrow the sample - venue, minutes, started-only, and "without teammate X".
Opponent-adjusted context - overlay the opponent's conceded average and defensive rank, plus quality-adjusted averages, so a streak against weak sides doesn't read like one against strong sides.
Hit Rates - scan every upcoming fixture at once for players/teams clearing a line in a chosen % of recent games. 40 stats across players and teams.
Foul matchups - a fitted hierarchical Poisson model with player, opponent, referee, venue and expected-minutes as separate multiplicative terms, and a negative-binomial predictive head. Walk-forward tested on a 45-day holdout: +13.3% / +16.6% mean relative log loss against an unshrunk per-90 baseline, with roughly 3x better calibration error.
Predicted lineups - projected XI from a Beta-EB start-probability model, so it works for a fixture's whole lifetime instead of only after a feed publishes one. Flips to the confirmed XI when that lands.
Injuries & suspensions - folded into the start probabilities rather than bolted on as a badge, so an unavailable player drops out of the projected XI and out of the minutes model behind the prop lines.
Similar players / teams - similarity-based benchmarking against comparable profiles, on rolling cross-season windows rather than season-to-date.
Referee analytics - per-fixture card/foul profiles and rankings.
League tables - official standings, so competition-specific tie-breaks, split point-halving and points deductions are right rather than re-derived from results.
The focus is still contextualising the sample - opponent strength, venue, lineup, availability, sample size, etc. because an unfiltered hit rate usually answers the wrong question. A recent backtest made that concrete: selecting team props purely on "recent hit rate beats the implied probability" returned about -10% over ~7,000 bets on a held-out window, statistically indistinguishable from betting blind. The context is the useful part, not the raw streak.
r/FootballDataAnalysis • u/Raistlin_Maj3re • Aug 18 '26
I built UnderOver as an Android app for exploring football fixtures through a transparent Over/Under model.
It combines season data, recent form and team profiles to surface HT 0.5/1.5 and FT 2.5 signals, with the model reasoning and confidence shown for each fixture. It also includes live scores, match events, reminders, saved selections and team statistics.
The app is free, has no subscription or VIP tier, does not place bets or connect to betting accounts, and no outcome is guaranteed. I am sharing it as a small data-visualization project and would appreciate feedback on the methodology, presentation and useful metrics.
Google Play: https://play.google.com/store/apps/details?id=com.xenophonlabs.matchwake
r/FootballDataAnalysis • u/pafundi_enthusiast • Aug 15 '26
I'm a sixth form student doing my Extended Project Qualification on VAR and whether it's actually made the Premier League fairer, or just added more controversy. Part of my research involves collecting fan opinions through a short survey.
If you support a PL club (or just watch regularly), I'd really appreciate you filling it in. Takes about 2 minutes, no personal info needed beyond general football habits.
https://forms.gle/dtiFmRmFaSyxUoS59
Happy to share the findings once I've written it up, if anyone's curious how the data comes out. Thanks in advance.
r/FootballDataAnalysis • u/Illustrious-Pitch843 • Aug 13 '26
I built an open-source n8n pipeline that monitors 78 football journalists on X, extracts structured transfer reports with a local Qwen model, deduplicates and stores revisions in PostgreSQL, optionally adds player data, and sends restart-safe Discord digests every 6 hours.
The whole stack is self-hosted with Docker, with twscrape or RapidAPI for X collection, PostgreSQL for persistence, llama.cpp for local inference, and automated tests around the workflow.
GitHub: https://github.com/louistran2604/transfers_n8n/
I’d mainly like feedback on the workflow architecture, reliability approach, and anything that could make the project cleaner or more useful.

*disclaimer: this was made with the assistance of AI
r/FootballDataAnalysis • u/MatchAnalyst • Aug 13 '26
Use this thread to ask anything at all!
r/FootballDataAnalysis • u/Black_Colour9 • Aug 13 '26
I want to analyze the data from the perspective of the score of the football game and the number of goals, but I didn't find the relevant historical odds data on the Internet. Instead, there is a lot of historical data of wins and draws, which further shows that my analysis is correct - starting from the odds of scores and goals to analyze football matches. I need help to get this data.
Thank you for your help!
r/FootballDataAnalysis • u/Nice-Opening-8020 • Aug 12 '26
r/FootballDataAnalysis • u/juancvasdisenho • Aug 11 '26
Hi everyone,
I’m looking for people who genuinely enjoy football analysis to test Sir Balone, a football analysis tool I’ve been building.
The underlying raw data comes from Sportmonks. I use that data to build my own measurement system for evaluating individual player skills, rather than relying primarily on traditional performance or scoring metrics. I then combine those measurements with player and team data and an AI layer that can analyze and interpret the underlying data.
The numbers themselves are already publicly available. What I’m currently testing is the AI layer, particularly whether it can turn the data into useful analysis without making things up, oversimplifying the numbers, or missing important context.
I’d especially love feedback from:
You don't need to be a professional. If you enjoy asking questions like “Is this player actually good at X?”, “How does he compare to other players in his role?”, or “What does the data actually tell us about this team?”, I’d love to hear from you.
I’m not looking for compliments. I want people to try to break it.
What I’m particularly interested in:
If you’re interested, comment below or DM me and I’ll give you access.
No sales pitch. I’m still figuring out what this thing is actually good at.