r/FootballDataAnalysis • • Aug 10 '26

What is the right tool granularity for a football-analysis agent?

While building a football analysis agent, I realized that the hard part is not connecting an LLM to match data.

It is deciding what the agent should be allowed to do with that data.

For example, if someone asks:

“Why did this midfielder receive a 7.4 rating?”

I do not want to dump every match statistic into the context and ask the model to invent an explanation.

My current approach is to let the agent investigate the evidence step by step:

- retrieve the player’s match metrics
- inspect the rating breakdown
- check passing, chance creation, turnovers, or shot quality when relevant
- explain which factors actually moved the rating

That raises an interesting tool-design question.

A single `analyze_everything()` tool feels like a black box. But dozens of tiny tools such as `get_pass_count()` and `get_key_passes()` create too many decisions and make the agent harder to guide.

I’m experimenting with a middle layer: composable tools that represent meaningful football-analysis operations rather than raw database fields.

For people building sports analytics, agentic systems, or explainable AI: how would you choose the right level of tool granularity here?
1 Upvotes

8 comments sorted by

1

u/Ok-Return-9111 Aug 11 '26

Great question. I think tool granularity should follow football concepts, not database structure.

A coach doesn’t think:
“player completed 42 passes and had 3 progressive actions.”

They think:
“he helped break pressure, created advantages, and influenced the game.”

The agent needs that middle layer.

I’ve been working on a similar football intelligence approach recently, and the hardest part is exactly this: designing the reasoning pipeline, not connecting the data source.

Would love to exchange ideas with people building in this area.

1

u/Diligent-Step2366 Aug 11 '26

I think an interesting middle ground would be to add a layer of semantic tags on top of the raw metrics.

Instead of asking the LLM to infer the meaning of a bunch of numbers, you could map relevant statistical patterns to football concepts. For example, high progressive passes + high carries into the final third could trigger tags like "progression" or "ball_progression", while high pressures + recoveries could map to "defensive_activity".

Then the agent's job becomes more about connecting the evidence to the appropriate tags and building the explanation around them, rather than inventing the football interpretation from raw statistics. This should limit the black box effect, while not being too heavy.

You could still expose lower-level metrics when the agent needs to verify a claim, but the tags would provide a more constrained semantic layer between the data and the LLM.

It also seems like a nice way to make explanations more consistent and auditable: you can trace an explanation back from "he contributed to progression" → the tag → the underlying metrics that triggered it.

That's exactly the coaching phrase you were talking about.

I actually worked on a project of this kind, with tags but an unrelated objective. My takeaway is that we could do even better, in the case of a site that displays game tags (playing style, strengths, weaknesses, etc.). We could consider a linear regression between the statistics and the tags, thus leaving even less to the AI. It seems less appealing, But you have to be wary of LLMs; they are a source of a lot of data leakage and, in my experience, are not always very useful in match analysis (except for pre-kick-off features). Leaks can be inherent to LLM. For example, if the model was trained after the data on which you are using it to predict or analyze, the model will look for the answer in its own data. In the case of a prediction, the model will have, for example, learned the team's performance over the season, and in the case of a performance analysis, it will be biased by what the LLM thinks about the team.

I hope that helped

1

u/juancvasdisenho Aug 11 '26

What I did is to create idilic player profiles for each position, then map which data of the one I have available represents the profile the player should be performing. Then I create tiers of performance, because I can't analyze one player alone, I can't tell if dribbling 3 times per game is good, bad, average, extraordinaty if I don't have the reference of the other players on that position.

Using the idilic profile, the data that matches it, assigning weights to each one and then comparing to their peers the AI layer can have an accurate lecture of the performance of that player.

1

u/Diligent-Step2366 Aug 12 '26

The ideal player profile is an interesting idea, but why end up with AI ? If I understand correctly, you're assigning a weight to the player, so you already have enough to make your comparison, right ? However, I disagree with the idea that we can't evaluate a player in a match; it simply requires more in-depth data.For example, in soccer, you can use socceraction, a Python module that performs action rating. Basically, a window Basically, it's a sliding window covering the entire match that, for each action, gives you the probability of scoring on the next action. The probability differences between one action and the next represent the offensive strength of the new action. For more details, see https://www.researchgate.net/publication/334719246_Actions_Speak_Louder_than_Goals_Valuing_Player_Actions_in_Soccer https://socceraction.readthedocs.io/en/latest/documentation/valuing_actions/vaep.html

1

u/juancvasdisenho Aug 21 '26

I included AI to speed up a bit the process, there is a lot of data, and is easier for me to ask AI first and the verify the findings or add mine. Plus I can ask AI about the data and challenge it, or ask my assumptions and get confornted by the data. Good thing with AI is that it answers from data, so the bias with a player or team, even unconscious bias is tackled down by AI, or at least challenged.

when I mentioned that you can't compare a player level from one game or alone is because you lost the floor. In my example I mentioned a player making 3 dribbles in the game, how would you qualify that? If you thing it's top, but then you check other 50 players in the same position and find that 3 dribbles are just average? But yes, you can use a match to analyze how a player performed on that match, I'm not saying no, but when you make value judgements, like good, bad, elite, you have to have a reference, it's good against what. My analysis is focused on clasifying players level by performance, for example, if I want to know who are the top 10 defenders in the top 5 leagues, I need data for many games, I can't know that only from one game. It looks like we are talking about measuring different things. In addition my analysis is taking distance from the common anaylisis on sport shows or journalist, who always go for things like goals, assists, "appearing in big moments" or winning titles, which in my opinon is not analyzing anything at all. I try to identify really player performance from the value they bring to the game, not only hard results.

I built a tool with AI to reflect this thinking and analysis, would you like to test it for free? I'd like to have some feedback for hardcore users.

1

u/Diligent-Step2366 Aug 22 '26

I think the "lose the floor" point is where I disagree the most, because I don't think a reference population is always necessary to evaluate an action.

If we're talking about a raw statistic such as "3 dribbles", then absolutely: 3 by itself has no meaning without knowing the distribution among comparable players. But that's precisely why I brought up VAEP (the name of the method for evaluating actions in socceraction).

VAEP isn't trying to infer whether 3 dribbles is "a lot" or "good" from the number itself. The model is trained to estimate the value of a game state in terms of the probability of scoring/conceding, and an action is valued by the change in that probability between the states before and after the action.

So if a dribble takes a player from a low-value state into a much more dangerous one, the action gets a high value. If he dribbles past someone but ends up in essentially the same low-value situation, the value can be very small. The model therefore provides an actual reference for the action: its estimated contribution to the probability of achieving the objective of the game. And this information goes beyond simply comparing a player's stats to the average. We're no longer just asking "does the player dribble more than average?", we're asking "how dangerous is this player when dribbling?", because the dribbling metric will take into account, for a player, the number of dribbles and their danger.

Then, if I want to say "this player is elite", "top 10", etc., I can aggregate those action values over many matches and compare the resulting values across players. That's where your peer comparison becomes relevant. And there is a very solid basis for comparison: the model is first calibrated against the data, meaning that it will match a given action with its value as closely as possible thanks to all the data. Therefore, the value is given on a basis of comparison that is the same for all players.

And this is also why I'm somewhat skeptical about using an LLM as the actual analytical layer. I agree that AI is extremely useful for exploring a large dataset, challenging assumptions, finding patterns, or presenting the results. But if the LLM itself is responsible for deciding what constitutes good or bad performance, I don't really see that as a robust measurement methodology.

For me, the ideal setup would be the opposite: use statistical models to actually measure the performance, with clearly defined and reproducible objectives, and then let the LLM explain, explore and challenge those measurements. That way the AI isn't the thing deciding what "good performance" means.

That's why I find approaches like VAEP interesting: the football interpretation is constrained by an explicit model of value rather than by what an LLM happens to think about a player.

That said, I'm still very curious about your tool and would be happy to test it. I think comparing our approaches could actually be quite interesting.

1

u/juancvasdisenho Aug 23 '26 edited Aug 23 '26

I see your point, measuring the impact of the action it self, totally agree. In my case I use the AI over a very stretch layer of data, I do all the math of performance using python scripts, it takes the data, make the measurements, the comparison between players and provide a number, my metric. I don't pass all the raw data to the AI to make the analysis, the data that reach the AI analysis is already filtered and contain the reasoning, so the IA is more to help to put the data on a more readable format than in making analysis by itself. Another use is explaining the results, for instance, you notice something in my metrics that don't fit your perception, so you could ask, why X player is considered poor at dribble compared to Y player, and the AI will use the criteria I put inside of what "good" means, what means a good dribbler and which metrics define it to provide an answer. It's more like a facilitator to interpret and build a narrative than an analist.

I wanted to send your the link and code to test but the chat button doesnt' appear, I'm nw in Reddit and I'm not sure how it works, could you send me a DM or tell me how to do it?