FootBallAnalysis
How much is a football player worth, according to the data?
A football recruitment platform built with CRISP-DM on more than 5.7 million rows of Transfermarkt data. I delivered objective 3: estimating a player’s market value. Data preparation, a comparison of 5 regression models, a LightGBM pipeline served by a FastAPI API, and a search for undervalued players.
- My role
- Objective 3: market value estimation, from data preparation to the deployed LightGBM model
- Platforms
- AI, Web
- Stack
- Python
- LightGBM
- Scikit-learn
- FastAPI
- XGBoost
- pandas
- NumPy
- SHAP
- Machine learning
- Groq API
- Matplotlib
- Seaborn
- Jupyter
- JavaScript
Situation
5,7 MTransfermarkt rows (92,671 players, 2,175 clubs)
Academic team project (2026). Clubs spend billions on transfers with no reliable tool to value a player, spot a talent or anticipate a departure. How can players’ historical data help recruiters decide faster, with less risk?
The project follows CRISP-DM in 6 phases. The team defined 8 business objectives and delivered 4: segment players by profile, spot high-potential young players, estimate market value, anticipate a transfer.
My part: objective 3, estimating what a player is worth, to help a club avoid overpaying and sell at the right price. Intended users: recruiters, sporting directors, agents.
Data: the Transfermarkt datalake, 10 CSV files, 5,673,773 rows, 92,671 players and 2,175 clubs.
Product
- Value a player: the recruiter enters age, height, position, preferred foot and statistics (matches, goals, assists, minutes, cards). The API returns a readable value, e.g. “12.50 M€” or “850 K€”.
- Spot undervalued players: predicted and current values are compared. Players whose predicted value most exceeds their current value are recruitment opportunities.
- An AI scouting agent (Llama 3.3 70B via Groq) runs the team’s models and writes a report that includes the market value estimate.
- The same web interface covers the 4 objectives: player profile, potential, market value, transfer probability.
Estimate a player’s value
≈2.9 M€
likely between 710 k€ and 7.4 M€LightGBM retrained for this demo on 33,508 active players from the project data (last known Transfermarkt value), from their profile, career statistics and club. On 6,702 test players: R² 0.74 (log scale), median error 45 %, 66 % of estimates within a factor of 2; the real value falls inside the range for 78 % of them. An indicative estimate, computed in your browser.
System
Notebooks train the models, which are exported (joblib) and loaded at start-up by a FastAPI API (Uvicorn). The web interface (HTML, CSS, JavaScript) is served by the same application. The Groq key is read from the .env file.
For objective 3, preprocessing is part of the same Pipeline as the model: it is fitted on the training set only (no data leakage) and the API replays it exactly.
| Route | Role |
|---|---|
POST /predict_market_value |
Market value in euros (objective 3) |
POST /predict |
Player profile (objective 1) |
POST /predict_potential |
Potential (objective 2) |
POST /predict_transfer |
Transfer probability (objective 4) |
POST /football_agent_analysis |
Scouting report written by the agent |
GET /health |
API and models loaded |
One request, end to end
Model & data
Data preparation
- Target: each player’s latest known value, transformed with
log1pto reduce skew (many small values, a few stars above €100M). - Performance aggregated per player (matches, goals, assists, minutes, cards), joined to profiles; a player without statistics gets 0.
- Preprocessing (
ColumnTransformer): 8 numeric features imputed with the median then standardised; position and preferred foot one-hot encoded. - Split 80% / 20%.
5 models compared, same preprocessing
| Model | What it brings |
|---|---|
| Linear regression | Simple, readable baseline |
| Random Forest | Robust, captures non-linear effects |
| XGBoost | Regularised gradient boosting |
| LightGBM (selected) | Best R²; fast and lean on tens of thousands of players |
| MLP | Neural network with 2 hidden layers |
The model with the best R² is selected automatically, then analysed: predicted vs actual, residuals, feature importance. A variant with more features (age squared, minutes per match, etc.) explains its predictions with SHAP.
Results, without rounding the truth
The only recorded score is the variant’s: R² = 0.641 (linear regression, MAE 2.850 and RMSE 3.835 on the log scale). The main notebook was not run to the end: the 5 models’ scores are not published; we only know LightGBM had the best R², since it is the one that was exported. An R² around 0.64 is realistic: value also depends on contract, club, league and media exposure, which are not in the data.
Limits and next steps
- Statistics summed over the whole career, and no league: a goal in the first and the second division count the same. Next: stats from the last 1–2 seasons, per-90 ratios, competition level.
- A single train / test split. Next: cross-validation and hyperparameter tuning (Optuna).
- Metrics on the log scale. Next: MAE and MAPE in euros, by value band, and SHAP in the API.
The demo model
For the demo on this page, the model was retrained from the project’s prepared data. Testing the exported model revealed a flaw: trained on retired players too (whose last value is often €0), it had learned to read age as “retired or not”, and valued a 33-year-old at a few euros.
- Data: only the 33,508 active players with a known, positive last value are kept.
- Features: the project’s (age, height, position, strong foot, career goals, assists and minutes), plus the level of the current club: the smoothed mean value of its players. Career totals do not say at what level a player plays; the club does. It is target encoding: computed out of fold on the training set and from the training set only for the test set, so a player never sees his own value.
| On 6,702 test players | R² (log) | Within a factor of 2 |
|---|---|---|
| Median per position (reference) | ≈ 0 | |
| Project features only | 0.39 | 46 % |
| Ridge, with the club | 0.64 | 59 % |
| Random Forest, with the club | 0.74 | 65 % |
| XGBoost, with the club | 0.75 | 67 % |
| LightGBM, with the club (kept) | 0.74 | 66 % |
- Tuning: randomized search (30 draws, 5-fold cross-validation); a second, wider search (up to 2,500 trees) did no better. Monotone constraints make sure more goals or assists never lower the estimate.
- Final result: R² 0.74, 45 % median error in euros, 66 % of estimates within a factor of 2.
- Range: two quantile models (10th and 90th percentiles) frame the estimate; the real value falls inside for 78 % of test players, for 80 % intended.
- Guards: beyond what 99 % of training players show (by age and by position), the demo flags an unusual profile. A club missing from the data gets the average level, and the demo says so.
- The model is exported to JSON and computed in the browser; a check confirms on 3,000 players that its predictions are Python’s (difference below 0.00001 on the log scale).