An ML-powered dynamic pricing and decision support system for ticket pricing in a sports stadium. The engine forecasts demand at any candidate price, grid-searches for the revenue-maximizing recommendation, and surfaces it on a one-click approval dashboard. The same architecture was deployed on a real club's ticketing data and delivered +6% revenue per match and 86% recommendation adoption by the commercial team.
Pick a match and zone, see the recommended price and the revenue-vs-price curve, then approve. βΆ Try the live demo.
pip install -r requirements.txt
make all # synthetic data, train, holdout evaluation, elasticity sanity check
make app # open the human-in-the-loop Streamlit page locallyThe challenge: transform a static, manual pricing strategy into a responsive, automated system with a human-in-the-loop, creating a market-driven approach to setting ticket prices per match. Prices were historically set once per season in rigid categories and updated weekly or monthly. The team would manually pull data from several systems, propose changes, then push them to the live ticketing platform by hand. The engine collapses that loop into "data β recommendation β one-click approval β live price".
Fig. 1: A standard stadium ticket pricing by zone during the checkout process.
Click for the full problem β solution breakdown
| π© Problem | π‘ Solution |
|---|---|
| Static pricing: prices set once per season in rigid categories (A++, A, B), updated weekly/monthly. | Dynamic recommendations: price proposals per seating zone based on near real-time data analysis, allowing daily updates. |
| Manual adjustments: slow analysis to propose changes. | Impact simulation: instantly model projected impact of any price change on revenue and ticket sales. |
| Data bottleneck: manual extraction from fragmented systems. | Centralized data: aggregates sales, web analytics, contextual data into one place. |
| Slow implementation: disconnected from the sales platform. | Seamless integration: one-click approval on a dashboard pushes a price update to the live ticketing system. |
The Dynamic Pricing Engine ingests historical data from the Club's Data Systems and real-time sales from the Ticketing System, recommends prices that are simulated and approved by the Club's Pricing Team, and pushes approved prices back to the live ticketing system where fans complete a purchase. That last step closes the feedback loop.
Fig. 2: System Context Diagram β Dynamic Pricing System.
Click for the high-level market-dynamics view
The engine acts as the central brain balancing the club and the fan, ingesting internal and external factors to forecast demand at various price points.
make evaluate runs a 14-day per-series holdout. The ensemble is re-trained from scratch on the train split before predicting on the holdout, so the numbers below are leakage-free.
| Metric | Ensemble (Prophet + XGBoost) | Baseline (mean) |
|---|---|---|
| WAPE | 26.7% | 79.7% |
| RΒ² | 0.726 | -0.428 |
| MAE | 6.2 tickets | 18.6 tickets |
| RMSE | 11.8 tickets | 27.0 tickets |
The ensemble's WAPE is 67% lower than the mean baseline. Reproducible with RANDOM_SEED=42 in src/data/make_dataset.py.
A common failure mode for demand models is to learn everything except the price-to-sales relationship, leaving the optimizer to recommend the price cap on every row. make sanity defends against this: it samples 20 historical rows, sweeps each through the optimizer's price range, and asserts that predicted sales move with price by at least 20% relative spread. Latest run: mean spread 11.6 tickets across the per-zone range, median optimal price β¬182 (well inside the search band), 0% violation rate.
βΉοΈ These numbers are not reproduced by this repository β they come from a deployment on a confidential real-world dataset. Treat them as case-study evidence, not a benchmark.
| Metric | Result | Description |
|---|---|---|
| π Revenue uplift | +6% avg. revenue per match | Achieved by dynamically adjusting prices to match real-time demand forecasts. Validated via A/B testing. |
| ποΈ Optimized sales | +4% sell-through rate | Improved occupancy alongside revenue, which positively affects atmosphere and in-stadium sales. |
| βοΈ Operational efficiency | 7Γ faster price changes | From weekly to daily updates by automating data aggregation and analysis. |
| π€ Recommendation adoption | 86% of proposals approved | Commercial team reviewed and approved the model's price proposals at a high rate, indicating trust. |
The engine was validated via segment-based A/B tests: a subset of seating zones used the dynamic engine (treatment), the rest stayed on static pricing (control). Tests ran across matches of varying importance to ensure the lift wasn't an artifact of any single event. The +6% revenue lift held alongside a +4% sell-through rate, confirming the engine found market equilibrium rather than simply over-charging.
Two stages: first predict, then optimize.
Fig. 3: Dynamic Pricing Engine component.
A Prophet + XGBoost residual ensemble. One Prophet model per (match_id, seat_zone) series captures temporal structure (trend, weekly seasonality, weekday/holiday effects). A single XGBoost regressor then fits Prophet's in-sample residuals using the full feature set (price, demand signals, external factors), picking up the non-linear interactions Prophet misses. The combination beats either alone on the holdout. The unified prediction surface lives in src/models/predict_demand.py as DemandModel.predict().
Click for design choices and trade-offs (Stage 1)
| Aspect | Description |
|---|---|
| Stage A β Prophet | One model per (match_id, seat_zone) series captures trend, weekly seasonality, and weekday/holiday effects via Prophet regressors. |
| Stage B β XGBoost | A single XGBoost regressor is fit on Prophet's in-sample residuals using the full feature set. |
| Prediction | final = clip(prophet_yhat + xgb_residual, 0, β). |
| Why this split | Prophet handles temporal structure cleanly; XGBoost picks up complex non-linear interactions Prophet cannot. The combination beats either alone on the holdout. |
Grid search over a zone-aware range of prices: [0.5 Γ base_price, 2.5 Γ base_price]. Outside this band the model would be extrapolating beyond the training distribution and the residual XGBoost can't be trusted; the band is configurable in src/decision_engine/constants.py (PRICE_SEARCH_RANGE_RATIO). For each candidate price, the engine builds a row, predicts sales with DemandModel, computes revenue = price Γ predicted_sales, and returns the argmax. The result becomes a Price Variation Proposal sent to the commercial team for approval.
Click for design choices and trade-offs (Stage 2)
| Aspect | Description |
|---|---|
| Why grid search | Pricing is a critical business decision; grid search guarantees the revenue-maximizing price within the search space, at modest compute cost (one vectorized prediction batch per match-zone). |
| Process | For each candidate price in the range, build a row, predict sales with DemandModel, compute revenue = price Γ predicted_sales, return the argmax. |
| Why not Bayesian opt. | Bayesian optimization would converge faster but doesn't guarantee the maximum. For pricing decisions, the guarantee is worth the modest extra cost. |
The repository ships a synthetically generated dataset engineered to mirror the complexity and statistical properties of a real ticketing environment: 10 matches of varied importance, a 90-day daily sales window per match, and up to 5 stadium zones per match (a small per-zone dropout probability removes some pairs to mimic real data gaps, yielding ~37-43 of the 50 possible series).
β οΈ web_conversion_rateis deliberately dropped from the model's feature set (src/features/build_features.py). It is defined assales / web_visitsin the generator, which is target leakage at decision time: when we propose a price, the conversion rate at that price is precisely what we are trying to predict. Including it would inflate the headline metrics while breaking the optimizer's price-elasticity signal.
Click for the full feature schema
| Category | Features | Description |
|---|---|---|
| Match & Opponent | match_id, days_until_match, is_weekday, opponent_tier, ea_opponent_strength, is_international |
Core details about the match, its timing, and opponent quality. |
| Team Status | team_position, top_player_injured, league_winner_known |
Current performance, player status, and league context. |
| Ticket & Zone | seat_zone, ticket_price, ticket_availability_pct, zone_seats_availability |
Attributes of the specific ticket and seating area. |
| Demand & Hype | internal_search_trends, google_trends_index, social_media_sentiment, web_visits |
Digital signals measuring interest and purchase intent. |
| External Factors | is_holiday, popular_concert_in_city, competitor_avg_price, flights_to_barcelona_index |
External events, competition, and tourism proxies. |
| Weather | weather_forecast |
Forecasted weather conditions for the match day. |
zone_historical_sales[Target Variable] β the historical number of tickets sold in a given zone-day. This is what the model predicts.
Click for how prices and sales are generated
Two pieces of the data generator carry the model's job:
- Price has a wide, partly-independent distribution. Real-world prices reflect both a strategic baseline tied to match excitement and operational variation (A/B tests, promotions, last-minute discounts). The generator implements both, so price varies roughly between
0.5Γand2.5Γthe zone's base price even within a single excitement level. Without this, the model cannot identify price elasticity from historical observations alone. - Demand follows a linear-elasticity curve with a hard ceiling at
2.5Γbase price. This yields a clean interior revenue optimum near~1.25Γbase, rather than a runaway "pick the price cap" recommendation.
Click for the Match Excitement Factor
The generation script unifies non-price demand drivers under a single "Match Excitement Factor":
- Starts with the opponent: a top-tier opponent generates more interest.
- Adjusts for context: league position, player injuries, match importance (e.g. league winner already decided), proximity to holidays, weekday/weekend.
- Drives the demand signals:
google_trends_index,social_media_sentiment,internal_search_trendsall scale with this factor.
dynamic-pricing/
βββ Makefile # Pipeline: data β train β evaluate β sanity β app
βββ README.md
βββ app.py # Streamlit HiTL page
βββ config.py # Paths and reference dates
βββ requirements.txt
βββ assets/ # Diagrams and images
βββ data/
β βββ 03_synthetic/
β βββ synthetic_match_data.csv # Generated by make_dataset.py
βββ models/ # Trained artifacts (regenerable, gitignored)
β βββ prophet_models.joblib
β βββ xgb_residual_model.joblib
β βββ feature_pipeline.joblib
βββ src/
βββ data/
β βββ make_dataset.py # Synthetic data generator (seeded)
βββ features/
β βββ build_features.py # Pipeline factory: drops, scales, one-hot encodes
βββ models/
β βββ train_demand_model.py # Fits Prophet + XGBoost ensemble
β βββ predict_demand.py # DemandModel: unified predict() surface
β βββ evaluate.py # Leakage-free holdout metrics
β βββ sanity_check.py # Asserts the model actually responds to price
βββ decision_engine/
βββ simulate.py # What-if for a single price
βββ optimize.py # Zone-aware grid-search optimal price
βββ constants.py # Sample feature row + zone base pricespip install -r requirements.txt
make clean # remove cached artifacts
make all # data + train + evaluate + sanity check
make app # launch the Streamlit page
# Or run the CLI examples directly:
python -m src.decision_engine.simulate
python -m src.decision_engine.optimizeThe Streamlit page lets you pick a match and seat zone, see the recommended price with a revenue-vs-price curve, simulate any hypothetical price, and approve the recommendation β which appends a JSON line to proposals.jsonl. That's the HiTL loop in miniature.
MIT