Formula 1 race strategy is a temporal decision problem in which teams must act before the final outcome is known, pit stops are a central example: they can restore tyre performance and create strategic opportunities, but they also impose an immediate time loss. This thesis addresses pit decision support from historical Formula 1 data through a reproducible replay and evaluation pipeline, historical races are extracted, enriched, and transformed into replayable event streams, published through Kafka and processed by a Flink based Strategy Engine, which builds race context, emits rule based pit recommendations, and persists JSONL artifacts for downstream learning. These artifacts are then converted into datasets for a Batch XGBoost model and a MOA Online Learner. The evaluation is organized around two contracts; the first contract, pit_any_h2, checks whether an eligible pit stop occurs within a short horizon after a positive call, the second contract, pit_success_h2, is stricter: it requires the matched future pit stop to be classified as successful. This separates anticipating a pit window from recommending a pit that later produces a favourable outcome. The experiments compare three decision paradigms: the deterministic Flink Strategy Engine, Batch XGBoost, and the MOA Online Learner. The results show that pit_any_h2 is the cleaner and more learnable task. The learning systems improve over the deterministic rule baseline, with Batch providing the largest event coverage and MOA the strongest precision-oriented balance. In contrast, pit_success_h2 is substantially harder: successful pit calls are sparse and threshold-sensitive. Overall, this work shows that historical race state contains useful signal for pit timing, while successful pit-call prediction requires stricter accounting and remains harder to cover reliably. The main contribution is an executable pipeline for replaying historical races and comparing deterministic, Batch, and online decision systems.
La strategia di gara in Formula 1 è un problema decisionale in cui i team devono agire prima che l’esito finale sia noto, i pit stop sono un esempio centrale: possono ripristinare la prestazione degli pneumatici, ma impongono una perdita di tempo immediata. Questa tesi affronta il supporto decisionale per le chiamate ai box da dati storici, costruendo una pipeline riproducibile, le gare storiche diventano stream replayabili, pubblicati tramite Kafka e processati da uno Strategy Engine basato su Flink, che costruisce il contesto, emette raccomandazioni e persiste artefatti JSONL. Questi artefatti diventano dataset per un modello Batch XGBoost e un MOA Online Learner. La valutazione è organizzata attorno a due contratti; il primo contratto, pit_any_h2, controlla se un pit stop eleggibile avviene entro un breve orizzonte dopo una chiamata positiva, il secondo contratto, pit_success_h2, è più severo: richiede che il pit stop futuro associato venga classificato di successo. Questa distinzione separa l’anticipazione di una finestra di pit stop dalla raccomandazione di un pit stop con esito favorevole. Gli esperimenti confrontano tre paradigmi: il Flink Strategy Engine deterministico, Batch XGBoost e il MOA Online Learner. I risultati mostrano che pit_any_h2 è il compito più pulito e apprendibile. I sistemi basati su apprendimento migliorano rispetto al baseline deterministico, con Batch che fornisce la maggiore copertura degli eventi e MOA il miglior bilanciamento orientato alla precisione. Al contrario, pit_success_h2 è più difficile: le chiamate a pit stop di successo sono sparse e sensibili alla soglia scelta. Nel complesso, lo stato storico della gara contiene segnale utile per anticipare il timing dei pit stop, mentre la previsione dei pit stop di successo richiede una contabilità più severa e rimane difficile da coprire. Il contributo principale è una pipeline per replayare gare storiche e confrontare sistemi deterministici, Batch e online.
From historical F1 strategy streams to pit stop decision models
Pizzoccheri, Pietro
2025/2026
Abstract
Formula 1 race strategy is a temporal decision problem in which teams must act before the final outcome is known, pit stops are a central example: they can restore tyre performance and create strategic opportunities, but they also impose an immediate time loss. This thesis addresses pit decision support from historical Formula 1 data through a reproducible replay and evaluation pipeline, historical races are extracted, enriched, and transformed into replayable event streams, published through Kafka and processed by a Flink based Strategy Engine, which builds race context, emits rule based pit recommendations, and persists JSONL artifacts for downstream learning. These artifacts are then converted into datasets for a Batch XGBoost model and a MOA Online Learner. The evaluation is organized around two contracts; the first contract, pit_any_h2, checks whether an eligible pit stop occurs within a short horizon after a positive call, the second contract, pit_success_h2, is stricter: it requires the matched future pit stop to be classified as successful. This separates anticipating a pit window from recommending a pit that later produces a favourable outcome. The experiments compare three decision paradigms: the deterministic Flink Strategy Engine, Batch XGBoost, and the MOA Online Learner. The results show that pit_any_h2 is the cleaner and more learnable task. The learning systems improve over the deterministic rule baseline, with Batch providing the largest event coverage and MOA the strongest precision-oriented balance. In contrast, pit_success_h2 is substantially harder: successful pit calls are sparse and threshold-sensitive. Overall, this work shows that historical race state contains useful signal for pit timing, while successful pit-call prediction requires stricter accounting and remains harder to cover reliably. The main contribution is an executable pipeline for replaying historical races and comparing deterministic, Batch, and online decision systems.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Pizzoccheri_Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Sommario
Dimensione
1.19 MB
Formato
Adobe PDF
|
1.19 MB | Adobe PDF | Visualizza/Apri |
|
2026_07_Pizzoccheri_Tesi.pdf
accessibile in internet per tutti
Descrizione: Testo della Tesi
Dimensione
9.4 MB
Formato
Adobe PDF
|
9.4 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/259258