Large Language Models (LLMs) increasingly power natural language interfaces to structured data, yet deploying them in production raises concerns beyond accuracy, namely inference latency and energy cost. To study these, this thesis builds a Sales Agent: an LLM-powered system, orchestrated with LangGraph, that translates a business question into a SQL query, interprets the results in text, and optionally renders a chart. A hierarchical YAML configuration exposes the studied inference parameters (sampling parameters, best-of-N, token budgets, and Chain-of-Thought refinement depth). The core contribution is a systematic profiling that jointly measures output quality, execution time, and per-step energy consumption (via CodeCarbon). Quality is assessed over 20 analytical questions, stratified across four SQL difficulty levels, along three dimensions: a graded ground-truth table-similarity score for retrieval and two LLM-as-judge scores for analysis and visualization. Five step-isolated experiments establish a baseline for three open-weight models (Gemma 4 E4B, Gemma 4 26B, Mistral Small 3.2 24B) and then vary one agent step at a time, closing with a Pareto confirmation run. They are also designed to test a common assumption in local LLM deployment, that larger thinking models are necessarily the best solution. The results show instead that model choice alone does not explain performance: the dominant effects arise from the interaction between model family, agent step, prompt difficulty, and inference-time settings. Mistral is the most efficient untuned model but collapses on the hardest aggregation prompts; once its lookup step is repaired through Chain-of-Thought refinement, however, it attains the highest strict quality and defines the final Pareto frontier. Gemma 26B is highly tunable yet energetically costly, and Gemma E4B, repaired through lookup Chain-of-Thought refinement at a lower temperature, remains a very close high-quality alternative but is dominated once energy is counted. The sensitivity analysis further shows that hyperparameter effects are strongly step-specific, that larger token budgets help only genuinely truncated model-step pairs, and that Chain-of-Thought refinement pays off only when its feedback matches the failure mode: it is the strongest intervention for verifiable SQL lookup, but harmful for visualization, where errors are semantic rather than executable. The thesis distills three recommendations for energy-aware LLM data agents: tune hyperparameters per agent step rather than globally, allocate inference-time compute by failure mode, and report accuracy and energy by prompt difficulty, since the best model for simple tasks is not the best for hard ones.
I Large Language Model (LLM) alimentano sempre più spesso interfacce in linguaggio naturale verso dati strutturati, ma il loro impiego in produzione solleva questioni che vanno oltre l'accuratezza, ovvero la latenza di inferenza e il costo energetico. Per studiare questi aspetti, la tesi realizza un Sales Agent: un sistema basato su LLM, orchestrato con LangGraph, che traduce una domanda di business in una query SQL, ne interpreta i risultati in forma testuale e, opzionalmente, genera un grafico. Una configurazione gerarchica in YAML espone i parametri di inferenza studiati (parametri di campionamento, best-of-N, budget di token e profondità del raffinamento Chain-of-Thought). Il contributo principale è un profiling sistematico che misura congiuntamente la qualità dell'output, il tempo di esecuzione e il consumo energetico per singolo step (tramite CodeCarbon). La qualità è valutata su 20 domande analitiche, stratificate su quattro livelli di difficoltà SQL, lungo tre dimensioni: un punteggio graduato di similarità tabellare rispetto alla ground truth per il recupero dei dati e due punteggi LLM-as-judge per l'analisi testuale e la visualizzazione. Cinque esperimenti a step isolato definiscono una baseline per tre modelli open-weight (Gemma 4 E4B, Gemma 4 26B, Mistral Small 3.2 24B) e poi variano uno step dell'agente alla volta, concludendo con una conferma di Pareto. Sono inoltre progettati per verificare un'assunzione comune nel deployment locale di LLM, ovvero che i modelli thinking più grandi siano necessariamente la soluzione migliore. I risultati mostrano invece che la sola scelta del modello non spiega le prestazioni: gli effetti dominanti emergono dall'interazione tra famiglia del modello, step dell'agente, difficoltà del prompt e impostazioni di inferenza. Mistral è il modello non ottimizzato più efficiente ma crolla sui prompt di aggregazione più difficili; tuttavia, una volta corretto il suo step di lookup con il raffinamento Chain-of-Thought, raggiunge la qualità strict più alta e definisce la frontiera di Pareto finale. Gemma 26B è altamente ottimizzabile ma energeticamente costoso; Gemma E4B, corretto con raffinamento Chain-of-Thought sullo step di lookup a temperatura più bassa, rimane un'alternativa ad alta qualità molto vicina ma è dominato una volta considerata l'energia. L'analisi di sensitività mostra inoltre che gli effetti degli iperparametri sono fortemente specifici per step, che budget di token più ampi aiutano solo le coppie modello-step realmente troncate, e che il raffinamento Chain-of-Thought conviene solo quando il suo feedback corrisponde al tipo di errore: è l'intervento più efficace per il lookup SQL verificabile, ma dannoso per la visualizzazione, dove gli errori sono semantici e non eseguibili. La tesi distilla tre raccomandazioni per agenti dati LLM attenti all'energia: ottimizzare gli iperparametri per singolo step dell'agente anziché globalmente, allocare il compute a tempo di inferenza in base al tipo di errore, e riportare accuratezza ed energia per livello di difficoltà del prompt, poiché il modello migliore per i compiti semplici non è il migliore per quelli difficili.
Profiling framework for LLM-based data agents: quality, latency, and energy under parameter sensitivity
VALENZISI, DAVIDE;El Oukili, Ossama
2025/2026
Abstract
Large Language Models (LLMs) increasingly power natural language interfaces to structured data, yet deploying them in production raises concerns beyond accuracy, namely inference latency and energy cost. To study these, this thesis builds a Sales Agent: an LLM-powered system, orchestrated with LangGraph, that translates a business question into a SQL query, interprets the results in text, and optionally renders a chart. A hierarchical YAML configuration exposes the studied inference parameters (sampling parameters, best-of-N, token budgets, and Chain-of-Thought refinement depth). The core contribution is a systematic profiling that jointly measures output quality, execution time, and per-step energy consumption (via CodeCarbon). Quality is assessed over 20 analytical questions, stratified across four SQL difficulty levels, along three dimensions: a graded ground-truth table-similarity score for retrieval and two LLM-as-judge scores for analysis and visualization. Five step-isolated experiments establish a baseline for three open-weight models (Gemma 4 E4B, Gemma 4 26B, Mistral Small 3.2 24B) and then vary one agent step at a time, closing with a Pareto confirmation run. They are also designed to test a common assumption in local LLM deployment, that larger thinking models are necessarily the best solution. The results show instead that model choice alone does not explain performance: the dominant effects arise from the interaction between model family, agent step, prompt difficulty, and inference-time settings. Mistral is the most efficient untuned model but collapses on the hardest aggregation prompts; once its lookup step is repaired through Chain-of-Thought refinement, however, it attains the highest strict quality and defines the final Pareto frontier. Gemma 26B is highly tunable yet energetically costly, and Gemma E4B, repaired through lookup Chain-of-Thought refinement at a lower temperature, remains a very close high-quality alternative but is dominated once energy is counted. The sensitivity analysis further shows that hyperparameter effects are strongly step-specific, that larger token budgets help only genuinely truncated model-step pairs, and that Chain-of-Thought refinement pays off only when its feedback matches the failure mode: it is the strongest intervention for verifiable SQL lookup, but harmful for visualization, where errors are semantic rather than executable. The thesis distills three recommendations for energy-aware LLM data agents: tune hyperparameters per agent step rather than globally, allocate inference-time compute by failure mode, and report accuracy and energy by prompt difficulty, since the best model for simple tasks is not the best for hard ones.| File | Dimensione | Formato | |
|---|---|---|---|
|
Thesis.pdf
accessibile in internet per tutti
Descrizione: Thesis
Dimensione
4.88 MB
Formato
Adobe PDF
|
4.88 MB | Adobe PDF | Visualizza/Apri |
|
Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Executive summary
Dimensione
653.22 kB
Formato
Adobe PDF
|
653.22 kB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/261164