Deploying Large Language Models at scale is expensive and technically demanding, mak- ing inference optimization a critical priority for any organization self-hosting AI work- loads. This thesis benchmarks two cloud hardware accelerators — the NVIDIA T4 GPU (g4dn.xlarge) and the AWS Inferentia2 ASIC (inf2.xlarge) — across a controlled grid of 198 configurations combining two open-source transformer models (Meta’s Llama and Al- ibaba’s Qwen) with a systematic sweep of vLLM inference engine hyperparameters. Each configuration was evaluated under synthetic load generated by GuideLLM, collecting high- resolution telemetry on throughput, Time to First Token, Time Per Output Token, KV cache utilization and infrastructure cost per million output tokens. The experimental pipeline was designed for reproducibility and generality, pairing an automated bench- mark orchestration script with a structured telemetry collection layer that produced a unified dataset for the subsequent analysis. Performance results were interpreted through Pareto efficiency frontiers mapping each metric against cost, enabling both hardware com- parison and optimal hyperparameter identification across different deployment objectives: throughput maximization, latency minimization and cost reduction. The GPU shows a modest cost and performance edge over the ASIC, which proves ideal in scenarios where workload predictability and resource stability outweigh the pursuit of peak efficiency. At the optimal frontier, the GPU achieves a 2.37x lower cost per million tokens. At their respective optimal configurations, this advantage extends to raw performance: 64% higher throughput, 96% lower Time to First Token and 56% lower inter-token latency than the Inferentia2 ASIC. The analysis also validated the impact of individual hyperparameters on inference performance and identified the optimal configurations for each specific use case, with the request concurrency level being the main driver for performance tuning along the optimal curves. This research was conducted in collaboration with Moviri, an IT consulting company specializing in performance engineering, analytics and cybersecu- rity.
Il deployment su larga scala di Large Language Models è costoso e tecnicamente com- plesso, rendendo l’ottimizzazione dell’inferenza una priorità per qualsiasi organizzazione che gestisca internamente carichi di lavoro AI. Questa tesi confronta due acceleratori hard- ware cloud — la GPU NVIDIA T4 (g4dn.xlarge) e l’ASIC AWS Inferentia2 (inf2.xlarge) — attraverso una griglia controllata di 198 configurazioni che combinano due modelli transformer open-source (Llama di Meta e Qwen di Alibaba) con una ricerca sistematica sugli iperparametri del motore di inferenza vLLM. Ogni configurazione è stata valutata sotto carico sintetico generato da GuideLLM, raccogliendo telemetria ad alta risoluzione su throughput, Time to First Token, Time Per Output Token, utilizzo della KV cache e costo infrastrutturale per milione di token generati. La pipeline sperimentale è stata pro- gettata per garantire riproducibilità e generalità, abbinando uno script di orchestrazione automatizzata dei benchmark a uno strato strutturato di raccolta telemetrica che ha prodotto un dataset unificato per l’analisi successiva. I risultati di performance sono stati interpretati attraverso frontiere di efficienza di Pareto che mettono in relazione ciascuna metrica con il costo, consentendo sia il confronto tra architetture che l’identificazione della configurazione ottimale per diversi obiettivi di deployment: massimizzazione del through- put, minimizzazione della latenza e riduzione dei costi. La GPU mostra un modesto vantaggio in termini di costo e performance rispetto all’ASIC, che risulta ideale negli sce- nari in cui la prevedibilità del carico di lavoro e la stabilità delle risorse hanno la priorità rispetto alla ricerca dell’efficienza massima. Sulla frontiera ottimale, la GPU raggiunge un costo per milione di token 2,37 volte inferiore. Nelle rispettive configurazioni ottimali, questo vantaggio si estende alle metriche di performance: throughput superiore del 64%, Time to First Token inferiore del 96% e latenza inter-token inferiore del 56% rispetto all’ASIC Inferentia2. L’analisi ha inoltre validato l’impatto dei singoli iperparametri sulle performance di inferenza e identificato le configurazioni ottimali per ciascun caso d’uso specifico, individuando nel livello di concorrenza delle richieste il principale contributore all’ottimizzazione delle prestazioni lungo le curve di ottimo. La ricerca è stata condotta in collaborazione con Moviri, società di consulenza IT specializzata in performance engi- neering, analytics e cybersecurity.
Benchmarking cloud instances for generative ai inference: the impact of accelerators and serving hyperparameters on cost-performance frontiers
ANESA, NICCOLÒ
2025/2026
Abstract
Deploying Large Language Models at scale is expensive and technically demanding, mak- ing inference optimization a critical priority for any organization self-hosting AI work- loads. This thesis benchmarks two cloud hardware accelerators — the NVIDIA T4 GPU (g4dn.xlarge) and the AWS Inferentia2 ASIC (inf2.xlarge) — across a controlled grid of 198 configurations combining two open-source transformer models (Meta’s Llama and Al- ibaba’s Qwen) with a systematic sweep of vLLM inference engine hyperparameters. Each configuration was evaluated under synthetic load generated by GuideLLM, collecting high- resolution telemetry on throughput, Time to First Token, Time Per Output Token, KV cache utilization and infrastructure cost per million output tokens. The experimental pipeline was designed for reproducibility and generality, pairing an automated bench- mark orchestration script with a structured telemetry collection layer that produced a unified dataset for the subsequent analysis. Performance results were interpreted through Pareto efficiency frontiers mapping each metric against cost, enabling both hardware com- parison and optimal hyperparameter identification across different deployment objectives: throughput maximization, latency minimization and cost reduction. The GPU shows a modest cost and performance edge over the ASIC, which proves ideal in scenarios where workload predictability and resource stability outweigh the pursuit of peak efficiency. At the optimal frontier, the GPU achieves a 2.37x lower cost per million tokens. At their respective optimal configurations, this advantage extends to raw performance: 64% higher throughput, 96% lower Time to First Token and 56% lower inter-token latency than the Inferentia2 ASIC. The analysis also validated the impact of individual hyperparameters on inference performance and identified the optimal configurations for each specific use case, with the request concurrency level being the main driver for performance tuning along the optimal curves. This research was conducted in collaboration with Moviri, an IT consulting company specializing in performance engineering, analytics and cybersecu- rity.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Anesa.pdf
accessibile in internet per tutti
Descrizione: testo della tesi
Dimensione
1.9 MB
Formato
Adobe PDF
|
1.9 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/261524