For the past few years, language models have improved mainly by getting bigger, with each generation built with more parameters, more data, and more computation, and served from large data centers. This thesis follows the opposite direction, asking how little capability is lost as a model is made small enough to run on an edge device such as a wearable, a sensor, or a battery-powered board. At this scale memory and energy matter more than raw capability, and the attention mechanism that made the Transformer successful becomes ill-suited to the small accelerators found at the edge. The work therefore adopts Mamba, a state-space architecture whose computational cost grows sub-quadratically with the length of the input. Because a model with few parameters cannot hold broad world knowledge, it has to specialize on a narrow domain instead. To measure depth, four dedicated question-answering benchmarks are built, covering two narrow domains, medicine and finance, each cast into two task formats: a yes/no benchmark and a multiple-choice one. The existing small models BabyLlama, MobileLLM, and a 130M Mamba are adapted to these benchmarks and compared with BabyMamba, a Mamba model of 50 million parameters trained from scratch for this work through pretraining and offline distillation. The results are encouraging: the architecture can be pushed well below the smallest size of the original Mamba study and still be specialized on a narrow domain, remaining competitive with a model of equal size built on a different architecture.
Negli ultimi anni i modelli linguistici sono migliorati principalmente diventando più grandi; ogni generazione è costruita con più parametri, più dati, più potenza di calcolo, ed è eseguita in grandi data center. Questa tesi segue la direzione opposta, chiedendosi quali capacità sopravvivono quando un modello viene ridotto fino a renderlo abbastanza piccolo da poter essere eseguito su un edge device, come un dispositivo indossabile, un sensore o una scheda embedded alimentata a batteria. A questa scala la memoria e l'energia contano più delle capacità pure, e il meccanismo di attention che ha decretato il successo del Transformer si rivela poco adatto ai piccoli acceleratori. Il lavoro adotta quindi Mamba, un'architettura state-space il cui costo computazionale cresce in modo sub-quadratico rispetto alla lunghezza dell'input. Poiché un modello con pochi parametri non può racchiudere un'ampia conoscenza del mondo, deve quindi specializzarsi su un dominio ristretto. Per misurare quanto a fondo un modello riesca a specializzarsi vengono costruiti quattro benchmark di domande e risposte, che coprono due domini ristretti, la medicina e la finanza, ciascuno articolato in due task: un benchmark a risposta sì/no e uno a scelta multipla. I modelli piccoli già esistenti BabyLlama, MobileLLM e Mamba da 130M vengono adattati a questi benchmark e confrontati con BabyMamba, un modello Mamba da 50 milioni di parametri addestrato da zero per questo lavoro tramite pretraining e distillazione. I risultati sono incoraggianti: l'architettura può essere ridotta ben al di sotto della dimensione più piccola dello studio originale su Mamba e può comunque essere specializzata su un dominio ristretto, tenendo testa a un modello di pari dimensione costruito su un'architettura diversa.
Beyond attention: efficient small language models for domain-specific tasks
CECCARELLI, VALERIO
2025/2026
Abstract
For the past few years, language models have improved mainly by getting bigger, with each generation built with more parameters, more data, and more computation, and served from large data centers. This thesis follows the opposite direction, asking how little capability is lost as a model is made small enough to run on an edge device such as a wearable, a sensor, or a battery-powered board. At this scale memory and energy matter more than raw capability, and the attention mechanism that made the Transformer successful becomes ill-suited to the small accelerators found at the edge. The work therefore adopts Mamba, a state-space architecture whose computational cost grows sub-quadratically with the length of the input. Because a model with few parameters cannot hold broad world knowledge, it has to specialize on a narrow domain instead. To measure depth, four dedicated question-answering benchmarks are built, covering two narrow domains, medicine and finance, each cast into two task formats: a yes/no benchmark and a multiple-choice one. The existing small models BabyLlama, MobileLLM, and a 130M Mamba are adapted to these benchmarks and compared with BabyMamba, a Mamba model of 50 million parameters trained from scratch for this work through pretraining and offline distillation. The results are encouraging: the architecture can be pushed well below the smallest size of the original Mamba study and still be specialized on a narrow domain, remaining competitive with a model of equal size built on a different architecture.| File | Dimensione | Formato | |
|---|---|---|---|
|
Valerio_Ceccarelli_Master_Thesis___Small_Language_Models.pdf
accessibile in internet per tutti
Descrizione: testo tesi
Dimensione
1.73 MB
Formato
Adobe PDF
|
1.73 MB | Adobe PDF | Visualizza/Apri |
|
Executive_Summary_Valerio_Ceccarelli.pdf
accessibile in internet per tutti
Descrizione: testo executive summary
Dimensione
482.63 kB
Formato
Adobe PDF
|
482.63 kB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260379