Parameter-Efficient Fine-Tuning (PEFT), and in particular Low-Rank Adaptation (LoRA), has made it cheap to specialize a Large Language Model (LLM) for a single task, and open-source specialists are now widely available. Serving many tasks from one model calls for a way to combine these specialists, and model merging offers a training-free route to do so. Most state-of-the-art merging methods, however, were designed for fully fine-tuned (FFT) models and transfer poorly to LoRA adapters. The obstacle is task interference, which on LoRA is dominated by a large and uneven variance in the magnitude of the task updates that a single global merge cannot reconcile. This thesis presents DMS (Deconflicting Magnitude Steering), a data-light framework that resolves this interference per module rather than with one global coefficient. Its guiding observation is that a single scale cannot serve tasks whose updates pull in opposite directions: adding magnitude everywhere helps some tasks and harms others, so what matters is where magnitude is added, not how much. DMS acts in three stages: an analytical profiling stage that locates where the specialists conflict, a multi-objective evolutionary search over per-module merge policies, and a parameter-free magnitude reallocation. On an eight-task commonsense-reasoning and question-answering suite with the Llama-3.2-1B backbone, evaluated zero-shot against the strongest Model-Merging and LoRA-Merging baselines, DMS reaches a mean accuracy of 65.75%, +10.9 points over the un-adapted backbone and +2.1 over a linear merge. This is parity with the best SVD-based methods (TSV-Merge 65.64%, KnOTS 65.53%), reached with no per-tensor SVD and an explicit worst-task objective those methods do not optimize; the per-module operator alone recovers the lowest-norm task, boolq, by +16 points. A conflict analysis traces the interference to architecture rather than semantics: the residual directional conflict sits in the early attention layers, while the magnitude imbalance DMS corrects concentrates in the deep feed-forward projections.
Il Parameter-Efficient Fine-Tuning (PEFT), e in particolare la Low-Rank Adaptation (LoRA), ha reso economico specializzare un Large Language Model (LLM) su un singolo task, e oggi gli specialisti open-source sono ampiamente disponibili. Servire molti task da un unico modello richiede un modo per combinare questi specialisti, e la fusione di modelli offre una via priva di addestramento per farlo. La maggior parte dei metodi di fusione allo stato dell'arte, tuttavia, è stata progettata per modelli con fine-tuning completo (FFT) e rende male sugli adattatori LoRA. L'ostacolo è l'interferenza tra task, che su LoRA è dominata da un'ampia e irregolare varianza nella magnitudine degli aggiornamenti, che una singola fusione globale non riesce a conciliare. Questa tesi presenta DMS (Deconflicting Magnitude Steering), un framework parco di dati che risolve questa interferenza per modulo anziché con un unico coefficiente globale. L'osservazione guida è che una singola scala non può servire task i cui aggiornamenti tirano in direzioni opposte: aggiungere magnitudine ovunque aiuta alcuni task e ne danneggia altri, per cui conta dove collocare la magnitudine, non quanta aggiungerne. DMS agisce in tre fasi: una fase analitica di profilazione che individua dove gli specialisti confliggono, una ricerca evolutiva multi-obiettivo sulle politiche di fusione per-modulo, e una riallocazione della magnitudine senza parametri. Su una suite di otto task di ragionamento di senso comune e question answering con backbone Llama-3.2-1B, valutata in zero-shot contro i più forti baseline di Model Merging e LoRA Merging, DMS raggiunge un'accuratezza media del 65,75%, +10,9 punti sul backbone non adattato e +2,1 su una fusione lineare. È la parità con i migliori metodi basati su SVD (TSV-Merge 65,64%, KnOTS 65,53%), ottenuta senza alcuna SVD per-tensore e con un esplicito obiettivo sul task peggiore che tali metodi non ottimizzano; il solo operatore per-modulo recupera il task a norma più bassa, boolq, di +16 punti. Un'analisi del conflitto riconduce l'interferenza all'architettura più che alla semantica: il conflitto direzionale residuo si colloca nei primi blocchi di attenzione, mentre lo squilibrio di magnitudine che DMS corregge si concentra nelle proiezioni feed-forward profonde.
DMS: LoRA multi-task merging via per-module magnitude reallocation
BOSSI, NICOLÓ
2025/2026
Abstract
Parameter-Efficient Fine-Tuning (PEFT), and in particular Low-Rank Adaptation (LoRA), has made it cheap to specialize a Large Language Model (LLM) for a single task, and open-source specialists are now widely available. Serving many tasks from one model calls for a way to combine these specialists, and model merging offers a training-free route to do so. Most state-of-the-art merging methods, however, were designed for fully fine-tuned (FFT) models and transfer poorly to LoRA adapters. The obstacle is task interference, which on LoRA is dominated by a large and uneven variance in the magnitude of the task updates that a single global merge cannot reconcile. This thesis presents DMS (Deconflicting Magnitude Steering), a data-light framework that resolves this interference per module rather than with one global coefficient. Its guiding observation is that a single scale cannot serve tasks whose updates pull in opposite directions: adding magnitude everywhere helps some tasks and harms others, so what matters is where magnitude is added, not how much. DMS acts in three stages: an analytical profiling stage that locates where the specialists conflict, a multi-objective evolutionary search over per-module merge policies, and a parameter-free magnitude reallocation. On an eight-task commonsense-reasoning and question-answering suite with the Llama-3.2-1B backbone, evaluated zero-shot against the strongest Model-Merging and LoRA-Merging baselines, DMS reaches a mean accuracy of 65.75%, +10.9 points over the un-adapted backbone and +2.1 over a linear merge. This is parity with the best SVD-based methods (TSV-Merge 65.64%, KnOTS 65.53%), reached with no per-tensor SVD and an explicit worst-task objective those methods do not optimize; the per-module operator alone recovers the lowest-norm task, boolq, by +16 points. A conflict analysis traces the interference to architecture rather than semantics: the residual directional conflict sits in the early attention layers, while the magnitude imbalance DMS corrects concentrates in the deep feed-forward projections.| File | Dimensione | Formato | |
|---|---|---|---|
|
01_Bossi_Nicolo_DMS_Thesis.pdf
accessibile in internet per tutti
Descrizione: Thesis
Dimensione
4.69 MB
Formato
Adobe PDF
|
4.69 MB | Adobe PDF | Visualizza/Apri |
|
02_Bossi_Nicolo_DMS_Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Executive Summary
Dimensione
1.51 MB
Formato
Adobe PDF
|
1.51 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260744