Data intensive architectures integrate multiple specialized systems to acquire, process, and deliver data at scale, but designing and deploying them remains a largely manual process requiring deep technical expertise. Dragoni and Margara propose a semi-automated methodology that translates an application scenario into an architecture description expressed in an Architecture Description Language (ADL), but the ADL remains a conceptual model and the methodology stops short of producing a deployable system, a gap that the authors identify as future work. This thesis addresses that gap along two directions. First, it introduces a formal grammar for the ADL that makes architecture descriptions machine-readable, enabling automatic parsing, semantic validation and translation to an extensible intermediate representation. Second, given an architecture description enriched with concrete systems assignments for its elements, it explores whether an infrastructure-from-code approach can automate the generation and deployment of the infrastructure required to run the described system. Nitric, an infrastructure-from-code framework, is adopted as the target platform and extended to better represent data-intensive processing pipelines. Three fundamentally different execution models are tested on top of this extended framework: Spark Structured Streaming for continuous processing, Spark Batch for scheduled bounded jobs and PostgreSQL Materialized Views for lightweight declarative recomputation. The grammar is shown to express the reference architectures from the literature, and the resulting infrastructure generation layer reduces deployment code while remaining portable across execution models, suggesting that infrastructure-from-code tools (if suitably extended) are a viable foundation for automating the deployment of data-intensive architectures.
Le architetture data-intensive integrano molteplici sistemi specializzati per acquisire, elaborare e distribuire dati su larga scala, ma la loro progettazione e il loro deployment rimangono in gran parte manuali e richiedono una profonda competenza tecnica in diversi domini. Dragoni e Margara propongono una metodologia semi-automatizzata che traduce uno scenario applicativo in una descrizione architetturale espressa tramite un Architecture Description Language (ADL), ma l'ADL proposto rimane un modello concettuale e la metodologia si ferma prima di produrre un sistema pronto all'uso, un limite che gli autori stessi indicano come direzione per lavori futuri. Questa tesi affronta questo limite lungo due direzioni. La prima introduce una grammatica formale per l'ADL che rende le descrizioni architetturali leggibili da una macchina, abilitando il parsing automatico, la validazione semantica e la traduzione verso una rappresentazione intermedia estendibile. La seconda, data una descrizione architetturale arricchita con assegnazioni concrete dei sistemi ai suoi elementi, esplora se un approccio Infrastructure-from-Code (IfC) possa automatizzare la generazione e il deployment dell'infrastruttura necessaria per eseguire il sistema descritto. Nitric, un framework IfC, viene adottato ed esteso per rappresentare al meglio le architetture data-intensive. Tre modelli di esecuzione diversi sono testati su questo framework esteso: Spark Structured Streaming per l'elaborazione continua, Spark Batch per job periodici, e PostgreSQL Materialized Views per rielaborazioni leggere. La grammatica si dimostra in grado di esprimere le architetture di riferimento della letteratura e l'approccio IfC riduce il codice infrastrutturale mantenendo la portabilità tra i modelli di esecuzione, suggerendo che gli strumenti IfC, se opportunamente estesi, rappresentano una base concretamente utilizzabile per automatizzare il deployment di architetture data-intensive.
Towards a framework for semi-automated deployment of data-intensive architectures
MARTELLOSIO, FRANCESCO
2025/2026
Abstract
Data intensive architectures integrate multiple specialized systems to acquire, process, and deliver data at scale, but designing and deploying them remains a largely manual process requiring deep technical expertise. Dragoni and Margara propose a semi-automated methodology that translates an application scenario into an architecture description expressed in an Architecture Description Language (ADL), but the ADL remains a conceptual model and the methodology stops short of producing a deployable system, a gap that the authors identify as future work. This thesis addresses that gap along two directions. First, it introduces a formal grammar for the ADL that makes architecture descriptions machine-readable, enabling automatic parsing, semantic validation and translation to an extensible intermediate representation. Second, given an architecture description enriched with concrete systems assignments for its elements, it explores whether an infrastructure-from-code approach can automate the generation and deployment of the infrastructure required to run the described system. Nitric, an infrastructure-from-code framework, is adopted as the target platform and extended to better represent data-intensive processing pipelines. Three fundamentally different execution models are tested on top of this extended framework: Spark Structured Streaming for continuous processing, Spark Batch for scheduled bounded jobs and PostgreSQL Materialized Views for lightweight declarative recomputation. The grammar is shown to express the reference architectures from the literature, and the resulting infrastructure generation layer reduces deployment code while remaining portable across execution models, suggesting that infrastructure-from-code tools (if suitably extended) are a viable foundation for automating the deployment of data-intensive architectures.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Martellosio_executive_summary.pdf
accessibile in internet per tutti
Dimensione
363.74 kB
Formato
Adobe PDF
|
363.74 kB | Adobe PDF | Visualizza/Apri |
|
2026_07_Martellosio_thesis.pdf
accessibile in internet per tutti
Dimensione
1.15 MB
Formato
Adobe PDF
|
1.15 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260750