Human behaviour recognition in real-world educational and care settings requires models that are not only technically effective, but also privacy-preserving, interpretable, and compatible with the observation of vulnerable populations. This thesis investigates whether pose-based representations can support the automatic recognition of observable body behaviours in a daycare centre attended by adults with intellectual and developmental disabilities. First, we design an end-to-end pipeline that transforms multi-view video recordings into anonymous skeletal representations through camera calibration, video undistortion, 2D pose extraction, anonymous tracking, track stitching, annotation alignment, and dataset construction. We then define an ELAN-based annotation protocol focused on observable individual and relational behaviours, explicitly avoiding facial recognition, voice identification, emotion recognition, and biometric profiling. Next, we evaluate several learning-based approaches, including SVM baselines on 2D tabular features, a multi-view 3D reconstruction branch, PoseC3D experiments on skeleton-based representations, and exploratory relational and annotation-aware models. The results show that pose-derived features contain a learnable behavioural signal, but also reveal strong limitations due to class imbalance, tracking fragmentation, incomplete visibility, and session-dependent generalization. Overall, the thesis provides a privacy-preserving methodological framework for studying body behaviour recognition in naturalistic inclusive settings, highlighting both the potential of skeletal representations and the need for cautious, human-centred interpretation.
Il riconoscimento automatico del comportamento umano in contesti educativi e di cura reali richiede modelli non solo tecnicamente efficaci, ma anche privacy-preserving, interpretabili e compatibili con l’osservazione di popolazioni vulnerabili. Questa tesi studia se rappresentazioni basate su pose scheletriche possano supportare il riconoscimento automatico di comportamenti corporei osservabili in un centro diurno frequentato da adulti con disabilità intellettiva e dello sviluppo. In primo luogo, viene progettata una pipeline end-to-end che trasforma registrazioni video multi-view in rappresentazioni scheletriche anonime attraverso calibrazione delle camere, undistortion, estrazione di pose 2D, tracking anonimo, stitching delle tracce, allineamento con le annotazioni e costruzione dei dataset. Successivamente, viene definito un protocollo di annotazione in ELAN centrato su comportamenti individuali e relazionali osservabili, evitando esplicitamente riconoscimento facciale, identificazione vocale, riconoscimento delle emozioni e profilazione biometrica. Vengono poi valutati diversi approcci di apprendimento automatico, tra cui baseline SVM su feature tabellari 2D, un ramo di ricostruzione 3D multi-view, esperimenti PoseC3D su rappresentazioni skeleton-based e modelli esplorativi relazionali e annotation-aware. I risultati mostrano che le feature derivate dalle pose contengono un segnale comportamentale apprendibile, ma evidenziano anche limiti significativi dovuti a sbilanciamento delle classi, frammentazione del tracking, visibilità incompleta e generalizzazione dipendente dalla sessione. Nel complesso, la tesi propone un framework metodologico privacy-preserving per studiare il riconoscimento del comportamento corporeo in contesti inclusivi naturalistici, evidenziando sia il potenziale delle rappresentazioni scheletriche sia la necessità di un’interpretazione cauta e human-centred.
A multi-view skeleton-based pipeline for body behaviour recognition in inclusive care settings
MASTROBERARDINO, SIMONA
2025/2026
Abstract
Human behaviour recognition in real-world educational and care settings requires models that are not only technically effective, but also privacy-preserving, interpretable, and compatible with the observation of vulnerable populations. This thesis investigates whether pose-based representations can support the automatic recognition of observable body behaviours in a daycare centre attended by adults with intellectual and developmental disabilities. First, we design an end-to-end pipeline that transforms multi-view video recordings into anonymous skeletal representations through camera calibration, video undistortion, 2D pose extraction, anonymous tracking, track stitching, annotation alignment, and dataset construction. We then define an ELAN-based annotation protocol focused on observable individual and relational behaviours, explicitly avoiding facial recognition, voice identification, emotion recognition, and biometric profiling. Next, we evaluate several learning-based approaches, including SVM baselines on 2D tabular features, a multi-view 3D reconstruction branch, PoseC3D experiments on skeleton-based representations, and exploratory relational and annotation-aware models. The results show that pose-derived features contain a learnable behavioural signal, but also reveal strong limitations due to class imbalance, tracking fragmentation, incomplete visibility, and session-dependent generalization. Overall, the thesis provides a privacy-preserving methodological framework for studying body behaviour recognition in naturalistic inclusive settings, highlighting both the potential of skeletal representations and the need for cautious, human-centred interpretation.| File | Dimensione | Formato | |
|---|---|---|---|
|
Tesi_Magistrale___Simona_Mastroberardino.pdf
accessibile in internet per tutti
Descrizione: Master’s thesis on privacy-preserving recognition of bodily behaviors in adults with intellectual disabilities, based on anonymous 2D/3D poses extracted from multi-camera video. The work includes data acquisition and processing, tracking, annotation with educators, construction of temporal datasets, and evaluation of SVM and PoseC3D models.
Dimensione
4.48 MB
Formato
Adobe PDF
|
4.48 MB | Adobe PDF | Visualizza/Apri |
|
Executive_Summary___Simona_Mastroberardino.pdf
accessibile in internet per tutti
Descrizione: This Executive Summary outlines the main contributions of the thesis on privacy-preserving behaviour recognition in adults with intellectual and developmental disabilities. It summarizes the multi-view pose-processing pipeline, the annotation and dataset construction process, and the evaluation of 2D, 3D, and skeleton-based approaches, highlighting performance, generalization, and practical limitations.
Dimensione
1.31 MB
Formato
Adobe PDF
|
1.31 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260591