Imagined speech decoding from EEG is a promising route toward naturalistic brain--computer interfaces, while traditional closed-set classifiers fail to exploit the inherent topological structure of language, limiting their scalability to flexible communication. We propose a multimodal contrastive learning framework that aligns imagined speech EEG trials with language representations by mapping EEG into the same embedding space and dimensionality as a frozen pre-trained text encoder. An EEG encoder is trained with an InfoNCE-style objective that pulls each EEG representation toward its corresponding sentence embedding and pushes it away from mismatched sentences, encouraging semantically structured representations rather than purely class-specific decision boundaries. We evaluate the approach on the Chisco imagined speech dataset with 39 semantic categories across multiple subjects, reporting Top-k accuracy alongside semantic and statistical validation metrics. In addition to the established balanced pairwise matching accuracy (2v2), we introduce a permutation-tested Semantic Probability Alignment (SPA) score and an Average Semantic Quality (ASQ) measure for selective decoding. ASQ quantifies the semantic coherence of retrieved candidates in the frozen text space, and provides a controllable accuracy–coverage trade-off via abstention on low-quality trials. Results indicate non-trivial alignment between EEG and language embeddings within the fixed 39-intent Chisco setting, and suggest that semantic alignment can support meaning-aware retrieval-style decoding and reliability-oriented operation under constrained vocabularies.
La decodifica dell’imagined speech a partire dai segnali EEG rappresenta una via promettente verso lo sviluppo di interfacce cervello-computer più naturalistiche. Tuttavia, i tradizionali classificatori a set chiuso non riescono a sfruttare la struttura topologica del linguaggio, limitando la loro scalabilità verso una comunicazione flessibile. In questo lavoro proponiamo un framework di contrastive learning multimodale che allinea i trial EEG di imagined speech con le rappresentazioni linguistiche, mappando l’EEG nello stesso spazio di embedding e con la stessa dimensionalità di un encoder testuale pre-addestrato e mantenuto congelato (frozen). Un encoder EEG viene addestrato con un obiettivo in stile InfoNCE che avvicina ogni rappresentazione EEG al corrispondente embedding della frase e l’allontana dalle frasi non corrispondenti, incoraggiando rappresentazioni semanticamente strutturate piuttosto che semplici confini decisionali specifici per classe. Valutiamo l’approccio sul dataset di imagined speech Chisco, con 39 categorie semantiche e più soggetti, riportando l’accuratezza Top-k insieme a metriche di validazione semantica e statistica. Oltre alla consolidata accuratezza di matching bilanciata a coppie (2v2), introduciamo un punteggio di Semantic Probability Alignment (SPA) validato tramite test di permutazione e una misura di Average Semantic Quality (ASQ) per la decodifica selettiva. ASQ quantifica la coerenza semantica dei candidati recuperati nello spazio testuale frozen e fornisce un compromesso controllabile tra accuratezza e copertura tramite l’astensione sui trial di bassa qualità. I risultati indicano un allineamento non banale tra EEG ed embedding linguistici nel setting con 39 intenti di Chisco, suggerendo che l’allineamento semantico possa supportare una decodifica in stile retrieval sensibile al significato e un funzionamento orientato all’affidabilità in contesti con vocabolari vincolati.
Semantic alignment of EEG and text: a contrastive learning framework for decoding imagined speech
NAMDARI GHAREGHANI, SHAHRYAR
2025/2026
Abstract
Imagined speech decoding from EEG is a promising route toward naturalistic brain--computer interfaces, while traditional closed-set classifiers fail to exploit the inherent topological structure of language, limiting their scalability to flexible communication. We propose a multimodal contrastive learning framework that aligns imagined speech EEG trials with language representations by mapping EEG into the same embedding space and dimensionality as a frozen pre-trained text encoder. An EEG encoder is trained with an InfoNCE-style objective that pulls each EEG representation toward its corresponding sentence embedding and pushes it away from mismatched sentences, encouraging semantically structured representations rather than purely class-specific decision boundaries. We evaluate the approach on the Chisco imagined speech dataset with 39 semantic categories across multiple subjects, reporting Top-k accuracy alongside semantic and statistical validation metrics. In addition to the established balanced pairwise matching accuracy (2v2), we introduce a permutation-tested Semantic Probability Alignment (SPA) score and an Average Semantic Quality (ASQ) measure for selective decoding. ASQ quantifies the semantic coherence of retrieved candidates in the frozen text space, and provides a controllable accuracy–coverage trade-off via abstention on low-quality trials. Results indicate non-trivial alignment between EEG and language embeddings within the fixed 39-intent Chisco setting, and suggest that semantic alignment can support meaning-aware retrieval-style decoding and reliability-oriented operation under constrained vocabularies.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_03_Namdari_Executive Summary_02..pdf
accessibile in internet per tutti a partire dal 27/02/2027
Descrizione: Executive Summary
Dimensione
5.59 MB
Formato
Adobe PDF
|
5.59 MB | Adobe PDF | Visualizza/Apri |
|
2026_03_Namdari_Thesis_01.pdf
accessibile in internet per tutti a partire dal 27/02/2027
Descrizione: Main File
Dimensione
9.47 MB
Formato
Adobe PDF
|
9.47 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/253201