今日论文合集:cs.SD语音8篇,eess.AS音频处理5篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】823-OLT @ BUET DL Sprint 4.0: Context-Aware Windowing for ASR and Fine-Tuned Speaker Diarization in Bengali Long Form Audio
标题:823-OLLT @ BUET DL Sprint 4.0:在孟加拉语长格式音频中针对SVR和微调扬声器拨号的上下文感知窗口
链接:https://arxiv.org/abs/2602.21183

作者:Ratnajit Dhar,Arpita Mallik
摘要:孟加拉语尽管是全球使用最广泛的语言之一,但在长格式语音技术中仍然代表性不足,特别是在解决转录和说话者归属的系统中。我们提出了一个框架,长期形式的孟加拉语语音智能,解决自动语音识别使用耳语媒体为基础的模型和扬声器diarization使用微调分割模型。ASR流水线结合了语音分离、语音活动检测和间隙感知窗口策略来构建上下文保留段以进行稳定解码。对于日志化,预先训练的说话者分割模型在官方比赛数据集上进行微调(作为BUET CSE Fest组织的DL Sprint 4.0比赛的一部分提供),以更好地捕捉孟加拉语会话模式。由此产生的系统提供了长格式音频的高效转录和说话人感知转录,为低资源语言提供可扩展的语音技术解决方案。
摘要:Bengali, despite being one of the most widely spoken languages globally, remains underrepresented in long form speech technology, particularly in systems addressing transcription and speaker attribution. We present frameworks for long form Bengali speech intelligence that address automatic speech recognition using a Whisper Medium based model and speaker diarization using a finetuned segmentation model. The ASR pipeline incorporates vocal separation, voice activity detection, and a gap aware windowing strategy to construct context preserving segments for stable decoding. For diarization, a pretrained speaker segmentation model is finetuned on the official competition dataset (provided as part of the DL Sprint 4.0 competition organized under BUET CSE Fest), to better capture Bengali conversational patterns. The resulting systems deliver both efficient transcription of long form audio and speaker aware transcription to provide scalable speech technology solutions for low resource languages.


【2】Geometric Analysis of Speech Representation Spaces: Topological Disentanglement and Confound Detection
标题:语音表示空间的几何分析:布局解纠缠和混淆检测
链接:https://arxiv.org/abs/2602.20823

作者:Bipasha Kashyap,Pubudu N. Pathirana
备注:Submitted to INTERSPEECH 2026
摘要:基于语音的临床工具越来越多地部署在多语言环境中,但病理性语音标记是否仍然与口音变化在几何上分离仍不清楚。系统可能会错误分类健康的非母语人士或错过多语言患者的病理。我们提出了一个四指标聚类框架,以评估几何解纠缠的情感,语言和病理语音功能在六个语料库和八个数据集的组合。出现了一致的等级制度:情感特征形成最紧密的聚类(轮廓0.250),其次是病理(0.141)和语言(0.077)。混淆分析显示,病理-语言重叠仍低于0.21,高于零排列,但对于临床部署有界。可信性分析证实了嵌入保真度和鲁棒性的几何结论。我们的框架为不同人群的公平和可靠的言语健康系统提供了可操作的指导方针。
摘要:Speech-based clinical tools are increasingly deployed in multilingual settings, yet whether pathological speech markers remain geometrically separable from accent variation remains unclear. Systems may misclassify healthy non-native speakers or miss pathology in multilingual patients. We propose a four-metric clustering framework to evaluate geometric disentanglement of emotional, linguistic, and pathological speech features across six corpora and eight dataset combinations. A consistent hierarchy emerges: emotional features form the tightest clusters (Silhouette 0.250), followed by pathological (0.141) and linguistic (0.077). Confound analysis shows pathological-linguistic overlap remains below 0.21, which is above the permutation null but bounded for clinical deployment. Trustworthiness analysis confirms embedding fidelity and robustness of the geometric conclusions. Our framework provides actionable guidelines for equitable and reliable speech health systems across diverse populations.


【3】Assessing the Impact of Speaker Identity in Speech Spoofing Detection
标题:评估说话者身份在语音欺骗检测中的影响
链接:https://arxiv.org/abs/2602.20805

作者:Anh-Tuan Dao,Driss Matrouf,Nicholas Evans
摘要:欺骗检测系统通常使用来自多个说话者的不同录音进行训练,通常假设所得到的嵌入与说话者身份无关。然而,这一假设仍未得到证实。在本文中,我们研究了说话人信息对欺骗检测系统的影响。我们在Speaker-Invariant Multi-Task框架中提出了两种方法,一种是在嵌入中对说话人身份进行建模,另一种是将其删除。SInMT集成了多任务学习,用于联合说话人识别和欺骗检测,并结合了梯度反转层。使用四个数据集进行评估,与基线相比,我们的说话者不变模型将平均等错误率降低了17%,对于最具挑战性的攻击(例如,A11)。
摘要:Spoofing detection systems are typically trained using diverse recordings from multiple speakers, often assuming that the resulting embeddings are independent of speaker identity. However, this assumption remains unverified. In this paper, we investigate the impact of speaker information on spoofing detection systems. We propose two approaches within our Speaker-Invariant Multi-Task framework, one that models speaker identity within the embeddings and another that removes it. SInMT integrates multi-task learning for joint speaker recognition and spoofing detection, incorporating a gradient reversal layer. Evaluated using four datasets, our speaker-invariant model reduces the average equal error rate by 17% compared to the baseline, with up to 48% reduction for the most challenging attacks (e.g., A11).


【4】Voices of the Mountains: Deep Learning-Based Vocal Error Detection System for Kurdish Maqams
标题:山之声:基于深度学习的库尔德Maqams声音错误检测系统
链接:https://arxiv.org/abs/2602.20744

作者:Darvan Shvan Khairaldeen,Hossein Hassani
摘要:Maqam是一种歌唱类型,是库尔德音乐的重要组成部分。木卡姆歌手接受传统的面对面培训或通过自我培训。自动歌唱评估(ASA)使用机器学习(ML)来提供歌唱风格的准确性,并可以帮助学习者通过错误检测来提高他们的表现。目前,可用的ASA工具遵循西方音乐规则。音乐作品要求所有音符从开始到结束都保持在预期的音高范围内。该系统无法检测到微音程和音高弯曲,因此即使歌手根据传统规则表演,它也会将库尔德人的maqam演唱识别为不正确。库尔德木卡姆要求在微色调空间内识别性能错误,这超出了西方的平等气质。这项研究是解决上述差距的第一次尝试。虽然许多错误类型发生在唱歌,我们的重点是音高,节奏和模态的稳定性错误的背景下,巴亚提库尔德。我们收集了13位歌手的50首歌曲(2-3小时),并注释了221个错误跨度(150个细音高,46个节奏,25个模态漂移)。数据被分割成15,199个重叠窗口,并转换为对数梅尔光谱图。我们开发了一个具有注意力模式的双头CNN-BiLSTM,以确定窗口是否包含错误,并根据所选错误对其进行分类。经过20个时期的训练,在第10个时期提前停止,该模型达到了0.468的验证宏F1。在0.750阈值的完整50首歌曲评估中,召回率为39.4%,准确率为25.8%。在检测窗口内,宏类型F1为0.387,F1为0.492(细音高),0.536(节奏)和0.133(模态漂移);模态漂移回忆率为8.0%。对常见错误类型的更好性能表明该方法有效,而较差的模式漂移召回表明需要更多数据和平衡。
摘要:Maqam, a singing type, is a significant component of Kurdish music. A maqam singer receives training in a traditional face-to-face or through self-training. Automatic Singing Assessment (ASA) uses machine learning (ML) to provide the accuracy of singing styles and can help learners to improve their performance through error detection. Currently, the available ASA tools follow Western music rules. The musical composition requires all notes to stay within their expected pitch range from start to finish. The system fails to detect micro-intervals and pitch bends, so it identifies Kurdish maqam singing as incorrect even though the singer performs according to traditional rules. Kurdish maqam requires recognizing performance errors within microtonal spaces, which is beyond Western equal temperament. This research is the first attempt to address the mentioned gap. While many error types happen during singing, our focus is on pitch, rhythm, and modal stability errors in the context of Bayati-Kurd. We collected 50 songs from 13 vocalists ( 2-3 hours) and annotated 221 error spans (150 fine pitch, 46 rhythm, 25 modal drift). The data was segmented into 15,199 overlapping windows and converted to log-mel spectrograms. We developed a two-headed CNN-BiLSTM with attention mode to decide whether a window contains an error and to classify it based on the chosen errors. Trained for 20 epochs with early stopping at epoch 10, the model reached a validation macro-F1 of 0.468. On the full 50-song evaluation at a 0.750 threshold, recall was 39.4% and precision 25.8% . Within detected windows, type macro-F1 was 0.387, with F1 of 0.492 (fine pitch), 0.536 (rhythm), and 0.133 (modal drift); modal drift recall was 8.0%. The better performance on common error types shows that the method works, while the poor modal-drift recall shows that more data and balancing are needed.


【5】Quantifying Dimensional Independence in Speech: An Information-Theoretic Framework for Disentangled Representation Learning
标题:量化语音中的维度独立性:理清表示学习的信息理论框架
链接:https://arxiv.org/abs/2602.20592

作者:Bipasha Kashyap,Björn W. Schuller,Pubudu N. Pathirana
摘要:语音信号在共享的声学通道内编码情感、语言和病理信息;然而,解缠通常通过下游任务性能间接评估。我们引入了一个信息理论框架,通过将有界神经互信息(MI)估计与非参数验证相结合来量化手工制作的声学特征中的跨维度统计依赖。在六个语料库中,跨维度MI仍然很低,估计界限很紧($< 0.15$ nats),表明所考虑的数据中的统计耦合较弱,而源-过滤器MI显著较高(0.47 nats)。归因分析(定义为归因于源成分与过滤成分的总MI的比例)揭示了情感维度的源优势(80%)和语言和病理维度的过滤优势(分别为60%和58%)。这些发现提供了一个原则性的框架,量化语音的维度独立性。
摘要:Speech signals encode emotional, linguistic, and pathological information within a shared acoustic channel; however, disentanglement is typically assessed indirectly through downstream task performance. We introduce an information-theoretic framework to quantify cross-dimension statistical dependence in handcrafted acoustic features by integrating bounded neural mutual information (MI) estimation with non-parametric validation. Across six corpora, cross-dimension MI remains low, with tight estimation bounds ($< 0.15$ nats), indicating weak statistical coupling in the data considered, whereas Source--Filter MI is substantially higher (0.47 nats). Attribution analysis, defined as the proportion of total MI attributable to source versus filter components, reveals source dominance for emotional dimensions (80\%) and filter dominance for linguistic and pathological dimensions (60\% and 58\%, respectively). These findings provide a principled framework for quantifying dimensional independence in speech.


【6】Memory-guided Prototypical Co-occurrence Learning for Mixed Emotion Recognition
标题:用于混合情绪识别的记忆引导原型同现学习
链接:https://arxiv.org/abs/2602.20530

作者:Ming Li,Yong-Jin Liu,Fang Liu,Huankun Sheng,Yeying Fan,Yixiang Wei,Minnan Luo,Weizhan Zhang,Wenping Wang
摘要:从多模态生理和行为信号中识别情感在情感计算中起着关键作用,然而大多数现有模型仍然局限于在受控实验室环境中预测奇异情感。相比之下,现实世界中的人类情感体验的特点往往是同时存在多个情感状态,刺激了最近的兴趣,混合情感识别作为一个情感分布学习问题。然而,目前的研究方法往往忽视了共存情绪之间内在的效价一致性和结构相关性。为了解决这个问题,我们提出了一个记忆引导的原型共现学习(MPCL)框架,明确的情感共现模式模型。具体来说,我们首先通过多尺度联想记忆机制融合多模态信号。为了捕捉跨模态语义关系,我们构建情感特定的原型记忆库,产生丰富的生理和行为表示,并采用原型关系蒸馏,以确保跨模态对齐的潜在原型空间。此外,受人类认知记忆系统的启发,我们引入了一种记忆检索策略来提取跨情感类别的语义级共现关联。通过这种自下而上的分层抽象过程,我们的模型学习情感信息表示,以实现准确的情感分布预测。在两个公开数据集上的综合实验表明,MPCL在混合情感识别方面无论是定量还是定性都始终优于最先进的方法。
摘要:Emotion recognition from multi-modal physiological and behavioral signals plays a pivotal role in affective computing, yet most existing models remain constrained to the prediction of singular emotions in controlled laboratory settings. Real-world human emotional experiences, by contrast, are often characterized by the simultaneous presence of multiple affective states, spurring recent interest in mixed emotion recognition as an emotion distribution learning problem. Current approaches, however, often neglect the valence consistency and structured correlations inherent among coexisting emotions. To address this limitation, we propose a Memory-guided Prototypical Co-occurrence Learning (MPCL) framework that explicitly models emotion co-occurrence patterns. Specifically, we first fuse multi-modal signals via a multi-scale associative memory mechanism. To capture cross-modal semantic relationships, we construct emotion-specific prototype memory banks, yielding rich physiological and behavioral representations, and employ prototype relation distillation to ensure cross-modal alignment in the latent prototype space. Furthermore, inspired by human cognitive memory systems, we introduce a memory retrieval strategy to extract semantic-level co-occurrence associations across emotion categories. Through this bottom-up hierarchical abstraction process, our model learns affectively informative representations for accurate emotion distribution prediction. Comprehensive experiments on two public datasets demonstrate that MPCL consistently outperforms state-of-the-art methods in mixed emotion recognition, both quantitatively and qualitatively.


【7】Graph Modelling Analysis of Speech-Gesture Interaction for Aphasia Severity Estimation
标题:失语严重程度估计的言语与手势相互作用的图模型分析
链接:https://arxiv.org/abs/2602.20163

作者:Navya Martin Kollapally,Christa Akers,Renjith Nelson Joseph
备注:IJCAI
摘要:失语症是一种后天性语言障碍,由负责语言的大脑区域损伤引起。失语症可能会损害书面和口头语言的使用和理解。西方失语症成套测验修订版(WAB-R)是一种由言语语言病理学家(SLP)管理的评估工具,用于评估失语症的类型和严重程度。由于WAB-R测量的是孤立的语言技能,因此人们越来越关注对话语产出的评估,将其作为日常语言能力的一种更全面的表现。语音分析的最新进展集中在自动估计失语症的严重程度,从自发的讲话,主要依赖于孤立的语言或声学特征。在这项工作中,我们提出了一个基于图神经网络的框架来估计失语症的严重程度。我们将每个参与者的话语表示为有向多模态图,其中节点表示词汇项和手势,边缘编码单词-单词,手势-单词和单词-手势转换。GraphSAGE用于学习参与者级嵌入,从而整合来自直接邻居和整体图结构的信息。我们的研究结果表明,失语症的严重程度不是编码在孤立的词汇分布,而是出现在结构化的相互作用之间的讲话和手势。所提出的架构提供了一个可靠的自动化失语症评估,可能用于床边筛查和远程医疗为基础的监测。
摘要:Aphasia is an acquired language disorder caused by injury to the regions of the brain that are responsible for language. Aphasia may impair the use and comprehension of written and spoken language. The Western Aphasia Battery-Revised (WAB-R) is an assessment tool administered by speech-language pathologists (SLPs) to evaluate the aphasia type and severity. Because the WAB-R measures isolated linguistic skills, there has been growing interest in the assessment of discourse production as a more holistic representation of everyday language abilities. Recent advancements in speech analysis focus on automated estimation of aphasia severity from spontaneous speech, relying mostly in isolated linguistic or acoustical features. In this work, we propose a graph neural network-based framework for estimating aphasia severity. We represented each participant's discourse as a directed multi-modal graph, where nodes represent lexical items and gestures and edges encode word-word, gesture-word, and word-gesture transitions. GraphSAGE is employed to learn participant-level embeddings, thus integrating information from immediate neighbors and overall graph structure. Our results suggest that aphasia severity is not encoded in isolated lexical distribution, but rather emerges from structured interactions between speech and gesture. The proposed architecture offers a reliable automated aphasia assessment, with possible uses in bedside screening and telehealth-based monitoring.


【8】Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
标题:用于噪声ASR的免训练智能引导观察添加
链接:https://arxiv.org/abs/2602.20967

作者:Haoyang Li,Changsong Liu,Wei Rao,Hao Shi,Sakriani Sakti,Eng Siong Chng
摘要:自动语音识别(ASR)在噪声环境中性能严重下降。虽然语音增强(SE)前端有效地抑制了背景噪声,但它们通常会引入损害识别的伪影。观察添加(OA)通过融合有噪语音和SE增强语音来解决这个问题,在不修改SE或ASR模型参数的情况下提高识别率。本文提出了一种可懂度引导的OA方法,其中融合权重来自直接从后端ASR获得的可懂度估计。与之前基于训练神经预测器的OA方法不同,该方法无需训练,降低了复杂度并增强了泛化能力。在不同的SE-ASR组合和数据集上进行的广泛实验表明,与现有的OA基线相比,具有强大的鲁棒性和改进。对基于可理解性引导的切换的替代方案和框架与话语级OA的其他分析进一步验证了所提出的设计。
摘要:Automatic speech recognition (ASR) degrades severely in noisy environments. Although speech enhancement (SE) front-ends effectively suppress background noise, they often introduce artifacts that harm recognition. Observation addition (OA) addressed this issue by fusing noisy and SE enhanced speech, improving recognition without modifying the parameters of the SE or ASR models. This paper proposes an intelligibility-guided OA method, where fusion weights are derived from intelligibility estimates obtained directly from the backend ASR. Unlike prior OA methods based on trained neural predictors, the proposed method is training-free, reducing complexity and enhances generalization. Extensive experiments across diverse SE-ASR combinations and datasets demonstrate strong robustness and improvements over existing OA baselines. Additional analyses of intelligibility-guided switching-based alternatives and frame versus utterance-level OA further validate the proposed design.


eess.AS音频处理


【1】Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
标题:用于噪声ASR的免训练智能引导观察添加
链接:https://arxiv.org/abs/2602.20967

作者:Haoyang Li,Changsong Liu,Wei Rao,Hao Shi,Sakriani Sakti,Eng Siong Chng
摘要:自动语音识别(ASR)在噪声环境中性能严重下降。虽然语音增强(SE)前端有效地抑制了背景噪声,但它们通常会引入损害识别的伪影。观察添加(OA)通过融合有噪语音和SE增强语音来解决这个问题,在不修改SE或ASR模型参数的情况下提高识别率。本文提出了一种可懂度引导的OA方法,其中融合权重来自直接从后端ASR获得的可懂度估计。与之前基于训练神经预测器的OA方法不同,该方法无需训练,降低了复杂度并增强了泛化能力。在不同的SE-ASR组合和数据集上进行的广泛实验表明,与现有的OA基线相比,具有强大的鲁棒性和改进。对基于可理解性引导的切换的替代方案和框架与话语级OA的其他分析进一步验证了所提出的设计。
摘要:Automatic speech recognition (ASR) degrades severely in noisy environments. Although speech enhancement (SE) front-ends effectively suppress background noise, they often introduce artifacts that harm recognition. Observation addition (OA) addressed this issue by fusing noisy and SE enhanced speech, improving recognition without modifying the parameters of the SE or ASR models. This paper proposes an intelligibility-guided OA method, where fusion weights are derived from intelligibility estimates obtained directly from the backend ASR. Unlike prior OA methods based on trained neural predictors, the proposed method is training-free, reducing complexity and enhances generalization. Extensive experiments across diverse SE-ASR combinations and datasets demonstrate strong robustness and improvements over existing OA baselines. Additional analyses of intelligibility-guided switching-based alternatives and frame versus utterance-level OA further validate the proposed design.


【2】Geometric Analysis of Speech Representation Spaces: Topological Disentanglement and Confound Detection
标题:语音表示空间的几何分析:布局解纠缠和混淆检测
链接:https://arxiv.org/abs/2602.20823

作者:Bipasha Kashyap,Pubudu N. Pathirana
备注:Submitted to INTERSPEECH 2026
摘要:基于语音的临床工具越来越多地部署在多语言环境中,但病理性语音标记是否仍然与口音变化在几何上分离仍不清楚。系统可能会错误分类健康的非母语人士或错过多语言患者的病理。我们提出了一个四指标聚类框架,以评估几何解纠缠的情感,语言和病理语音功能在六个语料库和八个数据集的组合。出现了一致的等级制度:情感特征形成最紧密的聚类(轮廓0.250),其次是病理(0.141)和语言(0.077)。混淆分析显示,病理-语言重叠仍低于0.21,高于零排列,但对于临床部署有界。可信性分析证实了嵌入保真度和鲁棒性的几何结论。我们的框架为不同人群的公平和可靠的言语健康系统提供了可操作的指导方针。
摘要:Speech-based clinical tools are increasingly deployed in multilingual settings, yet whether pathological speech markers remain geometrically separable from accent variation remains unclear. Systems may misclassify healthy non-native speakers or miss pathology in multilingual patients. We propose a four-metric clustering framework to evaluate geometric disentanglement of emotional, linguistic, and pathological speech features across six corpora and eight dataset combinations. A consistent hierarchy emerges: emotional features form the tightest clusters (Silhouette 0.250), followed by pathological (0.141) and linguistic (0.077). Confound analysis shows pathological-linguistic overlap remains below 0.21, which is above the permutation null but bounded for clinical deployment. Trustworthiness analysis confirms embedding fidelity and robustness of the geometric conclusions. Our framework provides actionable guidelines for equitable and reliable speech health systems across diverse populations.


【3】Quantifying Dimensional Independence in Speech: An Information-Theoretic Framework for Disentangled Representation Learning
标题:量化语音中的维度独立性:理清表示学习的信息理论框架
链接:https://arxiv.org/abs/2602.20592

作者:Bipasha Kashyap,Björn W. Schuller,Pubudu N. Pathirana
摘要:语音信号在共享的声学通道内编码情感、语言和病理信息;然而,解缠通常通过下游任务性能间接评估。我们引入了一个信息理论框架,通过将有界神经互信息(MI)估计与非参数验证相结合来量化手工制作的声学特征中的跨维度统计依赖。在六个语料库中,跨维度MI仍然很低,估计界限很紧($< 0.15$ nats),表明所考虑的数据中的统计耦合较弱,而源-过滤器MI显著较高(0.47 nats)。归因分析(定义为归因于源成分与过滤成分的总MI的比例)揭示了情感维度的源优势(80%)和语言和病理维度的过滤优势(分别为60%和58%)。这些发现提供了一个原则性的框架,量化语音的维度独立性。
摘要:Speech signals encode emotional, linguistic, and pathological information within a shared acoustic channel; however, disentanglement is typically assessed indirectly through downstream task performance. We introduce an information-theoretic framework to quantify cross-dimension statistical dependence in handcrafted acoustic features by integrating bounded neural mutual information (MI) estimation with non-parametric validation. Across six corpora, cross-dimension MI remains low, with tight estimation bounds ($< 0.15$ nats), indicating weak statistical coupling in the data considered, whereas Source--Filter MI is substantially higher (0.47 nats). Attribution analysis, defined as the proportion of total MI attributable to source versus filter components, reveals source dominance for emotional dimensions (80\%) and filter dominance for linguistic and pathological dimensions (60\% and 58\%, respectively). These findings provide a principled framework for quantifying dimensional independence in speech.


【4】Memory-guided Prototypical Co-occurrence Learning for Mixed Emotion Recognition
标题:用于混合情绪识别的记忆引导原型同现学习
链接:https://arxiv.org/abs/2602.20530

作者:Ming Li,Yong-Jin Liu,Fang Liu,Huankun Sheng,Yeying Fan,Yixiang Wei,Minnan Luo,Weizhan Zhang,Wenping Wang
摘要:从多模态生理和行为信号中识别情感在情感计算中起着关键作用,然而大多数现有模型仍然局限于在受控实验室环境中预测奇异情感。相比之下,现实世界中的人类情感体验的特点往往是同时存在多个情感状态,刺激了最近的兴趣,混合情感识别作为一个情感分布学习问题。然而,目前的研究方法往往忽视了共存情绪之间内在的效价一致性和结构相关性。为了解决这个问题,我们提出了一个记忆引导的原型共现学习(MPCL)框架,明确的情感共现模式模型。具体来说,我们首先通过多尺度联想记忆机制融合多模态信号。为了捕捉跨模态语义关系,我们构建情感特定的原型记忆库,产生丰富的生理和行为表示,并采用原型关系蒸馏,以确保跨模态对齐的潜在原型空间。此外,受人类认知记忆系统的启发,我们引入了一种记忆检索策略来提取跨情感类别的语义级共现关联。通过这种自下而上的分层抽象过程,我们的模型学习情感信息表示,以实现准确的情感分布预测。在两个公开数据集上的综合实验表明,MPCL在混合情感识别方面无论是定量还是定性都始终优于最先进的方法。
摘要:Emotion recognition from multi-modal physiological and behavioral signals plays a pivotal role in affective computing, yet most existing models remain constrained to the prediction of singular emotions in controlled laboratory settings. Real-world human emotional experiences, by contrast, are often characterized by the simultaneous presence of multiple affective states, spurring recent interest in mixed emotion recognition as an emotion distribution learning problem. Current approaches, however, often neglect the valence consistency and structured correlations inherent among coexisting emotions. To address this limitation, we propose a Memory-guided Prototypical Co-occurrence Learning (MPCL) framework that explicitly models emotion co-occurrence patterns. Specifically, we first fuse multi-modal signals via a multi-scale associative memory mechanism. To capture cross-modal semantic relationships, we construct emotion-specific prototype memory banks, yielding rich physiological and behavioral representations, and employ prototype relation distillation to ensure cross-modal alignment in the latent prototype space. Furthermore, inspired by human cognitive memory systems, we introduce a memory retrieval strategy to extract semantic-level co-occurrence associations across emotion categories. Through this bottom-up hierarchical abstraction process, our model learns affectively informative representations for accurate emotion distribution prediction. Comprehensive experiments on two public datasets demonstrate that MPCL consistently outperforms state-of-the-art methods in mixed emotion recognition, both quantitatively and qualitatively.


【5】Graph Modelling Analysis of Speech-Gesture Interaction for Aphasia Severity Estimation
标题:失语严重程度估计的言语与手势相互作用的图模型分析
链接:https://arxiv.org/abs/2602.20163

作者:Navya Martin Kollapally,Christa Akers,Renjith Nelson Joseph
备注:IJCAI
摘要:失语症是一种后天性语言障碍,由负责语言的大脑区域损伤引起。失语症可能会损害书面和口头语言的使用和理解。西方失语症成套测验修订版(WAB-R)是一种由言语语言病理学家(SLP)管理的评估工具,用于评估失语症的类型和严重程度。由于WAB-R测量的是孤立的语言技能,因此人们越来越关注对话语产出的评估,将其作为日常语言能力的一种更全面的表现。语音分析的最新进展集中在自动估计失语症的严重程度,从自发的讲话,主要依赖于孤立的语言或声学特征。在这项工作中,我们提出了一个基于图神经网络的框架来估计失语症的严重程度。我们将每个参与者的话语表示为有向多模态图,其中节点表示词汇项和手势,边缘编码单词-单词,手势-单词和单词-手势转换。GraphSAGE用于学习参与者级嵌入,从而整合来自直接邻居和整体图结构的信息。我们的研究结果表明,失语症的严重程度不是编码在孤立的词汇分布,而是出现在结构化的相互作用之间的讲话和手势。所提出的架构提供了一个可靠的自动化失语症评估,可能用于床边筛查和远程医疗为基础的监测。
摘要:Aphasia is an acquired language disorder caused by injury to the regions of the brain that are responsible for language. Aphasia may impair the use and comprehension of written and spoken language. The Western Aphasia Battery-Revised (WAB-R) is an assessment tool administered by speech-language pathologists (SLPs) to evaluate the aphasia type and severity. Because the WAB-R measures isolated linguistic skills, there has been growing interest in the assessment of discourse production as a more holistic representation of everyday language abilities. Recent advancements in speech analysis focus on automated estimation of aphasia severity from spontaneous speech, relying mostly in isolated linguistic or acoustical features. In this work, we propose a graph neural network-based framework for estimating aphasia severity. We represented each participant's discourse as a directed multi-modal graph, where nodes represent lexical items and gestures and edges encode word-word, gesture-word, and word-gesture transitions. GraphSAGE is employed to learn participant-level embeddings, thus integrating information from immediate neighbors and overall graph structure. Our results suggest that aphasia severity is not encoded in isolated lexical distribution, but rather emerges from structured interactions between speech and gesture. The proposed architecture offers a reliable automated aphasia assessment, with possible uses in bedside screening and telehealth-based monitoring.


机器翻译由腾讯交互翻译提供,仅供参考