微信公众号:arXiv_Daily
cs.SD语音
【1】An Investigation Into Various Approaches For Bengali Long-Form Speech Transcription and Bengali Speaker Diarization
标题:孟加拉语长形式语音转录和孟加拉语说话人数字化的各种方法调查
链接:https://arxiv.org/abs/2603.03158
备注:5 pages, 2 figures
摘要:孟加拉语在语音技术中仍然是一种低资源语言,特别是对于长格式转录和说话人日记等复杂任务。本文提出了一种多阶段的方法开发的“DL Sprint 4.0 -孟加拉语长格式语音识别”和“DL Sprint 4.0 -孟加拉语发言者日记”的Kaggle上的比赛,解决了挑战“谁说什么时候/什么”在长达一小时的录音。我们在孟加拉语数据(bengaliAI/tugstugi bengaliai-asr whisper-medium)上实现了微调的Whisper Medium,用于转录,并将pyannote/speaker-diarization-community-1与我们的定制训练分割模型集成,以处理多样化和嘈杂的声学环境。使用带有超参数调整的两次通过方法,我们在私人排行榜上实现了0.27的DER,在公共排行榜上实现了0.19的DER。对于转录,分块,背景噪音清理和算法后处理在私人排行榜上产生了0.38的WER。这些结果表明,有针对性的调优和战略性的数据利用可以显着提高南亚语言的AI包容性。所有相关代码可在https://github.com/Short-Potatoes/Bengali-long-form-transcription-and-diarization.git上获得 索引术语:孟加拉语语音识别,说话人日记,Whisper,ASR,低资源语言,pyannote,语音活动检测
摘要:Bengali remains a low-resource language in speech technology, especially for complex tasks like long-form transcription and speaker diarization. This paper presents a multistage approach developed for the "DL Sprint 4.0 - Bengali Long-Form Speech Recognition" and "DL Sprint 4.0 - Bengali Speaker Diarization" competitions on Kaggle, addressing the challenge of "who spoke when/what" in hour-long recordings. We implemented Whisper Medium fine-tuned on Bengali data (bengaliAI/tugstugi bengaliai-asr whisper-medium) for transcription and integrated pyannote/speaker-diarization-community-1 with our custom-trained segmentation model to handle diverse and noisy acoustic environments. Using a two-pass method with hyperparameter tuning, we achieved a DER of 0.27 on the private leaderboard and 0.19 on the public leaderboard. For transcription, chunking, background noise cleaning, and algorithmic post-processing yielded a WER of 0.38 on the private leaderboard. These results show that targeted tuning and strategic data utilization can significantly improve AI inclusivity for South Asian languages. All relevant code is available at: https://github.com/Short-Potatoes/Bengali-long-form-transcription-and-diarization.git Index Terms: Bengali speech recognition, speaker diarization, Whisper, ASR, low-resource languages, pyannote, voice activity detection
【2】Differentiable Time-Varying IIR Filtering for Real-Time Speech Denoising
标题:用于实时语音去噪的可区分时变IRR过滤
链接:https://arxiv.org/abs/2603.02794
备注:Submitted to Interspeech 2026
摘要:我们提出了TVF(时变滤波),一个低延迟的语音增强模型与1万个参数。TVF将数字信号处理(DSP)的可解释性与深度学习的适应性相结合,弥合了传统滤波与现代神经语音建模之间的差距。该模型利用轻量级神经网络主干来实时预测可微分35频段IIR滤波器级联的系数,使其能够动态适应非平稳噪声。与“黑盒”深度学习方法不同,TVF提供了一个完全可解释的处理链,其中频谱修改是明确的和可调整的。我们使用Valentini-Botinhao数据集证明了这种方法在语音去噪任务中的有效性,并将结果与静态DDSP方法和完全基于深度学习的解决方案进行了比较,表明TVF能够有效地适应不断变化的噪声条件。
摘要:We present TVF (Time-Varying Filtering), a low-latency speech enhancement model with 1 million parameters. Combining the interpretability of Digital Signal Processing (DSP) with the adaptability of deep learning, TVF bridges the gap between traditional filtering and modern neural speech modeling. The model utilizes a lightweight neural network backbone to predict the coefficients of a differentiable 35-band IIR filter cascade in real time, allowing it to adapt dynamically to non-stationary noise. Unlike ``black-box'' deep learning approaches, TVF offers a completely interpretable processing chain, where spectral modifications are explicit and adjustable. We demonstrate the efficacy of this approach on a speech denoising task using the Valentini-Botinhao dataset and compare the results to a static DDSP approach and a fully deep-learning-based solution, showing that TVF achieves effective adaptation to changing noise conditions.
【3】Single Microphone Own Voice Detection based on Simulated Transfer Functions for Hearing Aids
标题:基于模拟传递函数的助听器单麦克风自身语音检测
链接:https://arxiv.org/abs/2603.02724
摘要:本文提出了一种基于仿真的方法,使用单个麦克风的助听器中自己的语音检测(OVD)。虽然OVD可以显著提高用户舒适度和语音清晰度,但现有的解决方案通常依赖于多个麦克风或额外的传感器,从而增加了设备的复杂性和成本。为了使基于ML的OVD,而不需要昂贵的传递函数测量,我们提出了一个数据增强策略的基础上模拟的声学传递函数(ATF),暴露模型的空间传播条件的范围很广。基于变换器的分类器首先在分析生成的ATF上进行训练,然后使用数值模拟的ATF进行逐步微调,从刚性球体模型过渡到详细的头部和躯干表示。这种分层适应使模型能够在保持泛化的同时改善其空间理解。实验结果表明,95.52%的准确率模拟头和躯干测试数据。在短持续时间条件下,该模型保持90.02%的准确率与一秒的话语。在真实的助听器录音中,该模型在没有微调的情况下达到了80%的准确率,并得到了轻量级测试时间特征补偿的帮助。这突出了该模型从模拟到真实世界条件的推广能力,证明了实际可行性,并为未来的助听器设计指明了一个有希望的方向。
摘要:This paper presents a simulation-based approach to own voice detection (OVD) in hearing aids using a single microphone. While OVD can significantly improve user comfort and speech intelligibility, existing solutions often rely on multiple microphones or additional sensors, increasing device complexity and cost. To enable ML-based OVD without requiring costly transfer-function measurements, we propose a data augmentation strategy based on simulated acoustic transfer functions (ATFs) that expose the model to a wide range of spatial propagation conditions. A transformer-based classifier is first trained on analytically generated ATFs and then progressively fine-tuned using numerically simulated ATFs, transitioning from a rigid-sphere model to a detailed head-and-torso representation. This hierarchical adaptation enabled the model to refine its spatial understanding while maintaining generalization. Experimental results show 95.52% accuracy on simulated head-and-torso test data. Under short-duration conditions, the model maintained 90.02% accuracy with one-second utterances. On real hearing aid recordings, the model achieved 80% accuracy without fine-tuning, aided by lightweight test-time feature compensation. This highlights the model's ability to generalize from simulated to real-world conditions, demonstrating practical viability and pointing toward a promising direction for future hearing aid design.
【4】Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement
标题:重新思考通用语音增强的训练目标、架构和数据质量
链接:https://arxiv.org/abs/2603.02641
摘要:通用语音增强(Universal Speech Enhancement,USE)的目标是在各种退化条件下恢复语音质量,同时保持信号保真度。尽管最近取得了进展,但在训练目标选择、失真-感知权衡和数据管理方面的关键挑战仍然没有得到解决。在本书中,我们系统地解决了这三个被忽视的问题。首先,我们回顾了使用早期反射语音作为去混响目标的传统做法,并表明它会降低感知质量和下游ASR性能。相反,我们证明,时移无回声干净的语音提供了一个优越的学习目标。其次,失真-感知权衡理论的指导下,我们提出了一个简单的两阶段的框架,在给定的感知质量水平下实现最小的失真。第三,我们分析了训练数据规模和使用质量之间的权衡,揭示了在大型未经策划的语料库上进行训练会带来性能上限,因为模型很难去除细微的伪影。我们的方法在URGENT 2025非盲测试集上实现了最先进的性能,并具有很强的语言不可知泛化能力,使其能够有效地改善TTS训练数据。代码和模型将在验收后发布。
摘要:Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in training target selection, the distortion--perception tradeoff, and data curation remain unresolved. In this work, we systematically address these three overlooked problems. First, we revisit the conventional practice of using early-reflected speech as the dereverberation target and show that it can degrade perceptual quality and downstream ASR performance. We instead demonstrate that time-shifted anechoic clean speech provides a superior learning target. Second, guided by the distortion--perception tradeoff theory, we propose a simple two-stage framework that achieves minimal distortion under a given level of perceptual quality. Third, we analyze the trade-off between training data scale and quality for USE, revealing that training on large uncurated corpora imposes a performance ceiling, as models struggle to remove subtle artifacts. Our method achieves state-of-the-art performance on the URGENT 2025 non-blind test set and exhibits strong language-agnostic generalization, making it effective for improving TTS training data. Code and models will be released upon acceptance.
【5】MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
标题:MUSE:一个以运行为中心的平台,用于大型语言模型的多模式统一安全评估
链接:https://arxiv.org/abs/2603.02482
备注:Submitted to ACL 2026 System Demonstration Track
摘要:大型语言模型的安全评估和红队仍然主要以文本为中心,现有的框架缺乏系统地测试对齐是否适用于音频、图像和视频输入的基础设施。我们提出了MUSE(多模态统一安全评估),一个开源的,以运行为中心的平台,集成了自动跨模态有效载荷生成,三个多回合攻击算法(Crescendo,PAIR,Violent Durian),提供商不可知的模型路由,和一个LLM判断与五级安全分类到一个基于浏览器的系统。双指标框架将硬攻击成功率(仅合规性)与软ASR(包括部分合规性)区分开来,捕获二进制指标遗漏的部分信息泄漏。为了探索对齐是否跨越模态边界,我们引入了匝间模态切换(ITMS),它通过每匝模态旋转来增强多匝攻击。来自四家供应商的六个多模态LLM的实验表明,多转向策略可以实现高达90-100%的ASR,而不是接近完美的单转向拒绝模型。ITMS不会在已经饱和的基线上统一提高最终ASR,但会通过破坏早期防御来加速收敛,并且消融揭示了模态效应的方向是特定于模型系列而不是通用的,这强调了提供商感知的跨模态安全测试的必要性。
摘要:Safety evaluation and red-teaming of large language models remain predominantly text-centric, and existing frameworks lack the infrastructure to systematically test whether alignment generalizes to audio, image, and video inputs. We present MUSE (Multimodal Unified Safety Evaluation), an open-source, run-centric platform that integrates automatic cross-modal payload generation, three multi-turn attack algorithms (Crescendo, PAIR, Violent Durian), provider-agnostic model routing, and an LLM judge with a five-level safety taxonomy into a single browser-based system. A dual-metric framework distinguishes hard Attack Success Rate (Compliance only) from soft ASR (including Partial Compliance), capturing partial information leakage that binary metrics miss. To probe whether alignment generalizes across modality boundaries, we introduce Inter-Turn Modality Switching (ITMS), which augments multi-turn attacks with per-turn modality rotation. Experiments across six multimodal LLMs from four providers show that multi-turn strategies can achieve up to 90-100% ASR against models with near-perfect single-turn refusal. ITMS does not uniformly raise final ASR on already-saturated baselines, but accelerates convergence by destabilizing early-turn defenses, and ablation reveals that the direction of modality effects is model-family-specific rather than universal, underscoring the need for provider-aware cross-modal safety testing.
【6】RO-N3WS: Enhancing Generalization in Low-Resource ASR with Diverse Romanian Speech Benchmarks
标题:RO-N3 WS:通过多样化的罗马尼亚语音基准增强低资源ASB的概括性
链接:https://arxiv.org/abs/2603.02368
摘要:我们介绍RO-N3 WS,一个基准罗马尼亚语音数据集,旨在提高自动语音识别(ASR)的泛化,特别是在低资源和分布(OOD)条件。RO-N3 WS包括超过126小时的转录音频,这些音频来自广播新闻,文学有声读物,电影对话,儿童故事和播客对话。这种多样性使得能够在风格不同的领域进行强大的训练和微调。我们评估了几个国家的最先进的ASR系统(耳语,Wav 2 Vec 2.0)在zero-shot和微调设置,并进行控制比较,使用合成数据与表达TTS模型。我们的研究结果表明,即使是有限的微调对真正的语音从RO-N3 WS产生实质性的WER改善超过zero-shot基线。我们将发布所有的模型、脚本和数据分割,以支持多语言ASR、域适配和轻量级部署中的可重复研究。
摘要:We introduce RO-N3WS, a benchmark Romanian speech dataset designed to improve generalization in automatic speech recognition (ASR), particularly in low-resource and out-of-distribution (OOD) conditions. RO-N3WS comprises over 126 hours of transcribed audio collected from broadcast news, literary audiobooks, film dialogue, children's stories, and conversational podcast speech. This diversity enables robust training and fine-tuning across stylistically distinct domains. We evaluate several state-of-the-art ASR systems (Whisper, Wav2Vec 2.0) in both zero-shot and fine-tuned settings, and conduct controlled comparisons using synthetic data generated with expressive TTS models. Our results show that even limited fine-tuning on real speech from RO-N3WS yields substantial WER improvements over zero-shot baselines. We will release all models, scripts, and data splits to support reproducible research in multilingual ASR, domain adaptation, and lightweight deployment.
【7】When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus
标题:当欺骗检测器旅行时:低资源语言欺骗数据库中66种语言的评估
链接:https://arxiv.org/abs/2603.02364
备注:This paper has been submitted to Interspeech 2026 for review
摘要:我们介绍LRLspoof,这是一个用于跨语言欺骗检测的大规模多语言合成语音语料库,包含2,732小时的音频,由24个开源TTS系统生成,涵盖66种语言,其中包括我们操作定义下的45种低资源语言。为了在不需要目标域善意语音的情况下评估鲁棒性,我们使用阈值转移对11个公开可用的对策进行基准测试:对于每个模型,我们在合并的外部基准上校准EER操作点,并应用所产生的阈值,报告欺骗拒绝率(SRR)。结果表明,依赖于模型的跨语言的差异,欺骗拒绝显着不同的语言,即使在受控条件下,突出显示语言作为一个独立的源域转移欺骗检测。该数据集可在\href{https://huggingface.co/martets/MTUCI/LRLspoof}{\textbf{\underline{\textit{HuggingFace}和\href{https://modelscope.cn/martets/lab260/LRLspoof}{\textbf{\underline{\textit{ModelScope}上公开获取
摘要:We introduce LRLspoof, a large-scale multilingual synthetic-speech corpus for cross-lingual spoof detection, comprising 2,732 hours of audio generated with 24 open-source TTS systems across 66 languages, including 45 low-resource languages under our operational definition. To evaluate robustness without requiring target-domain bonafide speech, we benchmark 11 publicly available countermeasures using threshold transfer: for each model we calibrate an EER operating point on pooled external benchmarks and apply the resulting threshold, reporting spoof rejection rate (SRR). Results show model-dependent cross-lingual disparity, with spoof rejection varying markedly across languages even under controlled conditions, highlighting language as an independent source of domain shift in spoof detection. The dataset is publicly available at \href{https://huggingface.co/datasets/MTUCI/LRLspoof}{\textbf{\underline{\textit{HuggingFace}}}} and \href{https://modelscope.cn/datasets/lab260/LRLspoof}{\textbf{\underline{\textit{ModelScope}}}}
【8】Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
标题:语音识别中的序列级无监督训练:理论研究
链接:https://arxiv.org/abs/2603.02285
备注:accepted to ICASSP 2026
摘要:无监督语音识别是一项利用非配对数据训练语音识别模型的任务。为了确定无监督语音识别何时以及如何成功,以及分类错误如何与候选训练目标相关,我们开发了一个基于分类错误界限的无监督语音识别理论框架。我们介绍了两个条件下,无监督语音识别是可能的。并对这些条件的必要性进行了讨论。在这些条件下,我们推导出无监督语音识别的分类误差界,并在模拟中验证了这一界。受此限制,我们提出了一个单阶段的序列级交叉熵损失的无监督语音识别。
摘要:Unsupervised speech recognition is a task of training a speech recognition model with unpaired data. To determine when and how unsupervised speech recognition can succeed, and how classification error relates to candidate training objectives, we develop a theoretical framework for unsupervised speech recognition grounded in classification error bounds. We introduce two conditions under which unsupervised speech recognition is possible. The necessity of these conditions are also discussed. Under these conditions, we derive a classification error bound for unsupervised speech recognition and validate this bound in simulations. Motivated by this bound, we propose a single-stage sequence-level cross-entropy loss for unsupervised speech recognition.
【9】When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning
标题:当缩放失败时:通过多步感知推理缓解LALM的音频感知衰退
链接:https://arxiv.org/abs/2603.02266
备注:Under Review
摘要:测试时间缩放在通过缩放推理计算来解决复杂问题方面显示出显著的效果。然而,在大型音频语言模型(LALM)中,存在一种不直观的现象:与直接回答的后训练相比,结构化推理轨迹的后训练模型会产生边际甚至负增益。为了研究它,我们引入CAFE,一个评估框架,旨在精确量化音频推理错误。评估结果显示,LALM在推理过程中与感知斗争,并遇到了一个关键的瓶颈:推理性能随着推理长度的延长而受到音频感知衰减的影响。为了解决这个问题,我们提出了MPAR $^2 $,一个范式,鼓励动态感知推理和分解成感知丰富的子问题的复杂问题。利用强化学习,MPAR $^2 $将CAFE的感知性能从31.74%提高到63.51%,并有效地减轻了感知衰减,同时增强了推理能力,在MMAU基准测试中达到了74.59%的准确率。进一步的分析表明,MPAR $^2 $加强LALM出席音频输入,并动态地适应推理预算,以匹配任务的复杂性。
摘要:Test-Time Scaling has shown notable efficacy in addressing complex problems through scaling inference compute. However, within Large Audio-Language Models (LALMs), an unintuitive phenomenon exists: post-training models for structured reasoning trajectories results in marginal or even negative gains compared to post-training for direct answering. To investigate it, we introduce CAFE, an evaluation framework designed to precisely quantify audio reasoning errors. Evaluation results reveal LALMs struggle with perception during reasoning and encounter a critical bottleneck: reasoning performance suffers from audio perception decay as reasoning length extends. To address it, we propose MPAR$^2$, a paradigm that encourages dynamic perceptual reasoning and decomposes complex questions into perception-rich sub-problems. Leveraging reinforcement learning, MPAR$^2$ improves perception performance on CAFE from 31.74% to 63.51% and effectively mitigates perception decay, concurrently enhancing reasoning capabilities to achieve a significant 74.59% accuracy on the MMAU benchmark. Further analysis demonstrates that MPAR$^2$ reinforces LALMs to attend to audio input and dynamically adapts reasoning budget to match task complexity.
【10】MEBM-Speech: Multi-scale Enhanced BrainMagic for Robust MEG Speech Detection
标题:MEBM-Speech:用于稳健MEG语音检测的多尺度增强BrainMagic
链接:https://arxiv.org/abs/2603.02255
备注:5 pages, 1 figure. To appear in the PNPL Competition Workshop at NeurIPS 2025
摘要:我们提出MEBM语音,多尺度增强神经解码器的语音活动检测从非侵入性脑磁图(MEG)信号。MEBM-Speech建立在BrainMagic主干上,集成了三种互补的时间建模机制:用于短期模式提取的多尺度卷积模块,用于长期上下文建模的双向LSTM(BiLSTM),以及用于有效跨尺度特征融合的深度可分离卷积层。轻量级的时间抖动策略和平均池化进一步提高了起始鲁棒性和边界稳定性。该模型对MEG信号进行连续概率解码,从而实现语音与沉默状态的细粒度检测-这是认知神经科学和临床应用的关键能力。LibriBrain Competition 2025 Track 1基准测试的综合评估显示了强劲的性能,在验证集上实现了89.3%的平均F1宏,并在官方测试排行榜上获得了可比的结果。这些发现突出了多尺度时间表示学习的有效性,鲁棒的基于MEG的语音解码。
摘要:We propose MEBM-Speech, a multi-scale enhanced neural decoder for speech activity detection from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, MEBM-Speech integrates three complementary temporal modeling mechanisms: a multi-scale convolutional module for short-term pattern extraction, a bidirectional LSTM (BiLSTM) for long-range context modeling, and a depthwise separable convolutional layer for efficient cross-scale feature fusion. A lightweight temporal jittering strategy and average pooling further improve onset robustness and boundary stability. The model performs continuous probabilistic decoding of MEG signals, enabling fine-grained detection of speech versus silence states - an ability crucial for both cognitive neuroscience and clinical applications. Comprehensive evaluations on the LibriBrain Competition 2025 Track1 benchmark demonstrate strong performance, achieving an average F1 macro of 89.3% on the validation set and comparable results on the official test leaderboard. These findings highlight the effectiveness of multi-scale temporal representation learning for robust MEG-based speech decoding.
【11】MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification
标题:MEBM音素:用于端到端MEG音素分类的多尺度增强BrainMagic
链接:https://arxiv.org/abs/2603.02254
备注:5 pages, 1 figure. To appear in the PNPL Competition Workshop at NeurIPS 2025
摘要:我们提出了MEBM-Phoneme,这是一种多尺度增强神经解码器,用于从非侵入性脑磁图(MEG)信号中进行音素分类。MEBM-Phoneme基于BrainMagic主干构建,集成了一个短期多尺度卷积模块,以增强本地中期编码器,并通过深度可分离卷积进行融合表示,以实现高效的跨尺度集成。卷积注意力层动态地对时间依赖性进行加权以细化特征聚合。为了解决类不平衡和会话特定的分布变化,我们引入了一个基于堆栈的本地验证集,以及加权交叉熵损失和随机时间增强。对LibriBrain Competition 2025 Track 2的全面评估显示出强大的泛化能力,在验证和官方测试排行榜上实现了具有竞争力的音素解码准确性。这些结果强调了分层时间建模和训练稳定性对于推进基于MEG的语音感知分析的价值。
摘要:We propose MEBM-Phoneme, a multi-scale enhanced neural decoder for phoneme classification from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, MEBM-Phoneme integrates a short-term multi-scale convolutional module to augment the native mid-term encoder, with fused representations via depthwise separable convolution for efficient cross-scale integration. A convolutional attention layer dynamically weights temporal dependencies to refine feature aggregation. To address class imbalance and session-specific distributional shifts, we introduce a stacking-based local validation set alongside weighted cross-entropy loss and random temporal augmentation. Comprehensive evaluations on LibriBrain Competition 2025 Track2 demonstrate robust generalization, achieving competitive phoneme decoding accuracy on the validation and official test leaderboard. These results underscore the value of hierarchical temporal modeling and training stabilization for advancing MEG-based speech perception analysis.
【12】SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models
标题:SGPA:多模式大型语言模型中可行的Shapley值解释的谱图引导语音对齐
链接:https://arxiv.org/abs/2603.02250
备注:Submitted for admission in Interspeech 2026 conference
摘要:通过Shapley值属性解释端到端音频语言模型的行为在原生标记化下是棘手的:典型的话语产生超过150 $$的编码器帧,相对于文本将联盟空间膨胀大约10 ^{42}$;单个音频帧缺乏独立意义;平分语音过渡的标记边界引入掩蔽伪像。我们引入了频谱图引导的语音对齐(SGPA),这是一个四阶段的流水线,它将联结主义时间分类强制对齐与频谱边界细化相结合,以产生声学稳定的单词对齐的音频片段。使用VoiceBench对LFM 2-Audio-1.5 B进行的受控诊断显示,SGPA在模型评估中产生43$\times$减少。统计测试证实,SGPA显着改变归因集中,同时保持全球累积配置文件,建立它作为一个可行性,使层音频可解释性。
摘要:Explaining the behavior of end-to-end audio language models via Shapley value attribution is intractable under native tokenization: a typical utterance yields over $150$ encoder frames, inflating the coalition space by roughly $10^{42}$ relative to text; individual audio frames lack standalone meaning; and token boundaries that bisect phonetic transitions introduce masking artifacts. We introduce Spectrogram-Guided Phonetic Alignment (SGPA), a four-stage pipeline that combines Connectionist Temporal Classification forced alignment with spectral boundary refinement to produce acoustically stable, word-aligned audio segments. Controlled diagnostics on LFM2-Audio-1.5B with VoiceBench show that SGPA yields a 43$\times$ reduction in model evaluations. Statistical testing confirms that SGPA significantly alters attribution concentration while preserving the global cumulative profile, establishing it as a feasibility-enabling layer for audio explainability.
【13】Decomposing the Influence of Physical Acoustic Modeling on Neural Personal Sound Zone Rendering: An Ablation Study
标题:分解物理声学建模对神经个人音区渲染的影响:一项消融研究
链接:https://arxiv.org/abs/2603.02508
摘要:基于深度学习的个人声音区域(PSZ)依赖于模拟的声学传递函数(ATF)进行训练,但理想化的点源模型表现出很大的模拟与真实差距。虽然物理上知情的组成部分改善了泛化,但个人的贡献仍然不清楚。本文提出了一种控制消融研究的头部姿势条件双耳PSZ渲染器使用双耳空间音频神经网络(BSANN)。我们逐步丰富模拟的ATF与三个组成部分:(一)消声测量的频率响应的特定扬声器(FR),(ii)分析圆形活塞方向性(ESTA),和(iii)刚性球头相关的传递函数(RS-HRTF)。四个配置进行了评估,通过现场测量与两个虚拟头。性能指标包括区间隔离(IZI)、程序间干扰(IPI)和100-20000 Hz范围内的串扰消除(XTC)。结果显示,FR提供了频谱校准,产生了适度的XTC改进并减少了听众间IPI不平衡。降噪提供最一致的声区分离增益(平均IZI/IPI为10.05 dB)。RS-HRTF主导双耳分离,将XTC提高+2.38/+2.89 dB(平均4.51至7.91 dB),主要高于2 kHz,同时引入轻微的依赖于信标的IZI/IPI偏移。这些研究结果指导优先级的测量和模型时,在有限的预算下构建培训ATF。
摘要:Deep learning-based Personal Sound Zones (PSZs) rely on simulated acoustic transfer functions (ATFs) for training, yet idealized point-source models exhibit large sim-to-real gaps. While physically informed components improve generalization, individual contributions remain unclear. This paper presents a controlled ablation study on a head-pose-conditioned binaural PSZ renderer using the Binaural Spatial Audio Neural Network (BSANN). We progressively enrich simulated ATFs with three components: (i) anechoically measured frequency responses of the particular loudspeakers(FR), (ii) analytic circular-piston directivity (DIR), and (iii) rigid-sphere head-related transfer functions (RS-HRTF). Four configurations are evaluated via in-situ measurements with two dummy heads. Performance metrics include inter-zone isolation (IZI), inter-program interference (IPI), and crosstalk cancellation (XTC) over 100-20000 Hz. Results show FR provides spectral calibration, yielding modest XTC improvements and reduced inter-listener IPI imbalance. DIR delivers the most consistent sound-zone separation gains (10.05 dB average IZI/IPI). RS-HRTF dominates binaural separation, boosting XTC by +2.38/+2.89 dB (average 4.51 to 7.91 dB), primarily above 2 kHz, while introducing mild listener-dependent IZI/IPI shifts. These findings guide prioritization of measurements and models when constructing training ATFs under limited budgets.
【14】Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
标题:Whisper-RIR-Mega:ASB对房间声学鲁棒性的配对清洁-回响语音基准
链接:https://arxiv.org/abs/2603.02252
摘要:我们介绍Whisper-RIR-Mega,一个用于评估自动语音识别(ASR)对室内声学的鲁棒性的配对干净和混响语音的基准数据集。每个样本将干净的LibriSpeech话语与来自RIR-Mega语料库的真实房间脉冲响应卷积的相同话语配对,其中通过混响时间(RT 60)和直接混响比(DRR)分层分割。我们在1600个测试样本上评估了五种Whisper模型(从小到大v3),并在干净和混响条件下报告了单词错误率(WER)和字符错误率(CER)。混响会持续降低所有型号的性能; WER中的混响损失范围从0.12到1.07个百分点,具体取决于型号。我们发布了数据集,评估代码和基线结果,以支持对强大ASR的可重复研究。
摘要:We introduce Whisper-RIR-Mega, a benchmark dataset of paired clean and reverberant speech for evaluating automatic speech recognition (ASR) robustness to room acoustics. Each sample pairs a clean LibriSpeech utterance with the same utterance convolved with a real room impulse response from the RIR-Mega corpus, with stratified splits by reverberation time (RT60) and direct-to-reverberant ratio (DRR). We evaluate five Whisper models (tiny through large-v3) on 1600 test samples and report word error rate (WER) and character error rate (CER) under clean and reverberant conditions. Reverberation consistently degrades performance across all model sizes; the reverb penalty in WER ranges from 0.12 to 1.07 percentage points depending on the model. We release the dataset, evaluation code, and baseline results to support reproducible research on robust ASR.
【15】OnDA: On-device Channel Pruning for Efficient Personalized Keyword Spotting
标题:OnDA:设备上频道修剪,以实现高效的个性化关键词发现
链接:https://arxiv.org/abs/2603.02247
备注:Submitted for review at Interspeech2026
摘要:始终在线的关键字识别(KWS)需要设备自适应,以在严格的延迟和能源预算下应对特定于用户和环境的分布变化。本文首次提出了耦合权重自适应(即,设备上的训练)与架构自适应,在线结构化通道修剪的形式,个性化的设备上的KWS。从最先进的自学习个性化KWS管道开始,我们比较了应用于现场伪标记用户数据的数据不可知和数据感知修剪标准。在HeySnips和HeySnapdragon数据集上,我们在iso-task性能下相对于未修剪的基线实现了高达9.63倍的模型大小压缩,以每小时0.5个错误警报的准确度来衡量。当在Jetson Orin Nano嵌入式GPU上部署我们的适配管道时,与仅权重适配相比,我们在在线训练/推理期间实现了高达1.52倍/1.57倍和1.64倍/1.77倍的延迟和能耗改进。
摘要:Always-on keyword spotting (KWS) demands on-device adaptation to cope with user- and environment-specific distribution shifts under tight latency and energy budgets. This paper proposes, for the first time, coupling weight adaptation (i.e., on-device training) with architectural adaptation, in the form of online structured channel pruning, for personalized on-device KWS. Starting from a state-of-the-art self-learning personalized KWS pipeline, we compare data-agnostic and data-aware pruning criteria applied on in-field pseudo-labelled user data. On the HeySnips and HeySnapdragon datasets, we achieve up to 9.63x model-size compression with respect to unpruned baselines at iso-task performance, measured as the accuracy at 0.5 false alarms per hour. When deploying our adaptation pipeline on a Jetson Orin Nano embedded GPU, we achieve up to 1.52x/1.57x and 1.64x/1.77x latency and energy-consumption improvements during online training/inference compared to weights-only adaptation.
【16】Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
标题:自动语音识别的质量--波兰语案例研究--从Wave 2 Vec到Scribe ElevenLabs
链接:https://arxiv.org/abs/2603.02246
摘要:本文涉及的比较研究的自动语音识别(ASR)模型与大语言模型(LLM)用于医疗采访。提出的解决方案进行了测试波兰语言基准和数据集与医疗采访。最新的ASR技术基于卷积神经网络(CNN)、递归神经网络(RNN)和Transformers。大多数都是端到端的解决方案。在Whisper模型的情况下,所提出的方法显示了一个两阶段的解决方案,端到端ASR和LLM在管道中一起工作。ASR输出是LLM的输入。LLM是校正和改善ASR输出的组件。对现代端到端深度学习架构和ASR混合模型之间的波兰语自动识别进行了比较研究。医学访谈测试使用两种最先进的ASR模型进行:与LLM和Scribe ElevenLabs合并的OpenAI Whisper。此外,还将结果与Mozilla Common Voice和VoxPopuli数据库上的五个端到端模型(QuartzNet,FastConformer,Wav2Vec 2.0 XLSR和ESPnet Model Zoo)进行了比较。对干净的音频信号、带宽受限的信号和降级的信号进行了测试。测试模型进行了评估的基础上的字错误率(WER)和字符错误率(CER)。结果表明,Whisper模型在开源模型中表现最好。另一方面,ElevenLabs Scribe模型在一般基准和医疗数据上对波兰语表现最好。
摘要:This article concerns comparative studies on the Automatic Speech Recognition (ASR) model incorporated with the Large Language Model (LLM) used for medical interviews. The proposed solution is tested on polish language benchmarks and dataset with medical interviews. The latest ASR technologies are based on convolutional neural networks (CNNs), recurrent neural networks (RNNs) and Transformers. Most of them work as end-to-end solutions. The presented approach in the case of the Whisper model shows a two-stage solution with End-To-End ASR and LLM working together in a pipeline. The ASR output is an input for LLM. The LLM is a component by which the output from ASR is corrected and improved. Comparative studies for automatic recognition of the Polish language between modern End-To-End deep learning architectures and the ASR hybrid model were performed. The medical interview tests were performed with two state-of-the-art ASR models: OpenAI Whisper incorporated with LLM and Scribe ElevenLabs. Additionally, the results were compared with five more end-to-end models (QuartzNet, FastConformer, Wav2Vec 2.0 XLSR and ESPnet Model Zoo) on Mozilla Common Voice and VoxPopuli databases. Tests were conducted for clean audio signal, signal with bandwidth limitation, and degraded. The tested models were evaluated on the basis of Word Error Rate (WER) and Character Error Rate (CER). The results show that the Whisper model performs by far the best among the open-source models. ElevenLabs Scribe model, on the other hand, performs best for Polish on both general benchmark and medical data.
【17】LMU-Based Sequential Learning and Posterior Ensemble Fusion for Cross-Domain Infant Cry Classification
标题:基于LMU的序列学习和后验集融合用于跨领域婴儿哭声分类
链接:https://arxiv.org/abs/2603.02245
备注:7 pages
摘要:由于短的非平稳信号、有限的注释以及婴儿和数据集之间的强域偏移,解码婴儿哭泣原因对于医疗保健监测仍然具有挑战性。我们提出了一个紧凑的声学框架,融合MFCC,STFT和音高功能的多分支CNN编码器和模型的时间动态使用增强的勒让德记忆单元(LMU)。与LSTM相比,LMU主干提供稳定的序列建模,具有显著更少的重复参数,支持高效部署。为了提高跨数据集的泛化能力,我们引入了具有熵门控加权的校准后验集成融合,以保留特定领域的专业知识,同时减轻数据集偏差。在Baby2020和Baby Crying上的实验证明了跨域评估下改进的宏F1,以及泄漏软件分裂和实时设备监控的可行性。
摘要:Decoding infant cry causes remains challenging for healthcare monitoring due to short nonstationary signals, limited annotations, and strong domain shifts across infants and datasets. We propose a compact acoustic framework that fuses MFCC, STFT, and pitch features within a multi-branch CNN encoder and models temporal dynamics using an enhanced Legendre Memory Unit (LMU). Compared to LSTMs, the LMU backbone provides stable sequence modeling with substantially fewer recurrent parameters, supporting efficient deployment. To improve cross-dataset generalization, we introduce calibrated posterior ensemble fusion with entropy-gated weighting to preserve domain-specific expertise while mitigating dataset bias. Experiments on Baby2020 and Baby Crying demonstrate improved macro-F1 under cross-domain evaluation, along with leakageaware splits and real-time feasibility for on-device monitoring.
【1】Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
标题:从自我监督言语特征维度解读说话者特征
链接:https://arxiv.org/abs/2603.03096
备注:5 pages, 7 figures, submitted to IEEE Signal Processing Letters
摘要:通过自监督学习训练的语音模型如何构建其表示?以前的研究已经研究了信息如何在不同层的特征向量中编码。但是很少有研究考虑语音特征是否在SSL特征的各个维度内被捕获。在本文中,我们专门研究说话人信息使用PCA的话语平均表示。使用WavLM,我们发现解释大多数方差的主维度编码音高和相关特征,如性别。其他个别的主要尺寸与强度,噪声水平,第二共振峰,和更高的频率特性。最后,在合成实验中,我们表明,大多数特性可以通过改变相应的尺寸进行控制。这提供了在合成应用中控制输出语音的特性的简单方法。
摘要:How do speech models trained through self-supervised learning structure their representations? Previous studies have looked at how information is encoded in feature vectors across different layers. But few studies have considered whether speech characteristics are captured within individual dimensions of SSL features. In this paper we specifically look at speaker information using PCA on utterance-averaged representations. Using WavLM, we find that the principal dimension that explains most variance encodes pitch and associated characteristics like gender. Other individual principal dimensions correlate with intensity, noise levels, the second formant, and higher frequency characteristics. Finally, in synthesis experiments we show that most characteristics can be controlled by changing the corresponding dimensions. This provides a simple method to control characteristics of the output voice in synthesis applications.
【2】DLIOS: An LLM-Augmented Real-Time Multi-Modal Interactive Enhancement Overlay System for Douyin Live Streaming
标题:DLIOS:用于抖音直播的LLM增强实时多模式互动增强叠加系统
链接:https://arxiv.org/abs/2603.03060
备注:14 pages, 13 figures, 6 tables, 7 algorithms, 16 references, submitted to ACM/IEEE International Conference on Systems and Software Engineering
摘要:我们提出了DLIOS,一个大语言模型(LLM)增强的实时多模态交互增强覆盖系统,用于抖音(TikTok)直播。DLIOS采用三层透明窗口架构,用于独立渲染danmaku(滚动文本)、礼物和类似粒子效果以及VIP入口动画,围绕事件驱动的WebView 2捕获管道和线程安全事件总线构建。在此基础上,我们提出了一个LLM广播自动化框架,包括:(1)一个每首歌曲四段提示调度系统(T1开场/过渡,T2移情,T3时代故事/制作笔记,T4结束),其从歌词元数据生成情感上连贯的广播风格评论;(2)支持热插拔多角色广播的JSON可串行化的RadioPersonaConfig模式;(3)实时danmaku快速反应引擎,关键字路由到静态紧急语音或LLM生成的移情响应;以及(4)Suwan Li AI创作歌手人物角色案例研究-Suno制作的100多首AI生成的歌曲。36小时的压力测试表明:零danmaku重叠,零死锁崩溃,礼物效应P95延迟<= 180 ms,LLM到TTS段P95延迟<= 2.1 s,TTS综合响度增益为9.5 LUFS。直播; danmaku;大型语言模型;迅速的工程设计;虚拟人物; WebView 2; WINMM; TTS; Suno;响度归一化;实时调度
摘要:We present DLIOS, a Large Language Model (LLM)-augmented real-time multi-modal interactive enhancement overlay system for Douyin (TikTok) live streaming. DLIOS employs a three-layer transparent window architecture for independent rendering of danmaku (scrolling text), gift and like particle effects, and VIP entrance animations, built around an event-driven WebView2 capture pipeline and a thread-safe event bus. On top of this foundation we contribute an LLM broadcast automation framework comprising: (1) a per-song four-segment prompt scheduling system (T1 opening/transition, T2 empathy, T3 era story/production notes, T4 closing) that generates emotionally coherent radio-style commentary from lyric metadata; (2) a JSON-serializable RadioPersonaConfig schema supporting hot-swap multi-persona broadcasting; (3) a real-time danmaku quick-reaction engine with keyword routing to static urgent speech or LLM-generated empathetic responses; and (4) the Suwan Li AI singer-songwriter persona case study -- over 100 AI-generated songs produced with Suno. A 36-hour stress test demonstrates: zero danmaku overlap, zero deadlock crashes, gift effect P95 latency <= 180 ms, LLM-to-TTS segment P95 latency <= 2.1 s, and TTS integrated loudness gain of 9.5 LUFS. live streaming; danmaku; large language model; prompt engineering; virtual persona; WebView2; WINMM; TTS; Suno; loudness normalization; real-time scheduling
【3】Bias and Fairness in Self-Supervised Acoustic Representations for Cognitive Impairment Detection
标题:用于认知障碍检测的自我监督声学表示中的偏差和公平性
链接:https://arxiv.org/abs/2603.02937
备注:12 pages, 4 figures, 6 tables, Journal paper
摘要:基于语音的认知障碍(CI)检测为早期诊断提供了一种很有前途的非侵入性方法,但人口统计学和临床亚组之间的性能差异仍未得到充分研究,引起了人们对公平性和普遍性的担忧。本研究提出了一个系统的偏见分析声学为基础的CI和抑郁症分类使用DementiaBank皮特语料库。我们比较了传统的声学特征(MFCC,eGeMAPS)与Wav 2 Vec 2.0(W2 V2)的上下文语音嵌入,并评估了性别,年龄和抑郁状态亚组的分类性能。对于CI检测,更高层的W2 V2嵌入优于基线特征(UAR高达80.6%),但表现出性能差异;特别是,女性和年轻参与者表现出较低的辨别力(AUC:0.769和0.746)和显著的特异性差异(Δ spec分别高达18%和15%),导致误分类的风险高于其对应物。这些差异反映了代表性偏倚,定义为人口统计学或临床亚组之间模型性能的系统性差异。CI受试者中的抑郁检测产生较低的整体性能,从低和中等水平的W2 V2层轻微改善。CI和抑郁症分类之间的跨任务概括是有限的,这表明每个任务依赖于不同的表征。这些研究结果强调,需要公平意识的模型评估和亚组特定的分析,在临床语音应用程序,特别是在现实世界的应用程序中的人口和临床异质性。
摘要:Speech-based detection of cognitive impairment (CI) offers a promising non-invasive approach for early diagnosis, yet performance disparities across demographic and clinical subgroups remain underexplored, raising concerns around fairness and generalizability. This study presents a systematic bias analysis of acoustic-based CI and depression classification using the DementiaBank Pitt Corpus. We compare traditional acoustic features (MFCCs, eGeMAPS) with contextualized speech embeddings from Wav2Vec 2.0 (W2V2), and evaluate classification performance across gender, age, and depression-status subgroups. For CI detection, higher-layer W2V2 embeddings outperform baseline features (UAR up to 80.6\%), but exhibit performance disparities; specifically, females and younger participants demonstrate lower discriminative power (\(AUC\): 0.769 and 0.746, respectively) and substantial specificity disparities (\(Δ_{spec}\) up to 18\% and 15\%, respectively), leading to a higher risk of misclassifications than their counterparts. These disparities reflect representational biases, defined as systematic differences in model performance across demographic or clinical subgroups. Depression detection within CI subjects yields lower overall performance, with mild improvements from low and mid-level W2V2 layers. Cross-task generalization between CI and depression classification is limited, indicating that each task depends on distinct representations. These findings emphasize the need for fairness-aware model evaluation and subgroup-specific analysis in clinical speech applications, particularly in light of demographic and clinical heterogeneity in real-world applications.
【4】Does Fine-tuning by Reinforcement Learning Improve Generalization in Binary Speech Deepfake Detection?
标题:通过强化学习进行的微调是否能改善二进制语音Deepfake检测的概括?
链接:https://arxiv.org/abs/2603.02914
备注:Submitted to Interspeech 2026; put on arxiv based on requirement of paper open-access rule; quote from Interspeech: "Interspeech no longer enforces an anonymity period for submissions. While uploading a version online is permitted, your official submission to Interspeech must not contain any author-identifying information"
摘要:构建可推广到看不见的攻击的语音deepfake检测模型仍然是一个具有挑战性的问题。虽然该领域已经转向使用语音基础模型的预训练和微调范式,但大多数方法仅依赖于监督微调(SFT)。受大型语言模型领域的启发,其中强化学习(RL)用于模型微调,我们研究了RL的影响,特别是组相对策略优化(GRPO)。使用多个检测器和测试集的实验结果表明,纯基于GRPO的微调提高了域外测试集的性能,同时保持了目标域测试数据的性能。这种方法优于仅SFT和混合设置。我们的消融研究进一步表明,GRPO中的负奖励可能是这种改善的关键因素。
摘要:Building speech deepfake detection models that are generalizable to unseen attacks remains a challenging problem. Although the field has shifted toward a pre-training and fine-tuning paradigm using speech foundation models, most approaches rely solely on supervised fine-tuning (SFT). Inspired by the field of large language models, wherein reinforcement learning (RL) is used for model fine-tuning, we investigate the impact of RL, specifically Group Relative Policy Optimization (GRPO). The results from experiments using multiple detectors and test sets indicate that pure GRPO-based fine-tuning improves performance on out-of-domain test sets while maintaining performance on target-domain test data. This approach outperforms both SFT-only and hybrid setups. Our ablation studies further suggest that the negative reward in GRPO may be a key factor in this improvement.
【5】DBMIF: a deep balanced multimodal iterative fusion framework for air- and bone-conduction speech enhancement
标题:DB米非司酮:用于空气和骨导语音增强的深度平衡多模式迭代融合框架
链接:https://arxiv.org/abs/2603.02877
备注:10 pages, 7 figures, Applied Intelligence
摘要:传统的语音增强系统的性能急剧下降,在极低的信噪比(SNR)的环境中,空气传导(AC)麦克风被淹没的环境噪声。虽然骨传导(BC)传感器提供互补的,噪声容忍的信息,现有的融合方法的斗争,以保持一致的性能在广泛的SNR条件。为了解决这个问题,我们提出了深度平衡多模态迭代融合框架(DBMIF),这是一个三分支架构,旨在通过严格的跨模态交互来重建高保真语音。具体而言,接地在多尺度交互式编码器-解码器骨干,框架编排迭代注意模块和交叉分支门控模块,以促进自适应加权和双向交换。为了补充这种动态交互,进一步集成平衡交互瓶颈以学习紧凑,稳定的融合表示。大量的实验表明,DBMIF实现竞争力的性能相比,最近的单峰和多模态基线在语音质量和可懂度在不同的噪声类型。在下游ASR任务中,与竞争方法相比,该方法将字符错误率降低了至少2.5%。这些结果证实了DBMIF有效地利用了BC语音的鲁棒性,同时保留了AC语音的自然性,确保了在现实世界场景中的可靠性。源代码可在github.com/wyl516w/dbmif上公开获得。
摘要:The performance of conventional speech enhancement systems degrades sharply in extremely low signal-to-noise ratio (SNR) environments where air-conduction (AC) microphones are overwhelmed by ambient noise. Although bone-conduction (BC) sensors offer complementary, noise-tolerant information, existing fusion approaches struggle to maintain consistent performance across a wide range of SNR conditions. To address this limitation, we propose the Deep Balanced Multimodal Iterative Fusion Framework (DBMIF), a three-branch architecture designed to reconstruct high-fidelity speech through rigorous cross-modal interaction. Specifically, grounded in a multi-scale interactive encoder-decoder backbone, the framework orchestrates an iterative attention module and a cross-branch gated module to facilitate adaptive weighting and bidirectional exchange. To complement this dynamic interaction, a balanced-interaction bottleneck is further integrated to learn a compact, stable fused representation. Extensive experiments demonstrate that DBMIF achieves competitive performance compared with recent unimodal and multimodal baselines in both speech quality and intelligibility across diverse noise types. In downstream ASR tasks, the proposed method reduces the character error rate by at least 2.5 percent compared to competing approaches. These results confirm that DBMIF effectively harnesses the robustness of BC speech while preserving the naturalness of AC speech, ensuring reliability in real-world scenarios. The source code is publicly available at github.com/wyl516w/dbmif.
【6】Benchmarking Speech Systems for Frontline Health Conversations: The DISPLACE-M Challenge
标题:一线健康对话语音系统基准测试:DISPLACE-M挑战
链接:https://arxiv.org/abs/2603.02813
备注:Submitted for review to Interspeech 2026
摘要:对话环境中语言理解的DIarization和语音处理-医疗(DISPLACE-M)挑战引入了一个对话AI基准,专注于理解在该领域收集的目标导向的真实世界医疗对话。该挑战解决了医疗保健工作者和寻求者之间的多扬声器互动,其特点是印度语言和方言之间的自发,嘈杂和重叠的语音。作为挑战的一部分,发布了包括25小时开发数据和10小时盲测评估记录的医学对话数据集。我们在一个统一的端到端管道中提供了4个任务的基线系统-说话人日志化,自动语音识别,主题识别和对话摘要-以实现一致的基准测试。系统性能的评估使用既定的指标,如日记错误率(DER),时间约束的最小排列字错误率(tcpWER),和ROUGE-L。在这次评估(第一阶段)中,全球12个团队积极参与推动这些指标的基线系统。然而,即使有来自不同参与者的6-8周的专门努力,该任务也被证明是相当具有挑战性的,并且现有系统明显缺乏医疗保健部署准备。
摘要:The DIarization and Speech Processing for LAnguage understanding in Conversational Environments - Medical (DISPLACE-M) challenge introduces a conversational AI benchmark focused on understanding goal-oriented, real-world medical dialogues collected in the field. The challenge addresses multi-speaker interactions between healthcare workers and seekers characterized by spontaneous, noisy and overlapping speech across Indian languages and dialects. As part of the challenge, medical conversational dataset comprising 25 hours of development data and 10 hours of blind evaluation recordings was released. We provided baseline systems within a unified end-to-end pipeline across 4 tasks - speaker diarization, automatic speech recognition, topic identification and dialogue summarization - to enable consistent benchmarking. System performance is evaluated using established metrics such as diarization error rate (DER), time-constrained minimum-permutation word error rate (tcpWER), and ROUGE-L. During this evaluation (Phase-I), 12 teams, across the globe, actively participated pushing the baseline systems on these metrics. However, even with a 6-8 week dedicated effort from various participants, the task is shown to be substantially challenging, and the existing systems are significantly short of healthcare deployment readiness.
【7】Decomposing the Influence of Physical Acoustic Modeling on Neural Personal Sound Zone Rendering: An Ablation Study
标题:分解物理声学建模对神经个人音区渲染的影响:一项消融研究
链接:https://arxiv.org/abs/2603.02508
摘要:基于深度学习的个人声音区域(PSZ)依赖于模拟的声学传递函数(ATF)进行训练,但理想化的点源模型表现出很大的模拟与真实差距。虽然物理上知情的组成部分改善了泛化,但个人的贡献仍然不清楚。本文提出了一种控制消融研究的头部姿势条件双耳PSZ渲染器使用双耳空间音频神经网络(BSANN)。我们逐步丰富模拟的ATF与三个组成部分:(一)消声测量的频率响应的特定扬声器(FR),(ii)分析圆形活塞方向性(ESTA),和(iii)刚性球头相关的传递函数(RS-HRTF)。四个配置进行了评估,通过现场测量与两个虚拟头。性能指标包括区间隔离(IZI)、程序间干扰(IPI)和100-20000 Hz范围内的串扰消除(XTC)。结果显示,FR提供了频谱校准,产生了适度的XTC改进并减少了听众间IPI不平衡。降噪提供最一致的声区分离增益(平均IZI/IPI为10.05 dB)。RS-HRTF主导双耳分离,将XTC提高+2.38/+2.89 dB(平均4.51至7.91 dB),主要高于2 kHz,同时引入轻微的依赖于信标的IZI/IPI偏移。这些研究结果指导优先级的测量和模型时,在有限的预算下构建培训ATF。
摘要:Deep learning-based Personal Sound Zones (PSZs) rely on simulated acoustic transfer functions (ATFs) for training, yet idealized point-source models exhibit large sim-to-real gaps. While physically informed components improve generalization, individual contributions remain unclear. This paper presents a controlled ablation study on a head-pose-conditioned binaural PSZ renderer using the Binaural Spatial Audio Neural Network (BSANN). We progressively enrich simulated ATFs with three components: (i) anechoically measured frequency responses of the particular loudspeakers(FR), (ii) analytic circular-piston directivity (DIR), and (iii) rigid-sphere head-related transfer functions (RS-HRTF). Four configurations are evaluated via in-situ measurements with two dummy heads. Performance metrics include inter-zone isolation (IZI), inter-program interference (IPI), and crosstalk cancellation (XTC) over 100-20000 Hz. Results show FR provides spectral calibration, yielding modest XTC improvements and reduced inter-listener IPI imbalance. DIR delivers the most consistent sound-zone separation gains (10.05 dB average IZI/IPI). RS-HRTF dominates binaural separation, boosting XTC by +2.38/+2.89 dB (average 4.51 to 7.91 dB), primarily above 2 kHz, while introducing mild listener-dependent IZI/IPI shifts. These findings guide prioritization of measurements and models when constructing training ATFs under limited budgets.
【8】Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
标题:Whisper-RIR-Mega:ASB对房间声学鲁棒性的配对清洁-回响语音基准
链接:https://arxiv.org/abs/2603.02252
摘要:我们介绍Whisper-RIR-Mega,一个用于评估自动语音识别(ASR)对室内声学的鲁棒性的配对干净和混响语音的基准数据集。每个样本将干净的LibriSpeech话语与来自RIR-Mega语料库的真实房间脉冲响应卷积的相同话语配对,其中通过混响时间(RT 60)和直接混响比(DRR)分层分割。我们在1600个测试样本上评估了五种Whisper模型(从小到大v3),并在干净和混响条件下报告了单词错误率(WER)和字符错误率(CER)。混响会持续降低所有型号的性能; WER中的混响损失范围从0.12到1.07个百分点,具体取决于型号。我们发布了数据集,评估代码和基线结果,以支持对强大ASR的可重复研究。
摘要:We introduce Whisper-RIR-Mega, a benchmark dataset of paired clean and reverberant speech for evaluating automatic speech recognition (ASR) robustness to room acoustics. Each sample pairs a clean LibriSpeech utterance with the same utterance convolved with a real room impulse response from the RIR-Mega corpus, with stratified splits by reverberation time (RT60) and direct-to-reverberant ratio (DRR). We evaluate five Whisper models (tiny through large-v3) on 1600 test samples and report word error rate (WER) and character error rate (CER) under clean and reverberant conditions. Reverberation consistently degrades performance across all model sizes; the reverb penalty in WER ranges from 0.12 to 1.07 percentage points depending on the model. We release the dataset, evaluation code, and baseline results to support reproducible research on robust ASR.
【9】OnDA: On-device Channel Pruning for Efficient Personalized Keyword Spotting
标题:OnDA:设备上频道修剪,以实现高效的个性化关键词发现
链接:https://arxiv.org/abs/2603.02247
备注:Submitted for review at Interspeech2026
摘要:始终在线的关键字识别(KWS)需要设备自适应,以在严格的延迟和能源预算下应对特定于用户和环境的分布变化。本文首次提出了耦合权重自适应(即,设备上的训练)与架构自适应,在线结构化通道修剪的形式,个性化的设备上的KWS。从最先进的自学习个性化KWS管道开始,我们比较了应用于现场伪标记用户数据的数据不可知和数据感知修剪标准。在HeySnips和HeySnapdragon数据集上,我们在iso-task性能下相对于未修剪的基线实现了高达9.63倍的模型大小压缩,以每小时0.5个错误警报的准确度来衡量。当在Jetson Orin Nano嵌入式GPU上部署我们的适配管道时,与仅权重适配相比,我们在在线训练/推理期间实现了高达1.52倍/1.57倍和1.64倍/1.77倍的延迟和能耗改进。
摘要:Always-on keyword spotting (KWS) demands on-device adaptation to cope with user- and environment-specific distribution shifts under tight latency and energy budgets. This paper proposes, for the first time, coupling weight adaptation (i.e., on-device training) with architectural adaptation, in the form of online structured channel pruning, for personalized on-device KWS. Starting from a state-of-the-art self-learning personalized KWS pipeline, we compare data-agnostic and data-aware pruning criteria applied on in-field pseudo-labelled user data. On the HeySnips and HeySnapdragon datasets, we achieve up to 9.63x model-size compression with respect to unpruned baselines at iso-task performance, measured as the accuracy at 0.5 false alarms per hour. When deploying our adaptation pipeline on a Jetson Orin Nano embedded GPU, we achieve up to 1.52x/1.57x and 1.64x/1.77x latency and energy-consumption improvements during online training/inference compared to weights-only adaptation.
【10】Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
标题:自动语音识别的质量--波兰语案例研究--从Wave 2 Vec到Scribe ElevenLabs
链接:https://arxiv.org/abs/2603.02246
摘要:本文涉及的比较研究的自动语音识别(ASR)模型与大语言模型(LLM)用于医疗采访。提出的解决方案进行了测试波兰语言基准和数据集与医疗采访。最新的ASR技术基于卷积神经网络(CNN)、递归神经网络(RNN)和Transformers。大多数都是端到端的解决方案。在Whisper模型的情况下,所提出的方法显示了一个两阶段的解决方案,端到端ASR和LLM在管道中一起工作。ASR输出是LLM的输入。LLM是校正和改善ASR输出的组件。对现代端到端深度学习架构和ASR混合模型之间的波兰语自动识别进行了比较研究。医学访谈测试使用两种最先进的ASR模型进行:与LLM和Scribe ElevenLabs合并的OpenAI Whisper。此外,还将结果与Mozilla Common Voice和VoxPopuli数据库上的五个端到端模型(QuartzNet,FastConformer,Wav2Vec 2.0 XLSR和ESPnet Model Zoo)进行了比较。对干净的音频信号、带宽受限的信号和降级的信号进行了测试。测试模型进行了评估的基础上的字错误率(WER)和字符错误率(CER)。结果表明,Whisper模型在开源模型中表现最好。另一方面,ElevenLabs Scribe模型在一般基准和医疗数据上对波兰语表现最好。
摘要:This article concerns comparative studies on the Automatic Speech Recognition (ASR) model incorporated with the Large Language Model (LLM) used for medical interviews. The proposed solution is tested on polish language benchmarks and dataset with medical interviews. The latest ASR technologies are based on convolutional neural networks (CNNs), recurrent neural networks (RNNs) and Transformers. Most of them work as end-to-end solutions. The presented approach in the case of the Whisper model shows a two-stage solution with End-To-End ASR and LLM working together in a pipeline. The ASR output is an input for LLM. The LLM is a component by which the output from ASR is corrected and improved. Comparative studies for automatic recognition of the Polish language between modern End-To-End deep learning architectures and the ASR hybrid model were performed. The medical interview tests were performed with two state-of-the-art ASR models: OpenAI Whisper incorporated with LLM and Scribe ElevenLabs. Additionally, the results were compared with five more end-to-end models (QuartzNet, FastConformer, Wav2Vec 2.0 XLSR and ESPnet Model Zoo) on Mozilla Common Voice and VoxPopuli databases. Tests were conducted for clean audio signal, signal with bandwidth limitation, and degraded. The tested models were evaluated on the basis of Word Error Rate (WER) and Character Error Rate (CER). The results show that the Whisper model performs by far the best among the open-source models. ElevenLabs Scribe model, on the other hand, performs best for Polish on both general benchmark and medical data.
【11】LMU-Based Sequential Learning and Posterior Ensemble Fusion for Cross-Domain Infant Cry Classification
标题:基于LMU的序列学习和后验集融合用于跨领域婴儿哭声分类
链接:https://arxiv.org/abs/2603.02245
备注:7 pages
摘要:由于短的非平稳信号、有限的注释以及婴儿和数据集之间的强域偏移,解码婴儿哭泣原因对于医疗保健监测仍然具有挑战性。我们提出了一个紧凑的声学框架,融合MFCC,STFT和音高功能的多分支CNN编码器和模型的时间动态使用增强的勒让德记忆单元(LMU)。与LSTM相比,LMU主干提供稳定的序列建模,具有显著更少的重复参数,支持高效部署。为了提高跨数据集的泛化能力,我们引入了具有熵门控加权的校准后验集成融合,以保留特定领域的专业知识,同时减轻数据集偏差。在Baby2020和Baby Crying上的实验证明了跨域评估下改进的宏F1,以及泄漏软件分裂和实时设备监控的可行性。
摘要:Decoding infant cry causes remains challenging for healthcare monitoring due to short nonstationary signals, limited annotations, and strong domain shifts across infants and datasets. We propose a compact acoustic framework that fuses MFCC, STFT, and pitch features within a multi-branch CNN encoder and models temporal dynamics using an enhanced Legendre Memory Unit (LMU). Compared to LSTMs, the LMU backbone provides stable sequence modeling with substantially fewer recurrent parameters, supporting efficient deployment. To improve cross-dataset generalization, we introduce calibrated posterior ensemble fusion with entropy-gated weighting to preserve domain-specific expertise while mitigating dataset bias. Experiments on Baby2020 and Baby Crying demonstrate improved macro-F1 under cross-domain evaluation, along with leakageaware splits and real-time feasibility for on-device monitoring.
【12】Differentiable Time-Varying IIR Filtering for Real-Time Speech Denoising
标题:用于实时语音去噪的可区分时变IRR过滤
链接:https://arxiv.org/abs/2603.02794
备注:Submitted to Interspeech 2026
摘要:我们提出了TVF(时变滤波),一个低延迟的语音增强模型与1万个参数。TVF将数字信号处理(DSP)的可解释性与深度学习的适应性相结合,弥合了传统滤波与现代神经语音建模之间的差距。该模型利用一个轻量级的神经网络骨干,实时预测可微分35波段IIR滤波器级联的系数,使其能够动态适应非平稳噪声。与“黑盒”深度学习方法不同,TVF提供了一个完全可解释的处理链,其中频谱修改是明确的和可调整的。我们使用Valentini-Botinhao数据集证明了这种方法在语音去噪任务中的有效性,并将结果与静态DDSP方法和完全基于深度学习的解决方案进行了比较,表明TVF能够有效地适应不断变化的噪声条件。
摘要:We present TVF (Time-Varying Filtering), a low-latency speech enhancement model with 1 million parameters. Combining the interpretability of Digital Signal Processing (DSP) with the adaptability of deep learning, TVF bridges the gap between traditional filtering and modern neural speech modeling. The model utilizes a lightweight neural network backbone to predict the coefficients of a differentiable 35-band IIR filter cascade in real time, allowing it to adapt dynamically to non-stationary noise. Unlike ``black-box'' deep learning approaches, TVF offers a completely interpretable processing chain, where spectral modifications are explicit and adjustable. We demonstrate the efficacy of this approach on a speech denoising task using the Valentini-Botinhao dataset and compare the results to a static DDSP approach and a fully deep-learning-based solution, showing that TVF achieves effective adaptation to changing noise conditions.
【13】MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
标题:MUSE:一个以运行为中心的平台,用于大型语言模型的多模式统一安全评估
链接:https://arxiv.org/abs/2603.02482
备注:Submitted to ACL 2026 System Demonstration Track
摘要:大型语言模型的安全评估和红队仍然主要以文本为中心,现有的框架缺乏系统地测试对齐是否适用于音频、图像和视频输入的基础设施。我们提出了MUSE(多模态统一安全评估),一个开源的,以运行为中心的平台,集成了自动跨模态有效载荷生成,三个多回合攻击算法(Crescendo,PAIR,Violent Durian),提供商不可知的模型路由,和一个LLM判断与五级安全分类到一个基于浏览器的系统。双指标框架将硬攻击成功率(仅合规性)与软ASR(包括部分合规性)区分开来,捕获二进制指标遗漏的部分信息泄漏。为了探索对齐是否跨越模态边界,我们引入了匝间模态切换(ITMS),它通过每匝模态旋转来增强多匝攻击。来自四家供应商的六个多模态LLM的实验表明,多转向策略可以实现高达90-100%的ASR,而不是接近完美的单转向拒绝模型。ITMS不会在已经饱和的基线上统一提高最终ASR,但会通过破坏早期防御来加速收敛,并且消融揭示了模态效应的方向是特定于模型系列而不是通用的,这强调了提供商感知的跨模态安全测试的必要性。
摘要:Safety evaluation and red-teaming of large language models remain predominantly text-centric, and existing frameworks lack the infrastructure to systematically test whether alignment generalizes to audio, image, and video inputs. We present MUSE (Multimodal Unified Safety Evaluation), an open-source, run-centric platform that integrates automatic cross-modal payload generation, three multi-turn attack algorithms (Crescendo, PAIR, Violent Durian), provider-agnostic model routing, and an LLM judge with a five-level safety taxonomy into a single browser-based system. A dual-metric framework distinguishes hard Attack Success Rate (Compliance only) from soft ASR (including Partial Compliance), capturing partial information leakage that binary metrics miss. To probe whether alignment generalizes across modality boundaries, we introduce Inter-Turn Modality Switching (ITMS), which augments multi-turn attacks with per-turn modality rotation. Experiments across six multimodal LLMs from four providers show that multi-turn strategies can achieve up to 90-100% ASR against models with near-perfect single-turn refusal. ITMS does not uniformly raise final ASR on already-saturated baselines, but accelerates convergence by destabilizing early-turn defenses, and ablation reveals that the direction of modality effects is model-family-specific rather than universal, underscoring the need for provider-aware cross-modal safety testing.
【14】When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus
标题:当欺骗检测器旅行时:低资源语言欺骗数据库中66种语言的评估
链接:https://arxiv.org/abs/2603.02364
备注:This paper has been submitted to Interspeech 2026 for review
摘要:我们介绍LRLspoof,这是一个用于跨语言欺骗检测的大规模多语言合成语音语料库,包含2,732小时的音频,由24个开源TTS系统生成,涵盖66种语言,其中包括我们操作定义下的45种低资源语言。为了在不需要目标域善意语音的情况下评估鲁棒性,我们使用阈值转移对11个公开可用的对策进行基准测试:对于每个模型,我们在合并的外部基准上校准EER操作点,并应用所产生的阈值,报告欺骗拒绝率(SRR)。结果表明,依赖于模型的跨语言的差异,欺骗拒绝显着不同的语言,即使在受控条件下,突出显示语言作为一个独立的源域转移欺骗检测。该数据集可在\href{https://huggingface.co/martets/MTUCI/LRLspoof}{\textbf{\underline{\textit{HuggingFace}和\href{https://modelscope.cn/martets/lab260/LRLspoof}{\textbf{\underline{\textit{ModelScope}上公开获取
摘要:We introduce LRLspoof, a large-scale multilingual synthetic-speech corpus for cross-lingual spoof detection, comprising 2,732 hours of audio generated with 24 open-source TTS systems across 66 languages, including 45 low-resource languages under our operational definition. To evaluate robustness without requiring target-domain bonafide speech, we benchmark 11 publicly available countermeasures using threshold transfer: for each model we calibrate an EER operating point on pooled external benchmarks and apply the resulting threshold, reporting spoof rejection rate (SRR). Results show model-dependent cross-lingual disparity, with spoof rejection varying markedly across languages even under controlled conditions, highlighting language as an independent source of domain shift in spoof detection. The dataset is publicly available at \href{https://huggingface.co/datasets/MTUCI/LRLspoof}{\textbf{\underline{\textit{HuggingFace}}}} and \href{https://modelscope.cn/datasets/lab260/LRLspoof}{\textbf{\underline{\textit{ModelScope}}}}
【15】Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
标题:语音识别中的序列级无监督训练:理论研究
链接:https://arxiv.org/abs/2603.02285
备注:accepted to ICASSP 2026
摘要:无监督语音识别是一项利用非配对数据训练语音识别模型的任务。为了确定无监督语音识别何时以及如何成功,以及分类错误如何与候选训练目标相关,我们开发了一个基于分类错误界限的无监督语音识别理论框架。我们介绍了两个条件下,无监督语音识别是可能的。并对这些条件的必要性进行了讨论。在这些条件下,我们推导出无监督语音识别的分类误差界,并在模拟中验证了这一界。受此限制,我们提出了一个单阶段的序列级交叉熵损失的无监督语音识别。
摘要:Unsupervised speech recognition is a task of training a speech recognition model with unpaired data. To determine when and how unsupervised speech recognition can succeed, and how classification error relates to candidate training objectives, we develop a theoretical framework for unsupervised speech recognition grounded in classification error bounds. We introduce two conditions under which unsupervised speech recognition is possible. The necessity of these conditions are also discussed. Under these conditions, we derive a classification error bound for unsupervised speech recognition and validate this bound in simulations. Motivated by this bound, we propose a single-stage sequence-level cross-entropy loss for unsupervised speech recognition.
【16】When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning
标题:当缩放失败时:通过多步感知推理缓解LALM的音频感知衰退
链接:https://arxiv.org/abs/2603.02266
备注:Under Review
摘要:测试时间缩放在通过缩放推理计算来解决复杂问题方面显示出显著的效果。然而,在大型音频语言模型(LALM)中,存在一种不直观的现象:与直接回答的后训练相比,结构化推理轨迹的后训练模型会产生边际甚至负增益。为了研究它,我们引入CAFE,一个评估框架,旨在精确量化音频推理错误。评估结果显示,LALM在推理过程中与感知斗争,并遇到了一个关键的瓶颈:推理性能随着推理长度的延长而受到音频感知衰减的影响。为了解决这个问题,我们提出了MPAR $^2 $,一个范式,鼓励动态感知推理和分解成感知丰富的子问题的复杂问题。利用强化学习,MPAR $^2 $将CAFE的感知性能从31.74%提高到63.51%,并有效地减轻了感知衰减,同时增强了推理能力,在MMAU基准测试中达到了74.59%的准确率。进一步的分析表明,MPAR $^2 $加强LALM出席音频输入,并动态地适应推理预算,以匹配任务的复杂性。
摘要:Test-Time Scaling has shown notable efficacy in addressing complex problems through scaling inference compute. However, within Large Audio-Language Models (LALMs), an unintuitive phenomenon exists: post-training models for structured reasoning trajectories results in marginal or even negative gains compared to post-training for direct answering. To investigate it, we introduce CAFE, an evaluation framework designed to precisely quantify audio reasoning errors. Evaluation results reveal LALMs struggle with perception during reasoning and encounter a critical bottleneck: reasoning performance suffers from audio perception decay as reasoning length extends. To address it, we propose MPAR$^2$, a paradigm that encourages dynamic perceptual reasoning and decomposes complex questions into perception-rich sub-problems. Leveraging reinforcement learning, MPAR$^2$ improves perception performance on CAFE from 31.74% to 63.51% and effectively mitigates perception decay, concurrently enhancing reasoning capabilities to achieve a significant 74.59% accuracy on the MMAU benchmark. Further analysis demonstrates that MPAR$^2$ reinforces LALMs to attend to audio input and dynamically adapts reasoning budget to match task complexity.
【17】MEBM-Speech: Multi-scale Enhanced BrainMagic for Robust MEG Speech Detection
标题:MEBM-Speech:用于稳健MEG语音检测的多尺度增强BrainMagic
链接:https://arxiv.org/abs/2603.02255
备注:5 pages, 1 figure. To appear in the PNPL Competition Workshop at NeurIPS 2025
摘要:我们提出MEBM语音,多尺度增强神经解码器的语音活动检测从非侵入性脑磁图(MEG)信号。MEBM-Speech建立在BrainMagic主干上,集成了三种互补的时间建模机制:用于短期模式提取的多尺度卷积模块,用于长期上下文建模的双向LSTM(BiLSTM),以及用于有效跨尺度特征融合的深度可分离卷积层。轻量级的时间抖动策略和平均池化进一步提高了起始鲁棒性和边界稳定性。该模型对MEG信号进行连续概率解码,从而实现语音与沉默状态的细粒度检测-这是认知神经科学和临床应用的关键能力。LibriBrain Competition 2025 Track 1基准测试的综合评估显示了强劲的性能,在验证集上实现了89.3%的平均F1宏,并在官方测试排行榜上获得了可比的结果。这些发现突出了多尺度时间表示学习的有效性,鲁棒的基于MEG的语音解码。
摘要:We propose MEBM-Speech, a multi-scale enhanced neural decoder for speech activity detection from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, MEBM-Speech integrates three complementary temporal modeling mechanisms: a multi-scale convolutional module for short-term pattern extraction, a bidirectional LSTM (BiLSTM) for long-range context modeling, and a depthwise separable convolutional layer for efficient cross-scale feature fusion. A lightweight temporal jittering strategy and average pooling further improve onset robustness and boundary stability. The model performs continuous probabilistic decoding of MEG signals, enabling fine-grained detection of speech versus silence states - an ability crucial for both cognitive neuroscience and clinical applications. Comprehensive evaluations on the LibriBrain Competition 2025 Track1 benchmark demonstrate strong performance, achieving an average F1 macro of 89.3% on the validation set and comparable results on the official test leaderboard. These findings highlight the effectiveness of multi-scale temporal representation learning for robust MEG-based speech decoding.
【18】MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification
标题:MEBM音素:用于端到端MEG音素分类的多尺度增强BrainMagic
链接:https://arxiv.org/abs/2603.02254
备注:5 pages, 1 figure. To appear in the PNPL Competition Workshop at NeurIPS 2025
摘要:我们提出了MEBM-Phoneme,这是一种多尺度增强神经解码器,用于从非侵入性脑磁图(MEG)信号中进行音素分类。MEBM-Phoneme基于BrainMagic主干构建,集成了一个短期多尺度卷积模块,以增强本地中期编码器,并通过深度可分离卷积进行融合表示,以实现高效的跨尺度集成。卷积注意力层动态地对时间依赖性进行加权以细化特征聚合。为了解决类不平衡和会话特定的分布变化,我们引入了一个基于堆栈的本地验证集,以及加权交叉熵损失和随机时间增强。对LibriBrain Competition 2025 Track 2的全面评估显示出强大的泛化能力,在验证和官方测试排行榜上实现了具有竞争力的音素解码准确性。这些结果强调了分层时间建模和训练稳定性对于推进基于MEG的语音感知分析的价值。
摘要:We propose MEBM-Phoneme, a multi-scale enhanced neural decoder for phoneme classification from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, MEBM-Phoneme integrates a short-term multi-scale convolutional module to augment the native mid-term encoder, with fused representations via depthwise separable convolution for efficient cross-scale integration. A convolutional attention layer dynamically weights temporal dependencies to refine feature aggregation. To address class imbalance and session-specific distributional shifts, we introduce a stacking-based local validation set alongside weighted cross-entropy loss and random temporal augmentation. Comprehensive evaluations on LibriBrain Competition 2025 Track2 demonstrate robust generalization, achieving competitive phoneme decoding accuracy on the validation and official test leaderboard. These results underscore the value of hierarchical temporal modeling and training stabilization for advancing MEG-based speech perception analysis.
【19】SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models
标题:SGPA:多模式大型语言模型中可行的Shapley值解释的谱图引导语音对齐
链接:https://arxiv.org/abs/2603.02250
备注:Submitted for admission in Interspeech 2026 conference
摘要:通过Shapley值属性解释端到端音频语言模型的行为在原生标记化下是棘手的:典型的话语产生超过150 $$的编码器帧,相对于文本将联盟空间膨胀大约10 ^{42}$;单个音频帧缺乏独立意义;平分语音过渡的标记边界引入掩蔽伪像。我们引入了频谱图引导的语音对齐(SGPA),这是一个四阶段的流水线,它将联结主义时间分类强制对齐与频谱边界细化相结合,以产生声学稳定的单词对齐的音频片段。使用VoiceBench对LFM 2-Audio-1.5 B进行的受控诊断显示,SGPA在模型评估中产生43$\times$减少。统计测试证实,SGPA显着改变归因集中,同时保持全球累积配置文件,建立它作为一个可行性,使层音频可解释性。
摘要:Explaining the behavior of end-to-end audio language models via Shapley value attribution is intractable under native tokenization: a typical utterance yields over $150$ encoder frames, inflating the coalition space by roughly $10^{42}$ relative to text; individual audio frames lack standalone meaning; and token boundaries that bisect phonetic transitions introduce masking artifacts. We introduce Spectrogram-Guided Phonetic Alignment (SGPA), a four-stage pipeline that combines Connectionist Temporal Classification forced alignment with spectral boundary refinement to produce acoustically stable, word-aligned audio segments. Controlled diagnostics on LFM2-Audio-1.5B with VoiceBench show that SGPA yields a 43$\times$ reduction in model evaluations. Statistical testing confirms that SGPA significantly alters attribution concentration while preserving the global cumulative profile, establishing it as a feasibility-enabling layer for audio explainability.
机器翻译由腾讯交互翻译提供,仅供参考
