2026年1月27日的论文合集:cs.SD语音38篇,eess.AS音频处理33篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings
标题:低资源环境下尼泊尔神经多扬声器语音克隆
链接:https://arxiv.org/abs/2601.18694

作者:Aayush M. Shrestha,Aditya Bajracharya,Projan Shakya,Dinesh B. Kshatri
备注:16 pages with appendix included
摘要:本研究提出了一个Few-Shot语音克隆系统,尼泊尔扬声器,旨在合成语音在一个特定的扬声器的声音从梵文文本使用最少的数据。尼泊尔语的语音克隆由于其资源匮乏的性质,在很大程度上仍未得到探索。为了解决这个问题,我们构建了单独的数据集:用于训练扬声器编码器的未转录音频和用于训练基于Tacotron 2的合成器的配对文本音频数据。扬声器编码器,优化生成End2End损失,生成嵌入捕获扬声器的声音身份,通过均匀流形近似和投影(UMAP)降维可视化验证。这些嵌入与Tacotron 2的文本嵌入融合以产生mel频谱图,然后使用WaveRNN声码器将其转换为音频。音频数据从各种来源收集,包括自我录音,并进行了彻底的质量和对齐预处理。在多个超参数设置下使用mel和gate损失函数进行训练。该系统有效地克隆说话人的特征,即使是看不见的声音,证明了Few-Shot语音克隆的尼泊尔语的可行性,并建立了在低资源的情况下的个性化语音合成的基础。
摘要:This research presents a few-shot voice cloning system for Nepali speakers, designed to synthesize speech in a specific speaker's voice from Devanagari text using minimal data. Voice cloning in Nepali remains largely unexplored due to its low-resource nature. To address this, we constructed separate datasets: untranscribed audio for training a speaker encoder and paired text-audio data for training a Tacotron2-based synthesizer. The speaker encoder, optimized with Generative End2End loss, generates embeddings that capture the speaker's vocal identity, validated through Uniform Manifold Approximation and Projection (UMAP) for dimension reduction visualizations. These embeddings are fused with Tacotron2's text embeddings to produce mel-spectrograms, which are then converted into audio using a WaveRNN vocoder. Audio data were collected from various sources, including self-recordings, and underwent thorough preprocessing for quality and alignment. Training was performed using mel and gate loss functions under multiple hyperparameter settings. The system effectively clones speaker characteristics even for unseen voices, demonstrating the feasibility of few-shot voice cloning for the Nepali language and establishing a foundation for personalized speech synthesis in low-resource scenarios.


【2】Geneses: Unified Generative Speech Enhancement and Separation
标题:起源:统一生成语音增强和分离
链接:https://arxiv.org/abs/2601.18456

作者:Kohei Asai,Wataru Nakata,Yuki Saito,Hiroshi Saruwatari
备注:Accepted to ICASSP 2025 workshop
摘要:真实世界的音频记录通常包含多个扬声器和各种降级,这限制了可用于构建最先进的语音处理模型的语音数据的数量和质量。虽然连接语音增强(SE)和语音分离(SS)以获得每个说话者的干净语音信号的端到端方法是有希望的,但是传统的SE-SS方法遭受超过加性噪声的复杂劣化。为此,我们提出了\textbf{Geneses},这是一个生成框架,用于实现统一、高质量的SE-SS。我们的Geneses利用潜在流匹配来估计每个说话人的干净的语音特征,使用多模态扩散Transformer条件下的自监督学习表示从嘈杂的混合。我们使用LibriTTS-R的两个扬声器混合物在两种条件下进行实验评估:仅加性噪声和复杂的退化。结果表明,基因显着优于传统的掩模为基础的SE-SS方法在各种客观的指标,对复杂的退化具有很高的鲁棒性。音频样本可在我们的演示页面。
摘要:Real-world audio recordings often contain multiple speakers and various degradations, which limit both the quantity and quality of speech data available for building state-of-the-art speech processing models. Although end-to-end approaches that concatenate speech enhancement (SE) and speech separation (SS) to obtain a clean speech signal for each speaker are promising, conventional SE-SS methods suffer from complex degradations beyond additive noise. To this end, we propose \textbf{Geneses}, a generative framework to achieve unified, high-quality SE--SS. Our Geneses leverages latent flow matching to estimate each speaker's clean speech features using multi-modal diffusion Transformer conditioned on self-supervised learning representation from noisy mixture. We conduct experimental evaluation using two-speaker mixtures from LibriTTS-R under two conditions: additive-noise-only and complex degradations. The results demonstrate that Geneses significantly outperforms a conventional mask-based SE--SS method across various objective metrics with high robustness against complex degradations. Audio samples are available in our demo page.


【3】3DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action Control
标题:3DGesPolicy:基于动作控制的音素感知整体合声手势生成
链接:https://arxiv.org/abs/2601.18451

作者:Xuanmeng Sha,Liyun Zhang,Tomohiro Mashita,Naoya Chiba,Yuki Uranishi
备注:13 pages, 5 figures
摘要:由于现有的部分分解或帧级回归方法,生成整合全身运动与面部表情的整体协同语音手势存在身体运动和空间不稳定的无意义运动的语义不一致协调,我们引入了3DGesPolicy,一种新的基于动作的框架,通过机器人的扩散策略将整体手势生成重新表述为连续轨迹控制问题。通过将帧到帧的变化建模为统一的整体动作,我们的方法有效地学习帧间整体手势运动模式,并确保空间和语义上一致的运动轨迹符合现实的运动流形。为了进一步弥合表达对齐方面的差距,我们提出了一个手势-音频-音素(GAP)融合模块,该模块可以深度整合和细化多模态信号,确保语音语义,身体运动和面部表情之间的结构化和细粒度对齐。在BEAT 2数据集上进行的大量定量和定性实验证明了我们的3DGesPolicy在其他最先进的方法中生成自然,富有表现力和高度语音对齐的整体手势的有效性。
摘要:Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstable meaningless movements due to existing part-decomposed or frame-level regression methods, We introduce 3DGesPolicy, a novel action-based framework that reformulates holistic gesture generation as a continuous trajectory control problem through diffusion policy from robotics. By modeling frame-to-frame variations as unified holistic actions, our method effectively learns inter-frame holistic gesture motion patterns and ensures both spatially and semantically coherent movement trajectories that adhere to realistic motion manifolds. To further bridge the gap in expressive alignment, we propose a Gesture-Audio-Phoneme (GAP) fusion module that can deeply integrate and refine multi-modal signals, ensuring structured and fine-grained alignment between speech semantics, body motion, and facial expressions. Extensive quantitative and qualitative experiments on the BEAT2 dataset demonstrate the effectiveness of our 3DGesPolicy across other state-of-the-art methods in generating natural, expressive, and highly speech-aligned holistic gestures.


【4】UrgentMOS: Unified Multi-Metric and Preference Learning for Robust Speech Quality Assessment
标题:UrgentMOS:统一的多度量和偏好学习用于鲁棒语音质量评估
链接:https://arxiv.org/abs/2601.18438

作者:Wei Wang,Wangyou Zhang,Chenda Li,Jiahe Wang,Samuele Cornell,Marvin Sach,Kohei Saijo,Yihui Fu,Zhaoheng Ni,Bing Han,Xun Gong,Mengxiao Bi,Tim Fingscheidt,Shinji Watanabe,Yanmin Qian
摘要:随着现代语音生成系统的不断发展,自动语音质量评估变得越来越重要,而人类听力测试仍然昂贵,耗时且难以扩展。大多数现有的基于学习的评估模型主要依赖于稀缺的人类注释的平均意见得分(MOS)数据,这限制了鲁棒性和泛化能力,特别是在跨异构数据集进行训练时。在这项工作中,我们提出了UrgentMOS,一个统一的语音质量评估框架,共同学习不同的客观和感知质量指标,同时明确容忍在训练过程中没有任意子集的指标。通过在异构监督下利用互补的质量方面,UrgentMOS能够有效利用部分注释的数据,并在大规模多源数据集上训练时提高鲁棒性。除了绝对分数预测之外,UrgentMOS还通过直接预测比较MOS(CMOS)来显式地对成对质量偏好进行建模,使其非常适合系统基准测试中常用的基于偏好的评估场景。在广泛的语音质量数据集(包括模拟失真、语音增强和语音合成)中进行的广泛实验表明,UrgentMOS在绝对和比较评估设置中始终实现最先进的性能。
摘要:Automatic speech quality assessment has become increasingly important as modern speech generation systems continue to advance, while human listening tests remain costly, time-consuming, and difficult to scale. Most existing learning-based assessment models rely primarily on scarce human-annotated mean opinion score (MOS) data, which limits robustness and generalization, especially when training across heterogeneous datasets. In this work, we propose UrgentMOS, a unified speech quality assessment framework that jointly learns from diverse objective and perceptual quality metrics, while explicitly tolerating the absence of arbitrary subsets of metrics during training. By leveraging complementary quality facets under heterogeneous supervision, UrgentMOS enables effective utilization of partially annotated data and improves robustness when trained on large-scale, multi-source datasets. Beyond absolute score prediction, UrgentMOS explicitly models pairwise quality preferences by directly predicting comparative MOS (CMOS), making it well suited for preference-based evaluation scenarios commonly adopted in system benchmarking. Extensive experiments across a wide range of speech quality datasets, including simulated distortions, speech enhancement, and speech synthesis, demonstrate that UrgentMOS consistently achieves state-of-the-art performance in both absolute and comparative evaluation settings.


【5】Pisets: A Robust Speech Recognition System for Lectures and Interviews
标题:Pipets:用于讲座和面试的强大语音识别系统
链接:https://arxiv.org/abs/2601.18415

作者:Ivan Bondarenko,Daniil Grebenkin,Oleg Sedukhin,Mikhail Klementev,Roman Derunets,Lyudmila Budneva
摘要:这项工作为科学家和记者提供了一个语音到文本系统“Pisets”,该系统基于三组件架构,旨在提高语音识别的准确性,同时最大限度地减少与Whisper模型相关的错误和幻觉。该架构包括使用Wav2Vec2的初级识别,通过音频频谱图Transformer(AST)的假阳性过滤,以及通过Whisper的最终语音识别。课程学习方法的实施和多种俄语语音语料库的利用显着提高了系统的有效性。此外,还引入了先进的不确定性建模技术,有助于进一步提高转录质量。与WhisperX和通常的Whisper模型相比,所提出的方法确保了在各种声学条件下对长音频数据的稳健转录。“Pisets”系统的源代码可在GitHub上公开获取:https://github.com/bond005/pisets。
摘要:This work presents a speech-to-text system "Pisets" for scientists and journalists which is based on a three-component architecture aimed at improving speech recognition accuracy while minimizing errors and hallucinations associated with the Whisper model. The architecture comprises primary recognition using Wav2Vec2, false positive filtering via the Audio Spectrogram Transformer (AST), and final speech recognition through Whisper. The implementation of curriculum learning methods and the utilization of diverse Russian-language speech corpora significantly enhanced the system's effectiveness. Additionally, advanced uncertainty modeling techniques were introduced, contributing to further improvements in transcription quality. The proposed approaches ensure robust transcribing of long audio data across various acoustic conditions compared to WhisperX and the usual Whisper model. The source code of "Pisets" system is publicly available at GitHub: https://github.com/bond005/pisets.


【6】OCR-Enhanced Multimodal ASR Can Read While Listening
标题:OCR增强型多模式ASB可以边听边阅读
链接:https://arxiv.org/abs/2601.18393

作者:Junli Chen,Changli Tang,Yixuan Li,Guangzhi Sun,Chao Zhang
备注:4 pages, 2 figures. Submitted to ICASSP 2026
摘要:视觉信息(例如电影中的字幕)通常有助于自动语音识别。在本文中,我们提出了甜甜圈耳语,视听ASR模型与双编码器,利用视觉信息,以提高语音识别性能的英语和汉语。Donut-Whisper通过交叉注意模块结合了线性和基于Q-Former的模态对齐结构的优点,生成更强大的视听特征。同时,我们提出了一个轻量级的知识蒸馏计划展示了使用视听模型教音频模型,以实现更好的性能的潜力。此外,我们提出了一个新的多语种视听语音识别数据集的基础上的电影剪辑包含中文和英文的分区。因此,与Donut和Whisper大型V3基线相比,Donut-Whisper在数据集的英语和中文分区上都取得了显着更好的性能。特别是,与Whisper ASR基线相比,英语和中文集分别实现了绝对WER减少5.75%和绝对CER减少16.5%。
摘要:Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve speech recognition performance in both English and Chinese. Donut-Whisper combines the advantage of the linear and the Q-Former-based modality alignment structures via a cross-attention module, generating more powerful audio-visual features. Meanwhile, we propose a lightweight knowledge distillation scheme showcasing the potential of using audio-visual models to teach audio-only models to achieve better performance. Moreover, we propose a new multilingual audio-visual speech recognition dataset based on movie clips containing both Chinese and English partitions. As a result, Donut-Whisper achieved significantly better performance on both English and Chinese partition of the dataset compared to both Donut and Whisper large V3 baselines. In particular, an absolute 5.75% WER reduction and a 16.5% absolute CER reduction were achieved on the English and Chinese sets respectively compared to the Whisper ASR baseline.


【7】A Dataset for Automatic Vocal Mode Classification
标题:自动人声模式分类数据集
链接:https://arxiv.org/abs/2601.18339

作者:Reemt Hinrichs,Sonja Stephan,Alexander Lange,Jörn Ostermann
备注:Part of the proceedings of the EvoMUSART 2026: 15th International Conference on Artificial Intelligence in Music, Sound, Art and Design
摘要:完整的声乐技巧(CVT)是在过去的几十年里由凯瑟琳Sadolin等人发展起来的一个歌唱流派。CVT将声音的使用分为所谓的声音模式,即中性,抑制,超速和边缘。了解所需的发声模式可以帮助歌唱学生。因此,自动分类的声乐模式可以是重要的技术辅助歌唱教学。以前,已经尝试了语音模式的自动分类,但没有取得重大成功,这可能是由于缺乏数据。因此,我们记录了一个新的声乐模式数据集,包括从四个歌手,其中三个专业歌手超过五年的CVT经验记录的持续元音。该数据集涵盖了受试者的整个声音范围,总共有3,752个独特的样本。通过使用四个麦克风,从而提供自然的数据增强,数据集由超过13,000个样本组成。注释是使用三名经验丰富的CVT注释者创建的,每个注释者都提供一个单独的注释。合并的注释以及三个单独的注释随发布的数据集一起提供。此外,我们还提供了一些基线分类结果。ResNet18在5倍交叉验证中实现了81.3%的最佳平衡准确度。该数据集可以在https://zenodo.org/records/14276415下载。
摘要:The Complete Vocal Technique (CVT) is a school of singing developed in the past decades by Cathrin Sadolin et al.. CVT groups the use of the voice into so called vocal modes, namely Neutral, Curbing, Overdrive and Edge. Knowledge of the desired vocal mode can be helpful for singing students. Automatic classification of vocal modes can thus be important for technology-assisted singing teaching. Previously, automatic classification of vocal modes has been attempted without major success, potentially due to a lack of data. Therefore, we recorded a novel vocal mode dataset consisting of sustained vowels recorded from four singers, three of which professional singers with more than five years of CVT-experience. The dataset covers the entire vocal range of the subjects, totaling 3,752 unique samples. By using four microphones, thereby offering a natural data augmentation, the dataset consists of more than 13,000 samples combined. An annotation was created using three CVT-experienced annotators, each providing an individual annotation. The merged annotation as well as the three individual annotations come with the published dataset. Additionally, we provide some baseline classification results. The best balanced accuracy across a 5-fold cross validation of 81.3\,\% was achieved with a ResNet18. The dataset can be downloaded under https://zenodo.org/records/14276415.


【8】Analytic Incremental Learning For Sound Source Localization With Imbalance Rectification
标题:具有不平衡纠正的声音源定位的分析增量学习
链接:https://arxiv.org/abs/2601.18335

作者:Zexia Fan,Yu Chen,Qiquan Zhang,Kainan Chen,Xinyuan Qian
备注:Accepted by ICASSP26
摘要:声源定位(SSL)在受控环境中表现出显着的效果,但由于双重不平衡的挑战,在现实世界中的部署斗争:由长尾到达方向(DoA)分布引起的任务内不平衡,以及由跨任务偏斜和重叠引起的任务间不平衡。这些通常会导致灾难性的遗忘,显著降低定位精度。为了缓解这些问题,我们提出了一个统一的框架,有两个关键的创新。具体来说,我们设计了一种基于GCC-PHAT的数据增强(GDA)方法,该方法利用峰值特征来缓解任务内分布偏差。我们还提出了一个分析动态不平衡整流器(ANODC)与任务自适应正则化,使分析更新,适应任务间的动态。在SSLR基准测试中,我们的提案实现了89.0%准确度、5.3°平均绝对误差和1.6反向传输的最新(SoTA)结果,证明了在没有样本存储的情况下对不断变化的不平衡的鲁棒性。
摘要:Sound source localization (SSL) demonstrates remarkable results in controlled settings but struggles in real-world deployment due to dual imbalance challenges: intra-task imbalance arising from long-tailed direction-of-arrival (DoA) distributions, and inter-task imbalance induced by cross-task skews and overlaps. These often lead to catastrophic forgetting, significantly degrading the localization accuracy. To mitigate these issues, we propose a unified framework with two key innovations. Specifically, we design a GCC-PHAT-based data augmentation (GDA) method that leverages peak characteristics to alleviate intra-task distribution skews. We also propose an Analytic dynamic imbalance rectifier (ADIR) with task-adaption regularization, which enables analytic updates that adapt to inter-task dynamics. On the SSLR benchmark, our proposal achieves state-of-the-art (SoTA) results of 89.0% accuracy, 5.3° mean absolute error, and 1.6 backward transfer, demonstrating robustness to evolving imbalances without exemplar storage.


【9】Reflecting Twice before Speaking with Empathy: Self-Reflective Alternating Inference for Empathy-Aware End-to-End Spoken Dialogue
标题:带着同理心说话前反思两次:自我反思的交替推理以实现同理心的端到端口语对话
链接:https://arxiv.org/abs/2601.18281

作者:Yuhang Jia,Pei Liu,Haoqin Sun,Jiaming Zhou,Xuxin Cheng,Cao Liu,Ke Zeng,Xunliang Cai,Yong Qin
摘要:端到端的口语模型(SLM)具有很大的潜力,为非语言感知,许多研究旨在提高他们的能力,特别是移情对话。然而,目前的方法在很大程度上依赖于严格的监督信号,例如监督微调中的地面实况响应或强化学习中的偏好得分。这种依赖在模拟复杂的同理心时基本上是有限的,因为没有单一的“正确”反应,简单的数字分数不能完全捕捉情感表达的细微差别或同理心行为的适当性。为了解决这些局限性,我们依次介绍EmpathyEval,一个描述性的基于自然语言的评价模型,用于评估口语对话中的移情质量。在EmpathyEval的基础上,我们提出了ReEmpathy,一个端到端的SLM,它通过一种新的移情自我反思交替推理机制来增强移情对话,该机制将口语反应生成与自由形式的移情相关的反思推理交织在一起。大量的实验表明,ReEmpathy通过实现反思推理,大大提高了对同理心敏感的口语对话,为更情绪化的智能和同理心感知的人机交互提供了一种有前途的方法。
摘要:End-to-end Spoken Language Models (SLMs) hold great potential for paralinguistic perception, and numerous studies have aimed to enhance their capabilities, particularly for empathetic dialogue. However, current approaches largely depend on rigid supervised signals, such as ground-truth response in supervised fine-tuning or preference scores in reinforcement learning. Such reliance is fundamentally limited for modeling complex empathy, as there is no single "correct" response and a simple numerical score cannot fully capture the nuances of emotional expression or the appropriateness of empathetic behavior. To address these limitations, we sequentially introduce EmpathyEval, a descriptive natural-language-based evaluation model for assessing empathetic quality in spoken dialogues. Building upon EmpathyEval, we propose ReEmpathy, an end-to-end SLM that enhances empathetic dialogue through a novel Empathetic Self-Reflective Alternating Inference mechanism, which interleaves spoken response generation with free-form, empathy-related reflective reasoning. Extensive experiments demonstrate that ReEmpathy substantially improves empathy-sensitive spoken dialogue by enabling reflective reasoning, offering a promising approach toward more emotionally intelligent and empathy-aware human-computer interactions.


【10】LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
标题:LLM-ForcedAligner:一种基于非自回归和精确LLM-ForcedAligner的多语言和长格式语音强制对齐器
链接:https://arxiv.org/abs/2601.18220

作者:Bingshen Mu,Xian Shi,Xiong Wang,Hexin Liu,Jin Xu,Lei Xie
摘要:强制对齐(FA)预测语音中单词或字符的开始和结束时间戳,但现有的方法是特定于语言的,并且易于累积时间偏移。语音大语言模型(SLLM)的多语种语音理解和长序列处理能力,使其有望在多语种,跨语言,长格式语音设置FA。然而,直接将SLLM的下一个令牌预测范式应用于FA会导致幻觉和缓慢的推理。为了弥合这一差距,我们提出了LLM-ForcedAligner,将FA重新定义为一个插槽填充范式:时间戳被视为离散索引,特殊的时间戳标记作为插槽插入到成绩单中。SLLM以语音嵌入和带时隙的文本为条件,直接预测时隙处的时间索引。在训练期间,使用非移位输入和标签序列的因果注意力掩蔽允许每个时隙基于其自身和先前上下文预测其自己的时间戳索引,仅在时隙位置处计算损失。动态插槽插入使FA在任意位置。此外,支持非自回归推理,避免幻觉并提高速度。多语种、跨语种和长格式语音场景的实验表明,与现有方法相比,LLM-ForcedAligner实现了69%~78%的累积平均偏移相对减少。检查点和推理代码将在稍后发布。
摘要:Forced alignment (FA) predicts start and end timestamps for words or characters in speech, but existing methods are language-specific and prone to cumulative temporal shifts. The multilingual speech understanding and long-sequence processing abilities of speech large language models (SLLMs) make them promising for FA in multilingual, crosslingual, and long-form speech settings. However, directly applying the next-token prediction paradigm of SLLMs to FA results in hallucinations and slow inference. To bridge the gap, we propose LLM-ForcedAligner, reformulating FA as a slot-filling paradigm: timestamps are treated as discrete indices, and special timestamp tokens are inserted as slots into the transcript. Conditioned on the speech embeddings and the transcript with slots, the SLLM directly predicts the time indices at slots. During training, causal attention masking with non-shifted input and label sequences allows each slot to predict its own timestamp index based on itself and preceding context, with loss computed only at slot positions. Dynamic slot insertion enables FA at arbitrary positions. Moreover, non-autoregressive inference is supported, avoiding hallucinations and improving speed. Experiments across multilingual, crosslingual, and long-form speech scenarios show that LLM-ForcedAligner achieves a 69%~78% relative reduction in accumulated averaging shift compared with prior methods. The checkpoint and inference code will be released later.


【11】VIBEVOICE-ASR Technical Report
标题:VIBEVOICE-ASR技术报告
链接:https://arxiv.org/abs/2601.18184

作者:Zhiliang Peng,Jianwei Yu,Yaoyao Chang,Zilong Wang,Li Dong,Yingbo Hao,Yujie Tu,Chenyu Yang,Wenhui Wang,Songchen Xu,Yutao Sun,Hangbo Bao,Weijiang Xu,Yi Zhu,Zehua Wang,Ting Song,Yan Xia,Zewen Chi,Shaohan Huang,Liang Wang,Chuang Ding,Shuai Wang,Xie Chen,Furu Wei
摘要:本报告介绍了VibeVoice-ASR,这是一个基于VibeVoice的通用语音理解框架,旨在解决长格式音频中上下文碎片和多说话者复杂性的持续挑战(例如,会议,播客),尽管最近在短形式语音识别方面取得了进展,但仍然存在。与依赖音频分块的传统流水线方法不同,VibeVoice-ASR支持对长达60分钟的音频进行单次处理。它将自动语音识别、扬声器日志化和时间戳统一到一个端到端的生成任务中。此外,VibeVoice-ASR支持超过50种语言,不需要明确的语言设置,并且可以原生地处理话语内和话语间的代码切换。此外,我们引入了一个基于文本的上下文注入机制,允许用户提供定制的conetxt,显着提高特定领域的术语和多音字消歧的准确性。
摘要:This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation and multi-speaker complexity in long-form audio (e.g., meetings, podcasts) that remain despite recent advancements in short-form speech recognition. Unlike traditional pipelined approaches that rely on audio chunking, VibeVoice-ASRsupports single-pass processing for up to 60 minutes of audio. It unifies Automatic Speech Recognition, Speaker Diarization, and Timestamping into a single end-to-end generation task. In addition, VibeVoice-ASR supports over 50 languages, requires no explicit language setting, and natively handles code-switching within and across utterances. Furthermore, we introduce a prompt-based context injection mechanism that allows users to supply customized conetxt, significantly improving accuracy on domain-specific terminology and polyphonic character disambiguation.


【12】From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
标题:从人的语音到海洋信号:用于水声目标识别的语音大模型转换
链接:https://arxiv.org/abs/2601.18086

作者:Mengcheng Huang,Xue Zhou,Chen Xu,Dapeng Man
摘要:水声目标识别在海洋应用中起着重要的作用,但由于标记数据的有限性和海洋环境的复杂性,水声目标识别仍然具有挑战性。本文探讨了一个中心问题:语音大模型(SLM),训练在大量的人类语音语料库,可以有效地转移到水下声学?为了研究这一点,我们提出了UATR-SLM,一个简单的框架,重用语音特征管道,适应SLM作为声学编码器,并添加了一个轻量级的分类器。在DeepShip和ShipsEar基准测试上的实验表明,UATR-SLM实现了超过99%的域内准确率,在不同的信号长度上保持了强大的鲁棒性,并在跨域评估中达到了96.67%的准确率。这些结果突出了SLM到UATR的强大可转移性,建立了一个很有前途的范例,利用语音基础模型在水下声学。
摘要:Underwater acoustic target recognition (UATR) plays a vital role in marine applications but remains challenging due to limited labeled data and the complexity of ocean environments. This paper explores a central question: can speech large models (SLMs), trained on massive human speech corpora, be effectively transferred to underwater acoustics? To investigate this, we propose UATR-SLM, a simple framework that reuses the speech feature pipeline, adapts the SLM as an acoustic encoder, and adds a lightweight classifier.Experiments on the DeepShip and ShipsEar benchmarks show that UATR-SLM achieves over 99% in-domain accuracy, maintains strong robustness across variable signal lengths, and reaches up to 96.67% accuracy in cross-domain evaluation. These results highlight the strong transferability of SLMs to UATR, establishing a promising paradigm for leveraging speech foundation models in underwater acoustics.


【13】dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
标题:dLLM-ASB:一种基于LLM的更快扩散语音识别框架
链接:https://arxiv.org/abs/2601.17902

作者:Wenjie Tian,Bingshen Mu,Guobin Ma,Xuelong Geng,Zhixian Zhao,Lei Xie
摘要:基于大型语言模型(LLM)的自动语音识别(ASR)系统通过利用预训练的LLM作为解码器来实现卓越的性能,但其逐个令牌生成机制导致推理延迟随序列长度线性增长。与此同时,离散扩散大语言模型(dLLM)提供了一种有前途的替代方案,可以使用预训练的解码器生成高质量的并行序列。然而,直接应用本地面向文本的DLLM到ASR导致开放式文本生成和ASR所需的声学条件转录范式之间的根本不匹配。因此,它引入了不必要的困难和计算冗余,例如从纯噪声中去噪,不灵活的生成长度和固定的去噪步骤。我们提出了dLLM-ASR,一个有效的dLLM为基础的ASR框架,制定dLLM的解码作为一个事先指导和自适应去噪过程。它在初始化去噪过程之前利用ASR,并为序列长度提供锚点。在此基础上,长度自适应修剪动态删除冗余令牌,而基于置信度的去噪允许收敛令牌提前退出去噪循环,从而实现令牌级自适应计算。实验表明,dLLM-ASR实现的识别精度与基于自回归LLM的ASR系统相当,并提供了4.44$\times$的推理加速比,为ASR建立了一个实用而有效的范例。
摘要:Automatic speech recognition (ASR) systems based on large language models (LLMs) achieve superior performance by leveraging pretrained LLMs as decoders, but their token-by-token generation mechanism leads to inference latency that grows linearly with sequence length. Meanwhile, discrete diffusion large language models (dLLMs) offer a promising alternative, enabling high-quality parallel sequence generation with pretrained decoders. However, directly applying native text-oriented dLLMs to ASR leads to a fundamental mismatch between open-ended text generation and the acoustically conditioned transcription paradigm required by ASR. As a result, it introduces unnecessary difficulty and computational redundancy, such as denoising from pure noise, inflexible generation lengths, and fixed denoising steps. We propose dLLM-ASR, an efficient dLLM-based ASR framework that formulates dLLM's decoding as a prior-guided and adaptive denoising process. It leverages an ASR prior to initialize the denoising process and provide an anchor for sequence length. Building upon this prior, length-adaptive pruning dynamically removes redundant tokens, while confidence-based denoising allows converged tokens to exit the denoising loop early, enabling token-level adaptive computation. Experiments demonstrate that dLLM-ASR achieves recognition accuracy comparable to autoregressive LLM-based ASR systems and delivers a 4.44$\times$ inference speedup, establishing a practical and efficient paradigm for ASR.


【14】CaSNet: Compress-and-Send Network Based Multi-Device Speech Enhancement Model for Distributed Microphone Arrays
标题:CaSNet:基于压缩发送网络的分布式麦克风阵列多设备语音增强模型
链接:https://arxiv.org/abs/2601.17711

作者:Chengqian Jiang,Jie Zhang,Haoyin Yan
备注:this paper has been accept by ICASSP2026
摘要:分布式麦克风阵列(DMA)是下一代语音交互平台,但在噪声环境下仍需要语音增强来提高语音质量。现有的SE方法通常首先在融合中心(FC)从所有设备收集原始波形,然后设计多麦克风模型,这导致高带宽和能量成本。在这项工作中,我们为资源受限的DMA提出了\{压缩发送网络(CaSNet)},其中一个麦克风用作FC和参考。每个其他设备将测量的原始数据编码为特征矩阵,然后通过奇异值分解(SVD)压缩该特征矩阵以产生更紧凑的表示。在FC处接收的特征通过相对于参考的交叉窗口查询来对齐,随后进行神经解码以产生空间相干增强的语音。在多个数据集上的实验表明,与未压缩的情况相比,CaSNet可以节省数据量,对性能的影响可以忽略不计。可复制的代码可在https://github.com/Jokejiangv/CaSNet上获得。
摘要:Distributed microphone array (DMA) is a promising next-generation platform for speech interaction, where speech enhancement (SE) is still required to improve the speech quality in noisy cases. Existing SE methods usually first gather raw waveforms at a fusion center (FC) from all devices and then design a multi-microphone model, causing high bandwidth and energy costs. In this work, we propose a \emph{Compress-and-Send Network (CaSNet)} for resource-constrained DMAs, where one microphone serves as the FC and reference. Each of other devices encodes the measured raw data into a feature matrix, which is then compressed by singular value decomposition (SVD) to produce a more compact representation. The received features at the FC are aligned via cross window query with respect to the reference, followed by neural decoding to yield spatially coherent enhanced speech. Experiments on multiple datasets show that the proposed CaSNet can save the data amount with a negligible impact on the performance compared to the uncompressed case. The reproducible code is available at https://github.com/Jokejiangv/CaSNet.


【15】Segment Length Matters: A Study of Segment Lengths on Audio Fingerprinting Performance
标题:段长度很重要:段长度对音频指纹性能的研究
链接:https://arxiv.org/abs/2601.17690

作者:Ziling Gong,Yunyan Ouyang,Iram Kamdar,Melody Ma,Hongjie Chen,Franck Dernoncourt,Ryan A. Rossi,Nesreen K. Ahmed
摘要:音频指纹识别提供了声学信号的可识别表示,其可以稍后用于识别和检索系统。为了获得有区别的表示,输入音频通常被分割成较短的时间间隔,允许提取和分析局部声学特征。现代神经方法通常对短的、固定持续时间的音频片段进行操作,然而片段持续时间的选择通常是抽象的,很少深入研究。在本文中,我们研究了段长度如何影响音频指纹的性能。我们扩展了现有的神经指纹架构,以采用不同的段长度,并评估不同段长度和查询持续时间的检索准确性。我们的研究结果表明,较短的段长度(0.5秒)通常可以实现更好的性能。此外,我们评估了LLM在推荐最佳片段长度方面的能力,这表明GPT-5-mini在三个研究的LLM中,在五个考虑因素中始终提供最佳建议。我们的研究结果为大规模神经音频检索系统中片段持续时间的选择提供了实际指导。
摘要:Audio fingerprinting provides an identifiable representation of acoustic signals, which can be later used for identification and retrieval systems. To obtain a discriminative representation, the input audio is usually segmented into shorter time intervals, allowing local acoustic features to be extracted and analyzed. Modern neural approaches typically operate on short, fixed-duration audio segments, yet the choice of segment duration is often made heuristically and rarely examined in depth. In this paper, we study how segment length affects audio fingerprinting performance. We extend an existing neural fingerprinting architecture to adopt various segment lengths and evaluate retrieval accuracy across different segment lengths and query durations. Our results show that short segment lengths (0.5-second) generally achieve better performance. Moreover, we evaluate LLM capacity in recommending the best segment length, which shows that GPT-5-mini consistently gives the best suggestions across five considerations among three studied LLMs. Our findings provide practical guidance for selecting segment duration in large-scale neural audio retrieval systems.


【16】BanglaRobustNet: A Hybrid Denoising-Attention Architecture for Robust Bangla Speech Recognition
标题:BanglaRobustNet:一种用于稳健孟加拉语语音识别的混合去噪-注意架构
链接:https://arxiv.org/abs/2601.17679

作者:Md Sazzadul Islam Ridoy,Mubaswira Ibnat Zidney,Sumi Akter,Md. Aminur Rahman
摘要:孟加拉语是使用最广泛的语言之一,在最先进的自动语音识别(ASR)研究中仍然代表性不足,特别是在嘈杂和说话者多样化的条件下。本文介绍了BanglaRobustNet,这是一个基于Wav 2 Vec-BERT的混合去噪-注意力框架,旨在解决这些挑战。该架构集成了一个基于扩散的去噪模块,以抑制环境噪声,同时保留孟加拉语特定的语音线索,以及一个上下文交叉注意模块,该模块可以在说话人嵌入上进行识别,以实现跨性别,年龄和方言的鲁棒性。经过端到端的训练,结合了CTC损失,语音一致性和扬声器对齐的复合目标,与Wav 2 Vec-BERT和Whisper基线相比,BanglaRobustNet实现了字错误率(WER)和字符错误率(CER)的大幅降低。对Mozilla Common Voice Bangla和增强噪声语音的评估证实了我们方法的有效性,将BanglaRobustNet建立为针对低资源,易受噪声影响的语言环境的强大ASR系统。
摘要:Bangla, one of the most widely spoken languages, remains underrepresented in state-of-the-art automatic speech recognition (ASR) research, particularly under noisy and speaker-diverse conditions. This paper presents BanglaRobustNet, a hybrid denoising-attention framework built on Wav2Vec-BERT, designed to address these challenges. The architecture integrates a diffusion-based denoising module to suppress environmental noise while preserving Bangla-specific phonetic cues, and a contextual cross-attention module that conditions recognition on speaker embeddings for robustness across gender, age, and dialects. Trained end-to-end with a composite objective combining CTC loss, phonetic consistency, and speaker alignment, BanglaRobustNet achieves substantial reductions in word error rate (WER) and character error rate (CER) compared to Wav2Vec-BERT and Whisper baselines. Evaluations on Mozilla Common Voice Bangla and augmented noisy speech confirm the effectiveness of our approach, establishing BanglaRobustNet as a robust ASR system tailored to low-resource, noise-prone linguistic settings.


【17】AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
标题:AVMeme考试:针对法学硕士背景和文化知识和思维的多模式多语言多文化基准
链接:https://arxiv.org/abs/2601.17645

作者:Xilin Jiang,Qiaolin Wang,Junkai Wu,Xiaomin He,Zhongweiyang Xu,Yinghao Ma,Minshuo Piao,Kaiyi Yang,Xiuwen Zheng,Riki Shimizu,Yicong Chen,Arsalan Firoozi,Gavin Mischler,Sukru Samet Dindar,Richard Antonello,Linyang He,Tsun-An Hsieh,Xulin Fan,Yulun Wu,Yuesheng Ma,Chaitanya Amballa,Weixiong Chen,Jiarui Hai,Ruisi Li,Vishal Choudhari,Cong Han,Yinghao Aaron Li,Adeen Flinker,Mounya Elhilali,Emmanouil Benetos,Mark Hasegawa-Johnson,Romit Roy Choudhury,Nima Mesgarani
备注:avmemeexam.github.io/public
摘要:互联网视听剪辑通过随时间变化的声音和动作传达意义,这超出了文本本身所能代表的范围。为了检验人工智能模型是否能够在人类文化背景下理解这些信号,我们引入了AVMeme Exam,这是一个人工策划的基准测试,包含一千多个标志性的互联网声音和视频,包括语音、歌曲、音乐和音效。每个模因都配有一个独特的问答,评估从表面内容到上下文和情感到用法和世界知识的理解水平,以及原始年份,成绩单,摘要和敏感性等元数据。我们使用这个基准系统地评估了最先进的多模态大型语言模型(MLLM)以及人类参与者。我们的研究结果揭示了一个一致的局限性:目前的模型在无文本音乐和声音效果上表现不佳,并且与表面内容相比,很难在上下文和文化中进行思考。这些发现突出了与人类一致的多模态智能的一个关键差距,并呼吁模型能够超越他们所听到和看到的表面进行上下文和文化感知。项目页面:avmemeexam.github.io/public
摘要:Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand such signals in human cultural contexts, we introduce AVMeme Exam, a human-curated benchmark of over one thousand iconic Internet sounds and videos spanning speech, songs, music, and sound effects. Each meme is paired with a unique Q&A assessing levels of understanding from surface content to context and emotion to usage and world knowledge, along with metadata such as original year, transcript, summary, and sensitivity. We systematically evaluate state-of-the-art multimodal large language models (MLLMs) alongside human participants using this benchmark. Our results reveal a consistent limitation: current models perform poorly on textless music and sound effects, and struggle to think in context and in culture compared to surface content. These findings highlight a key gap in human-aligned multimodal intelligence and call for models that can perceive contextually and culturally beyond the surface of what they hear and see. Project page: avmemeexam.github.io/public


【18】Home Health System Deployment Experience for Geriatric Care Remote Monitoring
标题:老年护理远程监控的家庭健康系统部署经验
链接:https://arxiv.org/abs/2601.17608

作者:Dong Yoon Lee,Alyssa Weakley,Hui Wei,Daniel Cardona,Shijia Pan
摘要:为了支持就地养老,成年子女经常从远处照顾年迈的父母。这些非正式的护理人员需要即插即用的远程护理解决方案,以保护隐私,实现实时活动监控和直观的可操作信息。这篇简短的论文介绍了远程监控系统部署经验的三次迭代以及在老年4M框架(最重要的是心理状态,移动性和药物治疗)指导下硬件,建模和用户界面的迭代改进。一个LLM辅助解决方案的开发,以平衡用户体验(隐私保护,即插即用)和系统性能。
摘要:To support aging-in-place, adult children often provide care to their aging parents from a distance. These informal caregivers desire plug-and-play remote care solutions for privacy-preserving continuous monitoring that enabling real-time activity monitoring and intuitive, actionable information. This short paper presents insights from three iterations of deployment experience for remote monitoring system and the iterative improvement in hardware, modeling, and user interface guided by the Geriatric 4Ms framework (matters most, mentation, mobility, and medication). An LLM-assisted solution is developed to balance user experience (privacy-preserving, plug-and-play) and system performance.


【19】EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding
标题:EuleroDec:一种用于高效鲁棒音频编码的复值RVQ-VAE
链接:https://arxiv.org/abs/2601.17517

作者:Luca Cerovaz,Michele Mancusi,Emanuele Rodolà
备注:Accepted at ICASSP 2026
摘要:音频编解码器通过将PCM音频压缩到带宽友好的比特率,为离散音乐生成建模、音乐流和沉浸式媒体提供支持。最近的工作已经倾向于在谱域中进行处理;然而,谱图域通常与相位建模相斗争,相位建模自然是复值的。大多数频域神经编解码器要么忽略相位信息,要么将其编码为两个独立的实值通道,从而限制了空间保真度。这就需要引入对抗性鉴别器,以牺牲收敛速度和训练稳定性来补偿音频信号的不足表示能力。在这项工作中,我们引入了一个端到端的复值RVQ-VAE音频编解码器,它在整个分析-量化-合成流水线中保留了幅度-相位耦合,并去除了对抗性鉴别器和扩散后滤波器。在没有GANs或扩散的情况下,我们在域内匹配或超过更长时间的训练基线,并在相位相干性和波形保真度方面达到SOTA域外性能。与训练数十万步的标准基线相比,我们的模型将训练预算降低了一个数量级,在保持高感知质量的同时,计算效率明显提高。
摘要:Audio codecs power discrete music generative modelling, music streaming, and immersive media by shrinking PCM audio to bandwidth-friendly bitrates. Recent works have gravitated towards processing in the spectral domain; however, spectrogram domains typically struggle with phase modeling, which is naturally complex-valued. Most frequency-domain neural codecs either disregard phase information or encode it as two separate real-valued channels, limiting spatial fidelity. This entails the need to introduce adversarial discriminators at the expense of convergence speed and training stability to compensate for the inadequate representation power of the audio signal. In this work we introduce an end-to-end complex-valued RVQ-VAE audio codec that preserves magnitude-phase coupling across the entire analysis-quantization-synthesis pipeline and removes adversarial discriminators and diffusion post-filters. Without GANs or diffusion, we match or surpass much longer-trained baselines in-domain and reach SOTA out-of-domain performance on phase coherence and waveform fidelity. Compared to standard baselines that train for hundreds of thousands of steps, our model, which reduces the training budget by an order of magnitude, is markedly more compute-efficient while preserving high perceptual quality.


【20】Window Size Versus Accuracy Experiments in Voice Activity Detectors
标题:语音活动检测器中的窗口大小与准确性实验
链接:https://arxiv.org/abs/2601.17270

作者:Max McKinnon,Samir Khaki,Chandan KA Reddy,William Huang
摘要:语音活动检测(VAD)在实现语音识别等应用方面起着至关重要的作用。我们分析了窗口大小对三种VAD算法的准确性的影响:Silero,WebRTC和均方根(RMS)在一组不同的现实世界的数字音频流。我们还探讨了在每个VAD输出端上使用迟滞。研究结果为优化VAD系统提供了实际参考。Silero的性能明显优于WebRTC和RMS,滞后为WebRTC提供了优势。
摘要:Voice activity detection (VAD) plays a vital role in enabling applications such as speech recognition. We analyze the impact of window size on the accuracy of three VAD algorithms: Silero, WebRTC, and Root Mean Square (RMS) across a set of diverse real-world digital audio streams. We additionally explore the use of hysteresis on top of each VAD output. Our results offer practical references for optimizing VAD systems. Silero significantly outperforms WebRTC and RMS, and hysteresis provides a benefit for WebRTC.


【21】Sink or SWIM: Tackling Real-Time ASR at Scale
标题:水槽或游泳:大规模解决实时ZR问题
链接:https://arxiv.org/abs/2601.17097

作者:Federico Bruzzone,Walter Cazzola,Matteo Brancaleoni,Dario Pellegrino
备注:14 pages, 7 figures
摘要:实时自动语音识别系统越来越多地集成到交互式应用中,从语音助理到实时转录服务。然而,扩展这些系统以支持多个并发客户端,同时保持低延迟和高准确性仍然是一个重大挑战。在这项工作中,我们提出了SWIM,这是一种建立在OpenAI Whisper模型之上的新型实时ASR系统,可以实现真正的模型级并行化,以实现可扩展的多语言转录。SWIM支持多个并发音频流,而无需修改底层模型。它引入了一种缓冲区合并策略,在确保高效资源使用的同时保持转录保真度。我们评估SWIM在多客户端设置-扩展到20个并发用户-并表明,它提供了准确的实时transmittance在英语,意大利语和西班牙语,同时保持低延迟和高吞吐量。虽然Whisper-Streaming在单客户端、仅英语设置中实现了约8.2%的单词错误率和约3.4秒的平均延迟,但SWIM将此功能扩展到多语言、多客户端环境。它保持了相当的准确性和显著更低的延迟(5个客户端约2.4秒),并继续有效地扩展到20个并发客户端,而不会降低转录质量和提高整体吞吐量。我们的方法通过提高动态多用户环境中的鲁棒性和效率来推进可扩展的ASR。
摘要:Real-time automatic speech recognition systems are increasingly integrated into interactive applications, from voice assistants to live transcription services. However, scaling these systems to support multiple concurrent clients while maintaining low latency and high accuracy remains a major challenge. In this work, we present SWIM, a novel real-time ASR system built on top of OpenAI's Whisper model that enables true model-level parallelization for scalable, multilingual transcription. SWIM supports multiple concurrent audio streams without modifying the underlying model. It introduces a buffer merging strategy that maintains transcription fidelity while ensuring efficient resource usage. We evaluate SWIM in multi-client settings -- scaling up to 20 concurrent users -- and show that it delivers accurate real-time transcriptions in English, Italian, and Spanish, while maintaining low latency and high throughput. While Whisper-Streaming achieves a word error rate of approximately 8.2% with an average delay of approximately 3.4 s in a single-client, English-only setting, SWIM extends this capability to multilingual, multi-client environments. It maintains comparable accuracy with significantly lower delay -- around 2.4 s with 5 clients -- and continues to scale effectively up to 20 concurrent clients without degrading transcription quality and increasing overall throughput. Our approach advances scalable ASR by improving robustness and efficiency in dynamic, multi-user environments.


【22】SonoEdit: Null-Space Constrained Knowledge Editing for Pronunciation Correction in LLM-Based TTS
标题:SonoEdit:零空间约束知识编辑,用于基于LLM的TTC中的发音纠正
链接:https://arxiv.org/abs/2601.17086

作者:Ayush Pratap Singh,Harshit Singh,Nityanand Mathur,Akshat Mandloi,Sudarshan Kamath
摘要:神经文本到语音(TTS)系统系统地错误发音低资源专有名词,特别是非英语名称,品牌和地理位置,由于他们在主要是英语培训语料库的代表性不足。现有的解决方案通常依赖于昂贵的多语言数据收集、监督微调或手动语音注释,这限制了TTS系统在语言多样性环境中的部署。我们介绍SonoEdit,一种模型编辑技术,它可以在不进行再训练的情况下外科手术地纠正预训练TTS模型中的发音错误。而不是昂贵的微调或显式音素注入,我们提出了一种基于零空间发音编辑的简约替代方案,它执行单次参数更新以修改特定单词的发音,同时可证明保留所有其他模型行为。我们首先适应声学因果跟踪,以确定负责文本到发音映射的Transformer层。然后,我们应用零空间约束编辑来计算一个封闭形式的权重更新,该权重更新纠正目标发音,同时保持与控制一般语音生成的子空间在数学上正交。这种受约束的更新将模型的声学输出转向所需的发音样本,同时保证在保留的语音语料库上的零一阶变化。
摘要:Neural text-to-speech (TTS) systems systematically mispronounce low-resource proper nouns, particularly non-English names, brands, and geographic locations, due to their underrepresentation in predominantly English training corpora. Existing solutions typically rely on expensive multilingual data collection, supervised finetuning, or manual phonetic annotation, which limits the deployment of TTS systems in linguistically diverse settings. We introduce SonoEdit, a model editing technique that surgically corrects pronunciation errors in pre-trained TTS models without retraining. Instead of costly finetuning or explicit phoneme injection, we propose a parsimonious alternative based on Null-Space Pronunciation Editing, which performs a single-shot parameter update to modify the pronunciation of specific words while provably preserving all other model behavior. We first adapt Acoustic Causal Tracing to identify the Transformer layers responsible for text-to-pronunciation mapping. We then apply Null-Space Constrained Editing to compute a closed-form weight update that corrects the target pronunciation while remaining mathematically orthogonal to the subspace governing general speech generation. This constrained update steers the model's acoustic output toward a desired pronunciation exemplar while guaranteeing zero first-order change on a preserved speech corpus.


【23】Audio Inpainting in Time-Frequency Domain with Phase-Aware Prior
标题:具有相感知先验的时频域音频修复
链接:https://arxiv.org/abs/2601.18535

作者:Peter Balušík,Pavel Rajmic
备注:submitted to IEEE Transactions on Audio, Speech and Language Processing
摘要:时域中所谓的音频修复问题是指估计信号内样本的缺失片段。多年来,已经开发了几种用于这种类型的音频修复的方法。与这种情况相反,在文献中出现了修复的时频变体,其中的挑战是用可靠的信息重建丢失的频谱图列。我们提出了一种方法来解决这个时频音频修复问题。我们的方法是基于最近推出的相位感知信号之前,利用瞬时频率的估计。使用广义Chambolle-Pock算法制定和解决优化问题。该方法与其他时频修复方法进行了客观和主观评估,特别是深度先验神经网络和基于自回归的方法Janssen-TF。我们提出的方法超越了这些方法在客观评价,以及在进行听力测试。此外,与替代方法相比,该结果是以大幅降低的计算要求实现的。
摘要:The so-called audio inpainting problem in the time domain refers to estimating missing segments of samples within a signal. Over the years, several methods have been developed for such type of audio inpainting. In contrast to this case, a time-frequency variant of inpainting appeared in the literature, where the challenge is to reconstruct missing spectrogram columns with reliable information. We propose a method to address this time-frequency audio inpainting problem. Our approach is based on the recently introduced phase-aware signal prior that exploits an estimate of the instantaneous frequency. An optimization problem is formulated and solved using the generalized Chambolle-Pock algorithm. The proposed method is evaluated both objectively and subjectively against other time-frequency inpainting methods, specifically a deep-prior neural network and the autoregression-based approach known as Janssen-TF. Our proposed approach surpassed these methods in the objective evaluation as well as in the conducted listening test. Moreover, this outcome is achieved with a substantially reduced computational requirement compared to alternative methods.


【24】Noise-Robust AV-ASR Using Visual Features Both in the Whisper Encoder and Decoder
标题:在Whisper编码器和解码器中使用视觉特征的噪音稳健AV-ASB
链接:https://arxiv.org/abs/2601.18396

作者:Zhengyang Li,Thomas Graave,Björn Möller,Zehang Wu,Matthias Franz,Tim Fingscheidt
备注:accepted at ICASSP2026
摘要:在视听自动语音识别(AV-ASR)系统中,预训练的语音识别中视觉特征的信息融合已被证明是提高噪声鲁棒性的一种有前途的方法。在这项工作中,基于突出的Whisper ASR,首先,我们提出了一种简单有效的视觉融合方法-在编码器和解码器(两用)中使用视觉特征-学习编码器中的视听交互并在解码器中加权模态。其次,我们比较了各种尺寸的Whisper模型中的视觉融合方法。我们提出的两用方法显示出一致的噪声鲁棒性改善,例如,与信噪比(SNR)为0 dB的多路重合噪声中的典型参考中间融合相比,基于Whisper small的35%相对改进(WER:4.41%对6.83%),以及基于Whisper medium的57%相对改进(WER:4.07%对9.53%)。第三,我们进行消融研究,检查各种模块设计和融合选项的影响。在1929小时的视听数据上进行微调后,我们使用Whisper介质的两用方法在各种SNR下实现了4.08%(MUSAN串音噪声)和4.43%(NoiseX串音噪声)的平均WER,从而在LRS 3 AV-ASR基准测试中建立了噪声条件下的最新技术水平。我们的代码在https://github.com/ifnspaml/Dual-Use-AVASR
摘要:In audiovisual automatic speech recognition (AV-ASR) systems, information fusion of visual features in a pre-trained ASR has been proven as a promising method to improve noise robustness. In this work, based on the prominent Whisper ASR, first, we propose a simple and effective visual fusion method -- use of visual features both in encoder and decoder (dual-use) -- to learn the audiovisual interactions in the encoder and to weigh modalities in the decoder. Second, we compare visual fusion methods in Whisper models of various sizes. Our proposed dual-use method shows consistent noise robustness improvement, e.g., a 35% relative improvement (WER: 4.41% vs. 6.83%) based on Whisper small, and a 57% relative improvement (WER: 4.07% vs. 9.53%) based on Whisper medium, compared to typical reference middle fusion in babble noise with a signal-to-noise ratio (SNR) of 0dB. Third, we conduct ablation studies examining the impact of various module designs and fusion options. Fine-tuned on 1929 hours of audiovisual data, our dual-use method using Whisper medium achieves 4.08% (MUSAN babble noise) and 4.43% (NoiseX babble noise) average WER across various SNRs, thereby establishing a new state-of-the-art in noisy conditions on the LRS3 AV-ASR benchmark. Our code is at https://github.com/ifnspaml/Dual-Use-AVASR


【25】Residual Learning for Neural Ambisonics Encoders
标题:神经立体声编码器的剩余学习
链接:https://arxiv.org/abs/2601.18322

作者:Thomas Deppisch,Yang Gao,Manan Mittal,Benjamin Stahl,Christoph Hold,David Alon,Zamir Ben-Hur
摘要:智能眼镜和延展实境耳机等新兴可穿戴设备需要从紧凑的头戴式麦克风阵列捕获高质量的空间音频。高保真度立体声复制通过将阵列信号映射到球面谐波(SH)系数来提供与设备无关的空间音频表示。然而,在实践中,准确的编码仍然具有挑战性。虽然传统的线性编码器是信号独立和鲁棒的,但它们会放大低频噪声并遭受高频空间混叠。另一方面,神经网络方法可以优于线性编码器,但它们通常假设理想化的麦克风,并且在现实世界的场景中可能表现不一致。为了利用它们的互补优势,我们引入了一个剩余学习框架,该框架通过神经网络的校正来改进线性编码器。使用从智能眼镜测量的阵列传递函数,我们比较了一个基于UNet的编码器从文献中的一个新的经常性注意力模型。我们的分析表明,这两个神经编码器只有在集成到残差学习框架中时才始终优于线性基线。在残差配置中,两个神经模型在域内数据的所有测试指标上都实现了一致且显著的改进,并在域外数据上实现了适度的增益。然而,相干性分析表明,所有的神经编码器配置继续与方向准确的高频编码斗争。
摘要:Emerging wearable devices such as smartglasses and extended reality headsets demand high-quality spatial audio capture from compact, head-worn microphone arrays. Ambisonics provides a device-agnostic spatial audio representation by mapping array signals to spherical harmonic (SH) coefficients. In practice, however, accurate encoding remains challenging. While traditional linear encoders are signal-independent and robust, they amplify low-frequency noise and suffer from high-frequency spatial aliasing. On the other hand, neural network approaches can outperform linear encoders but they often assume idealized microphones and may perform inconsistently in real-world scenarios. To leverage their complementary strengths, we introduce a residual-learning framework that refines a linear encoder with corrections from a neural network. Using measured array transfer functions from smartglasses, we compare a UNet-based encoder from the literature with a new recurrent attention model. Our analysis reveals that both neural encoders only consistently outperform the linear baseline when integrated within the residual learning framework. In the residual configuration, both neural models achieve consistent and significant improvements across all tested metrics for in-domain data and moderate gains for out-of-domain data. Yet, coherence analysis indicates that all neural encoder configurations continue to struggle with directionally accurate high-frequency encoding.


【26】Noise-Robust Contrastive Learning with an MFCC-Conformer For Coronary Artery Disease Detection
标题:用于冠状动脉疾病检测的MFCC一致性的噪音稳健对比学习
链接:https://arxiv.org/abs/2601.18295

作者:Milan Marocchi,Matthew Fynn,Yue Rong
备注:This paper has been accepted for presentation at ICASSP 2026. \c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses. 5 pages, 1 figure
摘要:心血管疾病(CVD)是全球死亡的主要原因,其中冠状动脉疾病(CAD)是CVD的最大亚类。最近,人们越来越关注使用心音图(PCG)信号检测CAD,在低噪声和最佳传感器放置的临床环境中取得了很大成功。已经发现多通道技术对噪声更鲁棒;然而,在真实世界数据上实现鲁棒性能仍然是一个挑战。这项工作利用了一种新的基于多通道能量的噪声段抑制算法,使用心脏和噪声参考麦克风,在训练深度学习分类器之前丢弃具有大量非平稳噪声的音频段。这种基于一致性的分类器从多个通道中提取梅尔频率倒谱系数(MFCC),进一步帮助提高模型的噪声鲁棒性。所提出的方法在297个受试者上实现了78.4%的准确率和78.2%的平衡准确率,与没有噪声段拒绝的训练相比,分别提高了4.1%和4.3%。
摘要:Cardiovascular diseases (CVD) are the leading cause of death worldwide, with coronary artery disease (CAD) comprising the largest subcategory of CVDs. Recently, there has been increased focus on detecting CAD using phonocardiogram (PCG) signals, with high success in clinical environments with low noise and optimal sensor placement. Multichannel techniques have been found to be more robust to noise; however, achieving robust performance on real-world data remains a challenge. This work utilises a novel multichannel energy-based noisy-segment rejection algorithm, using heart and noise-reference microphones, to discard audio segments with large amounts of nonstationary noise before training a deep learning classifier. This conformer-based classifier takes mel-frequency cepstral coefficients (MFCCs) from multiple channels, further helping improve the model's noise robustness. The proposed method achieved 78.4% accuracy and 78.2% balanced accuracy on 297 subjects, representing improvements of 4.1% and 4.3%, respectively, compared to training without noisy-segment rejection.


【27】Efficient Rehearsal for Continual Learning in ASR via Singular Value Tuning
标题:通过奇异值调整在ASB中进行连续学习的高效排练
链接:https://arxiv.org/abs/2601.18266

作者:Steven Vander Eeckt,Hugo Van hamme
备注:Accepted for publication in IEEE Transactions on Audio, Speech, and Language Processing
摘要:自动语音识别(ASR)中的持续学习(CL)在适应新任务、领域或说话人时会遭受灾难性遗忘。缓解这种情况的常见策略是将过去数据的子集存储在内存中以供排练。然而,基于排练的方法面临着关键的限制:存储数据通常成本高昂,预先训练的模型不可行,或者受到隐私法规的限制。使用较小的内存大小运行现有的基于排练的方法来缓解这些问题通常会导致性能下降。   我们提出了一个排练为基础的CL方法,即使在最小的内存仍然有效。它分为两个阶段:首先,对新任务进行微调;其次,将奇异值分解(SVD)应用于线性层中的变化,并且以参数有效的方式,仅重新训练奇异值上的门控向量,其控制使用排练接受来自第一阶段的更新的程度。我们广泛的测试和分析我们的方法在两个单语和两个多语种的基准。我们的方法减少了遗忘,并优于最先进的CL方法的ASR,即使是在限制为一个单一的话语每个先前的任务。
摘要:Continual Learning (CL) in Automatic Speech Recognition (ASR) suffers from catastrophic forgetting when adapting to new tasks, domains, or speakers. A common strategy to mitigate this is to store a subset of past data in memory for rehearsal. However, rehearsal-based methods face key limitations: storing data is often costly, infeasible with pre-trained models, or restricted by privacy regulations. Running existing rehearsal-based methods with smaller memory sizes to alleviate these issues usually leads to degraded performance.   We propose a rehearsal-based CL method that remains effective even with minimal memory. It operates in two stages: first, fine-tuning on the new task; second, applying Singular Value Decomposition (SVD) to the changes in linear layers and, in a parameter-efficient manner, retraining only gating vectors on the singular values, which control to extent to which updates from the first stage are accepted, using rehearsal. We extensively test and analyze our method on two monolingual and two multilingual benchmarks. Our method reduces forgetting and outperforms state-of-the-art CL approaches for ASR, even when limited to a single utterance per previous task.


【28】OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
标题:OneVoice:一种模式、三种场景--迈向统一的Zero-Shot语音转换
链接:https://arxiv.org/abs/2601.18094

作者:Zhichao Wang,Tao Li,Wenshuo Ge,Zihao Cui,Shilei Zhang,Junlan Feng
备注:Work in progress
摘要:语音转换技术的最新进展在说话人克隆和语言保存方面取得了新的里程碑。但该领域仍然支离破碎,依赖于专门的模型来实现语言保留、表达和歌唱场景。我们提出了OneVoice,一个统一的zero-shot框架,能够在一个模型中处理所有三种情况。OneVoice建立在一个连续的语言模型上,该模型使用无VAE的下一个补丁扩散进行训练,确保了高保真度和高效的序列建模。其统一的核心设计在于一个混合专家(MoE),旨在明确建模共享的转换知识和具体的表达能力。专家的选择是协调的双路径路由机制,包括共享的专家隔离和基于全局-局部线索的领域专家分配。为了精确的调节,特定的韵律特征通过门控机制融合到每一层中,允许韵律信息的自适应使用。此外,为了实现核心思想并缓解不平衡问题(丰富的语音与稀缺的歌声),我们采用了两阶段渐进式训练,包括基础预训练和基于LoRA的领域专家的场景增强。实验表明,OneVoice在所有三种场景下都匹配或超越了专业模型,同时验证了对场景的灵活控制,并提供了只需2步的快速解码版本。代码和模型将很快发布。
摘要:Recent progress of voice conversion~(VC) has achieved a new milestone in speaker cloning and linguistic preservation. But the field remains fragmented, relying on specialized models for linguistic-preserving, expressive, and singing scenarios. We propose OneVoice, a unified zero-shot framework capable of handling all three scenarios within a single model. OneVoice is built upon a continuous language model trained with VAE-free next-patch diffusion, ensuring high fidelity and efficient sequence modeling. Its core design for unification lies in a Mixture-of-Experts (MoE) designed to explicitly model shared conversion knowledge and scenario-specific expressivity. Expert selection is coordinated by a dual-path routing mechanism, including shared expert isolation and scenario-aware domain expert assignment with global-local cues. For precise conditioning, scenario-specific prosodic features are fused into each layer via a gated mechanism, allowing adaptive usage of prosody information. Furthermore, to enable the core idea and alleviate the imbalanced issue (abundant speech vs. scarce singing), we adopt a two-stage progressive training that includes foundational pre-training and scenario enhancement with LoRA-based domain experts. Experiments show that OneVoice matches or surpasses specialized models across all three scenarios, while verifying flexible control over scenarios and offering a fast decoding version as few as 2 steps. Code and model will be released soon.


【29】SpatialEmb: Extract and Encode Spatial Information for 1-Stage Multi-channel Multi-speaker ASR on Arbitrary Microphone Arrays
标题:SpatialEmb:提取和编码任意麦克风阵列上的1级多通道多扬声器ASB的空间信息
链接:https://arxiv.org/abs/2601.18037

作者:Yiwen Shao,Yong Xu,Sanjeev Khudanpur,Dong Yu
备注:SLT 2024
摘要:空间信息是多通道多说话人目标语音识别的重要线索。大多数最先进的多通道自动语音识别(ASR)系统只在语音分离阶段提取空间特征,然后对分离的语音进行标准的单通道ASR。由于预处理模块的累积错误,这种方法导致低效、冗长的流水线和次优的ASR性能。此外,大多数空间特征提取方法都依赖于扬声器位置和麦克风拓扑结构的知识,这使得系统依赖于特定设置,并且难以适应新设备。在这项工作中,我们提出了一个解决这些问题的轻量级嵌入模块SpatialEmb,提取和编码空间信息直接为ASR模型,支持固定和任意麦克风拓扑结构。我们进行全面的实验AliMeeting,一个真正的会议语料库,以确定最佳的模型设计SpatialEmb的性能和效率。我们用105小时Train-Ali-far训练的最佳模型在Eval和Test集上实现了17.04%和20.32%的字符错误率(CER),用相同的训练数据建立了一个新的最先进的结果。
摘要:Spatial information is a critical clue for multi-channel multi-speaker target speech recognition. Most state-of-the-art multi-channel Automatic Speech Recognition (ASR) systems extract spatial features only during the speech separation stage, followed by standard single-channel ASR on the separated speech. This approach results in an inefficient, lengthy pipeline and sub-optimal ASR performance due to the accumulated errors from preprocessing modules. Furthermore, most spatial feature extraction methods depend on the knowledge of speaker positions and microphone topology, making the systems reliant on specific settings and challenging to adapt to new equipment. In this work, we propose a solution to these issues with a lightweight embedding module named SpatialEmb, which extracts and encodes spatial information directly for the ASR model, supporting both fixed and arbitrary microphone topology. We conduct comprehensive experiments on AliMeeting, a real meeting corpus, to determine the optimal model design for SpatialEmb in terms of both performance and efficiency. Our best model trained with 105 hours Train-Ali-far achieves 17.04% and 20.32% character error rates (CER) on the Eval and Test sets, establishing a new state-of-the-art result with the same training data.


【30】AmbER$^2$: Dual Ambiguity-Aware Emotion Recognition Applied to Speech and Text
标题:AmbER $' 2 $:应用于语音和文本的双重模糊感知情感识别
链接:https://arxiv.org/abs/2601.18010

作者:Jingyao Wu,Grace Lin,Yinuo Song,Rosalind Picard
备注:Accepted in ICASSP 2026
摘要:情感识别本质上是模糊的,不确定性既来自评分者的分歧,也来自语音和文本等模态之间的差异。有越来越多的兴趣建模评分员模糊使用标签分布。然而,模态模糊性仍然未被充分研究,多模态方法通常依赖于简单的特征融合,而没有明确解决模态之间的冲突。在这项工作中,我们提出了AmbER$^2$,一个双重的模糊意识的框架,同时模型的评分水平和模态水平的模糊性,通过一个教师-学生架构与分布明智的培训目标。对IEMOCAP和MSP播客的评估表明,AmbER$^2$始终提高了传统交叉熵基线的分布保真度,并实现了与最近最先进的系统竞争或优于其的性能。例如,在IEMOCAP上,AmbER$^2$在Bhattacharyya系数上实现了20.3%的相对改进(0.83 vs. 0.69),在R$^2$上实现了13.6%的相对改进(0.67 vs. 0.59),在准确性上实现了3.8%的相对改进(0.683 vs. 0.658),在F1上实现了4.5%的相对改进(0.675 vs. 0.646)。跨模糊度水平的进一步分析表明,显式建模模糊度对于高度不确定的样本特别有益。这些发现强调了在建立强大的情感识别系统时,联合解决评分者和模态模糊的重要性。
摘要:Emotion recognition is inherently ambiguous, with uncertainty arising both from rater disagreement and from discrepancies across modalities such as speech and text. There is growing interest in modeling rater ambiguity using label distributions. However, modality ambiguity remains underexplored, and multimodal approaches often rely on simple feature fusion without explicitly addressing conflicts between modalities. In this work, we propose AmbER$^2$, a dual ambiguity-aware framework that simultaneously models rater-level and modality-level ambiguity through a teacher-student architecture with a distribution-wise training objective. Evaluations on IEMOCAP and MSP-Podcast show that AmbER$^2$ consistently improves distributional fidelity over conventional cross-entropy baselines and achieves performance competitive with, or superior to, recent state-of-the-art systems. For example, on IEMOCAP, AmbER$^2$ achieves relative improvements of 20.3% on Bhattacharyya coefficient (0.83 vs. 0.69), 13.6% on R$^2$ (0.67 vs. 0.59), 3.8% on accuracy (0.683 vs. 0.658), and 4.5% on F1 (0.675 vs. 0.646). Further analysis across ambiguity levels shows that explicitly modeling ambiguity is particularly beneficial for highly uncertain samples. These findings highlight the importance of jointly addressing rater and modality ambiguity when building robust emotion recognition systems.


【31】Speech Emotion Recognition with ASR Integration
标题:与ASB集成的语音情感识别
链接:https://arxiv.org/abs/2601.17901

作者:Yuanchao Li
备注:PhD Thesis
摘要:语音情感识别(SER)在理解人类交流、实现情感智能系统以及作为通用人工智能(AGI)发展的基本组成部分方面发挥着关键作用。然而,由于情感表达的复杂性以及当前语音和语言技术的局限性,在真实世界、自发和低资源场景中部署SER仍然是一个重大挑战。本文研究了自动语音识别(ASR)与SER的集成,旨在提高口语情感识别的鲁棒性、可扩展性和实用性。
摘要:Speech Emotion Recognition (SER) plays a pivotal role in understanding human communication, enabling emotionally intelligent systems, and serving as a fundamental component in the development of Artificial General Intelligence (AGI). However, deploying SER in real-world, spontaneous, and low-resource scenarios remains a significant challenge due to the complexity of emotional expression and the limitations of current speech and language technologies. This thesis investigates the integration of Automatic Speech Recognition (ASR) into SER, with the goal of enhancing the robustness, scalability, and practical applicability of emotion recognition from spoken language.


【32】End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
标题:具有儿童与成人互动的端到端联合ASB和说话者角色扩展
链接:https://arxiv.org/abs/2601.17640

作者:Anfeng Xu,Tiantian Feng,Somer Bishop,Catherine Lord,Shrikanth Narayanan
备注:Under review for IEEE
摘要:儿童与成人之间口语交流的准确转录和说话人日记化对于发展研究和临床研究至关重要。然而,手动注释既耗时又难以扩展。现有的自动化系统通常依赖于级联的说话人日志和语音识别管道,这可能导致错误传播。本文提出了一个统一的端到端的框架,扩展了耳语编码器-解码器架构,以联合建模ASR和儿童-成人扬声器角色日记。拟议的办法包括:(i)发出说话者标签和开始/结束时间戳的串行化输出训练方案,(ii)增强说话者区分编码器表示的轻量级帧级日志化头,(iii)用于改进的时间精度的日志化引导的静音抑制,以及(iv)保证结构上有效的输出的基于状态机的强制解码过程。对两个数据集的综合评估表明,在两个级联基线上有了一致和实质性的改进,实现了更低的多说话者单词错误率,并在Whisper-small和Whisper-large模型中展示了具有竞争力的日记准确性。这些研究结果突出的有效性和实际效用的建议联合建模框架,以产生可靠的,扬声器归因于转录的儿童与成人的互动规模。代码和模型权重是公开的
摘要:Accurate transcription and speaker diarization of child-adult spoken interactions are crucial for developmental and clinical research. However, manual annotation is time-consuming and challenging to scale. Existing automated systems typically rely on cascaded speaker diarization and speech recognition pipelines, which can lead to error propagation. This paper presents a unified end-to-end framework that extends the Whisper encoder-decoder architecture to jointly model ASR and child-adult speaker role diarization. The proposed approach integrates: (i) a serialized output training scheme that emits speaker tags and start/end timestamps, (ii) a lightweight frame-level diarization head that enhances speaker-discriminative encoder representations, (iii) diarization-guided silence suppression for improved temporal precision, and (iv) a state-machine-based forced decoding procedure that guarantees structurally valid outputs. Comprehensive evaluations on two datasets demonstrate consistent and substantial improvements over two cascaded baselines, achieving lower multi-talker word error rates and demonstrating competitive diarization accuracy across both Whisper-small and Whisper-large models. These findings highlight the effectiveness and practical utility of the proposed joint modeling framework for generating reliable, speaker-attributed transcripts of child-adult interactions at scale. The code and model weights are publicly available


【33】ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video
标题:ToS:一个专家团队集成框架,用于立体声事件定位和检测,并在视频中进行距离估计
链接:https://arxiv.org/abs/2601.17611

作者:Davide Berghi,Philip J. B. Jackson
摘要:视频中的声音事件定位和距离估计检测(3D SELD)涉及在每个时间帧识别活动声音事件,同时估计它们的空间坐标。这种多模态任务需要跨语义、空间和时间维度的联合推理,这是单个模型通常难以有效解决的挑战。为了解决这个问题,我们引入了专家团队(ToS)集成框架,它集成了三个互补的子网络:空间语言模型,时空模型和时间语言模型。每个子网络都专注于一对独特的维度,为最终预测提供独特的见解,类似于一个拥有不同专业知识的协作团队。ToS已在DCASE2025 Task 3 Stereo SELD开发集上与最先进的3D SELD视听模型进行了基准测试,在关键指标上始终优于现有方法。未来的工作将通过加强专家的适当任务,培训和培训前课程来扩展这一概念的证明。
摘要:Sound event localization and detection with distance estimation (3D SELD) in video involves identifying active sound events at each time frame while estimating their spatial coordinates. This multimodal task requires joint reasoning across semantic, spatial, and temporal dimensions, a challenge that single models often struggle to address effectively. To tackle this, we introduce the Team of Specialists (ToS) ensemble framework, which integrates three complementary sub-networks: a spatio-linguistic model, a spatio-temporal model, and a tempo-linguistic model. Each sub-network specializes in a unique pair of dimensions, contributing distinct insights to the final prediction, akin to a collaborative team with diverse expertise. ToS has been benchmarked against state-of-the-art audio-visual models for 3D SELD on the DCASE2025 Task 3 Stereo SELD development set, consistently outperforming existing methods across key metrics. Future work will extend this proof of concept by strengthening the specialists with appropriate tasks, training, and pre-training curricula.


【34】Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
标题:通过子波提示调整和多模型集成进行欺骗意识说话者验证
链接:https://arxiv.org/abs/2601.17557

作者:Aref Farhadipour,Ming Jin,Valeriia Vyshnevetska,Xiyang Li,Elisa Pellegrino,Srikanth Madikeri
备注:System description of the T03 team in the WildSpoof Challenge at ICASSP 2026
摘要:本文描述了提交给WildSpoof 2026挑战赛SASV部分的UZH-CL系统。挑战的重点是通过要求同时验证说话者身份和音频真实性来综合防御生成式欺骗攻击。我们提出了一个级联的欺骗感知说话人确认框架,集成了小波滤波器调谐XLSR-AASIST对策与多模型集成。ASV组件使用ResNet 34、ResNet 293和WavLM-ECAPA-TDNN架构,Z分数归一化后进行分数平均。在VoxCeleb 2和SpoofCeleb上训练,系统获得了0.2017的Macro a-DCF和2.08%的SASV EER。虽然该系统在域内数据的欺骗检测中实现了0.16%的EER,但在ASVspoof 5等看不见的数据集上的结果突出了跨域泛化的关键挑战。
摘要:This paper describes the UZH-CL system submitted to the SASV section of the WildSpoof 2026 challenge. The challenge focuses on the integrated defense against generative spoofing attacks by requiring the simultaneous verification of speaker identity and audio authenticity. We proposed a cascaded Spoofing-Aware Speaker Verification framework that integrates a Wavelet Prompt-Tuned XLSR-AASIST countermeasure with a multi-model ensemble. The ASV component utilizes the ResNet34, ResNet293, and WavLM-ECAPA-TDNN architectures, with Z-score normalization followed by score averaging. Trained on VoxCeleb2 and SpoofCeleb, the system obtained a Macro a-DCF of 0.2017 and a SASV EER of 2.08%. While the system achieved a 0.16% EER in spoof detection on the in-domain data, results on unseen datasets, such as the ASVspoof5, highlight the critical challenge of cross-domain generalization.


【35】Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration
标题:通过多层融合和副语言特征集成从离散令牌恢复语音情感识别的性能
链接:https://arxiv.org/abs/2601.17085

作者:Esther Sun,Abinay Reddy Naini,Carlos Busso
备注:Accepted to ICASSP 2026
摘要:离散语音标记在存储和语言模型集成方面具有显著的优势,但量化过程中的语言信息丢失限制了其在语音情感识别(SER)中的应用。本文提出了一个全面的调查离散令牌的SER。使用微调WavLM大模型,我们系统地量化性能下降在不同的层配置和k-means量化粒度。为了恢复信息丢失,我们提出了两个关键策略:(1)基于注意力的多层融合,以重新捕获来自不同层的互补信息,以及(2)集成openSMILE功能,以显式地重新引入非语言提示。我们还比较了主流的神经编解码器标记器(SpeechTokenizer,DAC,EnCodec),并分析了它们与声学特征融合时的行为。我们的研究结果表明,通过多层融合和声学特征集成,离散令牌可以缩小SER任务中连续表征的性能差距。
摘要:Discrete speech tokens offer significant advantages for storage and language model integration, but their application in speech emotion recognition (SER) is limited by paralinguistic information loss during quantization. This paper presents a comprehensive investigation of discrete tokens for SER. Using a fine-tuned WavLM-Large model, we systematically quantify performance degradation across different layer configurations and k-means quantization granularities. To recover the information loss, we propose two key strategies: (1) attention-based multi-layer fusion to recapture complementary information from different layers, and (2) integration of openSMILE features to explicitly reintroduce paralinguistic cues. We also compare mainstream neural codec tokenizers (SpeechTokenizer, DAC, EnCodec) and analyze their behaviors when fused with acoustic features. Our findings demonstrate that through multi-layer fusion and acoustic feature integration, discrete tokens can close the performance gap with continuous representations in SER tasks.


【36】PC-MCL: Patient-Consistent Multi-Cycle Learning with multi-label bias correction for respiratory sound classification
标题:PC-MCL:患者一致的多周期学习,具有呼吸音分类的多标签偏差纠正
链接:https://arxiv.org/abs/2601.17080

作者:Seung Gyu Jeong,Seong-Eun Kim
摘要:自动呼吸音分类支持肺部疾病的诊断。然而,许多深度模型仍然依赖于周期水平分析,并遭受患者特异性过拟合。我们提出了PC-MCL(患者一致性多周期学习),以解决这些限制,利用三个关键组成部分:多周期串联,3标签配方,和患者匹配的辅助任务。我们的工作解决了呼吸音分类中的多标签分布偏差,这是将多循环串联与传统的2标签制剂(爆裂声,喘息声)相结合所固有的关键问题。当正常和异常周期结合时,这种偏差表现为正常信号信息的系统性丢失。我们提出的3-标签配方(正常,裂纹,喘息)纠正这一点,保留信息的所有组成周期的混合样本。此外,患者匹配辅助任务充当多任务正则化器,鼓励模型学习更鲁棒的特征并提高泛化能力。在ICBHI 2017基准测试中,PC-MCL的ICBHI得分为65.37%,优于现有基准。消融研究证实,所有三个组成部分是必不可少的,协同工作,以提高异常呼吸事件的检测。
摘要:Automated respiratory sound classification supports the diagnosis of pulmonary diseases. However, many deep models still rely on cycle-level analysis and suffer from patient-specific overfitting. We propose PC-MCL (Patient-Consistent Multi-Cycle Learning) to address these limitations by utilizing three key components: multi-cycle concatenation, a 3-label formulation, and a patient-matching auxiliary task. Our work resolves a multi-label distributional bias in respiratory sound classification, a critical issue inherent to applying multi-cycle concatenation with the conventional 2-label formulation (crackle, wheeze). This bias manifests as a systematic loss of normal signal information when normal and abnormal cycles are combined. Our proposed 3-label formulation (normal, crackle, wheeze) corrects this by preserving information from all constituent cycles in mixed samples. Furthermore, the patient-matching auxiliary task acts as a multi-task regularizer, encouraging the model to learn more robust features and improving generalization. On the ICBHI 2017 benchmark, PC-MCL achieves an ICBHI Score of 65.37%, outperforming existing baselines. Ablation studies confirm that all three components are essential, working synergistically to improve the detection of abnormal respiratory events.


【37】BickGraphing: Web-Based Application for Visual Inspection of Audio Recordings
标题:BickGraphing:基于Web的音频记录视觉检查应用程序
链接:https://arxiv.org/abs/2601.17014

作者:Kayley Seow,Alexander Arovas,Grace Steinmetz,Emily Bick
备注:11 pages, 4 figures for submission in Journal of Open Research Software
摘要:BickGraphing是一个基于浏览器的研究工具,可以对声学记录进行视觉检查。该工具的建立是为了支持可视化作物喂养害虫的声音,以支持昆虫羽化滴管项目;然而,它广泛适用于研究中的所有音频可视化。它允许多次上传大型.wav文件,在本地计算波形和频谱图,并支持在时间和频率上交互式探索音频事件。该应用程序被实现为SvelteKit和TypeScript Web应用程序,具有使用WebAssembly编译的FFmpeg和自定义FFT实用程序的客户端信号处理管道。该软件在开放的Git存储库(https://github.com/bicklabuw/BickGraphing)上发布,并在标准MIT许可证下存档,可用于昆虫生物声学和相关领域的.wav记录的快速视觉质量检查。BickGraphing有可能成为一个本地的,易于使用的编码免费可视化平台,用于研究中的音频数据。
摘要:BickGraphing is a browser based research tool that enables visual inspection of acoustic recordings. The tool was built in support of visualizing crop feeding pest sounds in support of the Insect Eavesdropper project; however, it is widely applicable to all audiovisualizations in research. It allows multiple uploads of large .wav files, computes waveforms and spectrograms locally, and supports interactive exploration of audio events in time and frequency. The application is implemented as a SvelteKit and TypeScript web app with a client side signal processing pipeline using WebAssembly compiled FFmpeg and custom FFT utilities. The software is released on an open Git repository (https://github.com/bicklabuw/BickGraphing) and archived under a standard MIT license and can be reused for rapid visual quality checks of .wav recordings in insect bioacoustics and related fields. BickGraphing has the potential to be a local, easy to use coding free visualization platform for audio data in research.


【38】The Voice of Equity: A Systematic Evaluation of Bias Mitigation Techniques for Speech-Based Cognitive Impairment Detection Across Architectures and Demographics
标题:公平之声:跨架构和人口统计学的基于言语的认知障碍检测的偏见缓解技术的系统评估
链接:https://arxiv.org/abs/2601.16989

作者:Yasaman Haghbin,Sina Rashidi,Ali Zolnour,Maryam Zolnoori
摘要:基于语音的认知障碍检测提供了一种可扩展的非侵入性筛查,但人口统计学和语言亚组的算法偏差仍然严重不足。我们提出了第一个全面的公平性分析框架,用于基于语音的多类认知障碍检测,系统地评估跨架构和人口统计分组的偏见缓解。我们开发了两个基于transformer的架构,SpeechCARE-AGF和Whisper-LWF-LoRA,在多语言NIA挑战数据集上。与以前通常检查单一缓解技术的工作不同,我们比较了预处理,处理中和后处理方法,通过性别,年龄,教育和语言的机会均等和均等化几率来评估公平性。这两种模型都取得了很好的性能(F1:SpeechCARE-AGF 70.87,Whisper-LWF-LoRA 71.46),但表现出很大的公平性差异。与年轻组相比,≥ 80岁的成年人表现出较低的敏感性;与英语使用者相比,西班牙语使用者表现出降低的TPR。缓解效果因架构而异:过采样改善了老年人的SpeechCARE-AGF(80+ TPR:46.19%=>49.97%),但对Whisper-LWF-LoRA的影响最小。这项研究通过证明架构设计从根本上塑造了偏见模式和缓解有效性,解决了关键的医疗保健AI差距。自适应融合机制能够灵活地响应数据干预,而频率重新加权则可以在整个架构中提供强大的改进。我们的研究结果表明,公平性干预措施必须针对模型架构和人口统计学特征进行定制,为开发公平的基于语音的筛查工具提供系统框架,这些工具对于减少认知医疗中的诊断差异至关重要。
摘要:Speech-based detection of cognitive impairment offers a scalable, non-invasive screening, yet algorithmic bias across demographic and linguistic subgroups remains critically underexplored. We present the first comprehensive fairness analysis framework for speech-based multi-class cognitive impairment detection, systematically evaluating bias mitigation across architectures, and demographic subgroups. We developed two transformer-based architectures, SpeechCARE-AGF and Whisper-LWF-LoRA, on the multilingual NIA PREPARE Challenge dataset. Unlike prior work that typically examines single mitigation techniques, we compared pre-processing, in-processing, and post-processing approaches, assessing fairness via Equality of Opportunity and Equalized Odds across gender, age, education, and language. Both models achieved strong performance (F1: SpeechCARE-AGF 70.87, Whisper-LWF-LoRA 71.46) but exhibited substantial fairness disparities. Adults >=80 showed lower sensitivity versus younger groups; Spanish speakers demonstrated reduced TPR versus English speakers. Mitigation effectiveness varied by architecture: oversampling improved SpeechCARE-AGF for older adults (80+ TPR: 46.19%=>49.97%) but minimally affected Whisper-LWF-LoRA. This study addresses a critical healthcare AI gap by demonstrating that architectural design fundamentally shapes bias patterns and mitigation effectiveness. Adaptive fusion mechanisms enable flexible responses to data interventions, while frequency reweighting offers robust improvements across architectures. Our findings establish that fairness interventions must be tailored to both model architecture and demographic characteristics, providing a systematic framework for developing equitable speech-based screening tools essential for reducing diagnostic disparities in cognitive healthcare.


eess.AS音频处理


【1】Learning to Discover: A Generalized Framework for Raga Identification without Forgetting
标题:学会发现:不忘记Raga识别的通用框架
链接:https://arxiv.org/abs/2601.18766

作者:Parampreet Singh,Somya Kumar,Chaitanya Shailendra Nitawe,Vipul Arora
备注:Accepted at NCC 2026 conference
摘要:印度艺术音乐(IAM)中的Raga识别仍然具有挑战性,因为存在许多很少执行的Raga,这些Raga在可用的训练数据集中没有表现出来。传统的分类模型在这种情况下挣扎,因为它们假设一组封闭的已知类别,因此无法识别或有意义地将以前看不见的Ragas分组。最近的研究试图对看不见的Ragas进行分类,但他们遇到了灾难性遗忘的问题,即以前看到的Ragas的知识减少了。为了解决这个问题,我们采用了一个统一的学习框架,利用标记和未标记的音频,使模型能够发现与看不见的Ragas对应的连贯类别,同时保留先前已知的知识。我们在基准Raga Identification数据集上测试了我们的模型,并展示了它在对以前看到的、看不见的和所有Raga类进行分类时的性能。所提出的方法甚至在发现看不见的Raga类别方面超越了以前基于NCD的管道,为IAM任务的表示学习提供了新的见解。
摘要:Raga identification in Indian Art Music (IAM) remains challenging due to the presence of numerous rarely performed Ragas that are not represented in available training datasets. Traditional classification models struggle in this setting, as they assume a closed set of known categories and therefore fail to recognise or meaningfully group previously unseen Ragas. Recent works have tried categorizing unseen Ragas, but they run into a problem of catastrophic forgetting, where the knowledge of previously seen Ragas is diminished. To address this problem, we adopt a unified learning framework that leverages both labeled and unlabeled audio, enabling the model to discover coherent categories corresponding to the unseen Ragas, while retaining the knowledge of previously known ones. We test our model on benchmark Raga Identification datasets and demonstrate its performance in categorizing previously seen, unseen, and all Raga classes. The proposed approach surpasses the previous NCD-based pipeline even in discovering the unseen Raga categories, offering new insights into representation learning for IAM tasks.


【2】Audio Inpainting in Time-Frequency Domain with Phase-Aware Prior
标题:具有相感知先验的时频域音频修复
链接:https://arxiv.org/abs/2601.18535

作者:Peter Balušík,Pavel Rajmic
备注:submitted to IEEE Transactions on Audio, Speech and Language Processing
摘要:时域中所谓的音频修复问题是指估计信号内样本的缺失片段。多年来,已经开发了几种用于这种类型的音频修复的方法。与这种情况相反,在文献中出现了修复的时频变体,其中的挑战是用可靠的信息重建丢失的频谱图列。我们提出了一种方法来解决这个时频音频修复问题。我们的方法是基于最近推出的相位感知信号之前,利用瞬时频率的估计。使用广义Chambolle-Pock算法制定和解决优化问题。该方法与其他时频修复方法进行了客观和主观评估,特别是深度先验神经网络和基于自回归的方法Janssen-TF。我们提出的方法超越了这些方法在客观评价,以及在进行听力测试。此外,与替代方法相比,该结果是以大幅降低的计算要求实现的。
摘要:The so-called audio inpainting problem in the time domain refers to estimating missing segments of samples within a signal. Over the years, several methods have been developed for such type of audio inpainting. In contrast to this case, a time-frequency variant of inpainting appeared in the literature, where the challenge is to reconstruct missing spectrogram columns with reliable information. We propose a method to address this time-frequency audio inpainting problem. Our approach is based on the recently introduced phase-aware signal prior that exploits an estimate of the instantaneous frequency. An optimization problem is formulated and solved using the generalized Chambolle-Pock algorithm. The proposed method is evaluated both objectively and subjectively against other time-frequency inpainting methods, specifically a deep-prior neural network and the autoregression-based approach known as Janssen-TF. Our proposed approach surpassed these methods in the objective evaluation as well as in the conducted listening test. Moreover, this outcome is achieved with a substantially reduced computational requirement compared to alternative methods.


【3】Noise-Robust AV-ASR Using Visual Features Both in the Whisper Encoder and Decoder
标题:在Whisper编码器和解码器中使用视觉特征的噪音稳健AV-ASB
链接:https://arxiv.org/abs/2601.18396

作者:Zhengyang Li,Thomas Graave,Björn Möller,Zehang Wu,Matthias Franz,Tim Fingscheidt
备注:accepted at ICASSP2026
摘要:在视听自动语音识别(AV-ASR)系统中,预先训练的ASR中的视觉特征的信息融合已被证明是一种有前途的方法,以提高噪声鲁棒性。在这项工作中,基于突出的Whisper ASR,首先,我们提出了一种简单有效的视觉融合方法-在编码器和解码器(两用)中使用视觉特征-学习编码器中的视听交互并在解码器中加权模态。其次,我们比较了各种尺寸的Whisper模型中的视觉融合方法。我们提出的两用方法显示出一致的噪声鲁棒性改善,例如,与信噪比(SNR)为0 dB的多路重合噪声中的典型参考中间融合相比,基于Whisper small的35%相对改进(WER:4.41%对6.83%),以及基于Whisper medium的57%相对改进(WER:4.07%对9.53%)。第三,我们进行消融研究,检查各种模块设计和融合选项的影响。通过对1929小时的视听数据进行微调,我们使用Whisper介质的两用方法在各种SNR下实现了4.08%(MUSAN串音噪声)和4.43%(NoiseX串音噪声)的平均WER,从而在LRS 3 AV-ASR基准测试中建立了噪声条件下的最新技术水平。我们的代码在https://github.com/ifnspaml/Dual-Use-AVASR
摘要:In audiovisual automatic speech recognition (AV-ASR) systems, information fusion of visual features in a pre-trained ASR has been proven as a promising method to improve noise robustness. In this work, based on the prominent Whisper ASR, first, we propose a simple and effective visual fusion method -- use of visual features both in encoder and decoder (dual-use) -- to learn the audiovisual interactions in the encoder and to weigh modalities in the decoder. Second, we compare visual fusion methods in Whisper models of various sizes. Our proposed dual-use method shows consistent noise robustness improvement, e.g., a 35% relative improvement (WER: 4.41% vs. 6.83%) based on Whisper small, and a 57% relative improvement (WER: 4.07% vs. 9.53%) based on Whisper medium, compared to typical reference middle fusion in babble noise with a signal-to-noise ratio (SNR) of 0dB. Third, we conduct ablation studies examining the impact of various module designs and fusion options. Fine-tuned on 1929 hours of audiovisual data, our dual-use method using Whisper medium achieves 4.08% (MUSAN babble noise) and 4.43% (NoiseX babble noise) average WER across various SNRs, thereby establishing a new state-of-the-art in noisy conditions on the LRS3 AV-ASR benchmark. Our code is at https://github.com/ifnspaml/Dual-Use-AVASR


【4】Residual Learning for Neural Ambisonics Encoders
标题:神经立体声编码器的剩余学习
链接:https://arxiv.org/abs/2601.18322

作者:Thomas Deppisch,Yang Gao,Manan Mittal,Benjamin Stahl,Christoph Hold,David Alon,Zamir Ben-Hur
摘要:智能眼镜和延展实境耳机等新兴可穿戴设备需要从紧凑的头戴式麦克风阵列捕获高质量的空间音频。高保真度立体声复制通过将阵列信号映射到球面谐波(SH)系数来提供与设备无关的空间音频表示。然而,在实践中,准确的编码仍然具有挑战性。虽然传统的线性编码器是信号独立和鲁棒的,但它们会放大低频噪声并遭受高频空间混叠。另一方面,神经网络方法可以优于线性编码器,但它们通常假设理想化的麦克风,并且在现实世界的场景中可能表现不一致。为了利用它们的互补优势,我们引入了一个剩余学习框架,该框架通过神经网络的校正来改进线性编码器。使用从智能眼镜测量的阵列传递函数,我们比较了一个基于UNet的编码器从文献中的一个新的经常性注意力模型。我们的分析表明,这两个神经编码器只有在集成到残差学习框架中时才始终优于线性基线。在残差配置中,两个神经模型在域内数据的所有测试指标上都实现了一致且显著的改进,并在域外数据上实现了适度的增益。然而,相干性分析表明,所有的神经编码器配置继续与方向准确的高频编码斗争。
摘要:Emerging wearable devices such as smartglasses and extended reality headsets demand high-quality spatial audio capture from compact, head-worn microphone arrays. Ambisonics provides a device-agnostic spatial audio representation by mapping array signals to spherical harmonic (SH) coefficients. In practice, however, accurate encoding remains challenging. While traditional linear encoders are signal-independent and robust, they amplify low-frequency noise and suffer from high-frequency spatial aliasing. On the other hand, neural network approaches can outperform linear encoders but they often assume idealized microphones and may perform inconsistently in real-world scenarios. To leverage their complementary strengths, we introduce a residual-learning framework that refines a linear encoder with corrections from a neural network. Using measured array transfer functions from smartglasses, we compare a UNet-based encoder from the literature with a new recurrent attention model. Our analysis reveals that both neural encoders only consistently outperform the linear baseline when integrated within the residual learning framework. In the residual configuration, both neural models achieve consistent and significant improvements across all tested metrics for in-domain data and moderate gains for out-of-domain data. Yet, coherence analysis indicates that all neural encoder configurations continue to struggle with directionally accurate high-frequency encoding.


【5】Noise-Robust Contrastive Learning with an MFCC-Conformer For Coronary Artery Disease Detection
标题:用于冠状动脉疾病检测的MFCC一致性的噪音稳健对比学习
链接:https://arxiv.org/abs/2601.18295

作者:Milan Marocchi,Matthew Fynn,Yue Rong
备注:This paper has been accepted for presentation at ICASSP 2026. \c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses. 5 pages, 1 figure
摘要:心血管疾病(CVD)是全球死亡的主要原因,其中冠状动脉疾病(CAD)是CVD的最大亚类。最近,人们越来越关注使用心音图(PCG)信号检测CAD,在低噪声和最佳传感器放置的临床环境中取得了很大成功。已经发现多通道技术对噪声更鲁棒;然而,在真实世界数据上实现鲁棒性能仍然是一个挑战。这项工作利用了一种新的基于多通道能量的噪声段抑制算法,使用心脏和噪声参考麦克风,在训练深度学习分类器之前丢弃具有大量非平稳噪声的音频段。这种基于一致性的分类器从多个通道中提取梅尔频率倒谱系数(MFCC),进一步帮助提高模型的噪声鲁棒性。所提出的方法在297个受试者上实现了78.4%的准确率和78.2%的平衡准确率,与没有噪声段拒绝的训练相比,分别提高了4.1%和4.3%。
摘要:Cardiovascular diseases (CVD) are the leading cause of death worldwide, with coronary artery disease (CAD) comprising the largest subcategory of CVDs. Recently, there has been increased focus on detecting CAD using phonocardiogram (PCG) signals, with high success in clinical environments with low noise and optimal sensor placement. Multichannel techniques have been found to be more robust to noise; however, achieving robust performance on real-world data remains a challenge. This work utilises a novel multichannel energy-based noisy-segment rejection algorithm, using heart and noise-reference microphones, to discard audio segments with large amounts of nonstationary noise before training a deep learning classifier. This conformer-based classifier takes mel-frequency cepstral coefficients (MFCCs) from multiple channels, further helping improve the model's noise robustness. The proposed method achieved 78.4% accuracy and 78.2% balanced accuracy on 297 subjects, representing improvements of 4.1% and 4.3%, respectively, compared to training without noisy-segment rejection.


【6】Efficient Rehearsal for Continual Learning in ASR via Singular Value Tuning
标题:通过奇异值调整在ASB中进行连续学习的高效排练
链接:https://arxiv.org/abs/2601.18266

作者:Steven Vander Eeckt,Hugo Van hamme
备注:Accepted for publication in IEEE Transactions on Audio, Speech, and Language Processing
摘要:自动语音识别(ASR)中的持续学习(CL)在适应新任务、领域或说话人时会遭受灾难性遗忘。缓解这种情况的常见策略是将过去数据的子集存储在内存中以供排练。然而,基于排练的方法面临着关键的限制:存储数据通常成本高昂,预先训练的模型不可行,或者受到隐私法规的限制。使用较小的内存大小运行现有的基于排练的方法来缓解这些问题通常会导致性能下降。   我们提出了一个排练为基础的CL方法,即使在最小的内存仍然有效。它分为两个阶段:首先,对新任务进行微调;其次,将奇异值分解(SVD)应用于线性层中的变化,并且以参数有效的方式,仅重新训练奇异值上的门控向量,其控制使用排练接受来自第一阶段的更新的程度。我们广泛的测试和分析我们的方法在两个单语和两个多语种的基准。我们的方法减少了遗忘,并优于最先进的CL方法的ASR,即使是在限制为一个单一的话语每个先前的任务。
摘要:Continual Learning (CL) in Automatic Speech Recognition (ASR) suffers from catastrophic forgetting when adapting to new tasks, domains, or speakers. A common strategy to mitigate this is to store a subset of past data in memory for rehearsal. However, rehearsal-based methods face key limitations: storing data is often costly, infeasible with pre-trained models, or restricted by privacy regulations. Running existing rehearsal-based methods with smaller memory sizes to alleviate these issues usually leads to degraded performance.   We propose a rehearsal-based CL method that remains effective even with minimal memory. It operates in two stages: first, fine-tuning on the new task; second, applying Singular Value Decomposition (SVD) to the changes in linear layers and, in a parameter-efficient manner, retraining only gating vectors on the singular values, which control to extent to which updates from the first stage are accepted, using rehearsal. We extensively test and analyze our method on two monolingual and two multilingual benchmarks. Our method reduces forgetting and outperforms state-of-the-art CL approaches for ASR, even when limited to a single utterance per previous task.


【7】OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
标题:OneVoice:一种模式、三种场景--迈向统一的Zero-Shot语音转换
链接:https://arxiv.org/abs/2601.18094

作者:Zhichao Wang,Tao Li,Wenshuo Ge,Zihao Cui,Shilei Zhang,Junlan Feng
备注:Work in progress
摘要:语音转换技术的最新进展在说话人克隆和语言保存方面取得了新的里程碑。但该领域仍然是分散的,依赖于语言保护,表达和歌唱场景的专门模型。我们提出了OneVoice,一个统一的zero-shot框架,能够在一个模型中处理所有三种情况。OneVoice建立在一个连续的语言模型上,该模型使用无VAE的下一个补丁扩散进行训练,确保了高保真度和高效的序列建模。其统一的核心设计在于一个混合专家(MoE),旨在明确建模共享的转换知识和具体的表达能力。专家的选择是协调的双路径路由机制,包括共享的专家隔离和基于全局-局部线索的领域专家分配。为了精确的调节,特定的韵律特征通过门控机制融合到每一层中,允许韵律信息的自适应使用。此外,为了实现核心思想并缓解不平衡问题(丰富的语音与稀缺的歌声),我们采用了两阶段渐进式训练,包括基础预训练和基于LoRA的领域专家的场景增强。实验表明,OneVoice在所有三种场景下都匹配或超越了专业模型,同时验证了对场景的灵活控制,并提供了只需2步的快速解码版本。代码和模型将很快发布。
摘要:Recent progress of voice conversion~(VC) has achieved a new milestone in speaker cloning and linguistic preservation. But the field remains fragmented, relying on specialized models for linguistic-preserving, expressive, and singing scenarios. We propose OneVoice, a unified zero-shot framework capable of handling all three scenarios within a single model. OneVoice is built upon a continuous language model trained with VAE-free next-patch diffusion, ensuring high fidelity and efficient sequence modeling. Its core design for unification lies in a Mixture-of-Experts (MoE) designed to explicitly model shared conversion knowledge and scenario-specific expressivity. Expert selection is coordinated by a dual-path routing mechanism, including shared expert isolation and scenario-aware domain expert assignment with global-local cues. For precise conditioning, scenario-specific prosodic features are fused into each layer via a gated mechanism, allowing adaptive usage of prosody information. Furthermore, to enable the core idea and alleviate the imbalanced issue (abundant speech vs. scarce singing), we adopt a two-stage progressive training that includes foundational pre-training and scenario enhancement with LoRA-based domain experts. Experiments show that OneVoice matches or surpasses specialized models across all three scenarios, while verifying flexible control over scenarios and offering a fast decoding version as few as 2 steps. Code and model will be released soon.


【8】SpatialEmb: Extract and Encode Spatial Information for 1-Stage Multi-channel Multi-speaker ASR on Arbitrary Microphone Arrays
标题:SpatialEmb:提取和编码任意麦克风阵列上的1级多通道多扬声器ASB的空间信息
链接:https://arxiv.org/abs/2601.18037

作者:Yiwen Shao,Yong Xu,Sanjeev Khudanpur,Dong Yu
备注:SLT 2024
摘要:空间信息是多通道多说话人目标语音识别的重要线索。大多数最先进的多通道自动语音识别(ASR)系统只在语音分离阶段提取空间特征,然后对分离的语音进行标准的单通道ASR。由于预处理模块的累积错误,这种方法导致低效、冗长的流水线和次优的ASR性能。此外,大多数空间特征提取方法依赖于扬声器位置和麦克风拓扑结构的知识,使得系统依赖于特定的设置,并且难以适应新设备。在这项工作中,我们提出了一个解决这些问题的轻量级嵌入模块SpatialEmb,提取和编码空间信息直接为ASR模型,支持固定和任意麦克风拓扑结构。我们进行全面的实验AliMeeting,一个真正的会议语料库,以确定最佳的模型设计SpatialEmb的性能和效率。我们用105小时Train-Ali-far训练的最佳模型在Eval和Test集上实现了17.04%和20.32%的字符错误率(CER),用相同的训练数据建立了一个新的最先进的结果。
摘要:Spatial information is a critical clue for multi-channel multi-speaker target speech recognition. Most state-of-the-art multi-channel Automatic Speech Recognition (ASR) systems extract spatial features only during the speech separation stage, followed by standard single-channel ASR on the separated speech. This approach results in an inefficient, lengthy pipeline and sub-optimal ASR performance due to the accumulated errors from preprocessing modules. Furthermore, most spatial feature extraction methods depend on the knowledge of speaker positions and microphone topology, making the systems reliant on specific settings and challenging to adapt to new equipment. In this work, we propose a solution to these issues with a lightweight embedding module named SpatialEmb, which extracts and encodes spatial information directly for the ASR model, supporting both fixed and arbitrary microphone topology. We conduct comprehensive experiments on AliMeeting, a real meeting corpus, to determine the optimal model design for SpatialEmb in terms of both performance and efficiency. Our best model trained with 105 hours Train-Ali-far achieves 17.04% and 20.32% character error rates (CER) on the Eval and Test sets, establishing a new state-of-the-art result with the same training data.


【9】AmbER$^2$: Dual Ambiguity-Aware Emotion Recognition Applied to Speech and Text
标题:AmbER $' 2 $:应用于语音和文本的双重模糊感知情感识别
链接:https://arxiv.org/abs/2601.18010

作者:Jingyao Wu,Grace Lin,Yinuo Song,Rosalind Picard
备注:Accepted in ICASSP 2026
摘要:情感识别本质上是模糊的,不确定性既来自评分者的分歧,也来自语音和文本等模态之间的差异。有越来越多的兴趣建模评分员模糊使用标签分布。然而,模态模糊性仍然未被充分研究,多模态方法通常依赖于简单的特征融合,而没有明确解决模态之间的冲突。在这项工作中,我们提出了AmbER$^2$,一个双重的模糊意识的框架,同时模型的评分水平和模态水平的模糊性,通过一个教师-学生架构与分布明智的培训目标。对IEMOCAP和MSP播客的评估表明,AmbER$^2$始终提高了传统交叉熵基线的分布保真度,并实现了与最近最先进的系统竞争或优于其的性能。例如,在IEMOCAP上,AmbER$^2$在Bhattacharyya系数上实现了20.3%的相对改进(0.83 vs. 0.69),在R$^2$上实现了13.6%的相对改进(0.67 vs. 0.59),在准确性上实现了3.8%的相对改进(0.683 vs. 0.658),在F1上实现了4.5%的相对改进(0.675 vs. 0.646)。跨模糊度水平的进一步分析表明,显式建模模糊度对于高度不确定的样本特别有益。这些发现强调了在建立强大的情感识别系统时,联合解决评分者和模态模糊的重要性。
摘要:Emotion recognition is inherently ambiguous, with uncertainty arising both from rater disagreement and from discrepancies across modalities such as speech and text. There is growing interest in modeling rater ambiguity using label distributions. However, modality ambiguity remains underexplored, and multimodal approaches often rely on simple feature fusion without explicitly addressing conflicts between modalities. In this work, we propose AmbER$^2$, a dual ambiguity-aware framework that simultaneously models rater-level and modality-level ambiguity through a teacher-student architecture with a distribution-wise training objective. Evaluations on IEMOCAP and MSP-Podcast show that AmbER$^2$ consistently improves distributional fidelity over conventional cross-entropy baselines and achieves performance competitive with, or superior to, recent state-of-the-art systems. For example, on IEMOCAP, AmbER$^2$ achieves relative improvements of 20.3% on Bhattacharyya coefficient (0.83 vs. 0.69), 13.6% on R$^2$ (0.67 vs. 0.59), 3.8% on accuracy (0.683 vs. 0.658), and 4.5% on F1 (0.675 vs. 0.646). Further analysis across ambiguity levels shows that explicitly modeling ambiguity is particularly beneficial for highly uncertain samples. These findings highlight the importance of jointly addressing rater and modality ambiguity when building robust emotion recognition systems.


【10】Speech Emotion Recognition with ASR Integration
标题:与ASB集成的语音情感识别
链接:https://arxiv.org/abs/2601.17901

作者:Yuanchao Li
备注:PhD Thesis
摘要:语音情感识别(SER)在理解人类交流、实现情感智能系统以及作为通用人工智能(AGI)发展的基本组成部分方面发挥着关键作用。然而,由于情感表达的复杂性以及当前语音和语言技术的局限性,在真实世界、自发和低资源场景中部署SER仍然是一个重大挑战。本文研究了自动语音识别(ASR)与SER的集成,旨在提高口语情感识别的鲁棒性、可扩展性和实用性。
摘要:Speech Emotion Recognition (SER) plays a pivotal role in understanding human communication, enabling emotionally intelligent systems, and serving as a fundamental component in the development of Artificial General Intelligence (AGI). However, deploying SER in real-world, spontaneous, and low-resource scenarios remains a significant challenge due to the complexity of emotional expression and the limitations of current speech and language technologies. This thesis investigates the integration of Automatic Speech Recognition (ASR) into SER, with the goal of enhancing the robustness, scalability, and practical applicability of emotion recognition from spoken language.


【11】End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
标题:具有儿童与成人互动的端到端联合ASB和说话者角色扩展
链接:https://arxiv.org/abs/2601.17640

作者:Anfeng Xu,Tiantian Feng,Somer Bishop,Catherine Lord,Shrikanth Narayanan
备注:Under review for IEEE
摘要:儿童与成人之间口语交流的准确转录和说话人日记化对于发展研究和临床研究至关重要。然而,手动注释是耗时的,并且难以扩展。现有的自动化系统通常依赖于级联的说话人日志和语音识别管道,这可能导致错误传播。本文提出了一个统一的端到端的框架,扩展了耳语编码器-解码器架构,以联合建模ASR和儿童-成人扬声器角色日记。拟议的办法包括:(i)发出说话者标签和开始/结束时间戳的串行化输出训练方案,(ii)增强说话者区分编码器表示的轻量级帧级日志化头,(iii)用于改进的时间精度的日志化引导的静音抑制,以及(iv)保证结构上有效的输出的基于状态机的强制解码过程。对两个数据集的综合评估表明,在两个级联基线上有了一致和实质性的改进,实现了更低的多说话者单词错误率,并在Whisper-small和Whisper-large模型中展示了具有竞争力的日记准确性。这些研究结果突出的有效性和实际效用的建议联合建模框架,以产生可靠的,扬声器归因于转录的儿童与成人的互动规模。代码和模型权重是公开的
摘要:Accurate transcription and speaker diarization of child-adult spoken interactions are crucial for developmental and clinical research. However, manual annotation is time-consuming and challenging to scale. Existing automated systems typically rely on cascaded speaker diarization and speech recognition pipelines, which can lead to error propagation. This paper presents a unified end-to-end framework that extends the Whisper encoder-decoder architecture to jointly model ASR and child-adult speaker role diarization. The proposed approach integrates: (i) a serialized output training scheme that emits speaker tags and start/end timestamps, (ii) a lightweight frame-level diarization head that enhances speaker-discriminative encoder representations, (iii) diarization-guided silence suppression for improved temporal precision, and (iv) a state-machine-based forced decoding procedure that guarantees structurally valid outputs. Comprehensive evaluations on two datasets demonstrate consistent and substantial improvements over two cascaded baselines, achieving lower multi-talker word error rates and demonstrating competitive diarization accuracy across both Whisper-small and Whisper-large models. These findings highlight the effectiveness and practical utility of the proposed joint modeling framework for generating reliable, speaker-attributed transcripts of child-adult interactions at scale. The code and model weights are publicly available


【12】ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video
标题:ToS:一个专家团队集成框架,用于立体声事件定位和检测,并在视频中进行距离估计
链接:https://arxiv.org/abs/2601.17611

作者:Davide Berghi,Philip J. B. Jackson
摘要:视频中的声音事件定位和距离估计检测(3D SELD)涉及在每个时间帧识别活动声音事件,同时估计它们的空间坐标。这种多模态任务需要跨语义、空间和时间维度的联合推理,这是单个模型通常难以有效解决的挑战。为了解决这个问题,我们引入了专家团队(ToS)集成框架,它集成了三个互补的子网络:空间语言模型,时空模型和时间语言模型。每个子网络都专注于一对独特的维度,为最终预测提供独特的见解,类似于一个拥有不同专业知识的协作团队。ToS已在DCASE2025 Task 3 Stereo SELD开发集上与最先进的3D SELD视听模型进行了基准测试,在关键指标上始终优于现有方法。未来的工作将通过加强专家的适当任务,培训和培训前课程来扩展这一概念的证明。
摘要:Sound event localization and detection with distance estimation (3D SELD) in video involves identifying active sound events at each time frame while estimating their spatial coordinates. This multimodal task requires joint reasoning across semantic, spatial, and temporal dimensions, a challenge that single models often struggle to address effectively. To tackle this, we introduce the Team of Specialists (ToS) ensemble framework, which integrates three complementary sub-networks: a spatio-linguistic model, a spatio-temporal model, and a tempo-linguistic model. Each sub-network specializes in a unique pair of dimensions, contributing distinct insights to the final prediction, akin to a collaborative team with diverse expertise. ToS has been benchmarked against state-of-the-art audio-visual models for 3D SELD on the DCASE2025 Task 3 Stereo SELD development set, consistently outperforming existing methods across key metrics. Future work will extend this proof of concept by strengthening the specialists with appropriate tasks, training, and pre-training curricula.


【13】Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
标题:通过子波提示调整和多模型集成进行欺骗意识说话者验证
链接:https://arxiv.org/abs/2601.17557

作者:Aref Farhadipour,Ming Jin,Valeriia Vyshnevetska,Xiyang Li,Elisa Pellegrino,Srikanth Madikeri
备注:System description of the T03 team in the WildSpoof Challenge at ICASSP 2026
摘要:本文描述了提交给WildSpoof 2026挑战赛SASV部分的UZH-CL系统。挑战的重点是通过要求同时验证说话者身份和音频真实性来综合防御生成式欺骗攻击。我们提出了一个级联的欺骗感知说话人确认框架,集成了小波滤波器调谐XLSR-AASIST对策与多模型集成。ASV组件使用ResNet 34、ResNet 293和WavLM-ECAPA-TDNN架构,Z分数归一化后进行分数平均。在VoxCeleb 2和SpoofCeleb上训练,系统获得了0.2017的Macro a-DCF和2.08%的SASV EER。虽然该系统在域内数据的欺骗检测中实现了0.16%的EER,但在ASVspoof 5等看不见的数据集上的结果突出了跨域泛化的关键挑战。
摘要:This paper describes the UZH-CL system submitted to the SASV section of the WildSpoof 2026 challenge. The challenge focuses on the integrated defense against generative spoofing attacks by requiring the simultaneous verification of speaker identity and audio authenticity. We proposed a cascaded Spoofing-Aware Speaker Verification framework that integrates a Wavelet Prompt-Tuned XLSR-AASIST countermeasure with a multi-model ensemble. The ASV component utilizes the ResNet34, ResNet293, and WavLM-ECAPA-TDNN architectures, with Z-score normalization followed by score averaging. Trained on VoxCeleb2 and SpoofCeleb, the system obtained a Macro a-DCF of 0.2017 and a SASV EER of 2.08%. While the system achieved a 0.16% EER in spoof detection on the in-domain data, results on unseen datasets, such as the ASVspoof5, highlight the critical challenge of cross-domain generalization.


【14】Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration
标题:通过多层融合和副语言特征集成从离散令牌恢复语音情感识别的性能
链接:https://arxiv.org/abs/2601.17085

作者:Esther Sun,Abinay Reddy Naini,Carlos Busso
备注:Accepted to ICASSP 2026
摘要:离散语音标记在存储和语言模型集成方面具有显著的优势,但量化过程中的语言信息丢失限制了其在语音情感识别(SER)中的应用。本文提出了一个全面的调查离散令牌的SER。使用微调WavLM大模型,我们系统地量化性能下降在不同的层配置和k-means量化粒度。为了恢复信息丢失,我们提出了两个关键策略:(1)基于注意力的多层融合,以重新捕获来自不同层的互补信息,以及(2)集成openSMILE功能,以显式地重新引入非语言提示。我们还比较了主流的神经编解码器标记器(SpeechTokenizer,DAC,EnCodec),并分析了它们与声学特征融合时的行为。我们的研究结果表明,通过多层融合和声学特征集成,离散令牌可以缩小SER任务中连续表征的性能差距。
摘要:Discrete speech tokens offer significant advantages for storage and language model integration, but their application in speech emotion recognition (SER) is limited by paralinguistic information loss during quantization. This paper presents a comprehensive investigation of discrete tokens for SER. Using a fine-tuned WavLM-Large model, we systematically quantify performance degradation across different layer configurations and k-means quantization granularities. To recover the information loss, we propose two key strategies: (1) attention-based multi-layer fusion to recapture complementary information from different layers, and (2) integration of openSMILE features to explicitly reintroduce paralinguistic cues. We also compare mainstream neural codec tokenizers (SpeechTokenizer, DAC, EnCodec) and analyze their behaviors when fused with acoustic features. Our findings demonstrate that through multi-layer fusion and acoustic feature integration, discrete tokens can close the performance gap with continuous representations in SER tasks.


【15】PC-MCL: Patient-Consistent Multi-Cycle Learning with multi-label bias correction for respiratory sound classification
标题:PC-MCL:患者一致的多周期学习,具有呼吸音分类的多标签偏差纠正
链接:https://arxiv.org/abs/2601.17080

作者:Seung Gyu Jeong,Seong-Eun Kim
摘要:自动呼吸音分类支持肺部疾病的诊断。然而,许多深度模型仍然依赖于周期水平分析,并遭受患者特异性过拟合。我们提出了PC-MCL(患者一致性多周期学习),以解决这些限制,利用三个关键组成部分:多周期串联,3标签配方,和患者匹配的辅助任务。我们的工作解决了呼吸音分类中的多标签分布偏差,这是将多循环串联与传统的2标签制剂(爆裂声,喘息声)相结合所固有的关键问题。当正常和异常周期结合时,这种偏差表现为正常信号信息的系统性丢失。我们提出的3-标签配方(正常,裂纹,喘息)纠正这一点,保留信息的所有组成周期的混合样本。此外,患者匹配辅助任务充当多任务正则化器,鼓励模型学习更鲁棒的特征并提高泛化能力。在ICBHI 2017基准测试中,PC-MCL的ICBHI得分为65.37%,优于现有基准。消融研究证实,所有三个组成部分是必不可少的,协同工作,以提高异常呼吸事件的检测。
摘要:Automated respiratory sound classification supports the diagnosis of pulmonary diseases. However, many deep models still rely on cycle-level analysis and suffer from patient-specific overfitting. We propose PC-MCL (Patient-Consistent Multi-Cycle Learning) to address these limitations by utilizing three key components: multi-cycle concatenation, a 3-label formulation, and a patient-matching auxiliary task. Our work resolves a multi-label distributional bias in respiratory sound classification, a critical issue inherent to applying multi-cycle concatenation with the conventional 2-label formulation (crackle, wheeze). This bias manifests as a systematic loss of normal signal information when normal and abnormal cycles are combined. Our proposed 3-label formulation (normal, crackle, wheeze) corrects this by preserving information from all constituent cycles in mixed samples. Furthermore, the patient-matching auxiliary task acts as a multi-task regularizer, encouraging the model to learn more robust features and improving generalization. On the ICBHI 2017 benchmark, PC-MCL achieves an ICBHI Score of 65.37%, outperforming existing baselines. Ablation studies confirm that all three components are essential, working synergistically to improve the detection of abnormal respiratory events.


【16】BickGraphing: Web-Based Application for Visual Inspection of Audio Recordings
标题:BickGraphing:基于Web的音频记录视觉检查应用程序
链接:https://arxiv.org/abs/2601.17014

作者:Kayley Seow,Alexander Arovas,Grace Steinmetz,Emily Bick
备注:11 pages, 4 figures for submission in Journal of Open Research Software
摘要:BickGraphing是一个基于浏览器的研究工具,可以对声学记录进行视觉检查。该工具的建立是为了支持可视化作物喂养害虫的声音,以支持昆虫羽化滴管项目;然而,它广泛适用于研究中的所有音频可视化。它允许多次上传大型.wav文件,在本地计算波形和频谱图,并支持在时间和频率上交互式探索音频事件。该应用程序被实现为SvelteKit和TypeScript Web应用程序,具有使用WebAssembly编译的FFmpeg和自定义FFT实用程序的客户端信号处理管道。该软件在开放的Git存储库(https://github.com/bicklabuw/BickGraphing)上发布,并在标准MIT许可证下存档,可用于昆虫生物声学和相关领域的.wav记录的快速视觉质量检查。BickGraphing有可能成为一个本地的,易于使用的编码免费可视化平台,用于研究中的音频数据。
摘要:BickGraphing is a browser based research tool that enables visual inspection of acoustic recordings. The tool was built in support of visualizing crop feeding pest sounds in support of the Insect Eavesdropper project; however, it is widely applicable to all audiovisualizations in research. It allows multiple uploads of large .wav files, computes waveforms and spectrograms locally, and supports interactive exploration of audio events in time and frequency. The application is implemented as a SvelteKit and TypeScript web app with a client side signal processing pipeline using WebAssembly compiled FFmpeg and custom FFT utilities. The software is released on an open Git repository (https://github.com/bicklabuw/BickGraphing) and archived under a standard MIT license and can be reused for rapid visual quality checks of .wav recordings in insect bioacoustics and related fields. BickGraphing has the potential to be a local, easy to use coding free visualization platform for audio data in research.


【17】The Voice of Equity: A Systematic Evaluation of Bias Mitigation Techniques for Speech-Based Cognitive Impairment Detection Across Architectures and Demographics
标题:公平之声:跨架构和人口统计学的基于言语的认知障碍检测的偏见缓解技术的系统评估
链接:https://arxiv.org/abs/2601.16989

作者:Yasaman Haghbin,Sina Rashidi,Ali Zolnour,Maryam Zolnoori
摘要:基于语音的认知障碍检测提供了一种可扩展的非侵入性筛查,但人口统计学和语言亚组的算法偏差仍然严重不足。我们提出了第一个全面的公平性分析框架,用于基于语音的多类认知障碍检测,系统地评估跨架构和人口统计分组的偏见缓解。我们开发了两个基于transformer的架构,SpeechCARE-AGF和Whisper-LWF-LoRA,在多语言NIA挑战数据集上。与以前通常检查单一缓解技术的工作不同,我们比较了预处理,处理中和后处理方法,通过性别,年龄,教育和语言的机会均等和均等化几率来评估公平性。这两种模型都取得了很好的性能(F1:SpeechCARE-AGF 70.87,Whisper-LWF-LoRA 71.46),但表现出很大的公平性差异。与年轻组相比,≥ 80岁的成年人表现出较低的敏感性;与英语使用者相比,西班牙语使用者表现出降低的TPR。缓解效果因架构而异:过采样改善了老年人的SpeechCARE-AGF(80+ TPR:46.19%=>49.97%),但对Whisper-LWF-LoRA的影响最小。这项研究通过证明架构设计从根本上塑造了偏见模式和缓解有效性,解决了关键的医疗保健AI差距。自适应融合机制能够灵活地响应数据干预,而频率重新加权则可以在整个架构中提供强大的改进。我们的研究结果表明,公平性干预措施必须针对模型架构和人口统计学特征进行定制,为开发公平的基于语音的筛查工具提供系统框架,这些工具对于减少认知医疗中的诊断差异至关重要。
摘要:Speech-based detection of cognitive impairment offers a scalable, non-invasive screening, yet algorithmic bias across demographic and linguistic subgroups remains critically underexplored. We present the first comprehensive fairness analysis framework for speech-based multi-class cognitive impairment detection, systematically evaluating bias mitigation across architectures, and demographic subgroups. We developed two transformer-based architectures, SpeechCARE-AGF and Whisper-LWF-LoRA, on the multilingual NIA PREPARE Challenge dataset. Unlike prior work that typically examines single mitigation techniques, we compared pre-processing, in-processing, and post-processing approaches, assessing fairness via Equality of Opportunity and Equalized Odds across gender, age, education, and language. Both models achieved strong performance (F1: SpeechCARE-AGF 70.87, Whisper-LWF-LoRA 71.46) but exhibited substantial fairness disparities. Adults >=80 showed lower sensitivity versus younger groups; Spanish speakers demonstrated reduced TPR versus English speakers. Mitigation effectiveness varied by architecture: oversampling improved SpeechCARE-AGF for older adults (80+ TPR: 46.19%=>49.97%) but minimally affected Whisper-LWF-LoRA. This study addresses a critical healthcare AI gap by demonstrating that architectural design fundamentally shapes bias patterns and mitigation effectiveness. Adaptive fusion mechanisms enable flexible responses to data interventions, while frequency reweighting offers robust improvements across architectures. Our findings establish that fairness interventions must be tailored to both model architecture and demographic characteristics, providing a systematic framework for developing equitable speech-based screening tools essential for reducing diagnostic disparities in cognitive healthcare.


【18】Geneses: Unified Generative Speech Enhancement and Separation
标题:起源:统一生成语音增强和分离
链接:https://arxiv.org/abs/2601.18456

作者:Kohei Asai,Wataru Nakata,Yuki Saito,Hiroshi Saruwatari
备注:Accepted to ICASSP 2025 workshop
摘要:真实世界的音频记录通常包含多个扬声器和各种降级,这限制了可用于构建最先进的语音处理模型的语音数据的数量和质量。虽然连接语音增强(SE)和语音分离(SS)以获得每个说话者的干净语音信号的端到端方法是有希望的,但是传统的SE-SS方法遭受超过加性噪声的复杂劣化。为此,我们提出了\textbf{Geneses},一个生成框架,以实现统一的,高质量的SE-SS。我们的Geneses利用潜在流匹配来估计每个说话人的干净的语音特征,使用多模态扩散Transformer条件下的自监督学习表示从嘈杂的混合。我们使用LibriTTS-R的两个扬声器混合物在两种条件下进行实验评估:仅加性噪声和复杂的退化。结果表明,基因显着优于传统的掩模为基础的SE-SS方法在各种客观的指标,对复杂的退化具有很高的鲁棒性。音频样本可在我们的演示页面。
摘要:Real-world audio recordings often contain multiple speakers and various degradations, which limit both the quantity and quality of speech data available for building state-of-the-art speech processing models. Although end-to-end approaches that concatenate speech enhancement (SE) and speech separation (SS) to obtain a clean speech signal for each speaker are promising, conventional SE-SS methods suffer from complex degradations beyond additive noise. To this end, we propose \textbf{Geneses}, a generative framework to achieve unified, high-quality SE--SS. Our Geneses leverages latent flow matching to estimate each speaker's clean speech features using multi-modal diffusion Transformer conditioned on self-supervised learning representation from noisy mixture. We conduct experimental evaluation using two-speaker mixtures from LibriTTS-R under two conditions: additive-noise-only and complex degradations. The results demonstrate that Geneses significantly outperforms a conventional mask-based SE--SS method across various objective metrics with high robustness against complex degradations. Audio samples are available in our demo page.


【19】Pisets: A Robust Speech Recognition System for Lectures and Interviews
标题:Pipets:用于讲座和面试的强大语音识别系统
链接:https://arxiv.org/abs/2601.18415

作者:Ivan Bondarenko,Daniil Grebenkin,Oleg Sedukhin,Mikhail Klementev,Roman Derunets,Lyudmila Budneva
摘要:这项工作为科学家和记者提供了一个语音到文本系统“Pisets”,该系统基于三组件架构,旨在提高语音识别的准确性,同时最大限度地减少与Whisper模型相关的错误和幻觉。该架构包括使用Wav2Vec2的初级识别,通过音频频谱图Transformer(AST)的假阳性过滤,以及通过Whisper的最终语音识别。课程学习方法的实施和多种俄语语音语料库的利用显着提高了系统的有效性。此外,还引入了先进的不确定性建模技术,有助于进一步提高转录质量。与WhisperX和通常的Whisper模型相比,所提出的方法确保了在各种声学条件下对长音频数据的稳健转录。“Pisets”系统的源代码可在GitHub上公开获取:https://github.com/bond005/pisets。
摘要:This work presents a speech-to-text system "Pisets" for scientists and journalists which is based on a three-component architecture aimed at improving speech recognition accuracy while minimizing errors and hallucinations associated with the Whisper model. The architecture comprises primary recognition using Wav2Vec2, false positive filtering via the Audio Spectrogram Transformer (AST), and final speech recognition through Whisper. The implementation of curriculum learning methods and the utilization of diverse Russian-language speech corpora significantly enhanced the system's effectiveness. Additionally, advanced uncertainty modeling techniques were introduced, contributing to further improvements in transcription quality. The proposed approaches ensure robust transcribing of long audio data across various acoustic conditions compared to WhisperX and the usual Whisper model. The source code of "Pisets" system is publicly available at GitHub: https://github.com/bond005/pisets.


【20】OCR-Enhanced Multimodal ASR Can Read While Listening
标题:OCR增强型多模式ASB可以边听边阅读
链接:https://arxiv.org/abs/2601.18393

作者:Junli Chen,Changli Tang,Yixuan Li,Guangzhi Sun,Chao Zhang
备注:4 pages, 2 figures. Submitted to ICASSP 2026
摘要:视觉信息,如电影中的字幕,通常有助于自动语音识别。在本文中,我们提出了甜甜圈耳语,视听ASR模型与双编码器,利用视觉信息,以提高语音识别性能的英语和汉语。Donut-Whisper通过交叉注意模块结合了线性和基于Q-Former的模态对齐结构的优点,生成更强大的视听特征。同时,我们提出了一个轻量级的知识蒸馏计划展示了使用视听模型教音频模型,以实现更好的性能的潜力。此外,我们提出了一个新的多语种视听语音识别数据集的基础上的电影剪辑包含中文和英文的分区。因此,与Donut和Whisper大型V3基线相比,Donut-Whisper在数据集的英语和中文分区上都取得了显着更好的性能。特别是,与Whisper ASR基线相比,英语和中文集分别实现了绝对WER减少5.75%和绝对CER减少16.5%。
摘要:Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve speech recognition performance in both English and Chinese. Donut-Whisper combines the advantage of the linear and the Q-Former-based modality alignment structures via a cross-attention module, generating more powerful audio-visual features. Meanwhile, we propose a lightweight knowledge distillation scheme showcasing the potential of using audio-visual models to teach audio-only models to achieve better performance. Moreover, we propose a new multilingual audio-visual speech recognition dataset based on movie clips containing both Chinese and English partitions. As a result, Donut-Whisper achieved significantly better performance on both English and Chinese partition of the dataset compared to both Donut and Whisper large V3 baselines. In particular, an absolute 5.75% WER reduction and a 16.5% absolute CER reduction were achieved on the English and Chinese sets respectively compared to the Whisper ASR baseline.


【21】LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
标题:LLM-ForcedAligner:一种基于非自回归和精确LLM-ForcedAligner的多语言和长格式语音强制对齐器
链接:https://arxiv.org/abs/2601.18220

作者:Bingshen Mu,Xian Shi,Xiong Wang,Hexin Liu,Jin Xu,Lei Xie
摘要:强制对齐(FA)预测语音中单词或字符的开始和结束时间戳,但现有的方法是特定于语言的,并且易于累积时间偏移。语音大语言模型(SLLM)的多语种语音理解和长序列处理能力,使其有望在多语种,跨语言,长格式语音设置FA。然而,直接将SLLM的下一个令牌预测范式应用于FA会导致幻觉和缓慢的推理。为了弥合这一差距,我们提出了LLM-ForcedAligner,将FA重新定义为一个插槽填充范式:时间戳被视为离散索引,特殊的时间戳标记作为插槽插入到成绩单中。SLLM以语音嵌入和带时隙的文本为条件,直接预测时隙处的时间索引。在训练期间,使用非移位输入和标签序列的因果注意力掩蔽允许每个时隙基于其自身和先前上下文预测其自己的时间戳索引,仅在时隙位置处计算损失。动态插槽插入使FA在任意位置。此外,支持非自回归推理,避免幻觉并提高速度。多语种、跨语种和长格式语音场景的实验表明,与现有方法相比,LLM-ForcedAligner实现了69%~78%的累积平均偏移相对减少。检查点和推理代码将在稍后发布。
摘要:Forced alignment (FA) predicts start and end timestamps for words or characters in speech, but existing methods are language-specific and prone to cumulative temporal shifts. The multilingual speech understanding and long-sequence processing abilities of speech large language models (SLLMs) make them promising for FA in multilingual, crosslingual, and long-form speech settings. However, directly applying the next-token prediction paradigm of SLLMs to FA results in hallucinations and slow inference. To bridge the gap, we propose LLM-ForcedAligner, reformulating FA as a slot-filling paradigm: timestamps are treated as discrete indices, and special timestamp tokens are inserted as slots into the transcript. Conditioned on the speech embeddings and the transcript with slots, the SLLM directly predicts the time indices at slots. During training, causal attention masking with non-shifted input and label sequences allows each slot to predict its own timestamp index based on itself and preceding context, with loss computed only at slot positions. Dynamic slot insertion enables FA at arbitrary positions. Moreover, non-autoregressive inference is supported, avoiding hallucinations and improving speed. Experiments across multilingual, crosslingual, and long-form speech scenarios show that LLM-ForcedAligner achieves a 69%~78% relative reduction in accumulated averaging shift compared with prior methods. The checkpoint and inference code will be released later.


【22】VIBEVOICE-ASR Technical Report
标题:VIBEVOICE-ASR技术报告
链接:https://arxiv.org/abs/2601.18184

作者:Zhiliang Peng,Jianwei Yu,Yaoyao Chang,Zilong Wang,Li Dong,Yingbo Hao,Yujie Tu,Chenyu Yang,Wenhui Wang,Songchen Xu,Yutao Sun,Hangbo Bao,Weijiang Xu,Yi Zhu,Zehua Wang,Ting Song,Yan Xia,Zewen Chi,Shaohan Huang,Liang Wang,Chuang Ding,Shuai Wang,Xie Chen,Furu Wei
摘要:本报告介绍了VibeVoice-ASR,这是一个基于VibeVoice的通用语音理解框架,旨在解决长格式音频中上下文碎片和多说话者复杂性的持续挑战(例如,会议,播客),尽管最近在短形式语音识别方面取得了进展,但仍然存在。与依赖音频分块的传统流水线方法不同,VibeVoice-ASR支持对长达60分钟的音频进行单次处理。它将自动语音识别、扬声器日志化和时间戳统一到一个端到端的生成任务中。此外,VibeVoice-ASR支持超过50种语言,不需要明确的语言设置,并且可以原生地处理话语内和话语间的代码切换。此外,我们引入了一个基于文本的上下文注入机制,允许用户提供定制的conetxt,显着提高特定领域的术语和多音字消歧的准确性。
摘要:This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation and multi-speaker complexity in long-form audio (e.g., meetings, podcasts) that remain despite recent advancements in short-form speech recognition. Unlike traditional pipelined approaches that rely on audio chunking, VibeVoice-ASRsupports single-pass processing for up to 60 minutes of audio. It unifies Automatic Speech Recognition, Speaker Diarization, and Timestamping into a single end-to-end generation task. In addition, VibeVoice-ASR supports over 50 languages, requires no explicit language setting, and natively handles code-switching within and across utterances. Furthermore, we introduce a prompt-based context injection mechanism that allows users to supply customized conetxt, significantly improving accuracy on domain-specific terminology and polyphonic character disambiguation.


【23】From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
标题:从人的语音到海洋信号:用于水声目标识别的语音大模型转换
链接:https://arxiv.org/abs/2601.18086

作者:Mengcheng Huang,Xue Zhou,Chen Xu,Dapeng Man
摘要:水声目标识别在海洋应用中起着重要的作用,但由于标记数据的有限性和海洋环境的复杂性,水声目标识别仍然具有挑战性。本文探讨了一个中心问题:语音大模型(SLM),训练在大量的人类语音语料库,可以有效地转移到水下声学?为了研究这一点,我们提出了UATR-SLM,一个简单的框架,重用语音特征管道,适应SLM作为声学编码器,并添加了一个轻量级的分类器。在DeepShip和ShipsEar基准测试上的实验表明,UATR-SLM实现了超过99%的域内准确率,在不同的信号长度上保持了强大的鲁棒性,并在跨域评估中达到了96.67%的准确率。这些结果突出了SLM到UATR的强大可转移性,建立了一个很有前途的范例,利用语音基础模型在水下声学。
摘要:Underwater acoustic target recognition (UATR) plays a vital role in marine applications but remains challenging due to limited labeled data and the complexity of ocean environments. This paper explores a central question: can speech large models (SLMs), trained on massive human speech corpora, be effectively transferred to underwater acoustics? To investigate this, we propose UATR-SLM, a simple framework that reuses the speech feature pipeline, adapts the SLM as an acoustic encoder, and adds a lightweight classifier.Experiments on the DeepShip and ShipsEar benchmarks show that UATR-SLM achieves over 99% in-domain accuracy, maintains strong robustness across variable signal lengths, and reaches up to 96.67% accuracy in cross-domain evaluation. These results highlight the strong transferability of SLMs to UATR, establishing a promising paradigm for leveraging speech foundation models in underwater acoustics.


【24】dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
标题:dLLM-ASB:一种基于LLM的更快扩散语音识别框架
链接:https://arxiv.org/abs/2601.17902

作者:Wenjie Tian,Bingshen Mu,Guobin Ma,Xuelong Geng,Zhixian Zhao,Lei Xie
摘要:基于大型语言模型(LLM)的自动语音识别(ASR)系统通过利用预训练的LLM作为解码器来实现卓越的性能,但其逐个令牌生成机制导致推理延迟随序列长度线性增长。与此同时,离散扩散大语言模型(dLLM)提供了一种有前途的替代方案,可以使用预训练的解码器生成高质量的并行序列。然而,直接应用本地面向文本的DLLM到ASR导致开放式文本生成和ASR所需的声学条件转录范式之间的根本不匹配。因此,它引入了不必要的困难和计算冗余,例如从纯噪声中去噪,不灵活的生成长度和固定的去噪步骤。我们提出了dLLM-ASR,一个有效的dLLM为基础的ASR框架,制定dLLM的解码作为一个事先指导和自适应去噪过程。它在初始化去噪过程之前利用ASR,并为序列长度提供锚点。在此基础上,长度自适应修剪动态删除冗余令牌,而基于置信度的去噪允许收敛令牌提前退出去噪循环,从而实现令牌级自适应计算。实验表明,dLLM-ASR实现的识别精度与基于自回归LLM的ASR系统相当,并提供了4.44$\times $的推理加速比,为ASR建立了一个实用而有效的范例。
摘要:Automatic speech recognition (ASR) systems based on large language models (LLMs) achieve superior performance by leveraging pretrained LLMs as decoders, but their token-by-token generation mechanism leads to inference latency that grows linearly with sequence length. Meanwhile, discrete diffusion large language models (dLLMs) offer a promising alternative, enabling high-quality parallel sequence generation with pretrained decoders. However, directly applying native text-oriented dLLMs to ASR leads to a fundamental mismatch between open-ended text generation and the acoustically conditioned transcription paradigm required by ASR. As a result, it introduces unnecessary difficulty and computational redundancy, such as denoising from pure noise, inflexible generation lengths, and fixed denoising steps. We propose dLLM-ASR, an efficient dLLM-based ASR framework that formulates dLLM's decoding as a prior-guided and adaptive denoising process. It leverages an ASR prior to initialize the denoising process and provide an anchor for sequence length. Building upon this prior, length-adaptive pruning dynamically removes redundant tokens, while confidence-based denoising allows converged tokens to exit the denoising loop early, enabling token-level adaptive computation. Experiments demonstrate that dLLM-ASR achieves recognition accuracy comparable to autoregressive LLM-based ASR systems and delivers a 4.44$\times$ inference speedup, establishing a practical and efficient paradigm for ASR.


【25】CaSNet: Compress-and-Send Network Based Multi-Device Speech Enhancement Model for Distributed Microphone Arrays
标题:CaSNet:基于压缩发送网络的分布式麦克风阵列多设备语音增强模型
链接:https://arxiv.org/abs/2601.17711

作者:Chengqian Jiang,Jie Zhang,Haoyin Yan
备注:this paper has been accept by ICASSP2026
摘要:分布式麦克风阵列(DMA)是下一代语音交互平台,但在噪声环境下仍需要语音增强来提高语音质量。现有的SE方法通常首先在融合中心(FC)从所有设备收集原始波形,然后设计多麦克风模型,这导致高带宽和能量成本。在这项工作中,我们提出了一个压缩和发送网络(CaSNet)的资源受限的DMA,其中一个麦克风作为FC和参考。每个其他设备将测量的原始数据编码为特征矩阵,然后通过奇异值分解(SVD)压缩该特征矩阵以产生更紧凑的表示。在FC处接收的特征通过相对于参考的交叉窗口查询来对齐,随后进行神经解码以产生空间相干增强的语音。在多个数据集上的实验表明,与未压缩的情况相比,CaSNet可以节省数据量,对性能的影响可以忽略不计。可复制的代码可在https://github.com/Jokejiangv/CaSNet上获得。
摘要:Distributed microphone array (DMA) is a promising next-generation platform for speech interaction, where speech enhancement (SE) is still required to improve the speech quality in noisy cases. Existing SE methods usually first gather raw waveforms at a fusion center (FC) from all devices and then design a multi-microphone model, causing high bandwidth and energy costs. In this work, we propose a \emph{Compress-and-Send Network (CaSNet)} for resource-constrained DMAs, where one microphone serves as the FC and reference. Each of other devices encodes the measured raw data into a feature matrix, which is then compressed by singular value decomposition (SVD) to produce a more compact representation. The received features at the FC are aligned via cross window query with respect to the reference, followed by neural decoding to yield spatially coherent enhanced speech. Experiments on multiple datasets show that the proposed CaSNet can save the data amount with a negligible impact on the performance compared to the uncompressed case. The reproducible code is available at https://github.com/Jokejiangv/CaSNet.


【26】Segment Length Matters: A Study of Segment Lengths on Audio Fingerprinting Performance
标题:段长度很重要:段长度对音频指纹性能的研究
链接:https://arxiv.org/abs/2601.17690

作者:Ziling Gong,Yunyan Ouyang,Iram Kamdar,Melody Ma,Hongjie Chen,Franck Dernoncourt,Ryan A. Rossi,Nesreen K. Ahmed
摘要:音频指纹识别提供了声学信号的可识别表示,其可以稍后用于识别和检索系统。为了获得有区别的表示,输入音频通常被分割成较短的时间间隔,允许提取和分析局部声学特征。现代神经方法通常对短的、固定持续时间的音频片段进行操作,然而片段持续时间的选择通常是抽象的,很少深入研究。在本文中,我们研究了段长度如何影响音频指纹的性能。我们扩展了现有的神经指纹架构,以采用不同的段长度,并评估不同段长度和查询持续时间的检索准确性。我们的研究结果表明,较短的段长度(0.5秒)通常可以实现更好的性能。此外,我们评估了LLM在推荐最佳片段长度方面的能力,这表明GPT-5-mini在三个研究的LLM中,在五个考虑因素中始终提供最佳建议。我们的研究结果为大规模神经音频检索系统中片段持续时间的选择提供了实际指导。
摘要:Audio fingerprinting provides an identifiable representation of acoustic signals, which can be later used for identification and retrieval systems. To obtain a discriminative representation, the input audio is usually segmented into shorter time intervals, allowing local acoustic features to be extracted and analyzed. Modern neural approaches typically operate on short, fixed-duration audio segments, yet the choice of segment duration is often made heuristically and rarely examined in depth. In this paper, we study how segment length affects audio fingerprinting performance. We extend an existing neural fingerprinting architecture to adopt various segment lengths and evaluate retrieval accuracy across different segment lengths and query durations. Our results show that short segment lengths (0.5-second) generally achieve better performance. Moreover, we evaluate LLM capacity in recommending the best segment length, which shows that GPT-5-mini consistently gives the best suggestions across five considerations among three studied LLMs. Our findings provide practical guidance for selecting segment duration in large-scale neural audio retrieval systems.


【27】BanglaRobustNet: A Hybrid Denoising-Attention Architecture for Robust Bangla Speech Recognition
标题:BanglaRobustNet:一种用于稳健孟加拉语语音识别的混合去噪-注意架构
链接:https://arxiv.org/abs/2601.17679

作者:Md Sazzadul Islam Ridoy,Mubaswira Ibnat Zidney,Sumi Akter,Md. Aminur Rahman
摘要:孟加拉语是使用最广泛的语言之一,在最先进的自动语音识别(ASR)研究中仍然代表性不足,特别是在嘈杂和说话者多样化的条件下。本文介绍了BanglaRobustNet,这是一个基于Wav 2 Vec-BERT的混合去噪-注意力框架,旨在解决这些挑战。该架构集成了一个基于扩散的去噪模块,以抑制环境噪声,同时保留孟加拉语特定的语音线索,以及一个上下文交叉注意模块,该模块可以在说话人嵌入上进行识别,以实现跨性别,年龄和方言的鲁棒性。经过端到端的训练,结合了CTC损失,语音一致性和扬声器对齐的复合目标,与Wav 2 Vec-BERT和Whisper基线相比,BanglaRobustNet实现了字错误率(WER)和字符错误率(CER)的大幅降低。对Mozilla Common Voice Bangla和增强噪声语音的评估证实了我们方法的有效性,将BanglaRobustNet建立为针对低资源,易受噪声影响的语言环境的强大ASR系统。
摘要:Bangla, one of the most widely spoken languages, remains underrepresented in state-of-the-art automatic speech recognition (ASR) research, particularly under noisy and speaker-diverse conditions. This paper presents BanglaRobustNet, a hybrid denoising-attention framework built on Wav2Vec-BERT, designed to address these challenges. The architecture integrates a diffusion-based denoising module to suppress environmental noise while preserving Bangla-specific phonetic cues, and a contextual cross-attention module that conditions recognition on speaker embeddings for robustness across gender, age, and dialects. Trained end-to-end with a composite objective combining CTC loss, phonetic consistency, and speaker alignment, BanglaRobustNet achieves substantial reductions in word error rate (WER) and character error rate (CER) compared to Wav2Vec-BERT and Whisper baselines. Evaluations on Mozilla Common Voice Bangla and augmented noisy speech confirm the effectiveness of our approach, establishing BanglaRobustNet as a robust ASR system tailored to low-resource, noise-prone linguistic settings.


【28】AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
标题:AVMeme考试:针对法学硕士背景和文化知识和思维的多模式多语言多文化基准
链接:https://arxiv.org/abs/2601.17645

作者:Xilin Jiang,Qiaolin Wang,Junkai Wu,Xiaomin He,Zhongweiyang Xu,Yinghao Ma,Minshuo Piao,Kaiyi Yang,Xiuwen Zheng,Riki Shimizu,Yicong Chen,Arsalan Firoozi,Gavin Mischler,Sukru Samet Dindar,Richard Antonello,Linyang He,Tsun-An Hsieh,Xulin Fan,Yulun Wu,Yuesheng Ma,Chaitanya Amballa,Weixiong Chen,Jiarui Hai,Ruisi Li,Vishal Choudhari,Cong Han,Yinghao Aaron Li,Adeen Flinker,Mounya Elhilali,Emmanouil Benetos,Mark Hasegawa-Johnson,Romit Roy Choudhury,Nima Mesgarani
备注:avmemeexam.github.io/public
摘要:互联网视听剪辑通过随时间变化的声音和动作传达意义,这超出了文本本身所能代表的范围。为了检验人工智能模型是否能够在人类文化背景下理解这些信号,我们引入了AVMeme Exam,这是一个人工策划的基准测试,包含一千多个标志性的互联网声音和视频,包括语音、歌曲、音乐和音效。每个模因都配有一个独特的问答,评估从表面内容到上下文和情感到用法和世界知识的理解水平,以及原始年份,成绩单,摘要和敏感性等元数据。我们使用这个基准系统地评估了最先进的多模态大型语言模型(MLLM)以及人类参与者。我们的研究结果揭示了一个一致的局限性:目前的模型在无文本音乐和声音效果上表现不佳,并且与表面内容相比,很难在上下文和文化中进行思考。这些发现突出了与人类一致的多模态智能的一个关键差距,并呼吁模型能够超越他们所听到和看到的表面进行上下文和文化感知。项目页面:avmemeexam.github.io/public
摘要:Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand such signals in human cultural contexts, we introduce AVMeme Exam, a human-curated benchmark of over one thousand iconic Internet sounds and videos spanning speech, songs, music, and sound effects. Each meme is paired with a unique Q&A assessing levels of understanding from surface content to context and emotion to usage and world knowledge, along with metadata such as original year, transcript, summary, and sensitivity. We systematically evaluate state-of-the-art multimodal large language models (MLLMs) alongside human participants using this benchmark. Our results reveal a consistent limitation: current models perform poorly on textless music and sound effects, and struggle to think in context and in culture compared to surface content. These findings highlight a key gap in human-aligned multimodal intelligence and call for models that can perceive contextually and culturally beyond the surface of what they hear and see. Project page: avmemeexam.github.io/public


【29】Home Health System Deployment Experience for Geriatric Care Remote Monitoring
标题:老年护理远程监控的家庭健康系统部署经验
链接:https://arxiv.org/abs/2601.17608

作者:Dong Yoon Lee,Alyssa Weakley,Hui Wei,Daniel Cardona,Shijia Pan
摘要:为了支持就地养老,成年子女经常从远处照顾年迈的父母。这些非正式的护理人员需要即插即用的远程护理解决方案,以保护隐私,实现实时活动监控和直观的可操作信息。这篇简短的论文介绍了远程监控系统部署经验的三次迭代以及在老年4M框架(最重要的是心理状态,移动性和药物治疗)指导下硬件,建模和用户界面的迭代改进。一个LLM辅助解决方案的开发,以平衡用户体验(隐私保护,即插即用)和系统性能。
摘要:To support aging-in-place, adult children often provide care to their aging parents from a distance. These informal caregivers desire plug-and-play remote care solutions for privacy-preserving continuous monitoring that enabling real-time activity monitoring and intuitive, actionable information. This short paper presents insights from three iterations of deployment experience for remote monitoring system and the iterative improvement in hardware, modeling, and user interface guided by the Geriatric 4Ms framework (matters most, mentation, mobility, and medication). An LLM-assisted solution is developed to balance user experience (privacy-preserving, plug-and-play) and system performance.


【30】EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding
标题:EuleroDec:一种用于高效鲁棒音频编码的复值RVQ-VAE
链接:https://arxiv.org/abs/2601.17517

作者:Luca Cerovaz,Michele Mancusi,Emanuele Rodolà
备注:Accepted at ICASSP 2026
摘要:音频编解码器通过将PCM音频压缩到带宽友好的比特率,为离散音乐生成建模、音乐流和沉浸式媒体提供支持。最近的工作已经倾向于在谱域中进行处理;然而,谱图域通常与相位建模相斗争,相位建模自然是复值的。大多数频域神经编解码器要么忽略相位信息,要么将其编码为两个独立的实值通道,从而限制了空间保真度。这就需要引入对抗性鉴别器,以牺牲收敛速度和训练稳定性来补偿音频信号的不足表示能力。在这项工作中,我们引入了一个端到端的复值RVQ-VAE音频编解码器,它在整个分析-量化-合成流水线中保留了幅度-相位耦合,并去除了对抗性鉴别器和扩散后滤波器。在没有GANs或扩散的情况下,我们在域内匹配或超过更长时间的训练基线,并在相位相干性和波形保真度方面达到SOTA域外性能。与训练数十万步的标准基线相比,我们的模型将训练预算降低了一个数量级,在保持高感知质量的同时,计算效率明显提高。
摘要:Audio codecs power discrete music generative modelling, music streaming, and immersive media by shrinking PCM audio to bandwidth-friendly bitrates. Recent works have gravitated towards processing in the spectral domain; however, spectrogram domains typically struggle with phase modeling, which is naturally complex-valued. Most frequency-domain neural codecs either disregard phase information or encode it as two separate real-valued channels, limiting spatial fidelity. This entails the need to introduce adversarial discriminators at the expense of convergence speed and training stability to compensate for the inadequate representation power of the audio signal. In this work we introduce an end-to-end complex-valued RVQ-VAE audio codec that preserves magnitude-phase coupling across the entire analysis-quantization-synthesis pipeline and removes adversarial discriminators and diffusion post-filters. Without GANs or diffusion, we match or surpass much longer-trained baselines in-domain and reach SOTA out-of-domain performance on phase coherence and waveform fidelity. Compared to standard baselines that train for hundreds of thousands of steps, our model, which reduces the training budget by an order of magnitude, is markedly more compute-efficient while preserving high perceptual quality.


【31】Window Size Versus Accuracy Experiments in Voice Activity Detectors
标题:语音活动检测器中的窗口大小与准确性实验
链接:https://arxiv.org/abs/2601.17270

作者:Max McKinnon,Samir Khaki,Chandan KA Reddy,William Huang
摘要:语音活动检测(VAD)在实现语音识别等应用方面起着至关重要的作用。我们分析了窗口大小对三种VAD算法的准确性的影响:Silero,WebRTC和均方根(RMS)在一组不同的现实世界的数字音频流。我们还探讨了在每个VAD输出端上使用迟滞。研究结果为优化VAD系统提供了实际参考。Silero的性能明显优于WebRTC和RMS,滞后为WebRTC提供了优势。
摘要:Voice activity detection (VAD) plays a vital role in enabling applications such as speech recognition. We analyze the impact of window size on the accuracy of three VAD algorithms: Silero, WebRTC, and Root Mean Square (RMS) across a set of diverse real-world digital audio streams. We additionally explore the use of hysteresis on top of each VAD output. Our results offer practical references for optimizing VAD systems. Silero significantly outperforms WebRTC and RMS, and hysteresis provides a benefit for WebRTC.


【32】Sink or SWIM: Tackling Real-Time ASR at Scale
标题:水槽或游泳:大规模解决实时ZR问题
链接:https://arxiv.org/abs/2601.17097

作者:Federico Bruzzone,Walter Cazzola,Matteo Brancaleoni,Dario Pellegrino
备注:14 pages, 7 figures
摘要:实时自动语音识别系统越来越多地集成到交互式应用中,从语音助理到实时转录服务。然而,扩展这些系统以支持多个并发客户端,同时保持低延迟和高准确性仍然是一个重大挑战。在这项工作中,我们提出了SWIM,这是一种建立在OpenAI Whisper模型之上的新型实时ASR系统,可以实现真正的模型级并行化,以实现可扩展的多语言转录。SWIM支持多个并发音频流,而无需修改底层模型。它引入了一种缓冲区合并策略,在确保高效资源使用的同时保持转录保真度。我们评估SWIM在多客户端设置-扩展到20个并发用户-并表明,它提供了准确的实时transmittance在英语,意大利语和西班牙语,同时保持低延迟和高吞吐量。虽然Whisper-Streaming在单客户端、仅英语设置中实现了约8.2%的单词错误率和约3.4秒的平均延迟,但SWIM将此功能扩展到多语言、多客户端环境。它保持了相当的准确性和显著更低的延迟(5个客户端约2.4秒),并继续有效地扩展到20个并发客户端,而不会降低转录质量和提高整体吞吐量。我们的方法通过提高动态多用户环境中的鲁棒性和效率来推进可扩展的ASR。
摘要:Real-time automatic speech recognition systems are increasingly integrated into interactive applications, from voice assistants to live transcription services. However, scaling these systems to support multiple concurrent clients while maintaining low latency and high accuracy remains a major challenge. In this work, we present SWIM, a novel real-time ASR system built on top of OpenAI's Whisper model that enables true model-level parallelization for scalable, multilingual transcription. SWIM supports multiple concurrent audio streams without modifying the underlying model. It introduces a buffer merging strategy that maintains transcription fidelity while ensuring efficient resource usage. We evaluate SWIM in multi-client settings -- scaling up to 20 concurrent users -- and show that it delivers accurate real-time transcriptions in English, Italian, and Spanish, while maintaining low latency and high throughput. While Whisper-Streaming achieves a word error rate of approximately 8.2% with an average delay of approximately 3.4 s in a single-client, English-only setting, SWIM extends this capability to multilingual, multi-client environments. It maintains comparable accuracy with significantly lower delay -- around 2.4 s with 5 clients -- and continues to scale effectively up to 20 concurrent clients without degrading transcription quality and increasing overall throughput. Our approach advances scalable ASR by improving robustness and efficiency in dynamic, multi-user environments.


【33】SonoEdit: Null-Space Constrained Knowledge Editing for Pronunciation Correction in LLM-Based TTS
标题:SonoEdit:零空间约束知识编辑,用于基于LLM的TTC中的发音纠正
链接:https://arxiv.org/abs/2601.17086

作者:Ayush Pratap Singh,Harshit Singh,Nityanand Mathur,Akshat Mandloi,Sudarshan Kamath
摘要:神经文本到语音(TTS)系统系统地错误发音低资源专有名词,特别是非英语名称,品牌和地理位置,由于他们在主要是英语培训语料库的代表性不足。现有的解决方案通常依赖于昂贵的多语言数据收集、监督微调或手动语音注释,这限制了TTS系统在语言多样性环境中的部署。我们介绍SonoEdit,一种模型编辑技术,它可以在不进行再训练的情况下外科手术地纠正预训练TTS模型中的发音错误。而不是昂贵的微调或显式音素注入,我们提出了一种基于零空间发音编辑的简约替代方案,它执行单次参数更新以修改特定单词的发音,同时可证明保留所有其他模型行为。我们首先适应声学因果跟踪,以确定负责文本到发音映射的Transformer层。然后,我们应用零空间约束编辑来计算一个封闭形式的权重更新,该权重更新纠正目标发音,同时保持与控制一般语音生成的子空间在数学上正交。这种受约束的更新将模型的声学输出转向所需的发音样本,同时保证在保留的语音语料库上的零一阶变化。
摘要:Neural text-to-speech (TTS) systems systematically mispronounce low-resource proper nouns, particularly non-English names, brands, and geographic locations, due to their underrepresentation in predominantly English training corpora. Existing solutions typically rely on expensive multilingual data collection, supervised finetuning, or manual phonetic annotation, which limits the deployment of TTS systems in linguistically diverse settings. We introduce SonoEdit, a model editing technique that surgically corrects pronunciation errors in pre-trained TTS models without retraining. Instead of costly finetuning or explicit phoneme injection, we propose a parsimonious alternative based on Null-Space Pronunciation Editing, which performs a single-shot parameter update to modify the pronunciation of specific words while provably preserving all other model behavior. We first adapt Acoustic Causal Tracing to identify the Transformer layers responsible for text-to-pronunciation mapping. We then apply Null-Space Constrained Editing to compute a closed-form weight update that corrects the target pronunciation while remaining mathematically orthogonal to the subspace governing general speech generation. This constrained update steers the model's acoustic output toward a desired pronunciation exemplar while guaranteeing zero first-order change on a preserved speech corpus.


机器翻译由腾讯交互翻译提供,仅供参考