今日论文合集:cs.SD语音12篇,eess.AS音频处理18篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Modeling Overlapped Speech with Shuffles
标题:用Shuffle建模重叠语音
链接:https://arxiv.org/abs/2603.17769

作者:Matthew Wiesner,Samuele Cornell,Alexander Polok,Lucas Ondel Yang,Lukáš Burget,Sanjeev Khudanpur
摘要:我们建议使用shuffles对并行数据流(如重叠语音)进行建模。具体而言,本文展示了如何洗牌产品和偏序有限状态自动机(FSA)可以用于对齐和扬声器属性的转录重叠的语音。我们使用这些FSA的总得分作为损失函数进行训练,在子字、字和短语级别上对重叠序列的所有可能的序列化进行边缘化。为了减少图的大小,我们施加时间约束,通过构建偏序FSA。我们通过直接建模(token,speaker)元组来解决说话人属性。通过混洗产物FSA的维特比对准直接实现一遍对准。我们评估合成LibriSpeech重叠的性能。据我们所知,这是第一个能够对多人录音进行单次对齐的算法。所有算法都是使用k2 / Icefall实现的。
摘要:We propose to model parallel streams of data, such as overlapped speech, using shuffles. Specifically, this paper shows how the shuffle product and partial order finite-state automata (FSAs) can be used for alignment and speaker-attributed transcription of overlapped speech. We train using the total score on these FSAs as a loss function, marginalizing over all possible serializations of overlapping sequences at subword, word, and phrase levels. To reduce graph size, we impose temporal constraints by constructing partial order FSAs. We address speaker attribution by modeling (token, speaker) tuples directly. Viterbi alignment through the shuffle product FSA directly enables one-pass alignment. We evaluate performance on synthetic LibriSpeech overlaps. To our knowledge, this is the first algorithm that enables single-pass alignment of multi-talker recordings. All algorithms are implemented using k2 / Icefall.


【2】Zipper-LoRA: Dynamic Parameter Decoupling for Speech-LLM based Multilingual Speech Recognition
标题:Zipper-LoRA:基于Speech-LLM的多语言语音识别的动态参数脱钩
链接:https://arxiv.org/abs/2603.17558

作者:Yuxiang Mei,Delai Qiu,Shengping Liu,Jiaen Liang,Yanhua Long
备注:13 pages, 8 figures
摘要:语音大语言模型(Speech-LLM)通过将语音编码器与大语言模型对齐,已经成为自动语音识别(ASR)的一种强大方法。然而,使这些系统适应数据分布不平衡的多语言环境仍然具有挑战性。在这种情况下,一个稳定性和可塑性的困境往往会出现:完全共享的参数高效微调(PEFT)可能会导致负面的语言间干扰的代表性不足的语言,而完全特定于语言的调整限制了跨语言的有益的知识转移所需的低资源的任务。为了解决这个问题,我们提出了Zipper-LoRA,这是一种新的等级级别解耦框架,具有三种变体(静态,硬和软),可以从共享和特定语言的子空间动态合成LoRA更新。通过使用轻量级的语言条件路由器,Zipper-LoRA在LoRA等级级别动态控制每个子空间的贡献,从而在语言兼容的情况下实现细粒度共享,并在发生冲突时严格解耦。为了进一步稳定不平衡数据下的优化,我们提出了一个两阶段的训练策略,其中包含一个初始B热启动,可以显着加速收敛。在12种语言混合资源环境下的实验表明,Zipper-LoRA的性能始终优于完全共享和独立的基线,特别是在资源极低的情况下。此外,我们证明,这些收益是强大的分块和非分块编码器配置,确认该框架的可靠性,为实际的,大规模的多语言ASR。我们的代码和数据将在https://github.com/YuCeong-May/Zipper-LoRA上提供,以供重复使用。
摘要:Speech Large Language Models (Speech-LLMs) have emerged as a powerful approach for automatic speech recognition (ASR) by aligning speech encoders with large language models. However, adapting these systems to multilingual settings with imbalanced data distributions remains challenging. In such scenarios, a stability-plasticity dilemma often arises: fully shared Parameter-Efficient Fine-Tuning (PEFT) can cause negative inter-lingual interference for under-represented languages, while fully language-specific tuning limits the cross-lingual beneficial knowledge transfer needed for low-resource tasks. To address this, we propose Zipper-LoRA, a novel rank-level decoupling framework with three variants (Static, Hard, and Soft) that dynamically synthesizes LoRA updates from shared and language-specific subspaces. By using a lightweight language-conditioned router, Zipper-LoRA dynamically controls the contribution of each subspace at the LoRA rank level, enabling fine-grained sharing where languages are compatible and strict decoupling when conflicts occur. To further stabilize optimization under imbalanced data, we propose a two-stage training strategy with an Initial-B warm start that significantly accelerates convergence. Experiments on a 12-language mixed-resource setting show that Zipper-LoRA consistently outperforms both fully shared and independent baselines, particularly in extremely low-resource scenarios. Moreover, we demonstrate that these gains are robust across both chunked and non-chunked encoder configurations, confirming the framework's reliability for practical, large-scale multilingual ASR. Our code and data will be available at https://github.com/YuCeong-May/Zipper-LoRA for reproducibility.


【3】CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization
标题:CineSRD:利用视觉、声学和语言线索进行开放世界视觉媒体扬声器对话
链接:https://arxiv.org/abs/2603.16966

作者:Liangbin Huang,Xiaohua Liao,Chaoqun Cui,Shijing Wang,Zhaolong Huang,Yanlong Du,Wenji Mao
备注:Accepted to CVPR 2026
摘要:传统的发言者日记系统主要集中在诸如会议和采访之类的受约束的场景,其中发言者的数量有限并且声学条件相对干净。为了探索开放世界的扬声器日记,我们将这项任务扩展到视觉媒体领域,包括复杂的视听节目,如电影和电视剧。这种新的设置带来了几个挑战,包括长格式视频理解,大量的扬声器,音频和视觉线索之间的跨模态干扰,以及不受控制的野外变化。为了解决这些挑战,我们提出了电影扬声器注册和日记(CineSRD),一个统一的多模式框架,利用视觉,声学和语言线索,从视频,语音和字幕的扬声器注释。CineSRD首先执行视觉锚点聚类以注册初始扬声器,然后集成音频语言模型以进行扬声器转向检测,改进注释并补充未注册的屏幕外扬声器。此外,我们构建并发布了一个专门针对视觉媒体的演讲者日志基准,包括中文和英文节目。实验结果表明,CineSRD在所提出的基准测试和传统数据集上取得了优异的性能,验证了其在开放世界视觉媒体环境中的鲁棒性和通用性。
摘要:Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore open-world speaker diarization, we extend this task to the visual media domain, encompassing complex audiovisual programs such as films and TV series. This new setting introduces several challenges, including long-form video understanding, a large number of speakers, cross-modal asynchrony between audio and visual cues, and uncontrolled in-the-wild variability. To address these challenges, we propose Cinematic Speaker Registration & Diarization (CineSRD), a unified multimodal framework that leverages visual, acoustic, and linguistic cues from video, speech, and subtitles for speaker annotation. CineSRD first performs visual anchor clustering to register initial speakers and then integrates an audio language model for speaker turn detection, refining annotations and supplementing unregistered off-screen speakers. Furthermore, we construct and release a dedicated speaker diarization benchmark for visual media that includes Chinese and English programs. Experimental results demonstrate that CineSRD achieves superior performance on the proposed benchmark and competitive results on conventional datasets, validating its robustness and generalizability in open-world visual media settings.


【4】Music Source Restoration with Ensemble Separation and Targeted Reconstruction
标题:整体分离和有针对性重建的音乐源泉恢复
链接:https://arxiv.org/abs/2603.16926

作者:Xinlong Deng,Yu Xia,Jie Jiang
摘要:首届音乐源恢复(MSR)挑战赛的目标是从完全混合和掌握的音乐中恢复原始的、未经处理的音乐。与传统的音乐源分离不同,MSR需要反转复杂的制作过程,例如均衡,压缩,混响和其他真实世界的降级。为了解决MSR,我们提出了一个两阶段系统。首先,预先训练的分离模型的集合产生初步的源估计。然后,一组预先训练的基于BSRNN的恢复模型执行有针对性的重建,以细化这些估计。在官方的MSR基准测试中,我们的系统在所有指标上都超过了基线,在所有提交中排名第二。该代码可在https://github.com/xinghour/Music-source-restoration-CUPAudioGroup上获得
摘要:The Inaugural Music Source Restoration (MSR) Challenge targets the recovery of original, unprocessed stems from fully mixed and mastered music. Unlike conventional music source separation, MSR requires reversing complex production processes such as equalization, compression, reverberation, and other real-world degradations. To address MSR, we propose a two-stage system. First, an ensemble of pre-trained separation models produces preliminary source estimates. Then a set of pre-trained BSRNN-based restoration models performs targeted reconstruction to refine these estimates. On the official MSR benchmark, our system surpasses the baselines on all metrics, ranking second among all submissions. The code is available at https://github.com/xinghour/Music-source-restoration-CUPAudioGroup


【5】Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
标题:用于语音深度伪造检测的量化器感知分层神经编解码器建模
链接:https://arxiv.org/abs/2603.16914

作者:Jinyang Wu,Zihan Pan,Qiquan Zhang,Sailor Hardik Bhupendra,Soumik Mondal
备注:5 pages, 3 figures
摘要:神经音频编解码器通过残差矢量量化(RVQ)离散化语音,在量化器之间形成从粗到细的层次结构。虽然编解码器模型已经被探索用于表示学习,但它们的离散结构在语音深度伪造检测中仍然没有得到充分利用。特别是,不同的量化级别捕获互补的声学线索,其中早期的量化器编码粗糙的结构,稍后的量化器细化残留的细节,揭示合成伪影。现有的系统要么依赖于连续编码器功能,要么忽略这个量化器级别的层次结构。我们提出了一个层次感知的表示学习框架,通过可学习的全局加权模型量化器级的贡献,使结构化的编解码器表示与法医线索对齐。保持语音编码器骨干冻结并仅更新4.4%的附加参数,我们的方法在ASVspoof 2019上实现了46.2%的相对EER降低,在ASVspoof 5上实现了13.9%的相对EER降低。
摘要:Neural audio codecs discretize speech via residual vector quantization (RVQ), forming a coarse-to-fine hierarchy across quantizers. While codec models have been explored for representation learning, their discrete structure remains underutilized in speech deepfake detection. In particular, different quantization levels capture complementary acoustic cues, where early quantizers encode coarse structure and later quantizers refine residual details that reveal synthesis artifacts. Existing systems either rely on continuous encoder features or ignore this quantizer-level hierarchy. We propose a hierarchy-aware representation learning framework that models quantizer-level contributions through learnable global weighting, enabling structured codec representations aligned with forensic cues. Keeping the speech encoder backbone frozen and updating only 4.4% additional parameters, our method achieves relative EER reductions of 46.2% on ASVspoof 2019 and 13.9% on ASVspoof5 over strong baselines.


【6】Amanous: Distribution-Switching for Superhuman Piano Density on Disklavier
标题:Amanous:《尤利西斯》上超人钢琴密度的分布转换
链接:https://arxiv.org/abs/2603.16890

作者:Joonhyung Bae
摘要:自动钢琴使音符密度,复调和寄存器的变化远远超出了人类的物理限制,但三个占主导地位的传统组成这样的纹理-南卡罗的速度佳能,Xenakis的随机分布,L系统语法-已经孤立地发展。本文介绍了Amanous,一个硬件感知的组成系统,雅马哈pullavier,统一这些方法,通过分布切换:L-系统符号选择不同的分布制度,而不仅仅是调制参数在一个固定的家庭。报告了四项贡献。(1)一个四层架构(符号,参数,数字,物理)产生统计上不同的部分与大的效果大小(d = 3.70-5.34),每层的退化和消融实验验证。(2)一个硬件抽象层形式化速度相关的延迟和密钥重置约束,将超人的纹理保持在机器人的可致信封内。(3)密度扫描揭示了在24-30个音符/秒处的计算饱和过渡(自举95%CI:23.3-50.0),超过该饱和过渡,单域旋律度量失去辨别能力,并且跨域耦合变得必要。(4)一个收敛点演算操作tempo-canon几何作为一个控制接口,使收敛事件触发分布开关连接宏观时间结构的微观层次的纹理。所有结果都是计算的,心理声学验证协议提出了未来的工作。该流水线已部署在一个物理平台上,展示了算法的自一致性和亚毫秒级的软件精度。补充材料(摘录1-4):https://www.amanous.xyz。源代码:https://github.com/joonhyungbae/Amanous。
摘要:The automated piano enables note densities, polyphony, and register changes far beyond human physical limits, yet the three dominant traditions for composing such textures--Nancarrow's tempo canons, Xenakis's stochastic distributions, and L-system grammars--have developed in isolation. This paper presents Amanous, a hardware-aware composition system for Yamaha Disklavier that unifies these methodologies through distribution-switching: L-system symbols select distinct distributional regimes rather than merely modulating parameters within a fixed family. Four contributions are reported. (1) A four-layer architecture (symbolic, parametric, numeric, physical) produces statistically distinct sections with large effect sizes (d = 3.70-5.34), validated by per-layer degradation and ablation experiments. (2) A hardware abstraction layer formalizes velocity-dependent latency and key reset constraints, keeping superhuman textures within the Disklavier's actuable envelope. (3) A density sweep reveals a computational saturation transition at 24-30 notes/s (bootstrap 95% CI: 23.3-50.0), beyond which single-domain melodic metrics lose discriminative power and cross-domain coupling becomes necessary. (4) A convergence point calculus operationalizes tempo-canon geometry as a control interface, enabling convergence events to trigger distribution switches linking macro-temporal structure to micro-level texture. All results are computational; a psychoacoustic validation protocol is proposed for future work. The pipeline has been deployed on a physical Disklavier, demonstrating algorithmic self-consistency and sub-millisecond software precision. Supplementary materials (Excerpts 1-4): https://www.amanous.xyz. Source code: https://github.com/joonhyungbae/Amanous.


【7】Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
标题:针对多方面、多评分者L2阅读言语评估的SpeechLLM的文字引导微调
链接:https://arxiv.org/abs/2603.16889

作者:Aditya Kamlesh Parikh,Cristian Tejedor-Garcia,Catia Cucchiarini,Helmer Strik
备注:Accepted to LREC 2026. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants, which is financed by the Dutch Research Council (NWO)
摘要:第二语言(L2)语音的可靠和可解释的自动评估仍然是一个核心挑战,因为大型语音语言模型(SpeechLLM)通常难以与人类评分员的细微变化保持一致。为了解决这个问题,我们引入了一个规则引导的推理框架,明确编码多方面的人类评估标准:准确性,流畅性和韵律,同时校准模型的不确定性,以捕捉自然的评级变化。我们微调Qwen 2-Audio-7 B-Instruct模型使用多评分人的判断,并开发了一个不确定性校准的回归方法,由保形校准支持可解释的置信区间。我们的高斯不确定性建模和共形校准方法实现了与人类评级最强的一致性,优于回归和分类基线。该模型可靠地评估流畅性和韵律,同时突出了评估准确性的固有困难。总之,这些结果表明,规则引导、不确定性校准的推理为基于SpeechLLM的可信且可解释的语音评估提供了一条原则性路径。
摘要:Reliable and interpretable automated assessment of second-language (L2) speech remains a central challenge, as large speech-language models (SpeechLLMs) often struggle to align with the nuanced variability of human raters. To address this, we introduce a rubric-guided reasoning framework that explicitly encodes multi-aspect human assessment criteria: accuracy, fluency, and prosody, while calibrating model uncertainty to capture natural rating variability. We fine-tune the Qwen2-Audio-7B-Instruct model using multi-rater human judgments and develop an uncertainty-calibrated regression approach supported by conformal calibration for interpretable confidence intervals. Our Gaussian uncertainty modeling and conformal calibration approach achieves the strongest alignment with human ratings, outperforming regression and classification baselines. The model reliably assesses fluency and prosody while highlighting the inherent difficulty of assessing accuracy. Together, these results demonstrate that rubric-guided, uncertainty-calibrated reasoning offers a principled path toward trustworthy and explainable SpeechLLM-based speech assessment.


【8】Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
标题:Wav2Vec语音识别神经网络的空中白盒攻击
链接:https://arxiv.org/abs/2603.16972

作者:Protopopov Alexey
备注:9 pages, 5 figures, 1 table
摘要:基于神经网络的自动语音识别系统容易受到以恶意方式改变传输的对抗性攻击。该领域最近的工作集中在使攻击在空中场景中工作,然而这种攻击通常可以通过人类听觉检测到,限制了它们的潜在应用。在目前的工作中,我们探讨了不同的方法,使空中攻击不易察觉,以及这些方法对攻击的有效性的影响。
摘要:Automatic speech recognition systems based on neural networks are vulnerable to adversarial attacks that alter transcriptions in a malicious way. Recent works in this field have focused on making attacks work in over-the-air scenarios, however such attacks are typically detectable by human hearing, limiting their potential applications. In the present work we explore different approaches of making over-the-air attacks less detectable, as well as the impact these approaches have on the attacks' effectiveness.


【9】The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
标题:言语背后的声音:量化SpeechLLM中的交叉偏见
链接:https://arxiv.org/abs/2603.16941

作者:Shree Harsha Bokkahalli Satish,Christoph Minixhofer,Maria Teleki,James Caverlee,Ondřej Klejch,Peter Bell,Gustav Eje Henter,Éva Székely
备注:5 pages, 3 figures, 1 table, Submitted to Interspeech 2026
摘要:语音大语言模型(SpeechLLM)直接处理口语输入,保留先前在级联管道中删除的口音和感知性别等线索。这在响应中引入了说话者身份相关的变化。我们在三个SpeechLLM中对口音和性别偏见进行了大规模的交叉评估,使用了六种英语口音和两种性别呈现的2,880种受控交互,通过语音克隆保持语言内容不变。使用逐点LLM判断评级,成对比较和最佳-最差缩放与人类验证,我们检测一致的差异。东欧口音的讲话得到较低的帮助分数,特别是女性提出的声音。这种偏见是隐性的:反应仍然是礼貌的,但在帮助方面有所不同。虽然LLM法官捕捉这些偏见的方向趋势,人类评估者表现出显着更高的敏感性,发现更尖锐的交叉差异。
摘要:Speech Large Language Models (SpeechLLMs) process spoken input directly, retaining cues such as accent and perceived gender that were previously removed in cascaded pipelines. This introduces speaker identity dependent variation in responses. We present a large-scale intersectional evaluation of accent and gender bias in three SpeechLLMs using 2,880 controlled interactions across six English accents and two gender presentations, keeping linguistic content constant through voice cloning. Using pointwise LLM-judge ratings, pairwise comparisons, and Best-Worst Scaling with human validation, we detect consistent disparities. Eastern European-accented speech receives lower helpfulness scores, particularly for female-presenting voices. The bias is implicit: responses remain polite but differ in helpfulness. While LLM judges capture the directional trend of these biases, human evaluators exhibit significantly higher sensitivity, uncovering sharper intersectional disparities.


【10】Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
标题:超越深度学习:使用神经组合的语音分割和电话分类
链接:https://arxiv.org/abs/2603.16923

作者:Trevor Adelson,Vidhyasaharan Sethu,Ting Dang
备注:Submitted to Interspeech 2026. 9 Pages
摘要:深度学习主导语音处理,但依赖于大量数据集、全局反向传播引导的权重更新,并产生纠缠表示。装配演算(AC),通过赫布可塑性和赢家通吃的竞争模型稀疏神经元组件,提供了一个生物接地替代,但以前的工作集中在离散的符号输入。我们引入了一个基于AC的语音处理框架,该框架通过结合三个关键贡献直接对连续语音进行操作:(i)神经编码,使用概率mel二进制化和人口编码的MFCC将语音转换为装配兼容的尖峰模式;(ii)多区域架构,跨层次时间尺度和类组织装配;(iii)跨区域更新方案,用于下游任务。应用于两个核心任务的边界检测和段分类,我们的框架检测电话(F1=0.69)和单词(F1=0.61)的边界没有任何权重训练,并达到47.5%和45.1%的准确率电话和命令识别。这些结果表明,基于AC的动力系统是深度学习语音处理的可行替代方案。
摘要:Deep learning dominates speech processing but relies on massive datasets, global backpropagation-guided weight updates, and produces entangled representations. Assembly Calculus (AC), which models sparse neuronal assemblies via Hebbian plasticity and winner-take-all competition, offers a biologically grounded alternative, yet prior work focused on discrete symbolic inputs. We introduce an AC-based speech processing framework that operates directly on continuous speech by combining three key contributions:(i) neural encoding that converts speech into assembly-compatible spike patterns using probabilistic mel binarisation and population-coded MFCCs; (ii) a multi-area architecture organising assemblies across hierarchical timescales and classes; and (iii) cross-area update schemes for downstream tasks. Applied to two core tasks of boundary detection and segment classification, our framework detects phone (F1=0.69) and word (F1=0.61) boundaries without any weight training, and achieves 47.5% and 45.1% accuracy on phone and command recognition. These results show that AC-based dynamical systems are a viable alternative to deep learning for speech processing.


【11】Learnable Pulse Accumulation for On-Device Speech Recognition: How Much Attention Do You Need?
标题:设备上语音识别的可学习脉搏累积:您需要多少关注?
链接:https://arxiv.org/abs/2603.16922

作者:Yakov Pyotr Shkolnikov
摘要:自我注意力与序列长度成二次关系,限制了边缘设备上基于变换器的语音模型。我们引入了可学习脉冲累加器(LPA),这是一个O(n)的替换,它用学习的门函数(内容相关的矩形脉冲,周期性窗口和位置相关的基函数)代替了关键字查询点积。MSE诊断扫描确定每层替换难度和顺序。替换12个wav 2 vec 2-base层中的8个,在LibriSpeech测试中产生了10.61%的单词错误率(WER),比3.37%的基线高出7.24个百分点(pp),通过优化的MLX推理路径,在Apple M4 Pro上以120秒的音频速度获得3.27倍的加速。SepFormer语音增强的跨域验证显示了所有16个帧内块注意层可以被替换而不会崩溃,这表明深度墙来自语言计算而不是LPA限制。LPA在推理时的近二进制门可以实现密集的GPU计算,而无需CPU-GPU同步,所有操作都映射到移动神经加速器。
摘要:Self-attention scales quadratically with sequence length, limiting transformer-based speech models on edge devices. We introduce the Learnable Pulse Accumulator (LPA), an O(n) replacement that substitutes key-query dot products with learned gating functions: content-dependent rectangular pulses, periodic windows, and position-dependent basis functions. An MSE diagnostic sweep determines per-layer replacement difficulty and ordering. Replacing 8 of 12 wav2vec2-base layers yields 10.61% word error rate (WER) on LibriSpeech test-clean, +7.24 percentage points (pp) over the 3.37% baseline, with 3.27x speedup at 120s audio on Apple M4 Pro via an optimized MLX inference path. Cross-domain validation on SepFormer speech enhancement shows all 16 intra-chunk attention layers can be replaced without collapse, suggesting the depth wall arises from linguistic computation rather than an LPA limitation. LPA's near-binary gates at inference enable dense GPU computation with no CPU-GPU synchronization, and all operations map to mobile neural accelerators.


【12】Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
标题:通过基于LLM的文本和音素呼吸增强对ASB进行合成数据域自适应
链接:https://arxiv.org/abs/2603.16920

作者:Natsuo Yamashita,Koichi Nagatsuka,Hiroaki Kokubo,Kota Dohi,Tuan Vu Ho
备注:accepted by ICASSP 2026
摘要:端到端的自动语音识别通常会由于域内资源的稀缺而在特定于域的数据上降级。我们提出了一个基于合成数据的领域自适应框架,它有两个贡献:(1)一个基于大型语言模型(LLM)的文本增强管道,具有平衡词汇多样性,困惑和领域术语覆盖的过滤策略,以及(2)语音重新拼写增强(PRA),一种通过LLM生成的正字法伪拼写引入发音变化的新方法。与SpecAugment等传统声学级方法不同,PRA在语音合成之前提供语音多样性,使合成语音能够更好地近似真实世界的变化。四个特定领域数据集的实验结果表明,单词错误率的一致降低,证实了将特定领域的词汇覆盖率与现实的发音变化相结合,显着提高了ASR的鲁棒性。
摘要:End-to-end automatic speech recognition often degrades on domain-specific data due to scarce in-domain resources. We propose a synthetic-data-based domain adaptation framework with two contributions: (1) a large language model (LLM)-based text augmentation pipeline with a filtering strategy that balances lexical diversity, perplexity, and domain-term coverage, and (2) phonetic respelling augmentation (PRA), a novel method that introduces pronunciation variability through LLM-generated orthographic pseudo-spellings. Unlike conventional acoustic-level methods such as SpecAugment, PRA provides phonetic diversity before speech synthesis, enabling synthetic speech to better approximate real-world variability. Experimental results across four domain-specific datasets demonstrate consistent reductions in word error rate, confirming that combining domain-specific lexical coverage with realistic pronunciation variation significantly improves ASR robustness.


eess.AS音频处理


【1】The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning
标题:沉默的思想:通过潜在推理在全复式口语对话模型中建模内部认知
链接:https://arxiv.org/abs/2603.17837

作者:Donghang Wu,Tianyu Zhang,Yuxin Li,Hexin Liu,Chen Chen,Eng Siong Chng,Yoshua Bengio
摘要:在对话互动过程中,人们在听演讲者说话时会下意识地进行并发思考。虽然这种内部认知过程可能并不总是表现为明确的语言结构,但它有助于形成高质量的反应。受这种认知现象的启发,我们提出了一种新的全双工内隐推理方法FLAIR,它可以在言语感知的同时进行内隐思维。与NLP中需要事后生成的传统“思考”机制不同,我们的方法与口语对话系统无缝对齐:在用户的说话阶段,它递归地将前一步的潜在嵌入输出馈送到下一步,从而实现严格遵守因果关系的连续推理,而不会引入额外的延迟。为了实现这种潜在的推理,我们设计了一个基于证据下限的目标,通过教师强制支持有效的监督微调,规避了显式推理注释的需要。实验证明了这种边听边想的设计的有效性,在一系列语音基准测试中取得了有竞争力的结果。此外,FLAIR强大地处理会话动态,并在全双工交互指标上获得有竞争力的性能。
摘要:During conversational interactions, humans subconsciously engage in concurrent thinking while listening to a speaker. Although this internal cognitive processing may not always manifest as explicit linguistic structures, it is instrumental in formulating high-quality responses. Inspired by this cognitive phenomenon, we propose a novel Full-duplex LAtent and Internal Reasoning method named FLAIR that conducts latent thinking simultaneously with speech perception. Unlike conventional "thinking" mechanisms in NLP, which require post-hoc generation, our approach aligns seamlessly with spoken dialogue systems: during the user's speaking phase, it recursively feeds the latent embedding output from the previous step into the next step, enabling continuous reasoning that strictly adheres to causality without introducing additional latency. To enable this latent reasoning, we design an Evidence Lower Bound-based objective that supports efficient supervised finetuning via teacher forcing, circumventing the need for explicit reasoning annotations. Experiments demonstrate the effectiveness of this think-while-listening design, which achieves competitive results on a range of speech benchmarks. Furthermore, FLAIR robustly handles conversational dynamics and attains competitive performance on full-duplex interaction metrics.


【2】Multi-Source Evidence Fusion for Audio Question Answering
标题:音频问题回答的多源证据融合
链接:https://arxiv.org/abs/2603.17822

作者:Aivo Olev,Tanel Alumäe
摘要:大型音频语言模型(LALM)可以回答有关语音、音乐和环境声音的问题,但它们的内部推理在很大程度上是不透明的,难以验证。我们描述了TalTech对Interspeech 2026音频推理挑战赛的代理跟踪的解决方案,在该挑战赛中,系统根据推理过程质量进行评估,特别是事实准确性,逻辑可靠性和推理链的完整性。我们的多源集成管道使用两个LALM生成独立的观察结果,而一个单独的纯文本推理模型将这些结果与25个声学工具的输出进行交叉检查,这些工具被组织成可靠性等级。通过将每个推理步骤都建立在明确的、可靠性标记的证据中,该系统产生了密集的、可验证的推理链。我们的系统在挑战中排名第一,在挑战的推理质量度量方面远远优于所有竞争系统。
摘要:Large audio language models (LALMs) can answer questions about speech, music, and environmental sounds, yet their internal reasoning is largely opaque and difficult to validate. We describe TalTech's solution to the Agent Track of the Interspeech 2026 Audio Reasoning Challenge, in which systems are evaluated on reasoning process quality, specifically the factual accuracy, logical soundness, and completeness of their reasoning chains. Our multi-source ensemble pipeline uses two LALMs that generate independent observations, while a separate text-only reasoning model cross-checks these against outputs from 25 acoustic tools organized into reliability tiers. By grounding every inference step in explicit, reliability-tagged evidence, the system produces dense, verifiable reasoning chains. Our system ranked first in the challenge, outperforming all competing systems by a wide margin in challenge's reasoning quality metric.


【3】Robust Nasality Representation Learning for Cleft Palate-Related Velopharyngeal Dysfunction Screening in Real-World Settings
标题:在现实世界环境中用于唇裂相关喉咽功能障碍筛查的稳健鼻腔表示学习
链接:https://arxiv.org/abs/2603.17383

作者:Weixin Liu,Bowen Qu,Amy Stone,Maria E. Powell,Shama Dufresne,Stephane Braun,Izabela Galdyn,Michael Golinko,Bradley Malin,Zhijun Yin,Matthew E. Pontell
备注:2 figures. Machine learning for speech-based VPD screening under domain shift
摘要:腭咽功能障碍(VPD)的特征是在讲话时腭咽闭合不充分,通常会导致鼻音过强和清晰度降低。虽然基于语音的机器学习模型在标准化的临床记录条件下可以表现良好,但由于设备、通道、噪声和室内声学差异导致的域偏移,它们的性能在现实环境中往往会下降。为了提高鲁棒性,我们提出了一个两阶段的框架VPD筛选。首先,一个鼻音为重点的语音表示是通过监督对比预训练的辅助语料库与音素对齐,使用口腔上下文与鼻上下文监督。其次,编码器被冻结,并与0.5秒语音块上的轻量级分类器一起使用,其概率被聚合以产生具有固定阈值的记录级决策。在82名受试者的域内临床队列中,所提出的方法实现了完美的记录水平筛选性能(宏F1 = 1.000,准确度= 1.000)。在131个异构公共互联网录音的单独域外集合上,大型预训练语音表示大幅下降,而MFCC是最强的基线(宏F1 = 0.612,准确度= 0.641)。所提出的方法实现了最佳的域外性能(宏F1 = 0.679,准确度= 0.695),在相同的评估协议下,最强的基线上有所改善。这些结果表明,在临床分类之前学习以鼻音为中心的表示可以降低对记录伪影的敏感性,并提高可部署的基于语音的VPD筛查的鲁棒性。
摘要:Velopharyngeal dysfunction (VPD) is characterized by inadequate velopharyngeal closure during speech and often causes hypernasality and reduced intelligibility. Although speech-based machine learning models can perform well under standardized clinical recording conditions, their performance often drops in real-world settings because of domain shift caused by differences in devices, channels, noise, and room acoustics. To improve robustness, we propose a two-stage framework for VPD screening. First, a nasality-focused speech representation is learned by supervised contrastive pre-training on an auxiliary corpus with phoneme alignments, using oral-context versus nasal-context supervision. Second, the encoder is frozen and used with lightweight classifiers on 0.5-second speech chunks, whose probabilities are aggregated to produce recording-level decisions with a fixed threshold. On an in-domain clinical cohort of 82 subjects, the proposed method achieved perfect recording-level screening performance (macro-F1 = 1.000, accuracy = 1.000). On a separate out-of-domain set of 131 heterogeneous public Internet recordings, large pretrained speech representations degraded substantially, while MFCC was the strongest baseline (macro-F1 = 0.612, accuracy = 0.641). The proposed method achieved the best out-of-domain performance (macro-F1 = 0.679, accuracy = 0.695), improving on the strongest baseline under the same evaluation protocol. These results suggest that learning a nasality-focused representation before clinical classification can reduce sensitivity to recording artifacts and improve robustness for deployable speech-based VPD screening.


【4】Uncertainty Quantification and Risk Control for Multi-Speaker Sound Source Localization
标题:多扬声器声音源定位的不确定性量化与风险控制
链接:https://arxiv.org/abs/2603.17377

作者:Vadim Rozenfeld,Bracha Laufer Goldshtein
备注:13 pages, 4 figures. Code available at: https://github.com/vadimroz/UQ_in_multi_SSL
摘要:可靠的声源定位(SSL)在许多下游任务中起着至关重要的作用,其中明智的决策不仅取决于准确的定位,还取决于对每个估计的信心。这种对可靠性的需求在具有挑战性的条件下变得更加明显,例如混响环境和多源场景。然而,现有的SSL方法通常仅提供点估计,提供有限或不提供不确定性量化(UQ)。我们利用共形预测(CP)框架及其扩展控制一般风险函数开发两个互补的UQ方法SSL。第一个假设活跃源的数量是已知的,并构造覆盖真实源位置的预测区域。第二个解决了更具挑战性的设置,其中源计数是未知的,首先可靠地估计活动源的数量,然后形成相应的预测区域。我们评估所提出的方法在广泛的模拟和现实世界的录音在不同的混响水平和源配置。结果表明,可靠的有限样本保证和一致的性能为已知和未知的源计数的情况下,突出了不确定性感知SSL的实际效用的建议框架。
摘要:Reliable Sound Source Localization (SSL) plays an essential role in many downstream tasks, where informed decision making depends not only on accurate localization but also on the confidence in each estimate. This need for reliability becomes even more pronounced in challenging conditions, such as reverberant environments and multi-source scenarios. However, existing SSL methods typically provide only point estimates, offering limited or no Uncertainty Quantification (UQ). We leverage the Conformal Prediction (CP) framework and its extensions for controlling general risk functions to develop two complementary UQ approaches for SSL. The first assumes that the number of active sources is known and constructs prediction regions that cover the true source locations. The second addresses the more challenging setting where the source count is unknown, first reliably estimating the number of active sources and then forming corresponding prediction regions. We evaluate the proposed methods on extensive simulations and real-world recordings across varying reverberation levels and source configurations. Results demonstrate reliable finite-sample guarantees and consistent performance for both known and unknown source-count scenarios, highlighting the practical utility of the proposed frameworks for uncertainty-aware SSL.


【5】Shared Representation Learning for Reference-Guided Targeted Sound Detection
标题:用于参考引导目标声音检测的共享表示学习
链接:https://arxiv.org/abs/2603.17025

作者:Shubham Gupta,Adarsh Arigala,B. R. Dilleswari,Sri Rama Murty Kodukula
备注:Accepted to IEEE ICASSP 2026
摘要:人类听者表现出通过选择性听觉注意从复杂的声学场景中分离出所需声音的显著能力,激发了目标声音检测(TSD)的研究。该任务需要在提供目标声音的参考音频时检测和定位混合中的该声音。先前的方法依赖于为参考生成声音判别条件嵌入向量,并将其与混合编码器配对,利用多任务学习方法联合优化。在这项工作中,我们提出了一个统一的编码器架构,在一个共享的表示空间内处理参考和混合音频,促进更强的对齐,同时降低架构的复杂性。这种设计选择不仅简化了整个框架,而且还增强了对不可见类的泛化。遵循多任务训练范式,我们的方法比以前的方法实现了实质性的改进,超越了现有的方法,并建立了一个新的最先进的目标声音检测基准,在URBAN-SED数据集上,分段级F1得分为83.15%,总体准确率为95.17%。
摘要:Human listeners exhibit the remarkable ability to segregate a desired sound from complex acoustic scenes through selective auditory attention, motivating the study of Targeted Sound Detection (TSD). The task requires detecting and localizing a target sound in a mixture when a reference audio of that sound is provided. Prior approaches, rely on generating a sound-discriminative conditional embedding vector for the reference and pairing it with a mixture encoder, jointly optimized with a multi-task learning approach. In this work, we propose a unified encoder architecture that processes both the reference and mixture audio within a shared representation space, promoting stronger alignment while reducing architectural complexity. This design choice not only simplifies the overall framework but also enhances generalization to unseen classes. Following the multi-task training paradigm, our method achieves substantial improvements over prior approaches, surpassing existing methods and establishing a new state-of-the-art benchmark for targeted sound detection, with a segment-level F1 score of 83.15% and an overall accuracy of 95.17% on the URBAN-SED dataset.


【6】Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
标题:Wav2Vec语音识别神经网络的空中白盒攻击
链接:https://arxiv.org/abs/2603.16972

作者:Protopopov Alexey
备注:9 pages, 5 figures, 1 table
摘要:基于神经网络的自动语音识别系统容易受到以恶意方式改变传输的对抗性攻击。该领域最近的工作集中在使攻击在空中场景中工作,然而这种攻击通常可以通过人类听觉检测到,限制了它们的潜在应用。在目前的工作中,我们探讨了不同的方法,使空中攻击不易察觉,以及这些方法对攻击的有效性的影响。
摘要:Automatic speech recognition systems based on neural networks are vulnerable to adversarial attacks that alter transcriptions in a malicious way. Recent works in this field have focused on making attacks work in over-the-air scenarios, however such attacks are typically detectable by human hearing, limiting their potential applications. In the present work we explore different approaches of making over-the-air attacks less detectable, as well as the impact these approaches have on the attacks' effectiveness.


【7】The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
标题:言语背后的声音:量化SpeechLLM中的交叉偏见
链接:https://arxiv.org/abs/2603.16941

作者:Shree Harsha Bokkahalli Satish,Christoph Minixhofer,Maria Teleki,James Caverlee,Ondřej Klejch,Peter Bell,Gustav Eje Henter,Éva Székely
备注:5 pages, 3 figures, 1 table, Submitted to Interspeech 2026
摘要:语音大语言模型(SpeechLLM)直接处理口语输入,保留先前在级联管道中删除的口音和感知性别等线索。这在响应中引入了说话者身份相关的变化。我们在三个SpeechLLM中对口音和性别偏见进行了大规模的交叉评估,使用了六种英语口音和两种性别呈现的2,880种受控交互,通过语音克隆保持语言内容不变。使用逐点LLM判断评级,成对比较和最佳-最差缩放与人类验证,我们检测一致的差异。东欧口音的讲话得到较低的帮助分数,特别是女性提出的声音。这种偏见是隐性的:反应仍然是礼貌的,但在帮助方面有所不同。虽然LLM法官捕捉这些偏见的方向趋势,人类评估者表现出显着更高的敏感性,发现更尖锐的交叉差异。
摘要:Speech Large Language Models (SpeechLLMs) process spoken input directly, retaining cues such as accent and perceived gender that were previously removed in cascaded pipelines. This introduces speaker identity dependent variation in responses. We present a large-scale intersectional evaluation of accent and gender bias in three SpeechLLMs using 2,880 controlled interactions across six English accents and two gender presentations, keeping linguistic content constant through voice cloning. Using pointwise LLM-judge ratings, pairwise comparisons, and Best-Worst Scaling with human validation, we detect consistent disparities. Eastern European-accented speech receives lower helpfulness scores, particularly for female-presenting voices. The bias is implicit: responses remain polite but differ in helpfulness. While LLM judges capture the directional trend of these biases, human evaluators exhibit significantly higher sensitivity, uncovering sharper intersectional disparities.


【8】SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
标题:SimulU:长格式语音同步翻译免训练政策
链接:https://arxiv.org/abs/2603.16924

作者:Amirbek Djanibekov,Luisa Bentivogli,Matteo Negri,Sara Papi
摘要:同时语音到语音翻译(SimulS 2S)对于实时多语言通信至关重要,并且越来越多地集成到会议和流媒体平台中。尽管如此,SimulS 2S在研究中仍然没有得到充分的探索,目前的解决方案往往依赖于资源密集型的训练程序,并对短形式的预分段话语进行操作,无法推广到连续语音。为了弥合这一差距,我们提出了SimulU,这是长格式SimulS 2S的第一个免训练策略。SimulU采用历史管理和语音输出选择策略,利用预先训练的端到端模型中的交叉注意来调节输入历史和输出生成。跨8种语言的MuST-C评估表明,SimulU实现了与强大级联模型相比更好或相当的质量-延迟权衡。通过消除对ad-hoc培训的需求,SimulU为现实的长形式场景中的端到端SimulS 2S提供了一条有前途的道路。
摘要:Simultaneous speech-to-speech translation (SimulS2S) is essential for real-time multilingual communication, with increasing integration into meeting and streaming platforms. Despite this, SimulS2S remains underexplored in research, where current solutions often rely on resource-intensive training procedures and operate on short-form, pre-segmented utterances, failing to generalize to continuous speech. To bridge this gap, we propose SimulU, the first training-free policy for long-form SimulS2S. SimulU adopts history management and speech output selection strategies that exploit cross-attention in pre-trained end-to-end models to regulate both input history and output generation. Evaluations on MuST-C across 8 languages show that SimulU achieves a better or comparable quality-latency trade-off against strong cascaded models. By eliminating the need for ad-hoc training, SimulU offers a promising path to end-to-end SimulS2S in realistic, long-form scenarios.


【9】Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
标题:超越深度学习:使用神经组合的语音分割和电话分类
链接:https://arxiv.org/abs/2603.16923

作者:Trevor Adelson,Vidhyasaharan Sethu,Ting Dang
备注:Submitted to Interspeech 2026. 9 Pages
摘要:深度学习主导语音处理,但依赖于大量数据集、全局反向传播引导的权重更新,并产生纠缠表示。装配演算(AC),通过赫布可塑性和赢家通吃的竞争模型稀疏神经元组件,提供了一个生物接地替代,但以前的工作集中在离散的符号输入。我们引入了一个基于AC的语音处理框架,该框架通过结合三个关键贡献直接对连续语音进行操作:(i)神经编码,使用概率mel二进制化和人口编码的MFCC将语音转换为装配兼容的尖峰模式;(ii)多区域架构,跨层次时间尺度和类组织装配;(iii)跨区域更新方案,用于下游任务。应用于两个核心任务的边界检测和段分类,我们的框架检测电话(F1=0.69)和单词(F1=0.61)的边界没有任何权重训练,并达到47.5%和45.1%的准确率电话和命令识别。这些结果表明,基于AC的动态系统是深度学习语音处理的可行替代方案。
摘要:Deep learning dominates speech processing but relies on massive datasets, global backpropagation-guided weight updates, and produces entangled representations. Assembly Calculus (AC), which models sparse neuronal assemblies via Hebbian plasticity and winner-take-all competition, offers a biologically grounded alternative, yet prior work focused on discrete symbolic inputs. We introduce an AC-based speech processing framework that operates directly on continuous speech by combining three key contributions:(i) neural encoding that converts speech into assembly-compatible spike patterns using probabilistic mel binarisation and population-coded MFCCs; (ii) a multi-area architecture organising assemblies across hierarchical timescales and classes; and (iii) cross-area update schemes for downstream tasks. Applied to two core tasks of boundary detection and segment classification, our framework detects phone (F1=0.69) and word (F1=0.61) boundaries without any weight training, and achieves 47.5% and 45.1% accuracy on phone and command recognition. These results show that AC-based dynamical systems are a viable alternative to deep learning for speech processing.


【10】Learnable Pulse Accumulation for On-Device Speech Recognition: How Much Attention Do You Need?
标题:设备上语音识别的可学习脉搏累积:您需要多少关注?
链接:https://arxiv.org/abs/2603.16922

作者:Yakov Pyotr Shkolnikov
摘要:自我注意力与序列长度成二次关系,限制了边缘设备上基于transformer的语音模型。我们引入了可学习脉冲累加器(LPA),这是一个O(n)的替换,它用学习的门函数(内容相关的矩形脉冲,周期性窗口和位置相关的基函数)代替了关键字查询点积。MSE诊断扫描确定每层替换难度和顺序。替换12个wav 2 vec 2-base层中的8个,在LibriSpeech测试中产生了10.61%的单词错误率(WER),比3.37%的基线高出7.24个百分点(pp),通过优化的MLX推理路径,在Apple M4 Pro上以120秒的音频速度获得3.27倍的加速。SepFormer语音增强的跨域验证显示了所有16个帧内块注意层可以被替换而不会崩溃,这表明深度墙来自语言计算而不是LPA限制。LPA在推理时的近二进制门可以实现密集的GPU计算,而无需CPU-GPU同步,所有操作都映射到移动神经加速器。
摘要:Self-attention scales quadratically with sequence length, limiting transformer-based speech models on edge devices. We introduce the Learnable Pulse Accumulator (LPA), an O(n) replacement that substitutes key-query dot products with learned gating functions: content-dependent rectangular pulses, periodic windows, and position-dependent basis functions. An MSE diagnostic sweep determines per-layer replacement difficulty and ordering. Replacing 8 of 12 wav2vec2-base layers yields 10.61% word error rate (WER) on LibriSpeech test-clean, +7.24 percentage points (pp) over the 3.37% baseline, with 3.27x speedup at 120s audio on Apple M4 Pro via an optimized MLX inference path. Cross-domain validation on SepFormer speech enhancement shows all 16 intra-chunk attention layers can be replaced without collapse, suggesting the depth wall arises from linguistic computation rather than an LPA limitation. LPA's near-binary gates at inference enable dense GPU computation with no CPU-GPU synchronization, and all operations map to mobile neural accelerators.


【11】Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
标题:通过基于LLM的文本和音素呼吸增强对ASB进行合成数据域自适应
链接:https://arxiv.org/abs/2603.16920

作者:Natsuo Yamashita,Koichi Nagatsuka,Hiroaki Kokubo,Kota Dohi,Tuan Vu Ho
备注:accepted by ICASSP 2026
摘要:端到端的自动语音识别通常会由于域内资源的稀缺而在特定于域的数据上降级。我们提出了一个基于合成数据的领域自适应框架,它有两个贡献:(1)一个基于大型语言模型(LLM)的文本增强管道,具有平衡词汇多样性,困惑和领域术语覆盖的过滤策略,以及(2)语音重新拼写增强(PRA),一种通过LLM生成的正字法伪拼写引入发音变化的新方法。与SpecAugment等传统声学级方法不同,PRA在语音合成之前提供语音多样性,使合成语音能够更好地近似真实世界的变化。四个特定领域数据集的实验结果表明,单词错误率的一致降低,证实了将特定领域的词汇覆盖率与现实的发音变化相结合,显着提高了ASR的鲁棒性。
摘要:End-to-end automatic speech recognition often degrades on domain-specific data due to scarce in-domain resources. We propose a synthetic-data-based domain adaptation framework with two contributions: (1) a large language model (LLM)-based text augmentation pipeline with a filtering strategy that balances lexical diversity, perplexity, and domain-term coverage, and (2) phonetic respelling augmentation (PRA), a novel method that introduces pronunciation variability through LLM-generated orthographic pseudo-spellings. Unlike conventional acoustic-level methods such as SpecAugment, PRA provides phonetic diversity before speech synthesis, enabling synthetic speech to better approximate real-world variability. Experimental results across four domain-specific datasets demonstrate consistent reductions in word error rate, confirming that combining domain-specific lexical coverage with realistic pronunciation variation significantly improves ASR robustness.


【12】Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
标题:语音生成大型音频语言模型中的神经元级情绪控制
链接:https://arxiv.org/abs/2603.17231

作者:Xiutian Zhao,Ismail Rasim Ulgen,Philipp Koehn,Björn Schuller,Berrak Sisman
备注:11 pages, 10 figures
摘要:大型音频语言模型(LALM)可以产生富有表达力的语音,但可靠的情感控制仍然难以捉摸:转换经常错过目标影响,并可能通过拒绝,幻觉或释义降低语言保真度。据我们所知,我们提出了第一个神经元水平的研究语音生成LALM的情绪控制,并证明了紧凑的情绪敏感神经元(ESN)是因果可操作的,使训练自由情绪转向推理时间。ESN通过成功过滤激活聚合来识别,从而实现情感实现和内容保存。在三种LAM(Qwen 2.5-Omni-7 B、MiniCPM-o 4.5、Kimi-Audio)中,ESN干预产生特定于情绪的增益,这些增益可以推广到看不见的说话者,并得到自动和人工评估的支持。可控性取决于选择器设计、掩码稀疏性、过滤和干预强度。我们的研究结果建立了一个机械的框架,在语音生成的训练自由情绪控制。
摘要:Large audio-language models (LALMs) can produce expressive speech, yet reliable emotion control remains elusive: conversions often miss the target affect and may degrade linguistic fidelity through refusals, hallucinations, or paraphrase. We present, to our knowledge, the first neuron-level study of emotion control in speech-generative LALMs and demonstrate that compact emotion-sensitive neurons (ESNs) are causally actionable, enabling training-free emotion steering at inference time. ESNs are identified via success-filtered activation aggregation enforcing both emotion realization and content preservation. Across three LALMs (Qwen2.5-Omni-7B, MiniCPM-o 4.5, Kimi-Audio), ESN interventions yield emotion-specific gains that generalize to unseen speakers and are supported by automatic and human evaluation. Controllability depends on selector design, mask sparsity, filtering, and intervention strength. Our results establish a mechanistic framework for training-free emotion control in speech generation.


【13】Collecting Prosody in the Wild: A Content-Controlled, Privacy-First Smartphone Protocol and Empirical Evaluation
标题:野外收集韵律:内容控制、隐私优先的智能手机协议和经验评估
链接:https://arxiv.org/abs/2603.17061

作者:Timo K. Koch,Florian Bemmann,Ramona Schoedel,Markus Buehner,Clemens Stachl
备注:Submitted to Interspeech 2026
摘要:由于韵律和语义、隐私约束和参与者依从性的混杂,收集日常语音数据用于韵律分析是具有挑战性的。我们介绍和经验评估的内容控制,隐私第一的智能手机协议,使用脚本朗读句子标准化的词汇内容(包括提示价),同时捕捉自然变化的韵律交付。该协议执行设备上的韵律特征提取,立即删除原始音频,并仅传输导出的特征进行分析。我们在一项大型研究中部署了该协议(N = 560; 9,877个记录),评估了依从性和数据质量,并对提取的特征进行了诊断预测任务,预测说话者性别并同时报告了瞬时情感状态(效价,唤醒)。我们讨论的影响和方向推进和部署的协议。
摘要:Collecting everyday speech data for prosodic analysis is challenging due to the confounding of prosody and semantics, privacy constraints, and participant compliance. We introduce and empirically evaluate a content-controlled, privacy-first smartphone protocol that uses scripted read-aloud sentences to standardize lexical content (including prompt valence) while capturing natural variation in prosodic delivery. The protocol performs on-device prosodic feature extraction, deletes raw audio immediately, and transmits only derived features for analysis. We deployed the protocol in a large study (N = 560; 9,877 recordings), evaluated compliance and data quality, and conducted diagnostic prediction tasks on the extracted features, predicting speaker sex and concurrently reported momentary affective states (valence, arousal). We discuss implications and directions for advancing and deploying the protocol.


【14】CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization
标题:CineSRD:利用视觉、声学和语言线索进行开放世界视觉媒体扬声器对话
链接:https://arxiv.org/abs/2603.16966

作者:Liangbin Huang,Xiaohua Liao,Chaoqun Cui,Shijing Wang,Zhaolong Huang,Yanlong Du,Wenji Mao
备注:Accepted to CVPR 2026
摘要:传统的发言者日记系统主要集中在诸如会议和采访之类的受约束的场景,其中发言者的数量有限并且声学条件相对干净。为了探索开放世界的扬声器日记,我们将这项任务扩展到视觉媒体领域,包括复杂的视听节目,如电影和电视剧。这种新的设置带来了几个挑战,包括长格式视频理解,大量的扬声器,音频和视觉线索之间的跨模态干扰,以及不受控制的野外变化。为了解决这些挑战,我们提出了电影扬声器注册和日记(CineSRD),一个统一的多模式框架,利用视觉,声学和语言线索,从视频,语音和字幕的扬声器注释。CineSRD首先执行视觉锚点聚类以注册初始扬声器,然后集成音频语言模型以进行扬声器转向检测,改进注释并补充未注册的屏幕外扬声器。此外,我们构建并发布了一个专门针对视觉媒体的演讲者日志基准,包括中文和英文节目。实验结果表明,CineSRD在提出的基准测试中取得了卓越的性能,在传统数据集上取得了有竞争力的结果,验证了其在开放世界视觉媒体环境中的鲁棒性和可推广性。
摘要:Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore open-world speaker diarization, we extend this task to the visual media domain, encompassing complex audiovisual programs such as films and TV series. This new setting introduces several challenges, including long-form video understanding, a large number of speakers, cross-modal asynchrony between audio and visual cues, and uncontrolled in-the-wild variability. To address these challenges, we propose Cinematic Speaker Registration & Diarization (CineSRD), a unified multimodal framework that leverages visual, acoustic, and linguistic cues from video, speech, and subtitles for speaker annotation. CineSRD first performs visual anchor clustering to register initial speakers and then integrates an audio language model for speaker turn detection, refining annotations and supplementing unregistered off-screen speakers. Furthermore, we construct and release a dedicated speaker diarization benchmark for visual media that includes Chinese and English programs. Experimental results demonstrate that CineSRD achieves superior performance on the proposed benchmark and competitive results on conventional datasets, validating its robustness and generalizability in open-world visual media settings.


【15】Music Source Restoration with Ensemble Separation and Targeted Reconstruction
标题:整体分离和有针对性重建的音乐源泉恢复
链接:https://arxiv.org/abs/2603.16926

作者:Xinlong Deng,Yu Xia,Jie Jiang
摘要:首届音乐源恢复(MSR)挑战赛的目标是从完全混合和掌握的音乐中恢复原始的、未经处理的音乐。与传统的音乐源分离不同,MSR需要反转复杂的制作过程,例如均衡,压缩,混响和其他真实世界的降级。为了解决MSR,我们提出了一个两阶段系统。首先,预先训练的分离模型的集合产生初步的源估计。然后,一组预先训练的基于BSRNN的恢复模型执行有针对性的重建,以细化这些估计。在官方的MSR基准测试中,我们的系统在所有指标上都超过了基线,在所有提交中排名第二。该代码可在https://github.com/xinghour/Music-source-restoration-CUPAudioGroup上获得
摘要:The Inaugural Music Source Restoration (MSR) Challenge targets the recovery of original, unprocessed stems from fully mixed and mastered music. Unlike conventional music source separation, MSR requires reversing complex production processes such as equalization, compression, reverberation, and other real-world degradations. To address MSR, we propose a two-stage system. First, an ensemble of pre-trained separation models produces preliminary source estimates. Then a set of pre-trained BSRNN-based restoration models performs targeted reconstruction to refine these estimates. On the official MSR benchmark, our system surpasses the baselines on all metrics, ranking second among all submissions. The code is available at https://github.com/xinghour/Music-source-restoration-CUPAudioGroup


【16】Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
标题:用于语音深度伪造检测的量化器感知分层神经编解码器建模
链接:https://arxiv.org/abs/2603.16914

作者:Jinyang Wu,Zihan Pan,Qiquan Zhang,Sailor Hardik Bhupendra,Soumik Mondal
备注:5 pages, 3 figures
摘要:神经音频编解码器通过残差矢量量化(RVQ)离散化语音,在量化器之间形成从粗到细的层次结构。虽然编解码器模型已经被探索用于表示学习,但它们的离散结构在语音深度伪造检测中仍然没有得到充分利用。特别是,不同的量化级别捕获互补的声学线索,其中早期的量化器编码粗糙的结构,稍后的量化器细化残留的细节,揭示合成伪影。现有的系统要么依赖于连续编码器功能,要么忽略这个量化器级别的层次结构。我们提出了一个层次感知的表示学习框架,通过可学习的全局加权模型量化器级的贡献,使结构化的编解码器表示与法医线索对齐。保持语音编码器骨干冻结并仅更新4.4%的附加参数,我们的方法在ASVspoof 2019上实现了46.2%的相对EER降低,在ASVspoof 5上实现了13.9%的相对EER降低。
摘要:Neural audio codecs discretize speech via residual vector quantization (RVQ), forming a coarse-to-fine hierarchy across quantizers. While codec models have been explored for representation learning, their discrete structure remains underutilized in speech deepfake detection. In particular, different quantization levels capture complementary acoustic cues, where early quantizers encode coarse structure and later quantizers refine residual details that reveal synthesis artifacts. Existing systems either rely on continuous encoder features or ignore this quantizer-level hierarchy. We propose a hierarchy-aware representation learning framework that models quantizer-level contributions through learnable global weighting, enabling structured codec representations aligned with forensic cues. Keeping the speech encoder backbone frozen and updating only 4.4% additional parameters, our method achieves relative EER reductions of 46.2% on ASVspoof 2019 and 13.9% on ASVspoof5 over strong baselines.


【17】Amanous: Distribution-Switching for Superhuman Piano Density on Disklavier
标题:Amanous:《尤利西斯》上超人钢琴密度的分布转换
链接:https://arxiv.org/abs/2603.16890

作者:Joonhyung Bae
摘要:自动钢琴使音符密度,复调和寄存器的变化远远超出了人类的物理限制,但三个占主导地位的传统组成这样的纹理-南卡罗的速度佳能,Xenakis的随机分布,L系统语法-已经孤立地发展。本文介绍了Amanous,一个硬件感知的组成系统,雅马哈pullavier,统一这些方法,通过分布切换:L-系统符号选择不同的分布制度,而不仅仅是调制参数在一个固定的家庭。报告了四项贡献。(1)一个四层架构(符号,参数,数字,物理)产生统计上不同的部分与大的效果大小(d = 3.70-5.34),每层的退化和消融实验验证。(2)一个硬件抽象层形式化速度相关的延迟和密钥重置约束,将超人的纹理保持在机器人的可致信封内。(3)密度扫描揭示了在24-30个音符/秒处的计算饱和过渡(自举95%CI:23.3-50.0),超过该饱和过渡,单域旋律度量失去辨别能力,并且跨域耦合变得必要。(4)一个收敛点演算操作tempo-canon几何作为一个控制接口,使收敛事件触发分布开关连接宏观时间结构的微观层次的纹理。所有结果都是计算的,心理声学验证协议提出了未来的工作。该流水线已部署在一个物理平台上,展示了算法的自一致性和亚毫秒级的软件精度。补充材料(摘录1-4):https://www.amanous.xyz。源代码:https://github.com/joonhyungbae/Amanous。
摘要:The automated piano enables note densities, polyphony, and register changes far beyond human physical limits, yet the three dominant traditions for composing such textures--Nancarrow's tempo canons, Xenakis's stochastic distributions, and L-system grammars--have developed in isolation. This paper presents Amanous, a hardware-aware composition system for Yamaha Disklavier that unifies these methodologies through distribution-switching: L-system symbols select distinct distributional regimes rather than merely modulating parameters within a fixed family. Four contributions are reported. (1) A four-layer architecture (symbolic, parametric, numeric, physical) produces statistically distinct sections with large effect sizes (d = 3.70-5.34), validated by per-layer degradation and ablation experiments. (2) A hardware abstraction layer formalizes velocity-dependent latency and key reset constraints, keeping superhuman textures within the Disklavier's actuable envelope. (3) A density sweep reveals a computational saturation transition at 24-30 notes/s (bootstrap 95% CI: 23.3-50.0), beyond which single-domain melodic metrics lose discriminative power and cross-domain coupling becomes necessary. (4) A convergence point calculus operationalizes tempo-canon geometry as a control interface, enabling convergence events to trigger distribution switches linking macro-temporal structure to micro-level texture. All results are computational; a psychoacoustic validation protocol is proposed for future work. The pipeline has been deployed on a physical Disklavier, demonstrating algorithmic self-consistency and sub-millisecond software precision. Supplementary materials (Excerpts 1-4): https://www.amanous.xyz. Source code: https://github.com/joonhyungbae/Amanous.


【18】Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
标题:针对多方面、多评分者L2阅读言语评估的SpeechLLM的文字引导微调
链接:https://arxiv.org/abs/2603.16889

作者:Aditya Kamlesh Parikh,Cristian Tejedor-Garcia,Catia Cucchiarini,Helmer Strik
备注:Accepted to LREC 2026. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants, which is financed by the Dutch Research Council (NWO)
摘要:第二语言(L2)语音的可靠和可解释的自动评估仍然是一个核心挑战,因为大型语音语言模型(SpeechLLM)通常难以与人类评分员的细微变化保持一致。为了解决这个问题,我们引入了一个规则引导的推理框架,明确编码多方面的人类评估标准:准确性,流畅性和韵律,同时校准模型的不确定性,以捕捉自然的评级变化。我们微调Qwen 2-Audio-7 B-Instruct模型使用多评分人的判断,并开发了一个不确定性校准的回归方法,由保形校准支持可解释的置信区间。我们的高斯不确定性建模和共形校准方法实现了与人类评级最强的一致性,优于回归和分类基线。该模型可靠地评估流畅性和韵律,同时突出了评估准确性的固有困难。总之,这些结果表明,标题引导,不确定性校准推理提供了一个原则性的路径,值得信赖的和可解释的SpeechLLM为基础的语音评估。
摘要:Reliable and interpretable automated assessment of second-language (L2) speech remains a central challenge, as large speech-language models (SpeechLLMs) often struggle to align with the nuanced variability of human raters. To address this, we introduce a rubric-guided reasoning framework that explicitly encodes multi-aspect human assessment criteria: accuracy, fluency, and prosody, while calibrating model uncertainty to capture natural rating variability. We fine-tune the Qwen2-Audio-7B-Instruct model using multi-rater human judgments and develop an uncertainty-calibrated regression approach supported by conformal calibration for interpretable confidence intervals. Our Gaussian uncertainty modeling and conformal calibration approach achieves the strongest alignment with human ratings, outperforming regression and classification baselines. The model reliably assesses fluency and prosody while highlighting the inherent difficulty of assessing accuracy. Together, these results demonstrate that rubric-guided, uncertainty-calibrated reasoning offers a principled path toward trustworthy and explainable SpeechLLM-based speech assessment.


机器翻译由腾讯交互翻译提供,仅供参考