今日论文合集:cs.SD语音14篇,eess.AS音频处理17篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases
标题:SCENEBBench:基于辅助和工业用例的音频理解基准
链接:https://arxiv.org/abs/2603.09853

作者:Laya Iyer,Angelina Wang,Sanmi Koyejo
备注:Accepted to EACL 2026 (Main Conference). 10 pages, 10 figures. Camera-ready version
摘要:大型语言模型(LLM)的进步已经实现了音频处理的重要功能,从而产生了现在被称为大型音频语言模型(LALM)的最先进的模型。然而,除了自动语音识别(ASR)之外,还做了很少的工作来测量音频理解。本文通过提出一个基准套件SCENEBench(空间,跨语言,环境,非语音评估)来缩小这一差距,该套件针对四个现实世界类别的广泛形式的音频理解:背景声音理解,噪声定位,跨语言语音理解和语音特征识别。这四个类别是根据对无障碍技术和工业噪声监测的未充分研究的需求而选择的。除了性能,我们还测量模型延迟。这个基准测试套件的目的是评估音频,而不仅仅是说了什么话-更确切地说,是如何说的,以及音频的非语音成分。因为我们的音频样本是合成构建的(例如,通过覆盖两个自然音频样本),我们进一步验证了我们的基准,每个任务20个自然音频项,从现有数据集进行二次采样以匹配我们的任务标准,以评估生态有效性。我们评估了五种最先进的LALM,并发现了关键的差距:不同任务的性能各不相同,有些任务的性能低于随机机会,而另一些任务则达到了很高的准确性。这些结果为模型能力的有针对性的改进提供了方向。
摘要:Advances in large language models (LLMs) have enabled significant capabilities in audio processing, resulting in state-of-the-art models now known as Large Audio Language Models (LALMs). However, minimal work has been done to measure audio understanding beyond automatic speech recognition (ASR). This paper closes that gap by proposing a benchmark suite, SCENEBench (Spatial, Cross-lingual, Environmental, Non-speech Evaluation), that targets a broad form of audio comprehension across four real-world categories: background sound understanding, noise localization, cross-linguistic speech understanding, and vocal characterizer recognition. These four categories are selected based on understudied needs from accessibility technology and industrial noise monitoring. In addition to performance, we also measure model latency. The purpose of this benchmark suite is to assess audio beyond just what words are said - rather, how they are said and the non-speech components of the audio. Because our audio samples are synthetically constructed (e.g., by overlaying two natural audio samples), we further validate our benchmark against 20 natural audio items per task, sub-sampled from existing datasets to match our task criteria, to assess ecological validity. We assess five state-of-the-art LALMs and find critical gaps: performance varies across tasks, with some tasks performing below random chance and others achieving high accuracy. These results provide direction for targeted improvements in model capabilities.


【2】EmoSURA: Towards Accurate Evaluation of Detailed and Long-Context Emotional Speech Captions
标题:CLARSURA:准确评估详细和长上下文情感言语说明
链接:https://arxiv.org/abs/2603.09820

作者:Xin Jing,Andreas Triantafyllopoulos,Jiadong Wang,Shahin Amiriparian,Jun Luo,Björn Schuller
备注:Submitted to Interspeech 2026
摘要:语音字幕模型的最新进展使得能够为情感语音生成丰富的细粒度字幕。然而,此类字幕的评估仍然是一个关键瓶颈:传统的N元语法指标无法捕捉语义细微差别,而LLM法官在处理长篇描述时经常遇到推理不一致和上下文崩溃的问题。在这项工作中,我们提出了一个新的评估框架,将范式从整体评分原子验证的RISSURA。CIMSURA将复杂的字幕分解为原子感知单元,这些单元是关于声音或情感属性的自包含语句,并采用基于音频的验证机制来验证每个单元与原始语音信号的对比。此外,我们通过引入SURABench(一个经过仔细平衡和分层的基准)来解决标准化评估资源的稀缺问题。我们的实验表明,MANUFACTURSURA实现了与人类判断的正相关性,提供了一个更可靠的评估长形式的字幕相比,传统的指标,表现出负相关性,由于其敏感性的字幕长度。
摘要:Recent advancements in speech captioning models have enabled the generation of rich, fine-grained captions for emotional speech. However, the evaluation of such captions remains a critical bottleneck: traditional N-gram metrics fail to capture semantic nuances, while LLM judges often suffer from reasoning inconsistency and context-collapse when processing long-form descriptions. In this work, we propose EmoSURA, a novel evaluation framework that shifts the paradigm from holistic scoring to atomic verification. EmoSURA decomposes complex captions into Atomic Perceptual Units, which are self-contained statements regarding vocal or emotional attributes, and employs an audio-grounded verification mechanism to validate each unit against the raw speech signal. Furthermore, we address the scarcity of standardized evaluation resources by introducing SURABench, a carefully balanced and stratified benchmark. Our experiments show that EmoSURA achieves a positive correlation with human judgments, offering a more reliable assessment for long-form captions compared to traditional metrics, which demonstrated negative correlations due to their sensitivity to caption length.


【3】MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
标题:MUgen:评估和改善大型音频语言模型的多音频理解
链接:https://arxiv.org/abs/2603.09714

作者:Chih-Kai Yang,Yun-Shao Tsai,Yu-Kai Guo,Ping-Le Tsai,Yen-Ting Piao,Hung-Wei Chen,Ting-Lin Hsiao,Yun-Man Hsu,Ke-Han Lu,Hung-yi Lee
备注:6 pages, 3 figures, 3 tables. Dataset: https://huggingface.co/Multi-Audio-Grounding
摘要:虽然多音频理解对于大型音频语言模型(LALM)至关重要,但它仍然没有得到充分的探索。我们介绍MUGEN,这是一个全面的基准测试,用于评估语音、通用音频和音乐的这种能力。我们的实验揭示了多音频设置中的一致弱点,并且随着并发音频输入数量的增加,性能急剧下降,将输入缩放确定为根本瓶颈。我们进一步研究了免训练策略,并观察到音频置换自一致性(Audio-Permutational Self-Consistency)使音频候选项的顺序多样化,有助于模型形成更强大的聚合预测,准确率提高了6.28%。将这种排列策略与Chain-of-Thought相结合,进一步将性能提高到6.74%。这些结果暴露了当前LALM中的盲点,并为评估复杂的听觉理解提供了基础。
摘要:While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and music. Our experiments reveal consistent weaknesses in multi-audio settings, and performance degrades sharply as the number of concurrent audio inputs increases, identifying input scaling as a fundamental bottleneck. We further investigate training-free strategies and observe that Audio-Permutational Self-Consistency, which diversifies the order of audio candidates, helps models form more robust aggregated predictions, yielding up to 6.28% accuracy gains. Combining this permutation strategy with Chain-of-Thought further improves performance to 6.74%. These results expose blind spots in current LALMs and provide a foundation for evaluating complex auditory comprehension.


【4】Physics-Informed Neural Engine Sound Modeling with Differentiable Pulse-Train Synthesis
标题:具有可区分脉冲串合成的物理信息神经引擎声音建模
链接:https://arxiv.org/abs/2603.09391

作者:Robin Doerfler,Lonce Wyse
备注:Preprint. 5 pages, 2 figures. Audio examples, code, and model weights available online
摘要:发动机声音来源于连续的排气压力脉冲,而不是持续的谐波振荡。虽然神经合成方法通常旨在近似得到的频谱特性,但我们建议直接对潜在的脉冲形状和时间结构进行建模。我们提出了脉冲串谐振器(PTR)模型,一个可微分的合成架构,生成发动机音频参数化脉冲串对齐发动机点火模式,并通过递归的Karplus-Strong谐振器模拟排气声学传播。该架构集成了物理信息的感应偏置,包括谐波衰减,热力学桨距调制,阀动态包络,排气系统共振和衍生的发动机操作模式,如油门操作和减速燃料切断(DCFO)。   在三种不同的引擎类型上进行了总计7.5小时的音频验证,PTR在谐波重建方面实现了21%的改进,并在谐波加噪声基线模型上减少了5.7%的总损失,同时提供了与物理现象相对应的可解释参数。   完整的代码、模型权重和音频示例都是公开的。
摘要:Engine sounds originate from sequential exhaust pressure pulses rather than sustained harmonic oscillations. While neural synthesis methods typically aim to approximate the resulting spectral characteristics, we propose directly modeling the underlying pulse shapes and temporal structure. We present the Pulse-Train-Resonator (PTR) model, a differentiable synthesis architecture that generates engine audio as parameterized pulse trains aligned to engine firing patterns and propagates them through recursive Karplus-Strong resonators simulating exhaust acoustics. The architecture integrates physics-informed inductive biases including harmonic decay, thermodynamic pitch modulation, valve-dynamics envelopes, exhaust system resonances and derived engine operating modes such as throttle operation and deceleration fuel cutoff (DCFO).   Validated on three diverse engine types totaling 7.5 hours of audio, PTR achieves a 21% improvement in harmonic reconstruction and a 5.7% reduction in total loss over a harmonic-plus-noise baseline model, while providing interpretable parameters corresponding to physical phenomena.   Complete code, model weights, and audio examples are openly available.


【5】TimberAgent: Gram-Guided Retrieval for Executable Music Effect Control
标题:TimberAgent:可执行音乐效果控制的文法引导检索
链接:https://arxiv.org/abs/2603.09332

作者:Shihao He,Yihan Xia,Fang Liu,Taotao Wang,Shengli Zhang
摘要:数字音频工作站暴露出丰富的效果链,但在感知用户意图和低级别信号处理参数之间仍然存在语义差距。我们研究检索接地音频效果控制,输出是一个可编辑的插件配置,而不是一个最终的波形。我们的重点是纹理共振检索(TRR),一个音频表示建立从革兰氏矩阵的预计中级Wav 2 Vec 2激活。这种设计保留了纹理相关的共激活结构。我们在吉他效果基准上评估了TRR,其中包含1,063个候选搜索和204个查询。评估遵循方案A,这是一种防止训练测试泄漏的交叉验证方案。我们比较TRR对CLAP和内部检索基线(Wav 2 Vec-RAG,Text-RAG,WavureNN-RAG),使用最小-最大归一化指标接地的物理DSP参数范围。消融研究验证TRR的核心设计选择:投影维度,层选择和投影类型。一个近乎重复的敏感性分析证实,结果是强大的平凡的知识库匹配。TRR实现了最低的归一化参数误差之间的评估方法。一项有26名参与者的多刺激听力研究提供了补充的感知证据。我们将这些结果解释为基准证据,即纹理感知检索对于可编辑音频效果控制是有用的,而更广泛的个性化和真实音频鲁棒性声明仍然在这里提供的验证证据之外。
摘要:Digital audio workstations expose rich effect chains, yet a semantic gap remains between perceptual user intent and low-level signal-processing parameters. We study retrieval-grounded audio effect control, where the output is an editable plugin configuration rather than a finalized waveform. Our focus is Texture Resonance Retrieval (TRR), an audio representation built from Gram matrices of projected mid-level Wav2Vec2 activations. This design preserves texture-relevant co-activation structure. We evaluate TRR on a guitar-effects benchmark with 1,063 candidate presets and 204 queries. The evaluation follows Protocol-A, a cross-validation scheme that prevents train-test leakage. We compare TRR against CLAP and internal retrieval baselines (Wav2Vec-RAG, Text-RAG, FeatureNN-RAG), using min-max normalized metrics grounded in physical DSP parameter ranges. Ablation studies validate TRR's core design choices: projection dimensionality, layer selection, and projection type. A near-duplicate sensitivity analysis confirms that results are robust to trivial knowledge-base matches. TRR achieves the lowest normalized parameter error among evaluated methods. A multiple-stimulus listening study with 26 participants provides complementary perceptual evidence. We interpret these results as benchmark evidence that texture-aware retrieval is useful for editable audio effect control, while broader personalization and real-audio robustness claims remain outside the verified evidence presented here.


【6】Paralinguistic Emotion-Aware Validation Timing Detection in Japanese Empathetic Spoken Dialogue
标题:日语同理心口语对话中副语言感知确认时间检测
链接:https://arxiv.org/abs/2603.09307

作者:Zi Haur Pang,Yahui Fu,Yuan Gao,Tatsuya Kawahara
备注:Accepted to ICASSP 2026
摘要:情绪验证是一种心理治疗沟通技术,涉及识别,理解和明确承认另一个人的感受和行为,这加强了联盟,减少了负面影响。为了最大限度地发挥验证所提供的情感支持,以适当的时间和频率提供它是至关重要的。本研究从语音的角度探讨验证时序检测。利用语言和情感信息,我们提出了一个语言和情感感知模型的验证时间检测,而不依赖于文本上下文。具体来说,我们首先在不同的HuBERT骨干上进行持续的自监督训练和微调,以获得(i)语言感知的自监督学习(SSL)编码器和(ii)多任务语音情感分类编码器。然后,我们融合这些编码器,并在下游验证定时检测任务上进一步微调组合模型。TUT情绪讲故事语料库(TESC)的实验评估比较多个模型,融合机制,和训练策略,并表明,所提出的方法实现了显着的改进,传统的语音基线。我们的研究结果表明,非语言的语音提示,当与影响相关的表示,携带足够的信号,以决定何时验证应表示,提供了一个语音第一的途径,更移情的人机交互。
摘要:Emotional Validation is a psychotherapy communication technique that involves recognizing, understanding, and explicitly acknowledging another person's feelings and actions, which strengthens alliance and reduces negative affect. To maximize the emotional support provided by validation, it is crucial to deliver it with appropriate timing and frequency. This study investigates validation timing detection from the speech perspective. Leveraging both paralinguistic and emotional information, we propose a paralinguistic- and emotion-aware model for validation timing detection without relying on textual context. Specifically, we first conduct continued self-supervised training and fine-tuning on different HuBERT backbones to obtain (i) a paralinguistics-aware Self-Supervised Learning (SSL) encoder and (ii) a multi-task speech emotion classification encoder. We then fuse these encoders and further fine-tune the combined model on the downstream validation timing detection task. Experimental evaluations on the TUT Emotional Storytelling Corpus (TESC) compare multiple models, fusion mechanisms, and training strategies, and demonstrate that the proposed approach achieves significant improvements over conventional speech baselines. Our results indicate that non-linguistic speech cues, when integrated with affect-related representations, carry sufficient signal to decide when validation should be expressed, offering a speech-first pathway toward more empathetic human-robot interaction.


【7】How Contrastive Decoding Enhances Large Audio Language Models?
标题:对比解码如何增强大型音频语言模型?
链接:https://arxiv.org/abs/2603.09232

作者:Tzu-Quan Lin,Wei-Ping Huang,Yi-Cheng Lin,Hung-yi Lee
备注:Submitted to INTERSPEECH 2026. Code and additional analysis results are provided in our repository: https://github.com/nervjack2/LALM-Contrastive-Decoding-Error-Profiles
摘要:虽然对比解码(CD)已被证明在增强大型音频语言模型(LALM)方面是有效的,但推动其成功的潜在机制以及不同策略的比较功效仍不清楚。本研究系统地评估了四个不同的CD策略在不同的LALM架构。我们认为音频感知解码和音频对比解码是最有效的方法。然而,其影响因模式而异。为了解释这种变化,我们引入了一个转换矩阵框架来映射推理过程中的错误模式转移。我们的分析表明,CD可靠地纠正模型错误地声称没有音频或诉诸不确定性驱动的猜测的错误。相反,它无法纠正有缺陷的推理或自信的错误断言。最终,这些发现提供了一个明确的指导方针,以确定哪些LALM架构是最适合的CD增强的基础上,他们的基线错误配置文件。
摘要:While Contrastive Decoding (CD) has proven effective at enhancing Large Audio Language Models (LALMs), the underlying mechanisms driving its success and the comparative efficacy of different strategies remain unclear. This study systematically evaluates four distinct CD strategies across diverse LALM architectures. We identify Audio-Aware Decoding and Audio Contrastive Decoding as the most effective methods. However, their impact varies significantly by model. To explain this variability, we introduce a Transition Matrix framework to map error pattern shifts during inference. Our analysis demonstrates that CD reliably rectifies errors in which models falsely claim an absence of audio or resort to uncertainty-driven guessing. Conversely, it fails to correct flawed reasoning or confident misassertions. Ultimately, these findings provide a clear guideline for determining which LALM architectures are most suitable for CD enhancement based on their baseline error profiles.


【8】The Costs of Reproducibility in Music Separation Research: a Replication of Band-Split RNN
标题:音乐分离研究中的复制成本:频段分裂RNN的复制
链接:https://arxiv.org/abs/2603.09187

作者:Paul Magron,Romain Serizel,Constance Douwes
摘要:音乐源分离是将乐器音轨从音乐歌曲中分离出来的任务。尽管最近取得了惊人的进展,但更复杂的架构和训练协议的趋势加剧了可重复性问题。带分裂递归神经网络(BSRNN)模型在这方面很有前途,因为它在公共数据集上产生了接近最先进的结果,并且需要合理的训练资源。不幸的是,它并不容易复制,因为它的完整代码不可用。在本文中,我们试图通过大量的实验尽可能接近地复制BSRNN,这使我们能够对这种可重复性问题进行批判性反思。我们的贡献是三方面的。首先,这项研究对模型设计和训练管道产生了一些见解,这些见解揭示了未来潜在的改进。特别是,由于我们在复制原始结果方面不成功,我们探索了其他变体,最终产生了优化的BSRNN模型,其性能大大提高了原始模型的性能。其次,我们从方法和实践的角度讨论了再现性问题。我们特别强调,如果整个管道可用,本可以节省大量时间和能源成本。第三,我们的代码和预训练模型是公开发布的,以促进可重复的研究。我们希望这项研究将有助于在音乐分离社区中传播对可复制研究重要性的认识,并有助于促进更透明和可持续的实践。
摘要:Music source separation is the task of isolating the instrumental tracks from a music song. Despite its spectacular recent progress, the trend towards more complex architectures and training protocols exacerbates reproducibility issues. The band-split recurrent neural networks (BSRNN) model is promising in this regard, since it yields close to state-of-the-art results on public datasets, and requires reasonable resources for training. Unfortunately, it is not straightforward to reproduce since its full code is not available. In this paper, we attempt to replicate BSRNN as closely as possible to the original paper through extensive experiments, which allows us to conduct a critical reflection on this reproducibility issue. Our contributions are three-fold. First, this study yields several insights on the model design and training pipeline, which sheds light on potential future improvements. In particular, since we were unsuccessful in reproducing the original results, we explore additional variants that ultimately yield an optimized BSRNN model, whose performance largely improves that of the original. Second, we discuss reproducibility issues from both methodological and practical perspectives. We notably underline how substantial time and energy costs could have been saved upon availability of the full pipeline. Third, our code and pre-trained models are released publicly to foster reproducible research. We hope that this study will contribute to spread awareness on the importance of reproducible research in the music separation community, and help promoting more transparent and sustainable practices.


【9】Gender Fairness in Audio Deepfake Detection: Performance and Disparity Analysis
标题:音频Deepfake检测中的性别公平:性能和差异分析
链接:https://arxiv.org/abs/2603.09007

作者:Aishwarya Fursule,Shruti Kshirsagar,Anderson R. Avila
备注:6 pages, 3 Figures
摘要:音频deepfake检测旨在从人工智能(AI)生成的声音中检测真实的人类声音,并已成为语音生物识别系统领域的一个重要问题。随着合成语音质量的不断提高,这种语音被用于身份盗用和模仿等非法行为的可能性也在增加。尽管近年来在音频Deepfake检测领域取得了重大进展,但性别偏见问题仍然没有得到充分研究,并且处于初期阶段。在本文中,我们尝试对音频Deepfake检测模型中的性别相关性能和公平性进行了全面分析。我们使用ASVspoof 5数据集和训练ResNet-18分类器,并评估四种不同音频特征的检测性能,并将性能与基线AASIST模型进行比较。除了等误差率(EER %)等传统指标外,我们还纳入了五个既定的公平性指标,以量化模型中的性别差异。我们的研究结果表明,即使当性别之间的整体EER差异似乎很低,公平意识的评价揭示了错误分布的差异,被掩盖的综合性能指标。这些研究结果表明,依赖于标准的指标是不可靠的,而公平性指标提供了重要的见解,人口特定的故障模式。这项工作强调了公平意识评估对于开发更公平,更强大,更值得信赖的音频deepfake检测系统的重要性。
摘要:Audio deepfake detection aims to detect real human voices from those generated by Artificial Intelligence (AI) and has emerged as a significant problem in the field of voice biometrics systems. With the ever-improving quality of synthetic voice, the probability of such a voice being exploited for illicit practices like identity thest and impersonation increases. Although significant progress has been made in the field of Audio Deepfake Detection in recent times, the issue of gender bias remains underexplored and in its nascent stage In this paper, we have attempted a thorough analysis of gender dependent performance and fairness in audio deepfake detection models. We have used the ASVspoof 5 dataset and train a ResNet-18 classifier and evaluate detection performance across four different audio features, and compared the performance with baseline AASIST model. Beyond conventional metrics such as Equal Error Rate (EER %), we incorporated five established fairness metrics to quantify gender disparities in the model. Our results show that even when the overall EER difference between genders appears low, fairness-aware evaluation reveals disparities in error distribution that are obscured by aggregate performance measures. These findings demonstrate that reliance on standard metrics is unreliable, whereas fairness metrics provide critical insights into demographic-specific failure modes. This work highlights the importance of fairness-aware evaluation for developing a more equitable, robust, and trustworthy audio deepfake detection system.


【10】VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
标题:VoxEmo:使用语音LLM对语音情感识别进行基准测试
链接:https://arxiv.org/abs/2603.08936

作者:Hezhao Zhang,Huang-Cheng Chou,Shrikanth Narayanan,Thomas Hain
备注:submitted to Interspeech 2026
摘要:语音大语言模型(LLM)通过生成接口在语音情感识别(SER)方面显示出巨大的潜力。然而,从闭集分类到开放文本生成的转变引入了zero-shot随机性,使得评估对提示高度敏感。此外,传统的语音LLM基准忽略了人类情感的固有模糊性。因此,我们提出了VoxEmo,一个全面的SER基准,包括15种语言的35个情感语料库的语音LLM。VoxEmo提供了一个标准化的工具包,具有不同的提示复杂性,从直接分类到语言推理。为了反映现实世界的感知/应用,我们引入了一个分布式感知的软标签协议和一个模仿注释者分歧的集成策略。实验表明,虽然zero-shot语音LLM在硬标签准确性方面落后于监督基线,但它们与人类主观分布唯一一致。
摘要:Speech Large Language Models (LLMs) show great promise for speech emotion recognition (SER) via generative interfaces. However, shifting from closed-set classification to open text generation introduces zero-shot stochasticity, making evaluation highly sensitive to prompts. Additionally, conventional speech LLMs benchmarks overlook the inherent ambiguity of human emotion. Hence, we present VoxEmo, a comprehensive SER benchmark encompassing 35 emotion corpora across 15 languages for Speech LLMs. VoxEmo provides a standardized toolkit featuring varying prompt complexities, from direct classification to paralinguistic reasoning. To reflect real-world perception/application, we introduce a distribution-aware soft-label protocol and a prompt-ensemble strategy that emulates annotator disagreement. Experiments reveal that while zero-shot speech LLMs trail supervised baselines in hard-label accuracy, they uniquely align with human subjective distributions.


【11】Fish Audio S2 Technical Report
标题:Fish Audio S2技术报告
链接:https://arxiv.org/abs/2603.08823

作者:Shijia Liao,Yuxuan Wang,Songting Liu,Yifan Cheng,Ruoyi Zhang,Tianyu Li,Shidong Li,Yisheng Zheng,Xingwei Liu,Qingzheng Wang,Zhizhuo Zhou,Jiahua Liu,Xin Chen,Dawei Han
摘要:我们介绍Fish Audio S2,这是一个开源的文本到语音转换系统,具有多扬声器,多回合生成功能,最重要的是,通过自然语言描述进行跟随控制。为了扩展训练,我们开发了一个多阶段的训练配方,以及一个阶段性的数据管道,包括视频字幕和语音字幕,语音质量评估和奖励建模。为了推动开源TTS的前沿,我们发布了我们的模型权重,微调代码和基于SGLang的推理引擎。推理引擎可以用于流媒体制作,实现了0.195的RTF和低于100 ms的首次音频时间。我们的代码和权重可以在GitHub(https://github.com/fishaudio/fish-speech)和Hugging Face(https://huggingface.co/fishaudio/s2-pro)上找到。我们强烈建议读者访问https://fish.audio来尝试自定义语音。
摘要:We introduce Fish Audio S2, an open-sourced text-to-speech system featuring multi-speaker, multi-turn generation, and, most importantly, instruction-following control via natural-language descriptions. To scale training, we develop a multi-stage training recipe together with a staged data pipeline covering video captioning and speech captioning, voice-quality assessment, and reward modeling. To push the frontier of open-source TTS, we release our model weights, fine-tuning code, and an SGLang-based inference engine. The inference engine is production-ready for streaming, achieving an RTF of 0.195 and a time-to-first-audio below 100 ms.Our code and weights are available on GitHub (https://github.com/fishaudio/fish-speech) and Hugging Face (https://huggingface.co/fishaudio/s2-pro). We highly encourage readers to visit https://fish.audio to try custom voices.


【12】EDMFormer: Genre-Specific Self-Supervised Learning for Music Structure Segmentation
标题:EDDMFormer:音乐结构分割的特定流派自我监督学习
链接:https://arxiv.org/abs/2603.08759

作者:Sahal Sajeer,Krish Patel,Oscar Chung,Joel Song Bae
备注:Published in CUCAI 2026 conference proceedings
摘要:音乐结构分割是音频分析中的一项关键任务,但现有的模型在电子舞曲(EDM)中的表现不佳。这个问题的存在是因为大多数方法依赖于抒情或和声的相似性,这对流行音乐很有效,但对EDM却不适用。相反,EDM结构由能量,节奏和音色的变化定义,具有不同的部分,如积累,下降和崩溃。我们介绍了EDMFormer,一个Transformer模型,它使用特定于EDM的数据集和分类法将自监督音频嵌入相结合。我们将此数据集作为EDM-98发布:一组98个专业注释的EDM轨道。与现有模型相比,EDMFormer改进了边界检测和截面标记,特别是对于跌落和堆积。研究结果表明,将学习到的表征与特定类型的数据和结构先验相结合对EDM是有效的,并且可以应用于其他专业音乐类型或更广泛的音频领域。
摘要:Music structure segmentation is a key task in audio analysis, but existing models perform poorly on Electronic Dance Music (EDM). This problem exists because most approaches rely on lyrical or harmonic similarity, which works well for pop music but not for EDM. EDM structure is instead defined by changes in energy, rhythm, and timbre, with different sections such as buildup, drop, and breakdown. We introduce EDMFormer, a transformer model that combines self-supervised audio embeddings using an EDM-specific dataset and taxonomy. We release this dataset as EDM-98: a group of 98 professionally annotated EDM tracks. EDMFormer improves boundary detection and section labelling compared to existing models, particularly for drops and buildups. The results suggest that combining learned representations with genre-specific data and structural priors is effective for EDM and could be applied to other specialized music genres or broader audio domains.


【13】Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
标题:神经音频编解码器容量和鲁棒性之间的权衡以实现对抗鲁棒语音识别
链接:https://arxiv.org/abs/2603.09034

作者:Jordan Prescott,Thanathai Lertpetchpun,Shrikanth Narayanan
备注:Submitted to Interspeech 2026
摘要:对抗扰动利用自动语音识别(ASR)系统中的漏洞,同时保留人类感知的语言内容。神经音频编解码器施加了一个离散的瓶颈,可以抑制与对抗性噪声相关的细粒度信号变化。我们研究了这个瓶颈的粒度,由残差矢量量化(RVQ)深度控制,形状对抗鲁棒性。在基于梯度的攻击下,我们观察到一个非单调的权衡:浅量化抑制了对抗性扰动,但降低了语音内容,而更深的量化保留了内容和扰动。中间深度平衡这些影响并最小化转录错误。我们进一步表明,adversarially引起的变化,离散码本令牌强烈相关的转录错误。这些增益在自适应攻击下持续存在,其中神经编解码器配置优于传统的压缩防御。
摘要:Adversarial perturbations exploit vulnerabilities in automatic speech recognition (ASR) systems while preserving human perceived linguistic content. Neural audio codecs impose a discrete bottleneck that can suppress fine-grained signal variations associated with adversarial noise. We examine how the granularity of this bottleneck, controlled by residual vector quantization (RVQ) depth, shapes adversarial robustness. We observe a non-monotonic trade-off under gradient-based attacks: shallow quantization suppresses adversarial perturbations but degrades speech content, while deeper quantization preserves both content and perturbations. Intermediate depths balance these effects and minimize transcription error. We further show that adversarially induced changes in discrete codebook tokens strongly correlate with transcription error. These gains persist under adaptive attacks, where neural codec configurations outperform traditional compression defenses.


【14】Universal Speech Content Factorization
标题:通用语音内容分解
链接:https://arxiv.org/abs/2603.08977

作者:Henry Li Xinyuan,Zexin Cai,Lin Zhang,Leibny Paola García-Perera,Berrak Sisman,Sanjeev Khudanpur,Nicholas Andrews,Matthew Wiesner
摘要:我们提出了通用语音内容因子分解(USCF),一个简单的和可逆的线性方法提取一个低秩的语音表示,其中扬声器音色被抑制,同时保留语音内容。USCF通过最小二乘优化学习通用语音到内容映射,并从几秒钟的目标语音中导出特定于说话者的转换,将语音内容因子分解(一种闭集语音转换(VC)方法)扩展到开集设置。我们通过嵌入分析表明,USCF有效地消除了说话人相关的变化。作为一个zero-shot VC系统,与需要更多目标说话人数据或额外神经训练的方法相比,USCF实现了具有竞争力的可懂度,自然度和说话人相似性。最后,我们证明,作为一个训练有效的音色解开语音功能,USCF功能可以作为训练音色提示的文本到语音模型的声学表示。语音样本和代码是公开的。
摘要:We propose Universal Speech Content Factorization (USCF), a simple and invertible linear method for extracting a low-rank speech representation in which speaker timbre is suppressed while phonetic content is preserved. USCF extends Speech Content Factorization, a closed-set voice conversion (VC) method, to an open-set setting by learning a universal speech-to-content mapping via least-squares optimization and deriving speaker-specific transformations from only a few seconds of target speech. We show through embedding analysis that USCF effectively removes speaker-dependent variation. As a zero-shot VC system, USCF achieves competitive intelligibility, naturalness, and speaker similarity compared to methods that require substantially more target-speaker data or additional neural training. Finally, we demonstrate that as a training-efficient timbre-disentangled speech feature, USCF features can serve as the acoustic representation for training timbre-prompted text-to-speech models. Speech samples and code are publicly available.


eess.AS音频处理


【1】Distributed Multichannel Wiener Filtering for Wireless Acoustic Sensor Networks
标题:无线声学传感器网络的分布式多通道维纳过滤
链接:https://arxiv.org/abs/2603.09735

作者:Paul Didier,Toon van Waterschoot,Simon Doclo,Jörg Bitzer,Pourya Behmandpoor,Henri Gode,Marc Moonen
摘要:在无线声学传感器网络(WESTERN)中,设备(即,节点)可以通过分布式算法协作以共同执行音频信号处理任务。本文主要研究了基于网络维纳滤波的节点特定期望语音信号的分布式估计。目标是匹配将访问所有麦克风信号的集中式系统的性能,同时减少算法的通信带宽使用。现有的解决方案,如分布式自适应节点特定信号估计(DANSE)算法,收敛到多通道维纳滤波器(MWF),解决了集中式线性最小均方误差(LMMSE)信号估计问题。然而,他们迭代地这样做,这可能是缓慢和不切实际的。许多解决方案还假设所有节点都观察到相同的感兴趣源集合,但实际情况往往并非如此。为了克服这些限制,我们提出了分布式多通道维纳滤波器(dMWF)的全连接WASNs。dMWF是非迭代和最佳的,即使当节点观察不同的源集。在该算法中,节点交换特定于邻居对的低维(融合)信号,估计由该对中的两个节点观察到的源的贡献。我们正式证明了dMWF的最优性,并在模拟语音增强实验中证明了其性能。所提出的算法表现出优于DANSE的客观指标后,较短的操作时间,突出其迭代设计的好处。
摘要:In a wireless acoustic sensor network (WASN), devices (i.e., nodes) can collaborate through distributed algorithms to collectively perform audio signal processing tasks. This paper focuses on the distributed estimation of node-specific desired speech signals using network-wide Wiener filtering. The objective is to match the performance of a centralized system that would have access to all microphone signals, while reducing the communication bandwidth usage of the algorithm. Existing solutions, such as the distributed adaptive node-specific signal estimation (DANSE) algorithm, converge towards the multichannel Wiener filter (MWF) which solves a centralized linear minimum mean square error (LMMSE) signal estimation problem. However, they do so iteratively, which can be slow and impractical. Many solutions also assume that all nodes observe the same set of sources of interest, which is often not the case in practice. To overcome these limitations, we propose the distributed multichannel Wiener filter (dMWF) for fully connected WASNs. The dMWF is non-iterative and optimal even when nodes observe different sets of sources. In this algorithm, nodes exchange neighbor-pair-specific, low-dimensional (fused) signals estimating the contribution of sources observed by both nodes in the pair. We formally prove the optimality of dMWF and demonstrate its performance in simulated speech enhancement experiments. The proposed algorithm is shown to outperform DANSE in terms of objective metrics after short operation times, highlighting the benefit of its iterationless design.


【2】A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
标题:用于语音增强和语音识别的半自发荷兰语音数据集
链接:https://arxiv.org/abs/2603.09725

作者:Dimme de Groot,Yuanyuan Zhang,Jorge Martinez,Odette Scharenborg
备注:Submitted to Interspeech 2026
摘要:我们提出DRES:一个1.5小时的荷兰现实引发(半自发)语音数据集从80扬声器记录在嘈杂的,公共室内环境。DRES被设计为一个测试集,用于评估现实世界场景中最先进的(SOTA)自动语音识别(ASR)和语音增强(SE)模型:一个人在有背景说话者和噪音的公共室内空间中说话。演讲是用四通道线性麦克风阵列记录的。在这项工作中,我们评估了五个著名的单通道SE算法的语音质量和8 SOTA现成的ASR模型的识别性能之前和之后,应用SE的语音DRES。我们发现,尽管条件具有挑战性,但八个ASR模型中有五个在DRES上的WER低于22%。与最近的工作相比,我们没有发现现代单通道SE对ASR性能的积极影响,强调了在现实条件下评估的重要性。
摘要:We present DRES: a 1.5-hour Dutch realistic elicited (semi-spontaneous) speech dataset from 80 speakers recorded in noisy, public indoor environments. DRES was designed as a test set for the evaluation of state-of-the-art (SOTA) automatic speech recognition (ASR) and speech enhancement (SE) models in a real-world scenario: a person speaking in a public indoor space with background talkers and noise. The speech was recorded with a four-channel linear microphone array. In this work we evaluate the speech quality of five well-known single-channel SE algorithms and the recognition performance of eight SOTA off-the-shelf ASR models before and after applying SE on the speech of DRES. We found that five out of the eight ASR models have WERs lower than 22% on DRES, despite the challenging conditions. In contrast to recent work, we did not find a positive effect of modern single-channel SE on ASR performance, emphasizing the importance of evaluating in realistic conditions.


【3】Finetuning a Text-to-Audio Model for Room Impulse Response Generation
标题:微调文本到音频模型以生成房间脉冲响应
链接:https://arxiv.org/abs/2603.09708

作者:Kirak Kim,Sungyoung Kim
备注:5 pages, 2 figures, submitted to Interspeech 2026
摘要:房间脉冲响应(RIR)实现了逼真的声学模拟,应用范围从多媒体制作到语音数据增强。然而,获取高质量的真实世界RIR是劳动密集型的,数据稀缺仍然是数据驱动的RIR生成方法的挑战。在本文中,我们提出了一种通过微调预训练的文本到音频模型来生成RIR的新方法,首次证明了大规模生成音频先验可以有效地用于该任务。为了解决文本RIR配对数据的缺乏,我们建立了一个标记管道,利用视觉语言模型从现有的图像RIR数据集提取声学描述。我们引入了一个上下文学习策略,以适应自由形式的用户提示推理。包括MUSHRA听力测试和下游ASR性能的评估表明,我们的模型生成合理的RIR,并作为语音数据增强的有效工具。
摘要:Room Impulse Responses (RIRs) enable realistic acoustic simulation, with applications ranging from multimedia production to speech data augmentation. However, acquiring high-quality real-world RIRs is labor-intensive, and data scarcity remains a challenge for data-driven RIR generation approaches. In this paper, we propose a novel approach to RIR generation by fine-tuning a pre-trained text-to-audio model, demonstrating for the first time that large-scale generative audio priors can be effectively leveraged for the task. To address the lack of text-RIR paired data, we establish a labeling pipeline utilizing vision-language models to extract acoustic descriptions from existing image-RIR datasets. We introduce an in-context learning strategy to accommodate free-form user prompts during inference. Evaluations involving MUSHRA listening tests and downstream ASR performance demonstrate that our model generates plausible RIRs and serves as an effective tool for speech data augmentation.


【4】Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models
标题:Speech-Omni-Lite:视觉语言模型的便携式语音界面
链接:https://arxiv.org/abs/2603.09627

作者:Dehua Tao,Xuan Luo,Daxin Tan,Kai Chen,Lanqing Hong,Jing Li,Ruifeng Xu,Xiao Chen
摘要:虽然大规模全模型在各种模态中表现出令人印象深刻的能力,但其强大的性能严重依赖于大量的多模态数据,并产生大量的计算成本。这项工作介绍了Speech-Omni-Lite,这是一个具有成本效益的框架,用于扩展具有语音理解和生成功能的预训练视觉语言(VL)骨干,同时完全保留骨干的视觉语言性能。具体来说,VL骨干配备了两个轻量级的,可训练的即插即用模块,一个语音投影仪和一个语音令牌生成器,同时保持VL骨干完全冻结。为了缓解口语QA语料库的稀缺性,提出了一种低成本的数据构建策略,以从现有的ASR语音-文本对生成文本-文本-语音(QTATS)数据,从而促进有效的语音生成训练。实验结果表明,即使只有数千小时的语音训练数据,Speech-Omni-Lite也能实现出色的口语QA性能,可与数百万小时语音数据训练的全模型相媲美。此外,学习的语音模块表现出很强的跨VL骨干的可转移性。
摘要:While large-scale omni-models have demonstrated impressive capabilities across various modalities, their strong performance heavily relies on massive multimodal data and incurs substantial computational costs. This work introduces Speech-Omni-Lite, a cost-efficient framework for extending pre-trained Visual-Language (VL) backbones with speech understanding and generation capabilities, while fully preserving the backbones' vision-language performance. Specifically, the VL backbone is equipped with two lightweight, trainable plug-and-play modules, a speech projector and a speech token generator, while keeping the VL backbone fully frozen. To mitigate the scarcity of spoken QA corpora, a low-cost data construction strategy is proposed to generate Question-Text Answer-Text-Speech (QTATS) data from existing ASR speech-text pairs, facilitating effective speech generation training. Experimental results show that, even with only thousands of hours of speech training data, Speech-Omni-Lite achieves excellent spoken QA performance, which is comparable to omni-models trained on millions of hours of speech data. Furthermore, the learned speech modules exhibit strong transferability across VL backbones.


【5】A Fast Solver for Interpolating Stochastic Differential Equation Diffusion Models for Speech Restoration
标题:语音恢复随机方程扩散模型内插的快速求解器
链接:https://arxiv.org/abs/2603.09508

作者:Bunlong Lay,Timo Gerkmann
摘要:扩散概率模型(DPM)是一种成熟的无条件图像生成扩散模型,而SGMSE+是一种成熟的语音增强条件扩散模型。扩散模型的缺点之一是解决反向过程需要对大型神经网络进行多次评估。虽然先进的快速采样求解器已开发的DPM,他们不直接适用于模型,如SGMSE+由于其扩散过程的差异。具体来说,DPM在数据分布和标准高斯分布之间进行转换,而SGMSE+在目标分布和噪声观测之间进行插值。这项工作首先开发了一个形式主义的插值随机微分方程(iSDEs),其中包括SGMSE+,第二,提出了一个求解器的iSDEs。所提出的求解器可以在多个语音恢复任务中使用最少10个神经网络评估进行快速采样。
摘要:Diffusion Probabilistic Models (DPMs) are a well-established class of diffusion models for unconditional image generation, while SGMSE+ is a well-established conditional diffusion model for speech enhancement. One of the downsides of diffusion models is that solving the reverse process requires many evaluations of a large Neural Network. Although advanced fast sampling solvers have been developed for DPMs, they are not directly applicable to models such as SGMSE+ due to differences in their diffusion processes. Specifically, DPMs transform between the data distribution and a standard Gaussian distribution, whereas SGMSE+ interpolates between the target distribution and a noisy observation. This work first develops a formalism of interpolating Stochastic Differential Equations (iSDEs) that includes SGMSE+, and second proposes a solver for iSDEs. The proposed solver enables fast sampling with as few as 10 Neural Network evaluations across multiple speech restoration tasks.


【6】End-to-End Direction-Aware Keyword Spotting with Spatial Priors in Noisy Environments
标题:噪声环境下基于空间先验的端到端方向感知关键词发现
链接:https://arxiv.org/abs/2603.09505

作者:Rui Wang,Zhifei Zhang,Yu Gao,Xiaofeng Mou,Yi Xu
备注:Submitted for review to Interspeech 2026
摘要:关键词定位(KWS)对于许多语音驱动的应用程序至关重要,但在嘈杂的环境中鲁棒的KWS仍然具有挑战性。传统系统通常依赖于单通道输入和将前端增强与KWS分离的级联管道。这妨碍了联合优化,从而固有地限制了性能。我们提出了一个端到端的多通道KWS框架,利用空间线索,以提高噪声鲁棒性。空间编码器学习通道间特征,而空间嵌入注入方向先验;融合表示由流骨干处理。在多个信噪比(SNR)的模拟噪声条件下的实验表明,空间建模和方向先验各自产生明显的收益超过基线,与它们的组合实现最佳效果。这些发现验证了端到端多通道空间建模,表明在复杂声学场景中目标说话人感知检测的强大潜力。
摘要:Keyword spotting (KWS) is crucial for many speech-driven applications, but robust KWS in noisy environments remains challenging. Conventional systems often rely on single-channel inputs and a cascaded pipeline separating front-end enhancement from KWS. This precludes joint optimization, inherently limiting performance. We present an end-to-end multi-channel KWS framework that exploits spatial cues to improve noise robustness. A spatial encoder learns inter-channel features, while a spatial embedding injects directional priors; the fused representation is processed by a streaming backbone. Experiments in simulated noisy conditions across multiple signal-to-noise ratios (SNRs) show that spatial modeling and directional priors each yield clear gains over baselines, with their combination achieving the best results. These findings validate end-to-end multi-channel spatial modeling, indicating strong potential for the target-speaker-aware detection in complex acoustic scenarios.


【7】StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
标题:StuPASE:迈向低幻觉演播室质量生成语音增强
链接:https://arxiv.org/abs/2603.09234

作者:Xiaobin Rong,Jun Gao,Zheng Wang,Mansur Yesilbursa,Kamil Wojcicki,Jing Lu
备注:Submitted to Interspeech 2026
摘要:在生成式语音增强中,实现高感知质量而不产生幻觉仍然是一个挑战。代表性的方法PASE对幻觉是鲁棒的,但是在不利条件下具有有限的感知质量。我们提出了Stupase,建立在PASE,以实现演播室级的质量,同时保留其低幻觉属性。首先,我们表明,微调PASE与干目标,而不是目标包含模拟早期反射大大提高了去混响。其次,为了解决强加性噪声下的性能限制,我们将PASE中基于GAN的生成模块替换为流匹配模块,即使在极具挑战性的条件下也能实现演播室质量的生成。实验表明,StupASE始终产生高质量的语音,同时保持低幻觉,优于最先进的SE方法。音频演示可在https://xiaobin-rong.github.io/stupase_demo/上获得。
摘要:Achieving high perceptual quality without hallucination remains a challenge in generative speech enhancement (SE). A representative approach, PASE, is robust to hallucination but has limited perceptual quality under adverse conditions. We propose StuPASE, built upon PASE to achieve studio-level quality while retaining its low-hallucination property. First, we show that finetuning PASE with dry targets rather than targets containing simulated early reflections substantially improves dereverberation. Second, to address performance limitations under strong additive noise, we replace the GAN-based generative module in PASE with a flow-matching module, enabling studio-quality generation even under highly challenging conditions. Experiments demonstrate that StuPASE consistently produces perceptually high-quality speech while maintaining low hallucination, outperforming state-of-the-art SE methods. Audio demos are available at: https://xiaobin-rong.github.io/stupase_demo/.


【8】Acoustic and Semantic Modeling of Emotion in Spoken Language
标题:口语情感的声学和语义建模
链接:https://arxiv.org/abs/2603.09212

作者:Soumya Dutta
备注:PhD thesis
摘要:情绪在人类沟通中发挥着核心作用,塑造信任,参与和社交互动。随着由大型语言模型驱动的人工智能系统越来越多地融入日常生活,使它们能够可靠地理解和生成人类情感仍然是一个重要的挑战。虽然情感表达本质上是多模态的,但本文主要关注通过口语传达的情感,并研究如何将声学和语义信息联合建模,以促进情感理解和情感合成。本文第一部分研究了通过预训练实现的情绪感知表征学习。我们提出的策略,将声学和语义监督学习表示,更好地捕捉语音中的情感线索。还引入了一个语音驱动的监督预训练框架,以实现大规模的情感感知文本建模,而不需要手动注释的文本语料库。第二部分讨论了会话环境中的情感识别。结合跨模态注意和混合专家融合的分层架构,开发了跨会话轮整合声学和语义信息。最后,本文介绍了一个无文本和非并行语音到语音的框架,情感风格转移,使可控的情感转换,同时保持说话人的身份和语言内容。结果表明,改进的情感转移,并表明风格转移的语音可以用于数据增强,以提高情感识别。
摘要:Emotions play a central role in human communication, shaping trust, engagement, and social interaction. As artificial intelligence systems powered by large language models become increasingly integrated into everyday life, enabling them to reliably understand and generate human emotions remains an important challenge. While emotional expression is inherently multimodal, this thesis focuses on emotions conveyed through spoken language and investigates how acoustic and semantic information can be jointly modeled to advance both emotion understanding and emotion synthesis from speech. The first part of the thesis studies emotion-aware representation learning through pre-training. We propose strategies that incorporate acoustic and semantic supervision to learn representations that better capture affective cues in speech. A speech-driven supervised pre-training framework is also introduced to enable large-scale emotion-aware text modeling without requiring manually annotated text corpora. The second part addresses emotion recognition in conversational settings. Hierarchical architectures combining cross-modal attention and mixture-of-experts fusion are developed to integrate acoustic and semantic information across conversational turns. Finally, the thesis introduces a textless and non-parallel speech-to-speech framework for emotion style transfer that enables controllable emotional transformations while preserving speaker identity and linguistic content. The results demonstrate improved emotion transfer and show that style-transferred speech can be used for data augmentation to improve emotion recognition.


【9】Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
标题:识别性前置:语音转换模型中的显式情感控制
链接:https://arxiv.org/abs/2603.09120

作者:Haoyuan Yang,Mu Yang,Jiamin Xie,Szu-Jui Chen,John H. L. Hansen
备注:Submitted to Interspeech 2026
摘要:zero-shot语音转换的最新进展在情感控制方面表现出了潜力,但由于其有限的表达能力,其性能并不理想或不一致。我们提出了一个两阶段的语音转换骨干中的显式情感控制的预防意识前缀。我们显著提高了情感转换性能,将基线情感转换准确率(ECA)从42.40%提高了一倍,达到85.50%,同时保持了语言完整性和语音质量,而不会影响说话人的身份。我们的消融研究表明,序列调制和声学实现的联合控制对于合成不同的情感至关重要。通过对比分析,验证了该方法的通用性,同时也揭示了声解耦在保持说话人身份方面的作用。
摘要:Recent advances in zero-shot voice conversion have exhibited potential in emotion control, yet the performance is suboptimal or inconsistent due to their limited expressive capacity. We propose Emotion-Aware Prefix for explicit emotion control in a two-stage voice conversion backbone. We significantly improve emotion conversion performance, doubling the baseline Emotion Conversion Accuracy (ECA) from 42.40% to 85.50% while maintaining linguistic integrity and speech quality, without compromising speaker identity. Our ablation study suggests that a joint control of both sequence modulation and acoustic realization is essential to synthesize distinct emotions. Furthermore, comparative analysis verifies the generalizability of proposed method, while it provides insights on the role of acoustic decoupling in maintaining speaker identity.


【10】Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
标题:神经音频编解码器容量和鲁棒性之间的权衡以实现对抗鲁棒语音识别
链接:https://arxiv.org/abs/2603.09034

作者:Jordan Prescott,Thanathai Lertpetchpun,Shrikanth Narayanan
备注:Submitted to Interspeech 2026
摘要:对抗扰动利用自动语音识别(ASR)系统中的漏洞,同时保留人类感知的语言内容。神经音频编解码器施加了一个离散的瓶颈,可以抑制与对抗性噪声相关的细粒度信号变化。我们研究了这个瓶颈的粒度,由残差矢量量化(RVQ)深度控制,形状对抗鲁棒性。在基于梯度的攻击下,我们观察到一个非单调的权衡:浅量化抑制了对抗性扰动,但降低了语音内容,而更深的量化保留了内容和扰动。中间深度平衡这些影响并最小化转录错误。我们进一步表明,adversarially引起的变化,离散码本令牌强烈相关的转录错误。这些增益在自适应攻击下持续存在,其中神经编解码器配置优于传统的压缩防御。
摘要:Adversarial perturbations exploit vulnerabilities in automatic speech recognition (ASR) systems while preserving human perceived linguistic content. Neural audio codecs impose a discrete bottleneck that can suppress fine-grained signal variations associated with adversarial noise. We examine how the granularity of this bottleneck, controlled by residual vector quantization (RVQ) depth, shapes adversarial robustness. We observe a non-monotonic trade-off under gradient-based attacks: shallow quantization suppresses adversarial perturbations but degrades speech content, while deeper quantization preserves both content and perturbations. Intermediate depths balance these effects and minimize transcription error. We further show that adversarially induced changes in discrete codebook tokens strongly correlate with transcription error. These gains persist under adaptive attacks, where neural codec configurations outperform traditional compression defenses.


【11】Universal Speech Content Factorization
标题:通用语音内容分解
链接:https://arxiv.org/abs/2603.08977

作者:Henry Li Xinyuan,Zexin Cai,Lin Zhang,Leibny Paola García-Perera,Berrak Sisman,Sanjeev Khudanpur,Nicholas Andrews,Matthew Wiesner
摘要:我们提出了通用语音内容因子分解(USCF),一个简单的和可逆的线性方法提取一个低秩的语音表示,其中扬声器音色被抑制,同时保留语音内容。USCF通过最小二乘优化学习通用语音到内容映射并仅从几秒的目标语音中推导出特定于说话者的转换,将语音内容分解(一种闭集语音转换(VC)方法)扩展到开放集设置。语音。我们通过嵌入分析表明,USCF有效地消除了说话人相关的变化。作为一个zero-shot VC系统,与需要更多目标说话人数据或额外神经训练的方法相比,USCF实现了具有竞争力的可懂度,自然度和说话人相似性。最后,我们证明,作为一个训练有效的音色解开语音功能,USCF功能可以作为训练音色提示的文本到语音模型的声学表示。语音样本和代码是公开的。
摘要:We propose Universal Speech Content Factorization (USCF), a simple and invertible linear method for extracting a low-rank speech representation in which speaker timbre is suppressed while phonetic content is preserved. USCF extends Speech Content Factorization, a closed-set voice conversion (VC) method, to an open-set setting by learning a universal speech-to-content mapping via least-squares optimization and deriving speaker-specific transformations from only a few seconds of target speech. We show through embedding analysis that USCF effectively removes speaker-dependent variation. As a zero-shot VC system, USCF achieves competitive intelligibility, naturalness, and speaker similarity compared to methods that require substantially more target-speaker data or additional neural training. Finally, we demonstrate that as a training-efficient timbre-disentangled speech feature, USCF features can serve as the acoustic representation for training timbre-prompted text-to-speech models. Speech samples and code are publicly available.


【12】MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
标题:MUgen:评估和改善大型音频语言模型的多音频理解
链接:https://arxiv.org/abs/2603.09714

作者:Chih-Kai Yang,Yun-Shao Tsai,Yu-Kai Guo,Ping-Le Tsai,Yen-Ting Piao,Hung-Wei Chen,Ting-Lin Hsiao,Yun-Man Hsu,Ke-Han Lu,Hung-yi Lee
备注:6 pages, 3 figures, 3 tables. Dataset: https://huggingface.co/Multi-Audio-Grounding
摘要:虽然多音频理解对于大型音频语言模型(LALM)至关重要,但它仍然没有得到充分的探索。我们介绍MUGEN,这是一个全面的基准测试,用于评估语音、通用音频和音乐的这种能力。我们的实验揭示了多音频设置中的一致弱点,并且随着并发音频输入数量的增加,性能急剧下降,将输入缩放确定为根本瓶颈。我们进一步研究了免训练策略,并观察到音频置换自一致性(Audio-Permutational Self-Consistency)使音频候选项的顺序多样化,有助于模型形成更强大的聚合预测,准确率提高了6.28%。将这种排列策略与Chain-of-Thought相结合,进一步将性能提高到6.74%。这些结果暴露了当前LALM中的盲点,并为评估复杂的听觉理解提供了基础。
摘要:While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and music. Our experiments reveal consistent weaknesses in multi-audio settings, and performance degrades sharply as the number of concurrent audio inputs increases, identifying input scaling as a fundamental bottleneck. We further investigate training-free strategies and observe that Audio-Permutational Self-Consistency, which diversifies the order of audio candidates, helps models form more robust aggregated predictions, yielding up to 6.28% accuracy gains. Combining this permutation strategy with Chain-of-Thought further improves performance to 6.74%. These results expose blind spots in current LALMs and provide a foundation for evaluating complex auditory comprehension.


【13】Physics-Informed Neural Engine Sound Modeling with Differentiable Pulse-Train Synthesis
标题:具有可区分脉冲串合成的物理信息神经引擎声音建模
链接:https://arxiv.org/abs/2603.09391

作者:Robin Doerfler,Lonce Wyse
备注:Preprint. 5 pages, 2 figures. Audio examples, code, and model weights available online
摘要:发动机声音来源于连续的排气压力脉冲,而不是持续的谐波振荡。虽然神经合成方法通常旨在近似得到的频谱特性,但我们建议直接对潜在的脉冲形状和时间结构进行建模。我们提出了脉冲串谐振器(PTR)模型,一个可微分的合成架构,生成发动机音频参数化脉冲串对齐发动机点火模式,并通过递归的Karplus-Strong谐振器模拟排气声学传播。该架构集成了物理信息的感应偏置,包括谐波衰减,热力学桨距调制,阀动态包络,排气系统共振和衍生的发动机操作模式,如油门操作和减速燃料切断(DCFO)。   在三种不同的引擎类型上进行了总计7.5小时的音频验证,PTR在谐波重建方面实现了21%的改进,并在谐波加噪声基线模型上减少了5.7%的总损失,同时提供了与物理现象相对应的可解释参数。   完整的代码、模型权重和音频示例都是公开的。
摘要:Engine sounds originate from sequential exhaust pressure pulses rather than sustained harmonic oscillations. While neural synthesis methods typically aim to approximate the resulting spectral characteristics, we propose directly modeling the underlying pulse shapes and temporal structure. We present the Pulse-Train-Resonator (PTR) model, a differentiable synthesis architecture that generates engine audio as parameterized pulse trains aligned to engine firing patterns and propagates them through recursive Karplus-Strong resonators simulating exhaust acoustics. The architecture integrates physics-informed inductive biases including harmonic decay, thermodynamic pitch modulation, valve-dynamics envelopes, exhaust system resonances and derived engine operating modes such as throttle operation and deceleration fuel cutoff (DCFO).   Validated on three diverse engine types totaling 7.5 hours of audio, PTR achieves a 21% improvement in harmonic reconstruction and a 5.7% reduction in total loss over a harmonic-plus-noise baseline model, while providing interpretable parameters corresponding to physical phenomena.   Complete code, model weights, and audio examples are openly available.


【14】How Contrastive Decoding Enhances Large Audio Language Models?
标题:对比解码如何增强大型音频语言模型?
链接:https://arxiv.org/abs/2603.09232

作者:Tzu-Quan Lin,Wei-Ping Huang,Yi-Cheng Lin,Hung-yi Lee
备注:Submitted to INTERSPEECH 2026. Code and additional analysis results are provided in our repository: https://github.com/nervjack2/LALM-Contrastive-Decoding-Error-Profiles
摘要:虽然对比解码(CD)已被证明在增强大型音频语言模型(LALM)方面是有效的,但推动其成功的潜在机制以及不同策略的比较功效仍不清楚。本研究系统地评估了四个不同的CD策略在不同的LALM架构。我们认为音频感知解码和音频对比解码是最有效的方法。然而,其影响因模式而异。为了解释这种变化,我们引入了一个转换矩阵框架来映射推理过程中的错误模式转移。我们的分析表明,CD可靠地纠正模型错误地声称没有音频或诉诸不确定性驱动的猜测的错误。相反,它无法纠正有缺陷的推理或自信的错误断言。最终,这些发现提供了一个明确的指导方针,以确定哪些LALM架构是最适合的CD增强的基础上,他们的基线错误配置文件。
摘要:While Contrastive Decoding (CD) has proven effective at enhancing Large Audio Language Models (LALMs), the underlying mechanisms driving its success and the comparative efficacy of different strategies remain unclear. This study systematically evaluates four distinct CD strategies across diverse LALM architectures. We identify Audio-Aware Decoding and Audio Contrastive Decoding as the most effective methods. However, their impact varies significantly by model. To explain this variability, we introduce a Transition Matrix framework to map error pattern shifts during inference. Our analysis demonstrates that CD reliably rectifies errors in which models falsely claim an absence of audio or resort to uncertainty-driven guessing. Conversely, it fails to correct flawed reasoning or confident misassertions. Ultimately, these findings provide a clear guideline for determining which LALM architectures are most suitable for CD enhancement based on their baseline error profiles.


【15】SPAR-K: Scheduled Periodic Alternating Early Exit for Spoken Language Models
标题:SPAR-K:口语模型计划定期交替提前退出
链接:https://arxiv.org/abs/2603.09215

作者:Hsiao-Ying Huang,Cheng-Han Chiang,Hung-yi Lee
备注:6 pages, 1 figures, 2 tables
摘要:交错口语模型(SLM)交替生成文本和语音标记,但每一步都以全Transformer深度解码变得成本高昂,特别是由于长语音序列。我们提出了SPAR-K,一个模态感知的早期退出框架,旨在加速交织SLM推理,同时保持感知质量。SPAR-K引入了语音交替深度调度:大多数语音位置在固定的中间层退出,而周期性的全深度“刷新”步骤减轻了由于提前退出而导致的分布偏移。我们使用Step-Audio-2-mini和GLM-4-Voice在四个数据集上评估我们的框架,这些数据集涵盖推理,事实QA和对话任务,测量ASR转录准确性和感知质量方面的性能。实验结果表明,SPAR-K在很大程度上保持了问答准确率,最大准确率下降了0.82%,而在Step-Audio-2-mini上平均语音解码深度降低了11%,在GLM-4-Voice上降低了5%,MOS和WER的变化可以忽略不计,并且没有辅助计算开销。我们进一步证明,基于信心的早期退出策略,广泛用于文本LLM,是次优的SLM,突出的语音令牌的独特的统计性质需要一个专门的早期退出设计。
摘要:Interleaved spoken language models (SLMs) alternately generate text and speech tokens, but decoding at full transformer depth for every step becomes costly, especially due to long speech sequences. We propose SPAR-K, a modality-aware early exit framework designed to accelerate interleaved SLM inference while preserving perceptual quality. SPAR-K introduces a speech alternating-depth schedule: most speech positions exit at a fixed intermediate layer, while periodic full-depth "refresh" steps mitigate distribution shift due to early exit. We evaluate our framework using Step-Audio-2-mini and GLM-4-Voice across four datasets spanning reasoning, factual QA, and dialogue tasks, measuring performance in terms of ASR transcription accuracy and perceptual quality. Experimental results demonstrate that SPAR-K largely preserves question-answering accuracy with a maximum accuracy drop of 0.82\% while reducing average speech decoding depth by up to 11\% on Step-Audio-2-mini and 5\% on GLM-4-Voice, both with negligible changes in MOS and WER and no auxiliary computation overhead. We further demonstrate that confidence-based early exit strategies, widely used in text LLMs, are suboptimal for SLMs, highlighting that the unique statistical nature of speech tokens necessitates a specialized early exit design.


【16】Can You Hear, Localize, and Segment Continually? An Exemplar-Free Continual Learning Benchmark for Audio-Visual Segmentation
标题:您能持续聆听、本地化和细分吗?视听分割的无范例持续学习基准
链接:https://arxiv.org/abs/2603.08967

作者:Siddeshwar Raghavan,Gautham Vinod,Bruce Coburn,Fengqing Zhu
摘要:视听分割(AVS)的目的是通过从音频和视频信号中联合学习,产生视频中声音产生对象的像素级掩模。然而,现实世界的环境本质上是动态的,导致音频和视频分布随着时间的推移而演变,这对现有的假设静态训练设置的AVS系统提出了挑战。为了解决这一差距,我们引入了第一个用于视听分割的无样本持续学习基准,包括跨单源和多源AVS数据集的四个学习协议。我们进一步提出了一个强大的基线,ATLAS,它使用音频引导的融合前条件反射,通过预测的音频上下文前跨模态注意调制视觉特征通道。最后,我们通过引入低秩排序(LRA)来减轻灾难性遗忘,该方法基于损失敏感性来稳定自适应权重。广泛的实验证明了在各种连续场景中的竞争力,为终身视听感知奠定了基础。代码可在${}^{*}$\footnote{Paper under review} - \hyperlink{https://gitlab.com/viper-purdue/atlas}{https://gitlab.com/viper-purdue/atlas}获得   关键词:连续学习、视听分割、多模态学习
摘要:Audio-Visual Segmentation (AVS) aims to produce pixel-level masks of sound producing objects in videos, by jointly learning from audio and visual signals. However, real-world environments are inherently dynamic, causing audio and visual distributions to evolve over time, which challenge existing AVS systems that assume static training settings. To address this gap, we introduce the first exemplar-free continual learning benchmark for Audio-Visual Segmentation, comprising four learning protocols across single-source and multi-source AVS datasets. We further propose a strong baseline, ATLAS, which uses audio-guided pre-fusion conditioning to modulate visual feature channels via projected audio context before cross-modal attention. Finally, we mitigate catastrophic forgetting by introducing Low-Rank Anchoring (LRA), which stabilizes adapted weights based on loss sensitivity. Extensive experiments demonstrate competitive performance across diverse continual scenarios, establishing a foundation for lifelong audio-visual perception. Code is available at${}^{*}$\footnote{Paper under review} - \hyperlink{https://gitlab.com/viper-purdue/atlas}{https://gitlab.com/viper-purdue/atlas}   \keywords{Continual Learning \and Audio-Visual Segmentation \and Multi-Modal Learning}


【17】VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
标题:VoxEmo:使用语音LLM对语音情感识别进行基准测试
链接:https://arxiv.org/abs/2603.08936

作者:Hezhao Zhang,Huang-Cheng Chou,Shrikanth Narayanan,Thomas Hain
备注:submitted to Interspeech 2026
摘要:语音大语言模型(LLM)通过生成接口在语音情感识别(SER)方面显示出巨大的潜力。然而,从闭集分类到开放文本生成的转变引入了zero-shot随机性,使得评估对提示高度敏感。此外,传统的语音LLM基准忽略了人类情感的固有模糊性。因此,我们提出了VoxEmo,一个全面的SER基准,包括15种语言的35个情感语料库的语音LLM。VoxEmo提供了一个标准化的工具包,具有不同的提示复杂性,从直接分类到语言推理。为了反映现实世界的感知/应用,我们引入了一个分布式感知的软标签协议和一个模仿注释者分歧的集成策略。实验表明,虽然zero-shot语音LLM在硬标签准确性方面落后于监督基线,但它们与人类主观分布唯一一致。
摘要:Speech Large Language Models (LLMs) show great promise for speech emotion recognition (SER) via generative interfaces. However, shifting from closed-set classification to open text generation introduces zero-shot stochasticity, making evaluation highly sensitive to prompts. Additionally, conventional speech LLMs benchmarks overlook the inherent ambiguity of human emotion. Hence, we present VoxEmo, a comprehensive SER benchmark encompassing 35 emotion corpora across 15 languages for Speech LLMs. VoxEmo provides a standardized toolkit featuring varying prompt complexities, from direct classification to paralinguistic reasoning. To reflect real-world perception/application, we introduce a distribution-aware soft-label protocol and a prompt-ensemble strategy that emulates annotator disagreement. Experiments reveal that while zero-shot speech LLMs trail supervised baselines in hard-label accuracy, they uniquely align with human subjective distributions.


机器翻译由腾讯交互翻译提供,仅供参考