今日论文合集:cs.SD语音25篇,eess.AS音频处理21篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
标题:高保真音频无损压缩的基准语言建模
链接:https://arxiv.org/abs/2603.08683

作者:Phillip Long,Zachary Novack,Chris Donahue
备注:Submitted for review at Interspeech 2026, 7 pages, 5 figures
摘要:在原始波形上训练的自回归“语言”模型(LM)可以重新用于无损音频压缩,但之前的工作仅限于8位音频,因此这些方法是否适用于实际设置(16/24位)以及是否可以与现有编解码器竞争尚不清楚。我们在不同领域(音乐,语音,生物声学),采样率(16 kHz-48 kHz)和位深度(8,16,24位)的全保真度音频上对基于LM的压缩进行基准测试。由于词汇量的大小(16位为65 K; 24位为16.7M),标准的样本级标记化在更高的位深度上变得难以处理。我们提出了Trilobyte,一种用于全分辨率音频的字节级标记化方案,将词汇扩展从$O(2^{b})$提高到$O(1)$,并实现了第一个易于处理的24位基于LM的无损压缩。虽然LM始终优于FLAC,并在8位和16位时产生最先进的压缩,但我们观察到,随着位深度增加到8位以上,压缩增益变得更加适度。
摘要:Autoregressive "language" models (LMs) trained on raw waveforms can be repurposed for lossless audio compression, but prior work is limited to 8-bit audio, leaving open whether such approaches work for practical settings (16/24-bit) and can compete with existing codecs. We benchmark LM-based compression on full-fidelity audio across diverse domains (music, speech, bioacoustics), sampling rates (16kHz-48kHz), and bit depths (8, 16, 24-bit). Standard sample-level tokenization becomes intractable at higher bit depths due to vocabulary size (65K for 16-bit; 16.7M for 24-bit). We propose Trilobyte, a byte-level tokenization schema for full resolution audio, improving vocabulary scaling from $O(2^{b})$ to $O(1)$ and enabling the first tractable 24-bit LM-based lossless compression. While LMs consistently outperform FLAC and yield state-of-the-art compression at 8-bit and 16-bit, we observe that compression gains become more modest as bit depth increases beyond 8-bit.


【2】Scalable Neural Vocoder from Range-Null Space Decomposition
标题:来自范围-范围空间分解的可扩展神经声码器
链接:https://arxiv.org/abs/2603.08574

作者:Andong Li,Tong Lei,Zhihang Sun,Rilin Chen,Xiaodong Li,Dong Yu,Chengshi Zheng
备注:30 pages, 30 figures, 21 tables, Extension journal
摘要:尽管近年来深度神经网络促进了神经声码器的重大进展,但它们通常面临固有的挑战,如不透明的建模,不同输入配置下的不灵活再训练以及参数性能权衡。这些固有的障碍会严重阻碍这一领域的发展。为了解决这些问题,在本文中,我们提出了一种新的神经声码器在时间-频率(T-F)域。具体而言,我们桥接经典的距离零分解(RND)理论和声码器任务之间的连接,其中目标频谱图的重建被制定为距离空间和零空间之间的叠加。前者旨在将原始mel域中的表示投影到目标线性尺度域中,后者可以通过神经网络实例化以进一步填充光谱细节。为了充分利用频谱先验,设计了一个精心设计的双路径框架,其中频谱被分层编码和解码,并且交叉和窄带模块被利用来沿着子带和时间维度有效地建模。为了在各种配置下进行推理,我们提出了一种简单而有效的策略,该策略将推理阶段的多条件自适应转换为训练阶段的数据增强。在各种基准上进行了综合实验。定量和定性的结果表明,同时享有轻量级的网络结构和可扩展的推理范式,所提出的框架达到了现有的先进方法的最先进的性能。代码可在https://github.com/Andong-Li-speech/RNDVoC上获得。
摘要:Although deep neural networks have facilitated significant progress of neural vocoders in recent years, they usually suffer from intrinsic challenges like opaque modeling, inflexible retraining under different input configurations, and parameter-performance trade-off. These inherent hurdles can heavily impede the development of this field. To resolve these problems, in this paper, we propose a novel neural vocoder in the time-frequency (T-F) domain. Specifically, we bridge the connection between the classical range-null decomposition (RND) theory and the vocoder task, where the reconstruction of the target spectrogram is formulated into the superimposition between range-space and null-space. The former aims to project the representation in the original mel-domain into the target linear-scale domain, and the latter can be instantiated via neural networks to further infill the spectral details. To fully leverage the spectrum prior, an elaborate dual-path framework is devised, where the spectrum is hierarchically encoded and decoded, and the cross- and narrow-band modules are leveraged for effectively modeling along sub-band and time dimensions. To enable inference under various configurations, we propose a simple yet effective strategy, which transforms the multi-condition adaption in the inference stage into the data augmentation in the training stage. Comprehensive experiments are conducted on various benchmarks. Quantitative and qualitative results show that while enjoying lightweight network structure and scalable inference paradigm, the proposed framework achieves state-ofthe-art performance among existing advanced methods. Code is available at https://github.com/Andong-Li-speech/RNDVoC.


【3】LoopLens: Supporting Search as Creation in Loop-Based Music Composition
标题:LoopLens:在基于循环的音乐创作中支持搜索作为创作
链接:https://arxiv.org/abs/2603.08571

作者:Sheng Long,Atsuya Kobayashi,Kei Tateno
摘要:创造力支持工具(CST)通常将搜索框定为信息检索,但在电子舞曲制作等实践中,搜索作为拼贴风格创作的创造性媒介。为了解决这一差距,我们提出了LoopLens,一个基于循环的音乐创作的研究探针,可视化音频搜索结果,以支持创造性的觅食和组装。我们在一项主题内用户研究中评估了LoopLens,其中有16名不同音乐领域专业知识的参与者,他们执行开放式(发散)和目标导向(收敛)任务。我们的研究结果揭示了一个明显的行为分裂:领域专业知识的参与者利用多模态线索快速利用一组狭窄的循环,而那些没有领域知识的参与者主要依赖于音频印象,参与广泛的探索,通常受到有限的音乐词汇的限制,用于查询制定。这种行为二分法为理解创造性搜索中探索和利用之间的平衡提供了一个新的视角,并为支持未来CST中的词汇独立发现提供了明确的设计含义。
摘要:Creativity support tools (CSTs) typically frame search as information retrieval, yet in practices like electronic dance music production, search serves as a creative medium for collage-style composition. To address this gap, we present LoopLens, a research probe for loop-based music composition that visualizes audio search results to support creative foraging and assembling. We evaluated LoopLens in a within-subject user study with 16 participants of diverse musical domain expertise, performing both open-ended (divergent) and goal-directed (convergent) tasks. Our results reveal a clear behavioral split: participants with domain expertise leveraged multimodal cues to quickly exploit a narrow set of loops, while those without domain knowledge relied primarily on audio impressions, engaging in broad exploration often constrained by limited musical vocabulary for query formulation. This behavioral dichotomy provides a new lens for understanding the balance between exploration and exploitation in creative search and offers clear design implications for supporting vocabulary-independent discovery in future CSTs.


【4】Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction
标题:用于模糊情绪预测的大型音频语言模型中的理清推理
链接:https://arxiv.org/abs/2603.08230

作者:Xiaofeng Yu,Jiaheng Dong,Jean Honorio,Abhirup Ghosh,Hong Jia,Ting Dang
备注:The paper was submitted to Interspeech for review
摘要:语音情感识别在各种应用中起着重要的作用。然而,大多数现有的方法预测一个单一的情感标签,过度简化了人类情感表达的内在模糊性。最近的大型音频语言模型在产生更丰富的输出方面显示出了希望,但它们对模糊情感理解的推理能力仍然有限。在这项工作中,我们重新制定了模糊的情感识别作为一个分布式推理问题,并提出了第一个系统的研究模糊意识推理LALM。我们的框架包括两个互补的组成部分:一个模糊意识的目标,使预测与人类的感知分布,和一个结构化的模糊意识的思想链的监督,指导推理的情感线索。IEMOCAP和CREMA-D上的实验证明了SFT、DPO和GRPO训练策略的一致改进。
摘要:Speech emotion recognition plays an important role in various applications. However, most existing approaches predict a single emotion label, oversimplifying the inherently ambiguous nature of human emotional expression. Recent large audio-language models show promise in generating richer outputs, but their reasoning ability for ambiguous emotional understanding remains limited. In this work, we reformulate ambiguous emotion recognition as a distributional reasoning problem and present the first systematic study of ambiguity-aware reasoning in LALMs. Our framework comprises two complementary components: an ambiguity-aware objective that aligns predictions with human perceptual distributions, and a structured ambiguity-aware chain-of-thought supervision that guides reasoning over emotional cues. Experiments on IEMOCAP and CREMA-D demonstrate consistent improvements across SFT, DPO, and GRPO training strategies.


【5】Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models
标题:基于进化策略的语音模型低位量化校准
链接:https://arxiv.org/abs/2603.08173

作者:Lucas Rakotoarivony
备注:Submitted to INTERSPEECH 2026
摘要:量化对于语音处理系统的有效部署已经变得至关重要。虽然被广泛研究,但大多数现有的量化方法都是为视觉和NLP架构开发的,而音频信号的特定挑战在很大程度上仍然被忽视。特别是,我们表明,音频激活可以表现出很大的校准范围,导致显着的信息丢失时,标准的校准技术。为了解决这个问题,我们提出了ESC,一种基于进化策略的校准方法,将激活缩放作为一个优化问题,并使用由进化策略驱动的两步局部全局方案来解决它。ESC可在全INT 8量化下保持不变的性能,是第一种在多个语音任务中实现全INT 4量化的近无损性能的校准方法。将ESC与PTQ方法集成进一步降低了性能损失,在AST模型上实现了1%的相对精度下降。
摘要:Quantization has become essential for the efficient deployment of speech processing systems. Although widely studied, most existing quantization methods were developed for vision and NLP architectures, while the specific challenges of audio signals remain largely overlooked. In particular, we show that audio activations can exhibit large calibration ranges, leading to significant information loss when standard calibration techniques are applied. To address this, we propose ESC, an Evolution Strategy-based Calibration method that formulates activation scaling as an optimization problem and solves it using a two-step local-global scheme driven by an evolution strategy. ESC enables unaltered performance under full INT8 quantization and is the first calibration method to achieve near-lossless performance for full INT4 quantization across multiple speech tasks. Integrating ESC with PTQ methods further reduces performance loss, achieving a 1% relative accuracy degradation on the AST model.


【6】Soundscapes in Spectrograms: Pioneering Multilabel Classification for South Asian Sounds
标题:声谱图中的音景:南亚声音的开创性多标签分类
链接:https://arxiv.org/abs/2603.08154

作者:Sudip Chakrabarty,Pappu Bishwas,Rajdeep Chatterjee,Tathagata Bandyopadhyay,Digonto Biswas,Bibek Howlader
摘要:环境声音分类是城市监测和文化声景分析的一个日益重要的领域,特别是在南亚声学丰富的环境中。这些区域提出了一个独特的挑战,因为多种自然,人类和文化的声音经常重叠,使经常依赖梅尔频率倒谱系数(MFCC)的传统方法变得紧张。这项研究介绍了一种新的基于频谱的方法,具有优越的能力来捕捉这些复杂的听觉模式。卷积神经网络(CNN)架构的实现,以解决一个苛刻的多标签,多类分类问题的SAS-KIIT数据集。为了证明鲁棒性和可比性,该方法还使用著名的UrbanSound 8 K数据集进行了验证。结果证实,所提出的基于频谱图的方法显着优于现有的基于MFCC的技术,实现更高的分类精度在两个数据集。这一改进为现实世界应用中更鲁棒和准确的音频分类系统奠定了基础。
摘要:Environmental sound classification is a field of growing importance for urban monitoring and cultural soundscape analysis, especially within the acoustically rich environments of South Asia. These regions present a unique challenge as multiple natural, human, and cultural sounds often overlap, straining traditional methods that frequently rely on Mel Frequency Cepstral Coefficients (MFCC). This study introduces a novel spectrogram-based methodology with a superior ability to capture these complex auditory patterns. A Convolutional Neural Network (CNN) architecture is implemented to solve a demanding multilabel, multiclass classification problem on the SAS-KIIT dataset. To demonstrate robustness and comparability, the approach is also validated using the renowned UrbanSound8K dataset. The results confirm that the proposed spectrogram-based method significantly outperforms existing MFCC-based techniques, achieving higher classification accuracy across both datasets. This improvement lays the groundwork for more robust and accurate audio classification systems in real-world applications.


【7】Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
标题:Foley-Flow:具有掩蔽视听对齐和动态条件流的协调视频到音频生成
链接:https://arxiv.org/abs/2603.08126

作者:Shentong Mo,Yibing Song
摘要:基于视频输入的协调音频生成通常需要严格的视听(AV)对齐,其中所生成的音频片段的语义和节奏都应对应于视频帧中的那些。先前的研究利用两阶段设计,其中AV编码器首先通过对比学习对齐,然后编码的视频表示引导音频生成过程。我们观察到,对比学习和全局视频引导都可以有效地调整整体AV语义,同时限制时间节奏同步。在这项工作中,我们提出FoleyFlow首先通过掩码建模训练对齐单峰AV编码器,其中在相应视频片段的指导下恢复掩码音频片段。在训练之后,仅使用单峰数据单独预训练的AV编码器与语义和节奏一致性对齐。然后,我们开发了一个动态的条件流的最终音频生成。建立在高效的速度流生成框架,我们的动态条件流利用随时间变化的视频特征作为动态条件,以指导相应的音频段生成。为此,我们提取连贯的语义和节奏表示在掩蔽AV对齐,并使用这种表示的视频片段,以指导音频生成时间。我们的音频结果在标准基准上进行评估,并在几个指标下大大超过现有结果。优越的性能表明,FoleyFlow是有效的生成协调的音频,语义和节奏连贯的各种视频序列。
摘要:Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall correspond to those in the video frames. Previous studies leverage a two-stage design where the AV encoders are firstly aligned via contrastive learning, then the encoded video representations guide the audio generation process. We observe that both contrastive learning and global video guidance are effective in aligning overall AV semantics while limiting temporally rhythmic synchronization. In this work, we propose FoleyFlow to first align unimodal AV encoders via masked modeling training, where the masked audio segments are recovered under the guidance of the corresponding video segments. After training, the AV encoders which are separately pretrained using only unimodal data are aligned with semantic and rhythmic consistency. Then, we develop a dynamic conditional flow for the final audio generation. Built upon the efficient velocity flow generation framework, our dynamic conditional flow utilizes temporally varying video features as the dynamic condition to guide corresponding audio segment generations. To this end, we extract coherent semantic and rhythmic representations during masked AV alignment, and use this representation of video segments to guide audio generation temporally. Our audio results are evaluated on the standard benchmarks and largely surpass existing results under several metrics. The superior performance indicates that FoleyFlow is effective in generating coordinated audios that are both semantically and rhythmically coherent to various video sequences.


【8】PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment
标题:PathBench:自动病理言语评估的言语可理解性基准
链接:https://arxiv.org/abs/2603.08097

作者:Bence Mark Halpern,Thomas Tienkamp,Defne Abur,Tomoki Toda
备注:5 pages, 1 table. Submitted to Interspeech 2026
摘要:语音清晰度的自动评估是监测语音障碍和治疗效果的关键。然而,现有的方法很难比较:研究分散在私有数据集上,协议不一致。我们介绍PathBench,一个使用公共数据集进行病理语音评估的统一基准。我们比较了三种协议(匹配内容,扩展和完整)的无参考,参考文本和参考音频方法,代表语言学家(受控刺激)与机器学习专家(最大数据)如何处理相同的数据。我们在六个数据集上建立了基准基线,从而能够系统地评估未来的方法学进步,并引入了Dual-ASR关节精度(DArtP),实现了无参考方法中最高的平均相关性。
摘要:Automatic speech intelligibility assessment is crucial for monitoring speech disorders and therapy efficacy. However, existing methods are difficult to compare: research is fragmented across private datasets with inconsistent protocols. We introduce PathBench, a unified benchmark for pathological speech assessment using public datasets. We compare reference-free, reference-text, and reference-audio methods across three protocols (Matched Content, Extended, and Full) representing how a linguist (controlled stimuli) versus machine learning specialist (maximum data) would approach the same data. We establish benchmark baselines across six datasets, enabling systematic evaluation of future methodological advances, and introduce Dual-ASR Articulatory Precision (DArtP), achieving the highest average correlation among reference-free methods.


【9】WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
标题:WhispEar:一种基于伪并行耳语生成的双向分级耳语语音转换框架
链接:https://arxiv.org/abs/2603.08046

作者:Zihao Fang,Yingda Shen,Zifan Guan,Tongtong Song,Zhenyi Liu,Zhizheng Wu
备注:Submitted to Interspeech 2026
摘要:耳语语音缺乏声带振动和基频,导致声学线索退化,并使耳语到正常(W2N)转换具有挑战性,特别是在有限的并行数据。我们提出了WhispEar,一个双向框架的基础上统一的语义表示,捕捉说话模式不变的信息共享的耳语和正常的语音。该框架包含W2N和正常到耳语(N2W)模型。值得注意的是,N2W模型能够从丰富的正常语音中实现zero-shot伪并行耳语生成,从而允许用于W2N训练的可扩展数据增强。不断增加生成的数据可持续提高性能。我们还发布了迄今为止最大的双语(汉英)正常耳语平行语料库。实验表明,WhispEar优于强基线,并从可扩展的伪并行数据中受益匪浅。
摘要:Whispered speech lacks vocal fold vibration and fundamental frequency, resulting in degraded acoustic cues and making whisper-to-normal (W2N) conversion challenging, especially with limited parallel data. We propose WhispEar, a bidirectional framework based on unified semantic representations that capture speaking-mode-invariant information shared by whispered and normal speech. The framework contains both W2N and normal-to-whisper (N2W) models. Notably, the N2W model enables zero-shot pseudo-parallel whisper generation from abundant normal speech, allowing scalable data augmentation for W2N training. Increasing generated data consistently improves performance. We also release the largest bilingual (Chinese-English) whispered-normal parallel corpus to date. Experiments demonstrate that WhispEar outperforms strong baselines and benefits significantly from scalable pseudo-parallel data.


【10】Not Like Transformers: Drop the Beat Representation for Dance Generation with Mamba-Based Diffusion Model
标题:不像Transformer:放弃基于曼巴舞的扩散模型的舞蹈世代的节拍表示
链接:https://arxiv.org/abs/2603.08023

作者:Sangjune Park,Inhyeok Choi,Donghyeon Soon,Youngwoo Jeon,Kyungdon Joo
备注:Accepted by WACV 2026
摘要:舞蹈是一种以情感表达和交流为特征的人体运动形式,在音乐、虚拟现实、内容创作等多个领域发挥着作用。现有的舞蹈生成方法往往不能充分捕捉舞蹈固有的顺序,节奏和音乐同步的特点。在本文中,我们提出了一个新的舞蹈生成方法,利用基于曼巴的扩散模型的\emba {MambaDance}。Mamba非常适合处理长序列和自回归序列,它被集成到我们的两阶段扩散架构中,取代了现成的Transformer。此外,考虑到音乐节拍在舞蹈编排中的关键作用,我们提出了一个基于高斯的节拍表示,明确指导舞蹈序列的解码。在AIST++和FineDance数据集上对每个序列长度进行的实验表明,与以前的方法相比,我们提出的方法有效地生成了合理的舞蹈动作,同时反映了从短到长的舞蹈的基本特征。其他定性结果和演示视频可在\small{https://vision3d-lab.github.io/mambadance}获得。
摘要:Dance is a form of human motion characterized by emotional expression and communication, playing a role in various fields such as music, virtual reality, and content creation. Existing methods for dance generation often fail to adequately capture the inherently sequential, rhythmical, and music-synchronized characteristics of dance. In this paper, we propose \emph{MambaDance}, a new dance generation approach that leverages a Mamba-based diffusion model. Mamba, well-suited to handling long and autoregressive sequences, is integrated into our two-stage diffusion architecture, substituting off-the-shelf Transformer. Additionally, considering the critical role of musical beats in dance choreography, we propose a Gaussian-based beat representation to explicitly guide the decoding of dance sequences. Experiments on AIST++ and FineDance datasets for each sequence length show that our proposed method effectively generates plausible dance movements while reflecting essential characteristics, consistently from short to long dances, compared to the previous methods. Additional qualitative results and demo videos are available at \small{https://vision3d-lab.github.io/mambadance}.


【11】Unsupervised Domain Adaptation for Audio Deepfake Detection with Modular Statistical Transformations
标题:利用模块统计变换进行音频深度伪造检测的无监督域自适应
链接:https://arxiv.org/abs/2603.07935

作者:Urawee Thani,Gagandeep Singh,Priyanka Singh
备注:9 pages, 4 figures
摘要:在一个数据集上训练的音频deepfake检测系统在部署来自不同来源的数据时往往会失败,这是由于记录条件、合成方法和声学环境的分布变化。我们提出了一个用于无监督域自适应的模块化管道,该管道将预训练的Wav 2 Vec 2.0嵌入与统计变换相结合,以提高跨域泛化能力,而无需标记目标数据。我们的方法应用功率变换进行特征归一化,基于ANOVA的特征选择,联合PCA进行域不可知的降维,CORAL对齐在通过逻辑回归进行分类之前匹配源和目标协方差结构。我们对两种跨域传输场景进行了评估:ASVspoof 2019 LA到Fake-or-Real(FoR)和FoR到ASVspoof,实现了62.7- 63.6\ %的准确率,并且在真实和虚假类之间具有平衡的性能。系统消融实验表明,特征选择(+3.5%)和CORAL对齐(+3.2%)提供了最大的个体贡献,完整的管道将精度提高了10.7%。虽然性能与域内检测(94-96%)相比适中,但我们的管道提供了透明度和模块化,使其适合需要可解释决策的部署场景。
摘要:Audio deepfake detection systems trained on one dataset often fail when deployed on data from different sources due to distributional shifts in recording conditions, synthesis methods, and acoustic environments. We present a modular pipeline for unsupervised domain adaptation that combines pre-trained Wav2Vec 2.0 embeddings with statistical transformations to improve cross-domain generalization without requiring labeled target data. Our approach applies power transformation for feature normalization, ANOVA-based feature selection, joint PCA for domain-agnostic dimensionality reduction, and CORAL alignment to match source and target covariance structures before classification via logistic regression. We evaluate on two cross-domain transfer scenarios: ASVspoof 2019 LA to Fake-or-Real (FoR) and FoR to ASVspoof, achieving 62.7--63.6\% accuracy with balanced performance across real and fake classes. Systematic ablation experiments reveal that feature selection (+3.5%) and CORAL alignment (+3.2%) provide the largest individual contributions, with the complete pipeline improving accuracy by 10.7% over baseline. While performance is modest compared to within-domain detection (94-96%), our pipeline offers transparency and modularity, making it suitable for deployment scenarios requiring interpretable decisions.


【12】SoundWeaver: Semantic Warm-Starting for Text-to-Audio Diffusion Serving
标题:SoundWeaver:文本到音频扩散服务的语义热身
链接:https://arxiv.org/abs/2603.07865

作者:Ayush Barik,Sofia Stoica,Nikhil Sarda,Arnav Kethana,Abhinav Khanduja,Muchen Xu,Fan Lai
备注:Submitted to INTERSPEECH 2026
摘要:文本到音频扩散模型产生高保真音频,但需要数十个函数评估(NFE),导致多秒延迟和有限的吞吐量。我们提出了SoundWeaver,第一个无训练,模型不可知的服务系统,通过从语义相似的缓存音频热启动来加速文本到音频的扩散。SoundWeaver引入了三个组件:一个引用缓存器,通过语义和持续时间感知门控检索并在时间上对齐缓存的候选项;一个跳过门控器,动态确定要跳过的NFE的百分比;以及一个轻量级缓存管理器,通过质量感知驱逐和细化来维护缓存实用程序。在真实世界的音频跟踪上,SoundWeaver仅使用${\sim}$1K条目的缓存就实现了1.8- 3.0 $\times $延迟减少,同时保持或提高了感知质量。
摘要:Text-to-audio diffusion models produce high-fidelity audio but require tens of function evaluations (NFEs), incurring multi-second latency and limited throughput. We present SoundWeaver, the first training-free, model-agnostic serving system that accelerates text-to-audio diffusion by warm-starting from semantically similar cached audio. SoundWeaver introduces three components: a Reference Selector that retrieves and temporally aligns cached candidates via semantic and duration-aware gating; a Skip Gater that dynamically determines the percentage of NFEs to skip; and a lightweight Cache Manager that maintains cache utility through quality-aware eviction and refinement. On real-world audio traces, SoundWeaver achieves 1.8--3.0$ \times $ latency reduction with a cache of only ${\sim}$1K entries while preserving or improving perceptual quality.


【13】VoiceSHIELD-Small: Real-Time Malicious Speech Detection and Transcription
标题:SecureSHIELD-Small:实时恶意语音检测和转录
链接:https://arxiv.org/abs/2603.07708

作者:Sumit Ranjan,Sugandha Sharma,Ubaid Abbas,Puneeth N Ail
备注:17 pages, 9 figures
摘要:语音界面正在迅速成为人们与AI系统交互的常用方式。这也带来了新的安全风险,例如快速注入、社会工程和有害语音命令。传统的安全方法依赖于将语音转换为文本,然后过滤该文本,这会引入延迟并忽略重要的音频提示。本文介绍了VoiceSHIELD-Small,一个实时工作的轻量级模型。它可以转录语音,并检测它是安全的还是有害的,所有这些都在一个步骤中完成。VoiceSHIELD基于OpenAI的Whisper-small编码器构建,增加了一个均值池层和一个简单的分类头。在中端GPU上对音频进行分类只需90-120毫秒,而转录同时发生。在一组平衡的947个音频片段上进行测试,该模型达到了99.16%的准确率和0.9865的F1得分。在默认设置下,它错过了2.33%的有害输入。交叉验证显示性能一致(F1标准差= 0.0026)。本文还介绍了模型的设计、训练数据、性能权衡和负责任的使用指南。VoiceSHIELD是在MIT许可下发布的,以鼓励进一步研究和采用语音AI安全。
摘要:Voice interfaces are quickly becoming a common way for people to interact with AI systems. This also brings new security risks, such as prompt injection, social engineering, and harmful voice commands. Traditional security methods rely on converting speech to text and then filtering that text, which introduces delays and can ignore important audio cues. This paper introduces VoiceSHIELD-Small, a lightweight model that works in real time. It can transcribe speech and detect whether it is safe or harmful, all in one step. Built on OpenAI's Whisper-small encoder, VoiceSHIELD adds a mean-pooling layer and a simple classification head. It takes just 90-120 milliseconds to classify audio on mid-tier GPUs, while transcription happens at the same time. Tested on a balanced set of 947 audio clips, the model achieved 99.16 percent accuracy and an F1 score of 0.9865. At the default setting, it missed 2.33 percent of harmful inputs. Cross-validation showed consistent performance (F1 standard deviation = 0.0026). The paper also covers the model's design, training data, performance trade-offs, and responsible use guidelines. VoiceSHIELD is released under the MIT license to encourage further research and adoption in voice AI security.


【14】Analysis-Driven Procedural Generation of an Engine Sound Dataset with Embedded Control Annotations
标题:分析驱动的带有嵌入式控制注释的发动机声音数据集的过程生成
链接:https://arxiv.org/abs/2603.07584

作者:Robin Doerfler,Lonce Wyse
备注:Preprint. 19 hours of engine audio, 5,935 files, sample-accurate annotations. Dataset publicly available at https://doi.org/10.5281/zenodo.16883336 and https://huggingface.co/datasets/rdoerfler/procedural-engine-sounds
摘要:计算引擎声音建模是汽车音响行业的核心,特别是对于主动声音设计、虚拟原型和新兴的数据驱动引擎声音合成方法。这些应用需要大量的标准化、清晰的音频记录,并具有精确的时间对齐的操作状态注释:由于成本高、专业测量设备要求和不可避免的噪声污染,这些数据很难获得。   我们提出了一个分析驱动的框架,用于生成具有样本精确控制注释的引擎音频。该方法通过音调自适应频谱分析从实际录音中提取谐波结构,然后驱动扩展的参数谐波加噪声合成器。   通过这个框架,我们生成了程序引擎声音数据集(19小时,5,935个文件),这是一组具有样本准确的RPM和扭矩注释的引擎音频信号,涵盖了广泛的操作条件,信号复杂性和谐波特征。与真实记录的比较验证了合成数据保留了特征谐波结构,基线实验证实了其适用于基于学习的参数估计和合成任务。   该数据集公开发布,以支持发动机音色分析,控制参数估计,声学建模和神经生成网络的研究。
摘要:Computational engine sound modeling is central to the automotive audio industry, particularly for active sound design, virtual prototyping, and emerging data-driven engine sound synthesis methods. These applications require large volumes of standardized, clean audio recordings with precisely time-aligned operating-state annotations: data that is difficult to obtain due to high costs, specialized measurement equipment requirements, and inevitable noise contamination.   We present an analysis-driven framework for generating engine audio with sample-accurate control annotations. The method extracts harmonic structures from real recordings through pitch-adaptive spectral analysis, which then drive an extended parametric harmonic-plus-noise synthesizer.   With this framework, we generate the Procedural Engine Sounds Dataset (19 hours, 5,935 files), a set of engine audio signals with sample-accurate RPM and torque annotations, spanning a wide range of operating conditions, signal complexities, and harmonic profiles. Comparison against real recordings validates that the synthesized data preserves characteristic harmonic structures, and baseline experiments confirm its suitability for learning-based parameter estimation and synthesis tasks.   The dataset is released publicly to support research on engine timbre analysis, control parameter estimation, acoustic modeling and neural generative networks.


【15】Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
标题:NwâchðMun:尼泊尔Bhasha ASB的Devanagari言语库和近端转会基准
链接:https://arxiv.org/abs/2603.07554

作者:Rishikesh Kumar Sharma,Safal Narshing Shrestha,Jenny Poudel,Rupak Tiwari,Arju Shrestha,Rupak Raj Ghimire,Bal Krishna Bal
摘要:尼泊尔Bhasha(Newari)是加德满都谷地的一种濒危语言,由于注释语音资源严重匮乏,仍然处于数字边缘化状态。在这项工作中,我们介绍了NwāchMunnak,一个新策划的5.39小时手动转录的梵文语音语料库尼泊尔Bhasha,并建立了第一个基准使用脚本保留声学建模。我们研究了从地理和语言上相邻的语言(尼泊尔语)的近端跨语言迁移是否可以在超低资源自动语音识别(ASR)设置中与大规模多语言预训练相媲美。微调尼泊尔语一致性模型将字符错误率(CER)从52.54% zero-shot基线降低到17.59%,并进行数据增强,有效地匹配多语言Whisper-Small模型的性能,尽管使用的参数显著减少。我们的研究结果表明,南亚语言集群内的近端迁移是大规模多语言模型的一种计算效率高的替代方案。我们公开发布数据集和基准,以数字化方式支持Newari社区,并促进尼泊尔Bhasha的进一步研究。
摘要:Nepal Bhasha (Newari), an endangered language of the Kathmandu Valley, remains digitally marginalized due to the severe scarcity of annotated speech resources. In this work, we introduce Nwāchā Munā, a newly curated 5.39-hour manually transcribed Devanagari speech corpus for Nepal Bhasha, and establish the first benchmark using script-preserving acoustic modeling. We investigate whether proximal cross-lingual transfer from a geographically and linguistically adjacent language (Nepali) can rival large-scale multilingual pretraining in an ultra-low-resource Automatic Speech Recognition (ASR) setting. Fine-tuning a Nepali Conformer model reduces the Character Error Rate (CER) from a 52.54% zero-shot baseline to 17.59% with data augmentation, effectively matching the performance of the multilingual Whisper-Small model despite utilizing significantly fewer parameters. Our findings demonstrate that proximal transfer within South Asian language clusters serves as a computationally efficient alternative to massive multilingual models. We openly release the dataset and benchmarks to digitally enable the Newari community and foster further research in Nepal Bhasha.


【16】Targeted Speaker Poisoning Framework in Zero-Shot Text-to-Speech
标题:Zero-Shot文本转语音中的目标说话者中毒框架
链接:https://arxiv.org/abs/2603.07551

作者:Thanapat Trachu,Thanathai Lertpetchpun,Sai Praneeth Karimireddy,Shrikanth Narayanan
备注:Submitted to Interspeech2026
摘要:Zero-shot文本到语音(TTS)语音克隆带来了严重的隐私风险,需要从训练的TTS模型中删除特定的说话者身份。在这种情况下,传统的机器非学习是不够的,因为zero-shot TTS可以仅从参考提示动态地重建语音。我们将此任务形式化为语音生成扬声器中毒(SGSP),其中我们修改训练模型以防止生成特定身份,同时保留其他扬声器的实用性。我们评估了1,15和100个被遗忘的说话者的推理时间过滤和参数修改基线。性能通过效用(WER)和隐私之间的权衡进行评估,使用AUC和忘记说话人相似性(FSSIM)进行量化。我们实现了强大的隐私多达15个扬声器,但显示可扩展性限制在100个扬声器,由于身份重叠增加。因此,我们的研究引入了一个新的问题和评估框架,以进一步促进生成语音隐私。
摘要:Zero-shot Text-to-Speech (TTS) voice cloning poses severe privacy risks, demanding the removal of specific speaker identities from trained TTS models. Conventional machine unlearning is insufficient in this context, as zero-shot TTS can dynamically reconstruct voices from just reference prompts. We formalize this task as Speech Generation Speaker Poisoning (SGSP), in which we modify trained models to prevent the generation of specific identities while preserving utility for other speakers. We evaluate inference-time filtering and parameter-modification baselines across 1, 15, and 100 forgotten speakers. Performance is assessed through the trade-off between utility (WER) and privacy, quantified using AUC and Forget Speaker Similarity (FSSIM). We achieve strong privacy for up to 15 speakers but reveal scalability limits at 100 speakers due to increased identity overlap. Our study thus introduces a novel problem and evaluation framework toward further advances in generative voice privacy.


【17】Evaluating Parkinson's Disease Detection in Anonymized Speech: A Performance and Acoustic Analysis
标题:评估语音中帕金森病检测:性能和声学分析
链接:https://arxiv.org/abs/2603.07544

作者:Carlos Franzreb,Francisco Teixeira,Ben Luks,Sebastian Möller,Alberto Abad
备注:Submitted to Interspeech 2026
摘要:从语音中自动检测帕金森病(PD)是一种很有前途的非侵入性诊断工具,但它引起了严重的隐私问题。说话者匿名化减轻了这些风险,但它可能会抑制PD检测所需的病理信息。我们使用两个西班牙语数据集评估两个匿名者(STT-TTS和kNN-VC)的隐私和PD检测之间的权衡。STT-TTS提供了更好的隐私,但通过消除韵律信息严重降低了PD检测。kNN-VC保留了宏韵律特征,如持续时间和F0轮廓,实现F1分数仅比原始基线低3- 7%,表明在使用适当的匿名化时,隐私保护PD检测是可行的。最后,声学失真分析表征了kNN-VC中的特定弱点,为设计更好地保留PD信息的匿名器提供了见解。
摘要:Automatic detection of Parkinson's disease (PD) from speech is a promising non-invasive diagnostic tool, but it raises significant privacy concerns. Speaker anonymization mitigates these risks, but it may suppress the pathological information necessary for PD detection. We assess the trade-off between privacy and PD detection for two anonymizers (STT-TTS and kNN-VC) using two Spanish datasets. STT-TTS provides better privacy but severely degrades PD detection by eradicating prosodic information. kNN-VC preserves macro-prosodic features such as duration and F0 contours, achieving F1 scores only 3-7\% lower than original baselines, demonstrating that privacy-preserving PD detection is viable when using appropriate anonymization. Finally, an acoustic distortion analysis characterizes specific weaknesses in kNN-VC, offering insights for designing anonymizers that better preserve PD information.


【18】Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
标题:查看上下文:通过多模式推理的丰富视觉上下文感知语音识别
链接:https://arxiv.org/abs/2603.07263

作者:Wenjie Tian,Mingchen Shao,Bingshen Mu,Xuelong Geng,Chengyou Wang,Yujie Liao,Zhixian Zhao,Ziyu Zhang,Jingbin Hu,Mengqi Wei,Lei Xie
摘要:视听语音识别(AVSR)是结合视觉信号的ASR的扩展。目前的AVSR方法主要集中在嘴唇运动,在很大程度上忽略了丰富的背景下,目前在视频中,如说话场景和屏幕上的文字。为了解决这样的CAVSR(AVSR包括丰富的视觉上下文),我们提出了VASR,旨在“看到”和原因的视觉上下文,以提高语音识别。具体来说,我们构建了一个视听思想链(AV-CoT),明确执行声学信号和视觉证据之间的中间跨模态接地。这种证据驱动的推理缓解了“单一模态优势”问题,即模型过度依赖视觉上下文或无法利用视觉上下文。此外,为了解决数据稀缺问题,我们构建并发布了相应的数据管道和测试集。实验表明,AV-CoT有效地缓解了单模态优势,在CAVSR中实现了最先进的性能。该项目是开源的。
摘要:Audio-visual speech recognition (AVSR) is an extension of ASR that incorporates visual signals. Current AVSR approaches primarily focus on lip motion, largely overlooking rich context present in the video such as speaking scene and on-screen text. To tackle such CAVSR (AVSR including rich visual Context), we propose VASR designed to "see" and reason the visual context to improve speech recognition. Specifically, we construct an Audio-Visual Chain-of-Thought (AV-CoT) that explicitly enforces intermediate cross-modal grounding between acoustic signals and visual evidence. This evidence-driven reasoning mitigates the "single-modality dominance" problem, where models either over-rely on visual context or fail to utilize it. Besides, to address the data scarcity, we construct and release a corresponding data pipeline and test set. Experiments show that AV-CoT effectively mitigates the single-modality dominance, achieving state-of-the-art performance in CAVSR. The project is open-sourced.


【19】Towards Objective Gastrointestinal Auscultation: Automated Segmentation and Annotation of Bowel Sound Patterns
标题:走向客观的胃肠道听诊:肠道声音模式的自动分割和注释
链接:https://arxiv.org/abs/2603.07215

作者:Zahra Mansour,Verena Uslar,Dirk Weyhe,Danilo Hollosi,Nils Strodthoff
摘要:肠鸣音(BS)通常是瞬时的并且具有低振幅,使得它们难以通过手动听诊准确地检测。这导致临床评估的显著差异。数字声学传感器允许采集高质量的BS并实现自动信号分析,从而有可能为临床医生提供有关肠道活动的客观和定量反馈。这项研究提出了一个自动管道肠声音分割和分类使用可穿戴声学SonicGuard传感器。使用SonicGuard传感器记录来自83名受试者的BS信号。来自40名受试者的数据由临床专家手动注释并用于训练自动注释算法,而其余受试者用于进一步的模型评估。提出了一种基于能量的事件检测算法。然后使用预训练的音频频谱图Transformer(AST)模型将检测到的声音片段分类为BS模式。分别评价健康个体和患者的模型性能。最佳配置使用两个专门的模型,一个在健康受试者上训练,一个在患者上训练,健康组和患者组分别达到(准确度:0.97,AUROC:0.98)和(准确度:0.96,AUROC:0.98)。自动注释方法将手动标记时间减少了约70%,专家评审表明,只有不到12%的自动检测片段需要校正。所提出的自动分割和分类系统能够定量评估肠道活动,为临床医生提供客观的诊断工具,可以改善胃肠道功能的诊断,并支持大规模数据集的注释。
摘要:Bowel sounds (BS) are typically momentary and have low amplitude, making them difficult to detect accurately through manual auscultation. This leads to significant variability in clinical assessment. Digital acoustic sensors allow the acquisition of high-quality BS and enable automated signal analysis, offering the potential to provide clinicians with both objective and quantitative feedback on bowel activity. This study presents an automated pipeline for bowel sound segmentation and classification using a wearable acoustic SonicGuard sensor. BS signals from 83 subjects were recorded using a SonicGuard sensor. Data from 40 subjects were manually annotated by clinical experts and used to train an automatic annotation algorithm, while the remaining subjects were used for further model evaluation. An energy-based event detection algorithm was developed to detect BS events. Detected sound segments were then classified into BS patterns using a pretrained Audio Spectrogram Transformer (AST) model. Model performance was evaluated separately for healthy individuals and patients. The best configuration used two specialized models, one trained on healthy subjects and one on patients, achieving (accuracy: 0.97, AUROC: 0.98) for healthy group and (accuracy: 0.96, AUROC: 0.98) for patient group. The auto-annotation method reduced manual labeling time by approximately 70%, and expert review showed that less than 12% of automatically detected segments required correction. The proposed automated segmentation and classification system enables quantitative assessment of bowel activity, providing clinicians with an objective diagnostic tool that may improve the diagnostic of gastrointestinal function and support the annotation of large-scale datasets.


【20】Toward Multimodal Industrial Fault Analysis: A Single-Speed Chain Conveyor Dataset with Audio and Vibration Signals
标题:走向多模式工业故障分析:具有音频和振动信号的单速链条输送机数据集
链接:https://arxiv.org/abs/2603.07130

作者:Zhang Chen,Yucong Zhang,Xiaoxiao Miao,Ming Li
备注:Submitted to Interspeech 2026
摘要:我们介绍了从单速链式输送机(SSCC)系统收集的多模态工业故障分析数据集,针对生产线中的系统级故障检测。该数据集由多模态信号组成,包括三个音频通道和四个振动通道。它涵盖了在多种速度、负载以及现场再现的干净和真实的工厂噪声条件下的正常运行和四种代表性故障类型。它被明确设计为支持通道分析和多模态融合研究。我们建立了标准化的评估协议,用于无监督的故障检测,仅进行正常训练,并在不同的操作条件和故障类型之间进行平衡的数据集分割,进行有监督的故障分类。提供了统一的通道kNN基线,以便在没有特定任务训练的情况下公平比较表示质量。该数据集为强大的多模态工业故障分析提供了一个实用和可扩展的基准。
摘要:We introduce a multimodal industrial fault analysis dataset collected from a single-speed chain conveyor (SSCC) system, targeting system-level fault detection in production lines. The dataset consists of multimodal signals, including three audio and four vibration channels. It covers normal operation and four representative fault types under multiple speeds, loads, and both clean and realistic factory-noise conditions reproduced on-site. It is explicitly designed to support channel-wise analysis and multimodal fusion research. We establish standardized evaluation protocols for unsupervised fault detection with normal-only training and supervised fault classification with balanced dataset splits across different operating conditions and fault types. A unified channel-wise kNN baseline is provided to enable fair comparison of representation quality without task-specific training. The dataset offers a practical and extensible benchmark for robust multimodal industrial fault analysis.


【21】Adaptive Discovery of Interpretable Audio Attributes with Multimodal LLMs for Low-Resource Classification
标题:利用多模式LLM自适应发现可解释音频属性以实现低资源分类
链接:https://arxiv.org/abs/2603.06991

作者:Kosuke Yoshimura,Hisashi Kashima
备注:5 pages, 1 figure
摘要:在低资源音频分类的预测建模中,提取高精度和可解释的属性是至关重要的。特别是在高可靠性应用中,可解释的音频属性是必不可少的。虽然人类驱动的属性发现是有效的,但其低吞吐量成为瓶颈。我们提出了一种自适应发现可解释的音频属性使用多模态大语言模型(MLLM)的方法。通过用MLLM替换AdaFlock框架中的人类,我们的方法实现了更快的属性发现。我们的方法动态识别显着的声学特征,通过提示和构造一个基于属性的集成分类器。在各种音频任务的实验结果表明,我们的方法优于直接MLLM预测在大多数评估的情况下。整个训练在11分钟内完成,证明它是一种实用的自适应解决方案,超越了传统的依赖人类的方法。
摘要:In predictive modeling for low-resource audio classification, extracting high-accuracy and interpretable attributes is critical. Particularly in high-reliability applications, interpretable audio attributes are indispensable. While human-driven attribute discovery is effective, its low throughput becomes a bottleneck. We propose a method for adaptively discovering interpretable audio attributes using Multimodal Large Language Models (MLLMs). By replacing humans in the AdaFlock framework with MLLMs, our method achieves significantly faster attribute discovery. Our method dynamically identifies salient acoustic characteristics via prompting and constructs an attribute-based ensemble classifier. Experimental results across various audio tasks demonstrate that our method outperforms direct MLLM prediction in the majority of evaluated cases. The entire training completes within 11 minutes, proving it a practical, adaptive solution that surpasses conventional human-reliant approaches.


【22】Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio Steering
标题:音频语言模型在听吗?自适应音频转向音频专家负责人
链接:https://arxiv.org/abs/2603.06854

作者:Neta Glazer,Lenny Aharon,Ethan Fetaya
摘要:多模态大型语言模型可以表现出文本优势,过度依赖语言先验,而不是在非文本输入中进行基础预测。一个例子是大型音频语言模型(LALM),即使它包含重要信息,也可能未充分利用决定性的音频证据。为了解决这个问题,我们使用机械的可解释性,以确定一个小的音频专家注意头的音频注意力产生一个“成功”的信号。我们发现,当音频证据影响模型的输出时,此信号会增加,从而提供了标准提示下音频参与的指标。利用这个本地化,我们构建了一个音频-静音转向方向,并将推理时间激活干预应用于最终表示,放大模型的音频效果。为了证明这种干预的实用性,我们在MMAU上显示,这在两个基于Qwen的LALM上将准确性提高了+8.0个百分点,而无需任何参数更新。
摘要:Multimodal large language models can exhibit text dominance, over-relying on linguistic priors instead of grounding predictions in non-text inputs. One example is large audio-language models (LALMs) where decisive audio evidence can be under-utilized even when it contains important information. To address this issue we use mechanistic interpretability to identify a small set of audio-specialist attention heads whose audio attention yields a ``listening'' signal. We show that this signal increases when audio evidence affects the model's output, providing an indicator of audio engagement under standard prompting. Leveraging this localization, we construct an audio--silence steering direction and apply an inference-time activation intervention to the final representation, amplifying the model's audio effect. To demonstrate the utility of this intervention, we show on MMAU that this improves accuracy by up to +8.0 percentage points on two Qwen-based LALMs, without any parameter updates.


【23】DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
标题:DualTurn:从双通道生成语音预训练中学习回合转换
链接:https://arxiv.org/abs/2603.08216

作者:Shangeth Rajaa
备注:Submitted to Interspeech 2026
摘要:语音到语音模型自然地处理话轮转换,但对工具调用或复杂推理的支持有限,而生产ASR-LLM-TTS语音管道提供了这些功能,但依赖于沉默超时,这会导致不自然的话轮转换。我们提出了DualTurn,它通过对双通道会话音频进行生成预训练来缩小这一差距。该模型生成两个扬声器的未来音频自回归,隐式学习会话动态没有任何标签,然后微调预测可解释的话轮转换信号,直接映射到代理动作。DualTurn持续监控两个通道,预测转弯边界并产生五个代理动作。在标准基准测试中,DualTurn(0.5B)在代理动作预测(wF 1 0.633 vs. 0.389)和单词级转向预测(AUC 0.930 vs. 0.880)上优于VAP和3.1B音频文本模型,同时更早地预测转向边界,中断更少。
摘要:Speech-to-speech models handle turn-taking naturally but offer limited support for tool-calling or complex reasoning, while production ASR-LLM-TTS voice pipelines offer these capabilities but rely on silence timeouts, which lead to unnatural turn-taking. We present DualTurn, which narrows this gap through generative pretraining on dual-channel conversational audio. The model generates both speakers' future audio autoregressively, implicitly learning conversational dynamics without any labels, and is then fine-tuned to predict interpretable turn-taking signals that map directly to agent actions. DualTurn monitors both channels continuously, anticipating turn boundaries and producing five agent actions. On standard benchmarks, DualTurn (0.5B) outperforms both VAP on agent action prediction (wF1 0.633 vs. 0.389) and a 3.1B audio-text model on word-level turn prediction (AUC 0.930 vs. 0.880), while anticipating turn boundaries earlier with fewer interruptions.


【24】Towards Lightweight Adaptation of Speech Enhancement Models in Real-World Environments
标题:实现现实环境中语音增强模型的轻量级适应
链接:https://arxiv.org/abs/2603.07471

作者:Longbiao Cheng,Shih-Chii Liu
备注:Accepted to ICASSP 2026
摘要:最近的研究表明,部署后自适应可以提高语音增强模型在不可见噪声条件下的鲁棒性。然而,现有的方法通常会导致高昂的计算和存储成本,限制了它们对设备上部署的适用性。在这项工作中,我们研究了具有动态声学场景变化的现实环境中的模型自适应,并提出了一个轻量级框架,该框架通过自监督训练更新低等级适配器来增强冻结的骨干。在37种噪声类型和三个信噪比范围(包括具有挑战性的[-8,0] dB范围)的111个环境中进行的连续场景评估实验表明,我们的方法更新的基本模型参数不到1%,同时每个场景仅在20次更新内实现了平均1.51 dB的SI-SDR改进。与最先进的方法相比,我们的框架实现了具有竞争力或卓越的感知质量,具有更平滑和更稳定的收敛,证明了其在现实世界声学条件下对语音增强模型的轻量级设备自适应的实用性。
摘要:Recent studies have shown that post-deployment adaptation can improve the robustness of speech enhancement models in unseen noise conditions. However, existing methods often incur prohibitive computational and memory costs, limiting their suitability for on-device deployment. In this work, we investigate model adaptation in realistic settings with dynamic acoustic scene changes and propose a lightweight framework that augments a frozen backbone with low-rank adapters updated via self-supervised training. Experiments on sequential scene evaluations spanning 111 environments across 37 noise types and three signal-to-noise ratio ranges, including the challenging [-8, 0] dB range, show that our method updates fewer than 1% of the base model's parameters while achieving an average 1.51 dB SI-SDR improvement within only 20 updates per scene. Compared to state-of-the-art approaches, our framework achieves competitive or superior perceptual quality with smoother and more stable convergence, demonstrating its practicality for lightweight on-device adaptation of speech enhancement models under real-world acoustic conditions.


【25】Fast and Flexible Audio Bandwidth Extension via Vocos
标题:通过Vocos快速灵活的音频带宽扩展
链接:https://arxiv.org/abs/2603.07285

作者:Yatharth Sharma
备注:5 pages, 2 figures, 5 tables. Submitted to INTERSPEECH 2026. Code available at https://github.com/ysharma3501/LavaSR.git
摘要:我们提出了一个基于Vocos的带宽扩展模型,通过生成丢失的高频内容来增强8-48 kHz的音频。输入被重新采样到48 kHz,并由神经声码器骨干处理,使单个网络能够支持任意上采样率。一个轻量级的Linkwitz-Riley启发细化合并原始低频带与生成的高频通过平滑的交叉。经过验证,该模型在NVIDIA A100 GPU上以0.0001的实时系数和在8核CPU上以0.0053的实时系数运行时,实现了具有竞争力的对数谱距离,证明了在极端吞吐量下的实用、高质量斗轮挖掘机。
摘要:We propose a Vocos-based bandwidth extension model that enhances audio at 8-48 kHz by generating missing high-frequency content. Inputs are resampled to 48 kHz and processed by a neural vocoder backbone, enabling a single network to support arbitrary upsampling ratios. A lightweight Linkwitz-Riley-inspired refiner merges the original low band with the generated high frequencies via a smooth crossover. On validation, the model achieves competitive log-spectral distance while running at a real-time factor of 0.0001 on an NVIDIA A100 GPU and 0.0053 on an 8-core CPU, demonstrating practical, high-quality BWE at extreme throughput.


eess.AS音频处理


【1】NLE: Non-autoregressive LLM-based ASR by Transcript Editing
标题:NLE:通过脚本编辑的基于LLM的非自回归的ASB
链接:https://arxiv.org/abs/2603.08397

作者:Avihu Dekel,Samuel Thomas,Takashi Fukada,George Saon
备注:Preprint
摘要:虽然基于自回归(AR)LLM的ASR系统实现了很高的准确性,但它们的顺序解码限制了并行性并导致高延迟。我们提出了NLE,一种非自回归(NAR)方法,将语音识别制定为有条件的转录编辑,从而实现完全并行的预测。NLE从预先训练的语音编码器中提取声学嵌入和初始假设,然后使用用潜在对齐目标训练的双向LLM编辑器来细化假设。交错填充策略利用Transformers的身份映射偏差,允许模型专注于校正而不是完全重建。在Open ASR排行榜上,NLE++的平均WER为5.67%,RTFx(逆实时因子)为1630。在单话语场景中,NLE在AR基线上实现了27倍的加速,使其适用于实时应用。
摘要:While autoregressive (AR) LLM-based ASR systems achieve strong accuracy, their sequential decoding limits parallelism and incurs high latency. We propose NLE, a non-autoregressive (NAR) approach that formulates speech recognition as conditional transcript editing, enabling fully parallel prediction. NLE extracts acoustic embeddings and an initial hypothesis from a pretrained speech encoder, then refines the hypothesis using a bidirectional LLM editor trained with a latent alignment objective. An interleaved padding strategy exploits the identity mapping bias of Transformers, allowing the model to focus on corrections rather than full reconstruction. On the Open ASR leaderboard, NLE++ achieves 5.67% average WER with an RTFx (inverse real-time factor) of 1630. In single-utterance scenarios, NLE achieves 27x speedup over the AR baseline, making it suitable for real-time applications.


【2】Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios with Synthetic Visual Data
标题:利用合成视觉数据在零AV资源场景中引导视听语音识别
链接:https://arxiv.org/abs/2603.08249

作者:Pol Buitrago,Pol Gàlvez,Oriol Pareras,Javier Hernando
备注:6 pages, 3 figures, Submitted to Interspeech 2026
摘要:视听语音识别(AVSR)结合了声学和视觉线索,以提高在具有挑战性的条件下的转录鲁棒性,但由于缺乏用于训练的标记视频语料库,对于大多数资源不足的语言来说仍然遥不可及。我们提出了一个零AV资源的AVSR框架,依赖于合成的视觉流与真实的音频唇同步静态面部图像产生的。   我们首先在西班牙基准上评估合成视觉增强,然后将其应用于加泰罗尼亚语,一种没有注释的视听语料库的语言。我们合成了超过700小时的头部视频,并对预训练的AV-HuBERT模型进行了微调。在手动注释的加泰罗尼亚基准测试中,我们的模型以更少的参数和训练数据实现了接近最先进的性能,优于相同训练的仅音频基线,并保留了噪声中的多模态优势。因此,可伸缩的合成视频提供了一个可行的替代真实记录在零AV资源AVSR。
摘要:Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but remains out of reach for most under-resourced languages due to the lack of labeled video corpora for training. We propose a zero-AV-resource AVSR framework that relies on synthetic visual streams generated by lip-syncing static facial images with real audio.   We first evaluate synthetic visual augmentation on Spanish benchmarks, then apply it to Catalan, a language with no annotated audiovisual corpora. We synthesize over 700 hours of talking-head video and fine-tune a pre-trained AV-HuBERT model. On a manually annotated Catalan benchmark, our model achieves near state-of-the-art performance with much fewer parameters and training data, outperforms an identically trained audio-only baseline, and preserves multimodal advantages in noise. Scalable synthetic video thus offers a viable substitute for real recordings in zero-AV-resource AVSR.


【3】Quantifying Cross-Lingual Transfer in Paralinguistic Speech Tasks
标题:量化副语言言语任务中的跨语言迁移
链接:https://arxiv.org/abs/2603.08231

作者:Pol Buitrago,Oriol Pareras,Federico Costa,Javier Hernando
备注:6 pages, 5 figures, Submitted to Interspeech 2026
摘要:副语言言语任务通常被认为是相对语言不可知的,因为它们依赖于语言外的声学线索,而不是词汇内容。然而,之前的研究报告了跨语言条件下的性能下降,表明语言依赖性不容忽视。尽管如此,这些研究通常集中在孤立的语言对或特定的任务设置,限制了可比性,并防止任务层面的语言依赖的系统评估。   我们介绍了跨语言传递矩阵(CLTM),一个系统的方法来量化给定任务中的语言对之间的跨语言的相互作用。我们将CLTM应用于两个语言学任务,性别识别和说话人验证,使用多语言的基于休伯特的编码器,分析捐助者的语言数据如何影响目标语言的性能在微调。我们的研究结果揭示了不同的任务和语言的迁移模式,反映了系统的,语言依赖的影响。
摘要:Paralinguistic speech tasks are often considered relatively language-agnostic, as they rely on extralinguistic acoustic cues rather than lexical content. However, prior studies report performance degradation under cross-lingual conditions, indicating non-negligible language dependence. Still, these studies typically focus on isolated language pairs or task-specific settings, limiting comparability and preventing a systematic assessment of task-level language dependence.   We introduce the Cross-Lingual Transfer Matrix (CLTM), a systematic method to quantify cross-lingual interactions between pairs of languages within a given task. We apply the CLTM to two paralinguistic tasks, gender identification and speaker verification, using a multilingual HuBERT-based encoder, to analyze how donor-language data affects target-language performance during fine-tuning. Our results reveal distinct transfer patterns across tasks and languages, reflecting systematic, language-dependent effects.


【4】DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
标题:DualTurn:从双通道生成语音预训练中学习回合转换
链接:https://arxiv.org/abs/2603.08216

作者:Shangeth Rajaa
备注:Submitted to Interspeech 2026
摘要:语音到语音模型自然地处理话轮转换,但对工具调用或复杂推理的支持有限,而生产ASR-LLM-TTS语音管道提供了这些功能,但依赖于沉默超时,这会导致不自然的话轮转换。我们提出了DualTurn,它通过对双通道会话音频进行生成预训练来缩小这一差距。该模型生成两个扬声器的未来音频自回归,隐式学习会话动态没有任何标签,然后微调预测可解释的话轮转换信号,直接映射到代理动作。DualTurn持续监控两个通道,预测转弯边界并产生五个代理动作。在标准基准测试中,DualTurn(0.5B)在代理动作预测(wF 1 0.633 vs. 0.389)和单词级转向预测(AUC 0.930 vs. 0.880)上优于VAP和3.1B音频文本模型,同时更早地预测转向边界,中断更少。
摘要:Speech-to-speech models handle turn-taking naturally but offer limited support for tool-calling or complex reasoning, while production ASR-LLM-TTS voice pipelines offer these capabilities but rely on silence timeouts, which lead to unnatural turn-taking. We present DualTurn, which narrows this gap through generative pretraining on dual-channel conversational audio. The model generates both speakers' future audio autoregressively, implicitly learning conversational dynamics without any labels, and is then fine-tuned to predict interpretable turn-taking signals that map directly to agent actions. DualTurn monitors both channels continuously, anticipating turn boundaries and producing five agent actions. On standard benchmarks, DualTurn (0.5B) outperforms both VAP on agent action prediction (wF1 0.633 vs. 0.389) and a 3.1B audio-text model on word-level turn prediction (AUC 0.930 vs. 0.880), while anticipating turn boundaries earlier with fewer interruptions.


【5】Privacy-Preserving End-to-End Full-Duplex Speech Dialogue Models
标题:保护隐私的端到端全复式语音对话模型
链接:https://arxiv.org/abs/2603.08179

作者:Nikita Kuzmin,Tao Zhong,Jiajun Deng,Yingke Zhu,Tristan Tsoi,Tianxiang Cao,Simon Lui,Kong Aik Lee,Eng Siong Chng
摘要:端到端全双工语音模型通过始终在线的LLM主干向用户提供音频,但其隐藏表示的扬声器隐私影响仍未得到研究。在VoicePrivacy 2024协议与懒惰通知攻击者之后,我们表明SALM双工和Moshi的隐藏状态在所有Transformer层中泄漏了大量的扬声器身份。逐层和逐圈分析表明,泄漏持续存在于所有层中,SALM-Duplex在早期层中显示出更强的泄漏,而Moshi均匀地泄漏,并且可连接性在前几圈内急剧上升。我们提出了两个流匿名设置使用Stream-Voice-Anon:波形级前端(Anon-W2 W)和特征域替换(Anon-W2 F)。Anon-W2 F相对于离散编码器基线(11.2%至41.0%)将EER提高了3.5倍以上,接近50%的随机概率上限,而Anon-W2 W在亚秒级响应延迟(FRL低于0.8 s)的设置中保持了基线sBERT的78-93%。
摘要:End-to-end full-duplex speech models feed user audio through an always-on LLM backbone, yet the speaker privacy implications of their hidden representations remain unexamined. Following the VoicePrivacy 2024 protocol with a lazy-informed attacker, we show that the hidden states of SALM-Duplex and Moshi leak substantial speaker identity across all transformer layers. Layer-wise and turn-wise analyses reveal that leakage persists across all layers, with SALM-Duplex showing stronger leakage in early layers while Moshi leaks uniformly, and that Linkability rises sharply within the first few turns. We propose two streaming anonymization setups using Stream-Voice-Anon: a waveform-level front-end (Anon-W2W) and a feature-domain replacement (Anon-W2F). Anon-W2F raises EER by over 3.5x relative to the discrete encoder baseline (11.2% to 41.0%), approaching the 50% random-chance ceiling, while Anon-W2W retains 78-93% of baseline sBERT across setups with sub-second response latency (FRL under 0.8 s).


【6】Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge
标题:TidyVoice 2026挑战赛的默认不变多语言说话者验证
链接:https://arxiv.org/abs/2603.08092

作者:Ze Li,Xiaoxiao Miao,Juan Liu,Ming Li
备注:submitted to Interspeech 2026
摘要:由于说话人嵌入中的跨语言数据和语言相关信息有限,多语言说话人确认(SV)仍然具有挑战性。本文提出了一种语言不变的多语言SV系统,用于TidyVoice 2026挑战赛。我们采用多语言自监督w2 v-BERT 2.0模型作为骨干,通过层适配器和多尺度特征聚合进行增强,以更好地利用多层表示。一种带有梯度反射层的语言对抗训练策略被应用于促进语言不变的说话人嵌入。此外,使用多语言zero-shot文本到语音系统来合成多种语言的语音,提高语言多样性。实验结果表明,微调大规模的预训练模型产生有竞争力的性能,而语言对抗训练进一步增强了鲁棒性。此外,合成语音增强在有限的训练数据条件下提供额外的增益。源代码可在https://github.com/ZXHY-82/LI-MSV-TidyVoice2026上获得。
摘要:Multilingual speaker verification (SV) remains challenging due to limited cross-lingual data and language-dependent information in speaker embeddings. This paper presents a language-invariant multilingual SV system for the TidyVoice 2026 Challenge. We adopt the multilingual self-supervised w2v-BERT 2.0 model as the backbone, enhanced with Layer Adapters and Multi-scale Feature Aggregation to better exploit multi-layer representations. A language-adversarial training strategy with a Gradient Reversal Layer is applied to promote language-invariant speaker embeddings. Moreover, a multilingual zero-shot text-to-speech system is used to synthesize speech in multiple languages, improving language diversity. Experimental results demonstrate that fine-tuning the large-scale pretrained model yields competitive performance, while language-adversarial training further enhances robustness. In addition, synthetic speech augmentation provides additional gains under limited training data conditions. Source code is available at https://github.com/ZXHY-82/LI-MSV-TidyVoice2026.


【7】Multi-View Based Audio Visual Target Speaker Extraction
标题:基于多视图的视听目标说话人提取
链接:https://arxiv.org/abs/2603.07696

作者:Peijun Yang,Zhan Jin,Juan Liu,Ming Li
备注:submitted to INTERSPEECH
摘要:视听目标说话人提取(AVTSE)的目的是使用相应的视觉线索从混合音频信号中分离出目标说话人的语音。虽然大多数现有的AVTSE方法完全依赖于正面视图视频,但这种限制限制了它们在非正面视图普遍存在的真实世界场景中的鲁棒性。这样的视觉视角往往包含补充发音信息,可以提高语音提取。在这项工作中,我们提出了多视图张量融合(MVTF),这是一种将多视图学习转换为单视图性能增益的新框架。在训练阶段,我们利用同步的多视角嘴唇视频来通过MVTF学习跨视图相关性,其中成对的外积显式地对输入嘴唇嵌入的不同视图之间的乘法交互进行建模。在推理阶段,系统支持单视图和多视图输入。实验结果表明,在单视图输入下,该框架利用多视图知识实现了显著的性能提升,而在多视图模式下,该框架进一步提高了整体性能,增强了鲁棒性。我们的演示、代码和数据可在https://anonymous.4open.science/w/MVTF-Gridnet-209C/上获得
摘要:Audio-Visual Target Speaker Extraction (AVTSE) aims to separate a target speaker's voice from a mixed audio signal using the corresponding visual cues. While most existing AVTSE methods rely exclusively on frontal-view videos, this limitation restricts their robustness in real-world scenarios where non-frontal views are prevalent. Such visual perspectives often contain complementary articulatory information that could enhance speech extraction. In this work, we propose Multi-View Tensor Fusion (MVTF), a novel framework that transforms multi-view learning into single-view performance gains. During the training stage, we leverage synchronized multi-perspective lip videos to learn cross-view correlations through MVTF, where pairwise outer products explicitly model multiplicative interactions between different views of input lip embeddings. At the inference stage, the system supports both single-view and multi-view inputs. Experimental results show that in the single-view inputs, our framework leverages multi-view knowledge to achieve significant performance gains, while in the multi-view mode, it further improves overall performance and enhances the robustness. Our demo, code and data are available at https://anonymous.4open.science/w/MVTF-Gridnet-209C/


【8】Towards Lightweight Adaptation of Speech Enhancement Models in Real-World Environments
标题:实现现实环境中语音增强模型的轻量级适应
链接:https://arxiv.org/abs/2603.07471

作者:Longbiao Cheng,Shih-Chii Liu
备注:Accepted to ICASSP 2026
摘要:最近的研究表明,部署后自适应可以提高语音增强模型在不可见噪声条件下的鲁棒性。然而,现有的方法通常会导致高昂的计算和存储成本,限制了它们对设备上部署的适用性。在这项工作中,我们研究了具有动态声学场景变化的现实环境中的模型自适应,并提出了一个轻量级框架,该框架通过自监督训练更新低等级适配器来增强冻结的骨干。在37种噪声类型和三个信噪比范围(包括具有挑战性的[-8,0] dB范围)的111个环境中进行的连续场景评估实验表明,我们的方法更新的基本模型参数不到1%,同时每个场景仅在20次更新内实现了平均1.51 dB的SI-SDR改进。与最先进的方法相比,我们的框架实现了具有竞争力或卓越的感知质量,具有更平滑和更稳定的收敛,证明了其在现实世界声学条件下对语音增强模型的轻量级设备自适应的实用性。
摘要:Recent studies have shown that post-deployment adaptation can improve the robustness of speech enhancement models in unseen noise conditions. However, existing methods often incur prohibitive computational and memory costs, limiting their suitability for on-device deployment. In this work, we investigate model adaptation in realistic settings with dynamic acoustic scene changes and propose a lightweight framework that augments a frozen backbone with low-rank adapters updated via self-supervised training. Experiments on sequential scene evaluations spanning 111 environments across 37 noise types and three signal-to-noise ratio ranges, including the challenging [-8, 0] dB range, show that our method updates fewer than 1% of the base model's parameters while achieving an average 1.51 dB SI-SDR improvement within only 20 updates per scene. Compared to state-of-the-art approaches, our framework achieves competitive or superior perceptual quality with smoother and more stable convergence, demonstrating its practicality for lightweight on-device adaptation of speech enhancement models under real-world acoustic conditions.


【9】Fast and Flexible Audio Bandwidth Extension via Vocos
标题:通过Vocos快速灵活的音频带宽扩展
链接:https://arxiv.org/abs/2603.07285

作者:Yatharth Sharma
备注:5 pages, 2 figures, 5 tables. Submitted to INTERSPEECH 2026. Code available at https://github.com/ysharma3501/LavaSR.git
摘要:我们提出了一个基于Vocos的带宽扩展模型,通过生成丢失的高频内容来增强8-48 kHz的音频。输入被重新采样到48 kHz,并由神经声码器骨干处理,使单个网络能够支持任意上采样率。一个轻量级的Linkwitz-Riley启发细化合并原始低频带与生成的高频通过平滑的交叉。经过验证,该模型在NVIDIA A100 GPU上以0.0001的实时系数和在8核CPU上以0.0053的实时系数运行时,实现了具有竞争力的对数谱距离,证明了在极端吞吐量下的实用、高质量斗轮挖掘机。
摘要:We propose a Vocos-based bandwidth extension model that enhances audio at 8-48 kHz by generating missing high-frequency content. Inputs are resampled to 48 kHz and processed by a neural vocoder backbone, enabling a single network to support arbitrary upsampling ratios. A lightweight Linkwitz-Riley-inspired refiner merges the original low band with the generated high frequencies via a smooth crossover. On validation, the model achieves competitive log-spectral distance while running at a real-time factor of 0.0001 on an NVIDIA A100 GPU and 0.0053 on an 8-core CPU, demonstrating practical, high-quality BWE at extreme throughput.


【10】Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
标题:高保真音频无损压缩的基准语言建模
链接:https://arxiv.org/abs/2603.08683

作者:Phillip Long,Zachary Novack,Chris Donahue
备注:Submitted for review at Interspeech 2026, 7 pages, 5 figures
摘要:在原始波形上训练的自回归“语言”模型(LM)可以重新用于无损音频压缩,但之前的工作仅限于8位音频,因此这些方法是否适用于实际设置(16/24位)以及是否可以与现有编解码器竞争尚不清楚。我们在不同领域(音乐,语音,生物声学),采样率(16 kHz-48 kHz)和位深度(8,16,24位)的全保真度音频上对基于LM的压缩进行基准测试。由于词汇量的大小(16位为65 K; 24位为16.7M),标准的样本级标记化在更高的位深度上变得难以处理。我们提出了Trilobyte,一种用于全分辨率音频的字节级标记化方案,将词汇扩展从$O(2^{b})$提高到$O(1)$,并实现了第一个易于处理的24位基于LM的无损压缩。虽然LM始终优于FLAC,并在8位和16位时产生最先进的压缩,但我们观察到,随着位深度增加到8位以上,压缩增益变得更加适度。
摘要:Autoregressive "language" models (LMs) trained on raw waveforms can be repurposed for lossless audio compression, but prior work is limited to 8-bit audio, leaving open whether such approaches work for practical settings (16/24-bit) and can compete with existing codecs. We benchmark LM-based compression on full-fidelity audio across diverse domains (music, speech, bioacoustics), sampling rates (16kHz-48kHz), and bit depths (8, 16, 24-bit). Standard sample-level tokenization becomes intractable at higher bit depths due to vocabulary size (65K for 16-bit; 16.7M for 24-bit). We propose Trilobyte, a byte-level tokenization schema for full resolution audio, improving vocabulary scaling from $O(2^{b})$ to $O(1)$ and enabling the first tractable 24-bit LM-based lossless compression. While LMs consistently outperform FLAC and yield state-of-the-art compression at 8-bit and 16-bit, we observe that compression gains become more modest as bit depth increases beyond 8-bit.


【11】Computational modeling of early language learning from acoustic speech and audiovisual input without linguistic priors
标题:从没有语言先验的声学语音和视听输入中进行早期语言学习的计算建模
链接:https://arxiv.org/abs/2603.08359

作者:Okko Räsänen
摘要:学习理解语音似乎几乎毫不费力的典型发展中的婴儿,但从信息处理的角度来看,从声学语音获得语言是一个巨大的挑战。本章回顾了最近的发展,使用计算模型来理解早期语言习得的语音和视听输入。重点是自我监督和视觉接地模型的感知学习。我们展示了这些模型是如何变得越来越强大,在学习语音的各个方面没有强大的语言先验,以及有多少功能的早期语言发展可以解释通过一套共享的学习原则广泛兼容的语言习得和人类认知的多种理论。我们还讨论了现代学习模拟如何逐渐变得更加现实,无论是在输入数据方面,还是在将模型行为与婴儿语言发展的实证研究结果联系起来方面。
摘要:Learning to understand speech appears almost effortless for typically developing infants, yet from an information-processing perspective, acquiring a language from acoustic speech is an enormous challenge. This chapter reviews recent developments in using computational models to understand early language acquisition from speech and audiovisual input. The focus is on self-supervised and visually grounded models of perceptual learning. We show how these models are becoming increasingly powerful in learning various aspects of speech without strong linguistic priors, and how many features of early language development can be explained through a shared set of learning principles-principles broadly compatible with multiple theories of language acquisition and human cognition. We also discuss how modern learning simulations are gradually becoming more realistic, both in terms of input data and in linking model behavior to empirical findings on infant language development.


【12】Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction
标题:用于模糊情绪预测的大型音频语言模型中的理清推理
链接:https://arxiv.org/abs/2603.08230

作者:Xiaofeng Yu,Jiaheng Dong,Jean Honorio,Abhirup Ghosh,Hong Jia,Ting Dang
备注:The paper was submitted to Interspeech for review
摘要:语音情感识别在各种应用中起着重要的作用。然而,大多数现有的方法预测一个单一的情感标签,过度简化了人类情感表达的内在模糊性。最近的大型音频语言模型在产生更丰富的输出方面显示出了希望,但它们对模糊情感理解的推理能力仍然有限。在这项工作中,我们重新制定了模糊的情感识别作为一个分布式推理问题,并提出了第一个系统的研究模糊意识推理LALM。我们的框架包括两个互补的组成部分:一个模糊意识的目标,使预测与人类的感知分布,和一个结构化的模糊意识的思想链的监督,指导推理的情感线索。IEMOCAP和CREMA-D上的实验证明了SFT、DPO和GRPO训练策略的一致改进。
摘要:Speech emotion recognition plays an important role in various applications. However, most existing approaches predict a single emotion label, oversimplifying the inherently ambiguous nature of human emotional expression. Recent large audio-language models show promise in generating richer outputs, but their reasoning ability for ambiguous emotional understanding remains limited. In this work, we reformulate ambiguous emotion recognition as a distributional reasoning problem and present the first systematic study of ambiguity-aware reasoning in LALMs. Our framework comprises two complementary components: an ambiguity-aware objective that aligns predictions with human perceptual distributions, and a structured ambiguity-aware chain-of-thought supervision that guides reasoning over emotional cues. Experiments on IEMOCAP and CREMA-D demonstrate consistent improvements across SFT, DPO, and GRPO training strategies.


【13】Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
标题:Foley-Flow:具有掩蔽视听对齐和动态条件流的协调视频到音频生成
链接:https://arxiv.org/abs/2603.08126

作者:Shentong Mo,Yibing Song
摘要:基于视频输入的协调音频生成通常需要严格的视听(AV)对齐,其中所生成的音频片段的语义和节奏都应对应于视频帧中的那些。先前的研究利用两阶段设计,其中AV编码器首先通过对比学习对齐,然后编码的视频表示引导音频生成过程。我们观察到,对比学习和全球视频指导是有效的调整整体AV语义,同时限制时间节奏同步。在这项工作中,我们提出FoleyFlow首先通过掩码建模训练对齐单峰AV编码器,其中在相应视频片段的指导下恢复掩码音频片段。在训练之后,仅使用单峰数据单独预训练的AV编码器与语义和节奏一致性对齐。然后,我们开发了一个动态的条件流的最终音频生成。建立在高效的速度流生成框架,我们的动态条件流利用随时间变化的视频特征作为动态条件,以指导相应的音频段生成。为此,我们提取连贯的语义和节奏表示在掩蔽AV对齐,并使用这种表示的视频片段,以指导音频生成时间。我们的音频结果在标准基准上进行评估,并在几个指标下大大超过现有结果。优越的性能表明,FoleyFlow是有效的生成协调的音频,语义和节奏连贯的各种视频序列。
摘要:Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall correspond to those in the video frames. Previous studies leverage a two-stage design where the AV encoders are firstly aligned via contrastive learning, then the encoded video representations guide the audio generation process. We observe that both contrastive learning and global video guidance are effective in aligning overall AV semantics while limiting temporally rhythmic synchronization. In this work, we propose FoleyFlow to first align unimodal AV encoders via masked modeling training, where the masked audio segments are recovered under the guidance of the corresponding video segments. After training, the AV encoders which are separately pretrained using only unimodal data are aligned with semantic and rhythmic consistency. Then, we develop a dynamic conditional flow for the final audio generation. Built upon the efficient velocity flow generation framework, our dynamic conditional flow utilizes temporally varying video features as the dynamic condition to guide corresponding audio segment generations. To this end, we extract coherent semantic and rhythmic representations during masked AV alignment, and use this representation of video segments to guide audio generation temporally. Our audio results are evaluated on the standard benchmarks and largely surpass existing results under several metrics. The superior performance indicates that FoleyFlow is effective in generating coordinated audios that are both semantically and rhythmically coherent to various video sequences.


【14】WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
标题:WhispEar:一种基于伪并行耳语生成的双向分级耳语语音转换框架
链接:https://arxiv.org/abs/2603.08046

作者:Zihao Fang,Yingda Shen,Zifan Guan,Tongtong Song,Zhenyi Liu,Zhizheng Wu
备注:Submitted to Interspeech 2026
摘要:耳语语音缺乏声带振动和基频,导致声学线索退化,并使耳语到正常(W2N)转换具有挑战性,特别是在有限的并行数据。我们提出了WhispEar,一个双向框架的基础上统一的语义表示,捕捉说话模式不变的信息共享的耳语和正常的语音。该框架包含W2N和正常到耳语(N2W)模型。值得注意的是,N2W模型能够从丰富的正常语音中实现zero-shot伪并行耳语生成,从而允许用于W2N训练的可扩展数据增强。不断增加生成的数据可持续提高性能。我们还发布了迄今为止最大的双语(汉英)正常耳语平行语料库。实验表明,WhispEar优于强基线,并从可扩展的伪并行数据中受益匪浅。
摘要:Whispered speech lacks vocal fold vibration and fundamental frequency, resulting in degraded acoustic cues and making whisper-to-normal (W2N) conversion challenging, especially with limited parallel data. We propose WhispEar, a bidirectional framework based on unified semantic representations that capture speaking-mode-invariant information shared by whispered and normal speech. The framework contains both W2N and normal-to-whisper (N2W) models. Notably, the N2W model enables zero-shot pseudo-parallel whisper generation from abundant normal speech, allowing scalable data augmentation for W2N training. Increasing generated data consistently improves performance. We also release the largest bilingual (Chinese-English) whispered-normal parallel corpus to date. Experiments demonstrate that WhispEar outperforms strong baselines and benefits significantly from scalable pseudo-parallel data.


【15】SoundWeaver: Semantic Warm-Starting for Text-to-Audio Diffusion Serving
标题:SoundWeaver:文本到音频扩散服务的语义热身
链接:https://arxiv.org/abs/2603.07865

作者:Ayush Barik,Sofia Stoica,Nikhil Sarda,Arnav Kethana,Abhinav Khanduja,Muchen Xu,Fan Lai
备注:Submitted to INTERSPEECH 2026
摘要:文本到音频扩散模型产生高保真音频,但需要数十个函数评估(NFE),导致多秒延迟和有限的吞吐量。我们提出了SoundWeaver,第一个无训练,模型不可知的服务系统,通过从语义相似的缓存音频热启动来加速文本到音频的扩散。SoundWeaver引入了三个组件:一个引用缓存器,通过语义和持续时间感知门控来检索和在时间上对齐缓存的候选项;一个跳过门控器,动态确定要跳过的NFE的百分比;以及一个轻量级缓存管理器,通过质量感知驱逐和细化来维护缓存实用程序。在真实世界的音频跟踪上,SoundWeaver仅使用${\sim}$1K条目的缓存就实现了1.8- 3.0 $\times $延迟减少,同时保持或提高了感知质量。
摘要:Text-to-audio diffusion models produce high-fidelity audio but require tens of function evaluations (NFEs), incurring multi-second latency and limited throughput. We present SoundWeaver, the first training-free, model-agnostic serving system that accelerates text-to-audio diffusion by warm-starting from semantically similar cached audio. SoundWeaver introduces three components: a Reference Selector that retrieves and temporally aligns cached candidates via semantic and duration-aware gating; a Skip Gater that dynamically determines the percentage of NFEs to skip; and a lightweight Cache Manager that maintains cache utility through quality-aware eviction and refinement. On real-world audio traces, SoundWeaver achieves 1.8--3.0$ \times $ latency reduction with a cache of only ${\sim}$1K entries while preserving or improving perceptual quality.


【16】Analysis-Driven Procedural Generation of an Engine Sound Dataset with Embedded Control Annotations
标题:分析驱动的带有嵌入式控制注释的发动机声音数据集的过程生成
链接:https://arxiv.org/abs/2603.07584

作者:Robin Doerfler,Lonce Wyse
备注:Preprint. 19 hours of engine audio, 5,935 files, sample-accurate annotations. Dataset publicly available at https://doi.org/10.5281/zenodo.16883336 and https://huggingface.co/datasets/rdoerfler/procedural-engine-sounds
摘要:计算引擎声音建模是汽车音响行业的核心,特别是对于主动声音设计、虚拟原型和新兴的数据驱动引擎声音合成方法。这些应用需要大量的标准化、清晰的音频记录,并具有精确的时间对齐的操作状态注释:由于成本高、专业测量设备要求和不可避免的噪声污染,这些数据很难获得。   我们提出了一个分析驱动的框架,用于生成具有样本准确控制注释的引擎音频。该方法通过音调自适应频谱分析从实际录音中提取谐波结构,然后驱动扩展的参数谐波加噪声合成器。   通过这个框架,我们生成了程序引擎声音数据集(19小时,5,935个文件),这是一组具有样本准确的RPM和扭矩注释的引擎音频信号,涵盖了广泛的操作条件,信号复杂性和谐波特征。与真实记录的比较验证了合成数据保留了特征谐波结构,基线实验证实了其适用于基于学习的参数估计和合成任务。   该数据集公开发布,以支持发动机音色分析,控制参数估计,声学建模和神经生成网络的研究。
摘要:Computational engine sound modeling is central to the automotive audio industry, particularly for active sound design, virtual prototyping, and emerging data-driven engine sound synthesis methods. These applications require large volumes of standardized, clean audio recordings with precisely time-aligned operating-state annotations: data that is difficult to obtain due to high costs, specialized measurement equipment requirements, and inevitable noise contamination.   We present an analysis-driven framework for generating engine audio with sample-accurate control annotations. The method extracts harmonic structures from real recordings through pitch-adaptive spectral analysis, which then drive an extended parametric harmonic-plus-noise synthesizer.   With this framework, we generate the Procedural Engine Sounds Dataset (19 hours, 5,935 files), a set of engine audio signals with sample-accurate RPM and torque annotations, spanning a wide range of operating conditions, signal complexities, and harmonic profiles. Comparison against real recordings validates that the synthesized data preserves characteristic harmonic structures, and baseline experiments confirm its suitability for learning-based parameter estimation and synthesis tasks.   The dataset is released publicly to support research on engine timbre analysis, control parameter estimation, acoustic modeling and neural generative networks.


【17】Evaluating Parkinson's Disease Detection in Anonymized Speech: A Performance and Acoustic Analysis
标题:评估语音中帕金森病检测:性能和声学分析
链接:https://arxiv.org/abs/2603.07544

作者:Carlos Franzreb,Francisco Teixeira,Ben Luks,Sebastian Möller,Alberto Abad
备注:Submitted to Interspeech 2026
摘要:从语音中自动检测帕金森病(PD)是一种很有前途的非侵入性诊断工具,但它引起了严重的隐私问题。说话者匿名化减轻了这些风险,但它可能会抑制PD检测所需的病理信息。我们使用两个西班牙语数据集评估两个匿名者(STT-TTS和kNN-VC)的隐私和PD检测之间的权衡。STT-TTS提供了更好的隐私,但通过消除韵律信息严重降低了PD检测。kNN-VC保留了宏韵律特征,如持续时间和F0轮廓,实现F1分数仅比原始基线低3- 7%,表明在使用适当的匿名化时,隐私保护PD检测是可行的。最后,声学失真分析表征了kNN-VC中的特定弱点,为设计更好地保留PD信息的匿名器提供了见解。
摘要:Automatic detection of Parkinson's disease (PD) from speech is a promising non-invasive diagnostic tool, but it raises significant privacy concerns. Speaker anonymization mitigates these risks, but it may suppress the pathological information necessary for PD detection. We assess the trade-off between privacy and PD detection for two anonymizers (STT-TTS and kNN-VC) using two Spanish datasets. STT-TTS provides better privacy but severely degrades PD detection by eradicating prosodic information. kNN-VC preserves macro-prosodic features such as duration and F0 contours, achieving F1 scores only 3-7\% lower than original baselines, demonstrating that privacy-preserving PD detection is viable when using appropriate anonymization. Finally, an acoustic distortion analysis characterizes specific weaknesses in kNN-VC, offering insights for designing anonymizers that better preserve PD information.


【18】Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
标题:查看上下文:通过多模式推理的丰富视觉上下文感知语音识别
链接:https://arxiv.org/abs/2603.07263

作者:Wenjie Tian,Mingchen Shao,Bingshen Mu,Xuelong Geng,Chengyou Wang,Yujie Liao,Zhixian Zhao,Ziyu Zhang,Jingbin Hu,Mengqi Wei,Lei Xie
摘要:视听语音识别(AVSR)是结合视觉信号的ASR的扩展。目前的AVSR方法主要集中在嘴唇运动,在很大程度上忽略了丰富的背景下,目前在视频中,如说话场景和屏幕上的文字。为了解决这样的CAVSR(AVSR包括丰富的视觉上下文),我们提出了VASR,旨在“看到”和原因的视觉上下文,以提高语音识别。具体来说,我们构建了一个视听思想链(AV-CoT),明确执行声学信号和视觉证据之间的中间跨模态接地。这种证据驱动的推理缓解了“单一模态优势”问题,即模型过度依赖视觉上下文或无法利用视觉上下文。此外,为了解决数据稀缺问题,我们构建并发布了相应的数据管道和测试集。实验表明,AV-CoT有效地缓解了单模态优势,在CAVSR中实现了最先进的性能。该项目是开源的。
摘要:Audio-visual speech recognition (AVSR) is an extension of ASR that incorporates visual signals. Current AVSR approaches primarily focus on lip motion, largely overlooking rich context present in the video such as speaking scene and on-screen text. To tackle such CAVSR (AVSR including rich visual Context), we propose VASR designed to "see" and reason the visual context to improve speech recognition. Specifically, we construct an Audio-Visual Chain-of-Thought (AV-CoT) that explicitly enforces intermediate cross-modal grounding between acoustic signals and visual evidence. This evidence-driven reasoning mitigates the "single-modality dominance" problem, where models either over-rely on visual context or fail to utilize it. Besides, to address the data scarcity, we construct and release a corresponding data pipeline and test set. Experiments show that AV-CoT effectively mitigates the single-modality dominance, achieving state-of-the-art performance in CAVSR. The project is open-sourced.


【19】Scaling Self-Supervised Speech Models Uncovers Deep Linguistic Relationships: Evidence from the Pacific Cluster
标题:扩展自我监督言语模型揭示深层语言关系:来自太平洋集群的证据
链接:https://arxiv.org/abs/2603.07238

作者:Minu Kim,Hoirin Kim,David R. Mortensen
备注:Submitted to Interspeech 2026
摘要:从自监督语音模型(S3M)导出的语言表示之间的相似性已被观察到,主要反映了最近的扩张或接触所驱动的地理接近或表面类型相似性,可能会丢失更深层次的谱系信号。我们研究了基于S3M的语言识别系统的语言覆盖范围从126到4,017种语言如何影响这种拓扑结构。我们的研究结果揭示了一种非线性效应:虽然系统发育恢复在1K尺度上仍然停滞不前,但4K模型显示出巨大的质的转变,解决了清晰的谱系和复杂的长期语言接触。值得注意的是,我们的分析揭示了太平洋地区(包括巴布亚语,大洋洲和澳大利亚语言)出现了一个强大的宏观集群,并调查了其潜在的驱动因素。我们发现,4K模型采用了更集中的编码,可以捕获共享的、鲁棒的声学签名,例如全局能量动态。这些发现表明,大量的S3Ms可以内化语言历史的多个层面,为计算遗传学和语言接触的研究提供了一个有前途的前景。
摘要:Similarities between language representations derived from Self-Supervised Speech Models (S3Ms) have been observed to primarily reflect geographic proximity or surface typological similarities driven by recent expansion or contact, potentially missing deeper genealogical signals. We investigate how scaling linguistic coverage of an S3M-based language identification system from 126 to 4,017 languages influences this topology. Our results reveal a non-linear effect: while phylogenetic recovery remains stagnant up to the 1K scale, the 4K model displays a dramatic qualitative shift, resolving both clear lineages and complex, long-term linguistic contact. Notably, our analysis reveals the emergence of a robust macro-cluster in the Pacific (comprising Papuan, Oceanic, and Australian languages) and investigates its latent drivers. We find that the 4K model utilizes a more concentrated encoding that captures shared, robust acoustic signatures such as global energy dynamics. These findings suggest that massive S3Ms can internalize multiple layers of language history, providing a promising perspective for computational phylogenetics and the study of language contact.


【20】Towards Objective Gastrointestinal Auscultation: Automated Segmentation and Annotation of Bowel Sound Patterns
标题:走向客观的胃肠道听诊:肠道声音模式的自动分割和注释
链接:https://arxiv.org/abs/2603.07215

作者:Zahra Mansour,Verena Uslar,Dirk Weyhe,Danilo Hollosi,Nils Strodthoff
摘要:肠鸣音(BS)通常是瞬时的并且具有低振幅,使得它们难以通过手动听诊准确地检测。这导致临床评估的显著差异。数字声学传感器允许采集高质量的BS并实现自动信号分析,从而有可能为临床医生提供有关肠道活动的客观和定量反馈。这项研究提出了一个自动管道肠声音分割和分类使用可穿戴声学SonicGuard传感器。使用SonicGuard传感器记录来自83名受试者的BS信号。来自40名受试者的数据由临床专家手动注释并用于训练自动注释算法,而其余受试者用于进一步的模型评估。提出了一种基于能量的事件检测算法。然后使用预训练的音频频谱图Transformer(AST)模型将检测到的声音片段分类为BS模式。分别评价健康个体和患者的模型性能。最佳配置使用两个专门的模型,一个在健康受试者上训练,一个在患者上训练,健康组和患者组分别达到(准确度:0.97,AUROC:0.98)和(准确度:0.96,AUROC:0.98)。自动注释方法将手动标记时间减少了约70%,专家评审表明,只有不到12%的自动检测片段需要校正。所提出的自动分割和分类系统能够定量评估肠道活动,为临床医生提供客观的诊断工具,可以改善胃肠道功能的诊断,并支持大规模数据集的注释。
摘要:Bowel sounds (BS) are typically momentary and have low amplitude, making them difficult to detect accurately through manual auscultation. This leads to significant variability in clinical assessment. Digital acoustic sensors allow the acquisition of high-quality BS and enable automated signal analysis, offering the potential to provide clinicians with both objective and quantitative feedback on bowel activity. This study presents an automated pipeline for bowel sound segmentation and classification using a wearable acoustic SonicGuard sensor. BS signals from 83 subjects were recorded using a SonicGuard sensor. Data from 40 subjects were manually annotated by clinical experts and used to train an automatic annotation algorithm, while the remaining subjects were used for further model evaluation. An energy-based event detection algorithm was developed to detect BS events. Detected sound segments were then classified into BS patterns using a pretrained Audio Spectrogram Transformer (AST) model. Model performance was evaluated separately for healthy individuals and patients. The best configuration used two specialized models, one trained on healthy subjects and one on patients, achieving (accuracy: 0.97, AUROC: 0.98) for healthy group and (accuracy: 0.96, AUROC: 0.98) for patient group. The auto-annotation method reduced manual labeling time by approximately 70%, and expert review showed that less than 12% of automatically detected segments required correction. The proposed automated segmentation and classification system enables quantitative assessment of bowel activity, providing clinicians with an objective diagnostic tool that may improve the diagnostic of gastrointestinal function and support the annotation of large-scale datasets.


【21】Toward Multimodal Industrial Fault Analysis: A Single-Speed Chain Conveyor Dataset with Audio and Vibration Signals
标题:走向多模式工业故障分析:具有音频和振动信号的单速链条输送机数据集
链接:https://arxiv.org/abs/2603.07130

作者:Zhang Chen,Yucong Zhang,Xiaoxiao Miao,Ming Li
备注:Submitted to Interspeech 2026
摘要:我们介绍了从单速链式输送机(SSCC)系统收集的多模态工业故障分析数据集,针对生产线中的系统级故障检测。该数据集由多模态信号组成,包括三个音频通道和四个振动通道。它涵盖了在多种速度、负载以及现场再现的干净和真实的工厂噪声条件下的正常运行和四种代表性故障类型。它被明确设计为支持通道分析和多模态融合研究。我们建立了标准化的评估协议,用于无监督的故障检测,仅进行正常训练,并在不同的操作条件和故障类型之间进行平衡的数据集分割,进行有监督的故障分类。提供了统一的通道kNN基线,以便在没有特定任务训练的情况下公平比较表示质量。该数据集为强大的多模态工业故障分析提供了一个实用和可扩展的基准。
摘要:We introduce a multimodal industrial fault analysis dataset collected from a single-speed chain conveyor (SSCC) system, targeting system-level fault detection in production lines. The dataset consists of multimodal signals, including three audio and four vibration channels. It covers normal operation and four representative fault types under multiple speeds, loads, and both clean and realistic factory-noise conditions reproduced on-site. It is explicitly designed to support channel-wise analysis and multimodal fusion research. We establish standardized evaluation protocols for unsupervised fault detection with normal-only training and supervised fault classification with balanced dataset splits across different operating conditions and fault types. A unified channel-wise kNN baseline is provided to enable fair comparison of representation quality without task-specific training. The dataset offers a practical and extensible benchmark for robust multimodal industrial fault analysis.


机器翻译由腾讯交互翻译提供,仅供参考