微信公众号:arXiv_Daily
cs.SD语音
【1】Building Enterprise Realtime Voice Agents from Scratch: A Technical Tutorial
标题:从Scratch构建企业实时语音代理:技术任务集
链接:https://arxiv.org/abs/2603.05413
摘要:我们提出了一个技术教程,从第一原则建立企业级实时语音代理。虽然存在超过25个开源语音到语音模型和许多语音代理框架,但没有一个资源解释从单个组件到具有函数调用功能的工作流语音代理的完整管道。通过系统的研究,我们发现:(1)像Qwen2.5-Omni这样的原生语音到语音模型,虽然能够生成高质量的音频,但对于实时交互来说太慢了(2)行业标准方法使用级联流管道:STT $\rightarrow$ LLM $\rightarrow$ TTS,其中每个组件将其输出流式传输到下一个组件;(3)“实时”的关键不是任何单一的快速模型,而是跨组件的流和流水线。我们使用Deepgram(流STT)、具有函数调用功能的vLLM-served LLM(流文本生成)和ElevenLabs(流TTS)构建了一个完整的语音代理,使用云LLM API实现了947 ms的测量P50首次音频时间(最佳情况为729 ms),并且与NVIDIA A10 G GPU上的自托管vLLM具有相当的延迟。我们将完整的代码库作为教程发布,其中包含每个组件的工作,测试代码。
摘要:We present a technical tutorial for building enterprise-grade realtime voice agents from first principles. While over 25 open-source speech-to-speech models and numerous voice agent frameworks exist, no single resource explains the complete pipeline from individual components to a working streaming voice agent with function calling capabilities. Through systematic investigation, we find that (1) native speech-to-speech models like Qwen2.5-Omni, while capable of high-quality audio generation, are too slow for realtime interaction ($\sim$13s time-to-first-audio); (2) the industry-standard approach uses a cascaded streaming pipeline: STT $\rightarrow$ LLM $\rightarrow$ TTS, where each component streams its output to the next; and (3) the key to ``realtime'' is not any single fast model but rather \textit{streaming and pipelining} across components. We build a complete voice agent using Deepgram (streaming STT), vLLM-served LLMs with function calling (streaming text generation), and ElevenLabs (streaming TTS), achieving a measured P50 time-to-first-audio of 947ms (best case 729ms) with cloud LLM APIs, and comparable latency with self-hosted vLLM on NVIDIA A10G GPU. We release the full codebase as a tutorial with working, tested code for every component.
【2】Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
标题:具有多分辨率欺骗检测的离散语音合成分层解码
链接:https://arxiv.org/abs/2603.05373
备注:7 pages, 3 figures, 3 tables, 2 algorithms
摘要:神经编解码器语言模型使高质量的离散语音合成,但他们的推理仍然容易受到令牌级文物和分布式漂移,降低感知的现实主义。而不是依赖于偏好优化或再训练,我们提出了MSpoof-TTS,一个无训练的推理框架,通过多分辨率欺骗指导提高了zero-shot合成。我们引入了一个基于多分辨率令牌的欺骗检测框架,该框架评估不同时间粒度的编解码器序列,以检测局部不一致或不自然的模式。然后,我们将欺骗检测器集成到一个分层解码策略中,逐步修剪低质量的候选人和重新排名的假设。这种鉴别器引导的生成在不修改模型参数的情况下增强了鲁棒性。实验验证了我们的框架的有效性,鲁棒性和高质量的基于编解码器的语音生成。
摘要:Neural codec language models enable high-quality discrete speech synthesis, yet their inference remains vulnerable to token-level artifacts and distributional drift that degrade perceptual realism. Rather than relying on preference optimization or retraining, we propose MSpoof-TTS, a training-free inference framework that improves zero-shot synthesis through multi-resolution spoof guidance. We introduce a Multi-Resolution Token-based Spoof Detection framework that evaluates codec sequences at different temporal granularities to detect locally inconsistent or unnatural patterns. We then integrate the spoof detectors into a hierarchical decoding strategy, progressively pruning low-quality candidates and re-ranking hypotheses. This discriminator-guided generation enhances robustness without modifying model parameters. Experiments validate the effectiveness of our framework for robust and high-quality codec-based speech generation.
【3】Latent-Mark: An Audio Watermark Robust to Neural Resynthesis
标题:潜伏标记:对神经再合成鲁棒的音频水印
链接:https://arxiv.org/abs/2603.05310
摘要:虽然现有的音频水印技术对传统的数字信号处理(DSP)攻击具有很强的鲁棒性,但它们仍然容易受到神经再合成的影响。这是因为现代神经音频编解码器充当语义过滤器并丢弃先前水印方法中使用的不可感知的波形变化。为了解决这个问题,我们提出了潜伏标记,第一个零比特音频水印框架,旨在生存语义压缩。我们的关键见解是,鲁棒性的编码-解码过程中需要嵌入水印的编解码器的不变的潜在空间。我们通过优化音频波形来实现这一点,以在其编码的潜在表示中引起可检测的方向性偏移,同时约束扰动以与自然音频流形对齐以确保不可感知性。为了防止过拟合到单个编解码器的量化规则,我们引入了跨编解码器优化,跨多个代理编解码器联合优化波形,以实现共享的潜在不变量。广泛的评估表明,强大的zero-shot可转移到看不见的神经编解码器,实现了最先进的弹性对传统的DSP攻击,同时保持感知的不可感知性。我们的工作启发了未来对通用水印框架的研究,这些框架能够在日益复杂和多样化的生成失真中保持完整性。
摘要:While existing audio watermarking techniques have achieved strong robustness against traditional digital signal processing (DSP) attacks, they remain vulnerable to neural resynthesis. This occurs because modern neural audio codecs act as semantic filters and discard the imperceptible waveform variations used in prior watermarking methods. To address this limitation, we propose Latent-Mark, the first zero-bit audio watermarking framework designed to survive semantic compression. Our key insight is that robustness to the encode-decode process requires embedding the watermark within the codec's invariant latent space. We achieve this by optimizing the audio waveform to induce a detectable directional shift in its encoded latent representation, while constraining perturbations to align with the natural audio manifold to ensure imperceptibility. To prevent overfitting to a single codec's quantization rules, we introduce Cross-Codec Optimization, jointly optimizing the waveform across multiple surrogate codecs to target shared latent invariants. Extensive evaluations demonstrate robust zero-shot transferability to unseen neural codecs, achieving state-of-the-art resilience against traditional DSP attacks while preserving perceptual imperceptibility. Our work inspires future research into universal watermarking frameworks capable of maintaining integrity across increasingly complex and diverse generative distortions.
【4】SLICE: Speech Enhancement via Layer-wise Injection of Conditioning Embeddings
标题:切片:通过分层注入条件嵌入的言语增强
链接:https://arxiv.org/abs/2603.05302
备注:5 pages, 1 figure, 4 tables, submitted to INTERSPEECH 2026
摘要:真实世界的语音常常同时受到多种退化的破坏,包括加性噪声、混响和非线性失真。基于扩散的增强方法在单一退化上表现良好,但在复合退化上表现不佳。先前的噪声感知方法仅在输入层注入调节,这会使性能降低到低于未调节模型的性能。为了解决这个问题,我们建议将来自预训练编码器的退化调节(具有用于噪声类型,混响和失真的多任务头)注入到时间步长嵌入中,以便它在没有架构变化的情况下通过所有残差块传播。在只有注入方法不同的受控实验中,输入级调节在复合降解方面的表现比没有编码器更差,而逐层注入则达到了最佳效果。该方法还推广到不同的现实世界的记录。
摘要:Real-world speech is often corrupted by multiple degradations simultaneously, including additive noise, reverberation, and nonlinear distortion. Diffusion-based enhancement methods perform well on single degradations but struggle with compound corruptions. Prior noise-aware approaches inject conditioning at the input layer only, which can degrade performance below that of an unconditioned model. To address this, we propose injecting degradation conditioning, derived from a pretrained encoder with multi-task heads for noise type, reverberation, and distortion, into the timestep embedding so that it propagates through all residual blocks without architectural changes. In controlled experiments where only the injection method varies, input-level conditioning performs worse than no encoder at all on compound degradations, while layer-wise injection achieves the best results. The method also generalizes to diverse real-world recordings.
【5】WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
标题:WavSLM:通过WavLM蒸馏的单流语音语言建模
链接:https://arxiv.org/abs/2603.05299
备注:6 pages, 1 figure
摘要:大型语言模型表明,简单的自回归训练可以产生可扩展和连贯的生成,但由于语义和声学信息的纠缠,将这种范式扩展到语音仍然具有挑战性。大多数现有的语音语言模型依赖于文本监督,分层令牌流或复杂的混合架构,偏离了已被证明在文本中有效的单流生成预训练范式。在这项工作中,我们引入了WavSLM,这是一种语音语言模型,通过将自监督WavLM表示量化和提取到单个码本中并优化自回归下一个块预测目标来训练。WavSLM在单个令牌流中联合建模语义和声学信息,无需文本监督或文本预训练。尽管它很简单,但它在一致性基准和语音生成方面具有竞争力的性能,同时使用更少的参数,更少的训练数据,并支持流式推理。演示示例可在https://lucadellalib.github.io/wavslm-web/上获得。
摘要:Large language models show that simple autoregressive training can yield scalable and coherent generation, but extending this paradigm to speech remains challenging due to the entanglement of semantic and acoustic information. Most existing speech language models rely on text supervision, hierarchical token streams, or complex hybrid architectures, departing from the single-stream generative pretraining paradigm that has proven effective in text. In this work, we introduce WavSLM, a speech language model trained by quantizing and distilling self-supervised WavLM representations into a single codebook and optimizing an autoregressive next-chunk prediction objective. WavSLM jointly models semantic and acoustic information within a single token stream without text supervision or text pretraining. Despite its simplicity, it achieves competitive performance on consistency benchmarks and speech generation while using fewer parameters, less training data, and supporting streaming inference. Demo samples are available at https://lucadellalib.github.io/wavslm-web/.
【6】SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning
标题:SarcasmMiner:用于稳健视听讽刺推理的双轨训练后框架
链接:https://arxiv.org/abs/2603.05275
摘要:多模态讽刺检测需要通过跨模态推理来解决文本、声学和视觉线索之间的语用不一致。为了使强大的讽刺推理与基础模型,我们提出了SarcasmMiner,一个基于强化学习的后训练框架,抵抗幻觉的多模态推理。我们将讽刺检测重新定义为结构化推理,并采用双轨蒸馏策略:高质量的教师轨迹初始化学生模型,而完整的轨迹集训练生成奖励模型(GenRM)来评估推理质量。学生使用组相对策略优化(GRPO)进行优化,使用精确度和推理质量的解耦奖励。在MUSTARD++上,SarcasmMiner将F1从59.83%(zero-shot),68.23%(监督微调)增加到70.22%。这些研究结果表明,理性意识的奖励模型提高了性能和多模态接地。
摘要:Multimodal sarcasm detection requires resolving pragmatic incongruity across textual, acoustic, and visual cues through cross-modal reasoning. To enable robust sarcasm reasoning with foundation models, we propose SarcasmMiner, a reinforcement learning based post-training framework that resists hallucination in multimodal reasoning. We reformulate sarcasm detection as structured reasoning and adopt a dual-track distillation strategy: high-quality teacher trajectories initialize the student model, while the full set of trajectories trains a generative reward model (GenRM) to evaluate reasoning quality. The student is optimized with group relative policy optimization (GRPO) using decoupled rewards for accuracy and reasoning quality. On MUStARD++, SarcasmMiner increases F1 from 59.83% (zero-shot), 68.23% (supervised finetuning) to 70.22%. These findings suggest that reasoning-aware reward modeling enhances both performance and multimodal grounding.
【7】Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards
标题:通过带有音频文本语义奖励的测试时强化学习来提高ASB的鲁棒性
链接:https://arxiv.org/abs/2603.05231
摘要:最近,自动语音识别(ASR)系统(例如,Whisper)已经取得了显着的准确性提高,但仍然对现实世界中看不见的数据(具有较大分布变化的数据)高度敏感,包括嘈杂的环境和不同的口音。为了解决这个问题,测试时自适应(TTA)在提高模型的适应性在推理时间没有地面真值标签,和现有的TTA方法往往依赖于伪标记或熵最小化的潜力。然而,通过将模型置信度视为学习信号,这些方法可能会强化高置信度错误,导致破坏自适应的确认偏差。为了克服这些局限性,我们提出了ASR-TRA,一个新的测试时强化自适应框架的灵感来自因果干预。更确切地说,我们的方法引入了一个可学习的解码器提示,并利用温度控制的随机解码来生成不同的转录候选人。这些由衡量音频-文本语义对齐的奖励模型进行评分,所得反馈用于通过强化学习更新模型和提示参数。综合实验LibriSpeech与合成噪声和L2北极口音英语数据集表明,我们的方法实现了更高的准确性,同时保持较低的延迟比现有的TTA基线。消融研究进一步证实了结合音频和基于语言的奖励的有效性,突出了我们的方法的增强的稳定性和可解释性。总的来说,我们的方法为在具有挑战性的现实条件下部署ASR系统提供了一个实用而强大的解决方案。
摘要:Recently, Automatic Speech Recognition (ASR) systems (e.g., Whisper) have achieved remarkable accuracy improvements but remain highly sensitive to real-world unseen data (data with large distribution shifts), including noisy environments and diverse accents. To address this issue, test-time adaptation (TTA) has shown great potential in improving the model adaptability at inference time without ground-truth labels, and existing TTA methods often rely on pseudo-labeling or entropy minimization. However, by treating model confidence as a learning signal, these methods may reinforce high-confidence errors, leading to confirmation bias that undermines adaptation. To overcome these limitations, we present ASR-TRA, a novel Test-time Reinforcement Adaptation framework inspired by causal intervention. More precisely, our method introduces a learnable decoder prompt and utilizes temperature-controlled stochastic decoding to generate diverse transcription candidates. These are scored by a reward model that measures audio-text semantic alignment, and the resulting feedback is used to update both model and prompt parameters via reinforcement learning. Comprehensive experiments on LibriSpeech with synthetic noise and L2 Arctic accented English datasets demonstrate that our method achieves higher accuracy while maintaining lower latency than existing TTA baselines. Ablation studies further confirm the effectiveness of combining audio and language-based rewards, highlighting our method's enhanced stability and interpretability. Overall, our approach provides a practical and robust solution for deploying ASR systems in challenging real-world conditions.
【8】TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling
标题:TW-Sound 580 K:一个用于本地化音频语言建模的验证引导的区域音频文本数据集
链接:https://arxiv.org/abs/2603.05094
摘要:由于缺乏专门的语料库,大型音频语言模型(LALM)通常难以处理本地化的方言韵律。我们提出了TW-Sound 580 K,一个通过验证生成批判(VGC)协议开发的台湾音频文本指令数据集。该管道利用Dual-ASR验证来过滤522 K原始片段,随后使用教师模型将其扩展为580,000个高保真指令对。该数据集的实用性通过Tai-LALM进行了展示,Tai-LALM对DeSTA 2.5-Audio初始化的主干进行了微调,并采用了动态Dual-ASR仲裁策略来优化推理过程中的转录选择。在TAU基准测试中,Tai-LALM达到了49.1%的准确率,比zero-shot基线(ASR文本调节为42.6%)提高了6.5%。这证实了将区域语料库与严格的策展和动态仲裁相结合显著增强了LALM在本地化语音上的性能。
摘要:Large Audio-Language Models (LALMs) typically struggle with localized dialectal prosody due to the scarcity of specialized corpora. We present TW-Sound580K, a Taiwanese audio-text instruction dataset developed through a Verify-Generate-Critique (VGC) protocol. This pipeline leverages Dual-ASR validation to filter 522K raw clips, subsequently expanding them into 580,000 high-fidelity instruction pairs using a teacher model. The dataset's utility is demonstrated through Tai-LALM, which fine-tunes a DeSTA 2.5-Audio-initialized backbone and incorporates a dynamic Dual-ASR Arbitration strategy to optimize transcription selection during inference. On the TAU Benchmark, Tai-LALM reaches 49.1% accuracy, marking a 6.5% absolute improvement over the zero-shot baseline (42.6% with ASR text conditioning). This confirms that integrating regional corpora with rigorous curation and dynamic arbitration significantly enhances LALM performance on localized speech.
【9】Training Dynamics-Aware Multi-Factor Curriculum Learning for Target Speaker Extraction
标题:用于目标说话人提取的训练动态感知多因素课程学习
链接:https://arxiv.org/abs/2603.04943
摘要:目标说话人提取(TSE)的目的是从多个说话人的混合语音中分离出特定说话人的语音。尽管有很强的基准测试结果,但由于不同的相互作用因素,真实世界的性能往往会下降。TSE以前的课程学习方法通常单独解决这些因素,未能捕捉它们复杂的相互作用,并依赖于可能与实际模型学习行为不一致的预定义难度因素。为了应对这一挑战,我们首先提出了一个多因素的课程学习策略,联合调度SNR阈值,扬声器计数,重叠率和合成/真实比例,使渐进式学习从简单到复杂的场景。然而,在没有预定义假设的情况下确定最优调度仍然具有挑战性。因此,我们引入了TSE-Datamap,这是一个可视化框架,通过跟踪训练时期的信心和变化,在观察到的训练动态中进行课程设计。我们的分析揭示了三个特征数据区域:(i)易于学习的示例,其中模型始终表现良好,(ii)模糊的示例,其中模型在替代预测之间振荡,以及(iii)难以学习的示例,其中模型持续挣扎。在这些数据驱动的见解的指导下,我们的方法改进了随机采样的提取结果,在具有挑战性的多说话者场景中具有特别强的增益。
摘要:Target speaker extraction (TSE) aims to isolate a specific speaker's voice from multi-speaker mixtures. Despite strong benchmark results, real-world performance often degrades due to different interacting factors. Previous curriculum learning approaches for TSE typically address these factors separately, failing to capture their complex interactions and relying on predefined difficulty factors that may not align with actual model learning behavior. To address this challenge, we first propose a multi-factor curriculum learning strategy that jointly schedules SNR thresholds, speaker counts, overlap ratios, and synthetic/real proportions, enabling progressive learning from simple to complex scenarios. However, determining optimal scheduling without predefined assumptions remains challenging. We therefore introduce TSE-Datamap, a visualization framework that grounds curriculum design in observed training dynamics by tracking confidence and variability across training epochs. Our analysis reveals three characteristic data regions: (i) easy-to-learn examples where models consistently perform well, (ii) ambiguous examples where models oscillate between alternative predictions, and (iii) hard-to-learn examples where models persistently struggle. Guided by these data-driven insights, our methods improve extraction results over random sampling, with particularly strong gains in challenging multi-speaker scenarios.
【10】The First Environmental Sound Deepfake Detection Challenge: Benchmarking Robustness, Evaluation, and Insights
标题:首届环境声音Deepfake检测挑战:对稳健性、评估和见解进行基准测试
链接:https://arxiv.org/abs/2603.04865
摘要:音频生成的最新进展使得创建高度逼真的环境音景变得越来越容易,这些音景可能被滥用来产生欺骗性内容,例如假警报,枪声和人群声音,从而引发对公共安全和信任的担忧。虽然对语音和歌声的深度伪造检测已经得到了广泛的研究,但环境声音深度伪造检测(ESDD)仍然没有得到充分的研究。为了推进ESDD,第一届ESDD挑战赛已经启动,吸引了97个注册团队,并收到了1,748份有效提交。本文介绍了任务制定,数据集构建,评估协议,基线系统和挑战结果的关键见解。此外,我们分析了最佳性能系统之间的常见架构选择和培训策略。最后,我们讨论了未来潜在的研究方向ESDD,概述了关键的机会和开放的问题,以指导在这一领域的后续研究。
摘要:Recent progress in audio generation has made it increasingly easy to create highly realistic environmental soundscapes, which can be misused to produce deceptive content, such as fake alarms, gunshots, and crowd sounds, raising concerns for public safety and trust. While deepfake detection for speech and singing voice has been extensively studied, environmental sound deepfake detection (ESDD) remains underexplored. To advance ESDD, the first edition of the ESDD challenge was launched, attracting 97 registered teams and receiving 1,748 valid submissions. This paper presents the task formulation, dataset construction, evaluation protocols, baseline systems, and key insights from the challenge results. Furthermore, we analyze common architectural choices and training strategies among top-performing systems. Finally, we discuss potential future research directions for ESDD, outlining key opportunities and open problems to guide subsequent studies in this field.
【11】Focus Then Listen: Exploring Plug-and-Play Audio Enhancer for Noise-Robust Large Audio Language Models
标题:专注然后倾听:探索即可插即用音频增强器,以实现噪音稳健的大型音频语言模型
链接:https://arxiv.org/abs/2603.04862
摘要:大音频语言模型是音频理解的一类基础模型。现有的LALM往往会在语音和非语音声音干扰的现实世界嘈杂的声学条件下显着降低。虽然噪声感知微调可以提高鲁棒性,但它需要特定于任务的噪声数据和昂贵的再训练,限制了可扩展性。为了解决这个问题,我们提出了FTL,即插即用的音频增强器,提高了LALM的噪声鲁棒性。具体地,FTL首先将输入波形分离成语音和非语音,并且应用模态路由器来预测目标音频模态(例如,语音)基于用户的指令。最后,模态感知融合块生成用于改进下游感知和推理的任务自适应增强信号。跨多个LALM和任务的实验表明,FTL在不同噪声水平下提高了性能,而无需对LALM进行微调。
摘要:Large audio language models (LALMs) are a class of foundation models for audio understanding. Existing LALMs tend to degrade significantly in real-world noisy acoustic conditions where speech and non-speech sounds interfere. While noise-aware fine-tuning can improve robustness, it requires task-specific noisy data and expensive retraining, limiting scalability. To address this issue, we propose Focus-Then-Listen (FTL), a plug-and-play audio enhancer that improves LALMs' noise robustness. Specifically, FTL first separates the input waveform into speech and non-speech, and a modality router is applied to predict the target audio modality (e.g., speech) based on the user's instruction. Finally, a modality-aware fusion block generates a task-adaptive enhanced signal for improved downstream perception and reasoning. Experiments across multiple LALMs and tasks show that FTL improves performance across different noise levels without fine-tuning on LALMs.
【12】WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech
标题:WhisperAlign:用于长格式孟加拉语语音的单词边界感知ASB和WhisperX锚定的Pyannote扩展
链接:https://arxiv.org/abs/2603.04809
摘要:本文介绍了我们的DL Sprint 4.0解决方案,解决孟加拉语长格式语音识别(任务1)和扬声器日记(任务2)的双重挑战。处理长格式的多扬声器孟加拉语音频在语音活动检测、重叠语音和上下文保存方面引入了重大障碍。为了解决长格式转录的挑战,我们利用耳语时间戳实现了一个强大的音频分块策略,使我们能够将精确的、上下文感知的片段输入到我们微调的声学模型中,以实现高准确度的转录。对于日志化任务,我们利用pyannote.audio和WhisperX开发了一个集成管道。我们的方法的一个关键贡献是在竞争数据集上对Pyannote分割模型进行了特定领域的微调。这种适应使模型能够更好地捕捉孟加拉语会话动态的细微差别,并准确地解决复杂的,重叠的扬声器边界。我们的方法表明,在低资源环境下,将智能时间戳分块应用于ASR和有针对性的分割微调应用于日记化可以显着降低单词错误率(WER)和日记化错误率(DER)。
摘要:This paper presents our solution for the DL Sprint 4.0, addressing the dual challenges of Bengali Long-Form Speech Recognition (Task 1) and Speaker Diarization (Task 2). Processing long-form, multi-speaker Bengali audio introduces significant hurdles in voice activity detection, overlapping speech, and context preservation. To solve the long-form transcription challenge, we implemented a robust audio chunking strategy utilizing whisper-timestamped, allowing us to feed precise, context-aware segments into our fine-tuned acoustic model for high-accuracy transcription. For the diarization task, we developed an integrated pipeline leveraging pyannote.audio and WhisperX. A key contribution of our approach is the domain-specific fine-tuning of the Pyannote segmentation model on the competition dataset. This adaptation allowed the model to better capture the nuances of Bengali conversational dynamics and accurately resolve complex, overlapping speaker boundaries. Our methodology demonstrates that applying intelligent timestamped chunking to ASR and targeted segmentation fine-tuning to diarization significantly drives down Word Error Rate (WER) and Diarization Error Rate (DER), in low-resource settings.
【13】When Denoising Hinders: Revisiting Zero-Shot ASR with SAM-Audio and Whisper
标题:当去噪阻碍:用SAM音频和Whisper重新审视零拍ASR
链接:https://arxiv.org/abs/2603.04710
备注:6 pages, 4 figures, 5 tables. IEEE Conference Paper
摘要:自动语音识别(ASR)和语音增强的最新进展导致了一个广泛的假设,即提高感知音频质量应直接受益于识别精度。在这项工作中,我们严格审查是否这一假设适用于现代zero-shot ASR系统。我们提出了一个系统的实证研究的影响段任何模型音频由Meta AI,最近的基础规模的语音增强模型提出的Meta,当用作预处理步骤的zero-shot转录与耳语。实验在多个Whisper模型变体和两个语言上不同的嘈杂语音数据集上进行:一个真实世界的孟加拉语YouTube语料库和一个公开的英语嘈杂数据集。与常见的直觉相反,我们的研究结果表明,SAM音频预处理一贯降低ASR性能,增加字错误率(WER)和字符错误率(CER)相比,原始嘈杂的语音,尽管在信号电平的质量大幅改善。对英语数据集的客观峰值信噪比分析证实,SAM音频产生声学上更清晰的信号,但这种改进未能转化为识别增益。因此,我们进行了详细的话语层面的分析,以理解这一违反直觉的结果。我们发现,识别性能下降是一个系统性问题,会影响大多数音频,而不仅仅是孤立的离群值,并且随着Whisper模型大小的增加,错误会恶化。这些发现暴露了一个根本的不匹配:对人类听众来说感知上更干净的音频不一定对机器识别来说是鲁棒的。这突出了盲目应用最先进的去噪作为zero-shot ASR管道中的预处理步骤的风险。
摘要:Recent advances in automatic speech recognition (ASR) and speech enhancement have led to a widespread assumption that improving perceptual audio quality should directly benefit recognition accuracy. In this work, we rigorously examine whether this assumption holds for modern zero-shot ASR systems. We present a systematic empirical study on the impact of Segment Anything Model Audio by Meta AI, a recent foundation-scale speech enhancement model proposed by Meta, when used as a preprocessing step for zero-shot transcription with Whisper. Experiments are conducted across multiple Whisper model variants and two linguistically distinct noisy speech datasets: a real-world Bengali YouTube corpus and a publicly available English noisy dataset. Contrary to common intuition, our results show that SAM-Audio preprocessing consistently degrades ASR performance, increasing both Word Error Rate (WER) and Character Error Rate (CER) compared to raw noisy speech, despite substantial improvements in signal-level quality. Objective Peak Signal-to-Noise Ratio analysis on the English dataset confirms that SAM-Audio produces acoustically cleaner signals, yet this improvement fails to translate into recognition gains. Therefore, we conducted a detailed utterance-level analysis to understand this counterintuitive result. We found that the recognition degradation is a systematic issue affecting the majority of the audio, not just isolated outliers, and that the errors worsen as the Whisper model size increases. These findings expose a fundamental mismatch: audio that is perceptually cleaner to human listeners is not necessarily robust for machine recognition. This highlights the risk of blindly applying state-of-the-art denoising as a preprocessing step in zero-shot ASR pipelines.
【14】PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio
标题:PolyBench:复音音频合成推理的基准
链接:https://arxiv.org/abs/2603.05128
摘要:大型音频语言模型(LALM)越来越能够通过音频进行推理。然而,现有的基准提供有限的复音音频,其中多个声音事件共同发生,并诱导组成结构的推理覆盖。在这项工作中,我们介绍了PolyBench,一个基准,旨在评估合成推理复调音频。PolyBench包含五个评估子集,涵盖计数,分类,检测,并发和持续时间估计,需要对多个并发事件及其关系进行推理。对最先进的LALM的评估揭示了复调音频的一致性能下降,表明当前LALM存在根本瓶颈。
摘要:Large Audio Language Models (LALMs) are increasingly capable of reasoning over audio. However, existing benchmarks provide limited coverage of reasoning in polyphonic audio, where multiple sound events co-occur and induce compositional structure. In this work, we introduce PolyBench, a benchmark designed to evaluate compositional reasoning in polyphonic audio. PolyBench comprises five evaluation subsets covering counting, classification, detection, concurrency, and duration estimation, requiring reasoning over multiple concurrent events and their relations. Evaluation of state-of-the-art LALMs reveals consistent performance degradation in polyphonic audio, indicating a fundamental bottleneck in current LALMs.
【15】Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings
标题:具有自我监督音频嵌入的免训练异常声音检测的时间池策略
链接:https://arxiv.org/abs/2603.04605
摘要:基于预训练音频嵌入模型的免训练异常声音检测(ASD)最近引起了极大的关注,因为它能够仅使用正常参考数据检测异常声音,同时在域偏移下提供更好的鲁棒性。然而,现有的基于嵌入的方法几乎完全依赖于时间平均池,而替代池化策略到目前为止仅被探索用于基于谱图的表示。因此,时间池在具有预训练嵌入的无训练ASD中的作用仍然没有得到充分的理解。在本文中,我们提出了一个系统的评估跨多个国家的最先进的音频嵌入模型的时间池策略。我们提出了相对偏差池(RDP),自适应池化方法,强调信息的时间偏差,并介绍了一种混合池化策略,结合RDP与广义均值池化。在五个基准数据集上的实验表明,所提出的方法始终优于均值池,并实现了最先进的无训练ASD性能,包括超过DCASE2025 ASD数据集上所有先前报告的训练系统和集合的结果。
摘要:Training-free anomalous sound detection (ASD) based on pre-trained audio embedding models has recently garnered significant attention, as it enables the detection of anomalous sounds using only normal reference data while offering improved robustness under domain shifts. However, existing embedding-based approaches almost exclusively rely on temporal mean pooling, while alternative pooling strategies have so far only been explored for spectrogram-based representations. Consequently, the role of temporal pooling in training-free ASD with pre-trained embeddings remains insufficiently understood. In this paper, we present a systematic evaluation of temporal pooling strategies across multiple state-of-the-art audio embedding models. We propose relative deviation pooling (RDP), an adaptive pooling method that emphasizes informative temporal deviations, and introduce a hybrid pooling strategy that combines RDP with generalized mean pooling. Experiments on five benchmark datasets demonstrate that the proposed methods consistently outperform mean pooling and achieve state-of-the-art performance for training-free ASD, including results that surpass all previously reported trained systems and ensembles on the DCASE2025 ASD dataset.
【1】Visual-Informed Speech Enhancement Using Attention-Based Beamforming
标题:使用基于注意力的射束形成的视觉信息语音增强
链接:https://arxiv.org/abs/2603.05270
备注:15 pages, 14 figures
摘要:最近的研究表明,结合辅助信息,如扬声器声纹或视觉提示,可以大大提高语音增强(SE)的性能。然而,单通道方法在低信噪比(SNR)条件下,当存在高混响时,或者在涉及动态扬声器、重叠语音或非平稳噪声的复杂场景中,通常会产生次优结果。为了解决这些问题,我们提出了一种新的视觉信息神经波束形成网络(VI-NBFNet),它集成了麦克风阵列信号处理和使用多模态输入特征的深度神经网络(DNN)。该网络利用预训练的视觉语音识别模型提取嘴唇运动作为输入特征,用于语音活动检测(VAD)和目标说话人识别。该系统旨在通过引入配备有注意力机制的监督端到端波束成形框架来处理静态和移动扬声器。实验结果表明,所提出的视听系统具有更好的SE性能和鲁棒性的固定和动态的扬声器的情况下,相比几个基线方法。
摘要:Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often yield suboptimal results in low signal-to-noise ratio (SNR) conditions, when there is high reverberation, or in complex scenarios involving dynamic speakers, overlapping speech, or non-stationary noise. To address these issues, we propose a novel Visual-Informed Neural Beamforming Network (VI-NBFNet), which integrates microphone array signal processing and deep neural networks (DNNs) using multimodal input features. The proposed network leverages a pretrained visual speech recognition model to extract lip movements as input features, which serve for voice activity detection (VAD) and target speaker identification. The system is intended to handle both static and moving speakers by introducing a supervised end-to-end beamforming framework equipped with an attention mechanism. The experimental results demonstrated that the proposed audiovisual system has achieved better SE performance and robustness for both stationary and dynamic speaker scenarios, compared to several baseline methods.
【2】BabAR: from phoneme recognition to developmental measures of young children's speech production
标题:BabAR:从音素识别到幼儿言语产生的发展测量
链接:https://arxiv.org/abs/2603.05213
摘要:大规模研究早期语音发展需要自动化工具,但自动音素识别,特别是对幼儿来说,在很大程度上仍然没有解决。基于数十年的数据收集,我们策划了TinyVox,这是一个包含超过50万个英语,法语,葡萄牙语,德语和西班牙语语音转录儿童发声的语料库。我们使用TinyVox来训练BabAR,这是一个针对儿童语音的跨语言音素识别系统。我们发现,在以多语言儿童为中心的全天录音上对系统进行预训练的效果大大优于其他选择,并且在微调期间提供20秒的周围音频上下文进一步提高了性能。错误分析表明,替代主要属于相同的广泛的语音类别,这表明适合粗粒度的发展分析。我们验证了BabAR的语音成熟度的自动测量与文献中的发展估计。
摘要:Studying early speech development at scale requires automatic tools, yet automatic phoneme recognition, especially for young children, remains largely unsolved. Building on decades of data collection, we curate TinyVox, a corpus of more than half a million phonetically transcribed child vocalizations in English, French, Portuguese, German, and Spanish. We use TinyVox to train BabAR, a cross-linguistic phoneme recognition system for child speech. We find that pretraining the system on multilingual child-centered daylong recordings substantially outperforms alternatives, and that providing 20 seconds of surrounding audio context during fine-tuning further improves performance. Error analyses show that substitutions predominantly fall within the same broad phonetic categories, suggesting suitability for coarse-grained developmental analyses. We validate BabAR by showing that its automatic measures of speech maturity align with developmental estimates from the literature.
【3】PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio
标题:PolyBench:复音音频合成推理的基准
链接:https://arxiv.org/abs/2603.05128
摘要:大型音频语言模型(LALM)越来越能够通过音频进行推理。然而,现有的基准提供有限的复音音频,其中多个声音事件共同发生,并诱导组成结构的推理覆盖。在这项工作中,我们介绍了PolyBench,一个基准,旨在评估合成推理复调音频。PolyBench包含五个评估子集,涵盖计数,分类,检测,并发和持续时间估计,需要对多个并发事件及其关系进行推理。对最先进的LALM的评估揭示了复调音频的一致性能下降,表明当前LALM存在根本瓶颈。
摘要:Large Audio Language Models (LALMs) are increasingly capable of reasoning over audio. However, existing benchmarks provide limited coverage of reasoning in polyphonic audio, where multiple sound events co-occur and induce compositional structure. In this work, we introduce PolyBench, a benchmark designed to evaluate compositional reasoning in polyphonic audio. PolyBench comprises five evaluation subsets covering counting, classification, detection, concurrency, and duration estimation, requiring reasoning over multiple concurrent events and their relations. Evaluation of state-of-the-art LALMs reveals consistent performance degradation in polyphonic audio, indicating a fundamental bottleneck in current LALMs.
【4】Voice Timbre Attribute Detection with Compact and Interpretable Training-Free Acoustic Parameters
标题:具有紧凑且可解释的无需训练的声学参数的语音音色属性检测
链接:https://arxiv.org/abs/2603.05091
备注:Under review
摘要:语音音色属性检测是确定语音话语之间音色属性的相对强度的任务。语音音色是语音感知的一个重要而又复杂的组成部分。虽然深度神经网络(DNN)嵌入在说话人建模中表现良好,但它们通常作为黑箱表示,具有有限的物理可解释性和高计算成本。在这项工作中,一个紧凑的声学参数集的研究。这套捕捉重要的声学措施和他们的时间动态被发现是至关重要的任务。尽管简单,但声学参数集具有竞争力,优于传统的倒谱特征和监督DNN嵌入,并接近最先进的自监督模型。重要的是,所研究的集合不需要可训练的参数,产生可忽略的计算,并提供明确的可解释性,用于分析人类音色感知背后的物理特征。
摘要:Voice timbre attribute detection (vTAD) is the task of determining the relative intensity of timbre attributes between speech utterances. Voice timbre is a crucial yet inherently complex component of speech perception. While deep neural network (DNN) embeddings perform well in speaker modelling, they often act as black-box representations with limited physical interpretability and high computational cost. In this work, a compact acoustic parameter set is investigated for vTAD. The set captures important acoustic measures and their temporal dynamics which are found to be crucial in the task. Despite its simplicity, the acoustic parameter set is competitive, outperforming conventional cepstral features and supervised DNN embeddings, and approaching state-of-the-art self-supervised models. Importantly, the studied set require no trainable parameters, incur negligible computation, and offer explicit interpretability for analysing physical traits behind human timbre perception.
【5】An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production
标题:语音产生过程中关节、大脑和肌肉活动同时采集实时MRI视频、脑电和表面EMG的方法
链接:https://arxiv.org/abs/2603.04840
摘要:言语产生是一个复杂的过程,包括神经规划、运动控制、肌肉激活和发音运动学。虽然声学语音信号是语音产生行为最容易获得的产物,但它并不直接揭示其因果神经生理学基础。我们提出了实时(动态)MRI,EEG和表面EMG的第一次同时采集,捕捉语音产生链的几个关键方面:大脑信号,肌肉激活和发音运动。这种多模态采集模式提出了重大的技术挑战,包括MRI引起的电磁干扰和肌源性伪影。为了缓解这些问题,我们引入了一个针对这种三模态设置的伪影抑制管道。一旦完全开发,这个框架将为语音神经科学提供一个前所未有的窗口,并导致脑机接口的进步。
摘要:Speech production is a complex process spanning neural planning, motor control, muscle activation, and articulatory kinematics. While the acoustic speech signal is the most accessible product of the speech production act, it does not directly reveal its causal neurophysiological substrates. We present the first simultaneous acquisition of real-time (dynamic) MRI, EEG, and surface EMG, capturing several key aspects of the speech production chain: brain signals, muscle activations, and articulatory movements. This multimodal acquisition paradigm presents substantial technical challenges, including MRI-induced electromagnetic interference and myogenic artifacts. To mitigate these, we introduce an artifact suppression pipeline tailored to this tri-modal setting. Once fully developed, this framework is poised to offer an unprecedented window into speech neuroscience and insights leading to brain-computer interface advances.
【6】Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings
标题:具有自我监督音频嵌入的免训练异常声音检测的时间池策略
链接:https://arxiv.org/abs/2603.04605
摘要:基于预训练音频嵌入模型的免训练异常声音检测(ASD)最近引起了极大的关注,因为它能够仅使用正常参考数据检测异常声音,同时在域偏移下提供更好的鲁棒性。然而,现有的基于嵌入的方法几乎完全依赖于时间平均池,而替代池化策略到目前为止仅被探索用于基于谱图的表示。因此,时间池在具有预训练嵌入的无训练ASD中的作用仍然没有得到充分的理解。在本文中,我们提出了一个系统的评估跨多个国家的最先进的音频嵌入模型的时间池策略。我们提出了相对偏差池(RDP),自适应池化方法,强调信息的时间偏差,并介绍了一种混合池化策略,结合RDP与广义均值池化。在五个基准数据集上的实验表明,所提出的方法始终优于均值池,并实现了最先进的无训练ASD性能,包括超过DCASE2025 ASD数据集上所有先前报告的训练系统和集合的结果。
摘要:Training-free anomalous sound detection (ASD) based on pre-trained audio embedding models has recently garnered significant attention, as it enables the detection of anomalous sounds using only normal reference data while offering improved robustness under domain shifts. However, existing embedding-based approaches almost exclusively rely on temporal mean pooling, while alternative pooling strategies have so far only been explored for spectrogram-based representations. Consequently, the role of temporal pooling in training-free ASD with pre-trained embeddings remains insufficiently understood. In this paper, we present a systematic evaluation of temporal pooling strategies across multiple state-of-the-art audio embedding models. We propose relative deviation pooling (RDP), an adaptive pooling method that emphasizes informative temporal deviations, and introduce a hybrid pooling strategy that combines RDP with generalized mean pooling. Experiments on five benchmark datasets demonstrate that the proposed methods consistently outperform mean pooling and achieve state-of-the-art performance for training-free ASD, including results that surpass all previously reported trained systems and ensembles on the DCASE2025 ASD dataset.
【7】Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
标题:具有多分辨率欺骗检测的离散语音合成分层解码
链接:https://arxiv.org/abs/2603.05373
备注:7 pages, 3 figures, 3 tables, 2 algorithms
摘要:神经编解码器语言模型使高质量的离散语音合成,但他们的推理仍然容易受到令牌级文物和分布式漂移,降低感知的现实主义。而不是依赖于偏好优化或再训练,我们提出了MSpoof-TTS,一个无训练的推理框架,通过多分辨率欺骗指导提高了zero-shot合成。我们引入了一个基于多分辨率令牌的欺骗检测框架,该框架评估不同时间粒度的编解码器序列,以检测局部不一致或不自然的模式。然后,我们将欺骗检测器集成到一个分层解码策略中,逐步修剪低质量的候选人和重新排名的假设。这种鉴别器引导的生成在不修改模型参数的情况下增强了鲁棒性。实验验证了我们的框架的有效性,鲁棒性和高质量的基于编解码器的语音生成。
摘要:Neural codec language models enable high-quality discrete speech synthesis, yet their inference remains vulnerable to token-level artifacts and distributional drift that degrade perceptual realism. Rather than relying on preference optimization or retraining, we propose MSpoof-TTS, a training-free inference framework that improves zero-shot synthesis through multi-resolution spoof guidance. We introduce a Multi-Resolution Token-based Spoof Detection framework that evaluates codec sequences at different temporal granularities to detect locally inconsistent or unnatural patterns. We then integrate the spoof detectors into a hierarchical decoding strategy, progressively pruning low-quality candidates and re-ranking hypotheses. This discriminator-guided generation enhances robustness without modifying model parameters. Experiments validate the effectiveness of our framework for robust and high-quality codec-based speech generation.
【8】Exploring the potential and limitations of Model Merging for Multi-Domain Adaptation in ASR
标题:探索ASB中多域自适应模型合并的潜力和局限性
链接:https://arxiv.org/abs/2603.05354
备注:submitted for review for INTERSPEECH2026 conference
摘要:模型合并是多任务训练的可扩展替代方案,将多个专业模型的功能组合到单个模型中。这对于大型语音基础模型特别有吸引力,大型语音基础模型通常通过特定于域的微调进行调整,从而产生多个定制的检查点,对于这些检查点,当新数据变得可用时重复完全微调在计算上是禁止的。在这项工作中,我们研究了多域ASR的模型合并和基准11个合并算法为10个欧洲葡萄牙语域,评估域的准确性,鲁棒性下的分布转移,以及英语和多语言的性能。我们进一步提出了BoostedTSV-M,一个新的合并算法的基础上TSV-M,通过奇异值提升减轻秩崩溃,提高数值稳定性。总的来说,我们的方法优于欧洲葡萄牙语的全面微调,同时保留在一个单一的模型中的分布泛化。
摘要:Model merging is a scalable alternative to multi-task training that combines the capabilities of multiple specialised models into a single model. This is particularly attractive for large speech foundation models, which are typically adapted through domain-specific fine-tuning, resulting in multiple customised checkpoints, for which repeating full fine-tuning when new data becomes available is computationally prohibitive. In this work, we study model merging for multi-domain ASR and benchmark 11 merging algorithms for 10 European Portuguese domains, evaluating in-domain accuracy, robustness under distribution shift, as well as English and multilingual performance. We further propose BoostedTSV-M, a new merging algorithm based on TSV-M that mitigates rank collapse via singular-value boosting and improves numerical stability. Overall, our approach outperforms full fine-tuning on European Portuguese while preserving out-of-distribution generalisation in a single model.
机器翻译由腾讯交互翻译提供,仅供参考
