今日论文合集:cs.SD语音16篇,eess.AS音频处理19篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
标题:InconVAR:用于多模式情绪不一致检测的两级双塔框架
链接:https://arxiv.org/abs/2509.20140
作者:Zongyi Li, Junchuan Zhao, Francis Bu Sung Lee, Andrew Zi Han Yee
备注:5 pages, 1 figure, 3 tables
摘要:检测情感不一致的模态是情感计算的一个关键挑战,因为语音和文本往往传达相互冲突的线索。现有的方法通常依赖于不完整的情感表示和无条件融合,这削弱了性能时,模态是不一致的。此外,很少有以前的工作明确解决不一致检测本身。我们提出了InconVAD,一个两阶段的框架,在效价/唤醒/优势(VAD)空间接地。在第一阶段,独立的不确定性感知模型产生鲁棒的单峰预测。在第二阶段,分类器识别跨模态不一致性,并选择性地整合一致的信号。大量的实验表明,InconVAD在多模态情感不一致检测和建模方面都优于现有方法,为情感分析提供了更可靠和可解释的解决方案。
摘要:Detecting emotional inconsistency across modalities is a key challenge in affective computing, as speech and text often convey conflicting cues. Existing approaches generally rely on incomplete emotion representations and employ unconditional fusion, which weakens performance when modalities are inconsistent. Moreover, little prior work explicitly addresses inconsistency detection itself. We propose InconVAD, a two-stage framework grounded in the Valence/Arousal/Dominance (VAD) space. In the first stage, independent uncertainty-aware models yield robust unimodal predictions. In the second stage, a classifier identifies cross-modal inconsistency and selectively integrates consistent signals. Extensive experiments show that InconVAD surpasses existing methods in both multimodal emotion inconsistency detection and modeling, offering a more reliable and interpretable solution for emotion analysis.

【2】Enabling Multi-Species Bird Classification on Low-Power Bioacoustic Loggers
标题:在低功耗生物声学记录仪上实现多物种鸟类分类
链接:https://arxiv.org/abs/2509.20103
作者:Stefano Ciapponi, Leonardo Mannini, Jarek Scanferla, Matteo Anderle, Elisabetta Farella
摘要:WrenNet是一种高效的神经网络,可在低功耗微控制器上实现实时多物种鸟类音频分类,用于可扩展的生物多样性监测。我们提出了一个半可学习的光谱特征提取器,适应鸟类发声,优于标准梅尔规模和完全可学习的替代品。在专家策划的70个物种数据集上,WrenNet在声学上独特的物种上达到了90.8%的准确率,在整个任务上达到了70.1%。当部署在AudioMoth设备(1MB RAM)上时,每个推理仅消耗77 mJ。此外,在Raspberry Pi 3B+上运行时,该模型的能效比Birdnet高出16倍以上。这项工作展示了第一个在低功耗边缘设备上进行连续、多物种声学监测的实用框架。
摘要:This paper introduces WrenNet, an efficient neural network enabling real-time multi-species bird audio classification on low-power microcontrollers for scalable biodiversity monitoring. We propose a semi-learnable spectral feature extractor that adapts to avian vocalizations, outperforming standard mel-scale and fully-learnable alternatives. On an expert-curated 70-species dataset, WrenNet achieves up to 90.8\% accuracy on acoustically distinctive species and 70.1\% on the full task. When deployed on an AudioMoth device ($\leq$1MB RAM), it consumes only 77mJ per inference. Moreover, the proposed model is over 16x more energy-efficient compared to Birdnet when running on a Raspberry Pi 3B+. This work demonstrates the first practical framework for continuous, multi-species acoustic monitoring on low-power edge devices.

【3】MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
标题:MultiSoundGen:通过慢速对比视听预训练和直接偏好优化为多事件场景生成视频到音频
链接:https://arxiv.org/abs/2509.19999
作者:Jianxuan Yang, Xiaoran Yang, Lipan Zhang, Xinyue Guo, Zhao Wang, Gongping Huang
摘要:由于两个关键限制,当前的视频到音频(V2 A)方法在复杂的多事件场景(涉及多个声源、声音事件或过渡的视频场景)中挣扎。首先,现有的方法面临的挑战,精确地对齐复杂的语义信息与快速的动态功能。其次,基础训练缺乏对语义-时间对齐和音频质量的定量偏好优化。因此,它不能提高综合生成质量,在杂乱的多事件场景。为了解决这些核心限制,本研究提出了一种新的V2 A框架:MultiSoundGen。它将直接偏好优化(DPO)引入V2 A领域,利用视听预训练(AVP)来增强复杂多事件场景中的性能。我们的贡献包括两个关键创新:第一个是慢快对比AVP(SF-CAVP),这是一个具有统一双流架构的开创性AVP模型。SF-CAVP显式地将核心语义表示和视听数据的快速动态特征结合起来,以处理多事件复杂性;其次,我们将DPO方法集成到V2 A任务中,并提出AVP排名偏好优化(AVP-RPO)。它使用SF-CAVP作为奖励模型来量化和优先考虑关键的语义-时间匹配,同时提高音频质量。实验表明,MultiSoundGen在多事件场景中实现了最先进的(SOTA)性能,在分布匹配、音频质量、语义对齐和时间同步方面提供了全面的收益。完整的代码和数据集将很快发布。
摘要:Current video-to-audio (V2A) methods struggle in complex multi-event scenarios (video scenarios involving multiple sound sources, sound events, or transitions) due to two critical limitations. First, existing methods face challenges in precisely aligning intricate semantic information together with rapid dynamic features. Second, foundational training lacks quantitative preference optimization for semantic-temporal alignment and audio quality. As a result, it fails to enhance integrated generation quality in cluttered multi-event scenes. To address these core limitations, this study proposes a novel V2A framework: MultiSoundGen. It introduces direct preference optimization (DPO) into the V2A domain, leveraging audio-visual pretraining (AVP) to enhance performance in complex multi-event scenarios. Our contributions include two key innovations: the first is SlowFast Contrastive AVP (SF-CAVP), a pioneering AVP model with a unified dual-stream architecture. SF-CAVP explicitly aligns core semantic representations and rapid dynamic features of audio-visual data to handle multi-event complexity; second, we integrate the DPO method into V2A task and propose AVP-Ranked Preference Optimization (AVP-RPO). It uses SF-CAVP as a reward model to quantify and prioritize critical semantic-temporal matches while enhancing audio quality. Experiments demonstrate that MultiSoundGen achieves state-of-the-art (SOTA) performance in multi-event scenarios, delivering comprehensive gains across distribution matching, audio quality, semantic alignment, and temporal synchronization. The complete code and dataset will be released soon.

【4】CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
标题:CoMelSinger:基于离散令牌的Zero-Shot歌唱合成,具有结构化旋律控制和引导
链接:https://arxiv.org/abs/2509.19883
作者:Junchuan Zhao, Wei Zeng, Tianle Lyu, Ye Wang
备注:13 pages, 5 figures, 5 tables
摘要:歌唱声音合成(SVS)旨在从结构化的音乐输入(如歌词和音高序列)生成富有表现力的声乐表演。虽然基于离散编解码器的语音合成的最新进展已经通过上下文学习实现了zero-shot生成,但由于需要精确的旋律控制,将这些技术直接扩展到SVS仍然是不平凡的。特别是,基于音调的生成经常引入韵律泄漏,其中音调信息无意中纠缠在音色提示内,从而损害可控性。我们提出了CoMelSinger,一个zero-shot SVS框架,使结构化和分离的旋律控制在一个离散的编解码器建模范例。基于非自回归MaskGCT架构,CoMelSinger用歌词和音高标记取代了传统的文本输入,在增强旋律调节的同时保留了上下文泛化。为了抑制韵律泄漏,我们提出了一个由粗到细的对比学习策略,明确规范之间的声学提示和旋律输入的音高冗余。此外,我们还集成了一个轻量级的仅编码器的歌唱语音转录(SVT)模块,以将声学标记与音高和持续时间对齐,提供细粒度的帧级监督。实验结果表明,CoMelSinger在音高准确性、音色一致性和zero-shot可移植性方面都有显著的改善。
摘要:Singing Voice Synthesis (SVS) aims to generate expressive vocal performances from structured musical inputs such as lyrics and pitch sequences. While recent progress in discrete codec-based speech synthesis has enabled zero-shot generation via in-context learning, directly extending these techniques to SVS remains non-trivial due to the requirement for precise melody control. In particular, prompt-based generation often introduces prosody leakage, where pitch information is inadvertently entangled within the timbre prompt, compromising controllability. We present CoMelSinger, a zero-shot SVS framework that enables structured and disentangled melody control within a discrete codec modeling paradigm. Built on the non-autoregressive MaskGCT architecture, CoMelSinger replaces conventional text inputs with lyric and pitch tokens, preserving in-context generalization while enhancing melody conditioning. To suppress prosody leakage, we propose a coarse-to-fine contrastive learning strategy that explicitly regularizes pitch redundancy between the acoustic prompt and melody input. Furthermore, we incorporate a lightweight encoder-only Singing Voice Transcription (SVT) module to align acoustic tokens with pitch and duration, offering fine-grained frame-level supervision. Experimental results demonstrate that CoMelSinger achieves notable improvements in pitch accuracy, timbre consistency, and zero-shot transferability over competitive baselines.

【5】SEA-Spoof: Bridging The Gap in Multilingual Audio Deepfake Detection for South-East Asian
标题:SEA-Spoof:弥合东南亚多语言音频Deepfake检测的差距
链接:https://arxiv.org/abs/2509.19865
作者:Jinyang Wu, Nana Hou, Zihan Pan, Qiquan Zhang, Sailor Hardik Bhupendra, Soumik Mondal
备注:5 pages, 1 figure, 3 tables
摘要:东南亚(SEA)数字经济的快速增长放大了音频深度伪造的风险,但目前的数据集只覆盖了很少的SEA语言,使得模型无法处理这一关键地区。这一遗漏至关重要:在高资源语言上训练的检测模型在应用于SEA时会崩溃,这是由于合成质量,语言特定特征和数据稀缺性的不匹配。为了弥补这一差距,我们提出了SEA-Spoof,这是第一个专门针对SEA语言的大规模音频深度伪造检测(ADD)数据集。SEA-Spoof跨越泰米尔语,印地语,泰语,印度尼西亚语,马来语和越南语的300多个小时的配对真实和恶搞语音。欺骗样本是从最先进的开源和商业系统的多样化组合中生成的,捕捉了风格和保真度的广泛变化。对最先进的检测模型进行基准测试,发现存在严重的跨语言性能下降,但对SEA-Spoof进行微调后,可以显著恢复跨语言和合成源的性能。这些结果突出了对以SEA为重点的研究的迫切需要,并将SEA-Spoof建立为开发强大的,跨语言的和欺诈弹性检测系统的基础。
摘要:The rapid growth of the digital economy in South-East Asia (SEA) has amplified the risks of audio deepfakes, yet current datasets cover SEA languages only sparsely, leaving models poorly equipped to handle this critical region. This omission is critical: detection models trained on high-resource languages collapse when applied to SEA, due to mismatches in synthesis quality, language-specific characteristics, and data scarcity. To close this gap, we present SEA-Spoof, the first large-scale Audio Deepfake Detection (ADD) dataset especially for SEA languages. SEA-Spoof spans 300+ hours of paired real and spoof speech across Tamil, Hindi, Thai, Indonesian, Malay, and Vietnamese. Spoof samples are generated from a diverse mix of state-of-the-art open-source and commercial systems, capturing wide variability in style and fidelity. Benchmarking state-of-the-art detection models reveals severe cross-lingual degradation, but fine-tuning on SEA-Spoof dramatically restores performance across languages and synthesis sources. These results highlight the urgent need for SEA-focused research and establish SEA-Spoof as a foundation for developing robust, cross-lingual, and fraud-resilient detection systems.

【6】Eliminating stability hallucinations in llm-based tts models via attention guidance
标题:通过注意力引导消除基于LLM的tts模型中的稳定性幻觉
链接:https://arxiv.org/abs/2509.19852
作者:ShiMing Wang, ZhiHao Du, Yang Xiang, TianYu Zhao, Han Zhao, Qian Chen, XianGang Li, HanJie Guo, ZhenHua Ling
备注:5 pages, submitted to ICASSP2026
摘要:本文重点讨论解决稳定性幻觉(例如,重复或省略的语音)在基于LLM的文本到语音(TTS)模型通过改进和利用注意力机制。首先,我们分析了LLM中文本标记和语音标记之间的对齐机制。然后,我们提出了一个度量称为最佳对齐分数(OAS),它采用维特比算法来评估文本语音对齐质量。随后,OAS被集成到CosyVoice 2的培训中,以帮助LLM学习持续,稳定的对齐。此外,预先训练的注意力值用于通过思想链(CoT)指导学生CosyVoice 2的训练,这进一步减少了合成语音中的稳定性幻觉。在Seed-TTS-Eval和CV 3-Eval测试集上的实验表明,该方法能有效地降低CosyVoice 2的稳定性幻听,且不会引入额外的负面影响。附录见https://wsmzzz.github.io/llm_attn。
摘要:This paper focuses on resolving stability hallucinations (e.g., repetitive or omitted speech) in LLM-based Text-to-Speech (TTS) models by improving and leveraging the attention mechanism. First, we analyzed the alignment mechanism between text tokens and speech tokens in LLMs. We then proposed a metric termed the Optimal Alignment Score (OAS), which employs the Viterbi algorithm to evaluate text-speech alignment quality. Subsequently, OAS was integrated into the training of CosyVoice2 to assist LLMs in learning continuous, stable alignment. Additionally, the pre-trained attention value is employed to guide the training of the student CosyVoice2 via chain-of-thought (CoT), which further reduces stability hallucinations in synthesized speech. Experiments on the Seed-TTS-Eval and CV3-Eval test sets demonstrate that the proposed methods can effectively reduce the stability hallucinations of CosyVoice2 without introducing additional negative effects. The appendix is available at https://wsmzzz.github.io/llm_attn.

【7】Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
标题:通过渐进知识蒸馏实现语音合成的高效语音水印
链接:https://arxiv.org/abs/2509.19812
作者: Yang Cui, Peter Pan, Lei He, Sheng Zhao
备注:6 pages of main text, 1 page of references, 2 figures, 2 tables, accepted at ASRU 2025
摘要:随着语音生成模型的快速发展,未经授权的语音克隆带来了重大的隐私和安全风险。语音水印技术为追踪语音来源和防止滥用提供了一种可行的解决方案。目前的水印技术主要分为两类:基于DSP的方法和基于深度学习的方法。基于DSP的方法是有效的,但容易受到攻击,而基于深度学习的方法提供了强大的保护,代价是显着更高的计算成本。为了提高计算效率和增强鲁棒性,我们提出了PKDMark,一种轻量级的基于深度学习的语音水印方法,利用渐进知识蒸馏(PKD)。我们的方法分两个阶段进行:(1)使用基于可逆神经网络的架构训练高性能教师模型,以及(2)通过渐进式知识蒸馏将教师的能力转移到紧凑的学生模型。该过程将计算成本降低了93.6%,同时保持了高水平的鲁棒性能和不可感知性。实验结果表明,我们的蒸馏模型实现了平均检测F1分数为99.6%,PESQ为4.30的高级失真,使有效的语音水印的实时语音合成应用。
摘要:With the rapid advancement of speech generative models, unauthorized voice cloning poses significant privacy and security risks. Speech watermarking offers a viable solution for tracing sources and preventing misuse. Current watermarking technologies fall mainly into two categories: DSP-based methods and deep learning-based methods. DSP-based methods are efficient but vulnerable to attacks, whereas deep learning-based methods offer robust protection at the expense of significantly higher computational cost. To improve the computational efficiency and enhance the robustness, we propose PKDMark, a lightweight deep learning-based speech watermarking method that leverages progressive knowledge distillation (PKD). Our approach proceeds in two stages: (1) training a high-performance teacher model using an invertible neural network-based architecture, and (2) transferring the teacher's capabilities to a compact student model through progressive knowledge distillation. This process reduces computational costs by 93.6% while maintaining high level of robust performance and imperceptibility. Experimental results demonstrate that our distilled model achieves an average detection F1 score of 99.6% with a PESQ of 4.30 in advanced distortions, enabling efficient speech watermarking for real-time speech synthesis applications.

【8】Can Audio Large Language Models Verify Speaker Identity?
标题:音频大语言模型可以验证说话人身份吗?
链接:https://arxiv.org/abs/2509.19755
作者:Yiming Ren, Xuenan Xu, Baoxiang Li, Shuai Wang, Chao Zhang
摘要:本文研究了适应音频大语言模型(ALLM)的说话人确认(SV)。我们将SV重新定义为音频问答任务,并在公共基准上进行全面的zero-shot评估,表明当前ALLM具有有限的zero-shot SV能力,并且经常在不同的声学条件下挣扎。为了应对这一挑战,我们对说话人验证数据进行了监督微调。提出了一种基于规则的硬对采样策略来构造更具挑战性的训练对。轻量化的微调大大提高了性能,尽管ALLM和传统模型之间仍然存在差距。然后,我们扩展到文本相关的SV,通过联合查询ALLM来验证说话人身份和口语内容,从而产生与级联ASR-SV系统竞争的结果。我们的研究结果表明,通过适当的适应,ALLM具有强大的说话人验证系统的统一模型,同时保持一般的音频理解能力的巨大潜力。
摘要:This paper investigates adapting Audio Large Language Models (ALLMs) for speaker verification (SV). We reformulate SV as an audio question-answering task and conduct comprehensive zero-shot evaluations on public benchmarks, showing that current ALLMs have limited zero-shot SV capability and often struggle in diverse acoustic conditions. To address this challenge, we perform supervised fine-tuning on speaker verification data. A rule-based hard pair sampling strategy is proposed to construct more challenging training pairs. Lightweight fine-tuning substantially improves the performance, though there is still a gap between ALLMs and conventional models. Then, we extend to text-dependent SV by jointly querying ALLMs to verify speaker identity and spoken content, yielding results competitive with cascaded ASR-SV systems. Our findings demonstrate that with proper adaptation, ALLMs hold substantial potential as a unified model for robust speaker verification systems, while maintaining the general audio understanding capabilities.

【9】PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
标题:部分:使用LLM进行多语言语音转文本的渐进对齐表示训练
链接:https://arxiv.org/abs/2509.19745
作者:Pei Zhang, Andong Chen, Xi Chen, Baosong Yang, Derek F. Wong, Fei Huang
摘要:大型语言模型(LLM)已经从文本扩展到语音,从而产生了支持识别,翻译和合成的语音大型模型(SLM)。一个关键的挑战是对齐语音和文本表示,这在多语言环境中变得更加困难。现有的方法经常冻结LLM参数和训练编码器的多语言数据,但这迫使跨语言的收敛和限制性能。我们介绍了渐进对齐表示训练(PART),这是一个多阶段和多任务的框架,它将语言内对齐与跨语言对齐分开。在跨语言训练期间,LLM参数被动态激活,随后引入基于文本的任务以增强多语言理解。在CommonVoice 15、Fleurs、Wenetspeech和CoVoST 2上的实验表明,PART超越了传统方法,分析证实了其平衡语言特定差异和跨语言泛化的能力。这些结果表明,部分的有效性和通用性的多语种语音模态对齐。
摘要:Large language models (LLMs) have expanded from text to speech, giving rise to Speech Large Models (SLMs) that support recognition, translation, and synthesis. A key challenge is aligning speech and text representations, which becomes harder in multilingual settings. Existing methods often freeze LLM parameters and train encoders on multilingual data, but this forces cross-language convergence and limits performance. We introduce Progressive Alignment Representation Training (PART), a multi-stage and multi-task framework that separates within-language from cross-language alignment. During cross-language training, LLM parameters are dynamically activated, and text-based tasks are later introduced to enhance multilingual understanding. Experiments on CommonVoice 15, Fleurs, Wenetspeech, and CoVoST2 show that PART surpasses conventional approaches, with analysis confirming its ability to balance language-specific distinctions and cross-language generalization. These results demonstrate PART's effectiveness and generality for multilingual speech modality alignment.

【10】Thinking While Listening: Simple Test Time Scaling For Audio Classification
标题:边听边思考:音频分类的简单测试时间缩放
链接:https://arxiv.org/abs/2509.19676
作者:Prateek Verma, Mert Pilanci
备注:6 pages, 3 figures, 2 Tables, ICASSP 2026
摘要:我们提出了一个框架,使神经模型能够“边听边思考”日常声音,从而提高音频分类性能。受大型语言模型推理能力的最新进展的启发,我们解决了两个核心问题:(i)如何将思维纳入现有的音频分类管道中,以实现类别空间中的推理并提高性能,以及(ii)可以从头开始设计一个新的架构来支持思维和测试时间缩放?我们证明,在这两种设置中,我们的模型表现出更高的分类精度。利用测试时间缩放,我们观察到一致的增益作为采样迹线的数量增加。此外,我们评估了两个开源的推理模型,GPT-OSS-20 B和Qwen 3 - 14 B,表明虽然这些模型能够进行zero-shot推理,但一种轻量级的方法-只重新训练冻结的嵌入矩阵,如GPT-2-可以超越基于十亿参数文本的推理模型的性能。
摘要:We propose a framework that enables neural models to "think while listening" to everyday sounds, thereby enhancing audio classification performance. Motivated by recent advances in the reasoning capabilities of large language models, we address two central questions: (i) how can thinking be incorporated into existing audio classification pipelines to enable reasoning in the category space and improve performance, and (ii) can a new architecture be designed from the ground up to support both thinking and test-time scaling? We demonstrate that in both settings, our models exhibit improved classification accuracy. Leveraging test-time scaling, we observe consistent gains as the number of sampled traces increases. Furthermore, we evaluate two open-source reasoning models, GPT-OSS-20B and Qwen3-14B, showing that while such models are capable of zero-shot reasoning, a lightweight approach--retraining only the embedding matrix of a frozen, smaller model like GPT-2--can surpass the performance of billion-parameter text-based reasoning models.

【11】ArtiFree: Detecting and Reducing Generative Artifacts in Diffusion-based Speech Enhancement
标题:ArtiFree:检测和减少基于扩散的语音增强中的生成伪影
链接:https://arxiv.org/abs/2509.19495
作者:Bhawana Chhaglani, Yang Gao, Julius Richter, Xilin Li, Syavosh Zadissa, Tarun Pruthi, Andrew Lovitt
摘要:基于扩散的语音增强(SE)实现了听起来自然的语音和强泛化,但受到生成伪影和高推理延迟等关键限制。在这项工作中,我们系统地研究伪影预测和减少扩散为基础的SE。我们发现,语音嵌入的方差可以用来预测语音错误推理过程中。基于这些发现,我们提出了一个集成推理方法的语义一致性多个扩散运行的指导。该技术在低信噪比条件下将WER降低了15%,有效地提高了语音准确性和语义可扩展性。最后,我们分析了扩散步骤的数量的影响,表明自适应扩散步骤平衡伪影抑制和延迟。我们的研究结果强调语义先验作为一个强大的工具,引导生成SE走向无伪影输出。
摘要:Diffusion-based speech enhancement (SE) achieves natural-sounding speech and strong generalization, yet suffers from key limitations like generative artifacts and high inference latency. In this work, we systematically study artifact prediction and reduction in diffusion-based SE. We show that variance in speech embeddings can be used to predict phonetic errors during inference. Building on these findings, we propose an ensemble inference method guided by semantic consistency across multiple diffusion runs. This technique reduces WER by 15% in low-SNR conditions, effectively improving phonetic accuracy and semantic plausibility. Finally, we analyze the effect of the number of diffusion steps, showing that adaptive diffusion steps balance artifact suppression and latency. Our findings highlight semantic priors as a powerful tool to guide generative SE toward artifact-free outputs.

【12】MusiCRS: Benchmarking Audio-Centric Conversational Recommendation
标题:MusiCRS:以音频为中心的对话推荐基准
链接:https://arxiv.org/abs/2509.19469
作者:Rohan Surana, Amit Namburi, Gagan Mundada, Abhay Lal, Zachary Novack, Julian McAuley, Junda Wu
备注:6 pages
摘要:会话式推荐随着大型语言模型(LLM)的发展而迅速发展,但音乐仍然是一个具有独特挑战性的领域,其中有效的推荐需要对音频内容进行推理,而不仅仅是文本或元数据可以捕获的内容。我们提出了MusiCRS,这是第一个以音频为中心的会话推荐基准,它将Reddit中的真实用户对话与相应的音轨联系起来。MusiCRS包含477个高质量的对话,跨越不同的流派(古典,嘻哈,电子,金属,流行,独立,爵士),3,589个独特的音乐实体和音频基础通过YouTube链接。MusiCRS支持三种输入模态配置的评估:仅音频,仅查询和音频+查询(多模态),允许系统比较音频LLM,检索模型和传统方法。我们的实验表明,目前的系统严重依赖于文本信号,并与细致入微的音频推理作斗争。这暴露了跨模态知识集成的根本局限性,其中模型擅长对话语义,但不能有效地将抽象音乐概念与实际音频内容结合起来。为了促进进展,我们发布了MusiCRS数据集(https://huggingface.co/martets/rohan2810/MusiCRS),评估代码(https://github.com/rohan2810/musiCRS)和全面的基线。
摘要:Conversational recommendation has advanced rapidly with large language models (LLMs), yet music remains a uniquely challenging domain where effective recommendations require reasoning over audio content beyond what text or metadata can capture. We present MusiCRS, the first benchmark for audio-centric conversational recommendation that links authentic user conversations from Reddit with corresponding audio tracks. MusiCRS contains 477 high-quality conversations spanning diverse genres (classical, hip-hop, electronic, metal, pop, indie, jazz) with 3,589 unique musical entities and audio grounding via YouTube links. MusiCRS enables evaluation across three input modality configurations: audio-only, query-only, and audio+query (multimodal), allowing systematic comparison of audio-LLMs, retrieval models, and traditional approaches. Our experiments reveal that current systems rely heavily on textual signals and struggle with nuanced audio reasoning. This exposes fundamental limitations in cross-modal knowledge integration where models excel at dialogue semantics but cannot effectively ground abstract musical concepts in actual audio content. To facilitate progress, we release the MusiCRS dataset (https://huggingface.co/datasets/rohan2810/MusiCRS), evaluation code (https://github.com/rohan2810/musiCRS), and comprehensive baselines.

【13】MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
标题:MAGE:具有掩蔽生成模型的从粗到细的语音增强器
链接:https://arxiv.org/abs/2509.19881
作者:The Hieu Pham, Tan Dat Nguyen, Phuong Thanh Tran, Joon Son Chun, Duc Dung Nguyen
备注:Submitted to ICASSP 2026
摘要:由于效率和感知质量之间的权衡,语音增强仍然具有挑战性。在本文中,我们介绍了MAGE,一个掩蔽的音频生成增强器,通过紧凑和强大的设计,先进的生成语音增强。与之前使用随机掩蔽的掩蔽生成模型不同,MAGE采用了一种稀缺感知的粗到细掩蔽策略,该策略在早期步骤中优先考虑频繁令牌,在后期改进中优先考虑稀有令牌,从而提高了效率和泛化能力。我们还提出了一个轻量级的校正器模块,通过检测低置信度预测并重新掩蔽它们以进行改进来进一步稳定推理。MAGE基于BigCodec构建,并从Qwen2.5-0.5B进行了微调,通过选择性层保留将MAGE参数减少到200 M。DNS Challenge和带噪LibriSpeech的实验表明,MAGE实现了最先进的感知质量,并显着降低了下游识别的单词错误率,表现优于较大的基线。音频示例可在https://hieugiaosu.github.io/MAGE/上获得。
摘要:Speech enhancement remains challenging due to the trade-off between efficiency and perceptual quality. In this paper, we introduce MAGE, a Masked Audio Generative Enhancer that advances generative speech enhancement through a compact and robust design. Unlike prior masked generative models with random masking, MAGE employs a scarcity-aware coarse-to-fine masking strategy that prioritizes frequent tokens in early steps and rare tokens in later refinements, improving efficiency and generalization. We also propose a lightweight corrector module that further stabilizes inference by detecting low-confidence predictions and re-masking them for refinement. Built on BigCodec and finetuned from Qwen2.5-0.5B, MAGE is reduced to 200M parameters through selective layer retention. Experiments on DNS Challenge and noisy LibriSpeech show that MAGE achieves state-of-the-art perceptual quality and significantly reduces word error rate for downstream recognition, outperforming larger baselines. Audio examples are available at https://hieugiaosu.github.io/MAGE/.

【14】Non-locally averaged pruned reassigned spectrograms: a tool for glottal pulse visualization and analysis
标题:非局部平均修剪重新分配频谱图:喉舌脉搏可视化和分析的工具
链接:https://arxiv.org/abs/2509.19686
作者:Gabriel J. Griswold, Mark A. Griswold
备注:Submitted to Speech Communications. 16 pages, 7 figs, 1 table
摘要:重新分配的声谱图在精确的共振峰测量和说话人间区分方面显示出优势。然而,重新分配的频谱图无法以容易理解和可再现的方式可视化大量数据。利用Fulop和Fitz开发的技术和工具,提出了重新分配的频谱图的变体。非局部平均修剪重新分配频谱图(NAPReS)提供了一个简化的视图到扬声器的声门脉动模式的特性,通过堆叠,求和,并修剪大量的声门脉冲的元音的质心。在这项探索性研究中,NAPReS已被证明以易于理解和量化的方式显示大量数据,同时也使低振幅周期性结构的观察更容易获得。NAPReS还允许替代共振峰拟合方法,例如高斯混合建模。在这项研究中,NAPReS与GMM与传统的LPC拟合共振峰值进行了比较,并显示出比传统的LPC拟合在高噪声的情况下更可重复。
摘要:Reassigned spectrograms have shown advantages in precise formant measuring and inter-speaker differentiation. However, reassigned spectrograms suffer from their inability to visualize larger amounts of data in an easily comprehensible and reproducible manner. Utilizing the techniques and tools developed by Fulop and Fitz, a variation of the reassigned spectrogram is proposed. Non-locally Averaged Pruned Reassigned Spectrograms (NAPReS) provide a simplified view into the characteristics of a speaker's glottal pulsation patterns throughout the centroid of a vowel through the stacking, summing, and pruning of large numbers of glottal pulses. In this exploratory study, NAPReS has been shown to display a large amount of data in an easily comprehensible and quantifiable manner, while also making the observation of low-amplitude cyclical structures more accessible. NAPReS also allows for alternative formant fitting methods such as Gaussian mixture modeling. In this study, NAPReS with GMM was compared against conventional LPC fitting of formant values and was shown to be more reproducible than conventional LPC fitting in high-noise situations.

【15】Selective Classifier-free Guidance for Zero-shot Text-to-speech
标题:Zero-Shot文本到语音的选择性无分类器指南
链接:https://arxiv.org/abs/2509.19668
作者:John Zheng, Farhad Maleki
备注:5 pages, 7 figures, 1 table. Submitted to ICASSP 2026
摘要:在zero-shot文本到语音中,实现对目标说话者的保真度和对文本内容的坚持之间的平衡仍然是一个挑战。虽然无分类器引导(CFG)策略在图像生成方面已经显示出了很好的效果,但它们在语音合成中的应用还有待探索。分离用于CFG的条件使得能够在语音合成中的不同期望特性之间进行权衡。在本文中,我们评估的适应性的CFG策略最初开发的图像生成语音合成和扩展分离条件的CFG方法在这一领域。我们的研究结果表明,在图像生成中有效的CFG策略通常无法提高语音合成。我们还发现,我们可以提高扬声器的相似性,同时限制退化的文本坚持应用标准的CFG在早期的时间步长和切换到选择性的CFG只有在以后的时间步长。令人惊讶的是,我们观察到,选择性CFG策略的有效性是高度文本表示依赖的,因为即使使用相同的模型,英语和普通话这两种语言之间的差异也会导致不同的结果。
摘要:In zero-shot text-to-speech, achieving a balance between fidelity to the target speaker and adherence to text content remains a challenge. While classifier-free guidance (CFG) strategies have shown promising results in image generation, their application to speech synthesis are underexplored. Separating the conditions used for CFG enables trade-offs between different desired characteristics in speech synthesis. In this paper, we evaluate the adaptability of CFG strategies originally developed for image generation to speech synthesis and extend separated-condition CFG approaches for this domain. Our results show that CFG strategies effective in image generation generally fail to improve speech synthesis. We also find that we can improve speaker similarity while limiting degradation of text adherence by applying standard CFG during early timesteps and switching to selective CFG only in later timesteps. Surprisingly, we observe that the effectiveness of a selective CFG strategy is highly text-representation dependent, as differences between the two languages of English and Mandarin can lead to different results even with the same model.

【16】Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation
标题:用于高效多码本语音生成的帧堆叠本地Transformer
链接:https://arxiv.org/abs/2509.19592
作者:Roy Fejgin, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Ryan Langman Jaehyeon Kim, Subhankar Ghosh, Shehzeen Hussain, Jason Li
备注:This work has been submitted to the IEEE for possible publication
摘要:基于大语言模型(LLM)的语音生成模型通常对离散声学代码进行操作,由于其多码本结构,这些声学代码与文本令牌有着根本的不同。在每个时间步,模型必须联合预测N个码本条目,引入了挑战简单并行预测方法的依赖性。并行预测假设码本之间的独立性,从而产生有效的解码,但通常以降低保真度为代价。为了解决这个问题,分层策略采用本地Transformer(LT)来细化预测并捕获时间步内的依赖关系。在这项工作中,我们系统地研究了两个LT架构:自回归Transformer,顺序生成码本,和一个基于MaskGIT的Transformer,执行迭代掩码预测。这两种设计都进一步实现了帧堆叠,其中初级Transformer联合预测多个帧,并且LT解码它们的码本,从而在不损害感知质量的情况下提供速度上的改进。通过广泛的分析,我们的特点在不同的吞吐量和质量制度的并行和迭代采样策略之间的权衡。最后,我们提出了实用的指导方针,选择解码策略的基础上部署的优先级,如计算效率和合成保真度。
摘要:Speech generation models based on large language models (LLMs) typically operate on discrete acoustic codes, which differ fundamentally from text tokens due to their multicodebook structure. At each timestep, models must predict N codebook entries jointly, introducing dependencies that challenge simple parallel prediction approaches. Parallel prediction assumes independence among codebooks, yielding efficient decoding but often at the cost of reduced fidelity. To address this, hierarchical strategies employ a local transformer (LT) to refine predictions and capture intra-timestep dependencies. In this work, we systematically investigate two LT architectures: an autoregressive transformer that generates codebooks sequentially, and a MaskGIT-based transformer that performs iterative masked prediction. Both designs further enable frame stacking, where the primary transformer predicts multiple frames jointly, and the LT decodes their codebooks, offering improvements in speed without compromising perceptual quality. Through extensive analysis, we characterize the tradeoffs between parallel and iterative sampling strategies across different throughput and quality regimes. Finally, we propose practical guidelines for selecting decoding strategies based on deployment priorities such as computational efficiency and synthesis fidelity.


eess.AS音频处理


【1】Discrete Diffusion for Generative Modeling of Text-Aligned Speech Tokens
标题:文本对齐语音标记生成建模的离散扩散
链接:https://arxiv.org/abs/2509.20060
作者:Pin-Jui Ku, He Huang, Jean-Marie Lemercier, Subham Sekhar Sahoo, Zhehuai Chen, Ante Jukić
备注:5 pages. submitted to ICASSP 2026
摘要:本文介绍了一种离散扩散模型(DDM)框架的文本对齐语音标记和重建。通过将自回归语音解码器替换为离散扩散解码器,我们的模型实现了更好的重建质量,更强的ASR性能和更快的推理。我们提供了一个全面的分析应用DDMs语音重建,检查采样器的选择,推理步骤,和鲁棒性的长度尺度估计误差。此外,我们通过系统地比较矢量量化模块来改进原始TASTE,结果表明FSQ比AR模型的RVQ产生了高达35%的相对WER减少和+0.14 UT-MOS改进,同时还增强了DDM性能。我们的模型只需10个去噪步骤就可以生成语音,甚至支持单步生成,只有轻微的质量下降。
摘要:This paper introduces a discrete diffusion model (DDM) framework for text-aligned speech tokenization and reconstruction. By replacing the auto-regressive speech decoder with a discrete diffusion counterpart, our model achieves significantly better reconstruction quality, stronger ASR performance, and faster inference. We provide a comprehensive analysis of applying DDMs to speech reconstruction, examining sampler choices, inference steps, and robustness to length-scale estimation errors. Furthermore, we improve the original TASTE by systematically comparing vector quantization modules, showing that FSQ yields up to a 35% relative WER reduction and +0.14 UT-MOS improvement over RVQ for AR models, while also enhancing DDM performance. Our model generates speech in just 10 denoising steps and even supports single-step generation with only minor quality degradation.

【2】On the Invariance of Cross-Correlation Peak Positions Under Monotonic Signal Transformations, with Application to Fast Time Difference Estimation
标题:单调信号变换下互相关峰值位置的不变性及其在快速时差估计中的应用
链接:https://arxiv.org/abs/2509.19974
作者:Natsuki Ueno, Ryotaro Sato, Nobutaka Ono
摘要:我们提出了一个定理的互相关峰位置的不变性,这提供了一个基础的时间差估计的新方法,可能比传统的快速傅立叶变换(FFT)的方法为实/复序列。这一理论结果表明,在输入信号的任意单调变换下,两个移位离散时间信号之间的互相关函数的峰值位置保持不变。利用这一特性,我们设计了一个有效的估计算法的基础上的互相关函数的信号量化成低比特整数。所提出的方法只需要整数运算,而不是实值运算,并通过数论算法可以实现进一步的计算效率。数值实验表明,该方法实现了较短的处理时间比传统的基于FFT的方法。
摘要:We present a theorem concerning the invariance of cross-correlation peak positions, which provides a foundation for a new method for time difference estimation that is potentially faster than the conventional fast Fourier transform (FFT) approach for real/complex sequences. This theoretical result shows that the peak position of the cross-correlation function between two shifted discrete-time signals remains unchanged under arbitrary monotonic transformations of the input signals. By exploiting this property, we design an efficient estimation algorithm based on the cross-correlation function between signals quantized into low-bit integers. The proposed method requires only integer arithmetic instead of real-valued operations, and further computational efficiency can be achieved through number-theoretic algorithms. Numerical experiments demonstrate that the proposed method achieves a shorter processing time than conventional FFT-based approaches.

【3】Evaluating pretrained speech embedding systems for dysarthria detection across heterogenous datasets
标题:评估预训练的语音嵌入系统以用于跨异类数据集的构音障碍检测
链接:https://arxiv.org/abs/2509.19946
作者:Lovisa Wihlborg, Jemima Goodall, David Wheatley, Jacob J. Webber, Johnny Tam, Christine Weaver, Suvankar Pal, Siddharthan Chandran, Sohan Seth, Oliver Watts, Cassia Valentini-Botinhao
备注:Submitted to ICASSP 2026. This work is supported by NEURii, a collaborative partnership involving the University of Edinburgh, Gates Ventures, Eisai, LifeArc and Health Data Research UK (HDR UK)
摘要:我们提出了一个全面的评估预训练的语音嵌入系统,使用现有的可访问的数据检测构音障碍的语音。构音障碍的语音数据集通常很小,并且可能会受到记录偏差和数据不平衡的影响。为了解决这些问题,我们选择了一系列涵盖相关条件的数据集,并采用了几个交叉验证运行来估计机会水平。为了证明结果高于偶然性,我们将这些运行中的分数分布与精心设计的零假设的分数分布进行比较。通过这种方式,我们评估了6个不同数据集的17个公开可用的语音嵌入系统,并报告了每个系统的交叉验证性能。我们还报告了使用一个特定数据集进行训练并使用另一个数据集进行测试时得出的跨数据集结果。我们观察到,数据集内的结果根据数据集的不同而有很大的差异,无论使用的嵌入如何,这就提出了关于应该使用哪些数据集进行基准测试的问题。我们发现,正如预期的那样,跨数据集的准确性低于数据集内,突出了系统泛化的挑战。这些发现对在同一数据集上训练和测试的系统的临床有效性具有重要意义。
摘要:We present a comprehensive evaluation of pretrained speech embedding systems for the detection of dysarthric speech using existing accessible data. Dysarthric speech datasets are often small and can suffer from recording biases as well as data imbalance. To address these we selected a range of datasets covering related conditions and adopt the use of several cross-validations runs to estimate the chance level. To certify that results are above chance, we compare the distribution of scores across these runs against the distribution of scores of a carefully crafted null hypothesis. In this manner, we evaluate 17 publicly available speech embedding systems across 6 different datasets, reporting the cross-validation performance on each. We also report cross-dataset results derived when training with one particular dataset and testing with another. We observed that within-dataset results vary considerably depending on the dataset, regardless of the embedding used, raising questions about which datasets should be used for benchmarking. We found that cross-dataset accuracy is, as expected, lower than within-dataset, highlighting challenges in the generalization of the systems. These findings have important implications for the clinical validity of systems trained and tested on the same dataset.

【4】Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration
标题:测量Zero-ShotTTC中的韵律多样性:新指标、基准和探索
链接:https://arxiv.org/abs/2509.19928
作者:Yifan Yang, Bing Han, Hui Wang, Long Zhou, Wei Wang, Mingyu Cui, Xu Tan, Xie Chen
摘要:韵律多样性对于在零镜头文本到语音(zero-shot text to speech,TTS)中实现自然性和表现力至关重要。然而,经常使用的声学指标捕捉韵律变化的部分意见,并与人类的感知差,留下的问题,可靠地量化韵律多样性的探索。为了弥补这一差距,我们引入了ProsodyEval,一个韵律多样性评估数据集,提供韵律平均意见得分(PMOS)以及传统的声学指标。ProsodyEval包括来自7个主流TTS系统的1000个语音样本,2000人的评级。在此基础上,我们提出了离散语音加权编辑距离(DS-WED),这是一种新的客观多样性度量,通过语义标记上的加权编辑距离来量化韵律变化。ProsodyEval上的实验表明,DS-WED实现了更高的相关性与人类的判断比现有的声学指标,同时保持高度鲁棒的语音标记从HuBERT和WavLM。利用DS-WED,我们在LibriSpeech test-clean和Seed-TTS test-en上对最先进的开源TTS系统进行了基准测试,并进一步探索了影响韵律多样性的几个因素,包括生成建模范式,持续时间控制和强化学习。此外,我们发现,目前的大型音频语言模型(LALM)仍然有限,在捕捉韵律的变化。音频样本可在https://prosodyeval.github.io上获得。
摘要:Prosody diversity is essential for achieving naturalness and expressiveness in zero-shot text-to-speech (TTS). However, frequently used acoustic metrics capture only partial views of prosodic variation and correlate poorly with human perception, leaving the problem of reliably quantifying prosody diversity underexplored. To bridge this gap, we introduce ProsodyEval, a prosody diversity assessment dataset that provides Prosody Mean Opinion Score (PMOS) alongside conventional acoustic metrics. ProsodyEval comprises 1000 speech samples derived from 7 mainstream TTS systems, with 2000 human ratings. Building on this, we propose the Discretized Speech Weighted Edit Distance (DS-WED), a new objective diversity metric that quantifies prosodic variation via weighted edit distance over semantic tokens. Experiments on ProsodyEval show that DS-WED achieves substantially higher correlation with human judgments than existing acoustic metrics, while remaining highly robust in speech tokenization from HuBERT and WavLM. Leveraging DS-WED, we benchmark state-of-the-art open-source TTS systems on LibriSpeech test-clean and Seed-TTS test-en, and further explorations uncover several factors that influence prosody diversity, including generative modeling paradigms, duration control, and reinforcement learning. Moreover, we find that current large audio language models (LALMs) remain limited in capturing prosodic variations. Audio samples are available at https://prosodyeval.github.io.

【5】Voice Privacy Preservation with Multiple Random Orthogonal Secret Keys: Attack Resistance Analysis
标题:基于多随机正交密钥的语音隐私保护:抗攻击性分析
链接:https://arxiv.org/abs/2509.19906
作者:Kohei Tanaka, Hitoshi Kiya, Sayaka Shiota
备注:Submitted to APSIPA ASC 2025
摘要:最近,将语音数据传输到在云中执行的深度学习模型的机会增加了。这导致人们越来越关注语音隐私,包括说话人特定信息和话语的语言内容。作为一种语音隐私保护方法,提出了一种基于随机正交矩阵加密的语音隐私保护方法。该方法能够实现基于云的模型推断,同时隐藏语音内容和说话者身份。然而,该方法具有有限的抗攻击能力,并且在加密可以应用的深度学习模型方面受到限制。在这项工作中,我们提出了一种方法,通过采用多个随机正交矩阵作为密钥,提高了传统的语音隐私保护技术的抗攻击能力。我们还介绍了放松模型约束的方法,使我们的方法能够应用于更广泛的深度学习模型。此外,我们调查所提出的方法对攻击的鲁棒性,使用扩展的攻击场景的基础上,在语音隐私的挑战。我们的实验结果证实,所提出的方法保持隐私保护性能的扬声器隐藏,即使在更强大的攻击情况下没有考虑在以前的工作。
摘要:Recently, opportunities to transmit speech data to deep learning models executed in the cloud have increased. This has led to growing concerns about speech privacy, including both speaker-specific information and the linguistic content of utterances. As an approach to preserving speech privacy, a speech privacy-preserving method based on encryption using a secret key with a random orthogonal matrix has been proposed. This method enables cloud-based model inference while concealing both the speech content and the speaker identity. However, the method has limited attack resistance and is constrained in terms of the deep learning models to which the encryption can be applied. In this work, we propose a method that enhances the attack resistance of the conventional speech privacy-preserving technique by employing multiple random orthogonal matrices as secret keys. We also introduce approaches to relax the model constraints, enabling the application of our method to a broader range of deep learning models. Furthermore, we investigate the robustness of the proposed method against attacks using extended attack scenarios based on the scenarios employed in the Voice Privacy Challenge. Our experimental results confirmed that the proposed method maintains privacy protection performance for speaker concealment, even under more powerful attack scenarios not considered in prior work.

【6】MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
标题:MAGE:具有掩蔽生成模型的从粗到细的语音增强器
链接:https://arxiv.org/abs/2509.19881
作者:The Hieu Pham, Tan Dat Nguyen, Phuong Thanh Tran, Joon Son Chun, Duc Dung Nguyen
备注:Submitted to ICASSP 2026
摘要:由于效率和感知质量之间的权衡,语音增强仍然具有挑战性。在本文中,我们介绍了MAGE,一个掩蔽的音频生成增强器,通过紧凑和强大的设计,先进的生成语音增强。与之前使用随机掩蔽的掩蔽生成模型不同,MAGE采用了一种稀缺感知的粗到细掩蔽策略,该策略在早期步骤中优先考虑频繁令牌,在后期改进中优先考虑稀有令牌,从而提高了效率和泛化能力。我们还提出了一个轻量级的校正器模块,通过检测低置信度预测并重新掩蔽它们以进行改进来进一步稳定推理。MAGE基于BigCodec构建,并从Qwen2.5-0.5B进行了微调,通过选择性层保留将MAGE参数减少到200 M。DNS Challenge和带噪LibriSpeech的实验表明,MAGE实现了最先进的感知质量,并显着降低了下游识别的单词错误率,表现优于较大的基线。音频示例可在https://hieugiaosu.github.io/MAGE/上获得。
摘要:Speech enhancement remains challenging due to the trade-off between efficiency and perceptual quality. In this paper, we introduce MAGE, a Masked Audio Generative Enhancer that advances generative speech enhancement through a compact and robust design. Unlike prior masked generative models with random masking, MAGE employs a scarcity-aware coarse-to-fine masking strategy that prioritizes frequent tokens in early steps and rare tokens in later refinements, improving efficiency and generalization. We also propose a lightweight corrector module that further stabilizes inference by detecting low-confidence predictions and re-masking them for refinement. Built on BigCodec and finetuned from Qwen2.5-0.5B, MAGE is reduced to 200M parameters through selective layer retention. Experiments on DNS Challenge and noisy LibriSpeech show that MAGE achieves state-of-the-art perceptual quality and significantly reduces word error rate for downstream recognition, outperforming larger baselines. Audio examples are available at https://hieugiaosu.github.io/MAGE/.

【7】Weakly Supervised Phonological Features for Pathological Speech Analysis
标题:用于病理语音分析的弱监督语音特征
链接:https://arxiv.org/abs/2509.19879
作者:Jenthe Thienpondt, Geoffroy Vanderreydt, Abdessalem Hammami, Kris Demuynck
备注:proceedings of ICASSP 2025
摘要:言语的副语言特性对于分析和选择言语障碍患者的最佳治疗方案至关重要。然而,这些特征的自动建模是困难的,由于缺乏标记的语音数据集描述的非语言学属性,特别是在帧级。在本文中,我们提出了一种弱监督的训练方法,利用已知的声学特性的音素训练的ASR模型与可解释的帧级语音特征瓶颈层。随后,我们评估这些语音特征的可行性,在语音病理分析,通过开发相应的模型,可懂度预测和语音病理分类。使用我们提出的语音特征的模型在这两项任务上的分类准确率为75%,语音清晰度预测的RMSE为8.43,与其他最先进的声学特征相似。与其他人相比,我们的语音特征是文本无关的,高度可解释的,提供了潜在的有用的见解,言语治疗师。
摘要:Paralinguistic properties of speech are essential in analyzing and choosing optimal treatment options for patients with speech disorders. However, automatic modeling of these characteristics is difficult due to the lack of labeled speech datasets describing paralinguistic properties, especially at the frame-level. In this paper, we propose a weakly supervised training method which exploits the known acoustic properties of phonemes by training an ASR model with an interpretable frame-level phonological feature bottleneck layer. Subsequently, we assess the viability of these phonological features in speech pathology analysis by developing corresponding models for intelligibility prediction and speech pathology classification. Models using our proposed phonological features perform similar to other state-of-the-art acoustic features on both tasks with a classification accuracy of 75% and a 8.43 RMSE on speech intelligibility prediction. In contrast to others, our phonological features are text-independent and highly interpretable, providing potentially useful insights for speech therapists.

【8】SCORE: Scaling audio generation using Standardized COmposite REwards
标题:评分:使用标准化COmbosite奖励扩展音频生成
链接:https://arxiv.org/abs/2509.19831
作者:Jaemin Jung, Jaehun Kim, Inkyu Shin, Joon Son Chung
摘要:本文的目标是增强文本到音频生成的推理,重点是生成逼真的音频,精确地与文本提示。尽管取得了快速的进步,现有的模型往往无法实现感知质量和文本对齐之间的可靠平衡。为了解决这个问题,我们采用了推理时间缩放,这是一种无需训练的方法,通过增加推理计算来提高性能。我们建立了其未开发的应用程序音频生成,并提出了一种新的多奖励指导,同样意味着每个组件的感知至关重要。通过将每个奖励值归一化为一个共同的尺度,并将它们与加权求和相结合,该方法不仅可以实现稳定的制导,而且还可以实现显式控制,以达到期望的方面。此外,我们引入了一个新的音频文本对齐度量使用音频语言模型更强大的评估。从经验上讲,我们的方法提高了语义对齐和感知质量,显着优于天真的一代和现有的奖励指导技术。合成样品可在我们的演示页面上获得:https://mm.kaist.ac.kr/projects/score
摘要:The goal of this paper is to enhance Text-to-Audio generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advancements, existing models often fail to achieve a reliable balance between perceptual quality and textual alignment. To address this, we adopt Inference-Time Scaling, a training-free method that improves performance by increasing inference computation. We establish its unexplored application to audio generation and propose a novel multi-reward guidance that equally signifies each component essential in perception. By normalizing each reward value into a common scale and combining them with a weighted summation, the method not only enforces stable guidance but also enables explicit control to reach desired aspects. Moreover, we introduce a new audio-text alignment metric using an audio language model for more robust evaluation. Empirically, our method improves both semantic alignment and perceptual quality, significantly outperforming naive generation and existing reward guidance techniques. Synthesized samples are available on our demo page: https://mm.kaist.ac.kr/projects/score

【9】MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition
标题:MMedFD:用于多圈全速自动语音识别的现实医疗保健基准
链接:https://arxiv.org/abs/2509.19817
作者:Hongzhao Chen, XiaoYang Wang, Jing Lan, Hexiao Ding, Yufeng Jiang MingHui Yang, DanHui Xu, Jun Luo, Nga-Chun Ng, Gerald W.Y. Cheng, Yunlin Mao, Jung Sun Yoo
摘要:临床对话中的自动语音识别(ASR)要求对全双工交互、说话人重叠和低延迟约束具有鲁棒性,但开放的基准仍然很少。我们提出了MMedFD,第一个真实世界的中国医疗保健ASR语料库设计的多轮,全双工设置。该数据集从部署的AI助手捕获,包括5,805个带注释的会话,其中包含同步的用户和混合通道视图,RTTM/CTM时序和角色标签。我们引入了一个与模型无关的流水线,用于流分割,说话者属性和对话记忆,并微调Whisper-small on role-concatenated audio用于长上下文识别。ASR评估包括WER、CER和HC-WER,用于测量整个医疗保健环境中的概念级准确性。LLM生成的响应使用基于规则的和成对的方案进行评估。MMedFD建立了一个可复制的框架,用于对医疗保健部署中的流ASR和端到端双工代理进行基准测试。数据集和相关资源可在https://github.com/Kinetics-JOJO/MMedFD上公开获取
摘要:Automatic speech recognition (ASR) in clinical dialogue demands robustness to full-duplex interaction, speaker overlap, and low-latency constraints, yet open benchmarks remain scarce. We present MMedFD, the first real-world Chinese healthcare ASR corpus designed for multi-turn, full-duplex settings. Captured from a deployed AI assistant, the dataset comprises 5,805 annotated sessions with synchronized user and mixed-channel views, RTTM/CTM timing, and role labels. We introduce a model-agnostic pipeline for streaming segmentation, speaker attribution, and dialogue memory, and fine-tune Whisper-small on role-concatenated audio for long-context recognition. ASR evaluation includes WER, CER, and HC-WER, which measures concept-level accuracy across healthcare settings. LLM-generated responses are assessed using rubric-based and pairwise protocols. MMedFD establishes a reproducible framework for benchmarking streaming ASR and end-to-end duplex agents in healthcare deployment. The dataset and related resources are publicly available at https://github.com/Kinetics-JOJO/MMedFD

【10】Short-Segment Speaker Verification with Pre-trained Models and Multi-Resolution Encoder
标题:使用预训练模型和多分辨率编码器进行短段说话人验证
链接:https://arxiv.org/abs/2509.19721
作者:Jisoo Myoung, Sangwook Han, Kihyuk Kim, Jong Won Shin
备注:Submitted to ICASSP 2026
摘要:说话人确认(SV)利用通过自监督学习预训练的模型获得的特征,最近表现出令人印象深刻的性能。然而,这些预训练模型(PTM)通常具有20 ms的时间分辨率,低于典型的滤波器组特征。这可能是有问题的,特别是对于输入段短于2 s的短段SV,其中我们需要从有限长度的输入中提取尽可能多的信息。虽然已经有方法利用来自HuBERT模型的多分辨率特征,但是当采样率为16 kHz时,窗口移位为320、640和1600个样本,因此仅考虑较低分辨率的特征。在这项研究中,我们提出了一个SV系统,它利用PTM功能以及滤波器组功能和那些从多分辨率时域编码器的窗口移位25,50,100,和200个样本。在具有各种输入长度的VoxCeleb数据集上的实验结果表明,与具有各种输入特征组合的系统相比,系统得到了一致的改进。
摘要:Speaker verification (SV) utilizing features obtained from models pre-trained via self-supervised learning has recently demonstrated impressive performances. However, these pre-trained models (PTMs) usually have a temporal resolution of 20 ms, which is lower than typical filterbank features. It may be problematic especially for short-segment SV with an input segment shorter than 2 s, in which we need to extract as much information as possible from the input with a limited length. Although there have been approaches to utilize multi-resolution features from the HuBERT models, the window shifts were 320, 640, and 1600 samples when the sampling rate was 16 kHz and thus only lower resolution features were considered. In this study, we propose an SV system which utilizes PTM features along with filterbank features and those from the multi-resolution time domain encoder with window shifts of 25, 50, 100, and 200 samples. Experimental results on the VoxCeleb dataset with various input lengths showed consistent improvements over systems with various combinations of input features.

【11】Non-locally averaged pruned reassigned spectrograms: a tool for glottal pulse visualization and analysis
标题:非局部平均修剪重新分配频谱图:喉舌脉搏可视化和分析的工具
链接:https://arxiv.org/abs/2509.19686
作者:Gabriel J. Griswold, Mark A. Griswold
备注:Submitted to Speech Communications. 16 pages, 7 figs, 1 table
摘要:重新分配的声谱图在精确的共振峰测量和说话人间区分方面显示出优势。然而,重新分配的频谱图无法以容易理解和可再现的方式可视化大量数据。利用Fulop和Fitz开发的技术和工具,提出了重新分配的频谱图的变体。非局部平均修剪重新分配频谱图(NAPReS)提供了一个简化的视图到扬声器的声门脉动模式的特性,通过堆叠,求和,并修剪大量的声门脉冲的元音的质心。在这项探索性研究中,NAPReS已被证明以易于理解和量化的方式显示大量数据,同时也使低振幅周期性结构的观察更容易获得。NAPReS还允许替代共振峰拟合方法,例如高斯混合建模。在这项研究中,NAPReS与GMM与传统的LPC拟合共振峰值进行了比较,并显示出比传统的LPC拟合在高噪声的情况下更可重复。
摘要:Reassigned spectrograms have shown advantages in precise formant measuring and inter-speaker differentiation. However, reassigned spectrograms suffer from their inability to visualize larger amounts of data in an easily comprehensible and reproducible manner. Utilizing the techniques and tools developed by Fulop and Fitz, a variation of the reassigned spectrogram is proposed. Non-locally Averaged Pruned Reassigned Spectrograms (NAPReS) provide a simplified view into the characteristics of a speaker's glottal pulsation patterns throughout the centroid of a vowel through the stacking, summing, and pruning of large numbers of glottal pulses. In this exploratory study, NAPReS has been shown to display a large amount of data in an easily comprehensible and quantifiable manner, while also making the observation of low-amplitude cyclical structures more accessible. NAPReS also allows for alternative formant fitting methods such as Gaussian mixture modeling. In this study, NAPReS with GMM was compared against conventional LPC fitting of formant values and was shown to be more reproducible than conventional LPC fitting in high-noise situations.

【12】Selective Classifier-free Guidance for Zero-shot Text-to-speech
标题:Zero-Shot文本到语音的选择性无分类器指南
链接:https://arxiv.org/abs/2509.19668
作者:John Zheng, Farhad Maleki
备注:5 pages, 7 figures, 1 table. Submitted to ICASSP 2026
摘要:在zero-shot文本到语音中,实现对目标说话者的保真度和对文本内容的坚持之间的平衡仍然是一个挑战。虽然无分类器引导(CFG)策略在图像生成方面已经显示出了很好的效果,但它们在语音合成中的应用还有待探索。分离用于CFG的条件使得能够在语音合成中的不同期望特性之间进行权衡。在本文中,我们评估的适应性的CFG策略最初开发的图像生成语音合成和扩展分离条件的CFG方法在这一领域。我们的研究结果表明,在图像生成中有效的CFG策略通常无法提高语音合成。我们还发现,我们可以提高扬声器的相似性,同时限制退化的文本坚持应用标准的CFG在早期的时间步长和切换到选择性的CFG只有在以后的时间步长。令人惊讶的是,我们观察到,选择性CFG策略的有效性是高度文本表示依赖的,因为即使使用相同的模型,英语和普通话这两种语言之间的差异也会导致不同的结果。
摘要:In zero-shot text-to-speech, achieving a balance between fidelity to the target speaker and adherence to text content remains a challenge. While classifier-free guidance (CFG) strategies have shown promising results in image generation, their application to speech synthesis are underexplored. Separating the conditions used for CFG enables trade-offs between different desired characteristics in speech synthesis. In this paper, we evaluate the adaptability of CFG strategies originally developed for image generation to speech synthesis and extend separated-condition CFG approaches for this domain. Our results show that CFG strategies effective in image generation generally fail to improve speech synthesis. We also find that we can improve speaker similarity while limiting degradation of text adherence by applying standard CFG during early timesteps and switching to selective CFG only in later timesteps. Surprisingly, we observe that the effectiveness of a selective CFG strategy is highly text-representation dependent, as differences between the two languages of English and Mandarin can lead to different results even with the same model.

【13】Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
标题:利用强化学习推进多模式LLM中的语音摘要
链接:https://arxiv.org/abs/2509.19631
作者:Shaoshi Ling, Gang Liu, Guoli Ye, Jinyu Li
摘要:语音摘要是口语内容理解的关键组成部分,特别是在口语和视听数据快速增长的时代。多模态大型语言模型(MLLM)的最新进展,利用LLM的强大功能,可以直接从语音中生成文本摘要,而无需中间转换,同时支持可控风格和zero-shot泛化。然而,开源MLLM仍然落后于最先进的基于文本的LLM,限制了它们在语音摘要中的实际部署。在这项工作中,我们提出了一种新的多阶段强化学习训练框架,以提高MLLM的语音摘要能力。我们的模型在强大的基线上提供了实质性的改进,优于更大的MLLM,并显着缩小了与最先进的基于文本的LLM的差距。
摘要:Speech summarization is a critical component of spoken content understanding, particularly in the era of rapidly growing spoken and audiovisual data. Recent advances in multi-modal large language models (MLLMs), leveraging the power of LLMs, enable generating textual summaries directly from speech without intermediate transcriptions, while supporting controllable styles and zero-shot generalization. However, open-source MLLMs continue to lag behind the state-of-the-art text-based LLMs, limiting their practical deployment for speech summarization. In this work, we present a novel multi-stage reinforcement learning training framework to enhance the speech summarization capabilities in MLLMs. Our model delivers substantial improvements over strong baselines, outperforms much larger MLLMs, and significantly narrows the gap with state-of-the-art text-based LLMs.

【14】Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation
标题:用于高效多码本语音生成的帧堆叠本地Transformer
链接:https://arxiv.org/abs/2509.19592
作者:Roy Fejgin, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Ryan Langman Jaehyeon Kim, Subhankar Ghosh, Shehzeen Hussain, Jason Li
备注:This work has been submitted to the IEEE for possible publication
摘要:基于大语言模型(LLM)的语音生成模型通常对离散声学代码进行操作,由于其多码本结构,这些声学代码与文本令牌有着根本的不同。在每个时间步,模型必须联合预测N个码本条目,引入了挑战简单并行预测方法的依赖性。并行预测假设码本之间的独立性,从而产生有效的解码,但通常以降低保真度为代价。为了解决这个问题,分层策略采用本地Transformer(LT)来细化预测并捕获时间步内的依赖关系。在这项工作中,我们系统地研究了两个LT架构:自回归Transformer,顺序生成码本,和一个基于MaskGIT的Transformer,执行迭代掩码预测。这两种设计都进一步实现了帧堆叠,其中初级Transformer联合预测多个帧,并且LT解码它们的码本,从而在不损害感知质量的情况下提供速度上的改进。通过广泛的分析,我们的特点在不同的吞吐量和质量制度的并行和迭代采样策略之间的权衡。最后,我们提出了实用的指导方针,选择解码策略的基础上部署的优先级,如计算效率和合成保真度。
摘要:Speech generation models based on large language models (LLMs) typically operate on discrete acoustic codes, which differ fundamentally from text tokens due to their multicodebook structure. At each timestep, models must predict N codebook entries jointly, introducing dependencies that challenge simple parallel prediction approaches. Parallel prediction assumes independence among codebooks, yielding efficient decoding but often at the cost of reduced fidelity. To address this, hierarchical strategies employ a local transformer (LT) to refine predictions and capture intra-timestep dependencies. In this work, we systematically investigate two LT architectures: an autoregressive transformer that generates codebooks sequentially, and a MaskGIT-based transformer that performs iterative masked prediction. Both designs further enable frame stacking, where the primary transformer predicts multiple frames jointly, and the LT decodes their codebooks, offering improvements in speed without compromising perceptual quality. Through extensive analysis, we characterize the tradeoffs between parallel and iterative sampling strategies across different throughput and quality regimes. Finally, we propose practical guidelines for selecting decoding strategies based on deployment priorities such as computational efficiency and synthesis fidelity.

【15】DRES: Benchmarking LLMs for Disfluency Removal
标题:DRES:对LLM进行不流利消除的基准
链接:https://arxiv.org/abs/2509.20321
作者:Maria Teleki, Sai Janjur, Haoran Liu, Oliver Grabner, Ketan Verma, Thomas Docog, Xiangjue Dong, Lingfeng Shi, Cong Wang, Stephanie Birkelbach, Jason Kim, Yin Zhang, James Caverlee
摘要:不流利-例如“嗯”、“呃”、感叹词、插入语和编辑语句-仍然是语音驱动系统的持续挑战,降低了命令解释、摘要和会话代理的准确性。我们介绍DRES(不流畅消除评估套件),一个受控的文本级基准,建立了一个可重复的语义上限这项任务。DRES建立在人类注释的Switchboard转录本上,将不流利消除与ASR错误和声学变化隔离开来。我们系统地评估专有和开源LLM的规模,提示策略和架构。我们的研究结果表明,(i)简单的分割始终提高性能,即使是长上下文模型;(ii)面向推理的模型往往会过度删除流畅的标记;(iii)微调实现了接近最先进的精度和召回率,但损害了泛化能力。我们进一步提出了一组特定于LLM的错误模式,并提供了九个实用的建议(R1-R9),用于在语音驱动的管道中部署不流利消除。DRES为推进健壮的口语系统提供了一个可重复的、与模型无关的基础。
摘要:Disfluencies -- such as "um," "uh," interjections, parentheticals, and edited statements -- remain a persistent challenge for speech-driven systems, degrading accuracy in command interpretation, summarization, and conversational agents. We introduce DRES (Disfluency Removal Evaluation Suite), a controlled text-level benchmark that establishes a reproducible semantic upper bound for this task. DRES builds on human-annotated Switchboard transcripts, isolating disfluency removal from ASR errors and acoustic variability. We systematically evaluate proprietary and open-source LLMs across scales, prompting strategies, and architectures. Our results reveal that (i) simple segmentation consistently improves performance, even for long-context models; (ii) reasoning-oriented models tend to over-delete fluent tokens; and (iii) fine-tuning achieves near state-of-the-art precision and recall but harms generalization abilities. We further present a set of LLM-specific error modes and offer nine practical recommendations (R1-R9) for deploying disfluency removal in speech-driven pipelines. DRES provides a reproducible, model-agnostic foundation for advancing robust spoken-language systems.

【16】Z-Scores: A Metric for Linguistically Assessing Disfluency Removal
标题:Z分数:语言上评估不流利消除的指标
链接:https://arxiv.org/abs/2509.20319
作者:Maria Teleki, Sai Janjur, Haoran Liu, Oliver Grabner, Ketan Verma, Thomas Docog, Xiangjue Dong, Lingfeng Shi, Cong Wang, Stephanie Birkelbach, Jason Kim, Yin Zhang, James Caverlee
摘要:评估言语中的不流利消除需要的不仅仅是总的标记级分数。传统的基于单词的指标,如精确度、召回率和F1(E分数),可以捕捉整体性能,但无法揭示模型成功或失败的原因。我们引入了Z分数,这是一个跨语言水平的评估指标,它将系统行为分为不同的不流利类型(EDITED,INTJ,PRN)。我们的确定性对齐模块可以在生成的文本和不流利的成绩单之间实现强大的映射,从而使Z分数能够暴露单词级指标所掩盖的系统性弱点。通过提供特定类别的诊断,Z-Scores使研究人员能够识别模型故障模式并设计有针对性的干预措施-例如量身定制的提示或数据增强-从而产生可衡量的性能改进。LLM的案例研究表明,Z分数揭示了隐藏在聚合F1中的INTJ和PRN不流利的挑战,直接为模型优化策略提供信息。
摘要:Evaluating disfluency removal in speech requires more than aggregate token-level scores. Traditional word-based metrics such as precision, recall, and F1 (E-Scores) capture overall performance but cannot reveal why models succeed or fail. We introduce Z-Scores, a span-level linguistically-grounded evaluation metric that categorizes system behavior across distinct disfluency types (EDITED, INTJ, PRN). Our deterministic alignment module enables robust mapping between generated text and disfluent transcripts, allowing Z-Scores to expose systematic weaknesses that word-level metrics obscure. By providing category-specific diagnostics, Z-Scores enable researchers to identify model failure modes and design targeted interventions -- such as tailored prompts or data augmentation -- yielding measurable performance improvements. A case study with LLMs shows that Z-Scores uncover challenges with INTJ and PRN disfluencies hidden in aggregate F1, directly informing model refinement strategies.

【17】Can Audio Large Language Models Verify Speaker Identity?
标题:音频大语言模型可以验证说话人身份吗?
链接:https://arxiv.org/abs/2509.19755
作者:Yiming Ren, Xuenan Xu, Baoxiang Li, Shuai Wang, Chao Zhang
摘要:本文研究了适应音频大语言模型(ALLM)的说话人确认(SV)。我们将SV重新定义为音频问答任务,并在公共基准上进行全面的zero-shot评估,表明当前ALLM具有有限的zero-shot SV能力,并且经常在不同的声学条件下挣扎。为了应对这一挑战,我们对说话人验证数据进行了监督微调。提出了一种基于规则的硬对采样策略来构造更具挑战性的训练对。轻量化的微调大大提高了性能,尽管ALLM和传统模型之间仍然存在差距。然后,我们扩展到文本相关的SV,通过联合查询ALLM来验证说话人身份和口语内容,从而产生与级联ASR-SV系统竞争的结果。我们的研究结果表明,通过适当的适应,ALLM具有强大的说话人验证系统的统一模型,同时保持一般的音频理解能力的巨大潜力。
摘要:This paper investigates adapting Audio Large Language Models (ALLMs) for speaker verification (SV). We reformulate SV as an audio question-answering task and conduct comprehensive zero-shot evaluations on public benchmarks, showing that current ALLMs have limited zero-shot SV capability and often struggle in diverse acoustic conditions. To address this challenge, we perform supervised fine-tuning on speaker verification data. A rule-based hard pair sampling strategy is proposed to construct more challenging training pairs. Lightweight fine-tuning substantially improves the performance, though there is still a gap between ALLMs and conventional models. Then, we extend to text-dependent SV by jointly querying ALLMs to verify speaker identity and spoken content, yielding results competitive with cascaded ASR-SV systems. Our findings demonstrate that with proper adaptation, ALLMs hold substantial potential as a unified model for robust speaker verification systems, while maintaining the general audio understanding capabilities.

【18】Thinking While Listening: Simple Test Time Scaling For Audio Classification
标题:边听边思考:音频分类的简单测试时间缩放
链接:https://arxiv.org/abs/2509.19676
作者:Prateek Verma, Mert Pilanci
备注:6 pages, 3 figures, 2 Tables, ICASSP 2026
摘要:我们提出了一个框架,使神经模型能够“边听边思考”日常声音,从而提高音频分类性能。受大型语言模型推理能力的最新进展的启发,我们解决了两个核心问题:(i)如何将思维纳入现有的音频分类管道中,以实现类别空间中的推理并提高性能,以及(ii)可以从头开始设计一个新的架构来支持思维和测试时间缩放?我们证明,在这两种设置中,我们的模型表现出更高的分类精度。利用测试时间缩放,我们观察到一致的增益作为采样迹线的数量增加。此外,我们评估了两个开源的推理模型,GPT-OSS-20 B和Qwen 3 - 14 B,表明虽然这些模型能够进行zero-shot推理,但一种轻量级的方法-只重新训练冻结的嵌入矩阵,如GPT-2-可以超越基于十亿参数文本的推理模型的性能。
摘要:We propose a framework that enables neural models to "think while listening" to everyday sounds, thereby enhancing audio classification performance. Motivated by recent advances in the reasoning capabilities of large language models, we address two central questions: (i) how can thinking be incorporated into existing audio classification pipelines to enable reasoning in the category space and improve performance, and (ii) can a new architecture be designed from the ground up to support both thinking and test-time scaling? We demonstrate that in both settings, our models exhibit improved classification accuracy. Leveraging test-time scaling, we observe consistent gains as the number of sampled traces increases. Furthermore, we evaluate two open-source reasoning models, GPT-OSS-20B and Qwen3-14B, showing that while such models are capable of zero-shot reasoning, a lightweight approach--retraining only the embedding matrix of a frozen, smaller model like GPT-2--can surpass the performance of billion-parameter text-based reasoning models.

【19】Retrieval Augmented Generation based context discovery for ASR
标题:基于检索增强生成的ASB上下文发现
链接:https://arxiv.org/abs/2509.19567
作者:Dimitrios Siskos, Stavros Papadopoulos, Pablo Peso Parada, Jisi Zhang, Karthikeyan Saravanan, Anastasios Drosou
备注:Accepted at EMNLP 2025
摘要:本文研究了在上下文感知的自动语音识别(ASR)系统中,检索增强生成作为一种有效的上下文自动发现策略,以提高在存在稀有或词汇表外术语的情况下的转录准确率。然而,自动识别正确的上下文仍然是一个开放的挑战。本文提出了一种基于嵌入的自动上下文发现方法。为了使其有效性上下文中,还评估了基于大语言模型(LLM)的两种替代方案:(1)通过提示的基于大语言模型(LLM)的上下文生成,以及(2)使用LLM的识别后转录校正。TED-LIUMv 3,Earnings 21和SPGISpeech上的实验表明,相对于使用无上下文,所提出的方法将WER降低了17%(百分比差异),而Oracle上下文导致降低了24.1%。
摘要:This work investigates retrieval augmented generation as an efficient strategy for automatic context discovery in context-aware Automatic Speech Recognition (ASR) system, in order to improve transcription accuracy in the presence of rare or out-of-vocabulary terms. However, identifying the right context automatically remains an open challenge. This work proposes an efficient embedding-based retrieval approach for automatic context discovery in ASR. To contextualize its effectiveness, two alternatives based on large language models (LLMs) are also evaluated: (1) large language model (LLM)-based context generation via prompting, and (2) post-recognition transcript correction using LLMs. Experiments on the TED-LIUMv3, Earnings21 and SPGISpeech demonstrate that the proposed approach reduces WER by up to 17% (percentage difference) relative to using no-context, while the oracle context results in a reduction of up to 24.1%.


机器翻译由腾讯交互翻译提供,仅供参考