微信公众号:arXiv_Daily
cs.SD语音
标题: MIKU-PAL:一种用于语音副语言和情感标签的自动化和标准化多模式方法
链接:https://arxiv.org/abs/2505.15772
备注:Accepted by Interspeech
摘要:获取大规模的具有强一致性的情感语音数据仍然是语音合成的一个挑战。本文介绍了MIKU-PAL,一个全自动的多模式管道,用于从未标记的视频数据中提取高一致性的情感语音。利用人脸检测和跟踪算法,我们开发了一个自动情感分析系统,使用多模态大语言模型(MLLM)。我们的研究结果表明,MIKU-PAL可以实现人类水平的准确性(MELD上为68.5%)和优越的一致性(0.93 Fleiss kappa评分),同时比人类注释更便宜和更快。借助MIKU-PAL的高质量、灵活和一致的注释,我们可以注释多达26种类型的细粒度语音情感类别,并通过人工注释器进行验证,合理性评级为83%。基于我们提出的系统,我们进一步发布了一个细粒度的情感语音数据集MIKU-MIBench Bench(131.2小时),作为情感文本到语音和视觉语音克隆的新基准。
摘要:Acquiring large-scale emotional speech data with strong consistency remains a challenge for speech synthesis. This paper presents MIKU-PAL, a fully automated multimodal pipeline for extracting high-consistency emotional speech from unlabeled video data. Leveraging face detection and tracking algorithms, we developed an automatic emotion analysis system using a multimodal large language model (MLLM). Our results demonstrate that MIKU-PAL can achieve human-level accuracy (68.5% on MELD) and superior consistency (0.93 Fleiss kappa score) while being much cheaper and faster than human annotation. With the high-quality, flexible, and consistent annotation from MIKU-PAL, we can annotate fine-grained speech emotion categories of up to 26 types, validated by human annotators with 83% rationality ratings. Based on our proposed system, we further released a fine-grained emotional speech dataset MIKU-EmoBench(131.2 hours) as a new benchmark for emotional text-to-speech and visual voice cloning.
【2】 Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
标题: 语音到语音语言模型的高效、直接的双重建模链接:https://arxiv.org/abs/2505.15670
备注:Accepted to Interspeech 2025
摘要:口语对话是人机交互的一种直观形式,但当前的语音语言模型往往仍局限于基于话轮的交流,缺乏用户插话等实时适应性。我们提出了一种新的双工语音到语音(S2S)的体系结构,具有连续的用户输入和编解码器代理输出与信道融合,直接模拟同时用户和代理流。使用预训练的流编码器用于用户输入,实现了第一双工S2S模型,而无需语音预训练。代理和用户建模的单独架构便于编解码器微调更好的代理语音和比特率减半(0.6 kbps)相比,以前的作品。实验结果表明,该模型在推理、话轮转换和插话能力方面优于以往的双工模型。该模型需要的语音数据显著减少,因为跳过了语音预训练,这显著简化了从任何LLM构建双工S2S模型的过程。最后,它是第一个公开可用的双S2S模型,具有训练和推理代码,以促进再现性。
摘要:Spoken dialogue is an intuitive form of human-computer interaction, yet current speech language models often remain constrained to turn-based exchanges, lacking real-time adaptability such as user barge-in. We propose a novel duplex speech to speech (S2S) architecture featuring continuous user inputs and codec agent outputs with channel fusion that directly models simultaneous user and agent streams. Using a pretrained streaming encoder for user input enables the first duplex S2S model without requiring speech pretrain. Separate architectures for agent and user modeling facilitate codec fine-tuning for better agent voices and halve the bitrate (0.6 kbps) compared to previous works. Experimental results show that the proposed model outperforms previous duplex models in reasoning, turn-taking, and barge-in abilities. The model requires significantly less speech data, as speech pretrain is skipped, which markedly simplifies the process of building a duplex S2S model from any LLMs. Finally, it is the first openly available duplex S2S model with training and inference code to foster reproducibility.
【3】 Word Level Timestamp Generation for Automatic Speech Recognition and Translation
标题: 自动语音识别和翻译的单词级时间戳生成链接:https://arxiv.org/abs/2505.15646
备注:Accepted to Interspeech 2025
摘要:我们介绍了一种数据驱动的方法,使单词级的时间戳预测的金丝雀模型。准确的时间戳信息对于语音内容检索和定时字幕等各种下游任务至关重要。虽然传统的混合系统和端到端(E2E)模型可能会采用外部模块进行时间戳预测,但我们的方法消除了对单独对齐机制的需求。通过利用NeMo Forced Aligner(NFA)作为教师模型,我们生成单词级时间戳并训练Canary模型以直接预测时间戳。我们介绍一个新的<|时间戳|> token,使Canary模型能够预测每个单词的开始和结束时间戳。我们的方法表明,精确率和召回率在80%到90%之间,时间戳预测误差范围从20到120毫秒,四种语言,WER退化最小。此外,我们将我们的系统扩展到自动语音翻译(AST)任务,实现了大约200毫秒的时间戳预测误差。
摘要:We introduce a data-driven approach for enabling word-level timestamp prediction in the Canary model. Accurate timestamp information is crucial for a variety of downstream tasks such as speech content retrieval and timed subtitles. While traditional hybrid systems and end-to-end (E2E) models may employ external modules for timestamp prediction, our approach eliminates the need for separate alignment mechanisms. By leveraging the NeMo Forced Aligner (NFA) as a teacher model, we generate word-level timestamps and train the Canary model to predict timestamps directly. We introduce a new <|timestamp|> token, enabling the Canary model to predict start and end timestamps for each word. Our method demonstrates precision and recall rates between 80% and 90%, with timestamp prediction errors ranging from 20 to 120 ms across four languages, with minimal WER degradation. Additionally, we extend our system to automatic speech translation (AST) tasks, achieving timestamp prediction errors around 200 milliseconds.
【4】 Moonbeam: A MIDI Foundation Model Using Both Absolute and Relative Music Attributes
标题: Moonbeam:使用绝对和相对音乐属性的MIDI基金会模型链接:https://arxiv.org/abs/2505.15559
摘要:Moonbeam是一个基于transformer的符号音乐基础模型,在大量不同的音乐数据集合上进行了预训练,总计81.6K小时的音乐和180亿个令牌。Moonbeam通过引入一种新的领域知识启发的标记化方法和多维相对注意力(MRA)来捕获绝对和相对音乐属性,从而结合了音乐领域的归纳偏差,该方法捕获相对音乐信息而无需额外的可训练参数。利用预先训练的Moonbeam,我们提出了2个具有完全预期能力的微调架构,针对2类下游任务:符号音乐理解和条件音乐生成(包括音乐填充)。在大多数情况下,我们的模型在4个数据集上的3个下游音乐分类任务的准确性和F1得分方面优于其他大规模预训练的音乐模型。此外,我们的微调条件音乐生成模型优于一个强大的Transformer基线与一个类似于REMI的标记。我们开源了代码,预训练了模型,并在Github上生成了样本。
摘要:Moonbeam is a transformer-based foundation model for symbolic music, pretrained on a large and diverse collection of MIDI data totaling 81.6K hours of music and 18 billion tokens. Moonbeam incorporates music-domain inductive biases by capturing both absolute and relative musical attributes through the introduction of a novel domain-knowledge-inspired tokenization method and Multidimensional Relative Attention (MRA), which captures relative music information without additional trainable parameters. Leveraging the pretrained Moonbeam, we propose 2 finetuning architectures with full anticipatory capabilities, targeting 2 categories of downstream tasks: symbolic music understanding and conditional music generation (including music infilling). Our model outperforms other large-scale pretrained music models in most cases in terms of accuracy and F1 score across 3 downstream music classification tasks on 4 datasets. Moreover, our finetuned conditional music generation model outperforms a strong transformer baseline with a REMI-like tokenizer. We open-source the code, pretrained model, and generated samples on Github.
【5】 Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
标题: Audio Jailbreak:一个用于越狱大型音频语言模型的开放综合基准测试链接:https://arxiv.org/abs/2505.15406
备注:We release AJailBench, including both static and optimized adversarial data, to facilitate future research: this https URL
摘要:大型音频语言模型(LAMs)的兴起带来了潜在的风险,因为它们的音频输出可能包含有害或不道德的内容。然而,目前的研究缺乏一个系统的,定量的评估LAM的安全性,特别是对越狱攻击,这是具有挑战性的,由于语音的时间和语义的性质。为了弥补这一差距,我们引入AJailBench,这是第一个专门用于评估LAM中越狱漏洞的基准测试。我们首先构建AJailBench-Base,这是一个包含1,495个对抗性音频提示的数据集,涵盖10个违反策略的类别,从使用真实文本的文本越狱攻击转换为语音合成。使用这个数据集,我们评估了几个最先进的LAM,并发现没有一个在攻击中表现出一致的鲁棒性。为了进一步加强越狱测试和模拟更真实的攻击条件,我们提出了一种生成动态对抗变体的方法。我们的Audio Perturbation Toolkit(APT)可在时域、频域和振幅域应用目标失真。为了保持原始的越狱意图,我们强制执行语义一致性约束,并采用贝叶斯优化来有效地搜索既微妙又高效的扰动。这导致AJailBench-APT,这是一个优化的对抗性音频样本的扩展数据集。我们的研究结果表明,即使是很小的,语义保留的扰动可以显着降低领先的LAM的安全性能,强调需要更强大的和语义感知的防御机制。
摘要:The rise of Large Audio Language Models (LAMs) brings both potential and risks, as their audio outputs may contain harmful or unethical content. However, current research lacks a systematic, quantitative evaluation of LAM safety especially against jailbreak attacks, which are challenging due to the temporal and semantic nature of speech. To bridge this gap, we introduce AJailBench, the first benchmark specifically designed to evaluate jailbreak vulnerabilities in LAMs. We begin by constructing AJailBench-Base, a dataset of 1,495 adversarial audio prompts spanning 10 policy-violating categories, converted from textual jailbreak attacks using realistic text to speech synthesis. Using this dataset, we evaluate several state-of-the-art LAMs and reveal that none exhibit consistent robustness across attacks. To further strengthen jailbreak testing and simulate more realistic attack conditions, we propose a method to generate dynamic adversarial variants. Our Audio Perturbation Toolkit (APT) applies targeted distortions across time, frequency, and amplitude domains. To preserve the original jailbreak intent, we enforce a semantic consistency constraint and employ Bayesian optimization to efficiently search for perturbations that are both subtle and highly effective. This results in AJailBench-APT, an extended dataset of optimized adversarial audio samples. Our findings demonstrate that even small, semantically preserved perturbations can significantly reduce the safety performance of leading LAMs, underscoring the need for more robust and semantically aware defense mechanisms.
【6】 Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
标题: 通过上下文学习实现零拍语音转换的韵律自适应音频编解码器链接:https://arxiv.org/abs/2505.15402
备注:5 pages, 3 figures
摘要:离散音频编解码器的最新进展已经显著地改进了语音表示建模,而编解码器语言模型已经实现了用于zero-shot语音合成的上下文学习。受此启发,我们提出了VALLE-X框架内的语音转换(VC)模型,利用其强大的上下文学习能力进行说话人自适应。为了增强韵律控制,我们引入了一个韵律感知音频编解码器编码器(PACE)模块,该模块将韵律与其他源隔离并细化,从而提高表现力和控制力。通过将PACE集成到我们的VC模型中,我们在保留扬声器音色的同时实现了韵律操作的更大灵活性。实验评估结果表明,我们的方法优于基线VC系统的韵律保留,音色一致性,和整体的自然性,超过基线VC系统。
摘要:Recent advances in discrete audio codecs have significantly improved speech representation modeling, while codec language models have enabled in-context learning for zero-shot speech synthesis. Inspired by this, we propose a voice conversion (VC) model within the VALLE-X framework, leveraging its strong in-context learning capabilities for speaker adaptation. To enhance prosody control, we introduce a prosody-aware audio codec encoder (PACE) module, which isolates and refines prosody from other sources, improving expressiveness and control. By integrating PACE into our VC model, we achieve greater flexibility in prosody manipulation while preserving speaker timbre. Experimental evaluation results demonstrate that our approach outperforms baseline VC systems in prosody preservation, timbre consistency, and overall naturalness, surpassing baseline VC systems.
【7】 Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
标题: 利用语音推测解码加速自回归语音合成推理链接:https://arxiv.org/abs/2505.15380
备注:5 pages, 4 figures
摘要:利用语言模型的现代自回归语音合成模型已经表现出显着的性能。然而,这些模型中下一个令牌预测的顺序性质导致了显著的延迟,阻碍了它们在推理速度至关重要的场景中的部署。在这项工作中,我们提出了语音推测解码(SSD),自回归语音合成加速的新框架。具体来说,我们的方法采用了一个轻量级的草案模型来生成候选令牌序列,随后使用建议的SSD框架的目标模型并行验证。实验结果表明,SSD实现了显着的加速比为1.4倍相比,传统的自回归解码,同时保持高保真度和自然。主观评估进一步验证了SSD在保持目标模型的感知质量同时加速推理的有效性。
摘要:Modern autoregressive speech synthesis models leveraging language models have demonstrated remarkable performance. However, the sequential nature of next token prediction in these models leads to significant latency, hindering their deployment in scenarios where inference speed is critical. In this work, we propose Speech Speculative Decoding (SSD), a novel framework for autoregressive speech synthesis acceleration. Specifically, our method employs a lightweight draft model to generate candidate token sequences, which are subsequently verified in parallel by the target model using the proposed SSD framework. Experimental results demonstrate that SSD achieves a significant speedup of 1.4x compared with conventional autoregressive decoding, while maintaining high fidelity and naturalness. Subjective evaluations further validate the effectiveness of SSD in preserving the perceptual quality of the target model while accelerating inference.
【8】 Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
标题: Neurodyne:利用表示学习和周期一致性GAN的神经音调操纵链接:https://arxiv.org/abs/2505.15368
摘要:音高控制是指制作者将音频片段的音高调整到特定的音调和语调的过程,这在音乐制作中是必不可少的。基于神经网络的音高控制系统由于其优于传统的DSP方法的合成质量,近年来得到了广泛的应用。然而,它们的性能仍然有限,因为它们使用源过滤器模型进行不准确的特征分解,并且缺乏成对的调内和调外训练数据。这项工作提出了Neurodyne来解决这些问题。具体来说,Neurodyne使用对抗性表示学习来学习音高独立的潜在表示,以避免不准确的解纠缠和循环一致性训练,以隐式地创建成对的训练数据。全局键和基于模板的音高操作的实验结果表明,该系统的有效性,标记提高合成质量,同时保持原来的歌手身份。
摘要:Pitch manipulation is the process of producers adjusting the pitch of an audio segment to a specific key and intonation, which is essential in music production. Neural-network-based pitch-manipulation systems have been popular in recent years due to their superior synthesis quality compared to classical DSP methods. However, their performance is still limited due to their inaccurate feature disentanglement using source-filter models and the lack of paired in- and out-of-tune training data. This work proposes Neurodyne to address these issues. Specifically, Neurodyne uses adversarial representation learning to learn a pitch-independent latent representation to avoid inaccurate disentanglement and cycle-consistency training to create paired training data implicitly. Experimental results on global-key and template-based pitch manipulation demonstrate the effectiveness of the proposed system, marking improved synthesis quality while maintaining the original singer identity.
【9】 MHANet: Multi-scale Hybrid Attention Network for Auditory Attention Detection
标题: MHANet:用于听觉注意力检测的多尺度混合注意力网络链接:https://arxiv.org/abs/2505.15364
摘要:听觉注意检测(AAD)旨在从脑电信号中检测出多人说话环境中的目标说话人,如脑电信号(EEG),已经取得了很大的进展。然而,大多数AAD方法仅顺序地利用注意机制,并且忽略了EEG信号内的有价值的多尺度上下文信息,限制了它们同时捕获长-短范围时空依赖性的能力。为了解决这些问题,本文提出了一种多尺度混合注意力网络(MHANet)的AAD,它包括多尺度混合注意力(MHA)模块和时空卷积(STC)模块。具体而言,MHA结合了通道注意力和多尺度时间和全局注意力机制。这有效地提取多尺度的时间模式内的EEG信号,并捕获长,短距离的时空依赖性同时。为了进一步提高AAD的性能,STC利用时间和空间卷积来聚合表达时空表示。实验结果表明,所提出的MHANet实现了最先进的性能,在三个数据集上具有更少的可训练参数,比最先进的模型低3倍。代码可从以下网址获得:https://github.com/fchest/MHANet。
摘要:Auditory attention detection (AAD) aims to detect the target speaker in a multi-talker environment from brain signals, such as electroencephalography (EEG), which has made great progress. However, most AAD methods solely utilize attention mechanisms sequentially and overlook valuable multi-scale contextual information within EEG signals, limiting their ability to capture long-short range spatiotemporal dependencies simultaneously. To address these issues, this paper proposes a multi-scale hybrid attention network (MHANet) for AAD, which consists of the multi-scale hybrid attention (MHA) module and the spatiotemporal convolution (STC) module. Specifically, MHA combines channel attention and multi-scale temporal and global attention mechanisms. This effectively extracts multi-scale temporal patterns within EEG signals and captures long-short range spatiotemporal dependencies simultaneously. To further improve the performance of AAD, STC utilizes temporal and spatial convolutions to aggregate expressive spatiotemporal representations. Experimental results show that the proposed MHANet achieves state-of-the-art performance with fewer trainable parameters across three datasets, 3 times lower than that of the most advanced model. Code is available at: https://github.com/fchest/MHANet.
【10】 Decoding Phone Pairs from MEG Signals Across Speech Modalities
标题: 从跨语音模式的MEG信号解码电话对链接:https://arxiv.org/abs/2505.15355
备注:21 pages, 4 figures, 1 graphical abstract, submitted to Computer Speech and Language (special issue on Iberian Languages)
摘要:理解言语产生的神经机制对于推进认知神经科学理论和开发实用的通信技术都是至关重要的。在这项研究中,我们研究了脑磁图信号解码语音产生和感知(被动倾听和语音回放)任务期间的大脑活动的电话。使用包括17名参与者的数据集,我们进行了成对的电话分类,将我们的分析扩展到15个语音对。比较了多种机器学习方法,包括正则化线性模型和神经网络架构,以确定它们在解码语音信息方面的有效性。我们的研究结果表明,与被动倾听和回放模式(~51%)相比,语音产生过程中的解码准确率(76.6%)显着更高,强调了公开语音过程中可用的更丰富的神经信息。在这些模型中,弹性网络分类器的性能始终优于更复杂的神经网络,突出了传统正则化技术在应用于有限和高维MEG数据集时的有效性。此外,对特定脑频带的分析显示,低频振荡,特别是Delta(0.2-3 Hz)和Theta(4-7 Hz),对解码准确性的贡献最大,这表明这些频带编码关键的语音产生相关的神经过程。尽管使用了先进的去噪方法,但仍不清楚解码是否仅反映了神经活动,或者残余的肌肉或运动伪影是否也有贡献,这表明需要进一步改进方法。总的来说,我们的研究结果强调了检查明显的语音产生范式的至关重要性,尽管它们很复杂,但它们提供了改善脑机接口以帮助严重语音障碍患者的机会。
摘要:Understanding the neural mechanisms underlying speech production is essential for both advancing cognitive neuroscience theory and developing practical communication technologies. In this study, we investigated magnetoencephalography signals to decode phones from brain activity during speech production and perception (passive listening and voice playback) tasks. Using a dataset comprising 17 participants, we performed pairwise phone classification, extending our analysis to 15 phonetic pairs. Multiple machine learning approaches, including regularized linear models and neural network architectures, were compared to determine their effectiveness in decoding phonetic information. Our results demonstrate significantly higher decoding accuracy during speech production (76.6%) compared to passive listening and playback modalities (~51%), emphasizing the richer neural information available during overt speech. Among the models, the Elastic Net classifier consistently outperformed more complex neural networks, highlighting the effectiveness of traditional regularization techniques when applied to limited and high-dimensional MEG datasets. Besides, analysis of specific brain frequency bands revealed that low-frequency oscillations, particularly Delta (0.2-3 Hz) and Theta (4-7 Hz), contributed the most substantially to decoding accuracy, suggesting that these bands encode critical speech production-related neural processes. Despite using advanced denoising methods, it remains unclear whether decoding solely reflects neural activity or if residual muscular or movement artifacts also contributed, indicating the need for further methodological refinement. Overall, our findings underline the critical importance of examining overt speech production paradigms, which, despite their complexity, offer opportunities to improve brain-computer interfaces to help individuals with severe speech impairments.
【11】 Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
标题: 利用单元语言指导推进无文本语音翻译中的语音建模链接:https://arxiv.org/abs/2505.15333
备注:Accepted to ACL 2025 Findings
摘要:无文本语音到语音翻译(S2 ST)模型的成功构建引起了人们的广泛关注。然而,S2 ST仍然面临两个主要挑战:1)提取各种语音信号的语言特征,称为跨模态(CM),以及2)学习长序列中不同语言的对齐,称为跨语言(CL)。我们提出的单位语言,以克服这两个建模的挑战。单元语言可以被认为是一种类似文本的表示格式,使用$n$-gram语言建模构建。我们实现了多任务学习,利用单位语言指导语音建模过程。我们的初步结果揭示了冲突时,同时应用源和目标单元语言。我们建议任务提示建模,以减轻这种冲突。我们在Voxpupil数据集的四种语言上进行实验。我们的方法在强大的基线上表现出显着的改进,并实现了与使用文本训练的模型相当的性能。
摘要:The success of building textless speech-to-speech translation (S2ST) models has attracted much attention. However, S2ST still faces two main challenges: 1) extracting linguistic features for various speech signals, called cross-modal (CM), and 2) learning alignment of difference languages in long sequences, called cross-lingual (CL). We propose the unit language to overcome the two modeling challenges. The unit language can be considered a text-like representation format, constructed using $n$-gram language modeling. We implement multi-task learning to utilize the unit language in guiding the speech modeling process. Our initial results reveal a conflict when applying source and target unit languages simultaneously. We propose task prompt modeling to mitigate this conflict. We conduct experiments on four languages of the Voxpupil dataset. Our method demonstrates significant improvements over a strong baseline and achieves performance comparable to models trained with text.
【12】 Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
标题: Voice-ENHANCE:使用基于扩散的语音转换框架进行语音恢复链接:https://arxiv.org/abs/2505.15254
备注:5 pages, 3 figures, Accepted to INTERSPEECH 2025
摘要:我们提出了一个语音增强系统,结合说话人不可知的语音恢复与语音转换(VC),以获得演播室级质量的语音信号。虽然语音转换模型通常用于改变说话者的特征,但当目标说话者与源说话者相同时,它们也可以用作语音恢复的手段。然而,由于VC模型容易受到噪声条件的影响,我们在我们提出的系统的前端包括了一个生成语音恢复(GSR)模型。GSR模型执行噪声抑制,并在不了解目标说话者的情况下恢复在该过程中发生的语音损伤。VC阶段然后使用来自干净说话者嵌入的指导来进一步恢复输出语音。通过采用这种两阶段的方法,我们在多个数据集上实现了与最先进的(SOTA)方法相当的语音质量客观度量分数。
摘要:We propose a speech enhancement system that combines speaker-agnostic speech restoration with voice conversion (VC) to obtain a studio-level quality speech signal. While voice conversion models are typically used to change speaker characteristics, they can also serve as a means of speech restoration when the target speaker is the same as the source speaker. However, since VC models are vulnerable to noisy conditions, we have included a generative speech restoration (GSR) model at the front end of our proposed system. The GSR model performs noise suppression and restores speech damage incurred during that process without knowledge about the target speaker. The VC stage then uses guidance from clean speaker embeddings to further restore the output speech. By employing this two-stage approach, we have achieved speech quality objective metric scores comparable to state-of-the-art (SOTA) methods across multiple datasets.
【13】 Hybrid Audio Detection Using Fine-Tuned Audio Spectrogram Transformers: A Dataset-Driven Evaluation of Mixed AI-Human Speech
标题: 使用微调音频频谱图变换器的混合音频检测:人工智能-人类混合语音的数据集驱动评估链接:https://arxiv.org/abs/2505.15136
备注:11 pages, 6 figures
摘要:人工智能(AI)的快速发展实现了复杂的音频生成和语音克隆技术,为依赖语音认证的应用带来了重大的安全风险。虽然现有的数据集和模型主要集中在区分人类和完全合成的语音,但现实世界的攻击通常涉及结合了真实和克隆片段的音频。为了解决这一差距,我们构建了一个新的混合音频数据集,其中包含人类,AI生成,克隆和混合音频样本。我们进一步提出了微调音频频谱图Transformer(AST)为基础的模型检测这些复杂的声学模式。大量的实验表明,我们的方法显着优于现有的混合音频检测基线,达到97%的分类准确率。我们的研究结果强调了混合数据集和定制模型在提高基于语音的认证系统的鲁棒性方面的重要性。
摘要:The rapid advancement of artificial intelligence (AI) has enabled sophisticated audio generation and voice cloning technologies, posing significant security risks for applications reliant on voice authentication. While existing datasets and models primarily focus on distinguishing between human and fully synthetic speech, real-world attacks often involve audio that combines both genuine and cloned segments. To address this gap, we construct a novel hybrid audio dataset incorporating human, AI-generated, cloned, and mixed audio samples. We further propose fine-tuned Audio Spectrogram Transformer (AST)-based models tailored for detecting these complex acoustic patterns. Extensive experiments demonstrate that our approach significantly outperforms existing baselines in mixed-audio detection, achieving 97\% classification accuracy. Our findings highlight the importance of hybrid datasets and tailored models in advancing the robustness of speech-based authentication systems.
【14】 SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
标题: SHEET:多用途开源语音人类评估估计工具包链接:https://arxiv.org/abs/2505.15061
备注:INTERSPEECH 2025. Codebase: this https URL
摘要:我们介绍SHEET,一个多用途的开源工具包,旨在加速主观语音质量评估(SSQA)的研究。SHEET代表Speech Human Evaluation Estimation Toolkit,它专注于数据驱动的基于深度神经网络的模型,这些模型经过训练,可以预测语音样本的人类标记质量分数。SHEET提供全面的训练和评估脚本,多数据集和多模型支持,以及通过Torch Hub和HuggingFace Spaces访问的预训练模型。为了证明它的能力,我们重新评估了SSL-MOS,这是一种基于语音自监督学习(SSL)的SSQA模型,在最近的科学论文中广泛使用。在两个具有代表性的SSQA数据集BVCC和NISQA上进行了实验,我们确定了最佳语音SSL模型,其性能超过了原始SSL-MOS实现,与最先进的方法相当。
摘要:We introduce SHEET, a multi-purpose open-source toolkit designed to accelerate subjective speech quality assessment (SSQA) research. SHEET stands for the Speech Human Evaluation Estimation Toolkit, which focuses on data-driven deep neural network-based models trained to predict human-labeled quality scores of speech samples. SHEET provides comprehensive training and evaluation scripts, multi-dataset and multi-model support, as well as pre-trained models accessible via Torch Hub and HuggingFace Spaces. To demonstrate its capabilities, we re-evaluated SSL-MOS, a speech self-supervised learning (SSL)-based SSQA model widely used in recent scientific papers, on an extensive list of speech SSL models. Experiments were conducted on two representative SSQA datasets named BVCC and NISQA, and we identified the optimal speech SSL model, whose performance surpassed the original SSL-MOS implementation and was comparable to state-of-the-art methods.
【15】 AsynFusion: Towards Asynchronous Latent Consistency Models for Decoupled Whole-Body Audio-Driven Avatars
标题: AsynFusion:面向脱钩全身音频驱动化身的同步性模型链接:https://arxiv.org/abs/2505.15058
备注:11pages, conference
摘要:全身音频驱动的化身姿态和表情生成是创建逼真的数字人和增强交互式虚拟代理的能力的关键任务,在虚拟现实,数字娱乐和远程通信中具有广泛的应用。现有的方法通常独立地生成音频驱动的面部表情和手势,这引入了一个显著的限制:面部和手势元素之间缺乏无缝协调,导致不太自然和有凝聚力的动画。为了解决这个问题,我们提出了AsynFusion,一个新的框架,利用扩散Transformers来实现和谐的表达和手势合成。所提出的方法是建立在一个双分支DiT架构,使面部表情和手势的并行生成。在该模型中,我们引入了一个协作同步模块,以促进两种模式之间的双向功能交互,以及一个异步LCM采样策略,以减少计算开销,同时保持高质量的输出。大量实验表明,AsynFusion在生成实时同步全身动画方面达到了最先进的性能,在定量和定性评估方面始终优于现有方法。
摘要:Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital entertainment, and remote communication. Existing approaches often generate audio-driven facial expressions and gestures independently, which introduces a significant limitation: the lack of seamless coordination between facial and gestural elements, resulting in less natural and cohesive animations. To address this limitation, we propose AsynFusion, a novel framework that leverages diffusion transformers to achieve harmonious expression and gesture synthesis. The proposed method is built upon a dual-branch DiT architecture, which enables the parallel generation of facial expressions and gestures. Within the model, we introduce a Cooperative Synchronization Module to facilitate bidirectional feature interaction between the two modalities, and an Asynchronous LCM Sampling strategy to reduce computational overhead while maintaining high-quality outputs. Extensive experiments demonstrate that AsynFusion achieves state-of-the-art performance in generating real-time, synchronized whole-body animations, consistently outperforming existing methods in both quantitative and qualitative evaluations.
【16】 Discrete Audio Representations for Automated Audio Captioning
标题: 用于自动音频字幕的离散音频表示链接:https://arxiv.org/abs/2505.14989
备注:Interspeech 2025
摘要:离散音频表示,称为音频令牌,被广泛地分类为语义和声学令牌,通常通过连续音频表示的无监督令牌化生成。然而,它们对自动音频字幕(AAC)的适用性仍然有待探索。本文通过对各种标记化方法的比较分析,系统地研究了AAC音频标记驱动模型的可行性。我们的研究结果表明,音频标记化导致AAC模型的性能下降相比,那些直接利用连续的音频表示。为了解决这个问题,我们引入了一个有监督的音频标记训练的音频标记目标。与缺乏明确语义理解的无监督标记器不同,所提出的标记器有效地捕获音频事件信息。在Clotho数据集上进行的实验表明,所提出的音频令牌在AAC任务中优于传统的音频令牌。
摘要:Discrete audio representations, termed audio tokens, are broadly categorized into semantic and acoustic tokens, typically generated through unsupervised tokenization of continuous audio representations. However, their applicability to automated audio captioning (AAC) remains underexplored. This paper systematically investigates the viability of audio token-driven models for AAC through comparative analyses of various tokenization methods. Our findings reveal that audio tokenization leads to performance degradation in AAC models compared to those that directly utilize continuous audio representations. To address this issue, we introduce a supervised audio tokenizer trained with an audio tagging objective. Unlike unsupervised tokenizers, which lack explicit semantic understanding, the proposed tokenizer effectively captures audio event information. Experiments conducted on the Clotho dataset demonstrate that the proposed audio tokens outperform conventional audio tokens in the AAC task.
【17】 Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
标题: 迈向包容性ASB:研究低资源语言中用于发音障碍语音识别的语音转换链接:https://arxiv.org/abs/2505.14874
备注:5 pages, 1 figure, Accepted to Interspeech 2025
摘要:由于数据稀缺,特别是在非英语语言中,构音障碍语音的自动语音识别(ASR)仍然具有挑战性。为了解决这个问题,我们微调英语构音障碍语音(UASpeech)的语音转换模型,对说话者特征和韵律失真进行编码,然后将其应用于将健康的非英语语音(FLEURS)转换为非英语构音障碍类语音。然后,生成的数据用于微调多语言ASR模型,大规模多语言语音(MMS),以改善构音障碍的语音识别。对PC-GITA(西班牙语),EasyCall(意大利语)和SSNCE(泰米尔语)的评估表明,VC与扬声器和韵律转换显着优于现成的MMS性能和传统的增强技术,如速度和节奏扰动。客观和主观的分析所产生的数据进一步证实,所产生的语音模拟构音障碍的特点。
摘要:Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, particularly in non-English languages. To address this, we fine-tune a voice conversion model on English dysarthric speech (UASpeech) to encode both speaker characteristics and prosodic distortions, then apply it to convert healthy non-English speech (FLEURS) into non-English dysarthric-like speech. The generated data is then used to fine-tune a multilingual ASR model, Massively Multilingual Speech (MMS), for improved dysarthric speech recognition. Evaluation on PC-GITA (Spanish), EasyCall (Italian), and SSNCE (Tamil) demonstrates that VC with both speaker and prosody conversion significantly outperforms the off-the-shelf MMS performance and conventional augmentation techniques such as speed and tempo perturbation. Objective and subjective analyses of the generated data further confirm that the generated speech simulates dysarthric characteristics.
【18】 Replay Attacks Against Audio Deepfake Detection
标题: 针对音频Deepfake检测的重播攻击链接:https://arxiv.org/abs/2505.14862
备注:None
摘要:我们展示了重放攻击如何破坏音频deepfake检测:通过各种扬声器和麦克风播放和重新录制deepfake音频,我们使欺骗样本在检测模型中看起来是真实的。为了更详细地研究这一现象,我们引入了ReplayDF,这是一个来自M-AILABS和MLAAD的录音数据集,其中包括六种语言和四种TTS模型的109种扬声器-麦克风组合。它包括各种声学条件,其中一些对检测具有高度挑战性。我们对五个数据集的六个开源检测模型的分析揭示了显著的漏洞,表现最好的W2 V2-AASIST模型的等错误率(EER)从4.7%飙升至18.2%。即使采用自适应房间脉冲响应(RIR)再训练,性能仍然受到11.0% EER的影响。我们发布ReplayDF用于非商业研究用途。
摘要:We show how replay attacks undermine audio deepfake detection: By playing and re-recording deepfake audio through various speakers and microphones, we make spoofed samples appear authentic to the detection model. To study this phenomenon in more detail, we introduce ReplayDF, a dataset of recordings derived from M-AILABS and MLAAD, featuring 109 speaker-microphone combinations across six languages and four TTS models. It includes diverse acoustic conditions, some highly challenging for detection. Our analysis of six open-source detection models across five datasets reveals significant vulnerability, with the top-performing W2V2-AASIST model's Equal Error Rate (EER) surging from 4.7% to 18.2%. Even with adaptive Room Impulse Response (RIR) retraining, performance remains compromised with an 11.0% EER. We release ReplayDF for non-commercial research use.
【19】 GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples
标题: GraphemeAug:一个系统化的方法来综合硬否定关键词发现的例子链接:https://arxiv.org/abs/2505.14814
备注:Accepted at Interspeech 2025
摘要:口语关键词识别(英语:Spoken Keyword Spotting,简称KWS)是一项区分音频中关键词的存在和不存在的任务。KWS模型的准确性取决于其正确分类接近关键字和非关键字边界的示例的能力。这些边界示例在训练数据中通常很少,限制了模型性能。在本文中,我们提出了一种方法,通过对关键字的字素进行插入/删除/替换编辑,系统地生成接近决策边界的对抗性示例。我们对流行关键字的保留数据进行了评估,并表明该技术将合成硬否定数据集的AUC提高了61%,同时保持了积极和环境负面音频数据的质量。
摘要:Spoken Keyword Spotting (KWS) is the task of distinguishing between the presence and absence of a keyword in audio. The accuracy of a KWS model hinges on its ability to correctly classify examples close to the keyword and non-keyword boundary. These boundary examples are often scarce in training data, limiting model performance. In this paper, we propose a method to systematically generate adversarial examples close to the decision boundary by making insertion/deletion/substitution edits on the keyword's graphemes. We evaluate this technique on held-out data for a popular keyword and show that the technique improves AUC on a dataset of synthetic hard negatives by 61% while maintaining quality on positives and ambient negative audio data.
【20】 Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
标题: 保存副语言和韵律信息的分段变体代码簿链接:https://arxiv.org/abs/2505.15667
备注:Accepted to Interspeech 2025
摘要:SSL语音模型中的量化(例如,HuBERT)在诸如语言建模、再合成和文本到语音的任务中改进了压缩和性能,但是经常丢弃韵律和非语言学信息(例如,情感,突出)。虽然增加码本大小减轻了一些损失,但它无效地提高了比特率。我们提出了分段变体码本(SVC),它在不同的语言单位(帧,电话,单词,话语)的语音,分解成多个流的段特定的离散功能。我们的研究结果表明,SVCs是显着更有效地保留韵律和跨探测任务的语言信息。此外,我们发现,池之前,而不是离散化后更好地保留段级信息。再合成实验进一步证实了改进的风格实现和略有改善的质量,同时保持可理解性。
摘要:Quantization in SSL speech models (e.g., HuBERT) improves compression and performance in tasks like language modeling, resynthesis, and text-to-speech but often discards prosodic and paralinguistic information (e.g., emotion, prominence). While increasing codebook size mitigates some loss, it inefficiently raises bitrates. We propose Segmentation-Variant Codebooks (SVCs), which quantize speech at distinct linguistic units (frame, phone, word, utterance), factorizing it into multiple streams of segment-specific discrete features. Our results show that SVCs are significantly more effective at preserving prosodic and paralinguistic information across probing tasks. Additionally, we find that pooling before rather than after discretization better retains segment-level information. Resynthesis experiments further confirm improved style realization and slightly improved quality while preserving intelligibility.
【21】 Towards Pre-training an Effective Respiratory Audio Foundation Model
标题: 迈向预训练有效的呼吸音频基础模型链接:https://arxiv.org/abs/2505.15307
备注:5 pages, 2 figures, 4 tables, Accepted by Interspeech 2025
摘要:基础模型的最新进展引发了对呼吸音频基础模型的兴趣。然而,将传统的预训练方案应用于小规模且缺乏多样性的数据集的有效性尚未得到充分验证。本研究旨在通过比较大量预训练的音频模型,探索更好的呼吸声预训练方法。我们的调查显示,在AudioSet(一个通用的音频数据集)上预训练的模型比专门在呼吸声上预训练的模型更有效。此外,结合AudioSet和呼吸声数据集进行进一步的预训练可以提高性能,并且在聚合特征时保留频率信息至关重要。随着实验中发现的更多见解,我们为OPERA基准建立了一个新的最先进的水平,有助于推进呼吸音频基础模型。我们的代码可从https://github.com/nttcslab/eval-audio-repr/tree/main/plugin/OPERA在线获取。
摘要:Recent advancements in foundation models have sparked interest in respiratory audio foundation models. However, the effectiveness of applying conventional pre-training schemes to datasets that are small-sized and lack diversity has not been sufficiently verified. This study aims to explore better pre-training practices for respiratory sounds by comparing numerous pre-trained audio models. Our investigation reveals that models pre-trained on AudioSet, a general audio dataset, are more effective than the models specifically pre-trained on respiratory sounds. Moreover, combining AudioSet and respiratory sound datasets for further pre-training enhances performance, and preserving the frequency-wise information when aggregating features is vital. Along with more insights found in the experiments, we establish a new state-of-the-art for the OPERA benchmark, contributing to advancing respiratory audio foundation models. Our code is available online at https://github.com/nttcslab/eval-audio-repr/tree/main/plugin/OPERA.
【22】 EASY: Emotion-aware Speaker Anonymization via Factorized Distillation
标题: 轻松:通过分解蒸馏实现语音感知的扬声器语音化链接:https://arxiv.org/abs/2505.15004
备注:Accepted by INTERSPEECH 2025
摘要:情感在语音交互中起着重要的作用,通过音调、音高和节奏来传达,使情感和意图的表达超越语言,创造出更加个性化的体验。然而,大多数现有的说话人匿名化系统采用并行解纠缠方法,仅将语音分离为语言内容和说话人身份,往往忽略了原始情感状态的保留。在这项研究中,我们介绍了EASY,一个情感感知的说话人匿名框架。EASY采用了一种新的顺序解开过程来解开说话人身份,语言内容和情感表征,通过因子分解蒸馏方法在不同的子空间中建模每个语音属性。通过独立约束说话人身份和情感表达,EASY最大限度地减少了信息泄露,增强了隐私保护,同时保留了原始的语言内容和情感状态。在VoicePrivacy Challenge官方数据集上的实验结果表明,我们提出的方法优于所有基线系统,有效地保护了说话者的隐私,同时保持了语言内容和情绪状态。
摘要:Emotion plays a significant role in speech interaction, conveyed through tone, pitch, and rhythm, enabling the expression of feelings and intentions beyond words to create a more personalized experience. However, most existing speaker anonymization systems employ parallel disentanglement methods, which only separate speech into linguistic content and speaker identity, often neglecting the preservation of the original emotional state. In this study, we introduce EASY, an emotion-aware speaker anonymization framework. EASY employs a novel sequential disentanglement process to disentangle speaker identity, linguistic content, and emotional representation, modeling each speech attribute in distinct subspaces through a factorized distillation approach. By independently constraining speaker identity and emotional representation, EASY minimizes information leakage, enhancing privacy protection while preserving original linguistic content and emotional state. Experimental results on the VoicePrivacy Challenge official datasets demonstrate that our proposed approach outperforms all baseline systems, effectively protecting speaker privacy while maintaining linguistic content and emotional state.
【23】 TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
标题: TC Singer 2:可定制多语言Zero-Shot歌唱语音合成链接:https://arxiv.org/abs/2505.14910
备注:Accepted by ACL 2025
摘要:可定制的多语种zero-shot演唱语音合成(SVS)在音乐创作和短视频配音方面有着多种潜在应用。然而,现有的SVS模型过度依赖于音素和音符边界注释,限制了它们在zero-shot场景中的鲁棒性,并且在音素和音符之间产生差的过渡。此外,他们也缺乏有效的多层次的风格控制,通过不同的提示。为了克服这些挑战,我们介绍了TCSinger 2,一个多任务的多语言zero-shot SVS模型,具有基于各种提示的风格转换和风格控制。TCSinger 2主要包括三个关键模块:1)模糊边界内容(BBC)编码器,预测持续时间,扩展内容嵌入,并对边界应用掩蔽以实现平滑过渡。2)自定义音频编码器,使用对比学习从歌唱,语音和文本提示中提取对齐的表示。3)基于流的自定义Transformer利用Cus-MOE,通过F0监督,增强了合成质量和生成的歌声的风格建模。实验结果表明,TCSinger 2在多个相关任务的主观和客观指标上都优于基线模型。
摘要:Customizable multilingual zero-shot singing voice synthesis (SVS) has various potential applications in music composition and short video dubbing. However, existing SVS models overly depend on phoneme and note boundary annotations, limiting their robustness in zero-shot scenarios and producing poor transitions between phonemes and notes. Moreover, they also lack effective multi-level style control via diverse prompts. To overcome these challenges, we introduce TCSinger 2, a multi-task multilingual zero-shot SVS model with style transfer and style control based on various prompts. TCSinger 2 mainly includes three key modules: 1) Blurred Boundary Content (BBC) Encoder, predicts duration, extends content embedding, and applies masking to the boundaries to enable smooth transitions. 2) Custom Audio Encoder, uses contrastive learning to extract aligned representations from singing, speech, and textual prompts. 3) Flow-based Custom Transformer, leverages Cus-MOE, with F0 supervision, enhancing both the synthesis quality and style modeling of the generated singing voice. Experimental results show that TCSinger 2 outperforms baseline models in both subjective and objective metrics across multiple related tasks.
【24】 QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
标题: QUADS:用于高效语音语言理解的QUANTized Distillation框架链接:https://arxiv.org/abs/2505.14723
备注:None
摘要:口语理解(SLU)系统必须平衡性能和效率,特别是在资源受限的环境中。现有的方法分别应用蒸馏和量化,导致次优压缩,因为蒸馏忽略了量化约束。我们提出了QUADS,这是一个统一的框架,通过预调整模型的多阶段训练来优化两者,在保持准确性的同时增强对低位制度的适应性。QUADS在SLURP上达到71.13\%的准确度,在FSC上达到99.20\%,与最先进的模型相比,只有高达5.56\%的轻微退化。此外,它减少了60- 73 $\times $(GMAC)的计算复杂度和83- 700 $\times $的模型大小,在极端量化下表现出强大的鲁棒性。这些结果建立QUADS作为一个高效的解决方案,为现实世界,资源受限的SLU应用。
摘要:Spoken Language Understanding (SLU) systems must balance performance and efficiency, particularly in resource-constrained environments. Existing methods apply distillation and quantization separately, leading to suboptimal compression as distillation ignores quantization constraints. We propose QUADS, a unified framework that optimizes both through multi-stage training with a pre-tuned model, enhancing adaptability to low-bit regimes while maintaining accuracy. QUADS achieves 71.13\% accuracy on SLURP and 99.20\% on FSC, with only minor degradations of up to 5.56\% compared to state-of-the-art models. Additionally, it reduces computational complexity by 60--73$\times$ (GMACs) and model size by 83--700$\times$, demonstrating strong robustness under extreme quantization. These results establish QUADS as a highly efficient solution for real-world, resource-constrained SLU applications.
标题: ToxicTone:一个带有毒性和毒性直言不讳音调注释的普通话音频数据集
链接:https://arxiv.org/abs/2505.15773
备注:Accepted by INTERSPEECH 2025. 5 pages
摘要:尽管对文本中的有毒语音检测进行了广泛的研究,但在处理普通话音频方面仍然存在关键差距。缺乏注释的数据集,捕捉独特的韵律线索和文化的具体表达,在普通话叶口语毒性探索不足。为了解决这个问题,我们引入了ToxicTone --同类最大的公共数据集--其具有区分两种形式毒性的详细注释(例如,亵渎,欺凌)和毒性来源(例如,愤怒、讽刺、轻蔑)。我们的数据来源于各种真实世界的音频,并分为13个主题类别,反映了真实的通信场景。我们还提出了一个多模态检测框架,集成了声学,语言和情感特征,使用最先进的语音和情感编码器。大量的实验表明,我们的方法优于纯文本和基线模型,强调了语音特定的线索在揭示隐藏的有毒表达的重要作用。
摘要:Despite extensive research on toxic speech detection in text, a critical gap remains in handling spoken Mandarin audio. The lack of annotated datasets that capture the unique prosodic cues and culturally specific expressions in Mandarin leaves spoken toxicity underexplored. To address this, we introduce ToxicTone -- the largest public dataset of its kind -- featuring detailed annotations that distinguish both forms of toxicity (e.g., profanity, bullying) and sources of toxicity (e.g., anger, sarcasm, dismissiveness). Our data, sourced from diverse real-world audio and organized into 13 topical categories, mirrors authentic communication scenarios. We also propose a multimodal detection framework that integrates acoustic, linguistic, and emotional features using state-of-the-art speech and emotion encoders. Extensive experiments show our approach outperforms text-only and baseline models, underscoring the essential role of speech-specific cues in revealing hidden toxic expressions.
【2】 Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
标题: 保存副语言和韵律信息的分段变体代码簿链接:https://arxiv.org/abs/2505.15667
备注:Accepted to Interspeech 2025
摘要:SSL语音模型中的量化(例如,HuBERT)在诸如语言建模、再合成和文本到语音的任务中改进了压缩和性能,但是经常丢弃韵律和非语言学信息(例如,情感,突出)。虽然增加码本大小减轻了一些损失,但它无效地提高了比特率。我们提出了分段变体码本(SVC),它在不同的语言单位(帧,电话,单词,话语)的语音,分解成多个流的段特定的离散功能。我们的研究结果表明,SVCs是显着更有效地保留韵律和跨探测任务的语言信息。此外,我们发现,池之前,而不是离散化后更好地保留段级信息。再合成实验进一步证实了改进的风格实现和略有改善的质量,同时保持可理解性。
摘要:Quantization in SSL speech models (e.g., HuBERT) improves compression and performance in tasks like language modeling, resynthesis, and text-to-speech but often discards prosodic and paralinguistic information (e.g., emotion, prominence). While increasing codebook size mitigates some loss, it inefficiently raises bitrates. We propose Segmentation-Variant Codebooks (SVCs), which quantize speech at distinct linguistic units (frame, phone, word, utterance), factorizing it into multiple streams of segment-specific discrete features. Our results show that SVCs are significantly more effective at preserving prosodic and paralinguistic information across probing tasks. Additionally, we find that pooling before rather than after discretization better retains segment-level information. Resynthesis experiments further confirm improved style realization and slightly improved quality while preserving intelligibility.
【3】 On the Relevance of Clinical Assessment Tasks for the Automatic Detection of Parkinson's Disease Medication State from Speech
标题: 从言语自动检测帕金森病用药状态的临床评估任务的相关性链接:https://arxiv.org/abs/2505.15378
备注:Accepted to Interspeech 2025
摘要:帕金森病(PD)患者的药物状态的自动识别可以帮助临床医生监测和安排个性化治疗,以及研究药物在缓解疾病特征的运动症状方面的作用。本文探讨了语音作为一种非侵入性和可访问的生物标志物,用于识别PD药物状态,介绍了一种新的方法,从说话者独立的角度来解决这一任务。虽然传统的机器学习模型取得了有竞争力的结果,但自我监督的语音表示被证明对最佳性能至关重要,显着超过基于知识的声学描述符。不同的语音评估任务的实验突出了韵律和连续语音在区分药物状态方面的相关性,达到了88.2%的F1分数。这些发现可能会简化临床医生的工作,并减少患者在语音记录方面的努力。
摘要:The automatic identification of medication states of Parkinson's disease (PD) patients can assist clinicians in monitoring and scheduling personalized treatments, as well as studying the effects of medication in alleviating the motor symptoms that characterize the disease. This paper explores speech as a non-invasive and accessible biomarker for identifying PD medication states, introducing a novel approach that addresses this task from a speaker-independent perspective. While traditional machine learning models achieve competitive results, self-supervised speech representations prove essential for optimal performance, significantly surpassing knowledge-based acoustic descriptors. Experiments across diverse speech assessment tasks highlight the relevance of prosody and continuous speech in distinguishing medication states, reaching an F1-score of 88.2%. These findings may streamline clinicians' work and reduce patient effort in voice recordings.
【4】 Analysis of ABC Frontend Audio Systems for the NIST-SRE24
标题: NIST-SRE 24的ABC前端音频系统分析链接:https://arxiv.org/abs/2505.15320
备注:Accepted at Interspeech 2025
摘要:我们对ABC团队为NIST SRE 2024的音轨开发的嵌入提取器(前端)进行了全面分析。我们遵循NIST规定的两种情况:仅使用提供的一组电话录音进行训练(固定)或添加公开可用的数据(开放条件)。在这些约束下,我们开发了最好的说话人嵌入提取器占主导地位的会话电话语音(CTS)域。我们探索了基于ResNet的具有不同池化机制的架构,最近引入了ReDimNet架构,以及基于XLS-R模型的系统,该模型代表了大型预训练自监督模型家族。在开放条件下,我们在VoxBlink 2数据集上进行训练,该数据集包含11万多语言的说话者。我们观察到VoxBlink训练模型的良好性能和鲁棒性,我们的实验显示了开发最先进的说话人识别前端的实用方法。
摘要:We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by NIST: using only a provided set of telephone recordings for training (fixed) or adding publicly available data (open condition). Under these constraints, we develop the best possible speaker embedding extractors for the pre-dominant conversational telephone speech (CTS) domain. We explored architectures based on ResNet with different pooling mechanisms, recently introduced ReDimNet architecture, as well as a system based on the XLS-R model, which represents the family of large pre-trained self-supervised models. In open condition, we train on VoxBlink2 dataset, containing 110 thousand speakers across multiple languages. We observed a good performance and robustness of VoxBlink-trained models, and our experiments show practical recipes for developing state-of-the-art frontends for speaker recognition.
【5】 Towards Pre-training an Effective Respiratory Audio Foundation Model
标题: 迈向预训练有效的呼吸音频基础模型链接:https://arxiv.org/abs/2505.15307
备注:5 pages, 2 figures, 4 tables, Accepted by Interspeech 2025
摘要:基础模型的最新进展引发了对呼吸音频基础模型的兴趣。然而,将传统的预训练方案应用于小规模且缺乏多样性的数据集的有效性尚未得到充分验证。本研究旨在通过比较大量预训练的音频模型,探索更好的呼吸声预训练方法。我们的调查显示,在AudioSet(一个通用的音频数据集)上预训练的模型比专门在呼吸声上预训练的模型更有效。此外,结合AudioSet和呼吸声数据集进行进一步的预训练可以提高性能,并且在聚合特征时保留频率信息至关重要。随着实验中发现的更多见解,我们为OPERA基准建立了一个新的最先进的水平,有助于推进呼吸音频基础模型。我们的代码可从www.example.com在线获取。
摘要:Recent advancements in foundation models have sparked interest in respiratory audio foundation models. However, the effectiveness of applying conventional pre-training schemes to datasets that are small-sized and lack diversity has not been sufficiently verified. This study aims to explore better pre-training practices for respiratory sounds by comparing numerous pre-trained audio models. Our investigation reveals that models pre-trained on AudioSet, a general audio dataset, are more effective than the models specifically pre-trained on respiratory sounds. Moreover, combining AudioSet and respiratory sound datasets for further pre-training enhances performance, and preserving the frequency-wise information when aggregating features is vital. Along with more insights found in the experiments, we establish a new state-of-the-art for the OPERA benchmark, contributing to advancing respiratory audio foundation models. Our code is available online at https://github.com/nttcslab/eval-audio-repr/tree/main/plugin/OPERA.
【6】 EASY: Emotion-aware Speaker Anonymization via Factorized Distillation
标题: 轻松:通过分解蒸馏实现语音感知的扬声器语音化链接:https://arxiv.org/abs/2505.15004
备注:Accepted by INTERSPEECH 2025
摘要:情感在语音交互中起着重要的作用,通过音调、音高和节奏来传达,使情感和意图的表达超越语言,创造出更加个性化的体验。然而,大多数现有的说话人匿名化系统采用并行解纠缠方法,仅将语音分离为语言内容和说话人身份,往往忽略了原始情感状态的保留。在这项研究中,我们介绍了EASY,一个情感感知的说话人匿名框架。EASY采用了一种新的顺序解开过程来解开说话人身份,语言内容和情感表征,通过因子分解蒸馏方法在不同的子空间中建模每个语音属性。通过独立约束说话人身份和情感表达,EASY最大限度地减少了信息泄露,增强了隐私保护,同时保留了原始的语言内容和情感状态。在VoicePrivacy Challenge官方数据集上的实验结果表明,我们提出的方法优于所有基线系统,有效地保护了说话者的隐私,同时保持了语言内容和情绪状态。
摘要:Emotion plays a significant role in speech interaction, conveyed through tone, pitch, and rhythm, enabling the expression of feelings and intentions beyond words to create a more personalized experience. However, most existing speaker anonymization systems employ parallel disentanglement methods, which only separate speech into linguistic content and speaker identity, often neglecting the preservation of the original emotional state. In this study, we introduce EASY, an emotion-aware speaker anonymization framework. EASY employs a novel sequential disentanglement process to disentangle speaker identity, linguistic content, and emotional representation, modeling each speech attribute in distinct subspaces through a factorized distillation approach. By independently constraining speaker identity and emotional representation, EASY minimizes information leakage, enhancing privacy protection while preserving original linguistic content and emotional state. Experimental results on the VoicePrivacy Challenge official datasets demonstrate that our proposed approach outperforms all baseline systems, effectively protecting speaker privacy while maintaining linguistic content and emotional state.
【7】 TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
标题: TC Singer 2:可定制多语言Zero-Shot歌唱语音合成链接:https://arxiv.org/abs/2505.14910
备注:Accepted by ACL 2025
摘要:可定制的多语种zero-shot演唱语音合成(SVS)在音乐创作和短视频配音方面有着多种潜在应用。然而,现有的SVS模型过度依赖于音素和音符边界注释,限制了它们在zero-shot场景中的鲁棒性,并且在音素和音符之间产生差的过渡。此外,他们也缺乏有效的多层次的风格控制,通过不同的提示。为了克服这些挑战,我们介绍了TCSinger 2,一个多任务的多语言zero-shot SVS模型,具有基于各种提示的风格转换和风格控制。TCSinger 2主要包括三个关键模块:1)模糊边界内容(BBC)编码器,预测持续时间,扩展内容嵌入,并对边界应用掩蔽以实现平滑过渡。2)自定义音频编码器,使用对比学习从歌唱,语音和文本提示中提取对齐的表示。3)基于流的自定义Transformer利用Cus-MOE,通过F0监督,增强了合成质量和生成的歌声的风格建模。实验结果表明,TCSinger 2在多个相关任务的主观和客观指标上都优于基线模型。
摘要:Customizable multilingual zero-shot singing voice synthesis (SVS) has various potential applications in music composition and short video dubbing. However, existing SVS models overly depend on phoneme and note boundary annotations, limiting their robustness in zero-shot scenarios and producing poor transitions between phonemes and notes. Moreover, they also lack effective multi-level style control via diverse prompts. To overcome these challenges, we introduce TCSinger 2, a multi-task multilingual zero-shot SVS model with style transfer and style control based on various prompts. TCSinger 2 mainly includes three key modules: 1) Blurred Boundary Content (BBC) Encoder, predicts duration, extends content embedding, and applies masking to the boundaries to enable smooth transitions. 2) Custom Audio Encoder, uses contrastive learning to extract aligned representations from singing, speech, and textual prompts. 3) Flow-based Custom Transformer, leverages Cus-MOE, with F0 supervision, enhancing both the synthesis quality and style modeling of the generated singing voice. Experimental results show that TCSinger 2 outperforms baseline models in both subjective and objective metrics across multiple related tasks.
【8】 QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
标题: QUADS:用于高效语音语言理解的QUANTized Distillation框架链接:https://arxiv.org/abs/2505.14723
备注:None
摘要:口语理解(SLU)系统必须平衡性能和效率,特别是在资源受限的环境中。现有的方法分别应用蒸馏和量化,导致次优压缩,因为蒸馏忽略了量化约束。我们提出了QUADS,这是一个统一的框架,通过预调整模型的多阶段训练来优化两者,在保持准确性的同时增强对低位制度的适应性。QUADS在SLURP上达到71.13\%的准确度,在FSC上达到99.20\%,与最先进的模型相比,只有高达5.56\%的轻微退化。此外,它减少了60- 73 $\times $(GMAC)的计算复杂度和83- 700 $\times $的模型大小,在极端量化下表现出强大的鲁棒性。这些结果建立QUADS作为一个高效的解决方案,为现实世界,资源受限的SLU应用。
摘要:Spoken Language Understanding (SLU) systems must balance performance and efficiency, particularly in resource-constrained environments. Existing methods apply distillation and quantization separately, leading to suboptimal compression as distillation ignores quantization constraints. We propose QUADS, a unified framework that optimizes both through multi-stage training with a pre-tuned model, enhancing adaptability to low-bit regimes while maintaining accuracy. QUADS achieves 71.13\% accuracy on SLURP and 99.20\% on FSC, with only minor degradations of up to 5.56\% compared to state-of-the-art models. Additionally, it reduces computational complexity by 60--73$\times$ (GMACs) and model size by 83--700$\times$, demonstrating strong robustness under extreme quantization. These results establish QUADS as a highly efficient solution for real-world, resource-constrained SLU applications.
【9】 MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
标题: MIKU-PAL:一种用于语音副语言和情感标签的自动化和标准化多模式方法链接:https://arxiv.org/abs/2505.15772
备注:Accepted by Interspeech
摘要:获取大规模的具有强一致性的情感语音数据仍然是语音合成的一个挑战。本文介绍了MIKU-PAL,一个全自动的多模式管道,用于从未标记的视频数据中提取高一致性的情感语音。利用人脸检测和跟踪算法,我们开发了一个自动情感分析系统,使用多模态大语言模型(MLLM)。我们的研究结果表明,MIKU-PAL可以实现人类水平的准确性(MELD上为68.5%)和优越的一致性(0.93 Fleiss kappa评分),同时比人类注释更便宜和更快。借助MIKU-PAL的高质量、灵活和一致的注释,我们可以注释多达26种类型的细粒度语音情感类别,并通过人工注释器进行验证,合理性评级为83%。基于我们提出的系统,我们进一步发布了一个细粒度的情感语音数据集MIKU-MIBench Bench(131.2小时),作为情感文本到语音和视觉语音克隆的新基准。
摘要:Acquiring large-scale emotional speech data with strong consistency remains a challenge for speech synthesis. This paper presents MIKU-PAL, a fully automated multimodal pipeline for extracting high-consistency emotional speech from unlabeled video data. Leveraging face detection and tracking algorithms, we developed an automatic emotion analysis system using a multimodal large language model (MLLM). Our results demonstrate that MIKU-PAL can achieve human-level accuracy (68.5% on MELD) and superior consistency (0.93 Fleiss kappa score) while being much cheaper and faster than human annotation. With the high-quality, flexible, and consistent annotation from MIKU-PAL, we can annotate fine-grained speech emotion categories of up to 26 types, validated by human annotators with 83% rationality ratings. Based on our proposed system, we further released a fine-grained emotional speech dataset MIKU-EmoBench(131.2 hours) as a new benchmark for emotional text-to-speech and visual voice cloning.
【10】 "Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
链接:https://arxiv.org/abs/2505.15700摘要:机器非学习,即从机器学习模型中有效去除特定信息的过程,是负责任人工智能越来越感兴趣的领域。然而,很少有研究探讨遗忘方法在复杂任务,特别是与语音相关的任务上的有效性。本文介绍了UnSLU-BENCH,这是口语理解(SLU)中机器非学习的第一个基准测试,重点关注四种语言的四个数据集。我们解决了从特定的扬声器数据的遗忘作为一种方式来评估潜在的“被遗忘的权利”的请求的质量。我们评估了八种非学习技术,并提出了一种新的度量标准,以同时更好地捕捉它们的功效、效用和效率。UnSLU-BENCH为SLU中的非学习奠定了基础,并揭示了各种技术的有效性和计算可行性的显着差异。
摘要:Machine unlearning, the process of efficiently removing specific information from machine learning models, is a growing area of interest for responsible AI. However, few studies have explored the effectiveness of unlearning methods on complex tasks, particularly speech-related ones. This paper introduces UnSLU-BENCH, the first benchmark for machine unlearning in spoken language understanding (SLU), focusing on four datasets spanning four languages. We address the unlearning of data from specific speakers as a way to evaluate the quality of potential "right to be forgotten" requests. We assess eight unlearning techniques and propose a novel metric to simultaneously better capture their efficacy, utility, and efficiency. UnSLU-BENCH sets a foundation for unlearning in SLU and reveals significant differences in the effectiveness and computational feasibility of various techniques.
【11】 Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
标题: 语音到语音语言模型的高效、直接的双重建模链接:https://arxiv.org/abs/2505.15670
备注:Accepted to Interspeech 2025
摘要:口语对话是一种直观的人机交互形式,然而目前的语音语言模型往往局限于基于话轮的交流,缺乏实时适应性,如用户闯入。我们提出了一种新的双工语音到语音(S2S)的体系结构,具有连续的用户输入和编解码器代理输出与信道融合,直接模拟同时用户和代理流。使用预训练的流编码器用于用户输入,使得第一双工S2S模型不需要语音预训练。代理和用户建模的单独架构便于编解码器微调更好的代理语音和比特率减半(0.6 kbps)相比,以前的作品。实验结果表明,该模型在推理、话轮转换和插话能力方面优于以往的双工模型。该模型需要的语音数据显著减少,因为跳过了语音预训练,这显著简化了从任何LLM构建双工S2S模型的过程。最后,它是第一个公开可用的双S2S模型,具有训练和推理代码,以促进再现性。
摘要:Spoken dialogue is an intuitive form of human-computer interaction, yet current speech language models often remain constrained to turn-based exchanges, lacking real-time adaptability such as user barge-in. We propose a novel duplex speech to speech (S2S) architecture featuring continuous user inputs and codec agent outputs with channel fusion that directly models simultaneous user and agent streams. Using a pretrained streaming encoder for user input enables the first duplex S2S model without requiring speech pretrain. Separate architectures for agent and user modeling facilitate codec fine-tuning for better agent voices and halve the bitrate (0.6 kbps) compared to previous works. Experimental results show that the proposed model outperforms previous duplex models in reasoning, turn-taking, and barge-in abilities. The model requires significantly less speech data, as speech pretrain is skipped, which markedly simplifies the process of building a duplex S2S model from any LLMs. Finally, it is the first openly available duplex S2S model with training and inference code to foster reproducibility.
【12】 Word Level Timestamp Generation for Automatic Speech Recognition and Translation
标题: 自动语音识别和翻译的单词级时间戳生成链接:https://arxiv.org/abs/2505.15646
备注:Accepted to Interspeech 2025
摘要:我们介绍了一种数据驱动的方法,使单词级的时间戳预测的金丝雀模型。准确的时间戳信息对于语音内容检索和定时字幕等各种下游任务至关重要。虽然传统的混合系统和端到端(E2E)模型可能会采用外部模块进行时间戳预测,但我们的方法消除了对单独对齐机制的需求。通过利用NeMo Forced Aligner(NFA)作为教师模型,我们生成单词级时间戳并训练Canary模型以直接预测时间戳。我们介绍一个新的<|时间戳|> token,使Canary模型能够预测每个单词的开始和结束时间戳。我们的方法表明,精确率和召回率在80%到90%之间,时间戳预测误差范围从20到120毫秒,四种语言,WER退化最小。此外,我们将我们的系统扩展到自动语音翻译(AST)任务,实现了大约200毫秒的时间戳预测误差。
摘要:We introduce a data-driven approach for enabling word-level timestamp prediction in the Canary model. Accurate timestamp information is crucial for a variety of downstream tasks such as speech content retrieval and timed subtitles. While traditional hybrid systems and end-to-end (E2E) models may employ external modules for timestamp prediction, our approach eliminates the need for separate alignment mechanisms. By leveraging the NeMo Forced Aligner (NFA) as a teacher model, we generate word-level timestamps and train the Canary model to predict timestamps directly. We introduce a new <|timestamp|> token, enabling the Canary model to predict start and end timestamps for each word. Our method demonstrates precision and recall rates between 80% and 90%, with timestamp prediction errors ranging from 20 to 120 ms across four languages, with minimal WER degradation. Additionally, we extend our system to automatic speech translation (AST) tasks, achieving timestamp prediction errors around 200 milliseconds.
【13】 Moonbeam: A MIDI Foundation Model Using Both Absolute and Relative Music Attributes
标题: Moonbeam:使用绝对和相对音乐属性的MIDI基金会模型链接:https://arxiv.org/abs/2505.15559
摘要:Moonbeam是一个基于transformer的符号音乐基础模型,在大量不同的音乐数据集合上进行了预训练,总计81.6K小时的音乐和180亿个令牌。Moonbeam通过引入一种新的领域知识启发的标记化方法和多维相对注意力(MRA)来捕获绝对和相对音乐属性,从而结合了音乐领域的归纳偏差,该方法捕获相对音乐信息而无需额外的可训练参数。利用预先训练的Moonbeam,我们提出了2个具有完全预期能力的微调架构,针对2类下游任务:符号音乐理解和条件音乐生成(包括音乐填充)。在大多数情况下,我们的模型在4个数据集上的3个下游音乐分类任务的准确性和F1得分方面优于其他大规模预训练的音乐模型。此外,我们的微调条件音乐生成模型优于一个强大的Transformer基线与一个类似于REMI的标记。我们开源了代码,预训练了模型,并在Github上生成了样本。
摘要:Moonbeam is a transformer-based foundation model for symbolic music, pretrained on a large and diverse collection of MIDI data totaling 81.6K hours of music and 18 billion tokens. Moonbeam incorporates music-domain inductive biases by capturing both absolute and relative musical attributes through the introduction of a novel domain-knowledge-inspired tokenization method and Multidimensional Relative Attention (MRA), which captures relative music information without additional trainable parameters. Leveraging the pretrained Moonbeam, we propose 2 finetuning architectures with full anticipatory capabilities, targeting 2 categories of downstream tasks: symbolic music understanding and conditional music generation (including music infilling). Our model outperforms other large-scale pretrained music models in most cases in terms of accuracy and F1 score across 3 downstream music classification tasks on 4 datasets. Moreover, our finetuned conditional music generation model outperforms a strong transformer baseline with a REMI-like tokenizer. We open-source the code, pretrained model, and generated samples on Github.
【14】 Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
标题: Audio Jailbreak:一个用于越狱大型音频语言模型的开放综合基准测试链接:https://arxiv.org/abs/2505.15406
备注:We release AJailBench, including both static and optimized adversarial data, to facilitate future research: this https URL
摘要:大型音频语言模型(LAMs)的兴起带来了潜在的风险,因为它们的音频输出可能包含有害或不道德的内容。然而,目前的研究缺乏一个系统的,定量的评估LAM的安全性,特别是对越狱攻击,这是具有挑战性的,由于语音的时间和语义的性质。为了弥补这一差距,我们引入AJailBench,这是第一个专门用于评估LAM中越狱漏洞的基准测试。我们首先构建AJailBench-Base,这是一个包含1,495个对抗性音频提示的数据集,涵盖10个违反策略的类别,从使用真实文本的文本越狱攻击转换为语音合成。使用这个数据集,我们评估了几个最先进的LAM,并发现没有一个在攻击中表现出一致的鲁棒性。为了进一步加强越狱测试和模拟更真实的攻击条件,我们提出了一种生成动态对抗变体的方法。我们的Audio Perturbation Toolkit(APT)可在时域、频域和振幅域应用目标失真。为了保持原始的越狱意图,我们强制执行语义一致性约束,并采用贝叶斯优化来有效地搜索既微妙又高效的扰动。这导致AJailBench-APT,这是一个优化的对抗性音频样本的扩展数据集。我们的研究结果表明,即使是很小的,语义保留的扰动可以显着降低领先的LAM的安全性能,强调需要更强大的和语义感知的防御机制。
摘要:The rise of Large Audio Language Models (LAMs) brings both potential and risks, as their audio outputs may contain harmful or unethical content. However, current research lacks a systematic, quantitative evaluation of LAM safety especially against jailbreak attacks, which are challenging due to the temporal and semantic nature of speech. To bridge this gap, we introduce AJailBench, the first benchmark specifically designed to evaluate jailbreak vulnerabilities in LAMs. We begin by constructing AJailBench-Base, a dataset of 1,495 adversarial audio prompts spanning 10 policy-violating categories, converted from textual jailbreak attacks using realistic text to speech synthesis. Using this dataset, we evaluate several state-of-the-art LAMs and reveal that none exhibit consistent robustness across attacks. To further strengthen jailbreak testing and simulate more realistic attack conditions, we propose a method to generate dynamic adversarial variants. Our Audio Perturbation Toolkit (APT) applies targeted distortions across time, frequency, and amplitude domains. To preserve the original jailbreak intent, we enforce a semantic consistency constraint and employ Bayesian optimization to efficiently search for perturbations that are both subtle and highly effective. This results in AJailBench-APT, an extended dataset of optimized adversarial audio samples. Our findings demonstrate that even small, semantically preserved perturbations can significantly reduce the safety performance of leading LAMs, underscoring the need for more robust and semantically aware defense mechanisms.
【15】 Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
标题: 通过上下文学习实现零拍语音转换的韵律自适应音频编解码器链接:https://arxiv.org/abs/2505.15402
备注:5 pages, 3 figures
摘要:离散音频编解码器的最新进展已经显著地改进了语音表示建模,而编解码器语言模型已经实现了用于zero-shot语音合成的上下文学习。受此启发,我们提出了VALLE-X框架内的语音转换(VC)模型,利用其强大的上下文学习能力进行说话人自适应。为了增强韵律控制,我们引入了一个韵律感知音频编解码器编码器(PACE)模块,该模块将韵律与其他源隔离并细化,从而提高表现力和控制力。通过将PACE集成到我们的VC模型中,我们在保留扬声器音色的同时实现了韵律操作的更大灵活性。实验评估结果表明,我们的方法优于基线VC系统的韵律保留,音色一致性,和整体的自然性,超过基线VC系统。
摘要:Recent advances in discrete audio codecs have significantly improved speech representation modeling, while codec language models have enabled in-context learning for zero-shot speech synthesis. Inspired by this, we propose a voice conversion (VC) model within the VALLE-X framework, leveraging its strong in-context learning capabilities for speaker adaptation. To enhance prosody control, we introduce a prosody-aware audio codec encoder (PACE) module, which isolates and refines prosody from other sources, improving expressiveness and control. By integrating PACE into our VC model, we achieve greater flexibility in prosody manipulation while preserving speaker timbre. Experimental evaluation results demonstrate that our approach outperforms baseline VC systems in prosody preservation, timbre consistency, and overall naturalness, surpassing baseline VC systems.
【16】 Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
标题: 利用语音推测解码加速自回归语音合成推理链接:https://arxiv.org/abs/2505.15380
备注:5 pages, 4 figures
摘要:利用语言模型的现代自回归语音合成模型已经表现出显着的性能。然而,这些模型中下一个令牌预测的顺序性质导致了显著的延迟,阻碍了它们在推理速度至关重要的场景中的部署。在这项工作中,我们提出了语音推测解码(SSD),自回归语音合成加速的新框架。具体来说,我们的方法采用了一个轻量级的草案模型来生成候选令牌序列,随后使用建议的SSD框架的目标模型并行验证。实验结果表明,SSD实现了显着的加速比为1.4倍相比,传统的自回归解码,同时保持高保真度和自然。主观评估进一步验证了SSD在保持目标模型的感知质量同时加速推理的有效性。
摘要:Modern autoregressive speech synthesis models leveraging language models have demonstrated remarkable performance. However, the sequential nature of next token prediction in these models leads to significant latency, hindering their deployment in scenarios where inference speed is critical. In this work, we propose Speech Speculative Decoding (SSD), a novel framework for autoregressive speech synthesis acceleration. Specifically, our method employs a lightweight draft model to generate candidate token sequences, which are subsequently verified in parallel by the target model using the proposed SSD framework. Experimental results demonstrate that SSD achieves a significant speedup of 1.4x compared with conventional autoregressive decoding, while maintaining high fidelity and naturalness. Subjective evaluations further validate the effectiveness of SSD in preserving the perceptual quality of the target model while accelerating inference.
【17】 Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
标题: Neurodyne:利用表示学习和周期一致性GAN的神经音调操纵链接:https://arxiv.org/abs/2505.15368
摘要:音高控制是指制作者将音频片段的音高调整到特定的音调和语调的过程,这在音乐制作中是必不可少的。基于神经网络的音高控制系统由于其优于传统的DSP方法的合成质量,近年来得到了广泛的应用。然而,它们的性能仍然有限,因为它们使用源过滤器模型进行不准确的特征分解,并且缺乏成对的调内和调外训练数据。这项工作提出了Neurodyne来解决这些问题。具体来说,Neurodyne使用对抗性表示学习来学习音高独立的潜在表示,以避免不准确的解纠缠和循环一致性训练,以隐式地创建成对的训练数据。全局键和基于模板的音高操作的实验结果表明,该系统的有效性,标记提高合成质量,同时保持原来的歌手身份。
摘要:Pitch manipulation is the process of producers adjusting the pitch of an audio segment to a specific key and intonation, which is essential in music production. Neural-network-based pitch-manipulation systems have been popular in recent years due to their superior synthesis quality compared to classical DSP methods. However, their performance is still limited due to their inaccurate feature disentanglement using source-filter models and the lack of paired in- and out-of-tune training data. This work proposes Neurodyne to address these issues. Specifically, Neurodyne uses adversarial representation learning to learn a pitch-independent latent representation to avoid inaccurate disentanglement and cycle-consistency training to create paired training data implicitly. Experimental results on global-key and template-based pitch manipulation demonstrate the effectiveness of the proposed system, marking improved synthesis quality while maintaining the original singer identity.
【18】 MHANet: Multi-scale Hybrid Attention Network for Auditory Attention Detection
标题: MHANet:用于听觉注意力检测的多尺度混合注意力网络链接:https://arxiv.org/abs/2505.15364
摘要:听觉注意检测(AAD)旨在从脑电信号中检测出多人说话环境中的目标说话人,如脑电信号(EEG),已经取得了很大的进展。然而,大多数AAD方法仅顺序地利用注意机制,并且忽略了EEG信号内的有价值的多尺度上下文信息,限制了它们同时捕获长-短范围时空依赖性的能力。为了解决这些问题,本文提出了一种多尺度混合注意力网络(MHANet)的AAD,它包括多尺度混合注意力(MHA)模块和时空卷积(STC)模块。具体而言,MHA结合了通道注意力和多尺度时间和全局注意力机制。这有效地提取多尺度的时间模式内的EEG信号,并捕获长,短距离的时空依赖性同时。为了进一步提高AAD的性能,STC利用时间和空间卷积来聚合表达时空表示。实验结果表明,所提出的MHANet实现了最先进的性能,在三个数据集上具有更少的可训练参数,比最先进的模型低3倍。代码可从以下网址获得:https://github.com/fchest/MHANet。
摘要:Auditory attention detection (AAD) aims to detect the target speaker in a multi-talker environment from brain signals, such as electroencephalography (EEG), which has made great progress. However, most AAD methods solely utilize attention mechanisms sequentially and overlook valuable multi-scale contextual information within EEG signals, limiting their ability to capture long-short range spatiotemporal dependencies simultaneously. To address these issues, this paper proposes a multi-scale hybrid attention network (MHANet) for AAD, which consists of the multi-scale hybrid attention (MHA) module and the spatiotemporal convolution (STC) module. Specifically, MHA combines channel attention and multi-scale temporal and global attention mechanisms. This effectively extracts multi-scale temporal patterns within EEG signals and captures long-short range spatiotemporal dependencies simultaneously. To further improve the performance of AAD, STC utilizes temporal and spatial convolutions to aggregate expressive spatiotemporal representations. Experimental results show that the proposed MHANet achieves state-of-the-art performance with fewer trainable parameters across three datasets, 3 times lower than that of the most advanced model. Code is available at: https://github.com/fchest/MHANet.
【19】 Decoding Phone Pairs from MEG Signals Across Speech Modalities
标题: 从跨语音模式的MEG信号解码电话对链接:https://arxiv.org/abs/2505.15355
备注:21 pages, 4 figures, 1 graphical abstract, submitted to Computer Speech and Language (special issue on Iberian Languages)
摘要:理解言语产生的神经机制对于推进认知神经科学理论和开发实用的通信技术都是至关重要的。在这项研究中,我们研究了脑磁图信号解码语音产生和感知(被动倾听和语音回放)任务期间的大脑活动的电话。使用包括17名参与者的数据集,我们进行了成对的电话分类,将我们的分析扩展到15个语音对。比较了多种机器学习方法,包括正则化线性模型和神经网络架构,以确定它们在解码语音信息方面的有效性。我们的研究结果表明,与被动倾听和回放模式(~51%)相比,语音产生过程中的解码准确率(76.6%)显着更高,强调了公开语音过程中可用的更丰富的神经信息。在这些模型中,弹性网络分类器的性能始终优于更复杂的神经网络,突出了传统正则化技术在应用于有限和高维MEG数据集时的有效性。此外,对特定脑频带的分析显示,低频振荡,特别是Delta(0.2-3 Hz)和Theta(4-7 Hz),对解码准确性的贡献最大,这表明这些频带编码关键的语音产生相关的神经过程。尽管使用了先进的去噪方法,但仍不清楚解码是否仅反映了神经活动,或者残余的肌肉或运动伪影是否也有贡献,这表明需要进一步改进方法。总的来说,我们的研究结果强调了检查明显的语音产生范式的至关重要性,尽管它们很复杂,但它们提供了改善脑机接口以帮助严重语音障碍患者的机会。
摘要:Understanding the neural mechanisms underlying speech production is essential for both advancing cognitive neuroscience theory and developing practical communication technologies. In this study, we investigated magnetoencephalography signals to decode phones from brain activity during speech production and perception (passive listening and voice playback) tasks. Using a dataset comprising 17 participants, we performed pairwise phone classification, extending our analysis to 15 phonetic pairs. Multiple machine learning approaches, including regularized linear models and neural network architectures, were compared to determine their effectiveness in decoding phonetic information. Our results demonstrate significantly higher decoding accuracy during speech production (76.6%) compared to passive listening and playback modalities (~51%), emphasizing the richer neural information available during overt speech. Among the models, the Elastic Net classifier consistently outperformed more complex neural networks, highlighting the effectiveness of traditional regularization techniques when applied to limited and high-dimensional MEG datasets. Besides, analysis of specific brain frequency bands revealed that low-frequency oscillations, particularly Delta (0.2-3 Hz) and Theta (4-7 Hz), contributed the most substantially to decoding accuracy, suggesting that these bands encode critical speech production-related neural processes. Despite using advanced denoising methods, it remains unclear whether decoding solely reflects neural activity or if residual muscular or movement artifacts also contributed, indicating the need for further methodological refinement. Overall, our findings underline the critical importance of examining overt speech production paradigms, which, despite their complexity, offer opportunities to improve brain-computer interfaces to help individuals with severe speech impairments.
【20】 Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
标题: 利用单元语言指导推进无文本语音翻译中的语音建模链接:https://arxiv.org/abs/2505.15333
备注:Accepted to ACL 2025 Findings
摘要:无文本语音到语音翻译(S2 ST)模型的成功构建引起了人们的广泛关注。然而,S2 ST仍然面临两个主要挑战:1)提取各种语音信号的语言特征,称为跨模态(CM),以及2)学习长序列中不同语言的对齐,称为跨语言(CL)。我们提出的单位语言,以克服这两个建模的挑战。单元语言可以被认为是一种类似文本的表示格式,使用$n$-gram语言建模构建。我们实现了多任务学习,利用单位语言指导语音建模过程。我们的初步结果揭示了冲突时,同时应用源和目标单元语言。我们建议任务提示建模,以减轻这种冲突。我们在Voxpupil数据集的四种语言上进行实验。我们的方法在强大的基线上表现出显着的改进,并实现了与使用文本训练的模型相当的性能。
摘要:The success of building textless speech-to-speech translation (S2ST) models has attracted much attention. However, S2ST still faces two main challenges: 1) extracting linguistic features for various speech signals, called cross-modal (CM), and 2) learning alignment of difference languages in long sequences, called cross-lingual (CL). We propose the unit language to overcome the two modeling challenges. The unit language can be considered a text-like representation format, constructed using $n$-gram language modeling. We implement multi-task learning to utilize the unit language in guiding the speech modeling process. Our initial results reveal a conflict when applying source and target unit languages simultaneously. We propose task prompt modeling to mitigate this conflict. We conduct experiments on four languages of the Voxpupil dataset. Our method demonstrates significant improvements over a strong baseline and achieves performance comparable to models trained with text.
【21】 Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
标题: Voice-ENHANCE:使用基于扩散的语音转换框架进行语音恢复链接:https://arxiv.org/abs/2505.15254
备注:5 pages, 3 figures, Accepted to INTERSPEECH 2025
摘要:我们提出了一个语音增强系统,结合说话人不可知的语音恢复与语音转换(VC),以获得演播室级质量的语音信号。虽然语音转换模型通常用于改变说话者的特征,但当目标说话者与源说话者相同时,它们也可以用作语音恢复的手段。然而,由于VC模型容易受到噪声条件的影响,我们在我们提出的系统的前端包括了一个生成语音恢复(GSR)模型。GSR模型执行噪声抑制,并在不了解目标说话者的情况下恢复在该过程中发生的语音损伤。VC阶段然后使用来自干净说话者嵌入的指导来进一步恢复输出语音。通过采用这种两阶段的方法,我们在多个数据集上实现了与最先进的(SOTA)方法相当的语音质量客观度量分数。
摘要:We propose a speech enhancement system that combines speaker-agnostic speech restoration with voice conversion (VC) to obtain a studio-level quality speech signal. While voice conversion models are typically used to change speaker characteristics, they can also serve as a means of speech restoration when the target speaker is the same as the source speaker. However, since VC models are vulnerable to noisy conditions, we have included a generative speech restoration (GSR) model at the front end of our proposed system. The GSR model performs noise suppression and restores speech damage incurred during that process without knowledge about the target speaker. The VC stage then uses guidance from clean speaker embeddings to further restore the output speech. By employing this two-stage approach, we have achieved speech quality objective metric scores comparable to state-of-the-art (SOTA) methods across multiple datasets.
【22】 Hybrid Audio Detection Using Fine-Tuned Audio Spectrogram Transformers: A Dataset-Driven Evaluation of Mixed AI-Human Speech
标题: 使用微调音频频谱图变换器的混合音频检测:人工智能-人类混合语音的数据集驱动评估链接:https://arxiv.org/abs/2505.15136
备注:11 pages, 6 figures
摘要:人工智能(AI)的快速发展使复杂的音频生成和语音克隆技术成为可能,这对依赖语音认证的应用程序构成了重大的安全风险。虽然现有的数据集和模型主要集中在区分人类和完全合成的语音,但现实世界的攻击通常涉及结合了真实和克隆片段的音频。为了解决这一差距,我们构建了一个新的混合音频数据集,其中包含人类,AI生成,克隆和混合音频样本。我们进一步提出了微调音频频谱图Transformer(AST)为基础的模型检测这些复杂的声学模式。大量的实验表明,我们的方法显着优于现有的混合音频检测基线,达到97%的分类准确率。我们的研究结果强调了混合数据集和定制模型在提高基于语音的认证系统的鲁棒性方面的重要性。
摘要:The rapid advancement of artificial intelligence (AI) has enabled sophisticated audio generation and voice cloning technologies, posing significant security risks for applications reliant on voice authentication. While existing datasets and models primarily focus on distinguishing between human and fully synthetic speech, real-world attacks often involve audio that combines both genuine and cloned segments. To address this gap, we construct a novel hybrid audio dataset incorporating human, AI-generated, cloned, and mixed audio samples. We further propose fine-tuned Audio Spectrogram Transformer (AST)-based models tailored for detecting these complex acoustic patterns. Extensive experiments demonstrate that our approach significantly outperforms existing baselines in mixed-audio detection, achieving 97\% classification accuracy. Our findings highlight the importance of hybrid datasets and tailored models in advancing the robustness of speech-based authentication systems.
【23】 SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
标题: SHEET:多用途开源语音人类评估估计工具包链接:https://arxiv.org/abs/2505.15061
备注:INTERSPEECH 2025. Codebase: this https URL
摘要:我们介绍SHEET,一个多用途的开源工具包,旨在加速主观语音质量评估(SSQA)的研究。SHEET代表Speech Human Evaluation Estimation Toolkit,它专注于数据驱动的基于深度神经网络的模型,这些模型经过训练,可以预测语音样本的人类标记质量分数。SHEET提供全面的训练和评估脚本,多数据集和多模型支持,以及通过Torch Hub和HuggingFace Spaces访问的预训练模型。为了证明它的能力,我们重新评估了SSL-MOS,这是一种基于语音自监督学习(SSL)的SSQA模型,在最近的科学论文中广泛使用。在两个具有代表性的SSQA数据集BVCC和NISQA上进行了实验,我们确定了最佳语音SSL模型,其性能超过了原始SSL-MOS实现,与最先进的方法相当。
摘要:We introduce SHEET, a multi-purpose open-source toolkit designed to accelerate subjective speech quality assessment (SSQA) research. SHEET stands for the Speech Human Evaluation Estimation Toolkit, which focuses on data-driven deep neural network-based models trained to predict human-labeled quality scores of speech samples. SHEET provides comprehensive training and evaluation scripts, multi-dataset and multi-model support, as well as pre-trained models accessible via Torch Hub and HuggingFace Spaces. To demonstrate its capabilities, we re-evaluated SSL-MOS, a speech self-supervised learning (SSL)-based SSQA model widely used in recent scientific papers, on an extensive list of speech SSL models. Experiments were conducted on two representative SSQA datasets named BVCC and NISQA, and we identified the optimal speech SSL model, whose performance surpassed the original SSL-MOS implementation and was comparable to state-of-the-art methods.
【24】 AsynFusion: Towards Asynchronous Latent Consistency Models for Decoupled Whole-Body Audio-Driven Avatars
标题: AsynFusion:面向脱钩全身音频驱动化身的同步性模型链接:https://arxiv.org/abs/2505.15058
备注:11pages, conference
摘要:全身音频驱动的化身姿态和表情生成是创建逼真的数字人和增强交互式虚拟代理的能力的关键任务,在虚拟现实,数字娱乐和远程通信中具有广泛的应用。现有的方法通常独立地生成音频驱动的面部表情和手势,这引入了一个显著的限制:面部和手势元素之间缺乏无缝协调,导致不太自然和有凝聚力的动画。为了解决这个问题,我们提出了AsynFusion,一个新的框架,利用扩散Transformers来实现和谐的表达和手势合成。所提出的方法是建立在一个双分支DiT架构,使面部表情和手势的并行生成。在该模型中,我们引入了一个协作同步模块,以促进两种模式之间的双向功能交互,以及一个异步LCM采样策略,以减少计算开销,同时保持高质量的输出。大量实验表明,AsynFusion在生成实时同步全身动画方面达到了最先进的性能,在定量和定性评估方面始终优于现有方法。
摘要:Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital entertainment, and remote communication. Existing approaches often generate audio-driven facial expressions and gestures independently, which introduces a significant limitation: the lack of seamless coordination between facial and gestural elements, resulting in less natural and cohesive animations. To address this limitation, we propose AsynFusion, a novel framework that leverages diffusion transformers to achieve harmonious expression and gesture synthesis. The proposed method is built upon a dual-branch DiT architecture, which enables the parallel generation of facial expressions and gestures. Within the model, we introduce a Cooperative Synchronization Module to facilitate bidirectional feature interaction between the two modalities, and an Asynchronous LCM Sampling strategy to reduce computational overhead while maintaining high-quality outputs. Extensive experiments demonstrate that AsynFusion achieves state-of-the-art performance in generating real-time, synchronized whole-body animations, consistently outperforming existing methods in both quantitative and qualitative evaluations.
【25】 Discrete Audio Representations for Automated Audio Captioning
标题: 用于自动音频字幕的离散音频表示链接:https://arxiv.org/abs/2505.14989
备注:Interspeech 2025
摘要:离散音频表示,称为音频令牌,被广泛地分类为语义和声学令牌,通常通过连续音频表示的无监督令牌化生成。然而,它们对自动音频字幕(AAC)的适用性仍然有待探索。本文通过对各种标记化方法的比较分析,系统地研究了AAC音频标记驱动模型的可行性。我们的研究结果表明,音频标记化导致AAC模型的性能下降相比,那些直接利用连续的音频表示。为了解决这个问题,我们引入了一个有监督的音频标记训练的音频标记目标。与缺乏明确语义理解的无监督标记器不同,所提出的标记器有效地捕获音频事件信息。在Clotho数据集上进行的实验表明,所提出的音频令牌在AAC任务中优于传统的音频令牌。
摘要:Discrete audio representations, termed audio tokens, are broadly categorized into semantic and acoustic tokens, typically generated through unsupervised tokenization of continuous audio representations. However, their applicability to automated audio captioning (AAC) remains underexplored. This paper systematically investigates the viability of audio token-driven models for AAC through comparative analyses of various tokenization methods. Our findings reveal that audio tokenization leads to performance degradation in AAC models compared to those that directly utilize continuous audio representations. To address this issue, we introduce a supervised audio tokenizer trained with an audio tagging objective. Unlike unsupervised tokenizers, which lack explicit semantic understanding, the proposed tokenizer effectively captures audio event information. Experiments conducted on the Clotho dataset demonstrate that the proposed audio tokens outperform conventional audio tokens in the AAC task.
【26】 In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
标题: 上下文学习通过对说话者和语言多样性的类人适应来促进语音识别链接:https://arxiv.org/abs/2505.14887
备注:15 pages; 3 figures
摘要:人类听众很容易通过接触来适应不熟悉的说话者和语言变体,但这些适应益处是否延伸到最先进的口语模型?我们介绍了一个可扩展的框架,允许在Phi-4多模态使用交错的任务提示和音频文本对的上下文学习(ICL),并发现在推理时间只有12个例子(~50秒)减少单词错误率相对19.7%(1.2页)。在不同英语语料库中的平均值。这些改进在低资源品种中最为明显,当上下文和目标说话者匹配时,以及当提供更多示例时-尽管缩放我们的过程会使上下文长度的边际收益递减。总的来说,我们发现我们的新型ICL自适应方案(1)揭示了与人类听众相似的性能特征,(2)在不同的说话人和语言背景下,自动语音识别(ASR)的鲁棒性得到了一致的改善。虽然适应取得了广泛的成功,但某些品种仍然存在重大差距,揭示了目前的模式仍然缺乏人类的灵活性。我们在GitHub上发布提示和代码。
摘要:Human listeners readily adjust to unfamiliar speakers and language varieties through exposure, but do these adaptation benefits extend to state-of-the-art spoken language models? We introduce a scalable framework that allows for in-context learning (ICL) in Phi-4 Multimodal using interleaved task prompts and audio-text pairs, and find that as few as 12 example utterances (~50 seconds) at inference time reduce word error rates by a relative 19.7% (1.2 pp.) on average across diverse English corpora. These improvements are most pronounced in low-resource varieties, when the context and target speaker match, and when more examples are provided--though scaling our procedure yields diminishing marginal returns to context length. Overall, we find that our novel ICL adaptation scheme (1) reveals a similar performance profile to human listeners, and (2) demonstrates consistent improvements to automatic speech recognition (ASR) robustness across diverse speakers and language backgrounds. While adaptation succeeds broadly, significant gaps remain for certain varieties, revealing where current models still fall short of human flexibility. We release our prompts and code on GitHub.
【27】 Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
标题: 迈向包容性ASB:研究低资源语言中用于发音障碍语音识别的语音转换链接:https://arxiv.org/abs/2505.14874
备注:5 pages, 1 figure, Accepted to Interspeech 2025
摘要:由于数据稀缺,特别是在非英语语言中,构音障碍语音的自动语音识别(ASR)仍然具有挑战性。为了解决这个问题,我们微调英语构音障碍语音(UASpeech)的语音转换模型,对说话者特征和韵律失真进行编码,然后将其应用于将健康的非英语语音(FLEURS)转换为非英语构音障碍类语音。然后,生成的数据用于微调多语言ASR模型,大规模多语言语音(MMS),以改善构音障碍的语音识别。对PC-GITA(西班牙语),EasyCall(意大利语)和SSNCE(泰米尔语)的评估表明,VC与扬声器和韵律转换显着优于现成的MMS性能和传统的增强技术,如速度和节奏扰动。客观和主观的分析所产生的数据进一步证实,所产生的语音模拟构音障碍的特点。
摘要:Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, particularly in non-English languages. To address this, we fine-tune a voice conversion model on English dysarthric speech (UASpeech) to encode both speaker characteristics and prosodic distortions, then apply it to convert healthy non-English speech (FLEURS) into non-English dysarthric-like speech. The generated data is then used to fine-tune a multilingual ASR model, Massively Multilingual Speech (MMS), for improved dysarthric speech recognition. Evaluation on PC-GITA (Spanish), EasyCall (Italian), and SSNCE (Tamil) demonstrates that VC with both speaker and prosody conversion significantly outperforms the off-the-shelf MMS performance and conventional augmentation techniques such as speed and tempo perturbation. Objective and subjective analyses of the generated data further confirm that the generated speech simulates dysarthric characteristics.
【28】 Replay Attacks Against Audio Deepfake Detection
标题: 针对音频Deepfake检测的重播攻击链接:https://arxiv.org/abs/2505.14862
备注:None
摘要:我们展示了重放攻击如何破坏音频deepfake检测:通过各种扬声器和麦克风播放和重新录制deepfake音频,我们使欺骗样本在检测模型中看起来是真实的。为了更详细地研究这一现象,我们引入了ReplayDF,这是一个来自M-AILABS和MLAAD的录音数据集,其中包括六种语言和四种TTS模型的109种扬声器-麦克风组合。它包括各种声学条件,其中一些对检测具有高度挑战性。我们对五个数据集的六个开源检测模型的分析揭示了显著的漏洞,表现最好的W2 V2-AASIST模型的等错误率(EER)从4.7%飙升至18.2%。即使采用自适应房间脉冲响应(RIR)再训练,性能仍然受到11.0% EER的影响。我们发布ReplayDF用于非商业研究用途。
摘要:We show how replay attacks undermine audio deepfake detection: By playing and re-recording deepfake audio through various speakers and microphones, we make spoofed samples appear authentic to the detection model. To study this phenomenon in more detail, we introduce ReplayDF, a dataset of recordings derived from M-AILABS and MLAAD, featuring 109 speaker-microphone combinations across six languages and four TTS models. It includes diverse acoustic conditions, some highly challenging for detection. Our analysis of six open-source detection models across five datasets reveals significant vulnerability, with the top-performing W2V2-AASIST model's Equal Error Rate (EER) surging from 4.7% to 18.2%. Even with adaptive Room Impulse Response (RIR) retraining, performance remains compromised with an 11.0% EER. We release ReplayDF for non-commercial research use.
【29】 GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples
标题: GraphemeAug:一个系统化的方法来综合硬否定关键词发现的例子链接:https://arxiv.org/abs/2505.14814
备注:Accepted at Interspeech 2025
摘要:口语关键词识别(英语:Spoken Keyword Spotting,简称KWS)是一项区分音频中关键词的存在和不存在的任务。KWS模型的准确性取决于其正确分类接近关键字和非关键字边界的示例的能力。这些边界示例在训练数据中通常很少,限制了模型性能。在本文中,我们提出了一种方法,通过对关键字的字素进行插入/删除/替换编辑,系统地生成接近决策边界的对抗性示例。我们对流行关键字的保留数据进行了评估,并表明该技术将合成硬否定数据集的AUC提高了61%,同时保持了阳性和环境负面音频数据的质量。
摘要:Spoken Keyword Spotting (KWS) is the task of distinguishing between the presence and absence of a keyword in audio. The accuracy of a KWS model hinges on its ability to correctly classify examples close to the keyword and non-keyword boundary. These boundary examples are often scarce in training data, limiting model performance. In this paper, we propose a method to systematically generate adversarial examples close to the decision boundary by making insertion/deletion/substitution edits on the keyword's graphemes. We evaluate this technique on held-out data for a popular keyword and show that the technique improves AUC on a dataset of synthetic hard negatives by 61% while maintaining quality on positives and ambient negative audio data.
机器翻译由腾讯交互翻译提供,仅供参考
