微信公众号:arXiv_Daily
cs.SD语音
标题: Vox-Profile:描述不同说话者和言语特征的言语基金会模型基准
链接:https://arxiv.org/abs/2505.14648
摘要:我们介绍Vox-Profile,一个全面的基准测试,使用语音基础模型来表征丰富的扬声器和语音特征。与现有的专注于说话者特质的单一维度的作品不同,Vox-Profile提供了反映静态说话者特质(例如,年龄、性别、口音)和动态语音属性(例如,情感、言语流)。该基准以语音科学和语言学为基础,与领域专家一起开发,以准确地索引说话者和语音特征。我们使用超过15个公开的语音数据集和几个广泛使用的语音基础模型,针对各种静态和动态扬声器和语音属性的基准实验报告。除了基准测试实验,我们展示了几个下游应用程序支持的Vox-Profile。首先,我们证明了Vox-Profile可以增强现有的语音识别数据集,以分析ASR性能的变化。Vox-Profile也被用作评估语音生成系统性能的工具。最后,我们评估我们的自动配置文件的质量,通过与人类的评价比较,并显示收敛的有效性。Vox-Profile可在https://github.com/tiantiaf0627/vox-profile-release上公开获取。
摘要:We introduce Vox-Profile, a comprehensive benchmark to characterize rich speaker and speech traits using speech foundation models. Unlike existing works that focus on a single dimension of speaker traits, Vox-Profile provides holistic and multi-dimensional profiles that reflect both static speaker traits (e.g., age, sex, accent) and dynamic speech properties (e.g., emotion, speech flow). This benchmark is grounded in speech science and linguistics, developed with domain experts to accurately index speaker and speech characteristics. We report benchmark experiments using over 15 publicly available speech datasets and several widely used speech foundation models that target various static and dynamic speaker and speech properties. In addition to benchmark experiments, we showcase several downstream applications supported by Vox-Profile. First, we show that Vox-Profile can augment existing speech recognition datasets to analyze ASR performance variability. Vox-Profile is also used as a tool to evaluate the performance of speech generation systems. Finally, we assess the quality of our automated profiles through comparison with human evaluation and show convergent validity. Vox-Profile is publicly available at: https://github.com/tiantiaf0627/vox-profile-release.
标题: 语言、音频和视觉模式语义对齐的表示学习
链接:https://arxiv.org/abs/2505.14562
备注:Accepted to European Signal Processing Conference (EUSIPCO 2025)
摘要:本文提出了一种单阶段训练方法,使用对比学习框架在语义上对齐三种模态-音频,视觉和文本。对比训练在多模态对齐方面取得了突出的成就,利用大规模未标记数据来学习共享表示。现有的三模态对齐的深度学习方法包括两个阶段,分别对齐视觉-文本和音频-文本模态。这种方法遭受不匹配的数据分布,导致次优对齐。利用AVCaps数据集,它为视频剪辑提供音频,视频和视听字幕,我们的方法使用对比训练联合优化了所有模态的表示。我们的研究结果表明,单阶段的方法优于两阶段的方法,实现了基于音频的视觉检索的两倍的改善,突出了统一的多模态表示学习的优势。
摘要:This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive training has gained prominence for multimodal alignment, utilizing large-scale unlabeled data to learn shared representations. Existing deep learning approach for trimodal alignment involves two-stages, that separately align visual-text and audio-text modalities. This approach suffers from mismatched data distributions, resulting in suboptimal alignment. Leveraging the AVCaps dataset, which provides audio, visual and audio-visual captions for video clips, our method jointly optimizes the representation of all the modalities using contrastive training. Our results demonstrate that the single-stage approach outperforms the two-stage method, achieving a two-fold improvement in audio based visual retrieval, highlighting the advantages of unified multimodal representation learning.
标题: 过去:声控语音令牌器
链接:https://arxiv.org/abs/2505.14470
摘要:我们提出了PAST,这是一种新型的端到端框架,它在信号重建的同时联合对语音信息进行建模,从而消除了对外部预训练模型的需求。与以前依赖于预训练的自监督模型的方法不同,PAST采用监督语音数据,通过辅助任务将领域知识直接集成到标记化过程中。此外,我们引入了一个流,因果的PAST变体,使实时语音应用程序。结果表明,PAST超越了现有的评估基线标记在共同的评估指标,包括语音表示和语音重建。值得注意的是,PAST在作为语音语言模型的语音表示时也实现了卓越的性能,进一步突出了其作为口语生成基础的有效性。为了促进进一步的研究,我们发布了完整的实现。有关代码、模型检查点和示例,请参见:https://pages.cs.huji.ac.il/adiyoss-lab/PAST
摘要:We present PAST, a novel end-to-end framework that jointly models phonetic information alongside signal reconstruction, eliminating the need for external pretrained models. Unlike previous approaches that rely on pretrained self-supervised models, PAST employs supervised phonetic data, directly integrating domain knowledge into the tokenization process via auxiliary tasks. Additionally, we introduce a streamable, causal variant of PAST, enabling real-time speech applications. Results demonstrate that PAST surpasses existing evaluated baseline tokenizers across common evaluation metrics, including phonetic representation and speech reconstruction. Notably, PAST also achieves superior performance when serving as a speech representation for speech language models, further highlighting its effectiveness as a foundation for spoken language generation. To foster further research, we release the full implementation. For code, model checkpoints, and samples see: https://pages.cs.huji.ac.il/adiyoss-lab/PAST
标题: 低音中提琴中提琴频率波动的复杂性和诠释风格
链接:https://arxiv.org/abs/2505.14448
备注:15 pages, 5 figures
摘要:将一组音乐作品中的音频信号建模为复杂网络,以研究低音中提琴频率波动的复杂性与演奏风格之间的关系。基于跨学科的科学和音乐方法,我们计算频谱分解并将其频率分量转换为声音网络。我们应用最佳拟合分析来识别更精确地描述此类频率行为的统计分布,并计算中心性度量并识别用于表征此类网络的集团。研究结果表明,统计分布的类型,最好地描述了频率波动的统计分布。中心性测度确定了一段音乐中最有影响力和最稳定的声音组,同时最大集团的识别表明了声音的功能组,这些功能组密切相互作用,以识别复杂的频率波动的出现。因此,通过将声音建模为复杂网络,我们可以清楚地将大规模统计波动的存在与同一音乐家演奏的不同音乐事件相关的类似频率波动的存在联系起来。
摘要:Audio signals in a set of musical pieces are modeled as a complex network for studying the relationship between the complexity of frequency fluctuations and the interpretive style of the bass viola da gamba. Based on interdisciplinary scientific and music approaches, we compute the spectral decomposition and translated its frequency components to a network of sounds. We applied a best fit analysis for identifying the statistical distributions that describe more precisely the behavior of such frequencies and computed the centrality measures and identify cliques for characterizing such a network. Findings suggested statistical regularities in the type of statistical distribution that best describes frequency fluctuations. The centrality measure confirmed the most influential and stable group of sounds in a piece of music, meanwhile the identification of the largest clique indicated functional groups of sounds that interact closely for identifying the emergence of complex frequency fluctuations. Therefore, by modeling the sound as a complex network, we can clearly associate the presence of large-scale statistical regularities with the presence of similar frequency fluctuations related to different musical events played by a same musician.
标题: S2 SBench:量化语音到语音大型语言模型中智力退化的基准
链接:https://arxiv.org/abs/2505.14438
摘要:端到端语音大语言模型(LLM)扩展了基于文本的模型的功能,可以直接处理和生成音频令牌。然而,与文本输入相比,这通常会导致推理和生成性能下降,这种现象被称为智能退化。为了系统地评估这一差距,我们提出了S2SBench,一个旨在量化语音LLM性能下降的基准。它包括针对音频输入下的句子延续和常识推理的诊断数据集。我们进一步介绍了一个成对的评估协议的基础上的困惑之间的差异似是而非的和难以置信的样本来衡量退化相对于文本输入。我们应用S2SBench对百川音频的训练过程进行了分析,进一步验证了该基准的有效性。所有数据集和评估代码都可以在https://github.com/undobug/S2SBench上找到。
摘要:End-to-end speech large language models ((LLMs)) extend the capabilities of text-based models to directly process and generate audio tokens. However, this often leads to a decline in reasoning and generation performance compared to text input, a phenomenon referred to as intelligence degradation. To systematically evaluate this gap, we propose S2SBench, a benchmark designed to quantify performance degradation in Speech LLMs. It includes diagnostic datasets targeting sentence continuation and commonsense reasoning under audio input. We further introduce a pairwise evaluation protocol based on perplexity differences between plausible and implausible samples to measure degradation relative to text input. We apply S2SBench to analyze the training process of Baichuan-Audio, which further demonstrates the benchmark's effectiveness. All datasets and evaluation code are available at https://github.com/undobug/S2SBench.
标题: PersonaTab:在全复式语音对话中使用文本、声学和行为线索预测性格特征
链接:https://arxiv.org/abs/2505.14356
备注:This is accepted to Interspeech 2025; Added an extra page for supplementary figures; Project page: this https URL
摘要:尽管在神经口语对话系统中取得了重大进展,但由于语音数据集中缺乏个性注释,个性感知对话代理(能够根据个性调整行为)仍然未得到充分研究。我们提出了一个预处理原始音频记录的管道,以创建一个带有时间戳,响应类型和情绪/情感标签的对话数据集。我们采用自动语音识别(ASR)系统来提取成绩单和时间戳,然后生成会话级注释。利用这些注释,我们设计了一个系统,采用大型语言模型来预测会话个性。人类评估人员参与识别会话特征并分配个性标签。我们的分析表明,该系统实现了更强的对齐与人类的判断相比,现有的方法。
摘要:Despite significant progress in neural spoken dialog systems, personality-aware conversation agents -- capable of adapting behavior based on personalities -- remain underexplored due to the absence of personality annotations in speech datasets. We propose a pipeline that preprocesses raw audio recordings to create a dialogue dataset annotated with timestamps, response types, and emotion/sentiment labels. We employ an automatic speech recognition (ASR) system to extract transcripts and timestamps, then generate conversation-level annotations. Leveraging these annotations, we design a system that employs large language models to predict conversational personality. Human evaluators were engaged to identify conversational characteristics and assign personality labels. Our analysis demonstrates that the proposed system achieves stronger alignment with human judgments compared to existing approaches.
标题: FMSD-TTC:用于Deliver-Tsang、Amdo和Kham语音数据集生成的Few-Shot多说话人多方言文本到语音合成
链接:https://arxiv.org/abs/2505.14351
备注:13 pages
摘要:藏语是一种低资源的语言,只有最少的平行语音语料库跨越其三大方言--“乌斯藏语、安多语和康语--限制了语音建模的进展。为了解决这个问题,我们提出了FMSD-TTS,一个Few-Shot,多扬声器,多方言的文本到语音的框架,合成并行方言语音从有限的参考音频和明确的方言标签。我们的方法具有一个新的扬声器方言融合模块和方言专用动态路由网络(DSDR-Net),以捕获细粒度的声音和语言的方言变化,同时保留扬声器的身份。大量的客观和主观评估表明,FMSD-TTS显着优于基线的方言表达和说话人相似性。我们通过一个具有挑战性的语音到语音的方言转换任务,进一步验证合成语音的质量和效用。我们的贡献包括:(1)一个为藏语多方言语音合成量身定制的新型Few-Shot TTS系统,(2)公开发布由FMSD-TTS生成的大规模合成藏语语音语料库,(3)一个开源评估工具包,用于标准化评估说话人相似性、方言一致性和音频质量。
摘要:Tibetan is a low-resource language with minimal parallel speech corpora spanning its three major dialects-\"U-Tsang, Amdo, and Kham-limiting progress in speech modeling. To address this issue, we propose FMSD-TTS, a few-shot, multi-speaker, multi-dialect text-to-speech framework that synthesizes parallel dialectal speech from limited reference audio and explicit dialect labels. Our method features a novel speaker-dialect fusion module and a Dialect-Specialized Dynamic Routing Network (DSDR-Net) to capture fine-grained acoustic and linguistic variations across dialects while preserving speaker identity. Extensive objective and subjective evaluations demonstrate that FMSD-TTS significantly outperforms baselines in both dialectal expressiveness and speaker similarity. We further validate the quality and utility of the synthesized speech through a challenging speech-to-speech dialect conversion task. Our contributions include: (1) a novel few-shot TTS system tailored for Tibetan multi-dialect speech synthesis, (2) the public release of a large-scale synthetic Tibetan speech corpus generated by FMSD-TTS, and (3) an open-source evaluation toolkit for standardized assessment of speaker similarity, dialect consistency, and audio quality.
标题: 语音灵活控制的通用声学对抗攻击-LLM
链接:https://arxiv.org/abs/2505.14286
摘要:预先训练的语音编码器与大型语言模型的组合使得能够开发能够处理各种口语处理任务的语音LLM。虽然这些模型功能强大且灵活,但这种灵活性可能使它们更容易受到对抗性攻击。为了研究这个问题的程度,在这项工作中,我们研究了通用的声音对抗攻击语音LLM。在这里,原始输入音频中预先添加了固定的、通用的、对抗性的音频片段。我们最初调查的攻击,导致模型要么不产生输出或执行修改后的任务覆盖原来的提示。然后,我们将攻击的性质扩展为选择性的,以便仅当存在特定的输入属性(例如说话者性别或口语)时才激活。没有目标属性的输入应该不受影响,允许对模型输出进行细粒度控制。我们的研究结果揭示了Qwen 2-Audio和Granite-Speech中的关键漏洞,并表明类似的语音LLM可能容易受到普遍对抗性攻击。这突出表明需要更强大的培训策略和提高对抗性攻击的抵抗力。
摘要:The combination of pre-trained speech encoders with large language models has enabled the development of speech LLMs that can handle a wide range of spoken language processing tasks. While these models are powerful and flexible, this very flexibility may make them more vulnerable to adversarial attacks. To examine the extent of this problem, in this work we investigate universal acoustic adversarial attacks on speech LLMs. Here a fixed, universal, adversarial audio segment is prepended to the original input audio. We initially investigate attacks that cause the model to either produce no output or to perform a modified task overriding the original prompt. We then extend the nature of the attack to be selective so that it activates only when specific input attributes, such as a speaker gender or spoken language, are present. Inputs without the targeted attribute should be unaffected, allowing fine-grained control over the model outputs. Our findings reveal critical vulnerabilities in Qwen2-Audio and Granite-Speech and suggest that similar speech LLMs may be susceptible to universal adversarial attacks. This highlights the need for more robust training strategies and improved resistance to adversarial attacks.
标题: AcquaSignal:鲁棒水下声学分析的集成框架
链接:https://arxiv.org/abs/2505.14285
备注:8 pages; 9 figures
摘要:本文介绍了AquaSignal,这是一个模块化和可扩展的管道,用于水声信号的预处理,去噪,分类和新颖性检测。AquaSignal旨在在嘈杂和动态的海洋环境中有效运行,集成了最先进的深度学习架构,以提高声学信号分析的可靠性和准确性。该系统在Deepship和Ocean Networks Canada(ONC)基准的组合数据集上进行评估,提供了一组不同的真实水下场景。AquaSignal采用U-Net架构进行去噪,使用ResNet 18卷积神经网络对已知声学事件进行分类,并使用基于AutoEncoder的模型对新信号或异常信号进行无监督检测。据我们所知,这是第一次全面的研究,应用和评估这种技术组合的海上船舶声学数据。实验结果表明,AquaSignal提高了信号的清晰度和任务性能,实现了71%的分类准确率和91%的新奇检测准确率。尽管与一些最先进的模型相比,分类性能略低,但数据划分策略的差异限制了直接比较。总体而言,AquaSignal在科学,环境和海洋领域的实时水声监测方面表现出强大的潜力。
摘要:This paper presents AquaSignal, a modular and scalable pipeline for preprocessing, denoising, classification, and novelty detection of underwater acoustic signals. Designed to operate effectively in noisy and dynamic marine environments, AquaSignal integrates state-of-the-art deep learning architectures to enhance the reliability and accuracy of acoustic signal analysis. The system is evaluated on a combined dataset from the Deepship and Ocean Networks Canada (ONC) benchmarks, providing a diverse set of real-world underwater scenarios. AquaSignal employs a U-Net architecture for denoising, a ResNet18 convolutional neural network for classifying known acoustic events, and an AutoEncoder-based model for unsupervised detection of novel or anomalous signals. To our knowledge, this is the first comprehensive study to apply and evaluate this combination of techniques on maritime vessel acoustic data. Experimental results show that AquaSignal improves signal clarity and task performance, achieving 71% classification accuracy and 91% accuracy in novelty detection. Despite slightly lower classification performance compared to some state-of-the-art models, differences in data partitioning strategies limit direct comparisons. Overall, AquaSignal demonstrates strong potential for real-time underwater acoustic monitoring in scientific, environmental, and maritime domains.
标题: MatchDance:Mamba-Transformer协作架构,匹配高质量3D舞蹈合成
链接:https://arxiv.org/abs/2505.14222
摘要:从音乐到舞蹈的生成是一项具有挑战性但又是关键的任务,它涉及编舞、虚拟现实和创意内容生成的交叉点。尽管它的意义,现有的方法面临着很大的限制,在实现编排的一致性。为了应对这一挑战,我们提出了MatchDance,一个新的框架,音乐舞蹈生成,构建了一个潜在的代表性,以提高编舞的一致性。MatchDance采用两阶段设计:(1)基于运动学-动态的量化阶段(KDQS),其通过具有运动学-动态约束的有限标量量化(FSQ)将舞蹈动作编码成潜在表示,并且以高保真度重构它们,以及(2)混合音乐到舞蹈生成阶段(HMDGS),其使用Mamba-Transformer混合架构来将音乐映射到潜在表示,然后通过KDQS解码器生成3D舞蹈动作。此外,音乐舞蹈检索框架和全面的指标进行评估。在FineDance数据集上进行的大量实验展示了最先进的性能。代码将在接受后发布。
摘要:Music-to-dance generation represents a challenging yet pivotal task at the intersection of choreography, virtual reality, and creative content generation. Despite its significance, existing methods face substantial limitation in achieving choreographic consistency. To address the challenge, we propose MatchDance, a novel framework for music-to-dance generation that constructs a latent representation to enhance choreographic consistency. MatchDance employs a two-stage design: (1) a Kinematic-Dynamic-based Quantization Stage (KDQS), which encodes dance motions into a latent representation by Finite Scalar Quantization (FSQ) with kinematic-dynamic constraints and reconstructs them with high fidelity, and (2) a Hybrid Music-to-Dance Generation Stage(HMDGS), which uses a Mamba-Transformer hybrid architecture to map music into the latent representation, followed by the KDQS decoder to generate 3D dance motions. Additionally, a music-dance retrieval framework and comprehensive metrics are introduced for evaluation. Extensive experiments on the FineDance dataset demonstrate state-of-the-art performance. Code will be released upon acceptance.
标题: Speech Deepfakes的源验证
链接:https://arxiv.org/abs/2505.14188
备注:Accepted at INTERSPEECH 2025
摘要:随着语音deepfake生成器的普及,不仅要评估合成音频的真实性,还要追踪其来源。虽然源归因模型试图解决这一挑战,但它们往往在开放集条件下与看不见的生成器作斗争。在本文中,我们介绍了源验证任务,它的灵感来自说话人验证,确定是否使用相同的模型作为一组参考信号的测试轨道。我们的方法利用了一个经过源属性训练的分类器的嵌入,计算轨道之间的距离分数,以评估它们是否来自同一个源。我们在不同的场景中评估多个模型,分析说话者多样性,语言不匹配和后处理操作的影响。这项工作首次探索了来源验证,突出了其潜力和漏洞,并为现实世界的法医应用提供了见解。
摘要:With the proliferation of speech deepfake generators, it becomes crucial not only to assess the authenticity of synthetic audio but also to trace its origin. While source attribution models attempt to address this challenge, they often struggle in open-set conditions against unseen generators. In this paper, we introduce the source verification task, which, inspired by speaker verification, determines whether a test track was produced using the same model as a set of reference signals. Our approach leverages embeddings from a classifier trained for source attribution, computing distance scores between tracks to assess whether they originate from the same source. We evaluate multiple models across diverse scenarios, analyzing the impact of speaker diversity, language mismatch, and post-processing operations. This work provides the first exploration of source verification, highlighting its potential and vulnerabilities, and offers insights for real-world forensic applications.
标题: AudSemThinker:通过声音语义推理增强音频语言模型
链接:https://arxiv.org/abs/2505.14142
摘要:音频语言模型已经在各种声音理解任务中显示出有希望的结果,但它们在细粒度声音语义上的推理能力仍然有限。在本文中,我们提出了AudSemThinker,一个模型,其推理是围绕一个框架的听觉语义启发人类认知。为了支持这一点,我们引入了AudSem,这是一个专门为音频语言模型中的语义描述符推理而策划的新数据集。AudSem通过提供经过仔细过滤的音频样本集合以及通过强大的多级管道生成的字幕,解决了zero-shot评估中数据污染的持续挑战。我们的实验表明,AudSemThinker在多个训练设置中的表现优于最先进的模型,突出了其在语义音频推理方面的优势。AudSemThinker和AudSem数据集都是公开发布的。
摘要:Audio-language models have shown promising results in various sound understanding tasks, yet they remain limited in their ability to reason over the fine-grained semantics of sound. In this paper, we present AudSemThinker, a model whose reasoning is structured around a framework of auditory semantics inspired by human cognition. To support this, we introduce AudSem, a novel dataset specifically curated for semantic descriptor reasoning in audio-language models. AudSem addresses the persistent challenge of data contamination in zero-shot evaluations by providing a carefully filtered collection of audio samples paired with captions generated through a robust multi-stage pipeline. Our experiments demonstrate that AudSemThinker outperforms state-of-the-art models across multiple training settings, highlighting its strength in semantic audio reasoning. Both AudSemThinker and the AudSem dataset are released publicly.
标题: AudioJailbreak:针对端到端大型音频语言模型的越狱攻击
链接:https://arxiv.org/abs/2505.14103
摘要:最近研究了针对大型音频语言模型(LALM)的越狱攻击,但它们实现了次优的有效性,适用性和实用性,特别是假设对手可以完全操纵用户提示。在这项工作中,我们首先进行了广泛的实验表明,先进的文本越狱攻击不能很容易地通过文本到语音(TTS)技术移植到端到端的LALM。然后,我们提出了AudioJailbreak,一种新颖的音频越狱攻击,具有(1)隐蔽性:通过制作后缀越狱音频,越狱音频不需要在时间轴上与用户提示对齐;(2)通用性:通过将多个提示合并到扰动生成中,单个越狱扰动对不同的提示有效;(3)隐蔽性:越狱音频的恶意意图不会通过提出各种意图隐藏策略来提高受害者的意识;以及(4)空中鲁棒性:越狱音频通过将混响失真效果与房间脉冲响应结合到扰动的产生。相比之下,所有先前的音频越狱攻击都不能提供隐蔽性、普遍性、隐蔽性或空中鲁棒性。此外,AudioJailbreak还适用于无法完全操纵用户提示的对手,因此具有更广泛的攻击场景。迄今为止,大多数LALM的广泛实验证明了AudioJailbreak的高效率。我们强调,我们的工作窥视到对LALM的音频越狱攻击的安全影响,并切实促进提高其安全鲁棒性。实现和音频示例可在我们的网站https://audiojailbreak.github.io/AudioJailbreak上获得。
摘要:Jailbreak attacks to Large audio-language models (LALMs) are studied recently, but they achieve suboptimal effectiveness, applicability, and practicability, particularly, assuming that the adversary can fully manipulate user prompts. In this work, we first conduct an extensive experiment showing that advanced text jailbreak attacks cannot be easily ported to end-to-end LALMs via text-to speech (TTS) techniques. We then propose AudioJailbreak, a novel audio jailbreak attack, featuring (1) asynchrony: the jailbreak audio does not need to align with user prompts in the time axis by crafting suffixal jailbreak audios; (2) universality: a single jailbreak perturbation is effective for different prompts by incorporating multiple prompts into perturbation generation; (3) stealthiness: the malicious intent of jailbreak audios will not raise the awareness of victims by proposing various intent concealment strategies; and (4) over-the-air robustness: the jailbreak audios remain effective when being played over the air by incorporating the reverberation distortion effect with room impulse response into the generation of the perturbations. In contrast, all prior audio jailbreak attacks cannot offer asynchrony, universality, stealthiness, or over-the-air robustness. Moreover, AudioJailbreak is also applicable to the adversary who cannot fully manipulate user prompts, thus has a much broader attack scenario. Extensive experiments with thus far the most LALMs demonstrate the high effectiveness of AudioJailbreak. We highlight that our work peeks into the security implications of audio jailbreak attacks against LALMs, and realistically fosters improving their security robustness. The implementation and audio samples are available at our website https://audiojailbreak.github.io/AudioJailbreak.
标题: 用语言和语音模型嵌入重建语音产生期间的神经活动
链接:https://arxiv.org/abs/2505.14074
备注:Accepted for presentation at Interspeech2025
摘要:了解神经活动如何编码语音和语言产生是神经科学和人工智能的一个基本挑战。这项研究调查了来自大规模自监督语言和语音模型的嵌入是否可以有效地重建语音产生过程中捕获的神经活动记录。我们利用基于语言和声学数据训练的深度学习模型的预训练嵌入来表示高级语音特征,并将其映射到神经信号上。我们分析了这些嵌入在多大程度上保留了大脑活动的时空动态。我们使用相关性度量和信号重建质量评估来评估重建的神经信号与地面真实记录的关系。结果表明,神经活动可以有效地重建使用嵌入从大型语言和语音模型在所有研究参与者,产生皮尔逊相关系数范围从0.79到0.99。
摘要:Understanding how neural activity encodes speech and language production is a fundamental challenge in neuroscience and artificial intelligence. This study investigates whether embeddings from large-scale, self-supervised language and speech models can effectively reconstruct neural activity recordings captured during speech production. We leverage pre-trained embeddings from deep learning models trained on linguistic and acoustic data to represent high-level speech features and map them onto neural signals. We analyze the extent to which these embeddings preserve the spatio-temporal dynamics of brain activity. We evaluate reconstructed neural signals against ground truth recordings using correlation metrics and signal reconstruction quality assessments. The results indicate that neural activity can be effectively reconstructed using embeddings from large language and speech models across all study participants, yielding Pearson correlation coefficients ranging from 0.79 to 0.99.
标题: 将确定性增强条件与双流编码相结合用于基于扩散的语音增强
链接:https://arxiv.org/abs/2505.13983
摘要:基于扩散的语音增强(SE)模型需要将正确的先验知识作为可靠的条件来生成准确的预测。然而,使用噪声特征提供可靠的条件是具有挑战性的。一种解决方案是使用由确定性方法增强的特征作为条件。然而,确定性方法所造成的信息失真和损失可能会影响扩散过程。在本文中,我们首先研究使用不同的确定性SE模型作为扩散条件的影响。我们验证两个条件,这取决于噪声特征是否被用作条件的一部分:一个仅使用确定性特征(仅确定性),另一个同时使用确定性和噪声特征(确定性噪声)。初步调查发现,使用确定性增强条件可以改善真实数据的听力体验,而使用仅确定性条件或确定性噪声条件之间的选择取决于确定性模型。基于这些发现,我们提出了一个双码流编码修复扩散模型SE(DERDM-SE),更有效地利用这两个条件。此外,我们发现细粒度的确定性模型在客观评价指标方面具有更大的潜力,而基于UNet的确定性模型提供了更稳定的扩散性能。因此,在DERDM-SE中,我们提出了一个确定性模型,它结合了粗粒度和细粒度的处理。CHiME 4上的实验结果表明,所提出的模型有效地利用确定性模型,以实现更好的SE评估分数,以及更稳定的性能相比,其他基于扩散的SE模型。
摘要:Diffusion-based speech enhancement (SE) models need to incorporate correct prior knowledge as reliable conditions to generate accurate predictions. However, providing reliable conditions using noisy features is challenging. One solution is to use features enhanced by deterministic methods as conditions. However, the information distortion and loss caused by deterministic methods might affect the diffusion process. In this paper, we first investigate the effects of using different deterministic SE models as conditions for diffusion. We validate two conditions depending on whether the noisy feature was used as part of the condition: one using only the deterministic feature (deterministic-only), and the other using both deterministic and noisy features (deterministic-noisy). Preliminary investigation found that using deterministic enhanced conditions improves hearing experiences on real data, while the choice between using deterministic-only or deterministic-noisy conditions depends on the deterministic models. Based on these findings, we propose a dual-streaming encoding Repair-Diffusion Model for SE (DERDM-SE) to more effectively utilize both conditions. Moreover, we found that fine-grained deterministic models have greater potential in objective evaluation metrics, while UNet-based deterministic models provide more stable diffusion performance. Therefore, in the DERDM-SE, we propose a deterministic model that combines coarse- and fine-grained processing. Experimental results on CHiME4 show that the proposed models effectively leverage deterministic models to achieve better SE evaluation scores, along with more stable performance compared to other diffusion-based SE models.
标题: 语音情感识别与个性的桥梁:数据集和时间交互条件网络
链接:https://arxiv.org/abs/2505.13978
摘要:本研究探讨人格特质与情绪表达之间的交互作用,探讨人格信息如何提高言语情绪识别能力。我们收集了IEMOCAP数据集的人格注释,统计分析发现人格特质和情绪表达之间存在显著相关性。为了提取细粒度的个性特征,我们提出了一个时间交互条件网络(TICN),其中个性特征与基于Hubert的声学特征相结合的SER。实验表明,将地面真实的个性特征显着提高效价识别,提高一致性相关系数(CCC)从0.698到0.785相比,没有个性信息的基线。针对对话系统中无法获取用户个性信息的实际应用,我们开发了一个自动个性识别的前端模块。使用这些自动预测的性状作为我们提出的TICN模型的输入,我们实现了CCC为0.776的效价识别,比基线相对提高了11.17%。这些发现证实了个性感知SER的有效性,并为进一步探索个性感知语音处理应用提供了坚实的基础。
摘要:This study investigates the interaction between personality traits and emotional expression, exploring how personality information can improve speech emotion recognition (SER). We collected personality annotation for the IEMOCAP dataset, and the statistical analysis identified significant correlations between personality traits and emotional expressions. To extract finegrained personality features, we propose a temporal interaction condition network (TICN), in which personality features are integrated with Hubert-based acoustic features for SER. Experiments show that incorporating ground-truth personality traits significantly enhances valence recognition, improving the concordance correlation coefficient (CCC) from 0.698 to 0.785 compared to the baseline without personality information. For practical applications in dialogue systems where personality information about the user is unavailable, we develop a front-end module of automatic personality recognition. Using these automatically predicted traits as inputs to our proposed TICN model, we achieve a CCC of 0.776 for valence recognition, representing an 11.17% relative improvement over the baseline. These findings confirm the effectiveness of personality-aware SER and provide a solid foundation for further exploration in personality-aware speech processing applications.
标题: 基于多模式信息的语音处理(MISP)2025挑战:视听数字化和识别
链接:https://arxiv.org/abs/2505.13971
备注:Accepted by Interspeech 2025. Camera-ready version
摘要:由于复杂的声学条件,会议是语音应用的一个有价值但具有挑战性的场景。本文总结了在Interspeech 2025上举办的MISP 2025挑战赛的成果,该挑战赛的重点是通过将视频模态与音频相结合来实现多模态、多设备会议转录。这些任务包括视听扬声器日记(AVSD),视听语音识别(AVSR)和视听日记和识别(AVDR)。我们提出了挑战的目标,任务,数据集,基线系统和参与者提出的解决方案。表现最好的系统在基线上取得了显著的改善:最好的AVSD模型实现了8.09%的日记错误率(DER),改善了7.43%;最好的AVSR系统实现了9.48%的字符错误率(CER),改善了10.62%;最佳AVDR系统的级联最小排列字符错误率(cpCER)为11.56%,提高了72.49%。
摘要:Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on multi-modal, multi-device meeting transcription by incorporating video modality alongside audio. The tasks include Audio-Visual Speaker Diarization (AVSD), Audio-Visual Speech Recognition (AVSR), and Audio-Visual Diarization and Recognition (AVDR). We present the challenge's objectives, tasks, dataset, baseline systems, and solutions proposed by participants. The best-performing systems achieved significant improvements over the baseline: the top AVSD model achieved a Diarization Error Rate (DER) of 8.09%, improving by 7.43%; the top AVSR system achieved a Character Error Rate (CER) of 9.48%, improving by 10.62%; and the best AVDR system achieved a concatenated minimum-permutation Character Error Rate (cpCER) of 11.56%, improving by 72.49%.
标题: BiCrossMamba-ST:采用双向Mamba光谱-时间交叉注意力的语音深度伪造检测
链接:https://arxiv.org/abs/2505.13930
备注:Accepted Interspeech 2025
摘要:我们提出了BiCrossMamba-ST,这是一个强大的语音deepfake检测框架,它利用了由双向Mamba块和相互交叉注意力驱动的双分支频谱-时间架构。通过分别处理频谱子带和时间间隔,然后整合它们的表示,BiCrossMamba-ST有效地捕捉合成语音的微妙线索。此外,我们提出的框架利用基于卷积的2D注意力地图来关注特定的频谱-时间区域,从而实现强大的深度伪造检测。BiCrossMamba-ST直接在原始特征上操作,实现了显着的性能改进,在ASVSpoof LA 21和ASVSpoof DF 21基准测试中分别比最先进的AASIST相对增益67.74%和26.3%,在ASVSpoof DF 21上比RawBMamba提高了6.80%。代码和模型将公开提供。
摘要:We propose BiCrossMamba-ST, a robust framework for speech deepfake detection that leverages a dual-branch spectro-temporal architecture powered by bidirectional Mamba blocks and mutual cross-attention. By processing spectral sub-bands and temporal intervals separately and then integrating their representations, BiCrossMamba-ST effectively captures the subtle cues of synthetic speech. In addition, our proposed framework leverages a convolution-based 2D attention map to focus on specific spectro-temporal regions, enabling robust deepfake detection. Operating directly on raw features, BiCrossMamba-ST achieves significant performance improvements, a 67.74% and 26.3% relative gain over state-of-the-art AASIST on ASVSpoof LA21 and ASVSpoof DF21 benchmarks, respectively, and a 6.80% improvement over RawBMamba on ASVSpoof DF21. Code and models will be made publicly available.
标题: 使用分段语音特征进行取证深度伪造音频检测
链接:https://arxiv.org/abs/2505.13847
摘要:这项研究探讨了使用分段语音的声学特征来检测deepfake音频的潜力。这些特征具有高度的可解释性,因为它们与人类的发音过程密切相关,预计Deepfake模型将更难以复制。实验结果表明,在语音比对中常用的某些分段特征在识别深度假声方面是有效的,而一些全局特征的识别价值不大。这些发现强调了在法医语音比较中以不同方式进行音频deepfake检测的必要性,并为利用分段特征提供了一个新的视角。
摘要:This study explores the potential of using acoustic features of segmental speech sounds to detect deepfake audio. These features are highly interpretable because of their close relationship with human articulatory processes and are expected to be more difficult for deepfake models to replicate. The results demonstrate that certain segmental features commonly used in forensic voice comparison are effective in identifying deep-fakes, whereas some global features provide little value. These findings underscore the need to approach audio deepfake detection differently for forensic voice comparison and offer a new perspective on leveraging segmental features for this purpose.
标题: ClapFM-EVC:高保真和灵活的情感语音转换,具有自然语言和语音的双重控制
链接:https://arxiv.org/abs/2505.13805
备注:Accepted by InterSpeech 2025
摘要:尽管取得了巨大的进步,但实现具有灵活和可解释控制的高保真情感语音转换(EVC)仍然具有挑战性。本文介绍了ClapFM-EVC,一种新的EVC框架,能够生成高质量的转换语音驱动的自然语言提示或参考语音与可调的情感强度。我们首先提出了EVC-CLAP,这是一种情感对比语言音频预训练模型,由自然语言提示和分类标签指导,以提取和对齐语音和文本模态中的细粒度情感元素。然后,提出了一种具有自适应强度门的FuEncoder,用于将情感特征与来自预训练ASR模型的语音后验图无缝融合。为了进一步提高情感表达能力和语音自然度,我们提出了一种基于这些特征的流匹配模型来重建源语音的Mel谱图。主观和客观评价验证了ClapFM-EVC的有效性。
摘要:Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a novel EVC framework capable of generating high-quality converted speech driven by natural language prompts or reference speech with adjustable emotion intensity. We first propose EVC-CLAP, an emotional contrastive language-audio pre-training model, guided by natural language prompts and categorical labels, to extract and align fine-grained emotional elements across speech and text modalities. Then, a FuEncoder with an adaptive intensity gate is presented to seamless fuse emotional features with Phonetic PosteriorGrams from a pre-trained ASR model. To further improve emotion expressiveness and speech naturalness, we propose a flow matching model conditioned on these captured features to reconstruct Mel-spectrogram of source speech. Subjective and objective evaluations validate the effectiveness of ClapFM-EVC.
标题: Sat2Sound:一个用于零拍声景映射的统一框架
链接:https://arxiv.org/abs/2505.13777
摘要:我们提出了Sat2Sound,一个用于声景映射的多模态表示学习框架,旨在预测地球上任何位置的声音分布。用于此任务的现有方法依赖于卫星图像和成对的地理标记的音频样本,其通常无法捕获给定位置处的声源的多样性。为了解决这一限制,我们利用视觉语言模型(VLM)来增强现有的数据集,以生成语义丰富的音景描述的卫星图像中描绘的位置。我们的方法结合了音频,音频字幕,卫星图像和卫星图像字幕的对比学习。我们假设,有一个固定的一套跨模态共享的声景概念。为此,我们学习了一个共享的音景概念码本,并将每个样本表示为这些概念的加权平均值。Sat2Sound在GeoSound和SoundingEarth两个数据集上实现了卫星图像和音频之间的跨模态检索。此外,Sat2Sound的检索详细的声景字幕的能力的基础上,我们引入了一个新的应用程序:基于位置的声景合成,它使沉浸式的声学体验。我们的代码和模型将公开发布。
摘要:We present Sat2Sound, a multimodal representation learning framework for soundscape mapping, designed to predict the distribution of sounds at any location on Earth. Existing methods for this task rely on satellite image and paired geotagged audio samples, which often fail to capture the diversity of sound sources at a given location. To address this limitation, we enhance existing datasets by leveraging a Vision-Language Model (VLM) to generate semantically rich soundscape descriptions for locations depicted in satellite images. Our approach incorporates contrastive learning across audio, audio captions, satellite images, and satellite image captions. We hypothesize that there is a fixed set of soundscape concepts shared across modalities. To this end, we learn a shared codebook of soundscape concepts and represent each sample as a weighted average of these concepts. Sat2Sound achieves state-of-the-art performance in cross-modal retrieval between satellite image and audio on two datasets: GeoSound and SoundingEarth. Additionally, building on Sat2Sound's ability to retrieve detailed soundscape captions, we introduce a novel application: location-based soundscape synthesis, which enables immersive acoustic experiences. Our code and models will be publicly available.
标题: 基于能量的TTS模型的分数训练
链接:https://arxiv.org/abs/2505.13771
摘要:噪声对比估计(NCE)是一种用于训练具有难以处理的归一化项的基于能量的模型(EBM)的流行方法。NCE的关键思想是通过比较参考样本和噪声样本的非归一化对数似然来学习,从而避免显式计算归一化项。然而,NCE严重依赖于噪声样本的质量。最近,切片分数匹配(SSM)已被密切相关的扩散模型(DM)普及。与NCE不同,SSM通过学习其在随机选择的方向上的投影分布来学习对数似然或分数的梯度。然而,NCE和SSM都忽略了对数似然函数的形式,这是有问题的,因为EBM和DM在推理期间使用一阶优化。本文提出了一种新的标准,学习分数更适合于一阶格式。实验对比这些方法用于训练EBM。
摘要:Noise contrastive estimation (NCE) is a popular method for training energy-based models (EBM) with intractable normalisation terms. The key idea of NCE is to learn by comparing unnormalised log-likelihoods of the reference and noisy samples, thus avoiding explicitly computing normalisation terms. However, NCE critically relies on the quality of noisy samples. Recently, sliced score matching (SSM) has been popularised by closely related diffusion models (DM). Unlike NCE, SSM learns a gradient of log-likelihood, or score, by learning distribution of its projections on randomly chosen directions. However, both NCE and SSM disregard the form of log-likelihood function, which is problematic given that EBMs and DMs make use of first-order optimisation during inference. This paper proposes a new criterion that learns scores more suitable for first-order schemes. Experiments contrasts these approaches for training EBMs.
标题: VocalAgent:具有安全意识评估的声乐健康诊断大型语言模型
链接:https://arxiv.org/abs/2505.13577
摘要:None
摘要:Vocal health plays a crucial role in peoples' lives, significantly impacting their communicative abilities and interactions. However, despite the global prevalence of voice disorders, many lack access to convenient diagnosis and treatment. This paper introduces VocalAgent, an audio large language model (LLM) to address these challenges through vocal health diagnosis. We leverage Qwen-Audio-Chat fine-tuned on three datasets collected in-situ from hospital patients, and present a multifaceted evaluation framework encompassing a safety assessment to mitigate diagnostic biases, cross-lingual performance analysis, and modality ablation studies. VocalAgent demonstrates superior accuracy on voice disorder classification compared to state-of-the-art baselines. Its LLM-based method offers a scalable solution for broader adoption of health diagnostics, while underscoring the importance of ethical and technical validation.
标题: 收听、分析和适应以学习新攻击:用于音频Deepfake源跟踪的无示例类增量学习方法
链接:https://arxiv.org/abs/2505.14601
备注:Accepted by Interspeech 2025
摘要:随着Deepfake语音变得普遍且难以检测,追踪其来源至关重要。最近关于音频deepfake源跟踪(ST)的工作旨在找到合成或操纵语音的起源。然而,ST模型必须适应学习新的deepfake攻击,同时保留以前的知识。一个主要挑战是灾难性的遗忘,模型失去了识别之前学习的攻击的能力。一些持续学习方法有助于深度伪造检测,但随着类数量的增加,ST等多类任务会带来额外的挑战。为了解决这个问题,我们提出了一种分析类增量学习方法,称为AnaST。当新的攻击出现时,特征提取器保持固定,分类器在一个时期内用封闭形式的解析解更新。这种方法可以确保数据隐私,优化内存使用,并且适合在线培训。在这项工作中进行的实验表明,我们的方法优于基线。
摘要:As deepfake speech becomes common and hard to detect, it is vital to trace its source. Recent work on audio deepfake source tracing (ST) aims to find the origins of synthetic or manipulated speech. However, ST models must adapt to learn new deepfake attacks while retaining knowledge of the previous ones. A major challenge is catastrophic forgetting, where models lose the ability to recognize previously learned attacks. Some continual learning methods help with deepfake detection, but multi-class tasks such as ST introduce additional challenges as the number of classes grows. To address this, we propose an analytic class incremental learning method called AnaST. When new attacks appear, the feature extractor remains fixed, and the classifier is updated with a closed-form analytical solution in one epoch. This approach ensures data privacy, optimizes memory usage, and is suitable for online training. The experiments carried out in this work show that our method outperforms the baselines.
标题: AdaKWS:通过测试时自适应实现稳健的关键词发现
链接:https://arxiv.org/abs/2505.14600
备注:Accepted by Interspeech 2025
摘要:口语关键词识别(KWS)旨在识别音频中的关键词,以实现广泛的应用,特别是在边缘设备上。目前的小型KWS系统专注于高效的模型设计。然而,在看不见的环境或嘈杂的背景下,它们的推理性能可能会下降。测试时自适应(TTA)帮助模型适应测试样本,而无需原始训练数据。在这项研究中,我们提出了AdaKWS,第一TTA方法强大的KWS,以我们所知。具体来说,1)我们首先通过基于预测熵最小化选择可靠样本并调整每个批次中的归一化统计量来优化模型的置信度。2)我们引入了伪关键字一致性(PKC)来识别关键的、可靠的特征,而不会过度拟合噪声。我们的实验表明,AdaKWS优于其他方法在各种条件下,包括高斯噪声和真实场景噪声。代码将在适当的时候发布。
摘要:Spoken keyword spotting (KWS) aims to identify keywords in audio for wide applications, especially on edge devices. Current small-footprint KWS systems focus on efficient model designs. However, their inference performance can decline in unseen environments or noisy backgrounds. Test-time adaptation (TTA) helps models adapt to test samples without needing the original training data. In this study, we present AdaKWS, the first TTA method for robust KWS to the best of our knowledge. Specifically, 1) We initially optimize the model's confidence by selecting reliable samples based on prediction entropy minimization and adjusting the normalization statistics in each batch. 2) We introduce pseudo-keyword consistency (PKC) to identify critical, reliable features without overfitting to noise. Our experiments show that AdaKWS outperforms other methods across various conditions, including Gaussian noise and real-scenario noises. The code will be released in due course.
标题: SSPS:用于鲁棒自监督说话人确认的自监督正采样
链接:https://arxiv.org/abs/2505.14561
备注:accepted at Interspeech 2025
摘要:自监督学习(SSL)在说话人确认(SV)方面取得了长足的进步。标准框架使用相同的话语积极采样和数据增强来生成同一说话人的正锚对。这是一个主要的限制,因为这种策略主要编码来自记录条件的信道信息,由锚点和正共享。我们提出了一种新的正采样技术来解决这个瓶颈:自监督正采样(SSPS)。对于给定的锚点,SSPS旨在找到适当的积极因素,即,相同的发言人身份,但不同的录音条件,在潜在的空间中使用聚类分配和正嵌入的记忆队列。SSPS提高了Simplified和DINO的SV性能,达到2.57%和2.53%的EER,优于VoxCeleb 1-O上的SOTA SSL方法。特别是,SimCLR-SSPS通过降低扬声器内方差实现了58%的EER降低,提供了与DINO-SSPS相当的性能。
摘要:Self-Supervised Learning (SSL) has led to considerable progress in Speaker Verification (SV). The standard framework uses same-utterance positive sampling and data-augmentation to generate anchor-positive pairs of the same speaker. This is a major limitation, as this strategy primarily encodes channel information from the recording condition, shared by the anchor and positive. We propose a new positive sampling technique to address this bottleneck: Self-Supervised Positive Sampling (SSPS). For a given anchor, SSPS aims to find an appropriate positive, i.e., of the same speaker identity but a different recording condition, in the latent space using clustering assignments and a memory queue of positive embeddings. SSPS improves SV performance for both SimCLR and DINO, reaching 2.57% and 2.53% EER, outperforming SOTA SSL methods on VoxCeleb1-O. In particular, SimCLR-SSPS achieves a 58% EER reduction by lowering intra-speaker variance, providing comparable performance to DINO-SSPS.
标题: 教授有声大型语言模型听不到的内容:通过合成阴性样本减轻幻觉
链接:https://arxiv.org/abs/2505.14518
备注:Accepted to Interspeech 2025
摘要:音频感知大语言模型(ALLM)的最新进展使它们能够处理和理解音频输入。然而,这些模型经常会产生不存在的声音事件,降低了它们在现实世界应用中的可靠性。为了解决这个问题,我们提出了LISTEN(Learning to Identify Sounds Through Extended Negative Samples),这是一种类似对比的训练方法,可以增强ALLM使用骨干LLM的合成数据区分存在和不存在声音的能力。与以前的方法不同,我们的方法不需要修改LLM参数,并通过一个轻量级的适配器有效地集成了音频表示。实验表明,LISTEN有效地减轻了幻觉,同时保持了现有音频问题和推理基准的令人印象深刻的性能。同时,它在数据和计算方面都更有效。
摘要:Recent advancements in audio-aware large language models (ALLMs) enable them to process and understand audio inputs. However, these models often hallucinate non-existent sound events, reducing their reliability in real-world applications. To address this, we propose LISTEN (Learning to Identify Sounds Through Extended Negative Samples), a contrastive-like training method that enhances ALLMs' ability to distinguish between present and absent sounds using synthesized data from the backbone LLM. Unlike prior approaches, our method requires no modification to LLM parameters and efficiently integrates audio representations via a lightweight adapter. Experiments show that LISTEN effectively mitigates hallucinations while maintaining impressive performance on existing audio question and reasoning benchmarks. At the same time, it is more efficient in both data and computation.
标题: 引导深度非线性空间选择性过滤器以弱引导提取动态场景中的移动扬声器
链接:https://arxiv.org/abs/2505.14517
备注:Accepted at Interspeech 2025
摘要:最近的说话人提取方法,使用深度非线性空间滤波执行非常好的目标方向是已知的和固定的。然而,由于时变空间特征和产生的模糊性,例如当移动扬声器交叉时,空间动态场景更具挑战性。虽然在静态场景中,用户可以容易地指向目标的方向,但手动跟踪移动的说话者是不切实际的。而不是依赖于准确的时间相关的方向线索,我们称之为强指导,在本文中,我们提出了一个弱指导提取方法,仅依赖于目标的初始位置,以应付空间动态场景。通过结合我们自己的深度跟踪算法并在合成数据集上开发联合训练策略,我们证明了我们的方法在解决空间模糊性方面的能力,甚至优于不匹配但强烈引导的提取方法。
摘要:Recent speaker extraction methods using deep non-linear spatial filtering perform exceptionally well when the target direction is known and stationary. However, spatially dynamic scenarios are considerably more challenging due to time-varying spatial features and arising ambiguities, e.g. when moving speakers cross. While in a static scenario it may be easy for a user to point to the target's direction, manually tracking a moving speaker is impractical. Instead of relying on accurate time-dependent directional cues, which we refer to as strong guidance, in this paper we propose a weakly guided extraction method solely depending on the target's initial position to cope with spatial dynamic scenarios. By incorporating our own deep tracking algorithm and developing a joint training strategy on a synthetic dataset, we demonstrate the proficiency of our approach in resolving spatial ambiguities and even outperform a mismatched, but strongly guided extraction method.
标题: 利用距离和房间线索的单通道目标语音提取
链接:https://arxiv.org/abs/2505.14433
备注:5 pages, 3 figures, accepted by Eusipco 2025
摘要:本文旨在利用距离线索和房间信息实现封闭环境下的单通道目标语音提取。最近的工作已经验证了TSE任务的距离线索的可行性,其可以暗示声源的直达混响比(DRR),因此可以用于语音分离和TSE系统。然而,这样的距离线索受到房间的声学特性(诸如尺寸和混响时间)的显著影响,使得仅依赖于距离线索的TSE系统具有挑战性以概括各种不同的房间。为了解决这个问题,我们建议为基于距离的TSE提供房间环境信息(房间尺寸和混响时间),以获得更好的泛化能力。特别是,我们提出了一个基于距离和环境的时频(TF)域的TSE模型与学习距离和房间嵌入。模拟和真实数据集上的结果证明了该方法的可行性。演示材料可在https://runwushi.github.io/distance-room-demo-page/上获得。
摘要:This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of distance clues for the TSE task, which can imply the sound source's direct-to-reverberation ratio (DRR) and thus can be utilized for speech separation and TSE systems. However, such distance clue is significantly influenced by the room's acoustic characteristics, such as dimension and reverberation time, making it challenging for TSE systems that rely solely on distance clues to generalize across a variety of different rooms. To solve this, we suggest providing room environmental information (room dimensions and reverberation time) for distance-based TSE for better generalization capabilities. Especially, we propose a distance and environment-based TSE model in the time-frequency (TF) domain with learnable distance and room embedding. Results on both simulated and real collected datasets demonstrate its feasibility. Demonstration materials are available at https://runwushi.github.io/distance-room-demo-page/.
标题: 语音合成中口音相似度的成对评估
链接:https://arxiv.org/abs/2505.14410
备注:Accepted by INTERSPEECH 2025
摘要:尽管人们对生成高保真口音的兴趣越来越大,但在语音合成中评估口音相似性的研究还不够深入。我们的目标是加强口音相似度的主观和客观评估方法。主观上,我们通过添加组件来改进XAB听力测试,这些组件以更少的听众和更低的成本实现更高的统计显著性。我们的方法包括为听众提供transmittance,让他们突出感知的口音差异,并实施细致的可靠性筛选。客观地说,我们利用发音相关的指标,元音共振峰和语音posteriorgrams之间的距离的基础上,评估口音的产生。比较实验表明,这些指标,口音相似性,扬声器相似性,和梅尔倒谱失真,可以使用。此外,我们的研究结果强调了单词错误率等常用指标在评估代表性不足的口音方面的重大局限性。
摘要:Despite growing interest in generating high-fidelity accents, evaluating accent similarity in speech synthesis has been underexplored. We aim to enhance both subjective and objective evaluation methods for accent similarity. Subjectively, we refine the XAB listening test by adding components that achieve higher statistical significance with fewer listeners and lower costs. Our method involves providing listeners with transcriptions, having them highlight perceived accent differences, and implementing meticulous screening for reliability. Objectively, we utilise pronunciation-related metrics, based on distances between vowel formants and phonetic posteriorgrams, to evaluate accent generation. Comparative experiments reveal that these metrics, alongside accent similarity, speaker similarity, and Mel Cepstral Distortion, can be used. Moreover, our findings underscore significant limitations of common metrics like Word Error Rate in assessing underrepresented accents.
标题: 扩展和增强基于LLM的AVSR:投影仪的稀疏混合方法
链接:https://arxiv.org/abs/2505.14336
摘要:视听语音识别(AVSR)通过整合视觉线索来增强噪声环境中的鲁棒性。虽然最近的进展将大型语言模型(LLM)集成到AVSR中,但其高计算成本阻碍了在资源受限环境中的部署。为了解决这个问题,我们提出了Llama-SMoP,这是一种高效的多模态LLM,它采用稀疏混合投影仪(SMoP)模块来扩展模型容量,而不会增加推理成本。通过整合稀疏选通专家混合(MoE)投影仪,Llama-SMoP能够在保持强大性能的同时使用更小的LLM。我们探讨了三种SMoP配置,并表明Llama-SMoP DEDR(分离专家,分离路由器),它使用特定于模态的路由器和专家,在ASR,VSR和AVSR任务上实现了卓越的性能。消融研究证实了其在专家激活、可扩展性和噪声鲁棒性方面的有效性。
摘要:Audio-Visual Speech Recognition (AVSR) enhances robustness in noisy environments by integrating visual cues. While recent advances integrate Large Language Models (LLMs) into AVSR, their high computational cost hinders deployment in resource-constrained settings. To address this, we propose Llama-SMoP, an efficient Multimodal LLM that employs a Sparse Mixture of Projectors (SMoP) module to scale model capacity without increasing inference costs. By incorporating sparsely-gated mixture-of-experts (MoE) projectors, Llama-SMoP enables the use of smaller LLMs while maintaining strong performance. We explore three SMoP configurations and show that Llama-SMoP DEDR (Disjoint-Experts, Disjoint-Routers), which uses modality-specific routers and experts, achieves superior performance on ASR, VSR, and AVSR tasks. Ablation studies confirm its effectiveness in expert activation, scalability, and noise robustness.
标题: 无障碍编辑:背景噪音感知Zero-Shot语音编辑,具有上下文增强
链接:https://arxiv.org/abs/2505.14066
备注:5 pages, 3 figures
摘要:随着zero-shot文语转换技术的迅速发展,产生高质量的、与真实语音难以区分的语音信号成为可能。语音编辑,包括语音插入和替换,由于其潜在的应用吸引了研究人员。然而,现有的研究只考虑了干净的语音场景。在实际应用中,环境噪声的存在会显著降低生成的质量。在这项研究中,我们提出了一个噪声弹性语音编辑框架,无噪声编辑,嘈杂的语音编辑。无噪声编辑采用了一个频段感知的噪声抑制模块和一个内容细化策略。它可以很好地解决语音和背景噪声的频带不分离的情况。所提出的无障碍编辑框架在多个定量和定性评价中优于最先进的方法。
摘要:With the fast development of zero-shot text-to-speech technologies, it is possible to generate high-quality speech signals that are indistinguishable from the real ones. Speech editing, including speech insertion and replacement, appeals to researchers due to its potential applications. However, existing studies only considered clean speech scenarios. In real-world applications, the existence of environmental noise could significantly degrade the quality of the generation. In this study, we propose a noise-resilient speech editing framework, SeamlessEdit, for noisy speech editing. SeamlessEdit adopts a frequency-band-aware noise suppression module and an in-content refinement strategy. It can well address the scenario where the frequency bands of voice and background noise are not separated. The proposed SeamlessEdit framework outperforms state-of-the-art approaches in multiple quantitative and qualitative evaluations.
标题: 具有动态温度的自然感知课程学习用于语音深度伪造检测
链接:https://arxiv.org/abs/2505.13976
备注:Accepted by Interspeech 2025
摘要:语音深度伪造检测(SDD)的最新进展显着改善了欺骗语音中基于伪影的检测。然而,大多数模型忽略了语音自然度,这是区分真实语音和欺骗语音的关键线索。本研究提出了自然感知课程学习,一种新的训练框架,利用语音自然度来增强SDD的鲁棒性和泛化能力。这种方法使用地面实况标签和平均意见得分来测量样本难度,并调整训练时间表以逐步引入更具挑战性的样本。为了进一步提高泛化能力,在训练过程中引入了基于语音自然度的动态温度缩放方法。在ASVspoof 2021 DF数据集上的实验中,EER相对降低了23%,而无需修改模型架构。消融研究证实了SDD任务的自然意识训练策略的有效性。
摘要:Recent advances in speech deepfake detection (SDD) have significantly improved artifacts-based detection in spoofed speech. However, most models overlook speech naturalness, a crucial cue for distinguishing bona fide speech from spoofed speech. This study proposes naturalness-aware curriculum learning, a novel training framework that leverages speech naturalness to enhance the robustness and generalization of SDD. This approach measures sample difficulty using both ground-truth labels and mean opinion scores, and adjusts the training schedule to progressively introduce more challenging samples. To further improve generalization, a dynamic temperature scaling method based on speech naturalness is incorporated into the training process. A 23% relative reduction in the EER was achieved in the experiments on the ASVspoof 2021 DF dataset, without modifying the model architecture. Ablation studies confirmed the effectiveness of naturalness-aware training strategies for SDD tasks.
标题: U-Sam:统一语音、音频和音乐理解的音频语言模型
链接:https://arxiv.org/abs/2505.13880
备注:Accepted to Interspeech 2025
摘要:音频任务的文本生成范式为统一的音频理解开辟了新的可能性。然而,现有的模型在实现对不同音频类型(如语音、一般音频事件和音乐)的全面理解方面面临着重大挑战。此外,它们完全依赖交叉熵损失进行对齐往往是不够的,因为它平等地对待所有标记,并且无法考虑冗余的音频特征,导致较弱的交叉模态对齐。为了应对上述挑战,本文介绍了U-SAM,这是一种先进的音频语言模型,它将语音,音频和音乐的专用编码器与预训练的大型语言模型(LLM)集成在一起。U-SAM采用混合专家(MoE)投影仪进行任务感知特征融合,动态路由和集成特定于域的编码器输出。此外,U-SAM还集成了一个语义感知对比损失模块,该模块在语言监督下明确识别冗余音频特征,并纠正其语义和频谱表示,以增强跨模态对齐。大量的实验表明,U-SAM在多个基准测试中始终优于专业模型和现有的音频语言模型。此外,它在看不见的任务上表现出紧急能力,展示了它的泛化潜力。代码可用(https://github.com/Honee-W/U-SAM/)。
摘要:The text generation paradigm for audio tasks has opened new possibilities for unified audio understanding. However, existing models face significant challenges in achieving a comprehensive understanding across diverse audio types, such as speech, general audio events, and music. Furthermore, their exclusive reliance on cross-entropy loss for alignment often falls short, as it treats all tokens equally and fails to account for redundant audio features, leading to weaker cross-modal alignment. To deal with the above challenges, this paper introduces U-SAM, an advanced audio language model that integrates specialized encoders for speech, audio, and music with a pre-trained large language model (LLM). U-SAM employs a Mixture of Experts (MoE) projector for task-aware feature fusion, dynamically routing and integrating the domain-specific encoder outputs. Additionally, U-SAM incorporates a Semantic-Aware Contrastive Loss Module, which explicitly identifies redundant audio features under language supervision and rectifies their semantic and spectral representations to enhance cross-modal alignment. Extensive experiments demonstrate that U-SAM consistently outperforms both specialized models and existing audio language models across multiple benchmarks. Moreover, it exhibits emergent capabilities on unseen tasks, showcasing its generalization potential. Code is available (https://github.com/Honee-W/U-SAM/).
标题: 一种基于语义信息的分层语音增强方法
链接:https://arxiv.org/abs/2505.13843
备注:Accepted by interspeech 2025
摘要:目前的语音增强(SE)方法大多通过直接估计时频掩模或频谱来从噪声输入中恢复干净的语音。然而,这些方法往往忽略了语音信号中固有的不同属性,如语义内容和声学细节,这可能会阻碍下游任务的性能。此外,它们的有效性往往会在复杂的声学环境中降低。为了克服这些挑战,我们提出了一种新的,语义信息为基础的,一步一步的因子分解SE方法使用因子分解编解码器和扩散模型。与传统的SE方法不同,我们的语义和声学属性的分层建模能够实现更强大的干净语音恢复,特别是在具有挑战性的声学场景中。此外,该方法为下游TTS任务提供了进一步的优势。实验结果表明,我们的算法不仅优于SOTA基线的语音质量,但也提高了TTS在噪声环境中的性能。
摘要:Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as semantic content and acoustic details, inherent in speech signals, which can hinder performance in downstream tasks. Moreover, their effectiveness tends to degrade in complex acoustic environments. To overcome these challenges, we propose a novel, semantic information-based, step-by-step factorized SE method using factorized codec and diffusion model. Unlike traditional SE methods, our hierarchical modeling of semantic and acoustic attributes enables more robust clean speech recovery, particularly in challenging acoustic scenarios. Moreover, this method offers further advantages for downstream TTS tasks. Experimental results demonstrate that our algorithm not only outperforms SOTA baselines in terms of speech quality but also enhances TTS performance in noisy environments.
标题: 基于离散声学标记去噪的LLM零拍TTS抗噪性研究
链接:https://arxiv.org/abs/2505.13830
备注:Accepted by Interspeech 2025
摘要:基于大语言模型(LLM)的zero-shot文本到语音(TTS)方法倾向于保留音频提示的声学环境,当音频提示包含噪声时,导致合成语音质量下降。在本文中,我们提出了一种新的基于神经编解码器的语音去噪器,并将其与先进的基于LLM的TTS模型,LauraTTS,实现噪声鲁棒的zero-shot TTS。建议的编解码器去噪器包括一个音频编解码器,令牌去噪器,和嵌入细化。令牌去噪器从噪声中预测出前两组干净的声学令牌,这些令牌可以作为声学提示,以使LauraTTS合成高质量的个性化语音或通过嵌入细化器和编解码器解码器转换为干净的语音波形。实验结果表明,我们提出的编解码器去噪优于国家的最先进的语音增强(SE)的方法,和建议的噪声鲁棒性LauraTTS超越使用额外的SE模型的方法。
摘要:Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech quality when the audio prompt contains noise. In this paper, we propose a novel neural codec-based speech denoiser and integrate it with the advanced LLM-based TTS model, LauraTTS, to achieve noise-robust zero-shot TTS. The proposed codec denoiser consists of an audio codec, a token denoiser, and an embedding refiner. The token denoiser predicts the first two groups of clean acoustic tokens from the noisy ones, which can serve as the acoustic prompt for LauraTTS to synthesize high-quality personalized speech or be converted to clean speech waveforms through the embedding refiner and codec decoder. Experimental results show that our proposed codec denoiser outperforms state-of-the-art speech enhancement (SE) methods, and the proposed noise-robust LauraTTS surpasses the approach using additional SE models.
标题: 用维度正规化和分数正规化推动自蒸馏原型网络的前沿
链接:https://arxiv.org/abs/2505.13826
摘要:开发没有说话人标签的鲁棒说话人验证(SV)系统一直是一个长期的挑战。早期的研究强调了自我监督和完全监督方法之间的相当大的性能差距。在本文中,我们通过引入维度正则化来增强非对比自监督框架自蒸馏原型网络(SDPN),该维度正则化通过将正则化项应用于说话人嵌入来明确解决崩溃问题。此外,我们整合了来自全监督SV的分数归一化技术,以进一步弥合监督验证性能的差距。具有维度正则化和分数归一化的SDPN在VoxCeleb 1说话人验证评估基准上设置了一个新的最先进的水平,对于试验VoxCeleb 1-{O,E,H}分别实现了1.29%,1.60%和2.80%的等错误率。这些结果表明,与当前最好的自监督方法相比,相对提高了28.3%,19.6%和22.6%,从而推进了SV技术的前沿。
摘要:Developing robust speaker verification (SV) systems without speaker labels has been a longstanding challenge. Earlier research has highlighted a considerable performance gap between self-supervised and fully supervised approaches. In this paper, we enhance the non-contrastive self-supervised framework, Self-Distillation Prototypes Network (SDPN), by introducing dimension regularization that explicitly addresses the collapse problem through the application of regularization terms to speaker embeddings. Moreover, we integrate score normalization techniques from fully supervised SV to further bridge the gap toward supervised verification performance. SDPN with dimension regularization and score normalization sets a new state-of-the-art on the VoxCeleb1 speaker verification evaluation benchmark, achieving Equal Error Rate 1.29%, 1.60%, and 2.80% for trial VoxCeleb1-{O,E,H} respectively. These results demonstrate relative improvements of 28.3%, 19.6%, and 22.6% over the current best self-supervised methods, thereby advancing the frontiers of SV technology.
标题: 言语产生过程中表面肌电信号的关节特征预测
链接:https://arxiv.org/abs/2505.13814
备注:Accepted for Interspeech2025
摘要:我们提出了一个模型,用于预测发音功能的表面肌电图(EMG)信号在语音生产。该模型集成了卷积层和Transformer块,然后是发音特征的单独预测器。对于大多数发音特征,我们的方法实现了约0.9的高预测相关性。此外,我们证明,这些预测发音功能可以解码成可理解的语音波形。据我们所知,这是第一种通过发音特征从表面EMG解码语音波形的方法,为基于EMG的语音合成提供了一种新方法。此外,我们分析了EMG电极放置和发音特征可预测性之间的关系,为优化EMG电极配置提供了知识驱动的见解。源代码和解码的语音样本是公开的。
摘要:We present a model for predicting articulatory features from surface electromyography (EMG) signals during speech production. The proposed model integrates convolutional layers and a Transformer block, followed by separate predictors for articulatory features. Our approach achieves a high prediction correlation of approximately 0.9 for most articulatory features. Furthermore, we demonstrate that these predicted articulatory features can be decoded into intelligible speech waveforms. To our knowledge, this is the first method to decode speech waveforms from surface EMG via articulatory features, offering a novel approach to EMG-based speech synthesis. Additionally, we analyze the relationship between EMG electrode placement and articulatory feature predictability, providing knowledge-driven insights for optimizing EMG electrode configurations. The source code and decoded speech samples are publicly available.
标题: 基于方向感知的神经声场在高保真度脉冲响应少频内插中的应用
链接:https://arxiv.org/abs/2505.13617
备注:Accepted at Interspeech 2025
摘要:声场的特征与声源和听众周围环境的几何和空间属性有本质联系。声音传播的物理过程被捕获在称为房间脉冲响应(RIR)的时域信号中。之前使用神经场(NF)的工作允许从有限的RIR测量中学习RIR的空间连续表示。然而,先前的基于NF的方法集中于单声道全向或至多双耳收听者,其不能精确地捕获单个点处的真实声场的方向特性。我们提出了一个方向感知神经场(DANF),更明确地采用了立体声格式RIR的方向信息。虽然DANF固有地捕获源和听众之间的空间关系,我们进一步提出了方向感知损失。此外,我们调查的能力DANF适应新的房间,以各种方式,包括低级别的适应。
摘要:The characteristics of a sound field are intrinsically linked to the geometric and spatial properties of the environment surrounding a sound source and a listener. The physics of sound propagation is captured in a time-domain signal known as a room impulse response (RIR). Prior work using neural fields (NFs) has allowed learning spatially-continuous representations of RIRs from finite RIR measurements. However, previous NF-based methods have focused on monaural omnidirectional or at most binaural listeners, which does not precisely capture the directional characteristics of a real sound field at a single point. We propose a direction-aware neural field (DANF) that more explicitly incorporates the directional information by Ambisonic-format RIRs. While DANF inherently captures spatial relations between sources and listeners, we further propose a direction-aware loss. In addition, we investigate the ability of DANF to adapt to new rooms in various ways including low-rank adaptation.
标题: 收听、分析和适应以学习新攻击:用于音频Deepfake源跟踪的无示例类增量学习方法
链接:https://arxiv.org/abs/2505.14601
备注:Accepted by Interspeech 2025
摘要:随着Deepfake语音变得普遍且难以检测,追踪其来源至关重要。最近关于音频deepfake源跟踪(ST)的工作旨在找到合成或操纵语音的起源。然而,ST模型必须适应学习新的deepfake攻击,同时保留以前的知识。一个主要挑战是灾难性的遗忘,模型失去了识别之前学习的攻击的能力。一些持续学习方法有助于深度伪造检测,但随着类数量的增加,ST等多类任务会带来额外的挑战。为了解决这个问题,我们提出了一种分析类增量学习方法,称为AnaST。当新的攻击出现时,特征提取器保持固定,分类器在一个时期内用封闭形式的解析解更新。这种方法可以确保数据隐私,优化内存使用,并且适合在线培训。在这项工作中进行的实验表明,我们的方法优于基线。
摘要:As deepfake speech becomes common and hard to detect, it is vital to trace its source. Recent work on audio deepfake source tracing (ST) aims to find the origins of synthetic or manipulated speech. However, ST models must adapt to learn new deepfake attacks while retaining knowledge of the previous ones. A major challenge is catastrophic forgetting, where models lose the ability to recognize previously learned attacks. Some continual learning methods help with deepfake detection, but multi-class tasks such as ST introduce additional challenges as the number of classes grows. To address this, we propose an analytic class incremental learning method called AnaST. When new attacks appear, the feature extractor remains fixed, and the classifier is updated with a closed-form analytical solution in one epoch. This approach ensures data privacy, optimizes memory usage, and is suitable for online training. The experiments carried out in this work show that our method outperforms the baselines.
标题: AdaKWS:通过测试时自适应实现稳健的关键词发现
链接:https://arxiv.org/abs/2505.14600
备注:Accepted by Interspeech 2025
摘要:口语关键词识别(KWS)旨在识别音频中的关键词,以实现广泛的应用,特别是在边缘设备上。目前的小型KWS系统专注于高效的模型设计。然而,在看不见的环境或嘈杂的背景下,它们的推理性能可能会下降。测试时自适应(TTA)帮助模型适应测试样本,而无需原始训练数据。在这项研究中,我们提出了AdaKWS,第一TTA方法强大的KWS,以我们所知。具体来说,1)我们首先通过基于预测熵最小化选择可靠样本并调整每个批次中的归一化统计量来优化模型的置信度。2)我们引入了伪关键字一致性(PKC)来识别关键的、可靠的特征,而不会过度拟合噪声。我们的实验表明,AdaKWS优于其他方法在各种条件下,包括高斯噪声和真实场景噪声。代码将在适当的时候发布。
摘要:Spoken keyword spotting (KWS) aims to identify keywords in audio for wide applications, especially on edge devices. Current small-footprint KWS systems focus on efficient model designs. However, their inference performance can decline in unseen environments or noisy backgrounds. Test-time adaptation (TTA) helps models adapt to test samples without needing the original training data. In this study, we present AdaKWS, the first TTA method for robust KWS to the best of our knowledge. Specifically, 1) We initially optimize the model's confidence by selecting reliable samples based on prediction entropy minimization and adjusting the normalization statistics in each batch. 2) We introduce pseudo-keyword consistency (PKC) to identify critical, reliable features without overfitting to noise. Our experiments show that AdaKWS outperforms other methods across various conditions, including Gaussian noise and real-scenario noises. The code will be released in due course.
标题: SSPS:用于鲁棒自监督说话人确认的自监督正采样
链接:https://arxiv.org/abs/2505.14561
备注:accepted at Interspeech 2025
摘要:自监督学习(SSL)在说话人确认(SV)方面取得了长足的进步。标准框架使用相同的话语积极采样和数据增强来生成同一说话人的正锚对。这是一个主要的限制,因为这种策略主要编码来自记录条件的信道信息,由锚点和正共享。我们提出了一种新的正采样技术来解决这个瓶颈:自监督正采样(SSPS)。对于给定的锚点,SSPS旨在找到适当的积极因素,即,相同的发言人身份,但不同的录音条件,在潜在的空间中使用聚类分配和正嵌入的记忆队列。SSPS提高了Simplified和DINO的SV性能,达到2.57%和2.53%的EER,优于VoxCeleb 1-O上的SOTA SSL方法。特别是,SimCLR-SSPS通过降低扬声器内方差实现了58%的EER降低,提供了与DINO-SSPS相当的性能。
摘要:Self-Supervised Learning (SSL) has led to considerable progress in Speaker Verification (SV). The standard framework uses same-utterance positive sampling and data-augmentation to generate anchor-positive pairs of the same speaker. This is a major limitation, as this strategy primarily encodes channel information from the recording condition, shared by the anchor and positive. We propose a new positive sampling technique to address this bottleneck: Self-Supervised Positive Sampling (SSPS). For a given anchor, SSPS aims to find an appropriate positive, i.e., of the same speaker identity but a different recording condition, in the latent space using clustering assignments and a memory queue of positive embeddings. SSPS improves SV performance for both SimCLR and DINO, reaching 2.57% and 2.53% EER, outperforming SOTA SSL methods on VoxCeleb1-O. In particular, SimCLR-SSPS achieves a 58% EER reduction by lowering intra-speaker variance, providing comparable performance to DINO-SSPS.
标题: 教授有声大型语言模型听不到的内容:通过合成阴性样本减轻幻觉
链接:https://arxiv.org/abs/2505.14518
备注:Accepted to Interspeech 2025
摘要:音频感知大语言模型(ALLM)的最新进展使它们能够处理和理解音频输入。然而,这些模型经常会产生不存在的声音事件,降低了它们在现实世界应用中的可靠性。为了解决这个问题,我们提出了LISTEN(Learning to Identify Sounds Through Extended Negative Samples),这是一种类似对比的训练方法,可以增强ALLM使用骨干LLM的合成数据区分存在和不存在声音的能力。与以前的方法不同,我们的方法不需要修改LLM参数,并通过一个轻量级的适配器有效地集成了音频表示。实验表明,LISTEN有效地减轻了幻觉,同时保持了现有音频问题和推理基准的令人印象深刻的性能。同时,它在数据和计算方面都更有效。
摘要:Recent advancements in audio-aware large language models (ALLMs) enable them to process and understand audio inputs. However, these models often hallucinate non-existent sound events, reducing their reliability in real-world applications. To address this, we propose LISTEN (Learning to Identify Sounds Through Extended Negative Samples), a contrastive-like training method that enhances ALLMs' ability to distinguish between present and absent sounds using synthesized data from the backbone LLM. Unlike prior approaches, our method requires no modification to LLM parameters and efficiently integrates audio representations via a lightweight adapter. Experiments show that LISTEN effectively mitigates hallucinations while maintaining impressive performance on existing audio question and reasoning benchmarks. At the same time, it is more efficient in both data and computation.
标题: 引导深度非线性空间选择性过滤器以弱引导提取动态场景中的移动扬声器
链接:https://arxiv.org/abs/2505.14517
备注:Accepted at Interspeech 2025
摘要:最近的说话人提取方法,使用深度非线性空间滤波执行非常好的目标方向是已知的和固定的。然而,由于时变空间特征和产生的模糊性,例如当移动扬声器交叉时,空间动态场景更具挑战性。虽然在静态场景中,用户可以容易地指向目标的方向,但手动跟踪移动的说话者是不切实际的。而不是依赖于准确的时间相关的方向线索,我们称之为强指导,在本文中,我们提出了一个弱指导提取方法,仅依赖于目标的初始位置,以应付空间动态场景。通过结合我们自己的深度跟踪算法并在合成数据集上开发联合训练策略,我们证明了我们的方法在解决空间模糊性方面的能力,甚至优于不匹配但强烈引导的提取方法。
摘要:Recent speaker extraction methods using deep non-linear spatial filtering perform exceptionally well when the target direction is known and stationary. However, spatially dynamic scenarios are considerably more challenging due to time-varying spatial features and arising ambiguities, e.g. when moving speakers cross. While in a static scenario it may be easy for a user to point to the target's direction, manually tracking a moving speaker is impractical. Instead of relying on accurate time-dependent directional cues, which we refer to as strong guidance, in this paper we propose a weakly guided extraction method solely depending on the target's initial position to cope with spatial dynamic scenarios. By incorporating our own deep tracking algorithm and developing a joint training strategy on a synthetic dataset, we demonstrate the proficiency of our approach in resolving spatial ambiguities and even outperform a mismatched, but strongly guided extraction method.
标题: FlowPSE:利用流量匹配的目标说话人提取
链接:https://arxiv.org/abs/2505.14465
备注:InterSpeech 2025
摘要:目标说话人提取(TSE)的目的是使用说话人登记作为参考,从混合语音中分离出特定说话人的语音。虽然大多数现有的方法是歧视性的,最近的生成方法TSE实现强大的结果。然而,TSE的生成方法仍然没有得到充分的探索,大多数现有的方法依赖于复杂的管道和预训练的组件,导致计算开销。在这项工作中,我们提出了FlowTSE,一个简单而有效的TSE方法的基础上的条件流匹配。我们的模型接收一个注册音频样本和一个混合语音信号,两者都表示为梅尔频谱图,与提取目标扬声器的干净的语音的目标。此外,对于相位重建是至关重要的任务,我们提出了一种新的声码器的混合信号的复杂的STFT的条件下,使改进的相位估计。在标准TSE基准上的实验结果表明,FlowTSE匹配或优于强基线。
摘要:Target speaker extraction (TSE) aims to isolate a specific speaker's speech from a mixture using speaker enrollment as a reference. While most existing approaches are discriminative, recent generative methods for TSE achieve strong results. However, generative methods for TSE remain underexplored, with most existing approaches relying on complex pipelines and pretrained components, leading to computational overhead. In this work, we present FlowTSE, a simple yet effective TSE approach based on conditional flow matching. Our model receives an enrollment audio sample and a mixed speech signal, both represented as mel-spectrograms, with the objective of extracting the target speaker's clean speech. Furthermore, for tasks where phase reconstruction is crucial, we propose a novel vocoder conditioned on the complex STFT of the mixed signal, enabling improved phase estimation. Experimental results on standard TSE benchmarks show that FlowTSE matches or outperforms strong baselines.
标题: 缓解多标签语音情感识别中的亚群差异:伪标签和无监督学习方法
链接:https://arxiv.org/abs/2505.14449
备注:Accepted by InterSpeech 2025. 7 pages including 2 pages of appendix
摘要:虽然亚组差异和性能偏差越来越多地在计算研究中研究,分类语音情感识别(SER)的公平性仍有待探索。现有的方法通常依赖于明确的人口统计标签,由于隐私问题而难以获得。为了解决这一限制,我们引入了一个隐式人口统计推断(IDI)模块,该模块利用来自预训练模型的伪标记和使用k均值聚类的无监督学习来减轻SER中的偏差。我们的实验表明,伪标记IDI减少了子组差异,将公平性指标提高了33%以上,SER准确性降低了不到3%。此外,无监督IDI在公平性指标上提高了26%以上,SER性能下降了不到4%。进一步的分析表明,无监督IDI一贯减轻种族和年龄差异,显示其在明确的人口统计信息不可用的情况下的潜力。
摘要:While subgroup disparities and performance bias are increasingly studied in computational research, fairness in categorical Speech Emotion Recognition (SER) remains underexplored. Existing methods often rely on explicit demographic labels, which are difficult to obtain due to privacy concerns. To address this limitation, we introduce an Implicit Demography Inference (IDI) module that leverages pseudo-labeling from a pre-trained model and unsupervised learning using k-means clustering to mitigate bias in SER. Our experiments show that pseudo-labeling IDI reduces subgroup disparities, improving fairness metrics by over 33% with less than a 3% decrease in SER accuracy. Also, the unsupervised IDI yields more than a 26% improvement in fairness metrics with a drop of less than 4% in SER performance. Further analyses reveal that the unsupervised IDI consistently mitigates race and age disparities, demonstrating its potential in scenarios where explicit demographic information is unavailable.
标题: 利用距离和房间线索的单通道目标语音提取
链接:https://arxiv.org/abs/2505.14433
备注:5 pages, 3 figures, accepted by Eusipco 2025
摘要:本文旨在利用距离线索和房间信息实现封闭环境下的单通道目标语音提取。最近的工作已经验证了TSE任务的距离线索的可行性,其可以暗示声源的直达混响比(DRR),因此可以用于语音分离和TSE系统。然而,这样的距离线索受到房间的声学特性(诸如尺寸和混响时间)的显著影响,使得仅依赖于距离线索的TSE系统具有挑战性以概括各种不同的房间。为了解决这个问题,我们建议为基于距离的TSE提供房间环境信息(房间尺寸和混响时间),以获得更好的泛化能力。特别是,我们提出了一个基于距离和环境的时频(TF)域的TSE模型与学习距离和房间嵌入。模拟和真实数据集上的结果证明了该方法的可行性。演示材料可在https://runwushi.github.io/distance-room-demo-page/上获得。
摘要:This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of distance clues for the TSE task, which can imply the sound source's direct-to-reverberation ratio (DRR) and thus can be utilized for speech separation and TSE systems. However, such distance clue is significantly influenced by the room's acoustic characteristics, such as dimension and reverberation time, making it challenging for TSE systems that rely solely on distance clues to generalize across a variety of different rooms. To solve this, we suggest providing room environmental information (room dimensions and reverberation time) for distance-based TSE for better generalization capabilities. Especially, we propose a distance and environment-based TSE model in the time-frequency (TF) domain with learnable distance and room embedding. Results on both simulated and real collected datasets demonstrate its feasibility. Demonstration materials are available at https://runwushi.github.io/distance-room-demo-page/.
标题: 语音合成中口音相似度的成对评估
链接:https://arxiv.org/abs/2505.14410
备注:Accepted by INTERSPEECH 2025
摘要:尽管人们对生成高保真口音的兴趣越来越大,但在语音合成中评估口音相似性的研究还不够深入。我们的目标是加强口音相似度的主观和客观评估方法。主观上,我们通过添加组件来改进XAB听力测试,这些组件以更少的听众和更低的成本实现更高的统计显著性。我们的方法包括为听众提供transmittance,让他们突出感知的口音差异,并实施细致的可靠性筛选。客观地说,我们利用发音相关的指标,元音共振峰和语音posteriorgrams之间的距离的基础上,评估口音的产生。比较实验表明,这些指标,口音相似性,扬声器相似性,和梅尔倒谱失真,可以使用。此外,我们的研究结果强调了单词错误率等常用指标在评估代表性不足的口音方面的重大局限性。
摘要:Despite growing interest in generating high-fidelity accents, evaluating accent similarity in speech synthesis has been underexplored. We aim to enhance both subjective and objective evaluation methods for accent similarity. Subjectively, we refine the XAB listening test by adding components that achieve higher statistical significance with fewer listeners and lower costs. Our method involves providing listeners with transcriptions, having them highlight perceived accent differences, and implementing meticulous screening for reliability. Objectively, we utilise pronunciation-related metrics, based on distances between vowel formants and phonetic posteriorgrams, to evaluate accent generation. Comparative experiments reveal that these metrics, alongside accent similarity, speaker similarity, and Mel Cepstral Distortion, can be used. Moreover, our findings underscore significant limitations of common metrics like Word Error Rate in assessing underrepresented accents.
标题: 扩展和增强基于LLM的AVSR:投影仪的稀疏混合方法
链接:https://arxiv.org/abs/2505.14336
摘要:视听语音识别(AVSR)通过整合视觉线索来增强噪声环境中的鲁棒性。虽然最近的进展将大型语言模型(LLM)集成到AVSR中,但其高计算成本阻碍了在资源受限环境中的部署。为了解决这个问题,我们提出了Llama-SMoP,这是一种高效的多模态LLM,它采用稀疏混合投影仪(SMoP)模块来扩展模型容量,而不会增加推理成本。通过整合稀疏选通专家混合(MoE)投影仪,Llama-SMoP能够在保持强大性能的同时使用更小的LLM。我们探讨了三种SMoP配置,并表明Llama-SMoP DEDR(分离专家,分离路由器),它使用特定于模态的路由器和专家,在ASR,VSR和AVSR任务上实现了卓越的性能。消融研究证实了其在专家激活、可扩展性和噪声鲁棒性方面的有效性。
摘要:Audio-Visual Speech Recognition (AVSR) enhances robustness in noisy environments by integrating visual cues. While recent advances integrate Large Language Models (LLMs) into AVSR, their high computational cost hinders deployment in resource-constrained settings. To address this, we propose Llama-SMoP, an efficient Multimodal LLM that employs a Sparse Mixture of Projectors (SMoP) module to scale model capacity without increasing inference costs. By incorporating sparsely-gated mixture-of-experts (MoE) projectors, Llama-SMoP enables the use of smaller LLMs while maintaining strong performance. We explore three SMoP configurations and show that Llama-SMoP DEDR (Disjoint-Experts, Disjoint-Routers), which uses modality-specific routers and experts, achieves superior performance on ASR, VSR, and AVSR tasks. Ablation studies confirm its effectiveness in expert activation, scalability, and noise robustness.
标题: 无障碍编辑:背景噪音感知Zero-Shot语音编辑,具有上下文增强
链接:https://arxiv.org/abs/2505.14066
备注:5 pages, 3 figures
摘要:随着zero-shot文语转换技术的迅速发展,产生高质量的、与真实语音难以区分的语音信号成为可能。语音编辑,包括语音插入和替换,由于其潜在的应用吸引了研究人员。然而,现有的研究只考虑了干净的语音场景。在实际应用中,环境噪声的存在会显著降低生成的质量。在这项研究中,我们提出了一个噪声弹性语音编辑框架,无噪声编辑,嘈杂的语音编辑。无噪声编辑采用了一个频段感知的噪声抑制模块和一个内容细化策略。它可以很好地解决语音和背景噪声的频带不分离的情况。所提出的无障碍编辑框架在多个定量和定性评价中优于最先进的方法。
摘要:With the fast development of zero-shot text-to-speech technologies, it is possible to generate high-quality speech signals that are indistinguishable from the real ones. Speech editing, including speech insertion and replacement, appeals to researchers due to its potential applications. However, existing studies only considered clean speech scenarios. In real-world applications, the existence of environmental noise could significantly degrade the quality of the generation. In this study, we propose a noise-resilient speech editing framework, SeamlessEdit, for noisy speech editing. SeamlessEdit adopts a frequency-band-aware noise suppression module and an in-content refinement strategy. It can well address the scenario where the frequency bands of voice and background noise are not separated. The proposed SeamlessEdit framework outperforms state-of-the-art approaches in multiple quantitative and qualitative evaluations.
标题: 具有动态温度的自然感知课程学习用于语音深度伪造检测
链接:https://arxiv.org/abs/2505.13976
备注:Accepted by Interspeech 2025
摘要:语音深度伪造检测(SDD)的最新进展显着改善了欺骗语音中基于伪影的检测。然而,大多数模型忽略了语音自然度,这是区分真实语音和欺骗语音的关键线索。本研究提出了自然感知课程学习,一种新的训练框架,利用语音自然度来增强SDD的鲁棒性和泛化能力。这种方法使用地面实况标签和平均意见得分来测量样本难度,并调整训练时间表以逐步引入更具挑战性的样本。为了进一步提高泛化能力,在训练过程中引入了基于语音自然度的动态温度缩放方法。在ASVspoof 2021 DF数据集上的实验中,EER相对降低了23%,而无需修改模型架构。消融研究证实了SDD任务的自然意识训练策略的有效性。
摘要:Recent advances in speech deepfake detection (SDD) have significantly improved artifacts-based detection in spoofed speech. However, most models overlook speech naturalness, a crucial cue for distinguishing bona fide speech from spoofed speech. This study proposes naturalness-aware curriculum learning, a novel training framework that leverages speech naturalness to enhance the robustness and generalization of SDD. This approach measures sample difficulty using both ground-truth labels and mean opinion scores, and adjusts the training schedule to progressively introduce more challenging samples. To further improve generalization, a dynamic temperature scaling method based on speech naturalness is incorporated into the training process. A 23% relative reduction in the EER was achieved in the experiments on the ASVspoof 2021 DF dataset, without modifying the model architecture. Ablation studies confirmed the effectiveness of naturalness-aware training strategies for SDD tasks.
标题: U-Sam:统一语音、音频和音乐理解的音频语言模型
链接:https://arxiv.org/abs/2505.13880
备注:Accepted to Interspeech 2025
摘要:音频任务的文本生成范式为统一的音频理解开辟了新的可能性。然而,现有的模型在实现对不同音频类型(如语音、一般音频事件和音乐)的全面理解方面面临着重大挑战。此外,它们完全依赖交叉熵损失进行对齐往往是不够的,因为它平等地对待所有标记,并且无法考虑冗余的音频特征,导致较弱的交叉模态对齐。为了应对上述挑战,本文介绍了U-SAM,这是一种先进的音频语言模型,它将语音,音频和音乐的专用编码器与预训练的大型语言模型(LLM)集成在一起。U-SAM采用混合专家(MoE)投影仪进行任务感知特征融合,动态路由和集成特定于域的编码器输出。此外,U-SAM还集成了一个语义感知对比损失模块,该模块在语言监督下明确识别冗余音频特征,并纠正其语义和频谱表示,以增强跨模态对齐。大量的实验表明,U-SAM在多个基准测试中始终优于专业模型和现有的音频语言模型。此外,它在看不见的任务上表现出紧急能力,展示了它的泛化潜力。代码可用(https://github.com/Honee-W/U-SAM/)。
摘要:The text generation paradigm for audio tasks has opened new possibilities for unified audio understanding. However, existing models face significant challenges in achieving a comprehensive understanding across diverse audio types, such as speech, general audio events, and music. Furthermore, their exclusive reliance on cross-entropy loss for alignment often falls short, as it treats all tokens equally and fails to account for redundant audio features, leading to weaker cross-modal alignment. To deal with the above challenges, this paper introduces U-SAM, an advanced audio language model that integrates specialized encoders for speech, audio, and music with a pre-trained large language model (LLM). U-SAM employs a Mixture of Experts (MoE) projector for task-aware feature fusion, dynamically routing and integrating the domain-specific encoder outputs. Additionally, U-SAM incorporates a Semantic-Aware Contrastive Loss Module, which explicitly identifies redundant audio features under language supervision and rectifies their semantic and spectral representations to enhance cross-modal alignment. Extensive experiments demonstrate that U-SAM consistently outperforms both specialized models and existing audio language models across multiple benchmarks. Moreover, it exhibits emergent capabilities on unseen tasks, showcasing its generalization potential. Code is available (https://github.com/Honee-W/U-SAM/).
标题: 一种基于语义信息的分层语音增强方法
链接:https://arxiv.org/abs/2505.13843
备注:Accepted by interspeech 2025
摘要:目前的语音增强(SE)方法大多通过直接估计时频掩模或频谱来从噪声输入中恢复干净的语音。然而,这些方法往往忽略了语音信号中固有的不同属性,如语义内容和声学细节,这可能会阻碍下游任务的性能。此外,它们的有效性往往会在复杂的声学环境中降低。为了克服这些挑战,我们提出了一种新的,语义信息为基础的,一步一步的因子分解SE方法使用因子分解编解码器和扩散模型。与传统的SE方法不同,我们的语义和声学属性的分层建模能够实现更强大的干净语音恢复,特别是在具有挑战性的声学场景中。此外,该方法为下游TTS任务提供了进一步的优势。实验结果表明,我们的算法不仅优于SOTA基线的语音质量,但也提高了TTS在噪声环境中的性能。
摘要:Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as semantic content and acoustic details, inherent in speech signals, which can hinder performance in downstream tasks. Moreover, their effectiveness tends to degrade in complex acoustic environments. To overcome these challenges, we propose a novel, semantic information-based, step-by-step factorized SE method using factorized codec and diffusion model. Unlike traditional SE methods, our hierarchical modeling of semantic and acoustic attributes enables more robust clean speech recovery, particularly in challenging acoustic scenarios. Moreover, this method offers further advantages for downstream TTS tasks. Experimental results demonstrate that our algorithm not only outperforms SOTA baselines in terms of speech quality but also enhances TTS performance in noisy environments.
标题: 基于离散声学标记去噪的LLM零拍TTS抗噪性研究
链接:https://arxiv.org/abs/2505.13830
备注:Accepted by Interspeech 2025
摘要:基于大语言模型(LLM)的zero-shot文本到语音(TTS)方法倾向于保留音频提示的声学环境,当音频提示包含噪声时,导致合成语音质量下降。在本文中,我们提出了一种新的基于神经编解码器的语音去噪器,并将其与先进的基于LLM的TTS模型,LauraTTS,实现噪声鲁棒的zero-shot TTS。建议的编解码器去噪器包括一个音频编解码器,令牌去噪器,和嵌入细化。令牌去噪器从噪声中预测出前两组干净的声学令牌,这些令牌可以作为声学提示,以使LauraTTS合成高质量的个性化语音或通过嵌入细化器和编解码器解码器转换为干净的语音波形。实验结果表明,我们提出的编解码器去噪优于国家的最先进的语音增强(SE)的方法,和建议的噪声鲁棒性LauraTTS超越使用额外的SE模型的方法。
摘要:Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech quality when the audio prompt contains noise. In this paper, we propose a novel neural codec-based speech denoiser and integrate it with the advanced LLM-based TTS model, LauraTTS, to achieve noise-robust zero-shot TTS. The proposed codec denoiser consists of an audio codec, a token denoiser, and an embedding refiner. The token denoiser predicts the first two groups of clean acoustic tokens from the noisy ones, which can serve as the acoustic prompt for LauraTTS to synthesize high-quality personalized speech or be converted to clean speech waveforms through the embedding refiner and codec decoder. Experimental results show that our proposed codec denoiser outperforms state-of-the-art speech enhancement (SE) methods, and the proposed noise-robust LauraTTS surpasses the approach using additional SE models.
标题: 用维度正规化和分数正规化推动自蒸馏原型网络的前沿
链接:https://arxiv.org/abs/2505.13826
摘要:开发没有说话人标签的鲁棒说话人验证(SV)系统一直是一个长期的挑战。早期的研究强调了自我监督和完全监督方法之间的相当大的性能差距。在本文中,我们通过引入维度正则化来增强非对比自监督框架自蒸馏原型网络(SDPN),该维度正则化通过将正则化项应用于说话人嵌入来明确解决崩溃问题。此外,我们整合了来自全监督SV的分数归一化技术,以进一步弥合监督验证性能的差距。具有维度正则化和分数归一化的SDPN在VoxCeleb 1说话人验证评估基准上设置了一个新的最先进的水平,对于试验VoxCeleb 1-{O,E,H}分别实现了1.29%,1.60%和2.80%的等错误率。这些结果表明,与当前最好的自监督方法相比,相对提高了28.3%,19.6%和22.6%,从而推进了SV技术的前沿。
摘要:Developing robust speaker verification (SV) systems without speaker labels has been a longstanding challenge. Earlier research has highlighted a considerable performance gap between self-supervised and fully supervised approaches. In this paper, we enhance the non-contrastive self-supervised framework, Self-Distillation Prototypes Network (SDPN), by introducing dimension regularization that explicitly addresses the collapse problem through the application of regularization terms to speaker embeddings. Moreover, we integrate score normalization techniques from fully supervised SV to further bridge the gap toward supervised verification performance. SDPN with dimension regularization and score normalization sets a new state-of-the-art on the VoxCeleb1 speaker verification evaluation benchmark, achieving Equal Error Rate 1.29%, 1.60%, and 2.80% for trial VoxCeleb1-{O,E,H} respectively. These results demonstrate relative improvements of 28.3%, 19.6%, and 22.6% over the current best self-supervised methods, thereby advancing the frontiers of SV technology.
标题: 言语产生过程中表面肌电信号的关节特征预测
链接:https://arxiv.org/abs/2505.13814
备注:Accepted for Interspeech2025
摘要:我们提出了一个模型,用于预测发音功能的表面肌电图(EMG)信号在语音生产。该模型集成了卷积层和Transformer块,然后是发音特征的单独预测器。对于大多数发音特征,我们的方法实现了约0.9的高预测相关性。此外,我们证明,这些预测发音功能可以解码成可理解的语音波形。据我们所知,这是第一种通过发音特征从表面肌电信号解码语音波形的方法,为基于肌电信号的语音合成提供了一种新的方法。此外,我们分析了EMG电极放置和发音特征可预测性之间的关系,为优化EMG电极配置提供了知识驱动的见解。源代码和解码的语音样本是公开的。
摘要:We present a model for predicting articulatory features from surface electromyography (EMG) signals during speech production. The proposed model integrates convolutional layers and a Transformer block, followed by separate predictors for articulatory features. Our approach achieves a high prediction correlation of approximately 0.9 for most articulatory features. Furthermore, we demonstrate that these predicted articulatory features can be decoded into intelligible speech waveforms. To our knowledge, this is the first method to decode speech waveforms from surface EMG via articulatory features, offering a novel approach to EMG-based speech synthesis. Additionally, we analyze the relationship between EMG electrode placement and articulatory feature predictability, providing knowledge-driven insights for optimizing EMG electrode configurations. The source code and decoded speech samples are publicly available.
标题: 基于方向感知的神经声场在高保真度脉冲响应少频内插中的应用
链接:https://arxiv.org/abs/2505.13617
备注:Accepted at Interspeech 2025
摘要:声场的特性与声源和收听者周围环境的几何和空间特性有内在联系。声音传播的物理过程被捕获在称为房间脉冲响应(RIR)的时域信号中。使用神经场(NF)的先前工作允许从有限RIR测量学习RIR的空间连续表示。然而,先前的基于NF的方法集中于单声道全向或至多双耳收听者,其不能精确地捕获单个点处的真实声场的方向特性。我们提出了一个方向感知神经场(DANF),更明确地采用了立体声格式RIR的方向信息。虽然DANF固有地捕获源和听众之间的空间关系,我们进一步提出了方向感知损失。此外,我们调查的能力DANF适应新的房间,以各种方式,包括低级别的适应。
摘要:The characteristics of a sound field are intrinsically linked to the geometric and spatial properties of the environment surrounding a sound source and a listener. The physics of sound propagation is captured in a time-domain signal known as a room impulse response (RIR). Prior work using neural fields (NFs) has allowed learning spatially-continuous representations of RIRs from finite RIR measurements. However, previous NF-based methods have focused on monaural omnidirectional or at most binaural listeners, which does not precisely capture the directional characteristics of a real sound field at a single point. We propose a direction-aware neural field (DANF) that more explicitly incorporates the directional information by Ambisonic-format RIRs. While DANF inherently captures spatial relations between sources and listeners, we further propose a direction-aware loss. In addition, we investigate the ability of DANF to adapt to new rooms in various ways including low-rank adaptation.
标题: 精神:修补语音语言模型以防止越狱攻击
链接:https://arxiv.org/abs/2505.13541
摘要:语音语言模型(SLM)通过口头指令实现自然交互,通过检测语音中的细微差别更有效地捕获用户意图。与基于文本的模型相比,更丰富的语音信号引入了新的安全风险,因为对手可以通过向语音中注入难以察觉的噪声来更好地绕过安全机制。我们分析了对抗性攻击,发现SLM更容易受到越狱攻击,在某些情况下可以达到完美的100%攻击成功率。为了提高安全性,我们提出了事后修补防御,用于通过修改SLM的激活来干预推理过程,从而提高高达99%的鲁棒性,(i)对效用的影响可以忽略不计,(ii)无需任何重新训练。我们进行消融研究,以最大限度地提高我们的防御功效,并改善实用性/安全性权衡,并通过SLM特有的大规模基准进行验证。
摘要:Speech Language Models (SLMs) enable natural interactions via spoken instructions, which more effectively capture user intent by detecting nuances in speech. The richer speech signal introduces new security risks compared to text-based models, as adversaries can better bypass safety mechanisms by injecting imperceptible noise to speech. We analyze adversarial attacks and find that SLMs are substantially more vulnerable to jailbreak attacks, which can achieve a perfect 100% attack success rate in some instances. To improve security, we propose post-hoc patching defenses used to intervene during inference by modifying the SLM's activations that improve robustness up to 99% with (i) negligible impact on utility and (ii) without any re-training. We conduct ablation studies to maximize the efficacy of our defenses and improve the utility/security trade-off, validated with large-scale benchmarks unique to SLMs.
标题: 探索二元互动中的情感同步性:言语条件在面部和声音情感对齐中的作用
链接:https://arxiv.org/abs/2505.13455
摘要:理解人类如何在多个通信渠道,特别是面部表情和语音中表达和同步情感,对情感识别系统和人机交互具有重要意义。动机的概念,非重叠的语音促进更清晰的情感协调,而重叠的语音破坏同步,本研究探讨这些会话的动态形状的空间和时间对齐的唤醒和效价在面部和发声方式。使用IEMOCAP数据集的二元交互,我们通过WaveNet(面部视频)和基于Wav2Vec2的模型(语音音频)提取了连续的情感估计。根据语音重叠对片段进行分类,并使用Pearson相关性、滞后调整分析和动态时间规整(DTW)对情感对齐进行评估。在分析中,非重叠语音比重叠语音与更稳定和可预测的情感同步相关。虽然零滞后相关性低,没有统计学差异,非重叠的语音表现出减少的变异性,特别是唤醒。滞后调整的相关性和最佳滞后分布显示更清晰,更一致的时间在这些部分对齐。与此相反,重叠的语音表现出更高的变异性和平坦的滞后配置文件,虽然DTW表示意外更紧密的对齐,建议不同的协调策略。值得注意的是,方向性模式表明,面部表情更经常之前的话轮转换过程中的讲话,而在同时发声的讲话导致。这些发现强调了会话结构在调节情感交流中的重要性,并为现实世界互动中多模态情感对齐的时空动态提供了新的见解。
摘要:Understanding how humans express and synchronize emotions across multiple communication channels particularly facial expressions and speech has significant implications for emotion recognition systems and human computer interaction. Motivated by the notion that non-overlapping speech promotes clearer emotional coordination, while overlapping speech disrupts synchrony, this study examines how these conversational dynamics shape the spatial and temporal alignment of arousal and valence across facial and vocal modalities. Using dyadic interactions from the IEMOCAP dataset, we extracted continuous emotion estimates via EmoNet (facial video) and a Wav2Vec2-based model (speech audio). Segments were categorized based on speech overlap, and emotional alignment was assessed using Pearson correlation, lag adjusted analysis, and Dynamic Time Warping (DTW). Across analyses, non overlapping speech was associated with more stable and predictable emotional synchrony than overlapping speech. While zero-lag correlations were low and not statistically different, non overlapping speech showed reduced variability, especially for arousal. Lag adjusted correlations and best-lag distributions revealed clearer, more consistent temporal alignment in these segments. In contrast, overlapping speech exhibited higher variability and flatter lag profiles, though DTW indicated unexpectedly tighter alignment suggesting distinct coordination strategies. Notably, directionality patterns showed that facial expressions more often preceded speech during turn-taking, while speech led during simultaneous vocalizations. These findings underscore the importance of conversational structure in regulating emotional communication and provide new insight into the spatial and temporal dynamics of multimodal affective alignment in real world interaction.
标题: Vox-Profile:描述不同说话者和言语特征的言语基金会模型基准
链接:https://arxiv.org/abs/2505.14648
摘要:我们介绍Vox-Profile,一个全面的基准测试,使用语音基础模型来表征丰富的扬声器和语音特征。与现有的专注于说话者特质的单一维度的作品不同,Vox-Profile提供了反映静态说话者特质(例如,年龄、性别、口音)和动态语音属性(例如,情感、言语流)。该基准以语音科学和语言学为基础,与领域专家一起开发,以准确地索引说话者和语音特征。我们使用超过15个公开的语音数据集和几个广泛使用的语音基础模型,针对各种静态和动态扬声器和语音属性的基准实验报告。除了基准测试实验,我们展示了几个下游应用程序支持的Vox-Profile。首先,我们证明了Vox-Profile可以增强现有的语音识别数据集,以分析ASR性能的变化。Vox-Profile也被用作评估语音生成系统性能的工具。最后,我们评估我们的自动配置文件的质量,通过与人类的评价比较,并显示收敛的有效性。Vox-Profile可在https://github.com/tiantiaf0627/vox-profile-release上公开获取。
摘要:We introduce Vox-Profile, a comprehensive benchmark to characterize rich speaker and speech traits using speech foundation models. Unlike existing works that focus on a single dimension of speaker traits, Vox-Profile provides holistic and multi-dimensional profiles that reflect both static speaker traits (e.g., age, sex, accent) and dynamic speech properties (e.g., emotion, speech flow). This benchmark is grounded in speech science and linguistics, developed with domain experts to accurately index speaker and speech characteristics. We report benchmark experiments using over 15 publicly available speech datasets and several widely used speech foundation models that target various static and dynamic speaker and speech properties. In addition to benchmark experiments, we showcase several downstream applications supported by Vox-Profile. First, we show that Vox-Profile can augment existing speech recognition datasets to analyze ASR performance variability. Vox-Profile is also used as a tool to evaluate the performance of speech generation systems. Finally, we assess the quality of our automated profiles through comparison with human evaluation and show convergent validity. Vox-Profile is publicly available at: https://github.com/tiantiaf0627/vox-profile-release.
标题: 语言、音频和视觉模式语义对齐的表示学习
链接:https://arxiv.org/abs/2505.14562
备注:Accepted to European Signal Processing Conference (EUSIPCO 2025)
摘要:本文提出了一种单阶段训练方法,使用对比学习框架在语义上对齐三种模态-音频,视觉和文本。对比训练在多模态对齐方面取得了突出的成就,利用大规模未标记数据来学习共享表示。现有的三模态对齐的深度学习方法包括两个阶段,分别对齐视觉-文本和音频-文本模态。这种方法遭受不匹配的数据分布,导致次优对齐。利用AVCaps数据集,它为视频剪辑提供音频,视频和视听字幕,我们的方法使用对比训练联合优化了所有模态的表示。我们的研究结果表明,单阶段的方法优于两阶段的方法,实现了基于音频的视觉检索的两倍的改善,突出了统一的多模态表示学习的优势。
摘要:This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive training has gained prominence for multimodal alignment, utilizing large-scale unlabeled data to learn shared representations. Existing deep learning approach for trimodal alignment involves two-stages, that separately align visual-text and audio-text modalities. This approach suffers from mismatched data distributions, resulting in suboptimal alignment. Leveraging the AVCaps dataset, which provides audio, visual and audio-visual captions for video clips, our method jointly optimizes the representation of all the modalities using contrastive training. Our results demonstrate that the single-stage approach outperforms the two-stage method, achieving a two-fold improvement in audio based visual retrieval, highlighting the advantages of unified multimodal representation learning.
标题: 过去:声控语音令牌器
链接:https://arxiv.org/abs/2505.14470
摘要:我们提出了PAST,这是一种新型的端到端框架,它在信号重建的同时联合对语音信息进行建模,从而消除了对外部预训练模型的需求。与以前依赖于预训练的自监督模型的方法不同,PAST采用监督语音数据,通过辅助任务将领域知识直接集成到标记化过程中。此外,我们引入了一个流,因果的PAST变体,使实时语音应用程序。结果表明,PAST超越了现有的评估基线标记在共同的评估指标,包括语音表示和语音重建。值得注意的是,PAST在作为语音语言模型的语音表示时也实现了卓越的性能,进一步突出了其作为口语生成基础的有效性。为了促进进一步的研究,我们发布了完整的实现。有关代码、模型检查点和示例,请参见:https://pages.cs.huji.ac.il/adiyoss-lab/PAST
摘要:We present PAST, a novel end-to-end framework that jointly models phonetic information alongside signal reconstruction, eliminating the need for external pretrained models. Unlike previous approaches that rely on pretrained self-supervised models, PAST employs supervised phonetic data, directly integrating domain knowledge into the tokenization process via auxiliary tasks. Additionally, we introduce a streamable, causal variant of PAST, enabling real-time speech applications. Results demonstrate that PAST surpasses existing evaluated baseline tokenizers across common evaluation metrics, including phonetic representation and speech reconstruction. Notably, PAST also achieves superior performance when serving as a speech representation for speech language models, further highlighting its effectiveness as a foundation for spoken language generation. To foster further research, we release the full implementation. For code, model checkpoints, and samples see: https://pages.cs.huji.ac.il/adiyoss-lab/PAST
标题: 低音中提琴中提琴频率波动的复杂性和诠释风格
链接:https://arxiv.org/abs/2505.14448
备注:15 pages, 5 figures
摘要:将一组音乐作品中的音频信号建模为复杂网络,以研究低音中提琴频率波动的复杂性与演奏风格之间的关系。基于跨学科的科学和音乐方法,我们计算频谱分解并将其频率分量转换为声音网络。我们应用最佳拟合分析来识别更精确地描述此类频率行为的统计分布,并计算中心性度量并识别用于表征此类网络的集团。研究结果表明,统计分布的类型,最好地描述了频率波动的统计分布。中心性测度确定了一段音乐中最有影响力和最稳定的声音组,同时最大集团的识别表明了声音的功能组,这些功能组密切相互作用,以识别复杂的频率波动的出现。因此,通过将声音建模为复杂网络,我们可以清楚地将大规模统计波动的存在与同一音乐家演奏的不同音乐事件相关的类似频率波动的存在联系起来。
摘要:Audio signals in a set of musical pieces are modeled as a complex network for studying the relationship between the complexity of frequency fluctuations and the interpretive style of the bass viola da gamba. Based on interdisciplinary scientific and music approaches, we compute the spectral decomposition and translated its frequency components to a network of sounds. We applied a best fit analysis for identifying the statistical distributions that describe more precisely the behavior of such frequencies and computed the centrality measures and identify cliques for characterizing such a network. Findings suggested statistical regularities in the type of statistical distribution that best describes frequency fluctuations. The centrality measure confirmed the most influential and stable group of sounds in a piece of music, meanwhile the identification of the largest clique indicated functional groups of sounds that interact closely for identifying the emergence of complex frequency fluctuations. Therefore, by modeling the sound as a complex network, we can clearly associate the presence of large-scale statistical regularities with the presence of similar frequency fluctuations related to different musical events played by a same musician.
标题: S2 SBench:量化语音到语音大型语言模型中智力退化的基准
链接:https://arxiv.org/abs/2505.14438
摘要:端到端语音大语言模型(LLM)扩展了基于文本的模型的功能,可以直接处理和生成音频令牌。然而,与文本输入相比,这通常会导致推理和生成性能下降,这种现象被称为智能退化。为了系统地评估这一差距,我们提出了S2SBench,一个旨在量化语音LLM性能下降的基准。它包括针对音频输入下的句子延续和常识推理的诊断数据集。我们进一步介绍了一个成对的评估协议的基础上的困惑之间的差异似是而非的和难以置信的样本来衡量退化相对于文本输入。我们应用S2SBench对百川音频的训练过程进行了分析,进一步验证了该基准的有效性。所有数据集和评估代码都可以在https://github.com/undobug/S2SBench上找到。
摘要:End-to-end speech large language models ((LLMs)) extend the capabilities of text-based models to directly process and generate audio tokens. However, this often leads to a decline in reasoning and generation performance compared to text input, a phenomenon referred to as intelligence degradation. To systematically evaluate this gap, we propose S2SBench, a benchmark designed to quantify performance degradation in Speech LLMs. It includes diagnostic datasets targeting sentence continuation and commonsense reasoning under audio input. We further introduce a pairwise evaluation protocol based on perplexity differences between plausible and implausible samples to measure degradation relative to text input. We apply S2SBench to analyze the training process of Baichuan-Audio, which further demonstrates the benchmark's effectiveness. All datasets and evaluation code are available at https://github.com/undobug/S2SBench.
标题: PersonaTab:在全复式语音对话中使用文本、声学和行为线索预测性格特征
链接:https://arxiv.org/abs/2505.14356
备注:This is accepted to Interspeech 2025; Added an extra page for supplementary figures; Project page: this https URL
摘要:尽管在神经口语对话系统中取得了重大进展,但由于语音数据集中缺乏个性注释,个性感知对话代理(能够根据个性调整行为)仍然未得到充分研究。我们提出了一个预处理原始音频记录的管道,以创建一个带有时间戳,响应类型和情绪/情感标签的对话数据集。我们采用自动语音识别(ASR)系统来提取成绩单和时间戳,然后生成会话级注释。利用这些注释,我们设计了一个系统,采用大型语言模型来预测会话个性。人类评估人员参与识别会话特征并分配个性标签。我们的分析表明,该系统实现了更强的对齐与人类的判断相比,现有的方法。
摘要:Despite significant progress in neural spoken dialog systems, personality-aware conversation agents -- capable of adapting behavior based on personalities -- remain underexplored due to the absence of personality annotations in speech datasets. We propose a pipeline that preprocesses raw audio recordings to create a dialogue dataset annotated with timestamps, response types, and emotion/sentiment labels. We employ an automatic speech recognition (ASR) system to extract transcripts and timestamps, then generate conversation-level annotations. Leveraging these annotations, we design a system that employs large language models to predict conversational personality. Human evaluators were engaged to identify conversational characteristics and assign personality labels. Our analysis demonstrates that the proposed system achieves stronger alignment with human judgments compared to existing approaches.
标题: FMSD-TTC:用于Deliver-Tsang、Amdo和Kham语音数据集生成的Few-Shot多说话人多方言文本到语音合成
链接:https://arxiv.org/abs/2505.14351
备注:13 pages
摘要:藏语是一种低资源的语言,只有最少的平行语音语料库跨越其三大方言--“乌斯藏语、安多语和康语--限制了语音建模的进展。为了解决这个问题,我们提出了FMSD-TTS,一个Few-Shot,多扬声器,多方言的文本到语音的框架,合成并行方言语音从有限的参考音频和明确的方言标签。我们的方法具有一个新的扬声器方言融合模块和方言专用动态路由网络(DSDR-Net),以捕获细粒度的声音和语言的方言变化,同时保留扬声器的身份。大量的客观和主观评估表明,FMSD-TTS显着优于基线的方言表达和说话人相似性。我们通过一个具有挑战性的语音到语音的方言转换任务,进一步验证合成语音的质量和效用。我们的贡献包括:(1)一个为藏语多方言语音合成量身定制的新型Few-Shot TTS系统,(2)公开发布由FMSD-TTS生成的大规模合成藏语语音语料库,(3)一个开源评估工具包,用于标准化评估说话人相似性、方言一致性和音频质量。
摘要:Tibetan is a low-resource language with minimal parallel speech corpora spanning its three major dialects-\"U-Tsang, Amdo, and Kham-limiting progress in speech modeling. To address this issue, we propose FMSD-TTS, a few-shot, multi-speaker, multi-dialect text-to-speech framework that synthesizes parallel dialectal speech from limited reference audio and explicit dialect labels. Our method features a novel speaker-dialect fusion module and a Dialect-Specialized Dynamic Routing Network (DSDR-Net) to capture fine-grained acoustic and linguistic variations across dialects while preserving speaker identity. Extensive objective and subjective evaluations demonstrate that FMSD-TTS significantly outperforms baselines in both dialectal expressiveness and speaker similarity. We further validate the quality and utility of the synthesized speech through a challenging speech-to-speech dialect conversion task. Our contributions include: (1) a novel few-shot TTS system tailored for Tibetan multi-dialect speech synthesis, (2) the public release of a large-scale synthetic Tibetan speech corpus generated by FMSD-TTS, and (3) an open-source evaluation toolkit for standardized assessment of speaker similarity, dialect consistency, and audio quality.
标题: 语音灵活控制的通用声学对抗攻击-LLM
链接:https://arxiv.org/abs/2505.14286
摘要:预先训练的语音编码器与大型语言模型的组合使得能够开发能够处理各种口语处理任务的语音LLM。虽然这些模型功能强大且灵活,但这种灵活性可能使它们更容易受到对抗性攻击。为了研究这个问题的程度,在这项工作中,我们研究了通用的声音对抗攻击语音LLM。在这里,固定的、通用的、对抗性的音频段被前置到原始输入音频。我们最初调查的攻击,导致模型要么不产生输出或执行修改后的任务覆盖原来的提示。然后,我们将攻击的性质扩展为选择性的,以便仅当存在特定的输入属性(例如说话者性别或口语)时才激活。没有目标属性的输入应该不受影响,允许对模型输出进行细粒度控制。我们的研究结果揭示了Qwen 2-Audio和Granite-Speech中的关键漏洞,并表明类似的语音LLM可能容易受到普遍对抗性攻击。这突出表明需要更强大的培训策略和提高对抗性攻击的抵抗力。
摘要:The combination of pre-trained speech encoders with large language models has enabled the development of speech LLMs that can handle a wide range of spoken language processing tasks. While these models are powerful and flexible, this very flexibility may make them more vulnerable to adversarial attacks. To examine the extent of this problem, in this work we investigate universal acoustic adversarial attacks on speech LLMs. Here a fixed, universal, adversarial audio segment is prepended to the original input audio. We initially investigate attacks that cause the model to either produce no output or to perform a modified task overriding the original prompt. We then extend the nature of the attack to be selective so that it activates only when specific input attributes, such as a speaker gender or spoken language, are present. Inputs without the targeted attribute should be unaffected, allowing fine-grained control over the model outputs. Our findings reveal critical vulnerabilities in Qwen2-Audio and Granite-Speech and suggest that similar speech LLMs may be susceptible to universal adversarial attacks. This highlights the need for more robust training strategies and improved resistance to adversarial attacks.
标题: AcquaSignal:鲁棒水下声学分析的集成框架
链接:https://arxiv.org/abs/2505.14285
备注:8 pages; 9 figures
摘要:本文介绍了AquaSignal,这是一个模块化和可扩展的管道,用于水声信号的预处理,去噪,分类和新颖性检测。AquaSignal旨在在嘈杂和动态的海洋环境中有效运行,集成了最先进的深度学习架构,以提高声学信号分析的可靠性和准确性。该系统在Deepship和Ocean Networks Canada(ONC)基准的组合数据集上进行评估,提供了一组不同的真实水下场景。AquaSignal采用U-Net架构进行去噪,使用ResNet 18卷积神经网络对已知声学事件进行分类,并使用基于AutoEncoder的模型对新信号或异常信号进行无监督检测。据我们所知,这是第一次全面的研究,应用和评估这种技术组合的海上船舶声学数据。实验结果表明,AquaSignal提高了信号的清晰度和任务性能,实现了71%的分类准确率和91%的新奇检测准确率。尽管与一些最先进的模型相比,分类性能略低,但数据划分策略的差异限制了直接比较。总体而言,AquaSignal在科学,环境和海洋领域的实时水声监测方面表现出强大的潜力。
摘要:This paper presents AquaSignal, a modular and scalable pipeline for preprocessing, denoising, classification, and novelty detection of underwater acoustic signals. Designed to operate effectively in noisy and dynamic marine environments, AquaSignal integrates state-of-the-art deep learning architectures to enhance the reliability and accuracy of acoustic signal analysis. The system is evaluated on a combined dataset from the Deepship and Ocean Networks Canada (ONC) benchmarks, providing a diverse set of real-world underwater scenarios. AquaSignal employs a U-Net architecture for denoising, a ResNet18 convolutional neural network for classifying known acoustic events, and an AutoEncoder-based model for unsupervised detection of novel or anomalous signals. To our knowledge, this is the first comprehensive study to apply and evaluate this combination of techniques on maritime vessel acoustic data. Experimental results show that AquaSignal improves signal clarity and task performance, achieving 71% classification accuracy and 91% accuracy in novelty detection. Despite slightly lower classification performance compared to some state-of-the-art models, differences in data partitioning strategies limit direct comparisons. Overall, AquaSignal demonstrates strong potential for real-time underwater acoustic monitoring in scientific, environmental, and maritime domains.
标题: MatchDance:Mamba-Transformer协作架构,匹配高质量3D舞蹈合成
链接:https://arxiv.org/abs/2505.14222
摘要:从音乐到舞蹈的生成是一项具有挑战性但又是关键的任务,它涉及编舞、虚拟现实和创意内容生成的交叉点。尽管它的意义,现有的方法面临着很大的限制,在实现编排的一致性。为了应对这一挑战,我们提出了MatchDance,一个新的框架,音乐舞蹈生成,构建了一个潜在的代表性,以提高编舞的一致性。MatchDance采用两阶段设计:(1)基于运动学-动态的量化阶段(KDQS),其通过具有运动学-动态约束的有限标量量化(FSQ)将舞蹈动作编码成潜在表示,并且以高保真度重构它们,以及(2)混合音乐到舞蹈生成阶段(HMDGS),其使用Mamba-Transformer混合架构来将音乐映射到潜在表示,然后通过KDQS解码器生成3D舞蹈动作。此外,音乐舞蹈检索框架和全面的指标进行评估。在FineDance数据集上进行的大量实验展示了最先进的性能。代码将在接受后发布。
摘要:Music-to-dance generation represents a challenging yet pivotal task at the intersection of choreography, virtual reality, and creative content generation. Despite its significance, existing methods face substantial limitation in achieving choreographic consistency. To address the challenge, we propose MatchDance, a novel framework for music-to-dance generation that constructs a latent representation to enhance choreographic consistency. MatchDance employs a two-stage design: (1) a Kinematic-Dynamic-based Quantization Stage (KDQS), which encodes dance motions into a latent representation by Finite Scalar Quantization (FSQ) with kinematic-dynamic constraints and reconstructs them with high fidelity, and (2) a Hybrid Music-to-Dance Generation Stage(HMDGS), which uses a Mamba-Transformer hybrid architecture to map music into the latent representation, followed by the KDQS decoder to generate 3D dance motions. Additionally, a music-dance retrieval framework and comprehensive metrics are introduced for evaluation. Extensive experiments on the FineDance dataset demonstrate state-of-the-art performance. Code will be released upon acceptance.
标题: Speech Deepfakes的源验证
链接:https://arxiv.org/abs/2505.14188
备注:Accepted at INTERSPEECH 2025
摘要:随着语音deepfake生成器的普及,不仅要评估合成音频的真实性,还要追踪其来源。虽然源归因模型试图解决这一挑战,但它们往往在开放集条件下与看不见的生成器作斗争。在本文中,我们介绍了源验证任务,它的灵感来自说话人验证,确定是否使用相同的模型作为一组参考信号的测试轨道。我们的方法利用了一个经过源属性训练的分类器的嵌入,计算轨道之间的距离分数,以评估它们是否来自同一个源。我们在不同的场景中评估多个模型,分析说话者多样性,语言不匹配和后处理操作的影响。这项工作首次探索了来源验证,突出了其潜力和漏洞,并为现实世界的法医应用提供了见解。
摘要:With the proliferation of speech deepfake generators, it becomes crucial not only to assess the authenticity of synthetic audio but also to trace its origin. While source attribution models attempt to address this challenge, they often struggle in open-set conditions against unseen generators. In this paper, we introduce the source verification task, which, inspired by speaker verification, determines whether a test track was produced using the same model as a set of reference signals. Our approach leverages embeddings from a classifier trained for source attribution, computing distance scores between tracks to assess whether they originate from the same source. We evaluate multiple models across diverse scenarios, analyzing the impact of speaker diversity, language mismatch, and post-processing operations. This work provides the first exploration of source verification, highlighting its potential and vulnerabilities, and offers insights for real-world forensic applications.
标题: AudSemThinker:通过声音语义推理增强音频语言模型
链接:https://arxiv.org/abs/2505.14142
摘要:音频语言模型已经在各种声音理解任务中显示出有希望的结果,但它们在细粒度声音语义上的推理能力仍然有限。在本文中,我们提出了AudSemThinker,一个模型,其推理是围绕一个框架的听觉语义启发人类认知。为了支持这一点,我们引入了AudSem,这是一个专门为音频语言模型中的语义描述符推理而策划的新数据集。AudSem通过提供经过仔细过滤的音频样本集合以及通过强大的多级管道生成的字幕,解决了zero-shot评估中数据污染的持续挑战。我们的实验表明,AudSemThinker在多个训练设置中的表现优于最先进的模型,突出了其在语义音频推理方面的优势。AudSemThinker和AudSem数据集都是公开发布的。
摘要:Audio-language models have shown promising results in various sound understanding tasks, yet they remain limited in their ability to reason over the fine-grained semantics of sound. In this paper, we present AudSemThinker, a model whose reasoning is structured around a framework of auditory semantics inspired by human cognition. To support this, we introduce AudSem, a novel dataset specifically curated for semantic descriptor reasoning in audio-language models. AudSem addresses the persistent challenge of data contamination in zero-shot evaluations by providing a carefully filtered collection of audio samples paired with captions generated through a robust multi-stage pipeline. Our experiments demonstrate that AudSemThinker outperforms state-of-the-art models across multiple training settings, highlighting its strength in semantic audio reasoning. Both AudSemThinker and the AudSem dataset are released publicly.
标题: AudioJailbreak:针对端到端大型音频语言模型的越狱攻击
链接:https://arxiv.org/abs/2505.14103
摘要:最近研究了针对大型音频语言模型(LALM)的越狱攻击,但它们实现了次优的有效性,适用性和实用性,特别是假设对手可以完全操纵用户提示。在这项工作中,我们首先进行了广泛的实验表明,先进的文本越狱攻击不能很容易地通过文本到语音(TTS)技术移植到端到端的LALM。然后,我们提出了AudioJailbreak,一种新颖的音频越狱攻击,具有(1)隐蔽性:通过制作后缀越狱音频,越狱音频不需要在时间轴上与用户提示对齐;(2)通用性:通过将多个提示合并到扰动生成中,单个越狱扰动对不同的提示有效;(3)隐蔽性:越狱音频的恶意意图不会通过提出各种意图隐藏策略来提高受害者的意识;以及(4)空中鲁棒性:越狱音频通过将混响失真效果与房间脉冲响应结合到扰动的产生。相比之下,所有先前的音频越狱攻击都不能提供隐蔽性、普遍性、隐蔽性或空中鲁棒性。此外,AudioJailbreak还适用于无法完全操纵用户提示的对手,因此具有更广泛的攻击场景。迄今为止,大多数LALM的广泛实验证明了AudioJailbreak的高效率。我们强调,我们的工作窥视到对LALM的音频越狱攻击的安全影响,并切实促进提高其安全鲁棒性。实现和音频示例可在我们的网站https://audiojailbreak.github.io/AudioJailbreak上获得。
摘要:Jailbreak attacks to Large audio-language models (LALMs) are studied recently, but they achieve suboptimal effectiveness, applicability, and practicability, particularly, assuming that the adversary can fully manipulate user prompts. In this work, we first conduct an extensive experiment showing that advanced text jailbreak attacks cannot be easily ported to end-to-end LALMs via text-to speech (TTS) techniques. We then propose AudioJailbreak, a novel audio jailbreak attack, featuring (1) asynchrony: the jailbreak audio does not need to align with user prompts in the time axis by crafting suffixal jailbreak audios; (2) universality: a single jailbreak perturbation is effective for different prompts by incorporating multiple prompts into perturbation generation; (3) stealthiness: the malicious intent of jailbreak audios will not raise the awareness of victims by proposing various intent concealment strategies; and (4) over-the-air robustness: the jailbreak audios remain effective when being played over the air by incorporating the reverberation distortion effect with room impulse response into the generation of the perturbations. In contrast, all prior audio jailbreak attacks cannot offer asynchrony, universality, stealthiness, or over-the-air robustness. Moreover, AudioJailbreak is also applicable to the adversary who cannot fully manipulate user prompts, thus has a much broader attack scenario. Extensive experiments with thus far the most LALMs demonstrate the high effectiveness of AudioJailbreak. We highlight that our work peeks into the security implications of audio jailbreak attacks against LALMs, and realistically fosters improving their security robustness. The implementation and audio samples are available at our website https://audiojailbreak.github.io/AudioJailbreak.
标题: 用语言和语音模型嵌入重建语音产生期间的神经活动
链接:https://arxiv.org/abs/2505.14074
备注:Accepted for presentation at Interspeech2025
摘要:了解神经活动如何编码语音和语言产生是神经科学和人工智能的一个基本挑战。这项研究调查了来自大规模自监督语言和语音模型的嵌入是否可以有效地重建语音产生过程中捕获的神经活动记录。我们利用基于语言和声学数据训练的深度学习模型的预训练嵌入来表示高级语音特征,并将其映射到神经信号上。我们分析了这些嵌入在多大程度上保留了大脑活动的时空动态。我们使用相关性度量和信号重建质量评估来评估重建的神经信号与地面真实记录的关系。结果表明,神经活动可以有效地重建使用嵌入从大型语言和语音模型在所有研究参与者,产生皮尔逊相关系数范围从0.79到0.99。
摘要:Understanding how neural activity encodes speech and language production is a fundamental challenge in neuroscience and artificial intelligence. This study investigates whether embeddings from large-scale, self-supervised language and speech models can effectively reconstruct neural activity recordings captured during speech production. We leverage pre-trained embeddings from deep learning models trained on linguistic and acoustic data to represent high-level speech features and map them onto neural signals. We analyze the extent to which these embeddings preserve the spatio-temporal dynamics of brain activity. We evaluate reconstructed neural signals against ground truth recordings using correlation metrics and signal reconstruction quality assessments. The results indicate that neural activity can be effectively reconstructed using embeddings from large language and speech models across all study participants, yielding Pearson correlation coefficients ranging from 0.79 to 0.99.
标题: 将确定性增强条件与双流编码相结合用于基于扩散的语音增强
链接:https://arxiv.org/abs/2505.13983
摘要:基于扩散的语音增强(SE)模型需要将正确的先验知识作为可靠的条件来生成准确的预测。然而,使用噪声特征提供可靠的条件是具有挑战性的。一种解决方案是使用由确定性方法增强的特征作为条件。然而,确定性方法所造成的信息失真和损失可能会影响扩散过程。在本文中,我们首先研究使用不同的确定性SE模型作为扩散条件的影响。我们验证两个条件,这取决于噪声特征是否被用作条件的一部分:一个仅使用确定性特征(仅确定性),另一个同时使用确定性和噪声特征(确定性噪声)。初步调查发现,使用确定性增强条件可以改善真实数据的听力体验,而使用仅确定性条件或确定性噪声条件之间的选择取决于确定性模型。基于这些发现,我们提出了一个双码流编码修复扩散模型SE(DERDM-SE),更有效地利用这两个条件。此外,我们发现细粒度的确定性模型在客观评价指标方面具有更大的潜力,而基于UNet的确定性模型提供了更稳定的扩散性能。因此,在DERDM-SE中,我们提出了一个确定性模型,它结合了粗粒度和细粒度的处理。CHiME 4上的实验结果表明,所提出的模型有效地利用确定性模型,以实现更好的SE评估分数,以及更稳定的性能相比,其他基于扩散的SE模型。
摘要:Diffusion-based speech enhancement (SE) models need to incorporate correct prior knowledge as reliable conditions to generate accurate predictions. However, providing reliable conditions using noisy features is challenging. One solution is to use features enhanced by deterministic methods as conditions. However, the information distortion and loss caused by deterministic methods might affect the diffusion process. In this paper, we first investigate the effects of using different deterministic SE models as conditions for diffusion. We validate two conditions depending on whether the noisy feature was used as part of the condition: one using only the deterministic feature (deterministic-only), and the other using both deterministic and noisy features (deterministic-noisy). Preliminary investigation found that using deterministic enhanced conditions improves hearing experiences on real data, while the choice between using deterministic-only or deterministic-noisy conditions depends on the deterministic models. Based on these findings, we propose a dual-streaming encoding Repair-Diffusion Model for SE (DERDM-SE) to more effectively utilize both conditions. Moreover, we found that fine-grained deterministic models have greater potential in objective evaluation metrics, while UNet-based deterministic models provide more stable diffusion performance. Therefore, in the DERDM-SE, we propose a deterministic model that combines coarse- and fine-grained processing. Experimental results on CHiME4 show that the proposed models effectively leverage deterministic models to achieve better SE evaluation scores, along with more stable performance compared to other diffusion-based SE models.
标题: 语音情感识别与个性的桥梁:数据集和时间交互条件网络
链接:https://arxiv.org/abs/2505.13978
摘要:本研究探讨人格特质与情绪表达之间的交互作用,探讨人格信息如何提高言语情绪识别能力。我们收集了IEMOCAP数据集的人格注释,统计分析发现人格特质和情绪表达之间存在显著相关性。为了提取细粒度的个性特征,我们提出了一个时间交互条件网络(TICN),其中个性特征与基于Hubert的声学特征相结合的SER。实验表明,将地面真实的个性特征显着提高效价识别,提高一致性相关系数(CCC)从0.698到0.785相比,没有个性信息的基线。针对对话系统中无法获取用户个性信息的实际应用,我们开发了一个自动个性识别的前端模块。使用这些自动预测的性状作为我们提出的TICN模型的输入,我们实现了CCC为0.776的效价识别,比基线相对提高了11.17%。这些发现证实了个性感知SER的有效性,并为进一步探索个性感知语音处理应用提供了坚实的基础。
摘要:This study investigates the interaction between personality traits and emotional expression, exploring how personality information can improve speech emotion recognition (SER). We collected personality annotation for the IEMOCAP dataset, and the statistical analysis identified significant correlations between personality traits and emotional expressions. To extract finegrained personality features, we propose a temporal interaction condition network (TICN), in which personality features are integrated with Hubert-based acoustic features for SER. Experiments show that incorporating ground-truth personality traits significantly enhances valence recognition, improving the concordance correlation coefficient (CCC) from 0.698 to 0.785 compared to the baseline without personality information. For practical applications in dialogue systems where personality information about the user is unavailable, we develop a front-end module of automatic personality recognition. Using these automatically predicted traits as inputs to our proposed TICN model, we achieve a CCC of 0.776 for valence recognition, representing an 11.17% relative improvement over the baseline. These findings confirm the effectiveness of personality-aware SER and provide a solid foundation for further exploration in personality-aware speech processing applications.
标题: 基于多模式信息的语音处理(MISP)2025挑战:视听数字化和识别
链接:https://arxiv.org/abs/2505.13971
备注:Accepted by Interspeech 2025. Camera-ready version
摘要:由于复杂的声学条件,会议是语音应用的一个有价值但具有挑战性的场景。本文总结了在Interspeech 2025上举办的MISP 2025挑战赛的成果,该挑战赛的重点是通过将视频模态与音频相结合来实现多模态、多设备会议转录。这些任务包括视听扬声器日记(AVSD),视听语音识别(AVSR)和视听日记和识别(AVDR)。我们提出了挑战的目标,任务,数据集,基线系统和参与者提出的解决方案。表现最好的系统在基线上取得了显著的改善:最好的AVSD模型实现了8.09%的日记错误率(DER),改善了7.43%;最好的AVSR系统实现了9.48%的字符错误率(CER),改善了10.62%;最佳AVDR系统的级联最小排列字符错误率(cpCER)为11.56%,提高了72.49%。
摘要:Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on multi-modal, multi-device meeting transcription by incorporating video modality alongside audio. The tasks include Audio-Visual Speaker Diarization (AVSD), Audio-Visual Speech Recognition (AVSR), and Audio-Visual Diarization and Recognition (AVDR). We present the challenge's objectives, tasks, dataset, baseline systems, and solutions proposed by participants. The best-performing systems achieved significant improvements over the baseline: the top AVSD model achieved a Diarization Error Rate (DER) of 8.09%, improving by 7.43%; the top AVSR system achieved a Character Error Rate (CER) of 9.48%, improving by 10.62%; and the best AVDR system achieved a concatenated minimum-permutation Character Error Rate (cpCER) of 11.56%, improving by 72.49%.
标题: BiCrossMamba-ST:采用双向Mamba光谱-时间交叉注意力的语音深度伪造检测
链接:https://arxiv.org/abs/2505.13930
备注:Accepted Interspeech 2025
摘要:我们提出了BiCrossMamba-ST,这是一个强大的语音deepfake检测框架,它利用了由双向Mamba块和相互交叉注意力驱动的双分支频谱-时间架构。通过分别处理频谱子带和时间间隔,然后整合它们的表示,BiCrossMamba-ST有效地捕捉合成语音的微妙线索。此外,我们提出的框架利用基于卷积的2D注意力地图来关注特定的频谱-时间区域,从而实现强大的深度伪造检测。BiCrossMamba-ST直接在原始特征上操作,实现了显着的性能改进,在ASVSpoof LA 21和ASVSpoof DF 21基准测试中分别比最先进的AASIST相对增益67.74%和26.3%,在ASVSpoof DF 21上比RawBMamba提高了6.80%。代码和模型将公开发布。
摘要:We propose BiCrossMamba-ST, a robust framework for speech deepfake detection that leverages a dual-branch spectro-temporal architecture powered by bidirectional Mamba blocks and mutual cross-attention. By processing spectral sub-bands and temporal intervals separately and then integrating their representations, BiCrossMamba-ST effectively captures the subtle cues of synthetic speech. In addition, our proposed framework leverages a convolution-based 2D attention map to focus on specific spectro-temporal regions, enabling robust deepfake detection. Operating directly on raw features, BiCrossMamba-ST achieves significant performance improvements, a 67.74% and 26.3% relative gain over state-of-the-art AASIST on ASVSpoof LA21 and ASVSpoof DF21 benchmarks, respectively, and a 6.80% improvement over RawBMamba on ASVSpoof DF21. Code and models will be made publicly available.
标题: 使用分段语音特征进行取证深度伪造音频检测
链接:https://arxiv.org/abs/2505.13847
摘要:这项研究探讨了使用分段语音的声学特征来检测deepfake音频的潜力。这些特征具有高度的可解释性,因为它们与人类的发音过程密切相关,预计Deepfake模型将更难以复制。实验结果表明,在语音比对中常用的某些分段特征在识别深度假声方面是有效的,而一些全局特征的识别价值不大。这些发现强调了在法医语音比较中以不同方式进行音频deepfake检测的必要性,并为利用分段特征提供了一个新的视角。
摘要:This study explores the potential of using acoustic features of segmental speech sounds to detect deepfake audio. These features are highly interpretable because of their close relationship with human articulatory processes and are expected to be more difficult for deepfake models to replicate. The results demonstrate that certain segmental features commonly used in forensic voice comparison are effective in identifying deep-fakes, whereas some global features provide little value. These findings underscore the need to approach audio deepfake detection differently for forensic voice comparison and offer a new perspective on leveraging segmental features for this purpose.
标题: ClapFM-EVC:高保真和灵活的情感语音转换,具有自然语言和语音的双重控制
链接:https://arxiv.org/abs/2505.13805
备注:Accepted by InterSpeech 2025
摘要:尽管取得了巨大的进步,但实现具有灵活和可解释控制的高保真情感语音转换(EVC)仍然具有挑战性。本文介绍了ClapFM-EVC,一种新的EVC框架,能够生成高质量的转换语音驱动的自然语言提示或参考语音与可调的情感强度。我们首先提出了EVC-CLAP,这是一种情感对比语言音频预训练模型,由自然语言提示和分类标签指导,以提取和对齐语音和文本模态中的细粒度情感元素。然后,提出了一种具有自适应强度门的FuEncoder,用于将情感特征与来自预训练ASR模型的语音后验图无缝融合。为了进一步提高情感表达能力和语音自然度,我们提出了一种基于这些特征的流匹配模型来重建源语音的Mel谱图。主观和客观评价验证了ClapFM-EVC的有效性。
摘要:Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a novel EVC framework capable of generating high-quality converted speech driven by natural language prompts or reference speech with adjustable emotion intensity. We first propose EVC-CLAP, an emotional contrastive language-audio pre-training model, guided by natural language prompts and categorical labels, to extract and align fine-grained emotional elements across speech and text modalities. Then, a FuEncoder with an adaptive intensity gate is presented to seamless fuse emotional features with Phonetic PosteriorGrams from a pre-trained ASR model. To further improve emotion expressiveness and speech naturalness, we propose a flow matching model conditioned on these captured features to reconstruct Mel-spectrogram of source speech. Subjective and objective evaluations validate the effectiveness of ClapFM-EVC.
标题: 基于能量的TTS模型的分数训练
链接:https://arxiv.org/abs/2505.13771
摘要:噪声对比估计(NCE)是一种用于训练具有难以处理的归一化项的基于能量的模型(EBM)的流行方法。NCE的关键思想是通过比较参考样本和噪声样本的非归一化对数似然来学习,从而避免显式计算归一化项。然而,NCE严重依赖于噪声样本的质量。最近,切片分数匹配(SSM)已被密切相关的扩散模型(DM)普及。与NCE不同,SSM通过学习其在随机选择的方向上的投影分布来学习对数似然或分数的梯度。然而,NCE和SSM都忽略了对数似然函数的形式,这是有问题的,因为EBM和DM在推理期间使用一阶优化。本文提出了一种新的标准,学习分数更适合于一阶格式。实验对比这些方法用于训练EBM。
摘要:Noise contrastive estimation (NCE) is a popular method for training energy-based models (EBM) with intractable normalisation terms. The key idea of NCE is to learn by comparing unnormalised log-likelihoods of the reference and noisy samples, thus avoiding explicitly computing normalisation terms. However, NCE critically relies on the quality of noisy samples. Recently, sliced score matching (SSM) has been popularised by closely related diffusion models (DM). Unlike NCE, SSM learns a gradient of log-likelihood, or score, by learning distribution of its projections on randomly chosen directions. However, both NCE and SSM disregard the form of log-likelihood function, which is problematic given that EBMs and DMs make use of first-order optimisation during inference. This paper proposes a new criterion that learns scores more suitable for first-order schemes. Experiments contrasts these approaches for training EBMs.
标题: VocalAgent:具有安全意识评估的声乐健康诊断大型语言模型
链接:https://arxiv.org/abs/2505.13577
摘要:声音健康在人们的生活中起着至关重要的作用,显著影响他们的沟通能力和互动。然而,尽管嗓音障碍在全球范围内普遍存在,但许多人无法获得方便的诊断和治疗。本文介绍了VocalAgent,一个音频大语言模型(LLM),以解决这些挑战,通过声乐健康诊断。我们利用Qwen-Audio-Chat对从医院患者现场收集的三个数据集进行了微调,并提出了一个多方面的评估框架,包括安全性评估,以减轻诊断偏倚,跨语言性能分析和模态消融研究。与最先进的基线相比,VocalAgent在语音障碍分类方面表现出更高的准确性。其基于LLM的方法为更广泛地采用健康诊断提供了可扩展的解决方案,同时强调了道德和技术验证的重要性。
摘要:Vocal health plays a crucial role in peoples' lives, significantly impacting their communicative abilities and interactions. However, despite the global prevalence of voice disorders, many lack access to convenient diagnosis and treatment. This paper introduces VocalAgent, an audio large language model (LLM) to address these challenges through vocal health diagnosis. We leverage Qwen-Audio-Chat fine-tuned on three datasets collected in-situ from hospital patients, and present a multifaceted evaluation framework encompassing a safety assessment to mitigate diagnostic biases, cross-lingual performance analysis, and modality ablation studies. VocalAgent demonstrates superior accuracy on voice disorder classification compared to state-of-the-art baselines. Its LLM-based method offers a scalable solution for broader adoption of health diagnostics, while underscoring the importance of ethical and technical validation.
机器翻译由腾讯交互翻译提供,仅供参考
