微信公众号:arXiv_Daily
cs.SD语音
【1】Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets
标题:使用MGPHot专家注释与通用标签数据集对音乐自动标记进行基准测试
链接:https://arxiv.org/abs/2509.06936
摘要:音乐自动标签旨在自动为音频记录分配描述性标签,例如流派,情绪或乐器。由于其挑战性、语义描述的多样性以及在各种应用中的实用价值,它已成为评估从音频数据中学习的通用音乐表示的性能的常见下游任务。我们介绍了一个新的基准数据集的基础上,最近公布的MGPHot数据集,其中包括专家音乐学注释,允许额外的见解和比较常见的通用标签数据集上获得的结果。虽然MGPHot注释已被证明对计算音乐学有用,但原始数据集既不包括音频,也不提供评估设置,以用作标准化的自动标记基准。为了解决这个问题,我们提供了一组精心策划的YouTube URL与可检索的音频,并提出了一个训练/验证/测试分裂标准化的评估,并预先计算表示为七个国家的最先进的模型。使用这些资源,我们在MGPHot和标准参考标签数据集中评估了这些模型,突出了专家和通用标签注释之间的关键差异。总之,我们的贡献提供了一个更先进的基准框架,为未来的研究在音乐理解。
摘要:Music autotagging aims to automatically assign descriptive tags, such as genre, mood, or instrumentation, to audio recordings. Due to its challenges, diversity of semantic descriptions, and practical value in various applications, it has become a common downstream task for evaluating the performance of general-purpose music representations learned from audio data. We introduce a new benchmarking dataset based on the recently published MGPHot dataset, which includes expert musicological annotations, allowing for additional insights and comparisons with results obtained on common generic tag datasets. While MGPHot annotations have been shown to be useful for computational musicology, the original dataset neither includes audio nor provides evaluation setups for its use as a standardized autotagging benchmark. To address this, we provide a curated set of YouTube URLs with retrievable audio, and propose a train/val/test split for standardized evaluation, and precomputed representations for seven state-of-the-art models. Using these resources, we evaluated these models in MGPHot and standard reference tag datasets, highlighting key differences between expert and generic tag annotations. Altogether, our contributions provide a more advanced benchmarking framework for future research in music understanding.
【2】Continuous Audio Language Models
标题:连续音频语言模型
链接:https://arxiv.org/abs/2509.06926
备注:17 pages, 3 figures
摘要:音频语言模型(ALM)通过将音频表示为离散令牌序列,已经成为语音和音乐生成的主导范式。然而,与可逆的文本令牌不同,音频令牌是从具有有限比特率的有损编解码器中提取的。因此,提高音频质量需要生成更多的令牌,这在保真度和计算成本之间强加了权衡。我们通过研究连续音频语言模型(CALM)来解决这个问题。这些模型实例化了一个大型的Transformer主干,它在每个时间步产生一个上下文嵌入。然后,该顺序信息调节MLP,该MLP通过一致性建模生成音频VAE的下一个连续帧。通过避免有损压缩,CALM以较低的计算成本实现了更高的质量。语音和音乐的实验表明,提高了效率和保真度的最先进的离散音频语言模型,促进轻量级,高质量的音频生成。样品可在https://continuous-audio-language-models.github.io上获得
摘要:Audio Language Models (ALM) have emerged as the dominant paradigm for speech and music generation by representing audio as sequences of discrete tokens. Yet, unlike text tokens, which are invertible, audio tokens are extracted from lossy codecs with a limited bitrate. As a consequence, increasing audio quality requires generating more tokens, which imposes a trade-off between fidelity and computational cost. We address this issue by studying Continuous Audio Language Models (CALM). These models instantiate a large Transformer backbone that produces a contextual embedding at every timestep. This sequential information then conditions an MLP that generates the next continuous frame of an audio VAE through consistency modeling. By avoiding lossy compression, CALM achieves higher quality at lower computational cost than their discrete counterpart. Experiments on speech and music demonstrate improved efficiency and fidelity over state-of-the-art discrete audio language models, facilitating lightweight, high-quality audio generation. Samples are available at https://continuous-audio-language-models.github.io
【3】AnalysisGNN: Unified Music Analysis with Graph Neural Networks
标题:AnalysisGNN:使用图神经网络的统一音乐分析
链接:https://arxiv.org/abs/2509.06654
备注:Accepted at the 17th International Symposium on Computer Music Multidisciplinary Research (CMMR) 2025
摘要:近年来,音乐分析的计算方法蓬勃发展,但每一种方法通常都是针对特定的分析领域。在这项工作中,我们引入了AnalysisGNN,这是一种新型的图神经网络框架,它利用数据洗牌策略,在特定任务分类器之间进行自定义加权多任务丢失和logit融合,以整合异质注释的符号数据集进行综合得分分析。我们还集成了一个非和弦音预测模块,它可以识别并排除所有任务中的传递和非功能性音符,从而提高标签信号的一致性。实验评估表明,AnalysisGNN实现了与传统静态数据集方法相当的性能,同时显示出对多个异构语料库中的域转移和注释不一致的弹性增强。
摘要:Recent years have seen a boom in computational approaches to music analysis, yet each one is typically tailored to a specific analytical domain. In this work, we introduce AnalysisGNN, a novel graph neural network framework that leverages a data-shuffling strategy with a custom weighted multi-task loss and logit fusion between task-specific classifiers to integrate heterogeneously annotated symbolic datasets for comprehensive score analysis. We further integrate a Non-Chord-Tone prediction module, which identifies and excludes passing and non-functional notes from all tasks, thereby improving the consistency of label signals. Experimental evaluations demonstrate that AnalysisGNN achieves performance comparable to traditional static-dataset approaches, while showing increased resilience to domain shifts and annotation inconsistencies across multiple heterogeneous corpora.
【4】The First Voice Timbre Attribute Detection Challenge
标题:首届语音音色属性检测挑战赛
链接:https://arxiv.org/abs/2509.06635
摘要:第一个语音音色属性检测挑战将在NCMMSC 2025的特别会议上进行。该方法主要研究语音音色的可解释性,并在一个特定的音色描述维度上比较两个语音的强度。在VCTK-RVA数据集上进行评价。参与者开发了自己的系统,并将其输出提交给组织者,组织者评估了绩效并向他们发送反馈。6个小组提交了产出,5个小组说明了方法。
摘要:The first voice timbre attribute detection challenge is featured in a special session at NCMMSC 2025. It focuses on the explainability of voice timbre and compares the intensity of two speech utterances in a specified timbre descriptor dimension. The evaluation was conducted on the VCTK-RVA dataset. Participants developed their systems and submitted their outputs to the organizer, who evaluated the performance and sent feedback to them. Six teams submitted their outputs, with five providing descriptions of their methodologies.
【5】Unveiling the Listener Structure Underlying K-pop's Global Success: A Large-Scale Listening Data Analysis
标题:揭示K-pop全球成功背后的结构:大规模听力数据分析
链接:https://arxiv.org/abs/2509.06606
摘要:从2000年代中期到2010年代,K-pop超越了其作为亚洲地区流行音乐的地位,并成为全球音乐流派,拥有世界各地的热情粉丝。然而,很少有人知道如何在全球范围内的广大音乐听众听和感知K-pop.本研究通过分析一个大规模的听力数据集从Last.fm解决这个问题.对播放次数分布的分析显示,在2005年至2019年期间,K-pop的播放次数显著增加,主要是由一小群重度听众支持。游戏数量的基尼系数明显大于现有的主流类型和其他正在增长的利基类型。此外,基于用户分配的流派标签的分析定量地表明,在2005年至2010年期间,K-pop摆脱了其作为亚洲本土流派的地位,并凭借自己的力量成为一种独特的音乐流派。
摘要:From the mid-2000s to the 2010s, K-pop moved beyond its status as a regionally popular genre in Asia and established itself as a global music genre with enthusiastic fans around the world. However, little is known about how the vast number of music listeners across the globe have listened to and perceived K-pop. This study addresses this question by analyzing a large-scale listening dataset from Last.fm. An analysis of the distribution of play counts reveals that K-pop experienced a significant increase in plays between 2005 and 2019, largely supported by a small group of heavy listeners. The Gini coefficient in play counts is notably greater than that of existing mainstream genres and other growing niche genres. Furthermore, an analysis based on user-assigned genre tags quantitatively demonstrates that between 2005 and 2010, K-pop shed its status as a local Asian genre and established itself as a distinct music genre in its own right.
【6】FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
标题:FireRedChat:一个具有级联和半级联实现的可插入、全Duplex语音交互系统
链接:https://arxiv.org/abs/2509.06502
备注:12 pages, 2 figures
摘要:全双工语音交互允许用户和代理同时说话,并可控制闯入,从而实现逼真的助理和客户服务。现有的解决方案要么是端到端的,难以设计和控制,要么是由轮流控制器管理的模块化管道,便于升级和每个模块的优化;然而,以前的模块化框架依赖于非开放组件和外部提供商,限制了整体优化。在这项工作中,我们提出了一个完整的,实用的全双工语音交互系统,包括一个话轮转换控制器,一个交互模块,和一个对话管理器。该控制器集成了流媒体个性化VAD(pVAD),以抑制来自噪声和非主要扬声器的错误闯入,精确地为主要扬声器段添加时间戳,并显式地启用主要扬声器闯入;语义结束回合检测器改善了停止决策。它将异构半双工流水线(级联、半级联和语音到语音)升级为全双工。使用内部模型,我们实现级联和半级联的变体;半级联的变体捕获情感和非语言线索,产生更连贯的响应,降低延迟和错误传播,并提高鲁棒性。对话管理器通过工具调用和上下文管理扩展功能。我们还提出了三个系统级的指标,闯入,回合结束检测精度和端到端的延迟,以评估自然度,控制精度和效率。实验表明,更少的错误中断,更准确的语义结束,更低的延迟接近工业系统,实现强大的,自然的,实时的全双工交互。演示:https://fireredteam.github.io/demos/firered_chat.
摘要:Full-duplex voice interaction allows users and agents to speak simultaneously with controllable barge-in, enabling lifelike assistants and customer service. Existing solutions are either end-to-end, difficult to design and hard to control, or modular pipelines governed by turn-taking controllers that ease upgrades and per-module optimization; however, prior modular frameworks depend on non-open components and external providers, limiting holistic optimization. In this work, we present a complete, practical full-duplex voice interaction system comprising a turn-taking controller, an interaction module, and a dialogue manager. The controller integrates streaming personalized VAD (pVAD) to suppress false barge-ins from noise and non-primary speakers, precisely timestamp primary-speaker segments, and explicitly enable primary-speaker barge-ins; a semantic end-of-turn detector improves stop decisions. It upgrades heterogeneous half-duplex pipelines, cascaded, semi-cascaded, and speech-to-speech, to full duplex. Using internal models, we implement cascaded and semi-cascaded variants; the semi-cascaded one captures emotional and paralinguistic cues, yields more coherent responses, lowers latency and error propagation, and improves robustness. A dialogue manager extends capabilities via tool invocation and context management. We also propose three system-level metrics, barge-in, end-of-turn detection accuracy, and end-to-end latency, to assess naturalness, control accuracy, and efficiency. Experiments show fewer false interruptions, more accurate semantic ends, and lower latency approaching industrial systems, enabling robust, natural, real-time full-duplex interaction. Demos: https://fireredteam.github.io/demos/firered_chat.
【7】AudioBoost: Increasing Audiobook Retrievability in Spotify Search with Synthetic Query Generation
标题:AudioBoost:通过合成查询生成提高Spotify搜索中有声读物检索能力
链接:https://arxiv.org/abs/2509.06452
备注:EARL Workshop @ RecSys25
摘要:Spotify最近推出了有声读物作为其目录的一部分,补充其音乐和播客产品。搜索通常是用户访问新项目的第一个入口点,Spotify的一个重要目标是支持用户探索有声读物目录。更具体地说,我们希望让没有特定项目的用户能够按主题、流派、故事比喻、年代进行广泛搜索,并发现他们可能喜欢的有声读物、作者和出版商。要做到这一点,我们需要1)激励用户输入更多的探索性查询有声读物和2)增强我们的检索系统,以更好地处理探索性有声读物查询。这在冷启动场景中是具有挑战性的,在冷启动场景中,由于与先前可用的项目(如音乐和播客内容)相比,用户与有声读物的交互量很少,因此我们存在可检索性偏差。为了解决这个问题,我们提出了AudioBoost,一个系统,以提高有声读物检索Spotify的搜索通过合成查询生成。AudioBoost利用大型语言模型(LLM)生成以有声读物元数据为条件的合成查询。合成查询在查询自动完成(QAC)和搜索检索引擎中都被索引,以同时改进查询制定和检索。我们通过离线评估表明,合成查询增加检索和高质量。此外,在线A/B测试的结果显示,AudioBoost在有声读物印象中获得了+0.7%,在有声读物点击中获得了+1.22%,在有声读物探索性查询完成中获得了+1.82%。
摘要:Spotify has recently introduced audiobooks as part of its catalog, complementing its music and podcast offering. Search is often the first entry point for users to access new items, and an important goal for Spotify is to support users in the exploration of the audiobook catalog. More specifically, we would like to enable users without a specific item in mind to broadly search by topic, genre, story tropes, decade, and discover audiobooks, authors and publishers they may like. To do this, we need to 1) inspire users to type more exploratory queries for audiobooks and 2) augment our retrieval systems to better deal with exploratory audiobook queries. This is challenging in a cold-start scenario, where we have a retrievabiliy bias due to the little amount of user interactions with audiobooks compared to previously available items such as music and podcast content. To address this, we propose AudioBoost, a system to boost audiobook retrievability in Spotify's Search via synthetic query generation. AudioBoost leverages Large Language Models (LLMs) to generate synthetic queries conditioned on audiobook metadata. The synthetic queries are indexed both in the Query AutoComplete (QAC) and in the Search Retrieval engine to improve query formulation and retrieval at the same time. We show through offline evaluation that synthetic queries increase retrievability and are of high quality. Moreover, results from an online A/B test show that AudioBoost leads to a +0.7% in audiobook impressions, +1.22% in audiobook clicks, and +1.82% in audiobook exploratory query completions.
【8】MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation
标题:通过一步生成的MeanFlow加速多模式视频到音频合成
链接:https://arxiv.org/abs/2509.06389
摘要:从无声视频合成音频的关键挑战是现有方法中合成质量和推理效率之间的固有权衡。例如,基于流匹配的模型依赖于对瞬时速度进行建模,固有地需要迭代采样过程,导致推理速度慢。为了解决这个效率瓶颈,我们引入了一个平均流加速模型,该模型使用平均速度表征流场,从而实现一步生成,从而显着加速多模式视频到音频(VTA)合成,同时保持音频质量,语义对齐和时间同步。此外,标量重新缩放机制,以平衡条件和无条件的预测时,无分类器的指导(CFG)的应用,有效地减轻CFG引起的失真在一步生成。由于音频合成网络是与多模态条件联合训练的,因此我们进一步在文本到音频(TTA)合成任务上对其进行评估。实验结果表明,将MeanFlow纳入网络显着提高推理速度,而不影响感知质量的VTA和TTA合成任务。
摘要:A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow matching based models rely on modeling instantaneous velocity, inherently require an iterative sampling process, leading to slow inference speeds. To address this efficiency bottleneck, we introduce a MeanFlow-accelerated model that characterizes flow fields using average velocity, enabling one-step generation and thereby significantly accelerating multimodal video-to-audio (VTA) synthesis while preserving audio quality, semantic alignment, and temporal synchronization. Furthermore, a scalar rescaling mechanism is employed to balance conditional and unconditional predictions when classifier-free guidance (CFG) is applied, effectively mitigating CFG-induced distortions in one step generation. Since the audio synthesis network is jointly trained with multimodal conditions, we further evaluate it on text-to-audio (TTA) synthesis task. Experimental results demonstrate that incorporating MeanFlow into the network significantly improves inference speed without compromising perceptual quality on both VTA and TTA synthesis tasks.
【9】DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
标题:DreamAudio:使用扩散模型的自定义文本到音频生成
链接:https://arxiv.org/abs/2509.06027
备注:Demos are available at this https URL
摘要:随着大规模基于扩散和基于语言建模的生成模型的发展,文本到音频生成已经取得了令人印象深刻的进展。尽管产生高质量的输出,现有的文本到音频模型主要旨在生成语义对齐的声音,并且在精确控制特定声音的细粒度声学特性方面存在不足。因此,需要特定声音内容的用户可能发现生成期望的音频剪辑具有挑战性。在本文中,我们提出了DreamAudio定制的文本到音频生成(CTTA)。具体来说,我们引入了一个新的框架,旨在使该模型识别听觉信息,从用户提供的参考概念的音频生成。给定一些包含个性化音频事件的参考音频样本,我们的系统可以生成包括这些特定事件的新音频样本。此外,开发了两种类型的数据集用于训练和测试定制系统。实验结果表明,该模型生成的音频样本与自定义音频特征高度一致,与输入文本提示对齐良好。此外,DreamAudio在一般的文本到音频任务中提供了相当的性能。我们还提供了一个人类参与的数据集,其中包含来自真实世界CTTA案例的音频事件,作为定制生成任务的基准。
摘要:With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite producing high-quality outputs, existing text-to-audio models mainly aim to generate semantically aligned sound and fall short on precisely controlling fine-grained acoustic characteristics of specific sounds. As a result, users that need specific sound content may find it challenging to generate the desired audio clips. In this paper, we present DreamAudio for customized text-to-audio generation (CTTA). Specifically, we introduce a new framework that is designed to enable the model to identify auditory information from user-provided reference concepts for audio generation. Given a few reference audio samples containing personalized audio events, our system can generate new audio samples that include these specific events. In addition, two types of datasets are developed for training and testing the customized systems. The experiments show that the proposed model, DreamAudio, generates audio samples that are highly consistent with the customized audio features and aligned well with the input text prompts. Furthermore, DreamAudio offers comparable performance in general text-to-audio tasks. We also provide a human-involved dataset containing audio events from real-world CTTA cases as the benchmark for customized generation tasks.
【10】Xi+: Uncertainty Supervision for Robust Speaker Embedding
标题:XI+:稳健的说话者嵌入的不确定性监督
链接:https://arxiv.org/abs/2509.05993
摘要:有各种因素会影响说话者识别系统的性能,例如情感、语言以及其他与说话者相关或与上下文相关的变化。由于单个语音帧对话语级表示的贡献不相等,因此必须估计每个帧的重要性或可靠性。xi向量模型通过基于不确定性估计向帧分配不同权重来解决这个问题。然而,其不确定性估计模型仅通过分类损失进行隐式训练,并且不考虑帧之间的时间关系,这可能导致次优监督。在本文中,我们提出了一个改进的架构,xi+。与xi-向量相比,xi+结合了时间注意力模块,以上下文感知的方式捕获帧级不确定性。此外,我们引入了一个新的损失函数,随机方差损失,它明确地监督学习的不确定性。结果表明,VoxCeleb 1-O集的性能提高了约10%,NIST SRE 2024评估集的性能提高了约11%。
摘要:There are various factors that can influence the performance of speaker recognition systems, such as emotion, language and other speaker-related or context-related variations. Since individual speech frames do not contribute equally to the utterance-level representation, it is essential to estimate the importance or reliability of each frame. The xi-vector model addresses this by assigning different weights to frames based on uncertainty estimation. However, its uncertainty estimation model is implicitly trained through classification loss alone and does not consider the temporal relationships between frames, which may lead to suboptimal supervision. In this paper, we propose an improved architecture, xi+. Compared to xi-vector, xi+ incorporates a temporal attention module to capture frame-level uncertainty in a context-aware manner. In addition, we introduce a novel loss function, Stochastic Variance Loss, which explicitly supervises the learning of uncertainty. Results demonstrate consistent performance improvements of about 10\% on the VoxCeleb1-O set and 11\% on the NIST SRE 2024 evaluation set.
【11】TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
标题:TSPC:一种以音素为中心的两级架构,用于代码切换越语-英语语音识别
链接:https://arxiv.org/abs/2509.05983
摘要:语码转换对一般的自动语音识别系统提出了重大挑战。现有的方法往往无法捕捉微妙的语音变化固有的CS场景。对于像越南语和英语这样的语言对来说,这一挑战尤其困难,因为它们既有明显的语音特征,又存在由相似的声音识别引起的歧义。在本文中,我们提出了一种新的架构,越南语-英语CS ASR,两阶段音素为中心的模型(TSPC)。TSPC采用以音素为中心的方法,建立在扩展的越南语音素集作为中间表示,以促进混合语言建模。实验结果表明,TSPC始终优于现有的基线,包括PhoWhisper-base,在越南语-英语CS ASR,实现了显着降低的单词错误率为20.8%,减少训练资源。此外,基于语音的两阶段架构使音素适应和语言转换,以提高在复杂的CS越南语-英语ASR场景中的ASR性能。
摘要:Code-switching (CS) presents a significant challenge for general Auto-Speech Recognition (ASR) systems. Existing methods often fail to capture the subtle phonological shifts inherent in CS scenarios. The challenge is particularly difficult for language pairs like Vietnamese and English, where both distinct phonological features and the ambiguity arising from similar sound recognition are present. In this paper, we propose a novel architecture for Vietnamese-English CS ASR, a Two-Stage Phoneme-Centric model (TSPC). The TSPC employs a phoneme-centric approach, built upon an extended Vietnamese phoneme set as an intermediate representation to facilitate mixed-lingual modeling. Experimental results demonstrate that TSPC consistently outperforms existing baselines, including PhoWhisper-base, in Vietnamese-English CS ASR, achieving a significantly lower word error rate of 20.8\% with reduced training resources. Furthermore, the phonetic-based two-stage architecture enables phoneme adaptation and language conversion to enhance ASR performance in complex CS Vietnamese-English ASR scenarios.
【12】Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
标题:通过净化语义相关联合建模增强上下文ASB对变化偏差信息的鲁棒性
链接:https://arxiv.org/abs/2509.05908
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing, 2025 (this https URL). DOI: https://doi.org/10.1109/TASLPRO.2025.3606198
摘要:最近,基于交叉注意的上下文自动语音识别(ASR)模型在识别个性化的偏见短语方面取得了显着的进步。然而,交叉注意的有效性受到偏见信息量变化的影响,特别是当偏见列表的长度显著增加时。我们发现,无论偏置列表的长度,只有有限数量的偏置信息是最相关的一个特定的ASR中间表示。因此,通过识别和整合最相关的偏置信息,而不是整个偏置列表,我们可以减轻上下文ASR的偏置信息量的变化的影响。为此,我们提出了一种纯化的语义相关联合建模(PSC-Joint)方法。在PSC-Joint中,我们定义并计算了ASR中间表示和偏置信息之间的三个语义相关性:从粗到细:列表级,短语级和令牌级。然后,这三个相关性被联合建模以产生它们的交集,使得跨各种粒度的最相关的偏置信息被突出显示并被集成用于上下文识别。此外,为了减少由三个语义相关性的联合建模引入的计算成本,我们还提出了一种基于分组和竞争策略的净化机制来过滤掉不相关的偏见短语。与基线相比,我们的PSC联合方法在不同长度的偏置列表中,AISHELL-1和KeSpeech上的平均相对F1分数分别提高了21.34%和28.46%。
摘要:Recently, cross-attention-based contextual automatic speech recognition (ASR) models have made notable advancements in recognizing personalized biasing phrases. However, the effectiveness of cross-attention is affected by variations in biasing information volume, especially when the length of the biasing list increases significantly. We find that, regardless of the length of the biasing list, only a limited amount of biasing information is most relevant to a specific ASR intermediate representation. Therefore, by identifying and integrating the most relevant biasing information rather than the entire biasing list, we can alleviate the effects of variations in biasing information volume for contextual ASR. To this end, we propose a purified semantic correlation joint modeling (PSC-Joint) approach. In PSC-Joint, we define and calculate three semantic correlations between the ASR intermediate representations and biasing information from coarse to fine: list-level, phrase-level, and token-level. Then, the three correlations are jointly modeled to produce their intersection, so that the most relevant biasing information across various granularities is highlighted and integrated for contextual recognition. In addition, to reduce the computational cost introduced by the joint modeling of three semantic correlations, we also propose a purification mechanism based on a grouped-and-competitive strategy to filter out irrelevant biasing phrases. Compared with baselines, our PSC-Joint approach achieves average relative F1 score improvements of up to 21.34% on AISHELL-1 and 28.46% on KeSpeech, across biasing lists of varying lengths.
【13】Yours or Mine? Overwriting Attacks against Neural Audio Watermarking
标题:你的还是我的?针对神经音频水印的覆盖攻击
链接:https://arxiv.org/abs/2509.05835
摘要:随着生成音频模型的快速发展,人工智能生成的音频越来越引起人们对侵犯版权和错误信息传播的担忧。音频水印作为一种主动防御技术,可以在音频中嵌入秘密信息,实现版权保护和来源验证。然而,目前的神经音频水印方法主要集中在水印的不可感知性和鲁棒性,而忽略了其对安全攻击的脆弱性。在本文中,我们开发了一个简单而强大的攻击:篡改攻击,覆盖合法的音频水印与伪造的,使原来的合法水印无法检测。基于攻击者所掌握的音频水印信息,我们提出了三种攻击方法,即,白盒、灰盒和黑盒攻击。我们还彻底评估了对最先进的神经音频水印方法的攻击。实验结果表明,所提出的水印攻击可以有效地破坏现有的水印方案在各种设置,并实现了近100%的攻击成功率。所提出的神经网络攻击的实用性和有效性暴露了现有神经音频水印系统的安全缺陷,强调了在未来的音频水印设计中需要提高安全性。
摘要:As generative audio models are rapidly evolving, AI-generated audios increasingly raise concerns about copyright infringement and misinformation spread. Audio watermarking, as a proactive defense, can embed secret messages into audio for copyright protection and source verification. However, current neural audio watermarking methods focus primarily on the imperceptibility and robustness of watermarking, while ignoring its vulnerability to security attacks. In this paper, we develop a simple yet powerful attack: the overwriting attack that overwrites the legitimate audio watermark with a forged one and makes the original legitimate watermark undetectable. Based on the audio watermarking information that the adversary has, we propose three categories of overwriting attacks, i.e., white-box, gray-box, and black-box attacks. We also thoroughly evaluate the proposed attacks on state-of-the-art neural audio watermarking methods. Experimental results demonstrate that the proposed overwriting attacks can effectively compromise existing watermarking schemes across various settings and achieve a nearly 100% attack success rate. The practicality and effectiveness of the proposed overwriting attacks expose security flaws in existing neural audio watermarking systems, underscoring the need to enhance security in future audio watermarking designs.
【14】Effectively obtaining acoustic, visual and textual data from videos
标题:从视频中有效获取声学、视觉和文本数据
链接:https://arxiv.org/abs/2509.05786
摘要:机器学习模型的使用越来越多,扩大了对高质量、大规模多模态数据集的需求。然而,这类数据集的可用性,特别是那些结合声学、视觉和文本数据的数据集,仍然有限。本文通过提出一种从视频中提取相关音频-图像-文本观察结果的方法来解决这一差距。我们详细介绍了选择合适的视频,提取相关数据对,并使用图像到文本模型生成描述性文本的过程。我们的方法确保了模态之间的强大语义连接,增强了所创建的数据集在各种应用程序中的实用性。我们还讨论了遇到的挑战,并提出了解决方案,以提高数据质量。由此产生的数据集是公开的,旨在支持和推进多模态数据分析和机器学习的研究。
摘要:The increasing use of machine learning models has amplified the demand for high-quality, large-scale multimodal datasets. However, the availability of such datasets, especially those combining acoustic, visual and textual data, remains limited. This paper addresses this gap by proposing a method to extract related audio-image-text observations from videos. We detail the process of selecting suitable videos, extracting relevant data pairs, and generating descriptive texts using image-to-text models. Our approach ensures a robust semantic connection between modalities, enhancing the utility of the created datasets for various applications. We also discuss the challenges encountered and propose solutions to improve data quality. The resulting datasets, publicly available, aim to support and advance research in multimodal data analysis and machine learning.
【15】Time-domain sound field estimation using kernel ridge regression
标题:基于核岭回归的时域声场估计
链接:https://arxiv.org/abs/2509.05720
摘要:基于核岭回归的声场估计方法已被证明是有效的,除了包括诸如声场的方向性的先验知识之外,还允许严格执行物理性质。这些方法是针对单频声场制定的,限制了可以使用的数据类型和先验知识。在本文中,核岭回归方法被推广到考虑离散时间声场。所提出的方法提供时域声场估计,其可以以封闭形式计算,保证是物理上可实现的,并且可以利用声场的时域特性来提高估计性能。利用房间脉冲响应的时域行为的先验信息,所提出的方法的估计性能被证明是使用时域数据加权,证明所提出的方法的有用性得到改善。它进一步示出使用模拟和真实数据,时域数据加权可以与方向加权相结合,利用空间和时间属性的房间脉冲响应的先验知识。所提出的方法的理论框架能够使用核岭回归来解决更广泛的声场估计问题,其中需要单独考虑每个频率的时域响应而不是频域响应。
摘要:Sound field estimation methods based on kernel ridge regression have proven effective, allowing for strict enforcement of physical properties, in addition to the inclusion of prior knowledge such as directionality of the sound field. These methods have been formulated for single-frequency sound fields, restricting the types of data and prior knowledge that can be used. In this paper, the kernel ridge regression approach is generalized to consider discrete-time sound fields. The proposed method provides time-domain sound field estimates that can be computed in closed form, are guaranteed to be physically realizable, and for which time-domain properties of the sound fields can be exploited to improve estimation performance. Exploiting prior information on the time-domain behaviour of room impulse responses, the estimation performance of the proposed method is shown to be improved using a time-domain data weighting, demonstrating the usefulness of the proposed approach. It is further shown using both simulated and real data that the time-domain data weighting can be combined with a directional weighting, exploiting prior knowledge of both spatial and temporal properties of the room impulse responses. The theoretical framework of the proposed method enables solving a broader class of sound field estimation problems using kernel ridge regression where it would be required to consider the time-domain response rather than the frequency-domain response of each frequency separately.
【16】On the Contribution of Lexical Features to Speech Emotion Recognition
标题:论词汇特征对语音情感识别的贡献
链接:https://arxiv.org/abs/2509.05634
备注:Accepted to 13th Conference on Speech Technology and Human-Computer Dialogue
摘要:虽然语言线索通常被认为是语音情感识别(SER)的主要驱动因素,我们调查的作用,从语音中提取的词汇内容,并表明它可以实现竞争力,在某些情况下,更高的性能相比,声学模型。在MELD数据集上,我们基于词汇的方法获得了51.5%的加权F1分数(WF1),而具有较大参数计数的仅声学管道的加权F1分数为49.3%。此外,我们分析了不同的自监督(SSL)语音和文本表示,进行基于变换器的编码器的逐层研究,并评估音频去噪的效果。
摘要:Although paralinguistic cues are often considered the primary drivers of speech emotion recognition (SER), we investigate the role of lexical content extracted from speech and show that it can achieve competitive and in some cases higher performance compared to acoustic models. On the MELD dataset, our lexical-based approach obtains a weighted F1-score (WF1) of 51.5%, compared to 49.3% for an acoustic-only pipeline with a larger parameter count. Furthermore, we analyze different self-supervised (SSL) speech and text representations, conduct a layer-wise study of transformer-based encoders, and evaluate the effect of audio denoising.
【1】Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
标题:集成空间和语义嵌入以实现视频中的立体声事件定位
链接:https://arxiv.org/abs/2509.06598
备注:arXiv admin note: substantial text overlap with arXiv:2507.04845
摘要:在这项研究中,我们解决的立体声声音事件定位和检测与源距离估计(3D SELD)在常规视频内容的多模态任务。3D SELD是一项复杂的任务,它将时间事件分类与空间定位相结合,需要跨空间,时间和语义维度进行推理。最后一个可以说是最具挑战性的模型。传统的SELD方法通常依赖于多通道输入,由于数据限制,限制了它们从大规模预训练中受益的能力。为了克服这一点,我们通过集成预训练的对比语言对齐模型来增强标准SELD架构的语义信息:CLAP用于音频,OWL-ViT用于视觉输入。这些嵌入被合并到一个修改后的Conformer模块中,该模块专为多模态融合而设计,我们称之为跨模态Conformer。我们对DCASE 2025 Task 3 Stereo SELD数据集的开发集进行了消融研究,以评估语言对齐模型和基准对DCASE任务3基线系统的单独贡献。此外,我们还详细介绍了用于模型预训练的大型合成音频和视听数据集的管理过程。这些数据集通过左右通道交换增强进一步扩展。我们的方法结合了广泛的预训练,模型集成和视觉后处理,在DCASE 2025挑战任务3(Track B)中获得了第二名,强调了我们方法的有效性。未来的工作将探索特定模式的贡献和架构的改进。
摘要:In this study, we address the multimodal task of stereo sound event localization and detection with source distance estimation (3D SELD) in regular video content. 3D SELD is a complex task that combines temporal event classification with spatial localization, requiring reasoning across spatial, temporal, and semantic dimensions. The last is arguably the most challenging to model. Traditional SELD approaches typically rely on multichannel input, limiting their capacity to benefit from large-scale pre-training due to data constraints. To overcome this, we enhance a standard SELD architecture with semantic information by integrating pre-trained, contrastive language-aligned models: CLAP for audio and OWL-ViT for visual inputs. These embeddings are incorporated into a modified Conformer module tailored for multimodal fusion, which we refer to as the Cross-Modal Conformer. We perform an ablation study on the development set of the DCASE2025 Task3 Stereo SELD Dataset to assess the individual contributions of the language-aligned models and benchmark against the DCASE Task 3 baseline systems. Additionally, we detail the curation process of large synthetic audio and audio-visual datasets used for model pre-training. These datasets were further expanded through left-right channel swapping augmentation. Our approach, combining extensive pre-training, model ensembling, and visual post-processing, achieved second rank in the DCASE 2025 Challenge Task 3 (Track B), underscoring the effectiveness of our method. Future work will explore the modality-specific contributions and architectural refinements.
【2】Speaker Privacy and Security in the Big Data Era: Protection and Defense against Deepfake
标题:演讲者大数据时代的隐私和安全:针对Deepfake的保护和防御
链接:https://arxiv.org/abs/2509.06361
摘要:在大数据时代,个性化语音生成技术取得了显着的进步,这些技术利用说话者的属性(包括语音和说话风格)来生成deepfake语音。这也加剧了deepfake语音滥用的全球安全风险,导致全球范围内的社会成本相当高。为了解决Deepfake语音带来的安全威胁,已经开发了专注于保护语音属性和防御Deepfake语音的技术。其中,语音匿名化技术已被开发用于保护语音属性不被提取以生成深度伪造,而深度伪造检测和水印已被用于防止深度伪造语音的滥用。本文提供了这三种技术的简短概述,描述了方法,进步和挑战。不久将出版一个全面的版本,提供更多的讨论。
摘要:In the era of big data, remarkable advancements have been achieved in personalized speech generation techniques that utilize speaker attributes, including voice and speaking style, to generate deepfake speech. This has also amplified global security risks from deepfake speech misuse, resulting in considerable societal costs worldwide. To address the security threats posed by deepfake speech, techniques have been developed focusing on both the protection of voice attributes and the defense against deepfake speech. Among them, the voice anonymization technique has been developed to protect voice attributes from extraction for deepfake generation, while deepfake detection and watermarking have been utilized to defend against the misuse of deepfake speech. This paper provides a short and concise overview of the three techniques, describing the methodologies, advancements, and challenges. A comprehensive version, offering additional discussions, will be published in the near future.
【3】Beamforming-LLM: What, Where and When Did I Miss?
标题:Beamforming-LLM:我错过了什么、在哪里以及何时?
链接:https://arxiv.org/abs/2509.06221
摘要:我们提出了波束成形LLM,一个系统,使用户能够语义回忆他们可能错过了多扬声器环境中的对话。该系统将使用麦克风阵列的空间音频捕获与检索增强生成(RAG)相结合,以支持自然语言查询,例如“当我跟踪狗的对话时,我错过了什么?定向音频流使用波束成形进行分离,用Whisper转录,并使用句子编码器嵌入到矢量数据库中。在接收到用户查询时,检索语义相关的片段,在时间上与无人值守的片段对齐,并使用轻量级大型语言模型(GPT-4 o-mini)进行总结。其结果是一个用户友好的界面,提供对比摘要,空间上下文和时间戳音频播放。这项工作为智能听觉记忆系统奠定了基础,并在辅助技术,会议摘要和上下文感知个人空间计算中具有广泛的应用。
摘要:We present Beamforming-LLM, a system that enables users to semantically recall conversations they may have missed in multi-speaker environments. The system combines spatial audio capture using a microphone array with retrieval-augmented generation (RAG) to support natural language queries such as, "What did I miss when I was following the conversation on dogs?" Directional audio streams are separated using beamforming, transcribed with Whisper, and embedded into a vector database using sentence encoders. Upon receiving a user query, semantically relevant segments are retrieved, temporally aligned with non-attended segments, and summarized using a lightweight large language model (GPT-4o-mini). The result is a user-friendly interface that provides contrastive summaries, spatial context, and timestamped audio playback. This work lays the foundation for intelligent auditory memory systems and has broad applications in assistive technology, meeting summarization, and context-aware personal spatial computing.
【4】From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model
标题:从感知到产生:声学不变性如何促进自我监督声乐模仿模型中的发音学习
链接:https://arxiv.org/abs/2509.05849
备注:Accepted at EMNLP 2025 (Main Conference)
摘要:人类婴儿在语音习得方面面临着一个艰巨的挑战:在没有明确指导的情况下将极其可变的声学输入映射到适当的发音运动。我们提出了一个计算模型,通过自我监督学习解决声学发音映射问题。我们的模型包括一个特征提取器,将语音转换为潜在的表示,一个逆模型,将这些表示映射到发音参数,和一个合成器,生成语音输出。在单扬声器和多扬声器设置中进行的实验表明,预训练的wav 2 vec 2.0模型的中间层为发音学习提供了最佳表示,显著优于MFCC特征。这些表示使我们的模型能够学习与人类模式相关的发音轨迹,区分发音位置,并产生可理解的语音。成功的发音学习的关键是平衡语音辨别力和说话人不变性的表征--这正是自我监督表征学习模型的特征。我们的研究结果提供了与发展理论相一致的计算证据,该理论提出语音类别的感知学习指导发音发展,尽管婴儿面临复杂的映射问题,但他们如何获得语音产生能力提供了见解。
摘要:Human infants face a formidable challenge in speech acquisition: mapping extremely variable acoustic inputs into appropriate articulatory movements without explicit instruction. We present a computational model that addresses the acoustic-to-articulatory mapping problem through self-supervised learning. Our model comprises a feature extractor that transforms speech into latent representations, an inverse model that maps these representations to articulatory parameters, and a synthesizer that generates speech outputs. Experiments conducted in both single- and multi-speaker settings reveal that intermediate layers of a pre-trained wav2vec 2.0 model provide optimal representations for articulatory learning, significantly outperforming MFCC features. These representations enable our model to learn articulatory trajectories that correlate with human patterns, discriminate between places of articulation, and produce intelligible speech. Critical to successful articulatory learning are representations that balance phonetic discriminability with speaker invariance -- precisely the characteristics of self-supervised representation learning models. Our findings provide computational evidence consistent with developmental theories proposing that perceptual learning of phonetic categories guides articulatory development, offering insights into how infants might acquire speech production capabilities despite the complex mapping problem they face.
【5】Time-domain sound field estimation using kernel ridge regression
标题:基于核岭回归的时域声场估计
链接:https://arxiv.org/abs/2509.05720
摘要:基于核岭回归的声场估计方法已被证明是有效的,除了包括诸如声场的方向性的先验知识之外,还允许严格执行物理性质。这些方法是针对单频声场制定的,限制了可以使用的数据类型和先验知识。在本文中,核岭回归方法被推广到考虑离散时间声场。所提出的方法提供时域声场估计,其可以以封闭形式计算,保证是物理上可实现的,并且可以利用声场的时域特性来提高估计性能。利用房间脉冲响应时域行为的先验信息,使用时域数据加权可以提高所提出方法的估计性能,证明了所提出方法的有用性。它进一步示出使用模拟和真实数据,时域数据加权可以与方向加权相结合,利用空间和时间属性的房间脉冲响应的先验知识。所提出的方法的理论框架使得能够使用核岭回归来解决更广泛的声场估计问题,其中需要单独考虑每个频率的时域响应而不是频域响应。
摘要:Sound field estimation methods based on kernel ridge regression have proven effective, allowing for strict enforcement of physical properties, in addition to the inclusion of prior knowledge such as directionality of the sound field. These methods have been formulated for single-frequency sound fields, restricting the types of data and prior knowledge that can be used. In this paper, the kernel ridge regression approach is generalized to consider discrete-time sound fields. The proposed method provides time-domain sound field estimates that can be computed in closed form, are guaranteed to be physically realizable, and for which time-domain properties of the sound fields can be exploited to improve estimation performance. Exploiting prior information on the time-domain behaviour of room impulse responses, the estimation performance of the proposed method is shown to be improved using a time-domain data weighting, demonstrating the usefulness of the proposed approach. It is further shown using both simulated and real data that the time-domain data weighting can be combined with a directional weighting, exploiting prior knowledge of both spatial and temporal properties of the room impulse responses. The theoretical framework of the proposed method enables solving a broader class of sound field estimation problems using kernel ridge regression where it would be required to consider the time-domain response rather than the frequency-domain response of each frequency separately.
【6】On the Contribution of Lexical Features to Speech Emotion Recognition
标题:论词汇特征对语音情感识别的贡献
链接:https://arxiv.org/abs/2509.05634
备注:Accepted to 13th Conference on Speech Technology and Human-Computer Dialogue
摘要:虽然语言线索通常被认为是语音情感识别(SER)的主要驱动因素,我们调查的作用,从语音中提取的词汇内容,并表明它可以实现竞争力,在某些情况下,更高的性能相比,声学模型。在MELD数据集上,我们基于词汇的方法获得了51.5%的加权F1分数(WF1),而具有较大参数计数的仅声学管道的加权F1分数为49.3%。此外,我们分析了不同的自监督(SSL)语音和文本表示,进行基于变换器的编码器的逐层研究,并评估音频去噪的效果。
摘要:Although paralinguistic cues are often considered the primary drivers of speech emotion recognition (SER), we investigate the role of lexical content extracted from speech and show that it can achieve competitive and in some cases higher performance compared to acoustic models. On the MELD dataset, our lexical-based approach obtains a weighted F1-score (WF1) of 51.5%, compared to 49.3% for an acoustic-only pipeline with a larger parameter count. Furthermore, we analyze different self-supervised (SSL) speech and text representations, conduct a layer-wise study of transformer-based encoders, and evaluate the effect of audio denoising.
【7】Graph Connectionist Temporal Classification for Phoneme Recognition
标题:用于音素识别的图连接论时态分类
链接:https://arxiv.org/abs/2509.05399
备注:Accepted to the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2025)
摘要:自动音素识别(APR)系统通常使用通过字形到音素(G2P)系统从文本生成的伪音素级注释来训练。这些G2P系统经常输出每个单词多个可能的发音,但标准的联结主义时间分类(CTC)损失不能解释训练过程中的这种模糊性。在这项工作中,我们适应图时间分类(GTC)的APR设置。GTC支持从替代音素序列图中进行训练,允许模型将每个单词的多个发音视为有效监督。我们在英语和荷兰语数据集上的实验表明,与CTC训练的基线相比,将每个单词的多个发音纳入训练损失中始终可以提高音素错误率。这些结果表明,将发音变化整合到损失函数中是一种很有前途的策略,用于从基于G2P的监督中训练APR系统。
摘要:Automatic Phoneme Recognition (APR) systems are often trained using pseudo phoneme-level annotations generated from text through Grapheme-to-Phoneme (G2P) systems. These G2P systems frequently output multiple possible pronunciations per word, but the standard Connectionist Temporal Classification (CTC) loss cannot account for such ambiguity during training. In this work, we adapt Graph Temporal Classification (GTC) to the APR setting. GTC enables training from a graph of alternative phoneme sequences, allowing the model to consider multiple pronunciations per word as valid supervision. Our experiments on English and Dutch data sets show that incorporating multiple pronunciations per word into the training loss consistently improves phoneme error rates compared to a baseline trained with CTC. These results suggest that integrating pronunciation variation into the loss function is a promising strategy for training APR systems from noisy G2P-based supervision.
【8】Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets
标题:使用MGPHot专家注释与通用标签数据集对音乐自动标记进行基准测试
链接:https://arxiv.org/abs/2509.06936
摘要:音乐自动标签旨在自动为音频记录分配描述性标签,例如流派,情绪或乐器。由于其挑战性、语义描述的多样性以及在各种应用中的实用价值,它已成为评估从音频数据中学习的通用音乐表示的性能的常见下游任务。我们介绍了一个新的基准数据集的基础上,最近公布的MGPHot数据集,其中包括专家音乐学注释,允许额外的见解和比较常见的通用标签数据集上获得的结果。虽然MGPHot注释已被证明对计算音乐学有用,但原始数据集既不包括音频,也不提供评估设置,以用作标准化的自动标记基准。为了解决这个问题,我们提供了一组精心策划的YouTube URL与可检索的音频,并提出了一个训练/验证/测试分裂标准化的评估,并预先计算表示为七个国家的最先进的模型。使用这些资源,我们在MGPHot和标准参考标签数据集中评估了这些模型,突出了专家和通用标签注释之间的关键差异。总之,我们的贡献提供了一个更先进的基准框架,为未来的研究在音乐理解。
摘要:Music autotagging aims to automatically assign descriptive tags, such as genre, mood, or instrumentation, to audio recordings. Due to its challenges, diversity of semantic descriptions, and practical value in various applications, it has become a common downstream task for evaluating the performance of general-purpose music representations learned from audio data. We introduce a new benchmarking dataset based on the recently published MGPHot dataset, which includes expert musicological annotations, allowing for additional insights and comparisons with results obtained on common generic tag datasets. While MGPHot annotations have been shown to be useful for computational musicology, the original dataset neither includes audio nor provides evaluation setups for its use as a standardized autotagging benchmark. To address this, we provide a curated set of YouTube URLs with retrievable audio, and propose a train/val/test split for standardized evaluation, and precomputed representations for seven state-of-the-art models. Using these resources, we evaluated these models in MGPHot and standard reference tag datasets, highlighting key differences between expert and generic tag annotations. Altogether, our contributions provide a more advanced benchmarking framework for future research in music understanding.
【9】Continuous Audio Language Models
标题:连续音频语言模型
链接:https://arxiv.org/abs/2509.06926
备注:17 pages, 3 figures
摘要:音频语言模型(ALM)通过将音频表示为离散令牌序列,已经成为语音和音乐生成的主导范式。然而,与可逆的文本令牌不同,音频令牌是从具有有限比特率的有损编解码器中提取的。因此,提高音频质量需要生成更多令牌,这需要在保真度和计算成本之间进行权衡。我们通过研究连续音频语言模型(CALM)来解决这个问题。这些模型实例化了一个大型的Transformer主干,它在每个时间步产生一个上下文嵌入。然后,该顺序信息调节MLP,该MLP通过一致性建模生成音频VAE的下一个连续帧。通过避免有损压缩,CALM以较低的计算成本实现了更高的质量。语音和音乐的实验表明,提高了效率和保真度的最先进的离散音频语言模型,促进轻量级,高质量的音频生成。样品可在https://continuous-audio-language-models.github.io上获得
摘要:Audio Language Models (ALM) have emerged as the dominant paradigm for speech and music generation by representing audio as sequences of discrete tokens. Yet, unlike text tokens, which are invertible, audio tokens are extracted from lossy codecs with a limited bitrate. As a consequence, increasing audio quality requires generating more tokens, which imposes a trade-off between fidelity and computational cost. We address this issue by studying Continuous Audio Language Models (CALM). These models instantiate a large Transformer backbone that produces a contextual embedding at every timestep. This sequential information then conditions an MLP that generates the next continuous frame of an audio VAE through consistency modeling. By avoiding lossy compression, CALM achieves higher quality at lower computational cost than their discrete counterpart. Experiments on speech and music demonstrate improved efficiency and fidelity over state-of-the-art discrete audio language models, facilitating lightweight, high-quality audio generation. Samples are available at https://continuous-audio-language-models.github.io
【10】DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
标题:DreamAudio:使用扩散模型的自定义文本到音频生成
链接:https://arxiv.org/abs/2509.06027
备注:Demos are available at this https URL
摘要:随着大规模基于扩散和基于语言建模的生成模型的发展,文本到音频生成方面取得了令人印象深刻的进展。尽管产生高质量的输出,现有的文本到音频模型主要旨在生成语义对齐的声音,并且在精确控制特定声音的细粒度声学特性方面存在不足。因此,需要特定声音内容的用户可能发现生成期望的音频剪辑具有挑战性。在本文中,我们提出了DreamAudio定制的文本到音频生成(CTTA)。具体来说,我们引入了一个新的框架,旨在使该模型识别听觉信息,从用户提供的参考概念的音频生成。给定一些包含个性化音频事件的参考音频样本,我们的系统可以生成包括这些特定事件的新音频样本。此外,开发了两种类型的数据集用于训练和测试定制系统。实验结果表明,该模型生成的音频样本与自定义音频特征高度一致,与输入文本提示对齐良好。此外,DreamAudio在一般的文本到音频任务中提供了相当的性能。我们还提供了一个人类参与的数据集,其中包含来自真实世界CTTA案例的音频事件,作为定制生成任务的基准。
摘要:With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite producing high-quality outputs, existing text-to-audio models mainly aim to generate semantically aligned sound and fall short on precisely controlling fine-grained acoustic characteristics of specific sounds. As a result, users that need specific sound content may find it challenging to generate the desired audio clips. In this paper, we present DreamAudio for customized text-to-audio generation (CTTA). Specifically, we introduce a new framework that is designed to enable the model to identify auditory information from user-provided reference concepts for audio generation. Given a few reference audio samples containing personalized audio events, our system can generate new audio samples that include these specific events. In addition, two types of datasets are developed for training and testing the customized systems. The experiments show that the proposed model, DreamAudio, generates audio samples that are highly consistent with the customized audio features and aligned well with the input text prompts. Furthermore, DreamAudio offers comparable performance in general text-to-audio tasks. We also provide a human-involved dataset containing audio events from real-world CTTA cases as the benchmark for customized generation tasks.
【11】Xi+: Uncertainty Supervision for Robust Speaker Embedding
标题:XI+:稳健的说话者嵌入的不确定性监督
链接:https://arxiv.org/abs/2509.05993
摘要:有各种因素会影响说话者识别系统的性能,例如情感、语言以及其他与说话者相关或与上下文相关的变化。由于单个语音帧对话语级表示的贡献不相等,因此必须估计每个帧的重要性或可靠性。xi向量模型通过基于不确定性估计向帧分配不同权重来解决这个问题。然而,其不确定性估计模型仅通过分类损失进行隐式训练,并且不考虑帧之间的时间关系,这可能导致次优监督。在本文中,我们提出了一个改进的架构,xi+。与xi-向量相比,xi+结合了时间注意力模块,以上下文感知的方式捕获帧级不确定性。此外,我们引入了一个新的损失函数,随机方差损失,它明确地监督学习的不确定性。结果表明,VoxCeleb 1-O集的性能提高了约10%,NIST SRE 2024评估集的性能提高了约11%。
摘要:There are various factors that can influence the performance of speaker recognition systems, such as emotion, language and other speaker-related or context-related variations. Since individual speech frames do not contribute equally to the utterance-level representation, it is essential to estimate the importance or reliability of each frame. The xi-vector model addresses this by assigning different weights to frames based on uncertainty estimation. However, its uncertainty estimation model is implicitly trained through classification loss alone and does not consider the temporal relationships between frames, which may lead to suboptimal supervision. In this paper, we propose an improved architecture, xi+. Compared to xi-vector, xi+ incorporates a temporal attention module to capture frame-level uncertainty in a context-aware manner. In addition, we introduce a novel loss function, Stochastic Variance Loss, which explicitly supervises the learning of uncertainty. Results demonstrate consistent performance improvements of about 10\% on the VoxCeleb1-O set and 11\% on the NIST SRE 2024 evaluation set.
【12】TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
标题:TSPC:一种以音素为中心的两级架构,用于代码切换越语-英语语音识别
链接:https://arxiv.org/abs/2509.05983
摘要:语码转换对一般的自动语音识别系统提出了重大挑战。现有的方法往往无法捕捉微妙的语音变化固有的CS场景。对于像越南语和英语这样的语言对来说,这一挑战尤其困难,因为它们既有明显的语音特征,又存在由相似的声音识别引起的歧义。在本文中,我们提出了一种新的架构,越南语-英语CS ASR,两阶段音素为中心的模型(TSPC)。TSPC采用以音素为中心的方法,建立在扩展的越南语音素集作为中间表示,以促进混合语言建模。实验结果表明,TSPC始终优于现有的基线,包括PhoWhisper-base,在越南语-英语CS ASR,实现了显着降低的单词错误率为20.8%,减少训练资源。此外,基于语音的两阶段架构使音素适应和语言转换,以提高在复杂的CS越南语-英语ASR场景中的ASR性能。
摘要:Code-switching (CS) presents a significant challenge for general Auto-Speech Recognition (ASR) systems. Existing methods often fail to capture the subtle phonological shifts inherent in CS scenarios. The challenge is particularly difficult for language pairs like Vietnamese and English, where both distinct phonological features and the ambiguity arising from similar sound recognition are present. In this paper, we propose a novel architecture for Vietnamese-English CS ASR, a Two-Stage Phoneme-Centric model (TSPC). The TSPC employs a phoneme-centric approach, built upon an extended Vietnamese phoneme set as an intermediate representation to facilitate mixed-lingual modeling. Experimental results demonstrate that TSPC consistently outperforms existing baselines, including PhoWhisper-base, in Vietnamese-English CS ASR, achieving a significantly lower word error rate of 20.8\% with reduced training resources. Furthermore, the phonetic-based two-stage architecture enables phoneme adaptation and language conversion to enhance ASR performance in complex CS Vietnamese-English ASR scenarios.
【13】Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
标题:通过净化语义相关联合建模增强上下文ASB对变化偏差信息的鲁棒性
链接:https://arxiv.org/abs/2509.05908
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing, 2025 (this https URL). DOI: https://doi.org/10.1109/TASLPRO.2025.3606198
摘要:最近,基于交叉注意的上下文自动语音识别(ASR)模型在识别个性化的偏见短语方面取得了显着的进步。然而,交叉注意的有效性受到偏见信息量变化的影响,特别是当偏见列表的长度显著增加时。我们发现,无论偏置列表的长度,只有有限数量的偏置信息是最相关的一个特定的ASR中间表示。因此,通过识别和整合最相关的偏置信息,而不是整个偏置列表,我们可以减轻上下文ASR的偏置信息量的变化的影响。为此,我们提出了一种纯化的语义相关联合建模(PSC-Joint)方法。在PSC-Joint中,我们定义并计算了ASR中间表示和偏置信息之间的三个语义相关性:从粗到细:列表级,短语级和令牌级。然后,这三个相关性被联合建模以产生它们的交集,使得跨各种粒度的最相关的偏置信息被突出显示并被集成用于上下文识别。此外,为了减少由三个语义相关性的联合建模引入的计算成本,我们还提出了一种基于分组和竞争策略的净化机制来过滤掉不相关的偏见短语。与基线相比,我们的PSC联合方法在不同长度的偏置列表中,AISHELL-1和KeSpeech上的平均相对F1分数分别提高了21.34%和28.46%。
摘要:Recently, cross-attention-based contextual automatic speech recognition (ASR) models have made notable advancements in recognizing personalized biasing phrases. However, the effectiveness of cross-attention is affected by variations in biasing information volume, especially when the length of the biasing list increases significantly. We find that, regardless of the length of the biasing list, only a limited amount of biasing information is most relevant to a specific ASR intermediate representation. Therefore, by identifying and integrating the most relevant biasing information rather than the entire biasing list, we can alleviate the effects of variations in biasing information volume for contextual ASR. To this end, we propose a purified semantic correlation joint modeling (PSC-Joint) approach. In PSC-Joint, we define and calculate three semantic correlations between the ASR intermediate representations and biasing information from coarse to fine: list-level, phrase-level, and token-level. Then, the three correlations are jointly modeled to produce their intersection, so that the most relevant biasing information across various granularities is highlighted and integrated for contextual recognition. In addition, to reduce the computational cost introduced by the joint modeling of three semantic correlations, we also propose a purification mechanism based on a grouped-and-competitive strategy to filter out irrelevant biasing phrases. Compared with baselines, our PSC-Joint approach achieves average relative F1 score improvements of up to 21.34% on AISHELL-1 and 28.46% on KeSpeech, across biasing lists of varying lengths.
【14】Yours or Mine? Overwriting Attacks against Neural Audio Watermarking
标题:你的还是我的?针对神经音频水印的覆盖攻击
链接:https://arxiv.org/abs/2509.05835
摘要:随着生成音频模型的快速发展,人工智能生成的音频越来越引起人们对侵犯版权和错误信息传播的担忧。音频水印作为一种主动防御技术,可以在音频中嵌入秘密信息,实现版权保护和来源验证。然而,目前的神经音频水印方法主要集中在水印的不可感知性和鲁棒性,而忽略了其对安全攻击的脆弱性。在本文中,我们开发了一个简单而强大的攻击:篡改攻击,覆盖合法的音频水印与伪造的,使原来的合法水印无法检测。基于攻击者所掌握的音频水印信息,我们提出了三种攻击方法,即,白盒、灰盒和黑盒攻击。我们还彻底评估了对最先进的神经音频水印方法的攻击。实验结果表明,所提出的水印攻击可以有效地破坏现有的水印方案在各种设置,并实现了近100%的攻击成功率。所提出的神经网络攻击的实用性和有效性暴露了现有神经音频水印系统的安全缺陷,强调了在未来的音频水印设计中需要提高安全性。
摘要:As generative audio models are rapidly evolving, AI-generated audios increasingly raise concerns about copyright infringement and misinformation spread. Audio watermarking, as a proactive defense, can embed secret messages into audio for copyright protection and source verification. However, current neural audio watermarking methods focus primarily on the imperceptibility and robustness of watermarking, while ignoring its vulnerability to security attacks. In this paper, we develop a simple yet powerful attack: the overwriting attack that overwrites the legitimate audio watermark with a forged one and makes the original legitimate watermark undetectable. Based on the audio watermarking information that the adversary has, we propose three categories of overwriting attacks, i.e., white-box, gray-box, and black-box attacks. We also thoroughly evaluate the proposed attacks on state-of-the-art neural audio watermarking methods. Experimental results demonstrate that the proposed overwriting attacks can effectively compromise existing watermarking schemes across various settings and achieve a nearly 100% attack success rate. The practicality and effectiveness of the proposed overwriting attacks expose security flaws in existing neural audio watermarking systems, underscoring the need to enhance security in future audio watermarking designs.
【15】Effectively obtaining acoustic, visual and textual data from videos
标题:从视频中有效获取声学、视觉和文本数据
链接:https://arxiv.org/abs/2509.05786
摘要:机器学习模型的使用越来越多,扩大了对高质量、大规模多模态数据集的需求。然而,这类数据集的可用性,特别是那些结合声学、视觉和文本数据的数据集,仍然有限。本文通过提出一种从视频中提取相关音频-图像-文本观察结果的方法来解决这一差距。我们详细介绍了选择合适的视频,提取相关数据对,并使用图像到文本模型生成描述性文本的过程。我们的方法确保了模态之间的强大语义连接,增强了所创建的数据集在各种应用程序中的实用性。我们还讨论了遇到的挑战,并提出了解决方案,以提高数据质量。由此产生的数据集是公开的,旨在支持和推进多模态数据分析和机器学习的研究。
摘要:The increasing use of machine learning models has amplified the demand for high-quality, large-scale multimodal datasets. However, the availability of such datasets, especially those combining acoustic, visual and textual data, remains limited. This paper addresses this gap by proposing a method to extract related audio-image-text observations from videos. We detail the process of selecting suitable videos, extracting relevant data pairs, and generating descriptive texts using image-to-text models. Our approach ensures a robust semantic connection between modalities, enhancing the utility of the created datasets for various applications. We also discuss the challenges encountered and propose solutions to improve data quality. The resulting datasets, publicly available, aim to support and advance research in multimodal data analysis and machine learning.
【16】An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
标题:语音语言建模预训练中离散单元表征的实证分析
链接:https://arxiv.org/abs/2509.05359
备注:Published in International Conference on Text, Speech, and Dialogue, 13-24
摘要:本文研究了语音语言模型(SLM)中的离散单元表示,重点是在连续预训练过程中优化语音建模。在本文中,我们系统地研究了模型架构,数据表示和训练鲁棒性如何影响预训练阶段,在此阶段,我们将现有的预训练语言模型适应语音模态。我们的实验突出了语音编码器和聚类粒度在不同模型尺度上的作用,显示了最佳离散化策略如何随模型容量而变化。通过研究集群分布和音素对齐,我们调查离散词汇的有效使用,揭示语言和非语言模式。此外,我们还探讨了聚类数据选择对模型鲁棒性的影响,强调了离散化训练和目标应用程序之间域匹配的重要性。
摘要:This paper investigates discrete unit representations in Speech Language Models (SLMs), focusing on optimizing speech modeling during continual pre-training. In this paper, we systematically examine how model architecture, data representation, and training robustness influence the pre-training stage in which we adapt existing pre-trained language models to the speech modality. Our experiments highlight the role of speech encoders and clustering granularity across different model scales, showing how optimal discretization strategies vary with model capacity. By examining cluster distribution and phonemic alignments, we investigate the effective use of discrete vocabulary, uncovering both linguistic and paralinguistic patterns. Additionally, we explore the impact of clustering data selection on model robustness, highlighting the importance of domain matching between discretization training and target applications.
机器翻译由腾讯交互翻译提供,仅供参考
