今日论文合集:cs.SD语音16篇,eess.AS音频处理14篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception
标题:UniAudio-Token:通过通用音频感知增强语义语音令牌器
链接:https://arxiv.org/abs/2605.31521
作者:Yuhan Song,Linhao Zhang,Aiwei Liu,Chuhan Wu,Sijun Zhang,Wei Jia,Yuan Liu,Houfeng Wang,Xiao Zhou
备注:19 pages, 10 figures
摘要:语义语音标记器由于其紧凑的单码本设计和强大的语言对齐,已成为Audio-LLM的广泛使用的接口。然而,他们专注于语言抽象导致声盲,限制了他们的适用性以外的语音为中心的任务。我们提出了UniAudio-Token,这是一个框架,它使语义标记器具有一般的音频感知,而不影响语音能力。UniAudio-Token没有改变语义范式,而是通过两个关键创新来减轻其信息丢失:(1)语义声学原语(SAP)通过将音频分解为语言内容,声音属性和语音场景原语来提供结构化监督;(2)语义声学均衡(SAE)引入了一种内容感知的门控机制,自适应地从浅层恢复细粒度的声学细节。广泛的评估表明,UniAudio-Token学习全面的通用表示,同时保持高保真的语音生成。当与下游LLM集成时,它在理解和生成任务方面优于所有单码本基线标记器,有效地充当统一的音频接口。我们在https://github.com/Tencent/Universal_Audio_Tokenizer上公开发布了所有代码,包括训练和推理脚本以及模型检查点。
摘要:Semantic speech tokenizers have become a widely used interface for Audio-LLMs, owing to their compact single-codebook design and strong linguistic alignment. However, their focus on linguistic abstraction induces acoustic blindness, limiting their applicability beyond speech-centric tasks. We propose UniAudio-Token, a framework that empowers semantic tokenizers with general audio perception without compromising speech ability. Instead of altering the semantic paradigm, UniAudio-Token mitigates its information loss through two key innovations: (1) Semantic-Acoustic Primitives (SAP) provide structured supervision by decomposing audio into linguistic content, vocal attributes, and auditory-scene primitives; and (2) Semantic-Acoustic Equilibrium (SAE) introduces a content-aware gating mechanism that adaptively restores fine-grained acoustic details from shallow layers. Extensive evaluations show that UniAudio-Token learns comprehensive universal representations while preserving high-fidelity speech generation. When integrated with downstream LLMs, it outperforms all single-codebook baseline tokenizers on both understanding and generation tasks, effectively serving as a unified audio interface. We publicly release all our code, including training and inference scripts, together with the model checkpoints at https://github.com/Tencent/Universal_Audio_Tokenizer.


【2】Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus

标题:缩放对话匈牙利SVR:BEA-Dialogue+ Corpus
链接:https://arxiv.org/abs/2605.31469
作者:Máté Gedeon,Piroska Zsófia Barta,Péter Mihajlik,Katalin Mády
摘要:匈牙利语的对话式自动语音识别受到公开可用的对话式训练数据数量有限的限制。BEA-Dialogue语料库满足了这一需求,但其严格的扬声器分离训练/开发/评估分离将可用材料减少到只有85小时。在本文中,我们介绍BEA-Dialogue+,一个扩展版本的语料库,放宽了实验者和对话伙伴的分裂标准,同时保持完全分离的主要发言人。这导致了200小时的转录自然对话,并使额外的训练数据和跨分裂的说话者重叠之间的权衡的受控研究成为可能。我们在两个语料库版本上评估了几个基于Whisper和FastConformer的模型,包括基于序列化输出训练(SOT)的对话转录微调。我们的研究结果表明,较大的语料库是更具有挑战性的模型没有微调,而基于SOT的适应产生一致的改善WER,CER,cpWER和cpCER。总体而言,BEA-Dialogue+为匈牙利对话ASR提供了一个更大但仍然要求严格的基准,以及用于培训和评估对话转录系统的实用资源。
摘要:Conversational automatic speech recognition in Hungarian is constrained by the limited amount of publicly available dialogue-style training data. The BEA-Dialogue corpus addresses this need, but its strictly speaker-disjoint train/dev/eval split reduces the usable material to only 85 hours. In this paper, we introduce BEA-Dialogue+, an expanded version of the corpus that relaxes the split criterion for experimenters and dialogue partners while preserving complete separation of the primary speakers. This results in 200 hours of transcribed natural conversations and enables a controlled study of the trade-off between additional training data and speaker overlap across the splits. We evaluate several Whisper- and FastConformer-based models on both corpus versions, including Serialized Output Training (SOT)-based fine-tuning for dialogue transcription. Our results show that the larger corpus is more challenging for models without fine-tuning, whereas SOT-based adaptation yields consistent improvements in WER, CER, cpWER, and cpCER. Overall, BEA-Dialogue+ provides a substantially larger yet still demanding benchmark for Hungarian dialogue ASR, and a practical resource for training and evaluating dialogue transcription systems.


【3】DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs

标题:DOE:具有SpeechLLM的长式同声翻译的免训练解码器关注政策
链接:https://arxiv.org/abs/2605.31432
作者:Sara Papi,Luisa Bentivogli
摘要:同步语音到文本翻译(SimulST)在语音仍在展开时生成翻译,需要一个流策略来决定何时读取和何时写入。现有技术的方法依赖于基于注意力的编码器-解码器模型,其中交叉注意力提供明确的对准信号。相比之下,语音大语言模型(SpeechLLM)是仅依赖于自注意力的解码器架构。这就提出了一个核心问题:解码器的自我关注是否包含足够稳定的对齐信号来指导流媒体策略。此外,现有的方法通常依赖于基于训练的适应或启发式等待$k$政策,并没有在长期的设置验证。为了填补这些空白,我们提出了仅解码器注意力(DOA),这是一种无需训练的策略,通过从自我注意力中获得代理对齐,可以使用现成的SpeechLLM实现长格式的同声翻译。在Phi 4-Multimodal和Qwen 3-Omni上的实验表明,DOA提供了一个有效的对齐信号,用于支持流决策,使低延迟的长格式SimulST具有接近离线解码的质量,而无需重新训练。
摘要:Simultaneous speech-to-text translation (SimulST) generates translations while speech is still unfolding, requiring a streaming policy that decides when to read and when to write. State-of-the-art approaches rely on attention-based encoder-decoder models where cross-attention provides explicit alignment signals. In contrast, Speech Large Language Models (SpeechLLMs) are decoder-only architectures relying solely on self-attention. This raises a central question: whether decoder self-attention contains sufficiently stable alignment signals to guide the streaming policy. Moreover, existing approaches typically rely on training-based adaptations or heuristic wait-$k$ policies and have not been validated in long-form settings. To fill these gaps, we propose Decoder-Only Attention (DOA), a training-free policy that enables long-form simultaneous translation with off-the-shelf SpeechLLMs by deriving a proxy alignment from self-attention. Experiments on Phi4-Multimodal and Qwen3-Omni show that DOA provides an effective alignment signal for supporting streaming decisions, enabling low-latency long-form SimulST with quality close to offline decoding without retraining.


【4】Latent Space Disentanglement via Activation Steering for Interpretable Attribute Control in Symbolic Music Generation

标题:通过激活引导实现潜在空间解纠缠以实现符号音乐生成中的可解释属性控制
链接:https://arxiv.org/abs/2605.31295
作者:Ioannis Prokopiou,Pantelis Vikatos,Maximos Kaliakatsos-Papakostas,Theodoros Giannakopoulos,Themos Stafylakis
备注:Accepted at EUSIPCO 2026 (34th European Signal Processing Conference), 5 pages, 2 figures
摘要:基于变换器的架构已经显著地推进了复杂符号序列的生成,但是在实现对离散信号属性的细粒度、可解释的控制方面仍然存在显著的差距。本文研究了多轨音乐Transformer(MMT)的机械可解释性,并提出了一个框架,确定性属性调制没有重新训练,以弥合这一差距,通过推理时间激活转向。利用差分均值(DiffMean)方法,我们隔离信号属性的潜在方向,特别是音高和持续时间,在残留流。我们验证了线性表示假设在这一领域,实现高相关性转向幅度和属性偏移。为了解决多属性转向中固有的特征纠缠,我们引入了一个利用Gram-Schmidt归一化的双转向框架。实验结果表明,这种几何解耦减少了概念干扰和信号退化相比,天真的矢量加法,使独立的确定性控制,即使对强自回归条件。
摘要:Transformer-based architectures have significantly advanced the generation of complex symbolic sequences, yet a significant gap remains in achieving fine-grained, interpretable control over discrete signal attributes. This paper investigates the mechanistic interpretability of the Multitrack Music Transformer (MMT) and proposes a framework for deterministic attribute modulation without retraining to bridge this gap via inference-time activation steering. Utilizing the Difference-in-Means (DiffMean) methodology, we isolate latent directions for signal attributes, specifically Pitch and Duration, within the residual stream. We validate the Linear Representation Hypothesis in this domain, achieving high correlation between steering magnitude and attribute shift. To address the inherent feature entanglement in multi-attribute steering, we introduce a Dual Steering framework utilizing Gram-Schmidt Orthogonalization. Experimental results demonstrate that this geometric decoupling reduces conceptual interference and signal degradation compared to naive vector addition, enabling independent deterministic control even against strong autoregressive conditioning.


【5】MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors

标题:MindVoice:利用预先训练的先验从无创神经信号重建可理解的语音
链接:https://arxiv.org/abs/2605.31173
作者:Guangyin Bao,Taiping Zeng,Jianfeng Feng,Xiangyang Xue
摘要:从非侵入性神经记录中重建连续语音是探索人类听觉感知和构建安全、可扩展的语音脑机接口的基本问题。尽管最近取得了进展,但可理解的重建仍然是难以捉摸的,因为非侵入性记录本身就有噪声,空间模糊,并且只能部分保留有关感知语音的信息。现有的方法在用神经声码器合成波形之前直接将神经活动映射到纠缠的语音表示,导致频谱相似但难以理解的结果。为了克服这些限制,我们引入了MindVoice,这是一个神经到语音重建框架,它使用预训练的模型来补偿神经记录中不完整的语义和声学信息。MindVoice将重建分解为两个互补的途径:一个恢复高级语义内容,而另一个估计细粒度的声学属性。然后,这些推断的表示与强大的语音生成模型和上下文语音克隆相融合,以合成自然和可理解的话语。对EEG和MEG的大量实验表明,MindVoice在各种指标上都大大优于现有方法。这些结果表明,预先训练的先验提供了一种原则性的方法来弥合嘈杂的神经记录和自然语音之间的差距,突出了听觉神经科学研究和非侵入性语音脑机接口的一个有前途的尝试。
摘要:Reconstructing continuous speech from non-invasive neural recordings is a fundamental problem for probing human auditory perception and building safe, scalable speech brain-computer interfaces. Despite recent progress, intelligible reconstruction remains elusive, as non-invasive recordings are inherently noisy, spatially blurred, and only partially preserve information about perceived speech. Existing methods directly map neural activity to entangled speech representations before synthesizing waveforms with neural vocoders, resulting in spectral-similar but unintelligible results. To overcome these limitations, we introduce MindVoice, a neuro-to-speech reconstruction framework that uses pretrained models to compensate for the incomplete semantic and acoustic information in neural recordings. MindVoice disentangles reconstruction into two complementary pathways: one recovers high-level semantic content, while the other estimates fine-grained acoustic attributes. These inferred representations are then fused with powerful speech generation models and in-context voice cloning to synthesize natural and intelligible utterances. Extensive experiments on EEG and MEG demonstrate that MindVoice substantially outperforms existing methods on various metrics. These results show that pretrained priors provide a principled way to bridge the gap between noisy neural recordings and natural speech, highlighting a promising attempt for auditory neuroscience research and non-invasive speech brain-computer interfaces.


【6】Sound effects in media:A comparative analysis of recorded and synthetic samples in live-action and animation

标题:媒体中的声音效果:对真人和动画中录制和合成样本的比较分析
链接:https://arxiv.org/abs/2605.31082
作者:Nelly Garcia,Joshua Reiss
备注:ArtsIT, Interactivity and Game Creation 2024
摘要:为讲故事创造声音对于在电影、电视剧和视频游戏等制作中建立环境至关重要。这个过程通常涉及重复,分层和记录真实对象或使用声音库,这可能是耗时和重复的。为了应对这些挑战,程序音频(也称为数字音频)提供了一种解决方案,允许声音设计师快速生成样本。尽管它的效率,问题仍然是关于合成样本的可信度相比,真正的。在我们的研究中,我们比较了由在线程序引擎生成的合成样本,并将它们与动画和真人视觉效果相结合。我们的结果表明,程序音频非常有效,并且在戏剧和科幻场景中被认为是可信的,特别是对于激光、打击、空气和火箭等声音模型,而合成声音在卡通制作中代表日常动作时并不那么可信。最后,我们确定了需要优化的特定模型,并根据音频专业人士的反馈强调了需要改进的音频功能。
摘要:Creating sound for storytelling is crucial to establishing the environment in productions such as films, TV series and video games. This process often involves repeating, layering and recording real objects or using sound libraries, which can be time-consuming and repetitive. To address these challenges, procedural audio, also known as digital foley, offers a solution by allowing sound designers to quickly generate samples. Despite its efficiency, questions remain about the believability of synthetic samples compared to real ones. In our study, we compared synthetic samples generated by an online procedural engine and integrated them with both animated and live-action visuals. Our results indicate that procedural audio is highly effective and perceived as believable in drama and sci-fi scenes, particularly for sound models such as lasers, hits, air and rockets, whereas synthetic sounds weren't as believable in cartoon productions when representing everyday actions. Finally, we identified specific models that needed optimisation and highlighted audio features that needed improvement with feedback from audio professionals.


【7】AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing

标题:AnchorSteer:自我发现的概念注入,用于保留结构的音乐编辑
链接:https://arxiv.org/abs/2605.31053
作者:Chih-Heng Chang,Keng-Seng Ho,Chih-Yu Tsai,Kuan-Lin Chen,Yi-Hsuan Yang,Jian-Jiun Ding
备注:Accepted by the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026)
摘要:可控音乐编辑是在严格保留节奏和旋律结构的同时修改高级属性。然而,这一任务的语义结构纠缠的挑战:转向方法往往会降低结构,以实现编辑性能,而结构适配器抑制语义响应。我们提出了锚转向,一个框架,解开这种紧张局势,耦合结构锚定与自我发现的语义转向。所提出的方法探测内部表示,通过自监督重建目标提取可解释的,无标签的概念向量,隔离属性,而无需策划数据。在编辑过程中,这些便携式,即插即用的概念向量被注入到扩散隐藏流形,而结构适配器强制执行一致性。提供了无条件和有条件注入的变体,以平衡鲁棒性和语义强度。ZoME-Bench和主观测试上的实验表明,该框架优于仅转向和仅锚定基线,实现了具有高保真结构保留的显著语义转换。
摘要:Controllable music editing is to modify high-level attributes while strictly preserving rhythmic and melodic structures. However, this task is challenged by a semantic-structural entanglement: steering methods often degrade structure to achieve editing performance, while structural adaptors suppress semantic responsiveness. We propose AnchorSteer, a framework that disentangles this tension by coupling structural anchoring with self-discovered semantic steering. The proposed approach probes internal representations to extract interpretable, label-free concept vectors via a self-supervised reconstruction objective, isolating attributes without curated data. During editing, these portable, plug-and-play concept vectors are injected into diffusion hidden manifolds while a structural adaptor enforces consistency. Variants for unconditioned and conditioned injections are provided to balance robustness and semantic strength. Experiments on ZoME-Bench and subjective tests show that the proposed framework outperforms both steering-only and anchoring-only baselines, enabling significant semantic transformations with high-fidelity structural preservation.


【8】GaMi: Geometry-Agnostic Material Identification via Cross-Modal Subtractive Disentanglement

标题:GaMi:通过跨模减法解纠缠进行几何不可知材料识别
链接:https://arxiv.org/abs/2605.30818
作者:Zhiwei Chen,Yijie Li,Yimo Zhang,Shiyun Shao,Yichao Chen,Dian Ding,Liang Wang,Haiwei Wu,Liwei Guo,Jie Yang,Xiaosong Zhang,Yongzhao Zhang
备注:17 pages, 18 figures
摘要:非接触式材料识别使得能够实现体现智能的自适应交互,但面临来自几何形状引起的变化的挑战(例如,方向、形状、距离)和单模态模糊性。在本文中,我们提出了GaMi,一个多模态材料识别系统集成毫米波和声学传感鲁棒性不受约束的几何条件下运行。通过利用共同定位的双峰传感器之间的共享几何一致性的洞察力,GaMi采用了样本内交叉模态减法解纠缠框架。通过语义对齐模态和减去共享的几何上下文,它隔离了内在的材料特征。此外,GaMi结合样本间对比学习来校正由交叉模态不对准引起的残余干扰。此外,两种模态之间基于配对的自适应策略实现了跨设备的Few-Shot泛化。对20种材料的广泛评估表明,GaMi达到了95.2%的准确率,在看不见的几何条件下优于单模态基线。
摘要:Non-contact material identification enables adaptive interaction for embodied intelligence yet faces challenges from geometry-induced variations (e.g., orientation, shape, distance) and single-modality ambiguities. In this paper, we present GaMi, a multimodal material identification system integrating mmWave and acoustic sensing to robustly operate under unconstrained geometric conditions. By leveraging the insight of shared geometric consistency between co-located bimodal sensors, GaMi employs an intra-sample cross-modal subtractive disentanglement framework. By semantically aligning modalities and subtracting the shared geometric context, it isolates intrinsic material features. Furthermore, GaMi incorporates inter-sample contrastive learning to correct the residual interference caused by cross-modal misalignment. Additionally, a pairing-based adaptation strategy between two modalities enables few-shot generalization across devices. Extensive evaluations on 20 materials show that GaMi achieves 95.2% accuracy, outperforming single-modality baselines across unseen geometric conditions.


【9】Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

标题:Chatterbox-Flash:用于流媒体Zero-ShotTTC的预先校准块扩散
链接:https://arxiv.org/abs/2605.30748
作者:Deokjin Seo,Gangin Park,Kihyun Nam
备注:8 pages, 4 figures, 9 tables
摘要:我们提出了Chatterbox-Flash,一个zero-shot文本到语音模型,通过将预训练的自回归TTS解码器微调为块扩散解码器,使每个块内的并行令牌生成,同时保留逐块流。我们发现,天真地将主流块扩散解码离散语音令牌降低质量,作为一个长尾令牌分布偏向于几个高频令牌的并行位置选择。为了在不修改架构的情况下减轻这种情况,我们引入了两种推理时间技术:事先校准的评分,它减去块级边缘令牌分布,和早期解码时间表,它自适应地终止迭代校准的信心的基础上。在标准的zero-shot TTS基准测试中,Chatterbox-Flash实现了与强自回归和非自回归基线相当的高保真合成,同时支持与流式AR系统相当的第一个数据包时间的流式推理,并大大降低了实时因素。代码和音频示例可在https://github.com/resemble-ai/chatterbox-flash上获得。
摘要:We present Chatterbox-Flash, a zero-shot text-to-speech model obtained by fine-tuning a pretrained autoregressive TTS decoder into a block-diffusion decoder, enabling parallel token generation within each block while retaining block-by-block streaming. We find that naively transferring mainstream block-diffusion decoding to discrete speech tokens degrades quality, as a long-tail token distribution biases parallel position selection toward a few high-frequency tokens. To mitigate this without architectural modification, we introduce two inference-time techniques: prior-calibrated scoring, which subtracts the block-level marginal token distribution, and an early-decoding schedule, which adaptively terminates iteration based on calibrated confidence. On standard zero-shot TTS benchmarks, Chatterbox-Flash attains high-fidelity synthesis comparable to strong autoregressive and non-autoregressive baselines, while supporting streaming inference with time-to-first-packet on par with streaming AR systems and substantially lower real-time factor. Code and audio samples are available at https://github.com/resemble-ai/chatterbox-flash.


【10】Audio Pirates: Black-box Audio Watermark Removal via Diffusion Priors

标题:音频海盗:通过扩散先验去除黑匣子音频水印
链接:https://arxiv.org/abs/2605.30614
作者:Lingfeng Yao,Xincong Zhong,Chenpei Huang,Xuandong Zhao,Hanqing Guo,Aohan Li,Jiang Liu,Tomoaki Ohtsuki,Miao Pan
摘要:随着人工智能生成的音频的兴起,水印已被广泛用于检测滥用和保护知识产权。然而,攻击者可能会试图删除这些水印,这使得评估水印方案如何抵御删除攻击变得至关重要。现有的攻击通常是不切实际的:它们要么明显降低感知质量,要么需要访问水印方案。我们提出了Difficult,黑盒水印去除攻击,假设没有知识的目标水印方案,同时保持感知质量。Diffusible将水印音频扰动到中间扩散噪声水平,并使用预训练的去噪模型重新生成它,有效地抑制水印信号。理论分析和大量的实验表明,听不见的音频水印是非常脆弱的:在多个音频域,Difficult一致地删除水印,同时保持感知质量。这些发现强调了未来音频水印设计需要考虑基于扩散的威胁。代码和演示可在https://differase.github.io/DiffErase/上获得。
摘要:With the rise of AI-generated audio, watermarking has become widely used for detecting misuse and protecting intellectual property. However, adversaries may try to remove these watermarks, making it critical to evaluate how well watermarking schemes withstand removal attacks. Existing attacks are often impractical: they either noticeably degrade perceptual quality or require access to the watermarking scheme. We propose DiffErase, a black-box watermark removal attack that assumes no knowledge of the target watermarking scheme while maintaining perceptual quality. DiffErase perturbs watermarked audio to an intermediate diffusion noise level and regenerates it using a pretrained denoising model, effectively suppressing watermark signals. Theoretical analysis and extensive experiments demonstrate that inaudible audio watermarks are highly vulnerable: across multiple audio domains, DiffErase consistently removes watermarks while preserving perceptual quality. These findings highlight the need for future audio watermarking designs to consider diffusion-based threats. Code and demos are available at https://differase.github.io/DiffErase/.


【11】3DAE: Binaural Quality Assessment for Audio Novel View Synthesis with Spatial Maps and Benchmark

标题:3DTE:利用空间地图和基准进行音频新颖视图合成的双耳质量评估
链接:https://arxiv.org/abs/2605.30469
作者:Jialu Xu,Yifan Zhou
摘要:3D音频和新视角声学合成模型通常采用全局度量进行评估,但全局度量往往隐藏了双耳预测失败的位置和原因。我们提出了一个完整的参考诊断框架,使用时频音频误差图的幅度,ILD,IPD,时间对齐,响度和高频故障,形成一个三维音频误差图(3DAE地图)的视觉检查。我们将这些诊断框成一个模型不可知的基准,空间音频误差基准(3DAE Bench),它采用任意的地面实况和预测的双耳对,并报告音频新视图合成模型的预测质量。Replay-NVAS和SoundSpaces上的ViGAS输出实验显示了不同的主要故障模式:Replay-NVAS上的时间不对准和SoundSpaces上的ILD不匹配。总体而言,该框架提供了可解释的故障模式摘要和直观的视觉地图,用于音频新视图合成模型开发优化。
摘要:3D audio and novel-view acoustic synthesis models are usually evaluated with global metrics.However, global metrics often hide where and why binaural prediction fails. We propose a full-reference diagnostic framework that uses time-frequency audio error maps for magnitude, ILD, IPD, temporal alignment, loudness, and high-frequency failures, forming a 3D Audio Error Map (3DAE Map) for visual inspection. We frame these diagnostics into a model-agnostic benchmark, Spatial Audio Error Bench (3DAE Bench), which takes arbitrary ground-truth and predicted binaural pairs and reports the prediction quality of audio novel-view synthesis models. Experiments on ViGAS outputs over Replay-NVAS and SoundSpaces show different dominant failure modes: temporal misalignment on Replay-NVAS and ILD mismatch on SoundSpaces. Overall, the framework provides interpretable failure-mode summaries and intuitive visual maps for audio Novel-view-synthesis model development optimization.


【12】Escaping the Linearity Trap: Manifold Detours for Black-Box Adversarial Attacks on Singing Audio Deepfake Detection

标题:摆脱线性陷阱:歌唱音频Deepfake检测黑匣子对抗攻击的多种迂回
链接:https://arxiv.org/abs/2605.30366
作者:Yifan Liao,Yule Liu,Zhen Sun,Zongmin Zhang,Yupeng He,Jiaheng Wei,Xinhu Zheng,Xinlei He
摘要:最近的歌唱声音合成(SVS)技术进步使高度逼真但可能具有恶意的AI封面成为可能,这使得歌唱声音深度伪造检测(SVDD)变得至关重要。基于自监督学习(SSL)的检测器通过微调语音SSL主干来捕获特定于唱歌的欺骗伪像,从而实现最先进的性能。现有的对抗性攻击通常无法抵抗SSL-SVDD,从而造成了固有鲁棒性的假象。我们发现这源于两个挑战。首先,在客观层面上,攻击优化了局部代理的交叉熵,跨越了代理特定的边界,而不是抑制共享的欺骗证据。其次,在方法级,攻击遵循代理的主导梯度方向。在SSL-SVDD中,这与微调的伪影敏感方向相一致,限制了不可见检测器的可转移性-我们称之为线性陷阱的几何故障。为了正确评估鲁棒性,我们提出了MARS(元对抗性语义回归),这是一个为SSL-SVDD量身定制的基于传输的黑盒框架。在结构上,MARS通过从预训练的SSL空间构建自然语义锚和从微调空间构建工件锚,转向假设-证据操作。从理论上讲,MARS通过双层优化摆脱了线性陷阱:内部阶段诱导切线探索,而外部阶段将音频引导到自然语义流形。在CtrSVDD基准测试上的实验表明,MARS提高了分发内传输(13%)、分发外传输(10%)和跨任务评估(36%)的攻击成功率(ASR),突出了对强大SVDD系统的迫切需求。
摘要:Recent Singing Voice Synthesis (SVS) advances enable highly realistic but potentially malicious AI covers, making singing voice deepfake detection (SVDD) crucial. Self-Supervised Learning (SSL)-based detectors achieve state-of-the-art performance by fine-tuning speech SSL backbones to capture singing-specific spoof artifacts. Existing adversarial attacks often fail against SSL-SVDD, creating a false impression of inherent robustness. We reveal this stems from two challenges. First, at the objective level, attacks optimize cross-entropy on local surrogates, crossing surrogate-specific boundaries rather than suppressing shared spoof evidence. Second, at the method level, attacks follow the surrogate's dominant gradient direction. In SSL-SVDD, this aligns with fine-tuned artifact-sensitive directions, limiting transferability to unseen detectors - a geometric failure we term the Linearity Trap. To properly evaluate robustness, we propose MARS (Meta-Adversarial Regression of Semantics), a transfer-based black-box framework tailored to SSL-SVDD. Structurally, MARS shifts to hypothesis-evidence manipulation by constructing a natural semantic anchor from the pre-trained SSL space and an artifact anchor from the fine-tuned space. Algorithmically, MARS escapes the Linearity Trap via bi-level optimization: the inner stage induces tangential exploration, while the outer stage guides the audio toward the natural semantic manifold. Experiments on the CtrSVDD benchmark show MARS improves Attack Success Rate (ASR) in in-distribution transfer (13%), out-of-distribution transfer (10%), and cross-task evaluation (36%), highlighting the urgent need for robust SVDD systems.


【13】Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation

标题:精神伤害:对检索增强文本到音乐生成的字幕中毒攻击
链接:https://arxiv.org/abs/2605.30365
作者:Yizhu Wen,Shuhao Zhang,Nan Zhang,Long Cheng,Hanqing Guo
备注:This paper was accepted by the S&P 2026 ArtSec Workshop
摘要:检索增强的文本到音乐(TTM)系统使用从音乐字幕数据集检索的字幕来增强未指定的用户提示。该设计引入了对音乐知识库的完整性依赖。我们表明,攻击者可以通过注入少量精心制作的音乐字幕来毒害数据库,导致系统检索恶意字幕,这些字幕会使提示增强和引导生成远离用户的预期功能,而无需修改用户提示,检索器或生成器。为了实现音乐字幕中毒攻击,我们提出了一个双层字幕中毒策略,保留高级别的检索锚,同时注入低级别的声学描述符,以引导即时增强和下游音乐生成攻击者选择的目标意图。在MusicCaps知识数据库、CLAP检索器和MusicGen管道中,中毒的代基本上更接近攻击者的目标,同时与原始用户查询保持一致。这些结果暴露了检索增强的创造性AI系统的实际完整性风险。我们的演示可以在https://yizhu-wen.github.io/Mental-Damage/上找到
摘要:Retrieval-augmented text-to-music (TTM) systems augment underspecified user prompts using captions retrieved from a music caption dataset. This design introduces an integrity dependency on the music knowledge database. We show that an attacker can poison the database by injecting a small number of crafted music captions, causing the system to retrieve malicious captions that bias prompt augmentation and steer generation away from the user's intended function, without modifying the user prompt, retriever, or generator. To achieve the music caption poisoning attack, we propose a dual-layer caption poisoning strategy that preserves high-level retrieval anchors while injecting low-level acoustic descriptors to steer prompt augmentation and downstream music generation toward an attacker-chosen target intent. In a MusicCaps knowledge database, CLAP retriever, and MusicGen pipeline, poisoned generations move substantially closer to the attacker's target, while remaining comparably aligned with the original user query. These results expose a practical integrity risk for retrieval-augmented creative AI systems. Our demo can be found at: https://yizhu-wen.github.io/Mental-Damage/


【14】UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion

标题:UNISON:通过Deep LLM Fusion的统一声音生成和编辑框架
链接:https://arxiv.org/abs/2605.31530
作者:Zhaoqing Li,Haoning Xu,Jingran Su,Yaofang Liu,Zhefan Rao,Huimeng Wang,Jiajun Deng,Tianzi Wang,Zengrui Jin,Rui Liu,Haoxuan Che,Xunying Liu
摘要:我们提出了UNISON,一个潜在的扩散框架,统一的语音生成,声音生成和音频编辑在一个单一的模型。单个模型处理文本到音频、文本到语音、zero-shot扬声器克隆、混合语音和声音生成、场景级音频编辑、场景中语音编辑和定时时间合成,所有这些都共享单个权重集。我们的架构具有两个核心设计:(1)逐层深度LLM融合,其通过学习投影将来自冻结MLLM的均匀采样层的隐藏状态注入到相应的MM-DiT块中,提供深度匹配的语义条件,从而改善单层基线上的指令跟随;以及(2)统一的多任务体系结构,其中任务身份仅由逐声道掩码编码,并且源音频通过VAE编码的声道级联来提供。训练通过在线GPU端多任务数据合成管道进行稳定,该管道具有任务同质化和两阶段课程。凭借621 M--732 M的可训练参数,UNISON在所评估的领域中实现了与任务专家模型竞争或超过任务专家模型的结果,同时比类似的统一系统大约便宜4倍。
摘要:We present UNISON, a latent diffusion framework that unifies speech generation, sound generation, and audio editing within a single model. A single model handles text-to-audio, text-to-speech, zero-shot speaker cloning, mixed speech-and-sound generation, scene-level audio editing, speech-in-scene editing, and timed temporal composition, all of which share a single set of weights. Our architecture features two core designs: (1) Layer-wise deep LLM fusion, which injects hidden states from uniformly sampled layers of a frozen MLLM into corresponding MM-DiT blocks via learned projections, providing depth-matched semantic conditioning that improves instruction following over single-layer baselines; and (2) a unified multi-task architecture where task identity is encoded solely by a channel-wise mask and source audio is provided through VAE-encoded channel concatenation. Training is stabilized by an online GPU-side multi-task data synthesis pipeline with task-homogeneous batching and a two-stage curriculum. With 621M--732M trainable parameters, UNISON achieves results competitive with or exceeding task-specialist models across evaluated domains, while being roughly $4\times$ smaller than comparable unified systems.


【15】Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

标题:通过自回归扩散Transformer实现流同步空间音频生成
链接:https://arxiv.org/abs/2605.30940
作者:Ke Lei,Yu Zhang,Changhao Pan,Xueyi Pu,Wenxiang Guo,Ruiqi Li,Zhou Zhao
备注:Accepted by ICML 2026
摘要:实时和准确的空间音频生成对于提供沉浸式体验至关重要。然而,现有的空间音频合成技术通常受到生成质量和高推理延迟之间的权衡以及难以从多模态输入捕获精确的空间信息的阻碍。为了解决这些挑战,我们提出了SwanSphere,这是一个统一的流媒体框架,用于从全景视频和文本提示生成高保真空间音频。SwanSphere主要做了以下贡献:1)我们引入了一个因果自回归扩散Transformer架构,使流高质量的空间音频生成。2)我们设计了一个空间视频音频对比(SVAC)学习策略,使视频编码器与声学域对齐,并进一步采用多目标在线直接偏好优化(ODPO)方案,从而实现强大的空间感知和鲁棒的多模态空间音频合成。3)为了缓解目前空间音频数据集的稀缺性,我们还开发了一个自动注释管道来生成详细的空间字幕。实验结果表明,SwanSphere在视频到空间和文本到空间的音频生成任务中实现了优异的性能。演示可以在https://swanaigc.github.io上找到。
摘要:Real-time and accurate spatial audio generation is pivotal for delivering an immersive experience. However, existing spatial audio synthesis technologies are often encumbered by a tradeoff between generation quality and high inference latency, as well as difficulty in capturing precise spatial information from multimodal inputs. To address these challenges, we propose SwanSphere, a unified streaming framework for high-fidelity spatial audio generation from panoramic videos and text prompts. SwanSphere mainly makes the following contributions: 1) We introduce a causal autoregressive diffusion transformer architecture that enables streaming high-quality spatial audio generation. 2) We design a Spatial Video-Audio Contrastive (SVAC) learning strategy to align the video encoder with the acoustic domain, and further employ a multi-objective online direct preference optimization (ODPO) scheme, resulting in strong spatial perception and robust multimodal spatial audio synthesis. 3) To alleviate the current scarcity of spatial audio datasets, we also develop an automated annotation pipeline for generating detailed spatial captions. Experimental results demonstrate that SwanSphere achieves superior performance in both video-to-spatial and text-to-spatial audio generation tasks. Demos can be found at: https://swanaigc.github.io.


【16】A Unified and Reproducible Experimentation Framework for Speech Understanding

标题:统一且可复制的言语理解实验框架
链接:https://arxiv.org/abs/2605.30899
作者:Jing Peng,Junhao Du,Chenghao Wang,Hanqi Li,Yi Yang,Yixuan Wang,Xiaoyu Gu,Guanyu Chen,Yucheng Wang,Jiang Li,Zhangjie Zhao,Haoran Wang,Wenming Tu,Haoyu Li,Duo Ma,Lirong Qian,Yu Xi,Wen Wen,Jiaqi Guo,Hui Zhang,Shuai Fan,Wenbin Jiang,Shuai Wang,Kai Yu
备注:This paper is submitted to INTERSPEECH 2026
摘要:语音基础模型和语音LLM具有先进的语音理解能力,但面向部署的模型选择受到不匹配的后处理导致的非可比评估以及难以跨数据规模和管道重现的训练结果的阻碍。我们提出了SURE,一个统一的实验框架,该框架将预测格式,归一化和评分结合起来。SURE评估跨范式的强大系统,从传统的管道到语音LLM,在现实的声学和语言压力下的代表性任务。除了评估之外,SURE还引入了一个代理辅助的训练转换流程,该流程将纸张和代码映射到版本化的、可运行的训练管道中,并在匹配的开放数据子集上使用统一协议。总体而言,SURE提高了部署导向评估的可比性和再现性。
摘要:Speech foundation models and Speech LLMs have advanced speech understanding, yet deployment-oriented model selection is hindered by non-comparable evaluations caused by mismatched post-processing, and by training results that are hard to reproduce across data scales and pipelines. We present SURE, a unified experimentation framework that standardizes prediction formats, normalization, and scoring. SURE evaluates strong systems across paradigms, from conventional pipelines to Speech LLMs, on representative tasks under realistic acoustic and linguistic stressors. Beyond evaluation, SURE introduces an agent-assisted training conversion flow that maps paper and code into versioned, runnable training pipelines under a unified protocol on matched open-data subsets. Overall, SURE improves comparability and reproducibility for deployment-oriented evaluation.


eess.AS音频处理


【1】UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion
标题:UNISON:通过Deep LLM Fusion的统一声音生成和编辑框架
链接:https://arxiv.org/abs/2605.31530
作者:Zhaoqing Li,Haoning Xu,Jingran Su,Yaofang Liu,Zhefan Rao,Huimeng Wang,Jiajun Deng,Tianzi Wang,Zengrui Jin,Rui Liu,Haoxuan Che,Xunying Liu
摘要:我们提出了UNISON,一个潜在的扩散框架,统一的语音生成,声音生成和音频编辑在一个单一的模型。单个模型处理文本到音频、文本到语音、zero-shot扬声器克隆、混合语音和声音生成、场景级音频编辑、场景中语音编辑和定时时间合成,所有这些都共享单个权重集。我们的架构具有两个核心设计:(1)逐层深度LLM融合,其通过学习投影将来自冻结MLLM的均匀采样层的隐藏状态注入到相应的MM-DiT块中,提供深度匹配的语义条件,从而改善单层基线上的指令跟随;以及(2)统一的多任务体系结构,其中任务身份仅由逐声道掩码编码,并且源音频通过VAE编码的声道级联来提供。训练通过在线GPU端多任务数据合成管道进行稳定,该管道具有任务同质化和两阶段课程。凭借621 M-732 M的可训练参数,UNISON在评估的领域中实现了与任务专家模型竞争或超过任务专家模型的结果,同时比可比的统一系统小大约4倍。
摘要:We present UNISON, a latent diffusion framework that unifies speech generation, sound generation, and audio editing within a single model. A single model handles text-to-audio, text-to-speech, zero-shot speaker cloning, mixed speech-and-sound generation, scene-level audio editing, speech-in-scene editing, and timed temporal composition, all of which share a single set of weights. Our architecture features two core designs: (1) Layer-wise deep LLM fusion, which injects hidden states from uniformly sampled layers of a frozen MLLM into corresponding MM-DiT blocks via learned projections, providing depth-matched semantic conditioning that improves instruction following over single-layer baselines; and (2) a unified multi-task architecture where task identity is encoded solely by a channel-wise mask and source audio is provided through VAE-encoded channel concatenation. Training is stabilized by an online GPU-side multi-task data synthesis pipeline with task-homogeneous batching and a two-stage curriculum. With 621M--732M trainable parameters, UNISON achieves results competitive with or exceeding task-specialist models across evaluated domains, while being roughly $4\times$ smaller than comparable unified systems.


【2】Improving acoustic drone detection generalization through pretraining and data augmentation

标题:通过预训练和数据增强提高声学无人机检测的通用性
链接:https://arxiv.org/abs/2605.31329
作者:Paul M. Reuter,Mattes Ohlenbusch,Christian Rollwage
备注:Accepted to Quiet Drones 2026
摘要:检测未经授权的无人机飞行对于监视、安全和空域管理至关重要。声学无人机探测依赖于无人机独特的螺旋桨和马达声音,提供了一种低成本、无源的解决方案,不需要视线。一个核心挑战是泛化:在看不见的记录设置、环境和无人机类型(域外)中可靠地区分无人机签名和环境噪声。受大规模音频预训练进步的启发,我们开发了一种紧凑的基于DNN的检测器,并通过以下方式提高其泛化能力:(1)在对各种内部和公共无人机录音进行微调之前,对广泛的声音事件分类模型进行预训练,以及(2)应用动态增强(音调偏移、噪声混合、麦克风传递函数模拟、声谱图增强)以将模型暴露于变化的声学条件。消融研究量化了每次增强的影响。为了进行评估,我们设定了与真实世界监测需求相一致的目标假阳性率(FPR),并报告了域内数据(公共IDMT Berne 2022)和域外数据(公共AuDroK)的真阳性率(TPR)。我们的研究结果表明,预训练是鲁棒检测的主导因素,在所有基准测试中,与从头开始的训练相比,TPR得到了实质性的改善。完整的增强链为声学不匹配的域外数据提供了额外的增益,在AuDroK子集上实现了最佳的平均TPR,并在最具挑战性的场景中实现了最大的改进。我们通过测量公共非无人机语料库(IDMT-TRAFFIC和ESC-50)的误报,进一步验证了真实世界的适用性,在不熟悉的背景下表现出同样低的FPR。对IDMT Berne 2022的距离依赖性分析显示,在高达150米的距离上进行有效检测。
摘要:Detecting unauthorized UAV flights is critical for surveillance, security, and airspace management. Acoustic drone detection, which relies on the distinctive propeller and motor sounds of UAVs, provides a low-cost, passive solution that requires no line of sight. A central challenge is generalization: reliably distinguishing drone signatures from ambient noise across unseen recording setups, environments, and UAV types (out-of-domain). Inspired by advances in large-scale audio pretraining, we develop a compact DNN-based detector and improve its generalization by (1) pretraining the model for broad sound-event classification before fine-tuning on diverse in-house and public drone recordings, and (2) applying on-the-fly augmentations (pitch shifting, noise mixing, microphone transfer function simulation, spectrogram augmentation) to expose the model to varied acoustic conditions. An ablation study quantifies the impact of each augmentation. For evaluation, we set target false-positive rates (FPR) aligned with real-world surveillance needs and report true-positive rates (TPR) on both in-domain data (public IDMT Berne 2022) and out-of-domain data (public AuDroK). Our results show that pretraining is the dominant factor for robust detection, yielding substantial TPR improvements over training from scratch on all benchmarks. The full augmentation chain provides additional gains on acoustically mismatched out-of-domain data, achieving the best mean TPR on the AuDroK subsets and the largest improvements on the most challenging scenarios. We further validate real-world applicability by measuring false positives on public non-drone corpora (IDMT-TRAFFIC and ESC-50), demonstrating equally low FPR on unfamiliar backgrounds. A distance-dependent analysis on IDMT Berne 2022 shows effective detection at distances up to 150 m.


【3】On the Use of Dereverberation for Acoustic Feedback Cancellation

标题:关于使用去回响消除声反馈
链接:https://arxiv.org/abs/2605.31101
作者:Basil Liekens,Arnout Roebben,Toon van Waterschoot,Marc Moonen
备注:Accepted for publication in proceedings of EUSIPCO 2026
摘要:在公共广播系统和助听器中,最大可实现的放大或增益受到声反馈的限制。因此,为了能够应用更高的增益,需要反馈消除方法。此外,在回放之前,通常还希望对记录的信号进行去混响,即去除信号的后期混响分量。在本文中,它表明,在两个温和的条件下,声反馈信号可以写为源信号的混响版本。因此,可以将联合去混响和声反馈消除问题视为仅去混响问题,这意味着去混响算法可以应用于联合问题。模拟证实了这一发现
摘要:In public address systems and hearing aids, the maximally achievable amplification or gain is limited by acoustic feedback. Therefore, in order to be able to apply a higher gain, feedback cancellation methods are required. In addition, it is oftentimes also desirable to dereverberate a recorded signal, that is, remove the late reverberation component of the signal, before playing it back. In this paper, it is shown that under two mild conditions, the acoustic feedback signal can be written as a reverberant version of the source signal. Therefore, it is possible to treat the joint dereverberation and acoustic feedback cancellation problem as a dereverberation-only problem, meaning that dereverberation algorithms can be applied to the joint problem. Simulations corroborate this finding


【4】SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

标题:SwanVoice:独白和对话的表达性长篇Zero-Shot语音合成
链接:https://arxiv.org/abs/2605.30993
作者:Ruiqi Li,Yu Zhang,Changhao Pan,Ke Lei,Xiang Yin,Cheng Yang
备注:Technical Report
摘要:Zero-shot文本到语音(TTS)已经大大改善了单扬声器合成,但表达长形式的多扬声器对话仍然很困难。一个常见的解决方法是用独白TTS模型合成每个回合,并将输出缝合在一起。这增加了推理成本,并经常打破声音的一致性,会话的连贯性和情感的连续性。最近的对话TTS系统已经开始解决这个问题,但他们仍然努力保持表达连贯性,可控的扬声器切换,并在同一时间独白质量。我们介绍SwanData-Speech和SwanVoice。SwanData-Speech从野外音频中构建独白和对话语料库,使用Swan Forced Aligner进行停顿感知单词级对齐,并使用RobustMegaTTS 3进行发音困难的情况。基于这些数据,SwanVoice是一个适用于1- 4个扬声器的zero-shot TTS模型,结合了25 Hz VAE、带有暂停感知符号的原始文本调节和暂停替换,以及带有扬声器转向调节的流匹配DiT。训练从独白语音开始,通过混合和真实的对话数据,然后使用DiffusionNFT后训练与电话级别和说话者相似性奖励。在SwanBench-Speech上,SwanVoice在独白和对话设置中获得了比所有评估的开源基线更高的丰富性和层次分数,而内容准确性仍然是主要限制。音频演示可在https://swanaigc.github.io//#swanvoice上获得。
摘要:Zero-shot text-to-speech (TTS) has improved substantially for single-speaker synthesis, yet expressive long-form multi-speaker dialogue remains difficult. A common workaround is to synthesize each turn with a monologue TTS model and stitch the outputs together. This adds inference cost and often breaks acoustic consistency, conversational coherence, and affective continuity across turns. Recent dialogue TTS systems have begun to address this setting, but they still struggle to keep expressive coherence, controllable speaker switching, and monologue quality at the same time. We present SwanData-Speech and SwanVoice. SwanData-Speech builds monologue and dialogue corpora from in-the-wild audio, using Swan Forced Aligner for pause-aware word-level alignment and RobustMegaTTS3 for pronunciation-hard cases. Built on these data, SwanVoice is a zero-shot TTS model for 1--4 speakers, combining a 25 Hz VAE, raw-text conditioning with pause-aware symbols and pinyin substitution, and a flow-matching DiT with speaker-turn conditioning. Training starts from monologue speech, moves through mixed and real dialogue data, and then uses DiffusionNFT post-training with phone-level and speaker-similarity rewards. On SwanBench-Speech, SwanVoice obtains higher richness and hierarchy scores than all evaluated open-source baselines in both monologue and dialogue settings, while content accuracy remains the main limitation. Audio demos are available at https://swanaigc.github.io//#swanvoice.


【5】ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

标题:ImmersiveTTC:具有多模式扩散Transformer和特定领域表示对齐的环境感知文本到语音
链接:https://arxiv.org/abs/2605.30965
作者:Jun-Hak Yun,Seung-Bin Kim,Seong-Whan Lee
备注:Accepted to ACL 2026 main conference. Code is available at https://github.com/jjunak-yun/ImmersiveTTS
摘要:文本引导音频生成的最新进展在包括声音效果、语音和音乐在内的各种领域都取得了可喜的成果。然而,联合生成语音与环境音频仍然具有挑战性,由于其声学模式和时间动态的固有差异。我们提出ImmersiveTTS,环境感知的文本到语音(TTS)模型,生成自然语音无缝集成在环境背景下明确建模跨模态的相互作用。我们的模型建立在一个多模态扩散Transformer和融合成绩单对齐的语音潜在的文本条件下的环境上下文通过联合注意。为了增强语义一致性,我们引入了一个特定于域的表示对齐目标,针对环境感知的TTS,利用语音和音频编码器的互补自监督表示。实验结果表明,ImmersiveTTS实现了更高的自然度,可懂度和音频保真度比现有的方法在客观指标和人类听力测试。
摘要:Recent advancements in text-guided audio generation have yielded promising results in diverse domains, including sound effects, speech, and music. However, jointly generating speech with environmental audio remains challenging due to the inherent disparities in their acoustic patterns and temporal dynamics. We propose ImmersiveTTS, an environment-aware text-to-speech (TTS) model that generates natural speech seamlessly integrated within environmental contexts by explicitly modeling cross-modal interactions. Our model builds on a multimodal diffusion transformer and fuses transcript-aligned speech latent with text-conditioned environmental context via joint attention. To enhance semantic consistency, we introduce a domain-specific representation alignment objective tailored to environment-aware TTS, leveraging complementary self-supervised representations from speech and audio encoders. Experimental results show that ImmersiveTTS achieves higher naturalness, intelligibility, and audio fidelity than existing approaches across objective metrics and human listening tests.


【6】Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

标题:通过自回归扩散Transformer实现流同步空间音频生成
链接:https://arxiv.org/abs/2605.30940
作者:Ke Lei,Yu Zhang,Changhao Pan,Xueyi Pu,Wenxiang Guo,Ruiqi Li,Zhou Zhao
备注:Accepted by ICML 2026
摘要:实时和准确的空间音频生成对于提供沉浸式体验至关重要。然而,现有的空间音频合成技术通常受到生成质量和高推理延迟之间的权衡以及难以从多模态输入捕获精确的空间信息的阻碍。为了解决这些挑战,我们提出了SwanSphere,这是一个统一的流媒体框架,用于从全景视频和文本提示生成高保真空间音频。SwanSphere主要做了以下贡献:1)我们引入了一个因果自回归扩散Transformer架构,使流高质量的空间音频生成。2)我们设计了一个空间视频音频对比(SVAC)学习策略,使视频编码器与声学域对齐,并进一步采用多目标在线直接偏好优化(ODPO)方案,从而实现强大的空间感知和鲁棒的多模态空间音频合成。3)为了缓解目前空间音频数据集的稀缺性,我们还开发了一个自动注释管道来生成详细的空间字幕。实验结果表明,SwanSphere在视频到空间和文本到空间的音频生成任务中实现了优异的性能。演示可以在https://swanaigc.github.io上找到。
摘要:Real-time and accurate spatial audio generation is pivotal for delivering an immersive experience. However, existing spatial audio synthesis technologies are often encumbered by a tradeoff between generation quality and high inference latency, as well as difficulty in capturing precise spatial information from multimodal inputs. To address these challenges, we propose SwanSphere, a unified streaming framework for high-fidelity spatial audio generation from panoramic videos and text prompts. SwanSphere mainly makes the following contributions: 1) We introduce a causal autoregressive diffusion transformer architecture that enables streaming high-quality spatial audio generation. 2) We design a Spatial Video-Audio Contrastive (SVAC) learning strategy to align the video encoder with the acoustic domain, and further employ a multi-objective online direct preference optimization (ODPO) scheme, resulting in strong spatial perception and robust multimodal spatial audio synthesis. 3) To alleviate the current scarcity of spatial audio datasets, we also develop an automated annotation pipeline for generating detailed spatial captions. Experimental results demonstrate that SwanSphere achieves superior performance in both video-to-spatial and text-to-spatial audio generation tasks. Demos can be found at: https://swanaigc.github.io.


【7】A Unified and Reproducible Experimentation Framework for Speech Understanding

标题:统一且可复制的言语理解实验框架
链接:https://arxiv.org/abs/2605.30899
作者:Jing Peng,Junhao Du,Chenghao Wang,Hanqi Li,Yi Yang,Yixuan Wang,Xiaoyu Gu,Guanyu Chen,Yucheng Wang,Jiang Li,Zhangjie Zhao,Haoran Wang,Wenming Tu,Haoyu Li,Duo Ma,Lirong Qian,Yu Xi,Wen Wen,Jiaqi Guo,Hui Zhang,Shuai Fan,Wenbin Jiang,Shuai Wang,Kai Yu
备注:This paper is submitted to INTERSPEECH 2026
摘要:语音基础模型和语音LLM具有先进的语音理解能力,但面向部署的模型选择受到不匹配的后处理导致的非可比评估以及难以跨数据规模和管道重现的训练结果的阻碍。我们提出了SURE,一个统一的实验框架,该框架将预测格式,归一化和评分结合起来。SURE评估跨范式的强大系统,从传统的管道到语音LLM,在现实的声学和语言压力下的代表性任务。除了评估之外,SURE还引入了一个代理辅助的训练转换流程,该流程将纸张和代码映射到版本化的、可运行的训练管道中,并在匹配的开放数据子集上使用统一协议。总体而言,SURE提高了部署导向评估的可比性和再现性。
摘要:Speech foundation models and Speech LLMs have advanced speech understanding, yet deployment-oriented model selection is hindered by non-comparable evaluations caused by mismatched post-processing, and by training results that are hard to reproduce across data scales and pipelines. We present SURE, a unified experimentation framework that standardizes prediction formats, normalization, and scoring. SURE evaluates strong systems across paradigms, from conventional pipelines to Speech LLMs, on representative tasks under realistic acoustic and linguistic stressors. Beyond evaluation, SURE introduces an agent-assisted training conversion flow that maps paper and code into versioned, runnable training pipelines under a unified protocol on matched open-data subsets. Overall, SURE improves comparability and reproducibility for deployment-oriented evaluation.


【8】OpenSTBench: Beyond Semantic Evaluation for Speech Translation

标题:OpenSTBench:超越语音翻译的语义评估
链接:https://arxiv.org/abs/2605.30792
作者:Yanjie An,Yuxiang Zhao,Yichi Zhang,Qixi Zheng,Yujie Tu,Keqi Deng,Kai Yu,Xie Chen
备注:Submitted to EMNLP 2026
摘要:语音翻译系统越来越多地跨越语音到文本翻译(S2TT)、语音到语音翻译(S2ST)、离线翻译和流生成,产生在模态、语音实现和定时行为方面不同的输出。现有的评估实践评估翻译质量、语音质量和时间质量等重要方面,但这些方面通常在单独的协议下进行评估,因此难以全面比较异构系统。为了解决这个差距,我们提出了OpenSTBench,一个统一的多维评估框架,将异构的语音翻译输出组织成一个共享的评估格式。OpenSTBench支持离线和流媒体设置中的S2TT和S2ST系统,并联合评估翻译质量、语音质量、说话人保留、情感和语言保真度、时间一致性和延迟。通过对有代表性的语音翻译系统的实验,我们表明,具有较强翻译质量的系统仍然可以在语音质量和时间质量上有很大的不同。OpenSTBench提供了一个可重现的协议,用于分析这些跨维度差异,并支持语音翻译系统的面向应用的比较。代码和数据集可在https://github.com/sjtuayj/OpenSTBench上获得。
摘要:Speech translation systems increasingly span speech-to-text translation (S2TT), speech-to-speech translation (S2ST), offline translation, and streaming generation, producing outputs that differ in modality, speech realization, and timing behavior. Existing evaluation practices assess important aspects such as translation quality, speech quality, and temporal quality, but these aspects are often evaluated under separate protocols, making it difficult to compare heterogeneous systems comprehensively. To address this gap, we present OpenSTBench, a unified multidimensional evaluation framework that organizes heterogeneous speech translation outputs into a shared evaluation format. OpenSTBench supports both S2TT and S2ST systems in offline and streaming settings, and jointly evaluates translation quality, speech quality, speaker preservation, emotion and paralinguistic fidelity, temporal consistency, and latency. Through experiments on representative speech translation systems, we show that systems with strong translation quality can still differ substantially in speech quality, as well as in temporal quality. OpenSTBench provides a reproducible protocol for analyzing these cross-dimensional differences and supporting application-oriented comparison of speech translation systems. The code and datasets are available at https://github.com/sjtuayj/OpenSTBench.


【9】FiPA-SR -- FiLM-Conditioned Perceptually Informed Audio Super-Resolution

标题:FiPA-SR --FiLM调节感知信息音频超分辨率
链接:https://arxiv.org/abs/2605.30594
作者:Wallace Abreu,Luiz W. P. Biscainho
备注:Submitted to the XLIV BRAZILIAN SYMPOSIUM ON TELECOMMUNICATIONS AND SIGNAL PROCESSING - SBrT 2026
摘要:音频带宽扩展的目的是从有限带宽的信号中重建丢失的高频内容。本文提出了FiPA-SR,这是一种基于GAN的感知架构,能够在单个模型中处理不同的输入带宽。在以前的$\textrm{AEROMamba}_\textrm{P}$框架的基础上,所提出的模型结合了薄膜层,以根据各自的带宽来适应重建过程。在MUSDB数据集上的实验表明,FiPA-SR在8、20和32 kHz输入采样率下的性能优于最先进的AudioSR模型。此外,所提出的架构使用约3$\times$更少的GPU内存和执行推理超过60$\times$比基于扩散的基线。
摘要:Audio bandwidth extension aims to reconstruct missing high-frequency content from bandlimited signals. This paper proposes FiPA-SR, a GAN-based perceptual architecture capable of handling different input bandwidths within a single model. Building upon the previous $\textrm{AEROMamba}_\textrm{P}$ framework, the proposed model incorporates FiLM layers to adapt the reconstruction process according to the respective bandwidth. Experiments on the MUSDB dataset show that FiPA-SR outperforms the state-of-the-art AudioSR model across 8, 20, and 32 kHz input sampling rates. Moreover, the proposed architecture uses approximately 3$\times$ less GPU memory and performs inference more than 60$\times$ faster than the diffusion-based baseline.


【10】Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

标题:提取巴西葡萄牙语口语中不带社会语言标签的口音特征
链接:https://arxiv.org/abs/2605.30457
作者:Pedro H. L. Leite,Pedro Benevenuto Valadares,Luiz W. P. Biscainho
备注:This work was submitted to the XLIV Brazilian Symposium on Telecommunications and Signal Processing (SBrT 2026)
摘要:巴西葡萄牙语(pt-BR)的地区口音分类需要可靠的标签。虽然大型自监督学习(SSL)语音模型功能强大,但它们的训练管道会稀释社会语音信息,因为口音标签通常不可靠或不用于训练目标。这项工作介绍了一种新的工作流程,只使用声学标签的特征提取。通过隔离明确的区域口音地标和使用基于音素的强制对齐器(ZIPA),我们的目标特征集捕获方言的变化比话语嵌入更有效,表明本地化的功能可以优于通用架构的口音相关的任务,使用最小和客观的数据标签。
摘要:Regional accent classification in Brazilian Portuguese (pt-BR) suffers from the need for reliable labeling. While large self-supervised learning (SSL) speech models are powerful, their training pipelines dilute sociophonetic information, since accent labels are generally not reliable or are not used in training objectives. This work introduces a novel workflow for feature extraction using only acoustic labels. By isolating explicit regional accent landmarks and using a phoneme-based forced aligner (ZIPA), our targeted feature set captures dialectal variance more effectively than utterance embeddings, demonstrating that localized features can outperform general-purpose architectures on accent-related tasks using minimal and objective data labels.


【11】Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus

标题:缩放对话匈牙利SVR:BEA-Dialogue+ Corpus
链接:https://arxiv.org/abs/2605.31469
作者:Máté Gedeon,Piroska Zsófia Barta,Péter Mihajlik,Katalin Mády
摘要:匈牙利语的对话式自动语音识别受到公开可用的对话式训练数据数量有限的限制。BEA-Dialogue语料库满足了这一需求,但其严格的扬声器分离训练/开发/评估分离将可用材料减少到只有85小时。在本文中,我们介绍BEA-Dialogue+,一个扩展版本的语料库,放宽了实验者和对话伙伴的分裂标准,同时保持完全分离的主要发言人。这导致了200小时的转录自然对话,并使额外的训练数据和跨分裂的说话者重叠之间的权衡的受控研究成为可能。我们在两个语料库版本上评估了几个基于Whisper和FastConformer的模型,包括基于序列化输出训练(SOT)的对话转录微调。我们的研究结果表明,较大的语料库是更具有挑战性的模型没有微调,而基于SOT的适应产生一致的改善WER,CER,cpWER和cpCER。总体而言,BEA-Dialogue+为匈牙利语对话ASR提供了一个更大但仍然要求严格的基准,以及培训和评估对话转录系统的实用资源。
摘要:Conversational automatic speech recognition in Hungarian is constrained by the limited amount of publicly available dialogue-style training data. The BEA-Dialogue corpus addresses this need, but its strictly speaker-disjoint train/dev/eval split reduces the usable material to only 85 hours. In this paper, we introduce BEA-Dialogue+, an expanded version of the corpus that relaxes the split criterion for experimenters and dialogue partners while preserving complete separation of the primary speakers. This results in 200 hours of transcribed natural conversations and enables a controlled study of the trade-off between additional training data and speaker overlap across the splits. We evaluate several Whisper- and FastConformer-based models on both corpus versions, including Serialized Output Training (SOT)-based fine-tuning for dialogue transcription. Our results show that the larger corpus is more challenging for models without fine-tuning, whereas SOT-based adaptation yields consistent improvements in WER, CER, cpWER, and cpCER. Overall, BEA-Dialogue+ provides a substantially larger yet still demanding benchmark for Hungarian dialogue ASR, and a practical resource for training and evaluating dialogue transcription systems.


【12】Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

标题:Chatterbox-Flash:用于流媒体Zero-ShotTTC的预先校准块扩散
链接:https://arxiv.org/abs/2605.30748
作者:Deokjin Seo,Gangin Park,Kihyun Nam
备注:8 pages, 4 figures, 9 tables
摘要:我们提出了Chatterbox-Flash,一个zero-shot文本到语音模型,通过将预训练的自回归TTS解码器微调为块扩散解码器,使每个块内的并行令牌生成,同时保留逐块流。我们发现,天真地将主流块扩散解码离散语音令牌降低质量,作为一个长尾令牌分布偏向于几个高频令牌的并行位置选择。为了在不修改架构的情况下减轻这种情况,我们引入了两种推理时间技术:事先校准的评分,它减去块级边缘令牌分布,和早期解码时间表,它自适应地终止迭代校准的信心的基础上。在标准的zero-shot TTS基准测试中,Chatterbox-Flash实现了与强自回归和非自回归基线相当的高保真合成,同时支持与流式AR系统相当的第一个数据包时间的流式推理,并大大降低了实时因素。代码和音频示例可在https://github.com/resemble-ai/chatterbox-flash上获得。
摘要:We present Chatterbox-Flash, a zero-shot text-to-speech model obtained by fine-tuning a pretrained autoregressive TTS decoder into a block-diffusion decoder, enabling parallel token generation within each block while retaining block-by-block streaming. We find that naively transferring mainstream block-diffusion decoding to discrete speech tokens degrades quality, as a long-tail token distribution biases parallel position selection toward a few high-frequency tokens. To mitigate this without architectural modification, we introduce two inference-time techniques: prior-calibrated scoring, which subtracts the block-level marginal token distribution, and an early-decoding schedule, which adaptively terminates iteration based on calibrated confidence. On standard zero-shot TTS benchmarks, Chatterbox-Flash attains high-fidelity synthesis comparable to strong autoregressive and non-autoregressive baselines, while supporting streaming inference with time-to-first-packet on par with streaming AR systems and substantially lower real-time factor. Code and audio samples are available at https://github.com/resemble-ai/chatterbox-flash.


【13】Escaping the Linearity Trap: Manifold Detours for Black-Box Adversarial Attacks on Singing Audio Deepfake Detection

标题:摆脱线性陷阱:歌唱音频Deepfake检测黑匣子对抗攻击的多种迂回
链接:https://arxiv.org/abs/2605.30366
作者:Yifan Liao,Yule Liu,Zhen Sun,Zongmin Zhang,Yupeng He,Jiaheng Wei,Xinhu Zheng,Xinlei He
摘要:最近的歌唱声音合成(SVS)技术进步使高度逼真但可能具有恶意的AI封面成为可能,这使得歌唱声音深度伪造检测(SVDD)变得至关重要。基于自监督学习(SSL)的检测器通过微调语音SSL主干来捕获特定于唱歌的欺骗伪像,从而实现最先进的性能。现有的对抗性攻击通常无法抵抗SSL-SVDD,从而造成了固有鲁棒性的假象。我们发现这源于两个挑战。首先,在客观层面上,攻击优化了局部代理的交叉熵,跨越了代理特定的边界,而不是抑制共享的欺骗证据。其次,在方法级,攻击遵循代理的主导梯度方向。在SSL-SVDD中,这与微调的伪影敏感方向相一致,限制了不可见检测器的可转移性-我们称之为线性陷阱的几何故障。为了正确评估鲁棒性,我们提出了MARS(元对抗性语义回归),这是一个为SSL-SVDD量身定制的基于传输的黑盒框架。在结构上,MARS通过从预训练的SSL空间构建自然语义锚和从微调空间构建工件锚,转向假设-证据操作。从理论上讲,MARS通过双层优化摆脱了线性陷阱:内部阶段诱导切线探索,而外部阶段将音频引导到自然语义流形。在CtrSVDD基准测试上的实验表明,MARS提高了分发内传输(13%)、分发外传输(10%)和跨任务评估(36%)的攻击成功率(ASR),突出了对强大SVDD系统的迫切需求。
摘要:Recent Singing Voice Synthesis (SVS) advances enable highly realistic but potentially malicious AI covers, making singing voice deepfake detection (SVDD) crucial. Self-Supervised Learning (SSL)-based detectors achieve state-of-the-art performance by fine-tuning speech SSL backbones to capture singing-specific spoof artifacts. Existing adversarial attacks often fail against SSL-SVDD, creating a false impression of inherent robustness. We reveal this stems from two challenges. First, at the objective level, attacks optimize cross-entropy on local surrogates, crossing surrogate-specific boundaries rather than suppressing shared spoof evidence. Second, at the method level, attacks follow the surrogate's dominant gradient direction. In SSL-SVDD, this aligns with fine-tuned artifact-sensitive directions, limiting transferability to unseen detectors - a geometric failure we term the Linearity Trap. To properly evaluate robustness, we propose MARS (Meta-Adversarial Regression of Semantics), a transfer-based black-box framework tailored to SSL-SVDD. Structurally, MARS shifts to hypothesis-evidence manipulation by constructing a natural semantic anchor from the pre-trained SSL space and an artifact anchor from the fine-tuned space. Algorithmically, MARS escapes the Linearity Trap via bi-level optimization: the inner stage induces tangential exploration, while the outer stage guides the audio toward the natural semantic manifold. Experiments on the CtrSVDD benchmark show MARS improves Attack Success Rate (ASR) in in-distribution transfer (13%), out-of-distribution transfer (10%), and cross-task evaluation (36%), highlighting the urgent need for robust SVDD systems.


【14】Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation

标题:精神伤害:对检索增强文本到音乐生成的字幕中毒攻击
链接:https://arxiv.org/abs/2605.30365
作者:Yizhu Wen,Shuhao Zhang,Nan Zhang,Long Cheng,Hanqing Guo
备注:This paper was accepted by the S&P 2026 ArtSec Workshop
摘要:检索增强的文本到音乐(TTM)系统使用从音乐字幕数据集检索的字幕来增强未指定的用户提示。该设计引入了对音乐知识库的完整性依赖。我们表明,攻击者可以通过注入少量精心制作的音乐字幕来毒害数据库,导致系统检索恶意字幕,这些字幕会使提示增强和引导生成远离用户的预期功能,而无需修改用户提示,检索器或生成器。为了实现音乐字幕中毒攻击,我们提出了一个双层字幕中毒策略,保留高级别的检索锚,同时注入低级别的声学描述符,以引导即时增强和下游音乐生成攻击者选择的目标意图。在MusicCaps知识数据库、CLAP检索器和MusicGen管道中,中毒的代基本上更接近攻击者的目标,同时与原始用户查询保持一致。这些结果暴露了检索增强的创造性AI系统的实际完整性风险。我们的演示可以在https://yizhu-wen.github.io/Mental-Damage/上找到
摘要:Retrieval-augmented text-to-music (TTM) systems augment underspecified user prompts using captions retrieved from a music caption dataset. This design introduces an integrity dependency on the music knowledge database. We show that an attacker can poison the database by injecting a small number of crafted music captions, causing the system to retrieve malicious captions that bias prompt augmentation and steer generation away from the user's intended function, without modifying the user prompt, retriever, or generator. To achieve the music caption poisoning attack, we propose a dual-layer caption poisoning strategy that preserves high-level retrieval anchors while injecting low-level acoustic descriptors to steer prompt augmentation and downstream music generation toward an attacker-chosen target intent. In a MusicCaps knowledge database, CLAP retriever, and MusicGen pipeline, poisoned generations move substantially closer to the attacker's target, while remaining comparably aligned with the original user query. These results expose a practical integrity risk for retrieval-augmented creative AI systems. Our demo can be found at: https://yizhu-wen.github.io/Mental-Damage/


机器翻译由腾讯交互翻译提供,仅供参考