今日论文合集:cs.SD语音13篇,eess.AS音频处理10篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】A Sociolinguistic Analysis of Automatic Speech Recognition Bias in Newcastle English
标题:纽卡斯尔英语自动语音识别偏差的社会语言学分析
链接:https://arxiv.org/abs/2603.24549

作者:Dana Serditova,Kevin Tang
备注:54 pages, 11 figures
摘要:自动语音识别(ASR)系统广泛应用于日常通信、教育、医疗保健和工业领域,但其性能在不同说话者之间仍然不均衡,特别是当方言变异与训练数据中代表的主流口音不同时。本研究通过对纽卡斯尔英语的社会语言学分析来研究ASR偏见,纽卡斯尔英语是英格兰东北部的一种区域性语言,已被证明是对当前语音识别技术的挑战。使用泰恩赛德英语历时电子语料库(DECTE)的自发语音,我们评估了最先进的商业ASR系统的输出,并对3,000多个转录错误进行了细粒度的分析。错误按语言领域分类,并与性别、年龄和社会经济地位等社会变量进行检查。此外,声学案例研究选定的元音功能演示了梯度语音变化如何直接导致识别错误。   结果表明,语音变化占大多数的错误,与经常性的失败与方言的具体特点,如元音质量和声门化,以及当地的词汇和非标准的语法形式。错误率也因社会群体而异,男性和年龄范围极端的说话者的错误频率较高。这些发现表明,ASR错误不是随机的,而是社会模式,可以从社会语言学的角度来解释。因此,这项研究表明,将社会语言学专业知识纳入语音技术的评估和发展的重要性,并认为,更公平的ASR系统需要明确关注方言的变化和基于社区的语音数据。
摘要:Automatic Speech Recognition (ASR) systems are widely used in everyday communication, education, healthcare, and industry, yet their performance remains uneven across speakers, particularly when dialectal variation diverges from the mainstream accents represented in training data. This study investigates ASR bias through a sociolinguistic analysis of Newcastle English, a regional variety of North-East England that has been shown to challenge current speech recognition technologies. Using spontaneous speech from the Diachronic Electronic Corpus of Tyneside English (DECTE), we evaluate the output of a state-of-the-art commercial ASR system and conduct a fine-grained analysis of more than 3,000 transcription errors. Errors are classified by linguistic domain and examined in relation to social variables including gender, age, and socioeconomic status. In addition, an acoustic case study of selected vowel features demonstrates how gradient phonetic variation contributes directly to misrecognition.   The results show that phonological variation accounts for the majority of errors, with recurrent failures linked to dialect-specific features like vowel quality and glottalisation, as well as local vocabulary and non-standard grammatical forms. Error rates also vary across social groups, with higher error frequencies observed for men and for speakers at the extremes of the age spectrum. These findings indicate that ASR errors are not random but socially patterned and can be explained from a sociolinguistic perspective. Thus, the study demonstrates the importance of incorporating sociolinguistic expertise into the evaluation and development of speech technologies and argues that more equitable ASR systems require explicit attention to dialectal variation and community-based speech data.


【2】What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification
标题:学习什么以及何时学习:大规模说话者验证的Currency排名损失
链接:https://arxiv.org/abs/2603.24432

作者:Massa Baali,Sarthak Bisht,Rita Singh,Bhiksha Raj
摘要:大规模的说话人验证仍然是一个开放的挑战,因为固定的边缘损失平等地对待所有样本,无论质量如何。我们假设错误标记或降级的样本引入噪声梯度,破坏紧凑的扬声器流形。我们提出了Curry(Currency Ranking),这是一种自适应损失,通过子中心ArcFace在线估计样本难度:使用运行的批量统计,从占主导地位的子中心余弦相似性排名样本到简单,中等和困难层的置信度得分,没有辅助注释。可学习的权重引导模型从稳定的身份基础,通过流形细化到边界锐化。据我们所知,这是迄今为止训练的最大规模的说话人确认系统。在VoxCeleb 1-O和SITW上进行评估,Curry在次中心ArcFace基线上将EER降低了86.8%和60.0%,为不完美大规模数据上的鲁棒说话人验证建立了新的范例。
摘要:Speaker verification at large scale remains an open challenge as fixed-margin losses treat all samples equally regardless of quality. We hypothesize that mislabeled or degraded samples introduce noisy gradients that disrupt compact speaker manifolds. We propose Curry (CURriculum Ranking), an adaptive loss that estimates sample difficulty online via Sub-center ArcFace: confidence scores from dominant sub-center cosine similarity rank samples into easy, medium, and hard tiers using running batch statistics, without auxiliary annotations. Learnable weights guide the model from stable identity foundations through manifold refinement to boundary sharpening. To our knowledge, this is the largest-scale speaker verification system trained to date. Evaluated on VoxCeleb1-O, and SITW, Curry reduces EER by 86.8\% and 60.0\% over the Sub-center ArcFace baseline, establishing a new paradigm for robust speaker verification on imperfect large-scale data.


【3】Iterate to Differentiate: Enhancing Discriminability and Reliability in Zero-Shot TTS Evaluation
标题:迭代区分:增强Zero-ShotTTC评估的区分性和可靠性
链接:https://arxiv.org/abs/2603.24430

作者:Shengfan Shen,Di Wu,Xingchen Song,Dinghao Zhou,Liumeng Xue,Meng Meng,Jian Luan,Shuai Wang
备注:submitted to Interspeech 2026, under review
摘要:None
摘要:Reliable evaluation of modern zero-shot text-to-speech (TTS) models remains challenging. Subjective tests are costly and hard to reproduce, while objective metrics often saturate, failing to distinguish SOTA systems. To address this, we propose Iterate to Differentiate (I2D), an evaluation framework that recursively synthesizes speech using the model's own outputs as references. Higher-quality models exhibit greater resilience to the distributional shift induced by iterative synthesis, resulting in slower performance degradation. I2D exploits this differential degradation to amplify performance gaps and reveal robustness. By aggregating objective metrics across iterations, I2D improves discriminability and alignment with human judgments, increasing system-level SRCC from 0.118 to 0.464 for UTMOSv2. Experiments on 11 models across Chinese, English, and emotion datasets demonstrate that I2D enables more reliable automated evaluation for zero-shot TTS.


【4】Enhancing Efficiency and Performance in Deepfake Audio Detection through Neuron-level dropin & Neuroplasticity Mechanisms
标题:通过神经元级下垂和神经可塑性机制提高Deepfake音频检测的效率和性能
链接:https://arxiv.org/abs/2603.24343

作者:Yupei Li,Shuaijie Shao,Manuel Milling,Björn Schuller
备注:Accepted at IJCNN 2026
摘要:当前的音频深度伪造检测使用ResNet等多种深度学习架构取得了显着的性能,并且随着Wav2 Vec等大型模型(LM)的引入,性能得到了进一步的改进。大型语言模型(LLM)的成功进一步证明了扩展模型参数的好处,但也突出了一个瓶颈,即性能增益受到参数计数的限制。简单地堆叠额外的层,如在当前的LLM中所做的,在计算上是昂贵的,并且需要完全重新训练。此外,现有的低秩自适应方法主要应用于基于注意力的架构,这限制了它们的范围。受哺乳动物大脑中观察到的神经元可塑性的启发,我们提出了新的算法,dropin和进一步的可塑性,动态调整某些层中的神经元数量,以灵活地调节模型参数。我们在多种架构上评估这些算法,包括ResNet,Gated Recurrent Neural Networks和Wav2Vec。使用广泛认可的ASVSpoof2019 LA,PA和FakeorReal数据集的实验结果表明,使用dropin方法的计算效率得到了一致的提高,并且使用dropin和plasticity方法在这些数据集中分别最多降低了约39%和66%的等错误率。代码和补充材料可以在Github链接上找到。
摘要:Current audio deepfake detection has achieved remarkable performance using diverse deep learning architectures such as ResNet, and has seen further improvements with the introduction of large models (LMs) like Wav2Vec. The success of large language models (LLMs) further demonstrates the benefits of scaling model parameters, but also highlights one bottleneck where performance gains are constrained by parameter counts. Simply stacking additional layers, as done in current LLMs, is computationally expensive and requires full retraining. Furthermore, existing low-rank adaptation methods are primarily applied to attention-based architectures, which limits their scope. Inspired by the neuronal plasticity observed in mammalian brains, we propose novel algorithms, dropin and further plasticity, that dynamically adjust the number of neurons in certain layers to flexibly modulate model parameters. We evaluate these algorithms on multiple architectures, including ResNet, Gated Recurrent Neural Networks, and Wav2Vec. Experimental results using the widely recognised ASVSpoof2019 LA, PA, and FakeorReal dataset demonstrate consistent improvements in computational efficiency with the dropin approach and a maximum of around 39% and 66% relative reduction in Equal Error Rate with the dropin and plasticity approach among these dataset, respectively. The code and supplementary material are available at Github link.


【5】Bridging Biological Hearing and Neuromorphic Computing: End-to-End Time-Domain Audio Signal Processing with Reservoir Computing
标题:生物听力和神经形态计算的桥梁:端到端的时间域音频信号处理与水库计算
链接:https://arxiv.org/abs/2603.24283

作者:Rinku Sebastian,Simon O'Keefe,Martin Trefzer
摘要:尽管尖端技术取得了进步,但音频信号处理仍然面临挑战,并且缺乏人类语音处理系统的精度。为了解决这些挑战,我们提出了一种新的方法来简化音频信号处理,利用时域技术和水库计算。通过我们的研究,我们已经开发了一个实时音频信号处理系统,通过简化音频信号处理,通过利用水库计算机,这是显着更容易训练。   特征提取是语音信号处理中的基本步骤,其中Mel频率倒谱系数(MFCC)由于其与人类听觉的感知相关性而成为主要选择。然而,传统的MFCC提取依赖于计算密集的时间-频率变换,限制了实时应用的效率。为了解决这个问题,我们提出了一种新的方法,利用水库计算简化MFCC提取。通过用卷积运算取代传统的频域转换,我们消除了对复杂变换的需要,同时保持了特征的可辨别性。我们提出了一个端到端的音频处理框架,集成了这种方法,展示了其高效和实时语音分析的潜力。我们的研究成果有助于推动节能音频处理技术的发展,实现嵌入式系统和语音驱动应用的无缝部署。这项工作弥合了生物启发的特征提取和现代神经形态计算之间的差距,为下一代语音识别系统提供了可扩展的解决方案。
摘要:Despite the advancements in cutting-edge technologies, audio signal processing continues to pose challenges and lacks the precision of a human speech processing system. To address these challenges, we propose a novel approach to simplify audio signal processing by leveraging time-domain techniques and reservoir computing. Through our research, we have developed a real-time audio signal processing system by simplifying audio signal processing through the utilization of reservoir computers, which are significantly easier to train.   Feature extraction is a fundamental step in speech signal processing, with Mel Frequency Cepstral Coefficients (MFCCs) being a dominant choice due to their perceptual relevance to human hearing. However, conventional MFCC extraction relies on computationally intensive time-frequency transformations, limiting efficiency in real-time applications. To address this, we propose a novel approach that leverages reservoir computing to streamline MFCC extraction. By replacing traditional frequency-domain conversions with convolution operations, we eliminate the need for complex transformations while maintaining feature discriminability. We present an end-to-end audio processing framework that integrates this method, demonstrating its potential for efficient and real-time speech analysis. Our results contribute to the advancement of energy-efficient audio processing technologies, enabling seamless deployment in embedded systems and voice-driven applications. This work bridges the gap between biologically inspired feature extraction and modern neuromorphic computing, offering a scalable solution for next-generation speech recognition systems.


【6】Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
标题:口语对话系统中的语义感知中断检测:基准、指标和模型
链接:https://arxiv.org/abs/2603.24144

作者:Kangxiang Xia,Bingshen Mu,Xian Shi,Jin Xu,Lei Xie
备注:Accepted by ICME 2026
摘要:在口语对话系统(SDS)中实现自然的全双工交互仍然是一个挑战,因为难以准确地检测用户中断。目前的解决方案是两极分化的“制造商快乐”VAD为基础的方法,误解反向通道和强大的端到端的模型,表现出不可接受的响应延迟。此外,缺乏现实世界的基准和整体衡量标准也阻碍了这一领域的进展。本文提出了一个全面的框架,以克服这些限制。我们首先介绍SID-Bench,这是第一个完全基于真实人类对话的语义感知中断检测基准。为了提供一个严格的评估的响应性鲁棒性的权衡,我们提出了平均惩罚时间(APT)度量,它分配一个时间成本的误报和延迟响应。在此框架的基础上,我们设计了一个基于LLM的检测模型,通过一种新的训练范式进行优化,以捕获意图的微妙语义线索。实验结果表明,我们的模型显着优于主流基线,实现了近三倍的APT减少。通过成功解决速度和稳定性之间的长期紧张关系,我们的工作为SDS中的智能中断处理建立了一个新的最先进的技术。为了方便未来的研究,SID-Bench和相关代码可在https://github.com/xkx-hub/SID-bench上获得。
摘要:Achieving natural full-duplex interaction in spoken dialogue systems (SDS) remains a challenge due to the difficulty of accurately detecting user interruptions. Current solutions are polarized between "trigger-happy" VAD-based methods that misinterpret backchannels and robust end-to-end models that exhibit unacceptable response delays. Moreover, the absence of real-world benchmarks and holistic metrics hinders progress in the field. This paper presents a comprehensive frame-work to overcome these limitations. We first introduce SID-Bench, the first benchmark for semantic-aware interruption detection built entirely from real-world human dialogues. To provide a rigorous assessment of the responsiveness-robustness trade-off, we propose the Average Penalty Time (APT) metric, which assigns a temporal cost to both false alarms and late responses. Building on this framework, we design an LLM-based detection model optimized through a novel training paradigm to capture subtle semantic cues of intent. Experimental results show that our model significantly outperforms mainstream baselines, achieving a nearly threefold reduction in APT. By successfully resolving the long-standing tension between speed and stability, our work establishes a new state-of-the-art for intelligent interruption handling in SDS. To facilitate future research, SID-Bench and the associated code are available at: https://github.com/xkx-hub/SID-bench.


【7】Variable-Length Audio Fingerprinting
标题:可变长度音频指纹识别
链接:https://arxiv.org/abs/2603.23947

作者:Hongjie Chen,Hanyu Meng,Huimin Zeng,Ryan A. Rossi,Lie Lu,Josh Kimball
摘要:音频指纹识别将音频转换为更低维的表示,允许失真的录音仍然通过类似的指纹被识别为原始录音。现有的深度学习方法严格识别固定长度的音频片段,从而忽略了分割过程中的时间动态。为了解决由于这种刚性的限制,我们提出了可变长度音频指纹(VLAFP),一种新的方法,支持可变长度指纹。据我们所知,VLAFP是第一个能够处理可变长度音频的深度音频指纹模型,用于训练和测试。我们的实验表明,VLAFP优于现有的最先进的实时音频识别和音频检索在三个真实世界的数据集。
摘要:Audio fingerprinting converts audio to much lower-dimensional representations, allowing distorted recordings to still be recognized as their originals through similar fingerprints. Existing deep learning approaches rigidly fingerprint fixed-length audio segments, thereby neglecting temporal dynamics during segmentation. To address limitations due to this rigidity, we propose Variable-Length Audio FingerPrinting (VLAFP), a novel method that supports variable-length fingerprinting. To the best of our knowledge, VLAFP is the first deep audio fingerprinting model capable of processing audio of variable length, for both training and testing. Our experiments show that VLAFP outperforms existing state-of-the-arts in live audio identification and audio retrieval across three real-world datasets.


【8】Echoes: A semantically-aligned music deepfake detection dataset
标题:Echoes:一个语义对齐的音乐Deepfake检测数据集
链接:https://arxiv.org/abs/2603.23667

作者:Octavian Pascu,Dan Oneata,Horia Cucu,Nicolas M. Muller
摘要:我们介绍了Echoes,这是一个用于音乐deepfake检测的新数据集,旨在在现实和提供商多样化的条件下训练和基准检测器。Echoes包含3,577首曲目(110小时的音频),跨越多种流派(流行,摇滚,电子),并包括由十个流行的AI音乐生成系统生成的内容。为了防止捷径学习和促进鲁棒的泛化,数据集被故意构造成具有挑战性,在欺骗音频和真实参考之间强制执行语义级对齐。这种对齐是通过直接在真正的波形或歌曲描述符上调节生成的音频样本来实现的。我们使用最先进的Wav 2 Vec 2 XLS-R 2B表示,在跨数据集设置中对三个现有的AI生成的音乐数据集进行了回声评估。结果表明:(i)Echoes是最难的域内数据集;(ii)在现有数据集上训练的检测器向Echoes的转移很差;(iii)在Echoes上训练产生最强的泛化性能。这些研究结果表明,供应商的多样性和语义对齐有助于学习更多的可转移检测线索。
摘要:We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provider-diverse conditions. Echoes comprises 3,577 tracks (110 hours of audio) spanning multiple genres (pop, rock, electronic), and includes content generated by ten popular AI music generation systems. To prevent shortcut learning and promote robust generalization, the dataset is deliberately constructed to be challenging, enforcing semantic-level alignment between spoofed audio and bona fide references. This alignment is achieved by conditioning generated audio samples directly on bona-fide waveforms or song descriptors. We evaluate Echoes in a cross-dataset setting against three existing AI-generated music datasets using state-of-the-art Wav2Vec2 XLS-R 2B representations. Results show that (i) Echoes is the hardest in-domain dataset; (ii) detectors trained on existing datasets transfer poorly to Echoes; (iii) training on Echoes yields the strongest generalization performance. These findings suggest that provider diversity and semantic alignment help learn more transferable detection cues.


【9】YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
标题:YingMusic-Singer:可控的歌声合成,灵活的歌词操纵和无注释的旋律指导
链接:https://arxiv.org/abs/2603.24589

作者:Chunbo Hao,Junjie Zheng,Guobin Ma,Yuepeng Jiang,Huakang Chen,Wenjie Tian,Gongyu Chen,Zihao Chen,Lei Xie
摘要:在保持旋律一致性的同时用改变的歌词再生歌声仍然具有挑战性,因为现有方法要么提供有限的可控性,要么需要费力的手动对齐。我们提出了YingMusic-Singer,一个完全基于扩散的模型,使旋律可控的歌声合成与灵活的歌词操作。该模型有三个输入:一个可选的音色参考,一个提供旋律的演唱片段,以及修改后的歌词,无需手动调整。经过课程学习和集团相关政策优化的培训,YingMusic-Singer实现了比Vevo 2更强的旋律保留和歌词坚持,这是支持无需手动对齐的旋律控制的最具可比性的基线。我们还介绍了LyricEditBench,第一个保留旋律的歌词修改评估基准。代码、权重、基准和演示可在https://github.com/ASLP-lab/YingMusic-Singer上公开获得。
摘要:Regenerating singing voices with altered lyrics while preserving melody consistency remains challenging, as existing methods either offer limited controllability or require laborious manual alignment. We propose YingMusic-Singer, a fully diffusion-based model enabling melody-controllable singing voice synthesis with flexible lyric manipulation. The model takes three inputs: an optional timbre reference, a melody-providing singing clip, and modified lyrics, without manual alignment. Trained with curriculum learning and Group Relative Policy Optimization, YingMusic-Singer achieves stronger melody preservation and lyric adherence than Vevo2, the most comparable baseline supporting melody control without manual alignment. We also introduce LyricEditBench, the first benchmark for melody-preserving lyric modification evaluation. The code, weights, benchmark, and demos are publicly available at https://github.com/ASLP-lab/YingMusic-Singer.


【10】ACAVCaps: Enabling large-scale training for fine-grained and diverse audio understanding
标题:ACAVCaps:支持大规模训练,以实现细粒度和多样化的音频理解
链接:https://arxiv.org/abs/2603.24038

作者:Yadong Niu,Tianzi Wang,Heinrich Dinkel,Xingwei Sun,Jiahao Zhou,Gang Li,Jizhong Liu,Junbo Zhang,Jian Luan
备注:accepted by ICASSP 2026
摘要:通用音频理解是大型音频语言模型的基本目标,音频字幕是其开发的基石任务。然而,这一领域的进展受到现有数据集的阻碍,这些数据集缺乏训练真正通用模型所需的规模和描述粒度。为了解决这一差距,我们引入了ACAVCaps,一个新的大规模,细粒度,多方面的音频字幕数据集。ACAVCaps源自ACAV100M系列,使用多专家管道构建,从不同角度分析音频,包括语音,音乐和声学属性,然后通过大型语言模型合成为丰富,详细的描述。实验结果表明,在ACAVCaps上预训练的模型在各种下游任务上表现出比在其他领先字幕数据集上训练的模型更强的泛化能力。该数据集可在https://github.com/xiaomi-research/acavcaps上获得。
摘要:General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in this domain is hindered by existing datasets, which lack the scale and descriptive granularity required to train truly versatile models. To address this gap, we introduce ACAVCaps, a new large-scale, fine-grained, and multi-faceted audio captioning dataset. Derived from the ACAV100M collection, ACAVCaps is constructed using a multi-expert pipeline that analyzes audio from diverse perspectives-including speech, music, and acoustic properties-which are then synthesized into rich, detailed descriptions by a large language model. Experimental results demonstrate that models pre-trained on ACAVCaps exhibit substantially stronger generalization capabilities on various downstream tasks compared to those trained on other leading captioning datasets. The dataset is available at https://github.com/xiaomi-research/acavcaps.


【11】Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
标题:重新思考基于掩蔽预测的音频自我监督学习的掩蔽策略
链接:https://arxiv.org/abs/2603.23810

作者:Daisuke Niizumi,Daiki Takeuchi,Masahiro Yasuda,Binh Thien Nguyen,Noboru Harada,Nobutaka Ono
备注:6+1 pages, 2 figures, 3 tables, accepted at IJCNN 2026
摘要:自从引入Masked Autoencoder以来,已经探索了对掩蔽技术的各种改进。在本文中,我们重新思考掩蔽策略的音频表示学习使用掩蔽预测为基础的自监督学习(SSL)的一般音频频谱。虽然最近的知情掩蔽技术引起了人们的注意,我们观察到,它们会产生大量的计算开销。出于这种观察,我们提出了分散加权掩蔽(DWM),一个轻量级的掩蔽策略,利用频谱稀疏固有的音频内容的频率结构。我们的实验表明,在最近的SSL框架中常用的逆块掩蔽,提高了音频事件的理解性能,同时引入了一个权衡泛化。建议的DWM消除了这些限制和计算复杂性,导致一致的性能改进。这项工作为基于掩蔽预测的音频表示学习的掩蔽策略设计提供了实际指导。
摘要:Since the introduction of Masked Autoencoders, various improvements to masking techniques have been explored. In this paper, we rethink masking strategies for audio representation learning using masked prediction-based self-supervised learning (SSL) on general audio spectrograms. While recent informed masking techniques have attracted attention, we observe that they incur substantial computational overhead. Motivated by this observation, we propose dispersion-weighted masking (DWM), a lightweight masking strategy that leverages the spectral sparsity inherent in the frequency structure of audio content. Our experiments show that inverse block masking, commonly used in recent SSL frameworks, improves audio event understanding performance while introducing a trade-off in generalization. The proposed DWM alleviates these limitations and computational complexity, leading to consistent performance improvements. This work provides practical guidance on masking strategy design for masked prediction-based audio representation learning.


【12】Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
标题:使用Bayesian跟踪对深度空间选择性过滤器进行自回归引导以有效提取移动说话者
链接:https://arxiv.org/abs/2603.23723

作者:Jakob Kienegger,Timo Gerkmann
备注:This work has been submitted to the IEEE for possible publication
摘要:深度空间选择性滤波器通过用于已知方向的固定扬声器的具有实时能力的架构实现高质量增强。为了在动态场景中保持这种水平的性能,当只给出扬声器的初始方向时,精确但计算量小的跟踪算法变得必要。假设逐帧因果处理风格,时间反馈允许利用增强的语音信号来改善跟踪性能。在这项工作中,我们研究了将增强信号纳入轻量级跟踪算法和自回归引导深度空间滤波器的策略。我们提出的贝叶斯跟踪算法与任意深度空间滤波器兼容。为了在开发和评估过程中增加模拟轨迹的真实性,我们提出并发布了一个基于社会力量模型的新数据集。结果验证,自回归结合显着提高了我们的贝叶斯跟踪器的准确性,从而在没有或只有可忽略的计算开销增加卓越的增强。真实世界的录音补充了这些发现,并证明了我们的方法对看不见的,具有挑战性的声学条件的普遍性。
摘要:Deep spatially selective filters achieve high-quality enhancement with real-time capable architectures for stationary speakers of known directions. To retain this level of performance in dynamic scenarios when only the speakers' initial directions are given, accurate, yet computationally lightweight tracking algorithms become necessary. Assuming a frame-wise causal processing style, temporal feedback allows for leveraging the enhanced speech signal to improve tracking performance. In this work, we investigate strategies to incorporate the enhanced signal into lightweight tracking algorithms and autoregressively guide deep spatial filters. Our proposed Bayesian tracking algorithms are compatible with arbitrary deep spatial filters. To increase the realism of simulated trajectories during development and evaluation, we propose and publish a novel dataset based on the social force model. Results validate that the autoregressive incorporation significantly improves the accuracy of our Bayesian trackers, resulting in superior enhancement with none or only negligibly increased computational overhead. Real-world recordings complement these findings and demonstrate the generalizability of our methods to unseen, challenging acoustic conditions.


【13】Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
标题:Crab:多层对比监督提高动作和自然语音条件下的语音情感识别
链接:https://arxiv.org/abs/2603.23673

作者:Lucas H. Ueda,João G. T. Lima,Paula D. P. Costa
备注:IEEE Transactions on Affective Computing submission
摘要:语音情感识别(SER)在现实世界中的场景仍然具有挑战性,由于严重的类不平衡和自发的,自然的语音的流行。虽然最近的方法利用自监督学习(SSL)表示和语音和文本的多模态融合,但大多数现有方法仅在最终分类层应用监督,限制了中间表示的区分能力。在这项工作中,我们提出了螃蟹(对比表示和多模态对齐瓶颈),一个双峰跨模态Transformer架构,集成了语音表示从WavLM和文本表示从RoBERTA,连同一个新的\textit{多层对比监督}(MLCS)的战略。MLCS在网络的多个层注入多正对比学习信号,在整个模型中鼓励情感区分表示,而不会在推理时引入额外的参数。为了进一步解决数据不平衡的问题,我们在训练过程中采用了加权交叉熵。我们评估了三个基准数据集,涵盖不同程度的情感自然:IEMOCAP,MELD,和MSP播客2.0的方法。实验结果表明,Crab在所有数据集上的表现始终优于强单峰和多峰基线,在自然和高度不平衡的条件下尤其有很大的收益。这些发现突出了\textit{多层对比监督}作为SER的一般和强大的策略的有效性。官方实施可以在https://github.com/AI-Unicamp/Crab中找到。
摘要:Speech Emotion Recognition (SER) in real-world scenarios remains challenging due to severe class imbalance and the prevalence of spontaneous, natural speech. While recent approaches leverage self-supervised learning (SSL) representations and multimodal fusion of speech and text, most existing methods apply supervision only at the final classification layer, limiting the discriminative power of intermediate representations. In this work, we propose Crab (Contrastive Representation and Multimodal Aligned Bottleneck), a bimodal Cross-Modal Transformer architecture that integrates speech representations from WavLM and textual representations from RoBERTa, together with a novel \textit{Multi Layer Contrastive Supervision} (MLCS) strategy. MLCS injects multi-positive contrastive learning signals at multiple layers of the network, encouraging emotionally discriminative representations throughout the model without introducing additional parameters at inference time. To further address data imbalance, we adopt weighted cross-entropy during training. We evaluate the proposed approach on three benchmark datasets covering different degrees of emotional naturalness: IEMOCAP, MELD, and MSP-Podcast 2.0. Experimental results demonstrate that Crab consistently outperforms strong unimodal and multimodal baselines across all datasets, with particularly large gains under naturalistic and highly imbalanced conditions. These findings highlight the effectiveness of \textit{Multi Layer Contrastive Supervision} as a general and robust strategy for SER. Official implementation can be found in https://github.com/AI-Unicamp/Crab.


eess.AS音频处理


【1】YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
标题:YingMusic-Singer:可控的歌声合成,灵活的歌词操纵和无注释的旋律指导
链接:https://arxiv.org/abs/2603.24589

作者:Chunbo Hao,Junjie Zheng,Guobin Ma,Yuepeng Jiang,Huakang Chen,Wenjie Tian,Gongyu Chen,Zihao Chen,Lei Xie
摘要:在保持旋律一致性的同时用改变的歌词再生歌声仍然具有挑战性,因为现有方法要么提供有限的可控性,要么需要费力的手动对齐。我们提出了YingMusic-Singer,一个完全基于扩散的模型,使旋律可控的歌声合成与灵活的歌词操作。该模型有三个输入:一个可选的音色参考,一个提供旋律的演唱片段,以及修改后的歌词,无需手动调整。经过课程学习和集团相关政策优化的培训,YingMusic-Singer实现了比Vevo 2更强的旋律保留和歌词坚持,这是支持无需手动对齐的旋律控制的最具可比性的基线。我们还介绍了LyricEditBench,第一个保留旋律的歌词修改评估基准。代码、权重、基准和演示可在https://github.com/ASLP-lab/YingMusic-Singer上公开获得。
摘要:Regenerating singing voices with altered lyrics while preserving melody consistency remains challenging, as existing methods either offer limited controllability or require laborious manual alignment. We propose YingMusic-Singer, a fully diffusion-based model enabling melody-controllable singing voice synthesis with flexible lyric manipulation. The model takes three inputs: an optional timbre reference, a melody-providing singing clip, and modified lyrics, without manual alignment. Trained with curriculum learning and Group Relative Policy Optimization, YingMusic-Singer achieves stronger melody preservation and lyric adherence than Vevo2, the most comparable baseline supporting melody control without manual alignment. We also introduce LyricEditBench, the first benchmark for melody-preserving lyric modification evaluation. The code, weights, benchmark, and demos are publicly available at https://github.com/ASLP-lab/YingMusic-Singer.


【2】ArrayDPS-Refine: Generative Refinement of Discriminative Multi-Channel Speech Enhancement
标题:ArrayDPS-Refine:区分性多通道语音增强的生成式细化
链接:https://arxiv.org/abs/2603.24385

作者:Zhongweiyang Xu,Ashutosh Pandey,Juan Azcarreta,Zhaoheng Ni,Sanjeel Parekh,Buye Xu
备注:Accepted to ICASSP 2026
摘要:多通道语音增强的目的是从多通道噪声中恢复出纯净的语音。大多数深度学习方法都采用判别式训练,这可能导致基于回归的目标的非线性失真,特别是在具有挑战性的环境噪声条件下。受ArrayDPS用于无监督多通道源分离的启发,我们引入了ArrayDPS-Refine,这是一种旨在使用干净的语音扩散先验来增强判别模型的输出的方法。ArrayDPS-Refine无需训练,生成,并且与阵列无关。该算法首先从由判别模型产生的增强语音中估计出噪声空间协方差矩阵,然后利用估计出的噪声空间协方差矩阵进行扩散后验采样。这种方法允许直接细化任何判别模型的输出,而无需重新训练。我们的研究结果表明,ArrayDPS-Refine始终提高了各种判别模型的性能,包括最先进的波形和STFT域模型。音频演示在https://xzwy.github.io/ArrayDPSRefineDemo/上提供。
摘要:Multi-channel speech enhancement aims to recover clean speech from noisy multi-channel recordings. Most deep learning methods employ discriminative training, which can lead to non-linear distortions from regression-based objectives, especially under challenging environmental noise conditions. Inspired by ArrayDPS for unsupervised multi-channel source separation, we introduce ArrayDPS-Refine, a method designed to enhance the outputs of discriminative models using a clean speech diffusion prior. ArrayDPS-Refine is training-free, generative, and array-agnostic. It first estimates the noise spatial covariance matrix (SCM) from the enhanced speech produced by a discriminative model, then uses this estimated noise SCM for diffusion posterior sampling. This approach allows direct refinement of any discriminative model's output without retraining. Our results show that ArrayDPS-Refine consistently improves the performance of various discriminative models, including state-of-the-art waveform and STFT domain models. Audio demos are provided at https://xzwy.github.io/ArrayDPSRefineDemo/.


【3】How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools for Romanian
标题:Open TTC有多开放?罗马尼亚语开源RTS工具的实用评估
链接:https://arxiv.org/abs/2603.24116

作者:Teodora Răgman,Adrian Bogdan Stânea,Horia Cucu,Adriana Stan
备注:Published in IEEE Access
摘要:开源文本到语音(TTS)框架已经成为开发跨多种语言的语音合成系统的高度适应性平台。然而,它们的适用性并不统一-特别是当目标语言资源不足或计算资源受限时。在这项研究中,我们系统地评估了使用四种广泛采用的开源架构构建新型TTS模型的可行性:FastPitch,VITS,Grad-TTS和Matcha-TTS。我们的评估涵盖多个方面,包括定性方面,如安装的方便性,数据集准备和硬件要求,以及对罗马尼亚语合成质量的定量评估。我们采用客观指标和主观听力测试,以评估可懂度,扬声器相似性,和生成的语音的自然度。结果揭示了工具链设置,数据预处理和计算效率方面的重大挑战,这些挑战可能会阻碍在低资源环境中的采用。通过将分析建立在可重复的协议和可访问的评估标准上,这项工作旨在为最佳实践提供信息,并促进更具包容性,语言多样性的TTS开发。   复制本研究所需的所有信息(即代码和数据)都可以在我们的git存储库中找到:https://gitlab.com/opentts_ragman/OpenTTS
摘要:Open-source text-to-speech (TTS) frameworks have emerged as highly adaptable platforms for developing speech synthesis systems across a wide range of languages. However, their applicability is not uniform -- particularly when the target language is under-resourced or when computational resources are constrained. In this study, we systematically assess the feasibility of building novel TTS models using four widely adopted open-source architectures: FastPitch, VITS, Grad-TTS, and Matcha-TTS. Our evaluation spans multiple dimensions, including qualitative aspects such as ease of installation, dataset preparation, and hardware requirements, as well as quantitative assessments of synthesis quality for Romanian. We employ both objective metrics and subjective listening tests to evaluate intelligibility, speaker similarity, and naturalness of the generated speech. The results reveal significant challenges in tool chain setup, data preprocessing, and computational efficiency, which can hinder adoption in low-resource contexts. By grounding the analysis in reproducible protocols and accessible evaluation criteria, this work aims to inform best practices and promote more inclusive, language-diverse TTS development.   All information needed to reproduce this study (i.e. code and data) are available in our git repository: https://gitlab.com/opentts_ragman/OpenTTS


【4】Photogrammetry-Reconstructed 3D Head Meshes for Accessible Individual Head-Related Transfer Functions
标题:摄影测量重建的3D头部网格,可实现单个头部相关传递函数
链接:https://arxiv.org/abs/2603.24104

作者:Ludovic Pirard,Lorenzo Picinali,Katarina C. Poole
备注:Submitted to Acta Acustica Topical Issue - Spatial and binaural hearing: From neural processes to applications
摘要:个体头部相关传递函数(HRTF)对于精确的空间音频双耳渲染是必不可少的,但由于测量复杂性而仍然难以获得。本研究调查摄影测量重建(PR)的头部和耳朵网格,获得与消费者的硬件,可以提供一个实际有用的基线为个人的HRTF合成。使用SONITARY HRTF数据集,使用Apple的Object Capture API处理每个受试者的72个图像摄影测量捕获,以生成150个受试者的PR网格。Mesh2HRTF用于计算PR合成HRTF,通过数值评估、听觉模型和行为声音定位实验(N = 27),将其与测量的HRTF、高分辨率3D扫描衍生的HRTF、KEMAR和随机HRTF进行比较。PR合成HRTF保留ITD线索,但表现出增加ILD和光谱误差。听觉模型的预测和行为数据显示,象限误差率大大提高,海拔精度降低,更大的前后混淆比测得的HRTF,表现不如随机HRTF的感知指标。目前的摄影测量管道支持个人HRTF合成,但受到耳廓形态细节不足和包含单声道线索的准确个人HRTF所需的高频频谱保真度的限制。
摘要:Individual head-related transfer functions (HRTFs) are essential for accurate spatial audio binaural rendering but remain difficult to obtain due to measurement complexity. This study investigates whether photogrammetry-reconstructed (PR) head and ear meshes, acquired with consumer hardware, can provide a practically useful baseline for individual HRTF synthesis. Using the SONICOM HRTF dataset, 72-image photogrammetry captures per subject were processed with Apple's Object Capture API to generate PR meshes for 150 subjects. Mesh2HRTF was used to compute PR synthetic HRTFs, which were compared against measured HRTFs, high-resolution 3D scan-derived HRTFs, KEMAR, and random HRTFs through numerical evaluation, auditory models, and a behavioural sound localisation experiment (N = 27). PR synthetic HRTFs preserved ITD cues but exhibited increased ILD and spectral errors. Auditory-model predictions and behavioural data showed substantially higher quadrant error rates, reduced elevation accuracy, and greater front-back confusions than measured HRTFs, performing worse than random HRTFs on perceptual metrics. Current photogrammetry pipelines support individual HRTF synthesis but are limited by insufficient pinna morphology details and high-frequency spectral fidelity needed for accurate individual HRTFs containing monaural cues.


【5】ACAVCaps: Enabling large-scale training for fine-grained and diverse audio understanding
标题:ACAVCaps:支持大规模训练,以实现细粒度和多样化的音频理解
链接:https://arxiv.org/abs/2603.24038

作者:Yadong Niu,Tianzi Wang,Heinrich Dinkel,Xingwei Sun,Jiahao Zhou,Gang Li,Jizhong Liu,Junbo Zhang,Jian Luan
备注:accepted by ICASSP 2026
摘要:通用音频理解是大型音频语言模型的基本目标,音频字幕是其开发的基石任务。然而,这一领域的进展受到现有数据集的阻碍,这些数据集缺乏训练真正通用模型所需的规模和描述粒度。为了解决这一差距,我们引入了ACAVCaps,一个新的大规模,细粒度,多方面的音频字幕数据集。ACAVCaps源自ACAV100M系列,使用多专家管道构建,从不同角度分析音频,包括语音,音乐和声学属性,然后通过大型语言模型合成为丰富,详细的描述。实验结果表明,在ACAVCaps上预训练的模型在各种下游任务上表现出比在其他领先字幕数据集上训练的模型更强的泛化能力。该数据集可在https://github.com/xiaomi-research/acavcaps上获得。
摘要:General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in this domain is hindered by existing datasets, which lack the scale and descriptive granularity required to train truly versatile models. To address this gap, we introduce ACAVCaps, a new large-scale, fine-grained, and multi-faceted audio captioning dataset. Derived from the ACAV100M collection, ACAVCaps is constructed using a multi-expert pipeline that analyzes audio from diverse perspectives-including speech, music, and acoustic properties-which are then synthesized into rich, detailed descriptions by a large language model. Experimental results demonstrate that models pre-trained on ACAVCaps exhibit substantially stronger generalization capabilities on various downstream tasks compared to those trained on other leading captioning datasets. The dataset is available at https://github.com/xiaomi-research/acavcaps.


【6】Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
标题:重新思考基于掩蔽预测的音频自我监督学习的掩蔽策略
链接:https://arxiv.org/abs/2603.23810

作者:Daisuke Niizumi,Daiki Takeuchi,Masahiro Yasuda,Binh Thien Nguyen,Noboru Harada,Nobutaka Ono
备注:6+1 pages, 2 figures, 3 tables, accepted at IJCNN 2026
摘要:自从引入Masked Autoencoder以来,已经探索了对掩蔽技术的各种改进。在本文中,我们重新思考掩蔽策略的音频表示学习使用掩蔽预测为基础的自监督学习(SSL)的一般音频频谱。虽然最近的知情掩蔽技术引起了人们的注意,我们观察到,它们会产生大量的计算开销。出于这种观察,我们提出了分散加权掩蔽(DWM),一个轻量级的掩蔽策略,利用频谱稀疏固有的音频内容的频率结构。我们的实验表明,在最近的SSL框架中常用的逆块掩蔽,提高了音频事件的理解性能,同时引入了一个权衡泛化。建议的DWM消除了这些限制和计算复杂性,导致一致的性能改进。这项工作为基于掩蔽预测的音频表示学习的掩蔽策略设计提供了实际指导。
摘要:Since the introduction of Masked Autoencoders, various improvements to masking techniques have been explored. In this paper, we rethink masking strategies for audio representation learning using masked prediction-based self-supervised learning (SSL) on general audio spectrograms. While recent informed masking techniques have attracted attention, we observe that they incur substantial computational overhead. Motivated by this observation, we propose dispersion-weighted masking (DWM), a lightweight masking strategy that leverages the spectral sparsity inherent in the frequency structure of audio content. Our experiments show that inverse block masking, commonly used in recent SSL frameworks, improves audio event understanding performance while introducing a trade-off in generalization. The proposed DWM alleviates these limitations and computational complexity, leading to consistent performance improvements. This work provides practical guidance on masking strategy design for masked prediction-based audio representation learning.


【7】Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
标题:使用Bayesian跟踪对深度空间选择性过滤器进行自回归引导以有效提取移动说话者
链接:https://arxiv.org/abs/2603.23723

作者:Jakob Kienegger,Timo Gerkmann
备注:This work has been submitted to the IEEE for possible publication
摘要:深度空间选择性滤波器通过用于已知方向的固定扬声器的具有实时能力的架构实现高质量增强。为了在动态场景中保持这种水平的性能,当只给出扬声器的初始方向时,精确但计算量小的跟踪算法变得必要。假设逐帧因果处理风格,时间反馈允许利用增强的语音信号来改善跟踪性能。在这项工作中,我们研究了将增强信号纳入轻量级跟踪算法和自回归引导深度空间滤波器的策略。我们提出的贝叶斯跟踪算法与任意深度空间滤波器兼容。为了在开发和评估过程中增加模拟轨迹的真实性,我们提出并发布了一个基于社会力量模型的新数据集。结果验证,自回归结合显着提高了我们的贝叶斯跟踪器的准确性,从而在没有或只有可忽略的计算开销增加卓越的增强。真实世界的录音补充了这些发现,并证明了我们的方法对看不见的,具有挑战性的声学条件的普遍性。
摘要:Deep spatially selective filters achieve high-quality enhancement with real-time capable architectures for stationary speakers of known directions. To retain this level of performance in dynamic scenarios when only the speakers' initial directions are given, accurate, yet computationally lightweight tracking algorithms become necessary. Assuming a frame-wise causal processing style, temporal feedback allows for leveraging the enhanced speech signal to improve tracking performance. In this work, we investigate strategies to incorporate the enhanced signal into lightweight tracking algorithms and autoregressively guide deep spatial filters. Our proposed Bayesian tracking algorithms are compatible with arbitrary deep spatial filters. To increase the realism of simulated trajectories during development and evaluation, we propose and publish a novel dataset based on the social force model. Results validate that the autoregressive incorporation significantly improves the accuracy of our Bayesian trackers, resulting in superior enhancement with none or only negligibly increased computational overhead. Real-world recordings complement these findings and demonstrate the generalizability of our methods to unseen, challenging acoustic conditions.


【8】Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
标题:Crab:多层对比监督提高动作和自然语音条件下的语音情感识别
链接:https://arxiv.org/abs/2603.23673

作者:Lucas H. Ueda,João G. T. Lima,Paula D. P. Costa
备注:IEEE Transactions on Affective Computing submission
摘要:语音情感识别(SER)在现实世界中的场景仍然具有挑战性,由于严重的类不平衡和自发的,自然的语音的流行。虽然最近的方法利用自监督学习(SSL)表示和语音和文本的多模态融合,但大多数现有方法仅在最终分类层应用监督,限制了中间表示的区分能力。在这项工作中,我们提出了螃蟹(对比表示和多模态对齐瓶颈),一个双峰跨模态Transformer架构,集成了语音表示从WavLM和文本表示从RoBERTA,连同一个新的\textit{多层对比监督}(MLCS)的战略。MLCS在网络的多个层注入多正对比学习信号,在整个模型中鼓励情感区分表示,而不会在推理时引入额外的参数。为了进一步解决数据不平衡的问题,我们在训练过程中采用了加权交叉熵。我们评估了三个基准数据集,涵盖不同程度的情感自然:IEMOCAP,MELD,和MSP播客2.0的方法。实验结果表明,Crab在所有数据集上的表现始终优于强单峰和多峰基线,在自然和高度不平衡的条件下尤其有很大的收益。这些发现突出了\textit{多层对比监督}作为SER的一般和强大的策略的有效性。官方实施可以在https://github.com/AI-Unicamp/Crab中找到。
摘要:Speech Emotion Recognition (SER) in real-world scenarios remains challenging due to severe class imbalance and the prevalence of spontaneous, natural speech. While recent approaches leverage self-supervised learning (SSL) representations and multimodal fusion of speech and text, most existing methods apply supervision only at the final classification layer, limiting the discriminative power of intermediate representations. In this work, we propose Crab (Contrastive Representation and Multimodal Aligned Bottleneck), a bimodal Cross-Modal Transformer architecture that integrates speech representations from WavLM and textual representations from RoBERTa, together with a novel \textit{Multi Layer Contrastive Supervision} (MLCS) strategy. MLCS injects multi-positive contrastive learning signals at multiple layers of the network, encouraging emotionally discriminative representations throughout the model without introducing additional parameters at inference time. To further address data imbalance, we adopt weighted cross-entropy during training. We evaluate the proposed approach on three benchmark datasets covering different degrees of emotional naturalness: IEMOCAP, MELD, and MSP-Podcast 2.0. Experimental results demonstrate that Crab consistently outperforms strong unimodal and multimodal baselines across all datasets, with particularly large gains under naturalistic and highly imbalanced conditions. These findings highlight the effectiveness of \textit{Multi Layer Contrastive Supervision} as a general and robust strategy for SER. Official implementation can be found in https://github.com/AI-Unicamp/Crab.


【9】Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
标题:口语对话系统中的语义感知中断检测:基准、指标和模型
链接:https://arxiv.org/abs/2603.24144

作者:Kangxiang Xia,Bingshen Mu,Xian Shi,Jin Xu,Lei Xie
备注:Accepted by ICME 2026
摘要:在口语对话系统(SDS)中实现自然的全双工交互仍然是一个挑战,因为难以准确地检测用户中断。当前的解决方案在误解反向通道的“触发式”基于VAD的方法和表现出不可接受的响应延迟的强大端到端模型之间存在两极分化。此外,缺乏现实世界的基准和整体衡量标准也阻碍了这一领域的进展。本文提出了一个全面的框架,以克服这些限制。我们首先介绍SID-Bench,这是第一个完全基于真实人类对话的语义感知中断检测基准。为了提供一个严格的评估的响应性鲁棒性的权衡,我们提出了平均惩罚时间(APT)度量,它分配一个时间成本的误报和延迟响应。在此框架的基础上,我们设计了一个基于LLM的检测模型,通过一种新的训练范式进行优化,以捕获意图的微妙语义线索。实验结果表明,我们的模型显着优于主流基线,实现了近三倍的APT减少。通过成功解决速度和稳定性之间的长期紧张关系,我们的工作为SDS中的智能中断处理建立了一个新的最先进的技术。为了方便未来的研究,SID-Bench和相关代码可在www.example.com上获得。
摘要:Achieving natural full-duplex interaction in spoken dialogue systems (SDS) remains a challenge due to the difficulty of accurately detecting user interruptions. Current solutions are polarized between "trigger-happy" VAD-based methods that misinterpret backchannels and robust end-to-end models that exhibit unacceptable response delays. Moreover, the absence of real-world benchmarks and holistic metrics hinders progress in the field. This paper presents a comprehensive frame-work to overcome these limitations. We first introduce SID-Bench, the first benchmark for semantic-aware interruption detection built entirely from real-world human dialogues. To provide a rigorous assessment of the responsiveness-robustness trade-off, we propose the Average Penalty Time (APT) metric, which assigns a temporal cost to both false alarms and late responses. Building on this framework, we design an LLM-based detection model optimized through a novel training paradigm to capture subtle semantic cues of intent. Experimental results show that our model significantly outperforms mainstream baselines, achieving a nearly threefold reduction in APT. By successfully resolving the long-standing tension between speed and stability, our work establishes a new state-of-the-art for intelligent interruption handling in SDS. To facilitate future research, SID-Bench and the associated code are available at: https://github.com/xkx-hub/SID-bench.


【10】Echoes: A semantically-aligned music deepfake detection dataset
标题:Echoes:一个语义对齐的音乐Deepfake检测数据集
链接:https://arxiv.org/abs/2603.23667

作者:Octavian Pascu,Dan Oneata,Horia Cucu,Nicolas M. Muller
摘要:我们介绍了Echoes,这是一个用于音乐deepfake检测的新数据集,旨在在现实和提供商多样化的条件下训练和基准检测器。Echoes包含3,577首曲目(110小时的音频),跨越多种流派(流行,摇滚,电子),并包括由十种流行的AI音乐生成系统生成的内容。为了防止捷径学习和促进鲁棒的泛化,数据集被故意构造成具有挑战性,在欺骗音频和真实参考之间强制执行语义级对齐。这种对齐是通过直接在真正的波形或歌曲描述符上调节生成的音频样本来实现的。我们使用最先进的Wav 2 Vec 2 XLS-R 2B表示,在跨数据集设置中对三个现有的AI生成的音乐数据集进行了回声评估。结果表明:(i)Echoes是最难的域内数据集;(ii)在现有数据集上训练的检测器向Echoes的转移很差;(iii)在Echoes上训练产生最强的泛化性能。这些研究结果表明,供应商的多样性和语义对齐有助于学习更多的可转移检测线索。
摘要:We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provider-diverse conditions. Echoes comprises 3,577 tracks (110 hours of audio) spanning multiple genres (pop, rock, electronic), and includes content generated by ten popular AI music generation systems. To prevent shortcut learning and promote robust generalization, the dataset is deliberately constructed to be challenging, enforcing semantic-level alignment between spoofed audio and bona fide references. This alignment is achieved by conditioning generated audio samples directly on bona-fide waveforms or song descriptors. We evaluate Echoes in a cross-dataset setting against three existing AI-generated music datasets using state-of-the-art Wav2Vec2 XLS-R 2B representations. Results show that (i) Echoes is the hardest in-domain dataset; (ii) detectors trained on existing datasets transfer poorly to Echoes; (iii) training on Echoes yields the strongest generalization performance. These findings suggest that provider diversity and semantic alignment help learn more transferable detection cues.


机器翻译由腾讯交互翻译提供,仅供参考