微信公众号:arXiv_Daily
cs.SD语音
【1】Masked Autoencoders as Universal Speech Enhancer
标题:屏蔽自动编码器作为通用语音增强器
链接:https://arxiv.org/abs/2602.02413
摘要:有监督的语音增强方法已经非常成功。然而,在实际场景中,缺乏干净的语音,并且期望提供相当的增强性能并且可以应用于其他语音相关下游应用的基于自监督学习(SSL)的语音增强方法。在这项工作中,我们开发了一种基于掩蔽自动编码器的通用语音增强器,该增强器对影响语音的失真类型不可知,可以同时处理多个失真,并且以自我监督的方式进行训练。增强堆栈向噪声输入数据添加进一步的失真。掩蔽的自动编码器模型在预训练期间学习去除添加的失真以及重建频谱图的掩蔽区域。然后,预训练的嵌入被用于微调模型,这些模型是在特定下游任务的少量配对数据上训练的。我们评估了用于去噪和去混响下游任务的预训练特征。我们探索了预训练增强堆栈中的不同增强(如单个或多个扬声器)以及不同噪声输入特征表示(如$log1p$压缩)对预训练嵌入和下游微调增强性能的影响。我们表明,所提出的方法不仅优于基线,但也达到了最先进的性能域内和域外的评估数据集。
摘要:Supervised speech enhancement methods have been very successful. However, in practical scenarios, there is a lack of clean speech, and self-supervised learning-based (SSL) speech enhancement methods that offer comparable enhancement performance and can be applied to other speech-related downstream applications are desired. In this work, we develop a masked autoencoder based universal speech enhancer that is agnostic to the type of distortion affecting speech, can handle multiple distortions simultaneously, and is trained in a self-supervised manner. An augmentation stack adds further distortions to the noisy input data. The masked autoencoder model learns to remove the added distortions along with reconstructing the masked regions of the spectrogram during pre-training. The pre-trained embeddings are then used by fine-tuning models trained on a small amount of paired data for specific downstream tasks. We evaluate the pre-trained features for denoising and dereverberation downstream tasks. We explore different augmentations (like single or multi-speaker) in the pre-training augmentation stack and the effect of different noisy input feature representations (like $log1p$ compression) on pre-trained embeddings and downstream fine-tuning enhancement performance. We show that the proposed method not only outperforms the baseline but also achieves state-of-the-art performance for both in-domain and out-of-domain evaluation datasets.
【2】DFKI-Speech System for WildSpoof Challenge: A robust framework for SASV In-the-Wild
标题:用于WildSpoof挑战的DFKI语音系统:SASV野外的强大框架
链接:https://arxiv.org/abs/2602.02286
摘要:本文介绍了DFKI语音系统开发的WildSpoof挑战赛下的欺骗感知自动说话人确认(SASV)的轨道。我们提出了一个强大的SASV框架,其中欺骗检测器和说话人验证(SV)网络协同工作。欺骗检测器采用自监督语音嵌入提取器作为前端,结合最先进的图神经网络后端。此外,基于前3层的混合专家(MoE)被用来融合高级别和低级别的功能,有效的欺骗话语检测。对于说话人验证,我们采用了一种低复杂度的卷积神经网络,该网络在多个尺度上融合了2D和1D特征,并使用SphereFace损失进行了训练。此外,对比循环损失被应用于每个训练批次中的自适应权重正对和负对,使网络能够更好地区分困难和容易的样本对。最后,采用基于固定冒名顶替者队列的AS Norm分数归一化和模型集成方法进一步提高了说话人确认系统的鉴别能力。
摘要:This paper presents the DFKI-Speech system developed for the WildSpoof Challenge under the Spoofing aware Automatic Speaker Verification (SASV) track. We propose a robust SASV framework in which a spoofing detector and a speaker verification (SV) network operate in tandem. The spoofing detector employs a self-supervised speech embedding extractor as the frontend, combined with a state-of-the-art graph neural network backend. In addition, a top-3 layer based mixture-of-experts (MoE) is used to fuse high-level and low-level features for effective spoofed utterance detection. For speaker verification, we adapt a low-complexity convolutional neural network that fuses 2D and 1D features at multiple scales, trained with the SphereFace loss. Additionally, contrastive circle loss is applied to adaptively weight positive and negative pairs within each training batch, enabling the network to better distinguish between hard and easy sample pairs. Finally, fixed imposter cohort based AS Norm score normalization and model ensembling are used to further enhance the discriminative capability of the speaker verification system.
【3】Evaluating Acoustic Data Transmission Schemes for Ad-Hoc Communication Between Nearby Smart Devices
标题:评估附近智能设备之间自组织通信的声学数据传输方案
链接:https://arxiv.org/abs/2602.02249
备注:31 pages, 9 figures, the dataset is available at https://doi.org/10.5281/zenodo.17661991
摘要:声学数据传输通过利用智能手机和物联网设备中无处不在的扬声器和麦克风,为蓝牙和NFC提供了一种引人注目的替代方案。然而,这一领域的大多数研究都依赖于模拟或有限的设备测试,这使得所提出的方案的真实可靠性难以评估。我们系统地回顾了31项针对商品设备的声学通信研究,发现没有一项提供可访问的源代码。在与作者联系并重新实现三个有前途的方案后,我们组装了一个由八个代表性声学通信系统组成的测试平台。使用超过11000智能手机传输在现实的室内环境和消声室,我们提供了一个系统的和可重复的方法来评估这些计划在现实条件下的可靠性和可推广性。我们的研究结果表明,许多现有的计划面临的挑战,在实际使用中,主要是由于严重的多径传播室内和不同的音频特性跨设备模型。为了支持未来的研究并促进更强大的评估,我们发布了我们的重新实现以及第一个真实世界声学传输的综合数据集。总的来说,我们的研究结果强调了严格的设备测试的重要性,并强调了强大的设计策略的必要性,以弥合模拟结果和可靠的物联网部署之间的差距。
摘要:Acoustic data transmission offers a compelling alternative to Bluetooth and NFC by leveraging the ubiquitous speakers and microphones in smartphones and IoT devices. However, most research in this field relies on simulations or limited on-device testing, which makes the real-world reliability of proposed schemes difficult to assess. We systematically reviewed 31 acoustic communication studies for commodity devices and found that none provided accessible source code. After contacting authors and re-implementing three promising schemes, we assembled a testbed of eight representative acoustic communication systems. Using over 11000 smartphone transmissions in both realistic indoor environments and an anechoic chamber, we provide a systematic and repeatable methodology for evaluating the reliability and generalizability of these schemes under real-world conditions. Our results show that many existing schemes face challenges in practical usage, largely due to severe multipath propagation indoors and varying audio characteristics across device models. To support future research and foster more robust evaluations, we release our re-implementations alongside the first comprehensive dataset of real-world acoustic transmissions. Overall, our findings highlight the importance of rigorous on-device testing and underscore the need for robust design strategies to bridge the gap between simulation results and reliable IoT deployments.
【4】LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency
标题:LipSody:具有增强的韵律一致性的唇语音合成
链接:https://arxiv.org/abs/2602.01908
备注:This paper has been accepted to ICASSP 2026
摘要:唇到语音合成旨在通过从嘴唇运动重建语言内容来直接从无声的面部视频生成语音音频,在音频信号不可用或退化的情况下提供有价值的应用。虽然最近的扩散为基础的模型,如LipVoicer在重建语言内容表现出令人印象深刻的性能,他们往往缺乏韵律的一致性。在这项工作中,我们提出了LipSody,唇语音框架增强韵律一致性。LipSody引入了一种韵律指导策略,该策略利用了三个互补的线索:从面部图像中提取的说话人身份,从嘴唇运动中获得的语言内容,以及从面部视频中推断的情感背景。实验结果表明,LipSody大大提高了韵律相关的指标,包括全球和本地的音高偏差,能量一致性,说话人相似性,与以前的方法相比。
摘要:Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded. While recent diffusion-based models such as LipVoicer have demonstrated impressive performance in reconstructing linguistic content, they often lack prosodic consistency. In this work, we propose LipSody, a lip-to-speech framework enhanced for prosody consistency. LipSody introduces a prosody-guiding strategy that leverages three complementary cues: speaker identity extracted from facial images, linguistic content derived from lip movements, and emotional context inferred from face video. Experimental results demonstrate that LipSody substantially improves prosody-related metrics, including global and local pitch deviations, energy consistency, and speaker similarity, compared to prior approaches.
【5】Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
标题:无声音说话:仅使用面部输入的多扬声器无声语音
链接:https://arxiv.org/abs/2602.01879
备注:This paper was presented at ICASSP 2025
摘要:在本文中,我们介绍了一种新的框架,用于生成多扬声器语音,而不依赖于任何可听输入。我们的方法利用无声的肌电图(EMG)信号来捕获语言内容,而面部图像用于与目标说话人的声音身份相匹配。值得注意的是,我们提出了一个音高解开的内容嵌入,增强了从EMG信号中提取的语言内容。大量的分析表明,我们的方法可以产生多扬声器语音没有任何可听输入,并证实了所提出的音高解纠缠方法的有效性。
摘要:In this paper, we introduce a novel framework for generating multi-speaker speech without relying on any audible inputs. Our approach leverages silent electromyography (EMG) signals to capture linguistic content, while facial images are used to match with the vocal identity of the target speaker. Notably, we present a pitch-disentangled content embedding that enhances the extraction of linguistic content from EMG signals. Extensive analysis demonstrates that our method can generate multi-speaker speech without any audible inputs and confirms the effectiveness of the proposed pitch-disentanglement approach.
【6】ParaGSE: Parallel Generative Speech Enhancement with Group-Vector-Quantization-based Neural Speech Codec
标题:ParaGSE:采用基于群向量化的神经语音编解码器的并行生成语音增强
链接:https://arxiv.org/abs/2602.01793
备注:Accepted by ICASSP 2026
摘要:最近,生成式语音增强已经获得了相当大的兴趣,然而,现有的方法受到过度的复杂性,有限的效率,和次优的语音质量。为了克服这些挑战,本文提出了一种新的并行生成语音增强(ParaGSE)框架,利用组矢量量化(GVQ)为基础的神经语音编解码器。基于GVQ的编解码器采用单独的VQ来产生相互独立的令牌,从而在ParaGSE中实现高效的并行令牌预测。具体而言,ParaGSE利用基于GVQ的编解码器将降级语音编码成不同的令牌,通过以降级频谱特征为条件的并行分支预测相应的干净令牌,并最终通过编解码器解码器重建干净语音。实验结果表明,ParaGSE一贯产生优越的增强语音相比,歧视和生成基线,在广泛的失真,包括噪声,混响,频带限制,及其混合物。此外,授权并行计算的令牌预测,ParaGSE实现了约1.5倍的提高,在CPU上的生成效率相比,串行生成语音增强方法。
摘要:Recently, generative speech enhancement has garnered considerable interest; however, existing approaches are hindered by excessive complexity, limited efficiency, and suboptimal speech quality. To overcome these challenges, this paper proposes a novel parallel generative speech enhancement (ParaGSE) framework that leverages a group vector quantization (GVQ)-based neural speech codec. The GVQ-based codec adopts separate VQs to produce mutually independent tokens, enabling efficient parallel token prediction in ParaGSE. Specifically, ParaGSE leverages the GVQ-based codec to encode degraded speech into distinct tokens, predicts the corresponding clean tokens through parallel branches conditioned on degraded spectral features, and ultimately reconstructs clean speech via the codec decoder. Experimental results demonstrate that ParaGSE consistently produces superior enhanced speech compared to both discriminative and generative baselines, under a wide range of distortions including noise, reverberation, band-limiting, and their mixtures. Furthermore, empowered by parallel computation in token prediction, ParaGSE attains about a 1.5-fold improvement in generation efficiency on CPU compared with serial generative speech enhancement approaches.
【7】Voting-based Pitch Estimation with Temporal and Frequential Alignment and Correlation Aware Selection
标题:具有时间和频率对齐以及相关感知选择的基于投票的音调估计
链接:https://arxiv.org/abs/2602.01727
备注:Accepted for ICASSP 2026
摘要:投票法,一种用于基频估计的集成方法,经验上以其鲁棒性而闻名,但缺乏深入的研究。本文对该技术进行了原理性分析和改进。首先,我们提供了一个理论基础,其有效性,解释的误差方差减少基频估计和调用孔多塞的陪审团定理有声/无声检测精度。为了解决其实际局限性,我们提出了两个关键的改进:1)预投票对齐过程,以纠正估计量之间的时间和频率偏差,以及2)贪婪算法,以选择一个紧凑而有效的估计量子集的基础上误差相关性。在语音、歌唱和音乐的不同数据集上进行的实验表明,我们提出的对齐方法在干净条件下优于单个最先进的估计器,并在嘈杂环境中保持鲁棒的有声/无声检测。
摘要:The voting method, an ensemble approach for fundamental frequency estimation, is empirically known for its robustness but lacks thorough investigation. This paper provides a principled analysis and improvement of this technique. First, we offer a theoretical basis for its effectiveness, explaining the error variance reduction for fundamental frequency estimation and invoking Condorcet's jury theorem for voiced/unvoiced detection accuracy. To address its practical limitations, we propose two key improvements: 1) a pre-voting alignment procedure to correct temporal and frequential biases among estimators, and 2) a greedy algorithm to select a compact yet effective subset of estimators based on error correlation. Experiments on a diverse dataset of speech, singing, and music show that our proposed method with alignment outperforms individual state-of-the-art estimators in clean conditions and maintains robust voiced/unvoiced detection in noisy environments.
【8】Membership Inference Attack Against Music Diffusion Models via Generative Manifold Perturbation
标题:通过生成型Manifold扰动对音乐扩散模型的隶属度推理攻击
链接:https://arxiv.org/abs/2602.01645
摘要:成员推理攻击(MIA)测试特定的音频片段是否用于训练模型,使其成为审计生成音乐模型是否符合版权的关键工具。然而,基于损失的信号(例如,重建误差)在实践中与人类感知弱一致,在取证所需的低假阳性率(FPR)下产生差的可分离性。我们提出了潜在稳定性对抗探针(LSA探针),白盒方法,测量的几何属性的反向扩散:最小的时间归一化的扰动预算需要跨越一个固定的感知退化阈值在中间扩散状态。我们表明,培训成员,居住在更稳定的地区,表现出显着更高的退化成本。
摘要:Membership inference attacks (MIAs) test whether a specific audio clip was used to train a model, making them a key tool for auditing generative music models for copyright compliance. However, loss-based signals (e.g., reconstruction error) are weakly aligned with human perception in practice, yielding poor separability at the low false-positive rates (FPRs) required for forensics. We propose the Latent Stability Adversarial Probe (LSA-Probe), a white-box method that measures a geometric property of the reverse diffusion: the minimal time-normalized perturbation budget needed to cross a fixed perceptual degradation threshold at an intermediate diffusion state. We show that training members, residing in more stable regions, exhibit a significantly higher degradation cost.
【9】Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
标题:用于语音情感识别的大型音频语言模型中的注意力加权中心核对齐知识提取
链接:https://arxiv.org/abs/2602.01547
备注:Accepted to 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)
摘要:大型音频语言模型(LALM)的出现促进了语音情感识别(SER)的发展,但其规模限制了在资源受限环境中的部署。虽然知识蒸馏对于LALM压缩是有效的,但是现有方法在提取跨模态投影模块(投影仪)方面仍然探索不足,并且由于特征维度的差异而经常与对齐斗争。我们提出PL-Distill,一个KD框架,它结合了投影仪级蒸馏(PDist)来对齐音频嵌入和逻辑级蒸馏(LDist)来对齐输出逻辑。PDist介绍了注意力加权中心内核对齐,这是我们提出的一种新方法,可以突出重要的时间步长并解决维度不匹配问题。与此同时,LDist最大限度地减少了教师和学生之间的Kullback-Leibler分歧从音频和文本形式的逻辑。在IEMOCAP、RAVDESS和SAVEE上,PL-Distill将8.4B参数的教师压缩为紧凑的1.1B参数的学生,在所有指标上始终优于教师、最先进的预训练模型和其他KD基线。
摘要:The emergence of Large Audio-Language Models (LALMs) has advanced Speech Emotion Recognition (SER), but their size limits deployment in resource-constrained environments. While Knowledge Distillation is effective for LALM compression, existing methods remain underexplored in distilling the cross-modal projection module (Projector), and often struggle with alignment due to differences in feature dimensions. We propose PL-Distill, a KD framework that combines Projector-Level Distillation (PDist) to align audio embeddings and Logits-Level Distillation (LDist) to align output logits. PDist introduces Attention-weighted Centered Kernel Alignment, a novel approach we propose to highlight important time steps and address dimension mismatches. Meanwhile, LDist minimizes the Kullback-Leibler divergence between teacher and student logits from audio and text modalities. On IEMOCAP, RAVDESS, and SAVEE, PL-Distill compresses an 8.4B-parameter teacher to a compact 1.1B-parameter student, consistently outperforming the teacher, state-of-the-art pretrained models, and other KD baselines across all metrics.
【10】Causally Disentangled Contrastive Learning for Multilingual Speaker Embeddings
标题:多语言说话人嵌入的因果分离对比学习
链接:https://arxiv.org/abs/2602.01363
摘要:自监督说话人嵌入被广泛用于说话人验证系统,但先前的工作表明,他们往往编码敏感的人口统计属性,提高公平性和隐私问题。本文研究了人口统计信息,特别是性别,年龄和口音,在SimCLR训练的说话人嵌入中存在的程度,以及是否可以在不严重降低说话人验证性能的情况下减轻这种泄漏。我们研究了两种去偏策略:通过梯度反转进行对抗性训练,以及明确分离人口统计信息和剩余信息的因果瓶颈架构。使用线性和非线性探测分类器来量化人口统计泄漏,而使用ROC-AUC和EER来评估说话人验证性能。我们的研究结果表明,性别信息是强烈的,线性编码的基线嵌入,而年龄和口音较弱,主要是非线性表示。对抗性去偏见减少了性别泄漏,但对年龄和口音的影响有限,并与验证准确性进行了明确的权衡。因果瓶颈进一步抑制了人口统计信息,特别是在残差表示中,但会导致性能大幅下降。这些研究结果强调了在减轻自我监督的说话人嵌入中的人口泄漏方面的根本局限性,并澄清了当前去偏方法中固有的权衡。
摘要:Self-supervised speaker embeddings are widely used in speaker verification systems, but prior work has shown that they often encode sensitive demographic attributes, raising fairness and privacy concerns. This paper investigates the extent to which demographic information, specifically gender, age, and accent, is present in SimCLR-trained speaker embeddings and whether such leakage can be mitigated without severely degrading speaker verification performance. We study two debiasing strategies: adversarial training through gradient reversal and a causal bottleneck architecture that explicitly separates demographic and residual information. Demographic leakage is quantified using both linear and nonlinear probing classifiers, while speaker verification performance is evaluated using ROC-AUC and EER. Our results show that gender information is strongly and linearly encoded in baseline embeddings, whereas age and accent are weaker and primarily nonlinearly represented. Adversarial debiasing reduces gender leakage but has limited effect on age and accent and introduces a clear trade-off with verification accuracy. The causal bottleneck further suppresses demographic information, particularly in the residual representation, but incurs substantial performance degradation. These findings highlight fundamental limitations in mitigating demographic leakage in self-supervised speaker embeddings and clarify the trade-offs inherent in current debiasing approaches.
【11】TLDiffGAN: A Latent Diffusion-GAN Framework with Temporal Information Fusion for Anomalous Sound Detection
标题:TLDifGAN:一个具有时间信息融合的潜在扩散-GAN框架,用于异常声音检测
链接:https://arxiv.org/abs/2602.01060
备注:Accepted by ICASSP 2026
摘要:现有的生成模型的无监督异常声音检测是有限的,他们无法完全捕捉正常声音的复杂特征分布,而在这一领域的强大的扩散模型的潜力仍然在很大程度上未被开发。为了应对这一挑战,我们提出了一个新的框架,TLDiffGAN,它由两个互补的分支。一个分支将潜在扩散模型并入GAN生成器中进行对抗训练,从而使训练器的任务更具挑战性并提高生成样本的质量。另一个分支利用预训练的音频模型编码器直接从原始音频波形中提取特征用于辅助辨别。该框架有效地从原始音频和Mel频谱图中捕获正常声音的特征表示。此外,我们引入了TMixup声谱图增强技术,以提高敏感性,往往被忽视的微妙和局部的时间模式。在DCASE 2020 Challenge Task 2数据集上的大量实验证明了TLDiffGAN的优越检测性能,以及其在异常时频定位方面的强大能力。
摘要:Existing generative models for unsupervised anomalous sound detection are limited by their inability to fully capture the complex feature distribution of normal sounds, while the potential of powerful diffusion models in this domain remains largely unexplored. To address this challenge, we propose a novel framework, TLDiffGAN, which consists of two complementary branches. One branch incorporates a latent diffusion model into the GAN generator for adversarial training, thereby making the discriminator's task more challenging and improving the quality of generated samples. The other branch leverages pretrained audio model encoders to extract features directly from raw audio waveforms for auxiliary discrimination. This framework effectively captures feature representations of normal sounds from both raw audio and Mel spectrograms. Moreover, we introduce a TMixup spectrogram augmentation technique to enhance sensitivity to subtle and localized temporal patterns that are often overlooked. Extensive experiments on the DCASE 2020 Challenge Task 2 dataset demonstrate the superior detection performance of TLDiffGAN, as well as its strong capability in anomalous time-frequency localization.
【12】HierCon: Hierarchical Contrastive Attention for Audio Deepfake Detection
标题:HierCon:音频Deepfake检测的分层对比注意力
链接:https://arxiv.org/abs/2602.01032
备注:Proceedings of The Web Conference 2026 (WWW'26), short track
摘要:现代TTS和语音转换系统生成的音频深度伪造越来越难以与真实语音区分开来,这给安全和在线信任带来了严重风险。虽然最先进的自监督模型提供了丰富的多层表示,但现有的检测器独立地处理层,忽略了对识别合成伪影至关重要的时间和层次依赖性。我们提出了HierCon,一个分层注意力框架,结合了基于边缘的对比学习,该框架对时间帧,相邻层和层组之间的依赖关系进行建模,同时鼓励域不变嵌入。在ASVspoof 2021 DF和In-the-Wild数据集上进行评估,我们的方法实现了最先进的性能(1.93%和6.87%EER),比独立层加权分别提高了36.6%和22.5%。结果和注意力可视化证实,分层建模增强了跨域生成技术和记录条件的概括。
摘要:Audio deepfakes generated by modern TTS and voice conversion systems are increasingly difficult to distinguish from real speech, raising serious risks for security and online trust. While state-of-the-art self-supervised models provide rich multi-layer representations, existing detectors treat layers independently and overlook temporal and hierarchical dependencies critical for identifying synthetic artefacts. We propose HierCon, a hierarchical layer attention framework combined with margin-based contrastive learning that models dependencies across temporal frames, neighbouring layers, and layer groups, while encouraging domain-invariant embeddings. Evaluated on ASVspoof 2021 DF and In-the-Wild datasets, our method achieves state-of-the-art performance (1.93% and 6.87% EER), improving over independent layer weighting by 36.6% and 22.5% respectively. The results and attention visualisations confirm that hierarchical modelling enhances generalisation to cross-domain generation techniques and recording conditions.
【13】Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
标题:耳朵里的偏见:评估跨语言、人口和位置变化的音频语言模型的敏感性
链接:https://arxiv.org/abs/2602.01030
备注:Accepted as a long findings paper at EACL 2026
摘要:这项工作提出了第一个系统的调查,在多语言MLLM的言语偏见。我们构建并发布了BiasInEar数据集,这是一个基于Global MMLU Lite的语音增强基准测试,涵盖英语,中文和韩语,按性别和口音平衡,总计70.8小时(约4,249分钟)的语音,包含11,200个问题。使用四个互补的指标(准确性,熵,APES和Fleiss的$κ$),我们评估了九个代表性的模型在语言(语言和口音),人口(性别)和结构(选项顺序)扰动。我们的研究结果表明,MLLM是相对强大的人口因素,但高度敏感的语言和选项顺序,这表明,讲话可以放大现有的结构性偏见。此外,架构设计和推理策略大大影响跨语言的鲁棒性。总的来说,本研究建立了一个统一的框架,用于评估语音集成LLM的公平性和鲁棒性,弥合了基于文本和基于语音的评估之间的差距。这些资源可以在https://github.com/ntunlplab/BiasInEar上找到。
摘要:This work presents the first systematic investigation of speech bias in multilingual MLLMs. We construct and release the BiasInEar dataset, a speech-augmented benchmark based on Global MMLU Lite, spanning English, Chinese, and Korean, balanced by gender and accent, and totaling 70.8 hours ($\approx$4,249 minutes) of speech with 11,200 questions. Using four complementary metrics (accuracy, entropy, APES, and Fleiss' $κ$), we evaluate nine representative models under linguistic (language and accent), demographic (gender), and structural (option order) perturbations. Our findings reveal that MLLMs are relatively robust to demographic factors but highly sensitive to language and option order, suggesting that speech can amplify existing structural biases. Moreover, architectural design and reasoning strategy substantially affect robustness across languages. Overall, this study establishes a unified framework for assessing fairness and robustness in speech-integrated LLMs, bridging the gap between text- and speech-based evaluation. The resources can be found at https://github.com/ntunlplab/BiasInEar.
【14】A Baseline Multimodal Approach to Emotion Recognition in Conversations
标题:对话中情感识别的基线多模式方法
链接:https://arxiv.org/abs/2602.00914
备注:10 pages
摘要:我们提出了一个轻量级的多模态基线的情感识别会话中使用SemEval-2024任务3数据集从情景喜剧朋友。本报告的目标不是提出一种新的最先进的方法,而是记录一种可访问的参考实现,该实现结合了(i)基于transformer的文本分类器和(ii)自监督语音表示模型,以及一个简单的后期融合集成。我们报告了在有限训练协议下获得的基线设置和经验结果,强调了多峰融合何时优于单峰模型。提供此预印本是为了提高透明度,并支持未来更严格的比较。
摘要:We present a lightweight multimodal baseline for emotion recognition in conversations using the SemEval-2024 Task 3 dataset built from the sitcom Friends. The goal of this report is not to propose a novel state-of-the-art method, but to document an accessible reference implementation that combines (i) a transformer-based text classifier and (ii) a self-supervised speech representation model, with a simple late-fusion ensemble. We report the baseline setup and empirical results obtained under a limited training protocol, highlighting when multimodal fusion improves over unimodal models. This preprint is provided for transparency and to support future, more rigorous comparisons.
【15】ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation
标题:ACE-步骤1.5:突破开源音乐生成的界限
链接:https://arxiv.org/abs/2602.00744
摘要:我们介绍ACE-Step v1.5,这是一个高效的开源音乐基础模型,为消费硬件带来了商业级的一代。在常用的评估指标上,ACE-Step v1.5实现了超越大多数商业音乐模型的质量,同时保持了极高的速度-在A100上每首完整歌曲不到2秒,在RTX 3090上不到10秒。该模型在本地运行,VRAM不到4GB,并支持轻量级个性化:用户可以通过几首歌曲训练LoRA来捕捉自己的风格。它的核心是一个新颖的混合架构,其中语言模型(LM)作为一个全能的规划器:它将简单的用户查询转换为全面的歌曲蓝图-从短循环扩展到10分钟的作品-同时通过思想链合成元数据,歌词和标题,以指导扩散Transformer(DiT)。独特的是,这种对齐是通过仅依赖于模型内部机制的内在强化学习实现的,从而消除了外部奖励模型或人类偏好中固有的偏见。除了标准的合成,ACE-Step v1.5将精确的风格控制与多功能的编辑功能(如封面生成、重绘和声音到BGM的转换)统一起来,同时严格遵守50多种语言的提示。这为无缝集成到音乐艺术家、制作人和内容创作者的创作工作流程中的强大工具铺平了道路。代码、模型重量和演示可在https://ace-step.github.io/ace-step-v1.5.github.io/上获得
摘要:We present ACE-Step v1.5, a highly efficient open-source music foundation model that brings commercial-grade generation to consumer hardware. On commonly used evaluation metrics, ACE-Step v1.5 achieves quality beyond most commercial music models while remaining extremely fast -- under 2 seconds per full song on an A100 and under 10 seconds on an RTX 3090. The model runs locally with less than 4GB of VRAM, and supports lightweight personalization: users can train a LoRA from just a few songs to capture their own style. At its core lies a novel hybrid architecture where the Language Model (LM) functions as an omni-capable planner: it transforms simple user queries into comprehensive song blueprints -- scaling from short loops to 10-minute compositions -- while synthesizing metadata, lyrics, and captions via Chain-of-Thought to guide the Diffusion Transformer (DiT). Uniquely, this alignment is achieved through intrinsic reinforcement learning relying solely on the model's internal mechanisms, thereby eliminating the biases inherent in external reward models or human preferences. Beyond standard synthesis, ACE-Step v1.5 unifies precise stylistic control with versatile editing capabilities -- such as cover generation, repainting, and vocal-to-BGM conversion -- while maintaining strict adherence to prompts across 50+ languages. This paves the way for powerful tools that seamlessly integrate into the creative workflows of music artists, producers, and content creators. The code, the model weights and the demo are available at: https://ace-step.github.io/ace-step-v1.5.github.io/
【16】Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning
标题:跨模态二元注意:一种能量有效的视听学习融合框架
链接:https://arxiv.org/abs/2602.00701
摘要:有效的多模态融合需要能够捕获复杂的跨模态依赖关系的机制,同时保持计算可扩展性以用于现实世界的部署。现有的视听融合方法面临着一个根本的权衡:基于注意力的方法有效地建模跨模态关系,但产生二次计算复杂性,防止分层,多尺度架构,而有效的融合策略依赖于简单的级联,无法提取互补的跨模态信息。我们引入CMQKA,一种新的跨模态融合机制,通过有效的二进制运算实现线性O(N)的复杂度,使可扩展的层次融合以前不可行的传统的注意。CMQKA采用双向跨模态查询-关键注意力来提取互补的时空特征,并使用可学习的残差融合来保留特定于模态的特征,同时用跨模态信息丰富表示。在CMQKA的基础上,我们提出了SNNergy,一个节能的多模态融合框架,具有分层架构,通过逐步降低空间分辨率和增加语义抽象来处理输入。这种多尺度融合能力允许框架跨模态捕获局部模式和全局上下文。SNNergy采用事件驱动的二进制尖峰操作实现,在保持融合有效性的同时实现了卓越的能效,并在具有挑战性的视听基准上建立了新的最先进的结果,包括CREMA-D,AVE和UrbanSound 8 K-AV,显著优于现有的多模式融合基准。我们的框架通过引入可扩展的融合机制来推进多模态融合,该融合机制能够为现实世界的视听智能系统实现具有实用能源效率的分层跨模态集成。
摘要:Effective multimodal fusion requires mechanisms that can capture complex cross-modal dependencies while remaining computationally scalable for real-world deployment. Existing audio-visual fusion approaches face a fundamental trade-off: attention-based methods effectively model cross-modal relationships but incur quadratic computational complexity that prevents hierarchical, multi-scale architectures, while efficient fusion strategies rely on simplistic concatenation that fails to extract complementary cross-modal information. We introduce CMQKA, a novel cross-modal fusion mechanism that achieves linear O(N) complexity through efficient binary operations, enabling scalable hierarchical fusion previously infeasible with conventional attention. CMQKA employs bidirectional cross-modal Query-Key attention to extract complementary spatiotemporal features and uses learnable residual fusion to preserve modality-specific characteristics while enriching representations with cross-modal information. Building upon CMQKA, we present SNNergy, an energy-efficient multimodal fusion framework with a hierarchical architecture that processes inputs through progressively decreasing spatial resolutions and increasing semantic abstraction. This multi-scale fusion capability allows the framework to capture both local patterns and global context across modalities. Implemented with event-driven binary spike operations, SNNergy achieves remarkable energy efficiency while maintaining fusion effectiveness and establishing new state-of-the-art results on challenging audio-visual benchmarks, including CREMA-D, AVE, and UrbanSound8K-AV, significantly outperforming existing multimodal fusion baselines. Our framework advances multimodal fusion by introducing a scalable fusion mechanism that enables hierarchical cross-modal integration with practical energy efficiency for real-world audio-visual intelligence systems.
【17】Audio-to-Image Bird Species Retrieval without Audio-Image Pairs via Text Distillation
标题:通过文本蒸馏进行无音频图像对的音频到图像鸟类物种检索
链接:https://arxiv.org/abs/2602.00681
摘要:音频到图像检索提供了一个可解释的替代音频的生物声学物种识别的分类,但学习对齐的音频图像表示是具有挑战性的,由于成对的音频图像数据的稀缺性。我们提出了一个简单的和数据高效的方法,使音频到图像检索没有任何音频图像的监督。我们提出的方法使用文本作为语义中介:我们通过使用对比目标微调其音频编码器,将预训练的图像-文本模型(BioCLIP-2)的文本嵌入空间提取到预训练的音频-文本模型(BioLingual)中,该模型编码丰富的视觉和分类结构。这种提炼将视觉上的语义转移到音频表示中,在训练过程中不使用图像的情况下,诱导音频和图像嵌入之间的紧急对齐。我们评估多个生物声学基准模型。提取的音频编码器保留了音频辨别能力,同时大大提高了焦点录音和音景数据集上的音频-文本对齐。最重要的是,在SSW60基准测试中,所提出的方法实现了强大的音频到图像检索性能,超过了基于zero-shot模型组合或文本嵌入之间的学习映射的基线,尽管没有对配对的音频图像数据进行训练。这些结果表明,通过文本的间接语义转移足以诱导有意义的音频图像对齐,提供了一个实用的解决方案,在数据稀缺的生物声学环境中的视觉接地物种识别。
摘要:Audio-to-image retrieval offers an interpretable alternative to audio-only classification for bioacoustic species recognition, but learning aligned audio-image representations is challenging due to the scarcity of paired audio-image data. We propose a simple and data-efficient approach that enables audio-to-image retrieval without any audio-image supervision. Our proposed method uses text as a semantic intermediary: we distill the text embedding space of a pretrained image-text model (BioCLIP-2), which encodes rich visual and taxonomic structure, into a pretrained audio-text model (BioLingual) by fine-tuning its audio encoder with a contrastive objective. This distillation transfers visually grounded semantics into the audio representation, inducing emergent alignment between audio and image embeddings without using images during training. We evaluate the resulting model on multiple bioacoustic benchmarks. The distilled audio encoder preserves audio discriminative power while substantially improving audio-text alignment on focal recordings and soundscape datasets. Most importantly, on the SSW60 benchmark, the proposed approach achieves strong audio-to-image retrieval performance exceeding baselines based on zero-shot model combinations or learned mappings between text embeddings, despite not training on paired audio-image data. These results demonstrate that indirect semantic transfer through text is sufficient to induce meaningful audio-image alignment, providing a practical solution for visually grounded species recognition in data-scarce bioacoustic settings.
【18】MTAVG-Bench: A Comprehensive Benchmark for Evaluating Multi-Talker Dialogue-Centric Audio-Video Generation
标题:MTAVG-Bench:评估以多人对话为中心的音频视频生成的综合基准
链接:https://arxiv.org/abs/2602.00607
摘要:文本到音频视频(T2 AV)生成的最新进展使模型能够合成具有多参与者对话的视听视频。然而,现有的评估基准仍然主要是针对人类录制的视频或单扬声器设置而设计的。结果,不能有效地捕获和分析在所生成的多说话者对话视频中发生的潜在错误,诸如身份漂移、不自然的转向转换和视听未对准。为了解决这个问题,我们引入MTAVG-Bench,用于评估视听多说话者对话生成的基准。MTAVG-Bench是通过半自动管道构建的,其中使用多个流行模型生成1.8k视频,并精心设计提示,产生2.4k手动注释的QA对。该基准从四个层面评估多说话人对话生成:视听信号保真度、时间属性一致性、社会交互和电影表达。我们在MTAVG-Bench上对12种专有和开源全方位模型进行了基准测试,其中Gemini 3 Pro实现了最强的整体性能,而领先的开源模型在信号保真度和一致性方面仍然具有竞争力。总的来说,MTAVG-Bench能够进行细粒度的故障分析,以进行严格的模型比较和有针对性的视频生成优化。
摘要:Recent advances in text-to-audio-video (T2AV) generation have enabled models to synthesize audio-visual videos with multi-participant dialogues. However, existing evaluation benchmarks remain largely designed for human-recorded videos or single-speaker settings. As a result, potential errors that occur in generated multi-talker dialogue videos, such as identity drift, unnatural turn transitions, and audio-visual misalignment, cannot be effectively captured and analyzed. To address this issue, we introduce MTAVG-Bench, a benchmark for evaluating audio-visual multi-speaker dialogue generation. MTAVG-Bench is built via a semi-automatic pipeline, where 1.8k videos are generated using multiple popular models with carefully designed prompts, yielding 2.4k manually annotated QA pairs. The benchmark evaluates multi-speaker dialogue generation at four levels: audio-visual signal fidelity, temporal attribute consistency, social interaction, and cinematic expression. We benchmark 12 proprietary and open-source omni-models on MTAVG-Bench, with Gemini 3 Pro achieving the strongest overall performance, while leading open-source models remain competitive in signal fidelity and consistency. Overall, MTAVG-Bench enables fine-grained failure analysis for rigorous model comparison and targeted video generation refinement.
【19】The TMU System for the XACLE Challenge: Training Large Audio Language Models with CLAP Pseudo-Labels
标题:XACLE挑战赛TMU系统:使用CLAP伪标签训练大型音频语言模型
链接:https://arxiv.org/abs/2602.00604
备注:3 pages; 2 figures; 2 tables; Accepted at ICASSP 2026 Workshop (SP Grand Challenges, GC-12: XACLE)
摘要:在本文中,我们提出了一个提交x到音频对齐(XACLE)的挑战。目标是预测给定的一般音频和文本对的语义对齐。该系统是基于一个大的音频语言模型(LALM)架构。我们采用了三个阶段的训练管道:自动音频字幕预训练,CLAP伪标签预训练,以及对XACLE数据集进行微调。我们的实验表明,使用CLAP伪标签进行预训练是主要的性能驱动因素。在XACLE测试集上,我们的系统达到了0.632的SRCC,显著优于基线系统(0.334),并在挑战团队排名中获得第三名。代码和模型可以在https://github.com/shiotalab-tmu/tmu-xacle2026上找到
摘要:In this paper, we propose a submission to the x-to-audio alignment (XACLE) challenge. The goal is to predict semantic alignment of a given general audio and text pair. The proposed system is based on a large audio language model (LALM) architecture. We employ a three-stage training pipeline: automated audio captioning pretraining, pretraining with CLAP pseudo-labels, and fine-tuning on the XACLE dataset. Our experiments show that pretraining with CLAP pseudo-labels is the primary performance driver. On the XACLE test set, our system reaches an SRCC of 0.632, significantly outperforming the baseline system (0.334) and securing third place in the challenge team ranking. Code and models can be found at https://github.com/shiotalab-tmu/tmu-xacle2026
【20】Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
标题:Kanade:一个用于口语建模的简单解开令牌器
链接:https://arxiv.org/abs/2602.00594
摘要:一个好的语言模型始于一个好的标记器。标记化对于语音建模尤其重要,因为语音建模必须处理混合了语言和非语言信息的连续信号。语音分词器应该提取语音和韵律,抑制语言上不相关的信息,如说话人身份,并实现高质量的合成。我们提出了Kanade,一个单层的解开语音标记器,实现了这一理想。Kanade将声学常数分离出来,创建一个单一的令牌流,可以捕获丰富的语音和韵律。实验表明,Kanade在保持良好重建质量的同时,实现了最先进的说话人解缠和词汇可用性。
摘要:A good language model starts with a good tokenizer. Tokenization is especially important for speech modeling, which must handle continuous signals that mix linguistic and non-linguistic information. A speech tokenizer should extract phonetics and prosody, suppress linguistically irrelevant information like speaker identity, and enable high-quality synthesis. We present Kanade, a single-layer disentangled speech tokenizer that realizes this ideal. Kanade separates out acoustic constants to create a single stream of tokens that captures rich phonetics and prosody. It does so without the need for auxiliary methods that existing disentangled codecs often rely on. Experiments show that Kanade achieves state-of-the-art speaker disentanglement and lexical availability, while maintaining excellent reconstruction quality.
【21】Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy
标题:双视图预测扩散:通过谱图-图像协同实现轻量级语音增强
链接:https://arxiv.org/abs/2602.00568
摘要:扩散模型最近在语音增强(SE)中设定了新的基准。然而,大多数现有的基于分数的模型仅将语音频谱图视为通用的2D图像,应用忽略音频的固有结构稀疏性的统一处理,这导致低效的频谱表示和过高的计算复杂度。为了弥合这一差距,我们提出了DVPD,一个非常轻量级的双视图预测扩散模型,它独特地利用了频谱图的双重性质,即在训练和推理阶段都是视觉纹理和物理频域表示。具体来说,在训练过程中,我们通过频率自适应非均匀压缩(FANC)编码器优化频谱利用率,该编码器在修剪高频冗余的同时保留了关键的低频谐波。同时,我们引入了一个轻量级的基于图像的光谱感知(LISA)模块,以最小的开销从视觉角度捕捉功能。在推理过程中,我们提出了一个无训练的无损提升(TLB)策略,利用相同的双视图先验来改进生成质量,而无需任何额外的微调。在各种基准测试中进行的大量实验表明,与SOTA轻量级模型PGUSE相比,DVPD实现了最先进的性能,同时只需要35%的参数和40%的推理MAC。这些结果突出了DVPD在平衡高保真语音质量与极端架构效率方面的卓越能力。代码和音频示例可在匿名网站上获得:{https://anonymous.4open.science/r/dvpd_demo-E630}
摘要:Diffusion models have recently set new benchmarks in Speech Enhancement (SE). However, most existing score-based models treat speech spectrograms merely as generic 2D images, applying uniform processing that ignores the intrinsic structural sparsity of audio, which results in inefficient spectral representation and prohibitive computational complexity. To bridge this gap, we propose DVPD, an extremely lightweight Dual-View Predictive Diffusion model, which uniquely exploits the dual nature of spectrograms as both visual textures and physical frequency-domain representations across both training and inference stages. Specifically, during training, we optimize spectral utilization via the Frequency-Adaptive Non-uniform Compression (FANC) encoder, which preserves critical low-frequency harmonics while pruning high-frequency redundancies. Simultaneously, we introduce a Lightweight Image-based Spectro-Awareness (LISA) module to capture features from a visual perspective with minimal overhead. During inference, we propose a Training-free Lossless Boost (TLB) strategy that leverages the same dual-view priors to refine generation quality without any additional fine-tuning. Extensive experiments across various benchmarks demonstrate that DVPD achieves state-of-the-art performance while requiring only 35% of the parameters and 40% of the inference MACs compared to SOTA lightweight model, PGUSE. These results highlight DVPD's superior ability to balance high-fidelity speech quality with extreme architectural efficiency. Code and audio samples are available at the anonymous website: {https://anonymous.4open.science/r/dvpd_demo-E630}
【22】Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
标题:编辑内容,保留声音:通过自一致性奖励进行基于文本的语音编辑
链接:https://arxiv.org/abs/2602.00560
摘要:不可感知的基于文本的语音编辑允许用户通过改变转录来修改口语内容。它要求修改后的片段与周围的上下文无缝融合。在声学空间中操作的流行方法遭受固有的内容风格纠缠,导致生成不稳定性和边界伪影。在本文中,我们提出了一个新的框架接地的原则“编辑内容,保留声学”。我们的方法依赖于两个核心组件:(1)结构基础,它将编辑纳入稳定的语义空间,同时将声学重建委托给流匹配解码器;(2)感知对齐,它采用了一种新的自一致性奖励组相对策略优化。通过利用预先训练的文本到语音转换模型作为隐式批评者-辅以严格的可理解性和持续时间限制-我们有效地将编辑的语义标记序列与原始上下文对齐。经验评估表明,我们的方法显着优于最先进的自回归和非自回归基线,实现卓越的可懂度,鲁棒性和感知质量。
摘要:Imperceptible text-based speech editing allows users to modify spoken content by altering the transcript. It demands that modified segments fuse seamlessly with the surrounding context. Prevalent methods operating in the acoustic space suffer from inherent content-style entanglement, leading to generation instability and boundary artifacts. In this paper, we propose a novel framework grounded in the principle of "Edit Content, Preserve Acoustics". Our approach relies on two core components: (1) Structural Foundations, which decouples editing into a stable semantic space while delegating acoustic reconstruction to a Flow Matching decoder; and (2) Perceptual Alignment, which employs a novel Self-Consistency Rewards Group Relative Policy Optimization. By leveraging a pre-trained Text-to-Speech model as an implicit critic -- complemented by strict intelligibility and duration constraints -- we effectively align the edited semantic token sequence with the original context. Empirical evaluations demonstrate that our method significantly outperforms state-of-the-art autoregressive and non-autoregressive baselines, achieving superior intelligibility, robustness, and perceptual quality.
【23】RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
标题:RVCBench:对现代音频生成模型中语音克隆的稳健性进行基准测试
链接:https://arxiv.org/abs/2602.00443
备注:40 pages, 12figures
摘要:现代语音克隆(VC)可以仅从几秒钟的参考音频中合成与目标说话人非常匹配的语音,从而实现个性化语音界面和配音等应用。在实际部署中,现代音频生成模型不可避免地会遇到嘈杂的参考音频、不完美的文本提示和多样化的下游处理,这会严重损害鲁棒性。尽管VC在自回归编解码器-令牌语言模型和基于扩散的模型的驱动下取得了快速进展,但现实部署变化下的鲁棒性仍然未得到充分探索。本文介绍了RVCBench,这是一个全面的基准测试,用于评估整个生成管道中VC的鲁棒性,包括输入变化,生成挑战,输出后处理和对抗性扰动,涵盖10个鲁棒性任务,225个扬声器,14,370个话语和11个代表性的现代VC模型。我们的评估揭示了VC中的大量鲁棒性差距:在常见的输入移位和后处理下,性能可能会急剧恶化;长上下文和跨语言场景进一步暴露了稳定性限制;被动噪声和主动扰动都影响生成鲁棒性。总的来说,这些发现提供了一个统一的图片,目前的VC模型在实践中失败,并引入了一个标准化的,开源的测试平台,以支持更强大的和可部署的VC模型的开发。我们在https://github.com/Nanboy-Ronan/RVCBench上开源了我们的项目。
摘要:Modern voice cloning (VC) can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling applications such as personalized speech interfaces and dubbing. In practical deployments, modern audio generation models inevitably encounter noisy reference audios, imperfect text prompts, and diverse downstream processing, which can significantly hurt robustness. Despite rapid progress in VC driven by autoregressive codec-token language models and diffusion-based models, robustness under realistic deployment shifts remains underexplored. This paper introduces RVCBench, a comprehensive benchmark that evaluates Robustness in VC across the full generation pipeline, including input variation, generation challenges, output post-processing, and adversarial perturbations, covering 10 robustness tasks, 225 speakers, 14,370 utterances, and 11 representative modern VC models. Our evaluation uncovers substantial robustness gaps in VC: performance can deteriorate sharply under common input shifts and post-processing; long-context and cross-lingual scenarios further expose stability limitations; and both passive noise and proactive perturbation influence generation robustness. Collectively, these findings provide a unified picture of how current VC models fail in practice and introduce a standardized, open-source testbed to support the development of more robust and deployable VC models. We open-source our project at https://github.com/Nanboy-Ronan/RVCBench.
【24】Multi-Speaker Conversational Audio Deepfake: Taxonomy, Dataset and Pilot Study
标题:多扬声器对话音频Deepfake:分类、数据集和试点研究
链接:https://arxiv.org/abs/2602.00295
备注:This work was presented at the 2025 IEEE International Conference on Data Mining, ICDM 2025, November 12-15,2025, Washington DC, USA
摘要:文本到语音(TTS)技术的快速发展使音频deepfake变得越来越真实和可访问,引发了重大的安全和信任问题。虽然现有的研究主要集中在检测单扬声器音频deepfake,但具有多扬声器会话设置的真实恶意应用程序也正在成为一个未被充分研究的主要威胁。为了解决这一差距,我们提出了一种多说话者对话音频deepfakes的概念分类,区分部分操作(一个或多个说话者改变)和完全操作(整个对话合成)。作为第一步,我们引入了一个新的多扬声器会话音频Deepfakes数据集(MsCADD),其中包含2,830个音频片段,这些音频片段包含真实和完全合成的两个扬声器对话,使用基于VITS和SoundStorm的NotebookLM模型生成,以模拟具有扬声器性别变化的自然对话和会话自发性。MsCADD仅限于文本到语音(TTS)类型的deepfake。我们在这个数据集上对三个神经基线模型进行了基准测试:LFCC-LCNN,RawNet 2和Wav 2 Vec 2.0,并报告了F1评分,准确性,真阳性率(TPR)和真阴性率(TNR)方面的性能。结果表明,这些基线模型提供了一个有用的基准,然而,结果也强调了多说话者deepfake研究在不同会话动态下可靠检测合成语音方面存在显着差距。我们的数据集和基准测试为未来在对话场景中进行deepfake检测的研究奠定了基础,这是一个高度未开发的研究领域,但也是对音频设置中的可信信息构成威胁的主要领域。MsCADD数据集是公开的,以支持研究社区的可重复性和基准测试。
摘要:The rapid advances in text-to-speech (TTS) technologies have made audio deepfakes increasingly realistic and accessible, raising significant security and trust concerns. While existing research has largely focused on detecting single-speaker audio deepfakes, real-world malicious applications with multi-speaker conversational settings is also emerging as a major underexplored threat. To address this gap, we propose a conceptual taxonomy of multi-speaker conversational audio deepfakes, distinguishing between partial manipulations (one or multiple speakers altered) and full manipulations (entire conversations synthesized). As a first step, we introduce a new Multi-speaker Conversational Audio Deepfakes Dataset (MsCADD) of 2,830 audio clips containing real and fully synthetic two-speaker conversations, generated using VITS and SoundStorm-based NotebookLM models to simulate natural dialogue with variations in speaker gender, and conversational spontaneity. MsCADD is limited to text-to-speech (TTS) types of deepfake. We benchmark three neural baseline models; LFCC-LCNN, RawNet2, and Wav2Vec 2.0 on this dataset and report performance in terms of F1 score, accuracy, true positive rate (TPR), and true negative rate (TNR). Results show that these baseline models provided a useful benchmark, however, the results also highlight that there is a significant gap in multi-speaker deepfake research in reliably detecting synthetic voices under varied conversational dynamics. Our dataset and benchmarks provide a foundation for future research on deepfake detection in conversational scenarios, which is a highly underexplored area of research but also a major area of threat to trustworthy information in audio settings. The MsCADD dataset is publicly available to support reproducibility and benchmarking by the research community.
【25】VoxServe: Streaming-Centric Serving System for Speech Language Models
标题:VoxServe:以流媒体为中心的语音语言模型服务系统
链接:https://arxiv.org/abs/2602.00269
备注:The code is available at https://github.com/vox-serve/vox-serve
摘要:在流媒体环境中部署现代语音语言模型(SpeechLM)需要系统提供低延迟、高吞吐量和强有力的流媒体性保证。现有的系统不能灵活有效地支持不同的模型。我们提出VoxServe,一个统一的SpeechLM服务系统,优化流媒体性能。VoxServe引入了一种模型执行抽象,将模型架构与系统级优化相结合,从而在单个框架内支持多种SpeechLM架构。在此抽象的基础上,VoxServe实现了流感知调度和异步推理管道,以提高端到端的效率。对多个现代SpeechLM的评估表明,VoxServe在相当延迟的情况下实现了比现有实现高10- 20倍的吞吐量,同时保持了高的流可行性。VoxServe的代码可在https://github.com/vox-serve/vox-serve上获得。
摘要:Deploying modern Speech Language Models (SpeechLMs) in streaming settings requires systems that provide low latency, high throughput, and strong guarantees of streamability. Existing systems fall short of supporting diverse models flexibly and efficiently. We present VoxServe, a unified serving system for SpeechLMs that optimizes streaming performance. VoxServe introduces a model-execution abstraction that decouples model architecture from system-level optimizations, thereby enabling support for diverse SpeechLM architectures within a single framework. Building on this abstraction, VoxServe implements streaming-aware scheduling and an asynchronous inference pipeline to improve end-to-end efficiency. Evaluations across multiple modern SpeechLMs show that VoxServe achieves 10-20x higher throughput than existing implementations at comparable latency while maintaining high streaming viability. The code of VoxServe is available at https://github.com/vox-serve/vox-serve.
【26】LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
标题:LPIPS-AttnWave 2Lip:适合野外会说话的头一代的通用音频驱动嘴唇同步
链接:https://arxiv.org/abs/2602.00189
备注:This paper has been accepted by Elsevier's \textit{Speech Communication} journal. Official publication link: https://doi.org/10.1016/j.specom.2023.103028 The code for the paper is available at the following link: https://github.com/FelixChan9527/LPIPS-AttnWav2Lip
摘要:研究人员对音频驱动的Talking Head Generation越来越感兴趣。说话头部生成的主要挑战是实现嘴唇和音频之间的视听一致性,称为嘴唇同步。本文提出了一种通用的方法,LPIPS-AttnWav 2Lip,用于重建任何说话人的人脸图像的基础上的音频。我们使用基于残差CBAM的U-Net架构来更好地编码和融合音频和视觉模态信息。此外,语义对齐模块扩展生成器网络的感受野,有效获取视觉特征的空间和通道信息;将视觉特征的统计信息与音频特征向量进行匹配,实现音频内容信息对视觉信息的调整和注入。为了实现精确的唇同步并生成逼真的高质量图像,我们的方法采用LPIPS损失,它模拟人类对图像质量的判断,并减少训练过程中不稳定的可能性。主观和客观评价结果表明,该方法在唇同步精度和视觉质量方面取得了优异的性能。论文的代码可在以下链接中获得:https://github.com/FelixChan9527/LPIPS-AttnWav2Lip
摘要:Researchers have shown a growing interest in Audio-driven Talking Head Generation. The primary challenge in talking head generation is achieving audio-visual coherence between the lips and the audio, known as lip synchronization. This paper proposes a generic method, LPIPS-AttnWav2Lip, for reconstructing face images of any speaker based on audio. We used the U-Net architecture based on residual CBAM to better encode and fuse audio and visual modal information. Additionally, the semantic alignment module extends the receptive field of the generator network to obtain the spatial and channel information of the visual features efficiently; and match statistical information of visual features with audio latent vector to achieve the adjustment and injection of the audio content information to the visual information. To achieve exact lip synchronization and to generate realistic high-quality images, our approach adopts LPIPS Loss, which simulates human judgment of image quality and reduces instability possibility during the training process. The proposed method achieves outstanding performance in terms of lip synchronization accuracy and visual quality as demonstrated by subjective and objective evaluation results. The code for the paper is available at the following link: https://github.com/FelixChan9527/LPIPS-AttnWav2Lip
【27】SSNAPS: Audio-Visual Separation of Speech and Background Noise with Diffusion Inverse Sampling
标题:SSNAPS:使用扩散反采样的语音和背景噪音的视听分离
链接:https://arxiv.org/abs/2602.01394
摘要:本文讨论了在真实环境噪声存在下的视听单麦克风语音分离和增强的挑战。我们的方法是基于生成式逆采样,我们用专用的扩散先验对干净的语音和环境噪声进行建模,并联合利用它们来恢复所有潜在的源。为了实现这一点,我们重新制定了最近的逆采样器,以匹配我们的设置。我们对1,2和3个扬声器与噪声的混合物进行了评估,并表明,尽管完全无监督,但我们的方法在所有条件下都始终优于\ac{WER}中的领先监督基线。我们进一步扩展我们的框架来处理屏幕外扬声器分离。此外,分离的噪声分量的高保真度使其适合于下游声学场景检测。演示页面:https://ssnapsicml.github.io/ssnapsicml2026/
摘要:This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise. Our approach is based on generative inverse sampling, where we model clean speech and ambient noise with dedicated diffusion priors and jointly leverage them to recover all underlying sources. To achieve this, we reformulate a recent inverse sampler to match our setting. We evaluate on mixtures of 1, 2, and 3 speakers with noise and show that, despite being entirely unsupervised, our method consistently outperforms leading supervised baselines in \ac{WER} across all conditions. We further extend our framework to handle off-screen speaker separation. Moreover, the high fidelity of the separated noise component makes it suitable for downstream acoustic scene detection. Demo page: https://ssnapsicml.github.io/ssnapsicml2026/
【28】Adapting Where It Matters: Depth-Aware Adaptation for Efficient Multilingual Speech Recognition in Low-Resource Languages
标题:适应重要的地方:深度感知适应,在低资源语言中实现高效多语言语音识别
链接:https://arxiv.org/abs/2602.01008
备注:13 pages
摘要:最近的语音基础模型在高资源语言的多语言自动语音识别(ASR)方面表现出色,但由于数据稀缺和效率限制,使其适应低资源语言仍然具有挑战性。全模型微调在计算上是昂贵的,并且容易过拟合,而像LoRA这样的参数高效方法在层间均匀地应用自适应,忽略了内部表示,从而影响了有效性和效率。我们分析了多语言ASR模型,并揭示了一个U形的适应性模式:早期和晚期层是语言特定的,需要更多的调整,而中间层保留共享的语义,需要更少。基于这一观察,我们提出了DAMA,深度感知模型适应框架,根据每个层的角色分配适应能力。DAMA还引入了基于奇异值分解(SVD)的初始化来约束自适应并保留U形模式,以及冻结的中间层基础以进一步提高效率。在两个基准数据集上对18种低资源语言进行了评估,DAMA匹配或超越了最先进的准确性,可训练参数减少了80%,在极端数据稀缺的情况下实现了29%的错误减少,并显着提高了内存,训练时间和计算效率。这些结果突出了结构感知适应高效,可扩展的多语言ASR的好处。
摘要:Recent speech foundation models excel at multilingual automatic speech recognition (ASR) for high-resource languages, but adapting them to low-resource languages remains challenging due to data scarcity and efficiency constraints. Full-model fine-tuning is computationally expensive and prone to overfitting, while parameter-efficient methods like LoRA apply adaptation uniformly across layers, overlooking internal representations thus compromising effectiveness and efficiency. We analyze multilingual ASR models and reveal a U-shaped adaptability pattern: early and late layers are language-specific and require more adaptation, while intermediate layers retain shared semantics and need less. Building on this observation, we propose DAMA, a Depth-Aware Model Adaptation framework that allocates adaptation capacity according to each layer's role. DAMA also introduces Singular Value Decomposition (SVD)-based initialization to constrain adaptation and preserve the U-shaped pattern, as well as a frozen middle-layer basis for further efficiency. Evaluated on 18 low-resource languages across two benchmark datasets, DAMA matches or surpasses state-of-the-art accuracy with 80% fewer trainable parameters, achieves a 29% error reduction under extreme data scarcity, and significantly improves memory, training time, and computational efficiency over baselines. These results highlight the benefits of structure-aware adaptation for efficient, scalable multilingual ASR.
【29】High-Fidelity Generative Audio Compression at 0.275kbps
标题:高保真生成音频压缩,速度为0.275kMbps
链接:https://arxiv.org/abs/2602.00648
备注:Technical Report
摘要:超低比特率的高保真通用音频压缩对于从低带宽通信到生成音频语言建模的应用至关重要。传统的音频压缩方法和当代神经编解码器基本上是为波形重构而设计的。因此,当在超低比特率下操作时,这些方法迅速降级并且通常无法保留基本信息,导致严重的声学伪影和明显的语义失真。为了克服这些限制,我们引入了生成音频压缩(GAC),一种新的范式转变,从信号保真度的任务导向的有效性。GAC在AI Flow框架内实现,理论上以信息容量定律为基础。这些基础表明,丰富的计算能力可以在接收器处被利用来抵消极端的通信瓶颈--这体现了更多计算,更少带宽的理念。通过将发送方的语义理解与接收方的可扩展生成合成相结合,GAC将信息负担卸载到强大的模型先验中。我们的1.8B参数模型在0.275kbps的比特率下实现了32 kHz普通音频的高保真重建。即使在0.175kbps下,它仍然保持了强大的可理解的音频传输能力,代表了约3000倍的压缩比,在保持感知质量和语义一致性方面明显优于当前最先进的神经编解码器。
摘要:High-fidelity general audio compression at ultra-low bitrates is crucial for applications ranging from low-bandwidth communication to generative audio-language modeling. Traditional audio compression methods and contemporary neural codecs are fundamentally designed for waveform reconstruction. As a result, when operating at ultra-low bitrates, these methods degrade rapidly and often fail to preserve essential information, leading to severe acoustic artifacts and pronounced semantic distortion. To overcome these limitations, we introduce Generative Audio Compression (GAC), a novel paradigm shift from signal fidelity to task-oriented effectiveness. Implemented within the AI Flow framework, GAC is theoretically grounded in the Law of Information Capacity. These foundations posit that abundant computational power can be leveraged at the receiver to offset extreme communication bottlenecks--exemplifying the More Computation, Less Bandwidth philosophy. By integrating semantic understanding at the transmitter with scalable generative synthesis at the receiver, GAC offloads the information burden to powerful model priors. Our 1.8B-parameter model achieves high-fidelity reconstruction of 32kHz general audio at an unprecedented bitrate of 0.275kbps. Even at 0.175kbps, it still preserves a strong intelligible audio transmission capability, which represents an about 3000x compression ratio, significantly outperforming current state-of-the-art neural codecs in maintaining both perceptual quality and semantic consistency.
【1】RIR-Former: Coordinate-Guided Transformer for Continuous Reconstruction of Room Impulse Responses
标题:RIR-former:用于连续重建房间冲击响应的坐标引导Transformer
链接:https://arxiv.org/abs/2602.01861
备注:Accepted to International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2026. Equal contribution: Shaoheng Xu and Chunyi Sun
摘要:房间冲激响应(RIR)是许多声学信号处理任务的基础,但在空间上密集测量它们往往是不切实际的。在这项工作中,我们提出了RIR-Former,一个无网格,一步前馈模型RIR重建。通过将正弦编码模块引入到Transformer骨干中,我们的方法有效地结合了麦克风位置信息,从而能够在任意阵列位置进行插值。此外,分段的多分支解码器被设计为分别处理早期反射和后期混响,从而改善整个RIR的重建。不同的模拟声学环境的实验表明,RIR-Former一贯优于国家的最先进的基线方面的归一化均方误差(NMSE)和余弦距离(CD),在不同的缺失率和阵列配置。这些结果突出了我们的方法在实际部署中的潜力,并激发了未来从随机间隔线性阵列扩展到复杂阵列几何形状,动态声学场景和真实环境的工作。
摘要:Room impulse responses (RIRs) are essential for many acoustic signal processing tasks, yet measuring them densely across space is often impractical. In this work, we propose RIR-Former, a grid-free, one-step feed-forward model for RIR reconstruction. By introducing a sinusoidal encoding module into a transformer backbone, our method effectively incorporates microphone position information, enabling interpolation at arbitrary array locations. Furthermore, a segmented multi-branch decoder is designed to separately handle early reflections and late reverberation, improving reconstruction across the entire RIR. Experiments on diverse simulated acoustic environments demonstrate that RIR-Former consistently outperforms state-of-the-art baselines in terms of normalized mean square error (NMSE) and cosine distance (CD), under varying missing rates and array configurations. These results highlight the potential of our approach for practical deployment and motivate future work on scaling from randomly spaced linear arrays to complex array geometries, dynamic acoustic scenes, and real-world environments.
【2】Short-wave admittance correction for a time-domain cochlear transmission line model
标题:时间域人工耳传输线模型的长波电导率修正
链接:https://arxiv.org/abs/2602.01758
备注:22 pages, 7 figures
摘要:在时域中实现的传输线(TL)模型可以有效地模拟响应于瞬态或非平稳声音的基底膜(BM)位移。通过设计,TL模型非常适合于行波的一维(1-D)表征,但耳蜗的真实配置也引入了更高维的效应。这种效应包括BM周围的压力集中和横向粘性阻尼,这两者都在短波区域被放大。这两种效应取决于波长,并且更容易在频域中表达。在本文中,我们介绍了一个数值修正BM导纳占2-D的影响,在时域中使用自回归滤波和回归技术。校正需要实施的TL模型适合沙鼠耳蜗生理。该模型,其中包括瞬时非线性形式的可变阻尼,最初提出了不足的压缩与增加声级。这种限制被解释为强耦合增益和频率选择性之间的1-D非线性TL模型假设,而耳蜗频率选择性显示只有适度的依赖于声音水平在小型哺乳动物。在沙鼠模型中实现校正因子,并使用反馈回路使其与水平相关。更新后的模型实现了频率选择性和增益之间的一些解耦,提供了5 dB的额外增益,并将压缩状态的声级范围扩展了10 dB。我们讨论了这项工作的相关性,通过两个关键功能:集成的分析和回归方法表征BM导纳,瞬时和非瞬时非线性的组合。
摘要:Transmission line (TL) models implemented in the time domain can efficiently simulate basilar-membrane (BM) displacement in response to transient or non-stationary sounds. By design, a TL model is well-suited for an one-dimensional (1-D) characterization of the traveling wave, but the real configuration of the cochlea also introduces higher-dimensional effects. Such effects include the focusing of the pressure around the BM and transverse viscous damping, both of which are magnified in the short-wave region. The two effects depend on the wavelength and are more readily expressed in the frequency domain. In this paper, we introduce a numerical correction for the BM admittance to account for 2-D effects in the time domain using autoregressive filtering and regression techniques. The correction was required for the implementation of a TL model tailored to the gerbil cochlear physiology. The model, which includes instantaneous nonlinearities in the form of variable damping, initially presented insufficient compression with increasing sound levels. This limitation was explained by the strong coupling between gain and frequency selectivity assumed in the 1-D nonlinear TL model, whereas cochlear frequency selectivity shows only a moderate dependence on sound level in small mammals. The correction factor was implemented in the gerbil model and made level-dependent using a feedback loop. The updated model achieved some decoupling between frequency selectivity and gain, providing 5 dB of additional gain and extending the range of sound levels of the compressive regime by 10 dB. We discuss the relevance of this work through two key features: the integration of both analytical and regression methods for characterizing BM admittance, and the combination of instantaneous and non-instantaneous nonlinearities.
【3】Joint Optimization of ASV and CM tasks: BTUEF Team's Submission for WildSpoof Challenge
标题:ASV和CM任务的联合优化:BTUF团队提交WildSpoof挑战赛
链接:https://arxiv.org/abs/2602.01722
摘要:欺骗感知说话人验证(SASV)联合解决自动说话人验证和欺骗对策,以提高对抗性攻击的鲁棒性。在本文中,我们研究了我们最近提出的模块化SASV框架,该框架通过非线性融合,显式建模它们的相互作用,并使用依赖于操作条件的可训练a-DCF损失进行优化,从而有效地重用公开可用的ASV和CM系统。该框架使用ECAPA-TDNN和ReDimNet作为ASV嵌入提取器,SSL-AASIST作为CM模型进行评估,并在WildSpoof SASV训练数据上进行了微调和未微调的实验。结果表明,通过将基于ReDimNet的ASV嵌入与微调的SSL-AASIST表示相结合,可以实现最佳性能,在进度评估集上产生0.0515的a-DCF,在最终评估集上产生0.2163的a-DCF。
摘要:Spoofing-aware speaker verification (SASV) jointly addresses automatic speaker verification and spoofing countermeasures to improve robustness against adversarial attacks. In this paper, we investigate our recently proposed modular SASV framework that enables effective reuse of publicly available ASV and CM systems through non-linear fusion, explicitly modeling their interaction, and optimization with an operating-condition-dependent trainable a-DCF loss. The framework is evaluated using ECAPA-TDNN and ReDimNet as ASV embedding extractors and SSL-AASIST as the CM model, with experiments conducted both with and without fine-tuning on the WildSpoof SASV training data. Results show that the best performance is achieved by combining ReDimNet-based ASV embeddings with fine-tuned SSL-AASIST representations, yielding an a-DCF of 0.0515 on the progress evaluation set and 0.2163 on the final evaluation set.
【4】HuPER: A Human-Inspired Framework for Phonetic Perception
标题:HuPER:一个以人为本的语音感知框架
链接:https://arxiv.org/abs/2602.01634
摘要:我们提出了HuPER,一个人类启发的框架,语音感知模型的声学语音证据和语言知识的自适应推理。仅用100小时的训练数据,HuPER就在五种英语基准测试中实现了最先进的语音错误率,并对95种未知语言实现了强大的zero-shot迁移。HuPER也是第一个在不同声学条件下实现自适应多路径语音感知的框架。所有的训练数据、模型和代码都是开源的。代码和演示可在https://github.com/HuPER29/HuPER。
摘要:We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic error rates on five English benchmarks and strong zero-shot transfer to 95 unseen languages. HuPER is also the first framework to enable adaptive, multi-path phonetic perception under diverse acoustic conditions. All training data, models, and code are open-sourced. Code and demo avaliable at https://github.com/HuPER29/HuPER.
【5】SSNAPS: Audio-Visual Separation of Speech and Background Noise with Diffusion Inverse Sampling
标题:SSNAPS:使用扩散反采样的语音和背景噪音的视听分离
链接:https://arxiv.org/abs/2602.01394
摘要:本文讨论了在真实环境噪声存在下的视听单麦克风语音分离和增强的挑战。我们的方法是基于生成式逆采样,我们用专用的扩散先验对干净的语音和环境噪声进行建模,并联合利用它们来恢复所有潜在的源。为了实现这一点,我们重新制定了最近的逆采样器,以匹配我们的设置。我们对1,2和3个扬声器与噪声的混合物进行了评估,并表明,尽管完全无监督,但我们的方法在所有条件下都始终优于\ac{WER}中的领先监督基线。我们进一步扩展我们的框架来处理屏幕外扬声器分离。此外,分离的噪声分量的高保真度使其适合于下游声学场景检测。演示页面:https://ssnapsicml.github.io/ssnapsicml2026/
摘要:This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise. Our approach is based on generative inverse sampling, where we model clean speech and ambient noise with dedicated diffusion priors and jointly leverage them to recover all underlying sources. To achieve this, we reformulate a recent inverse sampler to match our setting. We evaluate on mixtures of 1, 2, and 3 speakers with noise and show that, despite being entirely unsupervised, our method consistently outperforms leading supervised baselines in \ac{WER} across all conditions. We further extend our framework to handle off-screen speaker separation. Moreover, the high fidelity of the separated noise component makes it suitable for downstream acoustic scene detection. Demo page: https://ssnapsicml.github.io/ssnapsicml2026/
【6】Generative AI in Signal Processing Education: An Audio Foundation Model Based Approach
标题:信号处理教育中的生成性人工智能:基于音频基础模型的方法
链接:https://arxiv.org/abs/2602.01249
备注:accepted at IEEE EDUCON 2026
摘要:音频基础模型(AFMs)是生成人工智能(GenAI)的一个专门类别,通过集成语音和音频增强、去噪、源分离、特征提取、自动等核心应用程序,有可能改变信号处理(SP)教育。分类和实时信号分析到学习和研究中。本文介绍了SPEduAFM,一个概念性的AFM为SP教育量身定制,桥接传统的SP原则与GenAI驱动的创新。通过一个设想的案例研究,我们概述了AFMs如何能够实现一系列应用,包括自动演讲转录,交互式演示和包容性学习工具,展示了它们将抽象概念转化为引人入胜的实践经验的潜力。本文还通过强调动态,实时的听觉互动,促进体验式和真实的学习,解决了道德,可解释性和定制等挑战。通过将SPEduAFM作为一个前瞻性的愿景,我们的目标是激发GenAI在工程教育中的更广泛采用,增强课堂内外的可访问性,参与度和创新。
摘要:Audio Foundation Models (AFMs), a specialized category of Generative AI (GenAI), have the potential to transform signal processing (SP) education by integrating core applications such as speech and audio enhancement, denoising, source separation, feature extraction, automatic classification, and real-time signal analysis into learning and research. This paper introduces SPEduAFM, a conceptual AFM tailored for SP education, bridging traditional SP principles with GenAI-driven innovations. Through an envisioned case study, we outline how AFMs can enable a range of applications, including automated lecture transcription, interactive demonstrations, and inclusive learning tools, showcasing their potential to transform abstract concepts into engaging, practical experiences. This paper also addresses challenges such as ethics, explainability, and customization by highlighting dynamic, real-time auditory interactions that foster experiential and authentic learning. By presenting SPEduAFM as a forward-looking vision, we aim to inspire broader adoption of GenAI in engineering education, enhancing accessibility, engagement, and innovation in the classroom and beyond.
【7】Adapting Where It Matters: Depth-Aware Adaptation for Efficient Multilingual Speech Recognition in Low-Resource Languages
标题:适应重要的地方:深度感知适应,在低资源语言中实现高效多语言语音识别
链接:https://arxiv.org/abs/2602.01008
备注:13 pages
摘要:最近的语音基础模型在高资源语言的多语言自动语音识别(ASR)方面表现出色,但由于数据稀缺和效率限制,使其适应低资源语言仍然具有挑战性。全模型微调在计算上是昂贵的,并且容易过拟合,而像LoRA这样的参数高效方法在层间均匀地应用自适应,忽略了内部表示,从而影响了有效性和效率。我们分析了多语言ASR模型,并揭示了一个U形的适应性模式:早期和晚期层是语言特定的,需要更多的调整,而中间层保留共享的语义,需要更少。基于这一观察,我们提出了DAMA,深度感知模型适应框架,根据每个层的角色分配适应能力。DAMA还引入了基于奇异值分解(SVD)的初始化来约束自适应并保留U形模式,以及冻结的中间层基础以进一步提高效率。在两个基准数据集上对18种低资源语言进行了评估,DAMA匹配或超越了最先进的准确性,可训练参数减少了80%,在极端数据稀缺的情况下实现了29%的错误减少,并显着提高了内存,训练时间和计算效率。这些结果突出了结构感知适应高效,可扩展的多语言ASR的好处。
摘要:Recent speech foundation models excel at multilingual automatic speech recognition (ASR) for high-resource languages, but adapting them to low-resource languages remains challenging due to data scarcity and efficiency constraints. Full-model fine-tuning is computationally expensive and prone to overfitting, while parameter-efficient methods like LoRA apply adaptation uniformly across layers, overlooking internal representations thus compromising effectiveness and efficiency. We analyze multilingual ASR models and reveal a U-shaped adaptability pattern: early and late layers are language-specific and require more adaptation, while intermediate layers retain shared semantics and need less. Building on this observation, we propose DAMA, a Depth-Aware Model Adaptation framework that allocates adaptation capacity according to each layer's role. DAMA also introduces Singular Value Decomposition (SVD)-based initialization to constrain adaptation and preserve the U-shaped pattern, as well as a frozen middle-layer basis for further efficiency. Evaluated on 18 low-resource languages across two benchmark datasets, DAMA matches or surpasses state-of-the-art accuracy with 80% fewer trainable parameters, achieves a 29% error reduction under extreme data scarcity, and significantly improves memory, training time, and computational efficiency over baselines. These results highlight the benefits of structure-aware adaptation for efficient, scalable multilingual ASR.
【8】Solving Room Impulse Response Inverse Problems Using Flow Matching with Analytic Wiener Denoiser
标题:利用解析维纳降噪器的流量匹配解决房间脉冲响应反问题
链接:https://arxiv.org/abs/2602.00652
备注:Submitted to the Journal of the Acoustical Society of America (JASA)
摘要:房间脉冲响应(RIR)估计自然地作为一类逆问题出现,包括去噪和反卷积。虽然最近的方法通常依赖于监督学习或学习的生成先验,但这些方法需要大量的训练数据,并且在训练分布之外可能推广得很差。在这项工作中,我们提出了RIRFlow,使用流匹配的RIR逆问题的无训练贝叶斯框架。我们从RIR的统计结构中推导出一个流一致的分析先验,消除了对数据驱动先验的需要。具体来说,我们的模型RIR作为一个高斯过程的指数衰减方差,从而产生一个封闭形式的最小均方误差(MMSE)维纳降噪。该分析去噪器被集成为现有的基于流的逆求解器中的先验,其中逆问题通过引导后验采样来解决。此外,我们通过局部高斯近似的指导后,扩展求解器的非线性和非高斯逆问题,并实证证明,这种近似在实践中仍然有效。在不同逆问题上的真实RIR实验表明了鲁棒性能,突出了经典RIR模型与最近基于流的生成推理相结合的有效性。
摘要:Room impulse response (RIR) estimation naturally arises as a class of inverse problems, including denoising and deconvolution. While recent approaches often rely on supervised learning or learned generative priors, such methods require large amounts of training data and may generalize poorly outside the training distribution. In this work, we present RIRFlow, a training-free Bayesian framework for RIR inverse problems using flow matching. We derive a flow-consistent analytic prior from the statistical structure of RIRs, eliminating the need for data-driven priors. Specifically, we model RIR as a Gaussian process with exponentially decaying variance, which yields a closed-form minimum mean squared error (MMSE) Wiener denoiser. This analytic denoiser is integrated as a prior in an existing flow-based inverse solver, where inverse problems are solved via guided posterior sampling. Furthermore, we extend the solver to nonlinear and non-Gaussian inverse problems via a local Gaussian approximation of the guided posterior, and empirically demonstrate that this approximation remains effective in practice. Experiments on real RIRs across different inverse problems demonstrate robust performance, highlighting the effectiveness of combining a classic RIR model with the recent flow-based generative inference.
【9】High-Fidelity Generative Audio Compression at 0.275kbps
标题:高保真生成音频压缩,速度为0.275kMbps
链接:https://arxiv.org/abs/2602.00648
备注:Technical Report
摘要:超低比特率的高保真通用音频压缩对于从低带宽通信到生成音频语言建模的应用至关重要。传统的音频压缩方法和当代神经编解码器基本上是为波形重构而设计的。因此,当在超低比特率下操作时,这些方法迅速降级并且通常无法保留基本信息,导致严重的声学伪影和明显的语义失真。为了克服这些限制,我们引入了生成音频压缩(GAC),一种新的范式转变,从信号保真度的任务导向的有效性。GAC在AI Flow框架内实现,理论上以信息容量定律为基础。这些基础表明,丰富的计算能力可以在接收器处被利用来抵消极端的通信瓶颈--这体现了更多计算,更少带宽的理念。通过将发送方的语义理解与接收方的可扩展生成合成相结合,GAC将信息负担卸载到强大的模型先验中。我们的1.8B参数模型在0.275kbps的比特率下实现了32 kHz普通音频的高保真重建。即使在0.175kbps下,它仍然保持了强大的可理解的音频传输能力,代表了约3000倍的压缩比,在保持感知质量和语义一致性方面明显优于当前最先进的神经编解码器。
摘要:High-fidelity general audio compression at ultra-low bitrates is crucial for applications ranging from low-bandwidth communication to generative audio-language modeling. Traditional audio compression methods and contemporary neural codecs are fundamentally designed for waveform reconstruction. As a result, when operating at ultra-low bitrates, these methods degrade rapidly and often fail to preserve essential information, leading to severe acoustic artifacts and pronounced semantic distortion. To overcome these limitations, we introduce Generative Audio Compression (GAC), a novel paradigm shift from signal fidelity to task-oriented effectiveness. Implemented within the AI Flow framework, GAC is theoretically grounded in the Law of Information Capacity. These foundations posit that abundant computational power can be leveraged at the receiver to offset extreme communication bottlenecks--exemplifying the More Computation, Less Bandwidth philosophy. By integrating semantic understanding at the transmitter with scalable generative synthesis at the receiver, GAC offloads the information burden to powerful model priors. Our 1.8B-parameter model achieves high-fidelity reconstruction of 32kHz general audio at an unprecedented bitrate of 0.275kbps. Even at 0.175kbps, it still preserves a strong intelligible audio transmission capability, which represents an about 3000x compression ratio, significantly outperforming current state-of-the-art neural codecs in maintaining both perceptual quality and semantic consistency.
【10】QuietPrint: Protecting 3D Printers Against Acoustic Side-Channel Attacks
标题:QuietPrint:保护3D收件箱免受声学侧通道攻击
链接:https://arxiv.org/abs/2602.02198
摘要:近年来,3D打印市场经历了显着增长,预计2025年的收入将达到150亿美元。针对3D打印过程的网络攻击,无论是通过机器本身、供应链还是制造组件,都变得越来越普遍。一个主要的问题是知识产权(IP)盗窃,其中恶意攻击者获得对设计文件的访问权。进行这种盗窃的一种方法是通过侧信道攻击。在这项工作中,我们调查了通过声学侧信道进行IP盗窃的可能性,并提出了一种新的方法来保护3D打印机免受此类攻击。我们的方法的主要优点是它不需要额外的硬件,如大型扬声器或降噪设备。相反,它通过对G代码进行最小限度的修改来保护打印部件。
摘要:The 3D printing market has experienced significant growth in recent years, with an estimated revenue of 15 billion USD for 2025. Cyber-attacks targeting the 3D printing process whether through the machine itself, the supply chain, or the fabricated components are becoming increasingly common. One major concern is intellectual property (IP) theft, where a malicious attacker gains access to the design file. One method for carrying out such theft is through side-channel attacks. In this work, we investigate the possibility of IP theft via acoustic side channels and propose a novel method to protect 3D printers against such attacks. The primary advantage of our approach is that it requires no additional hardware, such as large speakers or noise-canceling devices. Instead, it secures printed parts by minimal modifications to the G-code.
【11】Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
标题:用于语音情感识别的大型音频语言模型中的注意力加权中心核对齐知识提取
链接:https://arxiv.org/abs/2602.01547
备注:Accepted to 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)
摘要:大型音频语言模型(LALM)的出现促进了语音情感识别(SER)的发展,但其规模限制了在资源受限环境中的部署。虽然知识蒸馏对于LALM压缩是有效的,但是现有方法在提取跨模态投影模块(投影仪)方面仍然探索不足,并且由于特征维度的差异而经常与对齐斗争。我们提出PL-Distill,一个KD框架,它结合了投影仪级蒸馏(PDist)来对齐音频嵌入和逻辑级蒸馏(LDist)来对齐输出逻辑。PDist介绍了注意力加权中心内核对齐,这是我们提出的一种新方法,可以突出重要的时间步长并解决维度不匹配问题。与此同时,LDist最大限度地减少了教师和学生之间的Kullback-Leibler分歧从音频和文本形式的逻辑。在IEMOCAP、RAVDESS和SAVEE上,PL-Distill将8.4B参数的教师压缩为紧凑的1.1B参数的学生,在所有指标上始终优于教师、最先进的预训练模型和其他KD基线。
摘要:The emergence of Large Audio-Language Models (LALMs) has advanced Speech Emotion Recognition (SER), but their size limits deployment in resource-constrained environments. While Knowledge Distillation is effective for LALM compression, existing methods remain underexplored in distilling the cross-modal projection module (Projector), and often struggle with alignment due to differences in feature dimensions. We propose PL-Distill, a KD framework that combines Projector-Level Distillation (PDist) to align audio embeddings and Logits-Level Distillation (LDist) to align output logits. PDist introduces Attention-weighted Centered Kernel Alignment, a novel approach we propose to highlight important time steps and address dimension mismatches. Meanwhile, LDist minimizes the Kullback-Leibler divergence between teacher and student logits from audio and text modalities. On IEMOCAP, RAVDESS, and SAVEE, PL-Distill compresses an 8.4B-parameter teacher to a compact 1.1B-parameter student, consistently outperforming the teacher, state-of-the-art pretrained models, and other KD baselines across all metrics.
【12】Causally Disentangled Contrastive Learning for Multilingual Speaker Embeddings
标题:多语言说话人嵌入的因果分离对比学习
链接:https://arxiv.org/abs/2602.01363
摘要:自监督说话人嵌入被广泛用于说话人验证系统,但先前的工作表明,他们往往编码敏感的人口统计属性,提高公平性和隐私问题。本文研究了人口统计信息,特别是性别,年龄和口音,在SimCLR训练的说话人嵌入中存在的程度,以及是否可以在不严重降低说话人验证性能的情况下减轻这种泄漏。我们研究了两种去偏策略:通过梯度反转进行对抗性训练,以及明确分离人口统计信息和剩余信息的因果瓶颈架构。人口泄漏量化使用线性和非线性探测分类器,而扬声器验证性能使用ROC AUC和EER进行评估。我们的研究结果表明,性别信息是强烈的,线性编码的基线嵌入,而年龄和口音较弱,主要是非线性表示。对抗性去偏见减少了性别泄漏,但对年龄和口音的影响有限,并与验证准确性进行了明确的权衡。因果瓶颈进一步抑制了人口统计信息,特别是在残差表示中,但会导致性能大幅下降。这些研究结果强调了在减轻自我监督的说话人嵌入中的人口泄漏方面的根本局限性,并澄清了当前去偏方法中固有的权衡。
摘要:Self-supervised speaker embeddings are widely used in speaker verification systems, but prior work has shown that they often encode sensitive demographic attributes, raising fairness and privacy concerns. This paper investigates the extent to which demographic information, specifically gender, age, and accent, is present in SimCLR-trained speaker embeddings and whether such leakage can be mitigated without severely degrading speaker verification performance. We study two debiasing strategies: adversarial training through gradient reversal and a causal bottleneck architecture that explicitly separates demographic and residual information. Demographic leakage is quantified using both linear and nonlinear probing classifiers, while speaker verification performance is evaluated using ROC-AUC and EER. Our results show that gender information is strongly and linearly encoded in baseline embeddings, whereas age and accent are weaker and primarily nonlinearly represented. Adversarial debiasing reduces gender leakage but has limited effect on age and accent and introduces a clear trade-off with verification accuracy. The causal bottleneck further suppresses demographic information, particularly in the residual representation, but incurs substantial performance degradation. These findings highlight fundamental limitations in mitigating demographic leakage in self-supervised speaker embeddings and clarify the trade-offs inherent in current debiasing approaches.
【13】TLDiffGAN: A Latent Diffusion-GAN Framework with Temporal Information Fusion for Anomalous Sound Detection
标题:TLDifGAN:一个具有时间信息融合的潜在扩散-GAN框架,用于异常声音检测
链接:https://arxiv.org/abs/2602.01060
备注:Accepted by ICASSP 2026
摘要:现有的生成模型的无监督异常声音检测是有限的,他们无法完全捕捉正常声音的复杂特征分布,而在这一领域的强大的扩散模型的潜力仍然在很大程度上未被开发。为了应对这一挑战,我们提出了一个新的框架,TLDiffGAN,它由两个互补的分支。一个分支将潜在扩散模型并入GAN生成器中进行对抗训练,从而使训练器的任务更具挑战性并提高生成样本的质量。另一个分支利用预训练的音频模型编码器直接从原始音频波形中提取特征用于辅助辨别。该框架有效地从原始音频和Mel频谱图中捕获正常声音的特征表示。此外,我们引入了TMixup声谱图增强技术,以提高敏感性,往往被忽视的微妙和局部的时间模式。在DCASE 2020 Challenge Task 2数据集上的大量实验证明了TLDiffGAN的优越检测性能,以及其在异常时频定位方面的强大能力。
摘要:Existing generative models for unsupervised anomalous sound detection are limited by their inability to fully capture the complex feature distribution of normal sounds, while the potential of powerful diffusion models in this domain remains largely unexplored. To address this challenge, we propose a novel framework, TLDiffGAN, which consists of two complementary branches. One branch incorporates a latent diffusion model into the GAN generator for adversarial training, thereby making the discriminator's task more challenging and improving the quality of generated samples. The other branch leverages pretrained audio model encoders to extract features directly from raw audio waveforms for auxiliary discrimination. This framework effectively captures feature representations of normal sounds from both raw audio and Mel spectrograms. Moreover, we introduce a TMixup spectrogram augmentation technique to enhance sensitivity to subtle and localized temporal patterns that are often overlooked. Extensive experiments on the DCASE 2020 Challenge Task 2 dataset demonstrate the superior detection performance of TLDiffGAN, as well as its strong capability in anomalous time-frequency localization.
【14】HierCon: Hierarchical Contrastive Attention for Audio Deepfake Detection
标题:HierCon:音频Deepfake检测的分层对比注意力
链接:https://arxiv.org/abs/2602.01032
备注:Proceedings of The Web Conference 2026 (WWW'26), short track
摘要:现代TTS和语音转换系统生成的音频深度伪造越来越难以与真实语音区分开来,这给安全和在线信任带来了严重风险。虽然最先进的自监督模型提供了丰富的多层表示,但现有的检测器独立地处理层,忽略了对识别合成伪影至关重要的时间和层次依赖性。我们提出了HierCon,一个分层注意力框架,结合了基于边缘的对比学习,该框架对时间帧,相邻层和层组之间的依赖关系进行建模,同时鼓励域不变嵌入。在ASVspoof 2021 DF和In-the-Wild数据集上进行评估,我们的方法实现了最先进的性能(1.93%和6.87%EER),比独立层加权分别提高了36.6%和22.5%。结果和注意力可视化证实,分层建模增强了跨域生成技术和记录条件的概括。
摘要:Audio deepfakes generated by modern TTS and voice conversion systems are increasingly difficult to distinguish from real speech, raising serious risks for security and online trust. While state-of-the-art self-supervised models provide rich multi-layer representations, existing detectors treat layers independently and overlook temporal and hierarchical dependencies critical for identifying synthetic artefacts. We propose HierCon, a hierarchical layer attention framework combined with margin-based contrastive learning that models dependencies across temporal frames, neighbouring layers, and layer groups, while encouraging domain-invariant embeddings. Evaluated on ASVspoof 2021 DF and In-the-Wild datasets, our method achieves state-of-the-art performance (1.93% and 6.87% EER), improving over independent layer weighting by 36.6% and 22.5% respectively. The results and attention visualisations confirm that hierarchical modelling enhances generalisation to cross-domain generation techniques and recording conditions.
【15】Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
标题:耳朵里的偏见:评估跨语言、人口和位置变化的音频语言模型的敏感性
链接:https://arxiv.org/abs/2602.01030
备注:Accepted as a long findings paper at EACL 2026
摘要:这项工作提出了第一个系统的调查,在多语言MLLM的言语偏见。我们构建并发布了BiasInEar数据集,这是一个基于Global MMLU Lite的语音增强基准测试,涵盖英语,中文和韩语,按性别和口音平衡,总计70.8小时(约4,249分钟)的语音,包含11,200个问题。使用四个互补的指标(准确性,熵,APES和Fleiss的$κ$),我们评估了九个代表性的模型在语言(语言和口音),人口(性别)和结构(选项顺序)扰动。我们的研究结果表明,MLLM是相对强大的人口因素,但高度敏感的语言和选项顺序,这表明,讲话可以放大现有的结构性偏见。此外,架构设计和推理策略大大影响跨语言的鲁棒性。总的来说,本研究建立了一个统一的框架,用于评估语音集成LLM的公平性和鲁棒性,弥合了基于文本和基于语音的评估之间的差距。这些资源可以在https://github.com/ntunlplab/BiasInEar上找到。
摘要:This work presents the first systematic investigation of speech bias in multilingual MLLMs. We construct and release the BiasInEar dataset, a speech-augmented benchmark based on Global MMLU Lite, spanning English, Chinese, and Korean, balanced by gender and accent, and totaling 70.8 hours ($\approx$4,249 minutes) of speech with 11,200 questions. Using four complementary metrics (accuracy, entropy, APES, and Fleiss' $κ$), we evaluate nine representative models under linguistic (language and accent), demographic (gender), and structural (option order) perturbations. Our findings reveal that MLLMs are relatively robust to demographic factors but highly sensitive to language and option order, suggesting that speech can amplify existing structural biases. Moreover, architectural design and reasoning strategy substantially affect robustness across languages. Overall, this study establishes a unified framework for assessing fairness and robustness in speech-integrated LLMs, bridging the gap between text- and speech-based evaluation. The resources can be found at https://github.com/ntunlplab/BiasInEar.
【16】A Baseline Multimodal Approach to Emotion Recognition in Conversations
标题:对话中情感识别的基线多模式方法
链接:https://arxiv.org/abs/2602.00914
备注:10 pages
摘要:我们提出了一个轻量级的多模态基线的情感识别会话中使用SemEval-2024任务3数据集从情景喜剧朋友。本报告的目标不是提出一种新的最先进的方法,而是记录一种可访问的参考实现,该实现结合了(i)基于transformer的文本分类器和(ii)自监督语音表示模型,以及一个简单的后期融合集成。我们报告了在有限的训练协议下获得的基线设置和经验结果,突出了多模态融合在单峰模型上的改进。提供此预印本是为了提高透明度,并支持未来更严格的比较。
摘要:We present a lightweight multimodal baseline for emotion recognition in conversations using the SemEval-2024 Task 3 dataset built from the sitcom Friends. The goal of this report is not to propose a novel state-of-the-art method, but to document an accessible reference implementation that combines (i) a transformer-based text classifier and (ii) a self-supervised speech representation model, with a simple late-fusion ensemble. We report the baseline setup and empirical results obtained under a limited training protocol, highlighting when multimodal fusion improves over unimodal models. This preprint is provided for transparency and to support future, more rigorous comparisons.
【17】The TMU System for the XACLE Challenge: Training Large Audio Language Models with CLAP Pseudo-Labels
标题:XACLE挑战赛TMU系统:使用CLAP伪标签训练大型音频语言模型
链接:https://arxiv.org/abs/2602.00604
备注:3 pages; 2 figures; 2 tables; Accepted at ICASSP 2026 Workshop (SP Grand Challenges, GC-12: XACLE)
摘要:在本文中,我们提出了一个提交x到音频对齐(XACLE)的挑战。目标是预测给定的一般音频和文本对的语义对齐。该系统是基于一个大的音频语言模型(LALM)架构。我们采用了三个阶段的训练管道:自动音频字幕预训练,CLAP伪标签预训练,以及对XACLE数据集进行微调。我们的实验表明,使用CLAP伪标签进行预训练是主要的性能驱动因素。在XACLE测试集上,我们的系统达到了0.632的SRCC,显著优于基线系统(0.334),并在挑战团队排名中获得第三名。代码和模型可以在https://github.com/shiotalab-tmu/tmu-xacle2026上找到
摘要:In this paper, we propose a submission to the x-to-audio alignment (XACLE) challenge. The goal is to predict semantic alignment of a given general audio and text pair. The proposed system is based on a large audio language model (LALM) architecture. We employ a three-stage training pipeline: automated audio captioning pretraining, pretraining with CLAP pseudo-labels, and fine-tuning on the XACLE dataset. Our experiments show that pretraining with CLAP pseudo-labels is the primary performance driver. On the XACLE test set, our system reaches an SRCC of 0.632, significantly outperforming the baseline system (0.334) and securing third place in the challenge team ranking. Code and models can be found at https://github.com/shiotalab-tmu/tmu-xacle2026
【18】Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
标题:Kanade:一个用于口语建模的简单解开令牌器
链接:https://arxiv.org/abs/2602.00594
摘要:一个好的语言模型始于一个好的标记器。标记化对于语音建模尤其重要,因为语音建模必须处理混合了语言和非语言信息的连续信号。语音分词器应该提取语音和韵律,抑制语言上不相关的信息,如说话人身份,并实现高质量的合成。我们提出了Kanade,一个单层的解开语音标记器,实现了这一理想。Kanade将声学常数分离出来,创建一个单一的令牌流,可以捕获丰富的语音和韵律。实验表明,Kanade在保持良好重建质量的同时,实现了最先进的说话人解缠和词汇可用性。
摘要:A good language model starts with a good tokenizer. Tokenization is especially important for speech modeling, which must handle continuous signals that mix linguistic and non-linguistic information. A speech tokenizer should extract phonetics and prosody, suppress linguistically irrelevant information like speaker identity, and enable high-quality synthesis. We present Kanade, a single-layer disentangled speech tokenizer that realizes this ideal. Kanade separates out acoustic constants to create a single stream of tokens that captures rich phonetics and prosody. It does so without the need for auxiliary methods that existing disentangled codecs often rely on. Experiments show that Kanade achieves state-of-the-art speaker disentanglement and lexical availability, while maintaining excellent reconstruction quality.
【19】Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy
标题:双视图预测扩散:通过谱图-图像协同实现轻量级语音增强
链接:https://arxiv.org/abs/2602.00568
摘要:扩散模型最近在语音增强(SE)中设定了新的基准。然而,大多数现有的基于分数的模型仅将语音频谱图视为通用的2D图像,应用忽略音频的固有结构稀疏性的统一处理,这导致低效的频谱表示和过高的计算复杂度。为了弥合这一差距,我们提出了DVPD,一个非常轻量级的双视图预测扩散模型,它独特地利用了频谱图的双重性质,即在训练和推理阶段都是视觉纹理和物理频域表示。具体来说,在训练过程中,我们通过频率自适应非均匀压缩(FANC)编码器优化频谱利用率,该编码器在修剪高频冗余的同时保留了关键的低频谐波。同时,我们引入了一个轻量级的基于图像的光谱感知(LISA)模块,以最小的开销从视觉角度捕捉功能。在推理过程中,我们提出了一个无训练的无损提升(TLB)策略,利用相同的双视图先验来改进生成质量,而无需任何额外的微调。在各种基准测试中进行的大量实验表明,与SOTA轻量级模型PGUSE相比,DVPD实现了最先进的性能,同时只需要35%的参数和40%的推理MAC。这些结果突出了DVPD在平衡高保真语音质量与极端架构效率方面的卓越能力。代码和音频示例可在匿名网站上获得:{https://anonymous.4open.science/r/dvpd_demo-E630}
摘要:Diffusion models have recently set new benchmarks in Speech Enhancement (SE). However, most existing score-based models treat speech spectrograms merely as generic 2D images, applying uniform processing that ignores the intrinsic structural sparsity of audio, which results in inefficient spectral representation and prohibitive computational complexity. To bridge this gap, we propose DVPD, an extremely lightweight Dual-View Predictive Diffusion model, which uniquely exploits the dual nature of spectrograms as both visual textures and physical frequency-domain representations across both training and inference stages. Specifically, during training, we optimize spectral utilization via the Frequency-Adaptive Non-uniform Compression (FANC) encoder, which preserves critical low-frequency harmonics while pruning high-frequency redundancies. Simultaneously, we introduce a Lightweight Image-based Spectro-Awareness (LISA) module to capture features from a visual perspective with minimal overhead. During inference, we propose a Training-free Lossless Boost (TLB) strategy that leverages the same dual-view priors to refine generation quality without any additional fine-tuning. Extensive experiments across various benchmarks demonstrate that DVPD achieves state-of-the-art performance while requiring only 35% of the parameters and 40% of the inference MACs compared to SOTA lightweight model, PGUSE. These results highlight DVPD's superior ability to balance high-fidelity speech quality with extreme architectural efficiency. Code and audio samples are available at the anonymous website: {https://anonymous.4open.science/r/dvpd_demo-E630}
【20】Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
标题:编辑内容,保留声音:通过自一致性奖励进行基于文本的语音编辑
链接:https://arxiv.org/abs/2602.00560
摘要:不可感知的基于文本的语音编辑允许用户通过改变转录来修改口语内容。它要求修改后的片段与周围的上下文无缝融合。在声学空间中操作的流行方法遭受固有的内容风格纠缠,导致生成不稳定性和边界伪影。在本文中,我们提出了一个新的框架接地的原则“编辑内容,保留声学”。我们的方法依赖于两个核心组件:(1)结构基础,它将编辑纳入稳定的语义空间,同时将声学重建委托给流匹配解码器;(2)感知对齐,它采用了一种新的自一致性奖励组相对策略优化。通过利用预先训练的文本到语音转换模型作为隐式批评者-辅以严格的可理解性和持续时间限制-我们有效地将编辑的语义标记序列与原始上下文对齐。经验评估表明,我们的方法显着优于最先进的自回归和非自回归基线,实现卓越的可懂度,鲁棒性和感知质量。
摘要:Imperceptible text-based speech editing allows users to modify spoken content by altering the transcript. It demands that modified segments fuse seamlessly with the surrounding context. Prevalent methods operating in the acoustic space suffer from inherent content-style entanglement, leading to generation instability and boundary artifacts. In this paper, we propose a novel framework grounded in the principle of "Edit Content, Preserve Acoustics". Our approach relies on two core components: (1) Structural Foundations, which decouples editing into a stable semantic space while delegating acoustic reconstruction to a Flow Matching decoder; and (2) Perceptual Alignment, which employs a novel Self-Consistency Rewards Group Relative Policy Optimization. By leveraging a pre-trained Text-to-Speech model as an implicit critic -- complemented by strict intelligibility and duration constraints -- we effectively align the edited semantic token sequence with the original context. Empirical evaluations demonstrate that our method significantly outperforms state-of-the-art autoregressive and non-autoregressive baselines, achieving superior intelligibility, robustness, and perceptual quality.
【21】RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
标题:RVCBench:对现代音频生成模型中语音克隆的稳健性进行基准测试
链接:https://arxiv.org/abs/2602.00443
备注:40 pages, 12figures
摘要:现代语音克隆(VC)可以仅从几秒钟的参考音频中合成与目标说话人非常匹配的语音,从而实现个性化语音界面和配音等应用。在实际部署中,现代音频生成模型不可避免地会遇到嘈杂的参考音频、不完美的文本提示和多样化的下游处理,这会严重损害鲁棒性。尽管VC在自回归编解码器-令牌语言模型和基于扩散的模型的驱动下取得了快速进展,但现实部署变化下的鲁棒性仍然未得到充分探索。本文介绍了RVCBench,这是一个全面的基准测试,用于评估整个生成管道中VC的鲁棒性,包括输入变化,生成挑战,输出后处理和对抗性扰动,涵盖10个鲁棒性任务,225个扬声器,14,370个话语和11个代表性的现代VC模型。我们的评估揭示了VC中的大量鲁棒性差距:在常见的输入移位和后处理下,性能可能会急剧恶化;长上下文和跨语言场景进一步暴露了稳定性限制;被动噪声和主动扰动都影响生成鲁棒性。总的来说,这些发现提供了一个统一的图片,目前的VC模型在实践中失败,并引入了一个标准化的,开源的测试平台,以支持更强大的和可部署的VC模型的开发。我们在https://github.com/Nanboy-Ronan/RVCBench上开源了我们的项目。
摘要:Modern voice cloning (VC) can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling applications such as personalized speech interfaces and dubbing. In practical deployments, modern audio generation models inevitably encounter noisy reference audios, imperfect text prompts, and diverse downstream processing, which can significantly hurt robustness. Despite rapid progress in VC driven by autoregressive codec-token language models and diffusion-based models, robustness under realistic deployment shifts remains underexplored. This paper introduces RVCBench, a comprehensive benchmark that evaluates Robustness in VC across the full generation pipeline, including input variation, generation challenges, output post-processing, and adversarial perturbations, covering 10 robustness tasks, 225 speakers, 14,370 utterances, and 11 representative modern VC models. Our evaluation uncovers substantial robustness gaps in VC: performance can deteriorate sharply under common input shifts and post-processing; long-context and cross-lingual scenarios further expose stability limitations; and both passive noise and proactive perturbation influence generation robustness. Collectively, these findings provide a unified picture of how current VC models fail in practice and introduce a standardized, open-source testbed to support the development of more robust and deployable VC models. We open-source our project at https://github.com/Nanboy-Ronan/RVCBench.
【22】Multi-Speaker Conversational Audio Deepfake: Taxonomy, Dataset and Pilot Study
标题:多扬声器对话音频Deepfake:分类、数据集和试点研究
链接:https://arxiv.org/abs/2602.00295
备注:This work was presented at the 2025 IEEE International Conference on Data Mining, ICDM 2025, November 12-15,2025, Washington DC, USA
摘要:文本到语音(TTS)技术的快速发展使音频deepfake变得越来越真实和可访问,引发了重大的安全和信任问题。虽然现有的研究主要集中在检测单扬声器音频deepfake,但具有多扬声器会话设置的真实恶意应用程序也正在成为一个未被充分研究的主要威胁。为了解决这一差距,我们提出了一种多说话者对话音频deepfakes的概念分类,区分部分操作(一个或多个说话者改变)和完全操作(整个对话合成)。作为第一步,我们引入了一个新的多扬声器会话音频Deepfakes数据集(MsCADD),其中包含2,830个音频片段,这些音频片段包含真实和完全合成的两个扬声器对话,使用基于VITS和SoundStorm的NotebookLM模型生成,以模拟具有扬声器性别变化的自然对话和会话自发性。MsCADD仅限于文本到语音(TTS)类型的deepfake。我们在这个数据集上对三个神经基线模型进行了基准测试:LFCC-LCNN,RawNet 2和Wav 2 Vec 2.0,并报告了F1评分,准确性,真阳性率(TPR)和真阴性率(TNR)方面的性能。结果表明,这些基线模型提供了一个有用的基准,然而,结果也强调了多说话者deepfake研究在不同会话动态下可靠检测合成语音方面存在显着差距。我们的数据集和基准测试为未来在对话场景中进行deepfake检测的研究奠定了基础,这是一个高度未开发的研究领域,但也是对音频设置中的可信信息构成威胁的主要领域。MsCADD数据集是公开的,以支持研究社区的可重复性和基准测试。
摘要:The rapid advances in text-to-speech (TTS) technologies have made audio deepfakes increasingly realistic and accessible, raising significant security and trust concerns. While existing research has largely focused on detecting single-speaker audio deepfakes, real-world malicious applications with multi-speaker conversational settings is also emerging as a major underexplored threat. To address this gap, we propose a conceptual taxonomy of multi-speaker conversational audio deepfakes, distinguishing between partial manipulations (one or multiple speakers altered) and full manipulations (entire conversations synthesized). As a first step, we introduce a new Multi-speaker Conversational Audio Deepfakes Dataset (MsCADD) of 2,830 audio clips containing real and fully synthetic two-speaker conversations, generated using VITS and SoundStorm-based NotebookLM models to simulate natural dialogue with variations in speaker gender, and conversational spontaneity. MsCADD is limited to text-to-speech (TTS) types of deepfake. We benchmark three neural baseline models; LFCC-LCNN, RawNet2, and Wav2Vec 2.0 on this dataset and report performance in terms of F1 score, accuracy, true positive rate (TPR), and true negative rate (TNR). Results show that these baseline models provided a useful benchmark, however, the results also highlight that there is a significant gap in multi-speaker deepfake research in reliably detecting synthetic voices under varied conversational dynamics. Our dataset and benchmarks provide a foundation for future research on deepfake detection in conversational scenarios, which is a highly underexplored area of research but also a major area of threat to trustworthy information in audio settings. The MsCADD dataset is publicly available to support reproducibility and benchmarking by the research community.
【23】VoxServe: Streaming-Centric Serving System for Speech Language Models
标题:VoxServe:以流媒体为中心的语音语言模型服务系统
链接:https://arxiv.org/abs/2602.00269
备注:The code is available at https://github.com/vox-serve/vox-serve
摘要:在流媒体环境中部署现代语音语言模型(SpeechLM)需要系统提供低延迟、高吞吐量和强有力的流媒体性保证。现有的系统不能灵活有效地支持不同的模型。我们提出VoxServe,一个统一的SpeechLM服务系统,优化流媒体性能。VoxServe引入了一种模型执行抽象,将模型架构与系统级优化相结合,从而在单个框架内支持多种SpeechLM架构。在此抽象的基础上,VoxServe实现了流感知调度和异步推理管道,以提高端到端的效率。对多个现代SpeechLM的评估表明,VoxServe在相当延迟的情况下实现了比现有实现高10- 20倍的吞吐量,同时保持了高的流可行性。VoxServe的代码可在https://github.com/vox-serve/vox-serve上获得。
摘要:Deploying modern Speech Language Models (SpeechLMs) in streaming settings requires systems that provide low latency, high throughput, and strong guarantees of streamability. Existing systems fall short of supporting diverse models flexibly and efficiently. We present VoxServe, a unified serving system for SpeechLMs that optimizes streaming performance. VoxServe introduces a model-execution abstraction that decouples model architecture from system-level optimizations, thereby enabling support for diverse SpeechLM architectures within a single framework. Building on this abstraction, VoxServe implements streaming-aware scheduling and an asynchronous inference pipeline to improve end-to-end efficiency. Evaluations across multiple modern SpeechLMs show that VoxServe achieves 10-20x higher throughput than existing implementations at comparable latency while maintaining high streaming viability. The code of VoxServe is available at https://github.com/vox-serve/vox-serve.
【24】LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
标题:LPIPS-AttnWave 2Lip:适合野外会说话的头一代的通用音频驱动嘴唇同步
链接:https://arxiv.org/abs/2602.00189
备注:This paper has been accepted by Elsevier's \textit{Speech Communication} journal. Official publication link: https://doi.org/10.1016/j.specom.2023.103028 The code for the paper is available at the following link: https://github.com/FelixChan9527/LPIPS-AttnWav2Lip
摘要:研究人员对音频驱动的Talking Head Generation越来越感兴趣。说话头部生成的主要挑战是实现嘴唇和音频之间的视听一致性,称为嘴唇同步。本文提出了一种通用的方法,LPIPS-AttnWav 2Lip,用于重建任何说话人的人脸图像的基础上的音频。我们使用基于残差CBAM的U-Net架构来更好地编码和融合音频和视觉模态信息。此外,语义对齐模块扩展生成器网络的感受野,有效获取视觉特征的空间和通道信息;将视觉特征的统计信息与音频特征向量进行匹配,实现音频内容信息对视觉信息的调整和注入。为了实现精确的唇同步并生成逼真的高质量图像,我们的方法采用LPIPS损失,它模拟人类对图像质量的判断,并减少训练过程中不稳定的可能性。主观和客观评价结果表明,该方法在唇同步精度和视觉质量方面取得了优异的性能。论文的代码可在以下链接中获得:https://github.com/FelixChan9527/LPIPS-AttnWav2Lip
摘要:Researchers have shown a growing interest in Audio-driven Talking Head Generation. The primary challenge in talking head generation is achieving audio-visual coherence between the lips and the audio, known as lip synchronization. This paper proposes a generic method, LPIPS-AttnWav2Lip, for reconstructing face images of any speaker based on audio. We used the U-Net architecture based on residual CBAM to better encode and fuse audio and visual modal information. Additionally, the semantic alignment module extends the receptive field of the generator network to obtain the spatial and channel information of the visual features efficiently; and match statistical information of visual features with audio latent vector to achieve the adjustment and injection of the audio content information to the visual information. To achieve exact lip synchronization and to generate realistic high-quality images, our approach adopts LPIPS Loss, which simulates human judgment of image quality and reduces instability possibility during the training process. The proposed method achieves outstanding performance in terms of lip synchronization accuracy and visual quality as demonstrated by subjective and objective evaluation results. The code for the paper is available at the following link: https://github.com/FelixChan9527/LPIPS-AttnWav2Lip
机器翻译由腾讯交互翻译提供,仅供参考
