微信公众号:arXiv_Daily
cs.SD语音
【1】Phonological Tokenizer: Prosody-Aware Phonetic Token via Multi-Objective Fine-Tuning with Differentiable K-Means
标题:音素令牌器:通过具有可区分K均值的多目标微调来感知韵律的音素令牌
链接:https://arxiv.org/abs/2601.19781
备注:Accepted to ICASSP 2026
摘要:近年来,人们对用离散令牌表示语音越来越感兴趣,离散令牌用作语音语言模型(speechLM)的伪文本和下游任务的有效中间表示。这些标记通常被分类为声学标记和语音标记:前者保存用于重建的详细声学信息,而后者主要捕获语言内容。然而,在人类语音通信中,不必要的声学细节,如说话人信息被抽象,而语言和韵律信息用于语音理解和生产。考虑到这一点,两种类型的令牌似乎都不是对韵律敏感的任务的理想表示,例如speechLM。在这项研究中,我们提出了语音标记器,一种通过可微k均值对语音标记进行微调的方法,该方法具有ASR和语音再合成的多任务目标。不同任务的实验验证证实,我们的令牌保留语音(语言和韵律)的信息,同时适当地丢弃扬声器的身份。
摘要:In recent years, there has been growing interest in representing speech with discrete tokens, which serve as pseudo-text for speech language models (speechLMs) and as efficient intermediate representations for downstream tasks. These tokens are typically categorized as acoustic and phonetic tokens: the former holds detailed acoustic information for reconstruction while the latter mainly captures linguistic content. In human speech communication, however, unnecessary acoustic details such as speaker information are abstracted, while both linguistic and prosodic information are utilized for speech comprehension and production. Given this, neither type of token seems an ideal representation for tasks sensitive to prosody, such as speechLMs. In this study, we propose the Phonological Tokenizer, a method that fine-tunes phonetic tokens via differentiable k-means with a multi-task objective of ASR and speech resynthesis. Experimental validation on diverse tasks confirms that our tokens retain phonological (both linguistic and prosodic) information while appropriately discarding speaker identity.
【2】Advanced Modeling of Interlanguage Speech Intelligibility Benefit with L1-L2 Multi-Task Learning Using Differentiable K-Means for Accent-Robust Discrete Token-Based ASR
标题:使用可区分K均值的L1-L2多任务学习对中介语语音可理解性的高级建模,用于口音稳健的基于离散令牌的ASB
链接:https://arxiv.org/abs/2601.19767
备注:Accepted to ICASSP 2026
摘要:在当今全球化的世界中,建立对外国口音语音鲁棒的ASR系统是一个重要的挑战。先前的一项研究探讨了通过再现被称为中介语语音可懂度效益(ISIB)的现象来增强基于语音标记的ASR对口音语音的性能的方法,其中外国口音语音对共享说话者的母语的听众比对母语听众更容易理解。ISIB在技术上是通过使用说话人的L1来学习SSL特征空间中的k-means聚类质心以获得语音标记来实现的。在这项研究中,我们提出了一个更先进的建模ISIB。通过采用可微分的k均值并优化L1和L2 ASR的整个模块,所提出的方法在仅使用母语和另外包含有限数量的口音语音时都优于基线。值得注意的是,在后一种情况下,我们的方法在识别准确性方面实现了大约20%的相对提高。
摘要:Building ASR systems robust to foreign-accented speech is an important challenge in today's globalized world. A prior study explored the way to enhance the performance of phonetic token-based ASR on accented speech by reproducing the phenomenon known as interlanguage speech intelligibility benefit (ISIB), where foreign-accented speech is more intelligible to listeners sharing the speaker's native language than to native listeners. ISIB was technically implemented by using the speaker's L1 to learn k-means cluster centroids in an SSL feature space to obtain phonetic tokens. In this study, we propose a more advanced modeling of ISIB. By employing differentiable k-means and optimizing the entire module for both L1 and L2 ASR, the proposed method outperformed the baselines, both when using only native speech and when additionally incorporating a limited amount of accented speech. Notably, in the latter scenario, our method achieved approximately a 20% relative improvement in recognition accuracy.
【3】Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
标题:具有视觉语言先验和3D声学环境建模的物理感知新视图声学合成
链接:https://arxiv.org/abs/2601.19712
备注:ICASSP 2026 Accept, Project page: https://physnvas.github.io/
摘要:空间音频对于沉浸式体验至关重要,但由于反射、衍射和材料吸收等复杂的物理现象,新颖视图声学合成(NVAS)仍然具有挑战性。基于单视图或全景输入的现有方法提高了空间保真度,但未能捕获全局几何和语义线索,如对象布局和材料属性。为了解决这个问题,我们提出了物理NVAS,第一个物理感知NVAS框架,集成了空间几何建模与视觉语言语义先验。从多视图图像和深度图重建全局3D声学环境,以估计房间大小和形状,增强声音传播的空间意识。同时,视觉语言模型提取物体、布局和材料的物理感知先验,捕捉几何结构之外的吸收和反射。声学特征融合适配器将这些线索统一到用于双耳生成的物理感知表示中。RWAVS上的实验表明,物理NVAS产生的双耳音频具有改善的真实感和物理一致性。
摘要:Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction, and material absorption. Existing methods based on single-view or panoramic inputs improve spatial fidelity but fail to capture global geometry and semantic cues such as object layout and material properties. To address this, we propose Phys-NVAS, the first physics-aware NVAS framework that integrates spatial geometry modeling with vision-language semantic priors. A global 3D acoustic environment is reconstructed from multi-view images and depth maps to estimate room size and shape, enhancing spatial awareness of sound propagation. Meanwhile, a vision-language model extracts physics-aware priors of objects, layouts, and materials, capturing absorption and reflection beyond geometry. An acoustic feature fusion adapter unifies these cues into a physics-aware representation for binaural generation. Experiments on RWAVS demonstrate that Phys-NVAS yields binaural audio with improved realism and physical consistency.
【4】Hyperbolic Additive Margin Softmax with Hierarchical Information for Speaker Verification
标题:具有分层信息的双曲加性余量Softmax用于说话人验证
链接:https://arxiv.org/abs/2601.19709
备注:5 pages, 3 figures, Accepted at ICASSP 2026
摘要:基于欧氏空间的说话人嵌入学习已经取得了很大的进展,但在对说话人特征中的层次信息建模方面仍然存在不足。双曲空间具有负曲率的几何特性,可以有效地表示有限体积内的层次信息,更适合于说话人嵌入的特征分布。本文在双曲空间的基础上提出了双曲型Softmax(H-Softmax)和双曲型可加裕度Softmax(HAM-Softmax)。H-Softmax通过将嵌入和扬声器中心投影到双曲空间并计算双曲距离,将分层信息纳入扬声器嵌入。HAM-Softmax在此基础上通过引入边缘约束进一步增强了类间可分性。实验结果表明,与标准Softmax和AM-Softmax相比,H-Softmax和HAM-Softmax的平均相对EER分别降低了27.84%和14.23%,表明该方法在保持层次结构建模能力的同时,有效地提高了说话人确认性能.该代码将在https://github.com/PunkMale/HAM-Softmax上发布。
摘要:Speaker embedding learning based on Euclidean space has achieved significant progress, but it is still insufficient in modeling hierarchical information within speaker features. Hyperbolic space, with its negative curvature geometric properties, can efficiently represent hierarchical information within a finite volume, making it more suitable for the feature distribution of speaker embeddings. In this paper, we propose Hyperbolic Softmax (H-Softmax) and Hyperbolic Additive Margin Softmax (HAM-Softmax) based on hyperbolic space. H-Softmax incorporates hierarchical information into speaker embeddings by projecting embeddings and speaker centers into hyperbolic space and computing hyperbolic distances. HAM-Softmax further enhances inter-class separability by introducing margin constraint on this basis. Experimental results show that H-Softmax and HAM-Softmax achieve average relative EER reductions of 27.84% and 14.23% compared with standard Softmax and AM-Softmax, respectively, demonstrating that the proposed methods effectively improve speaker verification performance and at the same time preserve the capability of hierarchical structure modeling. The code will be released at https://github.com/PunkMale/HAM-Softmax.
【5】A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
标题:多模式大型语言模型音频推理能力的基准
链接:https://arxiv.org/abs/2601.19673
备注:31 pages, 2 figures, accepted to EACL 2026
摘要:目前用于测试多模态大语言模型的音频模态的基准集中于测试各种音频任务,例如孤立地进行说话人日记化或性别识别。多模态模型是否可以回答需要推理技能来组合不同类别的音频任务的问题,无法通过其使用来验证。为了解决这个问题,我们提出了音频推理任务(ART),一个新的基准评估多模态模型的能力,以解决问题,需要推理的音频信号。
摘要:The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identification in isolation. Whether a multimodal model can answer the questions that require reasoning skills to combine audio tasks of different categories, cannot be verified with their use. To address this issue, we propose Audio Reasoning Tasks (ART), a new benchmark for assessing the ability of multimodal models to solve problems that require reasoning over audio signal.
【6】GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
标题:GMS-CAVP:通过多尺度对比和生成性预训练改善音视频通信
链接:https://arxiv.org/abs/2601.19606
摘要:最近的进展,在视频-音频(V-A)的理解和生成越来越多地依赖于联合V-A嵌入,作为跨模态检索和生成等任务的基础。虽然像CAVP这样的现有方法使用对比目标有效地对模态之间的语义和时间对应进行建模,但它们的性能仍然不理想。一个关键的限制是对视频和音频信号的密集、多尺度性质的建模不足,对应关系通常跨越细粒度到粗粒度的时空结构,这在现有框架中未得到充分利用。为此,我们提出了GMS-CAVP,这是一种新的框架,它结合了多尺度视频-音频对齐和基于多尺度时空扩散的预训练目标,以增强V-A对应建模。首先,GMS-CAVP引入了一种多尺度对比学习策略,可以捕获不同粒度的语义和时间关系。其次,我们超越了传统的对比学习,结合了基于扩散的生成目标,实现了视频和音频之间的模态翻译和合成。这种统一的判别生成公式促进了更深入的跨模态理解,并为高保真生成铺平了道路。在VGGSound、AudioSet和Panda 70 M上的大量实验表明,GMS-CAVP在生成和检索方面优于以前的方法。
摘要:Recent advances in video-audio (V-A) understanding and generation have increasingly relied on joint V-A embeddings, which serve as the foundation for tasks such as cross-modal retrieval and generation. While prior methods like CAVP effectively model semantic and temporal correspondences between modalities using contrastive objectives, their performance remains suboptimal. A key limitation is the insufficient modeling of the dense, multi-scale nature of both video and audio signals, correspondences often span fine- to coarse-grained spatial-temporal structures, which are underutilized in existing frameworks. To this end, we propose GMS-CAVP, a novel framework that combines Multi-Scale Video-Audio Alignment and Multi-Scale Spatial-Temporal Diffusion-based pretraining objectives to enhance V-A correspondence modeling. First, GMS-CAVP introduces a multi-scale contrastive learning strategy that captures semantic and temporal relations across varying granularities. Second, we go beyond traditional contrastive learning by incorporating a diffusion-based generative objective, enabling modality translation and synthesis between video and audio. This unified discriminative-generative formulation facilitates deeper cross-modal understanding and paves the way for high-fidelity generation. Extensive experiments on VGGSound, AudioSet, and Panda70M demonstrate that GMS-CAVP outperforms previous methods in generation and retrieval.
【7】SLM-SS: Speech Language Model for Generative Speech Separation
标题:SL M-SS:生成式语音分离的语音语言模型
链接:https://arxiv.org/abs/2601.19533
摘要:语音分离(SS)的进步显着与神经网络为基础的方法,显示出改善的信号电平指标的性能。然而,这些方法通常难以保持分离信号中的语音清晰度,这可能会对下游任务(如语音识别)的性能产生负面影响。在这项工作中,我们提出了SLM-SS,一种新的方法,适用于语音语言模型SS,旨在提高分离信号的可懂度和连贯性。我们将SS定义为离散的多码本序列生成,使用编码器-解码器模型将量化的语音混合映射到目标令牌。除了自回归建模策略之外,我们还引入了一个非自回归模型来提高剩余令牌的解码效率。LibriMix数据集上的实验结果表明,与现有方法相比,我们的方法显着更好地保留了语音清晰度,从而提高了各种下游任务的语言一致性。
摘要:Speech separation (SS) has advanced significantly with neural network-based methods, showing improved performance on signal-level metrics. However, these methods often struggle to maintain speech intelligibility in the separated signals, which can negatively affect the performance of downstream tasks such as speech recognition. In this work, we propose SLM-SS, a novel approach that applies speech language models to SS, aiming to enhance the intelligibility and coherence of the separated signals. We frame SS as discrete multi-codebook sequence generation, using Encoder-Decoder models to map quantized speech mixtures to target tokens. In addition to the autoregressive modeling strategy, we introduce a non-autoregressive model to improve decoding efficiency for residual tokens. Experimental results on the LibriMix dataset demonstrate that our approach shows significantly better preservation of speech intelligibility, leading to improved linguistic consistency in a variety of downstream tasks compared to existing approaches.
【8】Dual-Strategy-Enhanced ConBiMamba for Neural Speaker Diarization
标题:用于神经扬声器扩展的双策略增强ConBiMamba
链接:https://arxiv.org/abs/2601.19472
备注:Accepted at ICASSP 2026
摘要:Conformer和Mamba在语音建模方面取得了很好的效果,但在说话人日志化方面存在局限性。Mamba是高效的,但与局部细节和非线性模式的斗争。Conformer的自我注意会导致长语音序列的高内存开销,并可能导致长距离依赖建模的不稳定性。这些限制对于日志化是至关重要的,日志化需要对局部变化进行精确建模,并在扩展的跨度上保持稳健的说话人一致性。为了解决这些挑战,我们首先将ConBiMamba应用于说话人日记。我们遵循Pyannote管道,提出了双策略增强的ConBiMamba神经说话人日记系统。ConBiMamba集成了Conformer和Mamba的优势,其中Conformer的卷积和前馈结构用于改进局部特征提取。通过用ExtBiMamba替换Conformer的自我注意力,ConBiMamba有效地处理长音频序列,同时减轻自我注意力的高内存成本。此外,为了解决说话人变化点附近的较高DER的问题,我们引入边界增强的转换损失来增强说话人变化点的检测。我们还提出了逐层特征聚合,以提高利用多层表示。该系统在六个日志数据集上进行了评估,并在其中四个数据集上实现了最先进的性能。我们研究的源代码可在https://github.com/lz-hust/DSE-CBM上获得。
摘要:Conformer and Mamba have achieved strong performance in speech modeling but face limitations in speaker diarization. Mamba is efficient but struggles with local details and nonlinear patterns. Conformer's self-attention incurs high memory overhead for long speech sequences and may cause instability in long-range dependency modeling. These limitations are critical for diarization, which requires both precise modeling of local variations and robust speaker consistency over extended spans. To address these challenges, we first apply ConBiMamba for speaker diarization. We follow the Pyannote pipeline and propose the Dual-Strategy-Enhanced ConBiMamba neural speaker diarization system. ConBiMamba integrates the strengths of Conformer and Mamba, where Conformer's convolutional and feed-forward structures are utilized to improve local feature extraction. By replacing Conformer's self-attention with ExtBiMamba, ConBiMamba efficiently handles long audio sequences while alleviating the high memory cost of self-attention. Furthermore, to address the problem of the higher DER around speaker change points, we introduce the Boundary-Enhanced Transition Loss to enhance the detection of speaker change points. We also propose Layer-wise Feature Aggregation to enhance the utilization of multi-layer representations. The system is evaluated on six diarization datasets and achieves state-of-the-art performance on four of them. The source code of our study is available at https://github.com/lz-hust/DSE-CBM.
【9】Residual Tokens Enhance Masked Autoencoders for Speech Modeling
标题:剩余标记增强语音建模中的掩蔽自编码器
链接:https://arxiv.org/abs/2601.19399
备注:Submitted to ICASSP 2026 (accepted)
摘要:最近的语音建模依赖于明确的属性,如音高,内容和扬声器的身份,但这些单独不能捕捉自然语音的全部丰富性。我们介绍了RT-MAE,这是一种新型的掩码自动编码器框架,它用无监督的残余可训练令牌增强了基于监督属性的建模,旨在对没有被显式标记因子解释的信息进行编码(例如,音色变化、噪声、情绪等)。实验表明,RT-MAE提高了重建质量,保持内容和说话人的相似性,同时增强了表现力。我们进一步证明了它的适用性语音增强,在推理中去除噪声,同时保持可控性和自然性。
摘要:Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce RT-MAE, a novel masked autoencoder framework that augments the supervised attributes-based modeling with unsupervised residual trainable tokens, designed to encode the information not explained by explicit labeled factors (e.g., timbre variations, noise, emotion etc). Experiments show that RT-MAE improves reconstruction quality, preserving content and speaker similarity while enhancing expressivity. We further demonstrate its applicability to speech enhancement, removing noise at inference while maintaining controllability and naturalness.
【10】Phase-Retrieval-Based Physics-Informed Neural Networks For Acoustic Magnitude Field Reconstruction
标题:基于相检索的物理信息神经网络用于声学幅度场重建
链接:https://arxiv.org/abs/2601.19297
备注:Accepted to International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2026
摘要:我们提出了一种从空间稀疏幅度测量中估计声场幅度分布的方法。当相位测量不可靠或不可访问时,这种方法是有用的。物理信息神经网络(PINN)通过将从偏微分方程(PDE)导出的约束合并到神经网络中,已经显示出用于声场估计的前景。然而,它们不扩展到相位测量不可用的设置,因为基于支配PDE的损失函数依赖于相位信息。为了弥补这一点,我们提出了一个相位检索为基础的PINN的幅度场估计。通过用单独的网络表示幅度和相位分布,可以基于重构的复振幅来计算PDE损耗。我们证明了我们的相位检索为基础的PINN的有效性,通过实验评估。
摘要:We propose a method for estimating the magnitude distribution of an acoustic field from spatially sparse magnitude measurements. Such a method is useful when phase measurements are unreliable or inaccessible. Physics-informed neural networks (PINNs) have shown promise for sound field estimation by incorporating constraints derived from governing partial differential equations (PDEs) into neural networks. However, they do not extend to settings where phase measurements are unavailable, as the loss function based on the governing PDE relies on phase information. To remedy this, we propose a phase-retrieval-based PINN for magnitude field estimation. By representing the magnitude and phase distributions with separate networks, the PDE loss can be computed based on the reconstructed complex amplitude. We demonstrate the effectiveness of our phase-retrieval-based PINN through experimental evaluation.
【11】A Hybrid Discriminative and Generative System for Universal Speech Enhancement
标题:一种混合判别生成的通用语音增强系统
链接:https://arxiv.org/abs/2601.19113
备注:Accepted by ICASSP 2026.This work was submitted to the ICASSP 2026 URGENT Challenge (Track 1)
摘要:通用语音增强旨在处理具有各种语音失真和记录条件的输入。在这项工作中,我们提出了一种新的混合架构,协同信号保真度的判别建模与生成建模的重建能力。我们的系统采用了判别式TF-GridNet模型与采样频率无关的策略,以处理可变的采样率普遍。同时,自回归模型结合频谱映射建模生成细节丰富的语音,同时有效地抑制生成伪影。最后,融合网络学习的信号电平损失和综合语音质量评估(SQA)损失的优化下的两个输出的自适应权重。我们提出的系统在ICASSP 2026紧急挑战(轨道1)中进行了评估,并排名第三。
摘要:Universal speech enhancement aims at handling inputs with various speech distortions and recording conditions. In this work, we propose a novel hybrid architecture that synergizes the signal fidelity of discriminative modeling with the reconstruction capabilities of generative modeling. Our system utilizes the discriminative TF-GridNet model with the Sampling-Frequency-Independent strategy to handle variable sampling rates universally. In parallel, an autoregressive model combined with spectral mapping modeling generates detail-rich speech while effectively suppressing generative artifacts. Finally, a fusion network learns adaptive weights of the two outputs under the optimization of signal-level losses and the comprehensive Speech Quality Assessment (SQA) loss. Our proposed system is evaluated in the ICASSP 2026 URGENT Challenge (Track 1) and ranks the third place.
【12】Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation
标题:具有不确定性的3D情感说话面部合成与情感优先蒸馏
链接:https://arxiv.org/abs/2601.19112
备注:Accepted by ICASSP 2026
摘要:情感说话人脸合成是多媒体和信号处理中的关键,然而现有的3D方法面临两个关键挑战:音频-视觉情感对齐不良,表现为难以提取音频情感和对情感微表情的控制不足;以及一刀切的多视图融合策略,忽略了不确定性和特征质量差异,破坏了渲染质量。我们提出了UA-3DTalk,基于情感先验提取的不确定性感知的3D情感说话人脸合成算法,该算法有三个核心模块:先验提取模块将音频分解为内容同步的特征,用于对齐;情感提取模块引入了多模态注意力加权融合机制和具有多分辨率码本的4D高斯编码,实现细粒度音频情感提取和情感微表情的精确控制;基于不确定性的变形部署不确定性块来估计视图特定的任意(输入噪声)和认知(模型参数)不确定性,实现自适应多视图融合,并结合用于高斯基元优化的多头解码器以减轻均匀权重融合的限制。在常规和情感数据集上进行的大量实验表明,UA-3DTalk在E-FID的情感对齐方面优于DEGSTalk和EDTalk等最先进的方法5.2%,在SyncC的嘴唇同步方面优于3.1%,在LPIPS的渲染质量方面优于0.015。项目页面:https://mrask999.github.io/UA-3DTalk
摘要:Emotional Talking Face synthesis is pivotal in multimedia and signal processing, yet existing 3D methods suffer from two critical challenges: poor audio-vision emotion alignment, manifested as difficult audio emotion extraction and inadequate control over emotional micro-expressions; and a one-size-fits-all multi-view fusion strategy that overlooks uncertainty and feature quality differences, undermining rendering quality. We propose UA-3DTalk, Uncertainty-Aware 3D Emotional Talking Face Synthesis with emotion prior distillation, which has three core modules: the Prior Extraction module disentangles audio into content-synchronized features for alignment and person-specific complementary features for individualization; the Emotion Distillation module introduces a multi-modal attention-weighted fusion mechanism and 4D Gaussian encoding with multi-resolution code-books, enabling fine-grained audio emotion extraction and precise control of emotional micro-expressions; the Uncertainty-based Deformation deploys uncertainty blocks to estimate view-specific aleatoric (input noise) and epistemic (model parameters) uncertainty, realizing adaptive multi-view fusion and incorporating a multi-head decoder for Gaussian primitive optimization to mitigate the limitations of uniform-weight fusion. Extensive experiments on regular and emotional datasets show UA-3DTalk outperforms state-of-the-art methods like DEGSTalk and EDTalk by 5.2% in E-FID for emotion alignment, 3.1% in SyncC for lip synchronization, and 0.015 in LPIPS for rendering quality. Project page: https://mrask999.github.io/UA-3DTalk
【13】Interpretable and Perceptually-Aligned Music Similarity with Pretrained Embeddings
标题:具有预训练嵌入的可解释和感知一致的音乐相似性
链接:https://arxiv.org/abs/2601.19109
摘要:感知相似性表示使音乐检索系统能够确定哪些歌曲听起来与听众最相似。基于通过自监督度量学习进行特定任务训练的最新方法显示出与人类判断的良好一致性,但由于数据集可用性有限,很难解释或概括。我们表明,预训练的文本音频嵌入(CLAP和MuQ-MuLan)在相似性任务上提供了相当的感知对齐,而无需任何额外的微调。为了超越这一基线,我们引入了一种新的方法,将预训练的嵌入与来自听力测试的ABX偏好数据的源分离和线性优化进行感知对齐。我们的模型提供了可解释和可控的仪器明智的权重,允许音乐制作人检索干级循环和样本的基础上混合参考歌曲。
摘要:Perceptual similarity representations enable music retrieval systems to determine which songs sound most similar to listeners. State-of-the-art approaches based on task-specific training via self-supervised metric learning show promising alignment with human judgment, but are difficult to interpret or generalize due to limited dataset availability. We show that pretrained text-audio embeddings (CLAP and MuQ-MuLan) offer comparable perceptual alignment on similarity tasks without any additional fine-tuning. To surpass this baseline, we introduce a novel method to perceptually align pretrained embeddings with source separation and linear optimization on ABX preference data from listening tests. Our model provides interpretable and controllable instrument-wise weights, allowing music producers to retrieve stem-level loops and samples based on mixed reference songs.
【14】Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
标题:利用人工智能反馈的强化学习优化口语对话系统中的对话质量
链接:https://arxiv.org/abs/2601.19063
摘要:针对语音输入/语音输出对话系统(SDS)的人类或AI反馈(RLHF/RLAIF)的强化学习仍然没有得到充分研究,先前的工作主要限于在话语层面应用的单一语义奖励。这样的设置忽略了会话质量的多维和多模态性质,其包括语义连贯性、音频自然性、说话者一致性、情感一致性和话轮转换行为。此外,他们是从根本上不匹配的双工口语对话系统,产生的反应逐步增加,代理人必须根据部分话语作出决定。我们解决了这些限制与SDS的第一个多奖励RLAIF框架,结合语义,音频质量和情感一致性奖励。为了将话语级别偏好与双重模型中的增量、分块解码相匹配,我们应用回合级别偏好采样并在单个DPO目标内聚合每个块的对数概率。我们提出了第一个系统的研究偏好学习,以提高SDS质量的多轮思想链和块双工模型,并发布了多奖励DPO数据集,以支持可重复的研究。实验表明,单奖励RLAIF选择性地提高其目标度量,而联合多奖励训练在语义质量和音频自然度方面产生一致的增益。这些结果突出了整体的重要性,多奖励调整实用的会话SDS。
摘要:Reinforcement learning from human or AI feedback (RLHF/RLAIF) for speech-in/speech-out dialogue systems (SDS) remains underexplored, with prior work largely limited to single semantic rewards applied at the utterance level. Such setups overlook the multi-dimensional and multi-modal nature of conversational quality, which encompasses semantic coherence, audio naturalness, speaker consistency, emotion alignment, and turn-taking behavior. Moreover, they are fundamentally mismatched with duplex spoken dialogue systems that generate responses incrementally, where agents must make decisions based on partial utterances. We address these limitations with the first multi-reward RLAIF framework for SDS, combining semantic, audio-quality, and emotion-consistency rewards. To align utterance-level preferences with incremental, blockwise decoding in duplex models, we apply turn-level preference sampling and aggregate per-block log-probabilities within a single DPO objective. We present the first systematic study of preference learning for improving SDS quality in both multi-turn Chain-of-Thought and blockwise duplex models, and release a multi-reward DPO dataset to support reproducible research. Experiments show that single-reward RLAIF selectively improves its targeted metric, while joint multi-reward training yields consistent gains across semantic quality and audio naturalness. These results highlight the importance of holistic, multi-reward alignment for practical conversational SDS.
【15】Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
标题:音频基础模型在钢琴演奏评估中优于符号表示
链接:https://arxiv.org/abs/2601.19029
备注:6 pages, 4 figures, 2 tables. Code available at https://github.com/Jai-Dhiman/crescendai
摘要:传统上,自动钢琴演奏评估依赖于符号(symbolic)表示,它捕获音符级别的信息,但错过了表现性演奏的声学细微差别。我建议使用预先训练的音频基础模型,特别是MuQ和MERT,来预测钢琴演奏质量的19个感知维度。使用PercePiano的合成音频文件(通过Pianoteq渲染),我比较音频和符号的方法在受控条件下,两者都来自相同的源数据。最好的模型,MuQ层9-12与Pianoteq soundfont增强,达到R^2 = 0.537(95% CI:[0.465,0.575]),表示比符号基线(R^2 = 0.347)提高了55%。统计分析证实,音频在所有19个维度上的表现都优于符号,具有显著性(p < 10^-25)。我通过跨音字体泛化(R^2 = 0.534 +/- 0.075)、与外部数据集的难度相关性(rho = 0.623)和多执行者一致性分析来验证该方法。音频-符号融合的分析揭示了高误差相关性(r = 0.738),解释了为什么融合提供的益处最小:单独的音频表示就足够了。我发布了完整的训练管道、预训练模型和推理代码。
摘要:Automated piano performance evaluation traditionally relies on symbolic (MIDI) representations, which capture note-level information but miss the acoustic nuances that characterize expressive playing. I propose using pre-trained audio foundation models, specifically MuQ and MERT, to predict 19 perceptual dimensions of piano performance quality. Using synthesized audio from PercePiano MIDI files (rendered via Pianoteq), I compare audio and symbolic approaches under controlled conditions where both derive from identical source data. The best model, MuQ layers 9-12 with Pianoteq soundfont augmentation, achieves R^2 = 0.537 (95% CI: [0.465, 0.575]), representing a 55% improvement over the symbolic baseline (R^2 = 0.347). Statistical analysis confirms significance (p < 10^-25) with audio outperforming symbolic on all 19 dimensions. I validate the approach through cross-soundfont generalization (R^2 = 0.534 +/- 0.075), difficulty correlation with an external dataset (rho = 0.623), and multi-performer consistency analysis. Analysis of audio-symbolic fusion reveals high error correlation (r = 0.738), explaining why fusion provides minimal benefit: audio representations alone are sufficient. I release the complete training pipeline, pretrained models, and inference code.
【16】A Framework for Evaluating Faithfulness in Explainable AI for Machine Anomalous Sound Detection Using Frequency-Band Perturbation
标题:使用频段微扰检测机器异常声音的可解释人工智能可靠性评估框架
链接:https://arxiv.org/abs/2601.19017
备注:16 pages, 24 figures
摘要:可解释AI(XAI)通常应用于异常声音检测(ASD)模型,以识别音频信号的哪些时频区域有助于异常决策。然而,大多数音频解释依赖于显着图的定性检查,留下了这些属性是否准确反映模型使用的频谱线索的问题。在这项工作中,我们引入了一个新的定量框架,通过系统的频带去除直接将归因相关性与模型行为联系起来,来评估XAI在机器声音分析中的忠诚度。这种方法提供了一个客观的衡量标准,用于机器ASD的XAI方法是否正确识别影响ASD模型预测的频率区域。通过使用四种被广泛采用的方法,即集成反射,遮挡,Grad-CAM和SmoothGrad,我们证明了XAI技术在可靠性上的不同,遮挡显示出与真实模型灵敏度最强的对齐,基于梯度+的方法通常无法准确捕获光谱依赖性。所提出的框架提供了一种可重复的方式来基准音频解释,并使基于频谱的ASD系统的解释更值得信赖。
摘要:Explainable AI (XAI) is commonly applied to anomalous sound detection (ASD) models to identify which time-frequency regions of an audio signal contribute to an anomaly decision. However, most audio explanations rely on qualitative inspection of saliency maps, leaving open the question of whether these attributions accurately reflect the spectral cues the model uses. In this work, we introduce a new quantitative framework for evaluating XAI faithfulness in machine-sound analysis by directly linking attribution relevance to model behaviour through systematic frequency-band removal. This approach provides an objective measure of whether an XAI method for machine ASD correctly identifies frequency regions that influence an ASD model's predictions. By using four widely adopted methods, namely Integrated Gradients, Occlusion, Grad-CAM and SmoothGrad, we show that XAI techniques differ in reliability, with Occlusion demonstrating the strongest alignment with true model sensitivity and gradient-+based methods often failing to accurately capture spectral dependencies. The proposed framework offers a reproducible way to benchmark audio explanations and enables more trustworthy interpretation of spectrogram-based ASD systems.
【17】Enhancing Speech Emotion Recognition using Dynamic Spectral Features and Kalman Smoothing
标题:基于动态谱特征和卡尔曼平滑的语音情感识别
链接:https://arxiv.org/abs/2601.18908
摘要:语音情感识别系统通常使用静态特征,如梅尔频率倒谱系数(MFCC),过零率(ZCR)和均方根能量(RMSE)。正因为如此,当声音信号中存在声学噪声时,他们可能会对情绪进行错误分类。为了解决这个问题,我们使用动态光谱特征(Delta和Delta-Delta)以及卡尔曼平滑算法添加了动态特征。这种方法减少了噪音,提高了情感分类。由于情绪随时间变化,卡尔曼平滑滤波器也有助于使分类器输出更稳定。在RAVDESS数据集上的测试表明,该方法达到了87%的最先进的准确率,并减少了具有相似声学特征的情感之间的错误分类
摘要:Speech Emotion Recognition systems often use static features like Mel-Frequency Cepstral Coefficients (MFCCs), Zero Crossing Rate (ZCR), and Root Mean Square Energy (RMSE). Because of this, they can misclassify emotions when there is acoustic noise in vocal signals. To address this, we added dynamic features using Dynamic Spectral features (Deltas and Delta-Deltas) along with the Kalman Smoothing algorithm. This approach reduces noise and improves emotion classification. Since emotion changes over time, the Kalman Smoothing filter also helped make the classifier outputs more stable. Tests on the RAVDESS dataset showed that this method achieved a state-of-the-art accuracy of 87\% and reduced misclassification between emotions with similar acoustic features
【18】SICL-AT: Another way to adapt Auditory LLM to low-resource task
标题:SICL-AT:将Auditory LLM适应低资源任务的另一种方法
链接:https://arxiv.org/abs/2601.18904
摘要:听觉大语言模型(LLM)在广泛的语音和音频理解任务中表现出强大的性能。然而,当应用于低资源或不熟悉的任务时,它们往往会遇到困难。在标记的域内数据很少或与真实测试分布不匹配的情况下,直接微调可能是脆弱的。在上下文学习(ICL)提供了一个训练免费,推理时间的解决方案,通过调整听觉LLM通过在一些领域内的示范条件。在这项工作中,我们首先表明,\n {香草ICL},提高zero-shot性能在不同的语音和音频任务的选择模型,这表明这种ICL适应能力可以推广到多模态设置。在此基础上,我们提出了\textbf{Speech In-Context Learning Adaptation Training(SICL-AT)},这是一种仅利用高资源语音数据的后训练方法,旨在增强模型的上下文学习能力。该增强可以推广到音频理解/推理任务。实验表明,我们提出的方法始终优于直接微调在低资源的情况下。
摘要:Auditory Large Language Models (LLMs) have demonstrated strong performance across a wide range of speech and audio understanding tasks. Nevertheless, they often struggle when applied to low-resource or unfamiliar tasks. In case of labeled in-domain data is scarce or mismatched to the true test distribution, direct fine-tuning can be brittle. In-Context Learning (ICL) provides a training-free, inference-time solution by adapting auditory LLMs through conditioning on a few in-domain demonstrations. In this work, we first show that \emph{Vanilla ICL}, improves zero-shot performance across diverse speech and audio tasks for selected models which suggest this ICL adaptation capability can be generalized to multimodal setting. Building on this, we propose \textbf{Speech In-Context Learning Adaptation Training (SICL-AT)}, a post-training recipe utilizes only high resource speech data intending to strengthen model's in-context learning capability. The enhancement can generalize to audio understanding/reasoning task. Experiments indicate our proposed method consistently outperforms direct fine-tuning in low-resource scenario.
【19】Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries
标题:语言家族很重要:跨语言边界评估基于法学硕士的ASB
链接:https://arxiv.org/abs/2601.18899
摘要:大型语言模型(LLM)驱动的自动语音识别(ASR)系统通过将冻结的语音编码器通过轻量级连接器链接到预训练的LLM,从而在有限的资源下实现强大的性能。先前的工作训练每种语言的单独连接器,忽略了语言相关性。我们提出了一个有效的和新颖的连接器共享策略的基础上的语言家庭成员,使每个家庭一个连接器,并在两个多语言LLM和两个现实世界的语料库跨越策划和众包语音经验验证其有效性。我们的研究结果表明,基于家庭的连接器减少了参数计数,同时提高跨域的泛化,为多语言ASR部署提供了一个实用的和可扩展的策略。
摘要:Large Language Model (LLM)-powered Automatic Speech Recognition (ASR) systems achieve strong performance with limited resources by linking a frozen speech encoder to a pretrained LLM via a lightweight connector. Prior work trains a separate connector per language, overlooking linguistic relatedness. We propose an efficient and novel connector-sharing strategy based on linguistic family membership, enabling one connector per family, and empirically validate its effectiveness across two multilingual LLMs and two real-world corpora spanning curated and crowd-sourced speech. Our results show that family-based connectors reduce parameter count while improving generalization across domains, offering a practical and scalable strategy for multilingual ASR deployment.
【20】Echoes of the Land: An Interactive Installation Based on Physical Model of Earthquake
标题:土地的回声:基于地震物理模型的互动装置
链接:https://arxiv.org/abs/2507.14947
备注:7 pages, 8 figures, submitted to Leonardo
摘要:《大地的回声》是一个互动装置,通过一个有科学依据的弹簧块模型,将地震动力学转化为多感官体验。模拟地震复发和自组织临界性,工作产生实时的声音和光,通过运动捕捉和串联颗粒合成。每个区块都充当一个代理,产生紧急视听级联,可视化破裂和阈值行为的物理学。这项工作体现了科学知识和艺术实践的融合,为乐器和叙事媒介的新形式开辟了新的途径,同时邀请进一步研究新兴的复杂性,美学和互动性的交叉点。
摘要:Echoes of the Land is an interactive installation that transforms seismic dynamics into a multisensory experience through a scientifically grounded spring-block model. Simulating earthquake recurrence and self-organized criticality, the work generates real-time sound and light via motion capture and concatenative granular synthesis. Each block acts as an agent, producing emergent audiovisual cascades that visualize the physics of rupture and threshold behavior. This work exemplifies the amalgamation of scientific knowledge and artistic practice, opening new avenues for novel forms of musical instrument and narrative medium, while inviting further investigation into the intersection of emergent complexity, aesthetics and interactivity.
【21】Rethinking Discrete Speech Representation Tokens for Accent Generation
标题:重新思考用于口音生成的离散语音表示令牌
链接:https://arxiv.org/abs/2601.19786
摘要:离散语音表示令牌(DSRT)已成为语音生成的基础组件。虽然先前的工作已经广泛地研究了语音和说话者信息在DSRT,如何口音信息编码在DSRT仍然在很大程度上未被探索。在本文中,我们提出了第一个系统的调查口音信息的DSRT。我们提出了一个统一的评估框架,通过一个新的口音ABX任务和可恢复性通过跨口音语音转换(VC)再合成的口音信息的可访问性。使用这个框架,我们分析DSRT来自各种语音编码器。我们的研究结果表明,口音信息大大减少时,ASR监督是用来微调编码器,但不能有效地从语音和扬声器信息通过天真的码本大小减少。基于这些发现,我们提出了新的内容和内容口音DSRT显着优于现有的设计在可控的口音生成。我们的工作突出了口音感知评估的重要性,并为口音控制语音生成设计DSRT提供了实际指导。
摘要:Discrete Speech Representation Tokens (DSRTs) have become a foundational component in speech generation. While prior work has extensively studied phonetic and speaker information in DSRTs, how accent information is encoded in DSRTs remains largely unexplored. In this paper, we present the first systematic investigation of accent information in DSRTs. We propose a unified evaluation framework that measures both accessibility of accent information via a novel Accent ABX task and recoverability via cross-accent Voice Conversion (VC) resynthesis. Using this framework, we analyse DSRTs derived from a variety of speech encoders. Our results reveal that accent information is substantially reduced when ASR supervision is used to fine-tune the encoder, but cannot be effectively disentangled from phonetic and speaker information through naive codebook size reduction. Based on these findings, we propose new content-only and content-accent DSRTs that significantly outperform existing designs in controllable accent generation. Our work highlights the importance of accent-aware evaluation and provides practical guidance for designing DSRTs for accent-controlled speech generation.
【1】Rethinking Discrete Speech Representation Tokens for Accent Generation
标题:重新思考用于口音生成的离散语音表示令牌
链接:https://arxiv.org/abs/2601.19786
摘要:离散语音表示令牌(DSRT)已成为语音生成的基础组件。虽然先前的工作已经广泛地研究了语音和说话者信息在DSRT,如何口音信息编码在DSRT仍然在很大程度上未被探索。在本文中,我们提出了第一个系统的调查口音信息的DSRT。我们提出了一个统一的评估框架,通过一个新的口音ABX任务和可恢复性通过跨口音语音转换(VC)再合成的口音信息的可访问性。使用这个框架,我们分析了来自各种语音编码器的DSRT。我们的研究结果表明,口音信息大大减少时,ASR监督是用来微调编码器,但不能有效地从语音和扬声器信息通过天真的码本大小减少。基于这些发现,我们提出了新的内容和内容口音DSRT显着优于现有的设计在可控的口音生成。我们的工作突出了口音感知评估的重要性,并为口音控制语音生成设计DSRT提供了实际指导。
摘要:Discrete Speech Representation Tokens (DSRTs) have become a foundational component in speech generation. While prior work has extensively studied phonetic and speaker information in DSRTs, how accent information is encoded in DSRTs remains largely unexplored. In this paper, we present the first systematic investigation of accent information in DSRTs. We propose a unified evaluation framework that measures both accessibility of accent information via a novel Accent ABX task and recoverability via cross-accent Voice Conversion (VC) resynthesis. Using this framework, we analyse DSRTs derived from a variety of speech encoders. Our results reveal that accent information is substantially reduced when ASR supervision is used to fine-tune the encoder, but cannot be effectively disentangled from phonetic and speaker information through naive codebook size reduction. Based on these findings, we propose new content-only and content-accent DSRTs that significantly outperform existing designs in controllable accent generation. Our work highlights the importance of accent-aware evaluation and provides practical guidance for designing DSRTs for accent-controlled speech generation.
【2】SAM Audio Judge: A Unified Multimodal Framework for Perceptual Evaluation of Audio Separation
标题:Sam音频法官:用于音频分离感知评估的统一多模式框架
链接:https://arxiv.org/abs/2601.19702
摘要:性能评估仍然是音频分离中的一个复杂挑战,现有的评估指标通常与人类感知不一致,粗粒度,依赖于地面真实信号。另一方面,主观听力测试仍然是真实世界评估的黄金标准,但它们昂贵,耗时且难以扩展。本文讨论了日益增长的需要自动化系统能够评估音频分离,而无需人为干预。建议的评价指标,SAM音频法官(SAJ),是一个多模态细粒度的参考自由的客观度量,它显示出高度一致的人类感知。SAJ支持三个音频域(语音,音乐和一般声音事件)和三个提示输入(文本,视觉和跨度),涵盖四个不同的评估维度(回忆,精确,忠实和整体)。SAM Audio Judge还在数据过滤、伪标记大型数据集和音频分离模型中的重新排序方面显示了潜在的应用。我们在https://github.com/facebookresearch/sam-audio上发布我们的代码和预训练模型。
摘要:The performance evaluation remains a complex challenge in audio separation, and existing evaluation metrics are often misaligned with human perception, course-grained, relying on ground truth signals. On the other hand, subjective listening tests remain the gold standard for real-world evaluation, but they are expensive, time-consuming, and difficult to scale. This paper addresses the growing need for automated systems capable of evaluating audio separation without human intervention. The proposed evaluation metric, SAM Audio Judge (SAJ), is a multimodal fine-grained reference-free objective metric, which shows highly alignment with human perceptions. SAJ supports three audio domains (speech, music and general sound events) and three prompt inputs (text, visual and span), covering four different dimensions of evaluation (recall, percision, faithfulness, and overall). SAM Audio Judge also shows potential applications in data filtering, pseudo-labeling large datasets and reranking in audio separation models. We release our code and pre-trained models at: https://github.com/facebookresearch/sam-audio.
【3】Audio Deepfake Detection at the First Greeting: "Hi!"
链接:https://arxiv.org/abs/2601.19573
备注:Accepted at ICASSP 2026. Copyright 2026 IEEE. The final published version will be available via IEEE Xplore
摘要:本文重点关注现实世界通信降级下的音频深度伪造检测,重点是超短输入(0.5-2.0s),目标是在对话开始时检测合成语音的能力,例如,当一个骗子说“嗨。“我们提出了Short-MGAA(S-MGAA),这是多粒度自适应时频注意力的一种新型轻量级扩展,旨在增强对受到通信处理和扰动的短、降级输入的区分性表示学习。S-MGAA集成了两个量身定制的模块:像素通道增强模块(PCEM),用于放大细粒度的时间-频率显着性;频率补偿增强模块(FCEM),用于通过多尺度频率建模和自适应频率-时间交互来补充有限的时间证据。广泛的实验表明,S-MGAA始终超过9个最先进的基线,同时实现了对降级的强大鲁棒性和良好的效率-准确性权衡,包括低RTF,有竞争力的GFLOPs,紧凑的参数和降低的培训成本,突出了其在通信系统和边缘设备中实时部署的强大潜力。
摘要:This paper focuses on audio deepfake detection under real-world communication degradations, with an emphasis on ultra-short inputs (0.5-2.0s), targeting the capability to detect synthetic speech at a conversation opening, e.g., when a scammer says "Hi." We propose Short-MGAA (S-MGAA), a novel lightweight extension of Multi-Granularity Adaptive Time-Frequency Attention, designed to enhance discriminative representation learning for short, degraded inputs subjected to communication processing and perturbations. The S-MGAA integrates two tailored modules: a Pixel-Channel Enhanced Module (PCEM) that amplifies fine-grained time-frequency saliency, and a Frequency Compensation Enhanced Module (FCEM) to supplement limited temporal evidence via multi-scale frequency modeling and adaptive frequency-temporal interaction. Extensive experiments demonstrate that S-MGAA consistently surpasses nine state-of-the-art baselines while achieving strong robustness to degradations and favorable efficiency-accuracy trade-offs, including low RTF, competitive GFLOPs, compact parameters, and reduced training cost, highlighting its strong potential for real-time deployment in communication systems and edge devices.
【4】Permutation-Invariant Physics-Informed Neural Network for Region-to-Region Sound Field Reconstruction
标题:用于区域到区域声场重建的排列不变物理信息神经网络
链接:https://arxiv.org/abs/2601.19491
备注:Accepted to the 31st International Congress on Sound and Vibration (ICSV 2025)
摘要:大多数现有的声场重建方法的目标点到区域的重建,插值的声学传递函数(ATF)之间的固定位置的声源和接收器区域。这些方法的适用性是有限的,因为现实世界的ATF往往相对于声源和接收器区域的位置连续变化。本文提出了一种用于区域到区域声场重建的置换不变物理信息神经网络,其目的是在连续变化的声源和测量区域内插入ATF。所提出的方法采用深集架构来处理接收器和声源的位置作为一个无序集,保持声学互易性。此外,它将亥姆霍兹方程作为物理约束来指导网络训练,确保物理上一致的预测。
摘要:Most existing sound field reconstruction methods target point-to-region reconstruction, interpolating the Acoustic Transfer Functions (ATFs) between a fixed-position sound source and a receiver region. The applicability of these methods is limited because real-world ATFs tend to varying continuously with respect to the positions of sound sources and receiver regions. This paper presents a permutation-invariant physics-informed neural network for region-to-region sound field reconstruction, which aims to interpolate the ATFs across continuously varying sound sources and measurement regions. The proposed method employs a deep set architecture to process the receiver and sound source positions as an unordered set, preserving acoustic reciprocity. Furthermore, it incorporates the Helmholtz equation as a physical constraint to guide network training, ensuring physically consistent predictions.
【5】SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
标题:SE-DiCoW:自注册的日记调节耳语
链接:https://arxiv.org/abs/2601.19194
备注:Accepted to ICASSP 2026
摘要:多说话人环境下的说话人属性自动语音识别(ASR)仍然是一个重大挑战。虽然一些方法在特定领域进行微调时可以实现强大的性能,但很少有系统能够在域外数据集上进行良好的泛化。我们以前的工作,Diarization-Conditioned Whisper(DiCoW),利用扬声器日记输出作为条件信息,并以最小的微调,表现出强大的多语言和多域性能。在本文中,我们解决了一个关键的限制DiCoW:模糊的沉默目标非目标重叠(STNO)面具,其中两个或两个以上的完全重叠的扬声器可能有几乎相同的条件,尽管不同的transparency。我们引入SE-DiCoW(自注册的Diarization-Conditioned Whisper),它使用日记输出来定位目标说话者最活跃的对话中的任何地方的注册段。该登记段经由每个编码器层处的交叉注意用作固定条件。我们通过改进的数据分割、模型初始化和增强来进一步完善DiCoW。总之,这些进步产生了巨大的收益:SE-DiCoW在EMMA MT-ASR基准上相对于原始DiCoW将宏观平均tcpWER降低了52.4%。
摘要:Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a major challenge. While some approaches achieve strong performance when fine-tuned on specific domains, few systems generalize well across out-of-domain datasets. Our prior work, Diarization-Conditioned Whisper (DiCoW), leverages speaker diarization outputs as conditioning information and, with minimal fine-tuning, demonstrated strong multilingual and multi-domain performance. In this paper, we address a key limitation of DiCoW: ambiguity in Silence-Target-Non-target-Overlap (STNO) masks, where two or more fully overlapping speakers may have nearly identical conditioning despite differing transcriptions. We introduce SE-DiCoW (Self-Enrolled Diarization-Conditioned Whisper), which uses diarization output to locate an enrollment segment anywhere in the conversation where the target speaker is most active. This enrollment segment is used as fixed conditioning via cross-attention at each encoder layer. We further refine DiCoW with improved data segmentation, model initialization, and augmentation. Together, these advances yield substantial gains: SE-DiCoW reduces macro-averaged tcpWER by 52.4% relative to the original DiCoW on the EMMA MT-ASR benchmark.
【6】LuSeeL: Language-queried Binaural Universal Sound Event Extraction and Localization
标题:LuSeeL:百分比查询的双耳通用声音事件提取和定位
链接:https://arxiv.org/abs/2601.19153
备注:ICASSP 2026
摘要:大多数通用的声音提取算法集中于从单声道音频混合中分离目标声音事件。然而,现实世界是三维的,模仿人类听觉的双耳音频可以捕获更丰富的空间信息,包括声源位置。这种空间背景对于理解和建模复杂的听觉场景至关重要,因为它固有地通知声音检测和提取。在这项工作中,我们提出了一种语言驱动的通用声音提取网络,通过有效地利用双耳信号中存在的空间线索,将文本描述的声音事件从双耳混合物中分离出来。此外,我们共同预测的目标声音的到达方向(DoA)使用提取网络的空间特征。这种双任务方法利用互补的位置信息来提高提取性能,同时实现准确的DoA估计。在野外AudioCaps数据集上的实验结果表明,我们提出的LuSeeL模型显着优于单通道和单任务基线。
摘要:Most universal sound extraction algorithms focus on isolating a target sound event from single-channel audio mixtures. However, the real world is three-dimensional, and binaural audio, which mimics human hearing, can capture richer spatial information, including sound source location. This spatial context is crucial for understanding and modeling complex auditory scenes, as it inherently informs sound detection and extraction. In this work, we propose a language-driven universal sound extraction network that isolates text-described sound events from binaural mixtures by effectively leveraging the spatial cues present in binaural signals. Additionally, we jointly predict the direction of arrival (DoA) of the target sound using spatial features from the extraction network. This dual-task approach exploits complementary location information to improve extraction performance while enabling accurate DoA estimation. Experimental results on the in-the-wild AudioCaps dataset show that our proposed LuSeeL model significantly outperforms single-channel and uni-task baselines.
【7】Beyond Lips: Integrating Gesture and Lip Cues for Robust Audio-visual Speaker Extraction
标题:超越嘴唇:集成手势和嘴唇线索,实现稳健的视听说话者提取
链接:https://arxiv.org/abs/2601.19130
备注:ICASSP 2026
摘要:大多数视听说话人提取方法依赖于同步唇记录来从多说话人混合物中分离出目标说话人的语音。然而,在自然的人类交流中,共同语音手势也与语音在时间上对齐,通常强调特定的单词或音节。这些手势提供了互补的视觉提示,当面部或嘴唇区域被遮挡或远离时,这些视觉提示特别有价值。在这项工作中,我们超越了以嘴唇为中心的方法,并提出了SeLG,一个模型,集成了嘴唇和上半身的姿态信息,强大的扬声器提取。SeLG具有基于交叉注意的融合机制,该机制使每个视觉模态能够查询并选择性地关注混合物中的相关语音特征。为了改善手势表示与语音动态的对齐,SeLG还采用了对比性InfoNCE损失,该损失鼓励手势嵌入与相应的嘴唇嵌入更紧密地对齐,这与语音相关性更强。在包含TED演讲的YGD数据集上的实验结果表明,所提出的对比学习策略显着改善了基于手势的说话人提取,并且我们提出的SeLG模型通过有效地融合嘴唇和手势线索与注意力机制和InfoNCE损失,与基线相比,在完整和部分(即,missing-modality)条件。
摘要:Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human communication, co-speech gestures are also temporally aligned with speech, often emphasizing specific words or syllables. These gestures provide complementary visual cues that can be especially valuable when facial or lip regions are occluded or distant. In this work, we move beyond lip-centric approaches and propose SeLG, a model that integrates both lip and upper-body gesture information for robust speaker extraction. SeLG features a cross-attention-based fusion mechanism that enables each visual modality to query and selectively attend to relevant speech features in the mixture. To improve the alignment of gesture representations with speech dynamics, SeLG also employs a contrastive InfoNCE loss that encourages gesture embeddings to align more closely with corresponding lip embeddings, which are more strongly correlated with speech. Experimental results on the YGD dataset, containing TED talks, demonstrate that the proposed contrastive learning strategy significantly improves gesture-based speaker extraction, and that our proposed SeLG model, by effectively fusing lip and gesture cues with an attention mechanism and InfoNCE loss, achieves superior performance compared to baselines, across both complete and partial (i.e., missing-modality) conditions.
【8】GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
标题:GMS-CAVP:通过多尺度对比和生成性预训练改善音视频通信
链接:https://arxiv.org/abs/2601.19606
摘要:最近的进展,在视频-音频(V-A)的理解和生成越来越多地依赖于联合V-A嵌入,作为跨模态检索和生成等任务的基础。虽然像CAVP这样的现有方法使用对比目标有效地对模态之间的语义和时间对应进行建模,但它们的性能仍然不理想。一个关键的限制是对视频和音频信号的密集、多尺度性质的建模不足,对应关系通常跨越细粒度到粗粒度的时空结构,这在现有框架中未得到充分利用。为此,我们提出了GMS-CAVP,这是一种新的框架,它结合了多尺度视频-音频对齐和基于多尺度时空扩散的预训练目标,以增强V-A对应建模。首先,GMS-CAVP引入了一种多尺度对比学习策略,可以捕获不同粒度的语义和时间关系。其次,我们超越了传统的对比学习,结合了基于扩散的生成目标,实现了视频和音频之间的模态翻译和合成。这种统一的判别生成公式促进了更深入的跨模态理解,并为高保真生成铺平了道路。在VGGSound、AudioSet和Panda 70 M上的大量实验表明,GMS-CAVP在生成和检索方面优于以前的方法。
摘要:Recent advances in video-audio (V-A) understanding and generation have increasingly relied on joint V-A embeddings, which serve as the foundation for tasks such as cross-modal retrieval and generation. While prior methods like CAVP effectively model semantic and temporal correspondences between modalities using contrastive objectives, their performance remains suboptimal. A key limitation is the insufficient modeling of the dense, multi-scale nature of both video and audio signals, correspondences often span fine- to coarse-grained spatial-temporal structures, which are underutilized in existing frameworks. To this end, we propose GMS-CAVP, a novel framework that combines Multi-Scale Video-Audio Alignment and Multi-Scale Spatial-Temporal Diffusion-based pretraining objectives to enhance V-A correspondence modeling. First, GMS-CAVP introduces a multi-scale contrastive learning strategy that captures semantic and temporal relations across varying granularities. Second, we go beyond traditional contrastive learning by incorporating a diffusion-based generative objective, enabling modality translation and synthesis between video and audio. This unified discriminative-generative formulation facilitates deeper cross-modal understanding and paves the way for high-fidelity generation. Extensive experiments on VGGSound, AudioSet, and Panda70M demonstrate that GMS-CAVP outperforms previous methods in generation and retrieval.
【9】Phase-Retrieval-Based Physics-Informed Neural Networks For Acoustic Magnitude Field Reconstruction
标题:基于相检索的物理信息神经网络用于声学幅度场重建
链接:https://arxiv.org/abs/2601.19297
备注:Accepted to International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2026
摘要:我们提出了一种从空间稀疏幅度测量中估计声场幅度分布的方法。当相位测量不可靠或不可访问时,这种方法是有用的。物理信息神经网络(PINN)通过将从偏微分方程(PDE)导出的约束合并到神经网络中,已经显示出用于声场估计的前景。然而,它们不扩展到相位测量不可用的设置,因为基于支配PDE的损失函数依赖于相位信息。为了弥补这一点,我们提出了一个相位检索为基础的PINN的幅度场估计。通过用单独的网络表示幅度和相位分布,可以基于重构的复振幅来计算PDE损耗。我们证明了我们的相位检索为基础的PINN的有效性,通过实验评估。
摘要:We propose a method for estimating the magnitude distribution of an acoustic field from spatially sparse magnitude measurements. Such a method is useful when phase measurements are unreliable or inaccessible. Physics-informed neural networks (PINNs) have shown promise for sound field estimation by incorporating constraints derived from governing partial differential equations (PDEs) into neural networks. However, they do not extend to settings where phase measurements are unavailable, as the loss function based on the governing PDE relies on phase information. To remedy this, we propose a phase-retrieval-based PINN for magnitude field estimation. By representing the magnitude and phase distributions with separate networks, the PDE loss can be computed based on the reconstructed complex amplitude. We demonstrate the effectiveness of our phase-retrieval-based PINN through experimental evaluation.
【10】A Hybrid Discriminative and Generative System for Universal Speech Enhancement
标题:一种混合判别生成的通用语音增强系统
链接:https://arxiv.org/abs/2601.19113
备注:Accepted by ICASSP 2026.This work was submitted to the ICASSP 2026 URGENT Challenge (Track 1)
摘要:通用语音增强旨在处理具有各种语音失真和记录条件的输入。在这项工作中,我们提出了一种新的混合架构,协同信号保真度的判别建模与生成建模的重建能力。我们的系统采用了判别式TF-GridNet模型与采样频率无关的策略,以处理可变的采样率普遍。同时,自回归模型结合频谱映射建模生成细节丰富的语音,同时有效地抑制生成伪影。最后,融合网络学习的信号电平损失和综合语音质量评估(SQA)损失的优化下的两个输出的自适应权重。我们提出的系统在ICASSP 2026紧急挑战(轨道1)中进行了评估,并排名第三。
摘要:Universal speech enhancement aims at handling inputs with various speech distortions and recording conditions. In this work, we propose a novel hybrid architecture that synergizes the signal fidelity of discriminative modeling with the reconstruction capabilities of generative modeling. Our system utilizes the discriminative TF-GridNet model with the Sampling-Frequency-Independent strategy to handle variable sampling rates universally. In parallel, an autoregressive model combined with spectral mapping modeling generates detail-rich speech while effectively suppressing generative artifacts. Finally, a fusion network learns adaptive weights of the two outputs under the optimization of signal-level losses and the comprehensive Speech Quality Assessment (SQA) loss. Our proposed system is evaluated in the ICASSP 2026 URGENT Challenge (Track 1) and ranks the third place.
【11】Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
标题:音频基础模型在钢琴演奏评估中优于符号表示
链接:https://arxiv.org/abs/2601.19029
备注:6 pages, 4 figures, 2 tables. Code available at https://github.com/Jai-Dhiman/crescendai
摘要:传统上,自动钢琴演奏评估依赖于符号(symbolic)表示,它捕获音符级别的信息,但错过了表现性演奏的声学细微差别。我建议使用预先训练的音频基础模型,特别是MuQ和MERT,来预测钢琴演奏质量的19个感知维度。使用PercePiano的合成音频文件(通过Pianoteq渲染),我比较音频和符号的方法在受控条件下,两者都来自相同的源数据。最好的模型,MuQ层9-12与Pianoteq soundfont增强,达到R^2 = 0.537(95% CI:[0.465,0.575]),表示比符号基线(R^2 = 0.347)提高了55%。统计分析证实,音频在所有19个维度上的表现都优于符号,具有显著性(p < 10^-25)。我通过跨音字体泛化(R^2 = 0.534 +/- 0.075)、与外部数据集的难度相关性(rho = 0.623)和多执行者一致性分析来验证该方法。音频-符号融合的分析揭示了高误差相关性(r = 0.738),解释了为什么融合提供的益处最小:单独的音频表示就足够了。我发布了完整的训练管道、预训练模型和推理代码。
摘要:Automated piano performance evaluation traditionally relies on symbolic (MIDI) representations, which capture note-level information but miss the acoustic nuances that characterize expressive playing. I propose using pre-trained audio foundation models, specifically MuQ and MERT, to predict 19 perceptual dimensions of piano performance quality. Using synthesized audio from PercePiano MIDI files (rendered via Pianoteq), I compare audio and symbolic approaches under controlled conditions where both derive from identical source data. The best model, MuQ layers 9-12 with Pianoteq soundfont augmentation, achieves R^2 = 0.537 (95% CI: [0.465, 0.575]), representing a 55% improvement over the symbolic baseline (R^2 = 0.347). Statistical analysis confirms significance (p < 10^-25) with audio outperforming symbolic on all 19 dimensions. I validate the approach through cross-soundfont generalization (R^2 = 0.534 +/- 0.075), difficulty correlation with an external dataset (rho = 0.623), and multi-performer consistency analysis. Analysis of audio-symbolic fusion reveals high error correlation (r = 0.738), explaining why fusion provides minimal benefit: audio representations alone are sufficient. I release the complete training pipeline, pretrained models, and inference code.
【12】Enhancing Speech Emotion Recognition using Dynamic Spectral Features and Kalman Smoothing
标题:基于动态谱特征和卡尔曼平滑的语音情感识别
链接:https://arxiv.org/abs/2601.18908
摘要:语音情感识别系统通常使用静态特征,如梅尔频率倒谱系数(MFCC),过零率(ZCR)和均方根能量(RMSE)。正因为如此,当声音信号中存在声学噪声时,他们可能会对情绪进行错误分类。为了解决这个问题,我们使用动态光谱特征(Delta和Delta-Delta)以及卡尔曼平滑算法添加了动态特征。这种方法减少了噪音,提高了情感分类。由于情绪随时间变化,卡尔曼平滑滤波器也有助于使分类器输出更稳定。在RAVDESS数据集上的测试表明,该方法达到了87%的最先进的准确率,并减少了具有相似声学特征的情感之间的错误分类
摘要:Speech Emotion Recognition systems often use static features like Mel-Frequency Cepstral Coefficients (MFCCs), Zero Crossing Rate (ZCR), and Root Mean Square Energy (RMSE). Because of this, they can misclassify emotions when there is acoustic noise in vocal signals. To address this, we added dynamic features using Dynamic Spectral features (Deltas and Delta-Deltas) along with the Kalman Smoothing algorithm. This approach reduces noise and improves emotion classification. Since emotion changes over time, the Kalman Smoothing filter also helped make the classifier outputs more stable. Tests on the RAVDESS dataset showed that this method achieved a state-of-the-art accuracy of 87\% and reduced misclassification between emotions with similar acoustic features
机器翻译由腾讯交互翻译提供,仅供参考
