微信公众号:arXiv_Daily
cs.SD语音
【1】SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding
标题:SpotSound:以细粒度的时间基础增强大型音频语言模型
链接:https://arxiv.org/abs/2604.13023
摘要:大型音频语言模型(ALM)最近在整体音频理解方面表现出了卓越的能力,但它们对于时间基础仍然不可靠,即,精确定位长格式音频中事件发生的时间的任务。这种限制源于两个因素:训练数据由剪辑级监督主导,缺乏精确的时间戳,以及基准测试无法模拟真实世界的场景,其中短事件被密集的背景声音所掩盖。在本文中,我们介绍了SpotSound,音频语言模型接地音频事件设计。SpotSound采用了一种新颖的训练目标,专门用于抑制输入中不存在的事件的幻觉时间戳。此外,我们提出了SpotSound-Bench,这是一个具有挑战性的时间基础基准,目标事件占每个片段的不到10%,创建了一个严格的“大海捞针”评估。实验表明,SpotSound在时间接地基准上实现了最先进的结果,同时在一般下游音频语言任务中保持了稳健的性能。代码、模型和基准测试在https://loiesun.github.io/spotsound/上发布
摘要:Large Audio-Language Models (ALMs) have recently demonstrated remarkable capabilities in holistic audio understanding, yet they remain unreliable for temporal grounding, i.e., the task of pinpointing exactly when an event occurs within long-form audio. This limitation stems from two factors: training data dominated by clip-level supervision lacking precise timestamps, and benchmarks that fail to simulate real-world scenarios where short events are obscured by dense background sounds. In this paper, we introduce SpotSound, an audio language model designed for grounding audio events. SpotSound incorporates a novel training objective, specifically designed to suppress hallucinated timestamps for events absent from the input. Additionally, we present SpotSound-Bench, a challenging temporal grounding benchmark where target events occupy less than ~10\% of each clip, creating a rigorous `needle-in-a-haystack' evaluation. Experiments demonstrate that SpotSound achieves state-of-the-art results on temporal grounding benchmarks while maintaining robust performance across general downstream audio-language tasks. Code, models and benchmark are released on https://loiesun.github.io/spotsound/
【2】Transformer Based Machine Fault Detection From Audio Input
标题:基于Transformer的音频输入机器故障检测
链接:https://arxiv.org/abs/2604.12733
摘要:近年来,Sound AI越来越多地用于预测机器故障。通过将麦克风连接到感兴趣的机器上,可以从现场获得机器行为的实时数据。传统上,卷积神经网络(CNN)架构已用于分析从捕获的声音生成的频谱图图像,并预测机器是否按预期运行。CNN架构似乎在经验上工作得很好,即使它们有像局部性和参数共享这样的偏差,这些偏差可能与频谱图分析不完全相关。从2020年的Vision Transformer(ViT)开始,基于transformer的模型在图像处理领域的成功应用,人们对在Sound AI领域利用这些模型产生了浓厚的兴趣。由于基于变压器的架构具有显著较低的电感偏差,因此在足够的数据下,它们在谱图分析方面的表现预计将优于CNN。本文展示了变压器驱动架构在分析声音数据方面的有效性,并比较了它们与CNN在机器故障检测的特定任务上生成的嵌入。
摘要:In recent years, Sound AI is being increasingly used to predict machine failures. By attaching a microphone to the machine of interest, one can get real time data on machine behavior from the field. Traditionally, Convolutional Neural Net (CNN) architectures have been used to analyze spectrogram images generated from the sounds captured and predict if the machine is functioning as expected. CNN architectures seem to work well empirically even though they have biases like locality and parameter-sharing which may not be completely relevant for spectrogram analysis. With the successful application of transformer-based models in the field of image processing starting with Vision Transformer (ViT) in 2020, there has been significant interest in leveraging these in the field of Sound AI. Since transformer-based architectures have significantly lower inductive biases, they are expected to perform better than CNNs at spectrogram analysis given enough data. This paper demonstrates the effectiveness of transformer-driven architectures in analyzing Sound data and compares the embeddings they generate with CNNs on the specific task of machine fault detection.
【3】Adaptive Test-Time Scaling for Zero-Shot Respiratory Audio Classification
标题:用于零激发呼吸音分类的自适应测试时间缩放
链接:https://arxiv.org/abs/2604.12647
备注:Accepted at AHLI CHIL 2026
摘要:自动呼吸音分析有望实现可扩展的非侵入性疾病筛查,但进展受到稀缺的标记数据和昂贵的专家注释的限制。Zero-shot推理消除了特定于任务的监督,但现有方法对每个输入应用统一的计算,无论其难度如何。我们介绍了TRIAGE,一个分层的zero-shot框架,通过逐步丰富的推理阶段路由每个音频样本,自适应地扩展测试时的计算:在联合音频文本嵌入空间(Tier-L)中的快速标签余弦评分,与临床医生风格的描述符(Tier-M)的结构化匹配,以及检索增强的大型语言模型推理(Tier-H)。基于置信度的路由器可以提前完成简单的预测,同时将额外的计算分配给模糊的输入,使近一半的样本能够在最便宜的层退出。在没有特定任务训练的九个呼吸分类任务中,TRIAGE实现了0.744的平均AUROC,优于先前的zero-shot方法,并在多个任务上匹配或超过监督基线。我们的分析表明,测试时间缩放将收益集中在重要的地方:不确定的情况下可以看到高达19%的相对改善,而有信心的预测以最小的成本保持不变。
摘要:Automated respiratory audio analysis promises scalable, non-invasive disease screening, yet progress is limited by scarce labeled data and costly expert annotation. Zero-shot inference eliminates task-specific supervision, but existing methods apply uniform computation to every input regardless of difficulty. We introduce TRIAGE, a tiered zero-shot framework that adaptively scales test-time compute by routing each audio sample through progressively richer reasoning stages: fast label-cosine scoring in a joint audio-text embedding space (Tier-L), structured matching with clinician-style descriptors (Tier-M), and retrieval-augmented large language model reasoning (Tier-H). A confidence-based router finalizes easy predictions early while allocating additional computation to ambiguous inputs, enabling nearly half of all samples to exit at the cheapest tier. Across nine respiratory classification tasks without task-specific training, TRIAGE achieves a mean AUROC of 0.744, outperforming prior zero-shot methods and matching or exceeding supervised baselines on multiple tasks. Our analysis show that test-time scaling concentrates gains where they matter: uncertain cases see up to 19% relative improvement while confident predictions remain unchanged at minimal cost.
【4】Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs
标题:超越转录:感知音频的统一音频模式LLM
链接:https://arxiv.org/abs/2604.12506
备注:Accepted to ACL 2026 Findings
摘要:最近的音频大语言模型(AudioLLM)表现出惊人的性能反转:虽然在复杂的推理任务中表现出色,但它们在细粒度的声学感知上一直表现不佳。我们将这一差距归因于以ASR为中心的训练的根本局限性,它提供了精确的语言目标,但隐含地教导模型将非语言线索和声学事件作为噪声来抑制。为了解决这个问题,我们提出了统一音频模式(UAS),这是一个整体和结构化的监督框架,它将音频信息组织成三个明确的组件-转录,副语言和非语言事件-在统一的JSON格式中。这种设计实现了全面的声学覆盖,而不牺牲紧密的音频文本对齐,使推理。我们通过将其应用于离散和连续AudioLLM架构来验证这种监督策略的有效性。在MMSU、MMAR和MMAU上进行的大量实验表明,UAS-Audio产生了一致的改进,与相同大小的最先进模型相比,MMSU上的细粒度感知提高了10.9%,同时保留了强大的推理能力。我们的代码和模型可在https://github.com/Tencent/Unified_Audio_Schema上公开获取。
摘要:Recent Audio Large Language Models (AudioLLMs) exhibit a striking performance inversion: while excelling at complex reasoning tasks, they consistently underperform on fine-grained acoustic perception. We attribute this gap to a fundamental limitation of ASR-centric training, which provides precise linguistic targets but implicitly teaches models to suppress paralinguistic cues and acoustic events as noise. To address this, we propose Unified Audio Schema (UAS), a holistic and structured supervision framework that organizes audio information into three explicit components -- Transcription, Paralinguistics, and Non-linguistic Events -- within a unified JSON format. This design achieves comprehensive acoustic coverage without sacrificing the tight audio-text alignment that enables reasoning. We validate the effectiveness of this supervision strategy by applying it to both discrete and continuous AudioLLM architectures. Extensive experiments on MMSU, MMAR, and MMAU demonstrate that UAS-Audio yields consistent improvements, boosting fine-grained perception by 10.9% on MMSU over the same-size state-of-the-art models while preserving robust reasoning capabilities. Our code and model are publicly available at https://github.com/Tencent/Unified_Audio_Schema.
【5】Elastic Net Regularization and Gabor Dictionary for Classification of Heart Sound Signals using Deep Learning
标题:弹性网络正规化和Gabor词典用于使用深度学习对心弦信号进行分类
链接:https://arxiv.org/abs/2604.12483
摘要:在这篇文章中,我们提出了优化的时间-频率原子的分辨率和拟合模型的正则化,以获得更好的表示心音信号。这是通过评估深度学习(DL)网络在基于从拟合模型导出的一类新的时频特征矩阵来区分五种心脏瓣膜疾病方面的分类性能来完成的。我们检查了分辨率和正则化的几种组合,最佳组合是提供最高性能的组合。为此,使用线性模型的弹性网络正则化,基于心音信号和Gabor原子的过完备字典来获得拟合模型。我们考虑两种不同的DL架构,第一种主要由1D卷积神经网络(CNN)层和长短期记忆(LSTM)层组成,而第二种由1D和2D CNN层以及LSTM层组成。该网络使用两种算法进行训练,即动量随机梯度下降(SGDM)和自适应矩(ADAM)。已经使用包含五种心脏瓣膜状况的心音信号的数据库进行了广泛的实验。最好的分类精度为98.95美元的第二个架构时,训练与ADAM和特征矩阵来自最佳模型获得的Gabor字典组成的原子与高时间低频分辨率和施加稀疏的模型。
摘要:In this article, we propose the optimization of the resolution of time-frequency atoms and the regularization of fitting models to obtain better representations of heart sound signals. This is done by evaluating the classification performance of deep learning (DL) networks in discriminating five heart valvular conditions based on a new class of time-frequency feature matrices derived from the fitting models. We inspect several combinations of resolution and regularization, and the optimal one is that provides the highest performance. To this end, a fitting model is obtained based on a heart sound signal and an overcomplete dictionary of Gabor atoms using elastic net regularization of linear models. We consider two different DL architectures, the first mainly consisting of a 1D convolutional neural network (CNN) layer and a long short-term memory (LSTM) layer, while the second is composed of 1D and 2D CNN layers followed by an LSTM layer. The networks are trained with two algorithms, namely stochastic gradient descent with momentum (SGDM) and adaptive moment (ADAM). Extensive experimentation has been conducted using a database containing heart sound signals of five heart valvular conditions. The best classification accuracy of $98.95\%$ is achieved with the second architecture when trained with ADAM and feature matrices derived from optimal models obtained with a Gabor dictionary consisting of atoms with high-time low-frequency resolution and imposing sparsity on the models.
【6】Audio Source Separation in Reverberant Environments using $β$-divergence based Nonnegative Factorization
标题:使用基于$β$-分歧的非负因式分解在回响环境中分离音频源
链接:https://arxiv.org/abs/2604.12480
摘要:在基于高斯模型的多声道音频源分离中,源信号的观测混合的似然性由源频谱方差和相关联的空间协方差矩阵参数化。这些参数通过期望最大化算法通过最大化似然来估计,并用于通过多通道维纳滤波来分离信号。 我们建议估计这些参数应用非负因子分解的基础上源方差的先验信息。在非负分解中,谱基矩阵可以定义为先验信息。矩阵可以通过预先训练的冗余库提取或间接提供。在一个单独的步骤中,应用非负张量因式分解,提出了两种算法,以提取或检测的基础矩阵,最好地表示在所观察到的混合物的源信号的功率谱。通过乘性更新规则最小化$β$-发散来实现因子分解。因子分解的稀疏性可以通过调整$β$的值来控制。 实验表明,稀疏性,而不是在训练中分配给$β$的值,是至关重要的,以提高分离性能。在几种混合条件下评价了所提出的方法。它提供了更好的分离质量相对于其他类似的算法。
摘要:In Gaussian model-based multichannel audio source separation, the likelihood of observed mixtures of source signals is parametrized by source spectral variances and by associated spatial covariance matrices. These parameters are estimated by maximizing the likelihood through an Expectation-Maximization algorithm and used to separate the signals by means of multichannel Wiener filtering. We propose to estimate these parameters by applying nonnegative factorization based on prior information on source variances. In the nonnegative factorization, spectral basis matrices can be defined as the prior information. The matrices can be either extracted or indirectly made available through a redundant library that is trained in advance. In a separate step, applying nonnegative tensor factorization, two algorithms are proposed in order to either extract or detect the basis matrices that best represent the power spectra of the source signals in the observed mixtures. The factorization is achieved by minimizing the $β$-divergence through multiplicative update rules. The sparsity of factorization can be controlled by tuning the value of $β$. Experiments show that sparsity, rather than the value assigned to $β$ in the training, is crucial in order to increase the separation performance. The proposed method was evaluated in several mixing conditions. It provides better separation quality with respect to other comparable algorithms.
【7】On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
标题:统一重建、理解和生成语音VAE的蒸馏损失函数
链接:https://arxiv.org/abs/2604.12383
备注:Submitted to Interspeech 2026
摘要:基于变分自编码器(VAE)的连续语音表示已经成为传统的频谱图或基于离散令牌的语音生成和重建功能的一种有前途的替代方案。最近的研究试图通过与自监督学习(SSL)特征相结合来丰富VAE潜在表示中的结构信息,以获得更好的生成性能。然而,目前还不清楚广泛使用的对齐方法的基础上的时间轴蒸馏是最佳的,当考虑更多的任务。为了解决这个问题,本文系统地探讨了不同的对齐方法,并分析了它们对三个轴上的性能的影响:重建,理解和生成。我们研究了蒸馏损失的各种设计选择。大量的实验表明,联合边缘对齐方法与自适应加权可以实现最佳的整体性能,同时允许一个可控的平衡。
摘要:Continuous speech representations based on Variational Autoencoders (VAEs) have emerged as a promising alternative to traditional spectrogram or discrete token based features for speech generation and reconstruction. Recent research has tried to enrich the structural information in VAE latent representations by aligning with self-supervised learning (SSL) features, aiming for better generation performance. However, it remains unclear whether the widely-used alignment approach based on time-axis distillation is optimal when considering more tasks. To address this problem, this paper systematically explores different alignment approaches and analyzes their impact on the performances over three axes: reconstruction, understanding, and generation. We investigate various design choices in the distillation loss. Extensive experiments show that the joint-marginal alignment approach with adaptive weighting can achieve the best overall performance while allowing for a controllable balance.
【8】CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
标题:CoSyncDiT:电影配音的认知同步扩散Transformer
链接:https://arxiv.org/abs/2604.12292
摘要:电影配音的目标是合成语音,保留参考音频的声音身份,同时与目标视频中的嘴唇运动同步。现有的方法无法实现精确的唇同步,缺乏自然,由于明确的对齐在持续时间水平。虽然隐式对齐解决方案已经出现,但它们仍然容易受到参考音频的干扰,从而在野外场景中引发音色和发音退化。在本文中,我们提出了一种新的基于流匹配的电影配音框架驱动的认知同步扩散Transformer(CoSync-DiT),灵感来自专业演员的认知过程。该架构通过执行声学风格自适应、细粒度视觉校准和时间感知上下文对齐来逐步引导噪声到语音生成轨迹。此外,我们设计了联合语义和对齐正则化(JSAR)机制,同时约束帧级的时间一致性的上下文输出和语义一致性的流隐藏状态,确保鲁棒的对齐。在标准基准测试和具有挑战性的野外配音基准测试上进行的大量实验表明,我们的方法在多个指标上都达到了最先进的性能。
摘要:Movie dubbing aims to synthesize speech that preserves the vocal identity of a reference audio while synchronizing with the lip movements in a target video. Existing methods fail to achieve precise lip-sync and lack naturalness due to explicit alignment at the duration level. While implicit alignment solutions have emerged, they remain susceptible to interference from the reference audio, triggering timbre and pronunciation degradation in in-the-wild scenarios. In this paper, we propose a novel flow matching-based movie dubbing framework driven by the Cognitive Synchronous Diffusion Transformer (CoSync-DiT), inspired by the cognitive process of professional actors. This architecture progressively guides the noise-to-speech generative trajectory by executing acoustic style adapting, fine-grained visual calibrating, and time-aware context aligning. Furthermore, we design the Joint Semantic and Alignment Regularization (JSAR) mechanism to simultaneously constrain frame-level temporal consistency on the contextual outputs and semantic consistency on the flow hidden states, ensuring robust alignment. Extensive experiments on both standard benchmarks and challenging in-the-wild dubbing benchmarks demonstrate that our method achieves the state-of-the-art performance across multiple metrics.
【9】StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
标题:StableToken:用于弹性SpeechLLM的噪音稳健语义语音令牌化器
链接:https://arxiv.org/abs/2509.22220
备注:Accepted to ICLR 2026
摘要:流行的语义语音标记,旨在捕捉语言内容,是令人惊讶的脆弱。我们发现它们对意义无关的声学扰动不鲁棒;即使在语音完全可理解的高信噪比(SNR)下,它们的输出令牌序列也会急剧变化,增加下游LLM的学习负担。这种不稳定性源于两个缺陷:脆弱的单路径量化架构和对中间令牌稳定性漠不关心的远程训练信号。为了解决这个问题,我们引入了StableToken,这是一种通过共识驱动机制实现稳定性的令牌化器。它的多分支架构并行处理音频,这些表示通过强大的逐位投票机制合并,以形成单个稳定的令牌序列。StableToken在令牌稳定性方面树立了新的先进水平,在各种噪声条件下大幅降低了单位编辑距离(UED)。这种基础稳定性直接转化为下游优势,显著提高了SpeechLLM在各种任务上的鲁棒性。我们的代码和模型可在https://github.com/Tencent/StableToken上公开获取。
摘要:Prevalent semantic speech tokenizers, designed to capture linguistic content, are surprisingly fragile. We find they are not robust to meaning-irrelevant acoustic perturbations; even at high Signal-to-Noise Ratios (SNRs) where speech is perfectly intelligible, their output token sequences can change drastically, increasing the learning burden for downstream LLMs. This instability stems from two flaws: a brittle single-path quantization architecture and a distant training signal indifferent to intermediate token stability. To address this, we introduce StableToken, a tokenizer that achieves stability through a consensus-driven mechanism. Its multi-branch architecture processes audio in parallel, and these representations are merged via a powerful bit-wise voting mechanism to form a single, stable token sequence. StableToken sets a new state-of-the-art in token stability, drastically reducing Unit Edit Distance (UED) under diverse noise conditions. This foundational stability translates directly to downstream benefits, significantly improving the robustness of SpeechLLMs on a variety of tasks. Our code and model are publicly available at https://github.com/Tencent/StableToken.
【10】Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization
标题:为什么您的代币器在信息融合中失败:用于视频增强音频代币化的时间感知预量化融合
链接:https://arxiv.org/abs/2604.12145
摘要:音频标记化已成为端到端音频语言模型中的一个关键组件,为音频理解和生成任务提供了高效的离散表示学习。然而,现有的音频tokenizer面临的理解任务,由于单模态的约束,特别是当音频信号包含模糊或不完整的信息的基本限制。虽然结合额外的模态信息可以显着提高音频的理解,目前的多模态融合方法总是降低重建质量。这种降级对于需要高保真音频生成能力的端到端音频系统是不可接受的。在这项工作中,我们调查了视频增强音频标记化中重建质量下降的根本原因,并提出了三个关键发现。首先,融合在标记器架构中的位置对于保持重建质量至关重要。第二,我们表明,对比学习,虽然有效的连续表示融合,是不适合离散标记,因为它不能提高下游任务的性能。第三,虽然特征维度融合方法取得了一定的成功,但我们发现,在独特特征概念的指导下,沿时间轴融合会产生更好的结果。基于这些见解,我们引入了用于视频增强音频令牌化的定时感知预量化融合,这是第一种成功将视觉信息集成到音频令牌化器架构中同时保持重建保真度的方法。我们的方法不仅保持了高保真重建,而且与仅音频标记器和已建立的多模态融合基线相比,在下游理解任务上实现了卓越的性能。
摘要:Audio tokenization has emerged as a critical component in end-to-end audio language models, enabling efficient discrete representation learning for both audio understanding and generation tasks. However, existing audio tokenizers face fundamental limitations in understanding tasks due to single-modality constraints, particularly when audio signals contain ambiguous or incomplete information. While incorporating additional modality information can significantly enhance audio understanding, current multimodal fusion approaches invariably degrade reconstruction quality. This degradation is unacceptable for end-to-end audio systems that require high-fidelity audio generation capabilities. In this work, we investigate the root causes of reconstruction quality degradation in video-enhanced audio tokenization and present three key findings. First, the location of fusion within the tokenizer architecture is crucial for preserving reconstruction quality. Second, we show that contrastive learning, though effective in continuous representation fusion, is unsuitable for discrete tokenizers as it fails to enhance downstream task performance. Third, while feature-dimension fusion approaches achieve moderate success, we discover that fusing along the temporal axis -- guided by the concept of distinctive features -- yields significantly better results. Building on these insights, we introduce the Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization, the first approach to successfully integrate visual information into audio tokenizer architectures while preserving reconstruction fidelity. Our approach not only maintains high-fidelity reconstruction but also achieves superior performance on downstream understanding tasks compared with audio-only tokenizers and established multimodal fusion baselines.
【1】Four Decades of Digital Waveguides
标题:四十年的数字光导管
链接:https://arxiv.org/abs/2604.12878
摘要:与计算物理中常用的一般有限差分法相比,数字波导物理建模提供了声波传播的有效模拟。这种效率使得能够实时实现物理建模的乐器和声音效果,以及实时声乐模型和人工混响。本文概述了数字波导建模的历史演变和应用,并重点介绍了该领域的最新进展。使用经典的,进化的和神经的方法参数优化进行了讨论和比较。数字波导提供了物理上精确的模拟,降低了计算成本,现在可以用现代机器学习和可微分数字信号处理技术进行优化。
摘要:Digital waveguide physical modeling offers efficient simulation of acoustic wave propagation as compared to general finite-difference schemes commonly used in computational physics. This efficiency has enabled the real-time implementation of physically modeled musical instruments and sound effects, as well as real-time vocal models and artificial reverberation. This paper provides an overview of the historical evolution and applications of digital waveguide modeling and highlights recent advances in the field. Parametric optimization using classical, evolutionary and neural approaches are also discussed and compared. Digital waveguides provide physically accurate simulations with reduced computational cost, and can now be optimized with modern machine learning and differentiable digital signal processing techniques.
【2】Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models
标题:Audio-Cogito:在大型音频语言模型中实现深度音频推理
链接:https://arxiv.org/abs/2604.12527
备注:Submitted to Interspeech 2026
摘要:推理模型的最新进展推动了文本和多模态领域的重大进展,但音频推理仍然相对有限。只有少数大型音频语言模型(LALM)包含显式的思想链(CoT)推理,并且它们的能力通常不一致,不足以完成复杂的任务。为了弥补这一差距,我们引入了Audio-Cogito,这是一种用于深度音频推理的完全开源解决方案。我们开发了Cogito-pipe,用于高质量的音频推理数据管理,产生了545 k个推理样本,这些样本将在审查后发布。基于此数据集,我们采用自蒸馏策略进行模型微调。MMAR基准测试是唯一评估CoT过程的音频基准测试,实验表明,我们的模型在开源模型中达到了最佳性能,并在特定指标上匹配或超过某些闭源模型。我们的方法也在Interspeech 2026音频推理挑战赛中跻身顶级系统之列。
摘要:Recent advances in reasoning models have driven significant progress in text and multimodal domains, yet audio reasoning remains relatively limited. Only a few Large Audio Language Models (LALMs) incorporate explicit Chain-of-Thought (CoT) reasoning, and their capabilities are often inconsistent and insufficient for complex tasks. To bridge this gap, we introduce Audio-Cogito, a fully open-source solution for deep audio reasoning. We develop Cogito-pipe for high-quality audio reasoning data curation, producing 545k reasoning samples that will be released after review. Based on this dataset, we adopt a self-distillation strategy for model fine-tuning. Experiments on the MMAR benchmark, the only audio benchmark evaluating the CoT process, show that our model achieves the best performance among open-source models and matches or surpasses certain closed-source models in specific metrics. Our approach also ranks among the top-tier systems in the Interspeech 2026 Audio Reasoning Challenge.
【3】X-VC: Zero-shot Streaming Voice Conversion in Codec Space
标题:X-VC:编解码器空间中的Zero-Shot流语音转换
链接:https://arxiv.org/abs/2604.12456
摘要:Zero-shot语音转换(VC)的目的是将源话语转换成一个看不见的目标说话人的声音,同时保持其语言内容。虽然最近的系统已经提高了转换质量,但构建用于交互式场景的zero-shot VC系统仍然具有挑战性,因为高保真扬声器传输和低延迟流推理难以同时实现。在这项工作中,我们提出了X-VC,一个zero-shot流VC系统,在预训练的神经编解码器的潜在空间中执行一步转换。X-VC使用双调节声学转换器,该转换器联合建模源编解码器潜伏期和从目标参考语音导出的帧级声学条件,同时通过自适应归一化注入话语级目标说话者信息。为了减少训练和推理之间的不匹配,我们使用生成的成对数据和结合标准、重构和反向模式的角色分配策略来训练模型。对于流式推理,我们进一步采用了具有重叠平滑的分块推理方案,该方案与编解码器的基于段的训练范例相一致。在Seed-TTS-Eval上的实验表明,X-VC在英语和汉语中均获得了最佳的流媒体WER,在相同语言和跨语言设置中具有很强的说话人相似性,并且离线实时因子显著低于比较基线。这些结果表明,编解码器空间的一步转换是一种实用的方法,用于构建高质量的低延迟zero-shot VC系统。音频样本可在https://x-vc.github.io上获得。我们的代码和检查点也将被释放。
摘要:Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content. Although recent systems have improved conversion quality, building zero-shot VC systems for interactive scenarios remains challenging because high-fidelity speaker transfer and low-latency streaming inference are difficult to achieve simultaneously. In this work, we present X-VC, a zero-shot streaming VC system that performs one-step conversion in the latent space of a pretrained neural codec. X-VC uses a dual-conditioning acoustic converter that jointly models source codec latents and frame-level acoustic conditions derived from target reference speech, while injecting utterance-level target speaker information through adaptive normalization. To reduce the mismatch between training and inference, we train the model with generated paired data and a role-assignment strategy that combines standard, reconstruction, and reversed modes. For streaming inference, we further adopt a chunkwise inference scheme with overlap smoothing that is aligned with the segment-based training paradigm of the codec. Experiments on Seed-TTS-Eval show that X-VC achieves the best streaming WER in both English and Chinese, strong speaker similarity in same-language and cross-lingual settings, and substantially lower offline real-time factor than the compared baselines. These results suggest that codec-space one-step conversion is a practical approach for building high-quality low-latency zero-shot VC systems. Audio samples are available at https://x-vc.github.io. Our code and checkpoints will also be released.
【4】Sky-Ear: An Unmanned Aerial Vehicle-Enabled Victim Sound Detection and Localization System
标题:Sky-Ear:一种无人机受害者声音检测和定位系统
链接:https://arxiv.org/abs/2604.12455
摘要:无人机(UAV)越来越多地用于搜救(SAR)任务,但由于机载硬件的限制,持续可靠的受害者检测和定位仍然具有挑战性。本文设计了一个无人机辅助的受害者声音检测和定位系统(简称“天耳”),以实现节能的声传感和声音检测的合成孔径雷达。基于圆形麦克风阵列,开发了两级(Sentinel和Responder)音频处理,用于高能耗和高可靠性的声音检测。在Sentinel阶段设计了一种基于掩蔽自动编码器(MAE)的声音检测方法,用于分析频率-时间声学特征。为了提高定位精度,设计了一种连续定位方法,通过优化检测方向从多个观测。大量的仿真实验进行了验证系统的受害者检测精度和定位误差方面的性能。
摘要:Unmanned Aerial Vehicles (UAVs) are increasingly deployed in search-and-rescue (SAR) missions, yet continuous and reliable victim detection and localization remain challenging due to on-board hardware constraints. This paper designs an UAV-Enabled Victim Sound Detection and Localization System (called ``Sky-Ear'' for brevity) to achieve energy-efficient acoustic sensing and sound detection for SAR. Based on a circular-shaped microphone array, two-stage (Sentinel and Responder) audio processing is developed for energy-consuming and highly reliable sound detection. A Masking autoencoder (MAE)-based sound detection method is designed in the Sentinel stage to analyze frequency-time acoustic features. For improved precision, a continuous localization method is designed by optimizing detected directions from multiple observations. Extensive simulation experiments are conducted to validate the system's performance in terms of victim detection accuracy and localization error.
【5】Room compensation for loudspeaker reproduction using a supporting source
标题:使用支持源的扬声器再现的房间补偿
链接:https://arxiv.org/abs/2604.12439
摘要:房间补偿旨在提高混响环境中扬声器再现的准确性。然而,传统的方法仅限于提高频谱(音色)和时间精度,忽略了扬声器再现的空间精度。提出了一种通过使用延迟的次级支持源以频率选择性方式向感知的混响声场添加能量来补偿扬声器再现的频谱和空间属性的方法。这种方法允许根据频率修改直达混响比,从而改变空间和频谱再现。所提出的方法进行了感知评估,证明其能够改变一个主扬声器的感知,而无需听众感知的支持源。结果表明,该方法执行的一个成熟的商业房间补偿算法,并有几个优于传统的房间补偿方法。
摘要:Room compensation aims to improve the accuracy of loudspeaker reproduction in reverberant environments. Traditional methods, however, are limited to improving only spectral (timbral) and temporal accuracy, neglecting the spatial accuracy of loudspeaker reproduction. Proposed is a method that compensates for both spectral and spatial properties of loudspeaker reproduction, by adding energy to the perceived reverberant sound field in a frequency-selective manner using a delayed secondary supporting source. This approach allows for the modification of the direct to reverberant ratio as a function of frequency, altering spatial and spectral reproduction. The proposed method is perceptually evaluated, demonstrating its ability to alter the perception of a primary loudspeaker without the listener perceiving the supporting source. The results show that the proposed method performs comparably to a well-established commercial room compensation algorithm and has several advantages over traditional room compensation methods.
【6】An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding
标题:通过逐块生成和逐深度编解码的超低延迟、端到端流语音合成架构
链接:https://arxiv.org/abs/2604.12438
备注:29 pages, 5 figures
摘要:实时语音合成需要在交互式应用中平衡推理延迟和声音保真度。传统的连续文本到语音流水线需要计算密集型神经声码器来重建相位信息,从而产生显著的流传输瓶颈。此外,基于回归的声学建模经常引起频谱过平滑伪影。为了解决这些局限性,本文提出了一种新型的端到端非自回归架构,该架构针对超低延迟逐块生成进行了优化,直接对Mimi神经音频编解码器的高度压缩离散潜在空间进行建模。该架构将改进的FastSpeech 2骨干与渐进式深度顺序解码策略相结合,动态调节32层残差矢量量化码。该机制解决了语音对齐退化问题,并在没有时间自回归开销的情况下管理高保真离散表示的复杂性。在英语和马来语数据集上的实验评估验证了其独立于语言的部署能力。与传统的连续回归模型相比,所提出的架构在基本发声准确性方面表现出定量的改进,并减轻了高频频谱退化。它实现了超低延迟推理,与传统级联流水线相比,绝对加速提高了10.6倍。最重要的是,该系统实现了48.99毫秒的平均首字节延迟时间,大大低于实时交互流的人类感知阈值。这些结果坚定地建立了所提出的架构作为一个高度优化的解决方案,部署实时流语音接口。
摘要:Real-time speech synthesis requires balancing inference latency and acoustic fidelity for interactive applications. Conventional continuous text-to-speech pipelines require computationally intensive neural vocoders to reconstruct phase information, creating a significant streaming bottleneck. Furthermore, regression-based acoustic modeling frequently induces spectral over-smoothing artifacts. To address these limitations, this paper proposes a novel end-to-end non-autoregressive architecture optimized for ultra-low latency block-wise generation, directly modeling the highly compressed discrete latent space of the Mimi neural audio codec. Integrating a modified FastSpeech 2 backbone with a progressive depth-wise sequential decoding strategy, the architecture dynamically conditions 32 layers of residual vector quantization codes. This mechanism resolves phonetic alignment degradation and manages the complexity of high-fidelity discrete representations without temporal autoregressive overhead. Experimental evaluations on English and Malay datasets validate its language-independent deployment capability. Compared to conventional continuous regression models, the proposed architecture demonstrates quantitative improvements in fundamental voicing accuracy and mitigates high-frequency spectral degradation. It achieves ultra-low latency inference, translating to a 10.6-fold absolute acceleration over conventional cascaded pipelines. Crucially, the system achieves an average time-to-first-byte latency of 48.99 milliseconds, falling significantly below the human perception threshold for real-time interactive streaming. These results firmly establish the proposed architecture as a highly optimized solution for deploying real-time streaming speech interfaces.
【7】Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction
标题:具有常见词线索和偏向词位置预测的语音LLM中ASB的上下文偏向
链接:https://arxiv.org/abs/2604.12398
备注:Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:语音感知LLM(SLLM)最近已经达到了最先进的ASR性能;然而,它们仍然无法准确地转录训练数据中很少或从未出现的偏见词。上下文偏置机制通常通过经由文本提示或附加模块将预定义的偏置词列表引入到模型中来实现。为了进一步改进,可以将预定义的偏置词与它们的音素表示配对作为发音线索。通常,音素序列是通过G2 P系统生成的,该系统覆盖目标语言和偏置词的域。因此,当兼容的G2 P系统不可用时,音素辅助的上下文偏置变得难以执行。此外,手动添加准确的音素序列需要高级的语音知识。在本文中,我们探讨了语境偏置SLLM基于声学线索与一组常见的单词,其发音是部分类似的目标偏见的话。我们假设ASR应用程序中,最终用户不需要特殊的语音知识或利用G2 P工具进行推理。为了增强鲁棒性,我们还引入了以多输出学习方式实现的偏置词位置预测。与基线系统相比,我们的方法将偏见单词识别错误减少了16.3%,包括域外数据。
摘要:Speech-aware LLMs (SLLMs) have recently achieved state-of-the-art ASR performance; however, they still fail to accurately transcribe bias words that appear rarely or never in the training data. Contextual biasing mechanisms are commonly implemented by introducing a predefined bias word list into the model via a text prompt or additional module. For further improvement, predefined bias words can be paired with their phoneme representations as pronunciation cues. Typically, phoneme sequences are generated through a G2P system that covers the target languages and domains of the bias words. Therefore, when a compatible G2P system is unavailable, phoneme-assisted contextual biasing becomes difficult to perform. Moreover, manually adding accurate phoneme sequences requires advanced phonetic knowledge. In this paper, we explore contextual biasing in SLLM based on acoustic cues associated with a set of common words whose pronunciations are partially similar to those of the target bias words. We assume ASR applications in which end users do not require special knowledge of phonetics or utilize G2P tools for inference. For enhanced robustness, we also introduce bias word positional prediction implemented in a multi-output learning fashion. Our method reduces bias word recognition errors by 16.3% compared to baseline systems, including on out-of-domain data.
【8】VoxEffects: A Speech-Oriented Audio Effects Dataset and Benchmark
标题:VoxEffects:面向语音的音效数据集和基准
链接:https://arxiv.org/abs/2604.12389
摘要:野外语音音频通常经过后期制作效果处理,但现有的语音数据集很少提供精确的效果和参数注释,限制了系统研究。我们介绍VoxEffects,这是一个语音音频效果数据集,它将产生的语音与多粒度的精确效果链监督配对。VoxEffects支持面向语音的音频效果识别:给定生成的波形,推断存在哪些效果以及如何应用它们。它由最少编辑的干净语音构建而成,为离线合成和实时渲染提供了可扩展的渲染管道,以实现高效的训练和评估。音频效果识别基准包括效果存在检测、预设分类和强度预测,具有覆盖捕获端和平台端降级的鲁棒性协议。我们提供了一个基于AudioMAE的多任务基线和域转移,鲁棒性,输入持续时间和性别公平性的分析。
摘要:Speech audio in the wild is often processed by post-production effects, but existing speech datasets rarely provide precise annotations of effects and parameters, limiting systematic study. We introduce VoxEffects, a speech audio effects dataset that pairs produced speech with exact effect-chain supervision at multiple granularities. VoxEffects supports speech-oriented audio effect identification: given a produced waveform, infer which effects are present and how they are applied. Built from minimally edited clean speech, it provides an extensible rendering pipeline for both offline synthesis and on-the-fly rendering for efficient training and evaluation. The audio effect identification benchmark includes effect presence detection, preset classification, and intensity prediction, with a robustness protocol covering capture-side and platform-side degradations. We provide an AudioMAE-based multi-task baseline and analyses of domain shift, robustness, input duration, and gender fairness.
【9】TokenSE: a Mamba-based discrete token speech enhancement framework for cochlear implants
标题:TokenSE:一个基于Mamba的用于人工智能的离散令牌语音增强框架
链接:https://arxiv.org/abs/2604.12246
摘要:语音增强(SE)对于提高真实世界环境中的语音可懂度和质量至关重要,特别是对于在噪声和混响条件下语音理解严重退化的人工耳蜗(CI)用户。在这项研究中,我们提出了TokenSE,一个离散的基于令牌的SE框架,在神经音频编解码器空间中运行,它使用基于Mamba的模型从降级的语音中预测干净的编解码器令牌索引。与早期的Transformer架构不同,其自注意机制的计算复杂度与序列长度成二次方增长,Mamba的输入依赖选择机制实现了线性复杂度,使其成为Transformers的引人注目的替代方案,特别是对于CI和助听器(HA)应用。客观评估表明,TokenSE在域内和域外数据集上的性能始终优于基线方法。此外,与CI用户的主观听力实验表明,在不利的噪声和混响环境下的语音清晰度明显的好处。
摘要:Speech enhancement (SE) is critical for improving speech intelligibility and quality in real-world environments, particularly for cochlear implant (CI) users who experience severe degradations in speech understanding under noisy and reverberant conditions. In this study, we propose TokenSE, a discrete token-based SE framework operating in the neural audio codec space, which predicts clean codec token indices from degraded speech using a Mamba-based model. Unlike the earlier Transformer architecture, whose self-attention mechanism has a computational complexity that grows quadratically with sequence length, the input-dependent selection mechanism of Mamba achieves linear complexity, making it a compelling alternative to Transformers, especially for CI and hearing-aid (HA) applications. Objective evaluations show that TokenSE consistently outperforms baseline methods on both in-domain and out-of-domain datasets. Moreover, subjective listening experiments with CI users indicate clear benefit in speech intelligibility under adverse noisy and reverberant environments.
【10】Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization
标题:为什么您的代币器在信息融合中失败:用于视频增强音频代币化的时间感知预量化融合
链接:https://arxiv.org/abs/2604.12145
摘要:音频标记化已成为端到端音频语言模型中的一个关键组件,为音频理解和生成任务提供了高效的离散表示学习。然而,现有的音频tokenizer面临的理解任务,由于单模态的约束,特别是当音频信号包含模糊或不完整的信息的基本限制。虽然结合额外的模态信息可以显着提高音频的理解,目前的多模态融合方法总是降低重建质量。这种降级对于需要高保真音频生成能力的端到端音频系统是不可接受的。在这项工作中,我们调查了视频增强音频标记化中重建质量下降的根本原因,并提出了三个关键发现。首先,融合在标记器架构中的位置对于保持重建质量至关重要。第二,我们表明,对比学习,虽然有效的连续表示融合,是不适合离散标记,因为它不能提高下游任务的性能。第三,虽然特征维度融合方法取得了一定的成功,但我们发现,在独特特征概念的指导下,沿时间轴融合会产生更好的结果。基于这些见解,我们引入了用于视频增强音频令牌化的定时感知预量化融合,这是第一种成功将视觉信息集成到音频令牌化器架构中同时保持重建保真度的方法。我们的方法不仅保持了高保真重建,而且与仅音频标记器和已建立的多模态融合基线相比,在下游理解任务上实现了卓越的性能。
摘要:Audio tokenization has emerged as a critical component in end-to-end audio language models, enabling efficient discrete representation learning for both audio understanding and generation tasks. However, existing audio tokenizers face fundamental limitations in understanding tasks due to single-modality constraints, particularly when audio signals contain ambiguous or incomplete information. While incorporating additional modality information can significantly enhance audio understanding, current multimodal fusion approaches invariably degrade reconstruction quality. This degradation is unacceptable for end-to-end audio systems that require high-fidelity audio generation capabilities. In this work, we investigate the root causes of reconstruction quality degradation in video-enhanced audio tokenization and present three key findings. First, the location of fusion within the tokenizer architecture is crucial for preserving reconstruction quality. Second, we show that contrastive learning, though effective in continuous representation fusion, is unsuitable for discrete tokenizers as it fails to enhance downstream task performance. Third, while feature-dimension fusion approaches achieve moderate success, we discover that fusing along the temporal axis -- guided by the concept of distinctive features -- yields significantly better results. Building on these insights, we introduce the Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization, the first approach to successfully integrate visual information into audio tokenizer architectures while preserving reconstruction fidelity. Our approach not only maintains high-fidelity reconstruction but also achieves superior performance on downstream understanding tasks compared with audio-only tokenizers and established multimodal fusion baselines.
【11】StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection
标题:StreamMark:一种基于深度学习的半脆弱音频水印,用于主动式Deepfake检测
链接:https://arxiv.org/abs/2604.11917
备注:ICASSP 2026
摘要:生成式人工智能的快速发展使得区分deepfake音频和真实的人类语音变得越来越具有挑战性。为了克服被动检测方法的局限性,我们提出了StreamMark,一种新的基于深度学习的半脆弱音频水印系统。StreamMark被设计为对保留语义含义的良性音频转换(例如,压缩、噪声),同时对恶意的、语义改变的操纵(例如,语音转换、语音编辑)。我们的方法在一个独特的编码器-失真-解码器架构中引入了一种复域嵌入技术,该架构经过明确训练以区分这两类变换。全面的基准测试表明,StreamMark实现了高不可感知性(SNR 24.16 dB,PESQ 4.20),对Opus编码等现实世界的失真具有弹性,并且对一系列deepfake攻击表现出原则性的脆弱性,消息恢复准确性下降到偶然水平(~50%),同时对基于AI的良性传输保持稳健(ACC >98%)。
摘要:The rapid advancement of generative AI has made it increasingly challenging to distinguish between deepfake audio and authentic human speech. To overcome the limitations of passive detection methods, we propose StreamMark, a novel deep learning-based, semi-fragile audio watermarking system. StreamMark is designed to be robust against benign audio conversions that preserve semantic meaning (e.g., compression, noise) while remaining fragile to malicious, semantics-altering manipulations (e.g., voice conversion, speech editing). Our method introduces a complex-domain embedding technique within a unique Encoder-Distortion-Decoder architecture, trained explicitly to differentiate between these two classes of transformations. Comprehensive benchmarks demonstrate that StreamMark achieves high imperceptibility (SNR 24.16 dB, PESQ 4.20), is resilient to real-world distortions like Opus encoding, and exhibits principled fragility against a suite of deepfake attacks, with message recovery accuracy dropping to chance levels (~50%), while remaining robust to benign AI-based style transfers (ACC >98%).
【12】MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
标题:MoshiRAG:适用于高保真语音语言模型的同步知识检索
链接:https://arxiv.org/abs/2604.12928
摘要:最近出现了语音到语音语言模型,以增强对话AI的自然性。特别是,全双工模式的特点是它们的实时交互性,包括暂停,中断和反向通道的处理。然而,提高其真实性仍然是一个公开的挑战。虽然缩放模型大小可以解决这一差距,但它会使实时推理变得过于昂贵。在这项工作中,我们提出了MoshiRAG,一个模块化的方法,结合了一个紧凑的全双工接口与选择性检索访问更强大的知识源。我们的异步框架使模型能够识别知识需求的查询,并在外部信息的基础上做出响应。通过利用响应开始和核心信息传递之间的自然时间间隙,可以在保持自然会话流的同时完成检索过程。通过这种方法,MoshiRAG实现了与最好的公开发布的非双工语音语言模型相媲美的真实性,同时保留了全双工系统固有的交互性。此外,我们灵活的设计支持即插即用的检索方法,而无需重新训练,并在域外数学推理任务上表现出强大的性能。
摘要:Speech-to-speech language models have recently emerged to enhance the naturalness of conversational AI. In particular, full-duplex models are distinguished by their real-time interactivity, including handling of pauses, interruptions, and backchannels. However, improving their factuality remains an open challenge. While scaling the model size could address this gap, it would make real-time inference prohibitively expensive. In this work, we propose MoshiRAG, a modular approach that combines a compact full-duplex interface with selective retrieval to access more powerful knowledge sources. Our asynchronous framework enables the model to identify knowledge-demanding queries and ground its responses in external information. By leveraging the natural temporal gap between response onset and the delivery of core information, the retrieval process can be completed while maintaining a natural conversation flow. With this approach, MoshiRAG achieves factuality comparable to the best publicly released non-duplex speech language models while preserving the interactivity inherent to full-duplex systems. Moreover, our flexible design supports plug-and-play retrieval methods without retraining and demonstrates strong performance on out-of-domain mathematical reasoning tasks.
机器翻译由腾讯交互翻译提供,仅供参考
