今日论文合集:cs.SD语音10篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Precise and Simple Audio-to-Score Alignment
标题:精确简单的音频与乐谱对齐
链接:https://arxiv.org/abs/2605.20014
作者:Silvan Peter,Patricia Hu,Gerhard Widmer
备注:published at the Music Encoding Conference (MEC) 2026
摘要:音频-乐谱对齐是音乐信息检索中一个长期存在的挑战,也是音乐研究中应用最广泛的对齐任务。对齐算法匹配一段音乐的两个版本,为了实现这一点,这些版本需要具有可比的格式。音频到音频对齐匹配音频特征;当将音频文件与乐谱匹配时,它们必须通过钢琴卷或类似的特征序列来合成乐谱或导出类似音频的特征。相比之下,符号对齐匹配符号编码的音符;在音频到乐谱的场景中,这些将通过音频文件的转录来获得。在这篇文章中,我们提出了一个算法,桥梁音频和符号级功能直接。编码起始和频谱激活的顺序音频特征通过来自符号对准方法的定制的基于动态规划的匹配算法与得分位置相匹配。由此产生的方法既精确-超越了广泛使用的基于合成分数的音频到音频方法-又在其数字信号处理组件中保持灵活,即,该方法适用于不同的音色特征,而不需要单独的转录模型。此外,它继承了一些符号对齐运行时优点,算法复杂度在最坏情况下在符号分数(通常短)和音频特征序列(通常长)的长度上是线性的。在下面的章节中,我们将提供详细的算法描述,并在大规模钢琴独奏录音数据集上评估其对齐质量。
摘要:Audio-to-score alignment is a long-standing challenge in music information retrieval and arguably the most widely applicable alignment task for music research. Alignment algorithms match two versions of a piece of music, and for this to work these versions need to be in comparable formats. Audio-to-audio alignment matches audio features; when matching audio files to scores, they must either synthesize the score or derive audio-like features by means of piano rolls or similar feature sequences. Symbolic alignment, by contrast, matches symbolically encoded notes; in an audio-to-score scenario these would be obtained by a transcription of the audio file. In this article, we present an algorithm that bridges audio-like and symbol-level features directly. Sequential audio features encoding onset and spectral activation are matched to score positions by a bespoke dynamic programming-based matching algorithm derived from symbolic alignment methods. The resulting method is both precise - surpassing widely used audio-to-audio approaches based on synthesized scores -, and remains flexible in its digital signal processing components, i.e., the method is adaptable to diverse timbral characteristics without requiring a separate transcription model. Furthermore it inherits some of the symbolic alignment runtime advantages with an algorithmic complexity that is at worst linear in the length of the (typically short) symbolic score and (typically long) audio feature sequence. In the following sections, we provide a detailed algorithm description and evaluate its alignment quality on a large-scale dataset of solo piano recordings.


【2】A conceptual framework for learning to listen by reward: Curiosity-driven search for novel sources

标题:学习通过奖励倾听的概念框架:好奇心驱动的新颖来源搜索
链接:https://arxiv.org/abs/2605.19984
作者:Andreas Triantafyllopoulos,Jakub Šťastný,Alexios Terpinas,Tianyi Liu,Yuanqi Wang,Björn W. Schuller
摘要:强化学习是一种强大的学习范式,在许多领域都取得了进展。它的核心承诺在于通过高级目标进行学习,而不需要粒度标签。然而,它在音频领域仍然难以捉摸,在那里它受到的关注远远少于计算机视觉或其他领域。关键问题仍然是:代理人如何通过奖励驱动的探索学习纯粹的倾听?在这篇文章中,我们提出了一个概述以前的尝试和一个新的概念框架,学习听奖励。我们的方法依赖于不断寻找新的声源。我们制定我们的框架,讨论开放的技术挑战,并提出了第一个概念验证的实施,展示了我们的方法的可行性。
摘要:Reinforcement learning is a powerful learning paradigm that has spearheaded progress in numerous domains. Its core promise lies in learning through high-level goals without the need for granular labels. However, it still remains elusive in the realm of audio, where it has received substantially less attention than in computer vision or other domains. The key question remains: how can agents learn to listen purely via reward-driven exploration? In this contribution, we present an overview of previous attempts and a new conceptual framework for learning to listen by reward. Our approach depends on the continuous search for novel sound sources. We formulate our framework, discuss open technical challenges, and present a first proof-of-concept implementation that showcases the feasibility of our approach.


【3】DASM: Domain-Aware Sharpness Minimization for Multi-Domain Voice Stream Steganalysis

标题:DASM:多域语音流隐写分析的域感知清晰度最小化
链接:https://arxiv.org/abs/2605.19955
作者:Pengcheng Zhou,Pianran Guo,Shuhua Chen,Mengqin Zhao,Zhongliang Yang,Linna Zhou
摘要:网络流媒体中信息隐藏技术的日益广泛应用对隐蔽通信构成了严重的安全威胁,因此需要开发鲁棒的检测技术。然而,现有的网络语音流隐写分析方法大多依赖于特定场景下的数据分布,难以适应非同源数据分布的实际检测需求。通过Hessian分析,我们发现主流模型的损失景观由众多鞍点和尖锐的局部极小值主导,使它们对数据分布变化高度敏感,并从根本上限制了泛化。因此,我们提出了一个新的优化器,域感知共享最小化(DASM)。DASM的核心机制包括两个方面:第一,它集成了域监督对比学习与锐度感知优化,在寻求平坦最小值的同时显式地保持域间特征分离;第二,我们设计了一种自适应域间隙调制策略,该策略通过感知不同域的实时特征可分性来动态校准优化损失权重。大量的实验结果表明,我们的方法优于国家的最先进的方法由一个大的利润,并实现了良好的泛化和鲁棒性。
摘要:The growing use of information hiding in network streaming media for covert communication poses a significant security threat, necessitating the development of robust detection technologies. However, existing steganalysis methods for network voice streams mostly rely on data distributions in specific scenarios, making it difficult to adapt to the practical detection needs of non-homologous data distributions. Through Hessian analysis, we find that the loss landscapes of mainstream models are dominated by numerous saddle points and sharp local minima, rendering them highly sensitive to data distribution shifts and fundamentally limiting generalization. Therefore, we propose a new optimizer, Domain-Aware Sharpness Minimization (DASM). The core mechanisms of DASM consist of two aspects: first, it integrates domain-supervised contrastive learning with sharpness-aware optimization, explicitly preserving inter-domain feature separation while seeking flat minima; second, we design an adaptive domain gap modulation strategy that dynamically calibrates the optimization loss weights by sensing the real-time feature separability of different domains. Extensive experimental results demonstrate that our method outperforms the state-of-the-art methods by a large margin and achieves excellent generalization and robustness.


【4】Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

标题:Mega-ASR:通过放大真实世界声学模拟实现野外语音识别
链接:https://arxiv.org/abs/2605.19833
作者:Zhifei Xie,Kaiyu Pang,Haobin Zhang,Deheng Ye,Xiaobin Hu,Shuicheng Yan,Chunyan Miao
备注:Project page: https://xzf-thu.github.io/Mega-ASR/. Code, models, and dataset will be released. A robust ASR framework targeting in-the-wild and compositional acoustic scenarios where conventional ASR systems fail
摘要:尽管自动语音识别(ASR)和大型音频语言模型的快速发展,但现实环境中的鲁棒性识别仍然受到“声学鲁棒性瓶颈”的限制:模型通常会失去声学基础,并在严重的合成失真下产生遗漏或幻觉。我们提出了Mega-ASR,一个统一的ASR-in-the-wild框架,它将可扩展的复合数据构建与渐进的声学到语义优化相结合。我们引入Voices-in-the-Wild-2 M,涵盖7个经典声学现象和54个物理上合理的复合场景,并使用声学到语义渐进式监督微调和双粒度WER-Gated策略优化训练Mega-ASR。大量实验表明,Mega-ASR在不利条件下的ASR基准测试中比现有技术系统具有显著优势(VOiCES R4-B-F上为45.69% vs. 54.01%,NOIZEUS Sta-0上为21.49% vs. 29.34%)。在复杂的合成声学场景中,Mega-ASR进一步提供了超过30%的相对WER降低,相对于强大的开源和闭源基线,建立了一个可扩展的范例,用于在野外实现强大的ASR。
摘要:Despite rapid advances in automatic speech recognition (ASR) and large audio-language models, robust recognition in real-world environments remains limited by an "acoustic robustness bottleneck": models often lose acoustic grounding and produce omissions or hallucinations under severe, compositional distortions. We propose Mega-ASR, a unified ASR-in-the-wild framework that combines scalable compound-data construction with progressive acoustic-to-semantic optimization. We introduce Voices-in-the-Wild-2M, covering 7 classic acoustic phenomena and 54 physically plausible compound scenarios, and train Mega-ASR with Acoustic-to-Semantic Progressive Supervised Fine-Tuning and Dual-Granularity WER-Gated Policy Optimization. Extensive experiments demonstrate that Mega-ASR achieves significant advantages over prior state-of-the-art systems on adverse-condition ASR benchmarks (45.69% vs. 54.01% on VOiCES R4-B-F, and 21.49% vs. 29.34% on NOIZEUS Sta-0). On complex compositional acoustic scenarios, Mega-ASR further delivers over 30% relative WER reduction against strong open- and closed-source baselines, establishing a scalable paradigm for robust ASR in-the-wild.


【5】Executable Boundary Contracts for Sound Event Traces

标题:健全事件痕迹的可执行边界合同
链接:https://arxiv.org/abs/2605.19632
作者:Faruk Alpay,Hamdi Alakkad
备注:39 pages. Finite frame core code, tables, manifests, and Lean checks are ancillary material
摘要:声音事件报告通常将定时边界行为压缩到帧、片段或事件分数中。本文定义了有限声音事件轨迹的可执行边界契约。框架片段是一个有界的布尔片段,在网格投影后可嵌入STL中。事件层添加声明的间隔匹配、持续时间子句、分段子句和义务限制向量评分。其目的是测量,而不是一个新的通用时间逻辑,也不是一个挑战排行榜。该人工制品评估受控的Mini LibriSpeech种子场景,MAESTRO Real音景,冻结的预训练定时探头和官方DCASE 2024任务4基线轨道。在这些轨道上,标准分数和合同坐标在可解释的方式上不一致。最强的真实语料库发现是联合活动可以隐藏类型化边界失败,而外部DCASE输出提供了一个类索引的挑战级别参考。代码、生成的表格、清单和有限框架核心的精益检查作为辅助材料提供。
摘要:Sound event reports often compress timed boundary behavior into frame, segment, or event scores. This paper defines executable boundary contracts for finite sound event traces. The frame fragment is a bounded Boolean fragment embeddable in STL after grid projection. The event layer adds declared interval matching, duration clauses, fragmentation clauses, and obligation restricted vector scoring. The aim is measurement, not a new general temporal logic and not a challenge leaderboard. The artifact evaluates controlled Mini LibriSpeech seeded scenes, MAESTRO Real soundscapes, frozen pretrained timing probes, and an official DCASE 2024 Task 4 baseline track. Across these tracks, standard scores and contract coordinates disagree in interpretable ways. The strongest real corpus finding is that union activity can hide typed boundary failure, while external DCASE outputs provide a class indexed challenge level reference. Code, generated tables, manifests, and Lean checks for the finite frame core are supplied as ancillary material.


【6】Optimising Neural Speech Codecs for 300bps Communication using Reinforcement Learning

标题:使用强化学习优化神经语音编解码器以实现300 Mbps通信
链接:https://arxiv.org/abs/2605.19541
作者:Junyi Wang,Chi Zhang,Jing Qian,Haifeng Luo,Hao Wang,Zengrui Jin,Chao Zhang
摘要:在卫星和水下信道等带宽受限的通信中,语音通常必须以超低比特率传输,其中可懂度是主要目标。在这种极端的压缩水平下,用声学重建损失训练的编解码器倾向于将比特分配给感知细节,导致字错误率(WER)的大幅下降。本文提出了ClariCodec,这是一种以300比特每秒(bps)的速度运行的神经语音编解码器,它将量化重新制定为随机策略,从而实现基于强化学习(RL)的可懂度优化。具体而言,编码器使用WER驱动的奖励进行微调,而声学重建管道保持冻结。即使没有RL,ClariCodec在LibriSpeech测试中以300 bps的速度实现了4.64%的WER,已经与以更高比特率运行的编解码器竞争。进一步的RL微调将测试清洁的WER降低至3.55%,测试其他的WER降低至10.4%,相当于相对降低23%,同时保持感知质量。
摘要:In bandwidth-constrained communication such as satellite and underwater channels, speech must often be transmitted at ultra-low bitrates where intelligibility is the primary objective. At such extreme compression levels, codecs trained with acoustic reconstruction losses tend to allocate bits to perceptual detail, leading to substantial degradation in word error rate (WER). This paper proposes ClariCodec, a neural speech codec operating at 300 bit per second (bps) that reformulates quantisation as a stochastic policy, enabling reinforcement learning (RL)-based optimisation of intelligibility. Specifically, the encoder is fine-tuned using WER-driven rewards while the acoustic reconstruction pipeline remains frozen. Even without RL, ClariCodec achieves 4.64% WER on the LibriSpeech test-clean set at 300 bps, already competitive with codecs operating at higher bitrates. Further RL fine-tuning reduces WER to 3.55% on test-clean and 10.4% on test-other, corresponding to a 23% relative reduction while preserving perceptual quality.


【7】Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training

标题:用于高效音频大语言模型训练的异类感知数据集调度
链接:https://arxiv.org/abs/2605.19101
作者:Yanru Wu,Jianning Wang,Chongxin Gan,Yang Li
摘要:跨不同数据集训练通用音频大语言模型(ALLM)对于整体音频理解至关重要,但由于数据集的异质性,它面临着重大挑战,这通常会导致梯度冲突和收敛缓慢。尽管它的影响,如何明确地管理这种异质性在培训过程中仍然没有得到充分的探索,目前的做法主要依赖于统一的混合。在这项工作中,我们从收敛的角度分析了多数据集AudioQA训练,并提出了分组顺序训练(GST)。GST策略性地将数据集组织成具有亲和力的组,并通过渐进式调度协议引入它们,有效地平衡了并行训练的稳定性和顺序优化的效率。为了确保可扩展性,我们开发了基于梯度的亲和度度量,捕获数据集间的关系,而无需高昂的经验可转移性估计成本。对涵盖语音、音乐和环境声音的14个AudioQA数据集的广泛评估表明,GST比标准并行训练快30- 40%,同时保持甚至超过混合训练的性能。我们的研究结果提供了有效的大规模ALLM优化的理论见解和实用的,模型无关的框架。
摘要:Training general-purpose Audio Large Language Models (ALLMs) across diverse datasets is essential for holistic audio understanding, yet it faces significant challenges due to dataset heterogeneity, which often leads to conflicting gradients and slow convergence. Despite its impact, how to explicitly manage this heterogeneity during training remains underexplored, with current practices relying primarily on uniform mixture. In this work, we analyze multi-dataset AudioQA training from a convergence perspective and propose Grouped Sequential Training (GST). GST strategically organizes datasets into affinity-aware groups and introduces them via a progressive scheduling protocol, effectively balancing the stability of parallel training with the efficiency of sequential optimization. To ensure scalability, we develop gradient-based affinity metrics that capture inter-dataset relationships without the prohibitive cost of empirical transferability estimation. Extensive evaluations on 14 AudioQA datasets spanning speech, music, and environmental sounds demonstrate that GST achieves 30--40\% faster convergence than standard parallel training while maintaining or even surpassing the performance of mix-all training. Our results provide both theoretical insights and a practical, model-agnostic framework for efficient large-scale ALLM optimization.


【8】CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation

标题:CounterFlow:反事实视频Foley生成的两阶段推理时间采样
链接:https://arxiv.org/abs/2605.18916
作者:Gyubin Lee,Junwon Lee,Juhan Nam
备注:accepted to CVPR 2026 Workshop on Sight and Sound
摘要:我们调查反事实视频福利一代,其目的是通过一个声源的身份,矛盾的视觉证据,同时保持时间同步到一个无声的视频。现有的视频和文本到音频(VT 2A)模型与此斗争,当视频和文本内容不一致时,通常保持锚定到视觉暗示的声源。我们提出了ConterFlow,这是一种用于预训练流匹配VT 2A模型的推理时间双相采样方案。第一阶段建立了一个视频衍生的时间结构,同时抑制视觉暗示的来源;第二阶段放弃视频调节,完全专注于塑造音频音色的目标提示。与天真的负面提示和最先进的基线相比,ConterFlow大大提高了反事实视频Foley生成。为了评估替换质量,我们提出了一个利用文本-音频共嵌入空间来测量目标提示证据和残余视觉暗示源泄漏的度量。视频演示和代码可在https://gyubin-lee.github.io/counterflow-demo/上获得
摘要:We investigate Counterfactual Video Foley Generation, which aims to adopt a sound-source identity that contradicts the visual evidence while remaining temporally synchronized to a silent video. Existing Video&Text-to-Audio (VT2A) models struggle with this, often remaining anchored to the visually implied sound source when video and text contents disagree. We present ConterFlow, an inference-time dual-phase sampling scheme for pretrained flow-matching VT2A models. Phase 1 builds a video-derived temporal structure while suppressing the visually implied source; Phase 2 drops video conditioning to focus entirely on shaping audio timbre toward the target prompt. ConterFlow substantially improves counterfactual Video Foley generation compared to naive negative prompting and state-of-the-art baselines. To evaluate replacement quality, we propose a metric leveraging a text-audio co-embedding space to measure both target-prompt evidence and residual visually implied source leakage. Video demonstrations and code are available at https://gyubin-lee.github.io/counterflow-demo/


【9】Cross-Talk Speech Reduction, by Separation, for Separation

标题:通过分离来减少串话语音,用于分离
链接:https://arxiv.org/abs/2605.19695
作者:Zhong-Qiu Wang,Samuele Cornell
备注:in submission
摘要:在会话语音分离和识别任务中,除了使用远场麦克风来记录远场混合信号之外,在训练数据收集期间,近距离通话麦克风通常被附接到每个说话者以捕获近场、近距离通话混合信号。每个这样的近距离谈话混合物对于佩戴者表现出相当高的能量水平,并且可以直观地用作直接在真实记录的远场信号上训练远场语音分离模型的弱监督。然而,它们对于这个目的来说不够干净,因为它们除了背景噪声之外还经常包含来自其他说话者的强串扰语音。为了解决这个问题,我们提出了减少串扰(CTR),这是一项旨在将佩戴者的语音从每个近距离说话混合物中分离出来的任务,以及一种名为CTRnet的新方法,该方法可以直接在近距离说话和远场混合物的真实记录对上进行训练,以实现CTR。在CTRnet的基础上,我们进一步提出了基于伪标签的远场语音分离(PuLSS),它使用CTRnet的估计干净的语音作为伪标签来训练分离远场混合的模型。所提出的框架的一个关键优势是,CTRnet和PuLSS都可以在目标域的真实记录数据上进行训练,解决了模型仅在模拟数据上训练时通常观察到的泛化差距。在CHiME-6数据集上,我们的框架在oracle和估计的说话者日志化下都实现了最先进的ASR性能,超过了所有CHiME-{7,8}挑战提交。据我们所知,这是第一个神经语音分离方法,大大优于真正的会话“语音在野外”数据的指导源分离。
摘要:In conversational speech separation and recognition tasks, close-talk microphones are typically attached to each speaker during training data collection to capture near-field, close-talk mixture signals, in addition to using far-field microphones to record far-field mixture signals. Each such close-talk mixture exhibits a reasonably high energy level for the wearer and could intuitively serve as weak supervision for training far-field speech separation models directly on real-recorded far-field signals. However, they are not sufficiently clean for this purpose, as they often contain strong cross-talk speech from other speakers in addition to background noise. To address this, we propose cross-talk reduction (CTR), a task aiming to isolate the wearer's speech from each close-talk mixture, and a novel method called CTRnet, which can be trained directly on real-recorded pairs of close-talk and far-field mixtures to accomplish CTR. Building on CTRnet, we further propose pseudo-label based far-field speech separation (PuLSS), which uses CTRnet's estimated clean speech as pseudo-labels to train models for separating far-field mixtures. A key advantage of the proposed framework is that both CTRnet and PuLSS can be trained on real-recorded data from the target domain, addressing the generalization gap commonly observed when models are trained exclusively on simulated data. On the CHiME-6 dataset, our framework achieves state-of-the-art ASR performance under both oracle and estimated speaker diarization, surpassing all CHiME-{7,8} challenge submissions. To our knowledge, it is the first neural speech separation method that substantially outperforms guided source separation on real conversational "speech-in-the-wild" data.


【10】A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models

标题:将音频超分辨率和带宽扩展从区分模型推进到生成模型的综述
链接:https://arxiv.org/abs/2605.16681
作者:Ningyuan Yang,Yize Li,Diego A. Cuji,Ryan M. Corey,Pu Zhao,Xue Lin,Andrew C. Singer
备注:Under review
摘要:音频超分辨率(SR),也称为带宽扩展(BWE),旨在从低分辨率(LR)或带限(BL)观测中重建高保真信号,这是一项由于丢失高频(HF)内容的模糊性而固有的不适定任务。该调查提供了该领域的全面概述,特别关注从判别映射到现代生成建模的范式转变。我们首先回顾早期的判别式深度神经网络(DNN)模型,这些模型将BWE/SR公式化为确定性映射问题,并且易于回归到均值效应和频谱过度平滑。然后,我们系统地回顾了生成方法,包括自回归(AR)模型,变分自编码器(VAE),生成对抗网络(GANs),基于扩散和分数的模型,基于流的方法和薛定谔桥。在这些方法中,我们研究了关键的设计方面,包括表示域,架构,调节机制,以及重建保真度,感知质量,鲁棒性和计算效率之间的权衡。此外,我们还讨论了涉及大型语言模型(LLM)和多模态基础模型的新兴方向,并强调了感知评估、相位建模和现实世界泛化方面的开放性挑战。通过提供结构化的分类和统一的视角,本调查建立了一个全面的基础,并为推进BWE/SR从确定性点估计到分布感知生成建模提供了实用的路线图。
摘要:Audio super-resolution (SR), also referred to as bandwidth extension (BWE), aims to reconstruct high-fidelity signals from low-resolution (LR) or band-limited (BL) observations, an inherently ill-posed task due to the ambiguity of missing high-frequency (HF) content. This survey provides a comprehensive overview of the field, with a particular focus on the paradigm shift from discriminative mapping to modern generative modeling. We first review early discriminative deep neural network (DNN) models, which formulate BWE/SR as a deterministic mapping problem and are prone to regression-to-the-mean effects and spectral over-smoothing. We then systematically review generative approaches, including autoregressive (AR) models, variational autoencoders (VAEs), generative adversarial networks (GANs), diffusion and score-based models, flow-based methods, and Schrödinger bridges. Across these approaches, we examine key design aspects, including representation domain, architecture, conditioning mechanisms, and trade-offs among reconstruction fidelity, perceptual quality, robustness, and computational efficiency. Furthermore, we discuss emerging directions involving large language models (LLMs) and multimodal foundation models, and highlight open challenges in perceptual evaluation, phase modeling, and real-world generalization. By providing a structured taxonomy and unified perspective, this survey establishes a comprehensive foundation and offers a practical roadmap for advancing BWE/SR from deterministic point estimation toward distribution-aware generative modeling.


eess.AS音频处理


【1】Cross-Talk Speech Reduction, by Separation, for Separation
标题:通过分离来减少串话语音,用于分离
链接:https://arxiv.org/abs/2605.19695
作者:Zhong-Qiu Wang,Samuele Cornell
备注:in submission
摘要:在会话语音分离和识别任务中,除了使用远场麦克风来记录远场混合信号之外,在训练数据收集期间,近距离通话麦克风通常被附接到每个说话者以捕获近场、近距离通话混合信号。每个这样的近距离谈话混合物对于佩戴者表现出相当高的能量水平,并且可以直观地用作直接在真实记录的远场信号上训练远场语音分离模型的弱监督。然而,它们对于这个目的来说不够干净,因为它们除了背景噪声之外还经常包含来自其他说话者的强串扰语音。为了解决这个问题,我们提出了减少串扰(CTR),这是一项旨在将佩戴者的语音从每个近距离说话混合物中分离出来的任务,以及一种名为CTRnet的新方法,该方法可以直接在近距离说话和远场混合物的真实记录对上进行训练,以实现CTR。在CTRnet的基础上,我们进一步提出了基于伪标签的远场语音分离(PuLSS),它使用CTRnet的估计干净的语音作为伪标签来训练分离远场混合的模型。所提出的框架的一个关键优势是,CTRnet和PuLSS都可以在目标域的真实记录数据上进行训练,解决了模型仅在模拟数据上训练时通常观察到的泛化差距。在CHiME-6数据集上,我们的框架在oracle和估计的说话者日志化下都实现了最先进的ASR性能,超过了所有CHiME-{7,8}挑战提交。据我们所知,这是第一个神经语音分离方法,大大优于真正的会话“语音在野外”数据的指导源分离。
摘要:In conversational speech separation and recognition tasks, close-talk microphones are typically attached to each speaker during training data collection to capture near-field, close-talk mixture signals, in addition to using far-field microphones to record far-field mixture signals. Each such close-talk mixture exhibits a reasonably high energy level for the wearer and could intuitively serve as weak supervision for training far-field speech separation models directly on real-recorded far-field signals. However, they are not sufficiently clean for this purpose, as they often contain strong cross-talk speech from other speakers in addition to background noise. To address this, we propose cross-talk reduction (CTR), a task aiming to isolate the wearer's speech from each close-talk mixture, and a novel method called CTRnet, which can be trained directly on real-recorded pairs of close-talk and far-field mixtures to accomplish CTR. Building on CTRnet, we further propose pseudo-label based far-field speech separation (PuLSS), which uses CTRnet's estimated clean speech as pseudo-labels to train models for separating far-field mixtures. A key advantage of the proposed framework is that both CTRnet and PuLSS can be trained on real-recorded data from the target domain, addressing the generalization gap commonly observed when models are trained exclusively on simulated data. On the CHiME-6 dataset, our framework achieves state-of-the-art ASR performance under both oracle and estimated speaker diarization, surpassing all CHiME-{7,8} challenge submissions. To our knowledge, it is the first neural speech separation method that substantially outperforms guided source separation on real conversational "speech-in-the-wild" data.


【2】Fast Multichannel NMF with Block-Diagonal Spatial Covariance Matrices for Efficient Blind Source Separation Using Distributed Microphone Arrays

标题:基于块对角空间协方差矩阵的快速多通道NMF算法及其在分布式麦克风阵列盲分离中的应用
链接:https://arxiv.org/abs/2605.19388
作者:Hirotaka Nishikori,Nobutaka Ito,Kouei Yamaoka,Norihiro Takamune,Hiroshi Saruwatari
摘要:由多个子阵组成的分布式麦克风阵列可以在很宽的空间范围内实现盲源分离。直接将快速多通道非负矩阵分解(FastMNMF)应用于所有子阵列可以利用来自所有子阵列的观察结果,但它需要对跨越所有麦克风的大型矩阵进行重复求逆,从而导致计算成本随着麦克风数量的增加而迅速增加。相比之下,将FastMNMF应用于一个子阵列减小了矩阵大小,但不能利用来自其他子阵列的观察。我们提出分布式FastMNMF,它对源空间协方差矩阵施加块对角结构,以便在子阵列内执行矩阵求逆。基于NMF的源谱图模型在子阵列之间共享,允许该方法聚合源活动信息,同时丢弃子阵列间协方差。在同步,无噪声的模拟与固定的房间和阵列/源的几何形状,该方法需要更少的计算时间比传统的FastMNMF使用所有的子阵列,实现了更高的平均源失真比比传统的FastMNMF使用一个子阵列,并适用于在测试的五个源的条件下,每个四个麦克风子阵列是局部欠定。
摘要:Distributed microphone arrays composed of multiple subarrays enable blind source separation over a wide spatial area. Directly applying fast multichannel nonnegative matrix factorization (FastMNMF) to all subarrays can exploit observations from all subarrays, but it requires repeated inversions of large matrices spanning all microphones, causing the computational cost to increase rapidly as the number of microphones grows. In contrast, applying FastMNMF to one subarray reduces the matrix size but cannot exploit observations from other subarrays. We propose distributed FastMNMF, which imposes a block-diagonal structure on the source spatial covariance matrices, so that matrix inversions are performed within subarrays. The NMF-based source spectrogram model is shared across subarrays, allowing the method to aggregate source activity information while discarding inter-subarray covariance. In synchronized, noiseless simulations with fixed room and array/source geometry, the method required less computation time than conventional FastMNMF using all subarrays, achieved a higher average source-to-distortion ratio than conventional FastMNMF using one subarray, and was applicable in the tested five-source condition, where each four-microphone subarray was locally underdetermined.


【3】Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

标题:Mega-ASR:通过放大真实世界声学模拟实现野外语音识别
链接:https://arxiv.org/abs/2605.19833
作者:Zhifei Xie,Kaiyu Pang,Haobin Zhang,Deheng Ye,Xiaobin Hu,Shuicheng Yan,Chunyan Miao
备注:Project page: https://xzf-thu.github.io/Mega-ASR/. Code, models, and dataset will be released. A robust ASR framework targeting in-the-wild and compositional acoustic scenarios where conventional ASR systems fail
摘要:尽管自动语音识别(ASR)和大型音频语言模型的快速发展,但现实环境中的鲁棒性识别仍然受到“声学鲁棒性瓶颈”的限制:模型通常会失去声学基础,并在严重的合成失真下产生遗漏或幻觉。我们提出了Mega-ASR,一个统一的ASR-in-the-wild框架,它将可扩展的复合数据构建与渐进的声学到语义优化相结合。我们引入Voices-in-the-Wild-2 M,涵盖7个经典声学现象和54个物理上合理的复合场景,并使用声学到语义渐进式监督微调和双粒度WER-Gated策略优化训练Mega-ASR。大量实验表明,Mega-ASR在不利条件下的ASR基准测试中比现有技术系统具有显著优势(VOiCES R4-B-F上为45.69% vs. 54.01%,NOIZEUS Sta-0上为21.49% vs. 29.34%)。在复杂的合成声学场景中,Mega-ASR进一步提供了超过30%的相对WER降低,相对于强大的开源和闭源基线,建立了一个可扩展的范例,用于在野外实现强大的ASR。
摘要:Despite rapid advances in automatic speech recognition (ASR) and large audio-language models, robust recognition in real-world environments remains limited by an "acoustic robustness bottleneck": models often lose acoustic grounding and produce omissions or hallucinations under severe, compositional distortions. We propose Mega-ASR, a unified ASR-in-the-wild framework that combines scalable compound-data construction with progressive acoustic-to-semantic optimization. We introduce Voices-in-the-Wild-2M, covering 7 classic acoustic phenomena and 54 physically plausible compound scenarios, and train Mega-ASR with Acoustic-to-Semantic Progressive Supervised Fine-Tuning and Dual-Granularity WER-Gated Policy Optimization. Extensive experiments demonstrate that Mega-ASR achieves significant advantages over prior state-of-the-art systems on adverse-condition ASR benchmarks (45.69% vs. 54.01% on VOiCES R4-B-F, and 21.49% vs. 29.34% on NOIZEUS Sta-0). On complex compositional acoustic scenarios, Mega-ASR further delivers over 30% relative WER reduction against strong open- and closed-source baselines, establishing a scalable paradigm for robust ASR in-the-wild.


【4】CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation

标题:CounterFlow:反事实视频Foley生成的两阶段推理时间采样
链接:https://arxiv.org/abs/2605.18916
作者:Gyubin Lee,Junwon Lee,Juhan Nam
备注:accepted to CVPR 2026 Workshop on Sight and Sound
摘要:我们调查反事实视频福利一代,其目的是通过一个声源的身份,矛盾的视觉证据,同时保持时间同步到一个无声的视频。现有的视频和文本到音频(VT 2A)模型与此斗争,当视频和文本内容不一致时,通常保持锚定到视觉暗示的声源。我们提出了ConterFlow,这是一种用于预训练流匹配VT 2A模型的推理时间双相采样方案。第一阶段建立了一个视频衍生的时间结构,同时抑制视觉暗示的来源;第二阶段放弃视频调节,完全专注于塑造音频音色的目标提示。与天真的负面提示和最先进的基线相比,ConterFlow大大提高了反事实视频Foley生成。为了评估替换质量,我们提出了一个利用文本-音频共嵌入空间来测量目标提示证据和残余视觉暗示源泄漏的度量。视频演示和代码可在https://gyubin-lee.github.io/counterflow-demo/上获得
摘要:We investigate Counterfactual Video Foley Generation, which aims to adopt a sound-source identity that contradicts the visual evidence while remaining temporally synchronized to a silent video. Existing Video&Text-to-Audio (VT2A) models struggle with this, often remaining anchored to the visually implied sound source when video and text contents disagree. We present ConterFlow, an inference-time dual-phase sampling scheme for pretrained flow-matching VT2A models. Phase 1 builds a video-derived temporal structure while suppressing the visually implied source; Phase 2 drops video conditioning to focus entirely on shaping audio timbre toward the target prompt. ConterFlow substantially improves counterfactual Video Foley generation compared to naive negative prompting and state-of-the-art baselines. To evaluate replacement quality, we propose a metric leveraging a text-audio co-embedding space to measure both target-prompt evidence and residual visually implied source leakage. Video demonstrations and code are available at https://gyubin-lee.github.io/counterflow-demo/


机器翻译由腾讯交互翻译提供,仅供参考