微信公众号:arXiv_Daily
cs.SD语音
【1】Localizing Speech Deepfakes Beyond Transitions via Segment-Aware Learning
标题:通过分段感知学习本地化语音深度造假超越转型
链接:https://arxiv.org/abs/2601.21925
摘要:由于这些修改的微妙和分散的性质,定位部分deepfake音频(其中只有语音片段被操纵)仍然具有挑战性。现有的方法通常依赖于帧级预测来识别欺骗片段,并且最近的一些方法通过专注于真实和虚假音频之间的过渡来提高性能。然而,我们观察到,这些模型往往过度依赖于边界工件,而忽略了随后的操纵内容。我们认为,有效的本地化需要了解整个段不仅仅是检测转换。因此,我们提出了段感知学习(SAL),一个框架,鼓励模型专注于段的内部结构。SAL引入了两种核心技术:段位置标记,它提供基于段内相对位置的细粒度帧监督;跨段混合,一种生成不同段模式的数据增强方法。在多个deepfake本地化数据集上的实验表明,SAL在域内和域外设置中始终实现了强大的性能,在非边界区域中具有显着的增益,并减少了对过渡伪影的依赖。该代码可在https://github.com/SentryMao/SAL上获得。
摘要:Localizing partial deepfake audio, where only segments of speech are manipulated, remains challenging due to the subtle and scattered nature of these modifications. Existing approaches typically rely on frame-level predictions to identify spoofed segments, and some recent methods improve performance by concentrating on the transitions between real and fake audio. However, we observe that these models tend to over-rely on boundary artifacts while neglecting the manipulated content that follows. We argue that effective localization requires understanding the entire segments beyond just detecting transitions. Thus, we propose Segment-Aware Learning (SAL), a framework that encourages models to focus on the internal structure of segments. SAL introduces two core techniques: Segment Positional Labeling, which provides fine-grained frame supervision based on relative position within a segment; and Cross-Segment Mixing, a data augmentation method that generates diverse segment patterns. Experiments across multiple deepfake localization datasets show that SAL consistently achieves strong performance in both in-domain and out-of-domain settings, with notable gains in non-boundary regions and reduced reliance on transition artifacts. The code is available at https://github.com/SentryMao/SAL.
【2】MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding
标题:MIDI-LLaMA:一门遵循指令的多模式LLM,用于理解象征性音乐
链接:https://arxiv.org/abs/2601.21740
备注:Accepted for publication at International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026
摘要:音频音乐的多模态大语言模型(MLLM)的最新进展已经证明了音乐理解的强大能力,但符号音乐,音乐结构的基本表示,仍然是未开发的。在这项工作中,我们介绍MIDI-LLaMA,第一个解释以下MLLM符号音乐理解。我们的方法通过一个两阶段的流水线,包括功能对齐和指令调整,对齐的编码器MusicBERT和Llama-3-8B。为了支持训练,我们设计了一个可扩展的注释管道,用细粒度的元数据注释GiantMIDI-Piano,从而生成MIDI文本数据集。与在相同的注释调整过程下将注释转换为ABC符号的基线训练相比,MIDI-LLaMA在问题回答中的标题和语义对齐方面表现明显优于基线。人类评估进一步证实了MIDI-LLaMA在音乐理解,情感识别,创造力和整体偏好方面的优势。这些发现表明,将符号音乐纳入大型语言模型可以增强他们对音乐的理解能力。
摘要:Recent advances in multimodal large language models (MLLM) for audio music have demonstrated strong capabilities in music understanding, yet symbolic music, a fundamental representation of musical structure, remains unexplored. In this work, we introduce MIDI-LLaMA, the first instruction-following MLLM for symbolic music understanding. Our approach aligns the MIDI encoder MusicBERT and Llama-3-8B via a two-stage pipeline comprising feature alignment and instruction tuning. To support training, we design a scalable annotation pipeline that annotates GiantMIDI-Piano with fine-grained metadata, resulting in a MIDI-text dataset. Compared with the baseline trained on converting MIDI into ABC notation under the same instruction-tuning procedure, MIDI-LLaMA substantially outperforms in captioning and semantic alignment in question answering. Human evaluation further confirms the advantages of MIDI-LLaMA in music understanding, emotion recognition, creativity, and overall preference. These findings demonstrate that incorporating symbolic music into large language models enhances their capacity for musical understanding.
【3】Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs
标题:通过优先增强的音频LLM统一语音编辑检测和内容定位
链接:https://arxiv.org/abs/2601.21463
摘要:语音编辑通过对原始话语进行细粒度的分段级操作来实现语义反转,同时保持全局感知自然度。现有的检测研究主要集中在手动编辑的语音与显式拼接文物,因此,努力应付新兴的端到端的神经语音编辑技术,产生无缝的声学过渡。为了应对这一挑战,我们首先构建了一个大规模双语数据集AiEdit,该数据集利用大型语言模型驱动精确的语义篡改逻辑,并采用多种先进的神经语音编辑方法进行数据合成,从而填补了高质量语音编辑数据集的空白。在此基础上,我们提出了PELM(先前增强的音频大语言模型),这是第一个大模型框架,通过将语音编辑检测和内容本地化统一为音频问答任务。为了减轻现有音频大模型中观察到的固有伪造偏差和语义优先级偏差,PELM结合单词级概率先验来提供明确的声学线索,并进一步设计了基于质心聚合的声学一致性感知损失来明确地执行细微的局部分布异常的建模。大量的实验结果表明,PELM显着优于国家的最先进的方法在HumanEdit和AiEdit数据集,实现等错误率(EER)分别为0.57%和9.28%(本地化)。
摘要:Speech editing achieves semantic inversion by performing fine-grained segment-level manipulation on original utterances, while preserving global perceptual naturalness. Existing detection studies mainly focus on manually edited speech with explicit splicing artifacts, and therefore struggle to cope with emerging end-to-end neural speech editing techniques that generate seamless acoustic transitions. To address this challenge, we first construct a large-scale bilingual dataset, AiEdit, which leverages large language models to drive precise semantic tampering logic and employs multiple advanced neural speech editing methods for data synthesis, thereby filling the gap of high-quality speech editing datasets. Building upon this foundation, we propose PELM (Prior-Enhanced Audio Large Language Model), the first large-model framework that unifies speech editing detection and content localization by formulating them as an audio question answering task. To mitigate the inherent forgery bias and semantic-priority bias observed in existing audio large models, PELM incorporates word-level probability priors to provide explicit acoustic cues, and further designs a centroid-aggregation-based acoustic consistency perception loss to explicitly enforce the modeling of subtle local distribution anomalies. Extensive experimental results demonstrate that PELM significantly outperforms state-of-the-art methods on both the HumanEdit and AiEdit datasets, achieving equal error rates (EER) of 0.57\% and 9.28\% (localization), respectively.
【4】Understanding Frechet Speech Distance for Synthetic Speech Quality Evaluation
标题:了解弗雷切特语音距离以进行合成语音质量评估
链接:https://arxiv.org/abs/2601.21386
备注:accepted to ICASSP 2026
摘要:合成语音质量的客观评价仍然是一个关键的挑战。人类听力测试是黄金标准,但成本高昂,规模不切实际。Fréchet距离已经成为一个很有前途的替代方案,但它的可靠性在很大程度上取决于嵌入和实验设置的选择。在这项工作中,我们全面评估Fréchet语音距离(FSD)和它的变体语音最大平均离散(SMMD)在不同的嵌入和条件下。我们进一步将人类听力评估与TTS可懂度和合成训练的ASR WER结合起来,以验证这些指标的感知相关性。我们的研究结果表明,WavLM Base+功能与人类评级的一致性最稳定。虽然FSD和SMMD不能完全取代主观评价,但我们表明,它们可以作为补充,具有成本效益和可重复的措施,特别是在大规模或直接听力评估不可行时。代码可在https://github.com/kaen2891/FrechetSpeechDistance上获得。
摘要:Objective evaluation of synthetic speech quality remains a critical challenge. Human listening tests are the gold standard, but costly and impractical at scale. Fréchet Distance has emerged as a promising alternative, yet its reliability depends heavily on the choice of embeddings and experimental settings. In this work, we comprehensively evaluate Fréchet Speech Distance (FSD) and its variant Speech Maximum Mean Discrepancy (SMMD) under varied embeddings and conditions. We further incorporate human listening evaluations alongside TTS intelligibility and synthetic-trained ASR WER to validate the perceptual relevance of these metrics. Our findings show that WavLM Base+ features yield the most stable alignment with human ratings. While FSD and SMMD cannot fully replace subjective evaluation, we show that they can serve as complementary, cost-efficient, and reproducible measures, particularly useful when large-scale or direct listening assessments are infeasible. Code is available at https://github.com/kaen2891/FrechetSpeechDistance.
【5】Qwen3-ASR Technical Report
标题:Qwen 3-ASR技术报告
链接:https://arxiv.org/abs/2601.21337
备注:https://github.com/QwenLM/Qwen3-ASR
摘要:在这份报告中,我们介绍了Qwen 3-ASR家族,其中包括两个强大的一体化语音识别模型和一个新的非自回归语音强制对齐模型。Qwen 3-ASR-1.7B和Qwen 3-ASR-0.6B是ASR模型,支持52种语言和方言的语言识别和ASR。它们都利用了大规模的语音训练数据和基础模型Qwen 3-Omni强大的音频理解能力。除了开源基准之外,我们还进行了全面的内部评估,因为ASR模型在开源基准得分上可能差异不大,但在真实场景中表现出显着的质量差异。实验表明,1.7B版本在开源ASR模型中实现了SOTA性能,并与最强的专有API竞争,而0.6B版本提供了最佳的准确性-效率权衡。Qwen 3-ASR-0.6B可以实现平均TTFT低至92 ms,并在128并发的情况下在1秒内转录2000秒的语音。Qwen 3-ForcedAligner-0.6B是一个基于LLM的NAR时间戳预测器,能够对齐11种语言的文本-语音对。时间戳精度实验表明,该模型优于三种最强的力对齐模型,在效率和通用性方面具有更多优势。为了进一步加速ASR和音频理解的社区研究,我们在Apache 2.0许可下发布了这些模型。
摘要:In this report, we introduce Qwen3-ASR family, which includes two powerful all-in-one speech recognition models and a novel non-autoregressive speech forced alignment model. Qwen3-ASR-1.7B and Qwen3-ASR-0.6B are ASR models that support language identification and ASR for 52 languages and dialects. Both of them leverage large-scale speech training data and the strong audio understanding ability of their foundation model Qwen3-Omni. We conduct comprehensive internal evaluation besides the open-sourced benchmarks as ASR models might differ little on open-sourced benchmark scores but exhibit significant quality differences in real-world scenarios. The experiments reveal that the 1.7B version achieves SOTA performance among open-sourced ASR models and is competitive with the strongest proprietary APIs while the 0.6B version offers the best accuracy-efficiency trade-off. Qwen3-ASR-0.6B can achieve an average TTFT as low as 92ms and transcribe 2000 seconds speech in 1 second at a concurrency of 128. Qwen3-ForcedAligner-0.6B is an LLM based NAR timestamp predictor that is able to align text-speech pairs in 11 languages. Timestamp accuracy experiments show that the proposed model outperforms the three strongest force alignment models and takes more advantages in efficiency and versatility. To further accelerate the community research of ASR and audio understanding, we release these models under the Apache 2.0 license.
【6】Evaluating Spatialized Auditory Cues for Rapid Attention Capture in XR
标题:评估XR中快速注意力捕获的空间化听觉线索
链接:https://arxiv.org/abs/2601.21264
备注:8 pages, 4 figures. This is the author's version of the article that will appear at the IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (IEEE VRW) 2026
摘要:在时间关键延展实境(XR)场景中,用户必须在从事主要任务时迅速将注意力重新定位到危险,警报或指令,空间音频可以提供即时的方向提示,而不会占用视觉带宽。然而,这样的场景只能提供短暂的听觉暴露,需要用户快速解释声音方向,而无需延长收听或头部驱动的细化。本文报告了一个控制的探索性研究,快速空间音频定位在XR。使用HRTF呈现的宽带刺激从一个半密集的方向围绕听众,我们量化如何准确地用户可以推断出粗糙的方向,从简短的音频单独。我们进一步研究了短期视听反馈训练作为轻量级校准机制的效果。我们的研究结果表明,简短的空间线索可以传达粗糙的方向信息,即使是短校准可以提高用户的听觉信号的感知。虽然这些结果强调了空间音频在快速注意力引导方面的潜力,但它们也表明,单独的听觉提示可能无法为复杂或高风险的任务提供足够的精度,并且空间音频在与其他感觉方式或视觉提示补充时可能是最有效的,而不依赖于头部驱动的细化。我们利用这项关于空间音频的研究作为对可穿戴XR的第一阶段注意力引导通道的初步调查(例如,VR头戴式显示器和AR智能眼镜),并提供有关刺激选择和校准的设计见解,用于时间关键型用途。
摘要:In time-critical eXtended reality (XR) scenarios where users must rapidly reorient their attention to hazards, alerts, or instructions while engaged in a primary task, spatial audio can provide an immediate directional cue without occupying visual bandwidth. However, such scenarios can afford only a brief auditory exposure, requiring users to interpret sound direction quickly and without extended listening or head-driven refinement. This paper reports a controlled exploratory study of rapid spatial-audio localization in XR. Using HRTF-rendered broadband stimuli presented from a semi-dense set of directions around the listener, we quantify how accurately users can infer coarse direction from brief audio alone. We further examine the effects of short-term visuo-auditory feedback training as a lightweight calibration mechanism. Our findings show that brief spatial cues can convey coarse directional information, and that even short calibration can improve users' perception of aural signals. While these results highlight the potential of spatial audio for rapid attention guidance, they also show that auditory cues alone may not provide sufficient precision for complex or high-stakes tasks, and that spatial audio may be most effective when complemented by other sensory modalities or visual cues, without relying on head-driven refinement. We leverage this study on spatial audio as a preliminary investigation into a first-stage attention-guidance channel for wearable XR (e.g., VR head-mounted displays and AR smart glasses), and provide design insights on stimulus selection and calibration for time-critical use.
【7】Music Plagiarism Detection: Problem Formulation and a Segment-based Solution
标题:音乐抄袭检测:问题制定和基于片段的解决方案
链接:https://arxiv.org/abs/2601.21260
摘要:近年来,音乐剽窃问题已成为一个更加紧迫的社会问题。随着音乐信息检索研究的进展,人们越来越努力地解决与音乐剽窃相关的问题。然而,许多研究,包括我们以前的工作,进行了研究,没有明确定义音乐剽窃检测任务实际上涉及什么。这种缺乏明确定义的情况减缓了研究进展,并使其难以将结果应用于现实世界的场景。为了解决这种情况,我们定义了音乐剽窃检测与其他MIR任务的不同之处,并解释了需要解决的问题。我们引入了相似音乐对数据集来支持这个新定义的任务。此外,我们提出了一种基于片段转录的方法来解决这个问题。我们的演示和数据集可以在https://github.com/Mippia/ICASSP2026-MPD上找到。
摘要:Recently, the problem of music plagiarism has emerged as an even more pressing social issue. As music information retrieval research advances, there is a growing effort to address issues related to music plagiarism. However, many studies, including our previous work, have conducted research without clearly defining what the music plagiarism detection task actually involves. This lack of a clear definition has slowed research progress and made it hard to apply results to real-world scenarios. To fix this situation, we defined how Music Plagiarism Detection is different from other MIR tasks and explained what problems need to be solved. We introduce the Similar Music Pair dataset to support this newly defined task. In addition, we propose a method based on segment transcription as one way to solve the task. Our demo and dataset are available at https://github.com/Mippia/ICASSP2026-MPD.
【8】Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
标题:使用通用电话识别和语音特定音素对比建模的多语言发音障碍言语评估
链接:https://arxiv.org/abs/2601.21205
备注:10 pages, 4 figures
摘要:与构音障碍相关的神经系统疾病的日益普遍,激发了对适用于各种语言的自动可懂度评估方法的需求。然而,大多数现有的方法要么局限于一种语言或未能捕捉语言特定的因素塑造可理解性。我们提出了一个多语种的音素生产评估框架,集成了通用的电话识别与语言特定的音素解释,使用对比的音位特征距离的电话到音素映射和序列对齐。该框架产生三个指标:音素错误率(PER),语音特征错误率(PFER),和一个新提出的无干扰措施,音素覆盖率(PhonCov)。对英语、西班牙语、意大利语和泰米尔语的分析表明,PER受益于映射和对齐的组合,PFER受益于单独的对齐,PhonCov受益于映射。进一步的分析表明,所提出的框架捕捉临床上有意义的模式,与构音障碍语音的既定观察一致的清晰度下降。
摘要:The growing prevalence of neurological disorders associated with dysarthria motivates the need for automated intelligibility assessment methods that are applicalbe across languages. However, most existing approaches are either limited to a single language or fail to capture language-specific factors shaping intelligibility. We present a multilingual phoneme-production assessment framework that integrates universal phone recognition with language-specific phoneme interpretation using contrastive phonological feature distances for phone-to-phoneme mapping and sequence alignment. The framework yields three metrics: phoneme error rate (PER), phonological feature error rate (PFER), and a newly proposed alignment-free measure, phoneme coverage (PhonCov). Analysis on English, Spanish, Italian, and Tamil show that PER benefits from the combination of mapping and alignment, PFER from alignment alone, and PhonCov from mapping. Further analyses demonstrate that the proposed framework captures clinically meaningful patterns of intelligibility degradation consistent with established observations of dysarthric speech.
【9】PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
标题:PhaseCoder:多模式LLM的麦克风几何不可知的空间音频理解
链接:https://arxiv.org/abs/2601.21124
摘要:当前的多模态LLM将音频作为单声道流处理,忽略了嵌入式AI所必需的丰富空间信息。相反,现有的空间音频模型受限于固定的麦克风几何形状,从而阻止了跨不同设备的部署。我们提出了PhaseCoder,一个仅变换器的空间音频编码器,它与麦克风几何形状无关。PhaseCoder将原始多通道音频和麦克风坐标作为输入来执行定位,并产生鲁棒的空间嵌入。我们证明了Gemma 3n LLM可以被微调到由PhaseCoder产生的“空间音频令牌”。我们展示了我们的编码器在麦克风不变的定位基准上实现了最先进的结果,并且首次使LLM能够从任意麦克风阵列执行复杂的空间推理和有针对性的转录任务。
摘要:Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained to fixed microphone geometries, preventing deployment across diverse devices. We present PhaseCoder, a transformer-only spatial audio encoder that is agnostic to microphone geometry. PhaseCoder takes raw multichannel audio and microphone coordinates as inputs to perform localization and produces robust spatial embeddings. We demonstrate that Gemma 3n LLM can be fine-tuned to reason over "Spatial Audio Tokens" produced by PhaseCoder. We show our encoder achieves state-of-the-art results on microphone-invariant localization benchmarks and, for the first time, enables an LLM to perform complex spatial reasoning and targeted transcription tasks from an arbitrary microphone array.
【10】Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
标题:具有自监督语音表示的语音增强模型的位置不变微调
链接:https://arxiv.org/abs/2601.21084
备注:Accepted to ICASSP 2026
摘要:将前端语音增强(SE)模型与基于自监督学习(SSL)的语音模型相结合,对于噪声条件下的下游任务是有效的。SE模型通常使用SSL表示进行微调,增强语音和干净语音之间的均方误差(MSE)损失。然而,MSE倾向于利用SSL模型中的位置嵌入,允许通过位置相关性而不是内容相关信息来最小化目标。这项工作框架的问题作为自我监督表示微调的一般限制,并通过代表指导SE调查。考虑了两种战略:(1)零填充,以前在SSL预训练中探索过,但在微调设置中进行了检查,以及(2)具有软DTW损失的速度扰动。实验表明,基于软DTW的方法实现了更快的收敛速度和改善下游性能,强调了位置不变的微调在基于SSL的语音建模的重要性。
摘要:Integrating front-end speech enhancement (SE) models with self-supervised learning (SSL)-based speech models is effective for downstream tasks in noisy conditions. SE models are commonly fine-tuned using SSL representations with mean squared error (MSE) loss between enhanced and clean speech. However, MSE is prone to exploiting positional embeddings in SSL models, allowing the objective to be minimised through positional correlations instead of content-related information. This work frames the problem as a general limitation of self-supervised representation fine-tuning and investigates it through representation-guided SE. Two strategies are considered: (1) zero-padding, previously explored in SSL pre-training but here examined in the fine-tuning setting, and (2) speed perturbations with a soft-DTW loss. Experiments show that the soft-DTW-based approach achieves faster convergence and improved downstream performance, underscoring the importance of position-invariant fine-tuning in SSL-based speech modelling.
【11】asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
标题:asr_eval:用于多参考和流语音识别评估的算法和工具
链接:https://arxiv.org/abs/2601.20992
摘要:我们提出了几个改进的语音识别评估。首先,我们提出了一个字符串对齐算法,支持多引用标记,任意长度的插入和更好的词对齐。这对于非拉丁语的语言特别有用,这些语言具有丰富的构词法,可以标记杂乱或冗长的语音。其次,我们收集了一个新的测试集DiverseSpeech-Ru的长形式在野生俄语语音仔细多参考标记。我们还执行多参考重新标记流行的俄罗斯测试集和研究微调动力学在其相应的训练集。我们证明了该模型往往采用特定于特定于网络的标签,造成一种错觉的度量改进。基于改进的词对齐,我们开发的工具来评估流语音识别和对齐多个transmits进行视觉比较。此外,我们为许多离线和流式语音识别模型提供了统一的包装器。我们的代码将公开提供。
摘要:We propose several improvements to the speech recognition evaluation. First, we propose a string alignment algorithm that supports both multi-reference labeling, arbitrary-length insertions and better word alignment. This is especially useful for non-Latin languages, those with rich word formation, to label cluttered or longform speech. Secondly, we collect a novel test set DiverseSpeech-Ru of longform in-the-wild Russian speech with careful multi-reference labeling. We also perform multi-reference relabeling of popular Russian tests set and study fine-tuning dynamics on its corresponding train set. We demonstrate that the model often adopts to dataset-specific labeling, causing an illusion of metric improvement. Based on the improved word alignment, we develop tools to evaluate streaming speech recognition and to align multiple transcriptions to compare them visually. Additionally, we provide uniform wrappers for many offline and streaming speech recognition models. Our code will be made publicly available.
【12】Text-only adaptation in LLM-based ASR through text denoising
标题:通过文本去噪在基于LLM的ASR中进行纯文本自适应
链接:https://arxiv.org/abs/2601.20900
备注:Paper accepted at ICASSP 2026
摘要:基于大型语言模型(LLM)的自动语音识别(ASR)系统适应使用纯文本数据的新领域是一个重要但尚未探索的挑战。LLM对目标域文本的标准微调通常会破坏投影仪学习的语音和文本模态之间的关键对齐,从而降低性能。我们引入了一种新的纯文本自适应方法,通过将其视为文本去噪任务来模拟音频投影任务。因此,我们的方法训练LLM从嘈杂的输入中恢复干净的成绩单。该过程有效地使模型适应目标域,同时保持跨模态对齐。我们的解决方案是轻量级的,不需要架构更改或其他参数。对两个数据集的广泛评估表明,相对改善高达22.1%,优于最近最先进的纯文本自适应方法。
摘要:Adapting automatic speech recognition (ASR) systems based on large language models (LLMs) to new domains using text-only data is a significant yet underexplored challenge. Standard fine-tuning of the LLM on target-domain text often disrupts the critical alignment between speech and text modalities learned by the projector, degrading performance. We introduce a novel text-only adaptation method that emulates the audio projection task by treating it as a text denoising task. Our approach thus trains the LLM to recover clean transcripts from noisy inputs. This process effectively adapts the model to a target domain while preserving cross-modal alignment. Our solution is lightweight, requiring no architectural changes or additional parameters. Extensive evaluation on two datasets demonstrates up to 22.1% relative improvement, outperforming recent state-of-the-art text-only adaptation methods.
【13】A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
标题:预训练自我监督语音模型的数据选择策略研究
链接:https://arxiv.org/abs/2601.20896
备注:Accepted for publication in the 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2026)
摘要:自监督学习(SSL)已经改变了语音处理,但其对大量预训练数据集的依赖仍然是一个瓶颈。虽然鲁棒性通常归因于规模和多样性,但数据分布的作用却鲜为人知。我们系统地研究了预训练数据的精选子集如何影响自动语音识别(ASR)性能。令人惊讶的是,优化声学,扬声器或语言多样性并没有产生明显的改善随机采样。相反,我们发现,优先考虑最长的话语可以获得更好的ASR结果,同时只使用一半的原始数据集,在大型语料库上减少了24%的预训练时间。这些发现表明,对于预训练语音SSL模型,数据长度是比数据多样性或整体数据量更关键的因素,为SSL语音处理中的数据选择策略提供了新的视角。
摘要:Self-supervised learning (SSL) has transformed speech processing, yet its reliance on massive pre-training datasets remains a bottleneck. While robustness is often attributed to scale and diversity, the role of the data distribution is less understood. We systematically examine how curated subsets of pre-training data influence Automatic Speech Recognition (ASR) performance. Surprisingly, optimizing for acoustic, speaker, or linguistic diversity yields no clear improvements over random sampling. Instead, we find that prioritizing the longest utterances achieves superior ASR results while using only half the original dataset, reducing pre-training time by 24% on a large corpora. These findings suggest that for pre-training speech SSL models, data length is a more critical factor than either data diversity or overall data quantity for performance and efficiency, offering a new perspective for data selection strategies in SSL speech processing.
【14】SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
标题:SW-ASR:一种用于鲁棒单字语音识别的上下文感知混合ASR管道
链接:https://arxiv.org/abs/2601.20890
摘要:单字自动语音识别(ASR)是一项具有挑战性的任务,因为它缺乏语言背景,对噪声、发音变化和通道伪影敏感,特别是在低资源、通信关键的领域,如医疗保健和应急响应。本文回顾了最近的深度学习方法,并提出了一个用于鲁棒单字检测的模块化框架。该系统将去噪和归一化与混合ASR前端(Whisper + Vosk)和验证层相结合,旨在处理词汇表外的单词和降级的音频。验证层支持多种匹配策略,包括嵌入相似性,编辑距离和基于LLM的匹配以及可选的上下文指导。我们在Google Speech Commands数据集和从带宽受限条件下的电话和消息传递平台收集的真实数据集上评估了该框架。结果表明,虽然混合ASR前端在干净音频上表现良好,但验证层显着提高了噪声和压缩通道的准确性。上下文引导和基于LLM的匹配产生最大的收益,表明轻量级验证和上下文机制可以大大提高单字ASR的鲁棒性,而不会牺牲实时电话应用所需的延迟。
摘要:Single-word Automatic Speech Recognition (ASR) is a challenging task due to the lack of linguistic context and sensitivity to noise, pronunciation variation, and channel artifacts, especially in low-resource, communication-critical domains such as healthcare and emergency response. This paper reviews recent deep learning approaches and proposes a modular framework for robust single-word detection. The system combines denoising and normalization with a hybrid ASR front end (Whisper + Vosk) and a verification layer designed to handle out-of-vocabulary words and degraded audio. The verification layer supports multiple matching strategies, including embedding similarity, edit distance, and LLM-based matching with optional contextual guidance. We evaluate the framework on the Google Speech Commands dataset and a curated real-world dataset collected from telephony and messaging platforms under bandwidth-limited conditions. Results show that while the hybrid ASR front end performs well on clean audio, the verification layer significantly improves accuracy on noisy and compressed channels. Context-guided and LLM-based matching yield the largest gains, demonstrating that lightweight verification and context mechanisms can substantially improve single-word ASR robustness without sacrificing latency required for real-time telephony applications.
【15】VoxMorph: Scalable Zero-shot Voice Identity Morphing via Disentangled Embeddings
标题:VoxMorph:通过分解嵌入实现可扩展的Zero-Shot语音身份变形
链接:https://arxiv.org/abs/2601.20883
备注:Accepted to IEEE ICASSP 2026 (51st International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2026). 5 pages, 1 figure, 3 tables. Project page: https://vcbsl.github.io/VoxMorph/
摘要:变形技术生成人工生物特征样本,这些样本结合了来自多个个体的特征,允许每个贡献者根据单个注册模板进行验证。虽然在人脸识别中得到了广泛的研究,但在语音生物识别中,这种漏洞在很大程度上仍未被探索。语音变形的先前工作是计算昂贵的,不可扩展的,并限于声学相似的身份对,限制实际部署。此外,现有的声音变形方法针对音频纹理、音乐或环境声音,并且不能转移到语音身份操纵。我们提出了VoxMorph,一个zero-shot框架,它可以从每个主题的五秒音频中产生高保真的语音变体,而无需模型重新训练。我们的方法将声音特征分解为韵律和音色嵌入,从而实现说话风格和身份的细粒度插值。这些嵌入通过球面线性插值(Slerp)进行融合,并使用自回归语言模型与条件流匹配网络进行合成。VoxMorph实现了最先进的性能,在严格的安全阈值下,在自动说话人验证系统上提供了2.6倍的音频质量增益,73%的可理解性错误减少,以及67.8%的变形攻击成功率。这项工作建立了一个实用的和可扩展的范例语音变形与生物识别安全的重大影响。代码和数据集可以在我们的项目页面上找到:https://vcbsl.github.io/VoxMorph/
摘要:Morphing techniques generate artificial biometric samples that combine features from multiple individuals, allowing each contributor to be verified against a single enrolled template. While extensively studied in face recognition, this vulnerability remains largely unexplored in voice biometrics. Prior work on voice morphing is computationally expensive, non-scalable, and limited to acoustically similar identity pairs, constraining practical deployment. Moreover, existing sound-morphing methods target audio textures, music, or environmental sounds and are not transferable to voice identity manipulation. We propose VoxMorph, a zero-shot framework that produces high-fidelity voice morphs from as little as five seconds of audio per subject without model retraining. Our method disentangles vocal traits into prosody and timbre embeddings, enabling fine-grained interpolation of speaking style and identity. These embeddings are fused via Spherical Linear Interpolation (Slerp) and synthesized using an autoregressive language model coupled with a Conditional Flow Matching network. VoxMorph achieves state-of-the-art performance, delivering a 2.6x gain in audio quality, a 73% reduction in intelligibility errors, and a 67.8% morphing attack success rate on automated speaker verification systems under strict security thresholds. This work establishes a practical and scalable paradigm for voice morphing with significant implications for biometric security. The code and dataset are available on our project page: https://vcbsl.github.io/VoxMorph/
【16】Generalizable Prompt Tuning for Audio-Language Models via Semantic Expansion
标题:通过语义扩展对音频语言模型进行可概括的即时调优
链接:https://arxiv.org/abs/2601.20867
摘要:提示调整在视觉语言模型(VLM)中取得了显着的进展,最近正在被用于音频语言模型(ALM)。然而,它的泛化能力在ALMs仍然在很大程度上未被探索。我们观察到,传统的快速调整ALM也遭受新的基础权衡,我们确定这个问题源于嵌入空间的语义结构的破坏。为了解决这个问题,我们提出了语义扩展提示调优(SEPT)-一个即插即用的框架,明确规范化的提示嵌入空间,将大型语言模型生成的语义邻居。SEPT引入了一种新的语义扩展损失与利润限制,促进类内的紧凑性和类间的可分性,从而增强提示嵌入空间的语义结构。为了进行全面的评估,我们建立了第一个在ALM中进行快速泛化的基准设置,包括基础到新的泛化和跨数据集的可移植性。大量的实验表明,SEPT在多个提示调优基线上持续提高泛化性能,同时在推理过程中保持计算成本。代码可在https://github.com/jhyukjang/SEPT上找到。
摘要:Prompt tuning has achieved remarkable progress in vision-language models (VLMs) and is recently being adopted for audio-language models (ALMs). However, its generalization ability in ALMs remains largely underexplored. We observe that conventional prompt tuning for ALMs also suffers from the Base-New Tradeoff, and we identify that this issue stems from the disrupted semantic structure of the embedding space. To address this issue, we propose Semantically Expanded Prompt Tuning (SEPT)-a plug-and-play framework that explicitly regularizes the prompt embedding space by incorporating semantic neighbors generated by large language models. SEPT introduces a novel semantic expansion loss with margin constraints that promote intra-class compactness and inter-class separability, thereby enhancing the semantic structure of the prompt embedding space. For comprehensive evaluation, we establish the first benchmark setup for prompt generalization in ALMs, covering both base-to-new generalization and cross-dataset transferability. Extensive experiments demonstrate that SEPT consistently improves generalization performance across multiple prompt tuning baselines, while maintaining computational cost during inference. Codes are available in https://github.com/jhyukjang/SEPT.
【17】TidyVoice 2026 Challenge Evaluation Plan
标题:TidyVoice 2026挑战评估计划
链接:https://arxiv.org/abs/2601.21960
备注:https://tidyvoice2026.github.io/
摘要:说话人确认系统的性能显着下降下的语言不匹配,一个关键的挑战加剧了该领域的依赖以英语为中心的数据。为了解决这个问题,我们提出了TidyVoice Challenge用于跨语言说话人验证。这项挑战利用了来自新的TidyVoice基准测试的TidyVoiceX数据集,这是一个来自Mozilla Common Voice的大规模多语言语料库,专门用于隔离大约40种语言之间的语言切换效果。参与者的任务是构建对这种不匹配具有鲁棒性的系统,主要使用跨语言试验的等错误率来评估性能。通过提供标准化数据、开源基线和严格的评估协议,这项挑战旨在推动研究朝着更公平、更具包容性和独立于语言的说话人识别技术发展,直接与Interspeech 2026主题“一起说话”保持一致。"
摘要:The performance of speaker verification systems degrades significantly under language mismatch, a critical challenge exacerbated by the field's reliance on English-centric data. To address this, we propose the TidyVoice Challenge for cross-lingual speaker verification. The challenge leverages the TidyVoiceX dataset from the novel TidyVoice benchmark, a large-scale, multilingual corpus derived from Mozilla Common Voice, and specifically curated to isolate the effect of language switching across approximately 40 languages. Participants will be tasked with building systems robust to this mismatch, with performance primarily evaluated using the Equal Error Rate on cross-language trials. By providing standardized data, open-source baselines, and a rigorous evaluation protocol, this challenge aims to drive research towards fairer, more inclusive, and language-independent speaker recognition technologies, directly aligning with the Interspeech 2026 theme, "Speaking Together."
【18】Representation-Regularized Convolutional Audio Transformer for Audio Understanding
标题:用于音频理解的表示规则化卷积音频Transformer
链接:https://arxiv.org/abs/2601.21612
备注:12 pages, 3 figures
摘要:基于Bootstrap的自监督学习(SSL)在音频理解方面取得了显着进展。然而,现有的方法通常在单个粒度级别上操作,限制了它们对复杂音频信号中固有的不同时间和频谱结构进行建模的能力。此外,从头开始的自举表示在计算上是昂贵的,通常需要大量的训练来收敛。在这项工作中,我们提出了卷积音频Transformer(CAT),一个统一的框架,旨在解决这些挑战。首先,为了捕获分层音频特征,CAT采用了多分辨率块,该块在不同粒度上聚合信息。其次,为了提高训练效率,我们引入了Representation Regularization目标。从生成建模中汲取灵感,这个辅助任务通过将其预测与来自冻结的预训练外部编码器的高质量语义表示对齐来指导学生模型。实验结果表明,CAT在音频理解基准上的性能明显优于基线。值得注意的是,它在AudioSet 20k数据集上实现了具有竞争力的性能,收敛速度比现有方法快5倍。代码和检查点将很快在https://github.com/realzhouchushu/CAT上发布。
摘要:Bootstrap-based Self-Supervised Learning (SSL) has achieved remarkable progress in audio understanding. However, existing methods typically operate at a single level of granularity, limiting their ability to model the diverse temporal and spectral structures inherent in complex audio signals. Furthermore, bootstrapping representations from scratch is computationally expensive, often requiring extensive training to converge. In this work, we propose the Convolutional Audio Transformer (CAT), a unified framework designed to address these challenges. First, to capture hierarchical audio features, CAT incorporates a Multi-resolution Block that aggregates information across varying granularities. Second, to enhance training efficiency, we introduce a Representation Regularization objective. Drawing inspiration from generative modeling, this auxiliary task guides the student model by aligning its predictions with high-quality semantic representations from frozen, pre-trained external encoders. Experimental results demonstrate that CAT significantly outperforms baselines on audio understanding benchmarks. Notably, it achieves competitive performance on the AudioSet 20k dataset with 5 times faster convergence than existing methods. Codes and checkpoints will be released soon at https://github.com/realzhouchushu/CAT.
【19】SemanticAudio: Audio Generation and Editing in Semantic Space
标题:SemanticAudio:语义空间中的音频生成和编辑
链接:https://arxiv.org/abs/2601.21402
摘要:近年来,文本到音频生成已经取得了显着的进展,为声音创作者提供了强大的工具,将文本灵感转化为生动的音频。然而,现有的模型主要直接在变分自动编码器(VAE)的声学潜在空间中操作,通常导致生成的音频和文本描述之间的次优对齐。在本文中,我们介绍了SemanticAudio,一个新的框架,直接在一个高层次的语义空间进行音频生成和编辑。我们定义这个语义空间作为一个紧凑的表示捕捉全球身份和时间序列的声音事件,不同于细粒度的声学细节。SemanticAudio采用两阶段流匹配架构:语义规划器首先生成这些紧凑的语义特征以勾画全局语义布局,声学合成器随后根据此语义规划生成高保真声学潜伏。利用这种解耦的设计,我们进一步引入了一种无需训练的文本引导编辑机制,该机制可以在不重新训练的情况下对通用音频进行精确的属性级修改。具体地,这是通过经由从源文本提示和目标文本提示导出的速度场的差异来操纵语义生成轨迹来实现的。大量的实验表明,SemanticAudio超过现有的主流方法在语义对齐。演示网址:https://semanticaudio1.github.io/
摘要:In recent years, Text-to-Audio Generation has achieved remarkable progress, offering sound creators powerful tools to transform textual inspirations into vivid audio. However, existing models predominantly operate directly in the acoustic latent space of a Variational Autoencoder (VAE), often leading to suboptimal alignment between generated audio and textual descriptions. In this paper, we introduce SemanticAudio, a novel framework that conducts both audio generation and editing directly in a high-level semantic space. We define this semantic space as a compact representation capturing the global identity and temporal sequence of sound events, distinct from fine-grained acoustic details. SemanticAudio employs a two-stage Flow Matching architecture: the Semantic Planner first generates these compact semantic features to sketch the global semantic layout, and the Acoustic Synthesizer subsequently produces high-fidelity acoustic latents conditioned on this semantic plan. Leveraging this decoupled design, we further introduce a training-free text-guided editing mechanism that enables precise attribute-level modifications on general audio without retraining. Specifically, this is achieved by steering the semantic generation trajectory via the difference of velocity fields derived from source and target text prompts. Extensive experiments demonstrate that SemanticAudio surpasses existing mainstream approaches in semantic alignment. Demo available at: https://semanticaudio1.github.io/
【20】Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
标题:迈向稳健的合成障碍语音识别:超越WER的LLM-Agent后ASR纠正
链接:https://arxiv.org/abs/2601.21347
备注:Accepted to ICASSP 2026
摘要:虽然自动语音识别(ASR)通常以单词错误率(WER)为基准,但实际应用最终取决于语义保真度。这种不匹配对于构音障碍的言语来说尤其成问题,在构音障碍的言语中,发音不精确和不流利会导致严重的语义扭曲。为了弥合这一差距,我们引入了一个基于大语言模型(LLM)的代理后ASR校正:一个判断编辑器的前k个ASR假设,保持高置信度的跨度,重写不确定的部分,并在zero-shot和微调模式。与此同时,我们发布了SAP-Hypo 5,这是构音障碍语音矫正的最大基准,以实现可重复性和未来探索。在多角度评估下,我们的代理实现了14.51%的WER减少以及大量的语义增益,包括MENLI中的+7.59 pp改进和Slot Micro F1中的+7.66 pp改进。我们的分析进一步表明,WER是高度敏感的域转移,而语义指标与下游任务的性能更密切相关。
摘要:While Automatic Speech Recognition (ASR) is typically benchmarked by word error rate (WER), real-world applications ultimately hinge on semantic fidelity. This mismatch is particularly problematic for dysarthric speech, where articulatory imprecision and disfluencies can cause severe semantic distortions. To bridge this gap, we introduce a Large Language Model (LLM)-based agent for post-ASR correction: a Judge-Editor over the top-k ASR hypotheses that keeps high-confidence spans, rewrites uncertain segments, and operates in both zero-shot and fine-tuned modes. In parallel, we release SAP-Hypo5, the largest benchmark for dysarthric speech correction, to enable reproducibility and future exploration. Under multi-perspective evaluation, our agent achieves a 14.51% WER reduction alongside substantial semantic gains, including a +7.59 pp improvement in MENLI and +7.66 pp in Slot Micro F1 on challenging samples. Our analysis further reveals that WER is highly sensitive to domain shift, whereas semantic metrics correlate more closely with downstream task performance.
【21】DNN-Based Online Source Counting Based on Spatial Generalized Magnitude Squared Coherence
标题:基于空间广义幅度平方相关性的DNN在线源计数
链接:https://arxiv.org/abs/2601.21114
备注:in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2026, Barcelona, Spain
摘要:在声源定位、声源分离和多麦克风语音增强等声学信号处理任务中,有源声源的数量是一个关键参数。提出了一种基于空间相干性检测有源信号数目变化的在线源计数方法。所提出的方法利用了这样一个事实:空间白背景噪声中的单个相干源产生高空间相干性,而只有噪声会导致低空间相干性。通过应用空间白化操作,源计数问题被重新表述为变化检测任务,旨在识别活动源数量变化时的时间帧。该方法利用广义幅度平方相干性作为量化空间相干性的度量,为训练用于检测逐帧源计数变化的紧凑神经网络提供特征。双耳助听器在混响声场景中的仿真结果表明,多达4个扬声器和背景噪声的在线源计数所提出的方法的有效性。
摘要:The number of active sound sources is a key parameter in many acoustic signal processing tasks, such as source localization, source separation, and multi-microphone speech enhancement. This paper proposes a novel method for online source counting by detecting changes in the number of active sources based on spatial coherence. The proposed method exploits the fact that a single coherent source in spatially white background noise yields high spatial coherence, whereas only noise results in low spatial coherence. By applying a spatial whitening operation, the source counting problem is reformulated as a change detection task, aiming to identify the time frames when the number of active sources changes. The method leverages the generalized magnitude-squared coherence as a measure to quantify spatial coherence, providing features for a compact neural network trained to detect source count changes framewise. Simulation results with binaural hearing aids in reverberant acoustic scenes with up to 4 speakers and background noise demonstrate the effectiveness of the proposed method for online source counting.
【22】Unseen but not Unknown: Using Dataset Concealment to Robustly Evaluate Speech Quality Estimation Models
标题:未知但并非未知:使用数据集隐藏来稳健地评估语音质量估计模型
链接:https://arxiv.org/abs/2601.21110
备注:To be appear in Proc. ICASSP 2026
摘要:我们介绍数据集Concealment(DSC),一个严格的新的程序,用于评估和解释客观的语音质量估计模型。DSC量化和分解研究结果与实际应用需求之间的性能差距,同时提供对模型行为和数据集特征的上下文和其他见解。我们还展示了在使用多个数据集训练模型时,通过使用AlignNet的数据集Aligner来解决语料库效应的好处。我们使用九个训练数据集和九个看不见的数据集以及三个经过充分研究的模型来展示DSC和Aligner的改进:MOSNet、NISQA和基于Wav2Vec2.0的模型。DSC提供了模型泛化能力和局限性的可解释视图,同时允许在训练时使用所有可用数据。另一个结果是,在训练过程中将1000个参数的数据集Aligner添加到9400万个参数的Wav2Vec模型中,确实显著提高了所得模型估计未知数据的语音质量的能力。
摘要:We introduce Dataset Concealment (DSC), a rigorous new procedure for evaluating and interpreting objective speech quality estimation models. DSC quantifies and decomposes the performance gap between research results and real-world application requirements, while offering context and additional insights into model behavior and dataset characteristics. We also show the benefits of addressing the corpus effect by using the dataset Aligner from AlignNet when training models with multiple datasets. We demonstrate DSC and the improvements from the Aligner using nine training datasets and nine unseen datasets with three well-studied models: MOSNet, NISQA, and a Wav2Vec2.0-based model. DSC provides interpretable views of the generalization capabilities and limitations of models, while allowing all available data to be used at training. An additional result is that adding the 1000 parameter dataset Aligner to the 94 million parameter Wav2Vec model during training does significantly improve the resulting model's ability to estimate speech quality for unseen data.
【1】TidyVoice 2026 Challenge Evaluation Plan
标题:TidyVoice 2026挑战评估计划
链接:https://arxiv.org/abs/2601.21960
备注:https://tidyvoice2026.github.io/
摘要:说话人确认系统的性能显着下降下的语言不匹配,一个关键的挑战加剧了该领域的依赖以英语为中心的数据。为了解决这个问题,我们提出了TidyVoice Challenge用于跨语言说话人验证。这项挑战利用了来自新的TidyVoice基准测试的TidyVoiceX数据集,这是一个来自Mozilla Common Voice的大规模多语言语料库,专门用于隔离大约40种语言之间的语言切换效果。参与者的任务是构建对这种不匹配具有鲁棒性的系统,主要使用跨语言试验的等错误率来评估性能。通过提供标准化数据、开源基线和严格的评估协议,这项挑战旨在推动研究朝着更公平、更具包容性和独立于语言的说话人识别技术发展,直接与Interspeech 2026主题“一起说话”保持一致。"
摘要:The performance of speaker verification systems degrades significantly under language mismatch, a critical challenge exacerbated by the field's reliance on English-centric data. To address this, we propose the TidyVoice Challenge for cross-lingual speaker verification. The challenge leverages the TidyVoiceX dataset from the novel TidyVoice benchmark, a large-scale, multilingual corpus derived from Mozilla Common Voice, and specifically curated to isolate the effect of language switching across approximately 40 languages. Participants will be tasked with building systems robust to this mismatch, with performance primarily evaluated using the Equal Error Rate on cross-language trials. By providing standardized data, open-source baselines, and a rigorous evaluation protocol, this challenge aims to drive research towards fairer, more inclusive, and language-independent speaker recognition technologies, directly aligning with the Interspeech 2026 theme, "Speaking Together."
【2】DisContSE: Single-Step Diffusion Speech Enhancement Based on Joint Discrete and Continuous Embeddings
标题:DisContSE:基于离散和连续联合嵌入的一步扩散语音增强
链接:https://arxiv.org/abs/2601.21940
备注:Accepted by IEEE ICASSP 2026
摘要:离散音频编解码器特征的扩散语音增强由于其改进的语音分量重建能力而受到广泛关注。然而,由于多次反向过程迭代,它们通常遭受高推理计算复杂度。此外,它们通常在非侵入性指标上取得有希望的结果,但在侵入性指标上表现不佳,因为它们可能难以重建正确的电话。在本文中,我们提出了DisContSE,一个有效的基于扩散的语音增强模型联合离散编解码器令牌和连续嵌入。我们的贡献是三方面的。首先,我们制定了一个离散的和连续的增强模块上操作的离散音频编解码器令牌和连续嵌入,分别实现同时提高保真度和可懂度。其次,进一步采用语义增强模块,以达到最佳的语音准确度。第三,我们实现了一个单步有效的反向过程中的推理与一个新的量化误差掩模初始化策略,这是第一个成功的单步扩散语音增强的基础上的音频编解码器。在URGENT 2024语音增强挑战赛数据分割上进行了训练和评估,提出的DisContSE优于PESQ,POLQA,UTMOS和主观ITU-T P.808听力测试中报告的最佳时域和频域扩散基线方法,显然达到了整体排名第一。
摘要:Diffusion speech enhancement on discrete audio codec features gain immense attention due to their improved speech component reconstruction capability. However, they usually suffer from high inference computational complexity due to multiple reverse process iterations. Furthermore, they generally achieve promising results on non-intrusive metrics but show poor performance on intrusive metrics, as they may struggle in reconstructing the correct phones. In this paper, we propose DisContSE, an efficient diffusion-based speech enhancement model on joint discrete codec tokens and continuous embeddings. Our contributions are three-fold. First, we formulate both a discrete and a continuous enhancement module operating on discrete audio codec tokens and continuous embeddings, respectively, to achieve improved fidelity and intelligibility simultaneously. Second, a semantic enhancement module is further adopted to achieve optimal phonetic accuracy. Third, we achieve a single-step efficient reverse process in inference with a novel quantization error mask initialization strategy, which, according to our knowledge, is the first successful single-step diffusion speech enhancement based on an audio codec. Trained and evaluated on URGENT 2024 Speech Enhancement Challenge data splits, the proposed DisContSE excels top-reported time- and frequency-domain diffusion baseline methods in PESQ, POLQA, UTMOS, and in a subjective ITU-T P.808 listening test, clearly achieving an overall top rank.
【3】Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
标题:基于语音质量的低质量语音定位和文本到语音合成制品
链接:https://arxiv.org/abs/2601.21886
备注:Accepted at ICASSP 2026
摘要:大量的作品从话语或系统级的角度来看待语音的自动评估。虽然这些方法在判断整体质量方面很好,但它们不能充分解释为什么给一个话语分配一定的分数。帧级分数可以提供更好的可解释性,但预测它们的模型更难调整和正则化,因为在训练期间没有强目标可用。在这项工作中,我们表明,话语级语音质量预测可以正则化与基于段的一致性约束,显着降低帧级的随机性。然后,我们展示了两个应用程序,涉及帧级分数:部分欺骗的情况下,在两个国家的最先进的文本到语音系统的合成文物的检测。对于后者,我们进行了听力测试,并确认听众率段质量差更经常在低帧级分数定义的集合中比在随机控制集。
摘要:A large number of works view the automatic assessment of speech from an utterance- or system-level perspective. While such approaches are good in judging overall quality, they cannot adequately explain why a certain score was assigned to an utterance. frame-level scores can provide better interpretability, but models predicting them are harder to tune and regularize since no strong targets are available during training. In this work, we show that utterance-level speech quality predictors can be regularized with a segment-based consistency constraint which notably reduces frame-level stochasticity. We then demonstrate two applications involving frame-level scores: The partial spoof scenario and the detection of synthesis artefacts in two state-of-the-art text-to-speech systems. For the latter, we perform listening tests and confirm that listeners rate segments to be of poor quality more often in the set defined by low frame-level scores than in a random control set.
【4】Representation-Regularized Convolutional Audio Transformer for Audio Understanding
标题:用于音频理解的表示规则化卷积音频Transformer
链接:https://arxiv.org/abs/2601.21612
备注:12 pages, 3 figures
摘要:基于Bootstrap的自监督学习(SSL)在音频理解方面取得了显着进展。然而,现有的方法通常在单个粒度级别上操作,限制了它们对复杂音频信号中固有的不同时间和频谱结构进行建模的能力。此外,从头开始的自举表示在计算上是昂贵的,通常需要大量的训练来收敛。在这项工作中,我们提出了卷积音频Transformer(CAT),一个统一的框架,旨在解决这些挑战。首先,为了捕获分层音频特征,CAT采用了多分辨率块,该块在不同粒度上聚合信息。其次,为了提高训练效率,我们引入了Representation Regularization目标。从生成建模中汲取灵感,这个辅助任务通过将其预测与来自冻结的预训练外部编码器的高质量语义表示对齐来指导学生模型。实验结果表明,CAT在音频理解基准上的性能明显优于基线。值得注意的是,它在AudioSet 20k数据集上实现了具有竞争力的性能,收敛速度比现有方法快5倍。代码和检查点将很快在https://github.com/realzhouchushu/CAT上发布。
摘要:Bootstrap-based Self-Supervised Learning (SSL) has achieved remarkable progress in audio understanding. However, existing methods typically operate at a single level of granularity, limiting their ability to model the diverse temporal and spectral structures inherent in complex audio signals. Furthermore, bootstrapping representations from scratch is computationally expensive, often requiring extensive training to converge. In this work, we propose the Convolutional Audio Transformer (CAT), a unified framework designed to address these challenges. First, to capture hierarchical audio features, CAT incorporates a Multi-resolution Block that aggregates information across varying granularities. Second, to enhance training efficiency, we introduce a Representation Regularization objective. Drawing inspiration from generative modeling, this auxiliary task guides the student model by aligning its predictions with high-quality semantic representations from frozen, pre-trained external encoders. Experimental results demonstrate that CAT significantly outperforms baselines on audio understanding benchmarks. Notably, it achieves competitive performance on the AudioSet 20k dataset with 5 times faster convergence than existing methods. Codes and checkpoints will be released soon at https://github.com/realzhouchushu/CAT.
【5】SemanticAudio: Audio Generation and Editing in Semantic Space
标题:SemanticAudio:语义空间中的音频生成和编辑
链接:https://arxiv.org/abs/2601.21402
摘要:近年来,文本到音频生成已经取得了显着的进展,为声音创作者提供了强大的工具,将文本灵感转化为生动的音频。然而,现有的模型主要直接在变分自动编码器(VAE)的声学潜在空间中操作,通常导致生成的音频和文本描述之间的次优对齐。在本文中,我们介绍了SemanticAudio,一个新的框架,直接在一个高层次的语义空间进行音频生成和编辑。我们定义这个语义空间作为一个紧凑的表示捕捉全球身份和时间序列的声音事件,不同于细粒度的声学细节。SemanticAudio采用两阶段流匹配架构:语义规划器首先生成这些紧凑的语义特征以勾画全局语义布局,声学合成器随后根据此语义规划生成高保真声学潜伏。利用这种解耦的设计,我们进一步引入了一种无需训练的文本引导编辑机制,该机制可以在不重新训练的情况下对通用音频进行精确的属性级修改。具体地,这是通过经由从源文本提示和目标文本提示导出的速度场的差异来操纵语义生成轨迹来实现的。大量的实验表明,SemanticAudio超过现有的主流方法在语义对齐。演示网址:https://semanticaudio1.github.io/
摘要:In recent years, Text-to-Audio Generation has achieved remarkable progress, offering sound creators powerful tools to transform textual inspirations into vivid audio. However, existing models predominantly operate directly in the acoustic latent space of a Variational Autoencoder (VAE), often leading to suboptimal alignment between generated audio and textual descriptions. In this paper, we introduce SemanticAudio, a novel framework that conducts both audio generation and editing directly in a high-level semantic space. We define this semantic space as a compact representation capturing the global identity and temporal sequence of sound events, distinct from fine-grained acoustic details. SemanticAudio employs a two-stage Flow Matching architecture: the Semantic Planner first generates these compact semantic features to sketch the global semantic layout, and the Acoustic Synthesizer subsequently produces high-fidelity acoustic latents conditioned on this semantic plan. Leveraging this decoupled design, we further introduce a training-free text-guided editing mechanism that enables precise attribute-level modifications on general audio without retraining. Specifically, this is achieved by steering the semantic generation trajectory via the difference of velocity fields derived from source and target text prompts. Extensive experiments demonstrate that SemanticAudio surpasses existing mainstream approaches in semantic alignment. Demo available at: https://semanticaudio1.github.io/
【6】Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
标题:迈向稳健的合成障碍语音识别:超越WER的LLM-Agent后ASR纠正
链接:https://arxiv.org/abs/2601.21347
备注:Accepted to ICASSP 2026
摘要:虽然自动语音识别(ASR)通常以单词错误率(WER)为基准,但实际应用最终取决于语义保真度。这种不匹配对于构音障碍的言语来说尤其成问题,在构音障碍的言语中,发音不精确和不流利会导致严重的语义扭曲。为了弥合这一差距,我们引入了一个基于大语言模型(LLM)的代理后ASR校正:一个判断编辑器的前k个ASR假设,保持高置信度的跨度,重写不确定的部分,并在zero-shot和微调模式。与此同时,我们发布了SAP-Hypo 5,这是构音障碍语音矫正的最大基准,以实现可重复性和未来探索。在多角度评估下,我们的代理实现了14.51%的WER减少以及大量的语义增益,包括MENLI中的+7.59 pp改进和Slot Micro F1中的+7.66 pp改进。我们的分析进一步表明,WER是高度敏感的域转移,而语义指标与下游任务的性能更密切相关。
摘要:While Automatic Speech Recognition (ASR) is typically benchmarked by word error rate (WER), real-world applications ultimately hinge on semantic fidelity. This mismatch is particularly problematic for dysarthric speech, where articulatory imprecision and disfluencies can cause severe semantic distortions. To bridge this gap, we introduce a Large Language Model (LLM)-based agent for post-ASR correction: a Judge-Editor over the top-k ASR hypotheses that keeps high-confidence spans, rewrites uncertain segments, and operates in both zero-shot and fine-tuned modes. In parallel, we release SAP-Hypo5, the largest benchmark for dysarthric speech correction, to enable reproducibility and future exploration. Under multi-perspective evaluation, our agent achieves a 14.51% WER reduction alongside substantial semantic gains, including a +7.59 pp improvement in MENLI and +7.66 pp in Slot Micro F1 on challenging samples. Our analysis further reveals that WER is highly sensitive to domain shift, whereas semantic metrics correlate more closely with downstream task performance.
【7】DNN-Based Online Source Counting Based on Spatial Generalized Magnitude Squared Coherence
标题:基于空间广义幅度平方相关性的DNN在线源计数
链接:https://arxiv.org/abs/2601.21114
备注:in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2026, Barcelona, Spain
摘要:在声源定位、声源分离和多麦克风语音增强等声学信号处理任务中,有源声源的数量是一个关键参数。提出了一种基于空间相干性检测有源信号数目变化的在线源计数方法。所提出的方法利用的事实是,一个单一的相干源在空间白背景噪声产生高的空间相干性,而只有噪声导致低的空间相干性。通过应用空间白化操作,源计数问题被重新表述为变化检测任务,旨在识别活动源数量变化时的时间帧。该方法利用广义幅度平方相干性作为量化空间相干性的度量,为训练用于检测逐帧源计数变化的紧凑神经网络提供特征。双耳助听器在混响声场景中的仿真结果表明,多达4个扬声器和背景噪声的在线源计数所提出的方法的有效性。
摘要:The number of active sound sources is a key parameter in many acoustic signal processing tasks, such as source localization, source separation, and multi-microphone speech enhancement. This paper proposes a novel method for online source counting by detecting changes in the number of active sources based on spatial coherence. The proposed method exploits the fact that a single coherent source in spatially white background noise yields high spatial coherence, whereas only noise results in low spatial coherence. By applying a spatial whitening operation, the source counting problem is reformulated as a change detection task, aiming to identify the time frames when the number of active sources changes. The method leverages the generalized magnitude-squared coherence as a measure to quantify spatial coherence, providing features for a compact neural network trained to detect source count changes framewise. Simulation results with binaural hearing aids in reverberant acoustic scenes with up to 4 speakers and background noise demonstrate the effectiveness of the proposed method for online source counting.
【8】Unseen but not Unknown: Using Dataset Concealment to Robustly Evaluate Speech Quality Estimation Models
标题:未知但并非未知:使用数据集隐藏来稳健地评估语音质量估计模型
链接:https://arxiv.org/abs/2601.21110
备注:To be appear in Proc. ICASSP 2026
摘要:我们介绍数据集Concealment(DSC),一个严格的新的程序,用于评估和解释客观的语音质量估计模型。DSC量化和分解研究结果与实际应用需求之间的性能差距,同时提供对模型行为和数据集特征的上下文和其他见解。我们还展示了在使用多个数据集训练模型时,通过使用AlignNet的数据集Aligner来解决语料库效应的好处。我们使用九个训练数据集和九个看不见的数据集以及三个经过充分研究的模型来展示DSC和Aligner的改进:MOSNet,NISQA和基于Wav2Vec2.0的模型。DSC提供了模型泛化能力和局限性的可解释视图,同时允许在训练时使用所有可用数据。另一个结果是,在训练过程中将1000个参数的数据集Aligner添加到9400万个参数的Wav2Vec模型中,确实显著提高了所得模型估计未知数据的语音质量的能力。
摘要:We introduce Dataset Concealment (DSC), a rigorous new procedure for evaluating and interpreting objective speech quality estimation models. DSC quantifies and decomposes the performance gap between research results and real-world application requirements, while offering context and additional insights into model behavior and dataset characteristics. We also show the benefits of addressing the corpus effect by using the dataset Aligner from AlignNet when training models with multiple datasets. We demonstrate DSC and the improvements from the Aligner using nine training datasets and nine unseen datasets with three well-studied models: MOSNet, NISQA, and a Wav2Vec2.0-based model. DSC provides interpretable views of the generalization capabilities and limitations of models, while allowing all available data to be used at training. An additional result is that adding the 1000 parameter dataset Aligner to the 94 million parameter Wav2Vec model during training does significantly improve the resulting model's ability to estimate speech quality for unseen data.
【9】Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
标题:通过可学习投影降低基于LLM的语音识别中的提示敏感性
链接:https://arxiv.org/abs/2601.20898
备注:Paper accepted at ICASSP 2026
摘要:基于LLM的自动语音识别(ASR)是一种成熟的方法,通过语音到LLM投影仪将语音基础模型连接到大型语言模型(LLM),产生了有希望的结果。这些架构中的一个常见设计选择是在训练和推理期间使用固定的手动定义的提示。这种设置不仅能够在一系列实际场景中实现适用性,还有助于最大限度地提高模型性能。然而,及时设计的影响仍然没有得到充分的探讨。本文对不同数据集上常用的提示进行了全面的分析,结果表明提示选择会显著影响ASR性能并引入不稳定性,没有一个提示在所有情况下都表现最好。受语音到LLM投影仪的启发,我们提出了一个提示投影仪模块,这是一个简单的模型不可知扩展,可以学习将提示嵌入投影到LLM输入空间的更有效区域,而无需修改底层的基于LLM的ASR模型。在四个数据集上的实验表明,添加提示投影仪始终提高性能,减少变异性,并优于最佳手动选择的提示。
摘要:LLM-based automatic speech recognition (ASR), a well-established approach, connects speech foundation models to large language models (LLMs) through a speech-to-LLM projector, yielding promising results. A common design choice in these architectures is the use of a fixed, manually defined prompt during both training and inference. This setup not only enables applicability across a range of practical scenarios, but also helps maximize model performance. However, the impact of prompt design remains underexplored. This paper presents a comprehensive analysis of commonly used prompts across diverse datasets, showing that prompt choice significantly affects ASR performance and introduces instability, with no single prompt performing best across all cases. Inspired by the speech-to-LLM projector, we propose a prompt projector module, a simple, model-agnostic extension that learns to project prompt embeddings to more effective regions of the LLM input space, without modifying the underlying LLM-based ASR model. Experiments on four datasets show that the addition of a prompt projector consistently improves performance, reduces variability, and outperforms the best manually selected prompts.
【10】Qwen3-ASR Technical Report
标题:Qwen 3-ASR技术报告
链接:https://arxiv.org/abs/2601.21337
备注:https://github.com/QwenLM/Qwen3-ASR
摘要:在这份报告中,我们介绍了Qwen 3-ASR家族,其中包括两个强大的一体化语音识别模型和一个新的非自回归语音强制对齐模型。Qwen 3-ASR-1.7B和Qwen 3-ASR-0.6B是ASR模型,支持52种语言和方言的语言识别和ASR。它们都利用了大规模的语音训练数据和基础模型Qwen 3-Omni强大的音频理解能力。除了开源基准之外,我们还进行了全面的内部评估,因为ASR模型在开源基准得分上可能差异不大,但在真实场景中表现出显着的质量差异。实验表明,1.7B版本在开源ASR模型中实现了SOTA性能,并与最强的专有API竞争,而0.6B版本提供了最佳的准确性-效率权衡。Qwen 3-ASR-0.6B可以实现平均TTFT低至92 ms,并在128并发的情况下在1秒内转录2000秒的语音。Qwen 3-ForcedAligner-0.6B是一个基于LLM的NAR时间戳预测器,能够对齐11种语言的文本-语音对。时间戳精度实验表明,该模型优于三种最强的力对齐模型,在效率和通用性方面具有更多优势。为了进一步加速ASR和音频理解的社区研究,我们在Apache 2.0许可下发布了这些模型。
摘要:In this report, we introduce Qwen3-ASR family, which includes two powerful all-in-one speech recognition models and a novel non-autoregressive speech forced alignment model. Qwen3-ASR-1.7B and Qwen3-ASR-0.6B are ASR models that support language identification and ASR for 52 languages and dialects. Both of them leverage large-scale speech training data and the strong audio understanding ability of their foundation model Qwen3-Omni. We conduct comprehensive internal evaluation besides the open-sourced benchmarks as ASR models might differ little on open-sourced benchmark scores but exhibit significant quality differences in real-world scenarios. The experiments reveal that the 1.7B version achieves SOTA performance among open-sourced ASR models and is competitive with the strongest proprietary APIs while the 0.6B version offers the best accuracy-efficiency trade-off. Qwen3-ASR-0.6B can achieve an average TTFT as low as 92ms and transcribe 2000 seconds speech in 1 second at a concurrency of 128. Qwen3-ForcedAligner-0.6B is an LLM based NAR timestamp predictor that is able to align text-speech pairs in 11 languages. Timestamp accuracy experiments show that the proposed model outperforms the three strongest force alignment models and takes more advantages in efficiency and versatility. To further accelerate the community research of ASR and audio understanding, we release these models under the Apache 2.0 license.
【11】Evaluating Spatialized Auditory Cues for Rapid Attention Capture in XR
标题:评估XR中快速注意力捕获的空间化听觉线索
链接:https://arxiv.org/abs/2601.21264
备注:8 pages, 4 figures. This is the author's version of the article that will appear at the IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (IEEE VRW) 2026
摘要:在时间关键延展实境(XR)场景中,用户必须在从事主要任务时迅速将注意力重新定位到危险,警报或指令,空间音频可以提供即时的方向提示,而不会占用视觉带宽。然而,这样的场景只能提供短暂的听觉暴露,需要用户快速解释声音方向,而无需延长收听或头部驱动的细化。本文报告了一个控制的探索性研究,快速空间音频定位在XR。使用HRTF呈现的宽带刺激从一个半密集的方向围绕听众,我们量化如何准确地用户可以推断出粗糙的方向,从简短的音频单独。我们进一步研究了短期视听反馈训练作为轻量级校准机制的效果。我们的研究结果表明,简短的空间线索可以传达粗糙的方向信息,即使是短校准可以提高用户的听觉信号的感知。虽然这些结果强调了空间音频在快速注意力引导方面的潜力,但它们也表明,单独的听觉提示可能无法为复杂或高风险的任务提供足够的精度,并且空间音频在与其他感觉方式或视觉提示补充时可能是最有效的,而不依赖于头部驱动的细化。我们利用这项关于空间音频的研究作为对可穿戴XR的第一阶段注意力引导通道的初步调查(例如,VR头戴式显示器和AR智能眼镜),并提供有关刺激选择和校准的设计见解,用于时间关键型用途。
摘要:In time-critical eXtended reality (XR) scenarios where users must rapidly reorient their attention to hazards, alerts, or instructions while engaged in a primary task, spatial audio can provide an immediate directional cue without occupying visual bandwidth. However, such scenarios can afford only a brief auditory exposure, requiring users to interpret sound direction quickly and without extended listening or head-driven refinement. This paper reports a controlled exploratory study of rapid spatial-audio localization in XR. Using HRTF-rendered broadband stimuli presented from a semi-dense set of directions around the listener, we quantify how accurately users can infer coarse direction from brief audio alone. We further examine the effects of short-term visuo-auditory feedback training as a lightweight calibration mechanism. Our findings show that brief spatial cues can convey coarse directional information, and that even short calibration can improve users' perception of aural signals. While these results highlight the potential of spatial audio for rapid attention guidance, they also show that auditory cues alone may not provide sufficient precision for complex or high-stakes tasks, and that spatial audio may be most effective when complemented by other sensory modalities or visual cues, without relying on head-driven refinement. We leverage this study on spatial audio as a preliminary investigation into a first-stage attention-guidance channel for wearable XR (e.g., VR head-mounted displays and AR smart glasses), and provide design insights on stimulus selection and calibration for time-critical use.
【12】Music Plagiarism Detection: Problem Formulation and a Segment-based Solution
标题:音乐抄袭检测:问题制定和基于片段的解决方案
链接:https://arxiv.org/abs/2601.21260
摘要:近年来,音乐剽窃问题已成为一个更加紧迫的社会问题。随着音乐信息检索研究的进展,人们越来越努力地解决与音乐剽窃相关的问题。然而,许多研究,包括我们以前的工作,进行了研究,没有明确定义音乐剽窃检测任务实际上涉及什么。这种缺乏明确定义的情况减缓了研究进展,并使其难以将结果应用于现实世界的场景。为了解决这种情况,我们定义了音乐剽窃检测与其他MIR任务的不同之处,并解释了需要解决的问题。我们引入了相似音乐对数据集来支持这个新定义的任务。此外,我们提出了一种基于片段转录的方法来解决这个问题。我们的演示和数据集可以在https://github.com/Mippia/ICASSP2026-MPD上找到。
摘要:Recently, the problem of music plagiarism has emerged as an even more pressing social issue. As music information retrieval research advances, there is a growing effort to address issues related to music plagiarism. However, many studies, including our previous work, have conducted research without clearly defining what the music plagiarism detection task actually involves. This lack of a clear definition has slowed research progress and made it hard to apply results to real-world scenarios. To fix this situation, we defined how Music Plagiarism Detection is different from other MIR tasks and explained what problems need to be solved. We introduce the Similar Music Pair dataset to support this newly defined task. In addition, we propose a method based on segment transcription as one way to solve the task. Our demo and dataset are available at https://github.com/Mippia/ICASSP2026-MPD.
【13】Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
标题:使用通用电话识别和语音特定音素对比建模的多语言发音障碍言语评估
链接:https://arxiv.org/abs/2601.21205
备注:10 pages, 4 figures
摘要:与构音障碍相关的神经系统疾病的日益普遍,激发了对适用于各种语言的自动可懂度评估方法的需求。然而,大多数现有的方法要么局限于一种语言或未能捕捉语言特定的因素塑造可理解性。我们提出了一个多语种的音素生产评估框架,集成了通用的电话识别与语言特定的音素解释,使用对比的音位特征距离的电话到音素映射和序列对齐。该框架产生三个指标:音素错误率(PER),语音特征错误率(PFER),和一个新提出的无干扰措施,音素覆盖率(PhonCov)。对英语、西班牙语、意大利语和泰米尔语的分析表明,PER受益于映射和对齐的组合,PFER受益于单独的对齐,PhonCov受益于映射。进一步的分析表明,所提出的框架捕捉临床上有意义的模式,与构音障碍语音的既定观察一致的清晰度下降。
摘要:The growing prevalence of neurological disorders associated with dysarthria motivates the need for automated intelligibility assessment methods that are applicalbe across languages. However, most existing approaches are either limited to a single language or fail to capture language-specific factors shaping intelligibility. We present a multilingual phoneme-production assessment framework that integrates universal phone recognition with language-specific phoneme interpretation using contrastive phonological feature distances for phone-to-phoneme mapping and sequence alignment. The framework yields three metrics: phoneme error rate (PER), phonological feature error rate (PFER), and a newly proposed alignment-free measure, phoneme coverage (PhonCov). Analysis on English, Spanish, Italian, and Tamil show that PER benefits from the combination of mapping and alignment, PFER from alignment alone, and PhonCov from mapping. Further analyses demonstrate that the proposed framework captures clinically meaningful patterns of intelligibility degradation consistent with established observations of dysarthric speech.
【14】PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
标题:PhaseCoder:多模式LLM的麦克风几何不可知的空间音频理解
链接:https://arxiv.org/abs/2601.21124
摘要:当前的多模态LLM将音频作为单声道流处理,忽略了嵌入式AI所必需的丰富空间信息。相反,现有的空间音频模型受限于固定的麦克风几何形状,从而阻止了跨不同设备的部署。我们提出了PhaseCoder,一个仅变换器的空间音频编码器,它与麦克风几何形状无关。PhaseCoder将原始多通道音频和麦克风坐标作为输入来执行定位,并产生鲁棒的空间嵌入。我们证明了Gemma 3n LLM可以被微调到由PhaseCoder产生的“空间音频令牌”。我们展示了我们的编码器在麦克风不变的定位基准上实现了最先进的结果,并且首次使LLM能够从任意麦克风阵列执行复杂的空间推理和有针对性的转录任务。
摘要:Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained to fixed microphone geometries, preventing deployment across diverse devices. We present PhaseCoder, a transformer-only spatial audio encoder that is agnostic to microphone geometry. PhaseCoder takes raw multichannel audio and microphone coordinates as inputs to perform localization and produces robust spatial embeddings. We demonstrate that Gemma 3n LLM can be fine-tuned to reason over "Spatial Audio Tokens" produced by PhaseCoder. We show our encoder achieves state-of-the-art results on microphone-invariant localization benchmarks and, for the first time, enables an LLM to perform complex spatial reasoning and targeted transcription tasks from an arbitrary microphone array.
【15】Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
标题:具有自监督语音表示的语音增强模型的位置不变微调
链接:https://arxiv.org/abs/2601.21084
备注:Accepted to ICASSP 2026
摘要:将前端语音增强(SE)模型与基于自监督学习(SSL)的语音模型相结合,对于噪声条件下的下游任务是有效的。SE模型通常使用SSL表示进行微调,增强语音和干净语音之间的均方误差(MSE)损失。然而,MSE倾向于利用SSL模型中的位置嵌入,允许通过位置相关性而不是内容相关信息来最小化目标。这项工作框架的问题作为自我监督表示微调的一般限制,并通过代表指导SE调查。考虑了两种战略:(1)零填充,以前在SSL预训练中探索过,但在微调设置中进行了检查,以及(2)具有软DTW损失的速度扰动。实验表明,基于软DTW的方法实现了更快的收敛速度和改善下游性能,强调了位置不变的微调在基于SSL的语音建模的重要性。
摘要:Integrating front-end speech enhancement (SE) models with self-supervised learning (SSL)-based speech models is effective for downstream tasks in noisy conditions. SE models are commonly fine-tuned using SSL representations with mean squared error (MSE) loss between enhanced and clean speech. However, MSE is prone to exploiting positional embeddings in SSL models, allowing the objective to be minimised through positional correlations instead of content-related information. This work frames the problem as a general limitation of self-supervised representation fine-tuning and investigates it through representation-guided SE. Two strategies are considered: (1) zero-padding, previously explored in SSL pre-training but here examined in the fine-tuning setting, and (2) speed perturbations with a soft-DTW loss. Experiments show that the soft-DTW-based approach achieves faster convergence and improved downstream performance, underscoring the importance of position-invariant fine-tuning in SSL-based speech modelling.
【16】asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
标题:asr_eval:用于多参考和流语音识别评估的算法和工具
链接:https://arxiv.org/abs/2601.20992
摘要:我们提出了几个改进的语音识别评估。首先,我们提出了一个字符串对齐算法,支持多引用标记,任意长度的插入和更好的词对齐。这对于非拉丁语的语言特别有用,这些语言具有丰富的构词法,可以标记杂乱或冗长的语音。其次,我们收集了一个新的测试集DiverseSpeech-Ru的长形式在野生俄语语音仔细多参考标记。我们还执行多参考重新标记流行的俄罗斯测试集和研究微调动力学在其相应的训练集。我们证明了该模型往往采用特定于特定于网络的标签,造成一种错觉的度量改进。基于改进的词对齐,我们开发的工具来评估流语音识别和对齐多个transmits进行视觉比较。此外,我们为许多离线和流式语音识别模型提供了统一的包装器。我们的代码将公开提供。
摘要:We propose several improvements to the speech recognition evaluation. First, we propose a string alignment algorithm that supports both multi-reference labeling, arbitrary-length insertions and better word alignment. This is especially useful for non-Latin languages, those with rich word formation, to label cluttered or longform speech. Secondly, we collect a novel test set DiverseSpeech-Ru of longform in-the-wild Russian speech with careful multi-reference labeling. We also perform multi-reference relabeling of popular Russian tests set and study fine-tuning dynamics on its corresponding train set. We demonstrate that the model often adopts to dataset-specific labeling, causing an illusion of metric improvement. Based on the improved word alignment, we develop tools to evaluate streaming speech recognition and to align multiple transcriptions to compare them visually. Additionally, we provide uniform wrappers for many offline and streaming speech recognition models. Our code will be made publicly available.
【17】Text-only adaptation in LLM-based ASR through text denoising
标题:通过文本去噪在基于LLM的ASR中进行纯文本自适应
链接:https://arxiv.org/abs/2601.20900
备注:Paper accepted at ICASSP 2026
摘要:基于大型语言模型(LLM)的自动语音识别(ASR)系统适应使用纯文本数据的新领域是一个重要但尚未探索的挑战。LLM对目标域文本的标准微调通常会破坏投影仪学习的语音和文本模态之间的关键对齐,从而降低性能。我们引入了一种新的纯文本自适应方法,通过将其视为文本去噪任务来模拟音频投影任务。因此,我们的方法训练LLM从嘈杂的输入中恢复干净的成绩单。该过程有效地使模型适应目标域,同时保持跨模态对齐。我们的解决方案是轻量级的,不需要架构更改或其他参数。对两个数据集的广泛评估表明,相对改善高达22.1%,优于最近最先进的纯文本自适应方法。
摘要:Adapting automatic speech recognition (ASR) systems based on large language models (LLMs) to new domains using text-only data is a significant yet underexplored challenge. Standard fine-tuning of the LLM on target-domain text often disrupts the critical alignment between speech and text modalities learned by the projector, degrading performance. We introduce a novel text-only adaptation method that emulates the audio projection task by treating it as a text denoising task. Our approach thus trains the LLM to recover clean transcripts from noisy inputs. This process effectively adapts the model to a target domain while preserving cross-modal alignment. Our solution is lightweight, requiring no architectural changes or additional parameters. Extensive evaluation on two datasets demonstrates up to 22.1% relative improvement, outperforming recent state-of-the-art text-only adaptation methods.
【18】A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
标题:预训练自我监督语音模型的数据选择策略研究
链接:https://arxiv.org/abs/2601.20896
备注:Accepted for publication in the 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2026)
摘要:自监督学习(SSL)已经改变了语音处理,但其对大量预训练数据集的依赖仍然是一个瓶颈。虽然鲁棒性通常归因于规模和多样性,但数据分布的作用却鲜为人知。我们系统地研究了预训练数据的精选子集如何影响自动语音识别(ASR)性能。令人惊讶的是,优化声学,扬声器或语言多样性并没有产生明显的改善随机采样。相反,我们发现,优先考虑最长的话语可以获得更好的ASR结果,同时只使用一半的原始数据集,在大型语料库上减少了24%的预训练时间。这些发现表明,对于预训练语音SSL模型,数据长度是比数据多样性或整体数据量更关键的因素,为SSL语音处理中的数据选择策略提供了新的视角。
摘要:Self-supervised learning (SSL) has transformed speech processing, yet its reliance on massive pre-training datasets remains a bottleneck. While robustness is often attributed to scale and diversity, the role of the data distribution is less understood. We systematically examine how curated subsets of pre-training data influence Automatic Speech Recognition (ASR) performance. Surprisingly, optimizing for acoustic, speaker, or linguistic diversity yields no clear improvements over random sampling. Instead, we find that prioritizing the longest utterances achieves superior ASR results while using only half the original dataset, reducing pre-training time by 24% on a large corpora. These findings suggest that for pre-training speech SSL models, data length is a more critical factor than either data diversity or overall data quantity for performance and efficiency, offering a new perspective for data selection strategies in SSL speech processing.
【19】SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
标题:SW-ASR:一种用于鲁棒单字语音识别的上下文感知混合ASR管道
链接:https://arxiv.org/abs/2601.20890
摘要:由于缺乏语言上下文以及对噪音、发音变化和通道伪影的敏感性,单字自动语音识别(ASR)是一项具有挑战性的任务,特别是在医疗保健和紧急响应等资源匮乏、通信关键的领域。本文回顾了最近的深度学习方法,并提出了一个用于鲁棒单字检测的模块化框架。该系统将去噪和归一化与混合ASR前端(Whisper + Vosk)和验证层相结合,旨在处理词汇表外的单词和降级的音频。验证层支持多种匹配策略,包括嵌入相似性,编辑距离和基于LLM的匹配以及可选的上下文指导。我们在Google Speech Commands数据集和从带宽受限条件下的电话和消息传递平台收集的真实数据集上评估了该框架。结果表明,虽然混合ASR前端在干净音频上表现良好,但验证层显着提高了噪声和压缩通道的准确性。上下文引导和基于LLM的匹配产生最大的收益,表明轻量级验证和上下文机制可以大大提高单字ASR的鲁棒性,而不会牺牲实时电话应用所需的延迟。
摘要:Single-word Automatic Speech Recognition (ASR) is a challenging task due to the lack of linguistic context and sensitivity to noise, pronunciation variation, and channel artifacts, especially in low-resource, communication-critical domains such as healthcare and emergency response. This paper reviews recent deep learning approaches and proposes a modular framework for robust single-word detection. The system combines denoising and normalization with a hybrid ASR front end (Whisper + Vosk) and a verification layer designed to handle out-of-vocabulary words and degraded audio. The verification layer supports multiple matching strategies, including embedding similarity, edit distance, and LLM-based matching with optional contextual guidance. We evaluate the framework on the Google Speech Commands dataset and a curated real-world dataset collected from telephony and messaging platforms under bandwidth-limited conditions. Results show that while the hybrid ASR front end performs well on clean audio, the verification layer significantly improves accuracy on noisy and compressed channels. Context-guided and LLM-based matching yield the largest gains, demonstrating that lightweight verification and context mechanisms can substantially improve single-word ASR robustness without sacrificing latency required for real-time telephony applications.
【20】VoxMorph: Scalable Zero-shot Voice Identity Morphing via Disentangled Embeddings
标题:VoxMorph:通过分解嵌入实现可扩展的Zero-Shot语音身份变形
链接:https://arxiv.org/abs/2601.20883
备注:Accepted to IEEE ICASSP 2026 (51st International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2026). 5 pages, 1 figure, 3 tables. Project page: https://vcbsl.github.io/VoxMorph/
摘要:变形技术生成人工生物特征样本,这些样本结合了来自多个个体的特征,允许每个贡献者根据单个注册模板进行验证。虽然在人脸识别中得到了广泛的研究,但在语音生物识别中,这种漏洞在很大程度上仍未被探索。语音变形的先前工作是计算昂贵的,不可扩展的,并限于声学相似的身份对,限制实际部署。此外,现有的声音变形方法针对音频纹理、音乐或环境声音,并且不能转移到语音身份操纵。我们提出了VoxMorph,一个zero-shot框架,它可以从每个主题的五秒音频中产生高保真的语音变体,而无需模型重新训练。我们的方法将声音特征分解为韵律和音色嵌入,从而实现说话风格和身份的细粒度插值。这些嵌入通过球面线性插值(Slerp)进行融合,并使用自回归语言模型与条件流匹配网络进行合成。VoxMorph实现了最先进的性能,在严格的安全阈值下,在自动说话人验证系统上提供了2.6倍的音频质量增益,73%的可理解性错误减少,以及67.8%的变形攻击成功率。这项工作建立了一个实用的和可扩展的范例语音变形与生物识别安全的重大影响。代码和数据集可以在我们的项目页面上找到:https://vcbsl.github.io/VoxMorph/
摘要:Morphing techniques generate artificial biometric samples that combine features from multiple individuals, allowing each contributor to be verified against a single enrolled template. While extensively studied in face recognition, this vulnerability remains largely unexplored in voice biometrics. Prior work on voice morphing is computationally expensive, non-scalable, and limited to acoustically similar identity pairs, constraining practical deployment. Moreover, existing sound-morphing methods target audio textures, music, or environmental sounds and are not transferable to voice identity manipulation. We propose VoxMorph, a zero-shot framework that produces high-fidelity voice morphs from as little as five seconds of audio per subject without model retraining. Our method disentangles vocal traits into prosody and timbre embeddings, enabling fine-grained interpolation of speaking style and identity. These embeddings are fused via Spherical Linear Interpolation (Slerp) and synthesized using an autoregressive language model coupled with a Conditional Flow Matching network. VoxMorph achieves state-of-the-art performance, delivering a 2.6x gain in audio quality, a 73% reduction in intelligibility errors, and a 67.8% morphing attack success rate on automated speaker verification systems under strict security thresholds. This work establishes a practical and scalable paradigm for voice morphing with significant implications for biometric security. The code and dataset are available on our project page: https://vcbsl.github.io/VoxMorph/
【21】Generalizable Prompt Tuning for Audio-Language Models via Semantic Expansion
标题:通过语义扩展对音频语言模型进行可概括的即时调优
链接:https://arxiv.org/abs/2601.20867
摘要:提示调整在视觉语言模型(VLM)中取得了显着的进展,最近正在被用于音频语言模型(ALM)。然而,它的泛化能力在ALMs仍然在很大程度上未被探索。我们观察到,传统的快速调整ALM也遭受新的基础权衡,我们确定这个问题源于嵌入空间的语义结构的破坏。为了解决这个问题,我们提出了语义扩展提示调优(SEPT)-一个即插即用的框架,明确规范化的提示嵌入空间,将大型语言模型生成的语义邻居。SEPT引入了一种新的语义扩展损失与利润限制,促进类内的紧凑性和类间的可分性,从而增强提示嵌入空间的语义结构。为了进行全面的评估,我们建立了第一个在ALM中进行快速泛化的基准设置,包括基础到新的泛化和跨数据集的可移植性。大量的实验表明,SEPT在多个提示调优基线上持续提高泛化性能,同时在推理过程中保持计算成本。代码可在https://github.com/jhyukjang/SEPT上找到。
摘要:Prompt tuning has achieved remarkable progress in vision-language models (VLMs) and is recently being adopted for audio-language models (ALMs). However, its generalization ability in ALMs remains largely underexplored. We observe that conventional prompt tuning for ALMs also suffers from the Base-New Tradeoff, and we identify that this issue stems from the disrupted semantic structure of the embedding space. To address this issue, we propose Semantically Expanded Prompt Tuning (SEPT)-a plug-and-play framework that explicitly regularizes the prompt embedding space by incorporating semantic neighbors generated by large language models. SEPT introduces a novel semantic expansion loss with margin constraints that promote intra-class compactness and inter-class separability, thereby enhancing the semantic structure of the prompt embedding space. For comprehensive evaluation, we establish the first benchmark setup for prompt generalization in ALMs, covering both base-to-new generalization and cross-dataset transferability. Extensive experiments demonstrate that SEPT consistently improves generalization performance across multiple prompt tuning baselines, while maintaining computational cost during inference. Codes are available in https://github.com/jhyukjang/SEPT.
机器翻译由腾讯交互翻译提供,仅供参考
