微信公众号:arXiv_Daily
cs.SD语音
【1】Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection
标题:研究在音频深度伪造检测环境中使用多模式大型语言模型的可行性
链接:https://arxiv.org/abs/2601.00777
备注:Accepted at IJCB 2025
摘要:虽然视觉语言模型(VLM)和多模态大型语言模型(MLLM)在检测图像和视频深度伪造方面表现出了很强的泛化能力,但它们在音频深度伪造检测中的应用在很大程度上尚未得到探索。在这项工作中,我们的目标是探索MLLM用于音频深度伪造检测的潜力。将音频输入与一系列文本提示相结合作为查询,以找出MLLM的可行性,从而学习跨模态的鲁棒表示,以进行音频deepfake检测。因此,我们试图探索文本感知和上下文丰富的,基于问答的提示与二元决策。我们假设,这种特征引导的推理将有助于促进更深入的多模态理解,并为音频deepfake检测提供强大的特征学习。我们评估两个MLLM,Qwen 2-Audio-7 B-Instruct和SALMONN的性能,在两种评估模式:(a)zero-shot和(b)微调。我们的实验表明,将音频与多提示方法相结合可能是音频deepfake检测的可行方法。我们的实验表明,在没有特定任务训练的情况下,模型的性能很差,并且很难推广到域外数据。然而,它们在域内数据上实现了良好的性能,并具有最小的监督,这表明音频深度伪造检测具有很大的潜力。
摘要:While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored. In this work, we aim to explore the potential of MLLMs for audio deepfake detection. Combining audio inputs with a range of text prompts as queries to find out the viability of MLLMs to learn robust representations across modalities for audio deepfake detection. Therefore, we attempt to explore text-aware and context-rich, question-answer based prompts with binary decisions. We hypothesise that such a feature-guided reasoning will help in facilitating deeper multimodal understanding and enable robust feature learning for audio deepfake detection. We evaluate the performance of two MLLMs, Qwen2-Audio-7B-Instruct and SALMONN, in two evaluation modes: (a) zero-shot and (b) fine-tuned. Our experiments demonstrate that combining audio with a multi-prompt approach could be a viable way forward for audio deepfake detection. Our experiments show that the models perform poorly without task-specific training and struggle to generalise to out-of-domain data. However, they achieve good performance on in-domain data with minimal supervision, indicating promising potential for audio deepfake detection.
【2】A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
标题:用于基于ATC的多语言ASB的分层LoRA-MoE架构
链接:https://arxiv.org/abs/2601.00557
备注:5 pages, submitted to IEEE Signal Processing Letters
摘要:Whisper等大规模多语言ASR(mASR)模型可实现强大的性能,但会产生较高的计算和延迟成本,从而限制了它们在资源受限的边缘设备上的部署。在这项研究中,我们提出了一个轻量级的和语言无关的多语言ASR系统的基础上CTC架构与域适应。具体来说,我们引入了一个与存储无关的分层LoRA-MoE(HLoRA)框架,该框架集成到mHuBERT-CTC模型中,通过LID后驱动的LoRA路由实现端到端解码。分层设计包括一个多语言共享LoRA,用于学习语言不变的声学表示和语言特定的LoRA专家,用于建模语言相关的特征。所提出的路由机制消除了在推理过程中对先前的语言身份信息或显式语言标签的需要,实现了真正的语言不可知解码。在MSR-86 K和MLC-SLM 2025 Challenge数据集上的实验表明,HLoRA仅使用单遍解码就可以实现最先进的两阶段推理方法,从而显著提高低资源mASR应用的解码效率。
摘要:Large-scale multilingual ASR (mASR) models such as Whisper achieve strong performance but incur high computational and latency costs, limiting their deployment on resource-constrained edge devices. In this study, we propose a lightweight and language-agnostic multilingual ASR system based on a CTC architecture with domain adaptation. Specifically, we introduce a Language-agnostic Hierarchical LoRA-MoE (HLoRA) framework integrated into an mHuBERT-CTC model, enabling end-to-end decoding via LID-posterior-driven LoRA routing. The hierarchical design consists of a multilingual shared LoRA for learning language-invariant acoustic representations and language-specific LoRA experts for modeling language-dependent characteristics. The proposed routing mechanism removes the need for prior language identity information or explicit language labels during inference, achieving true language-agnostic decoding. Experiments on MSR-86K and the MLC-SLM 2025 Challenge datasets demonstrate that HLoRA achieves competitive performance with state-of-the-art two-stage inference methods using only single-pass decoding, significantly improving decoding efficiency for low-resource mASR applications.
【3】MR-DAW: Towards Collaborative Digital Audio Workstations in Mixed Reality
标题:MR-M2 C:面向混合现实的协作数字音频工作站
链接:https://arxiv.org/abs/2601.00326
摘要:数字音频工作站(DW)是现代音乐制作的核心,但通常会阻碍音乐家的工作流程,将他们束缚在桌子上,阻碍与乐器的自然互动。此外,有效的远程协作仍然是一个重大挑战,现有解决方案受到网络延迟和异步文件共享的阻碍。本文探讨了混合现实(MR)克服这些障碍的潜力,为实时,远程音乐协作创建一个直观的环境。我们采用定性和推测性设计技术来更好地了解:1)玩家当前如何使用DW,以及2)想象协作MR-DW的推测性未来。为了促进这一讨论,我们开发和评估的设计探头,MR-100的可用性。一种MR系统,其使得多个地理上分散的用户能够控制单个共享的DRM实例,同时在其本地空间中自由移动。我们的网络系统使每个远程音乐家能够使用物理脚踏板进行协作循环,将熟悉的免提交互与共享的虚拟会话合并。基于20位音乐家的访谈和系统评估,我们分析了当前的做法,报告与我们的MR系统的用户体验,并推测未来的音乐合作在MR。我们的研究结果突出了无障碍的音乐互动的启示MR,并提供了一个推测性的展望未来的远程协作的DAWS在音乐元宇宙。
摘要:Digital Audio Workstations (DAWs) are central to modern music production but often encumber the musician's workflow, tethering them to a desk and hindering natural interaction with their instrument. Furthermore, effective remote collaboration remains a significant challenge, with existing solutions hampered by network latency and asynchronous file sharing. This paper investigates the potential of Mixed Reality (MR) to overcome these barriers, creating an intuitive environment for real-time, remote musical collaboration. We employ qualitative and speculative design techniques to better understand: 1) how players currently use DAWs, and 2) to imagine a speculative future of collaborative MR-DAWs. To facilitate this discussion, we developed and evaluated the usability of a design probe, MR-DAW. An MR system enabling multiple, geographically dispersed users to control a single, shared DAW instance while moving freely in their local spaces. Our networked system enables each remote musician to use a physical foot pedal for collaborative looping, merging a familiar, hands-free interaction with a shared virtual session. Based on interviews and system evaluations with 20 musicians, we analyze current practices, report on the user experience with our MR system, and speculate on the future of musical collaboration in MR. Our results highlight the affordances of MR for unencumbered musical interaction and provide a speculative outlook on the future of remote collaborative DAWs in the Musical Metaverse.
【4】Timed text extraction from Taiwanese Kua-á-hì TV series
标题:台湾夸人电视剧的定时文本提取
链接:https://arxiv.org/abs/2601.00299
备注:Accepted to ISMIR 2025 Late-Breaking Demo (LBD)
摘要:歌仔戏是当地戏剧传统的主要形式,经过广泛的电视改编,尤其是像伊莲·吕华这样的先驱。这些影片,虽然对深入研究歌仔戏有潜在的价值,但往往质量不高,在数据准备过程中需要大量的手工劳动。为了简化这一过程,我们开发了一个交互式系统,用于实时OCR校正和两步方法,将OCR驱动的分割与语音和音乐活动检测(SMAD)集成在一起,以高精度有效地从存档剧集中识别声乐片段。由此产生的数据集,包括声乐片段和相应的歌词,可以潜在地支持各种MIR任务,如歌词识别和曲调检索。代码可在https://github.com/z-huang/ocr-subtitle-editor上获得。
摘要:Taiwanese opera (Kua-á-hì), a major form of local theatrical tradition, underwent extensive television adaptation notably by pioneers like Iûnn Lē-hua. These videos, while potentially valuable for in-depth studies of Taiwanese opera, often have low quality and require substantial manual effort during data preparation. To streamline this process, we developed an interactive system for real-time OCR correction and a two-step approach integrating OCR-driven segmentation with Speech and Music Activity Detection (SMAD) to efficiently identify vocal segments from archival episodes with high precision. The resulting dataset, consisting of vocal segments and corresponding lyrics, can potentially supports various MIR tasks such as lyrics identification and tune retrieval. Code is available at https://github.com/z-huang/ocr-subtitle-editor .
【5】Latent Flow Matching for Expressive Singing Voice Synthesis
标题:表达性歌唱声音合成的潜在流匹配
链接:https://arxiv.org/abs/2601.00217
摘要:基于条件变分自编码器(cVAE)的歌唱声音合成通过学习分数条件先验和记录条件后验潜在空间来提供有效的推理和强音频质量。然而,由于合成依赖于先前的样本,而训练使用从真实记录推断的后验潜伏期,不完美的分布匹配可能导致先验-后验失配,从而降低细粒度的表现力,例如颤音和微韵律。我们提出了FM-Singer,它在潜在空间中引入条件流匹配(CFM)来学习一个连续的向量场,该向量场沿着最优传输启发的路径将先前的潜在流传输到后一个潜在流。在推理时,学习的潜在流通过在波形生成之前求解常微分方程(ODE)来细化先前的样本,从而在保持并行解码效率的同时提高表达能力。在韩国和中国歌唱数据集上的实验表明,在强基线上有一致的改进,包括在韩国数据集上较低的梅尔倒频谱失真和基频误差以及较高的感知分数。代码、预先训练的检查点和音频演示可在https://github.com/alsgur9368/FM-Singer上获得
摘要:Conditional variational autoencoder (cVAE)-based singing voice synthesis provides efficient inference and strong audio quality by learning a score-conditioned prior and a recording-conditioned posterior latent space. However, because synthesis relies on prior samples while training uses posterior latents inferred from real recordings, imperfect distribution matching can cause a prior-posterior mismatch that degrades fine-grained expressiveness such as vibrato and micro-prosody. We propose FM-Singer, which introduces conditional flow matching (CFM) in latent space to learn a continuous vector field transporting prior latents toward posterior latents along an optimal-transport-inspired path. At inference time, the learned latent flow refines a prior sample by solving an ordinary differential equation (ODE) before waveform generation, improving expressiveness while preserving the efficiency of parallel decoding. Experiments on Korean and Chinese singing datasets demonstrate consistent improvements over strong baselines, including lower mel-cepstral distortion and fundamental-frequency error and higher perceptual scores on the Korean dataset. Code, pretrained checkpoints, and audio demos are available at https://github.com/alsgur9368/FM-Singer
【6】IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
标题:IKFST:IOO和KOO算法用于加速和精确的基于WFST的端到端自动语音识别
链接:https://arxiv.org/abs/2601.00160
摘要:端到端自动语音识别已经成为学术界和工业界的主导范式。为了提高识别性能,广泛采用了加权语音状态转换器(WFST),通过静态图合成来集成声学和语言模型,提供鲁棒的解码和有效的纠错。然而,WFST解码依赖于CTC后验概率上的逐帧自回归搜索,这严重限制了推理效率。出于建立WFST解码和CTC建模之间更原则的兼容性,我们系统地研究了CTC输出的两个基本组成部分,即空白和非空白帧,并确定了一个关键的见解:空白帧主要编码位置信息,而非空白帧携带语义内容。在此基础上,我们引入了Keep-Only-One和Insert-Only-One,这两种解码算法明确利用空白和非空白帧的结构作用,以实现更快的基于WFST的推理,而不会影响识别精度。在大规模内部、AISHELL-1和LibriSpeech数据集上进行的实验证明了最先进的识别准确性,同时大幅降低了解码延迟,从而在现代语音识别系统中实现真正高效和高性能的WFST解码。
摘要:End-to-end automatic speech recognition has become the dominant paradigm in both academia and industry. To enhance recognition performance, the Weighted Finite-State Transducer (WFST) is widely adopted to integrate acoustic and language models through static graph composition, providing robust decoding and effective error correction. However, WFST decoding relies on a frame-by-frame autoregressive search over CTC posterior probabilities, which severely limits inference efficiency. Motivated by establishing a more principled compatibility between WFST decoding and CTC modeling, we systematically study the two fundamental components of CTC outputs, namely blank and non-blank frames, and identify a key insight: blank frames primarily encode positional information, while non-blank frames carry semantic content. Building on this observation, we introduce Keep-Only-One and Insert-Only-One, two decoding algorithms that explicitly exploit the structural roles of blank and non-blank frames to achieve significantly faster WFST-based inference without compromising recognition accuracy. Experiments on large-scale in-house, AISHELL-1, and LibriSpeech datasets demonstrate state-of-the-art recognition accuracy with substantially reduced decoding latency, enabling truly efficient and high-performance WFST decoding in modern speech recognition systems.
【1】Learning Speech Representations with Variational Predictive Coding
标题:使用变分预测编码学习语音表示
链接:https://arxiv.org/abs/2601.00100
备注:Accepted to Transactions of the Association for Computational Linguistics (TACL); Pre MIT Press version
摘要:尽管HuBERT目标是学习语音表示的最著名的目标,但它并没有得到进一步的发展和改进。我们认为,这是缺乏一个基本的原则,阻碍了发展,在本文中,我们表明,预测编码下的变分的观点是背后的原则,休伯特目标。由于其通用性,我们的配方提供了机会,以改善参数化和优化,我们展示了两个简单的修改,带来立即改善的休伯特目标。此外,预测编码公式与各种其他目标(诸如APC、CPC、wav 2 vec和BEST-RQ)有紧密联系。从经验上讲,预训练的改进为四个下游任务带来了显着的改进:电话分类,f0跟踪,说话人识别和自动语音识别,突出了预测编码解释的重要性。
摘要:Despite being the best known objective for learning speech representations, the HuBERT objective has not been further developed and improved. We argue that it is the lack of an underlying principle that stalls the development, and, in this paper, we show that predictive coding under a variational view is the principle behind the HuBERT objective. Due to its generality, our formulation provides opportunities to improve parameterization and optimization, and we show two simple modifications that bring immediate improvements to the HuBERT objective. In addition, the predictive coding formulation has tight connections to various other objectives, such as APC, CPC, wav2vec, and BEST-RQ. Empirically, the improvement in pre-training brings significant improvements to four downstream tasks: phone classification, f0 tracking, speaker recognition, and automatic speech recognition, highlighting the importance of the predictive coding interpretation.
【2】Neural Brain Fields: A NeRF-Inspired Approach for Generating Nonexistent EEG Electrodes
标题:Neural Brain Fields:A NeRF Inspired Approach for Generating Nonexisting EEG Electrodes神经脑场:一种基于NeRF的生成不存在EEG电极的方法
链接:https://arxiv.org/abs/2601.00012
摘要:脑电图(EEG)数据提出了独特的建模挑战,因为记录长度不同,信噪比非常低,参与者之间存在显著差异,会话内随时间推移而漂移,并且很少在大型和干净的数据集中可用。因此,开发能够有效处理EEG信号的深度学习方法仍然是一个开放且重要的研究问题。为了解决这个问题,这项工作提出了一种受神经辐射场(NeRF)启发的新方法。在计算机视觉中,NeRF技术训练神经网络来记忆3D场景的外观,然后使用其学习的参数从任何视点渲染和编辑场景。我们在从不同视角捕获的离散图像之间进行类比,这些图像用于学习NeRF中的连续3D场景,而EEG电极位于头皮上的不同位置,用于推断连续神经活动的潜在表示。基于这种连接,我们证明了神经网络可以以NeRF风格的方式在单个EEG样本上进行训练,以产生一个固定大小和信息量的权重向量,对整个信号进行编码。此外,通过这种表示,我们可以在以前看不见的时间步长和空间电极位置呈现EEG信号。我们证明,这种方法能够以任何所需的分辨率(包括超高分辨率)连续可视化大脑活动,并重建原始脑电信号。最后,我们的实证分析表明,该方法可以有效地模拟不存在的电极数据在EEG记录,允许重建的信号被送入标准的EEG处理网络,以提高性能。
摘要:Electroencephalography (EEG) data present unique modeling challenges because recordings vary in length, exhibit very low signal to noise ratios, differ significantly across participants, drift over time within sessions, and are rarely available in large and clean datasets. Consequently, developing deep learning methods that can effectively process EEG signals remains an open and important research problem. To tackle this problem, this work presents a new method inspired by Neural Radiance Fields (NeRF). In computer vision, NeRF techniques train a neural network to memorize the appearance of a 3D scene and then uses its learned parameters to render and edit the scene from any viewpoint. We draw an analogy between the discrete images captured from different viewpoints used to learn a continuous 3D scene in NeRF, and EEG electrodes positioned at different locations on the scalp, which are used to infer the underlying representation of continuous neural activity. Building on this connection, we show that a neural network can be trained on a single EEG sample in a NeRF style manner to produce a fixed size and informative weight vector that encodes the entire signal. Moreover, via this representation we can render the EEG signal at previously unseen time steps and spatial electrode positions. We demonstrate that this approach enables continuous visualization of brain activity at any desired resolution, including ultra high resolution, and reconstruction of raw EEG signals. Finally, our empirical analysis shows that this method can effectively simulate nonexistent electrodes data in EEG recordings, allowing the reconstructed signal to be fed into standard EEG processing networks to improve performance.
【3】A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
标题:用于基于ATC的多语言ASB的分层LoRA-MoE架构
链接:https://arxiv.org/abs/2601.00557
备注:5 pages, submitted to IEEE Signal Processing Letters
摘要:Whisper等大规模多语言ASR(mASR)模型可实现强大的性能,但会产生较高的计算和延迟成本,从而限制了它们在资源受限的边缘设备上的部署。在这项研究中,我们提出了一个轻量级的和语言无关的多语言ASR系统的基础上CTC架构与域适应。具体来说,我们引入了一个集成到mHuBERT-CTC模型中的语言不可知分层LoRA-MoE(HLoRA)框架,通过LID后验驱动的LoRA路由实现端到端解码。分层设计包括一个多语言共享LoRA,用于学习语言不变的声学表示和语言特定的LoRA专家,用于建模语言相关的特征。所提出的路由机制消除了在推理过程中对先前的语言身份信息或显式语言标签的需要,实现了真正的语言不可知解码。在MSR-86 K和MLC-SLM 2025 Challenge数据集上的实验表明,HLoRA仅使用单遍解码就可以实现最先进的两阶段推理方法,从而显著提高低资源mASR应用的解码效率。
摘要:Large-scale multilingual ASR (mASR) models such as Whisper achieve strong performance but incur high computational and latency costs, limiting their deployment on resource-constrained edge devices. In this study, we propose a lightweight and language-agnostic multilingual ASR system based on a CTC architecture with domain adaptation. Specifically, we introduce a Language-agnostic Hierarchical LoRA-MoE (HLoRA) framework integrated into an mHuBERT-CTC model, enabling end-to-end decoding via LID-posterior-driven LoRA routing. The hierarchical design consists of a multilingual shared LoRA for learning language-invariant acoustic representations and language-specific LoRA experts for modeling language-dependent characteristics. The proposed routing mechanism removes the need for prior language identity information or explicit language labels during inference, achieving true language-agnostic decoding. Experiments on MSR-86K and the MLC-SLM 2025 Challenge datasets demonstrate that HLoRA achieves competitive performance with state-of-the-art two-stage inference methods using only single-pass decoding, significantly improving decoding efficiency for low-resource mASR applications.
【4】MR-DAW: Towards Collaborative Digital Audio Workstations in Mixed Reality
标题:MR-M2 C:面向混合现实的协作数字音频工作站
链接:https://arxiv.org/abs/2601.00326
摘要:数字音频工作站(DW)是现代音乐制作的核心,但通常会阻碍音乐家的工作流程,将他们束缚在桌子上,阻碍与乐器的自然互动。此外,有效的远程协作仍然是一个重大挑战,现有解决方案受到网络延迟和异步文件共享的阻碍。本文探讨了混合现实(MR)克服这些障碍的潜力,为实时,远程音乐协作创建一个直观的环境。我们采用定性和推测性的设计技术来更好地理解:1)玩家目前如何使用DAO,以及2)想象协作MR-Daws的推测性未来。为了促进这一讨论,我们开发和评估的设计探头,MR-100的可用性。一种MR系统,其使得多个地理上分散的用户能够控制单个共享的DRM实例,同时在其本地空间中自由移动。我们的网络系统使每个远程音乐家能够使用物理脚踏板进行协作循环,将熟悉的免提交互与共享的虚拟会话合并。基于20位音乐家的访谈和系统评估,我们分析了目前的做法,报告与我们的MR系统的用户体验,并推测未来的音乐合作在MR。我们的研究结果突出了无障碍的音乐互动的启示MR,并提供了一个推测性的展望未来的远程协作的DAWS在音乐元宇宙。
摘要:Digital Audio Workstations (DAWs) are central to modern music production but often encumber the musician's workflow, tethering them to a desk and hindering natural interaction with their instrument. Furthermore, effective remote collaboration remains a significant challenge, with existing solutions hampered by network latency and asynchronous file sharing. This paper investigates the potential of Mixed Reality (MR) to overcome these barriers, creating an intuitive environment for real-time, remote musical collaboration. We employ qualitative and speculative design techniques to better understand: 1) how players currently use DAWs, and 2) to imagine a speculative future of collaborative MR-DAWs. To facilitate this discussion, we developed and evaluated the usability of a design probe, MR-DAW. An MR system enabling multiple, geographically dispersed users to control a single, shared DAW instance while moving freely in their local spaces. Our networked system enables each remote musician to use a physical foot pedal for collaborative looping, merging a familiar, hands-free interaction with a shared virtual session. Based on interviews and system evaluations with 20 musicians, we analyze current practices, report on the user experience with our MR system, and speculate on the future of musical collaboration in MR. Our results highlight the affordances of MR for unencumbered musical interaction and provide a speculative outlook on the future of remote collaborative DAWs in the Musical Metaverse.
【5】Latent Flow Matching for Expressive Singing Voice Synthesis
标题:表达性歌唱声音合成的潜在流匹配
链接:https://arxiv.org/abs/2601.00217
摘要:基于条件变分自编码器(cVAE)的歌唱声音合成通过学习分数条件先验和记录条件后验潜在空间来提供有效的推理和强音频质量。然而,由于合成依赖于先前的样本,而训练使用从真实记录推断的后验潜伏期,不完美的分布匹配可能导致先验-后验失配,从而降低细粒度的表现力,例如颤音和微韵律。我们提出了FM-Singer,它在潜在空间中引入条件流匹配(CFM)来学习一个连续的向量场,该向量场沿着最优传输启发的路径将先前的潜在流传输到后一个潜在流。在推理时,学习的潜在流通过在波形生成之前求解常微分方程(ODE)来细化先前的样本,从而在保持并行解码效率的同时提高表达能力。在韩国和中国歌唱数据集上的实验表明,在强基线上有一致的改进,包括在韩国数据集上较低的梅尔倒频谱失真和基频误差以及较高的感知分数。代码、预先训练的检查点和音频演示可在https://github.com/alsgur9368/FM-Singer上获得
摘要:Conditional variational autoencoder (cVAE)-based singing voice synthesis provides efficient inference and strong audio quality by learning a score-conditioned prior and a recording-conditioned posterior latent space. However, because synthesis relies on prior samples while training uses posterior latents inferred from real recordings, imperfect distribution matching can cause a prior-posterior mismatch that degrades fine-grained expressiveness such as vibrato and micro-prosody. We propose FM-Singer, which introduces conditional flow matching (CFM) in latent space to learn a continuous vector field transporting prior latents toward posterior latents along an optimal-transport-inspired path. At inference time, the learned latent flow refines a prior sample by solving an ordinary differential equation (ODE) before waveform generation, improving expressiveness while preserving the efficiency of parallel decoding. Experiments on Korean and Chinese singing datasets demonstrate consistent improvements over strong baselines, including lower mel-cepstral distortion and fundamental-frequency error and higher perceptual scores on the Korean dataset. Code, pretrained checkpoints, and audio demos are available at https://github.com/alsgur9368/FM-Singer
【6】IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
标题:IKFST:IOO和KOO算法用于加速和精确的基于WFST的端到端自动语音识别
链接:https://arxiv.org/abs/2601.00160
摘要:端到端自动语音识别已经成为学术界和工业界的主导范式。为了提高识别性能,广泛采用了加权语音状态转换器(WFST),通过静态图合成来集成声学和语言模型,提供鲁棒的解码和有效的纠错。然而,WFST解码依赖于CTC后验概率上的逐帧自回归搜索,这严重限制了推理效率。出于建立WFST解码和CTC建模之间更原则的兼容性,我们系统地研究了CTC输出的两个基本组成部分,即空白和非空白帧,并确定了一个关键的见解:空白帧主要编码位置信息,而非空白帧携带语义内容。在此基础上,我们引入了Keep-Only-One和Insert-Only-One,这两种解码算法明确利用空白和非空白帧的结构作用,以实现更快的基于WFST的推理,而不会影响识别精度。在大规模内部、AISHELL-1和LibriSpeech数据集上进行的实验证明了最先进的识别准确性,同时大幅降低了解码延迟,从而在现代语音识别系统中实现真正高效和高性能的WFST解码。
摘要:End-to-end automatic speech recognition has become the dominant paradigm in both academia and industry. To enhance recognition performance, the Weighted Finite-State Transducer (WFST) is widely adopted to integrate acoustic and language models through static graph composition, providing robust decoding and effective error correction. However, WFST decoding relies on a frame-by-frame autoregressive search over CTC posterior probabilities, which severely limits inference efficiency. Motivated by establishing a more principled compatibility between WFST decoding and CTC modeling, we systematically study the two fundamental components of CTC outputs, namely blank and non-blank frames, and identify a key insight: blank frames primarily encode positional information, while non-blank frames carry semantic content. Building on this observation, we introduce Keep-Only-One and Insert-Only-One, two decoding algorithms that explicitly exploit the structural roles of blank and non-blank frames to achieve significantly faster WFST-based inference without compromising recognition accuracy. Experiments on large-scale in-house, AISHELL-1, and LibriSpeech datasets demonstrate state-of-the-art recognition accuracy with substantially reduced decoding latency, enabling truly efficient and high-performance WFST decoding in modern speech recognition systems.
机器翻译由腾讯交互翻译提供,仅供参考
