
标题:CTC Variations Through New WFST Topologies
链接:https://arxiv.org/pdf/2110.03098.pdf
作者:Aleksandr Laptev, Somshubra Majumdar, Boris Ginsburg
摘要:This paper presents novel Weighted Finite-State Transducer (WFST) topologies to implement Connectionist Temporal Classification (CTC)-like algorithms for automatic speech recognition. Three new CTC variants are proposed: (1) the “compactCTC”, in which direct transitions between units are replaced with hi back-off transitions; (2) the “minimal-CTC”, that only adds hblanki self-loops when used in WFST-composition; and (3) the “selfless-CTC” variants, which disallows self-loop for non-blank units. Compact-CTC allows for 1.5 times smaller WFST decoding graphs and reduces memory consumption by two times when training CTC models with the LF-MMI objective without hurting the recognition accuracy. Minimal-CTC reduces graph size and memory consumption by two and four times for the cost of a small accuracy drop. Using selfless-CTC can improve the accuracy for wide context window models.
标题:Consistent Training and Decoding for End-to-End Speech Recognition Using Lattice-Free MMI
链接:https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=9746579
代码:https://github.com/jctian98/e2e_lfmmi
作者:Jinchuan Tian、Jianwei Yu、Chao Weng、Shi-Xiong Zhang、Dan Su、Dong Yu、Yuexian Zou
摘要:本文由腾讯AI Lab主导,与北京大学合作完成。近年来,端到端语音识别系统在多项语音识别任务上取得长足进展。然而,广泛应用于Hybrid语音识别系统的LF-MMI训练准则极少被应用于端到端语音识别系统中。本文提出一种新的方法,将LF-MMI准则应用于端到端语音识别系统的训练和解码中。该方法的有效性在两种常见的端到端语音识别系统(Attention-based Encoder-Decoder和Neural Transducer)上均得到验证。实验表明,在端到端语音识别系统中引入LF-MMI准则能为两种端到端语音识别系统带来一致性性能提升。该方法最佳模型在Aishell-1的dev/test集上实现了4.1%/4.4%的字错误率(CER),在其他数据集Aishell-2和Librispeech上也实现了显著提升。标题:Improving Mandarin End-to-End Speech Recognition with Word N-gram Language Model
链接:https://arxiv.org/pdf/2201.01995.pdf
作者:Jinchuan Tian, Jianwei Yu, Chao Weng, Yuexian Zou, Senior Member, Dong Yu
摘要:Despite the rapid progress of end-to-end (E2E) automatic speech recognition (ASR), it has been shown that incorporating external language models (LMs) into the decoding can further improve the recognition performance of E2E ASR systems. To align with the modeling units adopted in E2E ASR systems, subword-level (e.g., characters, BPE) LMs are usually used to cooperate with current E2E ASR systems. However, the use of subword-level LMs will ignore the word-level information, which may limit the strength of the external LMs in E2E ASR. Although several methods have been proposed to incorporate word-level external LMs in E2E ASR, these methods are mainly designed for languages with clear word boundaries such as English and cannot be directly applied to languages like Mandarin, in which each character sequence can have multiple corresponding word sequences. To this end, we propose a novel decoding algorithm where a word-level lattice is constructed onthe-fly to consider all possible word sequences for each partial hypothesis. Then, the LM score of the hypothesis is obtained by intersecting the generated lattice with an external word Ngram LM. The proposed method is examined on both Attentionbased Encoder-Decoder (AED) and Neural Transducer (NT) frameworks. Experiments suggest that our method consistently outperforms subword-level LMs, including N-gram LM and neural network LM. We achieve state-of-the-art results on both Aishell-1 (CER 4.18%) and Aishell-2 (CER 5.06%) datasets and reduce CER by 14.8% relatively on a 21K-hour Mandarin dataset.标题:Integrate Lattice-Free MMI into End-to-End Speech Recognition
链接:https://arxiv.org/pdf/2203.15614.pdf
作者:Jinchuan Tian, Student Member, Jianwei Yu, Chao Weng, Yuexian Zou, Senior Member,Dong Yu,
摘要:In automatic speech recognition (ASR) research, discriminative criteria have achieved superior performance in DNN-HMM systems. Given this success, the adoption of discriminative criteria is promising to boost the performance of end-to-end (E2E) ASR systems. With this motivation, previous works have introduced the minimum Bayesian risk (MBR, one of the discriminative criteria) into E2E ASR systems. However, the effectiveness and efficiency of the MBR-based methods are compromised: the MBR criterion is only used in system training, which creates a mismatch between training and decoding; the onthe-fly decoding process in MBR-based methods results in the need for pre-trained models and slow training speeds. To this end, novel algorithms are proposed in this work to integrate another widely used discriminative criterion, lattice-free maximum mutual information (LF-MMI), into E2E ASR systems not only in the training stage but also in the decoding process. The proposed LF-MMI training and decoding methods show their effectiveness on two widely used E2E frameworks: Attention-Based EncoderDecoders (AEDs) and Neural Transducers (NTs). Compared with MBR-based methods, the proposed LF-MMI method: maintains the consistency between training and decoding; eschews the on-the-fly decoding process; trains from randomly initialized models with superior training efficiency. Experiments suggest that the LF-MMI method outperforms its MBR counterparts and consistently leads to statistically significant performance improvements on various frameworks and datasets from 30 hours to 14.3k hours. The proposed method achieves state-of-the-art (SOTA) results on Aishell-1 (CER 4.10%) and Aishell-2 (CER 5.02%) datasets. Code is released1 .
https://github.com/k2-fsa/k2https://github.com/lhotse-speech/lhotsehttps://github.com/k2-fsa/icefallhttps://github.com/k2-fsa/sherpa