原场语音识别(Distant ASR)

原场语音识别中的挑战


  • 混响
  • 背景噪声
  • 其他说话人的插入、打断等



远场语音识别系统


    主要有这几个模块构成:

    • 语音增强前端
    • 特征提取
    • 模型自适应
    • 解码器
      • 声学模型
      • 字典
      • 语言模型

    下面主要讲加粗的模块中涉及到的相关技术。

    未涉及以下方面:

    • VAD
    • Keyword spotting
    • Multi-speaker / Speaker diarization / Speech Separation
    • Online processing


    语音增强前端

      语音增强前端(SE Front-end)由三部分构成:

      1. 多通道去混响
      2. 多通道降噪
      3. 单通道降噪

      做语音增强的目的是减小真实语音和ASR后端收到做过降噪或去混响的语音的不匹配程度。


        处理方式:

        • 线性
          • 对长句采用固定线性滤波器。[公式]
          • 非线性
            • 做线性滤波器在每一帧都变化。[公式]
          • 非线性变换。[公式]

        技术总结如下:



        去混响

          常用方法:

          • Linear filtering
          • Weighted prediction error
          • Non-linear filtering
          • Spectral subtraction using a statistical model of late reverberation
          • Neural network-based dereverberation


          波束形成

          做波束形成的目的:

          • Pickup signals in the direction of the target speaker
          • Attenuate signals in the direction of the noise sources



          主流方法1:Delay and Sum beamformer


          主流方法2:MVDR beamformer


            其他方法

            • Max-SNR beamformer
            • Multi-channel Wiener filter

            基于DNN的增强

            基本结构:回归问题。训练一个能映射带噪语音到干净语音的网络。



            目标函数

            1.Regression based DNN,以MMSE作为目标函数。


            [公式]

            • [公式]clean speech feature (output)
            • Log power spectrum
            • [公式]noisy speech feature (input)
            • Log power spectrum + Context
            • [公式]network output
            • [公式]can be unbounded (i.e.,[公式], which is considered to be difficult
            • Normalize the output by[公式]
            • Use tanh() as an activation function
            • [公式]network parameters


            2.Mask-estimation based DNN (Cross entropy)


            3.Mask estimation based DNN (MMSE)


              使用循环(神经网络)结构

              • 简单RNN
              • LSTM

              进一步,用LSTM来做Mask估计。在CHiME2实验上相对于DNN有进一步的提升(29.7%->26.1%)。


              Back-end techniques for distant ASR


              提特征

                特征变换方法:

                • Linear Discriminant Analysis (LDA)
                • Maximum Likelihood Linear Transformation (MLLT)
                • Feature-space Maximum Likelihood Linear Regression (fMLLR)
                • factored CLP(Google采用的方法)


                鲁棒声学模型

                常用的一些方法:

                • Long frame context(L5R5)
                • Sequence discriminative criterion(sMBR损失函数)
                • Multi-task objectives(MMSE+CE双损失函数)
                • 使用复杂模型
                • TDNN
                • CNN
                • LSTM, Grid-LSTM, ...
                • GRU, mGRP, ...
                • resNet, skip-connection, ...
                • End-to-End & Attention
                • GAN


                声学模型自适应


                  模型自适应方面

                  • Retraining
                  • Linear transformation of input or hidden layers (fDLR, LIN, LHN, LHUC)
                  • Adaptive training (Cluster/Speaker adaptive training)
                  • Teacher/Student


                    辅助特征

                    • Speaker aware (i-vector, Bottleneck feat.)
                    • Noise aware (noise estimate)
                    • Room aware (RT60, Distance, …)


                    结合前端与后端,联合训练网络模型



                    Data Simulation

                    由于缺少在真实场景中的远场数据,为了训练鲁棒声学模型,需要来生成一些模拟远场的数据来训练LVCSR模型。


                    常用方法:

                    • 加入混响(RIR)
                    • 改变信噪比
                    • 改变麦克风分布
                    • 改变声源位置、噪声位置、房间大小