原场语音识别(Distant ASR)
原场语音识别中的挑战
- 混响
- 背景噪声
- 其他说话人的插入、打断等

远场语音识别系统

主要有这几个模块构成:
- 语音增强前端
- 特征提取
- 模型自适应
- 解码器
- 声学模型
- 字典
- 语言模型
下面主要讲加粗的模块中涉及到的相关技术。
未涉及以下方面:
- VAD
- Keyword spotting
- Multi-speaker / Speaker diarization / Speech Separation
- Online processing
语音增强前端
语音增强前端(SE Front-end)由三部分构成:
- 多通道去混响
- 多通道降噪
- 单通道降噪
做语音增强的目的是减小真实语音和ASR后端收到做过降噪或去混响的语音的不匹配程度。

处理方式:
- 线性
- 对长句采用固定线性滤波器。
- 非线性
- 做线性滤波器在每一帧都变化。
- 非线性变换。
技术总结如下:

去混响
常用方法:
- Linear filtering
- Weighted prediction error
- Non-linear filtering
- Spectral subtraction using a statistical model of late reverberation
- Neural network-based dereverberation
波束形成
做波束形成的目的:
- Pickup signals in the direction of the target speaker
- Attenuate signals in the direction of the noise sources

主流方法1:Delay and Sum beamformer

主流方法2:MVDR beamformer

其他方法
- Max-SNR beamformer
- Multi-channel Wiener filter
基于DNN的增强
基本结构:回归问题。训练一个能映射带噪语音到干净语音的网络。

目标函数
1.Regression based DNN,以MMSE作为目标函数。

clean speech feature (output)
- Log power spectrum
noisy speech feature (input)
- Log power spectrum + Context
network output
can be unbounded (i.e.,
, which is considered to be difficult
- Normalize the output by
- Use tanh() as an activation function
network parameters
2.Mask-estimation based DNN (Cross entropy)

3.Mask estimation based DNN (MMSE)

使用循环(神经网络)结构
- 简单RNN
- LSTM
进一步,用LSTM来做Mask估计。在CHiME2实验上相对于DNN有进一步的提升(29.7%->26.1%)。
Back-end techniques for distant ASR
提特征
特征变换方法:
- Linear Discriminant Analysis (LDA)
- Maximum Likelihood Linear Transformation (MLLT)
- Feature-space Maximum Likelihood Linear Regression (fMLLR)
- factored CLP(Google采用的方法)

鲁棒声学模型
常用的一些方法:
- Long frame context(L5R5)
- Sequence discriminative criterion(sMBR损失函数)
- Multi-task objectives(MMSE+CE双损失函数)
- 使用复杂模型
- TDNN
- CNN
- LSTM, Grid-LSTM, ...
- GRU, mGRP, ...
- resNet, skip-connection, ...
- End-to-End & Attention
- GAN
声学模型自适应
模型自适应方面
- Retraining
- Linear transformation of input or hidden layers (fDLR, LIN, LHN, LHUC)
- Adaptive training (Cluster/Speaker adaptive training)
- Teacher/Student
辅助特征
- Speaker aware (i-vector, Bottleneck feat.)
- Noise aware (noise estimate)
- Room aware (RT60, Distance, …)
结合前端与后端,联合训练网络模型


Data Simulation
由于缺少在真实场景中的远场数据,为了训练鲁棒声学模型,需要来生成一些模拟远场的数据来训练LVCSR模型。
常用方法:
- 加入混响(RIR)
- 改变信噪比
- 改变麦克风分布
- 改变声源位置、噪声位置、房间大小
