作者吴晓龙

ICASSP (International Conference on Acoustics, Speech and Signal Processing) 即国际声学、语音与信号处理会议,是IEEE主办的全世界最大、最全面的信号处理及其应用方面的顶级会议,在国际上享有盛誉并具有广泛的学术影响力。据我们统计,今年入选 ICASSP 2022 的论文中,说话人识别(声纹识别)方向约有56篇,初步划分为Speaker Verification(28篇)、Speaker Recognition(6篇)、Speaker Representation Learning(8篇)、Speaker Diarization(8篇)、Others(6篇)等五种类型。


以下为Speaker Verification部分28篇的论文简述。
1.Attention Back-End For Automatic Speaker Verification With Multiple Enrollment Utterances
标题:用于多注册语音自动说话人验证的注意力后端

作者:Chang Zeng1;2, Xin Wang1, Erica Cooper1, Xiaoxiao Miao1, Junichi Yamagishi1;2

单位:1National Institute of Informatics, Japan

2SOKENDAI, Japan

链接:https://arxiv.org/abs/2104.01541

摘要:概率线性鉴别分析(PLDA)或余弦相似度在传统的说话人确认系统中被广泛用作衡量成对相似度的后端技术。为了更好地利用多个注册语音,我们提出了一种新的注意后端模型,该模型可用于文本无关(TI)和文本相关(TD)说话人确认,并且我们使用缩放点积自注意力结构(scaled-dot self-attention)与前馈自注意力结构(feed-forward self-attention)作为学习注册话语内部关系的架构。为了验证所提出的模型,我们在CNCeleb和VoxCeleb数据集上进行了一系列实验,将其与包括TDNN和ResNet在内的几种最先进的扬声器编码器相结合。在CNCeleb上使用多个注册话语获得的实验结果表明,对于每个说话人编码器,所提出的注意后端模型的EER和minDCF分数低于其PLDA和余弦相似性对应项,并且在VoxCeleb上的实验表明,我们的模型甚至可以用于单个注册语音的情况。

Probabilistic linear discriminant analysis (PLDA) or cosine similarity have been widely used in traditional speaker verification systems as back-end techniques to measure pairwise similarities. To make better use of multiple enrollment utterances, we propose a novel attention back-end model that can be used for both textindependent (TI) and text-dependent (TD) speaker verification, and we use scaled-dot self-attention and feed-forward self-attention networks as architectures that learn the intra-relationships of enrollment utterances. To verify the proposed model, we conduct a series of experiments on the CNCeleb and VoxCeleb datasets by combining it with several state-of-the-art speaker encoders including TDNN and ResNet. Experimental results obtained using multiple enrollment utterances on CNCeleb show that the proposed attention back-end model leads to lower EER and minDCF scores than its PLDA and cosine similarity counterparts for each speaker encoder, and an experiment on VoxCeleb demonstrates that our model can be used even for a single enrollment case.


2.Simple Attention Module based Speaker Verification with Iterative noisy label detection

标题:基于简单注意模块的迭代噪声标签检测技术下的说话人验证

作者:Xiaoyi Qin1,2, Na Li3, Chao Weng3, Dan Su3, Ming Li1,2

单位:

1School of Computer Science, Wuhan University, Wuhan, China

2Data Science Research Center, Duke Kunshan University, Kunshan, China

3Tencent AI Lab, Shenzhen, China

链接:https://arxiv.org/abs/2110.06534

摘要:近年来,在基于深度学习的说话人确认系统中,诸如压缩和激励模块(SE)和卷积块注意模块(CBAM)等注意机制取得了巨大的成功。本文介绍了一种有效而简单的说话人验证方法,即简单注意模块(SimAM)。SimAM模块是一个即插即用模块,没有额外的模态参数。此外,考虑到带有人工标注或其他自动化过程的大规模数据集可能包含噪声标签,我们提出了一种噪声标签检测方法,以迭代方式从训练数据中过滤出带有噪声标签的数据样本。带有噪声标签的数据可能会过度参数化深度神经网络(DNN),并由于DNN的记忆效应导致性能下降。在VoxCeleb数据集上进行了实验。

使用SimAM的说话人验证模型在VoxCeleb1原始测试中达到了0.675%的等错误率(EER)。我们提出的迭代噪声标签检测方法进一步将EER降低到0.643%。

Recently, the attention mechanism such as squeeze-andexcitation module (SE) and convolutional block attention module (CBAM) has achieved great success in deep learningbased speaker verification system. This paper introduces an alternative effective yet simple one, i.e., simple attention module (SimAM), for speaker verification. The SimAM module is a plug-and-play module without extra modal parameters. In addition, we propose a noisy label detection method to iteratively filter out the data samples with a noisy label from the training data, considering that a large-scale dataset labeled with human annotation or other automated processes may contain noisy labels. Data with the noisy label may over parameterize a deep neural network (DNN) and result in a performance drop due to the memorization effect of the DNN. Experiments are conducted on VoxCeleb dataset.

The speaker verification model with SimAM achieves the 0.675% equal error rate (EER) on VoxCeleb1 original test trials. Our proposed iterative noisy label detection method further reduces the EER to 0.643%.


3.Local Information Modeling with Self-Attention for Speaker Verification

标题:基于自我注意的说话人确认局部信息建模

作者:Bing Han, Zhengyang Chen, Yanmin Qian

单位:

MoE Key Lab of Artificial Intelligence, AI Institute

X-LANCE Lab, Department of Computer Science and Engineering

Shanghai Jiao Tong University, Shanghai, China

链接:

https://ieeexplore.ieee.org/document/9746050

摘要:基于自我注意机制的Transformer在很多数自然语言处理(NLP)任务中表现出了最先进的性能,但在以前的工作中,它在用于说话人确认时并没有很强的竞争力。通常,说话人身份主要由相邻token之间的关系来反映,其提取主要依赖于局部建模能力。然而,作为transformer关键组件的self-attention模块可以帮助模型充分利用全局信息,但不足以捕获局部信息。为了解决这一局限性,本文从两个不同的方面加强了局部信息建模:将注意上下文限制为局部的和在transformer中引入卷积运算。在Voxceleb上进行的实验表明,我们提出的方法可以显著提高系统性能,验证局部信息对说话人验证任务的重要性。

Transformer based on self attention mechanism has demonstrated its state-of-the-art performance in most natural language processing (NLP) tasks, but it’s not very competitive when applied for speaker verification in previous works. Generally, speaker identity is mostly reflected by the relationship between adjacent tokens, whose extraction mainly depends on local modeling ability. However, the selfattention module, as the key component of transformer, can help the model make full use of global information but insufficient to capture the local information. To tackle this limitation, in this paper, we strengthen the local information modeling from two different aspects: restricting the attention context to be local and introducing convolution operation into transformer. Experiments conducted on Voxceleb illustrate that our proposed methods can notably improve system performance, verifying the significance of local information for speaker verification task.


4.Multi-query multi-head attention pooling and Inter-topK penalty for speaker verification

标题:用于说话人确认的多查询多头注意池化和TOP K惩罚方法

作者:Miao Zhao1, Yufeng Ma1, Yiwei Ding1;2, Yu Zheng1, Min Liu1, Minqiang Xu1

单位:

1 SpeakIn Technologies Co. Ltd.

2 Fudan University

链接:https://arxiv.org/abs/2110.05042

摘要:本文讲述了2021 VoxCeleb说话人识别挑战(VoxSRC)系统描述中首次提出的多查询多头注意(MQMHA)池和top K类间惩罚方法。大多数多头注意力池化(multi-head attention pooling)机制要么通过多个头注意整个特征,要么关注整个特征的几个分割部分。我们提出的MQMHA结合了这两种机制,并获得了更加多样化的信息。通常采用基于边缘的softmax损失函数来获得有区别的说话人表示。为了进一步增强类间鉴别能力,我们提出了一种方法,对一些易混淆的说话人增加额外的类间top K惩罚。通过采用MQMHA和类间top K惩罚,我们在所有公共VoxCeleb测试集中都实现了最先进的性能。

This paper describes the multi-query multi-head attention (MQMHA) pooling and inter-topK penalty methods which were first proposed in our submitted system description for VoxCeleb speaker recognition challenge (VoxSRC) 2021. Most multi-head attention pooling mechanisms either attend to the whole feature through multiple heads or attend to several split parts of the whole feature. Our proposed MQMHA combines both these two mechanisms and gain more diversified information. The margin-based softmax loss functions are commonly adopted to obtain discriminative speaker representations. To further enhance the inter-class discriminability, we propose a method that adds an extra inter-topK penalty on some confused speakers. By adopting both the MQMHA and inter-topK penalty, we achieved state-of-the-art performance in all of the public VoxCeleb test sets.


5.Temporal Dynamic Convolutional Neural Network for Text-Independent Speaker Verification and Phonemic Analysis

标题:用于文本无关说话人确认和音素分析的时域动态卷积神经网络

作者:Seong-Hu Kim, Hyeonuk Nam, Yong-Hwa Park

单位:Department of Mechanical Engineering, Korea Advanced Institute of Science and Technology, Korea

链接:

https://ieeexplore.ieee.org/document/9747421

摘要:在与文本无关的说话人识别领域,人们提出了沿时间轴自适应的动态模型来考虑语音的音素变化特征。然而,对动态模型如何根据音素工作的详细分析是不够的。在本文中,我们提出了时间动态CNN(TDY-CNN),该CNN通过将核函数最好地应用于每个时间段来考虑音素的时态变化。这些核通过应用训练基核的加权和来适应时间仓位。然后,分析了自适应核在不同层次上对不同音素的作用。TDY-ResNet-38(×0.5)使用六个基核,使说话人验证性能等误率(EER)提高了17.3%与基线模型ResNet-38相比(×0.5)。此外,我们还表明,自适应核依赖于音素组,并且在早期层更具音素特异性。在训练过程中,时间动态模型能够适应没有明确给出音素信息的音素,结果表明有必要考虑话语中的音素变化,以实现更准确和鲁棒的文本无关说话人验证。

In the field of text-independent speaker recognition, dynamic models that adapt along the time axis have been proposed to consider the phoneme-varying characteristics of speech. However, a detailed analysis of how dynamic models work depending on phonemes is insufficient. In this paper, we propose temporal dynamic CNN (TDY-CNN) that considers temporal variation of phonemes by applying kernels optimally adapting to each time bin. These kernels adapt to time bins by applying weighted sum of trained basis kernels. Then, an analysis of how adaptive kernels work on different phonemes in various layers is carried out. TDYResNet-38(×0.5) using six basis kernels improved an equal error rate (EER), the speaker verification performance, by 17.3% compared to the baseline model ResNet-38(×0.5). In addition, we showed that adaptive kernels depend on phoneme groups and are

more phoneme-specific at early layers. The temporal dynamic model adapts itself to phonemes without explicitly given phoneme information during training, and results show the necessity to consider phoneme variation within utterances for more accurate and robust text-independent speaker verification.


6.Towards Lightweight Applications: Asymmetric Enroll-Verify Structure for Speaker Verification

标题:面向轻量级应用:用于说话人验证的非对称ENROLL-VERIFY结构

作者:Qingjian Li1,Lin Yang1, Xuyang Wang1, Xiaoyi Qin2, Junjie Wang1, Ming Li2

单位:

1AI Lab, Lenovo Research, Beijing, China

2Data Science Research Center, Duke Kunshan University, Kunshan, China

链接:https://arxiv.org/abs/2110.04438

摘要:随着深度学习的发展,自动说话人确认在过去几年中取得了长足的进步。然而,在有限的计算资源下设计一个轻量级和鲁棒的系统仍然是一个具有挑战性的问题。传统上,说话人验证系统是对称的,这表明相同的嵌入提取模型适用于推理中的注册和验证。在本文中,我们提出了一种创新的非对称结构,该结构以大规模ECAPA-TDNN模型进行注册,以小规模ECAPA-TDNNLite模型进行验证。作为一个对称系统,我们提出的ECAPA-TDNNLite模型实现了3.07%的能效比,在Voxceleb1原始测试集上,只有11.6M的FLOPS 。此外,非对称结构进一步将EER降低到2.31%,而不会增加验证期间的任何计算成本。

With the development of deep learning, automatic speaker verification has made considerable progress over the past few years. However, to design a lightweight and robust system with limited computational resources is still a challenging problem. Traditionally, a speaker verification system is symmetrical, indicating that the same embedding extraction model is applied for both enrollment and verification in inference. In this paper, we come up with an innovative asymmetric structure, which takes the large-scale ECAPA-TDNN model for enrollment and the small-scale ECAPA-TDNNLite model for verification. As a symmetrical system, our proposed ECAPA-TDNNLite model achieves an EER of 3.07% on the Voxceleb1 original test set with only 11.6M FLOPS. Moreover, the asymmetric structure further reduces the EER to 2.31%, without increasing any computational costs during verification.


7.Improving Fairness in Speaker Verification via Group-Adapted Fusion Network

标题:基于组自适应融合网络的说话人确认公平性改进

作者:Hua Shen1;2, Yuguang Yang2, Guoli Sun2, Ryan Langman2, Eunjung Han2,Jasha Droppo2, Andreas Stolcke2

单位:

1The Pennsylvania State University

2Amazon Alexa AI

链接:https://arxiv.org/abs/2202.11323

摘要:现代说话人确认模型使用深度神经网络将话语音频编码为有区别的嵌入向量。在训练过程中,通常会对这些网络进行优化,以区分任意说话者。这种学习过程会使语音特征的学习偏向于占主导地位的人口统计学群体,这可能导致不同群体之间的不公平表现差异。这一点在代表性不足的人口群体中尤其明显,他们的声音特征相似。在这项工作中,我们研究了在性别分布不平衡的受控数据集上说话人验证模型的公平性,为代表性不足的群体的模型性能受损提供了直接证据。

为了缓解这种差异,我们提出了组自适应融合网络(GFN)体系结构,这是一种基于组嵌入自适应和分数融合的模块化体系结构。我们表明,我们的方法通过改进说话人验证整体和单个组来缓解模型不公平。考虑到训练中的群体代表性不平衡,我们提出的方法实现了相对9.6%-29.0%的总体等错误率(EER)降低,少数群体EER降低了13.7%-18.6%,与基线相比,EER差异减少了20.0%-25.4%。该方法适用于说话人识别系统中其他类型的训练数据倾斜情况。

Modern speaker verification models use deep neural networks to encode utterance audio into discriminative embedding vectors. During the training process, these networks are typically optimized to differentiate arbitrary speakers. This learning process biases the learning of fine voice characteristics towards dominant demographic groups, which can lead to an unfair performance disparity across different groups. This is observed especially with underrepresented demographic groups sharing similar voice characteristics. In this work, we investigate the fairness of speaker verification models on controlled datasets with imbalanced gender distributions, providing direct evidence that model performance suffers for underrepresented groups. To mitigate this disparity we propose the group-adapted fusion network (GFN) architecture, a modular architecture based on group embedding adaptation and score fusion. We show that our method alleviates model unfairness by improving speaker verification both overall and for individual groups. Given imbalanced group representation in training, our proposed method achieves overall equal error rate (EER) reduction of 9.6% to 29.0% relative, reduces minority group EER by 13.7% to 18.6%, and results in 20.0% to 25.4% less EER disparity, compared to baselines. The approach is applicable to other types of training data skew in speaker recognition systems.


8.CS-Rep: Making Speaker Verification Networks Embracing Re-parameterization

标题:CS-REP:使说话人验证网络支持重新参数化

作者:Ruiteng Zhang1, Jianguo Wei1;2, Wenhuan Lu1, Lin Zhang3, Yantao Ji4, Junhai Xu1, Xugang Lu5

单位:

1College of Intelligence and Computing, Tianjin University, Tianjin, China

2Computer College, Qinghai Nationalities University, Xining, China

3National Institute of Informatics, Tokyo, Japan

4School of Software Engineering, Xi’an Jiaotong University, Xi’an, China

5National Institute of Information and Communications Technology, Kyoto, Japan

链接:https://arxiv.org/abs/2110.13465

摘要:自动说话人验证(ASV)系统主要关注确认的准确性,而忽略了推理速度。然而,在实际应用中,推理速度和验证精度都至关重要。为了提高模型的推理速度和验证精度,本研究提出了一种新的多类型网络拓扑重复参数化策略&交叉顺序重参数化(CS-Rep)。CS-Rep解决了现有重参数化方法不适用于典型ASV主干网络的问题。当模型应用CS-Rep时,训练周期网络利用多分支拓扑来捕获说话人信息,而推理周期模型则转换为具有堆叠TDNN层的类时延神经网络(TDNN)的普通主干,以实现快速确认。在CS-Rep的基础上,提出了一种改进的测试部署友好的TDNN-Rep。与最先进的ECAPA-TDNN模型相比,Rep-TDNN将实际推理速度提高了约50%,EER降低了10%。

Automatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and verification accuracy are essential. This study proposes cross-sequential re-parameterization

(CS-Rep), a novel topology re-parameterization strategy for multitype networks, to increase the inference speed and verification accuracy of models. CS-Rep solves the problem that existing reparameterization methods are not suitable for typical ASV backbones. When a model applies CS-Rep, the training-period network  utilizes a multi-branch topology to capture speaker information, whereas the inference-period model converts to a time-delay neural network (TDNN)-like plain backbone with stacked TDNN layers to achieve the fast inference speed. Based on CS-Rep, an improved TDNN with friendly test and deployment called Rep-TDNN is proposed. Compared with the state-of-the-art model ECAPA-TDNN, Rep-TDNN increases the actual inference speed by about 50% and reduces the EER by 10%. The code and trained models are available at https://github.com/zrtlemontree/CS-Rep.


9.Learning Domain-Invariant Transformation for Speaker Verification

标题:用于说话人确认的学习域不变性变换方法

作者:Hanyi Zhang1, Longbiao Wang1, Kong Aik Lee2, Meng Liu1, Jianwu Dang1;3, Hui Chen1

单位:

1Tianjin Key Laboratory of Cognitive Computing and Application,College of Intelligence and Computing, Tianjin University, Tianjin, China

2Institute for Infocomm Research, A*STAR, Singapore

3Japan Advanced Institute of Science and Technology, Ishikawa, Japan

链接:

https://ieeexplore.ieee.org/abstract/document/9747514

摘要:自动说话人验证(ASV)在实际应用中面临着由于记录设备和说话风格等内在和外在因素的不匹配而导致的领域转移,从而导致性能不理想。为此,我们提出了通过元学习的元广义变换来构建域不变嵌入空间。具体而言,转换模块通过对元训练集和元测试集执行元优化来学习领域泛化知识,元训练集和元测试集旨在模拟领域转移。此外,还引入了分布优化来监控嵌入的度量结构。在转换模块方面,我们研究了各种实例,并观察到带门控的多层感知器(gMLP)在其外推能力方面是最有效的。跨体裁和跨数据集设置的实验结果表明,元广义变换显著提高了ASV系统对域转移的鲁棒性,同时优于现有的方法。

Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors such as recording device and speaking style in real-world applications, which leads to unsatisfactory performance. To this end, we propose the meta generalized transformation via meta-learning to build a domain-invariant

embedding space. Specifically, the transformation module is motivated to learn the domain generalization knowledge by executing meta-optimization on the meta-train and meta-test sets which are designed to simulate domain shift. Furthermore, distribution optimization is incorporated to supervise the metric structure of embeddings. In terms of the transformation module, we investigate various instantiations and observe the multilayer perceptron with gating (gMLP) is the most effective given its extrapolation capability. The experimental results on cross-genre and cross-dataset settings demonstrate that the meta generalized transformation dramatically improves the robustness of ASV systems to domain shift, while outperforms the state-of-the-art methods.


10.Tackling the Score Shift in Cross-Lingual Speaker Verification by Exploiting Language Information

标题:利用语言信息解决跨语言说话人确认中的分数偏移问题

作者:Jenthe Thienpondt, Brecht Desplanques, Kris Demuynck

单位:

IDLab, Department of Electronics and Information Systems, Ghent University - imec, Belgium

链接:https://arxiv.org/abs/2110.09150

摘要:本文对2021 VoxCeleb说话人识别挑战赛(VoxSRC-21)IDLab提交的跨语言说话人验证进行了赛后性能分析。我们发现,当前的说话人嵌入提取器在说话人内跨语言试验中始终低估了说话人相似度。因此,典型的训练和评分协议没有对说话人内部语言可变性的补偿给予足够的重视。我们提出了两种技术来提高跨语言说话人确认的稳健性。首先,我们使用一种小批量采样策略来增强之前提出的大幅度微调(Large-Margin Fine-Tuning,LM-FT)训练阶段,该策略增加了小批量中说话人内部跨语言样本的数量。其次,我们将语言信息纳入逻辑回归校准阶段。我们基于VoxLingua107语言识别模型的软决策和硬决策来集成质量度量。所提出的技术在VoxSRC-21测试集上比基线模型相对提高了11.7%,并帮助我们在相应的挑战中获得第三名。

This paper contains a post-challenge performance analysis on crosslingual speaker verification of the IDLab submission to the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). We show that current speaker embedding extractors consistently underestimate speaker similarity in within-speaker cross-lingual trials. Consequently, the typical training and scoring protocols do not put enough emphasis on the compensation of intra-speaker language variability. We propose two techniques to increase cross-lingual speaker verification robustness. First, we enhance our previously proposed Large-Margin Fine-Tuning (LM-FT) training stage with a minibatch sampling strategy which increases the amount of intra-speaker cross-lingual samples within the mini-batch. Second, we incorporate language information in the logistic regression calibration stage. We integrate quality metrics based on soft and hard decisions of a VoxLingua107 language identification model. The proposed techniques result in a 11.7% relative improvement over the baseline model on the VoxSRC-21 test set and contributed to our third place finish in the corresponding challenge.


11.CDMA: Cross-Domain Distance Metric Adaptation for Speaker Verification

标题:CDMA:用于说话人确认的跨域距离度量自适应

作者:Jianchen Li, Jiqing Han, Hongwei Song

单位:

School of Computer Science and Technology, Harbin Institute of Technology, Harbin, China

链接:

https://ieeexplore.ieee.org/abstract/document/9747907

摘要:为了解决说话人确认中的域转移问题,一种有效的域自适应方法是通过对齐嵌入空间中的源域分布和目标域分布来学习域不变嵌入。然而,当源域和目标域来自不相交的说话人标签空间时,由于不同说话人的嵌入分布无法对齐,这种方法可能会出现问题。在本文中,我们提出了一种跨域距离度量自适应(CDMA)方法来缓解距离度量空间中的域偏移,其中源域和目标域共享相同的类,即说话人内部和说话人之间。具体而言,将两个目标的成对距离分布与源的成对距离分布对齐,并进一步分离以学习一个域不变度量,该度量更适合基于度量学习的说话人验证。实验表明,CDMA在嵌入空间的性能明显优于提出的方法。

To solve the domain shift problem in speaker verification, one effective domain adaptation approach is to learn domaininvariant embeddings via aligning the source and target distributions in the embedding space. However, this approach could be problematic when the source and target domains are from the disjoint speaker label spaces as the embedding distributions of different speakers cannot be aligned. In this paper, we propose a Cross-domain Distance Metric Adaptation (CDMA) approach to alleviate the domain shift in the distance metric space, where the source and target domains share the same classes, i.e., within- and between-speaker. Specifically, the two target pairwise distance distributions are aligned with the source pairwise distance distributions and further separated to learn a domain-invariant metric, which is more suitable for speaker verification based on metric learning. Experiments indicate that CDMA significantly outperforms the approach proposed in the embedding space.


12.MFA: TDNN with Multi-scale Frequency-channel Attention for Text-independent Speaker Verification with Short Utterances

标题:MFA:基于多尺度频率通道注意的TDNN短话语文本无关说话人验证

作者:Tianchi Liu1;2, Rohan Kumar Das3, Kong Aik Lee1 ,Haizhou Li2;4

单位:

1Institute for Infocomm Research, A*STAR, Singapore

2Department of Electrical and Computer Engineering, National University of Singapore, Singapore

3Fortemedia Singapore, Singapore 4The Chinese University of Hong Kong, Shenzhen, China

链接:https://arxiv.org/abs/2202.01624

摘要:时延神经网络(TDNN)代表了文本无关说话人确认的最新神经解决方案之一。然而,它们需要大量的滤波器来捕获任何局部频率区域的扬声器特性。此外,这种系统的性能在短话语场景下可能会下降。为了解决这些问题,我们提出了一种多尺度频率通道注意(MFA),通过一种由卷积神经网络和TDNN组成的新型双路径设计,在不同尺度上对说话人进行表征。我们在VoxCeleb数据库上对所提出的MFA进行了评估,并观察到所提出的带有MFA的框架可以实现最先进的性能,同时降低参数和计算复杂度。此外,MFA机制对于短测试话语的说话人验证是有效的。

The time delay neural network (TDNN) represents one of the state-of-the-art of neural solutions to text-independent speaker verification. However, they require a large number of filters to capture the speaker characteristics at any local frequency region. In addition, the performance of such systems may degrade under short utterance scenarios. To address these issues, we propose a multi-scale frequency-channel attention (MFA), where we characterize speakers at different scales through a novel dual-path design which consists of a

convolutional neural network and TDNN. We evaluate the proposed MFA on the VoxCeleb database and observe that the proposed framework with MFA can achieve state-of-theart performance while reducing parameters and computation complexity. Further, the MFA mechanism is found to be effective for speaker verification with short test utterances.


13.MLP-SVNET: A Multi-Layer Perceptrons Based Network for Speaker Verification

标题:MLP-SVNET:一种基于多层感知器的说话人确认网络

作者:Bing Han, Zhengyang Chen, Bei Liu, Yanmin Qian

单位:

MoE Key Lab of Artificial Intelligence, AI Institute

X-LANCE Lab, Department of Computer Science and Engineering

Shanghai Jiao Tong University, Shanghai, China

链接:https://ieeexplore.ieee.org/document/9747172

摘要:概率线性鉴别分析(PLDA)或余弦相似度在传统的说话人确认系统中被广泛用作衡量成对相似度的后端技术。为了更好地利用多个注册语音,我们提出了一种新的注意后端模型,该模型可用于文本无关(TI)和文本相关(TD)说话人确认,并且我们使用缩放点积自注意力结构(scaled-dot self-attention)与前馈自注意力结构(feed-forward self-attention)作为学习注册话语内部关系的架构。为了验证所提出的模型,我们在CNCeleb和VoxCeleb数据集上进行了一系列实验,将其与包括TDNN和ResNet在内的几种最先进的扬声器编码器相结合。在CNCeleb上使用多个注册话语获得的实验结果表明,对于每个说话人编码器,所提出的注意后端模型的EER和minDCF分数低于其PLDA和余弦相似性对应项,并且在VoxCeleb上的实验表明,我们的模型甚至可以用于单个注册语音的情况。

Probabilistic linear discriminant analysis (PLDA) or cosine similarity have been widely used in traditional speaker verification systems as back-end techniques to measure pairwise similarities. To make better use of multiple enrollment utterances, we propose a novel attention back-end model that can be used for both textindependent (TI) and text-dependent (TD) speaker verification, and we use scaled-dot self-attention and feed-forward self-attention networks as architectures that learn the intra-relationships of enrollment utterances. To verify the proposed model, we conduct a series of experiments on the CNCeleb and VoxCeleb datasets by combining it with several state-of-the-art speaker encoders including TDNN and ResNet. Experimental results obtained using multiple enrollment utterances on CNCeleb show that the proposed attention back-end model leads to lower EER and minDCF scores than its PLDA and cosine similarity counterparts for each speaker encoder, and an experiment on VoxCeleb demonstrates that our model can be used even for a single enrollment case.


14.Real Additive Margin Softmax for Speaker Verification

标题:用于说话人验证的实际附加边界SOFTMAX

作者:Lantian Li, Ruiqian Nai, Dong Wang

单位:

Center for Speech and Language Technologies, BNRist, Tsinghua University, China

链接:https://arxiv.org/abs/2110.09116

摘要:加性间隔softmax(additive margin softmax,AM softmax)损失函数在说话人确认中提供了良好的性能。AM Softmax的一个假定行为是,它可以通过强调目标注册来缩小类内变化,从而提高目标类和非目标类之间的差距。在本文中,我们对AM Softmax损耗的行为进行了仔细的分析,并表明该损失并没有实现真正的最大边界训练。基于这一观察结果,我们提出了一个真实的AM Softmax损耗,该损耗涉及Softmax训练中的真实边界函数。在VoxCeleb1、SITW和CNCeleb上进行的实验表明,修正后的AM Softmax损耗始终优于原始损失。

The additive margin softmax (AM-Softmax) loss has delivered remarkable performance in speaker verification.

A supposed behavior of AM-Softmax is that it can shrink within-class variation by putting emphasis on target logits, which in turn improves margin between target and non-target classes. In this paper, we conduct a careful analysis on the behavior of AM-Softmax loss, and show that this loss does not implement real max-margin training. Based on this observation, we present a Real AM-Softmax loss which involves a true margin function in the softmax training. Experiments conducted on VoxCeleb1, SITW and CNCeleb demonstrated that the corrected AM-Softmax loss consistently outperforms the original one. The code has been released at https://gitlab.com/csltstu/sunine.


审阅丨程星亮


15.Adversarial Sample Detection for Speaker Verification by Neural Vocoders

标题:神经声码器用于说话人验证的对抗性样本检测

作者:Haibin Wu1 , Po-chun Hsu1 , Ji Gao2 , Shanshan Zhang2 , Shen Huang2 , Jian Kang2 , Zhiyong Wu3 , Helen Meng4 , Hung-yi Lee1

单位:

1 Graduate Institute of Communication Engineering, National Taiwan University

4 Centre for Perceptual and Interactive Intelligence, The Chinese University of Hong Kong

3 Shenzhen International Graduate School, Tsinghua University

2 Tencent Research, Beijing, China

链接:https://arxiv.org/abs/2107.00309

摘要:语音自动验证(ASV)技术是生物特征识别的重要技术之一,已广泛应用于至关重要的安全领域。然而,ASV在最近出现的对抗性攻击中非常脆弱,有效的应对措施却有限。本文采用神经声码器对ASV的对抗样本进行识别。我们使用神经声码器重新合成音频,发现原始音频和重新合成音频之间的ASV分数差异是区分真实样本和敌对样本的一个很好的指标。据我们所知,这项工作是在ASV检测时域对抗样本的技术方向上进行的第一次尝试,所以缺乏用于比较的基线。因此,我们将Griffin-Lim算法作为检测基线。在所有的实验和参数配置下,该方法的检测性能都优于基线。我们还表明,检测框架中采用的神经声码器是独立于数据集的。我们的代码开源,以便将来的工作进行公平的比较。

Automatic speaker verification (ASV), one of the most important technology for biometric identification, has been widely adopted in security-critical applications. However, ASV is seriously vulnerable to recently emerged adversarial attacks, yet effective countermeasures against them are limited. In this paper, we adopt neural vocoders to spot adversarial samples for ASV. We use the neural vocoder to re-synthesize audio and find that the difference between the ASV scores for the original and re-synthesized audio is a good indicator for discrimination between genuine and adversarial samples. This effort is, to the best of our knowledge, among the first to pursue such a technical direction for detecting time-domain adversarial samples for ASV, and hence there is a lack of established baselines for comparison. Consequently, we implement the Griffin-Lim algorithm as the detection baseline. The proposed approach achieves effective detection performance that outperforms the baselines in all the settings. We also show that the neural vocoder adopted in the detection framework is dataset-independent. Our codes will be made open-source for future works to do fair comparison.


16.Self-Supervised Speaker Verification with Simple Siamese Network and Self-Supervised Regularization

标题:基于简单孪生网络和自监督正则化的自监督说话人验证

作者:Mufan Sang1, Haoqi Li2 , Fang Liu2 , Andrew O. Arnold2 , Li Wan2

单位:

1The University of Texas at Dallas, TX, USA

2Amazon AWS AI, USA

链接:https://arxiv.org/abs/2112.04459

摘要:在缺少说话人标签的情况下训练说话人识别和鲁棒的说话人验证系统仍然具有挑战性。在本研究中,我们提出了一个有效的自监督学习框架和一种新的正则化策略来促进自监督说话人表征学习。与基于对比学习的自监督学习方法不同,所提出的自监督正则化方法(SSReg)只关注正例数据对潜在表示之间的相似性。我们还探讨了替代在线数据扩充策略在时域和频域上的有效性。通过强大的在线数据扩充策略,所提出的SSReg显示了无需使用负样本对的自监督学习的潜力,并且它可以通过简单的孪生网络结构显著提高自监督说话人表示学习的性能。在VoxCeleb数据集上的综合实验表明,通过添加有效的自监督正则化,我们提出的自监督方法获得了23.4%的相对改进,性能超过了其他研究工作。

Training speaker-discriminative and robust speaker verification systems without speaker labels is still challenging and worthwhile to explore. In this study, we propose an effective self-supervised learning framework and a novel regularization strategy to facilitate self-supervised speaker representation learning. Different from contrastive learning-based self-supervised learning methods, the proposed self-supervised regularization (SSReg) focuses exclusively on the similarity between the latent representations of positive data pairs. We also explore the effectiveness of alternative online data augmentation strategies on both the time domain and frequency domain. With our strong online data augmentation strategy, the proposed SSReg shows the potential of self-supervised learning without using negative pairs and it can significantly improve the performance of self-supervised speaker representation learning with a simple Siamese network architecture. Comprehensive experiments on the VoxCeleb datasets demonstrate that our proposed self-supervised approach obtains a 23.4% relative improvement by adding the effective self-supervised regularization and outperforms other previous works.


17.Robust Speaker Verification with Joint Self-Supervised and Supervised Learning

标题:基于联合自监督和监督学习的鲁棒说话人验证

作者:Kai Wang1 , Xiaolei Zhang1 , Miao Zhang1 , Yuguang Li1 , Jaeyun Lee2 , Kiho Cho2 , Sung-UN Park

单位:

1Samsung R&D Institute China Xian,

2Samsung Advanced Institute of Technology

链接:https://ieeexplore.ieee.org/document/9747209

摘要:监督学习和自我监督学习涉及不同的方面。监督学习具有较高的精度,但它确实需要大量昂贵的标记数据。相应地,自监督学习利用大量的未标记数据进行学习,但其性能落后于监督学习。为了克服获取标注数据的困难,并在说话人验证的背景下保持较高的性能,本文提出了一种自我监督联合学习(SS-JL)框架,该框架在联合训练中用自我监督辅助任务来补充监督的主任务。这些辅助任务有助于说话人验证管道生成与声纹密切相关的鲁棒说话人表示。我们的模型在英语数据集上进行了训练,并在多语言数据集(包括英语、汉语和韩语数据集)上进行了测试,与基线相比,平均错误率(EER)分别提高了13.6%、12.7%和13.5%。

upervised learning and self-supervised learning address different facets. Supervised learning achieves high accuracy, but it requires numerous expensive labeled data indeed. Correspondingly, self-supervised learning, makes use of abundant unlabeled data to learn, but the performance lags behind that of the supervised counterpart. To overcome the difficulty of acquiring annotated data and contain the high performance in the context of speaker verification, we propose in this work a self-supervised joint learning (SS-JL) framework which complements the supervised main task with self-supervised auxiliary tasks in joint training. These auxiliary tasks help the speaker verification pipeline to generate robust speaker representation that is closely relevant to voiceprints. Our model is trained on English dataset and tested on multilingual datasets, including English, Chinese and Korean datasets, and 13.6%, 12.7% and 13.5% improvement is achieved respectively in terms of equal error rate (EER) compared with the baselines.


18.Statistical Pyramid Dense Time Delay Neural Network for Speaker Verification

标题:用于说话人确认的统计金字塔密集时延神经网络

作者:Zi-Kai Wan1 , Qing-Hua Ren1 , You-Cai Qin1 , Qi-Rong Mao1,2

单位:

1School of Computer Science and Communication Engineering, Jiangsu University, China

2 Jiangsu Engineering Research Center of Big Data Ubiquitous Perception and Intelligent Agriculture Applications, China

链接:https://ieeexplore.ieee.org/document/9746650

摘要:近年来,说话人验证(SV)技术依赖于深度学习框架来提取信息量更大的嵌入向量,与传统的机器学习方法相比,大大提高了准确性。众所周知的x-vector体系结构是一种时延神经网络(TDNN),广泛适用于SV任务。然而,现有的大多数变体很少结合全局和子区域上下文信息,并且受到标准卷积运算产生的局部感受野的影响。在本文中,我们提出了统计金字塔密集TDNN(SPDTDNN),其中包含捕获上下文信息的统计金字塔池化模块。具体而言,所开发的模块从不同的角度自适应地在上下文区域之间交换信息,这些部分区域对应于多个并行分支。全局分支收集的统计数据由时域的平均值和标准差组成,以获取更多的全局上下文信息。在VoxCeleb1&2数据集上的大量实验表明,所提出的PSD-TDNN优于相应的D-TDNN、D-TDNN-SS和ECAPA-TDNN,它们在SV任务上实现了最先进的性能,并且具有相似的模型复杂度。

Recently, speaker verification (SV) techniques relay on deep learning frameworks to extract more informative embedding vectors, which greatly improves the accuracy compared with traditional machine learning methods. The well-known xvector architecture, a time delay neural network (TDNN), is widely adapted for SV tasks. However, most of existing variants rarely combines the global and sub-region context information and suffer from the local receptive field that is engendered by the standard convolutional operation. In this paper, we propose statistical pyramid dense TDNN (SPDTDNN) with the statistical pyramid pooling module which captures the context information. Specifically, the developed module adaptively exchanges information among contextual regions from different perspectives, which correspond to multiple parallel branches. The statistics collected by the global-region branch are comprised of mean and standard deviation across the time domain to acquire the more global context information. Extensive experiments on the VoxCeleb1&2 datasets demonstrate that the proposed PSDTDNN outperforms corresponding D-TDNN, D-TDNN-SS and ECAPA-TDNN which achieve the state-of-the-art performances on the SV task, with similar model complexity.


19.On the Importance of Different Frequency Bins for Speaker Verification

标题:不同频率单元对说话人验证的重要性

作者:Aiwen Deng1 , Shuai Wang2 , Wenxiong Kang1 , Feiqi Deng1

单位:

1South China University of Technology, Guangzhou, China

2 Shanghai Jiao Tong University, Shanghai, China

链接:https://ieeexplore.ieee.org/document/9746084

摘要:大多数现代说话人验证系统都以基于频谱分析的特征作为输入,其中包含多个频率单元。当然,会有一个问题,即所有不同的频率单元是否对说话人验证系统的性能有同等的贡献?在本文中,我们提出了频率重新加权层(FRL)来自动学习和平衡不同频率单元的重要性。该层可以在不同的层上自由地插入到原始的说话人嵌入学习器中一次或多次,引入新参数的数量极少。基于所提出的新体系结构,在VoxCeleb1数据集上完成了实验,该数据集不仅实现了优异的性能,而且还显示了有意义的权重分布——较低的频率更重要。

The majority of modern speaker verification systems take spectral analysis-based features as input, which contains mul[1]tiple frequency bins. Naturally, there would be a question of whether all different frequency bins contribute equally to the speaker verification system performance? In this paper, we propose the frequency reweighting layer (FRL) to automati[1]cally learn and balance the importance of different frequency bins. This new layer can be freely inserted into the original speaker embedding learner once or multiple times at different layers, with an ignorable number of new parameters. Based on the proposed novel architecture, a set of experiments are designed and carried out on the VoxCeleb1 dataset, which not only achieves superior performance but also exhibits an interesting weight distribution – the lower frequencies matter more.


20.Self-Knowledge Distillation via Feature Enhancement for Speaker Verification

标题:说话人验证中的特征增强自知识蒸馏方法

作者:Bei Liu, Haoyu Wang, Zhengyang Chen, Shuai Wang, Yanmin Qian

单位:

MoE Key Lab of Artificial Intelligence, AI Institute

X-LANCE Lab, Department of Computer Science and Engineering

Shanghai Jiao Tong University, Shanghai, China

链接:https://ieeexplore.ieee.org/document/9746529

摘要:近些年来,深度说话人嵌入学习已经成为说话验证任务中主流且应用最广泛的技术。ECAPA-TDNN和ResNet等大型神经网络可以实现最先进的性能。然而,大型模型通常不利于计算,需要大量的存储和计算资源。模型压缩一直是研究的热点。参数量化通常会导致性能显著下降。知识蒸馏需要一个经过预先训练的复杂教师模型。本文介绍了一种新的自知识提取方法,即基于特征增强的自知识蒸馏方法(SKDFE)。它利用一个辅助自教师网络来提取自己的精炼知识,而不需要预先训练好的教师网络。此外,我们在标签级和特征级两个不同的层次上应用了自知识提取。在Voxceleb数据集上的实验表明,我们提出的自知识提取方法可以使小模型具有与大模型相当甚至更好的性能。应用我们的方法可以进一步改进大型模型。

As the most widely used technique, deep speaker embedding learning has become predominant in speaker verification task recently. Very large neural networks such as ECAPA-TDNN and ResNet can achieve the state-of-the-art performance. However, large models are computationally unfriendly in general, which require massive storage and computation resources. Model compression has been a hot research topic. Parameter quantization usually results in significant performance degradation. Knowledge distillation demands a pretrained complex teacher model. In this paper, we introduce a novel self-knowledge distillation method, namely Self-Knowledge Distillation via Feature Enhancement (SKDFE). It utilizes an auxiliary self-teacher network to distill its own refined knowledge without the need of a pretrained teacher network. Additionally, we apply the self-knowledge distillation at two different levels: label level and feature level. Experiments on Voxceleb dataset show that our proposed self-knowledge distillation method can make small models have comparable or even better performance than large ones. Large models can also be further improved when applying our method.


21.Robust Speaker Verification Using Population-Based Data Augmentation

标题:基于群体数据增强的鲁棒说话人验证

作者:Weiwei Lin and Man-Wai Ma

单位:Dept. of Electronic and Information Engineering, The Hong Kong Polytechnic University Hong Kong SAR

链接:

https://ieeexplore.ieee.org/abstract/document/9746956/

摘要:低信噪比和高混响环境下的说话人识别是一个重要挑战。数据增强可以用来模拟说话人识别系统可能遇到的不利环境。通常扩增参数是手动设置的。近年来,基于群体学习的超参数自动优化方法取得了良好的效果。这篇文章提出了一种基于群体的增广参数优化搜索策略。我们将由此产生的扩增称为基于群体的扩增(PBA)。PBA学习一个策略器来设置超参数,而不是查找一组固定的超参数。与网格搜索相比,该策略提供了相当大的计算优势。我们仅使用六个网络模块就获得了高性能的增强策略。使用PBA,我们在VOiCES19评估集上实现了3.98%的EER。

Speaker recognition under environments with a low signalto-noise ratio (SNR) and high reverberation level has always been challenging. Data augmentation can be applied to simulate the adverse environments that a speaker recognition system may encounter. Typically, the augmentation parameters are manually set. Recently, automatic hyper-parameter optimization using population-based learning has shown promising results. This paper proposes a population-based searching strategy for optimizing the augmentation parameters. We refer to the resulting augmentation as populationbased augmentation (PBA). Instead of finding a fixed set of hyper-parameters, PBA learns a scheduler for setting the hyper-parameters. This strategy offers a considerable computation advantage over the grid search. We obtained highperformance augmentation policies using a population of six networks only. With PBA, we achieved an EER of 3.98% on the VOiCES19 evaluation set.


22.RawNeXt: Speaker Verification System For Variable-Duration Utterances With Deep Layer Aggregation And Extended Dynamic Scaling Policies

标题:RAWNEXT:具有深层聚合和扩展动态缩放策略的可变时长语音说话人验证系统

作者:Ju-ho Kim, Hye-jin Shim, Jungwoo Heo, and Ha-Jin Yu

单位:School of Computer Science, University of Seoul

链接:https://arxiv.org/abs/2112.07935

摘要:尽管使用深度神经网络在说话人确认方面取得了令人满意的性能,但可变时长的话语仍然是一个挑战,威胁着系统的鲁棒性。为了解决这个问题,我们提出了一个名为RawNeXt的说话人验证系统,该系统可以通过使用以下两个组件来处理任意长度的输入原始波形:(1)深层聚合策略通过迭代和分层聚合来自块输出的各种时间尺度和频谱通道的特征来增强说话人信息。(2) 扩展的动态缩放策略通过有选择地合并每个块中不同分辨率分支的激活特征,根据话语的长度灵活地处理。由于这两个因素,我们提出的模型可以提取出丰富的时间谱信息的说话人嵌入,并对长度变化进行动态操作。在包含各种持续时间话语的VoxCeleb1测试集上的实验结果表明,与最近提出的系统相比,RawNeXt实现了最先进的性能。我们的代码和经过训练的模型权重可在https://github.com/wngh1187/RawNeXt.

Despite achieving satisfactory performance in speaker verification using deep neural networks, variable-duration utterances remain a challenge that threatens the robustness of systems. To deal with this issue, we propose a speaker verification system called RawNeXt that can handle input raw waveforms of arbitrary length by employing the following two components: (1) A deep layer aggregation strategy enhances speaker information by iteratively and hierarchically aggregating features of various time scales and spectral channels output from blocks. (2) An extended dynamic scaling policy flexibly processes features according to the length of the utterance by selectively merging the activations of different resolution branches in each block. Owing to these two components, our proposed model can extract speaker embeddings rich in time-spectral information and operate dynamically on length variations. Experimental results on the VoxCeleb1 test set consisting of various duration utterances demonstrate that RawNeXt achieves state-of-the-art performance compared to the recently proposed systems. Our code and trained model weights are available at https://github.com/wngh1187/RawNeXt.


23.Contrastive-mixup Learning for Improved Speaker Verification

标题:用于改进说话人验证的对比mixup学习

作者:Xin Zhang1,Minho Jin2 , Roger Cheng2, Ruirui Li2, Eunjung Han2, Andreas Stolcke2

单位:

1 Texas A&M University Dept. of Computer Science & Engineering College Station, TX, USA

2 Amazon Alexa AI Sunnyvale, CA, USA

链接:https://arxiv.org/abs/2202.10672

摘要:本文提出了一种用于说话人验证的带有mixup的原型损失公式,Mixup是一种简单而有效的数据增强技术,它将随机数据点和标签对进行加权组合,用于深度神经网络训练。由于mixup能够提高深层神经网络的鲁棒性和泛化能力,因此它受到了越来越多的关注。尽管mixup在不同领域取得了成功,但大多数应用程序都围绕着闭集分类任务展开。在这项工作中,我们提出了对比mixup,这是一种新的增强策略,可以学习基于距离度量的区分性表示。在训练期间,mixup操作生成输入和虚拟标签的凸插值。此外,我们重新制定了原型损失函数,以便在度量学习目标上实现混合。为了在有限的训练数据下证明其泛化性,我们通过改变VoxCeleb数据库中每个说话人的可用话语数来进行实验。实验结果表明,对比mixup的性能优于现有的基线,尤其是在每个说话人的训练话语数有限的情况下,错误率相对降低了16%。

This paper proposes a novel formulation of prototypical loss with mixup for speaker verification. Mixup is a simple yet efficient data augmentation technique that fabricates a weighted combination of random data point and label pairs for deep neural network training. Mixup has attracted increasing attention due to its ability to improve robustness and generalization of deep neural networks. Although mixup has shown success in diverse domains, most applications have centered around closed-set classification tasks. In this work, we propose contrastive-mixup, a novel augmentation strategy that learns distinguishing representations based on a distance metric. During training, mixup operations generate convex interpolations of both inputs and virtual labels. Moreover, we have reformulated the prototypical loss function such that mixup is enabled on metric learning objectives. To demonstrate its generalization given limited training data, we conduct experiments by varying the number of available utterances from each speaker in the VoxCeleb database. Experimental results show that applying contrastive-mixup outperforms the existing baseline, reducing error rate by 16% relatively, especially when the number of training utterances per speaker is limited.


24.Learnable Nonlinear Compression for Robust Speaker Verification

标题:用于鲁棒说话人验证的可学习非线性压缩方法

作者:Xuechen Liu1,2 , Md Sahidullah2 , Tomi Kinnunen1

单位:

1School of Computing, University of Eastern Finland, Joensuu, Finland 2Universite de Lorraine, CNRS, Inria, LORIA, F-54000, Nancy, France

链接:https://arxiv.org/abs/2202.05236

摘要:在这项工作中,我们主要研究基于深度神经网络的说话人确认中频谱特征的非线性压缩方法。我们考虑了以数据驱动方式优化的不同类型的信道相关(CD)非线性压缩方法。我们的方法基于幂函数非线性和动态范围压缩(DRC)。我们还提出了基于非线性的多区域(MR)设计,以提高鲁棒性。在VoxCeleb1和Vox-Movies数据上的结果表明,所提出的压缩方法比常用的对数和静态对数都有所改进,尤其是基于幂函数的压缩方法。虽然CD泛化提高了VoxEleb1上的性能,但MR在Vox-Movies上提供了更高的鲁棒性,最大相对等错误率降低了21.6%。

In this study, we focus on nonlinear compression methods in spectral features for speaker verification based on deep neural network. We consider different kinds of channel-dependent (CD) nonlinear compression methods optimized in a data-driven manner. Our methods are based on power nonlinearities and dynamic range compression (DRC). We also propose multi-regime (MR) design on the nonlinearities, at improving robustness. Results on VoxCeleb1 and VoxMovies data demonstrate improvements brought by proposed compression methods over both the commonly-used logarithm and their static counterparts, especially for ones based on power function. While CD generalization improves performance on VoxCeleb1, MR provides more robustness on VoxMovies, with a maximum relative equal error rate reduction of 21.6%.


25.Graph Attentive Feature Aggregation for Text-Independent Speaker Verification

标题:用于文本无关说话人验证的图注意特征聚合方法

作者:Hye-jin Shim1 , Jungwoo Heo1 , Jae-han Park2 , Ga-hui Lee2 , and Ha-Jin Yu1

单位:

1School of Computer Science, University of Seoul,

2KT Corporation

链接:https://arxiv.org/abs/2112.12343

摘要:本文的目的是在考虑成对关系的情况下,将多个帧级特征组合成单个话语级表示。为此,我们提出了一种新的图注意力特征聚合模块,将每个帧级特征理解为图的一个节点。通常间接利用的所有可能的特征对之间的相互关系可以使用图直接建模。该模块包括图注意层和图池化层,然后是读出操作。图注意层首先对不同节点之间的非欧几里德数据流形进行建模。然后,考虑到节点的重要性,图池层丢弃信息较少的节点。最后,读出操作将剩余节点组合成单个表示。我们使用了两个最新的系统,SEResNet和RawNet2,它们具有不同的输入特性和体系结构,并证明了所提出的特性聚合模块与基线相比始终显示出超过10%的相对改进。

The objective of this paper is to combine multiple frame-level features into a single utterance-level representation considering pairwise relationships. For this purpose, we propose a novel graph attentive feature aggregation module by interpreting each frame-level feature as a node of a graph. The inter-relationship between all possible pairs of features, typically exploited indirectly, can be directly modeled using a graph. The module comprises a graph attention layer and a graph pooling layer followed by a readout operation. The graph attention layer first models the non-Euclidean data manifold between different nodes. Then, the graph pooling layer discards less informative nodes considering the significance of the nodes. Finally, the readout operation combines the remaining nodes into a single representation. We employ two recent systems, SEResNet and RawNet2, with different input features and architectures and demonstrate that the proposed feature aggregation module consistently shows a relative improvement over 10%, compared to the baseline.


26.Multisv: Dataset for Far-Field Multi-Channel Speaker Verification

标题:MULTISV:用于远场多通道说话人验证的数据集

作者:Ladislav Mosner, Old ˇ rich Plchot, Luk ˇ a´s Burget, Jan “Honza” ˇ Cernock ˇ y

单位:Brno University of Technology, Faculty of Information Technology, Speech@FIT, Czechia

链接:https://arxiv.org/abs/2111.06458

摘要:受数据稀疏和该领域缺乏标准基准的影响,我们在先前工作的基础上提出了一个用于训练和评估文本无关的多通道说话人验证系统的综合语料库。它将用于去冗余、去噪和语音增强的实验。我们通过在Voxceleb语料库的干净部分上进行数据模拟,解决了一直存在的缺少多通道训练数据的问题。开发和评估试验基于在复杂环境设置(VOiCES)语料库中模糊的重发语音,我们对其进行了修改,以提供多通道试验。我们发布了从公共来源创建数据集的完整集合,称为MultiSV数据集,并提供了两个基于神经网络波束形成的多通道说话人验证系统的结果,该系统基于预测理想二进制掩码或最近的Conv-TasNet。

Motivated by unconsolidated data situation and the lack of a stan[1]dard benchmark in the field, we complement our previous efforts and present a comprehensive corpus designed for training and evaluating text-independent multi-channel speaker verification systems. It can be readily used also for experiments with dereverberation, denois[1]ing, and speech enhancement. We tackled the ever-present problem of the lack of multi-channel training data by utilizing data simula[1]tion on top of clean parts of the Voxceleb corpus. The development and evaluation trials are based on a retransmitted Voices Obscured in Complex Environmental Settings (VOiCES) corpus, which we mod[1]ified to provide multi-channel trials. We publish full recipes that create the dataset from public sources as the MultiSV dataset, and we provide results with two of our multi-channel speaker verifica[1]tion systems with neural network-based beamforming based either on predicting ideal binary masks or the more recent Conv-TasNet.


27.Multi-Channel Speaker Verification with Conv-Tasnet Based Beamformer

标题:使用基于Conv-Tasnet的波束形成器进行多通道说话人验证

作者:Ladislav Mosner, Old ˇ rich Plchot, Luk ˇ a´s Burget, Jan “Honza” ˇ Cernock ˇ y

单位:Brno University of Technology, Faculty of Information Technology, Speech@FIT, Czechia

链接:https://arxiv.org/abs/2111.06458

摘要:我们主要研究远场多通道数据中的说话人识别问题。主要贡献是介绍了一种从时域信号预测波束形成器空间协方差矩阵(SCM)的替代方法。我们建议使用著名的源分离模型Conv-TasNet,并通过强制其分离语音和加性噪声来对其进行语音增强。我们使用Conv-TasNet输出的STFT进行实验,以获得语音和噪声的SCMs,最后,我们对这个多通道前端w.r.t.说话人验证目标进行了微调。我们利用MultiSV语料库的模拟数据成功地解决了缺乏真实多通道训练集的问题。对其转播和模拟测试部件进行了分析。我们使用比基于掩码估计神经网络的方案的基线模型小2.7倍的模型得到了同样的结果。

We focus on the problem of speaker recognition in far-field multichannel data. The main contribution is introducing an alternative way of predicting spatial covariance matrices (SCMs) for a beamformer from the time domain signal. We propose to use ConvTasNet, a well-known source separation model, and we adapt it to perform speech enhancement by forcing it to separate speech and additive noise. We experiment with using the STFT of Conv-TasNet outputs to obtain SCMs of speech and noise, and finally, we finetune this multi-channel frontend w.r.t. speaker verification objective. We successfully tackle the problem of the lack of a realistic multichannel training set by using simulated data of MultiSV corpus. The analysis is performed on its retransmitted and simulated test parts. We achieve consistent improvements with a 2.7 times smaller model than the baseline based on a scheme with mask estimating NN.


28.Disentangled Speaker Embedding for Robust Speaker Verification

标题:用于鲁棒说话人验证的解耦性说话人嵌入

作者:Lu YI and Man-Wai MAK

单位:Department of Electronic and Information Engineering The Hong Kong Polytechnic University, Hong Kong SAR

链接:

https://ieeexplore.ieee.org/abstract/document/9747778

摘要:当在一个未知域上评估说话人验证系统时,说话人特征和冗余特征的纠缠可能会导致性能不佳。为了解决这个问题,我们提出了一个InfoMax域分离和自适应网络(InfoMax–DSAN),以基于域自适应技术分离特定于域的特征和域不变的说话人特征。提出了一种基于帧的互信息神经估计器,以最大化帧级特征和输入声学特征之间的互信息,从而保留更多有用信息。此外,我们建议采用基于自监督学习思想的三重损失来克服标签失配问题。在VOiCES Challenge 2019上的实验结果表明,我们提出的方法可以帮助学习更具辨别力和鲁棒性的说话人嵌入。

Entanglement of speaker features and redundant features may lead to poor performance when evaluating speaker verification systems on an unseen domain. To address this issue, we propose an InfoMax domain separation and adaptation network (InfoMax–DSAN) to disentangle the domain-specific features and domain-invariant speaker features based on domain adaptation techniques. A frame-based mutual information neural estimator is proposed to maximize the mutual information between frame-level features and input acoustic features, which can help retain more useful information. Furthermore, we propose adopting triplet loss based on the idea of self-supervised learning to overcome the label mismatch problem. Experimental results on VOiCES Challenge 2019 demonstrate that our proposed method can help learn more discriminative and robust speaker embeddings.