【1】 Transferability of Adversarial Attacks on Synthetic Speech Detection作者:Jiacheng Deng,Shunyi Chen,Li Dong,Diqun Yan,Rangding Wang机构:Department of Information Science and Engineering, Ningbo University备注:5 pages, submit to Interspeech2022摘要:合成语音检测是音频安全领域最重要的研究课题之一。同时,深层神经网络容易受到敌对攻击。因此,我们建立了一个全面的基准来评估对抗性攻击在合成语音检测任务中的可转移性。具体来说,我们试图调查:1)不同功能之间的对抗性攻击的可转移性。2) 特征提取参数的变化对对抗性攻击可转移性的影响。3) 剪切或自填充操作对敌对攻击可转移性的影响。通过这些分析,我们总结了合成语音检测器的弱点以及对抗性攻击的可转移性行为,为未来的研究提供了见解。更多详情请访问https://gitee.com/djc_QRICK/Attack-Transferability-On-Synthetic-Detection.摘要:Synthetic speech detection is one of the most important research problems in audio security. Meanwhile, deep neural networks are vulnerable to adversarial attacks. Therefore, we establish a comprehensive benchmark to evaluate the transferability of adversarial attacks on the synthetic speech detection task. Specifically, we attempt to investigate: 1) The transferability of adversarial attacks between different features. 2) The influence of varying extraction hyperparameters of features on the transferability of adversarial attacks. 3) The effect of clipping or self-padding operation on the transferability of adversarial attacks. By performing these analyses, we summarise the weaknesses of synthetic speech detectors and the transferability behaviours of adversarial attacks, which provide insights for future research. More details can be found at https://gitee.com/djc_QRICK/Attack-Transferability-On-Synthetic-Detection.
【2】 L3-Net Deep Audio Embeddings to Improve COVID-19 Detection from Smartphone Data
标题:L3-Net深度音频嵌入改进智能手机数据中的新冠肺炎检测
链接:https://arxiv.org/abs/2205.07682
作者:Mattia Giovanni Campana,Andrea Rovati,Franca Delmastro,Elena Pagani机构:∗Institute for Informatics and Telematics of the National Research Council of Italy (IIT-CNR), Pisa, Italy, †Computer Science Department, University of Milano, Milan, Italy备注:accepted for IEEE SMARTCOMP 2022摘要:智能手机和可穿戴设备,以及人工智能,可以通过实施低成本、普及的解决方案,在早期阶段识别新疾病的发展,并有可能避免新疫情的出现,在大流行控制中成为一个游戏规则改变者。最近的一些研究表明,通过使用机器学习和手工制作的声学特征,有望从声音和咳嗽中检测出2019冠状病毒疾病的诊断信号。在本文中,我们决定研究最近提出的深度嵌入模型L3网络从原始呼吸音频记录中自动提取有意义的特征的能力,以提高标准机器学习分类器从智能手机数据中区分2019冠状病毒疾病阳性和阴性受试者的性能。我们在3个数据集上评估了该模型,并将所得结果与两个参考文献的结果进行了比较。结果表明,在一组独立于受试者的实验中,L3网络与手工制作的功能相结合,在AUC方面超过了其他作品28.57%的性能。这一结果为进一步研究不同深度音频嵌入以及不同疾病的自动检测奠定了基础。摘要:Smartphones and wearable devices, along with Artificial Intelligence, can represent a game-changer in the pandemic control, by implementing low-cost and pervasive solutions to recognize the development of new diseases at their early stages and by potentially avoiding the rise of new outbreaks. Some recent works show promise in detecting diagnostic signals of COVID-19 from voice and coughs by using machine learning and hand-crafted acoustic features. In this paper, we decided to investigate the capabilities of the recently proposed deep embedding model L3-Net to automatically extract meaningful features from raw respiratory audio recordings in order to improve the performances of standard machine learning classifiers in discriminating between COVID-19 positive and negative subjects from smartphone data. We evaluated the proposed model on 3 datasets, comparing the obtained results with those of two reference works. Results show that the combination of L3-Net with hand-crafted features overcomes the performance of the other works of 28.57% in terms of AUC in a set of subject-independent experiments. This result paves the way to further investigation on different deep audio embeddings, also for the automatic detection of different diseases.
【3】 A Fast Attention Network for Joint Intent Detection and Slot Filling on Edge Devices
标题:边缘设备联合意图检测和空位填充的快速注意力网络
链接:https://arxiv.org/abs/2205.07646
作者:Liang Huang,Senjie Liang,Feiyang Ye,Nan Gao摘要:意图检测和时隙填充是自然语言理解中的两个主要任务,在面向任务的对话系统中起着至关重要的作用。这两个任务的联合学习可以提高推理的准确性,在最近的工作中很流行。然而,大多数联合模型忽略了推理延迟,无法满足在边缘部署对话系统的需要。在本文中,我们提出了一种快速注意网络(FAN),用于联合意图检测和时隙填充任务,同时保证准确性和延迟。具体来说,我们引入了一个干净且参数优化的注意模块,以增强意图和时隙之间的信息交换,将语义准确性提高了2%以上。风扇可以在不同的编码器上实现,并在每个速度级别提供更精确的模型。我们在Jetson Nano平台上的实验表明,FAN每秒能推断出15次话语,但准确度下降很小,这表明了它在边缘设备上的有效性和效率。摘要:Intent detection and slot filling are two main tasks in natural language understanding and play an essential role in task-oriented dialogue systems. The joint learning of both tasks can improve inference accuracy and is popular in recent works. However, most joint models ignore the inference latency and cannot meet the need to deploy dialogue systems at the edge. In this paper, we propose a Fast Attention Network (FAN) for joint intent detection and slot filling tasks, guaranteeing both accuracy and latency. Specifically, we introduce a clean and parameter-refined attention module to enhance the information exchange between intent and slot, improving semantic accuracy by more than 2%. FAN can be implemented on different encoders and delivers more accurate models at every speed level. Our experiments on the Jetson Nano platform show that FAN inferences fifteen utterances per second with a small accuracy drop, showing its effectiveness and efficiency on edge devices.
【4】 PRISM: Pre-trained Indeterminate Speaker Representation Model for Speaker Diarization and Speaker Verification
标题:PRISM:用于说话人二值化和说话人确认的预训练不确定说话人表示模型
链接:https://arxiv.org/abs/2205.07450
作者:Siqi Zheng,Hongbin Suo,Qian Chen机构:Speech Lab, Alibaba Group摘要:说话人嵌入是与说话人相关的任务(如验证、聚类和日记化)的一个基本特征。传统上,说话人嵌入被表示为高维空间中的固定向量。这可能会导致有偏见的估计,尤其是在处理较短的话语时。在本文中,我们建议将说话人的话语表示为“浮动”向量,其状态在不知道上下文的情况下是不确定的。说话人陈述的状态由说话人自身、同一说话人的其他讲话以及与之进行比较的其他说话人共同决定。演讲的内容也有助于确定说话人陈述的最终状态。我们预先训练了一个不确定的说话人表示模型,该模型根据上下文估计话语的状态。预训练的模型可以针对下游任务进行微调,例如说话人验证、说话人聚类和说话人二值化。在所有下游任务中都观察到了实质性的改进。摘要:Speaker embedding has been a fundamental feature for speaker-related tasks such as verification, clustering, and diarization. Traditionally, speaker embeddings are represented as fixed vectors in high-dimensional space. This could lead to biased estimations, especially when handling shorter utterances. In this paper we propose to represent a speaker utterance as "floating" vector whose state is indeterminate without knowing the context. The state of a speaker representation is jointly determined by itself, other speech from the same speaker, as well as other speakers it is being compared to. The content of the speech also contributes to determining the final state of a speaker representation. We pre-train an indeterminate speaker representation model that estimates the state of an utterance based on the context. The pre-trained model can be fine-tuned for downstream tasks such as speaker verification, speaker clustering, and speaker diarization. Substantial improvements are observed across all downstream tasks.
【5】 cMelGAN: An Efficient Conditional Generative Model Based on Mel Spectrograms
标题:CMelGAN:一种基于Mel谱图的高效条件生成模型
链接:https://arxiv.org/abs/2205.07319
作者:Tracy Qian,Jackson Kaunismaa,Tony Chung机构:___________________________________________________________________________________________________________ 1Department of Engineering Science Machine Intelligence, University of Toronto 1摘要:在机器学习领域分析音乐是一个非常困难的问题,需要考虑许多约束条件。音频数据的性质,具有非常高的维度和广泛变化的结构尺度,是建模如此困难的主要原因之一。机器学习在音乐中有很多应用,比如对一段音乐的情绪进行分类、有条件的音乐生成或流行预测。该项目的目标是开发一个基于Mel谱图的音乐类型条件生成模型,并通过将其与使用基于音符表示的现有生成音乐模型进行比较来评估其性能。我们最初实现了一个基于RNN的自回归生成模型,称为MelNet。然而,由于其速度慢、输出保真度低,我们决定创建一种基于MelGAN[4]和条件GAN结构的新的完全卷积结构,称为cMelGAN。摘要:Analysing music in the field of machine learning is a very difficult problem with numerous constraints to consider. The nature of audio data, with its very high dimensionality and widely varying scales of structure, is one of the primary reasons why it is so difficult to model. There are many applications of machine learning in music, like the classifying the mood of a piece of music, conditional music generation, or popularity prediction. The goal for this project was to develop a genre-conditional generative model of music based on Mel spectrograms and evaluate its performance by comparing it to existing generative music models that use note-based representations. We initially implemented an autoregressive, RNN-based generative model called MelNet . However, due to its slow speed and low fidelity output, we decided to create a new, fully convolutional architecture that is based on the MelGAN [4] and conditional GAN architectures, called cMelGAN.
【6】 Conditional Vector Graphics Generation for Music Cover Images
标题:音乐封面图像的条件向量图形生成
链接:https://arxiv.org/abs/2205.07301
作者:Valeria Efimova,Ivan Jarsky,Ilya Bizyaev,Andrey Filchenkov摘要:生成性对抗网络(GAN)推动了计算机图像合成领域的快速发展。由于几乎所有现有的图像合成算法都将图像视为像素矩阵,因此高分辨率图像合成非常复杂。一个很好的替代方法是矢量图像。然而,它们属于高度复杂的参数空间,这是GANs解决矢量图形合成任务的一个限制。在本文中,我们考虑了一个特定的应用领域,它极大地软化了这一限制,允许使用矢量图像合成。音乐封面图像应符合互联网流媒体服务和打印标准的要求,这意味着图形材料的高分辨率,而不需要对此类图像的内容提出任何额外要求。现有的音乐封面图像生成服务本身不分析曲目;然而,一些服务大多只考虑类型标签。为了将音乐封面生成为反映音乐并由简单几何对象组成的矢量图像,我们提出了一种基于GAN的算法CoverGAN。结果图像的评估基于它们与音乐的对应关系,并与根据标题或歌词生成的AttnGAN和DALL-E文本图像进行比较。此外,根据生成的封面图像与音乐曲目的对应关系,对CoverGAN发现的图案的意义进行了评估。听众对所提出的算法生成的音乐封面的评价非常满意,并且与曲目相对应。音乐封面图片生成代码和演示可在https://github.com/IzhanVarsky/CoverGAN.摘要:Generative Adversarial Networks (GAN) have motivated a rapid growth of the domain of computer image synthesis. As almost all the existing image synthesis algorithms consider an image as a pixel matrix, the high-resolution image synthesis is complicated.A good alternative can be vector images. However, they belong to the highly sophisticated parametric space, which is a restriction for solving the task of synthesizing vector graphics by GANs. In this paper, we consider a specific application domain that softens this restriction dramatically allowing the usage of vector image synthesis. Music cover images should meet the requirements of Internet streaming services and printing standards, which imply high resolution of graphic materials without any additional requirements on the content of such images. Existing music cover image generation services do not analyze tracks themselves; however, some services mostly consider only genre tags. To generate music covers as vector images that reflect the music and consist of simple geometric objects, we suggest a GAN-based algorithm called CoverGAN. The assessment of resulting images is based on their correspondence to the music compared with AttnGAN and DALL-E text-to-image generation according to title or lyrics. Moreover, the significance of the patterns found by CoverGAN has been evaluated in terms of the correspondence of the generated cover images to the musical tracks. Listeners evaluate the music covers generated by the proposed algorithm as quite satisfactory and corresponding to the tracks. Music cover images generation code and demo are available at https://github.com/IzhanVarsky/CoverGAN.
【7】 Multiformer: A Head-Configurable Transformer-Based Model for Direct Speech Translation
标题:多转换器:一种基于头部可配置转换器的直接语音翻译模型
链接:https://arxiv.org/abs/2205.07100
作者:Gerard Sant,Gerard I. Gállego,Belen Alastruey,Marta R. Costa-Jussà机构:TALP Research Center, Universitat Politècnica de Catalunya, Barcelona摘要:基于Transformer的模型已经在自然语言处理的几个领域取得了最新的成果。然而,它在语音任务中的直接应用并非微不足道。这种序列的性质带来了诸如长序列长度和相邻令牌之间的冗余等问题。因此,我们认为,常规的自我注意机制可能并不适合这种情况。人们提出了不同的方法来克服这些问题,例如使用有效的注意机制。然而,使用这些方法通常会带来成本,这是由于信息丢失导致的性能降低。在这项研究中,我们提出了Multiformer,这是一个基于Transformer的模型,允许在每个头部使用不同的注意机制。通过这样做,该模型能够将自我关注偏向于提取更多样化的令牌交互,并且减少了信息损失。最后,我们对head贡献进行了分析,并观察到所有head相关性均匀分布的架构可以获得更好的结果。我们的研究结果表明,不同头部和层次的混合注意力模式比我们的基线高出0.7 BLEU。摘要:Transformer-based models have been achieving state-of-the-art results in several fields of Natural Language Processing. However, its direct application to speech tasks is not trivial. The nature of this sequences carries problems such as long sequence lengths and redundancy between adjacent tokens. Therefore, we believe that regular self-attention mechanism might not be well suited for it. Different approaches have been proposed to overcome these problems, such as the use of efficient attention mechanisms. However, the use of these methods usually comes with a cost, which is a performance reduction caused by information loss. In this study, we present the Multiformer, a Transformer-based model which allows the use of different attention mechanisms on each head. By doing this, the model is able to bias the self-attention towards the extraction of more diverse token interactions, and the information loss is reduced. Finally, we perform an analysis of the head contributions, and we observe that those architectures where all heads relevance is uniformly distributed obtain better results. Our results show that mixing attention patterns along the different heads and layers outperforms our baseline by up to 0.7 BLEU.
【8】 Improved Consistency Training for Semi-Supervised Sequence-to-Sequence ASR via Speech Chain Reconstruction and Self-Transcribing
标题:基于语音链重构和自转录的半监督序列间ASR一致性训练
链接:https://arxiv.org/abs/2205.06963
作者:Heli Qi,Sashi Novitasari,Sakriani Sakti,Satoshi Nakamura机构:Nara Institute of Science and Technology, Japan, Japan Advanced Institute of Science and Technology, Japan备注:Submitted to INTERSPEECH 2022摘要:一致性正则化最近被应用于半监督序列对序列(S2S)自动语音识别(ASR)。这一原理鼓励ASR模型对具有不同扰动的相同输入语音输出类似的预测。现有的半监督S2S ASR范式利用SPECAMULTE作为数据增强,需要静态教师模型为未翻译的语音生成伪转录本。然而,这种范式未能充分利用一致性正则化。首先,SpecAugment的掩蔽操作可能会破坏语音的语言内容,从而影响伪标签的质量。其次,S2S ASR需要输入语音和前缀标记来进行下一次预测。在一致性训练期间,离线教师模型生成的静态前缀标记无法与动态伪标签匹配。在这项工作中,我们提出了一种改进的半监督S2S ASR一致性训练范式。我们利用语音链重建作为弱增强来生成高质量的伪标签。此外,我们还证明了学生ASR模型产生的动态伪转录本有利于一致性训练。在LJSpeech和LibriSpeech语料库上的实验表明,与监督基线相比,我们改进的范例在单说话人环境下的CER提高了12.2%,在多说话人环境下的CER提高了38.6%。摘要:Consistency regularization has recently been applied to semi-supervised sequence-to-sequence (S2S) automatic speech recognition (ASR). This principle encourages an ASR model to output similar predictions for the same input speech with different perturbations. The existing paradigm of semi-supervised S2S ASR utilizes SpecAugment as data augmentation and requires a static teacher model to produce pseudo transcripts for untranscribed speech. However, this paradigm fails to take full advantage of consistency regularization. First, the masking operations of SpecAugment may damage the linguistic contents of the speech, thus influencing the quality of pseudo labels. Second, S2S ASR requires both input speech and prefix tokens to make the next prediction. The static prefix tokens made by the offline teacher model cannot match dynamic pseudo labels during consistency training. In this work, we propose an improved consistency training paradigm of semi-supervised S2S ASR. We utilize speech chain reconstruction as the weak augmentation to generate high-quality pseudo labels. Moreover, we demonstrate that dynamic pseudo transcripts produced by the student ASR model benefit the consistency training. Experiments on LJSpeech and LibriSpeech corpora show that compared to supervised baselines, our improved paradigm achieves a 12.2% CER improvement in the single-speaker setting and 38.6% in the multi-speaker setting.
【9】 Learning Representations for New Sound Classes With Continual Self-Supervised Learning
标题:具有连续自监督学习的新声音类的学习表征
链接:https://arxiv.org/abs/2205.07390
作者:Zhepei Wang,Cem Subakan,Xilin Jiang,Junkai Wu,Efthymios Tzinis,Mirco Ravanelli,Paris Smaragdis机构:♯ University of Illinois at Urbana-Champaign, USA; ♭ Mila-Quebec AIInstitute, Canada; ♮ Universit´e de Sherbrooke, Canada; † Concordia University, USA; χ Universit´e de Montr´eal备注:Submitted to IEEE Signal Processing Letters摘要:在本文中,我们提出了一个自监督学习框架,用于不断学习新声音类的表示。该系统依赖于一个不断训练的神经编码器,该编码器以基于相似性的学习目标进行训练,而不使用标签。我们发现,与完全监督的方法相比,用该方法学习的表征具有更好的泛化性,并且不太容易发生灾难性遗忘。值得注意的是,我们的技术不存储过去的数据或模型,并且比基于蒸馏的方法计算效率更高。为了准确地评估系统性能,除了使用现有的协议外,我们还提出了两个现实的评估协议,它们只使用少量的标记数据来模拟实际用例。摘要:In this paper, we present a self-supervised learning framework for continually learning representations for new sound classes. The proposed system relies on a continually trained neural encoder that is trained with similarity-based learning objectives without using labels. We show that representations learned with the proposed method generalize better and are less susceptible to catastrophic forgetting than fully-supervised approaches. Remarkably, our technique does not store past data or models and is more computationally efficient than distillation-based methods. To accurately assess the system performance, in addition to using existing protocols, we propose two realistic evaluation protocols that use only a small amount of labeled data to simulate practical use cases.
【10】 GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech Synthesis
标题:GenerSpeech:面向通用型域外文语合成的风格转换
链接:https://arxiv.org/abs/2205.07211
作者:Rongjie Huang,Yi Ren,Jinglin Liu,Chenye Cui,Zhou Zhao机构:which attracts broad interest in the machine learn-Equal contribution 1Zhejiang University摘要:面向域外(OOD)语音合成的风格转换(Style transfer for out-of-domain,Style transfer for out-domain,简称Style transfer for out-domain)旨在从声学参考中生成具有不可见风格(例如说话人身份、情感和韵律)的语音样本,同时面临以下挑战:1)表达性语音中高度动态的风格特征难以建模和转换;2)TTS模型应足够稳健,能够处理与源数据不同的各种OOD条件。本文提出了一种面向高保真零炮风格的定制语音传输的文本到语音模型GenerSpeech。GenerSpeech通过引入两个组件将语音变化分解为风格不可知和风格特定的部分:1)一个多级风格适配器,用于有效地建模大量风格条件,包括全局说话人和情感特征,以及局部(话语、音素和词级)细粒度韵律表示;2)具有混合样式层规范化的可概括内容适配器,以消除语言内容表示中的样式信息,从而提高模型的泛化能力。我们对Zero-Shot风格转换的评估表明,GeneralSpeech在音频质量和风格相似性方面超过了最先进的模型。对自适应风格转换的扩展研究进一步表明,GenerSpeech在Few-Shot数据环境下表现强劲。音频样本可在\url上获取{https://GenerSpeech.github.io/}摘要:Style transfer for out-of-domain (OOD) speech synthesis aims to generate speech samples with unseen style (e.g., speaker identity, emotion, and prosody) derived from an acoustic reference, while facing the following challenges: 1) The highly dynamic style features in expressive voice are difficult to model and transfer; and 2) the TTS models should be robust enough to handle diverse OOD conditions that differ from the source data. This paper proposes GenerSpeech, a text-to-speech model towards high-fidelity zero-shot style transfer of OOD custom voice. GenerSpeech decomposes the speech variation into the style-agnostic and style-specific parts by introducing two components: 1) a multi-level style adaptor to efficiently model a large range of style conditions, including global speaker and emotion characteristics, and the local (utterance, phoneme, and word-level) fine-grained prosodic representations; and 2) a generalizable content adaptor with Mix-Style Layer Normalization to eliminate style information in the linguistic content representation and thus improve model generalization. Our evaluations on zero-shot style transfer demonstrate that GenerSpeech surpasses the state-of-the-art models in terms of audio quality and style similarity. The extension studies to adaptive style transfer further show that GenerSpeech performs robustly in the few-shot data setting. Audio samples are available at \url{https://GenerSpeech.github.io/}
【11】 Learning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERT
标题:利用AV-Hubert学习基于Lip的视听说话人嵌入
链接:https://arxiv.org/abs/2205.07180
作者:Bowen Shi,Abdelrahman Mohamed,Wei-Ning Hsu机构:Toyota Technological Institute at Chicago, Meta AI备注:Submitted to Interspeech摘要:本文研究了视听说话人表征学习中的自我监督预训练,其中显示说话人口腔区域的视觉流与语音一起用作输入。我们的研究集中在视听隐藏单元BERT(AV HuBERT)方法上,这是一种最近开发的通用视听语音预训练框架。我们进行了广泛的实验,探索训练前和视觉模式的有效性。实验结果表明,AV-HuBERT可以很好地推广到与说话人相关的下游任务,在纯音频和视听说话人验证中,标签效率提高了大约十倍。在嘈杂的环境中,我们将噪声和噪声降低了75%,甚至在视觉环境中,我们将噪声和噪声降低了75%。我们的代码和模型将公开。摘要:This paper investigates self-supervised pre-training for audio-visual speaker representation learning where a visual stream showing the speaker's mouth area is used alongside speech as inputs. Our study focuses on the Audio-Visual Hidden Unit BERT (AV-HuBERT) approach, a recently developed general-purpose audio-visual speech pre-training framework. We conducted extensive experiments probing the effectiveness of pre-training and visual modality. Experimental results suggest that AV-HuBERT generalizes decently to speaker related downstream tasks, improving label efficiency by roughly ten fold for both audio-only and audio-visual speaker verification. We also show that incorporating visual information, even just the lip area, greatly improves the performance and noise robustness, reducing EER by 38% in the clean condition and 75% in noisy conditions. Our code and models will be publicly available.
【12】 Collar-aware Training for Streaming Speaker Change Detection in Broadcast Speech
标题:广播语音中流说话人变化检测的项圈感知训练
链接:https://arxiv.org/abs/2205.07086
作者:Joonas Kalda,Tanel Alumäe机构:Department of Software Science, Tallinn University of Technology, Estonia备注:Accepted to Speaker Odyssey 2022摘要:本文提出了一种新的说话人变化检测模型训练方法。说话人变化检测通常被视为一个二进制序列标签问题。这种方法的主要挑战是,由于大多数帧不包括说话人的变化,说话人轮次之间的沉默和数据不平衡导致注释的变化点模糊。传统的训练方法通过人为增加训练数据中正面标签的比例来解决这些问题。相反,提出的方法使用了一个目标函数,该函数鼓励模型预测指定项圈内的单个阳性标签。这是通过边缘化所有可能的子序列来实现的,这些子序列在衣领内只有一个阳性标签。在英语和爱沙尼亚语数据集上的实验表明,与传统的训练方法相比有了很大的改进。此外,模型输出的峰值集中在单个帧上,无需进行后处理以找到准确的预测变化点,这对流媒体应用程序特别有用。摘要:In this paper, we present a novel training method for speaker change detection models. Speaker change detection is often viewed as a binary sequence labelling problem. The main challenges with this approach are the vagueness of annotated change points caused by the silences between speaker turns and imbalanced data due to the majority of frames not including a speaker change. Conventional training methods tackle these by artificially increasing the proportion of positive labels in the training data. Instead, the proposed method uses an objective function which encourages the model to predict a single positive label within a specified collar. This is done by marginalizing over all possible subsequences that have exactly one positive label within the collar. Experiments on English and Estonian datasets show large improvements over the conventional training method. Additionally, the model outputs have peaks concentrated to a single frame, removing the need for post-processing to find the exact predicted change point which is particularly useful for streaming applications.
【13】 Task splitting for DNN-based acoustic echo and noise removal
标题:基于DNN的声学回波和噪声去除任务分解
链接:https://arxiv.org/abs/2205.06931
作者:Sebastian Braun,Maria Luis Valero机构:Microsoft Corporation, USA备注:submitted to IWAENC 2022摘要:神经网络为单任务语音增强带来了巨大的性能提升,如噪声抑制和回声消除(AEC)。在这项工作中,我们评估使用单个关节或单独的模块来解决这些问题是否更有用。我们描述了不同的可能实现,并深入了解了它们的性能和效率。我们发现,使用一个单独的回声消除模块和一个用于去除噪声和残余回声的模块,可以减少近端语音失真和更好的回声抑制,尤其是对于双重通话。摘要:Neural networks have led to tremendous performance gains for single-task speech enhancement, such as noise suppression and acoustic echo cancellation (AEC). In this work, we evaluate whether it is more useful to use a single joint or separate modules to tackle these problems. We describe different possible implementations and give insights into their performance and efficiency. We show that using a separate echo cancellation module and a module for noise and residual echo removal results in less near-end speech distortion and better echo suppression, especially for double-talk.
【1】 Learning Representations for New Sound Classes With Continual Self-Supervised Learning标题:具有连续自监督学习的新声音类的学习表征
链接:https://arxiv.org/abs/2205.07390
作者:Zhepei Wang,Cem Subakan,Xilin Jiang,Junkai Wu,Efthymios Tzinis,Mirco Ravanelli,Paris Smaragdis机构:♯ University of Illinois at Urbana-Champaign, USA; ♭ Mila-Quebec AIInstitute, Canada; ♮ Universit´e de Sherbrooke, Canada; † Concordia University, USA; χ Universit´e de Montr´eal备注:Submitted to IEEE Signal Processing Letters摘要:在本文中,我们提出了一个自监督学习框架,用于不断学习新声音类的表示。该系统依赖于一个不断训练的神经编码器,该编码器以基于相似性的学习目标进行训练,而不使用标签。我们发现,与完全监督的方法相比,用该方法学习的表征具有更好的泛化性,并且不太容易发生灾难性遗忘。值得注意的是,我们的技术不存储过去的数据或模型,并且比基于蒸馏的方法计算效率更高。为了准确地评估系统性能,除了使用现有的协议外,我们还提出了两个现实的评估协议,它们只使用少量的标记数据来模拟实际用例。摘要:In this paper, we present a self-supervised learning framework for continually learning representations for new sound classes. The proposed system relies on a continually trained neural encoder that is trained with similarity-based learning objectives without using labels. We show that representations learned with the proposed method generalize better and are less susceptible to catastrophic forgetting than fully-supervised approaches. Remarkably, our technique does not store past data or models and is more computationally efficient than distillation-based methods. To accurately assess the system performance, in addition to using existing protocols, we propose two realistic evaluation protocols that use only a small amount of labeled data to simulate practical use cases.
【2】 GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech Synthesis
标题:GenerSpeech:面向通用型域外文语合成的风格转换
链接:https://arxiv.org/abs/2205.07211
作者:Rongjie Huang,Yi Ren,Jinglin Liu,Chenye Cui,Zhou Zhao机构:which attracts broad interest in the machine learn-Equal contribution 1Zhejiang University摘要:面向域外(OOD)语音合成的风格转换(Style transfer for out-of-domain,Style transfer for out-domain,简称Style transfer for out-domain)旨在从声学参考中生成具有不可见风格(例如说话人身份、情感和韵律)的语音样本,同时面临以下挑战:1)表达性语音中高度动态的风格特征难以建模和转换;2)TTS模型应足够稳健,能够处理与源数据不同的各种OOD条件。本文提出了一种面向高保真零炮风格的定制语音传输的文本到语音模型GenerSpeech。GenerSpeech通过引入两个组件将语音变化分解为风格不可知和风格特定的部分:1)一个多级风格适配器,用于有效地建模大量风格条件,包括全局说话人和情感特征,以及局部(话语、音素和词级)细粒度韵律表示;2)具有混合样式层规范化的可概括内容适配器,以消除语言内容表示中的样式信息,从而提高模型的泛化能力。我们对Zero-Shot风格转换的评估表明,GeneralSpeech在音频质量和风格相似性方面超过了最先进的模型。对自适应风格转换的扩展研究进一步表明,GenerSpeech在Few-Shot数据环境下表现强劲。音频样本可在\url上获取{https://GenerSpeech.github.io/}摘要:Style transfer for out-of-domain (OOD) speech synthesis aims to generate speech samples with unseen style (e.g., speaker identity, emotion, and prosody) derived from an acoustic reference, while facing the following challenges: 1) The highly dynamic style features in expressive voice are difficult to model and transfer; and 2) the TTS models should be robust enough to handle diverse OOD conditions that differ from the source data. This paper proposes GenerSpeech, a text-to-speech model towards high-fidelity zero-shot style transfer of OOD custom voice. GenerSpeech decomposes the speech variation into the style-agnostic and style-specific parts by introducing two components: 1) a multi-level style adaptor to efficiently model a large range of style conditions, including global speaker and emotion characteristics, and the local (utterance, phoneme, and word-level) fine-grained prosodic representations; and 2) a generalizable content adaptor with Mix-Style Layer Normalization to eliminate style information in the linguistic content representation and thus improve model generalization. Our evaluations on zero-shot style transfer demonstrate that GenerSpeech surpasses the state-of-the-art models in terms of audio quality and style similarity. The extension studies to adaptive style transfer further show that GenerSpeech performs robustly in the few-shot data setting. Audio samples are available at \url{https://GenerSpeech.github.io/}
【3】 Learning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERT
标题:利用AV-Hubert学习基于Lip的视听说话人嵌入
链接:https://arxiv.org/abs/2205.07180
作者:Bowen Shi,Abdelrahman Mohamed,Wei-Ning Hsu机构:Toyota Technological Institute at Chicago, Meta AI备注:Submitted to Interspeech摘要:本文研究了视听说话人表征学习中的自我监督预训练,其中显示说话人口腔区域的视觉流与语音一起用作输入。我们的研究集中在视听隐藏单元BERT(AV HuBERT)方法上,这是一种最近开发的通用视听语音预训练框架。我们进行了广泛的实验,探索训练前和视觉模式的有效性。实验结果表明,AV-HuBERT可以很好地推广到与说话人相关的下游任务,在纯音频和视听说话人验证中,标签效率提高了大约十倍。在嘈杂的环境中,我们将噪声和噪声降低了75%,甚至在视觉环境中,我们将噪声和噪声降低了75%。我们的代码和模型将公开。摘要:This paper investigates self-supervised pre-training for audio-visual speaker representation learning where a visual stream showing the speaker's mouth area is used alongside speech as inputs. Our study focuses on the Audio-Visual Hidden Unit BERT (AV-HuBERT) approach, a recently developed general-purpose audio-visual speech pre-training framework. We conducted extensive experiments probing the effectiveness of pre-training and visual modality. Experimental results suggest that AV-HuBERT generalizes decently to speaker related downstream tasks, improving label efficiency by roughly ten fold for both audio-only and audio-visual speaker verification. We also show that incorporating visual information, even just the lip area, greatly improves the performance and noise robustness, reducing EER by 38% in the clean condition and 75% in noisy conditions. Our code and models will be publicly available.
【4】 Collar-aware Training for Streaming Speaker Change Detection in Broadcast Speech
标题:广播语音中流说话人变化检测的项圈感知训练
链接:https://arxiv.org/abs/2205.07086
作者:Joonas Kalda,Tanel Alumäe机构:Department of Software Science, Tallinn University of Technology, Estonia备注:Accepted to Speaker Odyssey 2022摘要:本文提出了一种新的说话人变化检测模型训练方法。说话人变化检测通常被视为一个二进制序列标签问题。这种方法的主要挑战是,由于大多数帧不包括说话人的变化,说话人轮次之间的沉默和数据不平衡导致注释的变化点模糊。传统的训练方法通过人为增加训练数据中正面标签的比例来解决这些问题。相反,提出的方法使用了一个目标函数,该函数鼓励模型预测指定项圈内的单个阳性标签。这是通过边缘化所有可能的子序列来实现的,这些子序列在衣领内只有一个阳性标签。在英语和爱沙尼亚语数据集上的实验表明,与传统的训练方法相比有了很大的改进。此外,模型输出的峰值集中在单个帧上,无需进行后处理以找到准确的预测变化点,这对流媒体应用程序特别有用。摘要:In this paper, we present a novel training method for speaker change detection models. Speaker change detection is often viewed as a binary sequence labelling problem. The main challenges with this approach are the vagueness of annotated change points caused by the silences between speaker turns and imbalanced data due to the majority of frames not including a speaker change. Conventional training methods tackle these by artificially increasing the proportion of positive labels in the training data. Instead, the proposed method uses an objective function which encourages the model to predict a single positive label within a specified collar. This is done by marginalizing over all possible subsequences that have exactly one positive label within the collar. Experiments on English and Estonian datasets show large improvements over the conventional training method. Additionally, the model outputs have peaks concentrated to a single frame, removing the need for post-processing to find the exact predicted change point which is particularly useful for streaming applications.
【5】 Pretraining Approaches for Spoken Language Recognition: TalTech Submission to the OLR 2021 Challenge
标题:口语识别的预训练方法:TalTech提交给OLR 2021挑战赛
链接:https://arxiv.org/abs/2205.07083
作者:Tanel Alumäe,Kunnar Kukk机构:Department of Software Science, Tallinn University of Technology, Estonia备注:Accepted to Speaker Odyssey 2022摘要:本文研究了口语识别的不同预训练方法。这篇论文是基于我们对2021东方语言识别挑战的提交。我们参与了挑战的两个方面:受限语言识别和无约束语言识别。对于受约束的轨迹,我们首先使用提供的训练数据训练了一个基于一致性的多语言自动语音识别(ASR)编码器-解码器模型,该模型具有可用的文本。然后,针对语言识别任务,对多语言ASR模型的共享编码器进行了微调。对于无约束任务,我们既依赖外部可用的预训练模型,也依赖外部数据:多语言XLSR-53 wav2vec2。0模型在VoxLingua107语料库上进行了微调,用于语言识别任务,最后在提供的目标语言训练数据上进行微调,并添加了公共语音数据。我们的主要指标$C_{\rm avg}$在测试集中的值是0.0079(对于受约束的任务)和0.0119(对于无约束的任务),这导致了两个排名中的第二名。在后评估实验中,我们研究了训练准确的后端模型所需的目标语言数据量,多语言训练前数据的重要性,并将不同的模型作为微调起点进行比较。摘要:This paper investigates different pretraining approaches to spoken language identification. The paper is based on our submission to the Oriental Language Recognition 2021 Challenge. We participated in two tracks of the challenge: constrained and unconstrained language recognition. For the constrained track, we first trained a Conformer-based encoder-decoder model for multilingual automatic speech recognition (ASR), using the provided training data that had transcripts available. The shared encoder of the multilingual ASR model was then finetuned for the language identification task. For the unconstrained task, we relied on both externally available pretrained models as well as external data: the multilingual XLSR-53 wav2vec2.0 model was finetuned on the VoxLingua107 corpus for the language recognition task, and finally finetuned on the provided target language training data, augmented with CommonVoice data. Our primary metric $C_{\rm avg}$ values on the Test set are 0.0079 for the constrained task and 0.0119 for the unconstrained task which resulted in the second place in both rankings. In post-evaluation experiments, we study the amount of target language data needed for training an accurate backend model, the importance of multilingual pretraining data, and compare different models as finetuning starting points.
【6】 Task splitting for DNN-based acoustic echo and noise removal
标题:基于DNN的声学回波和噪声去除任务分解
链接:https://arxiv.org/abs/2205.06931
作者:Sebastian Braun,Maria Luis Valero机构:Microsoft Corporation, USA备注:submitted to IWAENC 2022摘要:神经网络为单任务语音增强带来了巨大的性能提升,如噪声抑制和回声消除(AEC)。在这项工作中,我们评估使用单个关节或单独的模块来解决这些问题是否更有用。我们描述了不同的可能实现,并深入了解了它们的性能和效率。我们发现,使用一个单独的回声消除模块和一个用于去除噪声和残余回声的模块,可以减少近端语音失真和更好的回声抑制,尤其是对于双重通话。摘要:Neural networks have led to tremendous performance gains for single-task speech enhancement, such as noise suppression and acoustic echo cancellation (AEC). In this work, we evaluate whether it is more useful to use a single joint or separate modules to tackle these problems. We describe different possible implementations and give insights into their performance and efficiency. We show that using a separate echo cancellation module and a module for noise and residual echo removal results in less near-end speech distortion and better echo suppression, especially for double-talk.
【7】 Transferability of Adversarial Attacks on Synthetic Speech Detection
标题:合成语音检测中对抗性攻击的可转移性
链接:https://arxiv.org/abs/2205.07711
作者:Jiacheng Deng,Shunyi Chen,Li Dong,Diqun Yan,Rangding Wang机构:Department of Information Science and Engineering, Ningbo University备注:5 pages, submit to Interspeech2022摘要:合成语音检测是音频安全领域最重要的研究课题之一。同时,深层神经网络容易受到敌对攻击。因此,我们建立了一个全面的基准来评估对抗性攻击在合成语音检测任务中的可转移性。具体来说,我们试图调查:1)不同功能之间的对抗性攻击的可转移性。2) 特征提取参数的变化对对抗性攻击可转移性的影响。3) 剪切或自填充操作对敌对攻击可转移性的影响。通过这些分析,我们总结了合成语音检测器的弱点以及对抗性攻击的可转移性行为,为未来的研究提供了见解。更多详情请访问https://gitee.com/djc_QRICK/Attack-Transferability-On-Synthetic-Detection.摘要:Synthetic speech detection is one of the most important research problems in audio security. Meanwhile, deep neural networks are vulnerable to adversarial attacks. Therefore, we establish a comprehensive benchmark to evaluate the transferability of adversarial attacks on the synthetic speech detection task. Specifically, we attempt to investigate: 1) The transferability of adversarial attacks between different features. 2) The influence of varying extraction hyperparameters of features on the transferability of adversarial attacks. 3) The effect of clipping or self-padding operation on the transferability of adversarial attacks. By performing these analyses, we summarise the weaknesses of synthetic speech detectors and the transferability behaviours of adversarial attacks, which provide insights for future research. More details can be found at https://gitee.com/djc_QRICK/Attack-Transferability-On-Synthetic-Detection.
【8】 L3-Net Deep Audio Embeddings to Improve COVID-19 Detection from Smartphone Data
标题:L3-Net深度音频嵌入改进智能手机数据中的新冠肺炎检测
链接:https://arxiv.org/abs/2205.07682
作者:Mattia Giovanni Campana,Andrea Rovati,Franca Delmastro,Elena Pagani机构:∗Institute for Informatics and Telematics of the National Research Council of Italy (IIT-CNR), Pisa, Italy, †Computer Science Department, University of Milano, Milan, Italy备注:accepted for IEEE SMARTCOMP 2022摘要:智能手机和可穿戴设备,以及人工智能,可以通过实施低成本、普及的解决方案,在早期阶段识别新疾病的发展,并有可能避免新疫情的出现,在大流行控制中成为一个游戏规则改变者。最近的一些研究表明,通过使用机器学习和手工制作的声学特征,有望从声音和咳嗽中检测出2019冠状病毒疾病的诊断信号。在本文中,我们决定研究最近提出的深度嵌入模型L3网络从原始呼吸音频记录中自动提取有意义的特征的能力,以提高标准机器学习分类器从智能手机数据中区分2019冠状病毒疾病阳性和阴性受试者的性能。我们在3个数据集上评估了该模型,并将所得结果与两个参考文献的结果进行了比较。结果表明,在一组独立于受试者的实验中,L3网络与手工制作的功能相结合,在AUC方面超过了其他作品28.57%的性能。这一结果为进一步研究不同深度音频嵌入以及不同疾病的自动检测奠定了基础。摘要:Smartphones and wearable devices, along with Artificial Intelligence, can represent a game-changer in the pandemic control, by implementing low-cost and pervasive solutions to recognize the development of new diseases at their early stages and by potentially avoiding the rise of new outbreaks. Some recent works show promise in detecting diagnostic signals of COVID-19 from voice and coughs by using machine learning and hand-crafted acoustic features. In this paper, we decided to investigate the capabilities of the recently proposed deep embedding model L3-Net to automatically extract meaningful features from raw respiratory audio recordings in order to improve the performances of standard machine learning classifiers in discriminating between COVID-19 positive and negative subjects from smartphone data. We evaluated the proposed model on 3 datasets, comparing the obtained results with those of two reference works. Results show that the combination of L3-Net with hand-crafted features overcomes the performance of the other works of 28.57% in terms of AUC in a set of subject-independent experiments. This result paves the way to further investigation on different deep audio embeddings, also for the automatic detection of different diseases.
【9】 A Fast Attention Network for Joint Intent Detection and Slot Filling on Edge Devices
标题:边缘设备联合意图检测和空位填充的快速注意力网络
链接:https://arxiv.org/abs/2205.07646
作者:Liang Huang,Senjie Liang,Feiyang Ye,Nan Gao摘要:意图检测和时隙填充是自然语言理解中的两个主要任务,在面向任务的对话系统中起着至关重要的作用。这两个任务的联合学习可以提高推理的准确性,在最近的工作中很流行。然而,大多数联合模型忽略了推理延迟,无法满足在边缘部署对话系统的需要。在本文中,我们提出了一种快速注意网络(FAN),用于联合意图检测和时隙填充任务,同时保证准确性和延迟。具体来说,我们引入了一个干净且参数优化的注意模块,以增强意图和时隙之间的信息交换,将语义准确性提高了2%以上。风扇可以在不同的编码器上实现,并在每个速度级别提供更精确的模型。我们在Jetson Nano平台上的实验表明,FAN每秒能推断出15次话语,但准确度下降很小,这表明了它在边缘设备上的有效性和效率。摘要:Intent detection and slot filling are two main tasks in natural language understanding and play an essential role in task-oriented dialogue systems. The joint learning of both tasks can improve inference accuracy and is popular in recent works. However, most joint models ignore the inference latency and cannot meet the need to deploy dialogue systems at the edge. In this paper, we propose a Fast Attention Network (FAN) for joint intent detection and slot filling tasks, guaranteeing both accuracy and latency. Specifically, we introduce a clean and parameter-refined attention module to enhance the information exchange between intent and slot, improving semantic accuracy by more than 2%. FAN can be implemented on different encoders and delivers more accurate models at every speed level. Our experiments on the Jetson Nano platform show that FAN inferences fifteen utterances per second with a small accuracy drop, showing its effectiveness and efficiency on edge devices.
【10】 PRISM: Pre-trained Indeterminate Speaker Representation Model for Speaker Diarization and Speaker Verification
标题:PRISM:用于说话人二值化和说话人确认的预训练不确定说话人表示模型
链接:https://arxiv.org/abs/2205.07450
作者:Siqi Zheng,Hongbin Suo,Qian Chen机构:Speech Lab, Alibaba Group摘要:说话人嵌入是与说话人相关的任务(如验证、聚类和日记化)的一个基本特征。传统上,说话人嵌入被表示为高维空间中的固定向量。这可能会导致有偏见的估计,尤其是在处理较短的话语时。在本文中,我们建议将说话人的话语表示为“浮动”向量,其状态在不知道上下文的情况下是不确定的。说话人陈述的状态由说话人自身、同一说话人的其他讲话以及与之进行比较的其他说话人共同决定。演讲的内容也有助于确定说话人陈述的最终状态。我们预先训练了一个不确定的说话人表示模型,该模型根据上下文估计话语的状态。预训练的模型可以针对下游任务进行微调,例如说话人验证、说话人聚类和说话人二值化。在所有下游任务中都观察到了实质性的改进。摘要:Speaker embedding has been a fundamental feature for speaker-related tasks such as verification, clustering, and diarization. Traditionally, speaker embeddings are represented as fixed vectors in high-dimensional space. This could lead to biased estimations, especially when handling shorter utterances. In this paper we propose to represent a speaker utterance as "floating" vector whose state is indeterminate without knowing the context. The state of a speaker representation is jointly determined by itself, other speech from the same speaker, as well as other speakers it is being compared to. The content of the speech also contributes to determining the final state of a speaker representation. We pre-train an indeterminate speaker representation model that estimates the state of an utterance based on the context. The pre-trained model can be fine-tuned for downstream tasks such as speaker verification, speaker clustering, and speaker diarization. Substantial improvements are observed across all downstream tasks.
【11】 cMelGAN: An Efficient Conditional Generative Model Based on Mel Spectrograms
标题:CMelGAN:一种基于Mel谱图的高效条件生成模型
链接:https://arxiv.org/abs/2205.07319
作者:Tracy Qian,Jackson Kaunismaa,Tony Chung机构:___________________________________________________________________________________________________________ 1Department of Engineering Science Machine Intelligence, University of Toronto 1摘要:在机器学习领域分析音乐是一个非常困难的问题,需要考虑许多约束条件。音频数据的性质,具有非常高的维度和广泛变化的结构尺度,是建模如此困难的主要原因之一。机器学习在音乐中有很多应用,比如对一段音乐的情绪进行分类、有条件的音乐生成或流行预测。该项目的目标是开发一个基于Mel谱图的音乐类型条件生成模型,并通过将其与使用基于音符表示的现有生成音乐模型进行比较来评估其性能。我们最初实现了一个基于RNN的自回归生成模型,称为MelNet。然而,由于其速度慢、输出保真度低,我们决定创建一种基于MelGAN[4]和条件GAN结构的新的完全卷积结构,称为cMelGAN。摘要:Analysing music in the field of machine learning is a very difficult problem with numerous constraints to consider. The nature of audio data, with its very high dimensionality and widely varying scales of structure, is one of the primary reasons why it is so difficult to model. There are many applications of machine learning in music, like the classifying the mood of a piece of music, conditional music generation, or popularity prediction. The goal for this project was to develop a genre-conditional generative model of music based on Mel spectrograms and evaluate its performance by comparing it to existing generative music models that use note-based representations. We initially implemented an autoregressive, RNN-based generative model called MelNet . However, due to its slow speed and low fidelity output, we decided to create a new, fully convolutional architecture that is based on the MelGAN [4] and conditional GAN architectures, called cMelGAN.
【12】 Conditional Vector Graphics Generation for Music Cover Images
标题:音乐封面图像的条件向量图形生成
链接:https://arxiv.org/abs/2205.07301
作者:Valeria Efimova,Ivan Jarsky,Ilya Bizyaev,Andrey Filchenkov摘要:生成性对抗网络(GAN)推动了计算机图像合成领域的快速发展。由于几乎所有现有的图像合成算法都将图像视为像素矩阵,因此高分辨率图像合成非常复杂。一个很好的替代方法是矢量图像。然而,它们属于高度复杂的参数空间,这是GANs解决矢量图形合成任务的一个限制。在本文中,我们考虑了一个特定的应用领域,它极大地软化了这一限制,允许使用矢量图像合成。音乐封面图像应符合互联网流媒体服务和打印标准的要求,这意味着图形材料的高分辨率,而不需要对此类图像的内容提出任何额外要求。现有的音乐封面图像生成服务本身不分析曲目;然而,一些服务大多只考虑类型标签。为了将音乐封面生成为反映音乐并由简单几何对象组成的矢量图像,我们提出了一种基于GAN的算法CoverGAN。结果图像的评估基于它们与音乐的对应关系,并与根据标题或歌词生成的AttnGAN和DALL-E文本图像进行比较。此外,根据生成的封面图像与音乐曲目的对应关系,对CoverGAN发现的图案的意义进行了评估。听众对所提出的算法生成的音乐封面的评价非常满意,并且与曲目相对应。音乐封面图片生成代码和演示可在https://github.com/IzhanVarsky/CoverGAN.摘要:Generative Adversarial Networks (GAN) have motivated a rapid growth of the domain of computer image synthesis. As almost all the existing image synthesis algorithms consider an image as a pixel matrix, the high-resolution image synthesis is complicated.A good alternative can be vector images. However, they belong to the highly sophisticated parametric space, which is a restriction for solving the task of synthesizing vector graphics by GANs. In this paper, we consider a specific application domain that softens this restriction dramatically allowing the usage of vector image synthesis. Music cover images should meet the requirements of Internet streaming services and printing standards, which imply high resolution of graphic materials without any additional requirements on the content of such images. Existing music cover image generation services do not analyze tracks themselves; however, some services mostly consider only genre tags. To generate music covers as vector images that reflect the music and consist of simple geometric objects, we suggest a GAN-based algorithm called CoverGAN. The assessment of resulting images is based on their correspondence to the music compared with AttnGAN and DALL-E text-to-image generation according to title or lyrics. Moreover, the significance of the patterns found by CoverGAN has been evaluated in terms of the correspondence of the generated cover images to the musical tracks. Listeners evaluate the music covers generated by the proposed algorithm as quite satisfactory and corresponding to the tracks. Music cover images generation code and demo are available at https://github.com/IzhanVarsky/CoverGAN.
【13】 The VoicePrivacy 2020 Challenge Evaluation Plan
标题:VoicePrivacy 2020挑战评估计划
链接:https://arxiv.org/abs/2205.07123
作者:Natalia Tomashenko,Brij Mohan Lal Srivastava,Xin Wang,Emmanuel Vincent,Andreas Nautsch,Junichi Yamagishi,Nicholas Evans,Jose Patino,Jean-François Bonastre,Paul-Gauthier Noé,Massimiliano Todisco机构:Laboratoire Informatique d’Avignon (LIA), Avignon Université, France, Inria, France, National Institute of Informatics, Tokyo, Japan, Université de Lorraine, CNRS, Inria, LORIA, France, Audio Security and Privacy Group, EURECOM, France, University of Edinburgh, UK备注:arXiv admin note: text overlap with arXiv:2203.12468摘要:VoicePrivacy Challenge旨在通过聚集一个新的社区来定义感兴趣的任务和评估方法,并通过一系列挑战对解决方案进行基准测试,从而促进语音技术隐私保护工具的发展。在本文中,我们制定了为VoicePrivacy 2020挑战选择的语音匿名任务,并描述了用于系统开发和评估的数据集。我们还介绍了攻击模型以及相关的客观和主观评估指标。我们引入了两个匿名基线,并报告了客观的评估结果。摘要:The VoicePrivacy Challenge aims to promote the development of privacy preservation tools for speech technology by gathering a new community to define the tasks of interest and the evaluation methodology, and benchmarking solutions through a series of challenges. In this document, we formulate the voice anonymization task selected for the VoicePrivacy 2020 Challenge and describe the datasets used for system development and evaluation. We also present the attack models and the associated objective and subjective evaluation metrics. We introduce two anonymization baselines and report objective evaluation results.
【14】 Multiformer: A Head-Configurable Transformer-Based Model for Direct Speech Translation
标题:多转换器:一种基于头部可配置转换器的直接语音翻译模型
链接:https://arxiv.org/abs/2205.07100
作者:Gerard Sant,Gerard I. Gállego,Belen Alastruey,Marta R. Costa-Jussà机构:TALP Research Center, Universitat Politècnica de Catalunya, Barcelona摘要:基于Transformer的模型已经在自然语言处理的几个领域取得了最新的成果。然而,它在语音任务中的直接应用并非微不足道。这种序列的性质带来了诸如长序列长度和相邻令牌之间的冗余等问题。因此,我们认为,常规的自我注意机制可能并不适合这种情况。人们提出了不同的方法来克服这些问题,例如使用有效的注意机制。然而,使用这些方法通常会带来成本,这是由于信息丢失导致的性能降低。在这项研究中,我们提出了Multiformer,这是一个基于Transformer的模型,允许在每个头部使用不同的注意机制。通过这样做,该模型能够将自我关注偏向于提取更多样化的令牌交互,并且减少了信息损失。最后,我们对head贡献进行了分析,并观察到所有head相关性均匀分布的架构可以获得更好的结果。我们的研究结果表明,不同头部和层次的混合注意力模式比我们的基线高出0.7 BLEU。摘要:Transformer-based models have been achieving state-of-the-art results in several fields of Natural Language Processing. However, its direct application to speech tasks is not trivial. The nature of this sequences carries problems such as long sequence lengths and redundancy between adjacent tokens. Therefore, we believe that regular self-attention mechanism might not be well suited for it. Different approaches have been proposed to overcome these problems, such as the use of efficient attention mechanisms. However, the use of these methods usually comes with a cost, which is a performance reduction caused by information loss. In this study, we present the Multiformer, a Transformer-based model which allows the use of different attention mechanisms on each head. By doing this, the model is able to bias the self-attention towards the extraction of more diverse token interactions, and the information loss is reduced. Finally, we perform an analysis of the head contributions, and we observe that those architectures where all heads relevance is uniformly distributed obtain better results. Our results show that mixing attention patterns along the different heads and layers outperforms our baseline by up to 0.7 BLEU.
【15】 Improved Consistency Training for Semi-Supervised Sequence-to-Sequence ASR via Speech Chain Reconstruction and Self-Transcribing
标题:基于语音链重构和自转录的半监督序列间ASR一致性训练
链接:https://arxiv.org/abs/2205.06963
作者:Heli Qi,Sashi Novitasari,Sakriani Sakti,Satoshi Nakamura机构:Nara Institute of Science and Technology, Japan, Japan Advanced Institute of Science and Technology, Japan备注:Submitted to INTERSPEECH 2022摘要:一致性正则化最近被应用于半监督序列对序列(S2S)自动语音识别(ASR)。这一原理鼓励ASR模型对具有不同扰动的相同输入语音输出类似的预测。现有的半监督S2S ASR范式利用SPECAMULTE作为数据增强,需要静态教师模型为未翻译的语音生成伪转录本。然而,这种范式未能充分利用一致性正则化。首先,SpecAugment的掩蔽操作可能会破坏语音的语言内容,从而影响伪标签的质量。其次,S2S ASR需要输入语音和前缀标记来进行下一次预测。在一致性训练期间,离线教师模型生成的静态前缀标记无法与动态伪标签匹配。在这项工作中,我们提出了一种改进的半监督S2S ASR一致性训练范式。我们利用语音链重建作为弱增强来生成高质量的伪标签。此外,我们还证明了学生ASR模型产生的动态伪转录本有利于一致性训练。在LJSpeech和LibriSpeech语料库上的实验表明,与监督基线相比,我们改进的范例在单说话人环境下的CER提高了12.2%,在多说话人环境下的CER提高了38.6%。摘要:Consistency regularization has recently been applied to semi-supervised sequence-to-sequence (S2S) automatic speech recognition (ASR). This principle encourages an ASR model to output similar predictions for the same input speech with different perturbations. The existing paradigm of semi-supervised S2S ASR utilizes SpecAugment as data augmentation and requires a static teacher model to produce pseudo transcripts for untranscribed speech. However, this paradigm fails to take full advantage of consistency regularization. First, the masking operations of SpecAugment may damage the linguistic contents of the speech, thus influencing the quality of pseudo labels. Second, S2S ASR requires both input speech and prefix tokens to make the next prediction. The static prefix tokens made by the offline teacher model cannot match dynamic pseudo labels during consistency training. In this work, we propose an improved consistency training paradigm of semi-supervised S2S ASR. We utilize speech chain reconstruction as the weak augmentation to generate high-quality pseudo labels. Moreover, we demonstrate that dynamic pseudo transcripts produced by the student ASR model benefit the consistency training. Experiments on LJSpeech and LibriSpeech corpora show that compared to supervised baselines, our improved paradigm achieves a 12.2% CER improvement in the single-speaker setting and 38.6% in the multi-speaker setting