【1】Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations标题:通过自监督语音表征的对应训练改进声学单词嵌入链接:https://arxiv.org/abs/2403.08738作者:Amit Meghanani,Thomas Hain备注:Accepted to EACL 2024 Main Conference, Long paper摘要:声学词嵌入(AWE)是口语词的矢量表示。获得AWE的有效方法是对应自动编码器(CAE)。在过去,CAE方法一直与传统的MFCC功能。从基于自监督学习(SSL)的语音模型(如HuBERT、Wav 2 vec 2等)获得的表示,在许多下游任务中表现优于MFCC。然而,他们还没有得到很好的研究,在学习AWE的背景下。这项工作探讨了CAE与SSL为基础的语音表示,以获得改进的AWE的有效性。此外,基于SSL的语音模型的能力,探讨在跨语言的情况下获得AWE。实验在五种语言上进行:波兰语、葡萄牙语、西班牙语、法语和英语。基于HuBERT的CAE模型在所有语言中的单词识别方面都取得了最好的结果,尽管Hu-BERT只在英语上进行了预训练。此外,基于休伯特的CAE模型在跨语言设置中工作良好。当在一种源语言上训练并在目标语言上测试时,它优于在目标语言上训练的基于MFCC的CAE模型。摘要:Acoustic word embeddings (AWEs) are vector representations of spoken words. An effective method for obtaining AWEs is the Correspondence Auto-Encoder (CAE). In the past, the CAE method has been associated with traditional MFCC features. Representations obtained from self-supervised learning (SSL)-based speech models such as HuBERT, Wav2vec2, etc., are outperforming MFCC in many downstream tasks. However, they have not been well studied in the context of learning AWEs. This work explores the effectiveness of CAE with SSL-based speech representations to obtain improved AWEs. Additionally, the capabilities of SSL-based speech models are explored in cross-lingual scenarios for obtaining AWEs. Experiments are conducted on five languages: Polish, Portuguese, Spanish, French, and English. HuBERT-based CAE model achieves the best results for word discrimination in all languages, despite Hu-BERT being pre-trained on English only. Also, the HuBERT-based CAE model works well in cross-lingual settings. It outperforms MFCC-based CAE models trained on the target languages when trained on one source language and tested on target languages. 【2】 End-to-End Amp Modeling: From Data to Controllable Guitar Amplifier Models标题:端到端放大器建模:从数据到可控吉他放大器模型链接:https://arxiv.org/abs/2403.08559作者:Lauri Juvela,Eero-Pekka Damskägg,Aleksi Peussa,Jaakko Mäkinen,Thomas Sherson,Stylianos I. Mimilakis,Athanasios Gotsopoulos备注:Presented at ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)摘要:本文描述了一种数据驱动的方法来创建吉他放大器的实时神经网络模型,在物理设备上的全范围控制下重新创建放大器对任意输入的声音响应。虽然本文的重点是数据收集管道,但我们通过训练LSTM模型来证明这种条件黑盒方法的有效性,并将其性能与离线白盒SPICE电路仿真进行比较。我们的听力测试结果表明,神经放大器建模方法可以匹配高质量SPICE模型的主观性能,同时使用自动化,非侵入式数据收集过程和端到端可训练,实时可行的神经网络模型。摘要:This paper describes a data-driven approach to creating real-time neural network models of guitar amplifiers, recreating the amplifiers' sonic response to arbitrary inputs at the full range of controls present on the physical device. While the focus on the paper is on the data collection pipeline, we demonstrate the effectiveness of this conditioned black-box approach by training an LSTM model to the task, and comparing its performance to an offline white-box SPICE circuit simulation. Our listening test results demonstrate that the neural amplifier modeling approach can match the subjective performance of a high-quality SPICE model, all while using an automated, non-intrusive data collection process, and an end-to-end trainable, real-time feasible neural network model.
【3】 From Weak to Strong Sound Event Labels using Adaptive Change-Point Detection and Active Learning标题:基于自适应变点检测和主动学习的从弱声到强声事件标记链接:https://arxiv.org/abs/2403.08525作者:John Martinsson,Olof Mogren,Maria Sandsten,Tuomas Virtanen备注:Under review at EUSIPCO 2024摘要:在这项工作中,我们提出了一种基于自适应变点检测(A-CPD)的音频记录分割方法,用于机器引导的音频记录片段的弱标签标注。我们的目标是最大限度地获得有关目标声音的时间激活的信息量。对于每个未标记的音频记录,我们使用预测模型来导出用于指导注释的概率曲线。该预测模型最初是在可用的带注释的声音事件数据上进行预训练的,该声音事件数据具有与未标记数据集中的类不相交的类。然后,预测模型在主动学习循环中逐渐适应注释器提供的注释。用于引导弱标签注释器朝向强标签的查询是使用这些概率上的变化点检测来导出的。我们表明,即使在有限的注释预算下,也可以获得高质量的强标签,并且与两种基线查询策略相比,A-CPD显示出有利的结果。摘要:In this work we propose an audio recording segmentation method based on an adaptive change point detection (A-CPD) for machine guided weak label annotation of audio recording segments. The goal is to maximize the amount of information gained about the temporal activation's of the target sounds. For each unlabeled audio recording, we use a prediction model to derive a probability curve used to guide annotation. The prediction model is initially pre-trained on available annotated sound event data with classes that are disjoint from the classes in the unlabeled dataset. The prediction model then gradually adapts to the annotations provided by the annotator in an active learning loop. The queries used to guide the weak label annotator towards strong labels are derived using change point detection on these probabilities. We show that it is possible to derive strong labels of high quality even with a limited annotation budget, and show favorable results for A-CPD when compared to two baseline query strategies.
【4】 Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children标题:自动语音识别(ASR)对韩国儿童言语发音障碍的诊断链接:https://arxiv.org/abs/2403.08187作者:Taekyung Ahn,Yeonjung Hong,Younggon Im,Do Hyung Kim,Dayoung Kang,Joo Won Jeong,Jae Won Kim,Min Jung Kim,Ah-ra Cho,Dae-Hyun Jang,Hosung Nam备注:12 pages, 2 figures摘要:本研究提出了一种自动语音识别(ASR)模型,旨在诊断语音障碍(SSD)儿童的发音问题,以取代临床程序中的手动transmittance。由于为通用目的训练的ASR模型主要将输入语音预测为真实单词,因此采用众所周知的高性能ASR模型来评估患有SSD的儿童的发音是不切实际的。我们对wav 2 vec 2.0 XLS-R模型进行了微调,以将语音识别为发音而不是现有单词。该模型与来自137名语音产生不足的儿童的语音数据集进行了微调,这些儿童发音73个韩语单词,用于实际临床诊断。该模型对单词发音的预测与人类注释的准确率约为90%。虽然该模型在识别不清楚的发音方面仍需改进,但本研究表明,ASR模型可以简化临床领域复杂的发音错误诊断程序。摘要:This study presents a model of automatic speech recognition (ASR) designed to diagnose pronunciation issues in children with speech sound disorders (SSDs) to replace manual transcriptions in clinical procedures. Since ASR models trained for general purposes primarily predict input speech into real words, employing a well-known high-performance ASR model for evaluating pronunciation in children with SSDs is impractical. We fine-tuned the wav2vec 2.0 XLS-R model to recognize speech as pronounced rather than as existing words. The model was fine-tuned with a speech dataset from 137 children with inadequate speech production pronouncing 73 Korean words selected for actual clinical diagnosis. The model's predictions of the pronunciations of the words matched the human annotations with about 90% accuracy. While the model still requires improvement in recognizing unclear pronunciation, this study demonstrates that ASR models can streamline complex pronunciation error diagnostic procedures in clinical fields.
【5】 EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech标题:EM-TTS:高效训练的低资源蒙古语轻量级文语转换系统链接:https://arxiv.org/abs/2403.08164作者:Ziqi Liang,Haoxiang Shi,Jiawei Wang,Keda Lu备注:Accepted by the 27th IEEE International Conference on Computer Supported Cooperative Work in Design (IEEE CSCWD 2024). arXiv admin note: substantial text overlap with arXiv:2211.01948摘要:最近,基于深度学习的文本到语音(TTS)系统已经实现了高质量的语音合成结果。递归神经网络已成为TTS系统中序列数据的标准建模技术,并得到广泛应用。然而,训练包含RNN组件的TTS模型需要强大的GPU性能,并且需要很长时间。相比之下,基于CNN的序列合成技术可以显着减少TTS模型的参数和训练时间,同时由于其高度并行性而保证一定的性能,从而减轻这些训练的经济成本。在本文中,我们提出了一个基于深度卷积神经网络的轻量级TTS系统,它是一个两阶段训练的端到端TTS模型,不使用任何递归单元。我们的模型包括两个阶段:Text 2Spectrum和SSRN。前者用于将音素编码成粗梅尔频谱图,后者用于从粗梅尔频谱图合成完整频谱。同时,我们通过一系列的数据增强,如噪声抑制,时间弯曲,频率掩蔽和时间掩蔽,以提高我们的模型的鲁棒性,解决低资源蒙古问题。实验表明,与主流TTS模型相比,该模型在保证合成语音质量和自然度的同时,减少了训练时间和参数。我们的方法使用NCMMSC 2022-MTTSC Challenge数据集进行验证,在保持一定准确性的同时显著减少了训练时间。摘要:Recently, deep learning-based Text-to-Speech (TTS) systems have achieved high-quality speech synthesis results. Recurrent neural networks have become a standard modeling technique for sequential data in TTS systems and are widely used. However, training a TTS model which includes RNN components requires powerful GPU performance and takes a long time. In contrast, CNN-based sequence synthesis techniques can significantly reduce the parameters and training time of a TTS model while guaranteeing a certain performance due to their high parallelism, which alleviate these economic costs of training. In this paper, we propose a lightweight TTS system based on deep convolutional neural networks, which is a two-stage training end-to-end TTS model and does not employ any recurrent units. Our model consists of two stages: Text2Spectrum and SSRN. The former is used to encode phonemes into a coarse mel spectrogram and the latter is used to synthesize the complete spectrum from the coarse mel spectrogram. Meanwhile, we improve the robustness of our model by a series of data augmentations, such as noise suppression, time warping, frequency masking and time masking, for solving the low resource mongolian problem. Experiments show that our model can reduce the training time and parameters while ensuring the quality and naturalness of the synthesized speech compared to using mainstream TTS models. Our method uses NCMMSC2022-MTTSC Challenge dataset for validation, which significantly reduces training time while maintaining a certain accuracy.
【6】 Motifs, Phrases, and Beyond: The Modelling of Structure in Symbolic Music Generation标题:母题、词组与超越:符号音乐生成中的结构造型链接:https://arxiv.org/abs/2403.07995作者:Keshav Bhandari,Simon Colton备注:Accepted to 13th International Conference on Artificial Intelligence in Music, Sound, Art and Design (EvoMUSART) 2024摘要:音乐结构建模对于生成符号音乐作品的人工智能系统来说至关重要,但也具有挑战性。这篇文献综述剖析了整合连贯结构的技术的演变,从符号方法到基础性和变革性的深度学习方法,这些方法利用了各种各样的训练范式中的计算和数据的力量。在后面的阶段,我们回顾了一个新兴的技术,我们称之为“子任务分解”,涉及分解音乐生成到单独的高层次的结构规划和内容创作阶段。这种系统通过提取旋律骨架或结构模板来指导生成,从而结合了某种形式的音乐知识或神经符号方法。在捕捉所回顾的所有三个时代的主题和重复方面取得了明显的进展,但要以人类作曲家的风格在扩展作品中对主题的细微发展进行建模仍然很困难。我们概述了几个关键的未来方向,以实现协同效益相结合的方法,从所有时代的检查。摘要:Modelling musical structure is vital yet challenging for artificial intelligence systems that generate symbolic music compositions. This literature review dissects the evolution of techniques for incorporating coherent structure, from symbolic approaches to foundational and transformative deep learning methods that harness the power of computation and data across a wide variety of training paradigms. In the later stages, we review an emerging technique which we refer to as "sub-task decomposition" that involves decomposing music generation into separate high-level structural planning and content creation stages. Such systems incorporate some form of musical knowledge or neuro-symbolic methods by extracting melodic skeletons or structural templates to guide the generation. Progress is evident in capturing motifs and repetitions across all three eras reviewed, yet modelling the nuanced development of themes across extended compositions in the style of human composers remains difficult. We outline several key future directions to realize the synergistic benefits of combining approaches from all eras examined.
【7】 Text-to-Audio Generation Synchronized with Videos标题:与视频同步的文本到音频生成链接:https://arxiv.org/abs/2403.07938作者:Shentong Mo,Jing Shi,Yapeng Tian备注:arXiv admin note: text overlap with arXiv:2305.12903摘要:近年来,随着研究人员努力从文本描述合成音频,对文本到音频(TTA)生成的关注已经加强。然而,大多数现有的方法,虽然利用潜在的扩散模型来学习音频和文本嵌入之间的相关性,但在保持所产生的音频和视频之间的无缝同步方面都存在不足。这往往导致明显的视听不匹配。为了弥合这一差距,我们引入了一个突破性的文本到音频生成基准,与视频保持一致,名为T2 AV-Bench。该基准与三个致力于评估视觉对齐和时间一致性的新指标不同。为了补充这一点,我们还提出了一个简单而有效的视频对齐TTA生成模型,即T2 AV。超越传统方法,T2 AV通过整合视觉对齐的文本嵌入作为其条件基础来改进潜在扩散方法。它采用了一个时间多头注意力Transformer来提取和理解视频数据中的时间细微差别,这一壮举被我们的视听控制网放大,它巧妙地将时间视觉表示与文本嵌入相结合。为了进一步加强这种整合,我们编织了一个对比学习目标,旨在确保视觉对齐的文本嵌入与音频功能密切相关。对AudioCaps和T2 AV-Bench的广泛评估表明,我们的T2 AV为视频对齐TTA生成设定了新标准,以确保视觉对齐和时间一致性。摘要:In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models to learn the correlation between audio and text embeddings, fall short when it comes to maintaining a seamless synchronization between the produced audio and its video. This often results in discernible audio-visual mismatches. To bridge this gap, we introduce a groundbreaking benchmark for Text-to-Audio generation that aligns with Videos, named T2AV-Bench. This benchmark distinguishes itself with three novel metrics dedicated to evaluating visual alignment and temporal consistency. To complement this, we also present a simple yet effective video-aligned TTA generation model, namely T2AV. Moving beyond traditional methods, T2AV refines the latent diffusion approach by integrating visual-aligned text embeddings as its conditional foundation. It employs a temporal multi-head attention transformer to extract and understand temporal nuances from video data, a feat amplified by our Audio-Visual ControlNet that adeptly merges temporal visual representations with text embeddings. Further enhancing this integration, we weave in a contrastive learning objective, designed to ensure that the visual-aligned text embeddings resonate closely with the audio features. Extensive evaluations on the AudioCaps and T2AV-Bench demonstrate that our T2AV sets a new standard for video-aligned TTA generation in ensuring visual alignment and temporal consistency. 【8】 An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning标题:一种基于多任务学习的端到端噪声不变语音特征提取方法链接:https://arxiv.org/abs/2403.08654作者:Heitor R. Guimarães,Arthur Pimentel,Anderson R. Avila,Mehdi Rezagholizadeh,Boxing Chen,Tiago H. Falk备注:Under review on IEEE Transactions on Audio, Speech, and Language Processing (2024)摘要:自监督语音表示学习能够从原始波形中提取有意义的特征。然后,这些功能可以在多个下游任务中有效地使用。然而,当考虑这种方法的“野外”部署时,会出现两个重要问题:(i)它们的大尺寸,这对于边缘应用来说可能是禁止的;以及(ii)它们对有害因素(例如噪声和/或混响)的鲁棒性,这可能严重降低这种系统的性能。在这项工作中,我们提出了RobustDistiller,一种新的知识蒸馏机制,共同解决这两个问题。在蒸馏配方的同时,我们应用多任务学习目标来鼓励网络通过对输入进行降噪来学习噪声不变的表示。所提出的机制进行了评估12个不同的下游任务。它优于几个基准,无论噪声类型,或噪声和混响水平。实验结果表明,新的23 M参数的学生模型可以达到与95 M参数的教师模型相当的结果。最后,我们表明,所提出的配方可以应用于其他蒸馏方法,如最近的DPWavLM。为了重现性,代码和模型检查点将在\mbox{\url{https://github.com/Hguimaraes/robustdistiller}}上提供。摘要:Self-supervised speech representation learning enables the extraction of meaningful features from raw waveforms. These features can then be efficiently used across multiple downstream tasks. However, two significant issues arise when considering the deployment of such methods ``in-the-wild": (i) Their large size, which can be prohibitive for edge applications; and (ii) their robustness to detrimental factors, such as noise and/or reverberation, that can heavily degrade the performance of such systems. In this work, we propose RobustDistiller, a novel knowledge distillation mechanism that tackles both problems jointly. Simultaneously to the distillation recipe, we apply a multi-task learning objective to encourage the network to learn noise-invariant representations by denoising the input. The proposed mechanism is evaluated on twelve different downstream tasks. It outperforms several benchmarks regardless of noise type, or noise and reverberation levels. Experimental results show that the new Student model with 23M parameters can achieve results comparable to the Teacher model with 95M parameters. Lastly, we show that the proposed recipe can be applied to other distillation methodologies, such as the recent DPWavLM. For reproducibility, code and model checkpoints will be made available at \mbox{\url{https://github.com/Hguimaraes/robustdistiller}}.
【9】 Speech Robust Bench: A Robustness Benchmark For Speech Recognition标题:语音鲁棒性基准:语音识别的鲁棒性基准链接:https://arxiv.org/abs/2403.07937作者:Muhammad A. Shah,David Solans Noguero,Mikko A. Heikkila,Nicolas Kourtellis摘要:随着自动语音识别(ASR)模型变得越来越普遍,重要的是要确保它们在物理和数字世界中存在的腐败情况下做出可靠的预测。我们提出了语音鲁棒性基准(SRB),一个全面的基准评估ASR模型的鲁棒性,以不同的腐败。SRB由69个输入扰动组成,旨在模拟ASR模型在物理和数字世界中可能遇到的各种损坏。我们使用SRB来评估几个最先进的ASR模型的鲁棒性,并观察到模型大小和某些建模选择,如离散表示和自我训练似乎有利于鲁棒性。我们扩展了这种分析,以衡量ASR模型的鲁棒性的数据,从不同的人口分组,即英语和西班牙语的发言者,男性和女性,并观察到显着的差异,在模型的鲁棒性在各小组。我们相信,SRB将促进未来的研究对强大的ASR模型,使其更容易进行全面和可比的鲁棒性评估。摘要:As Automatic Speech Recognition (ASR) models become ever more pervasive, it is important to ensure that they make reliable predictions under corruptions present in the physical and digital world. We propose Speech Robust Bench (SRB), a comprehensive benchmark for evaluating the robustness of ASR models to diverse corruptions. SRB is composed of 69 input perturbations which are intended to simulate various corruptions that ASR models may encounter in the physical and digital world. We use SRB to evaluate the robustness of several state-of-the-art ASR models and observe that model size and certain modeling choices such as discrete representations, and self-training appear to be conducive to robustness. We extend this analysis to measure the robustness of ASR models on data from various demographic subgroups, namely English and Spanish speakers, and males and females, and observed noticeable disparities in the model's robustness across subgroups. We believe that SRB will facilitate future research towards robust ASR models, by making it easier to conduct comprehensive and comparable robustness evaluations.
eess.AS音频处理【1】 An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning标题:一种基于多任务学习的端到端噪声不变语音特征提取方法链接:https://arxiv.org/abs/2403.08654作者:Heitor R. Guimarães,Arthur Pimentel,Anderson R. Avila,Mehdi Rezagholizadeh,Boxing Chen,Tiago H. Falk备注:Under review on IEEE Transactions on Audio, Speech, and Language Processing (2024)摘要:自监督语音表示学习能够从原始波形中提取有意义的特征。然后,这些功能可以在多个下游任务中有效地使用。然而,当考虑这种方法的“野外”部署时,会出现两个重要问题:(i)它们的大尺寸,这对于边缘应用来说可能是禁止的;以及(ii)它们对有害因素(例如噪声和/或混响)的鲁棒性,这可能严重降低这种系统的性能。在这项工作中,我们提出了RobustDistiller,一种新的知识蒸馏机制,共同解决这两个问题。在蒸馏配方的同时,我们应用多任务学习目标来鼓励网络通过对输入进行降噪来学习噪声不变的表示。所提出的机制进行了评估12个不同的下游任务。它优于几个基准,无论噪声类型,或噪声和混响水平。实验结果表明,新的23 M参数的学生模型可以达到与95 M参数的教师模型相当的结果。最后,我们表明,所提出的配方可以应用于其他蒸馏方法,如最近的DPWavLM。为了重现性,代码和模型检查点将在\mbox{\url{https://github.com/Hguimaraes/robustdistiller}}上提供。摘要:Self-supervised speech representation learning enables the extraction of meaningful features from raw waveforms. These features can then be efficiently used across multiple downstream tasks. However, two significant issues arise when considering the deployment of such methods ``in-the-wild": (i) Their large size, which can be prohibitive for edge applications; and (ii) their robustness to detrimental factors, such as noise and/or reverberation, that can heavily degrade the performance of such systems. In this work, we propose RobustDistiller, a novel knowledge distillation mechanism that tackles both problems jointly. Simultaneously to the distillation recipe, we apply a multi-task learning objective to encourage the network to learn noise-invariant representations by denoising the input. The proposed mechanism is evaluated on twelve different downstream tasks. It outperforms several benchmarks regardless of noise type, or noise and reverberation levels. Experimental results show that the new Student model with 23M parameters can achieve results comparable to the Teacher model with 95M parameters. Lastly, we show that the proposed recipe can be applied to other distillation methodologies, such as the recent DPWavLM. For reproducibility, code and model checkpoints will be made available at \mbox{\url{https://github.com/Hguimaraes/robustdistiller}}.
【2】 The evaluation of a code-switched Sepedi-English automatic speech recognition system标题:一种码转换的英语语音自动识别系统的评价链接:https://arxiv.org/abs/2403.07947作者:Amanda Phaladi,Thipe Modipa备注:13 pages,2 figures,2nd International Conference on NLP & AI (NLPAI 2024)摘要:语音技术是包含用于使机器能够与语音交互的各种技术和工具的领域,诸如自动语音识别(ASR)、口语对话系统等,其允许设备通过麦克风从人类说话者捕获口语单词。端到端的方法,如连接时间分类(CTC)和基于注意力的方法是最常用的ASR系统的开发。然而,这些技术通常用于许多高资源语言的研究和开发,需要大量的语音数据进行训练和评估,而低资源语言相对不发达。虽然CTC方法已成功用于其他语言,但其对Sepedi语言的有效性仍不确定。在这项研究中,我们提出了评估的Sepedi-English代码切换自动语音识别系统。这个端到端系统是使用Sepedi的代码切换语料库和CTC方法开发的。该系统的性能进行了评估,使用NCHLT Sepedi测试语料库和Sepedi标记的代码切换语料库。该模型产生的WER最低,为41.9%,然而,该模型在识别Sepedi唯一文本方面面临挑战。摘要:Speech technology is a field that encompasses various techniques and tools used to enable machines to interact with speech, such as automatic speech recognition (ASR), spoken dialog systems, and others, allowing a device to capture spoken words through a microphone from a human speaker. End-to-end approaches such as Connectionist Temporal Classification (CTC) and attention-based methods are the most used for the development of ASR systems. However, these techniques were commonly used for research and development for many high-resourced languages with large amounts of speech data for training and evaluation, leaving low-resource languages relatively underdeveloped. While the CTC method has been successfully used for other languages, its effectiveness for the Sepedi language remains uncertain. In this study, we present the evaluation of the Sepedi-English code-switched automatic speech recognition system. This end-to-end system was developed using the Sepedi Prompted Code Switching corpus and the CTC approach. The performance of the system was evaluated using both the NCHLT Sepedi test corpus and the Sepedi Prompted Code Switching corpus. The model produced the lowest WER of 41.9%, however, the model faced challenges in recognizing the Sepedi only text.
【3】 Speech Robust Bench: A Robustness Benchmark For Speech Recognition标题:语音稳健性基准:一种语音识别的稳健性基准链接:https://arxiv.org/abs/2403.07937作者:Muhammad A. Shah,David Solans Noguero,Mikko A. Heikkila,Nicolas Kourtellis摘要:随着自动语音识别(ASR)模型变得越来越普遍,重要的是要确保它们在物理和数字世界中存在的腐败情况下做出可靠的预测。我们提出了语音鲁棒性基准(SRB),一个全面的基准评估ASR模型的鲁棒性,以不同的腐败。SRB由69个输入扰动组成,旨在模拟ASR模型在物理和数字世界中可能遇到的各种损坏。我们使用SRB来评估几个最先进的ASR模型的鲁棒性,并观察到模型大小和某些建模选择,如离散表示和自我训练似乎有利于鲁棒性。我们扩展了这种分析,以衡量ASR模型的鲁棒性的数据,从不同的人口分组,即英语和西班牙语的发言者,男性和女性,并观察到显着的差异,在模型的鲁棒性在各小组。我们相信,SRB将促进未来的研究对强大的ASR模型,使其更容易进行全面和可比的鲁棒性评估。摘要:As Automatic Speech Recognition (ASR) models become ever more pervasive, it is important to ensure that they make reliable predictions under corruptions present in the physical and digital world. We propose Speech Robust Bench (SRB), a comprehensive benchmark for evaluating the robustness of ASR models to diverse corruptions. SRB is composed of 69 input perturbations which are intended to simulate various corruptions that ASR models may encounter in the physical and digital world. We use SRB to evaluate the robustness of several state-of-the-art ASR models and observe that model size and certain modeling choices such as discrete representations, and self-training appear to be conducive to robustness. We extend this analysis to measure the robustness of ASR models on data from various demographic subgroups, namely English and Spanish speakers, and males and females, and observed noticeable disparities in the model's robustness across subgroups. We believe that SRB will facilitate future research towards robust ASR models, by making it easier to conduct comprehensive and comparable robustness evaluations.
【4】 Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations标题:通过自监督语音表征的对应训练改进声学单词嵌入链接:https://arxiv.org/abs/2403.08738作者:Amit Meghanani,Thomas Hain备注:Accepted to EACL 2024 Main Conference, Long paper摘要:声学词嵌入(AWE)是口语词的矢量表示。获得AWE的有效方法是对应自动编码器(CAE)。在过去,CAE方法一直与传统的MFCC功能。从基于自监督学习(SSL)的语音模型(如HuBERT、Wav 2 vec 2等)获得的表示,在许多下游任务中表现优于MFCC。然而,他们还没有得到很好的研究,在学习AWE的背景下。这项工作探讨了CAE与SSL为基础的语音表示,以获得改进的AWE的有效性。此外,基于SSL的语音模型的能力,探讨在跨语言的情况下获得AWE。实验在五种语言上进行:波兰语、葡萄牙语、西班牙语、法语和英语。基于HuBERT的CAE模型在所有语言中的单词识别方面都取得了最好的结果,尽管Hu-BERT只在英语上进行了预训练。此外,基于休伯特的CAE模型在跨语言设置中工作良好。当在一种源语言上训练并在目标语言上测试时,它优于在目标语言上训练的基于MFCC的CAE模型。摘要:Acoustic word embeddings (AWEs) are vector representations of spoken words. An effective method for obtaining AWEs is the Correspondence Auto-Encoder (CAE). In the past, the CAE method has been associated with traditional MFCC features. Representations obtained from self-supervised learning (SSL)-based speech models such as HuBERT, Wav2vec2, etc., are outperforming MFCC in many downstream tasks. However, they have not been well studied in the context of learning AWEs. This work explores the effectiveness of CAE with SSL-based speech representations to obtain improved AWEs. Additionally, the capabilities of SSL-based speech models are explored in cross-lingual scenarios for obtaining AWEs. Experiments are conducted on five languages: Polish, Portuguese, Spanish, French, and English. HuBERT-based CAE model achieves the best results for word discrimination in all languages, despite Hu-BERT being pre-trained on English only. Also, the HuBERT-based CAE model works well in cross-lingual settings. It outperforms MFCC-based CAE models trained on the target languages when trained on one source language and tested on target languages. 【5】 End-to-End Amp Modeling: From Data to Controllable Guitar Amplifier Models标题:端到端放大器建模:从数据到可控吉他放大器模型链接:https://arxiv.org/abs/2403.08559作者:Lauri Juvela,Eero-Pekka Damskägg,Aleksi Peussa,Jaakko Mäkinen,Thomas Sherson,Stylianos I. Mimilakis,Athanasios Gotsopoulos备注:Presented at ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)摘要:本文描述了一种数据驱动的方法来创建吉他放大器的实时神经网络模型,在物理设备上的全范围控制下重新创建放大器对任意输入的声音响应。虽然本文的重点是数据收集管道,但我们通过训练LSTM模型来证明这种条件黑盒方法的有效性,并将其性能与离线白盒SPICE电路仿真进行比较。我们的听力测试结果表明,神经放大器建模方法可以匹配高质量SPICE模型的主观性能,同时使用自动化,非侵入式数据收集过程和端到端可训练,实时可行的神经网络模型。摘要:This paper describes a data-driven approach to creating real-time neural network models of guitar amplifiers, recreating the amplifiers' sonic response to arbitrary inputs at the full range of controls present on the physical device. While the focus on the paper is on the data collection pipeline, we demonstrate the effectiveness of this conditioned black-box approach by training an LSTM model to the task, and comparing its performance to an offline white-box SPICE circuit simulation. Our listening test results demonstrate that the neural amplifier modeling approach can match the subjective performance of a high-quality SPICE model, all while using an automated, non-intrusive data collection process, and an end-to-end trainable, real-time feasible neural network model. 【6】 SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation标题:SpeechColab Leadboard:一个开源的自动语音识别评测平台链接:https://arxiv.org/abs/2403.08196作者:Jiayu Du,Jinpeng Li,Guoguo Chen,Wei-Qiang Zhang摘要:在过去十年中,随着深度学习浪潮的高涨,自动语音识别(ASR)引起了人们的极大关注,导致了许多可公开访问的ASR系统的出现,这些系统正在积极融入我们的日常生活。然而,由于各种关键的微妙之处,这些ASR系统的公正和可复制的评估遇到了挑战。在本文中,我们介绍了SpeechColab排行榜,一个通用的,开源的ASR评估平台。通过这个平台:(i)我们报告了一个全面的基准,揭示了ASR系统目前最先进的全景,涵盖开源模型和工业商业服务。(ii)我们研究了评分流程中的不同细微差别如何影响最终的基准结果。这些问题包括与大写、标点、感叹词、缩写、同义词用法、复合词等有关的细微差别。这些问题在向端到端未来过渡的背景下变得突出。(iii)我们提出了一个实际的修改传统的令牌错误率(TER)的评价指标,从柯尔莫哥洛夫复杂性和归一化信息距离(NID)的灵感。这种适应,称为修改TER(mTER),实现适当的规范化和对称处理的参考和假设。通过利用这个平台作为一个大规模的测试场,这项研究表明,与TER相比,mTER的鲁棒性和向后兼容性。SpeechColab排行榜可访问https://github.com/SpeechColab/Leaderboard摘要:In the wake of the surging tide of deep learning over the past decade, Automatic Speech Recognition (ASR) has garnered substantial attention, leading to the emergence of numerous publicly accessible ASR systems that are actively being integrated into our daily lives. Nonetheless, the impartial and replicable evaluation of these ASR systems encounters challenges due to various crucial subtleties. In this paper we introduce the SpeechColab Leaderboard, a general-purpose, open-source platform designed for ASR evaluation. With this platform: (i) We report a comprehensive benchmark, unveiling the current state-of-the-art panorama for ASR systems, covering both open-source models and industrial commercial services. (ii) We quantize how distinct nuances in the scoring pipeline influence the final benchmark outcomes. These include nuances related to capitalization, punctuation, interjection, contraction, synonym usage, compound words, etc. These issues have gained prominence in the context of the transition towards an End-to-End future. (iii) We propose a practical modification to the conventional Token-Error-Rate (TER) evaluation metric, with inspirations from Kolmogorov complexity and Normalized Information Distance (NID). This adaptation, called modified-TER (mTER), achieves proper normalization and symmetrical treatment of reference and hypothesis. By leveraging this platform as a large-scale testing ground, this study demonstrates the robustness and backward compatibility of mTER when compared to TER. The SpeechColab Leaderboard is accessible at https://github.com/SpeechColab/Leaderboard
【7】 Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children标题:自动语音识别(ASR)对韩国儿童言语发音障碍的诊断链接:https://arxiv.org/abs/2403.08187作者:Taekyung Ahn,Yeonjung Hong,Younggon Im,Do Hyung Kim,Dayoung Kang,Joo Won Jeong,Jae Won Kim,Min Jung Kim,Ah-ra Cho,Dae-Hyun Jang,Hosung Nam备注:12 pages, 2 figures摘要:本研究提出了一种自动语音识别(ASR)模型,旨在诊断语音障碍(SSD)儿童的发音问题,以取代临床程序中的手动transmittance。由于为通用目的训练的ASR模型主要将输入语音预测为真实单词,因此采用众所周知的高性能ASR模型来评估患有SSD的儿童的发音是不切实际的。我们对wav 2 vec 2.0 XLS-R模型进行了微调,以将语音识别为发音而不是现有单词。该模型与来自137名语音产生不足的儿童的语音数据集进行了微调,这些儿童发音73个韩语单词,用于实际临床诊断。该模型对单词发音的预测与人类注释的准确率约为90%。虽然该模型在识别不清楚的发音方面仍需改进,但本研究表明,ASR模型可以简化临床领域复杂的发音错误诊断程序。摘要:This study presents a model of automatic speech recognition (ASR) designed to diagnose pronunciation issues in children with speech sound disorders (SSDs) to replace manual transcriptions in clinical procedures. Since ASR models trained for general purposes primarily predict input speech into real words, employing a well-known high-performance ASR model for evaluating pronunciation in children with SSDs is impractical. We fine-tuned the wav2vec 2.0 XLS-R model to recognize speech as pronounced rather than as existing words. The model was fine-tuned with a speech dataset from 137 children with inadequate speech production pronouncing 73 Korean words selected for actual clinical diagnosis. The model's predictions of the pronunciations of the words matched the human annotations with about 90% accuracy. While the model still requires improvement in recognizing unclear pronunciation, this study demonstrates that ASR models can streamline complex pronunciation error diagnostic procedures in clinical fields. 【8】 EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech标题:EM-TTS:高效训练的低资源蒙古语轻量级文语转换系统链接:https://arxiv.org/abs/2403.08164作者:Ziqi Liang,Haoxiang Shi,Jiawei Wang,Keda Lu备注:Accepted by the 27th IEEE International Conference on Computer Supported Cooperative Work in Design (IEEE CSCWD 2024). arXiv admin note: substantial text overlap with arXiv:2211.01948摘要:最近,基于深度学习的文本到语音(TTS)系统已经实现了高质量的语音合成结果。递归神经网络已成为TTS系统中序列数据的标准建模技术,并得到广泛应用。然而,训练包含RNN组件的TTS模型需要强大的GPU性能,并且需要很长时间。相比之下,基于CNN的序列合成技术可以显着减少TTS模型的参数和训练时间,同时由于其高度并行性而保证一定的性能,从而减轻这些训练的经济成本。在本文中,我们提出了一个基于深度卷积神经网络的轻量级TTS系统,它是一个两阶段训练的端到端TTS模型,不使用任何递归单元。我们的模型包括两个阶段:Text 2Spectrum和SSRN。前者用于将音素编码成粗梅尔频谱图,后者用于从粗梅尔频谱图合成完整频谱。同时,我们通过一系列的数据增强,如噪声抑制,时间弯曲,频率掩蔽和时间掩蔽,以提高我们的模型的鲁棒性,解决低资源蒙古问题。实验表明,与主流TTS模型相比,该模型在保证合成语音质量和自然度的同时,减少了训练时间和参数。我们的方法使用NCMMSC 2022-MTTSC Challenge数据集进行验证,在保持一定准确性的同时显著减少了训练时间。摘要:Recently, deep learning-based Text-to-Speech (TTS) systems have achieved high-quality speech synthesis results. Recurrent neural networks have become a standard modeling technique for sequential data in TTS systems and are widely used. However, training a TTS model which includes RNN components requires powerful GPU performance and takes a long time. In contrast, CNN-based sequence synthesis techniques can significantly reduce the parameters and training time of a TTS model while guaranteeing a certain performance due to their high parallelism, which alleviate these economic costs of training. In this paper, we propose a lightweight TTS system based on deep convolutional neural networks, which is a two-stage training end-to-end TTS model and does not employ any recurrent units. Our model consists of two stages: Text2Spectrum and SSRN. The former is used to encode phonemes into a coarse mel spectrogram and the latter is used to synthesize the complete spectrum from the coarse mel spectrogram. Meanwhile, we improve the robustness of our model by a series of data augmentations, such as noise suppression, time warping, frequency masking and time masking, for solving the low resource mongolian problem. Experiments show that our model can reduce the training time and parameters while ensuring the quality and naturalness of the synthesized speech compared to using mainstream TTS models. Our method uses NCMMSC2022-MTTSC Challenge dataset for validation, which significantly reduces training time while maintaining a certain accuracy. 【9】 Motifs, Phrases, and Beyond: The Modelling of Structure in Symbolic Music Generation标题:母题、词组与超越:符号音乐生成中的结构造型链接:https://arxiv.org/abs/2403.07995作者:Keshav Bhandari,Simon Colton备注:Accepted to 13th International Conference on Artificial Intelligence in Music, Sound, Art and Design (EvoMUSART) 2024摘要:音乐结构建模对于生成符号音乐作品的人工智能系统来说至关重要,但也具有挑战性。这篇文献综述剖析了整合连贯结构的技术的演变,从符号方法到基础性和变革性的深度学习方法,这些方法利用了各种各样的训练范式中的计算和数据的力量。在后面的阶段,我们回顾了一个新兴的技术,我们称之为“子任务分解”,涉及分解音乐生成到单独的高层次的结构规划和内容创作阶段。这种系统通过提取旋律骨架或结构模板来指导生成,从而结合了某种形式的音乐知识或神经符号方法。在捕捉所回顾的所有三个时代的主题和重复方面取得了明显的进展,但要以人类作曲家的风格在扩展作品中对主题的细微发展进行建模仍然很困难。我们概述了几个关键的未来方向,以实现协同效益相结合的方法,从所有时代的检查。摘要:Modelling musical structure is vital yet challenging for artificial intelligence systems that generate symbolic music compositions. This literature review dissects the evolution of techniques for incorporating coherent structure, from symbolic approaches to foundational and transformative deep learning methods that harness the power of computation and data across a wide variety of training paradigms. In the later stages, we review an emerging technique which we refer to as "sub-task decomposition" that involves decomposing music generation into separate high-level structural planning and content creation stages. Such systems incorporate some form of musical knowledge or neuro-symbolic methods by extracting melodic skeletons or structural templates to guide the generation. Progress is evident in capturing motifs and repetitions across all three eras reviewed, yet modelling the nuanced development of themes across extended compositions in the style of human composers remains difficult. We outline several key future directions to realize the synergistic benefits of combining approaches from all eras examined. 【10】 Text-to-Audio Generation Synchronized with Videos标题:与视频同步的文本到音频生成链接:https://arxiv.org/abs/2403.07938作者:Shentong Mo,Jing Shi,Yapeng Tian备注:arXiv admin note: text overlap with arXiv:2305.12903摘要:近年来,随着研究人员努力从文本描述合成音频,对文本到音频(TTA)生成的关注已经加强。然而,大多数现有的方法,虽然利用潜在的扩散模型来学习音频和文本嵌入之间的相关性,但在保持所产生的音频和视频之间的无缝同步方面都存在不足。这往往导致明显的视听不匹配。为了弥合这一差距,我们引入了一个突破性的文本到音频生成基准,与视频保持一致,名为T2 AV-Bench。该基准与三个致力于评估视觉对齐和时间一致性的新指标不同。为了补充这一点,我们还提出了一个简单而有效的视频对齐TTA生成模型,即T2 AV。超越传统方法,T2 AV通过整合视觉对齐的文本嵌入作为其条件基础来改进潜在扩散方法。它采用了一个时间多头注意力Transformer来提取和理解视频数据中的时间细微差别,这一壮举被我们的视听控制网放大,它巧妙地将时间视觉表示与文本嵌入相结合。为了进一步加强这种整合,我们编织了一个对比学习目标,旨在确保视觉对齐的文本嵌入与音频功能密切相关。对AudioCaps和T2 AV-Bench的广泛评估表明,我们的T2 AV为视频对齐TTA生成设定了新标准,以确保视觉对齐和时间一致性。摘要:In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models to learn the correlation between audio and text embeddings, fall short when it comes to maintaining a seamless synchronization between the produced audio and its video. This often results in discernible audio-visual mismatches. To bridge this gap, we introduce a groundbreaking benchmark for Text-to-Audio generation that aligns with Videos, named T2AV-Bench. This benchmark distinguishes itself with three novel metrics dedicated to evaluating visual alignment and temporal consistency. To complement this, we also present a simple yet effective video-aligned TTA generation model, namely T2AV. Moving beyond traditional methods, T2AV refines the latent diffusion approach by integrating visual-aligned text embeddings as its conditional foundation. It employs a temporal multi-head attention transformer to extract and understand temporal nuances from video data, a feat amplified by our Audio-Visual ControlNet that adeptly merges temporal visual representations with text embeddings. Further enhancing this integration, we weave in a contrastive learning objective, designed to ensure that the visual-aligned text embeddings resonate closely with the audio features. Extensive evaluations on the AudioCaps and T2AV-Bench demonstrate that our T2AV sets a new standard for video-aligned TTA generation in ensuring visual alignment and temporal consistency.