今日论文合集:cs.SD语音16篇,eess.AS音频处理18篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and Videos
标题: 音频匹配剪切:在电影和视频中查找和创建匹配的音频转换
作者:Dennis Fedorishin,Lie Lu,Srirangaraj Setlur,Venu Govindaraju
备注:Accepted to ICASSP 2024
链接:点击下载PDF文件
摘要:“匹配剪切”是一种常见的视频编辑技术,其中具有相似构图的一对镜头从一个流畅地过渡到另一个。虽然匹配剪辑通常是视觉的,但某些匹配剪辑涉及音频的流畅过渡,其中来自不同来源的声音合并成两个镜头之间的一个难以区分的过渡。在本文中,我们将探索在视频和电影中自动查找和创建“音频匹配剪辑”的能力。我们创建了一个用于音频匹配切割的自监督音频表示,并开发了一个从粗到精的音频匹配管道,该管道推荐匹配镜头并创建混合音频。我们进一步注释了所提出的音频匹配切割任务的数据集,并比较了多个音频表示找到音频匹配切割候选者的能力。最后,我们评估了多种方法来混合两个匹配的音频候选,以创建平滑过渡的目标。项目页面和示例可在https: denfed.github.io audiomatchcut 上获得摘要:A "match cut" is a common video editing technique where a pair of shots that have a similar composition transition fluidly from one to another. Although match cuts are often visual, certain match cuts involve the fluid transition of audio, where sounds from different sources merge into one indistinguishable transition between two shots. In this paper, we explore the ability to automatically find and create "audio match cuts" within videos and movies. We create a self-supervised audio representation for audio match cutting and develop a coarse-to-fine audio match pipeline that recommends matching shots and creates the blended audio. We further annotate a dataset for the proposed audio match cut task and compare the ability of multiple audio representations to find audio match cut candidates. Finally, we evaluate multiple methods to blend two matching audio candidates with the goal of creating a smooth transition. Project page and examples are available at: https: denfed.github.io audiomatchcut

【2】 Rage Music Classification and Analysis using K-Nearest Neighbour, Random Forest, Support Vector Machine, Convolutional Neural Networks, and Gradient Boosting
标题: 使用K近邻、随机森林、支持载体机、卷积神经网络和梯度增强进行愤怒音乐分类和分析
作者:Akul Kumar
链接:点击下载PDF文件
摘要:我们通过包括随机森林、支持向量机、K近邻、梯度提升和卷积神经网络在内的算法,对愤怒音乐(说唱的一个子流派,因对特定歌曲是否属于该流派的分歧而闻名)进行分类。我们比较了机器学习在音频分析应用中的分类方法,并确定了最佳模型。然后,我们分析了存在于愤怒音乐中的最有效的音频特征,同时也确定了关键的音频特征以及更广泛的分离声音变化和趋势。摘要:We classify rage music (a subgenre of rap well-known for disagreements on whether a particular song is part of the genre) with an extensive feature set through algorithms including Random Forest, Support Vector Machine, K-nearest Neighbour, Gradient Boosting, and Convolutional Neural Networks. We compare methods of classification in the application of audio analysis with machine learning and identify optimal models. We then analyze the significant audio features present in and most effective in categorizing rage music, while also identifying key audio features as well as broader separating sonic variations and trends.

【3】 Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?
标题: 当前的Deepfake音频检测模型能否有效检测基于ILM的Deepfake音频?
作者:Yuankun Xie,Chenxu Xiong,Xiaopeng Wang,Zhiyong Wang,Yi Lu,Xin Qi,Ruibo Fu,Yukun Liu,Zhengqi Wen,Jianhua Tao,Guanjun Li,Long Ye
链接:点击下载PDF文件
摘要:目前,由于大型语言模型和音频神经编解码器的发展,音频语言模型(ALM)正在迅速发展。这些ALM大大降低了创建deepfake音频的障碍,产生了高度逼真和多样化的deepfake音频,对社会构成了严重威胁。因此,检测基于ALM的音频的有效音频深度伪造检测技术变得越来越重要。本文研究了当前对策(CM)对基于ALM的音频的有效性。具体来说,我们收集了12种最新的基于ALM的deepfake音频,并利用最新的CM进行评估。我们的研究结果表明,最新的编解码器训练的CM可以有效地检测基于ALM的音频,在大多数ALM测试条件下实现0%的等错误率,这超出了我们的预期。这为基于ALM的deepfake音频检测的未来研究指明了有希望的方向。摘要:Currently, Audio Language Models (ALMs) are rapidly advancing due to the developments in large language models and audio neural codecs. These ALMs have significantly lowered the barrier to creating deepfake audio, generating highly realistic and diverse types of deepfake audio, which pose severe threats to society. Consequently, effective audio deepfake detection technologies to detect ALM-based audio have become increasingly critical. This paper investigate the effectiveness of current countermeasure (CM) against ALM-based audio. Specifically, we collect 12 types of the latest ALM-based deepfake audio and utilizing the latest CMs to evaluate. Our findings reveal that the latest codec-trained CM can effectively detect ALM-based audio, achieving 0% equal error rate under most ALM test conditions, which exceeded our expectations. This indicates promising directions for future research in ALM-based deepfake audio detection.

【4】 EELE: Exploring Efficient and Extensible LoRA Integration in Emotional Text-to-Speech
标题: EELE:探索情感文本到语音中高效且可扩展的LoRA集成
作者:Xin Qi,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Shuchen Shi,Yi Lu,Zhiyong Wang,Xiaopeng Wang,Yuankun Xie,Yukun Liu,Guanjun Li,Xuefei Liu,Yongwei Li
链接:点击下载PDF文件
摘要:在当前的人工智能生成内容(AIGC)时代,出现了一种低秩自适应(LoRA)方法。它采用基于插件的方法来学习新知识,具有较低的参数量和计算成本,并且可以根据特定的子任务插入和退出,具有很高的灵活性。然而,目前的应用方案主要是将LoRA到预先引入的语音模型的条件部分。这固定了LoRA的位置,限制了其应用程序的灵活性和可扩展性。因此,我们提出了探索情感文本到语音(EELE)方法中高效且可扩展的LoRA集成。从一般的中性语音模型开始,我们没有预先引入情感信息,而是使用LoRA插件来设计一个灵活的自适应方案,赋予模型情感生成能力。具体来说,我们最初只使用中性语音数据训练模型。训练完成后,我们将LoRA插入到不同的模块中,并使用情感语音数据对模型进行微调,以找到最佳插入方案。通过实验,我们比较和测试了在模型中不同位置插入LoRA的效果,并评估了LoRA学习各种情绪的能力,有效地证明了我们方法的有效性。此外,我们还探讨了LoRA的秩大小的影响以及与直接微调整个模型相比的差异。摘要:In the current era of Artificial Intelligence Generated Content (AIGC), a Low-Rank Adaptation (LoRA) method has emerged. It uses a plugin-based approach to learn new knowledge with lower parameter quantities and computational costs, and it can be plugged in and out based on the specific sub-tasks, offering high flexibility. However, the current application schemes primarily incorporate LoRA into the pre-introduced conditional parts of the speech models. This fixes the position of LoRA, limiting the flexibility and scalability of its application. Therefore, we propose the Exploring Efficient and Extensible LoRA Integration in Emotional Text-to-Speech (EELE) method. Starting from a general neutral speech model, we do not pre-introduce emotional information but instead use the LoRA plugin to design a flexible adaptive scheme that endows the model with emotional generation capabilities. Specifically, we initially train the model using only neutral speech data. After training is complete, we insert LoRA into different modules and fine-tune the model with emotional speech data to find the optimal insertion scheme. Through experiments, we compare and test the effects of inserting LoRA at different positions within the model and assess LoRA's ability to learn various emotions, effectively proving the validity of our method. Additionally, we explore the impact of the rank size of LoRA and the difference compared to directly fine-tuning the entire model.

【5】 A Noval Feature via Color Quantisation for Fake Audio Detection
标题: 通过颜色量化进行假音频检测的新功能
作者:Zhiyong Wang,Xiaopeng Wang,Yuankun Xie,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Yukun Liu,Guanjun Li,Xin Qi,Yi Lu,Xuefei Liu,Yongwei Li
备注:accepted by ISCSLP2024
链接:点击下载PDF文件
摘要:在deepfake检测领域,以前的研究集中在使用重建或掩码和预测方法来训练预训练模型,然后将其转移到虚假音频检测训练中,其中使用编码器来提取特征,例如wav2vec2.0和Masked Auto Encoder。这些方法已经证明,使用真实音频进行重建预训练可以更好地帮助模型区分假音频。然而,缺点在于可解释性差,这意味着很难直观地呈现deepfake和真实音频之间的差异。本文提出了一种新的特征提取方法,通过颜色量化的限制重建使用有限数量的颜色的光谱图像的输入。所提出的方法确保重建的输入不同于原始的,这允许在光谱重建中直观地观察聚焦区域。在ASVspoof2019数据集上进行的实验表明,与使用原始频谱作为输入相比,所提出的方法具有更好的分类性能,并且预训练重新着色网络也有利于虚假音频检测。摘要:In the field of deepfake detection, previous studies focus on using reconstruction or mask and prediction methods to train pre-trained models, which are then transferred to fake audio detection training where the encoder is used to extract features, such as wav2vec2.0 and Masked Auto Encoder. These methods have proven that using real audio for reconstruction pre-training can better help the model distinguish fake audio. However, the disadvantage lies in poor interpretability, meaning it is hard to intuitively present the differences between deepfake and real audio. This paper proposes a noval feature extraction method via color quantisation which constrains the reconstruction to use a limited number of colors for the spectral image-like input. The proposed method ensures reconstructed input differs from the original, which allows for intuitive observation of the focus areas in the spectral reconstruction. Experiments conducted on the ASVspoof2019 dataset demonstrate that the proposed method achieves better classification performance compared to using the original spectral as input and pretraining the recolor network can also benefit the fake audio detection.

【6】 DisMix: Disentangling Mixtures of Musical Instruments for Source-level Pitch and Timbre Manipulation
标题: DisMix:解开乐器混合物以实现源级音调和音色操纵
作者:Yin-Jyun Luo,Kin Wai Cheuk,Woosung Choi,Toshimitsu Uesaka,Keisuke Toyama,Koichi Saito,Chieh-Hsin Lai,Yuhta Takida,Wei-Hsiang Liao,Simon Dixon,Yuki Mitsufuji
链接:点击下载PDF文件
摘要:关于音高和音色解缠的现有工作主要集中在单乐器音乐音频上,不包括呈现多乐器的情况。为了填补这一空白,我们提出了DisMix,一个生成框架,在这个框架中,音高和音色表示作为构建源的旋律和乐器的模块化构建块,并且其集合形成了一组基于所观察到的混合物的每种乐器的潜在表示。通过操纵的表示,我们的模型样本的组成文书的音高和音色的新组合的混合物。我们可以共同学习解开的音高-音色表示和潜在的扩散Transformer,该扩散transformer重建以源级表示集为条件的混合。我们使用一个简单的数据集孤立的和弦和一个现实的四部分合唱风格的J.S.巴赫,识别成功的解纠缠的关键组件,并演示了基于源代码级属性操作的混合转换的应用。摘要:Existing work on pitch and timbre disentanglement has been mostly focused on single-instrument music audio, excluding the cases where multiple instruments are presented. To fill the gap, we propose DisMix, a generative framework in which the pitch and timbre representations act as modular building blocks for constructing the melody and instrument of a source, and the collection of which forms a set of per-instrument latent representations underlying the observed mixture. By manipulating the representations, our model samples mixtures with novel combinations of pitch and timbre of the constituent instruments. We can jointly learn the disentangled pitch-timbre representations and a latent diffusion transformer that reconstructs the mixture conditioned on the set of source-level representations. We evaluate the model using both a simple dataset of isolated chords and a realistic four-part chorales in the style of J.S. Bach, identify the key components for the success of disentanglement, and demonstrate the application of mixture transformation based on source-level attribute manipulation.

【7】 Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
标题: 迈向无需排练的多语言ASB:基于LoRA的Whisper案例研究
作者:Tianyi Xu,Kaixun Huang,Pengcheng Guo,Yu Zhou,Longtao Huang,Hui Xue,Lei Xie
链接:点击下载PDF文件
摘要:预训练的多语言语音基础模型,如Whisper,在不同语言中表现出令人印象深刻的性能。然而,使这些模型适应新的或特定的语言需要大量计算,并且面临灾难性的遗忘问题。为了解决这些问题,我们的研究调查了在没有原始训练数据的情况下增强新语言模型的策略,同时还保留了原始语言的既定性能。具体来说,我们首先比较各种基于LoRA的方法,以找出它们容易被遗忘的弱点。为了缓解这个问题,我们建议利用原始模型中的LoRA参数对新样本进行近似正交梯度下降。此外,我们还引入了一个可学习的秩系数来分配可训练的参数,以实现更有效的训练。我们的实验与中国耳语模型(维吾尔语和藏族)产生更好的结果与更紧凑的参数集。摘要:Pre-trained multilingual speech foundation models, like Whisper, have shown impressive performance across different languages. However, adapting these models to new or specific languages is computationally extensive and faces catastrophic forgetting problems. Addressing these issues, our study investigates strategies to enhance the model on new languages in the absence of original training data, while also preserving the established performance on the original languages. Specifically, we first compare various LoRA-based methods to find out their vulnerability to forgetting. To mitigate this issue, we propose to leverage the LoRA parameters from the original model for approximate orthogonal gradient descent on the new samples. Additionally, we also introduce a learnable rank coefficient to allocate trainable parameters for more efficient training. Our experiments with a Chinese Whisper model (for Uyghur and Tibetan) yield better results with a more compact parameter set.

【8】 ICSD: An Open-source Dataset for Infant Cry and Snoring Detection
标题: ICSD:用于婴儿哭声和打鼾检测的开源数据集
作者:Qingyu Liu,Longfei Song,Dongxing Xu,Yanhua Long
备注:11 pages, 6 figures
链接:点击下载PDF文件
摘要:婴儿啼哭和打鼾事件的检测和分析是音频信号处理领域内的关键任务。虽然现有的用于一般声音事件检测的数据集非常丰富,但它们通常无法提供足够的、针对婴儿哭声和打鼾的强标记数据。为了提供一个基准数据集,从而促进婴儿哭泣和打鼾检测的研究,本文介绍了婴儿哭泣和打鼾检测(ICSD)数据集,一个新的,公开的数据集专门为ICSD任务设计。ICSD包括三种类型的子集:一个真正的强标记的子集与基于事件的标签手动注释,一个弱标记的子集,只有剪辑级事件注释,和一个合成的子集生成和标记强注释。本文详细介绍了ICSD的创建过程,包括遇到的挑战和采取的解决方案。我们提供了数据集的全面表征,讨论了其局限性和ICSD使用的关键因素。此外,我们对ICSD数据集进行了广泛的实验,以建立基线系统,并在使用该数据集进行ICSD研究时提供对主要因素的见解。我们的目标是开发一个将被社区广泛采用的数据集,作为未来ICSD研究的新开放基准。摘要:The detection and analysis of infant cry and snoring events are crucial tasks within the field of audio signal processing. While existing datasets for general sound event detection are plentiful, they often fall short in providing sufficient, strongly labeled data specific to infant cries and snoring. To provide a benchmark dataset and thus foster the research of infant cry and snoring detection, this paper introduces the Infant Cry and Snoring Detection (ICSD) dataset, a novel, publicly available dataset specially designed for ICSD tasks. The ICSD comprises three types of subsets: a real strongly labeled subset with event-based labels annotated manually, a weakly labeled subset with only clip-level event annotations, and a synthetic subset generated and labeled with strong annotations. This paper provides a detailed description of the ICSD creation process, including the challenges encountered and the solutions adopted. We offer a comprehensive characterization of the dataset, discussing its limitations and key factors for ICSD usage. Additionally, we conduct extensive experiments on the ICSD dataset to establish baseline systems and offer insights into the main factors when using this dataset for ICSD research. Our goal is to develop a dataset that will be widely adopted by the community as a new open benchmark for future ICSD research.

【9】 XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
标题: XCB:语音识别中对跨语言短语进行偏误的有效上下文偏误方法
作者:Xucheng Wan,Naijun Zheng,Kai Liu,Huan Zhou
备注:accepted to NCMMSC 2024
链接:点击下载PDF文件
摘要:已经证明,当预定义的短语列表可用时,上下文化的ASR模型可以有效地提高不常见短语的识别准确率。然而,这些模型往往与双语设置,这是普遍的代码切换语音识别的斗争。在这项研究中,我们提出了初步的尝试,以解决这一挑战,通过引入跨语言的上下文偏置(XCB)模块。具体来说,我们通过集成辅助语言偏置模块和补充语言特定损失来增强主语言的预训练ASR模型,旨在增强对第二语言中短语的识别。在我们的内部代码转换数据集上进行的实验结果验证了我们的方法的有效性,即使没有任何额外的推理开销,在识别第二语言中的偏置短语方面也有显着的改进。此外,我们提出的系统在应用于看不见的ASRU-2019测试集时表现出效率和泛化能力。摘要:Contextualized ASR models have been demonstrated to effectively improve the recognition accuracy of uncommon phrases when a predefined phrase list is available. However, these models often struggle with bilingual settings, which are prevalent in code-switching speech recognition. In this study, we make the initial attempt to address this challenge by introducing a Cross-lingual Contextual Biasing(XCB) module. Specifically, we augment a pre-trained ASR model for the dominant language by integrating an auxiliary language biasing module and a supplementary language-specific loss, aimed at enhancing the recognition of phrases in the secondary language. Experimental results conducted on our in-house code-switching dataset have validated the efficacy of our approach, demonstrating significant improvements in the recognition of biasing phrases in the secondary language, even without any additional inference overhead. Additionally, our proposed system exhibits both efficiency and generalization when is applied by the unseen ASRU-2019 test set.

【10】 SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
标题: SZTU-CMU在MER 2024上:利用Conv-Attention改进描述-LLaMA以实现多模式情感识别
作者:Zebang Cheng,Shuyuan Tu,Dawei Huang,Minghan Li,Xiaojiang Peng,Zhi-Qi Cheng,Alexander G. Hauptmann
链接:点击下载PDF文件
摘要:本文介绍了我们在MER 2024多模态情感识别挑战赛的MER-NOISE和MER-OV轨道上的获胜方法。我们的系统利用了prediction-LLaMA的高级情感理解功能,为未标记的样本生成高质量的注释,解决了有限的标记数据的挑战。为了增强多模态融合,同时减轻特定模态的噪声,我们引入了Conv-Attention,这是一个轻量级且高效的混合框架。大量的实验验证了我们的方法的有效性。在MER-NOISE赛道中,我们的系统实现了最先进的加权平均F分数85.30%,分别超过第二名和第三名团队1.47%和1.65%。对于MER-OV轨道,我们利用扩展LLaMA进行开放词汇注释,与GPT-4V相比,平均准确率和召回率提高了8.52%,在所有参与的大型多模态模型中获得了最高分。可在https: github.com ZebangCheng Emotion-LLaMA上获得反渗透-LLaMA的代码和模型。摘要:This paper presents our winning approach for the MER-NOISE and MER-OV tracks of the MER2024 Challenge on multimodal emotion recognition. Our system leverages the advanced emotional understanding capabilities of Emotion-LLaMA to generate high-quality annotations for unlabeled samples, addressing the challenge of limited labeled data. To enhance multimodal fusion while mitigating modality-specific noise, we introduce Conv-Attention, a lightweight and efficient hybrid framework. Extensive experimentation vali-dates the effectiveness of our approach. In the MER-NOISE track, our system achieves a state-of-the-art weighted average F-score of 85.30%, surpassing the second and third-place teams by 1.47% and 1.65%, respectively. For the MER-OV track, our utilization of Emotion-LLaMA for open-vocabulary annotation yields an 8.52% improvement in average accuracy and recall compared to GPT-4V, securing the highest score among all participating large multimodal models. The code and model for Emotion-LLaMA are available at https: github.com ZebangCheng Emotion-LLaMA.

【11】 Adversarial training of Keyword Spotting to Minimize TTS Data Overfitting
标题: 关键词发现的对抗训练以最大限度地减少TTC数据过度匹配
作者:Hyun Jin Park,Dhruuv Agarwal,Neng Chen,Rentao Sun,Kurt Partridge,Justin Chen,Harry Zhang,Pai Zhu,Jacob Bartel,Kyle Kastner,Gary Wang,Andrew Rosenberg,Quan Wang
备注:to be published in a Workshop at Interspeech 2024, Synthetic Data's Transformative Role in Foundational Speech Models
链接:点击下载PDF文件
摘要:关键词识别(KWS)问题需要大量的真实语音训练数据,以在不同人群中实现高准确性。利用大量的文本到语音(TTS)合成数据可以减少与KWS开发相关的成本和时间。然而,TTS数据可能包含真实语音中不存在的伪影,KWS模型可以利用这些伪影(过拟合),从而导致真实语音的准确性降低。为了解决这个问题,我们建议应用一种对抗性训练方法来防止KWS模型在大量TTS数据上训练时学习TTS特定的特征。实验结果表明,KWS模型在真实语音数据上的准确性可以提高高达12%时,除了原来的KWS损失使用对抗损失。令人惊讶的是,我们还观察到对抗设置将准确率提高了8%,即使只在TTS和真正的负面语音数据上训练,而没有任何真正的正面示例。摘要:The keyword spotting (KWS) problem requires large amounts of real speech training data to achieve high accuracy across diverse populations. Utilizing large amounts of text-to-speech (TTS) synthesized data can reduce the cost and time associated with KWS development. However, TTS data may contain artifacts not present in real speech, which the KWS model can exploit (overfit), leading to degraded accuracy on real speech. To address this issue, we propose applying an adversarial training method to prevent the KWS model from learning TTS-specific features when trained on large amounts of TTS data. Experimental results demonstrate that KWS model accuracy on real speech data can be improved by up to 12% when adversarial loss is used in addition to the original KWS loss. Surprisingly, we also observed that the adversarial setup improves accuracy by up to 8%, even when trained solely on TTS and real negative speech data, without any real positive examples.

【12】 Federated Learning of Large ASR Models in the Real World
标题: 现实世界中大型ASB模型的联邦学习
作者:Yonghui Xiao,Yuxin Ding,Changwan Ryu,Petr Zadrazil,Francoise Beaufays
链接:点击下载PDF文件
摘要:联邦学习(FL)在训练具有隐私保护的机器学习模型方面表现出了很好的效果。然而,对于具有超过1亿个参数的大型模型,训练资源需求成为FL的障碍,因为普通设备没有足够的存储器和计算能力来完成FL任务。虽然已经提出了有效的训练方法,但对基于Conformer的ASR等大型模型的训练仍然是一个挑战。本文提出了一个系统的解决方案,训练的全尺寸ASR模型的130 M参数与FL。据我们所知,这是第一个现实世界的FL应用的Conformer模型,这也是迄今为止最大的模型与FL训练。这是第一篇表明FL可以通过一组改进数据质量和客户端标签的方法来提高ASR模型质量的论文。我们在真实实验中证明了训练效率和模型质量的提高。摘要:Federated learning (FL) has shown promising results on training machine learning models with privacy preservation. However, for large models with over 100 million parameters, the training resource requirement becomes an obstacle for FL because common devices do not have enough memory and computation power to finish the FL tasks. Although efficient training methods have been proposed, it is still a challenge to train the large models like Conformer based ASR. This paper presents a systematic solution to train the full-size ASR models of 130M parameters with FL. To our knowledge, this is the first real-world FL application of the Conformer model, which is also the largest model ever trained with FL so far. And this is the first paper showing FL can improve the ASR model quality with a set of proposed methods to refine the quality of data and labels of clients. We demonstrate both the training efficiency and the model quality improvement in real-world experiments.

【13】 BrewCLIP: A Bifurcated Representation Learning Framework for Audio-Visual Retrieval
标题: BrewCLIP:用于视听检索的分叉表示学习框架
作者:Zhenyu Lu,Lakshay Sethi
链接:点击下载PDF文件
摘要:以前的音频图像匹配方法通常分为两类:管道模型或端到端模型。流水线模型首先转录语音,然后对生成的文本进行编码;端到端模型直接对语音进行编码。通常,管道模型的性能优于端到端模型,但中间转录必然会丢弃一些潜在有用的非文本信息。除了文本信息之外,语音还可以传达诸如口音、情绪和强调之类的细节,这些细节应该在编码表示中被有效地捕获。在本文中,我们调查是否非文本信息,这是被忽视的基于流水线的模型,可以利用,以提高语音图像匹配性能。我们深入分析和比较了端到端模型,管道模型和我们提出的双通道模型,用于在各种数据集上进行稳健的音频图像检索。我们的方法通过利用强大的预训练模型、提示机制和分叉设计,实现了比以前最先进的方法更大的性能增益。摘要:Previous methods for audio-image matching generally fall into one of two categories: pipeline models or End-to-End models. Pipeline models first transcribe speech and then encode the resulting text; End-to-End models encode speech directly. Generally, pipeline models outperform end-to-end models, but the intermediate transcription necessarily discards some potentially useful non-textual information. In addition to textual information, speech can convey details such as accent, mood, and and emphasis, which should be effectively captured in the encoded representation. In this paper, we investigate whether non-textual information, which is overlooked by pipeline-based models, can be leveraged to improve speech-image matching performance. We thoroughly analyze and compare End-to-End models, pipeline models, and our proposed dual-channel model for robust audio-image retrieval on a variety of datasets. Our approach achieves a substantial performance gain over the previous state-of-the-art by leveraging strong pretrained models, a prompting mechanism and a bifurcated design.

【14】 Meta-Learning in Audio and Speech Processing: An End to End Comprehensive Review
标题: 音频和语音处理中的元学习:端到端全面评论
作者:Athul Raimon,Shubha Masti,Shyam K Sateesh,Siyani Vengatagiri,Bhaskarjyoti Das
备注:Survey Paper (15 pages, 1 figure)
链接:点击下载PDF文件
摘要:本调查概述了音频和语音处理场景中使用的各种元学习方法。元学习用于模型性能需要以最少的注释样本最大化的情况,使其适用于低样本音频处理。虽然该领域已经做出了一些重大贡献,音频元学习仍然缺乏全面的调查文件。我们对音频处理中的元学习方法进行了系统回顾。其中包括有关数据增强、特征提取、预处理技术、元学习者、任务选择策略的特定于音频的讨论,还展示了音频中的重要数据集以及关键的现实用例。通过这一广泛的审查,我们的目标是提供有价值的见解,并确定未来的研究方向,在元学习和音频处理的交叉。摘要:This survey overviews various meta-learning approaches used in audio and speech processing scenarios. Meta-learning is used where model performance needs to be maximized with minimum annotated samples, making it suitable for low-sample audio processing. Although the field has made some significant contributions, audio meta-learning still lacks the presence of comprehensive survey papers. We present a systematic review of meta-learning methodologies in audio processing. This includes audio-specific discussions on data augmentation, feature extraction, preprocessing techniques, meta-learners, task selection strategies and also presents important datasets in audio, together with crucial real-world use cases. Through this extensive review, we aim to provide valuable insights and identify future research directions in the intersection of meta-learning and audio processing.

【15】 SSL-TTS: Leveraging Self-Supervised Embeddings and kNN Retrieval for Zero-Shot Multi-speaker TTS
标题: SSL-TTC:利用自监督嵌入和kNN检索实现Zero-Shot多扬声器TTC
作者:Karl El Hajal,Ajinkya Kulkarni,Enno Hermann,Mathew Magimai. -Doss
备注:Submitted to IEEE Signal Processing Letters
链接:点击下载PDF文件
摘要:虽然最近的zero-shot多说话者文本到语音(TTS)模型取得了令人印象深刻的结果,但它们通常依赖于来自众多说话者的大量转录语音数据集和复杂的训练管道。同时,自监督学习(SSL)语音特征已经成为TTS的有效中间表示。还观察到,来自线性接近的不同说话者的SSL特征共享语音信息,同时保持个体说话者身份,这使得能够进行直接和鲁棒的语音克隆。在这项研究中,我们介绍了SSL-TTS,一个轻量级的和高效的zero-shot TTS框架训练从一个单一的扬声器的转录语音。SSL-TTS利用SSL特征和检索方法实现简单而鲁棒的zero-shot多扬声器合成。客观和主观评估表明,我们的方法实现了与需要更大训练数据集的最先进模型相当的性能。低的训练数据要求意味着SSL-TTS非常适合用于低资源领域和语言的多扬声器TTS系统的开发。我们还引入了一个插值参数,使精细控制的输出语音混合的声音。演示示例可在https: idiap.github.io ssl-tts上获得摘要:While recent zero-shot multispeaker text-to-speech (TTS) models achieve impressive results, they typically rely on extensive transcribed speech datasets from numerous speakers and intricate training pipelines. Meanwhile, self-supervised learning (SSL) speech features have emerged as effective intermediate representations for TTS. It was also observed that SSL features from different speakers that are linearly close share phonetic information while maintaining individual speaker identity, which enables straight-forward and robust voice cloning. In this study, we introduce SSL-TTS, a lightweight and efficient zero-shot TTS framework trained on transcribed speech from a single speaker. SSL-TTS leverages SSL features and retrieval methods for simple and robust zero-shot multi-speaker synthesis. Objective and subjective evaluations show that our approach achieves performance comparable to state-of-the-art models that require significantly larger training datasets. The low training data requirements mean that SSL-TTS is well suited for the development of multi-speaker TTS systems for low-resource domains and languages. We also introduce an interpolation parameter which enables fine control over the output speech by blending voices. Demo samples are available at https: idiap.github.io ssl-tts

【16】 ASASVIcomtech: The Vicomtech-UGR Speech Deepfake Detection and SASV Systems for the ASVspoof5 Challenge
标题: ASASVIVomtech:用于ASVspoof 5挑战赛的Vicomtech-UGR语音Deepfake检测和SASV系统
作者:Juan M. Martín-Doñas,Eros Roselló,Angel M. Gomez,Aitor Álvarez,Iván López-Espejo,Antonio M. Peinado
备注:This paper was accepted at ASVspoof Workshop 2024
链接:点击下载PDF文件
摘要:本文介绍了由Vicomtech和格拉纳达大学的研究人员组成的ASASVIcomtech团队为ASVspoof5挑战所做的工作。该团队参与了Track 1(语音deepfake检测)和Track 2(欺骗感知扬声器验证)。这项工作始于对挑战可用数据的分析,这被认为是避免训练模型后期潜在偏差的重要步骤,其主要结论在这里给出。关于所提出的方法,为轨道1开发了一个采用深度复杂卷积递归架构的封闭条件系统,尽管不幸的是,没有取得值得注意的结果。另一方面,开放条件系统的不同可能性,基于利用自我监督模型,从以前的挑战中增强训练数据,以及新颖的声码器,被探索用于两个轨道,最终实现了非常有竞争力的结果与合奏系统。摘要:This paper presents the work carried out by the ASASVIcomtech team, made up of researchers from Vicomtech and University of Granada, for the ASVspoof5 Challenge. The team has participated in both Track 1 (speech deepfake detection) and Track 2 (spoofing-aware speaker verification). This work started with an analysis of the challenge available data, which was regarded as an essential step to avoid later potential biases of the trained models, and whose main conclusions are presented here. With respect to the proposed approaches, a closed-condition system employing a deep complex convolutional recurrent architecture was developed for Track 1, although, unfortunately, no noteworthy results were achieved. On the other hand, different possibilities of open-condition systems, based on leveraging self-supervised models, augmented training data from previous challenges, and novel vocoders, were explored for both tracks, finally achieving very competitive results with an ensemble system.


eess.AS音频处理
【1】 SSL-TTS: Leveraging Self-Supervised Embeddings and kNN Retrieval for Zero-Shot Multi-speaker TTS
标题: SSL-TTC:利用自监督嵌入和kNN检索实现Zero-Shot多扬声器TTC
作者:Karl El Hajal,Ajinkya Kulkarni,Enno Hermann,Mathew Magimai. -Doss
备注:Submitted to IEEE Signal Processing Letters
链接:点击下载PDF文件
摘要:虽然最近的zero-shot多说话者文本到语音(TTS)模型取得了令人印象深刻的结果,但它们通常依赖于来自众多说话者的大量转录语音数据集和复杂的训练管道。同时,自监督学习(SSL)语音特征已经成为TTS的有效中间表示。还观察到,来自线性接近的不同说话者的SSL特征共享语音信息,同时保持个体说话者身份,这使得能够进行直接和鲁棒的语音克隆。在这项研究中,我们介绍了SSL-TTS,一个轻量级的和高效的zero-shot TTS框架训练从一个单一的扬声器的转录语音。SSL-TTS利用SSL特征和检索方法实现简单而鲁棒的zero-shot多扬声器合成。客观和主观评估表明,我们的方法实现了与需要更大训练数据集的最先进模型相当的性能。低的训练数据要求意味着SSL-TTS非常适合用于低资源领域和语言的多扬声器TTS系统的开发。我们还引入了一个插值参数,使精细控制的输出语音混合的声音。演示示例可在https: idiap.github.io ssl-tts上获得摘要:While recent zero-shot multispeaker text-to-speech (TTS) models achieve impressive results, they typically rely on extensive transcribed speech datasets from numerous speakers and intricate training pipelines. Meanwhile, self-supervised learning (SSL) speech features have emerged as effective intermediate representations for TTS. It was also observed that SSL features from different speakers that are linearly close share phonetic information while maintaining individual speaker identity, which enables straight-forward and robust voice cloning. In this study, we introduce SSL-TTS, a lightweight and efficient zero-shot TTS framework trained on transcribed speech from a single speaker. SSL-TTS leverages SSL features and retrieval methods for simple and robust zero-shot multi-speaker synthesis. Objective and subjective evaluations show that our approach achieves performance comparable to state-of-the-art models that require significantly larger training datasets. The low training data requirements mean that SSL-TTS is well suited for the development of multi-speaker TTS systems for low-resource domains and languages. We also introduce an interpolation parameter which enables fine control over the output speech by blending voices. Demo samples are available at https: idiap.github.io ssl-tts

【2】 ASASVIcomtech: The Vicomtech-UGR Speech Deepfake Detection and SASV Systems for the ASVspoof5 Challenge
标题: ASASVIVomtech:用于ASVspoof 5挑战赛的Vicomtech-UGR语音Deepfake检测和SASV系统
作者:Juan M. Martín-Doñas,Eros Roselló,Angel M. Gomez,Aitor Álvarez,Iván López-Espejo,Antonio M. Peinado
备注:This paper was accepted at ASVspoof Workshop 2024
链接:点击下载PDF文件
摘要:本文介绍了由Vicomtech和格拉纳达大学的研究人员组成的ASASVIcomtech团队为ASVspoof5挑战所做的工作。该团队参与了Track 1(语音deepfake检测)和Track 2(欺骗感知扬声器验证)。这项工作始于对挑战可用数据的分析,这被认为是避免训练模型后期潜在偏差的重要步骤,其主要结论在这里给出。关于所提出的方法,为轨道1开发了一个采用深度复杂卷积递归架构的封闭条件系统,尽管不幸的是,没有取得值得注意的结果。另一方面,开放条件系统的不同可能性,基于利用自我监督模型,从以前的挑战中增强训练数据,以及新颖的声码器,被探索用于两个轨道,最终实现了非常有竞争力的结果与合奏系统。摘要:This paper presents the work carried out by the ASASVIcomtech team, made up of researchers from Vicomtech and University of Granada, for the ASVspoof5 Challenge. The team has participated in both Track 1 (speech deepfake detection) and Track 2 (spoofing-aware speaker verification). This work started with an analysis of the challenge available data, which was regarded as an essential step to avoid later potential biases of the trained models, and whose main conclusions are presented here. With respect to the proposed approaches, a closed-condition system employing a deep complex convolutional recurrent architecture was developed for Track 1, although, unfortunately, no noteworthy results were achieved. On the other hand, different possibilities of open-condition systems, based on leveraging self-supervised models, augmented training data from previous challenges, and novel vocoders, were explored for both tracks, finally achieving very competitive results with an ensemble system.

【3】 Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and Videos
标题: 音频匹配剪切:在电影和视频中查找和创建匹配的音频转换
作者:Dennis Fedorishin,Lie Lu,Srirangaraj Setlur,Venu Govindaraju
备注:Accepted to ICASSP 2024
链接:点击下载PDF文件
摘要:“匹配剪切”是一种常见的视频编辑技术,其中具有相似构图的一对镜头从一个流畅地过渡到另一个。虽然匹配剪辑通常是视觉的,但某些匹配剪辑涉及音频的流畅过渡,其中来自不同来源的声音合并成两个镜头之间的一个难以区分的过渡。在本文中,我们将探索在视频和电影中自动查找和创建“音频匹配剪辑”的能力。我们创建了一个用于音频匹配切割的自监督音频表示,并开发了一个从粗到精的音频匹配管道,该管道推荐匹配镜头并创建混合音频。我们进一步注释了所提出的音频匹配切割任务的数据集,并比较了多个音频表示找到音频匹配切割候选者的能力。最后,我们评估了多种方法来混合两个匹配的音频候选,以创建平滑过渡的目标。项目页面和示例可在https: denfed.github.io audiomatchcut 上获得摘要:A "match cut" is a common video editing technique where a pair of shots that have a similar composition transition fluidly from one to another. Although match cuts are often visual, certain match cuts involve the fluid transition of audio, where sounds from different sources merge into one indistinguishable transition between two shots. In this paper, we explore the ability to automatically find and create "audio match cuts" within videos and movies. We create a self-supervised audio representation for audio match cutting and develop a coarse-to-fine audio match pipeline that recommends matching shots and creates the blended audio. We further annotate a dataset for the proposed audio match cut task and compare the ability of multiple audio representations to find audio match cut candidates. Finally, we evaluate multiple methods to blend two matching audio candidates with the goal of creating a smooth transition. Project page and examples are available at: https: denfed.github.io audiomatchcut

【4】 Disentangling segmental and prosodic factors to non-native speech comprehensibility
标题: 解析非母语言语可理解性的分段和韵律因素
作者:Waris Quamer,Ricardo Gutierrez-Osuna
链接:点击下载PDF文件
摘要:目前的口音转换系统没有解决非母语口音的两个主要来源:音段特征和韵律特征。能够独立地操纵非本族语者的音段和 或韵律通道对于量化这两个通道对言语可理解性和社会态度的贡献至关重要。我们提出了一个AC系统,不仅从口音的语音质量,但也解开后者到它的分段和韵律特征。该系统能够生成口音转换,该口音转换结合了(1)源话语的分段特征,(2)目标话语的语音特征,以及(3)参考话语的韵律。我们表明,矢量量化的声学嵌入和连续重复的码字的去除允许系统转移韵律和提高语音相似性。我们进行了感知听力测试,以量化的个人贡献的音段特征和韵律的感知可理解性的非本族语。我们的研究结果表明,与先前的研究在非母语的语音,段的功能有更大的影响比韵律的可理解性。建议AC系统也可以用来研究如何节段和韵律线索影响社会态度对非母语的讲话。摘要:Current accent conversion (AC) systems do not disentangle the two main sources of non-native accent: segmental and prosodic characteristics. Being able to manipulate a non-native speaker's segmental and or prosodic channels independently is critical to quantify how these two channels contribute to speech comprehensibility and social attitudes. We present an AC system that not only decouples voice quality from accent, but also disentangles the latter into its segmental and prosodic characteristics. The system is able to generate accent conversions that combine (1) the segmental characteristics from a source utterance, (2) the voice characteristics from a target utterance, and (3) the prosody of a reference utterance. We show that vector quantization of acoustic embeddings and removal of consecutive duplicated codewords allows the system to transfer prosody and improve voice similarity. We conduct perceptual listening tests to quantify the individual contributions of segmental features and prosody on the perceived comprehensibility of non-native speech. Our results indicate that, contrary to prior research in non-native speech, segmental features have a larger impact on comprehensibility than prosody. The proposed AC system may also be used to study how segmental and prosody cues affect social attitudes towards non-native speech.

【5】 Rage Music Classification and Analysis using K-Nearest Neighbour, Random Forest, Support Vector Machine, Convolutional Neural Networks, and Gradient Boosting
标题: 使用K近邻、随机森林、支持载体机、卷积神经网络和梯度增强进行愤怒音乐分类和分析
作者:Akul Kumar
链接:点击下载PDF文件
摘要:我们通过包括随机森林、支持向量机、K近邻、梯度提升和卷积神经网络在内的算法,对愤怒音乐(说唱的一个子流派,因对特定歌曲是否属于该流派的分歧而闻名)进行分类。我们比较了机器学习在音频分析应用中的分类方法,并确定了最佳模型。然后,我们分析了存在于愤怒音乐中的最有效的音频特征,同时也确定了关键的音频特征以及更广泛的分离声音变化和趋势。摘要:We classify rage music (a subgenre of rap well-known for disagreements on whether a particular song is part of the genre) with an extensive feature set through algorithms including Random Forest, Support Vector Machine, K-nearest Neighbour, Gradient Boosting, and Convolutional Neural Networks. We compare methods of classification in the application of audio analysis with machine learning and identify optimal models. We then analyze the significant audio features present in and most effective in categorizing rage music, while also identifying key audio features as well as broader separating sonic variations and trends.

【6】 Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?
标题: 当前的Deepfake音频检测模型能否有效检测基于ILM的Deepfake音频?
作者:Yuankun Xie,Chenxu Xiong,Xiaopeng Wang,Zhiyong Wang,Yi Lu,Xin Qi,Ruibo Fu,Yukun Liu,Zhengqi Wen,Jianhua Tao,Guanjun Li,Long Ye
链接:点击下载PDF文件
摘要:目前,由于大型语言模型和音频神经编解码器的发展,音频语言模型(ALM)正在迅速发展。这些ALM大大降低了创建deepfake音频的障碍,产生了高度逼真和多样化的deepfake音频,对社会构成了严重威胁。因此,检测基于ALM的音频的有效音频深度伪造检测技术变得越来越重要。本文研究了当前对策(CM)对基于ALM的音频的有效性。具体来说,我们收集了12种最新的基于ALM的deepfake音频,并利用最新的CM进行评估。我们的研究结果表明,最新的编解码器训练的CM可以有效地检测基于ALM的音频,在大多数ALM测试条件下实现0%的等错误率,这超出了我们的预期。这为基于ALM的deepfake音频检测的未来研究指明了有希望的方向。摘要:Currently, Audio Language Models (ALMs) are rapidly advancing due to the developments in large language models and audio neural codecs. These ALMs have significantly lowered the barrier to creating deepfake audio, generating highly realistic and diverse types of deepfake audio, which pose severe threats to society. Consequently, effective audio deepfake detection technologies to detect ALM-based audio have become increasingly critical. This paper investigate the effectiveness of current countermeasure (CM) against ALM-based audio. Specifically, we collect 12 types of the latest ALM-based deepfake audio and utilizing the latest CMs to evaluate. Our findings reveal that the latest codec-trained CM can effectively detect ALM-based audio, achieving 0% equal error rate under most ALM test conditions, which exceeded our expectations. This indicates promising directions for future research in ALM-based deepfake audio detection.

【7】 EELE: Exploring Efficient and Extensible LoRA Integration in Emotional Text-to-Speech
标题: EELE:探索情感文本到语音中高效且可扩展的LoRA集成
作者:Xin Qi,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Shuchen Shi,Yi Lu,Zhiyong Wang,Xiaopeng Wang,Yuankun Xie,Yukun Liu,Guanjun Li,Xuefei Liu,Yongwei Li
链接:点击下载PDF文件
摘要:在当前的人工智能生成内容(AIGC)时代,出现了一种低秩自适应(LoRA)方法。它采用基于插件的方法来学习新知识,具有较低的参数量和计算成本,并且可以根据特定的子任务插入和退出,具有很高的灵活性。然而,目前的应用方案主要是将LoRA到预先引入的语音模型的条件部分。这固定了LoRA的位置,限制了其应用程序的灵活性和可扩展性。因此,我们提出了在情感文本到语音(EELE)方法中探索有效和可扩展的LoRA集成。从一般的中性语音模型开始,我们没有预先引入情感信息,而是使用LoRA插件来设计一个灵活的自适应方案,赋予模型情感生成能力。具体来说,我们最初只使用中性语音数据训练模型。训练完成后,我们将LoRA插入到不同的模块中,并使用情感语音数据对模型进行微调,以找到最佳插入方案。通过实验,我们比较和测试了在模型内不同位置插入LoRA的效果,评估了LoRA学习各种情绪的能力,有效地证明了我们方法的有效性。此外,我们还探讨了LoRA的秩大小的影响以及与直接微调整个模型相比的差异。摘要:In the current era of Artificial Intelligence Generated Content (AIGC), a Low-Rank Adaptation (LoRA) method has emerged. It uses a plugin-based approach to learn new knowledge with lower parameter quantities and computational costs, and it can be plugged in and out based on the specific sub-tasks, offering high flexibility. However, the current application schemes primarily incorporate LoRA into the pre-introduced conditional parts of the speech models. This fixes the position of LoRA, limiting the flexibility and scalability of its application. Therefore, we propose the Exploring Efficient and Extensible LoRA Integration in Emotional Text-to-Speech (EELE) method. Starting from a general neutral speech model, we do not pre-introduce emotional information but instead use the LoRA plugin to design a flexible adaptive scheme that endows the model with emotional generation capabilities. Specifically, we initially train the model using only neutral speech data. After training is complete, we insert LoRA into different modules and fine-tune the model with emotional speech data to find the optimal insertion scheme. Through experiments, we compare and test the effects of inserting LoRA at different positions within the model and assess LoRA's ability to learn various emotions, effectively proving the validity of our method. Additionally, we explore the impact of the rank size of LoRA and the difference compared to directly fine-tuning the entire model.

【8】 A Noval Feature via Color Quantisation for Fake Audio Detection
标题: 通过颜色量化进行假音频检测的新功能
作者:Zhiyong Wang,Xiaopeng Wang,Yuankun Xie,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Yukun Liu,Guanjun Li,Xin Qi,Yi Lu,Xuefei Liu,Yongwei Li
备注:accepted by ISCSLP2024
链接:点击下载PDF文件
摘要:在deepfake检测领域,以前的研究集中在使用重建或掩码和预测方法来训练预训练模型,然后将其转移到虚假音频检测训练中,其中使用编码器来提取特征,例如wav2vec2.0和Masked Auto Encoder。这些方法已经证明,使用真实音频进行重建预训练可以更好地帮助模型区分假音频。然而,缺点在于可解释性差,这意味着很难直观地呈现deepfake和真实音频之间的差异。本文提出了一种新的特征提取方法,通过颜色量化的限制重建使用有限数量的颜色的光谱图像的输入。所提出的方法确保重建的输入不同于原始的,这允许在光谱重建中直观地观察聚焦区域。在ASVspoof2019数据集上进行的实验表明,与使用原始频谱作为输入相比,所提出的方法具有更好的分类性能,并且预训练重新着色网络也有利于虚假音频检测。摘要:In the field of deepfake detection, previous studies focus on using reconstruction or mask and prediction methods to train pre-trained models, which are then transferred to fake audio detection training where the encoder is used to extract features, such as wav2vec2.0 and Masked Auto Encoder. These methods have proven that using real audio for reconstruction pre-training can better help the model distinguish fake audio. However, the disadvantage lies in poor interpretability, meaning it is hard to intuitively present the differences between deepfake and real audio. This paper proposes a noval feature extraction method via color quantisation which constrains the reconstruction to use a limited number of colors for the spectral image-like input. The proposed method ensures reconstructed input differs from the original, which allows for intuitive observation of the focus areas in the spectral reconstruction. Experiments conducted on the ASVspoof2019 dataset demonstrate that the proposed method achieves better classification performance compared to using the original spectral as input and pretraining the recolor network can also benefit the fake audio detection.

【9】 DisMix: Disentangling Mixtures of Musical Instruments for Source-level Pitch and Timbre Manipulation
标题: DisMix:解开乐器混合物以实现源级音调和音色操纵
作者:Yin-Jyun Luo,Kin Wai Cheuk,Woosung Choi,Toshimitsu Uesaka,Keisuke Toyama,Koichi Saito,Chieh-Hsin Lai,Yuhta Takida,Wei-Hsiang Liao,Simon Dixon,Yuki Mitsufuji
链接:点击下载PDF文件
摘要:关于音高和音色解缠的现有工作主要集中在单乐器音乐音频上,不包括呈现多乐器的情况。为了填补这一空白,我们提出了DisMix,一个生成框架,在这个框架中,音高和音色表示作为构建源的旋律和乐器的模块化构建块,并且其集合形成了一组基于所观察到的混合物的每种乐器的潜在表示。通过操纵的表示,我们的模型样本的组成文书的音高和音色的新组合的混合物。我们可以共同学习解开的音高-音色表示和潜在的扩散Transformer,该扩散transformer重建以源级表示集为条件的混合。我们使用一个简单的数据集孤立的和弦和一个现实的四部分合唱风格的J.S.巴赫,识别成功的解纠缠的关键组件,并演示了基于源代码级属性操作的混合转换的应用。摘要:Existing work on pitch and timbre disentanglement has been mostly focused on single-instrument music audio, excluding the cases where multiple instruments are presented. To fill the gap, we propose DisMix, a generative framework in which the pitch and timbre representations act as modular building blocks for constructing the melody and instrument of a source, and the collection of which forms a set of per-instrument latent representations underlying the observed mixture. By manipulating the representations, our model samples mixtures with novel combinations of pitch and timbre of the constituent instruments. We can jointly learn the disentangled pitch-timbre representations and a latent diffusion transformer that reconstructs the mixture conditioned on the set of source-level representations. We evaluate the model using both a simple dataset of isolated chords and a realistic four-part chorales in the style of J.S. Bach, identify the key components for the success of disentanglement, and demonstrate the application of mixture transformation based on source-level attribute manipulation.

【10】 Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
标题: 迈向无需排练的多语言ASB:基于LoRA的Whisper案例研究
作者:Tianyi Xu,Kaixun Huang,Pengcheng Guo,Yu Zhou,Longtao Huang,Hui Xue,Lei Xie
链接:点击下载PDF文件
摘要:预训练的多语言语音基础模型,如Whisper,在不同语言中表现出令人印象深刻的性能。然而,使这些模型适应新的或特定的语言需要大量计算,并且面临灾难性的遗忘问题。为了解决这些问题,我们的研究调查了在没有原始训练数据的情况下增强新语言模型的策略,同时还保留了原始语言的既定性能。具体来说,我们首先比较各种基于LoRA的方法,以找出它们容易被遗忘的弱点。为了缓解这个问题,我们建议利用原始模型中的LoRA参数对新样本进行近似正交梯度下降。此外,我们还引入了一个可学习的秩系数来分配可训练的参数,以实现更有效的训练。我们的实验与中国耳语模型(维吾尔语和藏族)产生更好的结果与更紧凑的参数集。摘要:Pre-trained multilingual speech foundation models, like Whisper, have shown impressive performance across different languages. However, adapting these models to new or specific languages is computationally extensive and faces catastrophic forgetting problems. Addressing these issues, our study investigates strategies to enhance the model on new languages in the absence of original training data, while also preserving the established performance on the original languages. Specifically, we first compare various LoRA-based methods to find out their vulnerability to forgetting. To mitigate this issue, we propose to leverage the LoRA parameters from the original model for approximate orthogonal gradient descent on the new samples. Additionally, we also introduce a learnable rank coefficient to allocate trainable parameters for more efficient training. Our experiments with a Chinese Whisper model (for Uyghur and Tibetan) yield better results with a more compact parameter set.

【11】 ICSD: An Open-source Dataset for Infant Cry and Snoring Detection
标题: ICSD:用于婴儿哭声和打鼾检测的开源数据集
作者:Qingyu Liu,Longfei Song,Dongxing Xu,Yanhua Long
备注:11 pages, 6 figures
链接:点击下载PDF文件
摘要:婴儿啼哭和打鼾事件的检测和分析是音频信号处理领域内的关键任务。虽然现有的用于一般声音事件检测的数据集非常丰富,但它们通常无法提供足够的、针对婴儿哭声和打鼾的强标记数据。为了提供一个基准数据集,从而促进婴儿哭泣和打鼾检测的研究,本文介绍了婴儿哭泣和打鼾检测(ICSD)数据集,一个新的,公开的数据集专门为ICSD任务设计。ICSD包括三种类型的子集:一个真正的强标记的子集与基于事件的标签手动注释,一个弱标记的子集,只有剪辑级事件注释,和一个合成的子集生成和标记强注释。本文详细介绍了ICSD的创建过程,包括遇到的挑战和采取的解决方案。我们提供了数据集的全面表征,讨论了其局限性和ICSD使用的关键因素。此外,我们对ICSD数据集进行了广泛的实验,以建立基线系统,并在使用该数据集进行ICSD研究时提供对主要因素的见解。我们的目标是开发一个将被社区广泛采用的数据集,作为未来ICSD研究的新开放基准。摘要:The detection and analysis of infant cry and snoring events are crucial tasks within the field of audio signal processing. While existing datasets for general sound event detection are plentiful, they often fall short in providing sufficient, strongly labeled data specific to infant cries and snoring. To provide a benchmark dataset and thus foster the research of infant cry and snoring detection, this paper introduces the Infant Cry and Snoring Detection (ICSD) dataset, a novel, publicly available dataset specially designed for ICSD tasks. The ICSD comprises three types of subsets: a real strongly labeled subset with event-based labels annotated manually, a weakly labeled subset with only clip-level event annotations, and a synthetic subset generated and labeled with strong annotations. This paper provides a detailed description of the ICSD creation process, including the challenges encountered and the solutions adopted. We offer a comprehensive characterization of the dataset, discussing its limitations and key factors for ICSD usage. Additionally, we conduct extensive experiments on the ICSD dataset to establish baseline systems and offer insights into the main factors when using this dataset for ICSD research. Our goal is to develop a dataset that will be widely adopted by the community as a new open benchmark for future ICSD research.

【12】 XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
标题: XCB:语音识别中对跨语言短语进行偏误的有效上下文偏误方法
作者:Xucheng Wan,Naijun Zheng,Kai Liu,Huan Zhou
备注:accepted to NCMMSC 2024
链接:点击下载PDF文件
摘要:已经证明,当预定义的短语列表可用时,上下文化的ASR模型可以有效地提高不常见短语的识别准确率。然而,这些模型往往与双语设置,这是普遍的代码切换语音识别的斗争。在这项研究中,我们通过引入跨语言上下文偏置(XCB)模块来应对这一挑战。具体来说,我们通过集成辅助语言偏置模块和补充语言特定损失来增强主语言的预训练ASR模型,旨在增强对第二语言中短语的识别。在我们的内部代码转换数据集上进行的实验结果验证了我们的方法的有效性,即使没有任何额外的推理开销,在识别第二语言中的偏置短语方面也有显着的改进。此外,我们提出的系统在应用于看不见的ASRU-2019测试集时表现出效率和泛化能力。摘要:Contextualized ASR models have been demonstrated to effectively improve the recognition accuracy of uncommon phrases when a predefined phrase list is available. However, these models often struggle with bilingual settings, which are prevalent in code-switching speech recognition. In this study, we make the initial attempt to address this challenge by introducing a Cross-lingual Contextual Biasing(XCB) module. Specifically, we augment a pre-trained ASR model for the dominant language by integrating an auxiliary language biasing module and a supplementary language-specific loss, aimed at enhancing the recognition of phrases in the secondary language. Experimental results conducted on our in-house code-switching dataset have validated the efficacy of our approach, demonstrating significant improvements in the recognition of biasing phrases in the secondary language, even without any additional inference overhead. Additionally, our proposed system exhibits both efficiency and generalization when is applied by the unseen ASRU-2019 test set.

【13】 SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
标题: SZTU-CMU在MER 2024上:利用Conv-Attention改进描述-LLaMA以实现多模式情感识别
作者:Zebang Cheng,Shuyuan Tu,Dawei Huang,Minghan Li,Xiaojiang Peng,Zhi-Qi Cheng,Alexander G. Hauptmann
链接:点击下载PDF文件
摘要:本文介绍了我们在MER 2024多模态情感识别挑战赛的MER-NOISE和MER-OV轨道上的获胜方法。我们的系统利用了prediction-LLaMA的高级情感理解功能,为未标记的样本生成高质量的注释,解决了有限的标记数据的挑战。为了增强多模态融合,同时减轻特定模态的噪声,我们引入了Conv-Attention,这是一个轻量级且高效的混合框架。大量的实验验证了我们的方法的有效性。在MER-NOISE赛道中,我们的系统实现了最先进的加权平均F分数85.30%,分别超过第二名和第三名团队1.47%和1.65%。对于MER-OV曲目,与GPT-4V相比,我们利用Emotion-LLaMA进行开放词汇注释,平均准确率和召回率提高了8.52%,在所有参与的大型多模式模型中获得了最高分。可在https: github.com ZebangCheng Emotion-LLaMA上获得反渗透-LLaMA的代码和模型。摘要:This paper presents our winning approach for the MER-NOISE and MER-OV tracks of the MER2024 Challenge on multimodal emotion recognition. Our system leverages the advanced emotional understanding capabilities of Emotion-LLaMA to generate high-quality annotations for unlabeled samples, addressing the challenge of limited labeled data. To enhance multimodal fusion while mitigating modality-specific noise, we introduce Conv-Attention, a lightweight and efficient hybrid framework. Extensive experimentation vali-dates the effectiveness of our approach. In the MER-NOISE track, our system achieves a state-of-the-art weighted average F-score of 85.30%, surpassing the second and third-place teams by 1.47% and 1.65%, respectively. For the MER-OV track, our utilization of Emotion-LLaMA for open-vocabulary annotation yields an 8.52% improvement in average accuracy and recall compared to GPT-4V, securing the highest score among all participating large multimodal models. The code and model for Emotion-LLaMA are available at https: github.com ZebangCheng Emotion-LLaMA.

【14】 Adversarial training of Keyword Spotting to Minimize TTS Data Overfitting
标题: 关键词发现的对抗训练以最大限度地减少TTC数据过度匹配
作者:Hyun Jin Park,Dhruuv Agarwal,Neng Chen,Rentao Sun,Kurt Partridge,Justin Chen,Harry Zhang,Pai Zhu,Jacob Bartel,Kyle Kastner,Gary Wang,Andrew Rosenberg,Quan Wang
备注:to be published in a Workshop at Interspeech 2024, Synthetic Data's Transformative Role in Foundational Speech Models
链接:点击下载PDF文件
摘要:关键词识别(KWS)问题需要大量的真实语音训练数据,以在不同人群中实现高准确性。利用大量的文本到语音(TTS)合成数据可以减少与KWS开发相关的成本和时间。然而,TTS数据可能包含真实语音中不存在的伪影,KWS模型可以利用这些伪影(过拟合),从而导致真实语音的准确性降低。为了解决这个问题,我们建议应用一种对抗性训练方法来防止KWS模型在大量TTS数据上训练时学习TTS特定的特征。实验结果表明,KWS模型在真实语音数据上的准确性可以提高高达12%时,除了原来的KWS损失使用对抗损失。令人惊讶的是,我们还观察到对抗设置将准确率提高了8%,即使只在TTS和真正的负面语音数据上训练,而没有任何真正的正面示例。摘要:The keyword spotting (KWS) problem requires large amounts of real speech training data to achieve high accuracy across diverse populations. Utilizing large amounts of text-to-speech (TTS) synthesized data can reduce the cost and time associated with KWS development. However, TTS data may contain artifacts not present in real speech, which the KWS model can exploit (overfit), leading to degraded accuracy on real speech. To address this issue, we propose applying an adversarial training method to prevent the KWS model from learning TTS-specific features when trained on large amounts of TTS data. Experimental results demonstrate that KWS model accuracy on real speech data can be improved by up to 12% when adversarial loss is used in addition to the original KWS loss. Surprisingly, we also observed that the adversarial setup improves accuracy by up to 8%, even when trained solely on TTS and real negative speech data, without any real positive examples.

【15】 Federated Learning of Large ASR Models in the Real World
标题: 现实世界中大型ASB模型的联邦学习
作者:Yonghui Xiao,Yuxin Ding,Changwan Ryu,Petr Zadrazil,Francoise Beaufays
链接:点击下载PDF文件
摘要:联邦学习(FL)在训练具有隐私保护的机器学习模型方面表现出了很好的效果。然而,对于具有超过1亿个参数的大型模型,训练资源需求成为FL的障碍,因为普通设备没有足够的存储器和计算能力来完成FL任务。虽然已经提出了有效的训练方法,但对基于Conformer的ASR等大型模型的训练仍然是一个挑战。本文提出了一个系统的解决方案,训练的全尺寸ASR模型的130 M参数与FL。据我们所知,这是第一个现实世界的FL应用的Conformer模型,这也是迄今为止最大的模型与FL训练。这是第一篇表明FL可以通过一组改进数据质量和客户端标签的方法来提高ASR模型质量的论文。我们在真实实验中证明了训练效率和模型质量的提高。摘要:Federated learning (FL) has shown promising results on training machine learning models with privacy preservation. However, for large models with over 100 million parameters, the training resource requirement becomes an obstacle for FL because common devices do not have enough memory and computation power to finish the FL tasks. Although efficient training methods have been proposed, it is still a challenge to train the large models like Conformer based ASR. This paper presents a systematic solution to train the full-size ASR models of 130M parameters with FL. To our knowledge, this is the first real-world FL application of the Conformer model, which is also the largest model ever trained with FL so far. And this is the first paper showing FL can improve the ASR model quality with a set of proposed methods to refine the quality of data and labels of clients. We demonstrate both the training efficiency and the model quality improvement in real-world experiments.

【16】 BrewCLIP: A Bifurcated Representation Learning Framework for Audio-Visual Retrieval
标题: BrewCLIP:用于视听检索的分叉表示学习框架
作者:Zhenyu Lu,Lakshay Sethi
链接:点击下载PDF文件
摘要:以前的音频图像匹配方法通常分为两类:管道模型或端到端模型。流水线模型首先转录语音,然后对生成的文本进行编码;端到端模型直接对语音进行编码。通常,管道模型优于端到端模型,但中间转录必然会丢弃一些潜在有用的非文本信息。除了文本信息之外,语音还可以传达诸如口音、情绪和强调之类的细节,这些细节应该在编码表示中被有效地捕获。在本文中,我们调查是否非文本信息,这是被忽视的基于流水线的模型,可以利用,以提高语音图像匹配性能。我们深入分析和比较了端到端模型,管道模型和我们提出的双通道模型,用于在各种数据集上进行稳健的音频图像检索。我们的方法通过利用强大的预训练模型、提示机制和分叉设计,实现了比以前最先进的方法更大的性能增益。摘要:Previous methods for audio-image matching generally fall into one of two categories: pipeline models or End-to-End models. Pipeline models first transcribe speech and then encode the resulting text; End-to-End models encode speech directly. Generally, pipeline models outperform end-to-end models, but the intermediate transcription necessarily discards some potentially useful non-textual information. In addition to textual information, speech can convey details such as accent, mood, and and emphasis, which should be effectively captured in the encoded representation. In this paper, we investigate whether non-textual information, which is overlooked by pipeline-based models, can be leveraged to improve speech-image matching performance. We thoroughly analyze and compare End-to-End models, pipeline models, and our proposed dual-channel model for robust audio-image retrieval on a variety of datasets. Our approach achieves a substantial performance gain over the previous state-of-the-art by leveraging strong pretrained models, a prompting mechanism and a bifurcated design.

【17】 Meta-Learning in Audio and Speech Processing: An End to End Comprehensive Review
标题: 音频和语音处理中的元学习:端到端全面评论
作者:Athul Raimon,Shubha Masti,Shyam K Sateesh,Siyani Vengatagiri,Bhaskarjyoti Das
备注:Survey Paper (15 pages, 1 figure)
链接:点击下载PDF文件
摘要:本调查概述了音频和语音处理场景中使用的各种元学习方法。元学习用于模型性能需要以最少的注释样本最大化的情况,使其适用于低样本音频处理。虽然该领域已经做出了一些重大贡献,音频元学习仍然缺乏全面的调查文件。我们提出了一个系统的审查元学习方法在音频处理。这包括关于数据增强,特征提取,预处理技术,元学习者,任务选择策略的音频特定讨论,还介绍了音频中的重要数据集,以及关键的现实用例。通过这一广泛的审查,我们的目标是提供有价值的见解,并确定未来的研究方向,在元学习和音频处理的交叉。摘要:This survey overviews various meta-learning approaches used in audio and speech processing scenarios. Meta-learning is used where model performance needs to be maximized with minimum annotated samples, making it suitable for low-sample audio processing. Although the field has made some significant contributions, audio meta-learning still lacks the presence of comprehensive survey papers. We present a systematic review of meta-learning methodologies in audio processing. This includes audio-specific discussions on data augmentation, feature extraction, preprocessing techniques, meta-learners, task selection strategies and also presents important datasets in audio, together with crucial real-world use cases. Through this extensive review, we aim to provide valuable insights and identify future research directions in the intersection of meta-learning and audio processing.

【18】 VyAnG-Net: A Novel Multi-Modal Sarcasm Recognition Model by Uncovering Visual, Acoustic and Glossary Features
标题: VyAnG-Net:一种通过揭示视觉、声学和词汇特征的新型多模式讽刺识别模型
作者:Ananya Pandey,Dinesh Kumar Vishwakarma
链接:点击下载PDF文件
摘要:各种语言和非语言线索,如过分强调一个词,语气的变化,或尴尬的表达,经常传达讽刺。会话中讽刺识别的计算机视觉问题旨在识别隐藏在日常对话中的讽刺,批评和隐喻信息。以前,讽刺识别主要集中在文本上。尽管如此,关键是要考虑所有的文本信息,音频流,面部表情和身体位置的可靠的讽刺识别。因此,我们提出了一种新的方法,它结合了一个轻量级的深度注意力模块与一个自我调节的ConvNet集中在视觉数据的最关键的功能和一个基于注意力标记的策略,以提取最关键的上下文特定的信息从文本数据。以下是我们的实验在执行多模态讽刺识别任务时所做的主要贡献的列表:注意力标记器分支,用于从字幕提供的词汇表内容中获得有益的特征;视觉分支,用于从视频帧中获取最突出的特征;从声学内容中的话语级特征提取和用于混合从多个模态获得的特征的基于多头注意的特征融合分支。在基准视频数据集之一MUSTaRD上进行的广泛测试表明,对于扬声器相关配置和扬声器独立配置,我们的方法优于现有方法,准确率为79.86%和76.94%。我们还进行了跨数据集分析,以测试VyAnG-Net与另一个数据集MUSTARD++的未知样本的适应性。摘要:Various linguistic and non-linguistic clues, such as excessive emphasis on a word, a shift in the tone of voice, or an awkward expression, frequently convey sarcasm. The computer vision problem of sarcasm recognition in conversation aims to identify hidden sarcastic, criticizing, and metaphorical information embedded in everyday dialogue. Prior, sarcasm recognition has focused mainly on text. Still, it is critical to consider all textual information, audio stream, facial expression, and body position for reliable sarcasm identification. Hence, we propose a novel approach that combines a lightweight depth attention module with a self-regulated ConvNet to concentrate on the most crucial features of visual data and an attentional tokenizer based strategy to extract the most critical context-specific information from the textual data. The following is a list of the key contributions that our experimentation has made in response to performing the task of Multi-modal Sarcasm Recognition: an attentional tokenizer branch to get beneficial features from the glossary content provided by the subtitles; a visual branch for acquiring the most prominent features from the video frames; an utterance-level feature extraction from acoustic content and a multi-headed attention based feature fusion branch to blend features obtained from multiple modalities. Extensive testing on one of the benchmark video datasets, MUSTaRD, yielded an accuracy of 79.86% for speaker dependent and 76.94% for speaker independent configuration demonstrating that our approach is superior to the existing methods. We have also conducted a cross-dataset analysis to test the adaptability of VyAnG-Net with unseen samples of another dataset MUStARD++.


机器翻译,仅供参考