今日论文合集:cs.SD语音21篇,eess.AS音频处理29篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音
【1】 Generative AI-based data augmentation for improved bioacoustic  classification in noisy environments
标题: 基于人工智能的生成性数据增强,以改进噪音环境中的生物声学分类
链接:https://arxiv.org/abs/2412.01530
作者: Anthony Gibbons,  Emma King,  Ian Donohue,  Andrew Parnell
备注:18 pages, 3 tables, 5 figures
摘要:1.获取数据以训练基于强大人工智能(AI)的物种分类模型可能具有挑战性,特别是对于稀有物种。数据增强可以通过增加训练数据的多样性来提高分类准确性,并且比专家标记的数据更便宜。然而,许多经典的基于图像的增强技术不适合音频频谱图。2.我们研究了两种生成AI模型作为数据增强工具来合成声谱图和补充音频数据:辅助分类器生成对抗网络(ACGAN)和去噪扩散概率模型(DDPM)。后者在生成的频谱图的真实性和所得分类任务的准确性方面表现得特别好。3.除了这些新的方法,我们提出了一个新的音频数据集640小时的鸟叫声从风力发电场在爱尔兰,其中约800个样本已被标记的专家。考虑到背景风和涡轮机噪音,风电场数据对于分类模型来说尤其具有挑战性。4.与高度自信的BirdNET预测相比,在真实数据和合成数据上训练分类模型的集合给出了92.6%的准确率(仅使用真实数据时为90.5%)。5.我们的方法可用于增强更多物种和其他土地利用类型的声学信号,并有可能使我们开发可靠的基于人工智能的稀有物种检测能力发生重大变化。我们的代码可以在https://github.com/gibbona1/ SpectrogramGenAI上找到。
摘要:1. Obtaining data to train robust artificial intelligence (AI)-based modelsfor species classification can be challenging, particularly for rare species.Data augmentation can boost classification accuracy by increasing the diversityof training data and is cheaper to obtain than expert-labelled data. However,many classic image-based augmentation techniques are not suitable for audiospectrograms. 2. We investigate two generative AI models as data augmentationtools to synthesise spectrograms and supplement audio data: AuxiliaryClassifier Generative Adversarial Networks (ACGAN) and Denoising DiffusionProbabilistic Models (DDPMs). The latter performed particularly well in termsof both realism of generated spectrograms and accuracy in a resultingclassification task. 3. Alongside these new approaches, we present a new audiodata set of 640 hours of bird calls from wind farm sites in Ireland,approximately 800 samples of which have been labelled by experts. Wind farmdata are particularly challenging for classification models given thebackground wind and turbine noise. 4. Training an ensemble of classificationmodels on real and synthetic data combined gave 92.6% accuracy (and 90.5% withjust the real data) when compared with highly confident BirdNET predictions. 5.Our approach can be used to augment acoustic signals for more species and otherland-use types, and has the potential to bring about a step-change in ourcapacity to develop reliable AI-based detection of rare species. Our code isavailable at https://github.com/gibbona1/ SpectrogramGenAI.

【2】 Reject Threshold Adaptation for Open-Set Model Attribution of Deepfake  Audio
标题: Deepfake音频开集模型属性的预设阈值自适应
链接:https://arxiv.org/abs/2412.01425
作者: Xinrui Yan,  Jiangyan Yi,  Jianhua Tao,  Yujie Chen,  Hao Gu,  Guanjun Li,  Junzuo Zhou,  Yong Ren,  Tao Xu
备注:Accepted by ISCSLP 2024
摘要:面向开放环境的Deepfake音频的开集模型属性是一个新兴的研究课题,旨在识别Deepfake音频的生成模型。大多数以前的工作需要手动设置未知类的拒绝阈值,以与预测的概率进行比较。然而,模型往往过度拟合训练实例,并生成过于自信的预测。此外,有效区分当前数据集中的未知类别的阈值可能不适合于识别另一数据分布中的已知和未知类别。为了解决这些问题,我们提出了一个新的框架,用于Deepfake音频的开集模型属性,并具有拒绝阈值自适应(ReTA)。具体地,重构误差学习模块通过将系统指纹的表示与对应于目标类别或随机选择的其他类别标签的标签组合来进行训练。该过程生成匹配和非匹配的重构样本,建立每类的重构误差分布,并为拒绝阈值计算模块奠定基础。拒绝阈值计算模块利用高斯概率估计来拟合匹配和非匹配重建误差的分布。然后,它通过概率最小化准则计算所有类别的自适应拒绝阈值。实验结果证明了ReTA在改善Deepfake音频的开集模型属性方面的有效性。
摘要:Open environment oriented open set model attribution of deepfake audio is anemerging research topic, aiming to identify the generation models of deepfakeaudio. Most previous work requires manually setting a rejection threshold forunknown classes to compare with predicted probabilities. However, models oftenoverfit training instances and generate overly confident predictions. Moreover,thresholds that effectively distinguish unknown categories in the currentdataset may not be suitable for identifying known and unknown categories inanother data distribution. To address the issues, we propose a novel frameworkfor open set model attribution of deepfake audio with rejection thresholdadaptation (ReTA). Specifically, the reconstruction error learning moduletrains by combining the representation of system fingerprints with labelscorresponding to either the target class or a randomly chosen other classlabel. This process generates matching and non-matching reconstructed samples,establishing the reconstruction error distributions for each class and layingthe foundation for the reject threshold calculation module. The rejectthreshold calculation module utilizes gaussian probability estimation to fitthe distributions of matching and non-matching reconstruction errors. It thencomputes adaptive reject thresholds for all classes through probabilityminimization criteria. The experimental results demonstrate the effectivenessof ReTA in improving the open set model attributes of deepfake audio.

【3】 OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
标题: OmniFlow:具有多模式整流流的任何对任何一代
链接:https://arxiv.org/abs/2412.01169
作者: Shufan Li,  Konstantinos Kallidromitis,  Akash Gokul,  Zichun Liao,  Yusuke Kato,  Kazuki Kozuka,  Aditya Grover
备注:12 pages, 14 figures
摘要:我们介绍OmniFlow,一种新的生成模型,专为任何到任何生成任务,如文本到图像,文本到音频,音频到图像合成。OmniFlow推进了文本到图像模型中使用的整流(RF)框架,以处理多种模态的联合分布。它在广泛的任务上优于以前的任何对任何模型,例如文本到图像和文本到音频合成。我们的工作提供了三个关键贡献:首先,我们将RF扩展到多模态设置,并引入了一种新的指导机制,使用户能够灵活地控制生成的输出中不同模态之间的对齐。其次,我们提出了一种新的架构,扩展了稳定扩散3的文本到图像的MMDiT架构,并使音频和文本生成。扩展模块可以单独进行有效的预训练,并与普通的文本到图像MMDiT合并进行微调。最后,我们对用于大规模音频和文本生成的整流流量Transformers的设计选择进行了全面的研究,为优化不同模态的性能提供了有价值的见解。该守则可在https://github.com/jacklishufan/OmniFlows上查阅。
摘要:We introduce OmniFlow, a novel generative model designed for any-to-anygeneration tasks such as text-to-image, text-to-audio, and audio-to-imagesynthesis. OmniFlow advances the rectified flow (RF) framework used intext-to-image models to handle the joint distribution of multiple modalities.It outperforms previous any-to-any models on a wide range of tasks, such astext-to-image and text-to-audio synthesis. Our work offers three keycontributions: First, we extend RF to a multi-modal setting and introduce anovel guidance mechanism, enabling users to flexibly control the alignmentbetween different modalities in the generated outputs. Second, we propose anovel architecture that extends the text-to-image MMDiT architecture of StableDiffusion 3 and enables audio and text generation. The extended modules can beefficiently pretrained individually and merged with the vanilla text-to-imageMMDiT for fine-tuning. Lastly, we conduct a comprehensive study on the designchoices of rectified flow transformers for large-scale audio and textgeneration, providing valuable insights into optimizing performance acrossdiverse modalities. The Code will be available athttps://github.com/jacklishufan/OmniFlows.

【4】 The Codec Language Model-based Zero-Shot Spontaneous Style TTS System  for CoVoC Challenge 2024
标题: 用于2024年CoVoC挑战赛的基于Codec语言模型的Zero-Shot自发风格TTC系统
链接:https://arxiv.org/abs/2412.01100
作者: Shuoyi Zhou,  Yixuan Zhou,  Weiqing Li,  Jun Chen,  Runchuan Ye,  Weihao Wu,  Zijian Lin,  Shun Lei,  Zhiyong Wu
备注:Accepted by ISCSLP 2024
摘要:本文描述了用于ISCSLP 2024会话语音克隆挑战赛(CoVoC)的zero-shot自发式TTS系统。我们提出了一个基于LLaMA的编解码器语言模型与延迟模式,以实现自发风格的语音克隆。为了提高语音可懂度,我们在语言模型中引入了无分类器指导(CFG)策略,以加强对标记预测的条件指导。为了生成高质量的话语,我们采用了有效的数据预处理操作,并微调我们的模型与选定的高质量的自发语音数据。在CoVoC约束的轨道上的官方评估表明,我们的系统达到了最好的语音自然度MOS为3.80,并获得了可观的语音质量和说话人相似度的结果。
摘要:This paper describes the zero-shot spontaneous style TTS system for theISCSLP 2024 Conversational Voice Clone Challenge (CoVoC). We propose aLLaMA-based codec language model with a delay pattern to achieve spontaneousstyle voice cloning. To improve speech intelligibility, we introduce theClassifier-Free Guidance (CFG) strategy in the language model to strengthenconditional guidance on token prediction. To generate high-quality utterances,we adopt effective data preprocessing operations and fine-tune our model withselected high-quality spontaneous speech data. The official evaluations in theCoVoC constrained track show that our system achieves the best speechnaturalness MOS of 3.80 and obtains considerable speech quality and speakersimilarity results.

【5】 FreeCodec: A disentangled neural speech codec with fewer tokens
标题: FreeCodec:一个具有更少令牌的分离神经语音编解码器
链接:https://arxiv.org/abs/2412.01053
作者: Youqiang Zheng,  Weiping Tu,  Yueteng Kang,  Jie Chen,  Yike Zhang,  Li Xiao,  Yuhong Yang,  Long Ma
摘要:神经语音编解码器以其出色的离散表征重构能力而受到广泛关注。   它是语音编码和大型语言模型(LLM)等生成任务的关键组成部分。   然而,大多数基于残差矢量量化的作品表现较差,由于编码效率低,建模复杂的耦合信息的令牌较少。   在本文中,我们提出了一种名为FreeCodec的神经语音编解码器,它通过将语音的内在属性分解为不同的分量来采用更有效的编码框架:   1)提取全局向量作为音色信息,   2)使用具有长步幅级别的韵律编码器来对韵律信息进行建模,   3)内容信息来自内容编码器。   使用不同的训练策略,FreeCodec在重建和解纠缠场景中实现了最先进的性能。   从主观和客观的实验结果表明,我们的框架优于现有的方法。
摘要:Neural speech codecs have gained great attention for their outstandingreconstruction with discrete token representations. It is a crucial component in generative tasks such as speech coding and largelanguage models (LLM). However, most works based on residual vector quantization perform worse withfewer tokens due to low coding efficiency for modeling complex coupledinformation. In this paper, we propose a neural speech codec named FreeCodec which employsa more effective encoding framework by decomposing intrinsic properties ofspeech into different components: 1) a global vector is extracted as the timbre information, 2) a prosody encoder with a long stride level is used to model the prosodyinformation, 3) the content information is from a content encoder. Using different training strategies, FreeCodec achieves state-of-the-artperformance in reconstruction and disentanglement scenarios. Results from subjective and objective experiments demonstrate that ourframework outperforms existing methods.

【6】 Complexity boosted adaptive training for better low resource ASR  performance
标题: 复杂性增强自适应训练,以获得更好的低资源ASB性能
链接:https://arxiv.org/abs/2412.00877
作者: Hongxuan Lu,  Shenjian Wang,  Biao Li
摘要:在ASR模型的整个训练过程中,数据增强的强度和计算训练损失的方法以基于预设参数的调节方式应用。例如,SpecAugment采用预定义的增强强度来掩蔽时频域频谱的部分。类似地,在基于CTC的多层模型中,通常基于训练过程期间编码器的最终层的输出来确定损失。然而,忽略动态特性可能使训练模型次优。为了解决这个问题,我们提出了一个两阶段的训练方法,称为复杂性增强自适应(CBA)训练。它涉及根据训练样本的复杂性对数据增强策略和CTC损失传播进行动态调整。在第一阶段,我们训练模型与中间CTC为基础的正则化和数据增强没有任何自适应的政策。在第二阶段,我们提出了一种新的自适应策略,称为MinMax-IBF,它计算样本的复杂度。我们将MinMax-IBF策略与数据增强和中间CTC损失正则化相结合,以继续训练。所提出的CBA训练方法显示出相当大的改进,在LibriSpeech 100 h测试清洁和测试其他数据集上WER的相对减少高达13.4%和14.1%,在AISHELL-1测试集上的相对减少也高达6.3%,超过Wenet中的Conformer架构。
摘要:During the entire training process of the ASR model, the intensity of dataaugmentation and the approach of calculating training loss are applied in aregulated manner based on preset parameters. For example, SpecAugment employs apredefined strength of augmentation to mask parts of the time-frequency domainspectrum. Similarly, in CTC-based multi-layer models, the loss is generallydetermined based on the output of the encoder's final layer during the trainingprocess. However, ignoring dynamic characteristics may suboptimally trainmodels. To address the issue, we present a two-stage training method, known ascomplexity-boosted adaptive (CBA) training. It involves making dynamicadjustments to data augmentation strategies and CTC loss propagation based onthe complexity of the training samples. In the first stage, we train the modelwith intermediate-CTC-based regularization and data augmentation without anyadaptive policy. In the second stage, we propose a novel adaptive policy,called MinMax-IBF, which calculates the complexity of samples. We combine theMinMax-IBF policy to data augmentation and intermediate CTC loss regularizationto continue training. The proposed CBA training approach shows considerableimprovements, up to 13.4% and 14.1% relative reduction in WER on theLibriSpeech 100h test-clean and test-other dataset and also up to 6.3% relativereduction on AISHELL-1 test set, over the Conformer architecture in Wenet.

【7】 A Comparative Study of LLM-based ASR and Whisper in Low Resource and  Code Switching Scenario
标题: 低资源和代码交换场景下基于LLM的ASB和Whisper的比较研究
链接:https://arxiv.org/abs/2412.00721
作者: Zheshu Song,  Ziyang Ma,  Yifan Yang,  Jianheng Zhuo,  Xie Chen
备注:4 pages
摘要:大型语言模型(LLM)在各种NLP任务中表现出卓越的性能,它们与语音编码器的集成正在迅速成为自动语音识别(ASR)领域的主导趋势。以前的工作主要集中在利用LLM的语音识别在英语和汉语。然而,它们在低资源环境中解决语音识别挑战的潜力仍然没有得到充分的探索。因此,在这项工作中,我们的目标是探索能力的LLM在低资源ASR和普通话-英语码切换ASR。我们还评估和比较基于LLM的ASR系统对Whisper模型的识别性能。大量的实验表明,基于LLM的ASR在低资源ASR中比Whisper模型产生了12.8%的相对增益,而Whisper在中英文代码切换ASR中表现更好。我们希望这项研究可以阐明ASR低资源的情况下。
摘要:Large Language Models (LLMs) have showcased exceptional performance acrossdiverse NLP tasks, and their integration with speech encoder is rapidlyemerging as a dominant trend in the Automatic Speech Recognition (ASR) field.Previous works mainly concentrated on leveraging LLMs for speech recognition inEnglish and Chinese. However, their potential for addressing speech recognitionchallenges in low resource settings remains underexplored. Hence, in this work,we aim to explore the capability of LLMs in low resource ASR andMandarin-English code switching ASR. We also evaluate and compare therecognition performance of LLM-based ASR systems against Whisper model.Extensive experiments demonstrate that LLM-based ASR yields a relative gain of12.8\% over the Whisper model in low resource ASR while Whisper performs betterin Mandarin-English code switching ASR. We hope that this study could shedlight on ASR for low resource scenarios.

【8】 Audio Atlas: Visualizing and Exploring Audio Datasets
标题: 音频地图集:可视化和探索音频数据集
链接:https://arxiv.org/abs/2412.00591
作者: Luca A. Lanzendörfer,  Florian Grötschla,  Uzeyir Valizada,  Roger Wattenhofer
备注:Extended Abstract at ISMIR 2024
摘要:我们介绍音频地图集,一个交互式的Web应用程序,用于可视化音频数据使用文本音频嵌入。Audio Atlas旨在使用对比嵌入模型和矢量数据库来促进音频数据集的探索和分析,以实现高效的数据管理和语义搜索。该系统将音频嵌入映射到二维空间中,并利用DeepScatter进行动态可视化。Audio Atlas专为可扩展性而设计,可以轻松集成新数据集,使用户能够更好地了解其音频数据并识别模式和异常值。我们开源了Audio Atlas的代码库,并提供了包含各种音频和音乐数据集的初始实现。
摘要:We introduce Audio Atlas, an interactive web application for visualizingaudio data using text-audio embeddings. Audio Atlas is designed to facilitatethe exploration and analysis of audio datasets using a contrastive embeddingmodel and a vector database for efficient data management and semantic search.The system maps audio embeddings into a two-dimensional space and leveragesDeepScatter for dynamic visualization. Designed for extensibility, Audio Atlasallows easy integration of new datasets, enabling users to better understandtheir audio data and identify both patterns and outliers. We open-source thecodebase of Audio Atlas, and provide an initial implementation containingvarious audio and music datasets.

【9】 From Audio Deepfake Detection to AI-Generated Music Detection -- A  Pathway and Overview
标题: 从音频Deepfake检测到人工智能生成的音乐检测--路径和概述
链接:https://arxiv.org/abs/2412.00571
作者: Yupei Li,  Manuel Milling,  Lucia Specia,  Björn W. Schuller
摘要:随着人工智能(AI)技术的不断发展,它们在生成逼真的、适合上下文的内容方面的应用已经扩展到各个领域。音乐是一种艺术形式和娱乐媒介,深深植根于人类文化,人工智能越来越多地参与其制作。然而,人工智能音乐生成(AIGM)工具的不受管制的使用引起了人们对音乐产业、版权和艺术完整性潜在负面影响的担忧,强调了有效检测AIGM的重要性。本文综述了现有的AIGM检测方法。为了为AIGM检测的一般工作和挑战奠定基础,我们首先回顾了AIGM的一般原理,包括deepfake音频的最新进展以及多模态检测技术。我们进一步提出了一种潜在的途径,用于利用从音频deepfake检测到AIGM检测的基础模型。此外,我们还讨论了这些工具的影响,并提出了未来研究的方向,以解决该领域正在面临的挑战。
摘要:As Artificial Intelligence (AI) technologies continue to evolve, their use ingenerating realistic, contextually appropriate content has expanded intovarious domains. Music, an art form and medium for entertainment, deeply rootedinto human culture, is seeing an increased involvement of AI into itsproduction. However, the unregulated use of AI music generation (AIGM) toolsraises concerns about potential negative impacts on the music industry,copyright and artistic integrity, underscoring the importance of effective AIGMdetection. This paper provides an overview of existing AIGM detection methods.To lay a foundation to the general workings and challenges of AIGM detection,we first review general principles of AIGM, including recent advancements indeepfake audios, as well as multimodal detection techniques. We further proposea potential pathway for leveraging foundation models from audio deepfakedetection to AIGM detection. Additionally, we discuss implications of thesetools and propose directions for future research to address ongoing challengesin the field.

【10】 Personal Sound Zones and Shielded Localized Communication through Active  Acoustic Control
标题: 通过主动声学控制实现个人声区和屏蔽本地通信
链接:https://arxiv.org/abs/2412.00456
作者: Neil Jerome A. Egarguin,  Daniel Onofrei
摘要:在本文中,我们提出了一个时域扩展我们的策略上操纵辐射标量亥姆霍兹场,并讨论了两个重要的应用场景,即(1)创建一个有界域内的个人声音区域和(2)屏蔽本地化通信。我们的策略是基于作者以前的工作建立的可能性和稳定性控制声场使用一个阵列的几乎非辐射耦合源,并提出了一个详细的傅立叶合成方法对时域效果。我们要求声源阵列在控制区域上产生所需的场,同时在更大的外接球体之外保持零场。本文回顾了主要的理论结果,然后提出了基本的傅立叶合成范式,并显示,通过相关的模拟,我们的策略的性能。
摘要:In this paper, we present a time domain extension of our strategy onmanipulating radiated scalar Helmholtz fields and discuss two important appliedscenarios, namely (1) creating personal sound zones inside a bounded domain and(2) shielded localized communication. Our strategy is based on the authors'previous works establishing the possibility and stability of controllingacoustic fields using an array of almost non-radiating coupling sources andpresents a detailed Fourier synthesis approach towards a time-domain effect. Werequire that the array of acoustic sources creates the desired fields on thecontrol regions while maintaining a zero field beyond a larger circumscribedsphere. This paper recalls the main theoretical results then presents theunderlying Fourier synthesis paradigm and show, through relevant simulations,the performance of our strategy.

【11】 Sample adaptive data augmentation with progressive scheduling
标题: 采用渐进式调度的自适应数据增强示例
链接:https://arxiv.org/abs/2412.00415
作者: Hongxuan Lu,  Biao Li
摘要:数据增强是一种被广泛采用的提高自动语音识别(ASR)鲁棒性的技术。对所有训练数据采用固定的数据增强策略是一种常见的做法。然而,重要的是要注意,在单个训练批次内的不同样本之间,可能存在诸如背景噪声、语音速率等因素的变化。通过使用固定增广策略,存在模型可能达到次优状态的风险。除了采用固定增强策略的风险之外,模型的能力在不同的训练阶段可能会有所不同。为了解决这些问题,本文提出了样本自适应数据增强渐进调度(PS-SapAug)的方法。所提出的方法在两阶段训练方法中应用动态数据增强。它采用混合归一化来计算基于每个样本的损失的样本特定的增强参数。此外,增强的概率在整个训练过程中逐渐增加。我们的方法在流行的ASR基准数据集上进行了评估,包括Aishell-1和LibriSpeech-100 h,在LibriSpeech-100 h测试-clean上实现了高达8.13%的WER减少,在测试-其他上实现了6.23%,在AISHELL-1测试集上实现了5.26%,这证明了我们的方法在提高性能和最小化错误方面的有效性。
摘要:Data augmentation is a widely adopted technique utilized to improve therobustness of automatic speech recognition (ASR). Employing a fixed dataaugmentation strategy for all training data is a common practice. However, itis important to note that there can be variations in factors such as backgroundnoise, speech rate, etc. among different samples within a single trainingbatch. By using a fixed augmentation strategy, there is a risk that the modelmay reach a suboptimal state. In addition to the risks of employing a fixedaugmentation strategy, the model's capabilities may differ across varioustraining stages. To address these issues, this paper proposes the method ofsample-adaptive data augmentation with progressive scheduling(PS-SapAug). Theproposed method applies dynamic data augmentation in a two-stage trainingapproach. It employs hybrid normalization to compute sample-specificaugmentation parameters based on each sample's loss. Additionally, theprobability of augmentation gradually increases throughout the trainingprogression. Our method is evaluated on popular ASR benchmark datasets,including Aishell-1 and Librispeech-100h, achieving up to 8.13% WER reductionon LibriSpeech-100h test-clean, 6.23% on test-other, and 5.26% on AISHELL-1test set, which demonstrate the efficacy of our approach enhancing performanceand minimizing errors.

【12】 MusicGen-Chord: Advancing Music Generation through Chord Progressions  and Interactive Web-UI
标题: MusicGen-Chord:通过Chord进展和交互式Web UI推进音乐生成
链接:https://arxiv.org/abs/2412.00325
作者: Jongmin Jung,  Andreas Jansson,  Dasaem Jeong
备注:Late-breaking/demo (LBD) at ISMIR 2024. this https URL
摘要:MusicGen是一种音乐生成语言模型(LM),可以根据文本描述和旋律特征进行调节。我们介绍MusicGen和弦,它扩展了这种能力,将和弦进行功能。该模型将独热编码的旋律色度向量修改为多热编码的和弦色度向量,使得能够生成既反映和弦进行又反映文本描述的音乐。此外,我们开发了MusicGen-Remixer,这是一个利用MusicGen-Chord生成基于文本描述的输入音乐混音的应用程序。这两种模型都使用cog集成到Replicate的Web UI中,促进了广泛的可访问性和用户友好的可控交互,以创建和体验AI生成的音乐。
摘要:MusicGen is a music generation language model (LM) that can be conditioned ontextual descriptions and melodic features. We introduce MusicGen-Chord, whichextends this capability by incorporating chord progression features. This modelmodifies one-hot encoded melody chroma vectors into multi-hot encoded chordchroma vectors, enabling the generation of music that reflects both chordprogressions and textual descriptions. Furthermore, we developedMusicGen-Remixer, an application utilizing MusicGen-Chord to generate remixesof input music conditioned on textual descriptions. Both models are integratedinto Replicate's web-UI using cog, facilitating broad accessibility anduser-friendly controllable interaction for creating and experiencingAI-generated music.

【13】 Improving speaker verification robustness with synthetic emotional  utterances
标题: 利用合成情感话语提高说话者验证稳健性
链接:https://arxiv.org/abs/2412.00319
作者: Nikhil Kumar Koditala,  Chelsea Jui-Ting Ju,  Ruirui Li,  Minho Jin,  Aman Chadha,  Andreas Stolcke
摘要:说话人验证(SV)系统提供了一种验证服务,用于确认给定的语音样本是否来自特定的说话人。这项技术为满足个人偏好的各种个性化应用铺平了道路。SV系统面临的一个值得注意的挑战是它们在一系列情绪谱中一致表现的能力。大多数现有的模型表现出较高的错误率时,处理情绪的话语相比,中性的。因此,这种现象往往导致错过感兴趣的演讲。这个问题主要源于有限的可用性标记的情绪语音数据,阻碍了发展强大的扬声器表示,包括不同的情绪状态。   为了解决这个问题,我们提出了一种新的方法,采用CycleGAN框架作为数据增强方法。该技术为每个特定的说话者合成情感语音片段,同时保留独特的声音身份。我们的实验结果强调了将合成情感数据纳入训练过程的有效性。使用这个增强数据集训练的模型在情感语音场景中验证说话者的任务上始终优于基线模型,相对降低了3.64%的等错误率。
摘要:A speaker verification (SV) system offers an authentication service designedto confirm whether a given speech sample originates from a specific speaker.This technology has paved the way for various personalized applications thatcater to individual preferences. A noteworthy challenge faced by SV systems istheir ability to perform consistently across a range of emotional spectra. Mostexisting models exhibit high error rates when dealing with emotional utterancescompared to neutral ones. Consequently, this phenomenon often leads to missingout on speech of interest. This issue primarily stems from the limitedavailability of labeled emotional speech data, impeding the development ofrobust speaker representations that encompass diverse emotional states. To address this concern, we propose a novel approach employing the CycleGANframework to serve as a data augmentation method. This technique synthesizesemotional speech segments for each specific speaker while preserving the uniquevocal identity. Our experimental findings underscore the effectiveness ofincorporating synthetic emotional data into the training process. The modelstrained using this augmented dataset consistently outperform the baselinemodels on the task of verifying speakers in emotional speech scenarios,reducing equal error rate by as much as 3.64% relative.

【14】 Raw Audio Classification with Cosine Convolutional Neural Network  (CosCovNN)
标题: 使用CosCovNN卷积神经网络(CosCovNN)进行原始音频分类
链接:https://arxiv.org/abs/2412.00312
作者: Kazi Nazmul Haque,  Rajib Rana,  Tasnim Jarin,  Bjorn W. Schuller Jr
摘要:这项研究探索了使用卷积神经网络(CNN)从原始波形中进行音频分类的领域,这种方法消除了在预处理步骤中提取专业特征的需要。与文献中的最新趋势不同,文献中的趋势通常侧重于仅为CNN的初始层设计前端或滤波器,我们的研究引入了余弦卷积神经网络(CosCovNN),用余弦滤波器取代传统的CNN滤波器。CosCovNN超过了等效CNN架构的准确性,参数减少了约77\%。我们的研究进一步发展了一个名为矢量量化余弦卷积神经网络与记忆(VQCCM)的增强CosCovNN,结合了记忆和矢量量化层VQCCM实现了最先进的(SOTA)性能在五个不同的数据集与现有文献相比。我们的研究结果表明,余弦滤波器可以大大提高CNN在原始音频分类中的效率和准确性。
摘要:This study explores the field of audio classification from raw waveform usingConvolutional Neural Networks (CNNs), a method that eliminates the need forextracting specialised features in the pre-processing step. Unlike recenttrends in literature, which often focuses on designing frontends or filters foronly the initial layers of CNNs, our research introduces the CosineConvolutional Neural Network (CosCovNN) replacing the traditional CNN filterswith Cosine filters. The CosCovNN surpasses the accuracy of the equivalent CNNarchitectures with approximately $77\%$ less parameters. Our research furtherprogresses with the development of an augmented CosCovNN named Vector QuantisedCosine Convolutional Neural Network with Memory (VQCCM), incorporating a memoryand vector quantisation layer VQCCM achieves state-of-the-art (SOTA)performance across five different datasets in comparison with existingliterature. Our findings show that cosine filters can greatly improve theefficiency and accuracy of CNNs in raw audio classification.

【15】 Circumventing shortcuts in audio-visual deepfake detection datasets with  unsupervised learning
标题: 通过无监督学习避免视听深度伪造检测数据集中的捷径
链接:https://arxiv.org/abs/2412.00175
作者: Dragos-Alexandru Boldisor,  Stefan Smeu,  Dan Oneata,  Elisabeta Oneata
摘要:良好的数据集对于开发和基准测试任何机器学习系统都至关重要。对于深度造假检测(本文的重点)等安全关键应用程序来说,它们的重要性更加极端。在这里,我们揭示了两个最广泛使用的音频视频深度伪造数据集存在一个之前未识别的虚假功能:主导沉默。假视频从一个非常短暂的沉默开始,仅仅基于这个功能,我们就可以几乎完美地分离真实和假样本。因此,先前的仅音频模型和音频-视频模型利用假视频中存在的静音,因此当去除前导静音时表现更差。为了避免锁定这种不需要的工件和可能其他未揭示的,我们提出了从监督到无监督学习的转变,通过专门在真实数据上训练模型。我们表明,通过调整自监督的音频-视频表示,我们消除了依赖特定于网络的偏见的风险,并提高了deepfake检测的鲁棒性。
摘要:Good datasets are essential for developing and benchmarking any machinelearning system. Their importance is even more extreme for safety criticalapplications such as deepfake detection - the focus of this paper. Here wereveal that two of the most widely used audio-video deepfake datasets sufferfrom a previously unidentified spurious feature: the leading silence. Fakevideos start with a very brief moment of silence and based on this featurealone, we can separate the real and fake samples almost perfectly. As such,previous audio-only and audio-video models exploit the presence of silence inthe fake videos and consequently perform worse when the leading silence isremoved. To circumvent latching on such unwanted artifact and possibly otherunrevealed ones we propose a shift from supervised to unsupervised learning bytraining models exclusively on real data. We show that by aligningself-supervised audio-video representations we remove the risk of relying ondataset-specific biases and improve robustness in deepfake detection.

【16】 A Survey of Recent Advances and Challenges in Deep Audio-Visual  Correlation Learning
标题: 深度视听相关学习的最新进展和挑战概览
链接:https://arxiv.org/abs/2412.00049
作者: Luis Vilaca,  Yi Yu,  Paula Vinan
备注:arXiv admin note: text overlap with arXiv:2202.13673
摘要:视听相关学习旨在捕捉和理解视听数据之间的自然现象。深度学习的快速增长推动了处理视听数据的提案的发展,并且可以在过去几年的提案数量中观察到。从而鼓励开展全面调查。除了分析在这种情况下使用的模型,我们还讨论了一些任务的定义和范式应用于人工智能多媒体。此外,我们研究了经常使用的目标函数,并讨论了如何在优化过程中利用视听数据,即,在视听领域中表示知识的不同方法。事实上,我们关注的是人类可以理解的机制,即,结构化知识是可理解性知识的反映,能够指导学习过程。最重要的是,我们总结了视听相关学习(AVCL)的最新进展,并讨论了未来的研究方向。
摘要:Audio-visual correlation learning aims to capture and understand naturalphenomena between audio and visual data. The rapid growth of Deep Learningpropelled the development of proposals that process audio-visual data and canbe observed in the number of proposals in the past years. Thus encouraging thedevelopment of a comprehensive survey. Besides analyzing the models used inthis context, we also discuss some tasks of definition and paradigm applied inAI multimedia. In addition, we investigate objective functions frequently usedand discuss how audio-visual data is exploited in the optimization process,i.e., the different methodologies for representing knowledge in theaudio-visual domain. In fact, we focus on how human-understandable mechanisms,i.e., structured knowledge that reflects comprehensible knowledge, can guidethe learning process. Most importantly, we provide a summarization of therecent progress of Audio-Visual Correlation Learning (AVCL) and discuss thefuture research directions.

【17】 Linear stimulus reconstruction works on the KU Leuven audiovisual,  gaze-controlled auditory attention decoding dataset
标题: 线性刺激重建适用于KU Leuven视听、凝视控制的听觉注意力解码数据集
链接:https://arxiv.org/abs/2412.01401
作者: Simon Geirnaert,  Iustina Rotaru,  Tom Francart,  Alexander Bertrand
摘要:在最近的一篇论文中,我们介绍了KU Leuven视听,凝视控制听觉注意力解码(AV-GC-AAD)数据集,其中我们记录了参与者在各种视听条件下注意两个竞争扬声器中的一个的脑电图(EEG)信号。该数据集的主要目标是将凝视方向与听觉注意力方向分开,以揭示现有空间AAD算法中与凝视相关的捷径,这些算法旨在直接从EEG解码听觉注意力(方向)。基于空间AAD的各种方法在我们的AV-GC-AAD数据集上没有实现显著的机会以上性能,这表明先前报道的结果主要是由现有数据集中的眼睛注视混淆驱动的。尽管如此,这些不利的结果往往被丢弃的原因是归因于AV-GC-AAD数据集的局限性,如有限的数据量训练工作模型,太多的数据异质性,由于不同的视听条件,或参与者据称无法集中他们的听觉注意力在复杂的指令。在本文中,我们提出了线性刺激重建AAD算法的结果,并表明,高AAD精度可以在每个单独的条件下,该模型概括了跨条件,跨新的主题,甚至跨数据集。因此,我们消除了任何疑问,即AV-GC-AAD数据集的不足是(空间)AAD算法与其他数据集相比未能实现上述机会性能的主要原因。此外,该报告提供了一个简单的基线评估程序(包括源代码),可以作为在此数据集上评估的所有未来AAD算法的最低基准。
摘要:In a recent paper, we presented the KU Leuven audiovisual, gaze-controlledauditory attention decoding (AV-GC-AAD) dataset, in which we recordedelectroencephalography (EEG) signals of participants attending to one out oftwo competing speakers under various audiovisual conditions. The main goal ofthis dataset was to disentangle the direction of gaze from the direction ofauditory attention, in order to reveal gaze-related shortcuts in existingspatial AAD algorithms that aim to decode the (direction of) auditory attentiondirectly from the EEG. Various methods based on spatial AAD do not achievesignificant above-chance performances on our AV-GC-AAD dataset, indicating thatpreviously reported results were mainly driven by eye gaze confounds inexisting datasets. Still, these adverse outcomes are often discarded forreasons that are attributed to the limitations of the AV-GC-AAD dataset, suchas the limited amount of data to train a working model, too much dataheterogeneity due to different audiovisual conditions, or participantsallegedly being unable to focus their auditory attention under the complexinstructions. In this paper, we present the results of the linear stimulusreconstruction AAD algorithm and show that high AAD accuracy can be obtainedwithin each individual condition and that the model generalizes acrossconditions, across new subjects, and even across datasets. Therefore, weeliminate any doubts that the inadequacy of the AV-GC-AAD dataset is theprimary reason for the (spatial) AAD algorithms failing to achieve above-chanceperformance when compared to other datasets. Furthermore, this report providesa simple baseline evaluation procedure (including source code) that can serveas the minimal benchmark for all future AAD algorithms evaluated on thisdataset.

【18】 Memory-Efficient Training for Deep Speaker Embedding Learning in Speaker  Verification
标题: 说话人验证中深度说话人嵌入学习的记忆高效训练
链接:https://arxiv.org/abs/2412.01195
作者: Bei Liu,  Yanmin Qian
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:最近的说话人确认(SV)系统已经显示出采用更深的说话人嵌入提取器的趋势。尽管更深、更大的神经网络可以显着提高性能,但其大量的内存需求阻碍了在消费者GPU上的训练。在本文中,我们探索了一种用于资源受限场景中的深度说话人嵌入学习的记忆高效训练策略。首先,我们对SV系统训练过程中的GPU内存分配进行了系统的分析。经验观察表明,激活和优化器状态是内存消耗的主要来源。对于激活,我们设计了两种类型的可逆神经网络,它们消除了在反向传播过程中存储中间激活的需要,从而在不损失性能的情况下显着减少了内存使用。对于优化器状态,我们引入了一种动态量化方法,该方法将原始的32位浮点值替换为基于动态树的8位数据类型。VoxCeleb上的实验结果表明,ResNets和DF-ResNets的可逆变体可以执行训练,而无需在GPU内存中缓存激活。此外,与32位版本相比,SGD和Adam的8位版本在保持性能的同时节省了75%的内存成本。最后,内存使用和性能的详细比较表明,我们提出的模型实现了高达16.2倍的内存节省,与香草系统相比,几乎相同的参数和性能。与之前需要多个高端GPU(如A100)不同,我们只需一两个消费级2080 Ti GPU就可以有效地训练深度扬声器嵌入提取器。
摘要:Recent speaker verification (SV) systems have shown a trend toward adoptingdeeper speaker embedding extractors. Although deeper and larger neural networkscan significantly improve performance, their substantial memory requirementshinder training on consumer GPUs. In this paper, we explore a memory-efficienttraining strategy for deep speaker embedding learning in resource-constrainedscenarios. Firstly, we conduct a systematic analysis of GPU memory allocationduring SV system training. Empirical observations show that activations andoptimizer states are the main sources of memory consumption. For activations,we design two types of reversible neural networks which eliminate the need tostore intermediate activations during back-propagation, thereby significantlyreducing memory usage without performance loss. For optimizer states, weintroduce a dynamic quantization approach that replaces the original 32-bitfloating-point values with a dynamic tree-based 8-bit data type. Experimentalresults on VoxCeleb demonstrate that the reversible variants of ResNets andDF-ResNets can perform training without the need to cache activations in GPUmemory. In addition, the 8-bit versions of SGD and Adam save 75% of memorycosts while maintaining performance compared to their 32-bit counterparts.Finally, a detailed comparison of memory usage and performance indicates thatour proposed models achieve up to 16.2x memory savings, with nearly identicalparameters and performance compared to the vanilla systems. In contrast to theprevious need for multiple high-end GPUs such as the A100, we can effectivelytrain deep speaker embedding extractors with just one or two consumer-level2080Ti GPUs.

【19】 Deep Learning-Based Approach for Identification and Compensation of  Nonlinear Distortions in Parametric Array Loudspeakers
标题: 基于深度学习的参数阵列扬声器非线性失真识别和补偿方法
链接:https://arxiv.org/abs/2412.01092
作者: Mengtong Li,  Tao Zhuang,  Kai Chen,  Jia-Xin Zhong,  Jing Lu
备注:5 pages, 7 figures
摘要:与传统的电动扬声器相比,参量阵列扬声器(PAL)为音频应用提供了卓越的方向性,但由于其固有的复杂解调过程而遭受显著的非线性失真。基于Volterra滤波器的方法已被广泛用于减少这些失真,但其有效性受到其逆滤波器的能力的限制。具体而言,其p阶逆滤波器只能补偿最高p阶的非线性,而它引入的高阶非线性会继续生成低阶谐波。相比之下,本文首次引入现代深度学习方法来解决PAL系统的非线性识别和补偿。具体而言,WaveNet神经网络的前馈变体,其在音频非线性系统建模中的成功被认可,用于识别和补偿基于双边带幅度调制的PAL系统中的失真。从250 Hz到8 kHz的实验测量表明,我们提出的方法显着降低了总谐波失真和互调失真的PAL产生的音频声音,实现平均减少到4.55%和2.47%,分别。该性能明显优于使用当前最先进的基于Volterra滤波器的方法获得的结果。我们的工作为改善PAL的声音再现性能开辟了新的可能性。
摘要:Compared to traditional electrodynamic loudspeakers, the parametric arrayloudspeaker (PAL) offers exceptional directivity for audio applications butsuffers from significant nonlinear distortions due to its inherent intricatedemodulation process. The Volterra filter-based approaches have been widelyused to reduce these distortions, but the effectiveness is limited by itsinverse filter's capability. Specifically, its pth-order inverse filter canonly compensate for nonlinearities up to the pth order, while the higher-ordernonlinearities it introduces continue to generate lower-order harmonics. Incontrast, this paper introduces the modern deep learning methods for the firsttime to address nonlinear identification and compensation for PAL systems.Specifically, a feedforward variant of the WaveNet neural network, recognizedfor its success in audio nonlinear system modeling, is utilized to identify andcompensate for distortions in a double sideband amplitude modulation-based PALsystem. Experimental measurements from 250 Hz to 8 kHz demonstrate that ourproposed approach significantly reduces both total harmonic distortion andintermodulation distortion of audio sound generated by PALs, achieving averagereductions to 4.55% and 2.47%, respectively. This performance is notablysuperior to results obtained using the current state-of-the-art Volterrafilter-based methods. Our work opens new possibilities for improving the soundreproduction performance of PALs.

【20】 Feasibility of Mental Health Triage Call Priority Prediction Using  Machine Learning
标题: 使用机器学习进行心理健康分类呼叫优先级预测的可行性
链接:https://arxiv.org/abs/2412.00057
作者: Rajib Rana,  Niall Higgins,  Kazi Nazmul Haque,  John Reilly,  Kylie Burke,  Kathryn Turner,  Terry Stedman
摘要:确保准确的呼叫优先级对于优化心理健康诊所的效率和响应能力至关重要。目前,呼叫操作员完全依赖于呼叫者的陈述来确定呼叫的优先级。事实证明,完全主观的评估可能会导致错误。此外,如果在通话过程中不利用语音属性来帮助评估,则会错失机会。不正确的优先顺序可能会导致对高风险个人的延迟援助,资源分配不当,心理健康恶化加剧,失去信任以及潜在的法律后果。必须解决这些风险,以保证精神卫生服务的可靠性和有效性。这项研究深入研究了使用机器学习(人工智能的一个分支)的潜力,以估计呼叫优先级从呼叫者的声音为用户的心理健康电话号码。在分析了459个电话记录从心理健康诊所,我们达到了92%的平衡准确率,显示出承诺,帮助呼叫运营商的效率在呼叫处理过程中,提高客户满意度。
摘要:Ensuring accurate call prioritisation is essential for optimising theefficiency and responsiveness of mental health helplines. Currently, calloperators rely entirely on the caller's statements to determine the priority ofthe calls. It has been shown that entirely subjective assessment can lead toerrors. Furthermore, it is a missed opportunity not to utilise the voiceproperties readily available during the call to aid in the evaluation.Incorrect prioritisation can result in delayed assistance for high-riskindividuals, resource misallocation, increased mental health deterioration,loss of trust, and potential legal consequences. It is vital to address theserisks to guarantee the reliability and effectiveness of mental health services.This study delves into the potential of using machine learning, a branch ofArtificial Intelligence, to estimate call priority from the callers' voices forusers of mental health phone helplines. After analysing 459 call records from amental health helpline, we achieved a balanced accuracy of 92\%, showingpromise in aiding the call operators' efficiency in call handling processes andimproving customer satisfaction.

【21】 High-precision medical speech recognition through synthetic data and  semantic correction: UNITED-MEDASR
标题: 通过合成数据和语义纠正实现高精度医学语音识别:UNITED-MEDASS
链接:https://arxiv.org/abs/2412.00055
作者: Sourav Banerjee,  Ayushi Agarwal,  Promila Ghosh
备注:15 pages
摘要:临床领域的自动语音识别(ASR)系统面临着重大挑战,特别是需要准确识别专业医学词汇并满足严格的精度要求。我们介绍了United-MedASR,一种新的架构,通过集成合成数据生成,精确的ASR微调和高级语义增强技术来解决这些挑战。United-MedASR通过综合权威来源的数据构建专业医学词汇,如ICD-10(国际疾病分类,第10次修订),MIMS(医学专业每月索引)和FDA数据库。这种丰富的词汇有助于微调Whisper ASR模型,以更好地满足临床需求。为了提高处理速度,我们采用了Faster Whisper,确保精简和高速ASR性能。此外,我们采用定制的基于BART的语义增强器来处理复杂的医学术语,从而有效地提高准确性。我们的分层方法在ASR性能方面建立了新的基准,在LibriSpeech测试中实现了0.985%的单词错误率(WER),在Europarl-ASR EN Guest-test中达到了0.26%,并在Tedarp(0.29%WER)和FLEURS(0.336%WER)上表现出了稳健的性能。此外,我们提出了一个适应性强的架构,可以在不同的域复制,使其成为一个通用的解决方案,特定于域的ASR系统。
摘要:Automatic Speech Recognition (ASR) systems in the clinical domain facesignificant challenges, notably the need to recognise specialised medicalvocabulary accurately and meet stringent precision requirements. We introduceUnited-MedASR, a novel architecture that addresses these challenges byintegrating synthetic data generation, precision ASR fine-tuning, and advancedsemantic enhancement techniques. United-MedASR constructs a specialised medicalvocabulary by synthesising data from authoritative sources such as ICD-10(International Classification of Diseases, 10th Revision), MIMS (Monthly Indexof Medical Specialties), and FDA databases. This enriched vocabulary helpsfinetune the Whisper ASR model to better cater to clinical needs. To enhanceprocessing speed, we incorporate Faster Whisper, ensuring streamlined andhigh-speed ASR performance. Additionally, we employ a customised BART-basedsemantic enhancer to handle intricate medical terminology, thereby increasingaccuracy efficiently. Our layered approach establishes new benchmarks in ASRperformance, achieving a Word Error Rate (WER) of 0.985% on LibriSpeechtest-clean, 0.26% on Europarl-ASR EN Guest-test, and demonstrating robustperformance on Tedlium (0.29% WER) and FLEURS (0.336% WER). Furthermore, wepresent an adaptable architecture that can be replicated across differentdomains, making it a versatile solution for domain-specific ASR systems.

eess.AS音频处理

【1】 TACO: Training-free Sound Prompted Segmentation via Deep Audio-visual  CO-factorization
标题: TACO:通过深度视听协同分解的免训练声音预编码
链接:https://arxiv.org/abs/2412.01488
作者: Hugo Malard,  Michel Olvera,  Stephane Lathuiliere,  Slim Essid
摘要:大规模预训练的音频和图像模型表现出前所未有的泛化程度,使其适用于广泛的应用。在这里,我们处理声音提示分割的特定任务,旨在分割与音频信号中听到的对象对应的图像区域。大多数现有方法通过微调预训练模型或专门针对任务训练额外模块来解决这个问题。我们采取了不同的策略:我们引入了一种无需训练的方法,该方法利用非负矩阵分解(NMF)来对来自预训练模型的音频和视觉特征进行共分解,以揭示共享的可解释概念。这些概念被传递到一个开放词汇分割模型,以获得精确的分割图。通过使用冻结的预训练模型,我们的方法实现了高度的泛化,并在无监督的声音提示分割中建立了最先进的性能,显着超过了以前的无监督方法。
摘要:Large-scale pre-trained audio and image models demonstrate an unprecedenteddegree of generalization, making them suitable for a wide range ofapplications. Here, we tackle the specific task of sound-prompted segmentation,aiming to segment image regions corresponding to objects heard in an audiosignal. Most existing approaches tackle this problem by fine-tuning pre-trainedmodels or by training additional modules specifically for the task. We adopt adifferent strategy: we introduce a training-free approach that leveragesNon-negative Matrix Factorization (NMF) to co-factorize audio and visualfeatures from pre-trained models to reveal shared interpretable concepts. Theseconcepts are passed to an open-vocabulary segmentation model for precisesegmentation maps. By using frozen pre-trained models, our method achieves highgeneralization and establishes state-of-the-art performance in unsupervisedsound-prompted segmentation, significantly surpassing previous unsupervisedmethods.

【2】 Linear stimulus reconstruction works on the KU Leuven audiovisual,  gaze-controlled auditory attention decoding dataset
标题: 线性刺激重建适用于KU Leuven视听、凝视控制的听觉注意力解码数据集
链接:https://arxiv.org/abs/2412.01401
作者: Simon Geirnaert,  Iustina Rotaru,  Tom Francart,  Alexander Bertrand
摘要:在最近的一篇论文中,我们介绍了KU Leuven视听,凝视控制听觉注意力解码(AV-GC-AAD)数据集,其中我们记录了参与者在各种视听条件下注意两个竞争扬声器中的一个的脑电图(EEG)信号。该数据集的主要目标是从听觉注意力的方向中解开凝视的方向,以揭示现有空间AAD算法中与凝视相关的快捷方式,该算法旨在直接从EEG解码听觉注意力(方向)。基于空间AAD的各种方法在我们的AV-GC-AAD数据集上没有实现显著的机会以上性能,这表明先前报道的结果主要是由现有数据集中的眼睛注视混淆驱动的。尽管如此,这些不利的结果往往被丢弃的原因是归因于AV-GC-AAD数据集的局限性,如有限的数据量训练工作模型,太多的数据异质性,由于不同的视听条件,或参与者据称无法集中他们的听觉注意力在复杂的指令。在本文中,我们提出了线性刺激重建AAD算法的结果,并表明,高AAD精度可以在每个单独的条件下,该模型概括了跨条件,跨新的主题,甚至跨数据集。因此,我们消除了任何疑问,即AV-GC-AAD数据集的不足是(空间)AAD算法与其他数据集相比未能实现上述机会性能的主要原因。此外,该报告提供了一个简单的基线评估程序(包括源代码),可以作为在此数据集上评估的所有未来AAD算法的最低基准。
摘要:In a recent paper, we presented the KU Leuven audiovisual, gaze-controlledauditory attention decoding (AV-GC-AAD) dataset, in which we recordedelectroencephalography (EEG) signals of participants attending to one out oftwo competing speakers under various audiovisual conditions. The main goal ofthis dataset was to disentangle the direction of gaze from the direction ofauditory attention, in order to reveal gaze-related shortcuts in existingspatial AAD algorithms that aim to decode the (direction of) auditory attentiondirectly from the EEG. Various methods based on spatial AAD do not achievesignificant above-chance performances on our AV-GC-AAD dataset, indicating thatpreviously reported results were mainly driven by eye gaze confounds inexisting datasets. Still, these adverse outcomes are often discarded forreasons that are attributed to the limitations of the AV-GC-AAD dataset, suchas the limited amount of data to train a working model, too much dataheterogeneity due to different audiovisual conditions, or participantsallegedly being unable to focus their auditory attention under the complexinstructions. In this paper, we present the results of the linear stimulusreconstruction AAD algorithm and show that high AAD accuracy can be obtainedwithin each individual condition and that the model generalizes acrossconditions, across new subjects, and even across datasets. Therefore, weeliminate any doubts that the inadequacy of the AV-GC-AAD dataset is theprimary reason for the (spatial) AAD algorithms failing to achieve above-chanceperformance when compared to other datasets. Furthermore, this report providesa simple baseline evaluation procedure (including source code) that can serveas the minimal benchmark for all future AAD algorithms evaluated on thisdataset.

【3】 Text-based Audio Retrieval by Learning from Similarities between Audio  Captions
标题: 通过学习音频标题之间的相似性来实现基于文本的音频检索
链接:https://arxiv.org/abs/2412.01356
作者: Huang Xie,  Khazar Khorrami,  Okko Räsänen,  Tuomas Virtanen
摘要:本文提出了使用音频字幕的相似性来估计音频字幕的相关性,用于训练基于文本的音频检索系统。当前音频字幕数据集(例如,Clotho)包含与注释字幕配对的音频样本,但除了注释的音频样本和字幕之外,缺乏有关音频样本和字幕的相关信息。此外,主流方法(例如,CLAP)通常将注释对视为肯定,并将所有其他音频-字幕组合视为否定,假设音频样本和字幕之间的二元相关性。为了推断音频样本和任意字幕之间的相关性,我们提出了一种基于音频字幕的文本相似性来计算非二进制音频字幕相关性分数的方法。我们通过计算它们的Sentence-BERT嵌入的余弦相似性来测量音频字幕的文本相似性,然后使用逻辑函数将这些相似性转换为音频字幕相关性分数,从而通过它们的注释字幕将音频样本链接到数据集中的所有其他字幕。为了将计算出的相关性整合到训练中,我们采用了一个列表式排名目标,其中相关性得分被转换为给定文本查询的音频样本排名的概率。我们展示了所提出的方法的有效性,通过展示改进基于文本的音频检索相比,使用二进制音频字幕相关性的训练方法。
摘要:This paper proposes to use similarities of audio captions for estimatingaudio-caption relevances to be used for training text-based audio retrievalsystems. Current audio-caption datasets (e.g., Clotho) contain audio samplespaired with annotated captions, but lack relevance information about audiosamples and captions beyond the annotated ones. Besides, mainstream approaches(e.g., CLAP) usually treat the annotated pairs as positives and consider allother audio-caption combinations as negatives, assuming a binary relevancebetween audio samples and captions. To infer the relevance between audiosamples and arbitrary captions, we propose a method that computes non-binaryaudio-caption relevance scores based on the textual similarities of audiocaptions. We measure textual similarities of audio captions by calculating thecosine similarity of their Sentence-BERT embeddings and then transform thesesimilarities into audio-caption relevance scores using a logistic function,thereby linking audio samples through their annotated captions to all othercaptions in the dataset. To integrate the computed relevances into training, weemploy a listwise ranking objective, where relevance scores are converted intoprobabilities of ranking audio samples for a given textual query. We show theeffectiveness of the proposed method by demonstrating improvements intext-based audio retrieval compared to methods that use binary audio-captionrelevances for training.

【4】 Memory-Efficient Training for Deep Speaker Embedding Learning in Speaker  Verification
标题: 说话人验证中深度说话人嵌入学习的记忆高效训练
链接:https://arxiv.org/abs/2412.01195
作者: Bei Liu,  Yanmin Qian
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:最近的说话人确认(SV)系统已经显示出采用更深的说话人嵌入提取器的趋势。尽管更深、更大的神经网络可以显着提高性能,但其大量的内存需求阻碍了在消费者GPU上的训练。在本文中,我们探索了一种用于资源受限场景中的深度说话人嵌入学习的记忆高效训练策略。首先,我们对SV系统训练过程中的GPU内存分配进行了系统的分析。经验观察表明,激活和优化器状态是内存消耗的主要来源。对于激活,我们设计了两种类型的可逆神经网络,它们消除了在反向传播过程中存储中间激活的需要,从而在不损失性能的情况下显着减少了内存使用。对于优化器状态,我们引入了一种动态量化方法,该方法将原始的32位浮点值替换为基于动态树的8位数据类型。VoxCeleb上的实验结果表明,ResNets和DF-ResNets的可逆变体可以执行训练,而无需在GPU内存中缓存激活。此外,与32位版本相比,SGD和Adam的8位版本在保持性能的同时节省了75%的内存成本。最后,内存使用和性能的详细比较表明,我们提出的模型实现了高达16.2倍的内存节省,与香草系统相比,几乎相同的参数和性能。与之前需要多个高端GPU(如A100)不同,我们只需一两个消费级2080 Ti GPU就可以有效地训练深度扬声器嵌入提取器。
摘要:Recent speaker verification (SV) systems have shown a trend toward adoptingdeeper speaker embedding extractors. Although deeper and larger neural networkscan significantly improve performance, their substantial memory requirementshinder training on consumer GPUs. In this paper, we explore a memory-efficienttraining strategy for deep speaker embedding learning in resource-constrainedscenarios. Firstly, we conduct a systematic analysis of GPU memory allocationduring SV system training. Empirical observations show that activations andoptimizer states are the main sources of memory consumption. For activations,we design two types of reversible neural networks which eliminate the need tostore intermediate activations during back-propagation, thereby significantlyreducing memory usage without performance loss. For optimizer states, weintroduce a dynamic quantization approach that replaces the original 32-bitfloating-point values with a dynamic tree-based 8-bit data type. Experimentalresults on VoxCeleb demonstrate that the reversible variants of ResNets andDF-ResNets can perform training without the need to cache activations in GPUmemory. In addition, the 8-bit versions of SGD and Adam save 75% of memorycosts while maintaining performance compared to their 32-bit counterparts.Finally, a detailed comparison of memory usage and performance indicates thatour proposed models achieve up to 16.2x memory savings, with nearly identicalparameters and performance compared to the vanilla systems. In contrast to theprevious need for multiple high-end GPUs such as the A100, we can effectivelytrain deep speaker embedding extractors with just one or two consumer-level2080Ti GPUs.

【5】 AlignFormer: Modality Matching Can Achieve Better Zero-shot  Instruction-Following Speech-LLM
标题: AlignFormer:情态匹配可以实现更好的Zero-Shot教学-LLM
链接:https://arxiv.org/abs/2412.01145
作者: Ruchao Fan,  Bo Ren,  Yuxuan Hu,  Rui Zhao,  Shujie Liu,  Jinyu Li
摘要:将语音集成到LLM(语音LLM)最近越来越受到关注。主流的解决方案是将经过良好训练的语音编码器和LLM与神经适配器连接起来。然而,语音和文本序列之间的长度不匹配没有得到很好的处理,导致语音和文本之间的模态匹配不完美。在这项工作中,我们提出了一种新的神经适配器,AlignFormer,以减少两种模式之间的长度差距。AlignFormer由CTC和动态窗口QFormer层组成,其中CTC对齐为qFormer层提供动态窗口信息。LLM主干在训练中被冻结,以保留其文本能力,特别是指令遵循能力。当仅使用ASR数据进行训练时,所提出的AlignFormer解锁了语音LLM的指令跟随能力,并且模型可以执行zero-shot语音翻译(ST)和语音问答(SQA)任务。事实上,带有AlignFormer的speech-LLM理论上可以执行LLM主干在语音版本中可以处理的任何任务。为了评估的有效性的预防以下语音LLM,我们建议使用指令跟随率(IFR),并提供了一个系统的角度IFR评估。另外,我们发现音频在训练中的位置会影响语音LLM的指令跟随能力,并对其进行了深入的研究,结果表明音频优先训练比重复优先训练获得更高的IFR。AlignFormer可以通过音频优先训练实现接近100%的IFR,并通过预防优先训练在一些评估数据上实现从零到非零的IFR的改变游戏规则的改进。我们相信,这项研究是一个很大的一步,完美的语音和文本模态匹配在LLM嵌入空间。
摘要:Integrating speech into LLM (speech-LLM) has gaining increased attentionrecently. The mainstream solution is to connect a well-trained speech encoderand LLM with a neural adapter. However, the length mismatch between the speechand text sequences are not well handled, leading to imperfect modality matchingbetween the speech and text. In this work, we propose a novel neural adapter,AlignFormer, to reduce the length gap between the two modalities. AlignFormerconsists of CTC and dynamic-window QFormer layers, where the CTC alignmentprovides the dynamic window information for qformer layers. The LLM backbone isfrozen in training to preserve its text capability, especially the instructionfollowing capability. When training with only the ASR data, the proposedAlignFormer unlocks the instruction following capability for speech-LLM and themodel can perform zero-shot speech translation (ST) and speech questionanswering (SQA) tasks. In fact, speech-LLM with AlignFormer can theoreticallyperform any tasks that the LLM backbone can deal with in the speech version. Toevaluate the effectiveness of the instruction-following speech-LLM, we proposeto use instruction following rate (IFR) and offer a systematic perspective forthe IFR evaluation. In addition, we find that the audio position in trainingwould affect the instruction following capability of speech-LLM and conduct anin-depth study on it. Our findings show that audio-first training achieveshigher IFR than instruction-first training. The AlignFormer can achieve a near100% IFR with audio-first training and game-changing improvements from zero tonon-zero IFR on some evaluation data with instruction-first training. Webelieve that this study is a big step towards the perfect speech and textmodality matching in the LLM embedding space.

【6】 Deep Learning-Based Approach for Identification and Compensation of  Nonlinear Distortions in Parametric Array Loudspeakers
标题: 基于深度学习的参数阵列扬声器非线性失真识别和补偿方法
链接:https://arxiv.org/abs/2412.01092
作者: Mengtong Li,  Tao Zhuang,  Kai Chen,  Jia-Xin Zhong,  Jing Lu
备注:5 pages, 7 figures
摘要:与传统的电动扬声器相比,参量阵列扬声器(PAL)为音频应用提供了卓越的方向性,但由于其固有的复杂解调过程而遭受显著的非线性失真。基于Volterra滤波器的方法已被广泛用于减少这些失真,但其有效性受到其逆滤波器的能力的限制。具体而言,其p阶逆滤波器只能补偿最高p阶的非线性,而它引入的高阶非线性会继续生成低阶谐波。相比之下,本文首次引入现代深度学习方法来解决PAL系统的非线性识别和补偿。具体而言,WaveNet神经网络的前馈变体,其在音频非线性系统建模中的成功被认可,用于识别和补偿基于双边带幅度调制的PAL系统中的失真。从250 Hz到8 kHz的实验测量表明,我们提出的方法显着降低了总谐波失真和互调失真的PAL产生的音频声音,实现平均减少到4.55%和2.47%,分别。该性能明显优于使用当前最先进的基于Volterra滤波器的方法获得的结果。我们的工作为改善PAL的声音再现性能开辟了新的可能性。
摘要:Compared to traditional electrodynamic loudspeakers, the parametric arrayloudspeaker (PAL) offers exceptional directivity for audio applications butsuffers from significant nonlinear distortions due to its inherent intricatedemodulation process. The Volterra filter-based approaches have been widelyused to reduce these distortions, but the effectiveness is limited by itsinverse filter's capability. Specifically, its pth-order inverse filter canonly compensate for nonlinearities up to the pth order, while the higher-ordernonlinearities it introduces continue to generate lower-order harmonics. Incontrast, this paper introduces the modern deep learning methods for the firsttime to address nonlinear identification and compensation for PAL systems.Specifically, a feedforward variant of the WaveNet neural network, recognizedfor its success in audio nonlinear system modeling, is utilized to identify andcompensate for distortions in a double sideband amplitude modulation-based PALsystem. Experimental measurements from 250 Hz to 8 kHz demonstrate that ourproposed approach significantly reduces both total harmonic distortion andintermodulation distortion of audio sound generated by PALs, achieving averagereductions to 4.55% and 2.47%, respectively. This performance is notablysuperior to results obtained using the current state-of-the-art Volterrafilter-based methods. Our work opens new possibilities for improving the soundreproduction performance of PALs.

【7】 Detecting Spoof Voices in Asian Non-Native Speech: An Indonesian and  Thai Case Study
标题: 检测亚洲非母语语音中的恶搞声音:印度尼西亚和泰国案例研究
链接:https://arxiv.org/abs/2412.01040
作者: Aulia Adila,  Candy Olivia Mawalim,  Masashi Unoki
摘要:本研究的重点是建立有效的欺骗对策(CM)的非母语,特别是针对印尼语和泰语的发言者。我们构建了一个包括母语和非母语语音的数据集,以方便我们的研究。从语音数据中提取了三个关键特征(MFCC,LFCC和CQCC),并采用三个经典的基于机器学习的分类器(CatBoost,XGBoost和GMM)来开发鲁棒的欺骗检测系统,使用本地和组合(本地和非本地)语音数据。这导致了两种类型的CM:原生和组合。这些CM的性能进行了评估,在本地和非本地语音数据集。我们的研究结果揭示了母语CM在处理非母语语音方面面临的重大挑战,突出了特定领域解决方案的必要性。所提出的方法显示出更好的检测能力,证明了将非母语语音数据纳入训练过程的重要性。这项工作奠定了基础,更有效的欺骗检测系统在不同的语言环境。
摘要:This study focuses on building effective spoofing countermeasures (CMs) fornon-native speech, specifically targeting Indonesian and Thai speakers. Weconstructed a dataset comprising both native and non-native speech tofacilitate our research. Three key features (MFCC, LFCC, and CQCC) wereextracted from the speech data, and three classic machine learning-basedclassifiers (CatBoost, XGBoost, and GMM) were employed to develop robustspoofing detection systems using the native and combined (native andnon-native) speech data. This resulted in two types of CMs: Native andCombined. The performance of these CMs was evaluated on both native andnon-native speech datasets. Our findings reveal significant challenges faced byNative CM in handling non-native speech, highlighting the necessity fordomain-specific solutions. The proposed method shows improved detectioncapabilities, demonstrating the importance of incorporating non-native speechdata into the training process. This work lays the foundation for moreeffective spoofing detection systems in diverse linguistic contexts.

【8】 Automating Feedback Analysis in Surgical Training: Detection,  Categorization, and Assessment
标题: 手术训练中的自动反馈分析:检测、分类和评估
链接:https://arxiv.org/abs/2412.00760
作者: Firdavs Nasriddinov,  Rafal Kocielnik,  Arushi Gupta,  Cherine Yang,  Elyssa Wong,  Anima Anandkumar,  Andrew Hung
备注:Accepted as a proceedings paper at Machine Learning for Health 2024
摘要:这项工作介绍了第一个框架重建手术对话从非结构化的现实世界的录音,这是至关重要的教学任务的特点。在手术培训中,培训师在现场手术期间向受训者提供的形成性口头反馈对于确保安全、立即纠正行为和促进长期技能习得至关重要。然而,由于其非结构化和专业化的性质,分析和量化这种反馈是具有挑战性的。自动化系统对于大规模管理这些复杂性至关重要,允许创建结构化数据集,以增强反馈分析并改善外科教育。我们的框架集成了语音活动检测,说话人日记和自动语音识别,具有一种新的增强功能,1)消除幻觉(在手术室中由噪声驱动的语音识别过程中产生的不存在的话语),2)使用Few-Shot语音样本将语音与训练者和受训者分离。这些方面对于重建准确的手术对话和了解手术室参与者的角色至关重要。使用来自33个真实手术的数据,我们证明了该系统重建手术教学对话和有效检测反馈实例的能力(F1得分为0.79+/-0.07)。此外,我们的幻觉消除步骤将反馈检测性能提高了约14%。对预测受训者行为调整和分类技术反馈的下游临床相关任务的评价显示,性能与手动注释相当,F1评分分别为0.82+/0.03和0.81+/0.03。这些结果突出了我们的框架在支持临床相关任务和改进手动方法方面的有效性。
摘要:This work introduces the first framework for reconstructing surgical dialoguefrom unstructured real-world recordings, which is crucial for characterizingteaching tasks. In surgical training, the formative verbal feedback thattrainers provide to trainees during live surgeries is crucial for ensuringsafety, correcting behavior immediately, and facilitating long-term skillacquisition. However, analyzing and quantifying this feedback is challengingdue to its unstructured and specialized nature. Automated systems are essentialto manage these complexities at scale, allowing for the creation of structureddatasets that enhance feedback analysis and improve surgical education. Ourframework integrates voice activity detection, speaker diarization, andautomated speech recaognition, with a novel enhancement that 1) removeshallucinations (non-existent utterances generated during speech recognitionfueled by noise in the operating room) and 2) separates speech from trainersand trainees using few-shot voice samples. These aspects are vital forreconstructing accurate surgical dialogues and understanding the roles ofoperating room participants. Using data from 33 real-world surgeries, wedemonstrated the system's capability to reconstruct surgical teaching dialoguesand detect feedback instances effectively (F1 score of 0.79+/-0.07). Moreover,our hallucination removal step improves feedback detection performance by ~14%.Evaluation on downstream clinically relevant tasks of predicting BehavioralAdjustment of trainees and classifying Technical feedback, showed performancescomparable to manual annotations with F1 scores of 0.82+/0.03 and 0.81+/0.03respectively. These results highlight the effectiveness of our framework insupporting clinically relevant tasks and improving over manual methods.

【9】 SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies
标题: SSDP 2.0:时间准确且不流利的语音丰富转录
链接:https://arxiv.org/abs/2412.00265
作者: Jiachen Lian,  Xuanru Zhou,  Zoe Ezzes,  Jet Vonk,  Brittany Morin,  David Baquirin,  Zachary Mille,  Maria Luisa Gorno Tempini,  Gopala Krishna Anumanchipalli
摘要:语音是文本、韵律、情感、不流畅等的分层集合。超出文本(单词)的语音的自动转录是一个未充分探索的问题。我们专注于转录语音以及非流利(不流利)。当前最先进的流水线SSDM遭受复杂的架构设计、训练复杂性和局部序列比对器中的显著缺点,并且它没有探索上下文学习能力。在这项工作中,我们提出了SSDM 2.0,它通过四个主要贡献来解决这些缺点:(1)我们提出了一种新的\textit{神经发音流}来获得高度可扩展的语音表示。(2)我们开发了一个\textit{全栈连接主义子序列比对器},可以捕捉所有类型的不流畅。(3)我们在LLM中引入了一个错误发音提示管道和一致性学习模块,以利用不流利的发音能力。(4)我们策划了Libri-Dys,并开源了目前最大规模的共语障碍语料库,\textit{Libri-Co-Dys},用于未来的研究工作。在病理性语音转录的临床实验中,我们使用nfvPPA语料库测试了SSDM 2.0,主要特征是发音不流畅。总的来说,SSDM 2.0的性能远远优于SSDM和所有其他不流畅的转录模型。查看我们的项目演示页面,网址为\url{https://berkeley-speech-group.github.io/SSDM2.0/}。
摘要:Speech is a hierarchical collection of text, prosody, emotions, dysfluencies,etc. Automatic transcription of speech that goes beyond text (words) is anunderexplored problem. We focus on transcribing speech along with non-fluencies(dysfluencies). The current state-of-the-art pipeline SSDM suffers from complexarchitecture design, training complexity, and significant shortcomings in thelocal sequence aligner, and it does not explore in-context learning capacity.In this work, we propose SSDM 2.0, which tackles those shortcomings via fourmain contributions: (1) We propose a novel \textit{neural articulatory flow} toderive highly scalable speech representations. (2) We developed a\textit{full-stack connectionist subsequence aligner} that captures all typesof dysfluencies. (3) We introduced a mispronunciation prompt pipeline andconsistency learning module into LLM to leverage dysfluency \textit{in-contextpronunciation learning} abilities. (4) We curated Libri-Dys and open-sourcedthe current largest-scale co-dysfluency corpus, \textit{Libri-Co-Dys}, forfuture research endeavors. In clinical experiments on pathological speechtranscription, we tested SSDM 2.0 using nfvPPA corpus primarily characterizedby \textit{articulatory dysfluencies}. Overall, SSDM 2.0 outperforms SSDM andall other dysfluency transcription models by a large margin. See our projectdemo page at \url{https://berkeley-speech-group.github.io/SSDM2.0/}.

【10】 Feasibility of Mental Health Triage Call Priority Prediction Using  Machine Learning
标题: 使用机器学习进行心理健康分类呼叫优先级预测的可行性
链接:https://arxiv.org/abs/2412.00057
作者: Rajib Rana,  Niall Higgins,  Kazi Nazmul Haque,  John Reilly,  Kylie Burke,  Kathryn Turner,  Terry Stedman
摘要:确保准确的呼叫优先级对于优化心理健康诊所的效率和响应能力至关重要。目前,呼叫操作员完全依赖于呼叫者的陈述来确定呼叫的优先级。事实证明,完全主观的评估可能会导致错误。此外,如果在通话过程中不利用语音属性来帮助评估,则会错失机会。不正确的优先顺序可能会导致对高风险个人的延迟援助,资源分配不当,心理健康恶化加剧,失去信任以及潜在的法律后果。必须解决这些风险,以保证精神卫生服务的可靠性和有效性。这项研究深入研究了使用机器学习(人工智能的一个分支)的潜力,以估计呼叫优先级从呼叫者的声音为用户的心理健康电话号码。在分析了459个电话记录从心理健康诊所,我们达到了92%的平衡准确率,显示出承诺,帮助呼叫运营商的效率在呼叫处理过程中,提高客户满意度。
摘要:Ensuring accurate call prioritisation is essential for optimising theefficiency and responsiveness of mental health helplines. Currently, calloperators rely entirely on the caller's statements to determine the priority ofthe calls. It has been shown that entirely subjective assessment can lead toerrors. Furthermore, it is a missed opportunity not to utilise the voiceproperties readily available during the call to aid in the evaluation.Incorrect prioritisation can result in delayed assistance for high-riskindividuals, resource misallocation, increased mental health deterioration,loss of trust, and potential legal consequences. It is vital to address theserisks to guarantee the reliability and effectiveness of mental health services.This study delves into the potential of using machine learning, a branch ofArtificial Intelligence, to estimate call priority from the callers' voices forusers of mental health phone helplines. After analysing 459 call records from amental health helpline, we achieved a balanced accuracy of 92\%, showingpromise in aiding the call operators' efficiency in call handling processes andimproving customer satisfaction.

【11】 High-precision medical speech recognition through synthetic data and  semantic correction: UNITED-MEDASR
标题: 通过合成数据和语义纠正实现高精度医学语音识别:UNITED-MEDASS
链接:https://arxiv.org/abs/2412.00055
作者: Sourav Banerjee,  Ayushi Agarwal,  Promila Ghosh
备注:15 pages
摘要:临床领域的自动语音识别(ASR)系统面临着重大挑战,特别是需要准确识别专业医学词汇并满足严格的精度要求。我们介绍了United-MedASR,一种新的架构,通过集成合成数据生成,精确的ASR微调和高级语义增强技术来解决这些挑战。United-MedASR通过综合权威来源的数据构建专业医学词汇,如ICD-10(国际疾病分类,第10次修订),MIMS(医学专业每月索引)和FDA数据库。这种丰富的词汇有助于微调Whisper ASR模型,以更好地满足临床需求。为了提高处理速度,我们采用了Faster Whisper,确保精简和高速ASR性能。此外,我们采用定制的基于BART的语义增强器来处理复杂的医学术语,从而有效地提高准确性。我们的分层方法为ASR性能建立了新的基准,在LibriSpeech测试中实现了0.985%的单词错误率(WER),在Europarl-ASR EN Guest-Test中实现了0.26%,并在Tedlium(0.29% WER)和FLEURS(0.336% WER)上表现出稳健的性能。此外,我们提出了一个适应性强的架构,可以在不同的域复制,使其成为一个通用的解决方案,特定于域的ASR系统。
摘要:Automatic Speech Recognition (ASR) systems in the clinical domain facesignificant challenges, notably the need to recognise specialised medicalvocabulary accurately and meet stringent precision requirements. We introduceUnited-MedASR, a novel architecture that addresses these challenges byintegrating synthetic data generation, precision ASR fine-tuning, and advancedsemantic enhancement techniques. United-MedASR constructs a specialised medicalvocabulary by synthesising data from authoritative sources such as ICD-10(International Classification of Diseases, 10th Revision), MIMS (Monthly Indexof Medical Specialties), and FDA databases. This enriched vocabulary helpsfinetune the Whisper ASR model to better cater to clinical needs. To enhanceprocessing speed, we incorporate Faster Whisper, ensuring streamlined andhigh-speed ASR performance. Additionally, we employ a customised BART-basedsemantic enhancer to handle intricate medical terminology, thereby increasingaccuracy efficiently. Our layered approach establishes new benchmarks in ASRperformance, achieving a Word Error Rate (WER) of 0.985% on LibriSpeechtest-clean, 0.26% on Europarl-ASR EN Guest-test, and demonstrating robustperformance on Tedlium (0.29% WER) and FLEURS (0.336% WER). Furthermore, wepresent an adaptable architecture that can be replicated across differentdomains, making it a versatile solution for domain-specific ASR systems.

【12】 A Context-Based Numerical Format Prediction for a Text-To-Speech System
标题: 基于上下文的文本语音转换系统数字格式预测
链接:https://arxiv.org/abs/2412.00028
作者: Yaser Darwesh,  Lit Wei Wern,  Mumtaz Begum Mustafa
备注:21 pages, 6 tables, 1 figure
摘要:许多现有的TTS系统无法准确合成包含多种数字格式的文本,导致合成语音的清晰度降低。本研究旨在开发一种数字格式分类器,可以对六种类型的数字上下文进行分类。实验进行了使用建议的基于上下文的特征提取技术,这是集中在提取关键字,标点符号和符号作为数字的特征。支持向量机,K-近邻线性判别分析,决策树被用来作为分类器。我们使用了10倍交叉验证技术来确定在召回率和精度方面的分类准确度。可以发现,所提出的解决方案是优于现有的特征提取技术,提高了30%至37%的分类精度。数字格式分类的使用可以提高TTS系统的可懂度。
摘要:Many of the existing TTS systems cannot accurately synthesize text containinga variety of numerical formats, resulting in reduced intelligibility of thesynthesized speech. This research aims to develop a numerical format classifierthat can classify six types of numeric contexts. Experiments were carried outusing the proposed context-based feature extraction technique, which is focusedon extracting keywords, punctuation marks, and symbols as the features of thenumbers. Support Vector Machine, K-Nearest Neighbors Linear DiscriminantAnalysis, and Decision Tree were used as classifiers. We have used the 10-foldcross-validation technique to determine the classification accuracy in terms ofrecall and precision. It can be found that the proposed solution is better thanthe existing feature extraction technique with improvement to theclassification accuracy by 30% to 37%. The use of the number formatclassification can increase the intelligibility of the TTS systems.

【13】 Generative AI-based data augmentation for improved bioacoustic  classification in noisy environments
标题: 基于人工智能的生成性数据增强,以改进噪音环境中的生物声学分类
链接:https://arxiv.org/abs/2412.01530
作者: Anthony Gibbons,  Emma King,  Ian Donohue,  Andrew Parnell
备注:18 pages, 3 tables, 5 figures
摘要:1.获取数据以训练基于强大人工智能(AI)的物种分类模型可能具有挑战性,特别是对于稀有物种。数据增强可以通过增加训练数据的多样性来提高分类准确性,并且比专家标记的数据更便宜。然而,许多经典的基于图像的增强技术不适合音频频谱图。2.我们研究了两种生成AI模型作为数据增强工具来合成声谱图和补充音频数据:辅助分类器生成对抗网络(ACGAN)和去噪扩散概率模型(DDPM)。后者在生成的频谱图的真实性和所得分类任务的准确性方面表现得特别好。3.除了这些新的方法,我们提出了一个新的音频数据集640小时的鸟叫声从风力发电场在爱尔兰,其中约800个样本已被标记的专家。考虑到背景风和涡轮机噪声,风电场数据对于分类模型特别具有挑战性。4.与高度自信的BirdNET预测相比,在真实数据和合成数据上训练分类模型的集合给出了92.6%的准确率(仅使用真实数据时为90.5%)。5.我们的方法可用于增强更多物种和其他土地利用类型的声学信号,并有可能使我们开发可靠的基于人工智能的稀有物种检测能力发生重大变化。我们的代码可以在https://github.com/gibbona1/ SpectrogramGenAI上找到。
摘要:1. Obtaining data to train robust artificial intelligence (AI)-based modelsfor species classification can be challenging, particularly for rare species.Data augmentation can boost classification accuracy by increasing the diversityof training data and is cheaper to obtain than expert-labelled data. However,many classic image-based augmentation techniques are not suitable for audiospectrograms. 2. We investigate two generative AI models as data augmentationtools to synthesise spectrograms and supplement audio data: AuxiliaryClassifier Generative Adversarial Networks (ACGAN) and Denoising DiffusionProbabilistic Models (DDPMs). The latter performed particularly well in termsof both realism of generated spectrograms and accuracy in a resultingclassification task. 3. Alongside these new approaches, we present a new audiodata set of 640 hours of bird calls from wind farm sites in Ireland,approximately 800 samples of which have been labelled by experts. Wind farmdata are particularly challenging for classification models given thebackground wind and turbine noise. 4. Training an ensemble of classificationmodels on real and synthetic data combined gave 92.6% accuracy (and 90.5% withjust the real data) when compared with highly confident BirdNET predictions. 5.Our approach can be used to augment acoustic signals for more species and otherland-use types, and has the potential to bring about a step-change in ourcapacity to develop reliable AI-based detection of rare species. Our code isavailable at https://github.com/gibbona1/ SpectrogramGenAI.

【14】 Reject Threshold Adaptation for Open-Set Model Attribution of Deepfake  Audio
标题: Deepfake音频开集模型属性的预设阈值自适应
链接:https://arxiv.org/abs/2412.01425
作者: Xinrui Yan,  Jiangyan Yi,  Jianhua Tao,  Yujie Chen,  Hao Gu,  Guanjun Li,  Junzuo Zhou,  Yong Ren,  Tao Xu
备注:Accepted by ISCSLP 2024
摘要:面向开放环境的Deepfake音频的开集模型属性是一个新兴的研究课题,旨在识别Deepfake音频的生成模型。大多数以前的工作需要手动设置未知类的拒绝阈值,以与预测的概率进行比较。然而,模型往往过度拟合训练实例,并生成过于自信的预测。此外,有效区分当前数据集中的未知类别的阈值可能不适合于识别另一数据分布中的已知和未知类别。为了解决这些问题,我们提出了一个新的框架,用于Deepfake音频的开集模型属性,并具有拒绝阈值自适应(ReTA)。具体地,重构误差学习模块通过将系统指纹的表示与对应于目标类别或随机选择的其他类别标签的标签相组合来进行训练。该过程生成匹配和非匹配的重构样本,建立每类的重构误差分布,并为拒绝阈值计算模块奠定基础。拒绝阈值计算模块利用高斯概率估计来拟合匹配和非匹配重建误差的分布。然后,它通过概率最小化准则计算所有类别的自适应拒绝阈值。实验结果证明了ReTA在改善Deepfake音频的开集模型属性方面的有效性。
摘要:Open environment oriented open set model attribution of deepfake audio is anemerging research topic, aiming to identify the generation models of deepfakeaudio. Most previous work requires manually setting a rejection threshold forunknown classes to compare with predicted probabilities. However, models oftenoverfit training instances and generate overly confident predictions. Moreover,thresholds that effectively distinguish unknown categories in the currentdataset may not be suitable for identifying known and unknown categories inanother data distribution. To address the issues, we propose a novel frameworkfor open set model attribution of deepfake audio with rejection thresholdadaptation (ReTA). Specifically, the reconstruction error learning moduletrains by combining the representation of system fingerprints with labelscorresponding to either the target class or a randomly chosen other classlabel. This process generates matching and non-matching reconstructed samples,establishing the reconstruction error distributions for each class and layingthe foundation for the reject threshold calculation module. The rejectthreshold calculation module utilizes gaussian probability estimation to fitthe distributions of matching and non-matching reconstruction errors. It thencomputes adaptive reject thresholds for all classes through probabilityminimization criteria. The experimental results demonstrate the effectivenessof ReTA in improving the open set model attributes of deepfake audio.

【15】 OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
标题: OmniFlow:具有多模式整流流的任何对任何一代
链接:https://arxiv.org/abs/2412.01169
作者: Shufan Li,  Konstantinos Kallidromitis,  Akash Gokul,  Zichun Liao,  Yusuke Kato,  Kazuki Kozuka,  Aditya Grover
备注:12 pages, 14 figures
摘要:我们介绍OmniFlow,一种新的生成模型,专为任何到任何生成任务,如文本到图像,文本到音频,音频到图像合成。OmniFlow推进了文本到图像模型中使用的整流(RF)框架,以处理多种模态的联合分布。它在广泛的任务上优于以前的任何对任何模型,例如文本到图像和文本到音频合成。我们的工作提供了三个关键贡献:首先,我们将RF扩展到多模态设置,并引入了一种新的指导机制,使用户能够灵活地控制生成的输出中不同模态之间的对齐。其次,我们提出了一种新的架构,扩展了稳定扩散3的文本到图像的MMDiT架构,并使音频和文本生成。扩展模块可以单独进行有效的预训练,并与普通的文本到图像MMDiT合并进行微调。最后,我们对用于大规模音频和文本生成的整流流量Transformers的设计选择进行了全面的研究,为优化不同模态的性能提供了有价值的见解。该守则可在https://github.com/jacklishufan/OmniFlows上查阅。
摘要:We introduce OmniFlow, a novel generative model designed for any-to-anygeneration tasks such as text-to-image, text-to-audio, and audio-to-imagesynthesis. OmniFlow advances the rectified flow (RF) framework used intext-to-image models to handle the joint distribution of multiple modalities.It outperforms previous any-to-any models on a wide range of tasks, such astext-to-image and text-to-audio synthesis. Our work offers three keycontributions: First, we extend RF to a multi-modal setting and introduce anovel guidance mechanism, enabling users to flexibly control the alignmentbetween different modalities in the generated outputs. Second, we propose anovel architecture that extends the text-to-image MMDiT architecture of StableDiffusion 3 and enables audio and text generation. The extended modules can beefficiently pretrained individually and merged with the vanilla text-to-imageMMDiT for fine-tuning. Lastly, we conduct a comprehensive study on the designchoices of rectified flow transformers for large-scale audio and textgeneration, providing valuable insights into optimizing performance acrossdiverse modalities. The Code will be available athttps://github.com/jacklishufan/OmniFlows.

【16】 HumekaFL: Automated Detection of Neonatal Asphyxia Using Federated  Learning
标题: HumekaFL:使用联邦学习自动检测新生儿窒息
链接:https://arxiv.org/abs/2412.01167
作者: Pamely Zantou,  Blessed Guda,  Bereket Retta,  Gladys Inabeza,  Carlee Joe-Wong,  Assane Gueye
备注:Poster at ACM compass 2024
摘要:出生性失语症(BA)是一种严重的疾病,其特征是在分娩过程中新生儿的氧气供应不足。BA是世界上新生儿死亡的主要原因之一。虽然过去二十年来新生儿死亡率有所下降,但发展中世界,特别是撒哈拉以南非洲,五岁以下儿童死亡率仍然最高。虽然循证方法通常用于在非洲医疗环境中检测BA,但它们可能会受到医生错误或诊断延迟的影响,从而无法及时干预。集中式机器学习(ML)方法在早期检测BA方面表现出良好的性能,但需要敏感的健康数据在训练之前离开其场所,这不能保证隐私和安全。因此,非洲的医疗机构不愿意采用这种解决方案。为了应对这一挑战,我们提出了一种基于联邦学习(FL)的软件架构,这是一种通过设计优先考虑隐私和安全的分布式学习方法。我们开发了一个用户友好和具有成本效益的移动应用程序,嵌入FL管道,用于早期检测BA。我们的联邦SVM模型优于现有文献中的集中式SVM管道和基于神经网络(NN)的方法
摘要:Birth Apshyxia (BA) is a severe condition characterized by insufficientsupply of oxygen to a newborn during the delivery. BA is one of the primarycauses of neonatal death in the world. Although there has been a decline inneonatal deaths over the past two decades, the developing world, particularlysub-Saharan Africa, continues to experience the highest under-five (<5)mortality rates. While evidence-based methods are commonly used to detect BA inAfrican healthcare settings, they can be subject to physician errors or delaysin diagnosis, preventing timely interventions. Centralized Machine Learning(ML) methods demonstrated good performance in early detection of BA but requiresensitive health data to leave their premises before training, which does notguarantee privacy and security. Healthcare institutions are therefore reluctantto adopt such solutions in Africa. To address this challenge, we suggest afederated learning (FL)-based software architecture, a distributed learningmethod that prioritizes privacy and security by design. We have developed auser-friendly and cost-effective mobile application embedding the FL pipelinefor early detection of BA. Our Federated SVM model outperformed centralized SVMpipelines and Neural Networks (NN)-based methods in the existing literature

【17】 The Codec Language Model-based Zero-Shot Spontaneous Style TTS System  for CoVoC Challenge 2024
标题: 用于2024年CoVoC挑战赛的基于Codec语言模型的Zero-Shot自发风格TTC系统
链接:https://arxiv.org/abs/2412.01100
作者: Shuoyi Zhou,  Yixuan Zhou,  Weiqing Li,  Jun Chen,  Runchuan Ye,  Weihao Wu,  Zijian Lin,  Shun Lei,  Zhiyong Wu
备注:Accepted by ISCSLP 2024
摘要:本文描述了用于ISCSLP 2024会话语音克隆挑战赛(CoVoC)的zero-shot自发式TTS系统。我们提出了一个基于LLaMA的编解码器语言模型与延迟模式,以实现自发风格的语音克隆。为了提高语音可懂度,我们在语言模型中引入了无分类器指导(CFG)策略,以加强对标记预测的条件指导。为了生成高质量的话语,我们采用了有效的数据预处理操作,并微调我们的模型与选定的高质量的自发语音数据。在CoVoC约束的轨道上的官方评估表明,我们的系统达到了最好的语音自然度MOS为3.80,并获得了可观的语音质量和说话人相似度的结果。
摘要:This paper describes the zero-shot spontaneous style TTS system for theISCSLP 2024 Conversational Voice Clone Challenge (CoVoC). We propose aLLaMA-based codec language model with a delay pattern to achieve spontaneousstyle voice cloning. To improve speech intelligibility, we introduce theClassifier-Free Guidance (CFG) strategy in the language model to strengthenconditional guidance on token prediction. To generate high-quality utterances,we adopt effective data preprocessing operations and fine-tune our model withselected high-quality spontaneous speech data. The official evaluations in theCoVoC constrained track show that our system achieves the best speechnaturalness MOS of 3.80 and obtains considerable speech quality and speakersimilarity results.

【18】 FreeCodec: A disentangled neural speech codec with fewer tokens
标题: FreeCodec:一个具有更少令牌的分离神经语音编解码器
链接:https://arxiv.org/abs/2412.01053
作者: Youqiang Zheng,  Weiping Tu,  Yueteng Kang,  Jie Chen,  Yike Zhang,  Li Xiao,  Yuhong Yang,  Long Ma
摘要:神经语音编解码器以其出色的离散表征重构能力而受到广泛关注。   它是语音编码和大型语言模型(LLM)等生成任务的关键组成部分。   然而,大多数基于残差矢量量化的作品表现较差,由于编码效率低,建模复杂的耦合信息的令牌较少。   在本文中,我们提出了一种名为FreeCodec的神经语音编解码器,它通过将语音的内在属性分解为不同的分量来采用更有效的编码框架:   1)提取全局向量作为音色信息,   2)使用具有长步幅级别的韵律编码器来对韵律信息进行建模,   3)内容信息来自内容编码器。   使用不同的训练策略,FreeCodec在重建和解纠缠场景中实现了最先进的性能。   从主观和客观的实验结果表明,我们的框架优于现有的方法。
摘要:Neural speech codecs have gained great attention for their outstandingreconstruction with discrete token representations. It is a crucial component in generative tasks such as speech coding and largelanguage models (LLM). However, most works based on residual vector quantization perform worse withfewer tokens due to low coding efficiency for modeling complex coupledinformation. In this paper, we propose a neural speech codec named FreeCodec which employsa more effective encoding framework by decomposing intrinsic properties ofspeech into different components: 1) a global vector is extracted as the timbre information, 2) a prosody encoder with a long stride level is used to model the prosodyinformation, 3) the content information is from a content encoder. Using different training strategies, FreeCodec achieves state-of-the-artperformance in reconstruction and disentanglement scenarios. Results from subjective and objective experiments demonstrate that ourframework outperforms existing methods.

【19】 Complexity boosted adaptive training for better low resource ASR  performance
标题: 复杂性增强自适应训练,以获得更好的低资源ASB性能
链接:https://arxiv.org/abs/2412.00877
作者: Hongxuan Lu,  Shenjian Wang,  Biao Li
摘要:在ASR模型的整个训练过程中,数据增强的强度和计算训练损失的方法以基于预设参数的调节方式应用。例如,SpecAugment采用预定义的增强强度来掩蔽时频域频谱的部分。类似地,在基于CTC的多层模型中,通常基于训练过程期间编码器的最终层的输出来确定损失。然而,忽略动态特性可能使训练模型次优。为了解决这个问题,我们提出了一个两阶段的训练方法,称为复杂性增强自适应(CBA)训练。它涉及根据训练样本的复杂性对数据增强策略和CTC损失传播进行动态调整。在第一阶段,我们训练模型与中间CTC为基础的正则化和数据增强没有任何自适应的政策。在第二阶段,我们提出了一种新的自适应策略,称为MinMax-IBF,它计算样本的复杂度。我们将MinMax-IBF策略与数据增强和中间CTC损失正则化相结合,以继续训练。所提出的CBA训练方法显示出相当大的改进,在LibriSpeech 100 h测试清洁和测试其他数据集上WER的相对减少高达13.4%和14.1%,在AISHELL-1测试集上的相对减少也高达6.3%,超过Wenet中的Conformer架构。
摘要:During the entire training process of the ASR model, the intensity of dataaugmentation and the approach of calculating training loss are applied in aregulated manner based on preset parameters. For example, SpecAugment employs apredefined strength of augmentation to mask parts of the time-frequency domainspectrum. Similarly, in CTC-based multi-layer models, the loss is generallydetermined based on the output of the encoder's final layer during the trainingprocess. However, ignoring dynamic characteristics may suboptimally trainmodels. To address the issue, we present a two-stage training method, known ascomplexity-boosted adaptive (CBA) training. It involves making dynamicadjustments to data augmentation strategies and CTC loss propagation based onthe complexity of the training samples. In the first stage, we train the modelwith intermediate-CTC-based regularization and data augmentation without anyadaptive policy. In the second stage, we propose a novel adaptive policy,called MinMax-IBF, which calculates the complexity of samples. We combine theMinMax-IBF policy to data augmentation and intermediate CTC loss regularizationto continue training. The proposed CBA training approach shows considerableimprovements, up to 13.4% and 14.1% relative reduction in WER on theLibriSpeech 100h test-clean and test-other dataset and also up to 6.3% relativereduction on AISHELL-1 test set, over the Conformer architecture in Wenet.

【20】 A Comparative Study of LLM-based ASR and Whisper in Low Resource and  Code Switching Scenario
标题: 低资源和代码交换场景下基于LLM的ASB和Whisper的比较研究
链接:https://arxiv.org/abs/2412.00721
作者: Zheshu Song,  Ziyang Ma,  Yifan Yang,  Jianheng Zhuo,  Xie Chen
备注:4 pages
摘要:大型语言模型(LLM)在各种NLP任务中表现出卓越的性能,它们与语音编码器的集成正在迅速成为自动语音识别(ASR)领域的主导趋势。以前的工作主要集中在利用LLM的语音识别在英语和汉语。然而,它们在低资源环境中解决语音识别挑战的潜力仍然没有得到充分的探索。因此,在这项工作中,我们的目标是探索能力的LLM在低资源ASR和普通话-英语码切换ASR。我们还评估和比较基于LLM的ASR系统对Whisper模型的识别性能。大量的实验表明,基于LLM的ASR在低资源ASR中比Whisper模型产生了12.8%的相对增益,而Whisper在中英文代码切换ASR中表现更好。我们希望这项研究可以阐明ASR低资源的情况下。
摘要:Large Language Models (LLMs) have showcased exceptional performance acrossdiverse NLP tasks, and their integration with speech encoder is rapidlyemerging as a dominant trend in the Automatic Speech Recognition (ASR) field.Previous works mainly concentrated on leveraging LLMs for speech recognition inEnglish and Chinese. However, their potential for addressing speech recognitionchallenges in low resource settings remains underexplored. Hence, in this work,we aim to explore the capability of LLMs in low resource ASR andMandarin-English code switching ASR. We also evaluate and compare therecognition performance of LLM-based ASR systems against Whisper model.Extensive experiments demonstrate that LLM-based ASR yields a relative gain of12.8\% over the Whisper model in low resource ASR while Whisper performs betterin Mandarin-English code switching ASR. We hope that this study could shedlight on ASR for low resource scenarios.

【21】 Audio Atlas: Visualizing and Exploring Audio Datasets
标题: 音频地图集:可视化和探索音频数据集
链接:https://arxiv.org/abs/2412.00591
作者: Luca A. Lanzendörfer,  Florian Grötschla,  Uzeyir Valizada,  Roger Wattenhofer
备注:Extended Abstract at ISMIR 2024
摘要:我们介绍音频地图集,一个交互式的Web应用程序,用于可视化音频数据使用文本音频嵌入。Audio Atlas旨在使用对比嵌入模型和矢量数据库来促进音频数据集的探索和分析,以实现高效的数据管理和语义搜索。该系统将音频嵌入映射到二维空间中,并利用DeepScatter进行动态可视化。Audio Atlas专为可扩展性而设计,允许轻松集成新数据集,使用户能够更好地理解他们的音频数据并识别模式和异常值。我们开源了Audio Atlas的代码库,并提供了包含各种音频和音乐数据集的初始实现。
摘要:We introduce Audio Atlas, an interactive web application for visualizingaudio data using text-audio embeddings. Audio Atlas is designed to facilitatethe exploration and analysis of audio datasets using a contrastive embeddingmodel and a vector database for efficient data management and semantic search.The system maps audio embeddings into a two-dimensional space and leveragesDeepScatter for dynamic visualization. Designed for extensibility, Audio Atlasallows easy integration of new datasets, enabling users to better understandtheir audio data and identify both patterns and outliers. We open-source thecodebase of Audio Atlas, and provide an initial implementation containingvarious audio and music datasets.

【22】 From Audio Deepfake Detection to AI-Generated Music Detection -- A  Pathway and Overview
标题: 从音频Deepfake检测到人工智能生成的音乐检测--路径和概述
链接:https://arxiv.org/abs/2412.00571
作者: Yupei Li,  Manuel Milling,  Lucia Specia,  Björn W. Schuller
摘要:随着人工智能(AI)技术的不断发展,它们在生成逼真的、适合上下文的内容方面的应用已经扩展到各个领域。音乐是一种艺术形式和娱乐媒介,深深植根于人类文化,人工智能越来越多地参与其制作。然而,人工智能音乐生成(AIGM)工具的不受管制的使用引起了人们对音乐产业、版权和艺术完整性潜在负面影响的担忧,强调了有效检测AIGM的重要性。本文综述了现有的AIGM检测方法。为了为AIGM检测的一般工作和挑战奠定基础,我们首先回顾了AIGM的一般原理,包括deepfake音频的最新进展以及多模态检测技术。我们进一步提出了一种潜在的途径,用于利用从音频deepfake检测到AIGM检测的基础模型。此外,我们还讨论了这些工具的影响,并提出了未来研究的方向,以解决该领域正在面临的挑战。
摘要:As Artificial Intelligence (AI) technologies continue to evolve, their use ingenerating realistic, contextually appropriate content has expanded intovarious domains. Music, an art form and medium for entertainment, deeply rootedinto human culture, is seeing an increased involvement of AI into itsproduction. However, the unregulated use of AI music generation (AIGM) toolsraises concerns about potential negative impacts on the music industry,copyright and artistic integrity, underscoring the importance of effective AIGMdetection. This paper provides an overview of existing AIGM detection methods.To lay a foundation to the general workings and challenges of AIGM detection,we first review general principles of AIGM, including recent advancements indeepfake audios, as well as multimodal detection techniques. We further proposea potential pathway for leveraging foundation models from audio deepfakedetection to AIGM detection. Additionally, we discuss implications of thesetools and propose directions for future research to address ongoing challengesin the field.

【23】 Personal Sound Zones and Shielded Localized Communication through Active  Acoustic Control
标题: 通过主动声学控制实现个人声区和屏蔽本地通信
链接:https://arxiv.org/abs/2412.00456
作者: Neil Jerome A. Egarguin,  Daniel Onofrei
摘要:在本文中,我们提出了一个时域扩展我们的策略上操纵辐射标量亥姆霍兹场,并讨论了两个重要的应用场景,即(1)创建一个有界域内的个人声音区域和(2)屏蔽本地化通信。我们的策略是基于作者以前的工作建立的可能性和稳定性控制声场使用一个阵列的几乎非辐射耦合源,并提出了一个详细的傅立叶合成方法对时域效果。我们要求声源阵列在控制区域上产生所需的场,同时在更大的外接球体之外保持零场。本文回顾了主要的理论结果,然后提出了基本的傅立叶合成范式,并显示,通过相关的模拟,我们的策略的性能。
摘要:In this paper, we present a time domain extension of our strategy onmanipulating radiated scalar Helmholtz fields and discuss two important appliedscenarios, namely (1) creating personal sound zones inside a bounded domain and(2) shielded localized communication. Our strategy is based on the authors'previous works establishing the possibility and stability of controllingacoustic fields using an array of almost non-radiating coupling sources andpresents a detailed Fourier synthesis approach towards a time-domain effect. Werequire that the array of acoustic sources creates the desired fields on thecontrol regions while maintaining a zero field beyond a larger circumscribedsphere. This paper recalls the main theoretical results then presents theunderlying Fourier synthesis paradigm and show, through relevant simulations,the performance of our strategy.

【24】 Sample adaptive data augmentation with progressive scheduling
标题: 采用渐进式调度的自适应数据增强示例
链接:https://arxiv.org/abs/2412.00415
作者: Hongxuan Lu,  Biao Li
摘要:数据增强是一种被广泛采用的提高自动语音识别(ASR)鲁棒性的技术。对所有训练数据采用固定的数据增强策略是一种常见的做法。然而,重要的是要注意,在单个训练批次内的不同样本之间,可能存在诸如背景噪声、语音速率等因素的变化。通过使用固定增广策略,存在模型可能达到次优状态的风险。除了采用固定增强策略的风险之外,模型的能力在不同的训练阶段可能会有所不同。为了解决这些问题,本文提出了渐进调度的样本自适应数据增强方法(PS-SapAug)。所提出的方法在两阶段训练方法中应用动态数据增强。它采用混合归一化来计算基于每个样本的损失的样本特定的增强参数。此外,增强的概率在整个训练过程中逐渐增加。我们的方法在流行的ASR基准数据集上进行了评估,包括Aishell-1和LibriSpeech-100 h,在LibriSpeech-100 h测试-clean上实现了高达8.13%的WER减少,在测试-其他上实现了6.23%,在AISHELL-1测试集上实现了5.26%,这证明了我们的方法在提高性能和最小化错误方面的有效性。
摘要:Data augmentation is a widely adopted technique utilized to improve therobustness of automatic speech recognition (ASR). Employing a fixed dataaugmentation strategy for all training data is a common practice. However, itis important to note that there can be variations in factors such as backgroundnoise, speech rate, etc. among different samples within a single trainingbatch. By using a fixed augmentation strategy, there is a risk that the modelmay reach a suboptimal state. In addition to the risks of employing a fixedaugmentation strategy, the model's capabilities may differ across varioustraining stages. To address these issues, this paper proposes the method ofsample-adaptive data augmentation with progressive scheduling(PS-SapAug). Theproposed method applies dynamic data augmentation in a two-stage trainingapproach. It employs hybrid normalization to compute sample-specificaugmentation parameters based on each sample's loss. Additionally, theprobability of augmentation gradually increases throughout the trainingprogression. Our method is evaluated on popular ASR benchmark datasets,including Aishell-1 and Librispeech-100h, achieving up to 8.13% WER reductionon LibriSpeech-100h test-clean, 6.23% on test-other, and 5.26% on AISHELL-1test set, which demonstrate the efficacy of our approach enhancing performanceand minimizing errors.

【25】 MusicGen-Chord: Advancing Music Generation through Chord Progressions  and Interactive Web-UI
标题: MusicGen-Chord:通过Chord进展和交互式Web UI推进音乐生成
链接:https://arxiv.org/abs/2412.00325
作者: Jongmin Jung,  Andreas Jansson,  Dasaem Jeong
备注:Late-breaking/demo (LBD) at ISMIR 2024. this https URL
摘要:MusicGen是一种音乐生成语言模型(LM),可以根据文本描述和旋律特征进行调节。我们介绍MusicGen和弦,它扩展了这种能力,将和弦进行功能。该模型将独热编码的旋律色度向量修改为多热编码的和弦色度向量,使得能够生成既反映和弦进行又反映文本描述的音乐。此外,我们开发了MusicGen-Remixer,这是一个利用MusicGen-Chord生成基于文本描述的输入音乐混音的应用程序。这两种模型都使用cog集成到Replicate的Web UI中,促进了广泛的可访问性和用户友好的可控交互,以创建和体验AI生成的音乐。
摘要:MusicGen is a music generation language model (LM) that can be conditioned ontextual descriptions and melodic features. We introduce MusicGen-Chord, whichextends this capability by incorporating chord progression features. This modelmodifies one-hot encoded melody chroma vectors into multi-hot encoded chordchroma vectors, enabling the generation of music that reflects both chordprogressions and textual descriptions. Furthermore, we developedMusicGen-Remixer, an application utilizing MusicGen-Chord to generate remixesof input music conditioned on textual descriptions. Both models are integratedinto Replicate's web-UI using cog, facilitating broad accessibility anduser-friendly controllable interaction for creating and experiencingAI-generated music.

【26】 Improving speaker verification robustness with synthetic emotional  utterances
标题: 利用合成情感话语提高说话者验证稳健性
链接:https://arxiv.org/abs/2412.00319
作者: Nikhil Kumar Koditala,  Chelsea Jui-Ting Ju,  Ruirui Li,  Minho Jin,  Aman Chadha,  Andreas Stolcke
摘要:说话人验证(SV)系统提供了一种验证服务,用于确认给定的语音样本是否来自特定的说话人。这项技术为满足个人偏好的各种个性化应用铺平了道路。SV系统面临的一个值得注意的挑战是它们在一系列情绪谱中一致表现的能力。大多数现有的模型表现出较高的错误率时,处理情绪的话语相比,中性的。因此,这种现象往往导致错过感兴趣的演讲。这个问题主要源于有限的可用性标记的情绪语音数据,阻碍了发展强大的扬声器表示,包括不同的情绪状态。   为了解决这个问题,我们提出了一种新的方法,采用CycleGAN框架作为数据增强方法。该技术为每个特定的说话者合成情感语音片段,同时保留独特的声音身份。我们的实验结果强调了将合成情感数据纳入训练过程的有效性。使用这个增强数据集训练的模型在情感语音场景中验证说话者的任务上始终优于基线模型,相对降低了3.64%的等错误率。
摘要:A speaker verification (SV) system offers an authentication service designedto confirm whether a given speech sample originates from a specific speaker.This technology has paved the way for various personalized applications thatcater to individual preferences. A noteworthy challenge faced by SV systems istheir ability to perform consistently across a range of emotional spectra. Mostexisting models exhibit high error rates when dealing with emotional utterancescompared to neutral ones. Consequently, this phenomenon often leads to missingout on speech of interest. This issue primarily stems from the limitedavailability of labeled emotional speech data, impeding the development ofrobust speaker representations that encompass diverse emotional states. To address this concern, we propose a novel approach employing the CycleGANframework to serve as a data augmentation method. This technique synthesizesemotional speech segments for each specific speaker while preserving the uniquevocal identity. Our experimental findings underscore the effectiveness ofincorporating synthetic emotional data into the training process. The modelstrained using this augmented dataset consistently outperform the baselinemodels on the task of verifying speakers in emotional speech scenarios,reducing equal error rate by as much as 3.64% relative.

【27】 Raw Audio Classification with Cosine Convolutional Neural Network  (CosCovNN)
标题: 使用CosCovNN卷积神经网络(CosCovNN)进行原始音频分类
链接:https://arxiv.org/abs/2412.00312
作者: Kazi Nazmul Haque,  Rajib Rana,  Tasnim Jarin,  Bjorn W. Schuller Jr
摘要:这项研究探索了使用卷积神经网络(CNN)从原始波形中进行音频分类的领域,这种方法消除了在预处理步骤中提取专业特征的需要。与文献中的最新趋势不同,文献中的趋势通常侧重于仅为CNN的初始层设计前端或滤波器,我们的研究引入了余弦卷积神经网络(CosCovNN),用余弦滤波器取代传统的CNN滤波器。CosCovNN超过了等效CNN架构的准确性,参数减少了约77\%。我们的研究进一步发展了一个名为矢量量化余弦卷积神经网络与记忆(VQCCM)的增强CosCovNN,结合了记忆和矢量量化层VQCCM实现了最先进的(SOTA)性能在五个不同的数据集与现有文献相比。我们的研究结果表明,余弦滤波器可以大大提高CNN在原始音频分类中的效率和准确性。
摘要:This study explores the field of audio classification from raw waveform usingConvolutional Neural Networks (CNNs), a method that eliminates the need forextracting specialised features in the pre-processing step. Unlike recenttrends in literature, which often focuses on designing frontends or filters foronly the initial layers of CNNs, our research introduces the CosineConvolutional Neural Network (CosCovNN) replacing the traditional CNN filterswith Cosine filters. The CosCovNN surpasses the accuracy of the equivalent CNNarchitectures with approximately $77\%$ less parameters. Our research furtherprogresses with the development of an augmented CosCovNN named Vector QuantisedCosine Convolutional Neural Network with Memory (VQCCM), incorporating a memoryand vector quantisation layer VQCCM achieves state-of-the-art (SOTA)performance across five different datasets in comparison with existingliterature. Our findings show that cosine filters can greatly improve theefficiency and accuracy of CNNs in raw audio classification.

【28】 Circumventing shortcuts in audio-visual deepfake detection datasets with  unsupervised learning
标题: 通过无监督学习避免视听深度伪造检测数据集中的捷径
链接:https://arxiv.org/abs/2412.00175
作者: Dragos-Alexandru Boldisor,  Stefan Smeu,  Dan Oneata,  Elisabeta Oneata
摘要:良好的数据集对于开发和基准测试任何机器学习系统都至关重要。它们对于安全关键应用的重要性甚至更为极端,例如深度伪造检测-本文的重点。在这里,我们揭示了两个最广泛使用的音频-视频deepfake数据集遭受了以前未识别的虚假功能:前导沉默。假视频从一个非常短暂的沉默开始,仅仅基于这个功能,我们就可以几乎完美地分离真实和假样本。因此,先前的仅音频模型和音频-视频模型利用假视频中存在的静音,因此当去除前导静音时表现更差。为了避免锁定这种不需要的工件和可能其他未揭示的,我们提出了从监督到无监督学习的转变,通过专门在真实数据上训练模型。我们表明,通过调整自监督的音频-视频表示,我们消除了依赖特定于网络的偏见的风险,并提高了deepfake检测的鲁棒性。
摘要:Good datasets are essential for developing and benchmarking any machinelearning system. Their importance is even more extreme for safety criticalapplications such as deepfake detection - the focus of this paper. Here wereveal that two of the most widely used audio-video deepfake datasets sufferfrom a previously unidentified spurious feature: the leading silence. Fakevideos start with a very brief moment of silence and based on this featurealone, we can separate the real and fake samples almost perfectly. As such,previous audio-only and audio-video models exploit the presence of silence inthe fake videos and consequently perform worse when the leading silence isremoved. To circumvent latching on such unwanted artifact and possibly otherunrevealed ones we propose a shift from supervised to unsupervised learning bytraining models exclusively on real data. We show that by aligningself-supervised audio-video representations we remove the risk of relying ondataset-specific biases and improve robustness in deepfake detection.

【29】 A Survey of Recent Advances and Challenges in Deep Audio-Visual  Correlation Learning
标题: 深度视听相关学习的最新进展和挑战概览
链接:https://arxiv.org/abs/2412.00049
作者: Luis Vilaca,  Yi Yu,  Paula Vinan
备注:arXiv admin note: text overlap with arXiv:2202.13673
摘要:视听相关学习旨在捕捉和理解视听数据之间的自然现象。深度学习的快速增长推动了处理视听数据的提案的发展,并且可以在过去几年的提案数量中观察到。从而鼓励开展全面调查。除了分析在这种情况下使用的模型,我们还讨论了一些任务的定义和范式应用于人工智能多媒体。此外,我们研究了经常使用的目标函数,并讨论了如何在优化过程中利用视听数据,即,在视听领域中表示知识的不同方法。事实上,我们关注的是人类可以理解的机制,即,结构化知识是可理解性知识的反映,能够指导学习过程。最重要的是,我们总结了视听相关学习(AVCL)的最新进展,并讨论了未来的研究方向。
摘要:Audio-visual correlation learning aims to capture and understand naturalphenomena between audio and visual data. The rapid growth of Deep Learningpropelled the development of proposals that process audio-visual data and canbe observed in the number of proposals in the past years. Thus encouraging thedevelopment of a comprehensive survey. Besides analyzing the models used inthis context, we also discuss some tasks of definition and paradigm applied inAI multimedia. In addition, we investigate objective functions frequently usedand discuss how audio-visual data is exploited in the optimization process,i.e., the different methodologies for representing knowledge in theaudio-visual domain. In fact, we focus on how human-understandable mechanisms,i.e., structured knowledge that reflects comprehensible knowledge, can guidethe learning process. Most importantly, we provide a summarization of therecent progress of Audio-Visual Correlation Learning (AVCL) and discuss thefuture research directions.

机器翻译由腾讯交互翻译提供,仅供参考