今日论文合集:cs.SD语音18篇,eess.AS音频处理17篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation
标题: 两只耳朵大开:走向数字驱动的空间音频生成
作者: Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye, Huadai Liu, Honggang Zhang, Wei Xue, Yike Guo
链接:点击下载PDF文件
摘要:近年来,扩散模型在单声道音频生成中取得了巨大的成功。然而,当涉及到立体声音频生成时,音景通常具有多个对象和方向的复杂场景。由于高数据成本和不稳定的生成模型,控制具有空间上下文的立体声音频仍然具有挑战性。据我们所知,这项工作是解决这些问题的第一次尝试。我们首先构建了一个大规模的,基于模拟的,GPT辅助的数据集,BEWO-1 M,具有丰富的音景和描述,甚至包括移动和多个源。除了文本模态,我们还通过检索获得了一组图像和合理配对的立体声音频,以推进多模态生成。现有的音频生成模型倾向于生成相当随机和模糊的空间音频。为了给潜在扩散模型提供准确的指导,我们引入了SpatialSonic模型,该模型利用空间感知编码器和方位角状态矩阵来揭示合理的空间指导。通过利用空间指导,我们的统一模型不仅实现了从文本和图像生成沉浸式和可控的空间音频的目标,而且还实现了在推理过程中生成交互式音频。最后,在公平的设置下,我们对模拟和真实世界的数据进行主观和客观的评估,将我们的方法与主流方法进行比较。结果表明,我们的方法的有效性,突出其能力,以产生空间音频,坚持物理规则。摘要:Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. Controlling stereo audio with spatial contexts remains challenging due to high data costs and unstable generative models. To the best of our knowledge, this work represents the first attempt to address these issues. We first construct a large-scale, simulation-based, and GPT-assisted dataset, BEWO-1M, with abundant soundscapes and descriptions even including moving and multiple sources. Beyond text modality, we have also acquired a set of images and rationally paired stereo audios through retrieval to advance multimodal generation. Existing audio generation models tend to generate rather random and indistinct spatial audio. To provide accurate guidance for latent diffusion models, we introduce the SpatialSonic model utilizing spatial-aware encoders and azimuth state matrices to reveal reasonable spatial guidance. By leveraging spatial guidance, our unified model not only achieves the objective of generating immersive and controllable spatial audio from text and image but also enables interactive audio generation during inference. Finally, under fair settings, we conduct subjective and objective evaluations on simulated and real-world data to compare our approach with prevailing methods. The results demonstrate the effectiveness of our method, highlighting its capability to generate spatial audio that adheres to physical rules.

【2】 Reproducible Machine Learning-based Voice Pathology Detection: Introducing the Pitch Difference Feature
标题: 可重复的基于机器学习的语音病理检测:引入音调差异特征
作者: Jan Vrba, Jakub Steinbach, Tomáš Jirsa, Laura Verde, Roberta De Fazio, Noriyasu Homma, Yuwen Zeng, Key Ichiji, Lukáš Hájek, Zuzana Sedláková, Jan Mareš
备注:33 pages, 8 figures, code repository: this https URL
链接:点击下载PDF文件
摘要:在这项研究中,我们提出了一套强大的功能,来自当代的实践在语音病理检测的彻底研究。该功能集基于声学手工制作功能的组合。此外,我们引入音高差异作为一个新的功能。我们结合这个功能集,包含数据从公开的Saarbrucken语音数据库(SVD),预处理使用K均值合成少数过采样技术算法,以解决类不平衡。 此外,我们应用多个ML模型作为二进制分类器。我们使用支持向量机,k-近邻,朴素贝叶斯,决策树,随机森林和AdaBoost分类器。为了确定最佳的分类方法,我们对各个分类器和特征子部分的可行超参数进行网格搜索。 通过SVD数据库上语音病理检测的未加权平均召回率来衡量,我们的方法已经达到了最先进的性能。我们故意省略了准确性,因为与上述指标相比,在数据不平衡的情况下,它是一个高度偏差的指标。通过重复分层交叉验证消除对结果的潜在高估,进一步增强了结果。这一进展表明ML方法的临床部署具有巨大潜力,为客观检查语音病理提供了一个有价值的工具。为了支持我们的主张,我们提供了一个公开可用的GitHub存储库,其DOI为10.5281 zenodo.13771573。最后,我们提供了改革清单。摘要:In this study, we propose a robust set of features derived from a thorough research of contemporary practices in voice pathology detection. The feature set is based on the combination of acoustic handcrafted features. Additionally, we introduce pitch difference as a novel feature. We combine this feature set, containing data from the publicly available Saarbr "ucken Voice Database (SVD), with preprocessing using the K-Means Synthetic Minority Over-Sampling Technique algorithm to address class imbalance. Moreover, we applied multiple ML models as binary classifiers. We utilized support vector machine, k-nearest neighbors, naive Bayes, decision tree, random forest and AdaBoost classifiers. To determine the best classification approach, we performed grid search on feasible hyperparameters of respective classifiers and subsections of features. Our approach has achieved the state-of-the-art performance, measured by unweighted average recall in voice pathology detection on SVD database. We intentionally omit accuracy as it is highly biased metric in case of unbalanced data compared to aforementioned metrics. The results are further enhanced by eliminating the potential overestimation of the results with repeated stratified cross-validation. This advancement demonstrates significant potential for the clinical deployment of ML methods, offering a valuable tool for an objective examination of voice pathologies. To support our claims, we provide a publicly available GitHub repository with DOI 10.5281 zenodo.13771573. Finally, we provide REFORMS checklist.

【3】 Do we need more complex representations for structure? A comparison of note duration representation for Music Transformers
标题: 我们是否需要更复杂的结构表示?音乐Transformer音符持续时间表示的比较
作者: Gabriel Souza, Flavio Figueiredo, Alexei Machado, Deborah Guimarães
备注:Presented at the Music for Machine Learning Workshop with ECMLPKDD. To be published by Springer
链接:点击下载PDF文件
摘要:近年来,深度学习在创造性计算方面取得了令人瞩目的成果。在音乐方面,一个可行的音乐生成模型是基于Transformer的模型。然而,虽然Transformers模型在音乐生成中很受欢迎,但它们通常依赖于注释的结构信息。在这项工作中,我们询问,如果现成的音乐Transformer模型执行以及结构相似性度量,仅使用未注释的数据库信息。我们表明,最常见的表示轻微的调整产生小,但显着的改善。我们还主张,寻找更好的未注释的音乐表示比产生大量的策划和注释数据更具成本效益。摘要:In recent years, deep learning has achieved formidable results in creative computing. When it comes to music, one viable model for music generation are Transformer based models. However, while transformers models are popular for music generation, they often rely on annotated structural information. In this work, we inquire if the off-the-shelf Music Transformer models perform just as well on structural similarity metrics using only unannotated MIDI information. We show that a slight tweak to the most common representation yields small but significant improvements. We also advocate that searching for better unannotated musical representations is more cost-effective than producing large amounts of curated and annotated data.

【4】 Everyday Speech in the Indian Subcontinent
标题: 印度次大陆的日常演讲
作者: Utkarsh Pathak (1), Chandra Sai Krishna Gunda (1), Sujitha Sathiyamoorthy (1), Keshav Agarwal (1), Hema A. Murthy (1) ((1) Indian Institute of Technology, Madras)
备注:5 Pages, 1 Figure, Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:印度有1369种语言,其中22种是官方语言。大约有13种不同的脚本用于表示这些语言。为了解决端到端(E2E)多语言合成框架中单元词汇量大的问题,提出了一种基于语音学的通用标签集(CLS)。这减少了合成器的占地面积,也使快速适应新的语言,有类似的音位结构,提供语言脚本属于同一个家庭。在本文中,我们提供了新的见解语音合成,脚本属于一个家庭,而音位来自另一个。首先将印度语文本转换为CLS,然后使用与该语言的语音匹配的合成器。质量类似于母语的梵语和康卡尼语获得零适应数据,分别使用卡纳达语和马拉地语合成器。此外,这种方法还可以在13种印度语言和英语之间以给定的母语语音进行无缝代码切换。摘要:India has 1369 languages of which 22 are official. About 13 different scripts are used to represent these languages. A Common Label Set (CLS) was developed based on phonetics to address the issue of large vocabulary of units required in the End to End (E2E) framework for multilingual synthesis. This reduced the footprint of the synthesizer and also enabled fast adaptation to new languages which had similar phonotactics, provided language scripts belonged to the same family. In this paper, we provide new insights into speech synthesis, where the script belongs to one family, while the phonotactics comes from another. Indian language text is first converted to CLS, and then a synthesizer that matches the phonotactics of the language is used. Quality akin to that of a native speaker is obtained for Sanskrit and Konkani with zero adaptation data, using Kannada and Marathi synthesizers respectively. Further, this approach also lends itself seamless code switching across 13 Indian languages and English in a given native speaker's voice.

【5】 Generative Deep Learning and Signal Processing for Data Augmentation of Cardiac Auscultation Signals: Improving Model Robustness Using Synthetic Audio
标题: 用于心脏听诊信号数据增强的生成性深度学习和信号处理:使用合成音频提高模型稳健性
作者: Leigh Abbott, Milan Marocchi, Matthew Fynn, Yue Rong, Sven Nordholm
备注:21 pages, 8 figures, 10 tables
链接:点击下载PDF文件
摘要:准确地解释心脏听诊信号在诊断和管理心血管疾病中起着至关重要的作用。然而,标记数据的缺乏抑制了分类模型的训练。研究人员已经转向结合信号处理的生成式深度学习技术,以增强现有数据并改进心脏听诊分类模型,以克服这一挑战。然而,以前的研究主要集中在模型性能,而不是模型的鲁棒性。在这种情况下,鲁棒性被定义为分布内和分布外的性能,如马修的相关系数。这项工作表明,可以使用增强的数据集来训练更强大的异常心音分类器。增强包括传统的音频方法和使用WaveGrad和DiffWave扩散模型有条件生成的合成音频的创建。研究发现,当用这种增强的数据集训练基于卷积神经网络的分类模型时,分布内和分布外的性能都可以在各种数据集上得到改善。随着性能的提高,不仅包括准确性,还包括平衡的准确性和马修的相关系数,增强的数据集显着有助于解决不平衡数据集的问题。这反过来又有助于提供更通用和鲁棒的分类器。摘要:Accurately interpreting cardiac auscultation signals plays a crucial role in diagnosing and managing cardiovascular diseases. However, the paucity of labelled data inhibits classification models' training. Researchers have turned to generative deep learning techniques combined with signal processing to augment the existing data and improve cardiac auscultation classification models to overcome this challenge. However, the primary focus of prior studies has been on model performance as opposed to model robustness. Robustness, in this case, is defined as both the in-distribution and out-of-distribution performance by measures such as Matthew's correlation coefficient. This work shows that more robust abnormal heart sound classifiers can be trained using an augmented dataset. The augmentations consist of traditional audio approaches and the creation of synthetic audio conditionally generated using the WaveGrad and DiffWave diffusion models. It is found that both the in-distribution and out-of-distribution performance can be improved over various datasets when training a convolutional neural network-based classification model with this augmented dataset. With the performance increase encompassing not only accuracy but also balanced accuracy and Matthew's correlation coefficient, an augmented dataset significantly contributes to resolving issues of imbalanced datasets. This, in turn, helps provide a more general and robust classifier.

【6】 M2M-Gen: A Multimodal Framework for Automated Background Music Generation in Japanese Manga Using Large Language Models
标题: M2 M-Gen:使用大型语言模型自动生成日本漫画背景音乐的多模式框架
作者: Megha Sharma, Muhammad Taimoor Haseeb, Gus Xia, Yoshimasa Tsuruoka
链接:点击下载PDF文件
摘要:本文介绍了M2M Gen,一个多模态的框架,为日本漫画背景音乐生成量身定制。这项任务的主要挑战是缺乏可用的数据集或基线。为了解决这些挑战,我们提出了一个自动化的音乐生成管道,产生输入漫画书的背景音乐。首先,我们使用漫画中的对话来检测场景边界,并使用场景中的人物面部进行情感分类。然后,我们使用GPT4o将此低级场景信息转换为高级音乐指令。在场景信息和音乐指令的条件下,GPT 4o的另一个实例生成页面级别的音乐标题,以将文本引导到音乐模型。这产生了与漫画不断发展的叙事相一致的音乐。M2M Gen的有效性通过广泛的主观评估得到证实,展示了其生成更高质量、更相关和一致的音乐的能力,与我们的基线相比,可以补充特定场景。摘要:This paper introduces M2M Gen, a multi modal framework for generating background music tailored to Japanese manga. The key challenges in this task are the lack of an available dataset or a baseline. To address these challenges, we propose an automated music generation pipeline that produces background music for an input manga book. Initially, we use the dialogues in a manga to detect scene boundaries and perform emotion classification using the characters faces within a scene. Then, we use GPT4o to translate this low level scene information into a high level music directive. Conditioned on the scene information and the music directive, another instance of GPT 4o generates page level music captions to guide a text to music model. This produces music that is aligned with the mangas evolving narrative. The effectiveness of M2M Gen is confirmed through extensive subjective evaluations, showcasing its capability to generate higher quality, more relevant and consistent music that complements specific scenes when compared to our baselines.

【7】 Prompt Tuning for Audio Deepfake Detection: Computationally Efficient Test-time Domain Adaptation with Limited Target Dataset
标题: 音频Deepfake检测的提示调整:具有有限目标数据集的计算高效的测试时域自适应
作者: Hideyuki Oiso, Yuto Matsunaga, Kazuya Kakizaki, Taiki Miyagawa
备注:Accepted at Interspeech 2024. Hideyuki Oiso and Yuto Matsunaga contributed equally
链接:点击下载PDF文件
摘要:我们研究了音频深度伪造检测(ADD)的测试时域自适应,解决了三个挑战:(i)源-目标域间隙,(ii)有限的目标数据集大小,以及(iii)高计算成本。我们提出了一个ADD方法,使用插件风格的提示调整。它通过将其与最先进的Transformer模型和 或其他微调方法无缝集成来弥合域差距,从而提高目标数据的性能(挑战(i))。此外,我们的方法可以适合小的目标数据集,因为它不需要大量的额外参数(挑战(ii))。该功能还有助于提高计算效率,对抗ADD中通常与大规模预训练模型相关的高计算成本(挑战(iii))。我们的结论是,及时调整域差距下的ADD提出了一个很有前途的途径,以最小的目标数据和可忽略不计的额外计算负担,以提高精度。摘要:We study test-time domain adaptation for audio deepfake detection (ADD), addressing three challenges: (i) source-target domain gaps, (ii) limited target dataset size, and (iii) high computational costs. We propose an ADD method using prompt tuning in a plug-in style. It bridges domain gaps by integrating it seamlessly with state-of-the-art transformer models and or with other fine-tuning methods, boosting their performance on target data (challenge (i)). In addition, our method can fit small target datasets because it does not require a large number of extra parameters (challenge (ii)). This feature also contributes to computational efficiency, countering the high computational costs typically associated with large-scale pre-trained models in ADD (challenge (iii)). We conclude that prompt tuning for ADD under domain gaps presents a promising avenue for enhancing accuracy with minimal target data and negligible extra computational burden.

【8】 LEAD Dataset: How Can Labels for Sound Event Detection Vary Depending on Annotators?
标题: LEAD数据集:声音事件检测的标签如何根据注释者而变化?
作者: Naoki Koga, Yoshiaki Bando, Keisuke Imoto
备注:Accepted to APSIPA ASC 2024
链接:点击下载PDF文件
摘要:在本文中,我们介绍了一个用于声音事件检测(LEAD)数据集的大规模注释器标签,该数据集用于更好地理解声音事件检测(SED)中强标签的变化。在SED中,收集大规模强标签非常耗时,并且在大多数情况下,多个工作人员将注释划分以创建单个数据集。一般来说,由多个注释器创建的强标签在声音事件的类型和时间起始 偏移方面具有很大的变化。通过多个工作者的注释,唯一地确定强标签是相当困难的,因为数据集包含可能被误认为相似类的声音和时间起始 偏移难以区分的声音。如果SED的强标签因注释者而异,则在多个注释者创建的数据集上训练的SED模型将有偏差。此外,如果训练数据和评估数据之间的注释器不同,则存在无法正确评估模型的风险。为了研究强标签的变化,我们发布了LEAD数据集,该数据集为20个不同注释器注释的每个剪辑提供了不同的强标签。LEAD数据集允许我们研究强标签在注释者之间的变化,并考虑对强标签变化具有鲁棒性的SED模型。LEAD数据集由分配给TUT Sound Events 2016 2017,TUT Acoustic Scenes 2016和URBAN-SED的声音片段的强标签组成。我们还分析了LEAD数据集中强标签的变化,并提供了对变化的见解。摘要:In this paper, we introduce a LargE-scale Annotator's labels for sound event Detection (LEAD) dataset, which is the dataset used to gain a better understanding of the variation in strong labels in sound event detection (SED). In SED, it is very time-consuming to collect large-scale strong labels, and in most cases, multiple workers divide up the annotations to create a single dataset. In general, strong labels created by multiple annotators have large variations in the type of sound events and temporal onset offset. Through the annotations of multiple workers, uniquely determining the strong label is quite difficult because the dataset contains sounds that can be mistaken for similar classes and sounds whose temporal onset offset is difficult to distinguish. If the strong labels of SED vary greatly depending on the annotator, the SED model trained on a dataset created by multiple annotators will be biased. Moreover, if annotators differ between training and evaluation data, there is a risk that the model cannot be evaluated correctly. To investigate the variation in strong labels, we release the LEAD dataset, which provides distinct strong labels for each clip annotated by 20 different annotators. The LEAD dataset allows us to investigate how strong labels vary from annotator to annotator and consider SED models that are robust to the variation of strong labels. The LEAD dataset consists of strong labels assigned to sound clips from TUT Sound Events 2016 2017, TUT Acoustic Scenes 2016, and URBAN-SED. We also analyze variations in the strong labels in the LEAD dataset and provide insights into the variations.

【9】 Objective Measurements of Voice Quality
标题: 语音质量的客观测量
作者: Hira Dhamyal, Rita Singh
链接:点击下载PDF文件
摘要:人类声音的质量在音乐、言语治疗和交流等各个领域都发挥着重要作用,但它缺乏一个普遍接受的客观定义。相反,语音质量是指使用主观描述符,如“粗糙”,“呼吸”等,尽管这种主观性,跨学科的广泛研究已经将这些语音质量与扬声器的特定信息联系起来,如健康,生理特征等。目前用于语音分析的机器学习方法依赖于数据驱动的分析,由于其定性性质,没有完全结合这些已建立的相关性。本文的目的是客观地量化语音质量合成公式化的陈述,从过去的研究结果,相关的语音质量信号处理指标。我们基于科学文献,基于25个信号属性引入了24个语音子质量的公式。这些公式进行了测试,对数据集与主观标记的语音质量,证明其有效性。摘要:The quality of human voice plays an important role across various fields like music, speech therapy, and communication, yet it lacks a universally accepted, objective definition. Instead, voice quality is referred to using subjective descriptors like "rough", "breathy" etc. Despite this subjectivity, extensive research across disciplines has linked these voice qualities to specific information about the speaker, such as health, physiological traits, and others. Current machine learning approaches for voice profiling rely on data-driven analysis without fully incorporating these established correlations, due to their qualitative nature. This paper aims to objectively quantify voice quality by synthesizing formulaic representations from past findings that correlate voice qualities to signal-processing metrics. We introduce formulae for 24 voice sub-qualities based on 25 signal properties, grounded in scientific literature. These formulae are tested against datasets with subjectively labeled voice qualities, demonstrating their validity.

【10】 Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
标题: 利用多模式多尺度上下文建模实现对话文本到语音的重点渲染
作者: Rui Liu, Zhenqi Jia, Jie Yang, Yifan Hu, Haizhou Li
备注:submitted to IEEE Transaction
链接:点击下载PDF文件
摘要:会话式文语转换(CTTS)旨在以恰当的风格准确地表达会话中的话语,受到越来越多的关注。在认识到CTTS任务的重要性的同时,由于会话强调数据集的稀缺和上下文理解的困难,先前的研究没有彻底研究语音强调表达,这对于传达人机交互场景中的潜在意图和态度至关重要。本文提出了一种新的CTTS模型的强调渲染方案ER-CTTS,该方案包括两个主要部分:1)同时考虑文本和声学语境,通过全局和局部语义建模来全面理解会话语境;(2)深入整合多模态、多尺度语境,研究语境对当前话语重点表达的影响。最后,将推断出的强调特征馈送到神经语音合成器中以生成会话语音。为了解决数据稀缺问题,我们在现有的会话数据集(DailyTalk)上创建了强调强度注释。客观和主观的评价表明,我们的模型优于基线模型在会话设置内的重点渲染。代码和音频示例可在https: github.com CodeStoreTTS ER-CTTS上获得。摘要:Conversational Text-to-Speech (CTTS) aims to accurately express an utterance with the appropriate style within a conversational setting, which attracts more attention nowadays. While recognizing the significance of the CTTS task, prior studies have not thoroughly investigated speech emphasis expression, which is essential for conveying the underlying intention and attitude in human-machine interaction scenarios, due to the scarcity of conversational emphasis datasets and the difficulty in context understanding. In this paper, we propose a novel Emphasis Rendering scheme for the CTTS model, termed ER-CTTS, that includes two main components: 1) we simultaneously take into account textual and acoustic contexts, with both global and local semantic modeling to understand the conversation context comprehensively; 2) we deeply integrate multi-modal and multi-scale context to learn the influence of context on the emphasis expression of the current utterance. Finally, the inferred emphasis feature is fed into the neural speech synthesizer to generate conversational speech. To address data scarcity, we create emphasis intensity annotations on the existing conversational dataset (DailyTalk). Both objective and subjective evaluations suggest that our model outperforms the baseline models in emphasis rendering within a conversational setting. The code and audio samples are available at https: github.com CodeStoreTTS ER-CTTS.

【11】 DRCap: Decoding CLAP Latents with Retrieval-augmented Generation for Zero-shot Audio Captioning
标题: DRCAP:利用Zero-Shot音频字幕的检索增强生成解码CLAP潜伏
作者: Xiquan Li, Wenxi Chen, Ziyang Ma, Xuenan Xu, Yuzhe Liang, Zhisheng Zheng, Qiuqiang Kong, Xie Chen
链接:点击下载PDF文件
摘要:虽然自动音频字幕(AAC)已经取得了显着的进展,传统的完全监督AAC模型仍然面临着两个关键挑战:需要昂贵的音频-文本对数据进行训练,以及跨域传输时性能下降。为了克服这些限制,我们提出了DRCap,一个数据高效和灵活的zero-shot音频字幕系统,需要纯文本数据进行训练,可以快速适应新的领域,而无需额外的微调。DRCap集成了对比语言音频预训练(CLAP)模型和大语言模型(LLM)作为其骨干。在训练过程中,该模型使用CLAP中的固定文本编码器预测地面实况字幕,而在推理过程中,文本编码器被音频编码器替换,以零拍摄(zero-shot)方式生成音频剪辑的字幕。为了减轻CLAP模型的模态差距,我们同时使用编码器端的投影策略和解码器端的检索增强生成策略。具体来说,音频嵌入首先投影到文本嵌入支持,以吸收CLAP的联合多模态空间内的广泛的语义信息。与此同时,从一个搜索引擎中检索到的类似标题作为提示输入,以指导LLM,并结合外部知识,以充分利用其强大的生成能力。在预测CLAP嵌入和检索到的相似字幕的条件下,该模型能够产生更准确和语义丰富的文本描述。通过定制文本嵌入支持和标题匹配到目标域,DRCap获得了以免训练方式适应新域的强大能力。实验结果表明,DRCap优于所有其他zero-shot模型在域内的场景,并实现了最先进的性能在跨域的场景。摘要:While automated audio captioning (AAC) has made notable progress, traditional fully supervised AAC models still face two critical challenges: the need for expensive audio-text pair data for training and performance degradation when transferring across domains. To overcome these limitations, we present DRCap, a data-efficient and flexible zero-shot audio captioning system that requires text-only data for training and can quickly adapt to new domains without additional fine-tuning. DRCap integrates a contrastive language-audio pre-training (CLAP) model and a large-language model (LLM) as its backbone. During training, the model predicts the ground-truth caption with a fixed text encoder from CLAP, whereas, during inference, the text encoder is replaced with the audio encoder to generate captions for audio clips in a zero-shot manner. To mitigate the modality gap of the CLAP model, we use both the projection strategy from the encoder side and the retrieval-augmented generation strategy from the decoder side. Specifically, audio embeddings are first projected onto a text embedding support to absorb extensive semantic information within the joint multi-modal space of CLAP. At the same time, similar captions retrieved from a datastore are fed as prompts to instruct the LLM, incorporating external knowledge to take full advantage of its strong generative capability. Conditioned on both the projected CLAP embedding and the retrieved similar captions, the model is able to produce a more accurate and semantically rich textual description. By tailoring the text embedding support and the caption datastore to the target domain, DRCap acquires a robust ability to adapt to new domains in a training-free manner. Experimental results demonstrate that DRCap outperforms all other zero-shot models in in-domain scenarios and achieves state-of-the-art performance in cross-domain scenarios.

【12】 ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance
标题: ExpGest:使用扩散模型和混合音频文本引导的表达者生成
作者: Yongkang Cheng, Mingjiang Liang, Shaoli Huang, Jifeng Ning, Wei Liu
备注:Accepted by ICME 2024
链接:点击下载PDF文件
摘要:现有的手势生成方法主要集中在基于音频特征的上半身手势,忽略了语音内容、情感和运动。这些限制导致僵硬、机械的手势无法传达音频内容的真正含义。我们介绍ExpGest,一个新颖的框架,利用同步的文本和音频信息来生成富有表现力的全身手势。与AdaIN或one-hot编码方法不同,我们设计了一个噪声情感分类器,用于优化对抗性方向噪声,避免旋律失真并将结果引导到指定的情感。此外,在潜在空间中对齐语义和手势提供了更好的泛化能力。ExpGest是一个基于扩散模型的手势生成框架,它是第一个尝试提供混合生成模式,包括音频驱动的手势和文本形状的运动。实验表明,我们的框架有效地从组合的文本驱动的运动和音频诱导的手势数据集学习,初步结果表明,ExpGest实现了更有表现力,自然,可控的全局运动扬声器相比,国家的最先进的模型。摘要:Existing gesture generation methods primarily focus on upper body gestures based on audio features, neglecting speech content, emotion, and locomotion. These limitations result in stiff, mechanical gestures that fail to convey the true meaning of audio content. We introduce ExpGest, a novel framework leveraging synchronized text and audio information to generate expressive full-body gestures. Unlike AdaIN or one-hot encoding methods, we design a noise emotion classifier for optimizing adversarial direction noise, avoiding melody distortion and guiding results towards specified emotions. Moreover, aligning semantic and gestures in the latent space provides better generalization capabilities. ExpGest, a diffusion model-based gesture generation framework, is the first attempt to offer mixed generation modes, including audio-driven gestures and text-shaped motion. Experiments show that our framework effectively learns from combined text-driven motion and audio-induced gesture datasets, and preliminary results demonstrate that ExpGest achieves more expressive, natural, and controllable global motion in speakers compared to state-of-the-art models.

【13】 Towards the Synthesis of Non-speech Vocalizations
标题: 迈向非言语发声的合成
作者: Enjamamul Hoq, Ifeoma Nwogu
链接:点击下载PDF文件
摘要:在这份报告中,我们专注于使用DiffWave框架无条件生成婴儿哭声,该框架在从噪声中生成高质量音频方面表现出很大的潜力。我们使用两个不同的婴儿哭声数据集:Baby Chillanto和deBarbaro cry数据集。这些数据集用于训练DiffWave模型,以生成保持高保真度和多样性的新哭声。这里的重点是DiffWave处理无条件生成任务的能力。摘要:In this report, we focus on the unconditional generation of infant cry sounds using the DiffWave framework, which has shown great promise in generating high-quality audio from noise. We use two distinct datasets of infant cries: the Baby Chillanto and the deBarbaro cry dataset. These datasets are used to train the DiffWave model to generate new cry sounds that maintain high fidelity and diversity. The focus here is on DiffWave's capability to handle the unconditional generation task.

【14】 AuD-Former: A Hierarchical Transformer Network for Multimodal Audio-Based Disease Prediction
标题: AuD-Former:用于多模式音频疾病预测的分层Transformer网络
作者: Jinjin Cai, Ruiqi Wang, Dezhong Zhao, Ziqin Yuan, Victoria McKenna, Aaron Friedman, Rachel Foot, Susan Storey, Ryan Boente, Sudip Vhaduri, Byung-Cheol Min
链接:点击下载PDF文件
摘要:基于音频的疾病预测正在成为传统医学诊断方法的一个有前途的补充,有助于早期,方便和非侵入性的疾病检测和预防。多模态融合,它集成了来自生物声学模态内或跨生物声学模态的各个领域的特征,已被证明在提高诊断性能方面是有效的。然而,该领域中大多数现有的方法采用单方面的融合策略,只专注于模态内或模态间的融合。这种方法限制了对不同声学特征域和生物声学模态的互补性质的充分利用。此外,模态特定和模态共享空间内的潜在依赖性的不充分和孤立的探索限制了他们管理多模态数据中固有异质性的能力。为了填补这些空白,我们提出了AuD-Former,这是一种分层Transformer网络,旨在用于基于音频的一般多模态疾病预测。具体来说,我们无缝集成内模态和模态间融合的层次化方式,并熟练地编码必要的模态内和模态间的互补相关性,分别。综合实验表明,AuD-Former在预测三种疾病方面达到了最先进的性能:COVID-19,帕金森病和病理性构音障碍,在基于音频的疾病预测任务的广泛背景下展示了其有前途的潜力。此外,广泛的消融研究和定性分析强调了我们模型中每个主要组件的显著益处。摘要:Audio-based disease prediction is emerging as a promising supplement to traditional medical diagnosis methods, facilitating early, convenient, and non-invasive disease detection and prevention. Multimodal fusion, which integrates features from various domains within or across bio-acoustic modalities, has proven effective in enhancing diagnostic performance. However, most existing methods in the field employ unilateral fusion strategies that focus solely on either intra-modal or inter-modal fusion. This approach limits the full exploitation of the complementary nature of diverse acoustic feature domains and bio-acoustic modalities. Additionally, the inadequate and isolated exploration of latent dependencies within modality-specific and modality-shared spaces curtails their capacity to manage the inherent heterogeneity in multimodal data. To fill these gaps, we propose AuD-Former, a hierarchical transformer network designed for general multimodal audio-based disease prediction. Specifically, we seamlessly integrate intra-modal and inter-modal fusion in a hierarchical manner and proficiently encode the necessary intra-modal and inter-modal complementary correlations, respectively. Comprehensive experiments demonstrate that AuD-Former achieves state-of-the-art performance in predicting three diseases: COVID-19, Parkinson's disease, and pathological dysarthria, showcasing its promising potential in a broad context of audio-based disease prediction tasks. Additionally, extensive ablation studies and qualitative analyses highlight the significant benefits of each main component within our model.

【15】 Quantum-Trained Convolutional Neural Network for Deepfake Audio Detection
标题: 用于Deepfake音频检测的量子训练卷积神经网络
作者: Chu-Hsuan Abraham Lin, Chen-Yu Liu, Samuel Yen-Chi Chen, Kuan-Cheng Chen
链接:点击下载PDF文件
摘要:Deepfake技术的兴起对隐私、安全和信息完整性构成了重大挑战,特别是在音频和多媒体内容方面。本文介绍了一种量子训练卷积神经网络(QT-CNN)框架,旨在利用量子机器学习(QML)的计算能力来增强对deepfake音频的检测。QT-CNN采用混合量子-经典方法,将量子神经网络(QNN)与经典神经架构集成,以优化训练效率,同时减少可训练参数的数量。我们的方法采用了一种新的量子到经典的参数映射,有效地利用量子态,以提高模型的表达能力,实现高达70%的参数减少相比,经典模型,而不影响精度。数据预处理涉及提取基本音频特征、标签编码、特征缩放以及构建用于鲁棒模型评估的序列数据集。实验结果表明,QT-CNN实现了与传统CNN相当的性能,在不同配置的QNN块的训练和测试阶段保持了高精度。QT框架能够在保持性能的同时减少计算开销,这突显了它在深度伪造检测和其他资源受限场景中的实际应用潜力。这项工作突出了将量子计算集成到人工智能中的实际好处,为推进deepfake检测技术提供了一种可扩展的高效方法。摘要:The rise of deepfake technologies has posed significant challenges to privacy, security, and information integrity, particularly in audio and multimedia content. This paper introduces a Quantum-Trained Convolutional Neural Network (QT-CNN) framework designed to enhance the detection of deepfake audio, leveraging the computational power of quantum machine learning (QML). The QT-CNN employs a hybrid quantum-classical approach, integrating Quantum Neural Networks (QNNs) with classical neural architectures to optimize training efficiency while reducing the number of trainable parameters. Our method incorporates a novel quantum-to-classical parameter mapping that effectively utilizes quantum states to enhance the expressive power of the model, achieving up to 70% parameter reduction compared to classical models without compromising accuracy. Data pre-processing involved extracting essential audio features, label encoding, feature scaling, and constructing sequential datasets for robust model evaluation. Experimental results demonstrate that the QT-CNN achieves comparable performance to traditional CNNs, maintaining high accuracy during training and testing phases across varying configurations of QNN blocks. The QT framework's ability to reduce computational overhead while maintaining performance underscores its potential for real-world applications in deepfake detection and other resource-constrained scenarios. This work highlights the practical benefits of integrating quantum computing into artificial intelligence, offering a scalable and efficient approach to advancing deepfake detection technologies.

【16】 In-Materia Speech Recognition
标题: 内部语音识别
作者: Mohamadreza Zolfagharinejad, Julian Büchel, Lorenzo Cassola, Sachin Kinge, Ghazi Sarwat Syed, Abu Sebastian, Wilfred G. van der Wiel
链接:点击下载PDF文件
摘要:随着去中心化计算的兴起,例如物联网、自动驾驶和个性化医疗保健,在边缘有效处理时间相关信号变得越来越重要:就在收集时间数据的地方,避免耗时、不安全和昂贵的通信与集中式计算设施(或云)。然而,现代处理器通常无法满足边缘系统的受限功率和时间预算,因为其架构(冯诺依曼瓶颈)或域转换(模数转换和时频转换)所施加的内在限制。在这里,我们提出了一个边缘时间信号处理器的基础上两个在物质计算系统的特征提取和分类,达到软件级的准确性为96.2%的TI-46字的语音识别任务。首先,一个非线性的,室温掺杂剂网络处理单元(DNPU)层实现模拟,时域特征提取从原始音频信号,类似于人类耳蜗。其次,由忆阻交叉阵列组成的模拟内存计算(AIMC)芯片实现了一个紧凑的神经网络,该神经网络在提取的特征上进行训练以进行分类。由于DNPU特征提取消耗100 nW,基于AIMC的分类每个乘法累加运算的潜力小于10 fJ,我们的研究结果为通过物质计算硬件提高异构智能边缘处理器的紧凑性,效率和性能提供了一种有希望的途径。摘要:With the rise of decentralized computing, as in the Internet of Things, autonomous driving, and personalized healthcare, it is increasingly important to process time-dependent signals at the edge efficiently: right at the place where the temporal data are collected, avoiding time-consuming, insecure, and costly communication with a centralized computing facility (or cloud). However, modern-day processors often cannot meet the restrained power and time budgets of edge systems because of intrinsic limitations imposed by their architecture (von Neumann bottleneck) or domain conversions (analogue-to-digital and time-to-frequency). Here, we propose an edge temporal-signal processor based on two in-materia computing systems for both feature extraction and classification, reaching a software-level accuracy of 96.2% for the TI-46-Word speech-recognition task. First, a nonlinear, room-temperature dopant-network-processing-unit (DNPU) layer realizes analogue, time-domain feature extraction from the raw audio signals, similar to the human cochlea. Second, an analogue in-memory computing (AIMC) chip, consisting of memristive crossbar arrays, implements a compact neural network trained on the extracted features for classification. With the DNPU feature extraction consuming 100s nW and AIMC-based classification having the potential for less than 10 fJ per multiply-accumulate operation, our findings offer a promising avenue for advancing the compactness, efficiency, and performance of heterogeneous smart edge processors through in-materia computing hardware.

【17】 SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
标题: SLAM-AAC:通过LLM通过解释增强和CLAP-Refine增强音频字幕
作者: Wenxi Chen, Ziyang Ma, Xiquan Li, Xuenan Xu, Yuzhe Liang, Zhisheng Zheng, Kai Yu, Xie Chen
链接:点击下载PDF文件
摘要:自动音频字幕(AAC)旨在为输入的音频信号生成自然的文本描述。音频预训练模型和大型语言模型(LLM)的最新进展显着增强了音频理解和文本推理能力,使AAC的改进成为可能。在本文中,我们提出了SLAM-AAC,以进一步增强AAC的释义增强和CLAP-Refine通过LLM。我们的方法使用自监督EAT模型来提取细粒度的音频表示,然后通过轻量级线性层将其与文本嵌入对齐。字幕生成LLM使用LoRA适配器进行有效的微调。从机器翻译中的回译方法中汲取灵感,我们在预训练期间实现了释义增强以扩展Clotho数据集。这种策略有助于缓解稀缺的音频文本对的限制,并从一个小的音频剪辑集生成更多样化的字幕。在推理过程中,我们引入了即插即用的CLAP-Refine策略,以充分利用多个解码输出,类似于语音识别中的n-best rescoring策略。使用CLAP模型进行音频-文本相似度计算,我们可以选择由多个搜索波束生成的与输入音频最匹配的文本描述。实验结果表明,SLAM-AAC在Clotho V2和AudioCaps上实现了最先进的性能,超过了以前的主流模型。摘要:Automated Audio Captioning (AAC) aims to generate natural textual descriptions for input audio signals. Recent progress in audio pre-trained models and large language models (LLMs) has significantly enhanced audio understanding and textual reasoning capabilities, making improvements in AAC possible. In this paper, we propose SLAM-AAC to further enhance AAC with paraphrasing augmentation and CLAP-Refine through LLMs. Our approach uses the self-supervised EAT model to extract fine-grained audio representations, which are then aligned with textual embeddings via lightweight linear layers. The caption generation LLM is efficiently fine-tuned using the LoRA adapter. Drawing inspiration from the back-translation method in machine translation, we implement paraphrasing augmentation to expand the Clotho dataset during pre-training. This strategy helps alleviate the limitation of scarce audio-text pairs and generates more diverse captions from a small set of audio clips. During inference, we introduce the plug-and-play CLAP-Refine strategy to fully exploit multiple decoding outputs, akin to the n-best rescoring strategy in speech recognition. Using the CLAP model for audio-text similarity calculation, we could select the textual descriptions generated by multiple searching beams that best match the input audio. Experimental results show that SLAM-AAC achieves state-of-the-art performance on Clotho V2 and AudioCaps, surpassing previous mainstream models.

【18】 Enhancing Infant Crying Detection with Gradient Boosting for Improved Emotional and Mental Health Diagnostics
标题: 通过梯度增强婴儿哭声检测,以改善情绪和心理健康诊断
作者: Kyunghun Lee, Lauren M. Henry, Eleanor Hansen, Elizabeth Tandilashvili, Lauren S. Wakschlag, Elizabeth Norton, Daniel S. Pine, Melissa A. Brotman, Francisco Pereira
链接:点击下载PDF文件
摘要:婴儿哭声可以作为各种生理和情绪状态的重要指标。本文介绍了一种全面的方法,用于检测音频数据中的婴儿哭声。我们集成了Meta的Wav2Vec与传统的音频特征,如梅尔频率倒谱系数(MFCC),色度和频谱对比度,采用梯度提升机(GBM)哭分类。我们在真实世界的数据集上验证了我们的方法,与现有方法相比,表现出显着的性能改进。摘要:Infant crying can serve as a crucial indicator of various physiological and emotional states. This paper introduces a comprehensive approach for detecting infant cries within audio data. We integrate Meta's Wav2Vec with traditional audio features, such as Mel-frequency cepstral coefficients (MFCCs), chroma, and spectral contrast, employing Gradient Boosting Machines (GBM) for cry classification. We validate our approach on a real-world dataset, demonstrating significant performance improvements over existing methods.


eess.AS音频处理
【1】 In-Materia Speech Recognition
标题: 内部语音识别
作者: Mohamadreza Zolfagharinejad, Julian Büchel, Lorenzo Cassola, Sachin Kinge, Ghazi Sarwat Syed, Abu Sebastian, Wilfred G. van der Wiel
链接:点击下载PDF文件
摘要:随着去中心化计算的兴起,例如物联网、自动驾驶和个性化医疗保健,在边缘有效处理时间相关信号变得越来越重要:就在收集时间数据的地方,避免耗时、不安全和昂贵的通信与集中式计算设施(或云)。然而,现代处理器通常无法满足边缘系统的受限功率和时间预算,因为其架构(冯诺依曼瓶颈)或域转换(模数转换和时频转换)所施加的内在限制。在这里,我们提出了一个边缘时间信号处理器的基础上两个在物质计算系统的特征提取和分类,达到软件级的准确性为96.2%的TI-46字的语音识别任务。首先,一个非线性的,室温掺杂剂网络处理单元(DNPU)层实现模拟,时域特征提取从原始音频信号,类似于人类耳蜗。其次,由忆阻交叉阵列组成的模拟内存计算(AIMC)芯片实现了一个紧凑的神经网络,该神经网络在提取的特征上进行训练以进行分类。由于DNPU特征提取消耗100 nW,基于AIMC的分类每个乘法累加运算的潜力小于10 fJ,我们的研究结果为通过物质计算硬件提高异构智能边缘处理器的紧凑性,效率和性能提供了一种有希望的途径。摘要:With the rise of decentralized computing, as in the Internet of Things, autonomous driving, and personalized healthcare, it is increasingly important to process time-dependent signals at the edge efficiently: right at the place where the temporal data are collected, avoiding time-consuming, insecure, and costly communication with a centralized computing facility (or cloud). However, modern-day processors often cannot meet the restrained power and time budgets of edge systems because of intrinsic limitations imposed by their architecture (von Neumann bottleneck) or domain conversions (analogue-to-digital and time-to-frequency). Here, we propose an edge temporal-signal processor based on two in-materia computing systems for both feature extraction and classification, reaching a software-level accuracy of 96.2% for the TI-46-Word speech-recognition task. First, a nonlinear, room-temperature dopant-network-processing-unit (DNPU) layer realizes analogue, time-domain feature extraction from the raw audio signals, similar to the human cochlea. Second, an analogue in-memory computing (AIMC) chip, consisting of memristive crossbar arrays, implements a compact neural network trained on the extracted features for classification. With the DNPU feature extraction consuming 100s nW and AIMC-based classification having the potential for less than 10 fJ per multiply-accumulate operation, our findings offer a promising avenue for advancing the compactness, efficiency, and performance of heterogeneous smart edge processors through in-materia computing hardware.

【2】 Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
标题: 我们可以基于Zero-Shot语音情感识别来估计购买意图吗?
作者: Ryotaro Nagase, Takashi Sumiyoshi, Natsuo Yamashita, Kota Dohi, Yohei Kawaguchi
备注:5 pages, 3 figures, accepted for APSIPA 2024 ASC
链接:点击下载PDF文件
摘要:本文提出了一种zero-shot语音情感识别(SER)方法,估计情感没有预先定义在SER模型训练。传统的方法仅限于识别由单个单词定义的情感。此外,我们有动机去识别未知的两极情绪,比如“我想买--我不想买”。为了允许模型自由地使用句子来定义类并估计未知的双极情感,我们提出的方法通过引入多类和多任务设置来扩展对比语言音频预训练(CLAP)框架。我们还将购买意向作为一种双极性情感,研究了该模型在zero-shot估计购买意向时的性能。实验结果表明,该方法的zero-shot估计结果与有监督学习训练的模型的估计结果处于同一水平。摘要:This paper proposes a zero-shot speech emotion recognition (SER) method that estimates emotions not previously defined in the SER model training. Conventional methods are limited to recognizing emotions defined by a single word. Moreover, we have the motivation to recognize unknown bipolar emotions such as I want to buy - I do not want to buy.'' In order to allow the model to define classes using sentences freely and to estimate unknown bipolar emotions, our proposed method expands upon the contrastive language-audio pre-training (CLAP) framework by introducing multi-class and multi-task settings. We also focus on purchase intention as a bipolar emotion and investigate the model's performance to zero-shot estimate it. This study is the first attempt to estimate purchase intention from speech directly. Experiments confirm that the results of zero-shot estimation by the proposed method are at the same level as those of the model trained by supervised learning.

【3】 SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
标题: SLAM-AAC:通过LLM通过解释增强和CLAP-Refine增强音频字幕
作者: Wenxi Chen, Ziyang Ma, Xiquan Li, Xuenan Xu, Yuzhe Liang, Zhisheng Zheng, Kai Yu, Xie Chen
链接:点击下载PDF文件
摘要:自动音频字幕(AAC)旨在为输入的音频信号生成自然的文本描述。音频预训练模型和大型语言模型(LLM)的最新进展显着增强了音频理解和文本推理能力,使AAC的改进成为可能。在本文中,我们提出了SLAM-AAC,以进一步增强AAC的释义增强和CLAP-Refine通过LLM。我们的方法使用自监督EAT模型来提取细粒度的音频表示,然后通过轻量级线性层将其与文本嵌入对齐。字幕生成LLM使用LoRA适配器进行有效的微调。从机器翻译中的回译方法中汲取灵感,我们在预训练期间实现了释义增强以扩展Clotho数据集。这种策略有助于缓解稀缺的音频文本对的限制,并从一个小的音频剪辑集生成更多样化的字幕。在推理过程中,我们引入了即插即用的CLAP-Refine策略,以充分利用多个解码输出,类似于语音识别中的n-best rescoring策略。使用CLAP模型进行音频-文本相似度计算,我们可以选择由多个搜索波束生成的与输入音频最匹配的文本描述。实验结果表明,SLAM-AAC在Clotho V2和AudioCaps上实现了最先进的性能,超过了以前的主流模型。摘要:Automated Audio Captioning (AAC) aims to generate natural textual descriptions for input audio signals. Recent progress in audio pre-trained models and large language models (LLMs) has significantly enhanced audio understanding and textual reasoning capabilities, making improvements in AAC possible. In this paper, we propose SLAM-AAC to further enhance AAC with paraphrasing augmentation and CLAP-Refine through LLMs. Our approach uses the self-supervised EAT model to extract fine-grained audio representations, which are then aligned with textual embeddings via lightweight linear layers. The caption generation LLM is efficiently fine-tuned using the LoRA adapter. Drawing inspiration from the back-translation method in machine translation, we implement paraphrasing augmentation to expand the Clotho dataset during pre-training. This strategy helps alleviate the limitation of scarce audio-text pairs and generates more diverse captions from a small set of audio clips. During inference, we introduce the plug-and-play CLAP-Refine strategy to fully exploit multiple decoding outputs, akin to the n-best rescoring strategy in speech recognition. Using the CLAP model for audio-text similarity calculation, we could select the textual descriptions generated by multiple searching beams that best match the input audio. Experimental results show that SLAM-AAC achieves state-of-the-art performance on Clotho V2 and AudioCaps, surpassing previous mainstream models.

【4】 Enhancing Infant Crying Detection with Gradient Boosting for Improved Emotional and Mental Health Diagnostics
标题: 通过梯度增强婴儿哭声检测,以改善情绪和心理健康诊断
作者: Kyunghun Lee, Lauren M. Henry, Eleanor Hansen, Elizabeth Tandilashvili, Lauren S. Wakschlag, Elizabeth Norton, Daniel S. Pine, Melissa A. Brotman, Francisco Pereira
链接:点击下载PDF文件
摘要:婴儿哭声可以作为各种生理和情绪状态的重要指标。本文介绍了一种全面的方法,用于检测音频数据中的婴儿哭声。我们集成了Meta的Wav2Vec与传统的音频特征,如梅尔频率倒谱系数(MFCC),色度和频谱对比度,采用梯度提升机(GBM)哭分类。我们在真实世界的数据集上验证了我们的方法,与现有方法相比,表现出显着的性能改进。摘要:Infant crying can serve as a crucial indicator of various physiological and emotional states. This paper introduces a comprehensive approach for detecting infant cries within audio data. We integrate Meta's Wav2Vec with traditional audio features, such as Mel-frequency cepstral coefficients (MFCCs), chroma, and spectral contrast, employing Gradient Boosting Machines (GBM) for cry classification. We validate our approach on a real-world dataset, demonstrating significant performance improvements over existing methods.

【5】 Everyday Speech in the Indian Subcontinent
标题: 印度次大陆的日常演讲
作者: Utkarsh Pathak (1), Chandra Sai Krishna Gunda (1), Sujitha Sathiyamoorthy (1), Keshav Agarwal (1), Hema A. Murthy (1) ((1) Indian Institute of Technology, Madras)
备注:5 Pages, 1 Figure, Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:印度有1369种语言,其中22种是官方语言。大约有13种不同的脚本用于表示这些语言。为了解决端到端(E2E)多语言合成框架中单元词汇量大的问题,提出了一种基于语音学的通用标签集(CLS)。这减少了合成器的占地面积,也使快速适应新的语言,有类似的音位结构,提供语言脚本属于同一个家庭。在本文中,我们提供了新的见解语音合成,脚本属于一个家庭,而音位来自另一个。首先将印度语文本转换为CLS,然后使用与该语言的语音匹配的合成器。质量类似于母语的梵语和康卡尼语获得零适应数据,分别使用卡纳达语和马拉地语合成器。此外,这种方法还可以在13种印度语言和英语之间以给定的母语语音进行无缝代码切换。摘要:India has 1369 languages of which 22 are official. About 13 different scripts are used to represent these languages. A Common Label Set (CLS) was developed based on phonetics to address the issue of large vocabulary of units required in the End to End (E2E) framework for multilingual synthesis. This reduced the footprint of the synthesizer and also enabled fast adaptation to new languages which had similar phonotactics, provided language scripts belonged to the same family. In this paper, we provide new insights into speech synthesis, where the script belongs to one family, while the phonotactics comes from another. Indian language text is first converted to CLS, and then a synthesizer that matches the phonotactics of the language is used. Quality akin to that of a native speaker is obtained for Sanskrit and Konkani with zero adaptation data, using Kannada and Marathi synthesizers respectively. Further, this approach also lends itself seamless code switching across 13 Indian languages and English in a given native speaker's voice.

【6】 Generative Deep Learning and Signal Processing for Data Augmentation of Cardiac Auscultation Signals: Improving Model Robustness Using Synthetic Audio
标题: 用于心脏听诊信号数据增强的生成性深度学习和信号处理:使用合成音频提高模型稳健性
作者: Leigh Abbott, Milan Marocchi, Matthew Fynn, Yue Rong, Sven Nordholm
备注:21 pages, 8 figures, 10 tables
链接:点击下载PDF文件
摘要:准确地解释心脏听诊信号在诊断和管理心血管疾病中起着至关重要的作用。然而,标记数据的缺乏抑制了分类模型的训练。研究人员已经转向结合信号处理的生成式深度学习技术,以增强现有数据并改进心脏听诊分类模型,以克服这一挑战。然而,以前的研究主要集中在模型性能,而不是模型的鲁棒性。在这种情况下,鲁棒性被定义为分布内和分布外的性能,如马修的相关系数。这项工作表明,可以使用增强的数据集来训练更强大的异常心音分类器。增强包括传统的音频方法和使用WaveGrad和DiffWave扩散模型有条件生成的合成音频的创建。研究发现,当用这种增强的数据集训练基于卷积神经网络的分类模型时,分布内和分布外的性能都可以在各种数据集上得到改善。随着性能的提高,不仅包括准确性,还包括平衡的准确性和马修的相关系数,增强的数据集显着有助于解决不平衡数据集的问题。这反过来又有助于提供更通用和鲁棒的分类器。摘要:Accurately interpreting cardiac auscultation signals plays a crucial role in diagnosing and managing cardiovascular diseases. However, the paucity of labelled data inhibits classification models' training. Researchers have turned to generative deep learning techniques combined with signal processing to augment the existing data and improve cardiac auscultation classification models to overcome this challenge. However, the primary focus of prior studies has been on model performance as opposed to model robustness. Robustness, in this case, is defined as both the in-distribution and out-of-distribution performance by measures such as Matthew's correlation coefficient. This work shows that more robust abnormal heart sound classifiers can be trained using an augmented dataset. The augmentations consist of traditional audio approaches and the creation of synthetic audio conditionally generated using the WaveGrad and DiffWave diffusion models. It is found that both the in-distribution and out-of-distribution performance can be improved over various datasets when training a convolutional neural network-based classification model with this augmented dataset. With the performance increase encompassing not only accuracy but also balanced accuracy and Matthew's correlation coefficient, an augmented dataset significantly contributes to resolving issues of imbalanced datasets. This, in turn, helps provide a more general and robust classifier.

【7】 M2M-Gen: A Multimodal Framework for Automated Background Music Generation in Japanese Manga Using Large Language Models
标题: M2 M-Gen:使用大型语言模型自动生成日本漫画背景音乐的多模式框架
作者: Megha Sharma, Muhammad Taimoor Haseeb, Gus Xia, Yoshimasa Tsuruoka
链接:点击下载PDF文件
摘要:本文介绍了M2M Gen,一个多模态的框架,为日本漫画背景音乐生成量身定制。这项任务的主要挑战是缺乏可用的数据集或基线。为了解决这些挑战,我们提出了一个自动化的音乐生成管道,产生输入漫画书的背景音乐。首先,我们使用漫画中的对话来检测场景边界,并使用场景中的人物面部进行情感分类。然后,我们使用GPT4o将此低级场景信息转换为高级音乐指令。在场景信息和音乐指令的条件下,GPT 4o的另一个实例生成页面级别的音乐标题,以将文本引导到音乐模型。这产生了与漫画不断发展的叙事相一致的音乐。M2M Gen的有效性通过广泛的主观评估得到证实,展示了其生成更高质量,更相关和一致的音乐的能力,与我们的基线相比,这些音乐补充了特定的场景。摘要:This paper introduces M2M Gen, a multi modal framework for generating background music tailored to Japanese manga. The key challenges in this task are the lack of an available dataset or a baseline. To address these challenges, we propose an automated music generation pipeline that produces background music for an input manga book. Initially, we use the dialogues in a manga to detect scene boundaries and perform emotion classification using the characters faces within a scene. Then, we use GPT4o to translate this low level scene information into a high level music directive. Conditioned on the scene information and the music directive, another instance of GPT 4o generates page level music captions to guide a text to music model. This produces music that is aligned with the mangas evolving narrative. The effectiveness of M2M Gen is confirmed through extensive subjective evaluations, showcasing its capability to generate higher quality, more relevant and consistent music that complements specific scenes when compared to our baselines.

【8】 Prompt Tuning for Audio Deepfake Detection: Computationally Efficient Test-time Domain Adaptation with Limited Target Dataset
标题: 音频Deepfake检测的提示调整:具有有限目标数据集的计算高效的测试时域自适应
作者: Hideyuki Oiso, Yuto Matsunaga, Kazuya Kakizaki, Taiki Miyagawa
备注:Accepted at Interspeech 2024. Hideyuki Oiso and Yuto Matsunaga contributed equally
链接:点击下载PDF文件
摘要:我们研究了音频深度伪造检测(ADD)的测试时域自适应,解决了三个挑战:(i)源-目标域间隙,(ii)有限的目标数据集大小,以及(iii)高计算成本。我们提出了一个ADD方法,使用插件风格的提示调整。它通过将其与最先进的Transformer模型和 或其他微调方法无缝集成来弥合域差距,从而提高目标数据的性能(挑战(i))。此外,我们的方法可以适合小的目标数据集,因为它不需要大量的额外参数(挑战(ii))。该功能还有助于提高计算效率,从而抵消ADD中通常与大规模预训练模型相关的高计算成本(挑战(iii))。我们的结论是,及时调整域差距下的ADD提出了一个很有前途的途径,以最小的目标数据和可忽略不计的额外计算负担,以提高精度。摘要:We study test-time domain adaptation for audio deepfake detection (ADD), addressing three challenges: (i) source-target domain gaps, (ii) limited target dataset size, and (iii) high computational costs. We propose an ADD method using prompt tuning in a plug-in style. It bridges domain gaps by integrating it seamlessly with state-of-the-art transformer models and or with other fine-tuning methods, boosting their performance on target data (challenge (i)). In addition, our method can fit small target datasets because it does not require a large number of extra parameters (challenge (ii)). This feature also contributes to computational efficiency, countering the high computational costs typically associated with large-scale pre-trained models in ADD (challenge (iii)). We conclude that prompt tuning for ADD under domain gaps presents a promising avenue for enhancing accuracy with minimal target data and negligible extra computational burden.

【9】 LEAD Dataset: How Can Labels for Sound Event Detection Vary Depending on Annotators?
标题: LEAD数据集:声音事件检测的标签如何根据注释者而变化?
作者: Naoki Koga, Yoshiaki Bando, Keisuke Imoto
备注:Accepted to APSIPA ASC 2024
链接:点击下载PDF文件
摘要:在本文中,我们介绍了一个用于声音事件检测(LEAD)数据集的大规模注释器标签,该数据集用于更好地理解声音事件检测(SED)中强标签的变化。在SED中,收集大规模强标签非常耗时,并且在大多数情况下,多个工作人员将注释划分为单个数据集。一般来说,由多个注释器创建的强标签在声音事件的类型和时间起始 偏移方面具有很大的变化。通过多个工作者的注释,唯一地确定强标签是相当困难的,因为数据集包含可能被误认为相似类的声音和时间起始 偏移难以区分的声音。如果SED的强标签因注释者而异,则在多个注释者创建的数据集上训练的SED模型将有偏差。此外,如果训练数据和评估数据之间的注释器不同,则存在无法正确评估模型的风险。为了研究强标签的变化,我们发布了LEAD数据集,该数据集为20个不同注释器注释的每个剪辑提供了不同的强标签。LEAD数据集允许我们研究强标签在注释者之间的变化,并考虑对强标签变化具有鲁棒性的SED模型。LEAD数据集由分配给TUT Sound Events 2016 2017,TUT Acoustic Scenes 2016和URBAN-SED的声音片段的强标签组成。我们还分析了LEAD数据集中强标签的变化,并提供了对变化的见解。摘要:In this paper, we introduce a LargE-scale Annotator's labels for sound event Detection (LEAD) dataset, which is the dataset used to gain a better understanding of the variation in strong labels in sound event detection (SED). In SED, it is very time-consuming to collect large-scale strong labels, and in most cases, multiple workers divide up the annotations to create a single dataset. In general, strong labels created by multiple annotators have large variations in the type of sound events and temporal onset offset. Through the annotations of multiple workers, uniquely determining the strong label is quite difficult because the dataset contains sounds that can be mistaken for similar classes and sounds whose temporal onset offset is difficult to distinguish. If the strong labels of SED vary greatly depending on the annotator, the SED model trained on a dataset created by multiple annotators will be biased. Moreover, if annotators differ between training and evaluation data, there is a risk that the model cannot be evaluated correctly. To investigate the variation in strong labels, we release the LEAD dataset, which provides distinct strong labels for each clip annotated by 20 different annotators. The LEAD dataset allows us to investigate how strong labels vary from annotator to annotator and consider SED models that are robust to the variation of strong labels. The LEAD dataset consists of strong labels assigned to sound clips from TUT Sound Events 2016 2017, TUT Acoustic Scenes 2016, and URBAN-SED. We also analyze variations in the strong labels in the LEAD dataset and provide insights into the variations.

【10】 Objective Measurements of Voice Quality
标题: 语音质量的客观测量
作者: Hira Dhamyal, Rita Singh
链接:点击下载PDF文件
摘要:人类声音的质量在音乐、言语治疗和交流等各个领域都发挥着重要作用,但它缺乏一个普遍接受的客观定义。相反,语音质量是指使用主观描述符,如“粗糙”,“呼吸”等,尽管这种主观性,跨学科的广泛研究已经将这些语音质量与扬声器的特定信息联系起来,如健康,生理特征等。目前用于语音分析的机器学习方法依赖于数据驱动的分析,由于其定性性质,没有完全结合这些已建立的相关性。本文的目的是客观地量化语音质量合成公式化的陈述,从过去的研究结果,相关的语音质量信号处理指标。我们介绍了24个语音子质量公式的基础上25个信号特性,接地在科学文献中。这些公式进行了测试,对数据集与主观标记的语音质量,证明其有效性。摘要:The quality of human voice plays an important role across various fields like music, speech therapy, and communication, yet it lacks a universally accepted, objective definition. Instead, voice quality is referred to using subjective descriptors like "rough", "breathy" etc. Despite this subjectivity, extensive research across disciplines has linked these voice qualities to specific information about the speaker, such as health, physiological traits, and others. Current machine learning approaches for voice profiling rely on data-driven analysis without fully incorporating these established correlations, due to their qualitative nature. This paper aims to objectively quantify voice quality by synthesizing formulaic representations from past findings that correlate voice qualities to signal-processing metrics. We introduce formulae for 24 voice sub-qualities based on 25 signal properties, grounded in scientific literature. These formulae are tested against datasets with subjectively labeled voice qualities, demonstrating their validity.

【11】 Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
标题: 利用多模式多尺度上下文建模实现对话文本到语音的重点渲染
作者: Rui Liu, Zhenqi Jia, Jie Yang, Yifan Hu, Haizhou Li
备注:submitted to IEEE Transaction
链接:点击下载PDF文件
摘要:会话式文语转换(CTTS)旨在以恰当的风格准确地表达会话中的话语,受到越来越多的关注。在认识到CTTS任务的重要性的同时,由于会话强调数据集的稀缺和上下文理解的困难,先前的研究没有彻底研究语音强调表达,这对于传达人机交互场景中的潜在意图和态度至关重要。本文提出了一种新的CTTS模型的强调渲染方案ER-CTTS,该方案包括两个主要部分:1)同时考虑文本和声学语境,通过全局和局部语义建模来全面理解会话语境;(2)深入整合多模态、多尺度语境,研究语境对当前话语重点表达的影响。最后,将推断出的强调特征馈送到神经语音合成器中以生成会话语音。为了解决数据稀缺问题,我们在现有的会话数据集(DailyTalk)上创建了强调强度注释。客观和主观的评价表明,我们的模型优于基线模型在会话设置内的重点渲染。代码和音频示例可在https: github.com CodeStoreTTS ER-CTTS上获得。摘要:Conversational Text-to-Speech (CTTS) aims to accurately express an utterance with the appropriate style within a conversational setting, which attracts more attention nowadays. While recognizing the significance of the CTTS task, prior studies have not thoroughly investigated speech emphasis expression, which is essential for conveying the underlying intention and attitude in human-machine interaction scenarios, due to the scarcity of conversational emphasis datasets and the difficulty in context understanding. In this paper, we propose a novel Emphasis Rendering scheme for the CTTS model, termed ER-CTTS, that includes two main components: 1) we simultaneously take into account textual and acoustic contexts, with both global and local semantic modeling to understand the conversation context comprehensively; 2) we deeply integrate multi-modal and multi-scale context to learn the influence of context on the emphasis expression of the current utterance. Finally, the inferred emphasis feature is fed into the neural speech synthesizer to generate conversational speech. To address data scarcity, we create emphasis intensity annotations on the existing conversational dataset (DailyTalk). Both objective and subjective evaluations suggest that our model outperforms the baseline models in emphasis rendering within a conversational setting. The code and audio samples are available at https: github.com CodeStoreTTS ER-CTTS.

【12】 DRCap: Decoding CLAP Latents with Retrieval-augmented Generation for Zero-shot Audio Captioning
标题: DRCAP:利用Zero-Shot音频字幕的检索增强生成解码CLAP潜伏
作者: Xiquan Li, Wenxi Chen, Ziyang Ma, Xuenan Xu, Yuzhe Liang, Zhisheng Zheng, Qiuqiang Kong, Xie Chen
链接:点击下载PDF文件
摘要:虽然自动音频字幕(AAC)已经取得了显着的进展,传统的完全监督AAC模型仍然面临着两个关键挑战:需要昂贵的音频-文本对数据进行训练,以及跨域传输时性能下降。为了克服这些限制,我们提出了DRCap,一个数据高效和灵活的zero-shot音频字幕系统,需要纯文本数据进行训练,可以快速适应新的领域,而无需额外的微调。DRCap集成了对比语言音频预训练(CLAP)模型和大语言模型(LLM)作为其骨干。在训练期间,模型使用CLAP中的固定文本编码器预测地面实况字幕,而在推断期间,文本编码器被音频编码器替换,以zero-shot方式生成音频片段的字幕。为了减轻CLAP模型的模态差距,我们使用的投影策略从编码器端和检索增强生成策略从解码器端。具体来说,音频嵌入首先投影到文本嵌入支持,以吸收CLAP的联合多模态空间内的广泛的语义信息。与此同时,从一个搜索引擎中检索到的类似标题作为提示输入,以指导LLM,并结合外部知识,以充分利用其强大的生成能力。在预测CLAP嵌入和检索到的相似字幕的条件下,该模型能够产生更准确和语义丰富的文本描述。通过定制文本嵌入支持和标题匹配到目标域,DRCap获得了以免训练方式适应新域的强大能力。实验结果表明,DRCap优于所有其他zero-shot模型在域内的场景,并实现了最先进的性能在跨域的场景。摘要:While automated audio captioning (AAC) has made notable progress, traditional fully supervised AAC models still face two critical challenges: the need for expensive audio-text pair data for training and performance degradation when transferring across domains. To overcome these limitations, we present DRCap, a data-efficient and flexible zero-shot audio captioning system that requires text-only data for training and can quickly adapt to new domains without additional fine-tuning. DRCap integrates a contrastive language-audio pre-training (CLAP) model and a large-language model (LLM) as its backbone. During training, the model predicts the ground-truth caption with a fixed text encoder from CLAP, whereas, during inference, the text encoder is replaced with the audio encoder to generate captions for audio clips in a zero-shot manner. To mitigate the modality gap of the CLAP model, we use both the projection strategy from the encoder side and the retrieval-augmented generation strategy from the decoder side. Specifically, audio embeddings are first projected onto a text embedding support to absorb extensive semantic information within the joint multi-modal space of CLAP. At the same time, similar captions retrieved from a datastore are fed as prompts to instruct the LLM, incorporating external knowledge to take full advantage of its strong generative capability. Conditioned on both the projected CLAP embedding and the retrieved similar captions, the model is able to produce a more accurate and semantically rich textual description. By tailoring the text embedding support and the caption datastore to the target domain, DRCap acquires a robust ability to adapt to new domains in a training-free manner. Experimental results demonstrate that DRCap outperforms all other zero-shot models in in-domain scenarios and achieves state-of-the-art performance in cross-domain scenarios.

【13】 Automatic Speech Recognition with BERT and CTC Transformers: A Review
标题: 使用BERT和CIC Transformers的自动语音识别:评论
作者: Noussaiba Djeffal, Hamza Kheddar, Djamel Addou, Ahmed Cherif Mazari, Yassine Himeur
Journal-ref:2023 2nd International Conference on Electronics, Energy and Measurement (IC2EM)
链接:点击下载PDF文件
摘要:这篇综述论文提供了一个全面的分析,在自动语音识别(ASR)的最新进展与双向编码器表示从Transformers BERT和连接时间分类(CTC)Transformers。本文首先介绍了ASR的基本概念和面临的挑战,然后介绍了BERT和CTC Transformers的结构及其在ASR中的潜在应用。本文回顾了几项研究,使用这些模型的语音识别任务,并讨论了所获得的结果。此外,本文强调了这些模型的局限性,并概述了进一步研究的潜在领域。总之,这篇综述为研究人员和从业人员提供了有价值的见解,他们对BERT和CTC Transformers的ASR感兴趣。摘要:This review paper provides a comprehensive analysis of recent advances in automatic speech recognition (ASR) with bidirectional encoder representations from transformers BERT and connectionist temporal classification (CTC) transformers. The paper first introduces the fundamental concepts of ASR and discusses the challenges associated with it. It then explains the architecture of BERT and CTC transformers and their potential applications in ASR. The paper reviews several studies that have used these models for speech recognition tasks and discusses the results obtained. Additionally, the paper highlights the limitations of these models and outlines potential areas for further research. All in all, this review provides valuable insights for researchers and practitioners who are interested in ASR with BERT and CTC transformers.

【14】 ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance
标题: ExpGest:使用扩散模型和混合音频文本引导的表达者生成
作者: Yongkang Cheng, Mingjiang Liang, Shaoli Huang, Jifeng Ning, Wei Liu
备注:Accepted by ICME 2024
链接:点击下载PDF文件
摘要:现有的手势生成方法主要集中在基于音频特征的上半身手势,忽略了语音内容、情感和运动。这些限制导致僵硬、机械的手势无法传达音频内容的真正含义。我们介绍ExpGest,一个新颖的框架,利用同步的文本和音频信息来生成富有表现力的全身手势。与AdaIN或one-hot编码方法不同,我们设计了一个噪声情感分类器,用于优化对抗性方向噪声,避免旋律失真并将结果引导到指定的情感。此外,在潜在空间中对齐语义和手势提供了更好的泛化能力。ExpGest是一个基于扩散模型的手势生成框架,它是第一个尝试提供混合生成模式,包括音频驱动的手势和文本形状的运动。实验表明,我们的框架有效地从组合的文本驱动的运动和音频诱导的手势数据集学习,初步结果表明,ExpGest实现了更有表现力,自然,可控的全局运动扬声器相比,国家的最先进的模型。摘要:Existing gesture generation methods primarily focus on upper body gestures based on audio features, neglecting speech content, emotion, and locomotion. These limitations result in stiff, mechanical gestures that fail to convey the true meaning of audio content. We introduce ExpGest, a novel framework leveraging synchronized text and audio information to generate expressive full-body gestures. Unlike AdaIN or one-hot encoding methods, we design a noise emotion classifier for optimizing adversarial direction noise, avoiding melody distortion and guiding results towards specified emotions. Moreover, aligning semantic and gestures in the latent space provides better generalization capabilities. ExpGest, a diffusion model-based gesture generation framework, is the first attempt to offer mixed generation modes, including audio-driven gestures and text-shaped motion. Experiments show that our framework effectively learns from combined text-driven motion and audio-induced gesture datasets, and preliminary results demonstrate that ExpGest achieves more expressive, natural, and controllable global motion in speakers compared to state-of-the-art models.

【15】 Towards the Synthesis of Non-speech Vocalizations
标题: 迈向非言语发声的合成
作者: Enjamamul Hoq, Ifeoma Nwogu
链接:点击下载PDF文件
摘要:在这份报告中,我们专注于使用DiffWave框架无条件生成婴儿哭声,该框架在从噪声中生成高质量音频方面表现出很大的潜力。我们使用两个不同的婴儿哭声数据集:Baby Chillanto和deBarbaro cry数据集。这些数据集用于训练DiffWave模型,以生成保持高保真度和多样性的新哭声。这里的重点是DiffWave处理无条件生成任务的能力。摘要:In this report, we focus on the unconditional generation of infant cry sounds using the DiffWave framework, which has shown great promise in generating high-quality audio from noise. We use two distinct datasets of infant cries: the Baby Chillanto and the deBarbaro cry dataset. These datasets are used to train the DiffWave model to generate new cry sounds that maintain high fidelity and diversity. The focus here is on DiffWave's capability to handle the unconditional generation task.

【16】 AuD-Former: A Hierarchical Transformer Network for Multimodal Audio-Based Disease Prediction
标题: AuD-Former:用于多模式音频疾病预测的分层Transformer网络
作者: Jinjin Cai, Ruiqi Wang, Dezhong Zhao, Ziqin Yuan, Victoria McKenna, Aaron Friedman, Rachel Foot, Susan Storey, Ryan Boente, Sudip Vhaduri, Byung-Cheol Min
链接:点击下载PDF文件
摘要:基于音频的疾病预测正在成为传统医学诊断方法的一个有前途的补充,有助于早期,方便和非侵入性的疾病检测和预防。多模态融合,它集成了来自生物声学模态内或跨生物声学模态的各个领域的特征,已被证明在提高诊断性能方面是有效的。然而,该领域中大多数现有的方法采用单方面的融合策略,只专注于模态内或模态间的融合。这种方法限制了对不同声学特征域和生物声学模态的互补性质的充分利用。此外,模态特定和模态共享空间内的潜在依赖性的不充分和孤立的探索限制了他们管理多模态数据中固有异质性的能力。为了填补这些空白,我们提出了AuD-Former,这是一种分层Transformer网络,旨在用于基于音频的通用多模式疾病预测。具体来说,我们无缝集成内模态和模态间融合的层次化方式,并熟练地编码必要的模态内和模态间的互补相关性,分别。综合实验表明,AuD-Former在预测三种疾病方面达到了最先进的性能:COVID-19,帕金森病和病理性构音障碍,在基于音频的疾病预测任务的广泛背景下展示了其有前途的潜力。此外,广泛的消融研究和定性分析强调了我们模型中每个主要组件的显著益处。摘要:Audio-based disease prediction is emerging as a promising supplement to traditional medical diagnosis methods, facilitating early, convenient, and non-invasive disease detection and prevention. Multimodal fusion, which integrates features from various domains within or across bio-acoustic modalities, has proven effective in enhancing diagnostic performance. However, most existing methods in the field employ unilateral fusion strategies that focus solely on either intra-modal or inter-modal fusion. This approach limits the full exploitation of the complementary nature of diverse acoustic feature domains and bio-acoustic modalities. Additionally, the inadequate and isolated exploration of latent dependencies within modality-specific and modality-shared spaces curtails their capacity to manage the inherent heterogeneity in multimodal data. To fill these gaps, we propose AuD-Former, a hierarchical transformer network designed for general multimodal audio-based disease prediction. Specifically, we seamlessly integrate intra-modal and inter-modal fusion in a hierarchical manner and proficiently encode the necessary intra-modal and inter-modal complementary correlations, respectively. Comprehensive experiments demonstrate that AuD-Former achieves state-of-the-art performance in predicting three diseases: COVID-19, Parkinson's disease, and pathological dysarthria, showcasing its promising potential in a broad context of audio-based disease prediction tasks. Additionally, extensive ablation studies and qualitative analyses highlight the significant benefits of each main component within our model.

【17】 Quantum-Trained Convolutional Neural Network for Deepfake Audio Detection
标题: 用于Deepfake音频检测的量子训练卷积神经网络
作者: Chu-Hsuan Abraham Lin, Chen-Yu Liu, Samuel Yen-Chi Chen, Kuan-Cheng Chen
链接:点击下载PDF文件
摘要:Deepfake技术的兴起对隐私、安全和信息完整性构成了重大挑战,特别是在音频和多媒体内容方面。本文介绍了一种量子训练卷积神经网络(QT-CNN)框架,旨在利用量子机器学习(QML)的计算能力来增强对deepfake音频的检测。QT-CNN采用混合量子-经典方法,将量子神经网络(QNN)与经典神经架构集成,以优化训练效率,同时减少可训练参数的数量。我们的方法采用了一种新的量子到经典的参数映射,有效地利用量子态,以提高模型的表达能力,实现高达70%的参数减少相比,经典模型,而不影响精度。数据预处理涉及提取基本音频特征、标签编码、特征缩放以及构建用于鲁棒模型评估的序列数据集。实验结果表明,QT-CNN实现了与传统CNN相当的性能,在不同配置的QNN块的训练和测试阶段保持了高精度。QT框架能够在保持性能的同时减少计算开销,这突显了它在深度伪造检测和其他资源受限场景中的实际应用潜力。这项工作突出了将量子计算集成到人工智能中的实际好处,为推进deepfake检测技术提供了一种可扩展的高效方法。摘要:The rise of deepfake technologies has posed significant challenges to privacy, security, and information integrity, particularly in audio and multimedia content. This paper introduces a Quantum-Trained Convolutional Neural Network (QT-CNN) framework designed to enhance the detection of deepfake audio, leveraging the computational power of quantum machine learning (QML). The QT-CNN employs a hybrid quantum-classical approach, integrating Quantum Neural Networks (QNNs) with classical neural architectures to optimize training efficiency while reducing the number of trainable parameters. Our method incorporates a novel quantum-to-classical parameter mapping that effectively utilizes quantum states to enhance the expressive power of the model, achieving up to 70% parameter reduction compared to classical models without compromising accuracy. Data pre-processing involved extracting essential audio features, label encoding, feature scaling, and constructing sequential datasets for robust model evaluation. Experimental results demonstrate that the QT-CNN achieves comparable performance to traditional CNNs, maintaining high accuracy during training and testing phases across varying configurations of QNN blocks. The QT framework's ability to reduce computational overhead while maintaining performance underscores its potential for real-world applications in deepfake detection and other resource-constrained scenarios. This work highlights the practical benefits of integrating quantum computing into artificial intelligence, offering a scalable and efficient approach to advancing deepfake detection technologies.


机器翻译,仅供参考