今日论文合集:cs.SD语音16篇,eess.AS音频处理22篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Unlocking Potential in Pre-Trained Music Language Models for Versatile Multi-Track Music Arrangement
标题: 释放预训练音乐语言模型的潜力,用于多功能多轨音乐编曲
作者:Longshen Ou,Jingwei Zhao,Ziyu Wang,Gus Xia,Ye Wang
备注:Submitted to AAAI 2025
链接:点击下载PDF文件
摘要:大型语言模型在各个领域都表现出了显著的能力,包括符号音乐生成。然而,利用这些预先训练的模型进行可控的音乐编排任务,每个任务都需要不同形式的音乐信息作为控制,仍然是一个新的挑战。在本文中,我们提出了一个统一的序列到序列的框架,使微调的符号音乐语言模型的多个多轨道的安排任务,包括乐队安排,钢琴减少,鼓安排,和语音分离。我们的实验表明,所提出的方法始终实现更高的音乐质量相比,在所有四个任务的任务特定的基线。此外,通过对探测分析的额外实验,我们表明预训练阶段为模型提供了理解音乐条件的基本知识,这是很难仅仅通过特定任务的微调获得的。摘要:Large language models have shown significant capabilities across various domains, including symbolic music generation. However, leveraging these pre-trained models for controllable music arrangement tasks, each requiring different forms of musical information as control, remains a novel challenge. In this paper, we propose a unified sequence-to-sequence framework that enables the fine-tuning of a symbolic music language model for multiple multi-track arrangement tasks, including band arrangement, piano reduction, drum arrangement, and voice separation. Our experiments demonstrate that the proposed approach consistently achieves higher musical quality compared to task-specific baselines across all four tasks. Furthermore, through additional experiments on probing analysis, we show the pre-training phase equips the model with essential knowledge to understand musical conditions, which is hard to acquired solely through task-specific fine-tuning.

【2】 Knowledge Discovery in Optical Music Recognition: Enhancing Information Retrieval with Instance Segmentation
标题: 光学音乐识别中的知识发现:通过实例分割增强信息检索
作者:Elona Shatri,George Fazekas
备注:8 pages content and one references, accepted version at the International Conference on Knowledge Discovery and Information Retrieval 2024, Porto, Portugal
链接:点击下载PDF文件
摘要:光学音乐识别(OMR)可自动将音乐符号从图像转录为机器可读的格式,如MusicXML、MEI或XML,从而显著降低手动转录的成本和时间。本研究通过使用Mask R-CNN应用实例分割来探索OMR中的知识发现,以增强乐谱中音乐符号的检测和描绘。与光学字符识别(OCR)不同,OMR必须处理通用西方音乐记谱法(CWMN)的复杂语义,其中符号含义取决于形状,位置和上下文。我们的方法利用实例分割来管理音乐符号的密度和重叠,从而促进从乐谱中进行更精确的信息检索。在DoReMi和MUSCIMA++数据集上的评估表明了实质性的改进,我们的方法在密集符号环境中实现了高达59.70%的平均精度(mAP),实现了与对象检测相当的结果。此外,使用传统的计算机视觉技术,我们增加了一个并行的步骤,工作人员检测,以推断所识别的符号的音高。这项研究强调了逐像素分割在推进准确的音乐符号识别中的作用,有助于OMR中的知识发现。我们的研究结果表明,实例分割提供了更精确的音乐符号表示,特别是在人口密集的分数,推进OMR技术。我们公开我们的实施、预处理脚本、训练模型和评估结果,以支持进一步的研究和开发。摘要:Optical Music Recognition (OMR) automates the transcription of musical notation from images into machine-readable formats like MusicXML, MEI, or MIDI, significantly reducing the costs and time of manual transcription. This study explores knowledge discovery in OMR by applying instance segmentation using Mask R-CNN to enhance the detection and delineation of musical symbols in sheet music. Unlike Optical Character Recognition (OCR), OMR must handle the intricate semantics of Common Western Music Notation (CWMN), where symbol meanings depend on shape, position, and context. Our approach leverages instance segmentation to manage the density and overlap of musical symbols, facilitating more precise information retrieval from music scores. Evaluations on the DoReMi and MUSCIMA++ datasets demonstrate substantial improvements, with our method achieving a mean Average Precision (mAP) of up to 59.70 % in dense symbol environments, achieving comparable results to object detection. Furthermore, using traditional computer vision techniques, we add a parallel step for staff detection to infer the pitch for the recognised symbols. This study emphasises the role of pixel-wise segmentation in advancing accurate music symbol recognition, contributing to knowledge discovery in OMR. Our findings indicate that instance segmentation provides more precise representations of musical symbols, particularly in densely populated scores, advancing OMR technology. We make our implementation, pre-processing scripts, trained models, and evaluation results publicly available to support further research and development.

【3】 Speech Recognition Transformers: Topological-lingualism Perspective
标题: 语音识别变形者:话题语言主义视角
作者:Shruti Singh,Muskaan Singh,Virender Kadyan
链接:点击下载PDF文件
摘要:Transformers在各种人工智能任务中取得了巨大的成功。由于我们最近流行的自我注意机制,它捕获了长期的依赖性,在语音处理和识别任务中产生了惊人的结果。本文对面向语音模态的Transformer技术进行了综述。本次调查的主要内容包括:(1)传统ASR的背景,端到端的Transformer生态系统,以及语音Transformers(2)通过语言学范式的语音基础模型,即,单语、双语、多语言和跨语言(3)数据集和语言、声学特征、架构、解码和特定拓扑语言学视角的评估指标(4)用于构建端到端ASR系统的流行语音Transformer工具包。最后,重点讨论了开放的挑战和潜在的研究方向,为社区在这一领域进行进一步的研究。摘要:Transformers have evolved with great success in various artificial intelligence tasks. Thanks to our recent prevalence of self-attention mechanisms, which capture long-term dependency, phenomenal outcomes in speech processing and recognition tasks have been produced. The paper presents a comprehensive survey of transformer techniques oriented in speech modality. The main contents of this survey include (1) background of traditional ASR, end-to-end transformer ecosystem, and speech transformers (2) foundational models in a speech via lingualism paradigm, i.e., monolingual, bilingual, multilingual, and cross-lingual (3) dataset and languages, acoustic features, architecture, decoding, and evaluation metric from a specific topological lingualism perspective (4) popular speech transformer toolkit for building end-to-end ASR systems. Finally, highlight the discussion of open challenges and potential research directions for the community to conduct further research in this domain.

【4】 Morphogenesis of sound creates acoustic rainbows
标题: 声音的形态发生创造了声学彩虹
作者:Rasmus E. Christiansen,Ole Sigmund,Efren Fernandez-Grande
备注:8 Pages, 4 Figures, Supplementary information including text and four movies
链接:点击下载PDF文件
摘要:声音是自然界中许多生物体的基本传感元件,并且多个物种已经进化出产生复杂的声散射和分散现象的有机结构,以明确地发射和感知声音。迄今为止,还没有证明有可能设计出与有机结构中发现的那些结构的性能相媲美的人工散射结构。首先,大多数声音操纵依赖于流体介质中的主动转导,而不是依赖于自然界中经常发现的被动散射原理。在这项工作中,我们利用计算形态合成复杂的节能波长大小的单材料散射结构,被动地将辐射声分解成其空间频谱分量。具体来说,我们量身定制的声学彩虹结构与“以上统一”的效率和声学波长分离器。我们的工作为声场工程的新前沿铺平了道路,在转导,仿生学,能量收集,通信和传感方面具有潜在的应用。摘要:Sound is an essential sensing element for many organisms in nature, and multiple species have evolved organic structures that create complex acoustic scattering and dispersion phenomena to emit and perceive sound unambiguously. To date, it has not proven possible to design artificial scattering structures that rival the performance of those found in organic structures. Contrarily, most sound manipulation relies on active transduction in fluid media rather than relying on passive scattering principles, as are often found in nature. In this work, we utilize computational morphogenesis to synthesize complex energy-efficient wavelength-sized single-material scattering structures that passively decompose radiated sound into its spatio-spectral components. Specifically, we tailor an acoustic rainbow structure with "above unity" efficiency and an acoustic wavelength-splitter. Our work paves the way for a new frontier in sound-field engineering, with potential applications in transduction, bionics, energy harvesting, communications and sensing.

【5】 Deep learning classification system for coconut maturity levels based on acoustic signals
标题: 基于声信号的椰子成熟度深度学习分类系统
作者:June Anne Caladcad,Eduardo Jr Piedad
备注:7 pages, 4 figures, accepted in the 2024 IEEE 12th Region 10 Humanitarian Technology Conference (R10-HTC)
链接:点击下载PDF文件
摘要:随着计算机图像处理、模式识别、信号处理等技术的进步,水果分类的人工方法逐渐被计算机和机械方法所取代。在农业领域,收获后水果的智能分类使智能设备的使用成为可能,这对农民,特别是出口产品产生了直接影响。对于椰子分类,它仍然是传统的过程。本研究提出了一种基于声信号的椰子数据集的分类。为了解决不平衡的数据集,通过听觉和程序音频生成方法进行数据增强技术。未成熟、成熟和过成熟下的音频信号现在分别具有4,050、4,050和5,850个音频信号。为了解决分类系统的更新和分类准确性性能,利用深度学习模型对来自数据生成的所生成的音频信号进行分类。具体来说,RNN和LSTM模型进行了训练和测试,并将它们的性能与Caladcad等人(2020)使用的机器学习方法进行了比较。这两个DL模型表现出令人印象深刻的性能,两者的准确率均为97.42%,并且由于它们的分类性能没有显着差异,因此它们都没有优于另一个。摘要:The advancement of computer image processing, pattern recognition, signal processing, and other technologies has gradually replaced the manual methods of classifying fruit with computer and mechanical methods. In the field of agriculture, the intelligent classification of post-harvested fruit has enabled the use of smart devices that creates a direct impact on farmers, especially on export products. For coconut classification, it remains to be traditional in process. This study presents a classification of the coconut dataset based on acoustic signals. To address the imbalanced dataset, a data augmentation technique was conducted through audiomentation and procedural audio generation methods. Audio signals under premature, mature, and overmature now have 4,050, 4,050, and 5,850 audio signals, respectively. To address the updation of the classification system and the classification accuracy performance, deep learning models were utilized for classifying the generated audio signals from data generation. Specifically, RNN and LSTM models were trained and tested, and their performances were compared with each other and the machine learning methods used by Caladcad et al. (2020). The two DL models showed impressive performance with both having an accuracy of 97.42% and neither of them outperformed the other since there are no significant differences in their classification performance.

【6】 A Functional Trade-off between Prosodic and Semantic Cues in Conveying Sarcasm
标题: 讽刺传达中韵律线索与语义线索的功能权衡
作者:Zhu Li,Xiyuan Gao,Yuqing Zhang,Shekhar Nayak,Matt Coler
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:本研究探讨了讽刺的声学特征,并解开了话语被讽刺使用的倾向和韵律线索之间的相互作用。使用从电视节目中编译的讽刺话语的数据集,我们分析的韵律特征内的话语和关键短语属于三个不同的讽刺类别(嵌入式,命题,和言外之语),不同程度的语义线索存在,并比较它们的中性表达。结果表明,在短语的讽刺意义是突出的语义,韵律线索的相关性较低时,讽刺的意义是不明显的语义,这表明韵律和语义线索之间的权衡讽刺在短语水平。这些研究结果强调了减少依赖韵律调制在语义密集的讽刺表达和微妙的互动,塑造讽刺意图的沟通。摘要:This study investigates the acoustic features of sarcasm and disentangles the interplay between the propensity of an utterance being used sarcastically and the presence of prosodic cues signaling sarcasm. Using a dataset of sarcastic utterances compiled from television shows, we analyze the prosodic features within utterances and key phrases belonging to three distinct sarcasm categories (embedded, propositional, and illocutionary), which vary in the degree of semantic cues present, and compare them to neutral expressions. Results show that in phrases where the sarcastic meaning is salient from the semantics, the prosodic cues are less relevant than when the sarcastic meaning is not evident from the semantics, suggesting a trade-off between prosodic and semantic cues of sarcasm at the phrase level. These findings highlight a lessened reliance on prosodic modulation in semantically dense sarcastic expressions and a nuanced interaction that shapes the communication of sarcastic intent.

【7】 The VoxCeleb Speaker Recognition Challenge: A Retrospective
标题: VoxCeleb演讲者识别挑战:回顾
作者:Jaesung Huh,Joon Son Chung,Arsha Nagrani,Andrew Brown,Jee-weon Jung,Daniel Garcia-Romero,Andrew Zisserman
备注:TASLP 2024
链接:点击下载PDF文件
摘要:VoxCeleb Speaker Recognition Challenges(VoxSRC)是一系列挑战和研讨会,从2019年到2023年每年举办一次。这些挑战主要评估了各种设置下的说话人识别和日记的任务,包括:封闭和开放训练数据;以及用于域适应的监督,自监督和半监督训练。挑战赛还为每项任务和设置提供了公开的培训和评估数据集,每年发布新的测试集。在本文中,我们提供了一个审查这些挑战,包括:他们探讨了什么;挑战参与者开发的方法以及这些方法是如何演变的;以及扬声器验证和diarisation领域的现状。我们在一个共同的评估数据集上绘制了五个阶段的表现进展,并详细分析了每年的特别关注如何影响参与者的表现。本文的目的是研究人员谁想要一个概述的说话人识别和diarisation领域,也在挑战组织者谁想要受益于成功,避免错误的VoxSRC挑战。最后,我们讨论了该领域目前的优势和面临的挑战。项目页面:https: mm.kaist.ac.kr datasets voxceleb voxsrc workshop.html摘要:The VoxCeleb Speaker Recognition Challenges (VoxSRC) were a series of challenges and workshops that ran annually from 2019 to 2023. The challenges primarily evaluated the tasks of speaker recognition and diarisation under various settings including: closed and open training data; as well as supervised, self-supervised, and semi-supervised training for domain adaptation. The challenges also provided publicly available training and evaluation datasets for each task and setting, with new test sets released each year. In this paper, we provide a review of these challenges that covers: what they explored; the methods developed by the challenge participants and how these evolved; and also the current state of the field for speaker verification and diarisation. We chart the progress in performance over the five installments of the challenge on a common evaluation dataset and provide a detailed analysis of how each year's special focus affected participants' performance. This paper is aimed both at researchers who want an overview of the speaker recognition and diarisation field, and also at challenge organisers who want to benefit from the successes and avoid the mistakes of the VoxSRC challenges. We end with a discussion of the current strengths of the field and open challenges. Project page : https: mm.kaist.ac.kr datasets voxceleb voxsrc workshop.html

【8】 Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
标题: 利用自我监督的音频表示进行数据高效的声学场景分类
作者:Yiqiang Cai,Shengchen Li,Xi Shao
备注:Accepted by DCASE Workshop 2024
链接:点击下载PDF文件
摘要:声场景分类(ASC)主要依赖于监督的方法。然而,获取用于训练ASC模型的标记数据通常是昂贵且耗时的。最近,自监督学习(SSL)已经成为一种从未标记音频数据中提取特征的强大方法,使许多下游音频任务受益。本文提出了一个数据高效和低复杂度的ASC系统,利用自监督音频表示提取通用音频数据集。我们引入BEAT,一个音频SSL预训练模型,从AudioSet中提取一般表示。通过大量的实验,它已被证明,自监督音频表示可以帮助实现高ASC精度与有限的标记微调数据。此外,我们发现,集成的SSL模型微调不同的策略有助于进一步提高性能。为了满足低复杂度的要求,我们使用知识蒸馏将自监督知识从大型教师模型转移到高效的学生模型。实验结果表明,教师自我监督有效地提高了学生模型的分类精度。我们性能最好的系统获得了56.7%的平均准确率。摘要:Acoustic scene classification (ASC) predominantly relies on supervised approaches. However, acquiring labeled data for training ASC models is often costly and time-consuming. Recently, self-supervised learning (SSL) has emerged as a powerful method for extracting features from unlabeled audio data, benefiting many downstream audio tasks. This paper proposes a data-efficient and low-complexity ASC system by leveraging self-supervised audio representations extracted from general-purpose audio datasets. We introduce BEATs, an audio SSL pre-trained model, to extract the general representations from AudioSet. Through extensive experiments, it has been demonstrated that the self-supervised audio representations can help to achieve high ASC accuracy with limited labeled fine-tuning data. Furthermore, we find that ensembling the SSL models fine-tuned with different strategies contributes to a further performance improvement. To meet low-complexity requirements, we use knowledge distillation to transfer the self-supervised knowledge from large teacher models to an efficient student model. The experimental results suggest that the self-supervised teachers effectively improve the classification accuracy of the student model. Our best-performing system obtains an average accuracy of 56.7%.

【9】 CoopASD: Cooperative Machine Anomalous Sound Detection with Privacy Concerns
标题: CoopASD:存在隐私问题的协作机器异常声音检测
作者:Anbai Jiang,Yuchen Shi,Pingyi Fan,Wei-Qiang Zhang,Jia Liu
备注:Accepted by GLOBECOM 2024
链接:点击下载PDF文件
摘要:机器异常声音检测(ASD)已成为工业物联网(IIoT)中最有前途的应用之一,因为它在降低故障风险和提高生产效率方面具有前所未有的功效。以前的工作主要是研究机器ASD任务集中设置下。然而,在分散设置下开发ASD系统在实践中至关重要,因为机器数据分散在各个工厂中,并且由于隐私问题,数据不应明确共享。为了使这些工厂能够合作开发可扩展的ASD模型,同时保护他们的隐私,我们提出了一个名为CoopASD的新框架,每个工厂在其本地数据集上训练ASD模型,中央服务器定期聚合这些本地模型。我们采用预先训练的模型作为ASD模型的骨干,以提高其鲁棒性,并开发专门的技术来在完全非iid和域偏移设置下稳定模型。与以前在集中式环境中训练的最先进的(SOTA)模型相比,CoopASD展示了具有竞争力的结果,其退化可以忽略不计,仅为0.08%。我们还进行了广泛的消融研究,以证明CoopASD的有效性。摘要:Machine anomalous sound detection (ASD) has emerged as one of the most promising applications in the Industrial Internet of Things (IIoT) due to its unprecedented efficacy in mitigating risks of malfunctions and promoting production efficiency. Previous works mainly investigated the machine ASD task under centralized settings. However, developing the ASD system under decentralized settings is crucial in practice, since the machine data are dispersed in various factories and the data should not be explicitly shared due to privacy concerns. To enable these factories to cooperatively develop a scalable ASD model while preserving their privacy, we propose a novel framework named CoopASD, where each factory trains an ASD model on its local dataset, and a central server aggregates these local models periodically. We employ a pre-trained model as the backbone of the ASD model to improve its robustness and develop specialized techniques to stabilize the model under a completely non-iid and domain shift setting. Compared with previous state-of-the-art (SOTA) models trained in centralized settings, CoopASD showcases competitive results with negligible degradation of 0.08%. We also conduct extensive ablation studies to demonstrate the effectiveness of CoopASD.

【10】 VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
标题: SecureTailor:用于基于扩散的个性化文本到语音的轻量级插件适配器
作者:Heeseung Kim,Sang-gil Lee,Jiheum Yeom,Che Hyun Lee,Sungwon Kim,Sungroh Yoon
链接:点击下载PDF文件
摘要:我们提出了VoiceTailor,一个参数高效的扬声器自适应文本到语音(TTS)系统,通过配备一个预先训练的基于扩散的TTS模型与个性化的适配器。VoiceTailor根据权重变化率分析确定从适配器中受益的关键模块。我们利用低秩自适应(LoRA)作为一种参数有效的自适应方法,并将适配器纳入预训练扩散解码器的关键模块。为了实现强大的自适应性能,很少的参数,我们探讨了各种指导技术的说话人自适应和调查的最佳策略,以加强说话人的信息。VoiceTailor通过仅微调总参数的0.25%来展示与现有自适应TTS模型相当的说话人自适应性能。VoiceTailor在适应各种真实世界的扬声器时表现出强大的鲁棒性,如演示中所示。摘要:We propose VoiceTailor, a parameter-efficient speaker-adaptive text-to-speech (TTS) system, by equipping a pre-trained diffusion-based TTS model with a personalized adapter. VoiceTailor identifies pivotal modules that benefit from the adapter based on a weight change ratio analysis. We utilize Low-Rank Adaptation (LoRA) as a parameter-efficient adaptation method and incorporate the adapter into pivotal modules of the pre-trained diffusion decoder. To achieve powerful adaptation performance with few parameters, we explore various guidance techniques for speaker adaptation and investigate the best strategies to strengthen speaker information. VoiceTailor demonstrates comparable speaker adaptation performance to existing adaptive TTS models by fine-tuning only 0.25 % of the total parameters. VoiceTailor shows strong robustness when adapting to a wide range of real-world speakers, as shown in the demo.

【11】 Physics-Informed Machine Learning For Sound Field Estimation
标题: 基于物理知识的机器学习用于声学场估计
作者:Shoichi Koyama,Juliano G. C. Ribeiro,Tomohiko Nakamura,Natsuki Ueno,Mirco Pezzoli
备注:Accepted to IEEE Signal Processing Magazine, Special Issue on Model-based and Data-Driven Audio Signal Processing
链接:点击下载PDF文件
摘要:关于空间声音估计的研究领域,即,诸如声压的声音的物理量的分布被称为声场估计,其是与空间音频处理相关的各种应用技术的基础。声场估计问题在简化的场景中被公式化为机器学习中的函数插值问题。然而,高估计性能不能通过简单地应用仅依赖于数据的一般插值技术来预期。声场的物理性质是有用的先验信息,并且将它们纳入估计中被认为是极其重要的。在这篇文章中,我们介绍了用于声场估计的物理信息机器学习(PIML)的基本原理,并概述了当前基于PIML的声场估计方法。摘要:The area of study concerning the estimation of spatial sound, i.e., the distribution of a physical quantity of sound such as acoustic pressure, is called sound field estimation, which is the basis for various applied technologies related to spatial audio processing. The sound field estimation problem is formulated as a function interpolation problem in machine learning in a simplified scenario. However, high estimation performance cannot be expected by simply applying general interpolation techniques that rely only on data. The physical properties of sound fields are useful a priori information, and it is considered extremely important to incorporate them into the estimation. In this article, we introduce the fundamentals of physics-informed machine learning (PIML) for sound field estimation and overview current PIML-based sound field estimation methods.

【12】 StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
标题: StyleSpeech:预训练的可控文本到语音的参数高效微调
作者:Haowei Lou,Helen Paik,Wen Hu,Lina Yao
链接:点击下载PDF文件
摘要:本文介绍了一种新型的文本到语音转换系统StyleSpeech,它提高了合成语音的自然度和准确度。基于现有的TTS技术,StyleSpeech采用了独特的Style Decorator结构,使深度学习模型能够同时学习风格和音素特征,通过Lower Rank Adaptation~(LoRA)原则提高适应性和效率。LoRA允许在预先训练的模型中有效地适应风格特征。此外,我们引入了一种新的自动评估指标,LLM-Guided Mean Opinion Score(LLM-MOS),它采用大型语言模型,为自动评估TTS系统性能提供了一个客观而强大的协议。对基准数据集的广泛测试表明,我们的方法在产生自然、准确和高质量的语音方面显着优于现有的最先进的基线方法。这些进步不仅推动了当前TTS系统功能的边界,而且还促进了TTS系统在更动态和专业化方面的应用,例如交互式虚拟助手,自适应有声读物和游戏定制语音。语音样本可以在https: style-speech.vercel.app上找到摘要:This paper introduces StyleSpeech, a novel Text-to-Speech~(TTS) system that enhances the naturalness and accuracy of synthesized speech. Building upon existing TTS technologies, StyleSpeech incorporates a unique Style Decorator structure that enables deep learning models to simultaneously learn style and phoneme features, improving adaptability and efficiency through the principles of Lower Rank Adaptation~(LoRA). LoRA allows efficient adaptation of style features in pre-trained models. Additionally, we introduce a novel automatic evaluation metric, the LLM-Guided Mean Opinion Score (LLM-MOS), which employs large language models to offer an objective and robust protocol for automatically assessing TTS system performance. Extensive testing on benchmark datasets shows that our approach markedly outperforms existing state-of-the-art baseline methods in producing natural, accurate, and high-quality speech. These advancements not only pushes the boundaries of current TTS system capabilities, but also facilitate the application of TTS system in more dynamic and specialized, such as interactive virtual assistants, adaptive audiobooks, and customized voice for gaming. Speech samples can be found in https: style-speech.vercel.app

【13】 Global-Local Distillation Network-Based Audio-Visual Speaker Tracking with Incomplete Modalities
标题: 基于全球-本地蒸馏网络的不完整模式视听说话人跟踪
作者:Yidi Li,Yihan Li,Yixin Guo,Bin Ren,Zhenhuan Xu,Hao Guo,Hong Liu,Nicu Sebe
备注:Audio-Visual Speaker Tracking with Incomplete Modalities
链接:点击下载PDF文件
摘要:在说话人跟踪研究中,多模态数据的融合和互补是提高跟踪系统准确性和鲁棒性的关键策略。然而,不完整的方式跟踪仍然是一个具有挑战性的问题,由于噪声观测造成的闭塞,噪声和传感器故障。特别是当多模态数据存在缺失时,现有的多模态融合方法的性能往往会下降。为此,我们提出了一个全球本地蒸馏为基础的...(GLDTracker)强大的视听扬声器跟踪。GLDTracker由教师-学生蒸馏模型驱动,能够灵活融合来自每种模态的不完整信息。教师网络处理摄像头和麦克风阵列捕获的全局信号,学生网络处理受视觉遮挡和丢失音频通道影响的本地信息。通过将知识从教师转移到学生,学生网络可以更好地适应观察不完整的复杂动态场景。在学生网络中,构建了一个基于生成对抗网络的全局特征重构模块,用于从缺失局部信息的特征嵌入中重构全局特征。在此基础上,利用视听特征和全局-局部特征的互补性和一致性,引入多模态多层次融合注意力,将不完整特征和重构特征进行融合。在AV16.3数据集上的实验结果表明,所提出的GLDTracker优于现有的最先进的视听...,并在标准和不完整的模态数据集上实现了领先的性能,突出了其在复杂条件下的优越性和鲁棒性。代码和模型将可用。摘要:In speaker tracking research, integrating and complementing multi-modal data is a crucial strategy for improving the accuracy and robustness of tracking systems. However, tracking with incomplete modalities remains a challenging issue due to noisy observations caused by occlusion, acoustic noise, and sensor failures. Especially when there is missing data in multiple modalities, the performance of existing multi-modal fusion methods tends to decrease. To this end, we propose a Global-Local Distillation-based Tracker (GLDTracker) for robust audio-visual speaker tracking. GLDTracker is driven by a teacher-student distillation model, enabling the flexible fusion of incomplete information from each modality. The teacher network processes global signals captured by camera and microphone arrays, and the student network handles local information subject to visual occlusion and missing audio channels. By transferring knowledge from teacher to student, the student network can better adapt to complex dynamic scenes with incomplete observations. In the student network, a global feature reconstruction module based on the generative adversarial network is constructed to reconstruct global features from feature embedding with missing local information. Furthermore, a multi-modal multi-level fusion attention is introduced to integrate the incomplete feature and the reconstructed feature, leveraging the complementarity and consistency of audio-visual and global-local features. Experimental results on the AV16.3 dataset demonstrate that the proposed GLDTracker outperforms existing state-of-the-art audio-visual trackers and achieves leading performance on both standard and incomplete modalities datasets, highlighting its superiority and robustness in complex conditions. The code and models will be available.

【14】 Infusing Acoustic Pause Context into Text-Based Dementia Assessment
标题: 将声学预设背景融入基于文本的痴呆症评估中
作者:Franziska Braun,Sebastian P. Bayerl,Florian Hönig,Hartmut Lehfeld,Thomas Hillemacher,Tobias Bocklet,Korbinian Riedhammer
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:语音停顿,以及内容和结构,为检测痴呆症提供了一个有价值的非侵入性生物标志物。这项工作研究了使用暂停丰富的成绩单在transformer为基础的语言模型,以区分没有认知障碍,轻度认知障碍,阿尔茨海默氏症的基础上,他们的语音从临床评估的认知状态的主题。我们解决三个二元分类任务:发病,监测和痴呆症排除。通过对德语口语流利性测试和图片描述测试的实验,比较该模型在不同的语音生成环境中的有效性。从一个文本的基线,我们调查的停顿信息和声学上下文的结合的效果。我们发现测试应该根据任务来选择,同样,词汇停顿信息和声学交叉注意的贡献不同。摘要:Speech pauses, alongside content and structure, offer a valuable and non-invasive biomarker for detecting dementia. This work investigates the use of pause-enriched transcripts in transformer-based language models to differentiate the cognitive states of subjects with no cognitive impairment, mild cognitive impairment, and Alzheimer's dementia based on their speech from a clinical assessment. We address three binary classification tasks: Onset, monitoring, and dementia exclusion. The performance is evaluated through experiments on a German Verbal Fluency Test and a Picture Description Test, comparing the model's effectiveness across different speech production contexts. Starting from a textual baseline, we investigate the effect of incorporation of pause information and acoustic context. We show the test should be chosen depending on the task, and similarly, lexical pause information and acoustic cross-attention contribute differently.

【15】 Development of Large Annotated Music Datasets using HMM-based Forced Viterbi Alignment
标题: 使用基于HM的强制维特比对齐开发大型注释音乐数据集
作者:S. Johanan Joysingh,P. Vijayalakshmi,T. Nagarajan
Journal-ref:S. J. Joysingh, P. Vijayalakshmi and T. Nagarajan, "Development of Large Annotated Music Datasets using HMM based Forced Viterbi Alignment," TENCON 2019 - 2019 IEEE Region 10 Conference (TENCON), Kochi, India, 2019, pp. 1298-1302
链接:点击下载PDF文件
摘要:数据集对于任何机器学习任务都至关重要。自动音乐转录(AMT)就是这样一项任务,根据解决方案的实现方式,需要大量的数据。考虑到一个音乐数据集,包括音频及其时间对齐的transmittance,需要有音乐经验的人的努力,可以说这项任务变得更具挑战性。在演奏乐器,注释和验证音准时需要音乐经验。我们提出了一种方法,这将有助于简化这一过程,使从特定的仪器获得数据集的任务变得简单和高效。我们使用预定义的吉他练习和基于隐马尔可夫模型(HMM)的强制维特比对齐来实现这一点。吉他练习被设计成简单的。由于音符序列已经被定义,基于HMM的强制维特比对齐提供了这些音频文件的时间对齐的传输。手动验证瞬变的起始时间,标签准确度高达10 ms,平均值为5 ms。所提出的工作的贡献是两方面的,i)一个很好的流线型和高效的方法,用于生成任何乐器的数据集,特别是单声道的,ii)一个包含波文件和以标签文件形式的音准的声学拨弦吉他数据集。这种方法将有助于为不同的仪器建立具体的数据集,以建立AMT系统的初步步骤。摘要:Datasets are essential for any machine learning task. Automatic Music Transcription (AMT) is one such task, where considerable amount of data is required depending on the way the solution is achieved. Considering the fact that a music dataset, complete with audio and its time-aligned transcriptions would require the effort of people with musical experience, it could be stated that the task becomes even more challenging. Musical experience is required in playing the musical instrument(s), and in annotating and verifying the transcriptions. We propose a method that would help in streamlining this process, making the task of obtaining a dataset from a particular instrument easy and efficient. We use predefined guitar exercises and hidden Markov model(HMM) based forced viterbi alignment to accomplish this. The guitar exercises are designed to be simple. Since the note sequence are already defined, HMM based forced viterbi alignment provides time-aligned transcriptions of these audio files. The onsets of the transcriptions are manually verified and the labels are accurate up to 10ms, averaging at 5ms. The contributions of the proposed work is two fold, i) a well streamlined and efficient method for generating datasets for any instrument, especially monophonic and, ii) an acoustic plectrum guitar dataset containing wave files and transcriptions in the form of label files. This method will aid as a preliminary step towards building concrete datasets for building AMT systems for different instruments.

【16】 Comparative Analysis Of Discriminative Deep Learning-Based Noise Reduction Methods In Low SNR Scenarios
标题: 低SNR场景下基于区分性深度学习的降噪方法的比较分析
作者:Shrishti Saha Shetu,Emanuël A. P. Habets,Andreas Brendel
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:在这项研究中,我们对低信噪比(SNR)场景下基于深度学习的降噪方法进行了比较分析。我们的调查主要集中在五个关键方面:训练数据的影响,各种损失函数的影响,直接和间接语音估计技术的有效性,掩蔽,映射和深度滤波方法的有效性,以及不同模型容量对降噪性能和语音质量的探索。通过全面的实验,我们提供了这些方法在低信噪比环境中的优点,缺点和适用性的见解。从我们的分析得出的结果是为了帮助研究人员和从业人员在选择更好的技术,适合他们的特定应用领域内的低信噪比降噪。摘要:In this study, we conduct a comparative analysis of deep learning-based noise reduction methods in low signal-to-noise ratio (SNR) scenarios. Our investigation primarily focuses on five key aspects: The impact of training data, the influence of various loss functions, the effectiveness of direct and indirect speech estimation techniques, the efficacy of masking, mapping, and deep filtering methodologies, and the exploration of different model capacities on noise reduction performance and speech quality. Through comprehensive experimentation, we provide insights into the strengths, weaknesses, and applicability of these methods in low SNR environments. The findings derived from our analysis are intended to assist both researchers and practitioners in selecting better techniques tailored to their specific applications within the domain of low SNR noise reduction.


eess.AS音频处理
【1】 Infusing Acoustic Pause Context into Text-Based Dementia Assessment
标题: 将声学预设背景融入基于文本的痴呆症评估中
作者:Franziska Braun,Sebastian P. Bayerl,Florian Hönig,Hartmut Lehfeld,Thomas Hillemacher,Tobias Bocklet,Korbinian Riedhammer
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:语音停顿,以及内容和结构,为检测痴呆症提供了一个有价值的非侵入性生物标志物。这项工作研究了使用暂停丰富的成绩单在transformer为基础的语言模型,以区分没有认知障碍,轻度认知障碍,阿尔茨海默氏症的基础上,他们的语音从临床评估的认知状态的主题。我们解决三个二元分类任务:发病,监测和痴呆症排除。通过对德语口语流利性测试和图片描述测试的实验,比较该模型在不同的语音生成环境中的有效性。从一个文本的基线,我们调查的停顿信息和声学上下文的结合的效果。我们发现测试应该根据任务来选择,同样,词汇停顿信息和声学交叉注意的贡献不同。摘要:Speech pauses, alongside content and structure, offer a valuable and non-invasive biomarker for detecting dementia. This work investigates the use of pause-enriched transcripts in transformer-based language models to differentiate the cognitive states of subjects with no cognitive impairment, mild cognitive impairment, and Alzheimer's dementia based on their speech from a clinical assessment. We address three binary classification tasks: Onset, monitoring, and dementia exclusion. The performance is evaluated through experiments on a German Verbal Fluency Test and a Picture Description Test, comparing the model's effectiveness across different speech production contexts. Starting from a textual baseline, we investigate the effect of incorporation of pause information and acoustic context. We show the test should be chosen depending on the task, and similarly, lexical pause information and acoustic cross-attention contribute differently.

【2】 Integrating Continuous and Binary Relevances in Audio-Text Relevance Learning
标题: 音频文本相关性学习中的连续相关性和二元相关性的整合
作者:Huang Xie,Khazar Khorrami,Okko Räsänen,Tuomas Virtanen
备注:Accepted at DCASE 2024 Workshop
链接:点击下载PDF文件
摘要:音频文本相关性学习是指学习音频样本和文本描述的共享语义属性。标准方法使用从成对的音频样本及其人类提供的字幕中得出的二进制相关性,将每对分类为正面或负面。由于音频样本和字幕之间的相关性水平不同,这可能导致次优系统。相比之下,最近的一项研究使用了人类分配的相关性评级,即,连续的相关性,这些对,但没有获得性能增益音频文本相关性学习。这项工作介绍了一种相关性学习方法,利用人类分配的连续相关性评级和二进制相关性使用一个列表排序目标和对比学习目标的组合。实验结果表明,该方法的有效性,显示在基于语言的音频检索,在音频文本相关性学习的下游任务的改进。此外,我们分析了字幕或音频片段的属性如何有助于由人类提供或由机器学习的连续音频-文本相关性。摘要:Audio-text relevance learning refers to learning the shared semantic properties of audio samples and textual descriptions. The standard approach uses binary relevances derived from pairs of audio samples and their human-provided captions, categorizing each pair as either positive or negative. This may result in suboptimal systems due to varying levels of relevance between audio samples and captions. In contrast, a recent study used human-assigned relevance ratings, i.e., continuous relevances, for these pairs but did not obtain performance gains in audio-text relevance learning. This work introduces a relevance learning method that utilizes both human-assigned continuous relevance ratings and binary relevances using a combination of a listwise ranking objective and a contrastive learning objective. Experimental results demonstrate the effectiveness of the proposed method, showing improvements in language-based audio retrieval, a downstream task in audio-text relevance learning. In addition, we analyze how properties of the captions or audio clips contribute to the continuous audio-text relevances provided by humans or learned by the machine.

【3】 Development of Large Annotated Music Datasets using HMM-based Forced Viterbi Alignment
标题: 使用基于HM的强制维特比对齐开发大型注释音乐数据集
作者:S. Johanan Joysingh,P. Vijayalakshmi,T. Nagarajan
Journal-ref:S. J. Joysingh, P. Vijayalakshmi and T. Nagarajan, "Development of Large Annotated Music Datasets using HMM based Forced Viterbi Alignment," TENCON 2019 - 2019 IEEE Region 10 Conference (TENCON), Kochi, India, 2019, pp. 1298-1302
链接:点击下载PDF文件
摘要:数据集对于任何机器学习任务都至关重要。自动音乐转录(AMT)就是这样一项任务,根据解决方案的实现方式,需要大量的数据。考虑到一个音乐数据集,包括音频及其时间对齐的transmittance,需要有音乐经验的人的努力,可以说这项任务变得更具挑战性。在演奏乐器,注释和验证音准时需要音乐经验。我们提出了一种方法,这将有助于简化这一过程,使从特定的仪器获得数据集的任务变得简单和高效。我们使用预定义的吉他练习和基于隐马尔可夫模型(HMM)的强制维特比对齐来实现这一点。吉他练习被设计成简单的。由于音符序列已经被定义,基于HMM的强制维特比对齐提供了这些音频文件的时间对齐的传输。手动验证瞬变的起始时间,标签准确度高达10 ms,平均值为5 ms。所提出的工作的贡献是两方面的,i)一个很好的流线型和高效的方法,用于生成任何乐器的数据集,特别是单声道的,ii)一个包含波文件和以标签文件形式的音准的声学拨弦吉他数据集。这种方法将有助于为不同的仪器建立具体的数据集,以建立AMT系统的初步步骤。摘要:Datasets are essential for any machine learning task. Automatic Music Transcription (AMT) is one such task, where considerable amount of data is required depending on the way the solution is achieved. Considering the fact that a music dataset, complete with audio and its time-aligned transcriptions would require the effort of people with musical experience, it could be stated that the task becomes even more challenging. Musical experience is required in playing the musical instrument(s), and in annotating and verifying the transcriptions. We propose a method that would help in streamlining this process, making the task of obtaining a dataset from a particular instrument easy and efficient. We use predefined guitar exercises and hidden Markov model(HMM) based forced viterbi alignment to accomplish this. The guitar exercises are designed to be simple. Since the note sequence are already defined, HMM based forced viterbi alignment provides time-aligned transcriptions of these audio files. The onsets of the transcriptions are manually verified and the labels are accurate up to 10ms, averaging at 5ms. The contributions of the proposed work is two fold, i) a well streamlined and efficient method for generating datasets for any instrument, especially monophonic and, ii) an acoustic plectrum guitar dataset containing wave files and transcriptions in the form of label files. This method will aid as a preliminary step towards building concrete datasets for building AMT systems for different instruments.

【4】 Literary and Colloquial Dialect Identification for Tamil using Acoustic Features
标题: 利用声学特征识别泰米尔语的文学方言和口语方言
作者:M. Nanmalar,P. Vijayalakshmi,T. Nagarajan
Journal-ref:TENCON 2019 - 2019 IEEE Region 10 Conference (TENCON), Kochi, India, 2019, pp. 1303-1306
链接:点击下载PDF文件
摘要:一种语言的演变和多样性从它的各种方言中显而易见。如果各种方言在自动语音识别和语音合成等技术进步中得不到解决,这些方言就有可能消失。语音技术在保护一种语言的各种方言免于灭绝方面发挥着作用。为了建立一个完整的自动语音识别系统,解决各种方言,自动方言识别(ADI)系统作为前端是必需的。这类似于语言识别系统如何作为处理多种语言的自动语音识别系统的前端。目前的工作提出了一种方法来确定两种流行的和广泛分类的泰米尔方言,即文学和口语泰米尔语。使用声学特征而不是语音学和音位学,减轻了对依赖语言的语言工具的需求。因此,所提出的方法的一个主要优点是,它不需要注释的语料库,因此它可以很容易地适应其他语言。使用梅尔频率倒谱系数(MFCC)特征的高斯混合模型(GMM)用于执行分类任务。实验得出的错误率为12%。元音鼻化,作为这种良好的性能的原因,进行了讨论。GMM的混合模型的数量是不同的,并分析了性能。摘要:The evolution and diversity of a language is evident from it's various dialects. If the various dialects are not addressed in technological advancements like automatic speech recognition and speech synthesis, there is a chance that these dialects may disappear. Speech technology plays a role in preserving various dialects of a language from going extinct. In order to build a full fledged automatic speech recognition system that addresses various dialects, an Automatic Dialect Identification (ADI) system acting as the front end is required. This is similar to how language identification systems act as front ends to automatic speech recognition systems that handle multiple languages. The current work proposes a way to identify two popular and broadly classified Tamil dialects, namely literary and colloquial Tamil. Acoustical characteristics rather than phonetics and phonotactics are used, alleviating the requirement of language-dependant linguistic tools. Hence one major advantage of the proposed method is that it does not require an annotated corpus, hence it can be easily adapted to other languages. Gaussian Mixture Models (GMM) using Mel Frequency Cepstral Coefficient (MFCC) features are used to perform the classification task. The experiments yielded an error rate of 12%. Vowel nasalization, as being the reason for this good performance, is discussed. The number of mixture models for the GMM is varied and the performance is analysed.

【5】 Similarity Metrics For Late Reverberation
标题: 后期回响的相似性检查
作者:Gloria Dal Santo,Karolina Prawda,Sebastian J. Schlecht,Vesa Välimäki
链接:点击下载PDF文件
摘要:混响算法的自动调谐依赖于成本函数的优化。虽然一般的音频相似性度量是有用的,但是它们没有针对房间中混响的特定统计特性进行优化。本文提出了两种新的度量标准,用于评估后期混响的房间脉冲响应的相似性。这些指标是可区分的,可以在机器学习框架中使用。我们使用包含各种房间配置和麦克风位置的大型房间脉冲响应数据集将这些指标的性能与两个流行的音频指标进行比较。结果表明,基于平均功率和频带能量衰减的建议函数优于基线,前者表现出最合适的轮廓接近最小值。所提出的工作有希望作为混响相似性度量的设计和评估的改进。摘要:Automatic tuning of reverberation algorithms relies on the optimization of a cost function. While general audio similarity metrics are useful, they are not optimized for the specific statistical properties of reverberation in rooms. This paper presents two novel metrics for assessing the similarity of late reverberation in room impulse responses. These metrics are differentiable and can be utilized within a machine-learning framework. We compare the performance of these metrics to two popular audio metrics using a large dataset of room impulse responses encompassing various room configurations and microphone positions. The results indicate that the proposed functions based on averaged power and frequency-band energy decay outperform the baselines with the former exhibiting the most suitable profile towards the minimum. The proposed work holds promise as an improvement to the design and evaluation of reverberation similarity metrics.

【6】 MaskCycleGAN-based Whisper to Normal Speech Conversion
标题: 基于MaskCycleGAN的Whisper到正常语音转换
作者:K. Rohith Gupta,K. Ramnath,S. Johanan Joysingh,P. Vijayalakshmi,T. Nagarajan
备注:submitted to TENCON 2024
链接:点击下载PDF文件
摘要:耳语到正常语音的转换是一个活跃的研究领域。最近,已经提出了基于生成对抗网络的各种架构。特别是,最近的研究表明,MaskCycleGAN是一种掩码引导的、保持循环一致性的生成对抗网络,在从声谱图表示进行语音转换方面表现非常好。在目前的工作中,我们提出了一个MaskCycleGAN的方法转换的耳语语音到正常的语音。我们发现,调整掩模参数,并预处理的信号与语音活动检测器相比,现有的方法提供了优越的性能。wTIMIT数据集用于评估。PESQ和G-Loss等客观指标用于评估转换后的语音,同时使用平均意见得分进行主观评估。结果表明,所提出的方法提供了相当大的好处。摘要:Whisper to normal speech conversion is an active area of research. Various architectures based on generative adversarial networks have been proposed in the recent past. Especially, recent study shows that MaskCycleGAN, which is a mask guided, and cyclic consistency keeping, generative adversarial network, performs really well for voice conversion from spectrogram representations. In the current work we present a MaskCycleGAN approach for the conversion of whispered speech to normal speech. We find that tuning the mask parameters, and pre-processing the signal with a voice activity detector provides superior performance when compared to the existing approach. The wTIMIT dataset is used for evaluation. Objective metrics such as PESQ and G-Loss are used to evaluate the converted speech, along with subjective evaluation using mean opinion score. The results show that the proposed approach offers considerable benefits.

【7】 Quartered Chirp Spectral Envelope for Whispered vs Normal Speech Classification
标题: 用于耳语与正常语音分类的四分之一Chirp频谱信封
作者:S. Johanan Joysingh,P. Vijayalakshmi,T. Nagarajan
备注:submitted to TENCON 2024
链接:点击下载PDF文件
摘要:耳语作为一种可接受的人机交互形式正在获得牵引力。处理多种语音模式的系统需要一个鲁棒的前端语音分类器。在存在加性高斯白噪声的情况下,耳语音与正常语音的分类性能下降,因为正常语音具有耳语音的一些特征。在这项工作中,我们提出了一个新的功能命名为四分之一的啁啾频谱包络,啁啾频谱和四分之一的频谱包络的组合,分类耳和正常的语音。啁啾频谱可以微调,以获得定制的功能,为给定的任务,和四分之一的频谱包络已被证明是工作特别好,为当前的任务。该特征在一维卷积神经网络上训练,该网络捕获频谱包络中的趋势。所提出的系统在白噪声的存在下比现有技术的性能更好。摘要:Whispered speech as an acceptable form of human-computer interaction is gaining traction. Systems that address multiple modes of speech require a robust front-end speech classifier. Performance of whispered vs normal speech classification drops in the presence of additive white Gaussian noise, since normal speech takes on some of the characteristics of whispered speech. In this work, we propose a new feature named the quartered chirp spectral envelope, a combination of the chirp spectrum and the quartered spectral envelope, to classify whispered and normal speech. The chirp spectrum can be fine-tuned to obtain customized features for a given task, and the quartered spectral envelope has been proven to work especially well for the current task. The feature is trained on a one dimensional convolutional neural network, that captures the trends in the spectral envelope. The proposed system performs better than the state of the art, in the presence of white noise.

【8】 Impact of Noisy Labels on Sound Event Detection: Deletion Errors Are More Detrimental Than Insertion Errors
标题: 噪音标签对声音事件检测的影响:删除错误比插入错误更有害
作者:Yuliang Zhang,Defeng,Huang,Roberto Togneri
链接:点击下载PDF文件
摘要:本研究探讨了关键的,但未充分检查的标签噪声对声音事件检测(SED),这需要声音识别和精确的时间定位的影响。我们将标签噪声分为删除,插入,替换和主观类型,并使用合成和现实生活中的数据集系统地评估其对SED的影响。我们的分析表明,删除噪声显着降低性能,而插入噪声是相对良性的。此外,损失函数有效的分类噪声不执行以及由于类内的不平衡之间的前景声音事件和背景声音的SED。我们证明,旨在解决SED中的数据不平衡的损失函数可以有效地减少噪声标签对系统性能的影响。例如,将合成数据集中的背景声音的权重减半,将宏F1和微F1分数提高了约9%,错误率增加最小,在现实生活中的数据集中结果一致。这项研究强调了噪声标签对SED系统的细微影响,并提供了增强模型鲁棒性的实用策略,这对于构建新的SED数据集和提高模型性能都至关重要,包括有效利用软标签和众包标签。摘要:This study explores the critical but underexamined impact of label noise on Sound Event Detection (SED), which requires both sound identification and precise temporal localization. We categorize label noise into deletion, insertion, substitution, and subjective types and systematically evaluate their effects on SED using synthetic and real-life datasets. Our analysis shows that deletion noise significantly degrades performance, while insertion noise is relatively benign. Moreover, loss functions effective against classification noise do not perform well for SED due to intra-class imbalance between foreground sound events and background sounds. We demonstrate that loss functions designed to address data imbalance in SED can effectively reduce the impact of noisy labels on system performance. For instance, halving the weight of background sounds in a synthetic dataset improved macro-F1 and micro-F1 scores by approximately $9 %$ with minimal Error Rate increase, with consistent results in real-life datasets. This research highlights the nuanced effects of noisy labels on SED systems and provides practical strategies to enhance model robustness, which are pivotal for both constructing new SED datasets and improving model performance, including efficient utilization of soft and crowdsourced labels.

【9】 Is Audio Spoof Detection Robust to Laundering Attacks?
标题: 音频欺骗检测对洗钱攻击是否稳健?
作者:Hashim Ali,Surya Subramani,Shefali Sudhir,Raksha Varahamurthy,Hafiz Malik
备注:Conference Paper
链接:点击下载PDF文件
摘要:近年来,语音克隆(VC)系统在合成语音的真实性方面有了显著的提高。合成语音的高质量和低成本VC服务的可用性引起了该技术的许多潜在滥用。多年来,已经提出了几种检测方法,可以以相当好的准确度检测语音欺骗。然而,这些方法大多在干净的音频数据库上进行评估,例如ASVSpoof 2019。本文评估SOTA音频欺骗检测方法在洗钱攻击的存在。在这方面,创建了一个新的洗钱攻击数据库,称为ASVSpoof洗钱数据库。该数据库基于ASVSpoof 2019(LA)评估数据库,包括总计1388.22小时的音频记录。七SOTA音频欺骗检测方法进行评估,这个清洗数据库。结果表明,SOTA系统在存在攻击性洗钱攻击(尤其是混响和加性噪声攻击)的情况下表现不佳。这表明需要鲁棒的音频欺骗检测。摘要:Voice-cloning (VC) systems have seen an exceptional increase in the realism of synthesized speech in recent years. The high quality of synthesized speech and the availability of low-cost VC services have given rise to many potential abuses of this technology. Several detection methodologies have been proposed over the years that can detect voice spoofs with reasonably good accuracy. However, these methodologies are mostly evaluated on clean audio databases, such as ASVSpoof 2019. This paper evaluates SOTA Audio Spoof Detection approaches in the presence of laundering attacks. In that regard, a new laundering attack database, called the ASVSpoof Laundering Database, is created. This database is based on the ASVSpoof 2019 (LA) eval database comprising a total of 1388.22 hours of audio recordings. Seven SOTA audio spoof detection approaches are evaluated on this laundered database. The results indicate that SOTA systems perform poorly in the presence of aggressive laundering attacks, especially reverberation and additive noise attacks. This suggests the need for robust audio spoof detection.

【10】 Comparative Analysis Of Discriminative Deep Learning-Based Noise Reduction Methods In Low SNR Scenarios
标题: 低SNR场景下基于区分性深度学习的降噪方法的比较分析
作者:Shrishti Saha Shetu,Emanuël A. P. Habets,Andreas Brendel
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:在这项研究中,我们对低信噪比(SNR)场景下基于深度学习的降噪方法进行了比较分析。我们的调查主要集中在五个关键方面:训练数据的影响,各种损失函数的影响,直接和间接语音估计技术的有效性,掩蔽,映射和深度滤波方法的有效性,以及不同模型容量对降噪性能和语音质量的探索。通过全面的实验,我们提供了这些方法在低信噪比环境中的优点,缺点和适用性的见解。从我们的分析得出的结果是为了帮助研究人员和从业人员在选择更好的技术,适合他们的特定应用领域内的低信噪比降噪。摘要:In this study, we conduct a comparative analysis of deep learning-based noise reduction methods in low signal-to-noise ratio (SNR) scenarios. Our investigation primarily focuses on five key aspects: The impact of training data, the influence of various loss functions, the effectiveness of direct and indirect speech estimation techniques, the efficacy of masking, mapping, and deep filtering methodologies, and the exploration of different model capacities on noise reduction performance and speech quality. Through comprehensive experimentation, we provide insights into the strengths, weaknesses, and applicability of these methods in low SNR environments. The findings derived from our analysis are intended to assist both researchers and practitioners in selecting better techniques tailored to their specific applications within the domain of low SNR noise reduction.

【11】 Unlocking Potential in Pre-Trained Music Language Models for Versatile Multi-Track Music Arrangement
标题: 释放预训练音乐语言模型的潜力,用于多功能多轨音乐编曲
作者:Longshen Ou,Jingwei Zhao,Ziyu Wang,Gus Xia,Ye Wang
备注:Submitted to AAAI 2025
链接:点击下载PDF文件
摘要:大型语言模型在各个领域都表现出了显著的能力,包括符号音乐生成。然而,利用这些预先训练的模型进行可控的音乐编排任务,每个任务都需要不同形式的音乐信息作为控制,仍然是一个新的挑战。在本文中,我们提出了一个统一的序列到序列的框架,使微调的符号音乐语言模型的多个多轨道的安排任务,包括乐队安排,钢琴减少,鼓安排,和语音分离。我们的实验表明,所提出的方法始终实现更高的音乐质量相比,在所有四个任务的任务特定的基线。此外,通过对探测分析的额外实验,我们表明预训练阶段为模型提供了理解音乐条件的基本知识,这是很难仅仅通过特定任务的微调获得的。摘要:Large language models have shown significant capabilities across various domains, including symbolic music generation. However, leveraging these pre-trained models for controllable music arrangement tasks, each requiring different forms of musical information as control, remains a novel challenge. In this paper, we propose a unified sequence-to-sequence framework that enables the fine-tuning of a symbolic music language model for multiple multi-track arrangement tasks, including band arrangement, piano reduction, drum arrangement, and voice separation. Our experiments demonstrate that the proposed approach consistently achieves higher musical quality compared to task-specific baselines across all four tasks. Furthermore, through additional experiments on probing analysis, we show the pre-training phase equips the model with essential knowledge to understand musical conditions, which is hard to acquired solely through task-specific fine-tuning.

【12】 Speech Recognition Transformers: Topological-lingualism Perspective
标题: 语音识别变形者:话题语言主义视角
作者:Shruti Singh,Muskaan Singh,Virender Kadyan
链接:点击下载PDF文件
摘要:Transformers在各种人工智能任务中取得了巨大的成功。由于我们最近流行的自我注意机制,它捕获了长期的依赖性,在语音处理和识别任务中产生了惊人的结果。本文对面向语音模态的Transformer技术进行了综述。本次调查的主要内容包括:(1)传统ASR的背景,端到端的Transformer生态系统,以及语音Transformers(2)通过语言学范式的语音基础模型,即,单语、双语、多语言和跨语言(3)数据集和语言、声学特征、架构、解码和特定拓扑语言学视角的评估指标(4)用于构建端到端ASR系统的流行语音Transformer工具包。最后,重点讨论了开放的挑战和潜在的研究方向,为社区在这一领域进行进一步的研究。摘要:Transformers have evolved with great success in various artificial intelligence tasks. Thanks to our recent prevalence of self-attention mechanisms, which capture long-term dependency, phenomenal outcomes in speech processing and recognition tasks have been produced. The paper presents a comprehensive survey of transformer techniques oriented in speech modality. The main contents of this survey include (1) background of traditional ASR, end-to-end transformer ecosystem, and speech transformers (2) foundational models in a speech via lingualism paradigm, i.e., monolingual, bilingual, multilingual, and cross-lingual (3) dataset and languages, acoustic features, architecture, decoding, and evaluation metric from a specific topological lingualism perspective (4) popular speech transformer toolkit for building end-to-end ASR systems. Finally, highlight the discussion of open challenges and potential research directions for the community to conduct further research in this domain.

【13】 Morphogenesis of sound creates acoustic rainbows
标题: 声音的形态发生创造了声学彩虹
作者:Rasmus E. Christiansen,Ole Sigmund,Efren Fernandez-Grande
备注:8 Pages, 4 Figures, Supplementary information including text and four movies
链接:点击下载PDF文件
摘要:声音是自然界中许多生物体的基本传感元件,并且多个物种已经进化出产生复杂的声散射和分散现象的有机结构,以明确地发射和感知声音。迄今为止,还没有证明有可能设计出与有机结构中发现的那些结构的性能相媲美的人工散射结构。首先,大多数声音操纵依赖于流体介质中的主动转导,而不是依赖于自然界中经常发现的被动散射原理。在这项工作中,我们利用计算形态合成复杂的节能波长大小的单材料散射结构,被动地将辐射声分解成其空间频谱分量。具体来说,我们量身定制的声学彩虹结构与“以上统一”的效率和声学波长分离器。我们的工作为声场工程的新前沿铺平了道路,在转导,仿生学,能量收集,通信和传感方面具有潜在的应用。摘要:Sound is an essential sensing element for many organisms in nature, and multiple species have evolved organic structures that create complex acoustic scattering and dispersion phenomena to emit and perceive sound unambiguously. To date, it has not proven possible to design artificial scattering structures that rival the performance of those found in organic structures. Contrarily, most sound manipulation relies on active transduction in fluid media rather than relying on passive scattering principles, as are often found in nature. In this work, we utilize computational morphogenesis to synthesize complex energy-efficient wavelength-sized single-material scattering structures that passively decompose radiated sound into its spatio-spectral components. Specifically, we tailor an acoustic rainbow structure with "above unity" efficiency and an acoustic wavelength-splitter. Our work paves the way for a new frontier in sound-field engineering, with potential applications in transduction, bionics, energy harvesting, communications and sensing.

【14】 Deep learning classification system for coconut maturity levels based on acoustic signals
标题: 基于声信号的椰子成熟度深度学习分类系统
作者:June Anne Caladcad,Eduardo Jr Piedad
备注:7 pages, 4 figures, accepted in the 2024 IEEE 12th Region 10 Humanitarian Technology Conference (R10-HTC)
链接:点击下载PDF文件
摘要:随着计算机图像处理、模式识别、信号处理等技术的进步,水果分类的人工方法逐渐被计算机和机械方法所取代。在农业领域,收获后水果的智能分类使智能设备的使用成为可能,这对农民,特别是出口产品产生了直接影响。对于椰子分类,它仍然是传统的过程。本研究提出了一种基于声信号的椰子数据集的分类。为了解决不平衡的数据集,通过听觉和程序音频生成方法进行数据增强技术。未成熟、成熟和过成熟下的音频信号现在分别具有4,050、4,050和5,850个音频信号。为了解决分类系统的更新和分类准确性性能,利用深度学习模型对来自数据生成的所生成的音频信号进行分类。具体来说,RNN和LSTM模型进行了训练和测试,并将它们的性能与Caladcad等人(2020)使用的机器学习方法进行了比较。这两个DL模型表现出令人印象深刻的性能,两者的准确率均为97.42%,并且由于它们的分类性能没有显着差异,因此它们都没有优于另一个。摘要:The advancement of computer image processing, pattern recognition, signal processing, and other technologies has gradually replaced the manual methods of classifying fruit with computer and mechanical methods. In the field of agriculture, the intelligent classification of post-harvested fruit has enabled the use of smart devices that creates a direct impact on farmers, especially on export products. For coconut classification, it remains to be traditional in process. This study presents a classification of the coconut dataset based on acoustic signals. To address the imbalanced dataset, a data augmentation technique was conducted through audiomentation and procedural audio generation methods. Audio signals under premature, mature, and overmature now have 4,050, 4,050, and 5,850 audio signals, respectively. To address the updation of the classification system and the classification accuracy performance, deep learning models were utilized for classifying the generated audio signals from data generation. Specifically, RNN and LSTM models were trained and tested, and their performances were compared with each other and the machine learning methods used by Caladcad et al. (2020). The two DL models showed impressive performance with both having an accuracy of 97.42% and neither of them outperformed the other since there are no significant differences in their classification performance.

【15】 A Functional Trade-off between Prosodic and Semantic Cues in Conveying Sarcasm
标题: 讽刺传达中韵律线索与语义线索的功能权衡
作者:Zhu Li,Xiyuan Gao,Yuqing Zhang,Shekhar Nayak,Matt Coler
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:本研究探讨了讽刺的声学特征,并解开了话语被讽刺使用的倾向和韵律线索之间的相互作用。使用从电视节目中编译的讽刺话语的数据集,我们分析的韵律特征内的话语和关键短语属于三个不同的讽刺类别(嵌入式,命题,和言外之语),不同程度的语义线索存在,并比较它们的中性表达。结果表明,在短语的讽刺意义是突出的语义,韵律线索的相关性较低时,讽刺的意义是不明显的语义,这表明韵律和语义线索之间的权衡讽刺在短语水平。这些研究结果强调了减少依赖韵律调制在语义密集的讽刺表达和微妙的互动,塑造讽刺意图的沟通。摘要:This study investigates the acoustic features of sarcasm and disentangles the interplay between the propensity of an utterance being used sarcastically and the presence of prosodic cues signaling sarcasm. Using a dataset of sarcastic utterances compiled from television shows, we analyze the prosodic features within utterances and key phrases belonging to three distinct sarcasm categories (embedded, propositional, and illocutionary), which vary in the degree of semantic cues present, and compare them to neutral expressions. Results show that in phrases where the sarcastic meaning is salient from the semantics, the prosodic cues are less relevant than when the sarcastic meaning is not evident from the semantics, suggesting a trade-off between prosodic and semantic cues of sarcasm at the phrase level. These findings highlight a lessened reliance on prosodic modulation in semantically dense sarcastic expressions and a nuanced interaction that shapes the communication of sarcastic intent.

【16】 The VoxCeleb Speaker Recognition Challenge: A Retrospective
标题: VoxCeleb演讲者识别挑战:回顾
作者:Jaesung Huh,Joon Son Chung,Arsha Nagrani,Andrew Brown,Jee-weon Jung,Daniel Garcia-Romero,Andrew Zisserman
备注:TASLP 2024
链接:点击下载PDF文件
摘要:VoxCeleb Speaker Recognition Challenges(VoxSRC)是一系列挑战和研讨会,从2019年到2023年每年举办一次。这些挑战主要评估了各种设置下的说话人识别和日记的任务,包括:封闭和开放训练数据;以及用于域适应的监督,自监督和半监督训练。挑战赛还为每项任务和设置提供了公开的培训和评估数据集,每年发布新的测试集。在本文中,我们提供了一个审查这些挑战,包括:他们探讨了什么;挑战参与者开发的方法以及这些方法是如何演变的;以及扬声器验证和diarisation领域的现状。我们在一个共同的评估数据集上绘制了五个阶段的表现进展,并详细分析了每年的特别关注如何影响参与者的表现。本文的目的是研究人员谁想要一个概述的说话人识别和diarisation领域,也在挑战组织者谁想要受益于成功,避免错误的VoxSRC挑战。最后,我们讨论了该领域目前的优势和面临的挑战。项目页面:https: mm.kaist.ac.kr datasets voxceleb voxsrc workshop.html摘要:The VoxCeleb Speaker Recognition Challenges (VoxSRC) were a series of challenges and workshops that ran annually from 2019 to 2023. The challenges primarily evaluated the tasks of speaker recognition and diarisation under various settings including: closed and open training data; as well as supervised, self-supervised, and semi-supervised training for domain adaptation. The challenges also provided publicly available training and evaluation datasets for each task and setting, with new test sets released each year. In this paper, we provide a review of these challenges that covers: what they explored; the methods developed by the challenge participants and how these evolved; and also the current state of the field for speaker verification and diarisation. We chart the progress in performance over the five installments of the challenge on a common evaluation dataset and provide a detailed analysis of how each year's special focus affected participants' performance. This paper is aimed both at researchers who want an overview of the speaker recognition and diarisation field, and also at challenge organisers who want to benefit from the successes and avoid the mistakes of the VoxSRC challenges. We end with a discussion of the current strengths of the field and open challenges. Project page : https: mm.kaist.ac.kr datasets voxceleb voxsrc workshop.html

【17】 Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
标题: 利用自我监督的音频表示进行数据高效的声学场景分类
作者:Yiqiang Cai,Shengchen Li,Xi Shao
备注:Accepted by DCASE Workshop 2024
链接:点击下载PDF文件
摘要:声场景分类(ASC)主要依赖于监督的方法。然而,获取用于训练ASC模型的标记数据通常是昂贵且耗时的。最近,自监督学习(SSL)已经成为一种从未标记音频数据中提取特征的强大方法,使许多下游音频任务受益。本文提出了一个数据高效和低复杂度的ASC系统,利用自监督音频表示提取通用音频数据集。我们引入BEAT,一个音频SSL预训练模型,从AudioSet中提取一般表示。通过大量的实验,它已被证明,自监督音频表示可以帮助实现高ASC精度与有限的标记微调数据。此外,我们发现,集成的SSL模型微调不同的策略有助于进一步提高性能。为了满足低复杂度的要求,我们使用知识蒸馏将自监督知识从大型教师模型转移到高效的学生模型。实验结果表明,教师自我监督有效地提高了学生模型的分类精度。我们性能最好的系统获得了56.7%的平均准确率。摘要:Acoustic scene classification (ASC) predominantly relies on supervised approaches. However, acquiring labeled data for training ASC models is often costly and time-consuming. Recently, self-supervised learning (SSL) has emerged as a powerful method for extracting features from unlabeled audio data, benefiting many downstream audio tasks. This paper proposes a data-efficient and low-complexity ASC system by leveraging self-supervised audio representations extracted from general-purpose audio datasets. We introduce BEATs, an audio SSL pre-trained model, to extract the general representations from AudioSet. Through extensive experiments, it has been demonstrated that the self-supervised audio representations can help to achieve high ASC accuracy with limited labeled fine-tuning data. Furthermore, we find that ensembling the SSL models fine-tuned with different strategies contributes to a further performance improvement. To meet low-complexity requirements, we use knowledge distillation to transfer the self-supervised knowledge from large teacher models to an efficient student model. The experimental results suggest that the self-supervised teachers effectively improve the classification accuracy of the student model. Our best-performing system obtains an average accuracy of 56.7%.

【18】 CoopASD: Cooperative Machine Anomalous Sound Detection with Privacy Concerns
标题: CoopASD:存在隐私问题的协作机器异常声音检测
作者:Anbai Jiang,Yuchen Shi,Pingyi Fan,Wei-Qiang Zhang,Jia Liu
备注:Accepted by GLOBECOM 2024
链接:点击下载PDF文件
摘要:机器异常声音检测(ASD)已成为工业物联网(IIoT)中最有前途的应用之一,因为它在降低故障风险和提高生产效率方面具有前所未有的功效。以前的工作主要是研究机器ASD任务集中设置下。然而,在分散设置下开发ASD系统在实践中至关重要,因为机器数据分散在各个工厂中,并且由于隐私问题,数据不应明确共享。为了使这些工厂能够合作开发可扩展的ASD模型,同时保护他们的隐私,我们提出了一个名为CoopASD的新框架,每个工厂在其本地数据集上训练ASD模型,中央服务器定期聚合这些本地模型。我们采用预先训练的模型作为ASD模型的骨干,以提高其鲁棒性,并开发专门的技术来在完全非iid和域偏移设置下稳定模型。与以前在集中式环境中训练的最先进的(SOTA)模型相比,CoopASD展示了具有竞争力的结果,其退化可以忽略不计,仅为0.08%。我们还进行了广泛的消融研究,以证明CoopASD的有效性。摘要:Machine anomalous sound detection (ASD) has emerged as one of the most promising applications in the Industrial Internet of Things (IIoT) due to its unprecedented efficacy in mitigating risks of malfunctions and promoting production efficiency. Previous works mainly investigated the machine ASD task under centralized settings. However, developing the ASD system under decentralized settings is crucial in practice, since the machine data are dispersed in various factories and the data should not be explicitly shared due to privacy concerns. To enable these factories to cooperatively develop a scalable ASD model while preserving their privacy, we propose a novel framework named CoopASD, where each factory trains an ASD model on its local dataset, and a central server aggregates these local models periodically. We employ a pre-trained model as the backbone of the ASD model to improve its robustness and develop specialized techniques to stabilize the model under a completely non-iid and domain shift setting. Compared with previous state-of-the-art (SOTA) models trained in centralized settings, CoopASD showcases competitive results with negligible degradation of 0.08%. We also conduct extensive ablation studies to demonstrate the effectiveness of CoopASD.

【19】 VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
标题: SecureTailor:用于基于扩散的个性化文本到语音的轻量级插件适配器
作者:Heeseung Kim,Sang-gil Lee,Jiheum Yeom,Che Hyun Lee,Sungwon Kim,Sungroh Yoon
链接:点击下载PDF文件
摘要:我们提出了VoiceTailor,一个参数高效的扬声器自适应文本到语音(TTS)系统,通过配备一个预先训练的基于扩散的TTS模型与个性化的适配器。VoiceTailor根据权重变化率分析确定从适配器中受益的关键模块。我们利用低秩自适应(LoRA)作为一种参数有效的自适应方法,并将适配器纳入预训练扩散解码器的关键模块。为了实现强大的自适应性能,很少的参数,我们探讨了各种指导技术的说话人自适应和调查的最佳策略,以加强说话人的信息。VoiceTailor通过仅微调总参数的0.25%来展示与现有自适应TTS模型相当的说话人自适应性能。VoiceTailor在适应各种真实世界的扬声器时表现出强大的鲁棒性,如演示中所示。摘要:We propose VoiceTailor, a parameter-efficient speaker-adaptive text-to-speech (TTS) system, by equipping a pre-trained diffusion-based TTS model with a personalized adapter. VoiceTailor identifies pivotal modules that benefit from the adapter based on a weight change ratio analysis. We utilize Low-Rank Adaptation (LoRA) as a parameter-efficient adaptation method and incorporate the adapter into pivotal modules of the pre-trained diffusion decoder. To achieve powerful adaptation performance with few parameters, we explore various guidance techniques for speaker adaptation and investigate the best strategies to strengthen speaker information. VoiceTailor demonstrates comparable speaker adaptation performance to existing adaptive TTS models by fine-tuning only 0.25 % of the total parameters. VoiceTailor shows strong robustness when adapting to a wide range of real-world speakers, as shown in the demo.

【20】 Physics-Informed Machine Learning For Sound Field Estimation
标题: 基于物理知识的机器学习用于声学场估计
作者:Shoichi Koyama,Juliano G. C. Ribeiro,Tomohiko Nakamura,Natsuki Ueno,Mirco Pezzoli
备注:Accepted to IEEE Signal Processing Magazine, Special Issue on Model-based and Data-Driven Audio Signal Processing
链接:点击下载PDF文件
摘要:关于空间声音估计的研究领域,即,诸如声压的声音的物理量的分布被称为声场估计,其是与空间音频处理相关的各种应用技术的基础。声场估计问题在简化的场景中被公式化为机器学习中的函数插值问题。然而,高估计性能不能通过简单地应用仅依赖于数据的一般插值技术来预期。声场的物理性质是有用的先验信息,并且将它们纳入估计中被认为是极其重要的。在这篇文章中,我们介绍了用于声场估计的物理信息机器学习(PIML)的基本原理,并概述了当前基于PIML的声场估计方法。摘要:The area of study concerning the estimation of spatial sound, i.e., the distribution of a physical quantity of sound such as acoustic pressure, is called sound field estimation, which is the basis for various applied technologies related to spatial audio processing. The sound field estimation problem is formulated as a function interpolation problem in machine learning in a simplified scenario. However, high estimation performance cannot be expected by simply applying general interpolation techniques that rely only on data. The physical properties of sound fields are useful a priori information, and it is considered extremely important to incorporate them into the estimation. In this article, we introduce the fundamentals of physics-informed machine learning (PIML) for sound field estimation and overview current PIML-based sound field estimation methods.

【21】 StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
标题: StyleSpeech:预训练的可控文本到语音的参数高效微调
作者:Haowei Lou,Helen Paik,Wen Hu,Lina Yao
链接:点击下载PDF文件
摘要:本文介绍了一种新颖的文本到语音转换系统StyleSpeech,它提高了合成语音的自然度和准确度。基于现有的TTS技术,StyleSpeech采用了独特的Style Decorator结构,使深度学习模型能够同时学习风格和音素特征,通过Lower Rank Adaptation~(LoRA)原则提高适应性和效率。LoRA允许在预先训练的模型中有效地适应风格特征。此外,我们引入了一种新的自动评估指标,LLM-Guided Mean Opinion Score(LLM-MOS),它采用大型语言模型,为自动评估TTS系统性能提供了一个客观而强大的协议。对基准数据集的广泛测试表明,我们的方法在产生自然,准确和高质量的语音方面明显优于现有的最先进的基线方法。这些进步不仅突破了当前TTS系统功能的界限,还促进了TTS系统在更动态和专业化的应用,例如交互式虚拟助理、自适应有声读物和游戏定制语音。语音样本可以在https: style-speech.vercel.app上找到摘要:This paper introduces StyleSpeech, a novel Text-to-Speech~(TTS) system that enhances the naturalness and accuracy of synthesized speech. Building upon existing TTS technologies, StyleSpeech incorporates a unique Style Decorator structure that enables deep learning models to simultaneously learn style and phoneme features, improving adaptability and efficiency through the principles of Lower Rank Adaptation~(LoRA). LoRA allows efficient adaptation of style features in pre-trained models. Additionally, we introduce a novel automatic evaluation metric, the LLM-Guided Mean Opinion Score (LLM-MOS), which employs large language models to offer an objective and robust protocol for automatically assessing TTS system performance. Extensive testing on benchmark datasets shows that our approach markedly outperforms existing state-of-the-art baseline methods in producing natural, accurate, and high-quality speech. These advancements not only pushes the boundaries of current TTS system capabilities, but also facilitate the application of TTS system in more dynamic and specialized, such as interactive virtual assistants, adaptive audiobooks, and customized voice for gaming. Speech samples can be found in https: style-speech.vercel.app

【22】 Global-Local Distillation Network-Based Audio-Visual Speaker Tracking with Incomplete Modalities
标题: 基于全球-本地蒸馏网络的不完整模式视听说话人跟踪
作者:Yidi Li,Yihan Li,Yixin Guo,Bin Ren,Zhenhuan Xu,Hao Guo,Hong Liu,Nicu Sebe
备注:Audio-Visual Speaker Tracking with Incomplete Modalities
链接:点击下载PDF文件
摘要:在说话人跟踪研究中,多模态数据的融合和互补是提高跟踪系统准确性和鲁棒性的关键策略。然而,不完整的方式跟踪仍然是一个具有挑战性的问题,由于噪声观测造成的闭塞,噪声和传感器故障。特别是当多模态数据存在缺失时,现有的多模态融合方法的性能往往会下降。为此,我们提出了一个全球本地蒸馏为基础的...(GLDTracker)强大的视听扬声器跟踪。GLDTracker由师生蒸馏模型驱动,能够灵活融合来自每种模态的不完整信息。教师网络处理摄像头和麦克风阵列捕获的全局信号,学生网络处理受视觉遮挡和丢失音频通道影响的本地信息。通过将知识从教师转移到学生,学生网络可以更好地适应观察不完整的复杂动态场景。在学生网络中,构建了一个基于生成对抗网络的全局特征重构模块,用于从缺失局部信息的特征嵌入中重构全局特征。此外,引入多模态多层次融合注意力来整合不完整特征和重构特征,充分利用视听特征和全局-局部特征的互补性和一致性。在AV16.3数据集上的实验结果表明,所提出的GLDTracker优于现有的最先进的视听...,并在标准和不完整的模态数据集上实现了领先的性能,突出了其在复杂条件下的优越性和鲁棒性。代码和模型将可用。摘要:In speaker tracking research, integrating and complementing multi-modal data is a crucial strategy for improving the accuracy and robustness of tracking systems. However, tracking with incomplete modalities remains a challenging issue due to noisy observations caused by occlusion, acoustic noise, and sensor failures. Especially when there is missing data in multiple modalities, the performance of existing multi-modal fusion methods tends to decrease. To this end, we propose a Global-Local Distillation-based Tracker (GLDTracker) for robust audio-visual speaker tracking. GLDTracker is driven by a teacher-student distillation model, enabling the flexible fusion of incomplete information from each modality. The teacher network processes global signals captured by camera and microphone arrays, and the student network handles local information subject to visual occlusion and missing audio channels. By transferring knowledge from teacher to student, the student network can better adapt to complex dynamic scenes with incomplete observations. In the student network, a global feature reconstruction module based on the generative adversarial network is constructed to reconstruct global features from feature embedding with missing local information. Furthermore, a multi-modal multi-level fusion attention is introduced to integrate the incomplete feature and the reconstructed feature, leveraging the complementarity and consistency of audio-visual and global-local features. Experimental results on the AV16.3 dataset demonstrate that the proposed GLDTracker outperforms existing state-of-the-art audio-visual trackers and achieves leading performance on both standard and incomplete modalities datasets, highlighting its superiority and robustness in complex conditions. The code and models will be available.


机器翻译,仅供参考