今日论文合集:cs.SD语音16篇,eess.AS音频处理16篇。本文经arXiv每日学术速递授权转载
【1】Spiking Music: Audio Compression with Event Based Auto-encoders
标题:尖峰音乐:使用基于事件的自动编码器进行音频压缩
链接:https://arxiv.org/abs/2402.01571
作者:Martim Lisboa,Guillaume Bellec
摘要:大脑中的神经元通过称为尖峰的准时事件传递信息。尖峰的时间被认为携带了丰富的信息,但目前尚不清楚如何在数字系统中利用这一点。我们证明,基于事件的编码是有效的音频压缩。为了构建这种基于事件的表示,我们使用了深度二进制自动编码器,在高稀疏压力下,模型进入了一种使用稀疏矩阵存储算法更有效地存储二进制事件矩阵的状态。我们在大型MAESTRO钢琴录音数据集上对矢量量化自动编码器进行了测试。我们的“尖峰音乐压缩”算法不仅实现了竞争性的压缩/重建权衡,而且编码事件和钢琴键敲击之间的选择性和同步性在稀疏状态下出现而无需监督。
摘要:Neurons in the brain communicate information via punctual events called spikes. The timing of spikes is thought to carry rich information, but it is not clear how to leverage this in digital systems. We demonstrate that event-based encoding is efficient for audio compression. To build this event-based representation we use a deep binary auto-encoder, and under high sparsity pressure, the model enters a regime where the binary event matrix is stored more efficiently with sparse matrix storage algorithms. We test this on the large MAESTRO dataset of piano recordings against vector quantized auto-encoders. Not only does our "Spiking Music compression" algorithm achieve a competitive compression/reconstruction trade-off, but selectivity and synchrony between encoded events and piano key strikes emerge without supervision in the sparse regime.
【2】 Low-Resource Cross-Domain Singing Voice Synthesis via Reduced Self-Supervised Speech Representations标题:基于简化自监督语音表示的低资源跨域歌唱声合成作者:Panos Kakoulidis,Nikolaos Ellinas,Georgios Vamvoukakis,Myrsini Christidou,Alexandra Vioni,Georgia Maniati,Junkwang Oh,Gunu Jho,Inchul Hwang,Pirros Tsiakoulis,Aimilios Chalamandaris备注:Accepted to IEEE ICASSP SASB 2024摘要:在本文中,我们提出了一个唱歌的声音合成模型,Karaoker-SSL,这是一个典型的多扬声器声学模型,只训练文本和语音数据。它是一个低资源管道,不利用任何端到端的歌唱数据,因为它的声码器也是在语音数据上训练的。Karaoker-SSL以无监督的方式由自监督语音表示调节。我们通过只选择与任务相关的维度的子集来预处理这些表示。在训练过程中,通过多任务处理间接引导条件反射模块捕获风格信息。这是通过基于Conformer的模块实现的,该模块根据声学模型的输出预测音高。因此,Karaoker-SSL允许唱歌的声音合成,而不依赖于手工制作和特定领域的功能。也没有对文本对齐或歌词时间戳的要求。为了改善语音质量,我们采用了一个U-Net训练器,该训练器以目标说话人为条件,并遵循扩散GAN训练方案。摘要:In this paper, we propose a singing voice synthesis model, Karaoker-SSL, that is trained only on text and speech data as a typical multi-speaker acoustic model. It is a low-resource pipeline that does not utilize any singing data end-to-end, since its vocoder is also trained on speech data. Karaoker-SSL is conditioned by self-supervised speech representations in an unsupervised manner. We preprocess these representations by selecting only a subset of their task-correlated dimensions. The conditioning module is indirectly guided to capture style information during training by multi-tasking. This is achieved with a Conformer-based module, which predicts the pitch from the acoustic model's output. Thus, Karaoker-SSL allows singing voice synthesis without reliance on hand-crafted and domain-specific features. There are also no requirements for text alignments or lyrics timestamps. To refine the voice quality, we employ a U-Net discriminator that is conditioned on the target speaker and follows a Diffusion GAN training scheme.
【3】 A Data-Driven Analysis of Robust Automatic Piano Transcription作者:Drew Edwards,Simon Dixon,Emmanouil Benetos,Akira Maezawa,Yuta Kusaka备注:Accepted for publication in IEEE Signal Processing Letters on 31 Janurary, 2024摘要:近年来,由于新的数据集和建模技术,自动钢琴转录算法得到了显着改进。最近的发展主要集中在适应新的神经网络架构,如Transformer和Perceiver,以产生更精确的系统。在这项工作中,我们从训练数据的角度研究转录系统。通过测量它们在分布外注释钢琴数据上的性能,我们展示了这些模型如何严重过拟合训练数据的声学特性。我们为MAESTRO数据集创建了一组新的音频,这些音频通过Yamaha超豪华播放器在专业录音室录音环境中自动捕获。在使用MAESTRO数据集的原始和重新执行版本进行训练时,使用各种数据增强技术,我们在MAPS数据集上实现了最先进的音符起始准确度88.4 F1-score,而无需看到任何训练数据。随后,我们在一系列消融研究中分析了这些数据增强技术,以更好地了解它们对所得模型的影响。摘要:Algorithms for automatic piano transcription have improved dramatically in recent years due to new datasets and modeling techniques. Recent developments have focused primarily on adapting new neural network architectures, such as the Transformer and Perceiver, in order to yield more accurate systems. In this work, we study transcription systems from the perspective of their training data. By measuring their performance on out-of-distribution annotated piano data, we show how these models can severely overfit to acoustic properties of the training data. We create a new set of audio for the MAESTRO dataset, captured automatically in a professional studio recording environment via Yamaha Disklavier playback. Using various data augmentation techniques when training with the original and re-performed versions of the MAESTRO dataset, we achieve state-of-the-art note-onset accuracy of 88.4 F1-score on the MAPS dataset, without seeing any of its training data. We subsequently analyze these data augmentation techniques in a series of ablation studies to better understand their influence on the resulting models.
【4】 Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge标题:语音增强方法在第七届CHAME挑战UDASE任务中的主客观评价作者:Simon Leglaive,Matthieu Fraticelli,Hend ElGhazaly,Léonie Borne,Mostafa Sadeghi,Scott Wisdom,Manuel Pariente,John R. Hershey,Daniel Pressnitzer,Jon P. Barker摘要:使用人工生成的干净语音和噪声信号的混合物来训练用于语音增强的监督模型。然而,合成训练条件可能无法准确反映测试期间遇到的真实世界条件。当测试域与合成训练域显著不同时,这种差异可能导致性能不佳。为了解决这个问题,第七届CHiME挑战赛的UDASE任务旨在利用来自测试域的真实噪声语音记录进行语音增强模型的无监督域自适应。具体而言,该测试域对应于CHiME-5数据集,其特征在于在嘈杂和混响的家庭环境中进行的真实多扬声器和对话语音记录,其中地面真实干净语音信号不可用。在本文中,我们提出了提交给CHiME-7 UDASE任务的系统的客观和主观评价,并提供了结果分析。该分析揭示了有限的主观评级和最近提出的语音增强几个监督非侵入性的性能指标之间的相关性。相反,结果表明,更传统的侵入性客观指标可以用于使用为挑战开发的混响LibriCHiME-5数据集进行域内性能评估。主观评价表明,所有系统都成功地降低了背景噪声,但总是以增加失真为代价。在主观评估的四种语音增强方法中,与未处理的嘈杂语音相比,只有一种表现出整体质量的改善,突出了任务的难度。为CHiME-7 UDASE任务创建的工具和音频材料与社区共享。摘要:Supervised models for speech enhancement are trained using artificially generated mixtures of clean speech and noise signals. However, the synthetic training conditions may not accurately reflect real-world conditions encountered during testing. This discrepancy can result in poor performance when the test domain significantly differs from the synthetic training domain. To tackle this issue, the UDASE task of the 7th CHiME challenge aimed to leverage real-world noisy speech recordings from the test domain for unsupervised domain adaptation of speech enhancement models. Specifically, this test domain corresponds to the CHiME-5 dataset, characterized by real multi-speaker and conversational speech recordings made in noisy and reverberant domestic environments, for which ground-truth clean speech signals are not available. In this paper, we present the objective and subjective evaluations of the systems that were submitted to the CHiME-7 UDASE task, and we provide an analysis of the results. This analysis reveals a limited correlation between subjective ratings and several supervised nonintrusive performance metrics recently proposed for speech enhancement. Conversely, the results suggest that more traditional intrusive objective metrics can be used for in-domain performance evaluation using the reverberant LibriCHiME-5 dataset developed for the challenge. The subjective evaluation indicates that all systems successfully reduced the background noise, but always at the expense of increased distortion. Out of the four speech enhancement methods evaluated subjectively, only one demonstrated an improvement in overall quality compared to the unprocessed noisy speech, highlighting the difficulty of the task. The tools and audio material created for the CHiME-7 UDASE task are shared with the community.【5】 Bass Accompaniment Generation via Latent Diffusion作者:Marco Pasini,Maarten Grachten,Stefan Lattner摘要:自动生成适当匹配任意输入音轨的音乐的能力是一项具有挑战性的任务。我们提出了一种新的可控系统,用于产生单茎伴随音乐混合任意长度。在我们的方法的核心是音频自动编码器,有效地压缩音频波形样本到可逆的潜在表示,和一个条件的潜在扩散模型,作为输入的混合的潜在编码,并生成相应的干的潜在编码。为了提供对生成的样本的音色的控制,我们引入了一种技术,在扩散采样期间将潜在空间接地到用户提供的参考样式。为了进一步提高音频质量,我们采用无分类器指导,以避免在生成无界潜在空间时在高指导强度下的失真。我们训练我们的模型对混音和匹配低音干的数据集。定量实验表明,给定一个输入组合,该系统可以生成与用户指定的音色。我们的可控条件音频生成框架代表了创建生成AI工具以帮助音乐家进行音乐制作的重要一步。摘要:The ability to automatically generate music that appropriately matches an arbitrary input track is a challenging task. We present a novel controllable system for generating single stems to accompany musical mixes of arbitrary length. At the core of our method are audio autoencoders that efficiently compress audio waveform samples into invertible latent representations, and a conditional latent diffusion model that takes as input the latent encoding of a mix and generates the latent encoding of a corresponding stem. To provide control over the timbre of generated samples, we introduce a technique to ground the latent space to a user-provided reference style during diffusion sampling. For further improving audio quality, we adapt classifier-free guidance to avoid distortions at high guidance strengths when generating an unbounded latent space. We train our model on a dataset of pairs of mixes and matching bass stems. Quantitative experiments demonstrate that, given an input mix, the proposed system can generate basslines with user-specified timbres. Our controllable conditional audio generation framework represents a significant step forward in creating generative AI tools to assist musicians in music production.
【6】 On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification标题:论大规模自我监管向Few-Shot音频分类的可转移性作者:Calum Heggan,Sam Budgett,Timothy Hosepedales,Mehrdad Yeghoobi备注:Camera Ready version as submitted to ICASSP SASB Workshop 2024. 5 pages, 2 figures, 3 tables摘要:近年来,自监督学习因其从未标记数据中学习鲁棒特征表示的能力而表现出色。通过自我监督预训练的网络可以作为下游任务的有效特征提取器,包括Few-Shot学习。虽然Few-Shot学习的无监督方法的评估在图像中已经建立,但在声学中却明显缺乏。本研究通过评估大规模自监督模型在Few-Shot音频分类中的性能来解决这一差距。此外,我们探索模型的Few-Shot学习能力和其他下游任务基准之间的关系。我们的研究结果揭示了一些Few-Shot问题(如SpeechCommandsv2)的最新性能,以及基于语音的Few-Shot问题与各种下游音频任务之间的强相关性。摘要:In recent years, self-supervised learning has excelled for its capacity to learn robust feature representations from unlabelled data. Networks pretrained through self-supervision serve as effective feature extractors for downstream tasks, including Few-Shot Learning. While the evaluation of unsupervised approaches for few-shot learning is well-established in imagery, it is notably absent in acoustics. This study addresses this gap by assessing large-scale self-supervised models' performance in few-shot audio classification. Additionally, we explore the relationship between a model's few-shot learning capability and other downstream task benchmarks. Our findings reveal state-of-the-art performance in some few-shot problems such as SpeechCommandsv2, as well as strong correlations between speech-based few-shot problems and various downstream audio tasks.
【7】 STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition标题:Staa-Net:一种稀疏可转移的语音情感识别对抗性攻击作者:Yi Chang,Zhao Ren,Zixing Zhang,Xin Jing,Kun Qian,Xi Shao,Bin Hu,Tanja Schultz,Björn W. Schuller摘要:语音包含了丰富的人类情感信息,语音情感识别一直是人机交互领域的重要研究课题。SER模型的鲁棒性至关重要,特别是在私人医疗保健等隐私敏感和可靠性要求高的领域。最近,音频领域的深度神经网络对对抗性攻击的脆弱性已成为一个热门的研究领域。然而,先前在音频域中对抗性攻击的工作主要依赖于迭代的基于梯度的技术,这是耗时的,并且容易过度拟合特定的威胁模型。此外,具有更好的隐蔽性的稀疏扰动的探索在音频域中仍然是有限的。为了解决这些挑战,我们提出了一种基于生成器的攻击方法,以生成稀疏和可转移的对抗性示例,从而以端到端和有效的方式欺骗SER模型。我们在两个广泛使用的SER数据集上评估了我们的方法,即语音中的诱发情绪数据库(DEMoS)和交互式情感二元运动捕获(IEMOCAP),并证明了它能够以有效的方式生成成功的稀疏对抗性示例。此外,我们生成的对抗性示例具有与模型无关的可转移性,从而能够对高级受害者模型进行有效的对抗性攻击。摘要:Speech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has become a popular area of research. However, prior works on adversarial attacks in the audio domain primarily rely on iterative gradient-based techniques, which are time-consuming and prone to overfitting the specific threat model. Furthermore, the exploration of sparse perturbations, which have the potential for better stealthiness, remains limited in the audio domain. To address these challenges, we propose a generator-based attack method to generate sparse and transferable adversarial examples to deceive SER models in an end-to-end and efficient manner. We evaluate our method on two widely-used SER datasets, Database of Elicited Mood in Speech (DEMoS) and Interactive Emotional dyadic MOtion CAPture (IEMOCAP), and demonstrate its ability to generate successful sparse adversarial examples in an efficient manner. Moreover, our generated adversarial examples exhibit model-agnostic transferability, enabling effective adversarial attacks on advanced victim models.【8】 Streaming Sequence Transduction through Dynamic Compression作者:Weiting Tan,Yunmo Chen,Tongfei Chen,Guanghui Qin,Haoran Xu,Heidi C. Zhang,Benjamin Van Durme,Philipp Koehn摘要:我们介绍STAR(Stream Transduction with Anchor Representations),这是一种新的基于transformer的模型,旨在通过流进行有效的序列到序列的转导。STAR动态分割输入流以创建压缩的锚点表示,在自动语音识别(ASR)中实现了近乎无损的压缩(12倍),并优于现有方法。此外,STAR在同步语音到文本任务中展示了卓越的分割和延迟质量权衡,优化了延迟,内存占用和质量。摘要:We introduce STAR (Stream Transduction with Anchor Representations), a novel Transformer-based model designed for efficient sequence-to-sequence transduction over streams. STAR dynamically segments input streams to create compressed anchor representations, achieving nearly lossless compression (12x) in Automatic Speech Recognition (ASR) and outperforming existing methods. Moreover, STAR demonstrates superior segmentation and latency-quality trade-offs in simultaneous speech-to-text tasks, optimizing latency, memory footprint, and quality.【9】 AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents标题:AccentFold:非洲口音之旅--零距离ASR适应目标口音作者:Abraham Toluwase Owodunni,Aditya Yadavalli,Chris Chinenye Emezue,Tobi Olatunji,Clinton C Mbataku备注:Accepted to EACL Findings 2024摘要:尽管语音识别取得了进步,但口音语音仍然具有挑战性。虽然以前的方法都集中在建模技术或创建口音语音数据集,收集足够的数据,为众多的口音,特别是在非洲的情况下,仍然是不切实际的,由于其纯粹的多样性和相关的预算限制。为了解决这些挑战,我们提出了一种利用学习的口音嵌入之间的空间关系来改善下游自动语音识别(ASR)的方法。我们对代表100多种非洲口音的语音嵌入的探索性分析揭示了有趣的空间口音关系,突出了地理和谱系的相似性,捕获了一致的语音和形态学特征,所有这些都是从语音中经验性地学到的。此外,我们发现口音的关系,以前没有特点的民族语。通过实证评估,我们证明了AccentFold的有效性,对于分布外(OOD)口音,基于AccentFold信息进行训练的口音子集的采样优于强基线,相对WER提高了4.6%。AccentFold提出了一种很有前途的方法,可以提高口音语音的ASR性能,特别是在非洲口音的背景下,数据稀缺和预算限制带来了重大挑战。我们的研究结果强调了利用语言关系来改善zero-shot ASR对目标口音的适应的潜力。摘要:Despite advancements in speech recognition, accented speech remains challenging. While previous approaches have focused on modeling techniques or creating accented speech datasets, gathering sufficient data for the multitude of accents, particularly in the African context, remains impractical due to their sheer diversity and associated budget constraints. To address these challenges, we propose \textit{AccentFold}, a method that exploits spatial relationships between learned accent embeddings to improve downstream Automatic Speech Recognition (ASR). Our exploratory analysis of speech embeddings representing 100+ African accents reveals interesting spatial accent relationships highlighting geographic and genealogical similarities, capturing consistent phonological, and morphological regularities, all learned empirically from speech. Furthermore, we discover accent relationships previously uncharacterized by the Ethnologue. Through empirical evaluation, we demonstrate the effectiveness of AccentFold by showing that, for out-of-distribution (OOD) accents, sampling accent subsets for training based on AccentFold information outperforms strong baselines a relative WER improvement of 4.6%. AccentFold presents a promising approach for improving ASR performance on accented speech, particularly in the context of African accents, where data scarcity and budget constraints pose significant challenges. Our findings emphasize the potential of leveraging linguistic relationships to improve zero-shot ASR adaptation to target accents.【10】 Screening method for early dementia using sound objects as voice biomarkers标题:以声音对象为语音生物标志物的早期痴呆筛查方法作者:Adam Pluta,Zbigniew Pioch,Jędrzej Kardach,Piotr Zioło,Tomasz Kręcicki,Elżbieta Trypka摘要:引言:我们提出了一种早期痴呆症的筛查方法,使用基于声音对象作为声音生物标志物的特征。 研究方法:用于机器学习模型的最终数据集由266个观察结果组成,其中186个健康个体,46个被诊断患有阿尔茨海默氏症,34个患有MCI。这种方法基于受试者说出的持续元音/a/的六秒录音。这项工作的主要原始贡献是使用基于声音对象的精心制作的功能。这种方法允许人们首先以比标准频谱更准确的方式表示声音频谱,然后构建包含关于受试者对其声音的控制的相关信息的可解释特征。 结果:本研究获得的区分健康受试者和MCI受试者的ROC AUC为0.85,准确度为0.76。为了区分健康受试者和患有MCI或阿尔茨海默氏症的受试者,结果分别为0.84,0.77。 结论:基于声音对象的特征的使用使得即使在非常短的独立于语言的语音样本的记录上也能够筛查早期痴呆症。摘要:Introduction: We present a screening method for early dementia using features based on sound objects as voice biomarkers. Methods: The final dataset used for machine learning models consisted of 266 observations, with a distribution of 186 healthy individuals, 46 diagnosed with Alzheimer's, and 34 with MCI. This method is based on six-second recordings of the sustained vowel /a/ spoken by the subject. The main original contribution of this work is the use of carefully crafted features based on sound objects. This approach allows one to first represent the sound spectrum in a more accurate way than the standard spectrum, and then build interpretable features containing relevant information about subjects' control over their voice. Results: ROC AUC obtained in this work for distinguishing healthy subjects from those with MCI was 0.85, while accuracy was 0.76. For distinguishing between healthy subjects and those with either MCI or Alzheimer's the results were 0.84, 0.77, respectively. Conclusion: The use of features based on sound objects enables screening for early dementia even on very short recordings of language-independent voice samples.
【11】 EVA-GAN: Enhanced Various Audio Generation via Scalable Generative Adversarial Networks标题:EVA-GAN:通过可扩展的生成性对抗网络增强各种音频生成作者:Shijia Liao,Shiyi Lan,Arun George Zachariah摘要:大型模型的出现标志着机器学习的新时代,通过利用庞大的数据集来捕获和合成复杂的模式,它的性能明显优于小型模型。尽管取得了这些进步,但对缩放的探索,特别是在音频生成领域,仍然有限,以前的努力没有扩展到高保真(HiFi)44.1kHz域,并且在高频域中存在频谱不连续性和模糊性,同时缺乏对域外数据的鲁棒性。这些限制限制了模型对不同用例的适用性,包括音乐和歌唱生成。我们的工作通过可扩展生成对抗网络(EVA-GAN)引入了增强的各种音频生成,在频谱和高频重建以及域外数据性能的鲁棒性方面比以前的最先进技术有了显着的改进,通过采用36,000小时44.1kHz音频的广泛数据集,一个上下文感知模块,一个人在环工件测量工具包,并将模型扩展到大约2亿个参数。我们的工作演示可在https://double-blind-eva-gan.cc上查阅。摘要:The advent of Large Models marks a new era in machine learning, significantly outperforming smaller models by leveraging vast datasets to capture and synthesize complex patterns. Despite these advancements, the exploration into scaling, especially in the audio generation domain, remains limited, with previous efforts didn't extend into the high-fidelity (HiFi) 44.1kHz domain and suffering from both spectral discontinuities and blurriness in the high-frequency domain, alongside a lack of robustness against out-of-domain data. These limitations restrict the applicability of models to diverse use cases, including music and singing generation. Our work introduces Enhanced Various Audio Generation via Scalable Generative Adversarial Networks (EVA-GAN), yields significant improvements over previous state-of-the-art in spectral and high-frequency reconstruction and robustness in out-of-domain data performance, enabling the generation of HiFi audios by employing an extensive dataset of 36,000 hours of 44.1kHz audio, a context-aware module, a Human-In-The-Loop artifact measurement toolkit, and expands the model to approximately 200 million parameters. Demonstrations of our work are available at https://double-blind-eva-gan.cc.
【12】 BAT: Learning to Reason about Spatial Sounds with Large Language Models作者:Zhisheng Zheng,Puyuan Peng,Ziyang Ma,Xie Chen,Eunsol Choi,David Harwath备注:Preprint, work in progress摘要:空间声音推理是人类的一项基本技能,使我们能够根据声音导航和解释周围环境。在本文中,我们提出了BAT,它结合了双耳声学场景分析模型的空间声音感知能力与大型语言模型(LLM)的自然语言推理能力,以复制这种先天能力。为了解决现有空间声音数据集的缺乏问题,我们使用AudioSet和SoundSpaces 2.0合成了一个双耳音频数据集。接下来,我们开发了SpatialSoundQA,这是一个基于空间声音的问答数据集,提供了一系列QA任务,可以在空间声音感知和推理的各个方面对BAT进行训练。BAT的声学前端编码器是一种新型的空间音频编码器,称为空间音频频谱图Transformer或Spatial-AST,其本身在声音事件检测、空间定位和距离估计方面实现了强大的性能。通过将Spatial-AST与LLaMA-2 7 B模型集成,BAT超越了标准的声音事件定位和检测(SELD)任务,使模型能够推理其环境中声音之间的关系。我们的实验证明了BAT在空间声音感知和推理方面的卓越性能,展示了LLM在导航和解释复杂空间音频环境方面的巨大潜力。摘要:Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating Spatial-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT's superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments.【13】 How Paralingual are Paralinguistic Representations? A Case Study in Speech Emotion Recognition标题:副语言是怎样的副语言表达?语音情感识别的实例研究作者:Orchid Chetia Phukan,Gautam Siddharth Kashyap,Arun Balaji Buduru,Rajesh Sharma摘要:预训练模型(Pre-trained Models,PTM)在语音情感识别(Speech Emotion Recognition,SER)领域取得了重大进展。SER是一个应用范围从人机交互到医疗保健的领域。最近的研究已经利用各种PTM表示作为SER下游模型的输入特征。专门针对非语言任务进行预训练的PTM已经获得了SER的最先进(SOTA)性能。然而,这种PTM尚未在多语言环境中对SER进行评估,并且只对英语进行了实验。因此,我们填补了这一空白,通过对五种PTM(TRILLsson,wav 2 vec 2,XLS-R,x-vector,Whisper)进行全面的比较研究,以评估跨语言PTM(TRILLsson)对SER的有效性。TRILLsson的代表在所有PTM中取得了最佳性能。这表明,TRILLsson能够有效地捕获语音数据的各种语言特征,以获得更好的SER。我们还表明,使用TRILLsson表示的下游模型在各种多语言数据集的准确性方面实现了SOTA性能。摘要:Pre-trained Models (PTMs) have facilitated substantial progress in the field of Speech Emotion Recognition (SER). SER is an area with applications ranging from HumanComputer Interaction to Healthcare. Recent studies have leveraged various PTM representations as input features for downstream models for SER. PTM specifically pre-trained for paralinguistic tasks have obtained state-of-the-art (SOTA) performance for SER. However, such PTM haven't been evaluated for SER in multilingual settings and experimented only with English. So, we fill this gap, by performing a comprehensive comparative study of five PTMs (TRILLsson, wav2vec2, XLS-R, x-vector, Whisper) for assessing the effectiveness of paralingual PTM (TRILLsson) for SER across multiple languages. Representations from TRILLsson achieved the best performance among all the PTMs. This demonstrates that TRILLsson is able to effectively capture the various paralinguistic features from speech data for better SER. We also show that downstream models using TRILLsson representations achieve SOTA performance in terms of accuracy across various multi-lingual datasets.
【14】 Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen标题:Del VisualAuditivo:Sonorización de escenas Guiada Por Imagen作者:María Sánchez,Laura Fernández,Julián Arias,Mateo Cámara,Giulia Comini,Adam Gabrys,José Luis Blanco,Juan Ignacio Godino,Luis Alfonso Hernández备注:10 pages, in Spanish, Tecniac\'ustica摘要:图像、视频、文本和音频生成技术的最新进展以及公众对它们的使用正在导致新形式的内容生成。通常,每一种方式都是单独处理的,这造成了限制。视觉序列的自动声音记录是多模态内容自动生成的最大挑战之一。我们提出了一个处理流程,从视频中提取的图像开始,能够发出声音。我们使用预先训练的模型,这些模型采用复杂的编码器,对比学习和多种模态,允许序列的复杂表示用于其声化。所提出的方案提出了音频映射和文本指导的不同可能性。我们在从商业视频游戏中提取的帧和从Freesound平台提取的声音的数据集上评估了该方案。主观测试证明,该方案能够自动生成和分配音频图像和方便。此外,它很好地适应用户的喜好,和建议的客观指标表现出较高的相关性与主观评级。摘要:Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality was approached separately, which poses limitations. The automatic sound recording of visual sequences is one of the greatest challenges for the automatic generation of multimodal content. We present a processing flow that, starting from images extracted from videos, is able to sound them. We work with pre-trained models that employ complex encoders, contrastive learning, and multiple modalities, allowing complex representations of the sequences for their sonorization. The proposed scheme proposes different possibilities for audio mapping and text guidance. We evaluated the scheme on a dataset of frames extracted from a commercial video game and sounds extracted from the Freesound platform. Subjective tests have evidenced that the proposed scheme is able to generate and assign audios automatically and conveniently to images. Moreover, it adapts well to user preferences, and the proposed objective metrics show a high correlation with the subjective ratings.
【15】 Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations标题:同时使用上下文和语音表征从原始音频信号中学习语义信息作者:Jaeyeon Kim,Injune Hwang,Kyogu Lee备注:Accepted to ICASSP 2024摘要:我们提出了一个框架来学习语义从原始音频信号使用两种类型的表示,编码上下文和语音信息分别。具体来说,我们引入了一个语音到单元处理流水线,捕获两种类型的表示具有不同的时间分辨率。对于语言模型,我们采用了双通道架构,将这两种类型的表示。我们还提出了新的训练目标,掩码上下文重建和掩码上下文预测,推动模型有效地学习语义。在Zero Resource Speech Benchmark 2021和Fluent Speech Command数据集的sSIMI度量上的实验表明,我们的框架比仅使用一种表示训练的模型更好地学习语义。摘要:We propose a framework to learn semantics from raw audio signals using two types of representations, encoding contextual and phonetic information respectively. Specifically, we introduce a speech-to-unit processing pipeline that captures two types of representations with different time resolutions. For the language model, we adopt a dual-channel architecture to incorporate both types of representation. We also present new training objectives, masked context reconstruction and masked context prediction, that push models to learn semantics effectively. Experiments on the sSIMI metric of Zero Resource Speech Benchmark 2021 and Fluent Speech Command dataset show our framework learns semantics better than models trained with only one type of representation.
【16】 An Intra-BRNN and GB-RVQ Based END-TO-END Neural Audio Codec标题:一种基于帧内BRNN和GB-RVQ的端到端神经音频编解码器作者:Linping Xu,Jiawei Jiang,Dejun Zhang,Xianjun Xia,Li Chen,Yijian Xiao,Piao Ding,Shenyi Song,Sixing Yin,Ferdous Sohel摘要:最近,神经网络已被证明是有效的,在执行语音编码任务在低比特率。然而,帧内相关性的利用不足和量化器的误差具体降低了重建的音频质量。为了提高编码质量,我们提出了一种端到端的神经语音编解码器,即CBRC(卷积和双向递归神经编解码器)。使用1D-CNN和Intra-BRNN的交织结构被设计成更有效地利用帧内相关性。此外,分组和波束搜索残差矢量量化器(GB-RVQ)被用来减少量化噪声。CBRC每20 ms编码一次音频,没有额外的延迟,适合实时通信。实验结果表明,所提出的编解码器的优越性时,比较银监会在3 kbps的Opus在12 kbps。摘要:Recently, neural networks have proven to be effective in performing speech coding task at low bitrates. However, under-utilization of intra-frame correlations and the error of quantizer specifically degrade the reconstructed audio quality. To improve the coding quality, we present an end-to-end neural speech codec, namely CBRC (Convolutional and Bidirectional Recurrent neural Codec). An interleaved structure using 1D-CNN and Intra-BRNN is designed to exploit the intra-frame correlations more efficiently. Furthermore, Group-wise and Beam-search Residual Vector Quantizer (GB-RVQ) is used to reduce the quantization noise. CBRC encodes audio every 20ms with no additional latency, which is suitable for real-time communication. Experimental results demonstrate the superiority of the proposed codec when comparing CBRC at 3kbps with Opus at 12kbps.
【1】 BAT: Learning to Reason about Spatial Sounds with Large Language Models作者:Zhisheng Zheng,Puyuan Peng,Ziyang Ma,Xie Chen,Eunsol Choi,David Harwath备注:Preprint, work in progress摘要:空间声音推理是人类的一项基本技能,使我们能够根据声音导航和解释周围环境。在本文中,我们提出了BAT,它结合了双耳声学场景分析模型的空间声音感知能力与大型语言模型(LLM)的自然语言推理能力,以复制这种先天能力。为了解决现有空间声音数据集的缺乏问题,我们使用AudioSet和SoundSpaces 2.0合成了一个双耳音频数据集。接下来,我们开发了SpatialSoundQA,这是一个基于空间声音的问答数据集,提供了一系列QA任务,可以在空间声音感知和推理的各个方面对BAT进行训练。BAT的声学前端编码器是一种新型的空间音频编码器,称为空间音频频谱图Transformer或Spatial-AST,其本身在声音事件检测、空间定位和距离估计方面实现了强大的性能。通过将Spatial-AST与LLaMA-2 7 B模型集成,BAT超越了标准的声音事件定位和检测(SELD)任务,使模型能够推理其环境中声音之间的关系。我们的实验证明了BAT在空间声音感知和推理方面的卓越性能,展示了LLM在导航和解释复杂空间音频环境方面的巨大潜力。摘要:Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating Spatial-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT's superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments.【2】 How Paralingual are Paralinguistic Representations? A Case Study in Speech Emotion Recognition标题:副语言是怎样的副语言表达?语音情感识别的实例研究作者:Orchid Chetia Phukan,Gautam Siddharth Kashyap,Arun Balaji Buduru,Rajesh Sharma摘要:预训练模型(Pre-trained Models,PTM)在语音情感识别(Speech Emotion Recognition,SER)领域取得了重大进展。SER是一个应用范围从人机交互到医疗保健的领域。最近的研究已经利用各种PTM表示作为SER下游模型的输入特征。专门针对非语言任务进行预训练的PTM已经获得了SER的最先进(SOTA)性能。然而,这种PTM尚未在多语言环境中对SER进行评估,并且只对英语进行了实验。因此,我们填补了这一空白,通过对五种PTM(TRILLsson,wav 2 vec 2,XLS-R,x-vector,Whisper)进行全面的比较研究,以评估跨语言PTM(TRILLsson)对SER的有效性。TRILLsson的代表在所有PTM中取得了最佳性能。这表明,TRILLsson能够有效地捕获语音数据的各种语言特征,以获得更好的SER。我们还表明,使用TRILLsson表示的下游模型在各种多语言数据集的准确性方面实现了SOTA性能。摘要:Pre-trained Models (PTMs) have facilitated substantial progress in the field of Speech Emotion Recognition (SER). SER is an area with applications ranging from HumanComputer Interaction to Healthcare. Recent studies have leveraged various PTM representations as input features for downstream models for SER. PTM specifically pre-trained for paralinguistic tasks have obtained state-of-the-art (SOTA) performance for SER. However, such PTM haven't been evaluated for SER in multilingual settings and experimented only with English. So, we fill this gap, by performing a comprehensive comparative study of five PTMs (TRILLsson, wav2vec2, XLS-R, x-vector, Whisper) for assessing the effectiveness of paralingual PTM (TRILLsson) for SER across multiple languages. Representations from TRILLsson achieved the best performance among all the PTMs. This demonstrates that TRILLsson is able to effectively capture the various paralinguistic features from speech data for better SER. We also show that downstream models using TRILLsson representations achieve SOTA performance in terms of accuracy across various multi-lingual datasets.
【3】 Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen标题:Del VisualAuditivo:Sonorización de escenas Guiada Por Imagen作者:María Sánchez,Laura Fernández,Julián Arias,Mateo Cámara,Giulia Comini,Adam Gabrys,José Luis Blanco,Juan Ignacio Godino,Luis Alfonso Hernández备注:10 pages, in Spanish, Tecniac\'ustica摘要:图像、视频、文本和音频生成技术的最新进展以及公众对它们的使用正在导致新形式的内容生成。通常,每一种方式都是单独处理的,这造成了限制。视觉序列的自动声音记录是多模态内容自动生成的最大挑战之一。我们提出了一个处理流程,从视频中提取的图像开始,能够发出声音。我们使用预先训练的模型,这些模型采用复杂的编码器,对比学习和多种模态,允许序列的复杂表示用于其声化。所提出的方案提出了音频映射和文本指导的不同可能性。我们在从商业视频游戏中提取的帧和从Freesound平台提取的声音的数据集上评估了该方案。主观测试证明,该方案能够自动生成和分配音频图像和方便。此外,它很好地适应用户的喜好,和建议的客观指标表现出较高的相关性与主观评级。摘要:Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality was approached separately, which poses limitations. The automatic sound recording of visual sequences is one of the greatest challenges for the automatic generation of multimodal content. We present a processing flow that, starting from images extracted from videos, is able to sound them. We work with pre-trained models that employ complex encoders, contrastive learning, and multiple modalities, allowing complex representations of the sequences for their sonorization. The proposed scheme proposes different possibilities for audio mapping and text guidance. We evaluated the scheme on a dataset of frames extracted from a commercial video game and sounds extracted from the Freesound platform. Subjective tests have evidenced that the proposed scheme is able to generate and assign audios automatically and conveniently to images. Moreover, it adapts well to user preferences, and the proposed objective metrics show a high correlation with the subjective ratings.
【4】 Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations标题:同时使用上下文和语音表征从原始音频信号中学习语义信息作者:Jaeyeon Kim,Injune Hwang,Kyogu Lee备注:Accepted to ICASSP 2024摘要:我们提出了一个框架来学习语义从原始音频信号使用两种类型的表示,编码上下文和语音信息分别。具体来说,我们引入了一个语音到单元处理流水线,捕获两种类型的表示具有不同的时间分辨率。对于语言模型,我们采用了双通道架构,将这两种类型的表示。我们还提出了新的训练目标,掩码上下文重建和掩码上下文预测,推动模型有效地学习语义。在Zero Resource Speech Benchmark 2021和Fluent Speech Command数据集的sSIMI度量上的实验表明,我们的框架比仅使用一种表示训练的模型更好地学习语义。摘要:We propose a framework to learn semantics from raw audio signals using two types of representations, encoding contextual and phonetic information respectively. Specifically, we introduce a speech-to-unit processing pipeline that captures two types of representations with different time resolutions. For the language model, we adopt a dual-channel architecture to incorporate both types of representation. We also present new training objectives, masked context reconstruction and masked context prediction, that push models to learn semantics effectively. Experiments on the sSIMI metric of Zero Resource Speech Benchmark 2021 and Fluent Speech Command dataset show our framework learns semantics better than models trained with only one type of representation.【5】 An Intra-BRNN and GB-RVQ Based END-TO-END Neural Audio Codec标题:一种基于帧内BRNN和GB-RVQ的端到端神经音频编解码器作者:Linping Xu,Jiawei Jiang,Dejun Zhang,Xianjun Xia,Li Chen,Yijian Xiao,Piao Ding,Shenyi Song,Sixing Yin,Ferdous Sohel摘要:最近,神经网络已被证明是有效的,在执行语音编码任务在低比特率。然而,帧内相关性的利用不足和量化器的误差具体降低了重建的音频质量。为了提高编码质量,我们提出了一种端到端的神经语音编解码器,即CBRC(卷积和双向递归神经编解码器)。使用1D-CNN和Intra-BRNN的交织结构被设计成更有效地利用帧内相关性。此外,分组和波束搜索残差矢量量化器(GB-RVQ)被用来减少量化噪声。CBRC每20 ms编码一次音频,没有额外的延迟,适合实时通信。实验结果表明,所提出的编解码器的优越性时,比较银监会在3 kbps的Opus在12 kbps。摘要:Recently, neural networks have proven to be effective in performing speech coding task at low bitrates. However, under-utilization of intra-frame correlations and the error of quantizer specifically degrade the reconstructed audio quality. To improve the coding quality, we present an end-to-end neural speech codec, namely CBRC (Convolutional and Bidirectional Recurrent neural Codec). An interleaved structure using 1D-CNN and Intra-BRNN is designed to exploit the intra-frame correlations more efficiently. Furthermore, Group-wise and Beam-search Residual Vector Quantizer (GB-RVQ) is used to reduce the quantization noise. CBRC encodes audio every 20ms with no additional latency, which is suitable for real-time communication. Experimental results demonstrate the superiority of the proposed codec when comparing CBRC at 3kbps with Opus at 12kbps.
【6】 Spiking Music: Audio Compression with Event Based Auto-encoders标题:Spiking Music:使用基于事件的自动编码器进行音频压缩作者:Martim Lisboa,Guillaume Bellec摘要:大脑中的神经元通过称为尖峰的准时事件传递信息。尖峰的时间被认为携带了丰富的信息,但目前尚不清楚如何在数字系统中利用这一点。我们证明,基于事件的编码是有效的音频压缩。为了构建这种基于事件的表示,我们使用了深度二进制自动编码器,在高稀疏压力下,模型进入了一种使用稀疏矩阵存储算法更有效地存储二进制事件矩阵的状态。我们在大型MAESTRO钢琴录音数据集上对矢量量化自动编码器进行了测试。我们的“尖峰音乐压缩”算法不仅实现了竞争性的压缩/重建权衡,而且编码事件和钢琴键敲击之间的选择性和同步性在稀疏状态下出现而无需监督。摘要:Neurons in the brain communicate information via punctual events called spikes. The timing of spikes is thought to carry rich information, but it is not clear how to leverage this in digital systems. We demonstrate that event-based encoding is efficient for audio compression. To build this event-based representation we use a deep binary auto-encoder, and under high sparsity pressure, the model enters a regime where the binary event matrix is stored more efficiently with sparse matrix storage algorithms. We test this on the large MAESTRO dataset of piano recordings against vector quantized auto-encoders. Not only does our "Spiking Music compression" algorithm achieve a competitive compression/reconstruction trade-off, but selectivity and synchrony between encoded events and piano key strikes emerge without supervision in the sparse regime.【7】 Low-Resource Cross-Domain Singing Voice Synthesis via Reduced Self-Supervised Speech Representations标题:基于简化自监督语音表示的低资源跨域歌唱语音合成作者:Panos Kakoulidis,Nikolaos Ellinas,Georgios Vamvoukakis,Myrsini Christidou,Alexandra Vioni,Georgia Maniati,Junkwang Oh,Gunu Jho,Inchul Hwang,Pirros Tsiakoulis,Aimilios Chalamandaris备注:Accepted to IEEE ICASSP SASB 2024摘要:在本文中,我们提出了一个唱歌的声音合成模型,Karaoker-SSL,这是一个典型的多扬声器声学模型,只训练文本和语音数据。它是一个低资源管道,不利用任何端到端的歌唱数据,因为它的声码器也是在语音数据上训练的。Karaoker-SSL以无监督的方式由自监督语音表示调节。我们通过只选择与任务相关的维度的子集来预处理这些表示。在训练过程中,通过多任务处理间接引导条件反射模块捕获风格信息。这是通过基于Conformer的模块实现的,该模块根据声学模型的输出预测音高。因此,Karaoker-SSL允许唱歌的声音合成,而不依赖于手工制作和特定领域的功能。也没有对文本对齐或歌词时间戳的要求。为了改善语音质量,我们采用了一个U-Net训练器,该训练器以目标说话人为条件,并遵循扩散GAN训练方案。摘要:In this paper, we propose a singing voice synthesis model, Karaoker-SSL, that is trained only on text and speech data as a typical multi-speaker acoustic model. It is a low-resource pipeline that does not utilize any singing data end-to-end, since its vocoder is also trained on speech data. Karaoker-SSL is conditioned by self-supervised speech representations in an unsupervised manner. We preprocess these representations by selecting only a subset of their task-correlated dimensions. The conditioning module is indirectly guided to capture style information during training by multi-tasking. This is achieved with a Conformer-based module, which predicts the pitch from the acoustic model's output. Thus, Karaoker-SSL allows singing voice synthesis without reliance on hand-crafted and domain-specific features. There are also no requirements for text alignments or lyrics timestamps. To refine the voice quality, we employ a U-Net discriminator that is conditioned on the target speaker and follows a Diffusion GAN training scheme.【8】 A Data-Driven Analysis of Robust Automatic Piano Transcription作者:Drew Edwards,Simon Dixon,Emmanouil Benetos,Akira Maezawa,Yuta Kusaka备注:Accepted for publication in IEEE Signal Processing Letters on 31 Janurary, 2024摘要:近年来,由于新的数据集和建模技术,自动钢琴转录算法得到了显着改进。最近的发展主要集中在适应新的神经网络架构,如Transformer和Perceiver,以产生更精确的系统。在这项工作中,我们从训练数据的角度研究转录系统。通过测量它们在分布外注释钢琴数据上的性能,我们展示了这些模型如何严重过拟合训练数据的声学特性。我们为MAESTRO数据集创建了一组新的音频,这些音频通过Yamaha超豪华播放器在专业录音室录音环境中自动捕获。在使用MAESTRO数据集的原始和重新执行版本进行训练时,使用各种数据增强技术,我们在MAPS数据集上实现了最先进的音符起始准确度88.4 F1-score,而无需看到任何训练数据。随后,我们在一系列消融研究中分析了这些数据增强技术,以更好地了解它们对所得模型的影响。摘要:Algorithms for automatic piano transcription have improved dramatically in recent years due to new datasets and modeling techniques. Recent developments have focused primarily on adapting new neural network architectures, such as the Transformer and Perceiver, in order to yield more accurate systems. In this work, we study transcription systems from the perspective of their training data. By measuring their performance on out-of-distribution annotated piano data, we show how these models can severely overfit to acoustic properties of the training data. We create a new set of audio for the MAESTRO dataset, captured automatically in a professional studio recording environment via Yamaha Disklavier playback. Using various data augmentation techniques when training with the original and re-performed versions of the MAESTRO dataset, we achieve state-of-the-art note-onset accuracy of 88.4 F1-score on the MAPS dataset, without seeing any of its training data. We subsequently analyze these data augmentation techniques in a series of ablation studies to better understand their influence on the resulting models.
【9】 Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge标题:语音增强方法在第七届CHAME挑战UDASE任务中的主客观评价作者:Simon Leglaive,Matthieu Fraticelli,Hend ElGhazaly,Léonie Borne,Mostafa Sadeghi,Scott Wisdom,Manuel Pariente,John R. Hershey,Daniel Pressnitzer,Jon P. Barker摘要:使用人工生成的干净语音和噪声信号的混合物来训练用于语音增强的监督模型。然而,合成训练条件可能无法准确反映测试期间遇到的真实世界条件。当测试域与合成训练域显著不同时,这种差异可能导致性能不佳。为了解决这个问题,第七届CHiME挑战赛的UDASE任务旨在利用来自测试域的真实噪声语音记录进行语音增强模型的无监督域自适应。具体而言,该测试域对应于CHiME-5数据集,其特征在于在嘈杂和混响的家庭环境中进行的真实多扬声器和对话语音记录,其中地面真实干净语音信号不可用。在本文中,我们提出了提交给CHiME-7 UDASE任务的系统的客观和主观评价,并提供了结果分析。该分析揭示了有限的主观评级和最近提出的语音增强几个监督非侵入性的性能指标之间的相关性。相反,结果表明,更传统的侵入性客观指标可以用于使用为挑战开发的混响LibriCHiME-5数据集进行域内性能评估。主观评价表明,所有系统都成功地降低了背景噪声,但总是以增加失真为代价。在主观评估的四种语音增强方法中,与未处理的嘈杂语音相比,只有一种表现出整体质量的改善,突出了任务的难度。为CHiME-7 UDASE任务创建的工具和音频材料与社区共享。摘要:Supervised models for speech enhancement are trained using artificially generated mixtures of clean speech and noise signals. However, the synthetic training conditions may not accurately reflect real-world conditions encountered during testing. This discrepancy can result in poor performance when the test domain significantly differs from the synthetic training domain. To tackle this issue, the UDASE task of the 7th CHiME challenge aimed to leverage real-world noisy speech recordings from the test domain for unsupervised domain adaptation of speech enhancement models. Specifically, this test domain corresponds to the CHiME-5 dataset, characterized by real multi-speaker and conversational speech recordings made in noisy and reverberant domestic environments, for which ground-truth clean speech signals are not available. In this paper, we present the objective and subjective evaluations of the systems that were submitted to the CHiME-7 UDASE task, and we provide an analysis of the results. This analysis reveals a limited correlation between subjective ratings and several supervised nonintrusive performance metrics recently proposed for speech enhancement. Conversely, the results suggest that more traditional intrusive objective metrics can be used for in-domain performance evaluation using the reverberant LibriCHiME-5 dataset developed for the challenge. The subjective evaluation indicates that all systems successfully reduced the background noise, but always at the expense of increased distortion. Out of the four speech enhancement methods evaluated subjectively, only one demonstrated an improvement in overall quality compared to the unprocessed noisy speech, highlighting the difficulty of the task. The tools and audio material created for the CHiME-7 UDASE task are shared with the community.
【10】 Bass Accompaniment Generation via Latent Diffusion作者:Marco Pasini,Maarten Grachten,Stefan Lattner摘要:自动生成适当匹配任意输入音轨的音乐的能力是一项具有挑战性的任务。我们提出了一种新的可控系统,用于产生单茎伴随音乐混合任意长度。在我们的方法的核心是音频自动编码器,有效地压缩音频波形样本到可逆的潜在表示,和一个条件的潜在扩散模型,作为输入的混合的潜在编码,并生成相应的干的潜在编码。为了提供对生成的样本的音色的控制,我们引入了一种技术,在扩散采样期间将潜在空间接地到用户提供的参考样式。为了进一步提高音频质量,我们采用无分类器指导,以避免在生成无界潜在空间时在高指导强度下的失真。我们训练我们的模型对混音和匹配低音干的数据集。定量实验表明,给定一个输入组合,该系统可以生成与用户指定的音色。我们的可控条件音频生成框架代表了创建生成AI工具以帮助音乐家进行音乐制作的重要一步。摘要:The ability to automatically generate music that appropriately matches an arbitrary input track is a challenging task. We present a novel controllable system for generating single stems to accompany musical mixes of arbitrary length. At the core of our method are audio autoencoders that efficiently compress audio waveform samples into invertible latent representations, and a conditional latent diffusion model that takes as input the latent encoding of a mix and generates the latent encoding of a corresponding stem. To provide control over the timbre of generated samples, we introduce a technique to ground the latent space to a user-provided reference style during diffusion sampling. For further improving audio quality, we adapt classifier-free guidance to avoid distortions at high guidance strengths when generating an unbounded latent space. We train our model on a dataset of pairs of mixes and matching bass stems. Quantitative experiments demonstrate that, given an input mix, the proposed system can generate basslines with user-specified timbres. Our controllable conditional audio generation framework represents a significant step forward in creating generative AI tools to assist musicians in music production.
【11】 On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification标题:论大规模自我监管向Few-Shot音频分类的可转移性作者:Calum Heggan,Sam Budgett,Timothy Hosepedales,Mehrdad Yeghoobi备注:Camera Ready version as submitted to ICASSP SASB Workshop 2024. 5 pages, 2 figures, 3 tables摘要:近年来,自监督学习因其从未标记数据中学习鲁棒特征表示的能力而表现出色。通过自我监督预训练的网络可以作为下游任务的有效特征提取器,包括Few-Shot学习。虽然Few-Shot学习的无监督方法的评估在图像中已经建立,但在声学中却明显缺乏。本研究通过评估大规模自监督模型在Few-Shot音频分类中的性能来解决这一差距。此外,我们探索模型的Few-Shot学习能力和其他下游任务基准之间的关系。我们的研究结果揭示了一些Few-Shot问题(如SpeechCommandsv2)的最新性能,以及基于语音的Few-Shot问题与各种下游音频任务之间的强相关性。摘要:In recent years, self-supervised learning has excelled for its capacity to learn robust feature representations from unlabelled data. Networks pretrained through self-supervision serve as effective feature extractors for downstream tasks, including Few-Shot Learning. While the evaluation of unsupervised approaches for few-shot learning is well-established in imagery, it is notably absent in acoustics. This study addresses this gap by assessing large-scale self-supervised models' performance in few-shot audio classification. Additionally, we explore the relationship between a model's few-shot learning capability and other downstream task benchmarks. Our findings reveal state-of-the-art performance in some few-shot problems such as SpeechCommandsv2, as well as strong correlations between speech-based few-shot problems and various downstream audio tasks.【12】 STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition标题:Staa-Net:一种稀疏可转移的语音情感识别对抗性攻击作者:Yi Chang,Zhao Ren,Zixing Zhang,Xin Jing,Kun Qian,Xi Shao,Bin Hu,Tanja Schultz,Björn W. Schuller摘要:语音包含了丰富的人类情感信息,语音情感识别一直是人机交互领域的重要研究课题。SER模型的鲁棒性至关重要,特别是在私人医疗保健等隐私敏感和可靠性要求高的领域。最近,音频领域的深度神经网络对对抗性攻击的脆弱性已成为一个热门的研究领域。然而,先前在音频域中对抗性攻击的工作主要依赖于迭代的基于梯度的技术,这是耗时的,并且容易过度拟合特定的威胁模型。此外,具有更好的隐蔽性的稀疏扰动的探索在音频域中仍然是有限的。为了解决这些挑战,我们提出了一种基于生成器的攻击方法,以生成稀疏和可转移的对抗性示例,从而以端到端和有效的方式欺骗SER模型。我们在两个广泛使用的SER数据集上评估了我们的方法,即语音中的诱发情绪数据库(DEMoS)和交互式情感二元运动捕获(IEMOCAP),并证明了它能够以有效的方式生成成功的稀疏对抗性示例。此外,我们生成的对抗性示例具有与模型无关的可转移性,从而能够对高级受害者模型进行有效的对抗性攻击。摘要:Speech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has become a popular area of research. However, prior works on adversarial attacks in the audio domain primarily rely on iterative gradient-based techniques, which are time-consuming and prone to overfitting the specific threat model. Furthermore, the exploration of sparse perturbations, which have the potential for better stealthiness, remains limited in the audio domain. To address these challenges, we propose a generator-based attack method to generate sparse and transferable adversarial examples to deceive SER models in an end-to-end and efficient manner. We evaluate our method on two widely-used SER datasets, Database of Elicited Mood in Speech (DEMoS) and Interactive Emotional dyadic MOtion CAPture (IEMOCAP), and demonstrate its ability to generate successful sparse adversarial examples in an efficient manner. Moreover, our generated adversarial examples exhibit model-agnostic transferability, enabling effective adversarial attacks on advanced victim models.
【13】 Streaming Sequence Transduction through Dynamic Compression作者:Weiting Tan,Yunmo Chen,Tongfei Chen,Guanghui Qin,Haoran Xu,Heidi C. Zhang,Benjamin Van Durme,Philipp Koehn摘要:我们介绍STAR(Stream Transduction with Anchor Representations),这是一种新的基于transformer的模型,旨在通过流进行有效的序列到序列的转导。STAR动态分割输入流以创建压缩的锚点表示,在自动语音识别(ASR)中实现了近乎无损的压缩(12倍),并优于现有方法。此外,STAR在同步语音到文本任务中展示了卓越的分割和延迟质量权衡,优化了延迟,内存占用和质量。摘要:We introduce STAR (Stream Transduction with Anchor Representations), a novel Transformer-based model designed for efficient sequence-to-sequence transduction over streams. STAR dynamically segments input streams to create compressed anchor representations, achieving nearly lossless compression (12x) in Automatic Speech Recognition (ASR) and outperforming existing methods. Moreover, STAR demonstrates superior segmentation and latency-quality trade-offs in simultaneous speech-to-text tasks, optimizing latency, memory footprint, and quality.【14】 AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents标题:AccentFold:非洲口音之旅--零距离ASR适应目标口音作者:Abraham Toluwase Owodunni,Aditya Yadavalli,Chris Chinenye Emezue,Tobi Olatunji,Clinton C Mbataku备注:Accepted to EACL Findings 2024摘要:尽管语音识别取得了进步,但口音语音仍然具有挑战性。虽然以前的方法都集中在建模技术或创建口音语音数据集,收集足够的数据,为众多的口音,特别是在非洲的情况下,仍然是不切实际的,由于其纯粹的多样性和相关的预算限制。为了解决这些挑战,我们提出了一种利用学习的口音嵌入之间的空间关系来改善下游自动语音识别(ASR)的方法。我们对代表100多种非洲口音的语音嵌入的探索性分析揭示了有趣的空间口音关系,突出了地理和谱系的相似性,捕获了一致的语音和形态学特征,所有这些都是从语音中经验性地学到的。此外,我们发现口音的关系,以前没有特点的民族语。通过实证评估,我们证明了AccentFold的有效性,对于分布外(OOD)口音,基于AccentFold信息进行训练的口音子集的采样优于强基线,相对WER提高了4.6%。AccentFold提出了一种很有前途的方法,可以提高口音语音的ASR性能,特别是在非洲口音的背景下,数据稀缺和预算限制带来了重大挑战。我们的研究结果强调了利用语言关系来改善zero-shot ASR对目标口音的适应的潜力。摘要:Despite advancements in speech recognition, accented speech remains challenging. While previous approaches have focused on modeling techniques or creating accented speech datasets, gathering sufficient data for the multitude of accents, particularly in the African context, remains impractical due to their sheer diversity and associated budget constraints. To address these challenges, we propose \textit{AccentFold}, a method that exploits spatial relationships between learned accent embeddings to improve downstream Automatic Speech Recognition (ASR). Our exploratory analysis of speech embeddings representing 100+ African accents reveals interesting spatial accent relationships highlighting geographic and genealogical similarities, capturing consistent phonological, and morphological regularities, all learned empirically from speech. Furthermore, we discover accent relationships previously uncharacterized by the Ethnologue. Through empirical evaluation, we demonstrate the effectiveness of AccentFold by showing that, for out-of-distribution (OOD) accents, sampling accent subsets for training based on AccentFold information outperforms strong baselines a relative WER improvement of 4.6%. AccentFold presents a promising approach for improving ASR performance on accented speech, particularly in the context of African accents, where data scarcity and budget constraints pose significant challenges. Our findings emphasize the potential of leveraging linguistic relationships to improve zero-shot ASR adaptation to target accents.【15】 Screening method for early dementia using sound objects as voice biomarkers标题:以声音对象为语音生物标志物的早期痴呆筛查方法作者:Adam Pluta,Zbigniew Pioch,Jędrzej Kardach,Piotr Zioło,Tomasz Kręcicki,Elżbieta Trypka摘要:引言:我们提出了一种早期痴呆症的筛查方法,使用基于声音对象作为声音生物标志物的特征。 研究方法:用于机器学习模型的最终数据集由266个观察结果组成,其中186个健康个体,46个被诊断患有阿尔茨海默氏症,34个患有MCI。这种方法基于受试者说出的持续元音/a/的六秒录音。这项工作的主要原始贡献是使用基于声音对象的精心制作的功能。这种方法允许人们首先以比标准频谱更准确的方式表示声音频谱,然后构建包含关于受试者对其声音的控制的相关信息的可解释特征。 结果:本研究获得的区分健康受试者和MCI受试者的ROC AUC为0.85,准确度为0.76。为了区分健康受试者和患有MCI或阿尔茨海默氏症的受试者,结果分别为0.84,0.77。 结论:基于声音对象的特征的使用使得即使在非常短的独立于语言的语音样本的记录上也能够筛查早期痴呆症。摘要:Introduction: We present a screening method for early dementia using features based on sound objects as voice biomarkers. Methods: The final dataset used for machine learning models consisted of 266 observations, with a distribution of 186 healthy individuals, 46 diagnosed with Alzheimer's, and 34 with MCI. This method is based on six-second recordings of the sustained vowel /a/ spoken by the subject. The main original contribution of this work is the use of carefully crafted features based on sound objects. This approach allows one to first represent the sound spectrum in a more accurate way than the standard spectrum, and then build interpretable features containing relevant information about subjects' control over their voice. Results: ROC AUC obtained in this work for distinguishing healthy subjects from those with MCI was 0.85, while accuracy was 0.76. For distinguishing between healthy subjects and those with either MCI or Alzheimer's the results were 0.84, 0.77, respectively. Conclusion: The use of features based on sound objects enables screening for early dementia even on very short recordings of language-independent voice samples.
【16】 EVA-GAN: Enhanced Various Audio Generation via Scalable Generative Adversarial Networks标题:EVA-GAN:通过可扩展的生成性对抗网络增强各种音频生成作者:Shijia Liao,Shiyi Lan,Arun George Zachariah摘要:大型模型的出现标志着机器学习的新时代,通过利用庞大的数据集来捕获和合成复杂的模式,它的性能明显优于小型模型。尽管取得了这些进步,但对缩放的探索,特别是在音频生成领域,仍然有限,以前的努力没有扩展到高保真(HiFi)44.1kHz域,并且在高频域中存在频谱不连续性和模糊性,同时缺乏对域外数据的鲁棒性。这些限制限制了模型对不同用例的适用性,包括音乐和歌唱生成。我们的工作通过可扩展生成对抗网络(EVA-GAN)引入了增强的各种音频生成,在频谱和高频重建以及域外数据性能的鲁棒性方面比以前的最先进技术有了显着的改进,通过采用36,000小时44.1kHz音频的广泛数据集,一个上下文感知模块,一个人在环工件测量工具包,并将模型扩展到大约2亿个参数。我们的工作演示可在https://double-blind-eva-gan.cc上查阅。摘要:The advent of Large Models marks a new era in machine learning, significantly outperforming smaller models by leveraging vast datasets to capture and synthesize complex patterns. Despite these advancements, the exploration into scaling, especially in the audio generation domain, remains limited, with previous efforts didn't extend into the high-fidelity (HiFi) 44.1kHz domain and suffering from both spectral discontinuities and blurriness in the high-frequency domain, alongside a lack of robustness against out-of-domain data. These limitations restrict the applicability of models to diverse use cases, including music and singing generation. Our work introduces Enhanced Various Audio Generation via Scalable Generative Adversarial Networks (EVA-GAN), yields significant improvements over previous state-of-the-art in spectral and high-frequency reconstruction and robustness in out-of-domain data performance, enabling the generation of HiFi audios by employing an extensive dataset of 36,000 hours of 44.1kHz audio, a context-aware module, a Human-In-The-Loop artifact measurement toolkit, and expands the model to approximately 200 million parameters. Demonstrations of our work are available at https://double-blind-eva-gan.cc.