今日论文合集:cs.SD语音5篇,eess.AS音频处理3篇。本文经arXiv每日学术速递授权转载
【1】Learning Disentangled Audio Representations through Controlled Synthesis
作者:Yusuf Brima,Ulf Krumnack,Simone Pika,Gunther Heidemann备注:12 pages, 12 figures, accepted as a Tiny paper at ICLR 2024摘要:本文处理的基准数据的稀缺性,在解开听觉表征学习。我们介绍SynTone,一个合成数据集,具有明确的地面实况解释因素,用于评估解纠缠技术。在SynTone上对最先进的方法进行基准测试,突出了其在方法评估方面的实用性。我们的研究结果强调了音频解缠的优势和局限性,激励了未来的研究。摘要:This paper tackles the scarcity of benchmarking data in disentangled auditory representation learning. We introduce SynTone, a synthetic dataset with explicit ground truth explanatory factors for evaluating disentanglement techniques. Benchmarking state-of-the-art methods on SynTone highlights its utility for method evaluation. Our results underscore strengths and limitations in audio disentanglement, motivating future research.【2】 APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding标题:APCodec:一种并行幅相谱编解码的神经音频编解码器作者:Yang Ai,Xiao-Hang Jiang,Ye-Xin Lu,Hui-Peng Du,Zhen-Hua Ling备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing摘要:本文介绍了一种新的神经音频编解码器,它以高波形采样率和低比特率为目标,称为APCodec,它无缝地集成了参数编解码器和波形编解码器的优点。APCodec通过像参数编解码器一样同时处理作为音频参数特性的幅度和相位谱,彻底改变了音频编码和解码的过程。它由一个编码器和一个解码器组成,以改进的ConvNeXt v2网络为骨干,通过基于残差矢量量化(RVQ)机制的量化器连接。编码器并行压缩音频幅度和相位谱,以降低的时间分辨率将它们合并成连续的潜在代码。该代码随后由量化器量化。最后,解码器并行重构音频的幅度和相位谱,并通过短时傅里叶逆变换得到解码波形。为了确保解码音频的保真度,如波形编解码器、频谱级损失、量化损失和基于生成对抗网络(GAN)的损失被共同地用于训练APCodec。为了支持低延迟流式推理,我们在APCodec中采用前馈层和因果卷积层,并结合知识蒸馏训练策略来提高解码音频的质量。实验结果证实,我们提出的APCodec可以编码48 kHz的音频比特率仅为6 kbps,解码音频的质量没有显着下降。在相同的比特率,我们提出的APCodec也表现出优越的解码音频质量和更快的生成速度相比,众所周知的编解码器,如SoundStream,Encodec,HiFi-Codec和AudioDec。摘要:This paper introduces a novel neural audio codec targeting high waveform sampling rates and low bitrates named APCodec, which seamlessly integrates the strengths of parametric codecs and waveform codecs. The APCodec revolutionizes the process of audio encoding and decoding by concurrently handling the amplitude and phase spectra as audio parametric characteristics like parametric codecs. It is composed of an encoder and a decoder with the modified ConvNeXt v2 network as the backbone, connected by a quantizer based on the residual vector quantization (RVQ) mechanism. The encoder compresses the audio amplitude and phase spectra in parallel, amalgamating them into a continuous latent code at a reduced temporal resolution. This code is subsequently quantized by the quantizer. Ultimately, the decoder reconstructs the audio amplitude and phase spectra in parallel, and the decoded waveform is obtained by inverse short-time Fourier transform. To ensure the fidelity of decoded audio like waveform codecs, spectral-level loss, quantization loss, and generative adversarial network (GAN) based loss are collectively employed for training the APCodec. To support low-latency streamable inference, we employ feed-forward layers and causal convolutional layers in APCodec, incorporating a knowledge distillation training strategy to enhance the quality of decoded audio. Experimental results confirm that our proposed APCodec can encode 48 kHz audio at bitrate of just 6 kbps, with no significant degradation in the quality of the decoded audio. At the same bitrate, our proposed APCodec also demonstrates superior decoded audio quality and faster generation speed compared to well-known codecs, such as SoundStream, Encodec, HiFi-Codec and AudioDec.【3】 Evaluating and Improving Continual Learning in Spoken Language Understanding作者:Muqiao Yang,Xiang Li,Umberto Cappellazzo,Shinji Watanabe,Bhiksha Raj摘要:持续学习已经成为包括口语理解(SLU)在内的各种任务中越来越重要的挑战。在SLU中,其目标是有效地处理新概念和不断变化的环境的出现。持续学习算法的评估通常涉及评估模型的稳定性,可塑性和可推广性作为标准的基本方面。然而,现有的持续学习指标主要集中在一个或两个属性。他们忽略了所有任务的整体表现,并且没有充分地理清模型中的可塑性与稳定性/可推广性的权衡。在这项工作中,我们提出了一个评价方法,提供了一个统一的评价稳定性,可塑性和概括性的持续学习。通过采用所提出的度量,我们演示了如何引入各种知识蒸馏可以提高SLU模型的这三个属性的不同方面。我们进一步表明,我们提出的度量在捕获持续学习中任务排序的影响方面更敏感,使其更适合实际用例场景。摘要:Continual learning has emerged as an increasingly important challenge across various tasks, including Spoken Language Understanding (SLU). In SLU, its objective is to effectively handle the emergence of new concepts and evolving environments. The evaluation of continual learning algorithms typically involves assessing the model's stability, plasticity, and generalizability as fundamental aspects of standards. However, existing continual learning metrics primarily focus on only one or two of the properties. They neglect the overall performance across all tasks, and do not adequately disentangle the plasticity versus stability/generalizability trade-offs within the model. In this work, we propose an evaluation methodology that provides a unified evaluation on stability, plasticity, and generalizability in continual learning. By employing the proposed metric, we demonstrate how introducing various knowledge distillations can improve different aspects of these three properties of the SLU model. We further show that our proposed metric is more sensitive in capturing the impact of task ordering in continual learning, making it better suited for practical use-case scenarios.【4】 Engraving Oriented Joint Estimation of Pitch Spelling and Local and Global Keys标题:面向雕刻的基音拼写与局部和全局关键字联合估计作者:Augustin Bouquillard,Florent Jacquemard备注:International Conference on Technologies for Music Notation and Representation (TENOR), Apr 2024, Zurich (CH), Switzerland摘要:我们重新审视音高拼写和音调猜测的问题,他们的联合估计从一个包含测量边界信息的文件的一个新算法。我们的算法不仅确定了一个全球的关键,但也本地的所有沿分析片。它使用动态编程技术来搜索最佳拼写,大致上,在雕刻的分数中显示的意外符号的数量。这个数字的评估与全局密钥和一些局部密钥的估计相结合,每个度量一个。这三个信息中的每一个都用于在多步骤过程中估计另一个。在单声道和钢琴数据集上进行的评估,总共包括216464个音符,显示出高度的准确性,无论是音高拼写(在巴赫语料库上平均为99.5%,在整个数据集上为98.2%)还是全局键签名估计(平均为93.0%,在钢琴数据集上为95.58%)。最初设计为音乐转录框架中的后端工具,该方法在与音乐符号处理相关的其他任务中也应该是有用的。摘要:We revisit the problems of pitch spelling and tonality guessing with a new algorithm for their joint estimation from a MIDI file including information about the measure boundaries. Our algorithm does not only identify a global key but also local ones all along the analyzed piece. It uses Dynamic Programming techniques to search for an optimal spelling in term, roughly, of the number of accidental symbols that would be displayed in the engraved score. The evaluation of this number is coupled with an estimation of the global key and some local keys, one for each measure. Each of the three informations is used for the estimation of the other, in a multi-steps procedure. An evaluation conducted on a monophonic and a piano dataset, comprising 216 464 notes in total, shows a high degree of accuracy, both for pitch spelling (99.5% on average on the Bach corpus and 98.2% on the whole dataset) and global key signature estimation (93.0% on average, 95.58% on the piano dataset). Designed originally as a backend tool in a music transcription framework, this method should also be useful in other tasks related to music notation processing.
【5】 AntiDeepFake: AI for Deep Fake Speech Recognition标题:AntiDeepFake:用于深度假语音识别的AI作者:Enkhtogtokh Togootogtokh,Christian Klasen备注:arXiv admin note: text overlap with arXiv:2308.12734 by other authors摘要:在这项研究中,我们提出了一种现代人工智能(AI)方法来识别deepfake语音,也称为生成AI克隆合成语音。我们提出的人工智能技术称为AntiDeepFake,包括从数据到评估的所有主要管道。我们提供的实验结果和分数,我们提出的所有方法。我们的方法的主要源代码可以在提供的链接中获得:https://github.com/enkhtogtokh/antideepfake repository。摘要:In this research study, we propose a modern artificial intelligence (AI) approach to recognize deepfake voice, also known as generative AI cloned synthetic voice. Our proposed AI technology, called AntiDeepFake, consists of all main pipelines from data to evaluation in the whole picture. We provide experimental results and scores for all our proposed methods. The main source code for our approach is available in the provided link: https://github.com/enkhtogtokh/antideepfake repository.【1】 Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model标题:小波域语音:一种简单有效的语音扩散模型加速方法作者:Xiangyu Zhang,Daijiao Liu,Hexin Liu,Qiquan Zhang,Hanyu Meng,Leibny Paola Garcia,Eng Siong Chng,Lina Yao摘要:最近,去噪扩散概率模型(DDPMs)已经在各种生成任务中取得了领先的性能。然而,在语音合成领域,虽然DDPM表现出令人印象深刻的性能,其长的训练时间和大量的推理成本阻碍了实际部署。现有的方法主要集中在提高推理速度,而加速训练的方法是与添加或定制语音相关的成本的关键因素,通常需要对模型进行复杂的修改,从而影响其普遍适用性。为了解决上述挑战,我们提出了一个查询:是否有可能通过修改语音信号本身来提高DDPM的训练/推理速度和性能?在本文中,我们通过简单地将生成目标重定向到小波域,将语音DDPM的训练和推理速度提高了一倍。该方法不仅在语音合成任务中实现了与原始模型相当或更好的性能,而且还展示了其通用性。通过研究和利用不同的小波基,我们的方法被证明是有效的,不仅在语音合成,而且在语音增强。摘要:Recently, Denoising Diffusion Probabilistic Models (DDPMs) have attained leading performances across a diverse range of generative tasks. However, in the field of speech synthesis, although DDPMs exhibit impressive performance, their long training duration and substantial inference costs hinder practical deployment. Existing approaches primarily focus on enhancing inference speed, while approaches to accelerate training a key factor in the costs associated with adding or customizing voices often necessitate complex modifications to the model, compromising their universal applicability. To address the aforementioned challenges, we propose an inquiry: is it possible to enhance the training/inference speed and performance of DDPMs by modifying the speech signal itself? In this paper, we double the training and inference speed of Speech DDPMs by simply redirecting the generative target to the wavelet domain. This method not only achieves comparable or superior performance to the original model in speech synthesis tasks but also demonstrates its versatility. By investigating and utilizing different wavelet bases, our approach proves effective not just in speech synthesis, but also in speech enhancement.【2】 Engraving Oriented Joint Estimation of Pitch Spelling and Local and Global Keys作者:Augustin Bouquillard,Florent Jacquemard备注:International Conference on Technologies for Music Notation and Representation (TENOR), Apr 2024, Zurich (CH), Switzerland摘要:我们重新审视音高拼写和音调猜测的问题,他们的联合估计从一个包含测量边界信息的文件的一个新算法。我们的算法不仅确定了一个全球的关键,但也本地的所有沿分析片。它使用动态编程技术来搜索最佳拼写,大致上,在雕刻的分数中显示的意外符号的数量。这个数字的评估与全局密钥和一些局部密钥的估计相结合,每个度量一个。这三个信息中的每一个都用于在多步骤过程中估计另一个。在单声道和钢琴数据集上进行的评估,总共包括216464个音符,显示出高度的准确性,无论是音高拼写(在巴赫语料库上平均为99.5%,在整个数据集上为98.2%)还是全局键签名估计(平均为93.0%,在钢琴数据集上为95.58%)。最初设计为音乐转录框架中的后端工具,该方法在与音乐符号处理相关的其他任务中也应该是有用的。摘要:We revisit the problems of pitch spelling and tonality guessing with a new algorithm for their joint estimation from a MIDI file including information about the measure boundaries. Our algorithm does not only identify a global key but also local ones all along the analyzed piece. It uses Dynamic Programming techniques to search for an optimal spelling in term, roughly, of the number of accidental symbols that would be displayed in the engraved score. The evaluation of this number is coupled with an estimation of the global key and some local keys, one for each measure. Each of the three informations is used for the estimation of the other, in a multi-steps procedure. An evaluation conducted on a monophonic and a piano dataset, comprising 216 464 notes in total, shows a high degree of accuracy, both for pitch spelling (99.5% on average on the Bach corpus and 98.2% on the whole dataset) and global key signature estimation (93.0% on average, 95.58% on the piano dataset). Designed originally as a backend tool in a music transcription framework, this method should also be useful in other tasks related to music notation processing.
【3】 AntiDeepFake: AI for Deep Fake Speech Recognition标题:AntiDeepFake:用于深度虚假语音识别的人工智能作者:Enkhtogtokh Togootogtokh,Christian Klasen备注:arXiv admin note: text overlap with arXiv:2308.12734 by other authors摘要:在这项研究中,我们提出了一种现代人工智能(AI)方法来识别deepfake语音,也称为生成AI克隆合成语音。我们提出的人工智能技术称为AntiDeepFake,包括从数据到评估的所有主要管道。我们提供的实验结果和分数,我们提出的所有方法。我们的方法的主要源代码可以在提供的链接中获得:https://github.com/enkhtogtokh/antideepfake repository。摘要:In this research study, we propose a modern artificial intelligence (AI) approach to recognize deepfake voice, also known as generative AI cloned synthetic voice. Our proposed AI technology, called AntiDeepFake, consists of all main pipelines from data to evaluation in the whole picture. We provide experimental results and scores for all our proposed methods. The main source code for our approach is available in the provided link: https://github.com/enkhtogtokh/antideepfake repository.