【1】Spectrogram-Based Detection of Auto-Tuned Vocals in Music Recordings标题:基于谱图的音乐录音中自调谐人声检测链接:https://arxiv.org/abs/2403.05380作者:Mahyar Gohari,Paolo Bestagini,Sergio Benini,Nicola Adami摘要:在音乐制作和音频处理领域,歌唱声音的自动音高校正(也称为自动调音)的实现显著改变了声乐表演的面貌。虽然自动调音技术为音乐家提供了调整音高并达到所需精度的能力,但它的使用也引发了关于其对真实性和艺术完整性影响的辩论。因此,检测和分析音乐录音中的自动调音人声已成为音乐学者,制作人和听众必不可少的。然而,据我们所知,在这方面事先没有作出任何努力。这项研究介绍了一种数据驱动的方法,利用三元组网络检测自动调谐的歌曲,通过创建一个由原始和自动调谐的音频片段组成的数据集来支持。实验结果表明,与Rawnet2相比,该方法在准确性和鲁棒性方面都具有优势,Rawnet2是一种端到端的反欺骗模型,广泛用于其他音频取证任务。摘要:In the domain of music production and audio processing, the implementation of automatic pitch correction of the singing voice, also known as Auto-Tune, has significantly transformed the landscape of vocal performance. While auto-tuning technology has offered musicians the ability to tune their vocal pitches and achieve a desired level of precision, its use has also sparked debates regarding its impact on authenticity and artistic integrity. As a result, detecting and analyzing Auto-Tuned vocals in music recordings has become essential for music scholars, producers, and listeners. However, to the best of our knowledge, no prior effort has been made in this direction. This study introduces a data-driven approach leveraging triplet networks for the detection of Auto-Tuned songs, backed by the creation of a dataset composed of original and Auto-Tuned audio clips. The experimental results demonstrate the superiority of the proposed method in both accuracy and robustness compared to Rawnet2, an end-to-end model proposed for anti-spoofing and widely used for other audio forensic tasks.
【2】 RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction标题:RFWave:用于音频波形重建的多频带整流链接:https://arxiv.org/abs/2403.05010作者:Peng Liu,Dongyang Dai摘要:生成建模的最新进展已经导致从不同表示的音频波形重构的显著进展。尽管扩散模型已经被用于重构音频波形,但是它们倾向于表现出延迟问题,因为它们在各个采样点的级别上操作并且需要相对大量的采样步骤。在这项研究中,我们介绍了RFWave,一种新的多频带整流方法,从梅尔频谱图重建高保真音频波形。RFWave的独特之处在于生成复杂的频谱图,并在帧级运行,同时处理所有子带以提高效率。得益于旨在实现平坦传输轨迹的整流流,RFWave仅需10个采样步骤。经验评估表明,RFWave实现了卓越的重建质量和卓越的计算效率,能够以比实时快90倍的速度生成音频。摘要:Recent advancements in generative modeling have led to significant progress in audio waveform reconstruction from diverse representations. Although diffusion models have been used for reconstructing audio waveforms, they tend to exhibit latency issues because they operate at the level of individual sample points and require a relatively large number of sampling steps. In this study, we introduce RFWave, a novel multi-band Rectified Flow approach that reconstructs high-fidelity audio waveforms from Mel-spectrograms. RFWave is distinctive for generating complex spectrograms and operating at the frame level, processing all subbands concurrently to enhance efficiency. Thanks to Rectified Flow, which aims for a flat transport trajectory, RFWave requires only 10 sampling steps. Empirical evaluations demonstrate that RFWave achieves exceptional reconstruction quality and superior computational efficiency, capable of generating audio at a speed 90 times faster than real-time. eess.AS音频处理【1】 Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks标题:基于深复卷积变换网络的双耳语音增强链接:https://arxiv.org/abs/2403.05393作者:Vikas Tokala,Eric Grinstein,Mike Brookes,Simon Doclo,Jesper Jensen,Patrick A. Naylor备注:Accepted to ICASSP 2024摘要:研究表明,在嘈杂的声学环境中,向辅助听音设备的用户提供双耳信号可以提高语音清晰度和空间意识。本文提出了一种使用复杂卷积神经网络的双耳语音增强方法,该网络具有编码器-解码器结构和复杂的多头注意力Transformer。该模型被训练以估计双耳听力设备的左耳和右耳通道的时频域中的各个复比掩模。该模型使用一种新的损失函数进行训练,该函数结合了空间信息的保留以及语音清晰度的提高和降噪。对单个目标说话人和各种类型的各向同性噪声的声学场景的仿真结果表明,与几种基线算法相比,该方法提高了估计的双耳语音可懂度,并更好地保留了双耳线索。摘要:Studies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement method using a complex convolutional neural network with an encoder-decoder architecture and a complex multi-head attention transformer. The model is trained to estimate individual complex ratio masks in the time-frequency domain for the left and right-ear channels of binaural hearing devices. The model is trained using a novel loss function that incorporates the preservation of spatial information along with speech intelligibility improvement and noise reduction. Simulation results for acoustic scenarios with a single target speaker and isotropic noise of various types show that the proposed method improves the estimated binaural speech intelligibility and preserves the binaural cues better in comparison with several baseline algorithms. 【2】 Robust Semantic Communications for Speech-to-Text Translation标题:面向语音到文本翻译的健壮语义通信链接:https://arxiv.org/abs/2403.05187作者:Zhenzi Weng,Zhijin Qin,Xiaoming Tao摘要:在本文中,我们提出了一个强大的语义通信系统,以实现语音到文本的翻译任务,命名为Ross-S2 T,通过提供必要的语义信息。特别地,开发了深度语义编码器,以直接将源语言中的语音压缩并转换为与目标语言相关联的文本语义特征,从而鼓励设计用于语音到文本翻译的支持深度学习的语义通信系统,该系统可以以端到端的方式联合训练。此外,针对语音输入失真的实际通信场景,提出了一种基于生成对抗网络(GAN)的深度语义补偿器,预测源语音中丢失的语义信息,同时生成目标语言的文本语义特征,为动态语音输入建立了一种鲁棒的语义传输机制.根据仿真结果,所提出的Ross-S2 T实现显着的语音到文本的翻译性能相比,传统的方法,并表现出对损坏的语音输入的高鲁棒性。摘要:In this paper, we propose a robust semantic communication system to achieve the speech-to-text translation task, named Ross-S2T, by delivering the essential semantic information. Particularly, a deep semantic encoder is developed to directly condense and convert the speech in the source language to the textual semantic features associated with the target language, thus encouraging the design of a deep learning-enabled semantic communication system for speech-to-text translation that can be jointly trained in an end-to-end manner. Moreover, to cope with the practical communication scenario when the input speech is corrupted, a novel generative adversarial network (GAN)-enabled deep semantic compensator is proposed to predict the lost semantic information in the source speech and produce the textual semantic features in the target language simultaneously, which establishes a robust semantic transmission mechanism for dynamic speech input. According to the simulation results, the proposed Ross-S2T achieves significant speech-to-text translation performance compared to the conventional approach and exhibits high robustness against the corrupted speech input.
【3】 AttentionStitch: How Attention Solves the Speech Editing Problem标题:AttentionStitch:注意力如何解决语音编辑问题链接:https://arxiv.org/abs/2403.04804作者:Antonios Alexos,Pierre Baldi备注:Accepted in Machine Learning for Audio workship in NeurIPS 2023摘要:从文本中生成自然的、高质量的语音是自然语言处理领域的一个具有挑战性的问题。除了语音生成之外,语音编辑也是一项至关重要的任务,它要求将编辑后的语音无缝地、不易察觉地集成到合成语音中。我们提出了一种新的语音编辑方法,它利用预先训练的文本到语音(TTS)模型,如FastSpeech 2,并在其上集成了一个双注意力块网络,以自动将合成的mel频谱图与编辑文本的mel频谱图合并。我们将此模型称为AttentionStitch,因为它利用注意力将音频样本缝合在一起。我们评估了所提出的AttentionStitch模型对单说话者和多说话者数据集(即LJSpeech和VCTK)的最新基线。我们通过涉及15名人类参与者的客观和主观评价测试证明了其优越的性能。AttentionStitch能够产生高质量的语音,即使是在训练过程中看不到的单词,同时自动运行,无需人工干预。此外,AttentionStitch在训练和推理过程中都很快,并且能够生成听起来像人类的编辑语音。摘要:The generation of natural and high-quality speech from text is a challenging problem in the field of natural language processing. In addition to speech generation, speech editing is also a crucial task, which requires the seamless and unnoticeable integration of edited speech into synthesized speech. We propose a novel approach to speech editing by leveraging a pre-trained text-to-speech (TTS) model, such as FastSpeech 2, and incorporating a double attention block network on top of it to automatically merge the synthesized mel-spectrogram with the mel-spectrogram of the edited text. We refer to this model as AttentionStitch, as it harnesses attention to stitch audio samples together. We evaluate the proposed AttentionStitch model against state-of-the-art baselines on both single and multi-speaker datasets, namely LJSpeech and VCTK. We demonstrate its superior performance through an objective and a subjective evaluation test involving 15 human participants. AttentionStitch is capable of producing high-quality speech, even for words not seen during training, while operating automatically without the need for human intervention. Moreover, AttentionStitch is fast during both training and inference and is able to generate human-sounding edited speech. 【4】 (Un)paired signal-to-signal translation with 1D conditional GANs标题:(Un)使用1D条件GAN的成对信号到信号转换链接:https://arxiv.org/abs/2403.04800作者:Eric Easthope摘要:我证明了具有对抗训练架构的一维(1D)条件生成对抗网络(cGAN)能够进行非配对信号到信号(“sig 2sig”)转换。使用具有1D层和更宽卷积内核的简化CycleGAN模型,镜像WaveGAN以将二维(2D)图像生成重新构建为1D音频生成,我表明,将2D图像到图像转换任务重新转换为1D信号到信号转换任务,使用深度卷积GANs是可能的,而无需对传统的U-Net模型和作为CycleGAN开发的对抗架构进行实质性修改。有了这个,我展示了一个小的可调数据集,1D CycleGAN模型看不到的噪声测试信号,并且没有配对训练,从源域转换为与转换域中的配对测试信号相似的信号,特别是在频率方面,我量化了这些差异的相关性和误差。摘要:I show that a one-dimensional (1D) conditional generative adversarial network (cGAN) with an adversarial training architecture is capable of unpaired signal-to-signal ("sig2sig") translation. Using a simplified CycleGAN model with 1D layers and wider convolutional kernels, mirroring WaveGAN to reframe two-dimensional (2D) image generation as 1D audio generation, I show that recasting the 2D image-to-image translation task to a 1D signal-to-signal translation task with deep convolutional GANs is possible without substantial modification to the conventional U-Net model and adversarial architecture developed as CycleGAN. With this I show for a small tunable dataset that noisy test signals unseen by the 1D CycleGAN model and without paired training transform from the source domain to signals similar to paired test signals in the translated domain, especially in terms of frequency, and I quantify these differences in terms of correlation and error.
【5】 Spectrogram-Based Detection of Auto-Tuned Vocals in Music Recordings标题:基于谱图的音乐录音中自调谐人声检测链接:https://arxiv.org/abs/2403.05380作者:Mahyar Gohari,Paolo Bestagini,Sergio Benini,Nicola Adami摘要:在音乐制作和音频处理领域,歌唱声音的自动音高校正(也称为自动调音)的实现显著改变了声乐表演的面貌。虽然自动调音技术为音乐家提供了调整音高并达到所需精度的能力,但它的使用也引发了关于其对真实性和艺术完整性影响的辩论。因此,检测和分析音乐录音中的自动调音人声已成为音乐学者,制作人和听众必不可少的。然而,据我们所知,在这方面事先没有作出任何努力。这项研究介绍了一种数据驱动的方法,利用三元组网络检测自动调谐的歌曲,通过创建一个由原始和自动调谐的音频片段组成的数据集来支持。实验结果表明,与Rawnet2相比,该方法在准确性和鲁棒性方面都具有优势,Rawnet2是一种端到端的反欺骗模型,广泛用于其他音频取证任务。摘要:In the domain of music production and audio processing, the implementation of automatic pitch correction of the singing voice, also known as Auto-Tune, has significantly transformed the landscape of vocal performance. While auto-tuning technology has offered musicians the ability to tune their vocal pitches and achieve a desired level of precision, its use has also sparked debates regarding its impact on authenticity and artistic integrity. As a result, detecting and analyzing Auto-Tuned vocals in music recordings has become essential for music scholars, producers, and listeners. However, to the best of our knowledge, no prior effort has been made in this direction. This study introduces a data-driven approach leveraging triplet networks for the detection of Auto-Tuned songs, backed by the creation of a dataset composed of original and Auto-Tuned audio clips. The experimental results demonstrate the superiority of the proposed method in both accuracy and robustness compared to Rawnet2, an end-to-end model proposed for anti-spoofing and widely used for other audio forensic tasks.
【6】 RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction标题:RFWave:用于音频波形重建的多频带整流链接:https://arxiv.org/abs/2403.05010作者:Peng Liu,Dongyang Dai摘要:生成建模的最新进展已经导致从不同表示的音频波形重构的显著进展。尽管扩散模型已经被用于重构音频波形,但是它们倾向于表现出延迟问题,因为它们在各个采样点的级别上操作并且需要相对大量的采样步骤。在这项研究中,我们介绍了RFWave,一种新的多频带整流方法,从梅尔频谱图重建高保真音频波形。RFWave的独特之处在于生成复杂的频谱图,并在帧级运行,同时处理所有子带以提高效率。得益于旨在实现平坦传输轨迹的整流流,RFWave仅需10个采样步骤。经验评估表明,RFWave实现了卓越的重建质量和卓越的计算效率,能够以比实时快90倍的速度生成音频。摘要:Recent advancements in generative modeling have led to significant progress in audio waveform reconstruction from diverse representations. Although diffusion models have been used for reconstructing audio waveforms, they tend to exhibit latency issues because they operate at the level of individual sample points and require a relatively large number of sampling steps. In this study, we introduce RFWave, a novel multi-band Rectified Flow approach that reconstructs high-fidelity audio waveforms from Mel-spectrograms. RFWave is distinctive for generating complex spectrograms and operating at the frame level, processing all subbands concurrently to enhance efficiency. Thanks to Rectified Flow, which aims for a flat transport trajectory, RFWave requires only 10 sampling steps. Empirical evaluations demonstrate that RFWave achieves exceptional reconstruction quality and superior computational efficiency, capable of generating audio at a speed 90 times faster than real-time. 机器翻译由腾讯交互翻译提供,仅供参考