今日论文合集:cs.SD语音7篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 MCDubber: Multimodal Context-Aware Expressive Video Dubbing
标题: MCDubber:多模式上下文感知表达性视频配音
作者:Yuan Zhao,Zhenqi Jia,Rui Liu,De Hu,Feilong Bao,Guanglai Gao
链接:点击下载PDF文件
摘要:自动视频配音(AVD)的目的是采取给定的脚本,并生成语音,符合嘴唇运动和韵律的表现力。目前的AVD模型主要利用当前句子的视觉信息来增强合成语音的韵律。然而,关键是要考虑所生成的配音的韵律是否与多模态上下文一致,因为配音将与最终视频中的原始上下文相结合。这一点在以往的研究中被忽视了。为了解决这个问题,我们提出了一个多模态的上下文感知的视频配音模型,称为 textbf{MCDubber},从一个单一的句子到更长的序列与上下文信息的建模对象,以确保全局上下文韵律的一致性。MCDubber包括三个主要部分:(1)上下文持续时间对齐器,旨在学习文本和唇帧之间的上下文感知对齐;(2)上下文韵律预测器,试图读取全局上下文视觉序列并预测上下文感知全局能量和音高;(3)上下文声学解码器最终在相邻的地面实况梅尔频谱图的帮助下预测全局上下文梅尔频谱图。目标句子的声谱图。通过这一过程,MCDubber在配音时充分考虑了多模态语境对当前句子韵律表现力的影响。从输出的上下文梅尔语谱图中提取出的属于目标句子的梅尔语谱图即为最终所需的配音音频。在Chem基准数据集上进行的大量实验表明,与所有高级基线相比,我们的MCDubber显着提高了配音表现力。代码和演示可在https: github.com XiaoYuanJun-zy MCDubber上获得。摘要:Automatic Video Dubbing (AVD) aims to take the given script and generate speech that aligns with lip motion and prosody expressiveness. Current AVD models mainly utilize visual information of the current sentence to enhance the prosody of synthesized speech. However, it is crucial to consider whether the prosody of the generated dubbing aligns with the multimodal context, as the dubbing will be combined with the original context in the final video. This aspect has been overlooked in previous studies. To address this issue, we propose a Multimodal Context-aware video Dubbing model, termed textbf{MCDubber}, to convert the modeling object from a single sentence to a longer sequence with context information to ensure the consistency of the global context prosody. MCDubber comprises three main components: (1) A context duration aligner aims to learn the context-aware alignment between the text and lip frames; (2) A context prosody predictor seeks to read the global context visual sequence and predict the context-aware global energy and pitch; (3) A context acoustic decoder ultimately predicts the global context mel-spectrogram with the assistance of adjacent ground-truth mel-spectrograms of the target sentence. Through this process, MCDubber fully considers the influence of multimodal context on the prosody expressiveness of the current sentence when dubbing. The extracted mel-spectrogram belonging to the target sentence from the output context mel-spectrograms is the final required dubbing audio. Extensive experiments on the Chem benchmark dataset demonstrate that our MCDubber significantly improves dubbing expressiveness compared to all advanced baselines. The code and demos are available at https: github.com XiaoYuanJun-zy MCDubber.

【2】 A Joint Noise Disentanglement and Adversarial Training Framework for Robust Speaker Verification
标题: 用于鲁棒说话人验证的联合噪音解纠缠和对抗训练框架
作者:Xujiang Xing,Mingxing Xu,Thomas Fang Zheng
链接:点击下载PDF文件
摘要:自动说话人确认(ASV)在噪声条件下受到性能下降的影响。为了解决这个问题,我们提出了一种新的对抗性学习框架,该框架结合了噪声分解来建立一个与噪声无关的说话人不变嵌入空间。具体地,解纠缠模块包括两个编码器,分别用于分离说话者相关和不相关信息。重建模块用作正则化项以约束噪声。一个功能强大的损失也被用来监督扬声器编码器学习噪声无关的扬声器嵌入,而不会丢失扬声器信息。此外,引入对抗性训练来阻止说话者编码器编码声学条件信息以实现说话者不变的嵌入空间。在VoxCeleb1上的实验表明,该方法提高了说话人确认系统在无噪声和有噪声条件下的性能。摘要:Automatic Speaker Verification (ASV) suffers from performance degradation in noisy conditions. To address this issue, we propose a novel adversarial learning framework that incorporates noise-disentanglement to establish a noise-independent speaker invariant embedding space. Specifically, the disentanglement module includes two encoders for separating speaker related and irrelevant information, respectively. The reconstruction module serves as a regularization term to constrain the noise. A feature-robust loss is also used to supervise the speaker encoder to learn noise-independent speaker embeddings without losing speaker information. In addition, adversarial training is introduced to discourage the speaker encoder from encoding acoustic condition information for achieving a speaker-invariant embedding space. Experiments on VoxCeleb1 indicate that the proposed method improves the performance of the speaker verification system under both clean and noisy conditions.

【3】 Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
标题: 提高耳语和常规语音的Zero-Shot任意语音转换的说话者相似性
作者:Anastasia Avdeeva,Aleksei Gusev
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:Zero-shot语音转换的目的是将源说话人的语音转换为训练过程中未看到的说话人的语音,同时保留内容信息。虽然已经提出了各种方法来重建生成的语音中的说话者信息,但在实现生成的记录和地面实况记录之间的高相似性方面仍有改进的空间。此外,用于特定领域中的语音的zero-shot语音转换,例如耳语,仍然是一个未开发的领域。为了解决这个问题,我们提出了一个SpeakerVC模型,可以有效地执行zero-shot语音转换在有声和低声域,同时是轻量级的,能够在流模式下运行,没有显着的质量下降。此外,我们探讨的方法,以提高说话人身份转移的质量,并证明其有效性的各种语音转换系统。摘要:Zero-shot voice conversion aims to transfer the voice of a source speaker to that of a speaker unseen during training, while preserving the content information. Although various methods have been proposed to reconstruct speaker information in generated speech, there is still room for improvement in achieving high similarity between generated and ground truth recordings. Furthermore, zero-shot voice conversion for speech in specific domains, such as whispered, remains an unexplored area. To address this problem, we propose a SpeakerVC model that can effectively perform zero-shot speech conversion in both voiced and whispered domains, while being lightweight and capable of running in streaming mode without significant quality degradation. In addition, we explore methods to improve the quality of speaker identity transfer and demonstrate their effectiveness for a variety of voice conversion systems.

【4】 DDSP Guitar Amp: Interpretable Guitar Amplifier Modeling
标题: DDSP吉他音箱:可解释吉他放大器建模
作者:Yen-Tung Yeh,Yu-Hua Chen,Yuan-Chiao Cheng,Jui-Te Wu,Jun-Jie Fu,Yi-Fan Yeh,Yi-Hsuan Yang
备注:Preprint paper
链接:点击下载PDF文件
摘要:用于吉他放大器仿真的神经网络模型虽然有效,但通常需要高计算成本并且缺乏可解释性。从物理放大器设计中汲取灵感,本文旨在通过一种新的基于可微分数字信号处理(DDSP)的模型(称为“DDSP吉他放大器”)来解决这些问题,该模型对吉他放大器的四个组件(即,前置放大器、音调堆栈、功率放大器和输出Transformer)。通过一组时域和频域指标,我们证明了DDSP吉他放大器实现了与黑盒基线相当的性能,同时每个音频样本需要不到10%的计算操作,从而在实时应用中具有更大的潜力。摘要:Neural network models for guitar amplifier emulation, while being effective, often demand high computational cost and lack interpretability. Drawing ideas from physical amplifier design, this paper aims to address these issues with a new differentiable digital signal processing (DDSP)-based model, called DDSP guitar amp,'' that models the four components of a guitar amp (i.e., preamp, tone stack, power amp, and output transformer) using specific DSP-inspired designs. With a set of time- and frequency-domain metrics, we demonstrate that DDSP guitar amp achieves performance comparable with that of black-box baselines while requiring less than 10 % of the computational operations per audio sample, thereby holding greater potential for usages in real-time applications.

【5】 BUT Systems and Analyses for the ASVspoof 5 Challenge
标题: 但ASVspoof 5挑战赛的系统和分析
作者:Johan Rohdin,Lin Zhang,Oldřich Plchot,Vojtěch Staněk,David Mihola,Junyi Peng,Themos Stafylakis,Dmitriy Beveraki,Anna Silnova,Jan Brukner,Lukáš Burget
备注:8 pages, ASVspoof 5 Workshop (Interspeech2024 Satellite)
链接:点击下载PDF文件
摘要:本文介绍了BUT提交的ASVspoof 5挑战系统,以及分析。对于传统的deepfake检测任务,我们分别在封闭和开放条件下使用ResNet18和自监督模型。此外,我们分析和可视化的不同组合的说话人信息和欺骗信息的标签方案的训练。对于欺骗鲁棒的自动说话人验证(SASV),我们引入有效的先验知识,并提出使用逻辑回归联合训练仿射变换的对策分数和自动说话人验证分数的方式,SASV LLR优化。摘要:This paper describes the BUT submitted systems for the ASVspoof 5 challenge, along with analyses. For the conventional deepfake detection task, we use ResNet18 and self-supervised models for the closed and open conditions, respectively. In addition, we analyze and visualize different combinations of speaker information and spoofing information as label schemes for training. For spoofing-robust automatic speaker verification (SASV), we introduce effective priors and propose using logistic regression to jointly train affine transformations of the countermeasure scores and the automatic speaker verification scores in such a way that the SASV LLR is optimized.

【6】 Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
标题: 估计的音频字幕对应性改进基于格式的音频检索
作者:Paul Primus,Florian Schmid,Gerhard Widmer
备注:In Proceedings of the 9th Workshop on Detection and Classification of Acoustic Scenes and Events, DCASE, Tokyo, Japan, 2024. Implementation available on GitHub: this https URL
链接:点击下载PDF文件
摘要:基于双编码器的音频检索系统通常通过对一组匹配和不匹配的音频字幕对进行对比学习来优化。这导致共享的嵌入空间,其中来自两个模态的对应项最终靠近在一起。由于音频-字幕数据集通常仅包含记录和描述的匹配对,因此通过将音频与从数据集中随机抽取的字幕配对来创建不匹配对已经成为常见的做法。这是不理想的,因为随机采样的字幕可能只是偶然地部分或全部描述音频记录。然而,所有可能的对的对应信息是昂贵的注释,因此通常不可用;因此,我们建议用估计的对应来代替它。为此,我们提出了一个两阶段的训练过程,其中多个检索模型首先像往常一样训练,即,没有估计的对应关系。在第二阶段中,由这些模型预测的音频-字幕对应关系然后用作预测目标。我们在ClothoV 2和AudioCaps基准上评估了我们的方法,并表明它提高了检索性能,即使在限制性的自蒸馏设置中,单个模型生成并从估计的对应关系中学习。我们进一步表明,我们的方法优于目前的最先进的1.6 pp。在ClothoV 2基准上的mAP@10。摘要:Dual-encoder-based audio retrieval systems are commonly optimized with contrastive learning on a set of matching and mismatching audio-caption pairs. This leads to a shared embedding space in which corresponding items from the two modalities end up close together. Since audio-caption datasets typically only contain matching pairs of recordings and descriptions, it has become common practice to create mismatching pairs by pairing the audio with a caption randomly drawn from the dataset. This is not ideal because the randomly sampled caption could, just by chance, partly or entirely describe the audio recording. However, correspondence information for all possible pairs is costly to annotate and thus typically unavailable; we, therefore, suggest substituting it with estimated correspondences. To this end, we propose a two-staged training procedure in which multiple retrieval models are first trained as usual, i.e., without estimated correspondences. In the second stage, the audio-caption correspondences predicted by these models then serve as prediction targets. We evaluate our method on the ClothoV2 and the AudioCaps benchmark and show that it improves retrieval performance, even in a restricting self-distillation setting where a single model generates and then learns from the estimated correspondences. We further show that our method outperforms the current state of the art by 1.6 pp. mAP@10 on the ClothoV2 benchmark.

【7】 Near-Field Signal Processing: Unleashing the Power of Proximity
标题: 近场信号处理:释放接近性的力量
作者:Ahmet M. Elbir,Özlem Tuğfe Demir,Kumar Vijay Mishra,Symeon Chatzinotas,Martin Haardt
备注:12pages7figures, submitted to IEEE
链接:点击下载PDF文件
摘要:近一个世纪以来,近场电磁波在光学、遥感和声学等领域的应用日益受到人们的关注。这种新的关注是由在各种领域,如无线通信,全息术,医学成像和量子启发系统的应用前景的出现推动的。NF传感和无线通信环境中的信号处理需要解决与扩展散射体,范围相关的波束图案,球面波阵面,互耦合效应以及反应场和辐射场的存在相关的问题。最近的调查集中在这些方面的背景下,非常大的阵列和宽的带宽,在信道估计,波束形成,波束训练,传感和定位带来了新的挑战。虽然NF光学具有悠久的历史,但NF相位恢复技术及其应用的进步最近引起了重大的研究关注。类似地,利用NF定位与声学阵列表示NF声学阵列信号处理中的已建立原理的当代扩展。本文旨在概述NF域中最先进的信号处理技术,并对各种应用的最新进展提供全面的视角。摘要:After nearly a century of specialized applications in optics, remote sensing, and acoustics, the near-field (NF) electromagnetic propagation zone is experiencing a resurgence in research interest. This renewed attention is fueled by the emergence of promising applications in various fields such as wireless communications, holography, medical imaging, and quantum-inspired systems. Signal processing within NF sensing and wireless communications environments entails addressing issues related to extended scatterers, range-dependent beampatterns, spherical wavefronts, mutual coupling effects, and the presence of both reactive and radiative fields. Recent investigations have focused on these aspects in the context of extremely large arrays and wide bandwidths, giving rise to novel challenges in channel estimation, beamforming, beam training, sensing, and localization. While NF optics has a longstanding history, advancements in NF phase retrieval techniques and their applications have lately garnered significant research attention. Similarly, utilizing NF localization with acoustic arrays represents a contemporary extension of established principles in NF acoustic array signal processing. This article aims to provide an overview of state-of-the-art signal processing techniques within the NF domain, offering a comprehensive perspective on recent advances in diverse applications.


eess.AS音频处理
【1】 Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
标题: 估计的音频字幕对应性改进基于格式的音频检索
作者:Paul Primus,Florian Schmid,Gerhard Widmer
备注:In Proceedings of the 9th Workshop on Detection and Classification of Acoustic Scenes and Events, DCASE, Tokyo, Japan, 2024. Implementation available on GitHub: this https URL
链接:点击下载PDF文件
摘要:基于双编码器的音频检索系统通常通过对一组匹配和不匹配的音频字幕对进行对比学习来优化。这导致共享的嵌入空间,其中来自两个模态的对应项最终靠近在一起。由于音频-字幕数据集通常仅包含记录和描述的匹配对,因此通过将音频与从数据集中随机抽取的字幕配对来创建不匹配对已成为常见的做法。这是不理想的,因为随机采样的字幕可能只是偶然地部分或全部描述音频记录。然而,所有可能的对的对应信息是昂贵的注释,因此通常不可用;因此,我们建议用估计的对应来代替它。为此,我们提出了一个两阶段的训练过程,其中多个检索模型首先像往常一样训练,即,没有估计的对应关系。在第二阶段中,由这些模型预测的音频-字幕对应关系然后用作预测目标。我们在ClothoV 2和AudioCaps基准上评估了我们的方法,并表明它提高了检索性能,即使在限制性的自蒸馏设置中,单个模型生成并从估计的对应关系中学习。我们进一步表明,我们的方法优于目前的最先进的1.6 pp。在ClothoV 2基准上的mAP@10。摘要:Dual-encoder-based audio retrieval systems are commonly optimized with contrastive learning on a set of matching and mismatching audio-caption pairs. This leads to a shared embedding space in which corresponding items from the two modalities end up close together. Since audio-caption datasets typically only contain matching pairs of recordings and descriptions, it has become common practice to create mismatching pairs by pairing the audio with a caption randomly drawn from the dataset. This is not ideal because the randomly sampled caption could, just by chance, partly or entirely describe the audio recording. However, correspondence information for all possible pairs is costly to annotate and thus typically unavailable; we, therefore, suggest substituting it with estimated correspondences. To this end, we propose a two-staged training procedure in which multiple retrieval models are first trained as usual, i.e., without estimated correspondences. In the second stage, the audio-caption correspondences predicted by these models then serve as prediction targets. We evaluate our method on the ClothoV2 and the AudioCaps benchmark and show that it improves retrieval performance, even in a restricting self-distillation setting where a single model generates and then learns from the estimated correspondences. We further show that our method outperforms the current state of the art by 1.6 pp. mAP@10 on the ClothoV2 benchmark.

【2】 Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining
标题: 通过对比学习和音频预训练改进人声模仿查询
作者:Jonathan Greif,Florian Schmid,Paul Primus,Gerhard Widmer
备注:Accepted to the DCASE Workshop 2024. Source code available: this https URL
链接:点击下载PDF文件
摘要:Query-by-Vocal Imitation(QBV)是关于使用由用户的声音创建的声音模仿来搜索数据库中的音频文件。由于大多数人可以通过语音有效地传达声音概念,因此与基于文本的搜索相比,QBV提供了更直观和方便的方法。为了充分利用QBV,为人声模仿和原始声音开发鲁棒的音频特征表示至关重要。在本文中,我们提出了一种新的QBV系统,该系统利用了用大规模通用音频数据集预训练的卷积神经网络的特征提取能力。我们将这些预先训练的模型集成到双编码器架构中,并使用对比学习对其进行端到端的微调。我们提出的方法的一个独特之处是使用自适应的NT-Xent损失进行对比学习,为参考录音和声乐模仿创建共享的嵌入空间,对预训练模型进行微调。该系统显着提高了音频检索性能,建立了一个新的国家的艺术在粗粒度和细粒度QBV任务。摘要:Query-by-Vocal Imitation (QBV) is about searching audio files within databases using vocal imitations created by the user's voice. Since most humans can effectively communicate sound concepts through voice, QBV offers the more intuitive and convenient approach compared to text-based search. To fully leverage QBV, developing robust audio feature representations for both the vocal imitation and the original sound is crucial. In this paper, we present a new system for QBV that utilizes the feature extraction capabilities of Convolutional Neural Networks pre-trained with large-scale general-purpose audio datasets. We integrate these pre-trained models into a dual encoder architecture and fine-tune them end-to-end using contrastive learning. A distinctive aspect of our proposed method is the fine-tuning strategy of pre-trained models using an adapted NT-Xent loss for contrastive learning, creating a shared embedding space for reference recordings and vocal imitations. The proposed system significantly enhances audio retrieval performance, establishing a new state of the art on both coarse- and fine-grained QBV tasks.

【3】 Near-Field Signal Processing: Unleashing the Power of Proximity
标题: 近场信号处理:释放接近性的力量
作者:Ahmet M. Elbir,Özlem Tuğfe Demir,Kumar Vijay Mishra,Symeon Chatzinotas,Martin Haardt
备注:12pages7figures, submitted to IEEE
链接:点击下载PDF文件
摘要:近一个世纪以来,近场电磁波在光学、遥感和声学等领域的应用日益受到人们的关注。这种新的关注是由在各种领域,如无线通信,全息术,医学成像和量子启发系统的应用前景的出现推动的。NF传感和无线通信环境中的信号处理需要解决与扩展散射体,范围相关的波束图案,球面波阵面,互耦合效应以及反应场和辐射场的存在相关的问题。最近的调查集中在这些方面的背景下,非常大的阵列和宽的带宽,在信道估计,波束形成,波束训练,传感和定位带来了新的挑战。虽然NF光学具有悠久的历史,但NF相位恢复技术及其应用的进步最近引起了重大的研究关注。类似地,利用NF定位与声学阵列表示NF声学阵列信号处理中的已建立原理的当代扩展。本文旨在概述NF域中最先进的信号处理技术,全面介绍各种应用的最新进展。摘要:After nearly a century of specialized applications in optics, remote sensing, and acoustics, the near-field (NF) electromagnetic propagation zone is experiencing a resurgence in research interest. This renewed attention is fueled by the emergence of promising applications in various fields such as wireless communications, holography, medical imaging, and quantum-inspired systems. Signal processing within NF sensing and wireless communications environments entails addressing issues related to extended scatterers, range-dependent beampatterns, spherical wavefronts, mutual coupling effects, and the presence of both reactive and radiative fields. Recent investigations have focused on these aspects in the context of extremely large arrays and wide bandwidths, giving rise to novel challenges in channel estimation, beamforming, beam training, sensing, and localization. While NF optics has a longstanding history, advancements in NF phase retrieval techniques and their applications have lately garnered significant research attention. Similarly, utilizing NF localization with acoustic arrays represents a contemporary extension of established principles in NF acoustic array signal processing. This article aims to provide an overview of state-of-the-art signal processing techniques within the NF domain, offering a comprehensive perspective on recent advances in diverse applications.

【4】 MCDubber: Multimodal Context-Aware Expressive Video Dubbing
标题: MCDubber:多模式上下文感知表达性视频配音
作者:Yuan Zhao,Zhenqi Jia,Rui Liu,De Hu,Feilong Bao,Guanglai Gao
链接:点击下载PDF文件
摘要:自动视频配音(AVD)的目的是采取给定的脚本,并生成语音,符合嘴唇运动和韵律的表现力。目前的AVD模型主要利用当前句子的视觉信息来增强合成语音的韵律。然而,关键是要考虑所生成的配音的韵律是否与多模态上下文一致,因为配音将与最终视频中的原始上下文相结合。这一点在以往的研究中被忽视了。为了解决这个问题,我们提出了一个多模态的上下文感知的视频配音模型,称为 textbf{MCDubber},从一个单一的句子到更长的序列与上下文信息的建模对象,以确保全局上下文韵律的一致性。MCDubber包括三个主要部分:(1)上下文持续时间对齐器,旨在学习文本和唇帧之间的上下文感知对齐;(2)上下文韵律预测器,试图读取全局上下文视觉序列并预测上下文感知全局能量和音高;(3)上下文声学解码器最终在相邻的地面实况梅尔频谱图的帮助下预测全局上下文梅尔频谱图。目标句子的声谱图。通过这一过程,MCDubber在配音时充分考虑了多模态语境对当前句子韵律表现力的影响。从输出的上下文梅尔语谱图中提取出的属于目标句子的梅尔语谱图即为最终所需的配音音频。在Chem基准数据集上进行的大量实验表明,与所有高级基线相比,我们的MCDubber显着提高了配音表现力。代码和演示可在https: github.com XiaoYuanJun-zy MCDubber上获得。摘要:Automatic Video Dubbing (AVD) aims to take the given script and generate speech that aligns with lip motion and prosody expressiveness. Current AVD models mainly utilize visual information of the current sentence to enhance the prosody of synthesized speech. However, it is crucial to consider whether the prosody of the generated dubbing aligns with the multimodal context, as the dubbing will be combined with the original context in the final video. This aspect has been overlooked in previous studies. To address this issue, we propose a Multimodal Context-aware video Dubbing model, termed textbf{MCDubber}, to convert the modeling object from a single sentence to a longer sequence with context information to ensure the consistency of the global context prosody. MCDubber comprises three main components: (1) A context duration aligner aims to learn the context-aware alignment between the text and lip frames; (2) A context prosody predictor seeks to read the global context visual sequence and predict the context-aware global energy and pitch; (3) A context acoustic decoder ultimately predicts the global context mel-spectrogram with the assistance of adjacent ground-truth mel-spectrograms of the target sentence. Through this process, MCDubber fully considers the influence of multimodal context on the prosody expressiveness of the current sentence when dubbing. The extracted mel-spectrogram belonging to the target sentence from the output context mel-spectrograms is the final required dubbing audio. Extensive experiments on the Chem benchmark dataset demonstrate that our MCDubber significantly improves dubbing expressiveness compared to all advanced baselines. The code and demos are available at https: github.com XiaoYuanJun-zy MCDubber.

【5】 A Joint Noise Disentanglement and Adversarial Training Framework for Robust Speaker Verification
标题: 用于鲁棒说话人验证的联合噪音解纠缠和对抗训练框架
作者:Xujiang Xing,Mingxing Xu,Thomas Fang Zheng
链接:点击下载PDF文件
摘要:自动说话人确认(ASV)在噪声条件下受到性能下降的影响。为了解决这个问题,我们提出了一种新的对抗性学习框架,该框架结合了噪声分解来建立一个与噪声无关的说话人不变嵌入空间。具体地,解纠缠模块包括两个编码器,分别用于分离说话者相关和不相关信息。重建模块用作正则化项以约束噪声。一个功能强大的损失也被用来监督扬声器编码器学习噪声无关的扬声器嵌入,而不会丢失扬声器信息。此外,引入对抗性训练来阻止说话者编码器编码声学条件信息以实现说话者不变的嵌入空间。在VoxCeleb1上的实验表明,该方法提高了说话人确认系统在无噪声和有噪声条件下的性能。摘要:Automatic Speaker Verification (ASV) suffers from performance degradation in noisy conditions. To address this issue, we propose a novel adversarial learning framework that incorporates noise-disentanglement to establish a noise-independent speaker invariant embedding space. Specifically, the disentanglement module includes two encoders for separating speaker related and irrelevant information, respectively. The reconstruction module serves as a regularization term to constrain the noise. A feature-robust loss is also used to supervise the speaker encoder to learn noise-independent speaker embeddings without losing speaker information. In addition, adversarial training is introduced to discourage the speaker encoder from encoding acoustic condition information for achieving a speaker-invariant embedding space. Experiments on VoxCeleb1 indicate that the proposed method improves the performance of the speaker verification system under both clean and noisy conditions.

【6】 Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
标题: 提高耳语和常规语音的Zero-Shot任意语音转换的说话者相似性
作者:Anastasia Avdeeva,Aleksei Gusev
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:Zero-shot语音转换的目的是将源说话人的语音转换为训练过程中未看到的说话人的语音,同时保留内容信息。虽然已经提出了各种方法来重建生成的语音中的说话者信息,但在实现生成的记录和地面实况记录之间的高相似性方面仍有改进的空间。此外,用于特定领域中的语音的zero-shot语音转换,例如耳语,仍然是一个未开发的领域。为了解决这个问题,我们提出了一个SpeakerVC模型,可以有效地执行zero-shot语音转换在有声和低声域,同时是轻量级的,能够在流模式下运行,没有显着的质量下降。此外,我们探讨的方法,以提高说话人身份转移的质量,并证明其有效性的各种语音转换系统。摘要:Zero-shot voice conversion aims to transfer the voice of a source speaker to that of a speaker unseen during training, while preserving the content information. Although various methods have been proposed to reconstruct speaker information in generated speech, there is still room for improvement in achieving high similarity between generated and ground truth recordings. Furthermore, zero-shot voice conversion for speech in specific domains, such as whispered, remains an unexplored area. To address this problem, we propose a SpeakerVC model that can effectively perform zero-shot speech conversion in both voiced and whispered domains, while being lightweight and capable of running in streaming mode without significant quality degradation. In addition, we explore methods to improve the quality of speaker identity transfer and demonstrate their effectiveness for a variety of voice conversion systems.

【7】 DDSP Guitar Amp: Interpretable Guitar Amplifier Modeling
标题: DDSP吉他音箱:可解释吉他放大器建模
作者:Yen-Tung Yeh,Yu-Hua Chen,Yuan-Chiao Cheng,Jui-Te Wu,Jun-Jie Fu,Yi-Fan Yeh,Yi-Hsuan Yang
备注:Preprint paper
链接:点击下载PDF文件
摘要:用于吉他放大器仿真的神经网络模型虽然有效,但通常需要高计算成本并且缺乏可解释性。从物理放大器设计中汲取灵感,本文旨在通过一种新的基于可微分数字信号处理(DDSP)的模型(称为“DDSP吉他放大器”)来解决这些问题,该模型对吉他放大器的四个组件(即,前置放大器、音调堆栈、功率放大器和输出Transformer)。通过一组时域和频域指标,我们证明了DDSP吉他放大器实现了与黑盒基线相当的性能,同时每个音频样本需要不到10%的计算操作,从而在实时应用中具有更大的潜力。摘要:Neural network models for guitar amplifier emulation, while being effective, often demand high computational cost and lack interpretability. Drawing ideas from physical amplifier design, this paper aims to address these issues with a new differentiable digital signal processing (DDSP)-based model, called DDSP guitar amp,'' that models the four components of a guitar amp (i.e., preamp, tone stack, power amp, and output transformer) using specific DSP-inspired designs. With a set of time- and frequency-domain metrics, we demonstrate that DDSP guitar amp achieves performance comparable with that of black-box baselines while requiring less than 10 % of the computational operations per audio sample, thereby holding greater potential for usages in real-time applications.

【8】 BUT Systems and Analyses for the ASVspoof 5 Challenge
标题: 但ASVspoof 5挑战赛的系统和分析
作者:Johan Rohdin,Lin Zhang,Oldřich Plchot,Vojtěch Staněk,David Mihola,Junyi Peng,Themos Stafylakis,Dmitriy Beveraki,Anna Silnova,Jan Brukner,Lukáš Burget
备注:8 pages, ASVspoof 5 Workshop (Interspeech2024 Satellite)
链接:点击下载PDF文件
摘要:本文介绍了BUT提交的ASVspoof 5挑战系统,以及分析。对于传统的deepfake检测任务,我们分别在封闭和开放条件下使用ResNet18和自监督模型。此外,我们分析和可视化的不同组合的说话人信息和欺骗信息的标签方案的训练。对于欺骗鲁棒的自动说话人验证(SASV),我们引入有效的先验知识,并提出使用逻辑回归联合训练仿射变换的对策分数和自动说话人验证分数的方式,SASV LLR优化。摘要:This paper describes the BUT submitted systems for the ASVspoof 5 challenge, along with analyses. For the conventional deepfake detection task, we use ResNet18 and self-supervised models for the closed and open conditions, respectively. In addition, we analyze and visualize different combinations of speaker information and spoofing information as label schemes for training. For spoofing-robust automatic speaker verification (SASV), we introduce effective priors and propose using logistic regression to jointly train affine transformations of the countermeasure scores and the automatic speaker verification scores in such a way that the SASV LLR is optimized.


机器翻译,仅供参考