今日论文合集:cs.SD语音6篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Avoiding an AI-imposed Taylor's Version of all music history
标题:避免人工智能强加给泰勒的所有音乐史版本
链接:https://arxiv.org/abs/2402.14589
作者:Nick Collins,Mick Grierson
摘要:由于未来的音乐人工智能与人类音乐密切相关,它们可能会在数据库中对特定的人类艺术家形成自己的依恋,而这些偏见在最坏的情况下可能会导致对所有音乐历史的潜在存在威胁。人工智能的超级粉丝可能会为了自己的喜好而破坏历史记录和现存的唱片,保护世界音乐文化的多样性可能会成为比强加12音平均律或其他西方同质化更紧迫的问题。我们讨论了人工智能封面软件的技术能力,并制作了泰勒版本的西方流行音乐历史上著名的曲目作为挑衅性的例子;这些作品的质量并不影响整体论点(甚至可能看到未来的人工智能试图将曲别针的声音强加给所有现有的音频文件,更不用说泰勒斯威夫特了)。我们讨论了一些潜在的防御措施,对未来的音乐垄断的危险,同时分析了一个最大的“泰勒Swiftication”的完整的音乐记录的可行性。
摘要:As future musical AIs adhere closely to human music, they may form their own attachments to particular human artists in their databases, and these biases may in the worst case lead to potential existential threats to all musical history. AI super fans may act to corrupt the historical record and extant recordings in favour of their own preferences, and preservation of the diversity of world music culture may become even more of a pressing issue than the imposition of 12 tone equal temperament or other Western homogenisations. We discuss the technical capability of AI cover software and produce Taylor's Versions of famous tracks from Western pop history as provocative examples; the quality of these productions does not affect the overall argument (which might even see a future AI try to impose the sound of paperclips onto all existing audio files, let alone Taylor Swift). We discuss some potential defenses against the danger of future musical monopolies, whilst analysing the feasibility of a maximal 'Taylor Swiftication' of the complete musical record.


【2】Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding  Decomposition
标题:DAISY-TTS:通过韵律嵌入分解模拟更广泛的情绪谱
链接:https://arxiv.org/abs/2402.14523
作者:Rendi Chevi,Alham Fikri Aji
备注:Project Page: this https URL
摘要:我们经常以多方面的方式口头表达情绪,它们可能在强度上有所不同,并且可能不仅仅是作为单一的情绪表达,而是作为混合的情绪表达。情绪的结构模型充分研究了这种广泛的情绪,该模型将各种情绪表示为具有不同程度强度的主要情绪的衍生产品。在本文中,我们提出了一个情感的文本到语音的设计,以模拟更广泛的情感的结构模型为基础。我们提出的设计,雏菊TTS,采用了韵律编码器学习情感可分离的韵律嵌入作为情感的代理。这种情感表示允许模型模拟:(1)主要情感,如从训练样本中学习的,(2)次要情感,作为主要情感的混合物,(3)强度水平,通过缩放情感嵌入,以及(4)情感极性,通过否定情感嵌入。通过一系列的感知评估,Daisy-TTS表现出更高的整体情感语音自然度和情感感知能力相比,基线。
摘要:We often verbally express emotions in a multifaceted manner, they may vary in their intensities and may be expressed not just as a single but as a mixture of emotions. This wide spectrum of emotions is well-studied in the structural model of emotions, which represents variety of emotions as derivative products of primary emotions with varying degrees of intensity. In this paper, we propose an emotional text-to-speech design to simulate a wider spectrum of emotions grounded on the structural model. Our proposed design, Daisy-TTS, incorporates a prosody encoder to learn emotionally-separable prosody embedding as a proxy for emotion. This emotion representation allows the model to simulate: (1) Primary emotions, as learned from the training samples, (2) Secondary emotions, as a mixture of primary emotions, (3) Intensity-level, by scaling the emotion embedding, and (4) Emotions polarity, by negating the emotion embedding. Through a series of perceptual evaluations, Daisy-TTS demonstrated overall higher emotional speech naturalness and emotion perceiveability compared to the baseline.

【3】Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion
标题:不可微规则引导扩散的符号音乐生成
链接:https://arxiv.org/abs/2402.14285
作者:Yujia Huang,Adishree Ghatare,Yuanzhe Liu,Ziniu Hu,Qinsheng Zhang,Chandramouli S Sastry,Siddharth Gururani,Sageev Oore,Yisong Yue
摘要:我们研究了符号音乐生成的问题(例如,生成钢琴卷),其技术重点在于不可微的规则指导。音乐规则通常以符号形式表达在音符特征上,例如音符密度或和弦进行,其中许多是不可微的,这在使用它们进行引导扩散时构成了挑战。我们提出了随机控制制导(SCG),这是一种新的制导方法,只需要对规则函数进行前向评估,就可以以即插即用的方式与预先训练的扩散模型一起工作,从而首次实现了不可微规则的免训练制导。此外,我们引入了一个潜在的扩散架构的符号音乐生成具有高的时间分辨率,它可以组成与SCG在一个即插即用的方式。与符号音乐生成的标准强基线相比,该框架在音乐质量和基于规则的可控性方面取得了显着进步,在各种设置中表现优于当前最先进的生成器。有关详细演示,请访问我们的项目网站:https://scg-rule-guided-music.github.io/。
摘要:We study the problem of symbolic music generation (e.g., generating piano rolls), with a technical focus on non-differentiable rule guidance. Musical rules are often expressed in symbolic form on note characteristics, such as note density or chord progression, many of which are non-differentiable which pose a challenge when using them for guided diffusion. We propose Stochastic Control Guidance (SCG), a novel guidance method that only requires forward evaluation of rule functions that can work with pre-trained diffusion models in a plug-and-play way, thus achieving training-free guidance for non-differentiable rules for the first time. Additionally, we introduce a latent diffusion architecture for symbolic music generation with high time resolution, which can be composed with SCG in a plug-and-play fashion. Compared to standard strong baselines in symbolic music generation, this framework demonstrates marked advancements in music quality and rule-based controllability, outperforming current state-of-the-art generators in a variety of settings. For detailed demonstrations, please visit our project site: https://scg-rule-guided-music.github.io/.


【4】Compression Robust Synthetic Speech Detection Using Patched Spectrogram  Transformer
标题:基于拼接频谱变换的压缩鲁棒合成语音检测
链接:https://arxiv.org/abs/2402.14205
作者:Amit Kumar Singh Yadav,Ziyue Xiang,Kratika Bhagtani,Paolo Bestagini,Stefano Tubaro,Edward J. Delp
备注:Accepted as long oral paper at ICMLA 2023
摘要:许多深度学习合成语音生成工具都是现成的。合成语音的使用造成了金融诈骗、冒名顶替和错误信息的传播。出于这个原因,已经提出了可以检测合成语音的取证方法。现有的方法通常在一个数据集上过拟合,并且它们的性能在实际场景中大幅降低,例如检测社交平台上共享的合成语音。在本文中,我们提出,修补频谱图合成语音检测Transformer(PS3DT),一个合成语音检测器,转换的时域语音信号的梅尔频谱图,并处理它在补丁使用Transformer神经网络。我们评估了PS3DT在ASVspoof2019数据集上的检测性能。我们的实验表明,与使用频谱图进行合成语音检测的其他方法相比,PS3DT在ASVspoof2019数据集上表现良好。我们还研究了PS3DT在野外数据集上的泛化性能。PS3DT比现有的几种方法更好地从分布外的数据集中检测合成语音。我们还评估了PS3DT检测电话质量合成语音和社交平台上共享的合成语音(压缩语音)的鲁棒性。PS3DT对压缩具有鲁棒性,并且可以比现有的几种方法更好地检测电话质量合成语音。
摘要:Many deep learning synthetic speech generation tools are readily available. The use of synthetic speech has caused financial fraud, impersonation of people, and misinformation to spread. For this reason forensic methods that can detect synthetic speech have been proposed. Existing methods often overfit on one dataset and their performance reduces substantially in practical scenarios such as detecting synthetic speech shared on social platforms. In this paper we propose, Patched Spectrogram Synthetic Speech Detection Transformer (PS3DT), a synthetic speech detector that converts a time domain speech signal to a mel-spectrogram and processes it in patches using a transformer neural network. We evaluate the detection performance of PS3DT on ASVspoof2019 dataset. Our experiments show that PS3DT performs well on ASVspoof2019 dataset compared to other approaches using spectrogram for synthetic speech detection. We also investigate generalization performance of PS3DT on In-the-Wild dataset. PS3DT generalizes well than several existing methods on detecting synthetic speech from an out-of-distribution dataset. We also evaluate robustness of PS3DT to detect telephone quality synthetic speech and synthetic speech shared on social platforms (compressed speech). PS3DT is robust to compression and can detect telephone quality synthetic speech better than several existing methods.


【5】PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a  Diffusion Probabilistic Model标题:PerioGrad:基于扩散概率模型的基音可控神经声码器
链接:https://arxiv.org/abs/2402.14692
作者:Yukiya Hono,Kei Hashimoto,Yoshihiko Nankaku,Keiichi Tokuda
备注:5 pages, 4 figures, To appear in ICASSP 2024. Audio samples: this https URL
摘要:本文提出了一种基于去噪扩散概率模型(DDPM)的神经声码器,并将显式周期信号作为辅助调节信号。最近,基于DDPM的神经声码器作为可以生成高质量波形的非自回归模型而获得了突出地位。基于DDPM的神经元声码器具有训练简单、时域损失小的优点。在实际应用中,如歌唱声合成,需要神经声码器产生高保真语音波形与灵活的音高控制。然而,传统的基于DDPM的神经声码器难以在这种条件下生成语音波形。我们所提出的模型旨在准确地捕捉语音波形的周期性结构,将明确的周期信号。实验结果表明,我们的模型提高了声音质量,并提供了更好的音调控制比传统的DDPM为基础的神经声码器。
摘要:This paper presents a neural vocoder based on a denoising diffusion probabilistic model (DDPM) incorporating explicit periodic signals as auxiliary conditioning signals. Recently, DDPM-based neural vocoders have gained prominence as non-autoregressive models that can generate high-quality waveforms. The neural vocoders based on DDPM have the advantage of training with a simple time-domain loss. In practical applications, such as singing voice synthesis, there is a demand for neural vocoders to generate high-fidelity speech waveforms with flexible pitch control. However, conventional DDPM-based neural vocoders struggle to generate speech waveforms under such conditions. Our proposed model aims to accurately capture the periodic structure of speech waveforms by incorporating explicit periodic signals. Experimental results show that our model improves sound quality and provides better pitch control than conventional DDPM-based neural vocoders.


【6】SICRN: Advancing Speech Enhancement through State Space Model and  Inplace Convolution Techniques
标题:SICRN:基于状态空间模型和原位卷积技术的语音增强
链接:https://arxiv.org/abs/2402.14225
作者:Changjiang Zhao,Shulin He,Xueliang Zhang
摘要:语音增强旨在提高语音质量和可懂度,特别是在背景噪声使语音信号降级的嘈杂环境中。目前,深度学习方法在语音增强方面取得了巨大成功,例如代表性的卷积递归神经网络(CRN)及其变体。然而,CRN通常采用连续的下采样和上采样卷积来进行频率建模,这破坏了信号在频率上的固有结构。此外,卷积层缺乏时间建模能力。为了解决这些问题,我们提出了一个创新的模块结合状态空间模型和原地卷积(SIC),并取代传统的卷积CRN,称为SICRN。具体来说,双路径多维状态空间模型捕获全局频率依赖性和长期时间依赖性。同时,使用2D-inplace卷积来捕获局部结构,从而放弃了下采样和上采样。对公共INTERSPEECH 2020 DNS挑战数据集的系统评估证明了SICRN的有效性。与强基线相比,SICRN实现了接近最先进的性能,同时在模型参数,计算和算法延迟方面具有优势。所提出的SICRN显示出很大的改善语音增强的希望。
摘要:Speech enhancement aims to improve speech quality and intelligibility, especially in noisy environments where background noise degrades speech signals. Currently, deep learning methods achieve great success in speech enhancement, e.g. the representative convolutional recurrent neural network (CRN) and its variants. However, CRN typically employs consecutive downsampling and upsampling convolution for frequency modeling, which destroys the inherent structure of the signal over frequency. Additionally, convolutional layers lacks of temporal modelling abilities. To address these issues, we propose an innovative module combing a State space model and Inplace Convolution (SIC), and to replace the conventional convolution in CRN, called SICRN. Specifically, a dual-path multidimensional State space model captures the global frequencies dependency and long-term temporal dependencies. Meanwhile, the 2D-inplace convolution is used to capture the local structure, which abandons the downsampling and upsampling. Systematic evaluations on the public INTERSPEECH 2020 DNS challenge dataset demonstrate SICRN's efficacy. Compared to strong baselines, SICRN achieves performance close to state-of-the-art while having advantages in model parameters, computations, and algorithmic delay. The proposed SICRN shows great promise for improved speech enhancement.


eess.AS音频处理
【1】PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a  Diffusion Probabilistic Model标题:PerioGrad:基于扩散概率模型的基音可控神经声码器
链接:https://arxiv.org/abs/2402.14692
作者:Yukiya Hono,Kei Hashimoto,Yoshihiko Nankaku,Keiichi Tokuda
备注:5 pages, 4 figures, To appear in ICASSP 2024. Audio samples: this https URL
摘要:本文提出了一种基于去噪扩散概率模型(DDPM)的神经声码器,并将显式周期信号作为辅助调节信号。最近,基于DDPM的神经声码器作为可以生成高质量波形的非自回归模型而获得了突出地位。基于DDPM的神经元声码器具有训练简单、时域损失小的优点。在实际应用中,如歌唱声合成,需要神经声码器产生高保真语音波形与灵活的音高控制。然而,传统的基于DDPM的神经声码器难以在这种条件下生成语音波形。我们所提出的模型旨在准确地捕捉语音波形的周期性结构,将明确的周期信号。实验结果表明,我们的模型提高了声音质量,并提供了更好的音调控制比传统的DDPM为基础的神经声码器。
摘要:This paper presents a neural vocoder based on a denoising diffusion probabilistic model (DDPM) incorporating explicit periodic signals as auxiliary conditioning signals. Recently, DDPM-based neural vocoders have gained prominence as non-autoregressive models that can generate high-quality waveforms. The neural vocoders based on DDPM have the advantage of training with a simple time-domain loss. In practical applications, such as singing voice synthesis, there is a demand for neural vocoders to generate high-fidelity speech waveforms with flexible pitch control. However, conventional DDPM-based neural vocoders struggle to generate speech waveforms under such conditions. Our proposed model aims to accurately capture the periodic structure of speech waveforms by incorporating explicit periodic signals. Experimental results show that our model improves sound quality and provides better pitch control than conventional DDPM-based neural vocoders.

【2】SICRN: Advancing Speech Enhancement through State Space Model and  Inplace Convolution Techniques
标题:SICRN:基于状态空间模型和原位卷积技术的语音增强
链接:https://arxiv.org/abs/2402.14225
作者:Changjiang Zhao,Shulin He,Xueliang Zhang
摘要:语音增强旨在提高语音质量和可懂度,特别是在背景噪声使语音信号降级的嘈杂环境中。目前,深度学习方法在语音增强方面取得了巨大成功,例如代表性的卷积递归神经网络(CRN)及其变体。然而,CRN通常采用连续的下采样和上采样卷积来进行频率建模,这破坏了信号在频率上的固有结构。此外,卷积层缺乏时间建模能力。为了解决这些问题,我们提出了一个创新的模块结合状态空间模型和原地卷积(SIC),并取代传统的卷积CRN,称为SICRN。具体来说,双路径多维状态空间模型捕获全局频率依赖性和长期时间依赖性。同时,使用2D-inplace卷积来捕获局部结构,从而放弃了下采样和上采样。对公共INTERSPEECH 2020 DNS挑战数据集的系统评估证明了SICRN的有效性。与强基线相比,SICRN实现了接近最先进的性能,同时在模型参数,计算和算法延迟方面具有优势。所提出的SICRN显示出很大的改善语音增强的希望。
摘要:Speech enhancement aims to improve speech quality and intelligibility, especially in noisy environments where background noise degrades speech signals. Currently, deep learning methods achieve great success in speech enhancement, e.g. the representative convolutional recurrent neural network (CRN) and its variants. However, CRN typically employs consecutive downsampling and upsampling convolution for frequency modeling, which destroys the inherent structure of the signal over frequency. Additionally, convolutional layers lacks of temporal modelling abilities. To address these issues, we propose an innovative module combing a State space model and Inplace Convolution (SIC), and to replace the conventional convolution in CRN, called SICRN. Specifically, a dual-path multidimensional State space model captures the global frequencies dependency and long-term temporal dependencies. Meanwhile, the 2D-inplace convolution is used to capture the local structure, which abandons the downsampling and upsampling. Systematic evaluations on the public INTERSPEECH 2020 DNS challenge dataset demonstrate SICRN's efficacy. Compared to strong baselines, SICRN achieves performance close to state-of-the-art while having advantages in model parameters, computations, and algorithmic delay. The proposed SICRN shows great promise for improved speech enhancement.


【3】Avoiding an AI-imposed Taylor's Version of all music history
标题: 避免人工智能强加的泰勒版本的所有音乐历史
链接:https://arxiv.org/abs/2402.14589
作者:Nick Collins,Mick Grierson
摘要:由于未来的音乐人工智能与人类音乐密切相关,它们可能会在数据库中对特定的人类艺术家形成自己的依恋,而这些偏见在最坏的情况下可能会导致对所有音乐历史的潜在存在威胁。人工智能的超级粉丝可能会为了自己的喜好而破坏历史记录和现存的唱片,保护世界音乐文化的多样性可能会成为比强加12音平均律或其他西方同质化更紧迫的问题。我们讨论了人工智能封面软件的技术能力,并制作了泰勒版本的西方流行音乐历史上著名的曲目作为挑衅性的例子;这些作品的质量并不影响整体论点(甚至可能看到未来的人工智能试图将曲别针的声音强加给所有现有的音频文件,更不用说泰勒斯威夫特了)。我们讨论了一些潜在的防御措施,对未来的音乐垄断的危险,同时分析了一个最大的“泰勒Swiftication”的完整的音乐记录的可行性。
摘要:As future musical AIs adhere closely to human music, they may form their own attachments to particular human artists in their databases, and these biases may in the worst case lead to potential existential threats to all musical history. AI super fans may act to corrupt the historical record and extant recordings in favour of their own preferences, and preservation of the diversity of world music culture may become even more of a pressing issue than the imposition of 12 tone equal temperament or other Western homogenisations. We discuss the technical capability of AI cover software and produce Taylor's Versions of famous tracks from Western pop history as provocative examples; the quality of these productions does not affect the overall argument (which might even see a future AI try to impose the sound of paperclips onto all existing audio files, let alone Taylor Swift). We discuss some potential defenses against the danger of future musical monopolies, whilst analysing the feasibility of a maximal 'Taylor Swiftication' of the complete musical record.


【4】Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding  Decomposition
标题:DAISY-TTS:通过韵律嵌入分解模拟更广泛的情绪谱
链接:https://arxiv.org/abs/2402.14523
作者:Rendi Chevi,Alham Fikri Aji
备注:Project Page: this https URL
摘要:我们经常以多方面的方式口头表达情绪,它们可能在强度上有所不同,并且可能不仅仅是作为单一的情绪表达,而是作为混合的情绪表达。情绪的结构模型充分研究了这种广泛的情绪,该模型将各种情绪表示为具有不同程度强度的主要情绪的衍生产品。在本文中,我们提出了一个情感的文本到语音的设计,以模拟更广泛的情感的结构模型为基础。我们提出的设计,雏菊TTS,采用了韵律编码器学习情感可分离的韵律嵌入作为情感的代理。这种情感表示允许模型模拟:(1)主要情感,如从训练样本中学习的,(2)次要情感,作为主要情感的混合物,(3)强度水平,通过缩放情感嵌入,以及(4)情感极性,通过否定情感嵌入。通过一系列的感知评估,Daisy-TTS表现出更高的整体情感语音自然度和情感感知能力相比,基线。
摘要:We often verbally express emotions in a multifaceted manner, they may vary in their intensities and may be expressed not just as a single but as a mixture of emotions. This wide spectrum of emotions is well-studied in the structural model of emotions, which represents variety of emotions as derivative products of primary emotions with varying degrees of intensity. In this paper, we propose an emotional text-to-speech design to simulate a wider spectrum of emotions grounded on the structural model. Our proposed design, Daisy-TTS, incorporates a prosody encoder to learn emotionally-separable prosody embedding as a proxy for emotion. This emotion representation allows the model to simulate: (1) Primary emotions, as learned from the training samples, (2) Secondary emotions, as a mixture of primary emotions, (3) Intensity-level, by scaling the emotion embedding, and (4) Emotions polarity, by negating the emotion embedding. Through a series of perceptual evaluations, Daisy-TTS demonstrated overall higher emotional speech naturalness and emotion perceiveability compared to the baseline.


【5】Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion
标题:不可微规则引导扩散的符号音乐生成
链接:https://arxiv.org/abs/2402.14285
作者:Yujia Huang,Adishree Ghatare,Yuanzhe Liu,Ziniu Hu,Qinsheng Zhang,Chandramouli S Sastry,Siddharth Gururani,Sageev Oore,Yisong Yue
摘要:我们研究了符号音乐生成的问题(例如,生成钢琴卷),其技术重点在于不可微的规则指导。音乐规则通常以符号形式表达在音符特征上,例如音符密度或和弦进行,其中许多是不可微的,这在使用它们进行引导扩散时构成了挑战。我们提出了随机控制制导(SCG),这是一种新的制导方法,只需要对规则函数进行前向评估,就可以以即插即用的方式与预先训练的扩散模型一起工作,从而首次实现了不可微规则的免训练制导。此外,我们引入了一个潜在的扩散架构的符号音乐生成具有高的时间分辨率,它可以组成与SCG在一个即插即用的方式。与符号音乐生成的标准强基线相比,该框架在音乐质量和基于规则的可控性方面取得了显着进步,在各种设置中表现优于当前最先进的生成器。有关详细演示,请访问我们的项目网站:https://scg-rule-guided-music.github.io/。
摘要:We study the problem of symbolic music generation (e.g., generating piano rolls), with a technical focus on non-differentiable rule guidance. Musical rules are often expressed in symbolic form on note characteristics, such as note density or chord progression, many of which are non-differentiable which pose a challenge when using them for guided diffusion. We propose Stochastic Control Guidance (SCG), a novel guidance method that only requires forward evaluation of rule functions that can work with pre-trained diffusion models in a plug-and-play way, thus achieving training-free guidance for non-differentiable rules for the first time. Additionally, we introduce a latent diffusion architecture for symbolic music generation with high time resolution, which can be composed with SCG in a plug-and-play fashion. Compared to standard strong baselines in symbolic music generation, this framework demonstrates marked advancements in music quality and rule-based controllability, outperforming current state-of-the-art generators in a variety of settings. For detailed demonstrations, please visit our project site: https://scg-rule-guided-music.github.io/.

【6】Compression Robust Synthetic Speech Detection Using Patched Spectrogram  Transformer
标题:基于拼接频谱变换的压缩鲁棒合成语音检测
链接:https://arxiv.org/abs/2402.14205
作者:Amit Kumar Singh Yadav,Ziyue Xiang,Kratika Bhagtani,Paolo Bestagini,Stefano Tubaro,Edward J. Delp
备注:Accepted as long oral paper at ICMLA 2023
摘要:许多深度学习合成语音生成工具都是现成的。合成语音的使用造成了金融诈骗、冒名顶替和错误信息的传播。出于这个原因,已经提出了可以检测合成语音的取证方法。现有的方法通常在一个数据集上过拟合,并且它们的性能在实际场景中大幅降低,例如检测社交平台上共享的合成语音。在本文中,我们提出,修补频谱图合成语音检测Transformer(PS3DT),一个合成语音检测器,转换的时域语音信号的梅尔频谱图,并处理它在补丁使用Transformer神经网络。我们评估了PS3DT在ASVspoof2019数据集上的检测性能。我们的实验表明,与使用频谱图进行合成语音检测的其他方法相比,PS3DT在ASVspoof2019数据集上表现良好。我们还研究了PS3DT在野外数据集上的泛化性能。PS3DT比现有的几种方法更好地从分布外的数据集中检测合成语音。我们还评估了PS3DT检测电话质量合成语音和社交平台上共享的合成语音(压缩语音)的鲁棒性。PS3DT对压缩具有鲁棒性,并且可以比现有的几种方法更好地检测电话质量合成语音。
摘要:Many deep learning synthetic speech generation tools are readily available. The use of synthetic speech has caused financial fraud, impersonation of people, and misinformation to spread. For this reason forensic methods that can detect synthetic speech have been proposed. Existing methods often overfit on one dataset and their performance reduces substantially in practical scenarios such as detecting synthetic speech shared on social platforms. In this paper we propose, Patched Spectrogram Synthetic Speech Detection Transformer (PS3DT), a synthetic speech detector that converts a time domain speech signal to a mel-spectrogram and processes it in patches using a transformer neural network. We evaluate the detection performance of PS3DT on ASVspoof2019 dataset. Our experiments show that PS3DT performs well on ASVspoof2019 dataset compared to other approaches using spectrogram for synthetic speech detection. We also investigate generalization performance of PS3DT on In-the-Wild dataset. PS3DT generalizes well than several existing methods on detecting synthetic speech from an out-of-distribution dataset. We also evaluate robustness of PS3DT to detect telephone quality synthetic speech and synthetic speech shared on social platforms (compressed speech). PS3DT is robust to compression and can detect telephone quality synthetic speech better than several existing methods.


机器翻译由腾讯交互翻译提供,仅供参考