今日论文合集:cs.SD语音3篇,eess.AS音频处理4篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音

【1】EveGuard: Defeating Vibration-based Side-Channel Eavesdropping with  Audio Adversarial Perturbations

标题:EveGuard:通过音频对抗性扰动击败基于振动的侧通道发射器丢弃
链接:https://arxiv.org/abs/2411.10034

作者:Jung-Woo Chang,  Ke Sun,  David Xia,  Xinyu Zhang,  Farinaz Koushanfar
摘要:基于振动测量的侧信道会带来重大的隐私风险,利用毫米波雷达、光传感器和加速度计等传感器来检测声源或邻近物体的振动,从而实现语音窃听。尽管提出了各种防御措施,但这些措施涉及具有固有物理限制的昂贵硬件解决方案。本文介绍了EveGuard,这是一个软件驱动的防御框架,可以创建对抗性音频,保护语音隐私免受侧通道的影响,而不会影响人类的感知。我们利用侧通道和传统麦克风的独特传感能力,侧通道捕获振动,麦克风记录气压变化,从而产生不同的频率响应。EveGuard首先提出了一个扰动发生器模型(PGM),有效地抑制基于传感器的窃听,同时保持高音频质量。其次,为了实现PGM的端到端训练,我们引入了一个新的域转换任务,称为Eve-GAN,用于从给定的音频中推断窃听信号。我们进一步应用Few-Shot学习来减轻Eve-GAN训练的数据收集开销。我们广泛的实验表明,EveGuard对音频分类器的保护率超过97%,并显着阻碍了窃听音频重建。我们进一步验证了EveGuard在三种自适应攻击机制中的性能。我们进行了一项用户研究,以验证我们扰动音频的感知质量。
摘要:Vibrometry-based side channels pose a significant privacy risk, exploitingsensors like mmWave radars, light sensors, and accelerometers to detectvibrations from sound sources or proximate objects, enabling speecheavesdropping. Despite various proposed defenses, these involve costly hardwaresolutions with inherent physical limitations. This paper presents EveGuard, asoftware-driven defense framework that creates adversarial audio, protectingvoice privacy from side channels without compromising human perception. Weleverage the distinct sensing capabilities of side channels and traditionalmicrophones where side channels capture vibrations and microphones recordchanges in air pressure, resulting in different frequency responses. EveGuardfirst proposes a perturbation generator model (PGM) that effectively suppressessensor-based eavesdropping while maintaining high audio quality. Second, toenable end-to-end training of PGM, we introduce a new domain translation taskcalled Eve-GAN for inferring an eavesdropped signal from a given audio. Wefurther apply few-shot learning to mitigate the data collection overhead forEve-GAN training. Our extensive experiments show that EveGuard achieves aprotection rate of more than 97 percent from audio classifiers andsignificantly hinders eavesdropped audio reconstruction. We further validatethe performance of EveGuard across three adaptive attack mechanisms. We haveconducted a user study to verify the perceptual quality of our perturbed audio.


【2】 Zero-shot Voice Conversion with Diffusion Transformers
标题:使用扩散变形器实现Zero-Shot语音转换
链接:https://arxiv.org/abs/2411.09943
作者:Songting Liu
摘要:Zero-shot语音转换的目的是将源语音发音转换为与来自不可见说话人的参考语音的音色相匹配。传统的方法与音色泄漏、音色表示不足以及训练和推理任务之间的不匹配作斗争。我们提出了Seed-VC,这是一种新的框架,通过在训练过程中引入外部音色移位器来干扰源语音音色,减轻泄漏并将训练与推理对齐来解决这些问题。此外,我们采用了一个扩散转换器(diffusion Transformer),该转换器利用整个参考语音上下文,通过上下文学习捕获细粒度的音色特征。实验表明,Seed-VC的性能优于OpenVoice和CosyVoice等强基线,在零触发语音转换任务中实现了更高的说话者相似度和更低的单词错误率(zero-shot voice conversion tasks)。我们进一步扩展我们的方法,zero-shot唱歌的声音转换,通过将基频(F0)空调,导致比较性能,目前国家的最先进的方法。我们的研究结果强调了Seed-VC在克服核心挑战方面的有效性,为更准确和多功能的语音转换系统铺平了道路。
摘要:Zero-shot voice conversion aims to transform a source speech utterance tomatch the timbre of a reference speech from an unseen speaker. Traditionalapproaches struggle with timbre leakage, insufficient timbre representation,and mismatches between training and inference tasks. We propose Seed-VC, anovel framework that addresses these issues by introducing an external timbreshifter during training to perturb the source speech timbre, mitigating leakageand aligning training with inference. Additionally, we employ a diffusiontransformer that leverages the entire reference speech context, capturingfine-grained timbre features through in-context learning. Experimentsdemonstrate that Seed-VC outperforms strong baselines like OpenVoice andCosyVoice, achieving higher speaker similarity and lower word error rates inzero-shot voice conversion tasks. We further extend our approach to zero-shotsinging voice conversion by incorporating fundamental frequency (F0)conditioning, resulting in comparative performance to current state-of-the-artmethods. Our findings highlight the effectiveness of Seed-VC in overcoming corechallenges, paving the way for more accurate and versatile voice conversionsystems.

【3】 XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing  Attack Detection
标题:XLSR-Mamba:用于欺骗攻击检测的双列双向状态空间模型
链接:https://arxiv.org/abs/2411.10027
作者:Yang Xiao,  Rohan Kumar Das
备注:5 pages
摘要:Transformers及其变体在语音处理中取得了巨大的成功。然而,他们的多头自我注意机制在计算上是昂贵的。因此,一种新的选择性状态空间模型,曼巴,已被提出作为替代。基于其在自动语音识别中的成功,我们将Mamba应用于欺骗攻击检测。Mamba非常适合这项任务,因为它可以通过处理长序列来捕获欺骗语音信号中的伪像。然而,当使用有限的标记数据进行训练时,Mamba的性能可能会受到影响。为了缓解这一问题,我们建议使用预训练的wav 2 vec 2.0模型,将基于双列架构的Mamba新结构与自监督学习相结合。实验表明,我们提出的方法在ASVspoof 2021 LA和DF数据集上实现了有竞争力的结果和更快的推理,并且在更具挑战性的In-the-Wild数据集上,它成为欺骗攻击检测的最强候选者。该代码将在适当的时候公开发布。
摘要:Transformers and their variants have achieved great success in speechprocessing. However, their multi-head self-attention mechanism iscomputationally expensive. Therefore, one novel selective state space model,Mamba, has been proposed as an alternative. Building on its success inautomatic speech recognition, we apply Mamba for spoofing attack detection.Mamba is well-suited for this task as it can capture the artifacts in spoofedspeech signals by handling long-length sequences. However, Mamba's performancemay suffer when it is trained with limited labeled data. To mitigate this, wepropose combining a new structure of Mamba based on a dual-column architecturewith self-supervised learning, using the pre-trained wav2vec 2.0 model. Theexperiments show that our proposed approach achieves competitive results andfaster inference on the ASVspoof 2021 LA and DF datasets, and on the morechallenging In-the-Wild dataset, it emerges as the strongest candidate forspoofing attack detection. The code will be publicly released in due course.

eess.AS音频处理

【1】 Perceptual implications of simplifying geometrical acoustics models for  Ambisonics-based binaural reverberation
标题:简化基于Ambisonics的双耳回响几何声学模型的感知影响
链接:https://arxiv.org/abs/2411.10375
作者:Vincent Martin,  Isaac Engel,  Lorenzo Picinali
备注:15 pages, 10 figures, 5 tables, will be submitted to IEEE transactions on Audio, Speech and Language processing after revisions
摘要:可以采用不同的方法来渲染虚拟混响,通常需要关于房间的几何形状和表面的声学特性的大量信息。然而,考虑给定环境的所有方面的完全全面的方法可能在计算上是昂贵的并且从感知的角度来看是冗余的。对于这些方法,实现感知真实性和模型的复杂性之间的权衡成为一个相关的挑战。  本研究探讨这种妥协,通过使用几何声学,使高保真度立体声为基础的双耳混响。除其他因素外,其精度取决于其对房间几何形状和材料声学特性的保真度。  本研究的目的是调查的影响,简化房间的几何形状和频率分辨率的吸收系数的感知混响在一个虚拟的声音场景。采用多刺激比较法对基于单个房间的几个抽取模型进行了感知评价。此外,这些差异进行了数值评估,通过混响的声学参数的计算。  根据数值和感知评估,降低吸收系数的频率分辨率可以对混响的感知产生显著影响,而当抽取模型的几何形状时观察到不太显著的影响。
摘要:Different methods can be employed to render virtual reverberation, oftenrequiring substantial information about the room's geometry and the acousticcharacteristics of the surfaces. However, fully comprehensive approaches thataccount for all aspects of a given environment may be computationally costlyand redundant from a perceptual standpoint. For these methods, achieving atrade-off between perceptual authenticity and model's complexity becomes arelevant challenge. This study investigates this compromise through the use of geometricalacoustics to render Ambisonics-based binaural reverberation. Its precision isdetermined, among other factors, by its fidelity to the room's geometry and tothe acoustic properties of its materials. The purpose of this study is to investigate the impact of simplifying theroom geometry and the frequency resolution of absorption coefficients on theperception of reverberation within a virtual sound scene. Several decimatedmodels based on a single room were perceptually evaluated using the amulti-stimulus comparison method. Additionally, these differences werenumerically assessed through the calculation of acoustic parameters of thereverberation. According to numerical and perceptual evaluations, lowering the frequencyresolution of absorption coefficients can have a significant impact on theperception of reverberation, while a less notable impact was observed whendecimating the geometry of the model.

【2】 XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing  Attack Detection
标题:XLSR-Mamba:用于欺骗攻击检测的双列双向状态空间模型
链接:https://arxiv.org/abs/2411.10027
作者:Yang Xiao,  Rohan Kumar Das
备注:5 pages
摘要:Transformers及其变体在语音处理中取得了巨大的成功。然而,他们的多头自我注意机制在计算上是昂贵的。因此,一种新的选择性状态空间模型,曼巴,已被提出作为替代。基于其在自动语音识别中的成功,我们将Mamba应用于欺骗攻击检测。Mamba非常适合这项任务,因为它可以通过处理长序列来捕获欺骗语音信号中的伪像。然而,当使用有限的标记数据进行训练时,Mamba的性能可能会受到影响。为了缓解这一问题,我们建议使用预训练的wav 2 vec 2.0模型,将基于双列架构的Mamba新结构与自监督学习相结合。实验表明,我们提出的方法在ASVspoof 2021 LA和DF数据集上实现了有竞争力的结果和更快的推理,并且在更具挑战性的In-the-Wild数据集上,它成为欺骗攻击检测的最强候选者。该代码将在适当的时候公开发布。
摘要:Transformers and their variants have achieved great success in speechprocessing. However, their multi-head self-attention mechanism iscomputationally expensive. Therefore, one novel selective state space model,Mamba, has been proposed as an alternative. Building on its success inautomatic speech recognition, we apply Mamba for spoofing attack detection.Mamba is well-suited for this task as it can capture the artifacts in spoofedspeech signals by handling long-length sequences. However, Mamba's performancemay suffer when it is trained with limited labeled data. To mitigate this, wepropose combining a new structure of Mamba based on a dual-column architecturewith self-supervised learning, using the pre-trained wav2vec 2.0 model. Theexperiments show that our proposed approach achieves competitive results andfaster inference on the ASVspoof 2021 LA and DF datasets, and on the morechallenging In-the-Wild dataset, it emerges as the strongest candidate forspoofing attack detection. The code will be publicly released in due course.

【3】 EveGuard: Defeating Vibration-based Side-Channel Eavesdropping with  Audio Adversarial Perturbations
标题:EveGuard:通过音频对抗性扰动击败基于振动的侧通道发射器丢弃
链接:https://arxiv.org/abs/2411.10034
作者:Jung-Woo Chang,  Ke Sun,  David Xia,  Xinyu Zhang,  Farinaz Koushanfar
摘要:基于振动测量的侧信道会带来重大的隐私风险,利用毫米波雷达、光传感器和加速度计等传感器来检测声源或邻近物体的振动,从而实现语音窃听。尽管提出了各种防御措施,但这些措施涉及具有固有物理限制的昂贵硬件解决方案。本文介绍了EveGuard,这是一个软件驱动的防御框架,可以创建对抗性音频,保护语音隐私免受侧通道的影响,而不会影响人类的感知。我们利用侧通道和传统麦克风的独特传感能力,侧通道捕获振动,麦克风记录气压变化,从而产生不同的频率响应。EveGuard首先提出了一个扰动发生器模型(PGM),有效地抑制基于传感器的窃听,同时保持高音频质量。其次,为了实现PGM的端到端训练,我们引入了一个新的域转换任务,称为Eve-GAN,用于从给定的音频中推断窃听信号。我们进一步应用Few-Shot学习来减轻Eve-GAN训练的数据收集开销。我们广泛的实验表明,EveGuard对音频分类器的保护率超过97%,并显着阻碍了窃听音频重建。我们进一步验证了EveGuard在三种自适应攻击机制中的性能。我们进行了一项用户研究,以验证我们扰动音频的感知质量。
摘要:Vibrometry-based side channels pose a significant privacy risk, exploitingsensors like mmWave radars, light sensors, and accelerometers to detectvibrations from sound sources or proximate objects, enabling speecheavesdropping. Despite various proposed defenses, these involve costly hardwaresolutions with inherent physical limitations. This paper presents EveGuard, asoftware-driven defense framework that creates adversarial audio, protectingvoice privacy from side channels without compromising human perception. Weleverage the distinct sensing capabilities of side channels and traditionalmicrophones where side channels capture vibrations and microphones recordchanges in air pressure, resulting in different frequency responses. EveGuardfirst proposes a perturbation generator model (PGM) that effectively suppressessensor-based eavesdropping while maintaining high audio quality. Second, toenable end-to-end training of PGM, we introduce a new domain translation taskcalled Eve-GAN for inferring an eavesdropped signal from a given audio. Wefurther apply few-shot learning to mitigate the data collection overhead forEve-GAN training. Our extensive experiments show that EveGuard achieves aprotection rate of more than 97 percent from audio classifiers andsignificantly hinders eavesdropped audio reconstruction. We further validatethe performance of EveGuard across three adaptive attack mechanisms. We haveconducted a user study to verify the perceptual quality of our perturbed audio.

【4】 Zero-shot Voice Conversion with Diffusion Transformers
标题:使用扩散变形器实现Zero-Shot语音转换
链接:https://arxiv.org/abs/2411.09943
作者:Songting Liu
摘要:Zero-shot语音转换的目的是将源语音发音转换为与来自不可见说话人的参考语音的音色相匹配。传统的方法与音色泄漏、音色表示不足以及训练和推理任务之间的不匹配作斗争。我们提出了Seed-VC,这是一种新的框架,通过在训练过程中引入外部音色移位器来干扰源语音音色,减轻泄漏并将训练与推理对齐来解决这些问题。此外,我们采用了扩散Transformer,利用整个参考语音上下文,通过上下文学习捕获细粒度的音色特征。实验表明,Seed-VC的性能优于OpenVoice和CosyVoice等强基线,在零触发语音转换任务中实现了更高的说话者相似度和更低的单词错误率(zero-shot voice conversion tasks)。我们进一步扩展我们的方法,zero-shot唱歌的声音转换,通过将基频(F0)空调,导致比较性能,目前国家的最先进的方法。我们的研究结果强调了Seed-VC在克服核心挑战方面的有效性,为更准确和多功能的语音转换系统铺平了道路。
摘要:Zero-shot voice conversion aims to transform a source speech utterance tomatch the timbre of a reference speech from an unseen speaker. Traditionalapproaches struggle with timbre leakage, insufficient timbre representation,and mismatches between training and inference tasks. We propose Seed-VC, anovel framework that addresses these issues by introducing an external timbreshifter during training to perturb the source speech timbre, mitigating leakageand aligning training with inference. Additionally, we employ a diffusiontransformer that leverages the entire reference speech context, capturingfine-grained timbre features through in-context learning. Experimentsdemonstrate that Seed-VC outperforms strong baselines like OpenVoice andCosyVoice, achieving higher speaker similarity and lower word error rates inzero-shot voice conversion tasks. We further extend our approach to zero-shotsinging voice conversion by incorporating fundamental frequency (F0)conditioning, resulting in comparative performance to current state-of-the-artmethods. Our findings highlight the effectiveness of Seed-VC in overcoming corechallenges, paving the way for more accurate and versatile voice conversionsystems.

机器翻译由腾讯交互翻译提供,仅供参考