今日论文合集:cs.SD语音4篇,eess.AS音频处理4篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音

【1】 AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal  Audio-Video Generation

标题:AV-Link:用于跨模式音频视频生成的时间对齐扩散功能
链接:https://arxiv.org/abs/2412.15191
作者:Moayed Haji-Ali,  Willi Menapace,  Aliaksandr Siarohin,  Ivan Skorokhodov,  Alper Canberk,  Kwot Sin Lee,  Vicente Ordonez,  Sergey Tulyakov
备注:Project Page: snap-research.github.io/AVLink/
摘要:我们提出了AV-Link,一个统一的视频到音频和音频到视频生成框架,利用冻结的视频和音频扩散模型的激活时间对齐的跨模态条件反射。我们的框架的关键是一个融合块,通过时间对齐的自我注意操作,使我们的骨干视频和音频扩散模型之间的双向信息交换。与使用针对条件信号的其他任务预先训练的特征提取器的先前工作不同,AV-Link可以在单个框架中直接利用由互补模态获得的特征,即视频特征来生成音频,或音频特征来生成视频。我们广泛评估我们的设计选择,并展示我们的方法,以实现同步和高质量的视听内容的能力,展示其在沉浸式媒体生成的应用潜力。项目页面:snap-research.github.io/AVLink/
摘要:We propose AV-Link, a unified framework for Video-to-Audio and Audio-to-Videogeneration that leverages the activations of frozen video and audio diffusionmodels for temporally-aligned cross-modal conditioning. The key to ourframework is a Fusion Block that enables bidirectional information exchangebetween our backbone video and audio diffusion models through atemporally-aligned self attention operation. Unlike prior work that usesfeature extractors pretrained for other tasks for the conditioning signal,AV-Link can directly leverage features obtained by the complementary modalityin a single framework i.e. video features to generate audio, or audio featuresto generate video. We extensively evaluate our design choices and demonstratethe ability of our method to achieve synchronized and high-quality audiovisualcontent, showcasing its potential for applications in immersive mediageneration. Project Page: snap-research.github.io/AVLink/

【2】 GIRAFE: Glottal Imaging Dataset for Advanced Segmentation, Analysis, and  Facilitative Playbacks Evaluation
标题:GIRAFE:用于高级分割、分析和辅助回放评估的喉舌成像数据集
链接:https://arxiv.org/abs/2412.15054
作者:G. Andrade-Miranda,  K. Chatzipapas,  J.D. Arias-Londoño,  J. I. Godino-Llorente
备注:18 pages, 8 figures
摘要:从声带的高速视频内窥镜序列中提取的促进性回放的发展的进展受到阻碍,这是由于明显缺乏用对应于声门间隙区域的语义分割注释的公开可用的数据集。这一事实也限制了该领域现有研究的可重复性和进一步探索。  为了解决这一差距,GIRAFE是一个数据存储库,旨在促进先进技术的发展,用于声带的高速视频内窥镜序列的语义分割,分析和快速评估。存储库包括来自50名患者(30名女性,20名男性)的65个高速视频内窥镜记录。该数据集包括来自健康对照组的15个记录,来自诊断为语音障碍的患者的26个记录,以及24个未知健康状况的记录。所有这些都由专家手动注释,包括与声门间隙的语义分割相对应的掩模。该存储库还补充了使用不同的国家的最先进的方法自动分割声门区。  该数据集已经支持了多项研究,这证明了它对于根据高速视频内窥镜序列开发新的声门间隙分割算法以改进或创建新的促进性回放的有用性。尽管在该领域取得了这些进展和其他进展,但执行声门区的准确且完全自动的语义分割方法的更广泛挑战仍然是开放的。
摘要:The advances in the development of Facilitative Playbacks extracted fromHigh-Speed videoendoscopic sequences of the vocal folds are hindered by anotable lack of publicly available datasets annotated with the semanticsegmentations corresponding to the area of the glottal gap. This fact alsolimits the reproducibility and further exploration of existing research in thisfield. To address this gap, GIRAFE is a data repository designed to facilitate thedevelopment of advanced techniques for the semantic segmentation, analysis, andfast evaluation of High-Speed videoendoscopic sequences of the vocal folds. Therepository includes 65 high-speed videoendoscopic recordings from a cohort of50 patients (30 female, 20 male). The dataset comprises 15 recordings fromhealthy controls, 26 from patients with diagnosed voice disorders, and 24 withan unknown health condition. All of them were manually annotated by an expert,including the masks corresponding to the semantic segmentation of the glottalgap. The repository is also complemented with the automatic segmentation of theglottal area using different state-of-the-art approaches. This data set has already supported several studies, which demonstrates itsusefulness for the development of new glottal gap segmentation algorithms fromHigh-Speed-Videoendoscopic sequences to improve or create new FacilitativePlaybacks. Despite these advances and others in the field, the broaderchallenge of performing an accurate and completely automatic semanticsegmentation method of the glottal area remains open.

【3】 Stable-V2A: Synthesis of Synchronized Sound Effects with Temporal and  Semantic Controls
标题:Stable-V2 A:具有时间和语义控制的同步声音效果的合成
链接:https://arxiv.org/abs/2412.15023
作者:Riccardo Fosco Gramaccioni,  Christian Marinoni,  Emilian Postolache,  Marco Comunità,  Luca Cosmo,  Joshua D. Reiss,  Danilo Comminiello
摘要:声音设计师和Foley艺术家通常通过手动注释和声音化视频中感兴趣的每个动作来对场景进行声音化,例如来自电影或视频游戏的场景。在我们的案例中,目的是将完全的创造性控制权留给声音设计师,让他们能够绕过工作中更重复的部分,从而能够专注于声音制作的创造性方面。我们实现了这一目标,提出了Stable-V2 A,一个两阶段模型,包括:RMS映射器,估计与输入视频相关的音频特征的包络代表;和Stable-Foley,一个基于Stable Audio Open的扩散模型,生成与目标视频语义和时间对齐的音频。时间对齐是通过使用信封作为ControlNet输入来保证的,而语义对齐是通过使用设计师选择的声音表示作为扩散过程的交叉注意调节来实现的。我们在Greatest Hits上训练和测试我们的模型,Greatest Hits是一个通常用于评估V2 A模型的数据集。此外,为了在感兴趣的案例研究中测试我们的模型,我们引入了Walking The Maps,这是一个从视频游戏中提取的视频数据集,描述了在不同位置行走的动画角色。样品和代码可在我们的演示页面https://ispamm.github.io/Stable-V2A。
摘要:Sound designers and Foley artists usually sonorize a scene, such as from amovie or video game, by manually annotating and sonorizing each action ofinterest in the video. In our case, the intent is to leave full creativecontrol to sound designers with a tool that allows them to bypass the morerepetitive parts of their work, thus being able to focus on the creativeaspects of sound production. We achieve this presenting Stable-V2A, a two-stagemodel consisting of: an RMS-Mapper that estimates an envelope representative ofthe audio characteristics associated with the input video; and Stable-Foley, adiffusion model based on Stable Audio Open that generates audio semanticallyand temporally aligned with the target video. Temporal alignment is guaranteedby the use of the envelope as a ControlNet input, while semantic alignment isachieved through the use of sound representations chosen by the designer ascross-attention conditioning of the diffusion process. We train and test ourmodel on Greatest Hits, a dataset commonly used to evaluate V2A models. Inaddition, to test our model on a case study of interest, we introduce WalkingThe Maps, a dataset of videos extracted from video games depicting animatedcharacters walking in different locations. Samples and code available on ourdemo page at https://ispamm.github.io/Stable-V2A.

【4】 Scale This, Not That: Investigating Key Dataset Attributes for Efficient  Speech Enhancement Scaling
标题:缩放这个,而不是那个:调查关键数据集属性以实现高效的语音增强缩放
链接:https://arxiv.org/abs/2412.14890
作者:Leying Zhang,  Wangyou Zhang,  Chenda Li,  Yanmin Qian
摘要:最近的语音增强模型已经显示出令人印象深刻的性能增益,按比例增加模型的复杂性和训练数据。然而,数据集可变性(例如文本、语言、说话者和噪声)的影响尚未得到充分研究。单独分析每个属性通常具有挑战性,因为多个属性通常纠缠在常用的数据集中,这对理解每个属性对模型性能的不同贡献构成了重大障碍。为了解决这一挑战,我们提出了一个生成-训练-评估框架,利用zero-shot文本到语音系统来研究受控属性变化对语音增强性能的影响。它使我们能够以可扩展的方式合成训练数据集,同时仔细更改每个属性。基于所提出的框架,我们分析了各种数据集属性对判别式和生成式SE模型性能的缩放效应。在多领域语料库上的广泛实验表明,声学属性(例如,说话者和噪声)对于当前的语音增强模型比语义属性(例如,语言和文本),为未来的研究提供新的见解。
摘要:Recent speech enhancement models have shown impressive performance gains byscaling up model complexity and training data. However, the impact of datasetvariability (e.g. text, language, speaker, and noise) has been underexplored.Analyzing each attribute individually is often challenging, as multipleattributes are usually entangled in commonly used datasets, posing asignificant obstacle in understanding the distinct contributions of eachattribute to the model's performance. To address this challenge, we propose ageneration-training-evaluation framework that leverages zero-shottext-to-speech systems to investigate the impact of controlled attributevariations on speech enhancement performance. It enables us to synthesizetraining datasets in a scalable manner while carefully altering each attribute.Based on the proposed framework, we analyze the scaling effects of variousdataset attributes on the performance of both discriminative and generative SEmodels. Extensive experiments on multi-domain corpora imply that acousticattributes (e.g., speaker and noise) are much more important to current speechenhancement models than semantic attributes (e.g., language and text), offeringnew insights for future research.

eess.AS音频处理

【1】 Scale This, Not That: Investigating Key Dataset Attributes for Efficient  Speech Enhancement Scaling
标题:缩放这个,而不是那个:调查关键数据集属性以实现高效的语音增强缩放
链接:https://arxiv.org/abs/2412.14890
作者:Leying Zhang,  Wangyou Zhang,  Chenda Li,  Yanmin Qian
摘要:最近的语音增强模型已经显示出令人印象深刻的性能增益,按比例增加模型的复杂性和训练数据。然而,数据集可变性(例如文本、语言、说话者和噪声)的影响尚未得到充分研究。单独分析每个属性通常具有挑战性,因为多个属性通常纠缠在常用的数据集中,这对理解每个属性对模型性能的不同贡献构成了重大障碍。为了解决这一挑战,我们提出了一个生成-训练-评估框架,利用zero-shot文本到语音系统来研究受控属性变化对语音增强性能的影响。它使我们能够以可扩展的方式合成训练数据集,同时仔细更改每个属性。基于所提出的框架,我们分析了各种数据集属性对判别式和生成式SE模型性能的缩放效应。在多领域语料库上的广泛实验表明,声学属性(例如,说话者和噪声)对于当前的语音增强模型比语义属性(例如,语言和文本),为未来的研究提供新的见解。
摘要:Recent speech enhancement models have shown impressive performance gains byscaling up model complexity and training data. However, the impact of datasetvariability (e.g. text, language, speaker, and noise) has been underexplored.Analyzing each attribute individually is often challenging, as multipleattributes are usually entangled in commonly used datasets, posing asignificant obstacle in understanding the distinct contributions of eachattribute to the model's performance. To address this challenge, we propose ageneration-training-evaluation framework that leverages zero-shottext-to-speech systems to investigate the impact of controlled attributevariations on speech enhancement performance. It enables us to synthesizetraining datasets in a scalable manner while carefully altering each attribute.Based on the proposed framework, we analyze the scaling effects of variousdataset attributes on the performance of both discriminative and generative SEmodels. Extensive experiments on multi-domain corpora imply that acousticattributes (e.g., speaker and noise) are much more important to current speechenhancement models than semantic attributes (e.g., language and text), offeringnew insights for future research.

【2】 AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal  Audio-Video Generation
标题:AV-Link:用于跨模式音频视频生成的时间对齐扩散功能
链接:https://arxiv.org/abs/2412.15191
作者:Moayed Haji-Ali,  Willi Menapace,  Aliaksandr Siarohin,  Ivan Skorokhodov,  Alper Canberk,  Kwot Sin Lee,  Vicente Ordonez,  Sergey Tulyakov
备注:Project Page: snap-research.github.io/AVLink/
摘要:我们提出了AV-Link,一个统一的视频到音频和音频到视频生成框架,利用冻结的视频和音频扩散模型的激活时间对齐的跨模态条件反射。我们的框架的关键是一个融合块,通过时间对齐的自我注意操作,使我们的骨干视频和音频扩散模型之间的双向信息交换。与使用针对条件信号的其他任务预先训练的特征提取器的先前工作不同,AV-Link可以在单个框架中直接利用由互补模态获得的特征,即视频特征来生成音频,或音频特征来生成视频。我们广泛评估我们的设计选择,并展示我们的方法,以实现同步和高质量的视听内容的能力,展示其在沉浸式媒体生成的应用潜力。项目页面:snap-research.github.io/AVLink/
摘要:We propose AV-Link, a unified framework for Video-to-Audio and Audio-to-Videogeneration that leverages the activations of frozen video and audio diffusionmodels for temporally-aligned cross-modal conditioning. The key to ourframework is a Fusion Block that enables bidirectional information exchangebetween our backbone video and audio diffusion models through atemporally-aligned self attention operation. Unlike prior work that usesfeature extractors pretrained for other tasks for the conditioning signal,AV-Link can directly leverage features obtained by the complementary modalityin a single framework i.e. video features to generate audio, or audio featuresto generate video. We extensively evaluate our design choices and demonstratethe ability of our method to achieve synchronized and high-quality audiovisualcontent, showcasing its potential for applications in immersive mediageneration. Project Page: snap-research.github.io/AVLink/

【3】 GIRAFE: Glottal Imaging Dataset for Advanced Segmentation, Analysis, and  Facilitative Playbacks Evaluation
标题:GIRAFE:用于高级分割、分析和辅助回放评估的喉舌成像数据集
链接:https://arxiv.org/abs/2412.15054
作者:G. Andrade-Miranda,  K. Chatzipapas,  J.D. Arias-Londoño,  J. I. Godino-Llorente
备注:18 pages, 8 figures
摘要:从声带的高速视频内窥镜序列中提取的促进性回放的发展的进展受到阻碍,这是由于明显缺乏用对应于声门间隙区域的语义分割注释的公开可用的数据集。这一事实也限制了该领域现有研究的可重复性和进一步探索。  为了解决这一差距,GIRAFE是一个数据存储库,旨在促进先进技术的发展,语义分割,分析和快速评估的高速视频内窥镜序列的声带。存储库包括来自50名患者(30名女性,20名男性)的65个高速视频内窥镜记录。该数据集包括来自健康对照组的15个记录,来自诊断为语音障碍的患者的26个记录,以及24个未知健康状况的记录。所有这些都由专家手动注释,包括与声门间隙的语义分割相对应的掩模。该存储库还配备了使用不同最先进方法对声门区域进行自动分割的功能。  该数据集已经支持了多项研究,证明了其对开发高速视频内窥镜序列的新声门间隙分割算法以改善或创建新的促进性回放的有用性。尽管在该领域取得了这些进展和其他进展,但执行声门区的准确且完全自动的语义分割方法的更广泛挑战仍然是开放的。
摘要:The advances in the development of Facilitative Playbacks extracted fromHigh-Speed videoendoscopic sequences of the vocal folds are hindered by anotable lack of publicly available datasets annotated with the semanticsegmentations corresponding to the area of the glottal gap. This fact alsolimits the reproducibility and further exploration of existing research in thisfield. To address this gap, GIRAFE is a data repository designed to facilitate thedevelopment of advanced techniques for the semantic segmentation, analysis, andfast evaluation of High-Speed videoendoscopic sequences of the vocal folds. Therepository includes 65 high-speed videoendoscopic recordings from a cohort of50 patients (30 female, 20 male). The dataset comprises 15 recordings fromhealthy controls, 26 from patients with diagnosed voice disorders, and 24 withan unknown health condition. All of them were manually annotated by an expert,including the masks corresponding to the semantic segmentation of the glottalgap. The repository is also complemented with the automatic segmentation of theglottal area using different state-of-the-art approaches. This data set has already supported several studies, which demonstrates itsusefulness for the development of new glottal gap segmentation algorithms fromHigh-Speed-Videoendoscopic sequences to improve or create new FacilitativePlaybacks. Despite these advances and others in the field, the broaderchallenge of performing an accurate and completely automatic semanticsegmentation method of the glottal area remains open.

【4】 Stable-V2A: Synthesis of Synchronized Sound Effects with Temporal and  Semantic Controls
标题:Stable-V2 A:具有时间和语义控制的同步声音效果的合成
链接:https://arxiv.org/abs/2412.15023
作者:Riccardo Fosco Gramaccioni,  Christian Marinoni,  Emilian Postolache,  Marco Comunità,  Luca Cosmo,  Joshua D. Reiss,  Danilo Comminiello
摘要:声音设计师和Foley艺术家通常通过手动注释和声音化视频中感兴趣的每个动作来对场景进行声音化,例如来自电影或视频游戏的场景。在我们的案例中,目的是将完全的创造性控制权留给声音设计师,让他们能够绕过工作中更重复的部分,从而能够专注于声音制作的创造性方面。我们实现了这一目标,提出了Stable-V2 A,一个两阶段模型,包括:RMS映射器,估计与输入视频相关的音频特征的包络代表;和Stable-Foley,一个基于Stable Audio Open的扩散模型,生成与目标视频语义和时间对齐的音频。时间对齐是通过使用信封作为ControlNet输入来保证的,而语义对齐是通过使用设计师选择的声音表示作为扩散过程的交叉注意调节来实现的。我们在Greatest Hits上训练和测试我们的模型,Greatest Hits是一个通常用于评估V2 A模型的数据集。此外,为了在感兴趣的案例研究中测试我们的模型,我们引入了Walking The Maps,这是一个从视频游戏中提取的视频数据集,描述了在不同位置行走的动画角色。样品和代码可在我们的演示页面https://ispamm.github.io/Stable-V2A。
摘要:Sound designers and Foley artists usually sonorize a scene, such as from amovie or video game, by manually annotating and sonorizing each action ofinterest in the video. In our case, the intent is to leave full creativecontrol to sound designers with a tool that allows them to bypass the morerepetitive parts of their work, thus being able to focus on the creativeaspects of sound production. We achieve this presenting Stable-V2A, a two-stagemodel consisting of: an RMS-Mapper that estimates an envelope representative ofthe audio characteristics associated with the input video; and Stable-Foley, adiffusion model based on Stable Audio Open that generates audio semanticallyand temporally aligned with the target video. Temporal alignment is guaranteedby the use of the envelope as a ControlNet input, while semantic alignment isachieved through the use of sound representations chosen by the designer ascross-attention conditioning of the diffusion process. We train and test ourmodel on Greatest Hits, a dataset commonly used to evaluate V2A models. Inaddition, to test our model on a case study of interest, we introduce WalkingThe Maps, a dataset of videos extracted from video games depicting animatedcharacters walking in different locations. Samples and code available on ourdemo page at https://ispamm.github.io/Stable-V2A.

机器翻译由腾讯交互翻译提供,仅供参考