今日论文合集:cs.SD语音6篇,eess.AS音频处理5篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音
【1】SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale  Spectral Diffusion Model
标题:SINGER:采用多尺度光谱扩散模型的生动音频驱动歌唱视频生成
链接:https://arxiv.org/abs/2412.03430
作者:Yan Li,  Ziya Zhou,  Zhiqiang Wang,  Wei Xue,  Wenhan Luo,  Yike Guo
摘要:生成模型的最新进展显着增强了说话的脸视频生成,但唱歌视频生成仍然未被探索。人类说话和唱歌之间的差异限制了现有的说话人脸视频生成模型在应用于唱歌时的性能。说话和唱歌之间的根本差异,特别是在音频特性和行为表达,限制了现有模型的有效性。我们观察到唱歌和说话音频之间的差异体现在频率和幅度方面。为了解决这个问题,我们设计了一个多尺度谱模块来帮助模型学习谱域中的歌唱模式。此外,我们开发了一个频谱过滤模块,帮助模型学习与唱歌音频相关的人类行为。这两个模块被集成到扩散模型中以增强歌唱视频生成性能,从而产生了我们提出的模型SINGER。此外,缺乏高质量的真实世界歌唱人脸视频阻碍了歌唱视频生成社区的发展。为了解决这一差距,我们收集了一个野外视听歌唱数据集,以促进这一领域的研究。我们的实验表明,辛格是能够产生生动的歌唱视频,并优于国家的最先进的方法,在客观和主观评价。
摘要:Recent advancements in generative models have significantly enhanced talkingface video generation, yet singing video generation remains underexplored. Thedifferences between human talking and singing limit the performance of existingtalking face video generation models when applied to singing. The fundamentaldifferences between talking and singing-specifically in audio characteristicsand behavioral expressions-limit the effectiveness of existing models. Weobserve that the differences between singing and talking audios manifest interms of frequency and amplitude. To address this, we have designed amulti-scale spectral module to help the model learn singing patterns in thespectral domain. Additionally, we develop a spectral-filtering module that aidsthe model in learning the human behaviors associated with singing audio. Thesetwo modules are integrated into the diffusion model to enhance singing videogeneration performance, resulting in our proposed model, SINGER. Furthermore,the lack of high-quality real-world singing face videos has hindered thedevelopment of the singing video generation community. To address this gap, wehave collected an in-the-wild audio-visual singing dataset to facilitateresearch in this area. Our experiments demonstrate that SINGER is capable ofgenerating vivid singing videos and outperforms state-of-the-art methods inboth objective and subjective evaluations.

【2】 DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for  Text-to-Speech with Diverse and Controllable Styles
标题:DistStyleTTC:基于扩散的分层韵律建模,用于具有多样化和可控风格的文本到语音
链接:https://arxiv.org/abs/2412.03388
作者:Jiaxuan Liu,  Zhaoci Liu,  Yajun Hu,  Yingying Gao,  Shilei Zhang,  Zhenhua Ling
备注:COLING 2025
摘要:人类语言具有丰富而灵活的韵律变化。为了解决从文本到韵律的一对多映射问题,提出了一种基于条件扩散模型和改进的无分类器引导的多说话人声学模型DiffStyleTTS,该模型对语音韵律特征进行分层建模,并控制不同的韵律风格来指导韵律预测。实验表明,我们的方法优于所有基线的自然度和实现优越的合成速度相比,三个基于扩散的基线。此外,通过调整引导尺度,DiffStyleTTS有效地控制了合成韵律的引导强度。
摘要:Human speech exhibits rich and flexible prosodic variations. To address theone-to-many mapping problem from text to prosody in a reasonable and flexiblemanner, we propose DiffStyleTTS, a multi-speaker acoustic model based on aconditional diffusion module and an improved classifier-free guidance, whichhierarchically models speech prosodic features, and controls different prosodicstyles to guide prosody prediction. Experiments show that our methodoutperforms all baselines in naturalness and achieves superior synthesis speedcompared to three diffusion-based baselines. Additionally, by adjusting theguiding scale, DiffStyleTTS effectively controls the guidance intensity of thesynthetic prosody.

【3】 Exploring trends in audio mixes and masters: Insights from a dataset  analysis
标题:探索混音和母版的趋势:数据集分析的见解
链接:https://arxiv.org/abs/2412.03373
作者:Angeliki Mourgela,  Elio Quinton,  Spyridon Bissas,  Joshua D. Reiss,  David Ronan
备注:11 pages, 6 figures, Presented at the AES 157th Convention October 2024, New York, USA
摘要:我们提出了一个数据集的音频指标和美学考虑的混合和主机提供的网络平台MixCheck工作室的分析。该平台是为教育目的而设计的,主要针对业余音乐制作人,旨在在他们的录音发布之前对其进行分析。分析集中在以下数据点:综合响度,单声道兼容性,存在的削波和相位问题,压缩和音调配置文件在30个用户指定的流派。混合(混音)和母版音频(母版)都包含在分析中,其中混音指的是单个曲目的初始组合和平衡,母版指的是针对分发进行优化的最终改进版本。结果表明,与动态问题相关的问题是最普遍的,特别是在掌握音频。然而,掌握音频呈现更好的结果,在压缩比只是混合音频。此外,结果表明,掌握的音频具有较低的立体声场和相位问题的百分比。
摘要:We present an analysis of a dataset of audio metrics and aestheticconsiderations about mixes and masters provided by the web platform MixCheckstudio. The platform is designed for educational purposes, primarily targetingamateur music producers, and aimed at analysing their recordings prior to thembeing released. The analysis focuses on the following data points: integratedloudness, mono compatibility, presence of clipping and phase issues,compression and tonal profile across 30 user-specified genres. Both mixed(mixes) and mastered audio (masters) are included in the analysis, where mixesrefer to the initial combination and balance of individual tracks, and mastersrefer to the final refined version optimized for distribution. Results showthat loudness-related issues along with dynamics issues are the most prevalent,particularly in mastered audio. However mastered audio presents better resultsin compression than just mixed audio. Additionally, results show that masteredaudio has a lower percentage of stereo field and phase issues.

【4】 Detecting abnormal heart sound using mobile phones and on-device IConNet
标题:使用手机和设备上IConNet检测异常心弦
链接:https://arxiv.org/abs/2412.03267
作者:Linh Vu,  Thu Tran
备注:N2Women'24 Workshop, MobiSys 2024, Tokyo, Japan
摘要:鉴于心血管疾病的全球流行,迫切需要易于获得的早期筛查方法。通常情况下,这需要医生调查心脏听诊不规则的声音,其次是超声心动图和心电图测试。为了使早期诊断民主化,我们提出了一种用户友好的异常心音检测解决方案,利用手机和针对设备上推理优化的轻量级神经网络。与以前依赖专业听诊器的方法不同,我们的方法直接分析录音,由一种名为IConNet的新型架构提供便利。IConNet是一种可解释的卷积神经网络,它利用了音频信号处理的洞察力,提高了效率,并提供了从原始波形信号中提取神经模式的透明度。这是迈向医疗保健领域值得信赖的人工智能的重要一步,有助于远程健康监测工作。
摘要:Given the global prevalence of cardiovascular diseases, there is a pressingneed for easily accessible early screening methods. Typically, this requiresmedical practitioners to investigate heart auscultations for irregular sounds,followed by echocardiography and electrocardiography tests. To democratizeearly diagnosis, we present a user-friendly solution for abnormal heart sounddetection, utilizing mobile phones and a lightweight neural network optimizedfor on-device inference. Unlike previous approaches reliant on specializedstethoscopes, our method directly analyzes audio recordings, facilitated by anovel architecture known as IConNet. IConNet, an Interpretable ConvolutionalNeural Network, harnesses insights from audio signal processing, enhancingefficiency and providing transparency in neural pattern extraction from rawwaveform signals. This is a significant step towards trustworthy AI inhealthcare, aiding in remote health monitoring efforts.

【5】 ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error  Correction
标题:ASR-EC基准:评估大型语言模型的中文ASB错误纠正
链接:https://arxiv.org/abs/2412.03075
作者:Victor Junqiu Wei,  Weicheng Wang,  Di Jiang,  Yuanfeng Song,  Lu Wang
摘要:自动语音识别(ASR)是语音和自然语言处理领域的一项基础和重要的研究课题。它是许多应用程序中固有的构建块,如语音助理,语音翻译等,尽管近年来ASR技术的进步,它仍然是不可避免的现代ASR系统有大量的错误识别,由于环境噪声,歧义等,因此,在ASR的纠错是至关重要的。  基于此,本文研究了在世界上最流行的语言之一,拥有大量用户的汉语中的自动语音识别纠错。我们首先创建一个名为\ldblquote ASR-EC}的基准数据集,该数据集包含由行业级ASR系统生成的广泛的ASR错误。据我们所知,它是第一个中文ASR纠错基准。然后,受大语言模型(LLM)的最新进展的启发,我们研究如何利用LLM的力量来纠正ASR错误。我们将LLM应用于三种范式的ASR纠错。第一种范例是提示,其进一步分类为zero-shot、Few-Shot和多步。第二个范例是微调,它用ASR纠错数据微调LLM。第三个范例是多模态增强,它共同利用音频和ASR成绩单进行纠错。大量的实验表明,提示对ASR错误纠正是无效的。微调仅对一部分LLM有效。多模态增强是最有效的纠错方法,并实现了最先进的性能。
摘要:Automatic speech Recognition (ASR) is a fundamental and important task in thefield of speech and natural language processing. It is an inherent buildingblock in many applications such as voice assistant, speech translation, etc.Despite the advancement of ASR technologies in recent years, it is stillinevitable for modern ASR systems to have a substantial number of erroneousrecognition due to environmental noise, ambiguity, etc. Therefore, the errorcorrection in ASR is crucial. Motivated by this, this paper studies ASR error correction in the Chineselanguage, which is one of the most popular languages and enjoys a large numberof users in the world. We first create a benchmark dataset named \emph{ASR-EC}that contains a wide spectrum of ASR errors generated by industry-grade ASRsystems. To the best of our knowledge, it is the first Chinese ASR errorcorrection benchmark. Then, inspired by the recent advances in \emph{largelanguage models (LLMs)}, we investigate how to harness the power of LLMs tocorrect ASR errors. We apply LLMs to ASR error correction in three paradigms.The first paradigm is prompting, which is further categorized as zero-shot,few-shot, and multi-step. The second paradigm is finetuning, which finetunesLLMs with ASR error correction data. The third paradigm is multi-modalaugmentation, which collectively utilizes the audio and ASR transcripts forerror correction. Extensive experiments reveal that prompting is not effectivefor ASR error correction. Finetuning is effective only for a portion of LLMs.Multi-modal augmentation is the most effective method for error correction andachieves state-of-the-art performance.

【6】 Analytic Study of Text-Free Speech Synthesis for Raw Audio using a  Self-Supervised Learning Model
标题:使用自我监督学习模型的原始音频无文本语音合成分析研究
链接:https://arxiv.org/abs/2412.03074
作者:Joonyong Park,  Daisuke Saito,  Nobuaki Minematsu
备注:APSIPA ASC 2024
摘要:我们研究的文本无语音表示的原始音频获得自监督学习(SSL)模型通过分析合成语音使用SSL表示,而不是传统的文本表示。由于原始音频不像转录文本那样具有配对的语音表示,因此从未配对的语音中获得语音表示对于增强语音合成的可用数据集至关重要。具体而言,建议的语音合成进行使用离散符号表示从SSL模型与文本表示的比较,并已进行合成语音的分析检查。实验结果表明,使用文本表示是有利于保留语义信息,而使用离散符号表示是优于保留声学内容,包括韵律和语调信息。
摘要:We examine the text-free speech representations of raw audio obtained from aself-supervised learning (SSL) model by analyzing the synthesized speech usingthe SSL representations instead of conventional text representations. Since rawaudio does not have paired speech representations as transcribed texts do,obtaining speech representations from unpaired speech is crucial for augmentingavailable datasets for speech synthesis. Specifically, the proposed speechsynthesis is conducted using discrete symbol representations from the SSL modelin comparison with text representations, and analytical examinations of thesynthesized speech have been carried out. The results empirically show thatusing text representations is advantageous for preserving semantic information,while using discrete symbol representations is superior for preserving acousticcontent, including prosodic and intonational information.

eess.AS音频处理

【1】 DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for  Text-to-Speech with Diverse and Controllable Styles
标题:DistStyleTTC:基于扩散的分层韵律建模,用于具有多样化和可控风格的文本到语音
链接:https://arxiv.org/abs/2412.03388
作者:Jiaxuan Liu,  Zhaoci Liu,  Yajun Hu,  Yingying Gao,  Shilei Zhang,  Zhenhua Ling
备注:COLING 2025
摘要:人类语言具有丰富而灵活的韵律变化。为了解决从文本到韵律的一对多映射问题,提出了一种基于条件扩散模型和改进的无分类器引导的多说话人声学模型DiffStyleTTS,该模型对语音韵律特征进行分层建模,并控制不同的韵律风格来指导韵律预测。实验表明,我们的方法优于所有基线的自然度和实现优越的合成速度相比,三个基于扩散的基线。此外,通过调整引导尺度,DiffStyleTTS有效地控制了合成韵律的引导强度。
摘要:Human speech exhibits rich and flexible prosodic variations. To address theone-to-many mapping problem from text to prosody in a reasonable and flexiblemanner, we propose DiffStyleTTS, a multi-speaker acoustic model based on aconditional diffusion module and an improved classifier-free guidance, whichhierarchically models speech prosodic features, and controls different prosodicstyles to guide prosody prediction. Experiments show that our methodoutperforms all baselines in naturalness and achieves superior synthesis speedcompared to three diffusion-based baselines. Additionally, by adjusting theguiding scale, DiffStyleTTS effectively controls the guidance intensity of thesynthetic prosody.

【2】 Exploring trends in audio mixes and masters: Insights from a dataset  analysis
标题:探索混音和母版的趋势:数据集分析的见解
链接:https://arxiv.org/abs/2412.03373
作者:Angeliki Mourgela,  Elio Quinton,  Spyridon Bissas,  Joshua D. Reiss,  David Ronan
备注:11 pages, 6 figures, Presented at the AES 157th Convention October 2024, New York, USA
摘要:我们提出了一个数据集的音频指标和美学考虑的混合和主机提供的网络平台MixCheck工作室的分析。该平台是为教育目的而设计的,主要针对业余音乐制作人,旨在在他们的录音发布之前对其进行分析。分析集中在以下数据点:综合响度,单声道兼容性,存在的削波和相位问题,压缩和音调配置文件在30个用户指定的流派。混合(混音)和母版音频(母版)都包含在分析中,其中混音指的是单个曲目的初始组合和平衡,母版指的是针对分发进行优化的最终改进版本。结果表明,与动态问题相关的问题是最普遍的,特别是在掌握音频。然而,掌握音频呈现更好的结果,在压缩比只是混合音频。此外,结果表明,掌握的音频具有较低的立体声场和相位问题的百分比。
摘要:We present an analysis of a dataset of audio metrics and aestheticconsiderations about mixes and masters provided by the web platform MixCheckstudio. The platform is designed for educational purposes, primarily targetingamateur music producers, and aimed at analysing their recordings prior to thembeing released. The analysis focuses on the following data points: integratedloudness, mono compatibility, presence of clipping and phase issues,compression and tonal profile across 30 user-specified genres. Both mixed(mixes) and mastered audio (masters) are included in the analysis, where mixesrefer to the initial combination and balance of individual tracks, and mastersrefer to the final refined version optimized for distribution. Results showthat loudness-related issues along with dynamics issues are the most prevalent,particularly in mastered audio. However mastered audio presents better resultsin compression than just mixed audio. Additionally, results show that masteredaudio has a lower percentage of stereo field and phase issues.

【3】 Detecting abnormal heart sound using mobile phones and on-device IConNet
标题:使用手机和设备上IConNet检测异常心弦
链接:https://arxiv.org/abs/2412.03267
作者:Linh Vu,  Thu Tran
备注:N2Women'24 Workshop, MobiSys 2024, Tokyo, Japan
摘要:鉴于心血管疾病的全球流行,迫切需要易于获得的早期筛查方法。通常情况下,这需要医生调查心脏听诊不规则的声音,其次是超声心动图和心电图测试。为了使早期诊断民主化,我们提出了一种用户友好的异常心音检测解决方案,利用手机和针对设备上推理优化的轻量级神经网络。与以前依赖专业听诊器的方法不同,我们的方法直接分析录音,由一种名为IConNet的新型架构提供便利。IConNet是一种可解释的卷积神经网络,它利用了音频信号处理的洞察力,提高了效率,并提供了从原始波形信号中提取神经模式的透明度。这是迈向医疗保健领域值得信赖的人工智能的重要一步,有助于远程健康监测工作。
摘要:Given the global prevalence of cardiovascular diseases, there is a pressingneed for easily accessible early screening methods. Typically, this requiresmedical practitioners to investigate heart auscultations for irregular sounds,followed by echocardiography and electrocardiography tests. To democratizeearly diagnosis, we present a user-friendly solution for abnormal heart sounddetection, utilizing mobile phones and a lightweight neural network optimizedfor on-device inference. Unlike previous approaches reliant on specializedstethoscopes, our method directly analyzes audio recordings, facilitated by anovel architecture known as IConNet. IConNet, an Interpretable ConvolutionalNeural Network, harnesses insights from audio signal processing, enhancingefficiency and providing transparency in neural pattern extraction from rawwaveform signals. This is a significant step towards trustworthy AI inhealthcare, aiding in remote health monitoring efforts.

【4】 ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error  Correction
标题:ASR-EC基准:评估大型语言模型的中文ASB错误纠正
链接:https://arxiv.org/abs/2412.03075
作者:Victor Junqiu Wei,  Weicheng Wang,  Di Jiang,  Yuanfeng Song,  Lu Wang
摘要:自动语音识别(ASR)是语音和自然语言处理领域的一项基础和重要的研究课题。它是许多应用程序中固有的构建块,如语音助理,语音翻译等,尽管近年来ASR技术的进步,它仍然是不可避免的现代ASR系统有大量的错误识别,由于环境噪声,歧义等,因此,在ASR的纠错是至关重要的。  基于此,本文研究了在世界上最流行的语言之一,拥有大量用户的汉语中的自动语音识别纠错。我们首先创建一个名为\{ASR-EC}的基准数据集,其中包含工业级ASR系统生成的广泛的ASR错误。据我们所知,它是第一个中文ASR纠错基准。然后,受大语言模型(LLM)的最新进展的启发,我们研究如何利用LLM的力量来纠正ASR错误。我们将LLM应用于三种范式的ASR纠错。第一种范例是提示,其进一步分类为zero-shot、Few-Shot和多步。第二种范式是微调,它使用ASR纠错数据对LLM进行微调。第三个范例是多模态增强,它共同利用音频和ASR成绩单进行纠错。大量的实验表明,提示对ASR错误纠正是无效的。微调仅对一部分LLM有效。多模态增强是最有效的纠错方法,并实现了最先进的性能。
摘要:Automatic speech Recognition (ASR) is a fundamental and important task in thefield of speech and natural language processing. It is an inherent buildingblock in many applications such as voice assistant, speech translation, etc.Despite the advancement of ASR technologies in recent years, it is stillinevitable for modern ASR systems to have a substantial number of erroneousrecognition due to environmental noise, ambiguity, etc. Therefore, the errorcorrection in ASR is crucial. Motivated by this, this paper studies ASR error correction in the Chineselanguage, which is one of the most popular languages and enjoys a large numberof users in the world. We first create a benchmark dataset named \emph{ASR-EC}that contains a wide spectrum of ASR errors generated by industry-grade ASRsystems. To the best of our knowledge, it is the first Chinese ASR errorcorrection benchmark. Then, inspired by the recent advances in \emph{largelanguage models (LLMs)}, we investigate how to harness the power of LLMs tocorrect ASR errors. We apply LLMs to ASR error correction in three paradigms.The first paradigm is prompting, which is further categorized as zero-shot,few-shot, and multi-step. The second paradigm is finetuning, which finetunesLLMs with ASR error correction data. The third paradigm is multi-modalaugmentation, which collectively utilizes the audio and ASR transcripts forerror correction. Extensive experiments reveal that prompting is not effectivefor ASR error correction. Finetuning is effective only for a portion of LLMs.Multi-modal augmentation is the most effective method for error correction andachieves state-of-the-art performance.

【5】 Analytic Study of Text-Free Speech Synthesis for Raw Audio using a  Self-Supervised Learning Model
标题:使用自我监督学习模型的原始音频无文本语音合成分析研究
链接:https://arxiv.org/abs/2412.03074
作者:Joonyong Park,  Daisuke Saito,  Nobuaki Minematsu
备注:APSIPA ASC 2024
摘要:我们研究的文本无语音表示的原始音频获得自监督学习(SSL)模型通过分析合成语音使用SSL表示,而不是传统的文本表示。由于原始音频不像转录文本那样具有配对的语音表示,因此从未配对的语音中获得语音表示对于增强语音合成的可用数据集至关重要。具体而言,建议的语音合成进行使用离散符号表示从SSL模型与文本表示的比较,并已进行合成语音的分析检查。实验结果表明,使用文本表示是有利于保留语义信息,而使用离散符号表示是优于保留声学内容,包括韵律和语调信息。
摘要:We examine the text-free speech representations of raw audio obtained from aself-supervised learning (SSL) model by analyzing the synthesized speech usingthe SSL representations instead of conventional text representations. Since rawaudio does not have paired speech representations as transcribed texts do,obtaining speech representations from unpaired speech is crucial for augmentingavailable datasets for speech synthesis. Specifically, the proposed speechsynthesis is conducted using discrete symbol representations from the SSL modelin comparison with text representations, and analytical examinations of thesynthesized speech have been carried out. The results empirically show thatusing text representations is advantageous for preserving semantic information,while using discrete symbol representations is superior for preserving acousticcontent, including prosodic and intonational information.

机器翻译由腾讯交互翻译提供,仅供参考