本文经arXiv每日学术速递授权转载
【1】 Latent Watermarking of Audio Generative Models
标题: 音频生成模型的潜在水印
作者:Robin San Roman,Pierre Fernandez,Antoine Deleforge,Yossi Adi,Romain Serizel
链接:点击下载PDF文件
【2】 Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
标题: 多轨音乐LDM:利用潜在扩散模型实现多功能音乐生成
作者:Tornike Karchkhadze,Mohammad Rasool Izadi,Ke Chen,Gerard Assayag,Shlomo Dubnov
链接:点击下载PDF文件
【3】 Effects of Recording Condition and Number of Monitored Days on Discriminative Power of the Daily Phonotrauma Index
标题: 记录条件和监测天数对每日语音创伤指数区分力的影响
作者:Hamzeh Ghasemzadeh,Robert E. Hillman,Jarrad H. Van Stan,Daryush D. Mehta
备注:The paper is submitted to JSLHR
链接:点击下载PDF文件
【4】 An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
标题: 用BEST-PQ分析线性复杂性注意替代
作者:Ryan Whetten,Titouan Parcollet,Adel Moumen,Marco Dinarelli,Yannick Estève
备注:Accepted in the IEEE Soken Language Technology Workshop 2024
链接:点击下载PDF文件
【5】 Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems
标题: 使用基于特征平滑的增强方法训练通用声码器,用于高质量的TTC系统
作者:Jeongmin Liu,Eunwoo Song
备注:4 pages, 4 figures, for demo samples, see this https URL
链接:点击下载PDF文件
【6】 NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention
标题: NeuroSpex:具有跨模式注意力的神经引导说话者提取
作者:Dashanka De Silva,Siqi Cai,Saurav Pahuja,Tanja Schultz,Haizhou Li
链接:点击下载PDF文件
【7】 MusicMamba: A Dual-Feature Modeling Approach for Generating Chinese Traditional Music with Modal Precision
标题: MusicMamba:一种具有模式精确度的中国传统音乐的双特征建模方法
作者:Jiatao Chen,Tianming Xie,Xing Tang,Jing Wang,Wenjing Dong,Bing Shi
链接:点击下载PDF文件
【8】 STAB: Speech Tokenizer Assessment Benchmark
标题: STAB:语音令牌器评估基准
作者:Shikhar Vashishth,Harman Singh,Shikhar Bharadwaj,Sriram Ganapathy,Chulayuth Asawaroengchai,Kartik Audhkhasi,Andrew Rosenberg,Ankur Bapna,Bhuvana Ramabhadran
备注:5 pages
链接:点击下载PDF文件
【9】 LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement
标题: LSTMSE-Net:用于视听语音增强的长短期语音增强网络
作者:Arnav Jain,Jasmer Singh Sanjotra,Harshvardhan Choudhary,Krish Agrawal,Rupal Shah,Rohan Jha,M. Sajid,Amir Hussain,M. Tanveer
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
【10】 FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
标题: FastSecureGrad:基于扩散的一步语音转换,具有对抗性条件扩散蒸馏
作者:Takuhiro Kaneko,Hirokazu Kameoka,Kou Tanaka,Yuto Kondo
备注:Accepted to Interspeech 2024. Project page: this https URL
链接:点击下载PDF文件
【11】 Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR
标题: 基于时态保留的最优传输的ASB跨模式知识转移学习
作者:Xugang Lu,Peng Shen,Yu Tsao,Hisashi Kawai
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
【12】 USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
标题: USEF-PSE:通用说话人嵌入自由目标说话人提取
作者:Bang Zeng,Ming Li
备注:13 pages, 6 figures
链接:点击下载PDF文件
【13】 Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models
标题: 从自监督语音模型中高效提取噪音稳健的离散单元
作者:Jakob Poncelet,Yujun Wang,Hugo Van hamme
备注:Accepted at SLT2024
链接:点击下载PDF文件
【14】 CUEMPATHY: A Counseling Speech Dataset for Psychotherapy Research
标题: CuEMPathsy:心理治疗研究的咨询言语数据集
作者:Dehua Tao,Harold Chui,Sarah Luk,Tan Lee
备注:Accepted by ISCSLP 2022
链接:点击下载PDF文件
【15】 Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP
标题: 使用差异化DSP进行快速、高质量和参数高效的关节合成
作者:Yisi Liu,Bohan Yu,Drake Lin,Peter Wu,Cheol Jun Cho,Gopala Krishna Anumanchipalli
备注:accepted for Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
【16】 Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
标题: 2024年受控歌唱语音Deepfake检测(CtrSDDD)挑战赛的语音基金会模型集合
作者:Anmol Guragain,Tianchi Liu,Zihan Pan,Hardik B. Sailor,Qiongqiong Wang
备注:Accepted to the IEEE Spoken Language Technology Workshop (SLT) 2024. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
标题: USEF-PSE:通用说话人嵌入自由目标说话人提取
作者:Bang Zeng,Ming Li
备注:13 pages, 6 figures
链接:点击下载PDF文件
【2】 Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models
标题: 从自监督语音模型中高效提取噪音稳健的离散单元
作者:Jakob Poncelet,Yujun Wang,Hugo Van hamme
备注:Accepted at SLT2024
链接:点击下载PDF文件
【3】 CUEMPATHY: A Counseling Speech Dataset for Psychotherapy Research
标题: CuEMPathsy:心理治疗研究的咨询言语数据集
作者:Dehua Tao,Harold Chui,Sarah Luk,Tan Lee
备注:Accepted by ISCSLP 2022
链接:点击下载PDF文件
【4】 Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP
标题: 使用差异化DSP进行快速、高质量和参数高效的关节合成
作者:Yisi Liu,Bohan Yu,Drake Lin,Peter Wu,Cheol Jun Cho,Gopala Krishna Anumanchipalli
备注:accepted for Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
【5】 Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
标题: 2024年受控歌唱语音Deepfake检测(CtrSDDD)挑战赛的语音基金会模型集合
作者:Anmol Guragain,Tianchi Liu,Zihan Pan,Hardik B. Sailor,Qiongqiong Wang
备注:Accepted to the IEEE Spoken Language Technology Workshop (SLT) 2024. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
【6】 Latent Watermarking of Audio Generative Models
标题: 音频生成模型的潜在水印
作者:Robin San Roman,Pierre Fernandez,Antoine Deleforge,Yossi Adi,Romain Serizel
链接:点击下载PDF文件
【7】 Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
标题: 多轨音乐LDM:利用潜在扩散模型实现多功能音乐生成
作者:Tornike Karchkhadze,Mohammad Rasool Izadi,Ke Chen,Gerard Assayag,Shlomo Dubnov
链接:点击下载PDF文件
【8】 Effects of Recording Condition and Number of Monitored Days on Discriminative Power of the Daily Phonotrauma Index
标题: 记录条件和监测天数对每日语音创伤指数区分力的影响
作者:Hamzeh Ghasemzadeh,Robert E. Hillman,Jarrad H. Van Stan,Daryush D. Mehta
备注:The paper is submitted to JSLHR
链接:点击下载PDF文件
【9】 An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
标题: 用BEST-PQ分析线性复杂性注意替代
作者:Ryan Whetten,Titouan Parcollet,Adel Moumen,Marco Dinarelli,Yannick Estève
备注:Accepted in the IEEE Soken Language Technology Workshop 2024
链接:点击下载PDF文件
【10】 Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems
标题: 使用基于特征平滑的增强方法训练通用声码器,用于高质量的TTC系统
作者:Jeongmin Liu,Eunwoo Song
备注:4 pages, 4 figures, for demo samples, see this https URL
链接:点击下载PDF文件
【11】 NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention
标题: NeuroSpex:具有跨模式注意力的神经引导说话者提取
作者:Dashanka De Silva,Siqi Cai,Saurav Pahuja,Tanja Schultz,Haizhou Li
链接:点击下载PDF文件
【12】 MusicMamba: A Dual-Feature Modeling Approach for Generating Chinese Traditional Music with Modal Precision
标题: MusicMamba:一种具有模式精确度的中国传统音乐的双特征建模方法
作者:Jiatao Chen,Tianming Xie,Xing Tang,Jing Wang,Wenjing Dong,Bing Shi
链接:点击下载PDF文件
【13】 STAB: Speech Tokenizer Assessment Benchmark
标题: STAB:语音令牌器评估基准
作者:Shikhar Vashishth,Harman Singh,Shikhar Bharadwaj,Sriram Ganapathy,Chulayuth Asawaroengchai,Kartik Audhkhasi,Andrew Rosenberg,Ankur Bapna,Bhuvana Ramabhadran
备注:5 pages
链接:点击下载PDF文件
【14】 LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement
标题: LSTMSE-Net:用于视听语音增强的长短期语音增强网络
作者:Arnav Jain,Jasmer Singh Sanjotra,Harshvardhan Choudhary,Krish Agrawal,Rupal Shah,Rohan Jha,M. Sajid,Amir Hussain,M. Tanveer
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
【15】 FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
标题: FastSecureGrad:基于扩散的一步语音转换,具有对抗性条件扩散蒸馏
作者:Takuhiro Kaneko,Hirokazu Kameoka,Kou Tanaka,Yuto Kondo
备注:Accepted to Interspeech 2024. Project page: this https URL
链接:点击下载PDF文件
【16】 Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR
标题: 基于时态保留的最优传输的ASB跨模式知识转移学习
作者:Xugang Lu,Peng Shen,Yu Tsao,Hisashi Kawai
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
标题: 音频生成模型的潜在水印
作者:Robin San Roman,Pierre Fernandez,Antoine Deleforge,Yossi Adi,Romain Serizel
链接:点击下载PDF文件
摘要:音频生成模型的进步在其负责任的披露和检测其滥用方面带来了新的挑战。作为回应,我们介绍了一种方法,水印潜在的生成模型通过一个特定的水印的训练数据。所得到的水印模型产生潜在的表示,其解码输出被检测到具有高的信心,无论使用的解码方法。这种方法使得能够检测所生成的内容,而不需要事后水印步骤。它为开源模型提供了一个更安全的解决方案,并有助于识别那些在不遵守其许可条款的情况下微调或使用这些模型的衍生作品。我们的研究结果表明,例如,所产生的输出检测的准确性超过75%,误报率为10 ^{-3}$,即使微调后的潜在生成模型。摘要:The advancements in audio generative models have opened up new challenges in their responsible disclosure and the detection of their misuse. In response, we introduce a method to watermark latent generative models by a specific watermarking of their training data. The resulting watermarked models produce latent representations whose decoded outputs are detected with high confidence, regardless of the decoding method used. This approach enables the detection of the generated content without the need for a post-hoc watermarking step. It provides a more secure solution for open-sourced models and facilitates the identification of derivative works that fine-tune or use these models without adhering to their license terms. Our results indicate for instance that generated outputs are detected with an accuracy of more than 75% at a false positive rate of $10^{-3}$, even after fine-tuning the latent generative model.
【2】 Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
标题: 多轨音乐LDM:利用潜在扩散模型实现多功能音乐生成
作者:Tornike Karchkhadze,Mohammad Rasool Izadi,Ke Chen,Gerard Assayag,Shlomo Dubnov
链接:点击下载PDF文件
摘要:扩散模型在涉及音频和音乐的跨模态生成任务中显示出有希望的结果,例如文本到声音和文本到音乐生成。这些文本控制的音乐生成模型通常专注于通过捕获诸如流派和情绪之类的全局音乐属性来生成音乐。然而,音乐创作是一个复杂的,多层次的任务,往往涉及音乐安排作为一个不可分割的一部分的过程。这个过程涉及到将每种乐器组合成与现有乐器在节拍、力度、和声和旋律方面保持一致,这需要比文本提示通常提供的更高的精度和对曲目的控制。在这项工作中,我们通过扩展MusicLDM,一个潜在的音乐扩散模型,到一个多轨生成模型来解决这些挑战。通过学习共享上下文的曲目的联合概率,我们的模型能够有条件或无条件地在多个曲目中生成相互对应的音乐。此外,我们的模型能够生成排列,其中模型可以生成给定其他轨道的任何轨道子集(例如,产生补充给定低音和鼓音轨的钢琴音轨)。我们将我们的模型与现有的多轨道生成模型进行了比较,并证明了我们的模型在总体和安排生成任务的客观指标上都取得了相当大的改进。摘要:Diffusion models have shown promising results in cross-modal generation tasks involving audio and music, such as text-to-sound and text-to-music generation. These text-controlled music generation models typically focus on generating music by capturing global musical attributes like genre and mood. However, music composition is a complex, multilayered task that often involves musical arrangement as an integral part of the process. This process involves composing each instrument to align with existing ones in terms of beat, dynamics, harmony, and melody, requiring greater precision and control over tracks than text prompts usually provide. In this work, we address these challenges by extending the MusicLDM, a latent diffusion model for music, into a multi-track generative model. By learning the joint probability of tracks sharing a context, our model is capable of generating music across several tracks that correspond well to each other, either conditionally or unconditionally. Additionally, our model is capable of arrangement generation, where the model can generate any subset of tracks given the others (e.g., generating a piano track complementing given bass and drum tracks). We compared our model with an existing multi-track generative model and demonstrated that our model achieves considerable improvements across objective metrics for both total and arrangement generation tasks.
【3】 Effects of Recording Condition and Number of Monitored Days on Discriminative Power of the Daily Phonotrauma Index
标题: 记录条件和监测天数对每日语音创伤指数区分力的影响
作者:Hamzeh Ghasemzadeh,Robert E. Hillman,Jarrad H. Van Stan,Daryush D. Mehta
备注:The paper is submitted to JSLHR
链接:点击下载PDF文件
摘要:目的:每日语音创伤指数(DPI)可以量化与语音创伤性发声功能亢进(PVH)患者日常发声相关的病理生理机制。由于DPI是基于为期一周的动态语音监测开发的,因此本研究调查了DPI是否可以使用(1)短实验室语音任务和(2)少于7天的动态数据实现相当的性能。方法:动态语音监测系统记录了134名女性PVH和声音健康匹配的对照组在两种不同的条件下的声音功能 行为。在实验室中,参与者阅读彩虹通道的第一段,并产生自发的演讲(实验室数据)。然后对它们进行了7天的监测(现场数据)。使用前两个谐波(H1-H2)的幅度之间的差异的标准差和颈部表面加速度幅度的偏斜度,从实验室和现场数据中训练单独的DPI模型。首先,10倍交叉验证评估了实验室和现场DPI的分类性能。其次,量化了门诊监测天数对现场DPI分类准确性的影响。结果如下:从Rainbow通道和自发语音计算的平均实验室DPI准确度分别为57.9%和48.9%,接近偶然性能。在现场DPI的平均分类精度显着较高,具有非常大的效应大小(73.4%,科恩D = 1.8)。第二,实地DPI的平均准确度从一天的66.5%提高到七天的75.0%,增加一天的准确度在4天后下降到不到1个百分点。摘要:Objective: The Daily Phonotrauma Index (DPI) can quantify pathophysiological mechanisms associated with daily voice use in individuals with phonotraumatic vocal hyperfunction (PVH). Since DPI was developed based on week-long ambulatory voice monitoring, this study investigated if DPI can achieve comparable performance using (1) short laboratory speech tasks and (2) fewer than seven days of ambulatory data. Method: An ambulatory voice monitoring system recorded the vocal function behavior of 134 females with PVH and vocally healthy matched controls in two different conditions. In the lab, the participants read the first paragraph of the Rainbow Passage and produced spontaneous speech (in-lab data). They were then monitored for seven days (in-field data). Separate DPI models were trained from in-lab and in-field data using the standard deviation of the difference between the magnitude of the first two harmonics (H1-H2) and the skewness of neck-surface acceleration magnitude. First, 10-fold cross-validation evaluated classification performance of in-lab and in-field DPIs. Second, the effect of the number of ambulatory monitoring days on the accuracy of in-field DPI classification was quantified. Results: The average in-lab DPI accuracy computed from the Rainbow passage and spontaneous speech were, respectively, 57.9% and 48.9%, which are close to chance performance. The average classification accuracy of in-field DPI was significantly higher with a very large effect size (73.4%, Cohens D = 1.8). Second, the average in-field DPI accuracy increased from 66.5% for one day to 75.0% for seven days, with the gain of including an additional day on accuracy dropping below 1 percentage point after 4 days.
【4】 An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
标题: 用BEST-PQ分析线性复杂性注意替代
作者:Ryan Whetten,Titouan Parcollet,Adel Moumen,Marco Dinarelli,Yannick Estève
备注:Accepted in the IEEE Soken Language Technology Workshop 2024
链接:点击下载PDF文件
摘要:自监督学习(SSL)已被证明在包括语音处理在内的各个领域都是有效的。然而,SSL在计算和内存上是昂贵的。这部分是由于多头自注意(MHSA)的二次复杂性。MHSA的替代方案已被提出并用于语音领域,但尚未在SSL设置中进行适当的研究。在这项工作中,我们研究了用最新的具有线性复杂性的替代方案(即HyperMixing,Fastformer,SummaryMixing和Mamba)取代MHSA的效果。我们通过查看速度、消耗的VRAM量以及SSL MP3S基准测试的性能来评估这些方法。结果表明,与MHSA相比,这些线性替代方案保持了具有竞争力的性能,同时平均将VRAM消耗降低了约20%至60%,并将输入序列从20秒增加到80秒的速度从7%增加到65%。摘要:Self-Supervised Learning (SSL) has proven to be effective in various domains, including speech processing. However, SSL is computationally and memory expensive. This is in part due the quadratic complexity of multi-head self-attention (MHSA). Alternatives for MHSA have been proposed and used in the speech domain, but have yet to be investigated properly in an SSL setting. In this work, we study the effects of replacing MHSA with recent state-of-the-art alternatives that have linear complexity, namely, HyperMixing, Fastformer, SummaryMixing, and Mamba. We evaluate these methods by looking at the speed, the amount of VRAM consumed, and the performance on the SSL MP3S benchmark. Results show that these linear alternatives maintain competitive performance compared to MHSA while, on average, decreasing VRAM consumption by around 20% to 60% and increasing speed from 7% to 65% for input sequences ranging from 20 to 80 seconds.
【5】 Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems
标题: 使用基于特征平滑的增强方法训练通用声码器,用于高质量的TTC系统
作者:Jeongmin Liu,Eunwoo Song
备注:4 pages, 4 figures, for demo samples, see this https URL
链接:点击下载PDF文件
摘要:虽然通用声码器已经在不同的语音中实现了熟练的波形生成,但它们集成到文本到语音(TTS)任务中通常会导致合成质量下降。为了应对这一挑战,我们提出了一种用于训练通用声码器的新型增强技术。我们的训练方案随机地将线性平滑滤波器应用于输入声学特征,从而促进声码器在广泛的平滑范围内的泛化。它显著地减轻了训练-推理不匹配,增强了合成输出的自然性,即使声学模型产生过度平滑的特征。值得注意的是,我们的方法是适用于任何声码器,而不需要架构修改或依赖于特定的声学模型。实验结果验证了我们的声码器的优越性,传统的方法,实现11.99%和12.05%的平均意见分数时,与Tacotron 2和FastSpeech 2 TTS声学模型集成,分别提高。摘要:While universal vocoders have achieved proficient waveform generation across diverse voices, their integration into text-to-speech (TTS) tasks often results in degraded synthetic quality. To address this challenge, we present a novel augmentation technique for training universal vocoders. Our training scheme randomly applies linear smoothing filters to input acoustic features, facilitating vocoder generalization across a wide range of smoothings. It significantly mitigates the training-inference mismatch, enhancing the naturalness of synthetic output even when the acoustic model produces overly smoothed features. Notably, our method is applicable to any vocoder without requiring architectural modifications or dependencies on specific acoustic models. The experimental results validate the superiority of our vocoder over conventional methods, achieving 11.99% and 12.05% improvements in mean opinion scores when integrated with Tacotron 2 and FastSpeech 2 TTS acoustic models, respectively.
【6】 NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention
标题: NeuroSpex:具有跨模式注意力的神经引导说话者提取
作者:Dashanka De Silva,Siqi Cai,Saurav Pahuja,Tanja Schultz,Haizhou Li
链接:点击下载PDF文件
摘要:在听觉注意的研究中,已经发现在注意言语和诱发的神经反应之间存在鲁棒的相关性,其可以通过脑电图(EEG)来测量。因此,有可能使用EEG信号内可用的注意力信息来引导计算地提取鸡尾酒会中的目标说话者。在本文中,我们提出了一个神经引导的说话人提取模型,即NeuroSpex,使用的唯一辅助参考线索,从单耳语音混合的听众的EEG响应提取出席的语音。我们提出了一种新的脑电信号编码器,捕捉的注意力信息。此外,我们提出了一个交叉注意(CA)机制,以提高语音特征表示,生成一个扬声器提取掩码。在公开数据集上的实验结果表明,我们提出的模型在各种评估指标上优于两个基线模型。摘要:In the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is possible to use the attention information available within EEG signals to guide the extraction of the target speaker in a cocktail party computationally. In this paper, we present a neuro-guided speaker extraction model, i.e. NeuroSpex, using the EEG response of the listener as the sole auxiliary reference cue to extract attended speech from monaural speech mixtures. We propose a novel EEG signal encoder that captures the attention information. Additionally, we propose a cross-attention (CA) mechanism to enhance the speech feature representations, generating a speaker extraction mask. Experimental results on a publicly available dataset demonstrate that our proposed model outperforms two baseline models across various evaluation metrics.
【7】 MusicMamba: A Dual-Feature Modeling Approach for Generating Chinese Traditional Music with Modal Precision
标题: MusicMamba:一种具有模式精确度的中国传统音乐的双特征建模方法
作者:Jiatao Chen,Tianming Xie,Xing Tang,Jing Wang,Wenjing Dong,Bing Shi
链接:点击下载PDF文件
摘要:近年来,深度学习大大推进了音乐领域,巩固了音乐生成作为人工智能的关键应用。然而,现有的研究主要集中在西方音乐,在中国传统音乐的旋律生成方面遇到了挑战,特别是在模态特征和情感表达方面。为了解决这些问题,我们提出了一个新的体系结构,双特征建模模块,它集成了长期的依赖建模的曼巴块与全局结构捕捉能力的Transformer块。此外,我们还引入了双向Mamba融合层,该层通过双向扫描将局部细节和全局结构整合在一起,增强了复杂序列的建模能力。在此基础上,我们提出了REMI-M表示,它更准确地捕捉和生成旋律中的模态信息。为了支持这项研究,我们开发了FolkDB,这是一个高质量的中国传统音乐数据集,包含各种风格,总共超过11个小时的音乐。实验结果表明,该架构能够生成具有中国传统音乐特色的旋律,为音乐生成提供了一种新的有效解决方案。摘要:In recent years, deep learning has significantly advanced the MIDI domain, solidifying music generation as a key application of artificial intelligence. However, existing research primarily focuses on Western music and encounters challenges in generating melodies for Chinese traditional music, especially in capturing modal characteristics and emotional expression. To address these issues, we propose a new architecture, the Dual-Feature Modeling Module, which integrates the long-range dependency modeling of the Mamba Block with the global structure capturing capabilities of the Transformer Block. Additionally, we introduce the Bidirectional Mamba Fusion Layer, which integrates local details and global structures through bidirectional scanning, enhancing the modeling of complex sequences. Building on this architecture, we propose the REMI-M representation, which more accurately captures and generates modal information in melodies. To support this research, we developed FolkDB, a high-quality Chinese traditional music dataset encompassing various styles and totaling over 11 hours of music. Experimental results demonstrate that the proposed architecture excels in generating melodies with Chinese traditional music characteristics, offering a new and effective solution for music generation.
【8】 STAB: Speech Tokenizer Assessment Benchmark
标题: STAB:语音令牌器评估基准
作者:Shikhar Vashishth,Harman Singh,Shikhar Bharadwaj,Sriram Ganapathy,Chulayuth Asawaroengchai,Kartik Audhkhasi,Andrew Rosenberg,Ankur Bapna,Bhuvana Ramabhadran
备注:5 pages
链接:点击下载PDF文件
摘要:将语音表示为离散标记提供了一个框架,用于将语音转换为与文本非常相似的格式,从而使语音能够用作广泛成功的大型语言模型(LLM)的输入。目前,虽然已经提出了几种语音分词器,但对于特定下游任务的分词器所需的属性及其整体可推广性存在模糊性。跨不同下游任务评估标记器的性能是计算密集型工作,这对可扩展性提出了挑战。为了规避这一要求,我们提出了STAB(语音标记器评估基准),一个系统的评估框架,旨在全面评估语音标记器,并揭示其固有的特点。该框架提供了对语音标记化的底层机制的更深入理解,从而为加快未来标记器模型的发展提供了宝贵的资源,并使用标准化基准进行比较分析。我们评估STAB指标,并将其与一系列语音任务和标记器选择的下游任务性能相关联。摘要:Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the widely successful large language models (LLMs). Currently, while several speech tokenizers have been proposed, there is ambiguity regarding the properties that are desired from a tokenizer for specific downstream tasks and its overall generalizability. Evaluating the performance of tokenizers across different downstream tasks is a computationally intensive effort that poses challenges for scalability. To circumvent this requirement, we present STAB (Speech Tokenizer Assessment Benchmark), a systematic evaluation framework designed to assess speech tokenizers comprehensively and shed light on their inherent characteristics. This framework provides a deeper understanding of the underlying mechanisms of speech tokenization, thereby offering a valuable resource for expediting the advancement of future tokenizer models and enabling comparative analysis using a standardized benchmark. We evaluate the STAB metrics and correlate this with downstream task performance across a range of speech tasks and tokenizer choices.
【9】 LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement
标题: LSTMSE-Net:用于视听语音增强的长短期语音增强网络
作者:Arnav Jain,Jasmer Singh Sanjotra,Harshvardhan Choudhary,Krish Agrawal,Rupal Shah,Rohan Jha,M. Sajid,Amir Hussain,M. Tanveer
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:本文提出了一种基于长短期记忆的语音增强网络(LSTMSE-Net)的音视频语音增强方法。这种创新方法利用视觉和音频信息的互补性来提高语音信号的质量。视觉特征的提取与视觉神经网络(VFN),和音频特征的处理通过一个编码器和解码器。该系统缩放并连接视觉和音频特征,然后通过分离器网络进行处理,以优化语音增强。该架构突出了利用多模态数据和插值技术实现强大AVSE挑战系统的进步。LSTMSE-Net的性能超过了COG-MHEAR AVSE Challenge 2024的基线模型,其尺度不变信号失真比(SISDR)为0.06,短时客观可懂度(STOI)为0.03 $,语音质量感知评估(PESQ)为1.32 $。建议的LSTMSE-Net的源代码可在 url{https: github.com mtanveer1 AVSEC-3-Challenge}获得。摘要:In this paper, we propose long short term memory speech enhancement network (LSTMSE-Net), an audio-visual speech enhancement (AVSE) method. This innovative method leverages the complementary nature of visual and audio information to boost the quality of speech signals. Visual features are extracted with VisualFeatNet (VFN), and audio features are processed through an encoder and decoder. The system scales and concatenates visual and audio features, then processes them through a separator network for optimized speech enhancement. The architecture highlights advancements in leveraging multi-modal data and interpolation techniques for robust AVSE challenge systems. The performance of LSTMSE-Net surpasses that of the baseline model from the COG-MHEAR AVSE Challenge 2024 by a margin of 0.06 in scale-invariant signal-to-distortion ratio (SISDR), $0.03$ in short-time objective intelligibility (STOI), and $1.32$ in perceptual evaluation of speech quality (PESQ). The source code of the proposed LSTMSE-Net is available at url{https: github.com mtanveer1 AVSEC-3-Challenge}.
【10】 FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
标题: FastSecureGrad:基于扩散的一步语音转换,具有对抗性条件扩散蒸馏
作者:Takuhiro Kaneko,Hirokazu Kameoka,Kou Tanaka,Yuto Kondo
备注:Accepted to Interspeech 2024. Project page: this https URL
链接:点击下载PDF文件
摘要:基于扩散的语音转换(VC)技术,如VoiceGrad已经引起了人们的兴趣,因为他们的高VC性能的语音质量和扬声器的相似性。然而,一个值得注意的限制是由多步反向扩散引起的慢推理。因此,我们提出了FastVoiceGrad,一种新的一步扩散为基础的VC,减少了迭代次数从几十个到一个,同时继承了高VC性能的多步扩散为基础的VC。我们使用对抗条件扩散蒸馏(ACDD)获得模型,利用生成对抗网络和扩散模型的能力,同时重新考虑采样中的初始状态。对单次任意到任意VC的评估表明,FastVoiceGrad在提高推理速度的同时,实现了优于或可与以前的基于多步扩散的VC相媲美的VC性能。音频样本可在https: www.kecl.ntt.co.jp people kaneko.takuhiro projects fastvoicegrad 上获得。摘要:Diffusion-based voice conversion (VC) techniques such as VoiceGrad have attracted interest because of their high VC performance in terms of speech quality and speaker similarity. However, a notable limitation is the slow inference caused by the multi-step reverse diffusion. Therefore, we propose FastVoiceGrad, a novel one-step diffusion-based VC that reduces the number of iterations from dozens to one while inheriting the high VC performance of the multi-step diffusion-based VC. We obtain the model using adversarial conditional diffusion distillation (ACDD), leveraging the ability of generative adversarial networks and diffusion models while reconsidering the initial states in sampling. Evaluations of one-shot any-to-any VC demonstrate that FastVoiceGrad achieves VC performance superior to or comparable to that of previous multi-step diffusion-based VC while enhancing the inference speed. Audio samples are available at https: www.kecl.ntt.co.jp people kaneko.takuhiro projects fastvoicegrad .
【11】 Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR
标题: 基于时态保留的最优传输的ASB跨模式知识转移学习
作者:Xugang Lu,Peng Shen,Yu Tsao,Hisashi Kawai
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
摘要:将语言知识从预训练语言模型(PLM)转移到声学模型已被证明可以大大提高自动语音识别(ASR)的性能。然而,由于跨模态的异构特征分布,设计一个有效的模型,语言和声学序列之间的特征对齐和知识转移仍然是一个具有挑战性的任务。最佳传输(OT)可以有效地测量概率分布差异,在声学和语言模态之间对齐和传递知识方面具有巨大的潜力。然而,原始OT将声学和语言特征序列视为两个无序的对齐集合,并且在OT耦合估计过程中忽略了时间顺序信息。因此,需要一个耗时的预训练阶段来学习声学和语言表示之间的良好对齐。在本文中,我们提出了一个时间顺序保持OT(TOT)的跨模态对齐和知识转移(CAKT)(TOT-CAKT)的ASR。在TOT-CAKT中,声学序列的局部相邻帧平滑地映射到语言序列的相邻区域,在特征对齐和匹配中保持它们的时序关系。在TOT-CAKT模型框架下,我们使用预训练的中文PLM进行了普通话ASR实验,用于语言知识转移。我们的结果表明,与几种采用语言知识转移的最新模型相比,提出的TOT-CAKT显着提高了ASR性能,并解决了原始基于OT的方法在ASR顺序特征对齐中的弱点。摘要:Transferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR). However, due to the heterogeneous feature distributions in cross-modalities, designing an effective model for feature alignment and knowledge transfer between linguistic and acoustic sequences remains a challenging task. Optimal transport (OT), which efficiently measures probability distribution discrepancies, holds great potential for aligning and transferring knowledge between acoustic and linguistic modalities. Nonetheless, the original OT treats acoustic and linguistic feature sequences as two unordered sets in alignment and neglects temporal order information during OT coupling estimation. Consequently, a time-consuming pretraining stage is required to learn a good alignment between the acoustic and linguistic representations. In this paper, we propose a Temporal Order Preserved OT (TOT)-based Cross-modal Alignment and Knowledge Transfer (CAKT) (TOT-CAKT) for ASR. In the TOT-CAKT, local neighboring frames of acoustic sequences are smoothly mapped to neighboring regions of linguistic sequences, preserving their temporal order relationship in feature alignment and matching. With the TOT-CAKT model framework, we conduct Mandarin ASR experiments with a pretrained Chinese PLM for linguistic knowledge transfer. Our results demonstrate that the proposed TOT-CAKT significantly improves ASR performance compared to several state-of-the-art models employing linguistic knowledge transfer, and addresses the weaknesses of the original OT-based method in sequential feature alignment for ASR.
【12】 USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
标题: USEF-PSE:通用说话人嵌入自由目标说话人提取
作者:Bang Zeng,Ming Li
备注:13 pages, 6 figures
链接:点击下载PDF文件
摘要:目标说话人提取的目的是从混合语音中分离出特定说话人的声音。传统上,这个过程依赖于从参考语音中提取说话人嵌入,需要说话人识别模型。然而,识别合适的说话人识别模型可能具有挑战性,并且使用目标说话人嵌入作为参考信息可能对于目标说话人提取任务不是最佳的。本文介绍了一种通用的无说话人嵌入的目标说话人提取(USEF-TSE)框架,该框架不依赖于说话人嵌入。USEF-TSE利用多头交叉注意机制作为帧级目标说话人特征提取器。这种创新方法允许主流说话人提取解决方案绕过对说话人识别模型的依赖,并充分利用注册语音中可用的信息,包括说话人特征和上下文细节。此外,USEF-TSE可以与任何时域或时频域语音分离模型无缝集成,以实现有效的说话人提取。实验结果表明,我们提出的方法在WSJ 0 - 2 mix,WHAM!,和“威猛”!数据集,这是单声道消声,噪声和噪声混响两个扬声器语音分离和扬声器提取的标准基准。摘要:Target speaker extraction aims to isolate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a reference speech, necessitating a speaker recognition model. However, identifying an appropriate speaker recognition model can be challenging, and using the target speaker embedding as reference information may not be optimal for target speaker extraction tasks. This paper introduces a Universal Speaker Embedding-Free Target Speaker Extraction (USEF-TSE) framework that operates without relying on speaker embeddings. USEF-TSE utilizes a multi-head cross-attention mechanism as a frame-level target speaker feature extractor. This innovative approach allows mainstream speaker extraction solutions to bypass the dependency on speaker recognition models and to fully leverage the information available in the enrollment speech, including speaker characteristics and contextual details. Additionally, USEF-TSE can seamlessly integrate with any time-domain or time-frequency domain speech separation model to achieve effective speaker extraction. Experimental results show that our proposed method achieves state-of-the-art (SOTA) performance in terms of Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) on the WSJ0-2mix, WHAM!, and WHAMR! datasets, which are standard benchmarks for monaural anechoic, noisy and noisy-reverberant two-speaker speech separation and speaker extraction.
【13】 Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models
标题: 从自监督语音模型中高效提取噪音稳健的离散单元
作者:Jakob Poncelet,Yujun Wang,Hugo Van hamme
备注:Accepted at SLT2024
链接:点击下载PDF文件
摘要:通过从自监督学习(SSL)语音模型的隐藏特征中导出离散单元,可以将连续语音转换为离散序列。尽管SSL模型变得越来越大,并且在更多数据上进行了训练,但它们通常对实际失真(如加性噪声或混响)敏感,这会转化为离散单元的偏移。我们提出了一种参数有效的方法,通过训练一个小的编码器-解码器模型,使用或不使用适配器,同时去噪和离散化SSL模型的隐藏特征,从预先训练的SSL模型中生成噪声鲁棒的离散单元。该模型学习生成一个干净的离散序列的嘈杂的话语,条件的SSL功能。建议的去噪器在噪声离散化和噪声语音识别的任务上优于几种预训练方法,并且可以通过少量未标记目标数据的记录来微调目标环境。摘要:Continuous speech can be converted into a discrete sequence by deriving discrete units from the hidden features of self-supervised learned (SSL) speech models. Although SSL models are becoming larger and trained on more data, they are often sensitive to real-life distortions like additive noise or reverberation, which translates to a shift in discrete units. We propose a parameter-efficient approach to generate noise-robust discrete units from pre-trained SSL models by training a small encoder-decoder model, with or without adapters, to simultaneously denoise and discretise the hidden features of the SSL model. The model learns to generate a clean discrete sequence for a noisy utterance, conditioned on the SSL features. The proposed denoiser outperforms several pre-training methods on the tasks of noisy discretisation and noisy speech recognition, and can be finetuned to the target environment with a few recordings of unlabeled target data.
【14】 CUEMPATHY: A Counseling Speech Dataset for Psychotherapy Research
标题: CuEMPathsy:心理治疗研究的咨询言语数据集
作者:Dehua Tao,Harold Chui,Sarah Luk,Tan Lee
备注:Accepted by ISCSLP 2022
链接:点击下载PDF文件
摘要:心理治疗或咨询通常通过治疗师和客户之间的口头对话进行。分析心理治疗互动的言语特征有助于理解与有效心理治疗相关的因素。本文介绍了CUEMPATHY,一个大规模的语音数据集收集从实际咨询会议。该数据集包括156次咨询会议,涉及39个治疗师-客户二人组。语音数据收集,主观评级(一个观察员和两个客户端评级),和转录的过程中进行了描述。开发了一个自动语音和文本处理系统,用于定位每个会话中发言者回合的时间戳。检查三个主观评级之间的关系表明,观察员和客户评级没有显着的相关性,而客户评级的措施显着相关。治疗师和客户之间的强度相似性,测量的平均绝对差异的扬声器回合水平强度,与心理治疗的结果。介绍了CUEMPATHY的声学特征和语言学特征的最新研究成果。摘要:Psychotherapy or counseling is typically conducted through spoken conversation between a therapist and a client. Analyzing the speech characteristics of psychotherapeutic interactions can help understand the factors associated with effective psychotherapy. This paper introduces CUEMPATHY, a large-scale speech dataset collected from actual counseling sessions. The dataset consists of 156 counseling sessions involving 39 therapist-client dyads. The process of speech data collection, subjective ratings (one observer and two client ratings), and transcription are described. An automatic speech and text processing system is developed to locate the time stamps of speaker turns in each session. Examining the relationships among the three subjective ratings suggests that observer and client ratings have no significant correlation, while the client-rated measures are significantly correlated. The intensity similarity between the therapist and the client, measured by the averaged absolute difference of speaker-turn-level intensities, is associated with the psychotherapy outcomes. Recent studies on the acoustic and linguistic characteristics of the CUEMPATHY are introduced.
【15】 Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP
标题: 使用差异化DSP进行快速、高质量和参数高效的关节合成
作者:Yisi Liu,Bohan Yu,Drake Lin,Peter Wu,Cheol Jun Cho,Gopala Krishna Anumanchipalli
备注:accepted for Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
摘要:发音轨迹,如电磁发音记录(EMA)提供了声道滤波器的低维表示,并已被用作语音合成的自然,接地功能。可微分数字信号处理(DDSP)是一种参数高效的音频合成框架。因此,将低维EMA特征与DDSP相结合可以显著提高语音合成的计算效率。在本文中,我们提出了一种快速,高品质,和参数有效的DDSP发音声码器,可以合成语音EMA,F0和响度。我们结合了几种技术来解决谐波 噪声不平衡问题,并添加了多分辨率对抗性损失以获得更好的合成质量。我们的模型实现了6.67%的转录词错误率(WER)和3.74的平均意见得分(MOS),与最先进的(SOTA)基线相比,分别提高了1.63%和0.16。我们的DDSP声码器在推理过程中比CPU上的基线快4.9倍,并且仅用0.4M参数就可以生成质量相当的语音,而SOTA要求的参数为9 M。摘要:Articulatory trajectories like electromagnetic articulography (EMA) provide a low-dimensional representation of the vocal tract filter and have been used as natural, grounded features for speech synthesis. Differentiable digital signal processing (DDSP) is a parameter-efficient framework for audio synthesis. Therefore, integrating low-dimensional EMA features with DDSP can significantly enhance the computational efficiency of speech synthesis. In this paper, we propose a fast, high-quality, and parameter-efficient DDSP articulatory vocoder that can synthesize speech from EMA, F0, and loudness. We incorporate several techniques to solve the harmonics noise imbalance problem, and add a multi-resolution adversarial loss for better synthesis quality. Our model achieves a transcription word error rate (WER) of 6.67% and a mean opinion score (MOS) of 3.74, with an improvement of 1.63% and 0.16 compared to the state-of-the-art (SOTA) baseline. Our DDSP vocoder is 4.9x faster than the baseline on CPU during inference, and can generate speech of comparable quality with only 0.4M parameters, in contrast to the 9M parameters required by the SOTA.
【16】 Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
标题: 2024年受控歌唱语音Deepfake检测(CtrSDDD)挑战赛的语音基金会模型集合
作者:Anmol Guragain,Tianchi Liu,Zihan Pan,Hardik B. Sailor,Qiongqiong Wang
备注:Accepted to the IEEE Spoken Language Technology Workshop (SLT) 2024. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:这项工作详细介绍了我们的方法,以实现一个领先的系统,1.79%的合并等错误率(EER)的评估集的控制唱歌的声音Deepfake检测(CtrSVDD)。生成AI模型的快速发展为检测AI生成的Deepfake歌声带来了重大挑战,吸引了越来越多的研究关注。2024年歌唱声深度伪造检测(SVDD)挑战赛旨在解决这一复杂的任务。在这项工作中,我们探索集成方法,利用语音基础模型,开发强大的歌声反欺骗系统。我们还介绍了一种新的挤压和激励聚合(SEA)方法,该方法有效地集成了语音基础模型的表示特征,超越了我们其他单独系统的性能。评估结果证实了我们的方法在检测deepfake歌声方面的有效性。这些代码可以在https: github.com Anmol2059 SVDD2024上访问。摘要:This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generative AI models presents significant challenges for detecting AI-generated deepfake singing voices, attracting increased research attention. The Singing Voice Deepfake Detection (SVDD) Challenge 2024 aims to address this complex task. In this work, we explore the ensemble methods, utilizing speech foundation models to develop robust singing voice anti-spoofing systems. We also introduce a novel Squeeze-and-Excitation Aggregation (SEA) method, which efficiently and effectively integrates representation features from the speech foundation models, surpassing the performance of our other individual systems. Evaluation results confirm the efficacy of our approach in detecting deepfake singing voices. The codes can be accessed at https: github.com Anmol2059 SVDD2024.
eess.AS音频处理
【1】 USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction标题: USEF-PSE:通用说话人嵌入自由目标说话人提取
作者:Bang Zeng,Ming Li
备注:13 pages, 6 figures
链接:点击下载PDF文件
摘要:目标说话人提取的目的是从混合语音中分离出特定说话人的声音。传统上,这个过程依赖于从参考语音中提取说话人嵌入,需要说话人识别模型。然而,识别合适的说话人识别模型可能具有挑战性,并且使用目标说话人嵌入作为参考信息可能对于目标说话人提取任务不是最佳的。本文介绍了一种通用的无说话人嵌入的目标说话人提取(USEF-TSE)框架,该框架不依赖于说话人嵌入。USEF-TSE利用多头交叉注意机制作为帧级目标说话人特征提取器。这种创新方法允许主流说话人提取解决方案绕过对说话人识别模型的依赖,并充分利用注册语音中可用的信息,包括说话人特征和上下文细节。此外,USEF-TSE可以与任何时域或时频域语音分离模型无缝集成,以实现有效的说话人提取。实验结果表明,我们提出的方法在WSJ 0 - 2 mix,WHAM!,和“威猛”!数据集,这是单声道消声,噪声和噪声混响两个扬声器语音分离和扬声器提取的标准基准。摘要:Target speaker extraction aims to isolate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a reference speech, necessitating a speaker recognition model. However, identifying an appropriate speaker recognition model can be challenging, and using the target speaker embedding as reference information may not be optimal for target speaker extraction tasks. This paper introduces a Universal Speaker Embedding-Free Target Speaker Extraction (USEF-TSE) framework that operates without relying on speaker embeddings. USEF-TSE utilizes a multi-head cross-attention mechanism as a frame-level target speaker feature extractor. This innovative approach allows mainstream speaker extraction solutions to bypass the dependency on speaker recognition models and to fully leverage the information available in the enrollment speech, including speaker characteristics and contextual details. Additionally, USEF-TSE can seamlessly integrate with any time-domain or time-frequency domain speech separation model to achieve effective speaker extraction. Experimental results show that our proposed method achieves state-of-the-art (SOTA) performance in terms of Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) on the WSJ0-2mix, WHAM!, and WHAMR! datasets, which are standard benchmarks for monaural anechoic, noisy and noisy-reverberant two-speaker speech separation and speaker extraction.
【2】 Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models
标题: 从自监督语音模型中高效提取噪音稳健的离散单元
作者:Jakob Poncelet,Yujun Wang,Hugo Van hamme
备注:Accepted at SLT2024
链接:点击下载PDF文件
摘要:通过从自监督学习(SSL)语音模型的隐藏特征中导出离散单元,可以将连续语音转换为离散序列。尽管SSL模型变得越来越大,并且在更多数据上进行了训练,但它们通常对实际失真(如加性噪声或混响)敏感,这会转化为离散单元的偏移。我们提出了一种参数有效的方法,通过训练一个小的编码器-解码器模型,使用或不使用适配器,同时去噪和离散化SSL模型的隐藏特征,从预先训练的SSL模型生成噪声鲁棒的离散单元。该模型学习生成一个干净的离散序列的嘈杂的话语,条件的SSL功能。建议的去噪器在噪声离散化和噪声语音识别的任务上优于几种预训练方法,并且可以通过少量未标记目标数据的记录来微调目标环境。摘要:Continuous speech can be converted into a discrete sequence by deriving discrete units from the hidden features of self-supervised learned (SSL) speech models. Although SSL models are becoming larger and trained on more data, they are often sensitive to real-life distortions like additive noise or reverberation, which translates to a shift in discrete units. We propose a parameter-efficient approach to generate noise-robust discrete units from pre-trained SSL models by training a small encoder-decoder model, with or without adapters, to simultaneously denoise and discretise the hidden features of the SSL model. The model learns to generate a clean discrete sequence for a noisy utterance, conditioned on the SSL features. The proposed denoiser outperforms several pre-training methods on the tasks of noisy discretisation and noisy speech recognition, and can be finetuned to the target environment with a few recordings of unlabeled target data.
【3】 CUEMPATHY: A Counseling Speech Dataset for Psychotherapy Research
标题: CuEMPathsy:心理治疗研究的咨询言语数据集
作者:Dehua Tao,Harold Chui,Sarah Luk,Tan Lee
备注:Accepted by ISCSLP 2022
链接:点击下载PDF文件
摘要:心理治疗或咨询通常通过治疗师和客户之间的口头对话进行。分析心理治疗互动的言语特征有助于理解与有效心理治疗相关的因素。本文介绍了CUEMPATHY,一个大规模的语音数据集收集从实际咨询会议。该数据集包括156次咨询会议,涉及39个治疗师-客户二人组。语音数据收集,主观评级(一个观察员和两个客户端评级),和转录的过程中进行了描述。开发了一个自动语音和文本处理系统,用于定位每个会话中发言者回合的时间戳。检查三个主观评级之间的关系表明,观察员和客户评级没有显着的相关性,而客户评级的措施显着相关。治疗师和客户之间的强度相似性,测量的平均绝对差异的扬声器回合水平强度,与心理治疗的结果。介绍了CUEMPATHY的声学特征和语言学特征的最新研究成果。摘要:Psychotherapy or counseling is typically conducted through spoken conversation between a therapist and a client. Analyzing the speech characteristics of psychotherapeutic interactions can help understand the factors associated with effective psychotherapy. This paper introduces CUEMPATHY, a large-scale speech dataset collected from actual counseling sessions. The dataset consists of 156 counseling sessions involving 39 therapist-client dyads. The process of speech data collection, subjective ratings (one observer and two client ratings), and transcription are described. An automatic speech and text processing system is developed to locate the time stamps of speaker turns in each session. Examining the relationships among the three subjective ratings suggests that observer and client ratings have no significant correlation, while the client-rated measures are significantly correlated. The intensity similarity between the therapist and the client, measured by the averaged absolute difference of speaker-turn-level intensities, is associated with the psychotherapy outcomes. Recent studies on the acoustic and linguistic characteristics of the CUEMPATHY are introduced.
【4】 Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP
标题: 使用差异化DSP进行快速、高质量和参数高效的关节合成
作者:Yisi Liu,Bohan Yu,Drake Lin,Peter Wu,Cheol Jun Cho,Gopala Krishna Anumanchipalli
备注:accepted for Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
摘要:发音轨迹,如电磁发音记录(EMA)提供了声道滤波器的低维表示,并已被用作语音合成的自然,接地功能。可微分数字信号处理(DDSP)是一种参数高效的音频合成框架。因此,将低维EMA特征与DDSP相结合可以显著提高语音合成的计算效率。在本文中,我们提出了一种快速,高品质,和参数有效的DDSP发音声码器,可以合成语音EMA,F0和响度。我们结合了几种技术来解决谐波 噪声不平衡问题,并添加了多分辨率对抗性损失以获得更好的合成质量。我们的模型实现了6.67%的转录词错误率(WER)和3.74的平均意见得分(MOS),与最先进的(SOTA)基线相比,分别提高了1.63%和0.16。我们的DDSP声码器在推理过程中比CPU上的基线快4.9倍,并且仅用0.4M参数就可以生成质量相当的语音,而SOTA要求的参数为9 M。摘要:Articulatory trajectories like electromagnetic articulography (EMA) provide a low-dimensional representation of the vocal tract filter and have been used as natural, grounded features for speech synthesis. Differentiable digital signal processing (DDSP) is a parameter-efficient framework for audio synthesis. Therefore, integrating low-dimensional EMA features with DDSP can significantly enhance the computational efficiency of speech synthesis. In this paper, we propose a fast, high-quality, and parameter-efficient DDSP articulatory vocoder that can synthesize speech from EMA, F0, and loudness. We incorporate several techniques to solve the harmonics noise imbalance problem, and add a multi-resolution adversarial loss for better synthesis quality. Our model achieves a transcription word error rate (WER) of 6.67% and a mean opinion score (MOS) of 3.74, with an improvement of 1.63% and 0.16 compared to the state-of-the-art (SOTA) baseline. Our DDSP vocoder is 4.9x faster than the baseline on CPU during inference, and can generate speech of comparable quality with only 0.4M parameters, in contrast to the 9M parameters required by the SOTA.
【5】 Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
标题: 2024年受控歌唱语音Deepfake检测(CtrSDDD)挑战赛的语音基金会模型集合
作者:Anmol Guragain,Tianchi Liu,Zihan Pan,Hardik B. Sailor,Qiongqiong Wang
备注:Accepted to the IEEE Spoken Language Technology Workshop (SLT) 2024. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:这项工作详细介绍了我们的方法,以实现一个领先的系统,1.79%的合并等错误率(EER)的评估集的控制唱歌的声音Deepfake检测(CtrSVDD)。生成AI模型的快速发展为检测AI生成的deepfake歌声带来了重大挑战,吸引了越来越多的研究关注。2024年歌唱声深度伪造检测(SVDD)挑战赛旨在解决这一复杂的任务。在这项工作中,我们探索集成方法,利用语音基础模型,开发强大的歌声反欺骗系统。我们还介绍了一种新的挤压和激励聚合(SEA)方法,该方法有效地集成了语音基础模型的表示特征,超越了我们其他单独系统的性能。评估结果证实了我们的方法在检测deepfake歌声方面的有效性。这些代码可以在https: github.com Anmol2059 SVDD2024上访问。摘要:This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generative AI models presents significant challenges for detecting AI-generated deepfake singing voices, attracting increased research attention. The Singing Voice Deepfake Detection (SVDD) Challenge 2024 aims to address this complex task. In this work, we explore the ensemble methods, utilizing speech foundation models to develop robust singing voice anti-spoofing systems. We also introduce a novel Squeeze-and-Excitation Aggregation (SEA) method, which efficiently and effectively integrates representation features from the speech foundation models, surpassing the performance of our other individual systems. Evaluation results confirm the efficacy of our approach in detecting deepfake singing voices. The codes can be accessed at https: github.com Anmol2059 SVDD2024.
【6】 Latent Watermarking of Audio Generative Models
标题: 音频生成模型的潜在水印
作者:Robin San Roman,Pierre Fernandez,Antoine Deleforge,Yossi Adi,Romain Serizel
链接:点击下载PDF文件
摘要:音频生成模型的进步在其负责任的披露和检测其滥用方面带来了新的挑战。作为回应,我们介绍了一种方法,水印潜在的生成模型通过一个特定的水印的训练数据。所得到的水印模型产生潜在的表示,其解码输出被检测到具有高的信心,无论使用的解码方法。这种方法使得能够检测所生成的内容,而不需要事后水印步骤。它为开源模型提供了一个更安全的解决方案,并有助于识别那些在不遵守其许可条款的情况下微调或使用这些模型的衍生作品。我们的研究结果表明,例如,所产生的输出检测的准确性超过75%,误报率为10 ^{-3}$,即使微调后的潜在生成模型。摘要:The advancements in audio generative models have opened up new challenges in their responsible disclosure and the detection of their misuse. In response, we introduce a method to watermark latent generative models by a specific watermarking of their training data. The resulting watermarked models produce latent representations whose decoded outputs are detected with high confidence, regardless of the decoding method used. This approach enables the detection of the generated content without the need for a post-hoc watermarking step. It provides a more secure solution for open-sourced models and facilitates the identification of derivative works that fine-tune or use these models without adhering to their license terms. Our results indicate for instance that generated outputs are detected with an accuracy of more than 75% at a false positive rate of $10^{-3}$, even after fine-tuning the latent generative model.
【7】 Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
标题: 多轨音乐LDM:利用潜在扩散模型实现多功能音乐生成
作者:Tornike Karchkhadze,Mohammad Rasool Izadi,Ke Chen,Gerard Assayag,Shlomo Dubnov
链接:点击下载PDF文件
摘要:扩散模型在涉及音频和音乐的跨模态生成任务中显示出有希望的结果,例如文本到声音和文本到音乐生成。这些文本控制的音乐生成模型通常专注于通过捕获诸如流派和情绪之类的全局音乐属性来生成音乐。然而,音乐创作是一个复杂的,多层次的任务,往往涉及音乐安排作为一个不可分割的一部分的过程。这个过程涉及到组合每种乐器,使其在节拍、力度、和声和旋律方面与现有乐器保持一致,这需要比文本提示通常提供的更高的精度和对曲目的控制。在这项工作中,我们通过扩展MusicLDM,一个潜在的音乐扩散模型,到一个多轨生成模型来解决这些挑战。通过学习共享上下文的曲目的联合概率,我们的模型能够有条件或无条件地在多个相互对应的曲目中生成音乐。此外,我们的模型能够生成排列,其中模型可以生成给定其他轨道的任何轨道子集(例如,产生补充给定低音和鼓音轨的钢琴音轨)。我们将我们的模型与现有的多轨道生成模型进行了比较,并证明了我们的模型在总体和安排生成任务的客观指标上都取得了相当大的改进。摘要:Diffusion models have shown promising results in cross-modal generation tasks involving audio and music, such as text-to-sound and text-to-music generation. These text-controlled music generation models typically focus on generating music by capturing global musical attributes like genre and mood. However, music composition is a complex, multilayered task that often involves musical arrangement as an integral part of the process. This process involves composing each instrument to align with existing ones in terms of beat, dynamics, harmony, and melody, requiring greater precision and control over tracks than text prompts usually provide. In this work, we address these challenges by extending the MusicLDM, a latent diffusion model for music, into a multi-track generative model. By learning the joint probability of tracks sharing a context, our model is capable of generating music across several tracks that correspond well to each other, either conditionally or unconditionally. Additionally, our model is capable of arrangement generation, where the model can generate any subset of tracks given the others (e.g., generating a piano track complementing given bass and drum tracks). We compared our model with an existing multi-track generative model and demonstrated that our model achieves considerable improvements across objective metrics for both total and arrangement generation tasks.
【8】 Effects of Recording Condition and Number of Monitored Days on Discriminative Power of the Daily Phonotrauma Index
标题: 记录条件和监测天数对每日语音创伤指数区分力的影响
作者:Hamzeh Ghasemzadeh,Robert E. Hillman,Jarrad H. Van Stan,Daryush D. Mehta
备注:The paper is submitted to JSLHR
链接:点击下载PDF文件
摘要:目的:每日语音创伤指数(DPI)可以量化与语音创伤性发声功能亢进(PVH)患者日常发声相关的病理生理机制。由于DPI是基于为期一周的动态语音监测开发的,因此本研究调查了DPI是否可以使用(1)短实验室语音任务和(2)少于7天的动态数据实现相当的性能。方法:动态语音监测系统记录了134名女性PVH和声音健康匹配的对照组在两种不同的条件下的声音功能 行为。在实验室中,参与者阅读彩虹通道的第一段,并产生自发的演讲(实验室数据)。然后对它们进行了7天的监测(现场数据)。使用前两个谐波(H1-H2)的幅度之间的差异的标准差和颈部表面加速度幅度的偏斜度,从实验室和现场数据中训练单独的DPI模型。首先,10倍交叉验证评估了实验室和现场DPI的分类性能。其次,量化了门诊监测天数对现场DPI分类准确性的影响。结果如下:从Rainbow通道和自发语音计算的平均实验室DPI准确度分别为57.9%和48.9%,接近偶然性能。在现场DPI的平均分类精度显着较高,具有非常大的效应大小(73.4%,科恩D = 1.8)。第二,实地DPI的平均准确度从一天的66.5%提高到七天的75.0%,增加一天的准确度在4天后下降到不到1个百分点。摘要:Objective: The Daily Phonotrauma Index (DPI) can quantify pathophysiological mechanisms associated with daily voice use in individuals with phonotraumatic vocal hyperfunction (PVH). Since DPI was developed based on week-long ambulatory voice monitoring, this study investigated if DPI can achieve comparable performance using (1) short laboratory speech tasks and (2) fewer than seven days of ambulatory data. Method: An ambulatory voice monitoring system recorded the vocal function behavior of 134 females with PVH and vocally healthy matched controls in two different conditions. In the lab, the participants read the first paragraph of the Rainbow Passage and produced spontaneous speech (in-lab data). They were then monitored for seven days (in-field data). Separate DPI models were trained from in-lab and in-field data using the standard deviation of the difference between the magnitude of the first two harmonics (H1-H2) and the skewness of neck-surface acceleration magnitude. First, 10-fold cross-validation evaluated classification performance of in-lab and in-field DPIs. Second, the effect of the number of ambulatory monitoring days on the accuracy of in-field DPI classification was quantified. Results: The average in-lab DPI accuracy computed from the Rainbow passage and spontaneous speech were, respectively, 57.9% and 48.9%, which are close to chance performance. The average classification accuracy of in-field DPI was significantly higher with a very large effect size (73.4%, Cohens D = 1.8). Second, the average in-field DPI accuracy increased from 66.5% for one day to 75.0% for seven days, with the gain of including an additional day on accuracy dropping below 1 percentage point after 4 days.
【9】 An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
标题: 用BEST-PQ分析线性复杂性注意替代
作者:Ryan Whetten,Titouan Parcollet,Adel Moumen,Marco Dinarelli,Yannick Estève
备注:Accepted in the IEEE Soken Language Technology Workshop 2024
链接:点击下载PDF文件
摘要:自监督学习(SSL)已被证明在包括语音处理在内的各个领域都是有效的。然而,SSL在计算和内存上是昂贵的。这部分是由于多头自注意(MHSA)的二次复杂性。MHSA的替代方案已被提出并用于语音领域,但尚未在SSL设置中进行适当的研究。在这项工作中,我们研究了用最新的具有线性复杂性的替代品(即HyperMixing,Fastformer,SummaryMixing和Mamba)取代MHSA的影响。我们通过查看速度、消耗的VRAM量以及SSL MP3S基准测试的性能来评估这些方法。结果表明,与MHSA相比,这些线性替代方案保持了具有竞争力的性能,同时平均将VRAM消耗降低了约20%至60%,并将输入序列从20秒增加到80秒的速度从7%增加到65%。摘要:Self-Supervised Learning (SSL) has proven to be effective in various domains, including speech processing. However, SSL is computationally and memory expensive. This is in part due the quadratic complexity of multi-head self-attention (MHSA). Alternatives for MHSA have been proposed and used in the speech domain, but have yet to be investigated properly in an SSL setting. In this work, we study the effects of replacing MHSA with recent state-of-the-art alternatives that have linear complexity, namely, HyperMixing, Fastformer, SummaryMixing, and Mamba. We evaluate these methods by looking at the speed, the amount of VRAM consumed, and the performance on the SSL MP3S benchmark. Results show that these linear alternatives maintain competitive performance compared to MHSA while, on average, decreasing VRAM consumption by around 20% to 60% and increasing speed from 7% to 65% for input sequences ranging from 20 to 80 seconds.
【10】 Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems
标题: 使用基于特征平滑的增强方法训练通用声码器,用于高质量的TTC系统
作者:Jeongmin Liu,Eunwoo Song
备注:4 pages, 4 figures, for demo samples, see this https URL
链接:点击下载PDF文件
摘要:虽然通用声码器已经在不同的语音中实现了熟练的波形生成,但它们集成到文本到语音(TTS)任务中通常会导致合成质量下降。为了应对这一挑战,我们提出了一种用于训练通用声码器的新型增强技术。我们的训练方案随机地将线性平滑滤波器应用于输入声学特征,从而促进声码器在广泛的平滑范围内的泛化。它显著地减轻了训练-推理不匹配,增强了合成输出的自然性,即使声学模型产生过度平滑的特征。值得注意的是,我们的方法是适用于任何声码器,而不需要架构修改或依赖于特定的声学模型。实验结果验证了我们的声码器优于传统的方法,实现11.99%和12.05%的平均意见分数时,与Tacotron 2和FastSpeech 2 TTS声学模型集成,分别提高。摘要:While universal vocoders have achieved proficient waveform generation across diverse voices, their integration into text-to-speech (TTS) tasks often results in degraded synthetic quality. To address this challenge, we present a novel augmentation technique for training universal vocoders. Our training scheme randomly applies linear smoothing filters to input acoustic features, facilitating vocoder generalization across a wide range of smoothings. It significantly mitigates the training-inference mismatch, enhancing the naturalness of synthetic output even when the acoustic model produces overly smoothed features. Notably, our method is applicable to any vocoder without requiring architectural modifications or dependencies on specific acoustic models. The experimental results validate the superiority of our vocoder over conventional methods, achieving 11.99% and 12.05% improvements in mean opinion scores when integrated with Tacotron 2 and FastSpeech 2 TTS acoustic models, respectively.
【11】 NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention
标题: NeuroSpex:具有跨模式注意力的神经引导说话者提取
作者:Dashanka De Silva,Siqi Cai,Saurav Pahuja,Tanja Schultz,Haizhou Li
链接:点击下载PDF文件
摘要:在听觉注意的研究中,已经发现在注意言语和诱发的神经反应之间存在鲁棒的相关性,其可以通过脑电图(EEG)来测量。因此,可以使用脑电信号中可用的注意力信息来通过计算指导鸡尾酒会中目标说话者的提取。在本文中,我们提出了一个神经引导的说话人提取模型,即NeuroSpex,使用的唯一辅助参考线索,从单耳语音混合的听众的EEG响应提取出席的语音。我们提出了一种新的脑电信号编码器,捕捉的注意力信息。此外,我们提出了一个交叉注意(CA)机制,以提高语音特征表示,生成一个扬声器提取掩码。在公开数据集上的实验结果表明,我们提出的模型在各种评估指标上优于两个基线模型。摘要:In the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is possible to use the attention information available within EEG signals to guide the extraction of the target speaker in a cocktail party computationally. In this paper, we present a neuro-guided speaker extraction model, i.e. NeuroSpex, using the EEG response of the listener as the sole auxiliary reference cue to extract attended speech from monaural speech mixtures. We propose a novel EEG signal encoder that captures the attention information. Additionally, we propose a cross-attention (CA) mechanism to enhance the speech feature representations, generating a speaker extraction mask. Experimental results on a publicly available dataset demonstrate that our proposed model outperforms two baseline models across various evaluation metrics.
【12】 MusicMamba: A Dual-Feature Modeling Approach for Generating Chinese Traditional Music with Modal Precision
标题: MusicMamba:一种具有模式精确度的中国传统音乐的双特征建模方法
作者:Jiatao Chen,Tianming Xie,Xing Tang,Jing Wang,Wenjing Dong,Bing Shi
链接:点击下载PDF文件
摘要:近年来,深度学习大大推进了音乐领域,巩固了音乐生成作为人工智能的关键应用。然而,现有的研究主要集中在西方音乐,在中国传统音乐的旋律生成方面遇到了挑战,特别是在模态特征和情感表达方面。为了解决这些问题,我们提出了一个新的体系结构,双特征建模模块,它集成了长期的依赖建模的曼巴块与全局结构捕捉能力的Transformer块。此外,我们还引入了双向Mamba融合层,该层通过双向扫描将局部细节和全局结构整合在一起,增强了复杂序列的建模能力。在此基础上,我们提出了REMI-M表示,它更准确地捕捉和生成旋律中的模态信息。为了支持这项研究,我们开发了FolkDB,这是一个高质量的中国传统音乐数据集,包含各种风格,总共超过11个小时的音乐。实验结果表明,该架构能够生成具有中国传统音乐特色的旋律,为音乐生成提供了一种新的有效解决方案。摘要:In recent years, deep learning has significantly advanced the MIDI domain, solidifying music generation as a key application of artificial intelligence. However, existing research primarily focuses on Western music and encounters challenges in generating melodies for Chinese traditional music, especially in capturing modal characteristics and emotional expression. To address these issues, we propose a new architecture, the Dual-Feature Modeling Module, which integrates the long-range dependency modeling of the Mamba Block with the global structure capturing capabilities of the Transformer Block. Additionally, we introduce the Bidirectional Mamba Fusion Layer, which integrates local details and global structures through bidirectional scanning, enhancing the modeling of complex sequences. Building on this architecture, we propose the REMI-M representation, which more accurately captures and generates modal information in melodies. To support this research, we developed FolkDB, a high-quality Chinese traditional music dataset encompassing various styles and totaling over 11 hours of music. Experimental results demonstrate that the proposed architecture excels in generating melodies with Chinese traditional music characteristics, offering a new and effective solution for music generation.
【13】 STAB: Speech Tokenizer Assessment Benchmark
标题: STAB:语音令牌器评估基准
作者:Shikhar Vashishth,Harman Singh,Shikhar Bharadwaj,Sriram Ganapathy,Chulayuth Asawaroengchai,Kartik Audhkhasi,Andrew Rosenberg,Ankur Bapna,Bhuvana Ramabhadran
备注:5 pages
链接:点击下载PDF文件
摘要:将语音表示为离散标记提供了一个框架,用于将语音转换为与文本非常相似的格式,从而使语音能够用作广泛成功的大型语言模型(LLM)的输入。目前,虽然已经提出了几种语音分词器,但对于特定下游任务的分词器所需的属性及其整体可推广性存在模糊性。跨不同下游任务评估标记器的性能是计算密集型工作,这对可扩展性提出了挑战。为了规避这一要求,我们提出了STAB(语音标记器评估基准),一个系统的评估框架,旨在全面评估语音标记器,并揭示其固有的特点。该框架提供了对语音标记化的底层机制的更深入理解,从而为加快未来标记器模型的发展提供了宝贵的资源,并使用标准化基准进行比较分析。我们评估STAB指标,并将其与一系列语音任务和标记器选择的下游任务性能相关联。摘要:Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the widely successful large language models (LLMs). Currently, while several speech tokenizers have been proposed, there is ambiguity regarding the properties that are desired from a tokenizer for specific downstream tasks and its overall generalizability. Evaluating the performance of tokenizers across different downstream tasks is a computationally intensive effort that poses challenges for scalability. To circumvent this requirement, we present STAB (Speech Tokenizer Assessment Benchmark), a systematic evaluation framework designed to assess speech tokenizers comprehensively and shed light on their inherent characteristics. This framework provides a deeper understanding of the underlying mechanisms of speech tokenization, thereby offering a valuable resource for expediting the advancement of future tokenizer models and enabling comparative analysis using a standardized benchmark. We evaluate the STAB metrics and correlate this with downstream task performance across a range of speech tasks and tokenizer choices.
【14】 LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement
标题: LSTMSE-Net:用于视听语音增强的长短期语音增强网络
作者:Arnav Jain,Jasmer Singh Sanjotra,Harshvardhan Choudhary,Krish Agrawal,Rupal Shah,Rohan Jha,M. Sajid,Amir Hussain,M. Tanveer
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:本文提出了一种基于长短期记忆的语音增强网络(LSTMSE-Net)的音视频语音增强方法。这种创新方法利用视觉和音频信息的互补性来提高语音信号的质量。视觉特征的提取与视觉神经网络(VFN),和音频特征的处理通过一个编码器和解码器。该系统缩放并连接视觉和音频特征,然后通过分离器网络进行处理,以优化语音增强。该架构突出了利用多模态数据和插值技术实现强大AVSE挑战系统的进步。LSTMSE-Net的性能超过了COG-MHEAR AVSE Challenge 2024的基线模型,其尺度不变信号失真比(SISDR)为0.06,短时客观可懂度(STOI)为0.03 $,语音质量感知评估(PESQ)为1.32 $。建议的LSTMSE-Net的源代码可在 url{https: github.com mtanveer1 AVSEC-3-Challenge}获得。摘要:In this paper, we propose long short term memory speech enhancement network (LSTMSE-Net), an audio-visual speech enhancement (AVSE) method. This innovative method leverages the complementary nature of visual and audio information to boost the quality of speech signals. Visual features are extracted with VisualFeatNet (VFN), and audio features are processed through an encoder and decoder. The system scales and concatenates visual and audio features, then processes them through a separator network for optimized speech enhancement. The architecture highlights advancements in leveraging multi-modal data and interpolation techniques for robust AVSE challenge systems. The performance of LSTMSE-Net surpasses that of the baseline model from the COG-MHEAR AVSE Challenge 2024 by a margin of 0.06 in scale-invariant signal-to-distortion ratio (SISDR), $0.03$ in short-time objective intelligibility (STOI), and $1.32$ in perceptual evaluation of speech quality (PESQ). The source code of the proposed LSTMSE-Net is available at url{https: github.com mtanveer1 AVSEC-3-Challenge}.
【15】 FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
标题: FastSecureGrad:基于扩散的一步语音转换,具有对抗性条件扩散蒸馏
作者:Takuhiro Kaneko,Hirokazu Kameoka,Kou Tanaka,Yuto Kondo
备注:Accepted to Interspeech 2024. Project page: this https URL
链接:点击下载PDF文件
摘要:基于扩散的语音转换(VC)技术,如VoiceGrad已经引起了人们的兴趣,因为他们的高VC性能的语音质量和扬声器的相似性。然而,一个值得注意的限制是由多步反向扩散引起的慢推理。因此,我们提出了FastVoiceGrad,一种新的一步扩散为基础的VC,减少了迭代次数从几十个到一个,同时继承了高VC性能的多步扩散为基础的VC。我们使用对抗条件扩散蒸馏(ACDD)获得模型,利用生成对抗网络和扩散模型的能力,同时重新考虑采样中的初始状态。对单次任意到任意VC的评估表明,FastVoiceGrad在提高推理速度的同时,实现了优于或可与以前的基于多步扩散的VC相媲美的VC性能。音频样本可在https: www.kecl.ntt.co.jp people kaneko.takuhiro projects fastvoicegrad 上获得。摘要:Diffusion-based voice conversion (VC) techniques such as VoiceGrad have attracted interest because of their high VC performance in terms of speech quality and speaker similarity. However, a notable limitation is the slow inference caused by the multi-step reverse diffusion. Therefore, we propose FastVoiceGrad, a novel one-step diffusion-based VC that reduces the number of iterations from dozens to one while inheriting the high VC performance of the multi-step diffusion-based VC. We obtain the model using adversarial conditional diffusion distillation (ACDD), leveraging the ability of generative adversarial networks and diffusion models while reconsidering the initial states in sampling. Evaluations of one-shot any-to-any VC demonstrate that FastVoiceGrad achieves VC performance superior to or comparable to that of previous multi-step diffusion-based VC while enhancing the inference speed. Audio samples are available at https: www.kecl.ntt.co.jp people kaneko.takuhiro projects fastvoicegrad .
【16】 Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR
标题: 基于时态保留的最优传输的ASB跨模式知识转移学习
作者:Xugang Lu,Peng Shen,Yu Tsao,Hisashi Kawai
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
摘要:将语言知识从预训练语言模型(PLM)转移到声学模型已被证明可以大大提高自动语音识别(ASR)的性能。然而,由于跨模态的异构特征分布,设计一个有效的模型,语言和声学序列之间的特征对齐和知识转移仍然是一个具有挑战性的任务。最优传输(OT),有效地衡量概率分布的差异,具有很大的潜力,在声学和语言模态之间的对齐和转移知识。然而,原始OT将声学和语言特征序列视为两个无序的对齐集合,并且在OT耦合估计过程中忽略了时间顺序信息。因此,需要一个耗时的预训练阶段来学习声学和语言表示之间的良好对齐。在本文中,我们提出了一个时间顺序保持OT(TOT)的跨模态对齐和知识转移(CAKT)(TOT-CAKT)的ASR。在TOT-CAKT中,声学序列的局部相邻帧平滑地映射到语言序列的相邻区域,在特征对齐和匹配中保持它们的时序关系。在TOT-CAKT模型框架下,我们使用预训练的中文PLM进行了普通话ASR实验,用于语言知识转移。我们的研究结果表明,提出的TOT-CAKT显着提高ASR性能相比,几个国家的最先进的模型,采用语言知识转移,并解决了原来的OT为基础的方法在顺序特征对齐ASR的弱点。摘要:Transferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR). However, due to the heterogeneous feature distributions in cross-modalities, designing an effective model for feature alignment and knowledge transfer between linguistic and acoustic sequences remains a challenging task. Optimal transport (OT), which efficiently measures probability distribution discrepancies, holds great potential for aligning and transferring knowledge between acoustic and linguistic modalities. Nonetheless, the original OT treats acoustic and linguistic feature sequences as two unordered sets in alignment and neglects temporal order information during OT coupling estimation. Consequently, a time-consuming pretraining stage is required to learn a good alignment between the acoustic and linguistic representations. In this paper, we propose a Temporal Order Preserved OT (TOT)-based Cross-modal Alignment and Knowledge Transfer (CAKT) (TOT-CAKT) for ASR. In the TOT-CAKT, local neighboring frames of acoustic sequences are smoothly mapped to neighboring regions of linguistic sequences, preserving their temporal order relationship in feature alignment and matching. With the TOT-CAKT model framework, we conduct Mandarin ASR experiments with a pretrained Chinese PLM for linguistic knowledge transfer. Our results demonstrate that the proposed TOT-CAKT significantly improves ASR performance compared to several state-of-the-art models employing linguistic knowledge transfer, and addresses the weaknesses of the original OT-based method in sequential feature alignment for ASR.
机器翻译,仅供参考
