今日论文合集:cs.SD语音6篇,eess.AS音频处理7篇。本文经arXiv每日学术速递授权转载
【1】Real-time Speech Extraction Using Spatially Regularized Independent Low-rank Matrix Analysis and Rank-constrained Spatial Covariance Matrix Estimation
标题:基于空间正则化独立低秩阵分析和秩受限空间协方差矩阵估计的实时语音提取
链接:https://arxiv.org/abs/2403.12477
作者:Yuto Ishikawa,Kohei Konaka,Tomohiko Nakamura,Norihiro Takamune,Hiroshi Saruwatari备注:5 pages, 3 figures, accepted at HSCMA 2024
摘要:实时语音提取是各种应用的一个重要挑战,例如类人化身/机器人中的语音识别。本文提出了一种基于独立低秩矩阵分析(ILRMA)和秩约束空间协方差矩阵估计(RCSCME)的语音提取方法的实时扩展。基于RCSCME的方法是一种多通道盲语音提取方法,其在扩散噪声环境中表现出优异的语音提取性能。为了提高性能,我们将空间正则化引入到基于RCSCME的语音提取的ILRMA部分,并设计了两个正则化器。语音提取实验表明,所提出的方法可以实时工作,所设计的正则化器提高了语音提取性能。
摘要:Real-time speech extraction is an important challenge with various applications such as speech recognition in a human-like avatar/robot. In this paper, we propose the real-time extension of a speech extraction method based on independent low-rank matrix analysis (ILRMA) and rank-constrained spatial covariance matrix estimation (RCSCME). The RCSCME-based method is a multichannel blind speech extraction method that demonstrates superior speech extraction performance in diffuse noise environments. To improve the performance, we introduce spatial regularization into the ILRMA part of the RCSCME-based speech extraction and design two regularizers. Speech extraction experiments demonstrated that the proposed methods can function in real time and the designed regularizers improve the speech extraction performance.
【2】 Multimodal Fusion Method with Spatiotemporal Sequences and Relationship Learning for Valence-Arousal Estimation标题:基于时空序列和关系学习的多模式融合价态觉醒估计方法作者:Jun Yu,Gongpeng Zhao,Yongqi Wan,Zhihong Wei,Yang Zheng,Zerui Zhang,Zhongpeng Cai,Guochen Xie,Jichao Zhu,Wangyuan Zhu摘要:本文介绍了我们的方法在ABAW6比赛中的VA(效价唤醒)估计任务。我们设计了一个全面的模型,通过预处理视频帧和音频片段,以提取视觉和音频特征。通过利用时间卷积网络(TCN)模块,我们有效地捕捉到这些功能之间的时间和空间的相关性。随后,我们采用了一个Transformer编码器结构来学习长程依赖关系,从而提高了模型的性能和泛化能力。我们的方法利用了多模态数据融合方法,将预训练的音频和视频骨干集成到特征提取中,然后进行基于TCN的时空编码和基于Transformer的时间信息捕获。实验结果表明,我们的方法的有效性,在AffWild2数据集上的VA估计中实现了有竞争力的性能。摘要:This paper presents our approach for the VA (Valence-Arousal) estimation task in the ABAW6 competition. We devised a comprehensive model by preprocessing video frames and audio segments to extract visual and audio features. Through the utilization of Temporal Convolutional Network (TCN) modules, we effectively captured the temporal and spatial correlations between these features. Subsequently, we employed a Transformer encoder structure to learn long-range dependencies, thereby enhancing the model's performance and generalization ability. Our method leverages a multimodal data fusion approach, integrating pre-trained audio and video backbones for feature extraction, followed by TCN-based spatiotemporal encoding and Transformer-based temporal information capture. Experimental results demonstrate the effectiveness of our approach, achieving competitive performance in VA estimation on the AffWild2 dataset.【3】 MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation作者:Yifan Peng,Ilia Kulikov,Yilin Yang,Sravya Popuri,Hui Lu,Changhan Wang,Hongyu Gong摘要:语音到语音翻译(S2ST),将话语从一种语言翻译成另一种语言,已经出现了研究兴趣和进展。这项工作提出了多任务语音语言模型(MSLM),这是一个解码器的语音语言模型在多任务设置训练。在不依赖文本训练数据的情况下,我们的模型能够支持多语言S2ST,并保留说话人风格。摘要:There have been emerging research interest and advances in speech-to-speech translation (S2ST), translating utterances from one language to another. This work proposes Multitask Speech Language Model (MSLM), which is a decoder-only speech language model trained in a multitask setting. Without reliance on text training data, our model is able to support multilingual S2ST with speaker style preserved.
【4】 An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis作者:Yifan Peng,Ilia Kulikov,Yilin Yang,Sravya Popuri,Hui Lu,Changhan Wang,Hongyu Gong摘要:语音语言模型(LM)是有希望的高质量的语音合成,通过在上下文学习。典型的语音LM以离散的语义单元为内容,以一个简短的话语为提示,合成出保留内容语义但模仿提示风格的语音。然而,对于合成的音频如何由提示和内容控制,还没有系统的理解。在这项工作中,我们进行了广泛使用的自回归(AR)和非自回归(NAR)语音LM的实证研究,并提供了深入的提示设计和内容语义单位。我们的分析表明,异构和非平稳的提示损害了音频质量,而以前的发现,较长的提示总是导致更好的合成。此外,我们发现合成音频的扬声器风格除了受到提示语的影响外,还受到内容的影响。我们进一步表明,语义单位携带丰富的声学信息,如音高,节奏,音量和语音强调,这可能会泄漏的内容到合成的音频。摘要:Speech language models (LMs) are promising for high-quality speech synthesis through in-context learning. A typical speech LM takes discrete semantic units as content and a short utterance as prompt, and synthesizes speech which preserves the content's semantics but mimics the prompt's style. However, there is no systematic understanding on how the synthesized audio is controlled by the prompt and content. In this work, we conduct an empirical study of the widely used autoregressive (AR) and non-autoregressive (NAR) speech LMs and provide insights into the prompt design and content semantic units. Our analysis reveals that heterogeneous and nonstationary prompts hurt the audio quality in contrast to the previous finding that longer prompts always lead to better synthesis. Moreover, we find that the speaker style of the synthesized audio is also affected by the content in addition to the prompt. We further show that semantic units carry rich acoustic information such as pitch, tempo, volume and speech emphasis, which might be leaked from the content to the synthesized audio.【5】 Reproducing the Acoustic Velocity Vectors in a Circular Listening Area作者:Jiarui Wang,Thushara Abhayapala,Jihui Aimee Zhang,Prasanga Samarasinghe备注:Submitted to EUSIPCO 2024摘要:声速矢量是低频声定位的重要参数。提出了一种在圆形听音区域内匹配声速矢量的声场再现算法。在以前的工作中,声速矢量匹配无论是在甜蜜点或边界上的听区。甜蜜点限制了听者的运动,而测量边界上的声速矢量需要复杂的测量设置。本文提出了圆形区域内声速矢量的柱谐系数(CHV系数),它是由声场平移公式计算的全局声压的柱谐系数(全局CHP系数)。全局CHP系数可以通过圆形麦克风阵列来测量,该圆形麦克风阵列可以是现成的。通过匹配CHV系数,在整个收听区域中再现声速矢量。因此,允许听众的移动。仿真结果表明,在低频区域,声速矢量是定位的主导因素,与传统的基于全局CHP系数的方法相比,基于CHV系数的方法可以获得更高精度的声速矢量。摘要:Acoustic velocity vectors are important for human's localization of sound at low frequencies. This paper proposes a sound field reproduction algorithm, which matches the acoustic velocity vectors in a circular listening area. In previous work, acoustic velocity vectors are matched either at sweet spots or on the boundary of the listening area. Sweet spots restrict listener's movement, whereas measuring the acoustic velocity vectors on the boundary requires complicated measurement setup. This paper proposes the cylindrical harmonic coefficients of the acoustic velocity vectors in a circular area (CHV coefficients), which are calculated from the cylindrical harmonic coefficients of the global pressure (global CHP coefficients) by using the sound field translation formula. The global CHP coefficients can be measured by a circular microphone array, which can be bought off-the-shelf. By matching the CHV coefficients, the acoustic velocity vectors are reproduced throughout the listening area. Hence, listener's movements are allowed. Simulations show that at low frequency, where the acoustic velocity vectors are the dominant factor for localization, the proposed reproduction method based on the CHV coefficients results in higher accuracy in reproduced acoustic velocity vectors when compared with traditional method based on the global CHP coefficients.
【6】 A Multi-loudspeaker Binaural Room Impulse Response Dataset with High-Resolution Translational and Rotational Head Coordinates in a Listening Room标题:具有高分辨率平移和旋转头部坐标的多扬声器双耳房间脉冲响应数据集作者:Yue Qiao,Ryan Miguel Gonzales,Edgar Choueiri备注:Submitted to Frontiers in Signal Processing摘要:3D 3A实验室双耳房间脉冲响应(BRIR)数据集的数据报告(https://doi.org/10.34770/6gc9-5787)。摘要:Data report for the 3D3A Lab Binaural Room Impulse Response (BRIR) Dataset (https://doi.org/10.34770/6gc9-5787).【1】 Reproducing the Acoustic Velocity Vectors in a Circular Listening Area作者:Jiarui Wang,Thushara Abhayapala,Jihui Aimee Zhang,Prasanga Samarasinghe备注:Submitted to EUSIPCO 2024摘要:声速矢量是低频声定位的重要参数。提出了一种在圆形听音区域内匹配声速矢量的声场再现算法。在以前的工作中,声速矢量匹配无论是在甜蜜点或边界上的听区。甜蜜点限制了听者的运动,而测量边界上的声速矢量需要复杂的测量设置。本文提出了圆形区域内声速矢量的柱谐系数(CHV系数),它是由声场平移公式计算的全局声压的柱谐系数(全局CHP系数)。全局CHP系数可以通过圆形麦克风阵列来测量,该圆形麦克风阵列可以是现成的。通过匹配CHV系数,在整个收听区域中再现声速矢量。因此,允许听众的移动。仿真结果表明,在低频区域,声速矢量是定位的主导因素,与传统的基于全局CHP系数的方法相比,基于CHV系数的方法可以获得更高精度的声速矢量。摘要:Acoustic velocity vectors are important for human's localization of sound at low frequencies. This paper proposes a sound field reproduction algorithm, which matches the acoustic velocity vectors in a circular listening area. In previous work, acoustic velocity vectors are matched either at sweet spots or on the boundary of the listening area. Sweet spots restrict listener's movement, whereas measuring the acoustic velocity vectors on the boundary requires complicated measurement setup. This paper proposes the cylindrical harmonic coefficients of the acoustic velocity vectors in a circular area (CHV coefficients), which are calculated from the cylindrical harmonic coefficients of the global pressure (global CHP coefficients) by using the sound field translation formula. The global CHP coefficients can be measured by a circular microphone array, which can be bought off-the-shelf. By matching the CHV coefficients, the acoustic velocity vectors are reproduced throughout the listening area. Hence, listener's movements are allowed. Simulations show that at low frequency, where the acoustic velocity vectors are the dominant factor for localization, the proposed reproduction method based on the CHV coefficients results in higher accuracy in reproduced acoustic velocity vectors when compared with traditional method based on the global CHP coefficients.【2】 A Multi-loudspeaker Binaural Room Impulse Response Dataset with High-Resolution Translational and Rotational Head Coordinates in a Listening Room标题:高分辨率平移和旋转头部坐标的多扬声器双耳房间脉冲响应数据集作者:Yue Qiao,Ryan Miguel Gonzales,Edgar Choueiri备注:Submitted to Frontiers in Signal Processing摘要:3D 3A实验室双耳房间脉冲响应(BRIR)数据集的数据报告(https://doi.org/10.34770/6gc9-5787)。摘要:Data report for the 3D3A Lab Binaural Room Impulse Response (BRIR) Dataset (https://doi.org/10.34770/6gc9-5787).
【3】 Latent CLAP Loss for Better Foley Sound Synthesis标题:潜在的CLAP损耗,更好的Foley声音合成作者:Tornike Karchkhadze,Hassan Salami Kavaki,Mohammad Rasool Izadi,Bryce Irvin,Mikolaj Kegler,Ari Hertz,Shuo Zhang,Marko Stamenovic摘要:Foley声音生成,为多媒体创建音频的艺术,最近看到了显着的进步,通过文本条件的潜在扩散模型。这些系统使用多模态文本音频表示模型,例如对比音频预训练(CLAP),其目标是将相应的音频和文本提示映射到联合嵌入空间中。AudioLDM,一个文本到音频的模型,是DCASE 2023 task 7 Foley声音合成挑战的获胜者。获奖系统针对特定音频类别微调了模型,并在推理时使用输出音频和输入文本之间的CLAP相似性分数应用了后过滤方法,需要生成额外的样本,从而降低了数据生成效率。我们引入了一个新的损失项,以增强在AudioLDM没有后滤波的福利声音生成。该损失项使用基于CLAP模式的新模块-Latent CLAP编码-在共享CLAP嵌入空间中将潜在扩散输出与真实音频对齐。实验结果表明,该方法有效地降低了生成音频的Frechet音频距离(FAD)分数,消除了后期滤波的需要,从而提高了生成效率。摘要:Foley sound generation, the art of creating audio for multimedia, has recently seen notable advancements through text-conditioned latent diffusion models. These systems use multimodal text-audio representation models, such as Contrastive Language-Audio Pretraining (CLAP), whose objective is to map corresponding audio and text prompts into a joint embedding space. AudioLDM, a text-to-audio model, was the winner of the DCASE2023 task 7 Foley sound synthesis challenge. The winning system fine-tuned the model for specific audio classes and applied a post-filtering method using CLAP similarity scores between output audio and input text at inference time, requiring the generation of extra samples, thus reducing data generation efficiency. We introduce a new loss term to enhance Foley sound generation in AudioLDM without post-filtering. This loss term uses a new module based on the CLAP mode-Latent CLAP encode-to align the latent diffusion output with real audio in a shared CLAP embedding space. Our experiments demonstrate that our method effectively reduces the Frechet Audio Distance (FAD) score of the generated audio and eliminates the need for post-filtering, thus enhancing generation efficiency.
【4】 Real-time Speech Extraction Using Spatially Regularized Independent Low-rank Matrix Analysis and Rank-constrained Spatial Covariance Matrix Estimation标题:基于空间正则化独立低秩阵分析和秩受限空间协方差矩阵估计的实时语音提取作者:Yuto Ishikawa,Kohei Konaka,Tomohiko Nakamura,Norihiro Takamune,Hiroshi Saruwatari备注:5 pages, 3 figures, accepted at HSCMA 2024摘要:实时语音提取是各种应用的一个重要挑战,例如类人化身/机器人中的语音识别。本文提出了一种基于独立低秩矩阵分析(ILRMA)和秩约束空间协方差矩阵估计(RCSCME)的语音提取方法的实时扩展。基于RCSCME的方法是一种多通道盲语音提取方法,其在扩散噪声环境中表现出优异的语音提取性能。为了提高性能,我们将空间正则化引入到基于RCSCME的语音提取的ILRMA部分,并设计了两个正则化器。语音提取实验表明,所提出的方法可以实时工作,所设计的正则化器提高了语音提取性能。摘要:Real-time speech extraction is an important challenge with various applications such as speech recognition in a human-like avatar/robot. In this paper, we propose the real-time extension of a speech extraction method based on independent low-rank matrix analysis (ILRMA) and rank-constrained spatial covariance matrix estimation (RCSCME). The RCSCME-based method is a multichannel blind speech extraction method that demonstrates superior speech extraction performance in diffuse noise environments. To improve the performance, we introduce spatial regularization into the ILRMA part of the RCSCME-based speech extraction and design two regularizers. Speech extraction experiments demonstrated that the proposed methods can function in real time and the designed regularizers improve the speech extraction performance.【5】 Multimodal Fusion Method with Spatiotemporal Sequences and Relationship Learning for Valence-Arousal Estimation作者:Jun Yu,Gongpeng Zhao,Yongqi Wan,Zhihong Wei,Yang Zheng,Zerui Zhang,Zhongpeng Cai,Guochen Xie,Jichao Zhu,Wangyuan Zhu摘要:本文介绍了我们的方法在ABAW6比赛中的VA(效价唤醒)估计任务。我们设计了一个全面的模型,通过预处理视频帧和音频片段,以提取视觉和音频特征。通过利用时间卷积网络(TCN)模块,我们有效地捕捉到这些功能之间的时间和空间的相关性。随后,我们采用了一个Transformer编码器结构来学习长程依赖关系,从而提高了模型的性能和泛化能力。我们的方法利用了多模态数据融合方法,将预训练的音频和视频骨干集成到特征提取中,然后进行基于TCN的时空编码和基于Transformer的时间信息捕获。实验结果表明,我们的方法的有效性,在AffWild2数据集上的VA估计中实现了有竞争力的性能。摘要:This paper presents our approach for the VA (Valence-Arousal) estimation task in the ABAW6 competition. We devised a comprehensive model by preprocessing video frames and audio segments to extract visual and audio features. Through the utilization of Temporal Convolutional Network (TCN) modules, we effectively captured the temporal and spatial correlations between these features. Subsequently, we employed a Transformer encoder structure to learn long-range dependencies, thereby enhancing the model's performance and generalization ability. Our method leverages a multimodal data fusion approach, integrating pre-trained audio and video backbones for feature extraction, followed by TCN-based spatiotemporal encoding and Transformer-based temporal information capture. Experimental results demonstrate the effectiveness of our approach, achieving competitive performance in VA estimation on the AffWild2 dataset.
【6】 MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation作者:Yifan Peng,Ilia Kulikov,Yilin Yang,Sravya Popuri,Hui Lu,Changhan Wang,Hongyu Gong摘要:语音到语音翻译(S2ST),将话语从一种语言翻译成另一种语言,已经出现了研究兴趣和进展。这项工作提出了多任务语音语言模型(MSLM),这是一个解码器的语音语言模型在多任务设置训练。在不依赖文本训练数据的情况下,我们的模型能够支持多语言S2ST,并保留说话人风格。摘要:There have been emerging research interest and advances in speech-to-speech translation (S2ST), translating utterances from one language to another. This work proposes Multitask Speech Language Model (MSLM), which is a decoder-only speech language model trained in a multitask setting. Without reliance on text training data, our model is able to support multilingual S2ST with speaker style preserved.
【7】 An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis作者:Yifan Peng,Ilia Kulikov,Yilin Yang,Sravya Popuri,Hui Lu,Changhan Wang,Hongyu Gong摘要:语音语言模型(LM)是有希望的高质量的语音合成,通过在上下文学习。典型的语音LM以离散的语义单元为内容,以一个简短的话语为提示,合成出保留内容语义但模仿提示风格的语音。然而,对于合成的音频如何由提示和内容控制,还没有系统的理解。在这项工作中,我们进行了广泛使用的自回归(AR)和非自回归(NAR)语音LM的实证研究,并提供了深入的提示设计和内容语义单位。我们的分析表明,异构和非平稳的提示损害了音频质量,而以前的发现,较长的提示总是导致更好的合成。此外,我们发现合成音频的扬声器风格除了受到提示语的影响外,还受到内容的影响。我们进一步表明,语义单位携带丰富的声学信息,如音高,节奏,音量和语音强调,这可能会泄漏的内容到合成的音频。摘要:Speech language models (LMs) are promising for high-quality speech synthesis through in-context learning. A typical speech LM takes discrete semantic units as content and a short utterance as prompt, and synthesizes speech which preserves the content's semantics but mimics the prompt's style. However, there is no systematic understanding on how the synthesized audio is controlled by the prompt and content. In this work, we conduct an empirical study of the widely used autoregressive (AR) and non-autoregressive (NAR) speech LMs and provide insights into the prompt design and content semantic units. Our analysis reveals that heterogeneous and nonstationary prompts hurt the audio quality in contrast to the previous finding that longer prompts always lead to better synthesis. Moreover, we find that the speaker style of the synthesized audio is also affected by the content in addition to the prompt. We further show that semantic units carry rich acoustic information such as pitch, tempo, volume and speech emphasis, which might be leaked from the content to the synthesized audio.