今日论文合集:cs.SD语音6篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音

【1】 CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language  Models

标题:CosyVoice 2:具有大型语言模型的可扩展流语音合成
链接:https://arxiv.org/abs/2412.10117
作者:Zhihao Du,  Yuxuan Wang,  Qian Chen,  Xian Shi,  Xiang Lv,  Tianyu Zhao,  Zhifu Gao,  Yexin Yang,  Changfeng Gao,  Hui Wang,  Fan Yu,  Huadai Liu,  Zhengyan Sheng,  Yue Gu,  Chong Deng,  Wen Wang,  Shiliang Zhang,  Zhijie Yan,  Jingren Zhou
备注:Tech report, work in progress
摘要:在我们之前的工作中,我们介绍了CosyVoice,一个基于监督离散语音令牌的多语言语音合成模型。通过采用渐进式语义解码与两个流行的生成模型,语言模型(LM)和流匹配,CosyVoice表现出较高的韵律自然度,内容一致性,说话人相似性在语音上下文学习。近年来,多模态大型语言模型(LLM)取得了重大进展,其中语音合成的响应延迟和实时因素在交互体验中起着至关重要的作用。因此,在这份报告中,我们提出了一个改进的流式语音合成模型,CosyVoice 2,它采用了全面和系统的优化。具体来说,我们引入有限标量量化,以提高语音令牌的码本利用率。对于文本-语音LM,我们简化了模型架构,允许直接使用预先训练的LLM作为骨干。此外,我们开发了一个块感知的因果流匹配模型来支持各种合成场景,从而在单个模型中实现流和非流合成。通过在大规模多语言数据集上进行训练,CosyVoice 2在流模式下实现了人类平价自然度,最小的响应延迟和几乎无损的合成质量。我们邀请读者在https://funaudiollm.github.io/cosyvoice2上收听演示。
摘要:In our previous work, we introduced CosyVoice, a multilingual speechsynthesis model based on supervised discrete speech tokens. By employingprogressive semantic decoding with two popular generative models, languagemodels (LMs) and Flow Matching, CosyVoice demonstrated high prosodynaturalness, content consistency, and speaker similarity in speech in-contextlearning. Recently, significant progress has been made in multi-modal largelanguage models (LLMs), where the response latency and real-time factor ofspeech synthesis play a crucial role in the interactive experience. Therefore,in this report, we present an improved streaming speech synthesis model,CosyVoice 2, which incorporates comprehensive and systematic optimizations.Specifically, we introduce finite-scalar quantization to improve the codebookutilization of speech tokens. For the text-speech LM, we streamline the modelarchitecture to allow direct use of a pre-trained LLM as the backbone. Inaddition, we develop a chunk-aware causal flow matching model to supportvarious synthesis scenarios, enabling both streaming and non-streamingsynthesis within a single model. By training on a large-scale multilingualdataset, CosyVoice 2 achieves human-parity naturalness, minimal responselatency, and virtually lossless synthesis quality in the streaming mode. Weinvite readers to listen to the demos athttps://funaudiollm.github.io/cosyvoice2.

【2】 Enhanced Speech Emotion Recognition with Efficient Channel Attention  Guided Deep CNN-BiLSTM Framework
标题:利用高效通道注意力引导的深度CNN-BiLSTM框架增强语音情感识别
链接:https://arxiv.org/abs/2412.10011
作者:Niloy Kumar Kundu,  Sarah Kobir,  Md. Rayhan Ahmed,  Tahmina Aktar,  Niloya Roy
备注:42 pages,10 figures
摘要:语音情感识别对于增强情感计算和丰富人机交互领域具有重要意义。然而,SER的主要挑战在于从语音信号中选择具有较低计算成本的相关特征表示。在本文中,我们提出了一个轻量级的SER架构,集成了基于注意力的局部特征块(ALFB),以捕捉高层次的相关特征向量的语音信号。我们还采用了全局特征块(GFB)技术来捕获语音信号中的序列、全局信息和长期依赖性。通过聚合基于注意力的局部和全局上下文特征向量,我们的模型有效地捕获了反映复杂人类情感线索的显著特征之间的内部相关性。为了评估我们的方法,我们从语音音频样本中提取了四种类型的频谱特征:梅尔频率倒谱系数,梅尔频谱图,均方根值和过零率。通过5重交叉验证策略,我们在5个多语言标准基准数据集上测试了所提出的方法:TESS,RAVDESS,BanglaSER,SUBESCO和EST-DB,平均准确率分别为99.65%,94.88%,98.12%,97.94%和97.19%。结果表明,与大多数现有方法相比,我们的模型达到了最先进的(SOTA)性能。
摘要:Speech emotion recognition (SER) is crucial for enhancing affective computingand enriching the domain of human-computer interaction. However, the mainchallenge in SER lies in selecting relevant feature representations from speechsignals with lower computational costs. In this paper, we propose a lightweightSER architecture that integrates attention-based local feature blocks (ALFBs)to capture high-level relevant feature vectors from speech signals. We alsoincorporate a global feature block (GFB) technique to capture sequential,global information and long-term dependencies in speech signals. By aggregatingattention-based local and global contextual feature vectors, our modeleffectively captures the internal correlation between salient features thatreflect complex human emotional cues. To evaluate our approach, we extractedfour types of spectral features from speech audio samples: mel-frequencycepstral coefficients, mel-spectrogram, root mean square value, andzero-crossing rate. Through a 5-fold cross-validation strategy, we tested theproposed method on five multi-lingual standard benchmark datasets: TESS,RAVDESS, BanglaSER, SUBESCO, and Emo-DB, and obtained a mean accuracy of99.65%, 94.88%, 98.12%, 97.94%, and 97.19% respectively. The results indicatethat our model achieves state-of-the-art (SOTA) performance compared to mostexisting methods.

【3】 Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's  Disease Identification
标题:利用多模式方法和自发言语识别阿尔茨海默病
链接:https://arxiv.org/abs/2412.09928
作者:Yifan Gao,  Long Guo,  Hong Liu
备注:ICASSP 2025
摘要:通过自发言语检测认知障碍为阿尔茨海默病(AD)和轻度认知障碍(MCI)的早期诊断提供了可能。作为ICASSP 2025的一部分,PROCESS Grand Challenge专注于通过分类和回归任务的创新解决方案来推进这一领域。在这项工作中,我们通过多模态融合策略将可解释的特征与从预训练模型中提取的时间特征相结合。对于分类任务,我们的模型在预测认知状态(健康,MCI,痴呆)方面的F1得分为0.649。对于涉及MMSE评分预测的回归任务,我们获得的均方根误差(RMSE)为2.628。这些结果使我们的团队在比赛中获得了最高的总排名。
摘要:Cognitive impairment detection through spontaneous speech offers potentialfor early diagnosis of Alzheimer's disease (AD) and mild cognitive impairment(MCI). The PROCESS Grand Challenge, part of ICASSP 2025, focuses on advancingthis field with innovative solutions for classification and regression tasks.In this work, we integrate interpretable features with temporal featuresextracted from pre-trained models through a multimodal fusion strategy. For theclassification task, our model achieved an F1-score of 0.649 in predictingcognitive states (healthy, MCI, dementia). For the regression task, whichinvolves MMSE score prediction, we obtained a root-mean-square error (RMSE) of2.628. These results led to our team securing the top overall ranking in thecompetition.

【4】 SonicBoom: Contact Localization Using Array of Microphones
标题:SonicBoom:使用麦克风阵列的联系人定位
链接:https://arxiv.org/abs/2412.09878
作者:Moonyoung Lee,  Uksang Yoo,  Jean Oh,  Jeffrey Ichnowski,  George Kantor,  Oliver Kroemer
备注:8 pages
摘要:在视觉传感器遇到严重遮挡的杂乱环境中,例如在农业环境中,触觉信号可以为机器人提供关键的空间信息,以定位刚性物体并在其周围机动。我们推出SonicBoom,这是一个全面的硬件和学习管道,通过一系列接触式麦克风实现接触式定位。虽然传统的声源定位方法可以有效地对空气中的声源进行三角测量,但通过具有不规则几何形状和结构的固体介质进行定位提出了难以分析建模的挑战。我们通过基于特征工程和学习的方法来解决这一挑战,自主收集18,000个机器人交互声音对,以学习机器人末端执行器链接上的声学信号和碰撞位置之间的映射。通过利用麦克风之间的相对特性,SonicBoom在分配交互中实现了0.42 cm的定位误差,即使在新的物体和接触条件下也能保持2.22 cm的稳健性能。我们证明了该系统的实际效用,通过触觉映射闭塞的分支机构在模拟树冠设置,显示基于声学的传感可以使可靠的机器人导航在视觉上具有挑战性的环境。
摘要:In cluttered environments where visual sensors encounter heavy occlusion,such as in agricultural settings, tactile signals can provide crucial spatialinformation for the robot to locate rigid objects and maneuver around them. Weintroduce SonicBoom, a holistic hardware and learning pipeline that enablescontact localization through an array of contact microphones. Whileconventional sound source localization methods effectively triangulate sourcesin air, localization through solid media with irregular geometry and structurepresents challenges that are difficult to model analytically. We address thischallenge through a feature engineering and learning based approach,autonomously collecting 18,000 robot interaction sound pairs to learn a mappingbetween acoustic signals and collision locations on the robot end effectorlink. By leveraging relative features between microphones, SonicBoom achieveslocalization errors of 0.42cm for in distribution interactions and maintainsrobust performance of 2.22cm error even with novel objects and contactconditions. We demonstrate the system's practical utility through hapticmapping of occluded branches in mock canopy settings, showing that acousticbased sensing can enable reliable robot navigation in visually challengingenvironments.

【5】 SILA: Signal-to-Language Augmentation for Enhanced Control in  Text-to-Audio Generation
标题:SILA:用于文本到音频生成中增强控制的信号到语言增强
链接:https://arxiv.org/abs/2412.09789
作者:Sonal Kumar,  Prem Seetharaman,  Justin Salamon,  Dinesh Manocha,  Oriol Nieto
备注:Website: this https URL
摘要:文本到音频生成领域已经取得了重大进展,但精细控制所生成音频的声学特性的能力仍然未得到充分开发。在本文中,我们介绍了一种新颖而简单的方法来生成声音效果,控制关键的声学参数,如响度,音高,混响,褪色,亮度,噪音和持续时间,使创造性的应用程序在声音设计和内容创作。这些参数超越了传统的数字信号处理(DSP)技术,结合了学习到的表示,捕捉声音特征如何在上下文中形成的微妙之处,从而对生成的音频进行更丰富、更细致的控制。我们的方法是模型不可知的,是基于学习音频语义和它的声学特征之间的解开。我们的方法不仅增强了文本到音频生成的多功能性和表现力,而且为创造性的音频制作和声音设计开辟了新的途径。我们的客观和主观评估结果证明了我们的方法在生产高质量,可定制的音频输出,密切配合用户的规格的有效性。
摘要:The field of text-to-audio generation has seen significant advancements, andyet the ability to finely control the acoustic characteristics of generatedaudio remains under-explored. In this paper, we introduce a novel yet simpleapproach to generate sound effects with control over key acoustic parameterssuch as loudness, pitch, reverb, fade, brightness, noise and duration, enablingcreative applications in sound design and content creation. These parametersextend beyond traditional Digital Signal Processing (DSP) techniques,incorporating learned representations that capture the subtleties of how soundcharacteristics can be shaped in context, enabling a richer and more nuancedcontrol over the generated audio. Our approach is model-agnostic and is basedon learning the disentanglement between audio semantics and its acousticfeatures. Our approach not only enhances the versatility and expressiveness oftext-to-audio generation but also opens new avenues for creative audioproduction and sound design. Our objective and subjective evaluation resultsdemonstrate the effectiveness of our approach in producing high-quality,customizable audio outputs that align closely with user specifications.

【6】 CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on  Conditional Transformer with Fine-Grained Lyric and Musical Controls
标题:CSl-L2 M:基于条件Transformer的可控歌曲级歌词到旋律生成,具有细粒度歌词和音乐控制
链接:https://arxiv.org/abs/2412.09887
作者:Li Chai,  Donglin Wang
备注:Accepted at AAAI-25
摘要:歌词到旋律生成是AI音乐生成领域中极具挑战性的任务。由于难以学习歌词和旋律之间严格但弱的相关性,以前的方法具有可控性弱,质量低和结构不良的生成。为了解决这些挑战,我们提出了CSL-L2 M,一个可控的歌曲级的歌词到旋律生成方法的基础上的注意力Transformer解码器与细粒度的歌词和音乐控制,这是能够生成完整的歌曲旋律与给定的歌词和用户指定的音乐属性相匹配。具体来说,我们首先介绍REMI对齐,一种新的音乐表示,采用严格的音节和音节级的歌词和旋律之间的对齐,促进精确的对齐建模。随后,从逐行Transformer编码器独立提取的逐行级语义歌词嵌入与词级词性嵌入和音节级音调嵌入相结合,作为细粒度控制,以增强歌词对旋律生成的可控性。然后,我们引入人类标记的音乐标签,音乐级别的统计音乐属性,并从预训练的VQ-VAE中提取的学习音乐特征作为粗粒度,细粒度和高保真度控制,分别到生成过程中,从而使用户能够控制旋律生成。最后,一个注意力Transformer解码器技术被用来对具有上述歌词和音乐条件的全曲旋律生成施加细粒度控制。实验结果表明,我们提出的CSL-L2 M优于国家的最先进的模型,产生更高的质量,更好的可控性和增强的结构旋律。演示和源代码可在https://lichaiustc.github.io/CSL-L2M/上获得。
摘要:Lyric-to-melody generation is a highly challenging task in the field of AImusic generation. Due to the difficulty of learning strict yet weakcorrelations between lyrics and melodies, previous methods have suffered fromweak controllability, low-quality and poorly structured generation. To addressthese challenges, we propose CSL-L2M, a controllable song-level lyric-to-melodygeneration method based on an in-attention Transformer decoder withfine-grained lyric and musical controls, which is able to generate full-songmelodies matched with the given lyrics and user-specified musical attributes.Specifically, we first introduce REMI-Aligned, a novel music representationthat incorporates strict syllable- and sentence-level alignments between lyricsand melodies, facilitating precise alignment modeling. Subsequently,sentence-level semantic lyric embeddings independently extracted from asentence-wise Transformer encoder are combined with word-level part-of-speechembeddings and syllable-level tone embeddings as fine-grained controls toenhance the controllability of lyrics over melody generation. Then we introducehuman-labeled musical tags, sentence-level statistical musical attributes, andlearned musical features extracted from a pre-trained VQ-VAE as coarse-grained,fine-grained and high-fidelity controls, respectively, to the generationprocess, thereby enabling user control over melody generation. Finally, anin-attention Transformer decoder technique is leveraged to exert fine-grainedcontrol over the full-song melody generation with the aforementioned lyric andmusical conditions. Experimental results demonstrate that our proposed CSL-L2Moutperforms the state-of-the-art models, generating melodies with higherquality, better controllability and enhanced structure. Demos and source codeare available at https://lichaiustc.github.io/CSL-L2M/.

eess.AS音频处理

【1】 CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on  Conditional Transformer with Fine-Grained Lyric and Musical Controls
标题:CSl-L2 M:基于条件Transformer的可控歌曲级歌词到旋律生成,具有细粒度歌词和音乐控制
链接:https://arxiv.org/abs/2412.09887
作者:Li Chai,  Donglin Wang
备注:Accepted at AAAI-25
摘要:歌词到旋律生成是AI音乐生成领域中极具挑战性的任务。由于难以学习歌词和旋律之间严格但弱的相关性,以前的方法具有可控性弱,质量低和结构不良的生成。为了解决这些挑战,我们提出了CSL-L2 M,一个可控的歌曲级的歌词到旋律生成方法的基础上的注意力Transformer解码器与细粒度的歌词和音乐控制,这是能够生成完整的歌曲旋律与给定的歌词和用户指定的音乐属性相匹配。具体来说,我们首先介绍REMI对齐,一种新的音乐表示,采用严格的音节和音节级的歌词和旋律之间的对齐,促进精确的对齐建模。随后,从逐行Transformer编码器独立提取的逐行级语义歌词嵌入与词级词性嵌入和音节级音调嵌入相结合,作为细粒度控制,以增强歌词对旋律生成的可控性。然后,我们引入人类标记的音乐标签,音乐级别的统计音乐属性,并从预训练的VQ-VAE中提取的学习音乐特征作为粗粒度,细粒度和高保真度控制,分别到生成过程中,从而使用户能够控制旋律生成。最后,利用关注中的Transformer解码器技术对具有上述歌词和音乐条件的整首歌曲旋律生成进行细粒度控制。实验结果表明,我们提出的CSL-L2 M优于国家的最先进的模型,产生更高的质量,更好的可控性和增强的结构旋律。演示和源代码可在https://lichaiustc.github.io/CSL-L2M/上获得。
摘要:Lyric-to-melody generation is a highly challenging task in the field of AImusic generation. Due to the difficulty of learning strict yet weakcorrelations between lyrics and melodies, previous methods have suffered fromweak controllability, low-quality and poorly structured generation. To addressthese challenges, we propose CSL-L2M, a controllable song-level lyric-to-melodygeneration method based on an in-attention Transformer decoder withfine-grained lyric and musical controls, which is able to generate full-songmelodies matched with the given lyrics and user-specified musical attributes.Specifically, we first introduce REMI-Aligned, a novel music representationthat incorporates strict syllable- and sentence-level alignments between lyricsand melodies, facilitating precise alignment modeling. Subsequently,sentence-level semantic lyric embeddings independently extracted from asentence-wise Transformer encoder are combined with word-level part-of-speechembeddings and syllable-level tone embeddings as fine-grained controls toenhance the controllability of lyrics over melody generation. Then we introducehuman-labeled musical tags, sentence-level statistical musical attributes, andlearned musical features extracted from a pre-trained VQ-VAE as coarse-grained,fine-grained and high-fidelity controls, respectively, to the generationprocess, thereby enabling user control over melody generation. Finally, anin-attention Transformer decoder technique is leveraged to exert fine-grainedcontrol over the full-song melody generation with the aforementioned lyric andmusical conditions. Experimental results demonstrate that our proposed CSL-L2Moutperforms the state-of-the-art models, generating melodies with higherquality, better controllability and enhanced structure. Demos and source codeare available at https://lichaiustc.github.io/CSL-L2M/.

【2】 CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language  Models
标题:CosyVoice 2:具有大型语言模型的可扩展流语音合成
链接:https://arxiv.org/abs/2412.10117
作者:Zhihao Du,  Yuxuan Wang,  Qian Chen,  Xian Shi,  Xiang Lv,  Tianyu Zhao,  Zhifu Gao,  Yexin Yang,  Changfeng Gao,  Hui Wang,  Fan Yu,  Huadai Liu,  Zhengyan Sheng,  Yue Gu,  Chong Deng,  Wen Wang,  Shiliang Zhang,  Zhijie Yan,  Jingren Zhou
备注:Tech report, work in progress
摘要:在我们之前的工作中,我们介绍了CosyVoice,一个基于监督离散语音令牌的多语言语音合成模型。通过采用渐进式语义解码与两个流行的生成模型,语言模型(LM)和流匹配,CosyVoice表现出较高的韵律自然度,内容一致性,说话人相似性在语音上下文学习。近年来,多模态大型语言模型(LLM)取得了重大进展,其中语音合成的响应延迟和实时因素在交互体验中起着至关重要的作用。因此,在这份报告中,我们提出了一个改进的流式语音合成模型,CosyVoice 2,它采用了全面和系统的优化。具体来说,我们引入有限标量量化,以提高语音令牌的码本利用率。对于文本-语音LM,我们简化了模型架构,允许直接使用预先训练的LLM作为骨干。此外,我们开发了一个块感知的因果流匹配模型来支持各种合成场景,从而在单个模型中实现流和非流合成。通过在大规模多语言数据集上进行训练,CosyVoice 2在流模式下实现了人类平价自然度,最小的响应延迟和几乎无损的合成质量。我们邀请读者在https://funaudiollm.github.io/cosyvoice2上收听演示。
摘要:In our previous work, we introduced CosyVoice, a multilingual speechsynthesis model based on supervised discrete speech tokens. By employingprogressive semantic decoding with two popular generative models, languagemodels (LMs) and Flow Matching, CosyVoice demonstrated high prosodynaturalness, content consistency, and speaker similarity in speech in-contextlearning. Recently, significant progress has been made in multi-modal largelanguage models (LLMs), where the response latency and real-time factor ofspeech synthesis play a crucial role in the interactive experience. Therefore,in this report, we present an improved streaming speech synthesis model,CosyVoice 2, which incorporates comprehensive and systematic optimizations.Specifically, we introduce finite-scalar quantization to improve the codebookutilization of speech tokens. For the text-speech LM, we streamline the modelarchitecture to allow direct use of a pre-trained LLM as the backbone. Inaddition, we develop a chunk-aware causal flow matching model to supportvarious synthesis scenarios, enabling both streaming and non-streamingsynthesis within a single model. By training on a large-scale multilingualdataset, CosyVoice 2 achieves human-parity naturalness, minimal responselatency, and virtually lossless synthesis quality in the streaming mode. Weinvite readers to listen to the demos athttps://funaudiollm.github.io/cosyvoice2.

【3】 Enhanced Speech Emotion Recognition with Efficient Channel Attention  Guided Deep CNN-BiLSTM Framework
标题:利用高效通道注意力引导的深度CNN-BiLSTM框架增强语音情感识别
链接:https://arxiv.org/abs/2412.10011
作者:Niloy Kumar Kundu,  Sarah Kobir,  Md. Rayhan Ahmed,  Tahmina Aktar,  Niloya Roy
备注:42 pages,10 figures
摘要:语音情感识别对于增强情感计算和丰富人机交互领域具有重要意义。然而,SER的主要挑战在于从语音信号中选择具有较低计算成本的相关特征表示。在本文中,我们提出了一个轻量级的SER架构,集成了基于注意力的局部特征块(ALFB),以捕捉高层次的相关特征向量的语音信号。我们还采用了一个全局特征块(GFB)技术来捕获语音信号中的序列,全局信息和长期依赖关系。通过聚合基于注意力的局部和全局上下文特征向量,我们的模型有效地捕获了反映复杂人类情感线索的显著特征之间的内部相关性。为了评估我们的方法,我们从语音音频样本中提取了四种类型的频谱特征:梅尔频率倒谱系数,梅尔频谱图,均方根值和过零率。通过5重交叉验证策略,我们在5个多语言标准基准数据集上测试了所提出的方法:TESS,RAVDESS,BanglaSER,SUBESCO和EST-DB,并分别获得了99.65%,94.88%,98.12%,97.94%和97.19%的平均准确率。结果表明,与大多数现有方法相比,我们的模型达到了最先进的(SOTA)性能。
摘要:Speech emotion recognition (SER) is crucial for enhancing affective computingand enriching the domain of human-computer interaction. However, the mainchallenge in SER lies in selecting relevant feature representations from speechsignals with lower computational costs. In this paper, we propose a lightweightSER architecture that integrates attention-based local feature blocks (ALFBs)to capture high-level relevant feature vectors from speech signals. We alsoincorporate a global feature block (GFB) technique to capture sequential,global information and long-term dependencies in speech signals. By aggregatingattention-based local and global contextual feature vectors, our modeleffectively captures the internal correlation between salient features thatreflect complex human emotional cues. To evaluate our approach, we extractedfour types of spectral features from speech audio samples: mel-frequencycepstral coefficients, mel-spectrogram, root mean square value, andzero-crossing rate. Through a 5-fold cross-validation strategy, we tested theproposed method on five multi-lingual standard benchmark datasets: TESS,RAVDESS, BanglaSER, SUBESCO, and Emo-DB, and obtained a mean accuracy of99.65%, 94.88%, 98.12%, 97.94%, and 97.19% respectively. The results indicatethat our model achieves state-of-the-art (SOTA) performance compared to mostexisting methods.

【4】 Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's  Disease Identification
标题:利用多模式方法和自发言语识别阿尔茨海默病
链接:https://arxiv.org/abs/2412.09928
作者:Yifan Gao,  Long Guo,  Hong Liu
备注:ICASSP 2025
摘要:通过自发言语检测认知障碍为阿尔茨海默病(AD)和轻度认知障碍(MCI)的早期诊断提供了可能。作为ICASSP 2025的一部分,PROCESS Grand Challenge专注于通过分类和回归任务的创新解决方案来推进这一领域。在这项工作中,我们通过多模式融合策略将可解释的特征与从预训练模型中提取的时间特征集成在一起。对于分类任务,我们的模型在预测认知状态(健康,MCI,痴呆)方面的F1得分为0.649。对于涉及MMSE评分预测的回归任务,我们获得的均方根误差(RMSE)为2.628。这些结果使我们的团队在比赛中获得了最高的总排名。
摘要:Cognitive impairment detection through spontaneous speech offers potentialfor early diagnosis of Alzheimer's disease (AD) and mild cognitive impairment(MCI). The PROCESS Grand Challenge, part of ICASSP 2025, focuses on advancingthis field with innovative solutions for classification and regression tasks.In this work, we integrate interpretable features with temporal featuresextracted from pre-trained models through a multimodal fusion strategy. For theclassification task, our model achieved an F1-score of 0.649 in predictingcognitive states (healthy, MCI, dementia). For the regression task, whichinvolves MMSE score prediction, we obtained a root-mean-square error (RMSE) of2.628. These results led to our team securing the top overall ranking in thecompetition.

【5】 SonicBoom: Contact Localization Using Array of Microphones
标题:SonicBoom:使用麦克风阵列的联系人定位
链接:https://arxiv.org/abs/2412.09878
作者:Moonyoung Lee,  Uksang Yoo,  Jean Oh,  Jeffrey Ichnowski,  George Kantor,  Oliver Kroemer
备注:8 pages
摘要:在视觉传感器遇到严重遮挡的杂乱环境中,例如在农业环境中,触觉信号可以为机器人提供关键的空间信息,以定位刚性物体并在其周围机动。我们推出SonicBoom,这是一个全面的硬件和学习管道,通过一系列接触式麦克风实现接触式定位。虽然传统的声源定位方法可以有效地对空气中的声源进行三角测量,但通过具有不规则几何形状和结构的固体介质进行定位提出了难以分析建模的挑战。我们通过基于特征工程和学习的方法来解决这一挑战,自主收集18,000个机器人交互声音对,以学习机器人末端执行器链接上的声学信号和碰撞位置之间的映射。通过利用麦克风之间的相对特性,SonicBoom在分配交互中实现了0.42 cm的定位误差,即使在新的物体和接触条件下也能保持2.22 cm的稳健性能。我们证明了该系统的实际效用,通过触觉映射闭塞的分支机构在模拟树冠设置,显示基于声学的传感可以使可靠的机器人导航在视觉上具有挑战性的环境。
摘要:In cluttered environments where visual sensors encounter heavy occlusion,such as in agricultural settings, tactile signals can provide crucial spatialinformation for the robot to locate rigid objects and maneuver around them. Weintroduce SonicBoom, a holistic hardware and learning pipeline that enablescontact localization through an array of contact microphones. Whileconventional sound source localization methods effectively triangulate sourcesin air, localization through solid media with irregular geometry and structurepresents challenges that are difficult to model analytically. We address thischallenge through a feature engineering and learning based approach,autonomously collecting 18,000 robot interaction sound pairs to learn a mappingbetween acoustic signals and collision locations on the robot end effectorlink. By leveraging relative features between microphones, SonicBoom achieveslocalization errors of 0.42cm for in distribution interactions and maintainsrobust performance of 2.22cm error even with novel objects and contactconditions. We demonstrate the system's practical utility through hapticmapping of occluded branches in mock canopy settings, showing that acousticbased sensing can enable reliable robot navigation in visually challengingenvironments.

【6】 SILA: Signal-to-Language Augmentation for Enhanced Control in  Text-to-Audio Generation
标题:SILA:用于文本到音频生成中增强控制的信号到语言增强
链接:https://arxiv.org/abs/2412.09789
作者:Sonal Kumar,  Prem Seetharaman,  Justin Salamon,  Dinesh Manocha,  Oriol Nieto
备注:Website: this https URL
摘要:文本到音频生成领域已经取得了重大进展,但精细控制所生成音频的声学特性的能力仍然未得到充分开发。在本文中,我们介绍了一种新颖而简单的方法来生成声音效果,控制关键的声学参数,如响度,音高,混响,褪色,亮度,噪音和持续时间,使创造性的应用程序在声音设计和内容创作。这些参数超越了传统的数字信号处理(DSP)技术,结合了学习到的表示,捕捉声音特征如何在上下文中形成的微妙之处,从而对生成的音频进行更丰富、更细致的控制。我们的方法是模型不可知的,是基于学习音频语义和它的声学特征之间的解开。我们的方法不仅增强了文本到音频生成的多功能性和表现力,而且为创造性的音频制作和声音设计开辟了新的途径。我们的客观和主观评估结果证明了我们的方法在生产高质量,可定制的音频输出,密切配合用户的规格的有效性。
摘要:The field of text-to-audio generation has seen significant advancements, andyet the ability to finely control the acoustic characteristics of generatedaudio remains under-explored. In this paper, we introduce a novel yet simpleapproach to generate sound effects with control over key acoustic parameterssuch as loudness, pitch, reverb, fade, brightness, noise and duration, enablingcreative applications in sound design and content creation. These parametersextend beyond traditional Digital Signal Processing (DSP) techniques,incorporating learned representations that capture the subtleties of how soundcharacteristics can be shaped in context, enabling a richer and more nuancedcontrol over the generated audio. Our approach is model-agnostic and is basedon learning the disentanglement between audio semantics and its acousticfeatures. Our approach not only enhances the versatility and expressiveness oftext-to-audio generation but also opens new avenues for creative audioproduction and sound design. Our objective and subjective evaluation resultsdemonstrate the effectiveness of our approach in producing high-quality,customizable audio outputs that align closely with user specifications.

机器翻译由腾讯交互翻译提供,仅供参考