今日论文合集:cs.SD语音5篇,eess.AS音频处理5篇。
本文经arXiv每日学术速递授权转载
标题: MAT-MED:AMasked音频Transformer,具有基于掩蔽重建的预训练,用于声音事件检测
作者:Pengfei Cai,Yan Song,Kang Li,Haoyu Song,Ian McLoughlin
备注:Received by interspeech 2024
链接:点击下载PDF文件
摘要:利用大型预训练的Transformer编码器网络的声音事件检测(SED)方法在最近的DCASE挑战中显示出有希望的性能。然而,他们仍然依赖于基于RNN的上下文网络来建模时间依赖性,这主要是由于标记数据的稀缺性。在这项工作中,我们提出了一个纯粹的基于transformer的SED模型与掩蔽重建为基础的预训练,称为MAT-SED。具体而言,首先设计具有相对位置编码的Transformer作为上下文网络,通过掩蔽重建任务以自监督的方式对所有可用的目标数据进行预训练。编码器和上下文网络都以半监督的方式进行联合微调。此外,提出了一种全局-局部特征融合策略,以增强定位能力。MAT-SED在DCASE 2023 task 4上的评估超过了最先进的性能,分别达到0.587 0.896 PSDS 1 PSDS 2。摘要:Sound event detection (SED) methods that leverage a large pre-trained Transformer encoder network have shown promising performance in recent DCASE challenges. However, they still rely on an RNN-based context network to model temporal dependencies, largely due to the scarcity of labeled data. In this work, we propose a pure Transformer-based SED model with masked-reconstruction based pre-training, termed MAT-SED. Specifically, a Transformer with relative positional encoding is first designed as the context network, pre-trained by the masked-reconstruction task on all available target data in a self-supervised way. Both the encoder and the context network are jointly fine-tuned in a semi-supervised manner. Furthermore, a global-local feature fusion strategy is proposed to enhance the localization capability. Evaluation of MAT-SED on DCASE2023 task4 surpasses state-of-the-art performance, achieving 0.587 0.896 PSDS1 PSDS2 respectively.
【2】 HSDreport: Heart Sound Diagnosis with Echocardiography Reports
标题: HUD报告:使用超声心动图报告进行心脏声音诊断
作者:Zihan Zhao,Pingjie Wang,Liudan Zhao,Yuchen Yang,Ya Zhang,Kun Sun,Xin Sun,Xin Zhou,Yu Wang,Yanfeng Wang
链接:点击下载PDF文件
摘要:心音听诊对先天性心脏病的诊断具有重要意义。然而,现有的心音诊断(HSD)任务的方法主要局限于几个固定的类别,框架HSD任务作为一个严格的分类问题,不完全符合医疗实践,只提供有限的信息给医生。此外,这些方法没有利用超声心动图报告,这是诊断相关疾病的金标准。为了应对这一挑战,我们引入了HSDreport,这是HSD的一个新基准,它要求直接利用从听诊获得的心音来预测超声心动图报告。该基准旨在将听诊的便利性与超声心动图报告的全面性相结合。首先,我们为这个基准收集了一个新的数据集,包括2,275个心音样本及其相应的报告。随后,我们开发了一个知识感知的查询为基础的Transformer来处理这个任务。其目的是利用医学预训练模型的能力和大型语言模型(LLM)的内部知识来解决任务的固有复杂性和可变性,从而增强该方法的鲁棒性和科学有效性。此外,我们的实验结果表明,我们的方法显着优于传统的HSD方法和现有的多模态LLM在检测心音中的关键异常。摘要:Heart sound auscultation holds significant importance in the diagnosis of congenital heart disease. However, existing methods for Heart Sound Diagnosis (HSD) tasks are predominantly limited to a few fixed categories, framing the HSD task as a rigid classification problem that does not fully align with medical practice and offers only limited information to physicians. Besides, such methods do not utilize echocardiography reports, the gold standard in the diagnosis of related diseases. To tackle this challenge, we introduce HSDreport, a new benchmark for HSD, which mandates the direct utilization of heart sounds obtained from auscultation to predict echocardiography reports. This benchmark aims to merge the convenience of auscultation with the comprehensive nature of echocardiography reports. First, we collect a new dataset for this benchmark, comprising 2,275 heart sound samples along with their corresponding reports. Subsequently, we develop a knowledge-aware query-based transformer to handle this task. The intent is to leverage the capabilities of medically pre-trained models and the internal knowledge of large language models (LLMs) to address the task's inherent complexity and variability, thereby enhancing the robustness and scientific validity of the method. Furthermore, our experimental results indicate that our method significantly outperforms traditional HSD approaches and existing multimodal LLMs in detecting key abnormalities in heart sounds.
【3】 GAPS: A Large and Diverse Classical Guitar Dataset and Benchmark Transcription Model
标题: GAPS:一个庞大而多样化的古典吉他数据集和基准转录模型
作者:Xavier Riley,Zixun Guo,Drew Edwards,Simon Dixon
备注:ISMIR 2024
链接:点击下载PDF文件
摘要:我们介绍了GAPS(Guitar-Aligned Performance Scores),一个新的古典吉他表演数据集,以及一个基准吉他转录模型,该模型在GuitarSet上在监督和zero-shot设置中实现了最先进的性能。GAPS是最大的真实吉他音频数据集,包含14小时免费提供的音频-乐谱对齐对,由200多名表演者在不同条件下录制,以及高分辨率的音符级对齐和表演视频。这些使我们能够训练一个最先进的模型,用于吉他独奏录音的自动转录,可以很好地推广到训练期间看不到的真实世界音频。摘要:We introduce GAPS (Guitar-Aligned Performance Scores), a new dataset of classical guitar performances, and a benchmark guitar transcription model that achieves state-of-the-art performance on GuitarSet in both supervised and zero-shot settings. GAPS is the largest dataset of real guitar audio, containing 14 hours of freely available audio-score aligned pairs, recorded in diverse conditions by over 200 performers, together with high-resolution note-level MIDI alignments and performance videos. These enable us to train a state-of-the-art model for automatic transcription of solo guitar recordings which can generalise well to real world audio that is unseen during training.
【4】 ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale
标题: ASVspoof 5:众包语音数据、Deepfakes和大规模对抗性攻击
作者:Xin Wang,Hector Delgado,Hemlata Tak,Jee-weon Jung,Hye-jin Shim,Massimiliano Todisco,Ivan Kukanov,Xuechen Liu,Md Sahidullah,Tomi Kinnunen,Nicholas Evans,Kong Aik Lee,Junichi Yamagishi
备注:8 pages, ASVspoof 5 Workshop (Interspeech2024 Satellite)
链接:点击下载PDF文件
摘要:ASVspoof 5是一系列挑战中的第五版,旨在促进语音欺骗和deepfake攻击的研究以及检测解决方案的设计。与之前的挑战相比,ASVspoof 5数据库是根据从不同声学条件下大量扬声器收集的众包数据构建的。使用代理检测模型生成和测试了众包攻击,而对抗性攻击则首次被纳入其中。新的指标支持评估欺骗鲁棒自动说话人验证(SASV)以及独立的检测解决方案,即,没有ASV。我们描述了两个挑战轨道,新的数据库,评估指标,基线和评估平台,并提出了结果的摘要。攻击严重损害了基线系统,而提交则带来了实质性的改进。摘要:ASVspoof 5 is the fifth edition in a series of challenges that promote the study of speech spoofing and deepfake attacks, and the design of detection solutions. Compared to previous challenges, the ASVspoof 5 database is built from crowdsourced data collected from a vastly greater number of speakers in diverse acoustic conditions. Attacks, also crowdsourced, are generated and tested using surrogate detection models, while adversarial attacks are incorporated for the first time. New metrics support the evaluation of spoofing-robust automatic speaker verification (SASV) as well as stand-alone detection solutions, i.e., countermeasures without ASV. We describe the two challenge tracks, the new database, the evaluation metrics, baselines, and the evaluation platform, and present a summary of the results. Attacks significantly compromise the baseline systems, while submissions bring substantial improvements.
【5】 ConcateNet: Dialogue Separation Using Local And Global Feature Concatenation
标题: ContateNet:使用本地和全球功能连锁的对话分离
作者:Mhd Modar Halimeh,Matteo Torcoli,Emanuël Habets
链接:点击下载PDF文件
摘要:对话分离涉及将对话信号从诸如电影或电视节目的混合物中隔离。这可能是为广播相关应用程序启用对话增强的必要步骤。在本文中,ConcateNet的对话分离,这是基于一种新的方法来处理本地和全球的功能,旨在更好地推广域外信号。ConcateNet使用以降噪为重点的公开可用数据集进行训练,并使用三个数据集进行评估:两个以降噪为重点的数据集(域内),显示了ConcateNet的竞争性能,以及一个以广播为重点的数据集(域外),与考虑的最先进的降噪方法相比,验证了所提出的架构的更好的泛化性能。摘要:Dialogue separation involves isolating a dialogue signal from a mixture, such as a movie or a TV program. This can be a necessary step to enable dialogue enhancement for broadcast-related applications. In this paper, ConcateNet for dialogue separation is proposed, which is based on a novel approach for processing local and global features aimed at better generalization for out-of-domain signals. ConcateNet is trained using a noise reduction-focused, publicly available dataset and evaluated using three datasets: two noise reduction-focused datasets (in-domain), which show competitive performance for ConcateNet, and a broadcast-focused dataset (out-of-domain), which verifies the better generalization performance for the proposed architecture compared to considered state-of-the-art noise-reduction methods.
eess.AS音频处理
【1】 ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale标题: ASVspoof 5:众包语音数据、Deepfakes和大规模对抗性攻击
作者:Xin Wang,Hector Delgado,Hemlata Tak,Jee-weon Jung,Hye-jin Shim,Massimiliano Todisco,Ivan Kukanov,Xuechen Liu,Md Sahidullah,Tomi Kinnunen,Nicholas Evans,Kong Aik Lee,Junichi Yamagishi
备注:8 pages, ASVspoof 5 Workshop (Interspeech2024 Satellite)
链接:点击下载PDF文件
摘要:ASVspoof 5是一系列挑战中的第五版,旨在促进语音欺骗和deepfake攻击的研究以及检测解决方案的设计。与之前的挑战相比,ASVspoof 5数据库是根据从不同声学条件下大量扬声器收集的众包数据构建的。使用代理检测模型生成和测试了众包攻击,而对抗性攻击则首次被纳入其中。新的指标支持评估欺骗鲁棒自动说话人验证(SASV)以及独立的检测解决方案,即,没有ASV。我们描述了两个挑战轨道,新的数据库,评估指标,基线和评估平台,并提出了结果的摘要。攻击严重损害了基线系统,而提交则带来了实质性的改进。摘要:ASVspoof 5 is the fifth edition in a series of challenges that promote the study of speech spoofing and deepfake attacks, and the design of detection solutions. Compared to previous challenges, the ASVspoof 5 database is built from crowdsourced data collected from a vastly greater number of speakers in diverse acoustic conditions. Attacks, also crowdsourced, are generated and tested using surrogate detection models, while adversarial attacks are incorporated for the first time. New metrics support the evaluation of spoofing-robust automatic speaker verification (SASV) as well as stand-alone detection solutions, i.e., countermeasures without ASV. We describe the two challenge tracks, the new database, the evaluation metrics, baselines, and the evaluation platform, and present a summary of the results. Attacks significantly compromise the baseline systems, while submissions bring substantial improvements.
【2】 ConcateNet: Dialogue Separation Using Local And Global Feature Concatenation
标题: ContateNet:使用本地和全球功能连锁的对话分离
作者:Mhd Modar Halimeh,Matteo Torcoli,Emanuël Habets
链接:点击下载PDF文件
摘要:对话分离涉及将对话信号从诸如电影或电视节目的混合物中隔离。这可能是为广播相关应用程序启用对话增强的必要步骤。在本文中,ConcateNet的对话分离,这是基于一种新的方法来处理本地和全球的功能,旨在更好地推广域外信号。ConcateNet使用以降噪为重点的公开可用数据集进行训练,并使用三个数据集进行评估:两个以降噪为重点的数据集(域内),显示了ConcateNet的竞争性能,以及一个以广播为重点的数据集(域外),与考虑的最先进的降噪方法相比,验证了所提出的架构的更好的泛化性能。摘要:Dialogue separation involves isolating a dialogue signal from a mixture, such as a movie or a TV program. This can be a necessary step to enable dialogue enhancement for broadcast-related applications. In this paper, ConcateNet for dialogue separation is proposed, which is based on a novel approach for processing local and global features aimed at better generalization for out-of-domain signals. ConcateNet is trained using a noise reduction-focused, publicly available dataset and evaluated using three datasets: two noise reduction-focused datasets (in-domain), which show competitive performance for ConcateNet, and a broadcast-focused dataset (out-of-domain), which verifies the better generalization performance for the proposed architecture compared to considered state-of-the-art noise-reduction methods.
【3】 MAT-SED: AMasked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
标题: MAT-MED:AMasked音频Transformer,具有基于掩蔽重建的预训练,用于声音事件检测
作者:Pengfei Cai,Yan Song,Kang Li,Haoyu Song,Ian McLoughlin
备注:Received by interspeech 2024
链接:点击下载PDF文件
摘要:利用大型预训练的Transformer编码器网络的声音事件检测(SED)方法在最近的DCASE挑战中显示出有希望的性能。然而,他们仍然依赖于基于RNN的上下文网络来建模时间依赖性,这主要是由于标记数据的稀缺性。在这项工作中,我们提出了一个纯粹的基于transformer的SED模型与掩蔽重建为基础的预训练,称为MAT-SED。具体而言,首先设计具有相对位置编码的Transformer作为上下文网络,通过掩蔽重建任务以自监督的方式对所有可用的目标数据进行预训练。编码器和上下文网络都以半监督的方式进行联合微调。此外,提出了一种全局-局部特征融合策略,以增强定位能力。MAT-SED在DCASE 2023 task 4上的评估超过了最先进的性能,分别达到0.587 0.896 PSDS 1 PSDS 2。摘要:Sound event detection (SED) methods that leverage a large pre-trained Transformer encoder network have shown promising performance in recent DCASE challenges. However, they still rely on an RNN-based context network to model temporal dependencies, largely due to the scarcity of labeled data. In this work, we propose a pure Transformer-based SED model with masked-reconstruction based pre-training, termed MAT-SED. Specifically, a Transformer with relative positional encoding is first designed as the context network, pre-trained by the masked-reconstruction task on all available target data in a self-supervised way. Both the encoder and the context network are jointly fine-tuned in a semi-supervised manner. Furthermore, a global-local feature fusion strategy is proposed to enhance the localization capability. Evaluation of MAT-SED on DCASE2023 task4 surpasses state-of-the-art performance, achieving 0.587 0.896 PSDS1 PSDS2 respectively.
【4】 HSDreport: Heart Sound Diagnosis with Echocardiography Reports
标题: HUD报告:使用超声心动图报告进行心脏声音诊断
作者:Zihan Zhao,Pingjie Wang,Liudan Zhao,Yuchen Yang,Ya Zhang,Kun Sun,Xin Sun,Xin Zhou,Yu Wang,Yanfeng Wang
链接:点击下载PDF文件
摘要:心音听诊对先天性心脏病的诊断具有重要意义。然而,现有的心音诊断(HSD)任务的方法主要局限于几个固定的类别,框架HSD任务作为一个严格的分类问题,不完全符合医疗实践,只提供有限的信息给医生。此外,这些方法没有利用超声心动图报告,这是诊断相关疾病的金标准。为了应对这一挑战,我们引入了HSDreport,这是HSD的一个新基准,它要求直接利用从听诊获得的心音来预测超声心动图报告。该基准旨在将听诊的便利性与超声心动图报告的全面性相结合。首先,我们为这个基准收集了一个新的数据集,包括2,275个心音样本及其相应的报告。随后,我们开发了一个知识感知的查询为基础的Transformer来处理这个任务。其目的是利用医学预训练模型的能力和大型语言模型(LLM)的内部知识来解决任务的固有复杂性和可变性,从而增强该方法的鲁棒性和科学有效性。此外,我们的实验结果表明,我们的方法显着优于传统的HSD方法和现有的多模态LLM在检测心音中的关键异常。摘要:Heart sound auscultation holds significant importance in the diagnosis of congenital heart disease. However, existing methods for Heart Sound Diagnosis (HSD) tasks are predominantly limited to a few fixed categories, framing the HSD task as a rigid classification problem that does not fully align with medical practice and offers only limited information to physicians. Besides, such methods do not utilize echocardiography reports, the gold standard in the diagnosis of related diseases. To tackle this challenge, we introduce HSDreport, a new benchmark for HSD, which mandates the direct utilization of heart sounds obtained from auscultation to predict echocardiography reports. This benchmark aims to merge the convenience of auscultation with the comprehensive nature of echocardiography reports. First, we collect a new dataset for this benchmark, comprising 2,275 heart sound samples along with their corresponding reports. Subsequently, we develop a knowledge-aware query-based transformer to handle this task. The intent is to leverage the capabilities of medically pre-trained models and the internal knowledge of large language models (LLMs) to address the task's inherent complexity and variability, thereby enhancing the robustness and scientific validity of the method. Furthermore, our experimental results indicate that our method significantly outperforms traditional HSD approaches and existing multimodal LLMs in detecting key abnormalities in heart sounds.
【5】 GAPS: A Large and Diverse Classical Guitar Dataset and Benchmark Transcription Model
标题: GAPS:一个庞大而多样化的古典吉他数据集和基准转录模型
作者:Xavier Riley,Zixun Guo,Drew Edwards,Simon Dixon
备注:ISMIR 2024
链接:点击下载PDF文件
摘要:我们介绍了GAPS(Guitar-Aligned Performance Scores),一个新的古典吉他表演数据集,以及一个基准吉他转录模型,该模型在GuitarSet上在监督和zero-shot设置中实现了最先进的性能。GAPS是最大的真实吉他音频数据集,包含14小时免费提供的音频-乐谱对齐对,由200多名表演者在不同条件下录制,以及高分辨率的音符级对齐和表演视频。这些使我们能够训练一个最先进的模型,用于吉他独奏录音的自动转录,可以很好地推广到训练期间看不到的真实世界音频。摘要:We introduce GAPS (Guitar-Aligned Performance Scores), a new dataset of classical guitar performances, and a benchmark guitar transcription model that achieves state-of-the-art performance on GuitarSet in both supervised and zero-shot settings. GAPS is the largest dataset of real guitar audio, containing 14 hours of freely available audio-score aligned pairs, recorded in diverse conditions by over 200 performers, together with high-resolution note-level MIDI alignments and performance videos. These enable us to train a state-of-the-art model for automatic transcription of solo guitar recordings which can generalise well to real world audio that is unseen during training.
机器翻译,仅供参考
