今日论文合集:cs.SD语音10篇,eess.AS音频处理10篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音

【1】 Investigating the Effectiveness of Explainability Methods in Parkinson's  Detection from Speech

标题:研究言语检测帕金森病的可解释性方法的有效性
链接:https://arxiv.org/abs/2411.08013
作者:Eleonora Mancini,  Francesco Paissan,  Paolo Torroni,  Cem Subakan,  Mirco Ravanelli
备注:The first two authors contributed equally to this research: author order is alphabetical
摘要:帕金森病(PD)的言语障碍为诊断提供了重要的早期指标。虽然基于语音的PD检测模型表现出了很强的性能,但其可解释性仍有待探索。本研究系统地评估了几种识别PD特异性语音特征的可解释性方法,旨在支持开发准确,可解释的模型,用于PD诊断和监测的临床决策。我们的方法包括:(i)使用主流的可解释性技术获得属性和显着性图,(ii)通过一系列已建立的度量标准,定量评估这些图及其组合的忠实性,以及(iii)评估显着性图所传达的信息,用于辅助分类器的PD检测。我们的研究结果表明,虽然解释与分类器一致,但它们往往无法为领域专家提供有价值的信息。
摘要:Speech impairments in Parkinson's disease (PD) provide significant earlyindicators for diagnosis. While models for speech-based PD detection have shownstrong performance, their interpretability remains underexplored. This studysystematically evaluates several explainability methods to identify PD-specificspeech features, aiming to support the development of accurate, interpretablemodels for clinical decision-making in PD diagnosis and monitoring. Ourmethodology involves (i) obtaining attributions and saliency maps usingmainstream interpretability techniques, (ii) quantitatively evaluating thefaithfulness of these maps and their combinations obtained via union andintersection through a range of established metrics, and (iii) assessing theinformation conveyed by the saliency maps for PD detection from an auxiliaryclassifier. Our results reveal that, while explanations are aligned with theclassifier, they often fail to provide valuable information for domain experts.

【2】 Automatic Album Sequencing
标题:自动专辑排序
链接:https://arxiv.org/abs/2411.07772
作者:Vincent Herrmann,  Dylan R. Ashley,  Jürgen Schmidhuber
备注:presented as a late breaking demo in the 25th International Society for Music Information Retrieval Conference; 3 pages in main text, 3 figures in main text; source code available at this https URL
摘要:专辑排序是专辑制作过程中至关重要的一部分。最近,提出了一种数据驱动的方法,通过提取集合中项目的叙事本质来对独立媒体的一般集合进行排序。虽然这种方法意味着一种专辑排序技术,但它并不广泛适用于技术含量较低的受众,需要先进的机器学习技术知识才能使用。为了解决这个问题,我们引入了一个新的用户友好的基于Web的工具,允许技术含量较低的观众上传音乐曲目,只需单击一下即可执行此技术,随后将结果以清晰的可视化方式呈现给用户。为了增加用户可用的模板数量并解决以前工作的缺点,我们还引入了一种新的直接基于transformer的相册排序方法。我们发现,我们更直接的方法优于随机基线,但没有达到相同的性能作为叙事本质的方法。这两种方法都包含在我们基于Web的用户界面中,并且可以在https://github.com/dylanashley/automatic-album-sequencing上公开获得该界面以及我们实现的完整副本
摘要:Album sequencing is a critical part of the album production process.Recently, a data-driven approach was proposed that sequences generalcollections of independent media by extracting the narrative essence of theitems in the collections. While this approach implies an album sequencingtechnique, it is not widely accessible to a less technical audience, requiringadvanced knowledge of machine learning techniques to use. To address this, weintroduce a new user-friendly web-based tool that allows a less technicalaudience to upload music tracks, execute this technique in one click, andsubsequently presents the result in a clean visualization to the user. To bothincrease the number of templates available to the user and address shortcomingsof previous work, we also introduce a new direct transformer-based albumsequencing method. We find that our more direct method outperforms a randombaseline but does not reach the same performance as the narrative essenceapproach. Both methods are included in our web-based user interface, and this-- alongside a full copy of our implementation -- is publicly available athttps://github.com/dylanashley/automatic-album-sequencing

【3】 SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State  Space Model
标题:SAV-SE:采用选择性状态空间模型的场景感知视听语音增强
链接:https://arxiv.org/abs/2411.07751
作者:Xinyuan Qian,  Jiaran Gao,  Yaodan Zhang,  Qiquan Zhang,  Hexin Liu,  Leibny Paola Garcia,  Haizhou Li
摘要:语音增强在各种应用中起着至关重要的作用,而视觉信息的整合已被证明会带来实质性的优势。然而,目前的大部分研究集中在面部和嘴唇运动的检查,这可能是妥协或完全无法访问的情况下,发生闭塞或当相机视图是遥远的。然而,来自周围环境的上下文视觉线索却被忽视了:例如,当我们看到狗叫时,我们的大脑具有识别和过滤吠叫噪音的先天能力。为此,在本文中,我们介绍了一种新的任务,即SAV-SE。据我们所知,这是第一个使用来自同步视频的丰富上下文信息作为辅助线索来指示噪声类型的建议,这最终提高了语音增强性能。具体来说,我们提出了VC-S$^2$E方法,该方法结合了Conformer和Mamba模块,以实现其互补优势。在公开的MUSIC、AVSpeech和AudioSet数据集上进行了大量的实验,结果表明VC-S$^2$E优于其他竞争方法。我们将公开源代码。项目演示页面:https://AVSEPage.github.io/
摘要:Speech enhancement plays an essential role in various applications, and theintegration of visual information has been demonstrated to bring substantialadvantages. However, the majority of current research concentrates on theexamination of facial and lip movements, which can be compromised or entirelyinaccessible in scenarios where occlusions occur or when the camera view isdistant. Whereas contextual visual cues from the surrounding environment havebeen overlooked: for example, when we see a dog bark, our brain has the innateability to discern and filter out the barking noise. To this end, in thispaper, we introduce a novel task, i.e. SAV-SE. To our best knowledge, this isthe first proposal to use rich contextual information from synchronized videoas auxiliary cues to indicate the type of noise, which eventually improves thespeech enhancement performance. Specifically, we propose the VC-S$^2$E method,which incorporates the Conformer and Mamba modules for their complementarystrengths. Extensive experiments are conducted on public MUSIC, AVSpeech andAudioSet datasets, where the results demonstrate the superiority of VC-S$^2$Eover other competitive methods. We will make the source code publiclyavailable. Project demo page: https://AVSEPage.github.io/

【4】 Understanding Audiovisual Deepfake Detection: Techniques, Challenges,  Human Factors and Perceptual Insights
标题:了解视听Deepfake检测:技术、挑战、人为因素和感知洞察
链接:https://arxiv.org/abs/2411.07650
作者:Ammarah Hashmi,  Sahibzada Adil Shahzad,  Chia-Wen Lin,  Yu Tsao,  Hsin-Min Wang
摘要:深度学习已成功应用于各个领域,其对deepfake检测的影响也不例外。Deepfakes是虚假但真实的合成内容,可以欺骗性地用于政治模仿,网络钓鱼,诽谤或传播错误信息。尽管对单峰深度伪造检测进行了广泛的研究,但通过音频和视频流的联合分析来识别复杂的深度伪造仍然相对未被探索。为了填补这一空白,本调查首先概述了视听Deepfake生成技术,应用及其后果,然后全面回顾了结合音频和视觉模式以提高检测准确性的最新方法,总结并批判性地分析了它们的优势和局限性。此外,我们还讨论了现有的开源数据集,以便更深入地了解,这可以为研究社区做出贡献,并为想要分析基于深度学习的视频取证视听方法的初学者提供必要的信息。通过弥合单模态和多模态方法之间的差距,本文旨在提高deepfake检测策略的有效性,并指导未来在网络安全和媒体完整性方面的研究。
摘要:Deep Learning has been successfully applied in diverse fields, and its impacton deepfake detection is no exception. Deepfakes are fake yet realisticsynthetic content that can be used deceitfully for political impersonation,phishing, slandering, or spreading misinformation. Despite extensive researchon unimodal deepfake detection, identifying complex deepfakes through jointanalysis of audio and visual streams remains relatively unexplored. To fillthis gap, this survey first provides an overview of audiovisual deepfakegeneration techniques, applications, and their consequences, and then providesa comprehensive review of state-of-the-art methods that combine audio andvisual modalities to enhance detection accuracy, summarizing and criticallyanalyzing their strengths and limitations. Furthermore, we discuss existingopen source datasets for a deeper understanding, which can contribute to theresearch community and provide necessary information to beginners who want toanalyze deep learning-based audiovisual methods for video forensics. Bybridging the gap between unimodal and multimodal approaches, this paper aims toimprove the effectiveness of deepfake detection strategies and guide futureresearch in cybersecurity and media integrity.

【5】 AuscultaBase: A Foundational Step Towards AI-Powered Body Sound  Diagnostics
标题:AuscutaBase:迈向人工智能驱动的身体声音诊断的基础一步
链接:https://arxiv.org/abs/2411.07547
作者:Pingjie Wang,  Zihan Zhao,  Liudan Zhao,  Miao He,  Xin Sun,  Ya Zhang,  Kun Sun,  Yanfeng Wang,  Yu Wang
备注:26 pages
摘要:内部身体声音的听诊对于诊断一系列健康状况至关重要,但其有效性通常受到临床医生的专业知识和人类听力的声学限制的限制,限制了其在各种临床场景中的使用。为了应对这些挑战,我们引入了AuscultaBase,这是一个基础框架,旨在通过创新的数据集成和对比学习技术来推进身体声音诊断。我们的贡献包括以下几点:首先,我们编译了AuscultaBase-Corpus,这是一个大规模的多源身体声音数据库,包含11个数据集,40,317个音频记录,总计322.4小时的心脏,肺和肠道声音。其次,我们开发了AuscultaBase-Model,一个身体声音的基础诊断模型,利用编译语料库上的对比学习。第三,我们建立了AuscultaBase-Bench,一个包含16个子任务的综合基准,评估各种开源声学预训练模型的性能。评估结果表明,我们的模型在16个任务中的12个任务中优于所有其他开源模型,证明了我们的方法在提高身体声音分析诊断能力方面的有效性。
摘要:Auscultation of internal body sounds is essential for diagnosing a range ofhealth conditions, yet its effectiveness is often limited by clinicians'expertise and the acoustic constraints of human hearing, restricting its useacross various clinical scenarios. To address these challenges, we introduceAuscultaBase, a foundational framework aimed at advancing body sounddiagnostics through innovative data integration and contrastive learningtechniques. Our contributions include the following: First, we compileAuscultaBase-Corpus, a large-scale, multi-source body sound databaseencompassing 11 datasets with 40,317 audio recordings and totaling 322.4 hoursof heart, lung, and bowel sounds. Second, we develop AuscultaBase-Model, afoundational diagnostic model for body sounds, utilizing contrastive learningon the compiled corpus. Third, we establish AuscultaBase-Bench, a comprehensivebenchmark containing 16 sub-tasks, assessing the performance of variousopen-source acoustic pre-trained models. Evaluation results indicate that ourmodel outperforms all other open-source models in 12 out of 16 tasks,demonstrating the efficacy of our approach in advancing diagnostic capabilitiesfor body sound analysis.

【6】 Music Discovery Dialogue Generation Using Human Intent Analysis and  Large Language Models
标题:使用人类意图分析和大型语言模型的音乐发现对话生成
链接:https://arxiv.org/abs/2411.07439
作者:SeungHeon Doh,  Keunwoo Choi,  Daeyong Kwon,  Taesu Kim,  Juhan Nam
备注:Accepted for publication at the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
摘要:对话式音乐检索系统可以帮助用户通过对话发现符合他们偏好的音乐。为了实现这一点,会话式音乐检索系统应该通过1)理解用户查询和2)用自然语言和检索到的音乐进行响应来无缝地参与多轮会话。一个简单的解决方案是利用这种对话日志的数据驱动方法。然而,很少有数据集可用于研究,并且在数量和质量方面受到限制。在本文中,我们提出了一个数据生成框架,丰富的音乐发现对话使用大型语言模型(LLM)和用户意图,系统动作和音乐属性。这是通过i)使用扎根理论的对话意图分析,ii)经由级联数据库过滤生成属性序列,以及iii)使用大型语言模型生成话语来完成的。通过将此框架应用于Million Song数据集,我们创建了LP-MusicDialog,这是一个基于大型语言模型的伪音乐对话数据集,包含超过288 k的音乐对话,使用超过319 k的音乐项目。我们的评估表明,该合成数据集在对话一致性、项目相关性和自然性方面与现有的小型人类对话数据集具有竞争力。此外,使用该数据集,我们训练会话音乐检索模型,并显示出可喜的结果。
摘要:A conversational music retrieval system can help users discover music thatmatches their preferences through dialogue. To achieve this, a conversationalmusic retrieval system should seamlessly engage in multi-turn conversation by1) understanding user queries and 2) responding with natural language andretrieved music. A straightforward solution would be a data-driven approachutilizing such conversation logs. However, few datasets are available for theresearch and are limited in terms of volume and quality. In this paper, wepresent a data generation framework for rich music discovery dialogue using alarge language model (LLM) and user intents, system actions, and musicalattributes. This is done by i) dialogue intent analysis using grounded theory,ii) generating attribute sequences via cascading database filtering, and iii)generating utterances using large language models. By applying this frameworkto the Million Song dataset, we create LP-MusicDialog, a Large Language Modelbased Pseudo Music Dialogue dataset, containing over 288k music conversationsusing more than 319k music items. Our evaluation shows that the syntheticdataset is competitive with an existing, small human dialogue dataset in termsof dialogue consistency, item relevance, and naturalness. Furthermore, usingthe dataset, we train a conversational music retrieval model and show promisingresults.

【7】 Just Label the Repeats for In-The-Wild Audio-to-Score Alignment
标题:只需标记重复点以实现野外音频与分数对齐
链接:https://arxiv.org/abs/2411.07428
作者:Irmak Bukey,  Michael Feffer,  Chris Donahue
备注:25th International Society for Music Information Retrieval Conference, San Francisco, 2024
摘要:我们提出了一种高效的工作流程,用于高质量的离线对齐野外表演音频和相应的乐谱扫描(图像)。最近的工作音频到分数对齐扩展动态时间规整(DTW),理论上能够处理由重复符号引起的乐谱跳跃,这种方法不需要人工注释,但我们表明,它往往会产生低质量的对齐。作为替代方案,我们提出了一个工作流程和界面,允许用户快速注释跳转(通过点击重复标志),需要少量的人工监督,但平均产生更高质量的对齐。此外,我们通过以下方式改进音频和乐谱特征表示以提高对齐质量:(1)将测量检测集成到乐谱特征表示中,以及(2)使用来自音乐转录模型而不是钢琴卷的原始起始预测概率。我们提出了一个评估协议的音频到分数的对齐,计算估计和地面实况之间的距离对齐措施的单位。在此评估下,我们发现我们提出的跳跃注释工作流程和改进的特征表示一起将对齐精度提高了150%,相对于先前的工作(33%至82%)。
摘要:We propose an efficient workflow for high-quality offline alignment ofin-the-wild performance audio and corresponding sheet music scans (images).Recent work on audio-to-score alignment extends dynamic time warping (DTW) tobe theoretically able to handle jumps in sheet music induced by repeatsigns-this method requires no human annotations, but we show that it oftenyields low-quality alignments. As an alternative, we propose a workflow andinterface that allows users to quickly annotate jumps (by clicking on repeatsigns), requiring a small amount of human supervision but yielding much higherquality alignments on average. Additionally, we refine audio and score featurerepresentations to improve alignment quality by: (1) integrating measuredetection into the score feature representation, and (2) using raw onsetprediction probabilities from a music transcription model instead of pianoroll. We propose an evaluation protocol for audio-to-score alignment thatcomputes the distance between the estimated and ground truth alignment in unitsof measures. Under this evaluation, we find that our proposed jump annotationworkflow and improved feature representations together improve alignmentaccuracy by 150% relative to prior work (33% to 82%).

【8】 CJST: CTC Compressor based Joint Speech and Text Training for  Decoder-Only ASR
标题:CJST:针对仅解码器的ASB的基于CIC压缩器的联合语音和文本训练
链接:https://arxiv.org/abs/2411.07607
作者:Wei Zhou,  Junteng Jia,  Leda Sari,  Jay Mahadeokar,  Ozlem Kalinli
备注:submitted to ICASSP2025
摘要:CTC压缩器是一种将音频编码器集成到仅解码器模型中的有效方法,在不同的语音应用中得到了越来越多的关注。在这项工作中,我们提出了一种新的CTC压缩器为基础的联合语音和文本训练(CJST)框架解码器只ASR。CJST通过探索一个简单的模态适配器和CTC压缩器的几个功能,包括序列压缩、动态强制峰值对齐和CTC类嵌入,从两个方向匹配语音和文本模态。在Librispeech和TED-LIUM 2语料库上的实验结果表明,CJST在不需要持续时间处理的情况下实现了有效的文本注入,在域内和跨域情况下都具有最佳性能。我们还提供了对CTC压缩器的全面研究,涵盖了各种压缩模式,边缘情况处理以及在干净和嘈杂数据条件下的行为,这揭示了将CTC压缩器用于仅解码器模型的最稳健设置。
摘要:CTC compressor can be an effective approach to integrate audio encoders todecoder-only models, which has gained growing interest for different speechapplications. In this work, we propose a novel CTC compressor based jointspeech and text training (CJST) framework for decoder-only ASR. CJST matchesspeech and text modalities from both directions by exploring a simple modalityadaptor and several features of the CTC compressor, including sequencecompression, on-the-fly forced peaky alignment and CTC class embeddings.Experimental results on the Librispeech and TED-LIUM2 corpora show that theproposed CJST achieves an effective text injection without the need of durationhandling, leading to the best performance for both in-domain and cross-domainscenarios. We also provide a comprehensive study on CTC compressor, coveringvarious compression modes, edge case handling and behavior under both clean andnoisy data conditions, which reveals the most robust setting to use CTCcompressor for decoder-only models.

【9】 SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with  Silhouettes
标题:SoundSil-DS:用剪影对场图像进行深度去噪和分割
链接:https://arxiv.org/abs/2411.07517
作者:Risako Tanigawa,  Kenji Ishikawa,  Noboru Harada,  Yasuhiro Oikawa
备注:13 pages, 12 figures, 5 tables. Accepted by WACV 2025
摘要:光学技术的发展使得能够对二维(2D)声场进行成像。这种声光传感能够理解声音和物体之间的相互作用,例如反射和衍射。此外,它有望用于自动驾驶车辆和辅助机器人的声纳的先进测量技术。然而,声光传感的低声压灵敏度导致图像上的高强度噪声。因此,去噪是声场可视化和分析的一项重要任务。除了去噪之外,还需要分割声音和物体轮廓以分析它们之间的相互作用。在本文中,我们提出了声场图像与对象轮廓去噪和分割(SoundSil-DS),共同执行去噪和分割的声场和对象轮廓上的可视化图像。我们基于当前最先进的去噪网络开发了一种新模型。我们还创建了一个数据集,通过声学仿真来训练和评估所提出的方法。所提出的方法进行了评估,使用模拟和测量数据。我们证实了我们的方法可以应用于实验测量数据。这些结果表明,所提出的方法可以改善声场的后处理,例如基于物理模型的三维重建,因为它可以去除不需要的噪声并分离声场和其他对象轮廓。我们的代码可从https://github.com/nttcslab/soundsil-ds获得。
摘要:Development of optical technology has enabled imaging of two-dimensional (2D)sound fields. This acousto-optic sensing enables understanding of theinteraction between sound and objects such as reflection and diffraction.Moreover, it is expected to be used an advanced measurement technology forsonars in self-driving vehicles and assistive robots. However, the lowsound-pressure sensitivity of the acousto-optic sensing results in highintensity of noise on images. Therefore, denoising is an essential task tovisualize and analyze the sound fields. In addition to denoising, segmentationof sound and object silhouette is also required to analyze interactions betweenthem. In this paper, we propose sound-field-images-with-object-silhouettedenoising and segmentation (SoundSil-DS) that jointly perform denoising andsegmentation for sound fields and object silhouettes on a visualized image. Wedeveloped a new model based on the current state-of-the-art denoising network.We also created a dataset to train and evaluate the proposed method throughacoustic simulation. The proposed method was evaluated using both simulated andmeasured data. We confirmed that our method can applied to experimentallymeasured data. These results suggest that the proposed method may improve thepost-processing for sound fields, such as physical model-basedthree-dimensional reconstruction since it can remove unwanted noise andseparate sound fields and other object silhouettes. Our code is available athttps://github.com/nttcslab/soundsil-ds.

【10】 AEROMamba: An efficient architecture for audio super-resolution using  generative adversarial networks and state space models
标题:AEROMamba:使用生成式对抗网络和状态空间模型的音频超分辨率高效架构
链接:https://arxiv.org/abs/2411.07364
作者:Wallace Abreu,  Luiz Wagner Pereira Biscainho
备注:Accepted at LAMIR 2024 Workshop (ISMIR 2024 Satellite Event)
摘要:音频超分辨率旨在通过创建高频内容来增强低分辨率信号。在这项工作中,我们修改了架构的AERO(一个国家的最先进的系统,这项任务)的音乐超分辨率。特别是,我们在所有网络层中将其原始的Attention和LSTM层替换为Mamba,一种状态空间模型(SSM)。Mamba能够有效地替代上述模块,因为它提供了一种类似于Attention的机制,同时也作为一个循环网络。使用拟议的AEROMamba,训练需要的GPU内存减少2- 4倍,因为Mamba利用了卷积公式并利用了GPU内存层次结构。此外,在推理过程中,由于递归,Mamba在恒定的内存中运行,避免了与注意力相关的内存增长。这导致14倍的速度提高,使用5倍的GPU。主观听力测试(0 ~ 100标度)表明,该模型优于AERO模型。在MUSDB数据集中,退化信号的得分为38.22,而AERO和AEROMamba的得分分别为60.03和66.74。对于PianoEval数据集,降级信号的评分为72.92,AERO为76.89,AEROMamba为84.41。
摘要:Audio super-resolution aims to enhance low-resolution signals by creatinghigh-frequency content. In this work, we modify the architecture of AERO (astate-of-the-art system for this task) for music super-resolution.SPecifically, we replace its original Attention and LSTM layers with Mamba, aState Space Model (SSM), across all network layers. Mamba is capable ofeffectively substituting the mentioned modules, as it offers a mechanismsimilar to that of Attention while also functioning as a recurrent network.With the proposed AEROMamba, training requires 2-4x less GPU memory, sinceMamba exploits the convolutional formulation and leverages GPU memoryhierarchy. Additionally, during inference, Mamba operates in constant memorydue to recurrence, avoiding memory growth associated with Attention. Thisresults in a 14x speed improvement using 5x less GPU. Subjective listeningtests (0 to 100 scale) show that the proposed model surpasses the AERO model.In the MUSDB dataset, degraded signals scored 38.22, while AERO and AEROMambascored 60.03 and 66.74, respectively. For the PianoEval dataset, scores were72.92 for degraded signals, 76.89 for AERO, and 84.41 for AEROMamba.

eess.AS音频处理

【1】 Study on Inter and Intra Speaker Variability in Speaker Recognition
标题:说话人识别中说话人间和内变异性的研究
链接:https://arxiv.org/abs/2411.07754
作者:Anton Okhotnikov,  Nikita Torgashov,  Ivan Yakovlev,  Pavel Malov,  Rostislav Makarov
摘要:说话人的数量和他们的时间变化性(或会话多样性)之间的权衡的优化是至关重要的说话人识别系统的发展,从时间的角度来看,使数据收集过程可行。在这篇文章中,我们使用VoxTube数据集进行文本无关说话人识别任务,为基于现代神经网络的说话人识别系统提供了训练数据中说话人间和说话人内变异性之间的依赖性分析。此外,这项工作的辅助贡献是在VoxTube数据集中发布每个话语的上传日期元数据。我们希望这篇文章有助于从媒体托管平台收集和过滤数据的指导方针和最佳实践,以促进研究人员在开发说话人识别系统方面的努力。
摘要:Optimization of a trade-off between the number of speakers and their temporalvariability (or session diversity) is crucial for the development of a speakerrecognition system together with making the data collection process feasiblefrom a time perspective. In this article, we provide the analysis of dependencybetween inter and intra speaker variability in training data for the modernneural network-based speaker recognition system using the VoxTube dataset fortext-independent speaker recognition task. Besides, an auxiliary contributionof this work is a release of upload date metadata per utterance in a VoxTubedataset. We want this article to contribute to guidelines and best practicesfor collecting and filtering data from media hosting platforms to facilitatethe efforts of researchers in developing speaker recognition systems.

【2】 CJST: CTC Compressor based Joint Speech and Text Training for  Decoder-Only ASR
标题:CJST:针对仅解码器的ASB的基于CIC压缩器的联合语音和文本训练
链接:https://arxiv.org/abs/2411.07607
作者:Wei Zhou,  Junteng Jia,  Leda Sari,  Jay Mahadeokar,  Ozlem Kalinli
备注:submitted to ICASSP2025
摘要:CTC压缩器是一种将音频编码器集成到仅解码器模型中的有效方法,在不同的语音应用中得到了越来越多的关注。在这项工作中,我们提出了一种新的CTC压缩器为基础的联合语音和文本训练(CJST)框架解码器只ASR。CJST通过探索一个简单的模态适配器和CTC压缩器的几个功能,包括序列压缩、动态强制峰值对齐和CTC类嵌入,从两个方向匹配语音和文本模态。在Librispeech和TED-LIUM 2语料库上的实验结果表明,CJST在不需要持续时间处理的情况下实现了有效的文本注入,在域内和跨域情况下都具有最佳性能。我们还提供了对CTC压缩器的全面研究,涵盖了各种压缩模式,边缘情况处理以及在干净和嘈杂数据条件下的行为,这揭示了将CTC压缩器用于仅解码器模型的最稳健设置。
摘要:CTC compressor can be an effective approach to integrate audio encoders todecoder-only models, which has gained growing interest for different speechapplications. In this work, we propose a novel CTC compressor based jointspeech and text training (CJST) framework for decoder-only ASR. CJST matchesspeech and text modalities from both directions by exploring a simple modalityadaptor and several features of the CTC compressor, including sequencecompression, on-the-fly forced peaky alignment and CTC class embeddings.Experimental results on the Librispeech and TED-LIUM2 corpora show that theproposed CJST achieves an effective text injection without the need of durationhandling, leading to the best performance for both in-domain and cross-domainscenarios. We also provide a comprehensive study on CTC compressor, coveringvarious compression modes, edge case handling and behavior under both clean andnoisy data conditions, which reveals the most robust setting to use CTCcompressor for decoder-only models.

【3】 SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with  Silhouettes
标题:SoundSil-DS:用剪影对场图像进行深度去噪和分割
链接:https://arxiv.org/abs/2411.07517
作者:Risako Tanigawa,  Kenji Ishikawa,  Noboru Harada,  Yasuhiro Oikawa
备注:13 pages, 12 figures, 5 tables. Accepted by WACV 2025
摘要:光学技术的发展使得能够对二维(2D)声场进行成像。这种声光传感能够理解声音和物体之间的相互作用,例如反射和衍射。此外,它有望用于自动驾驶车辆和辅助机器人的声纳的先进测量技术。然而,声光传感的低声压灵敏度导致图像上的高强度噪声。因此,去噪是声场可视化和分析的一项重要任务。除了去噪之外,还需要分割声音和物体轮廓以分析它们之间的相互作用。在本文中,我们提出了声场图像与对象轮廓去噪和分割(SoundSil-DS),共同执行去噪和分割的声场和对象轮廓上的可视化图像。我们基于当前最先进的去噪网络开发了一种新模型。我们还创建了一个数据集,通过声学仿真来训练和评估所提出的方法。所提出的方法进行了评估,使用模拟和测量数据。我们证实了我们的方法可以应用于实验测量数据。这些结果表明,所提出的方法可以改善声场的后处理,例如基于物理模型的三维重建,因为它可以去除不需要的噪声并分离声场和其他对象轮廓。我们的代码可从https://github.com/nttcslab/soundsil-ds获得。
摘要:Development of optical technology has enabled imaging of two-dimensional (2D)sound fields. This acousto-optic sensing enables understanding of theinteraction between sound and objects such as reflection and diffraction.Moreover, it is expected to be used an advanced measurement technology forsonars in self-driving vehicles and assistive robots. However, the lowsound-pressure sensitivity of the acousto-optic sensing results in highintensity of noise on images. Therefore, denoising is an essential task tovisualize and analyze the sound fields. In addition to denoising, segmentationof sound and object silhouette is also required to analyze interactions betweenthem. In this paper, we propose sound-field-images-with-object-silhouettedenoising and segmentation (SoundSil-DS) that jointly perform denoising andsegmentation for sound fields and object silhouettes on a visualized image. Wedeveloped a new model based on the current state-of-the-art denoising network.We also created a dataset to train and evaluate the proposed method throughacoustic simulation. The proposed method was evaluated using both simulated andmeasured data. We confirmed that our method can applied to experimentallymeasured data. These results suggest that the proposed method may improve thepost-processing for sound fields, such as physical model-basedthree-dimensional reconstruction since it can remove unwanted noise andseparate sound fields and other object silhouettes. Our code is available athttps://github.com/nttcslab/soundsil-ds.

【4】 AEROMamba: An efficient architecture for audio super-resolution using  generative adversarial networks and state space models
标题:AEROMamba:使用生成式对抗网络和状态空间模型的音频超分辨率高效架构
链接:https://arxiv.org/abs/2411.07364
作者:Wallace Abreu,  Luiz Wagner Pereira Biscainho
备注:Accepted at LAMIR 2024 Workshop (ISMIR 2024 Satellite Event)
摘要:音频超分辨率旨在通过创建高频内容来增强低分辨率信号。在这项工作中,我们修改了架构的AERO(一个国家的最先进的系统,这项任务)的音乐超分辨率。特别是,我们在所有网络层中将其原始的Attention和LSTM层替换为Mamba,一种状态空间模型(SSM)。Mamba能够有效地替代上述模块,因为它提供了一种类似于Attention的机制,同时也作为一个循环网络。使用拟议的AEROMamba,训练需要的GPU内存减少2- 4倍,因为Mamba利用了卷积公式并利用了GPU内存层次结构。此外,在推理过程中,由于递归,Mamba在恒定的内存中运行,避免了与注意力相关的内存增长。这导致14倍的速度提高,使用5倍的GPU。主观听力测试(0 ~ 100标度)表明,该模型优于AERO模型。在MUSDB数据集中,退化信号的得分为38.22,而AERO和AEROMamba的得分分别为60.03和66.74。对于PianoEval数据集,降级信号的评分为72.92,AERO为76.89,AEROMamba为84.41。
摘要:Audio super-resolution aims to enhance low-resolution signals by creatinghigh-frequency content. In this work, we modify the architecture of AERO (astate-of-the-art system for this task) for music super-resolution.SPecifically, we replace its original Attention and LSTM layers with Mamba, aState Space Model (SSM), across all network layers. Mamba is capable ofeffectively substituting the mentioned modules, as it offers a mechanismsimilar to that of Attention while also functioning as a recurrent network.With the proposed AEROMamba, training requires 2-4x less GPU memory, sinceMamba exploits the convolutional formulation and leverages GPU memoryhierarchy. Additionally, during inference, Mamba operates in constant memorydue to recurrence, avoiding memory growth associated with Attention. Thisresults in a 14x speed improvement using 5x less GPU. Subjective listeningtests (0 to 100 scale) show that the proposed model surpasses the AERO model.In the MUSDB dataset, degraded signals scored 38.22, while AERO and AEROMambascored 60.03 and 66.74, respectively. For the PianoEval dataset, scores were72.92 for degraded signals, 76.89 for AERO, and 84.41 for AEROMamba.

【5】 Investigating the Effectiveness of Explainability Methods in Parkinson's  Detection from Speech
标题:研究言语检测帕金森病的可解释性方法的有效性
链接:https://arxiv.org/abs/2411.08013
作者:Eleonora Mancini,  Francesco Paissan,  Paolo Torroni,  Cem Subakan,  Mirco Ravanelli
备注:The first two authors contributed equally to this research: author order is alphabetical
摘要:帕金森病(PD)的言语障碍为诊断提供了重要的早期指标。虽然基于语音的PD检测模型表现出了很强的性能,但其可解释性仍有待探索。本研究系统地评估了几种识别PD特异性语音特征的可解释性方法,旨在支持开发准确,可解释的模型,用于PD诊断和监测的临床决策。我们的方法包括:(i)使用主流的可解释性技术获得属性和显着性图,(ii)通过一系列已建立的度量标准,定量评估这些图及其组合的忠实性,以及(iii)评估显着性图所传达的信息,用于辅助分类器的PD检测。我们的研究结果表明,虽然解释与分类器一致,但它们往往无法为领域专家提供有价值的信息。
摘要:Speech impairments in Parkinson's disease (PD) provide significant earlyindicators for diagnosis. While models for speech-based PD detection have shownstrong performance, their interpretability remains underexplored. This studysystematically evaluates several explainability methods to identify PD-specificspeech features, aiming to support the development of accurate, interpretablemodels for clinical decision-making in PD diagnosis and monitoring. Ourmethodology involves (i) obtaining attributions and saliency maps usingmainstream interpretability techniques, (ii) quantitatively evaluating thefaithfulness of these maps and their combinations obtained via union andintersection through a range of established metrics, and (iii) assessing theinformation conveyed by the saliency maps for PD detection from an auxiliaryclassifier. Our results reveal that, while explanations are aligned with theclassifier, they often fail to provide valuable information for domain experts.

【6】 SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State  Space Model
标题:SAV-SE:采用选择性状态空间模型的场景感知视听语音增强
链接:https://arxiv.org/abs/2411.07751
作者:Xinyuan Qian,  Jiaran Gao,  Yaodan Zhang,  Qiquan Zhang,  Hexin Liu,  Leibny Paola Garcia,  Haizhou Li
摘要:语音增强在各种应用中起着至关重要的作用,而视觉信息的整合已被证明会带来实质性的优势。然而,目前的大部分研究集中在面部和嘴唇运动的检查,这可能会受到损害或完全无法访问的情况下,发生闭塞或当相机视图是遥远的。然而,来自周围环境的上下文视觉线索却被忽视了:例如,当我们看到狗叫时,我们的大脑具有识别和过滤吠叫噪音的先天能力。为此,在本文中,我们介绍了一种新的任务,即SAV-SE。据我们所知,这是第一个建议,使用丰富的上下文信息同步视频作为辅助线索,以指示噪声的类型,这最终提高了语音增强性能。具体来说,我们提出了VC-S$^2$E方法,它结合了Conformer和Mamba模块的互补优势。在公开的MUSIC、AVSpeech和AudioSet数据集上进行了大量的实验,结果表明VC-S$^2$E优于其他竞争方法。我们将公开源代码。项目演示页面:https://AVSEPage.github.io/
摘要:Speech enhancement plays an essential role in various applications, and theintegration of visual information has been demonstrated to bring substantialadvantages. However, the majority of current research concentrates on theexamination of facial and lip movements, which can be compromised or entirelyinaccessible in scenarios where occlusions occur or when the camera view isdistant. Whereas contextual visual cues from the surrounding environment havebeen overlooked: for example, when we see a dog bark, our brain has the innateability to discern and filter out the barking noise. To this end, in thispaper, we introduce a novel task, i.e. SAV-SE. To our best knowledge, this isthe first proposal to use rich contextual information from synchronized videoas auxiliary cues to indicate the type of noise, which eventually improves thespeech enhancement performance. Specifically, we propose the VC-S$^2$E method,which incorporates the Conformer and Mamba modules for their complementarystrengths. Extensive experiments are conducted on public MUSIC, AVSpeech andAudioSet datasets, where the results demonstrate the superiority of VC-S$^2$Eover other competitive methods. We will make the source code publiclyavailable. Project demo page: https://AVSEPage.github.io/

【7】 AuscultaBase: A Foundational Step Towards AI-Powered Body Sound  Diagnostics
标题:AuscutaBase:迈向人工智能驱动的身体声音诊断的基础一步
链接:https://arxiv.org/abs/2411.07547
作者:Pingjie Wang,  Zihan Zhao,  Liudan Zhao,  Miao He,  Xin Sun,  Ya Zhang,  Kun Sun,  Yanfeng Wang,  Yu Wang
备注:26 pages
摘要:内部身体声音的听诊对于诊断一系列健康状况至关重要,但其有效性通常受到临床医生的专业知识和人类听力的声学限制的限制,限制了其在各种临床场景中的使用。为了应对这些挑战,我们引入了AuscultaBase,这是一个基础框架,旨在通过创新的数据集成和对比学习技术来推进身体声音诊断。我们的贡献包括以下几点:首先,我们编译了AuscultaBase-Corpus,这是一个大规模的多源身体声音数据库,包含11个数据集,40,317个音频记录,总计322.4小时的心脏,肺和肠道声音。其次,我们开发了AuscultaBase-Model,一个身体声音的基础诊断模型,利用编译语料库上的对比学习。第三,我们建立了AuscultaBase-Bench,一个包含16个子任务的综合基准,评估各种开源声学预训练模型的性能。评估结果表明,我们的模型在16个任务中的12个任务中优于所有其他开源模型,证明了我们的方法在提高身体声音分析诊断能力方面的有效性。
摘要:Auscultation of internal body sounds is essential for diagnosing a range ofhealth conditions, yet its effectiveness is often limited by clinicians'expertise and the acoustic constraints of human hearing, restricting its useacross various clinical scenarios. To address these challenges, we introduceAuscultaBase, a foundational framework aimed at advancing body sounddiagnostics through innovative data integration and contrastive learningtechniques. Our contributions include the following: First, we compileAuscultaBase-Corpus, a large-scale, multi-source body sound databaseencompassing 11 datasets with 40,317 audio recordings and totaling 322.4 hoursof heart, lung, and bowel sounds. Second, we develop AuscultaBase-Model, afoundational diagnostic model for body sounds, utilizing contrastive learningon the compiled corpus. Third, we establish AuscultaBase-Bench, a comprehensivebenchmark containing 16 sub-tasks, assessing the performance of variousopen-source acoustic pre-trained models. Evaluation results indicate that ourmodel outperforms all other open-source models in 12 out of 16 tasks,demonstrating the efficacy of our approach in advancing diagnostic capabilitiesfor body sound analysis.

【8】 Music Discovery Dialogue Generation Using Human Intent Analysis and  Large Language Models
标题:使用人类意图分析和大型语言模型的音乐发现对话生成
链接:https://arxiv.org/abs/2411.07439
作者:SeungHeon Doh,  Keunwoo Choi,  Daeyong Kwon,  Taesu Kim,  Juhan Nam
备注:Accepted for publication at the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
摘要:对话式音乐检索系统可以帮助用户通过对话发现符合他们偏好的音乐。为了实现这一点,会话式音乐检索系统应该通过1)理解用户查询和2)用自然语言和检索到的音乐进行响应来无缝地参与多轮会话。一个简单的解决方案是利用这种对话日志的数据驱动方法。然而,很少有数据集可用于研究,并且在数量和质量方面受到限制。在本文中,我们提出了一个数据生成框架,丰富的音乐发现对话使用大型语言模型(LLM)和用户意图,系统动作和音乐属性。这是通过i)使用扎根理论的对话意图分析,ii)经由级联数据库过滤生成属性序列,以及iii)使用大型语言模型生成话语来完成的。通过将此框架应用于Million Song数据集,我们创建了LP-MusicDialog,这是一个基于大型语言模型的伪音乐对话数据集,包含超过288 k的音乐对话,使用超过319 k的音乐项目。我们的评估表明,该合成数据集在对话一致性、项目相关性和自然性方面与现有的小型人类对话数据集具有竞争力。此外,使用该数据集,我们训练会话音乐检索模型,并显示出可喜的结果。
摘要:A conversational music retrieval system can help users discover music thatmatches their preferences through dialogue. To achieve this, a conversationalmusic retrieval system should seamlessly engage in multi-turn conversation by1) understanding user queries and 2) responding with natural language andretrieved music. A straightforward solution would be a data-driven approachutilizing such conversation logs. However, few datasets are available for theresearch and are limited in terms of volume and quality. In this paper, wepresent a data generation framework for rich music discovery dialogue using alarge language model (LLM) and user intents, system actions, and musicalattributes. This is done by i) dialogue intent analysis using grounded theory,ii) generating attribute sequences via cascading database filtering, and iii)generating utterances using large language models. By applying this frameworkto the Million Song dataset, we create LP-MusicDialog, a Large Language Modelbased Pseudo Music Dialogue dataset, containing over 288k music conversationsusing more than 319k music items. Our evaluation shows that the syntheticdataset is competitive with an existing, small human dialogue dataset in termsof dialogue consistency, item relevance, and naturalness. Furthermore, usingthe dataset, we train a conversational music retrieval model and show promisingresults.

【9】 Just Label the Repeats for In-The-Wild Audio-to-Score Alignment
标题:只需标记重复点以实现野外音频与分数对齐
链接:https://arxiv.org/abs/2411.07428
作者:Irmak Bukey,  Michael Feffer,  Chris Donahue
备注:25th International Society for Music Information Retrieval Conference, San Francisco, 2024
摘要:我们提出了一种高效的工作流程,用于高质量的离线对齐野外表演音频和相应的乐谱扫描(图像)。最近的工作音频到分数对齐扩展动态时间规整(DTW),理论上能够处理由重复符号引起的乐谱跳跃,这种方法不需要人工注释,但我们表明,它往往会产生低质量的对齐。作为替代方案,我们提出了一个工作流程和界面,允许用户快速注释跳转(通过点击重复标志),需要少量的人工监督,但平均产生更高质量的对齐。此外,我们通过以下方式改进音频和乐谱特征表示以提高对齐质量:(1)将测量检测集成到乐谱特征表示中,以及(2)使用来自音乐转录模型而不是钢琴卷的原始起始预测概率。我们提出了一个评估协议的音频到分数的对齐,计算估计和地面实况之间的距离对齐措施的单位。在此评估下,我们发现我们提出的跳跃注释工作流程和改进的特征表示一起将对齐精度提高了150%,相对于先前的工作(33%至82%)。
摘要:We propose an efficient workflow for high-quality offline alignment ofin-the-wild performance audio and corresponding sheet music scans (images).Recent work on audio-to-score alignment extends dynamic time warping (DTW) tobe theoretically able to handle jumps in sheet music induced by repeatsigns-this method requires no human annotations, but we show that it oftenyields low-quality alignments. As an alternative, we propose a workflow andinterface that allows users to quickly annotate jumps (by clicking on repeatsigns), requiring a small amount of human supervision but yielding much higherquality alignments on average. Additionally, we refine audio and score featurerepresentations to improve alignment quality by: (1) integrating measuredetection into the score feature representation, and (2) using raw onsetprediction probabilities from a music transcription model instead of pianoroll. We propose an evaluation protocol for audio-to-score alignment thatcomputes the distance between the estimated and ground truth alignment in unitsof measures. Under this evaluation, we find that our proposed jump annotationworkflow and improved feature representations together improve alignmentaccuracy by 150% relative to prior work (33% to 82%).

【10】 Isochrony-Controlled Speech-to-Text Translation: A study on translating  from Sino-Tibetan to Indo-European Languages
标题:等时控制语音转文本翻译:汉藏语到印欧语言的翻译研究
链接:https://arxiv.org/abs/2411.07387
作者:Midia Yousefi,  Yao Qian,  Junkun Chen,  Gang Wang,  Yanqing Liu,  Dongmei Wang,  Xiaofei Wang,  Jian Xue
摘要:端到端的语音翻译(ST),将源语言语音直接翻译成目标语言文本,近年来受到了极大的关注。许多ST应用需要严格的长度控制,以确保翻译持续时间与源音频的长度相匹配,包括语音和暂停段。以前的方法通常控制机器翻译模型生成的单词或字符的数量,以近似源句子的长度,而不考虑停顿和语音片段的等时性,因为不同语言之间的持续时间可能会有所不同。为了解决这个问题,我们提出了改进的持续时间对齐组件的序列到序列ST模型。我们的方法控制翻译长度预测的持续时间和停顿的翻译过程中。这是通过向解码器提供定时信息来实现的,确保它在生成翻译时跟踪语音和停顿的剩余持续时间。在CoVoST 2的Zh-En测试集上的评估表明,所提出的等时控制ST实现了0.92语音重叠和8.9 BLEU,与ST基线相比仅下降了1.4 BLEU。
摘要:End-to-end speech translation (ST), which translates source language speechdirectly into target language text, has garnered significant attention inrecent years. Many ST applications require strict length control to ensure thatthe translation duration matches the length of the source audio, including bothspeech and pause segments. Previous methods often controlled the number ofwords or characters generated by the Machine Translation model to approximatethe source sentence's length without considering the isochrony of pauses andspeech segments, as duration can vary between languages. To address this, wepresent improvements to the duration alignment component of oursequence-to-sequence ST model. Our method controls translation length bypredicting the duration of speech and pauses in conjunction with thetranslation process. This is achieved by providing timing information to thedecoder, ensuring it tracks the remaining duration for speech and pauses whilegenerating the translation. The evaluation on the Zh-En test set of CoVoST 2,demonstrates that the proposed Isochrony-Controlled ST achieves 0.92 speechoverlap and 8.9 BLEU, which has only a 1.4 BLEU drop compared to the STbaseline.

机器翻译由腾讯交互翻译提供,仅供参考