今日论文合集:cs.SD语音10篇,eess.AS音频处理10篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】PianoVAM: A Multimodal Piano Performance Dataset
标题:PianoVAM:多模式钢琴演奏数据集
链接:https://arxiv.org/abs/2509.08800

作者:Yonghyun Kim, Junhyung Park, Joonhyung Bae, Kirak Kim, Taegyun Kwon, Alexander Lerch, Juhan Nam
备注:Accepted to the 26th International Society for Music Information Retrieval (ISMIR) Conference, 2025
摘要:音乐表现的多模态性质已经推动了对音乐信息检索(MIR)社区内音频域之外的数据的越来越大的兴趣。本文介绍了PianoVAM,这是一个综合的钢琴演奏数据集,包括视频,音频,视频,手部标志,指法标签和丰富的元数据。该数据集是使用一架更大的钢琴录制的,在业余钢琴家的日常练习中捕捉音频和视频,以及在现实和不同的表演条件下同步的顶视图视频。使用预先训练的手部姿势估计模型和半自动指法注释算法提取手部标志和指法标签。我们讨论了数据收集过程中遇到的挑战和不同模式的调整过程。此外,我们描述了我们的指法标注方法的基础上,从视频中提取的手地标。最后,我们使用PianoVAM数据集提出了纯音频和视听钢琴转录的基准测试结果,并讨论了其他潜在的应用。
摘要:The multimodal nature of music performance has driven increasing interest in data beyond the audio domain within the music information retrieval (MIR) community. This paper introduces PianoVAM, a comprehensive piano performance dataset that includes videos, audio, MIDI, hand landmarks, fingering labels, and rich metadata. The dataset was recorded using a Disklavier piano, capturing audio and MIDI from amateur pianists during their daily practice sessions, alongside synchronized top-view videos in realistic and varied performance conditions. Hand landmarks and fingering labels were extracted using a pretrained hand pose estimation model and a semi-automated fingering annotation algorithm. We discuss the challenges encountered during data collection and the alignment process across different modalities. Additionally, we describe our fingering annotation method based on hand landmarks extracted from videos. Finally, we present benchmarking results for both audio-only and audio-visual piano transcription using the PianoVAM dataset and discuss additional potential applications.


【2】Explainability of CNN Based Classification Models for Acoustic Signal
标题:基于CNN的声信号分类模型的可解释性
链接:https://arxiv.org/abs/2509.08717

作者:Zubair Faruqui, Mackenzie S. McIntire, Rahul Dubey, Jay McEntee
备注:Accepted in IEEE ICTAI 2025
摘要:可解释人工智能(XAI)已经成为解释复杂深度学习模型预测的关键工具。虽然XAI已经越来越多地应用于声学的各个领域,但它在生物声学中的应用,包括分析来自生物体的音频信号,仍然相对不足。在本文中,我们调查了一种鸟类的发声,在北美的整个范围内具有很强的地理变异。音频记录被转换成声谱图图像,并用于训练深度卷积神经网络(CNN)进行分类,准确率达到94.8%。为了解释模型的预测,我们应用了模型不可知(LIME,SHAP)和模型特定(DeepLIFT,Grad-CAM)XAI技术。这些技术产生了不同但互补的解释,当它们的解释被考虑在一起时,它们为模型的决策提供了更完整和可解释的见解。这项工作强调了使用XAI技术组合来提高信任和互操作性的重要性,不仅在更广泛的声学信号分析中,而且还主张在不同领域的特定任务中具有更广泛的适用性。
摘要:Explainable Artificial Intelligence (XAI) has emerged as a critical tool for interpreting the predictions of complex deep learning models. While XAI has been increasingly applied in various domains within acoustics, its use in bioacoustics, which involves analyzing audio signals from living organisms, remains relatively underexplored. In this paper, we investigate the vocalizations of a bird species with strong geographic variation throughout its range in North America. Audio recordings were converted into spectrogram images and used to train a deep Convolutional Neural Network (CNN) for classification, achieving an accuracy of 94.8\%. To interpret the model's predictions, we applied both model-agnostic (LIME, SHAP) and model-specific (DeepLIFT, Grad-CAM) XAI techniques. These techniques produced different but complementary explanations, and when their explanations were considered together, they provided more complete and interpretable insights into the model's decision-making. This work highlights the importance of using a combination of XAI techniques to improve trust and interoperability, not only in broader acoustics signal analysis but also argues for broader applicability in different domain specific tasks.


【3】Behind the Scenes: Mechanistic Interpretability of LoRA-adapted Whisper for Speech Emotion Recognition
标题:幕后:用于语音情感识别的LoRA适应Whisper的机械解释性
链接:https://arxiv.org/abs/2509.08454

作者:Yujian Ma, Jinqiu Sang, Ruizhe Li
备注:Work in process
摘要:像Whisper这样的大型预训练语音模型提供了很强的泛化能力,但对资源有效的适应提出了重大挑战。低秩自适应(Low-Rank Adaptation,LoRA)已成为一种流行的参数有效的微调方法,但其在语音任务中的潜在机制仍然知之甚少。在这项工作中,我们进行了第一次系统的机械可解释性研究LoRA内的耳语编码器语音情感识别(SER)。使用一套分析工具,包括层贡献探测,logit镜头检查,并通过奇异值分解(SVD)和中心内核对齐(CKA)表示相似性,我们揭示了两个关键机制:延迟专业化过程,在巩固特定任务的信息之前保留早期层的一般特征,以及LoRA矩阵之间的前向对齐,后向差分动态。我们的研究结果阐明了LoRA如何重塑编码器层次结构,为在大型语音模型中设计高效和可解释的自适应策略提供了经验见解和更深层次的机械理解。
摘要:Large pre-trained speech models such as Whisper offer strong generalization but pose significant challenges for resource-efficient adaptation. Low-Rank Adaptation (LoRA) has become a popular parameter-efficient fine-tuning method, yet its underlying mechanisms in speech tasks remain poorly understood. In this work, we conduct the first systematic mechanistic interpretability study of LoRA within the Whisper encoder for speech emotion recognition (SER). Using a suite of analytical tools, including layer contribution probing, logit-lens inspection, and representational similarity via singular value decomposition (SVD) and centered kernel alignment (CKA), we reveal two key mechanisms: a delayed specialization process that preserves general features in early layers before consolidating task-specific information, and a forward alignment, backward differentiation dynamic between LoRA's matrices. Our findings clarify how LoRA reshapes encoder hierarchies, providing both empirical insights and a deeper mechanistic understanding for designing efficient and interpretable adaptation strategies in large speech models.


【4】CommonVoice-SpeechRE and RPG-MoGe: Advancing Speech Relation Extraction with a New Dataset and Multi-Order Generative Framework
标题:Commonvoice-SpeechRE和RPG-MoGee:利用新数据集和多阶生成框架推进语音关系提取
链接:https://arxiv.org/abs/2509.08438

作者:Jinzhong Ning, Paerhati Tulajiang, Yingying Le, Yijia Zhang, Yuanyuan Sun, Hongfei Lin, Haifeng Liu
摘要:语音关系抽取(SpeechRE)的目的是直接从语音中抽取关系三元组。然而,现有的基准数据集严重依赖于合成数据,缺乏足够的数量和真实人类语音的多样性。此外,现有模型还受到严格的单阶生成模板和弱语义对齐的影响,大大限制了它们的性能。为了应对这些挑战,我们引入了CommonVoice-SpeechRE,这是一个大规模的数据集,包含来自不同说话者的近20,000个真人语音样本,为SpeechRE研究建立了一个新的基准。此外,我们提出了关系引导的多阶生成集成(RPG-MoGe),这是一种新的框架,其特征在于:(1)多阶三元组生成集成策略,在训练和推理过程中通过不同的元素顺序利用数据多样性,以及(2)基于CNN的潜在关系预测头,生成显式关系提示,以指导跨模态对齐和准确的三元组生成。实验表明,我们的方法优于最先进的方法,为现实世界的SpeechRE提供了一个基准数据集和一个有效的解决方案。源代码和数据集可在https://github.com/NingJinzhong/SpeechRE_RPG_MoGe上公开获得。
摘要:Speech Relation Extraction (SpeechRE) aims to extract relation triplets directly from speech. However, existing benchmark datasets rely heavily on synthetic data, lacking sufficient quantity and diversity of real human speech. Moreover, existing models also suffer from rigid single-order generation templates and weak semantic alignment, substantially limiting their performance. To address these challenges, we introduce CommonVoice-SpeechRE, a large-scale dataset comprising nearly 20,000 real-human speech samples from diverse speakers, establishing a new benchmark for SpeechRE research. Furthermore, we propose the Relation Prompt-Guided Multi-Order Generative Ensemble (RPG-MoGe), a novel framework that features: (1) a multi-order triplet generation ensemble strategy, leveraging data diversity through diverse element orders during both training and inference, and (2) CNN-based latent relation prediction heads that generate explicit relation prompts to guide cross-modal alignment and accurate triplet generation. Experiments show our approach outperforms state-of-the-art methods, providing both a benchmark dataset and an effective solution for real-world SpeechRE. The source code and dataset are publicly available at https://github.com/NingJinzhong/SpeechRE_RPG_MoGe.


【5】LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
标题:LatentSecureGrad:采用潜伏扩散/流匹配模型的非并行语音转换
链接:https://arxiv.org/abs/2509.08379

作者:Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Yuto Kondo
备注:Submitted to IEEE-TASLP
摘要:之前,我们介绍了VoiceGrad,一种非并行语音转换(VC)技术,使用基于分数的扩散模型实现从源到目标扬声器的梅尔频谱图转换。这个概念涉及训练一个分数网络来预测来自不同说话者的梅尔频谱图的对数密度的梯度。VC是通过迭代地调整输入的梅尔频谱图,直到类似于目标说话人的。然而,挑战依然存在:音频质量需要改进,并且与设计用于以非常高的速度操作的现代VC方法相比,转换速度较慢。为了解决这些问题,我们将潜在扩散模型引入到VoiceGrad中,提出了一个在自动编码器瓶颈中具有反向扩散的改进版本。此外,我们建议使用流匹配模型作为扩散模型的替代方案,以进一步加快转换过程,而不影响转换质量。实验结果表明,增强的语音质量和加速转换相比,原来的。
摘要:Previously, we introduced VoiceGrad, a nonparallel voice conversion (VC) technique enabling mel-spectrogram conversion from source to target speakers using a score-based diffusion model. The concept involves training a score network to predict the gradient of the log density of mel-spectrograms from various speakers. VC is executed by iteratively adjusting an input mel-spectrogram until resembling the target speaker's. However, challenges persist: audio quality needs improvement, and conversion is slower compared to modern VC methods designed to operate at very high speeds. To address these, we introduce latent diffusion models into VoiceGrad, proposing an improved version with reverse diffusion in the autoencoder bottleneck. Additionally, we propose using a flow matching model as an alternative to the diffusion model to further speed up the conversion process without compromising the conversion quality. Experimental results show enhanced speech quality and accelerated conversion compared to the original.


【6】Segment Transformer: AI-Generated Music Detection via Music Structural Analysis
标题:片段Transformer:通过音乐结构分析的人工智能生成音乐检测
链接:https://arxiv.org/abs/2509.08283

作者:Yumin Kim, Seonghyeon Go
摘要:在音乐信息检索领域,音频和音乐生成系统得到了显著的发展。这些技术的进步引发了版权问题,因为人工智能生成音乐(AIGM)的所有权和作者身份仍不清楚。此外,很难清楚地确定一个作品是由人工智能生成的还是由人类创作的。为了应对这些挑战,我们的目标是通过分析音乐片段的结构模式来提高AIGM检测的准确性。具体来说,为了从短音频片段中提取音乐特征,我们集成了各种预训练模型,包括自监督学习(SSL)模型或音频效果编码器,每个模型都在我们建议的基于transformer的框架中。此外,对于长音频,我们开发了一个片段转换器Transformer,它将音乐分成片段并学习片段间的关系。我们使用了FakeMusicCaps和SONICS数据集,在短音频和全音频检测实验中都实现了高准确性。这些发现表明,将片段级音乐特征整合到长距离时间分析中可以有效地提高AIGM检测系统的性能和鲁棒性。
摘要:Audio and music generation systems have been remarkably developed in the music information retrieval (MIR) research field. The advancement of these technologies raises copyright concerns, as ownership and authorship of AI-generated music (AIGM) remain unclear. Also, it can be difficult to determine whether a piece was generated by AI or composed by humans clearly. To address these challenges, we aim to improve the accuracy of AIGM detection by analyzing the structural patterns of music segments. Specifically, to extract musical features from short audio clips, we integrated various pre-trained models, including self-supervised learning (SSL) models or an audio effect encoder, each within our suggested transformer-based framework. Furthermore, for long audio, we developed a segment transformer that divides music into segments and learns inter-segment relationships. We used the FakeMusicCaps and SONICS datasets, achieving high accuracy in both the short-audio and full-audio detection experiments. These findings suggest that integrating segment-level musical features into long-range temporal analysis can effectively enhance both the performance and robustness of AIGM detection systems.


【7】Real-world Music Plagiarism Detection With Music Segment Transcription System
标题:使用音乐片段转录系统检测现实世界的音乐抄袭
链接:https://arxiv.org/abs/2509.08282

作者:Seonghyeon Go
备注:Accepted in APSIPA 2025 but not published yet(will be published in 2 month..), Arxiv preprint ready for references in future-works
摘要:由于音乐信息检索(MIR)技术的不断进步,产生和分发音乐变得更加多样化和可访问。在这种情况下,人们对音乐知识产权保护的兴趣日益增加,以保护个人音乐版权。在这项工作中,我们提出了一个系统,检测音乐剽窃相结合的各种MIR技术。我们开发了一个音乐片段转录系统,从音频记录中提取音乐上有意义的片段,以检测不同音乐格式的剽窃。有了这个系统,我们计算相似性分数的基础上,可以通过全面的音乐分析进行评估的多个音乐特征。我们的方法在音乐剽窃检测实验中表现出了良好的效果,所提出的方法可以应用于现实世界的音乐场景。我们还收集了一个类似的音乐对(SMP)数据集,用于使用真实案例进行音乐相似性研究。该数据集是公开的。
摘要:As a result of continuous advances in Music Information Retrieval (MIR) technology, generating and distributing music has become more diverse and accessible. In this context, interest in music intellectual property protection is increasing to safeguard individual music copyrights. In this work, we propose a system for detecting music plagiarism by combining various MIR technologies. We developed a music segment transcription system that extracts musically meaningful segments from audio recordings to detect plagiarism across different musical formats. With this system, we compute similarity scores based on multiple musical features that can be evaluated through comprehensive musical analysis. Our approach demonstrated promising results in music plagiarism detection experiments, and the proposed method can be applied to real-world music scenarios. We also collected a Similar Music Pair (SMP) dataset for musical similarity research using real-world cases. The dataset are publicly available.


【8】LALM-Eval: An Open-Source Toolkit for Holistic Evaluation of Large Audio Language Models
标题:LALM-Eval:一个用于大型音频语言模型整体评估的开源工具包
链接:https://arxiv.org/abs/2509.08031

作者:Sidharth Surapaneni, Hoang Nguyen, Jash Mehta, Aman Tiwari, Oluwanifemi Bamgbose, Akshay Kalkunte, Sai Rajeswar, Sathwik Tejaswi Madhusudhan
摘要:大型音频语言模型(LALM)正在迅速发展,但由于效率低下的工具包限制了公平比较和系统评估,因此对其进行评估仍然具有挑战性。目前的框架存在三个关键问题:处理速度慢,阻碍了大规模研究,不一致的提示损害了可重复性,以及任务覆盖范围窄,错过了重要的音频推理能力。我们介绍了LALM评估,一个有效的和全面的评估框架LALM。我们的系统通过优化的批处理和并行执行实现了高达127%的速度比现有的工具包,使大规模的评估以前不切实际。我们提供标准化的提示协议和灵活的配置,在不同的场景中进行公平的模型比较。此外,我们还引入了两个新的评估类别:用于时间音频理解的LLM自适应Diarization和用于复杂的基于音频的认知任务的口语推理。通过对380多个任务的评估,我们揭示了当前LALM的显著差距,特别是在时间理解和复杂的口语推理任务方面。我们的研究结果还强调了音频基准中存在的教学模式缺乏标准化,这可能导致下游任务后具有挑战性的复杂教学的性能差异高达9.5个绝对点。LALM-Eval提供了实用的评估工具和对模型局限性的深入了解,推进了系统的LALM开发。
摘要:Large Audio Language Models (LALMs) are rapidly advancing, but evaluating them remains challenging due to inefficient toolkits that limit fair comparison and systematic assessment. Current frameworks suffer from three critical issues: slow processing that bottlenecks large-scale studies, inconsistent prompting that hurts reproducibility, and narrow task coverage that misses important audio reasoning capabilities. We introduce LALM-Eval, an efficient and comprehensive evaluation framework for LALMs. Our system achieves a speedup of up to 127% over existing toolkits through optimized batch processing and parallel execution, enabling large-scale evaluations previously impractical. We provide standardized prompting protocols and flexible configurations for fair model comparison across diverse scenarios. Additionally, we introduce two new evaluation categories: LLM-Adaptive Diarization for temporal audio understanding and Spoken Language Reasoning for complex audio-based cognitive tasks. Through evaluation across 380+ tasks, we reveal significant gaps in current LALMs, particularly in temporal understanding and complex spoken language reasoning tasks. Our findings also highlight a lack of standardization in instruction modality existent across audio benchmarks, which can lead up performance differences up to 9.5 absolute points on the challenging complex instruction following downstream tasks. LALM-Eval provides both practical evaluation tools and insights into model limitations, advancing systematic LALM development.


【9】Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching
标题:通过Transformer层缓存加速基于转换器的文本到语音的扩散
链接:https://arxiv.org/abs/2509.08696

作者:Siratish Sakpiboonchit
备注:9 pages, 2 tables, 5 figures
摘要:本文提出了一种方法,以加速基于扩散Transformer(DiT)的文本到语音(TTS)模型的推理过程中,通过应用选择性缓存机制的Transformer层。具体来说,我将SmoothCache集成到F5-TTS架构中,专注于缓存自注意和前馈网络层的输出,以减少去噪过程中的冗余计算。引入校准阶段来分析时间步长之间的L1相对误差,指导选择最小化质量降级的缓存调度。为了解决层间依赖性的问题,采用了统一的缓存调度,将来自自关注层的缓存模式应用于两种层类型。LibriSpeech-PC和Seed-TTS数据集上的实验评估了各种缓存阈值和去噪步骤配置。结果表明,在较高的去噪步骤缓存减少推理时间,而不影响输出质量,而在较低的步骤缓存可能会产生负面影响的合成质量类似于减少去噪步骤的总数。客观和主观指标证实了SmoothCache在保持性能的同时提高计算效率的有效性。缓存推理和缩减步长推理之间的比较进一步突出了选择性缓存的好处,特别是在高步长配置下。这项工作表明,Transformer层缓存是优化基于扩散transformer的TTS模型的实用解决方案,无需更改架构或重新训练。示例推理结果可以在https://siratish.github.io/F5-TTS_SmoothCache/上听到。
摘要:This paper presents a method to accelerate the inference process of diffusion transformer (DiT)-based text-to-speech (TTS) models by applying a selective caching mechanism to transformer layers. Specifically, I integrate SmoothCache into the F5-TTS architecture, focusing on caching outputs of self-attention and feed-forward network layers to reduce redundant computations during the denoising process. A calibration phase is introduced to analyze L1 relative errors between timesteps, guiding the selection of cache schedules that minimize quality degradation. To address the problem of inter-layer dependency, a unified caching schedule is adopted, applying the cache pattern derived from self-attention layers to both layer types. Experiments on LibriSpeech-PC and Seed-TTS datasets evaluate various cache thresholds and denoising step configurations. Results show that caching at higher denoising steps reduces inference time without compromising output quality, whereas caching at lower steps can negatively impact synthesis quality similarly to reducing the total number of denoising steps. Objective and subjective metrics confirm the effectiveness of SmoothCache in maintaining performance while improving computational efficiency. Comparisons between cached inference and reduced-step inference further highlight the benefits of selective caching, especially under high-step configurations. This work demonstrates that transformer layer caching is a practical solution for optimizing diffusion transformer-based TTS models without requiring architectural changes or retraining. Example inference results can be heard at https://siratish.github.io/F5-TTS_SmoothCache/ .


【10】Context-Aware Query Refinement for Target Sound Extraction: Handling Partially Matched Queries
标题:目标声音提取的上下文感知查询细化:处理部分匹配的收件箱
链接:https://arxiv.org/abs/2509.08292

作者: Ryo Sato, Chiho Haruta, Nobuhiko Hiruma, Keisuke Imoto
备注:Accepted to IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2025
摘要:目标声音提取(TSE)是从音频混合中提取由查询指定的目标声音的任务。许多先前的研究都集中在完全匹配查询(FMQ)条件下的问题设置,其中查询只指定混合中存在的主动声音。然而,在现实世界的场景中,查询可能包括不存在于混合中的非活动声音。这导致了诸如完全不匹配查询(FUQ)条件和部分匹配查询(PMQ)条件之类的场景,在完全不匹配查询(FUQ)条件下,在查询中仅指定不活动的声音,在部分匹配查询(PMQ)条件下,既指定活动的声音又指定不活动的声音。在这些条件中,PMQ条件下的性能退化在很大程度上被忽视了。为了在PMQ条件下实现鲁棒的TSE,我们提出了上下文感知的查询细化。该方法在基于估计的声音类活动的推断期间从查询中消除不活动类。实验结果表明,传统方法在PMQ条件下性能下降,而本文方法有效地缓解了这种下降,并在不同的查询条件下实现了高鲁棒性。
摘要:Target sound extraction (TSE) is the task of extracting a target sound specified by a query from an audio mixture. Much prior research has focused on the problem setting under the Fully Matched Query (FMQ) condition, where the query specifies only active sounds present in the mixture. However, in real-world scenarios, queries may include inactive sounds that are not present in the mixture. This leads to scenarios such as the Fully Unmatched Query (FUQ) condition, where only inactive sounds are specified in the query, and the Partially Matched Query (PMQ) condition, where both active and inactive sounds are specified. Among these conditions, the performance degradation under the PMQ condition has been largely overlooked. To achieve robust TSE under the PMQ condition, we propose context-aware query refinement. This method eliminates inactive classes from the query during inference based on the estimated sound class activity. Experimental results demonstrate that while conventional methods suffer from performance degradation under the PMQ condition, the proposed method effectively mitigates this degradation and achieves high robustness under diverse query conditions.


eess.AS音频处理


【1】Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching
标题:通过转换器层缓存加速基于转换器的文本到语音的扩散
链接:https://arxiv.org/abs/2509.08696

作者:Siratish Sakpiboonchit
备注:9 pages, 2 tables, 5 figures
摘要:本文提出了一种方法,以加速基于扩散Transformer(DiT)的文本到语音(TTS)模型的推理过程中,通过应用选择性缓存机制的Transformer层。具体来说,我将SmoothCache集成到F5-TTS架构中,专注于缓存自注意和前馈网络层的输出,以减少去噪过程中的冗余计算。引入校准阶段来分析时间步之间的L1相对误差,从而指导选择最大限度地减少质量下降的缓存调度。为了解决层间依赖性的问题,采用了统一的缓存调度,将来自自关注层的缓存模式应用于两种层类型。LibriSpeech-PC和Seed-TTS数据集上的实验评估了各种缓存阈值和去噪步骤配置。结果表明,在较高的去噪步骤缓存减少推理时间,而不影响输出质量,而在较低的步骤缓存可能会产生负面影响的合成质量类似于减少去噪步骤的总数。客观和主观指标证实了SmoothCache在保持性能的同时提高计算效率的有效性。缓存推理和缩减步长推理之间的比较进一步突出了选择性缓存的好处,特别是在高步长配置下。这项工作表明,Transformer层缓存是一个实用的解决方案,优化扩散变压器为基础的TTS模型,而不需要架构的变化或再培训。示例推理结果可以在https://siratish.github.io/F5-TTS_SmoothCache/上听到。
摘要:This paper presents a method to accelerate the inference process of diffusion transformer (DiT)-based text-to-speech (TTS) models by applying a selective caching mechanism to transformer layers. Specifically, I integrate SmoothCache into the F5-TTS architecture, focusing on caching outputs of self-attention and feed-forward network layers to reduce redundant computations during the denoising process. A calibration phase is introduced to analyze L1 relative errors between timesteps, guiding the selection of cache schedules that minimize quality degradation. To address the problem of inter-layer dependency, a unified caching schedule is adopted, applying the cache pattern derived from self-attention layers to both layer types. Experiments on LibriSpeech-PC and Seed-TTS datasets evaluate various cache thresholds and denoising step configurations. Results show that caching at higher denoising steps reduces inference time without compromising output quality, whereas caching at lower steps can negatively impact synthesis quality similarly to reducing the total number of denoising steps. Objective and subjective metrics confirm the effectiveness of SmoothCache in maintaining performance while improving computational efficiency. Comparisons between cached inference and reduced-step inference further highlight the benefits of selective caching, especially under high-step configurations. This work demonstrates that transformer layer caching is a practical solution for optimizing diffusion transformer-based TTS models without requiring architectural changes or retraining. Example inference results can be heard at https://siratish.github.io/F5-TTS_SmoothCache/ .


【2】Audio Deepfake Verification
标题:音频Deepfake验证
链接:https://arxiv.org/abs/2509.08476

作者:Li Wang, Junyi Ao, Linyong Gan, Yuancheng Wang, Xueyao Zhang, Zhizheng Wu
摘要:随着deepfake技术的快速发展,简单地对音频进行真假的二元判断已经不足以满足实际需求。准确确定特定的deepfake方法变得至关重要。本文介绍了音频Deepfake验证(ADV)任务,有效解决了现有deepfake源跟踪方法在闭集场景中的局限性,旨在实现开集deepfake源跟踪。同时,提出了Audity双分支架构,从音频结构和生成伪影两个维度提取deepfake特征。实验结果表明,双分支Audity架构的性能优于任何单分支配置,它可以同时在深度伪造检测和验证任务中实现出色的性能。
摘要:With the rapid development of deepfake technology, simply making a binary judgment of true or false on audio is no longer sufficient to meet practical needs. Accurately determining the specific deepfake method has become crucial. This paper introduces the Audio Deepfake Verification (ADV) task, effectively addressing the limitations of existing deepfake source tracing methods in closed-set scenarios, aiming to achieve open-set deepfake source tracing. Meanwhile, the Audity dual-branch architecture is proposed, extracting deepfake features from two dimensions: audio structure and generation artifacts. Experimental results show that the dual-branch Audity architecture outperforms any single-branch configuration, and it can simultaneously achieve excellent performance in both deepfake detection and verification tasks.


【3】Joint Learning using Mixture-of-Expert-Based Representation for Enhanced Speech Generation and Robust Emotion Recognition
标题:使用基于混合专家的表示的联合学习增强语音生成和鲁棒的情感识别
链接:https://arxiv.org/abs/2509.08470

作者: Jing-Tong Tzeng, Carlos Busso, Chi-Chun Lee
摘要:语音情感识别在构建情感感知语音系统中起着至关重要的作用,但其性能在噪声条件下会显著下降。虽然语音增强(SE)可以提高鲁棒性,但它通常会引入模糊情感线索的伪像,并增加流水线的计算开销。多任务学习(MTL)通过联合优化SE和SER任务提供了一种替代方案。然而,传统的共享骨干网模型经常遭受梯度干扰和任务之间的代表性冲突。为了解决这些挑战,我们提出了稀疏混合专家表示集成技术(稀疏MERIT),一个灵活的MTL框架,适用于自监督语音表示的逐帧专家路由。Sparse MERIT结合了特定于任务的门控网络,可以从每个帧的共享专家池中动态选择,从而实现参数高效和任务自适应的表示学习。在MSP-Podcast语料库上的实验表明,稀疏MERIT在SER和SE任务上的性能始终优于基线模型。在最具挑战性的-5 dB信噪比(SNR)条件下,Sparse MERIT将SER F1-macro比依赖于SE预处理策略的基线平均提高了12.0%,比原始MTL基线平均提高了3.4%,在不可见噪声条件下具有统计学显著性。对于SE,Sparse MERIT使分段SNR(SSNR)比SE预处理基线提高28.2%,比初始MTL基线提高20.0%。这些结果表明,稀疏MERIT提供了强大的和可推广的性能,在嘈杂的环境中的情感识别和增强任务。
摘要:Speech emotion recognition (SER) plays a critical role in building emotion-aware speech systems, but its performance degrades significantly under noisy conditions. Although speech enhancement (SE) can improve robustness, it often introduces artifacts that obscure emotional cues and adds computational overhead to the pipeline. Multi-task learning (MTL) offers an alternative by jointly optimizing SE and SER tasks. However, conventional shared-backbone models frequently suffer from gradient interference and representational conflicts between tasks. To address these challenges, we propose the Sparse Mixture-of-Experts Representation Integration Technique (Sparse MERIT), a flexible MTL framework that applies frame-wise expert routing over self-supervised speech representations. Sparse MERIT incorporates task-specific gating networks that dynamically select from a shared pool of experts for each frame, enabling parameter-efficient and task-adaptive representation learning. Experiments on the MSP-Podcast corpus show that Sparse MERIT consistently outperforms baseline models on both SER and SE tasks. Under the most challenging condition of -5 dB signal-to-noise ratio (SNR), Sparse MERIT improves SER F1-macro by an average of 12.0% over a baseline relying on a SE pre-processing strategy, and by 3.4% over a naive MTL baseline, with statistical significance on unseen noise conditions. For SE, Sparse MERIT improves segmental SNR (SSNR) by 28.2% over the SE pre-processing baseline and by 20.0% over the naive MTL baseline. These results demonstrate that Sparse MERIT provides robust and generalizable performance for both emotion recognition and enhancement tasks in noisy environments.


【4】Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model
标题:基于语音语言模型的语音情感识别通过上下文学习进行Few-Shot个性化
链接:https://arxiv.org/abs/2509.08344

作者:Mana Ihori, Taiga Yamane, Naotaka Kawata, Naoki Makishima, Tomohiro Tanaka, Satoshi Suzuki, Shota Orihashi, Ryo Masumura
备注:Accepted by ASRU 2025
摘要:提出了一种基于上下文学习的语音情感识别方法。由于情绪的表达因人而异,特定于说话人的适应对于提高SER性能至关重要。传统的SER方法已经使用目标说话者的情感话语来个性化,但是通常难以预先准备对应于所有情感标签的话语。我们的想法,以克服这个困难是获得说话人的特征,通过条件化的几个情感话语的目标说话人在ICL为基础的推理。ICL是一种通过在大型语言模型(LLM)中进行推理来调节一些输入输出示例来执行看不见的任务的方法。我们元训练从LLM扩展的语音语言模型,以学习如何通过ICL执行个性化SER。使用我们新收集的SER数据集的实验结果表明,该方法优于传统的方法。
摘要:This paper proposes a personalization method for speech emotion recognition (SER) through in-context learning (ICL). Since the expression of emotions varies from person to person, speaker-specific adaptation is crucial for improving the SER performance. Conventional SER methods have been personalized using emotional utterances of a target speaker, but it is often difficult to prepare utterances corresponding to all emotion labels in advance. Our idea to overcome this difficulty is to obtain speaker characteristics by conditioning a few emotional utterances of the target speaker in ICL-based inference. ICL is a method to perform unseen tasks by conditioning a few input-output examples through inference in large language models (LLMs). We meta-train a speech-language model extended from the LLM to learn how to perform personalized SER via ICL. Experimental results using our newly collected SER dataset demonstrate that the proposed method outperforms conventional methods.


【5】Context-Aware Query Refinement for Target Sound Extraction: Handling Partially Matched Queries
标题:目标声音提取的上下文感知查询细化:处理部分匹配的收件箱
链接:https://arxiv.org/abs/2509.08292

作者: Ryo Sato, Chiho Haruta, Nobuhiko Hiruma, Keisuke Imoto
备注:Accepted to IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2025
摘要:目标声音提取(TSE)是从音频混合中提取由查询指定的目标声音的任务。许多先前的研究都集中在完全匹配查询(FMQ)条件下的问题设置,其中查询只指定混合中存在的主动声音。然而,在现实世界的场景中,查询可能包括不存在于混合中的非活动声音。这导致了诸如完全不匹配查询(FUQ)条件和部分匹配查询(PMQ)条件之类的场景,在完全不匹配查询(FUQ)条件下,在查询中仅指定不活动的声音,在部分匹配查询(PMQ)条件下,既指定活动的声音又指定不活动的声音。在这些条件中,PMQ条件下的性能退化在很大程度上被忽视了。为了在PMQ条件下实现鲁棒的TSE,我们提出了上下文感知的查询细化。该方法在基于估计的声音类活动的推断期间从查询中消除不活动类。实验结果表明,传统方法在PMQ条件下性能下降,而本文方法有效地缓解了这种下降,并在不同的查询条件下实现了高鲁棒性。
摘要:Target sound extraction (TSE) is the task of extracting a target sound specified by a query from an audio mixture. Much prior research has focused on the problem setting under the Fully Matched Query (FMQ) condition, where the query specifies only active sounds present in the mixture. However, in real-world scenarios, queries may include inactive sounds that are not present in the mixture. This leads to scenarios such as the Fully Unmatched Query (FUQ) condition, where only inactive sounds are specified in the query, and the Partially Matched Query (PMQ) condition, where both active and inactive sounds are specified. Among these conditions, the performance degradation under the PMQ condition has been largely overlooked. To achieve robust TSE under the PMQ condition, we propose context-aware query refinement. This method eliminates inactive classes from the query during inference based on the estimated sound class activity. Experimental results demonstrate that while conventional methods suffer from performance degradation under the PMQ condition, the proposed method effectively mitigates this degradation and achieves high robustness under diverse query conditions.


【6】A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
标题:用于基于音节的ASB的自下而上框架,具有算术通用的语音属性建模
链接:https://arxiv.org/abs/2509.08173

作者:Hao Yen, Pin-Jui Ku, Sabato Marco Siniscalchi, Chin-Hui Lee
摘要:我们提出了一个自下而上的框架,自动语音识别(ASR)在音节为基础的语言统一的语言通用发音属性建模与音节级预测。该系统首先识别发音属性的序列或网格,作为一种语言通用的,可解释的发音表示,然后通过结构化的知识集成过程将它们转换为音节。我们引入两个评估指标,即发音错误率(PrER)和音节同音异义词错误率(SHER),以评估模型的能力,以捕捉发音和处理音节歧义。在AISHELL-1汉语语料库上的实验结果表明,与直接音节预测模型相比,该自底向上框架在低资源条件下具有较好的性能和鲁棒性。此外,我们研究了日语的zero-shot跨语言可转移性,并证明了基于字符和音素的基线的40%的错误率降低显着改善。
摘要:We propose a bottom-up framework for automatic speech recognition (ASR) in syllable-based languages by unifying language-universal articulatory attribute modeling with syllable-level prediction. The system first recognizes sequences or lattices of articulatory attributes that serve as a language-universal, interpretable representation of pronunciation, and then transforms them into syllables through a structured knowledge integration process. We introduce two evaluation metrics, namely Pronunciation Error Rate (PrER) and Syllable Homonym Error Rate (SHER), to evaluate the model's ability to capture pronunciation and handle syllable ambiguities. Experimental results on the AISHELL-1 Mandarin corpus demonstrate that the proposed bottom-up framework achieves competitive performance and exhibits better robustness under low-resource conditions compared to the direct syllable prediction model. Furthermore, we investigate the zero-shot cross-lingual transferability on Japanese and demonstrate significant improvements over character- and phoneme-based baselines by 40% error rate reduction.


【7】PianoVAM: A Multimodal Piano Performance Dataset
标题:PianoVAM:多模式钢琴演奏数据集
链接:https://arxiv.org/abs/2509.08800

作者:Yonghyun Kim, Junhyung Park, Joonhyung Bae, Kirak Kim, Taegyun Kwon, Alexander Lerch, Juhan Nam
备注:Accepted to the 26th International Society for Music Information Retrieval (ISMIR) Conference, 2025
摘要:音乐表现的多模态性质已经推动了对音乐信息检索(MIR)社区内音频域之外的数据的越来越大的兴趣。本文介绍了PianoVAM,这是一个综合的钢琴演奏数据集,包括视频,音频,视频,手部标志,指法标签和丰富的元数据。该数据集是使用一架更大的钢琴录制的,在业余钢琴家的日常练习中捕捉音频和视频,以及在现实和不同的表演条件下同步的顶视图视频。使用预先训练的手部姿势估计模型和半自动指法注释算法提取手部标志和指法标签。我们讨论了数据收集过程中遇到的挑战和不同模式的调整过程。此外,我们描述了我们的指法标注方法的基础上,从视频中提取的手地标。最后,我们使用PianoVAM数据集提出了纯音频和视听钢琴转录的基准测试结果,并讨论了其他潜在的应用。
摘要:The multimodal nature of music performance has driven increasing interest in data beyond the audio domain within the music information retrieval (MIR) community. This paper introduces PianoVAM, a comprehensive piano performance dataset that includes videos, audio, MIDI, hand landmarks, fingering labels, and rich metadata. The dataset was recorded using a Disklavier piano, capturing audio and MIDI from amateur pianists during their daily practice sessions, alongside synchronized top-view videos in realistic and varied performance conditions. Hand landmarks and fingering labels were extracted using a pretrained hand pose estimation model and a semi-automated fingering annotation algorithm. We discuss the challenges encountered during data collection and the alignment process across different modalities. Additionally, we describe our fingering annotation method based on hand landmarks extracted from videos. Finally, we present benchmarking results for both audio-only and audio-visual piano transcription using the PianoVAM dataset and discuss additional potential applications.


【8】Behind the Scenes: Mechanistic Interpretability of LoRA-adapted Whisper for Speech Emotion Recognition
标题:幕后:用于语音情感识别的LoRA适应Whisper的机械解释性
链接:https://arxiv.org/abs/2509.08454

作者:Yujian Ma, Jinqiu Sang, Ruizhe Li
备注:Work in process
摘要:像Whisper这样的大型预训练语音模型提供了很强的泛化能力,但对资源有效的适应提出了重大挑战。低秩自适应(Low-Rank Adaptation,LoRA)已成为一种流行的参数有效的微调方法,但其在语音任务中的潜在机制仍然知之甚少。在这项工作中,我们进行了第一次系统的机械可解释性研究LoRA内的耳语编码器语音情感识别(SER)。使用一套分析工具,包括层贡献探测,logit镜头检查,并通过奇异值分解(SVD)和中心内核对齐(CKA)表示相似性,我们揭示了两个关键机制:延迟专业化过程,在巩固特定任务的信息之前保留早期层的一般特征,以及LoRA矩阵之间的前向对齐,后向差分动态。我们的研究结果阐明了LoRA如何重塑编码器层次结构,为在大型语音模型中设计高效和可解释的自适应策略提供了经验见解和更深层次的机械理解。
摘要:Large pre-trained speech models such as Whisper offer strong generalization but pose significant challenges for resource-efficient adaptation. Low-Rank Adaptation (LoRA) has become a popular parameter-efficient fine-tuning method, yet its underlying mechanisms in speech tasks remain poorly understood. In this work, we conduct the first systematic mechanistic interpretability study of LoRA within the Whisper encoder for speech emotion recognition (SER). Using a suite of analytical tools, including layer contribution probing, logit-lens inspection, and representational similarity via singular value decomposition (SVD) and centered kernel alignment (CKA), we reveal two key mechanisms: a delayed specialization process that preserves general features in early layers before consolidating task-specific information, and a forward alignment, backward differentiation dynamic between LoRA's matrices. Our findings clarify how LoRA reshapes encoder hierarchies, providing both empirical insights and a deeper mechanistic understanding for designing efficient and interpretable adaptation strategies in large speech models.


【9】Real-world Music Plagiarism Detection With Music Segment Transcription System
标题:使用音乐片段转录系统检测现实世界的音乐抄袭
链接:https://arxiv.org/abs/2509.08282

作者:Seonghyeon Go
备注:Accepted in APSIPA 2025 but not published yet(will be published in 2 month..), Arxiv preprint ready for references in future-works
摘要:由于音乐信息检索(MIR)技术的不断进步,产生和分发音乐变得更加多样化和可访问。在这种情况下,人们对音乐知识产权保护的兴趣日益增加,以保护个人音乐版权。在这项工作中,我们提出了一个系统,检测音乐剽窃相结合的各种MIR技术。我们开发了一个音乐片段转录系统,从音频记录中提取音乐上有意义的片段,以检测不同音乐格式的剽窃。有了这个系统,我们计算相似性分数的基础上,可以通过全面的音乐分析进行评估的多个音乐特征。我们的方法在音乐剽窃检测实验中表现出了良好的效果,所提出的方法可以应用于现实世界的音乐场景。我们还收集了一个类似的音乐对(SMP)数据集,用于使用真实案例进行音乐相似性研究。该数据集是公开的。
摘要:As a result of continuous advances in Music Information Retrieval (MIR) technology, generating and distributing music has become more diverse and accessible. In this context, interest in music intellectual property protection is increasing to safeguard individual music copyrights. In this work, we propose a system for detecting music plagiarism by combining various MIR technologies. We developed a music segment transcription system that extracts musically meaningful segments from audio recordings to detect plagiarism across different musical formats. With this system, we compute similarity scores based on multiple musical features that can be evaluated through comprehensive musical analysis. Our approach demonstrated promising results in music plagiarism detection experiments, and the proposed method can be applied to real-world music scenarios. We also collected a Similar Music Pair (SMP) dataset for musical similarity research using real-world cases. The dataset are publicly available.


【10】LALM-Eval: An Open-Source Toolkit for Holistic Evaluation of Large Audio Language Models
标题:LALM-Eval:一个用于大型音频语言模型整体评估的开源工具包
链接:https://arxiv.org/abs/2509.08031

作者:Sidharth Surapaneni, Hoang Nguyen, Jash Mehta, Aman Tiwari, Oluwanifemi Bamgbose, Akshay Kalkunte, Sai Rajeswar, Sathwik Tejaswi Madhusudhan
摘要:大型音频语言模型(LALM)正在迅速发展,但由于效率低下的工具包限制了公平比较和系统评估,因此对其进行评估仍然具有挑战性。目前的框架存在三个关键问题:处理速度慢,阻碍了大规模研究,不一致的提示损害了可重复性,以及任务覆盖范围窄,错过了重要的音频推理能力。我们介绍了LALM评估,一个有效的和全面的评估框架LALM。我们的系统通过优化的批处理和并行执行实现了高达127%的速度比现有的工具包,使大规模的评估以前不切实际。我们提供标准化的提示协议和灵活的配置,在不同的场景中进行公平的模型比较。此外,我们还引入了两个新的评估类别:用于时间音频理解的LLM自适应Diarization和用于复杂的基于音频的认知任务的口语推理。通过对380多个任务的评估,我们揭示了当前LALM的显著差距,特别是在时间理解和复杂的口语推理任务方面。我们的研究结果还强调了音频基准中存在的教学模式缺乏标准化,这可能导致下游任务后具有挑战性的复杂教学的性能差异高达9.5个绝对点。LALM-Eval提供了实用的评估工具和对模型局限性的深入了解,推进了系统的LALM开发。
摘要:Large Audio Language Models (LALMs) are rapidly advancing, but evaluating them remains challenging due to inefficient toolkits that limit fair comparison and systematic assessment. Current frameworks suffer from three critical issues: slow processing that bottlenecks large-scale studies, inconsistent prompting that hurts reproducibility, and narrow task coverage that misses important audio reasoning capabilities. We introduce LALM-Eval, an efficient and comprehensive evaluation framework for LALMs. Our system achieves a speedup of up to 127% over existing toolkits through optimized batch processing and parallel execution, enabling large-scale evaluations previously impractical. We provide standardized prompting protocols and flexible configurations for fair model comparison across diverse scenarios. Additionally, we introduce two new evaluation categories: LLM-Adaptive Diarization for temporal audio understanding and Spoken Language Reasoning for complex audio-based cognitive tasks. Through evaluation across 380+ tasks, we reveal significant gaps in current LALMs, particularly in temporal understanding and complex spoken language reasoning tasks. Our findings also highlight a lack of standardization in instruction modality existent across audio benchmarks, which can lead up performance differences up to 9.5 absolute points on the challenging complex instruction following downstream tasks. LALM-Eval provides both practical evaluation tools and insights into model limitations, advancing systematic LALM development.


机器翻译由腾讯交互翻译提供,仅供参考