今日论文合集:cs.SD语音25篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】SonoWorld: From One Image to a 3D Audio-Visual Scene
标题:SonoWorld:从一张图像到3D视听场景
链接:https://arxiv.org/abs/2603.28757

作者:Derong Jin,Xiyi Chen,Ming C. Lin,Ruohan Gao
备注:Accepted by CVPR 2026, project page: https://humathe.github.io/sonoworld/
摘要:视觉场景生成方面取得了巨大的进步,现在可以将单个图像转化为可探索的3D世界,但如果没有声音,沉浸感仍然是不完整的。我们介绍Image2AVScene,从单个图像生成3D视听场景的任务,并提出SonoWorld,第一个框架来解决这一挑战。从一个图像中,我们的管道绘制出360°全景,将其提升到可导航的3D场景中,放置语言引导的声音锚点,并为点,面和环境源渲染立体混响,产生与场景几何和语义对齐的空间音频。对新策划的真实世界数据集和受控用户研究的定量评估证实了我们方法的有效性。除了自由视点的视听渲染,我们还展示了一次性声学学习和视听空间源分离的应用。项目网址:https://humathe.github.io/sonoworld/
摘要:Tremendous progress in visual scene generation now turns a single image into an explorable 3D world, yet immersion remains incomplete without sound. We introduce Image2AVScene, the task of generating a 3D audio-visual scene from a single image, and present SonoWorld, the first framework to tackle this challenge. From one image, our pipeline outpaints a 360° panorama, lifts it into a navigable 3D scene, places language-guided sound anchors, and renders ambisonics for point, areal, and ambient sources, yielding spatial audio aligned with scene geometry and semantics. Quantitative evaluations on a newly curated real-world dataset and a controlled user study confirm the effectiveness of our approach. Beyond free-viewpoint audio-visual rendering, we also demonstrate applications to one-shot acoustic learning and audio-visual spatial source separation. Project website: https://humathe.github.io/sonoworld/


【2】Constructing Composite Features for Interpretable Music-Tagging
标题:构建可解释音乐标签的复合特征
链接:https://arxiv.org/abs/2603.28644

作者:Chenhao Xue,Weitao Hu,Joyraj Chakraborty,Zhijin Guo,Kang Li,Tianyu Shi,Martin Reed,Nikolaos Thomos
备注:5 pages, 8 figures, accepted at ICASSP 2026
摘要:结合多种音频特征可以提高音乐标记的性能,但常见的基于深度学习的特征融合方法往往缺乏可解释性。为了解决这个问题,我们提出了一个遗传编程(GP)管道,自动演变的复合功能,通过数学组合的基础音乐功能,从而捕捉协同作用,同时保持可解释性。这种方法提供了类似于深度特征融合的代表性优势,而不会牺牲可解释性。MTG-Jamendo和GTZAN数据集上的实验表明,与不同抽象级别的基本特征集上的最先进系统相比,这些系统具有一致的改进。应该注意的是,大多数性能增益是在前几百个GP评估中注意到的,这表明可以在适度的搜索预算下识别有效的特征组合。顶级进化表达式包括线性,非线性和条件形式,具有各种低复杂度的解决方案,最高性能与简约压力相一致,以偏好更简单的表达式。分析这些复合特征进一步揭示了哪些交互和转换往往有利于标记,提供了在黑盒深度模型中仍然不透明的见解。
摘要:Combining multiple audio features can improve the performance of music tagging, but common deep learning-based feature fusion methods often lack interpretability. To address this problem, we propose a Genetic Programming (GP) pipeline that automatically evolves composite features by mathematically combining base music features, thereby capturing synergistic interactions while preserving interpretability. This approach provides representational benefits similar to deep feature fusion without sacrificing interpretability. Experiments on the MTG-Jamendo and GTZAN datasets demonstrate consistent improvements compared to state-of-the-art systems across base feature sets at different abstraction levels. It should be noted that most of the performance gains are noticed within the first few hundred GP evaluations, indicating that effective feature combinations can be identified under modest search budgets. The top evolved expressions include linear, nonlinear, and conditional forms, with various low-complexity solutions at top performance aligned with parsimony pressure to prefer simpler expressions. Analyzing these composite features further reveals which interactions and transformations tend to be beneficial for tagging, offering insights that remain opaque in black-box deep models.


【3】A Probabilistic Generative Model for Spectral Speech Enhancement
标题:频谱语音增强的概率生成模型
链接:https://arxiv.org/abs/2603.28436

作者:Marco Hidalgo-Araya,Raphaël Trésor,Bart Van Erp,Wouter W. L. Nuijten,Thijs Van De Laar,Bert De Vries
备注:Submitted to the IEEE Open Journal of Signal Processing
摘要:助听器中的语音增强在非平稳声学环境中仍然是一项艰巨的任务,这主要是因为当前的信号处理算法依赖于固定的手动调谐参数,这些参数不能原位适应不同的用户或收听环境。本文介绍了一个统一的模块化框架,制定信号处理,学习和个性化的贝叶斯推理与明确的不确定性跟踪。所提出的框架取代ad hoc算法设计与一个单一的概率生成模型,不断适应不断变化的声学条件和用户的喜好。它扩展了频谱减法与原则机制,在现场个性化和适应声学环境。该系统被实现为一个相互关联的概率状态空间模型,并通过在\texttt{RxInfer.jl}概率编程环境中传递变量消息进行推理,从而实现助听器约束下的实时贝叶斯处理。在VoiceBank+DEMAND语料库上进行的概念验证实验表明,该算法具有较好的语音质量和降噪效果,有效参数为85个。该框架为不确定性感知的自适应助听器处理提供了一个可解释的、数据高效的基础,并指向通过概率推理不断学习的设备。
摘要:Speech enhancement in hearing aids remains a difficult task in nonstationary acoustic environments, mainly because current signal processing algorithms rely on fixed, manually tuned parameters that cannot adapt in situ to different users or listening contexts. This paper introduces a unified modular framework that formulates signal processing, learning, and personalization as Bayesian inference with explicit uncertainty tracking. The proposed framework replaces ad hoc algorithm design with a single probabilistic generative model that continuously adapts to changing acoustic conditions and user preferences. It extends spectral subtraction with principled mechanisms for in-situ personalization and adaptation to acoustic context. The system is implemented as an interconnected probabilistic state-space model, and inference is performed via variational message passing in the \texttt{RxInfer.jl} probabilistic programming environment, enabling real-time Bayesian processing under hearing-aid constraints. Proof-of-concept experiments on the \emph{VoiceBank+DEMAND} corpus show competitive speech quality and noise reduction with 85 effective parameters. The framework provides an interpretable, data-efficient foundation for uncertainty-aware, adaptive hearing-aid processing and points toward devices that learn continuously through probabilistic inference.


【4】Membership Inference Attacks against Large Audio Language Models
标题:针对大型音频语言模型的成员推断攻击
链接:https://arxiv.org/abs/2603.28378

作者:Jia-Kai Dong,Yu-Xiang Lin,Hung-Yi Lee
备注:submitted to Interspeech 2026
摘要:我们首次对大型音频语言模型(LALM)进行了系统的成员推理攻击(MIA)评估。由于音频编码非语义信息,它会导致严重的训练和测试分布偏移,并可能导致虚假的MIA性能。使用基于文本,频谱和韵律特征的多模态盲基线,我们证明了即使没有模型推断,常见的语音数据集也表现出近乎完美的训练/测试可分性(AUC约为1.0),并且标准MIA分数与这些盲声学伪影强烈相关(相关性大于0.7)。使用这个盲基线,我们确定分布匹配的数据集,使可靠的MIA评估没有分布偏移混淆。我们对多个MIA方法进行基准测试,并在这些数据集上进行模态解纠缠实验。结果表明,LALM记忆是跨模态的,只产生于绑定扬声器的语音身份与其文本。这些发现建立了一个原则性的标准,审计LALM超越虚假的相关性。
摘要:We present the first systematic Membership Inference Attack (MIA) evaluation of Large Audio Language Models (LALMs). As audio encodes non-semantic information, it induces severe train and test distribution shifts and can lead to spurious MIA performance. Using a multi-modal blind baseline based on textual, spectral, and prosodic features, we demonstrate that common speech datasets exhibit near-perfect train/test separability (AUC approximately 1.0) even without model inference, and the standard MIA scores strongly correlate with these blind acoustic artifacts (correlation greater than 0.7). Using this blind baseline, we identify that distribution-matched datasets enable reliable MIA evaluation without distribution shift confounds. We benchmark multiple MIA methods and conduct modality disentanglement experiments on these datasets. The results reveal that LALM memorization is cross-modal, arising only from binding a speaker's vocal identity with its text. These findings establish a principled standard for auditing LALMs beyond spurious correlations.


【5】On the Usefulness of Diffusion-Based Room Impulse Response Interpolation to Microphone Array Processing
标题:基于扩散的房间脉冲响应插值对麦克风阵列处理的有用性
链接:https://arxiv.org/abs/2603.28209

作者:Sagi Della Torre,Mirco Pezzoli,Fabio Antonacci,Sharon Gannot
摘要:房间冲激响应估计是空间域音频处理和语音增强中的一个基本问题。在本文中,我们建立在我们以前介绍的基于扩散的修补框架的房间脉冲响应插值,并证明其适用性,以提高实际的多麦克风阵列处理任务的性能。此外,我们验证了该方法的鲁棒性插值真实世界的房间脉冲响应。
摘要:Room Impulse Responses estimation is a fundamental problem in spatial audio processing and speech enhancement. In this paper, we build upon our previously introduced diffusion-based inpainting framework for Room Impulse Response interpolation and demonstrate its applicability to enhancing the performance of practical multi-microphone array processing tasks. Furthermore, we validate the robustness of this method in interpolating real-world Room Impulse Responses.


【6】MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
标题:MOSS-VoiceGenerator:使用自然语言描述创建真实的声音
链接:https://arxiv.org/abs/2603.28086

作者:Kexin Huang,Liwei Fan,Botian Jiang,Yaozhou Jiang,Qian Tu,Jie Zhu,Yuqian Zhang,Yiwei Zhao,Chenchen Yang,Zhaoye Fei,Shimin Li,Xiaogui Yang,Qinyuan Cheng,Xipeng Qiu
摘要:基于自然语言的语音设计旨在直接从自由形式的文本描述中生成扬声器音色,允许用户创建针对特定角色,个性和情感的声音。这种可控的语音创建有利于广泛的下游应用程序,包括讲故事,游戏配音,角色扮演代理和会话助理,使其成为现代文本到语音模型的重要任务。然而,现有的模型在很大程度上是在精心录制的录音室数据上训练的,这些数据产生的语音清晰、清晰,但缺乏真实人声的生活品质。为了解决这些限制,我们提出了MOSS-VoiceGenerator,这是一个开源的语音生成模型,可以直接从自然语言提示中创建新的音色。受到暴露于现实世界声学变化会产生感知上更自然的声音这一假设的启发,我们对来自电影内容的大规模表达性语音数据进行了训练。主观偏好的研究表明,它的整体性能的优越性,遵循谨慎,和自然相比,其他的声音设计模型。
摘要:Voice design from natural language aims to generate speaker timbres directly from free-form textual descriptions, allowing users to create voices tailored to specific roles, personalities, and emotions. Such controllable voice creation benefits a wide range of downstream applications-including storytelling, game dubbing, role-play agents, and conversational assistants, making it a significant task for modern Text-to-Speech models. However, existing models are largely trained on carefully recorded studio data, which produces speech that is clean and well-articulated, yet lacks the lived-in qualities of real human voices. To address these limitations, we present MOSS-VoiceGenerator, an open-source instruction-driven voice generation model that creates new timbres directly from natural language prompts. Motivated by the hypothesis that exposure to real-world acoustic variation produces more perceptually natural voices, we train on large-scale expressive speech data sourced from cinematic content. Subjective preference studies demonstrate its superiority in overall performance, instruction-following, and naturalness compared to other voice design models.


【7】Audio Language Model for Deepfake Detection Grounded in Acoustic Chain-of-Thought
标题:基于声学思想链的Deepfake检测音频语言模型
链接:https://arxiv.org/abs/2603.28021

作者:Runkun Chen,Yixiong Fang,Pengyu Chang,Yuante Li,Massa Baali,Bhiksha Ramakrishnan
摘要:Deepfake语音检测系统通常仅限于二进制分类任务,并且难以生成可解释的推理或为其决策提供丰富的上下文解释。这些模型主要提取用于真实性检测的潜在嵌入,但未能以有意义的方式利用结构化声学证据,例如韵律,频谱和生理属性。本文介绍了CoLMbo-DF,这是一种特征引导的音频语言模型,它通过将鲁棒的深度伪造检测与显式的声学思维链推理相集成来解决这些限制。通过将低级别声学特征的结构化文本表示直接注入到模型提示中,我们的方法使模型的推理基于可解释的证据,并提高了检测精度。为了支持这个框架,我们引入了一个新的音频对数据集,并与思维链注释配对。实验表明,我们的方法在轻量级开源语言模型上训练,尽管规模较小,但显著优于现有的音频语言模型基线,标志着可解释的deepfake语音检测的重大进步。
摘要:Deepfake speech detection systems are often limited to binary classification tasks and struggle to generate interpretable reasoning or provide context-rich explanations for their decisions. These models primarily extract latent embeddings for authenticity detection but fail to leverage structured acoustic evidence such as prosodic, spectral, and physiological attributes in a meaningful manner. This paper introduces CoLMbo-DF, a Feature-Guided Audio Language Model that addresses these limitations by integrating robust deepfake detection with explicit acoustic chain-of-thought reasoning. By injecting structured textual representations of low-level acoustic features directly into the model prompt, our approach grounds the model's reasoning in interpretable evidence and improves detection accuracy. To support this framework, we introduce a novel dataset of audio pairs paired with chain-of-thought annotations. Experiments show that our method, trained on a lightweight open-source language model, significantly outperforms existing audio language model baselines despite its smaller scale, marking a significant advancement in explainable deepfake speech detection.


【8】On the Role of Encoder Depth: Pruning Whisper and LoRA Fine-Tuning in SLAM-ASR
标题:关于编码器深度的作用:SLAM-ASB中的修剪Whisper和LoRA微调
链接:https://arxiv.org/abs/2603.27981

作者:Ganesh Pavan Kartikeya Bharadwaj Kolluri,Michael Kampouridis,Ravi Shekhar
备注:Accepted at SPEAKABLE Workshop, LREC 2026
摘要:近年来,自动语音识别(ASR)在大规模预训练模型和端到端架构(如SLAM-ASR)的推动下发展迅速。SLAM-ASR系统的一个关键组件是Whisper语音编码器,它提供了强大的声学表示。虽然已经针对完整的Whisper编码器-解码器架构探索了模型修剪,但其在SLAM-ASR设置中的影响仍然未得到充分研究。在这项工作中,我们分析了层修剪的Whisper编码器时,作为SLAM-ASR的声学骨干的效果。我们进一步研究了基于LoRA的微调可以在多大程度上恢复修剪引起的性能下降。在三种Whisper变体(Small,Medium,Large-v2),三种代表不同资源级别的语言(丹麦语,荷兰语,英语)以及超过200次训练运行中进行的实验表明,修剪两个编码器层仅导致2-4%的WER降级,并且将这种修剪与LoRA自适应相结合始终优于未修剪的基线,同时将总参数减少7- 14%。此外,我们的错误分析表明,LoRA主要通过语言模型的语言先验进行补偿,将荷兰语和英语的单词错误总数减少了11-21%,其中替换和删除的减少幅度最大。然而,对于低资源丹麦语,减少较小(4-7%),LoRA引入了增加的插入错误,这表明补偿有效性取决于LLM的预先存在的语言能力和可用的培训数据。
摘要:Automatic speech recognition (ASR) has advanced rapidly in recent years, driven by large-scale pretrained models and end-to-end architectures such as SLAM-ASR. A key component of SLAM-ASR systems is the Whisper speech encoder, which provides robust acoustic representations. While model pruning has been explored for the full Whisper encoder-decoder architecture, its impact within the SLAM-ASR setting remains under-investigated. In this work, we analyze the effects of layer pruning in the Whisper encoder when used as the acoustic backbone of SLAM-ASR. We further examine the extent to which LoRA-based fine-tuning can recover performance degradation caused by pruning. Experiments conducted across three Whisper variants (Small, Medium, Large-v2), three languages representing distinct resource levels (Danish, Dutch, English), and over 200 training runs demonstrate that pruning two encoder layers causes only 2-4% WER degradation, and that combining this pruning with LoRA adaptation consistently outperforms the unpruned baseline while reducing total parameters by 7-14%. Moreover, our error analysis reveals that LoRA primarily compensates through the language model's linguistic priors, reducing total word errors by 11-21% for Dutch and English, with substitutions and deletions showing the largest reductions. However, for low-resource Danish, the reduction is smaller (4-7%), and LoRA introduces increased insertion errors, indicating that compensation effectiveness depends on the LLM's pre-existing language proficiency and available training data.


【9】HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
标题:HumMusQA:一个人类编写的音乐理解QA基准数据集
链接:https://arxiv.org/abs/2603.27877

作者:Benno Weck,Pablo Puentes,Andrea Poltronieri,Satyajeet Prabhu,Dmitry Bogdanov
备注:Dataset available at https://doi.org/10.5281/zenodo.18462523
摘要:大型音频语言模型(LALM)中的音乐理解评估需要一个严格定义的基准,以真正测试模型是否可以感知和解释音乐,而当前的数据方法常常无法满足这一标准。本文介绍了一种精心构造的音乐评估方法,提出了一个由受过音乐训练的专家策划和验证的320个手写问题的新数据集,认为这种集中的手动策划对于探索复杂的音频理解是优越的。为了演示数据集的使用,我们对六个最先进的LALM进行了基准测试,并测试了它们对单峰捷径的鲁棒性。
摘要:The evaluation of music understanding in Large Audio-Language Models (LALMs) requires a rigorously defined benchmark that truly tests whether models can perceive and interpret music, a standard that current data methodologies frequently fail to meet. This paper introduces a meticulously structured approach to music evaluation, proposing a new dataset of 320 hand-written questions curated and validated by experts with musical training, arguing that such focused, manual curation is superior for probing complex audio comprehension. To demonstrate the use of the dataset, we benchmark six state-of-the-art LALMs and additionally test their robustness to uni-modal shortcuts.


【10】EvA: An Evidence-First Audio Understanding Paradigm for LALMs
标题:EvA:LALM的证据优先音频理解范式
链接:https://arxiv.org/abs/2603.27667

作者:Xinyuan Xie,Shunian Chen,Zhiheng Liu,Yuhao Zhang,Zhiqiang Lv,Liyin Liang,Benyou Wang
摘要:大型音频语言模型(LALM)仍然在复杂的声学场景中挣扎,因为它们通常无法在推理开始之前保留任务相关的声学证据。我们称之为证据瓶颈:国家的最先进的系统表现出更大的赤字,在证据提取比下游推理,这表明主要的限制在于上游的看法,而不是推理政策。为了解决这个问题,我们提出了EvA(Evidence-First Audio),这是一种双路径架构,通过非压缩、时间对齐的融合将Whisper和CED-Base结合在一起。EvA首先聚合中间CED层以保留多尺度声学线索,然后将聚合的CED特征与Whisper时间轴对齐,并在不改变序列长度的情况下添加两个流。我们还构建了EvA-Perception,这是一个大规模的开源训练集,包含约54 K事件排序的标题(150 h)和约500 K QA对。在统一的zero-shot协议下,EvA在MMAU、MMAR和MMSU上获得了最佳的开源感知分数,并在所有报告的指标上都比Kimi-Audio-7 B有所改进,在感知重分裂方面获得了最大的收益。这些结果支持证据优先的假设:更强的音频理解取决于在推理之前保留声学证据。
摘要:Large Audio Language Models (LALMs) still struggle in complex acoustic scenes because they often fail to preserve task-relevant acoustic evidence before reasoning begins. We call this failure the evidence bottleneck: state-of-the-art systems show larger deficits in evidence extraction than in downstream reasoning, suggesting that the main limitation lies in upstream perception rather than reasoning policy. To address this problem, we propose EvA (Evidence-First Audio), a dual-path architecture that combines Whisper and CED-Base through non-compressive, time-aligned fusion. EvA first aggregates intermediate CED layers to preserve multi-scale acoustic cues, then aligns the aggregated CED features to the Whisper timeline and adds the two streams without changing sequence length. We also build EvA-Perception, a large-scale open-source training set with about 54K event-ordered captions (150 h) and about 500K QA pairs. Under a unified zero-shot protocol, EvA achieves the best open-source Perception scores on MMAU, MMAR, and MMSU, and improves over Kimi-Audio-7B on all reported metrics, with the largest gains on perception-heavy splits. These results support the evidence-first hypothesis: stronger audio understanding depends on preserving acoustic evidence before reasoning.


【11】A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators
标题:Deepfake语音检测的通用模型:多样化的Bonafide资源或多样化的基于AI的生成器
链接:https://arxiv.org/abs/2603.27557

作者:Lam Pham,Khoi Vu,Dat Tran,David Fischinger,Simon Freitter,Marcel Hasenbalg,Davide Antonutti,Alexander Schindler,Martin Boyer,Ian McLoughlin
摘要:在本文中,我们分析了影响Deepfake语音检测(DSD)模型的性能和通用性的两个主要因素:Bonafide资源(BR)或基于AI的生成器(AG)。为此,我们首先提出了一个基于深度学习的模型,称为基线。然后,我们在基线上进行了实验,通过这些实验,我们表明了Bonafide Resource(BR)和基于AI的Generator(AG)因素如何影响用于在推理过程中检测虚假或真实输入音频的阈值分数。鉴于实验结果,提出了一种数据集,该数据集重用公共Deepfake语音检测(DSD)数据集,并在Bonafide资源(BR)或基于AI的生成器(AG)之间显示平衡。然后,我们在提出的数据集上训练各种基于深度学习的模型,并在不同的基准数据集上进行跨数据集评估。跨数据集评估结果证明,Bonafide资源(BR)和基于AI的生成器(AG)的平衡是训练和实现通用Deepfake语音检测(DSD)模型的关键因素。
摘要:In this paper, we analyze two main factors of Bonafide Resource (BR) or AI-based Generator (AG) which affect the performance and the generality of a Deepfake Speech Detection (DSD) model. To this end, we first propose a deep-learning based model, referred to as the baseline. Then, we conducted experiments on the baseline by which we indicate how Bonafide Resource (BR) and AI-based Generator (AG) factors affect the threshold score used to detect fake or bonafide input audio in the inference process. Given the experimental results, a dataset, which re-uses public Deepfake Speech Detection (DSD) datasets and shows a balance between Bonafide Resource (BR) or AI-based Generator (AG), is proposed. We then train various deep-learning based models on the proposed dataset and conduct cross-dataset evaluation on different benchmark datasets. The cross-dataset evaluation results prove that the balance of Bonafide Resources (BR) and AI-based Generators (AG) is the key factor to train and achieve a general Deepfake Speech Detection (DSD) model.


【12】Advancing Multi-Instrument Music Transcription: Results from the 2025 AMT Challenge
标题:推进多乐器音乐转录:2025年AMT挑战赛的结果
链接:https://arxiv.org/abs/2603.27528

作者:Ojas Chaturvedi,Kayshav Bhardwaj,Tanay Gondil,Benjamin Shiue-Hal Chou,Kristen Yeon-Ji Yun,Yung-Hsiang Lu,Yujia Yan,Sungkyun Chang
备注:7 pages, 3 figures. Accepted to the AI for Music Workshop at NeurIPS 2025
摘要:本文介绍了2025年自动音乐转录(AMT)挑战赛的结果,这是一项在线比赛,旨在对多乐器转录的进展进行基准测试。八个团队提交了有效的解决方案;两个团队的表现优于MT3基线模型。结果突出了两个进步的转录准确性和剩余的困难,在处理复调和音色变化。我们的结论与未来的挑战方向:更广泛的体裁覆盖面和更强的重视仪器检测。
摘要:This paper presents the results of the 2025 Automatic Music Transcription (AMT) Challenge, an online competition to benchmark progress in multi-instrument transcription. Eight teams submitted valid solutions; two outperformed the baseline MT3 model. The results highlight both advances in transcription accuracy and the remaining difficulties in handling polyphony and timbre variation. We conclude with directions for future challenges: broader genre coverage and stronger emphasis on instrument detection.


【13】Investigation on the Robustness of Acoustic Foundation Models on Post Exercise Speech
标题:声学基础模型对运动后言语的鲁棒性研究
链接:https://arxiv.org/abs/2603.27508

作者:Xiangyuan Xue,Yuyu Wang,Ruijie Yao,Xiaoyue Ni,Xiaofan Jiang,Jingping Nie
摘要:自动语音识别(ASR)已被广泛研究的中性和平稳的语音,但其鲁棒性下运动后的生理变化仍然是探索不足。与静息言语相比,运动后言语通常包含微呼吸、非语义停顿、不稳定的发声以及由呼吸支持减少引起的重复,使得转录更加困难。在这项工作中,我们基准声学基础模型运动后的讲话下一个统一的评估协议。我们比较了序列到序列模型(Whisper和FunASR/Paraformer)和自监督编码器与CTC解码(Wav 2 Vec 2,HuBERT和WavLM),在现成的推理和运动后域内微调下。在Static/Post-All基准中,大多数模型在运动后语音上会降级,而FunASR在Post-All上显示出最强的基线鲁棒性,WER为14.57%,CER为8.21%。微调大大改善了几个基于CTC的模型,而Whisper显示不稳定的适应。作为探索性案例研究,我们进一步按流利和不流利的说话者对结果进行分层;尽管不流利的子集很小,但它始终比流利的子集更具挑战性。总的来说,我们的研究结果表明,运动后ASR的鲁棒性是强烈的模型依赖性,在域适应可以是非常有效的,但不是一致的稳定,未来的运动后ASR研究应该明确分离的流畅性相关的影响,从运动引起的语音变化。
摘要:Automatic speech recognition (ASR) has been extensively studied on neutral and stationary speech, yet its robustness under post-exercise physiological shift remains underexplored. Compared with resting speech, post-exercise speech often contains micro-breaths, non-semantic pauses, unstable phonation, and repetitions caused by reduced breath support, making transcription more difficult. In this work, we benchmark acoustic foundation models on post-exercise speech under a unified evaluation protocol. We compare sequence-to-sequence models (Whisper and FunASR/Paraformer) and self-supervised encoders with CTC decoding (Wav2Vec2, HuBERT, and WavLM), under both off-the-shelf inference and post-exercise in-domain fine-tuning. Across the Static/Post-All benchmark, most models degrade on post-exercise speech, while FunASR shows the strongest baseline robustness at 14.57% WER and 8.21% CER on Post-All. Fine-tuning substantially improves several CTC-based models, whereas Whisper shows unstable adaptation. As an exploratory case study, we further stratify results by fluent and non-fluent speakers; although the non-fluent subset is small, it is consistently more challenging than the fluent subset. Overall, our findings show that post-exercise ASR robustness is strongly model-dependent, that in-domain adaptation can be highly effective but not uniformly stable, and that future post-exercise ASR studies should explicitly separate fluency-related effects from exercise-induced speech variation.


【14】TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba
标题:TokenDance:使用双向曼巴的代币到代币音乐到舞蹈一代
链接:https://arxiv.org/abs/2603.27314

作者:Ziyue Yang,Kaixing Yang,Xulong Tang
备注:CVPR2026 Workshop on HuMoGen
摘要:音乐舞蹈生成在虚拟现实、舞蹈教育和数字角色动画中有着广泛的应用。然而,现有的3D舞蹈数据集的覆盖范围有限,将当前的模型限制在音乐风格和舞蹈模式的狭窄子集上,导致对真实世界音乐的泛化能力较差。因此,生成的舞蹈往往变得过于简单和重复,大大降低了表现力和现实主义。   为了解决这个问题,我们提出了TokenDance,这是一个两阶段的音乐到舞蹈生成框架,通过双模态标记化和高效的标记级生成来明确解决这个限制。在第一阶段中,我们使用有限标量量化来离散舞蹈和音乐,其中舞蹈动作被分解为具有运动学动态约束的上半身和下半身组件,音乐被分解为具有专用码本的语义和声学特征以捕获特定于编舞的结构。在第二阶段,我们引入了一个基于双向Mamba主干的Local-Global-Local令牌到令牌生成器,实现了连贯的运动合成,强大的音乐舞蹈对齐和高效的非自回归推理。大量的实验表明,TokenDance在生成质量和推理速度方面都达到了整体最先进的(SOTA)性能,突出了其在现实世界音乐到舞蹈应用中的有效性和实用价值。
摘要:Music-to-dance generation has broad applications in virtual reality, dance education, and digital character animation. However, the limited coverage of existing 3D dance datasets confines current models to a narrow subset of music styles and choreographic patterns, resulting in poor generalization to real-world music. Consequently, generated dances often become overly simplistic and repetitive, substantially degrading expressiveness and realism.   To tackle this problem, we present TokenDance, a two-stage music-to-dance generation framework that explicitly addresses this limitation through dual-modality tokenization and efficient token-level generation. In the first stage, we discretize both dance and music using Finite Scalar Quantization, where dance motions are factorized into upper and lower-body components with kinematic-dynamic constraints, and music is decomposed into semantic and acoustic features with dedicated codebooks to capture choreography-specific structures. In the second stage, we introduce a Local-Global-Local token-to-token generator built on a Bidirectional Mamba backbone, enabling coherent motion synthesis, strong music-dance alignment, and efficient non-autoregressive inference. Extensive experiments demonstrate that TokenDance achieves overall state-of-the-art (SOTA) performance in both generation quality and inference speed, highlighting its effectiveness and practical value for real-world music-to-dance applications.


【15】Can pre-trained Deep Learning models predict groove ratings?
标题:预训练的深度学习模型可以预测凹槽评级吗?
链接:https://arxiv.org/abs/2603.27237

作者:Axel Marmoret,Nicolas Farrugia,Jan Alexander Stupacher
备注:Submitted to the SMC 2026 conference. 3 figures and 2 tables
摘要:这项研究探讨了深度学习模型直接从音频信号预测凹槽及其相关感知维度的程度。我们批判性地研究了七种最先进的深度学习模型在通过提取音频嵌入来预测凹槽评级和凹槽相关查询响应方面的有效性。此外,我们将这些预测与传统的手工制作的音频功能进行了比较。为了更好地理解潜在的机制,我们扩展了这种方法来分析基于源分离乐器的预测,从而隔离了单个音乐元素的贡献。我们的分析揭示了一个明确的分离槽的特点驱动的潜在的音乐风格的轨道(放克,流行,摇滚)。这些发现表明,深度音频表示可以成功地编码复杂的,依赖于风格的凹槽组件,传统的功能往往错过。最终,这项工作突出了先进的深度学习模型捕捉凹槽的多方面概念的能力,展示了表征学习推进预测音乐信息检索方法的强大潜力。
摘要:This study explores the extent to which deep learning models can predict groove and its related perceptual dimensions directly from audio signals. We critically examine the effectiveness of seven state-of-the-art deep learning models in predicting groove ratings and responses to groove-related queries through the extraction of audio embeddings. Additionally, we compare these predictions with traditional handcrafted audio features. To better understand the underlying mechanics, we extend this methodology to analyze predictions based on source-separated instruments, thereby isolating the contributions of individual musical elements. Our analysis reveals a clear separation of groove characteristics driven by the underlying musical style of the tracks (funk, pop, and rock). These findings indicate that deep audio representations can successfully encode complex, style-dependent groove components that traditional features often miss. Ultimately, this work highlights the capacity of advanced deep learning models to capture the multifaceted concept of groove, demonstrating the strong potential of representation learning to advance predictive Music Information Retrieval methodologies.


【16】Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis
标题:用于音乐结构分析的深度音频嵌入的无监督评估
链接:https://arxiv.org/abs/2603.27218

作者:Axel Marmoret
备注:Submitted to the SMC 2026 conference. 2 figures and 2 tables in the main document, 7 figures in Appendix
摘要:音乐结构分析(MSA)旨在揭示音乐作品的高层次组织。最先进的方法通常基于有监督的深度学习,但这些方法需要大量注释的数据和固有的结构模糊性。在本文中,我们提出了对MSA上九个开源、通用的预训练深度音频模型的无监督评估。对于每个模型,我们提取条形嵌入并使用三种无监督分割算法(Foote的棋盘核,谱聚类和相关块匹配(CBM))对其进行分割,专门关注边界检索。我们的研究结果表明,现代通用的深度嵌入通常优于传统的基于频谱的基线,但不是系统性的。此外,我们的无监督边界估计方法通常比最近的线性探测基线产生更强的性能。在评估的技术中,CBM算法始终是最有效的下游分割方法。最后,我们强调标准评估指标的人为膨胀,并主张系统地采用“修剪”,甚至“双修剪”注释,以建立更严格的MSA评估标准。
摘要:Music Structure Analysis (MSA) aims to uncover the high-level organization of musical pieces. State-of-the-art methods are often based on supervised deep learning, but these methods are bottlenecked by the need for heavily annotated data and inherent structural ambiguities. In this paper, we propose an unsupervised evaluation of nine open-source, generic pre-trained deep audio models, on MSA. For each model, we extract barwise embeddings and segment them using three unsupervised segmentation algorithms (Foote's checkerboard kernels, spectral clustering, and Correlation Block-Matching (CBM)), focusing exclusively on boundary retrieval. Our results demonstrate that modern, generic deep embeddings generally outperform traditional spectrogram-based baselines, but not systematically. Furthermore, our unsupervised boundary estimation methodology generally yields stronger performance than recent linear probing baselines. Among the evaluated techniques, the CBM algorithm consistently emerges as the most effective downstream segmentation method. Finally, we highlight the artificial inflation of standard evaluation metrics and advocate for the systematic adoption of ``trimming'', or even ``double trimming'' annotations to establish more rigorous MSA evaluation standards.


【17】Two-Stage Acoustic Adaptation with Gated Cross-Attention Adapters for LLM-Based Multi-Talker Speech Recognition
标题:基于LLM的多说话者语音识别的两阶段声学自适应和门控交叉注意适配器
链接:https://arxiv.org/abs/2603.27205

作者:Hao Shi,Yuan Gao,Xugang Lu,Tatsuya Kawahara
摘要:大型语言模型(LLM)是两个说话人自动语音识别(ASR)中串行输出训练(SOT)的强大解码器,但它们的性能在具有挑战性的条件下(如三个说话人混合)会大幅下降。一个关键的限制是,当前的系统仅通过投影前缀注入声学证据,这可能是有损的并且与LLM输入空间不完全对齐,从而在解码期间提供不足的细粒度接地。解决这一限制对于鲁棒的多讲话者ASR至关重要,特别是在三讲话者混合中。本文改进了基于LLM的多说话者ASR,显式地注入说话者感知的声学证据到解码器。我们首先回顾了连接主义时间分类(CTC)派生的前缀提示和比较三个变量增加声学内容。CTC信息是使用我们以前的作品中提出的序列化CTC获得的。虽然声音丰富的提示优于SOT只有基线,前缀只有条件仍然不足以为三个说话者的混合物。因此,我们提出了一个轻量级的门控残余交叉注意适配器,并设计了一个两阶段的声学适应框架的基础上低秩更新(LoRA)。在第一阶段,我们在自我注意子层之后插入门控交叉注意适配器,以稳定地注入声学嵌入作为外部存储器。在第二阶段,我们使用参数有效的LoRA来改进交叉注意适配器和预训练的LLM的自我注意投影,从而在有限的数据下提高大型骨干的鲁棒性;学习的更新被合并到基本权重中进行推理。Libri 2 Mix/Libri 3 Mix在干净和嘈杂条件下的实验显示出一致的增益,特别是在三个谈话者设置中有很大的改善。
摘要:Large Language Models (LLMs) are strong decoders for Serialized Output Training (SOT) in two-talker Automatic Speech Recognition (ASR), yet their performance degrades substantially in challenging conditions such as three-talker mixtures. A key limitation is that current systems inject acoustic evidence only through a projected prefix, which can be lossy and imperfectly aligned with the LLM input space, providing insufficient fine-grained grounding during decoding. Addressing this limitation is crucial for robust multi-talker ASR, especially in three-talker mixtures. This paper improves LLM-based multi-talker ASR by explicitly injecting talker-aware acoustic evidence into the decoder. We first revisit Connectionist Temporal Classification (CTC)-derived prefix prompting and compare three variants with increasing acoustic content. The CTC information is obtained using the serialized CTC proposed in our previous works. While acoustic-enriched prompts outperform the SOT-only baseline, prefix-only conditioning remains inadequate for three-talker mixtures. We therefore propose a lightweight gated residual cross-attention adapter and design a two-stage acoustic adaptation framework based on low-rank updates (LoRA). In Stage 1, we insert gated cross-attention adapters after the self-attention sub-layer to stably inject acoustic embeddings as external memory. In Stage 2, we refine both the cross-attention adapters and the pretrained LLM's self-attention projections using parameter-efficient LoRA, improving robustness for large backbones under limited data; the learned updates are merged into the base weights for inference. Experiments on Libri2Mix/Libri3Mix under clean and noisy conditions show consistent gains, with particularly large improvements in three-talker settings.


【18】Diachronic Modeling of Tonal Coherence on the Tonnetz Across Classical and Popular Repertoires
标题:古典和流行音乐剧Tonnetz音调连贯性的历时建模
链接:https://arxiv.org/abs/2603.27035

作者:Weilun Xu,Edward Hall,Martin Rohrmeier
摘要:不同的音乐传统如何实现音调的连贯性?迄今为止,大多数计算方法都是从单一维度分析音调的连贯性,而多维分析还没有得到充分的探索。我们提出了一个新的模型,借鉴的概念Tonnetz -我们定义了两个部分独立的措施:\n {音调焦点},音调中心附近的音高内容的浓度;和\n {音调连接},在何种程度上音高内容反映结构化的音程路径回到该中心。通过分析2,800多件西方古典和流行传统的作品,我们发现这些传统占据了二维空间的重叠但可区分的区域。流行音乐表现出更高的音调焦点,而古典音乐则表现出更高的音调连接。我们的补充措施接地不同的音调风格之间的差异,在定量证据,并提供可解释的尺寸计算音乐分析和可控的生成。
摘要:How do different musical traditions achieve tonal coherence? Most computational measures to date have analysed tonal coherence in terms of a single dimension, whereas a multi-dimensional analyses have not been sufficiently explored. We propose a new model drawing on the concept of the Tonnetz -- we define two partially independent measures: \emph{tonal focus}, the concentration of pitch content near a tonal center; and \emph{tonal connection}, the degree to which pitch content reflects structured intervallic pathways back to that center. Analyzing over 2,800 pieces from Western classical and popular traditions, we find that these traditions occupy overlapping yet distinguishable regions of the two-dimensional space. Popular music shows higher tonal focus, while classical music exhibits higher tonal connection. Our complementary measures ground the differences between different tonal styles in quantitative evidence, and offer interpretable dimensions for computational music analysis and controllable generation.


【19】Algo Pärt: An Algorithmic Reconstruction of Arvo Pärt's Summa
标题:阿尔戈·佩特:阿尔沃·佩特《全集》的数学重建
链接:https://arxiv.org/abs/2603.26989

作者:Bas Cornelissen
备注:21 pages, 15 figures
摘要:阿尔沃·佩特是当代最受欢迎的作曲家之一,以其高度原创的tintinnabuli风格而闻名。这种风格的作品通常是根据精确的程序组成的,甚至被描述为算法组成。为了理解佩特的音乐究竟是如何算法的,本文提出了一种综合分析:它提出了一种算法,几乎完全重建了Summa的乐谱,根据佩特自己在1994年的说法,这是他“最严格构造和最加密的作品”。这件作品被分析,然后使用所谓的tintinnabuli过程形式化。由此产生的算法的实施产生的乐谱匹配Summa超过93%的音符。由于声音之间的相互依赖性,只有一半的错误(3.5%)需要纠正,以忠实地再现原始乐谱。这项研究表明,Summa在很大程度上是一种算法组成,并提供了对Arvo Pärt音乐的新视角。
摘要:Arvo Pärt is one of the most popular contemporary composers, known for his highly original tintinnabuli style. Works in this style are typically composed according to precise procedures and have even been described as algorithmic compositions. To understand how algorithmic Pärt's music exactly is, this paper presents an analysis by synthesis: it proposes an algorithm that almost completely reconstructs the score of Summa, his "most strictly constructed and most encrypted work," according to Pärt himself in 1994. The piece is analyzed and then formalized using so-called tintinnabuli processes. An implementation of the resulting algorithm generates a musical score matching Summa in over 93% of the notes. Due to interdependencies between the voices, only half of the mistakes (3.5%) need to be corrected to reproduce the original score faithfully. This study shows that Summa is a largely algorithmic composition and offers new perspectives on the music of Arvo Pärt.


【20】Rhythmic segment analysis: Conceptualizing, visualizing, and measuring rhythmic data
标题:节奏片段分析:概念化,可视化和测量节奏数据
链接:https://arxiv.org/abs/2603.26988

作者:Bas Cornelissen
备注:15 pages, 7 figures
摘要:本文开发了一个框架,概念化,可视化,并测量节奏数据的重复性。我建议用间隔片段来考虑节奏数据:固定长度的连续间隔组,可以分解为持续时间和模式(间隔之间的比率)。这个简单的概念框架统一了三种节奏可视化方法,并产生了第四种:模式持续时间图。当与聚类转换网络配对时,它直观地揭示了合成和真实世界节奏数据中的重复性。此外,该框架概括了两种常见的节奏结构的措施:节奏比和归一化成对变异指数(nPVI)。特别是,nPVI可以重建为平均距离等时性,我提出了一个更一般的措施,anisochrony取代it.Finally,新的概念的数量可能会揭示更广泛的辩论小整数比节奏。
摘要:This paper develops a framework for conceptualizing, visualizing, and measuring regularities in rhythmic data. I propose to think about rhythmic data in terms of interval segments: fixed-length groups of consecutive intervals, which can be decomposed into a duration and a pattern (the ratios between the intervals). This simple conceptual framework unifies three rhythmic visualization methods and yields a fourth: the pattern-duration plot. When paired with a cluster transition network, it intuitively reveals regularities in both synthetic and real-world rhythmic data. Moreover, the framework generalizes two common measures of rhythmic structure: rhythm ratios and the normalized pairwise variability index (nPVI). In particular, nPVI can be reconstructed as the average distance from isochrony, and I propose a more general measure of anisochrony to replace it. Finally, the novel concept of quantality may shed light on wider debates regarding small-integer-ratio rhythms.


【21】Multilingual Stutter Event Detection for English, German, and Mandarin Speech
标题:英语、德语和普通话语音的多语言口吃事件检测
链接:https://arxiv.org/abs/2603.26939

作者:Felix Haas,Sebastian P. Bayerl
摘要:本文提出了一个基于英语、德语和汉语的多语言语料库的多标签口吃检测系统,该模型利用来自三种语言和四种语料库的带注释的口吃数据,捕捉与语言无关的口吃特征,实现跨语言背景的鲁棒检测。实验结果表明,多语言训练达到的性能相当,在某些情况下,甚至超过了以前的系统。这些发现表明,口吃表现出跨语言的一致性,这支持语言不可知检测系统的发展。我们的工作证明了使用多语言数据来提高自动口吃检测的通用性和可靠性的可行性和优势。
摘要:This paper presents a multi-label stuttering detection system trained on multi-corpus, multilingual data in English, German, and Mandarin.By leveraging annotated stuttering data from three languages and four corpora, the model captures language-independent characteristics of stuttering, enabling robust detection across linguistic contexts. Experimental results demonstrate that multilingual training achieves performance comparable to and, in some cases, even exceeds that of previous systems. These findings suggest that stuttering exhibits cross-linguistic consistency, which supports the development of language-agnostic detection systems. Our work demonstrates the feasibility and advantages of using multilingual data to improve generalizability and reliability in automated stuttering detection.


【22】AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
标题:AFSS:以人为本的自合成,用于减轻音频深度伪造检测中的偏差
链接:https://arxiv.org/abs/2603.26856

作者:Hai-Son Nguyen-Le,Hung-Cuong Nguyen-Thanh,Nhien-An Le-Khac,Dinh-Thuc Nguyen,Hong-Hanh Nguyen-Le
备注:Accepted at International Joint Conference on Neural Networks 2026
摘要:生成模型的快速发展使高度逼真的音频deepfake成为可能,但目前的检测器存在严重的偏差问题,导致在看不见的数据集上泛化能力差。本文提出了伪像聚焦自合成(AFSS),一种旨在通过两种机制从真实音频生成伪假样本来减轻这种偏差的方法:自转换和自重建。AFSS的核心思想在于强制执行同一说话人约束,确保真实和伪假样本共享相同的说话人身份和语义内容。这迫使检测器专门关注生成伪影,而不是不相关的混淆因素。此外,我们引入了一个可学习的重新加权损失,在训练过程中动态强调合成样本。在7个数据集上进行的大量实验表明,AFSS实现了最先进的性能,平均EER为5.45%,包括在WaveFake上显著降低到1.23%,在In-the-Wild上降低到2.70%,同时消除了对预先收集的假数据集的依赖。我们的代码可在https://github.com/NguyenLeHaiSonGit/AFSS上公开获取。
摘要:The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused Self-Synthesis (AFSS), a method designed to mitigate this bias by generating pseudo-fake samples from real audio via two mechanisms: self-conversion and self-reconstruction. The core insight of AFSS lies in enforcing same-speaker constraints, ensuring that real and pseudo-fake samples share identical speaker identity and semantic content. This forces the detector to focus exclusively on generation artifacts rather than irrelevant confounding factors. Furthermore, we introduce a learnable reweighting loss to dynamically emphasize synthetic samples during training. Extensive experiments across 7 datasets demonstrate that AFSS achieves state-of-the-art performance with an average EER of 5.45\%, including a significant reduction to 1.23\% on WaveFake and 2.70\% on In-the-Wild, all while eliminating the dependency on pre-collected fake datasets. Our code is publicly available at https://github.com/NguyenLeHaiSonGit/AFSS.


【23】ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
标题:ParaSpeechCLAP:一种用于丰富风格语音音频预训练的双编码器语音文本模型
链接:https://arxiv.org/abs/2603.28737

作者:Anuj Diwan,Eunsol Choi,David Harwath
备注:Under review
摘要:我们引入ParaSpeechCLAP,一个双编码器对比模型,将语音和文本风格的字幕映射到一个共同的嵌入空间,支持广泛的内在(扬声器级)和情境(话语级)描述符(如音高,纹理和情感)远远超出了现有模型处理的狭窄集合。我们训练专门的ParaSpeechCLAP-Intrinsic和ParaSpeechCLAP-Situational模型以及统一的ParaSpeechCLAP-Combined模型,发现专业化在个人风格维度上产生更强的性能,而统一模型在成分评估上表现出色。我们进一步表明,ParaSpeechCLAP-Intrinsic受益于额外的分类损失和类平衡训练。我们展示了我们的模型在风格标题检索,语音属性分类和作为一个推理时间奖励模型,提高风格提示的TTS没有额外的训练的性能。ParaSpeechCLAP在所有三个应用程序的大多数指标上都优于基线。我们的模型和代码在https://github.com/ajd12342/paraspeechclap上发布。
摘要:We introduce ParaSpeechCLAP, a dual-encoder contrastive model that maps speech and text style captions into a common embedding space, supporting a wide range of intrinsic (speaker-level) and situational (utterance-level) descriptors (such as pitch, texture and emotion) far beyond the narrow set handled by existing models. We train specialized ParaSpeechCLAP-Intrinsic and ParaSpeechCLAP-Situational models alongside a unified ParaSpeechCLAP-Combined model, finding that specialization yields stronger performance on individual style dimensions while the unified model excels on compositional evaluation. We further show that ParaSpeechCLAP-Intrinsic benefits from an additional classification loss and class-balanced training. We demonstrate our models' performance on style caption retrieval, speech attribute classification and as an inference-time reward model that improves style-prompted TTS without additional training. ParaSpeechCLAP outperforms baselines on most metrics across all three applications. Our models and code are released at https://github.com/ajd12342/paraspeechclap .


【24】SHroom: A Python Framework for Ambisonics Room Acoustics Simulation and Binaural Rendering
标题:SHroom:用于Ambisonics房间声学模拟和双耳渲染的Python框架
链接:https://arxiv.org/abs/2603.27342

作者:Yhonatan Gayer
摘要:Spherical Harmonics ROOM),一个使用Ambisonics进行室内声学模拟的开源Python库,可在https://github.com/Yhonatangayer/shroom上获得,并可通过\texttt{pip install pyshroom}安装。\textbf{shroom}将图像源贡献投影到球面谐波(SH)基础上,从而产生用于双耳解码、球面阵列模拟和实时头部旋转的可组合管道。以$N=30$参考的\texttt{pyroomacoustics}为基准,\textbf{shroom}与幅度最小二乘法(MagLS)实现了感知透明性(2.02~dB对数频谱距离(LSD),在$N=5$时,在1- 2 ~dB最小可察觉差异(JND)内),而其固定一次解码在多个源上分摊($K=1$-to-8 $:减速从$7\times$缩小到$3.1\times$)。对于动态头部旋转,\textbf{shroom}以$<1$~ms/帧的速度应用Wigner-D乘法,使其成为唯一架构上可行的实时选择。
摘要:Spherical Harmonics ROOM), an open-source Python library for room acoustics simulation using Ambisonics, available at https://github.com/Yhonatangayer/shroom and installable via \texttt{pip install pyshroom}. \textbf{shroom} projects image-source contributions onto a Spherical Harmonics (SH) basis, yielding a composable pipeline for binaural decoding, spherical array simulation, and real-time head rotation. Benchmarked against \texttt{pyroomacoustics} with an $N=30$ reference, \textbf{shroom} with Magnitude Least Squares (MagLS) achieves perceptual transparency (2.02~dB Log Spectral Distance (LSD) at $N=5$, within the 1--2~dB Just Noticeable Difference (JND)) while its fixed-once decode amortises over multiple sources ($K=1$-to-$8$: slowdown narrows from $7\times$ to $3.1\times$). For dynamic head rotation, \textbf{shroom} applies a Wigner-D multiply at $<1$~ms/frame, making it the only architecturally viable real-time choice.


【25】HASS: Hierarchical Simulation of Logopenic Aphasic Speech for Scalable PPA Detection
标题:HASS:逻辑开放失语语音的分层模拟,用于可扩展PPA检测
链接:https://arxiv.org/abs/2603.26795

作者:Harrison Li,Kevin Wang,Cheol Jun Cho,Jiachen Lian,Rabab Rangwala,Chenxu Guo,Emma Yang,Lynn Kurteff,Zoe Ezzes,Willa Keegan-Rodewald,Jet Vonk,Siddarth Ramkrishnan,Giada Antonicelli,Zachary Miller,Marilu Gorno Tempini,Gopala Anumanchipalli
摘要:由于数据稀缺,建立原发性进行性失语症(PPA)的诊断模型一直具有挑战性。大规模收集临床数据受到临床人群的高度脆弱性和专家标签的高成本的限制。为了避免这一点,以前的研究模拟不流利的语音生成训练数据。然而,这些方法不够全面,无法将PPA模拟为整体的、多水平的表型,而是依赖于孤立的表达障碍。为了解决这个问题,我们提出了一个新的,临床接地模拟框架,层次失语症语音模拟(HASS)。HASS旨在模拟具有不同严重程度的PPA的逻辑缺失变体(lvPPA)的行为。为此,语义,语音,和时间赤字lvPPA系统地确定由临床专家,和模拟。我们证明了我们的框架可以实现更准确和更通用的检测模型。
摘要:Building a diagnosis model for primary progressive aphasia (PPA) has been challenging due to the data scarcity. Collecting clinical data at scale is limited by the high vulnerability of clinical population and the high cost of expert labeling. To circumvent this, previous studies simulate dysfluent speech to generate training data. However, those approaches are not comprehensive enough to simulate PPA as holistic, multi-level phenotypes, instead relying on isolated dysfluencies. To address this, we propose a novel, clinically grounded simulation framework, Hierarchical Aphasic Speech Simulation (HASS). HASS aims to simulate behaviors of logopenic variant of PPA (lvPPA) with varying degrees of severity. To this end, semantic, phonological, and temporal deficits of lvPPA are systematically identified by clinical experts, and simulated. We demonstrate that our framework enables more accurate and generalizable detection models.


eess.AS音频处理


【1】ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
标题:ParaSpeechCLAP:一种用于丰富风格语音音频预训练的双编码器语音文本模型
链接:https://arxiv.org/abs/2603.28737

作者:Anuj Diwan,Eunsol Choi,David Harwath
备注:Under review
摘要:我们引入ParaSpeechCLAP,一个双编码器对比模型,将语音和文本风格的字幕映射到一个共同的嵌入空间,支持广泛的内在(扬声器级)和情境(话语级)描述符(如音高,纹理和情感)远远超出了现有模型处理的狭窄集合。我们训练专门的ParaSpeechCLAP-Intrinsic和ParaSpeechCLAP-Situational模型以及统一的ParaSpeechCLAP-Combined模型,发现专业化在个人风格维度上产生更强的性能,而统一模型在成分评估上表现出色。我们进一步表明,ParaSpeechCLAP-Intrinsic受益于额外的分类损失和类平衡训练。我们展示了我们的模型在风格标题检索,语音属性分类和作为一个推理时间奖励模型,提高风格提示的TTS没有额外的训练的性能。ParaSpeechCLAP在所有三个应用程序的大多数指标上都优于基线。我们的模型和代码在https://github.com/ajd12342/paraspeechclap上发布。
摘要:We introduce ParaSpeechCLAP, a dual-encoder contrastive model that maps speech and text style captions into a common embedding space, supporting a wide range of intrinsic (speaker-level) and situational (utterance-level) descriptors (such as pitch, texture and emotion) far beyond the narrow set handled by existing models. We train specialized ParaSpeechCLAP-Intrinsic and ParaSpeechCLAP-Situational models alongside a unified ParaSpeechCLAP-Combined model, finding that specialization yields stronger performance on individual style dimensions while the unified model excels on compositional evaluation. We further show that ParaSpeechCLAP-Intrinsic benefits from an additional classification loss and class-balanced training. We demonstrate our models' performance on style caption retrieval, speech attribute classification and as an inference-time reward model that improves style-prompted TTS without additional training. ParaSpeechCLAP outperforms baselines on most metrics across all three applications. Our models and code are released at https://github.com/ajd12342/paraspeechclap .


【2】Acoustic-to-articulatory Inversion of the Complete Vocal Tract from RT-MRI with Various Audio Embeddings and Dataset Sizes
标题:具有各种音频嵌入和数据集大小的RT-MRI完整气道的声学-发音倒置
链接:https://arxiv.org/abs/2603.28723

作者:Sofiane Azzouz,Pierre-André Vuissoz,Yves Laprie
摘要:关节声波反演强烈依赖于所使用的数据类型。虽然大多数以前的研究依赖于EMA,这是有限的传感器的数量和限制可访问的发音器官,我们提出了一种方法,旨在在一个完整的反向声道,从声门到嘴唇。为此,我们使用了来自单个扬声器的大约3.5小时的RT-MRI数据。我们的方法的创新在于使用自动提取的咬合架轮廓从MRI图像,而不是依赖于原始图像本身。通过关注这些轮廓,该模型优先考虑声道的基本几何动态,同时丢弃冗余的像素级信息。然后使用Bi-LSTM架构处理这些轮廓以及去噪音频。进行了两个实验:(1)分析音频嵌入的影响,其中评估了三种类型的嵌入作为模型的输入(MFCC,LCC和HuBERT),以及(2)研究数据集大小的影响,我们从10分钟到3.5小时不等。使用RMSE、中位误差以及道变量对测试数据进行评估,我们向其添加了额外的测量:喉高度。与像素尺寸(1.62mm)相比,得到的平均RMSE为1.48mm。这些结果证实了使用RT-MRI数据进行完整的声道反转的可行性。
摘要:Articulatory-to-acoustic inversion strongly depends on the type of data used. While most previous studies rely on EMA, which is limited by the number of sensors and restricted to accessible articulators, we propose an approach aiming at a complete inversion of the vocal tract, from the glottis to the lips. To this end, we used approximately 3.5 hours of RT-MRI data from a single speaker. The innovation of our approach lies in the use of articulator contours automatically extracted from MRI images, rather than relying on the raw images themselves. By focusing on these contours, the model prioritizes the essential geometric dynamics of the vocal tract while discarding redundant pixel-level information. These contours, alongside denoised audio, were then processed using a Bi-LSTM architecture. Two experiments were conducted: (1) the analysis of the impact of the audio embedding, for which three types of embeddings were evaluated as input to the model (MFCCs, LCCs, and HuBERT), and (2) the study of the influence of the dataset size, which we varied from 10 minutes to 3.5 hours. Evaluation was performed on the test data using RMSE, median error, as well as Tract Variables, to which we added an additional measurement: the larynx height. The average RMSE obtained is 1.48\,mm, compared with the pixel size (1.62\,mm). These results confirm the feasibility of a complete vocal-tract inversion using RT-MRI data.


【3】Can Hierarchical Cross-Modal Fusion Predict Human Perception of AI Dubbed Content?
标题:分层跨模式融合能否预测人类对人工智能配音内容的感知?
链接:https://arxiv.org/abs/2603.28717

作者:Ashwini Dasare,Nirmesh Shah,Ashishkumar Gudmalwar,Pankaj Wasnik
备注:Accepted at ICASSP 2026
摘要:评估人工智能生成的配音内容本质上是多维的,由同步,可理解性,说话者一致性,情感对齐和语义上下文形成。人类平均意见评分(MOS)仍然是黄金标准,但成本高昂,规模不切实际。我们提出了一个层次化的多模式架构,感知有意义的配音评估,整合从音频,视频和文本的互补线索。该模型捕获细粒度的特征,如说话人身份,韵律和音频,面部表情和场景级线索从视频和文本的语义上下文,这是逐步融合通过内部和模态间层的内容。轻量级LoRA适配器可实现跨模态的参数高效微调。为了克服有限的主观标签,我们通过聚合客观指标与通过主动学习优化的权重来获得代理MOS。所提出的架构在12 k印地语-英语双向配音片段上进行了训练,然后使用人类MOS进行微调。我们的方法实现了强大的感知对齐(PCC > 0.75),为AI配音内容的自动评估提供了可扩展的解决方案。
摘要:Evaluating AI generated dubbed content is inherently multi-dimensional, shaped by synchronization, intelligibility, speaker consistency, emotional alignment, and semantic context. Human Mean Opinion Scores (MOS) remain the gold standard but are costly and impractical at scale. We present a hierarchical multimodal architecture for perceptually meaningful dubbing evaluation, integrating complementary cues from audio, video, and text. The model captures fine-grained features such as speaker identity, prosody, and content from audio, facial expressions and scene-level cues from video and semantic context from text, which are progressively fused through intra and inter-modal layers. Lightweight LoRA adapters enable parameter-efficient fine-tuning across modalities. To overcome limited subjective labels, we derive proxy MOS by aggregating objective metrics with weights optimized via active learning. The proposed architecture was trained on 12k Hindi-English bidirectional dubbed clips, followed by fine-tuning with human MOS. Our approach achieves strong perceptual alignment (PCC > 0.75), providing a scalable solution for automatic evaluation of AI-dubbed content.


【4】VAANI: Capturing the language landscape for an inclusive digital India
标题:VaANI:捕捉语言格局,打造包容性数字印度
链接:https://arxiv.org/abs/2603.28714

作者:Sujith Pulikodan,Abhayjeet Singh,Agneedh Basu,Lokesh Rady,Nihar Desai,Pavan Kumar J,Prajjwal Srivastav,Pranav D Bhat,Raghu Dharmaraju,Ritika Gupta,Sathvik Udupa,Saurabh Kumar,Sumit Sharma,Vaibhav Vishwakarma,Visruth Sanka,Dinesh Tewari,Harsh Dhand,Amrita Kamat,Sukhwinder Singh,Shikhar Vashishth,Partha Talukdar,Raj Acharya,Prasanta Kumar Ghosh
摘要:VAANI项目旨在创建一个具有印度代表性的多模态数据集,全面绘制印度的语言多样性,在前两个阶段从全国165个地区开始。语音数据是通过一个精心设计的过程收集的,该过程使用基于图像的提示来鼓励自发的反应。图像是通过一个单独的过程,其中包括广泛的主题,从内部和跨地区收集。收集的数据经过严格的多阶段质量评估,包括自动和手动检查,以确保音频质量和转录准确性的最高标准。经过彻底的验证,我们已经开放了大约289K的图像,大约31,270小时的录音,以及大约2,067小时的转录语音,包括来自31个州和联邦领土的165个地区的112种语言。值得注意的是,这些语言中的重要语言首次出现在这种规模的数据集中,使VAANI项目成为保护和促进语言包容性的开创性努力。这些数据有助于为印度构建包容性的语音模型,并推动语音、图像和多模式应用的研究和开发。
摘要:Project VAANI is an initiative to create an India-representative multi-modal dataset that comprehensively maps India's linguistic diversity, starting with 165 districts across the country in its first two phases. Speech data is collected through a carefully structured process that uses image-based prompts to encourage spontaneous responses. Images are captured through a separate process that encompasses a broad range of topics, gathered from both within and across districts. The collected data undergoes a rigorous multi-stage quality evaluation, including both automated and manual checks to ensure highest possible standards in audio quality and transcription accuracy. Following this thorough validation, we have open-sourced around 289K images, approximately 31,270 hours of audio recordings, and around 2,067 hours of transcribed speech, encompassing 112 languages from 165 districts from 31 States and Union territories. Notably, significant of these languages are being represented for the first time in a dataset of this scale, making the VAANI project a groundbreaking effort in preserving and promoting linguistic inclusivity. This data can be instrumental in building inclusive speech models for India, and in advancing research and development across speech, image, and multimodal applications.


【5】BiFormer3D: Grid-Free Time-Domain Reconstruction of Head-Related Impulse Responses with a Spatially Encoded Transformer
标题:BiFormer 3D:使用空间编码Transformer进行头部相关脉冲响应的无网格时间域重建
链接:https://arxiv.org/abs/2603.27998

作者:Shaoheng Xu,Chunyi Sun,Jihui Zhang,Amy Bastine,Prasanga N. Samarasinghe,Thushara D. Abhayapala,Hongdong Li
备注:The paper was submitted for review to Interspeech 2026
摘要:个性化的头部相关脉冲响应(HRIR)使双耳渲染,但密集的每个听众的测量是昂贵的。我们解决HRIR空间上采样从稀疏的每一个听众的测量:给定几个测量的HRIR的听众,预测HRIR在未测量的目标方向。现有的学习方法通常在频域工作,依赖于最小相位假设或单独的定时模型,并使用固定方向网格,这会降低时间保真度和空间连续性。我们提出了BiFormer 3D,一个时域,无网格双耳Transformer重建HRIR在任意方向从稀疏输入。它使用正弦空间特征、Conv 1D细化模块以及辅助耳间时间差(ITD)和耳间电平差(ILD)头。在超声心动图上,它改善了归一化均方误差(NMSE)、余弦距离和ITD/ILD误差;消融验证了模块,并表明不需要最小相位预处理。
摘要:Individualized head-related impulse responses (HRIRs) enable binaural rendering, but dense per-listener measurements are costly. We address HRIR spatial up-sampling from sparse per-listener measurements: given a few measured HRIRs for a listener, predict HRIRs at unmeasured target directions. Prior learning methods often work in the frequency domain, rely on minimum-phase assumptions or separate timing models, and use a fixed direction grid, which can degrade temporal fidelity and spatial continuity. We propose BiFormer3D, a time-domain, grid-free binaural Transformer for reconstructing HRIRs at arbitrary directions from sparse inputs. It uses sinusoidal spatial features, a Conv1D refinement module, and auxiliary interaural time difference (ITD) and interaural level difference (ILD) heads. On SONICOM, it improves normalized mean squared error (NMSE), cosine distance, and ITD/ILD errors over prior methods; ablations validate modules and show minimum-phase pre-processing is unnecessary.


【6】SHroom: A Python Framework for Ambisonics Room Acoustics Simulation and Binaural Rendering
标题:SHroom:用于Ambisonics房间声学模拟和双耳渲染的Python框架
链接:https://arxiv.org/abs/2603.27342

作者:Yhonatan Gayer
摘要:Spherical Harmonics ROOM),一个使用Ambisonics进行室内声学模拟的开源Python库,可在https://github.com/Yhonatangayer/shroom上获得,并可通过\texttt{pip install pyshroom}安装。\textbf{shroom}将图像源贡献投影到球面谐波(SH)基础上,从而产生用于双耳解码、球面阵列模拟和实时头部旋转的可组合管道。以$N=30$参考的\texttt{pyroomacoustics}为基准,\textbf{shroom}与幅度最小二乘法(MagLS)实现了感知透明性(2.02~dB对数频谱距离(LSD),$N=5$,在1- 2 ~dB最小可察觉差异(JND)内),而其固定一次解码在多个源上进行摊销($K=1$-to-8 $:减速从$7\times$缩小到$3.1\times$)。对于动态头部旋转,\textbf{shroom}以$<1$~ms/帧的速度应用Wigner-D乘法,使其成为唯一架构上可行的实时选择。
摘要:Spherical Harmonics ROOM), an open-source Python library for room acoustics simulation using Ambisonics, available at https://github.com/Yhonatangayer/shroom and installable via \texttt{pip install pyshroom}. \textbf{shroom} projects image-source contributions onto a Spherical Harmonics (SH) basis, yielding a composable pipeline for binaural decoding, spherical array simulation, and real-time head rotation. Benchmarked against \texttt{pyroomacoustics} with an $N=30$ reference, \textbf{shroom} with Magnitude Least Squares (MagLS) achieves perceptual transparency (2.02~dB Log Spectral Distance (LSD) at $N=5$, within the 1--2~dB Just Noticeable Difference (JND)) while its fixed-once decode amortises over multiple sources ($K=1$-to-$8$: slowdown narrows from $7\times$ to $3.1\times$). For dynamic head rotation, \textbf{shroom} applies a Wigner-D multiply at $<1$~ms/frame, making it the only architecturally viable real-time choice.


【7】PHONOS: PHOnetic Neutralization for Online Streaming Applications
标题:PHONOS:在线流媒体应用程序的PHOETIC中和
链接:https://arxiv.org/abs/2603.27001

作者:Waris Quamer,Mu-Ruei Tseng,Ghady Nasrallah,Ricardo Gutierrez-Osuna
备注:The paper is submitted to Interspeech 2026 and currently under review
摘要:说话人匿名化(SA)系统修改音色,同时保持区域或非本地口音不变,这是有问题的,因为口音可以缩小匿名集。为了解决这个问题,我们提出PHONOS,实时SA的流模块,中和非本地口音听起来像本地人。我们的方法预生成黄金扬声器的话语,保留源音色和节奏,但取代外国段与本地的沉默意识DTW对齐和zero-shot语音转换。这些话语监督一个因果口音翻译器,该翻译器将非原生内容令牌映射到原生等价物,最多40 ms的前瞻,使用联合交叉熵和CTC损失进行训练。我们的评估显示,非母语口音置信度降低了81%,语音测试评分与此转变一致,并且随着口音中立的话语在嵌入空间中远离原始说话者而降低了说话者的可链接性,同时在单个GPU上的延迟低于241 ms。
摘要:Speaker anonymization (SA) systems modify timbre while leaving regional or non-native accents intact, which is problematic because accents can narrow the anonymity set. To address this issue, we present PHONOS, a streaming module for real-time SA that neutralizes non-native accent to sound native-like. Our approach pre-generates golden speaker utterances that preserve source timbre and rhythm but replace foreign segmentals with native ones using silence-aware DTW alignment and zero-shot voice conversion. These utterances supervise a causal accent translator that maps non-native content tokens to native equivalents with at most 40ms look-ahead, trained using joint cross-entropy and CTC losses. Our evaluations show an 81% reduction in non-native accent confidence, with listening-test ratings consistent with this shift, and reduced speaker linkability as accent-neutralized utterances move away from the original speaker in embedding space while having latency under 241 ms on single GPU.


【8】Dual-branch Graph Domain Adaptation for Cross-scenario Multi-modal Emotion Recognition
标题:跨场景多模式情感识别的双分支图域自适应
链接:https://arxiv.org/abs/2603.26840

作者:Yuntao Shou,Jun Zhou,Tao Meng,Wei Ai,Keqin Li
备注:29 pages
摘要:会话多模态情感识别(MERC)旨在通过文本、音频和视觉线索来预测多轮对话中说话者的情感状态。在现实世界中,对话场景在扬声器,主题,风格和噪音水平方面存在显着差异。现有的MERC方法通常忽略了这些跨场景的变化,限制了它们将在源域上训练的模型转移到看不见的目标域的能力。为了解决这个问题,我们提出了一个双分支图域自适应框架(DGDA)跨场景条件下的多模态情感识别。我们首先构建了一个情感交互图来描述话语之间复杂的情感依赖关系。一个双分支编码器,由一个超图神经网络(HGNN)和一个路径神经网络(PathNN),然后设计显式建模多元关系和隐式捕获全局依赖。为了实现域外泛化,引入了域对抗学习器来学习跨域的不变表示。此外,正则化损失被纳入抑制噪声标签的负面影响。据我们所知,DGDA是第一个共同解决域偏移和标签噪声的MERC框架。理论分析提供了更严格的泛化边界,IEMOCAP和MELD上的大量实验表明,DGDA始终优于强基线,更好地适应跨场景会话。我们的代码可在www.example.com上获得。
摘要:Multimodal Emotion Recognition in Conversations (MERC) aims to predict speakers' emotional states in multi-turn dialogues through text, audio, and visual cues. In real-world settings, conversation scenarios differ significantly in speakers, topics, styles, and noise levels. Existing MERC methods generally neglect these cross-scenario variations, limiting their ability to transfer models trained on a source domain to unseen target domains. To address this issue, we propose a Dual-branch Graph Domain Adaptation framework (DGDA) for multimodal emotion recognition under cross-scenario conditions. We first construct an emotion interaction graph to characterize complex emotional dependencies among utterances. A dual-branch encoder, consisting of a hypergraph neural network (HGNN) and a path neural network (PathNN), is then designed to explicitly model multivariate relationships and implicitly capture global dependencies. To enable out-of-domain generalization, a domain adversarial discriminator is introduced to learn invariant representations across domains. Furthermore, a regularization loss is incorporated to suppress the negative influence of noisy labels. To the best of our knowledge, DGDA is the first MERC framework that jointly addresses domain shift and label noise. Theoretical analysis provides tighter generalization bounds, and extensive experiments on IEMOCAP and MELD demonstrate that DGDA consistently outperforms strong baselines and better adapts to cross-scenario conversations. Our code is available at https://github.com/Xudmm1239439/DGDA-Net.


【9】HASS: Hierarchical Simulation of Logopenic Aphasic Speech for Scalable PPA Detection
标题:HASS:逻辑开放失语语音的分层模拟,用于可扩展PPA检测
链接:https://arxiv.org/abs/2603.26795

作者:Harrison Li,Kevin Wang,Cheol Jun Cho,Jiachen Lian,Rabab Rangwala,Chenxu Guo,Emma Yang,Lynn Kurteff,Zoe Ezzes,Willa Keegan-Rodewald,Jet Vonk,Siddarth Ramkrishnan,Giada Antonicelli,Zachary Miller,Marilu Gorno Tempini,Gopala Anumanchipalli
摘要:由于数据稀缺,建立原发性进行性失语症(PPA)的诊断模型一直具有挑战性。大规模收集临床数据受到临床人群的高度脆弱性和专家标签的高成本的限制。为了避免这一点,以前的研究模拟不流利的语音生成训练数据。然而,这些方法不够全面,无法将PPA模拟为整体的、多水平的表型,而是依赖于孤立的表达障碍。为了解决这个问题,我们提出了一个新的,临床接地模拟框架,层次失语症语音模拟(HASS)。HASS旨在模拟具有不同严重程度的PPA的逻辑缺失变体(lvPPA)的行为。为此,语义,语音,和时间赤字lvPPA系统地确定由临床专家,和模拟。我们证明了我们的框架可以实现更准确和更通用的检测模型。
摘要:Building a diagnosis model for primary progressive aphasia (PPA) has been challenging due to the data scarcity. Collecting clinical data at scale is limited by the high vulnerability of clinical population and the high cost of expert labeling. To circumvent this, previous studies simulate dysfluent speech to generate training data. However, those approaches are not comprehensive enough to simulate PPA as holistic, multi-level phenotypes, instead relying on isolated dysfluencies. To address this, we propose a novel, clinically grounded simulation framework, Hierarchical Aphasic Speech Simulation (HASS). HASS aims to simulate behaviors of logopenic variant of PPA (lvPPA) with varying degrees of severity. To this end, semantic, phonological, and temporal deficits of lvPPA are systematically identified by clinical experts, and simulated. We demonstrate that our framework enables more accurate and generalizable detection models.


【10】Can pre-trained Deep Learning models predict groove ratings?
标题:预训练的深度学习模型可以预测凹槽评级吗?
链接:https://arxiv.org/abs/2603.27237

作者:Axel Marmoret,Nicolas Farrugia,Jan Alexander Stupacher
备注:Submitted to the SMC 2026 conference. 3 figures and 2 tables
摘要:这项研究探讨了深度学习模型直接从音频信号预测凹槽及其相关感知维度的程度。我们批判性地检查了七种最先进的深度学习模型在通过提取音频嵌入来预测凹槽评级和对凹槽相关查询的响应方面的有效性。此外,我们将这些预测与传统的手工制作的音频功能进行了比较。为了更好地理解潜在的机制,我们扩展了这种方法来分析基于源分离乐器的预测,从而隔离了单个音乐元素的贡献。我们的分析揭示了一个明确的分离槽的特点驱动的潜在的音乐风格的轨道(放克,流行,摇滚)。这些发现表明,深度音频表示可以成功地编码复杂的,依赖于风格的凹槽组件,传统的功能往往错过。最终,这项工作突出了先进的深度学习模型捕捉凹槽的多方面概念的能力,展示了表征学习推进预测音乐信息检索方法的强大潜力。
摘要:This study explores the extent to which deep learning models can predict groove and its related perceptual dimensions directly from audio signals. We critically examine the effectiveness of seven state-of-the-art deep learning models in predicting groove ratings and responses to groove-related queries through the extraction of audio embeddings. Additionally, we compare these predictions with traditional handcrafted audio features. To better understand the underlying mechanics, we extend this methodology to analyze predictions based on source-separated instruments, thereby isolating the contributions of individual musical elements. Our analysis reveals a clear separation of groove characteristics driven by the underlying musical style of the tracks (funk, pop, and rock). These findings indicate that deep audio representations can successfully encode complex, style-dependent groove components that traditional features often miss. Ultimately, this work highlights the capacity of advanced deep learning models to capture the multifaceted concept of groove, demonstrating the strong potential of representation learning to advance predictive Music Information Retrieval methodologies.


【11】Rhythmic segment analysis: Conceptualizing, visualizing, and measuring rhythmic data
标题:节奏片段分析:概念化,可视化和测量节奏数据
链接:https://arxiv.org/abs/2603.26988

作者:Bas Cornelissen
备注:15 pages, 7 figures
摘要:本文开发了一个框架,概念化,可视化,并测量节奏数据的重复性。我建议用间隔片段来考虑节奏数据:固定长度的连续间隔组,可以分解为持续时间和模式(间隔之间的比率)。这个简单的概念框架统一了三种节奏可视化方法,并产生了第四种:模式持续时间图。当与聚类转换网络配对时,它直观地揭示了合成和真实世界节奏数据中的重复性。此外,该框架概括了两种常见的节奏结构的措施:节奏比和归一化成对变异指数(nPVI)。特别是,nPVI可以重建为平均距离等时性,我提出了一个更一般的措施,anisochrony取代it.Finally,新的概念的数量可能会揭示更广泛的辩论小整数比节奏。
摘要:This paper develops a framework for conceptualizing, visualizing, and measuring regularities in rhythmic data. I propose to think about rhythmic data in terms of interval segments: fixed-length groups of consecutive intervals, which can be decomposed into a duration and a pattern (the ratios between the intervals). This simple conceptual framework unifies three rhythmic visualization methods and yields a fourth: the pattern-duration plot. When paired with a cluster transition network, it intuitively reveals regularities in both synthetic and real-world rhythmic data. Moreover, the framework generalizes two common measures of rhythmic structure: rhythm ratios and the normalized pairwise variability index (nPVI). In particular, nPVI can be reconstructed as the average distance from isochrony, and I propose a more general measure of anisochrony to replace it. Finally, the novel concept of quantality may shed light on wider debates regarding small-integer-ratio rhythms.


【12】Multilingual Stutter Event Detection for English, German, and Mandarin Speech
标题:英语、德语和普通话语音的多语言口吃事件检测
链接:https://arxiv.org/abs/2603.26939

作者:Felix Haas,Sebastian P. Bayerl
摘要:本文提出了一个基于英语、德语和汉语的多语言语料库的多标签口吃检测系统,该模型利用来自三种语言和四种语料库的带注释的口吃数据,捕捉与语言无关的口吃特征,实现跨语言背景的鲁棒检测。实验结果表明,多语言训练达到的性能相当,在某些情况下,甚至超过了以前的系统。这些发现表明,口吃表现出跨语言的一致性,这支持语言不可知检测系统的发展。我们的工作证明了使用多语言数据来提高自动口吃检测的通用性和可靠性的可行性和优势。
摘要:This paper presents a multi-label stuttering detection system trained on multi-corpus, multilingual data in English, German, and Mandarin.By leveraging annotated stuttering data from three languages and four corpora, the model captures language-independent characteristics of stuttering, enabling robust detection across linguistic contexts. Experimental results demonstrate that multilingual training achieves performance comparable to and, in some cases, even exceeds that of previous systems. These findings suggest that stuttering exhibits cross-linguistic consistency, which supports the development of language-agnostic detection systems. Our work demonstrates the feasibility and advantages of using multilingual data to improve generalizability and reliability in automated stuttering detection.


【13】AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
标题:AFSS:以人为本的自合成,用于减轻音频深度伪造检测中的偏差
链接:https://arxiv.org/abs/2603.26856

作者:Hai-Son Nguyen-Le,Hung-Cuong Nguyen-Thanh,Nhien-An Le-Khac,Dinh-Thuc Nguyen,Hong-Hanh Nguyen-Le
备注:Accepted at International Joint Conference on Neural Networks 2026
摘要:生成模型的快速发展使高度逼真的音频deepfake成为可能,但目前的检测器存在严重的偏差问题,导致在看不见的数据集上泛化能力差。本文提出了伪像聚焦自合成(AFSS),一种旨在通过两种机制从真实音频生成伪假样本来减轻这种偏差的方法:自转换和自重建。AFSS的核心思想在于强制执行同一说话人约束,确保真实和伪假样本共享相同的说话人身份和语义内容。这迫使检测器专门关注生成伪影,而不是不相关的混淆因素。此外,我们引入了一个可学习的重新加权损失,在训练过程中动态强调合成样本。在7个数据集上进行的大量实验表明,AFSS实现了最先进的性能,平均EER为5.45%,包括在WaveFake上显著降低到1.23%,在In-the-Wild上降低到2.70%,同时消除了对预先收集的假数据集的依赖。我们的代码可在https://github.com/NguyenLeHaiSonGit/AFSS上公开获取。
摘要:The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused Self-Synthesis (AFSS), a method designed to mitigate this bias by generating pseudo-fake samples from real audio via two mechanisms: self-conversion and self-reconstruction. The core insight of AFSS lies in enforcing same-speaker constraints, ensuring that real and pseudo-fake samples share identical speaker identity and semantic content. This forces the detector to focus exclusively on generation artifacts rather than irrelevant confounding factors. Furthermore, we introduce a learnable reweighting loss to dynamically emphasize synthetic samples during training. Extensive experiments across 7 datasets demonstrate that AFSS achieves state-of-the-art performance with an average EER of 5.45\%, including a significant reduction to 1.23\% on WaveFake and 2.70\% on In-the-Wild, all while eliminating the dependency on pre-collected fake datasets. Our code is publicly available at https://github.com/NguyenLeHaiSonGit/AFSS.


机器翻译由腾讯交互翻译提供,仅供参考