微信公众号:arXiv_Daily
cs.SD语音
【1】Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints
标题:通过基于规则的音乐约束调整歌词到旋律生成的语言模型
链接:https://arxiv.org/abs/2604.18489
备注:Accepted by IEEE ICASSP 2026
摘要:大型语言模型(LLM)在歌词到旋律生成方面表现出了希望,但使用监督微调(SFT)训练的模型通常会产生音乐上令人难以置信的旋律,例如节奏差和不合适的音域,这种现象我们称之为“约束违反”。为了解决这个问题,我们提出了一个新的对齐框架,灌输音乐知识,而无需人类注释。我们定义了基于规则的音乐约束,以自动生成一个偏好数据集从SFT模型的输出。然后通过顺序过程对模型进行对齐,首先对配对偏好数据使用直接偏好优化(DPO),然后对未配对的阴性样本使用Kahneman-Tversky优化(KTO)。实验结果表明,我们的对齐模型大大减少了违反规则的行为,并在客观和主观评估中优于强基线,生成具有显著改善的音乐性和连贯性的旋律。一个互动演示与音频比较可在https://arain233.github.io/AligningMelody-demo。
摘要:Large Language Models (LLMs) show promise in lyric-to-melody generation, but models trained with Supervised Fine-Tuning (SFT) often produce musically implausible melodies with issues like poor rhythm and unsuitable vocal ranges, a phenomenon we term "constraint violation". To address this, we propose a novel alignment framework that instills musical knowledge without human annotation. We define rule-based musical constraints to automatically generate a preference dataset from an SFT model's outputs. The model is then aligned through a sequential process, first using Direct Preference Optimization (DPO) on paired preference data, followed by Kahneman-Tversky Optimization (KTO) on unpaired negative samples. Experimental results demonstrate that our aligned model substantially reduces rule violations and outperforms strong baselines in both objective and subjective evaluations, generating melodies with substantially improved musicality and coherence. An interactive demo with audio comparisons is available at https://arain233.github.io/AligningMelody-demo.
【2】Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
标题:全嵌入音频:利用多模式LLM实现稳健的音频文本检索
链接:https://arxiv.org/abs/2604.18360
备注:Accepted at ACL 2026 Main Conference. Camera-ready version
摘要:基于对比存储音频预训练(CLAP)的音频文本检索系统在传统的基准测试中取得了很好的性能;然而,这些基准测试依赖于与真实世界搜索行为大不相同的标题风格查询,限制了它们对实际检索鲁棒性的评估。我们提出了Omni-Embed-Audio(OEA),这是一个面向检索的编码器,它利用具有本地音频理解的多模态LLM。为了系统地评估标题式查询之外的鲁棒性,我们引入了用户意图查询(UIQs)-反映自然搜索行为的五种公式:问题,命令,关键字标签,释义和基于排除的否定查询。对于否定查询,我们开发了一个硬否定挖掘管道,并提出了判别指标(HNSR,TFR)评估模型抑制声学相似干扰项的能力。在AudioCaps、Clotho和MECAT上的实验表明,OEA实现了与最先进的M2 D-CLAP相当的文本到音频检索性能,同时在两个关键领域表现出明显的优势:(1)显性文本到文本检索(+22%相对改善),和(2)实质上优于硬负面歧视(+4.3%p HNSR@10,+34.7%相对TFR@10),揭示了LLM主干提供了对复杂查询的高级语义理解。
摘要:Audio-text retrieval systems based on Contrastive Language-Audio Pretraining (CLAP) achieve strong performance on traditional benchmarks; however, these benchmarks rely on caption-style queries that differ substantially from real-world search behavior, limiting their assessment of practical retrieval robustness. We present Omni-Embed-Audio (OEA), a retrieval-oriented encoder leveraging multimodal LLMs with native audio understanding. To systematically evaluate robustness beyond caption-style queries, we introduce User-Intent Queries (UIQs) - five formulations reflecting natural search behaviors: questions, commands, keyword tags, paraphrases, and exclusion-based negative queries. For negative queries, we develop a hard negative mining pipeline and propose discrimination metrics (HNSR, TFR) assessing models' ability to suppress acoustically similar distractors. Experiments on AudioCaps, Clotho, and MECAT show that OEA achieves comparable text-to-audio retrieval performance to state-of-the-art M2D-CLAP, while demonstrating clear advantages in two critical areas: (1) dominant text-to-text retrieval (+22% relative improvement), and (2) substantially superior hard negative discrimination (+4.3%p HNSR@10, +34.7% relative TFR@10), revealing that LLM backbones provide superior semantic understanding of complex queries.
【3】Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models
标题:Audio-DeepThinker:渐进推理感知强化学习,以实现音频语言模型中的高质量思想链涌现
链接:https://arxiv.org/abs/2604.18187
摘要:大型音频语言模型(LALM)在音频理解方面取得了重大进展,但它们主要作为感知和回答系统运行,没有显式推理过程。用于增强音频推理的现有方法依赖于受训练数据质量限制的监督式思维链(CoT)微调,或者依赖于具有不直接评估推理质量的粗略奖励的强化学习(RL)。因此,生成的推理链通常看起来结构良好,但缺乏特定的声学基础。我们提出了Audio-DeepThinker,这是一个基于两个核心思想的框架。首先,我们引入了一个混合推理相似性奖励,直接监督生成的推理链的质量相结合的LLM评估器评估逻辑路径对齐,关键步骤的覆盖范围和分析深度与嵌入相似性组件执行语义对齐与参考推理链。其次,我们提出了一个渐进的两阶段课程,使高质量的CoT推理出现通过纯RL探索,没有任何监督推理微调,从一个不具备事先的思想链能力的模型。第1阶段使用混合奖励对基础音频QA进行训练,以培养基本的推理模式,而第2阶段则转向具有声学挑战性的边界案例,仅使用LLM奖励,以获得更大的推理多样性。Audio-DeepThinker在MMAR(74.0%)、MMAU-test-mini(78.5%)和MMSU(77.26%)上取得了最先进的成绩,在Interspeech 2026音频推理挑战赛(单一模型赛道)中获得第一名。可解释性分析进一步揭示,RL训练主要重塑上层MoE门控机制,推理标记在上层Transformer层中逐渐结晶,为音频推理如何通过探索出现提供了机械的见解。
摘要:Large Audio-Language Models (LALMs) have made significant progress in audio understanding, yet they primarily operate as perception-and-answer systems without explicit reasoning processes. Existing methods for enhancing audio reasoning rely either on supervised chain-of-thought (CoT) fine-tuning, which is limited by training data quality, or on reinforcement learning (RL) with coarse rewards that do not directly evaluate reasoning quality. As a result, the generated reasoning chains often appear well-structured yet lack specific acoustic grounding. We propose Audio-DeepThinker, a framework built on two core ideas. First, we introduce a hybrid reasoning similarity reward that directly supervises the quality of generated reasoning chains by combining an LLM evaluator assessing logical path alignment, key step coverage, and analytical depth with an embedding similarity component enforcing semantic alignment with reference reasoning chains. Second, we propose a progressive two-stage curriculum that enables high-quality CoT reasoning to emerge through pure RL exploration, without any supervised reasoning fine-tuning, from an instruction-tuned model that possesses no prior chain-of-thought capability. Stage 1 trains on foundational audio QA with the hybrid reward to foster basic reasoning patterns, while Stage 2 shifts to acoustically challenging boundary cases with an LLM-only reward for greater reasoning diversity. Audio-DeepThinker achieves state-of-the-art results on MMAR (74.0%), MMAU-test-mini (78.5%), and MMSU (77.26%), winning 1st Place in the Interspeech 2026 Audio Reasoning Challenge (Single Model Track). Interpretability analyses further reveal that RL training primarily reshapes upper-layer MoE gating mechanisms and that reasoning tokens crystallize progressively in the upper transformer layers, offering mechanistic insights into how audio reasoning emerges through exploration.
【4】FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings
标题:FLiP:理解和解释多模式多语言句子嵌入
链接:https://arxiv.org/abs/2604.18109
备注:Under review
摘要:本文提出了用于理解预训练句子嵌入空间的因子分解线性投影(FLiP)模型。我们训练FLiP模型,以恢复多语言(LaBSE),多模态(SONAR)和基于API(双子座)的句子嵌入空间中的词汇内容在几个高,中资源语言。我们发现,FLiP可以从嵌入中回忆起超过75%的词汇内容,显着优于现有的非因子分解基线。使用这个作为诊断工具,我们发现了所选句子编码器的模态和语言偏见,并为从业者提供了关于编码器的内在见解,而不依赖于传统的下游评估任务。我们的实现是公开的https://github.com/BUTSpeechFIT/FLiP。
摘要:This paper presents factorized linear projection (FLiP) models for understanding pretrained sentence embedding spaces. We train FLiP models to recover the lexical content from multilingual (LaBSE), multimodal (SONAR) and API-based (Gemini) sentence embedding spaces in several high- and mid-resource languages. We show that FLiP can recall more than 75% of lexical content from the embeddings, significantly outperforming existing non-factorized baselines. Using this as a diagnostic tool, we uncover the modality and language biases across the selected sentence encoders and provide practitioners with intrinsic insights about the encoders without relying on conventional downstream evaluation tasks. Our implementation is public https://github.com/BUTSpeechFIT/FLiP.
【5】Latent Fourier Transform
标题:潜傅里叶变换
链接:https://arxiv.org/abs/2604.17986
备注:ICLR 2026 Oral
摘要:我们介绍了潜在的傅立叶变换(LatentFT),一个框架,提供了新的频域控制生成音乐模型。LatentFT将扩散自动编码器与潜在空间傅立叶变换相结合,以按时间尺度分离音乐模式。通过在训练过程中掩蔽频域中的潜伏期,我们的方法产生了可以在推理时进行相干操作的表示。这使我们能够从参考示例中生成音乐变奏和混合,同时保留所需时间尺度的特征,这些特征被指定为潜在空间中的频率。LatentFT与均衡器在音乐制作中的作用相似:传统均衡器在可听频率上操作以塑造音色,而LatentFT在潜在空间频率上操作以塑造音乐结构。实验和听力测试表明,与基线相比,LatentFT提高了条件依从性和质量。我们还提出了一种技术,在隔离的潜在空间中的听觉频率,并显示不同的音乐属性驻留在不同区域的潜在频谱。我们的研究结果表明,潜在空间中的频域控制如何为调节和混合提供直观,连续的频率轴,使我们朝着更具可解释性和交互性的生成音乐模型前进。
摘要:We introduce the Latent Fourier Transform (LatentFT), a framework that provides novel frequency-domain controls for generative music models. LatentFT combines a diffusion autoencoder with a latent-space Fourier transform to separate musical patterns by timescale. By masking latents in the frequency domain during training, our method yields representations that can be manipulated coherently at inference. This allows us to generate musical variations and blends from reference examples while preserving characteristics at desired timescales, which are specified as frequencies in the latent space. LatentFT parallels the role of the equalizer in music production: while traditional equalizers operates on audible frequencies to shape timbre, LatentFT operates on latent-space frequencies to shape musical structure. Experiments and listening tests show that LatentFT improves condition adherence and quality compared to baselines. We also present a technique for hearing frequencies in the latent space in isolation, and show different musical attributes reside in different regions of the latent spectrum. Our results show how frequency-domain control in latent space provides an intuitive, continuous frequency axis for conditioning and blending, advancing us toward more interpretable and interactive generative music models.
【6】LLM-Codec: Neural Audio Codec Meets Language Model Objectives
标题:LLM-Codec:神经音频编解码器满足语言模型目标
链接:https://arxiv.org/abs/2604.17852
备注:ACL2026 Finding
摘要:神经音频编解码器被广泛用作口语模型的标记器,但它们被优化用于波形重建而不是自回归预测。这种不匹配将声学驱动的不确定性注入到离散标记空间中,并增加了语言模型的复杂性。我们提出了我们的,它增强了编解码器训练与语言模型面向目标,同时保持编解码器和LLM架构不变。\我们介绍了(i)未来的令牌预测与美杜莎风格的多步头,以鼓励多步的可预测性,以及(ii)语义对齐,通过记忆库对比损失匹配音频和文本表示。可微分Gumbel桥实现从这些目标到编解码器编码器的端到端梯度。在SALMon语音连贯性方面,在我们的平台上训练的令牌LM达到了61.6%的准确率(比AUV高出12.1个百分点),同时减少了困惑35。在Codec-SUPERB-tiny上,我们的语音Mel距离比AUV提高了5.0%,同时实现了可学习性增益,表明重建保真度和令牌可预测性可以一起提高。
摘要:Neural audio codecs are widely used as tokenizers for spoken language models, but they are optimized for waveform reconstruction rather than autoregressive prediction. This mismatch injects acoustically driven uncertainty into the discrete token space and increases language-model perplexity. We propose \ours, which augments codec training with language-model-facing objectives while keeping both codec and LLM architectures unchanged. \ours introduces (i) future token prediction with Medusa-style multi-step heads to encourage multi-step predictability, and (ii) semantic alignment that matches audio and text representations via a memory-bank contrastive loss. A differentiable Gumbel bridge enables end-to-end gradients from these objectives to the codec encoder. On SALMon speech coherence, token LMs trained on \ours reach 61.6% accuracy (+12.1 points over AUV) while reducing perplexity 35. On Codec-SUPERB-tiny, \ours improves speech Mel distance by 5.0% over AUV while simultaneously achieving the learnability gains, demonstrating that reconstruction fidelity and token predictability can be improved together.
【7】A novel LSTM music generator based on the fractional time-frequency feature extraction
标题:基于分数时频特征提取的新型LSTM音乐生成器
链接:https://arxiv.org/abs/2604.17823
备注:This work was supported by Hainan Provincial Natural Science Foundation of China (Grant No. 723QN238)
摘要:在本文中,我们提出了一种新的方法来生成音乐的人工智能(AI)系统的基础上。我们分析音乐的特征,并使用它们来拟合和预测音乐。分数阶傅里叶变换(FrFT)和长短期记忆(LSTM)网络是我们方法的基础。FrFT方法用于提取音乐作品的谱特征,其中音乐信号在时域和频域上表示。LSTM网络用于基于提取的特征生成新的音乐,其中我们使用GiantMIDI-Piano数据集根据隐藏层特征和实时输入来预测音乐。我们的实验结果表明,我们提出的系统能够生成与人类生成的音乐相当的高质量音乐。
摘要:In this paper, we propose a novel approach for generating music based on an artificial intelligence (AI) system. We analyze the features of music and use them to fit and predict the music. The fractional Fourier transform (FrFT) and the long short-term memory (LSTM) network are the foundations of our method. The FrFT method is used to extract the spectral features of a music piece, where the music signal is expressed on the time and frequency domains. The LSTM network is used to generate new music based on the extracted features, where we predict the music according to the hidden layer features and real-time inputs using GiantMIDI-Piano dataset. The results of our experiments show that our proposed system is capable of generating high-quality music that is comparable to human-generated music.
【8】Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
标题:Video-Robin:意向导向的视频到音乐一代的自回归传播规划
链接:https://arxiv.org/abs/2604.17656
摘要:视频到音乐(V2 M)是为输入视频创建背景音乐的基本任务。最近的V2 M模型通常仅依赖于视觉调节来实现视听对齐,并向最终用户提供有限的语义和风格可控性。在本文中,我们提出了Video-Robin,一种新的文本条件下的视频到音乐生成模型,使快速,高质量,语义对齐的视频内容的音乐生成。为了平衡音乐保真度和语义理解,Video-Robin集成了自回归规划和基于扩散的合成。具体而言,自回归模块通过语义上对齐视觉和文本输入来对全局结构进行建模,以产生高级别的音乐潜伏期。这些潜在的随后被精炼成连贯的,高保真的音乐使用本地扩散Transformers。通过将语义驱动的规划分解到基于扩散的合成中,Video-Robin可以在不牺牲音频真实感的情况下实现细粒度的创作者控制。我们提出的模型优于仅接受视频输入的基线和分布内和分布外基准上的附加功能条件基线,与SOTA相比,推理速度为2.21倍。我们将在文件接受后开放所有内容。
摘要:Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditioning alone and provide limited semantic and stylistic controllability to the end user. In this paper, we present Video-Robin, a novel text-conditioned video-to-music generation model that enables fast, high-quality, semantically aligned music generation for video content. To balance musical fidelity and semantic understanding, Video-Robin integrates autoregressive planning with diffusion-based synthesis. Specifically, an autoregressive module models global structure by semantically aligning visual and textual inputs to produce high-level music latents. These latents are subsequently refined into coherent, high-fidelity music using local Diffusion Transformers. By factoring semantically driven planning into diffusion-based synthesis, Video-Robin enables fine-grained creator control without sacrificing audio realism. Our proposed model outperforms baselines that solely accept video input and additional feature conditioned baselines on both in-distribution and out-of-distribution benchmarks with a 2.21x speed in inference compared to SOTA. We will open-source everything upon paper acceptance.
【9】MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
标题:动作:通过言语到言语翻译中的发声专家混合翻译笑声和泪水
链接:https://arxiv.org/abs/2604.17435
备注:Submitted to Interspeech. Audio Demo and Dataset: https://47zzz.github.io/MoVE/
摘要:最近的语音到语音翻译(S2ST)系统实现了很强的语义准确性,但始终剥离非言语发声(NV),如表达语用意图的笑声和哭泣,这严重限制了现实世界的效用。我们通过三个贡献来解决这个问题。首先,我们提出了一个合成管道来构建可扩展的表达数据集,以克服数据稀缺性的限制。其次,我们提出了MoVE,一个混合的LoRA专家架构与表达专用适配器和软加权路由器,混合专家捕捉混合表达状态。第三,我们展示了预训练的AudioLLM能够实现惊人的数据效率:30分钟的数据就足以实现强大的性能。在英汉S2ST上,与强基线相比,MoVE在76%的情况下再现了目标NVs,并在所有比较系统中实现了最高的人类评级自然度和情感保真度,而现有的S2ST系统最多保留了14%的NVs。
摘要:Recent Speech-to-Speech Translation (S2ST) systems achieve strong semantic accuracy yet consistently strip away non-verbal vocalizations (NVs), such as laughter and crying that convey pragmatic intent, which severely limits real-world utility. We address this via three contributions. First, we propose a synthesis pipeline for building scalable expressive datasets to overcome the data scarcity limitation. Second, we propose MoVE, a Mixture-of-LoRA-Experts architecture with expressive-specialized adapters and a soft-weighting router that blends experts for capturing hybrid expressive states. Third, we show pretrained AudioLLMs enable striking data efficiency: 30 minutes of curated data is enough for strong performance. On English-Chinese S2ST, while comparing with strong baselines, MoVE reproduces target NVs in 76% of cases and achieves the highest human-rated naturalness and emotional fidelity among all compared systems, where existing S2ST systems preserve at most 14% of NVs.
【10】Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
标题:还在我们之间吗?评估和改进语音助手对第三方中断的鲁棒性
链接:https://arxiv.org/abs/2604.17358
备注:ACL 2026 main conference
摘要:虽然最近的口语模型(SLM)已被积极部署在现实世界的情况下,他们缺乏辨别第三方中断(TPI)从主用户的正在进行的流的能力,使他们容易受到上下文故障。为了弥合这一差距,我们引入了TPI-Train,这是一个88 K实例的数据集,设计了说话者感知的硬否定,以执行中断处理的声学线索优先级,以及TPI-Bench,这是一个综合评估框架,旨在严格测量中断处理策略和欺骗性上下文中的精确说话者区分。实验表明,我们的数据集设计减轻了语义捷径学习的关键陷阱,模型利用语义上下文,而忽略了声学信号识别扬声器的变化至关重要。我们相信,我们的工作建立了一个基础资源,克服文本主导的单峰依赖SLM,铺平了道路,更强大的多方口语互动。该框架的代码可在https://tpi-va.github.io上公开获取
摘要:While recent Spoken Language Models (SLMs) have been actively deployed in real-world scenarios, they lack the capability to discern Third-Party Interruptions (TPI) from the primary user's ongoing flow, leaving them vulnerable to contextual failures. To bridge this gap, we introduce TPI-Train, a dataset of 88K instances designed with speaker-aware hard negatives to enforce acoustic cue prioritization for interruption handling, and TPI-Bench, a comprehensive evaluation framework designed to rigorously measure the interruption-handling strategy and precise speaker discrimination in deceptive contexts. Experiments demonstrate that our dataset design mitigates semantic shortcut learning-a critical pitfall where models exploit semantic context while neglecting acoustic signals essential for discerning speaker changes. We believe our work establishes a foundational resource for overcoming text-dominated unimodal reliance in SLMs, paving the way for more robust multi-party spoken interaction. The code for the framework is publicly available at https://tpi-va.github.io
【11】TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation
标题:TeMuDance:基于对比对齐的文本控制,用于音乐驱动的舞蹈生成
链接:https://arxiv.org/abs/2604.17005
摘要:现有的音乐驱动的舞蹈生成方法已经实现了很强的真实感和有效的音频-运动对齐。然而,它们通常缺乏语义可控性,因此很难通过自然语言描述来指导特定的动作。这种限制主要源于缺乏大规模的数据集,这些数据集可以联合对齐音乐,文本和运动,以进行文本条件控制的监督学习。为了解决这一挑战,我们提出了TeMuDance,一个框架,使基于文本的控制音乐条件下的舞蹈生成,而不需要任何手动注释的音乐文本运动三元组数据集。TeMuDance引入了一种以运动为中心的桥接范式,该范式利用运动作为共享的语义锚点,在统一的嵌入空间内对齐不相交的音乐舞蹈和文本运动数据集,从而能够跨模式检索缺失的模式,以进行端到端训练。然后,在冻结的音乐到舞蹈扩散主干之上训练轻量级文本控制分支,在保持节奏保真度的同时实现细粒度的语义指导。为了进一步抑制噪声固有的检索监督,我们设计了一个双流微调策略与基于信心的过滤。我们还提出了一种新的任务对齐的度量,量化文本提示是否诱导音乐条件下的预期运动属性。大量的实验表明,TeMuDance实现了有竞争力的舞蹈质量,同时大大提高了现有方法的文本条件控制。
摘要:Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semantic controllability, making it difficult to guide specific movements through natural language descriptions. This limitation primarily stems from the absence of large-scale datasets that jointly align music, text, and motion for supervised learning of text-conditioned control. To address this challenge, we propose TeMuDance, a framework that enables text-based control for music-conditioned dance generation without requiring any manually annotated music-text-motion triplet dataset. TeMuDance introduces a motion-centred bridging paradigm that leverages motion as a shared semantic anchor to align disjoint music-dance and text-motion datasets within a unified embedding space, enabling cross-modal retrieval of missing modalities for end-to-end training. A lightweight text control branch is then trained on top of a frozen music-to-dance diffusion backbone, preserving rhythmic fidelity while enabling fine-grained semantic guidance. To further suppress noise inherent in the retrieved supervision, we design a dual-stream fine-tuning strategy with confidence-based filtering. We also propose a novel task-aligned metric that quantifies whether textual prompts induce the intended kinematic attributes under music conditioning. Extensive experiments demonstrate that TeMuDance achieves competitive dance quality while substantially improving text-conditioned control over existing methods.
【12】ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection
标题:ICLAT:音频深度伪造检测的上下文学习和比较指导
链接:https://arxiv.org/abs/2604.16749
备注:To appear at ACL Findings 2026
摘要:音频deepfake构成了重大的安全威胁,但目前最先进的(SOTA)检测系统并不能很好地推广到现实的deepfake。我们引入了一种新的\textbf{I}n-\textbf{C}上下文\textbf{L}学习范式,用于\textbf{A}udio \textbf{D}伪检测(\textbf{ICLAD})。该框架允许使用音频语言模型(ALM)对看不见的deepfake进行无训练泛化,并提供检测结果的文本依据。ICLAD的核心是一种成对比较推理策略,该策略指导ALM发现和过滤幻觉和与deepfake无关的声学属性。ALM与专门的deepfake检测器一起工作,其中路由机制将分发样本馈送到ALM。在野外数据集上,ICLAD相对于专用检测器改进了宏F1,相对改进高达$2\times$。进一步的分析表明,国际土地退化问题中心具有灵活性,并有潜力在最近的开放源码土地退化模型上部署。
摘要:Audio deepfakes pose a significant security threat, yet current state-of-the-art (SOTA) detection systems do not generalize well to realistic in-the-wild deepfakes. We introduce a novel \textbf{I}n-\textbf{C}ontext \textbf{L}earning paradigm with comparison-guidance for \textbf{A}udio \textbf{D}eepfake detection (\textbf{ICLAD}). The framework enables the use of audio language models (ALMs) for training-free generalization to unseen deepfakes and provides textual rationales on the detection outcome. At the core of ICLAD is a pairwise comparative reasoning strategy that guides the ALM to discover and filter hallucinations and deepfake-irrelevant acoustic attributes. The ALM works alongside a specialized deepfake detector, whereby a routing mechanism feeds out-of-distribution samples to the ALM. On in-the-wild datasets, ICLAD improves macro F1 over the specialized detector, with up to $2\times$ relative improvement. Further analysis demonstrates the flexibility of ICLAD and its potential for deployment on recent open-source ALMs.
【13】Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
标题:良性微调打破了音频LLM的安全一致性
链接:https://arxiv.org/abs/2604.16659
摘要:先前的工作表明,在良性数据上微调对齐模型会降低文本和视觉模式的安全性,并且在表示空间中接近有害内容会预测哪些样本会造成最大的损害。然而,现有的分析是在一个单一的、无差别的嵌入空间内进行的,这就使得不同的输入属性是否会以不同的方式驱动漏洞成为一个未知数。音频引入了一个结构上更丰富的问题:一个良性的样本不仅可以通过所说的话,而且可以通过它的声音来接近有害的内容,即使它的话是完全无害的。我们首次系统地研究了Audio LLM中的良性微调安全性,评估了三种最先进的模型,该模型采用基于邻近度的过滤框架,通过嵌入空间距离到有害内容来选择良性音频。通过使用外部参考编码器以及每个模型自己的内部编码器将接近度分解为语义,声学和混合轴,我们表明,良性微调将越狱成功率(JSR)从个位数提高到高达87.12%。至关重要的是,主要的脆弱性轴和音频与文本微调的相对风险都是由架构决定的-由每个模型的编码器和投影仪如何将音频转换到LLM的输入空间来决定。我们提出了两种防御措施:过滤训练数据以最大化与有害嵌入的距离,以及文本系统提示推理,两者都在无需架构修改的情况下将JSR降低到接近于零。我们对两种架构的机制分析表明,微调选择性地抑制了后层拒绝电路,而冻结的编码器保留表示,甚至抑制模式是架构条件的,反映了跨模态的行为不对称。良性微调导致的安全性下降是音频LLM中的一个定性的独特风险。
摘要:Prior work shows that fine-tuning aligned models on benign data degrades safety in text and vision modalities, and that proximity to harmful content in representation space predicts which samples cause the most damage. However, existing analyses operate within a single, undifferentiated embedding space -- leaving open whether distinct input properties drive the vulnerability differently. Audio introduces a structurally richer problem: a benign sample can neighbor harmful content not only through what is said but through how it sounds, even when its words are entirely innocuous. We present the first systematic study of benign fine-tuning safety in Audio LLMs, evaluating three state-of-the-art models with a proximity-based filtering framework that selects benign audio by embedding-space distance to harmful content. By decomposing proximity into semantic, acoustic, and mixed axes using external reference encoders alongside each model's own internal encoder, we show that benign fine-tuning elevates Jailbreak Success Rate (JSR) from single digits to as high as 87.12%. Crucially, the dominant vulnerability axis and the relative risk of audio versus text fine-tuning are both architecture-conditioned -- determined by how each model's encoder and projector transform audio into the LLM's input space. We propose two defenses: filtering training data to maximize distance from harmful embeddings, and a textual system prompt at inference, both reducing JSR to near-zero without architectural modification. Our mechanistic analysis on two architectures reveals that fine-tuning selectively suppresses the late-layer refusal circuit while the frozen encoder preserves representations, and that even the suppression pattern is architecture-conditioned, mirroring the behavioral asymmetries across modalities. Safety degradation from benign fine-tuning is a qualitatively distinct risk in Audio LLMs.
【14】Coexisting Tempo Traditions in Beethoven's Piano and Cello Sonatas: A K-means Clustering Analysis of Recorded Performances, 1930-2012
标题:贝多芬钢琴和大提琴奏鸣曲中共存的节奏描述:1930-2012年录制表演的K均值集群分析
链接:https://arxiv.org/abs/2604.16658
摘要:传统上,对录音演奏的实证研究将节奏变化建模为一个单向的历史过程,将线性回归线拟合到针对录音年份绘制的节奏数据。本文认为,这种方法强加了一个统一的文体演变的虚假叙述,事实上,是一个共存的解释传统的多元化。将k-均值聚类(k=3)应用于来自贝多芬五首钢琴和大提琴奏鸣曲(Op.5 Nos. 1和2; Op. 69; Op. 102 Nos.这项研究显示,每一个乐章都支持至少两个,通常是三个离散的节奏传统(慢、中、快),其内部回归斜率可以忽略不计(除了一个案例外,所有案例的R平方都<= 0.25),表明每一个传统在80年内都是独立稳定的。中程集群在所有运动中占主导地位,通常占记录的55-70%。快速的角色动作(作品5回旋曲,作品69谐谑曲)中没有缓慢的集群,反映了对他们角色的共同修辞共识。显著的簇内漂移的单个情况(Op. 102 No. 1 Allegro con brio,R-squared=0.246,p=0.013)表明在整个研究期间约为3.2 BPM的中等范围减速。集群成员和表演者的世代,国家或教育背景之间没有相关性,这表明节奏传统反映了个人的解释选择,而不是集体的文化传承。本文提出了一种风格变化的生态模型-共存的传统相对流行的转变,而不是一个单一的传统不断发展-并认为,这种重构具有广泛的影响,实证性能研究如何解释语料库水平的节奏数据。
摘要:Empirical studies of recorded performance have conventionally modelled tempo change as a unidirectional historical process, fitting linear regression lines to tempo data plotted against recording year. This paper argues that such approaches impose a false narrative of uniform stylistic evolution on what is, in fact, a plurality of coexisting interpretive traditions. Applying k-means clustering (k=3) to bar-level BPM data from over one hundred recordings of Beethoven's five piano and cello sonatas (Op. 5 Nos. 1 and 2; Op. 69; Op. 102 Nos. 1 and 2) spanning 1930-2012, this study reveals that every movement supports at least two, and usually three, discrete tempo traditions (slow, mid-range, and fast), whose internal regression slopes are negligible (R-squared <= 0.25 in all but one case), demonstrating that each tradition is independently stable across eight decades. The mid-range cluster dominates in all movements, typically comprising 55-70% of recordings. A slow cluster is absent from fast-character movements (Op. 5 Rondos, Op. 69 Scherzo), reflecting a shared rhetorical consensus about their character. The single case of significant intra-cluster drift (Op. 102 No. 1 Allegro con brio, R-squared=0.246, p=0.013) indicates a moderate mid-range deceleration of approximately 3.2 BPM across the study period. No correlation is found between cluster membership and performers' generational, national, or pedagogical backgrounds, suggesting that tempo tradition reflects individual interpretive choice rather than collective cultural inheritance. The paper proposes an ecological model of stylistic change - coexisting traditions shifting in relative prevalence rather than a single tradition evolving - and argues that this reframing has broad implications for how empirical performance studies interpret corpus-level tempo data.
【15】AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
标题:AVRT:通过单一情态教师的视听推理转移
链接:https://arxiv.org/abs/2604.16617
摘要:推理模型的最新进展已经在基于文本的领域中显示出显著的进步,但是将这些能力转移到多模态设置,例如,允许对视听数据进行推理仍然是一个挑战,部分原因是在目标多模态组合中高质量推理数据的可用性有限。为了解决这个问题,我们引入AVRT,一个新的框架,从单模态教师模型生成高质量的视听推理痕迹。我们生成独立的视觉和音频推理痕迹通过专门的模型,以理由在各自的模态和合并所产生的痕迹与LLM合并模型。由此产生的多模态轨迹用于监督微调(SFT)冷启动,以首先使目标模型适应视听推理轨迹,然后在第二个强化学习阶段对更大规模的数据进行训练。在七个视听和音频基准上进行评估,我们的3B和7B参数模型在大小相当的模型中实现了最先进的结果,包括用于视听的OmniBench和DailyOmni以及用于仅音频推理的MMAR,这表明跨模态训练也转移到单模态任务,并为多模态推理模型建立了新的训练管道。
摘要:Recent advances in reasoning models have shown remarkable progress in text-based domains, but transferring those capabilities to multimodal settings, e.g., to allow reasoning over audio-visual data, still remains a challenge, in part because of the limited availability of high-quality reasoning data in targeted multimodal combinations. To address this problem, we introduce AVRT, a novel framework that generates high-quality audio-visual reasoning traces from single-modality teacher models. We generate independent vision- and audio-reasoning traces via models specialized to reason over their respective modalities and merge the resulting traces with an LLM merger model. The resulting multimodal traces are used in a supervised fine-tuning (SFT) cold start to adapt the target model to audio-visual reasoning traces first, before training it in a second reinforcement learning stage on larger-scale data. Evaluated on seven audio-visual and audio benchmarks, our 3B and 7B parameter models achieve state-of-the-art results among models of comparable size including OmniBench and DailyOmni for audio-visual and MMAR for audio-only reasoning, showing that cross-modal training also transfers to single-modality tasks and establishing a new training pipeline for multimodal reasoning models.
【16】EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions
标题:EchoChain:一个用于中断状态更新推理的全双工基准测试
链接:https://arxiv.org/abs/2604.16456
摘要:实时语音助手必须在用户中断响应时修改任务状态,但现有的口语对话基准在很大程度上评估了基于回合的交互,并错过了这种故障模式。我们介绍EchoChain,这是一个用于评估语音中断下全双工状态更新推理的受控基准。EchoChain确定了中断后延续中的三种反复出现的失败模式:上下文惯性,中断健忘症和目标位移。该基准测试生成了一个由语音驱动的对话,并在一个相对于辅助语音起始的标准化点上注入中断,从而实现了受控的跨模型比较。在成对的半双工控制,总故障下降了40.2%,相对于中断运行,这表明许多错误是由状态更新推理中断,而不是任务难度单独驱动。在评估的实时语音模型中,没有一个系统的通过率超过50%,这表明在中期状态修改方面有很大的改进空间。EchoChain为诊断全双工语音交互中的状态更新推理故障提供了一个可重复的基准。
摘要:Real-time voice assistants must revise task state when users interrupt mid-response, but existing spoken-dialog benchmarks largely evaluate turn-based interaction and miss this failure mode. We introduce EchoChain, a controlled benchmark for evaluating full-duplex state-update reasoning under mid-speech interruptions. EchoChain identifies three recurring failure patterns in post-interruption continuations: contextual inertia, interruption amnesia, and objective displacement. The benchmark generates scenario-driven conversations and injects interruptions at a standardized point relative to assistant speech onset, enabling controlled cross-model comparison. In a paired half-duplex control, total failures drop by 40.2% relative to interrupted runs, indicating that many errors are driven by state-update reasoning under interruption rather than task difficulty alone. Across evaluated real-time voice models, no system exceeds a 50% pass rate, showing substantial room for improvement in mid-generation state revision. EchoChain provides a reproducible benchmark for diagnosing state-update reasoning failures in full-duplex voice interaction.
【17】A High-Accuracy Optical Music Recognition Method Based on Bottleneck Residual Convolutions
标题:基于瓶颈剩余卷积的高准确度光学音乐识别方法
链接:https://arxiv.org/abs/2604.16446
备注:2 figs, and 13 tables
摘要:光学音乐识别(OMR)旨在将打印或手写的乐谱图像转换为可编辑的符号表示。本文提出了一个端到端的OMR框架,它结合了残留瓶颈卷积和基于双向门控递归单元(BiGRU)的序列建模。使用具有ResNet-v2风格的残留瓶颈块和多尺度扩张卷积的卷积神经网络来提取对细粒度符号细节和全局staff-line结构进行编码的特征。然后将提取的特征序列馈送到BiGRU网络中,以对音乐符号之间的时间依赖性进行建模。该模型使用连接主义时间分类损失进行训练,从而实现端到端预测,而无需显式对齐注释。在Camera-PrIMuS和PrIMuS数据集上的实验结果证明了该框架的有效性。在Camera-PrIMuS上,该方法的序列错误率(SeER)为7.52%,符号错误率(SyER)为0.45%,音高、类型和音符的准确率分别为99.33%、99.60%和99.28%.平均训练时间为1.74s/epoch,在保持较强识别性能的同时,显示出较高的计算效率。在PrIMuS上,该方法的SeER为8.11美元,SyER为0.49美元,音高、类型和音符准确度分别为99.27美元、99.58美元和99.21美元。细粒度的误差分析进一步证实了所提出的模型的有效性。
摘要:Optical Music Recognition (OMR) aims to convert printed or handwritten music score images into editable symbolic representations. This paper presents an end-to-end OMR framework that combines residual bottleneck convolutions with bidirectional gated recurrent unit (BiGRU)-based sequence modeling. A convolutional neural network with ResNet-v2-style residual bottleneck blocks and multi-scale dilated convolutions is used to extract features that encode both fine-grained symbol details and global staff-line structures. The extracted feature sequences are then fed into a BiGRU network to model temporal dependencies among musical symbols. The model is trained using the Connectionist Temporal Classification loss, enabling end-to-end prediction without explicit alignment annotations. Experimental results on the Camera-PrIMuS and PrIMuS datasets demonstrate the effectiveness of the proposed framework. On Camera-PrIMuS, the proposed method achieves a sequence error rate (SeER) of $7.52\%$ and a symbol error rate (SyER) of $0.45\%$, with pitch, type, and note accuracies of $99.33\%$, $99.60\%$, and $99.28\%$, respectively. The average training time is 1.74~s per epoch, demonstrating high computational efficiency while maintaining strong recognition performance. On PrIMuS, the method achieves a SeER of $8.11\%$ and a SyER of $0.49\%$, with pitch, type, and note accuracies of $99.27\%$, $99.58\%$, and $99.21\%$, respectively. A fine-grained error analysis further confirms the effectiveness of the proposed model.
【18】iPhoneme: Brain-to-Text Communication for ALS Using ConformerXL Decoding
标题:iPhoneme:使用ConformerXL解码的ALS脑到文本通信
链接:https://arxiv.org/abs/2604.16441
摘要:脑-机接口(BCI)的语音恢复具有变革的潜力,为全球约173000 - 232500人与ALS相关的构音障碍。尽管最近取得了进展,但全球仅在22- 31例患者中证明了高性能语音BCI,这主要是由于神经解码准确性和实际输入接口的限制。我们提出iPhoneme,一个大脑到文本的通信系统,通过集成建模和交互设计,共同应对这些挑战。该系统结合了基于修改后的Conformer架构(ConformerXL,192.9M参数)的深度学习音素解码器和凝视辅助音素输入界面,可缓解眼动跟踪系统中的Midas触摸问题。声学模型包含一个具有多尺度扩张卷积和双向GRU的时间预网,用于神经抖动校正,用于CTC稳定性的时间子采样,以及跨12个编码器块的Pre-RMS Norm稳定性,使用AdamW和余弦调度进行训练。在交互方面,iPhoneme引入了一个和弦凝视加无声语音模式,取代了停留时间选择,实现了更有效的输入。我们在T15数据集(45个会话,8,071次试验)上评估了该系统,该数据集是来自语音运动皮层区域的256通道颅内EEG。在3.1M序列上训练的6-gram音素语言模型,结合WFST波束搜索(波束=128),达到92.14%的音素准确率(7.86% PER)和73.39%单词准确率(26.61%WER),比现有技术水平高出约3%。该系统在CPU上运行,延迟为180 ms,展示了实时性,高精度的脑-文本交流
摘要:Brain-computer interfaces (BCIs) for speech restoration hold transformative potential for the approximately 173,000--232,500 individuals worldwide with ALS-related dysarthria. Despite recent progress, high-performance speech BCIs have been demonstrated in only 22--31 patients globally, largely due to limitations in neural decoding accuracy and practical input interfaces. We present iPhoneme, a brain-to-text communication system that jointly addresses these challenges through integrated modeling and interaction design. The system combines a deep learning phoneme decoder based on a modified Conformer architecture (ConformerXL, 192.9M parameters) with a gaze-assisted phoneme input interface that mitigates the Midas touch problem in eye-tracking systems. The acoustic model incorporates a temporal prenet with multi-scale dilated convolutions and bidirectional GRU for neural jitter correction, temporal subsampling for CTC stability, and Pre-RMSNorm stabilization across 12 encoder blocks, trained with AdamW and cosine scheduling. On the interaction side, iPhoneme introduces a chorded gaze-plus-silent-speech paradigm that replaces dwell-time selection, enabling more efficient input. We evaluate the system on the T15 dataset (45 sessions, 8,071 trials) of 256-channel intracranial EEG from speech motor cortex regions. A 6-gram phoneme language model trained on 3.1M sequences, combined with WFST beam search (beam=128), achieves 92.14% phoneme accuracy (7.86% PER) and 73.39% word accuracy (26.61% WER), approximately 3% above prior state-of-the-art. The system operates on CPU with 180 ms latency, demonstrating real-time, high-accuracy brain-to-text communication for ALS.
【19】NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
标题:NIM 4-ASB:迈向高效、稳健且可定制的实时基于LLM的ASB
链接:https://arxiv.org/abs/2604.18105
摘要:近年来,将大语言模型(LLM)集成到自动语音识别(ASR)中已成为主流模式。尽管现有的基于LLM的ASR模型在公共基准测试中表现出令人印象深刻的性能,但它们的训练仍然主要是数据驱动的,这使得关键的实际挑战没有得到充分解决-特别是在资源受限的部署中有限的向下可扩展性以及在声学挑战条件下的幻觉。为了解决这些问题,我们提出了NIM 4-ASR,这是一个面向生产的基于LLM的ASR框架,针对效率和鲁棒性进行了优化。基于编码器和LLM之间的功能角色的原则性划分,我们重新设计了多阶段训练范式,以使每个模块与其预期的能力边界保持一致。具体而言,我们重新制定了预训练架构和目标,以减轻模态差距并提高参数效率;引入迭代异步SFT阶段以保持声学保真度并约束表示漂移;并设计了ASR专用的强化学习阶段,以进一步提高识别质量和鲁棒性。此外,我们还结合了一套面向生产的优化,包括在嘈杂和安静条件下的鲁棒性,实时流推理,以及通过检索增强生成(RAG)的热词定制。实验表明,NIM 4-ASR在多个公共基准测试中仅用2.3B参数就实现了最先进的性能,同时在内部基准测试中大大优于大规模的竞争对手-特别是在实体密集型的现实世界场景中。NIM 4-ASR进一步支持通过RAG进行百万级热词定制,检索延迟为亚毫秒级,能够有效适应新兴实体和个性化用户需求。
摘要:Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a mainstream paradigm in recent years. Although existing LLM-based ASR models demonstrate impressive performance on public benchmarks, their training remains predominantly data-driven, leaving key practical challenges insufficiently addressed -- particularly limited downward scalability in resource-constrained deployments and hallucinations under acoustically challenging conditions. To address these issues, we present NIM4-ASR, a production-oriented LLM-based ASR framework optimized for both efficiency and robustness. Grounded in a principled delineation of functional roles between the encoder and the LLM, we redesign the multi-stage training paradigm to align each module with its intended capability boundary. Specifically, we reformulate the pre-training architecture and objective to mitigate the modality gap and improve parameter efficiency; introduce an iterative asynchronous SFT stage to preserve acoustic fidelity and constrain representation drift; and design an ASR-specialized reinforcement learning stage to further enhance recognition quality and robustness. We additionally incorporate a suite of production-oriented optimizations, including robustness under noisy and silent conditions, real-time streaming inference, and hotword customization via retrieval-augmented generation (RAG). Experiments show that NIM4-ASR achieves state-of-the-art performance on multiple public benchmarks with merely 2.3B parameters, while substantially outperforming larger-scale competitors on internal benchmarks -- particularly in entity-intensive real-world scenarios. NIM4-ASR further supports million-scale hotword customization via RAG with sub-millisecond retrieval latency, enabling efficient adaptation to emerging entities and personalized user requirements.
【20】MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
标题:MINT-Bench:一个综合性的多语言教学语音转换基准
链接:https://arxiv.org/abs/2604.17958
摘要:指令跟随的文本到语音(TTS)已经成为一个重要的能力,可控和表达的语音生成,但它的评估仍然欠发达,由于有限的基准覆盖范围,诊断粒度弱,并没有足够的多语言支持。我们提出了\textbf{MINT-Bench},一个全面的多语言基准测试,用于解释后续TTS。MINT-Bench建立在分层多轴分类法、可扩展的多阶段数据构建管道和分层混合评估协议之上,该协议联合评估内容一致性、指令遵循和感知质量。十种语言的实验表明,目前的系统还远未解决:前沿商业系统总体领先,而领先的开源模型具有很强的竞争力,甚至可以在中文等本地化环境中超越商业对手。基准测试进一步揭示了更难的成分和语言控制仍然是当前系统的主要瓶颈。我们发布MINT-Bench以及数据构建和评估工具包,以支持未来对可控,多语言和诊断性TTS评估的研究。排行榜和演示可在https://longwaytog0.github.io/MINT-Bench/上获得
摘要:Instruction-following text-to-speech (TTS) has emerged as an important capability for controllable and expressive speech generation, yet its evaluation remains underdeveloped due to limited benchmark coverage, weak diagnostic granularity, and insufficient multilingual support. We present \textbf{MINT-Bench}, a comprehensive multilingual benchmark for instruction-following TTS. MINT-Bench is built upon a hierarchical multi-axis taxonomy, a scalable multi-stage data construction pipeline, and a hierarchical hybrid evaluation protocol that jointly assesses content consistency, instruction following, and perceptual quality. Experiments across ten languages show that current systems remain far from solved: frontier commercial systems lead overall, while leading open-source models become highly competitive and can even outperform commercial counterparts in localized settings such as Chinese. The benchmark further reveals that harder compositional and paralinguistic controls remain major bottlenecks for current systems. We release MINT-Bench together with the data construction and evaluation toolkit to support future research on controllable, multilingual, and diagnostically grounded TTS evaluation. The leaderboard and demo are available at https://longwaytog0.github.io/MINT-Bench/
【21】VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
标题:VUTE:通过现实世界语音对大型音频语言模型进行语音诱导的开放式偏差评估
链接:https://arxiv.org/abs/2604.17248
备注:Submitted to INTERSPEECH 2026
摘要:大型音频语言模型(LALM)越来越多地集成到日常应用中,但其生成偏差仍然没有得到充分的研究。现有的语音公平性基准依赖于合成语音和多项选择题(MCQ),两者都提供了一个支离破碎的公平性视图。我们提出了VIBE,这是一个通过开放式任务(如个性化推荐)评估生成偏差的框架,使用真实世界的人类记录。与MCQs不同,我们的方法允许刻板的关联有机地表现出来,而无需预定义的选项,使其易于扩展到新的任务。评估11个国家的最先进的LALM揭示了系统的偏见,在现实的情况下。我们发现,性别线索往往引发更大的分布变化比口音线索,表明目前LALM再现社会刻板印象。
摘要:Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using real-world human recordings. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 11 state-of-the-art LALMs reveals systematic biases in realistic scenarios. We find that gender cues often trigger larger distributional shifts than accent cues, indicating that current LALMs reproduce social stereotypes.
【22】A state-space representation of the boundary integral equation for room acoustic modelling
标题:房间声学建模的边界积分方程的状态空间表示
链接:https://arxiv.org/abs/2604.16970
备注:14 pages, 6 figures
摘要:我们介绍了一个新的框架,室内声学建模的基础上的状态空间模型的边界积分方程表示在一个房间里的声场。而线性时不变系统的状态空间模型传统上是通过一个状态向量和一个4元组的系统矩阵,在这项工作中引入的状态空间表示由一个状态函数表示的压力分布在房间的边界,和一个4元组的积分算子。我们把这种表示作为一个边界积分算子状态空间(BIOSS)模型,并提供了一个物理解释的积分算子。由于向量和矩阵上的许多数学运算转化为函数和算子,因此可以操纵BIOSS表示以获得具有反馈或并行前馈结构的两个传递函数表示。因此,在BIOSS框架中,在时域或频域中,以及在连续或离散空间中,获得了室内声学的各种等效表示。我们讨论了两个未来的方向如何建议的框架可以肥沃的室内声学建模的研究。首先,我们确定的BIOSS框架和各种现有的房间声学模型(边界元模型,延迟网络,几何模型),这可能是用来建立现有的模型之间的关系,并开发新的房间声学模型之间的等效。其次,我们假设如何从状态空间理论的概念,如可观测性,可控性和状态实现,可以用于开发新的推理和控制方法,室内声学。
摘要:We introduce a new framework for room acoustics modelling based on a state-space model of the boundary integral equation representing the sound field in a room. Whereas state-space models of linear time-invariant systems are traditionally constructed by means of a state vector and a 4-tuple of system matrices, the state-space representation introduced in this work consists of a state function representing the pressure distribution at the room boundary, and a 4-tuple of integral operators. We refer to this representation as a boundary integral operator state-space (BIOSS) model and provide a physical interpretation for each of the integral operators. As many mathematical operations on vectors and matrices translate to functions and operators, the BIOSS representation can be manipulated to obtain two transfer function representations, having either a feedback or a parallel feedforward structure. Consequently, various equivalent representations for room acoustics are obtained in the BIOSS framework, in the time or frequency domain, and in continuous or discrete space. We discuss two future directions for how the proposed framework can be fertile for research on room acoustics modelling. Firstly, we identify equivalences between the BIOSS framework and various existing room acoustics models (boundary element models, delay networks, geometric models), which may be used to establish relations between existing models and to develop novel room acoustics models. Secondly, we postulate on how concepts from state-space theory, such as observability, controllability, and state realization, can be used for developing new inference and control methods for room acoustics.
【23】Deep Hierarchical Knowledge Loss for Fault Intensity Diagnosis
标题:故障强度诊断的深度分层知识丢失
链接:https://arxiv.org/abs/2604.16459
备注:The paper has been accepted by Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 (KDD 2026)
摘要:故障强度诊断(FID)在智能制造中起着关键的作用,但忽略目标类之间的依赖关系阻碍了其实际应用。本文提出了一种新的和一般的框架与深度层次知识损失(DHK),以实现层次一致性表示和预测。我们开发了一种新的层次树损失,使相同属性类的整体映射,利用基于树的积极和消极的层次知识约束。我们进一步设计了一个焦点层次树损失,以提高其可扩展性,并设计了两个自适应加权方案的基础上树的高度。此外,我们提出了一个组树三重损失分层动态余量,通过将分层组的概念和树的距离模型的边界结构知识跨类。联合两个损失显着提高识别的微妙故障。在来自不同工业领域的四个真实数据集(三个来自SAMSON AG的空化数据集和一个公开可用的数据集)上进行了大量的实验,所有这些数据集都显示出优越的结果,并优于最近最先进的FID方法。
摘要:Fault intensity diagnosis (FID) plays a pivotal role in intelligent manufacturing while neglecting dependencies among target classes hinders its practical deployment. This paper introduces a novel and general framework with deep hierarchical knowledge loss (DHK) to achieve hierarchical consistent representation and prediction. We develop a novel hierarchical tree loss to enable a holistic mapping of same-attribute classes, leveraging tree-based positive and negative hierarchical knowledge constraints. We further design a focal hierarchical tree loss to enhance its extensibility and devise two adaptive weighting schemes based on tree height. In addition, we propose a group tree triplet loss with hierarchical dynamic margin by incorporating hierarchical group concepts and tree distance to model boundary structural knowledge across classes. The joint two losses significantly improve the recognition of subtle faults. Extensive experiments are performed on four real-world datasets from various industrial domains (three cavitation datasets from SAMSON AG and one publicly available dataset) for FID, all showing superior results and outperforming recent state-of-the-art FID methods.
【24】TokenChain: A Discrete Speech Chain via Semantic Token Modeling
标题:TokenChain:通过语义代币建模的离散语音链
链接:https://arxiv.org/abs/2510.06201
备注:5 pages, 3 figures. Submitted to IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026
摘要:机器语音链,模拟人类的感知-生产循环,证明有效地联合改善ASR和TTS。我们提出了TokenChain,这是一个完全离散的语音链,它将语义令牌ASR与两阶段TTS耦合在一起:一个与ASR共同训练的自回归文本到语义模型和一个仅用于合成的掩蔽生成语义到声学模型。通过直通argmax/Gumbel-Softmax启用文本界面上的端到端反馈,并通过动态权重平均与监督式ASR进行平衡。烧蚀检查最佳的温度时间表内和跨域转移。评估显示,TokenChain在LibriSpeech上的T2 S稳定,比基线精度提前2-6个epoch,等时误差降低5-13%,在TED-LIUM上的相对ASR WER降低56%,T2 S WER降低31%,遗忘最小,表明链学习对令牌接口和模型仍然有效。
摘要:Machine Speech Chain, simulating the human perception-production loop, proves effective in jointly improving ASR and TTS. We propose TokenChain, a fully discrete speech chain coupling semantic-token ASR with a two-stage TTS: an autoregressive text-to-semantic model co-trained with ASR and a masked-generative semantic-to-acoustic model for synthesis only. End-to-end feedback across the text interface is enabled with straight-through argmax/Gumbel-Softmax and balanced with supervised ASR via dynamic weight averaging. Ablations examine optimal temperature schedules for in- and cross-domain transfer. Evaluation reveals TokenChain surpasses baseline accuracy 2-6 epochs earlier and yields 5-13% lower equal-epoch error with stable T2S on LibriSpeech, and reduces relative ASR WER by 56% and T2S WER by 31% on TED-LIUM with minimal forgetting, showing that chain learning remains effective with token interfaces and models.
【1】Incremental learning for audio classification with Hebbian Deep Neural Networks
标题:使用Hebbian深度神经网络进行音频分类增量学习
链接:https://arxiv.org/abs/2604.18270
备注:ICASSP 2026
摘要:人类终身学习的能力是深度学习方法的灵感,特别是持续学习。在这项工作中,我们应用Hebbian学习,生物启发的学习过程,声音分类。我们提出了一种内核可塑性方法,在增量学习过程中选择性地调节网络内核,作用于选定的内核以学习新信息,并作用于其他内核以保留以前的知识。使用ESC-50数据集,所提出的方法在五个增量步骤中实现了76.3%的整体准确率,优于没有内核可塑性的基线(68.7%),并在任务中表现出更大的稳定性。
摘要:The ability of humans for lifelong learning is an inspiration for deep learning methods and in particular for continual learning. In this work, we apply Hebbian learning, a biologically inspired learning process, to sound classification. We propose a kernel plasticity approach that selectively modulates network kernels during incremental learning, acting on selected kernels to learn new information and on others to retain previous knowledge. Using the ESC-50 dataset, the proposed method achieves 76.3% overall accuracy over five incremental steps, outperforming a baseline without kernel plasticity (68.7%) and demonstrating significantly greater stability across tasks.
【2】NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
标题:NIM 4-ASB:迈向高效、稳健且可定制的实时基于LLM的ASB
链接:https://arxiv.org/abs/2604.18105
摘要:近年来,将大语言模型(LLM)集成到自动语音识别(ASR)中已成为主流模式。尽管现有的基于LLM的ASR模型在公共基准测试中表现出令人印象深刻的性能,但它们的训练仍然主要是数据驱动的,这使得关键的实际挑战没有得到充分解决-特别是在资源受限的部署中有限的向下可扩展性以及在声学挑战条件下的幻觉。为了解决这些问题,我们提出了NIM 4-ASR,这是一个面向生产的基于LLM的ASR框架,针对效率和鲁棒性进行了优化。基于编码器和LLM之间的功能角色的原则性划分,我们重新设计了多阶段训练范式,以使每个模块与其预期的能力边界保持一致。具体而言,我们重新制定了预训练架构和目标,以减轻模态差距并提高参数效率;引入迭代异步SFT阶段以保持声学保真度并约束表示漂移;并设计了ASR专用的强化学习阶段,以进一步提高识别质量和鲁棒性。此外,我们还结合了一套面向生产的优化,包括在嘈杂和安静条件下的鲁棒性,实时流推理,以及通过检索增强生成(RAG)的热词定制。实验表明,NIM 4-ASR在多个公共基准测试中仅用2.3B参数就实现了最先进的性能,同时在内部基准测试中大大优于大规模的竞争对手-特别是在实体密集型的现实世界场景中。NIM 4-ASR进一步支持通过RAG进行百万级热词定制,检索延迟为亚毫秒级,能够有效适应新兴实体和个性化用户需求。
摘要:Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a mainstream paradigm in recent years. Although existing LLM-based ASR models demonstrate impressive performance on public benchmarks, their training remains predominantly data-driven, leaving key practical challenges insufficiently addressed -- particularly limited downward scalability in resource-constrained deployments and hallucinations under acoustically challenging conditions. To address these issues, we present NIM4-ASR, a production-oriented LLM-based ASR framework optimized for both efficiency and robustness. Grounded in a principled delineation of functional roles between the encoder and the LLM, we redesign the multi-stage training paradigm to align each module with its intended capability boundary. Specifically, we reformulate the pre-training architecture and objective to mitigate the modality gap and improve parameter efficiency; introduce an iterative asynchronous SFT stage to preserve acoustic fidelity and constrain representation drift; and design an ASR-specialized reinforcement learning stage to further enhance recognition quality and robustness. We additionally incorporate a suite of production-oriented optimizations, including robustness under noisy and silent conditions, real-time streaming inference, and hotword customization via retrieval-augmented generation (RAG). Experiments show that NIM4-ASR achieves state-of-the-art performance on multiple public benchmarks with merely 2.3B parameters, while substantially outperforming larger-scale competitors on internal benchmarks -- particularly in entity-intensive real-world scenarios. NIM4-ASR further supports million-scale hotword customization via RAG with sub-millisecond retrieval latency, enabling efficient adaptation to emerging entities and personalized user requirements.
【3】MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
标题:MINT-Bench:一个综合性的多语言教学语音转换基准
链接:https://arxiv.org/abs/2604.17958
摘要:指令跟随的文本到语音(TTS)已经成为一个重要的能力,可控和表达的语音生成,但它的评估仍然欠发达,由于有限的基准覆盖范围,诊断粒度弱,并没有足够的多语言支持。我们提出了\textbf{MINT-Bench},一个全面的多语言基准测试,用于解释后续TTS。MINT-Bench建立在分层多轴分类法、可扩展的多阶段数据构建管道和分层混合评估协议之上,该协议联合评估内容一致性、指令遵循和感知质量。十种语言的实验表明,目前的系统还远未解决:前沿商业系统总体领先,而领先的开源模型具有很强的竞争力,甚至可以在中文等本地化环境中超越商业对手。基准测试进一步揭示了更难的成分和语言控制仍然是当前系统的主要瓶颈。我们发布MINT-Bench以及数据构建和评估工具包,以支持未来对可控,多语言和诊断性TTS评估的研究。排行榜和演示可在https://longwaytog0.github.io/MINT-Bench/上获得
摘要:Instruction-following text-to-speech (TTS) has emerged as an important capability for controllable and expressive speech generation, yet its evaluation remains underdeveloped due to limited benchmark coverage, weak diagnostic granularity, and insufficient multilingual support. We present \textbf{MINT-Bench}, a comprehensive multilingual benchmark for instruction-following TTS. MINT-Bench is built upon a hierarchical multi-axis taxonomy, a scalable multi-stage data construction pipeline, and a hierarchical hybrid evaluation protocol that jointly assesses content consistency, instruction following, and perceptual quality. Experiments across ten languages show that current systems remain far from solved: frontier commercial systems lead overall, while leading open-source models become highly competitive and can even outperform commercial counterparts in localized settings such as Chinese. The benchmark further reveals that harder compositional and paralinguistic controls remain major bottlenecks for current systems. We release MINT-Bench together with the data construction and evaluation toolkit to support future research on controllable, multilingual, and diagnostically grounded TTS evaluation. The leaderboard and demo are available at https://longwaytog0.github.io/MINT-Bench/
【4】Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
标题:作为监督的韵律:跨越非言语与言语的桥梁,实现多语言语音情感识别
链接:https://arxiv.org/abs/2604.17647
备注:Accepted to ACL 2026 (main)
摘要:在这项工作中,我们引入了一个低资源的多语言语音情感识别(LRM-SER),利用非言语发声,以利用韵律为中心的情感线索的语言监督范式。与传统的SER系统,严重依赖于标记的口头语音和遭受不良的跨语言的传输,我们的方法重新制定LRM-SER作为非言语到言语的传输,其中监督从标记的非言语源域是适应于未标记的口头语音跨多个目标语言。为此,我们提出了NOVA ARC,一个几何感知的框架,模型的庞加莱球的情感结构,通过双曲矢量量化韵律码本离散化的语言模式,并通过双曲情感镜头捕捉情感强度。对于无监督自适应,NOVA-ARC在源情感原型和目标话语之间执行基于最佳传输的原型对齐,在通过一致性正则化稳定的同时诱导对未标记语音的软监督。实验表明,NOVA-ARC在非言语到言语适应和补充言语到言语迁移设置下都提供了最强的性能,始终优于欧几里得同行和强大的SSL基线。据我们所知,这项工作是第一个超越以言语为中心的监督,通过引入非言语到言语的转移范式SER。
摘要:In this work, we introduce a paralinguistic supervision paradigm for low-resource multilingual speech emotion recognition (LRM-SER) that leverages non-verbal vocalizations to exploit prosody-centric emotion cues. Unlike conventional SER systems that rely heavily on labeled verbal speech and suffer from poor cross-lingual transfer, our approach reformulates LRM-SER as non-verbal-to-verbal transfer, where supervision from a labeled non-verbal source domain is adapted to unlabeled verbal speech across multiple target languages. To this end, we propose NOVA ARC, a geometry-aware framework that models affective structure in the Poincaré ball, discretizes paralinguistic patterns via a hyperbolic vector-quantized prosody codebook, and captures emotion intensity through a hyperbolic emotion lens. For unsupervised adaptation, NOVA-ARC performs optimal transport based prototype alignment between source emotion prototypes and target utterances, inducing soft supervision for unlabeled speech while being stabilized through consistency regularization. Experiments show that NOVA-ARC delivers the strongest performance under both non-verbal-to-verbal adaptation and the complementary verbal-to-verbal transfer setting, consistently outperforming Euclidean counterparts and strong SSL baselines. To the best of our knowledge, this work is the first to move beyond verbal-speech-centric supervision by introducing a non-verbal-to-verbal transfer paradigm for SER.
【5】HCFD: A Benchmark for Audio Deepfake Detection in Healthcare
标题:HCFD:医疗保健领域音频深度伪造检测的基准
链接:https://arxiv.org/abs/2604.17642
备注:Accepted to ACL 2026
摘要:在这项研究中,我们提出了医疗编解码器假检测(HCFD),一个新的任务,在病理语音条件下检测编解码器假。在这项工作中,我们有意专注于基于编解码器的合成语音,因为神经编解码器解码形成了现代语音生成管道中的核心构建块。首先,我们发布了Healthcare CodecFake,这是第一个病理感知数据集,包含多个临床条件和编解码器系列的配对真实和NAC合成语音。我们的评估表明,主要在健康语音上训练的SOTA编解码伪检测器在Healthcare CodecFake上表现不佳,突出了对HCFD特定模型的需求。其次,我们证明了PaSST优于现有的基于语音的模型HCFD,受益于其基于补丁的频谱时间表示。最后,我们提出了PHOENIX-Mamba,这是一个几何感知框架,它将编解码器伪码建模为双曲空间中的多个自发现模式,并在临床条件和编解码器上实现了HCFD的最强性能。在HCFK上的实验表明,PHOENIX-Mamba(PaSST)实现了最好的整体性能,在E-Dep上达到97.04 Acc,在E-Alz上达到96.73,在E-Dys上达到96.57,同时在中国保持了94.41(Dep),94.40(Alz)和93.20(Dys)的强劲结果。这种几何感知的公式使得能够在双曲空间中自发现的异构编解码器伪模式的聚类,从而促进病理语音变化下的鲁棒辨别。PHOENIX-Mamba在各种临床条件和编解码器的HCFD任务中实现了最佳性能。
摘要:In this study, we present Healthcare Codec-Fake Detection (HCFD), a new task for detecting codec-fakes under pathological speech conditions. We intentionally focus on codec based synthetic speech in this work, since neural codec decoding forms a core building block in modern speech generation pipelines. First, we release Healthcare CodecFake, the first pathology-aware dataset containing paired real and NAC-synthesized speech across multipl clinical conditions and codec families. Our evaluations show that SOTA codec-fake detectors trained primarily on healthy speech perform poorly on Healthcare CodecFake, highlighting the need for HCFD-specific models. Second, we demonstrate that PaSST outperforms existing speech-based models for HCFD, benefiting from its patch-based spectro-temporal representation. Finally, we propose PHOENIX-Mamba, a geometry-aware framework that models codec-fakes as multiple self-discovered modes in hyperbolic space and achieves the strongest performance on HCFD across clinical conditions and codecs. Experiments on HCFK show that PHOENIX-Mamba (PaSST) achieves the best overall performance, reaching 97.04 Acc on E-Dep, 96.73 on E-Alz, and 96.57 on E-Dys, while maintaining strong results on Chinese with 94.41 (Dep), 94.40 (Alz), and 93.20 (Dys). This geometry-aware formulation enables self-discovered clustering of heterogeneous codec-fake modes in hyperbolic space, facilitating robust discrimination under pathological speech variability. PHOENIX-Mamba achieves topmost performance on the HCFD task across clinical conditions and codecs.
【6】VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
标题:VUTE:通过现实世界语音对大型音频语言模型进行语音诱导的开放式偏差评估
链接:https://arxiv.org/abs/2604.17248
备注:Submitted to INTERSPEECH 2026
摘要:大型音频语言模型(LALM)越来越多地集成到日常应用中,但其生成偏差仍然没有得到充分的研究。现有的语音公平性基准依赖于合成语音和多项选择题(MCQ),两者都提供了一个支离破碎的公平性视图。我们提出了VIBE,这是一个通过开放式任务(如个性化推荐)评估生成偏差的框架,使用真实世界的人类记录。与MCQs不同,我们的方法允许刻板的关联有机地表现出来,而无需预定义的选项,使其易于扩展到新的任务。评估11个国家的最先进的LALM揭示了系统的偏见,在现实的情况下。我们发现,性别线索往往引发更大的分布变化比口音线索,表明目前LALM再现社会刻板印象。
摘要:Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using real-world human recordings. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 11 state-of-the-art LALMs reveals systematic biases in realistic scenarios. We find that gender cues often trigger larger distributional shifts than accent cues, indicating that current LALMs reproduce social stereotypes.
【7】Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
标题:ananisation,而不是消除:保留效用的语音anisation
链接:https://arxiv.org/abs/2604.17000
摘要:对大规模语音数据的日益依赖使隐私保护成为一个关键问题。然而,现有的匿名化方法通常会降低数据效用,例如通过破坏声学连续性或减少声音多样性,这会损害语音数据用于下游任务的价值,例如自动语音识别(ASR),文本到语音(TTS)和语音情感识别(SER)。目前的评估实践也是有限的,因为它们主要依赖于使用预先训练的模型对匿名语音进行直接测试,仅提供了效用的部分视图。为了解决这些问题,我们提出了一个新的两阶段的框架,保护语言内容和声学身份,同时保持可用性。对于内容隐私,我们采用生成式语音编辑模型来无缝替换个人身份信息(PII),对于语音隐私,我们引入了F3-VA,这是一种基于流匹配的匿名化框架,具有三阶段设计,可产生多样化和独特的匿名扬声器。为了实现更全面的评估,我们使用基于声学和基于内容的说话人验证指标来评估隐私,并通过从头开始训练ASR,TTS和SER模型来评估实用性。实验结果表明,我们的框架实现了更强的隐私保护与最小的效用退化相比,基线从VoicePrivacy挑战,而建议的评估协议提供了一个更现实的反映隐私保护下的匿名语音的效用。
摘要:The growing reliance on large-scale speech data has made privacy protection a critical concern. However, existing anonymization approaches often degrade data utility, for example by disrupting acoustic continuity or reducing vocal diversity, which compromises the value of speech data for downstream tasks such as Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Speech Emotion Recognition (SER). Current evaluation practices are also limited, as they mainly rely on direct testing of anonymized speech with pretrained models, providing only a partial view of utility. To address these issues, we propose a novel two-stage framework that protects both linguistic content and acoustic identity while maintaining usability. For content privacy, we employ a generative speech editing model to seamlessly replace personally identifiable information (PII), and for voice privacy, we introduce F3-VA, a flow-matching-based anonymization framework with a three-stage design that produces diverse and distinct anonymized speakers. To enable a more comprehensive assessment, we evaluate privacy using both acoustic- and content-based speaker verification metrics, and assess utility by training ASR, TTS, and SER models from scratch. Experimental results show that our framework achieves stronger privacy protection with minimal utility degradation compared to baselines from the VoicePrivacy Challenge, while the proposed evaluation protocol provides a more realistic reflection of the utility of anonymized speech under privacy protection.
【8】A state-space representation of the boundary integral equation for room acoustic modelling
标题:房间声学建模的边界积分方程的状态空间表示
链接:https://arxiv.org/abs/2604.16970
备注:14 pages, 6 figures
摘要:我们介绍了一个新的框架,室内声学建模的基础上的状态空间模型的边界积分方程表示在一个房间里的声场。而线性时不变系统的状态空间模型传统上是通过一个状态向量和一个4元组的系统矩阵,在这项工作中引入的状态空间表示由一个状态函数表示的压力分布在房间的边界,和一个4元组的积分算子。我们把这种表示作为一个边界积分算子状态空间(BIOSS)模型,并提供了一个物理解释的积分算子。由于向量和矩阵上的许多数学运算转化为函数和算子,因此可以操纵BIOSS表示以获得具有反馈或并行前馈结构的两个传递函数表示。因此,在BIOSS框架中,在时域或频域中,以及在连续或离散空间中,获得了室内声学的各种等效表示。我们讨论了两个未来的方向如何建议的框架可以肥沃的室内声学建模的研究。首先,我们确定的BIOSS框架和各种现有的房间声学模型(边界元模型,延迟网络,几何模型),这可能是用来建立现有的模型之间的关系,并开发新的房间声学模型之间的等效。其次,我们假设如何从状态空间理论的概念,如可观测性,可控性和状态实现,可以用于开发新的推理和控制方法,室内声学。
摘要:We introduce a new framework for room acoustics modelling based on a state-space model of the boundary integral equation representing the sound field in a room. Whereas state-space models of linear time-invariant systems are traditionally constructed by means of a state vector and a 4-tuple of system matrices, the state-space representation introduced in this work consists of a state function representing the pressure distribution at the room boundary, and a 4-tuple of integral operators. We refer to this representation as a boundary integral operator state-space (BIOSS) model and provide a physical interpretation for each of the integral operators. As many mathematical operations on vectors and matrices translate to functions and operators, the BIOSS representation can be manipulated to obtain two transfer function representations, having either a feedback or a parallel feedforward structure. Consequently, various equivalent representations for room acoustics are obtained in the BIOSS framework, in the time or frequency domain, and in continuous or discrete space. We discuss two future directions for how the proposed framework can be fertile for research on room acoustics modelling. Firstly, we identify equivalences between the BIOSS framework and various existing room acoustics models (boundary element models, delay networks, geometric models), which may be used to establish relations between existing models and to develop novel room acoustics models. Secondly, we postulate on how concepts from state-space theory, such as observability, controllability, and state realization, can be used for developing new inference and control methods for room acoustics.
【9】Neural Encoding Detection is Not All You Need for Synthetic Speech Detection
标题:神经编码检测并不是合成语音检测所需的全部内容
链接:https://arxiv.org/abs/2604.16700
备注:To appear in the proceedings of the IEEE International Workshop on Biometrics and Forensics (IWBF), Sophia Antipolis (France), 2026. Supplementary material available online at: https://neural-isnt-deepfake.github.io/
摘要:本文综述了合成语音检测的现状和发展趋势。它概述了主要的数据驱动方法,讨论了未来研究仅关注神经编码检测的优点和缺点,并为有前途的研究方向提供了建议。与引入新检测方法或数据集的作品不同,本文旨在指导该领域未来的最新研究,并强调过度使用可能经不起时间考验的方法的风险。
摘要:This paper reviews the current state and emerging trends in synthetic speech detection. It outlines the main data-driven approaches, discusses the advantages and drawbacks of focusing future research solely on neural encoding detection, and offers recommendations for promising research directions. Unlike works that introduce new detection methods or datasets, this paper aims to guide future state-of-the-art research in the field and to highlight the risk of overcommitting to approaches that may not stand the test of time.
【10】Deep Hierarchical Knowledge Loss for Fault Intensity Diagnosis
标题:故障强度诊断的深度分层知识丢失
链接:https://arxiv.org/abs/2604.16459
备注:The paper has been accepted by Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 (KDD 2026)
摘要:故障强度诊断(FID)在智能制造中起着关键的作用,但忽略目标类之间的依赖关系阻碍了其实际应用。本文提出了一种新的和一般的框架与深度层次知识损失(DHK),以实现层次一致性表示和预测。我们开发了一种新的层次树损失,使相同属性类的整体映射,利用基于树的积极和消极的层次知识约束。我们进一步设计了一个焦点层次树损失,以提高其可扩展性,并设计了两个自适应加权方案的基础上树的高度。此外,我们提出了一个组树三重损失分层动态余量,通过将分层组的概念和树的距离模型的边界结构知识跨类。联合两个损失显着提高识别的微妙故障。在来自不同工业领域的四个真实数据集(三个来自SAMSON AG的空化数据集和一个公开可用的数据集)上进行了大量的实验,所有这些数据集都显示出优越的结果,并优于最近最先进的FID方法。
摘要:Fault intensity diagnosis (FID) plays a pivotal role in intelligent manufacturing while neglecting dependencies among target classes hinders its practical deployment. This paper introduces a novel and general framework with deep hierarchical knowledge loss (DHK) to achieve hierarchical consistent representation and prediction. We develop a novel hierarchical tree loss to enable a holistic mapping of same-attribute classes, leveraging tree-based positive and negative hierarchical knowledge constraints. We further design a focal hierarchical tree loss to enhance its extensibility and devise two adaptive weighting schemes based on tree height. In addition, we propose a group tree triplet loss with hierarchical dynamic margin by incorporating hierarchical group concepts and tree distance to model boundary structural knowledge across classes. The joint two losses significantly improve the recognition of subtle faults. Extensive experiments are performed on four real-world datasets from various industrial domains (three cavitation datasets from SAMSON AG and one publicly available dataset) for FID, all showing superior results and outperforming recent state-of-the-art FID methods.
【11】SAND: The Challenge on Speech Analysis for Neurodegenerative Disease Assessment
标题:SAND:言语分析用于神经退行性疾病评估的挑战
链接:https://arxiv.org/abs/2604.16445
摘要:人工智能(AI)的最新进展和对非侵入性客观生物标志物(如语音信号)的探索,鼓励了算法的开发,以支持神经退行性疾病(包括肌萎缩侧索硬化症(ALS))的早期诊断。患有ALS的受试者的声音变化通常表现为进行性构音障碍,这是一种突出的神经退行性症状,因为它会随着疾病的进展而影响患者。由于语音信号是复杂的数据,因此开发和使用先进的人工智能技术是从中提取独特模式的基础。使用语音信号验证用于ALS诊断和监测的AI算法具有挑战性,特别是由于缺乏注释的参考数据集。在这项工作中,我们提出了临床医生和机器学习专家的多学科团队之间的合作成果,以创建临床注释的验证数据集和基于它的“神经退行性疾病的语音分析”(SAND)挑战。具体来说,通过分析语音障碍,SAND挑战提供了开发,测试,并评估AI模型用于自动早期识别和预测ALS疾病进展。
摘要:Recent advances in Artificial Intelligence (AI) and the exploration of noninvasive, objective biomarkers, such as speech signals, have encouraged the development of algorithms to support the early diagnosis of neurodegenerative diseases, including Amyotrophic Lateral Sclerosis (ALS). Voice changes in subjects suffering from ALS typically manifest as progressive dysarthria, which is a prominent neurodegenerative symptom because it affects patients as the disease progresses. Since voice signals are complex data, the development and use of advanced AI techniques are fundamental to extracting distinctive patterns from them. Validating AI algorithms for ALS diagnosis and monitoring using voice signals is challenging, particularly due to the lack of annotated reference datasets. In this work, we present the outcome of a collaboration between a multidisciplinary team of clinicians and Machine Learning experts to create both a clinically annotated validation dataset and the "Speech Analysis for Neurodegenerative Diseases" (SAND) challenge based on it. Specifically, by analyzing voice disorders, the SAND challenge provides an opportunity to develop, test, and evaluate AI models for the automatic early identification and prediction of ALS disease progression.
【12】Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints
标题:通过基于规则的音乐约束调整歌词到旋律生成的语言模型
链接:https://arxiv.org/abs/2604.18489
备注:Accepted by IEEE ICASSP 2026
摘要:大型语言模型(LLM)在歌词到旋律生成方面表现出了希望,但使用监督微调(SFT)训练的模型通常会产生音乐上令人难以置信的旋律,例如节奏差和不合适的音域,这种现象我们称之为“约束违反”。为了解决这个问题,我们提出了一个新的对齐框架,灌输音乐知识,而无需人类注释。我们定义了基于规则的音乐约束,以自动生成一个偏好数据集从SFT模型的输出。然后通过顺序过程对模型进行对齐,首先对配对偏好数据使用直接偏好优化(DPO),然后对未配对的阴性样本使用Kahneman-Tversky优化(KTO)。实验结果表明,我们的对齐模型大大减少了违反规则的行为,并在客观和主观评估中优于强基线,生成具有显著改善的音乐性和连贯性的旋律。一个互动演示与音频比较可在https://arain233.github.io/AligningMelody-demo。
摘要:Large Language Models (LLMs) show promise in lyric-to-melody generation, but models trained with Supervised Fine-Tuning (SFT) often produce musically implausible melodies with issues like poor rhythm and unsuitable vocal ranges, a phenomenon we term "constraint violation". To address this, we propose a novel alignment framework that instills musical knowledge without human annotation. We define rule-based musical constraints to automatically generate a preference dataset from an SFT model's outputs. The model is then aligned through a sequential process, first using Direct Preference Optimization (DPO) on paired preference data, followed by Kahneman-Tversky Optimization (KTO) on unpaired negative samples. Experimental results demonstrate that our aligned model substantially reduces rule violations and outperforms strong baselines in both objective and subjective evaluations, generating melodies with substantially improved musicality and coherence. An interactive demo with audio comparisons is available at https://arain233.github.io/AligningMelody-demo.
【13】MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
标题:动作:通过言语到言语翻译中的发声专家混合翻译笑声和泪水
链接:https://arxiv.org/abs/2604.17435
备注:Submitted to Interspeech. Audio Demo and Dataset: https://47zzz.github.io/MoVE/
摘要:最近的语音到语音翻译(S2ST)系统实现了很强的语义准确性,但始终剥离非言语发声(NV),如表达语用意图的笑声和哭泣,这严重限制了现实世界的效用。我们通过三个贡献来解决这个问题。首先,我们提出了一个合成管道来构建可扩展的表达数据集,以克服数据稀缺性的限制。其次,我们提出了MoVE,一个混合的LoRA专家架构与表达专用适配器和软加权路由器,混合专家捕捉混合表达状态。第三,我们展示了预训练的AudioLLM能够实现惊人的数据效率:30分钟的数据就足以实现强大的性能。在英汉S2ST上,与强基线相比,MoVE在76%的情况下再现了目标NVs,并在所有比较系统中实现了最高的人类评级自然度和情感保真度,而现有的S2ST系统最多保留了14%的NVs。
摘要:Recent Speech-to-Speech Translation (S2ST) systems achieve strong semantic accuracy yet consistently strip away non-verbal vocalizations (NVs), such as laughter and crying that convey pragmatic intent, which severely limits real-world utility. We address this via three contributions. First, we propose a synthesis pipeline for building scalable expressive datasets to overcome the data scarcity limitation. Second, we propose MoVE, a Mixture-of-LoRA-Experts architecture with expressive-specialized adapters and a soft-weighting router that blends experts for capturing hybrid expressive states. Third, we show pretrained AudioLLMs enable striking data efficiency: 30 minutes of curated data is enough for strong performance. On English-Chinese S2ST, while comparing with strong baselines, MoVE reproduces target NVs in 76% of cases and achieves the highest human-rated naturalness and emotional fidelity among all compared systems, where existing S2ST systems preserve at most 14% of NVs.
【14】ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection
标题:ICLAT:音频深度伪造检测的上下文学习和比较指导
链接:https://arxiv.org/abs/2604.16749
备注:To appear at ACL Findings 2026
摘要:音频deepfake构成了重大的安全威胁,但目前最先进的(SOTA)检测系统并不能很好地推广到现实的deepfake。我们引入了一种新的\textbf{I}n-\textbf{C}上下文\textbf{L}学习范式,用于\textbf{A}udio \textbf{D}伪检测(\textbf{ICLAD})。该框架允许使用音频语言模型(ALM)对看不见的deepfake进行无训练泛化,并提供检测结果的文本依据。ICLAD的核心是一种成对比较推理策略,该策略指导ALM发现和过滤幻觉和与deepfake无关的声学属性。ALM与专门的deepfake检测器一起工作,其中路由机制将分发样本馈送到ALM。在野外数据集上,ICLAD相对于专用检测器改进了宏F1,相对改进高达$2\times$。进一步的分析表明,国际土地退化问题中心具有灵活性,并有潜力在最近的开放源码土地退化模型上部署。
摘要:Audio deepfakes pose a significant security threat, yet current state-of-the-art (SOTA) detection systems do not generalize well to realistic in-the-wild deepfakes. We introduce a novel \textbf{I}n-\textbf{C}ontext \textbf{L}earning paradigm with comparison-guidance for \textbf{A}udio \textbf{D}eepfake detection (\textbf{ICLAD}). The framework enables the use of audio language models (ALMs) for training-free generalization to unseen deepfakes and provides textual rationales on the detection outcome. At the core of ICLAD is a pairwise comparative reasoning strategy that guides the ALM to discover and filter hallucinations and deepfake-irrelevant acoustic attributes. The ALM works alongside a specialized deepfake detector, whereby a routing mechanism feeds out-of-distribution samples to the ALM. On in-the-wild datasets, ICLAD improves macro F1 over the specialized detector, with up to $2\times$ relative improvement. Further analysis demonstrates the flexibility of ICLAD and its potential for deployment on recent open-source ALMs.
【15】A High-Accuracy Optical Music Recognition Method Based on Bottleneck Residual Convolutions
标题:基于瓶颈剩余卷积的高准确度光学音乐识别方法
链接:https://arxiv.org/abs/2604.16446
备注:2 figs, and 13 tables
摘要:光学音乐识别(OMR)旨在将打印或手写的乐谱图像转换为可编辑的符号表示。本文提出了一个端到端的OMR框架,它结合了残留瓶颈卷积和基于双向门控递归单元(BiGRU)的序列建模。使用具有ResNet-v2风格的残留瓶颈块和多尺度扩张卷积的卷积神经网络来提取对细粒度符号细节和全局staff-line结构进行编码的特征。然后将提取的特征序列馈送到BiGRU网络中,以对音乐符号之间的时间依赖性进行建模。该模型使用连接主义时间分类损失进行训练,从而实现端到端预测,而无需显式对齐注释。在Camera-PrIMuS和PrIMuS数据集上的实验结果证明了该框架的有效性。在Camera-PrIMuS上,该方法的序列错误率(SeER)为7.52%,符号错误率(SyER)为0.45%,音高、类型和音符的准确率分别为99.33%、99.60%和99.28%.平均训练时间为1.74s/epoch,在保持较强识别性能的同时,显示出较高的计算效率。在PrIMuS上,该方法的SeER为8.11美元,SyER为0.49美元,音高、类型和音符准确度分别为99.27美元、99.58美元和99.21美元。细粒度的误差分析进一步证实了所提出的模型的有效性。
摘要:Optical Music Recognition (OMR) aims to convert printed or handwritten music score images into editable symbolic representations. This paper presents an end-to-end OMR framework that combines residual bottleneck convolutions with bidirectional gated recurrent unit (BiGRU)-based sequence modeling. A convolutional neural network with ResNet-v2-style residual bottleneck blocks and multi-scale dilated convolutions is used to extract features that encode both fine-grained symbol details and global staff-line structures. The extracted feature sequences are then fed into a BiGRU network to model temporal dependencies among musical symbols. The model is trained using the Connectionist Temporal Classification loss, enabling end-to-end prediction without explicit alignment annotations. Experimental results on the Camera-PrIMuS and PrIMuS datasets demonstrate the effectiveness of the proposed framework. On Camera-PrIMuS, the proposed method achieves a sequence error rate (SeER) of $7.52\%$ and a symbol error rate (SyER) of $0.45\%$, with pitch, type, and note accuracies of $99.33\%$, $99.60\%$, and $99.28\%$, respectively. The average training time is 1.74~s per epoch, demonstrating high computational efficiency while maintaining strong recognition performance. On PrIMuS, the method achieves a SeER of $8.11\%$ and a SyER of $0.49\%$, with pitch, type, and note accuracies of $99.27\%$, $99.58\%$, and $99.21\%$, respectively. A fine-grained error analysis further confirms the effectiveness of the proposed model.
【16】Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
标题:建立低资源语言中用于多任务理解的语音大型语言模型
链接:https://arxiv.org/abs/2509.14804
摘要:建立在语音编码器、适配器和LLM之上的语音大语言模型(SLLM)在英语和汉语等高资源语言中表现出显著的多任务理解性能。然而,它们的有效性在低资源语言(如泰语)中大幅下降。这种限制来自三个因素:(1)现有的常用语音编码器,如Whisper家族,在低资源语言中表现不佳,并且缺乏对更广泛的口语理解任务的支持;(2)基于ASR的对齐范式需要训练整个SLLM,导致计算成本高;(3)低资源语言中的配对语音文本数据稀缺。为了克服这些挑战,在低资源的语言泰国,我们介绍XLSR-Thai,第一个自我监督学习(SSL)的语音编码器泰国。它是通过在36,000小时的泰语语音数据上连续训练标准SSL XLSR模型获得的。此外,我们提出了U-Align,一种语音文本对齐方法,比典型的基于ASR的对齐方法更具资源效率和多任务效率。最后,我们介绍了Thai-Kings,这是一个从高资源语言中生成泰语口语理解数据的管道,产生了第一个超过1,000小时的泰语口语理解数据集。多个实验证明了我们的方法在建立泰国多任务理解SLLM的有效性。我们开源了XLSR-Thai和Thai-Thai,以促进未来的研究。
摘要:Speech large language models (SLLMs) built on speech encoders, adapters, and LLMs demonstrate remarkable multitask understanding performance in high-resource languages such as English and Chinese. However, their effectiveness substantially degrades in low-resource languages such as Thai. This limitation arises from three factors: (1) existing commonly used speech encoders, like the Whisper family, underperform in low-resource languages and lack support for broader spoken language understanding tasks; (2) the ASR-based alignment paradigm requires training the entire SLLM, leading to high computational cost; (3) paired speech-text data in low-resource languages is scarce. To overcome these challenges in the low-resource language Thai, we introduce XLSR-Thai, the first self-supervised learning (SSL) speech encoder for Thai. It is obtained by continuously training the standard SSL XLSR model on 36,000 hours of Thai speech data. Furthermore, we propose U-Align, a speech-text alignment method that is more resource-efficient and multitask-effective than typical ASR-based alignment. Finally, we present Thai-SUP, a pipeline for generating Thai spoken language understanding data from high-resource languages, yielding the first Thai spoken language understanding dataset of over 1,000 hours. Multiple experiments demonstrate the effectiveness of our methods in building a Thai multitask-understanding SLLM. We open-source XLSR-Thai and Thai-SUP to facilitate future research.
机器翻译由腾讯交互翻译提供,仅供参考
