微信公众号:arXiv_Daily
cs.SD语音
【1】A Controllable Perceptual Feature Generative Model for Melody Harmonization via Conditional Variational Autoencoder
标题:基于条件变分自动编码器的旋律协调可控感知特征生成模型
链接:https://arxiv.org/abs/2511.14600
备注:13 pages, 8 figures, 2 url links
摘要:虽然大型语言模型(LLM)使符号音乐的生成变得越来越容易,但制作具有独特成分和丰富表现力的音乐仍然是一个重大挑战。许多研究都引入了情感模型来指导生成过程。然而,这些方法仍然不能提供新颖性和创造性。在音乐信息检索(MIR)领域,听觉感知被认为是音乐体验的一个关键维度,它提供了对作曲意图和情感模式的洞察。为此,我们提出了一个名为CPFG-Net的神经网络,以及将感知特征值映射到和弦表示的转换算法,从而实现旋律协调。该系统可以从给定的旋律中可控地预测感知特征和音调结构的序列,并随后生成和谐连贯的和弦进行。我们的网络是在我们新构建的感知特征数据集BCPT-220 K上训练的,该数据集来自古典音乐。实验结果表明,国家的最先进的感知特征预测能力,我们的模型,以及展示我们的音乐表现力和创造性的和弦推理。这项工作提供了一个新的视角旋律协调,并有助于更广泛的音乐生成任务。我们的基于符号的模型可以很容易地扩展到基于音频的模型。
摘要:While Large Language Models (LLMs) make symbolic music generation increasingly accessible, producing music with distinctive composition and rich expressiveness remains a significant challenge. Many studies have introduced emotion models to guide the generative process. However, these approaches still fall short of delivering novelty and creativity. In the field of Music Information Retrieval (MIR), auditory perception is recognized as a key dimension of musical experience, offering insights into both compositional intent and emotional patterns. To this end, we propose a neural network named CPFG-Net, along with a transformation algorithm that maps perceptual feature values to chord representations, enabling melody harmonization. The system can controllably predict sequences of perceptual features and tonal structures from given melodies, and subsequently generate harmonically coherent chord progressions. Our network is trained on our newly constructed perceptual feature dataset BCPT-220K, derived from classical music. Experimental results show state-of-the-art perceptual feature prediction capability of our model as well as demonstrate our musical expressiveness and creativity in chord inference. This work offers a novel perspective on melody harmonization and contributes to broader music generation tasks. Our symbolic-based model can be easily extended to audio-based models.
【2】IMSE: Efficient U-Net-based Speech Enhancement using Inception Depthwise Convolution and Amplitude-Aware Linear Attention
标题:IMSE:使用初始相关卷积和幅度感知线性注意力的高效基于U-Net的语音增强
链接:https://arxiv.org/abs/2511.14515
摘要:实现轻量化设计和高性能之间的平衡仍然是资源受限设备上语音增强(SE)任务的重大挑战。现有的最先进的方法,如MUSE,通过引入多径增强泰勒(MET)Transformer和可变形嵌入(DE),建立了一个强大的基线,只有0.51 M的参数。然而,深入分析表明,MUSE仍然存在效率瓶颈:MET模块依赖于复杂的“近似补偿”机制来减轻基于泰勒扩展的注意力的限制,而可变形嵌入的偏移计算则引入了额外的计算负担。本文提出了IMSE,一个系统优化和超轻量级的网络。我们引入了两项核心创新:1)用振幅感知线性注意力(MALA)取代MET模块。MALA通过在注意力计算中显式地保留查询向量的范数信息,从根本上纠正了线性注意力中的“幅度忽略”问题,实现了高效的全局建模,而无需辅助补偿分支。2)将DE模块替换为Inception Dependency Convolution(IDConv)。IDConv借用了Inception的概念,将大内核操作分解为高效的并行分支(方形、水平和垂直条带),从而以极低的参数冗余捕获频谱图特征。在VoiceBank+DEMAND数据集上进行的大量实验表明,与MUSE基线相比,IMSE显著减少了16.8%的参数计数(从0.513M减少到0.427M),同时实现了与PESQ指标(3.373)的最新技术水平相当的竞争性能。该研究为超轻量语音增强中模型大小和语音质量之间的权衡设定了新的基准。
摘要:Achieving a balance between lightweight design and high performance remains a significant challenge for speech enhancement (SE) tasks on resource-constrained devices. Existing state-of-the-art methods, such as MUSE, have established a strong baseline with only 0.51M parameters by introducing a Multi-path Enhanced Taylor (MET) transformer and Deformable Embedding (DE). However, an in-depth analysis reveals that MUSE still suffers from efficiency bottlenecks: the MET module relies on a complex "approximate-compensate" mechanism to mitigate the limitations of Taylor-expansion-based attention, while the offset calculation for deformable embedding introduces additional computational burden. This paper proposes IMSE, a systematically optimized and ultra-lightweight network. We introduce two core innovations: 1) Replacing the MET module with Amplitude-Aware Linear Attention (MALA). MALA fundamentally rectifies the "amplitude-ignoring" problem in linear attention by explicitly preserving the norm information of query vectors in the attention calculation, achieving efficient global modeling without an auxiliary compensation branch. 2) Replacing the DE module with Inception Depthwise Convolution (IDConv). IDConv borrows the Inception concept, decomposing large-kernel operations into efficient parallel branches (square, horizontal, and vertical strips), thereby capturing spectrogram features with extremely low parameter redundancy. Extensive experiments on the VoiceBank+DEMAND dataset demonstrate that, compared to the MUSE baseline, IMSE significantly reduces the parameter count by 16.8\% (from 0.513M to 0.427M) while achieving competitive performance comparable to the state-of-the-art on the PESQ metric (3.373). This study sets a new benchmark for the trade-off between model size and speech quality in ultra-lightweight speech enhancement.
【3】Audio Question Answering with GRPO-Based Fine-Tuning and Calibrated Segment-Level Predictions
标题:基于GRPO的微调和校正的分段级预测的音频问题分类
链接:https://arxiv.org/abs/2511.14307
备注:Submission to Track 5 of the DCASE 2025 Challenge
摘要:在本报告中,我们描述了我们提交给DCASE 2025挑战赛第5轨道的音频问题分类(AQA)任务。我们的系统利用SSL骨干BEAT来提取帧级音频特征,然后由分类头处理这些特征,以生成声学事件的分段级预测,遵循Audioset本体。这些片段级预测随后在产生事件级预测之前被校准。最后,这些预测与问题和候选答案一起被合并到结构化提示中。然后,这个提示被馈送到Qwen2.5- 7 B-Instruct的微调版本,使用带有简单奖励函数的GRPO算法进行训练。我们的方法在开发集上实现了62.6%的准确率,证明了将声学事件推理与AQA的预防调整的大型语言模型相结合的有效性。
摘要:In this report, we describe our submission to Track 5 of the DCASE 2025 Challenge for the task of Audio Question Answering(AQA). Our system leverages the SSL backbone BEATs to extract frame-level audio features, which are then processed by a classification head to generate segment-level predictions of acoustic events, following the Audioset ontology. These segment-level predictions are subsequently calibrated before producing event-level predictions. Finally, these predictions are incorporated into a structured prompt, along with the question and candidate answers. This prompt is then fed to a fine-tuned version of Qwen2.5-7B-Instruct, trained using the GRPO algorithm with a simple reward function. Our method achieves an accuracy of 62.6 % on the development set, demonstrating the effectiveness of combining acoustic event reasoning with instruction-tuned large language models for AQA.
【4】Segmentwise Pruning in Audio-Language Models
标题:音频语言模型中的分段剪枝
链接:https://arxiv.org/abs/2511.14293
备注:Submitted to ICASSP 2026 (under review)
摘要:最近的音频语言模型在广泛的音频任务中表现出令人印象深刻的性能,并且越来越能够处理长音频输入。然而,这些模型中的计算成本在很大程度上取决于序列长度,考虑到音频数据的性质,序列长度可能变得非常大。在视觉语言领域,标记修剪方法已被证明可以有效地减少标记数量,同时在标准基准测试中保持强大的性能。在这项工作中,我们调查的相关性和有效性,这样的标记选择策略的背景下,音频语言模型。我们还通过提出考虑时间维度的轻量级策略来改进它们。虽然只保留了初始令牌的四分之一,但我们的方法导致Clotho v2上的CIDER相对最大下降2%,MMAU上的准确性相对最大下降4%。
摘要:Recent audio-language models have shown impressive performance across a wide range of audio tasks and are increasingly capable of handling long audio inputs. However, the computing costs in these models heavily depend on sequence length, which can become very large given the nature of audio data. In the vision-language domain, token pruning methods have proven effective in reducing token counts while preserving strong performance on standard benchmarks. In this work, we investigate the relevance and effectiveness of such token selection strategies in the context of audio-language models. We also improve them by proposing a lightweight strategy that takes the time dimension into account. While retaining only a quarter of the initial tokens, our approach results in a relative maximum decrease of 2% in CIDEr on Clotho v2 and a relative maximum decrease of 4% in accuracy on MMAU.
【5】Count The Notes: Histogram-Based Supervision for Automatic Music Transcription
标题:数数音符:基于柱状图的自动音乐转录监督
链接:https://arxiv.org/abs/2511.14250
备注:ISMIR 2025
摘要:自动音乐转录(AMT)将音频记录转换为符号音乐表示。为AMT训练深度神经网络(DNN)通常需要具有精确帧级注释的高度对齐的训练对。由于创建此类数据集对于许多音乐背景来说成本高昂且不切实际,因此使用片段级注释的弱对齐方法已获得关注。然而,现有的方法往往依赖于动态时间规整(DTW)或软对齐损失函数,这两个仍然需要本地语义对应,使它们容易出错和计算昂贵。在本文中,我们介绍了CountEM,这是一种新颖的AMT框架,它通过利用笔记事件直方图作为监督,消除了显式局部对齐的需要,从而实现了更轻的计算和更大的灵活性。使用期望最大化(EM)方法,CountEM仅基于音符出现次数迭代地细化预测,显著减少注释工作,同时保持高转录准确性。在钢琴、吉他和多乐器数据集上的实验表明,CountEM匹配或超越了现有的弱监督方法,提高了AMT的鲁棒性、可扩展性和效率。我们的项目页面可以在https://yoni-yaffe.github.io/count-the-notes上找到。
摘要:Automatic Music Transcription (AMT) converts audio recordings into symbolic musical representations. Training deep neural networks (DNNs) for AMT typically requires strongly aligned training pairs with precise frame-level annotations. Since creating such datasets is costly and impractical for many musical contexts, weakly aligned approaches using segment-level annotations have gained traction. However, existing methods often rely on Dynamic Time Warping (DTW) or soft alignment loss functions, both of which still require local semantic correspondences, making them error-prone and computationally expensive. In this article, we introduce CountEM, a novel AMT framework that eliminates the need for explicit local alignment by leveraging note event histograms as supervision, enabling lighter computations and greater flexibility. Using an Expectation-Maximization (EM) approach, CountEM iteratively refines predictions based solely on note occurrence counts, significantly reducing annotation efforts while maintaining high transcription accuracy. Experiments on piano, guitar, and multi-instrument datasets demonstrate that CountEM matches or surpasses existing weakly supervised methods, improving AMT's robustness, scalability, and efficiency. Our project page is available at https://yoni-yaffe.github.io/count-the-notes.
【6】Listen Like a Teacher: Mitigating Whisper Hallucinations using Adaptive Layer Attention and Knowledge Distillation
标题:像老师一样倾听:使用自适应层注意力和知识提炼减轻低语幻觉
链接:https://arxiv.org/abs/2511.14219
备注:Accepted at AAAI 2026 - Main Technical Track
摘要:Whisper模型是一个开源的自动语音识别系统,因其在多语言和zero-shot设置中的强大性能而被广泛采用。然而,它经常出现幻觉错误,尤其是在嘈杂的声学条件下。以前减少耳语式ASR系统中的幻觉的工作主要集中在音频预处理或transmittance的后处理上,以过滤出错误的内容。然而,对Whisper模型本身的修改在很大程度上仍未被探索以直接减轻幻觉。为了应对这一挑战,我们提出了一个两阶段的架构,首先通过自适应层注意力(ALA)增强编码器的鲁棒性,并使用多目标知识蒸馏(KD)框架进一步抑制幻觉。在第一阶段中,ALA经由层间相关性分析将编码器层分组为语义上相干的块。然后,一个可学习的多头注意力模块融合这些块表示,使模型能够联合利用低级和高级特征进行更鲁棒的编码。在第二阶段,我们的KD框架在嘈杂的音频上训练学生模型,使其语义和注意力分布与处理干净输入的教师模型保持一致。我们的实验在嘈杂的语音基准显示显着减少幻觉和单词错误率,同时保持对干净的语音性能。ALA和KD共同提供了一个原则性的策略,以提高Whisper在现实世界嘈杂条件下的可靠性。
摘要:The Whisper model, an open-source automatic speech recognition system, is widely adopted for its strong performance across multilingual and zero-shot settings. However, it frequently suffers from hallucination errors, especially under noisy acoustic conditions. Previous works to reduce hallucinations in Whisper-style ASR systems have primarily focused on audio preprocessing or post-processing of transcriptions to filter out erroneous content. However, modifications to the Whisper model itself remain largely unexplored to mitigate hallucinations directly. To address this challenge, we present a two-stage architecture that first enhances encoder robustness through Adaptive Layer Attention (ALA) and further suppresses hallucinations using a multi-objective knowledge distillation (KD) framework. In the first stage, ALA groups encoder layers into semantically coherent blocks via inter-layer correlation analysis. A learnable multi-head attention module then fuses these block representations, enabling the model to jointly exploit low- and high-level features for more robust encoding. In the second stage, our KD framework trains the student model on noisy audio to align its semantic and attention distributions with a teacher model processing clean inputs. Our experiments on noisy speech benchmarks show notable reductions in hallucinations and word error rates, while preserving performance on clean speech. Together, ALA and KD offer a principled strategy to improve Whisper's reliability under real-world noisy conditions.
【7】Preference-Based Learning in Audio Applications: A Systematic Analysis
标题:音频应用中的偏好学习:系统分析
链接:https://arxiv.org/abs/2511.13936
摘要:尽管音频和文本领域在评估生成模型输出方面面临着并行的挑战,但偏好学习在音频应用中仍然显着不足。通过PRISMA指导的约500篇论文的系统综述,我们发现只有30(6%)应用偏好学习音频任务。我们的分析揭示了一个正在转型的领域:2021年之前的研究集中在使用传统排名方法(rankSVM)的情感识别上,而2021年之后的研究则转向使用现代RLHF框架的生成任务。我们确定了三个关键模式:(1)结合合成,自动化和人类偏好的多维评估策略的出现;(2)传统指标(WER,PESQ)与不同背景下的人类判断之间的一致性不一致;(3)结合奖励信号的多阶段训练管道的收敛。我们的研究结果表明,虽然偏好学习显示出对音频的承诺,特别是在捕捉自然性和音乐性等主观品质方面,但该领域需要标准化的基准,更高质量的数据集,以及对音频特有的时间因素如何影响偏好学习框架的系统调查。
摘要:Despite the parallel challenges that audio and text domains face in evaluating generative model outputs, preference learning remains remarkably underexplored in audio applications. Through a PRISMA-guided systematic review of approximately 500 papers, we find that only 30 (6%) apply preference learning to audio tasks. Our analysis reveals a field in transition: pre-2021 works focused on emotion recognition using traditional ranking methods (rankSVM), while post-2021 studies have pivoted toward generation tasks employing modern RLHF frameworks. We identify three critical patterns: (1) the emergence of multi-dimensional evaluation strategies combining synthetic, automated, and human preferences; (2) inconsistent alignment between traditional metrics (WER, PESQ) and human judgments across different contexts; and (3) convergence on multi-stage training pipelines that combine reward signals. Our findings suggest that while preference learning shows promise for audio, particularly in capturing subjective qualities like naturalness and musicality, the field requires standardized benchmarks, higher-quality datasets, and systematic investigation of how temporal factors unique to audio impact preference learning frameworks.
【8】Segmenting Collision Sound Sources in Egocentric Videos
标题:在以自我为中心的视频中分割碰撞声音源
链接:https://arxiv.org/abs/2511.13863
备注:Under Review. Webpage: https://krantiparida.github.io/projects/cs3.html
摘要:人类擅长多感官感知,通常可以从它们相互作用的声音中识别物体属性。受此启发,我们提出了碰撞声源分割(CS3)的新任务,我们的目标是分割负责视觉输入中的碰撞声音的对象(即来自碰撞剪辑的视频帧),以音频为条件。这项任务提出了独特的挑战。与孤立的声音事件不同,碰撞声音来自两个物体之间的相互作用,碰撞的声学特征取决于两者。我们专注于以自我为中心的视频,其中声音通常很清晰,但视觉场景很混乱,对象很小,交互很简短。 为了解决这些挑战,我们提出了一种弱监督的音频条件分割方法,利用基础模型(CLIP和SAM 2)。我们还结合自我中心的线索,即手中的物体,找到可能是碰撞声源的动作物体。在我们为CS3任务引入的两个基准测试中,我们的方法在mIoU方面比竞争基准高出3\times $和4.7\times $:EPIC-CS3和Ego 4D-CS3。
摘要:Humans excel at multisensory perception and can often recognise object properties from the sound of their interactions. Inspired by this, we propose the novel task of Collision Sound Source Segmentation (CS3), where we aim to segment the objects responsible for a collision sound in visual input (i.e. video frames from the collision clip), conditioned on the audio. This task presents unique challenges. Unlike isolated sound events, a collision sound arises from interactions between two objects, and the acoustic signature of the collision depends on both. We focus on egocentric video, where sounds are often clear, but the visual scene is cluttered, objects are small, and interactions are brief. To address these challenges, we propose a weakly-supervised method for audio-conditioned segmentation, utilising foundation models (CLIP and SAM2). We also incorporate egocentric cues, i.e. objects in hands, to find acting objects that can potentially be collision sound sources. Our approach outperforms competitive baselines by $3\times$ and $4.7\times$ in mIoU on two benchmarks we introduce for the CS3 task: EPIC-CS3 and Ego4D-CS3.
【9】Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
标题:通过说话人识别、知识提炼和分层融合进行多说话人对话中的情感识别
链接:https://arxiv.org/abs/2511.13731
摘要:由于说话人歧义和严重的类别不平衡,多说话人会话中的情感识别面临着巨大的挑战。我们提出了一个新的框架,通过三个关键的创新来解决这些问题:(1)说话人识别模块,利用视听同步来准确地识别活跃的说话人,(2)知识蒸馏策略,将卓越的文本情感理解转移到音频和视觉模态,以及(3)分层注意力融合与复合损失函数来处理类不平衡。对MELD和IEMOCAP数据集的综合评估显示出优异的性能,分别达到67.75%和72.44%的加权F1分数,其中少数情绪类的改善尤为显着。
摘要:Emotion recognition in multi-speaker conversations faces significant challenges due to speaker ambiguity and severe class imbalance. We propose a novel framework that addresses these issues through three key innovations: (1) a speaker identification module that leverages audio-visual synchronization to accurately identify the active speaker, (2) a knowledge distillation strategy that transfers superior textual emotion understanding to audio and visual modalities, and (3) hierarchical attention fusion with composite loss functions to handle class imbalance. Comprehensive evaluations on MELD and IEMOCAP datasets demonstrate superior performance, achieving 67.75% and 72.44% weighted F1 scores respectively, with particularly notable improvements on minority emotion classes.
【10】FxSearcher: gradient-free text-driven audio transformation
标题:FxSearcher:无梯度文本驱动的音频转换
链接:https://arxiv.org/abs/2511.14138
摘要:从文本提示实现多样化和高质量的音频转换仍然具有挑战性,因为现有方法从根本上受到其依赖于有限的可区分音频效果集的限制。本文提出了\textbf{FxSearcher},一种新的无梯度框架,发现音频效果(FX)的最佳配置,以根据文本提示转换源信号。我们的方法采用贝叶斯优化和基于CLAP的评分函数来有效地执行此搜索。此外,一个指导提示,以防止不良的文物和提高人类的喜好。为了客观地评估我们的方法,我们提出了一个基于AI的评估框架。结果表明,我们的方法在这些指标上获得的最高分数与人类偏好密切相关。演示可在https://hojoonki.github.io/FxSearcher/上获得
摘要:Achieving diverse and high-quality audio transformations from text prompts remains challenging, as existing methods are fundamentally constrained by their reliance on a limited set of differentiable audio effects. This paper proposes \textbf{FxSearcher}, a novel gradient-free framework that discovers the optimal configuration of audio effects (FX) to transform a source signal according to a text prompt. Our method employs Bayesian Optimization and CLAP-based score function to perform this search efficiently. Furthermore, a guiding prompt is introduced to prevent undesirable artifacts and enhance human preference. To objectively evaluate our method, we propose an AI-based evaluation framework. The results demonstrate that the highest scores achieved by our method on these metrics align closely with human preferences. Demos are available at https://hojoonki.github.io/FxSearcher/
【11】Subject-Independent Imagined Speech Detection via Cross-Subject Generalization and Calibration
标题:通过跨主题概括和校准的与主题无关的想象语音检测
链接:https://arxiv.org/abs/2511.13739
备注:4 pages, 2 figures, Name of Conference: International Conference on Brain-Computer Interface
摘要:由于神经活动模式的显著变化,实现个体之间的鲁棒泛化仍然是基于脑电图的想象语音解码的主要挑战。本研究探讨了如何训练动态和轻量级的主题特定的适应影响跨学科的表现在神经解码框架。一种循环的受试者间训练方法,涉及每个受试者较短的训练段和受试者之间的频繁交替,导致在看不见的目标数据中解码性能的适度但一致的改进。此外,在受试者校准的排除一个受试者方案下,仅纳入10%的目标受试者数据进行校准,实现了0.781的准确度和0.801的AUC,证明了Few Shot适应的有效性。这些发现表明,将循环训练与最小校准相结合,为开发可扩展的、用户自适应的脑机接口系统提供了一种简单有效的策略,该系统可以平衡泛化和个性化。
摘要:Achieving robust generalization across individuals remains a major challenge in electroencephalogram based imagined speech decoding due to substantial variability in neural activity patterns. This study examined how training dynamics and lightweight subject specific adaptation influence cross subject performance in a neural decoding framework. A cyclic inter subject training approach, involving shorter per subject training segments and frequent alternation among subjects, led to modest yet consistent improvements in decoding performance across unseen target data. Furthermore, under the subject calibrated leave one subject out scheme, incorporating only 10 % of the target subjects data for calibration achieved an accuracy of 0.781 and an AUC of 0.801, demonstrating the effectiveness of few shot adaptation. These findings suggest that integrating cyclic training with minimal calibration provides a simple and effective strategy for developing scalable, user adaptive brain computer interface systems that balance generalization and personalization.
【1】TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
标题:TTA:跨语言语音表示的转录、翻译和对齐
链接:https://arxiv.org/abs/2511.14410
备注:Submitted to ICASSP2026
摘要:Speech-LLM模型在多模态和多任务语音理解中表现出很好的性能。一个典型的语音LLM范式是集成语音模态与大型语言模型(LLM)。虽然Whisper编码器在以前的语音输入研究中经常被采用,但它在输入格式,模型规模和语义性能方面存在局限性。为此,我们提出了一个轻量级的TTA模型专门在语音语义更有效的LLM集成。通过对多语言语音识别(ASR)、语音翻译(ST)和语音文本对齐任务的35.8万小时语音数据进行大规模训练,TTA能够生成强大的跨语言语音表示。广泛的评估在不同的基准,包括ASR/ST,语音检索,和ASR-LLM性能评估,证明TTA的优越性耳语。此外,我们严格验证跨语言能力和ASR/ST性能之间的相互作用。TTA的模型权重和训练配方将作为音频理解工具包Auden的一部分发布。
摘要:Speech-LLM models have demonstrated great performance in multi-modal and multi-task speech understanding. A typical speech-LLM paradigm is integrating speech modality with a large language model (LLM). While the Whisper encoder was frequently adopted in previous studies for speech input, it shows limitations regarding input format, model scale, and semantic performance. To this end, we propose a lightweight TTA model specialized in speech semantics for more effective LLM integration. With large-scale training of 358k hours of speech data on multilingual speech recognition (ASR), speech translation (ST) and speech-text alignment tasks, TTA is capable of producing robust cross-lingual speech representations. Extensive evaluations across diverse benchmarks, including ASR/ST, speech retrieval, and ASR-LLM performance assessments, demonstrate TTA's superiority over Whisper. Furthermore, we rigorously validate the interplay between cross-lingual capabilities and ASR/ST performance. The model weights and training recipes of TTA will be released as part of an audio understanding toolkit Auden.
【2】Accelerating Automatic Differentiation of Direct Form Digital Filters
标题:加速直接形式数字过滤器的自动区分
链接:https://arxiv.org/abs/2511.14390
备注:Accepted at the 1st Workshop on Differentiable Systems and Scientific Machine Learning @ EurIPS 2025
摘要:我们介绍了一个一般的配方自动微分通过直接形式的过滤器,产生一个封闭的形式,包括初始条件梯度的反向传播。结果是一个表达式,可以表示过滤器及其梯度计算,同时支持并行性。PyTorch中的C++/CUDA实现比简单的Python实现至少实现了1000倍的加速,并且在GPU上始终运行得最快。对于实际中常用的低阶滤波器,具有解析梯度的精确时域滤波在速度方面优于频域方法。源代码可在https://github.com/yoyolicoris/philtorch上获得。
摘要:We introduce a general formulation for automatic differentiation through direct form filters, yielding a closed-form backpropagation that includes initial condition gradients. The result is a single expression that can represent both the filter and its gradients computation while supporting parallelism. C++/CUDA implementations in PyTorch achieve at least 1000x speedup over naive Python implementations and consistently run fastest on the GPU. For the low-order filters commonly used in practice, exact time-domain filtering with analytical gradients outperforms the frequency-domain method in terms of speed. The source code is available at https://github.com/yoyolicoris/philtorch.
【3】FxSearcher: gradient-free text-driven audio transformation
标题:FxSearcher:无梯度文本驱动的音频转换
链接:https://arxiv.org/abs/2511.14138
摘要:从文本提示实现多样化和高质量的音频转换仍然具有挑战性,因为现有方法从根本上受到其依赖于有限的可区分音频效果集的限制。本文提出了\textbf{FxSearcher},一种新的无梯度框架,发现音频效果(FX)的最佳配置,以根据文本提示转换源信号。我们的方法采用贝叶斯优化和基于CLAP的评分函数来有效地执行此搜索。此外,一个指导提示,以防止不良的文物和提高人类的喜好。为了客观地评估我们的方法,我们提出了一个基于AI的评估框架。结果表明,我们的方法在这些指标上获得的最高分数与人类偏好密切相关。演示可在https://hojoonki.github.io/FxSearcher/上获得
摘要:Achieving diverse and high-quality audio transformations from text prompts remains challenging, as existing methods are fundamentally constrained by their reliance on a limited set of differentiable audio effects. This paper proposes \textbf{FxSearcher}, a novel gradient-free framework that discovers the optimal configuration of audio effects (FX) to transform a source signal according to a text prompt. Our method employs Bayesian Optimization and CLAP-based score function to perform this search efficiently. Furthermore, a guiding prompt is introduced to prevent undesirable artifacts and enhance human preference. To objectively evaluate our method, we propose an AI-based evaluation framework. The results demonstrate that the highest scores achieved by our method on these metrics align closely with human preferences. Demos are available at https://hojoonki.github.io/FxSearcher/
【4】Principled Coarse-Grained Acceptance for Speculative Decoding in Speech
标题:对言语中推测解码的原则粗粒度接受
链接:https://arxiv.org/abs/2511.13732
摘要:推测解码通过让快速草稿模型提出更大目标模型验证的令牌来加速自回归语音生成。然而,对于生成声学令牌的语音LLM,精确的令牌匹配是过度限制的:许多离散令牌在声学上或语义上是可互换的,从而降低了接受率并限制了加速。我们介绍了原则粗粒化(PCG),它验证了来自目标模型的嵌入空间的声学相似性组(ASGs)的水平上的建议。通过将每个令牌的概率质量拆分到包含它的重叠组中,我们定义了一个可感知的粗粒度分布,并对产生的组变量执行拒绝采样。这在组级别上产生了准确性保证,同时允许接受的草稿令牌在实践中代替组中的任何成员。在LibriTTS上,PCG相对于标准推测解码和先前的语音特定松弛提高了接受度和吞吐量,同时保持了可懂度和说话者相似性。这些结果表明,声学感知,组级接受作为一个简单而通用的方法来加速语音令牌生成,同时保持语音质量。
摘要:Speculative decoding accelerates autoregressive speech generation by letting a fast draft model propose tokens that a larger target model verifies. However, for speech LLMs that generate acoustic tokens, exact token matching is overly restrictive: many discrete tokens are acoustically or semantically interchangeable, reducing acceptance rates and limiting speedups. We introduce Principled Coarse-Graining (PCG), which verifies proposals at the level of Acoustic Similarity Groups (ASGs) derived from the target model's embedding space. By splitting each token's probability mass across the overlapping groups that contain it, we define an overlap-aware coarse-grained distribution and perform rejection sampling on the resulting group variable. This yields an exactness guarantee at the group level while allowing the accepted draft token to stand in for any member of the group in practice. On LibriTTS, PCG increases acceptance and throughput relative to standard speculative decoding and prior speech-specific relaxations while maintaining intelligibility and speaker similarity. These results suggest acoustically aware, group-level acceptance as a simple and general way to accelerate speech token generation while maintaining speech quality.
【5】Segmenting Collision Sound Sources in Egocentric Videos
标题:在以自我为中心的视频中分割碰撞声音源
链接:https://arxiv.org/abs/2511.13863
备注:Under Review. Webpage: https://krantiparida.github.io/projects/cs3.html
摘要:人类擅长多感官感知,通常可以从它们相互作用的声音中识别物体属性。受此启发,我们提出了碰撞声源分割(CS3)的新任务,我们的目标是分割负责视觉输入中的碰撞声音的对象(即来自碰撞剪辑的视频帧),以音频为条件。这项任务提出了独特的挑战。与孤立的声音事件不同,碰撞声音来自两个物体之间的相互作用,碰撞的声学特征取决于两者。我们专注于以自我为中心的视频,其中声音通常很清晰,但视觉场景很混乱,对象很小,交互很简短。 为了解决这些挑战,我们提出了一种弱监督的音频条件分割方法,利用基础模型(CLIP和SAM 2)。我们还结合自我中心的线索,即手中的物体,找到可能是碰撞声源的动作物体。在我们为CS3任务引入的两个基准测试中,我们的方法在mIoU方面比竞争基准高出3\times $和4.7\times $:EPIC-CS3和Ego 4D-CS3。
摘要:Humans excel at multisensory perception and can often recognise object properties from the sound of their interactions. Inspired by this, we propose the novel task of Collision Sound Source Segmentation (CS3), where we aim to segment the objects responsible for a collision sound in visual input (i.e. video frames from the collision clip), conditioned on the audio. This task presents unique challenges. Unlike isolated sound events, a collision sound arises from interactions between two objects, and the acoustic signature of the collision depends on both. We focus on egocentric video, where sounds are often clear, but the visual scene is cluttered, objects are small, and interactions are brief. To address these challenges, we propose a weakly-supervised method for audio-conditioned segmentation, utilising foundation models (CLIP and SAM2). We also incorporate egocentric cues, i.e. objects in hands, to find acting objects that can potentially be collision sound sources. Our approach outperforms competitive baselines by $3\times$ and $4.7\times$ in mIoU on two benchmarks we introduce for the CS3 task: EPIC-CS3 and Ego4D-CS3.
【6】Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
标题:通过说话人识别、知识提炼和分层融合进行多说话人对话中的情感识别
链接:https://arxiv.org/abs/2511.13731
摘要:由于说话人歧义和严重的类别不平衡,多说话人会话中的情感识别面临着巨大的挑战。我们提出了一个新的框架,通过三个关键的创新来解决这些问题:(1)说话人识别模块,利用视听同步来准确地识别活跃的说话人,(2)知识蒸馏策略,将卓越的文本情感理解转移到音频和视觉模态,以及(3)分层注意力融合与复合损失函数来处理类不平衡。对MELD和IEMOCAP数据集的综合评估显示出优异的性能,分别达到67.75%和72.44%的加权F1分数,其中少数情绪类的改善尤为显着。
摘要:Emotion recognition in multi-speaker conversations faces significant challenges due to speaker ambiguity and severe class imbalance. We propose a novel framework that addresses these issues through three key innovations: (1) a speaker identification module that leverages audio-visual synchronization to accurately identify the active speaker, (2) a knowledge distillation strategy that transfers superior textual emotion understanding to audio and visual modalities, and (3) hierarchical attention fusion with composite loss functions to handle class imbalance. Comprehensive evaluations on MELD and IEMOCAP datasets demonstrate superior performance, achieving 67.75% and 72.44% weighted F1 scores respectively, with particularly notable improvements on minority emotion classes.
机器翻译由腾讯交互翻译提供,仅供参考
