今日论文合集:cs.SD语音13篇,eess.AS音频处理8篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Few-shot Acoustic Synthesis with Multimodal Flow Matching
标题:利用多峰流匹配的少次声学合成
链接:https://arxiv.org/abs/2603.19176

作者:Amandine Brunetto
备注:To appear at CVPR 2026. 23 pages, 16 figures. Project Page: https://amandinebtto.github.io/FLAC/
摘要:生成与场景声学一致的音频对于沉浸式虚拟环境至关重要。最近的神经声场方法能够实现空间连续的声音渲染,但仍然是场景特定的,需要密集的音频测量和昂贵的培训,为每个环境。Few-Shot方法提高了跨房间的可扩展性,但仍然依赖于多个记录,并且是确定性的,无法捕获稀疏背景下场景声学的固有不确定性。我们介绍了流匹配声学生成(FLAC),概率方法的Few-Shot声学合成,模型的分布合理的房间脉冲响应(RIR)给定的最小场景上下文。FLAC利用一个扩散Transformer,经过流匹配目标训练,以空间、几何和声学线索为条件,在新场景中的任意位置生成RIR。FLAC在AcousticRooms和Hearing Anything Anywhere数据集上的表现优于最先进的八次基线和一次基线。为了补充标准的感知指标,我们进一步引入AGREE,一个联合声学几何嵌入,通过检索和分布指标,使几何一致的评价生成的RIR。这项工作是第一个应用生成流匹配显式RIR合成,建立了一个新的方向,强大的和数据有效的声学合成。
摘要:Generating audio that is acoustically consistent with a scene is essential for immersive virtual environments. Recent neural acoustic field methods enable spatially continuous sound rendering but remain scene-specific, requiring dense audio measurements and costly training for each environment. Few-shot approaches improve scalability across rooms but still rely on multiple recordings and, being deterministic, fail to capture the inherent uncertainty of scene acoustics under sparse context. We introduce flow-matching acoustic generation (FLAC), a probabilistic method for few-shot acoustic synthesis that models the distribution of plausible room impulse responses (RIRs) given minimal scene context. FLAC leverages a diffusion transformer trained with a flow-matching objective to generate RIRs at arbitrary positions in novel scenes, conditioned on spatial, geometric, and acoustic cues. FLAC outperforms state-of-the-art eight-shot baselines with one-shot on both the AcousticRooms and Hearing Anything Anywhere datasets. To complement standard perceptual metrics, we further introduce AGREE, a joint acoustic-geometry embedding, enabling geometry-consistent evaluation of generated RIRs through retrieval and distributional metrics. This work is the first to apply generative flow matching to explicit RIR synthesis, establishing a new direction for robust and data-efficient acoustic synthesis.


【2】Dual-Model Prediction of Affective Engagement and Vocal Attractiveness from Speaker Expressiveness in Video Learning
标题:视频学习中情感投入和声音吸引力的双模型预测
链接:https://arxiv.org/abs/2603.18758

作者:Hung-Yue Suen,Kuo-En Hung,Fan-Hsun Tseng
备注:Preprint. Accepted for publication in IEEE Transactions on Computational Social Systems
摘要:本文概述了一种支持机器学习的以说话者为中心的情感AI方法,该方法能够在基于异步视频的学习中预测观众情感参与和声音吸引力,仅依赖于说话者侧的情感表达。受可扩展的、保护隐私的情感计算应用程序需求的启发,这种以说话者为中心的情感AI方法结合了两种不同的回归模型,利用大规模开放式在线课程(MOOC)中开发的大量语料库来实现情感参与体验。预测情感参与的回归模型是通过吸收面部动态,视觉特征,韵律和认知语义产生的情感表达,同时结合第二个回归模型来预测声音吸引力,完全基于扬声器侧的声学特征。值得注意的是,在独立于说话者的测试集上,两个回归模型都产生了令人印象深刻的预测性能(情感参与的R2 = 0.85,声音吸引力的R2 = 0.88),证实了说话者侧的影响可以在功能上代表聚集的观众反馈。本文提供了一种以说话者为中心的情感AI方法,该方法通过实证研究得到证实,该研究发现说话者侧多模态特征(包括声学)可以前瞻性地预测观众反馈,而不需要使用观众侧输入信息。
摘要:This paper outlines a machine learning-enabled speaker-centric Emotion AI approach capable of predicting audience-affective engagement and vocal attractiveness in asynchronous video-based learning, relying solely on speaker-side affective expressions. Inspired by the demand for scalable, privacy-preserving affective computing applications, this speaker-centric Emotion AI approach incorporates two distinct regression models that leverage a massive corpus developed within Massive Open Online Courses (MOOCs) to enable affectively engaging experiences. The regression model predicting affective engagement is developed by assimilating emotional expressions emanating from facial dynamics, oculomotor features, prosody, and cognitive semantics, while incorporating a second regression model to predict vocal attractiveness based exclusively on speaker-side acoustic features. Notably, on speaker-independent test sets, both regression models yielded impressive predictive performance (R2 = 0.85 for affective engagement and R2 = 0.88 for vocal attractiveness), confirming that speaker-side affect can functionally represent aggregated audience feedback. This paper provides a speaker-centric Emotion AI approach substantiated by an empirical study discovering that speaker-side multimodal features, including acoustics, can prospectively forecast audience feedback without necessarily employing audience-side input information.


【3】Words at Play: Benchmarking Audio Pun Understanding in Large Audio-Language Models
标题:发挥作用的词:大型音频语言模型中的音频双关语理解基准
链接:https://arxiv.org/abs/2603.18678

作者:Yuchen Su,Shaoxin Zhong,Yonghua Zhu,Ruofan Wang,Zijian Huang,Qiqi Wang,Na Zhao,Diana Benavides-Prado,Michael Witbrock
备注:The paper is currently under review
摘要:双关语是一种典型的语言现象,它利用一词多义和语音歧义来产生幽默,给自然语言理解带来了独特的挑战。在双关语研究中,除了文本和图像之外,音频在人类交流中起着核心作用,而口语双关语的数据集和系统资源仍然很少,这使得这一关键形式在很大程度上未被探索。在本文中,我们提出了APUN-Bench,这是第一个致力于评估大型音频语言模型(LALM)对音频双关语理解的基准。我们的基准测试包含4,434个音频样本,分为三个阶段:双关语识别,双关语位置和双关语含义推断。我们通过系统地评估10个最先进的LALM,对APUN-Bench进行深入分析,发现在识别,本地化和解释音频双关语方面存在重大性能差距。该分析揭示了关键的挑战,例如音频双关语位置的位置偏差和意义推断中的错误情况,为推进幽默感知音频智能提供了可操作的见解。
摘要:Puns represent a typical linguistic phenomenon that exploits polysemy and phonetic ambiguity to generate humour, posing unique challenges for natural language understanding. Within pun research, audio plays a central role in human communication except text and images, while datasets and systematic resources for spoken puns remain scarce, leaving this crucial modality largely underexplored. In this paper, we present APUN-Bench, the first benchmark dedicated to evaluating large audio language models (LALMs) on audio pun understanding. Our benchmark contains 4,434 audio samples annotated across three stages: pun recognition, pun word location and pun meaning inference. We conduct a deep analysis of APUN-Bench by systematically evaluating 10 state-of-the-art LALMs, uncovering substantial performance gaps in recognizing, localizing, and interpreting audio puns. This analysis reveals key challenges, such as positional biases in audio pun location and error cases in meaning inference, offering actionable insights for advancing humour-aware audio intelligence.


【4】DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
标题:DiscoPhon:用离散语音单元对音素库存的无监督发现进行基准测试
链接:https://arxiv.org/abs/2603.18612

作者:Maxime Poli,Manel Khentout,Angelo Ortiz Tandazo,Ewan Dunbar,Emmanuel Chemla,Emmanuel Dupoux
备注:6 pages, 2 figures. Submitted to Interspeech 2026
摘要:我们介绍DiscoPhon,一个多语言的基准评估无监督音素发现离散语音单元。DiscoPhon涵盖了6种开发语言和6种测试语言,选择这些语言来跨越广泛的音素对比。给定一种以前看不见的语言,只有10个小时的语音,系统必须通过多对一或一对一的分配产生映射到预定义音素库存的离散单元。由此产生的序列进行评估的单位质量,识别和分割。我们提供了四个预训练的多语言HuBERT和SpidR基线,并表明音素信息在当前模型中足够用于派生单元,以与音素良好相关,尽管不同语言之间存在差异。
摘要:We introduce DiscoPhon, a multilingual benchmark for evaluating unsupervised phoneme discovery from discrete speech units. DiscoPhon covers 6 dev and 6 test languages, chosen to span a wide range of phonemic contrasts. Given only 10 hours of speech in a previously unseen language, systems must produce discrete units that are mapped to a predefined phoneme inventory, through either a many-to-one or a one-to-one assignment. The resulting sequences are evaluated for unit quality, recognition and segmentation. We provide four pretrained multilingual HuBERT and SpidR baselines, and show that phonemic information is available enough in current models for derived units to correlate well with phonemes, though with variations across languages.


【5】Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information
标题:通过稀疏自动编码器实现神经音频编解码器的可解释框架:口音信息的案例研究
链接:https://arxiv.org/abs/2603.18359

作者:Shih-Heng Wang,Tiantian Feng,Aditya Kommineni,Thanathai Lertpetchpun,Bowen Yi,Xuan Shi,Shrikanth Narayanan
摘要:神经音频编解码器(NAC)在现代语音系统中被广泛采用,但它们如何编码语言和非语言信息仍不清楚。提高NAC表示的可解释性对于在敏感应用中理解和部署它们至关重要。因此,我们采用稀疏自动编码器(SAE)将密集的NAC表示分解为稀疏的,可解释的激活。在这项工作中,我们重点关注具有挑战性的副语言属性-口音-并提出了一个量化NAC可解释性的框架。我们评估四个NAC模型下16 SAE配置使用相对性能指数。我们的研究结果表明,DAC和SpeechTokenizer实现了最高的可解释性。我们进一步揭示,声学为导向的NAC编码口音信息主要在激活幅度的稀疏表示,而语音为导向的NAC更依赖于激活位置,低比特率的EnCodec变体表现出更高的可解释性。
摘要:Neural Audio Codecs (NACs) are widely adopted in modern speech systems, yet how they encode linguistic and paralinguistic information remains unclear. Improving the interpretability of NAC representations is critical for understanding and deploying them in sensitive applications. Hence, we employ Sparse Autoencoders (SAEs) to decompose dense NAC representations into sparse, interpretable activations. In this work, we focus on a challenging paralinguistic attribute-accent-and propose a framework to quantify NAC interpretability. We evaluate four NAC models under 16 SAE configurations using a relative performance index. Our results show that DAC and SpeechTokenizer achieve the highest interpretability. We further reveal that acoustic-oriented NACs encode accent information primarily in activation magnitudes of sparse representations, whereas phonetic-oriented NACs rely more on activation positions, and that low-bitrate EnCodec variants show higher interpretability.


【6】ALIGN: Adversarial Learning for Generalizable Speech Neuroprosthesis
标题:ALIGN:可概括言语神经假体的对抗学习
链接:https://arxiv.org/abs/2603.18299

作者:Zhanqi Zhang,Shun Li,Bernardo L. Sabatini,Mikio Aoi,Gal Mishne
摘要:皮质内脑机接口(BCI)在对记录会话中汇集的数据进行训练时,可以高精度地从神经活动中解码语音。然而,在实际部署中,模型必须推广到没有标记数据的新会话,并且性能通常由于跨会话非平稳性(例如,电极移位、神经更替和用户策略的变化)。在本文中,我们提出了ALIGN,一个会话不变的学习框架,基于多域对抗神经网络的半监督跨会话自适应。ALIGN训练一个特征编码器与一个音素分类器和一个域分类器共同操作的潜在表示。通过对抗性优化,编码器被鼓励保留任务相关信息,同时抑制特定于会话的提示。我们评估ALIGN对皮质内语音解码,发现它一贯更好地推广到以前看不见的会议,提高音素错误率和单词错误率相对于基线。这些结果表明,对抗域对齐是一种有效的方法,用于减轻会话级分布偏移,并实现鲁棒的纵向BCI解码。
摘要:Intracortical brain-computer interfaces (BCIs) can decode speech from neural activity with high accuracy when trained on data pooled across recording sessions. In realistic deployment, however, models must generalize to new sessions without labeled data, and performance often degrades due to cross-session nonstationarities (e.g., electrode shifts, neural turnover, and changes in user strategy). In this paper, we propose ALIGN, a session-invariant learning framework based on multi-domain adversarial neural networks for semi-supervised cross-session adaptation. ALIGN trains a feature encoder jointly with a phoneme classifier and a domain classifier operating on the latent representation. Through adversarial optimization, the encoder is encouraged to preserve task-relevant information while suppressing session-specific cues. We evaluate ALIGN on intracortical speech decoding and find that it generalizes consistently better to previously unseen sessions, improving both phoneme error rate and word error rate relative to baselines. These results indicate that adversarial domain alignment is an effective approach for mitigating session-level distribution shift and enabling robust longitudinal BCI decoding.


【7】STEP: Detecting Audio Backdoor Attacks via Stability-based Trigger Exposure Profiling
标题:步骤:通过基于稳定性的触发暴露分析检测音频后门攻击
链接:https://arxiv.org/abs/2603.18103

作者:Kun Wang,Meng Chen,Junhao Wang,Yuli Wu,Li Lu,Chong Zhang,Peng Cheng,Jiaheng Zhang,Kui Ren
摘要:随着基于深度学习的语音模型在安全关键型应用程序中的广泛部署,后门攻击已成为一种严重的威胁:攻击者只要毒化一小部分训练数据,就可以植入一个隐藏的触发器,控制模型的输出,同时保持干净输入的正常行为。现有的推理时间防御不太适合音频领域,因为它们要么依赖于触发器过度鲁棒性假设,而这些假设在基于转换和语义的触发器上失败,要么依赖于特定于图像或文本模态的属性。在本文中,我们提出了STEP(基于稳定性的触发曝光剖析),一个黑盒子,再培训的后门检测器,在硬标签下只访问。其核心思想是利用后门触发器的特征双重异常:语义破坏扰动下的异常标签稳定性和语义保持扰动下的异常标签脆弱性。STEP用两个互补的扰动分支分别针对这两个属性对每个测试样本进行分析,用在良性参考上训练的一类异常检测器对所得的稳定性特征进行评分,并通过无监督加权融合两个分数。在七次后门攻击中进行的广泛实验表明,STEP的平均AUROC为97.92%,EER为4.54%,大大优于最先进的基线,并在模型架构,语音任务,开放集验证场景和无线物理世界设置中进行了推广。
摘要:With the widespread deployment of deep-learning-based speech models in security-critical applications, backdoor attacks have emerged as a serious threat: an adversary who poisons a small fraction of training data can implant a hidden trigger that controls the model's output while preserving normal behavior on clean inputs. Existing inference-time defenses are not well suited to the audio domain, as they either rely on trigger over-robustness assumptions that fail on transformation-based and semantic triggers, or depend on properties specific to image or text modalities. In this paper, we propose STEP (Stability-based Trigger Exposure Profiling), a black-box, retraining-free backdoor detector that operates under hard-label-only access. Its core idea is to exploit a characteristic dual anomaly of backdoor triggers: anomalous label stability under semantic-breaking perturbations, and anomalous label fragility under semantic-preserving perturbations. STEP profiles each test sample with two complementary perturbation branches that target these two properties respectively, scores the resulting stability features with one-class anomaly detectors trained on benign references, and fuses the two scores via unsupervised weighting. Extensive experiments across seven backdoor attacks show that STEP achieves an average AUROC of 97.92% and EER of 4.54%, substantially outperforming state-of-the-art baselines, and generalizes across model architectures, speech tasks, an open-set verification scenario, and over-the-air physical-world settings.


【8】MOSS-TTS Technical Report
标题:MOSS-TTC技术报告
链接:https://arxiv.org/abs/2603.18090

作者:Yitian Gong,Botian Jiang,Yiwei Zhao,Yucheng Yuan,Kuangwei Chen,Yaozhou Jiang,Cheng Chang,Dong Hong,Mingshu Chen,Ruixiao Li,Yiyang Zhang,Yang Gao,Hanfu Chen,Ke Chen,Songlin Wang,Xiaogui Yang,Yuqian Zhang,Kexin Huang,ZhengYuan Lin,Kang Yu,Ziqi Chen,Jin Wang,Zhaoye Fei,Qinyuan Cheng,Shimin Li,Xipeng Qiu
备注:Project page: https://github.com/OpenMOSS/MOSS-TTS
摘要:本技术报告介绍了MOSS-TTS,这是一种基于可扩展配方构建的语音生成基础模型:离散音频令牌,自回归建模和大规模预训练。基于MOSS-Audio-Tokenizer,一个因果Transformer标记器,使用可变比特率RVQ和统一的语义声学表示将24 kHz音频压缩到12.5 fps,我们发布了两个互补的生成器:MOSS-TTS,强调结构简单性、可扩展性和面向长上下文/控制的部署,以及MOSS-TTS-Local-Transformer,引入了帧局部自回归模块以提高建模效率,更强的扬声器保存,以及更短的首次音频时间。MOSS-TTS支持多语言和开放域环境下的zero-shot语音克隆、令牌级时长控制、音素/拼音级发音控制、流畅的语码切换和稳定的长格式生成。该报告总结了已发布模型的设计、训练配方和经验特征。
摘要:This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretraining. Built on MOSS-Audio-Tokenizer, a causal Transformer tokenizer that compresses 24 kHz audio to 12.5 fps with variable-bitrate RVQ and unified semantic-acoustic representations, we release two complementary generators: MOSS-TTS, which emphasizes structural simplicity, scalability, and long-context/control-oriented deployment, and MOSS-TTS-Local-Transformer, which introduces a frame-local autoregressive module for higher modeling efficiency, stronger speaker preservation, and a shorter time to first audio. Across multilingual and open-domain settings, MOSS-TTS supports zero-shot voice cloning, token-level duration control, phoneme-/pinyin-level pronunciation control, smooth code-switching, and stable long-form generation. This report summarizes the design, training recipe, and empirical characteristics of the released models.


【9】EgoAdapt: Enhancing Robustness in Egocentric Interactive Speaker Detection Under Missing Modalities
标题:EgoAdapt:增强缺失模式下自我中心交互式说话人检测的鲁棒性
链接:https://arxiv.org/abs/2603.18082

作者:Xinyuan Qian,Xinjia Zhu,Alessio Brutti,Dong Liang
摘要:TTM(Talking to Me)任务是理解人类社会互动的关键组成部分,旨在确定谁与相机佩戴者进行对话。传统的模型在现实世界的场景中经常面临挑战,因为缺少视觉数据,忽略了头部方向的作用和背景噪声。本研究通过引入EgoAdapt来解决这些限制,EgoAdapt是一种自适应框架,旨在用于丢失模态下的强大的以自我为中心的“跟我说话”说话人检测。具体来说,EgoAdapt包含三个关键模块:(1)视觉说话者目标识别(VSTR)模块,该模块捕获头部方向作为非语言线索,嘴唇运动作为语言线索,允许全面解释语言和非语言信号以解决TTM,将其与仅关注检测说话状态的任务区分开来;(2)用于噪声环境中的增强音频特征提取的并行共享权重音频(PSA)编码器;(3)视觉通道缺失感知(VMMA)该模块估计每帧中每种模态的存在或不存在,以动态调整系统响应。Ego 4D数据集的TTM基准测试表明,EgoAdapt实现了67.39%的平均精度(mAP)和62.01%的准确度(Acc),在准确度方面显著优于最先进的方法4.96%,在mAP方面显著优于最先进的方法1.56%。
摘要:TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to missing visual data, neglecting the role of head orientation, and background noise. This study addresses these limitations by introducing EgoAdapt, an adaptive framework designed for robust egocentric "Talking to Me" speaker detection under missing modalities. Specifically, EgoAdapt incorporates three key modules: (1) a Visual Speaker Target Recognition (VSTR) module that captures head orientation as a non-verbal cue and lip movement as a verbal cue, allowing a comprehensive interpretation of both verbal and non-verbal signals to address TTM, setting it apart from tasks focused solely on detecting speaking status; (2) a Parallel Shared-weight Audio (PSA) encoder for enhanced audio feature extraction in noisy environments; and (3) a Visual Modality Missing Awareness (VMMA) module that estimates the presence or absence of each modality at each frame to adjust the system response dynamically.Comprehensive evaluations on the TTM benchmark of the Ego4D dataset demonstrate that EgoAdapt achieves a mean Average Precision (mAP) of 67.39% and an Accuracy (Acc) of 62.01%, significantly outperforming the state-of-the-art method by 4.96% in Accuracy and 1.56% in mAP.


【10】DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models
标题:DEAF:音频语言模型声学忠实度诊断评估的基准
链接:https://arxiv.org/abs/2603.18048

作者:Jiaqi Xiong,Yunjia Qi,Qi Cao,Yu Zheng,Weisheng Xu,Ziteng Wang,Ruofan Liao,Yutong Zhang,Sichen Liu
备注:14 pages,6 figures
摘要:最近的音频多模态大型语言模型(Audio MLLM)在语音基准测试中表现出令人印象深刻的性能,但目前还不清楚这些模型是否真正处理声学信号或依赖于基于文本的语义推理。为了系统地研究这个问题,我们引入了DEAF(声学忠诚度的诊断评估),这是一个包含2,700多个冲突刺激的基准,涵盖了三个声学维度:情感韵律,背景声音和说话者身份。然后,我们设计了一个受控的多层次的评估框架,逐步增加文本的影响,从语义冲突的内容,误导性提示和它们的组合,使我们能够解开内容驱动的偏见从广告诱导的奉承。我们进一步引入诊断指标来量化模型对声学信号上的文本线索的依赖。我们对七个音频MLLM的评估揭示了文本主导的一致模式:模型对声学变化敏感,但预测主要由文本输入驱动,揭示了标准语音基准测试的高性能与真正的声学理解之间的差距。
摘要:Recent Audio Multimodal Large Language Models (Audio MLLMs) demonstrate impressive performance on speech benchmarks, yet it remains unclear whether these models genuinely process acoustic signals or rely on text-based semantic inference. To systematically study this question, we introduce DEAF (Diagnostic Evaluation of Acoustic Faithfulness), a benchmark of over 2,700 conflict stimuli spanning three acoustic dimensions: emotional prosody, background sounds, and speaker identity. Then, we design a controlled multi-level evaluation framework that progressively increases textual influence, ranging from semantic conflicts in the content to misleading prompts and their combination, allowing us to disentangle content-driven bias from prompt-induced sycophancy. We further introduce diagnostic metrics to quantify model reliance on textual cues over acoustic signals. Our evaluation of seven Audio MLLMs reveals a consistent pattern of text dominance: models are sensitive to acoustic variations, yet predictions are predominantly driven by textual inputs, revealing a gap between high performance on standard speech benchmarks and genuine acoustic understanding.


【11】How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
标题:LLM主干中的听觉知识如何验证音频语言模型:整体评估
链接:https://arxiv.org/abs/2603.19195

作者:Ke-Han Lu,Szu-Wei Fu,Chao-Han Huck Yang,Zhehuai Chen,Sung-Feng Huang,Chih-Kai Yang,Yi-Cheng Lin,Chi-Yuan Hsiao,Wenze Ren,En-Pei Hu,Yu-Han Huang,An-Yu Cheng,Cheng-Han Chiang,Yu Tsao,Yu-Chiang Frank Wang,Hung-yi Lee
备注:Project website: https://kehanlu.github.io/AKB
摘要:大型语言模型(LLM)已被广泛用作大型音频语言模型(LALM)的知识骨干,但它们通过纯文本预训练编码了多少听觉知识,以及这如何影响下游性能仍不清楚。我们通过在两种纯文本和一种音频基础设置下比较不同的LLM来研究这种差距:(1)直接探测AKB-2000,这是一个测试听觉知识广度和深度的策划基准;(2)级联评估,其中LLM对音频字幕的文本描述进行推理;以及(3)基于音频的评估,其中每个LLM被用音频编码器微调成大型音频语言模型(LALM)。我们的研究结果表明,听觉知识在不同的家庭中有很大的差异,纯文本的结果与音频性能密切相关。我们的工作为全面了解LLM在音频研究中的作用提供了经验基础。
摘要:Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-training and how this affects downstream performance remains unclear. We study this gap by comparing different LLMs under two text-only and one audio-grounded setting: (1) direct probing on AKB-2000, a curated benchmark testing the breadth and depth of auditory knowledge; (2) cascade evaluation, where LLMs reason over text descriptions from an audio captioner; and (3) audio-grounded evaluation, where each LLM is fine-tuned into a Large Audio Language Model (LALM) with an audio encoder. Our findings reveal that auditory knowledge varies substantially across families, and text-only results are strongly correlated with audio performance. Our work provides empirical grounding for a comprehensive understanding of LLMs in audio research.


【12】ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
标题:ProKWS:通过音素和韵律的协作学习进行个性化关键词发现
链接:https://arxiv.org/abs/2603.18024

作者:Jianan Pan,Yuanming Zhang,Kejie Huang
摘要:目前的关键词识别系统主要使用音素级匹配来区分易混淆的单词,但忽略了用户特定的发音特征,如韵律(语调,重音,节奏)。本文介绍了ProKWS,一个新的框架集成细粒度音素学习个性化韵律建模。我们设计了一个双流编码器,其中一个流通过对比学习获得强大的音素表示,而另一个提取扬声器特定的韵律模式。一个协同融合模块动态地结合音素和韵律信息,增强了跨声学环境的适应性。实验表明,ProKWS提供了非常有竞争力的性能,与标准基准测试的最先进的模型相媲美,并表现出很强的鲁棒性,个性化的关键字与语气和意图的变化。
摘要:Current keyword spotting systems primarily use phoneme-level matching to distinguish confusable words but ignore user-specific pronunciation traits like prosody (intonation, stress, rhythm). This paper presents ProKWS, a novel framework integrating fine-grained phoneme learning with personalized prosody modeling. We design a dual-stream encoder where one stream derives robust phonemic representations through contrastive learning, while the other extracts speaker-specific prosodic patterns. A collaborative fusion module dynamically combines phonemic and prosodic information, enhancing adaptability across acoustic environments. Experiments show ProKWS delivers highly competitive performance, comparable to state-of-the-art models on standard benchmarks and demonstrates strong robustness for personalized keywords with tone and intent variations.


【13】PCOV-KWS: Multi-task Learning for Personalized Customizable Open Vocabulary Keyword Spotting
标题:PCOV-KWS:用于个性化定制开放词汇关键词发现的多任务学习
链接:https://arxiv.org/abs/2603.18023

作者:Jianan Pan,Kejie Huang
摘要:随着物联网(IoT)、自动语音识别(ASR)、说话人验证(SV)和文本到语音(TTS)等技术的进步,智能语音助手的使用量不断增加,对隐私和个性化的需求也在不断升级。在本文中,我们介绍了一个多任务学习框架的个性化,可定制的开放词汇关键词发现(PCOV-KWS)。该框架采用了一个轻量级的网络,同时执行关键字定位(KWS)和SV,以满足个性化的KWS要求。我们集成了一个不同于基于softmax的损失的训练标准,将多类分类转换为多个二进制分类,消除了类别间的竞争,同时在训练过程中采用了多任务损失加权的优化策略。我们在多个数据集中评估了我们的PCOV-KWS系统,证明它在评估结果中优于基线,同时还需要更少的参数和更低的计算资源。
摘要:As advancements in technologies like Internet of Things (IoT), Automatic Speech Recognition (ASR), Speaker Verification (SV), and Text-to-Speech (TTS) lead to increased usage of intelligent voice assistants, the demand for privacy and personalization has escalated. In this paper, we introduce a multi-task learning framework for personalized, customizable open-vocabulary Keyword Spotting (PCOV-KWS). This framework employs a lightweight network to simultaneously perform Keyword Spotting (KWS) and SV to address personalized KWS requirements. We have integrated a training criterion distinct from softmax-based loss, transforming multi-class classification into multiple binary classifications, which eliminates inter-category competition, while an optimization strategy for multi-task loss weighting is employed during training. We evaluated our PCOV-KWS system in multiple datasets, demonstrating that it outperforms the baselines in evaluation results, while also requiring fewer parameters and lower computational resources.


eess.AS音频处理


【1】How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
标题:LLM主干中的听觉知识如何验证音频语言模型:整体评估
链接:https://arxiv.org/abs/2603.19195

作者:Ke-Han Lu,Szu-Wei Fu,Chao-Han Huck Yang,Zhehuai Chen,Sung-Feng Huang,Chih-Kai Yang,Yi-Cheng Lin,Chi-Yuan Hsiao,Wenze Ren,En-Pei Hu,Yu-Han Huang,An-Yu Cheng,Cheng-Han Chiang,Yu Tsao,Yu-Chiang Frank Wang,Hung-yi Lee
备注:Project website: https://kehanlu.github.io/AKB
摘要:大型语言模型(LLM)已被广泛用作大型音频语言模型(LALM)的知识骨干,但它们通过纯文本预训练编码了多少听觉知识,以及这如何影响下游性能仍不清楚。我们通过在两种纯文本和一种音频基础设置下比较不同的LLM来研究这种差距:(1)直接探测AKB-2000,这是一个测试听觉知识广度和深度的策划基准;(2)级联评估,其中LLM对音频字幕的文本描述进行推理;以及(3)基于音频的评估,其中每个LLM被用音频编码器微调成大型音频语言模型(LALM)。我们的研究结果表明,听觉知识在不同的家庭中有很大的差异,纯文本的结果与音频性能密切相关。我们的工作为全面了解LLM在音频研究中的作用提供了经验基础。
摘要:Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-training and how this affects downstream performance remains unclear. We study this gap by comparing different LLMs under two text-only and one audio-grounded setting: (1) direct probing on AKB-2000, a curated benchmark testing the breadth and depth of auditory knowledge; (2) cascade evaluation, where LLMs reason over text descriptions from an audio captioner; and (3) audio-grounded evaluation, where each LLM is fine-tuned into a Large Audio Language Model (LALM) with an audio encoder. Our findings reveal that auditory knowledge varies substantially across families, and text-only results are strongly correlated with audio performance. Our work provides empirical grounding for a comprehensive understanding of LLMs in audio research.


【2】ARTT: Augmented Reverberant-Target Training for Unsupervised Monaural Speech Dereverberation
标题:ARTT:无监督单耳语音去回响的增强回响目标训练
链接:https://arxiv.org/abs/2603.18485

作者:Siqi Song,Fulin Wu,Zhong-Qiu Wang
备注:in submission
摘要:由于缺乏干净的参考信号和空间线索,单声道无监督语音去混响是一个具有挑战性的不适定逆问题。为了实现它,我们提出了增强混响目标训练(ARTT),它包括两个阶段。在第一阶段中,提出混响目标训练(RTT)以首先进一步混响所观察到的混响混合信号,并且然后训练深度神经网络(DNN)以经由辨别性训练来恢复所观察到的混响混合。虽然要拟合的目标信号是混响的,但我们发现所得到的DNN可以有效地减少混响。在第二阶段,提出了一种基于均值-教师算法的在线自蒸馏机制,以进一步提高去混响效果。评估结果表明,ARTT实现了强大的无监督去混响性能,显着优于以前的基线。
摘要:Due to the absence of clean reference signals and spatial cues, monaural unsupervised speech dereverberation is a challenging ill-posed inverse problem. To realize it, we propose augmented reverberant-target training (ARTT), which consists of two stages. In the first stage, reverberant-target training (RTT) is proposed to first further reverberate the observed reverberant mixture signal, and then train a deep neural network (DNN) to recover the observed reverberant mixture via discriminative training. Although the target signal to fit is reverberant, we find that the resulting DNN can effectively reduce reverberation. In the second stage, an online self-distillation mechanism based on the mean-teacher algorithm is proposed to further improve dereverberation. Evaluation results demonstrate that ARTT achieves strong unsupervised dereverberation performance, significantly outperforming previous baselines.


【3】ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
标题:ProKWS:通过音素和韵律的协作学习进行个性化关键词发现
链接:https://arxiv.org/abs/2603.18024

作者:Jianan Pan,Yuanming Zhang,Kejie Huang
摘要:当前的关键词识别系统主要使用音素级别的匹配来区分易混淆的单词,但忽略了用户特定的发音特征,例如韵律(语调、重读、节奏)。本文介绍了ProKWS,一个新的框架集成细粒度音素学习个性化韵律建模。我们设计了一个双流编码器,其中一个流通过对比学习获得强大的音素表示,而另一个提取扬声器特定的韵律模式。一个协同融合模块动态地结合音素和韵律信息,增强了跨声学环境的适应性。实验表明,ProKWS提供了非常有竞争力的性能,与标准基准测试的最先进的模型相媲美,并表现出很强的鲁棒性,个性化的关键字与语气和意图的变化。
摘要:Current keyword spotting systems primarily use phoneme-level matching to distinguish confusable words but ignore user-specific pronunciation traits like prosody (intonation, stress, rhythm). This paper presents ProKWS, a novel framework integrating fine-grained phoneme learning with personalized prosody modeling. We design a dual-stream encoder where one stream derives robust phonemic representations through contrastive learning, while the other extracts speaker-specific prosodic patterns. A collaborative fusion module dynamically combines phonemic and prosodic information, enhancing adaptability across acoustic environments. Experiments show ProKWS delivers highly competitive performance, comparable to state-of-the-art models on standard benchmarks and demonstrates strong robustness for personalized keywords with tone and intent variations.


【4】PCOV-KWS: Multi-task Learning for Personalized Customizable Open Vocabulary Keyword Spotting
标题:PCOV-KWS:用于个性化定制开放词汇关键词发现的多任务学习
链接:https://arxiv.org/abs/2603.18023

作者:Jianan Pan,Kejie Huang
摘要:随着物联网(IoT)、自动语音识别(ASR)、说话人验证(SV)和文本到语音(TTS)等技术的进步,智能语音助手的使用量不断增加,对隐私和个性化的需求也在不断升级。在本文中,我们介绍了一个多任务学习框架的个性化,可定制的开放词汇关键词发现(PCOV-KWS)。该框架采用了一个轻量级的网络,同时执行关键字定位(KWS)和SV,以满足个性化的KWS要求。我们集成了一个不同于基于softmax的损失的训练标准,将多类分类转换为多个二进制分类,消除了类别间的竞争,同时在训练过程中采用了多任务损失加权的优化策略。我们在多个数据集中评估了我们的PCOV-KWS系统,证明它在评估结果中优于基线,同时还需要更少的参数和更低的计算资源。
摘要:As advancements in technologies like Internet of Things (IoT), Automatic Speech Recognition (ASR), Speaker Verification (SV), and Text-to-Speech (TTS) lead to increased usage of intelligent voice assistants, the demand for privacy and personalization has escalated. In this paper, we introduce a multi-task learning framework for personalized, customizable open-vocabulary Keyword Spotting (PCOV-KWS). This framework employs a lightweight network to simultaneously perform Keyword Spotting (KWS) and SV to address personalized KWS requirements. We have integrated a training criterion distinct from softmax-based loss, transforming multi-class classification into multiple binary classifications, which eliminates inter-category competition, while an optimization strategy for multi-task loss weighting is employed during training. We evaluated our PCOV-KWS system in multiple datasets, demonstrating that it outperforms the baselines in evaluation results, while also requiring fewer parameters and lower computational resources.


【5】Few-shot Acoustic Synthesis with Multimodal Flow Matching
标题:利用多峰流匹配的少次声学合成
链接:https://arxiv.org/abs/2603.19176

作者:Amandine Brunetto
备注:To appear at CVPR 2026. 23 pages, 16 figures. Project Page: https://amandinebtto.github.io/FLAC/
摘要:生成与场景声学一致的音频对于沉浸式虚拟环境至关重要。最近的神经声场方法能够实现空间连续的声音渲染,但仍然是场景特定的,需要密集的音频测量和昂贵的培训,为每个环境。Few-Shot方法提高了跨房间的可扩展性,但仍然依赖于多个记录,并且是确定性的,无法捕获稀疏背景下场景声学的固有不确定性。我们介绍了流匹配声学生成(FLAC),概率方法的Few-Shot声学合成,模型的分布合理的房间脉冲响应(RIR)给定的最小场景上下文。FLAC利用一个扩散Transformer,经过流匹配目标训练,以空间、几何和声学线索为条件,在新场景中的任意位置生成RIR。FLAC在AcousticRooms和Hearing Anything Anywhere数据集上的表现优于最先进的八次基线和一次基线。为了补充标准的感知指标,我们进一步引入AGREE,一个联合声学几何嵌入,通过检索和分布指标,使几何一致的评价生成的RIR。这项工作是第一个应用生成流匹配显式RIR合成,建立了一个新的方向,强大的和数据有效的声学合成。
摘要:Generating audio that is acoustically consistent with a scene is essential for immersive virtual environments. Recent neural acoustic field methods enable spatially continuous sound rendering but remain scene-specific, requiring dense audio measurements and costly training for each environment. Few-shot approaches improve scalability across rooms but still rely on multiple recordings and, being deterministic, fail to capture the inherent uncertainty of scene acoustics under sparse context. We introduce flow-matching acoustic generation (FLAC), a probabilistic method for few-shot acoustic synthesis that models the distribution of plausible room impulse responses (RIRs) given minimal scene context. FLAC leverages a diffusion transformer trained with a flow-matching objective to generate RIRs at arbitrary positions in novel scenes, conditioned on spatial, geometric, and acoustic cues. FLAC outperforms state-of-the-art eight-shot baselines with one-shot on both the AcousticRooms and Hearing Anything Anywhere datasets. To complement standard perceptual metrics, we further introduce AGREE, a joint acoustic-geometry embedding, enabling geometry-consistent evaluation of generated RIRs through retrieval and distributional metrics. This work is the first to apply generative flow matching to explicit RIR synthesis, establishing a new direction for robust and data-efficient acoustic synthesis.


【6】DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
标题:DiscoPhon:用离散语音单元对音素库存的无监督发现进行基准测试
链接:https://arxiv.org/abs/2603.18612

作者:Maxime Poli,Manel Khentout,Angelo Ortiz Tandazo,Ewan Dunbar,Emmanuel Chemla,Emmanuel Dupoux
备注:6 pages, 2 figures. Submitted to Interspeech 2026
摘要:我们介绍DiscoPhon,一个多语言的基准评估无监督音素发现离散语音单元。DiscoPhon涵盖了6种开发语言和6种测试语言,选择这些语言来跨越广泛的音素对比。给定一种以前看不见的语言,只有10个小时的语音,系统必须通过多对一或一对一的分配产生映射到预定义音素库存的离散单元。由此产生的序列进行评估的单位质量,识别和分割。我们提供了四个预训练的多语言HuBERT和SpidR基线,并表明音素信息在当前模型中足够用于派生单元,以与音素良好相关,尽管不同语言之间存在差异。
摘要:We introduce DiscoPhon, a multilingual benchmark for evaluating unsupervised phoneme discovery from discrete speech units. DiscoPhon covers 6 dev and 6 test languages, chosen to span a wide range of phonemic contrasts. Given only 10 hours of speech in a previously unseen language, systems must produce discrete units that are mapped to a predefined phoneme inventory, through either a many-to-one or a one-to-one assignment. The resulting sequences are evaluated for unit quality, recognition and segmentation. We provide four pretrained multilingual HuBERT and SpidR baselines, and show that phonemic information is available enough in current models for derived units to correlate well with phonemes, though with variations across languages.


【7】DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models
标题:DEAF:音频语言模型声学忠实度诊断评估的基准
链接:https://arxiv.org/abs/2603.18048

作者:Jiaqi Xiong,Yunjia Qi,Qi Cao,Yu Zheng,Weisheng Xu,Ziteng Wang,Ruofan Liao,Yutong Zhang,Sichen Liu
备注:14 pages,6 figures
摘要:最近的音频多模态大型语言模型(Audio MLLM)在语音基准测试中表现出令人印象深刻的性能,但目前还不清楚这些模型是否真正处理声学信号或依赖于基于文本的语义推理。为了系统地研究这个问题,我们引入了DEAF(声学忠诚度的诊断评估),这是一个包含2,700多个冲突刺激的基准,涵盖了三个声学维度:情感韵律,背景声音和说话者身份。然后,我们设计了一个受控的多层次的评估框架,逐步增加文本的影响,从语义冲突的内容,误导性提示和它们的组合,使我们能够解开内容驱动的偏见从广告诱导的奉承。我们进一步引入诊断指标来量化模型对声学信号上的文本线索的依赖。我们对七个音频MLLM的评估揭示了文本主导的一致模式:模型对声学变化敏感,但预测主要由文本输入驱动,揭示了标准语音基准测试的高性能与真正的声学理解之间的差距。
摘要:Recent Audio Multimodal Large Language Models (Audio MLLMs) demonstrate impressive performance on speech benchmarks, yet it remains unclear whether these models genuinely process acoustic signals or rely on text-based semantic inference. To systematically study this question, we introduce DEAF (Diagnostic Evaluation of Acoustic Faithfulness), a benchmark of over 2,700 conflict stimuli spanning three acoustic dimensions: emotional prosody, background sounds, and speaker identity. Then, we design a controlled multi-level evaluation framework that progressively increases textual influence, ranging from semantic conflicts in the content to misleading prompts and their combination, allowing us to disentangle content-driven bias from prompt-induced sycophancy. We further introduce diagnostic metrics to quantify model reliance on textual cues over acoustic signals. Our evaluation of seven Audio MLLMs reveals a consistent pattern of text dominance: models are sensitive to acoustic variations, yet predictions are predominantly driven by textual inputs, revealing a gap between high performance on standard speech benchmarks and genuine acoustic understanding.


【8】Modeling Overlapped Speech with Shuffles
标题:用Shuffle建模重叠语音
链接:https://arxiv.org/abs/2603.17769

作者:Matthew Wiesner,Samuele Cornell,Alexander Polok,Lucas Ondel Yang,Lukáš Burget,Sanjeev Khudanpur
摘要:我们建议使用shuffles对并行数据流(如重叠语音)进行建模。具体而言,本文展示了如何洗牌产品和偏序有限状态自动机(FSA)可以用于对齐和扬声器属性的转录重叠的语音。我们使用这些FSA的总得分作为损失函数进行训练,在子字、字和短语级别上对重叠序列的所有可能的序列化进行边缘化。为了减少图的大小,我们施加时间约束,通过构建偏序FSA。我们通过直接建模(token,speaker)元组来解决说话人属性。通过混洗产物FSA的维特比对准直接实现一遍对准。我们评估合成LibriSpeech重叠的性能。据我们所知,这是第一个能够对多人录音进行单次对齐的算法。所有算法都是使用k2 / Icefall实现的。
摘要:We propose to model parallel streams of data, such as overlapped speech, using shuffles. Specifically, this paper shows how the shuffle product and partial order finite-state automata (FSAs) can be used for alignment and speaker-attributed transcription of overlapped speech. We train using the total score on these FSAs as a loss function, marginalizing over all possible serializations of overlapping sequences at subword, word, and phrase levels. To reduce graph size, we impose temporal constraints by constructing partial order FSAs. We address speaker attribution by modeling (token, speaker) tuples directly. Viterbi alignment through the shuffle product FSA directly enables one-pass alignment. We evaluate performance on synthetic LibriSpeech overlaps. To our knowledge, this is the first algorithm that enables single-pass alignment of multi-talker recordings. All algorithms are implemented using k2 / Icefall.


机器翻译由腾讯交互翻译提供,仅供参考