微信公众号:arXiv_Daily
cs.SD语音
【1】Speaker-Aware Simulation Improves Conversational Speech Recognition
标题:说话者感知模拟改进对话语音识别
链接:https://arxiv.org/abs/2602.04776
摘要:由于大规模、注释良好的多说话人对话数据的有限可用性以及自然交互的复杂时间动态性,会话语音的自动语音识别(ASR)仍然具有挑战性。说话人感知的模拟对话(SASC)提供了一种有效的数据增强策略,通过将单说话人录音转换为真实的多说话人对话。然而,之前的工作主要集中在英语数据上,留下了对低资源语言的适用性的问题。在本文中,我们适应和实现匈牙利会话ASR的SASC框架。我们进一步提出了C-SASC,一个扩展的变体,它结合了基于话语持续时间的停顿建模,使人类对话中观察到的局部时间依赖性的更忠实的表示,同时保留了原始方法的简单性和效率。我们从BEA-Large语料库中生成合成的匈牙利语对话,并将它们与真实的会话数据相结合进行ASR训练。SASC和C-SASC广泛的模拟配置下进行了评估,使用来自CallHome,BEA对话,和GRASS语料库的会话统计。实验结果表明,说话人感知的会话模拟一致地提高了识别性能,天真的级联为基础的增强。虽然C-SASC中的额外持续时间调节产生适度但系统的增益-最明显的是在字符级错误率方面-其有效性取决于源会话统计数据和目标域之间的匹配。总体而言,我们的研究结果证实了匈牙利ASR的说话者感知会话模拟的鲁棒性,并强调了合成对话生成中越来越详细的时间建模的优点和局限性。
摘要:Automatic speech recognition (ASR) for conversational speech remains challenging due to the limited availability of large-scale, well-annotated multi-speaker dialogue data and the complex temporal dynamics of natural interactions. Speaker-aware simulated conversations (SASC) offer an effective data augmentation strategy by transforming single-speaker recordings into realistic multi-speaker dialogues. However, prior work has primarily focused on English data, leaving questions about the applicability to lower-resource languages. In this paper, we adapt and implement the SASC framework for Hungarian conversational ASR. We further propose C-SASC, an extended variant that incorporates pause modeling conditioned on utterance duration, enabling a more faithful representation of local temporal dependencies observed in human conversation while retaining the simplicity and efficiency of the original approach. We generate synthetic Hungarian dialogues from the BEA-Large corpus and combine them with real conversational data for ASR training. Both SASC and C-SASC are evaluated extensively under a wide range of simulation configurations, using conversational statistics derived from CallHome, BEA-Dialogue, and GRASS corpora. Experimental results show that speaker-aware conversational simulation consistently improves recognition performance over naive concatenation-based augmentation. While the additional duration conditioning in C-SASC yields modest but systematic gains--most notably in character-level error rates--its effectiveness depends on the match between source conversational statistics and the target domain. Overall, our findings confirm the robustness of speaker-aware conversational simulation for Hungarian ASR and highlight the benefits and limitations of increasingly detailed temporal modeling in synthetic dialogue generation.
【2】Fine-Grained Frame Modeling in Multi-head Self-Attention for Speech Deepfake Detection
标题:用于语音深度伪造检测的多头自我注意中的细粒度框架建模
链接:https://arxiv.org/abs/2602.04702
备注:Accepted by ICASSP 2026
摘要:基于transformer的模型在语音deepfake检测方面表现出了很强的性能,这主要归功于多头自注意(MHSA)机制的有效性。MHSA提供了帧级注意力分数,这是特别有价值的,因为deepfake伪影通常发生在语音时间维度上的小的局部区域。这使得细粒度的帧建模对于准确检测微妙的欺骗线索至关重要。在这项工作中,我们提出了用于基于MHSA的语音深度伪造检测的细粒度帧建模(FGFM),其中首先通过多头投票(MHV)模块选择信息量最大的帧。然后,这些选定的帧通过跨层细化(cross-layer refinement,缩写为ESTA)模块进行细化,以增强模型学习微妙欺骗线索的能力。实验结果表明,我们的方法优于基线模型,在LA 21,DF 21和ITW数据集上分别实现了0.90%,1.88%和6.64%的等错误率(EER)。这些在多个基准测试中的一致改进凸显了我们的细粒度建模对于鲁棒语音深度伪造检测的有效性。
摘要:Transformer-based models have shown strong performance in speech deepfake detection, largely due to the effectiveness of the multi-head self-attention (MHSA) mechanism. MHSA provides frame-level attention scores, which are particularly valuable because deepfake artifacts often occur in small, localized regions along the temporal dimension of speech. This makes fine-grained frame modeling essential for accurately detecting subtle spoofing cues. In this work, we propose fine-grained frame modeling (FGFM) for MHSA-based speech deepfake detection, where the most informative frames are first selected through a multi-head voting (MHV) module. These selected frames are then refined via a cross-layer refinement (CLR) module to enhance the model's ability to learn subtle spoofing cues. Experimental results demonstrate that our method outperforms the baseline model and achieves Equal Error Rate (EER) of 0.90%, 1.88%, and 6.64% on the LA21, DF21, and ITW datasets, respectively. These consistent improvements across multiple benchmarks highlight the effectiveness of our fine-grained modeling for robust speech deepfake detection.
【3】UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization
标题:UniAudio 2.0:一个统一的音频语言模型,具有文本对齐的分解音频令牌化
链接:https://arxiv.org/abs/2602.04683
摘要:我们研究了音频语言模型中的两个基本问题:(1)如何设计一个音频标记器,可以作为理解和生成的中间表示;(2)如何建立一个音频基础模型,概括在Few-Shot和zero-shot设置,类似于大型语言模型。为此,我们作出以下两项贡献。首先,我们提出了ReasoningCodec,这是一种离散音频编解码器,它将音频分解为(i)推理令牌,它编码文本对齐,高级分析和规划表示,用于音频理解和分层生成,以及(ii)重建令牌,它编码语义丰富的声学线索,用于高保真波形重建。这种设计实现了与强连续表示相媲美的理解性能,同时提高了生成质量和重建保真度。其次,我们引入了一个统一的文本和音频自回归架构,以及多阶段训练和多任务数据构建。使用这个框架,我们训练UniAudio 2.0上的100 B的文本令牌和60 B的音频令牌。在广泛的语音,声音和音乐任务中,UniAudio 2.0在域内评估方面具有竞争力,并对看不见的任务表现出强大的Few-Shot和zero-shot泛化能力。演示、代码和检查点将在\href{https://dongchaoyang.top/UniAudio2Demo/}{https://dongchaoyang.top/UniAudio2Demo/}上提供。
摘要:We study two foundational problems in audio language models: (1) how to design an audio tokenizer that can serve as an intermediate representation for both understanding and generation; and (2) how to build an audio foundation model that generalizes in few-shot and zero-shot settings, analogous to large language models. To this end, we make the following two contributions. First, we propose ReasoningCodec, a discrete audio codec that factorizes audio into (i) reasoning tokens, which encode text-aligned, high-level analysis and planning representations for audio understanding and hierarchical generation, and (ii) reconstruction tokens, which encode semantic-rich acoustic cues for high-fidelity waveform reconstruction. This design achieves understanding performance comparable to strong continuous representations while improving generation quality and reconstruction fidelity over prior discrete tokenizers. Second, we introduce a unified autoregressive architecture for text and audio, together with multi-stage training and multi-task data construction. Using this framework, we train UniAudio 2.0 on 100B text tokens and 60B audio tokens. Across a wide range of speech, sound, and music tasks, UniAudio 2.0 performs competitively on in-domain evaluations and demonstrates strong few-shot and zero-shot generalization to unseen tasks. Demo, code, and checkpoints will be available at \href{https://dongchaoyang.top/UniAudio2Demo/}{https://dongchaoyang.top/UniAudio2Demo/}.
【4】Audio ControlNet for Fine-Grained Audio Generation and Editing
标题:音频控制网用于细粒度音频生成和编辑
链接:https://arxiv.org/abs/2602.04680
摘要:我们研究了细粒度的文本到音频(T2 A)生成任务。虽然最近的模型可以从文本描述中合成高质量的音频,但它们通常缺乏对响度,音高和声音事件等属性的精确控制。与现有的方法,重新训练模型的特定控制类型,我们建议训练ControlNet模型上的预先训练的T2 A骨干,以实现可控的生成响度,音高,和事件滚动。我们介绍了两种设计,T2 A-ControlNet和T2 A-Adapter,并表明T2 A-Adapter模型提供了一个更有效的结构,具有强大的控制能力。仅需38 M额外参数,T2 A-Adapter在事件级和片段级F1评分方面均实现了AudioSet-Strong的最先进性能。我们进一步扩展这个框架的音频编辑,提出T2 A编辑器删除和插入音频事件在指定的时间位置的指令。模型、代码、数据集管道和基准测试将被发布,以支持未来可控音频生成和编辑的研究。
摘要:We study the fine-grained text-to-audio (T2A) generation task. While recent models can synthesize high-quality audio from text descriptions, they often lack precise control over attributes such as loudness, pitch, and sound events. Unlike prior approaches that retrain models for specific control types, we propose to train ControlNet models on top of pre-trained T2A backbones to achieve controllable generation over loudness, pitch, and event roll. We introduce two designs, T2A-ControlNet and T2A-Adapter, and show that the T2A-Adapter model offers a more efficient structure with strong control ability. With only 38M additional parameters, T2A-Adapter achieves state-of-the-art performance on the AudioSet-Strong in both event-level and segment-level F1 scores. We further extend this framework to audio editing, proposing T2A-Editor for removing and inserting audio events at time locations specified by instructions. Models, code, dataset pipelines, and benchmarks will be released to support future research on controllable audio generation and editing.
【5】HoliAntiSpoof: Audio LLM for Holistic Speech Anti-Spoofing
标题:HoliAntiSpoof:用于整体语音反欺骗的音频LLM
链接:https://arxiv.org/abs/2602.04535
摘要:语音合成和编辑的最新进展使得语音欺骗越来越具有挑战性。然而,大多数现有的方法处理欺骗作为二进制分类,忽略了不同的欺骗技术操纵多个,耦合的语音属性和它们的语义效果。在本文中,我们介绍HoliAntiSpoof,第一个音频大语言模型(ALLM)框架,用于整体语音反欺骗分析。HoliAntiSpoof将欺骗分析重新定义为统一的文本生成任务,支持对欺骗方法、受影响的语音属性及其语义影响进行联合推理。为了支持语义级分析,我们引入DailyTalkEdit,一个新的反欺骗基准,模拟现实的会话操作,并提供语义影响的注释。大量的实验表明,HoliAntiSpoof在多个设置中优于传统的基线,而初步结果表明,上下文学习进一步提高了域外泛化。这些研究结果表明,ALLM不仅提高了语音欺骗检测性能,而且还可以对欺骗行为及其语义效果进行可解释的分析,从而提高语音安全性。数据和代码是公开的。
摘要:Recent advances in speech synthesis and editing have made speech spoofing increasingly challenging. However, most existing methods treat spoofing as binary classification, overlooking that diverse spoofing techniques manipulate multiple, coupled speech attributes and their semantic effects. In this paper, we introduce HoliAntiSpoof, the first audio large language model (ALLM) framework for holistic speech anti-spoofing analysis. HoliAntiSpoof reformulates spoofing analysis as a unified text generation task, enabling joint reasoning over spoofing methods, affected speech attributes, and their semantic impacts. To support semantic-level analysis, we introduce DailyTalkEdit, a new anti-spoofing benchmark that simulates realistic conversational manipulations and provides annotations of semantic influence. Extensive experiments demonstrate that HoliAntiSpoof outperforms conventional baselines across multiple settings, while preliminary results show that in-context learning further improves out-of-domain generalization. These findings indicate that ALLMs not only enhance speech spoofing detection performance but also enable interpretable analysis of spoofing behaviors and their semantic effects, pointing towards more trustworthy and explainable speech security. Data and code are publicly available.
【6】DementiaBank-Emotion: A Multi-Rater Emotion Annotation Corpus for Alzheimer's Disease Speech (Version 1.0)
标题:痴呆银行情绪:阿尔茨海默病言语的多评分者情绪注释库(1.0版)
链接:https://arxiv.org/abs/2602.04247
备注:Accepted at HeaLING Workshop @ EACL 2026. 9 pages, 3 figures, 8 tables
摘要:我们提出DementiaBank-Emotion,第一个用于阿尔茨海默病(AD)语音的多评价者情感注释语料库。对来自108名说话者的1,492句话语进行了Ekman的六种基本情绪和中性情绪的注释,我们发现AD患者表达的非中性情绪(16.9%)明显多于健康对照组(5.7%; p <0.001)。探索性声学分析表明可能存在分离:对照组说话者表现出对悲伤的显著F0调制(Δ =基线的-3.45倍),而AD组说话者表现出最小的变化(Δ = +0.11倍;相互作用p = 0.023),尽管这一发现基于有限的样本(悲伤:n=5对照组,n=15 AD),需要重复。在AD语音中,响度区分情感类别,表明部分保留的情感韵律映射。我们发布了语料库,注释指南和校准研讨会材料,以支持临床人群的情绪识别研究。
摘要:We present DementiaBank-Emotion, the first multi-rater emotion annotation corpus for Alzheimer's disease (AD) speech. Annotating 1,492 utterances from 108 speakers for Ekman's six basic emotions and neutral, we find that AD patients express significantly more non-neutral emotions (16.9%) than healthy controls (5.7%; p < .001). Exploratory acoustic analysis suggests a possible dissociation: control speakers showed substantial F0 modulation for sadness (Delta = -3.45 semitones from baseline), whereas AD speakers showed minimal change (Delta = +0.11 semitones; interaction p = .023), though this finding is based on limited samples (sadness: n=5 control, n=15 AD) and requires replication. Within AD speech, loudness differentiates emotion categories, indicating partially preserved emotion-prosody mappings. We release the corpus, annotation guidelines, and calibration workshop materials to support research on emotion recognition in clinical populations.
【7】Frontend Token Enhancement for Token-Based Speech Recognition
标题:基于令牌的语音识别的前端令牌增强
链接:https://arxiv.org/abs/2602.04217
备注:Accepted at ICASSP 2026
摘要:语音信号的离散化表示是各种语音应用中连续特征的有效替代,包括自动语音识别(ASR)和语音语言模型。然而,这些表示,例如从自监督学习(SSL)语音模型的聚类输出导出的语义或语音标记,容易受到环境噪声的影响,这会降低后端任务性能。在这项工作中,我们介绍了一个前端系统,估计干净的语音令牌从嘈杂的语音和评估它的ASR后端使用语义令牌。我们考虑四种类型的增强模型的基础上,他们的输入/输出域:波到波,令牌令牌,连续SSL功能令牌,和波令牌。这些模型独立于ASR后端进行训练。在CHiME-4数据集上的实验表明,波到令牌增强在前端中实现了最佳性能。此外,它的性能大多优于基于连续SSL特征的ASR系统。
摘要:Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and speech language models. However, these representations, such as semantic or phonetic tokens derived from clustering outputs of self-supervised learning (SSL) speech models, are susceptible to environmental noise, which can degrade backend task performance. In this work, we introduce a frontend system that estimates clean speech tokens from noisy speech and evaluate it on an ASR backend using semantic tokens. We consider four types of enhancement models based on their input/output domains: wave-to-wave, token-to-token, continuous SSL features-to-token, and wave-to-token. These models are trained independently of ASR backends. Experiments on the CHiME-4 dataset demonstrate that wave-to-token enhancement achieves the best performance among the frontends. Moreover, it mostly outperforms the ASR system based on continuous SSL features.
【8】PFluxTTS: Hybrid Flow-Matching TTS with Robust Cross-Lingual Voice Cloning and Inference-Time Model Fusion
标题:PFluxTTC:具有稳健的跨语言语音克隆和推理时间模型融合的混合流匹配TTC
链接:https://arxiv.org/abs/2602.04160
备注:Accepted at ICASSP 2026
摘要:我们提出了PFluxTTS,一个混合文本到语音的系统解决三个差距流匹配TTS:稳定性自然度的权衡,弱跨语言的语音克隆,和有限的音频质量从低速率梅尔功能。我们的贡献是:(1)通过推理时间矢量场融合结合持续时间引导和无干扰模型的双解码器设计;(2)在基于FLUX的解码器中使用语音提示嵌入序列的鲁棒克隆,在没有提示转录的情况下保留跨语言的说话者特征;以及(3)具有48 kHz超分辨率的修改的PeriodWave声码器。在跨语言的野外数据中,PFluxTTS明显优于F5-TTS,FishSpeech和SparkTTS,在自然度方面与ChatterBox相匹配(MOS 4.11),同时WER降低了23%(6.9% vs. 9.0%),并且在说话人相似度方面超过了ElevenLabs(+0.32 SMOS)。该系统在大多数开源模型失败的挑战性场景中仍然强大,同时只需要简短的参考音频,无需额外的训练。音频演示可在https://braskai.github.io/pfluxtts/上获得
摘要:We present PFluxTTS, a hybrid text-to-speech system addressing three gaps in flow-matching TTS: the stability-naturalness trade-off, weak cross-lingual voice cloning, and limited audio quality from low-rate mel features. Our contributions are: (1) a dual-decoder design combining duration-guided and alignment-free models through inference-time vector-field fusion; (2) robust cloning using a sequence of speech-prompt embeddings in a FLUX-based decoder, preserving speaker traits across languages without prompt transcripts; and (3) a modified PeriodWave vocoder with super-resolution to 48 kHz. On cross-lingual in-the-wild data, PFluxTTS clearly outperforms F5-TTS, FishSpeech, and SparkTTS, matches ChatterBox in naturalness (MOS 4.11) while achieving 23% lower WER (6.9% vs. 9.0%), and surpasses ElevenLabs in speaker similarity (+0.32 SMOS). The system remains robust in challenging scenarios where most open-source models fail, while requiring only short reference audio and no extra training. Audio demos are available at https://braskai.github.io/pfluxtts/
【9】BASS: Benchmarking Audio LMs for Musical Structure and Semantic Reasoning
标题:BASS:音乐结构和语义推理的音频LM基准
链接:https://arxiv.org/abs/2602.04085
摘要:音乐理解是一项复杂的任务,通常需要对音频的结构和语义元素进行推理。我们介绍BASS,旨在评估音频语言模型中的音乐理解和推理,分为四大类:结构分割,歌词转录,音乐学分析和艺术家合作。BASS由2658个问题组成,涵盖12个任务,1993首独特的歌曲,涵盖了超过138小时的音乐,来自广泛的流派和曲目,旨在评估音乐学知识和现实世界中的推理。我们评估了14个开源和前沿的多模态LM,发现即使是最先进的模型也难以完成更高层次的推理任务,如结构分割和艺术家合作,而在歌词转录方面表现最好。我们的分析表明,目前的模型有效地利用语言先验知识,但在音乐结构,声乐和音乐学属性的推理仍然有限。BASS提供了一个在音乐推荐和搜索中具有广泛应用的评估框架,并具有指导音频LM开发的潜力。
摘要:Music understanding is a complex task that often requires reasoning over both structural and semantic elements of audio. We introduce BASS, designed to evaluate music understanding and reasoning in audio language models across four broad categories: structural segmentation, lyric transcription, musicological analysis, and artist collaboration. BASS comprises 2658 questions spanning 12 tasks, 1993 unique songs and covering over 138 hours of music from a wide range of genres and tracks, crafted to assess musicological knowledge and reasoning in real-world scenarios. We evaluate 14 open-source and frontier multimodal LMs, finding that even state-of-the-art models struggle on higher-level reasoning tasks such as structural segmentation and artist collaboration, while performing best on lyric transcription. Our analysis reveals that current models leverage linguistic priors effectively but remain limited in reasoning over musical structure, vocal, and musicological attributes. BASS provides an evaluation framework with widespread applications in music recommendation and search and has the potential to guide the development of audio LMs.
【10】Audit After Segmentation: Reference-Free Mask Quality Assessment for Language-Referred Audio-Visual Segmentation
标题:分割后的审核:参考视听分割的无参考面罩质量评估
链接:https://arxiv.org/abs/2602.03892
摘要:参考音视频分割(Ref-AVS)的目的是通过对视频、音频和文本的联合推理来分割由自然语言描述的目标对象。除了生成分割掩模之外,提供掩模质量的丰富和可解释的诊断在很大程度上仍然未被探索。在这项工作中,我们介绍了掩模质量评估在Ref-AVS上下文中(MQA-RefAVS),一个新的任务,评估候选分割掩模的质量,而不依赖于地面实况注释作为参考在推理时间。给定视听语言输入和每个提供的分割掩码,该任务需要使用未观察到的地面事实估计其IoU,识别相应的错误类型,并推荐可操作的质量控制决策。为了支持这一任务,我们构建MQ-RAVSBench,一个基准测试,具有不同的和代表性的掩模错误模式,跨越几何和语义问题。我们进一步提出了MQ-Auditor,一个基于多模态大语言模型(MLLM)的审计师,明确的原因多模态线索和掩码信息,以产生定量和定性的掩码质量评估。大量的实验表明,MQ-Auditor优于强大的开源和商业MLLM,可以与现有的Ref-AVS系统集成,以检测分割失败,并支持下游分割改进。数据和代码将在https://github.com/jasongief/MQA-RefAVS上发布。
摘要:Language-referred audio-visual segmentation (Ref-AVS) aims to segment target objects described by natural language by jointly reasoning over video, audio, and text. Beyond generating segmentation masks, providing rich and interpretable diagnoses of mask quality remains largely underexplored. In this work, we introduce Mask Quality Assessment in the Ref-AVS context (MQA-RefAVS), a new task that evaluates the quality of candidate segmentation masks without relying on ground-truth annotations as references at inference time. Given audio-visual-language inputs and each provided segmentation mask, the task requires estimating its IoU with the unobserved ground truth, identifying the corresponding error type, and recommending an actionable quality-control decision. To support this task, we construct MQ-RAVSBench, a benchmark featuring diverse and representative mask error modes that span both geometric and semantic issues. We further propose MQ-Auditor, a multimodal large language model (MLLM)-based auditor that explicitly reasons over multimodal cues and mask information to produce quantitative and qualitative mask quality assessments. Extensive experiments demonstrate that MQ-Auditor outperforms strong open-source and commercial MLLMs and can be integrated with existing Ref-AVS systems to detect segmentation failures and support downstream segmentation improvement. Data and codes will be released at https://github.com/jasongief/MQA-RefAVS.
【11】Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models
标题:通过音频语言模型中的测试时间缩放解码模糊情绪
链接:https://arxiv.org/abs/2602.03873
摘要:从人类语音中识别情感是社交意识会话AI的关键推动因素。然而,虽然大多数先前的工作框架的情感识别作为一个分类问题,现实世界的情感状态往往是模糊的,重叠的,和上下文相关的,提出了重大挑战的注释和自动建模。最近的大规模音频语言模型(ALMs)提供了新的机会,细致入微的情感推理没有明确的情感监督,但他们的能力,以处理模棱两可的情绪仍然未充分发掘。与此同时,推理时间技术的进步,如测试时间缩放(TTS),已经显示出改善硬NLP任务的泛化和适应性的希望,但它们与情感计算的相关性在很大程度上仍然未知。在这项工作中,我们介绍了第一个基准模糊的情感识别语音与ALMs下测试时间缩放。我们的评估系统地比较了八个国家的最先进的ALM和五个TTS策略在三个突出的语音情感数据集。我们进一步深入分析了模型容量、TTS和情感模糊之间的相互作用,为模糊情感理解的计算和表征挑战提供了新的见解。我们的基准为开发更强大,上下文感知和情感智能的基于语音的AI系统奠定了基础,并强调了弥合模型假设与现实世界人类情感复杂性之间差距的关键未来方向。
摘要:Emotion recognition from human speech is a critical enabler for socially aware conversational AI. However, while most prior work frames emotion recognition as a categorical classification problem, real-world affective states are often ambiguous, overlapping, and context-dependent, posing significant challenges for both annotation and automatic modeling. Recent large-scale audio language models (ALMs) offer new opportunities for nuanced affective reasoning without explicit emotion supervision, but their capacity to handle ambiguous emotions remains underexplored. At the same time, advances in inference-time techniques such as test-time scaling (TTS) have shown promise for improving generalization and adaptability in hard NLP tasks, but their relevance to affective computing is still largely unknown. In this work, we introduce the first benchmark for ambiguous emotion recognition in speech with ALMs under test-time scaling. Our evaluation systematically compares eight state-of-the-art ALMs and five TTS strategies across three prominent speech emotion datasets. We further provide an in-depth analysis of the interaction between model capacity, TTS, and affective ambiguity, offering new insights into the computational and representational challenges of ambiguous emotion understanding. Our benchmark establishes a foundation for developing more robust, context-aware, and emotionally intelligent speech-based AI systems, and highlights key future directions for bridging the gap between model assumptions and the complexity of real-world human emotion.
【12】LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
标题:LALM作为法官:在多轮口语对话中对大型音频语言模型进行安全评估的基准
链接:https://arxiv.org/abs/2602.04796
摘要:语音代理之间的对话越来越普遍,但评估它们的社会有害内容,如暴力,骚扰和仇恨仍然以文本为中心,未能考虑音频特定的线索和转录错误。我们提出了LALM作为法官,这是第一个对大型音频语言模型(LALM)作为多回合口语对话安全法官的受控基准和系统研究。我们生成了24,000个不安全和合成的英语口语对话,由3-10个回合组成,通过具有包括8个有害类别之一的内容的单个对话回合(例如,暴力)和5个等级之一,从非常轻微到严重。在160次对话中,5名人类评分员确认了可靠的不安全检测和有意义的严重程度量表。我们对三个开源LALM进行基准测试:Qwen 2-Audio,Audio Flamingo 3和MERaLiON作为zero-shot判断,其在仅音频,仅转录或多模态输入中输出[0,1]中的标量安全分数,以及仅转录的LLaMA基线。我们测量了法官的敏感性检测不安全的内容,在订购严重程度的特异性,并在对话轮得分的稳定性。结果显示架构和模态相关的权衡:最敏感的判断也是最不稳定的跨回合,而稳定的配置牺牲检测轻度有害内容。转录质量是一个关键瓶颈:Whisper-Large可能会显著降低仅转录模式的敏感性,同时在很大程度上保留严重性排序。当非语言线索或转录保真度是类别关键时,音频变得至关重要。我们总结了所有的调查结果,并为从业者提供可操作的指导。
摘要:Spoken dialogues with and between voice agents are becoming increasingly common, yet assessing them for their socially harmful content such as violence, harassment, and hate remains text-centric and fails to account for audio-specific cues and transcription errors. We present LALM-as-a-Judge, the first controlled benchmark and systematic study of large audio-language models (LALMs) as safety judges for multi-turn spoken dialogues. We generate 24,000 unsafe and synthetic spoken dialogues in English that consist of 3-10 turns, by having a single dialogue turn including content with one of 8 harmful categories (e.g., violence) and on one of 5 grades, from very mild to severe. On 160 dialogues, 5 human raters confirmed reliable unsafe detection and a meaningful severity scale. We benchmark three open-source LALMs: Qwen2-Audio, Audio Flamingo 3, and MERaLiON as zero-shot judges that output a scalar safety score in [0,1] across audio-only, transcription-only, or multimodal inputs, along with a transcription-only LLaMA baseline. We measure the judges' sensitivity to detecting unsafe content, the specificity in ordering severity levels, and the stability of the score in dialogue turns. Results reveal architecture- and modality-dependent trade-offs: the most sensitive judge is also the least stable across turns, while stable configurations sacrifice detection of mild harmful content. Transcription quality is a key bottleneck: Whisper-Large may significantly reduce sensitivity for transcription-only modes, while largely preserving severity ordering. Audio becomes crucial when paralinguistic cues or transcription fidelity are category-critical. We summarize all findings and provide actionable guidance for practitioners.
【13】Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
标题:用于跨域语音识别和增强的通用鲁棒语音自适应
链接:https://arxiv.org/abs/2602.04307
备注:Accepted to IEEE Transactions on Audio, Speech and Language Processing (IEEE TASLP)
摘要:用于自动语音识别(ASR)和语音增强(SE)的预训练模型在匹配的噪声和信道条件下表现出了显着的能力。然而,这些模型往往遭受严重的性能下降时,面对域偏移,特别是在存在看不见的噪声和信道失真。有鉴于此,我们在本文中提出了URSA-GAN,这是一个统一的、领域感知的生成框架,专门用于减轻噪声和信道条件下的失配。URSA-GAN利用双嵌入架构,该架构由噪声编码器和信道编码器组成,每个编码器都使用有限的域内数据进行预训练,以捕获域相关的表示。这些嵌入条件基于GAN的语音生成器,促进语音的合成,该语音在声学上与目标域对齐,同时保留语音内容。为了进一步提高泛化能力,我们提出了动态随机扰动,这是一种新的正则化技术,它在生成过程中将受控的可变性引入到嵌入中,从而提高了对未知域的鲁棒性。实验结果表明,URSA-GAN有效地降低了字符错误率在ASR和提高感知指标在SE在不同的噪声和不匹配的信道场景。值得注意的是,在信道和噪声退化的复合测试条件下的评估证实了URSA-GAN的泛化能力,ASR性能相对提高了16.16%,SE指标相对提高了15.58%。
摘要:Pre-trained models for automatic speech recognition (ASR) and speech enhancement (SE) have exhibited remarkable capabilities under matched noise and channel conditions. However, these models often suffer from severe performance degradation when confronted with domain shifts, particularly in the presence of unseen noise and channel distortions. In view of this, we in this paper present URSA-GAN, a unified and domain-aware generative framework specifically designed to mitigate mismatches in both noise and channel conditions. URSA-GAN leverages a dual-embedding architecture that consists of a noise encoder and a channel encoder, each pre-trained with limited in-domain data to capture domain-relevant representations. These embeddings condition a GAN-based speech generator, facilitating the synthesis of speech that is acoustically aligned with the target domain while preserving phonetic content. To enhance generalization further, we propose dynamic stochastic perturbation, a novel regularization technique that introduces controlled variability into the embeddings during generation, promoting robustness to unseen domains. Empirical results demonstrate that URSA-GAN effectively reduces character error rates in ASR and improves perceptual metrics in SE across diverse noisy and mismatched channel scenarios. Notably, evaluations on compound test conditions with both channel and noise degradations confirm the generalization ability of URSA-GAN, yielding relative improvements of 16.16% in ASR performance and 15.58% in SE metrics.
【14】Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection
标题:声音亮点:用于视听视频亮点检测的双路径音频编码器
链接:https://arxiv.org/abs/2602.03891
备注:5 pages, 2 figures, to appear in ICASSP 2026
摘要:视听视频亮点检测旨在通过利用视觉和听觉线索自动识别视频中最显著的时刻。然而,现有的模型往往没有充分利用音频模态,侧重于高层次的语义特征,而未能充分利用丰富的,动态的声音特性。为了解决这个问题,我们提出了一个新的框架,双通道音频编码器的视频亮点检测(DAViHD)。双通道音频编码器由用于内容理解的语义通道和捕获频谱-时间动态的动态通道组成。语义路径通过识别音频中的内容(例如语音、音乐或特定声音事件)来提取高级信息。动态路径采用频率自适应机制,随着时间的推移,以共同模拟这些动态,使其能够识别瞬态声事件通过显着的频谱带和快速的能量变化。我们将新颖的音频编码器集成到一个完整的视听框架中,并在大规模Mr.HiSum基准测试中实现了新的最先进的性能。我们的研究结果表明,一个复杂的,双方面的音频表示是关键,推进重点检测领域。
摘要:Audio-visual video highlight detection aims to automatically identify the most salient moments in videos by leveraging both visual and auditory cues. However, existing models often underutilize the audio modality, focusing on high-level semantic features while failing to fully leverage the rich, dynamic characteristics of sound. To address this limitation, we propose a novel framework, Dual-Pathway Audio Encoders for Video Highlight Detection (DAViHD). The dual-pathway audio encoder is composed of a semantic pathway for content understanding and a dynamic pathway that captures spectro-temporal dynamics. The semantic pathway extracts high-level information by identifying the content within the audio, such as speech, music, or specific sound events. The dynamic pathway employs a frequency-adaptive mechanism as time evolves to jointly model these dynamics, enabling it to identify transient acoustic events via salient spectral bands and rapid energy changes. We integrate the novel audio encoder into a full audio-visual framework and achieve new state-of-the-art performance on the large-scale Mr.HiSum benchmark. Our results demonstrate that a sophisticated, dual-faceted audio representation is key to advancing the field of highlight detection.
【15】Benchmarking Automatic Speech Recognition for Indian Languages in Agricultural Contexts
标题:农业背景下印度语言自动语音识别基准
链接:https://arxiv.org/abs/2602.03868
备注:9 pages, 6 figures
摘要:印度农业咨询服务的数字化需要强大的自动语音识别(ASR)系统,能够准确地将特定领域的术语翻译成多种印度语言。本文提出了一个基准框架评估ASR性能在农业背景下,在印地语,泰卢固语和Odia语言。我们引入评估指标,包括农业加权字错误率(AWWER)和特定领域的效用评分,以补充传统的指标。我们评估了10,934个音频记录,每个记录由多达10个ASR模型转录,揭示了不同语言和模型的性能差异,印地语的整体性能最好(WER:16.2%),而Odia的挑战最大(最佳WER:35.1%,仅通过扬声器日记实现)。我们描述了现实世界农业现场录音所固有的音频质量挑战,并证明了使用最佳扬声器选择的扬声器日记可以大幅降低多扬声器录音的WER(取决于多扬声器音频的比例,高达66%)。我们确定了农业术语中经常出现的错误模式,并为改善低资源农业领域的ASR系统提供了实用的建议。该研究为未来农业ASR的发展建立了基线基准。
摘要:The digitization of agricultural advisory services in India requires robust Automatic Speech Recognition (ASR) systems capable of accurately transcribing domain-specific terminology in multiple Indian languages. This paper presents a benchmarking framework for evaluating ASR performance in agricultural contexts across Hindi, Telugu, and Odia languages. We introduce evaluation metrics including Agriculture Weighted Word Error Rate (AWWER) and domain-specific utility scoring to complement traditional metrics. Our evaluation of 10,934 audio recordings, each transcribed by up to 10 ASR models, reveals performance variations across languages and models, with Hindi achieving the best overall performance (WER: 16.2%) while Odia presents the greatest challenges (best WER: 35.1%, achieved only with speaker diarization). We characterize audio quality challenges inherent to real-world agricultural field recordings and demonstrate that speaker diarization with best-speaker selection can substantially reduce WER for multi-speaker recordings (upto 66% depending on the proportion of multi-speaker audio). We identify recurring error patterns in agricultural terminology and provide practical recommendations for improving ASR systems in low-resource agricultural domains. The study establishes baseline benchmarks for future agricultural ASR development.
【1】LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
标题:LALM作为法官:在多轮口语对话中对大型音频语言模型进行安全评估的基准
链接:https://arxiv.org/abs/2602.04796
摘要:语音代理之间的对话越来越普遍,但评估它们的社会有害内容,如暴力,骚扰和仇恨仍然以文本为中心,未能考虑音频特定的线索和转录错误。我们提出了LALM作为一个法官,第一个控制基准和系统的研究大型音频语言模型(LALM)作为安全法官的多轮口语对话。我们生成了24,000个不安全和合成的英语口语对话,由3-10个回合组成,通过具有包括8个有害类别之一的内容的单个对话回合(例如,暴力)和5个等级之一,从非常轻微到严重。在160次对话中,5名人类评分员确认了可靠的不安全检测和有意义的严重程度量表。我们对三个开源LALM进行基准测试:Qwen 2-Audio,Audio Flamingo 3和MERaLiON作为zero-shot判断,其在仅音频,仅转录或多模态输入中输出[0,1]中的标量安全分数,以及仅转录的LLaMA基线。我们测量了法官的敏感性检测不安全的内容,在订购严重程度的特异性,并在对话轮得分的稳定性。结果显示架构和模态相关的权衡:最敏感的判断也是最不稳定的跨回合,而稳定的配置牺牲检测轻度有害内容。转录质量是一个关键瓶颈:Whisper-Large可能会显著降低仅转录模式的敏感性,同时在很大程度上保留严重性排序。当非语言线索或转录保真度是类别关键时,音频变得至关重要。我们总结了所有的调查结果,并为从业者提供可操作的指导。
摘要:Spoken dialogues with and between voice agents are becoming increasingly common, yet assessing them for their socially harmful content such as violence, harassment, and hate remains text-centric and fails to account for audio-specific cues and transcription errors. We present LALM-as-a-Judge, the first controlled benchmark and systematic study of large audio-language models (LALMs) as safety judges for multi-turn spoken dialogues. We generate 24,000 unsafe and synthetic spoken dialogues in English that consist of 3-10 turns, by having a single dialogue turn including content with one of 8 harmful categories (e.g., violence) and on one of 5 grades, from very mild to severe. On 160 dialogues, 5 human raters confirmed reliable unsafe detection and a meaningful severity scale. We benchmark three open-source LALMs: Qwen2-Audio, Audio Flamingo 3, and MERaLiON as zero-shot judges that output a scalar safety score in [0,1] across audio-only, transcription-only, or multimodal inputs, along with a transcription-only LLaMA baseline. We measure the judges' sensitivity to detecting unsafe content, the specificity in ordering severity levels, and the stability of the score in dialogue turns. Results reveal architecture- and modality-dependent trade-offs: the most sensitive judge is also the least stable across turns, while stable configurations sacrifice detection of mild harmful content. Transcription quality is a key bottleneck: Whisper-Large may significantly reduce sensitivity for transcription-only modes, while largely preserving severity ordering. Audio becomes crucial when paralinguistic cues or transcription fidelity are category-critical. We summarize all findings and provide actionable guidance for practitioners.
【2】Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
标题:用于跨域语音识别和增强的通用鲁棒语音自适应
链接:https://arxiv.org/abs/2602.04307
备注:Accepted to IEEE Transactions on Audio, Speech and Language Processing (IEEE TASLP)
摘要:用于自动语音识别(ASR)和语音增强(SE)的预训练模型在匹配噪声和信道条件下表现出显著的能力。然而,这些模型往往遭受严重的性能下降时,面对域偏移,特别是在存在看不见的噪声和信道失真。有鉴于此,我们在本文中提出了URSA-GAN,这是一个统一的、领域感知的生成框架,专门用于减轻噪声和信道条件下的失配。URSA-GAN利用双嵌入架构,该架构由噪声编码器和信道编码器组成,每个编码器都使用有限的域内数据进行预训练,以捕获域相关的表示。这些嵌入条件基于GAN的语音生成器,促进语音的合成,该语音在声学上与目标域对齐,同时保留语音内容。为了进一步提高泛化能力,我们提出了动态随机扰动,这是一种新的正则化技术,它在生成过程中将受控的可变性引入到嵌入中,从而提高了对未知域的鲁棒性。实验结果表明,URSA-GAN有效地降低了字符错误率在ASR和提高感知指标在SE在不同的噪声和不匹配的信道场景。值得注意的是,在信道和噪声退化的复合测试条件下的评估证实了URSA-GAN的泛化能力,ASR性能相对提高了16.16%,SE指标相对提高了15.58%。
摘要:Pre-trained models for automatic speech recognition (ASR) and speech enhancement (SE) have exhibited remarkable capabilities under matched noise and channel conditions. However, these models often suffer from severe performance degradation when confronted with domain shifts, particularly in the presence of unseen noise and channel distortions. In view of this, we in this paper present URSA-GAN, a unified and domain-aware generative framework specifically designed to mitigate mismatches in both noise and channel conditions. URSA-GAN leverages a dual-embedding architecture that consists of a noise encoder and a channel encoder, each pre-trained with limited in-domain data to capture domain-relevant representations. These embeddings condition a GAN-based speech generator, facilitating the synthesis of speech that is acoustically aligned with the target domain while preserving phonetic content. To enhance generalization further, we propose dynamic stochastic perturbation, a novel regularization technique that introduces controlled variability into the embeddings during generation, promoting robustness to unseen domains. Empirical results demonstrate that URSA-GAN effectively reduces character error rates in ASR and improves perceptual metrics in SE across diverse noisy and mismatched channel scenarios. Notably, evaluations on compound test conditions with both channel and noise degradations confirm the generalization ability of URSA-GAN, yielding relative improvements of 16.16% in ASR performance and 15.58% in SE metrics.
【3】Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection
标题:声音亮点:用于视听视频亮点检测的双路径音频编码器
链接:https://arxiv.org/abs/2602.03891
备注:5 pages, 2 figures, to appear in ICASSP 2026
摘要:视听视频亮点检测旨在通过利用视觉和听觉线索自动识别视频中最显著的时刻。然而,现有的模型往往没有充分利用音频模态,侧重于高层次的语义特征,而未能充分利用丰富的,动态的声音特性。为了解决这个问题,我们提出了一个新的框架,双通道音频编码器的视频亮点检测(DAViHD)。双通道音频编码器由用于内容理解的语义通道和捕获频谱-时间动态的动态通道组成。语义路径通过识别音频中的内容(例如语音、音乐或特定声音事件)来提取高级信息。动态路径采用频率自适应机制,随着时间的推移,以共同模拟这些动态,使其能够识别瞬态声事件通过显着的频谱带和快速的能量变化。我们将新颖的音频编码器集成到一个完整的视听框架中,并在大规模Mr.HiSum基准测试中实现了新的最先进的性能。我们的研究结果表明,一个复杂的,双方面的音频表示是关键,推进重点检测领域。
摘要:Audio-visual video highlight detection aims to automatically identify the most salient moments in videos by leveraging both visual and auditory cues. However, existing models often underutilize the audio modality, focusing on high-level semantic features while failing to fully leverage the rich, dynamic characteristics of sound. To address this limitation, we propose a novel framework, Dual-Pathway Audio Encoders for Video Highlight Detection (DAViHD). The dual-pathway audio encoder is composed of a semantic pathway for content understanding and a dynamic pathway that captures spectro-temporal dynamics. The semantic pathway extracts high-level information by identifying the content within the audio, such as speech, music, or specific sound events. The dynamic pathway employs a frequency-adaptive mechanism as time evolves to jointly model these dynamics, enabling it to identify transient acoustic events via salient spectral bands and rapid energy changes. We integrate the novel audio encoder into a full audio-visual framework and achieve new state-of-the-art performance on the large-scale Mr.HiSum benchmark. Our results demonstrate that a sophisticated, dual-faceted audio representation is key to advancing the field of highlight detection.
【4】Benchmarking Automatic Speech Recognition for Indian Languages in Agricultural Contexts
标题:农业背景下印度语言自动语音识别基准
链接:https://arxiv.org/abs/2602.03868
备注:9 pages, 6 figures
摘要:印度农业咨询服务的数字化需要强大的自动语音识别(ASR)系统,能够准确地将特定领域的术语翻译成多种印度语言。本文提出了一个基准框架评估ASR性能在农业背景下,在印地语,泰卢固语和Odia语言。我们引入评估指标,包括农业加权字错误率(AWWER)和特定领域的效用评分,以补充传统的指标。我们评估了10,934个音频记录,每个记录由多达10个ASR模型转录,揭示了不同语言和模型的性能差异,印地语的整体性能最好(WER:16.2%),而Odia的挑战最大(最佳WER:35.1%,仅通过扬声器日记实现)。我们描述了现实世界农业现场录音所固有的音频质量挑战,并证明了使用最佳扬声器选择的扬声器日记可以大幅降低多扬声器录音的WER(取决于多扬声器音频的比例,高达66%)。我们确定了农业术语中经常出现的错误模式,并为改善低资源农业领域的ASR系统提供了实用的建议。该研究为未来农业ASR的发展建立了基线基准。
摘要:The digitization of agricultural advisory services in India requires robust Automatic Speech Recognition (ASR) systems capable of accurately transcribing domain-specific terminology in multiple Indian languages. This paper presents a benchmarking framework for evaluating ASR performance in agricultural contexts across Hindi, Telugu, and Odia languages. We introduce evaluation metrics including Agriculture Weighted Word Error Rate (AWWER) and domain-specific utility scoring to complement traditional metrics. Our evaluation of 10,934 audio recordings, each transcribed by up to 10 ASR models, reveals performance variations across languages and models, with Hindi achieving the best overall performance (WER: 16.2%) while Odia presents the greatest challenges (best WER: 35.1%, achieved only with speaker diarization). We characterize audio quality challenges inherent to real-world agricultural field recordings and demonstrate that speaker diarization with best-speaker selection can substantially reduce WER for multi-speaker recordings (upto 66% depending on the proportion of multi-speaker audio). We identify recurring error patterns in agricultural terminology and provide practical recommendations for improving ASR systems in low-resource agricultural domains. The study establishes baseline benchmarks for future agricultural ASR development.
【5】Speaker-Aware Simulation Improves Conversational Speech Recognition
标题:说话者感知模拟改进对话语音识别
链接:https://arxiv.org/abs/2602.04776
摘要:由于大规模、注释良好的多说话人对话数据的有限可用性以及自然交互的复杂时间动态性,会话语音的自动语音识别(ASR)仍然具有挑战性。说话人感知的模拟对话(SASC)提供了一种有效的数据增强策略,通过将单说话人录音转换为真实的多说话人对话。然而,之前的工作主要集中在英语数据上,留下了对低资源语言的适用性的问题。在本文中,我们适应和实现匈牙利会话ASR的SASC框架。我们进一步提出了C-SASC,一个扩展的变体,它结合了基于话语持续时间的停顿建模,使人类对话中观察到的局部时间依赖性的更忠实的表示,同时保留了原始方法的简单性和效率。我们从BEA-Large语料库中生成合成的匈牙利语对话,并将它们与真实的会话数据相结合进行ASR训练。SASC和C-SASC广泛的模拟配置下进行了评估,使用来自CallHome,BEA对话,和GRASS语料库的会话统计。实验结果表明,说话人感知的会话模拟一致地提高了识别性能,天真的级联为基础的增强。虽然C-SASC中的额外持续时间调节产生适度但系统的增益-最明显的是在字符级错误率方面-其有效性取决于源会话统计数据和目标域之间的匹配。总体而言,我们的研究结果证实了匈牙利ASR的说话者感知会话模拟的鲁棒性,并强调了合成对话生成中越来越详细的时间建模的优点和局限性。
摘要:Automatic speech recognition (ASR) for conversational speech remains challenging due to the limited availability of large-scale, well-annotated multi-speaker dialogue data and the complex temporal dynamics of natural interactions. Speaker-aware simulated conversations (SASC) offer an effective data augmentation strategy by transforming single-speaker recordings into realistic multi-speaker dialogues. However, prior work has primarily focused on English data, leaving questions about the applicability to lower-resource languages. In this paper, we adapt and implement the SASC framework for Hungarian conversational ASR. We further propose C-SASC, an extended variant that incorporates pause modeling conditioned on utterance duration, enabling a more faithful representation of local temporal dependencies observed in human conversation while retaining the simplicity and efficiency of the original approach. We generate synthetic Hungarian dialogues from the BEA-Large corpus and combine them with real conversational data for ASR training. Both SASC and C-SASC are evaluated extensively under a wide range of simulation configurations, using conversational statistics derived from CallHome, BEA-Dialogue, and GRASS corpora. Experimental results show that speaker-aware conversational simulation consistently improves recognition performance over naive concatenation-based augmentation. While the additional duration conditioning in C-SASC yields modest but systematic gains--most notably in character-level error rates--its effectiveness depends on the match between source conversational statistics and the target domain. Overall, our findings confirm the robustness of speaker-aware conversational simulation for Hungarian ASR and highlight the benefits and limitations of increasingly detailed temporal modeling in synthetic dialogue generation.
【6】Frontend Token Enhancement for Token-Based Speech Recognition
标题:基于令牌的语音识别的前端令牌增强
链接:https://arxiv.org/abs/2602.04217
备注:Accepted at ICASSP 2026
摘要:语音信号的离散化表示是各种语音应用中连续特征的有效替代,包括自动语音识别(ASR)和语音语言模型。然而,这些表示,例如从自监督学习(SSL)语音模型的聚类输出导出的语义或语音标记,容易受到环境噪声的影响,这会降低后端任务性能。在这项工作中,我们介绍了一个前端系统,估计干净的语音令牌从嘈杂的语音和评估它的ASR后端使用语义令牌。我们考虑四种类型的增强模型的基础上,他们的输入/输出域:波到波,令牌令牌,连续SSL功能令牌,和波令牌。这些模型独立于ASR后端进行训练。在CHiME-4数据集上的实验表明,波到令牌增强在前端中实现了最佳性能。此外,它的性能大多优于基于连续SSL特征的ASR系统。
摘要:Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and speech language models. However, these representations, such as semantic or phonetic tokens derived from clustering outputs of self-supervised learning (SSL) speech models, are susceptible to environmental noise, which can degrade backend task performance. In this work, we introduce a frontend system that estimates clean speech tokens from noisy speech and evaluate it on an ASR backend using semantic tokens. We consider four types of enhancement models based on their input/output domains: wave-to-wave, token-to-token, continuous SSL features-to-token, and wave-to-token. These models are trained independently of ASR backends. Experiments on the CHiME-4 dataset demonstrate that wave-to-token enhancement achieves the best performance among the frontends. Moreover, it mostly outperforms the ASR system based on continuous SSL features.
【7】Audit After Segmentation: Reference-Free Mask Quality Assessment for Language-Referred Audio-Visual Segmentation
标题:分割后的审核:参考视听分割的无参考面罩质量评估
链接:https://arxiv.org/abs/2602.03892
摘要:参考音视频分割(Ref-AVS)的目的是通过对视频、音频和文本的联合推理来分割由自然语言描述的目标对象。除了生成分割掩模之外,提供掩模质量的丰富和可解释的诊断在很大程度上仍然未被探索。在这项工作中,我们引入了Ref-AVS上下文中的掩模质量评估(MQA-RefAVS),这是一项新任务,可以评估候选分割掩模的质量,而不依赖地面实况注释作为推理时的参考。给定视听语言输入和每个提供的分割掩码,该任务需要使用未观察到的地面事实估计其IoU,识别相应的错误类型,并推荐可操作的质量控制决策。为了支持这一任务,我们构建MQ-RAVSBench,一个基准测试,具有不同的和代表性的掩模错误模式,跨越几何和语义问题。我们进一步提出了MQ-Auditor,一个基于多模态大语言模型(MLLM)的审计师,明确的原因多模态线索和掩码信息,以产生定量和定性的掩码质量评估。大量的实验表明,MQ-Auditor优于强大的开源和商业MLLM,可以与现有的Ref-AVS系统集成,以检测分割失败,并支持下游分割改进。数据和代码将在https://github.com/jasongief/MQA-RefAVS上发布。
摘要:Language-referred audio-visual segmentation (Ref-AVS) aims to segment target objects described by natural language by jointly reasoning over video, audio, and text. Beyond generating segmentation masks, providing rich and interpretable diagnoses of mask quality remains largely underexplored. In this work, we introduce Mask Quality Assessment in the Ref-AVS context (MQA-RefAVS), a new task that evaluates the quality of candidate segmentation masks without relying on ground-truth annotations as references at inference time. Given audio-visual-language inputs and each provided segmentation mask, the task requires estimating its IoU with the unobserved ground truth, identifying the corresponding error type, and recommending an actionable quality-control decision. To support this task, we construct MQ-RAVSBench, a benchmark featuring diverse and representative mask error modes that span both geometric and semantic issues. We further propose MQ-Auditor, a multimodal large language model (MLLM)-based auditor that explicitly reasons over multimodal cues and mask information to produce quantitative and qualitative mask quality assessments. Extensive experiments demonstrate that MQ-Auditor outperforms strong open-source and commercial MLLMs and can be integrated with existing Ref-AVS systems to detect segmentation failures and support downstream segmentation improvement. Data and codes will be released at https://github.com/jasongief/MQA-RefAVS.
【8】Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models
标题:通过音频语言模型中的测试时间缩放解码模糊情绪
链接:https://arxiv.org/abs/2602.03873
摘要:从人类语音中识别情感是社交意识会话AI的关键推动因素。然而,虽然大多数先前的工作框架的情感识别作为一个分类问题,现实世界的情感状态往往是模糊的,重叠的,和上下文相关的,提出了重大挑战的注释和自动建模。最近的大规模音频语言模型(ALMs)提供了新的机会,细致入微的情感推理没有明确的情感监督,但他们的能力,以处理模棱两可的情绪仍然未充分发掘。与此同时,推理时间技术的进步,如测试时间缩放(TTS),已经显示出改善硬NLP任务的泛化和适应性的希望,但它们与情感计算的相关性在很大程度上仍然未知。在这项工作中,我们介绍了第一个基准模糊的情感识别语音与ALMs下测试时间缩放。我们的评估系统地比较了八个国家的最先进的ALM和五个TTS策略在三个突出的语音情感数据集。我们进一步深入分析了模型容量、TTS和情感模糊之间的相互作用,为模糊情感理解的计算和表征挑战提供了新的见解。我们的基准为开发更强大,上下文感知和情感智能的基于语音的AI系统奠定了基础,并强调了弥合模型假设与现实世界人类情感复杂性之间差距的关键未来方向。
摘要:Emotion recognition from human speech is a critical enabler for socially aware conversational AI. However, while most prior work frames emotion recognition as a categorical classification problem, real-world affective states are often ambiguous, overlapping, and context-dependent, posing significant challenges for both annotation and automatic modeling. Recent large-scale audio language models (ALMs) offer new opportunities for nuanced affective reasoning without explicit emotion supervision, but their capacity to handle ambiguous emotions remains underexplored. At the same time, advances in inference-time techniques such as test-time scaling (TTS) have shown promise for improving generalization and adaptability in hard NLP tasks, but their relevance to affective computing is still largely unknown. In this work, we introduce the first benchmark for ambiguous emotion recognition in speech with ALMs under test-time scaling. Our evaluation systematically compares eight state-of-the-art ALMs and five TTS strategies across three prominent speech emotion datasets. We further provide an in-depth analysis of the interaction between model capacity, TTS, and affective ambiguity, offering new insights into the computational and representational challenges of ambiguous emotion understanding. Our benchmark establishes a foundation for developing more robust, context-aware, and emotionally intelligent speech-based AI systems, and highlights key future directions for bridging the gap between model assumptions and the complexity of real-world human emotion.
机器翻译由腾讯交互翻译提供,仅供参考
