今日论文合集:cs.SD语音6篇,eess.AS音频处理5篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】CLAR: CIF-Localized Alignment for Retrieval-Augmented Speech LLM-Based Contextual ASR
标题:CLAR:基于检索增强语音LLM的上下文ASB的inf本地化对齐
链接:https://arxiv.org/abs/2603.25460

作者:Shangkun Huang, Huan Shen, Wei Zou, Yunzhang Chen
备注:Submitted to Interspeech 2026
摘要:
摘要:


【2】CoDeTT: A Context-Aware Decision Benchmark for Turn-Taking Evaluation
标题:CoDeTT:回合转换评估的上下文感知决策基准
链接:https://arxiv.org/abs/2603.25434

作者:Huan Shen, Yingao Wang, Shangkun Huang, Wei Zou, Yunzhang Chen
备注:Submitted to Interspeech 2026
摘要:
摘要:


【3】Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition
标题:联合学习全球-本地说话人分类以增强端到端说话人拨号和识别
链接:https://arxiv.org/abs/2603.25377

作者:Yuhang Dai, Haopeng Lin, Jiale Qian, Ruiqi Yan, Hao Meng, Hanke Xie, Hanlin Wen, Shunshun Yin, Ming Tao, Xie Chen, Lei Xie, Xinsheng Wang
备注:5 pages, 2 figures, 2 tables
摘要:
摘要:


【4】SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
标题:SAVe:利用视觉伪影和视听错位的自我监督视听Deepfake检测
链接:https://arxiv.org/abs/2603.25140

作者:Sahibzada Adil Shahzad, Ammarah Hashmi, Junichi Yamagishi, Yusuke Yasuda, Yu Tsao, Chia-Wen Lin, Yan-Tsung Peng, Hsin-Min Wang
摘要:
摘要:


【5】AVControl: Efficient Framework for Training Audio-Visual Controls
标题:AVControl:用于训练视听控件的高效框架
链接:https://arxiv.org/abs/2603.24793

作者:Matan Ben-Yosef, Tavi Halperin, Naomi Ken Korem, Mohammad Salama, Harel Cain, Asaf Joseph, Anthony Chen, Urska Jelercic, Ofir Bibi
备注:Project page: this https URL
摘要:
摘要:


【6】When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews
标题:当一致性成为偏见时:半结构化临床面试中的面试官效应
链接:https://arxiv.org/abs/2603.24651

作者:Hasindri Watawana, Sergio Burdisso, Diego A. Moreno-Galván, Fernando Sánchez-Vega, A. Pastor López-Monroy, Petr Motlicek, Esaú Villatoro-Tello
备注:Accepted to LREC 2026 Conference
摘要:
摘要:


eess.AS音频处理


【1】AdaLTM: Adaptive Layer-wise Task Vector Merging for Categorical Speech Emotion Recognition with ASR Knowledge Integration
标题:AdaTLR:自适应分层任务载体合并,用于分类语音情感识别和ASB知识集成
链接:https://arxiv.org/abs/2603.25041

作者:Chia-Yu Lee,Huang-Cheng Chou,Tzu-Quan Lin,Yuanchao Li,Ya-Tse Wu,Shrikanth Narayanan,Chi-Chun Lee
备注:Submitted to Interspeech 2026
摘要:将自动语音识别(ASR)集成到语音情感识别(SER)中,通过提供语言上下文来增强建模。然而,传统的特征融合面临着性能瓶颈,多任务学习往往遭受优化冲突。虽然任务向量和模型合并已经解决了NLP和CV中的这种冲突,但它们在语音任务中的潜力在很大程度上尚未开发。在这项工作中,我们提出了一个自适应分层任务向量合并(AdaLTM)框架的基础上WavLM-大。而不是联合优化,我们提取任务向量域ASR和SER模型微调情绪数据集。这些向量被集成到一个冻结的基础模型使用逐层学习系数。该策略实现了跨Transformer层的语言和非语言知识的深度感知平衡,而没有梯度干扰。在MSP-Podcast上的实验表明,该方法有效地缓解了ASR和SER之间的冲突。
摘要:Integrating Automatic Speech Recognition (ASR) into Speech Emotion Recognition (SER) enhances modeling by providing linguistic context. However, conventional feature fusion faces performance bottlenecks, and multi-task learning often suffers from optimization conflicts. While task vectors and model merging have addressed such conflicts in NLP and CV, their potential in speech tasks remains largely unexplored. In this work, we propose an Adaptive Layer-wise Task Vector Merging (AdaLTM) framework based on WavLM-Large. Instead of joint optimization, we extract task vectors from in-domain ASR and SER models fine-tuned on emotion datasets. These vectors are integrated into a frozen base model using layer-wise learnable coefficients. This strategy enables depth-aware balancing of linguistic and paralinguistic knowledge across transformer layers without gradient interference. Experiments on the MSP-Podcast demonstrate that the proposed approach effectively mitigates conflicts between ASR and SER.


【2】Unified Diffusion Refinement for Multi-Channel Speech Enhancement and Separation
标题:多通道语音增强和分离的统一扩散细化
链接:https://arxiv.org/abs/2603.24810

作者:Zhongweiyang Xu,Ashutosh Pandey,Juan Azcarreta,Zhaoheng Ni,Sanjeel Parekh,Buye Xu,Romit Roy Choudhury
备注:Paper in submission
摘要:我们提出了Uni-ArrayDPS,一种新的基于扩散的细化框架,用于统一的多通道语音增强和分离。用于多通道语音增强/分离的现有方法大多是有区别的,并且在产生高SNR输出方面非常有效。然而,它们仍然可以生成不自然的语音,其中包含由神经网络和基于回归的目标引起的非线性失真。为了解决这个问题,我们提出了Uni-ArrayDPS,它使用语音扩散先验来细化任何强判别模型的输出。Uni-ArrayDPS是生成的,阵列不可知的,无需训练,并支持增强和分离。给定一个判别模型的增强/分离的语音,我们使用它,连同嘈杂的混合物,估计噪声空间协方差矩阵(SCM)。然后,我们使用此SCM来计算干净语音源的扩散后验采样所需的似然性。Uni-ArrayDPS只需要一个预先训练的干净语音扩散模型作为先验,不需要额外的训练或微调,允许它直接在任务(增强/分离),麦克风阵列几何形状和判别模型骨干中推广。大量的实验表明,Uni-ArrayDPS一致提高了广泛的判别模型的增强和分离任务。我们还报告了真实世界数据集的强有力结果。音频演示在\href{https://xzwy.github.io/Uni-ArrayDPS/}{https://xzwy.github.io/Uni-ArrayDPS/}提供。
摘要:We propose Uni-ArrayDPS, a novel diffusion-based refinement framework for unified multi-channel speech enhancement and separation. Existing methods for multi-channel speech enhancement/separation are mostly discriminative and are highly effective at producing high-SNR outputs. However, they can still generate unnatural speech with non-linear distortions caused by the neural network and regression-based objectives. To address this issue, we propose Uni-ArrayDPS, which refines the outputs of any strong discriminative model using a speech diffusion prior. Uni-ArrayDPS is generative, array-agnostic, and training-free, and supports both enhancement and separation. Given a discriminative model's enhanced/separated speech, we use it, together with the noisy mixtures, to estimate the noise spatial covariance matrix (SCM). We then use this SCM to compute the likelihood required for diffusion posterior sampling of the clean speech source(s). Uni-ArrayDPS requires only a pre-trained clean-speech diffusion model as a prior and does not require additional training or fine-tuning, allowing it to generalize directly across tasks (enhancement/separation), microphone array geometries, and discriminative model backbones. Extensive experiments show that Uni-ArrayDPS consistently improves a wide range of discriminative models for both enhancement and separation tasks. We also report strong results on a real-world dataset. Audio demos are provided at \href{https://xzwy.github.io/Uni-ArrayDPS/}{https://xzwy.github.io/Uni-ArrayDPS/}.


【3】X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
标题:X-OPD:语音LLM中能力一致的跨模式政策提炼
链接:https://arxiv.org/abs/2603.24596

作者:Di Cao,Dongjie Fu,Hai Yu,Siqi Zheng,Xu Tan,Tao Jin
备注:5 pages
摘要:虽然从级联对话系统到端到端(E2 E)语音大语言模型(LLM)的转变改善了延迟和语言建模,但与基于文本的模型相比,E2 E模型通常表现出显着的性能下降。标准的监督微调(SFT)和强化学习(RL)训练方法无法弥补这一差距。为了解决这个问题,我们提出了X-OPD,这是一种新的跨模态策略蒸馏框架,旨在系统地将语音LLM的功能与基于文本的对应功能相匹配。X-OPD使Speech LLM能够通过政策推出探索自己的分布,其中基于文本的教师模型评估这些轨迹并提供令牌级反馈,有效地将教师的能力提取到学生的多模态表示中。跨多个基准的广泛实验表明,X-OPD显着缩小了复杂任务的差距,同时保留了模型的固有功能。
摘要:While the shift from cascaded dialogue systems to end-to-end (E2E) speech Large Language Models (LLMs) improves latency and paralinguistic modeling, E2E models often exhibit a significant performance degradation compared to their text-based counterparts. The standard Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) training methods fail to close this gap. To address this, we propose X-OPD, a novel Cross-Modal On-Policy Distillation framework designed to systematically align the capabilities of Speech LLMs to their text-based counterparts. X-OPD enables the Speech LLM to explore its own distribution via on-policy rollouts, where a text-based teacher model evaluates these trajectories and provides token-level feedback, effectively distilling teacher's capabilities into student's multi-modal representations. Extensive experiments across multiple benchmarks demonstrate that X-OPD significantly narrows the gap in complex tasks while preserving the model's inherent capabilities.


【4】When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews
标题:当一致性成为偏见时:半结构化临床面试中的面试官效应
链接:https://arxiv.org/abs/2603.24651

作者:Hasindri Watawana,Sergio Burdisso,Diego A. Moreno-Galván,Fernando Sánchez-Vega,A. Pastor López-Monroy,Petr Motlicek,Esaú Villatoro-Tello
备注:Accepted to LREC 2026 Conference
摘要:由于公共语料库的可用性和语言建模的进步,从医患对话中自动检测抑郁症已经获得了动力。然而,可解释性仍然有限:报告的强劲表现往往没有揭示是什么推动了预测。我们分析了三个数据集:ANDROIDS,DAIC-WOZ,E-DAIC,并从半结构化面试中的面试官提示中识别出系统性偏差。在面试官回合上训练的模型利用固定的提示和位置来区分抑郁症和对照组,通常在不使用参与者语言的情况下获得很高的分类分数。将模型限制在参与者的话语中,可以更广泛地分配决策证据,并反映真正的语言线索。虽然半结构化协议确保了一致性,但包括面试官提示通过利用脚本工件来提高性能。我们的研究结果突出了跨数据集、与架构无关的偏见,并强调了需要进行分析,按时间和说话者定位决策证据,以确保模型从参与者的语言中学习。
摘要:Automatic depression detection from doctor-patient conversations has gained momentum thanks to the availability of public corpora and advances in language modeling. However, interpretability remains limited: strong performance is often reported without revealing what drives predictions. We analyze three datasets: ANDROIDS, DAIC-WOZ, E-DAIC and identify a systematic bias from interviewer prompts in semi-structured interviews. Models trained on interviewer turns exploit fixed prompts and positions to distinguish depressed from control subjects, often achieving high classification scores without using participant language. Restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. While semi-structured protocols ensure consistency, including interviewer prompts inflates performance by leveraging script artifacts. Our results highlight a cross-dataset, architecture-agnostic bias and emphasize the need for analyses that localize decision evidence by time and speaker to ensure models learn from participants' language.


【5】Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
标题:调整自我监督的语音表达用于帕金森病的跨舌发音障碍检测
链接:https://arxiv.org/abs/2603.22225

作者:Abner Hernandez,Eunjung Yeo,Kwanghee Choi,Chin-Jou Li,Zhengjun Yue,Rohan Kumar Das,Jan Rusz,Mathew Magimai Doss,Juan Rafael Orozco-Arroyave,Tomás Arias-Vergara,Andreas Maier,Elmar Nöth,David R. Mortensen,David Harwath,Paula Andrea Perez-Toro
备注:Submitted to Interspeech 2026
摘要:构音障碍语音数据的有限性使得跨语言检测成为一个重要但具有挑战性的问题。一个关键的困难是,语音表征往往编码语言依赖的结构,可以混淆构音障碍检测。我们提出了一个代表级的语言转换(LS),对齐源语言自我监督的语音表示与目标语言分布使用基于质心的向量自适应估计健康控制语音。我们评估的方法,从帕金森氏病的语音数据集在捷克语,德语和西班牙语的跨语言和多语言设置下的口头DDK录音。LS大大提高了跨语言设置中的灵敏度和F1,同时在多语言设置中产生较小但一致的增益。表征分析进一步表明,LS减少了嵌入空间中的语言身份,支持LS消除语言依赖结构的解释。
摘要:The limited availability of dysarthric speech data makes cross-lingual detection an important but challenging problem. A key difficulty is that speech representations often encode language-dependent structure that can confound dysarthria detection. We propose a representation-level language shift (LS) that aligns source-language self-supervised speech representations with the target-language distribution using centroid-based vector adaptation estimated from healthy-control speech. We evaluate the approach on oral DDK recordings from Parkinson's disease speech datasets in Czech, German, and Spanish under both cross-lingual and multilingual settings. LS substantially improves sensitivity and F1 in cross-lingual settings, while yielding smaller but consistent gains in multilingual settings. Representation analysis further shows that LS reduces language identity in the embedding space, supporting the interpretation that LS removes language-dependent structure.


机器翻译由腾讯交互翻译提供,仅供参考