今日论文合集:cs.SD语音10篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization
标题:生成合成医患对话以进行长篇音频总结
链接:https://arxiv.org/abs/2604.06138

作者:Yanis Labrak,David Grünert,Séverin Baroudi,Jiyun Chun,Pawel Cyrta,Sergio Burdisso,Ahmed Hassoon,David Liu,Adam Rothschild,Reed Van Deusen,Petr Motlicek,Andrew Perrault,Ricard Marxer,Thomas Schaaf
备注:Submitted for review at Interspeech 2026
摘要:长上下文音频推理在训练数据和评估中都是不足的。现有的基准测试针对短上下文任务,而与长上下文推理最相关的开放式生成任务对自动评估提出了众所周知的挑战。我们提出了一个合成的数据生成管道,旨在作为一个培训资源,并作为一个受控的评估环境,并将其实例化为第一次访问的医生与患者的对话与SOAP说明生成的任务。该管道有三个阶段,人物驱动的对话生成,多扬声器音频合成与重叠/暂停建模,室内声学和声音事件,以及基于LLM的参考SOAP笔记生产,完全建立在开放权重模型上。我们发布了8,800个合成对话,其中包含1.3k小时的相应音频和参考笔记。评估当前的开放权重系统,我们发现级联方法仍然大大优于端到端模型。
摘要:Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant to long-context reasoning pose well-known challenges for automatic evaluation. We propose a synthetic data generation pipeline designed to serve both as a training resource and as a controlled evaluation environment, and instantiate it for first-visit doctor-patient conversations with SOAP note generation as the task. The pipeline has three stages, persona-driven dialogue generation, multi-speaker audio synthesis with overlap/pause modeling, room acoustics, and sound events, and LLM-based reference SOAP note production, built entirely on open-weight models. We release 8,800 synthetic conversations with 1.3k hours of corresponding audio and reference notes. Evaluating current open-weight systems, we find that cascaded approaches still substantially outperform end-to-end models.


【2】Time-Domain Voice Identity Morphing (TD-VIM): A Signal-Level Approach to Morphing Attacks on Speaker Verification Systems
标题:时间域语音身份变形(TD-VIM):对说话人验证系统进行变形攻击的信号级方法
链接:https://arxiv.org/abs/2604.05683

作者:Aravinda Reddy PN,Raghavendra Ramachandra,K. Sreenivasa Rao,Pabitra Mitra,Kunal Singh
摘要:在生物识别系统中,通常的做法是将每个样本或模板与特定的个体相关联。尽管如此,最近的研究已经证明了产生能够匹配多个身份的“变形”生物特征样本的可行性。这些变形攻击已被认为是生物识别系统的潜在安全风险。然而,大多数关于变形攻击的研究都集中在图像域中的生物特征模式上,例如面部,指纹和虹膜。在这项工作中,我们介绍了时域语音身份变形(TD-VIM),一种新的方法,基于语音的生物特征变形。这种方法能够在信号水平上混合来自两个不同身份的语音特征,从而创建变形样本,这些样本对于说话人验证系统来说具有很高的脆弱性。利用多语言视听智能手机数据库,我们的研究创建了四个不同的变形信号的变形因素的基础上,并评估其有效性使用全面的脆弱性分析。为了评估TD-VIM的安全影响,我们使用广义变形攻击潜力(G-MAP)指标对我们的方法进行了基准测试,测量了两个基于深度学习的说话人验证系统(SVS)和一个商业系统Verispeak的攻击成功率。我们的研究结果表明,变形语音样本实现了高攻击成功率,在文本依赖场景中,iPhone-11和三星S8上的G-MAP值分别达到99.40%和99.74%,错误匹配率为0.1%。
摘要:In biometric systems, it is a common practice to associate each sample or template with a specific individual. Nevertheless, recent studies have demonstrated the feasibility of generating "morphed" biometric samples capable of matching multiple identities. These morph attacks have been recognized as potential security risks for biometric systems. However, most research on morph attacks has focused on biometric modalities that operate within the image domain, such as the face, fingerprints, and iris. In this work, we introduce Time-domain Voice Identity Morphing (TD-VIM), a novel approach for voice-based biometric morphing. This method enables the blending of voice characteristics from two distinct identities at the signal level, creating morphed samples that present a high vulnerability for speaker verification systems. Leveraging the Multilingual Audio-Visual Smartphone database, our study created four distinct morphed signals based on morphing factors and evaluated their effectiveness using a comprehensive vulnerability analysis. To assess the security impact of TD-VIM, we benchmarked our approach using the Generalized Morphing Attack Potential (G-MAP) metric, measuring attack success across two deep-learning-based Speaker Verification Systems (SVS) and one commercial system, Verispeak. Our findings indicate that the morphed voice samples achieved a high attack success rate, with G-MAP values reaching 99.40% on iPhone-11 and 99.74% on Samsung S8 in text-dependent scenarios, at a false match rate of 0.1%.


【3】Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck
标题:具有边界感知信息瓶颈的可控歌唱风格转换
链接:https://arxiv.org/abs/2604.05526

作者:Zhetao Hu,Yiquan Zhou,Wenyu Wang,Zhiyu Wu,Xin Gao,Jihua Zhu
备注:8 pages, 5 figures
摘要:本文介绍了S4团队提交的歌唱声音转换挑战赛2025(SVCC 2025)-一种新颖的歌唱风格转换系统,在域内设置中推进细粒度的风格转换和控制。为了解决样式泄漏,动态渲染和有限数据的高保真生成的关键挑战,我们引入了三个关键的创新:边界感知的Whisper瓶颈,池音素跨度表示以抑制残留源样式,同时保留语言内容;显式帧级技术矩阵,通过在推理期间进行有针对性的F0处理来增强,以实现稳定和独特的动态样式渲染;以及感知激励的高频带完成策略,该策略利用辅助标准48 kHz SVC模型来增强高频频谱,从而在不过度拟合的情况下克服数据稀缺。在官方的SVCC 2025主观评价中,我们的系统在所有提交的作品中实现了最佳的自然度性能,同时在说话者相似性和技术控制方面保持了竞争力,尽管与其他表现最好的系统相比,使用的额外歌唱数据明显减少。音频样本可在线获取。
摘要:This paper presents the submission of the S4 team to the Singing Voice Conversion Challenge 2025 (SVCC2025)-a novel singing style conversion system that advances fine-grained style conversion and control within in-domain settings. To address the critical challenges of style leakage, dynamic rendering, and high-fidelity generation with limited data, we introduce three key innovations: a boundary-aware Whisper bottleneck that pools phoneme-span representations to suppress residual source style while preserving linguistic content; an explicit frame-level technique matrix, enhanced by targeted F0 processing during inference, for stable and distinct dynamic style rendering; and a perceptually motivated high-frequency band completion strategy that leverages an auxiliary standard 48kHz SVC model to augment the high-frequency spectrum, thereby overcoming data scarcity without overfitting. In the official SVCC2025 subjective evaluation, our system achieves the best naturalness performance among all submissions while maintaining competitive results in speaker similarity and technique control, despite using significantly less extra singing data than other top-performing systems. Audio samples are available online.


【4】Anchored Cyclic Generation: A Novel Paradigm for Long-Sequence Symbolic Music Generation
标题:锚定循环生成:长序列符号音乐生成的新范式
链接:https://arxiv.org/abs/2604.05343

作者:Boyu Cao,Lekai Qian,Dehan Li,Haoyu Gu,Mingda Xu,Qi Liu
备注:Accepted at ACL 2026 Findings
摘要:生成具有结构一致性的长序列仍然是自回归模型跨序列生成任务的根本挑战。在符号音乐生成中,这一挑战尤其突出,因为现有方法受到自回归模型固有的严重误差积累问题的限制,导致音乐质量和结构完整性的性能较差。在本文中,我们提出了锚定循环生成(ACG)的范例,它依赖于锚的功能,从已经确定的音乐,以指导后续的一代在自回归过程中,有效地减轻自回归方法中的误差积累。基于ACG范式,我们进一步提出了分层锚定循环生成(Hi-ACG)框架,该框架采用了系统的全局到局部生成策略,并与我们专门设计的钢琴令牌(一种高效的音乐表示)高度兼容。实验结果表明,与传统的自回归模型相比,ACG范式实现了减少预测特征向量和真实语义向量之间的余弦距离平均为34.7%。在长序列符号音乐生成任务中,Hi-ACG框架在主观和客观评价方面都明显优于现有的主流方法。此外,该框架具有出色的任务泛化能力,在音乐完成等相关任务中表现出色。
摘要:Generating long sequences with structural coherence remains a fundamental challenge for autoregressive models across sequential generation tasks. In symbolic music generation, this challenge is particularly pronounced, as existing methods are constrained by the inherent severe error accumulation problem of autoregressive models, leading to poor performance in music quality and structural integrity. In this paper, we propose the Anchored Cyclic Generation (ACG) paradigm, which relies on anchor features from already identified music to guide subsequent generation during the autoregressive process, effectively mitigating error accumulation in autoregressive methods. Based on the ACG paradigm, we further propose the Hierarchical Anchored Cyclic Generation (Hi-ACG) framework, which employs a systematic global-to-local generation strategy and is highly compatible with our specifically designed piano token, an efficient musical representation. The experimental results demonstrate that compared to traditional autoregressive models, the ACG paradigm achieves reduces cosine distance by an average of 34.7% between predicted feature vectors and ground-truth semantic vectors. In long-sequence symbolic music generation tasks, the Hi-ACG framework significantly outperforms existing mainstream methods in both subjective and objective evaluations. Furthermore, the framework exhibits excellent task generalization capabilities, achieving superior performance in related tasks such as music completion.


【5】GLANCE: A Global-Local Coordination Multi-Agent Framework for Music-Grounded Non-Linear Video Editing
标题:GLAANCE:一个用于基于音乐的非线性视频编辑的全球-本地协调多代理框架
链接:https://arxiv.org/abs/2604.05076

作者:Zihao Lin,Haibo Wang,Zhiyang Xu,Siyao Dai,Huanjie Dong,Xiaohan Wang,Yolo Y. Tang,Yixin Wang,Qifan Wang,Lifu Huang
备注:14 pages, 4 figures, under review
摘要:基于音乐的mashup视频创建是视频非线性编辑的一种具有挑战性的形式,其中系统必须从大量源视频集合中组成连贯的时间轴,同时与音乐节奏,用户意图,故事完整性和长期结构约束保持一致。现有的方法通常依赖于固定的管道或简化的检索和连接范例,限制了它们适应不同提示和异构源材料的能力。在本文中,我们提出了GLANCE,一个全球本地协调多智能体框架的音乐接地非线性视频编辑。为了更好地进行编辑实践,GLANCE采用了一种双循环架构:外循环执行长期规划和任务图构建,内循环采用“观察-思考-行动-验证”流程进行分段编辑任务及其细化。为了解决子时间线组成后出现的跨部门和全球冲突,我们引入了一个专门的全球-本地协调机制,包括预防和纠正组件,其中包括一个新颖设计的上下文控制器,冲突区域分解模块,和自下而上的动态协商机制。为了支持严格的评估,我们构建了MVEBench,一个新的基准,分解编辑难度沿任务类型,提示特异性和音乐长度,并提出了一个代理作为一个法官的评估框架,可扩展的多维评估。实验结果表明,在相同的主干模型下,GLANCE的性能始终优于先前的研究基线和开源产品基线。以GPT-4 o-mini为骨干,GLANCE在两个任务设置上分别比最强基线提高了33.2%和15.6%。人工评价进一步确认了生成视频的质量,并验证了所提出的评价框架的有效性。
摘要:Music-grounded mashup video creation is a challenging form of video non-linear editing, where a system must compose a coherent timeline from large collections of source videos while aligning with music rhythm, user intent, story completeness, and long-range structural constraints. Existing approaches typically rely on fixed pipelines or simplified retrieval-and-concatenation paradigms, limiting their ability to adapt to diverse prompts and heterogeneous source materials. In this paper, we present GLANCE, a global-local coordination multi-agent framework for music-grounded nonlinear video editing. GLANCE adopts a bi-loop architecture for better editing practice: an outer loop performs long-horizon planning and task-graph construction, and an inner loop adopts the "Observe-Think-Act-Verify" flow for segment-wise editing tasks and their refinements. To address the cross-segment and global conflict emerging after subtimelines composition, we introduce a dedicated global-local coordination mechanism with both preventive and corrective components, which includes a novelly designed context controller, conflict region decomposition module, and a bottom-up dynamic negotiation mechanism. To support rigorous evaluation, we construct MVEBench, a new benchmark that factorizes editing difficulty along task type, prompt specificity, and music length, and propose an agent-as-a-judge evaluation framework for scalable multi-dimensional assessment. Experimental results show that GLANCE consistently outperforms prior research baselines and open-source product baselines under the same backbone models. With GPT-4o-mini as the backbone, GLANCE improves over the strongest baseline by 33.2% and 15.6% on two task settings, respectively. Human evaluation further confirms the quality of the generated videos and validates the effectiveness of the proposed evaluation framework.


【6】YMIR: A new Benchmark Dataset and Model for Arabic Yemeni Music Genre Classification Using Convolutional Neural Networks
标题:YMIR:使用卷积神经网络进行阿拉伯也门音乐流派分类的新基准数据集和模型
链接:https://arxiv.org/abs/2604.05011

作者:Moeen AL-Makhlafi,Abdulrahman A. AlKannad,Eiad Almekhlafi,Nawaf Q. Othman Ahmed Mohammed,Saher Qaid
摘要:自动音乐流派分类是音乐信息检索中的一项主要任务;然而,大多数当前的基准和模型主要针对西方音乐开发,使得特定文化传统的代表性不足。在本文中,我们介绍了也门音乐信息检索(YMIR)数据集,其中包含1,475个精心挑选的音频片段,涵盖五种传统的也门流派:Sanaani,Hadhrami,Lahji,Tihami和Adeni。该数据集由五位也门音乐专家按照清晰和结构化的协议进行标记,从而产生了强烈的注释者间协议(Fleiss kappa = 0.85)。我们还提出了也门音乐分类模型(YMCM),一个基于卷积神经网络(CNN)的系统,旨在从时频特征对音乐流派进行分类。使用一致的预处理管道,我们在六个实验组和五个不同的架构中进行了系统的比较,总共进行了30次实验。具体来说,我们评估了几种特征表示,包括梅尔频谱图,色度,滤波器组和具有13,20和40个系数的MFCC,并在相同的实验条件下将YMCM与标准模型(AlexNet,VGG 16,MobileNet和基线CNN)进行基准测试。实验结果表明,YMCM是最有效的,达到最高的准确率为98.8%的梅尔频谱图特征。研究结果还提供了实用的见解之间的关系特征表示和模型的能力。调查结果建立YMIR作为一个有用的基准和YMCM作为一个强大的基线分类也门音乐流派。
摘要:Automatic music genre classification is a major task in music information retrieval; however, most current benchmarks and models have been developed primarily for Western music, leaving culturally specific traditions underrepresented. In this paper, we introduce the Yemeni Music Information Retrieval (YMIR) dataset, which contains 1,475 carefully selected audio clips covering five traditional Yemeni genres: Sanaani, Hadhrami, Lahji, Tihami, and Adeni. The dataset was labeled by five Yemeni music experts following a clear and structured protocol, resulting in strong inter-annotator agreement (Fleiss kappa = 0.85). We also propose the Yemeni Music Classification Model (YMCM), a convolutional neural network (CNN)-based system designed to classify music genres from time-frequency features. Using a consistent preprocessing pipeline, we perform a systematic comparison across six experimental groups and five different architectures, resulting in a total of 30 experiments. Specifically, we evaluate several feature representations, including Mel-spectrograms, Chroma, FilterBank, and MFCCs with 13, 20, and 40 coefficients, and benchmark YMCM against standard models (AlexNet, VGG16, MobileNet, and a baseline CNN) under the same experimental conditions. The experimental findings reveal that YMCM is the most effective, achieving the highest accuracy of 98.8% with Mel-spectrogram features. The results also provide practical insights into the relationship between feature representation and model capacity. The findings establish YMIR as a useful benchmark and YMCM as a strong baseline for classifying Yemeni music genres.


【7】Generalizable Audio-Visual Navigation via Binaural Difference Attention and Action Transition Prediction
标题:通过双耳差异注意力和动作转换预测的可推广视听导航
链接:https://arxiv.org/abs/2604.05007

作者:Jia Li,Yinfeng Yu
备注:Main paper (6 pages). Accepted for publication by the International Joint Conference on Neural Networks (IJCNN 2026)
摘要:在视听导航(AVN)中,智能体必须使用视觉和听觉线索在看不见的3D环境中定位声源。然而,现有的方法往往在看不见的场景中难以泛化,因为它们往往过拟合语义声音特征和特定的训练环境。为了应对这些挑战,我们提出了\textbf{双耳差异注意力与动作转换预测(BDATP)}框架,它联合优化感知和策略。具体来说,\textbf{双耳差异注意力(BDA)}模块明确地对双耳差异进行建模,以增强空间方向,减少对语义类别的依赖。同时,\textbf{Action Transition Prediction(ATP)}任务引入了一个辅助动作预测目标作为正则化项,以减轻特定于环境的过拟合。在KIDS和Matterport 3D数据集上进行的大量实验表明,BDATP可以无缝集成到各种主流基线中,从而获得一致且显著的性能提升。值得注意的是,我们的框架在大多数环境中都实现了最先进的成功率,对于未听到的声音,在可识别数据集中有高达21.6个百分点的显着绝对改善。这些结果强调了BDATP优越的泛化能力及其在不同导航架构中的鲁棒性。
摘要:In Audio-Visual Navigation (AVN), agents must locate sound sources in unseen 3D environments using visual and auditory cues. However, existing methods often struggle with generalization in unseen scenarios, as they tend to overfit to semantic sound features and specific training environments. To address these challenges, we propose the \textbf{Binaural Difference Attention with Action Transition Prediction (BDATP)} framework, which jointly optimizes perception and policy. Specifically, the \textbf{Binaural Difference Attention (BDA)} module explicitly models interaural differences to enhance spatial orientation, reducing reliance on semantic categories. Simultaneously, the \textbf{Action Transition Prediction (ATP)} task introduces an auxiliary action prediction objective as a regularization term, mitigating environment-specific overfitting. Extensive experiments on the Replica and Matterport3D datasets demonstrate that BDATP can be seamlessly integrated into various mainstream baselines, yielding consistent and significant performance gains. Notably, our framework achieves state-of-the-art Success Rates across most settings, with a remarkable absolute improvement of up to 21.6 percentage points in Replica dataset for unheard sounds. These results underscore BDATP's superior generalization capability and its robustness across diverse navigation architectures.


【8】Brain-to-Speech: Prosody Feature Engineering and Transformer-Based Reconstruction
标题:大脑到言语:韵律特征工程和基于转换器的重建
链接:https://arxiv.org/abs/2604.05751

作者:Mohammed Salah Al-Radhi,Géza Németh,Andon Tchechmedjiev,Binbin Xu
备注:OpenAccess chapter: 10.1007/978-3-032-10561-5_16. In: Curry, E., et al. Artificial Intelligence, Data and Robotics (2026)
摘要:本章提出了一种新的方法,脑语音合成(BTS)从颅内脑电图(iEEG)数据,强调韵律感知的特征工程和先进的基于转换器的模型高保真语音重建。由于人们对直接从大脑活动中解码语音的兴趣日益浓厚,这项工作整合了神经科学、人工智能和信号处理,以生成准确而自然的语音。我们介绍了一种新的管道,用于直接从复杂的大脑iEEG信号中提取关键的韵律特征,包括语调,音高和节奏。为了有效地利用这些关键特征来生成自然的语音,我们采用了先进的深度学习模型。此外,本章介绍了一种新的Transformer编码器架构,专门设计用于大脑到语音的任务。与传统模型不同,我们的架构集成了提取的韵律特征,以显着增强语音重建,从而产生具有更好的可懂度和表现力的语音。详细的评估表明,在定量和感知指标方面,它的性能优于现有的基线方法,如传统的Griffin-Lim和基于CNN的重建。通过展示这些在特征提取和基于transformer的学习方面的进步,本章为人工智能驱动的神经修复领域的不断发展做出了贡献,为恢复言语障碍患者沟通的辅助技术铺平了道路。最后,我们讨论了有前途的未来研究方向,包括扩散模型和实时推理系统的集成。
摘要:This chapter presents a novel approach to brain-to-speech (BTS) synthesis from intracranial electroencephalography (iEEG) data, emphasizing prosody-aware feature engineering and advanced transformer-based models for high-fidelity speech reconstruction. Driven by the increasing interest in decoding speech directly from brain activity, this work integrates neuroscience, artificial intelligence, and signal processing to generate accurate and natural speech. We introduce a novel pipeline for extracting key prosodic features directly from complex brain iEEG signals, including intonation, pitch, and rhythm. To effectively utilize these crucial features for natural-sounding speech, we employ advanced deep learning models. Furthermore, this chapter introduces a novel transformer encoder architecture specifically designed for brain-to-speech tasks. Unlike conventional models, our architecture integrates the extracted prosodic features to significantly enhance speech reconstruction, resulting in generated speech with improved intelligibility and expressiveness. A detailed evaluation demonstrates superior performance over established baseline methods, such as traditional Griffin-Lim and CNN-based reconstruction, across both quantitative and perceptual metrics. By demonstrating these advancements in feature extraction and transformer-based learning, this chapter contributes to the growing field of AI-driven neuroprosthetics, paving the way for assistive technologies that restore communication for individuals with speech impairments. Finally, we discuss promising future research directions, including the integration of diffusion models and real-time inference systems.


【9】Active noise cancellation on open-ear smart glasses
标题:开耳智能眼镜上的主动降噪
链接:https://arxiv.org/abs/2604.05519

作者:Kuang Yuan,Freddy Yifei Liu,Tong Xiao,Yiwen Song,Chengyi Shen,Saksham Bhutani,Justin Chan,Swarun Kumar
摘要:智能眼镜正在成为一种越来越流行的可穿戴平台,音频是一种关键的交互方式。然而,在嘈杂环境中的听力仍然具有挑战性,因为智能眼镜配备了不密封耳道的开放式扬声器。此外,开耳设计与传统的有源噪声消除(ANC)技术不兼容,传统的有源噪声消除(ANC)技术依赖于耳道内部或入口处的误差麦克风来测量消除后听到的残余声音。在这里,我们介绍了第一个用于开放式智能眼镜的实时ANC系统,该系统仅使用麦克风和嵌入眼镜框架中的小型开放式扬声器来抑制环境噪声。我们的低延迟计算管道从分布在眼镜框架周围的八个麦克风阵列中估计耳朵处的噪声,并实时生成抗噪声信号以消除环境噪声。我们开发了一款定制眼镜原型,并在100- 1000 Hz频率范围内的8种移动环境(环境噪音集中)的用户研究中对其进行评估。我们实现了9.6 dB的平均噪声降低没有任何校准,11.2 dB的简短的用户特定的校准。
摘要:Smart glasses are becoming an increasingly prevalent wearable platform, with audio as a key interaction modality. However, hearing in noisy environments remains challenging because smart glasses are equipped with open-ear speakers that do not seal the ear canal. Furthermore, the open-ear design is incompatible with conventional active noise cancellation (ANC) techniques, which rely on an error microphone inside or at the entrance of the ear canal to measure the residual sound heard after cancellation. Here we present the first real-time ANC system for open-ear smart glasses that suppresses environmental noise using only microphones and miniaturized open-ear speakers embedded in the glasses frame. Our low-latency computational pipeline estimates the noise at the ear from an array of eight microphones distributed around the glasses frame and generates an anti-noise signal in real-time to cancel environmental noise. We develop a custom glasses prototype and evaluate it in a user study across 8 environments under mobility in the 100--1000 Hz frequency range, where environmental noise is concentrated. We achieve a mean noise reduction of 9.6 dB without any calibration, and 11.2 dB with a brief user-specific calibration.


【10】StrADiff: A Structured Source-Wise Adaptive Diffusion Framework for Linear and Nonlinear Blind Source Separation
标题:StrADiff:一种用于线性和非线性盲源分离的结构化逐源自适应扩散框架
链接:https://arxiv.org/abs/2604.04973

作者:Yuan-Hao Wei
摘要:本文提出了一种用于线性和非线性盲源分离的结构化逐源自适应扩散框架。该框架将每个潜在维度解释为源组件,并为其分配单独的自适应扩散机制,从而建立基于源的潜在建模,而不是依赖于单个共享的潜在先验。由此产生的公式在统一的端到端目标内共同学习源恢复和混合/重建过程,允许模型参数和潜在源在训练期间同时适应。这为线性和非线性盲源分离提供了一个通用的框架。在本实例化中,每个源在对潜在轨迹施加逐源时间结构之前还配备有其自己的自适应高斯过程(GP),而整体框架不限于高斯过程先验,并且原则上可以容纳其他结构化源先验。因此,所提出的模型提供了一个一般的结构化的基于扩散的路线,以无监督的源恢复,具有潜在的相关性超越盲源分离,可解释的潜在建模,源明智的解纠缠,并在适当的结构条件下潜在的可识别的非线性潜变量学习。
摘要:This paper presents a Structured Source-Wise Adaptive Diffusion Framework for linear and nonlinear blind source separation. The framework interprets each latent dimension as a source component and assigns to it an individual adaptive diffusion mechanism, thereby establishing source-wise latent modeling rather than relying on a single shared latent prior. The resulting formulation learns source recovery and the mixing/reconstruction process jointly within a unified end-to-end objective, allowing model parameters and latent sources to adapt simultaneously during training. This yields a common framework for both linear and nonlinear blind source separation. In the present instantiation, each source is further equipped with its own adaptive Gaussian process (GP) prior to impose source-wise temporal structure on the latent trajectories, while the overall framework is not restricted to Gaussian process priors and can in principle accommodate other structured source priors. The proposed model thus provides a general structured diffusion-based route to unsupervised source recovery, with potential relevance beyond blind source separation to interpretable latent modeling, source-wise disentanglement, and potentially identifiable nonlinear latent-variable learning under appropriate structural conditions.


eess.AS音频处理


【1】Multimodal Deep Learning Method for Real-Time Spatial Room Impulse Response Computing
标题:实时空间房间脉冲响应计算的多模式深度学习方法
链接:https://arxiv.org/abs/2604.05545

作者:Zhiyu Li,Xinwen Yue,Shenghui Zhao,Jing Wang
备注:This work was accepted by ICASSP 2026
摘要:我们提出了一种用于VR听觉化的多模态深度学习模型,该模型实时生成空间房间脉冲响应(SRIR),以重建场景特定的听觉感知。采用SRIR作为输出降低了计算复杂度,并促进了与个性化头部相关传递函数的集成。该模型采用两种模态作为输入:场景信息和波形,其中波形对应于低阶反射(LoR)。LoR可以使用几何声学(GA)有效地计算,但深度学习模型仍然难以准确预测。场景几何形状、声学特性、源坐标和收听者坐标首先用于通过GA实时计算LoR,并且LoR和这些特征随后作为输入提供给模型。构建了一个新的数据集,由多个场景及其相应的SRIR组成。数据集显示出更大的多样性。实验结果证明了该模型的优越性能。
摘要:We propose a multimodal deep learning model for VR auralization that generates spatial room impulse responses (SRIRs) in real time to reconstruct scene-specific auditory perception. Employing SRIRs as the output reduces computational complexity and facilitates integration with personalized head-related transfer functions. The model takes two modalities as input: scene information and waveforms, where the waveform corresponds to the low-order reflections (LoR). LoR can be efficiently computed using geometrical acoustics (GA) but remains difficult for deep learning models to predict accurately. Scene geometry, acoustic properties, source coordinates, and listener coordinates are first used to compute LoR in real time via GA, and both LoR and these features are subsequently provided as inputs to the model. A new dataset was constructed, consisting of multiple scenes and their corresponding SRIRs. The dataset exhibits greater diversity. Experimental results demonstrate the superior performance of the proposed model.


【2】Active noise cancellation on open-ear smart glasses
标题:开耳智能眼镜上的主动降噪
链接:https://arxiv.org/abs/2604.05519

作者:Kuang Yuan,Freddy Yifei Liu,Tong Xiao,Yiwen Song,Chengyi Shen,Saksham Bhutani,Justin Chan,Swarun Kumar
摘要:智能眼镜正在成为一种越来越流行的可穿戴平台,音频是一种关键的交互方式。然而,在嘈杂环境中的听力仍然具有挑战性,因为智能眼镜配备了不密封耳道的开放式扬声器。此外,开耳设计与传统的有源噪声消除(ANC)技术不兼容,传统的有源噪声消除(ANC)技术依赖于耳道内部或入口处的误差麦克风来测量消除后听到的残余声音。在这里,我们介绍了第一个用于开放式智能眼镜的实时ANC系统,该系统仅使用麦克风和嵌入眼镜框架中的小型开放式扬声器来抑制环境噪声。我们的低延迟计算管道从分布在眼镜框架周围的八个麦克风阵列中估计耳朵处的噪声,并实时生成抗噪声信号以消除环境噪声。我们开发了一款定制眼镜原型,并在100- 1000 Hz频率范围内的8种移动环境(环境噪音集中)的用户研究中对其进行评估。我们实现了9.6 dB的平均噪声降低没有任何校准,11.2 dB的简短的用户特定的校准。
摘要:Smart glasses are becoming an increasingly prevalent wearable platform, with audio as a key interaction modality. However, hearing in noisy environments remains challenging because smart glasses are equipped with open-ear speakers that do not seal the ear canal. Furthermore, the open-ear design is incompatible with conventional active noise cancellation (ANC) techniques, which rely on an error microphone inside or at the entrance of the ear canal to measure the residual sound heard after cancellation. Here we present the first real-time ANC system for open-ear smart glasses that suppresses environmental noise using only microphones and miniaturized open-ear speakers embedded in the glasses frame. Our low-latency computational pipeline estimates the noise at the ear from an array of eight microphones distributed around the glasses frame and generates an anti-noise signal in real-time to cancel environmental noise. We develop a custom glasses prototype and evaluate it in a user study across 8 environments under mobility in the 100--1000 Hz frequency range, where environmental noise is concentrated. We achieve a mean noise reduction of 9.6 dB without any calibration, and 11.2 dB with a brief user-specific calibration.


【3】Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
标题:探索整个生命周期的说话者数字化的语音基础模型
链接:https://arxiv.org/abs/2604.05201

作者:Anfeng Xu,Tiantian Feng,Shrikanth Narayanan
备注:Under review for Interspeech 2026
摘要:语音基础模型在广泛的语音应用中表现出很强的可移植性。然而,他们的鲁棒性与年龄有关的域转移的发言人日记仍然未充分探索。在这项工作中,我们在统一的端到端神经日记化框架(EEND-VC)中提出了跨寿命评估,涵盖了涉及儿童、成人和老年人的对话中的语音样本。我们比较了zero-shot交叉年龄推断,联合多年龄训练和特定领域自适应下的模型。结果显示,当将成人特定语音训练的模型应用于儿童和老年人-成人对话数据时,性能会大幅下降。此外,跨不同年龄组的联合多年龄训练提高了鲁棒性,而不会降低规范成人对话中的日记化性能,而目标年龄组自适应则进一步提高了日记化性能,特别是在使用Whisper编码器时。
摘要:Speech foundation models have shown strong transferability across a wide range of speech applications. However, their robustness to age-related domain shift in speaker diarization remains underexplored. In this work, we present a cross-lifespan evaluation within a unified end-to-end neural diarization framework (EEND-VC), covering speech samples from conversations involving children, adults, and older adults. We compare models under zero-shot cross-age inference, joint multi-age training, and domain-specific adaptation. Results show substantial performance degradation when models trained on adult-specific speech are applied to child and older-adult conversational data. Moreover, joint multi-age training across different age groups improves robustness without reducing diarization performance in canonical adult conversations, while targeted age group adaptation yields further gains in diarization performance, particularly when using the Whisper encoder.


【4】Generalizable Audio-Visual Navigation via Binaural Difference Attention and Action Transition Prediction
标题:通过双耳差异注意力和动作转换预测的可推广视听导航
链接:https://arxiv.org/abs/2604.05007

作者:Jia Li,Yinfeng Yu
备注:Main paper (6 pages). Accepted for publication by the International Joint Conference on Neural Networks (IJCNN 2026)
摘要:在视听导航(AVN)中,智能体必须使用视觉和听觉线索在看不见的3D环境中定位声源。然而,现有的方法往往在看不见的场景中难以泛化,因为它们往往过拟合语义声音特征和特定的训练环境。为了应对这些挑战,我们提出了\textbf{双耳差异注意力与动作转换预测(BDATP)}框架,它联合优化感知和策略。具体来说,\textbf{双耳差异注意力(BDA)}模块明确地对双耳差异进行建模,以增强空间方向,减少对语义类别的依赖。同时,\textbf{Action Transition Prediction(ATP)}任务引入了一个辅助动作预测目标作为正则化项,以减轻特定于环境的过拟合。在KIDS和Matterport 3D数据集上进行的大量实验表明,BDATP可以无缝集成到各种主流基线中,从而获得一致且显著的性能提升。值得注意的是,我们的框架在大多数环境中都实现了最先进的成功率,对于未听到的声音,在可识别数据集中有高达21.6个百分点的显着绝对改善。这些结果强调了BDATP优越的泛化能力及其在不同导航架构中的鲁棒性。
摘要:In Audio-Visual Navigation (AVN), agents must locate sound sources in unseen 3D environments using visual and auditory cues. However, existing methods often struggle with generalization in unseen scenarios, as they tend to overfit to semantic sound features and specific training environments. To address these challenges, we propose the \textbf{Binaural Difference Attention with Action Transition Prediction (BDATP)} framework, which jointly optimizes perception and policy. Specifically, the \textbf{Binaural Difference Attention (BDA)} module explicitly models interaural differences to enhance spatial orientation, reducing reliance on semantic categories. Simultaneously, the \textbf{Action Transition Prediction (ATP)} task introduces an auxiliary action prediction objective as a regularization term, mitigating environment-specific overfitting. Extensive experiments on the Replica and Matterport3D datasets demonstrate that BDATP can be seamlessly integrated into various mainstream baselines, yielding consistent and significant performance gains. Notably, our framework achieves state-of-the-art Success Rates across most settings, with a remarkable absolute improvement of up to 21.6 percentage points in Replica dataset for unheard sounds. These results underscore BDATP's superior generalization capability and its robustness across diverse navigation architectures.


机器翻译由腾讯交互翻译提供,仅供参考