今日论文合集:cs.SD语音6篇,eess.AS音频处理7篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】MSR-HuBERT: Self-supervised Pre-training for Adaptation to Multiple Sampling Rates
标题:MSR-HuBERT:适应多个采样率的自我监督预训练
链接:https://arxiv.org/abs/2603.23048

作者:Zikang Huang,Meng Ge,Tianrui Wang,Xuanchen Li,Xiaobao Wang,Longbiao Wang,Jianwu Dang
摘要:自监督学习(SSL)具有先进的语音处理。然而,现有的语音SSL方法通常假设一个单一的采样率和斗争与混合速率的数据,由于时间分辨率不匹配。为了解决这个问题,我们提出了MSRHuBERT,多采样率自适应预训练方法。在HuBERT的基础上,我们将其单速率下采样CNN替换为多采样率自适应下采样CNN,该CNN将来自不同采样率的原始波形映射到共享的时间分辨率,而无需重新采样。该设计实现了统一的混合速率预训练和微调。在16至48 kHz的实验中,MSRHuBERT在语音识别和全频带语音重建方面优于HuBERT,在对低频语义结构建模的同时保留了高频细节。此外,MSRHuBERT保留了HuBERT的掩模预测目标和Transformer编码器,因此为HuBERT开发的现有分析和改进可以直接应用。
摘要:Self-supervised learning (SSL) has advanced speech processing. However, existing speech SSL methods typically assume a single sampling rate and struggle with mixed-rate data due to temporal resolution mismatch. To address this limitation, we propose MSRHuBERT, a multi-sampling-rate adaptive pre-training method. Building on HuBERT, we replace its single-rate downsampling CNN with a multi-sampling-rate adaptive downsampling CNN that maps raw waveforms from different sampling rates to a shared temporal resolution without resampling. This design enables unified mixed-rate pre-training and fine-tuning. In experiments spanning 16 to 48 kHz, MSRHuBERT outperforms HuBERT on speech recognition and full-band speech reconstruction, preserving high-frequency detail while modeling low-frequency semantic structure. Moreover, MSRHuBERT retains HuBERT's mask-prediction objective and Transformer encoder, so existing analyses and improvements that were developed for HuBERT can apply directly.


【2】The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
标题:大型音频语言模型的Interspeech 2026音频编码器能力挑战
链接:https://arxiv.org/abs/2603.22728

作者:Heinrich Dinkel,Jiahao Zhou,Guanbo Wang,Yadong Niu,Junbo Zhang,Yufeng Hao,Ying Liu,Ke Li,Wenwu Wang,Zhiyong Wu,Jian Luan
备注:Interspeech 2026 Challenge
摘要:本文介绍了Interspeech 2026音频编码器能力挑战赛,这是一个专门设计用于评估和提高预训练音频编码器作为大型音频语言模型(LALM)前端模块的性能的基准。虽然LALM已经显示出对复杂声学场景的显著理解,但它们的性能取决于底层音频编码器表示的语义丰富性。这一挑战通过提供统一的生成评估框架XARES-LLM来解决集成差距,该框架可评估各种下游分类和生成任务中提交的编码器。通过将编码器开发与LLM微调解耦,该挑战为通用音频表示建立了一个标准化协议,可有效地用于下一代多模态语言模型。
摘要:This paper presents the Interspeech 2026 Audio Encoder Capability Challenge, a benchmark specifically designed to evaluate and advance the performance of pre-trained audio encoders as front-end modules for Large Audio Language Models (LALMs). While LALMs have shown remarkable understanding of complex acoustic scenes, their performance depends on the semantic richness of the underlying audio encoder representations. This challenge addresses the integration gap by providing a unified generative evaluation framework, XARES-LLM, which assesses submitted encoders across a diverse suite of downstream classification and generation tasks. By decoupling encoder development from LLM fine-tuning, the challenge establishes a standardized protocol for general-purpose audio representations that can effectively be used for the next generation of multimodal language models.


【3】MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation
标题:MuQ-Eval:用于人工智能音乐生成评估的开源每样本质量指标
链接:https://arxiv.org/abs/2603.22677

作者:Di Zhu,Zixuan Li
备注:10 Pages, 6 figures
摘要:分布式度量(如Fréchet Audio Distance)无法对单个音乐片段进行评分,并且与人类判断的相关性很差,而唯一实现高度人类相关性的每个样本学习度量是闭源的。我们介绍了MUQ-EVAL,这是一种针对人工智能生成的音乐的开源每样本质量指标,通过使用MusicEval在冻结的MuQ-310 M功能上训练轻量级预测头来构建,MusicEval是一个来自31个文本到音乐系统的生成剪辑的数据集,具有专家质量评级。我们最简单的模型,具有注意力池和两层MLP的冻结特征,在人类平均意见评分的情况下实现了系统级SRCC = 0.957和话语级SRCC = 0.838。对训练目标和适应策略的系统性消融表明,没有增加有意义的改善冻结基线,表明冻结MuQ表示已经捕获质量相关的信息。编码器的选择是主要的设计因素,超过了所有的架构和培训决策。在150个片段上训练的LoRA适应模型已经实现了可用的相关性,使个性化的质量评估器能够从单个听众注释中获得。一个受控的退化分析揭示了选择性的信号电平的文物,但不敏感的音乐结构失真的敏感性。我们的指标MUQ-EVAL是完全开源的,优于现有的开放样本指标,并在单个消费者GPU上实时运行。代码、模型权重和评估脚本可在https://github.com/dgtql/MuQ-Eval上获得。
摘要:Distributional metrics such as Fréchet Audio Distance cannot score individual music clips and correlate poorly with human judgments, while the only per-sample learned metric achieving high human correlation is closed-source. We introduce MUQ-EVAL, an open-source per-sample quality metric for AIgenerated music built by training lightweight prediction heads on frozen MuQ-310M features using MusicEval, a dataset of generated clips from 31 text-to-music systems with expert quality ratings. Our simplest model, frozen features with attention pooling and a two-layer MLP, achieves system-level SRCC = 0.957 and utterance-level SRCC = 0.838 with human mean opinion scores. A systematic ablation over training objectives and adaptation strategies shows that no addition meaningfully improves the frozen baseline, indicating that frozen MuQ representations already capture quality-relevant information. Encoder choice is the dominant design factor, outweighing all architectural and training decisions. LoRA-adapted models trained on as few as 150 clips already achieve usable correlation, enabling personalized quality evaluators from individual listener annotations. A controlled degradation analysis reveals selective sensitivity to signal-level artifacts but insensitivity to musical-structural distortions. Our metric, MUQ-EVAL, is fully open-source, outperforms existing open per-sample metrics, and runs in real time on a single consumer GPU. Code, model weights, and evaluation scripts are available at https://github.com/dgtql/MuQ-Eval.


【4】Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
标题:用于高效立体声冲击响应建模的速度势神经场
链接:https://arxiv.org/abs/2603.22589

作者:Yoshiki Masuyama,Francois G. Germain,Gordon Wichern,Chiori Hori,Jonathan Le Roux
备注:Accepted to ICASSP 2026
摘要:一阶高保真度立体声(FOA)是一种基于球面谐波分解的标准空间音频格式。它的零阶和一阶分量分别捕获声压和质点速度。最近,物理信息神经网络已被应用于FOA信号的空间内插,基于从物理原理导出的软惩罚项来正则化网络输出,例如,线性化的动量方程在本文中,我们重新制定的任务,使预测的FOA信号自动满足线性化的动量方程。我们的网络近似一个称为速度势的标量函数,而不是FOA信号本身。然后,通过速度势相对于网络输入的偏导数(即,时间和麦克风位置)。由单通道速度势导出四通道FOA,重构信号在任意时刻、任意位置都遵循物理原理。室内脉冲响应重建实验结果证实了该框架的有效性。
摘要:First-order Ambisonics (FOA) is a standard spatial audio format based on spherical harmonic decomposition. Its zeroth- and first-order components capture the sound pressure and particle velocity, respectively. Recently, physics-informed neural networks have been applied to the spatial interpolation of FOA signals, regularizing the network outputs based on soft penalty terms derived from physical principles, e.g., the linearized momentum equation. In this paper, we reformulate the task so that the predicted FOA signal automatically satisfies the linearized momentum equation. Our network approximates a scalar function called velocity potential, rather than the FOA signal itself. Then, the FOA signal can be readily recovered through the partial derivatives of the velocity potential with respect to the network inputs (i.e., time and microphone position) according to physics of sound propagation. By deriving the four channels of FOA from the single-channel velocity potential, the reconstructed signal follows the physical principle at any time and position by construction. Experimental results on room impulse response reconstruction confirm the effectiveness of the proposed framework.


【5】ST-GDance++: A Scalable Spatial-Temporal Diffusion for Long-Duration Group Choreography
标题:ST-GDance++:长时间团体编舞的可扩展时空扩散
链接:https://arxiv.org/abs/2603.22316

作者:Jing Xu,Weiqiang Wang,Cunjian Chen,Jun Liu,Qiuhong Ke
摘要:从音乐生成群舞需要同步多个舞者,同时保持空间协调,使其与电影制作,游戏和动画等应用高度相关。最近的群舞生成模型已经取得了很好的生成质量,但由于双向注意依赖性,它们仍然难以部署在交互式场景中。随着舞者的数量和序列长度的增加,将音乐条件与运动序列对齐所需的注意力计算以二次方式增长,导致效率降低和运动冲突的风险增加。因此,有效地建模密集的时空交互是必不可少的,但现有的方法往往难以捕捉这种复杂性,导致有限的可扩展性和不稳定的多舞者协调。为了解决这些挑战,我们提出了ST-GDance++,一个可扩展的框架,它可以扩展空间和时间的依赖关系,以实现高效和冲突感知的组编排生成。对于空间建模,我们引入轻量级的距离感知图卷积来捕获舞者之间的关系,同时减少计算开销。对于时间建模,我们设计了一个扩散噪声调度策略,以及一个有效的时间对齐的注意掩模,使基于流的生成长运动序列,并提高在长时间的情况下的可扩展性。在AIOZ-GDance数据集上的实验表明,与现有方法相比,ST-GDance++实现了具有竞争力的生成质量,并显著降低了延迟。
摘要:Group dance generation from music requires synchronizing multiple dancers while maintaining spatial coordination, making it highly relevant to applications such as film production, gaming, and animation. Recent group dance generation models have achieved promising generation quality, but they remain difficult to deploy in interactive scenarios due to bidirectional attention dependencies. As the number of dancers and the sequence length increase, the attention computation required for aligning music conditions with motion sequences grows quadratically, leading to reduced efficiency and increased risk of motion collisions. Effectively modeling dense spatial-temporal interactions is therefore essential, yet existing methods often struggle to capture such complexity, resulting in limited scalability and unstable multi-dancer coordination. To address these challenges, we propose ST-GDance++, a scalable framework that decouples spatial and temporal dependencies to enable efficient and collision-aware group choreography generation. For spatial modeling, we introduce lightweight distance-aware graph convolutions to capture inter-dancer relationships while reducing computational overhead. For temporal modeling, we design a diffusion noise scheduling strategy together with an efficient temporal-aligned attention mask, enabling stream-based generation for long motion sequences and improving scalability in long-duration scenarios. Experiments on the AIOZ-GDance dataset show that ST-GDance++ achieves competitive generation quality with significantly reduced latency compared to existing methods.


【6】MSP-Conversation: A Corpus for Naturalistic, Time-Continuous Emotion Recognition
标题:MSP-Conversation:自然主义、时间连续情感识别的数据库
链接:https://arxiv.org/abs/2603.22536

作者:Luz Martinez-Lucas,Pravin Mote,Abinay Reddy Naini,Mohammed Abdelwahab,Carlos Busso
摘要:情感计算旨在为计算系统理解和建模人类情感。在这个领域中,语音情感识别(SER)专注于预测通过语音传达的情感。虽然早期的SER系统依赖于有限的数据集和传统的机器学习模型,但最近的深度学习方法需要大规模的自然情感语料库。为了满足这一需求,我们引入了MSP会话语料库:一个超过70小时的会话音频数据集,具有时间连续的情感注释和详细的说话者日记。时间连续的注释捕捉到了情感表达的动态性和上下文相关性。语料库中的注释包括价、唤醒和支配的细粒度时间痕迹。音频数据来源于公开可用的播客,并且与MSP-播客语料库中的孤立的说话回合的子集重叠,以便于注释方法之间的直接比较(即,上下文内注释与上下文外注释)。本文概述了语料库的发展,注释方法,注释的分析,和基线SER实验,建立MSP会话语料库作为一个宝贵的资源,推进研究动态SER在自然环境。
摘要:Affective computing aims to understand and model human emotions for computational systems. Within this field, speech emotion recognition (SER) focuses on predicting emotions conveyed through speech. While early SER systems relied on limited datasets and traditional machine learning models, recent deep learning approaches demand largescale, naturalistic emotional corpora. To address this need, we introduce the MSP-Conversation corpus: a dataset of more than 70 hours of conversational audio with time-continuous emotional annotations and detailed speaker diarizations. The time-continuous annotations capture the dynamic and contextdependent nature of emotional expression. The annotations in the corpus include fine-grained temporal traces of valence, arousal, and dominance. The audio data is sourced from publicly available podcasts and overlaps with a subset of the isolated speaking turns in the MSP-Podcast corpus to facilitate direct comparisons between annotation methods (i.e., in-context versus out-of-context annotations). The paper outlines the development of the corpus, annotation methodology, analyses of the annotations, and baseline SER experiments, establishing the MSP-Conversation corpus as a valuable resource for advancing research in dynamic SER in naturalistic settings.


eess.AS音频处理


【1】Prompt Amplification and Zero-Shot Late Fusion in Audio-Language Models for Speech Emotion Recognition
标题:语音情感识别的音频语言模型中的即时放大和Zero-Shot后期融合
链接:https://arxiv.org/abs/2603.23057

作者:Saurabh Kataria,Xiao Hu
摘要:音频语言模型(ALM)在理解语音和非语音音频方面取得了长足的进步。然而,领域专家的基础模型(FM)仍然是封闭式语音处理任务(如语音情感识别(SER))的最佳选择。将ALM用于Zero-shot SER是一种流行的选择,但它们与专家合作以实现最先进(SOTA)性能的潜力尚未开发。我们提出了一种后期融合方法,将来自双编码器ALM的zero-shot情感估计与专家FM相结合。为了处理情感的模糊性和对提示选择的敏感性,1)我们使用一个简单的提示集合,2)提出一种称为提示放大的新技术,它重复音频和文本查询,以发现更强的zero-shot能力。我们通过使用三个双编码器ALM和两个FM来评估T-S,并在三个语音情感识别数据集上报告了SOTA基线(如WavLM-Large)的改进来证明我们的技术的有效性。
摘要:Audio-Language Models (ALMs) are making strides in understanding speech and non-speech audio. However, domain-specialist Foundation Models (FMs) remain the best for closed-ended speech processing tasks such as Speech Emotion Recognition (SER). Using ALMs for Zero-shot SER is a popular choice, but their potential to work with specialists to achieve state-of-the-art (SOTA) performance remains unexplored. We propose ZS-Fuse, a late-fusion method that combines zero-shot emotion estimates from a dual-encoder ALM with specialist FMs. To handle ambiguity in emotions and sensitivity to prompt choice, 1) we use a simple prompt ensemble and 2) suggest a novel technique called prompt amplification, which repeats audio and text queries to discover stronger zero-shot capabilities. We demonstrate the efficacy of our technique by evaluating ZS-Fuse with three dual-encoder ALMs and two FMs, and report improvements over SOTA baselines, such as WavLM-Large, on three speech emotion recognition datasets.


【2】Modelling Emotions is an Elusive Pursuit in Affective Computing
标题:情感建模是情感计算中的一种难以捉摸的追求
链接:https://arxiv.org/abs/2603.23017

作者:Anders Rolighed Larsen,Sneha Das,Line Clemmensen
摘要:情感计算-结合传感器技术,机器学习和心理学-已经研究了三十多年,并用于人工智能技术,以增强人工智能系统的情感意识,并检测焦虑和抑郁等心理健康障碍的症状。然而,在这样的系统中的不确定性仍然很高,并且应用领域受到情感和情感概念的分类定义的限制。本文认为,分类情感标签模糊情感计算中的情感细微差别,因此需要连续的维度定义来推进该领域,增加应用程序的有用性,并降低不确定性。
摘要:Affective computing - combining sensor technology, machine learning, and psychology - have been studied for over three decades and is employed in AI-powered technologies to enhance emotional awareness in AI systems, and detect symptoms of mental health disorders such as anxiety and depression. However, the uncertainty in such systems remains high, and the application areas are limited by categorical definitions of emotions and emotional concepts. This paper argues that categorical emotion labels obscure emotional nuance in affective computing, and therefore continuous dimensional definitions are needed to advance the field, increase application usefulness, and lower uncertainties.


【3】MSP-Conversation: A Corpus for Naturalistic, Time-Continuous Emotion Recognition
标题:MSP-Conversation:自然主义、时间连续情感识别的数据库
链接:https://arxiv.org/abs/2603.22536

作者:Luz Martinez-Lucas,Pravin Mote,Abinay Reddy Naini,Mohammed Abdelwahab,Carlos Busso
摘要:情感计算旨在为计算系统理解和建模人类情感。在这个领域中,语音情感识别(SER)专注于预测通过语音传达的情感。虽然早期的SER系统依赖于有限的数据集和传统的机器学习模型,但最近的深度学习方法需要大规模的自然情感语料库。为了满足这一需求,我们引入了MSP会话语料库:一个超过70小时的会话音频数据集,具有时间连续的情感注释和详细的说话者日记。时间连续的注释捕捉到了情感表达的动态性和上下文相关性。语料库中的注释包括价、唤醒和支配的细粒度时间痕迹。音频数据来源于公开可用的播客,并且与MSP-播客语料库中的孤立的说话回合的子集重叠,以便于注释方法之间的直接比较(即,上下文内注释与上下文外注释)。本文概述了语料库的发展,注释方法,注释的分析,和基线SER实验,建立MSP会话语料库作为一个宝贵的资源,推进研究动态SER在自然环境。
摘要:Affective computing aims to understand and model human emotions for computational systems. Within this field, speech emotion recognition (SER) focuses on predicting emotions conveyed through speech. While early SER systems relied on limited datasets and traditional machine learning models, recent deep learning approaches demand largescale, naturalistic emotional corpora. To address this need, we introduce the MSP-Conversation corpus: a dataset of more than 70 hours of conversational audio with time-continuous emotional annotations and detailed speaker diarizations. The time-continuous annotations capture the dynamic and contextdependent nature of emotional expression. The annotations in the corpus include fine-grained temporal traces of valence, arousal, and dominance. The audio data is sourced from publicly available podcasts and overlaps with a subset of the isolated speaking turns in the MSP-Podcast corpus to facilitate direct comparisons between annotation methods (i.e., in-context versus out-of-context annotations). The paper outlines the development of the corpus, annotation methodology, analyses of the annotations, and baseline SER experiments, establishing the MSP-Conversation corpus as a valuable resource for advancing research in dynamic SER in naturalistic settings.


【4】The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
标题:大型音频语言模型的Interspeech 2026音频编码器能力挑战
链接:https://arxiv.org/abs/2603.22728

作者:Heinrich Dinkel,Jiahao Zhou,Guanbo Wang,Yadong Niu,Junbo Zhang,Yufeng Hao,Ying Liu,Ke Li,Wenwu Wang,Zhiyong Wu,Jian Luan
备注:Interspeech 2026 Challenge
摘要:本文介绍了Interspeech 2026音频编码器能力挑战赛,这是一个专门设计用于评估和提高预训练音频编码器作为大型音频语言模型(LALM)前端模块的性能的基准。虽然LALM已经显示出对复杂声学场景的显著理解,但它们的性能取决于底层音频编码器表示的语义丰富性。这一挑战通过提供统一的生成评估框架XARES-LLM来解决集成差距,该框架可评估各种下游分类和生成任务中提交的编码器。通过将编码器开发与LLM微调解耦,该挑战为通用音频表示建立了一个标准化协议,可有效地用于下一代多模态语言模型。
摘要:This paper presents the Interspeech 2026 Audio Encoder Capability Challenge, a benchmark specifically designed to evaluate and advance the performance of pre-trained audio encoders as front-end modules for Large Audio Language Models (LALMs). While LALMs have shown remarkable understanding of complex acoustic scenes, their performance depends on the semantic richness of the underlying audio encoder representations. This challenge addresses the integration gap by providing a unified generative evaluation framework, XARES-LLM, which assesses submitted encoders across a diverse suite of downstream classification and generation tasks. By decoupling encoder development from LLM fine-tuning, the challenge establishes a standardized protocol for general-purpose audio representations that can effectively be used for the next generation of multimodal language models.


【5】Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
标题:谁什么时候说了什么?使用语义和重叠感知量表评估对话式ASB的口语模型
链接:https://arxiv.org/abs/2603.22709

作者:Naohiro Tawara,Samuele Cornell,Alexander Polok,Marc Delcroix,Lukáš Burget,Shinji Watanabe
备注:Submitted to INTERSPEECH 2026
摘要:由于语音重叠、远场噪声和说话人数量的变化,会话自动语音识别仍然具有挑战性。虽然最近基于LLM的系统在单扬声器基准测试中表现良好,但它们在多扬声器设置中的鲁棒性尚不清楚。我们系统地比较了基于LLM和模块化管道方法沿着四个轴:重叠鲁棒性,语义保真度,扬声器计数,以及单声道与多声道输入。为了捕获传统度量错过的意义改变错误,我们引入了tcpSemER,它通过用基于嵌入的语义相似性替换Levenshtein距离来扩展tcpWER。我们进一步将tcpWER分解为重叠和非重叠组件,以进行更细粒度的分析。在三个数据集上的实验表明,基于LLM的系统在两个扬声器设置中具有竞争力,但随着扬声器数量和重叠的增加而降低,而模块化管道仍然更强大。
摘要:Conversational automatic speech recognition remains challenging due to overlapping speech, far-field noise, and varying speaker counts. While recent LLM-based systems perform well on single-speaker benchmarks, their robustness in multi-speaker settings is unclear. We systematically compare LLM-based and modular pipeline approaches along four axes: overlap robustness, semantic fidelity, speaker count, and single- versus multi-channel input. To capture meaning-altering errors that conventional metrics miss, we introduce tcpSemER, which extends tcpWER by replacing Levenshtein distance with embedding-based semantic similarity. We further decompose tcpWER into overlapping and non-overlapping components for finer-grained analysis. Experiments across three datasets show that LLM-based systems are competitive in two-speaker settings but degrade as speaker count and overlap increase, whereas modular pipelines remain more robust.


【6】Precision-Varying Prediction (PVP): Robustifying ASR systems against adversarial attacks
标题:精确变化预测(VP):增强ASB系统抵御对抗攻击
链接:https://arxiv.org/abs/2603.22590

作者:Matías Pizarro,Raghavan Narasimhan,Asja Fischer
摘要:随着自动化和代理系统的部署越来越多,确保自动语音识别(ASR)模型的对抗鲁棒性变得至关重要。我们观察到,在推理过程中改变ASR模型的精度可以降低对抗性攻击成功的可能性。我们利用这一事实,通过在预测过程中对精度进行简单的随机抽样,使模型更加鲁棒。此外,通过比较不同精度产生的输出并利用简单的高斯分类器,可以将这种洞察转化为对抗性示例检测策略。实验分析表明,在各种ASR模型和攻击类型的鲁棒性和竞争力的检测性能显着增加。
摘要:With the increasing deployment of automated and agentic systems, ensuring the adversarial robustness of automatic speech recognition (ASR) models has become critical. We observe that changing the precision of an ASR model during inference reduces the likelihood of adversarial attacks succeeding. We take advantage of this fact to make the models more robust by simple random sampling of the precision during prediction. Moreover, the insight can be turned into an adversarial example detection strategy by comparing outputs resulting from different precisions and leveraging a simple Gaussian classifier. An experimental analysis demonstrates a significant increase in robustness and competitive detection performance for various ASR models and attack types.


【7】Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
标题:用于高效立体声冲击响应建模的速度势神经场
链接:https://arxiv.org/abs/2603.22589

作者:Yoshiki Masuyama,Francois G. Germain,Gordon Wichern,Chiori Hori,Jonathan Le Roux
备注:Accepted to ICASSP 2026
摘要:一阶高保真度立体声(FOA)是一种基于球面谐波分解的标准空间音频格式。它的零阶和一阶分量分别捕获声压和质点速度。最近,物理信息神经网络已被应用于FOA信号的空间内插,基于从物理原理导出的软惩罚项来正则化网络输出,例如,线性化的动量方程在本文中,我们重新制定的任务,使预测的FOA信号自动满足线性化的动量方程。我们的网络近似一个称为速度势的标量函数,而不是FOA信号本身。然后,通过速度势相对于网络输入的偏导数(即,时间和麦克风位置)。由单通道速度势导出四通道FOA,重构信号在任意时刻、任意位置都遵循物理原理。室内脉冲响应重建实验结果证实了该框架的有效性。
摘要:First-order Ambisonics (FOA) is a standard spatial audio format based on spherical harmonic decomposition. Its zeroth- and first-order components capture the sound pressure and particle velocity, respectively. Recently, physics-informed neural networks have been applied to the spatial interpolation of FOA signals, regularizing the network outputs based on soft penalty terms derived from physical principles, e.g., the linearized momentum equation. In this paper, we reformulate the task so that the predicted FOA signal automatically satisfies the linearized momentum equation. Our network approximates a scalar function called velocity potential, rather than the FOA signal itself. Then, the FOA signal can be readily recovered through the partial derivatives of the velocity potential with respect to the network inputs (i.e., time and microphone position) according to physics of sound propagation. By deriving the four channels of FOA from the single-channel velocity potential, the reconstructed signal follows the physical principle at any time and position by construction. Experimental results on room impulse response reconstruction confirm the effectiveness of the proposed framework.


机器翻译由腾讯交互翻译提供,仅供参考