今日论文合集:cs.SD语音12篇,eess.AS音频处理14篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
标题:STAR-Bench:探索作为音频4D智能的深度时空推理
链接:http://arxiv.org/pdf/2510.24693v1

作者:Zihan Liu, Zhikang Niu, Qiuyang Xiao, Zhisheng Zheng, Ruoqi Yuan, Yuhang Zang, Yuhang Cao, Xiaoyi Dong, Jianze Liang, Xie Chen, Leilei Sun, Dahua Lin, Jiaqi Wang
备注:Homepage: this https URL
摘要:尽管多模态大型语言模型和大型音频语言模型取得了快速进展,但现有的音频基准主要测试可以从文本标题中恢复的语义,掩盖了细粒度感知推理的缺陷。我们将音频4D智能形式化,定义为在时间和3D空间中对声音动态进行推理,并引入STAR-Bench来测量它。STAR-Bench结合了基础声学感知设置(绝对和相对制度下的六个属性)与整体时空推理设置,包括连续和离散过程的段重新排序和跨越静态定位,多源关系,和动态轨迹。我们的数据管理管道使用两种方法来确保高质量的样本。对于基础任务,我们使用程序合成和物理模拟音频。对于整体数据,我们遵循一个四阶段的过程,包括人类注释和基于人类表现的最终选择。与之前的基准测试不同,只有标题的回答稍微降低了准确性,STAR-Bench引起了更大的下降(-31.5 %时间,-35.2 %空间),证明了它对语言学上难以描述的线索的关注。对19个模型的评估揭示了与人类和能力层次相比的巨大差距:闭源模型受到细粒度感知的检验,而开源模型在感知,知识和推理方面落后。我们的STAR-Bench为开发未来模型提供了重要的见解和明确的前进道路,并对物理世界有更深入的了解。摘要:Despite rapid progress in Multi-modal Large Language Models and Large Audio-Language Models, existing audio benchmarks largely test semantics that can be recovered from text captions, masking deficits in fine-grained perceptual reasoning. We formalize audio 4D intelligence that is defined as reasoning over sound dynamics in time and 3D space, and introduce STAR-Bench to measure it. STAR-Bench combines a Foundational Acoustic Perception setting (six attributes under absolute and relative regimes) with a Holistic Spatio-Temporal Reasoning setting that includes segment reordering for continuous and discrete processes and spatial tasks spanning static localization, multi-source relations, and dynamic trajectories. Our data curation pipeline uses two methods to ensure high-quality samples. For foundational tasks, we use procedurally synthesized and physics-simulated audio. For holistic data, we follow a four-stage process that includes human annotation and final selection based on human performance. Unlike prior benchmarks where caption-only answering reduces accuracy slightly, STAR-Bench induces far larger drops (-31.5 % temporal, -35.2 % spatial), evidencing its focus on linguistically hard-to-describe cues. Evaluating 19 models reveals substantial gaps compared with humans and a capability hierarchy: closed-source models are bottlenecked by fine-grained perception, while open-source models lag across perception, knowledge, and reasoning. Our STAR-Bench provides critical insights and a clear path forward for developing future models with a more robust understanding of the physical world.


【2】Audio Signal Processing Using Time Domain Mel-Frequency Wavelet Coefficient
标题:利用时间域Mel频率子波系数的音频信号处理
链接:http://arxiv.org/pdf/2510.24519v1

作者:Rinku Sebastian, Simon O'Keefe, Martin Trefzer
摘要:语音特征提取是语音信号处理中最关键的过程。梅尔频率倒谱系数(MFCC)是大多数说话人和语音识别应用中最广泛使用的特征,因为该特征中的滤波类似于人耳中发生的滤波。但是该特征的主要缺点是它仅提供信号的频率信息,而不提供关于在什么时间出现哪个频率的信息。小波变换具有灵活的时频窗,提供了信号的时频信息,是分析语音等非平稳信号的合适工具。另一方面,由于其均匀的频率缩放,典型的小波变换在分析语音信号时可能不太有效,在低频中具有较差的频率分辨率,并且不太符合人的听觉感知。因此,有必要开发一个功能,结合MFCC和小波变换的优点。大量的研究试图将这两个特征结合起来。现有的基于小波变换的Mel尺度特征提取方法在Mel尺度滤波的基础上再进行小波变换时,由于增加了额外的处理步骤,计算量较大。本文结合小波变换的概念,提出了一种在时域提取Mel尺度特征的方法,从而降低了时频转换的计算量和小波提取的复杂度。将时域梅尔频率小波系数(TMFWC)技术与水库计算方法相结合,显著提高了音频信号处理的效率。摘要:Extracting features from the speech is the most critical process in speech signal processing. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used features in the majority of the speaker and speech recognition applications, as the filtering in this feature is similar to the filtering taking place in the human ear. But the main drawback of this feature is that it provides only the frequency information of the signal but does not provide the information about at what time which frequency is present. The wavelet transform, with its flexible time-frequency window, provides time and frequency information of the signal and is an appropriate tool for the analysis of non-stationary signals like speech. On the other hand, because of its uniform frequency scaling, a typical wavelet transform may be less effective in analysing speech signals, have poorer frequency resolution in low frequencies, and be less in line with human auditory perception. Hence, it is necessary to develop a feature that incorporates the merits of both MFCC and wavelet transform. A great deal of studies are trying to combine both these features. The present Wavelet Transform based Mel-scaled feature extraction methods require more computation when a wavelet transform is applied on top of Mel-scale filtering, since it adds extra processing steps. Here we are proposing a method to extract Mel scale features in time domain combining the concept of wavelet transform, thus reducing the computational burden of time-frequency conversion and the complexity of wavelet extraction. Combining our proposed Time domain Mel frequency Wavelet Coefficient(TMFWC) technique with the reservoir computing methodology has significantly improved the efficiency of audio signal processing.


【3】Online neural fusion of distortionless differential beamformers for robust speech enhancement
标题:无失真差异束形成器的在线神经融合实现鲁棒语音增强
链接:http://arxiv.org/pdf/2510.24497v1

作者:Yuanhang Qian, Kunlong Zhao, Jilu Jin, Xueqin Luo, Gongping Huang, Jingdong Chen, Jacob Benesty
摘要:固定波束形成在实践中被广泛使用,因为它不依赖于噪声统计的估计并且提供相对稳定的性能。然而,单个波束形成器不能适应变化的声学条件,这限制了其干扰抑制能力。为了解决这个问题,已经引入了自适应凸组合(ACC)算法,其中多个固定波束形成器的输出被线性组合以提高鲁棒性。然而,ACC经常在高度非平稳的场景中失败,例如快速移动的干扰,因为其自适应更新不能可靠地跟踪快速变化。为了克服这一限制,我们提出了一个框架在线神经融合框架的多个无失真差分波束形成器,通过神经网络估计的组合权重。与传统的自适应校正方法相比,该方法能更有效地适应动态声环境,在保持无失真约束的前提下实现更强的干扰抑制。摘要:Fixed beamforming is widely used in practice since it does not depend on the estimation of noise statistics and provides relatively stable performance. However, a single beamformer cannot adapt to varying acoustic conditions, which limits its interference suppression capability. To address this, adaptive convex combination (ACC) algorithms have been introduced, where the outputs of multiple fixed beamformers are linearly combined to improve robustness. Nevertheless, ACC often fails in highly non-stationary scenarios, such as rapidly moving interference, since its adaptive updates cannot reliably track rapid changes. To overcome this limitation, we propose a frame-online neural fusion framework for multiple distortionless differential beamformers, which estimates the combination weights through a neural network. Compared with conventional ACC, the proposed method adapts more effectively to dynamic acoustic environments, achieving stronger interference suppression while maintaining the distortionless constraint.


【4】Your Microphone Array Retains Your Identity: A Robust Voice Liveness Detection System for Smart Speakers
标题:您的麦克风阵列保留您的身份:一个强大的智能扬声器语音活跃度检测系统
链接:http://arxiv.org/pdf/2510.24393v1

作者:Yan Meng, Jiachun Li, Matthew Pillari, Arjun Deopujari, Liam Brennan, Hafsah Shamsie, Haojin Zhu, Yuan Tian

备注:This is a paper accepted by USENIX Security 2022. See: this https URL

摘要:虽然智能音箱在智能家居系统中发挥着重要作用,但它很容易受到语音欺骗攻击。被动活体检测,它只利用收集的音频,而不是部署的传感器来区分真人和重放的声音,引起了越来越多的关注。然而,它面临着性能下降的挑战下,不同的环境因素,以及严格的要求,固定的用户手势。 在这项研究中,我们提出了一种新的活性特征,阵列指纹,它利用智能扬声器固有的麦克风阵列来确定收集的音频的身份。我们的理论分析表明,通过利用麦克风的圆形布局,与现有的方案相比,阵列指纹实现了更强大的性能下的环境变化和用户的运动。然后,利用这样的指纹,我们提出了阵列,一个轻量级的被动检测方案,并阐述了一系列的功能与阵列指纹。我们对包含32,780个音频样本和14个欺骗设备的数据集进行的评估表明,ARRAYID的准确率达到99.84%,优于现有的被动活性检测方案。摘要:Though playing an essential role in smart home systems, smart speakers are vulnerable to voice spoofing attacks. Passive liveness detection, which utilizes only the collected audio rather than the deployed sensors to distinguish between live-human and replayed voices, has drawn increasing attention. However, it faces the challenge of performance degradation under the different environmental factors as well as the strict requirement of the fixed user gestures. In this study, we propose a novel liveness feature, array fingerprint, which utilizes the microphone array inherently adopted by the smart speaker to determine the identity of collected audios. Our theoretical analysis demonstrates that by leveraging the circular layout of microphones, compared with existing schemes, array fingerprint achieves a more robust performance under the environmental change and user's movement. Then, to leverage such a fingerprint, we propose ARRAYID, a lightweight passive detection scheme, and elaborate a series of features working together with array fingerprint. Our evaluation on the dataset containing 32,780 audio samples and 14 spoofing devices shows that ARRAYID achieves an accuracy of 99.84%, which is superior to existing passive liveness detection schemes.


【5】Bayesian Speech synthesizers Can Learn from Multiple Teachers
标题:Bayesian语音合成器可以向多位老师学习
链接:http://arxiv.org/pdf/2510.24372v1

作者:Ziyang Zhang, Yifan Gao, Xuenan Xu, Baoxiangli, Wen Wu, Chao Zhang
摘要:基于编解码器的文本到语音(TTS)模型最近因其在语音克隆中的效率和强大性能而受到关注。然而,基于编解码器的TTS面临的限制,由于预训练鲁棒的语音编解码器和量化误差引入的质量下降的挑战。新出现的证据表明,连续值生成模型可以缓解这些问题,并作为一个有前途的替代方案。然而,有效地模拟不同的语音模式和开发可靠的采样策略的连续值自回归(AR)TTS仍然未被探索。在这项工作中,我们提出了BELLE,贝叶斯证据学习与语言建模TTS,一种新的连续值AR框架,直接预测梅尔频谱从文本输入。BELLE将每个梅尔频谱图帧视为从学习的超分布中采样的高斯分布,从而实现原则性的不确定性估计,特别是在具有并行数据的场景中(即,一个文本音频提示与多个语音样本配对)。为了获得这样的数据,不同的语音样本合成使用多个预先训练的TTS模型给定相同的文本音频提示,这是通过贝叶斯证据学习提炼成BELLE。实验结果表明,与目前最好的开源TTS模型相比,BELLE表现出很强的竞争力,即使BELLE是在大量的合成数据上训练的,并且只使用了大约十分之一的训练数据。贝尔生成的音频样本可在https: belletts.github.io Belle 上获得。论文被接受后,将发布代码、检查点和合成数据。摘要:Codec-based text-to-speech (TTS) models have recently gained traction for their efficiency and strong performance in voice cloning. However, codec-based TTS faces limitations due to the challenges of pretraining robust speech codecs and the quality degradation introduced by quantization errors. Emerging evidence suggests that continuous-valued generative models can alleviate these issues and serve as a promising alternative. Yet, effectively modelling diverse speech patterns and developing reliable sampling strategies for continuous-valued autoregressive (AR) TTS remains underexplored. In this work, we propose BELLE, Bayesian evidential learning with language modelling for TTS, a novel continuous-valued AR framework that directly predicts mel-spectrograms from textual input. BELLE treats each mel-spectrogram frame as a Gaussian distribution sampled from a learned hyper distribution, enabling principled uncertainty estimation, particularly in scenarios with parallel data (i.e., one text-audio prompt paired with multiple speech samples). To obtain such data, diverse speech samples are synthesized using multiple pre-trained TTS models given the same text-audio prompts, which are distilled into BELLE via Bayesian evidential learning. Experimental results indicate that BELLE demonstrates highly competitive performance compared with the current best open-source TTS models, even though BELLE is trained on a large amount of synthetic data and uses only approximately one-tenth of their training data. Audio samples generated by BELLE are available at https: belletts.github.io Belle . The code, checkpoints, and synthetic data will be released after the paper is accepted.


【6】Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes
标题:动态场景中手术动作空间映射的光源定位
链接:http://arxiv.org/pdf/2510.24332v1

作者:Jonas Hein, Lazaros Vlachopoulos, Maurits Geert Laurent Olthof, Bastian Sigrist, Philipp Fürnstahl, Matthias Seibold
摘要:目的:手术场景理解是推进计算机辅助和智能手术系统的关键。目前的方法主要依赖于可视化数据或端到端学习,这限制了细粒度的上下文建模。这项工作的目的是通过整合3D声学信息来增强手术场景表示,从而实现对手术环境的时间和空间感知多模态理解。 研究方法:我们提出了一种新的框架,用于通过将来自相控麦克风阵列的声学定位信息投影到来自RGB-D相机的动态点云上来生成手术场景的4D视听表示。基于变换器的声学事件检测模块识别包含工具-组织相互作用的相关时间片段,所述相关时间片段在视听场景表示中空间定位。在专家进行的模拟外科手术期间,在真实的手术室设置中对该系统进行了实验性评价。 结果:所提出的方法成功地定位在三维空间中的手术声事件,并将它们与视觉场景元素。实验评估表明,准确的空间声音定位和多模态数据的鲁棒融合,提供了一个全面的,动态的手术活动表示。 结论:这项工作介绍了第一种方法,在动态手术场景的空间声音定位,标志着一个显着的进步,多模式手术场景表示。通过整合声学和视觉数据,所提出的框架能够实现更丰富的上下文理解,并为未来的智能和自主手术系统提供基础。摘要:Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work aims to enhance surgical scene representations by integrating 3D acoustic information, enabling temporally and spatially aware multimodal understanding of surgical environments. Methods: We propose a novel framework for generating 4D audio-visual representations of surgical scenes by projecting acoustic localization information from a phased microphone array onto dynamic point clouds from an RGB-D camera. A transformer-based acoustic event detection module identifies relevant temporal segments containing tool-tissue interactions which are spatially localized in the audio-visual scene representation. The system was experimentally evaluated in a realistic operating room setup during simulated surgical procedures performed by experts. Results: The proposed method successfully localizes surgical acoustic events in 3D space and associates them with visual scene elements. Experimental evaluation demonstrates accurate spatial sound localization and robust fusion of multimodal data, providing a comprehensive, dynamic representation of surgical activity. Conclusion: This work introduces the first approach for spatial sound localization in dynamic surgical scenes, marking a significant advancement toward multimodal surgical scene representations. By integrating acoustic and visual data, the proposed framework enables richer contextual understanding and provides a foundation for future intelligent and autonomous surgical systems.


【7】TsetlinKWS: A 65nm 16.58uW, 0.63mm2 State-Driven Convolutional Tsetlin Machine-Based Accelerator For Keyword Spotting
标题:TsetlinKWS:65纳米16.58uW、0.63mm2状态驱动的卷积Tsetlin机器加速器,用于关键词发现
链接:http://arxiv.org/pdf/2510.24282v1

作者:Baizhou Lin, Yuetong Fang, Renjing Xu, Rishad Shafik, Jagmohan Chauhan
备注:12 pages, 17 figures. This work has been submitted to the IEEE for possible publication
摘要:Tsetlin Machine(TM)最近作为神经网络的低功耗替代方案引起了人们的关注,因为它具有简单和可解释的推理机制。然而,它在语音相关任务上的表现仍然有限。本文提出了TsetlinKWS,第一个算法硬件协同设计框架的卷积Tsetlin机(CTM)的12个关键字定位任务。首先,我们引入了一种新的梅尔频率谱系数和谱通量(MFSC-SF)的特征提取方案,连同频谱卷积,使CTM达到其有史以来第一次有竞争力的准确率为87.35%的12个关键字的定位任务。其次,我们开发了一个优化分组块压缩稀疏行(OG-BCSR)算法,实现了显着的9.84times $减少模型大小,显着提高存储效率的CTM。最后,我们提出了一个状态驱动的体系结构量身定制的CTM,同时利用数据重用和稀疏性,以实现高能源效率。整个系统采用65 nm工艺技术进行评估,在0.7 V时消耗16.58 $ mu$W,具有紧凑的0.63 mm$^2$核心面积。TsetlinKWS每次推理只需要907 k逻辑运算,与最先进的KWS加速器相比减少了1000倍,将CTM定位为超低功耗语音应用的高效候选者。摘要:The Tsetlin Machine (TM) has recently attracted attention as a low-power alternative to neural networks due to its simple and interpretable inference mechanisms. However, its performance on speech-related tasks remains limited. This paper proposes TsetlinKWS, the first algorithm-hardware co-design framework for the Convolutional Tsetlin Machine (CTM) on the 12-keyword spotting task. Firstly, we introduce a novel Mel-Frequency Spectral Coefficient and Spectral Flux (MFSC-SF) feature extraction scheme together with spectral convolution, enabling the CTM to reach its first-ever competitive accuracy of 87.35% on the 12-keyword spotting task. Secondly, we develop an Optimized Grouped Block-Compressed Sparse Row (OG-BCSR) algorithm that achieves a remarkable 9.84$ times$ reduction in model size, significantly improving the storage efficiency on CTMs. Finally, we propose a state-driven architecture tailored for the CTM, which simultaneously exploits data reuse and sparsity to achieve high energy efficiency. The full system is evaluated in 65 nm process technology, consuming 16.58 $ mu$W at 0.7 V with a compact 0.63 mm$^2$ core area. TsetlinKWS requires only 907k logic operations per inference, representing a 10$ times$ reduction compared to the state-of-the-art KWS accelerators, positioning the CTM as a highly-efficient candidate for ultra-low-power speech applications.


【8】HergNet: a Fast Neural Surrogate Model for Sound Field Predictions via Superposition of Plane Waves
标题:HergNet:一种通过平面波叠加进行声学预测的快速神经代理模型
链接:http://arxiv.org/pdf/2510.24279v1

作者:Matteo Calafà, Yuanxin Xia, Cheol-Ho Jeong
摘要:我们提出了一种新的神经网络架构,用于二维和三维声场的有效预测。该网络被设计为自动满足亥姆霍兹方程,确保输出在物理上有效。因此,该方法可以有效地学习解决各种波动现象中的边值问题,如声学,光学和电磁学。数值实验表明,所提出的策略可以潜在地优于国家的最先进的方法在室内声学模拟,特别是在中高频率范围。摘要:We present a novel neural network architecture for the efficient prediction of sound fields in two and three dimensions. The network is designed to automatically satisfy the Helmholtz equation, ensuring that the outputs are physically valid. Therefore, the method can effectively learn solutions to boundary-value problems in various wave phenomena, such as acoustics, optics, and electromagnetism. Numerical experiments show that the proposed strategy can potentially outperform state-of-the-art methods in room acoustics simulation, in particular in the range of mid to high frequencies.


【9】Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
标题:模型引导的双角色对齐,用于高保真开放域视频到音频生成
链接:http://arxiv.org/pdf/2510.24103v1

作者:Kang Zhang, Trung X. Pham, Suyeon Lee, Axi Niu, Arda Senocak, Joon Son Chung
备注:accepted by NeurIPS 2025
摘要:我们提出了MGAudio,一种新的基于流的框架开放域的视频到音频生成,它引入了模型引导的双重角色对齐作为中心设计原则。与依赖于基于分类器或无分类器的指导的先前方法不同,MGAudio使生成模型能够通过为视频调节音频生成而设计的专用训练目标来指导自己。该框架集成了三个主要组件:(1)一个可扩展的基于流的Transformer模型,(2)一个双重角色的对齐机制,其中视听编码器既作为一个条件模块,又作为一个功能对齐器,以提高生成质量,以及(3)一个模型引导的目标,增强跨模态的连贯性和音频的真实感。MGAudio在VGGSound上实现了最先进的性能,将FAD降低到0.40,大大超过了最好的无分类器指导基线,并在FD,IS和对齐指标上始终优于现有方法。它也很好地概括了具有挑战性的UnAV-100基准。这些结果突出了模型引导的双重角色对齐作为一个强大的和可扩展的范例有条件的视频到音频生成。代码可从以下网址获得:https: github.com pantheon5100 mgaudio摘要:We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guide itself through a dedicated training objective designed for video-conditioned audio generation. The framework integrates three main components: (1) a scalable flow-based Transformer model, (2) a dual-role alignment mechanism where the audio-visual encoder serves both as a conditioning module and as a feature aligner to improve generation quality, and (3) a model-guided objective that enhances cross-modal coherence and audio realism. MGAudio achieves state-of-the-art performance on VGGSound, reducing FAD to 0.40, substantially surpassing the best classifier-free guidance baselines, and consistently outperforms existing methods across FD, IS, and alignment metrics. It also generalizes well to the challenging UnAV-100 benchmark. These results highlight model-guided dual-role alignment as a powerful and scalable paradigm for conditional video-to-audio generation. Code is available at: https: github.com pantheon5100 mgaudio


【10】emg2speech: synthesizing speech from electromyography using self-supervised speech models
标题:emg 2 speech:使用自我监督语音模型从肌电合成语音
链接:http://arxiv.org/pdf/2510.23969v1

作者:Harshavardhana T. Gowda, Lee M. Miller
摘要:我们提出了一个神经肌肉语音接口,翻译肌电图(EMG)从口面部肌肉在语音清晰度直接到音频收集的信号。我们发现,自我监督的语音(SS)表示表现出很强的线性关系与肌肉动作电位的电功率:SS功能可以线性映射到EMG功率的相关性为r = 0.85$。此外,对应于不同发音姿势的肌电功率向量在特征空间中形成结构化和可分离的聚类。这种关系:$ text{SS features}$ $ xrightarrow{ texttt{linear mapping}}$ $ text{EMG power}$ $ xrightarrow{ texttt{gestree-specific clustering}}$ $ text{articulatory movements}$,强调SS模型隐式编码发音机制。利用这一特性,我们直接将EMG信号映射到SS特征空间并合成语音,从而实现端到端的EMG到语音生成,而无需显式的发音模型和声码器训练。摘要:We present a neuromuscular speech interface that translates electromyographic (EMG) signals collected from orofacial muscles during speech articulation directly into audio. We show that self-supervised speech (SS) representations exhibit a strong linear relationship with the electrical power of muscle action potentials: SS features can be linearly mapped to EMG power with a correlation of $r = 0.85$. Moreover, EMG power vectors corresponding to different articulatory gestures form structured and separable clusters in feature space. This relationship: $ text{SS features}$ $ xrightarrow{ texttt{linear mapping}}$ $ text{EMG power}$ $ xrightarrow{ texttt{gesture-specific clustering}}$ $ text{articulatory movements}$, highlights that SS models implicitly encode articulatory mechanisms. Leveraging this property, we directly map EMG signals to SS feature space and synthesize speech, enabling end-to-end EMG-to-speech generation without explicit articulatory models and vocoder training.


【11】Optimized Loudspeaker Panning for Adaptive Sound-Field Correction and Non-stationary Listening Areas
标题:优化扬声器平移以实现自适应磁场纠正和非静止收听区域
链接:http://arxiv.org/pdf/2510.23937v1

作者:Yuancheng Luo

Journal-ref:Luo, Yuancheng; Optimized Loudspeaker Panning for Adaptive Sound-Field Correction and Non-stationary Listening Areas; AES Long Beach: 159th Audio Engineering Society Convention 2025; Paper 385

摘要:环绕声系统通常沿着用于多声道音频再现的标准化布局分布扬声器。然而,在较少控制的环境中,实际布局在扬声器数量、放置和收听位置 区域方面变化。与标准布局的偏差会引入声场误差,从而降低音频内容再现的音质、成像和清晰度。这项工作介绍了贝叶斯扬声器归一化和内容平移优化方法的声场校正。共轭先验分布在不同的扬声器-收听者方向上更新非静止收听位置的估计布局;数字滤波器在没有声学测量的情况下使扬声器声学响应适应于估计收听区域处的共同参考目标。频域平移系数然后通过受空间、电和声学域约束的灵敏度 效率目标进行优化;标准化和平移的扬声器形成标准化布局中的虚拟扬声器,以实现准确的多声道再现。实验研究了贝叶斯自适应的鲁棒性,并在实际应用中进行了平移优化。摘要:Surround sound systems commonly distribute loudspeakers along standardized layouts for multichannel audio reproduction. However in less controlled environments, practical layouts vary in loudspeaker quantity, placement, and listening locations areas. Deviations from standard layouts introduce sound-field errors that degrade acoustic timbre, imaging, and clarity of audio content reproduction. This work introduces both Bayesian loudspeaker normalization and content panning optimization methods for sound-field correction. Conjugate prior distributions over loudspeaker-listener directions update estimated layouts for non-stationary listening locations; digital filters adapt loudspeaker acoustic responses to a common reference target at the estimated listening area without acoustic measurements. Frequency-domain panning coefficients are then optimized via sensitivity efficiency objectives subject to spatial, electrical, and acoustic domain constraints; normalized and panned loudspeakers form virtual loudspeakers in standardized layouts for accurate multichannel reproduction. Experiments investigate robustness of Bayesian adaptation, and panning optimizations in practical applications.


【12】A Neural Model for Contextual Biasing Score Learning and Filtering
标题:一种用于上下文偏差分数学习和过滤的神经模型
链接:http://arxiv.org/pdf/2510.23849v1

作者:Wanting Huang, Weiran Wang
备注:Accepted to IEEE ASRU 2025
摘要:上下文偏置通过在解码期间集成外部知识(诸如用户特定的短语或实体)来改进自动语音识别(ASR)。在这项工作中,我们使用一个基于注意力的偏置解码器来产生分数候选短语的基础上提取的声学信息的ASR编码器,它可以用来过滤掉不太可能的短语,并计算奖金浅融合偏置。我们引入了一个每令牌的歧视性目标,鼓励更高的分数地面真理短语,同时抑制干扰。在Librispeech偏置基准上的实验表明,该方法有效地滤除了大部分候选短语,并在不同偏置条件下显著提高了识别精度。我们的方法是模块化的,可以与任何ASR系统一起使用,并且过滤机制可以潜在地提高其他偏置方法的性能。摘要:Contextual biasing improves automatic speech recognition (ASR) by integrating external knowledge, such as user-specific phrases or entities, during decoding. In this work, we use an attention-based biasing decoder to produce scores for candidate phrases based on acoustic information extracted by an ASR encoder, which can be used to filter out unlikely phrases and to calculate bonus for shallow-fusion biasing. We introduce a per-token discriminative objective that encourages higher scores for ground-truth phrases while suppressing distractors. Experiments on the Librispeech biasing benchmark show that our method effectively filters out majority of the candidate phrases, and significantly improves recognition accuracy under different biasing conditions when the scores are used in shallow fusion biasing. Our approach is modular and can be used with any ASR system, and the filtering mechanism can potentially boost performance of other biasing methods.


eess.AS音频处理


【1】Forward Convolutive Prediction for Frame Online Monaural Speech Dereverberation Based on Kronecker Product Decomposition
标题:基于Kronecker积分解的帧在线单耳语音去回响前向卷积预测
链接:http://arxiv.org/pdf/2510.24471v1

作者:Yujie Zhu, Jilu Jin, Xueqin Luo, Wenxing Yang, Zhong-Qiu Wang, Gongping Huang, Jingdong Chen, Jacob Benesty
摘要:混响抑制一直是语音处理领域的一个重要研究课题,其目的是减轻混响对语音通信和语音交互系统的不利影响。在现有的方法中,前向卷积预测(FCP)最近引起了人们的注意。它通常采用深度神经网络来预测直接路径信号,并随后估计线性预测滤波器以抑制残余混响。然而,这种方法的主要缺点是所需的线性预测滤波器通常过长,导致相当大的计算复杂度。为了解决这个问题,我们的工作提出了一种新的FCP方法的基础上克罗内克积(KP)分解,其中长预测滤波器建模为两个更短的滤波器的KP。这种分解大大降低了计算成本。然后,提供了一种自适应算法来迭代地在线更新这些较短的滤波器。实验结果表明,与传统的方法相比,我们的方法实现了竞争力的去混响性能,同时大大降低了计算成本。摘要:Dereverberation has long been a crucial research topic in speech processing, aiming to alleviate the adverse effects of reverberation in voice communication and speech interaction systems. Among existing approaches, forward convolutional prediction (FCP) has recently attracted attention. It typically employs a deep neural network to predict the direct-path signal and subsequently estimates a linear prediction filter to suppress residual reverberation. However, a major drawback of this approach is that the required linear prediction filter is often excessively long, leading to considerable computational complexity. To address this, our work proposes a novel FCP method based on Kronecker product (KP) decomposition, in which the long prediction filter is modeled as the KP of two much shorter filters. This decomposition significantly reduces the computational cost. An adaptive algorithm is then provided to iteratively update these shorter filters online. Experimental results show that, compared to conventional methods, our approach achieves competitive dereverberation performance while substantially reducing computational cost.


【2】Listening without Looking: Modality Bias in Audio-Visual Captioning
标题:听而不看:视听字幕中的情态偏差
链接:http://arxiv.org/pdf/2510.24024v1

作者:Yuchi Ishikawa, Toranosuke Manabe, Tatsuya Komatsu, Yoshimitsu Aoki
备注:under review
摘要:视听字幕旨在通过联合建模声音和视觉来生成整体场景描述。虽然最近的方法通过复杂的模态融合提高了性能,但目前仍不清楚这两种模态在当前视听字幕模型中的互补程度以及当一种模态降级时这些模型的鲁棒性如何。我们通过对LAVCap(一种最先进的视听字幕模型)进行系统的模态鲁棒性测试来解决这些问题,在该模型中,我们选择性地抑制或破坏音频或视频流以量化敏感性和互补性。分析表明,在LAVCap的音频流的一个明显的偏见。为了评估视听字幕模型在使用这两种模式时的平衡程度,我们使用文本注释来增强AudioCaps,这些文本注释共同描述了音频和视频流,从而生成AudioVisualCaps数据集。在我们的实验中,我们在AudioVisualCaps上报告LAVCap基线结果。我们还在AudioVisualCaps的模态鲁棒性测试下评估了模型,结果表明,在AudioVisualCaps上训练的LAVCap比在AudioCaps上训练时表现出更少的模态偏差。摘要:Audio-visual captioning aims to generate holistic scene descriptions by jointly modeling sound and vision. While recent methods have improved performance through sophisticated modality fusion, it remains unclear to what extent the two modalities are complementary in current audio-visual captioning models and how robust these models are when one modality is degraded. We address these questions by conducting systematic modality robustness tests on LAVCap, a state-of-the-art audio-visual captioning model, in which we selectively suppress or corrupt the audio or visual streams to quantify sensitivity and complementarity. The analysis reveals a pronounced bias toward the audio stream in LAVCap. To evaluate how balanced audio-visual captioning models are in their use of both modalities, we augment AudioCaps with textual annotations that jointly describe the audio and visual streams, yielding the AudioVisualCaps dataset. In our experiments, we report LAVCap baseline results on AudioVisualCaps. We also evaluate the model under modality robustness tests on AudioVisualCaps and the results indicate that LAVCap trained on AudioVisualCaps exhibits less modality bias than when trained on AudioCaps.


【3】A Neural Model for Contextual Biasing Score Learning and Filtering
标题:一种用于上下文偏差分数学习和过滤的神经模型
链接:http://arxiv.org/pdf/2510.23849v1

作者:Wanting Huang, Weiran Wang
备注:Accepted to IEEE ASRU 2025
摘要:上下文偏置通过在解码期间集成外部知识(诸如用户特定的短语或实体)来改进自动语音识别(ASR)。在这项工作中,我们使用一个基于注意力的偏置解码器来产生分数候选短语的基础上提取的声学信息的ASR编码器,它可以用来过滤掉不太可能的短语,并计算奖金浅融合偏置。我们引入了一个每令牌的歧视性目标,鼓励更高的分数地面真理短语,同时抑制干扰。在Librispeech偏置基准上的实验表明,该方法有效地滤除了大部分候选短语,并在不同偏置条件下显著提高了识别精度。我们的方法是模块化的,可以与任何ASR系统一起使用,并且过滤机制可以潜在地提高其他偏置方法的性能。摘要:Contextual biasing improves automatic speech recognition (ASR) by integrating external knowledge, such as user-specific phrases or entities, during decoding. In this work, we use an attention-based biasing decoder to produce scores for candidate phrases based on acoustic information extracted by an ASR encoder, which can be used to filter out unlikely phrases and to calculate bonus for shallow-fusion biasing. We introduce a per-token discriminative objective that encourages higher scores for ground-truth phrases while suppressing distractors. Experiments on the Librispeech biasing benchmark show that our method effectively filters out majority of the candidate phrases, and significantly improves recognition accuracy under different biasing conditions when the scores are used in shallow fusion biasing. Our approach is modular and can be used with any ASR system, and the filtering mechanism can potentially boost performance of other biasing methods.


【4】STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
标题:STAR-Bench:探索作为音频4D智能的深度时空推理
链接:http://arxiv.org/pdf/2510.24693v1

作者:Zihan Liu, Zhikang Niu, Qiuyang Xiao, Zhisheng Zheng, Ruoqi Yuan, Yuhang Zang, Yuhang Cao, Xiaoyi Dong, Jianze Liang, Xie Chen, Leilei Sun, Dahua Lin, Jiaqi Wang
备注:Homepage: this https URL
摘要:尽管多模态大型语言模型和大型音频语言模型取得了快速进展,但现有的音频基准主要测试可以从文本标题中恢复的语义,掩盖了细粒度感知推理的缺陷。我们将音频4D智能形式化,定义为在时间和3D空间中对声音动态进行推理,并引入STAR-Bench来测量它。STAR-Bench结合了基础声学感知设置(绝对和相对制度下的六个属性)与整体时空推理设置,包括连续和离散过程的段重新排序和跨越静态定位,多源关系,和动态轨迹。我们的数据管理管道使用两种方法来确保高质量的样本。对于基础任务,我们使用程序合成和物理模拟音频。对于整体数据,我们遵循一个四阶段的过程,包括人类注释和基于人类表现的最终选择。与之前的基准测试不同,只有标题的回答稍微降低了准确性,STAR-Bench引起了更大的下降(-31.5 %时间,-35.2 %空间),证明了它对语言学上难以描述的线索的关注。对19个模型的评估揭示了与人类和能力层次相比的巨大差距:闭源模型受到细粒度感知的检验,而开源模型在感知,知识和推理方面落后。我们的STAR-Bench为开发未来模型提供了重要的见解和明确的前进道路,并对物理世界有更深入的了解。摘要:Despite rapid progress in Multi-modal Large Language Models and Large Audio-Language Models, existing audio benchmarks largely test semantics that can be recovered from text captions, masking deficits in fine-grained perceptual reasoning. We formalize audio 4D intelligence that is defined as reasoning over sound dynamics in time and 3D space, and introduce STAR-Bench to measure it. STAR-Bench combines a Foundational Acoustic Perception setting (six attributes under absolute and relative regimes) with a Holistic Spatio-Temporal Reasoning setting that includes segment reordering for continuous and discrete processes and spatial tasks spanning static localization, multi-source relations, and dynamic trajectories. Our data curation pipeline uses two methods to ensure high-quality samples. For foundational tasks, we use procedurally synthesized and physics-simulated audio. For holistic data, we follow a four-stage process that includes human annotation and final selection based on human performance. Unlike prior benchmarks where caption-only answering reduces accuracy slightly, STAR-Bench induces far larger drops (-31.5 % temporal, -35.2 % spatial), evidencing its focus on linguistically hard-to-describe cues. Evaluating 19 models reveals substantial gaps compared with humans and a capability hierarchy: closed-source models are bottlenecked by fine-grained perception, while open-source models lag across perception, knowledge, and reasoning. Our STAR-Bench provides critical insights and a clear path forward for developing future models with a more robust understanding of the physical world.


【5】Audio Signal Processing Using Time Domain Mel-Frequency Wavelet Coefficient
标题:利用时间域Mel频率子波系数的音频信号处理
链接:http://arxiv.org/pdf/2510.24519v1

作者:Rinku Sebastian, Simon O'Keefe, Martin Trefzer
摘要:语音特征提取是语音信号处理中最关键的过程。梅尔频率倒谱系数(MFCC)是大多数说话人和语音识别应用中最广泛使用的特征,因为该特征中的滤波类似于人耳中发生的滤波。但是该特征的主要缺点是它仅提供信号的频率信息,而不提供关于在什么时间出现哪个频率的信息。小波变换具有灵活的时频窗,提供了信号的时频信息,是分析语音等非平稳信号的合适工具。另一方面,由于其均匀的频率缩放,典型的小波变换在分析语音信号时可能不太有效,在低频中具有较差的频率分辨率,并且不太符合人的听觉感知。因此,有必要开发一个功能,结合MFCC和小波变换的优点。大量的研究试图将这两个特征结合起来。现有的基于小波变换的Mel尺度特征提取方法在Mel尺度滤波的基础上再进行小波变换时,由于增加了额外的处理步骤,计算量较大。本文结合小波变换的概念,提出了一种在时域提取Mel尺度特征的方法,从而降低了时频转换的计算量和小波提取的复杂度。将时域梅尔频率小波系数(TMFWC)技术与水库计算方法相结合,显著提高了音频信号处理的效率。摘要:Extracting features from the speech is the most critical process in speech signal processing. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used features in the majority of the speaker and speech recognition applications, as the filtering in this feature is similar to the filtering taking place in the human ear. But the main drawback of this feature is that it provides only the frequency information of the signal but does not provide the information about at what time which frequency is present. The wavelet transform, with its flexible time-frequency window, provides time and frequency information of the signal and is an appropriate tool for the analysis of non-stationary signals like speech. On the other hand, because of its uniform frequency scaling, a typical wavelet transform may be less effective in analysing speech signals, have poorer frequency resolution in low frequencies, and be less in line with human auditory perception. Hence, it is necessary to develop a feature that incorporates the merits of both MFCC and wavelet transform. A great deal of studies are trying to combine both these features. The present Wavelet Transform based Mel-scaled feature extraction methods require more computation when a wavelet transform is applied on top of Mel-scale filtering, since it adds extra processing steps. Here we are proposing a method to extract Mel scale features in time domain combining the concept of wavelet transform, thus reducing the computational burden of time-frequency conversion and the complexity of wavelet extraction. Combining our proposed Time domain Mel frequency Wavelet Coefficient(TMFWC) technique with the reservoir computing methodology has significantly improved the efficiency of audio signal processing.


【6】Online neural fusion of distortionless differential beamformers for robust speech enhancement
标题:无失真差异束形成器的在线神经融合实现鲁棒语音增强
链接:http://arxiv.org/pdf/2510.24497v1

作者:Yuanhang Qian, Kunlong Zhao, Jilu Jin, Xueqin Luo, Gongping Huang, Jingdong Chen, Jacob Benesty
摘要:固定波束形成在实践中被广泛使用,因为它不依赖于噪声统计的估计并且提供相对稳定的性能。然而,单个波束形成器不能适应变化的声学条件,这限制了其干扰抑制能力。为了解决这个问题,已经引入了自适应凸组合(ACC)算法,其中多个固定波束形成器的输出被线性组合以提高鲁棒性。然而,ACC经常在高度非平稳的场景中失败,例如快速移动的干扰,因为其自适应更新不能可靠地跟踪快速变化。为了克服这一限制,我们提出了一个框架在线神经融合框架的多个无失真差分波束形成器,通过神经网络估计的组合权重。与传统的自适应校正方法相比,该方法能更有效地适应动态声环境,在保持无失真约束的前提下实现更强的干扰抑制。摘要:Fixed beamforming is widely used in practice since it does not depend on the estimation of noise statistics and provides relatively stable performance. However, a single beamformer cannot adapt to varying acoustic conditions, which limits its interference suppression capability. To address this, adaptive convex combination (ACC) algorithms have been introduced, where the outputs of multiple fixed beamformers are linearly combined to improve robustness. Nevertheless, ACC often fails in highly non-stationary scenarios, such as rapidly moving interference, since its adaptive updates cannot reliably track rapid changes. To overcome this limitation, we propose a frame-online neural fusion framework for multiple distortionless differential beamformers, which estimates the combination weights through a neural network. Compared with conventional ACC, the proposed method adapts more effectively to dynamic acoustic environments, achieving stronger interference suppression while maintaining the distortionless constraint.


【7】Your Microphone Array Retains Your Identity: A Robust Voice Liveness Detection System for Smart Speakers
标题:您的麦克风阵列保留您的身份:一个强大的智能扬声器语音活跃度检测系统
链接:http://arxiv.org/pdf/2510.24393v1

作者:Yan Meng, Jiachun Li, Matthew Pillari, Arjun Deopujari, Liam Brennan, Hafsah Shamsie, Haojin Zhu, Yuan Tian

备注:This is a paper accepted by USENIX Security 2022. See: this https URL

摘要:虽然智能音箱在智能家居系统中发挥着重要作用,但它很容易受到语音欺骗攻击。被动活体检测,它只利用收集的音频,而不是部署的传感器来区分真人和重放的声音,引起了越来越多的关注。然而,它面临着性能下降的挑战下,不同的环境因素,以及严格的要求,固定的用户手势。 在这项研究中,我们提出了一种新的活性特征,阵列指纹,它利用智能扬声器固有的麦克风阵列来确定收集的音频的身份。我们的理论分析表明,通过利用麦克风的圆形布局,与现有的方案相比,阵列指纹实现了更强大的性能下的环境变化和用户的运动。然后,利用这样的指纹,我们提出了阵列,一个轻量级的被动检测方案,并阐述了一系列的功能与阵列指纹。我们对包含32,780个音频样本和14个欺骗设备的数据集进行的评估表明,ARRAYID的准确率达到99.84%,优于现有的被动活性检测方案。摘要:Though playing an essential role in smart home systems, smart speakers are vulnerable to voice spoofing attacks. Passive liveness detection, which utilizes only the collected audio rather than the deployed sensors to distinguish between live-human and replayed voices, has drawn increasing attention. However, it faces the challenge of performance degradation under the different environmental factors as well as the strict requirement of the fixed user gestures. In this study, we propose a novel liveness feature, array fingerprint, which utilizes the microphone array inherently adopted by the smart speaker to determine the identity of collected audios. Our theoretical analysis demonstrates that by leveraging the circular layout of microphones, compared with existing schemes, array fingerprint achieves a more robust performance under the environmental change and user's movement. Then, to leverage such a fingerprint, we propose ARRAYID, a lightweight passive detection scheme, and elaborate a series of features working together with array fingerprint. Our evaluation on the dataset containing 32,780 audio samples and 14 spoofing devices shows that ARRAYID achieves an accuracy of 99.84%, which is superior to existing passive liveness detection schemes.


【8】Bayesian Speech synthesizers Can Learn from Multiple Teachers
标题:Bayesian语音合成器可以向多位老师学习
链接:http://arxiv.org/pdf/2510.24372v1

作者:Ziyang Zhang, Yifan Gao, Xuenan Xu, Baoxiangli, Wen Wu, Chao Zhang
摘要:基于编解码器的文本到语音(TTS)模型最近因其在语音克隆中的效率和强大性能而受到关注。然而,基于编解码器的TTS面临的限制,由于预训练鲁棒的语音编解码器和量化误差引入的质量下降的挑战。新出现的证据表明,连续值生成模型可以缓解这些问题,并作为一个有前途的替代方案。然而,有效地模拟不同的语音模式和开发可靠的采样策略的连续值自回归(AR)TTS仍然未被探索。在这项工作中,我们提出了BELLE,贝叶斯证据学习与语言建模TTS,一种新的连续值AR框架,直接预测梅尔频谱从文本输入。BELLE将每个梅尔频谱图帧视为从学习的超分布中采样的高斯分布,从而实现原则性的不确定性估计,特别是在具有并行数据的场景中(即,一个文本音频提示与多个语音样本配对)。为了获得这样的数据,不同的语音样本合成使用多个预先训练的TTS模型给定相同的文本音频提示,这是通过贝叶斯证据学习提炼成BELLE。实验结果表明,与目前最好的开源TTS模型相比,BELLE表现出很强的竞争力,即使BELLE是在大量的合成数据上训练的,并且只使用了大约十分之一的训练数据。贝尔生成的音频样本可在https: belletts.github.io Belle 上获得。论文被接受后,将发布代码、检查点和合成数据。摘要:Codec-based text-to-speech (TTS) models have recently gained traction for their efficiency and strong performance in voice cloning. However, codec-based TTS faces limitations due to the challenges of pretraining robust speech codecs and the quality degradation introduced by quantization errors. Emerging evidence suggests that continuous-valued generative models can alleviate these issues and serve as a promising alternative. Yet, effectively modelling diverse speech patterns and developing reliable sampling strategies for continuous-valued autoregressive (AR) TTS remains underexplored. In this work, we propose BELLE, Bayesian evidential learning with language modelling for TTS, a novel continuous-valued AR framework that directly predicts mel-spectrograms from textual input. BELLE treats each mel-spectrogram frame as a Gaussian distribution sampled from a learned hyper distribution, enabling principled uncertainty estimation, particularly in scenarios with parallel data (i.e., one text-audio prompt paired with multiple speech samples). To obtain such data, diverse speech samples are synthesized using multiple pre-trained TTS models given the same text-audio prompts, which are distilled into BELLE via Bayesian evidential learning. Experimental results indicate that BELLE demonstrates highly competitive performance compared with the current best open-source TTS models, even though BELLE is trained on a large amount of synthetic data and uses only approximately one-tenth of their training data. Audio samples generated by BELLE are available at https: belletts.github.io Belle . The code, checkpoints, and synthetic data will be released after the paper is accepted.


【9】Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes
标题:动态场景中手术动作空间映射的光源定位
链接:http://arxiv.org/pdf/2510.24332v1

作者:Jonas Hein, Lazaros Vlachopoulos, Maurits Geert Laurent Olthof, Bastian Sigrist, Philipp Fürnstahl, Matthias Seibold
摘要:目的:手术场景理解是推进计算机辅助和智能手术系统的关键。目前的方法主要依赖于可视化数据或端到端学习,这限制了细粒度的上下文建模。这项工作的目的是通过整合3D声学信息来增强手术场景表示,从而实现对手术环境的时间和空间感知多模态理解。 研究方法:我们提出了一种新的框架,用于通过将来自相控麦克风阵列的声学定位信息投影到来自RGB-D相机的动态点云上来生成手术场景的4D视听表示。基于变换器的声学事件检测模块识别包含工具-组织相互作用的相关时间片段,所述相关时间片段在视听场景表示中空间定位。在专家进行的模拟外科手术期间,在真实的手术室设置中对该系统进行了实验性评价。 结果:所提出的方法成功地定位在三维空间中的手术声事件,并将它们与视觉场景元素。实验评估表明,准确的空间声音定位和多模态数据的鲁棒融合,提供了一个全面的,动态的手术活动表示。 结论:这项工作介绍了第一种方法,在动态手术场景的空间声音定位,标志着一个显着的进步,多模式手术场景表示。通过整合声学和视觉数据,所提出的框架能够实现更丰富的上下文理解,并为未来的智能和自主手术系统提供基础。摘要:Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work aims to enhance surgical scene representations by integrating 3D acoustic information, enabling temporally and spatially aware multimodal understanding of surgical environments. Methods: We propose a novel framework for generating 4D audio-visual representations of surgical scenes by projecting acoustic localization information from a phased microphone array onto dynamic point clouds from an RGB-D camera. A transformer-based acoustic event detection module identifies relevant temporal segments containing tool-tissue interactions which are spatially localized in the audio-visual scene representation. The system was experimentally evaluated in a realistic operating room setup during simulated surgical procedures performed by experts. Results: The proposed method successfully localizes surgical acoustic events in 3D space and associates them with visual scene elements. Experimental evaluation demonstrates accurate spatial sound localization and robust fusion of multimodal data, providing a comprehensive, dynamic representation of surgical activity. Conclusion: This work introduces the first approach for spatial sound localization in dynamic surgical scenes, marking a significant advancement toward multimodal surgical scene representations. By integrating acoustic and visual data, the proposed framework enables richer contextual understanding and provides a foundation for future intelligent and autonomous surgical systems.


【10】TsetlinKWS: A 65nm 16.58uW, 0.63mm2 State-Driven Convolutional Tsetlin Machine-Based Accelerator For Keyword Spotting
标题:TsetlinKWS:65纳米16.58uW、0.63mm2状态驱动的卷积Tsetlin机器加速器,用于关键词发现
链接:http://arxiv.org/pdf/2510.24282v1

作者:Baizhou Lin, Yuetong Fang, Renjing Xu, Rishad Shafik, Jagmohan Chauhan
备注:12 pages, 17 figures. This work has been submitted to the IEEE for possible publication
摘要:Tsetlin Machine(TM)最近作为神经网络的低功耗替代方案引起了人们的关注,因为它具有简单和可解释的推理机制。然而,它在语音相关任务上的表现仍然有限。本文提出了TsetlinKWS,第一个算法硬件协同设计框架的卷积Tsetlin机(CTM)的12个关键字定位任务。首先,我们引入了一种新的梅尔频率谱系数和谱通量(MFSC-SF)的特征提取方案,连同频谱卷积,使CTM达到其有史以来第一次有竞争力的准确率为87.35%的12个关键字的定位任务。其次,我们开发了一个优化分组块压缩稀疏行(OG-BCSR)算法,实现了显着的9.84times $减少模型大小,显着提高存储效率的CTM。最后,我们提出了一个状态驱动的体系结构量身定制的CTM,同时利用数据重用和稀疏性,以实现高能源效率。整个系统采用65 nm工艺技术进行评估,在0.7 V时消耗16.58 $ mu$W,具有紧凑的0.63 mm$^2$核心面积。TsetlinKWS每次推理只需要907 k逻辑运算,与最先进的KWS加速器相比减少了1000倍,将CTM定位为超低功耗语音应用的高效候选者。摘要:The Tsetlin Machine (TM) has recently attracted attention as a low-power alternative to neural networks due to its simple and interpretable inference mechanisms. However, its performance on speech-related tasks remains limited. This paper proposes TsetlinKWS, the first algorithm-hardware co-design framework for the Convolutional Tsetlin Machine (CTM) on the 12-keyword spotting task. Firstly, we introduce a novel Mel-Frequency Spectral Coefficient and Spectral Flux (MFSC-SF) feature extraction scheme together with spectral convolution, enabling the CTM to reach its first-ever competitive accuracy of 87.35% on the 12-keyword spotting task. Secondly, we develop an Optimized Grouped Block-Compressed Sparse Row (OG-BCSR) algorithm that achieves a remarkable 9.84$ times$ reduction in model size, significantly improving the storage efficiency on CTMs. Finally, we propose a state-driven architecture tailored for the CTM, which simultaneously exploits data reuse and sparsity to achieve high energy efficiency. The full system is evaluated in 65 nm process technology, consuming 16.58 $ mu$W at 0.7 V with a compact 0.63 mm$^2$ core area. TsetlinKWS requires only 907k logic operations per inference, representing a 10$ times$ reduction compared to the state-of-the-art KWS accelerators, positioning the CTM as a highly-efficient candidate for ultra-low-power speech applications.


【11】HergNet: a Fast Neural Surrogate Model for Sound Field Predictions via Superposition of Plane Waves
标题:HergNet:一种通过平面波叠加进行声学预测的快速神经代理模型
链接:http://arxiv.org/pdf/2510.24279v1

作者:Matteo Calafà, Yuanxin Xia, Cheol-Ho Jeong
摘要:我们提出了一种新的神经网络架构,用于二维和三维声场的有效预测。该网络被设计为自动满足亥姆霍兹方程,确保输出在物理上有效。因此,该方法可以有效地学习解决各种波动现象中的边值问题,如声学,光学和电磁学。数值实验表明,所提出的策略可以潜在地优于国家的最先进的方法在室内声学模拟,特别是在中高频率范围。摘要:We present a novel neural network architecture for the efficient prediction of sound fields in two and three dimensions. The network is designed to automatically satisfy the Helmholtz equation, ensuring that the outputs are physically valid. Therefore, the method can effectively learn solutions to boundary-value problems in various wave phenomena, such as acoustics, optics, and electromagnetism. Numerical experiments show that the proposed strategy can potentially outperform state-of-the-art methods in room acoustics simulation, in particular in the range of mid to high frequencies.


【12】Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
标题:模型引导的双角色对齐,用于高保真开放域视频到音频生成
链接:http://arxiv.org/pdf/2510.24103v1

作者:Kang Zhang, Trung X. Pham, Suyeon Lee, Axi Niu, Arda Senocak, Joon Son Chung
备注:accepted by NeurIPS 2025
摘要:我们提出了MGAudio,一种新的基于流的框架开放域的视频到音频生成,它引入了模型引导的双重角色对齐作为中心设计原则。与依赖于基于分类器或无分类器的指导的先前方法不同,MGAudio使生成模型能够通过为视频调节音频生成而设计的专用训练目标来指导自己。该框架集成了三个主要组件:(1)一个可扩展的基于流的Transformer模型,(2)一个双重角色的对齐机制,其中视听编码器既作为一个条件模块,又作为一个功能对齐器,以提高生成质量,以及(3)一个模型引导的目标,增强跨模态的连贯性和音频的真实感。MGAudio在VGGSound上实现了最先进的性能,将FAD降低到0.40,大大超过了最好的无分类器指导基线,并在FD,IS和对齐指标上始终优于现有方法。它也很好地概括了具有挑战性的UnAV-100基准。这些结果突出了模型引导的双重角色对齐作为一个强大的和可扩展的范例有条件的视频到音频生成。代码可从以下网址获得:https: github.com pantheon5100 mgaudio摘要:We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guide itself through a dedicated training objective designed for video-conditioned audio generation. The framework integrates three main components: (1) a scalable flow-based Transformer model, (2) a dual-role alignment mechanism where the audio-visual encoder serves both as a conditioning module and as a feature aligner to improve generation quality, and (3) a model-guided objective that enhances cross-modal coherence and audio realism. MGAudio achieves state-of-the-art performance on VGGSound, reducing FAD to 0.40, substantially surpassing the best classifier-free guidance baselines, and consistently outperforms existing methods across FD, IS, and alignment metrics. It also generalizes well to the challenging UnAV-100 benchmark. These results highlight model-guided dual-role alignment as a powerful and scalable paradigm for conditional video-to-audio generation. Code is available at: https: github.com pantheon5100 mgaudio


【13】emg2speech: synthesizing speech from electromyography using self-supervised speech models
标题:emg 2 speech:使用自我监督语音模型从肌电合成语音
链接:http://arxiv.org/pdf/2510.23969v1

作者:Harshavardhana T. Gowda, Lee M. Miller
摘要:我们提出了一个神经肌肉语音接口,翻译肌电图(EMG)从口面部肌肉在语音清晰度直接到音频收集的信号。我们发现,自我监督的语音(SS)表示表现出很强的线性关系与肌肉动作电位的电功率:SS功能可以线性映射到EMG功率的相关性为r = 0.85$。此外,对应于不同发音姿势的肌电功率向量在特征空间中形成结构化和可分离的聚类。这种关系:$ text{SS features}$ $ xrightarrow{ texttt{linear mapping}}$ $ text{EMG power}$ $ xrightarrow{ texttt{gestree-specific clustering}}$ $ text{articulatory movements}$,强调SS模型隐式编码发音机制。利用这一特性,我们直接将EMG信号映射到SS特征空间并合成语音,从而实现端到端的EMG到语音生成,而无需显式的发音模型和声码器训练。摘要:We present a neuromuscular speech interface that translates electromyographic (EMG) signals collected from orofacial muscles during speech articulation directly into audio. We show that self-supervised speech (SS) representations exhibit a strong linear relationship with the electrical power of muscle action potentials: SS features can be linearly mapped to EMG power with a correlation of $r = 0.85$. Moreover, EMG power vectors corresponding to different articulatory gestures form structured and separable clusters in feature space. This relationship: $ text{SS features}$ $ xrightarrow{ texttt{linear mapping}}$ $ text{EMG power}$ $ xrightarrow{ texttt{gesture-specific clustering}}$ $ text{articulatory movements}$, highlights that SS models implicitly encode articulatory mechanisms. Leveraging this property, we directly map EMG signals to SS feature space and synthesize speech, enabling end-to-end EMG-to-speech generation without explicit articulatory models and vocoder training.


【14】Optimized Loudspeaker Panning for Adaptive Sound-Field Correction and Non-stationary Listening Areas
标题:优化扬声器平移以实现自适应磁场纠正和非静止收听区域
链接:http://arxiv.org/pdf/2510.23937v1

作者:Yuancheng Luo

Journal-ref:Luo, Yuancheng; Optimized Loudspeaker Panning for Adaptive Sound-Field Correction and Non-stationary Listening Areas; AES Long Beach: 159th Audio Engineering Society Convention 2025; Paper 385

摘要:环绕声系统通常沿着用于多声道音频再现的标准化布局分布扬声器。然而,在较少控制的环境中,实际布局在扬声器数量、放置和收听位置 区域方面变化。与标准布局的偏差会引入声场误差,从而降低音频内容再现的音质、成像和清晰度。这项工作介绍了贝叶斯扬声器归一化和内容平移优化方法的声场校正。共轭先验分布在不同的扬声器-收听者方向上更新非静止收听位置的估计布局;数字滤波器在没有声学测量的情况下使扬声器声学响应适应于估计收听区域处的共同参考目标。频域平移系数然后通过受空间、电和声学域约束的灵敏度 效率目标进行优化;标准化和平移的扬声器形成标准化布局中的虚拟扬声器,以实现准确的多声道再现。实验研究了贝叶斯自适应的鲁棒性,并在实际应用中进行了平移优化。摘要:Surround sound systems commonly distribute loudspeakers along standardized layouts for multichannel audio reproduction. However in less controlled environments, practical layouts vary in loudspeaker quantity, placement, and listening locations areas. Deviations from standard layouts introduce sound-field errors that degrade acoustic timbre, imaging, and clarity of audio content reproduction. This work introduces both Bayesian loudspeaker normalization and content panning optimization methods for sound-field correction. Conjugate prior distributions over loudspeaker-listener directions update estimated layouts for non-stationary listening locations; digital filters adapt loudspeaker acoustic responses to a common reference target at the estimated listening area without acoustic measurements. Frequency-domain panning coefficients are then optimized via sensitivity efficiency objectives subject to spatial, electrical, and acoustic domain constraints; normalized and panned loudspeakers form virtual loudspeakers in standardized layouts for accurate multichannel reproduction. Experiments investigate robustness of Bayesian adaptation, and panning optimizations in practical applications.


机器翻译由腾讯交互翻译提供,仅供参考