今日论文合集:cs.SD语音14篇,eess.AS音频处理11篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
标题:迈向对话式匈牙利语音识别:引入BEA-Large和BEA-Dialogue数据集
链接:https://arxiv.org/pdf/2511.13529v1

作者:Máté Gedeon,Piroska Zsófia Barta,Péter Mihajlik,Tekla Etelka Gráczi,Anna Kohári,Katalin Mády
备注:Submitted to LREC 2026
摘要:自动语音识别(ASR)的进步在很大程度上得到了高资源语言的广泛数据集的增强,而匈牙利语等语言由于有限的自发和会话语料库而仍然代表性不足。为了解决这一差距,我们引入了两个新的数据集-- BEA-Large和BEA-Dialogue --它们是从名为BEA的匈牙利语音语料库中之前未处理的部分构建的。BEA-Large扩展了BEA-Base,提供了来自433位演讲者的255小时的自发演讲,并通过详细的片段级元数据进行了丰富。BEA-Dialogue包含85小时的自发对话,是一个匈牙利语语音语料库,其特征是将自然对话划分为与说话者无关的子集,支持会话ASR和说话者日记化的研究。我们使用公开可用的ASR模型在这些数据集上建立了可重复的基线,微调的Fast Conformer模型在自发语音和重复语音上的单词错误率分别低至14.18%和4.8%。实验结果表明,该算法的错误率在13.05%~ 18.26%之间,为进一步改进算法提供了参考。结果突出了会话ASR的持续困难,特别是由于不流利,重叠和非正式的语音模式。通过发布这些数据集和基线,我们的目标是推进匈牙利语语音技术,并为开发其他语言的自发和对话基准提供方法框架。摘要:The advancement of automatic speech recognition (ASR) has been largely enhanced by extensive datasets in high-resource languages, while languages such as Hungarian remain underrepresented due to limited spontaneous and conversational corpora. To address this gap, we introduce two new datasets -- BEA-Large and BEA-Dialogue -- constructed from the previously unprocessed portions of the Hungarian speech corpus named BEA. BEA-Large extends BEA-Base with 255 hours of spontaneous speech from 433 speakers, enriched with detailed segment-level metadata. BEA-Dialogue, comprising 85 hours of spontaneous conversations, is a Hungarian speech corpus featuring natural dialogues partitioned into speaker-independent subsets, supporting research in conversational ASR and speaker diarization. We establish reproducible baselines on these datasets using publicly available ASR models, with the fine-tuned Fast Conformer model achieving word error rates as low as 14.18 % on spontaneous and 4.8 % on repeated speech. Diarization experiments yield diarization error rates between 13.05 % and 18.26 %, providing reference points for future improvements. The results highlight the persistent difficulty of conversational ASR, particularly due to disfluencies, overlaps, and informal speech patterns. By releasing these datasets and baselines, we aim to advance Hungarian speech technology and offer a methodological framework for developing spontaneous and conversational benchmarks in other languages.


【2】Spatial Blind Spot: Auditory Motion Perception Deficits in Audio LLMs
标题:空间盲点:音频LLM中的听觉运动感知缺陷
链接:https://arxiv.org/pdf/2511.13273v1

作者:Zhe Sun,Yujun Cai,Jiayu Yao,Yiwei Wang
摘要:大型音频语言模型(LALM)最近在语音识别,音频字幕和听觉问答方面取得了令人印象深刻的进展。然而,这些模型是否可以感知空间动态,特别是声源的运动,仍然不清楚。在这项工作中,我们发现了一个系统的运动知觉缺陷,目前ALLM。为了研究这个问题,我们介绍了AMPBench,第一个明确设计用于评估听觉运动理解的基准。AMPBench引入了一个受控的问答基准测试,旨在评估音频语言模型(LALM)是否可以从双耳音频中推断出移动声源的方向和轨迹。全面的定量和定性分析表明,目前的模型难以可靠地识别运动线索或区分方向模式。平均准确率仍然低于50%,强调了听觉空间推理的基本限制。我们的研究突出了人类和模型听觉空间推理之间的根本差距,为未来的音频语言模型中增强空间认知提供了诊断工具和新的见解。摘要:Large Audio-Language Models (LALMs) have recently shown impressive progress in speech recognition, audio captioning, and auditory question answering. Yet, whether these models can perceive spatial dynamics, particularly the motion of sound sources, remains unclear. In this work, we uncover a systematic motion perception deficit in current ALLMs. To investigate this issue, we introduce AMPBench, the first benchmark explicitly designed to evaluate auditory motion understanding. AMPBench introduces a controlled question-answering benchmark designed to evaluate whether Audio-Language Models (LALMs) can infer the direction and trajectory of moving sound sources from binaural audio. Comprehensive quantitative and qualitative analyses reveal that current models struggle to reliably recognize motion cues or distinguish directional patterns. The average accuracy remains below 50%, underscoring a fundamental limitation in auditory spatial reasoning. Our study highlights a fundamental gap between human and model auditory spatial reasoning, providing both a diagnostic tool and new insight for enhancing spatial cognition in future Audio-Language Models.


【3】FoleyBench: A Benchmark For Video-to-Audio Models
标题:FoleyBench:视频到音频模型的基准
链接:https://arxiv.org/pdf/2511.13219v1

作者:Satvik Dixit,Koichi Saito,Zhi Zhong,Yuki Mitsufuji,Chris Donahue
摘要:视频到音频生成(V2 A)在电影后期制作,AR VR和声音设计等领域越来越重要,特别是用于创建与屏幕上动作同步的Foley声音效果。Foley要求生成既在语义上与可见事件对齐又在时间上与其时序对齐的音频。然而,由于缺乏针对Foley式情景的基准,评估和下游应用程序之间存在不匹配。我们发现,来自过去评估数据集的74%的视频具有较差的视听对应性。此外,它们被语音和音乐所主导,这些领域不在Foley的用例范围内。为了解决这一差距,我们引入了FoleyBench,这是第一个明确为Foley式V2 A评估而设计的大规模基准测试。FoleyBench包含5,000个(视频,地面实况音频,文本标题)三元组,每个三元组都具有可见的声源,音频与屏幕上的事件有因果关系。该数据集是使用自动化的,可扩展的管道构建的,该管道应用于来自基于YouTube和Vimeo的源的野外互联网视频。与过去的数据集相比,我们表明,视频从FoleyBench有更强的覆盖范围的声音类别,从专门为福利声音设计的分类。每个剪辑都进一步标记了捕获源复杂性、UCS AudioSet类别和视频长度的元数据,从而实现对模型性能和故障模式的细粒度分析。我们对几种最先进的V2 A模型进行了基准测试,在音频质量、音频-视频对齐、时间同步和音频-文本一致性方面对其进行了评估。样品可在https: gclef-cmu.org foleybench上获得摘要:Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Foley requires generating audio that is both semantically aligned with visible events and temporally aligned with their timing. Yet, there is a mismatch between evaluation and downstream applications due to the absence of a benchmark tailored to Foley-style scenarios. We find that 74% of videos from past evaluation datasets have poor audio-visual correspondence. Moreover, they are dominated by speech and music, domains that lie outside the use case for Foley. To address this gap, we introduce FoleyBench, the first large-scale benchmark explicitly designed for Foley-style V2A evaluation. FoleyBench contains 5,000 (video, ground-truth audio, text caption) triplets, each featuring visible sound sources with audio causally tied to on-screen events. The dataset is built using an automated, scalable pipeline applied to in-the-wild internet videos from YouTube-based and Vimeo-based sources. Compared to past datasets, we show that videos from FoleyBench have stronger coverage of sound categories from a taxonomy specifically designed for Foley sound. Each clip is further labeled with metadata capturing source complexity, UCS AudioSet category, and video length, enabling fine-grained analysis of model performance and failure modes. We benchmark several state-of-the-art V2A models, evaluating them on audio quality, audio-video alignment, temporal synchronization, and audio-text consistency. Samples are available at: https: gclef-cmu.org foleybench


【4】Towards Practical Real-Time Low-Latency Music Source Separation
标题:迈向实用的实时低延迟音乐源分离
链接:https://arxiv.org/pdf/2511.13146v1

作者:Junyu Wu,Jie Liu,Tianrui Pan,Jie Tang,Gangshan Wu
摘要:近年来,用于音乐混音的深度学习领域取得了重大进展。然而,对实时、低延迟的音乐混音的关注有限,其具有各种应用的潜力,例如助听器、音频流混音和现场表演。此外,还出现了一个明显的趋势,即发展更大的模型,限制了它们在某些情况下的适用性。在本文中,我们介绍了一个轻量级的实时低延迟模型称为实时单路径TFC-TDF UNET(RT-STT),它是基于双路径TFC-TDF UNET(DTTNet)。在RT-STT中,我们提出了一种基于通道扩展的特征融合技术。我们还证明了实时模型中单路径建模相对于双路径建模的优越性。此外,我们还研究了量化的方法,以进一步减少推理时间。与最先进的模型相比,RT-STT具有更少的参数和更短的推理时间,表现出卓越的性能。摘要:In recent years, significant progress has been made in the field of deep learning for music demixing. However, there has been limited attention on real-time, low-latency music demixing, which holds potential for various applications, such as hearing aids, audio stream remixing, and live performances. Additionally, a notable tendency has emerged towards the development of larger models, limiting their applicability in certain scenarios. In this paper, we introduce a lightweight real-time low-latency model called Real-Time Single-Path TFC-TDF UNET (RT-STT), which is based on the Dual-Path TFC-TDF UNET (DTTNet). In RT-STT, we propose a feature fusion technique based on channel expansion. We also demonstrate the superiority of single-path modeling over dual-path modeling in real-time models. Moreover, we investigate the method of quantization to further reduce inference time. RT-STT exhibits superior performance with significantly fewer parameters and shorter inference times compared to state-of-the-art models.


【5】SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs
标题:SynthGuard:一个使用多模式LLM检测人工智能生成的多媒体的开放平台
链接:https://arxiv.org/pdf/2511.12404v1

作者:Shail Desai,Aditya Pawar,Li Lin,Xin Wang,Shu Hu
摘要:人工智能(AI)使任何人都可以前所未有地轻松创建图像,音频和视频,丰富了教育,交流和创造性表达。与此同时,人工智能生成的媒体的迅速崛起带来了严重的风险,包括错误信息、身份滥用以及公众信任的侵蚀,因为合成内容与真实媒体越来越难以区分。尽管deepfake检测已经取得了进展,但许多现有的工具仍然是封闭源代码的,形式有限,或者缺乏透明度和教育价值,使得用户难以理解如何做出检测决策。为了解决这些差距,我们引入了SynthGuard,这是一个开放的,用户友好的平台,用于使用传统检测器和多模态大型语言模型(MLLM)检测和分析AI生成的多媒体。SynthGuard提供可解释的推理,统一的图像和音频支持,以及旨在使研究人员,教育工作者和公众可以访问法医分析的交互式界面。SynthGuard平台可从以下网址获得:https: in-engr-nova.it.purdue.edu 摘要:Artificial Intelligence (AI) has made it possible for anyone to create images, audio, and video with unprecedented ease, enriching education, communication, and creative expression. At the same time, the rapid rise of AI-generated media has introduced serious risks, including misinformation, identity misuse, and the erosion of public trust as synthetic content becomes increasingly indistinguishable from real media. Although deepfake detection has advanced, many existing tools remain closed-source, limited in modality, or lacking transparency and educational value, making it difficult for users to understand how detection decisions are made. To address these gaps, we introduce SynthGuard, an open, user-friendly platform for detecting and analyzing AI-generated multimedia using both traditional detectors and multimodal large language models (MLLMs). SynthGuard provides explainable inference, unified image and audio support, and an interactive interface designed to make forensic analysis accessible to researchers, educators, and the public. The SynthGuard platform is available at: https: in-engr-nova.it.purdue.edu


【6】MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement
标题:MF语音:通过因子解纠缠实现语音生成中的细粒度和成分控制
链接:https://arxiv.org/pdf/2511.12074v1

作者:Xinyue Yu,Youqing Fang,Pingyu Wu,Guoyang Ye,Wenbo Zhou,Weiming Zhang,Song Xiao
摘要:生成具有表达力和可控性的人类语音是生成式人工智能的核心目标之一,但其进展长期以来受到两个基本挑战的制约:语音因素的深度纠缠和现有控制机制的粗粒度。为了克服这些挑战,我们提出了一种新的框架,称为MF-Speech,它由两个核心组件组成:MF-SpeechEncoder和MF-SpeechGenerator。MF-SpeechEncoder作为一个因素净化器,采用多目标优化策略将原始语音信号分解为内容,音色和情感的高度纯净和独立的表示。随后,MF-SpeechGenerator作为指挥,通过动态融合和分层风格自适应归一化(HSAN)实现对这些因素的精确,可组合和细粒度控制。实验表明,在极具挑战性的多因素合成语音生成任务中,MF-Speech显著优于当前最先进的方法,实现了较低的单词错误率(WER=4.67%),优越的风格控制(SECS=0.5685,Corr=0.68)和最高的主观评价分数(nMOS=3.96,sMOS_emotion=3.86,sMOS_style=3.78)。此外,学习的离散因素表现出很强的可转移性,展示了它们作为通用语音表示的巨大潜力。摘要:Generating expressive and controllable human speech is one of the core goals of generative artificial intelligence, but its progress has long been constrained by two fundamental challenges: the deep entanglement of speech factors and the coarse granularity of existing control mechanisms. To overcome these challenges, we have proposed a novel framework called MF-Speech, which consists of two core components: MF-SpeechEncoder and MF-SpeechGenerator. MF-SpeechEncoder acts as a factor purifier, adopting a multi-objective optimization strategy to decompose the original speech signal into highly pure and independent representations of content, timbre, and emotion. Subsequently, MF-SpeechGenerator functions as a conductor, achieving precise, composable and fine-grained control over these factors through dynamic fusion and Hierarchical Style Adaptive Normalization (HSAN). Experiments demonstrate that in the highly challenging multi-factor compositional speech generation task, MF-Speech significantly outperforms current state-of-the-art methods, achieving a lower word error rate (WER=4.67%), superior style control (SECS=0.5685, Corr=0.68), and the highest subjective evaluation scores(nMOS=3.96, sMOS_emotion=3.86, sMOS_style=3.78). Furthermore, the learned discrete factors exhibit strong transferability, demonstrating their significant potential as a general-purpose speech representation.


【7】ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
标题:ProAV-DiT:用于高效同步音频视频生成的投影潜伏扩散Transformer
链接:https://arxiv.org/pdf/2511.12072v1

作者:Jiahui Sun,Weining Wang,Mingzhen Sun,Yirong Yang,Xinxin Zhu,Jing Liu
摘要:由于音频和视频之间固有的结构不一致以及多模态数据处理的高计算成本,探测视频生成(SVG)仍然是一项具有挑战性的任务。在本文中,我们介绍ProAV-DiT,一个投影潜在扩散Transformer设计的高效和同步的音频视频生成。为了解决结构上的不一致,我们将原始音频预处理成类似视频的表示,从而在音频和视频之间对齐时间和空间维度。在其核心,ProAV-DiT采用多尺度双流时空自动编码器(MDSA),它使用正交分解将两种模态投影到统一的潜在空间中,从而实现细粒度时空建模和语义对齐。为了进一步增强时间一致性和特定模态的融合,我们引入了多尺度注意机制,它包括多尺度时间自我注意和群体跨模态注意。此外,我们堆叠的二维潜在的MDSA到一个统一的三维潜在空间,这是由时空扩散Transformer。该设计有效地对时空依赖性进行建模,使得能够在减少计算开销的同时生成高保真同步的音频-视频内容。在标准基准上进行的大量实验表明,ProAV-DiT在生成质量和计算效率方面都优于现有方法。摘要:Sounding Video Generation (SVG) remains a challenging task due to the inherent structural misalignment between audio and video, as well as the high computational cost of multimodal data processing. In this paper, we introduce ProAV-DiT, a Projected Latent Diffusion Transformer designed for efficient and synchronized audio-video generation. To address structural inconsistencies, we preprocess raw audio into video-like representations, aligning both the temporal and spatial dimensions between audio and video. At its core, ProAV-DiT adopts a Multi-scale Dual-stream Spatio-Temporal Autoencoder (MDSA), which projects both modalities into a unified latent space using orthogonal decomposition, enabling fine-grained spatiotemporal modeling and semantic alignment. To further enhance temporal coherence and modality-specific fusion, we introduce a multi-scale attention mechanism, which consists of multi-scale temporal self-attention and group cross-modal attention. Furthermore, we stack the 2D latents from MDSA into a unified 3D latent space, which is processed by a spatio-temporal diffusion Transformer. This design efficiently models spatiotemporal dependencies, enabling the generation of high-fidelity synchronized audio-video content while reducing computational overhead. Extensive experiments conducted on standard benchmarks demonstrate that ProAV-DiT outperforms existing methods in both generation quality and computational efficiency.


【8】Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering
标题:通过多模式场景感知声学渲染增强XR听觉现实主义
链接:https://arxiv.org/pdf/2511.11930v1

作者:Hung-Yang Sung, Chien-Chun Wang, Kuan-Tang Huang, Tien-Hong Lo, Yu-Sheng Tsao, Yung-Chang Hsu, Berlin Chen
备注:Tianyu Xu,Jihan Li,Penghe Zu,Pranav Sahay,Maruchi Kim,Jack Obeng-Marnu,Farley Miller,Xun Qian,Katrina Passarella,Mahitha Rachumalla,Rajeev Nongpiur,D. Shin

Journal-ref:Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST '25), Article 17, 1-16, 2025

摘要:在延展实境(XR)中,渲染精确模拟真实世界声学的声音对于创建逼真可信的虚拟体验至关重要。然而,现有的XR空间音频渲染方法通常难以实时适应不同的物理场景,导致视觉和听觉线索之间的感官不匹配,破坏了用户的沉浸感。为了解决这个问题,我们介绍了SAMOSA,一种新型的设备上系统,通过动态适应其物理环境来呈现空间准确的声音。SAMOSA通过融合房间几何形状、表面材料和语义驱动的声学上下文的实时估计来利用协同多模态场景表示。然后,这种丰富的表示通过场景先验实现高效的声学校准,允许系统合成高度逼真的房间脉冲响应(RIR)。我们通过技术评估使用声学指标在各种房间配置和声音类型中进行RIR合成来验证我们的系统,同时进行专家评估(N=12)。评估结果表明,SAMOSA的可行性和有效性,提高XR听觉真实感。摘要:In Extended Reality (XR), rendering sound that accurately simulates real-world acoustics is pivotal in creating lifelike and believable virtual experiences. However, existing XR spatial audio rendering methods often struggle with real-time adaptation to diverse physical scenes, causing a sensory mismatch between visual and auditory cues that disrupts user immersion. To address this, we introduce SAMOSA, a novel on-device system that renders spatially accurate sound by dynamically adapting to its physical environment. SAMOSA leverages a synergistic multimodal scene representation by fusing real-time estimations of room geometry, surface materials, and semantic-driven acoustic context. This rich representation then enables efficient acoustic calibration via scene priors, allowing the system to synthesize a highly realistic Room Impulse Response (RIR). We validate our system through technical evaluation using acoustic metrics for RIR synthesis across various room configurations and sound types, alongside an expert evaluation (N=12). Evaluation results demonstrate SAMOSA's feasibility and efficacy in enhancing XR auditory realism.


【9】Real-Time Speech Enhancement via a Hybrid ViT: A Dual-Input Acoustic-Image Feature Fusion
标题:通过混合ViT实现实时语音增强:双输入声学图像特征融合
链接:https://arxiv.org/pdf/2511.11825v1

作者:Behnaz Bahmei,Siamak Arzanpour,Elina Birmingham
摘要:在嘈杂的环境中,语音质量和可懂度显著降低。本文提出了一种新的基于变压器的学习框架,以解决单通道噪声抑制问题的实时应用。尽管现有的深度学习网络在处理平稳噪声方面已经显示出显著的改进,但它们的性能在以非平稳噪声为特征的现实世界环境中通常会降低(例如,狗叫,婴儿哭)。所提出的双输入声图像特征融合使用混合ViT框架有效地建模噪声信号中的时间和频谱依赖性。针对真实世界的音频环境设计,所提出的框架是计算轻量级的,适合在嵌入式设备上实现。为了评估其有效性,四个标准和常用的质量测量,即PESQ,STOI,Seg SNR和LLR,被利用。实验结果表明,该方法在噪声输入信号的降噪效果、语音清晰度和感知质量方面都有显著提高,性能接近于干净的参考信号。摘要:Speech quality and intelligibility are significantly degraded in noisy environments. This paper presents a novel transformer-based learning framework to address the single-channel noise suppression problem for real-time applications. Although existing deep learning networks have shown remarkable improvements in handling stationary noise, their performance often diminishes in real-world environments characterized by non-stationary noise (e.g., dog barking, baby crying). The proposed dual-input acoustic-image feature fusion using a hybrid ViT framework effectively models both temporal and spectral dependencies in noisy signals. Designed for real-world audio environments, the proposed framework is computationally lightweight and suitable for implementation on embedded devices. To evaluate its effectiveness, four standard and commonly used quality measurements, namely PESQ, STOI, Seg SNR, and LLR, are utilized. Experimental results obtained using the Librispeech dataset as the clean speech source and the UrbanSound8K and Google Audioset datasets as the noise sources, demonstrate that the proposed method significantly improves noise reduction, speech intelligibility, and perceptual quality compared to the noisy input signal, achieving performance close to the clean reference.


【10】Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues
标题:超越显着性:利用专家参考的声学线索增强语音情感识别的解释
链接:https://arxiv.org/pdf/2511.11691v1

作者:Seham Nasr,Zhao Ren,David Johnson

备注:5 pages, 2 figures

摘要:用于语音情感识别(SER)的可解释AI(XAI)对于构建透明、可信的模型至关重要。目前的显着性为基础的方法,适应于视觉,突出声谱图区域,但未能显示这些区域是否对应于有意义的声学标记的情绪,限制了忠实性和可解释性。我们提出了一个框架,克服了这些限制,通过量化显着区域内的线索的幅度。这澄清了“什么”被强调,并将其与“为什么”重要联系起来,将显着性与专家参考的语音情感声学线索联系起来。基准SER数据集上的实验表明,我们的方法提高了解释质量显式连接的显着地区的理论驱动的语音情感专家参考声学。与标准的显着性方法相比,它提供了更容易理解和合理的解释SER模型,提供了一个值得信赖的基于语音的情感计算的基础步骤。摘要:Explainable AI (XAI) for Speech Emotion Recognition (SER) is critical for building transparent, trustworthy models. Current saliency-based methods, adapted from vision, highlight spectrogram regions but fail to show whether these regions correspond to meaningful acoustic markers of emotion, limiting faithfulness and interpretability. We propose a framework that overcomes these limitations by quantifying the magnitudes of cues within salient regions. This clarifies "what" is highlighted and connects it to "why" it matters, linking saliency to expert-referenced acoustic cues of speech emotions. Experiments on benchmark SER datasets show that our approach improves explanation quality by explicitly linking salient regions to theory-driven speech emotions expert-referenced acoustics. Compared to standard saliency methods, it provides more understandable and plausible explanations of SER models, offering a foundational step towards trustworthy speech-based affective computing.


【11】Regularized Schrödinger: Alleviating Distortion and Exposure Bias in Solving Inverse Problems
标题:正规薛定汉:减轻求解反问题时的失真和暴露偏差
链接:https://arxiv.org/pdf/2511.11686v1

作者:Qing Yao,Lijian Gao,Qirong Mao,Dong Ming
摘要:扩散模型作为一个强大的生成框架求解反问题。然而,它们仍然面临两个关键挑战:1)失真-感知权衡,其中改善感知质量通常会降低重建保真度,以及2)曝光偏差问题,其中训练-推理输入失配导致预测误差累积和重建质量降低。在这项工作中,我们提出了正则化薛定谔桥(RSB),薛定谔桥的适应量身定制的反问题,解决了上述限制。RSB采用了一种新的正则化训练策略,该策略扰动了输入状态和目标,通过将模型暴露于模拟预测误差来有效地减轻暴露偏差,并通过后验均值的精心设计的插值来减轻失真。在两个典型的语音增强逆问题上的大量实验表明,RSB优于最先进的方法,显着改善失真度量,并有效地减少曝光偏差。摘要:Diffusion models serve as a powerful generative framework for solving inverse problems. However, they still face two key challenges: 1) the distortion-perception tradeoff, where improving perceptual quality often degrades reconstruction fidelity, and 2) the exposure bias problem, where the training-inference input mismatch leads to prediction error accumulation and reduced reconstruction quality. In this work, we propose the Regularized Schrödinger Bridge (RSB), an adaptation of Schrödinger Bridge tailored for inverse problems that addresses the above limitations. RSB employs a novel regularized training strategy that perturbs both the input states and targets, effectively mitigating exposure bias by exposing the model to simulated prediction errors and also alleviating distortion by well-designed interpolation via the posterior mean. Extensive experiments on two typical inverse problems for speech enhancement demonstrate that RSB outperforms state-of-the-art methods, significantly improving distortion metrics and effectively reducing exposure bias.


【12】Lightweight Hopfield Neural Networks for Bioacoustic Detection and Call Monitoring of Captive Primates
标题:轻量级Hopfield神经网络用于圈养灵长类动物的生物声学检测和呼叫监测
链接:https://arxiv.org/pdf/2511.11615v1

作者:Wendy Lomas,Andrew Gascoyne,Colin Dubreuil,Stefano Vaglio,Liam Naughton

备注:16 pages, 3 figures, Proceedings of the Future Technologies Conference (FTC) 2025, Volume 1

Journal-ref:Proceedings of the Future Technologies Conference (FTC) 2025, Volume 1. FTC 2025. Lecture Notes in Networks and Systems, vol 1675. Springer, Cham

摘要:被动声学监测是一种可持续的野生动物和环境监测方法,导致生成大型数据集,目前,处理积压。自动化这一过程的学术研究主要集中在资源密集型卷积神经网络的应用上,这些网络需要大量的预标记数据集进行训练,并且在应用中缺乏灵活性。我们提出了一个可行的替代相关的野生和圈养设置;一个透明的,轻量级的和快速的训练联想记忆AI模型与Hopfield神经网络(HNN)架构。改编自一个模型开发检测蝙蝠回声定位电话,该模型监测圈养濒危黑白皱狐猴Varecia variegata发声。在监测福利时感兴趣的狐猴社会呼叫被存储在HNN中,以便检测更大的声学数据集上的其他呼叫实例。我们通过存储由运动引起的额外信号来对模型进行重大改进,并实现了0.94的总体准确度。该模型每秒可以执行340美元的分类,每分钟处理超过5.5小时的音频数据,在运行其他应用程序的标准笔记本电脑上。它具有广泛的适用性,并在毫秒内训练。我们的轻量级解决方案缩短了从数据到洞察的周转时间,并可在强制和野生环境中加快决策制定。摘要:Passive acoustic monitoring is a sustainable method of monitoring wildlife and environments that leads to the generation of large datasets and, currently, a processing backlog. Academic research into automating this process is focused on the application of resource intensive convolutional neural networks which require large pre-labelled datasets for training and lack flexibility in application. We present a viable alternative relevant in both wild and captive settings; a transparent, lightweight and fast-to-train associative memory AI model with Hopfield neural network (HNN) architecture. Adapted from a model developed to detect bat echolocation calls, this model monitors captive endangered black-and-white ruffed lemur Varecia variegata vocalisations. Lemur social calls of interest when monitoring welfare are stored in the HNN in order to detect other call instances across the larger acoustic dataset. We make significant model improvements by storing an additional signal caused by movement and achieve an overall accuracy of 0.94. The model can perform $340$ classifications per second, processing over 5.5 hours of audio data per minute, on a standard laptop running other applications. It has broad applicability and trains in milliseconds. Our lightweight solution reduces data-to-insight turnaround times and can accelerate decision making in both captive and wild settings.


【13】Systematic evaluation of time-frequency features for binaural sound source localization
标题:双耳声源定位时频特征的系统评估
链接:https://arxiv.org/pdf/2511.13487v1

作者:Davoud Shariat Panah,Alessandro Ragano,Dan Barry,Jan Skoglund,Andrew Hines

备注:Submitted to ICASSP 2026

摘要:本研究提出了一个系统的评估双耳声源定位(SSL)的时频特征设计,重点是如何在不同的条件下,功能选择影响模型的性能。我们研究了卷积神经网络(CNN)模型的性能,该模型使用基于幅度的特征(幅度谱图,耳间水平差- ILD)和基于相位的特征(相位谱图,耳间相位差- IPD)的各种组合。对具有不匹配的头部相关传递函数(HRTF)的域内和域外数据的评估表明,精心选择的特征组合往往优于模型复杂性的增加。虽然ILD + IPD等两个特征集足以用于域内SSL,但对不同内容的泛化需要更丰富的输入,将信道频谱图与ILD和IPD组合在一起。使用最佳特征集,我们的低复杂度CNN模型实现了具有竞争力的性能。我们的研究结果强调了双耳SSL功能设计的重要性,并为特定领域和通用本地化提供了实际指导。摘要:This study presents a systematic evaluation of time-frequency feature design for binaural sound source localization (SSL), focusing on how feature selection influences model performance across diverse conditions. We investigate the performance of a convolutional neural network (CNN) model using various combinations of amplitude-based features (magnitude spectrogram, interaural level difference - ILD) and phase-based features (phase spectrogram, interaural phase difference - IPD). Evaluations on in-domain and out-of-domain data with mismatched head-related transfer functions (HRTFs) reveal that carefully chosen feature combinations often outperform increases in model complexity. While two-feature sets such as ILD + IPD are sufficient for in-domain SSL, generalization to diverse content requires richer inputs combining channel spectrograms with both ILD and IPD. Using the optimal feature sets, our low-complexity CNN model achieves competitive performance. Our findings underscore the importance of feature design in binaural SSL and provide practical guidance for both domain-specific and general-purpose localization.


【14】VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
标题:SecureCraft-X:统一多语言、语音克隆语音合成和语音编辑
链接:https://arxiv.org/pdf/2511.12347v1

作者:Zhisheng Zheng,Puyuan Peng,Anuj Diwan,Cong Phuoc Huynh,Xiaohang Sun,Zhu Liu,Vimal Bhat,David Harwath
备注:EMNLP 2025. Demo and code are available at https:zhishengzheng.comvoicecraft-x
摘要:我们介绍了VoiceCraft-X,一种自回归神经编解码器语言模型,它统一了11种语言的多语言语音编辑和zero-shot文本到语音(TTS)合成:英语,普通话,韩语,日语,西班牙语,法语,德语,荷兰语,意大利语,葡萄牙语和波兰语。VoiceCraft-X利用Qwen 3大型语言模型进行无音素跨语言文本处理,并采用一种新颖的令牌重新排序机制,将文本和语音令牌按时间对齐,以将这两项任务作为单个序列生成问题来处理。该模型可以生成高质量、听起来自然的语音,在一个框架内无缝创建新音频或编辑现有录音。VoiceCraft-X在不同的语言环境中表现出强大的性能,即使每种语言的数据有限,也强调了统一自回归方法在推进复杂的真实多语言语音应用方面的强大功能。音频样本可在https: zhishengzheng.com voicecraft-x 上获得。摘要:We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot Text-to-Speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X utilizes the Qwen3 large language model for phoneme-free cross-lingual text processing and a novel token reordering mechanism with time-aligned text and speech tokens to handle both tasks as a single sequence generation problem. The model generates high-quality, natural-sounding speech, seamlessly creating new audio or editing existing recordings within one framework. VoiceCraft-X shows robust performance in diverse linguistic settings, even with limited per-language data, underscoring the power of unified autoregressive approaches for advancing complex, real-world multilingual speech applications. Audio samples are available at https: zhishengzheng.com voicecraft-x .


eess.AS音频处理


【1】Systematic evaluation of time-frequency features for binaural sound source localization
标题:双耳声源定位时频特征的系统评估
链接:https://arxiv.org/pdf/2511.13487v1

作者:Davoud Shariat Panah,Alessandro Ragano,Dan Barry,Jan Skoglund,Andrew Hines

备注:Submitted to ICASSP 2026

摘要:本研究提出了一个系统的评估双耳声源定位(SSL)的时频特征设计,重点是如何在不同的条件下,功能选择影响模型的性能。我们研究了卷积神经网络(CNN)模型的性能,该模型使用基于幅度的特征(幅度谱图,耳间水平差- ILD)和基于相位的特征(相位谱图,耳间相位差- IPD)的各种组合。对具有不匹配的头部相关传递函数(HRTF)的域内和域外数据的评估表明,精心选择的特征组合往往优于模型复杂性的增加。虽然ILD + IPD等两个特征集足以用于域内SSL,但对不同内容的泛化需要更丰富的输入,将信道频谱图与ILD和IPD组合在一起。使用最佳特征集,我们的低复杂度CNN模型实现了具有竞争力的性能。我们的研究结果强调了双耳SSL功能设计的重要性,并为特定领域和通用本地化提供了实际指导。摘要:This study presents a systematic evaluation of time-frequency feature design for binaural sound source localization (SSL), focusing on how feature selection influences model performance across diverse conditions. We investigate the performance of a convolutional neural network (CNN) model using various combinations of amplitude-based features (magnitude spectrogram, interaural level difference - ILD) and phase-based features (phase spectrogram, interaural phase difference - IPD). Evaluations on in-domain and out-of-domain data with mismatched head-related transfer functions (HRTFs) reveal that carefully chosen feature combinations often outperform increases in model complexity. While two-feature sets such as ILD + IPD are sufficient for in-domain SSL, generalization to diverse content requires richer inputs combining channel spectrograms with both ILD and IPD. Using the optimal feature sets, our low-complexity CNN model achieves competitive performance. Our findings underscore the importance of feature design in binaural SSL and provide practical guidance for both domain-specific and general-purpose localization.


【2】PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
标题:PASE:利用WavLM的音素先验进行低幻觉生成语音增强
链接:https://arxiv.org/pdf/2511.13300v1

作者:Xiaobin Rong,Qinwen Hu,Mansur Yesilbursa,Kamil Wojcicki,Jing Lu

备注:Accepted by AAAI 2026

摘要:生成模型在语音增强(SE)方面表现出了卓越的性能,与传统的判别方法相比,可以获得更好的感知质量。然而,现有的生成SE方法往往忽略了严重噪声下的幻觉风险,导致不正确的口语内容或不一致的扬声器特征,我们分别称之为语言和听觉幻觉。我们认为,语言幻觉源于模型的失败,以约束有效的语音结构,这是一个更根本的挑战。虽然语言模型(LM)非常适合通过对离散令牌的分布进行建模来捕获底层语音结构,但现有方法在从噪声损坏的表示中学习方面受到限制,这可能导致污染的先验和幻觉。为了克服这些限制,我们提出了语音锚定语音增强器(PASE),这是一个生成的SE框架,它利用嵌入在预训练的WavLM模型中的鲁棒的语音先验来减轻幻觉。首先,我们通过表示蒸馏将WavLM适配为去噪专家,以清理其最终层特征。在模型固有的语音先验的指导下,这个过程可以实现鲁棒的去噪,同时最大限度地减少语言幻觉。为了进一步减少幻听,我们用双流表示训练声码器:高级语音表示提供干净的语言内容,而低级声学表示保留说话者身份和韵律。实验结果表明,PASE不仅在感知质量上超过了最先进的判别模型,而且还显著优于先前的生成模型,具有显著更低的语言和听觉幻觉。摘要:Generative models have shown remarkable performance in speech enhancement (SE), achieving superior perceptual quality over traditional discriminative approaches. However, existing generative SE approaches often overlook the risk of hallucination under severe noise, leading to incorrect spoken content or inconsistent speaker characteristics, which we term linguistic and acoustic hallucinations, respectively. We argue that linguistic hallucination stems from models' failure to constrain valid phonological structures and it is a more fundamental challenge. While language models (LMs) are well-suited for capturing the underlying speech structure through modeling the distribution of discrete tokens, existing approaches are limited in learning from noise-corrupted representations, which can lead to contaminated priors and hallucinations. To overcome these limitations, we propose the Phonologically Anchored Speech Enhancer (PASE), a generative SE framework that leverages the robust phonological prior embedded in the pre-trained WavLM model to mitigate hallucinations. First, we adapt WavLM into a denoising expert via representation distillation to clean its final-layer features. Guided by the model's intrinsic phonological prior, this process enables robust denoising while minimizing linguistic hallucinations. To further reduce acoustic hallucinations, we train the vocoder with a dual-stream representation: the high-level phonetic representation provides clean linguistic content, while a low-level acoustic representation retains speaker identity and prosody. Experimental results demonstrate that PASE not only surpasses state-of-the-art discriminative models in perceptual quality, but also significantly outperforms prior generative models with substantially lower linguistic and acoustic hallucinations.


【3】Eardrum sound pressure prediction from ear canal reflectance based on the inverse solution of Webster's horn equation
标题:基于韦伯斯特号角方程逆解的耳道反射率预测耳膜压
链接:https://arxiv.org/pdf/2511.12552v1

作者:Reinhild Roden,Tobias Sankowsky-Rothe,Nick Wulbusch,Alexey Chernov,Matthias Blau
备注:Manuscript submitted to the Journal of the Acoustical Society of America (under minor revision)
摘要:为了推导用于入耳式听力系统的个性化均衡算法的耳道传递函数,需要个体耳道模型。在一维方法中,这需要估计耳道的个体面积函数。通过时域反射率的有限差分近似,面积函数可以有效地和可重复地计算为Webster喇叭方程的逆解。基于先前的研究,本研究进一步研究了在最佳空间分辨率下近似的终止,解决了典型耳道测量中缺乏较高频率的问题,并提高了逆解的准确性。与几何参考相比,通过将耳道几何形状的模拟输入阻抗外推到3.5 MHz的频率(对应于0.1 mm的空间分辨率),实现了更精确的面积函数。采用了先前工作的低通,但根据带限输入阻抗的最高频率对其截止频率进行了调整。在近似的耳道length.Finally,三维模拟和测量的耳道传递阻抗的终止面积函数的鲁棒性标准被发现,以及采用先前介绍的,并在此验证的一维电声模型馈送的面积函数复制。摘要:To derive ear canal transfer functions for individualized equalization algorithms of in-ear hearing systems, individual ear canal models are needed. In a one-dimensional approach, this requires the estimation of the individual area function of the ear canal. The area function can be effectively and reproducibly calculated as the inverse solution of Webster's horn equation by finite difference approximation of the time domain reflectance. Building upon previous research, the present study further investigates the termination of the approximation at an optimal spatial resolution, addressing the absence of higher frequencies in typical ear canal measurements and enhancing the accuracy of the inverse solution. Compared to the geometric reference, more precise area functions were achieved by extrapolating simulated input impedances of ear canal geometries up to a frequency of 3.5 MHz, corresponding to 0.1 mm spatial resolution. The low pass of the previous work was adopted but adjusted for its cut-off frequency depending on the highest frequency of the band-limited input impedance. Robust criteria for terminating the area function at the approximated ear canal length were found. Finally, three-dimensional simulated and measured ear canal transfer impedances were replicated well employing the previously introduced and herein validated one-dimensional electro-acoustic model fed by the area functions.


【4】VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
标题:SecureCraft-X:统一多语言、语音克隆语音合成和语音编辑
链接:https://arxiv.org/pdf/2511.12347v1

作者:Zhisheng Zheng,Puyuan Peng,Anuj Diwan,Cong Phuoc Huynh,Xiaohang Sun,Zhu Liu,Vimal Bhat,David Harwath

备注:EMNLP 2025. Demo and code are available at https:zhishengzheng.comvoicecraft-x

摘要:我们介绍了VoiceCraft-X,一种自回归神经编解码器语言模型,它统一了11种语言的多语言语音编辑和zero-shot文本到语音(TTS)合成:英语,普通话,韩语,日语,西班牙语,法语,德语,荷兰语,意大利语,葡萄牙语和波兰语。VoiceCraft-X利用Qwen 3大型语言模型进行无音素跨语言文本处理,并采用一种新颖的令牌重新排序机制,将文本和语音令牌按时间对齐,以将这两项任务作为单个序列生成问题来处理。该模型可以生成高质量、听起来自然的语音,在一个框架内无缝创建新音频或编辑现有录音。VoiceCraft-X在不同的语言环境中表现出强大的性能,即使每种语言的数据有限,也强调了统一自回归方法在推进复杂的真实多语言语音应用方面的强大功能。音频样本可在https: zhishengzheng.com voicecraft-x 上获得。摘要:We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot Text-to-Speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X utilizes the Qwen3 large language model for phoneme-free cross-lingual text processing and a novel token reordering mechanism with time-aligned text and speech tokens to handle both tasks as a single sequence generation problem. The model generates high-quality, natural-sounding speech, seamlessly creating new audio or editing existing recordings within one framework. VoiceCraft-X shows robust performance in diverse linguistic settings, even with limited per-language data, underscoring the power of unified autoregressive approaches for advancing complex, real-world multilingual speech applications. Audio samples are available at https: zhishengzheng.com voicecraft-x .


【5】How Far Do SSL Speech Models Listen for Tone? Temporal Focus of Tone Representation under Low-resource Transfer
标题:SSL语音模型监听音调的距离有多远?低资源传输下语气表示的时间焦点
链接:https://arxiv.org/pdf/2511.12285v1

作者:Minu Kim,Ji Sub Um,Hoirin Kim

备注:5 pages, 7 figures, submitted to ICASSP 2026

摘要:词汇音调是许多语言的核心,但在自监督学习(SSL)语音模型中仍然没有得到充分的研究,特别是在普通话之外。我们研究了四种具有复杂多样的音调系统的语言:缅甸语,泰语,老挝语和越南语,以研究这些模型在多大程度上倾听音调以及在低资源条件下如何转移。作为一个基线参考,我们估计的音调线索的时间跨度约为100毫秒,在缅甸语和泰国,老挝语和越南语约为180毫秒。对微调SSL模型的探测和梯度分析表明,音调传递因下游任务而异:自动语音识别微调将跨度与语言特定的音调线索对齐,而韵律和语音相关的任务使模型偏向于过长的跨度。这些发现表明,声调迁移是由下游任务,突出任务的时间焦点在声调建模的影响。摘要:Lexical tone is central to many languages but remains underexplored in self-supervised learning (SSL) speech models, especially beyond Mandarin. We study four languages with complex and diverse tone systems: Burmese, Thai, Lao, and Vietnamese, to examine how far such models listen for tone and how transfer operates in low-resource conditions. As a baseline reference, we estimate the temporal span of tone cues to be about 100 ms in Burmese and Thai, and about 180 ms in Lao and Vietnamese. Probes and gradient analyses on fine-tuned SSL models reveal that tone transfer varies by downstream task: automatic speech recognition fine-tuning aligns spans with language-specific tone cues, while prosody- and voice-related tasks bias the model toward overly long spans. These findings indicate that tone transfer is shaped by downstream task, highlighting task effects on temporal focus in tone modeling.


【6】Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
标题:迈向对话式匈牙利语音识别:引入BEA-Large和BEA-Dialogue数据集
链接:https://arxiv.org/pdf/2511.13529v1

作者:Máté Gedeon,Piroska Zsófia Barta,Péter Mihajlik,Tekla Etelka Gráczi,Anna Kohári,Katalin Mády
备注:Submitted to LREC 2026
摘要:自动语音识别(ASR)的进步在很大程度上得到了高资源语言的广泛数据集的增强,而匈牙利语等语言由于有限的自发和会话语料库而仍然代表性不足。为了解决这一差距,我们引入了两个新的数据集- BEA-Large和BEA-Dialogue -从以前未处理的部分匈牙利语语音语料库命名为GARNING。BEA-Large扩展了BEA-Base,提供了来自433位演讲者的255小时的自发演讲,并通过详细的片段级元数据进行了丰富。BEA-Dialogue包含85小时的自发对话,是一个匈牙利语语音语料库,其特征是将自然对话划分为与说话者无关的子集,支持会话ASR和说话者日记化的研究。我们使用公开可用的ASR模型在这些数据集上建立了可重复的基线,微调的Fast Conformer模型在自发语音和重复语音上的单词错误率分别低至14.18%和4.8%。实验结果表明,该算法的错误率在13.05%~ 18.26%之间,为进一步改进算法提供了参考。结果突出了会话ASR的持续困难,特别是由于不流利,重叠和非正式的语音模式。通过发布这些数据集和基线,我们的目标是推进匈牙利语语音技术,并为开发其他语言的自发和对话基准提供方法框架。摘要:The advancement of automatic speech recognition (ASR) has been largely enhanced by extensive datasets in high-resource languages, while languages such as Hungarian remain underrepresented due to limited spontaneous and conversational corpora. To address this gap, we introduce two new datasets -- BEA-Large and BEA-Dialogue -- constructed from the previously unprocessed portions of the Hungarian speech corpus named BEA. BEA-Large extends BEA-Base with 255 hours of spontaneous speech from 433 speakers, enriched with detailed segment-level metadata. BEA-Dialogue, comprising 85 hours of spontaneous conversations, is a Hungarian speech corpus featuring natural dialogues partitioned into speaker-independent subsets, supporting research in conversational ASR and speaker diarization. We establish reproducible baselines on these datasets using publicly available ASR models, with the fine-tuned Fast Conformer model achieving word error rates as low as 14.18 % on spontaneous and 4.8 % on repeated speech. Diarization experiments yield diarization error rates between 13.05 % and 18.26 %, providing reference points for future improvements. The results highlight the persistent difficulty of conversational ASR, particularly due to disfluencies, overlaps, and informal speech patterns. By releasing these datasets and baselines, we aim to advance Hungarian speech technology and offer a methodological framework for developing spontaneous and conversational benchmarks in other languages.


【7】FoleyBench: A Benchmark For Video-to-Audio Models
标题:FoleyBench:视频到音频模型的基准
链接:https://arxiv.org/pdf/2511.13219v1

作者:Satvik Dixit,Koichi Saito,Zhi Zhong,Yuki Mitsufuji,Chris Donahue
摘要:视频到音频生成(V2 A)在电影后期制作,AR VR和声音设计等领域越来越重要,特别是用于创建与屏幕上动作同步的Foley声音效果。Foley要求生成既在语义上与可见事件对齐又在时间上与其时序对齐的音频。然而,由于缺乏针对Foley式情景的基准,评估和下游应用程序之间存在不匹配。我们发现,来自过去评估数据集的74%的视频具有较差的视听对应性。此外,它们被语音和音乐所主导,这些领域不在Foley的用例范围内。为了解决这一差距,我们引入了FoleyBench,这是第一个明确为Foley式V2 A评估而设计的大规模基准测试。FoleyBench包含5,000个(视频,地面实况音频,文本标题)三元组,每个三元组都具有可见的声源,音频与屏幕上的事件有因果关系。该数据集是使用自动化的,可扩展的管道构建的,该管道应用于来自基于YouTube和Vimeo的源的野外互联网视频。与过去的数据集相比,我们表明,视频从FoleyBench有更强的覆盖范围的声音类别,从专门为福利声音设计的分类。每个剪辑都进一步标记了捕获源复杂性、UCS AudioSet类别和视频长度的元数据,从而实现对模型性能和故障模式的细粒度分析。我们对几种最先进的V2 A模型进行了基准测试,在音频质量、音频-视频对齐、时间同步和音频-文本一致性方面对其进行了评估。样品可在https: gclef-cmu.org foleybench上获得摘要:Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Foley requires generating audio that is both semantically aligned with visible events and temporally aligned with their timing. Yet, there is a mismatch between evaluation and downstream applications due to the absence of a benchmark tailored to Foley-style scenarios. We find that 74% of videos from past evaluation datasets have poor audio-visual correspondence. Moreover, they are dominated by speech and music, domains that lie outside the use case for Foley. To address this gap, we introduce FoleyBench, the first large-scale benchmark explicitly designed for Foley-style V2A evaluation. FoleyBench contains 5,000 (video, ground-truth audio, text caption) triplets, each featuring visible sound sources with audio causally tied to on-screen events. The dataset is built using an automated, scalable pipeline applied to in-the-wild internet videos from YouTube-based and Vimeo-based sources. Compared to past datasets, we show that videos from FoleyBench have stronger coverage of sound categories from a taxonomy specifically designed for Foley sound. Each clip is further labeled with metadata capturing source complexity, UCS AudioSet category, and video length, enabling fine-grained analysis of model performance and failure modes. We benchmark several state-of-the-art V2A models, evaluating them on audio quality, audio-video alignment, temporal synchronization, and audio-text consistency. Samples are available at: https: gclef-cmu.org foleybench


【8】A Smart-Glasses for Emergency Medical Services via Multimodal Multitask Learning
标题:通过多模式多任务学习为紧急医疗服务提供智能眼镜
链接:https://arxiv.org/pdf/2511.13078v1

作者:Liuyi Jin,Pasan Gunawardena,Amran Haroon,Runzhi Wang,Sangwoo Lee,Radu Stoleru,Michael Middleton,Zepeng Huo,Jeeeun Kim,Jason Moats
摘要:紧急医疗技术人员(EMT)在高压环境中工作,在沉重的认知和操作负荷下做出快速,关键的生命决策。我们介绍了EMSGlass(一种由EMSNet(第一个用于紧急医疗服务(EMS)的多模式多任务模型)和EMSServe(一种针对EMS场景定制的低延迟多模式服务框架)提供支持的智能眼镜系统)。EMSNet集成了文本、生命体征和现场图像,以构建对EMS事件的统一实时理解。EMSNet在真实世界的多模式EMS数据集上进行培训,同时支持多达五个关键EMS任务,与最先进的单峰基线相比具有更高的准确性。EMSServe建立在PyTorch之上,引入了一个模态感知模型拆分器和一个特征缓存机制,实现了跨异构硬件的自适应和高效推理,同时解决了异步模态到达现场的挑战。通过优化EMS场景中的多模态推理执行,EMSServe比直接PyTorch多模态推理实现了1.9 - 11.7倍的加速。对六名专业EMT的用户研究评估表明,EMSGlass通过直观的玻璃交互增强了实时态势感知、决策速度和运营效率。此外,来自用户研究的定性见解为将EMSGlass扩展到下一代支持AI的EMS系统提供了可操作的方向,将多模式智能与现实世界的应急响应工作流程联系起来。摘要:Emergency Medical Technicians (EMTs) operate in high-pressure environments, making rapid, life-critical decisions under heavy cognitive and operational loads. We present EMSGlass, a smart-glasses system powered by EMSNet, the first multimodal multitask model for Emergency Medical Services (EMS), and EMSServe, a low-latency multimodal serving framework tailored to EMS scenarios. EMSNet integrates text, vital signs, and scene images to construct a unified real-time understanding of EMS incidents. Trained on real-world multimodal EMS datasets, EMSNet simultaneously supports up to five critical EMS tasks with superior accuracy compared to state-of-the-art unimodal baselines. Built on top of PyTorch, EMSServe introduces a modality-aware model splitter and a feature caching mechanism, achieving adaptive and efficient inference across heterogeneous hardware while addressing the challenge of asynchronous modality arrival in the field. By optimizing multimodal inference execution in EMS scenarios, EMSServe achieves 1.9x -- 11.7x speedup over direct PyTorch multimodal inference. A user study evaluation with six professional EMTs demonstrates that EMSGlass enhances real-time situational awareness, decision-making speed, and operational efficiency through intuitive on-glass interaction. In addition, qualitative insights from the user study provide actionable directions for extending EMSGlass toward next-generation AI-enabled EMS systems, bridging multimodal intelligence with real-world emergency response workflows.


【9】Real-Time Speech Enhancement via a Hybrid ViT: A Dual-Input Acoustic-Image Feature Fusion
标题:通过混合ViT实现实时语音增强:双输入声学图像特征融合
链接:https://arxiv.org/pdf/2511.11825v1

作者:Behnaz Bahmei,Siamak Arzanpour,Elina Birmingham
摘要:在嘈杂的环境中,语音质量和可懂度显著降低。本文提出了一种新的基于变压器的学习框架,以解决单通道噪声抑制问题的实时应用。尽管现有的深度学习网络在处理平稳噪声方面已经显示出显著的改进,但它们的性能在以非平稳噪声为特征的现实世界环境中通常会降低(例如,狗叫,婴儿哭)。所提出的双输入声图像特征融合使用混合ViT框架有效地建模噪声信号中的时间和频谱依赖性。针对真实世界的音频环境设计,所提出的框架是计算轻量级的,适合在嵌入式设备上实现。为了评估其有效性,四个标准和常用的质量测量,即PESQ,STOI,Seg SNR和LLR,被利用。实验结果表明,该方法在噪声输入信号的降噪效果、语音清晰度和感知质量方面都有显著提高,性能接近于干净的参考信号。摘要:Speech quality and intelligibility are significantly degraded in noisy environments. This paper presents a novel transformer-based learning framework to address the single-channel noise suppression problem for real-time applications. Although existing deep learning networks have shown remarkable improvements in handling stationary noise, their performance often diminishes in real-world environments characterized by non-stationary noise (e.g., dog barking, baby crying). The proposed dual-input acoustic-image feature fusion using a hybrid ViT framework effectively models both temporal and spectral dependencies in noisy signals. Designed for real-world audio environments, the proposed framework is computationally lightweight and suitable for implementation on embedded devices. To evaluate its effectiveness, four standard and commonly used quality measurements, namely PESQ, STOI, Seg SNR, and LLR, are utilized. Experimental results obtained using the Librispeech dataset as the clean speech source and the UrbanSound8K and Google Audioset datasets as the noise sources, demonstrate that the proposed method significantly improves noise reduction, speech intelligibility, and perceptual quality compared to the noisy input signal, achieving performance close to the clean reference.


【10】Lessons Learned from Developing a Privacy-Preserving Multimodal Wearable for Local Voice-and-Vision Inference
标题:开发保护隐私的多模式可穿戴设备以进行本地语音和视觉推理的经验教训
链接:https://arxiv.org/pdf/2511.11811v1

作者:Yonatan Tussa,Andy Heredia,Nirupam Roy

备注:7 pages, 5 figures

摘要:多模式可穿戴设备的许多有前途的应用需要连续的传感和繁重的计算,但用户由于隐私问题而拒绝使用此类设备。本文分享了我们构建耳戴式语音和视觉可穿戴设备的经验,该设备使用配对的智能手机作为可信的个人边缘来执行本地AI推理。我们描述了这个隐私保护系统的硬件-软件协同设计,包括在30克形状因子内集成摄像头,麦克风和扬声器,实现唤醒词触发捕获以及完全离线运行量化视觉语言和大型语言模型的挑战。通过迭代原型设计,我们确定了在功率预算,连接性,延迟和社会可接受性的关键设计障碍。我们的初步评估表明,完全本地多模态推理是可行的商品移动硬件与交互延迟。最后,我们为开发嵌入式AI系统的研究人员提供了设计经验,这些系统可以在日常环境中平衡隐私,响应能力和可用性。摘要:Many promising applications of multimodal wearables require continuous sensing and heavy computation, yet users reject such devices due to privacy concerns. This paper shares our experiences building an ear-mounted voice-and-vision wearable that performs local AI inference using a paired smartphone as a trusted personal edge. We describe the hardware--software co-design of this privacy-preserving system, including challenges in integrating a camera, microphone, and speaker within a 30-gram form factor, enabling wake word-triggered capture, and running quantized vision-language and large-language models entirely offline. Through iterative prototyping, we identify key design hurdles in power budgeting, connectivity, latency, and social acceptability. Our initial evaluation shows that fully local multimodal inference is feasible on commodity mobile hardware with interactive latency. We conclude with design lessons for researchers developing embedded AI systems that balance privacy, responsiveness, and usability in everyday settings.


【11】Lightweight Hopfield Neural Networks for Bioacoustic Detection and Call Monitoring of Captive Primates
标题:轻量级Hopfield神经网络用于圈养灵长类动物的生物声学检测和呼叫监测
链接:https://arxiv.org/pdf/2511.11615v1

作者:Wendy Lomas,Andrew Gascoyne,Colin Dubreuil,Stefano Vaglio,Liam Naughton

备注:16 pages, 3 figures, Proceedings of the Future Technologies Conference (FTC) 2025, Volume 1

Journal-ref:Proceedings of the Future Technologies Conference (FTC) 2025, Volume 1. FTC 2025. Lecture Notes in Networks and Systems, vol 1675. Springer, Cham

摘要:被动声学监测是一种可持续的野生动物和环境监测方法,导致生成大型数据集,目前,处理积压。自动化这一过程的学术研究主要集中在资源密集型卷积神经网络的应用上,这些网络需要大量的预标记数据集进行训练,并且在应用中缺乏灵活性。我们提出了一个可行的替代相关的野生和圈养设置;一个透明的,轻量级的和快速的训练联想记忆AI模型与Hopfield神经网络(HNN)架构。改编自一个模型开发检测蝙蝠回声定位电话,该模型监测圈养濒危黑白皱狐猴Varecia variegata发声。在监测福利时感兴趣的狐猴社会呼叫被存储在HNN中,以便检测更大的声学数据集上的其他呼叫实例。我们通过存储由运动引起的额外信号来对模型进行重大改进,并实现了0.94的总体准确度。该模型每秒可以执行340美元的分类,每分钟处理超过5.5小时的音频数据,在运行其他应用程序的标准笔记本电脑上。它具有广泛的适用性,并在毫秒内训练。我们的轻量级解决方案缩短了从数据到洞察的周转时间,并可在强制和野生环境中加快决策制定。摘要:Passive acoustic monitoring is a sustainable method of monitoring wildlife and environments that leads to the generation of large datasets and, currently, a processing backlog. Academic research into automating this process is focused on the application of resource intensive convolutional neural networks which require large pre-labelled datasets for training and lack flexibility in application. We present a viable alternative relevant in both wild and captive settings; a transparent, lightweight and fast-to-train associative memory AI model with Hopfield neural network (HNN) architecture. Adapted from a model developed to detect bat echolocation calls, this model monitors captive endangered black-and-white ruffed lemur Varecia variegata vocalisations. Lemur social calls of interest when monitoring welfare are stored in the HNN in order to detect other call instances across the larger acoustic dataset. We make significant model improvements by storing an additional signal caused by movement and achieve an overall accuracy of 0.94. The model can perform $340$ classifications per second, processing over 5.5 hours of audio data per minute, on a standard laptop running other applications. It has broad applicability and trains in milliseconds. Our lightweight solution reduces data-to-insight turnaround times and can accelerate decision making in both captive and wild settings.


机器翻译由腾讯交互翻译提供,仅供参考