微信公众号:arXiv_Daily
cs.SD语音
【1】Bridging the Gap between Micro-scale Traffic Simulation and 4D Digital Cityscapes
标题:缩小微观交通仿真与4D数字城市景观之间的差距
链接:https://arxiv.org/abs/2604.08497
摘要:虽然微观交通模拟为城市规划提供了必要的数据,但它们很少与有效的利益相关者沟通所需的高保真可视化或可听化相结合。在这项工作中,我们提出了一个实时的4D可视化框架,耦合的相扑交通与逼真的,地理空间上准确的VR表示苏黎世在虚幻引擎5。我们的架构实现了一个强大的C++数据管道,用于同步车辆可视化,并具有一个开放的声音控制(OSC)接口,以支持外部可听化引擎。我们通过用户研究评估模拟交通动态和人类感知之间的相关性来验证该框架。结果表明,高度的感知对齐,用户正确地解释了4D模拟的安全风险。此外,我们的研究结果表明,包括空间化音频改变了用户的安全感,显示多模态在交通模拟的重要性。
摘要:While micro-scale traffic simulations provide essential data for urban planning, they are rarely coupled with the high-fidelity visualization or auralization necessary for effective stakeholder communication. In this work, we present a real-time 4D visualization framework that couples the SUMO traffic with a photorealistic, geospatially accurate VR representation of Zurich in Unreal Engine 5. Our architecture implements a robust C++ data pipeline for synchronized vehicle visualization and features an Open Sound Control (OSC) interface to support external auralization engines. We validate the framework through a user study assessing the correlation between simulated traffic dynamics and human perception. Results demonstrate a high degree of perceptual alignment, where users correctly interpret safety risks from the 4D simulation. Furthermore, our findings indicate that the inclusion of spatialized audio alters the user's sense of safety, showing the importance of multimodality in traffic simulations.
【2】DeepFense: A Unified, Modular, and Extensible Framework for Robust Deepfake Audio Detection
标题:DeepFense:用于稳健Deepfake音频检测的统一、模块化和可扩展框架
链接:https://arxiv.org/abs/2604.08450
备注:Deepfense Toolkit
摘要:语音Deepfake检测是一个成熟的研究领域,具有不同的模型,数据集和训练策略。然而,缺乏标准化的实施和评估协议限制了重复性,基准测试和跨研究的比较。在这项工作中,我们介绍了DeepFense,这是一个全面的开源PyTorch工具包,集成了最新的架构,损失函数和增强管道,以及100多个配方。使用DeepFense,我们对400多个模型进行了大规模评估。我们的研究结果表明,虽然精心策划的训练数据提高了跨域泛化,但预训练的前端特征提取器的选择主导了整体性能方差。至关重要的是,我们在高性能模型中表现出严重的偏见,包括音频质量,说话者性别和语言。预计DeepFense将通过必要的工具促进实际部署,以解决公平的训练数据选择和前端微调。
摘要:Speech deepfake detection is a well-established research field with different models, datasets, and training strategies. However, the lack of standardized implementations and evaluation protocols limits reproducibility, benchmarking, and comparison across studies. In this work, we present DeepFense, a comprehensive, open-source PyTorch toolkit integrating the latest architectures, loss functions, and augmentation pipelines, alongside over 100 recipes. Using DeepFense, we conducted a large-scale evaluation of more than 400 models. Our findings reveal that while carefully curated training data improves cross-domain generalization, the choice of pre-trained front-end feature extractor dominates overall performance variance. Crucially, we show severe biases in high-performing models regarding audio quality, speaker gender, and language. DeepFense is expected to facilitate real-world deployment with the necessary tools to address equitable training data selection and front-end fine-tuning.
【3】Selective Attention System (SAS): Device-Addressed Speech Detection for Real-Time On-Device Voice AI
标题:选择性注意系统(SAS):用于实时设备上语音AI的设备编址语音检测
链接:https://arxiv.org/abs/2604.08412
摘要:我们研究了在预ASR边缘部署约束下的设备寻址语音检测,其中系统必须在严格的延迟和计算限制下决定是否在转录之前转发音频。我们发现,在多说话人的环境中,时间模糊的话语,这个任务更有效地建模为一个顺序路由问题的互动历史比作为一个话语本地分类任务。我们将其形式化为顺序设备寻址路由(SDAR),并提出了选择性注意系统(SAS),一个设备上的实现,实例化这个配方。 在一个60小时的多人英语测试集上,主要的纯音频配置达到了F1=0.86(精度=0.89,召回率=0.83);可选的摄像头,音频+视频融合将F1提高到0.95(精度=0.97,召回率=0.93)。在我们的评估方案下,删除因果交互历史(第3阶段)将音频+视频配置中的F1从0.95降至0.57+/-0.03。在测试的组件中,这是最大的观察到的消融效果,表明短期的相互作用的历史进行大量的决策相关的信息,在评价设置。SAS在ARM Cortex-A类硬件上完全在设备上运行(<150 ms延迟,<20 MB占用空间)。所有结果均来自主要以英语评价的专有数据集的内部评价;可共享5小时评价子集以进行独立验证(第8.8节)。
摘要:We study device-addressed speech detection under pre-ASR edge deployment constraints, where systems must decide whether to forward audio before transcription under strict latency and compute limits. We show that, in multi-speaker environments with temporally ambiguous utterances, this task is more effectively modelled as a sequential routing problem over interaction history than as an utterance-local classification task. We formalize this as Sequential Device-Addressed Routing (SDAR) and present the Selective Attention System (SAS), an on-device implementation that instantiates this formulation. On a held-out 60-hour multi-speaker English test set, the primary audio-only configuration achieves F1=0.86 (precision=0.89, recall=0.83); with an optional camera, audio+video fusion raises F1 to 0.95 (precision=0.97, recall=0.93). Removing causal interaction history (Stage~3) reduced F1 from 0.95 to 0.57+/-0.03 in the audio+video configuration under our evaluation protocol. Among the tested components, this was the largest observed ablation effect, indicating that short-horizon interaction history carries substantial decision-relevant information in the evaluated setting. SAS runs fully on-device on ARM Cortex-A class hardware (<150 ms latency, <20 MB footprint). All results are from internal evaluation on a proprietary dataset evaluated primarily in English; a 5-hour evaluation subset may be shared for independent verification (Section 8.8).
【4】CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation
标题:CapTalk:单一话语和对话语音生成的统一语音设计
链接:https://arxiv.org/abs/2604.08363
备注:14 pages, 2 figures
摘要:基于自然语言描述的语音设计是文本到语音多模态生成中的一个新任务,其目标是在不依赖参考音频的情况下合成具有目标音色和说话风格的语音。然而,现有的方法主要集中在单话语生成,留下会话语音设计在很大程度上未被探索。在这项工作中,我们将语音设计扩展到对话,从而在自然对话环境中实现更好的目标扬声器建模和回合级别表达控制。我们提出了CapTalk,一个统一的标题条件的文本音频自回归框架,用于单话语和对话语音设计。CapTalk使用话语级字幕进行单话语语音设计,使用说话人级字幕进行对话说话人建模,并进一步在对话中引入CoT控制序列来显式规划回合级动态属性。为了解决稳定的音色保持和上下文自适应表达之间的冲突,我们提出了一个分层变分条件模块与话语级扬声器编码器,以更好地平衡稳定的音色保持和上下文自适应表达。这使得音色重用,同时保持表达适应当前的话语,并在对话中,周围的环境。我们还建立了一个全面的评估协议,为单一的话语和对话设置。实验表明,CapTalk实现了最先进的性能上的单话语语音设计基准,并提供更好的表达可控性和上下文适当的多轮对话。音频样本可在www.example.com上获得。
摘要:Voice design from natural language descriptions is emerging as a new task in text-to-speech multimodal generation, aiming to synthesize speech with target timbre and speaking style without relying on reference audio. However, existing methods mainly focus on single-utterance generation, leaving conversational voice design largely unexplored. In this work, we extend voice design to dialogue, enabling better target speaker modeling and turn-level expressive control in natural conversational settings. We propose CapTalk, a unified caption-conditioned text-audio autoregressive framework for both single-utterance and dialogue voice design. CapTalk uses utterance-level captions for single-utterance voice design and speaker-level captions for dialogue speaker modeling, and further introduces a CoT control sequence in dialogue to explicitly plan turn-level dynamic attributes. To resolve the conflict between stable timbre preservation and context-adaptive expression, we propose a hierarchical variational conditioning module with an utterance-level speaker encoder to better balance stable timbre preservation and context-adaptive expression. This enables timbre reuse while keeping expression adaptive to the current utterance and, in dialogue, the surrounding context. We also build a comprehensive evaluation protocol for both single-utterance and dialogue settings. Experiments show that CapTalk achieves state-of-the-art performance on a single-utterance voice design benchmark and delivers better expression controllability and contextual appropriateness in multi-turn dialogue. Audio samples are available at: https://anonymous.4open.science/api/repo/Captalk-D601/file/index.html.
【5】AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan
标题:AT-ADD:全类型音频深度造假检测挑战评估计划
链接:https://arxiv.org/abs/2604.08184
备注:Accepted to the ACM Multimedia 2026 Grand Challenge
摘要:音频大语言模型(ALLM)的快速发展使语音和非语音音频(包括音效、歌声和音乐)的成本效益高、高保真度的生成和操作成为可能。虽然这些功能促进了创造力和内容制作,但它们也带来了重大的安全和信任挑战,因为现在可以大规模生成和传播逼真的音频deepfake。然而,现有的音频深度伪造检测(ADD)对策(CM)和基准测试仍然主要以语音为中心,通常依赖于语音特定的伪影,对现实世界的失真表现出有限的鲁棒性,以及对异构音频类型和新兴欺骗技术的有限推广。为了解决这些差距,我们提出了ACM Multimedia 2026的全类型音频Deepfake检测(AT-ADD)大挑战,旨在将受控的学术评估与实际的多媒体取证联系起来。AT-ADD包括两个轨道:(1)鲁棒的语音Deepfake检测,它在真实世界的场景下评估检测器,并针对看不见的最先进的语音生成方法;(2)所有类型的音频Deepfake检测,它将检测扩展到语音之外的各种未知音频类型,并促进语音,声音,唱歌和音乐的类型不可知的泛化。通过提供标准化的数据集、严格的评估协议和可复制的基线,AT-ADD旨在加速强大且可推广的音频取证技术的开发,在普遍存在的合成音频时代支持安全通信、可靠的媒体验证和负责任的治理。
摘要:The rapid advancement of Audio Large Language Models (ALLMs) has enabled cost-effective, high-fidelity generation and manipulation of both speech and non-speech audio, including sound effects, singing voices, and music. While these capabilities foster creativity and content production, they also introduce significant security and trust challenges, as realistic audio deepfakes can now be generated and disseminated at scale. Existing audio deepfake detection (ADD) countermeasures (CMs) and benchmarks, however, remain largely speech-centric, often relying on speech-specific artifacts and exhibiting limited robustness to real-world distortions, as well as restricted generalization to heterogeneous audio types and emerging spoofing techniques. To address these gaps, we propose the All-Type Audio Deepfake Detection (AT-ADD) Grand Challenge for ACM Multimedia 2026, designed to bridge controlled academic evaluation with practical multimedia forensics. AT-ADD comprises two tracks: (1) Robust Speech Deepfake Detection, which evaluates detectors under real-world scenarios and against unseen, state-of-the-art speech generation methods; and (2) All-Type Audio Deepfake Detection, which extends detection beyond speech to diverse, unknown audio types and promotes type-agnostic generalization across speech, sound, singing, and music. By providing standardized datasets, rigorous evaluation protocols, and reproducible baselines, AT-ADD aims to accelerate the development of robust and generalizable audio forensic technologies, supporting secure communication, reliable media verification, and responsible governance in an era of pervasive synthetic audio.
【6】Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning
标题:基于教师引导的双路径视听表征学习的语义降噪
链接:https://arxiv.org/abs/2604.08147
摘要:视听表征学习的最新进展表明了对比对齐与掩蔽重建相结合的价值。然而,在单个前向传递中联合优化这些目标迫使对比分支依赖于为重建而设计的随机可见补丁,而不是跨模态对齐,从而引入语义噪声和优化干扰。我们提出了TG-DP,教师指导的双路径框架,将重建和对齐整合到单独的优化路径中。通过解开两个分支的掩蔽机制,TG-DP使对比路径能够使用更适合于跨模态对齐的可见性模式。教师模型进一步提供了组织该分支中可见标记的辅助指导,有助于减少干扰并稳定跨模态表示学习。TG-DP在zero-shot检索方面实现了最先进的性能。在AudioSet上,视频到音频检索的R@1从35.2%提高到37.4%,音频到视频检索的R@1从27.9%提高到37.1%。学习的表示也保持了语义上的鲁棒性,在AS 20 K和VGGSound上实现了最先进的线性探头性能。综上所述,我们的研究结果表明,解耦多模态目标和引入教师指导的结构到对比路径提供了一个有效的框架,提高大规模的视听预训练。代码可在https://github.com/wanglg20/TG-DP上获得。
摘要:Recent advances in audio-visual representation learning have shown the value of combining contrastive alignment with masked reconstruction. However, jointly optimizing these objectives in a single forward pass forces the contrastive branch to rely on randomly visible patches designed for reconstruction rather than cross-modal alignment, introducing semantic noise and optimization interference. We propose TG-DP, a Teacher-Guided Dual-Path framework that decouples reconstruction and alignment into separate optimization paths. By disentangling the masking regimes of the two branches, TG-DP enables the contrastive pathway to use a visibility pattern better suited to cross-modal alignment. A teacher model further provides auxiliary guidance for organizing visible tokens in this branch, helping reduce interference and stabilize cross-modal representation learning. TG-DP achieves state-of-the-art performance in zero-shot retrieval. On AudioSet, it improves R@1 from 35.2\% to 37.4\% for video-to-audio retrieval and from 27.9\% to 37.1\% for audio-to-video retrieval. The learned representations also remain semantically robust, achieving state-of-the-art linear-probe performance on AS20K and VGGSound. Taken together, our results suggest that decoupling multimodal objectives and introducing teacher-guided structure into the contrastive pathway provide an effective framework for improving large-scale audio-visual pretraining. Code is available at https://github.com/wanglg20/TG-DP.
【7】DeepForestSound: a multi-species automatic detector for passive acoustic monitoring in African tropical forests, a case study in Kibale National Park
标题:DeepForestSound:一种用于非洲热带森林被动声学监测的多物种自动探测器,基巴莱国家公园的案例研究
链接:https://arxiv.org/abs/2604.08087
备注:8 pages
摘要:被动声监测(PAM)技术在生物多样性评价中有着广泛的应用。它在非洲热带森林的应用是有限的稀缺注释数据,降低了性能的通用生态声学模型代表性不足的类群。在这项研究中,我们介绍了DeepForestSound(DFS),一个多物种的自动检测模型,专为非洲热带森林中的PAM。DFS依赖于半监督管道,将未注释录音的聚类与手动验证相结合,然后使用低等级自适应对音频谱图Transformer(AST)进行监督微调,并将其与冻结的主干线性基线(DFS-Linear)进行比较。该框架支持从长期声学记录中检测多个分类组,包括鸟类,灵长类动物和大象。DFS在乌干达Kibale国家公园的Sebitoli地区收集的声学数据上进行了培训,并在两年后在同一森林的不同地点记录的独立数据集上进行了评估。因此,本评价评估了单一热带森林生态系统内跨时间和记录地点的一般化。在12个分类群中的8个分类群中,DFS优于现有的自动检测工具,特别是对于非鸟类分类群,灵长类动物的平均AP值为0.964,大象为0.961。结果进一步表明,基于LoRA的微调大大优于跨分类群的线性探测。总体而言,这些结果表明,以任务为导向,区域特定的培训大大提高了在声学复杂的热带环境中的检测性能,并强调DFS作为非洲雨林生物多样性监测和保护的实用工具的潜力。
摘要:Passive Acoustic Monitoring (PAM) is widely used for biodiversity assessment. Its application in African tropical forests is limited by scarce annotated data, reducing the performance of general-purpose ecoacoustic models on underrepresented taxa. In this study, we introduce DeepForestSound (DFS), a multi-species automatic detection model designed for PAM in African tropical forests. DFS relies on a semi-supervised pipeline combining clustering of unannotated recordings with manual validation, followed by supervised fine-tuning of an Audio Spectrogram Transformer (AST) using low-rank adaptation, which is compared to a frozen-backbone linear baseline (DFS-Linear). The framework supports the detection of multiple taxonomic groups, including birds, primates, and elephants, from long-term acoustic recordings. DFS was trained on acoustic data collected in the Sebitoli area, in Kibale National Park, Uganda, and evaluated on an independent dataset recorded two years later at different locations within the same forest. This evaluation therefore assesses generalization across time and recording sites within a single tropical forest ecosystem. Across 8 out of 12 taxons, DFS outperforms existing automatic detection tools, particularly for non-avian taxa, achieving average AP values of 0.964 for primates and 0.961 for elephants. Results further show that LoRA-based fine-tuning substantially outperforms linear probing across taxa. Overall, these results demonstrate that task-oriented, region-specific training substantially improves detection performance in acoustically complex tropical environments, and highlight the potential of DFS as a practical tool for biodiversity monitoring and conservation in African rainforests.
【8】Towards Real-Time Human-AI Musical Co-Performance: Accompaniment Generation with Latent Diffusion Models and MAX/MSP
标题:迈向实时人类与人工智能音乐联合表演:具有潜在扩散模型和MAX/MSP的伴奏生成
链接:https://arxiv.org/abs/2604.07612
备注:12 pages, 6 figures
摘要:我们提出了一个实时人类-人工智能音乐协同表演的框架,其中潜在扩散模型响应于上下文音频的实时流生成乐器伴奏。该系统将MAX/MSP前端(处理实时音频输入、缓冲和回放)与运行生成模型的Python推理服务器相结合,通过OSC/UDP消息进行通信。这使得音乐家可以在MAX/MSP(一个成熟的实时环境)中表演,同时与基于Python的大规模生成模型进行交互,克服了实时音乐工具和最先进的AI模型之间的根本脱节。我们将伴奏生成制定为滑动窗口前瞻协议,训练模型从部分上下文预测未来音频,其中系统延迟是一个关键约束。为了减少延迟,我们将一致性蒸馏应用于我们的扩散模型,实现了采样时间的5.4倍减少,两个模型都实现了实时操作。通过对音乐连贯性、节拍对齐和音频质量的评估,这两种型号在回溯模式下都实现了强大的性能,并随着前瞻性的增加而优雅地降低。这些结果证明了基于扩散的实时伴奏的可行性,并揭示了任何此类系统必须导航的模型延迟,前瞻深度和生成质量之间的基本权衡。
摘要:We present a framework for real-time human-AI musical co-performance, in which a latent diffusion model generates instrumental accompaniment in response to a live stream of context audio. The system combines a MAX/MSP front-end-handling real-time audio input, buffering, and playback-with a Python inference server running the generative model, communicating via OSC/UDP messages. This allows musicians to perform in MAX/MSP - a well-established, real-time capable environment - while interacting with a large-scale Python-based generative model, overcoming the fundamental disconnect between real-time music tools and state-of-the-art AI models. We formulate accompaniment generation as a sliding-window look-ahead protocol, training the model to predict future audio from partial context, where system latency is a critical constraint. To reduce latency, we apply consistency distillation to our diffusion model, achieving a 5.4x reduction in sampling time, with both models achieving real-time operation. Evaluated on musical coherence, beat alignment, and audio quality, both models achieve strong performance in the Retrospective regime and degrade gracefully as look-ahead increases. These results demonstrate the feasibility of diffusion-based real-time accompaniment and expose the fundamental trade-off between model latency, look-ahead depth, and generation quality that any such system must navigate.
【9】Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
标题:语义-情感共鸣嵌入:跨语言语音情感识别的半监督范式
链接:https://arxiv.org/abs/2604.07417
备注:Main paper (6 pages). Accepted for publication by IEEE International conference on Multimedia and Expo 2026 (ICME 2026)
摘要:跨语言语音情感识别(CLSER)旨在识别看不见的语言中的情感状态。然而,现有的方法严重依赖于完整标签的语义同步和静态特征稳定性,阻碍了低资源语言达到高资源性能。为了解决这个问题,我们提出了一个半监督框架的基础上语义情感共振嵌入(SERE),跨语言的动态特征范例,既不需要目标语言标签,也不需要翻译对齐。具体来说,SERE使用少量的标记样本构建情感语义结构。它通过瞬时共振场(IRF)学习人类的情感体验,使未标记的样本能够自组织成这种结构。这实现了半监督语义指导和结构发现。此外,我们设计了一个三重共振交互链(TRIC)损失,使模型能够加强在情感亮点标记和未标记的样本之间的交互和嵌入能力。跨多种语言的广泛实验证明了我们的方法的有效性,只需要5杆标记的源语言。
摘要:Cross-lingual Speech Emotion Recognition (CLSER) aims to identify emotional states in unseen languages. However, existing methods heavily rely on the semantic synchrony of complete labels and static feature stability, hindering low-resource languages from reaching high-resource performance. To address this, we propose a semi-supervised framework based on Semantic-Emotional Resonance Embedding (SERE), a cross-lingual dynamic feature paradigm that requires neither target language labels nor translation alignment. Specifically, SERE constructs an emotion-semantic structure using a small number of labeled samples. It learns human emotional experiences through an Instantaneous Resonance Field (IRF), enabling unlabeled samples to self-organize into this structure. This achieves semi-supervised semantic guidance and structural discovery. Additionally, we design a Triple-Resonance Interaction Chain (TRIC) loss to enable the model to reinforce the interaction and embedding capabilities between labeled and unlabeled samples during emotional highlights. Extensive experiments across multiple languages demonstrate the effectiveness of our method, requiring only 5-shot labeling in the source language.
【10】Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition
标题:用于阿拉伯语语音情感识别的混合CNN-转换器架构
链接:https://arxiv.org/abs/2604.07357
备注:7 pages, 4 figures. Master's thesis work, University of Science and Technology of Oran - Mohamed Boudiaf (USTO-MB)
摘要:使用机器学习从语音中识别情感已经成为一个活跃的研究领域,因为它在构建以人为本的应用程序中非常重要。然而,尽管许多研究是用英语、德语和其他欧洲和亚洲语言进行的,但由于注释数据集的可用性有限,阿拉伯语的研究仍然很少。在本文中,我们提出了一个阿拉伯语语音情感识别(SER)系统的基础上的混合CNN-Transformer架构。该模型利用卷积层从Mel频谱图输入和Transformer编码器中提取有区别的频谱特征,以捕获语音中的长距离时间依赖性。在EYASE(埃及阿拉伯语语音情感)语料库上进行了实验,所提出的模型达到了97.8%的准确率和0.98的宏观F1分数。这些结果证明了将卷积特征提取与基于注意力的建模相结合对阿拉伯语SER的有效性,并突出了基于transformer的方法在低资源语言中的潜力。
摘要:Recognizing emotions from speech using machine learning has become an active research area due to its importance in building human-centered applications. However, while many studies have been conducted in English, German, and other European and Asian languages, research in Arabic remains scarce because of the limited availability of annotated datasets. In this paper, we present an Arabic Speech Emotion Recognition (SER) system based on a hybrid CNN-Transformer architecture. The model leverages convolutional layers to extract discriminative spectral features from Mel-spectrogram inputs and Transformer encoders to capture long-range temporal dependencies in speech. Experiments were conducted on the EYASE (Egyptian Arabic speech emotion) corpus, and the proposed model achieved 97.8% accuracy and a macro F1-score of 0.98. These results demonstrate the effectiveness of combining convolutional feature extraction with attention-based modeling for Arabic SER and highlight the potential of Transformer-based approaches in low-resource languages.
【11】Contextual Earnings-22: A Speech Recognition Benchmark with Custom Vocabulary in the Wild
标题:上下文收入-22:一个语音识别基准与自定义词汇在野外
链接:https://arxiv.org/abs/2604.07354
摘要:语音到文本系统的准确性前沿已经达到了学术基准的水平。相比之下,工业基准和高风险领域的采用则表明情况并非如此。我们假设,两者之间的主要区别是上下文条件:学术基准是由经常遇到的一般词汇,这是相对容易识别相比,罕见的和上下文定义的自定义词汇,有不成比例的影响的可用性的演讲成绩单。尽管上下文语音到文本的进展,没有标准化的基准。我们引入了Contextual Earnings-22,这是一个建立在Earnings-22基础上的开放数据集,具有现实的自定义词汇上下文,以促进研究并揭示潜在的进展。我们为两种主要方法设置了六个强大的基线:关键字提示和关键字提升。实验表明,当从概念验证扩展到大规模系统时,两者都达到了相当的和显着提高的准确性。
摘要:The accuracy frontier of speech-to-text systems has plateaued on academic benchmarks.1 In contrast, industrial benchmarks and adoption in high-stakes domains suggest otherwise. We hypothesize that the primary difference between the two is contextual conditioning: Academic benchmarks are dominated by frequently encountered general vocabulary that is relatively easy to recognize compared with rare and context-defined custom vocabulary that has disproportionate impact on the usability of speech transcripts. Despite progress on contextual speech-to-text, there is no standardized benchmark. We introduce Contextual Earnings-22, an open dataset built upon Earnings-22, with realistic custom vocabulary contexts to foster research and reveal latent progress. We set six strong baselines for two dominant approaches: keyword prompting and keyword boosting. Experiments show both reach comparable and significantly improved accuracy when scaled from proof-of-concept to large-scale systems.
【12】Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
标题:重新思考基于LLM的ASB中的熵分配:了解语音编码器和LLM之间的动态
链接:https://arxiv.org/abs/2604.08003
摘要:将大型语言模型(LLM)集成到自动语音识别(ASR)中已成为一种主导模式。尽管最近基于LLM的ASR模型在公共基准测试中表现出了良好的性能,但平衡识别质量与延迟和开销仍然具有挑战性,而幻觉进一步限制了现实世界的部署。在这项研究中,我们从熵分配的角度重新审视基于LLM的ASR,并引入三个指标来表征训练范式如何在语音编码器和LLM之间分配熵减少。为了弥补熵分配效率低下的流行方法,我们提出了一个原则性的多阶段训练策略,以能力边界意识为基础,优化参数效率和幻觉鲁棒性。具体来说,我们重新设计了预训练策略,以减轻语音-文本模态差距,并进一步在对齐和联合SFT之间引入迭代异步SFT阶段,以保持功能解耦并约束编码器表示漂移。在普通话和英语基准测试上的实验表明,我们的方法仅使用2.3B参数就可以与最先进的模型实现竞争性性能,同时还可以通过我们的面向混合的设计有效地减轻幻觉。
摘要:Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models have shown promising performance on public benchmarks, it remains challenging to balance recognition quality with latency and overhead, while hallucinations further limit real-world deployment. In this study, we revisit LLM-based ASR from an entropy allocation perspective and introduce three metrics to characterize how training paradigms allocate entropy reduction between the speech encoder and the LLM. To remedy entropy-allocation inefficiencies in prevailing approaches, we propose a principled multi-stage training strategy grounded in capability-boundary awareness, optimizing parameter efficiency and hallucination robustness. Specifically, we redesign the pretraining strategy to alleviate the speech-text modality gap, and further introduce an iterative asynchronous SFT stage between alignment and joint SFT to preserve functional decoupling and constrain encoder representation drift. Experiments on Mandarin and English benchmarks show that our method achieves competitive performance with state-of-the-art models using only 2.3B parameters, while also effectively mitigating hallucinations through our decoupling-oriented design.
【1】Ring Mixing with Auxiliary Signal-to-Consistency-Error Ratio Loss for Unsupervised Denoising in Speech Separation
标题:语音分离中无监督去噪的辅助信号一致性错误比损失环混合
链接:https://arxiv.org/abs/2604.08415
备注:Submitted to Interspeech 2026
摘要:噪声语音分离系统通常在完全合成的混合物上训练,限制了对真实世界场景的推广。虽然在域(因此往往嘈杂)语音的混合物的训练是可能的,我们表明,这导致不理想的最优混合噪声保留在估计,由于背景噪声和损失函数的对称性的不可分割性。为了解决这个问题,我们提出了环形混合,一种在两种混合物中使用每个源的批量策略,以及一种新的信号与一致性误差比(SCER)辅助损失,惩罚来自不同混合物的相同源的不一致估计,打破对称性并激励去噪。在一个重击!基于基准测试,我们的方法可以将残余噪声减少一半以上,有效地学习仅从噪声记录中去噪。这为使用野外数据训练更通用的系统打开了大门,我们通过使用VoxCeleb的自然噪声语音训练的系统演示了这一点。
摘要:Noisy speech separation systems are typically trained on fully-synthetic mixtures, limiting generalization to real-world scenarios. Though training on mixtures of in-domain (thus often noisy) speech is possible, we show that this leads to undesirable optima where mixture noise is retained in the estimates, due to the inseparability of the background noises and the loss function's symmetry. To address this, we propose ring mixing, a batch strategy of using each source in two mixtures, alongside a new Signal-to-Consistency-Error Ratio (SCER) auxiliary loss penalizing inconsistent estimates of the same source from different mixtures, breaking symmetry and incentivizing denoising. On a WHAM!-based benchmark, our method can reduce residual noise by upwards of half, effectively learning to denoise from only noisy recordings. This opens the door to training more generalizable systems using in-the-wild data, which we demonstrate via systems trained using naturally-noisy speech from VoxCeleb.
【2】TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs
标题:TASU 2:用于语音LLM对齐和低资源自适应的可控CTC仿真
链接:https://arxiv.org/abs/2604.08384
摘要:语音LLM后训练越来越依赖于有效的跨模态对齐和强大的低资源适应,但收集大规模的音频文本对仍然是昂贵的。纯文本对齐方法(如TASU)通过模拟来自成绩单的CTC后验来减轻这种负担,但它们对不确定性和错误率的控制有限,使得课程设计在很大程度上具有启发性。我们提出了\textbf{TASU 2},一个可控的CTC模拟框架,在指定的WER范围内模拟CTC后验分布,产生更好地匹配声学解码接口的文本衍生监督。这使得有原则的培训后课程,顺利地改变监督的难度,没有TTS。在多个源到目标自适应设置中,TASU 2比TASU提高了域内和域外识别,并始终优于强基线,包括纯文本微调和基于TTS的增强,同时减轻了源域性能下降。
摘要:Speech LLM post-training increasingly relies on efficient cross-modal alignment and robust low-resource adaptation, yet collecting large-scale audio-text pairs remains costly. Text-only alignment methods such as TASU reduce this burden by simulating CTC posteriors from transcripts, but they provide limited control over uncertainty and error rate, making curriculum design largely heuristic. We propose \textbf{TASU2}, a controllable CTC simulation framework that simulates CTC posterior distributions under a specified WER range, producing text-derived supervision that better matches the acoustic decoding interface. This enables principled post-training curricula that smoothly vary supervision difficulty without TTS. Across multiple source-to-target adaptation settings, TASU2 improves in-domain and out-of-domain recognition over TASU, and consistently outperforms strong baselines including text-only fine-tuning and TTS-based augmentation, while mitigating source-domain performance degradation.
【3】Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework
标题:跟踪收件箱注意力:目光引导视听语音增强框架
链接:https://arxiv.org/abs/2604.08359
备注:Accepted to IEEE ICASSP 2026
摘要:本文提出了一种基于目光引导的视听语音增强(GG-AVSE)框架来解决鸡尾酒会问题。传统AVSE中的主要挑战是在多说话者环境中识别收听者的预期说话者。GG-AVSE通过利用注视方向作为目标说话人选择的监督线索来解决这个问题。具体来说,我们提出了GG-VM模块,该模块将凝视信号与YOLO 5 Face检测器相结合,以提取目标说话人的面部特征,并通过两种策略将其与预训练的AVSEMamba模型相结合:zero-shot合并和部分视觉微调。为了进行评估,我们引入了AVSEC 2-Gaze数据集。实验结果表明,GG-AVSE在无注视基线上实现了显著的性能提升:PESQ提高了10.08%(2.370到2.609),STOI提高了5.18%(0.8802到0.9258),SI-SDR提高了23.69%(9.16到11.33)。这些结果证实,凝视提供了一个有效的线索,解决目标说话人的歧义,并突出了GG-AVSE的可扩展性,为现实世界的应用。
摘要:This paper presents a Gaze-Guided Audio-Visual Speech Enhancement (GG-AVSE) framework to address the cocktail party problem. A major challenge in conventional AVSE is identifying the listener's intended speaker in multi-talker environments. GG-AVSE addresses this issue by exploiting gaze direction as a supervisory cue for target-speaker selection. Specifically, we propose the GG-VM module, which combines gaze signals with a YOLO5Face detector to extract the target speaker's facial features and integrates them with the pretrained AVSEMamba model through two strategies: zero-shot merging and partial visual fine-tuning. For evaluation, we introduce the AVSEC2-Gaze dataset. Experimental results show that GG-AVSE achieves substantial performance gains over gaze-free baselines: a 10.08% improvement in PESQ (2.370 to 2.609), a 5.18% improvement in STOI (0.8802 to 0.9258), and a 23.69% improvement in SI-SDR (9.16 to 11.33). These results confirm that gaze provides an effective cue for resolving target-speaker ambiguity and highlight the scalability of GG-AVSE for real-world applications.
【4】Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
标题:重新思考基于LLM的ASB中的熵分配:了解语音编码器和LLM之间的动态
链接:https://arxiv.org/abs/2604.08003
摘要:将大型语言模型(LLM)集成到自动语音识别(ASR)中已成为一种主导模式。尽管最近基于LLM的ASR模型在公共基准测试中表现出了良好的性能,但平衡识别质量与延迟和开销仍然具有挑战性,而幻觉进一步限制了现实世界的部署。在这项研究中,我们从熵分配的角度重新审视基于LLM的ASR,并引入三个指标来表征训练范式如何在语音编码器和LLM之间分配熵减少。为了弥补熵分配效率低下的流行方法,我们提出了一个原则性的多阶段训练策略,以能力边界意识为基础,优化参数效率和幻觉鲁棒性。具体来说,我们重新设计了预训练策略,以减轻语音-文本模态差距,并进一步在对齐和联合SFT之间引入迭代异步SFT阶段,以保持功能解耦并约束编码器表示漂移。在普通话和英语基准测试上的实验表明,我们的方法仅使用2.3B参数就可以与最先进的模型实现竞争性性能,同时还可以通过我们的面向混合的设计有效地减轻幻觉。
摘要:Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models have shown promising performance on public benchmarks, it remains challenging to balance recognition quality with latency and overhead, while hallucinations further limit real-world deployment. In this study, we revisit LLM-based ASR from an entropy allocation perspective and introduce three metrics to characterize how training paradigms allocate entropy reduction between the speech encoder and the LLM. To remedy entropy-allocation inefficiencies in prevailing approaches, we propose a principled multi-stage training strategy grounded in capability-boundary awareness, optimizing parameter efficiency and hallucination robustness. Specifically, we redesign the pretraining strategy to alleviate the speech-text modality gap, and further introduce an iterative asynchronous SFT stage between alignment and joint SFT to preserve functional decoupling and constrain encoder representation drift. Experiments on Mandarin and English benchmarks show that our method achieves competitive performance with state-of-the-art models using only 2.3B parameters, while also effectively mitigating hallucinations through our decoupling-oriented design.
【5】DeepFense: A Unified, Modular, and Extensible Framework for Robust Deepfake Audio Detection
标题:DeepFense:用于稳健Deepfake音频检测的统一、模块化和可扩展框架
链接:https://arxiv.org/abs/2604.08450
备注:Deepfense Toolkit
摘要:语音Deepfake检测是一个成熟的研究领域,具有不同的模型,数据集和训练策略。然而,缺乏标准化的实施和评估协议限制了重复性,基准测试和跨研究的比较。在这项工作中,我们介绍了DeepFense,这是一个全面的开源PyTorch工具包,集成了最新的架构,损失函数和增强管道,以及100多个配方。使用DeepFense,我们对400多个模型进行了大规模评估。我们的研究结果表明,虽然精心策划的训练数据提高了跨域泛化,但预训练的前端特征提取器的选择主导了整体性能方差。至关重要的是,我们在高性能模型中表现出严重的偏见,包括音频质量,说话者性别和语言。预计DeepFense将通过必要的工具促进实际部署,以解决公平的训练数据选择和前端微调。
摘要:Speech deepfake detection is a well-established research field with different models, datasets, and training strategies. However, the lack of standardized implementations and evaluation protocols limits reproducibility, benchmarking, and comparison across studies. In this work, we present DeepFense, a comprehensive, open-source PyTorch toolkit integrating the latest architectures, loss functions, and augmentation pipelines, alongside over 100 recipes. Using DeepFense, we conducted a large-scale evaluation of more than 400 models. Our findings reveal that while carefully curated training data improves cross-domain generalization, the choice of pre-trained front-end feature extractor dominates overall performance variance. Crucially, we show severe biases in high-performing models regarding audio quality, speaker gender, and language. DeepFense is expected to facilitate real-world deployment with the necessary tools to address equitable training data selection and front-end fine-tuning.
【6】Selective Attention System (SAS): Device-Addressed Speech Detection for Real-Time On-Device Voice AI
标题:选择性注意系统(SAS):用于实时设备上语音AI的设备编址语音检测
链接:https://arxiv.org/abs/2604.08412
摘要:我们研究了在预ASR边缘部署约束下的设备寻址语音检测,其中系统必须在严格的延迟和计算限制下决定是否在转录之前转发音频。我们发现,在多说话人的环境中,时间模糊的话语,这个任务更有效地建模为一个顺序路由问题的互动历史比作为一个话语本地分类任务。我们将其形式化为顺序设备寻址路由(SDAR),并提出了选择性注意系统(SAS),一个设备上的实现,实例化这个配方。 在一个60小时的多人英语测试集上,主要的纯音频配置达到了F1=0.86(精度=0.89,召回率=0.83);可选的摄像头,音频+视频融合将F1提高到0.95(精度=0.97,召回率=0.93)。在我们的评估方案下,删除因果交互历史(第3阶段)将音频+视频配置中的F1从0.95降至0.57+/-0.03。在测试的组件中,这是最大的观察到的消融效果,表明短期的相互作用的历史进行大量的决策相关的信息,在评价设置。SAS在ARM Cortex-A类硬件上完全在设备上运行(<150 ms延迟,<20 MB占用空间)。所有结果均来自主要以英语评价的专有数据集的内部评价;可共享5小时评价子集以进行独立验证(第8.8节)。
摘要:We study device-addressed speech detection under pre-ASR edge deployment constraints, where systems must decide whether to forward audio before transcription under strict latency and compute limits. We show that, in multi-speaker environments with temporally ambiguous utterances, this task is more effectively modelled as a sequential routing problem over interaction history than as an utterance-local classification task. We formalize this as Sequential Device-Addressed Routing (SDAR) and present the Selective Attention System (SAS), an on-device implementation that instantiates this formulation. On a held-out 60-hour multi-speaker English test set, the primary audio-only configuration achieves F1=0.86 (precision=0.89, recall=0.83); with an optional camera, audio+video fusion raises F1 to 0.95 (precision=0.97, recall=0.93). Removing causal interaction history (Stage~3) reduced F1 from 0.95 to 0.57+/-0.03 in the audio+video configuration under our evaluation protocol. Among the tested components, this was the largest observed ablation effect, indicating that short-horizon interaction history carries substantial decision-relevant information in the evaluated setting. SAS runs fully on-device on ARM Cortex-A class hardware (<150 ms latency, <20 MB footprint). All results are from internal evaluation on a proprietary dataset evaluated primarily in English; a 5-hour evaluation subset may be shared for independent verification (Section 8.8).
【7】Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
标题:语义-情感共鸣嵌入:跨语言语音情感识别的半监督范式
链接:https://arxiv.org/abs/2604.07417
备注:Main paper (6 pages). Accepted for publication by IEEE International conference on Multimedia and Expo 2026 (ICME 2026)
摘要:跨语言语音情感识别(CLSER)旨在识别看不见的语言中的情感状态。然而,现有的方法严重依赖于完整标签的语义同步和静态特征稳定性,阻碍了低资源语言达到高资源性能。为了解决这个问题,我们提出了一个半监督框架的基础上语义情感共振嵌入(SERE),跨语言的动态特征范例,既不需要目标语言标签,也不需要翻译对齐。具体来说,SERE使用少量的标记样本构建情感语义结构。它通过瞬时共振场(IRF)学习人类的情感体验,使未标记的样本能够自组织成这种结构。这实现了半监督语义指导和结构发现。此外,我们设计了一个三重共振交互链(TRIC)损失,使模型能够加强在情感亮点标记和未标记的样本之间的交互和嵌入能力。跨多种语言的广泛实验证明了我们的方法的有效性,只需要5杆标记的源语言。
摘要:Cross-lingual Speech Emotion Recognition (CLSER) aims to identify emotional states in unseen languages. However, existing methods heavily rely on the semantic synchrony of complete labels and static feature stability, hindering low-resource languages from reaching high-resource performance. To address this, we propose a semi-supervised framework based on Semantic-Emotional Resonance Embedding (SERE), a cross-lingual dynamic feature paradigm that requires neither target language labels nor translation alignment. Specifically, SERE constructs an emotion-semantic structure using a small number of labeled samples. It learns human emotional experiences through an Instantaneous Resonance Field (IRF), enabling unlabeled samples to self-organize into this structure. This achieves semi-supervised semantic guidance and structural discovery. Additionally, we design a Triple-Resonance Interaction Chain (TRIC) loss to enable the model to reinforce the interaction and embedding capabilities between labeled and unlabeled samples during emotional highlights. Extensive experiments across multiple languages demonstrate the effectiveness of our method, requiring only 5-shot labeling in the source language.
机器翻译由腾讯交互翻译提供,仅供参考
