微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 1 篇
2. 说话人识别、验证与分离 1 篇
3. 语音增强、降噪与音频修复 1 篇
4. 音频事件检测与场景理解 2 篇
5. 音乐信息检索与音乐生成 1 篇
6. 其他/综合语音音频 1 篇
1. 语音识别与关键词检测 | 1 篇
1. H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR
H-SAGE: 基于MoE的多说话人ASR的整体说话人感知引导专家
AI 总结:针对多说话人ASR中高重叠语音识别难题,提出H-SAGE方法,通过说话人感知全局编码器、重叠感知损失和整体门控机制,提升专家协作,在LibriSpeechMix上取得一致改进。
链接:https://arxiv.org/abs/2607.01566
作者:Yujie Guo, Jiaming Zhou, Yuhang Jia, Yang chen, Yong Qin
英文摘要:Multi-talker Automatic Speech Recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech, particularly under complex high-overlap conditions. While recent Mixture-of-Experts (MoE) approaches have shown promise, they typically rely on frame-independent routing that leads to temporal myopia, and depend solely on the downstream ASR objective, which results in implicit and ungrounded representation learning. To address these limitations, we propose Holistic Speaker-Aware Guided Experts (H-SAGE) for MoE-based MTASR. Specifically, we introduce a Speaker-Aware Global Encoder to capture long-term dependencies, supervised by an auxiliary Overlap-Aware Loss that explicitly guides the model to discern acoustic states. Furthermore, we design a Holistic Gating Mechanism to arbitrate expert selection by jointly evaluating global context and local details. Experiments on LibriSpeechMix demonstrate that H-SAGE achieves consistent improvements over strong baselines, particularly in complex scenarios, validating that explicit acoustic guidance effectively enhances expert collaboration. Our code can be found at this https URL.
2. 说话人识别、验证与分离 | 1 篇
2. Speaker head orientation estimation with a single microphone array using phase spectrogram features
使用单个麦克风阵列基于相位谱图特征的说话人头部朝向估计
AI 总结:提出利用单个麦克风阵列的短时傅里叶变换相位分量作为深度神经网络输入,结合卷积、循环和自注意力层,在模拟和真实数据上学习,实现高精度头部朝向估计,平均角度误差达11.3度。
链接:https://arxiv.org/abs/2607.02129
作者:Balint Turi, Archontis Politis, Parthasaarathy Sudarsanam, Tuomas Virtanen
英文摘要:Estimating a speaker's head orientation from audio can provide valuable information in smart environments, meetings, and driver monitoring. We propose a novel approach that leverages the phase component of the short-time Fourier transform from a single microphone array as input to a deep neural network combining convolutional, recurrent, and self-attention layers. Unlike prior methods that use physics-informed handcrafted features or raw waveform inputs, our approach enables robust learning from simulated and real data. Trained on a large-scale dataset generated with voice directivity patterns and fine-tuned on real recordings, our model achieves state-of-the-art accuracy, outperforming baselines under both clean and noisy conditions. Personalization experiments further demonstrate significant gains, reaching a mean angular error of 11.3 degrees when adapting to individual users and environments.
3. 语音增强、降噪与音频修复 | 1 篇
3. RT-Tango: Real-Time Distributed Binaural Speech Enhancement for Low-Power Hearing Aid Devices
RT-Tango:面向低功耗助听器设备的实时分布式双耳语音增强
AI 总结:提出RT-Tango实时分布式双耳语音增强框架,通过ERB特征压缩、轻量级分组循环掩码估计和时间稀疏化降低计算成本,结合非对称STFT和因果推理实现8ms超低延迟,在资源受限平台上达到竞争性增强效果。
链接:https://arxiv.org/abs/2607.01834
机构:Université Paris-Saclay, CEA, List(巴黎-萨克雷大学,法国原子能委员会,List实验室); Université de Lorraine, CNRS, Inria, LORIA(洛林大学,法国国家科学研究中心,法国国家信息与自动化研究所,洛林计算机科学及其应用实验室)
作者:Z. Benslimane, P. Chouteau, M. Poreba, F. Auzanneau, M. Szczepanski, F. Chersi, R. Serizel
英文摘要:Real-time binaural speech enhancement is constrained by latency, computational cost, and inter-device communication, yet existing efficient solutions predominantly address single-channel settings. In this paper, we introduce RT-Tango, a real-time distributed binaural speech enhancement framework designed for streaming on resource-constrained platforms and specifically for hearing aids. RT-Tango relies on a two-stage distributed architecture combining perceptually motivated ERB feature compression, lightweight grouped recurrent mask estimation, and temporal sparsification to reduce computational cost. Stringent latency constraints are addressed by decoupling spectral resolution from algorithmic delay using an asymmetric STFT, together with causal recurrent inference and online estimation of spatial statistics. Experimental results show that RT-Tango achieves competitive speech enhancement while significantly reducing MACs operations and functioning at ultra-low latencies as low as 8 ms.
4. 音频事件检测与场景理解 | 2 篇
4. A Multi-Branch Hierarchy-Aware Framework for Heterogeneous Audio Classification
异构音频分类的多分支层次感知框架
AI 总结:提出基于CLAP音频文本表示的多分支层次感知框架,通过扩展训练集、特征分支和层次感知分类器及KNN后处理,在DCASE 2026任务1上实现层次F1分数80.84%。
链接:https://arxiv.org/abs/2607.01974
作者:Beile Ning, Jiayi Yu, Zitong Wang, Yufei Hu, Wenjun Xu, Yuanhang Qian, Zhongxin Bai, Gongping Huang
英文摘要:This technical report describes our system for Task 1 of the DCASE 2026 Challenge, which aims to classify heterogeneous audio recordings according to the Broad Sound Taxonomy (BST). The task requires both accurate second-level prediction and consistency with the top-level taxonomy. Our system is built on CLAP-based audio-text representations and is improved along three strategies: expanding the training set with a filtered subset of BSD35k, enhancing acoustic modeling with feature-specific branches, and refining predictions using hierarchy-aware classifiers and KNN-based post-processing. Among the acoustic features considered, the log-STFT branch provides the strongest single-model performance. With KNN-based post-processing, our best single system achieves a hierarchical F1 score (Hier. F1) of 80.84% on the BSD10k-v1.2 set under the same evaluation protocol as the baseline. We further construct ensemble systems by combining models with complementary acoustic features and classification heads, achieving Hier. F1 scores of 81.25% and 81.18%, respectively.
5. SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios
SelectTSL:复杂场景中提示引导的选择性目标声音定位
AI 总结:提出SelectTSL架构,通过提示引导的选择性注意力模块增强相位差,实现多源声场中仅定位用户指定目标,并估计到达方向和目标源数量。
链接:https://arxiv.org/abs/2607.02343
作者:Ziyang Jiang, Yu Chen, Zexu Pan, Xinyuan Qian, Bowen Xing, Ivor W. Tsang, Xu-Cheng Yin, Haizhou Li
英文摘要:Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learning-based systems. Sound source localization (SSL) has achieved remarkable success with deep learning, yet most methods localize all active sources without selectivity. Conversely, target sound extraction (TSE) extracts sources using multimodal prompts but typically fails to preserve the multichannel spatial information required for accurate localization. To bridge this gap, we formulate the task of prompt-guided selective target sound localization and propose SelectTSL, an end-to-end architecture that localizes only the user-specified target in multi-source acoustic scenes. Specifically, we design a target-aware selective localization strategy that employs a Prompt-Guided Selective Attention Module (PGSA) to generate prompt-informed embeddings. These embeddings guide an inter-channel phase difference (IPD) enhancer to refine raw phase cues, fusing with target magnitudes to jointly estimate direction of arrival (DoA) and target-source cardinality, i.e., the number of target sound sources. This coupled design effectively focuses on the user-specified target spatial cues for selective localization and also handles time-varying numbers of target sources. Extensive experiments on both synthetic data and real-world recordings demonstrate that our proposed method consistently outperforms other baselines and exhibits robust generalization to real acoustic environments.
5. 音乐信息检索与音乐生成 | 1 篇
6. UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation
UT-AISTimprt 提交至 ICME 2026 学术文本到音乐生成大挑战
AI 总结:本研究针对低数据和小模型设置下的文本到音乐生成,提出基于文本或音频嵌入的聚类批量采样策略,分析模态和聚类粒度的影响,发现文本嵌入聚类在客观指标上更优,中等聚类数在客观指标最佳而大聚类数在听觉测试中结构更连贯。
链接:https://arxiv.org/abs/2607.01669
作者:Shunsuke Yoshida, Yu-Hua Chen, Satoru Fukayama
英文摘要:This work investigates the effect of batch sampling strategies during training for text-to-audio music generation under low-data and small-scale model settings. This paper describes our approach and findings for the ICME 2026 Grand Challenge on Academic Text-to-Music Generation. Training data are clustered using either text embeddings or audio embeddings, and samples with similar characteristics are grouped within the same mini-batch to mitigate gradient interference. The effects of modality and cluster granularity on clustering are analyzed. Results show that clustering based on text embeddings achieves better performance on objective evaluation metrics than clustering based on audio embeddings. In addition, different cluster granularity leads to different behaviors across evaluation criteria: a moderate number of clusters performs best on objective metrics, while a larger number of clusters tends to exhibit music with more coherent structure in listening tests.
6. 其他/综合语音音频 | 1 篇
7. Quantifying the Uncertainty of Blindly Estimated Room Embeddings Using a Dispersion-Calibrated Score
量化盲估计房间嵌入的不确定性:基于分散校准的分数
AI 总结:提出一种从混响语音中学习鲁棒房间嵌入和不确定性分数的框架,通过多视图KL对齐和对比学习增强鲁棒性,利用基于排序的目标校准不确定性分数,实现单次话语下的有效选择性预测。
链接:https://arxiv.org/abs/2607.01527
机构:Centre for Vision Speech and Signal Processing, University of Surrey(萨里大学视觉、语音与信号处理中心); International Audio Laboratories Erlangen(埃尔朗根国际音频实验室); Fraunhofer Institute for Integrated Circuits IIS(弗劳恩霍夫集成电路研究所)
作者:Yang Xiang, Philipp Götz, Emanuël A. P. Habets, Andreas Walther, Wenwu Wang, Philip J. B. Jackson
英文摘要:Room embeddings derived from reverberant speech are often unreliable: speech content and recording degradation can alter the representation even when speaker, room, and source-receiver geometry remain unchanged, degrading downstream task performance. We propose a framework that learns room embeddings robust to speech-content variation and a representation-level uncertainty score from reverberant speech without downstream-task supervision. The embedding is anchored to a structured room impulse response (RIR) latent space and trained using a multi-view data structure with Kullback-Leibler (KL)-based alignment; a multi-positive contrastive term further refines robustness. A lightweight uncertainty head is calibrated using the dispersion of corruption-induced embeddings and optimized with a rank-based objective. Across waveform- and spectrogram-level corruptions, the score is consistent with representation dispersion and enables effective selective prediction while requiring only a single utterance at inference.
