微信公众号:arXiv_Daily
cs.SD语音
【1】Transformer Architectures for Respiratory Sound Analysis and Multimodal Diagnosis
标题:用于呼吸声分析和多模式诊断的Transformer架构
链接:https://arxiv.org/abs/2601.14227
备注:7 pages, 4 figures
摘要:呼吸音分析是筛查哮喘和其他肺部病变的重要工具,但传统的听诊仍然是主观的和经验依赖的。我们之前的研究使用DenseNet 201建立了CNN基线,该基线在分类呼吸音方面表现出高灵敏度。在这项工作中,我们(i)适应音频频谱图Transformer(AST)的呼吸声分析和(ii)评估多模态视觉语言模型(VLM),集成频谱图与结构化的患者元数据。 AST从公开可用的权重初始化,并在包含每个诊断数百个记录的医疗数据集上进行微调。VLM实验使用了一个紧凑的Moondream类型的模型,该模型处理谱图图像以及结构化的文本提示(性别,年龄,记录地点),以输出JSON格式的诊断。结果表明,AST达到约97%的准确性,F1评分约为97%,ROC AUC为0.98,用于哮喘检测,显著优于内部CNN基线和典型的外部基准。VLM达到86-87%的准确度,执行与CNN基线的匹配,同时证明了将临床背景整合到推理过程中的能力。这些结果证实了自我注意力在声学筛查中的有效性,并强调了多模式架构在整体诊断工具中的潜力。
摘要:Respiratory sound analysis is a crucial tool for screening asthma and other pulmonary pathologies, yet traditional auscultation remains subjective and experience-dependent. Our prior research established a CNN baseline using DenseNet201, which demonstrated high sensitivity in classifying respiratory sounds. In this work, we (i) adapt the Audio Spectrogram Transformer (AST) for respiratory sound analysis and (ii) evaluate a multimodal Vision-Language Model (VLM) that integrates spectrograms with structured patient metadata. AST is initialized from publicly available weights and fine-tuned on a medical dataset containing hundreds of recordings per diagnosis. The VLM experiment uses a compact Moondream-type model that processes spectrogram images alongside a structured text prompt (sex, age, recording site) to output a JSON-formatted diagnosis. Results indicate that AST achieves approximately 97% accuracy with an F1-score around 97% and ROC AUC of 0.98 for asthma detection, significantly outperforming both the internal CNN baseline and typical external benchmarks. The VLM reaches 86-87% accuracy, performing comparably to the CNN baseline while demonstrating the capability to integrate clinical context into the inference process. These results confirm the effectiveness of self-attention for acoustic screening and highlight the potential of multimodal architectures for holistic diagnostic tools.
【2】ConceptCaps -- a Distilled Concept Dataset for Interpretability in Music Models
标题:ConceptCaps --音乐模型可解释性的提炼概念数据集
链接:https://arxiv.org/abs/2601.14157
摘要:基于概念的可解释性方法(如TCAV)需要清晰的、分离良好的正例和反例。现有的音乐数据集缺乏这种结构:标签是稀疏的,嘈杂的,或不明确的。我们介绍ConceptCaps,这是一个包含23k音乐字幕音频三元组的数据集,具有来自200个属性分类的明确标签。我们的管道将语义建模与文本生成分开:VAE学习合理的属性共现模式,微调的LLM将属性列表转换为专业描述,MusicGen合成相应的音频。这种分离提高了端到端方法的一致性和可控性。我们通过音频文本对齐(CLAP),语言质量指标(BERTScore,MAUVE)和TCAV分析验证数据集,确认概念探针恢复音乐有意义的模式。数据集和代码可在线获取。
摘要:Concept-based interpretability methods like TCAV require clean, well-separated positive and negative examples for each concept. Existing music datasets lack this structure: tags are sparse, noisy, or ill-defined. We introduce ConceptCaps, a dataset of 23k music-caption-audio triplets with explicit labels from a 200-attribute taxonomy. Our pipeline separates semantic modeling from text generation: a VAE learns plausible attribute co-occurrence patterns, a fine-tuned LLM converts attribute lists into professional descriptions, and MusicGen synthesizes corresponding audio. This separation improves coherence and controllability over end-to-end approaches. We validate the dataset through audio-text alignment (CLAP), linguistic quality metrics (BERTScore, MAUVE), and TCAV analysis confirming that concept probes recover musically meaningful patterns. Dataset and code are available online.
【3】PRiSM: Benchmarking Phone Realization in Speech Models
标题:PRiSM:语音模型中的电话实现基准测试
链接:https://arxiv.org/abs/2601.14046
摘要:电话识别(PR)是跨语言语音处理和语音分析的语言无关建模的原子接口。尽管在开发PR系统方面进行了长期的努力,但目前的评估仅测量表面水平的转录准确性。我们介绍PRiSM,第一个开源基准,旨在通过PR系统的内在和外在评估暴露语音感知中的盲点。PRiSM通过转录和代表探针,简化了基于转录的评估,并评估了临床,教育和多语言环境中的下游效用。我们发现,在训练过程中,不同的语言接触是PR性能的关键,编码器-CTC模型是最稳定的,而专业的PR模型仍然优于大型音频语言模型。PRiSM发布代码、配方和数据集,将该领域推向具有强大语音能力的多语言语音模型:https://github.com/changelinglab/prism。
摘要:Phone recognition (PR) serves as the atomic interface for language-agnostic modeling for cross-lingual speech processing and phonetic analysis. Despite prolonged efforts in developing PR systems, current evaluations only measure surface-level transcription accuracy. We introduce PRiSM, the first open-source benchmark designed to expose blind spots in phonetic perception through intrinsic and extrinsic evaluation of PR systems. PRiSM standardizes transcription-based evaluation and assesses downstream utility in clinical, educational, and multilingual settings with transcription and representation probes. We find that diverse language exposure during training is key to PR performance, encoder-CTC models are the most stable, and specialized PR models still outperform Large Audio Language Models. PRiSM releases code, recipes, and datasets to move the field toward multilingual speech models with robust phonetic ability: https://github.com/changelinglab/prism.
【4】Towards Effective Negation Modeling in Joint Audio-Text Models for Music
标题:音乐联合音频文本模型中的有效否定建模
链接:https://arxiv.org/abs/2601.13931
备注:Accepted at IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026
摘要:联合音频文本模型被广泛用于音乐检索,但他们的斗争与语义现象,如否定。否定是区分音乐元素(例如,“有人声”对“没有人声”),但是当前的系统不能可靠地表示这一点。在这项工作中,我们通过在百万歌曲数据集上使用LP-MusicCaps-MSD字幕从头开始训练CLAP模型来研究和减轻这种限制。我们通过文本增强和基于差异的对比损失引入否定,旨在明确分离联合嵌入空间中的原始和否定字幕。为了评估进展,我们提出了两个协议,帧否定建模检索和二进制分类任务。实验表明,这两种方法,单独和组合,提高否定处理,同时在很大程度上保持检索性能。
摘要:Joint audio-text models are widely used for music retrieval, yet they struggle with semantic phenomena such as negation. Negation is fundamental for distinguishing the absence (or presence) of musical elements (e.g., "with vocals" vs. "without vocals"), but current systems fail to represent this reliably. In this work, we investigate and mitigate this limitation by training CLAP models from scratch on the Million Song Dataset with LP-MusicCaps-MSD captions. We introduce negation through text augmentation and a dissimilarity-based contrastive loss, designed to explicitly separate original and negated captions in the joint embedding space. To evaluate progress, we propose two protocols that frame negation modeling as retrieval and binary classification tasks. Experiments demonstrate that both methods, individually and combined, improve negation handling while largely preserving retrieval performance.
【5】Emotion and Acoustics Should Agree: Cross-Level Inconsistency Analysis for Audio Deepfake Detection
标题:情感和声学应该达成一致:音频深度伪造检测的跨级别不一致性分析
链接:https://arxiv.org/abs/2601.13847
备注:Accepted by ICASSP 2026
摘要:音频深度伪造检测(ADD)旨在从真实语音中检测欺骗语音。大多数先前的研究假设,更强的相关性内或跨声学和情感特征意味着真实性,因此专注于增强或测量这种相关性。然而,现有的方法往往孤立地对待声学和情感特征,或者依赖于相关性度量,这忽略了它们之间的细微去相干化,并平滑了突然的不连续性。为了解决这些问题,我们提出了EAI-ADD,它将跨级别的情感声学不一致作为主要的检测信号。我们首先将情感和声学表征投射到一个类似的空间中。然后,我们逐步整合帧级和话语级情感特征与声学特征,以捕获跨时间粒度的跨级别情感声学不一致。在ASVspoof 2019 LA和2021 LA数据集上的实验结果表明,所提出的EAI-ADD优于基线,为音频反欺骗检测提供了更有效的解决方案。
摘要:Audio Deepfake Detection (ADD) aims to detect spoof speech from bonafide speech. Most prior studies assume that stronger correlations within or across acoustic and emotional features imply authenticity, and thus focus on enhancing or measuring such correlations. However, existing methods often treat acoustic and emotional features in isolation or rely on correlation metrics, which overlook subtle desynchronization between them and smooth out abrupt discontinuities. To address these issues, we propose EAI-ADD, which treats cross level emotion acoustic inconsistency as the primary detection signal. We first project emotional and acoustic representations into a comparable space. Then we progressively integrate frame level and utterance level emotion features with acoustic features to capture cross level emotion acoustic inconsistencies across different temporal granularities. Experimental results on the ASVspoof 2019LA and 2021LA datasets demonstrate that the proposed EAI-ADD outperforms baselines, providing a more effective solution for audio anti spoofing detection.
【6】Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis
标题:哈比比:奠定统一方言阿拉伯语语音合成的开源基础
链接:https://arxiv.org/abs/2601.13802
摘要:一个显着的差距仍然存在于语音合成的研究和开发阿拉伯方言,特别是从统一的建模的角度来看。尽管阿拉伯方言具有很高的实用价值,但其固有的语言复杂性,加上缺乏标准化数据、基准和评估指南,使研究人员转向更安全的领域。为了弥合这一鸿沟,我们提出了Habibi,这是一套专门和统一的文本到语音模型,利用现有的开源ASR语料库,通过语言学知识的课程学习来支持各种高到低资源的阿拉伯方言。我们的方法在生成质量上优于领先的商业服务,同时通过有效的上下文学习保持可扩展性,而不需要文本变音符号。我们致力于开源该模型,并为多方言阿拉伯语语音合成创建第一个系统基准。此外,通过确定过程中的关键挑战和建立评估标准,我们的目标是为后续研究提供坚实的基础。资源请访问https://SWivid.github.io/Habibi/。
摘要:A notable gap persists in speech synthesis research and development for Arabic dialects, particularly from a unified modeling perspective. Despite its high practical value, the inherent linguistic complexity of Arabic dialects, further compounded by a lack of standardized data, benchmarks, and evaluation guidelines, steers researchers toward safer ground. To bridge this divide, we present Habibi, a suite of specialized and unified text-to-speech models that harnesses existing open-source ASR corpora to support a wide range of high- to low-resource Arabic dialects through linguistically-informed curriculum learning. Our approach outperforms the leading commercial service in generation quality, while maintaining extensibility through effective in-context learning, without requiring text diacritization. We are committed to open-sourcing the model, along with creating the first systematic benchmark for multi-dialect Arabic speech synthesis. Furthermore, by identifying the key challenges in and establishing evaluation standards for the process, we aim to provide a solid groundwork for subsequent research. Resources at https://SWivid.github.io/Habibi/ .
【7】GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation Tasks
标题:GOMPSNR:反映音频生成任务的信噪比度量
链接:https://arxiv.org/abs/2601.13758
备注:Accepted by AAAI 2026
摘要:在音频生成领域中,信噪比(SNR)长期以来一直充当用于评估音频质量的客观度量。然而,最近的研究表明,SNR及其变体并不总是与人类感知高度相关,这促使我们提出问题:为什么SNR无法测量音频质量?如何提高其作为客观指标的可靠性?在本文中,我们确定的相位距离的测量不充分的一个关键因素,并建议重新制定SNR与专门设计的相位距离条款,产生一个改进的度量GOMPSNR。我们进一步扩展了新提出的配方,得到两个新的类别的损失函数,分别对应于幅度引导相位细化和联合幅度相位优化。此外,对不同损失函数的最优组合进行了大量的实验。先进的神经声码器的实验结果表明,我们提出的GOMPSNR具有更可靠的误差测量比SNR。与此同时,我们提出的损失函数大大提高了模型性能,并且我们精心选择的不同损失函数的组合进一步优化了整体模型能力。
摘要:In the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR and its variants are not always highly correlated with human perception, prompting us to raise the questions: Why does SNR fail in measuring audio quality? And how to improve its reliability as an objective metric? In this paper, we identify the inadequate measurement of phase distance as a pivotal factor and propose to reformulate SNR with specially designed phase-distance terms, yielding an improved metric named GOMPSNR. We further extend the newly proposed formulation to derive two novel categories of loss function, corresponding to magnitude-guided phase refinement and joint magnitude-phase optimization, respectively. Besides, extensive experiments are conducted for an optimal combination of different loss functions. Experimental results on advanced neural vocoders demonstrate that our proposed GOMPSNR exhibits more reliable error measurement than SNR. Meanwhile, our proposed loss functions yield substantial improvements in model performance, and our wellchosen combination of different loss functions further optimizes the overall model capability.
【8】Performance and Complexity Trade-off Optimization of Speech Models During Training
标题:训练期间语音模型的性能和复杂性权衡优化
链接:https://arxiv.org/abs/2601.13704
摘要:在语音机器学习中,神经网络模型通常通过选择具有固定层大小和结构的架构来设计。然后训练这些模型,以最大限度地提高与任务目标一致的指标的性能。虽然总体架构通常由任务的先验知识指导,但各个层的大小通常是按顺序选择的。然而,这种方法不能保证性能和计算复杂度之间的最佳权衡;因此,通常采用诸如权重量化或模型修剪的事后方法来降低计算成本。这是因为随机梯度下降(SGD)方法只能优化可微函数,而影响计算复杂性的因素,如层大小和每秒浮点运算(FLOP/s),是不可微的,需要在训练过程中修改模型结构。我们提出了一种基于特征噪声注入的重新参数化技术,该技术能够在使用基于SGD的方法进行训练期间联合优化性能和计算复杂度。与传统的修剪方法不同,我们的方法允许模型大小动态优化的目标性能复杂性的权衡,而不依赖于启发式标准来选择删除的权重或结构。我们通过三个案例研究证明了我们的方法的有效性,包括一个合成的例子和两个实际的现实世界中的应用:语音活动检测和音频反欺骗。与我们的工作相关的代码是公开的,以鼓励进一步的研究。
摘要:In speech machine learning, neural network models are typically designed by choosing an architecture with fixed layer sizes and structure. These models are then trained to maximize performance on metrics aligned with the task's objective. While the overall architecture is usually guided by prior knowledge of the task, the sizes of individual layers are often chosen heuristically. However, this approach does not guarantee an optimal trade-off between performance and computational complexity; consequently, post hoc methods such as weight quantization or model pruning are typically employed to reduce computational cost. This occurs because stochastic gradient descent (SGD) methods can only optimize differentiable functions, while factors influencing computational complexity, such as layer sizes and floating-point operations per second (FLOP/s), are non-differentiable and require modifying the model structure during training. We propose a reparameterization technique based on feature noise injection that enables joint optimization of performance and computational complexity during training using SGD-based methods. Unlike traditional pruning methods, our approach allows the model size to be dynamically optimized for a target performance-complexity trade-off, without relying on heuristic criteria to select which weights or structures to remove. We demonstrate the effectiveness of our method through three case studies, including a synthetic example and two practical real-world applications: voice activity detection and audio anti-spoofing. The code related to our work is publicly available to encourage further research.
【9】DistilMOS: Layer-Wise Self-Distillation For Self-Supervised Learning Model-Based MOS Prediction
标题:DistilMOS:基于自监督学习模型的MOS预测的分层自蒸馏
链接:https://arxiv.org/abs/2601.13700
备注:Accepted to ICASSP 2026
摘要:随着自监督学习(SSL)的发展,微调预训练的SSL模型用于平均意见得分(MOS)预测已经取得了最先进的性能。然而,在微调过程中,这些基于SSL的MOS预测模型往往会灾难性地忘记预先训练的知识,并倾向于过拟合训练集,从而导致泛化性能差。在这项研究中,我们提出了DistilMOS,这是一种新的方法,它不仅学习预测MOS,还学习预测通过对预训练SSL模型中每层的隐藏表示进行聚类而获得的令牌ID。这些分层令牌目标作为自蒸馏信号,使MOS预测模型能够从SSL模型中提取丰富的内部知识,从而提高预测精度和泛化能力。实验结果表明,该方法在域内和域外的预测性能均优于标准的基于SSL的MOS预测模型,验证了该方法的有效性和实用性.
摘要:With the advancement of self-supervised learning (SSL), fine-tuning pretrained SSL models for mean opinion score (MOS) prediction has achieved state-of-the-art performance. However, during fine-tuning, these SSL-based MOS prediction models often suffer from catastrophic forgetting of the pretrained knowledge and tend to overfit the training set, resulting in poor generalization performance. In this study, we propose DistilMOS, a novel method that learns to predict not only MOS but also token IDs obtained by clustering the hidden representations of each layer in the pretrained SSL model. These layer-wise token targets serve as self-distillation signals that enables the MOS prediction model to extract rich internal knowledge from SSL models, enhancing both prediction accuracy and generalization capability. Experimental evaluations demonstrate that our method significantly outperforms standard SSL-based MOS prediction models on both in-domain and out-of-domain evaluations, verifying the effectiveness and practicality of the proposed method.
【10】Ultra-Lightweight Network for Ship-Radiated Sound Classification on Embedded Deployment
标题:嵌入式部署中船舶辐射声音分类的超轻量级网络
链接:https://arxiv.org/abs/2601.13679
备注:This manuscript is under review at IEEE Geoscience and Remote Sensing Letters
摘要:这封信介绍了ShuffleFAC,一个轻量级的声学模型,船舶辐射声分类资源有限的海事监测系统。ShuffleFAC使用可分离卷积、逐点组卷积和通道混洗将频率感知卷积集成到以效率为导向的主干中,从而实现低计算成本的频率敏感特征提取。在DeepShip数据集上的实验表明,ShuffleFAC实现了具有竞争力的性能,大大降低了复杂性。特别是,ShuffleFAC($γ=16$)使用39 K参数和3.06M MAC获得了71.45 $\pm $1.18%的宏F1分数,并在Raspberry Pi上实现了6.05 $\pm $0.95ms的推理延迟。与MicroNet 0相比,它将宏F1分数提高了1.82%,同时将模型大小减少了9.7倍,延迟减少了2.5倍。这些结果表明,ShuffleFAC是适合实时嵌入式UATR。
摘要:This letter presents ShuffleFAC, a lightweight acoustic model for ship-radiated sound classification in resource-constrained maritime monitoring systems. ShuffleFAC integrates Frequency-Aware convolution into an efficiency-oriented backbone using separable convolution, point-wise group convolution, and channel shuffle, enabling frequency-sensitive feature extraction with low computational cost. Experiments on the DeepShip dataset show that ShuffleFAC achieves competitive performance with substantially reduced complexity. In particular, ShuffleFAC ($γ=16$) attains a macro F1-score of 71.45 $\pm$ 1.18% using 39K parameters and 3.06M MACs, and achieves an inference latency of 6.05 $\pm$ 0.95ms on a Raspberry Pi. Compared with MicroNet0, it improves macro F1-score by 1.82 % while reducing model size by 9.7x and latency by 2.5x. These results indicate that ShuffleFAC is suitable for real-time embedded UATR.
【11】Fusion Segment Transformer: Bi-Directional Attention Guided Fusion Network for AI-Generated Music Detection
标题:融合片段Transformer:用于人工智能生成音乐检测的双向注意力引导融合网络
链接:https://arxiv.org/abs/2601.13647
摘要:随着人工智能生成技术的兴起,任何人现在都可以轻松地创建和部署人工智能生成的音乐,这加剧了对解决版权和所有权问题的技术解决方案的需求。虽然现有的工作主要集中在短音频,全音频检测的挑战,这需要建模长期的结构和上下文,仍然没有得到充分的探讨。为了解决这个问题,我们提出了一个改进版本的段Transformer,称为融合段Transformer。与我们以前的工作一样,我们使用不同的特征提取器从短音乐片段中提取内容嵌入。此外,我们通过引入门控融合层来增强全音频AI生成音乐检测的架构,该层有效地集成了内容和结构信息,从而能够捕获长期上下文。在SONICS和AIME数据集上的实验表明,我们的方法优于以前的模型和最近的基线,在人工智能生成的音乐检测中取得了最先进的结果。
摘要:With the rise of generative AI technology, anyone can now easily create and deploy AI-generated music, which has heightened the need for technical solutions to address copyright and ownership issues. While existing works mainly focused on short-audio, the challenge of full-audio detection, which requires modeling long-term structure and context, remains insufficiently explored. To address this, we propose an improved version of the Segment Transformer, termed the Fusion Segment Transformer. As in our previous work, we extract content embeddings from short music segments using diverse feature extractors. Furthermore, we enhance the architecture for full-audio AI-generated music detection by introducing a Gated Fusion Layer that effectively integrates content and structural information, enabling the capture of long-term context. Experiments on the SONICS and AIME datasets show that our approach outperforms the previous model and recent baselines, achieving state-of-the-art results in AI-generated music detection.
【12】Motion-to-Response Content Generation via Multi-Agent AI System with Real-Time Safety Verification
标题:通过具有实时安全验证的多智能体人工智能系统生成动作响应内容
链接:https://arxiv.org/abs/2601.13589
摘要:本文提出了一种多智能体人工智能系统,该系统基于音频导出的情感信号实时生成面向响应的媒体内容。与主要关注分类准确性的传统语音情感识别研究不同,我们的方法强调通过专业AI代理的结构化管道将推断的情感状态转换为安全,年龄合适和可控的响应内容。所提出的系统包括四个合作代理:(1)一个情感识别代理与基于CNN的声学特征提取,(2)响应策略决策代理的情感映射到响应模式,(3)内容参数生成代理产生媒体控制参数,和(4)安全验证代理执行年龄适当性和刺激约束。我们引入了一个显式的安全验证循环,在输出之前过滤生成的内容,确保符合预定义的规则。在公开数据集上的实验结果表明,该系统实现了73.2%的情感识别准确率,89.4%的响应模式一致性和100%的安全合规性,同时保持了低于100 ms的推理延迟,适合在设备上部署。模块化架构实现了可解释性和可扩展性,使其适用于儿童附近的媒体,治疗应用程序和情感响应智能设备。
摘要:This paper proposes a multi-agent artificial intelligence system that generates response-oriented media content in real time based on audio-derived emotional signals. Unlike conventional speech emotion recognition studies that focus primarily on classification accuracy, our approach emphasizes the transformation of inferred emotional states into safe, age-appropriate, and controllable response content through a structured pipeline of specialized AI agents. The proposed system comprises four cooperative agents: (1) an Emotion Recognition Agent with CNN-based acoustic feature extraction, (2) a Response Policy Decision Agent for mapping emotions to response modes, (3) a Content Parameter Generation Agent for producing media control parameters, and (4) a Safety Verification Agent enforcing age-appropriateness and stimulation constraints. We introduce an explicit safety verification loop that filters generated content before output, ensuring compliance with predefined rules. Experimental results on public datasets demonstrate that the system achieves 73.2% emotion recognition accuracy, 89.4% response mode consistency, and 100% safety compliance while maintaining sub-100ms inference latency suitable for on-device deployment. The modular architecture enables interpretability and extensibility, making it applicable to child-adjacent media, therapeutic applications, and emotionally responsive smart devices.
【13】LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
标题:LongSpeech:长言语转录、翻译和理解的可扩展基准
链接:https://arxiv.org/abs/2601.13539
备注:ICASSP 2026
摘要:音频语言模型的最新进展在短的、段级的语音任务上取得了显着的成功。然而,现实世界的应用,如会议转录,口语文档理解和会话分析,需要强大的模型,能够处理和推理长格式的音频。在这项工作中,我们提出了LongSpeech,一个大规模和可扩展的基准,专门用于评估和提高语音模型对长时间音频的能力。LongSpeech包含超过100,000个语音片段,每个片段大约10分钟长,具有丰富的ASR注释,语音翻译,摘要,语言检测,说话人计数,内容分离和问答。我们引入了一个可复制的管道,用于从不同的来源构建长格式的语音基准,使未来的扩展。我们对最先进的模型进行的初步实验揭示了显著的性能差距,模型通常专注于一项任务,而牺牲其他任务,并在更高层次的推理中苦苦挣扎。这些发现强调了我们的基准的挑战性。我们的基准将公开提供给研究界。
摘要:Recent advances in audio-language models have demonstrated remarkable success on short, segment-level speech tasks. However, real-world applications such as meeting transcription, spoken document understanding, and conversational analysis require robust models capable of processing and reasoning over long-form audio. In this work, we present LongSpeech, a large-scale and scalable benchmark specifically designed to evaluate and advance the capabilities of speech models on long-duration audio. LongSpeech comprises over 100,000 speech segments, each approximately 10 minutes long, with rich annotations for ASR, speech translation, summarization, language detection, speaker counting, content separation, and question answering. We introduce a reproducible pipeline for constructing long-form speech benchmarks from diverse sources, enabling future extensions. Our initial experiments with state-of-the-art models reveal significant performance gaps, with models often specializing in one task at the expense of others and struggling with higher-level reasoning. These findings underscore the challenging nature of our benchmark. Our benchmark will be made publicly available to the research community.
【14】Event Classification by Physics-informed Inpainting for Distributed Multichannel Acoustic Sensor with Partially Degraded Channels
标题:通过物理信息修复对具有部分退化通道的分布式多通道声传感器进行事件分类
链接:https://arxiv.org/abs/2601.13513
备注:Accepted to ICASSP 2026
摘要:分布式多通道声学感测(DMAS)支持大规模声音事件分类(SEC),但当许多通道降级以及测试时的传感器布局与训练布局不同时,性能会下降。我们提出了一个学习自由,物理知情的基于逆时迁移(RTM)的修复前端。在这种方法中,观察到的多通道频谱图首先使用解析格林函数在3D网格上反向传播以形成场景一致的图像,然后在对数梅尔特征提取和基于变换器的分类之前进行前向投影以重建修复的信号。我们评估的方法ESC-50与50个传感器和三种布局(圆形,线性,直角),其中每通道信噪比从-30至0分贝采样。与AST基线、缩放稀疏最大信道选择和信道交换增强相比,所提出的RTM前端在所有布局中实现了最佳或有竞争力的准确性,在直角布局上将准确性提高了13.1个点(从9.7%提高到22.8%)。相关性分析表明,空间权重对齐更强烈的SNR比通道源距离,更高的SNR权重相关性对应于更高的SEC精度。这些结果表明,一个重建,然后项目,基于物理的预处理有效地补充了DMAS布局开放配置和严重的信道退化下的学习方法。
摘要:Distributed multichannel acoustic sensing (DMAS) enables large-scale sound event classification (SEC), but performance drops when many channels are degraded and when sensor layouts at test time differ from training layouts. We propose a learning-free, physics-informed inpainting frontend based on reverse time migration (RTM). In this approach, observed multichannel spectrograms are first back-propagated on a 3D grid using an analytic Green's function to form a scene-consistent image, and then forward-projected to reconstruct inpainted signals before log-mel feature extraction and Transformer-based classification. We evaluate the method on ESC-50 with 50 sensors and three layouts (circular, linear, right-angle), where per-channel SNRs are sampled from -30 to 0 dB. Compared with an AST baseline, scaling-sparsemax channel selection, and channel-swap augmentation, the proposed RTM frontend achieves the best or competitive accuracy across all layouts, improving accuracy by 13.1 points on the right-angle layout (from 9.7% to 22.8%). Correlation analyses show that spatial weights align more strongly with SNR than with channel--source distance, and that higher SNR--weight correlation corresponds to higher SEC accuracy. These results demonstrate that a reconstruct-then-project, physics-based preprocessing effectively complements learning-only methods for DMAS under layout-open configurations and severe channel degradation.
【15】Context and Transcripts Improve Detection of Deepfake Audios of Public Figures
标题:上下文和文字记录改进了对公众人物Deepfake音频的检测
链接:https://arxiv.org/abs/2601.13464
摘要:人类使用上下文来评估信息的真实性。然而,目前的音频deepfake检测器只分析音频文件,而不考虑上下文或转录。我们创建并分析了一个由记者提供的Deepfake数据集(JDD),其中包含255个公开的deepfake,自2024年初以来,主要由70多名记者贡献。我们还生成了已故公众人物的合成音频数据集(SYN),并提出了一种新的基于上下文的音频深度伪造检测器(CADD)架构。此外,我们评估了两个大规模数据集的性能:ITW和P$^2$V。我们表明,足够的上下文和/或转录可以显着提高音频deepfake检测器的效率。多个基线音频deepfake检测器和传统分类器的性能(通过F1评分,AUC和EER测量)可以在F1评分中提高5%-37.58%,AUC提高3.77%-42.79%,EER提高6.17%-47.83%。我们还表明,CADD通过使用上下文和/或成绩单,对5种对抗性逃避策略更鲁棒,在所有实验中,将性能下降限制在平均仅为-0.71%。代码、模型和数据集可在我们的项目页面上获得:https://sites.northwestern.edu/nsail/cadd-context-based-audio-deepfake-detection(审查期间访问受限)。
摘要:Humans use context to assess the veracity of information. However, current audio deepfake detectors only analyze the audio file without considering either context or transcripts. We create and analyze a Journalist-provided Deepfake Dataset (JDD) of 255 public deepfakes which were primarily contributed by over 70 journalists since early 2024. We also generate a synthetic audio dataset (SYN) of dead public figures and propose a novel Context-based Audio Deepfake Detector (CADD) architecture. In addition, we evaluate performance on two large-scale datasets: ITW and P$^2$V. We show that sufficient context and/or the transcript can significantly improve the efficacy of audio deepfake detectors. Performance (measured via F1 score, AUC, and EER) of multiple baseline audio deepfake detectors and traditional classifiers can be improved by 5%-37.58% in F1-score, 3.77%-42.79% in AUC, and 6.17%-47.83% in EER. We additionally show that CADD, via its use of context and/or transcripts, is more robust to 5 adversarial evasion strategies, limiting performance degradation to an average of just -0.71% across all experiments. Code, models, and datasets are available at our project page: https://sites.northwestern.edu/nsail/cadd-context-based-audio-deepfake-detection (access restricted during review).
【16】The Achilles' Heel of Angular Margins: A Chebyshev Polynomial Fix for Speaker Verification
标题:角边缘的阿喀琉斯之踵:发言人验证的切比雪夫多元修复
链接:https://arxiv.org/abs/2601.13198
备注:Accepted for presentation at ICASSP 2026
摘要:角裕度损失,如AAM-Softmax,已成为说话人和人脸验证中的事实。它们的成功取决于直接操纵特性和类原型之间的角度。然而,这种操作依赖于arccos函数来恢复角度,引入了一个重要但被忽视的训练不稳定性来源。arccos的导数在其边界处爆炸,导致优化期间的梯度峰值。此外,该配方未能产生足够尖锐的梯度难以分类的例子。我们解决这些问题,提出ChebyAAM,损失,取代arccos操作与Chebyshev多项式近似。这种替换消除了梯度爆炸,并将更强的校正信号应用于困难的示例,从而实现更有效的优化。在三个基准测试(VoxCeleb,SITW和CN-Celeb)上的实验表明,我们的方法解决了不稳定性,并不断提高性能。我们的工作表明,近似的角度操作,而不是明确地计算它们,为设计未来的度量学习损失提供了一个更强大的路径。代码可在https://github.com/ExtraOrdinaryLab/vibe上获得。
摘要:Angular margin losses, such as AAM-Softmax, have become the de facto in speaker and face verification. Their success hinges on directly manipulating the angle between features and class prototypes. However, this manipulation relies on the arccos function to recover the angle, introducing a significant yet overlooked source of training instability. The derivative of arccos explodes at its boundaries, causing gradient peaks during optimisation. Furthermore, the formulation fails to generate a sufficiently sharp gradient for hard-to-classify examples. We address these issues by proposing ChebyAAM, a loss that replaces the arccos operation with its Chebyshev polynomial approximation. This substitution eliminates gradient explosion and applies a stronger corrective signal to hard examples, leading to more effective optimisation. Experiments on three benchmarks (VoxCeleb, SITW, and CN-Celeb) demonstrate that our method resolves the instability and consistently improves performance. Our work suggests that approximating angular operations, rather than calculating them explicitly, offers a more robust path for designing future metric learning losses. Code is available at https://github.com/ExtraOrdinaryLab/vibe.
【17】Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
标题:适合任何具有可控风格嵌入的声音的Lombard语音合成
链接:https://arxiv.org/abs/2601.12966
摘要:伦巴第效应在自然交流中起着关键作用,特别是在嘈杂的环境中或与听力受损的听众交谈时。我们提出了一个可控的文本到语音(TTS)系统,能够合成伦巴第语音的任何扬声器,而不需要明确的伦巴第数据在训练过程中。我们的方法利用从一个大型的,韵律多样的数据集学习的风格嵌入,并使用主成分分析(PCA)分析它们与伦巴底属性的相关性。通过移动相关的PCA组件,我们操纵的风格嵌入,并将它们纳入我们的TTS模型,以生成所需的隆巴德水平的语音。评估表明,我们的方法保留了自然性和扬声器的身份,提高了噪声下的可懂度,并提供了细粒度的韵律控制,提供了一个强大的解决方案,为任何扬声器的可控伦巴第文语转换。
摘要:The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-speech (TTS) system capable of synthesizing Lombard speech for any speaker without requiring explicit Lombard data during training. Our approach leverages style embeddings learned from a large, prosodically diverse dataset and analyzes their correlation with Lombard attributes using principal component analysis (PCA). By shifting the relevant PCA components, we manipulate the style embeddings and incorporate them into our TTS model to generate speech at desired Lombard levels. Evaluations demonstrate that our method preserves naturalness and speaker identity, enhances intelligibility under noise, and provides fine-grained control over prosody, offering a robust solution for controllable Lombard TTS for any speaker.
【18】Supervised Learning for Game Music Segmentation
标题:游戏音乐分段的监督学习
链接:https://arxiv.org/abs/2601.12961
摘要:目前,基于神经网络的模型,包括Transformers,由于缺乏对音乐结构的理解,很难从统一和重复的音乐材料中生成令人难忘和易于理解的音乐。因此,这些模型很少被游戏行业采用。许多学者假设,音乐结构的建模可以在更高的层次上通知模型,从而提高音乐生成的质量。本研究的目的是探索监督学习方法在结构分割任务中的性能,这是音乐结构建模的第一步。创建了一个具有309个结构注释的音频游戏音乐数据集来训练所提出的方法,该方法结合了卷积神经网络和递归神经网络,以更少的训练资源实现了与最先进的无监督学习方法相当的性能。
摘要:At present, neural network-based models, including transformers, struggle to generate memorable and readily comprehensible music from unified and repetitive musical material due to a lack of understanding of musical structure. Consequently, these models are rarely employed by the games industry. It is hypothesised by many scholars that the modelling of musical structure may inform models at a higher level, thereby enhancing the quality of music generation. The aim of this study is to explore the performance of supervised learning methods in the task of structural segmentation, which is the initial step in music structure modelling. An audio game music dataset with 309 structural annotations was created to train the proposed method, which combines convolutional neural networks and recurrent neural networks, achieving performance comparable to the state-of-the-art unsupervised learning methods with fewer training resources.
【19】UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
标题:XX:解开高度相关的歌唱声音混合物
链接:https://arxiv.org/abs/2601.12802
备注:Accepted by ICASSP 2026
摘要:我们介绍了一个新的框架,多个歌声分离(MSVS)的MXXX。虽然与语音分离相关,但MSVS面临着独特的挑战:数据稀缺和歌声混合的高度相关性。为了解决这些问题,我们提出了具有三个关键组成部分的混合策略:(1)音乐信息混合策略,以构建高度相关的,音乐般的混合物,(2)交叉源注意,通过反向注意驱动两个歌手的表示,以及(3)幅度惩罚损失惩罚错误分配的干扰能量。RISTOXX不仅通过模拟真实的训练数据来解决数据稀缺问题,而且还擅长通过架构和损失级别的跨源交互来分离高度相关的混合物。我们广泛的实验表明,SDRXX大大提高了性能,SDRi增益超过2.2 dB,比以前的工作。
摘要:We introduce UNMIXX, a novel framework for multiple singing voices separation (MSVS). While related to speech separation, MSVS faces unique challenges: data scarcity and the highly correlated nature of singing voices mixture. To address these issues, we propose UNMIXX with three key components: (1) musically informed mixing strategy to construct highly correlated, music-like mixtures, (2) cross-source attention that drives representations of two singers apart via reverse attention, and (3) magnitude penalty loss penalizing erroneously assigned interfering energy. UNMIXX not only addresses data scarcity by simulating realistic training data, but also excels at separating highly correlated mixtures through cross-source interactions at both the architectural and loss levels. Our extensive experiments demonstrate that UNMIXX greatly enhances performance, with SDRi gains exceeding 2.2 dB over prior work.
【20】SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization
标题:SoundPlot:一个用于鸟鸣声学分析和神经合成的开源框架,具有交互式3D可视化
链接:https://arxiv.org/abs/2601.12752
摘要:我们提出了SoundPlot,一个开源的框架,通过声学特征提取,降维和神经音频合成分析鸟类发声。该系统将音频信号转换为多维声学特征空间,使用基于Web的交互式图形实现3D时间动态的实时可视化。我们的框架实现了一个完整的分析合成管道,提取频谱特征(质心,带宽,对比度),通过概率YIN(pYIN)和梅尔频率倒谱系数(MFCC)的音高轮廓,将它们映射到一个统一的音色空间进行可视化。音频重建采用Griffin-Lim相位估计算法应用于梅尔频谱图。附带的基于Three.js的界面提供了双视口可视化,可以比较原始和合成的音频轨迹以及独立的播放控件。我们展示了该框架的能力,通过全面的波形分析,频谱图比较,并使用主成分分析(PCA)的特征空间评估。定量评估显示梅尔频谱图相关分数超过0.92,表示高保真度的感知声学结构的保存。SoundPlot在MIT许可下发布,以促进生物声学,音频信号处理和计算行为学的研究。
摘要:We present SoundPlot, an open-source framework for analyzing avian vocalizations through acoustic feature extraction, dimensionality reduction, and neural audio synthesis. The system transforms audio signals into a multi-dimensional acoustic feature space, enabling real-time visualization of temporal dynamics in 3D using web-based interactive graphics. Our framework implements a complete analysis-synthesis pipeline that extracts spectral features (centroid, bandwidth, contrast), pitch contours via probabilistic YIN (pYIN), and mel-frequency cepstral coefficients (MFCCs), mapping them to a unified timbre space for visualization. Audio reconstruction employs the Griffin-Lim phase estimation algorithm applied to mel spectrograms. The accompanying Three.js-based interface provides dual-viewport visualization comparing original and synthesized audio trajectories with independent playback controls. We demonstrate the framework's capabilities through comprehensive waveform analysis, spectrogram comparisons, and feature space evaluation using Principal Component Analysis (PCA). Quantitative evaluation shows mel spectrogram correlation scores exceeding 0.92, indicating high-fidelity preservation of perceptual acoustic structure. SoundPlot is released under the MIT License to facilitate research in bioacoustics, audio signal processing, and computational ethology.
【21】Toward Faithful Explanations in Acoustic Anomaly Detection
标题:声学异常检测中的忠实解释
链接:https://arxiv.org/abs/2601.12660
备注:Accepted at the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026. Code: https://github.com/Maab-Nimir/Faithful-Explanations-in-Acoustic-Anomaly-Detection
摘要:可解释性是用户在现实世界中的异常检测应用程序的信任至关重要。然而,深度学习模型尽管性能强大,但往往缺乏透明度。在这项工作中,我们研究了基于自动编码器的音频异常检测模型的可解释性,通过比较标准自动编码器(AE)与掩码自动编码器(MAE)的检测性能和可解释性。我们应用了几种归因方法,包括错误图,显着图,SmoothGrad,集成的一致性,GradSHAP和Grad-CAM。虽然MAE显示出略低的检测,但它始终提供更忠实和时间上更精确的解释,表明与真实异常更好地对齐。为了评估解释方法突出显示的区域的相关性,我们提出了一个基于扰动的忠实性度量,用它们的重建来代替它们以模拟正常输入。我们的研究结果基于真实工业场景中的实验,强调了将可解释性纳入异常检测管道的重要性,并表明掩蔽训练在不影响性能的情况下提高了解释质量。
摘要:Interpretability is essential for user trust in real-world anomaly detection applications. However, deep learning models, despite their strong performance, often lack transparency. In this work, we study the interpretability of autoencoder-based models for audio anomaly detection, by comparing a standard autoencoder (AE) with a mask autoencoder (MAE) in terms of detection performance and interpretability. We applied several attribution methods, including error maps, saliency maps, SmoothGrad, Integrated Gradients, GradSHAP, and Grad-CAM. Although MAE shows a slightly lower detection, it consistently provides more faithful and temporally precise explanations, suggesting a better alignment with true anomalies. To assess the relevance of the regions highlighted by the explanation method, we propose a perturbation-based faithfulness metric that replaces them with their reconstructions to simulate normal input. Our findings, based on experiments in a real industrial scenario, highlight the importance of incorporating interpretability into anomaly detection pipelines and show that masked training improves explanation quality without compromising performance.
【22】SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition
标题:SSVD-O:采用结构化MVD进行参数高效微调,用于语音识别
链接:https://arxiv.org/abs/2601.12600
备注:Accepted by IEEE ICASSP 2026
摘要:参数有效的微调(PEFT)是一种可扩展的方法,使大型语音基础模型适应新的领域。虽然LoRA及其最先进的变体等方法降低了自适应成本,但它们通常在模型子空间中均匀分配参数,这限制了它们在语音应用中的效率和可扩展性。在我们先前工作的基础上,本文介绍了结构SVD引导(SSVD)微调方法的扩展SSVD-Outer(SSVD-O)。SSVD-O将输入声学特征空间关联的内部变换与输出语义特征空间关联的外部变换相结合,以实现可扩展和平衡的自适应。我们进行了第一次系统的分析,参数预算分配跨模型子空间PEFT自动语音识别(ASR),并调查学习和遗忘之间的权衡资源受限。SSVD-O在ESPnet框架内的0.1B到2B的模型尺度上,针对LoRA、DoRA、PiSSA和SSVD进行了领域转移ASR任务的基准测试,包括儿童语音和区域口音。实验结果表明,SSVD-O始终缩小了性能差距,完全微调,同时提高泛化能力和减轻灾难性遗忘。
摘要:Parameter-efficient fine-tuning (PEFT) is a scalable approach for adapting large speech foundation models to new domains. While methods such as LoRA and its state-of-the-art variants reduce adaptation costs, they typically allocate parameters uniformly across model subspaces, which limits their efficiency and scalability in speech applications. Building on our prior work, this paper introduces SSVD-Outer (SSVD-O), an extension of the structured SVD-guided (SSVD) fine-tuning method. SSVD-O combines input acoustic feature space-associated inner transformations with output semantic feature space-associated outer transformations to enable scalable and balanced adaptation. We conduct the first systematic analysis of parameter budget allocation across model subspaces in PEFT for automatic speech recognition (ASR), and investigate the trade-off between learning and forgetting under constrained resources. SSVD-O is benchmarked against LoRA, DoRA, PiSSA, and SSVD on domain-shifted ASR tasks, including child speech and regional accents, across model scales from 0.1B to 2B within the ESPnet framework. Experimental results show that SSVD-O consistently narrows the performance gap to full fine-tuning while improving generalization and mitigating catastrophic forgetting.
【23】SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
链接:https://arxiv.org/abs/2601.12591
备注:5 pages, accepted by ICASSP 2026
摘要:人类情感的模糊性给机器学习模型带来了一些挑战,因为它们经常重叠,缺乏清晰的边界。对比语言-音频预训练(CLAP)已成为可泛化情感识别的关键技术。然而,由于传统的CLAP在成对的音频文本样本之间强制执行严格的一对一对齐,因此它忽略了模态内相似性,并将所有不匹配的对视为同样的否定。这与不同情感之间的模糊界限相冲突。为了解决这一限制,我们提出了SmoothCLAP,它引入了来自模态内相似性和非语言特征的软化目标。通过将这些软化的目标与传统的对比监督相结合,SmoothCLAP学习尊重分级情感关系的嵌入,同时保留与CLAP相同的推理管道。跨英语和德语的八个情感计算任务的实验表明,SmoothCLAP始终实现卓越的性能。我们的研究结果强调,利用软监督是一个很有前途的策略,建立情感感知的音频文本模型。
摘要:The ambiguity of human emotions poses several challenges for machine learning models, as they often overlap and lack clear delineating boundaries. Contrastive language-audio pretraining (CLAP) has emerged as a key technique for generalisable emotion recognition. However, as conventional CLAP enforces a strict one-to-one alignment between paired audio-text samples, it overlooks intra-modal similarity and treats all non-matching pairs as equally negative. This conflicts with the fuzzy boundaries between different emotions. To address this limitation, we propose SmoothCLAP, which introduces softened targets derived from intra-modal similarity and paralinguistic features. By combining these softened targets with conventional contrastive supervision, SmoothCLAP learns embeddings that respect graded emotional relationships, while retaining the same inference pipeline as CLAP. Experiments on eight affective computing tasks across English and German demonstrate that SmoothCLAP is consistently achieving superior performance. Our results highlight that leveraging soft supervision is a promising strategy for building emotion-aware audio-text models.
【24】Harmonizing the Arabic Audio Space with Data Scheduling
标题:协调阿拉伯语音频空间与数据调度
链接:https://arxiv.org/abs/2601.12494
备注:Foundation Models, Large Language Models, Native, Speech Models, Arabic
摘要:音频大语言模型(LLM)可以实现统一的语音理解和生成,但它们对语言复杂、方言丰富的环境的适应性仍然有待探索。本文提出了第一个系统的研究多任务指令调谐为阿拉伯语为中心的音频LLM,涵盖了层次结构的生成任务(ASR,语音摘要)和判别任务(方言和情感识别)。为了支持这项研究,我们介绍了AraMega-SSum,一种用于阿拉伯语语音摘要的新型数据集。我们对Qwen2.5-Omni(7 B)进行了微调,并提出了任务渐进式课程(TPC)以及基于对齐器的多样化采样(ADS),这是一种通过选择任务和标签平衡的示例来构建信息密集批次的策略。我们的研究结果揭示了一个关键的效率,鲁棒性权衡:虽然ADS加速了初始收敛并提高了语言学F1分数,但其固有的梯度波动性可能会在长时间训练下破坏生成解码。此外,虽然TPC稳定了核心声学映射,但它经常在下游任务中引起负迁移。我们证明了一个混合TPC+ADS策略提供了一个最佳的训练“配方”,首先建立一个强大的代表性基础,然后采用多样性意识的细化,以捕捉细粒度的细微差别。这些研究结果提供了实用的指导,在复杂的,低资源的多式联运环境中的全模型的有效适应。
摘要:Audio large language models (LLMs) enable unified speech understanding and generation, yet their adaptation to linguistically complex, dialect-rich settings remains underexplored. This paper presents the first systematic study of multi-task instruction tuning for an Arabic-centric audio LLM, covering a hierarchy of generative tasks (ASR, speech summarization) and discriminative tasks (dialect and emotion identification). To support this study, we introduce AraMega-SSum, a novel dataset for Arabic speech summarization. We fine-tune Qwen2.5-Omni (7B) and propose Task-Progressive Curriculum (TPC) along with Aligner-Based Diverse Sampling (ADS), a strategy that constructs information-dense batches by selecting task- and label-balanced examples. Our results reveal a critical efficiency, robustness trade-off: while ADS accelerates initial convergence and boosts paralinguistic F1-scores, its inherent gradient volatility can destabilize generative decoding under prolonged training. Furthermore, while the TPC stabilizes core acoustic mapping, it often induces negative transfer in downstream tasks. We demonstrate that a Hybrid TPC+ADS Strategy provides an optimal training ``recipe'', first establishing a robust representative foundation before employing diversity-aware refinement to capture fine-grained nuances. These findings offer practical guidance for the efficient adaptation of Omni-models in complex, low-resource multimodal environments.
【25】A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation
标题:用于选择性可编辑文本到语音生成的统一神经编解码语言模型
链接:https://arxiv.org/abs/2601.12480
摘要:神经编解码器语言模型通过完全模仿短语音提示的声学特征(包括音色、韵律和非语言信息)来实现令人印象深刻的zero-shot文语转换(TTS)。然而,这种整体模仿限制了他们隔离和控制个体属性的能力。在本文中,我们提出了一个统一的编解码器语言模型SpeechEdit扩展zero-shot TTS与选择性控制机制。默认情况下,SpeechEdit再现从语音提示推断的完整声学配置文件,但它有选择地仅覆盖由显式控制指令指定的属性。为了实现可控建模,SpeechEdit在我们新构建的LibriEdit数据集上进行训练,该数据集提供了从LibriHeavy派生的delta(差异感知)训练对。实验结果表明,我们的方法保持自然性和鲁棒性,同时提供灵活和本地化的控制所需的属性。音频样本可在https://speech-editing.github.io/speech-editing/上获得。
摘要:Neural codec language models achieve impressive zero-shot Text-to-Speech (TTS) by fully imitating the acoustic characteristics of a short speech prompt, including timbre, prosody, and paralinguistic information. However, such holistic imitation limits their ability to isolate and control individual attributes. In this paper, we present a unified codec language model SpeechEdit that extends zero-shot TTS with a selective control mechanism. By default, SpeechEdit reproduces the complete acoustic profile inferred from the speech prompt, but it selectively overrides only the attributes specified by explicit control instructions. To enable controllable modeling, SpeechEdit is trained on our newly constructed LibriEdit dataset, which provides delta (difference-aware) training pairs derived from LibriHeavy. Experimental results show that our approach maintains naturalness and robustness while offering flexible and localized control over desired attributes. Audio samples are available at https://speech-editing.github.io/speech-editing/.
【26】A Similarity Network for Correlating Musical Structure to Military Strategy
标题:将音乐结构与军事战略关联起来的相似网络
链接:https://arxiv.org/abs/2601.12314
备注:This paper was completed in 2024
摘要:音乐感知是以通感效应为基础的多感官过程,是音乐审美教育的重要组成部分。了解音乐结构有助于感知和审美教育。音乐结构包含了一系列信息,这些信息的协调形成了旋律,就像不同的军事行动合作产生军事战略一样。然而,从系统操作和信息管理的角度来评估音乐感知的方法很少。在本文中,我们探索音乐结构和军事战略之间的相似性,同时创建音乐片段相关网络(MCCN)的基础上梅尔频率倒谱系数(MFCC)。灵感来自于音乐会指挥家的乐谱和军事战争指挥官的沙盘练习之间的比较。具体来说,我们为各种战争电影配乐创建MCCN,然后将军事战术(孙子兵法等)和政治机构到军事行动网络。我们的初步研究结果表明,一些相似之处,这意味着音乐感知和审美教育可以接近从军事战略和管理的角度,通过这种跨学科的研究。同样,通过网络分析,我们可以发现军事谋略艺术与音乐结构艺术之间的相似之处,从而有助于理解技术与艺术之间的关系。
摘要:Music perception, a multi-sensory process based on the synesthesia effect, is an essential component of music aesthetic education. Understanding music structure helps both perception and aesthetic education. Music structure incorporates a range of information, the coordination of which forms the melody, just as different military actions cooperate to produce a military strategy. However, there are a few ways for assessing music perception from the perspectives of system operation and information management. In this paper, we explore the similarities between music structure and military strategy while creating the Music Clips Correlation Network (MCCN) based on Mel-frequency Cepstral Coefficients (MFCCs). The inspiration comes from the comparison between a concert conductor's musical score and a military war commander's sand table exercise. Specifically, we create MCCNs for various kinds of war movie soundtracks, then relate military tactics (Sun Tzu's Art of War, etc.) and political institutions to military operations networks. Our primary findings suggest a few similarities, implying that music perception and aesthetic education can be approached from a military strategy and management perspective through this interdisciplinary research. Similarly, we can discover similarities between the art of military scheming and the art of musical structure based on network analysis in order to facilitate the understanding of the relationship between technology and art.
【27】ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
标题:ParaMETA:学习从言语中分离出副语言说话风格的表达
链接:https://arxiv.org/abs/2601.12289
备注:9 pages, 7 figures, Accepted to AAAI-26 (Main Technical Track)
摘要:针对不同类型的说话风格(诸如情感、年龄和性别)学习代表性嵌入对于识别任务(例如,认知计算和人机交互)和生成任务(例如,风格可控的语音生成)。在这项工作中,我们介绍了ParaMETA,一个统一的和灵活的框架,直接从语音学习和控制说话风格。与依赖于单任务模型或跨模态对齐的现有方法不同,ParaMETA通过将语音投影到每种风格的专用子空间中来学习分解的特定于任务的嵌入。这种设计减少了任务间干扰,减轻了负迁移,并允许单个模型处理多个语言任务,如情感,性别,年龄和语言分类。除了识别之外,ParaMETA还可以在文本到语音(TTS)生成模型中实现细粒度的风格控制。它支持语音和文本提示,并允许用户修改一种说话风格,同时保留其他风格。大量的实验表明,ParaMETA在分类准确性方面优于强基线,并生成更自然和更有表现力的语音,同时保持适合现实世界应用的轻量级和高效的模型。
摘要:Learning representative embeddings for different types of speaking styles, such as emotion, age, and gender, is critical for both recognition tasks (e.g., cognitive computing and human-computer interaction) and generative tasks (e.g., style-controllable speech generation). In this work, we introduce ParaMETA, a unified and flexible framework for learning and controlling speaking styles directly from speech. Unlike existing methods that rely on single-task models or cross-modal alignment, ParaMETA learns disentangled, task-specific embeddings by projecting speech into dedicated subspaces for each type of style. This design reduces inter-task interference, mitigates negative transfer, and allows a single model to handle multiple paralinguistic tasks such as emotion, gender, age, and language classification. Beyond recognition, ParaMETA enables fine-grained style control in Text-To-Speech (TTS) generative models. It supports both speech- and text-based prompting and allows users to modify one speaking styles while preserving others. Extensive experiments demonstrate that ParaMETA outperforms strong baselines in classification accuracy and generates more natural and expressive speech, while maintaining a lightweight and efficient model suitable for real-world applications.
【28】Confidence-based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens
标题:基于置信度的语音数据集处理过滤,并使用离散令牌进行生成语音增强
链接:https://arxiv.org/abs/2601.12254
备注:Accepted for ICASSP 2026
摘要:生成式语音增强(GSE)模型在从噪声输入中生成高质量的干净语音方面表现出很大的潜力,从而实现了将噪声文本到语音(TTS)数据集转化为高质量数据集等应用。然而,GSE模型容易产生幻觉错误,如音素遗漏和说话人不一致,传统的基于非侵入性语音质量度量的错误过滤往往无法检测到。为了解决这个问题,我们提出了一种非侵入性的方法来过滤幻觉错误的离散令牌为基础的GSE模型。我们的方法利用生成的令牌的对数概率作为置信度分数来检测潜在的错误。实验结果表明,置信度分数与一套侵入性SE度量密切相关,并且我们的方法有效地识别了传统过滤方法遗漏的幻觉错误。此外,我们证明了我们的方法的实际效用:用我们基于置信度的过滤来管理野生TTS数据集,提高了随后训练的TTS模型的性能。
摘要:Generative speech enhancement (GSE) models show great promise in producing high-quality clean speech from noisy inputs, enabling applications such as curating noisy text-to-speech (TTS) datasets into high-quality ones. However, GSE models are prone to hallucination errors, such as phoneme omissions and speaker inconsistency, which conventional error filtering based on non-intrusive speech quality metrics often fails to detect. To address this issue, we propose a non-intrusive method for filtering hallucination errors from discrete token-based GSE models. Our method leverages the log-probabilities of generated tokens as confidence scores to detect potential errors. Experimental results show that the confidence scores strongly correlate with a suite of intrusive SE metrics, and that our method effectively identifies hallucination errors missed by conventional filtering methods. Furthermore, we demonstrate the practical utility of our method: curating an in-the-wild TTS dataset with our confidence-based filtering improves the performance of subsequently trained TTS models.
【29】Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
标题:Sound 2 Hap:从人类评级中学习音频到振动触觉生成
链接:https://arxiv.org/abs/2601.12245
摘要:环境声音(如脚步声、键盘敲击声或狗吠声)携带着丰富的信息和情感背景,这使得它们对用户应用程序中的触觉设计很有价值。然而,现有的音频到振动方法依赖于针对音乐或游戏调整的信号处理规则,并且通常无法在不同的声音中推广。为了解决这个问题,我们首先研究了用户对四种现有音频到触觉算法的感知,然后创建了一个环境声音的数据驱动模型。在研究1中,34名参与者对四种算法产生的1,000种声音的振动进行了评级,没有发现一致的算法偏好。使用这个数据集,我们训练了基于CNN的自动编码器Sound 2 Hap,以低延迟从不同的声音中生成感知上有意义的振动。在研究2中,15名参与者在音频振动匹配和触觉体验指数(HXI)方面的输出高于信号处理基线,发现它与不同的声音更和谐。这项工作展示了一个感知验证的方法来音频触觉翻译,扩大了声音驱动的触觉的范围。
摘要:Environmental sounds like footsteps, keyboard typing, or dog barking carry rich information and emotional context, making them valuable for designing haptics in user applications. Existing audio-to-vibration methods, however, rely on signal-processing rules tuned for music or games and often fail to generalize across diverse sounds. To address this, we first investigated user perception of four existing audio-to-haptic algorithms, then created a data-driven model for environmental sounds. In Study 1, 34 participants rated vibrations generated by the four algorithms for 1,000 sounds, revealing no consistent algorithm preferences. Using this dataset, we trained Sound2Hap, a CNN-based autoencoder, to generate perceptually meaningful vibrations from diverse sounds with low latency. In Study 2, 15 participants rated its output higher than signal-processing baselines on both audio-vibration match and Haptic Experience Index (HXI), finding it more harmonious with diverse sounds. This work demonstrates a perceptually validated approach to audio-haptic translation, broadening the reach of sound-driven haptics.
【30】Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
标题:多干注意力和层次不确定性建模下的歌曲美学评价
链接:https://arxiv.org/abs/2601.12222
摘要:音乐生成人工智能(AI)正在迅速扩展音乐内容,需要自动化的歌曲美学评估。然而,现有的研究主要集中在语音,音频或演唱质量,留下歌曲美学探索不足。此外,传统的方法往往预测一个精确的平均意见得分(MOS)值,这很难捕捉人类感知的细微差别,在歌曲美学评价。本文提出了一个面向歌曲的美学评价框架,包括两个新的模块:1)多干注意力融合(MSAF)在混音-人声和混音-伴奏对之间建立双向交叉注意,融合它们以捕捉复杂的音乐特征; 2)分级粒度感知区间聚合(HiGIA)学习多粒度分数概率分布,将它们聚合到分数区间中,并在该区间内应用回归以产生最终分数。我们对两个全长歌曲数据集进行了评估:SongEval数据集(AI生成)和内部美学数据集(人类创建),并与两个最先进的(SOTA)模型进行了比较。实验结果表明,该方法对歌曲的多维美学评价具有较强的性能。
摘要:Music generative artificial intelligence (AI) is rapidly expanding music content, necessitating automated song aesthetics evaluation. However, existing studies largely focus on speech, audio or singing quality, leaving song aesthetics underexplored. Moreover, conventional approaches often predict a precise Mean Opinion Score (MOS) value directly, which struggles to capture the nuances of human perception in song aesthetics evaluation. This paper proposes a song-oriented aesthetics evaluation framework, featuring two novel modules: 1) Multi-Stem Attention Fusion (MSAF) builds bidirectional cross-attention between mixture-vocal and mixture-accompaniment pairs, fusing them to capture complex musical features; 2) Hierarchical Granularity-Aware Interval Aggregation (HiGIA) learns multi-granularity score probability distributions, aggregates them into a score interval, and applies a regression within the interval to produce the final score. We evaluated on two datasets of full-length songs: SongEval dataset (AI-generated) and an internal aesthetics dataset (human-created), and compared with two state-of-the-art (SOTA) models. Results show that the proposed method achieves stronger performance for multi-dimensional song aesthetics evaluation.
【31】Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
标题:神经编解码器可以通用吗?一项针对不可见语言和非言语任务的对照研究
链接:https://arxiv.org/abs/2601.12205
摘要:本文研究了神经音频编解码器(NAC)泛化能力的三个关键但未充分探索的方面:(i)NACs是否可以在预训练期间推广到看不见的语言,(ii)仅语音预训练的NACs是否可以有效地推广到非语音应用,例如环境声音、音乐和动物发声,以及(iii)在预训练期间结合非语音数据是否可以提高语音和非语音任务两者的性能。现有的研究通常依赖于现成的NAC进行比较,由于实施的差异,这限制了洞察力。在这项工作中,我们使用严格控制的配置和精心策划的预训练数据从头开始训练NAC,以实现公平的比较。我们使用11个指标对NAC在信号重建质量和下游应用方面的性能进行了全面评估。我们的研究结果表明,NAC可以在预训练期间推广到看不见的语言,仅语音预训练的NAC在非语音任务上表现出性能下降,并且在预训练期间合并非语音数据可以提高非语音任务的性能,同时保持语音任务的性能相当。
摘要:This paper investigates three crucial yet underexplored aspects of the generalization capabilities of neural audio codecs (NACs): (i) whether NACs can generalize to unseen languages during pre-training, (ii) whether speech-only pre-trained NACs can effectively generalize to non-speech applications such as environmental sounds, music, and animal vocalizations, and (iii) whether incorporating non-speech data during pre-training can improve performance on both speech and non-speech tasks. Existing studies typically rely on off-the-shelf NACs for comparison, which limits insight due to variations in implementation. In this work, we train NACs from scratch using strictly controlled configurations and carefully curated pre-training data to enable fair comparisons. We conduct a comprehensive evaluation of NAC performance on both signal reconstruction quality and downstream applications using 11 metrics. Our results show that NACs can generalize to unseen languages during pre-training, speech-only pre-trained NACs exhibit degraded performance on non-speech tasks, and incorporating non-speech data during pre-training improves performance on non-speech tasks while maintaining comparable performance on speech tasks.
【32】Embryonic Exposure to VPA Influences Chick Vocalisations: A Computational Study
标题:胚胎暴露于VPA影响小鸡发声:计算研究
链接:https://arxiv.org/abs/2601.12203
备注:Main text (approx. 23 pages including references) with extensive Supplementary Material ( 20 pages) and multiple figures
摘要:在幼雏(Gallus gallus)等动物中,发声传达了有关情感和行为状态的信息。传统的发声分析方法依赖于手动注释和预定义的类别,引入了偏见,限制了可扩展性,并且无法捕获声乐曲目的全部复杂性。我们介绍了一个计算框架的自动检测,声学特征提取,和鸡发声的无监督学习。将此框架应用于新孵化的小鸡的数据集,我们确定了两个主要的声乐集群。然后,我们在胚胎发育期间暴露于媒介物或丙戊酸(VPA)的小鸡的独立数据集上测试了我们的计算框架,丙戊酸是一种破坏神经发育并与自闭症样症状有关的化合物。实验数据集的聚类分析证实了两个主要的声乐集群,并揭示了系统组之间的差异。暴露于VPA的小鸡表现出改变的剧目,与软呼叫的相对增加。VPA差异影响呼叫集群,调制时间,频率和能量域功能。总体而言,VPA暴露的小鸡产生发声持续时间较短,音高变异性降低,并修改能源配置文件,观察到最强的变化,在更响亮的电话。这项研究为分析动物发声提供了一个计算框架,推进了典型和非典型发声发展中的早期交流知识。
摘要:In young animals like poultry chicks (Gallus gallus), vocalisations convey information about affective and behavioural states. Traditional approaches to vocalisation analysis, relying on manual annotation and predefined categories, introduce biases, limit scalability, and fail to capture the full complexity of vocal repertoires. We introduce a computational framework for the automated detection, acoustic feature extraction, and unsupervised learning of chick vocalisations. Applying this framework to a dataset of newly hatched chicks, we identified two primary vocal clusters. We then tested our computational framework on an independent dataset of chicks exposed during embryonic development to vehicle or Valproic Acid (VPA), a compound that disrupts neural development and is linked to autistic-like symptoms. Clustering analysis on the experimental dataset confirmed two primary vocal clusters and revealed systematic differences between groups. VPA-exposed chicks showed an altered repertoire, with a relative increase in softer calls. VPA differentially affected call clusters, modulating temporal, frequency, and energy domain features. Overall, VPA-exposed chicks produced vocalisations with shorter duration, reduced pitch variability, and modified energy profiles, with the strongest alterations observed in louder calls. This study provides a computational framework for analysing animal vocalisations, advancing knowledge of early-life communication in typical and atypical vocal development.
【33】VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
标题:VidButton:使用生成性音乐和上下文缩略图创建视频配乐
链接:https://arxiv.org/abs/2601.12180
备注:Accepted to CHI 2026
摘要:音乐塑造了视频的基调,但创作者往往很难找到与视频的情绪和叙事相匹配的配乐。最近的文本到音乐模型让创作者从文本提示生成音乐,但我们的形成性研究(N=8)显示创作者很难构建不同的提示,快速审查和比较曲目,并了解它们对视频的影响。我们提出了VidTune,一个系统,支持配乐创作,从创作者的提示生成不同的音乐选项,并产生快速审查的上下文缩略图。VidTune提取代表性的视频主题,以背景中的缩略图为基础,将每个轨道的效价和能量映射到颜色和亮度等视觉线索上,并描绘突出的流派和乐器。创作者可以通过自然语言编辑来优化曲目,VidTune将其扩展到新一代。在一项对照用户研究(N=12)和一项探索性案例研究(N=6)中,参与者发现VidTune有助于有效地审查和比较音乐选项,并将此过程描述为有趣和丰富。
摘要:Music shapes the tone of videos, yet creators often struggle to find soundtracks that match their video's mood and narrative. Recent text-to-music models let creators generate music from text prompts, but our formative study (N=8) shows creators struggle to construct diverse prompts, quickly review and compare tracks, and understand their impact on the video. We present VidTune, a system that supports soundtrack creation by generating diverse music options from a creator's prompt and producing contextual thumbnails for rapid review. VidTune extracts representative video subjects to ground thumbnails in context, maps each track's valence and energy onto visual cues like color and brightness, and depicts prominent genres and instruments. Creators can refine tracks through natural language edits, which VidTune expands into new generations. In a controlled user study (N=12) and an exploratory case study (N=6), participants found VidTune helpful for efficiently reviewing and comparing music options and described the process as playful and enriching.
【34】Learning Audio-Visual Embeddings with Inferred Latent Interaction Graphs
标题:使用推断潜在交互图学习视听嵌入
链接:https://arxiv.org/abs/2601.11995
备注:16 pages, 5 figures, 2 tables
摘要:学习鲁棒的视听嵌入需要将真正相关的音频和视觉信号放在一起,同时过滤掉偶然的同现-背景噪音,不相关的元素或未注释的事件。大多数对比和三重丢失方法对每个片段使用稀疏注释标签,并将任何同现视为语义相似性。例如,标记为“训练”的视频可能还包含摩托车音频和视频,因为“摩托车”不是所选择的注释;标准方法将这些同现视为对其他地方的真正摩托车锚的否定,从而产生假否定并丢失真正的跨模态依赖性。我们提出了一个框架,利用软标签预测和推断的潜在交互来解决这些问题:(1)视听语义对齐损失(AV-SAL)训练教师网络,以产生跨模态的对齐软标签分布,为共同发生但未注释的事件分配非零概率,并丰富监督信号。(2)推断潜在交互图(ILI)将GRaSP算法应用于教师软标签,以推断类之间的稀疏有向依赖图。该图突出了方向依赖性(例如,“训练(视觉)”->“摩托车(音频)”),其暴露类之间可能的语义或条件关系;这些被解释为估计的依赖性模式。(3)潜在交互正则化器(LIR):学生网络使用度量损失和ILI图指导的正则化器进行训练,将依赖性链接但未标记的对的嵌入与其软标签概率成比例。在AVE和VEGAS基准测试上的实验表明,平均精度(mAP)得到了一致的提高,这表明将推断的潜在交互集成到嵌入学习中可以增强鲁棒性和语义一致性。
摘要:Learning robust audio-visual embeddings requires bringing genuinely related audio and visual signals together while filtering out incidental co-occurrences - background noise, unrelated elements, or unannotated events. Most contrastive and triplet-loss methods use sparse annotated labels per clip and treat any co-occurrence as semantic similarity. For example, a video labeled "train" might also contain motorcycle audio and visual, because "motorcycle" is not the chosen annotation; standard methods treat these co-occurrences as negatives to true motorcycle anchors elsewhere, creating false negatives and missing true cross-modal dependencies. We propose a framework that leverages soft-label predictions and inferred latent interactions to address these issues: (1) Audio-Visual Semantic Alignment Loss (AV-SAL) trains a teacher network to produce aligned soft-label distributions across modalities, assigning nonzero probability to co-occurring but unannotated events and enriching the supervision signal. (2) Inferred Latent Interaction Graph (ILI) applies the GRaSP algorithm to teacher soft labels to infer a sparse, directed dependency graph among classes. This graph highlights directional dependencies (e.g., "Train (visual)" -> "Motorcycle (audio)") that expose likely semantic or conditional relationships between classes; these are interpreted as estimated dependency patterns. (3) Latent Interaction Regularizer (LIR): A student network is trained with both metric loss and a regularizer guided by the ILI graph, pulling together embeddings of dependency-linked but unlabeled pairs in proportion to their soft-label probabilities. Experiments on AVE and VEGAS benchmarks show consistent improvements in mean average precision (mAP), demonstrating that integrating inferred latent interactions into embedding learning enhances robustness and semantic coherence.
【35】MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
标题:MuseAgent-1:音乐配乐和表演音频的交互式接地多模式理解
链接:https://arxiv.org/abs/2601.11968
备注:Tech Report
摘要:尽管最近在多模态大型语言模型(MLLM)方面取得了进展,但它们理解音乐和与音乐交互的能力仍然有限。音乐理解需要对符号分数和表现力表现音频进行接地推理,由于感知接地不足,通用MLLM通常无法处理。我们介绍MuseAgent,一个以音乐为中心的多模态代理,增强语言模型与结构化的符号表示来自乐谱图像和性能音频。通过集成光学音乐识别和自动音乐转录模块,MuseAgent可以对细粒度的音乐内容进行多步推理和交互。为了系统地评估音乐理解能力,我们进一步提出了MuseBench,这是一个涵盖音乐理论推理,乐谱解释和跨文本,图像和音频形式的性能级别分析的基准。实验表明,现有的MLLM在这些任务上表现不佳,而MuseAgent实现了实质性的改进,突出了交互式音乐理解的结构化多模态接地的重要性。
摘要:Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reasoning over symbolic scores and expressive performance audio, which general-purpose MLLMs often fail to handle due to insufficient perceptual grounding. We introduce MuseAgent, a music-centric multimodal agent that augments language models with structured symbolic representations derived from sheet music images and performance audio. By integrating optical music recognition and automatic music transcription modules, MuseAgent enables multi-step reasoning and interaction over fine-grained musical content. To systematically evaluate music understanding capabilities, we further propose MuseBench, a benchmark covering music theory reasoning, score interpretation, and performance-level analysis across text, image, and audio modalities. Experiments show that existing MLLMs perform poorly on these tasks, while MuseAgent achieves substantial improvements, highlighting the importance of structured multimodal grounding for interactive music understanding.
【36】The Third VoicePrivacy Challenge: Preserving Emotional Expressiveness and Linguistic Content in Voice Anonymization
标题:第三个语音隐私挑战:在语音分析中保留情感表达和语言内容
链接:https://arxiv.org/abs/2601.11846
备注:under review
摘要:我们介绍了2024年举行的第三届语音隐私挑战赛的结果和分析,该挑战赛的重点是推进语音匿名化技术。这项任务是开发一个语音匿名系统的语音数据,隐藏说话者的语音身份,同时保留语言内容和情感状态。我们提供了挑战框架的系统概述,包括匿名化任务和用于系统开发和评估的数据集的详细描述。我们概述了攻击模型和评估隐私保护(隐藏扬声器的声音身份)和实用程序(内容和情绪状态保存)的客观评估指标。我们描述了六个基线匿名化系统,并总结了挑战参与者开发的创新方法。最后,我们提供了关键的见解和意见,以指导未来的VoicePrivacy挑战的设计,并确定有前途的语音匿名化研究方向。
摘要:We present results and analyses from the third VoicePrivacy Challenge held in 2024, which focuses on advancing voice anonymization technologies. The task was to develop a voice anonymization system for speech data that conceals a speaker's voice identity while preserving linguistic content and emotional state. We provide a systematic overview of the challenge framework, including detailed descriptions of the anonymization task and datasets used for both system development and evaluation. We outline the attack model and objective evaluation metrics for assessing privacy protection (concealing speaker voice identity) and utility (content and emotional state preservation). We describe six baseline anonymization systems and summarize the innovative approaches developed by challenge participants. Finally, we provide key insights and observations to guide the design of future VoicePrivacy challenges and identify promising directions for voice anonymization research.
【37】CSyMR: Benchmarking Compositional Symbolic Muisc Reasoning With MIR Tool Integration
标题:CSyMR:通过MIR工具集成对合成符号Muisc推理进行基准测试
链接:https://arxiv.org/abs/2601.11556
摘要:大型语言模型(LLM)被用于符号音乐推理,但现有的基准强调孤立的知识或原子分析,而不是连接音乐结构所需的综合成分推理。为了解决这个问题,我们提出了作曲符号音乐推理基准(CSyMR-Bench),这是一个来自专家论坛和专业考试的126个问题的精选多项选择数据集。每一个问题都需要结合几个原子分析来得出最终答案。此外,我们引入了一个工具增强的代理框架,利用符号音乐分析工具的music 21库,以解决CSyMR-Bench所带来的挑战。实验验证了CSyMR-Bench在社区来源和考试风格的问题上都提出了不小的挑战,而我们的工具增强代理始终优于所有基线,实现了5-7%的绝对准确率增益。
摘要:Large Language Models (LLMs) are leveraged in symbolic music reasoning, yet existing benchmarks emphasize isolated knowledge or atomic analyses rather than the integrative compositional reasoning needed to connect musical structures. To address this, we present the Compositional Symbolic Music Reasoning Benchmark (CSyMR-Bench), a curated multiple-choice dataset of 126 questions derived from expert forums and professional examinations. Each item involves combining several atomic analyses to arrive at the final answer. Furthermore, we introduce a tool-augmented agent framework that leverages symbolic music analysis tools from the music21 library to address the challenges posed by CSyMR-Bench. Experiments validate that CSyMR-Bench poses a non-trivial challenge across both community-sourced and exam-style questions, while our tool-augmented agent consistently outperforms all baselines, achieving 5-7% absolute accuracy gains.
【38】ICASSP 2026 URGENT Speech Enhancement Challenge
标题:ICASP 2026紧急语音增强挑战赛
链接:https://arxiv.org/abs/2601.13531
备注:The overview paper of the ICASSP 2026 URGENT Speech Enhancement Challenge
摘要:ICASSP 2026紧急挑战赛通过关注处理不同失真、域和输入条件的通用语音增强(SE)系统来推进该系列。这篇综述论文详细介绍了挑战的动机、任务定义、数据集、基线系统、评估协议和结果。这项挑战分为两个相辅相成的轨道。Track 1侧重于通用语音增强,而Track 2介绍了增强语音的语音质量评估。该挑战赛吸引了80多个团队注册,其中29个提交了有效参赛作品,显示出社区对强大的SE技术的浓厚兴趣。
摘要:The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task definitions, datasets, baseline systems, evaluation protocols, and results. The challenge is divided into two complementary tracks. Track 1 focuses on universal speech enhancement, while Track 2 introduces speech quality assessment for enhanced speech. The challenge attracted over 80 team registrations, with 29 submitting valid entries, demonstrating significant community interest in robust SE technologies.
【39】Content Leakage in LibriSpeech and Its Impact on the Privacy Evaluation of Speaker Anonymization
标题:LibriSpeech中的内容泄露及其对说话者匿名化隐私评估的影响
链接:https://arxiv.org/abs/2601.13107
备注:Accepted to ICASSP 2026
摘要:说话人匿名的目的是隐藏说话人的身份,而不考虑语言内容。在这项研究中,我们揭示了Librispeech的一个弱点,该数据集通常用于评估匿名者:Librispeech说话者阅读的书籍是如此独特,以至于说话者可以通过他们的词汇来识别。即使是完美的匿名者也无法阻止这种身份泄露。EdAcc数据集在这方面更好:只有少数说话者可以通过他们的词汇表识别,鼓励攻击者在其他地方寻找匿名说话者的身份。EdAcc还包括自发的演讲和更多样化的演讲者,补充了Librispeech,并对匿名者的工作方式提供了更多的见解。
摘要:Speaker anonymization aims to conceal a speaker's identity, without considering the linguistic content. In this study, we reveal a weakness of Librispeech, the dataset that is commonly used to evaluate anonymizers: the books read by the Librispeech speakers are so distinct, that speakers can be identified by their vocabularies. Even perfect anonymizers cannot prevent this identity leakage. The EdAcc dataset is better in this regard: only a few speakers can be identified through their vocabularies, encouraging the attacker to look elsewhere for the identities of the anonymized speakers. EdAcc also comprises spontaneous speech and more diverse speakers, complementing Librispeech and giving more insights into how anonymizers work.
【40】Improving Audio Question Answering with Variational Inference
标题:利用变分推理改进音频问题回答
链接:https://arxiv.org/abs/2601.12700
备注:ICASSP 2026
摘要:变分推理(VI)提供了一个原则性的框架,用于估计模型参数的后验分布,从而在优化过程中对权重不确定性进行显式建模。通过捕捉这种不确定性,VI提高了预测的可靠性,产生更好的校准输出。在这项工作中,我们调查的好处,具有挑战性的多模态理解和推理,通过应用改进的变分在线牛顿(IVON),最近的VI优化器,微调多模态大型语言模型的音频问答任务。我们的研究结果表明,VI不仅提高了预测精度,而且还显着提高了校准,减少了模型的过度自信。这些进展进一步支持风险敏感的应用,如选择性预测,其中可靠的置信度估计至关重要。
摘要:Variational inference (VI) provides a principled framework for estimating posterior distributions over model parameters, enabling explicit modeling of weight uncertainty during optimization. By capturing this uncertainty, VI improves the reliability of predictions, yielding better calibrated outputs. In this work, we investigate the benefits of VI for challenging multimodal understanding and reasoning by applying the Improved Variational Online Newton (IVON), a recent VI optimizer, to fine-tuning a multimodal large language model on audio question answering tasks. Our results show that VI not only enhances predictive accuracy but also significantly improves calibration, reducing the model's overconfidence. These advances further support risk-sensitive applications such as selective prediction, where reliable confidence estimates are crucial.
【41】SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
标题:Scrum:具有可变持续时间音频和多目标训练的可扩展音频预训练
链接:https://arxiv.org/abs/2601.12594
备注:Accepted to ICASSP 2026
摘要:对比语言-音频预训练(CLAP)在学习语义丰富的音频表示方面取得了显着的成功,并被广泛用于各种与音频相关的任务。然而,目前的CLAP模型面临着几个关键的限制。首先,它们通常在相对较小的数据集上训练,通常包括几百万个音频样本。其次,现有的CLAP模型被限制为短且固定的持续时间,这限制了它们在具有可变持续时间音频的真实世界场景中的使用。第三,标准的对比训练目标对全局表示进行操作,这可能会阻碍密集的细粒度音频特征的学习。为了应对这些挑战,我们引入了可扩展的音频预训练(Scalable Audio-Pretraining,简称SSTO),它将语言-音频预训练扩展到1.09亿个音频-文本对,具有可变的音频持续时间,并包含多个训练目标。在单阶段训练中,Sort将对比度损失与额外的自我监督和字幕损失统一起来,促进了更丰富的密集音频表示的学习。该模型在音频文本检索和zero-shot音频分类任务上实现了新的最先进的性能,在不同的基准测试中证明了其有效性。
摘要:Contrastive language-audio pretraining (CLAP) has achieved notable success in learning semantically rich audio representations and is widely adopted for various audio-related tasks. However, current CLAP models face several key limitations. First, they are typically trained on relatively small datasets, often comprising a few million audio samples. Second, existing CLAP models are restricted to short and fixed duration, which constrains their usage in real-world scenarios with variable-duration audio. Third, the standard contrastive training objective operates on global representations, which may hinder the learning of dense, fine-grained audio features. To address these challenges, we introduce Scalable Language-Audio Pretraining (SLAP), which scales language-audio pretraining to 109 million audio-text pairs with variable audio durations and incorporates multiple training objectives. SLAP unifies contrastive loss with additional self-supervised and captioning losses in a single-stage training, facilitating the learning of richer dense audio representations. The proposed SLAP model achieves new state-of-the-art performance on audio-text retrieval and zero-shot audio classification tasks, demonstrating its effectiveness across diverse benchmarks.
【42】Robust Online Overdetermined Independent Vector Analysis Based on Bilinear Decomposition
标题:基于双线性分解的鲁棒在线超定独立载体分析
链接:https://arxiv.org/abs/2601.12485
摘要:在线盲源分离对于语音通信和人机交互都是必不可少的。在现有的方法中,超定独立向量分析(OverIVA)通过利用源信号的统计独立性和源与噪声子空间之间的正交性提供了强大的性能。然而,当应用于大型麦克风阵列时,参数的数量迅速增长,这会降低在线估计精度。为了克服这一挑战,我们建议将每个长分离滤波器分解为两个较短滤波器的双线性形式,从而减少参数的数量。由于这两个滤波器是紧密耦合的,我们设计了交替迭代投影算法来依次更新它们。仿真结果表明,该方法在参数少得多的情况下,获得了较好的性能和鲁棒性.
摘要:Online blind source separation is essential for both speech communication and human-machine interaction. Among existing approaches, overdetermined independent vector analysis (OverIVA) delivers strong performance by exploiting the statistical independence of source signals and the orthogonality between source and noise subspaces. However, when applied to large microphone arrays, the number of parameters grows rapidly, which can degrade online estimation accuracy. To overcome this challenge, we propose decomposing each long separation filter into a bilinear form of two shorter filters, thereby reducing the number of parameters. Because the two filters are closely coupled, we design an alternating iterative projection algorithm to update them in turn. Simulation results show that, with far fewer parameters, the proposed method achieves improved performance and robustness.
【43】Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
标题:融合前的净化:迈向无屏蔽语音增强以实现稳健的视听语音识别
链接:https://arxiv.org/abs/2601.12436
备注:Accepted by ICASSP2026
摘要:视听语音识别(AVSR)通常通过将抗噪声的视觉线索与音频信号相结合来提高噪声环境中的识别精度。然而,高噪声音频输入容易将不利干扰引入特征融合过程。为了缓解这一问题,最近的AVSR方法通常采用基于掩码的策略来过滤特征交互和融合期间的音频噪声,但这样的方法有可能丢弃与噪声一起的语义相关信息。在这项工作中,我们提出了一个端到端的噪声鲁棒AVSR框架加上语音增强,消除了显式噪声掩模生成的需要。该框架利用基于Conformer的瓶颈融合模块隐式地在视频辅助下细化噪声音频特征。通过减少模态冗余和增强模态间的相互作用,我们的方法保留语音语义的完整性,以实现强大的识别性能。在公共LRS3基准上的实验评估表明,我们的方法在噪声条件下优于先前的先进的基于掩模的基线。
摘要:Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-noise audio inputs are prone to introducing adverse interference into the feature fusion process. To mitigate this, recent AVSR methods often adopt mask-based strategies to filter audio noise during feature interaction and fusion, yet such methods risk discarding semantically relevant information alongside noise. In this work, we propose an end-to-end noise-robust AVSR framework coupled with speech enhancement, eliminating the need for explicit noise mask generation. This framework leverages a Conformer-based bottleneck fusion module to implicitly refine noisy audio features with video assistance. By reducing modality redundancy and enhancing inter-modal interactions, our method preserves speech semantic integrity to achieve robust recognition performance. Experimental evaluations on the public LRS3 benchmark suggest that our method outperforms prior advanced mask-based baselines under noisy conditions.
【44】Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models
标题:基于条件扩散模型的骨导引导多模式语音增强
链接:https://arxiv.org/abs/2601.12354
备注:Accepted to IEEE ICASSP 2026
摘要:单通道语音增强模型在极端噪声环境中面临显著的性能下降。虽然先前的工作已经表明,互补的骨传导语音可以引导增强,但这种噪声免疫模式的有效整合仍然是一个挑战。本文介绍了一种新的多模态语音增强框架,集成骨导传感器与空气传导麦克风使用条件扩散模型。我们提出的模型显着优于先前建立的多模态技术和强大的扩散为基础的单一模态基线在广泛的声学条件。
摘要:Single-channel speech enhancement models face significant performance degradation in extremely noisy environments. While prior work has shown that complementary bone-conducted speech can guide enhancement, effective integration of this noise-immune modality remains a challenge. This paper introduces a novel multimodal speech enhancement framework that integrates bone-conduction sensors with air-conducted microphones using a conditional diffusion model. Our proposed model significantly outperforms previously established multimodal techniques and a powerful diffusion-based single-modal baseline across a wide range of acoustic conditions.
【45】Adaptive Rotary Steering with Joint Autoregression for Robust Extraction of Closely Moving Speakers in Dynamic Scenarios
标题:具有联合自回归的自适应旋转转向,用于动态场景中近距离移动的扬声器的鲁棒提取
链接:https://arxiv.org/abs/2601.12345
备注:Accepted at IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026
摘要:高保真度立体声的深度空间滤波的最新进展通过在多声道增强之前将声场朝向目标扬声器旋转,在固定多扬声器场景中展示了强大的性能。在动态声学条件下与移动扬声器的适用性,我们建议自动化使用交错跟踪算法的目标的初始方向为条件的旋转转向。然而,对于附近或交叉的说话者来说,稳健的跟踪变得困难,并且空间线索对于增强的效果较差。通过将处理后的录音作为额外的指导到这两种算法中,我们的新的联合自回归框架利用语音的时间-频谱相关性来解决空间上具有挑战性的扬声器星座。因此,我们提出的方法显着提高了对密集扬声器的跟踪和增强,在合成数据集上始终优于可比的非自回归方法。真实世界的录音补充了这些发现在复杂的情况下,多个扬声器交叉和不同的扬声器到阵列的距离。
摘要:Latest advances in deep spatial filtering for Ambisonics demonstrate strong performance in stationary multi-speaker scenarios by rotating the sound field toward a target speaker prior to multi-channel enhancement. For applicability in dynamic acoustic conditions with moving speakers, we propose to automate this rotary steering using an interleaved tracking algorithm conditioned on the target's initial direction. However, for nearby or crossing speakers, robust tracking becomes difficult and spatial cues less effective for enhancement. By incorporating the processed recording as additional guide into both algorithms, our novel joint autoregressive framework leverages temporal-spectral correlations of speech to resolve spatially challenging speaker constellations. Consequently, our proposed method significantly improves tracking and enhancement of closely spaced speakers, consistently outperforming comparable non-autoregressive methods on a synthetic dataset. Real-world recordings complement these findings in complex scenarios with multiple speaker crossings and varying speaker-to-array distances.
【46】AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering
标题:AQUA-Bench:超越在音频问题回答中寻找知道何时没有答案
链接:https://arxiv.org/abs/2601.12248
备注:Accepted to ICASSP 2026. Project Website: https://kuan2jiu99.github.io/AQUA-Bench-demo/
摘要:音频感知的大型语言模型的最新进展在音频问题回答上表现出强大的性能。然而,现有的基准主要涵盖可回答的问题,而忽略了无法回答的问题的挑战,即无法从音频中推断出可靠的答案。这种情况在现实世界中很常见,问题可能会误导,不适定或与信息不相容。为了解决这一差距,我们提出了AQUA-Bench,音频问题无法回答性评估的基准。它系统地评估了三种情况:缺席答案检测(缺少正确选项),不兼容答案集检测(选择与问题完全不匹配)和不兼容音频问题检测(问题不相关或在音频中缺乏足够的基础)。通过评估这些案例,AQUA-Bench提供了一个严格的模型可靠性衡量标准,并促进了更强大和更值得信赖的音频语言系统的开发。我们的实验表明,虽然模型在标准的可回答任务上表现出色,但它们经常面临无法回答的挑战,这表明当前音频语言理解中存在盲点。
摘要:Recent advances in audio-aware large language models have shown strong performance on audio question answering. However, existing benchmarks mainly cover answerable questions and overlook the challenge of unanswerable ones, where no reliable answer can be inferred from the audio. Such cases are common in real-world settings, where questions may be misleading, ill-posed, or incompatible with the information. To address this gap, we present AQUA-Bench, a benchmark for Audio Question Unanswerability Assessment. It systematically evaluates three scenarios: Absent Answer Detection (the correct option is missing), Incompatible Answer Set Detection (choices are categorically mismatched with the question), and Incompatible Audio Question Detection (the question is irrelevant or lacks sufficient grounding in the audio). By assessing these cases, AQUA-Bench offers a rigorous measure of model reliability and promotes the development of audio-language systems that are more robust and trustworthy. Our experiments suggest that while models excel on standard answerable tasks, they often face notable challenges with unanswerable ones, pointing to a blind spot in current audio-language understanding.
【47】A Survey on 30+ Years of Automatic Singing Assessment and Singing Information Processing
标题:自动歌唱评估和歌唱信息处理30多年的概况
链接:https://arxiv.org/abs/2601.12153
摘要:在过去的三十年里,自动歌唱评估和歌唱信息处理已经发展到支持歌唱教学,表演分析和声乐训练。虽然第一种方法通过从实时视觉反馈和声学生物反馈到复杂的音高跟踪和频谱分析的计算指标客观地评估歌手的表现,但后一种方法将预测器声乐信号与目标参考进行比较,以捕获嵌入在歌声中的细微差别的数据。值得注意的进步包括开发了显著改善实时视觉反馈的交互式系统,以及集成了机器学习和深度神经网络架构,提高了语音信号处理的精度。本调查批判性地审查了文献,以绘制这些技术的历史演变,同时确定和讨论关键差距。分析揭示了持续存在的挑战,例如缺乏标准化的评估框架,难以可靠地将声音信号与各种噪声源分离,以及未充分利用先进的数字信号处理和人工智能方法来捕捉艺术表现力。通过详细介绍这些限制和相应的技术进步,本文的评论表明,解决这些问题可以弥合客观的计算评估和主观的人性化评价之间的差距歌唱表演,最终提高技术的准确性和教学的相关性自动歌唱评价系统。
摘要:Automatic Singing Assessment and Singing Information Processing have evolved over the past three decades to support singing pedagogy, performance analysis, and vocal training. While the first approach objectively evaluates a singer's performance through computational metrics ranging from real-time visual feedback and acoustical biofeedback to sophisticated pitch tracking and spectral analysis, the latter method compares a predictor vocal signal with a target reference to capture nuanced data embedded in the singing voice. Notable advancements include the development of interactive systems that have significantly improved real-time visual feedback, and the integration of machine learning and deep neural network architectures that enhance the precision of vocal signal processing. This survey critically examines the literature to map the historical evolution of these technologies, while identifying and discussing key gaps. The analysis reveals persistent challenges, such as the lack of standardized evaluation frameworks, difficulties in reliably separating vocal signals from various noise sources, and the underutilization of advanced digital signal processing and artificial intelligence methodologies for capturing artistic expressivity. By detailing these limitations and the corresponding technological advances, this review demonstrates how addressing these issues can bridge the gap between objective computational assessments and subjective human-like evaluations of singing performance, ultimately enhancing both the technical accuracy and pedagogical relevance of automated singing evaluation systems.
【48】Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
标题:单音音乐中基本频率和准确发声概率的轻量级自监督检测
链接:https://arxiv.org/abs/2601.11768
备注:12 pages, 6 figures, 3 tables, and an appendix, Accepted for publication at ICPRAM 2026 in Marbella, Spain, on March 2, 2026
摘要:可靠的基频(F0)和浊音估计对于神经合成是必不可少的,然而许多基音提取器依赖于大的标记语料库并且在真实的记录伪影下退化。我们提出了一个轻量级的,完全自我监督的框架,联合F 0估计和发声推理,旨在从有限的音频快速单仪器培训。使用CQT特征的转置等变学习,我们引入了EM风格的迭代重加权方案,该方案使用移位交叉熵(SCE)一致性作为可靠性信号来抑制无信息的噪声/无声帧。所得到的权重提供置信度分数,使得能够在没有手动注释的情况下对单独的轻量级发声分类器进行伪标记。在MedleyDB上训练并在MDB-stem-synth地面实况上进行评估,我们的方法实现了具有竞争力的跨语料库性能(RPA 95.84,RCA 96.24)并展示了跨仪器泛化。
摘要:Reliable fundamental frequency (F 0) and voicing estimation is essential for neural synthesis, yet many pitch extractors depend on large labeled corpora and degrade under realistic recording artifacts. We propose a lightweight, fully self-supervised framework for joint F 0 estimation and voicing inference, designed for rapid single-instrument training from limited audio. Using transposition-equivariant learning on CQT features, we introduce an EM-style iterative reweighting scheme that uses Shift Cross-Entropy (SCE) consistency as a reliability signal to suppress uninformative noisy/unvoiced frames. The resulting weights provide confidence scores that enable pseudo-labeling for a separate lightweight voicing classifier without manual annotations. Trained on MedleyDB and evaluated on MDB-stem-synth ground truth, our method achieves competitive cross-corpus performance (RPA 95.84, RCA 96.24) and demonstrates cross-instrument generalization.
【1】MATE: Matryoshka Audio-Text Embeddings for Open-Vocabulary Keyword Spotting
标题:MATE:Matryoshka音频文本嵌入,用于开放词汇关键词定位
链接:https://arxiv.org/abs/2601.14012
备注:5 pages, 1 figure, Accepted at ICASSP 2026
摘要:基于文本注册的开放式词汇表关键字识别(KWS)已成为固定短语触发器的灵活替代方案。从嵌入学习的角度来看,现有的话语级匹配方法在单个固定维度上学习嵌入。我们从这个设计出发,提出了Matryoshka音频文本嵌入(MATE),一个双编码器框架,通过嵌套子嵌入(“前缀”)在一个单一的向量编码多个嵌入粒度。具体来说,我们引入了PCA引导的前缀对齐:每个前缀大小的全文嵌入的PCA压缩版本作为教师目标来对齐音频和文本前缀。这种对齐将突出的关键字线索集中在低维前缀中,而更高的维度则添加细节。MATE使用音频文本KWS的标准深度度量学习目标进行训练,并且是损失不可知的。据我们所知,这是首次将matryoshka风格的嵌入应用于KWS,在WSJ和LibriPhrase上实现了最先进的结果,而没有任何推理开销。
摘要:Open-vocabulary keyword spotting (KWS) with text-based enrollment has emerged as a flexible alternative to fixed-phrase triggers. Prior utterance-level matching methods, from an embedding-learning standpoint, learn embeddings at a single fixed dimensionality. We depart from this design and propose Matryoshka Audio-Text Embeddings (MATE), a dual-encoder framework that encodes multiple embedding granularities within a single vector via nested sub-embeddings ("prefixes"). Specifically, we introduce a PCA-guided prefix alignment: PCA-compressed versions of the full text embedding for each prefix size serve as teacher targets to align both audio and text prefixes. This alignment concentrates salient keyword cues in lower-dimensional prefixes, while higher dimensions add detail. MATE is trained with standard deep metric learning objectives for audio-text KWS, and is loss-agnostic. To our knowledge, this is the first application of matryoshka-style embeddings to KWS, achieving state-of-the-art results on WSJ and LibriPhrase without any inference overhead.
【2】DAME: Duration-Aware Matryoshka Embedding for Duration-Robust Speaker Verification
标题:DAME:持续时间感知Matryoshka嵌入持续时间稳健的说话者验证
链接:https://arxiv.org/abs/2601.13999
备注:5 pages, 2 figures, Accepted at ICASSP 2026
摘要:短话语说话人确认仍然是具有挑战性的,由于有限的说话人判别线索,在短的语音段。虽然现有的方法专注于增强说话人编码器,但嵌入学习策略仍然强制将单个固定维度的表示重新用于任何长度的话语,从而使容量与不同持续时间的可用信息不一致。我们提出了持续时间感知Matryoshka嵌入(DAME),一个模型不可知的框架,建立一个嵌套的层次结构的子嵌入对齐的话语持续时间:低维表示捕捉紧凑的扬声器特征从短话语,而更高的维度编码更丰富的细节从较长的语音。DAME支持从头开始的训练和微调,并作为传统的大幅度微调的直接替代方案,持续提高整个持续时间的性能。在VoxCeleb 1-O/E/H和VOiCES评估集上,DAME始终降低了1-s和其他短时间试验的相等错误率,同时保持了全长性能,没有额外的推理成本。这些增益在一般训练和微调设置下在各种扬声器编码器架构中推广。
摘要:Short-utterance speaker verification remains challenging due to limited speaker-discriminative cues in short speech segments. While existing methods focus on enhancing speaker encoders, the embedding learning strategy still forces a single fixed-dimensional representation reused for utterances of any length, leaving capacity misaligned with the information available at different durations. We propose Duration-Aware Matryoshka Embedding (DAME), a model-agnostic framework that builds a nested hierarchy of sub-embeddings aligned to utterance durations: lower-dimensional representations capture compact speaker traits from short utterances, while higher dimensions encode richer details from longer speech. DAME supports both training from scratch and fine-tuning, and serves as a direct alternative to conventional large-margin fine-tuning, consistently improving performance across durations. On the VoxCeleb1-O/E/H and VOiCES evaluation sets, DAME consistently reduces the equal error rate on 1-s and other short-duration trials, while maintaining full-length performance with no additional inference cost. These gains generalize across various speaker encoder architectures under both general training and fine-tuning setups.
【3】Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization via Neural Audio Codec and Language Models
标题:Stream-Voice-Anon:通过神经音频编解码器和语言模型增强实时说话者语音化的实用性
链接:https://arxiv.org/abs/2601.13948
备注:Accepted by ICASSP2026
摘要:保护说话人身份对于在线语音应用至关重要,但流媒体说话人匿名化(SA)仍未得到充分研究。最近的研究表明,神经音频编解码器(NAC)提供了优越的扬声器功能解纠缠和语言保真度。NAC还可以与因果语言模型(LM)一起使用,以增强语言保真度并提示对流式任务的控制。然而,现有的基于NAC的在线LM系统被设计用于语音转换(VC)而不是匿名化,缺乏隐私保护所需的技术。在这些进展的基础上,我们提出了流语音匿名,它适应现代因果LM为基础的NAC架构,专门为流SA集成匿名化技术。我们的匿名化方法结合了伪扬声器表示采样,扬声器嵌入混合和LM调节的各种提示选择策略,利用量化内容代码的解纠缠特性,以防止扬声器信息泄漏。此外,我们比较了动态和固定延迟配置,以探索实时场景中的延迟隐私权衡。在VoicePrivacy 2024 Challenge协议下,Stream-Voice-Anon在可懂度方面实现了实质性的改进(相对WER降低高达46%)和情绪保护(相对UAR高达28%),与之前最先进的流媒体方法DarkStream相比,同时保持相当的延迟(180 ms vs 200 ms)和针对懒惰通知攻击者的隐私保护,尽管对半通知攻击者显示出15%的相对下降。
摘要:Protecting speaker identity is crucial for online voice applications, yet streaming speaker anonymization (SA) remains underexplored. Recent research has demonstrated that neural audio codec (NAC) provides superior speaker feature disentanglement and linguistic fidelity. NAC can also be used with causal language models (LM) to enhance linguistic fidelity and prompt control for streaming tasks. However, existing NAC-based online LM systems are designed for voice conversion (VC) rather than anonymization, lacking the techniques required for privacy protection. Building on these advances, we present Stream-Voice-Anon, which adapts modern causal LM-based NAC architectures specifically for streaming SA by integrating anonymization techniques. Our anonymization approach incorporates pseudo-speaker representation sampling, a speaker embedding mixing and diverse prompt selection strategies for LM conditioning that leverage the disentanglement properties of quantized content codes to prevent speaker information leakage. Additionally, we compare dynamic and fixed delay configurations to explore latency-privacy trade-offs in real-time scenarios. Under the VoicePrivacy 2024 Challenge protocol, Stream-Voice-Anon achieves substantial improvements in intelligibility (up to 46% relative WER reduction) and emotion preservation (up to 28% UAR relative) compared to the previous state-of-the-art streaming method DarkStream while maintaining comparable latency (180ms vs 200ms) and privacy protection against lazy-informed attackers, though showing 15% relative degradation against semi-informed attackers.
【4】Synthetic Singers: A Review of Deep-Learning-based Singing Voice Synthesis Approaches
标题:合成歌手:基于深度学习的歌唱声音合成方法回顾
链接:https://arxiv.org/abs/2601.13910
备注:Accepetd by IJCNLP-AACL 2025(Oral)
摘要:歌唱声合成技术的最新进展引起了学术界和工业界的广泛关注。随着大型语言模型和新的生成范式的出现,产生可控的、高保真的歌唱声音已经成为一个可以实现的目标。然而,该领域仍然缺乏系统分析基于深度学习的歌唱声音合成系统及其使能技术的全面调查。为了解决上述问题,本调查首先按任务类型对现有系统进行分类,然后将当前架构组织成两个主要范例:级联和端到端方法。此外,我们提供了一个深入的分析核心技术,涵盖歌唱建模和控制技术。最后,我们回顾了支持培训和评估的相关数据集、注释工具和评估基准。在附录中,我们介绍了SVS的培训策略和进一步的讨论。本文综述了国内外关于SVS模型的最新研究成果,为研究人员和工程技术人员提供了有益的参考。相关材料可在https://github.com/David-Pigeon/SyntheticSingers上查阅。
摘要:Recent advances in singing voice synthesis (SVS) have attracted substantial attention from both academia and industry. With the advent of large language models and novel generative paradigms, producing controllable, high-fidelity singing voices has become an attainable goal. Yet the field still lacks a comprehensive survey that systematically analyzes deep-learning-based singing voice synthesis systems and their enabling technologies. To address the aforementioned issue, this survey first categorizes existing systems by task type and then organizes current architectures into two major paradigms: cascaded and end-to-end approaches. Moreover, we provide an in-depth analysis of core technologies, covering singing modeling and control techniques. Finally, we review relevant datasets, annotation tools, and evaluation benchmarks that support training and assessment. In appendix, we introduce training strategies and further discussion of SVS. This survey provides an up-to-date review of the literature on SVS models, which would be a useful reference for both researchers and engineers. Related materials are available at https://github.com/David-Pigeon/SyntheticSingers.
【5】Co-Initialization of Control Filter and Secondary Path via Meta-Learning for Active Noise Control
标题:通过元学习协调控制过滤器和辅助路径以实现主动噪音控制
链接:https://arxiv.org/abs/2601.13849
摘要:有源噪声控制(ANC)必须在声学环境变化时快速适应,但早期性能在很大程度上取决于初始化。我们通过模型不可知元学习(MAML)共同初始化来解决这个问题,该共同初始化为基于FxLMS的ANC联合设置控制滤波器和次级路径模型,同时保持运行时算法不变。初始化器在一小组测量路径上使用短两相内环进行预训练,该内环模拟识别,然后进行残余噪声降低,并通过简单地设置学习的初始系数来应用。在在线次级路径建模FxLMS测试平台中,与无需重新初始化的基线相比,该方法可降低早期误差、缩短到达目标的时间、降低随机噪声能量,并在路径变化后实现更快的恢复。该方法在环境变化下为前馈ANC提供了一个简单的快速启动,需要一小组路径进行预训练。
摘要:Active noise control (ANC) must adapt quickly when the acoustic environment changes, yet early performance is largely dictated by initialization. We address this with a Model-Agnostic Meta-Learning (MAML) co-initialization that jointly sets the control filter and the secondary-path model for FxLMS-based ANC while keeping the runtime algorithm unchanged. The initializer is pre-trained on a small set of measured paths using short two-phase inner loops that mimic identification followed by residual-noise reduction, and is applied by simply setting the learned initial coefficients. In an online secondary path modeling FxLMS testbed, it yields lower early-stage error, shorter time-to-target, reduced auxiliary-noise energy, and faster recovery after path changes than a baseline without re-initialization. The method provides a simple fast start for feedforward ANC under environment changes, requiring a small set of paths to pre-train.
【6】S$^2$Voice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion
标题:S $' 2$Voice:风格感知自回归建模,具有增强的条件反射,用于歌唱风格转换
链接:https://arxiv.org/abs/2601.13629
备注:accepted to ICASSP 2026
摘要:我们介绍S$^2$Voice,这是2025年歌唱声音转换挑战赛(SVCC)的获奖系统,适用于域内和zero-shot演唱风格转换曲目。基于强大的两阶段Vevo基线,S$^2$Voice通过几项贡献推进了风格控制和鲁棒性。首先,我们通过FiLM风格的层规范条件和风格感知的交叉注意力将风格嵌入集成到自回归大语言模型(AR LLM)中,以增强细粒度的风格建模。其次,我们引入了一个全局的扬声器嵌入到流匹配Transformer,以提高音色的相似性。第三,我们通过自动化管道来策划一个大型,高质量的歌唱语料库,用于网络采集,声乐分离和转录改进。最后,我们采用了一个多阶段的训练策略相结合的监督微调(SFT)和直接偏好优化(DPO)。主观听力测试证实了我们的系统的优越性能:领先的风格相似性和歌手相似性的任务1,跨自然,风格相似性,歌手相似性的任务2。消融研究表明,我们的贡献,在提高风格的保真度,音色的保存和推广的有效性。音频示例可用~\footnote{https://honee-w.github.io/SVC-Socke-Demo/}。
摘要:We present S$^2$Voice, the winning system of the Singing Voice Conversion Challenge (SVCC) 2025 for both the in-domain and zero-shot singing style conversion tracks. Built on the strong two-stage Vevo baseline, S$^2$Voice advances style control and robustness through several contributions. First, we integrate style embeddings into the autoregressive large language model (AR LLM) via a FiLM-style layer-norm conditioning and a style-aware cross-attention for enhanced fine-grained style modeling. Second, we introduce a global speaker embedding into the flow-matching transformer to improve timbre similarity. Third, we curate a large, high-quality singing corpus via an automated pipeline for web harvesting, vocal separation, and transcript refinement. Finally, we employ a multi-stage training strategy combining supervised fine-tuning (SFT) and direct preference optimization (DPO). Subjective listening tests confirm our system's superior performance: leading in style similarity and singer similarity for Task 1, and across naturalness, style similarity, and singer similarity for Task 2. Ablation studies demonstrate the effectiveness of our contributions in enhancing style fidelity, timbre preservation, and generalization. Audio samples are available~\footnote{https://honee-w.github.io/SVC-Challenge-Demo/}.
【7】ICASSP 2026 URGENT Speech Enhancement Challenge
标题:ICASP 2026紧急语音增强挑战赛
链接:https://arxiv.org/abs/2601.13531
备注:The overview paper of the ICASSP 2026 URGENT Speech Enhancement Challenge
摘要:ICASSP 2026紧急挑战赛通过关注处理不同失真、域和输入条件的通用语音增强(SE)系统来推进该系列。这篇综述论文详细介绍了挑战的动机、任务定义、数据集、基线系统、评估协议和结果。这项挑战分为两个相辅相成的轨道。Track 1侧重于通用语音增强,而Track 2介绍了增强语音的语音质量评估。该挑战赛吸引了80多个团队注册,其中29个提交了有效参赛作品,显示出社区对强大的SE技术的浓厚兴趣。
摘要:The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task definitions, datasets, baseline systems, evaluation protocols, and results. The challenge is divided into two complementary tracks. Track 1 focuses on universal speech enhancement, while Track 2 introduces speech quality assessment for enhanced speech. The challenge attracted over 80 team registrations, with 29 submitting valid entries, demonstrating significant community interest in robust SE technologies.
【8】RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models
标题:WLBR:针对上下文语音大型语言模型的带有偏向奖励的强化学习
链接:https://arxiv.org/abs/2601.13409
备注:Accepted to the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)
摘要:语音大语言模型(LLM)在端到端语音理解和识别方面取得了重大进展,但它们仍然难以准确识别罕见单词和特定领域的术语。本文提出了一种新的微调方法,强化学习与偏置奖励(RLBR),它采用了一个专门的偏置词首选奖励显式强调偏置词的奖励计算。此外,我们引入了参考感知机制,该机制通过参考转录扩展了强化学习算法,以加强潜在的轨迹探索空间。在LibriSpeech语料库上进行的各种偏置列表大小的实验表明,RLBR在强监督微调(SFT)基线上提供了实质性的性能改进,并且始终优于最近发布的几种方法。所提出的方法在LibriSpeech测试干净和测试其他集上实现了出色的性能,对于100,500和1000的偏置列表大小,分别达到0.59% /2.11%,1.09% /3.24%和1.36% / 4.04%的偏置字错误率(BWRs),而不影响整体WRs。
摘要:Speech large language models (LLMs) have driven significant progress in end-to-end speech understanding and recognition, yet they continue to struggle with accurately recognizing rare words and domain-specific terminology. This paper presents a novel fine-tuning method, Reinforcement Learning with Biasing Rewards (RLBR), which employs a specialized biasing words preferred reward to explicitly emphasize biasing words in the reward calculation. In addition, we introduce reference-aware mechanisms that extend the reinforcement learning algorithm with reference transcription to strengthen the potential trajectory exploration space. Experiments on the LibriSpeech corpus across various biasing list sizes demonstrate that RLBR delivers substantial performance improvements over a strong supervised fine-tuning (SFT) baseline and consistently outperforms several recently published methods. The proposed approach achieves excellent performance on the LibriSpeech test-clean and test-other sets, reaching Biasing Word Error Rates (BWERs) of 0.59% / 2.11%, 1.09% / 3.24%, and 1.36% / 4.04% for biasing list sizes of 100, 500, and 1000, respectively, without compromising the overall WERs.
【9】AMDM-SE: Attention-based Multichannel Diffusion Model for Speech Enhancement
标题:AMDM-SE:基于注意力的语音增强多通道扩散模型
链接:https://arxiv.org/abs/2601.13140
摘要:扩散模型最近在从噪声输入重建图像方面取得了令人印象深刻的结果,并且通过将时频表示视为图像,类似的想法已被应用于语音增强。随着多麦克风设备的普及,我们扩展了最先进的基于扩散的方法,以利用多通道输入来提高性能。基于多通道扩散的增强仍处于起步阶段,之前的工作对空间建模注意力等高级机制的使用有限--本文解决了这一差距。我们提出了AMDM-SE,一个基于注意力的语音增强多通道扩散模型,专为降噪而设计。AMDM-SE通过新颖的跨通道时频注意块利用空间通道间信息,从而能够在生成扩散框架内忠实地重建细粒度信号细节。在CHiME-3基准测试中,AMDM-SE的性能优于单通道扩散基线和无注意力的多通道模型,以及基于DNN的强大预测方法。模拟数据实验进一步强调了所提出的多通道注意机制的重要性。总的来说,我们的研究结果表明,将有针对性的多通道注意力扩散模型大大提高了降噪。虽然基于多通道扩散的语音增强仍然是一个新兴的领域,我们的工作贡献了一个新的和互补的方法,在这个方向上不断增长的研究机构。
摘要:Diffusion models have recently achieved impressive results in reconstructing images from noisy inputs, and similar ideas have been applied to speech enhancement by treating time-frequency representations as images. With the ubiquity of multi-microphone devices, we extend state-of-the-art diffusion-based methods to exploit multichannel inputs for improved performance. Multichannel diffusion-based enhancement remains in its infancy, with prior work making limited use of advanced mechanisms such as attention for spatial modeling - a gap addressed in this paper. We propose AMDM-SE, an Attention-based Multichannel Diffusion Model for Speech Enhancement, designed specifically for noise reduction. AMDM-SE leverages spatial inter-channel information through a novel cross-channel time-frequency attention block, enabling faithful reconstruction of fine-grained signal details within a generative diffusion framework. On the CHiME-3 benchmark, AMDM-SE outperforms both a single-channel diffusion baseline and a multichannel model without attention, as well as a strong DNN-based predictive method. Simulated-data experiments further underscore the importance of the proposed multichannel attention mechanism. Overall, our results show that incorporating targeted multichannel attention into diffusion models substantially improves noise reduction. While multichannel diffusion-based speech enhancement is still an emerging field, our work contributes a new and complementary approach to the growing body of research in this direction.
【10】Content Leakage in LibriSpeech and Its Impact on the Privacy Evaluation of Speaker Anonymization
标题:LibriSpeech中的内容泄露及其对说话者匿名化隐私评估的影响
链接:https://arxiv.org/abs/2601.13107
备注:Accepted to ICASSP 2026
摘要:说话人匿名的目的是隐藏说话人的身份,而不考虑语言内容。在这项研究中,我们揭示了Librispeech的一个弱点,该数据集通常用于评估匿名者:Librispeech说话者阅读的书籍是如此独特,以至于说话者可以通过他们的词汇来识别。即使是完美的匿名者也无法阻止这种身份泄露。EdAcc数据集在这方面更好:只有少数说话者可以通过他们的词汇表识别,鼓励攻击者在其他地方寻找匿名说话者的身份。EdAcc还包括自发的演讲和更多样化的演讲者,补充了Librispeech,并对匿名者的工作方式提供了更多的见解。
摘要:Speaker anonymization aims to conceal a speaker's identity, without considering the linguistic content. In this study, we reveal a weakness of Librispeech, the dataset that is commonly used to evaluate anonymizers: the books read by the Librispeech speakers are so distinct, that speakers can be identified by their vocabularies. Even perfect anonymizers cannot prevent this identity leakage. The EdAcc dataset is better in this regard: only a few speakers can be identified through their vocabularies, encouraging the attacker to look elsewhere for the identities of the anonymized speakers. EdAcc also comprises spontaneous speech and more diverse speakers, complementing Librispeech and giving more insights into how anonymizers work.
【11】VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
标题:VoCodec:一种高效的轻量级低比特率语音编解码器
链接:https://arxiv.org/abs/2601.13055
摘要:端到端神经语音编解码器的最新进展使得能够以极低的比特率压缩音频,同时保持高保真重建。同时,低计算复杂度和低延迟对于实时通信至关重要。在本文中,我们提出VoCodec,语音编解码器模型具有计算复杂度只有349.29 M乘累加运算每秒(MAC/s)和30 ms的延迟。与竞争力的声码器Vocos作为其骨干,该模型排名第四的轨道1在2025年LRAC挑战,并取得了最高的主观评价分数(MUSHRA)的干净语音测试集。此外,我们在前端级联了一个轻量级的神经网络,以扩展其语音增强能力。实验结果表明,这两个系统在多个评价指标上都取得了有竞争力的性能。语音样本可以在https://acceleration123.github.io/上找到。
摘要:Recent advancements in end-to-end neural speech codecs enable compressing audio at extremely low bitrates while maintaining high-fidelity reconstruction. Meanwhile, low computational complexity and low latency are crucial for real-time communication. In this paper, we propose VoCodec, a speech codec model featuring a computational complexity of only 349.29M multiply-accumulate operations per second (MACs/s) and a latency of 30 ms. With the competitive vocoder Vocos as its backbone, the proposed model ranked fourth on Track 1 in the 2025 LRAC Challenge and achieved the highest subjective evaluation score (MUSHRA) on the clean speech test set. Additionally, we cascade a lightweight neural network at the front end to extend its capability of speech enhancement. Experimental results demonstrate that the two systems achieve competitive performance across multiple evaluation metrics. Speech samples can be found at https://acceleration123.github.io/.
【12】ImmersiveFlow: Stereo-to-7.1.4 spatial audio generation with flow matching
标题:ImmersiveFlow:立体声到7.1.4的空间音频生成,具有流匹配
链接:https://arxiv.org/abs/2601.12950
备注:5 pages, 3 figures, 2 tables
摘要:沉浸式空间音频对于从AR/VR到家庭娱乐和汽车音响系统的应用越来越重要。然而,现有的生成方法仍然受限于低维格式,诸如双耳音频和一阶高保真度立体声(FOA)。双耳渲染固有地限于耳机回放,而FOA遭受空间混叠和高频分辨率不足。为了克服这些限制,我们引入了ImmersiveFlow,这是第一个端到端的生成框架,可以直接从立体声输入合成离散的7.1.4格式空间音频。ImmersiveFlow利用Flow Matching来学习从立体声输入到预训练的VAE潜在空间内的多通道空间特征的轨迹。在推理时,流量匹配模型预测的潜在特征由VAE解码并转换为最终的7.1.4波形。综合的客观和主观评估表明,我们的方法产生感知丰富的声场和增强的外部化,显着优于传统的上混技术。代码实现和音频示例在https://github.com/violet-audio/ImmersiveFlow上提供。
摘要:Immersive spatial audio has become increasingly critical for applications ranging from AR/VR to home entertainment and automotive sound systems. However, existing generative methods remain constrained to low-dimensional formats such as binaural audio and First-Order Ambisonics (FOA). Binaural rendering is inherently limited to headphone playback, while FOA suffers from spatial aliasing and insufficient resolution for high-frequency. To overcome these limitations, we introduce ImmersiveFlow, the first end-to-end generative framework that directly synthesizes discrete 7.1.4 format spatial audio from stereo input. ImmersiveFlow leverages Flow Matching to learn trajectories from stereo inputs to multichannel spatial features within a pretrained VAE latent space. At inference, the Flow Matching model predicted latent features are decoded by the VAE and converted into the final 7.1.4 waveform. Comprehensive objective and subjective evaluations demonstrate that our method produces perceptually rich sound fields and enhanced externalization, significantly outperforming traditional upmixing techniques. Code implementations and audio samples are provided at: https://github.com/violet-audio/ImmersiveFlow.
【13】Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech
标题:自适应说话人嵌入自增强用于短注册语音的个人语音活动检测
链接:https://arxiv.org/abs/2601.12769
备注:Accepted by ICASSP 2026
摘要:个人语音活动检测(PVAD)是识别混合中的目标说话人段的关键,但其性能在很大程度上取决于说话人嵌入的质量。一个关键的实际限制是简短的注册语音-例如唤醒词-提供有限的提示。本文提出了一种新的自适应说话人嵌入自增强策略,增强PVAD的性能,通过添加融合的关键帧嵌入从混合语音中提取的原始注册嵌入。此外,我们引入了一个长期的自适应策略,在检测过程中迭代细化嵌入,减轻扬声器的时间变化。实验表明,在短注册条件下,召回率,准确率和F1分数显着提高,在五次迭代更新后匹配全长注册性能。源代码可在https://anonymous.4open.science/r/ASE-PVAD-E5D6上获得。
摘要:Personal Voice Activity Detection (PVAD) is crucial for identifying target speaker segments in the mixture, yet its performance heavily depends on the quality of speaker embeddings. A key practical limitation is the short enrollment speech--such as a wake-up word--which provides limited cues. This paper proposes a novel adaptive speaker embedding self-augmentation strategy that enhances PVAD performance by augmenting the original enrollment embeddings through additive fusion of keyframe embeddings extracted from mixed speech. Furthermore, we introduce a long-term adaptation strategy to iteratively refine embeddings during detection, mitigating speaker temporal variability. Experiments show significant gains in recall, precision, and F1-score under short enrollment conditions, matching full-length enrollment performance after five iterative updates. The source code is available at https://anonymous.4open.science/r/ASE-PVAD-E5D6 .
【14】CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction
标题:CodeSep:低比特率编解码驱动的语音分离,具有基本令牌解纠缠和辅助令牌序列预测
链接:https://arxiv.org/abs/2601.12757
备注:Accepted by ICASSP 2026
摘要:本文针对一个新的场景,集成语音分离与语音压缩,旨在解开多个扬声器,同时产生离散表示有效的传输或存储,在在线会议和对话存档的应用程序。为了解决这个问题,我们提出了CodeSep,一个编解码器驱动的模型,联合执行语音分离和低比特率压缩。CodeSep包括基于残差矢量量化器(RVQ)的普通神经语音编解码器、基本令牌解纠缠(BTD)模块和并行令牌串行预测(ATSP)模块。BTD模块将混合语音梅尔频谱图分解为每个说话人的基本标记,ATSP模块对基本标记进行细化,连续预测辅助标记,最后通过编解码器对所有标记进行解码,重构出分离的波形。在训练期间,编解码器的RVQ提供具有置换不变和基于教师强制的交叉熵损失的监督。由于只传输或存储基本令牌,CodeSep实现了低比特率压缩。实验结果表明,CodeSep算法在1 kbps的速度下就能获得令人满意的分离效果。
摘要:This paper targets a new scenario that integrates speech separation with speech compression, aiming to disentangle multiple speakers while producing discrete representations for efficient transmission or storage, with applications in online meetings and dialogue archiving. To address this scenario, we propose CodeSep, a codec-driven model that jointly performs speech separation and low-bitrate compression. CodeSep comprises a residual vector quantizer (RVQ)-based plain neural speech codec, a base-token disentanglement (BTD) module, and parallel auxiliary-token serial prediction (ATSP) modules. The BTD module disentangles mixed-speech mel-spectrograms into base tokens for each speaker, which are then refined by ATSP modules to serially predict auxiliary tokens, and finally, all tokens are decoded to reconstruct separated waveforms through the codec decoder. During training, the codec's RVQ provides supervision with permutation-invariant and teacher-forcing-based cross-entropy losses. As only base tokens are transmitted or stored, CodeSep achieves low-bitrate compression. Experimental results show that CodeSep attains satisfactory separation performance at only 1 kbps compared with baseline methods.
【15】Improving Audio Question Answering with Variational Inference
标题:利用变分推理改进音频问题回答
链接:https://arxiv.org/abs/2601.12700
备注:ICASSP 2026
摘要:变分推理(VI)提供了一个原则性的框架,用于估计模型参数的后验分布,从而在优化过程中对权重不确定性进行显式建模。通过捕捉这种不确定性,VI提高了预测的可靠性,产生更好的校准输出。在这项工作中,我们调查的好处,具有挑战性的多模态理解和推理,通过应用改进的变分在线牛顿(IVON),最近的VI优化器,微调多模态大型语言模型的音频问答任务。我们的研究结果表明,VI不仅提高了预测精度,而且还显着提高了校准,减少了模型的过度自信。这些进展进一步支持风险敏感的应用,如选择性预测,其中可靠的置信度估计至关重要。
摘要:Variational inference (VI) provides a principled framework for estimating posterior distributions over model parameters, enabling explicit modeling of weight uncertainty during optimization. By capturing this uncertainty, VI improves the reliability of predictions, yielding better calibrated outputs. In this work, we investigate the benefits of VI for challenging multimodal understanding and reasoning by applying the Improved Variational Online Newton (IVON), a recent VI optimizer, to fine-tuning a multimodal large language model on audio question answering tasks. Our results show that VI not only enhances predictive accuracy but also significantly improves calibration, reducing the model's overconfidence. These advances further support risk-sensitive applications such as selective prediction, where reliable confidence estimates are crucial.
【16】SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
标题:Scrum:具有可变持续时间音频和多目标训练的可扩展音频预训练
链接:https://arxiv.org/abs/2601.12594
备注:Accepted to ICASSP 2026
摘要:对比语言-音频预训练(CLAP)在学习语义丰富的音频表示方面取得了显着的成功,并被广泛用于各种与音频相关的任务。然而,目前的CLAP模型面临着几个关键的限制。首先,它们通常在相对较小的数据集上训练,通常包括几百万个音频样本。其次,现有的CLAP模型被限制为短且固定的持续时间,这限制了它们在具有可变持续时间音频的真实世界场景中的使用。第三,标准的对比训练目标对全局表示进行操作,这可能会阻碍密集的细粒度音频特征的学习。为了应对这些挑战,我们引入了可扩展的音频预训练(Scalable Audio-Pretraining,简称SSTO),它将语言-音频预训练扩展到1.09亿个音频-文本对,具有可变的音频持续时间,并包含多个训练目标。在单阶段训练中,Sort将对比度损失与额外的自我监督和字幕损失统一起来,促进了更丰富的密集音频表示的学习。该模型在音频文本检索和zero-shot音频分类任务上实现了新的最先进的性能,在不同的基准测试中证明了其有效性。
摘要:Contrastive language-audio pretraining (CLAP) has achieved notable success in learning semantically rich audio representations and is widely adopted for various audio-related tasks. However, current CLAP models face several key limitations. First, they are typically trained on relatively small datasets, often comprising a few million audio samples. Second, existing CLAP models are restricted to short and fixed duration, which constrains their usage in real-world scenarios with variable-duration audio. Third, the standard contrastive training objective operates on global representations, which may hinder the learning of dense, fine-grained audio features. To address these challenges, we introduce Scalable Language-Audio Pretraining (SLAP), which scales language-audio pretraining to 109 million audio-text pairs with variable audio durations and incorporates multiple training objectives. SLAP unifies contrastive loss with additional self-supervised and captioning losses in a single-stage training, facilitating the learning of richer dense audio representations. The proposed SLAP model achieves new state-of-the-art performance on audio-text retrieval and zero-shot audio classification tasks, demonstrating its effectiveness across diverse benchmarks.
【17】Robust Online Overdetermined Independent Vector Analysis Based on Bilinear Decomposition
标题:基于双线性分解的鲁棒在线超定独立载体分析
链接:https://arxiv.org/abs/2601.12485
摘要:在线盲源分离对于语音通信和人机交互都是必不可少的。在现有的方法中,超定独立向量分析(OverIVA)通过利用源信号的统计独立性和源与噪声子空间之间的正交性提供了强大的性能。然而,当应用于大型麦克风阵列时,参数的数量迅速增长,这会降低在线估计精度。为了克服这一挑战,我们建议将每个长分离滤波器分解为两个较短滤波器的双线性形式,从而减少参数的数量。由于这两个滤波器是紧密耦合的,我们设计了交替迭代投影算法来依次更新它们。仿真结果表明,该方法在参数少得多的情况下,获得了较好的性能和鲁棒性.
摘要:Online blind source separation is essential for both speech communication and human-machine interaction. Among existing approaches, overdetermined independent vector analysis (OverIVA) delivers strong performance by exploiting the statistical independence of source signals and the orthogonality between source and noise subspaces. However, when applied to large microphone arrays, the number of parameters grows rapidly, which can degrade online estimation accuracy. To overcome this challenge, we propose decomposing each long separation filter into a bilinear form of two shorter filters, thereby reducing the number of parameters. Because the two filters are closely coupled, we design an alternating iterative projection algorithm to update them in turn. Simulation results show that, with far fewer parameters, the proposed method achieves improved performance and robustness.
【18】Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
标题:融合前的净化:迈向无屏蔽语音增强以实现稳健的视听语音识别
链接:https://arxiv.org/abs/2601.12436
备注:Accepted by ICASSP2026
摘要:视听语音识别(AVSR)通常通过将抗噪声的视觉线索与音频信号相结合来提高噪声环境中的识别精度。然而,高噪声音频输入容易将不利干扰引入特征融合过程。为了缓解这一问题,最近的AVSR方法通常采用基于掩码的策略来过滤特征交互和融合期间的音频噪声,但这样的方法有可能丢弃与噪声一起的语义相关信息。在这项工作中,我们提出了一个端到端的噪声鲁棒AVSR框架加上语音增强,消除了显式噪声掩模生成的需要。该框架利用基于Conformer的瓶颈融合模块隐式地在视频辅助下细化噪声音频特征。通过减少模态冗余和增强模态间的相互作用,我们的方法保留语音语义的完整性,以实现强大的识别性能。在公共LRS3基准上的实验评估表明,我们的方法在噪声条件下优于先前的先进的基于掩模的基线。
摘要:Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-noise audio inputs are prone to introducing adverse interference into the feature fusion process. To mitigate this, recent AVSR methods often adopt mask-based strategies to filter audio noise during feature interaction and fusion, yet such methods risk discarding semantically relevant information alongside noise. In this work, we propose an end-to-end noise-robust AVSR framework coupled with speech enhancement, eliminating the need for explicit noise mask generation. This framework leverages a Conformer-based bottleneck fusion module to implicitly refine noisy audio features with video assistance. By reducing modality redundancy and enhancing inter-modal interactions, our method preserves speech semantic integrity to achieve robust recognition performance. Experimental evaluations on the public LRS3 benchmark suggest that our method outperforms prior advanced mask-based baselines under noisy conditions.
【19】Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models
标题:基于条件扩散模型的骨导引导多模式语音增强
链接:https://arxiv.org/abs/2601.12354
备注:Accepted to IEEE ICASSP 2026
摘要:单通道语音增强模型在极端噪声环境中面临显著的性能下降。虽然先前的工作已经表明,互补的骨传导语音可以引导增强,但这种噪声免疫模式的有效整合仍然是一个挑战。本文介绍了一种新的多模态语音增强框架,集成骨导传感器与空气传导麦克风使用条件扩散模型。我们提出的模型显着优于先前建立的多模态技术和强大的扩散为基础的单一模态基线在广泛的声学条件。
摘要:Single-channel speech enhancement models face significant performance degradation in extremely noisy environments. While prior work has shown that complementary bone-conducted speech can guide enhancement, effective integration of this noise-immune modality remains a challenge. This paper introduces a novel multimodal speech enhancement framework that integrates bone-conduction sensors with air-conducted microphones using a conditional diffusion model. Our proposed model significantly outperforms previously established multimodal techniques and a powerful diffusion-based single-modal baseline across a wide range of acoustic conditions.
【20】Adaptive Rotary Steering with Joint Autoregression for Robust Extraction of Closely Moving Speakers in Dynamic Scenarios
标题:具有联合自回归的自适应旋转转向,用于动态场景中近距离移动的扬声器的鲁棒提取
链接:https://arxiv.org/abs/2601.12345
备注:Accepted at IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026
摘要:高保真度立体声的深度空间滤波的最新进展通过在多声道增强之前将声场朝向目标扬声器旋转,在固定多扬声器场景中展示了强大的性能。在动态声学条件下与移动扬声器的适用性,我们建议自动化使用交错跟踪算法的目标的初始方向为条件的旋转转向。然而,对于附近或交叉的扬声器,鲁棒跟踪变得困难,空间线索的增强效果较差。通过将处理后的录音作为额外的指导到这两种算法中,我们的新的联合自回归框架利用语音的时间-频谱相关性来解决空间上具有挑战性的扬声器星座。因此,我们提出的方法显着提高了对密集扬声器的跟踪和增强,在合成数据集上始终优于可比的非自回归方法。真实世界的录音补充了这些发现在复杂的情况下,多个扬声器交叉和不同的扬声器到阵列的距离。
摘要:Latest advances in deep spatial filtering for Ambisonics demonstrate strong performance in stationary multi-speaker scenarios by rotating the sound field toward a target speaker prior to multi-channel enhancement. For applicability in dynamic acoustic conditions with moving speakers, we propose to automate this rotary steering using an interleaved tracking algorithm conditioned on the target's initial direction. However, for nearby or crossing speakers, robust tracking becomes difficult and spatial cues less effective for enhancement. By incorporating the processed recording as additional guide into both algorithms, our novel joint autoregressive framework leverages temporal-spectral correlations of speech to resolve spatially challenging speaker constellations. Consequently, our proposed method significantly improves tracking and enhancement of closely spaced speakers, consistently outperforming comparable non-autoregressive methods on a synthetic dataset. Real-world recordings complement these findings in complex scenarios with multiple speaker crossings and varying speaker-to-array distances.
【21】AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering
标题:AQUA-Bench:超越在音频问题回答中寻找知道何时没有答案
链接:https://arxiv.org/abs/2601.12248
备注:Accepted to ICASSP 2026. Project Website: https://kuan2jiu99.github.io/AQUA-Bench-demo/
摘要:音频感知的大型语言模型的最新进展在音频问题回答上表现出强大的性能。然而,现有的基准主要涵盖可回答的问题,而忽略了无法回答的问题的挑战,即无法从音频中推断出可靠的答案。这种情况在现实世界中很常见,问题可能会误导,不适定或与信息不相容。为了解决这一差距,我们提出了AQUA-Bench,音频问题无法回答性评估的基准。它系统地评估了三种情况:缺席答案检测(缺少正确选项),不兼容答案集检测(选择与问题完全不匹配)和不兼容音频问题检测(问题不相关或在音频中缺乏足够的基础)。通过评估这些案例,AQUA-Bench提供了一个严格的模型可靠性衡量标准,并促进了更强大和更值得信赖的音频语言系统的开发。我们的实验表明,虽然模型在标准的可回答任务上表现出色,但它们经常面临无法回答的挑战,这表明当前音频语言理解中存在盲点。
摘要:Recent advances in audio-aware large language models have shown strong performance on audio question answering. However, existing benchmarks mainly cover answerable questions and overlook the challenge of unanswerable ones, where no reliable answer can be inferred from the audio. Such cases are common in real-world settings, where questions may be misleading, ill-posed, or incompatible with the information. To address this gap, we present AQUA-Bench, a benchmark for Audio Question Unanswerability Assessment. It systematically evaluates three scenarios: Absent Answer Detection (the correct option is missing), Incompatible Answer Set Detection (choices are categorically mismatched with the question), and Incompatible Audio Question Detection (the question is irrelevant or lacks sufficient grounding in the audio). By assessing these cases, AQUA-Bench offers a rigorous measure of model reliability and promotes the development of audio-language systems that are more robust and trustworthy. Our experiments suggest that while models excel on standard answerable tasks, they often face notable challenges with unanswerable ones, pointing to a blind spot in current audio-language understanding.
【22】A Survey on 30+ Years of Automatic Singing Assessment and Singing Information Processing
标题:自动歌唱评估和歌唱信息处理30多年的概况
链接:https://arxiv.org/abs/2601.12153
摘要:在过去的三十年里,自动歌唱评估和歌唱信息处理已经发展到支持歌唱教学,表演分析和声乐训练。虽然第一种方法通过从实时视觉反馈和声学生物反馈到复杂的音高跟踪和频谱分析的计算指标客观地评估歌手的表现,但后一种方法将预测器声乐信号与目标参考进行比较,以捕获嵌入在歌声中的细微差别的数据。值得注意的进步包括开发了显著改善实时视觉反馈的交互式系统,以及集成了机器学习和深度神经网络架构,提高了语音信号处理的精度。本调查批判性地审查了文献,以绘制这些技术的历史演变,同时确定和讨论关键差距。分析揭示了持续存在的挑战,例如缺乏标准化的评估框架,难以可靠地将声音信号与各种噪声源分离,以及未充分利用先进的数字信号处理和人工智能方法来捕捉艺术表现力。通过详细介绍这些限制和相应的技术进步,本文的评论表明,解决这些问题可以弥合客观的计算评估和主观的人性化评价之间的差距歌唱表演,最终提高技术的准确性和教学的相关性自动歌唱评价系统。
摘要:Automatic Singing Assessment and Singing Information Processing have evolved over the past three decades to support singing pedagogy, performance analysis, and vocal training. While the first approach objectively evaluates a singer's performance through computational metrics ranging from real-time visual feedback and acoustical biofeedback to sophisticated pitch tracking and spectral analysis, the latter method compares a predictor vocal signal with a target reference to capture nuanced data embedded in the singing voice. Notable advancements include the development of interactive systems that have significantly improved real-time visual feedback, and the integration of machine learning and deep neural network architectures that enhance the precision of vocal signal processing. This survey critically examines the literature to map the historical evolution of these technologies, while identifying and discussing key gaps. The analysis reveals persistent challenges, such as the lack of standardized evaluation frameworks, difficulties in reliably separating vocal signals from various noise sources, and the underutilization of advanced digital signal processing and artificial intelligence methodologies for capturing artistic expressivity. By detailing these limitations and the corresponding technological advances, this review demonstrates how addressing these issues can bridge the gap between objective computational assessments and subjective human-like evaluations of singing performance, ultimately enhancing both the technical accuracy and pedagogical relevance of automated singing evaluation systems.
【23】Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving
标题:听、看、驾驶:为用户感知的基于VLA的自动驾驶耦合音频指令
链接:https://arxiv.org/abs/2601.12142
备注:Accepted by IV
摘要:视觉语言动作(VLA)模型承诺一个开放的词汇接口,可以将感知模糊性转化为语义基础的驾驶决策,但它们仍然将语言视为推理时固定的静态先验。因此,该模型必须仅从像素推断连续移动的目标,从而产生延迟或过于保守的机动。我们认为,有效的VLA自动驾驶需要一个在线渠道,用户可以影响驾驶与特定的意图。为此,我们提出EchoVLA,一个用户感知的VLA,耦合相机流与原位音频指令。我们通过将自我运动描述转换为合成音频生成的时间对齐的意图特定的语音命令来增强nuScenes数据集。此外,我们将情感语音轨迹对组成多模态思想链(CoT),用于微调基于Qwen2.5-Omni的多模态大型模型(MLM)。具体来说,我们将音频增强数据集与不同的情感类型与相应的驾驶行为相结合,利用嵌入在音调,音高和语音节奏中的情感线索来反映不同的用户状态,例如紧急或犹豫的意图,从而使我们的EchoVLA不仅能够解释语义内容,还能够解释音频命令的情感背景,以实现更细致入微和情感适应的驾驶行为。在开环基准测试中,与仅视觉感知的基线相比,我们的方法减少了平均L2错误$59.4\%$和碰撞率$74.4\%$。在nuScenes数据集上进行的更多实验验证了EchoVLA不仅可以通过音频指令来引导轨迹,还可以根据用户语音中检测到的情绪来调节驾驶行为。
摘要:Vision Language Action (VLA) models promise an open-vocabulary interface that can translate perceptual ambiguity into semantically grounded driving decisions, yet they still treat language as a static prior fixed at inference time. As a result, the model must infer continuously shifting objectives from pixels alone, yielding delayed or overly conservative maneuvers. We argue that effective VLAs for autonomous driving need an online channel in which users can influence driving with specific intentions. To this end, we present EchoVLA, a user-aware VLA that couples camera streams with in situ audio instructions. We augment the nuScenes dataset with temporally aligned, intent-specific speech commands generated by converting ego-motion descriptions into synthetic audios. Further, we compose emotional speech-trajectory pairs into a multimodal Chain-of-Thought (CoT) for fine-tuning a Multimodal Large Model (MLM) based on Qwen2.5-Omni. Specifically, we synthesize the audio-augmented dataset with different emotion types paired with corresponding driving behaviors, leveraging the emotional cues embedded in tone, pitch, and speech tempo to reflect varying user states, such as urgent or hesitant intentions, thus enabling our EchoVLA to interpret not only the semantic content but also the emotional context of audio commands for more nuanced and emotionally adaptive driving behavior. In open-loop benchmarks, our approach reduces the average L2 error by $59.4\%$ and the collision rate by $74.4\%$ compared to the baseline of vision-only perception. More experiments on nuScenes dataset validate that EchoVLA not only steers the trajectory through audio instructions, but also modulates driving behavior in response to the emotions detected in the user's speech.
【24】Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
标题:单音音乐中基本频率和准确发声概率的轻量级自监督检测
链接:https://arxiv.org/abs/2601.11768
备注:12 pages, 6 figures, 3 tables, and an appendix, Accepted for publication at ICPRAM 2026 in Marbella, Spain, on March 2, 2026
摘要:可靠的基频(F0)和浊音估计对于神经合成是必不可少的,然而许多基音提取器依赖于大的标记语料库并且在真实的记录伪影下退化。我们提出了一个轻量级的,完全自我监督的框架,联合F 0估计和发声推理,旨在从有限的音频快速单仪器培训。使用CQT特征的转置等变学习,我们引入了EM风格的迭代重加权方案,该方案使用移位交叉熵(SCE)一致性作为可靠性信号来抑制无信息的噪声/无声帧。所得到的权重提供置信度分数,其使得能够在没有手动注释的情况下对单独的轻量级发声分类器进行伪标记。在MedleyDB上训练并在MDB-stem-synth地面实况上进行评估,我们的方法实现了具有竞争力的跨语料库性能(RPA 95.84,RCA 96.24)并展示了跨仪器泛化。
摘要:Reliable fundamental frequency (F 0) and voicing estimation is essential for neural synthesis, yet many pitch extractors depend on large labeled corpora and degrade under realistic recording artifacts. We propose a lightweight, fully self-supervised framework for joint F 0 estimation and voicing inference, designed for rapid single-instrument training from limited audio. Using transposition-equivariant learning on CQT features, we introduce an EM-style iterative reweighting scheme that uses Shift Cross-Entropy (SCE) consistency as a reliability signal to suppress uninformative noisy/unvoiced frames. The resulting weights provide confidence scores that enable pseudo-labeling for a separate lightweight voicing classifier without manual annotations. Trained on MedleyDB and evaluated on MDB-stem-synth ground truth, our method achieves competitive cross-corpus performance (RPA 95.84, RCA 96.24) and demonstrates cross-instrument generalization.
【25】Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis
标题:哈比比:奠定统一方言阿拉伯语语音合成的开源基础
链接:https://arxiv.org/abs/2601.13802
摘要:一个显着的差距仍然存在于语音合成的研究和开发阿拉伯方言,特别是从统一的建模的角度来看。尽管阿拉伯方言具有很高的实用价值,但其固有的语言复杂性,加上缺乏标准化数据、基准和评估指南,使研究人员转向更安全的领域。为了弥合这一鸿沟,我们提出了Habibi,这是一套专门和统一的文本到语音模型,利用现有的开源ASR语料库,通过语言学知识的课程学习来支持各种高到低资源的阿拉伯方言。我们的方法在生成质量上优于领先的商业服务,同时通过有效的上下文学习保持可扩展性,而不需要文本变音符号。我们致力于开源该模型,并为多方言阿拉伯语语音合成创建第一个系统基准。此外,通过确定过程中的关键挑战和建立评估标准,我们的目标是为后续研究提供坚实的基础。资源请访问https://SWivid.github.io/Habibi/。
摘要:A notable gap persists in speech synthesis research and development for Arabic dialects, particularly from a unified modeling perspective. Despite its high practical value, the inherent linguistic complexity of Arabic dialects, further compounded by a lack of standardized data, benchmarks, and evaluation guidelines, steers researchers toward safer ground. To bridge this divide, we present Habibi, a suite of specialized and unified text-to-speech models that harnesses existing open-source ASR corpora to support a wide range of high- to low-resource Arabic dialects through linguistically-informed curriculum learning. Our approach outperforms the leading commercial service in generation quality, while maintaining extensibility through effective in-context learning, without requiring text diacritization. We are committed to open-sourcing the model, along with creating the first systematic benchmark for multi-dialect Arabic speech synthesis. Furthermore, by identifying the key challenges in and establishing evaluation standards for the process, we aim to provide a solid groundwork for subsequent research. Resources at https://SWivid.github.io/Habibi/ .
【26】Performance and Complexity Trade-off Optimization of Speech Models During Training
标题:训练期间语音模型的性能和复杂性权衡优化
链接:https://arxiv.org/abs/2601.13704
摘要:在语音机器学习中,神经网络模型通常通过选择具有固定层大小和结构的架构来设计。然后训练这些模型,以最大限度地提高与任务目标一致的指标的性能。虽然总体架构通常由任务的先验知识指导,但各个层的大小通常是按顺序选择的。然而,这种方法不能保证性能和计算复杂度之间的最佳权衡;因此,通常采用诸如权重量化或模型修剪的事后方法来降低计算成本。这是因为随机梯度下降(SGD)方法只能优化可微函数,而影响计算复杂性的因素,如层大小和每秒浮点运算(FLOP/s),是不可微的,需要在训练过程中修改模型结构。我们提出了一种基于特征噪声注入的重新参数化技术,该技术能够在使用基于SGD的方法进行训练期间联合优化性能和计算复杂度。与传统的修剪方法不同,我们的方法允许模型大小动态优化的目标性能复杂性的权衡,而不依赖于启发式标准来选择删除的权重或结构。我们通过三个案例研究证明了我们的方法的有效性,包括一个合成的例子和两个实际的现实世界中的应用:语音活动检测和音频反欺骗。与我们的工作相关的代码是公开的,以鼓励进一步的研究。
摘要:In speech machine learning, neural network models are typically designed by choosing an architecture with fixed layer sizes and structure. These models are then trained to maximize performance on metrics aligned with the task's objective. While the overall architecture is usually guided by prior knowledge of the task, the sizes of individual layers are often chosen heuristically. However, this approach does not guarantee an optimal trade-off between performance and computational complexity; consequently, post hoc methods such as weight quantization or model pruning are typically employed to reduce computational cost. This occurs because stochastic gradient descent (SGD) methods can only optimize differentiable functions, while factors influencing computational complexity, such as layer sizes and floating-point operations per second (FLOP/s), are non-differentiable and require modifying the model structure during training. We propose a reparameterization technique based on feature noise injection that enables joint optimization of performance and computational complexity during training using SGD-based methods. Unlike traditional pruning methods, our approach allows the model size to be dynamically optimized for a target performance-complexity trade-off, without relying on heuristic criteria to select which weights or structures to remove. We demonstrate the effectiveness of our method through three case studies, including a synthetic example and two practical real-world applications: voice activity detection and audio anti-spoofing. The code related to our work is publicly available to encourage further research.
【27】Event Classification by Physics-informed Inpainting for Distributed Multichannel Acoustic Sensor with Partially Degraded Channels
标题:通过物理信息修复对具有部分退化通道的分布式多通道声传感器进行事件分类
链接:https://arxiv.org/abs/2601.13513
备注:Accepted to ICASSP 2026
摘要:分布式多通道声学感测(DMAS)支持大规模声音事件分类(SEC),但当许多通道降级以及测试时的传感器布局与训练布局不同时,性能会下降。我们提出了一个学习自由,物理知情的基于逆时迁移(RTM)的修复前端。在这种方法中,观察到的多通道频谱图首先使用解析格林函数在3D网格上反向传播以形成场景一致的图像,然后在对数梅尔特征提取和基于变换器的分类之前进行前向投影以重建修复的信号。我们评估的方法ESC-50与50个传感器和三种布局(圆形,线性,直角),其中每通道信噪比从-30至0分贝采样。与AST基线、缩放稀疏最大信道选择和信道交换增强相比,所提出的RTM前端在所有布局中实现了最佳或有竞争力的准确性,在直角布局上将准确性提高了13.1个点(从9.7%提高到22.8%)。相关性分析表明,空间权重对齐更强烈的SNR比通道源距离,更高的SNR权重相关性对应于更高的SEC精度。这些结果表明,一个重建,然后项目,基于物理的预处理有效地补充了DMAS布局开放配置和严重的信道退化下的学习方法。
摘要:Distributed multichannel acoustic sensing (DMAS) enables large-scale sound event classification (SEC), but performance drops when many channels are degraded and when sensor layouts at test time differ from training layouts. We propose a learning-free, physics-informed inpainting frontend based on reverse time migration (RTM). In this approach, observed multichannel spectrograms are first back-propagated on a 3D grid using an analytic Green's function to form a scene-consistent image, and then forward-projected to reconstruct inpainted signals before log-mel feature extraction and Transformer-based classification. We evaluate the method on ESC-50 with 50 sensors and three layouts (circular, linear, right-angle), where per-channel SNRs are sampled from -30 to 0 dB. Compared with an AST baseline, scaling-sparsemax channel selection, and channel-swap augmentation, the proposed RTM frontend achieves the best or competitive accuracy across all layouts, improving accuracy by 13.1 points on the right-angle layout (from 9.7% to 22.8%). Correlation analyses show that spatial weights align more strongly with SNR than with channel--source distance, and that higher SNR--weight correlation corresponds to higher SEC accuracy. These results demonstrate that a reconstruct-then-project, physics-based preprocessing effectively complements learning-only methods for DMAS under layout-open configurations and severe channel degradation.
【28】On the Relation of State Space Models and Hidden Markov Models
标题:状态空间模型与隐马尔科夫模型的关系
链接:https://arxiv.org/abs/2601.13357
摘要:状态空间模型(SSM)和隐马尔可夫模型(HMRM)是对具有潜在变量的序列数据进行建模的基础框架,广泛用于信号处理,控制理论和机器学习。尽管它们共享时间结构,但它们在潜在状态、概率假设、推理程序和训练范式的性质上存在根本差异。最近,确定性状态空间模型通过S4和Mamba等架构重新出现在自然语言处理中,提出了关于经典概率SSM,Hacker和现代神经序列模型之间关系的新问题。 在本文中,我们提出了一个统一的和系统的比较,HISTORY,线性高斯状态空间模型,卡尔曼滤波,和当代NLP状态空间模型。我们通过概率图形模型的镜头分析它们的配方,检查它们的推理算法-包括向前向后推理和卡尔曼滤波-并通过期望最大化和基于梯度的优化来对比它们的学习过程。通过突出结构相似性和语义差异,我们澄清了这些模型何时是等效的,何时从根本上分歧,以及现代NLP SSM如何与经典概率模型相关。我们的分析将控制理论、概率建模和现代深度学习的观点联系起来。
摘要:State Space Models (SSMs) and Hidden Markov Models (HMMs) are foundational frameworks for modeling sequential data with latent variables and are widely used in signal processing, control theory, and machine learning. Despite their shared temporal structure, they differ fundamentally in the nature of their latent states, probabilistic assumptions, inference procedures, and training paradigms. Recently, deterministic state space models have re-emerged in natural language processing through architectures such as S4 and Mamba, raising new questions about the relationship between classical probabilistic SSMs, HMMs, and modern neural sequence models. In this paper, we present a unified and systematic comparison of HMMs, linear Gaussian state space models, Kalman filtering, and contemporary NLP state space models. We analyze their formulations through the lens of probabilistic graphical models, examine their inference algorithms -- including forward-backward inference and Kalman filtering -- and contrast their learning procedures via Expectation-Maximization and gradient-based optimization. By highlighting both structural similarities and semantic differences, we clarify when these models are equivalent, when they fundamentally diverge, and how modern NLP SSMs relate to classical probabilistic models. Our analysis bridges perspectives from control theory, probabilistic modeling, and modern deep learning.
【29】UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
标题:XX:解开高度相关的歌唱声音混合物
链接:https://arxiv.org/abs/2601.12802
备注:Accepted by ICASSP 2026
摘要:我们介绍了一个新的框架,多个歌声分离(MSVS)的MXXX。虽然与语音分离相关,但MSVS面临着独特的挑战:数据稀缺和歌声混合的高度相关性。为了解决这些问题,我们提出了具有三个关键组成部分的混合策略:(1)音乐信息混合策略,以构建高度相关的,音乐般的混合物,(2)交叉源注意,通过反向注意驱动两个歌手的表示,以及(3)幅度惩罚损失惩罚错误分配的干扰能量。RISTOXX不仅通过模拟真实的训练数据来解决数据稀缺问题,而且还擅长通过架构和损失级别的跨源交互来分离高度相关的混合物。我们广泛的实验表明,SDRXX大大提高了性能,SDRi增益超过2.2 dB,比以前的工作。
摘要:We introduce UNMIXX, a novel framework for multiple singing voices separation (MSVS). While related to speech separation, MSVS faces unique challenges: data scarcity and the highly correlated nature of singing voices mixture. To address these issues, we propose UNMIXX with three key components: (1) musically informed mixing strategy to construct highly correlated, music-like mixtures, (2) cross-source attention that drives representations of two singers apart via reverse attention, and (3) magnitude penalty loss penalizing erroneously assigned interfering energy. UNMIXX not only addresses data scarcity by simulating realistic training data, but also excels at separating highly correlated mixtures through cross-source interactions at both the architectural and loss levels. Our extensive experiments demonstrate that UNMIXX greatly enhances performance, with SDRi gains exceeding 2.2 dB over prior work.
【30】Toward Faithful Explanations in Acoustic Anomaly Detection
标题:声学异常检测中的忠实解释
链接:https://arxiv.org/abs/2601.12660
备注:Accepted at the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026. Code: https://github.com/Maab-Nimir/Faithful-Explanations-in-Acoustic-Anomaly-Detection
摘要:可解释性是用户在现实世界中的异常检测应用程序的信任至关重要。然而,深度学习模型尽管性能强大,但往往缺乏透明度。在这项工作中,我们研究了基于自动编码器的音频异常检测模型的可解释性,通过比较标准自动编码器(AE)与掩码自动编码器(MAE)的检测性能和可解释性。我们应用了几种归因方法,包括错误图,显着图,SmoothGrad,集成的一致性,GradSHAP和Grad-CAM。虽然MAE显示出略低的检测,但它始终提供更忠实和时间上更精确的解释,表明与真实异常更好地对齐。为了评估解释方法突出显示的区域的相关性,我们提出了一个基于扰动的忠实性度量,用它们的重建来代替它们以模拟正常输入。我们的研究结果基于真实工业场景中的实验,强调了将可解释性纳入异常检测管道的重要性,并表明掩蔽训练在不影响性能的情况下提高了解释质量。
摘要:Interpretability is essential for user trust in real-world anomaly detection applications. However, deep learning models, despite their strong performance, often lack transparency. In this work, we study the interpretability of autoencoder-based models for audio anomaly detection, by comparing a standard autoencoder (AE) with a mask autoencoder (MAE) in terms of detection performance and interpretability. We applied several attribution methods, including error maps, saliency maps, SmoothGrad, Integrated Gradients, GradSHAP, and Grad-CAM. Although MAE shows a slightly lower detection, it consistently provides more faithful and temporally precise explanations, suggesting a better alignment with true anomalies. To assess the relevance of the regions highlighted by the explanation method, we propose a perturbation-based faithfulness metric that replaces them with their reconstructions to simulate normal input. Our findings, based on experiments in a real industrial scenario, highlight the importance of incorporating interpretability into anomaly detection pipelines and show that masked training improves explanation quality without compromising performance.
【31】SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition
标题:SSVD-O:采用结构化MVD进行参数高效微调,用于语音识别
链接:https://arxiv.org/abs/2601.12600
备注:Accepted by IEEE ICASSP 2026
摘要:参数有效的微调(PEFT)是一种可扩展的方法,使大型语音基础模型适应新的领域。虽然LoRA及其最先进的变体等方法降低了自适应成本,但它们通常在模型子空间中均匀分配参数,这限制了它们在语音应用中的效率和可扩展性。在我们先前工作的基础上,本文介绍了结构SVD引导(SSVD)微调方法的扩展SSVD-Outer(SSVD-O)。SSVD-O将输入声学特征空间关联的内部变换与输出语义特征空间关联的外部变换相结合,以实现可扩展和平衡的自适应。我们进行了第一次系统的分析,参数预算分配跨模型子空间PEFT自动语音识别(ASR),并研究在有限的资源下学习和遗忘之间的权衡。SSVD-O在ESPnet框架内的0.1B到2B的模型尺度上,针对LoRA、DoRA、PiSSA和SSVD进行了领域转移ASR任务的基准测试,包括儿童语音和区域口音。实验结果表明,SSVD-O始终缩小了性能差距,完全微调,同时提高泛化能力和减轻灾难性遗忘。
摘要:Parameter-efficient fine-tuning (PEFT) is a scalable approach for adapting large speech foundation models to new domains. While methods such as LoRA and its state-of-the-art variants reduce adaptation costs, they typically allocate parameters uniformly across model subspaces, which limits their efficiency and scalability in speech applications. Building on our prior work, this paper introduces SSVD-Outer (SSVD-O), an extension of the structured SVD-guided (SSVD) fine-tuning method. SSVD-O combines input acoustic feature space-associated inner transformations with output semantic feature space-associated outer transformations to enable scalable and balanced adaptation. We conduct the first systematic analysis of parameter budget allocation across model subspaces in PEFT for automatic speech recognition (ASR), and investigate the trade-off between learning and forgetting under constrained resources. SSVD-O is benchmarked against LoRA, DoRA, PiSSA, and SSVD on domain-shifted ASR tasks, including child speech and regional accents, across model scales from 0.1B to 2B within the ESPnet framework. Experimental results show that SSVD-O consistently narrows the performance gap to full fine-tuning while improving generalization and mitigating catastrophic forgetting.
【32】SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
链接:https://arxiv.org/abs/2601.12591
备注:5 pages, accepted by ICASSP 2026
摘要:人类情感的模糊性给机器学习模型带来了一些挑战,因为它们经常重叠,缺乏清晰的边界。对比语言-音频预训练(CLAP)已成为可泛化情感识别的关键技术。然而,由于传统的CLAP在成对的音频文本样本之间强制执行严格的一对一对齐,因此它忽略了模态内相似性,并将所有不匹配的对视为同样的否定。这与不同情感之间的模糊界限相冲突。为了解决这一限制,我们提出了SmoothCLAP,它引入了来自模态内相似性和非语言特征的软化目标。通过将这些软化的目标与传统的对比监督相结合,SmoothCLAP学习尊重分级情感关系的嵌入,同时保留与CLAP相同的推理管道。跨英语和德语的八个情感计算任务的实验表明,SmoothCLAP始终实现卓越的性能。我们的研究结果强调,利用软监督是一个很有前途的策略,建立情感感知的音频文本模型。
摘要:The ambiguity of human emotions poses several challenges for machine learning models, as they often overlap and lack clear delineating boundaries. Contrastive language-audio pretraining (CLAP) has emerged as a key technique for generalisable emotion recognition. However, as conventional CLAP enforces a strict one-to-one alignment between paired audio-text samples, it overlooks intra-modal similarity and treats all non-matching pairs as equally negative. This conflicts with the fuzzy boundaries between different emotions. To address this limitation, we propose SmoothCLAP, which introduces softened targets derived from intra-modal similarity and paralinguistic features. By combining these softened targets with conventional contrastive supervision, SmoothCLAP learns embeddings that respect graded emotional relationships, while retaining the same inference pipeline as CLAP. Experiments on eight affective computing tasks across English and German demonstrate that SmoothCLAP is consistently achieving superior performance. Our results highlight that leveraging soft supervision is a promising strategy for building emotion-aware audio-text models.
【33】Harmonizing the Arabic Audio Space with Data Scheduling
标题:协调阿拉伯语音频空间与数据调度
链接:https://arxiv.org/abs/2601.12494
备注:Foundation Models, Large Language Models, Native, Speech Models, Arabic
摘要:音频大语言模型(LLM)可以实现统一的语音理解和生成,但它们对语言复杂、方言丰富的环境的适应性仍然有待探索。本文提出了第一个系统的研究多任务指令调谐为阿拉伯语为中心的音频LLM,涵盖了层次结构的生成任务(ASR,语音摘要)和判别任务(方言和情感识别)。为了支持这项研究,我们介绍了AraMega-SSum,一种用于阿拉伯语语音摘要的新型数据集。我们对Qwen2.5-Omni(7 B)进行了微调,并提出了任务渐进式课程(TPC)以及基于对齐器的多样化采样(ADS),这是一种通过选择任务和标签平衡的示例来构建信息密集批次的策略。我们的研究结果揭示了一个关键的效率,鲁棒性权衡:虽然ADS加速了初始收敛并提高了语言学F1分数,但其固有的梯度波动性可能会在长时间训练下破坏生成解码。此外,虽然TPC稳定了核心声学映射,但它经常在下游任务中引起负迁移。我们证明了一个混合TPC+ADS策略提供了一个最佳的训练“配方”,首先建立一个强大的代表性基础,然后采用多样性意识的细化,以捕捉细粒度的细微差别。这些研究结果提供了实用的指导,在复杂的,低资源的多式联运环境中的全模型的有效适应。
摘要:Audio large language models (LLMs) enable unified speech understanding and generation, yet their adaptation to linguistically complex, dialect-rich settings remains underexplored. This paper presents the first systematic study of multi-task instruction tuning for an Arabic-centric audio LLM, covering a hierarchy of generative tasks (ASR, speech summarization) and discriminative tasks (dialect and emotion identification). To support this study, we introduce AraMega-SSum, a novel dataset for Arabic speech summarization. We fine-tune Qwen2.5-Omni (7B) and propose Task-Progressive Curriculum (TPC) along with Aligner-Based Diverse Sampling (ADS), a strategy that constructs information-dense batches by selecting task- and label-balanced examples. Our results reveal a critical efficiency, robustness trade-off: while ADS accelerates initial convergence and boosts paralinguistic F1-scores, its inherent gradient volatility can destabilize generative decoding under prolonged training. Furthermore, while the TPC stabilizes core acoustic mapping, it often induces negative transfer in downstream tasks. We demonstrate that a Hybrid TPC+ADS Strategy provides an optimal training ``recipe'', first establishing a robust representative foundation before employing diversity-aware refinement to capture fine-grained nuances. These findings offer practical guidance for the efficient adaptation of Omni-models in complex, low-resource multimodal environments.
【34】A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation
标题:用于选择性可编辑文本到语音生成的统一神经编解码语言模型
链接:https://arxiv.org/abs/2601.12480
摘要:神经编解码器语言模型通过完全模仿短语音提示的声学特征(包括音色、韵律和非语言信息)来实现令人印象深刻的zero-shot文语转换(TTS)。然而,这种整体模仿限制了他们隔离和控制个体属性的能力。在本文中,我们提出了一个统一的编解码器语言模型SpeechEdit扩展zero-shot TTS与选择性控制机制。默认情况下,SpeechEdit再现从语音提示推断的完整声学配置文件,但它有选择地仅覆盖由显式控制指令指定的属性。为了实现可控建模,SpeechEdit在我们新构建的LibriEdit数据集上进行训练,该数据集提供了从LibriHeavy派生的delta(差异感知)训练对。实验结果表明,我们的方法保持自然性和鲁棒性,同时提供灵活和本地化的控制所需的属性。音频样本可在https://speech-editing.github.io/speech-editing/上获得。
摘要:Neural codec language models achieve impressive zero-shot Text-to-Speech (TTS) by fully imitating the acoustic characteristics of a short speech prompt, including timbre, prosody, and paralinguistic information. However, such holistic imitation limits their ability to isolate and control individual attributes. In this paper, we present a unified codec language model SpeechEdit that extends zero-shot TTS with a selective control mechanism. By default, SpeechEdit reproduces the complete acoustic profile inferred from the speech prompt, but it selectively overrides only the attributes specified by explicit control instructions. To enable controllable modeling, SpeechEdit is trained on our newly constructed LibriEdit dataset, which provides delta (difference-aware) training pairs derived from LibriHeavy. Experimental results show that our approach maintains naturalness and robustness while offering flexible and localized control over desired attributes. Audio samples are available at https://speech-editing.github.io/speech-editing/.
【35】ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
标题:ParaMETA:学习从言语中分离出副语言说话风格的表达
链接:https://arxiv.org/abs/2601.12289
备注:9 pages, 7 figures, Accepted to AAAI-26 (Main Technical Track)
摘要:针对不同类型的说话风格(诸如情感、年龄和性别)学习代表性嵌入对于识别任务(例如,认知计算和人机交互)和生成任务(例如,风格可控的语音生成)。在这项工作中,我们介绍了ParaMETA,一个统一的和灵活的框架,直接从语音学习和控制说话风格。与依赖于单任务模型或跨模态对齐的现有方法不同,ParaMETA通过将语音投影到每种风格的专用子空间中来学习分解的特定于任务的嵌入。这种设计减少了任务间干扰,减轻了负迁移,并允许单个模型处理多个语言任务,如情感,性别,年龄和语言分类。除了识别之外,ParaMETA还可以在文本到语音(TTS)生成模型中实现细粒度的风格控制。它支持语音和文本提示,并允许用户修改一种说话风格,同时保留其他风格。大量的实验表明,ParaMETA在分类准确性方面优于强基线,并生成更自然和更有表现力的语音,同时保持适合现实世界应用的轻量级和高效的模型。
摘要:Learning representative embeddings for different types of speaking styles, such as emotion, age, and gender, is critical for both recognition tasks (e.g., cognitive computing and human-computer interaction) and generative tasks (e.g., style-controllable speech generation). In this work, we introduce ParaMETA, a unified and flexible framework for learning and controlling speaking styles directly from speech. Unlike existing methods that rely on single-task models or cross-modal alignment, ParaMETA learns disentangled, task-specific embeddings by projecting speech into dedicated subspaces for each type of style. This design reduces inter-task interference, mitigates negative transfer, and allows a single model to handle multiple paralinguistic tasks such as emotion, gender, age, and language classification. Beyond recognition, ParaMETA enables fine-grained style control in Text-To-Speech (TTS) generative models. It supports both speech- and text-based prompting and allows users to modify one speaking styles while preserving others. Extensive experiments demonstrate that ParaMETA outperforms strong baselines in classification accuracy and generates more natural and expressive speech, while maintaining a lightweight and efficient model suitable for real-world applications.
【36】Confidence-based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens
标题:基于置信度的语音数据集处理过滤,并使用离散令牌进行生成语音增强
链接:https://arxiv.org/abs/2601.12254
备注:Accepted for ICASSP 2026
摘要:生成式语音增强(GSE)模型在从噪声输入中生成高质量的干净语音方面表现出很大的潜力,从而实现了将噪声文本到语音(TTS)数据集转化为高质量数据集等应用。然而,GSE模型容易产生幻觉错误,如音素遗漏和说话人不一致,传统的基于非侵入性语音质量度量的错误过滤往往无法检测到。为了解决这个问题,我们提出了一种非侵入性的方法来过滤幻觉错误的离散令牌为基础的GSE模型。我们的方法利用生成的令牌的对数概率作为置信度分数来检测潜在的错误。实验结果表明,置信度分数与一套侵入性SE度量密切相关,并且我们的方法有效地识别了传统过滤方法遗漏的幻觉错误。此外,我们证明了我们的方法的实际效用:用我们基于置信度的过滤来管理野生TTS数据集,提高了随后训练的TTS模型的性能。
摘要:Generative speech enhancement (GSE) models show great promise in producing high-quality clean speech from noisy inputs, enabling applications such as curating noisy text-to-speech (TTS) datasets into high-quality ones. However, GSE models are prone to hallucination errors, such as phoneme omissions and speaker inconsistency, which conventional error filtering based on non-intrusive speech quality metrics often fails to detect. To address this issue, we propose a non-intrusive method for filtering hallucination errors from discrete token-based GSE models. Our method leverages the log-probabilities of generated tokens as confidence scores to detect potential errors. Experimental results show that the confidence scores strongly correlate with a suite of intrusive SE metrics, and that our method effectively identifies hallucination errors missed by conventional filtering methods. Furthermore, we demonstrate the practical utility of our method: curating an in-the-wild TTS dataset with our confidence-based filtering improves the performance of subsequently trained TTS models.
【37】Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
标题:Sound 2 Hap:从人类评级中学习音频到振动触觉生成
链接:https://arxiv.org/abs/2601.12245
摘要:环境声音(如脚步声、键盘敲击声或狗吠声)携带着丰富的信息和情感背景,这使得它们对用户应用程序中的触觉设计很有价值。然而,现有的音频到振动方法依赖于针对音乐或游戏调整的信号处理规则,并且通常无法在不同的声音中推广。为了解决这个问题,我们首先研究了用户对四种现有音频到触觉算法的感知,然后创建了一个环境声音的数据驱动模型。在研究1中,34名参与者对四种算法产生的1,000种声音的振动进行了评级,没有发现一致的算法偏好。使用这个数据集,我们训练了基于CNN的自动编码器Sound 2 Hap,以低延迟从不同的声音中生成有感知意义的振动。在研究2中,15名参与者在音频振动匹配和触觉体验指数(HXI)方面的输出高于信号处理基线,发现它与不同的声音更和谐。这项工作展示了一个感知验证的方法来音频触觉翻译,扩大了声音驱动的触觉的范围。
摘要:Environmental sounds like footsteps, keyboard typing, or dog barking carry rich information and emotional context, making them valuable for designing haptics in user applications. Existing audio-to-vibration methods, however, rely on signal-processing rules tuned for music or games and often fail to generalize across diverse sounds. To address this, we first investigated user perception of four existing audio-to-haptic algorithms, then created a data-driven model for environmental sounds. In Study 1, 34 participants rated vibrations generated by the four algorithms for 1,000 sounds, revealing no consistent algorithm preferences. Using this dataset, we trained Sound2Hap, a CNN-based autoencoder, to generate perceptually meaningful vibrations from diverse sounds with low latency. In Study 2, 15 participants rated its output higher than signal-processing baselines on both audio-vibration match and Haptic Experience Index (HXI), finding it more harmonious with diverse sounds. This work demonstrates a perceptually validated approach to audio-haptic translation, broadening the reach of sound-driven haptics.
【38】Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
标题:多干注意力和层次不确定性建模下的歌曲美学评价
链接:https://arxiv.org/abs/2601.12222
摘要:音乐生成人工智能(AI)正在迅速扩展音乐内容,需要自动化的歌曲美学评估。然而,现有的研究主要集中在语音,音频或演唱质量,留下歌曲美学探索不足。此外,传统的方法往往预测一个精确的平均意见得分(MOS)值,这很难捕捉人类感知的细微差别,在歌曲美学评价。本文提出了一个面向歌曲的美学评价框架,包括两个新的模块:1)多干注意力融合(MSAF)在混音-人声和混音-伴奏对之间建立双向交叉注意,融合它们以捕捉复杂的音乐特征; 2)分级粒度感知区间聚合(HiGIA)学习多粒度分数概率分布,将它们聚合到分数区间中,并在该区间内应用回归以产生最终分数。我们对两个全长歌曲数据集进行了评估:SongEval数据集(AI生成)和内部美学数据集(人类创建),并与两个最先进的(SOTA)模型进行了比较。实验结果表明,该方法对歌曲的多维美学评价具有较强的性能。
摘要:Music generative artificial intelligence (AI) is rapidly expanding music content, necessitating automated song aesthetics evaluation. However, existing studies largely focus on speech, audio or singing quality, leaving song aesthetics underexplored. Moreover, conventional approaches often predict a precise Mean Opinion Score (MOS) value directly, which struggles to capture the nuances of human perception in song aesthetics evaluation. This paper proposes a song-oriented aesthetics evaluation framework, featuring two novel modules: 1) Multi-Stem Attention Fusion (MSAF) builds bidirectional cross-attention between mixture-vocal and mixture-accompaniment pairs, fusing them to capture complex musical features; 2) Hierarchical Granularity-Aware Interval Aggregation (HiGIA) learns multi-granularity score probability distributions, aggregates them into a score interval, and applies a regression within the interval to produce the final score. We evaluated on two datasets of full-length songs: SongEval dataset (AI-generated) and an internal aesthetics dataset (human-created), and compared with two state-of-the-art (SOTA) models. Results show that the proposed method achieves stronger performance for multi-dimensional song aesthetics evaluation.
【39】Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
标题:神经编解码器可以通用吗?一项针对不可见语言和非言语任务的对照研究
链接:https://arxiv.org/abs/2601.12205
摘要:本文研究了神经音频编解码器(NAC)泛化能力的三个关键但未充分探索的方面:(i)NAC是否可以在预训练期间推广到看不见的语言,(ii)仅语音预训练的NAC是否可以有效地推广到非语音应用,诸如环境声音、音乐和动物发声,以及(iii)在预训练期间结合非语音数据是否可以提高语音和非语音任务两者的性能。现有的研究通常依赖于现成的NAC进行比较,由于实施的差异,这限制了洞察力。在这项工作中,我们使用严格控制的配置和精心策划的预训练数据从头开始训练NAC,以实现公平的比较。我们使用11个指标对NAC在信号重建质量和下游应用方面的性能进行了全面评估。我们的研究结果表明,NAC可以在预训练期间推广到看不见的语言,仅语音预训练的NAC在非语音任务上表现出性能下降,并且在预训练期间合并非语音数据可以提高非语音任务的性能,同时保持语音任务的性能相当。
摘要:This paper investigates three crucial yet underexplored aspects of the generalization capabilities of neural audio codecs (NACs): (i) whether NACs can generalize to unseen languages during pre-training, (ii) whether speech-only pre-trained NACs can effectively generalize to non-speech applications such as environmental sounds, music, and animal vocalizations, and (iii) whether incorporating non-speech data during pre-training can improve performance on both speech and non-speech tasks. Existing studies typically rely on off-the-shelf NACs for comparison, which limits insight due to variations in implementation. In this work, we train NACs from scratch using strictly controlled configurations and carefully curated pre-training data to enable fair comparisons. We conduct a comprehensive evaluation of NAC performance on both signal reconstruction quality and downstream applications using 11 metrics. Our results show that NACs can generalize to unseen languages during pre-training, speech-only pre-trained NACs exhibit degraded performance on non-speech tasks, and incorporating non-speech data during pre-training improves performance on non-speech tasks while maintaining comparable performance on speech tasks.
【40】VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
标题:VidButton:使用生成性音乐和上下文缩略图创建视频配乐
链接:https://arxiv.org/abs/2601.12180
备注:Accepted to CHI 2026
摘要:音乐塑造了视频的基调,但创作者往往很难找到与视频的情绪和叙事相匹配的配乐。最近的文本到音乐模型让创作者从文本提示生成音乐,但我们的形成性研究(N=8)显示创作者很难构建不同的提示,快速审查和比较曲目,并了解它们对视频的影响。我们提出了VidTune,一个系统,支持配乐创作,从创作者的提示生成不同的音乐选项,并产生快速审查的上下文缩略图。VidTune提取代表性的视频主题,以背景中的缩略图为基础,将每个轨道的效价和能量映射到颜色和亮度等视觉线索上,并描绘突出的流派和乐器。创作者可以通过自然语言编辑来优化曲目,VidTune将其扩展到新一代。在一项对照用户研究(N=12)和一项探索性案例研究(N=6)中,参与者发现VidTune有助于有效地审查和比较音乐选项,并将此过程描述为有趣和丰富。
摘要:Music shapes the tone of videos, yet creators often struggle to find soundtracks that match their video's mood and narrative. Recent text-to-music models let creators generate music from text prompts, but our formative study (N=8) shows creators struggle to construct diverse prompts, quickly review and compare tracks, and understand their impact on the video. We present VidTune, a system that supports soundtrack creation by generating diverse music options from a creator's prompt and producing contextual thumbnails for rapid review. VidTune extracts representative video subjects to ground thumbnails in context, maps each track's valence and energy onto visual cues like color and brightness, and depicts prominent genres and instruments. Creators can refine tracks through natural language edits, which VidTune expands into new generations. In a controlled user study (N=12) and an exploratory case study (N=6), participants found VidTune helpful for efficiently reviewing and comparing music options and described the process as playful and enriching.
【41】MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
标题:MuseAgent-1:音乐配乐和表演音频的交互式接地多模式理解
链接:https://arxiv.org/abs/2601.11968
备注:Tech Report
摘要:尽管最近在多模态大型语言模型(MLLM)方面取得了进展,但它们理解音乐和与音乐交互的能力仍然有限。音乐理解需要对符号分数和表现力表现音频进行接地推理,由于感知接地不足,通用MLLM通常无法处理。我们介绍MuseAgent,一个以音乐为中心的多模态代理,增强语言模型与结构化的符号表示来自乐谱图像和性能音频。通过集成光学音乐识别和自动音乐转录模块,MuseAgent可以对细粒度的音乐内容进行多步推理和交互。为了系统地评估音乐理解能力,我们进一步提出了MuseBench,这是一个涵盖音乐理论推理,乐谱解释和跨文本,图像和音频形式的性能级别分析的基准。实验表明,现有的MLLM在这些任务上表现不佳,而MuseAgent实现了实质性的改进,突出了交互式音乐理解的结构化多模态接地的重要性。
摘要:Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reasoning over symbolic scores and expressive performance audio, which general-purpose MLLMs often fail to handle due to insufficient perceptual grounding. We introduce MuseAgent, a music-centric multimodal agent that augments language models with structured symbolic representations derived from sheet music images and performance audio. By integrating optical music recognition and automatic music transcription modules, MuseAgent enables multi-step reasoning and interaction over fine-grained musical content. To systematically evaluate music understanding capabilities, we further propose MuseBench, a benchmark covering music theory reasoning, score interpretation, and performance-level analysis across text, image, and audio modalities. Experiments show that existing MLLMs perform poorly on these tasks, while MuseAgent achieves substantial improvements, highlighting the importance of structured multimodal grounding for interactive music understanding.
【42】The Third VoicePrivacy Challenge: Preserving Emotional Expressiveness and Linguistic Content in Voice Anonymization
标题:第三个语音隐私挑战:在语音分析中保留情感表达和语言内容
链接:https://arxiv.org/abs/2601.11846
备注:under review
摘要:我们介绍了2024年举行的第三届语音隐私挑战赛的结果和分析,该挑战赛的重点是推进语音匿名化技术。这项任务是开发一个语音匿名系统的语音数据,隐藏说话者的语音身份,同时保留语言内容和情感状态。我们提供了挑战框架的系统概述,包括匿名化任务和用于系统开发和评估的数据集的详细描述。我们概述了攻击模型和评估隐私保护(隐藏扬声器的声音身份)和实用程序(内容和情绪状态保存)的客观评估指标。我们描述了六个基线匿名化系统,并总结了挑战参与者开发的创新方法。最后,我们提供了关键的见解和意见,以指导未来的VoicePrivacy挑战的设计,并确定有前途的语音匿名化研究方向。
摘要:We present results and analyses from the third VoicePrivacy Challenge held in 2024, which focuses on advancing voice anonymization technologies. The task was to develop a voice anonymization system for speech data that conceals a speaker's voice identity while preserving linguistic content and emotional state. We provide a systematic overview of the challenge framework, including detailed descriptions of the anonymization task and datasets used for both system development and evaluation. We outline the attack model and objective evaluation metrics for assessing privacy protection (concealing speaker voice identity) and utility (content and emotional state preservation). We describe six baseline anonymization systems and summarize the innovative approaches developed by challenge participants. Finally, we provide key insights and observations to guide the design of future VoicePrivacy challenges and identify promising directions for voice anonymization research.
【43】CSyMR: Benchmarking Compositional Symbolic Muisc Reasoning With MIR Tool Integration
标题:CSyMR:通过MIR工具集成对合成符号Muisc推理进行基准测试
链接:https://arxiv.org/abs/2601.11556
摘要:大型语言模型(LLM)被用于符号音乐推理,但现有的基准强调孤立的知识或原子分析,而不是连接音乐结构所需的综合成分推理。为了解决这个问题,我们提出了作曲符号音乐推理基准(CSyMR-Bench),这是一个来自专家论坛和专业考试的126个问题的精选多项选择数据集。每一个问题都需要结合几个原子分析来得出最终答案。此外,我们引入了一个工具增强的代理框架,利用符号音乐分析工具的music 21库,以解决CSyMR-Bench所带来的挑战。实验验证了CSyMR-Bench在社区来源和考试风格的问题上都提出了不小的挑战,而我们的工具增强代理始终优于所有基线,实现了5-7%的绝对准确率增益。
摘要:Large Language Models (LLMs) are leveraged in symbolic music reasoning, yet existing benchmarks emphasize isolated knowledge or atomic analyses rather than the integrative compositional reasoning needed to connect musical structures. To address this, we present the Compositional Symbolic Music Reasoning Benchmark (CSyMR-Bench), a curated multiple-choice dataset of 126 questions derived from expert forums and professional examinations. Each item involves combining several atomic analyses to arrive at the final answer. Furthermore, we introduce a tool-augmented agent framework that leverages symbolic music analysis tools from the music21 library to address the challenges posed by CSyMR-Bench. Experiments validate that CSyMR-Bench poses a non-trivial challenge across both community-sourced and exam-style questions, while our tool-augmented agent consistently outperforms all baselines, achieving 5-7% absolute accuracy gains.
【44】WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
标题:WenetSpeech-Wu:统一汉语吴方言语音处理生态系统的数据集、基准和模型
链接:https://arxiv.org/abs/2601.11027
摘要:低资源方言的语音处理仍然是开发包容性和鲁棒性语音技术的基本挑战。尽管吴方言在语言学上具有重要意义,使用者众多,但由于缺乏大规模的语音数据、标准化的评估基准和公开可用的模型,吴方言的研究长期受到阻碍。在这项工作中,我们提出了WenetSpeech-Wu,第一个大规模的,多维注释的开源语音语料库的吴方言,包括大约8,000小时的不同的语音数据。在此数据集的基础上,我们介绍了WenetSpeech-Wu-Bench,这是第一个标准化和公开访问的基准,用于系统评估吴方言语音处理,涵盖自动语音识别(ASR),吴到普通话翻译,说话人属性预测,语音情感识别,文本到语音(TTS)合成,以及解释遵循TTS(指令TTS)。此外,我们发布了一套在WenetSpeech-Wu上训练的强大开源模型,在多个任务中建立了具有竞争力的性能,并通过经验验证了所提出的数据集的有效性。总之,这些贡献为全面的吴方言语音处理生态系统奠定了基础,我们开源了建议的数据集,基准和模型,以支持未来对方言语音智能的研究。
摘要:Speech processing for low-resource dialects remains a fundamental challenge in developing inclusive and robust speech technologies. Despite its linguistic significance and large speaker population, the Wu dialect of Chinese has long been hindered by the lack of large-scale speech data, standardized evaluation benchmarks, and publicly available models. In this work, we present WenetSpeech-Wu, the first large-scale, multi-dimensionally annotated open-source speech corpus for the Wu dialect, comprising approximately 8,000 hours of diverse speech data. Building upon this dataset, we introduce WenetSpeech-Wu-Bench, the first standardized and publicly accessible benchmark for systematic evaluation of Wu dialect speech processing, covering automatic speech recognition (ASR), Wu-to-Mandarin translation, speaker attribute prediction, speech emotion recognition, text-to-speech (TTS) synthesis, and instruction-following TTS (instruct TTS). Furthermore, we release a suite of strong open-source models trained on WenetSpeech-Wu, establishing competitive performance across multiple tasks and empirically validating the effectiveness of the proposed dataset. Together, these contributions lay the foundation for a comprehensive Wu dialect speech processing ecosystem, and we open-source proposed datasets, benchmarks, and models to support future research on dialectal speech intelligence.
机器翻译由腾讯交互翻译提供,仅供参考
