今日论文合集:cs.SD语音16篇,eess.AS音频处理8篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Beyond Transcripts: A Renewed Perspective on Audio Chaptering
标题:超越文字记录:音频章节的新视角
链接:https://arxiv.org/pdf/2602.08979v1

作者:Fabian Retkowski,Maike Züfle,Thai Binh Nguyen,Jan Niehues,Alexander Waibel
摘要:音频分节,自动将长格式音频分割成连贯的部分,对于导航播客,讲座和视频越来越重要。尽管其相关性,研究仍然有限,基于文本,留下了关于利用音频信息,处理ASR错误和无成绩单评估的关键问题尚未解决。我们通过三个贡献来解决这些差距:(1)系统地比较基于文本的模型与声学特征,一种新颖的纯音频架构(2)影响表现的因素的经验分析,包括转录质量、声学特征、持续时间和说话者组成;以及(3)形式化的评估协议,将转录依赖的文本空间协议与转录不变的时间空间协议进行对比。我们对YTSeg的实验表明,AudioSeg大大优于基于文本的方法,停顿提供了最大的声学增益,MLLM仍然受到上下文长度和弱指令跟随的限制,但MLLM在较短的音频上很有希望。摘要:Audio chaptering, the task of automatically segmenting long-form audio into coherent sections, is increasingly important for navigating podcasts, lectures, and videos. Despite its relevance, research remains limited and text-based, leaving key questions unresolved about leveraging audio information, handling ASR errors, and transcript-free evaluation. We address these gaps through three contributions: (1) a systematic comparison between text-based models with acoustic features, a novel audio-only architecture (AudioSeg) operating on learned audio representations, and multimodal LLMs; (2) empirical analysis of factors affecting performance, including transcript quality, acoustic features, duration, and speaker composition; and (3) formalized evaluation protocols contrasting transcript-dependent text-space protocols with transcript-invariant time-space protocols. Our experiments on YTSeg reveal that AudioSeg substantially outperforms text-based approaches, pauses provide the largest acoustic gains, and MLLMs remain limited by context length and weak instruction following, yet MLLMs are promising on shorter audio.


【2】No Word Left Behind: Mitigating Prefix Bias in Open-Vocabulary Keyword Spotting
标题:不留任何单词:缓解开放词汇关键词发现中的后缀偏差
链接:https://arxiv.org/pdf/2602.08930v1

作者:Yi Liu,Chuan-Che,Huang,Xiao Quan

备注:Published in ICASSP 2026

摘要:开放词汇关键词识别(OV-KWS)可通过任意语音命令实现个性化设备控制。最近,研究人员探索使用音频文本联合嵌入,允许用户注册短语与文本,并提出了消除类似话语歧义的技术。我们发现,现有的OV-KWS解决方案往往过于偏向注册的开始音素,导致错误的触发时,负注册查询对共享一个前缀("把音量调高“与"把音量调低”)。我们将其归因于两个因素:训练数据偏差和位置偏差的跨模态评分。为了解决这些限制,我们引入了部分重叠基准(POB)与两个数据集,POB-Spark和POB-LibriPhrase(POB-LP),包含不匹配的音频文本对与共享前缀,并提出了一个轻量级的决策层,加权位置评分(EPS)。单独使用EPS将POB-Spark上的EER从64.4 %降低到29.3 %,并将POB-LP准确率从87.6 %提高到96.8 %,同时保持LibriPhrase和Google Speech Commands(GSC)的性能。通过在训练中添加POB数据,我们的工作实现了最佳的POB基准测试结果,同时在基线之间的先前指标上产生了最少的退化。这种退化在GSC中最为明显,它只包含一个单词的命令。我们表面减轻这种权衡作为未来的工作。摘要:Open-vocabulary keyword spotting (OV-KWS) enables personalized device control via arbitrary voice commands. Recently, researchers have explored using audio-text joint embeddings, allowing users to enroll phrases with text, and proposed techniques to disambiguate similar utterances. We find that existing OV-KWS solutions often overly bias the beginning phonemes of an enrollment, causing false triggers when negative enrollment-query-pairs share a prefix ( turn the volume up'' vs. turn the volume down''). We trace this to two factors: training data bias and position-biased cross-modal scoring. To address these limitations, we introduce the Partial Overlap Benchmark (POB) with two datasets, POB-Spark and POB-LibriPhrase (POB-LP), containing mismatched audio-text pairs with shared prefixes, and propose Equal-weighting Position Scoring (EPS), a lightweight decision layer. Using EPS alone reduces EER on POB-Spark from 64.4 % to 29.3 % and improves POB-LP accuracy from 87.6 % to 96.8 %, while maintaining performance on LibriPhrase and Google Speech Commands (GSC). With POB data added in training, our work achieves the best POB benchmark results while incurring the least amount of degradation on prior metrics among baselines. This degradation is most pronounced in GSC, which contains only one-word commands. We surface mitigating this trade-off as future work.


【3】MOVA: Towards Scalable and Synchronized Video-Audio Generation
标题:MOVA:迈向可扩展和同步的视频音频生成
链接:https://arxiv.org/pdf/2602.08794v1

作者:SII-OpenMOSS Team,:,Donghua Yu,Mingshu Chen,Qi Chen,Qi Luo,Qianyi Wu,Qinyuan Cheng,Ruixiao Li,Tianyi Liang,Wenbo Zhang,Wenming Tu,Xiangyu Peng,Yang Gao,Yanru Huo,Ying Zhu,Yinze Luo,Yiyang Zhang,Yuerong Song,Zhe Xu,Zhiyu Zhang,Chenchen Yang,Cheng Chang,Chushu Zhou,Hanfu Chen,Hongnan Ma,Jiaxi Li,Jingqi Tong,Junxi Liu,Ke Chen,Shimin Li,Songlin Wang,Wei Jiang,Zhaoye Fei,Zhiyuan Ning,Chunguo Li,Chenhui Li,Ziwei He,Zengfeng Huang,Xie Chen,Xipeng Qiu

备注:Technical report for MOVA (open-source video-audio generation model). 38 pages, 10 figures, 22 tables. Project page: https:mosi.cnmodelsmova Code: https:github.comOpenMOSSMOVA Models: https:huggingface.cocollectionsOpenMOSS-Teammova. Qinyuan Cheng and Tianyi Liang are project leader. Xie Chen and Xipeng Qiu are corresponding authors

摘要:音频对于真实世界的视频是不可或缺的,但生成模型在很大程度上忽略了音频组件。当前的视听内容制作方法通常依赖于级联流水线,这增加了成本,积累了错误,并降低了整体质量。虽然Veo 3和Sora 2等系统强调了同时生成的价值,但联合多模态建模在架构、数据和培训方面带来了独特的挑战。此外,现有系统的封闭源性质限制了这一领域的进展。在这项工作中,我们介绍MOVA(MOSS视频和音频),一个开源模型,能够生成高质量,同步的视听内容,包括逼真的对口型语音,环境感知的声音效果,和内容对齐的音乐。MOVA采用混合专家(MoE)架构,共有32 B个参数,其中18 B在推理过程中是活动的。它支持IT 2 VA(图像-文本到视频-音频)生成任务。通过发布模型权重和代码,我们的目标是推进研究并培养一个充满活力的创作者社区。发布的代码库具有对高效推理,LoRA微调和即时增强的全面支持。摘要:Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on cascaded pipelines, which increase cost, accumulate errors, and degrade overall quality. While systems such as Veo 3 and Sora 2 emphasize the value of simultaneous generation, joint multimodal modeling introduces unique challenges in architecture, data, and training. Moreover, the closed-source nature of existing systems limits progress in the field. In this work, we introduce MOVA (MOSS Video and Audio), an open-source model capable of generating high-quality, synchronized audio-visual content, including realistic lip-synced speech, environment-aware sound effects, and content-aligned music. MOVA employs a Mixture-of-Experts (MoE) architecture, with a total of 32B parameters, of which 18B are active during inference. It supports IT2VA (Image-Text to Video-Audio) generation task. By releasing the model weights and code, we aim to advance research and foster a vibrant community of creators. The released codebase features comprehensive support for efficient inference, LoRA fine-tuning, and prompt enhancement.


【4】Prototype-Based Disentanglement for Controllable Dysarthric Speech Synthesis
标题:基于原型的可控结构障碍语音合成解纠缠
链接:https://arxiv.org/pdf/2602.08696v1

作者:Haoshen Wang,Xueli Zhong,Bingbing Lin,Jia Huang,Xingduo Pan,Shengxiang Liang,Nizhuan Wang,Wai Ting Siok
摘要:构音障碍语音具有高度的可变性和有限的标记数据,这对自动语音识别(ASR)和辅助语音技术都构成了重大挑战。现有的方法依赖于合成数据增强或语音重建,但往往纠缠扬声器身份与病理清晰度,限制可控性和鲁棒性。 在本文中,我们提出了ProtoDisent-TTS,这是一个基于原型的解纠缠TTS框架,它建立在一个预先训练的文本到语音的主干上,可以在一个统一的潜在空间内分解扬声器音色和构音障碍的清晰度。病理学原型码本提供了健康和构音障碍语音模式的可解释和可控制的表示,而具有梯度反转层的双分类器目标强制扬声器嵌入病理属性的不变性。在TORGO数据集上的实验表明,这种设计能够实现健康语音和构音障碍语音之间的双向转换,从而获得一致的ASR性能增益和鲁棒的说话者感知语音重建。摘要:Dysarthric speech exhibits high variability and limited labeled data, posing major challenges for both automatic speech recognition (ASR) and assistive speech technologies. Existing approaches rely on synthetic data augmentation or speech reconstruction, yet often entangle speaker identity with pathological articulation, limiting controllability and robustness. In this paper, we propose ProtoDisent-TTS, a prototype-based disentanglement TTS framework built on a pre-trained text-to-speech backbone that factorizes speaker timbre and dysarthric articulation within a unified latent space. A pathology prototype codebook provides interpretable and controllable representations of healthy and dysarthric speech patterns, while a dual-classifier objective with a gradient reversal layer enforces invariance of speaker embeddings to pathological attributes. Experiments on the TORGO dataset demonstrate that this design enables bidirectional transformation between healthy and dysarthric speech, leading to consistent ASR performance gains and robust, speaker-aware speech reconstruction.


【5】VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
标题:VocalNet-MDM:通过自蒸馏掩蔽扩散模型加速流式语音LLM
链接:https://arxiv.org/pdf/2602.08607v1

作者:Ziyang Cheng,Yuhao Wang,Heyang Liu,Ronghua Wu,Qunshan Gu,Yanfeng Wang,Yu Wang
摘要:最近的语音大语言模型(LLM)在端到端语音交互方面取得了令人印象深刻的能力。然而,流行的自回归范式施加了严格的序列约束,限制了生成效率并引入了暴露偏差。在本文中,我们研究了掩蔽扩散模型~(MDM)作为一种非自回归的语音LLM范例,并介绍了VocalNet-MDM。为了使MDM适应流式语音交互,我们解决了两个关键挑战:训练推理不匹配和迭代开销。我们提出了分层分块掩蔽来将训练目标与块扩散解码过程中遇到的渐进掩蔽状态对齐,并提出了迭代自蒸馏来将多步细化压缩为更少的步骤,以实现低延迟推理。VocalNet-MDM仅在6 K小时的语音数据的有限规模上进行训练,与AR基线相比,VocalNet-MDM实现了3.7$ times$--10$ times $的解码加速,并将第一块延迟降低了34%。它保持了具有竞争力的识别准确性,同时实现了最先进的文本质量和语音自然度,表明MDM是低延迟,高效语音LLM的一种有前途和可扩展的替代方案。摘要:Recent Speech Large Language Models~(LLMs) have achieved impressive capabilities in end-to-end speech interaction. However, the prevailing autoregressive paradigm imposes strict serial constraints, limiting generation efficiency and introducing exposure bias. In this paper, we investigate Masked Diffusion Modeling~(MDM) as a non-autoregressive paradigm for speech LLMs and introduce VocalNet-MDM. To adapt MDM for streaming speech interaction, we address two critical challenges: training-inference mismatch and iterative overhead. We propose Hierarchical Block-wise Masking to align training objectives with the progressive masked states encountered during block diffusion decoding, and Iterative Self-Distillation to compress multi-step refinement into fewer steps for low-latency inference. Trained on a limited scale of only 6K hours of speech data, VocalNet-MDM achieves a 3.7$ times$--10$ times$ decoding speedup and reduces first-chunk latency by 34 % compared to AR baselines. It maintains competitive recognition accuracy while achieving state-of-the-art text quality and speech naturalness, demonstrating that MDM is a promising and scalable alternative for low-latency, efficient speech LLMs.


【6】Global Rotation Equivariant Phase Modeling for Speech Enhancement with Deep Magnitude-Phase Interaction
标题:深度幅度-相相互作用语音增强的全局旋转等变相建模
链接:https://arxiv.org/pdf/2602.08556v1

作者:Chengzhong Wang,Andong Li,Dingding Yao,Junfeng Li

备注:Submitted to IEEE TASLP

摘要:虽然深度学习具有先进的语音增强(SE),但有效的相位建模仍然具有挑战性,因为传统网络通常在平坦的欧几里得特征空间内运行,这不容易对相位的底层圆形拓扑进行建模。为了解决这个问题,我们提出了一个流形感知的幅度相位双流框架,通过强制执行全局旋转等方差(GRE)特性,使相位流与其固有的圆形几何形状对齐。具体来说,我们引入了用于基于模的信息交换的幅度-相位交互式卷积模块(MPICM)和用于统一特征融合的混合注意力双FFN(HADF)瓶颈,这两者都旨在保留相位流中的GRE。通过相位恢复、去噪、去混响和带宽扩展任务进行综合评估,以验证所提出的方法在多个高级基线上的优越性。值得注意的是,所提出的架构在相位检索任务中将相位距离降低了20%以上,并且在zero-shot交叉语料库去噪评估中将PESQ提高了0.1以上。在涉及混合失真的通用SE任务中也建立了整体优势。定性分析进一步揭示了学习的相位特征表现出明显的周期性模式,这与相位的内在循环性质是一致的。源代码可在https: github.com wangchengzhong RENet上获得。摘要:While deep learning has advanced speech enhancement (SE), effective phase modeling remains challenging, as conventional networks typically operate within a flat Euclidean feature space, which is not easy to model the underlying circular topology of the phase. To address this, we propose a manifold-aware magnitude-phase dual-stream framework that aligns the phase stream with its intrinsic circular geometry by enforcing Global Rotation Equivariance (GRE) characteristic. Specifically, we introduce a Magnitude-Phase Interactive Convolutional Module (MPICM) for modulus-based information exchange and a Hybrid-Attention Dual-FFN (HADF) bottleneck for unified feature fusion, both of which are designed to preserve GRE in the phase stream. Comprehensive evaluations are conducted across phase retrieval, denoising, dereverberation, and bandwidth extension tasks to validate the superiority of the proposed method over multiple advanced baselines. Notably, the proposed architecture reduces Phase Distance by over 20 % in the phase retrieval task and improves PESQ by more than 0.1 in zero-shot cross-corpus denoising evaluations. The overall superiority is also established in universal SE tasks involving mixed distortions. Qualitative analysis further reveals that the learned phase features exhibit distinct periodic patterns, which are consistent with the intrinsic circular nature of the phase. The source code is available at https: github.com wangchengzhong RENet.


【7】PTS-SNN: A Prompt-Tuned Temporal Shift Spiking Neural Networks for Efficient Speech Emotion Recognition
标题:PTS-SNN:一种用于高效语音情感识别的预算调谐时移尖峰神经网络
链接:https://arxiv.org/pdf/2602.08240v1

作者:Xun Su,Huamin Wang,Qi Zhang
摘要:语音情感识别(SER)广泛应用于人机交互中,但传统模型的高计算成本阻碍了其在资源受限的边缘设备上的实现。尖峰神经网络(SNN)由于其事件驱动的性质而提供了一种节能的替代方案;然而,它们与连续自监督学习(SSL)表示的集成从根本上受到分布失配的挑战,其中高动态范围嵌入降低了基于阈值的神经元的信息编码能力。为了解决这个问题,我们提出了一个参数有效的神经形态适应框架,它将冻结的SSL主干与尖峰动态对齐。具体来说,我们引入了一个时间移位尖峰编码器,通过无参数通道移位来捕获局部时间依赖性,建立一个稳定的特征基础。为了弥合域的差距,我们设计了一个上下文感知的膜电位校准策略。该机制利用Spiking Sparse Linear Attention模块将全局语义上下文聚合到可学习的软提示中,该软提示动态调节参数泄漏积分和激发(PLIF)神经元的偏置电压。这种调节有效地将异质输入分布集中在响应性激发范围内,减轻功能沉默或饱和。在五个多语言数据集上进行了广泛的实验(例如,IEMOCAP,CASIA,EMODB)证明,PTS-SNN在IEMOCAP上达到73.34%的准确度,与竞争对手的人工神经网络(ANN)相当,同时每个样本仅需要1.19M的可训练参数和0.35 mJ的推理能量。摘要:Speech Emotion Recognition (SER) is widely deployed in Human-Computer Interaction, yet the high computational cost of conventional models hinders their implementation on resource-constrained edge devices. Spiking Neural Networks (SNNs) offer an energy-efficient alternative due to their event-driven nature; however, their integration with continuous Self-Supervised Learning (SSL) representations is fundamentally challenged by distribution mismatch, where high-dynamic-range embeddings degrade the information coding capacity of threshold-based neurons. To resolve this, we propose Prompt-Tuned Spiking Neural Networks (PTS-SNN), a parameter-efficient neuromorphic adaptation framework that aligns frozen SSL backbones with spiking dynamics. Specifically, we introduce a Temporal Shift Spiking Encoder to capture local temporal dependencies via parameter-free channel shifts, establishing a stable feature basis. To bridge the domain gap, we devise a Context-Aware Membrane Potential Calibration strategy. This mechanism leverages a Spiking Sparse Linear Attention module to aggregate global semantic context into learnable soft prompts, which dynamically regulate the bias voltages of Parametric Leaky Integrate-and-Fire (PLIF) neurons. This regulation effectively centers the heterogeneous input distribution within the responsive firing range, mitigating functional silence or saturation. Extensive experiments on five multilingual datasets (e.g., IEMOCAP, CASIA, EMODB) demonstrate that PTS-SNN achieves 73.34 % accuracy on IEMOCAP, comparable to competitive Artificial Neural Networks (ANNs), while requiring only 1.19M trainable parameters and 0.35 mJ inference energy per sample.


【8】Tutti: Expressive Multi-Singer Synthesis via Structure-Level Timbre Control and Vocal Texture Modeling
标题:Tutti:通过结构级音色控制和人声纹理建模的表现性多歌手合成
链接:https://arxiv.org/pdf/2602.08233v1

作者:Jiatao Chen,Xing Tang,Xiaoyue Duan,Yutang Feng,Jinchao Zhang,Jie Zhou
摘要:虽然现有的歌唱声音合成系统实现了高保真的独唱表演,但它们受到全局音色控制的限制,无法解决单首歌曲中的动态多歌手安排和声乐纹理。为了解决这个问题,我们提出了Tutti,一个统一的框架,设计用于结构化的多歌手生成。具体来说,我们引入了结构感知歌手提示,以实现灵活的歌手调度与音乐结构的发展,并提出了通过条件引导VAE的互补纹理学习,以捕捉隐含的声学纹理(例如,空间混响和频谱融合),其与显式控制互补。实验表明,Tutti擅长于精确的多歌手调度,显着提高了合唱生成的声学真实感,为复杂的多歌手安排提供了一个新的范例。音频样本可在https: annoauth123-ctrl.github.io Tutii_Demo 上获得。摘要:While existing Singing Voice Synthesis systems achieve high-fidelity solo performances, they are constrained by global timbre control, failing to address dynamic multi-singer arrangement and vocal texture within a single song. To address this, we propose Tutti, a unified framework designed for structured multi-singer generation. Specifically, we introduce a Structure-Aware Singer Prompt to enable flexible singer scheduling evolving with musical structure, and propose Complementary Texture Learning via Condition-Guided VAE to capture implicit acoustic textures (e.g., spatial reverberation and spectral fusion) that are complementary to explicit controls. Experiments demonstrate that Tutti excels in precise multi-singer scheduling and significantly enhances the acoustic realism of choral generation, offering a novel paradigm for complex multi-singer arrangement. Audio samples are available at https: annoauth123-ctrl.github.io Tutii_Demo .


【9】SNC: A Stem-Native Codec for Efficient Lossless Audio Storage with Adaptive Playback Capabilities
标题:SNC:一款具有自适应播放功能的高效无损音频存储的干原生编解码器
链接:https://arxiv.org/pdf/2602.08148v1

作者:Shaad Sufi
摘要:当前的音频格式在文件大小和功能之间存在一个基本的权衡:FLAC等无损格式保留了质量但缺乏适应性,而有损格式以保真度为代价减小了大小,并且不提供茎级访问。我们引入了Stem-Native Codec(SNC),这是一种新颖的音频容器格式,将音乐存储为独立编码的茎加上低能量的母带残差。通过利用与混合音频相比分离的茎的较低信息熵,SNC实现了与FLAC相比38.2%的文件大小减少(7.76 MB与2:18测试轨道的12.55 MB),同时保持感知透明度(STOI = 0.996)。与现有格式不同,SNC支持上下文感知的自适应播放,空间音频渲染和用户控制的混音,而无需额外的存储。我们的实验验证表明,茎加残留架构成功地解决了压缩效率和功能丰富性的冲突要求,为下一代音频分配系统提供了一条实用的道路。摘要:Current audio formats present a fundamental trade-off between file size and functionality: lossless formats like FLAC preserve quality but lack adaptability, while lossy formats reduce size at the cost of fidelity and offer no stem-level access.We introduce the Stem-Native Codec (SNC), a novel audio container format that stores music as independently encoded stems plus a low-energy mastering residual. By exploiting the lower information entropy of separated stems compared to mixed audio, SNC achieves a 38.2% file size reduction versus FLAC (7.76 MB vs. 12.55 MB for a 2:18 test track) while maintaining perceptual transparency (STOI = 0.996). Unlike existing formats, SNC enables context-aware adaptive playback, spatial audio rendering, and user-controlled remixing without requiring additional storage. Our experimental validation demonstrates that the stems-plus residual architecture successfully decouples the conflicting requirements of compression efficiency and feature richness, offering a practical path toward next-generation audio distribution systems.


【10】Equipping LLM with Directional Multi-Talker Speech Understanding Capabilities
标题:为LLM配备定向多说话者语音理解能力
链接:https://arxiv.org/pdf/2602.07211v1

作者:Ju Lin,Jing Pan,Ruizhi Li,Ming Sun,Yuzong Liu,Alaa Hassan,Jing Zheng,Florian Metze
摘要:最近的研究表明,用音频编码提示大型语言模型(LLM)可以实现有效的语音理解能力。然而,大多数语音LLM都是在单通道,单说话者数据上训练的,这使得将它们直接应用于多说话者和多通道语音理解任务具有挑战性。在这项工作中,我们提出了一个全面的调查,如何使定向多说话者语音理解能力的LLM,特别是在智能眼镜用例。我们提出了两种将方向性集成到LLM中的新方法:(1)利用源分离前端模块的级联系统,以及(2)利用串行化输出训练的端到端系统。所有这些方法都利用嵌入在智能眼镜中的多麦克风阵列来以流式方式优化方向性解释和处理。实验结果表明,我们提出的方法赋予LLM方向性语音理解能力的有效性,在语音识别和语音翻译任务中都取得了很好的性能。摘要:Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech understanding capabilities. However, most speech LLMs are trained on single-channel, single-talker data, which makes it challenging to directly apply them to multi-talker and multi-channel speech understanding task. In this work, we present a comprehensive investigation on how to enable directional multi-talker speech understanding capabilities for LLMs, specifically in smart glasses usecase. We propose two novel approaches to integrate directivity into LLMs: (1) a cascaded system that leverages a source separation front-end module, and (2) an end-to-end system that utilizes serialized output training. All of the approaches utilize a multi-microphone array embedded in smart glasses to optimize directivity interpretation and processing in a streaming manner. Experimental results demonstrate the efficacy of our proposed methods in endowing LLMs with directional speech understanding capabilities, achieving strong performance in both speech recognition and speech translation tasks.


【11】Massive Sound Embedding Benchmark (MSEB)
标题:大规模声音嵌入基准(MSEB)
链接:https://arxiv.org/pdf/2602.07143v1

作者:Georg Heigold,Ehsan Variani,Tom Bagby,Cyril Allauzen,Ji Ma,Shankar Kumar,Michael Riley
摘要:音频是多模态感知的关键组成部分,任何真正的智能系统都必须具备广泛的听觉能力。这些功能包括转录、分类、检索、推理、分割、聚类、重新排序和重建。从根本上讲,每个任务都涉及将原始音频信号转换为有意义的“嵌入”-无论是单个向量,连续或离散表示的序列,还是其他结构化形式-然后作为生成任务最终响应的基础。为了加速实现强大的机器听觉智能,我们提出了大量的声音嵌入基准(MSEB):一个可扩展的框架,旨在评估任何多模态系统的听觉组件。在其第一个版本中,MSEB提供了一套由八个核心任务组成的综合套件,未来还计划有更多的任务,并得到各种数据集的支持,包括新的大规模简单语音问题(SVQ)数据集。我们的初步实验建立了明确的性能余量,突出了改善音频为核心信号的现实世界多模式体验的重要机会。我们鼓励研究界使用MSEB来评估他们的算法并为其增长做出贡献。该库在github上公开托管。摘要:Audio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, each task involves transforming a raw audio signal into a meaningful 'embedding' - be it a single vector, a sequence of continuous or discrete representations, or another structured form - which then serves as the basis for generating the task's final response. To accelerate progress towards robust machine auditory intelligence, we present the Massive Sound Embedding Benchmark (MSEB): an extensible framework designed to evaluate the auditory components of any multimodal system. In its first release, MSEB offers a comprehensive suite of eight core tasks, with more planned for the future, supported by diverse datasets, including the new, large-scale Simple Voice Questions (SVQ) dataset. Our initial experiments establish clear performance headrooms, highlighting the significant opportunity to improve real-world multimodal experiences where audio is a core signal. We encourage the research community to use MSEB to assess their algorithms and contribute to its growth. The library is publicly hosted at github.


【12】CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
标题:CALM:大型音频语言模型的类条件稀疏注意力载体
链接:https://arxiv.org/pdf/2602.07077v1

作者:Videet Mehta,Liming Wang,Hilde Kuehne,Rogerio Feris,James R. Glass,M. Jehanzeb Mirza

备注:11 pages, 6 figures

摘要:大型音频语言模型(LALM)在多个下游任务中表现出强大的zero-shot能力,例如音频问题回答(AQA)和抽象推理;然而,这些模型仍然落后于某些区分性任务的专用模型(例如,音频分类)。最近的研究表明,在LALM内的注意力头的稀疏子集可以作为强判别特征提取器的下游任务,如通过简单的投票方案分类。然而,这些方法为所有选定的头部分配统一的权重,隐含地假设每个头部在所有语义类别中的贡献相等。在这项工作中,我们提出了大型音频语言模型的类条件稀疏注意向量,一个Few-Shot分类方法,学习类相关的重要性权重超过注意头。这种提法允许个人头专门在不同的语义类别,并有助于合奏预测成比例的估计可靠性。在多个Few-Shot音频和视听分类基准测试和任务上的实验表明,我们的方法在音频分类、视听分类和欺骗检测方面的绝对增益分别高达14.52%、1.53%和8.35%,始终优于最先进的基于统一投票的方法。摘要:Large audio-language models (LALMs) exhibit strong zero-shot capabilities in multiple downstream tasks, such as audio question answering (AQA) and abstract reasoning; however, these models still lag behind specialized models for certain discriminative tasks (e.g., audio classification). Recent studies show that sparse subsets of attention heads within an LALM can serve as strong discriminative feature extractors for downstream tasks such as classification via simple voting schemes. However, these methods assign uniform weights to all selected heads, implicitly assuming that each head contributes equally across all semantic categories. In this work, we propose Class-Conditional Sparse Attention Vectors for Large Audio-Language Models, a few-shot classification method that learns class-dependent importance weights over attention heads. This formulation allows individual heads to specialize in distinct semantic categories and to contribute to ensemble predictions proportionally to their estimated reliability. Experiments on multiple few-shot audio and audiovisual classification benchmarks and tasks demonstrate that our method consistently outperforms state-of-the-art uniform voting-based approaches by up to 14.52%, 1.53%, 8.35% absolute gains for audio classification, audio-visual classification, and spoofing detection respectively.


【13】Video-based Music Generation
标题:基于视频的音乐生成
链接:https://arxiv.org/pdf/2602.07063v1

作者:Serkan Sulun

备注:PhD thesis, University of Porto

摘要:随着互联网上视频内容的快速增长,找到合适的配乐仍然是一个重大挑战。本文介绍了EMSYNC(EMotion和SYNChronization),一个快速,免费,自动的解决方案,生成量身定制的输入视频的音乐,使内容创作者能够提高他们的作品,而无需创作或许可的音乐。我们的模型创建的音乐在情感和节奏上与视频同步。EMSYNC的核心组件是一种新颖的视频情感分类器。通过利用预训练的深度神经网络进行特征提取,并在只训练融合层的同时保持它们冻结,我们降低了计算复杂性,同时提高了准确性。我们通过在Ekman-6和MovieNet上获得最先进的结果来展示我们方法的泛化能力。另一个关键贡献是用于情感音乐生成的大规模情感标记数据集。然后,我们提出了一个基于情感的音乐生成器,第一个条件连续的情感值,而不是离散的类别,使细致入微的音乐生成与复杂的情感内容。为了提高时间同步,我们引入了一种新的时间边界条件的方法,称为“边界偏移编码”,调整音乐和弦与场景的变化。结合视频情感分类,基于情感的音乐生成和时间边界调节,EMSYNC成为一个全自动的基于视频的音乐生成器。用户研究表明,它在音乐丰富性,情感对齐,时间同步和整体偏好方面始终优于现有方法,为基于视频的音乐生成提供了新的最先进的技术。摘要:As the volume of video content on the internet grows rapidly, finding a suitable soundtrack remains a significant challenge. This thesis presents EMSYNC (EMotion and SYNChronization), a fast, free, and automatic solution that generates music tailored to the input video, enabling content creators to enhance their productions without composing or licensing music. Our model creates music that is emotionally and rhythmically synchronized with the video. A core component of EMSYNC is a novel video emotion classifier. By leveraging pretrained deep neural networks for feature extraction and keeping them frozen while training only fusion layers, we reduce computational complexity while improving accuracy. We show the generalization abilities of our method by obtaining state-of-the-art results on Ekman-6 and MovieNet. Another key contribution is a large-scale, emotion-labeled MIDI dataset for affective music generation. We then present an emotion-based MIDI generator, the first to condition on continuous emotional values rather than discrete categories, enabling nuanced music generation aligned with complex emotional content. To enhance temporal synchronization, we introduce a novel temporal boundary conditioning method, called "boundary offset encodings," aligning musical chords with scene changes. Combining video emotion classification, emotion-based music generation, and temporal boundary conditioning, EMSYNC emerges as a fully automatic video-based music generator. User studies show that it consistently outperforms existing methods in terms of music richness, emotional alignment, temporal synchronization, and overall preference, setting a new state-of-the-art in video-based music generation.


【14】MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
标题:MENASpeechBank:一个参考语音银行,为AudioLLM提供基于角色条件的多轮对话
链接:https://arxiv.org/pdf/2602.07036v1

作者:Zien Sheikh Ali,Hunzalah Hassan Bhatti,Rabindra Nath Nandi,Shammur Absar Chowdhury,Firoj Alam
备注:Foundation Models, Large Language Models, Native, Speech Models, Arabic, AI-persona, Persona-conditioned-conversations
摘要:音频大型语言模型(AudioLLM)支持语音和一般音频的指令遵循,但由于缺乏多样化、对话式、与指令一致的语音文本数据,进展越来越受到限制。这一瓶颈对于基于人物角色的互动和方言覆盖尤其严重,因为收集和发布真实的多说话者录音既昂贵又缓慢。我们介绍MENASpeechBank,一个参考语音库,包括来自多个MENA国家的124位发言者的约18K高质量话语,涵盖英语,现代标准阿拉伯语(MSA)和区域阿拉伯语品种。在此基础上,我们开发了一个可控的合成数据管道,该管道:(i)构建富含世界价值观调查启发的属性的人物简档,(ii)定义约5K会话场景的分类,(iii)通过语义相似性将人物与场景匹配,(iv)与LLM生成大约417K个角色扮演对话,其中用户作为角色说话,助理作为有用的代理,以及(v)通过调节参考说话者音频来合成用户回合以保持说话者身份和多样性。我们评估合成和人工记录的对话,并提供详细的分析。我们将向社区公开发布MENASpeechBank和生成的对话。摘要:Audio large language models (AudioLLMs) enable instruction-following over speech and general audio, but progress is increasingly limited by the lack of diverse, conversational, instruction-aligned speech-text data. This bottleneck is especially acute for persona-grounded interactions and dialectal coverage, where collecting and releasing real multi-speaker recordings is costly and slow. We introduce MENASpeechBank, a reference speech bank comprising about 18K high-quality utterances from 124 speakers spanning multiple MENA countries, covering English, Modern Standard Arabic (MSA), and regional Arabic varieties. Building on this resource, we develop a controllable synthetic data pipeline that: (i) constructs persona profiles enriched with World Values Survey-inspired attributes, (ii) defines a taxonomy of about 5K conversational scenarios, (iii) matches personas to scenarios via semantic similarity, (iv) generates about 417K role-play conversations with an LLM where the user speaks as the persona and the assistant behaves as a helpful agent, and (v) synthesizes the user turns by conditioning on reference speaker audio to preserve speaker identity and diversity. We evaluate both synthetic and human-recorded conversations and provide detailed analysis. We will release MENASpeechBank and the generated conversations publicly for the community.


【15】Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods
标题:利用OpenAI的Whisper表示和细心的池化方法的语音情感识别
链接:https://arxiv.org/pdf/2602.06000v1

作者:Ali Shendabadi,Parnia Izadirad,Mostafa Salehi,Mahmoud Bijankhan
摘要:语音情感识别(SER)研究由于缺乏标准和足够大的数据集而面临限制。最近的研究利用预训练模型来提取下游任务的特征,如SER。这项工作探讨了Whisper,一个预训练的ASR系统,在语音情感识别中的能力,提出了两种基于注意力的池化方法,多头注意力平均池化和QKV池化,旨在有效地降低Whisper表示的维数,同时保留情感特征。我们分别使用IEMOCAP和ShEMO数据集,使用Whisper Tiny和Small对英语和波斯语进行实验。我们的多头QKV架构在ShEMO数据集上实现了最先进的结果,未加权精度提高了2.47%。我们进一步比较了不同Whisper编码器层的性能,发现中间层通常在波斯数据集上的SER表现更好,为HuBERT X-Large等更大的模型提供了轻量级和高效的替代方案。我们的研究结果突出了Whisper作为SER的表示提取器的潜力,并证明了基于注意力的池降维的有效性。摘要:Speech Emotion Recognition (SER) research has faced limitations due to the lack of standard and sufficiently large datasets. Recent studies have leveraged pre-trained models to extract features for downstream tasks such as SER. This work explores the capabilities of Whisper, a pre-trained ASR system, in speech emotion recognition by proposing two attention-based pooling methods, Multi-head Attentive Average Pooling and QKV Pooling, designed to efficiently reduce the dimensionality of Whisper representations while preserving emotional features. We experiment on English and Persian, using the IEMOCAP and ShEMO datasets respectively, with Whisper Tiny and Small. Our multi-head QKV architecture achieves state-of-the-art results on the ShEMO dataset, with a 2.47% improvement in unweighted accuracy. We further compare the performance of different Whisper encoder layers and find that intermediate layers often perform better for SER on the Persian dataset, providing a lightweight and efficient alternative to much larger models such as HuBERT X-Large. Our findings highlight the potential of Whisper as a representation extractor for SER and demonstrate the effectiveness of attention-based pooling for dimension reduction.


【16】SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
标题:SoulX-Singer:迈向高质量Zero-Shot歌唱语音合成
链接:https://arxiv.org/pdf/2602.07803v1

作者:Jiale Qian,Hao Meng,Tian Zheng,Pengcheng Zhu,Haopeng Lin,Yuhang Dai,Hanke Xie,Wenxiao Cao,Ruixuan Shang,Jun Wu,Hongmei Liu,Hanlin Wen,Jian Zhao,Zhonglin Jiang,Yong Chen,Shunshun Yin,Ming Tao,Jianguo Wei,Lei Xie,Xinsheng Wang
备注:Technical Report
摘要:虽然近年来语音合成取得了快速进展,但开源歌声合成(SVS)系统在工业部署方面仍然面临重大障碍,特别是在鲁棒性和zero-shot泛化方面。在这份报告中,我们介绍了SoulX-Singer,这是一个高质量的开源SVS系统,设计时考虑了实际部署因素。SoulX-Singer支持以符号乐谱(MIDI)或旋律表示为条件的可控歌唱生成,从而在现实世界的制作工作流程中实现灵活且富有表现力的控制。经过超过42,000小时的语音数据训练,该系统支持普通话、英语和粤语,并在不同的音乐条件下持续实现跨语言的最先进的合成质量。此外,为了能够在实际场景中可靠地评估zero-shot SVS性能,我们构建了SoulX-Singer-Eval,这是一个具有严格训练-测试分离的专用基准,便于在zero-shot设置中进行系统评估。摘要:While recent years have witnessed rapid progress in speech synthesis, open-source singing voice synthesis (SVS) systems still face significant barriers to industrial deployment, particularly in terms of robustness and zero-shot generalization. In this report, we introduce SoulX-Singer, a high-quality open-source SVS system designed with practical deployment considerations in mind. SoulX-Singer supports controllable singing generation conditioned on either symbolic musical scores (MIDI) or melodic representations, enabling flexible and expressive control in real-world production workflows. Trained on more than 42,000 hours of vocal data, the system supports Mandarin Chinese, English, and Cantonese and consistently achieves state-of-the-art synthesis quality across languages under diverse musical conditions. Furthermore, to enable reliable evaluation of zero-shot SVS performance in practical scenarios, we construct SoulX-Singer-Eval, a dedicated benchmark with strict training-test disentanglement, facilitating systematic assessment in zero-shot settings.


eess.AS音频处理


【1】Input-Adaptive Spectral Feature Compression by Sequence Modeling for Source Separation
标题:通过序列建模实现源分离的输入自适应光谱特征压缩
链接:https://arxiv.org/pdf/2602.08671v1

作者:Kohei Saijo,Yoshiaki Bando
备注:Accepted by IEEE TASLP. c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses
摘要:时频域双径模型在信号源分离中表现出很强的性能,被广泛应用。由于其计算成本随着频率仓的数量而增加,因此这些模型通常在高采样率任务中使用带分离(BS)模块,例如音乐源分离(MSS)和电影音频源分离(CASS)。BS编码器通过对每个预定义子带的特征进行编码来压缩频率信息。它通过引入电感偏置来实现有效的压缩,该偏置更加强调低频部分。尽管其成功,BS模块具有两个固有的限制:(i)它不是输入自适应的,防止使用依赖于输入的信息,以及(ii)参数计数很大,因为每个子带需要专用模块。为了解决这些问题,我们提出了光谱特征压缩(SFC)。SFC使用单个序列建模模块压缩输入,使其具有输入自适应性和参数高效性。我们调查两种变体的SFC,一个基于交叉注意和其他曼巴,并引入归纳偏见的启发BS模块,使它们适合于频率信息压缩。MSS和CASS任务的实验表明,SFC模块始终优于BS模块在不同的分离器大小和压缩比。我们还提供了一个分析表明,SFC自适应捕捉输入的频率模式。摘要:Time-frequency domain dual-path models have demonstrated strong performance and are widely used in source separation. Because their computational cost grows with the number of frequency bins, these models often use the band-split (BS) module in high-sampling-rate tasks such as music source separation (MSS) and cinematic audio source separation (CASS). The BS encoder compresses frequency information by encoding features for each predefined subband. It achieves effective compression by introducing an inductive bias that places greater emphasis on low-frequency parts. Despite its success, the BS module has two inherent limitations: (i) it is not input-adaptive, preventing the use of input-dependent information, and (ii) the parameter count is large, since each subband requires a dedicated module. To address these issues, we propose Spectral Feature Compression (SFC). SFC compresses the input using a single sequence modeling module, making it both input-adaptive and parameter-efficient. We investigate two variants of SFC, one based on cross-attention and the other on Mamba, and introduce inductive biases inspired by the BS module to make them suitable for frequency information compression. Experiments on MSS and CASS tasks demonstrate that the SFC module consistently outperforms the BS module across different separator sizes and compression ratios. We also provide an analysis showing that SFC adaptively captures frequency patterns from the input.


【2】Physics-Guided Variational Model for Unsupervised Sound Source Tracking
标题:无监督声源跟踪的物理引导变分模型
链接:https://arxiv.org/pdf/2602.08484v1

作者:Luan Vinícius Fiorio,Ivana Nikoloska,Bruno Defraene,Alex Young,Johan David,Ronald M. Aarts
备注:This work has been submitted to the IEEE for possible publication
摘要:声源跟踪通常使用经典的阵列处理算法来执行。替代方法,如机器学习,依赖于地面实况位置标签,这是昂贵的获得。我们提出了一个变分模型,可以执行单源无监督的声源跟踪在潜在的空间,基于物理的解码器的帮助。我们的实验表明,所提出的方法超越了传统的基线,并实现了性能和计算复杂度相媲美的国家的最先进的监督模型。我们还表明,该方法具有很大的鲁棒性改变麦克风阵列的几何形状和损坏的麦克风位置元数据。最后,将该方法推广到多声源跟踪,并提出了基本的理论变化。摘要:Sound source tracking is often performed using classical array-processing algorithms. Alternative methods, such as machine learning, rely on ground truth position labels, which are costly to obtain. We propose a variational model that can perform single-source unsupervised sound source tracking in latent space, aided by a physics-based decoder. Our experiments demonstrate that the proposed method surpasses traditional baselines and achieves performance and computational complexity comparable to state-of-the-art supervised models. We also show that the method presents substantial robustness to altered microphone array geometries and corrupted microphone position metadata. Finally, the method is extended to multi-source sound tracking and the basic theoretical changes are proposed.


【3】Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition
标题:用于噪音鲁棒视听语音识别的跨模式瓶颈融合
链接:https://arxiv.org/pdf/2602.08293v1

作者:Seaone Ok,Min Jun Choi,Eungbeom Kim,Seungu Han,Kyogu Lee
备注:5 pages, 3 figures, ICASSP 2026 Accepted
摘要:视听语音识别(AVSR)利用声学和视觉线索来改善噪声条件下的语音识别。一个核心问题是如何设计一个融合机制,使模型能够有效地利用视觉信息时,音频信号是退化的,同时保持强大的性能上干净的语音。我们提出了Cobra(Cross-modal Bottleneck for Robust AVSR),这是一个基于验证的融合框架,它引入了一组紧凑的可学习令牌来调解跨模态交换。通过调节通过这些令牌的信息流,即使在不利或域外噪声下,音频流也可以可靠地访问基本的视觉提示。尽管训练数据有限,但我们的模型超过了可比基线,并通过噪声自适应融合与大规模系统保持竞争力,证明了效率和鲁棒性。消融研究强调融合深度是最关键的因素,强调了其在设计强大AVSR系统中的重要性。摘要:Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual cues to improve speech recognition under noisy conditions. A central question is how to design a fusion mechanism that allows the model to effectively exploit visual information when the audio signal is degraded, while maintaining strong performance on clean speech. We propose CoBRA (Cross-modal Bottleneck for Robust AVSR), a bottleneck-based fusion framework that introduces a compact set of learnable tokens to mediate cross-modal exchange. By regulating information flow through these tokens, the audio stream can reliably access essential visual cues even under adverse or out-of-domain noise. Despite limited training data, our model surpasses comparable baselines and remains competitive with large-scale systems through noise-adaptive fusion, demonstrating both efficiency and robustness. Ablation studies highlight that the depth of fusion is the most critical factor, underscoring its importance in designing robust AVSR systems.


【4】Detect, Attend and Extract: Keyword Guided Target Speaker Extraction
标题:检测、参与和提取:关键词引导的目标说话人提取
链接:https://arxiv.org/pdf/2602.07977v1

作者:Haoyu Li,Yu Xi,Yidi Jiang,Shuai Wang,Kate Knill,Mark Gales,Haizhou Li,Kai Yu
备注:4 figures, 4 tables. Submitted to IJCAI-ECAI 2026
摘要:目标说话人提取(TSE)的目的是从包含多个竞争说话人的混合语音中提取目标说话人的语音。传统的TSE系统主要依赖于说话者线索,例如预先登记的语音,来识别和隔离目标说话者。然而,在许多实际场景中,干净的登记话语是不可用的,限制了现有方法的适用性。在这项工作中,我们提出了DAE-TSE,一个关键词引导的TSE框架,指定目标扬声器通过不同的关键字,他们说出。通过利用关键字(即,作为线索,我们的方法提供了一个灵活和实用的替代方案,以登记为基础的TSE。DAE-TSE遵循检测-参与-提取(DAE)范式:它首先检测给定关键词的存在,然后根据关键词内容关注相应的说话者,最后提取目标语音。实验结果表明,DAE-TSE优于依赖于干净的注册语音的标准TSE系统。据我们所知,这是第一项利用部分转录作为在TSE中指定目标说话者的线索的研究,为现实世界的场景提供了灵活实用的解决方案。我们的代码和演示页面现已公开。摘要:Target speaker extraction (TSE) aims to extract the speech of a target speaker from mixtures containing multiple competing speakers. Conventional TSE systems predominantly rely on speaker cues, such as pre-enrolled speech, to identify and isolate the target speaker. However, in many practical scenarios, clean enrollment utterances are unavailable, limiting the applicability of existing approaches. In this work, we propose DAE-TSE, a keyword-guided TSE framework that specifies the target speaker through distinct keywords they utter. By leveraging keywords (i.e., partial transcriptions) as cues, our approach provides a flexible and practical alternative to enrollment-based TSE. DAE-TSE follows the Detect-Attend-Extract (DAE) paradigm: it first detects the presence of the given keywords, then attends to the corresponding speaker based on the keyword content, and finally extracts the target speech. Experimental results demonstrate that DAE-TSE outperforms standard TSE systems that rely on clean enrollment speech. To the best of our knowledge, this is the first study to utilize partial transcription as a cue for specifying the target speaker in TSE, offering a flexible and practical solution for real-world scenarios. Our code and demo page are now publicly available.


【5】SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
标题:SoulX-Singer:迈向高质量Zero-Shot歌唱语音合成
链接:https://arxiv.org/pdf/2602.07803v1

作者:Jiale Qian,Hao Meng,Tian Zheng,Pengcheng Zhu,Haopeng Lin,Yuhang Dai,Hanke Xie,Wenxiao Cao,Ruixuan Shang,Jun Wu,Hongmei Liu,Hanlin Wen,Jian Zhao,Zhonglin Jiang,Yong Chen,Shunshun Yin,Ming Tao,Jianguo Wei,Lei Xie,Xinsheng Wang
备注:Technical Report
摘要:虽然近年来语音合成取得了快速进展,但开源歌声合成(SVS)系统在工业部署方面仍然面临重大障碍,特别是在鲁棒性和zero-shot泛化方面。在这份报告中,我们介绍了SoulX-Singer,这是一个高质量的开源SVS系统,设计时考虑了实际部署因素。SoulX-Singer支持以符号乐谱(MIDI)或旋律表示为条件的可控歌唱生成,从而在现实世界的制作工作流程中实现灵活且富有表现力的控制。经过超过42,000小时的语音数据训练,该系统支持普通话、英语和粤语,并在不同的音乐条件下持续实现跨语言的最先进的合成质量。此外,为了能够在实际场景中可靠地评估zero-shot SVS性能,我们构建了SoulX-Singer-Eval,这是一个具有严格训练-测试分离的专用基准,便于在zero-shot设置中进行系统评估。摘要:While recent years have witnessed rapid progress in speech synthesis, open-source singing voice synthesis (SVS) systems still face significant barriers to industrial deployment, particularly in terms of robustness and zero-shot generalization. In this report, we introduce SoulX-Singer, a high-quality open-source SVS system designed with practical deployment considerations in mind. SoulX-Singer supports controllable singing generation conditioned on either symbolic musical scores (MIDI) or melodic representations, enabling flexible and expressive control in real-world production workflows. Trained on more than 42,000 hours of vocal data, the system supports Mandarin Chinese, English, and Cantonese and consistently achieves state-of-the-art synthesis quality across languages under diverse musical conditions. Furthermore, to enable reliable evaluation of zero-shot SVS performance in practical scenarios, we construct SoulX-Singer-Eval, a dedicated benchmark with strict training-test disentanglement, facilitating systematic assessment in zero-shot settings.


【6】Rho-Perfect: Correlation Ceiling For Subjective Evaluation Datasets
标题:Rho-Perfect:主观评估数据集的相关性上限
链接:https://arxiv.org/pdf/2602.08552v1

作者:Fredrik Cumlin
摘要:主观评级包含固有的噪声,限制了模型-人的相关性,但这种可靠性问题很少被量化。在本文中,我们提出了$ρ$-Perfect,这是一个对主观评级数据集的模型的最高可实现相关性的实用估计。我们将$ρ$-Perfect定义为完美预测器与人类评级之间的相关性,并基于异方差噪声场景(主观评级数据集中常见的情况)得出该值的估计。我们证明了$ρ$-完美平方估计重测相关性,并使用它来验证估计。我们演示了在语音质量数据集上使用$ρ$-Perfect,并展示了该度量如何区分模型限制和数据质量问题。摘要:Subjective ratings contain inherent noise that limits the model-human correlation, but this reliability issue is rarely quantified. In this paper, we present $ρ$-Perfect, a practical estimation of the highest achievable correlation of a model on subjectively rated datasets. We define $ρ$-Perfect to be the correlation between a perfect predictor and human ratings, and derive an estimate of the value based on heteroscedastic noise scenarios, a common occurrence in subjectively rated datasets. We show that $ρ$-Perfect squared estimates test-retest correlation and use this to validate the estimate. We demonstrate the use of $ρ$-Perfect on a speech quality dataset and show how the measure can distinguish between model limitations and data quality issues.


【7】SNC: A Stem-Native Codec for Efficient Lossless Audio Storage with Adaptive Playback Capabilities
标题:SNC:一款具有自适应播放功能的高效无损音频存储的干原生编解码器
链接:https://arxiv.org/pdf/2602.08148v1

作者:Shaad Sufi
摘要:当前的音频格式在文件大小和功能之间存在一个基本的权衡:FLAC等无损格式保留了质量但缺乏适应性,而有损格式以保真度为代价减小了大小,并且不提供茎级访问。我们引入了Stem-Native Codec(SNC),这是一种新颖的音频容器格式,将音乐存储为独立编码的茎加上低能量的母带残差。通过利用与混合音频相比分离的茎的较低信息熵,SNC实现了与FLAC相比38.2%的文件大小减少(7.76 MB与2:18测试轨道的12.55 MB),同时保持感知透明度(STOI = 0.996)。与现有格式不同,SNC支持上下文感知的自适应播放,空间音频渲染和用户控制的混音,而无需额外的存储。我们的实验验证表明,茎加残留架构成功地解决了压缩效率和功能丰富性的冲突要求,为下一代音频分配系统提供了一条实用的道路。摘要:Current audio formats present a fundamental trade-off between file size and functionality: lossless formats like FLAC preserve quality but lack adaptability, while lossy formats reduce size at the cost of fidelity and offer no stem-level access.We introduce the Stem-Native Codec (SNC), a novel audio container format that stores music as independently encoded stems plus a low-energy mastering residual. By exploiting the lower information entropy of separated stems compared to mixed audio, SNC achieves a 38.2% file size reduction versus FLAC (7.76 MB vs. 12.55 MB for a 2:18 test track) while maintaining perceptual transparency (STOI = 0.996). Unlike existing formats, SNC enables context-aware adaptive playback, spatial audio rendering, and user-controlled remixing without requiring additional storage. Our experimental validation demonstrates that the stems-plus residual architecture successfully decouples the conflicting requirements of compression efficiency and feature richness, offering a practical path toward next-generation audio distribution systems.


【8】MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
标题:MENASpeechBank:一个参考语音银行,为AudioLLM提供基于角色条件的多轮对话
链接:https://arxiv.org/pdf/2602.07036v1

作者:Zien Sheikh Ali,Hunzalah Hassan Bhatti,Rabindra Nath Nandi,Shammur Absar Chowdhury,Firoj Alam
备注:Foundation Models, Large Language Models, Native, Speech Models, Arabic, AI-persona, Persona-conditioned-conversations
摘要:音频大型语言模型(AudioLLM)支持语音和一般音频的指令遵循,但由于缺乏多样化、对话式、与指令一致的语音文本数据,进展越来越受到限制。这一瓶颈对于基于人物角色的互动和方言覆盖尤其严重,因为收集和发布真实的多说话者录音既昂贵又缓慢。我们介绍MENASpeechBank,一个参考语音库,包括来自多个MENA国家的124位发言者的约18K高质量话语,涵盖英语,现代标准阿拉伯语(MSA)和区域阿拉伯语品种。在此基础上,我们开发了一个可控的合成数据管道,该管道:(i)构建富含世界价值观调查启发的属性的人物简档,(ii)定义约5K会话场景的分类,(iii)通过语义相似性将人物与场景匹配,(iv)与LLM生成大约417K个角色扮演对话,其中用户作为角色说话,助理作为有用的代理,以及(v)通过调节参考说话者音频来合成用户回合以保持说话者身份和多样性。我们评估合成和人工记录的对话,并提供详细的分析。我们将向社区公开发布MENASpeechBank和生成的对话。摘要:Audio large language models (AudioLLMs) enable instruction-following over speech and general audio, but progress is increasingly limited by the lack of diverse, conversational, instruction-aligned speech-text data. This bottleneck is especially acute for persona-grounded interactions and dialectal coverage, where collecting and releasing real multi-speaker recordings is costly and slow. We introduce MENASpeechBank, a reference speech bank comprising about 18K high-quality utterances from 124 speakers spanning multiple MENA countries, covering English, Modern Standard Arabic (MSA), and regional Arabic varieties. Building on this resource, we develop a controllable synthetic data pipeline that: (i) constructs persona profiles enriched with World Values Survey-inspired attributes, (ii) defines a taxonomy of about 5K conversational scenarios, (iii) matches personas to scenarios via semantic similarity, (iv) generates about 417K role-play conversations with an LLM where the user speaks as the persona and the assistant behaves as a helpful agent, and (v) synthesizes the user turns by conditioning on reference speaker audio to preserve speaker identity and diversity. We evaluate both synthetic and human-recorded conversations and provide detailed analysis. We will release MENASpeechBank and the generated conversations publicly for the community.


机器翻译由腾讯交互翻译提供,仅供参考