今日论文合集:cs.SD语音3篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Difficulty-Controlled Simplification of Piano Scores with Synthetic Data for Inclusive Music Education
标题:利用合成数据简化钢琴乐谱,实现包容性音乐教育的难度控制
链接:https://arxiv.org/abs/2511.16228

作者:Pedro Ramoneda,Emilia Parada-Cabaleiro,Dasaem Jeong,Xavier Serra
摘要:尽管有潜力,但人工智能在音乐教育领域的进步受到专有系统的阻碍,这些系统限制了该领域技术的民主化。特别是,人工智能驱动的音乐难度调整特别有前途,因为简化复杂的作品可以使音乐教育更具包容性,并使所有年龄和背景的学习者都能接受。然而,最近的努力依赖于专有的数据集,这阻止了研究社区复制,比较或扩展当前的最新技术。此外,虽然这些生成方法提供了巨大的潜力,但它们中的大多数使用的是XML格式,与其他格式不同,如MusicXML,缺乏可读性和布局信息,从而限制了它们对人类表演者的实际使用。本文介绍了一种基于transformer的方法来调整MusicXML钢琴乐谱的难度。与以前的方法不同,这些方法依赖于注释数据集,我们提出了一个由按估计难度排序的成对钢琴乐谱组成的合成数据集,每对乐谱都包含同一首曲子的更具有挑战性和更容易的排列。我们通过创建以相同旋律和和声为条件的变奏来生成这些配对,并利用预先训练的模型来评估难度和风格,确保适当的配对。实验结果表明,所提出的方法的有效性,显示出准确的控制的可玩性和目标的难度,突出通过定性和定量的评价。与以前的工作相比,我们公开发布所有资源(代码,数据集和模型),确保可重复性,同时促进开源创新,以帮助弥合数字鸿沟。
摘要:Despite its potential, AI advances in music education are hindered by proprietary systems that limit the democratization of technology in this domain. In particular, AI-driven music difficulty adjustment is especially promising, as simplifying complex pieces can make music education more inclusive and accessible to learners of all ages and contexts. Nevertheless, recent efforts have relied on proprietary datasets, which prevents the research community from reproducing, comparing, or extending the current state of the art. In addition, while these generative methods offer great potential, most of them use the MIDI format, which, unlike others, such as MusicXML, lacks readability and layout information, thereby limiting their practical use for human performers. This work introduces a transformer-based method for adjusting the difficulty of MusicXML piano scores. Unlike previous methods, which rely on annotated datasets, we propose a synthetic dataset composed of pairs of piano scores ordered by estimated difficulty, with each pair comprising a more challenging and easier arrangement of the same piece. We generate these pairs by creating variations conditioned on the same melody and harmony and leverage pretrained models to assess difficulty and style, ensuring appropriate pairing. The experimental results illustrate the validity of the proposed approach, showing accurate control of playability and target difficulty, as highlighted through qualitative and quantitative evaluations. In contrast to previous work, we openly release all resources (code, dataset, and models), ensuring reproducibility while fostering open-source innovation to help bridge the digital divide.


【2】SceneGuard: Training-Time Voice Protection with Scene-Consistent Audible Background Noise
标题:SceneGuard:训练时语音保护,具有场景一致的可听背景噪音
链接:https://arxiv.org/abs/2511.16114

作者:Rui Sang,Yuxuan Liu
摘要:语音克隆技术通过从有限的音频样本中进行未经授权的语音合成,构成了严重的隐私威胁。基于不可感知的对抗性扰动的现有防御容易受到常见音频预处理(如去噪和压缩)的影响。我们提出了SceneGuard,一种训练时间的语音保护方法,适用于场景一致的可听背景噪声的语音录音。与不可感知的扰动不同,SceneGuard利用自然发生的声学场景(例如,机场、街道、公园),以产生保护性噪声,该噪声适合于环境并且对于对策是鲁棒的。我们对SceneGuard进行了文本到语音训练攻击的评估,显示出5.5%的说话人相似性下降,具有极高的统计显著性(p < 10^{-15},Cohen's d = 2.18),同时保留了98.6%的语音清晰度(STOI = 0.986)。鲁棒性评估表明,SceneGuard在五种常见的对策下保持或增强保护,包括MP3压缩,频谱减法,低通滤波和下采样。我们的研究结果表明,可听的,场景一致的噪声提供了一个更强大的替代难以察觉的扰动训练时间的声音保护。源代码可在https://github.com/richael-sang/SceneGuard上获得。
摘要:Voice cloning technology poses significant privacy threats by enabling unauthorized speech synthesis from limited audio samples. Existing defenses based on imperceptible adversarial perturbations are vulnerable to common audio preprocessing such as denoising and compression. We propose SceneGuard, a training-time voice protection method that applies scene-consistent audible background noise to speech recordings. Unlike imperceptible perturbations, SceneGuard leverages naturally occurring acoustic scenes (e.g., airport, street, park) to create protective noise that is contextually appropriate and robust to countermeasures. We evaluate SceneGuard on text-to-speech training attacks, demonstrating 5.5% speaker similarity degradation with extremely high statistical significance (p < 10^{-15}, Cohen's d = 2.18) while preserving 98.6% speech intelligibility (STOI = 0.986). Robustness evaluation shows that SceneGuard maintains or enhances protection under five common countermeasures including MP3 compression, spectral subtraction, lowpass filtering, and downsampling. Our results suggest that audible, scene-consistent noise provides a more robust alternative to imperceptible perturbations for training-time voice protection. The source code are available at: https://github.com/richael-sang/SceneGuard.


【3】Step-Audio-R1 Technical Report
标题:Step-Audio-R1技术报告
链接:https://arxiv.org/abs/2511.15848

作者:Fei Tian,Xiangyu Tony Zhang,Yuxin Zhang,Haoyang Zhang,Yuxin Li,Daijiao Liu,Yayue Deng,Donghang Wu,Jun Chen,Liang Zhao,Chengyuan Yao,Hexin Liu,Eng Siong Chng,Xuerui Yang,Xiangyu Zhang,Daxin Jiang,Gang Yu
备注:15 pages, 5 figures. Technical Report
摘要:推理模型的最新进展表明,通过扩展的思想链审议在文本和视觉领域取得了显着的成功。然而,一个令人困惑的现象仍然存在于音频语言模型中:它们在最少或没有推理的情况下始终表现得更好,这提出了一个基本问题-音频智能真的能从深思熟虑中受益吗?我们介绍Step-Audio-R1,这是第一个成功解锁音频领域推理能力的音频推理模型。通过我们提出的模态接地推理蒸馏(MGRD)框架,Step-Audio-R1学习生成音频相关的推理链,这些推理链真正基于声学特征,而不是幻觉断开审议。我们的模型具有强大的音频推理能力,超越了Gemini 2.5 Pro,并在语音,环境声音和音乐的全面音频理解和推理基准方面实现了与最先进的Gemini 3 Pro相当的性能。这些结果表明,推理是一种跨模态的可转移能力,当适当地锚定时,将扩展的审议从一种负债转变为音频智能的强大资产。通过建立第一个成功的音频推理模型,Step-Audio-R1为构建真正的多模态推理系统开辟了新的途径,该系统可以深入思考所有感官模式。
摘要:Recent advances in reasoning models have demonstrated remarkable success in text and vision domains through extended chain-of-thought deliberation. However, a perplexing phenomenon persists in audio language models: they consistently perform better with minimal or no reasoning, raising a fundamental question - can audio intelligence truly benefit from deliberate thinking? We introduce Step-Audio-R1, the first audio reasoning model that successfully unlocks reasoning capabilities in the audio domain. Through our proposed Modality-Grounded Reasoning Distillation (MGRD) framework, Step-Audio-R1 learns to generate audio-relevant reasoning chains that genuinely ground themselves in acoustic features rather than hallucinating disconnected deliberations. Our model exhibits strong audio reasoning capabilities, surpassing Gemini 2.5 Pro and achieving performance comparable to the state-of-the-art Gemini 3 Pro across comprehensive audio understanding and reasoning benchmarks spanning speech, environmental sounds, and music. These results demonstrate that reasoning is a transferable capability across modalities when appropriately anchored, transforming extended deliberation from a liability into a powerful asset for audio intelligence. By establishing the first successful audio reasoning model, Step-Audio-R1 opens new pathways toward building truly multimodal reasoning systems that think deeply across all sensory modalities.


eess.AS音频处理


【1】Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
标题:Codec 2Vec:使用神经语音编解码器的自我监督语音表示学习
链接:https://arxiv.org/abs/2511.16639

作者:Wei-Cheng Tseng,David Harwath
备注:To be presented at ASRU 2025
摘要:神经音频编解码器的最新进展不仅实现了卓越的音频压缩,而且增强了语音合成技术。研究人员正在探索它们作为通用声学特征提取器的潜力,用于更广泛的语音处理任务。在这一趋势的基础上,我们推出了Codec 2Vec,这是第一个完全依赖于离散音频编解码器单元的语音表示学习框架。这种方法具有几个优点,包括提高数据存储和传输效率,更快的训练和增强的数据隐私。我们探索掩蔽预测与各种训练目标推导策略,以彻底了解这个框架的有效性。在SUPERB基准测试中,Codec2Vec与连续输入模型相比具有竞争力的性能,同时将存储需求减少了16.5倍,训练时间减少了2.3倍,展示了其可扩展性和效率。
摘要:Recent advancements in neural audio codecs have not only enabled superior audio compression but also enhanced speech synthesis techniques. Researchers are now exploring their potential as universal acoustic feature extractors for a broader range of speech processing tasks. Building on this trend, we introduce Codec2Vec, the first speech representation learning framework that relies exclusively on discrete audio codec units. This approach offers several advantages, including improved data storage and transmission efficiency, faster training, and enhanced data privacy. We explore masked prediction with various training target derivation strategies to thoroughly understand the effectiveness of this framework. Evaluated on the SUPERB benchmark, Codec2Vec achieves competitive performance compared to continuous-input models while reducing storage requirements by up to 16.5x and training time by 2.3x, showcasing its scalability and efficiency.


【2】SUNAC: Source-aware Unified Neural Audio Codec
标题:SUAC:源感知统一神经音频编解码器
链接:https://arxiv.org/abs/2511.16126

作者:Ryo Aihara,Yoshiki Masuyama,Francesco Paissan,François G. Germain,Gordon Wichern,Jonathan Le Roux
备注:Submitted to ICASSP 2026
摘要:神经音频编解码器(NAC)提供了紧凑的表示,可以在许多下游应用中使用,特别是大型语言模型。然而,大多数NAC以纠缠的方式对多个源的混合进行编码,这可能妨碍仅需要访问源的子集(例如,特定类型声音的分析、给定说话者的转录等)。为了解决这个问题,我们提出了一个源感知的编解码器,直接从混合物中编码单个源,源类型提示的条件。这使得能够实现对要编码的源的用户驱动的选择,包括单独地编码相同类型的多个源(例如,多个语音信号)。实验表明,我们的模型实现了竞争力的再合成和分离质量相对于级联的源分离,其次是一个传统的NAC,具有较低的计算成本。
摘要:Neural audio codecs (NACs) provide compact representations that can be leveraged in many downstream applications, in particular large language models. Yet most NACs encode mixtures of multiple sources in an entangled manner, which may impede efficient downstream processing in applications that need access to only a subset of the sources (e.g., analysis of a particular type of sound, transcription of a given speaker, etc). To address this, we propose a source-aware codec that encodes individual sources directly from mixtures, conditioned on source type prompts. This enables user-driven selection of which source(s) to encode, including separately encoding multiple sources of the same type (e.g., multiple speech signals). Experiments show that our model achieves competitive resynthesis and separation quality relative to a cascade of source separation followed by a conventional NAC, with lower computational cost.


【3】Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
标题:Train Short,Infer Long:Speech-LLM在长音频上实现Zero-Shot流媒体联合ASB和Dialization
链接:https://arxiv.org/abs/2511.16046

作者:Mohan Shi,Xiong Xiao,Ruchao Fan,Shaoshi Ling,Jinyu Li
备注:Submitted to ICASSP2026
摘要:联合自动语音识别(ASR)和说话人日记旨在回答多说话人场景中“谁说了什么”的问题。本文提出了一个端到端的语音大语言模型(Speech-LLM),用于联合可压缩的DIarization和aSr(JEDIS-LLM)。该模型仅在20秒以下的短音频上进行训练,但能够在不进行额外训练的情况下对长格式音频进行流式推理。这是通过在块式流推理期间引入具有动态更新机制的扬声器提示缓存(SPC)来实现的,其灵感来自LLM的自回归性质。SPC还允许无缝使用预先注册的发言者配置文件,这在会议转录等许多场景中很常见。为了进一步增强日记化能力,我们在训练过程中将单词级说话人监督纳入语音编码器。实验结果表明,我们的系统优于强大的基线,包括Sortformer和Meta-Cat在本地设置的音频高达20秒,以及DiarizationLM在长格式音频,尽管是完全端到端和流,而DiarizationLM遵循级联离线管道。据我们所知,这是第一个使用仅在短音频上训练的Speech-LLM在长音频上实现zero-shot可流式传输的联合ASR和日记化的工作,实现了最先进的性能。
摘要:Joint automatic speech recognition (ASR) and speaker diarization aim to answer the question "who spoke what" in multi-speaker scenarios. In this paper, we present an end-to-end speech large language model (Speech-LLM) for Joint strEamable DIarization and aSr (JEDIS-LLM). The model is trained only on short audio under 20s but is capable of streamable inference on long-form audio without additional training. This is achieved by introducing a Speaker Prompt Cache (SPC) with an on-the-fly update mechanism during chunk-wise streaming inference, inspired by the autoregressive nature of LLMs. The SPC also allows the seamless use of pre-enrolled speaker profiles which is common in many scenarios like meeting transcription. To further enhance diarization capability, we incorporate word-level speaker supervision into the speech encoder during training. Experimental results demonstrate that our system outperforms strong baselines, including Sortformer and Meta-Cat in the local setting on audio up to 20s, and DiarizationLM on long-form audio, despite being fully end-to-end and streamable while DiarizationLM follows a cascaded offline pipeline. To the best of our knowledge, this is the first work enabling zero-shot streamable joint ASR and diarization on long audio using a Speech-LLM trained only on short audio, achieving state-of-the-art performance.


【4】A Generalized Weighted Overlap-Add (WOLA) Filter Bank for Improved Subband System Identification
标题:用于改进子带系统识别的广义加权重叠添加(WOLA)过滤器组
链接:https://arxiv.org/abs/2511.15766

作者:Mohit Sharma,Robbe Van Rompaey,Wouter Lanneer,Marc Moonen
备注:For associated MatLab script: https://github.com/mohit-nith/GeneralizedWOLA-SystemIdentification.git
摘要:本文讨论了短时傅立叶变换(STFT)域子带自适应滤波,特别是子带系统辨识的挑战。这一领域的先前研究主要集中在以下采样率进行子带滤波的设置上,该设置使用加权相加(WOLA)滤波器组来实现,该滤波器组因其降低的复杂性而在音频和语音处理中流行。然而,这种传统的方法施加约束的子带滤波器时,变换到他们的全速率表示。本文做出了三个重要贡献。首先,它引入了一个广义的WOLA滤波器组,重新定位子带滤波器之前的下采样操作,消除了传统的WOLA滤波器组中固有的子带滤波器的约束。其次,它研究了均方误差(MSE)性能的广义WOLA滤波器组的全带系统识别,建立子带滤波器的阶数,全带系统的脉冲响应长度,抽取因子,和原型滤波器之间的分析关系。第三,为了解决广义WOLA的计算复杂度增加,本文提出了一种低复杂度的实现,称为每音加权相加(PT-WOLA),它保持了与传统WOLA相当的计算复杂度。理论分析和实验结果表明,所提出的广义WOLA滤波器组显著提高了子带系统辨识的性能。
摘要:This paper addresses the challenges in short-time Fourier transform (STFT) domain subband adaptive filtering, in particular, subband system identification. Previous studies in this area have primarily focused on setups with subband filtering at a downsampled rate, implemented using the weighted overlap-add (WOLA) filter bank, popular in audio and speech-processing for its reduced complexity. However, this traditional approach imposes constraints on the subband filters when transformed to their full-rate representation. This paper makes three key contributions. First, it introduces a generalized WOLA filter bank that repositions subband filters before the downsampling operation, eliminating the constraints on subband filters inherent in the conventional WOLA filter bank. Second, it investigates the mean square error (MSE) performance of the generalized WOLA filter bank for full-band system identification, establishing analytical ties between the order of subband filters, the full-band system impulse response length, the decimation factor, and the prototype filters. Third, to address the increased computational complexity of the generalized WOLA, the paper proposes a low-complexity implementation termed per-tone weighted overlap-add (PT-WOLA), which maintains computational complexity on par with conventional WOLA. Analytical and empirical evidence demonstrates that the proposed generalized WOLA filter bank significantly enhances the performance of subband system identification.


机器翻译由腾讯交互翻译提供,仅供参考