今日论文合集:cs.SD语音6篇,eess.AS音频处理3篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Scalable Music Cover Retrieval Using Lyrics-Aligned Audio Embeddings
标题:使用歌词对齐的音频嵌入的可扩展音乐封面检索
链接:https://arxiv.org/abs/2601.11262

作者:Joanne Affolter,Benjamin Martin,Elena V. Epure,Gabriel Meseguer-Brocal,Frédéric Kaplan
备注:Published at ECIR 2026 (European Conference of Information Retrieval)
摘要:音乐封面检索,也被称为版本识别,旨在识别相同基础音乐作品的不同版本,这是目录管理,版权执法和音乐检索的核心任务。最先进的方法主要集中在和声和旋律特征上,采用越来越复杂的音频管道,这些管道被设计为对音乐属性保持不变,这些音乐属性通常在封面上变化很大。虽然有效,但这些方法需要大量的训练时间和计算资源。相比之下,歌词构成了一个强大的跨封面不变,虽然他们的使用受到了限制,准确和有效地从复调音频提取他们的困难。早期的方法依赖于限制下游性能的简单框架,而最近的系统提供更强大的结果,但需要在复杂的多模式架构中集成大型模型。我们介绍LIVI(歌词知情版本识别),一种方法,旨在平衡检索精度与计算效率。首先,LIVI在训练期间利用最先进的转录和文本嵌入模型的监督,以实现与基于谐波的系统相当或优于基于谐波的系统的检索准确性。其次,LIVI通过去除推理中的转录步骤保持了轻量级和高效,挑战了复杂性高的管道的主导地位。
摘要:Music Cover Retrieval, also known as Version Identification, aims to recognize distinct renditions of the same underlying musical work, a task central to catalog management, copyright enforcement, and music retrieval. State-of-the-art approaches have largely focused on harmonic and melodic features, employing increasingly complex audio pipelines designed to be invariant to musical attributes that often vary widely across covers. While effective, these methods demand substantial training time and computational resources. By contrast, lyrics constitute a strong invariant across covers, though their use has been limited by the difficulty of extracting them accurately and efficiently from polyphonic audio. Early methods relied on simple frameworks that limited downstream performance, while more recent systems deliver stronger results but require large models integrated within complex multimodal architectures. We introduce LIVI (Lyrics-Informed Version Identification), an approach that seeks to balance retrieval accuracy with computational efficiency. First, LIVI leverages supervision from state-of-the-art transcription and text embedding models during training to achieve retrieval accuracy on par with--or superior to--harmonic-based systems. Second, LIVI remains lightweight and efficient by removing the transcription step at inference, challenging the dominance of complexity-heavy pipelines.


【2】FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning
标题:Flash Labs Chroma 1.0:具有个性化语音克隆的实时端到端语音对话模型
链接:https://arxiv.org/abs/2601.11141

作者:Tanyu Chen,Tairan Chen,Kai Shen,Zhenghua Bao,Zhihui Zhang,Man Yuan,Yi Shi
摘要:最近的端到端口语对话系统利用语音标记器和神经音频编解码器来使LLM能够直接对离散语音表示进行操作。然而,这些模型往往表现出有限的说话人身份保护,阻碍个性化的语音交互。在这项工作中,我们提出了Chroma 1.0,这是第一个开源的,实时的,端到端的口语对话模型,它实现了低延迟交互和高保真的个性化语音克隆。Chroma通过支持流媒体生成的交错文本音频令牌调度(1:2)实现亚秒级的端到端延迟,同时在多轮对话中保持高质量的个性化语音合成。我们的实验结果表明,Chroma在说话人相似度方面比人类基线提高了10.96%,实时因子(RTF)为0.43,同时保持了强大的推理和对话能力。我们的代码和模型可在https://github.com/FlashLabs-AI-Corp/FlashLabs-Chroma和https://huggingface.co/FlashLabs/Chroma-4B上公开获取。
摘要:Recent end-to-end spoken dialogue systems leverage speech tokenizers and neural audio codecs to enable LLMs to operate directly on discrete speech representations. However, these models often exhibit limited speaker identity preservation, hindering personalized voice interaction. In this work, we present Chroma 1.0, the first open-source, real-time, end-to-end spoken dialogue model that achieves both low-latency interaction and high-fidelity personalized voice cloning. Chroma achieves sub-second end-to-end latency through an interleaved text-audio token schedule (1:2) that supports streaming generation, while maintaining high-quality personalized voice synthesis across multi-turn conversations. Our experimental results demonstrate that Chroma achieves a 10.96% relative improvement in speaker similarity over the human baseline, with a Real-Time Factor (RTF) of 0.43, while maintaining strong reasoning and dialogue capabilities. Our code and models are publicly available at https://github.com/FlashLabs-AI-Corp/FlashLabs-Chroma and https://huggingface.co/FlashLabs/Chroma-4B .


【3】SonicBench: Dissecting the Physical Perception Bottleneck in Large Audio Language Models
标题:SonicBench:剖析大型音频语言模型中的物理感知瓶颈
链接:https://arxiv.org/abs/2601.11039

作者:Yirong Sun,Yanjun Chen,Xin Qiu,Gang Zhang,Hongyu Chen,Daokuan Wu,Chengming Li,Min Yang,Dawei Zhu,Wei Zhang,Xiaoyu Shen
摘要:大型音频语言模型(LALM)擅长语义和非语言任务,但它们感知音频的基本物理属性(如音高、响度和空间位置)的能力仍有待研究。为了弥合这一差距,我们引入了SonicBench,这是一个基于心理学的基准测试,系统地评估了五个感知维度的12个核心物理属性。与以前的数据集不同,SonicBench使用可控生成工具箱来构建两个互补范式的刺激:识别(绝对判断)和比较(相对判断)。这种设计使我们不仅能够探测感官的精确度,还能探测关系推理能力,这是人类通常表现出更高熟练度的领域。我们的评估揭示了LALM基础听觉理解的实质性缺陷;大多数模型执行近随机猜测,与人类模式相反,在比较任务中未能显示出预期的优势。此外,显式推理产生的收益最小。然而,我们的线性探测分析至关重要地表明,冻结音频编码器确实成功地捕获了这些物理线索(准确率至少为60%),这表明主要瓶颈在于对齐和解码阶段,其中模型未能利用它们已经捕获的感官信号。
摘要:Large Audio Language Models (LALMs) excel at semantic and paralinguistic tasks, yet their ability to perceive the fundamental physical attributes of audio such as pitch, loudness, and spatial location remains under-explored. To bridge this gap, we introduce SonicBench, a psychophysically grounded benchmark that systematically evaluates 12 core physical attributes across five perceptual dimensions. Unlike previous datasets, SonicBench uses a controllable generation toolbox to construct stimuli for two complementary paradigms: recognition (absolute judgment) and comparison (relative judgment). This design allows us to probe not only sensory precision but also relational reasoning capabilities, a domain where humans typically exhibit greater proficiency. Our evaluation reveals a substantial deficiency in LALMs' foundational auditory understanding; most models perform near random guessing and, contrary to human patterns, fail to show the expected advantage on comparison tasks. Furthermore, explicit reasoning yields minimal gains. However, our linear probing analysis demonstrates crucially that frozen audio encoders do successfully capture these physical cues (accuracy at least 60%), suggesting that the primary bottleneck lies in the alignment and decoding stages, where models fail to leverage the sensory signals they have already captured.


【4】WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
标题:WenetSpeech-Wu:统一汉语吴方言语音处理生态系统的数据集、基准和模型
链接:https://arxiv.org/abs/2601.11027

作者:Chengyou Wang,Mingchen Shao,Jingbin Hu,Zeyu Zhu,Hongfei Xue,Bingshen Mu,Xin Xu,Xingyi Duan,Binbin Zhang,Pengcheng Zhu,Chuang Ding,Xiaojun Zhang,Hui Bu,Lei Xie
摘要:低资源方言的语音处理仍然是开发包容性和鲁棒性语音技术的基本挑战。尽管吴方言在语言学上具有重要意义,使用者众多,但由于缺乏大规模的语音数据、标准化的评估基准和公开可用的模型,吴方言的研究长期受到阻碍。在这项工作中,我们提出了WenetSpeech-Wu,第一个大规模的,多维注释的开源语音语料库的吴方言,包括大约8,000小时的不同的语音数据。在此数据集的基础上,我们介绍了WenetSpeech-Wu-Bench,这是第一个标准化和公开访问的基准,用于系统评估吴方言语音处理,涵盖自动语音识别(ASR),吴到普通话翻译,说话人属性预测,语音情感识别,文本到语音(TTS)合成,以及解释遵循TTS(指令TTS)。此外,我们发布了一套在WenetSpeech-Wu上训练的强大开源模型,在多个任务中建立了具有竞争力的性能,并通过经验验证了所提出的数据集的有效性。总之,这些贡献为全面的吴方言语音处理生态系统奠定了基础,我们开源了建议的数据集,基准和模型,以支持未来对方言语音智能的研究。
摘要:Speech processing for low-resource dialects remains a fundamental challenge in developing inclusive and robust speech technologies. Despite its linguistic significance and large speaker population, the Wu dialect of Chinese has long been hindered by the lack of large-scale speech data, standardized evaluation benchmarks, and publicly available models. In this work, we present WenetSpeech-Wu, the first large-scale, multi-dimensionally annotated open-source speech corpus for the Wu dialect, comprising approximately 8,000 hours of diverse speech data. Building upon this dataset, we introduce WenetSpeech-Wu-Bench, the first standardized and publicly accessible benchmark for systematic evaluation of Wu dialect speech processing, covering automatic speech recognition (ASR), Wu-to-Mandarin translation, speaker attribute prediction, speech emotion recognition, text-to-speech (TTS) synthesis, and instruction-following TTS (instruct TTS). Furthermore, we release a suite of strong open-source models trained on WenetSpeech-Wu, establishing competitive performance across multiple tasks and empirically validating the effectiveness of the proposed dataset. Together, these contributions lay the foundation for a comprehensive Wu dialect speech processing ecosystem, and we open-source proposed datasets, benchmarks, and models to support future research on dialectal speech intelligence.


【5】Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
标题:用自回归变换器统一语音识别、合成和转换
链接:https://arxiv.org/abs/2601.10770

作者:Runyuan Cai,Yu Lin,Yiming Wang,Chunlin Fu,Xiaodong Zeng
摘要:传统的语音系统通常依赖于单独的、特定于任务的文本到语音(TTS)、自动语音识别(ASR)和语音转换(VC)模型,从而导致分散的管道限制了可扩展性、效率和跨任务泛化。在本文中,我们提出了通用音频(GPA),一个统一的音频基础模型,在一个单一的大语言模型(LLM)架构中集成了多个核心语音任务。GPA在共享的离散音频令牌空间上运行,并支持推理驱动的任务诱导,使单个自回归模型能够灵活地执行TTS,ASR和VC,而无需修改架构。这种统一的设计结合了离散语音令牌上的完全自回归公式、跨语音域的联合多任务训练以及可实现高并发性和吞吐量的可扩展推理管道。由此产生的模型系列支持高效的多尺度部署,包括针对边缘和资源受限环境优化的轻量级0.3B参数变体。总之,这些设计选择表明,统一的自回归架构可以在不同的语音任务中实现有竞争力的性能,同时保持低延迟,实际部署的可行性。
摘要:Traditional speech systems typically rely on separate, task-specific models for text-to-speech (TTS), automatic speech recognition (ASR), and voice conversion (VC), resulting in fragmented pipelines that limit scalability, efficiency, and cross-task generalization. In this paper, we present General-Purpose Audio (GPA), a unified audio foundation model that integrates multiple core speech tasks within a single large language model (LLM) architecture. GPA operates on a shared discrete audio token space and supports instruction-driven task induction, enabling a single autoregressive model to flexibly perform TTS, ASR, and VC without architectural modifications. This unified design combines a fully autoregressive formulation over discrete speech tokens, joint multi-task training across speech domains, and a scalable inference pipeline that achieves high concurrency and throughput. The resulting model family supports efficient multi-scale deployment, including a lightweight 0.3B-parameter variant optimized for edge and resource-constrained environments. Together, these design choices demonstrate that a unified autoregressive architecture can achieve competitive performance across diverse speech tasks while remaining viable for low-latency, practical deployment.


【6】DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion
标题:DSA-令牌化器:通过基于流匹配的分层融合来解开语义-声学令牌化
链接:https://arxiv.org/abs/2601.09239

作者:Hanlin Zhang,Daxin Tan,Dehua Tao,Xiao Chen,Haochen Tan,Yunhe Li,Yuchen Cao,Jianping Wang,Linqi Song
备注:Submit to ACL ARR 2026 Jaunary
摘要:语音标记器是离散语音大语言模型(Speech LLM)的基石。现有的分词器要么优先进行语义编码,要么将语义内容与声学风格不可分割地融合在一起,要么实现了不完全的语义-声学分离。为了实现更好的解纠缠,我们提出了DSA-Tokenizer,它通过不同的优化约束明确地将语音解纠缠成离散的语义和声学令牌。具体而言,语义令牌由ASR监督以捕获语言内容,而声学令牌专注于梅尔频谱图恢复以编码风格。为了消除两个序列之间的刚性长度约束,我们引入了一个分层流匹配解码器,进一步提高语音生成质量。此外,我们采用联合重建重组训练策略来加强这种分离。DSA标记器通过强大的解纠缠实现高保真重建和灵活重组,促进语音LLM中的可控生成。我们的分析强调了解开标记作为未来语音建模的关键范式。音频样本可在https://anonymous.4open.science/w/DSA_Tokenizer_demo/上获得。该代码和模型将在论文被接受后公开提供。
摘要:Speech tokenizers serve as the cornerstone of discrete Speech Large Language Models (Speech LLMs). Existing tokenizers either prioritize semantic encoding, fuse semantic content with acoustic style inseparably, or achieve incomplete semantic-acoustic disentanglement. To achieve better disentanglement, we propose DSA-Tokenizer, which explicitly disentangles speech into discrete semantic and acoustic tokens via distinct optimization constraints. Specifically, semantic tokens are supervised by ASR to capture linguistic content, while acoustic tokens focus on mel-spectrograms restoration to encode style. To eliminate rigid length constraints between the two sequences, we introduce a hierarchical Flow-Matching decoder that further improve speech generation quality. Furthermore, We employ a joint reconstruction-recombination training strategy to enforce this separation. DSA-Tokenizer enables high fidelity reconstruction and flexible recombination through robust disentanglement, facilitating controllable generation in speech LLMs. Our analysis highlights disentangled tokenization as a pivotal paradigm for future speech modeling. Audio samples are avaialble at https://anonymous.4open.science/w/DSA_Tokenizer_demo/. The code and model will be made publicly available after the paper has been accepted.


eess.AS音频处理


【1】FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning
标题:Flash Labs Chroma 1.0:具有个性化语音克隆的实时端到端语音对话模型
链接:https://arxiv.org/abs/2601.11141

作者:Tanyu Chen,Tairan Chen,Kai Shen,Zhenghua Bao,Zhihui Zhang,Man Yuan,Yi Shi
摘要:最近的端到端口语对话系统利用语音标记器和神经音频编解码器来使LLM能够直接对离散语音表示进行操作。然而,这些模型通常表现出有限的说话者身份保留,阻碍了个性化语音交互。在这项工作中,我们提出了Chroma 1.0,这是第一个开源的,实时的,端到端的口语对话模型,它实现了低延迟交互和高保真的个性化语音克隆。Chroma通过支持流媒体生成的交错文本音频令牌调度(1:2)实现亚秒级的端到端延迟,同时在多轮对话中保持高质量的个性化语音合成。我们的实验结果表明,Chroma在说话人相似度方面比人类基线提高了10.96%,实时因子(RTF)为0.43,同时保持了强大的推理和对话能力。我们的代码和模型可在https://github.com/FlashLabs-AI-Corp/FlashLabs-Chroma和https://huggingface.co/FlashLabs/Chroma-4B上公开获取。
摘要:Recent end-to-end spoken dialogue systems leverage speech tokenizers and neural audio codecs to enable LLMs to operate directly on discrete speech representations. However, these models often exhibit limited speaker identity preservation, hindering personalized voice interaction. In this work, we present Chroma 1.0, the first open-source, real-time, end-to-end spoken dialogue model that achieves both low-latency interaction and high-fidelity personalized voice cloning. Chroma achieves sub-second end-to-end latency through an interleaved text-audio token schedule (1:2) that supports streaming generation, while maintaining high-quality personalized voice synthesis across multi-turn conversations. Our experimental results demonstrate that Chroma achieves a 10.96% relative improvement in speaker similarity over the human baseline, with a Real-Time Factor (RTF) of 0.43, while maintaining strong reasoning and dialogue capabilities. Our code and models are publicly available at https://github.com/FlashLabs-AI-Corp/FlashLabs-Chroma and https://huggingface.co/FlashLabs/Chroma-4B .


【2】Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
标题:用自回归变换器统一语音识别、合成和转换
链接:https://arxiv.org/abs/2601.10770

作者:Runyuan Cai,Yu Lin,Yiming Wang,Chunlin Fu,Xiaodong Zeng
摘要:传统的语音系统通常依赖于单独的、特定于任务的文本到语音(TTS)、自动语音识别(ASR)和语音转换(VC)模型,从而导致分散的管道限制了可扩展性、效率和跨任务泛化。在本文中,我们提出了通用音频(GPA),一个统一的音频基础模型,在一个单一的大语言模型(LLM)架构中集成了多个核心语音任务。GPA在共享的离散音频令牌空间上运行,并支持推理驱动的任务诱导,使单个自回归模型能够灵活地执行TTS,ASR和VC,而无需修改架构。这种统一的设计结合了离散语音令牌上的完全自回归公式、跨语音域的联合多任务训练以及可实现高并发性和吞吐量的可扩展推理管道。由此产生的模型系列支持高效的多尺度部署,包括针对边缘和资源受限环境优化的轻量级0.3B参数变体。总之,这些设计选择表明,统一的自回归架构可以在不同的语音任务中实现有竞争力的性能,同时保持低延迟,实际部署的可行性。
摘要:Traditional speech systems typically rely on separate, task-specific models for text-to-speech (TTS), automatic speech recognition (ASR), and voice conversion (VC), resulting in fragmented pipelines that limit scalability, efficiency, and cross-task generalization. In this paper, we present General-Purpose Audio (GPA), a unified audio foundation model that integrates multiple core speech tasks within a single large language model (LLM) architecture. GPA operates on a shared discrete audio token space and supports instruction-driven task induction, enabling a single autoregressive model to flexibly perform TTS, ASR, and VC without architectural modifications. This unified design combines a fully autoregressive formulation over discrete speech tokens, joint multi-task training across speech domains, and a scalable inference pipeline that achieves high concurrency and throughput. The resulting model family supports efficient multi-scale deployment, including a lightweight 0.3B-parameter variant optimized for edge and resource-constrained environments. Together, these design choices demonstrate that a unified autoregressive architecture can achieve competitive performance across diverse speech tasks while remaining viable for low-latency, practical deployment.


【3】DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion
标题:DSA-令牌化器:通过基于流匹配的分层融合来解开语义-声学令牌化
链接:https://arxiv.org/abs/2601.09239

作者:Hanlin Zhang,Daxin Tan,Dehua Tao,Xiao Chen,Haochen Tan,Yunhe Li,Yuchen Cao,Jianping Wang,Linqi Song
备注:Submit to ACL ARR 2026 Jaunary
摘要:语音标记器是离散语音大语言模型(Speech LLM)的基石。现有的分词器要么优先进行语义编码,要么将语义内容与声学风格不可分割地融合在一起,要么实现了不完全的语义-声学分离。为了实现更好的解纠缠,我们提出了DSA-Tokenizer,它通过不同的优化约束明确地将语音解纠缠成离散的语义和声学令牌。具体而言,语义令牌由ASR监督以捕获语言内容,而声学令牌专注于梅尔频谱图恢复以编码风格。为了消除两个序列之间的刚性长度约束,我们引入了一个分层流匹配解码器,进一步提高语音生成质量。此外,我们采用联合重建重组训练策略来加强这种分离。DSA标记器通过强大的解纠缠实现高保真重建和灵活重组,促进语音LLM中的可控生成。我们的分析强调了解开标记作为未来语音建模的关键范式。音频样本可在https://anonymous.4open.science/w/DSA_Tokenizer_demo/上获得。该代码和模型将在论文被接受后公开提供。
摘要:Speech tokenizers serve as the cornerstone of discrete Speech Large Language Models (Speech LLMs). Existing tokenizers either prioritize semantic encoding, fuse semantic content with acoustic style inseparably, or achieve incomplete semantic-acoustic disentanglement. To achieve better disentanglement, we propose DSA-Tokenizer, which explicitly disentangles speech into discrete semantic and acoustic tokens via distinct optimization constraints. Specifically, semantic tokens are supervised by ASR to capture linguistic content, while acoustic tokens focus on mel-spectrograms restoration to encode style. To eliminate rigid length constraints between the two sequences, we introduce a hierarchical Flow-Matching decoder that further improve speech generation quality. Furthermore, We employ a joint reconstruction-recombination training strategy to enforce this separation. DSA-Tokenizer enables high fidelity reconstruction and flexible recombination through robust disentanglement, facilitating controllable generation in speech LLMs. Our analysis highlights disentangled tokenization as a pivotal paradigm for future speech modeling. Audio samples are avaialble at https://anonymous.4open.science/w/DSA_Tokenizer_demo/. The code and model will be made publicly available after the paper has been accepted.


机器翻译由腾讯交互翻译提供,仅供参考