今日论文合集:cs.SD语音8篇,eess.AS音频处理10篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】PianoCoRe: Combined and Refined Piano MIDI Dataset
标题:PianoCoRe:组合和改进的钢琴收件箱数据集
链接:https://arxiv.org/abs/2605.06627
作者:Ilya Borovik
备注:Published in TISMIR. Project repository: https://github.com/ilya16/PianoCoRe
摘要:具有匹配的分数和性能的符号音乐数据集对于许多音乐信息检索(MIR)任务是必不可少的。然而,现有的资源往往涵盖范围狭窄的作曲家,缺乏性能多样性,省略音符级对齐,或使用不一致的命名格式。这项工作介绍了PianoCore,这是一个大规模的钢琴数据集,它统一和改进了主要的开源钢琴语料库。该数据集包含由483位作曲家创作的5,625首作品的250,046场演出,总计21,763小时的演出音乐。PianoCore以分层子集形式发布,以支持不同的应用程序:从大规模分析和预训练(PianoCore-C和去重复PianoCore-B)到具有音符级别分数对齐的表现力表现建模(PianoCore-A/A*)。音符对齐子集PianoCore-A提供了迄今为止最大的157,207个演奏的开源集合,与1,591个分数对齐。除了数据集之外,贡献还有:(1)用于检测损坏的和类似分数的transmittance的质量分类器,以及(2)RAScop,一个清除时间对齐错误并插入缺失音符的对齐细化管道。分析表明,改进减少了时间噪声,消除了节奏离群值。此外,与在原始或较小数据集上训练的模型相比,在PianoCoRe上训练的表现力表现出色的渲染模型表现出对不可见片段的鲁棒性。PianoCore为下一代表现力钢琴演奏研究提供了一个现成的基础。
摘要:Symbolic music datasets with matched scores and performances are essential for many music information retrieval (MIR) tasks. Yet, existing resources often cover a narrow range of composers, lack performance variety, omit note-level alignments, or use inconsistent naming formats. This work presents PianoCoRe, a large-scale piano MIDI dataset that unifies and refines major open-source piano corpora. The dataset contains 250,046 performances of 5,625 pieces written by 483 composers, totaling 21,763 h of performed music. PianoCoRe is released in tiered subsets to support different applications: from large-scale analysis and pre-training (PianoCoRe-C and deduplicated PianoCoRe-B) to expressive performance modeling with note-level score alignment (PianoCoRe-A/A*). The note-aligned subset, PianoCoRe-A, provides the largest open-source collection of 157,207 performances aligned to 1,591 scores to date. In addition to the dataset, the contributions are: (1) a MIDI quality classifier for detecting corrupted and score-like transcriptions and (2) RAScoP, an alignment refinement pipeline that cleans temporal alignment errors and interpolates missing notes. The analysis shows that the refinement reduces temporal noise and eliminates tempo outliers. Moreover, an expressive performance rendering model trained on PianoCoRe demonstrates improved robustness to unseen pieces compared to models trained on raw or smaller datasets. PianoCoRe provides a ready-to-use foundation for the next generation of expressive piano performance research.


【2】PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization

标题:PairAlign:一种基于自对齐的序列标记化框架及其在音频标记化中的应用
链接:https://arxiv.org/abs/2605.06582
作者:Adhiraj Banerjee,Vipul Arora
备注:101 pages, 7 Figures, pre-print, Under Review
摘要:对感觉数据的许多操作--比较、记忆、检索和推理--自然地通过离散的符号结构来表达。在语言中,这个接口是由标记给出的;在音频中,它必须学习。现有的音频标记器依赖于量化、聚类或编解码器重构,在本地分配标记,因此序列一致性、紧凑性、长度控制、终止和编辑相似性很少直接优化。   我们介绍PairAlign,这是一个通过序列级自对齐进行紧凑音频标记化的框架。PairAlign将令牌化视为条件序列生成:编码器将语音映射到连续条件,自回归解码器从BOS生成令牌,学习令牌身份,顺序,长度和EOS放置。给定两个内容保留视图,每个视图的序列都被训练成可能在另一个视图的表示下,而不相关的示例提供竞争序列。这为编辑距离保留提供了一个可伸缩的代理,同时阻止了多对一的崩溃。   PairAlign从VQ风格的标记化开始,并通过EMA教师目标,交叉配对教师强制,前缀损坏,可能性对比和长度控制对其进行改进。   在3秒的语音中,PairAlign学习紧凑的非退化序列,具有广泛的词汇使用和强大的跨视图一致性。在TIMIT检索上,它保留了编辑距离搜索,同时将归档令牌数量减少了55%。一个连续扫描探头显示较低的局部重叠比密集的几何标记,但更强的长度控制和有界编辑轨迹下100毫秒的移位。PairAlign是一个序列符号预测学习器:像JEPA风格的目标一样,它从另一个角度预测抽象目标,作为学习的可变长度符号序列,而不是连续的潜在目标。
摘要:Many operations on sensory data -- comparison, memory, retrieval, and reasoning -- are naturally expressed over discrete symbolic structures. In language this interface is given by tokens; in audio, it must be learned. Existing audio tokenizers rely on quantization, clustering, or codec reconstruction, assigning tokens locally, so sequence consistency, compactness, length control, termination, and edit similarity are rarely optimized directly.   We introduce PairAlign, a framework for compact audio tokenization through sequence-level self-alignment. PairAlign treats tokenization as conditional sequence generation: an encoder maps speech to a continuous condition, and an autoregressive decoder generates tokens from BOS, learning token identity, order, length, and EOS placement. Given two content-preserving views, each view's sequence is trained to be likely under the other's representation, while unrelated examples provide competing sequences. This gives a scalable surrogate for edit-distance preservation while discouraging many-to-one collapse.   PairAlign starts from VQ-style tokenization and refines it with EMA-teacher targets, cross-paired teacher forcing, prefix corruption, likelihood contrast, and length control.   On 3-second speech, PairAlign learns compact, non-degenerate sequences with broad vocabulary usage and strong cross-view consistency. On TIMIT retrieval, it preserves edit-distance search while reducing archive token count by 55%. A continuous-sweep probe shows lower local overlap than a dense geometric tokenizer, but stronger length control and bounded edit trajectories under 100 ms shifts. PairAlign is a sequence-symbolic predictive learner: like JEPA-style objectives, it predicts an abstract target from another view as a learned variable-length symbolic sequence, not a continuous latent.


【3】Quantum Kernels for Audio Deepfake Detection Using Spectrogram Patch Features

标题:使用频谱图补丁特征进行音频深度伪造检测的量子核
链接:https://arxiv.org/abs/2605.06035
作者:Lisan Al Amin,Rakib Hossain,Mahbubul Islam,Faisal Quader,Thanh Thi Nguyen
摘要:量子机器学习已经成为模式识别的一种有前途的工具,但许多以音频为重点的方法仍然将频谱图视为通用图像,并且没有明确地利用它们的时频结构。我们提出了Q-Patch,这是一种为音频量身定制的量子特征图,它使用具有邻接感知纠缠的浅硬件高效电路将来自梅尔频谱图的本地时频补丁编码为量子态。每个选定的补丁总结了一个紧凑的四维声学描述符,并映射到一个四量子位电路的深度最多为3,使实际的量子内核的建设在近期的限制。我们评估Q-Patch的音频欺骗检测任务,使用一个控制,平衡的协议,并将其与大小匹配的经典基线。Q-Patch提高了真实样本和欺骗样本之间的区分度,实现了0.87的接收器工作特征曲线(AUROC)下的面积,而在相同的补丁级别特征上训练的径向基函数支持向量机(RBF-SVM)为0.82。核空间分析进一步揭示了一个清晰的类结构,类间相似性约为0.615,类内自相似性为1.00。总的来说,Q-Patch提供了一个实用的框架,用于将时频感知表示纳入量子内核学习,以在低资源环境中进行音频真实性评估。
摘要:Quantum machine learning has emerged as a promising tool for pattern recognition, yet many audio-focused approaches still treat spectrograms as generic images and do not explicitly exploit their time-frequency structure. We propose Q-Patch, a quantum feature map tailored to audio that encodes local time-frequency patches from mel-spectrograms into quantum states using shallow, hardware-efficient circuits with adjacency-aware entanglement. Each selected patch is summarized by a compact four-dimensional acoustic descriptor and mapped to a four-qubit circuit with depth at most three, enabling practical quantum kernel construction under near-term constraints. We evaluate Q-Patch on an audio spoofing detection task using a controlled, balanced protocol and compare it with size-matched classical baselines. Q-Patch improves discrimination between bona fide and spoofed samples, achieving an area under the receiver operating characteristic curve (AUROC) of 0.87, compared with 0.82 for a radial basis function support vector machine (RBF-SVM) trained on the same patch-level features. Kernel-space analysis further reveals a clear class structure, with cross-class similarity around 0.615 and within-class self-similarity of 1.00. Overall, Q-Patch provides a practical framework for incorporating time-frequency-aware representations into quantum kernel learning for audio authenticity assessment in low-resource settings.


【4】Do Melody and Rhythm Coevolve?

标题:旋律和节奏相互进化吗?
链接:https://arxiv.org/abs/2605.05982
作者:Harin Lee,Rainer Polak,Manuel Anglada-Tort,Marc Schönwiesner,Minsu Park,Nori Jacoby
备注:6 pages, 3 figures, to be included in Proceedings of the Annual Meeting of the Cognitive Science Society
摘要:音乐包括两个核心结构组成部分,旋律和节奏,在不同文化中差异很大。这些成分是以耦合的方式共同进化,还是遵循独立的轨迹,目前尚不清楚。我们引入了一种新的计算管道,从59个国家的27,628首流行歌曲中提取声乐旋律音高间隔和重复的起始时间分布,从而实现了绕过传统音乐注释的大规模跨文化比较。国家之间的音乐相似性与地理和语言关系一致,验证了我们的方法。各国的旋律和节奏结构都出现了很大的变化,但这两个组成部分的多样性并没有显着的相关性,这对耦合进化的假设提出了挑战。只有节奏的多样性与种族和语言的异质性显着相关,而旋律的多样性没有表现出这样的关联。这些发现表明,旋律和节奏构成了部分独立的系统,由不同的文化和进化压力塑造,而不是单一的整体音乐风格的组成部分。
摘要:Music comprises two core structural components, melody and rhythm, that vary widely across cultures. Whether these components coevolve in a coupled way or follow independent trajectories remains unclear. We introduce a novel computational pipeline to extract vocal melodic pitch-interval and percussive inter-onset timing distributions from 27,628 popular songs across 59 countries, enabling large-scale cross-cultural comparison that bypasses traditional music annotations. Musical similarities between countries aligned with geographic and linguistic relationships, validating our approach. Substantial variation emerged in both melodic and rhythmic structures across countries, yet the diversity of the two components was not significantly correlated, challenging assumptions of coupled evolution. Only rhythmic diversity was significantly associated with ethnic and linguistic heterogeneity, while melodic diversity showed no such association. These findings suggest that melody and rhythm constitute partially independent systems shaped by distinct cultural and evolutionary pressures, rather than components of a single monolithic musical style.


【5】Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

标题:从输入端最大限度地减少情态差距:您的语音LLM可以是韵律感知文本LLM
链接:https://arxiv.org/abs/2605.05927
作者:Wenqian Cui,Xiao-Hui Li,Daxin Tan,Qiyong Zheng,Irwin King
备注:Work in progress
摘要:语音大语言模型(SLMs)通常是从文本大语言模型(TLM)检查点构建的,但它们仍然存在很大的模态差距。先前的工作主要试图通过使语音生成更像文本来从输出侧减少这种差距,但差距仍然存在。我们认为,关键的剩余瓶颈在于输入侧。我们提出了TextPro-SLM,一个SLM,使口语输入更接近的韵律感知文本LLM。TextPro-SLM结合了WhisperPro,一个统一的语音编码器,产生同步的文本标记和韵律嵌入,与一个LLM骨干训练,以保留原始TLM的语义能力,同时学习非语言理解。实验表明,TextPro-SLM在3B和7 B尺度上都实现了领先SLM中最低的模态差距,同时在非语言理解任务上也提供了强大的整体性能。这些增益仅用大约1,000小时的LLM训练音频就实现了,这表明从输入端减少模态差距既有效又有效。
摘要:Speech large language models (SLMs) are typically built from text large language model (TLM) checkpoints, yet they still suffer from a substantial modality gap. Prior work has mainly attempted to reduce this gap from the output side by making speech generation more text-like, but the gap remains. We argue that the key remaining bottleneck lies on the input side. We propose TextPro-SLM, an SLM that makes spoken input more closely resemble that of a prosody-aware text LLM. TextPro-SLM combines WhisperPro, a unified speech encoder that produces synchronized text tokens and prosody embeddings, with an LLM backbone trained to preserve the semantic capabilities of the original TLM while learning paralinguistic understanding. Experiments show that TextPro-SLM achieves the lowest modality gap among leading SLMs at both 3B and 7B scales, while also delivering strong overall performance on paralinguistic understanding tasks. These gains are achieved with only roughly 1,000 hours of LLM training audio, suggesting that reducing the modality gap from the input side is both effective and data-efficient.


【6】X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

标题:X-Voice:通过Zero-Shot跨语言语音克隆使每个人都能说30种语言
链接:https://arxiv.org/abs/2605.05611
作者:Rixi Xu,Qingyu Liu,Haitao Li,Yushen Chen,Zhikang Niu,Yunting Yang,Jian Zhao,Ke Li,Berrak Sisman,Qinyuan Cheng,Xipeng Qiu,Kai Yu,Xie Chen
备注:16 pages, 4 figures, 9 tables
摘要:在本文中,我们提出了X-Voice,一个0.4B的多语种zero-shot语音克隆模型,克隆任意的声音,使每个人都能说30种语言。X-Voice在420 K小时的多语言语料库上进行训练,使用国际音标(IPA)作为统一表示。为了消除对提示文本的依赖,而不需要像强制对齐这样复杂的预处理,我们设计了一个两阶段的训练范例。在第一阶段,我们通过标准的条件流匹配训练建立X-Voice$_{\text{s1}}$,并使用它来合成10 K小时的说话者一致性片段作为音频提示。在阶段2中,我们对这些音频对进行微调,并屏蔽提示文本以获得X-Voice$_{\text{s2}}$,这使得zero-shot语音克隆不需要音频提示的转录。在架构上,我们通过实现语言标识符的双层注入和无分类器指导的解耦和调度来扩展F5-TTS,以促进多语言语音合成。主观和客观的评估结果表明,X-Voice优于现有的流匹配的多语言系统,如LEMAS-TTS和实现zero-shot跨语言克隆能力媲美十亿规模的模型,如Qwen 3-TTS。为了促进研究透明度和社区发展,我们开放了所有相关资源。
摘要:In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to speak 30 languages. X-Voice is trained on a 420K-hour multilingual corpus using the International Phonetic Alphabet (IPA) as a unified representation. To eliminate the reliance on prompt text without complex preprocessing like forced alignment, we design a two-stage training paradigm. In Stage 1, we establish X-Voice$_{\text{s1}}$ through standard conditional flow-matching training and use it to synthesize 10K hours of speaker-consistent segments as audio prompts. In Stage 2, we fine-tune on these audio pairs with prompt text masked to derive X-Voice$_{\text{s2}}$, which enables zero-shot voice cloning without requiring transcripts of audio prompts. Architecturally, we extend F5-TTS by implementing a dual-level injection of language identifiers and decoupling and scheduling of Classifier-Free Guidance to facilitate multilingual speech synthesis. Subjective and objective evaluation results demonstrate that X-Voice outperforms existing flow-matching based multilingual systems like LEMAS-TTS and achieves zero-shot cross-lingual cloning capabilities comparable to billion-scale models such as Qwen3-TTS. To facilitate research transparency and community advancement, we open-source all related resources.


【7】Optimal Transport Audio Distance with Learned Riemannian Ground Metrics

标题:通过习得的Riemann地面搜索器实现最佳传输音频距离
链接:https://arxiv.org/abs/2605.05554
作者:Wonwoo Jeong
备注:21 pages, 4 figures, 10 tables. The otadtk toolkit is available at https://github.com/wonwoo-jeong/otadtk
摘要:在音频生成评估中,Fréchet Audio Distance(FAD)是一个2-Wasserstein距离,对两个原语都有结构约束:成本是冻结的嵌入拉回,其不变性集隐藏了严重的伪影,耦合是高斯拟合,相对于离散OT稀释了秩-1污染。我们提出了最佳传输音频距离(OSTO),它纠正每个原语与一个专用的机制-一个剩余的黎曼地面度量适配器的成本和熵Sinkhorn最佳运输的耦合。在四轴协议下的八个编码器中,仅在ε= 0.05$处进行耦合比较,表明Sinkhorn的秩1灵敏度超过FAD的1.9至3.6倍。此外,OTAD与音频质量MOS的平均Spearman相关性(DCASE 2023任务7)比基线指标更高。作为离散传输计划的内在好处,OSCOM产生每个样本的诊断与AUROC $\ge 0.86$,标量或内核聚合的指标结构上缺乏的能力。
摘要:In audio generation evaluation, Fréchet Audio Distance (FAD) is a 2-Wasserstein distance with structural constraints for both primitives: the cost is a frozen embedding pullback whose invariance set hides severe artifacts, and the coupling is a Gaussian fit that dilutes rank-1 contamination relative to discrete OT. We propose Optimal Transport Audio Distance (OTAD), which corrects each primitive with one dedicated mechanism -- a residual Riemannian ground-metric adapter for the cost and entropic Sinkhorn optimal transport for the coupling. Across eight encoders under a four-axis protocol, coupling-only comparisons at $ε= 0.05$ show that Sinkhorn's rank-1 sensitivity exceeds FAD's by a factor of 1.9 to 3.6. Furthermore, OTAD achieves a higher mean Spearman correlation with audio-quality MOS (DCASE 2023 Task 7) than baseline metrics. As an intrinsic benefit of the discrete transport plan, OTAD yields per-sample diagnostics with AUROC $\ge 0.86$, a capability that scalar- or kernel-aggregated metrics structurally lack.


【8】Prompting Whisper for Joint Speech Transcription and Diarization

标题:联合语音转录和数字化的签名耳语
链接:https://arxiv.org/abs/2605.05231
作者:Mariia Zamyrova,Henk van den Heuvel
备注:To be presented at the Joint Workshop on HSCMA and CHiME 2026
摘要:作为MediSpeech项目的一部分,我们的目标是开发一个系统,实时转录和记录医生和患者之间的荷兰语对话。在这项研究中(正在进行中),我们探索有效地结合耳语与扬声器日记(SD)的方法。在尝试用包含扬声器标签的文本提示Whisper之后,我们观察到它能够以有希望的准确性将标签插入到转录中。我们通过微调Whisper与扬声器标记的提示来继续这一系列的研究,以类似于序列化输出训练(SOT)的格式生成transmittance。微调Whisper在长格式音频块中产生了更一致的扬声器ID,并改进了逐字转录。该研究发现了新的挑战,因为Whisper的SD性能受到了错误的影响,这些错误通过提示和分配给重叠语音的不准确时间戳传播。
摘要:As part of the MediSpeech project, we aim to develop a system that transcribes and diarizes Dutch conversations between doctors and patients in real-time. In this research (in-progress) we explore ways of efficiently combining Whisper with speaker diarization (SD). After trying to prompt Whisper with text that contains speaker labels, we observed that it is able to insert labels into the transcription with promising accuracy. We continued this line of research by fine-tuning Whisper with speaker-labelled prompts to generate transcriptions in a format similar to that of Serialized Output Training (SOT). Fine-tuning Whisper yielded more consistent speaker IDs across the chunks of long-form audio and improved verbatim transcription. The study uncovered new challenges as Whisper's SD performance suffers because of mistakes that get propagated through prompts and inaccurate timestamps assigned to overlapping speech.


eess.AS音频处理


【1】Task-Aware Answer Preservation under Audio Compression for Large Audio Language Models
标题:大型音频语言模型音频压缩下的任务感知答案保留
链接:https://arxiv.org/abs/2605.06631
作者:Amir Ivry
备注:Preprint
摘要:
摘要:


【2】LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation

标题:LiVeAction:一个轻量级、通用和非对称的实时神经编解码器设计
链接:https://arxiv.org/abs/2605.06628
作者:Dan Jacobellis, Neeraja J. Yadwadkar
备注:DCC 2026
摘要:
摘要:


【3】WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

标题:WavCube:通过语义-声学联合建模统一语音表示以实现理解和生成
链接:https://arxiv.org/abs/2605.06407
作者:Guanrou Yang, Tian Tan, Qian Chen, Zhikang Niu, Yakun Song, Ziyang Ma, Yushen Chen, Zeyu Xie, Tianrui Wang, Yifan Yang, Wenxi Chen, Qi Chen, Wenrui Liu, Shan Yang, Xie Chen
摘要:
摘要:


【4】Predictive-Generative Drift Decomposition for Speech Enhancement and Separation

标题:预测-生成漂移分解语音增强与分离
链接:https://arxiv.org/abs/2605.06189
作者:Julius Richter, Yoshiki Masuyama, Christoph Boeddeker, Takahiro Edo, Gordon Wichern, Jonathan Le Roux
备注:Submitted to NeurIPS 2026
摘要:
摘要:


【5】NDF+: Joint Neural Directional Filtering and Diffuse Sound Extraction

标题:NDF+:联合神经方向过滤和漫音提取
链接:https://arxiv.org/abs/2605.06108
作者:Weilong Huang, Le Nhat Tam Huynh, Oliver Thiergart, Emanuël A. P. Habets
摘要:
摘要:


【6】Optimal Transport Audio Distance with Learned Riemannian Ground Metrics

标题:通过习得的Riemann地面搜索器实现最佳传输音频距离
链接:https://arxiv.org/abs/2605.05554
作者:Wonwoo Jeong
备注:21 pages, 4 figures, 10 tables. The otadtk toolkit is available at this https URL
摘要:
摘要:


【7】Prompting Whisper for Joint Speech Transcription and Diarization

标题:联合语音转录和数字化的签名耳语
链接:https://arxiv.org/abs/2605.05231
作者:Mariia Zamyrova, Henk van den Heuvel
备注:To be presented at the Joint Workshop on HSCMA and CHiME 2026
摘要:
摘要:


【8】Weight-Decay Turns Transformer Loss Landscapes Villani: Functional-Analytic Foundations for Optimization and Generalization

标题:重量衰减匝数Transformer损耗景观:优化和推广的泛函分析基础
链接:https://arxiv.org/abs/2605.06599
作者:Abhijit Das, Sayantan Dutta
备注:17 pages, 10 figures
摘要:
摘要:


【9】Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

标题:从输入端最大限度地减少情态差距:您的语音LLM可以是韵律感知文本LLM
链接:https://arxiv.org/abs/2605.05927
作者:Wenqian Cui,Xiao-Hui Li,Daxin Tan,Qiyong Zheng,Irwin King
备注:Work in progress
摘要:语音大语言模型(SLMs)通常是从文本大语言模型(TLM)检查点构建的,但它们仍然存在很大的模态差距。先前的工作主要试图通过使语音生成更像文本来从输出侧减少这种差距,但差距仍然存在。我们认为,关键的剩余瓶颈在于输入侧。我们提出了TextPro-SLM,一个SLM,使口语输入更接近的韵律感知文本LLM。TextPro-SLM结合了WhisperPro,一个统一的语音编码器,产生同步的文本标记和韵律嵌入,与一个LLM骨干训练,以保留原始TLM的语义能力,同时学习非语言理解。实验表明,TextPro-SLM在3B和7 B尺度上都实现了领先SLM中最低的模态差距,同时在非语言理解任务上也提供了强大的整体性能。这些增益仅用大约1,000小时的LLM训练音频就实现了,这表明从输入端减少模态差距既有效又有效。
摘要:Speech large language models (SLMs) are typically built from text large language model (TLM) checkpoints, yet they still suffer from a substantial modality gap. Prior work has mainly attempted to reduce this gap from the output side by making speech generation more text-like, but the gap remains. We argue that the key remaining bottleneck lies on the input side. We propose TextPro-SLM, an SLM that makes spoken input more closely resemble that of a prosody-aware text LLM. TextPro-SLM combines WhisperPro, a unified speech encoder that produces synchronized text tokens and prosody embeddings, with an LLM backbone trained to preserve the semantic capabilities of the original TLM while learning paralinguistic understanding. Experiments show that TextPro-SLM achieves the lowest modality gap among leading SLMs at both 3B and 7B scales, while also delivering strong overall performance on paralinguistic understanding tasks. These gains are achieved with only roughly 1,000 hours of LLM training audio, suggesting that reducing the modality gap from the input side is both effective and data-efficient.


【10】X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

标题:X-Voice:通过Zero-Shot跨语言语音克隆使每个人都能说30种语言
链接:https://arxiv.org/abs/2605.05611
作者:Rixi Xu,Qingyu Liu,Haitao Li,Yushen Chen,Zhikang Niu,Yunting Yang,Jian Zhao,Ke Li,Berrak Sisman,Qinyuan Cheng,Xipeng Qiu,Kai Yu,Xie Chen
备注:16 pages, 4 figures, 9 tables
摘要:在本文中,我们提出了X-Voice,一个0.4B的多语种zero-shot语音克隆模型,克隆任意的声音,使每个人都能说30种语言。X-Voice在420 K小时的多语言语料库上进行训练,使用国际音标(IPA)作为统一表示。为了消除对提示文本的依赖,而不需要像强制对齐这样复杂的预处理,我们设计了一个两阶段的训练范例。在第一阶段,我们通过标准的条件流匹配训练建立X-Voice$_{\text{s1}}$,并使用它来合成10 K小时的说话者一致性片段作为音频提示。在阶段2中,我们对这些音频对进行微调,并屏蔽提示文本以获得X-Voice$_{\text{s2}}$,这使得zero-shot语音克隆不需要音频提示的转录。在架构上,我们通过实现语言标识符的双层注入和无分类器指导的解耦和调度来扩展F5-TTS,以促进多语言语音合成。主观和客观的评估结果表明,X-Voice优于现有的流匹配的多语言系统,如LEMAS-TTS和实现zero-shot跨语言克隆能力媲美十亿规模的模型,如Qwen 3-TTS。为了促进研究透明度和社区发展,我们开放了所有相关资源。
摘要:In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to speak 30 languages. X-Voice is trained on a 420K-hour multilingual corpus using the International Phonetic Alphabet (IPA) as a unified representation. To eliminate the reliance on prompt text without complex preprocessing like forced alignment, we design a two-stage training paradigm. In Stage 1, we establish X-Voice$_{\text{s1}}$ through standard conditional flow-matching training and use it to synthesize 10K hours of speaker-consistent segments as audio prompts. In Stage 2, we fine-tune on these audio pairs with prompt text masked to derive X-Voice$_{\text{s2}}$, which enables zero-shot voice cloning without requiring transcripts of audio prompts. Architecturally, we extend F5-TTS by implementing a dual-level injection of language identifiers and decoupling and scheduling of Classifier-Free Guidance to facilitate multilingual speech synthesis. Subjective and objective evaluation results demonstrate that X-Voice outperforms existing flow-matching based multilingual systems like LEMAS-TTS and achieves zero-shot cross-lingual cloning capabilities comparable to billion-scale models such as Qwen3-TTS. To facilitate research transparency and community advancement, we open-source all related resources.


机器翻译由腾讯交互翻译提供,仅供参考