今日论文合集:cs.SD语音8篇,eess.AS音频处理6篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】TAC: Timestamped Audio Captioning
标题:TAC:时间戳音频字幕
链接:https://arxiv.org/abs/2602.15766

作者:Sonal Kumar,Prem Seetharaman,Ke Chen,Oriol Nieto,Jiaqi Su,Zhepei Wang,Rithesh Kumar,Dinesh Manocha,Nicholas J. Bryan,Zeyu Jin,Justin Salamon
摘要:大型音频语言模型很难在复杂的声学场景中理清重叠的事件,产生时间上不一致的字幕和频繁的幻觉。我们介绍时间戳音频字幕(TAC),一个模型,产生时间接地音频描述在不同程度的细节和分辨率。TAC使用合成数据管道进行训练,该管道从真实世界的音频源构建具有挑战性的动态混合,从而在现实复调条件下实现强大的学习。在事件检测和密集字幕方面,TAC优于所有竞争方法,具有低幻觉率和准确的时间基础。我们还介绍了TAC-V,一个视听管道生成语义丰富的视听描述。然后,我们表明,TAC和TAC-V作为一个“语义桥梁”的纯文本推理:一个简单的TAC$\rightarrow$LLM和TAC-V$\rightarrow$LLM级联达到最先进的分数基准音频(MMAU-Pro,MMSU,MMAR)和视听(DailyOmni,VideoHolmes)的理解和推理。
摘要:Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introduce Timestamped Audio Captioner (TAC), a model that produces temporally grounded audio descriptions at varying degrees of detail and resolution. TAC is trained with a synthetic data pipeline that constructs challenging and dynamic mixtures from real-world audio sources, enabling robust learning under realistic polyphonic conditions. Across event detection and dense captioning, TAC outperforms all competing methods, with a low hallucination rate and accurate temporal grounding. We also introduce TAC-V, an audio-visual pipeline to generate semantically rich audio-visual descriptions. We then show that TAC and TAC-V serves as a "semantic bridge" for a text-only reasoner: a simple TAC$\rightarrow$LLM and TAC-V$\rightarrow$LLM cascade achieves state-of-the-art scores on benchmarks for both audio (MMAU-Pro, MMSU, MMAR) and audio-visual (DailyOmni, VideoHolmes) understanding and reasoning respectively.


【2】A Generative-First Neural Audio Autoencoder
标题:生成优先的神经音频自动编码器
链接:https://arxiv.org/abs/2602.15749

作者:Jonah Casebeer,Ge Zhu,Zhepei Wang,Nicholas J. Bryan
备注:ICASSP 2026
摘要:神经自动编码器支持生成模型。实际上,大规模使用神经自编码器进行生成建模需要快速编码,低潜在速率和跨表示的单个模型。现有的方法是先重建:它们会导致高潜伏率,慢编码,以及离散与连续潜伏和不同音频通道格式的单独架构,阻碍了从预处理到推理调节的工作流程。我们引入了一种用于音频自动编码的生成优先架构,该架构将时间下采样从2048x增加到3360x,并在一个模型中支持连续和离散表示以及常见的音频通道格式。通过平衡压缩、质量和速度,它提供了10倍的编码速度,1.6倍的低速率,并消除了通道格式特定的变体,同时保持了具有竞争力的重建质量。这使得以前受处理成本限制的应用程序能够实现:60秒的单声道信号压缩到788个令牌,使生成建模更加易于处理。
摘要:Neural autoencoders underpin generative models. Practical, large-scale use of neural autoencoders for generative modeling necessitates fast encoding, low latent rates, and a single model across representations. Existing approaches are reconstruction-first: they incur high latent rates, slow encoding, and separate architectures for discrete vs. continuous latents and for different audio channel formats, hindering workflows from preprocessing to inference conditioning. We introduce a generative-first architecture for audio autoencoding that increases temporal downsampling from 2048x to 3360x and supports continuous and discrete representations and common audio channel formats in one model. By balancing compression, quality, and speed, it delivers 10x faster encoding, 1.6x lower rates, and eliminates channel-format-specific variants while maintaining competitive reconstruction quality. This enables applications previously constrained by processing costs: a 60-second mono signal compresses to 788 tokens, making generative modeling more tractable.


【3】UniTAF: A Modular Framework for Joint Text-to-Speech and Audio-to-Face Modeling
标题:UniTAF:文本到语音和音频到面部联合建模的模块化框架
链接:https://arxiv.org/abs/2602.15651

作者:Qiangong Zhou,Nagasaka Tomohiro
备注:16 pages, 12 figures
摘要:这项工作考虑将两个独立的模型TTS和A2 F合并为一个统一的模型,以实现内部特征传输,从而提高从文本生成的音频和面部表情之间的一致性。我们还讨论了情感控制机制从TTS到联合模型的扩展。这项工作的目的不是展示生成质量;而是从系统设计的角度验证了重用TTS中间表示进行语音和面部表情联合建模的可行性,并为后续语音表情协同设计提供工程实践参考。该项目代码已在https://github.com/GoldenFishes/UniTAF上开源
摘要:This work considers merging two independent models, TTS and A2F, into a unified model to enable internal feature transfer, thereby improving the consistency between audio and facial expressions generated from text. We also discuss the extension of the emotion control mechanism from TTS to the joint model. This work does not aim to showcase generation quality; instead, from a system design perspective, it validates the feasibility of reusing intermediate representations from TTS for joint modeling of speech and facial expressions, and provides engineering practice references for subsequent speech expression co-design. The project code has been open source at: https://github.com/GoldenFishes/UniTAF


【4】The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs
标题:均衡器:在神经音频编解码器中引入形状-收益分解
链接:https://arxiv.org/abs/2602.15491

作者:Samir Sadok,Laurent Girin,Xavier Alameda-Pineda
备注:Neural audio codecs, shape-gain decomposition, vector quantization, speech coding
摘要:神经音频编解码器(NAC)通常在同一潜在空间内联合编码语音/音频信号的短期能量(增益)和归一化结构(形状)。结果,它们对输入信号电平的全局变化的鲁棒性差,因为这种变化对编码器的输出处的嵌入向量及其量化具有很强的影响。这种方法本质上是低效的,导致码本冗余和次优比特率失真性能。为了解决这些限制,我们建议引入形状增益分解,广泛用于经典的语音/音频编码,到NAC框架。所提出的量化器方法的原理是在NAC编码器之前将输入信号分解为短期增益和归一化形状向量。形状矢量由NAC处理,而增益用标量量化进行量化并单独传输。输出(解码)信号从NAC的归一化输出和量化增益重构。我们对语音信号进行的实验表明,这种通用的方法,很容易适用于任何NAC,使比特率失真性能的大幅提高,以及复杂性的大幅降低。
摘要:Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly within the same latent space. As a result, they are poorly robust to a global variation of the input signal level in the sense that such variation has strong influence on the embedding vectors at the output of the encoder and their quantization. This methodology is inherently inefficient, leading to codebook redundancy and suboptimal bitrate-distortion performance. To address these limitations, we propose to introduce shape-gain decomposition, widely used in classical speech/audio coding, into the NAC framework. The principle of the proposed Equalizer methodology is to decompose the input signal -- before the NAC encoder -- into gain and normalized shape vector on a short-term basis. The shape vector is processed by the NAC, while the gain is quantized with scalar quantization and transmitted separately. The output (decoded) signal is reconstructed from the normalized output of the NAC and the quantized gain. Our experiments conducted on speech signals show that this general methodology, easily applicable to any NAC, enables a substantial gain in bitrate-distortion performance, as well as a massive reduction in complexity.


【5】S-PRESSO: Ultra Low Bitrate Sound Effect Compression With Diffusion Autoencoders And Offline Quantization
标题:S-CLARIO:采用扩散自动编码器和离线量化的超低比特率音效压缩
链接:https://arxiv.org/abs/2602.15082

作者:Zineb Lahrichi,Gaëtan Hadjeres,Gaël Richard,Geoffroy Peeters
摘要:神经音频压缩模型最近实现了极高的压缩率,从而实现了高效的潜在生成建模。相反,潜在生成模型已被应用于压缩,推动了连续和离散方法的极限。然而,现有的方法仍然局限于低分辨率音频,并且在非常低的比特率下大幅降级,其中可听伪像是突出的。在本文中,我们提出了S-STOO,一个48 kHz的声音效果压缩模型,产生连续和离散嵌入在超低比特率,低至0.096 kbps,通过离线量化。我们的模型依赖于预训练的潜在扩散模型来解码由潜在编码器学习的压缩音频嵌入。利用扩散解码器的生成先验,我们实现了极低的帧速率,低至1Hz(750倍压缩率),以精确的保真度为代价产生令人信服的逼真重建。尽管在高压缩率下运行,我们证明了S-STRO在音频质量,声学相似性和重建指标方面优于连续和离散基线。
摘要:Neural audio compression models have recently achieved extreme compression rates, enabling efficient latent generative modeling. Conversely, latent generative models have been applied to compression, pushing the limits of continuous and discrete approaches. However, existing methods remain constrained to low-resolution audio and degrade substantially at very low bitrates, where audible artifacts are prominent. In this paper, we present S-PRESSO, a 48kHz sound effect compression model that produces both continuous and discrete embeddings at ultra-low bitrates, down to 0.096 kbps, via offline quantization. Our model relies on a pretrained latent diffusion model to decode compressed audio embeddings learned by a latent encoder. Leveraging the generative priors of the diffusion decoder, we achieve extremely low frame rates, down to 1Hz (750x compression rate), producing convincing and realistic reconstructions at the cost of exact fidelity. Despite operating at high compression rates, we demonstrate that S-PRESSO outperforms both continuous and discrete baselines in audio quality, acoustic similarity and reconstruction metrics.


【6】Structure-Aware Piano Accompaniment via Style Planning and Dataset-Aligned Pattern Retrieval
标题:通过风格规划和数据集对齐模式检索的结构感知钢琴伴奏
链接:https://arxiv.org/abs/2602.15074

作者:Wanyu Zang,Yang Yu,Meng Yu
备注:12 pages
摘要:我们介绍了一种结构感知的方法,象征性的钢琴伴奏,从音符级实现的高层次规划。一个轻量级的Transformer预测一个可解释的,每措施的风格计划的条件部分/短语结构和功能的和谐,和检索器,然后选择和重新协调人类表演的钢琴模式从语料库。我们制定检索模式匹配下的一个明确的能源与谐波的可行性,结构角色的兼容性,语音领先的连续性,风格偏好,和重复控制。给定一个结构化的引导表和可选的关键字提示,系统生成钢琴伴奏谱。在我们的实验中,Transformer风格规划器引导的检索产生了具有强风格实现的多种长形式的副本。我们进一步分析计划消融和量化风格间的隔离。实验结果表明,我们的钢琴伴奏生成的推理时间的方法的有效性。
摘要:We introduce a structure-aware approach for symbolic piano accompaniment that decouples high-level planning from note-level realization. A lightweight transformer predicts an interpretable, per-measure style plan conditioned on section/phrase structure and functional harmony, and a retriever then selects and reharmonizes human-performed piano patterns from a corpus. We formulate retrieval as pattern matching under an explicit energy with terms for harmonic feasibility, structural-role compatibility, voice-leading continuity, style preferences, and repetition control. Given a structured lead sheet and optional keyword prompts, the system generates piano-accompaniment MIDI. In our experiments, transformer style-planner-guided retrieval produces diverse long-form accompaniments with strong style realization. We further analyze planner ablations and quantify inter-style isolation. Experimental results demonstrate the effectiveness of our inference-time approach for piano accompaniment generation.


【7】Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios
标题:唤醒时注册:真实嘈杂人机对话场景中用于无缝交互的目标语音提取的首次比较研究
链接:https://arxiv.org/abs/2602.15519

作者:Yiming Yang,Guangyong Wang,Haixin Guan,Yanhua Long
备注:This paper is submitted to Interspeech 2026
摘要:目标语音提取(TSE)通常依赖于预先录制的高质量注册语音,这破坏了用户体验并限制了自发交互的可行性。在本文中,我们提出了注册唤醒(EoW),一种新的框架,唤醒词段,自然捕获的人机交互过程中,自动利用作为注册参考。这消除了对预先收集的语音的需要,以实现无缝体验。我们进行了第一次系统的EoW-TSE研究,在真实的不同声学条件下评估先进的判别和生成模型。鉴于唤醒词段的短和嘈杂的性质,我们研究使用基于LLM的TTS的招生扩增。结果表明,虽然目前的TSE模型面临EoW-TSE性能下降,但基于TTS的辅助显着增强了听力体验,尽管语音识别准确性仍存在差距。
摘要:Target speech extraction (TSE) typically relies on pre-recorded high-quality enrollment speech, which disrupts user experience and limits feasibility in spontaneous interaction. In this paper, we propose Enroll-on-Wakeup (EoW), a novel framework where the wake-word segment, captured naturally during human-machine interaction, is automatically utilized as the enrollment reference. This eliminates the need for pre-collected speech to enable a seamless experience. We perform the first systematic study of EoW-TSE, evaluating advanced discriminative and generative models under real diverse acoustic conditions. Given the short and noisy nature of wake-word segments, we investigate enrollment augmentation using LLM-based TTS. Results show that while current TSE models face performance degradation in EoW-TSE, TTS-based assistance significantly enhances the listening experience, though gaps remain in speech recognition accuracy.


【8】What Do Neurons Listen To? A Neuron-level Dissection of a General-purpose Audio Model
标题:神经元听什么?通用音频模型的神经元级剖析
链接:https://arxiv.org/abs/2602.15307

作者:Takao Kawamura,Daisuke Niizumi,Nobutaka Ono
备注:5 pages, 8 figures. Submitted to EUSIPCO 2026
摘要:在本文中,我们分析了一个通用的音频自监督学习(SSL)模型的内部表示从神经元级的角度。尽管它们作为特征提取器具有强大的经验性能,但SSL音频模型鲁棒泛化的内部机制仍然不清楚。利用机械可解释性的框架,我们通过分析不同任务的条件激活模式来识别和检查类特异性神经元。我们的分析表明,SSL模型促进了类特异性神经元的出现,这些神经元在新的任务类别中提供了广泛的覆盖范围。这些神经元在不同的语义类别和声学相似性(如语音属性和音高)之间表现出共享的反应。我们还证实,这些神经元有分类性能的功能影响。据我们所知,这是对通用音频SSL模型的第一次系统的神经元级分析,为其内部表示提供了新的见解。
摘要:In this paper, we analyze the internal representations of a general-purpose audio self-supervised learning (SSL) model from a neuron-level perspective. Despite their strong empirical performance as feature extractors, the internal mechanisms underlying the robust generalization of SSL audio models remain unclear. Drawing on the framework of mechanistic interpretability, we identify and examine class-specific neurons by analyzing conditional activation patterns across diverse tasks. Our analysis reveals that SSL models foster the emergence of class-specific neurons that provide extensive coverage across novel task classes. These neurons exhibit shared responses across different semantic categories and acoustic similarities, such as speech attributes and musical pitch. We also confirm that these neurons have a functional impact on classification performance. To our knowledge, this is the first systematic neuron-level analysis of a general-purpose audio SSL model, providing new insights into its internal representation.


eess.AS音频处理


【1】Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios
标题:唤醒时注册:真实嘈杂人机对话场景中用于无缝交互的目标语音提取的首次比较研究
链接:https://arxiv.org/abs/2602.15519

作者:Yiming Yang,Guangyong Wang,Haixin Guan,Yanhua Long
备注:This paper is submitted to Interspeech 2026
摘要:目标语音提取(TSE)通常依赖于预先录制的高质量注册语音,这破坏了用户体验并限制了自发交互的可行性。在本文中,我们提出了注册唤醒(EoW),一种新的框架,唤醒词段,自然捕获的人机交互过程中,自动利用作为注册参考。这消除了对预先收集的语音的需要,以实现无缝体验。我们进行了第一次系统的EoW-TSE研究,在真实的不同声学条件下评估先进的判别和生成模型。鉴于唤醒词段的短和嘈杂的性质,我们研究使用基于LLM的TTS的招生扩增。结果表明,虽然目前的TSE模型面临EoW-TSE性能下降,但基于TTS的辅助显着增强了听力体验,尽管语音识别准确性仍存在差距。
摘要:Target speech extraction (TSE) typically relies on pre-recorded high-quality enrollment speech, which disrupts user experience and limits feasibility in spontaneous interaction. In this paper, we propose Enroll-on-Wakeup (EoW), a novel framework where the wake-word segment, captured naturally during human-machine interaction, is automatically utilized as the enrollment reference. This eliminates the need for pre-collected speech to enable a seamless experience. We perform the first systematic study of EoW-TSE, evaluating advanced discriminative and generative models under real diverse acoustic conditions. Given the short and noisy nature of wake-word segments, we investigate enrollment augmentation using LLM-based TTS. Results show that while current TSE models face performance degradation in EoW-TSE, TTS-based assistance significantly enhances the listening experience, though gaps remain in speech recognition accuracy.


【2】Bottleneck Transformer-Based Approach for Improved Automatic STOI Score Prediction
标题:基于瓶颈转换器的改进自动STIOI评分预测方法
链接:https://arxiv.org/abs/2602.15484

作者:Amartyaveer,Murali Kadambi,Chandra Mohan Sharma,Anupam Mondal,Prasanta Kumar Ghosh
备注:7 pages, 7 tables, 2 figures, ASRU 2025
摘要:在这项研究中,我们提出了一种新的方法来预测短期目标可懂度(STOI)度量使用瓶颈Transformer架构。用于计算STOI的传统方法通常需要干净的参考语音,这限制了它们在现实世界中的适用性。为了解决这个问题,许多基于深度学习的非侵入式语音评估模型引起了人们的极大兴趣。许多研究取得了令人称道的成绩,但仍有进一步改进的余地。   我们建议使用瓶颈Transformer,将卷积块用于学习帧级特征和多头自注意(MHSA)层来聚合信息。这些组件使Transformer能够专注于输入数据的关键方面。与使用自监督学习(SSL)和光谱特征作为输入的最先进模型相比,我们的模型在可见和不可见的场景中表现出更高的相关性和更低的均方误差。
摘要:In this study, we have presented a novel approach to predict the Short-Time Objective Intelligibility (STOI) metric using a bottleneck transformer architecture. Traditional methods for calculating STOI typically requires clean reference speech, which limits their applicability in the real world. To address this, numerous deep learning-based nonintrusive speech assessment models have garnered significant interest. Many studies have achieved commendable performance, but there is room for further improvement.   We propose the use of bottleneck transformer, incorporating convolution blocks for learning frame-level features and a multi-head self-attention (MHSA) layer to aggregate the information. These components enable the transformer to focus on the key aspects of the input data. Our model has shown higher correlation and lower mean squared error for both seen and unseen scenarios compared to the state-of-the-art model using self-supervised learning (SSL) and spectral features as inputs.


【3】What Do Neurons Listen To? A Neuron-level Dissection of a General-purpose Audio Model
标题:神经元听什么?通用音频模型的神经元级剖析
链接:https://arxiv.org/abs/2602.15307

作者:Takao Kawamura,Daisuke Niizumi,Nobutaka Ono
备注:5 pages, 8 figures. Submitted to EUSIPCO 2026
摘要:在本文中,我们分析了一个通用的音频自监督学习(SSL)模型的内部表示从神经元级的角度。尽管它们作为特征提取器具有强大的经验性能,但SSL音频模型鲁棒泛化的内部机制仍然不清楚。利用机械可解释性的框架,我们通过分析不同任务的条件激活模式来识别和检查类特异性神经元。我们的分析表明,SSL模型促进了类特异性神经元的出现,这些神经元在新的任务类别中提供了广泛的覆盖范围。这些神经元在不同的语义类别和声学相似性(如语音属性和音高)之间表现出共享的反应。我们还证实,这些神经元有分类性能的功能影响。据我们所知,这是对通用音频SSL模型的第一次系统的神经元级分析,为其内部表示提供了新的见解。
摘要:In this paper, we analyze the internal representations of a general-purpose audio self-supervised learning (SSL) model from a neuron-level perspective. Despite their strong empirical performance as feature extractors, the internal mechanisms underlying the robust generalization of SSL audio models remain unclear. Drawing on the framework of mechanistic interpretability, we identify and examine class-specific neurons by analyzing conditional activation patterns across diverse tasks. Our analysis reveals that SSL models foster the emergence of class-specific neurons that provide extensive coverage across novel task classes. These neurons exhibit shared responses across different semantic categories and acoustic similarities, such as speech attributes and musical pitch. We also confirm that these neurons have a functional impact on classification performance. To our knowledge, this is the first systematic neuron-level analysis of a general-purpose audio SSL model, providing new insights into its internal representation.


【4】A Generative-First Neural Audio Autoencoder
标题:生成优先的神经音频自动编码器
链接:https://arxiv.org/abs/2602.15749

作者:Jonah Casebeer,Ge Zhu,Zhepei Wang,Nicholas J. Bryan
备注:ICASSP 2026
摘要:神经自动编码器支持生成模型。实际上,大规模使用神经自编码器进行生成建模需要快速编码,低潜在速率和跨表示的单个模型。现有的方法是先重建:它们会导致高潜伏率,慢编码,以及离散与连续潜伏和不同音频通道格式的单独架构,阻碍了从预处理到推理调节的工作流程。我们引入了一种用于音频自动编码的生成优先架构,该架构将时间下采样从2048x增加到3360x,并在一个模型中支持连续和离散表示以及常见的音频通道格式。通过平衡压缩、质量和速度,它提供了10倍的编码速度,1.6倍的低速率,并消除了通道格式特定的变体,同时保持了具有竞争力的重建质量。这使得以前受处理成本限制的应用程序能够实现:60秒的单声道信号压缩到788个令牌,使生成建模更加易于处理。
摘要:Neural autoencoders underpin generative models. Practical, large-scale use of neural autoencoders for generative modeling necessitates fast encoding, low latent rates, and a single model across representations. Existing approaches are reconstruction-first: they incur high latent rates, slow encoding, and separate architectures for discrete vs. continuous latents and for different audio channel formats, hindering workflows from preprocessing to inference conditioning. We introduce a generative-first architecture for audio autoencoding that increases temporal downsampling from 2048x to 3360x and supports continuous and discrete representations and common audio channel formats in one model. By balancing compression, quality, and speed, it delivers 10x faster encoding, 1.6x lower rates, and eliminates channel-format-specific variants while maintaining competitive reconstruction quality. This enables applications previously constrained by processing costs: a 60-second mono signal compresses to 788 tokens, making generative modeling more tractable.


【5】UniTAF: A Modular Framework for Joint Text-to-Speech and Audio-to-Face Modeling
标题:UniTAF:文本到语音和音频到面部联合建模的模块化框架
链接:https://arxiv.org/abs/2602.15651

作者:Qiangong Zhou,Nagasaka Tomohiro
备注:16 pages, 12 figures
摘要:这项工作考虑将两个独立的模型TTS和A2F合并为一个统一的模型,以实现内部特征传输,从而提高从文本生成的音频和面部表情之间的一致性。我们还讨论了情感控制机制从TTS到联合模型的扩展。本文的工作并不是为了展示生成质量,而是从系统设计的角度,验证了将TTS中间表示复用到语音和面部表情联合建模的可行性,并为后续的语音表情协同设计提供了工程实践参考。该项目代码已在https://github.com/GoldenFishes/UniTAF上开源
摘要:This work considers merging two independent models, TTS and A2F, into a unified model to enable internal feature transfer, thereby improving the consistency between audio and facial expressions generated from text. We also discuss the extension of the emotion control mechanism from TTS to the joint model. This work does not aim to showcase generation quality; instead, from a system design perspective, it validates the feasibility of reusing intermediate representations from TTS for joint modeling of speech and facial expressions, and provides engineering practice references for subsequent speech expression co-design. The project code has been open source at: https://github.com/GoldenFishes/UniTAF


【6】ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
标题:ZeroSyl:用于口语建模的简单零资源音节令牌化
链接:https://arxiv.org/abs/2602.15537

作者:Nicol Visser,Simon Malan,Danel Slabbert,Herman Kamper
备注:3 figures, 2 tables
摘要:纯语音语言模型旨在直接从原始音频中学习语言,而无需文本资源。一个关键的挑战是,来自自监督语音编码器的离散令牌会导致过长的序列,这激发了最近对音节类单元的研究。然而,像Sylber和SyllableLM这样的方法依赖于复杂的多阶段训练管道。我们提出了ZeroSyl,这是一种简单的免训练方法,可以直接从冻结的WavLM模型中提取音节边界和嵌入。在WavLM的中间层中使用L2范数的特征,ZeroSyl实现了具有竞争力的音节分割性能。将得到的片段进行均值池化,使用K均值进行离散化,并用于训练语言模型。ZeroSyl在词汇、句法和叙事基准方面优于之前的音节标记器。缩放实验表明,虽然细粒度的单位是有益的词汇任务,我们发现的音节单位表现出更好的缩放行为的句法建模。
摘要:Pure speech language models aim to learn language directly from raw audio without textual resources. A key challenge is that discrete tokens from self-supervised speech encoders result in excessively long sequences, motivating recent work on syllable-like units. However, methods like Sylber and SyllableLM rely on intricate multi-stage training pipelines. We propose ZeroSyl, a simple training-free method to extract syllable boundaries and embeddings directly from a frozen WavLM model. Using L2 norms of features in WavLM's intermediate layers, ZeroSyl achieves competitive syllable segmentation performance. The resulting segments are mean-pooled, discretized using K-means, and used to train a language model. ZeroSyl outperforms prior syllabic tokenizers across lexical, syntactic, and narrative benchmarks. Scaling experiments show that while finer-grained units are beneficial for lexical tasks, our discovered syllabic units exhibit better scaling behavior for syntactic modeling.


机器翻译由腾讯交互翻译提供,仅供参考