今日论文合集:cs.SD语音13篇,eess.AS音频处理10篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
标题:SemanticVooder:通过语义潜伏连接音频生成和音频理解
链接:https://arxiv.org/abs/2602.23333

作者:Zeyu Xie,Chenxing Li,Qiao Jin,Xuenan Xu,Guanrou Yang,Wenfu Wang,Mengyue Wu,Dong Yu,Yuexian Zou
备注:Demo: https://zeyuxie29.github.io/SemanticVocoder/
摘要:最近的音频生成模型通常依赖于变分自动编码器(VAE),并在VAE潜在空间内执行生成。虽然VAE擅长压缩和重建,但它们的潜在特性本质上编码了低级别的声学细节,而不是语义上有区别的信息,导致纠缠的事件语义,并使生成模型的训练复杂化。为了解决这些问题,我们放弃VAE声学潜伏期,并引入语义编码器潜伏期,从而提出SemanticVocoder,生成声码器,直接合成波形的语义潜伏期。配备了SemanticVocoder,我们的文本到音频生成模型在AudioCaps测试集上实现了12.823的Frechet距离和1.709的Frechet音频距离,因为引入的语义潜伏期与声学VAE潜伏期相比具有更好的可辨别性。除了改进生成性能之外,它还可以作为在共享语义空间内统一音频理解和生成的有希望的尝试。生成的示例可在https://zeyuxie29.github.io/SemanticVocoder/上获得。
摘要:Recent audio generation models typically rely on Variational Autoencoders (VAEs) and perform generation within the VAE latent space. Although VAEs excel at compression and reconstruction, their latents inherently encode low-level acoustic details rather than semantically discriminative information, leading to entangled event semantics and complicating the training of generative models. To address these issues, we discard VAE acoustic latents and introduce semantic encoder latents, thereby proposing SemanticVocoder, a generative vocoder that directly synthesizes waveforms from semantic latents. Equipped with SemanticVocoder, our text-to-audio generation model achieves a Frechet Distance of 12.823 and a Frechet Audio Distance of 1.709 on the AudioCaps test set, as the introduced semantic latents exhibit superior discriminability compared to acoustic VAE latents. Beyond improved generation performance, it also serves as a promising attempt towards unifying audio understanding and generation within a shared semantic space. Generated samples are available at https://zeyuxie29.github.io/SemanticVocoder/.


【2】Make It Hard to Hear, Easy to Learn: Long-Form Bengali ASR and Speaker Diarization via Extreme Augmentation and Perfect Alignment
标题:让它听起来很难,易于学习:通过极端增强和完美对齐实现长格式孟加拉语ASB和扬声器扩展
链接:https://arxiv.org/abs/2602.23070

作者:Sanjid Hasan,Risalat Labib,A H M Fuad,Bayazid Hasan
备注:4 pages, 2 figures
摘要:虽然孟加拉语的自动语音识别(ASR)已经取得了重大进展,但处理长时间音频和执行鲁棒的说话人日志仍然是关键的研究空白。为了解决这种语言的联合ASR和日记资源的严重缺乏,我们引入了Lipi-Ghor-882,这是一个全面的882小时多说话者孟加拉语数据集。在本文中,详细介绍了我们提交的DL Sprint 4.0比赛,我们系统地评估各种架构和方法的长格式孟加拉语讲话。对于ASR,我们证明了原始数据缩放是无效的;相反,利用完美对齐的注释与合成声学降级(噪音和混响)相结合的有针对性的微调成为最有效的方法。相反,对于说话人日记,我们观察到全球开源的最先进的模型(如Diarizen)在这个复杂的数据集上表现得非常差。广泛的模型再训练产生了微不足道的改进;相反,基线模型输出的战略性启发式后处理被证明是提高准确性的主要驱动力。最终,这项工作概述了一个高度优化的双管道实现了0.019美元的实时因子(RTF),建立了一个实用的,经验支持的基准低资源,长形式的语音处理。
摘要:Although Automatic Speech Recognition (ASR) in Bengali has seen significant progress, processing long-duration audio and performing robust speaker diarization remain critical research gaps. To address the severe scarcity of joint ASR and diarization resources for this language, we introduce Lipi-Ghor-882, a comprehensive 882-hour multi-speaker Bengali dataset. In this paper, detailing our submission to the DL Sprint 4.0 competition, we systematically evaluate various architectures and approaches for long-form Bengali speech. For ASR, we demonstrate that raw data scaling is ineffective; instead, targeted fine-tuning utilizing perfectly aligned annotations paired with synthetic acoustic degradation (noise and reverberation) emerges as the singular most effective approach. Conversely, for speaker diarization, we observed that global open-source state-of-the-art models (such as Diarizen) performed surprisingly poorly on this complex dataset. Extensive model retraining yielded negligible improvements; instead, strategic, heuristic post-processing of baseline model outputs proved to be the primary driver for increasing accuracy. Ultimately, this work outlines a highly optimized dual pipeline achieving a $\sim$0.019 Real-Time Factor (RTF), establishing a practical, empirically backed benchmark for low-resource, long-form speech processing.


【3】TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment
标题:TADA:通过文本-声学双重对齐的语音建模生成框架
链接:https://arxiv.org/abs/2602.23068

作者:Trung Dang,Sharath Rao,Ananya Gupta,Christopher Gagne,Panagiotis Tzirakis,Alice Baird,Jakub Piotr Cłapa,Peter Chin,Alan Cowen
摘要:现代文语转换(TTS)系统越来越多地利用大型语言模型(LLM)架构来实现可扩展、高保真、zero-shot生成。然而,这些系统通常依赖于固定帧速率的声学标记化,导致语音序列明显长于其对应的文本并且与其对应的文本异步。除了计算效率低之外,这种序列长度差异经常引发TTS中的幻觉,并放大了口语建模(SLM)中的模态间隙。在本文中,我们提出了一种新的标记化方案,该方案在连续声学特征和文本标记之间建立一对一的同步,从而在LLM内实现统一的单流建模。我们证明,这些同步令牌保持高保真的音频重建,并可以有效地在一个潜在的空间建模的大型语言模型与流匹配头。此外,在上下文中无缝切换语音模态的能力使纯文本指导成为可能--这种技术将纯文本和文本语音模式的logits混合在一起,以灵活地弥合纯文本LLM智能的差距。实验结果表明,我们的方法实现了性能竞争力与国家的最先进的TTS和SLM系统,同时几乎消除内容幻觉和保留语言的完整性,所有在一个显着降低的推理成本。
摘要:Modern Text-to-Speech (TTS) systems increasingly leverage Large Language Model (LLM) architectures to achieve scalable, high-fidelity, zero-shot generation. However, these systems typically rely on fixed-frame-rate acoustic tokenization, resulting in speech sequences that are significantly longer than, and asynchronous with their corresponding text. Beyond computational inefficiency, this sequence length disparity often triggers hallucinations in TTS and amplifies the modality gap in spoken language modeling (SLM). In this paper, we propose a novel tokenization scheme that establishes one-to-one synchronization between continuous acoustic features and text tokens, enabling unified, single-stream modeling within an LLM. We demonstrate that these synchronous tokens maintain high-fidelity audio reconstruction and can be effectively modeled in a latent space by a large language model with a flow matching head. Moreover, the ability to seamlessly toggle speech modality within the context enables text-only guidance--a technique that blends logits from text-only and text-speech modes to flexibly bridge the gap toward text-only LLM intelligence. Experimental results indicate that our approach achieves performance competitive with state-of-the-art TTS and SLM systems while virtually eliminating content hallucinations and preserving linguistic integrity, all at a significantly reduced inference cost.


【4】A Holistic Framework for Robust Bangla ASR and Speaker Diarization with Optimized VAD and CTC Alignment
标题:具有优化的VAD和CTC对齐的鲁棒孟加拉语ASR和扬声器日志化的整体框架
链接:https://arxiv.org/abs/2602.22935

作者:Zarif Ishmam,Zarif Mahir,Shafnan Wasif,Md. Ishtiak Moin
备注:5 pages
摘要:尽管孟加拉语是全球使用最广泛的语言之一,但在自然语言处理(NLP)领域,孟加拉语仍然是一种低资源语言。孟加拉语的主流自动语音识别(ASR)和扬声器日记系统在处理超过3060秒的长格式音频时会遇到困难。本文提出了一个强大的框架,专门为扩展孟加拉语内容,利用现有的模型增强了新的优化管道的DL Sprint 4.0比赛。我们的方法利用语音活动检测(VAD)优化和连接主义时间分类(CTC)分割通过强制词对齐,以保持时间的准确性和转录的完整性在很长一段时间。此外,我们采用了几种微调技术,并使用增强技术和噪声去除对数据进行预处理。通过弥合在复杂的多扬声器环境中的性能差距,这项工作提供了一个可扩展的解决方案,为现实世界中,长格式孟加拉语语音应用。
摘要:Despite being one of the most widely spoken languages globally, Bangla remains a low-resource language in the field of Natural Language Processing (NLP). Mainstream Automatic Speech Recognition (ASR) and Speaker Diarization systems for Bangla struggles when processing longform audio exceeding 3060 seconds. This paper presents a robust framework specifically engineered for extended Bangla content by leveraging preexisting models enhanced with novel optimization pipelines for the DL Sprint 4.0 contest. Our approach utilizes Voice Activity Detection (VAD) optimization and Connectionist Temporal Classification (CTC) segmentation via forced word alignment to maintain temporal accuracy and transcription integrity over long durations. Additionally, we employed several finetuning techniques and preprocessed the data using augmentation techniques and noise removal. By bridging the performance gap in complex, multi-speaker environments, this work provides a scalable solution for real-world, longform Bangla speech applications.


【5】Same Words, Different Judgments: Modality Effects on Preference Alignment
标题:相同的词语,不同的判断:情态对偏好一致的影响
链接:https://arxiv.org/abs/2602.22710

作者:Aaron Broukhim,Nadir Weibel,Eshin Jolly
备注:Submitted to Interspeech 2026 for review
摘要:基于偏好的强化学习(PbRL)是将人工智能系统与人类偏好相匹配的主要框架,但其在语音方面的应用仍有待探索。我们提出了一个控制的跨模态研究人类和合成偏好注释,比较文本和音频评价相同的语义内容在100个提示。音频偏好被证明与文本一样可靠,评分者之间的一致性在$\sim$9评分者中达到良好水平(ICC(2,k)$\约$.80)-这是偏好注释文献中第一个基于ICC的可靠性表征。然而,模态重塑了人们的判断方式:音频评分员表现出更窄的决策阈值,减少长度偏差,更面向用户的评价标准,几乎有机会跨模态协议。合成评级进一步与人类判断保持一致,并预测评级者之间的一致性,支持它们用于分类模糊对和完全替代人类注释。
摘要:Preference-based reinforcement learning (PbRL) is the dominant framework for aligning AI systems to human preferences, but its application to speech remains underexplored. We present a controlled cross-modal study of human and synthetic preference annotations, comparing text and audio evaluations of identical semantic content across 100 prompts. Audio preferences prove as reliable as text, with inter-rater agreement reaching good levels (ICC(2,k) $\approx$ .80) at $\sim$9 raters -- the first ICC-based reliability characterization in the preference annotation literature for either modality. However, modality reshapes how people judge: audio raters exhibit narrower decision thresholds, reduced length bias, and more user-oriented evaluation criteria, with near-chance cross-modality agreement. Synthetic ratings further align with human judgments and predict inter-rater agreement, supporting their use both for triaging ambiguous pairs and as full replacements for human annotations.


【6】Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech
标题:将发声、模仿和想象语音的神经表示联系起来
链接:https://arxiv.org/abs/2602.22597

作者:Maryam Maghsoudi,Rupesh Chillale,Shihab A. Shamma
摘要:我们研究了使用公开的立体定向脑电图记录的发声,默拟和想象的语音神经表征之间的关系。大多数先前的研究都集中在解码语音反应在每个条件下分别。相反,在这里,我们通过为每个条件训练线性谱图重建模型来探索不同条件下的响应如何相关,并评估它们在不同条件下的泛化。我们证明了在一个条件下训练的线性解码器通常成功地转移到其他人,这意味着共享的语音表示。这种共性进行了评估与刺激水平的可辨别性,通过执行基于秩的分析,证明在内部和跨条件下的刺激特异性结构的保存。最后,我们将线性重建与非线性神经网络的重建进行了比较。虽然两者都表现出交叉条件转移,但线性模型实现了更好的刺激水平辨别力。
摘要:We investigated the relationship among neural representations of vocalized, mimed, and imagined speech recorded using publicly available stereotactic EEG recordings. Most prior studies have focused on decoding speech responses within each condition separately. Here, instead, we explore how responses across conditions relate by training linear spectrogram reconstruction models for each condition and evaluate their generalization across conditions. We demonstrate that linear decoders trained on one condition generally transfer successfully to others, implying shared speech representations. This commonality was assessed with stimulus-level discriminability by performing a rank-based analysis demonstrating preservation of stimulus-specific structure in both within- and across-conditions. Finally, we compared linear reconstructions to those from a nonlinear neural network. While both exhibited cross-condition transfer, linear models achieve superior stimulus-level discriminability.


【7】Efficient Dialect-Aware Modeling and Conditioning for Low-Resource Taiwanese Hakka Speech Processing
标题:低资源台湾客语语音处理的高效方言感知建模和条件处理
链接:https://arxiv.org/abs/2602.22522

作者:An-Ci Peng,Kuan-Tang Huang,Tien-Hong Lo,Hung-Shin Lee,Hsin-Min Wang,Berlin Chen
备注:Accepted to LREC 2026
摘要:台湾客家语是一种资源匮乏的濒危语言,对自动语音识别(ASR)提出了重大挑战,包括方言的高度变异性和两种不同的书写系统(汉字和拼音)的存在。传统的ASR模型在这种情况下往往会遇到困难,因为它们往往会将基本的语言内容与方言特有的语音和词汇方面的变化混为一谈。为了应对这些挑战,我们提出了一个基于递归神经网络传感器(RNN-T)的统一框架。我们的方法的核心是引入方言感知的建模策略,旨在解开方言的“风格”从语言的“内容”,这提高了模型的能力,学习强大的和广义的表示。此外,该框架采用参数有效的预测网络,同时建模ASR(汉字和拼音)。我们证明,这些任务创建一个强大的协同作用,其中的交叉脚本目标作为一个相互正则化,以提高主要ASR任务。在HAT语料库上进行的实验表明,该模型对汉字和拼音ASR的相对错误率分别降低了57.00%和40.41%。据我们所知,这是第一次系统地调查客家方言变异对ASR的影响,也是第一个能够共同解决这些任务的单一模型。
摘要:Taiwanese Hakka is a low-resource, endangered language that poses significant challenges for automatic speech recognition (ASR), including high dialectal variability and the presence of two distinct writing systems (Hanzi and Pinyin). Traditional ASR models often encounter difficulties in this context, as they tend to conflate essential linguistic content with dialect-specific variations across both phonological and lexical dimensions. To address these challenges, we propose a unified framework grounded in the Recurrent Neural Network Transducers (RNN-T). Central to our approach is the introduction of dialect-aware modeling strategies designed to disentangle dialectal "style" from linguistic "content", which enhances the model's capacity to learn robust and generalized representations. Additionally, the framework employs parameter-efficient prediction networks to concurrently model ASR (Hanzi and Pinyin). We demonstrate that these tasks create a powerful synergy, wherein the cross-script objective serves as a mutual regularizer to improve the primary ASR tasks. Experiments conducted on the HAT corpus reveal that our model achieves 57.00% and 40.41% relative error rate reduction on Hanzi and Pinyin ASR, respectively. To our knowledge, this is the first systematic investigation into the impact of Hakka dialectal variations on ASR and the first single model capable of jointly addressing these tasks.


【8】mmWave Radar Aware Dual-Conditioned GAN for Speech Reconstruction of Signals With Low SNR
标题:毫米波雷达感知双条件GAN用于低SNR信号的语音重建
链接:https://arxiv.org/abs/2602.22431

作者:Jash Karani,Adithya Chittem,Deepan Roy,Sandeep Joshi
备注:Under review at Interspeech 2026
摘要:毫米波(mmWave)雷达捕获的信号带宽有限且有噪声,使得难以重建可理解的全带宽语音。在这项工作中,我们提出了一个两阶段的语音重建管道毫米波使用雷达感知双条件生成对抗网络(RAD-GAN),这是能够执行带宽扩展的信号与低信噪比(-5 dB到-1 dB),通过玻璃墙捕获。我们提出了一个毫米波定制的多梅尔鉴别器(MMD)和残余融合门(RFG),以提高发电机输入处理多个条件通道。所提出的两阶段流水线涉及在合成剪切的干净语音上预训练模型,以及在RFG生成的融合梅尔频谱图上微调。我们的经验表明,所提出的方法,在有限的数据集上训练,没有预先训练的模块,也没有数据增强,在这个特定的任务中表现优于最先进的方法。RAD-GAN的音频示例可在https://rad-gan-demo-site.vercel.app/上在线获得。
摘要:Millimeter-wave (mmWave) radar captures are band-limited and noisy, making for difficult reconstruction of intelligible full-bandwidth speech. In this work, we propose a two-stage speech reconstruction pipeline for mmWave using a Radar-Aware Dual-conditioned Generative Adversarial Network (RAD-GAN), which is capable of performing bandwidth extension on signals with low signal-to-noise ratios (-5 dB to -1 dB), captured through glass walls. We propose an mmWave-tailored Multi-Mel Discriminator (MMD) and a Residual Fusion Gate (RFG) to enhance the generator input to process multiple conditioning channels. The proposed two-stage pipeline involves pretraining the model on synthetically clipped clean speech and finetuning on fused mel spectrograms generated by the RFG. We empirically show that the proposed method, trained on a limited dataset, with no pre-trained modules, and no data augmentations, outperformed state-of-the-art approaches for this specific task. Audio examples of RAD-GAN are available online at https://rad-gan-demo-site.vercel.app/.


【9】Absorbing Discrete Diffusion for Speech Enhancement
标题:吸收离散扩散进行语音增强
链接:https://arxiv.org/abs/2602.22417

作者:Philippe Gonzalez
备注:Submitted to Interspeech 2026
摘要:受神经语音编码和基于扩散的语言建模领域最新发展的启发,我们通过使用吸收离散扩散对给定有噪语音代码的干净语音代码的条件分布进行建模来解决语音增强问题。所提出的方法,我们称之为ADDSE,利用神经音频编解码器的表达潜在空间和扩散模型的非自回归采样过程。为了有效地建模残差矢量量化码的层次结构,我们提出了RQDiT,它结合了RQ-变压器和扩散Transformers的非自回归建模技术。结果表明,在两个数据集上的非侵入性客观指标方面具有竞争力的性能,特别是在低信噪比和很少的采样步骤。代码和音频示例可在线获得。
摘要:Inspired by recent developments in neural speech coding and diffusion-based language modeling, we tackle speech enhancement by modeling the conditional distribution of clean speech codes given noisy speech codes using absorbing discrete diffusion. The proposed approach, which we call ADDSE, leverages both the expressive latent space of neural audio codecs and the non-autoregressive sampling procedure of diffusion models. To efficiently model the hierarchical structure of residual vector quantization codes, we propose RQDiT, which combines techniques from RQ-Transformer and diffusion Transformers for non-autoregressive modeling. Results show competitive performance in terms of non-intrusive objective metrics on two datasets, especially at low signal-to-noise ratios and with few sampling steps. Code and audio examples are available online.


【10】WaveSSM: Multiscale State-Space Models for Non-stationary Signal Attention
标题:WaveRSM:非平稳信号注意力的多尺度状态空间模型
链接:https://arxiv.org/abs/2602.22266

作者:Ruben Solozabal,Velibor Bojkovic,Hilal Alquabeh,Klea Ziu,Kentaro Inui,Martin Takac
摘要:状态空间模型(SSM)已经成为远程序列建模的强大基础,HiPPO框架表明连续时间投影算子可以用于导出稳定的,内存高效的动态系统,该系统对输入信号的过去历史进行编码。然而,现有的基于投影的SSM通常依赖于具有全局时间支持的多项式基,其归纳偏差与表现出局部或瞬态结构的信号匹配不良。在这项工作中,我们介绍\n {WaveSSM},一个集合的SSM构建小波框架。我们的主要观察是,小波框架产生本地化的时间维度上的支持,用于需要精确定位的任务。从经验上讲,我们表明,在同等条件下,\textit{WaveSSM}在具有瞬态动态的真实世界数据集上的性能优于正交对应物S4,包括PTB-XL数据集上的生理信号和语音命令上的原始音频。
摘要:State-space models (SSMs) have emerged as a powerful foundation for long-range sequence modeling, with the HiPPO framework showing that continuous-time projection operators can be used to derive stable, memory-efficient dynamical systems that encode the past history of the input signal. However, existing projection-based SSMs often rely on polynomial bases with global temporal support, whose inductive biases are poorly matched to signals exhibiting localized or transient structure. In this work, we introduce \emph{WaveSSM}, a collection of SSMs constructed over wavelet frames. Our key observation is that wavelet frames yield a localized support on the temporal dimension, useful for tasks requiring precise localization. Empirically, we show that on equal conditions, \textit{WaveSSM} outperforms orthogonal counterparts as S4 on real-world datasets with transient dynamics, including physiological signals on the PTB-XL dataset and raw audio on Speech Commands.


【11】AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs
标题:AR & D:检索和描述解释音频LLM概念的框架
链接:https://arxiv.org/abs/2602.22253

作者:Townim Faisal Chowdhury,Ta Duc Huy,Siqi Pan,Jeremy Stoddard,Zhibin Liao
备注:Accepted at International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026
摘要:尽管在音频感知任务中表现出色,但大型音频语言模型(AudioLLM)仍然无法解释。这种缺乏可解释性背后的一个主要因素是,这些模型中的单个神经元经常在响应几个不相关的概念时激活。我们引入了AudioLLM的第一个机械可解释性框架,利用稀疏自动编码器(SAE)将多语义激活分解为单语义特征。我们的管道识别代表性的音频片段,通过自动字幕分配有意义的名称,并通过人工评估和指导来验证概念。实验表明,AudioLLM编码结构化和可解释的功能,增强透明度和控制。这项工作为高风险领域的可靠部署提供了基础,并使未来扩展到更大的模型,多语言音频和更细粒度的语言特征。项目网址:https://townim-faisal.github.io/AutoInterpret-AudioLLM/
摘要:Despite strong performance in audio perception tasks, large audio-language models (AudioLLMs) remain opaque to interpretation. A major factor behind this lack of interpretability is that individual neurons in these models frequently activate in response to several unrelated concepts. We introduce the first mechanistic interpretability framework for AudioLLMs, leveraging sparse autoencoders (SAEs) to disentangle polysemantic activations into monosemantic features. Our pipeline identifies representative audio clips, assigns meaningful names via automated captioning, and validates concepts through human evaluation and steering. Experiments show that AudioLLMs encode structured and interpretable features, enhancing transparency and control. This work provides a foundation for trustworthy deployment in high-stakes domains and enables future extensions to larger models, multilingual audio, and more fine-grained paralinguistic features. Project URL: https://townim-faisal.github.io/AutoInterpret-AudioLLM/


【12】Moving Speaker Separation via Parallel Spectral-Spatial Processing
标题:通过并行谱空间处理实现移动说话人分离
链接:https://arxiv.org/abs/2602.22487

作者:Yuzhu Wang,Archontis Politis,Konstantinos Drossos,Tuomas Virtanen
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing
摘要:动态环境中的多通道语音分离是具有挑战性的时变空间和频谱特征演变在不同的时间尺度。现有方法通常采用顺序架构,迫使单个网络流同时对两种特征类型进行建模,从而产生固有的建模冲突。在本文中,我们提出了一个双分支并行光谱空间(PS2)架构,分别处理光谱和空间特征,通过并行流。频谱分支使用基于双向长短期记忆(BLSTM)的频率模块、基于Mamba的时间模块和自注意模块来对频谱特征进行建模。空间分支采用双向门控递归单元(BGRU)网络来处理空间特征,该空间特征对源和麦克风之间的演变几何关系进行编码。来自两个分支的特征通过交叉注意融合机制进行集成,该机制自适应地对它们的贡献进行加权。实验结果表明,PS2优于现有的国家的最先进的(SOTA)方法的1.6-2.2 dB的尺度不变的信号失真比(SI-SDR)的移动扬声器的情况下,具有强大的分离质量在不同的混响时间(RT 60),噪声水平,和源移动速度。即使在快速源移动的情况下,所提出的模型也保持了超过13 dB的SI-SDR改进。这些改进在包括WHAMR在内的多个数据集上得到了一致的观察!和我们生成的WSJ 0-Demand-6ch-Move数据集。
摘要:Multi-channel speech separation in dynamic environments is challenging as time-varying spatial and spectral features evolve at different temporal scales. Existing methods typically employ sequential architectures, forcing a single network stream to simultaneously model both feature types, creating an inherent modeling conflict. In this paper, we propose a dual-branch parallel spectral-spatial (PS2) architecture that separately processes spectral and spatial features through parallel streams. The spectral branch uses a bi-directional long short-term memory (BLSTM)-based frequency module, a Mamba-based temporal module, and a self-attention module to model spectral features. The spatial branch employs bi-directional gated recurrent unit (BGRU) networks to process spatial features that encode the evolving geometric relationships between sources and microphones. Features from both branches are integrated through a cross-attention fusion mechanism that adaptively weights their contributions. Experimental results demonstrate that the PS2 outperforms existing state-of-the-art (SOTA) methods by 1.6-2.2 dB in scale-invariant signal-to-distortion ratio (SI-SDR) for moving speaker scenarios, with robust separation quality under different reverberation times (RT60), noise levels, and source movement speeds. Even with fast source movements, the proposed model maintains SI-SDR improvements of over 13 dB. These improvements are consistently observed across multiple datasets, including WHAMR! and our generated WSJ0-Demand-6ch-Move dataset.


【13】Learning to reconstruct from saturated data: audio declipping and high-dynamic range imaging
标题:学习从饱和数据中重建:音频去唇和高动态范围成像
链接:https://arxiv.org/abs/2602.22279

作者:Victor Sechaud,Laurent Jacques,Patrice Abry,Julián Tachella
摘要:基于学习的方法现在普遍用于解决逆问题,但它们在实际应用中的部署往往受到缺乏地面真值参考的阻碍。最近的自我监督学习策略提供了一个有希望的替代方案,避免了对地面真相的需要。然而,大多数现有的方法仅限于线性反问题。这项工作将自监督学习扩展到从剪切测量恢复音频和图像的非线性问题,假设信号分布对幅度的变化近似不变。我们提供了充分的条件,学习重建饱和信号单独和自我监督的损失,可用于训练重建网络。在音频和图像数据上的实验表明,所提出的方法几乎与完全监督的方法一样有效,尽管仅依赖于裁剪的测量值进行训练。
摘要:Learning based methods are now ubiquitous for solving inverse problems, but their deployment in real-world applications is often hindered by the lack of ground truth references for training. Recent self-supervised learning strategies offer a promising alternative, avoiding the need for ground truth. However, most existing methods are limited to linear inverse problems. This work extends self-supervised learning to the non-linear problem of recovering audio and images from clipped measurements, by assuming that the signal distribution is approximately invariant to changes in amplitude. We provide sufficient conditions for learning to reconstruct from saturated signals alone and a self-supervised loss that can be used to train reconstruction networks. Experiments on both audio and image data show that the proposed approach is almost as effective as fully supervised approaches, despite relying solely on clipped measurements for training.


eess.AS音频处理


【1】Align-Consistency: Improving Non-autoregressive and Semi-supervised ASR with Consistency Regularization
标题:对齐一致性:通过一致性正规化改进非自回归和半监督的ASB
链接:https://arxiv.org/abs/2602.23171

作者:Wanting Huang,Weiran Wang
备注:In submission to Interspeech 2026
摘要:一致性正则化(CR)通过确保预测在输入扰动下保持稳定,提高了连接主义时间分类(CTC)的鲁棒性和准确性。在这项工作中,我们提出了Align-Consistency,这是CR的扩展,专为Align-Refine设计-一种非自回归(非AR)模型,可对帧级假设进行迭代改进。这种方法利用了并行推理的速度,同时显著提高了识别性能。Align-Consistency的有效性在两种设置中得到证明。首先,在全监督设置中,我们的结果表明,将CR应用于基础CTC模型和后续的细化步骤是至关重要的,并且非AR解码和CR的准确性提高是相互叠加的。其次,对于半监督ASR,我们采用快速非AR解码来在未标记数据上生成在线伪标签,这些伪标签用于进一步改进监督模型并带来实质性收益。
摘要:Consistency regularization (CR) improves the robustness and accuracy of Connectionist Temporal Classification (CTC) by ensuring predictions remain stable across input perturbations. In this work, we propose Align-Consistency, an extension of CR designed for Align-Refine -- a non-autoregressive (non-AR) model that performs iterative refinement of frame-level hypotheses. This method leverages the speed of parallel inference while significantly boosting recognition performance. The effectiveness of Align-Consistency is demonstrated in two settings. First, in the fully supervised setting, our results indicate that applying CR to both the base CTC model and the subsequent refinement steps is critical, and the accuracy improvements from non-AR decoding and CR are mutually additive. Second, for semi-supervised ASR, we employ fast non-AR decoding to generate online pseudo-labels on unlabeled data, which are used to further refine the supervised model and lead to substantial gains.


【2】A Directional-Derivative-Constrained Method for Continuously Steerable Differential Beamformers with Uniform Circular Arrays
标题:均匀圆阵连续可调差差束形成器的方向求导约束方法
链接:https://arxiv.org/abs/2602.23119

作者:Tiantian Xiong,Yongyi Deng,Kunlong Zhao,Jilu Jin,Xueqin Luo,Gongping Huang,Jingdong Chen,Jacob Benesty
摘要:差分传声器阵列由于其高空间指向性和紧凑的阵列结构,为远场声信号采集提供了一种很有前途的解决方案。一个关键的挑战在于设计差分波束形成器,是连续可操纵的,能够增强来自任意方向的目标信号。本文研究了圆阵差分波束形成器的设计,提出了一种新的框架,结合方向导数约束。通过将波束图案在期望的转向方向上的一阶导数约束为零并将合适的值分配给高阶导数,确保波束形成器在目标方向上实现其最大响应并提供足够的波束转向。这种方法不仅提高了转向灵活性,而且还实现了更直观和鲁棒的波束图案设计。仿真结果表明,所提出的方法产生连续可控的波束方向图。
摘要:Differential microphone arrays offer a promising solution for far-field acoustic signal acquisition due to their high spatial directivity and compact array structure. A key challenge lies in designing differential beamformers that are continuously steerable and capable of enhancing target signals arriving from arbitrary directions. This paper studies the design of differential beamformers for circular arrays and proposes a novel framework that incorporates directional derivative constraints. By constraining the first-order derivatives of the beampattern at the desired steering direction to zero and assigning suitable values to higher-order derivatives, the beamformer is ensured to achieve its maximum response in the target direction and provide sufficient beam steering. This approach not only improves steering flexibility but also enables a more intuitive and robust beampattern design. Simulation results demonstrate that the proposed method produces continuously steerable beampatterns.


【3】Scattering Transform for Auditory Attention Decoding
标题:听觉注意力解码的散射变换
链接:https://arxiv.org/abs/2602.23003

作者:René Pallenberg,Fabrice Katzberg,Alfred Mertins,Marco Maass
备注:This work has been submitted to the IEEE for possible publication
摘要:在未来几年,由于人口结构的变化,助听器的使用将增加。新一代助听器仍有待解决的一个开放性问题是鸡尾酒会问题。一个可能的解决方案是基于脑电图的听觉注意解码。这是近年来几项研究的主题,它们的共同点是在大多数情况下使用相同的预处理方法。在这项工作中,为了实现的优势,提出了使用散射变换作为替代这些预处理方法。比较了两层散射变换与常规滤波器组、同步压缩短时傅里叶变换和常用预处理方法。为了证明性能,已知的和建议的预处理方法进行了比较,为不同的分类任务上的两个广泛使用的数据集,由KU鲁汶(KUL)和丹麦技术大学(DTU)。已建立的和新的基于神经网络的模型、CNN、LSTM和最近的基于Transformer/图形的模型都用于分类。不同的评价策略进行了比较,重点是谁是未知的训练的发言人进行分类的任务。我们表明,双层散射变换可以显着提高性能的主题相关的条件下,特别是在KUL数据集。然而,在DTU数据集上,这仅适用于某些模型,或者当提供大量训练数据时,如10倍交叉验证。这表明散射变换能够提取额外的相关信息。
摘要:The use of hearing aids will increase in the coming years due to demographic change. One open problem that remains to be solved by a new generation of hearing aids is the cocktail party problem. A possible solution is electroencephalography-based auditory attention decoding. This has been the subject of several studies in recent years, which have in common that they use the same preprocessing methods in most cases. In this work, in order to achieve an advantage, the use of a scattering transform is proposed as an alternative to these preprocessing methods. The two-layer scattering transform is compared with a regular filterbank, the synchrosqueezing short-time Fourier transform and the common preprocessing. To demonstrate the performance, the known and the proposed preprocessing methods are compared for different classification tasks on two widely used datasets, provided by the KU Leuven (KUL) and the Technical University of Denmark (DTU). Both established and new neural-network-based models, CNNs, LSTMs, and recent Transformer/graph-based models are used for classification. Various evaluation strategies were compared, with a focus on the task of classifying speakers who are unknown from the training. We show that the two-layer scattering transform can significantly improve the performance for subject-related conditions, especially on the KUL dataset. However, on the DTU dataset, this only applies to some of the models, or when larger amounts of training data are provided, as in 10-fold cross-validation. This suggests that the scattering transform is capable of extracting additional relevant information.


【4】Deepfake Word Detection by Next-token Prediction using Fine-tuned Whisper
标题:使用微调Whisper通过下一个令牌预测进行Deepfake单词检测
链接:https://arxiv.org/abs/2602.22658

作者:Hoan My Tran,Xin Wang,Wanying Ge,Xuechen Liu,Junichi Yamagishi
摘要:Deepfake语音话语可以通过用语音生成模型合成的语义不同的词替换真实话语中的一个或多个词来伪造。虽然可以开发专用的合成单词检测器,但我们研究了一种具有成本效益的方法,该方法可以微调预训练的Whisper模型,以检测合成单词,同时通过下一个令牌预测转录输入话语。我们进一步研究使用部分声码的话语作为微调数据,从而降低数据收集的成本。我们的实验表明,在域内测试数据,微调耳语产生低的合成词检测错误率和转录错误率。在由看不见的语音生成模型生成的合成词的域外测试数据上,微调的Whisper仍然与专用的基于ResNet的检测模型相当;然而,整体性能下降需要提高其泛化能力的策略。
摘要:Deepfake speech utterances can be forged by replacing one or more words in a bona fide utterance with semantically different words synthesized by speech generative models. While a dedicated synthetic word detector could be developed, we investigate a cost-effective method that fine-tunes a pre-trained Whisper model to detect synthetic words while transcribing the input utterance via next-token prediction. We further investigate using partially vocoded utterances as the fine-tuning data, thereby reducing the cost of data collection. Our experiments demonstrate that, on in-domain test data, the fine-tuned Whisper yields low synthetic-word detection error rates and transcription error rates. On out-of-domain test data with synthetic words produced by unseen speech generative models, the fine-tuned Whisper remains on par with a dedicated ResNet-based detection model; however, the overall performance degradation calls for strategies to improve its generalization capability.


【5】Moving Speaker Separation via Parallel Spectral-Spatial Processing
标题:通过并行谱空间处理实现移动说话人分离
链接:https://arxiv.org/abs/2602.22487

作者:Yuzhu Wang,Archontis Politis,Konstantinos Drossos,Tuomas Virtanen
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing
摘要:动态环境中的多通道语音分离是具有挑战性的时变空间和频谱特征演变在不同的时间尺度。现有方法通常采用顺序架构,迫使单个网络流同时对两种特征类型进行建模,从而产生固有的建模冲突。在本文中,我们提出了一个双分支并行光谱空间(PS2)架构,分别处理光谱和空间特征,通过并行流。频谱分支使用基于双向长短期记忆(BLSTM)的频率模块、基于Mamba的时间模块和自注意模块来对频谱特征进行建模。空间分支采用双向门控递归单元(BGRU)网络来处理空间特征,该空间特征对源和麦克风之间的演变几何关系进行编码。来自两个分支的特征通过交叉注意融合机制进行集成,该机制自适应地对它们的贡献进行加权。实验结果表明,PS2优于现有的国家的最先进的(SOTA)方法的1.6-2.2 dB的尺度不变的信号失真比(SI-SDR)的移动扬声器的情况下,具有强大的分离质量在不同的混响时间(RT 60),噪声水平,和源移动速度。即使在快速源移动的情况下,所提出的模型也保持了超过13 dB的SI-SDR改进。这些改进在包括WHAMR在内的多个数据集上得到了一致的观察!和我们生成的WSJ 0-Demand-6ch-Move数据集。
摘要:Multi-channel speech separation in dynamic environments is challenging as time-varying spatial and spectral features evolve at different temporal scales. Existing methods typically employ sequential architectures, forcing a single network stream to simultaneously model both feature types, creating an inherent modeling conflict. In this paper, we propose a dual-branch parallel spectral-spatial (PS2) architecture that separately processes spectral and spatial features through parallel streams. The spectral branch uses a bi-directional long short-term memory (BLSTM)-based frequency module, a Mamba-based temporal module, and a self-attention module to model spectral features. The spatial branch employs bi-directional gated recurrent unit (BGRU) networks to process spatial features that encode the evolving geometric relationships between sources and microphones. Features from both branches are integrated through a cross-attention fusion mechanism that adaptively weights their contributions. Experimental results demonstrate that the PS2 outperforms existing state-of-the-art (SOTA) methods by 1.6-2.2 dB in scale-invariant signal-to-distortion ratio (SI-SDR) for moving speaker scenarios, with robust separation quality under different reverberation times (RT60), noise levels, and source movement speeds. Even with fast source movements, the proposed model maintains SI-SDR improvements of over 13 dB. These improvements are consistently observed across multiple datasets, including WHAMR! and our generated WSJ0-Demand-6ch-Move dataset.


【6】A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations
标题:对话中多模式情感识别的专家混合模型
链接:https://arxiv.org/abs/2602.23300

作者:Soumya Dutta,Smruthi Balaji,Sriram Ganapathy
备注:Accepted to Elsevier Computer Speech and Language. 30 pages, 9 figures, 5 tables
摘要:会话中的情感识别(ERC)提出了独特的挑战,需要模型捕捉多轮对话的时间流,并有效地整合来自多种模态的线索。我们提出了混合语音文本专家识别的情绪(MISTER-E),一个模块化的混合专家(MoE)框架,旨在解耦ERC的两个核心挑战:特定模态的上下文建模和多模态信息融合。MiSTER-E利用针对语音和文本进行微调的大型语言模型(LLM)来提供丰富的话语级嵌入,然后通过卷积递归上下文建模层进行增强。该系统集成了来自三个专家的预测-语音,纯文本和跨模态-使用一个学习门控机制,动态权衡他们的输出。为了进一步鼓励跨模态的一致性和对齐,我们在成对的语音-文本表示之间引入了有监督的对比损失,并在专家预测中引入了基于KL发散的正则化。重要的是,MISTER-E在任何阶段都不依赖于说话者身份。在IEMOCAP、MELD和MOSI三个标准测试集上的实验表明,该算法分别获得了70.9%、69.5%和87.9%的加权F1分数,优于几个基本的语音-文本ERC系统。我们还提供了各种消融,以突出所提出的方法中所做的贡献。
摘要:Emotion Recognition in Conversations (ERC) presents unique challenges, requiring models to capture the temporal flow of multi-turn dialogues and to effectively integrate cues from multiple modalities. We propose Mixture of Speech-Text Experts for Recognition of Emotions (MiSTER-E), a modular Mixture-of-Experts (MoE) framework designed to decouple two core challenges in ERC: modality-specific context modeling and multimodal information fusion. MiSTER-E leverages large language models (LLMs) fine-tuned for both speech and text to provide rich utterance-level embeddings, which are then enhanced through a convolutional-recurrent context modeling layer. The system integrates predictions from three experts-speech-only, text-only, and cross-modal-using a learned gating mechanism that dynamically weighs their outputs. To further encourage consistency and alignment across modalities, we introduce a supervised contrastive loss between paired speech-text representations and a KL-divergence-based regulariza-tion across expert predictions. Importantly, MiSTER-E does not rely on speaker identity at any stage. Experiments on three benchmark datasets-IEMOCAP, MELD, and MOSI-show that our proposal achieves 70.9%, 69.5%, and 87.9% weighted F1-scores respectively, outperforming several baseline speech-text ERC systems. We also provide various ablations to highlight the contributions made in the proposed approach.


【7】Make It Hard to Hear, Easy to Learn: Long-Form Bengali ASR and Speaker Diarization via Extreme Augmentation and Perfect Alignment
标题:让它听起来很难,易于学习:通过极端增强和完美对齐实现长格式孟加拉语ASB和扬声器扩展
链接:https://arxiv.org/abs/2602.23070

作者:Sanjid Hasan,Risalat Labib,A H M Fuad,Bayazid Hasan
备注:4 pages, 2 figures
摘要:虽然孟加拉语的自动语音识别(ASR)已经取得了重大进展,但处理长时间音频和执行鲁棒的说话人日志仍然是关键的研究空白。为了解决这种语言的联合ASR和日记资源的严重缺乏,我们引入了Lipi-Ghor-882,一个全面的882小时多说话者孟加拉语数据集。在本文中,详细介绍了我们提交的DL Sprint 4.0比赛,我们系统地评估各种架构和方法的长格式孟加拉语讲话。对于ASR,我们证明了原始数据缩放是无效的;相反,利用完美对齐的注释与合成声学退化(噪声和混响)配对的有针对性的微调成为最有效的方法。相反,对于说话人日记,我们观察到全球开源的最先进的模型(如Diarizen)在这个复杂的数据集上表现得非常差。广泛的模型再训练产生了微不足道的改进;相反,基线模型输出的战略性启发式后处理被证明是提高准确性的主要驱动力。最终,这项工作概述了一个高度优化的双管道,实现了0.019美元的实时因子(RTF),建立了一个实用的,经验支持的低资源,长形式的语音处理基准。
摘要:Although Automatic Speech Recognition (ASR) in Bengali has seen significant progress, processing long-duration audio and performing robust speaker diarization remain critical research gaps. To address the severe scarcity of joint ASR and diarization resources for this language, we introduce Lipi-Ghor-882, a comprehensive 882-hour multi-speaker Bengali dataset. In this paper, detailing our submission to the DL Sprint 4.0 competition, we systematically evaluate various architectures and approaches for long-form Bengali speech. For ASR, we demonstrate that raw data scaling is ineffective; instead, targeted fine-tuning utilizing perfectly aligned annotations paired with synthetic acoustic degradation (noise and reverberation) emerges as the singular most effective approach. Conversely, for speaker diarization, we observed that global open-source state-of-the-art models (such as Diarizen) performed surprisingly poorly on this complex dataset. Extensive model retraining yielded negligible improvements; instead, strategic, heuristic post-processing of baseline model outputs proved to be the primary driver for increasing accuracy. Ultimately, this work outlines a highly optimized dual pipeline achieving a $\sim$0.019 Real-Time Factor (RTF), establishing a practical, empirically backed benchmark for low-resource, long-form speech processing.


【8】Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech
标题:将发声、模仿和想象语音的神经表示联系起来
链接:https://arxiv.org/abs/2602.22597

作者:Maryam Maghsoudi,Rupesh Chillale,Shihab A. Shamma
摘要:我们研究了使用公开的立体定向脑电图记录的发声,默拟和想象的语音神经表征之间的关系。大多数先前的研究都集中在解码语音反应在每个条件下分别。相反,在这里,我们通过为每个条件训练线性谱图重建模型来探索不同条件下的响应如何相关,并评估它们在不同条件下的泛化。我们证明了在一个条件下训练的线性解码器通常成功地转移到其他人,这意味着共享的语音表示。这种共性进行了评估与刺激水平的可辨别性,通过执行基于秩的分析,证明在内部和跨条件下的刺激特异性结构的保存。最后,我们将线性重建与非线性神经网络的重建进行了比较。虽然两者都表现出交叉条件转移,但线性模型实现了更好的刺激水平辨别力。
摘要:We investigated the relationship among neural representations of vocalized, mimed, and imagined speech recorded using publicly available stereotactic EEG recordings. Most prior studies have focused on decoding speech responses within each condition separately. Here, instead, we explore how responses across conditions relate by training linear spectrogram reconstruction models for each condition and evaluate their generalization across conditions. We demonstrate that linear decoders trained on one condition generally transfer successfully to others, implying shared speech representations. This commonality was assessed with stimulus-level discriminability by performing a rank-based analysis demonstrating preservation of stimulus-specific structure in both within- and across-conditions. Finally, we compared linear reconstructions to those from a nonlinear neural network. While both exhibited cross-condition transfer, linear models achieve superior stimulus-level discriminability.


【9】Efficient Dialect-Aware Modeling and Conditioning for Low-Resource Taiwanese Hakka Speech Processing
标题:低资源台湾客语语音处理的高效方言感知建模和条件处理
链接:https://arxiv.org/abs/2602.22522

作者:An-Ci Peng,Kuan-Tang Huang,Tien-Hong Lo,Hung-Shin Lee,Hsin-Min Wang,Berlin Chen
备注:Accepted to LREC 2026
摘要:台湾客家语是一种资源匮乏的濒危语言,对自动语音识别(ASR)提出了重大挑战,包括方言的高度变异性和两种不同的书写系统(汉字和拼音)的存在。传统的ASR模型在这种情况下往往会遇到困难,因为它们往往会将基本的语言内容与方言特有的语音和词汇方面的变化混为一谈。为了应对这些挑战,我们提出了一个基于递归神经网络传感器(RNN-T)的统一框架。我们的方法的核心是引入方言感知的建模策略,旨在解开方言的“风格”从语言的“内容”,这提高了模型的能力,学习强大的和广义的表示。此外,该框架采用参数有效的预测网络,同时建模ASR(汉字和拼音)。我们证明,这些任务创建一个强大的协同作用,其中的交叉脚本目标作为一个相互正则化,以提高主要ASR任务。在HAT语料库上进行的实验表明,该模型对汉字和拼音ASR的相对错误率分别降低了57.00%和40.41%。据我们所知,这是第一次系统地调查客家方言变异对ASR的影响,也是第一个能够共同解决这些任务的单一模型。
摘要:Taiwanese Hakka is a low-resource, endangered language that poses significant challenges for automatic speech recognition (ASR), including high dialectal variability and the presence of two distinct writing systems (Hanzi and Pinyin). Traditional ASR models often encounter difficulties in this context, as they tend to conflate essential linguistic content with dialect-specific variations across both phonological and lexical dimensions. To address these challenges, we propose a unified framework grounded in the Recurrent Neural Network Transducers (RNN-T). Central to our approach is the introduction of dialect-aware modeling strategies designed to disentangle dialectal "style" from linguistic "content", which enhances the model's capacity to learn robust and generalized representations. Additionally, the framework employs parameter-efficient prediction networks to concurrently model ASR (Hanzi and Pinyin). We demonstrate that these tasks create a powerful synergy, wherein the cross-script objective serves as a mutual regularizer to improve the primary ASR tasks. Experiments conducted on the HAT corpus reveal that our model achieves 57.00% and 40.41% relative error rate reduction on Hanzi and Pinyin ASR, respectively. To our knowledge, this is the first systematic investigation into the impact of Hakka dialectal variations on ASR and the first single model capable of jointly addressing these tasks.


【10】Absorbing Discrete Diffusion for Speech Enhancement
标题:吸收离散扩散进行语音增强
链接:https://arxiv.org/abs/2602.22417

作者:Philippe Gonzalez
备注:Submitted to Interspeech 2026
摘要:受神经语音编码和基于扩散的语言建模领域最新发展的启发,我们通过使用吸收离散扩散对给定有噪语音代码的干净语音代码的条件分布进行建模来解决语音增强问题。所提出的方法,我们称之为ADDSE,利用神经音频编解码器的表达潜在空间和扩散模型的非自回归采样过程。为了有效地建模残差矢量量化码的层次结构,我们提出了RQDiT,它结合了RQ-变压器和扩散Transformers的非自回归建模技术。结果表明,在两个数据集上的非侵入性客观指标方面具有竞争力的性能,特别是在低信噪比和很少的采样步骤。代码和音频示例可在线获得。
摘要:Inspired by recent developments in neural speech coding and diffusion-based language modeling, we tackle speech enhancement by modeling the conditional distribution of clean speech codes given noisy speech codes using absorbing discrete diffusion. The proposed approach, which we call ADDSE, leverages both the expressive latent space of neural audio codecs and the non-autoregressive sampling procedure of diffusion models. To efficiently model the hierarchical structure of residual vector quantization codes, we propose RQDiT, which combines techniques from RQ-Transformer and diffusion Transformers for non-autoregressive modeling. Results show competitive performance in terms of non-intrusive objective metrics on two datasets, especially at low signal-to-noise ratios and with few sampling steps. Code and audio examples are available online.


机器翻译由腾讯交互翻译提供,仅供参考