今日论文合集:cs.SD语音10篇,eess.AS音频处理6篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake Detection
标题:HyperPotter:诠释音频深度伪造检测中高级交互的魅力
链接:https://arxiv.org/pdf/2602.05670v1

作者:Qing Wen,Haohao Li,Zhongjie Ba,Peng Cheng,Miao He,Li Lu,Kui Ren

备注:20 pages, 8 figures

摘要:AIGC技术的进步使得能够合成能够欺骗人类听觉感知的高度逼真的音频deepfake。虽然已经开发了许多音频深度伪造检测(ADD)方法,但大多数依赖于局部时间 频谱特征或成对关系,忽略了高阶交互(HOI)。HOI捕获从超出其单独贡献的多个特征分量中出现的区分模式。我们提出了HyperPotter,一个基于超图的框架,明确地通过基于聚类的hyperedges与类感知原型初始化这些协同HOI模型。大量的实验表明,HyperPotter在11个数据集上的平均相对增益超过其基线22.15%,在4个具有挑战性的跨域数据集上的性能优于最先进的方法13.96%,表现出对不同攻击和说话者的卓越泛化能力。摘要:Advances in AIGC technologies have enabled the synthesis of highly realistic audio deepfakes capable of deceiving human auditory perception. Although numerous audio deepfake detection (ADD) methods have been developed, most rely on local temporal spectral features or pairwise relations, overlooking high-order interactions (HOIs). HOIs capture discriminative patterns that emerge from multiple feature components beyond their individual contributions. We propose HyperPotter, a hypergraph-based framework that explicitly models these synergistic HOIs through clustering-based hyperedges with class-aware prototype initialization. Extensive experiments demonstrate that HyperPotter surpasses its baseline by an average relative gain of 22.15% across 11 datasets and outperforms state-of-the-art methods by 13.96% on 4 challenging cross-domain datasets, demonstrating superior generalization to diverse attacks and speakers.


【2】Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language
标题:启用自动无序语音识别:阿坎语言中受损的语音数据集
链接:https://arxiv.org/pdf/2602.05406v1

作者:Isaac Wiafe,Akon Obu Ekpezu,Sumaya Ahmed Salihs,Elikem Doe Atsakpo,Fiifi Baffoe Payin Winful,Jamal-Deen Abdulai
摘要:受损语音数据的缺乏阻碍了包容性语音技术的发展,特别是在阿肯语等低资源语言中。为了解决这一差距,本研究提出了一个策划语料库的语音样本,从母语阿肯语说话的语音障碍。该数据集包括50.01小时的音频记录,跨越四类受损的语音,即口吃,脑瘫,腭裂和中风引起的语音障碍。记录是在受控的监督环境中完成的,参与者用自己的话描述预先选择的图像。所得到的数据集是音频记录、转录和关于说话者人口统计、损伤类别、记录环境和设备的相关元数据的集合。该数据集旨在支持低资源自动无序语音识别系统和辅助语音技术的研究。摘要:The lack of impaired speech data hinders advancements in the development of inclusive speech technologies, particularly in low-resource languages such as Akan. To address this gap, this study presents a curated corpus of speech samples from native Akan speakers with speech impairment. The dataset comprises of 50.01 hours of audio recordings cutting across four classes of impaired speech namely stammering, cerebral palsy, cleft palate, and stroke induced speech disorder. Recordings were done in controlled supervised environments were participants described pre-selected images in their own words. The resulting dataset is a collection of audio recordings, transcriptions, and associated metadata on speaker demographics, class of impairment, recording environment and device. The dataset is intended to support research in low-resource automatic disordered speech recognition systems and assistive speech technology.


【3】Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
标题:Speech-XL:在大型语音语言模型中实现长篇语音理解
链接:https://arxiv.org/pdf/2602.05373v1

作者:Haoqin Sun,Chenyang Lyu,Shiwan Zhao,Xuanfan Ni,Xiangyu Kong,Longyue Wang,Weihua Luo,Yong Qin
摘要:尽管大型语音语言模型(LSLMs)在处理短期声学信号方面取得了越来越大的成功,但它们对长形式音频理解的扩展受到了严重的挑战。这种限制源于有限的上下文长度和长形式推理所需的过高的内存占用。在这项工作中,我们提出了Speech-XL,一种新的模型,利用大语言模型(LLM)的内在键值(KV)稀疏化能力,以实现高比率的语音输入压缩。具体来说,我们引入了一种新的特殊令牌,语音摘要令牌(SST),为每个语音间隔封装的区间内语音信息到其相关的KV对。SST模块通过指令微调进行训练,采用课程学习策略,SST学习以渐进的方式压缩信息-从低比率(简单)到高比率(具有挑战性)压缩。尽管使用的训练数据比其他基线少得多,但我们的模型在主要基准测试(包括LongSpeech和AUDIOMARATHON)上表现出非常有竞争力的性能。通过解决长期存在的瓶颈,在长形式的音频建模,我们的方法提供了一个新的视角上的冷凝广泛的声学序列。摘要:Despite the growing success of Large Speech Language Models (LSLMs) in processing short-term acoustic signals, their extension to long-form audio understanding is severely bottlenecked. This limitation stems from the limited context length and the exorbitant memory footprints required for long-form inference. In this work, we propose Speech-XL, a new model that capitalizes on the intrinsic key-value (KV) sparsification capacity of Large Language Models (LLMs) to achieve high-ratio speech input compression. Specifically, we introduce a novel special token, the Speech Summarization Token (SST), for each speech interval to encapsulate the intra-interval speech information into its associated KV pairs. The SST module is trained via instruction fine-tuning, employing a curriculum learning strategy where the SST learns to compress information in a progressive manner--advancing from low-ratio (simple) to high-ratio (challenging) compression. Despite utilizing significantly less training data than other baselines, our model achieves highly competitive performance on major benchmarks, including LongSpeech and AUDIOMARATHON. By addressing the long-standing bottlenecks in long-form audio modeling, our approach offers a novel perspective on the condensation of extensive acoustic sequences.


【4】Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
标题:风笛手:通过丰富的字幕解决开放式音频任务
链接:https://arxiv.org/pdf/2602.05220v1

作者:Jinchuan Tian,Haoran Wang,Bo-Hao Su,Chien-yu Huang,Qingzheng Wang,Jiatong Shi,William Chen,Xun Gong,Siddhant Arora,Chin-Jou Li,Masao Someki,Takashi Maekaku,Yusuke Shinohara,Jin Sakuma,Chao-Han Huck Yang,Shinji Watanabe
摘要:目前的音频基础模型通常依赖于严格的、特定于任务的监督,解决音频的孤立因素,而不是整体。相比之下,人类智能整体地处理音频,无缝地将物理信号与抽象的认知概念联系起来,以执行复杂的任务。基于这一理念,我们推出了Bagpiper,一个8B音频基础模型,通过丰富的字幕来解释物理音频,即,封装信号中固有的关键认知概念的综合自然语言描述(例如,转录、音频事件)。通过在600 B令牌的大规模语料库上进行预训练,该模型在原始音频和这个高级概念空间之间建立了一个强大的双向映射。在微调过程中,Bagpiper采用先字幕后处理的工作流程,模拟中间认知推理步骤来解决不同的任务,而无需特定于任务的先验知识。在实验上,Bagpiper在MMAU和AIRBench上的音频理解性能优于Qwen-2.5-Omni,在生成质量上超过CosyVoice 3和TangoFlux,能够合成语音、音乐和音效的任意组合。据我们所知,Bagpiper是最早实现通用音频统一理解生成的作品之一。模型、数据和代码可在Bagpiper主页上获得。摘要:Current audio foundation models typically rely on rigid, task-specific supervision, addressing isolated factors of audio rather than the whole. In contrast, human intelligence processes audio holistically, seamlessly bridging physical signals with abstract cognitive concepts to execute complex tasks. Grounded in this philosophy, we introduce Bagpiper, an 8B audio foundation model that interprets physical audio via rich captions, i.e., comprehensive natural language descriptions that encapsulate the critical cognitive concepts inherent in the signal (e.g., transcription, audio events). By pre-training on a massive corpus of 600B tokens, the model establishes a robust bidirectional mapping between raw audio and this high-level conceptual space. During fine-tuning, Bagpiper adopts a caption-then-process workflow, simulating an intermediate cognitive reasoning step to solve diverse tasks without task-specific priors. Experimentally, Bagpiper outperforms Qwen-2.5-Omni on MMAU and AIRBench for audio understanding and surpasses CosyVoice3 and TangoFlux in generation quality, capable of synthesizing arbitrary compositions of speech, music, and sound effects. To the best of our knowledge, Bagpiper is among the first works that achieve unified understanding generation for general audio. Model, data, and code are available at Bagpiper Home Page.


【5】AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders
标题:AudioSAE:理解具有稀疏自动编码器的音频处理模型
链接:https://arxiv.org/pdf/2602.05027v1

作者:Georgii Aparin,Tasnima Sadekova,Alexey Rukhovich,Assel Yermekova,Laida Kushnareva,Vadim Popov,Kristian Kuznetsov,Irina Piontkovskaya
摘要:稀疏自动编码器(SAE)是解释神经表征的强大工具,但它们在音频中的使用仍有待探索。我们在Whisper和HuBERT的所有编码器层上训练SAE,对它们的稳定性和可解释性进行广泛的评估,并展示它们的实用性。超过50%的特征在随机种子中保持一致,并且保留了重建质量。SAE特征捕捉一般的声学和语义信息以及特定的事件,包括环境噪音和非语言声音(例如笑声,耳语),并有效地将它们分离,只需要删除19-27%的特征就可以删除一个概念。特征转向将Whisper的错误语音检测减少了70%,WER的增加可以忽略不计,证明了现实世界的适用性。最后,我们发现SAE功能与人类脑电活动在语音感知,表明对齐与人类神经处理。代码和检查点可在https: github.com audiosae audiosae_demo上获得。摘要:Sparse Autoencoders (SAEs) are powerful tools for interpreting neural representations, yet their use in audio remains underexplored. We train SAEs across all encoder layers of Whisper and HuBERT, provide an extensive evaluation of their stability, interpretability, and show their practical utility. Over 50% of the features remain consistent across random seeds, and reconstruction quality is preserved. SAE features capture general acoustic and semantic information as well as specific events, including environmental noises and paralinguistic sounds (e.g. laughter, whispering) and disentangle them effectively, requiring removal of only 19-27% of features to erase a concept. Feature steering reduces Whisper's false speech detections by 70% with negligible WER increase, demonstrating real-world applicability. Finally, we find SAE features correlated with human EEG activity during speech perception, indicating alignment with human neural processing. The code and checkpoints are available at https: github.com audiosae audiosae_demo.


【6】Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering
标题:知道何时回答:可靠的视听问题回答的自适应信心细化
链接:https://arxiv.org/pdf/2602.04924v1

作者:Dinh Phu Tran,Jihoon Jeong,Saad Wazir,Seongah Kim,Thao Do,Cem Subakan,Daeyoung Kim
备注:Technical Report
摘要:我们提出了一个正式的问题制定 textit{可靠}视听问题问答($ mathcal{R}$-AVQA),我们更喜欢回答错误。虽然最近的AVQA模型具有很高的准确性,但它们识别何时可能出错以及随后避免回答的能力仍然是研究领域的不足之处。为了填补这一空白,我们探索了几种方法,然后提出了自适应置信度细化(ACR),一个轻量级的方法,以进一步提高性能的$ mathcal{R}$-AVQA。我们的关键见解是,最大软最大概率(MSP)只有在强校准下才是贝叶斯最优的,这是深度神经网络通常不满足的条件,特别是在多模态模型中。而不是取代MSP,我们的ACR保持它作为一个主要的信心信号,并适用于输入自适应残差校正时,MSP被视为不可靠。ACR引入了两个学习头:i)预测MSP未捕获的低幅度正确性残差的残差风险头,以及ii)确定MSP可信度的置信度门控头。我们的实验和理论分析表明,ACR一致优于现有的方法,在三个不同的AVQA架构的分布内和分布外,和数据偏差设置,建立了坚实的基础,$ mathcal{R}$-AVQA任务。代码和检查点将在接受后提供 href{https: github.com R-AVQA}{at here}摘要:We present a formal problem formulation for textit{Reliable} Audio-Visual Question Answering ($ mathcal{R}$-AVQA), where we prefer abstention over answering incorrectly. While recent AVQA models have high accuracy, their ability to identify when they are likely wrong and their consequent abstention from answering remain underexplored areas of research. To fill this gap, we explore several approaches and then propose Adaptive Confidence Refinement (ACR), a lightweight method to further enhance the performance of $ mathcal{R}$-AVQA. Our key insight is that the Maximum Softmax Probability (MSP) is Bayes-optimal only under strong calibration, a condition usually not met in deep neural networks, particularly in multimodal models. Instead of replacing MSP, our ACR maintains it as a primary confidence signal and applies input-adaptive residual corrections when MSP is deemed unreliable. ACR introduces two learned heads: i) a Residual Risk Head that predicts low-magnitude correctness residuals that MSP does not capture, and ii) a Confidence Gating Head to determine MSP trustworthiness. Our experiments and theoretical analysis show that ACR consistently outperforms existing methods on in- and out-of-disrtibution, and data bias settings across three different AVQA architectures, establishing a solid foundation for $ mathcal{R}$-AVQA task. The code and checkpoints will be available upon acceptance href{https: github.com PhuTran1005 R-AVQA}{at here}


【7】CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal Learning
标题:CyIN:用于桥梁完全和不完全多模式学习的循环信息潜在空间
链接:https://arxiv.org/pdf/2602.04920v1

作者:Ronghao Lin,Qiaolin He,Sijie Mai,Ying Zeng,Aolin Xiong,Li Huang,Yap-Peng Tan,Haifeng Hu
备注:Accepted by NeurIPS 2025
摘要:多模态机器学习,模仿人类大脑整合各种模态的能力,已经出现了快速增长。大多数以前的多模态模型都是在完美配对的多模态输入上训练的,以达到最佳性能。然而,在现实世界的部署中,模态的存在是高度可变和不可预测的,导致预先训练的模型遭受显著的性能下降,并且在动态缺失模态的情况下无法保持鲁棒性。在本文中,我们提出了一种新的循环信息学习框架(CyIN),以弥合完整和不完整的多模态学习之间的差距。具体来说,我们首先建立一个信息的潜在空间,采用令牌和标签级的信息瓶颈(IB)之间的各种形式循环。利用变分近似捕获任务相关的特征,净化信息瓶颈潜伏期,以实现更有效的跨模态交互和多模态融合。此外,为了补充由于不完整的多模态输入而造成的信息缺失,我们提出了跨模态循环翻译,通过正向和反向传播过程用剩余的模态重建缺失的模态。借助提取和重构的信息潜势,CyIN成功地在一个统一的模型中联合优化了完整和不完整的多模态学习。在4个多模态数据集上进行的大量实验表明,我们的方法在完整和不同的不完整场景下都具有优异的性能。摘要:Multimodal machine learning, mimicking the human brain's ability to integrate various modalities has seen rapid growth. Most previous multimodal models are trained on perfectly paired multimodal input to reach optimal performance. In real-world deployments, however, the presence of modality is highly variable and unpredictable, causing the pre-trained models in suffering significant performance drops and fail to remain robust with dynamic missing modalities circumstances. In this paper, we present a novel Cyclic INformative Learning framework (CyIN) to bridge the gap between complete and incomplete multimodal learning. Specifically, we firstly build an informative latent space by adopting token- and label-level Information Bottleneck (IB) cyclically among various modalities. Capturing task-related features with variational approximation, the informative bottleneck latents are purified for more efficient cross-modal interaction and multimodal fusion. Moreover, to supplement the missing information caused by incomplete multimodal input, we propose cross-modal cyclic translation by reconstruct the missing modalities with the remained ones through forward and reverse propagation process. With the help of the extracted and reconstructed informative latents, CyIN succeeds in jointly optimizing complete and incomplete multimodal learning in one unified model. Extensive experiments on 4 multimodal datasets demonstrate the superior performance of our method in both complete and diverse incomplete scenarios.


【8】A$^2$-LLM: An End-to-end Conversational Audio Avatar Large Language Model
标题:A$#2 $-LLM:端到端对话音频化身大型语言模型
链接:https://arxiv.org/pdf/2602.04913v1

作者:Xiaolin Hu,Hang Yuan,Xinzhu Sang,Binbin Yan,Zhou Yu,Cong Huang,Kai Chen
备注:13 pages, 3 figures
摘要:开发富有表现力和响应性的对话数字人是下一代人机交互的基石。虽然大型语言模型(LLM)显著增强了对话能力,但大多数当前系统仍然依赖于连接独立模块的级联架构。这些流水线经常受到累积错误、高延迟和低实时性能的困扰。由于无法访问底层的会话上下文,这些管道固有地优先考虑严格的对口型而不是情感深度。为了解决这些挑战,我们提出了A$^2$-LLM,一个端到端的会话音频化身大语言模型,在一个统一的框架内联合的原因,语言,音频韵律和3D面部运动。为了便于训练,我们引入了FLAME-QA,这是一个高质量的多模态数据集,旨在将语义意图与QA格式中的表情面部动态相结合。通过利用深度语义理解,A$^2$-LLM生成情感丰富的面部动作,而不仅仅是简单的嘴唇同步。实验结果表明,我们的系统实现了卓越的情感表现力,同时保持实时效率(500毫秒的延迟,0.7 RTF)。摘要:Developing expressive and responsive conversational digital humans is a cornerstone of next-generation human-computer interaction. While large language models (LLMs) have significantly enhanced dialogue capabilities, most current systems still rely on cascaded architectures that connect independent modules. These pipelines are often plagued by accumulated errors, high latency, and poor real-time performance. Lacking access to the underlying conversational context, these pipelines inherently prioritize rigid lip-sync over emotional depth. To address these challenges, we propose A$^2$-LLM, an end-to-end conversational audio avatar large language model that jointly reasons about language, audio prosody, and 3D facial motion within a unified framework. To facilitate training, we introduce FLAME-QA, a high-quality multimodal dataset designed to align semantic intent with expressive facial dynamics within a QA format. By leveraging deep semantic understanding, A$^2$-LLM generates emotionally rich facial movements beyond simple lip-synchronization. Experimental results demonstrate that our system achieves superior emotional expressiveness while maintaining real-time efficiency (500 ms latency, 0.7 RTF).


【9】Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech Generation from SSL features
标题:Wave-Trainer-Fit:具有可训练先验和定点迭代的神经声码器,从SSL功能生成高质量语音
链接:https://arxiv.org/pdf/2602.05443v1

作者:Hien Ohnaka,Yuma Shirahata,Masaya Kawamura

备注:Accepted by IEEE ICASSP 2026. 5 pages, 3 figures, and 2 tables

摘要:我们提出了WaveTrainerFit,一种神经声码器,它可以从数据驱动的功能(如SSL功能)中生成高质量的波形。WaveTrainerFit建立在WaveFit声码器的基础上,它集成了扩散模型和生成对抗网络。此外,所提出的方法包括以下关键改进:1。通过引入可训练的先验知识,推理过程从接近目标语音的噪声而不是高斯噪声开始。2.通过在匹配语音能量之前对可训练对象施加约束来执行参考感知增益调整。这些改进预计将降低数据驱动特征的波形建模的复杂性,从而能够以更少的推理步骤生成高质量的波形。通过实验,我们证明了WaveTrainerFit可以从数据驱动的特征中生成高度自然的波形,提高了说话人的相似性,同时需要比WaveFit更少的迭代次数。此外,我们表明,所提出的方法相对于SSL特征提取的深度是鲁棒的。代码和预训练模型可从https: github.com line WaveTrainerFit获得。摘要:We propose WaveTrainerFit, a neural vocoder that performs high-quality waveform generation from data-driven features such as SSL features. WaveTrainerFit builds upon the WaveFit vocoder, which integrates diffusion model and generative adversarial network. Furthermore, the proposed method incorporates the following key improvements: 1. By introducing trainable priors, the inference process starts from noise close to the target speech instead of Gaussian noise. 2. Reference-aware gain adjustment is performed by imposing constraints on the trainable prior to matching the speech energy. These improvements are expected to reduce the complexity of waveform modeling from data-driven features, enabling high-quality waveform generation with fewer inference steps. Through experiments, we showed that WaveTrainerFit can generate highly natural waveforms with improved speaker similarity from data-driven features, while requiring fewer iterations than WaveFit. Moreover, we showed that the proposed method works robustly with respect to the depth at which SSL features are extracted. Code and pre-trained models are available from https: github.com line WaveTrainerFit.


【10】Instantaneous Spectra Analysis of Pulse Series - Application to Lung Sounds with Abnormalities
标题:脉冲序列的瞬时谱分析--应用于异常肺音
链接:https://arxiv.org/pdf/2602.03680v1

作者:Fumihiko Ishiyama

备注:10 pages, 6 figures.To appear Proc. IEEE CSPA 2026

摘要:“傅里叶分析时频分辨率的理论极限”的起源是它的数值实现,特别是一个世纪前提出的“周期边界条件”的假设。我们之前建议用“线性外插条件(LXC)”代替这个条件,它不需要周期性。这一特性使脉冲序列的瞬时谱分析成为可能,从而取代了短时傅里叶变换(STFT)。我们应用瞬时频谱分析两个肺音异常(爆裂音和喘息)和正常肺音,作为示范。其中,爆裂声包含一个随机脉冲序列。每个脉冲的频谱都是可用的,将各个频谱组合起来就可以得到脉冲序列的频谱图。结果显示了给定脉冲序列的时频结构。摘要:The origin of the "theoretical limit of time-frequency resolution of Fourier analysis" is from its numerical implementation, especially from an assumption of "Periodic Boundary Condition (PBC)," which was introduced a century ago. We previously proposed to replace this condition with "Linear eXtrapolation Condition (LXC)," which does not require periodicity. This feature makes instantaneous spectra analysis of pulse series available, which replaces the short time Fourier transform (STFT). We applied the instantaneous spectra analysis to two lung sounds with abnormalities (crackles and wheezing) and to a normal lung sound, as a demonstration. Among them, crackles contains a random pulse series. The spectrum of each pulse is available, and the spectrogram of pulse series is available with assembling each spectrum. As a result, the time-frequency structure of given pulse series is visualized.


eess.AS音频处理


【1】Zero-Shot TTS With Enhanced Audio Prompts: Bsc Submission For The 2026 Wildspoof Challenge TTS Track
标题:具有增强音频提示的Zero-ShotTTC:2026年Wildspoof挑战TTC曲目的BSC提交
链接:https://arxiv.org/pdf/2602.05770v1

作者:Jose Giraldo,Alex Peiró-Lilja,Rodolfo Zevallos,Cristina España-Bonet

备注:Accepted to ICASSP 2026

摘要:我们评估了两种非自回归架构StyleTTS 2和F5-TTS,以解决野外语音的自发性。我们的模型利用灵活的持续时间建模,以提高韵律自然。为了处理声学噪声,我们使用Sidon模型实现了一个多级增强管道,它在信号质量上明显优于标准Demucs。实验结果表明,微调增强的音频产生卓越的鲁棒性,实现高达4.21 UTMOS和3.47 DNSMOS。此外,我们分析了参考提示的质量和长度对zero-shot合成性能的影响,证明了我们的方法用于真实感语音生成的有效性。摘要:We evaluate two non-autoregressive architectures, StyleTTS2 and F5-TTS, to address the spontaneous nature of in-the-wild speech. Our models utilize flexible duration modeling to improve prosodic naturalness. To handle acoustic noise, we implement a multi-stage enhancement pipeline using the Sidon model, which significantly outperforms standard Demucs in signal quality. Experimental results show that finetuning enhanced audios yields superior robustness, achieving up to 4.21 UTMOS and 3.47 DNSMOS. Furthermore, we analyze the impact of reference prompt quality and length on zero-shot synthesis performance, demonstrating the effectiveness of our approach for realistic speech generation.


【2】Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech Generation from SSL features
标题:Wave-Trainer-Fit:具有可训练先验和定点迭代的神经声码器,从SSL功能生成高质量语音
链接:https://arxiv.org/pdf/2602.05443v1

作者:Hien Ohnaka,Yuma Shirahata,Masaya Kawamura

备注:Accepted by IEEE ICASSP 2026. 5 pages, 3 figures, and 2 tables

摘要:我们提出了WaveTrainerFit,一种神经声码器,它可以从数据驱动的功能(如SSL功能)中生成高质量的波形。WaveTrainerFit建立在WaveFit声码器的基础上,它集成了扩散模型和生成对抗网络。此外,所提出的方法包括以下关键改进:1。通过引入可训练的先验知识,推理过程从接近目标语音的噪声而不是高斯噪声开始。2.通过在匹配语音能量之前对可训练对象施加约束来执行参考感知增益调整。这些改进预计将降低数据驱动特征的波形建模的复杂性,从而能够以更少的推理步骤生成高质量的波形。通过实验,我们证明了WaveTrainerFit可以从数据驱动的特征中生成高度自然的波形,提高了说话人的相似性,同时需要比WaveFit更少的迭代次数。此外,我们表明,所提出的方法相对于SSL特征提取的深度是鲁棒的。代码和预训练模型可从https: github.com line WaveTrainerFit获得。摘要:We propose WaveTrainerFit, a neural vocoder that performs high-quality waveform generation from data-driven features such as SSL features. WaveTrainerFit builds upon the WaveFit vocoder, which integrates diffusion model and generative adversarial network. Furthermore, the proposed method incorporates the following key improvements: 1. By introducing trainable priors, the inference process starts from noise close to the target speech instead of Gaussian noise. 2. Reference-aware gain adjustment is performed by imposing constraints on the trainable prior to matching the speech energy. These improvements are expected to reduce the complexity of waveform modeling from data-driven features, enabling high-quality waveform generation with fewer inference steps. Through experiments, we showed that WaveTrainerFit can generate highly natural waveforms with improved speaker similarity from data-driven features, while requiring fewer iterations than WaveFit. Moreover, we showed that the proposed method works robustly with respect to the depth at which SSL features are extracted. Code and pre-trained models are available from https: github.com line WaveTrainerFit.


【3】Exterior sound field estimation based on physics-constrained kernel
标题:基于物理约束核的外部声场估计
链接:https://arxiv.org/pdf/2602.05236v1

作者:Juliano G. C. Ribeiro,Ryo Matsuda,Jorge Trevino
备注:This paper has been accepted to the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026
摘要:外部声场插值是一个具有挑战性的问题,通常需要特定的阵列配置和源条件的先验知识。我们提出了一种基于高斯过程的插值方法,该方法使用点源再生核,并使用可训练的内积公式来拟合外部声场。虽然这种估计没有封闭的公式,但它允许定义灵活的估计器,该估计器不受麦克风分布的限制,并且利用直接从记录优化的参数自动衰减高次谐波,这意味着可以使用麦克风的任意分布。在模拟实验中,将所提出的核估计器与使用球面波函数和已建立的物理信息机器学习模型的传统方法进行比较,在100 Hz和2.5 kHz的分析频率内平均实现约2 dB的较低插值误差,并在目标区域内更一致地重建地面实况声场。摘要:Exterior sound field interpolation is a challenging problem that often requires specific array configurations and prior knowledge on the source conditions. We propose an interpolation method based on Gaussian processes using a point source reproducing kernel with a trainable inner product formulation made to fit exterior sound fields. While this estimation does not have a closed formula, it allows for the definition of a flexible estimator that is not restricted by microphone distribution and attenuates higher harmonic orders automatically with parameters directly optimized from the recordings, meaning an arbitrary distribution of microphones can be used. The proposed kernel estimator is compared in simulated experiments to the conventional method using spherical wave functions and an established physics-informed machine learning model, achieving lower interpolation error by approximately 2 dB on average within the analyzed frequencies of 100 Hz and 2.5 kHz and reconstructing the ground truth sound field more consistently within the target region.


【4】ARCHI-TTS: A flow-matching-based Text-to-Speech Model with Self-supervised Semantic Aligner and Accelerated Inference
标题:ARCHI-TTC:一种基于流匹配的文本到语音模型,具有自监督语义对齐器和加速推理
链接:https://arxiv.org/pdf/2602.05207v1

作者:Chunyat Wu,Jiajun Deng,Zhengxi Liu,Zheqi Dai,Haolin He,Qiuqiang Kong
备注:Accepted by ICASSP 2026
摘要:虽然基于扩散的非自回归文本到语音(TTS)系统已经展示了令人印象深刻的zero-shot合成能力,但是它们的功效仍然受到两个关键挑战的阻碍:文本语音对齐建模的困难和迭代去噪过程的高计算开销。为了解决这些限制,我们提出了ARCHI-TTS,其具有专用的语义对齐器,以确保文本和音频之间的鲁棒的时间和语义一致性。为了克服高计算推理成本,ARCHI-TTS采用了一种高效的推理策略,该策略在去噪步骤中重用编码器功能,大大加快了合成速度,而不会降低性能。应用于条件编码器的辅助CTC损失进一步增强了语义理解。实验结果表明,ARCHI-TTS在LibriSpeech-PC test-clean上实现了1.98%的WER,在SeedTTS test-en test-zh上实现了1.47% 1.42%的WER,具有较高的推理效率,始终优于最近最先进的TTS系统。摘要:Although diffusion-based, non-autoregressive text-to-speech (TTS) systems have demonstrated impressive zero-shot synthesis capabilities, their efficacy is still hindered by two key challenges: the difficulty of text-speech alignment modeling and the high computational overhead of the iterative denoising process. To address these limitations, we propose ARCHI-TTS that features a dedicated semantic aligner to ensure robust temporal and semantic consistency between text and audio. To overcome high computational inference costs, ARCHI-TTS employs an efficient inference strategy that reuses encoder features across denoising steps, drastically accelerating synthesis without performance degradation. An auxiliary CTC loss applied to the condition encoder further enhances the semantic understanding. Experimental results demonstrate that ARCHI-TTS achieves a WER of 1.98% on LibriSpeech-PC test-clean, and 1.47% 1.42% on SeedTTS test-en test-zh with a high inference efficiency, consistently outperforming recent state-of-the-art TTS systems.


【5】Phase-Only Positioning in Distributed MIMO Under Phase Impairments: AP Selection Using Deep Learning
标题:相损伤下分布式多输出中的纯相定位:使用深度学习的AP选择
链接:https://arxiv.org/pdf/2602.05034v1

作者:Fatih Ayten,Musa Furkan Keskin,Akshay Jain,Mehmet C. Ilter,Ossi Kaltiokallio,Jukka Talvitie,Elena Simona Lohan,Mikko Valkama
摘要:载波相位定位(CPP)可以在下一代无线系统中实现厘米级的精度,而最近的文献表明,在分布式MIMO(D-MIMO)中使用纯相位测量仍然具有很高的精度。然而,相位同步误差对这种系统的影响仍然没有得到充分的探讨。为了解决这一差距,我们首先表明,所提出的双曲线相交方法实现了高度准确的定位,即使在相位同步误差的存在下,当训练适当的数据反映这种损害。然后,我们介绍了一个基于深度学习(DL)的D-MIMO天线点(AP)选择框架,该框架确保了相位同步误差下的高精度定位。仿真结果表明,该框架提高了定位精度相比,现有技术的方法,同时减少推理复杂度约19.7%。摘要:Carrier phase positioning (CPP) can enable cm-level accuracy in next-generation wireless systems, while recent literature shows that accuracy remains high using phase-only measurements in distributed MIMO (D-MIMO). However, the impact of phase synchronization errors on such systems remains insufficiently explored. To address this gap, we first show that the proposed hyperbola intersection method achieves highly accurate positioning even in the presence of phase synchronization errors, when trained on appropriate data reflecting such impairments. We then introduce a deep learning (DL)-based D-MIMO antenna point (AP) selection framework that ensures high-precision localization under phase synchronization errors. Simulation results show that the proposed framework improves positioning accuracy compared to prior-art methods, while reducing inference complexity by approximately 19.7%.


【6】HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake Detection
标题:HyperPotter:诠释音频深度伪造检测中高级交互的魅力
链接:https://arxiv.org/pdf/2602.05670v1

作者:Qing Wen,Haohao Li,Zhongjie Ba,Peng Cheng,Miao He,Li Lu,Kui Ren

备注:20 pages, 8 figures

摘要:AIGC技术的进步使得能够合成能够欺骗人类听觉感知的高度逼真的音频deepfake。虽然已经开发了许多音频深度伪造检测(ADD)方法,但大多数依赖于局部时间 频谱特征或成对关系,忽略了高阶交互(HOI)。HOI捕获从超出其单独贡献的多个特征分量中出现的区分模式。我们提出了HyperPotter,一个基于超图的框架,明确地通过基于聚类的hyperedges与类感知原型初始化这些协同HOI模型。大量的实验表明,HyperPotter在11个数据集上的平均相对增益超过其基线22.15%,在4个具有挑战性的跨域数据集上的性能优于最先进的方法13.96%,表现出对不同攻击和说话者的卓越泛化能力。摘要:Advances in AIGC technologies have enabled the synthesis of highly realistic audio deepfakes capable of deceiving human auditory perception. Although numerous audio deepfake detection (ADD) methods have been developed, most rely on local temporal spectral features or pairwise relations, overlooking high-order interactions (HOIs). HOIs capture discriminative patterns that emerge from multiple feature components beyond their individual contributions. We propose HyperPotter, a hypergraph-based framework that explicitly models these synergistic HOIs through clustering-based hyperedges with class-aware prototype initialization. Extensive experiments demonstrate that HyperPotter surpasses its baseline by an average relative gain of 22.15% across 11 datasets and outperforms state-of-the-art methods by 13.96% on 4 challenging cross-domain datasets, demonstrating superior generalization to diverse attacks and speakers.


机器翻译由腾讯交互翻译提供,仅供参考