今日论文合集:cs.SD语音16篇,eess.AS音频处理10篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks
标题:StylePitcher:为多功能歌唱任务生成遵循风格和表现力的音调曲线
链接:https://arxiv.org/abs/2510.21685

作者:Jingyue Huang, Qihui Yang, Fei Yueh Chen, Julian McAuley, Randal Leistikow, Perry R. Cook, Yongyi Zang
备注:Submitted to ICASSP 2026
摘要:现有的音高曲线生成器面临着两个主要挑战:他们往往忽略了歌手特定的表现力,降低了他们捕捉个人演唱风格的能力。它们通常被开发为用于特定任务的辅助模块,例如音高校正,歌唱声音合成或声音转换,这限制了它们的泛化能力。我们提出了StylePitcher,一个通用的音高曲线生成器,学习歌手的风格,从参考音频,同时保持与预期的旋律。StylePitcher基于整流匹配架构,灵活地将象征性乐谱和音高背景作为生成条件,无需重新训练即可无缝适应各种歌唱任务。各种歌唱任务的客观和主观评估表明,StylePitcher提高了风格相似性和音频质量,同时保持了与特定任务基线相当的音高准确性。
摘要:Existing pitch curve generators face two main challenges: they often neglect singer-specific expressiveness, reducing their ability to capture individual singing styles. And they are typically developed as auxiliary modules for specific tasks such as pitch correction, singing voice synthesis, or voice conversion, which restricts their generalization capability. We propose StylePitcher, a general-purpose pitch curve generator that learns singer style from reference audio while preserving alignment with the intended melody. Built upon a rectified flow matching architecture, StylePitcher flexibly incorporates symbolic music scores and pitch context as conditions for generation, and can seamlessly adapt to diverse singing tasks without retraining. Objective and subjective evaluations across various singing tasks demonstrate that StylePitcher improves style similarity and audio quality while maintaining pitch accuracy comparable to task-specific baselines.


【2】FlowSynth: Instrument Generation Through Distributional Flow Matching and Test-Time Search
标题:FlowSynth:通过分布流量匹配和测试时间搜索生成仪器
链接:https://arxiv.org/abs/2510.21667

作者:Qihui Yang, Randal Leistikow, Yongyi Zang
备注:Submitted to ICASSP 2026
摘要:虚拟乐器生成需要在不同的音高和速度上保持一致的音色,这是现有音符级模型难以解决的挑战。我们提出了FlowSynth,它结合了分布式流量匹配(DFM)与测试时间优化,用于高质量的仪器合成。与学习确定性映射的标准流匹配不同,DFM将速度场参数化为高斯分布,并通过负对数似然进行优化,使模型能够在其预测中表达不确定性。这种概率公式允许有原则的测试时搜索:我们对按模型置信度加权的多个轨迹进行采样,并选择最大化音色一致性的输出。FlowSynth在单音符质量和跨音符一致性方面都优于当前最先进的TokenSynth基线。我们的方法表明,建模预测的不确定性流匹配,结合音乐特定的一致性目标,提供了一个有效的路径,以专业品质的虚拟仪器,适合实时性能。
摘要:Virtual instrument generation requires maintaining consistent timbre across different pitches and velocities, a challenge that existing note-level models struggle to address. We present FlowSynth, which combines distributional flow matching (DFM) with test-time optimization for high-quality instrument synthesis. Unlike standard flow matching that learns deterministic mappings, DFM parameterizes the velocity field as a Gaussian distribution and optimizes via negative log-likelihood, enabling the model to express uncertainty in its predictions. This probabilistic formulation allows principled test-time search: we sample multiple trajectories weighted by model confidence and select outputs that maximize timbre consistency. FlowSynth outperforms the current state-of-the-art TokenSynth baseline in both single-note quality and cross-note consistency. Our approach demonstrates that modeling predictive uncertainty in flow matching, combined with music-specific consistency objectives, provides an effective path to professional-quality virtual instruments suitable for real-time performance.


【3】Smule Renaissance Small: Efficient General-Purpose Vocal Restoration
标题:Smule Renaissance Small:高效的通用声音恢复
链接:https://arxiv.org/abs/2510.21659

作者:Yongyi Zang, Chris Manchester, David Young, Ivan Ivanov, Jeffrey Lufkin, Martin Vladimirov, PJ Solomon, Svetoslav Kepchelev, Fei Yueh Chen, Dongting Cai, Teodor Naydenov, Randal Leistikow
备注:Technical Report
摘要:消费者设备上的语音录音通常遭受多种并发降级:噪声、混响、频带限制和削波。我们提出了Smule文艺复兴小(SRS),一个紧凑的单级模型,直接在复杂的STFT域进行端到端的声乐恢复。通过引入相位感知损耗,SRS可以实现更大的分析窗口,从而提高频率分辨率,同时在48 kHz的iPhone 12 CPU上实现10.5倍的实时推理。在DNS 5 Challenge盲集上,尽管没有语音训练,SRS仍优于强GAN基线,并且与计算昂贵的流匹配系统非常接近。为了在真实的多降级场景下进行评估,我们引入了极端降级台架(EDB):在恶劣的声学条件下捕获的87个唱歌和语音录音。在EDB上,SRS在唱歌方面超过了所有开源基线,并与商业系统相匹配,尽管没有专门的语音培训,但在语音方面仍然具有竞争力。我们在MIT许可证下发布SRS和EDB。
摘要:Vocal recordings on consumer devices commonly suffer from multiple concurrent degradations: noise, reverberation, band-limiting, and clipping. We present Smule Renaissance Small (SRS), a compact single-stage model that performs end-to-end vocal restoration directly in the complex STFT domain. By incorporating phase-aware losses, SRS enables large analysis windows for improved frequency resolution while achieving 10.5x real-time inference on iPhone 12 CPU at 48 kHz. On the DNS 5 Challenge blind set, despite no speech training, SRS outperforms a strong GAN baseline and closely matches a computationally expensive flow-matching system. To enable evaluation under realistic multi-degradation scenarios, we introduce the Extreme Degradation Bench (EDB): 87 singing and speech recordings captured under severe acoustic conditions. On EDB, SRS surpasses all open-source baselines on singing and matches commercial systems, while remaining competitive on speech despite no speech-specific training. We release both SRS and EDB under the MIT License.


【4】Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video
标题:Foley Control:将冻结的潜在文本到音频模型与视频对齐
链接:https://arxiv.org/abs/2510.21581

作者:Ciara Rowles, Varun Jampani, Simon Donné, Shimon Vainer, Julian Parker, Zach Evans
备注:Project Page: this https URL
摘要:Foley Control是视频引导Foley的一种轻量级方法,它可以保持预训练的单模态模型冻结,并且只学习它们之间的一个小的交叉注意力桥梁。我们将V-JEPA 2视频嵌入连接到冻结的Stable Audio Open DiT文本到音频(T2 A)模型,方法是在模型现有的文本交叉关注之后插入紧凑的视频交叉关注,因此提示设置全局语义,而视频细化时序和局部动态。冻结的主干保留强边缘(视频;给定文本的音频),桥学习同步所需的音频-视频依赖性-而无需重新训练音频先验。为了减少记忆和稳定训练,我们在条件反射之前将视频令牌集中起来。在策划的视频-音频基准测试中,Foley Control提供了具有竞争力的时间和语义对齐,其可训练参数远远少于最近的多模态系统,同时保留了可编程驱动的可控性和生产友好的模块化(交换/升级编码器或T2 A骨干网,而无需端到端再训练)。尽管我们专注于视频到Foley,但相同的桥设计可能会扩展到其他音频模式(例如,演讲)。
摘要:Foley Control is a lightweight approach to video-guided Foley that keeps pretrained single-modality models frozen and learns only a small cross-attention bridge between them. We connect V-JEPA2 video embeddings to a frozen Stable Audio Open DiT text-to-audio (T2A) model by inserting compact video cross-attention after the model's existing text cross-attention, so prompts set global semantics while video refines timing and local dynamics. The frozen backbones retain strong marginals (video; audio given text) and the bridge learns the audio-video dependency needed for synchronization -- without retraining the audio prior. To cut memory and stabilize training, we pool video tokens before conditioning. On curated video-audio benchmarks, Foley Control delivers competitive temporal and semantic alignment with far fewer trainable parameters than recent multi-modal systems, while preserving prompt-driven controllability and production-friendly modularity (swap/upgrade encoders or the T2A backbone without end-to-end retraining). Although we focus on Video-to-Foley, the same bridge design can potentially extend to other audio modalities (e.g., speech).


【5】FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
标题:FlexIO:灵活的单通道和多通道语音分离和增强
链接:https://arxiv.org/abs/2510.21485

作者:Yoshiki Masuyama, Kohei Saijo, Francesco Paissan, Jiangyu Han, Marc Delcroix, Ryo Aihara, François G. Germain, Gordon Wichern, Jonathan Le Roux
备注:Submitted to ICASSP 2026
摘要:语音分离与增强(SSE)技术已经取得了显著的进步,并在控制环境中取得了可喜的成果,如固定数量的扬声器和固定的阵列配置。对于通用SSE系统,单声道系统已经被扩展以处理可变数量的扬声器(即,产出)。同时,适应各种阵列配置的多通道系统(即,投入)已经开发。然而,这些尝试是分开进行的。在本文中,我们提出了一个灵活的输入和输出SSE系统,名为FlexIO。它使用提示向量执行条件分离,每个扬声器作为条件,允许分离任意数量的扬声器。通过阵列不可知信道通信机制,多信道混合物与提示向量一起被处理。我们的实验表明,FlexIO成功地覆盖了一到五个麦克风和一到三个扬声器的各种条件。我们还证实了FlexIO在CHiME-4真实数据上的鲁棒性。
摘要:Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel systems have been extended to deal with a variable number of speakers (i.e., outputs). Meanwhile, multi-channel systems accommodating various array configurations (i.e., inputs) have been developed. However, these attempts have been pursued separately. In this paper, we propose a flexible input and output SSE system, named FlexIO. It performs conditional separation using prompt vectors, one per speaker as a condition, allowing separation of an arbitrary number of speakers. Multi-channel mixtures are processed together with the prompt vectors via an array-agnostic channel communication mechanism. Our experiments demonstrate that FlexIO successfully covers diverse conditions with one to five microphones and one to three speakers. We also confirm the robustness of FlexIO on CHiME-4 real data.


【6】HiFi-HARP: A High-Fidelity 7th-Order Ambisonic Room Impulse Response Dataset
标题:HiFi-HARP:高保真七阶立体声房间脉冲响应数据集
链接:https://arxiv.org/abs/2510.21257

作者:Shivam Saini, Jürgen Peissig
备注:Under review for ICASSP 2026
摘要:我们介绍了HiFi-HARP,这是一个7阶高阶Ambisonic房间脉冲响应(HOA-RIR)的大规模数据集,由通过混合声学模拟在真实室内场景中生成的超过100,000个RIR组成。HiFi-HARP将3D-FRONT存储库中几何形状复杂的家具房间模型与混合仿真管道相结合:使用高达900 Hz的低频基于波的仿真(时域有限差分),而900 Hz以上的高频则使用光线跟踪方法进行仿真。将组合的原始RIR编码到球谐域(Ambix ACN)中以用于直接可听化。我们的数据集通过提供7阶Ambisonic RIR扩展了先前的工作,该RIR将波理论精度与现实的房间内容相结合。我们详细介绍了生成管道(场景和材料选择,阵列设计,混合仿真,立体声编码),并提供数据集统计(房间体积,RT 60分布,吸收特性)。对比表突出了HiFi-HARP相对于现有RIR集合的新颖性。最后,我们概述了潜在的基准,如FOA到HOA上采样,源定位和去混响。我们讨论机器学习用例(空间音频渲染,声学参数估计)和限制(例如,模拟近似、静态场景)。总体而言,HiFi-HARP为在复杂环境中开发空间音频和声学算法提供了丰富的资源。
摘要:We introduce HiFi-HARP, a large-scale dataset of 7th-order Higher-Order Ambisonic Room Impulse Responses (HOA-RIRs) consisting of more than 100,000 RIRs generated via a hybrid acoustic simulation in realistic indoor scenes. HiFi-HARP combines geometrically complex, furnished room models from the 3D-FRONT repository with a hybrid simulation pipeline: low-frequency wave-based simulation (finite-difference time-domain) up to 900 Hz is used, while high frequencies above 900 Hz are simulated using a ray-tracing approach. The combined raw RIRs are encoded into the spherical-harmonic domain (AmbiX ACN) for direct auralization. Our dataset extends prior work by providing 7th-order Ambisonic RIRs that combine wave-theoretic accuracy with realistic room content. We detail the generation pipeline (scene and material selection, array design, hybrid simulation, ambisonic encoding) and provide dataset statistics (room volumes, RT60 distributions, absorption properties). A comparison table highlights the novelty of HiFi-HARP relative to existing RIR collections. Finally, we outline potential benchmarks such as FOA-to-HOA upsampling, source localization, and dereverberation. We discuss machine learning use cases (spatial audio rendering, acoustic parameter estimation) and limitations (e.g., simulation approximations, static scenes). Overall, HiFi-HARP offers a rich resource for developing spatial audio and acoustics algorithms in complex environments.


【7】Robust Distortion-Free Watermark for Autoregressive Audio Generation Models
标题:自回归音频生成模型的鲁棒无失真水印
链接:https://arxiv.org/abs/2510.21115

作者:Yihan Wu, Georgios Milis, Ruibo Chen, Heng Huang
摘要:下一个令牌预测模型的快速发展导致了各种模式的广泛采用,从而能够创建逼真的合成媒体。在音频领域,虽然自回归语音模型推动了会话交互的发展,但滥用的可能性也在增加,例如网络钓鱼计划中的模仿或制作误导性语音录音。因此,诸如水印之类的安全措施对于确保数字媒体的真实性变得至关重要。传统的用于自回归语言模型的统计水印方法在应用于自回归音频模型时面临挑战,这是由于不可避免的"重新标记化失配“-原始和重新标记化的离散音频标记序列之间的差异。为了解决这个问题,我们引入了Aligned-IS,这是一种新颖的无失真水印,专为音频生成模型而设计。该技术利用了一种聚类方法,该方法等同地处理同一集群内的令牌,有效地解决了重新令牌化不匹配问题。我们在流行的音频生成平台上进行的全面测试表明,与最先进的无失真水印适配相比,Aligned-IS不仅保留了生成音频的质量,而且显著提高了水印的可检测性,为安全音频技术应用建立了新的基准。
摘要:The rapid advancement of next-token-prediction models has led to widespread adoption across modalities, enabling the creation of realistic synthetic media. In the audio domain, while autoregressive speech models have propelled conversational interactions forward, the potential for misuse, such as impersonation in phishing schemes or crafting misleading speech recordings, has also increased. Security measures such as watermarking have thus become essential to ensuring the authenticity of digital media. Traditional statistical watermarking methods used for autoregressive language models face challenges when applied to autoregressive audio models, due to the inevitable ``retokenization mismatch'' - the discrepancy between original and retokenized discrete audio token sequences. To address this, we introduce Aligned-IS, a novel, distortion-free watermark, specifically crafted for audio generation models. This technique utilizes a clustering approach that treats tokens within the same cluster equivalently, effectively countering the retokenization mismatch issue. Our comprehensive testing on prevalent audio generation platforms demonstrates that Aligned-IS not only preserves the quality of generated audio but also significantly improves the watermark detectability compared to the state-of-the-art distortion-free watermarking adaptations, establishing a new benchmark in secure audio technology applications.


【8】Can Current Detectors Catch Face-to-Voice Deepfake Attacks?
标题:当前的检测器可以捕捉面对面语音Deepfake攻击吗?
链接:https://arxiv.org/abs/2510.21004

作者:Nguyen Linh Bao Nguyen, Alsharif Abuadbba, Kristen Moore, Tingming Wu
备注:8 pages, Accepted at Workshop on AI for Cyber Threat Intelligence, co-located with ACSAC 2025
摘要:生成模型的快速发展使人们能够创建越来越隐秘的合成语音,通常被称为音频deepfake。最近的一项技术,FOICE [USENIX'24],展示了一种特别令人担忧的能力:从单个面部图像生成受害者的声音,而不需要任何声音样本。通过利用面部和声音特征之间的相关性,FOICE产生的合成声音足够逼真,可以绕过行业标准的身份验证系统,包括微信声纹和Microsoft Azure。这引起了严重的安全问题,因为面部图像比语音样本更容易被对手获得,大大降低了大规模攻击的门槛。在这项工作中,我们研究了两个核心研究问题:(RQ 1)最先进的音频deepfake检测器能否在干净和嘈杂的条件下可靠地检测FOICE生成的语音,以及(RQ 2)在FOICE数据上微调这些检测器是否可以在不过度拟合的情况下改善检测,从而保持对SpeechT5等看不见的语音生成器的鲁棒性。   我们的研究有三个贡献。首先,我们提出了第一个系统的评估FOICE检测,显示领先的检测器在标准和嘈杂的条件下始终失败。其次,我们引入了有针对性的微调策略,捕捉FOICE特定的工件,产生显着的准确性提高。第三,我们在微调后评估泛化,揭示了FOICE专业化和不可见合成管道的鲁棒性之间的权衡。这些发现暴露了当今防御的根本弱点,并激发了下一代音频deepfake检测的新架构和训练协议。
摘要:The rapid advancement of generative models has enabled the creation of increasingly stealthy synthetic voices, commonly referred to as audio deepfakes. A recent technique, FOICE [USENIX'24], demonstrates a particularly alarming capability: generating a victim's voice from a single facial image, without requiring any voice sample. By exploiting correlations between facial and vocal features, FOICE produces synthetic voices realistic enough to bypass industry-standard authentication systems, including WeChat Voiceprint and Microsoft Azure. This raises serious security concerns, as facial images are far easier for adversaries to obtain than voice samples, dramatically lowering the barrier to large-scale attacks. In this work, we investigate two core research questions: (RQ1) can state-of-the-art audio deepfake detectors reliably detect FOICE-generated speech under clean and noisy conditions, and (RQ2) whether fine-tuning these detectors on FOICE data improves detection without overfitting, thereby preserving robustness to unseen voice generators such as SpeechT5.   Our study makes three contributions. First, we present the first systematic evaluation of FOICE detection, showing that leading detectors consistently fail under both standard and noisy conditions. Second, we introduce targeted fine-tuning strategies that capture FOICE-specific artifacts, yielding significant accuracy improvements. Third, we assess generalization after fine-tuning, revealing trade-offs between specialization to FOICE and robustness to unseen synthesis pipelines. These findings expose fundamental weaknesses in today's defenses and motivate new architectures and training protocols for next-generation audio deepfake detection.


【9】Compressing Quaternion Convolutional Neural Networks for Audio Classification
标题:压缩四元数卷积神经网络音频分类
链接:https://arxiv.org/abs/2510.21388

作者:Arshdeep Singh, Vinayak Abrol, Mark D. Plumbley
备注:Under review in IEEE TASLPRO
摘要:传统的实域卷积神经网络(CNN)已被广泛用于音频分类。然而,它们的卷积运算独立地处理多通道输入,限制了捕获通道之间相关性的能力。这可能导致次优的特征学习,特别是对于复杂的音频模式,如多声道声谱图表示。四元数卷积神经网络(QCNN)通过采用四元数代数来联合捕获声道间依赖性来解决这一限制,从而实现具有更少可学习参数的更紧凑模型,同时更好地利用音频信号的多维特性。然而,由于四元数运算的开销,QCNN表现出更高的计算复杂度,与传统CNN相比,导致推理延迟增加和效率降低,这对资源受限平台上的部署提出了挑战。为了应对这一挑战,本研究探索了知识蒸馏(KD)和修剪,以降低QCNN的计算复杂度,同时保持性能。我们的音频分类实验表明,与KD相比,修剪QCNN可以实现类似或更优的性能,同时需要更少的计算量。与传统的CNN和基于Transformer的架构相比,修剪的QCNN实现了具有竞争力的性能,减少了可学习的参数数量和计算复杂度。在AudioSet数据集上,修剪后的QCNN将计算成本降低了50%,参数计数减少了80%,同时保持了与传统CNN相当的性能。此外,修剪后的QCNN在多个音频分类基准中具有良好的泛化能力,包括用于音乐流派识别的GTZAN,用于环境声音分类的ESC-50和用于语音情感识别的RAVDESS。
摘要:Conventional Convolutional Neural Networks (CNNs) in the real domain have been widely used for audio classification. However, their convolution operations process multi-channel inputs independently, limiting the ability to capture correlations among channels. This can lead to suboptimal feature learning, particularly for complex audio patterns such as multi-channel spectrogram representations. Quaternion Convolutional Neural Networks (QCNNs) address this limitation by employing quaternion algebra to jointly capture inter-channel dependencies, enabling more compact models with fewer learnable parameters while better exploiting the multi-dimensional nature of audio signals. However, QCNNs exhibit higher computational complexity due to the overhead of quaternion operations, resulting in increased inference latency and reduced efficiency compared to conventional CNNs, posing challenges for deployment on resource-constrained platforms. To address this challenge, this study explores knowledge distillation (KD) and pruning, to reduce the computational complexity of QCNNs while maintaining performance. Our experiments on audio classification reveal that pruning QCNNs achieves similar or superior performance compared to KD while requiring less computational effort. Compared to conventional CNNs and Transformer-based architectures, pruned QCNNs achieve competitive performance with a reduced learnable parameter count and computational complexity. On the AudioSet dataset, pruned QCNNs reduce computational cost by 50\% and parameter count by 80\%, while maintaining performance comparable to the conventional CNNs. Furthermore, pruned QCNNs generalize well across multiple audio classification benchmarks, including GTZAN for music genre recognition, ESC-50 for environmental sound classification and RAVDESS for speech emotion recognition.


【10】Are These Even Words? Quantifying the Gibberishness of Generative Speech Models
标题:这些词是偶数吗?量化生成语音模型的胡言乱语
链接:https://arxiv.org/abs/2510.21317

作者:Danilo de Oliveira, Tal Peer, Jonas Rochdi, Timo Gerkmann
摘要:目前正在致力于非侵入性质量和可懂度评估的重要研究工作,特别是考虑到它如何能够管理野外语音数据的大规模数据集。然而,随着生成模型合成高质量语音的能力不断提高,新类型的伪影变得相关,例如生成幻觉。虽然侵入性度量能够从参考信号中发现这种差异,但目前尚不清楚当前的非侵入性方法如何对高质量的音素混淆或更极端的胡言乱语做出反应。在本文中,我们将探讨如何利用语言模型在完全无监督的环境下考虑这方面的因素。此外,我们还发布了一个高质量的合成胡言乱语语音数据集,以进一步开发评估口语中不可信句子的措施,以及用于计算各种语音语言模型得分的代码。
摘要:Significant research efforts are currently being dedicated to non-intrusive quality and intelligibility assessment, especially given how it enables curation of large scale datasets of in-the-wild speech data. However, with the increasing capabilities of generative models to synthesize high quality speech, new types of artifacts become relevant, such as generative hallucinations. While intrusive metrics are able to spot such sort of discrepancies from a reference signal, it is not clear how current non-intrusive methods react to high-quality phoneme confusions or, more extremely, gibberish speech. In this paper we explore how to factor in this aspect under a fully unsupervised setting by leveraging language models. Additionally, we publish a dataset of high-quality synthesized gibberish speech for further development of measures to assess implausible sentences in spoken language, alongside code for calculating scores from a variety of speech language models.


【11】WhaleVAD-BPN: Improving Baleen Whale Call Detection with Boundary Proposal Networks and Post-processing Optimisation
标题:WhaleVAD-BPN:利用边界提议网络和后处理优化改进须鲸叫声检测
链接:https://arxiv.org/abs/2510.21280

作者:Christiaan M. Geldenhuys, Günther Tonitz, Thomas R. Niesler
摘要:虽然最近的声音事件检测(SED)系统可以识别海洋音频中的须鲸叫声,但与误报和少数群体检测相关的挑战仍然存在。我们提出了边界建议网络(BPN),它扩展了现有的轻量级SED系统。BPN的灵感来自图像对象检测的工作,旨在减少误报检测的数量。它通过使用在主干分类模型内计算的中间潜在表示来门控最终输出来实现这一点。当添加到现有的SED系统中时,BPN实现了16.8%的精确度绝对提高,以及少数类d-调用和bp-调用的F1分数分别提高了21.3%和9.4%。我们进一步考虑两种方法来选择后处理超参数:向前搜索和向后搜索。通过分别优化事件级和帧级超参数,这两种方法导致使用经验方法选择的参数的相当大的性能改进。完整的WhaleVAD-BPN系统实现了0.475的交叉验证开发F1评分,比基线绝对改善了9.8%。
摘要:While recent sound event detection (SED) systems can identify baleen whale calls in marine audio, challenges related to false positive and minority-class detection persist. We propose the boundary proposal network (BPN), which extends an existing lightweight SED system. The BPN is inspired by work in image object detection and aims to reduce the number of false positive detections. It achieves this by using intermediate latent representations computed within the backbone classification model to gate the final output. When added to an existing SED system, the BPN achieves a 16.8 % absolute increase in precision, as well as 21.3 % and 9.4 % improvements in the F1-score for minority-class d-calls and bp-calls, respectively. We further consider two approaches to the selection of post-processing hyperparameters: a forward-search and a backward-search. By separately optimising event-level and frame-level hyperparameters, these two approaches lead to considerable performance improvements over parameters selected using empirical methods. The complete WhaleVAD-BPN system achieves a cross-validated development F1-score of 0.475, which is a 9.8 % absolute improvement over the baseline.


【12】SpecTokenizer: A Lightweight Streaming Codec in the Compressed Spectrum Domain
标题:SpecTokenizer:压缩频谱域中的轻量级流编解码器
链接:https://arxiv.org/abs/2510.21209

作者:Zixiang Wan, Guochang Zhang, Yifeng He, Jianqiang Wei
备注:Accepted by Interspeech 2025; 5 pages, 1 figure, 5 tables
摘要:近年来,神经音频编解码器(NAC)作为语音语言模型中的音频压缩和音频表示技术得到了越来越多的关注。虽然主流的NAC通常需要G级计算和M级参数,但轻量级和流式NAC的性能仍有待研究。本文提出了SpecTokenizer,一个轻量级的流媒体编解码器,在压缩频谱域。SpecTokenizer仅由交替的CNN和RNN层组成,通过压缩频谱域中的多尺度建模实现了更高的效率和更好的表示能力。在4 kbps的速度下,与具有最先进的轻量级架构的编解码器相比,所提出的SpecTokenizer实现了相当或更高的性能,同时只需要20%的计算和10%的参数。此外,当使用类似的计算和存储资源时,它的性能明显优于编解码器。
摘要:Neural Audio Codecs (NACs) have gained growing attention in recent years as technologies for audio compression and audio representation in speech language models. While mainstream NACs typically require G-level computation and M-level parameters, the performance of lightweight and streaming NACs remains underexplored. This paper proposes SpecTokenizer, a lightweight streaming codec that operates in the compressed spectral domain. Composed solely of alternating CNN and RNN layers, SpecTokenizer achieves greater efficiency and better representational capability through multi-scale modeling in the compressed spectrum domain. At 4 kbps, the proposed SpecTokenizer achieves comparable or superior performance compared to the codec with state-of-the-art lightweight architecture while requiring only 20% of the computation and 10% of the parameters. Furthermore, it significantly outperforms the codec when using similar computational and storage resources.


【13】PhoenixCodec: Taming Neural Speech Coding for Extreme Low-Resource Scenarios
标题:PhoenixCodec:针对极低资源场景驯服神经语音编码
链接:https://arxiv.org/abs/2510.21196

作者:Zixiang Wan, Haoran Zhao, Guochang Zhang, Runqiang Han, Jianqiang Wei, Yuexian Zou
备注:5 pages, 1 figure, 4 tables
摘要:本文介绍了PhoenixCodec,一个全面的神经语音编码和解码框架,专为极低的资源条件。该系统集成了优化的非对称频率-时间架构,循环校准和细化(CCR)的训练策略,和噪声不变的微调过程。在严格的限制下-计算低于700 MFLOPs,延迟小于30 ms,以及支持1 kbps和6 kbps的双速率-现有方法面临效率和质量之间的权衡。PhoenixCodec通过减轻传统解码器的资源分散、采用CCR来避免局部最优以及通过噪声样本微调来增强鲁棒性来解决这些挑战。在LRAC 2025 Challenge Track 1中,该系统总体排名第三,并在清洁测试中在1 kbps的真实世界噪声和混响以及清晰度方面表现出最佳性能,证实了其有效性。
摘要:This paper presents PhoenixCodec, a comprehensive neural speech coding and decoding framework designed for extremely low-resource conditions. The proposed system integrates an optimized asymmetric frequency-time architecture, a Cyclical Calibration and Refinement (CCR) training strategy, and a noise-invariant fine-tuning procedure. Under stringent constraints - computation below 700 MFLOPs, latency less than 30 ms, and dual-rate support at 1 kbps and 6 kbps - existing methods face a trade-off between efficiency and quality. PhoenixCodec addresses these challenges by alleviating the resource scattering of conventional decoders, employing CCR to escape local optima, and enhancing robustness through noisy-sample fine-tuning. In the LRAC 2025 Challenge Track 1, the proposed system ranked third overall and demonstrated the best performance at 1 kbps in both real-world noise and reverberation and intelligibility in clean tests, confirming its effectiveness.


【14】refess-qi: reference-free evaluation for speech separation with joint quality and intelligibility scoring
标题:refess-qi:具有联合质量和可理解性评分的语音分离无参考评估
链接:https://arxiv.org/abs/2510.21014

作者:Ari Frummer, Helin Wang, Tianyu Cao, Adi Arbel, Yuval Sieradzki, Oren Gal, Jesús Villalba, Thomas Thebaud, Najim Dehak
摘要:源分离是各种语音处理任务(如自动语音识别(ASR))的关键预处理步骤。传统上,语音分离的评价指标依赖于匹配的参考音频和相应的瞬态来评估音频质量和可懂度。然而,它们不能用于评估没有参考存在的真实世界的混合物。本文介绍了一种基于自监督学习(SSL)表示的无文本无参考评估框架。建议的框架利用混合和分离的轨道来预测联合音频质量,通过标度不变的信噪比(SI-SNR)度量,和语音清晰度通过字错误率(WER)度量。我们在WHAMR上进行了实验!数据集,其示出了WER估计,平均绝对误差(MAE)为17%,皮尔逊相关系数(PCC)为0.77;以及SI-SNR估计,MAE为1.38,PCC为0.95。我们进一步证明了我们的估计的鲁棒性,通过使用各种SSL表示。
摘要:Source separation is a crucial pre-processing step for various speech processing tasks, such as automatic speech recognition (ASR). Traditionally, the evaluation metrics for speech separation rely on the matched reference audios and corresponding transcriptions to assess audio quality and intelligibility. However, they cannot be used to evaluate real-world mixtures for which no reference exists. This paper introduces a text-free reference-free evaluation framework based on self-supervised learning (SSL) representations. The proposed framework utilize the mixture and separated tracks to predict jointly audio quality, through the Scale Invariant Signal to Noise Ratio (SI-SNR) metric, and speech intelligibility through the Word Error Rate (WER) metric. We conducted experiments on the WHAMR! dataset, which shows a WER estimation with a mean absolute error (MAE) of 17\% and a Pearson correlation coefficient (PCC) of 0.77; and SI-SNR estimation with an MAE of 1.38 and PCC of 0.95. We further demonstrate the robustness of our estimator by using various SSL representations.


【15】Beyond Hearing: Learning Task-agnostic ExG Representations from Earphones via Physiology-informed Tokenization
标题:超越听力:通过基于生理学的代币化从耳机学习任务不可知的ExG表示
链接:https://arxiv.org/abs/2510.20853

作者:Hyungjun Yoon, Seungjoo Lee, Yu Yvonne Wu, Xiaomeng Chen, Taiting Lu, Freddy Yifei Liu, Taeckyung Lee, Hyeongheon Cha, Haochen Zhao, Gaoteng Zhao, Sung-Ju Lee, Cecilia Mascolo, Dongyao Chen, Lili Qiu
备注:19 pages, 9 figures
摘要:电生理学(ExG)信号提供了对人类生理学的有价值的见解,但是由于两个关键限制,构建跨日常任务推广的基础模型仍然具有挑战性:(i)数据多样性不足,因为大多数ExG记录是在具有庞大昂贵设备的受控实验室中收集的;以及(ii)需要定制处理的任务特定模型设计(即,目标频率滤波器)和架构,这限制了跨任务的泛化。为了应对这些挑战,我们引入了一种可扩展的、与任务无关的ExG监控方法。我们使用基于耳机的硬件原型收集了50小时的不显眼的自由生活ExG数据,以缩小数据多样性差距。我们的方法的核心是生理信息多波段令牌化(PiMT),它将ExG信号分解为12个生理信息令牌,然后进行重建任务以学习鲁棒的表示。这使得能够在捕获任务相关信息的同时在整个频谱上进行自适应特征识别。在我们新的DailySense智能卡上进行的实验-第一个在五种人类感官中实现基于ExG的分析-以及四个公共ExG基准测试,证明PiMT在各种任务中始终优于最先进的方法。
摘要:Electrophysiological (ExG) signals offer valuable insights into human physiology, yet building foundation models that generalize across everyday tasks remains challenging due to two key limitations: (i) insufficient data diversity, as most ExG recordings are collected in controlled labs with bulky, expensive devices; and (ii) task-specific model designs that require tailored processing (i.e., targeted frequency filters) and architectures, which limit generalization across tasks. To address these challenges, we introduce an approach for scalable, task-agnostic ExG monitoring in the wild. We collected 50 hours of unobtrusive free-living ExG data with an earphone-based hardware prototype to narrow the data diversity gap. At the core of our approach is Physiology-informed Multi-band Tokenization (PiMT), which decomposes ExG signals into 12 physiology-informed tokens, followed by a reconstruction task to learn robust representations. This enables adaptive feature recognition across the full frequency spectrum while capturing task-relevant information. Experiments on our new DailySense dataset-the first to enable ExG-based analysis across five human senses-together with four public ExG benchmarks, demonstrate that PiMT consistently outperforms state-of-the-art methods across diverse tasks.


【16】Can large audio language models understand child stuttering speech? speech summarization, and source separation
标题:大型音频语言模型能理解儿童口吃的言语吗?语音摘要和源分离
链接:https://arxiv.org/abs/2510.20850

作者:Chibuzor Okocha, Maya Bakri, Christan Grant
备注:7 pages, 1 Figure, 8 tables, Under review ICASSP 2026
摘要:儿童语音在声学、韵律和语言发展方面与成人语音不同,不流利(重复、重复、块)进一步挑战了自动语音识别(ASR)和下游自然语言处理(NLP)。最近的大型音频语言模型(LALM)表现出强大的跨模态音频理解,然而,他们的行为在不流利的儿童语音仍然未充分探索。我们在两种情况下评估了几种最先进的LALM:采访(混合扬声器)和阅读任务(独生子女)。这些任务是(i)单通道源分离,以隔离儿童和(ii)仅儿童总结,保留临床相关的不流利,并避免成人语音泄漏。   评估结合了大型语言模型(LLM)作为判断,人类专家评级和BERTScore(F1),我们报告了模型之间以及模型与人类之间的一致性,以评估可靠性。我们的研究结果描述了LALM从混合音频中产生忠实的儿童摘要的条件以及它们失败的地方,为临床和教育部署提供了实用的指导。我们提供提示和评估脚本来支持复制。
摘要:Child speech differs from adult speech in acoustics, prosody, and language development, and disfluencies (repetitions, prolongations, blocks) further challenge Automatic Speech Recognition (ASR) and downstream Natural Language Processing (NLP). Recent large audio-language models (LALMs) demonstrate strong cross-modal audio understanding; however, their behavior in disfluent child speech remains underexplored. We evaluate several state-of-the-art LALMs in two settings: an interview (mixed speakers) and a reading task (single child). The tasks are (i) single-channel source separation to isolate the child and (ii) child-only summarization that preserves clinically relevant disfluencies and avoids adult-speech leakage.   Evaluation combines Large Language Model (LLM) as a judge, human expert ratings, and BERTScore (F1), and we report agreement between models and between models and humans to assess reliability. Our findings delineate the conditions under which LALMs produce faithful child-only summaries from mixed audio and where they fail, offering practical guidance for clinical and educational deployments. We provide prompts and evaluation scripts to support replication.


eess.AS音频处理


【1】Compressing Quaternion Convolutional Neural Networks for Audio Classification
标题:压缩四元数卷积神经网络音频分类
链接:https://arxiv.org/abs/2510.21388

作者:Arshdeep Singh, Vinayak Abrol, Mark D. Plumbley
备注:Under review in IEEE TASLPRO
摘要:传统的实域卷积神经网络(CNN)已被广泛用于音频分类。然而,它们的卷积运算独立地处理多通道输入,限制了捕获通道之间相关性的能力。这可能导致次优的特征学习,特别是对于复杂的音频模式,如多声道声谱图表示。四元数卷积神经网络(QCNN)通过采用四元数代数来联合捕获声道间依赖性来解决这一限制,从而实现具有更少可学习参数的更紧凑模型,同时更好地利用音频信号的多维特性。然而,由于四元数运算的开销,QCNN表现出更高的计算复杂度,与传统CNN相比,导致推理延迟增加和效率降低,这对资源受限平台上的部署提出了挑战。为了应对这一挑战,本研究探索了知识蒸馏(KD)和修剪,以降低QCNN的计算复杂度,同时保持性能。我们的音频分类实验表明,与KD相比,修剪QCNN可以实现类似或更优的性能,同时需要更少的计算量。与传统的CNN和基于Transformer的架构相比,修剪的QCNN实现了具有竞争力的性能,减少了可学习的参数数量和计算复杂度。在AudioSet数据集上,修剪后的QCNN将计算成本降低了50%,参数计数减少了80%,同时保持了与传统CNN相当的性能。此外,修剪后的QCNN在多个音频分类基准中具有良好的泛化能力,包括用于音乐流派识别的GTZAN,用于环境声音分类的ESC-50和用于语音情感识别的RAVDESS。
摘要:Conventional Convolutional Neural Networks (CNNs) in the real domain have been widely used for audio classification. However, their convolution operations process multi-channel inputs independently, limiting the ability to capture correlations among channels. This can lead to suboptimal feature learning, particularly for complex audio patterns such as multi-channel spectrogram representations. Quaternion Convolutional Neural Networks (QCNNs) address this limitation by employing quaternion algebra to jointly capture inter-channel dependencies, enabling more compact models with fewer learnable parameters while better exploiting the multi-dimensional nature of audio signals. However, QCNNs exhibit higher computational complexity due to the overhead of quaternion operations, resulting in increased inference latency and reduced efficiency compared to conventional CNNs, posing challenges for deployment on resource-constrained platforms. To address this challenge, this study explores knowledge distillation (KD) and pruning, to reduce the computational complexity of QCNNs while maintaining performance. Our experiments on audio classification reveal that pruning QCNNs achieves similar or superior performance compared to KD while requiring less computational effort. Compared to conventional CNNs and Transformer-based architectures, pruned QCNNs achieve competitive performance with a reduced learnable parameter count and computational complexity. On the AudioSet dataset, pruned QCNNs reduce computational cost by 50\% and parameter count by 80\%, while maintaining performance comparable to the conventional CNNs. Furthermore, pruned QCNNs generalize well across multiple audio classification benchmarks, including GTZAN for music genre recognition, ESC-50 for environmental sound classification and RAVDESS for speech emotion recognition.


【2】Are These Even Words? Quantifying the Gibberishness of Generative Speech Models
标题:这些词是偶数吗?量化生成语音模型的胡言乱语
链接:https://arxiv.org/abs/2510.21317

作者:Danilo de Oliveira, Tal Peer, Jonas Rochdi, Timo Gerkmann
摘要:目前正在致力于非侵入性质量和可懂度评估的重要研究工作,特别是考虑到它如何能够管理野外语音数据的大规模数据集。然而,随着生成模型合成高质量语音的能力不断提高,新类型的伪影变得相关,例如生成幻觉。虽然侵入性度量能够从参考信号中发现这种差异,但目前尚不清楚当前的非侵入性方法如何对高质量的音素混淆或更极端的胡言乱语做出反应。在本文中,我们将探讨如何利用语言模型在完全无监督的环境下考虑这方面的因素。此外,我们还发布了一个高质量的合成胡言乱语语音数据集,以进一步开发评估口语中不可信句子的措施,以及用于计算各种语音语言模型得分的代码。
摘要:Significant research efforts are currently being dedicated to non-intrusive quality and intelligibility assessment, especially given how it enables curation of large scale datasets of in-the-wild speech data. However, with the increasing capabilities of generative models to synthesize high quality speech, new types of artifacts become relevant, such as generative hallucinations. While intrusive metrics are able to spot such sort of discrepancies from a reference signal, it is not clear how current non-intrusive methods react to high-quality phoneme confusions or, more extremely, gibberish speech. In this paper we explore how to factor in this aspect under a fully unsupervised setting by leveraging language models. Additionally, we publish a dataset of high-quality synthesized gibberish speech for further development of measures to assess implausible sentences in spoken language, alongside code for calculating scores from a variety of speech language models.


【3】WhaleVAD-BPN: Improving Baleen Whale Call Detection with Boundary Proposal Networks and Post-processing Optimisation
标题:WhaleVAD-BPN:利用边界提议网络和后处理优化改进须鲸叫声检测
链接:https://arxiv.org/abs/2510.21280

作者:Christiaan M. Geldenhuys, Günther Tonitz, Thomas R. Niesler
摘要:虽然最近的声音事件检测(SED)系统可以识别海洋音频中的须鲸叫声,但与误报和少数群体检测相关的挑战仍然存在。我们提出了边界建议网络(BPN),它扩展了现有的轻量级SED系统。BPN的灵感来自图像对象检测的工作,旨在减少误报检测的数量。它通过使用在主干分类模型内计算的中间潜在表示来门控最终输出来实现这一点。当添加到现有的SED系统中时,BPN实现了16.8%的精确度绝对提高,以及少数类d-调用和bp-调用的F1分数分别提高了21.3%和9.4%。我们进一步考虑两种方法来选择后处理超参数:向前搜索和向后搜索。通过分别优化事件级和帧级超参数,这两种方法导致使用经验方法选择的参数的相当大的性能改进。完整的WhaleVAD-BPN系统实现了0.475的交叉验证开发F1评分,比基线绝对改善了9.8%。
摘要:While recent sound event detection (SED) systems can identify baleen whale calls in marine audio, challenges related to false positive and minority-class detection persist. We propose the boundary proposal network (BPN), which extends an existing lightweight SED system. The BPN is inspired by work in image object detection and aims to reduce the number of false positive detections. It achieves this by using intermediate latent representations computed within the backbone classification model to gate the final output. When added to an existing SED system, the BPN achieves a 16.8 % absolute increase in precision, as well as 21.3 % and 9.4 % improvements in the F1-score for minority-class d-calls and bp-calls, respectively. We further consider two approaches to the selection of post-processing hyperparameters: a forward-search and a backward-search. By separately optimising event-level and frame-level hyperparameters, these two approaches lead to considerable performance improvements over parameters selected using empirical methods. The complete WhaleVAD-BPN system achieves a cross-validated development F1-score of 0.475, which is a 9.8 % absolute improvement over the baseline.


【4】SpecTokenizer: A Lightweight Streaming Codec in the Compressed Spectrum Domain
标题:SpecTokenizer:压缩频谱域中的轻量级流编解码器
链接:https://arxiv.org/abs/2510.21209

作者:Zixiang Wan, Guochang Zhang, Yifeng He, Jianqiang Wei
备注:Accepted by Interspeech 2025; 5 pages, 1 figure, 5 tables
摘要:近年来,神经音频编解码器(NAC)作为语音语言模型中的音频压缩和音频表示技术得到了越来越多的关注。虽然主流的NAC通常需要G级计算和M级参数,但轻量级和流式NAC的性能仍有待研究。本文提出了SpecTokenizer,一个轻量级的流媒体编解码器,在压缩频谱域。SpecTokenizer仅由交替的CNN和RNN层组成,通过压缩频谱域中的多尺度建模实现了更高的效率和更好的表示能力。在4 kbps的速度下,与具有最先进的轻量级架构的编解码器相比,所提出的SpecTokenizer实现了相当或更高的性能,同时只需要20%的计算和10%的参数。此外,当使用类似的计算和存储资源时,它的性能明显优于编解码器。
摘要:Neural Audio Codecs (NACs) have gained growing attention in recent years as technologies for audio compression and audio representation in speech language models. While mainstream NACs typically require G-level computation and M-level parameters, the performance of lightweight and streaming NACs remains underexplored. This paper proposes SpecTokenizer, a lightweight streaming codec that operates in the compressed spectral domain. Composed solely of alternating CNN and RNN layers, SpecTokenizer achieves greater efficiency and better representational capability through multi-scale modeling in the compressed spectrum domain. At 4 kbps, the proposed SpecTokenizer achieves comparable or superior performance compared to the codec with state-of-the-art lightweight architecture while requiring only 20% of the computation and 10% of the parameters. Furthermore, it significantly outperforms the codec when using similar computational and storage resources.


【5】PhoenixCodec: Taming Neural Speech Coding for Extreme Low-Resource Scenarios
标题:PhoenixCodec:针对极低资源场景驯服神经语音编码
链接:https://arxiv.org/abs/2510.21196

作者:Zixiang Wan, Haoran Zhao, Guochang Zhang, Runqiang Han, Jianqiang Wei, Yuexian Zou
备注:5 pages, 1 figure, 4 tables
摘要:本文介绍了PhoenixCodec,一个全面的神经语音编码和解码框架,专为极低的资源条件。该系统集成了优化的非对称频率-时间架构,循环校准和细化(CCR)的训练策略,和噪声不变的微调过程。在严格的限制下-计算低于700 MFLOPs,延迟小于30 ms,以及支持1 kbps和6 kbps的双速率-现有方法面临效率和质量之间的权衡。PhoenixCodec通过减轻传统解码器的资源分散、采用CCR来避免局部最优以及通过噪声样本微调来增强鲁棒性来解决这些挑战。在LRAC 2025 Challenge Track 1中,该系统总体排名第三,并在清洁测试中在1 kbps的真实世界噪声和混响以及清晰度方面表现出最佳性能,证实了其有效性。
摘要:This paper presents PhoenixCodec, a comprehensive neural speech coding and decoding framework designed for extremely low-resource conditions. The proposed system integrates an optimized asymmetric frequency-time architecture, a Cyclical Calibration and Refinement (CCR) training strategy, and a noise-invariant fine-tuning procedure. Under stringent constraints - computation below 700 MFLOPs, latency less than 30 ms, and dual-rate support at 1 kbps and 6 kbps - existing methods face a trade-off between efficiency and quality. PhoenixCodec addresses these challenges by alleviating the resource scattering of conventional decoders, employing CCR to escape local optima, and enhancing robustness through noisy-sample fine-tuning. In the LRAC 2025 Challenge Track 1, the proposed system ranked third overall and demonstrated the best performance at 1 kbps in both real-world noise and reverberation and intelligibility in clean tests, confirming its effectiveness.


【6】refess-qi: reference-free evaluation for speech separation with joint quality and intelligibility scoring
标题:refess-qi:具有联合质量和可理解性评分的语音分离无参考评估
链接:https://arxiv.org/abs/2510.21014

作者:Ari Frummer, Helin Wang, Tianyu Cao, Adi Arbel, Yuval Sieradzki, Oren Gal, Jesús Villalba, Thomas Thebaud, Najim Dehak
摘要:源分离是各种语音处理任务(如自动语音识别(ASR))的关键预处理步骤。传统上,语音分离的评价指标依赖于匹配的参考音频和相应的瞬态来评估音频质量和可懂度。然而,它们不能用于评估没有参考存在的真实世界的混合物。本文介绍了一种基于自监督学习(SSL)表示的无文本无参考评估框架。建议的框架利用混合和分离的轨道来预测联合音频质量,通过标度不变的信噪比(SI-SNR)度量,和语音清晰度通过字错误率(WER)度量。我们在WHAMR上进行了实验!数据集,其示出了WER估计,平均绝对误差(MAE)为17%,皮尔逊相关系数(PCC)为0.77;以及SI-SNR估计,MAE为1.38,PCC为0.95。我们进一步证明了我们的估计的鲁棒性,通过使用各种SSL表示。
摘要:Source separation is a crucial pre-processing step for various speech processing tasks, such as automatic speech recognition (ASR). Traditionally, the evaluation metrics for speech separation rely on the matched reference audios and corresponding transcriptions to assess audio quality and intelligibility. However, they cannot be used to evaluate real-world mixtures for which no reference exists. This paper introduces a text-free reference-free evaluation framework based on self-supervised learning (SSL) representations. The proposed framework utilize the mixture and separated tracks to predict jointly audio quality, through the Scale Invariant Signal to Noise Ratio (SI-SNR) metric, and speech intelligibility through the Word Error Rate (WER) metric. We conducted experiments on the WHAMR! dataset, which shows a WER estimation with a mean absolute error (MAE) of 17\% and a Pearson correlation coefficient (PCC) of 0.77; and SI-SNR estimation with an MAE of 1.38 and PCC of 0.95. We further demonstrate the robustness of our estimator by using various SSL representations.


【7】Data-Centric Lessons To Improve Speech-Language Pretraining
标题:以数据为中心的课程,以提高语音语言预训练
链接:https://arxiv.org/abs/2510.20860

作者:Vishaal Udandarao, Zhiyun Lu, Xuankai Chang, Yongqiang Wang, Violet Z. Yao, Albin Madapally Jose, Fartash Faghri, Josh Gardner, Chung-Cheng Chiu
备注:Tech Report
摘要:语音识别(SQA)是有用的和交互式人工智能系统的核心能力。最近,已经发布了几个语音语言模型(SpeechLM),重点是提高其SQA性能。然而,由于缺乏对预训练数据处理和策展的控制,尽管从其他数据模式的类似研究中获得了大量收益,但理解哪些因素影响了性能仍然具有挑战性。在这项工作中,我们通过对预训练SpeechLM进行以数据为中心的探索来解决这一差距。我们专注于语音语言预训练数据的三个基本研究问题:(1)如何处理原始网络抓取的音频内容进行语音文本预训练,(2)如何构建合成预训练数据集来增强网络抓取的数据,以及(3)如何将(文本,音频)片段交织到训练序列中。我们应用我们控制的以数据为中心的消融的见解来预训练3.8B参数的SpeechLM,称为SpeLangy,其性能优于高达3倍的模型,绝对性能为10.2%。我们希望我们的研究结果能够突出有效的数据策展对语音语言预训练的影响,并指导SpeechLM未来以数据为中心的探索。
摘要:Spoken Question-Answering (SQA) is a core capability for useful and interactive artificial intelligence systems. Recently, several speech-language models (SpeechLMs) have been released with a specific focus on improving their SQA performance. However, a lack of controlled ablations of pretraining data processing and curation makes it challenging to understand what factors account for performance, despite substantial gains from similar studies in other data modalities. In this work, we address this gap by conducting a data-centric exploration for pretraining SpeechLMs. We focus on three research questions fundamental to speech-language pretraining data: (1) how to process raw web-crawled audio content for speech-text pretraining, (2) how to construct synthetic pretraining datasets to augment web-crawled data and (3) how to interleave (text, audio) segments into training sequences. We apply the insights from our controlled data-centric ablations to pretrain a 3.8B-parameter SpeechLM, called SpeLangy, that outperforms models that are up to 3x larger by 10.2% absolute performance. We hope our findings highlight the impact of effective data curation for speech-language pretraining and guide future data-centric exploration in SpeechLMs.


【8】Beyond Hearing: Learning Task-agnostic ExG Representations from Earphones via Physiology-informed Tokenization
标题:超越听力:通过基于生理学的代币化从耳机学习任务不可知的ExG表示
链接:https://arxiv.org/abs/2510.20853

作者:Hyungjun Yoon, Seungjoo Lee, Yu Yvonne Wu, Xiaomeng Chen, Taiting Lu, Freddy Yifei Liu, Taeckyung Lee, Hyeongheon Cha, Haochen Zhao, Gaoteng Zhao, Sung-Ju Lee, Cecilia Mascolo, Dongyao Chen, Lili Qiu
备注:19 pages, 9 figures
摘要:电生理学(ExG)信号提供了对人类生理学的有价值的见解,但是由于两个关键限制,构建跨日常任务推广的基础模型仍然具有挑战性:(i)数据多样性不足,因为大多数ExG记录是在具有庞大昂贵设备的受控实验室中收集的;以及(ii)需要定制处理的任务特定模型设计(即,目标频率滤波器)和架构,这限制了跨任务的泛化。为了应对这些挑战,我们引入了一种可扩展的、与任务无关的ExG监控方法。我们使用基于耳机的硬件原型收集了50小时的不显眼的自由生活ExG数据,以缩小数据多样性差距。我们的方法的核心是生理信息多波段令牌化(PiMT),它将ExG信号分解为12个生理信息令牌,然后进行重建任务以学习鲁棒的表示。这使得能够在捕获任务相关信息的同时在整个频谱上进行自适应特征识别。我们的新DailySense数据集(第一个支持对五种人类感官进行基于ExG的分析的数据集)以及四个公共ExG基准测试的实验表明,PiMT在不同任务中的表现始终优于最先进的方法。
摘要:Electrophysiological (ExG) signals offer valuable insights into human physiology, yet building foundation models that generalize across everyday tasks remains challenging due to two key limitations: (i) insufficient data diversity, as most ExG recordings are collected in controlled labs with bulky, expensive devices; and (ii) task-specific model designs that require tailored processing (i.e., targeted frequency filters) and architectures, which limit generalization across tasks. To address these challenges, we introduce an approach for scalable, task-agnostic ExG monitoring in the wild. We collected 50 hours of unobtrusive free-living ExG data with an earphone-based hardware prototype to narrow the data diversity gap. At the core of our approach is Physiology-informed Multi-band Tokenization (PiMT), which decomposes ExG signals into 12 physiology-informed tokens, followed by a reconstruction task to learn robust representations. This enables adaptive feature recognition across the full frequency spectrum while capturing task-relevant information. Experiments on our new DailySense dataset-the first to enable ExG-based analysis across five human senses-together with four public ExG benchmarks, demonstrate that PiMT consistently outperforms state-of-the-art methods across diverse tasks.


【9】Can large audio language models understand child stuttering speech? speech summarization, and source separation
标题:大型音频语言模型能理解儿童口吃的言语吗?语音摘要和源分离
链接:https://arxiv.org/abs/2510.20850

作者:Chibuzor Okocha, Maya Bakri, Christan Grant
备注:7 pages, 1 Figure, 8 tables, Under review ICASSP 2026
摘要:儿童语音在声学、韵律和语言发展方面与成人语音不同,不流利(重复、重复、块)进一步挑战了自动语音识别(ASR)和下游自然语言处理(NLP)。最近的大型音频语言模型(LALM)表现出强大的跨模态音频理解,然而,他们的行为在不流利的儿童语音仍然未充分探索。我们在两种情况下评估了几种最先进的LALM:采访(混合扬声器)和阅读任务(独生子女)。这些任务是(i)单通道源分离,以隔离儿童和(ii)仅儿童总结,保留临床相关的不流利,并避免成人语音泄漏。   评估结合了大型语言模型(LLM)作为判断,人类专家评级和BERTScore(F1),我们报告了模型之间以及模型与人类之间的一致性,以评估可靠性。我们的研究结果描述了LALM从混合音频中产生忠实的儿童摘要的条件以及它们失败的地方,为临床和教育部署提供了实用的指导。我们提供提示和评估脚本来支持复制。
摘要:Child speech differs from adult speech in acoustics, prosody, and language development, and disfluencies (repetitions, prolongations, blocks) further challenge Automatic Speech Recognition (ASR) and downstream Natural Language Processing (NLP). Recent large audio-language models (LALMs) demonstrate strong cross-modal audio understanding; however, their behavior in disfluent child speech remains underexplored. We evaluate several state-of-the-art LALMs in two settings: an interview (mixed speakers) and a reading task (single child). The tasks are (i) single-channel source separation to isolate the child and (ii) child-only summarization that preserves clinically relevant disfluencies and avoids adult-speech leakage.   Evaluation combines Large Language Model (LLM) as a judge, human expert ratings, and BERTScore (F1), and we report agreement between models and between models and humans to assess reliability. Our findings delineate the conditions under which LALMs produce faithful child-only summaries from mixed audio and where they fail, offering practical guidance for clinical and educational deployments. We provide prompts and evaluation scripts to support replication.


【10】FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
标题:FlexIO:灵活的单通道和多通道语音分离和增强
链接:https://arxiv.org/abs/2510.21485

作者:Yoshiki Masuyama, Kohei Saijo, Francesco Paissan, Jiangyu Han, Marc Delcroix, Ryo Aihara, François G. Germain, Gordon Wichern, Jonathan Le Roux
备注:Submitted to ICASSP 2026
摘要:语音分离与增强(SSE)技术已经取得了显著的进步,并在控制环境中取得了可喜的成果,如固定数量的扬声器和固定的阵列配置。对于通用SSE系统,单声道系统已经被扩展以处理可变数量的扬声器(即,产出)。同时,适应各种阵列配置的多通道系统(即,投入)已经开发。然而,这些尝试是分开进行的。在本文中,我们提出了一个灵活的输入和输出SSE系统,名为FlexIO。它使用提示向量执行条件分离,每个扬声器一个作为条件,允许分离任意数量的扬声器。通过阵列不可知信道通信机制,多信道混合物与提示向量一起被处理。我们的实验表明,FlexIO成功地覆盖了一到五个麦克风和一到三个扬声器的各种条件。我们还证实了FlexIO在CHiME-4真实数据上的鲁棒性。
摘要:Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel systems have been extended to deal with a variable number of speakers (i.e., outputs). Meanwhile, multi-channel systems accommodating various array configurations (i.e., inputs) have been developed. However, these attempts have been pursued separately. In this paper, we propose a flexible input and output SSE system, named FlexIO. It performs conditional separation using prompt vectors, one per speaker as a condition, allowing separation of an arbitrary number of speakers. Multi-channel mixtures are processed together with the prompt vectors via an array-agnostic channel communication mechanism. Our experiments demonstrate that FlexIO successfully covers diverse conditions with one to five microphones and one to three speakers. We also confirm the robustness of FlexIO on CHiME-4 real data.


机器翻译由腾讯交互翻译提供,仅供参考