微信公众号:arXiv_Daily
cs.SD语音
标题:调整自我监督的语音表达用于帕金森病的跨舌发音障碍检测
链接:https://arxiv.org/abs/2603.22225
备注:Submitted to Interspeech 2026
摘要:构音障碍语音数据的有限性使得跨语言检测成为一个重要但具有挑战性的问题。一个关键的困难是,语音表征往往编码语言依赖的结构,可以混淆构音障碍检测。我们提出了一个代表级的语言转换(LS),对齐源语言自我监督的语音表示与目标语言分布使用基于质心的向量自适应估计健康控制语音。我们评估的方法,从帕金森氏病的语音数据集在捷克语,德语和西班牙语的跨语言和多语言设置下的口头DDK录音。LS大大提高了跨语言设置中的灵敏度和F1,同时在多语言设置中产生较小但一致的增益。表征分析进一步表明,LS减少了嵌入空间中的语言身份,支持LS消除语言依赖结构的解释。
摘要:The limited availability of dysarthric speech data makes cross-lingual detection an important but challenging problem. A key difficulty is that speech representations often encode language-dependent structure that can confound dysarthria detection. We propose a representation-level language shift (LS) that aligns source-language self-supervised speech representations with the target-language distribution using centroid-based vector adaptation estimated from healthy-control speech. We evaluate the approach on oral DDK recordings from Parkinson's disease speech datasets in Czech, German, and Spanish under both cross-lingual and multilingual settings. LS substantially improves sensitivity and F1 in cross-lingual settings, while yielding smaller but consistent gains in multilingual settings. Representation analysis further shows that LS reduces language identity in the embedding space, supporting the interpretation that LS removes language-dependent structure.
标题:AnimalCLAP:用于物种识别和特征推断的分类感知音频预训练
链接:https://arxiv.org/abs/2603.22053
备注:ICASSP 2026
摘要:动物发声为野生动物评估提供了重要的见解,特别是在森林等复杂环境中,有助于物种识别和生态监测。深度学习的最新进展使人们能够根据它们的发声进行自动物种分类。然而,对训练期间看不见的物种进行分类仍然具有挑战性。为了解决这个问题,我们引入AnimalCLAP,一个分类感知的语言音频框架,包括一个新的数据集和模型,其中包含分层的生物信息。具体来说,我们的发声数据集包括4,225小时的录音,涵盖6,823个物种,注释了22个生态特征。AnimalCLAP模型在此数据集上进行训练,以使用分类结构对齐音频和文本表示,从而提高对未知物种的识别。我们证明,我们提出的模型有效地推断生态和生物属性的物种直接从他们的发声,实现优越的性能相比,CLAP。我们的数据集,代码和模型将在https://dahlian00.github.io/AnimalCLAP_Page/上公开。
摘要:Animal vocalizations provide crucial insights for wildlife assessment, particularly in complex environments such as forests, aiding species identification and ecological monitoring. Recent advances in deep learning have enabled automatic species classification from their vocalizations. However, classifying species unseen during training remains challenging. To address this limitation, we introduce AnimalCLAP, a taxonomy-aware language-audio framework comprising a new dataset and model that incorporate hierarchical biological information. Specifically, our vocalization dataset consists of 4,225 hours of recordings covering 6,823 species, annotated with 22 ecological traits. The AnimalCLAP model is trained on this dataset to align audio and textual representations using taxonomic structures, improving the recognition of unseen species. We demonstrate that our proposed model effectively infers ecological and biological attributes of species directly from their vocalizations, achieving superior performance compared to CLAP. Our dataset, code, and models will be publicly available at https://dahlian00.github.io/AnimalCLAP_Page/.
标题:LipsAM:用于音频信号处理的Lipschitz连续幅度修改器及其在即插即用去回响中的应用
链接:https://arxiv.org/abs/2603.21684
备注:Accepted for IEEE ICASSP 2026
摘要:深度神经网络(DNN)的鲁棒性可以通过其Lipschitz连续性来证明,这使得Lipschitz连续DNN的构建成为一个活跃的研究领域。然而,用于音频处理的DNN由于与现有结果的兼容性差而没有成为主要焦点。在本文中,我们考虑的幅度修改器(AM),一个流行的架构处理音频信号,并提出其Lipschitz连续的变种,我们称之为LipsAM。我们证明了一个充分条件的AM是Lipschitz连续的,并提出了两个架构的例子LipsAM。所提出的架构被应用到一个即插即用的语音去混响算法,并通过数值实验证明其改善的稳定性。
摘要:The robustness of deep neural networks (DNNs) can be certified through their Lipschitz continuity, which has made the construction of Lipschitz-continuous DNNs an active research field. However, DNNs for audio processing have not been a major focus due to their poor compatibility with existing results. In this paper, we consider the amplitude modifier (AM), a popular architecture for handling audio signals, and propose its Lipschitz-continuous variants, which we refer to as LipsAM. We prove a sufficient condition for an AM to be Lipschitz continuous and propose two architectures as examples of LipsAM. The proposed architectures were applied to a Plug-and-Play algorithm for speech dereverberation, and their improved stability is demonstrated through numerical experiments.
标题:企业销售副驾驶:在实时销售电话中通过自动信息检索实现实时人工智能支持
链接:https://arxiv.org/abs/2603.21416
摘要:在现场销售电话中,客户经常询问详细的产品问题,这需要代表手动搜索内部数据库和CRM系统。这个过程通常需要25 - 65秒的查询,创建尴尬的停顿,损害客户体验,降低销售效率。我们介绍了SalesCopilot,这是一个实时的人工智能助手,它通过自动检测客户问题,从产品数据库中检索相关信息,并在几秒钟内在代表的仪表板上显示简洁的答案来消除这一瓶颈。该系统集成了流语音到文本转录,大语言模型(LLM)为基础的问题检测,检索增强生成(RAG)在一个结构化的产品数据库到一个统一的实时管道。我们在一个保险销售场景中演示了SalesCopilot,其中包含10个类别的50种产品(2,490个常见问题解答,290个覆盖范围详细信息和162个定价层)。在我们的基准评估中,SalesCopilot的实测平均响应时间为2.8秒,问题检测率为100%,与内部研究中的手动CRM搜索相比,速度提高了14倍。该系统与领域无关,可以通过替换产品数据库来适应任何企业销售领域。
摘要:During live sales calls, customers frequently ask detailed product questions that require representatives to manually search internal databases and CRM systems. This process typically takes 25-65 seconds per query, creating awkward pauses that hurt customer experience and reduce sales efficiency. We present SalesCopilot, a real-time AI-powered assistant that eliminates this bottleneck by automatically detecting customer questions, retrieving relevant information from the product database, and displaying concise answers on the representative's dashboard in seconds. The system integrates streaming speech-to-text transcription, large language model (LLM)-based question detection, and retrieval-augmented generation (RAG) over a structured product database into a unified real-time pipeline. We demonstrate SalesCopilot on an insurance sales scenario with 50 products spanning 10 categories (2,490 FAQs, 290 coverage details, and 162 pricing tiers). In our benchmark evaluation, SalesCopilot achieves a measured mean response time of 2.8 seconds with 100% question detection rate, representing a 14xspeedup compared to manual CRM search in an internal study. The system is domain-agnostic and can be adapted to any enterprise sales domain by replacing the product database.
标题:HEIX:利用混合曼巴-注意力超越二次极限来扩展原始音频理解
链接:https://arxiv.org/abs/2603.21316
备注:10 Pages, 8 Figures
摘要:音频表示学习通常评估设计选择,例如输入前端,序列主干和序列长度。我们表明,这些轴是耦合的,从一个设置的结论往往不会转移到其他。我们介绍HELIX,一个控制框架比较纯曼巴,纯注意力,和一个最小的混合动力与一个单一的注意力瓶颈。所有模型都在大约8.3M参数下进行参数匹配,以隔离建筑效果。在六个数据集上,我们发现首选的输入表示取决于主干,并且注意力会损害短的固定音频的性能,但在较长的序列长度上变得重要。在一个5分钟的说话人识别任务中,有30,000个标记,纯注意力由于内存不足而失败,而HELIX比纯Mamba缩小了11.5分的差距。
摘要:Audio representation learning typically evaluates design choices such as input frontend, sequence backbone, and sequence length in isolation. We show that these axes are coupled, and conclusions from one setting often do not transfer to others. We introduce HELIX, a controlled framework comparing pure Mamba, pure attention, and a minimal hybrid with a single attention bottleneck. All models are parameter-matched at about 8.3M parameters to isolate architectural effects. Across six datasets, we find that the preferred input representation depends on the backbone, and that attention hurts performance on short, stationary audio but becomes important at longer sequence lengths. On a 5-minute speaker identification task with 30,000 tokens, pure attention fails with out-of-memory errors, while HELIX closes an 11.5-point gap over pure Mamba.
标题:融合记忆和注意力:LSTM、Transformer和符号音乐生成混合架构的研究
链接:https://arxiv.org/abs/2603.21282
备注:20 pages, 6 figures. Published in Expert Systems with Applications (Elsevier), 2026. DOI: https://doi.org/10.1016/j.eswa.2026.131173
摘要:机器学习技术,如Transformers和长短期记忆(LSTM)网络,在符号音乐生成(SMG)中发挥着至关重要的作用。现有的文献表明LSTM和Transformers在模拟局部旋律连续性与维持全局结构一致性的能力方面存在差异。然而,它们在SMG背景下的具体性质尚未得到系统的研究。本文通过对SMG的LSTM与Transformers进行细粒度的比较分析,使用Deutschl数据集上的17个音乐质量指标详细检查本地和全局属性,从而解决了这一差距。我们发现LSTM网络擅长捕捉局部模式,但无法保持长期依赖关系,而Transformers有效地建模全局结构,但往往会产生不规则的措辞。基于这种分析并利用它们各自的优势,我们提出了一种将Transformer编码器与LSTM解码器相结合的混合架构,并根据两个基线对其进行评估。我们在Deutschl数据集上评估了三种架构中每种架构生成的1,000首旋律。结果表明,与基线相比,混合方法实现了更好的局部和全局连续性和一致性。我们的工作突出了这些模型的关键特征,并演示了如何利用它们的属性来设计卓越的模型。我们还支持消融研究和人类感知评估的实验,这些实验在统计学上支持这些发现,并为这项工作提供了有力的验证。
摘要:Machine learning techniques, such as Transformers and Long Short-Term Memory (LSTM) networks, play a crucial role in Symbolic Music Generation (SMG). Existing literature indicates a difference between LSTMs and Transformers regarding their ability to model local melodic continuity versus maintaining global structural coherence. However, their specific properties within the context of SMG have not been systematically studied. This paper addresses this gap by providing a fine-grained comparative analysis of LSTMs versus Transformers for SMG, examining local and global properties in detail using 17 musical quality metrics on the Deutschl dataset. We find that LSTM networks excel at capturing local patterns but fail to preserve long-range dependencies, while Transformers model global structure effectively but tend to produce irregular phrasing. Based on this analysis and leveraging their respective strengths, we propose a Hybrid architecture combining a Transformer Encoder with an LSTM Decoder and evaluate it against both baselines. We evaluated 1,000 generated melodies from each of the three architectures on the Deutschl dataset. The results show that the hybrid method achieves better local and global continuity and coherence compared to the baselines. Our work highlights the key characteristics of these models and demonstrates how their properties can be leveraged to design superior models. We also supported the experiments with ablation studies and human perceptual evaluations, which statistically support the findings and provide robust validation for this work.
标题:离散语音表达的情感感知量化分析
链接:https://arxiv.org/abs/2603.21224
摘要:现代语音系统越来越多地使用离散化的自监督语音表示来压缩和集成基于令牌的模型,但它们对情感信息的影响仍然不清楚。我们研究了如何残余矢量量化(RVQ)重塑情感信息的离散语音表示从代表性和任务层面的角度来看。我们的分析表明,积极的压缩不成比例地降低情感,在情感类和模型架构之间的损失不均匀。为了解决这个问题,我们引入了情绪感知量化使用情绪特定的和情绪偏见的码本,提高硬和软的情感感知的保存。我们进一步提出了Emo-Q,这是一种轻量级的路由量化方法,可以选择情感专用码本,从而提高较低比特率下的情感识别性能。这些结果突出了情感感知离散化的重要性,强大的情感语音处理。
摘要:Modern speech systems increasingly use discretized self-supervised speech representations for compression and integration with token-based models, yet their impact on emotional information remains unclear. We study how residual vector quantization (RVQ) reshapes emotional information in discrete speech representations from both representation- and task-level perspectives. Our analysis shows that aggressive compression disproportionately degrades emotion, with uneven loss across emotion classes and model architectures. To address this, we introduce emotion-aware quantization using emotion-specific and emotion-biased codebooks, improving the preservation of both hard and soft emotion perception. We further propose Emo-Q, a lightweight routed quantization method that selects emotion-specialized codebooks, improving emotion recognition performance at lower bitrates. These results highlight the importance of emotion-aware discretization for robust affective speech processing.
标题:评估神经TTC系统建模辅音引起的F0扰动的能力
链接:https://arxiv.org/abs/2603.21078
备注:Accepted for publication in Computer Speech & Language
摘要:本研究提出了一个段级韵律探测框架来评估神经TTS模型再现辅音引起的f0扰动的能力,这是一种反映局部发音机制的细粒度段韵律效应。我们比较合成和自然语音实现数千个单词,分层的词频,使用Tacotron 2和FastSpeech 2训练相同的语音语料库(LJ语音)。这些控制分析,然后补充了一个大规模的评估,跨越多个先进的TTS系统。结果表明,准确再现高频词,但低频率项目的泛化能力差,这表明所研究的TTS架构更依赖于词汇级的记忆比抽象的分段韵律编码。这一发现突出了这样的TTS系统的能力,概括韵律细节以外看到的数据的限制。建议的探针提供了一个语言学上知情的诊断框架,可能会告知未来的TTS评估方法,并在合成语音的可解释性和真实性评估的影响。
摘要:This study proposes a segmental-level prosodic probing framework to evaluate neural TTS models' ability to reproduce consonant-induced f0 perturbation, a fine-grained segmental-prosodic effect that reflects local articulatory mechanisms. We compare synthetic and natural speech realizations for thousands of words, stratified by lexical frequency, using Tacotron 2 and FastSpeech 2 trained on the same speech corpus (LJ Speech). These controlled analyses are then complemented by a large-scale evaluation spanning multiple advanced TTS systems. Results show accurate reproduction for high-frequency words but poor generalization to low-frequency items, suggesting that the examined TTS architectures rely more on lexical-level memorization than on abstract segmental-prosodic encoding. This finding highlights a limitation in such TTS systems' ability to generalize prosodic detail beyond seen data. The proposed probe offers a linguistically informed diagnostic framework that may inform future TTS evaluation methods, and has implications for interpretability and authenticity assessment in synthetic speech.
标题:ERM-MinMaxGAP:多语言多模式语音中的基准和缓解性别偏见-LLM情感识别
链接:https://arxiv.org/abs/2603.21050
摘要:语音情感识别(SER)系统可以表现出与性别相关的性能差异,但这种偏见如何体现在跨语言和模态的多语言语音LLM中尚不清楚。我们介绍了一种新的多语言,多模式的基准建立在MELD-ST,跨越英语,日语和德语,量化语言特定的SER性能和性别差距。我们发现偏见是强烈的语言依赖性,多模态融合并不能可靠地提高公平性。为了解决这些问题,我们提出了ERM-MinMaxGAP,一个公平的训练目标,它增强了经验风险最小化(ERM)与建议的自适应公平权重机制和一种新的MinMaxGAP正则化器的最大男女损失差距在每种语言和模态。基于Qwen 2-Audio主干,我们的ERM-MinMaxGAP方法将多语言SER性能提高了5.5%和5.0%,同时在单模态和多模态设置中分别将整体性别偏见差距降低了0.1%和1.4%。
摘要:Speech emotion recognition (SER) systems can exhibit gender-related performance disparities, but how such bias manifests in multilingual speech LLMs across languages and modalities is unclear. We introduce a novel multilingual, multimodal benchmark built on MELD-ST, spanning English, Japanese, and German, to quantify language-specific SER performance and gender gaps. We find bias is strongly language-dependent, and multimodal fusion does not reliably improve fairness. To address these, we propose ERM-MinMaxGAP, a fairness-informed training objective, which augments empirical risk minimization (ERM) with a proposed adaptive fairness weight mechanism and a novel MinMaxGAP regularizer on the maximum male-female loss gap within each language and modality. Building upon the Qwen2-Audio backbone, our ERM-MinMaxGAP approach improves multilingual SER performance by 5.5% and 5.0% while reducing the overall gender bias gap by 0.1% and 1.4% in the unimodal and multimodal settings, respectively.
标题:SNAP:语音Deepfake检测中语音投影的说话者为零
链接:https://arxiv.org/abs/2603.20686
备注:9 pages, 3 figures, 2 tables
摘要:文本到语音技术的最新进展使得能够生成与真实人类声音几乎无法区分的高保真合成语音。虽然最近的研究显示了基于自我监督学习的语音编码器在深度伪造检测中的有效性,但这些模型很难在看不见的说话者中推广。我们的定量分析表明,这些编码器表示的扬声器信息的影响很大,导致检测器利用扬声器特定的相关性,而不是文物相关的线索。我们称这种现象为扬声器纠缠。为了减轻这种依赖,我们引入了SNAP,一个说话者置零框架。我们估计一个扬声器子空间,并应用正交投影来抑制扬声器相关的组件,隔离合成文物内的残留功能。通过减少扬声器纠缠,SNAP鼓励检测器专注于与伪影相关的模式,从而实现最先进的性能。
摘要:Recent advancements in text-to-speech technologies enable generating high-fidelity synthetic speech nearly indistinguishable from real human voices. While recent studies show the efficacy of self-supervised learning-based speech encoders for deepfake detection, these models struggle to generalize across unseen speakers. Our quantitative analysis suggests these encoder representations are substantially influenced by speaker information, causing detectors to exploit speaker-specific correlations rather than artifact-related cues. We call this phenomenon speaker entanglement. To mitigate this reliance, we introduce SNAP, a speaker-nulling framework. We estimate a speaker subspace and apply orthogonal projection to suppress speaker-dependent components, isolating synthesis artifacts within the residual features. By reducing speaker entanglement, SNAP encourages detectors to focus on artifact-related patterns, leading to state-of-the-art performance.
标题:ALICE:大型音频语言模型上下文学习能力的多方面评估框架
链接:https://arxiv.org/abs/2603.20433
备注:Submitted to Interspeech 2026
摘要:虽然大型音频语言模型(LALM)已被证明表现出退化的推理跟随能力,但它们从音频条件下的上下文示例中推断任务模式的能力尚未得到研究。为了解决这个问题,我们提出了ALICE,一个三阶段的框架,逐步减少文本指导,系统地评估LALM在音频条件下的上下文学习能力。在两个输出约束类别下评估四个音频理解任务中的六个LALM,我们发现所有阶段和LALM之间存在一致的不对称性:上下文演示可靠地提高了格式合规性,但未能提高,而且往往会降低核心任务的性能。这表明LALM可以从演示中收集表面级别的格式化模式,但可能难以利用跨模态语义基础来可靠地从音频条件示例中推断任务目标,突出了当前跨模态整合的潜在局限性。
摘要:While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples under audio conditioning remains unstudied. To address this gap, we present ALICE, a three-stage framework that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability under audio conditioning. Evaluating six LALMs across four audio understanding tasks under two output constraint categories, we uncover a consistent asymmetry across all stages and LALMs: in-context demonstrations reliably improve format compliance but fail to improve, and often degrade, the core task performance. This suggests that LALMs can glean surface-level formatting patterns from demonstrations but may struggle to leverage cross-modal semantic grounding to reliably infer task objectives from audio-conditioned examples, highlighting potential limitations in current cross-modal integration.
标题:EARTalking:具有逐帧控制的端到端GPT式自回归说话头部合成
链接:https://arxiv.org/abs/2603.20307
摘要:音频驱动的说话头部生成旨在从静态肖像和语音中创建生动逼真的视频。现有的基于AR的方法依赖于中间的面部表示,这限制了它们的表现力和真实感。同时,基于扩散的方法生成逐个剪辑,缺乏细粒度控制,并且由于整个窗口的整体去噪而导致固有延迟。为了解决这些局限性,我们提出了EARTalking,一种新颖的端到端,GPT风格的自回归模型,用于交互式音频驱动的说话头生成。我们的方法介绍了一种新的逐帧,在上下文中,音频驱动的流生成范例。为了在本质上支持可变长度视频生成和身份一致性,我们提出了Sink Frame Window Attention(SFA)机制。此外,为了避免复杂的,单独的网络,以前的作品需要不同的控制信号,我们提出了一个流帧条件上下文(FCIC)计划。该方案有效地注入不同的控制信号,在一个流,在上下文的方式,使交互式控制在每一帧和任意时刻。实验表明,EARTalking优于现有的自回归方法,并实现性能与基于扩散的方法。我们的工作证明了上下文流自回归控制的可行性,为灵活,高效的生成解锁了可扩展的方向。将发布代码以进行再现。
摘要:Audio-driven talking head generation aims to create vivid and realistic videos from a static portrait and speech. Existing AR-based methods rely on intermediate facial representations, which limit their expressiveness and realism. Meanwhile, diffusion-based methods generate clip-by-clip, lacking fine-grained control and causing inherent latency due to overall denoising across the window. To address these limitations, we propose EARTalking, a novel end-to-end, GPT-style autoregressive model for interactive audio-driven talking head generation. Our method introduces a novel frame-by-frame, in-context, audio-driven streaming generation paradigm. For inherently supporting variable-length video generation with identity consistency, we propose the Sink Frame Window Attention (SFA) mechanism. Furthermore, to avoid the complex, separate networks that prior works required for diverse control signals, we propose a streaming Frame Condition In-Context (FCIC) scheme. This scheme efficiently injects diverse control signals in a streaming, in-context manner, enabling interactive control at every frame and at arbitrary moments. Experiments demonstrate that EARTalking outperforms existing autoregressive methods and achieves performance comparable to diffusion-based methods. Our work demonstrates the feasibility of in-context streaming autoregressive control, unlocking a scalable direction for flexible, efficient generation. The code will be released for reproducibility.
标题:基于属性的视角下的语音隐私
链接:https://arxiv.org/abs/2603.20301
备注:Submitted to InterSpeech 2026
摘要:语音隐私保护方法,保持匿名的发言者修改讲话,试图打破与真实身份的发言者的联系。当前的基准测试基于信号对信号的比较来测量扬声器保护。在本文中,我们介绍了一个基于属性的角度来看,我们衡量隐私保护的扬声器属性集之间的比较。首先,我们分析隐私的影响,通过计算扬声器的唯一性地面真理属性,属性推断的原始语音,和属性推断的语音保护与标准匿名化。接下来,我们研究一个威胁的情况下,每个扬声器只涉及一个单一的话语,并计算攻击错误率。总体而言,我们观察到,推断的属性仍然存在风险,尽管属性推断错误。我们的研究指出了在未来的语音隐私研究中考虑属性相关威胁和保护机制的重要性。
摘要:Voice privacy approaches that preserve the anonymity of speakers modify speech in an attempt to break the link with the true identity of the speaker. Current benchmarks measure speaker protection based on signal-to-signal comparisons. In this paper, we introduce an attribute-based perspective, where we measure privacy protection in terms of comparisons between sets of speaker attributes. First, we analyze privacy impact by calculating speaker uniqueness for ground truth attributes, attributes inferred on the original speech, and attributes inferred on speech protected with standard anonymization. Next, we examine a threat scenario involving only a single utterance per speaker and calculate attack error rates. Overall, we observe that inferred attributes still present a risk despite attribute inference errors. Our research points to the importance of considering both attribute-related threats and protection mechanisms in future voice privacy research.
标题:Abjad-Kids:小学教育阿拉伯语语音分类数据集
链接:https://arxiv.org/abs/2603.20255
摘要:近年来,基于语音的人工智能教育应用引起了人们的极大兴趣,尤其是对儿童。然而,儿童语音研究仍然有限,由于缺乏公开的数据集,特别是低资源的语言,如Arabic.This本文介绍了Abjad-Kids,阿拉伯语语音数据集设计的幼儿园和小学教育,重点是字母,数字和颜色的基本学习。该数据集包括从3 - 12岁儿童中收集的46397个音频样本,覆盖141个班级。所有样品均按照受控质量标准记录,以确保持续时间、采样率和格式的一致性。为了解决阿拉伯语音素之间的高类内相似性和每个类的有限样本,我们提出了一种基于CNN-LSTM架构的分层音频分类。我们提出的方法将字母识别分解为两个阶段的过程:一个初始的分组分类模型,然后为每组专门的分类器。这两种策略:静态语言为基础的分组和动态聚类为基础的分组,进行了评估。实验结果表明,静态的基于语言的分组取得了优异的性能。传统机器学习与深度学习方法之间的比较突出了CNN-LSTM模型与数据增强相结合的有效性。尽管取得了令人鼓舞的结果,但我们的大多数实验都表明存在过拟合的挑战,这可能是由于样本数量有限,即使在数据增强和模型正则化之后。因此,今后的工作可能侧重于收集更多的数据,以解决这一问题。Abjad-Kids将向公众开放。我们希望Abjad-Kids能够丰富语音数据库中的儿童语音表示,并为未来的阿拉伯语儿童语音分类研究提供一个很好的资源。
摘要:Speech-based AI educational applications have gained significant interest in recent years, particularly for children. However, children speech research remains limited due to the lack of publicly available datasets, especially for low-resource languages such as Arabic.This paper presents Abjad-Kids, an Arabic speech dataset designed for kindergarten and primary education, focusing on fundamental learning of alphabets, numbers, and colors. The dataset consists of 46397 audio samples collected from children aged 3 - 12 years, covering 141 classes. All samples were recorded under controlled specifications to ensure consistency in duration, sampling rate, and format. To address high intra-class similarity among Arabic phonemes and the limited samples per class, we propose a hierarchical audio classification based on CNN-LSTM architectures. Our proposed methodology decomposes alphabet recognition into a two-stage process: an initial grouping classification model followed by specialized classifiers for each group. Both strategies: static linguistic-based grouping and dynamic clustering-based grouping, were evaluated. Experimental results demonstrate that static linguistic-based grouping achieves superior performance. Comparisons between traditional machine learning with deep learning approaches, highlight the effectiveness of CNN-LSTM models combined with data augmentation. Despite achieving promising results, most of our experiments indicate a challenge with overfitting, which is likely due to the limited number of samples, even after data augmentation and model regularization. Thus, future work may focus on collecting additional data to address this issue. Abjad-Kids will be publicly available. We hope that Abjad-Kids enrich children representation in speech dataset, and be a good resource for future research in Arabic speech classification for kids.
标题:LL-SDR:通过离散表示实现低延迟语音增强
链接:https://arxiv.org/abs/2603.20242
备注:5 pages, 1 figure
摘要:许多语音增强(SE)方法依赖于连续表示。最近,已经探索了离散音频令牌以实现SE的自回归生成。然而,目前尚不清楚离散化本身是否始终提高SE性能。在本文中,我们介绍了LL-SDR,一个基于令牌的语音增强框架,明确利用离散化更好地分离语音和噪声。我们的第一个贡献是方差排序的残差矢量量化器(VO-RVQ),旨在解开语音和噪声分布在令牌化。其次,我们提出了一个潜在的空间嵌入,以更好地调整增强嵌入与语义嵌入。实验表明,LL-SDR优于连续基线,并匹配基于自回归令牌的方法的性能,同时在混响和非混响噪声环境中实现轻量级,低延迟的语音增强。演示和源代码可在我们的项目网站上获得。
摘要:Many speech enhancement (SE) methods rely on continuous representations. Recently, discrete audio tokens have been explored to enable autoregressive generation for SE. However, it remains unclear whether discretization itself consistently improves SE performance. In this paper, we introduce LL-SDR, a token-based speech enhancement framework that explicitly leverages discretization to better separate speech and noise. Our first contribution is a Variance-Ordered Residual Vector Quantizer (VO-RVQ), designed to disentangle speech and noise distributions during tokenization. Second, we propose a latent-space discriminator to better align enhanced embeddings with semantic embeddings. Experiments show that LL-SDR outperforms continuous baselines and matches the performance of autoregressive token-based approaches, while enabling lightweight, low-latency speech enhancement in both reverberant and non-reverberant noisy environments. Demos and source code are available at our project websites.
标题:SelfTTC:通过显式嵌入解纠缠和使用自我增强的自我完善来实现跨说话者风格转移
链接:https://arxiv.org/abs/2603.22252
备注:Submitted to Interspeech 2026
摘要:本文介绍了SelfTTS,一个文本到语音(TTS)模型,设计用于跨扬声器风格的传输,消除了对外部预先训练的扬声器或情感编码器的需要。该架构实现了情感表达的中性扬声器通过明确的解纠缠策略,利用梯度反射层(GRL)结合余弦相似性损失解耦扬声器和情感信息。我们引入了多正对比学习(MPCL),以基于各自的标签来诱导说话者和情感嵌入的聚类表示。此外,SelfTTS通过自增强采用自细化策略,利用模型的语音转换能力来增强合成语音的自然度。实验结果表明,与最先进的基线相比,SelfTTS在目标音色和情感方面实现了卓越的情感自然度(eMOS)和鲁棒稳定性。
摘要:This paper presents SelfTTS, a text-to-speech (TTS) model designed for cross-speaker style transfer that eliminates the need for external pre-trained speaker or emotion encoders. The architecture achieves emotional expressivity in neutral speakers through an explicit disentanglement strategy utilizing Gradient Reversal Layers (GRL) combined with cosine similarity loss to decouple speaker and emotion information. We introduce Multi Positive Contrastive Learning (MPCL) to induce clustered representations of speaker and emotion embeddings based on their respective labels. Furthermore, SelfTTS employs a self-refinement strategy via Self-Augmentation, exploiting the model's voice conversion capabilities to enhance the naturalness of synthesized speech. Experimental results demonstrate that SelfTTS achieves superior emotional naturalness (eMOS) and robust stability in target timbre and emotion compared to state-of-the-art baselines.
标题:通过Chebyshev多元性和Riemannian度量学习解开说话者特征进行Deepfake源验证
链接:https://arxiv.org/abs/2603.21875
备注:Submitted to Interspeech 2026; The code, evaluation protocols and demo website are available at https://github.com/xxuan-acoustics/RiemannSD-Net
摘要:语音deepfake源验证系统旨在确定两个合成语音话语是否源自同一个源生成器,通常假设所产生的源嵌入与说话者特征无关。然而,这一假设仍未得到证实。在本文中,我们首先研究说话人因素对源验证的影响。我们提出了一个扬声器解纠缠度量学习(SDML)框架,其中包含两个新的损失函数。第一个利用切比雪夫多项式,以减轻解纠缠优化过程中的梯度不稳定性。第二个项目的源和扬声器嵌入到双曲空间,利用黎曼度量距离,以减少扬声器信息和学习更多的歧视性源功能。MLAAD基准测试的实验结果表明,SDML框架的有效性,在四个新提出的协议设计的源说话人解纠缠的情况下进行评估。代码、评估协议和演示网站可在https://github.com/xxuan-acoustics/RiemannSD-Net上获得。
摘要:Speech deepfake source verification systems aims to determine whether two synthetic speech utterances originate from the same source generator, often assuming that the resulting source embeddings are independent of speaker traits. However, this assumption remains unverified. In this paper, we first investigate the impact of speaker factors on source verification. We propose a speaker-disentangled metric learning (SDML) framework incorporating two novel loss functions. The first leverages Chebyshev polynomial to mitigate gradient instability during disentanglement optimization. The second projects source and speaker embeddings into hyperbolic space, leveraging Riemannian metric distances to reduce speaker information and learn more discriminative source features. Experimental results on MLAAD benchmark, evaluated under four newly proposed protocols designed for source-speaker disentanglement scenarios, demonstrate the effectiveness of SDML framework. The code, evaluation protocols and demo website are available at https://github.com/xxuan-acoustics/RiemannSD-Net.
标题:DiT-Flow:基于潜在空间流匹配和扩散变换的抗多失真语音增强
链接:https://arxiv.org/abs/2603.21608
摘要:生成模型的最新进展,如扩散和流量匹配,在音频任务中表现出强大的性能。然而,语音增强(SE)模型通常在有限的数据集上训练,并在狭窄的条件下进行评估,限制了现实世界的适用性。为了解决这个问题,我们提出了DiT-Flow,这是一个基于流匹配的SE框架,建立在潜在的扩散Transformer(DiT)骨干上,并经过训练,以适应各种失真的鲁棒性,包括噪声,混响和压缩。DiT-Flow对紧凑变分自动编码器(VAE)衍生的潜在特征进行操作。我们在StillSonicSet上验证了我们的方法,StillSonicSet是一个由LibriSpeech,FSD 50 K,FMA和90个Matterport 3D场景组成的合成但声学逼真的数据集。实验表明,DiT-Flow的性能始终优于最先进的生成SE模型,证明了流匹配在多条件语音增强中的有效性。尽管正在努力扩大合成数据的真实性,但SE中的一个持续瓶颈是培训和部署条件之间不可避免的不匹配。通过将LoRA与MoE框架相集成,我们实现了DiT-Flow对多种失真的鲁棒性的参数高效和高性能训练,使用总参数的4.9%来获得对五种不可见失真的更好性能。
摘要:Recent advances in generative models, such as diffusion and flow matching, have shown strong performance in audio tasks. However, speech enhancement (SE) models are typically trained on limited datasets and evaluated under narrow conditions, limiting real-world applicability. To address this, we propose DiT-Flow, a flow matching-based SE framework built on the latent Diffusion Transformer (DiT) backbone and trained for robustness across diverse distortions, including noise, reverberation, and compression. DiT-Flow operates on compact variational auto-encoders (VAEs)-derived latent features. We validated our approach on StillSonicSet, a synthetic yet acoustically realistic dataset composed of LibriSpeech, FSD50K, FMA, and 90 Matterport3D scenes. Experiments show that DiT-Flow consistently outperforms state-of-the-art generative SE models, demonstrating the effectiveness of flow matching in multi-condition speech enhancement. Despite ongoing efforts to expand synthetic data realism, a persistent bottleneck in SE is the inevitable mismatch between training and deployment conditions. By integrating LoRA with the MoE framework, we achieve both parameter-efficient and high-performance training for DiT-Flow robust to multiple distortions with using 4.9% percentage of the total parameters to obtain a better performance on five unseen distortions.
标题:SqueezeComposer:时间加速是长篇音乐创作的简单技巧
链接:https://arxiv.org/abs/2603.21073
备注:Under Review
摘要:由于建模长距离依赖关系的复杂性以及与冗长的音频表示相关联的过高的存储器和计算要求,创作连贯的长形式音乐仍然是一个重大挑战。在这项工作中,我们提出了一个简单而强大的技巧:我们假设AI模型可以理解并以2x,4x甚至8x的速率生成时间加速(加速)音频。通过首先生成音乐的高速版本,我们大大减少了时间长度和资源需求,使得处理长格式音乐变得可行,否则会超过内存或计算限制。然后将生成的音频恢复到其原始速度,恢复完整的时间结构。这种时间加速和减速策略自然遵循从抽象到详细内容的分层生成原则,并且可以方便地应用于现有的音乐生成模型,以实现长格式音乐生成。我们在SqueezeComposer中实例化了这个想法,该框架采用扩散模型在加速域中生成并在恢复域中细化。我们验证了这种方法在两个任务上的有效性:长形式的音乐生成,它评估时间方面的控制(包括延续,完成,并从零开始生成),和整首歌演唱伴奏生成,它评估轨道方面的控制。实验结果表明,我们简单的时间加速技巧,使高效,可扩展,高品质的长格式音乐生成。音频样本可在https://SqueezeComposer.github.io/上获得。
摘要:Composing coherent long-form music remains a significant challenge due to the complexity of modeling long-range dependencies and the prohibitive memory and computational requirements associated with lengthy audio representations. In this work, we propose a simple yet powerful trick: we assume that AI models can understand and generate time-accelerated (speeded-up) audio at rates such as 2x, 4x, or even 8x. By first generating a high-speed version of the music, we greatly reduce the temporal length and resource requirements, making it feasible to handle long-form music that would otherwise exceed memory or computational limits. The generated audio is then restored to its original speed, recovering the full temporal structure. This temporal speed-up and slow-down strategy naturally follows the principle of hierarchical generation from abstract to detailed content, and can be conveniently applied to existing music generation models to enable long-form music generation. We instantiate this idea in SqueezeComposer, a framework that employs diffusion models for generation in the accelerated domain and refinement in the restored domain. We validate the effectiveness of this approach on two tasks: long-form music generation, which evaluates temporal-wise control (including continuation, completion, and generation from scratch), and whole-song singing accompaniment generation, which evaluates track-wise control. Experimental results demonstrate that our simple temporal speed-up trick enables efficient, scalable, and high-quality long-form music generation. Audio samples are available at https://SqueezeComposer.github.io/.
标题:捆绑效应:多元线索如何形成TTS教学中的性别偏见
链接:https://arxiv.org/abs/2603.20743
备注:5 pages, 1 figure, 6 tables, Submitted to INTERSPEECH 2026
摘要:目前的教学文本到语音(ITTS)的偏见评估往往依赖于单变量测试,忽视了社会线索的组成结构。在这项工作中,我们调查性别偏见的社会地位,职业刻板印象和人格描述符的组合建模提示。通过分析开源ITTS模型,我们发现了系统的相互作用效应,其中社会维度相互调节,创造了单变量基线所错过的复杂偏差模式。至关重要的是,我们的研究结果表明,这些偏见超出了表面层面的工件,表现出与预先训练的文本编码器的语义先验和训练数据中固有的偏斜分布的强烈关联。我们进一步证明,通用多样性提示是不足以推翻这些根深蒂固的模式,强调成分分析,以诊断潜在的风险生成语音的需要。
摘要:Current bias evaluations in Instruction Text-to-Speech (ITTS) often rely on univariate testing, overlooking the compositional structure of social cues. In this work, we investigate gender bias by modeling prompts as combinations of Social Status, Career stereotypes, and Persona descriptors. Analyzing open-source ITTS models, we uncover systematic interaction effects where social dimensions modulate one another, creating complex bias patterns missed by univariate baselines. Crucially, our findings indicate that these biases extend beyond surface-level artifacts, demonstrating strong associations with the semantic priors of pre-trained text encoders and the skewed distributions inherent in training data. We further demonstrate that generic diversity prompting is insufficient to override these entrenched patterns, underscoring the need for compositional analysis to diagnose latent risks in generative speech.
标题:端到端多任务学习可调节关节降噪和听力损失补偿
链接:https://arxiv.org/abs/2603.20387
摘要:提出了一种多任务学习框架,用于优化单个深度神经网络(DNN),以实现联合降噪(NR)和听力损失补偿(HLC)。为每个任务定义一个不同的训练目标,DNN预测两个时频掩码。在推断期间,NR和HLC的量可以通过在组合它们之前对每个掩码取幂来独立地调整。与最近依赖于训练听觉模型仿真器来定义可微分训练目标的方法相比,我们提出了一种固有可微分的听觉模型,从而允许端到端优化。听力图作为DNN的输入提供,从而实现特定于听众的个性化,而无需再培训。结果表明,所提出的方法不仅允许单独调整NR和HLC的量,而且与优化单个训练目标相比,还提高了目标度量。它还优于分别针对NR和HLC训练的两个DNN的级联,并且与传统助听器处方相比,显示出具有竞争力的HLC性能。据我们所知,这是第一项使用听觉模型在广泛的听众配置文件中为NR和HLC训练单个DNN的研究。
摘要:A multi-task learning framework is proposed for optimizing a single deep neural network (DNN) for joint noise reduction (NR) and hearing loss compensation (HLC). A distinct training objective is defined for each task, and the DNN predicts two time-frequency masks. During inference, the amounts of NR and HLC can be adjusted independently by exponentiating each mask before combining them. In contrast to recent approaches that rely on training an auditory-model emulator to define a differentiable training objective, we propose an auditory model that is inherently differentiable, thus allowing end-to-end optimization. The audiogram is provided as an input to the DNN, thereby enabling listener-specific personalization without the need for retraining. Results show that the proposed approach not only allows adjusting the amounts of NR and HLC individually, but also improves objective metrics compared to optimizing a single training objective. It also outperforms a cascade of two DNNs that were separately trained for NR and HLC, and shows competitive HLC performance compared to a traditional hearing-aid prescription. To the best of our knowledge, this is the first study that uses an auditory model to train a single DNN for both NR and HLC across a wide range of listener profiles.
标题:太赫兹多用户大规模CDMA上行链路中的半盲信道估计和混合接收机束整形
链接:https://arxiv.org/abs/2603.22258
摘要:我们开发了一个实用的多用户(MU)大规模多输入多输出(MIMO)信道模型量身定制的太赫兹波段,包括分子吸收,反射损耗和多径漫射射线组件等因素。接下来,我们提出了一种新的基于半盲的信道状态信息(CSI)获取技术,即MU白化去相关半盲(MU-WD-SB),其利用与未知数据符号以及导频向量相对应的二阶统计。一个受约束的克拉美-罗下界(C-CRLB)推导出约束的归一化均方误差(NMSE)的性能所提出的半盲学习技术。我们提出的方案有效地减少了训练开销,同时提高了信道学习过程的整体精度。此外,一种新的混合接收机组合器框架设计的MU太赫兹大规模MIMO系统,利用多个测量矢量的稀疏贝叶斯学习(MMV-SBL),依赖于通过我们提出的半盲技术,依赖于低分辨率模数转换器(ADC)获得的估计CSI。最后,我们提出了一种基于MMV-SBL的最优混合合路器,它直接降低了MU干扰。进行了大量的模拟,以评估所提出的MU-WD-SB计划的性能增益超过传统的基于训练和其他半盲学习技术的实际太赫兹信道从高分辨率传输(HITRAN)数据库。用于量化改进的度量包括NMSE、误码率(BER)和频谱效率(SE)。
摘要:We develop a pragmatic multi-user (MU) massive multiple-input multiple-output (MIMO) channel model tailored to the THz band, encompassing factors such as molecular absorption, reflection losses and multipath diffused ray components. Next, we propose a novel semi-blind based channel state information (CSI) acquisition technique i.e. MU whitening decorrelation semi-blind (MU-WD-SB) that exploits the second order statistics corresponding to the unknown data symbols along with pilot vectors. A constrained Cramer-Rao Lower Bound (C-CRLB) is derived to bound the normalized mean square error (NMSE) performance of the proposed semi-blind learning technique. Our proposed scheme efficiently reduces the training overheads while enhancing the overall accuracy of the channel learning process. Furthermore, a novel hybrid receiver combiner framework is devised for MU THz massive MIMO systems, leveraging multiple measurement vector based sparse Bayesian learning (MMV-SBL) that relies on the estimated CSI acquired through our proposed semi-blind technique relying on low resolution analog-to-digital converters (ADCs). Finally, we propose an optimal hybrid combiner based on MMV-SBL, which directly reduces the MU interference. Extensive simulations are conducted to evaluate the performance gain of the proposed MU-WD-SB scheme over conventional training-based and other semi-blind learning techniques for a practical THz channel obtained from the high-resolution transmission (HITRAN) database. The metrics considered for quantifying the improvements include the NMSE, bit error rate (BER) and spectral-efficiency (SE).
标题:SelfTTC:通过显式嵌入解纠缠和使用自我增强的自我完善来实现跨说话者风格转移
链接:https://arxiv.org/abs/2603.22252
备注:Submitted to Interspeech 2026
摘要:本文介绍了SelfTTS,一个文本到语音(TTS)模型,设计用于跨扬声器风格的传输,消除了对外部预先训练的扬声器或情感编码器的需要。该架构实现了情感表达的中性扬声器通过明确的解纠缠策略,利用梯度反射层(GRL)结合余弦相似性损失解耦扬声器和情感信息。我们引入了多正对比学习(MPCL),以基于各自的标签来诱导说话者和情感嵌入的聚类表示。此外,SelfTTS通过自增强采用自细化策略,利用模型的语音转换能力来增强合成语音的自然度。实验结果表明,与最先进的基线相比,SelfTTS在目标音色和情感方面实现了卓越的情感自然度(eMOS)和鲁棒稳定性。
摘要:This paper presents SelfTTS, a text-to-speech (TTS) model designed for cross-speaker style transfer that eliminates the need for external pre-trained speaker or emotion encoders. The architecture achieves emotional expressivity in neutral speakers through an explicit disentanglement strategy utilizing Gradient Reversal Layers (GRL) combined with cosine similarity loss to decouple speaker and emotion information. We introduce Multi Positive Contrastive Learning (MPCL) to induce clustered representations of speaker and emotion embeddings based on their respective labels. Furthermore, SelfTTS employs a self-refinement strategy via Self-Augmentation, exploiting the model's voice conversion capabilities to enhance the naturalness of synthesized speech. Experimental results demonstrate that SelfTTS achieves superior emotional naturalness (eMOS) and robust stability in target timbre and emotion compared to state-of-the-art baselines.
标题:WiRD-Gest:在COTS硬件上使用距离多普勒Wi-Fi传感在现实世界中进行手势识别
链接:https://arxiv.org/abs/2603.22131
摘要:Wi-Fi传感已成为手势识别的一种有前途的技术,但其实际部署受到环境敏感性和设备放置挑战的阻碍。为了克服这些限制,我们提出了Wi-Fi测距和多普勒(WiRD)-Gest,这是一种新颖的系统,它使用商用现成(COTS)笔记本电脑上的单个未经修改的Wi-Fi收发器进行手势识别。该系统利用能够提取距离-多普勒(RD)信息的单基地全双工感测管线。利用这一点,我们提出了基于单站传感的手势识别深度学习模型的第一个基准。关键的创新在于,与现有方法相比,单基地传感和空间(范围)信息如何从根本上改变准确性,鲁棒性和通用性。我们在拥挤的、看不见的公共空间中表现出了出色的性能,即使只在受控环境中训练数据时,也有动态干扰和额外的移动目标。在这些场景中,先前的Wi-Fi传感方法经常失败,然而,我们的系统会遭受轻微的降级。WiRD-Gest基准测试和数据集也将作为开源发布。
摘要:Wi-Fi sensing has emerged as a promising technique for gesture recognition, yet its practical deployment is hindered by environmental sensitivity and device placement challenges. To overcome these limitations we propose Wi-Fi Range and Doppler (WiRD)-Gest, a novel system that performs gesture recognition using a single, unmodified Wi-Fi transceiver on a commercial off-the-shelf (COTS) laptop. The system leverages an monostatic full duplex sensing pipeline capable of extracting Range-Doppler (RD) information. Utilizing this, we present the first benchmark of deep learning models for gesture recognition based on monostatic sensing. The key innovation lies in how monostatic sensing and spatial (range) information fundamentally transforms accuracy, robustness and generalization compared to prior approaches. We demonstrate excellent performance in crowded, unseen public spaces with dynamic interference and additional moving targets even when trained on data from controlled environments only. These are scenarios where prior Wi-Fi sensing approaches often fail, however, our system suffers minor degradation. The WiRD-Gest benchmark and dataset will also be released as open source.
标题:自我监督语音表示的自适应联邦微调
链接:https://arxiv.org/abs/2603.21888
备注:Submitted to Interspeech 2026
摘要:将联邦学习(FL)与自监督学习(SSL)相结合,可以对语音任务进行隐私保护微调。然而,联邦环境表现出显著的异质性:客户端的计算能力不同,导致统一微调下的离散效应,而不同的下游任务需要不同的表示深度,使全模型更新效率低下。为了解决这些挑战,我们提出了一个自适应联邦微调框架与早期退出。轻量级预测头被插入SSL主干的中间层,允许客户端基于本地约束和任务要求终止计算。我们还引入了一个分层的,深度感知的部分聚合策略,以更好地利用来自不同网络深度的表示。实验表明,该框架减少了边缘开销,支持异构硬件,并在资源受限的联邦环境中保持有竞争力的性能。
摘要:Integrating Federated Learning (FL) with self-supervised learning (SSL) enables privacy-preserving fine-tuning for speech tasks. However, federated environments exhibit significant heterogeneity: clients differ in computational capacity, causing straggler effects under unified fine-tuning, while diverse downstream tasks require different representation depths, making full-model updates inefficient. To address these challenges, we propose an adaptive federated fine-tuning framework with early exits. Lightweight prediction heads are inserted at intermediate layers of the SSL backbone, allowing clients to terminate computation based on local constraints and task requirements. We further introduce a layer-wise, depth-aware partial aggregation strategy to better utilize representations from different network depths. Experiments show that the framework reduces edge overhead, supports heterogeneous hardware, and maintains competitive performance in resource-constrained federated environments.
标题:通过Chebyshev多元性和Riemannian度量学习解开说话者特征进行Deepfake源验证
链接:https://arxiv.org/abs/2603.21875
备注:Submitted to Interspeech 2026; The code, evaluation protocols and demo website are available at https://github.com/xxuan-acoustics/RiemannSD-Net
摘要:语音deepfake源验证系统旨在确定两个合成语音话语是否源自同一个源生成器,通常假设所产生的源嵌入与说话者特征无关。然而,这一假设仍未得到证实。在本文中,我们首先研究说话人因素对源验证的影响。我们提出了一个扬声器解纠缠度量学习(SDML)框架,其中包含两个新的损失函数。第一个利用切比雪夫多项式,以减轻解纠缠优化过程中的梯度不稳定性。第二个项目的源和扬声器嵌入到双曲空间,利用黎曼度量距离,以减少扬声器信息和学习更多的歧视性源功能。MLAAD基准测试的实验结果表明,SDML框架的有效性,在四个新提出的协议设计的源说话人解纠缠的情况下进行评估。代码、评估协议和演示网站可在https://github.com/xxuan-acoustics/RiemannSD-Net上获得。
摘要:Speech deepfake source verification systems aims to determine whether two synthetic speech utterances originate from the same source generator, often assuming that the resulting source embeddings are independent of speaker traits. However, this assumption remains unverified. In this paper, we first investigate the impact of speaker factors on source verification. We propose a speaker-disentangled metric learning (SDML) framework incorporating two novel loss functions. The first leverages Chebyshev polynomial to mitigate gradient instability during disentanglement optimization. The second projects source and speaker embeddings into hyperbolic space, leveraging Riemannian metric distances to reduce speaker information and learn more discriminative source features. Experimental results on MLAAD benchmark, evaluated under four newly proposed protocols designed for source-speaker disentanglement scenarios, demonstrate the effectiveness of SDML framework. The code, evaluation protocols and demo website are available at https://github.com/xxuan-acoustics/RiemannSD-Net.
标题:DiT-Flow:基于潜在空间流匹配和扩散变换的抗多失真语音增强
链接:https://arxiv.org/abs/2603.21608
摘要:生成模型的最新进展,如扩散和流量匹配,在音频任务中表现出强大的性能。然而,语音增强(SE)模型通常在有限的数据集上进行训练,并在狭窄的条件下进行评估,从而限制了现实世界的适用性。为了解决这个问题,我们提出了DiT-Flow,这是一个基于流匹配的SE框架,建立在潜在的扩散Transformer(DiT)骨干上,并经过训练,以适应各种失真的鲁棒性,包括噪声,混响和压缩。DiT-Flow对紧凑变分自动编码器(VAE)衍生的潜在特征进行操作。我们在StillSonicSet上验证了我们的方法,StillSonicSet是一个由LibriSpeech,FSD 50 K,FMA和90个Matterport 3D场景组成的合成但声学逼真的数据集。实验表明,DiT-Flow的性能始终优于最先进的生成SE模型,证明了流匹配在多条件语音增强中的有效性。尽管正在努力扩大合成数据的真实性,但SE中的一个持续瓶颈是培训和部署条件之间不可避免的不匹配。通过将LoRA与MoE框架相集成,我们实现了DiT-Flow对多种失真的鲁棒性的参数高效和高性能训练,使用总参数的4.9%来获得对五种不可见失真的更好性能。
摘要:Recent advances in generative models, such as diffusion and flow matching, have shown strong performance in audio tasks. However, speech enhancement (SE) models are typically trained on limited datasets and evaluated under narrow conditions, limiting real-world applicability. To address this, we propose DiT-Flow, a flow matching-based SE framework built on the latent Diffusion Transformer (DiT) backbone and trained for robustness across diverse distortions, including noise, reverberation, and compression. DiT-Flow operates on compact variational auto-encoders (VAEs)-derived latent features. We validated our approach on StillSonicSet, a synthetic yet acoustically realistic dataset composed of LibriSpeech, FSD50K, FMA, and 90 Matterport3D scenes. Experiments show that DiT-Flow consistently outperforms state-of-the-art generative SE models, demonstrating the effectiveness of flow matching in multi-condition speech enhancement. Despite ongoing efforts to expand synthetic data realism, a persistent bottleneck in SE is the inevitable mismatch between training and deployment conditions. By integrating LoRA with the MoE framework, we achieve both parameter-efficient and high-performance training for DiT-Flow robust to multiple distortions with using 4.9% percentage of the total parameters to obtain a better performance on five unseen distortions.
标题:SqueezeComposer:时间加速是长篇音乐创作的简单技巧
链接:https://arxiv.org/abs/2603.21073
备注:Under Review
摘要:由于建模长距离依赖关系的复杂性以及与冗长的音频表示相关联的过高的存储器和计算要求,创作连贯的长形式音乐仍然是一个重大挑战。在这项工作中,我们提出了一个简单而强大的技巧:我们假设AI模型可以理解并以2x,4x甚至8x的速率生成时间加速(加速)音频。通过首先生成音乐的高速版本,我们大大减少了时间长度和资源需求,使得处理长格式音乐变得可行,否则会超过内存或计算限制。然后将生成的音频恢复到其原始速度,恢复完整的时间结构。这种时间加速和减速策略自然遵循从抽象到详细内容的分层生成原则,并且可以方便地应用于现有的音乐生成模型,以实现长格式音乐生成。我们在SqueezeComposer中实例化了这个想法,该框架采用扩散模型在加速域中生成并在恢复域中细化。我们验证了这种方法在两个任务上的有效性:长形式的音乐生成,它评估时间方面的控制(包括延续,完成,并从零开始生成),和整首歌演唱伴奏生成,它评估轨道方面的控制。实验结果表明,我们简单的时间加速技巧,使高效,可扩展,高品质的长格式音乐生成。音频样本可在https://SqueezeComposer.github.io/上获得。
摘要:Composing coherent long-form music remains a significant challenge due to the complexity of modeling long-range dependencies and the prohibitive memory and computational requirements associated with lengthy audio representations. In this work, we propose a simple yet powerful trick: we assume that AI models can understand and generate time-accelerated (speeded-up) audio at rates such as 2x, 4x, or even 8x. By first generating a high-speed version of the music, we greatly reduce the temporal length and resource requirements, making it feasible to handle long-form music that would otherwise exceed memory or computational limits. The generated audio is then restored to its original speed, recovering the full temporal structure. This temporal speed-up and slow-down strategy naturally follows the principle of hierarchical generation from abstract to detailed content, and can be conveniently applied to existing music generation models to enable long-form music generation. We instantiate this idea in SqueezeComposer, a framework that employs diffusion models for generation in the accelerated domain and refinement in the restored domain. We validate the effectiveness of this approach on two tasks: long-form music generation, which evaluates temporal-wise control (including continuation, completion, and generation from scratch), and whole-song singing accompaniment generation, which evaluates track-wise control. Experimental results demonstrate that our simple temporal speed-up trick enables efficient, scalable, and high-quality long-form music generation. Audio samples are available at https://SqueezeComposer.github.io/.
标题:OmniCodec:具有语义声学解纠缠的低帧率通用音频编解码器
链接:https://arxiv.org/abs/2603.20638
摘要:大型语言模型(LLM)通过离散表示学习来改进音频生成。然而,大多数现有的神经编解码器专注于语音并强调重建保真度,忽略了跨不同音频域(包括语音,音乐和一般声音)的统一低帧速率建模。此外,高重建质量不一定产生语义信息表示,限制了下游生成任务的有效性。我们提出了OmniCodec,一个通用的神经音频编解码器为低帧速率量身定制。它采用了一种分层的多码本设计,通过利用预训练的理解模型的音频编码器进行语义-声学解耦,以及一种自指导策略来提高码本的利用率和重构。与Mimi编解码器相比,实验表明,OmniCodec在相同的比特率下实现了出色的性能,提供了卓越的重建质量,同时还提供了更多的语义信息表示,有利于下游生成任务。我们的模型和代码将是开源的。我们的演示页面可用。
摘要:Large Language Models (LLMs) have advanced audio generation through discrete representation learning. However, most existing neural codecs focus on speech and emphasize reconstruction fidelity, overlooking unified low frame rate modeling across diverse audio domains, including speech, music, and general sound. Moreover, high reconstruction quality does not necessarily yield semantically informative representations, limiting effectiveness in downstream generation tasks. We propose OmniCodec, a universal neural audio codec tailored for low frame rate. It adopts a hierarchical multi-codebook design with semantic-acoustic decoupling by leveraging the audio encoder of the pre-trained understanding model, along with a self-guidance strategy to improve codebook utilization and reconstruction. Compared with the Mimi codec, experiments show that OmniCodec achieves outstanding performance at the same bitrate, delivering superior reconstruction quality while also providing more semantically informative representations that benefit downstream generation tasks. Our model and code will be open-sourced. Our demo page is available.
标题:端到端多任务学习可调节关节降噪和听力损失补偿
链接:https://arxiv.org/abs/2603.20387
摘要:提出了一种多任务学习框架,用于优化单个深度神经网络(DNN),以实现联合降噪(NR)和听力损失补偿(HLC)。为每个任务定义一个不同的训练目标,DNN预测两个时频掩码。在推断期间,NR和HLC的量可以通过在组合它们之前对每个掩码取幂来独立地调整。与最近依赖于训练听觉模型仿真器来定义可微分训练目标的方法相比,我们提出了一种固有可微分的听觉模型,从而允许端到端优化。听力图作为DNN的输入提供,从而实现特定于听众的个性化,而无需再培训。结果表明,所提出的方法不仅允许单独调整NR和HLC的量,而且与优化单个训练目标相比,还提高了目标度量。它还优于分别针对NR和HLC训练的两个DNN的级联,并且与传统助听器处方相比,显示出具有竞争力的HLC性能。据我们所知,这是第一项使用听觉模型在广泛的听众配置文件中为NR和HLC训练单个DNN的研究。
摘要:A multi-task learning framework is proposed for optimizing a single deep neural network (DNN) for joint noise reduction (NR) and hearing loss compensation (HLC). A distinct training objective is defined for each task, and the DNN predicts two time-frequency masks. During inference, the amounts of NR and HLC can be adjusted independently by exponentiating each mask before combining them. In contrast to recent approaches that rely on training an auditory-model emulator to define a differentiable training objective, we propose an auditory model that is inherently differentiable, thus allowing end-to-end optimization. The audiogram is provided as an input to the DNN, thereby enabling listener-specific personalization without the need for retraining. Results show that the proposed approach not only allows adjusting the amounts of NR and HLC individually, but also improves objective metrics compared to optimizing a single training objective. It also outperforms a cascade of two DNNs that were separately trained for NR and HLC, and shows competitive HLC performance compared to a traditional hearing-aid prescription. To the best of our knowledge, this is the first study that uses an auditory model to train a single DNN for both NR and HLC across a wide range of listener profiles.
标题:TiCo:口语对话模型的时间可控训练
链接:https://arxiv.org/abs/2603.22267
摘要:我们提出了TiCo,一个简单的后训练方法,使口语对话模型(SDM)遵循时间约束的指令,并生成具有可控持续时间的响应。这种能力对于语音助手和交互式代理等真实世界的口语系统非常有价值,控制响应持续时间可以提高交互质量。然而,尽管现有模型具有很强的生成自然口头响应的能力,但它们缺乏时间意识,并且难以遵循与持续时间相关的指令(例如,“请生成一个持续约15秒的响应”)。通过对开源和商业SDM的实证评估,我们表明它们经常无法满足此类时间控制要求。TiCo通过使模型能够在生成期间通过口语时间标记(STM)(例如,<10.6 seconds>).这些标记帮助模型保持时间意识,并调整剩余内容以满足目标持续时间。TiCo简单高效:它只需要少量的数据,不需要额外的问答对,而是依靠自我生成和强化学习。实验结果表明,TiCo显着提高坚持持续时间的限制,同时保持响应质量。
摘要:We propose TiCo, a simple post-training method for enabling spoken dialogue models (SDMs) to follow time-constrained instructions and generate responses with controllable duration. This capability is valuable for real-world spoken language systems such as voice assistants and interactive agents, where controlling response duration can improve interaction quality. However, despite their strong ability to generate natural spoken responses, existing models lack time awareness and struggle to follow duration-related instructions (e.g., "Please generate a response lasting about 15 seconds"). Through an empirical evaluation of both open-source and commercial SDMs, we show that they frequently fail to satisfy such time-control requirements. TiCo addresses this limitation by enabling models to estimate elapsed speaking time during generation through Spoken Time Markers (STM) (e.g., <10.6 seconds>). These markers help the model maintain awareness of time and adjust the remaining content to meet the target duration. TiCo is simple and efficient: it requires only a small amount of data and no additional question-answer pairs, relying instead on self-generation and reinforcement learning. Experimental results show that TiCo significantly improves adherence to duration constraints while preserving response quality.
标题:TaigiSpeech:一个低资源的现实世界语音意图数据集和具有可扩展数据挖掘的初步结果
链接:https://arxiv.org/abs/2603.21478
备注:submitted to Interspeech 2026
摘要:语音技术发展迅速,服务于世界各地的不同人群。然而,由于资源有限,许多语文的代表性仍然不足。在本文中,我们介绍了\textbf{TaigiSpeech},这是一个真实世界的语音意图数据集,在台湾台语(又名台湾闽南语/闽南语),这是一个低资源和主要口语。该数据集是从老年人中收集的,包括21个说话者,总共有3 k个话语。它专为实际的意图检测场景而设计,包括医疗保健和家庭助理应用。为了解决标记数据的稀缺性,我们探索了两个数据挖掘策略,两个层次的监督:关键字匹配数据挖掘与LLM伪标签通过中间语言和视听框架,利用多模态线索与最小的文本监督。这种设计使低资源和非书面口语的可扩展数据集构造成为可能。TaigiSpeech将在CC BY 4.0许可下发布,以促进对低资源和非书面语言的广泛采用和研究。项目网站和数据集可在https://kwchang.org/taigispeech上找到。
摘要:Speech technologies have advanced rapidly and serve diverse populations worldwide. However, many languages remain underrepresented due to limited resources. In this paper, we introduce \textbf{TaigiSpeech}, a real-world speech intent dataset in Taiwanese Taigi (aka Taiwanese Hokkien/Southern Min), which is a low-resource and primarily spoken language. The dataset is collected from older adults, comprising 21 speakers with a total of 3k utterances. It is designed for practical intent detection scenarios, including healthcare and home assistant applications. To address the scarcity of labeled data, we explore two data mining strategies with two levels of supervision: keyword match data mining with LLM pseudo labeling via an intermediate language and an audio-visual framework that leverages multimodal cues with minimal textual supervision. This design enables scalable dataset construction for low-resource and unwritten spoken languages. TaigiSpeech will be released under the CC BY 4.0 license to facilitate broad adoption and research on low-resource and unwritten languages. The project website and the dataset can be found on https://kwchang.org/taigispeech.
标题:HEIX:利用混合曼巴-注意力超越二次极限来扩展原始音频理解
链接:https://arxiv.org/abs/2603.21316
备注:10 Pages, 8 Figures
摘要:音频表示学习通常评估设计选择,例如输入前端,序列主干和序列长度。我们表明,这些轴是耦合的,从一个设置的结论往往不会转移到其他。我们介绍HELIX,一个控制框架比较纯曼巴,纯注意力,和一个最小的混合动力与一个单一的注意力瓶颈。所有模型都在大约8.3M参数下进行参数匹配,以隔离建筑效果。在六个数据集上,我们发现首选的输入表示取决于主干,并且注意力会损害短的固定音频的性能,但在较长的序列长度上变得重要。在一个5分钟的说话人识别任务中,有30,000个标记,纯注意力由于内存不足而失败,而HELIX比纯Mamba缩小了11.5分的差距。
摘要:Audio representation learning typically evaluates design choices such as input frontend, sequence backbone, and sequence length in isolation. We show that these axes are coupled, and conclusions from one setting often do not transfer to others. We introduce HELIX, a controlled framework comparing pure Mamba, pure attention, and a minimal hybrid with a single attention bottleneck. All models are parameter-matched at about 8.3M parameters to isolate architectural effects. Across six datasets, we find that the preferred input representation depends on the backbone, and that attention hurts performance on short, stationary audio but becomes important at longer sequence lengths. On a 5-minute speaker identification task with 30,000 tokens, pure attention fails with out-of-memory errors, while HELIX closes an 11.5-point gap over pure Mamba.
标题:ALICE:大型音频语言模型上下文学习能力的多方面评估框架
链接:https://arxiv.org/abs/2603.20433
备注:Submitted to Interspeech 2026
摘要:虽然大型音频语言模型(LALM)已被证明表现出退化的推理跟随能力,但它们从音频条件下的上下文示例中推断任务模式的能力尚未得到研究。为了解决这个问题,我们提出了ALICE,一个三阶段的框架,逐步减少文本指导,系统地评估LALM在音频条件下的上下文学习能力。在两个输出约束类别下评估四个音频理解任务中的六个LALM,我们发现所有阶段和LALM之间存在一致的不对称性:上下文演示可靠地提高了格式合规性,但未能提高,而且往往会降低核心任务的性能。这表明LALM可以从演示中收集表面级别的格式化模式,但可能难以利用跨模态语义基础来可靠地从音频条件示例中推断任务目标,突出了当前跨模态整合的潜在局限性。
摘要:While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples under audio conditioning remains unstudied. To address this gap, we present ALICE, a three-stage framework that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability under audio conditioning. Evaluating six LALMs across four audio understanding tasks under two output constraint categories, we uncover a consistent asymmetry across all stages and LALMs: in-context demonstrations reliably improve format compliance but fail to improve, and often degrade, the core task performance. This suggests that LALMs can glean surface-level formatting patterns from demonstrations but may struggle to leverage cross-modal semantic grounding to reliably infer task objectives from audio-conditioned examples, highlighting potential limitations in current cross-modal integration.
标题:Abjad-Kids:小学教育阿拉伯语语音分类数据集
链接:https://arxiv.org/abs/2603.20255
摘要:近年来,基于语音的人工智能教育应用引起了人们的极大兴趣,尤其是对儿童。然而,儿童语音研究仍然有限,由于缺乏公开的数据集,特别是低资源的语言,如Arabic.This本文介绍了Abjad-Kids,阿拉伯语语音数据集设计的幼儿园和小学教育,重点是字母,数字和颜色的基本学习。该数据集包括从3 - 12岁儿童中收集的46397个音频样本,覆盖141个班级。所有样品均按照受控质量标准记录,以确保持续时间、采样率和格式的一致性。为了解决阿拉伯音素之间的高类内相似性和每个类的有限样本问题,我们提出了一种基于CNN-LSTM架构的分层音频分类。我们提出的方法将字母识别分解为两个阶段的过程:一个初始的分组分类模型,然后为每组专门的分类器。这两种策略:静态语言为基础的分组和动态聚类为基础的分组,进行了评估。实验结果表明,静态的基于语言的分组取得了优异的性能。传统机器学习与深度学习方法之间的比较突出了CNN-LSTM模型与数据增强相结合的有效性。尽管取得了令人鼓舞的结果,但我们的大多数实验都表明存在过拟合的挑战,这可能是由于样本数量有限,即使在数据增强和模型正则化之后。因此,今后的工作可能侧重于收集更多的数据,以解决这一问题。Abjad-Kids将向公众开放。我们希望Abjad-Kids能够丰富语音数据库中的儿童语音表示,并为未来的阿拉伯语儿童语音分类研究提供一个很好的资源。
摘要:Speech-based AI educational applications have gained significant interest in recent years, particularly for children. However, children speech research remains limited due to the lack of publicly available datasets, especially for low-resource languages such as Arabic.This paper presents Abjad-Kids, an Arabic speech dataset designed for kindergarten and primary education, focusing on fundamental learning of alphabets, numbers, and colors. The dataset consists of 46397 audio samples collected from children aged 3 - 12 years, covering 141 classes. All samples were recorded under controlled specifications to ensure consistency in duration, sampling rate, and format. To address high intra-class similarity among Arabic phonemes and the limited samples per class, we propose a hierarchical audio classification based on CNN-LSTM architectures. Our proposed methodology decomposes alphabet recognition into a two-stage process: an initial grouping classification model followed by specialized classifiers for each group. Both strategies: static linguistic-based grouping and dynamic clustering-based grouping, were evaluated. Experimental results demonstrate that static linguistic-based grouping achieves superior performance. Comparisons between traditional machine learning with deep learning approaches, highlight the effectiveness of CNN-LSTM models combined with data augmentation. Despite achieving promising results, most of our experiments indicate a challenge with overfitting, which is likely due to the limited number of samples, even after data augmentation and model regularization. Thus, future work may focus on collecting additional data to address this issue. Abjad-Kids will be publicly available. We hope that Abjad-Kids enrich children representation in speech dataset, and be a good resource for future research in Arabic speech classification for kids.
标题:LL-SDR:通过离散表示实现低延迟语音增强
链接:https://arxiv.org/abs/2603.20242
备注:5 pages, 1 figure
摘要:许多语音增强(SE)方法依赖于连续表示。最近,已经探索了离散音频令牌以实现SE的自回归生成。然而,目前尚不清楚离散化本身是否始终提高SE性能。在本文中,我们介绍了LL-SDR,一个基于令牌的语音增强框架,明确利用离散化更好地分离语音和噪声。我们的第一个贡献是方差排序的残差矢量量化器(VO-RVQ),旨在解开语音和噪声分布在令牌化。其次,我们提出了一个潜在的空间嵌入,以更好地调整增强嵌入与语义嵌入。实验表明,LL-SDR优于连续基线,并匹配基于自回归令牌的方法的性能,同时在混响和非混响噪声环境中实现轻量级,低延迟的语音增强。演示和源代码可在我们的项目网站上获得。
摘要:Many speech enhancement (SE) methods rely on continuous representations. Recently, discrete audio tokens have been explored to enable autoregressive generation for SE. However, it remains unclear whether discretization itself consistently improves SE performance. In this paper, we introduce LL-SDR, a token-based speech enhancement framework that explicitly leverages discretization to better separate speech and noise. Our first contribution is a Variance-Ordered Residual Vector Quantizer (VO-RVQ), designed to disentangle speech and noise distributions during tokenization. Second, we propose a latent-space discriminator to better align enhanced embeddings with semantic embeddings. Experiments show that LL-SDR outperforms continuous baselines and matches the performance of autoregressive token-based approaches, while enabling lightweight, low-latency speech enhancement in both reverberant and non-reverberant noisy environments. Demos and source code are available at our project websites.
机器翻译由腾讯交互翻译提供,仅供参考
