微信公众号:arXiv_Daily
cs.SD语音
【1】RAMoEA-QA: Hierarchical Specialization for Robust Respiratory Audio Question Answering
标题:RAMoEA-QA:稳健呼吸音频问题解答的分层专业化
链接:https://arxiv.org/abs/2603.06542
摘要:会话生成AI正在迅速进入医疗保健领域,通用模型必须集成异构的患者信号,并支持多种交互风格,同时产生有临床意义的输出。在呼吸护理中,非侵入性音频(例如通过移动麦克风捕获的录音)可以实现可扩展的筛选和纵向监测,但异质性挑战特别严重:记录在设备,环境和采集协议中差异很大,并且问题跨越多种意图和问题格式。现有的生物医学音频语言QA系统通常是单一的,没有任何专门的机制来处理不同的呼吸语料库和查询意图。它们也只在有限的环境中得到验证,因此尚不清楚它们在处理现实环境中遇到的变化时的可靠性。 为了解决这些局限性,我们引入RAMoEA-QA,一个分层路由生成模型的呼吸音频问答,统一了多种问题类型,并支持在一个单一的多模态系统中的离散和连续的目标。RAMoEA-QA应用两阶段条件专业化:音频专家混合将每个录音路由到合适的预训练音频编码器,语言混合适配器在共享冻结LLM上选择LoRA适配器以匹配查询意图和答案格式。通过对每个示例的声学表示和生成行为进行专门化,RAMoEA-QA始终以最小的参数开销优于强基线和路由消融,将域内测试准确度提高到0.72(对于最先进的基线为0.61和0.67),并在域,模态和任务转移下表现出最强的诊断泛化能力。
摘要:Conversational generative AI is rapidly entering healthcare, where general-purpose models must integrate heterogeneous patient signals and support diverse interaction styles while producing clinically meaningful outputs. In respiratory care, non-invasive audio, such as recordings captured via mobile microphones, enables scalable screening and longitudinal monitoring, but the heterogeneity challenge is particularly acute: recordings vary widely across devices, environments, and acquisition protocols, and questions span multiple intents and question formats. Existing biomedical audio-language QA systems are typically monolithic, without any specialization mechanisms for tackling diverse respiratory corpora and query intents. They are also only validated in limited settings, leaving it unclear how reliably they handle the shifts encountered in real-world settings. To address these limitations, we introduce RAMoEA-QA, a hierarchically routed generative model for respiratory audio question answering that unifies multiple question types and supports both discrete and continuous targets within a single multimodal system. RAMoEA-QA applies two-stage conditional specialization: an Audio Mixture-of-Experts routes each recording to a suitable pre-trained audio encoder, and a Language Mixture-of-Adapters selects a LoRA adapter on a shared frozen LLM to match the query intent and answer format. By specializing both acoustic representations and generation behaviour per example, RAMoEA-QA consistently outperforms strong baselines and routing ablations with minimal parameter overhead, improving in-domain test accuracy to 0.72 (vs. 0.61 and 0.67 for state-of-the-art baselines) and exhibiting the strongest generalization for diagnosis under domain, modality, and task shifts.
【2】Prosodic Boundary-Aware Streaming Generation for LLM-Based TTS with Streaming Text Input
标题:具有流文本输入的基于LLM的TTC的韵律边界感知流生成
链接:https://arxiv.org/abs/2603.06444
摘要:流TTS接收流文本是必不可少的交互式系统,但这个计划面临着两个主要挑战:不自然的韵律,由于缺少前瞻性和长形式的崩溃,由于无界的上下文。我们提出了一种韵律边界感知的后训练策略,使用弱时间对齐数据适应预训练的基于LLM的TTS模型。具体地,该模型适于在提供有限的未来文本时学习在指定的内容边界处的早期停止。在推理过程中,滑动窗口提示符会将先前的文本和语音标记向前推进,确保有界上下文和无缝连接。评估表明,我们的方法优于CosyVoice-Style交织基线在短和长形式的情况下。特别是在长文本合成中,它实现了66.2%的单词错误率绝对降低(从71.0%到4.8%),并将说话者和情感相似性相对提高了16.1%和1.5%,为带有增量文本的流TTS提供了一个强大的解决方案。
摘要:Streaming TTS that receives streaming text is essential for interactive systems, yet this scheme faces two major challenges: unnatural prosody due to missing lookahead and long-form collapse due to unbounded context. We propose a prosodic-boundary-aware post-training strategy, adapting a pretrained LLM-based TTS model using weakly time-aligned data. Specifically, the model is adapted to learn early stopping at specified content boundaries when provided with limited future text. During inference, a sliding-window prompt carries forward previous text and speech tokens, ensuring bounded context and seamless concatenation. Evaluations show our method outperforms CosyVoice-Style interleaved baseline in both short and long-form scenarios. In long-text synthesis, especially, it achieves a 66.2% absolute reduction in word error rate (from 71.0% to 4.8%) and increases speaker and emotion similarity by 16.1% and 1.5% relatively, offering a robust solution for streaming TTS with incremental text.
【3】Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding
标题:Whisper-CD:使用多负对比解码的准确长形式语音识别
链接:https://arxiv.org/abs/2603.06193
备注:Submitted to Interspeech 2026
摘要:使用大型编码器-解码器模型(如Whisper)的长形式语音识别通常会出现幻觉,重复循环和内容遗漏。当先前片段的转录被用作解码上下文时,这些错误可以累积并被进一步放大。我们提出了Whisper-CD,一个免训练的对比解码框架,将干净的音频logits与从三个声学激励扰动计算的负logits进行对比:高斯噪声注入,沉默信号和音频时间偏移。我们通过对数和exp算子聚合这些否定,为逐个令牌解码建立统一的多否定目标。在五个英语长格式基准测试中,Whisper-CD在CORAAL上将WER降低了24.3pp,并且比beam search的令牌生成吞吐量快了48%。由于Whisper-CD纯粹在推理时间运行,因此它可以作为替代品应用于已经部署的Whisper系统,而无需重新培训。
摘要:Long-form speech recognition with large encoder-decoder models such as Whisper often exhibit hallucinations, repetition loops, and content omissions. These errors can accumulate and be further amplified when the previous segment's transcription is used as decoding context. We propose Whisper-CD, a training-free contrastive decoding framework that contrasts clean-audio logits against negative logits computed from three acoustically motivated perturbations: Gaussian noise injection, silence signal, and audio temporal shift. We aggregate these negatives via the log-sum-exp operator, building a unified multi-negative objective for token-by-token decoding. Across five English long-form benchmarks, Whisper-CD reduces WER by up to 24.3pp on CORAAL and shows 48% faster token generation throughput than beam search. Because Whisper-CD operates purely at inference time, it can be applied as a drop-in replacement to already-deployed Whisper systems without retraining.
【4】Do Compact SSL Backbones Matter for Audio Deepfake Detection? A Controlled Study with RAPTOR
标题:紧凑的SSL主干对音频Deepfake检测有影响吗?RAPTOR对照研究
链接:https://arxiv.org/abs/2603.06164
备注:Submitted to Interspeech 2026, 4 pages, 2 figures
摘要:自监督学习(SSL)是现代音频深度伪造检测的基础,但大多数先前的工作都集中在一个大型的wav 2 vec 2-XLSR主干上,而紧凑型则有待研究。我们提出了RAPTOR,表示感知成对门控Transformer域外识别紧凑的SSL骨干的控制研究,从一个统一的成对门控融合检测器内的HuBERT和WavLM,在14个跨域基准评估。我们发现,多语言HuBERT预训练是跨域鲁棒性的主要驱动力,使1亿模型能够匹配更大的商业系统。除了EER,我们引入了一个测试时间增强协议与扰动为基础的任意不确定性暴露校准差异看不见的标准指标:WavLM变量表现出过度自信的失调扰动下,而迭代mHuBERT保持稳定。这些发现表明,SSL预训练轨迹,而不是模型规模,驱动可靠的音频deepfake检测。
摘要:Self-supervised learning (SSL) underpins modern audio deepfake detection, yet most prior work centers on a single large wav2vec2-XLSR backbone, leaving compact under studied. We present RAPTOR, Representation Aware Pairwise-gated Transformer for Out-of-domain Recognition a controlled study of compact SSL backbones from the HuBERT and WavLM within a unified pairwise-gated fusion detector, evaluated across 14 cross-domain benchmarks. We show that multilingual HuBERT pre-training is the primary driver of cross-domain robustness, enabling 100M models to match larger and commercial systems. Beyond EER, we introduce a test-time augmentation protocol with perturbation-based aleatoric uncertainty to expose calibration differences invisible to standard metrics: WavLM variants exhibit overconfident miscalibration under perturbation, whereas iterative mHuBERT remains stable. These findings indicate that SSL pre-training trajectory, not model scale, drives reliable audio deepfake detection.
【5】TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head Generation
标题:TempoSyncDiff:为低延迟音频驱动的会说话的头部生成提取时间一致的扩散
链接:https://arxiv.org/abs/2603.06057
摘要:扩散模型最近先进的逼真的人类合成,虽然实际的说话头生成(THG)仍然受到高推理延迟,时间不稳定性,如闪烁和身份漂移,以及不完美的视听对齐在具有挑战性的语音条件下。本文介绍了TempoSyncDiff,一个参考条件的潜在扩散框架,探讨了高效的音频驱动的说话头生成的几步推理。该方法采用了教师-学生蒸馏配方,其中一个标准的噪声预测目标训练的扩散教师指导一个轻量级的学生去噪器,能够以显着更少的推理步骤来提高生成稳定性。该框架采用身份锚定和时间正则化设计,以减轻身份漂移和帧到帧的闪烁在合成过程中,而基于视位的音频调节提供粗略的嘴唇运动控制。LRS 3数据集上的实验报告了与VAE重建和初步延迟表征相关的降噪阶段组件级指标,包括仅CPU和边缘计算测量以及边缘部署的可行性估计。结果表明,蒸馏扩散模型可以保留一个更强大的教师的重建行为,同时使大大降低延迟推理。这项研究被定位为第一步,实际的扩散为基础的说话头生成下的约束计算设置。GitHub:https://mazumdarsoumya.github.io/TempoSyncDiff
摘要:Diffusion models have recently advanced photorealistic human synthesis, although practical talking-head generation (THG) remains constrained by high inference latency, temporal instability such as flicker and identity drift, and imperfect audio-visual alignment under challenging speech conditions. This paper introduces TempoSyncDiff, a reference-conditioned latent diffusion framework that explores few-step inference for efficient audio-driven talking-head generation. The approach adopts a teacher-student distillation formulation in which a diffusion teacher trained with a standard noise prediction objective guides a lightweight student denoiser capable of operating with significantly fewer inference steps to improve generation stability. The framework incorporates identity anchoring and temporal regularization designed to mitigate identity drift and frame-to-frame flicker during synthesis, while viseme-based audio conditioning provides coarse lip motion control. Experiments on the LRS3 dataset report denoising-stage component-level metrics relative to VAE reconstructions and preliminary latency characterization, including CPU-only and edge computing measurements and feasibility estimates for edge deployment. The results suggest that distilled diffusion models can retain much of the reconstruction behaviour of a stronger teacher while enabling substantially lower latency inference. The study is positioned as an initial step toward practical diffusion-based talking-head generation under constrained computational settings. GitHub: https://mazumdarsoumya.github.io/TempoSyncDiff
【6】How Well Do Current Speech Deepfake Detection Methods Generalize to the Real World?
标题:当前的语音Deepfake检测方法如何推广到现实世界?
链接:https://arxiv.org/abs/2603.05852
备注:Submitted to Interspeech 2026
摘要:语音合成和语音转换的最新进展极大地提高了所生成音频的自然度和真实性。与此同时,社交媒体平台上不断发展的编码、压缩和传输机制进一步模糊了deepfake伪影。这些因素使真实环境中的可靠检测复杂化,强调了对代表性评估基准的需求。为此,我们引入了ML-ITW(Multilingual In-The-Wild),这是一个多语言数据集,涵盖14种语言,7个主要平台和180个公众人物,总计28.39小时的音频。我们评估了三种检测范式:端到端神经模型、基于自监督特征(SSL)的方法和音频大语言模型(Audio LLM)。实验结果显示,在不同的语言和现实世界的声学条件下,显着的性能下降,突出了现有的检测器在实际场景中的泛化能力有限。ML-ITW数据集是公开的。
摘要:Recent advances in speech synthesis and voice conversion have greatly improved the naturalness and authenticity of generated audio. Meanwhile, evolving encoding, compression, and transmission mechanisms on social media platforms further obscure deepfake artifacts. These factors complicate reliable detection in real-world environments, underscoring the need for representative evaluation benchmarks. To this end, we introduce ML-ITW (Multilingual In-The-Wild), a multilingual dataset covering 14 languages, seven major platforms, and 180 public figures, totaling 28.39 hours of audio. We evaluate three detection paradigms: end-to-end neural models, self-supervised feature-based (SSL) methods, and audio large language models (Audio LLMs). Experimental results reveal significant performance degradation across diverse languages and real-world acoustic conditions, highlighting the limited generalization ability of existing detectors in practical scenarios. The ML-ITW dataset is publicly available.
【7】Which Data Matter? Embedding-Based Data Selection for Speech Recognition
标题:哪些数据重要?基于嵌入的语音识别数据选择
链接:https://arxiv.org/abs/2603.05819
摘要:现代ASR系统通常在跨越多个域的大规模伪标记的野外数据上进行训练。虽然这种异构数据有利于为广泛部署而设计的通用模型,但它们对针对特定领域的专家模型提出了挑战:专家模型缺乏从所有可用数据中学习的能力,并且必须更加关注解决训练和测试条件之间的不匹配。在这项工作中,我们研究了有针对性的数据选择作为解决这些挑战的策略,从10万小时的野外训练数据中选择相关子集,以优化目标域的性能。我们使用嵌入来表示语音样本,这些嵌入捕获互补特征-说话者属性,语音内容和语义-并分析在执行数据选择时沿这些轴的相关性和多样性如何影响下游ASR性能。我们对基于CTC的Conformer模型的实验表明,在战略选择的5%子集上进行训练,可以超过在完整数据集上训练的模型的性能,目标域的相对WER减少高达36.8%。
摘要:Modern ASR systems are typically trained on large-scale pseudo-labeled, in-the-wild data spanning multiple domains. While such heterogeneous data benefit generalist models designed for broad deployment, they pose challenges for specialist models targeting specific domains: specialist models lack the capacity to learn from all available data, and one must pay closer attention to addressing the mismatch between training and test conditions. In this work, we study targeted data selection as a strategy to address these challenges, selecting relevant subsets from 100k hours of in-the-wild training data to optimize performance on target domains. We represent speech samples using embeddings that capture complementary characteristic--speaker attributes, phonetic content, and semantic meaning--and analyze how relevance and diversity along these axes when performing data selection affect downstream ASR performance. Our experiments with CTC-based Conformer models show that training on a strategically selected 5% subset can exceed the performance of models trained on the full dataset by up to 36.8% relative WER reduction on target domains.
【8】Koopman Regularized Deep Speech Disentanglement for Speaker Verification
标题:用于说话人验证的Koopman正规化深度语音解纠缠
链接:https://arxiv.org/abs/2603.05577
备注:This work has been submitted to the IEEE for possible publication
摘要:人类语音既包含语言内容,又包含说话人相关特征,这使得说话人确认成为身份关键应用中的关键技术。现代深度学习说话人验证系统旨在学习对语义内容和环境噪声等干扰因素不变的说话人表示。然而,许多现有的方法依赖于标记数据,文本监督或大型预训练模型作为特征提取器,限制了可扩展性和实际部署,引发了可持续性问题。我们提出了Deep Koopman Speech Disentanglement Autoencoder(DKSD-AE),这是一种结构化的自动编码器,它将一种新颖的多步Koopman算子学习模块与实例归一化相结合,以解开扬声器和内容动态。跨多个数据集的定量实验表明,与最先进的基线相比,DKSD-AE实现了改进或有竞争力的说话人验证性能,同时保持了高内容EER,证实了有效的解纠缠。这些结果是在没有文本监督的情况下用更少的参数获得的。此外,性能保持稳定的评估规模增加,突出表示的鲁棒性和泛化。我们的研究结果表明,基于Koopman的时间建模,结合实例规范化,为以说话者为中心的表示学习提供了一个有效的和原则性的解决方案。
摘要:Human speech contains both linguistic content and speaker dependent characteristics making speaker verification a key technology in identity critical applications. Modern deep learning speaker verification systems aim to learn speaker representations that are invariant to semantic content and nuisance factors such as ambient noise. However, many existing approaches depend on labelled data, textual supervision or large pretrained models as feature extractors, limiting scalability and practical deployment, raising sustainability concerns. We propose Deep Koopman Speech Disentanglement Autoencoder (DKSD-AE), a structured autoencoder that combines a novel multi-step Koopman operator learning module with instance normalization to disentangle speaker and content dynamics. Quantitative experiments across multiple datasets demonstrate that DKSD-AE achieves improved or competitive speaker verification performance compared to state-of-the-art baselines while maintaining high content EER, confirming effective disentanglement. These results are obtained with substantially fewer parameters and without textual supervision. Moreover, performance remains stable under increased evaluation scale, highlighting representation robustness and generalization. Our findings suggest that Koopman-based temporal modelling, when combined with instance normalization, provides an efficient and principled solution for speaker-focused representation learning.
【9】Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
标题:Omni-C:将异类模式压缩到单个密集编码器中
链接:https://arxiv.org/abs/2603.05528
摘要:最近的多模态系统通常依赖于单独的专家模态编码器,这导致线性缩放的复杂性和计算开销与添加的模态。虽然统一的Omni模型通过具有专门专家和路由的专家混合(MoE)架构来解决这个问题,但它们仍然会增加参数计数并引入路由开销。在本文中,我们提出了Omni-C(Omni-Compress),这是一种基于transformer的单一密集编码器,它通过对大规模未对齐数据进行单峰对比预训练来学习跨异构模式(图像,音频和文本)的竞争性共享表示。通过最大化主干中的参数共享和使用轻量的模态特定投影头,Omni-C有效地减轻了模态间冲突,而无需MoE、配对监督或路由。该设计通过顺序模态处理和低内存推理支持在内存受限系统上的高效部署,从而消除了对并行专家加载或专用硬件的需求。实验表明,Omni-C在单峰和跨模型任务中实现了与专家模型相当的性能,对音频和文本的适度zero-shot降级在很大程度上通过轻量级线性探测或参数有效微调来恢复。与多编码器基线相比,统一架构大大减少了推理内存的使用,从而提高了高效和可扩展的多模式学习。
摘要:Recent multimodal systems often rely on separate expert modality encoders which cause linearly scaling complexity and computational overhead with added modalities. While unified Omni-models address this via Mixture-of-Expert (MoE) architectures with specialized experts and routing, they still inflate parameter counts and introduce routing overhead. In this paper, we propose Omni-C (Omni-Compress), a single dense Transformer-based encoder that learns competitive shared representations across heterogeneous modalities--images, audio, and text--through unimodal contrastive pretraining on large-scale unaligned data. By maximizing parameter sharing in the backbone and using lightweight modality-specific projection heads, Omni-C effectively mitigates inter-modality conflicts without requiring MoE, paired supervision, or routing. This design supports efficient deployment on memory-constrained systems via sequential modality processing and low-memory inference, eliminating the need for parallel expert loading or specialized hardware. Experiments show Omni-C achieves performance comparable to expert models in unimodal and cross-model tasks, with modest zero-shot degradation on audio and text that is largely recovered through lightweight linear probing or parameter efficient fine-tuning. The unified architecture substantially reduces inference memory usage compared to multi-encoder baselines, advancing efficient and scalable multimodal learning.
【10】Continual Adaptation for Pacific Indigenous Speech Recognition
标题:持续适应太平洋原住民语音识别
链接:https://arxiv.org/abs/2603.06310
备注:Submitted to Interspeech
摘要:语音基础模型与低资源的太平洋土著语言斗争,因为严重的数据稀缺。此外,完全微调有可能导致灾难性的遗忘。为了解决这一差距,我们提出了一个实证研究,使模型适应现实世界的太平洋数据集。我们研究了数据量和语言特征如何影响适应成功。具体来说,我们评估的策略包括全微调和低秩自适应(LoRA)。此外,我们分析了一个持续学习的框架,顺序获得多种语言。我们证明,适应这些遥远的语言会导致严重的内部代表性漂移。因此,这些模型面临着严格的塑性和稳定性的困境。虽然LoRA最初适应得很好,但它在顺序学习过程中会遭受灾难性的遗忘。最终,这项研究强调了迫切需要针对代表性不足的语言制定强有力的适应策略。
摘要:Speech foundation models struggle with low-resource Pacific Indigenous languages because of severe data scarcity. Furthermore, full fine-tuning risks catastrophic forgetting. To address this gap, we present an empirical study adapting models to real-world Pacific datasets. We investigate how data volume and linguistic features affect adaptation success. Specifically, we evaluate strategies including Full Fine-Tuning and Low-Rank Adaptation (LoRA). Additionally, we analyze a continual learning framework for sequentially acquiring multiple languages. We demonstrate that adapting to these distant languages causes severe internal representational drift. Consequently, these models face a strict plasticity and stability dilemma. While LoRA adapts well initially, it suffers from catastrophic forgetting during sequential learning. Ultimately, this study highlights the urgent need for robust adaptation strategies tailored to underrepresented languages.
【1】Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
标题:医生还是病人?协同扩张和ASB用于代码切换Hinglish医学疾病提取
链接:https://arxiv.org/abs/2603.06373
备注:Submitted for review at Interspeech 2026
摘要:由于快速的话轮转换和高度重叠的语音,从代码切换的临床口语对话中提取患者的医疗状况是具有挑战性的。我们提出了一个强大的系统上评估的位移-M数据集的真实世界的印度式英语医疗会话。我们提出了一种端到端神经日记与向量聚类方法(EEND-VC)来准确解决医患对话(DoPaCo)中的密集和说话者重叠问题。对于转录,我们通过特定于域的微调,梵文脚本规范化和对话级LLM错误校正来适应Qwen 3 ASR模型,实现了18.59%的tcpWER。我们基准开放和专有LLM的医疗条件提取,比较我们的基于文本的级联系统对多模式端到端(E2 E)音频框架。虽然专有的E2 E模型设定了性能上限,但我们的开放级联架构具有很强的竞争力,因为它在25名参与者中获得了第一名。所有实现都是公开发布的。
摘要:Extracting patient medical conditions from code-switched clinical spoken dialogues is challenging due to rapid turn-taking and highly overlapped speech. We present a robust system evaluated on the DISPLACE-M dataset of real-world Hinglish medical conversations. We propose an End-to-End Neural Diarization with Vector Clustering approach (EEND-VC) to accurately resolve dense and speaker overlaps in Doctor-Patient Conversations (DoPaCo). For transcription, we adapt a Qwen3 ASR model via domain-specific fine-tuning, Devanagari script normalization, and dialogue-level LLM error correction, achieving an 18.59% tcpWER. We benchmark open and proprietary LLMs on medical condition extraction, comparing our text-based cascade system against a multimodal End-to-End (E2E) audio framework. While proprietary E2E models set the performance ceiling, our open cascaded architecture is highly competitive, as it achieved first place out of 25 participants in the DISPLACE-M challenge. All implementations are publicly released.
【2】Cross-linguistic Prosodic Analysis of Autistic and Non-autistic Child Speech in Finnish, French and Slovak
标题:芬兰语、法语和斯洛伐克语自闭症和非自闭症儿童言语的跨语言韵律分析
链接:https://arxiv.org/abs/2603.06332
备注:Accepted to Speech Prosody 2026
摘要:孤独症患者的韵律差异已经得到了很好的证明,但跨语言的证据仍然有限。本研究探讨韵律在自闭症跨芬兰语,法语和斯洛伐克语的多语种语料库。从5,000多个间歇单位中提取了88个声学特征,并通过主成分分析(PCA)对数据进行了简化,并使用线性混合效应模型(LIFE)进行了分析。 跨语言,自闭症扬声器表现出增加的一般强度变化和更清晰,更少的呼吸声音质量(更高的谐波噪声比和α比),以及降低的时间强度动态和较低的中央f0。单语分析揭示了语言特定的细微差别:斯洛伐克语的结果与跨语言的f0模式一致,但在语音质量上存在差异,而芬兰语的结果反映了更广泛的语音质量结果。 这些结果强调,除了传统的音高测量外,还包括语音质量和强度动态,以研究自闭症的可能的语言独立标记。研究结果挑战了基于缺陷的模型,而是提出了一个复杂的,声学上不同的韵律分布在不同的语言。
摘要:Prosodic differences in autism are well-documented, but cross-linguistic evidence remains limited. This study investigates prosody in autism across a multilingual corpus of Finnish, French, and Slovak speakers. 88 acoustic features from over 5,000 inter-pausal units were extracted, and data were reduced via Principal Component Analysis (PCA) and analyzed using Linear Mixed-Effects Models (LMMs). Cross-linguistically, autistic speakers exhibited increased general intensity variability and a clearer, less breathy voice quality (higher Harmonics-to-Noise Ratio and alpha ratio), alongside reduced temporal intensity dynamics and lower central f0. Monolingual analyses revealed language-specific nuances: Slovak results aligned with cross-linguistic f0 patterns but diverged on voice quality, while Finnish results mirrored the broader voice quality findings. These results emphasize including voice quality and intensity dynamics in the study of possible language-independent markers of autism, alongside traditional pitch measures. The findings challenge deficiency-based models, suggesting instead a complex, acoustically distinct prosodic profile across languages.
【3】Classification of Autistic and Non-Autistic Children's Speech: A Cross-Linguistic Study in Finnish, French, and Slovak
标题:自闭症和非自闭症儿童言语的分类:芬兰语、法语和斯洛伐克语的跨语言研究
链接:https://arxiv.org/abs/2603.06327
备注:Accepted to Speech Prosody 2026
摘要:我们提出了一个跨语言的研究在自闭症和非自闭症儿童讲芬兰语,法语和斯洛伐克语的讲话。我们将监督分类与语言内和跨语料库迁移实验相结合,以评估语言内和跨语言的分类性能,并探索哪些声学线索是特定于语言的,而不是一般的语言。使用大量的声学韵律特征,我们实现扬声器级分类基准作为一种分析工具,而不是寻求最先进的性能。 内语言模型,与扬声器级交叉验证评估,产生了异质性的结果。芬兰模型表现最好(准确度0.84,F1 0.88),其次是斯洛伐克(准确度0.63,F1 0.68)和法国(准确度0.68,F1 0.56)。然后,我们测试了跨语言泛化。在所有合并语料库上训练的模型达到了0.61的总体准确度和0.68的F1。Leave-one-corpus-out实验,测试迁移到一种看不见的语言,在斯洛伐克语(F1 0.70)和芬兰语(F1 0.78)测试时显示出中等程度的成功,但迁移到法语(F1 0.42)的效果很差。跨语言的听觉重要性分析强调了自闭症的部分共享,但不是完全语言不变的声学标记。 这些研究结果表明,一些自闭症相关的语音线索概括了不同类型的语言,但强大的跨语言分类器可能需要语言感知建模和更均匀的记录条件。
摘要:We present a cross-linguistic study of speech in autistic and non-autistic children speaking Finnish, French, and Slovak. We combine supervised classification with within-language and cross-corpus transfer experiments to evaluate classification performance within and across languages and to probe which acoustic cues are language-specific versus language-general. Using a large set of acoustic-prosodic features, we implement speaker-level classification benchmarks as an analytical tool rather than to seek state-of-the-art performance. Within-language models, evaluated with speaker-level cross-validation, yielded heterogeneous results. The Finnish model performed best (Accuracy 0.84, F1 0.88), followed by Slovak (Accuracy 0.63, F1 0.68) and French (Accuracy 0.68, F1 0.56). We then tested cross-language generalization. A model trained on all pooled corpora reached an overall Accuracy of 0.61 and F1 0.68. Leave-one-corpus-out experiments, which test transfer to an unseen language, showed moderate success when testing on Slovak (F1 0.70) and Finnish (F1 0.78), but poor transfer to French (F1 0.42). Feature-importance analyses across languages highlighted partially shared, but not fully language-invariant, acoustic markers of autism. These findings suggest that some autism-related speech cues generalize across typologically distinct languages, but robust cross-linguistic classifiers will likely require language-aware modeling and more homogeneous recording conditions.
【4】Continual Adaptation for Pacific Indigenous Speech Recognition
标题:持续适应太平洋原住民语音识别
链接:https://arxiv.org/abs/2603.06310
备注:Submitted to Interspeech
摘要:语音基础模型与低资源的太平洋土著语言斗争,因为严重的数据稀缺。此外,完全微调有可能导致灾难性的遗忘。为了解决这一差距,我们提出了一个实证研究,使模型适应现实世界的太平洋数据集。我们研究了数据量和语言特征如何影响适应成功。具体来说,我们评估的策略包括全微调和低秩自适应(LoRA)。此外,我们分析了一个持续学习的框架,顺序获得多种语言。我们证明,适应这些遥远的语言会导致严重的内部代表性漂移。因此,这些模型面临着严格的塑性和稳定性的困境。虽然LoRA最初适应得很好,但它在顺序学习过程中会遭受灾难性的遗忘。最终,这项研究强调了迫切需要针对代表性不足的语言制定强有力的适应策略。
摘要:Speech foundation models struggle with low-resource Pacific Indigenous languages because of severe data scarcity. Furthermore, full fine-tuning risks catastrophic forgetting. To address this gap, we present an empirical study adapting models to real-world Pacific datasets. We investigate how data volume and linguistic features affect adaptation success. Specifically, we evaluate strategies including Full Fine-Tuning and Low-Rank Adaptation (LoRA). Additionally, we analyze a continual learning framework for sequentially acquiring multiple languages. We demonstrate that adapting to these distant languages causes severe internal representational drift. Consequently, these models face a strict plasticity and stability dilemma. While LoRA adapts well initially, it suffers from catastrophic forgetting during sequential learning. Ultimately, this study highlights the urgent need for robust adaptation strategies tailored to underrepresented languages.
【5】StreamVoiceAnon+: Emotion-Preserving Streaming Speaker Anonymization via Frame-Level Acoustic Distillation
标题:StreamPoker Anon+:通过帧级声学蒸馏保持流媒体扬声器音频化
链接:https://arxiv.org/abs/2603.06079
摘要:我们解决了在流媒体扬声器匿名化(SA)中保存情感内容的挑战。为音频延续训练的神经音频编解码器语言模型倾向于降低源情感:内容令牌丢弃情感信息,并且模型默认为占主导地位的声学模式,而不是保留非语言属性。我们提出了监督微调与中性情感话语对从同一个扬声器,结合帧级情感蒸馏声学令牌隐藏状态。所有修改都仅限于微调,在4个GPU上花费不到2小时,并增加零推理延迟开销,同时保持具有竞争力的180 ms流延迟。在VoicePrivacy 2024协议上,我们的方法实现了49.2%的UAR(情感保留)和5.77%的WER(可理解性),比基线(39.7%->49.2%)和情感提示变体(44.6% UAR)的相对UAR提高了+24%,同时保持了强隐私(EER 49.0%)。演示和代码可用:https://anonymous3842031239.github.io/
摘要:We address the challenge of preserving emotional content in streaming speaker anonymization (SA). Neural audio codec language models trained for audio continuation tend to degrade source emotion: content tokens discard emotional information, and the model defaults to dominant acoustic patterns rather than preserving paralinguistic attributes. We propose supervised finetuning with neutral-emotion utterance pairs from the same speaker, combined with frame-level emotion distillation on acoustic token hidden states. All modifications are confined to finetuning, which takes less than 2 hours on 4 GPUs and adds zero inference latency overhead, while maintaining a competitive 180ms streaming latency. On the VoicePrivacy 2024 protocol, our approach achieves a 49.2% UAR (emotion preservation) with 5.77% WER (intelligibility), a +24% relative UAR improvement over the baseline (39.7%->49.2%) and +10% over the emotion-prompt variant (44.6% UAR), while maintaining strong privacy (EER 49.0%). Demo and code are available: https://anonymous3842031239.github.io/
【6】Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech
标题:Accent中和Zero-Shot文本到语音的激活转向
链接:https://arxiv.org/abs/2603.05977
备注:Submitted to Interspeech 2026
摘要:Zero-shot文本到语音(TTS)模型可以生成捕获参考说话者的语音音色和口音的语音。然而,解开这些属性仍然具有挑战性,因为输出通常会从参考中继承重音和音色。在这项研究中,我们引入了一种新的,事后的,和培训免费的方法来中和口音,同时保留扬声器的原始音色,利用推理时间激活转向。我们首先离线提取特定于层的“导向向量”,这是来自内部激活的TTS模型之间的口音和母语的差异。在推理过程中,应用导向向量来引导模型产生口音中立、音色保留的语音。实验结果表明,所提出的导向矢量有效地抑制了输出口音,并对未被识别的口音说话人表现出很强的泛化能力,为无口音语音克隆提供了一种实用的解决方案。
摘要:Zero-shot Text-to-Speech (TTS) models can generate speech that captures both the voice timbre and accent of a reference speaker. However, disentangling these attributes remains challenging, as the output often inherits both the accent and timbre from the reference. In this study, we introduce a novel, post-hoc, and training-free approach to neutralize accent while preserving the speaker's original timbre, utilizing inference-time activation steering. We first extract layer-specific "steering vectors" offline, which are derived from the internal activation differences within the TTS model between accented and native speech. During inference, the steering vectors are applied to guide the model to produce accent-neutralized, timbre-preserving speech. Empirical results demonstrate that the proposed steering vectors effectively mitigate the output accent and exhibit strong generalizability to unseen accented speakers, offering a practical solution for accent-free voice cloning.
【7】Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
标题:重建!不编码:高可理解性和低延迟流媒体神经音频编解码器的自我监督表示重建丢失
链接:https://arxiv.org/abs/2603.05887
备注:Submitted to Interspeech 2026
摘要:针对梅尔频谱图重建优化的神经音频编解码器通常无法保持可懂度。虽然语义编码器提取改进了编码表示,但它不能保证重构语音中的内容保留。在这项工作中,我们证明了自监督表示重建(SSRR)损失从根本上提高了编解码器的训练和性能。首先,SSRR显著加快了收敛速度,仅使用单个GPU就能获得有竞争力的结果。其次,它通过从编解码器输出重建蒸馏的自监督表示来增强可理解性。第三,SSRR实现了高清晰度,而无需在基于Transformer的流编解码器中进行额外的前瞻,从而实现了用于实时部署的零前瞻架构。因此,我们的JHCodec实现了最先进的性能,同时保持最小的延迟和降低的训练成本。我们在Github www.example.com上开源了完整的实现、培训管道和演示。
摘要:Neural audio codecs optimized for mel-spectrogram reconstruction often fail to preserve intelligibility. While semantic encoder distillation improves encoded representations, it does not guarantee content preservation in reconstructed speech. In this work, we demonstrate that self-supervised representation reconstruction (SSRR) loss fundamentally improves codec training and performance. First, SSRR significantly accelerates convergence, enabling competitive results using only a single GPU. Second, it enhances intelligibility by reconstructing distilled self-supervised representations from codec outputs. Third, SSRR enables high intelligibility without additional lookahead in streaming Transformer-based codecs, allowing a zero-lookahead architecture for real-time deployment. As a result, our JHCodec achieves state-of-the-art performance while maintaining minimal latency and reduced training cost. We open-source the full implementation, training pipeline, and demo on Github https://github.com/jhcodec843/jhcodec.
【8】ImKWS: Test-Time Adaptation for Keyword Spotting with Class Imbalance
标题:ImKWS:针对类不平衡的关键词发现的测试时适应
链接:https://arxiv.org/abs/2603.05821
备注:Submitted to Interspeech
摘要:关键词识别(KWS)可以为语音助手识别单词,但环境噪音经常会降低准确性。标准改编解决了这个问题,并严格要求原始或标记的音频。测试时间自适应(TTA)仅使用未标记的测试音频来解决该数据约束。然而,目前的方法无法处理罕见的关键字和频繁的背景声音之间的严重不平衡。因此,标准的熵最小化(EM)变得过于自信,严重偏向于频繁的背景类。为了克服这个问题,我们提出了一个名为ImKWS的TTA方法。我们的方法将熵过程分成具有单独更新强度的奖励分支和惩罚分支。此外,我们在多个音频转换中强制执行一致性,以确保稳定的模型更新。在Google Speech Commands数据集上的实验表明,ImKWS在现实的不平衡场景中实现了可靠的自适应。代码可以在GitHub上找到。
摘要:Keyword spotting (KWS) identifies words for voice assistants, but environmental noise frequently reduces accuracy. Standard adaptation fixes this issue and strictly requires original or labeled audio. Test time adaptation (TTA) solves this data constraint using only unlabeled test audio. However, current methods fail to handle the severe imbalance between rare keywords and frequent background sounds. Consequently, standard entropy minimization (EM) becomes overconfident and heavily biased toward the frequent background class. To overcome this problem, we propose a TTA method named ImKWS. Our approach splits the entropy process into a reward branch and a penalty branch with separate update strengths. Furthermore, we enforce consistency across multiple audio transformations to ensure stable model updates. Experiments on the Google Speech Commands dataset indicate ImKWS achieves reliable adaptation in realistic imbalanced scenarios. The code is available on GitHub.
【9】Activation Steering for Accent Adaptation in Speech Foundation Models
标题:语音基础模型中口音适应的激活引导
链接:https://arxiv.org/abs/2603.05813
备注:Submitted to Interspeech. 5 pages
摘要:口音变异仍然是自动语音识别中的一个主要错误,然而大多数自适应方法依赖于参数微调,而不了解口音信息编码的位置。我们把重音变化作为隐藏表示中的可解释子空间,并研究它是否可以在激活空间中直接识别和控制。我们提取逐层编码器激活和估计平均偏移方向捕获口音引起的表示移位。通过将这些方向注入到各个层中并测量它们如何对齐重音和标准嵌入,我们得到了逐层的重音敏感度配置文件,揭示了重音信息集中在中间编码器层的窄带中。利用这种结构,我们进一步引入了无参数的口音转向,在推理过程中修改表示,而不更新模型权重。八种口音的实验表明,一致的单词错误率降低。
摘要:Accent variability remains a major errors in automatic speech recognition, yet most adaptation methods rely on parameter fine-tuning without understanding where accent information is encoded. We treat accent variation as an interpretable subspace in hidden representations and investigate whether it can be identified and controlled directly in activation space. We extract layer-wise encoder activations and estimate mean-shift directions capturing accent-induced representation shifts. By injecting these directions into individual layers and measuring how they align accented and standard embeddings, we derive a layer-wise accent sensitivity profile, revealing that accent information concentrates in a narrow band of middle encoder layers. Leveraging this structure, we further introduce parameter-free accent steering that modifies representations during inference without updating model weights. Experiments across eight accents show consistent word error rate reductions.
【10】Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding
标题:Whisper-CD:使用多负对比解码的准确长形式语音识别
链接:https://arxiv.org/abs/2603.06193
备注:Submitted to Interspeech 2026
摘要:使用大型编码器-解码器模型(如Whisper)的长形式语音识别通常会出现幻觉,重复循环和内容遗漏。当先前片段的转录被用作解码上下文时,这些错误可以累积并被进一步放大。我们提出了Whisper-CD,一个免训练的对比解码框架,将干净的音频logits与从三个声学激励扰动计算的负logits进行对比:高斯噪声注入,沉默信号和音频时间偏移。我们通过对数和exp算子聚合这些否定,为逐个令牌解码建立统一的多否定目标。在五个英语长格式基准测试中,Whisper-CD在CORAAL上将WER降低了24.3pp,并且比beam search的令牌生成吞吐量快了48%。由于Whisper-CD纯粹在推理时间运行,因此它可以作为替代品应用于已经部署的Whisper系统,而无需重新培训。
摘要:Long-form speech recognition with large encoder-decoder models such as Whisper often exhibit hallucinations, repetition loops, and content omissions. These errors can accumulate and be further amplified when the previous segment's transcription is used as decoding context. We propose Whisper-CD, a training-free contrastive decoding framework that contrasts clean-audio logits against negative logits computed from three acoustically motivated perturbations: Gaussian noise injection, silence signal, and audio temporal shift. We aggregate these negatives via the log-sum-exp operator, building a unified multi-negative objective for token-by-token decoding. Across five English long-form benchmarks, Whisper-CD reduces WER by up to 24.3pp on CORAAL and shows 48% faster token generation throughput than beam search. Because Whisper-CD operates purely at inference time, it can be applied as a drop-in replacement to already-deployed Whisper systems without retraining.
【11】Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
标题:Omni-C:将异类模式压缩到单个密集编码器中
链接:https://arxiv.org/abs/2603.05528
摘要:最近的多模态系统通常依赖于单独的专家模态编码器,这导致线性缩放的复杂性和计算开销与添加的模态。虽然统一的Omni模型通过具有专门专家和路由的专家混合(MoE)架构来解决这个问题,但它们仍然会增加参数计数并引入路由开销。在本文中,我们提出了Omni-C(Omni-Compress),这是一种基于transformer的单一密集编码器,它通过对大规模未对齐数据进行单峰对比预训练来学习跨异构模式(图像,音频和文本)的竞争性共享表示。通过最大化主干中的参数共享和使用轻量的模态特定投影头,Omni-C有效地减轻了模态间冲突,而无需MoE、配对监督或路由。该设计通过顺序模态处理和低内存推理支持在内存受限系统上的高效部署,从而消除了对并行专家加载或专用硬件的需求。实验表明,Omni-C在单峰和跨模型任务中实现了与专家模型相当的性能,对音频和文本的适度zero-shot降级在很大程度上通过轻量级线性探测或参数有效微调来恢复。与多编码器基线相比,统一架构大大减少了推理内存的使用,从而提高了高效和可扩展的多模式学习。
摘要:Recent multimodal systems often rely on separate expert modality encoders which cause linearly scaling complexity and computational overhead with added modalities. While unified Omni-models address this via Mixture-of-Expert (MoE) architectures with specialized experts and routing, they still inflate parameter counts and introduce routing overhead. In this paper, we propose Omni-C (Omni-Compress), a single dense Transformer-based encoder that learns competitive shared representations across heterogeneous modalities--images, audio, and text--through unimodal contrastive pretraining on large-scale unaligned data. By maximizing parameter sharing in the backbone and using lightweight modality-specific projection heads, Omni-C effectively mitigates inter-modality conflicts without requiring MoE, paired supervision, or routing. This design supports efficient deployment on memory-constrained systems via sequential modality processing and low-memory inference, eliminating the need for parallel expert loading or specialized hardware. Experiments show Omni-C achieves performance comparable to expert models in unimodal and cross-model tasks, with modest zero-shot degradation on audio and text that is largely recovered through lightweight linear probing or parameter efficient fine-tuning. The unified architecture substantially reduces inference memory usage compared to multi-encoder baselines, advancing efficient and scalable multimodal learning.
机器翻译由腾讯交互翻译提供,仅供参考
