微信公众号:arXiv_Daily
cs.SD语音
【1】Resurfacing Paralinguistic Awareness in Large Audio Language Models
标题:大型音频语言模型中重塑副语言意识
链接:https://arxiv.org/abs/2603.11947
备注:Submitted to Interspeech 2026
摘要:大型音频语言模型(LALM)已经将与人类的交互扩展到语音模态,这引入了巨大的交互潜力,由于隐含地指示用户上下文的非语言线索。然而,建立在当前的内容为中心的范例,LALM通常忽略这样的非语言线索和响应仅基于查询内容。在这项工作中,重新出现在LALMs的语言意识,我们引入了五个不同的逐层分析,共同确定语言层和语义理解层。基于这些见解,我们提出了一个语言增强的微调(PE-FT)协议,相应地配备LALM与语言感知的能力,包括(1)选择层微调,和(2)一个辅助的双层分类头。我们的实验表明,PE-FT协议有效地,有效地重现了语言意识,甚至超过了性能的全层微调策略。
摘要:Large Audio Language Models (LALMs) have expanded the interaction with human to speech modality, which introduces great interactive potential, due to the paralinguistic cues implicitly indicating the user context. However, building on the current content-centred paradigm, LALMs usually neglect such paralinguistic cues and respond solely based on query content. In this work, to resurface the paralinguistic awareness in LALMs, we introduce five diverse layer-wise analyses to jointly identify paralinguistic layers and semantic understanding layers. Based on these insights, we propose a paralinguistic-enhanced fine-tuning (PE-FT) protocol accordingly to equip LALMs with paralinguistic-aware capabilities, including (1) selective-layer fine-tuning, and (2) an auxiliary dual-level classification head. Our experiments demonstrate that PE-FT protocol efficiently and effectively resurfaces the paralinguistic awareness, even surpassing the performance of the all-layer fine-tuning strategy.
【2】Causal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2
标题:文本到言语的因果韵律调解:FastSpeech中持续时间、音调和精力的反事实训练2
链接:https://arxiv.org/abs/2603.11683
摘要:我们提出了一个新的因果韵律调解框架表达的文本到语音(TTS)合成。我们的方法通过显式情绪调节来增强FastSpeech 2架构,并引入反事实训练目标,以将情绪韵律与语言内容分开。通过制定一个结构因果模型的文本(内容),情感,和扬声器如何共同影响韵律(持续时间,音高,能量),并最终语音波形,我们得到两个互补的损失条款:间接路径约束(IPC),以执行情感影响语音只通过韵律,和反事实韵律约束(CPC),以鼓励不同的韵律模式,不同的情绪。由此产生的模型在多说话者情感语料库(LibriTTS,VCTK)上进行训练,其组合目标包括标准声谱图重建和方差预测损失以及我们的因果损失。在表达性语音合成的评估中,我们的方法实现了显着改进的韵律操作和情感渲染,具有更高的平均意见分数(MOS)和情感准确性比基线FastSpeech2变体。我们还观察到更好的可懂度(低WER)和扬声器的一致性时,跨扬声器转移情绪。广泛的消融证实,因果目标成功地分离韵律属性,产生一个可解释的模型,允许控制反事实的韵律编辑(如“相同的话语,不同的情感”),而不损害自然。我们讨论了韵律建模和轮廓的局限性,如假设情绪效果完全捕获音高,持续时间和能量的可识别性的影响。我们的工作证明了如何将因果学习原则融入TTS可以提高生成语音的可控性和表现力。
摘要:We propose a novel causal prosody mediation framework for expressive text-to-speech (TTS) synthesis. Our approach augments the FastSpeech2 architecture with explicit emotion conditioning and introduces counterfactual training objectives to disentangle emotional prosody from linguistic content. By formulating a structural causal model of how text (content), emotion, and speaker jointly influence prosody (duration, pitch, energy) and ultimately the speech waveform, we derive two complementary loss terms: an Indirect Path Constraint (IPC) to enforce that emotion affects speech only through prosody, and a Counterfactual Prosody Constraint (CPC) to encourage distinct prosody patterns for different emotions. The resulting model is trained on multi-speaker emotional corpora (LibriTTS, EmoV-DB, VCTK) with a combined objective that includes standard spectrogram reconstruction and variance prediction losses alongside our causal losses. In evaluations on expressive speech synthesis, our method achieves significantly improved prosody manipulation and emotion rendering, with higher mean opinion scores (MOS) and emotion accuracy than baseline FastSpeech2 variants. We also observe better intelligibility (low WER) and speaker consistency when transferring emotions across speakers. Extensive ablations confirm that the causal objectives successfully separate prosody attribution, yielding an interpretable model that allows controlled counterfactual prosody editing (e.g. "same utterance, different emotion") without compromising naturalness. We discuss the implications for identifiability in prosody modeling and outline limitations such as the assumption that emotion effects are fully captured by pitch, duration, and energy. Our work demonstrates how integrating causal learning principles into TTS can improve controllability and expressiveness in generated speech.
【3】Resonate: Reinforcing Text-to-Audio Generation via Online Feedback from Large Audio Language Models
标题:Resonate:通过大型音频语言模型的在线反馈加强文本到音频的生成
链接:https://arxiv.org/abs/2603.11661
摘要:强化学习(RL)已经成为增强大型语言模型(LLM)和视觉生成模型的有效范例。然而,它在文本到音频(TTA)生成中的应用在很大程度上仍未得到充分开发。以前的工作通常采用离线方法,如直接偏好优化(DPO),并利用对比音频预训练(CLAP)模型作为奖励函数。在这项研究中,我们研究了在线组相对策略优化(GRPO)到TTA生成的集成。我们适应基于流匹配的音频模型的算法,并证明在线RL显着优于其离线同行。此外,我们结合了来自大型音频语言模型(LALM)的奖励,它可以提供更好地与人类感知保持一致的细粒度评分信号。仅使用470 M参数,我们的最终模型\textbf{Resonate}就在音频质量和语义对齐方面在TTA-Bench上建立了一个新的SOTA。
摘要:Reinforcement Learning (RL) has become an effective paradigm for enhancing Large Language Models (LLMs) and visual generative models. However, its application in text-to-audio (TTA) generation remains largely under-explored. Prior work typically employs offline methods like Direct Preference Optimization (DPO) and leverages Contrastive Language-Audio Pretraining (CLAP) models as reward functions. In this study, we investigate the integration of online Group Relative Policy Optimization (GRPO) into TTA generation. We adapt the algorithm for Flow Matching-based audio models and demonstrate that online RL significantly outperforms its offline counterparts. Furthermore, we incorporate rewards derived from Large Audio Language Models (LALMs), which can provide fine-grained scoring signals that are better aligned with human perception. With only 470M parameters, our final model, \textbf{Resonate}, establishes a new SOTA on TTA-Bench in terms of both audio quality and semantic alignment.
【4】OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
标题:OmniForcing:释放实时联合视听生成
链接:https://arxiv.org/abs/2603.11647
备注:14 pages
摘要:最近的联合视听扩散模型实现了显着的生成质量,但遭受高延迟,由于其双向注意依赖性,阻碍实时应用。我们提出了OmniForcing,这是第一个将离线双流双向扩散模型提取为高保真流自回归生成器的框架。然而,天真地将因果蒸馏应用于这种双流架构会引发严重的训练不稳定性,这是由于模态之间的极端时间不对称性和由此产生的令牌稀疏性。我们通过引入一个零截断全局前缀的非对称块因果对齐来解决固有的信息密度差距,以防止多模态同步漂移。在因果转移期间由极端音频令牌稀疏性引起的梯度爆炸通过配备有身份RoPE约束的音频宿令牌机制进一步解决。最后,联合自强制蒸馏范例使模型能够动态地自校正长期推出期间来自曝光偏差的累积跨模态误差。OmniForcing采用独立于模态的滚动KV缓存推理方案,在单个GPU上以$\sim$25 FPS的速度实现最先进的流媒体生成,保持与双向教师同等的多模态同步和视觉质量。textbf{项目页面:} \href{https:omniforcing.com}{https://omniforcing.com}
摘要:Recent joint audio-visual diffusion models achieve remarkable generation quality but suffer from high latency due to their bidirectional attention dependencies, hindering real-time applications. We propose OmniForcing, the first framework to distill an offline, dual-stream bidirectional diffusion model into a high-fidelity streaming autoregressive generator. However, naively applying causal distillation to such dual-stream architectures triggers severe training instability, due to the extreme temporal asymmetry between modalities and the resulting token sparsity. We address the inherent information density gap by introducing an Asymmetric Block-Causal Alignment with a zero-truncation Global Prefix that prevents multi-modal synchronization drift. The gradient explosion caused by extreme audio token sparsity during the causal shift is further resolved through an Audio Sink Token mechanism equipped with an Identity RoPE constraint. Finally, a Joint Self-Forcing Distillation paradigm enables the model to dynamically self-correct cumulative cross-modal errors from exposure bias during long rollouts. Empowered by a modality-independent rolling KV-cache inference scheme, OmniForcing achieves state-of-the-art streaming generation at $\sim$25 FPS on a single GPU, maintaining multi-modal synchronization and visual quality on par with the bidirectional teacher.\textbf{Project Page:} \href{https://omniforcing.com}{https://omniforcing.com}
【5】Toward Complex-Valued Neural Networks for Waveform Generation
标题:走向复值神经网络以生成波
链接:https://arxiv.org/abs/2603.11589
备注:ICLR 2026 (accepted)
摘要:神经声码器最近先进的波形生成,产生自然和富有表现力的音频。在这些方法中,基于ISTFT的声码器最近受到关注。它们预测复值频谱图,然后通过iSTFT合成波形,从而避免可能增加计算成本的学习上采样阶段。然而,目前的方法使用独立处理实部和虚部的实值网络。这种分离限制了它们捕获复杂光谱图的固有结构的能力。我们提出了ComVo,一个复值神经声码器,其生成器和解码器使用本机复杂的算术。这使得对抗性训练框架能够以复值表示提供结构化反馈。为了以结构化的方式指导相位变换,我们引入了相位量化,它将相位值离散化并将训练过程正则化。最后,我们提出了一个块矩阵计算方案,通过减少冗余操作来提高训练效率。实验表明,ComVo实现了比可比的实值基线更高的合成质量,其块矩阵方案减少了25%的训练时间。音频样本和代码可在https://hs-oh-prml.github.io/ComVo/上获得。
摘要:Neural vocoders have recently advanced waveform generation, yielding natural and expressive audio. Among these approaches, iSTFT-based vocoders have recently gained attention. They predict a complex-valued spectrogram and then synthesize the waveform via iSTFT, thereby avoiding learned upsampling stages that can increase computational cost. However, current approaches use real-valued networks that process the real and imaginary parts independently. This separation limits their ability to capture the inherent structure of complex spectrograms. We present ComVo, a Complex-valued neural Vocoder whose generator and discriminator use native complex arithmetic. This enables an adversarial training framework that provides structured feedback in complex-valued representations. To guide phase transformations in a structured manner, we introduce phase quantization, which discretizes phase values and regularizes the training process. Finally, we propose a block-matrix computation scheme to improve training efficiency by reducing redundant operations. Experiments demonstrate that ComVo achieves higher synthesis quality than comparable real-valued baselines, and that its block-matrix scheme reduces training time by 25%. Audio samples and code are available at https://hs-oh-prml.github.io/ComVo/.
【6】AnimeScore: A Preference-Based Dataset and Framework for Evaluating Anime-Like Speech Style
标题:AnimeScore:用于评估类似动画的语音风格的基于偏好的数据集和框架
链接:https://arxiv.org/abs/2603.11482
摘要:目前,评估“动画般”的声音依赖于昂贵的主观判断,但没有标准化的客观指标存在。一个关键的挑战是,与自然不同,动画相似性缺乏共享的绝对尺度,使得传统的平均意见评分(MOS)协议不可靠。为了解决这个差距,我们提出了AnimeScore,一个基于偏好的框架,通过成对排名自动评估动画相似性。我们收集了15,000成对的判断,从187个评价与自由形式的描述,声学分析表明,感知动画相似性是由受控的共振成形,韵律连续性和故意清晰度,而不是简单的发音,如高音。我们发现,手工制作的声学特征达到了69.3%的AUC上限,而基于SSL的排名模型达到了90.8%的AUC,提供了一个实用的指标,也可以作为基于偏好的生成语音模型优化的奖励信号。
摘要:Evaluating 'anime-like' voices currently relies on costly subjective judgments, yet no standardized objective metric exists. A key challenge is that anime-likeness, unlike naturalness, lacks a shared absolute scale, making conventional Mean Opinion Score (MOS) protocols unreliable. To address this gap, we propose AnimeScore, a preference-based framework for automatic anime-likeness evaluation via pairwise ranking. We collect 15,000 pairwise judgments from 187 evaluators with free-form descriptions, and acoustic analysis reveals that perceived anime-likeness is driven by controlled resonance shaping, prosodic continuity, and deliberate articulation rather than simple heuristics such as high pitch. We show that handcrafted acoustic features reach a 69.3% AUC ceiling, while SSL-based ranking models achieve up to 90.8% AUC, providing a practical metric that can also serve as a reward signal for preference-based optimization of generative speech models.
【7】Stage-Adaptive Reliability Modeling for Continuous Valence-Arousal Estimation
标题:连续价唤起估计的阶段自适应可靠性建模
链接:https://arxiv.org/abs/2603.11468
备注:8 pages, 3 figures, 2 pages
摘要:在真实世界环境中的连续价唤醒估计是具有挑战性的,由于不一致的模态可靠性和交互依赖的视听信号的可变性。现有的方法主要集中在建模的时间动态,往往忽略了一个事实,即模态的可靠性可以大大不同的互动阶段。为了解决这个问题,我们提出了SAGE,这是一个阶段自适应可靠性建模框架,可以显式估计和校准多模态集成期间的模态置信度。SAGE引入了一种可靠性感知的融合机制,该机制根据音频和视觉表示的阶段相关信息量动态地重新平衡音频和视觉表示,防止不可靠的信号主导预测过程。通过将可靠性估计与特征表示分离,所提出的框架能够在交叉模态噪声、遮挡和不同交互条件下实现更稳定的情感估计。在Aff-Wild 2基准上进行的大量实验表明,与现有的多模态融合方法相比,SAGE始终提高了一致性相关系数得分,突出了可靠性驱动的连续影响预测建模的有效性。
摘要:Continuous valence-arousal estimation in real-world environments is challenging due to inconsistent modality reliability and interaction-dependent variability in audio-visual signals. Existing approaches primarily focus on modeling temporal dynamics, often overlooking the fact that modality reliability can vary substantially across interaction stages. To address this issue, we propose SAGE, a Stage-Adaptive reliability modeling framework that explicitly estimates and calibrates modality-wise confidence during multimodal integration. SAGE introduces a reliability-aware fusion mechanism that dynamically rebalances audio and visual representations according to their stage-dependent informativeness, preventing unreliable signals from dominating the prediction process. By separating reliability estimation from feature representation, the proposed framework enables more stable emotion estimation under cross-modal noise, occlusion, and varying interaction conditions. Extensive experiments on the Aff-Wild2 benchmark demonstrate that SAGE consistently improves concordance correlation coefficient scores compared with existing multimodal fusion approaches, highlighting the effectiveness of reliability-driven modeling for continuous affect prediction.
【8】Edge-Cloud Collaborative Speech Emotion Captioning via Token-Level Speculative Decoding in Audio-Language Models
标题:通过音频语言模型中的令牌级推测解码的边缘云协作语音情感字幕
链接:https://arxiv.org/abs/2603.11397
摘要:语音情感字幕(SEC)利用大型音频语言模型从语音中生成丰富的、上下文感知的情感描述。然而,由于对资源受限的边缘设备的大量计算需求以及传输生物特征音频的隐私风险,实际部署仍然具有挑战性。虽然较小的音频语言模型可以实现高效的设备上SEC,但其有限的容量往往会削弱微妙的非语言建模和细粒度的情感基础。我们提出了一个基于不确定性引导的推测性解码(UGSD)的边缘云协作框架。轻量级边缘模型在本地起草标题,只有高不确定性的令牌块才有选择地升级到更强大的云验证器进行验证。在MER 2024基准测试上的实验表明,BLEU性能的显著改善高达62.7%。与仅边缘模型相比,UGSD进一步实现了1.4倍的延迟和8.5倍的令牌吞吐量。这些结果经验性地表征了可部署SEC系统中的质量-效率-隐私权衡。
摘要:Speech Emotion Captioning (SEC) leverages large audio-language models to generate rich, context-aware affective descriptions from speech. However, real-world deployment remains challenging due to the substantial computational demands on resource-constrained edge devices and the privacy risks of transmitting biometric audio. While smaller audio-language models enable efficient on-device SEC, their limited capacity often weakens subtle paralinguistic modeling and fine-grained affective grounding. We propose an edge-cloud collaborative framework based on Uncertainty-Guided Speculative Decoding (UGSD). A lightweight edge model drafts captions locally, and only high-uncertainty token blocks are selectively escalated to a stronger cloud verifier for validation. Experiments on the MER2024 benchmark demonstrate substantial BLEU improvements up to 62.7%. UGSD further achieves 1.4x lower latency and 8.5x higher token throughput compared to an edge-only model. These results empirically characterize the quality-efficiency-privacy trade-off in deployable SEC systems.
【9】Continued Pretraining for Low-Resource Swahili ASR: Achieving State-of-the-Art Performance with Minimal Labeled Data
标题:低资源斯瓦希里语SVR的持续预训练:利用最少的标签数据实现最先进的性能
链接:https://arxiv.org/abs/2603.11378
摘要:我们调查持续预训练(CPT)适应wav 2 vec 2-BERT-2.0斯瓦希里语自动语音识别(ASR)。我们的方法通过伪标记CPT将未标记的音频与有限的标记数据结合起来,然后进行监督微调。有20,000个标记的样本,我们实现了3.24%的WER的共同语音斯瓦希里语-一个82%的相对改善基线。这一结果超过了以前报道的最好的学术系统(8.3% WER从XLS-R)61%的相对改善。我们提供了具体的数据要求和适用于其他低资源语言的可复制方法。
摘要:We investigate continued pretraining (CPT) for adapting wav2vec2-bert-2.0 to Swahili automatic speech recognition (ASR). Our approach combines unlabeled audio with limited labeled data through pseudo-labeled CPT followed by supervised finetuning. With 20,000 labeled samples, we achieve 3.24% WER on Common Voice Swahili-an 82% relative improvement over the baseline. This result surpasses the best previously reported academic system (8.3% WER from XLS-R) by 61% relative improvement. We provide concrete data requirements and a replicable methodology applicable to other low-resource languages.
【10】Fair-Gate: Fairness-Aware Interpretable Risk Gating for Sex-Fair Voice Biometrics
标题:Fair-Gate:性别公平语音生物识别技术的公平性可解释风险门控
链接:https://arxiv.org/abs/2603.11360
摘要:语音生物识别系统可能会表现出与性别相关的性能差距,即使整体验证精度很高。我们将这些差距归因于两个实际机制:(i)人口统计学捷径学习,其中说话人分类训练利用性别和说话人身份之间的虚假相关性,以及(ii)特征纠缠,其中与性别相关的声学变化与身份线索重叠,并且在不降低说话人歧视的情况下无法去除。我们提出了公平门,一个公平意识和可解释的风险门控框架,在一个单一的管道中解决这两种机制。公平门应用风险外推,以减少代理性别群体的扬声器分类风险的变化,并引入了一个本地的互补门,路由中间功能到一个身份分支和性别分支。该门通过产生一个明确的路由掩码来提供可解释性,可以检查该掩码以了解哪些特征被分配给身份与性别相关的路径。在VoxCeleb 1上的实验表明,Fair-Gate改进了效用-公平性权衡,在具有挑战性的评估条件下产生更多的性别公平ASV性能。
摘要:Voice biometric systems can exhibit sex-related performance gaps even when overall verification accuracy is strong. We attribute these gaps to two practical mechanisms: (i) demographic shortcut learning, where speaker classification training exploits spurious correlations between sex and speaker identity, and (ii) feature entanglement, where sex-linked acoustic variation overlaps with identity cues and cannot be removed without degrading speaker discrimination. We propose Fair-Gate, a fairness-aware and interpretable risk-gating framework that addresses both mechanisms in a single pipeline. Fair-Gate applies risk extrapolation to reduce variation in speaker-classification risk across proxy sex groups, and introduces a local complementary gate that routes intermediate features into an identity branch and a sex branch. The gate provides interpretability by producing an explicit routing mask that can be inspected to understand which features are allocated to identity versus sex-related pathways. Experiments on VoxCeleb1 show that Fair-Gate improves the utility--fairness trade-off, yielding more sex-fair ASV performance under challenging evaluation conditions.
【11】Huntington Disease Automatic Speech Recognition with Biomarker Supervision
标题:具有生物标志物监督的亨廷顿病自动语音识别
链接:https://arxiv.org/abs/2603.11168
摘要:病理性语音的自动语音识别(ASR)仍然未被探索,特别是对于亨廷顿氏病(HD),其中不规则的定时、不稳定的发声和发音失真挑战当前模型。我们提出了一个系统的HD-ASR研究,使用高保真临床语音语料库,以前没有用于端到端的ASR训练。我们比较多个ASR家庭下一个统一的评价,分析WER以及取代,缺失和插入模式。HD语音会导致特定于架构的错误机制,Parakeet-TDT优于编码器-解码器和CTC基线。HD特定的适应减少WER从6.99%到4.95%,我们还提出了一种方法,使用基于生物标记的辅助监督和分析错误行为是如何重塑严重程度相关的方式,而不是统一提高WER。我们开放所有代码和模型。
摘要:Automatic speech recognition (ASR) for pathological speech remains underexplored, especially for Huntington's disease (HD), where irregular timing, unstable phonation, and articulatory distortion challenge current models. We present a systematic HD-ASR study using a high-fidelity clinical speech corpus not previously used for end-to-end ASR training. We compare multiple ASR families under a unified evaluation, analyzing WER as well as substitution, deletion, and insertion patterns. HD speech induces architecture-specific error regimes, with Parakeet-TDT outperforming encoder-decoder and CTC baselines. HD-specific adaptation reduces WER from 6.99% to 4.95% and we also propose a method for using biomarker-based auxiliary supervision and analyze how error behavior is reshaped in severity-dependent ways rather than uniformly improving WER. We open-source all code and models.
【12】Uni-ASR: Unified LLM-Based Architecture for Non-Streaming and Streaming Automatic Speech Recognition
标题:Uni-ASB:基于LLM的统一架构,用于非流媒体和流媒体自动语音识别
链接:https://arxiv.org/abs/2603.11123
备注:Submitted to Interspeech 2026
摘要:尽管自动语音识别(ASR)系统与大型语言模型(LLM)的深度集成显著提高了准确性,但在低延迟流媒体场景中部署此类系统仍然具有挑战性。在本文中,我们提出了Uni-ASR,一个统一的框架,基于LLM,集成了非流和流语音识别功能。我们提出了一个联合训练范式,使系统能够在两种识别模式之间无缝转换,而无需任何架构修改。此外,我们引入了一个上下文感知的训练范式和一个共同设计的回退解码策略,它可以提高流识别的准确性,而不会引入额外的延迟。实验结果表明,Uni-ASR不仅在非流模式下具有较好的性能,而且在不同延迟约束下的流场景中也具有很强的有效性。
摘要:Although the deep integration of the Automatic Speech Recognition (ASR) system with Large Language Models (LLMs) has significantly improved accuracy, the deployment of such systems in low-latency streaming scenarios remains challenging. In this paper, we propose Uni-ASR, a unified framework based on LLMs that integrates both non-streaming and streaming speech recognition capabilities. We propose a joint training paradigm that enables the system to seamlessly transition between two recognition modes without any architectural modifications. Furthermore, we introduce a context-aware training paradigm and a co-designed fallback decoding strategy, which can enhance streaming recognition accuracy without introducing additional latency. The experimental results demonstrate that Uni-ASR not only achieves competitive performance within non-streaming mode, but also demonstrates strong effectiveness in streaming scenarios under diverse latency constraints.
【13】Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition
标题:用于视听情绪识别的具有时间对齐的多模式自我注意网络
链接:https://arxiv.org/abs/2603.11095
备注:5 pages, 3 figures, accepted to ICASSP 2026
摘要:视听情感识别(AVER)方法通常融合话语级特征,甚至帧级注意力模型也很少解决跨模态的帧速率不匹配问题。在本文中,我们提出了一个基于transformer的框架,专注于多模态特征的时间对齐。我们的设计采用了多模态自我注意编码器,同时捕获内和模态间的依赖关系在一个共享的特征空间。为了解决异构采样率,我们采用了时间对齐旋转位置嵌入(TaRoPE),隐式同步音频和视频令牌。此外,我们引入了一个跨时间匹配(CTM)的损失,强制执行时间上接近的对之间的一致性,引导编码器更好地对齐。CREMA-D和RAVDESS数据集上的实验表明,与最近的基线相比,一致的改进,表明明确解决帧速率不匹配有助于保留时间线索并增强跨模态融合。
摘要:Audio-visual emotion recognition (AVER) methods typically fuse utterance-level features, and even frame-level attention models seldom address the frame-rate mismatch across modalities. In this paper, we propose a Transformer-based framework focusing on the temporal alignment of multimodal features. Our design employs a multimodal self-attention encoder that simultaneously captures intra- and inter-modal dependencies within a shared feature space. To address heterogeneous sampling rates, we incorporate Temporally-aligned Rotary Position Embeddings (TaRoPE), which implicitly synchronize audio and video tokens. Furthermore, we introduce a Cross-Temporal Matching (CTM) loss that enforces consistency among temporally proximate pairs, guiding the encoder toward better alignment. Experiments on CREMA-D and RAVDESS datasets demonstrate consistent improvements over recent baselines, suggesting that explicitly addressing frame-rate mismatch helps preserve temporal cues and enhances cross-modal fusion.
【14】V2A-DPO: Omni-Preference Optimization for Video-to-Audio Generation
标题:V2 A-DPO:视频到音频生成的全偏好优化
链接:https://arxiv.org/abs/2603.11089
备注:Accepted at ICASSP2026
摘要:本文介绍了V2 A-DPO,一种新的直接偏好优化(DPO)框架,专为基于流的视频到音频生成(V2 A)模型量身定制,结合关键的调整,以有效地将生成的音频与人类的偏好相匹配。我们的方法结合了三个核心创新:(1)AudioScore-一个全面的人类偏好对齐评分系统,用于评估合成音频的语义一致性,时间对齐和感知质量;(2)自动化AudioScore驱动的管道,用于生成用于DPO优化的大规模偏好对数据;(3)课程学习授权的DPO优化策略,专门为基于流的生成模型量身定制。在基准VGGSound数据集上的实验表明,使用V2 A-DPO的人类偏好对齐的Frieren和MMAudio优于使用去噪扩散策略优化(DDPO)以及预训练基线优化的同行。此外,我们的DPO优化MMAudio在多个指标上实现了最先进的性能,超过了已发布的V2 A型号。
摘要:This paper introduces V2A-DPO, a novel Direct Preference Optimization (DPO) framework tailored for flow-based video-to-audio generation (V2A) models, incorporating key adaptations to effectively align generated audio with human preferences. Our approach incorporates three core innovations: (1) AudioScore-a comprehensive human preference-aligned scoring system for assessing semantic consistency, temporal alignment, and perceptual quality of synthesized audio; (2) an automated AudioScore-driven pipeline for generating large-scale preference pair data for DPO optimization; (3) a curriculum learning-empowered DPO optimization strategy specifically tailored for flow-based generative models. Experiments on benchmark VGGSound dataset demonstrate that human-preference aligned Frieren and MMAudio using V2A-DPO outperform their counterparts optimized using Denoising Diffusion Policy Optimization (DDPO) as well as pre-trained baselines. Furthermore, our DPO-optimized MMAudio achieves state-of-the-art performance across multiple metrics, surpassing published V2A models.
【15】Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
标题:SHAP-AV博士:通过Shapley归因解码视听语音识别中的相对情态贡献
链接:https://arxiv.org/abs/2603.12046
备注:Project website: https://umbertocappellazzo.github.io/Dr-SHAP-AV
摘要:视听语音识别(AVSR)利用声学和视觉信息在噪声下进行鲁棒识别。然而,模型如何平衡这些模式仍不清楚。我们提出了博士SHAP-AV,一个框架,使用Shapley值来分析AVSR中的模态贡献。通过在两个基准和不同SNR水平的六个模型上进行实验,我们介绍了三种分析:用于整体模态平衡的全局SHAP,用于解码过程中贡献动态的生成SHAP,以及用于输入输出对应的时间对齐SHAP。我们的研究结果表明,模型在噪声下转向视觉依赖,但即使在严重退化的情况下也能保持较高的音频贡献。模态平衡在生成过程中演变,在噪声下保持时间对齐,并且SNR是驱动模态加权的主导因素。这些研究结果暴露了一个持久的音频偏见,激励特设模态加权机制和Shapley为基础的归因作为一个标准的AVSR诊断。
摘要:Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual information for robust recognition under noise. However, how models balance these modalities remains unclear. We present Dr. SHAP-AV, a framework using Shapley values to analyze modality contributions in AVSR. Through experiments on six models across two benchmarks and varying SNR levels, we introduce three analyses: Global SHAP for overall modality balance, Generative SHAP for contribution dynamics during decoding, and Temporal Alignment SHAP for input-output correspondence. Our findings reveal that models shift toward visual reliance under noise yet maintain high audio contributions even under severe degradation. Modality balance evolves during generation, temporal alignment holds under noise, and SNR is the dominant factor driving modality weighting. These findings expose a persistent audio bias, motivating ad-hoc modality-weighting mechanisms and Shapley-based attribution as a standard AVSR diagnostic.
【16】Affect Decoding in Phonated and Silent Speech Production from Surface EMG
标题:影响表面EMG发声和无声语音产生的解码
链接:https://arxiv.org/abs/2603.11715
备注:Submitted to Interspeech 2026
摘要:情感的表达是口语交流的一部分,然而,它与潜在的发音执行的联系仍然不清楚。发音肌肉活动的措施,如肌电图可以揭示语音生产是如何调制的情绪与声学语音分析。我们研究了发声和无声言语产生过程中面部和颈部表面肌电图(sEMG)的情感解码。为此,我们引入了一个数据集,该数据集包括来自12名参与者的2,780个话语,跨越3个任务,在此基础上,我们使用一系列特征和模型嵌入来评估主体内和主体间解码。我们的研究结果表明,EMG表征可靠地区分挫折与高达0.845 AUC,并概括以及跨发音模式。我们的消融研究进一步表明,情感签名嵌入在面部运动活动中,并在没有发声的情况下持续存在,突出了EMG传感的影响意识到沉默的语音接口的潜力。
摘要:The expression of affect is integral to spoken communication, yet, its link to underlying articulatory execution remains unclear. Measures of articulatory muscle activity such as EMG could reveal how speech production is modulated by emotion alongside acoustic speech analyses. We investigate affect decoding from facial and neck surface electromyography (sEMG) during phonated and silent speech production. For this purpose, we introduce a dataset comprising 2,780 utterances from 12 participants across 3 tasks, on which we evaluate both intra- and inter-subject decoding using a range of features and model embeddings. Our results reveal that EMG representations reliably discriminate frustration with up to 0.845 AUC, and generalize well across articulation modes. Our ablation study further demonstrates that affective signatures are embedded in facial motor activity and persist in the absence of phonation, highlighting the potential of EMG sensing for affect-aware silent speech interfaces.
【17】RAF: Relativistic Adversarial Feedback For Universal Speech Synthesis
标题:英国皇家空军:通用语音合成的相对论对抗反馈
链接:https://arxiv.org/abs/2603.11678
备注:Submitted to Interspeech 2026
摘要:我们提出了相对论对抗反馈(RAF),一种新的训练目标的GAN声码器,提高域保真度和泛化到看不见的场景。虽然现代GAN声码器采用先进的体系结构,但它们的训练目标往往不能促进可推广的表示。RAF通过利用语音自监督学习模型来帮助鉴别器评估样本质量,鼓励生成器学习更丰富的表示来解决这个问题。此外,我们利用真实和假波形的相对论配对来改进训练数据分布的建模。跨多个数据集的实验显示,基于GAN的声码器在客观和主观指标上都有一致的增益。重要的是,RAF训练的BigVGAN基础在感知质量方面优于LSGAN训练的BigVGAN,仅使用12%的参数。比较研究进一步证实了RAF作为GAN声码器训练框架的有效性。
摘要:We propose Relativistic Adversarial Feedback (RAF), a novel training objective for GAN vocoders that improves in-domain fidelity and generalization to unseen scenarios. Although modern GAN vocoders employ advanced architectures, their training objectives often fail to promote generalizable representations. RAF addresses this problem by leveraging speech self-supervised learning models to assist discriminators in evaluating sample quality, encouraging the generator to learn richer representations. Furthermore, we utilize relativistic pairing for real and fake waveforms to improve the modeling of the training data distribution. Experiments across multiple datasets show consistent gains in both objective and subjective metrics on GAN-based vocoders. Importantly, the RAF-trained BigVGAN-base outperforms the LSGAN-trained BigVGAN in perceptual quality using only 12\% of the parameters. Comparative studies further confirm the effectiveness of RAF as a training framework for GAN vocoders.
【18】SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
标题:SEmamba++:利用全局、局部和周期性频谱模式的通用语音恢复框架
链接:https://arxiv.org/abs/2603.11669
备注:Submitted to Interspeech 2026
摘要:一般的语音恢复需要能够在各种失真下解释复杂语音结构的技术。虽然像SEAmba这样的状态空间模型在语音去噪方面取得了先进的进展,但它们并没有针对关键语音特征进行固有优化,例如频谱周期性或多分辨率频率分析。在这项工作中,我们引入了一个架构,将语音特定的功能作为归纳偏见。特别是,我们提出了频率GLP,频率特征提取块,有效地和高效地利用频率箱的属性。然后,我们设计了一个多分辨率并行时频双处理模块来捕获不同的频谱模式,并设计了一个可学习的映射来进一步提高模型的性能。结合我们所有的想法,提出的SEMamba++在保持计算效率的同时,在多个基线模型中实现了最佳性能。
摘要:General speech restoration demands techniques that can interpret complex speech structures under various distortions. While State-Space Models like SEMamba have advanced the state-of-the-art in speech denoising, they are not inherently optimized for critical speech characteristics, such as spectral periodicity or multi-resolution frequency analysis. In this work, we introduce an architecture tailored to incorporate speech-specific features as inductive biases. In particular, we propose Frequency GLP, a frequency feature extraction block that effectively and efficiently leverages the properties of frequency bins. Then, we design a multi-resolution parallel time-frequency dual-processing block to capture diverse spectral patterns, and a learnable mapping to further enhance model performance. With all our ideas combined, the proposed SEMamba++ achieves the best performance among multiple baseline models while remaining computationally efficient.
【19】Cough activity detection for automatic tuberculosis screening
标题:自动结核病筛查的咳嗽活动检测
链接:https://arxiv.org/abs/2603.11241
摘要:通过确定开始点和结束点来自动识别音频中的咳嗽片段对于在用于肺部相关疾病的健康技术中构建可扩展的筛查工具至关重要。我们提出了两个当前的预训练架构的应用程序的咳嗽活动检测的任务。采用了一个记录数据集,该数据集包含来自南非和乌干达社区一级护理中心的结核病(TB)患者的咳嗽症状。当使用XLS-R确定自动开始和结束点时,测试集的平均精度为0.96,接收器操作特性下的面积为0.99。我们表明,最好的平均精度是通过只利用前三层的网络,这具有减少计算和内存需求的双重好处,基于智能手机的应用程序的关键。该XLS-R配置被示出在测试集平均精度方面分别以9%和27%的绝对值优于音频频谱图Transformer(AST)以及逻辑回归基线。此外,使用由XLS-R自动隔离的咳嗽训练的下游TB分类模型轻松地优于在由AST隔离的咳嗽上训练的模型,并且仅勉强优于在地面真实咳嗽上训练的分类器。我们得出结论,应用大型预训练的Transformer模型是识别咳嗽终点的有效方法,并且将这种模型集成到筛选工具中是可行的。
摘要:The automatic identification of cough segments in audio through the determination of start and end points is pivotal to building scalable screening tools in health technologies for pulmonary related diseases. We propose the application of two current pre-trained architectures to the task of cough activity detection. A dataset of recordings containing cough from patients symptomatic for tuberculosis (TB) who self-present at community-level care centres in South Africa and Uganda is employed. When automatic start and end points are determined using XLS-R, an average precision of 0.96 and an area under the receiver-operating characteristic of 0.99 are achieved for the test set. We show that best average precision is achieved by utilising only the first three layers of the network, which has the dual benefits of reduced computational and memory requirements, pivotal for smartphone-based applications. This XLS-R configuration is shown to outperform an audio spectrogram transformer (AST) as well as a logistic regression baseline by 9% and 27% absolute in test set average precision respectively. Furthermore, a downstream TB classification model trained using the coughs automatically isolated by XLS-R comfortably outperforms a model trained on the coughs isolated by AST, and is only narrowly outperformed by a classifier trained on the ground truth coughs. We conclude that the application of large pre-trained transformer models is an effective approach to identifying cough end-points and that the integration of such a model into a screening tool is feasible.
【20】Can LLMs Help Localize Fake Words in Partially Fake Speech?
标题:LLM能否帮助本地化部分虚假言语中的虚假词语?
链接:https://arxiv.org/abs/2603.11205
备注:Submitted to Interspeech 2026; put on arxiv based on requirement from Interspeech: "Interspeech no longer enforces an anonymity period for submissions." and "For authors that prefer to upload their paper online, a note indicating that the paper was submitted for review to Interspeech should be included in the posting."
摘要:在大规模文本上训练的大型语言模型(LLM)最近因其在许多任务中的出色表现而引起了人们的极大关注。出于这一动机,我们研究了文本训练的LLM是否可以帮助定位部分假语音中的假单词,其中只有语音中的特定单词被编辑。我们建立了一个语音LLM通过下一个令牌预测执行假词定位。在AV-Deepfake 1 M和PartialEdit上的实验和分析表明,该模型经常利用从训练数据中学习到的编辑风格模式,特别是我们讨论的这两个数据库的单词级极性替换,作为定位假单词的线索。尽管这些特定模式在域内场景中提供了有用的信息,但如何避免过度依赖这些特定模式并改进对不可见编辑样式的泛化仍然是一个悬而未决的问题。
摘要:Large language models (LLMs), trained on large-scale text, have recently attracted significant attention for their strong performance across many tasks. Motivated by this, we investigate whether a text-trained LLM can help localize fake words in partially fake speech, where only specific words within a speech are edited. We build a speech LLM to perform fake word localization via next token prediction. Experiments and analyses on AV-Deepfake1M and PartialEdit indicates that the model frequently leverages editing-style pattern learned from the training data, particularly word-level polarity substitutions for those two databases we discussed, as cues for localizing fake words. Although such particular patterns provide useful information in an in-domain scenario, how to avoid over-reliance on such particular pattern and improve generalization to unseen editing styles remains an open question.
【1】Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
标题:SHAP-AV博士:通过Shapley归因解码视听语音识别中的相对情态贡献
链接:https://arxiv.org/abs/2603.12046
备注:Project website: https://umbertocappellazzo.github.io/Dr-SHAP-AV
摘要:视听语音识别(AVSR)利用声学和视觉信息在噪声下进行鲁棒识别。然而,模型如何平衡这些模式仍不清楚。我们提出了博士SHAP-AV,一个框架,使用Shapley值来分析AVSR中的模态贡献。通过在两个基准和不同SNR水平的六个模型上进行实验,我们介绍了三种分析:用于整体模态平衡的全局SHAP,用于解码过程中贡献动态的生成SHAP,以及用于输入输出对应的时间对齐SHAP。我们的研究结果表明,模型在噪声下转向视觉依赖,但即使在严重退化的情况下也能保持较高的音频贡献。模态平衡在生成过程中演变,在噪声下保持时间对齐,并且SNR是驱动模态加权的主导因素。这些研究结果暴露了一个持久的音频偏见,激励特设模态加权机制和Shapley为基础的归因作为一个标准的AVSR诊断。
摘要:Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual information for robust recognition under noise. However, how models balance these modalities remains unclear. We present Dr. SHAP-AV, a framework using Shapley values to analyze modality contributions in AVSR. Through experiments on six models across two benchmarks and varying SNR levels, we introduce three analyses: Global SHAP for overall modality balance, Generative SHAP for contribution dynamics during decoding, and Temporal Alignment SHAP for input-output correspondence. Our findings reveal that models shift toward visual reliance under noise yet maintain high audio contributions even under severe degradation. Modality balance evolves during generation, temporal alignment holds under noise, and SNR is the dominant factor driving modality weighting. These findings expose a persistent audio bias, motivating ad-hoc modality-weighting mechanisms and Shapley-based attribution as a standard AVSR diagnostic.
【2】Silent Speech Interfaces in the Era of Large Language Models: A Comprehensive Taxonomy and Systematic Review
标题:大型语言模型时代的无声语音界面:全面分类和系统回顾
链接:https://arxiv.org/abs/2603.11877
备注:20 pages, 4 figures
摘要:人机交互传统上依赖于声学通道,这是一种依赖性,它会对环境噪声、隐私限制和生理语音障碍带来系统漏洞。无声言语接口(SSIs)是一种跨越声学阶段的变革性范式,它直接从神经-肌肉-发音连续体中解码语言意图。这篇综述提供了SSI景观的高层次综合,从传统的以传感器为中心的分析过渡到整体的意图到执行分类。我们系统地评估了四个关键生理拦截点的感知方式:神经振荡,神经肌肉激活,关节运动学(超声/磁力测量),以及通过声学或射频感知的普遍主动探测。关键是,我们分析了当前的范式转变,从启发式信号处理潜在语义对齐。在这个新时代,大型语言模型(LLM)和深度生成架构作为高级语言先验来解决生物信号的“信息稀疏性”和非平稳性。通过将碎片化的生理手势映射到结构化的语义潜在空间中,现代SSI框架首次接近了现实世界部署所需的单词错误率可用性阈值。我们进一步研究了SSI从笨重的实验室仪器到集成到商品级可穿戴设备(如earables和智能眼镜)中的“隐形界面”的转变。最后,我们概述了一个战略路线图,通过自我监督的基础模型解决“用户依赖悖论”,并定义了“神经安全”的道德边界,以保护认知自由,在日益接口的世界。
摘要:Human-computer interaction has traditionally relied on the acoustic channel, a dependency that introduces systemic vulnerabilities to environmental noise, privacy constraints, and physiological speech impairments. Silent Speech Interfaces (SSIs) emerge as a transformative paradigm that bypasses the acoustic stage by decoding linguistic intent directly from the neuro-muscular-articulatory continuum. This review provides a high-level synthesis of the SSI landscape, transitioning from traditional transducer-centric analysis to a holistic intent-to-execution taxonomy. We systematically evaluate sensing modalities across four critical physiological interception points: neural oscillations, neuromuscular activation, articulatory kinematics (ultrasound/magnetometry), and pervasive active probing via acoustic or radio-frequency sensing. Critically, we analyze the current paradigm shift from heuristic signal processing to Latent Semantic Alignment. In this new era, Large Language Models (LLMs) and deep generative architectures serve as high-level linguistic priors to resolve the ``informational sparsity'' and non-stationarity of biosignals. By mapping fragmented physiological gestures into structured semantic latent spaces, modern SSI frameworks have, for the first time, approached the Word Error Rate usability threshold required for real-world deployment. We further examine the transition of SSIs from bulky laboratory instrumentation to ``invisible interfaces'' integrated into commodity-grade wearables, such as earables and smart glasses. Finally, we outline a strategic roadmap addressing the ``user-dependency paradox'' through self-supervised foundation models and define the ethical boundaries of ``neuro-security'' to protect cognitive liberty in an increasingly interfaced world.
【3】Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data
标题:利用MRI数据通过语音表示从语音重建声道
链接:https://arxiv.org/abs/2603.11847
摘要:发音声学反演旨在从语音信号中重建声道的完整几何结构。在本文中,我们提出了一个比较研究的几个层次的语音分割精度,连同比较基线介绍了我们以前的工作,这是基于梅尔频率倒谱系数(MFCC)。所有考虑的方法都是基于去噪语音信号,目的是调查通过三个连续的水平将语音信息的影响:未校正的自动转录,时间对齐的语音分割,和专家手动校正后对齐。训练模型预测发音轮廓提取的声道MRI图像使用自动轮廓跟踪方法。结果表明,在依赖于语音表示的模型中,对齐后的手动校正产生了最好的性能,接近基线。
摘要:Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a comparison to the baseline introduced in our previous work, which is based on Mel-Frequency Cepstral Coefficients (MFCCs). All the approaches considered are based on a denoised speech signal and aim to investigate the impact of incorporating phonetic information through three successive levels: an uncorrected automatic transcription, a temporally aligned phonetic segmentation, and an expert manual correction following alignment. The models are trained to predict articulatory contours extracted from vocal tract MRI images using an automatic contour tracking method. The results show that, among the models relying on phonetic representations, manual correction after alignment yields the best performance, approaching that of the baseline.
【4】Acoustic-to-Articulatory Inversion of Clean Speech Using an MRI-Trained Model
标题:使用核磁共振训练模型进行干净言语的声学到关节翻转
链接:https://arxiv.org/abs/2603.11845
摘要:发音声学反转从语音重建声道形状。实时磁共振成像(rt-MRI)允许同时采集声学语音信号和发音信息。除了rt-MRI采集的复杂性外,记录的音频还受到扫描仪噪声的严重破坏,需要进行降噪才能使用。为了实际使用,必须能够在没有MRI噪声的情况下反转记录的语音。在这项研究中,我们调查使用的语音记录在一个干净的声学环境中作为一种替代去噪MRI语音。为此,我们比较了两个信号,从同一个扬声器相同的句子,使用语音分割对齐。在去噪MRI语音上训练的模型在去噪MRI和干净语音上进行评估。我们还评估了一个只在干净语音上训练和测试的模型。结果表明,干净的语音有效地支持发音反转,实现了1.56 mm的RMSE,接近基于MRI的性能。
摘要:Articulatory acoustic inversion reconstructs vocal tract shapes from speech. Real-time magnetic resonance imaging (rt-MRI) allows simultaneous acquisition of both the acoustic speech signal and articulatory information. Besides the complexity of rt-MRI acquisition, the recorded audio is heavily corrupted by scanner noise and requires denoising to be usable. For practical use, it must be possible to invert speech recorded without MRI noise. In this study, we investigate the use of speech recorded in a clean acoustic environment as an alternative to denoised MRI speech. To this end we compare two signals from the same speaker with identical sentences which are aligned using phonetic segmentation. A model trained on denoised MRI speech is evaluated on both denoised MRI and clean speech. We also assess a model trained and tested only on clean speech. Results show that clean speech supports articulatory inversion effectively, achieving an RMSE of 1.56 mm, close to MRI-based performance.
【5】ReDimNet2: Scaling Speaker Verification via Time-Pooled Dimension Reshaping
标题:ReDimNet 2:通过时间合并维度重塑扩展说话者验证
链接:https://arxiv.org/abs/2603.11841
备注:Submitted to Interspeech 2026
摘要:我们提出了ReDimNet 2,这是一种改进的神经网络架构,用于提取建立在ReDimNet维度重塑框架基础上的话语级说话人表示。ReDimNet 2中的关键修改是在1D处理路径中引入了时间维度上的池化。该操作保留了1D特征空间的性质,因为1D特征仍然是2D特征的重塑版本,而不管时间分辨率如何,同时能够在不成比例地增加计算的情况下显著更积极地缩放通道维度。我们介绍了一个家庭的七个模型配置(B 0-B6),从1.1 M到12.3 M的参数和0.33至13 GMACS。在VoxCeleb 1基准测试上的实验结果表明,与ReDimNet相比,ReDimNet 2在每个尺度点上都提高了计算成本与精度的Pareto前沿,在Vox 1-O上实现了0.287%的EER,具有12.3M参数和13个GMACS。
摘要:We present ReDimNet2, an improved neural network architecture for extracting utterance-level speaker representations that builds upon the ReDimNet dimension-reshaping framework. The key modification in ReDimNet2 is the introduction of pooling over the time dimension within the 1D processing pathway. This operation preserves the nature of the 1D feature space, since 1D features remain a reshaped version of 2D features regardless of temporal resolution, while enabling significantly more aggressive scaling of the channel dimension without proportional compute increase. We introduce a family of seven model configurations (B0-B6) ranging from 1.1M to 12.3M parameters and 0.33 to 13 GMACS. Experimental results on VoxCeleb1 benchmarks demonstrate that ReDimNet2 improves the Pareto front of computational cost versus accuracy at every scale point compared to ReDimNet, achieving 0.287% EER on Vox1-O with 12.3M parameters and 13 GMACS.
【6】Affect Decoding in Phonated and Silent Speech Production from Surface EMG
标题:影响表面EMG发声和无声语音产生的解码
链接:https://arxiv.org/abs/2603.11715
备注:Submitted to Interspeech 2026
摘要:情感的表达是口语交流的一部分,然而,它与潜在的发音执行的联系仍然不清楚。发音肌肉活动的措施,如肌电图可以揭示语音生产是如何调制的情绪与声学语音分析。我们研究了发声和无声言语产生过程中面部和颈部表面肌电图(sEMG)的情感解码。为此,我们引入了一个数据集,该数据集包括来自12名参与者的2,780个话语,跨越3个任务,在此基础上,我们使用一系列特征和模型嵌入来评估主体内和主体间解码。我们的研究结果表明,EMG表征可靠地区分挫折与高达0.845 AUC,并概括以及跨发音模式。我们的消融研究进一步表明,情感签名嵌入在面部运动活动中,并在没有发声的情况下持续存在,突出了EMG传感的影响意识到沉默的语音接口的潜力。
摘要:The expression of affect is integral to spoken communication, yet, its link to underlying articulatory execution remains unclear. Measures of articulatory muscle activity such as EMG could reveal how speech production is modulated by emotion alongside acoustic speech analyses. We investigate affect decoding from facial and neck surface electromyography (sEMG) during phonated and silent speech production. For this purpose, we introduce a dataset comprising 2,780 utterances from 12 participants across 3 tasks, on which we evaluate both intra- and inter-subject decoding using a range of features and model embeddings. Our results reveal that EMG representations reliably discriminate frustration with up to 0.845 AUC, and generalize well across articulation modes. Our ablation study further demonstrates that affective signatures are embedded in facial motor activity and persist in the absence of phonation, highlighting the potential of EMG sensing for affect-aware silent speech interfaces.
【7】RAF: Relativistic Adversarial Feedback For Universal Speech Synthesis
标题:英国皇家空军:通用语音合成的相对论对抗反馈
链接:https://arxiv.org/abs/2603.11678
备注:Submitted to Interspeech 2026
摘要:我们提出了相对论对抗反馈(RAF),一种新的训练目标的GAN声码器,提高域保真度和泛化到看不见的场景。虽然现代GAN声码器采用先进的体系结构,但它们的训练目标往往不能促进可推广的表示。RAF通过利用语音自监督学习模型来帮助鉴别器评估样本质量,鼓励生成器学习更丰富的表示来解决这个问题。此外,我们利用真实和假波形的相对论配对来改进训练数据分布的建模。跨多个数据集的实验显示,基于GAN的声码器在客观和主观指标上都有一致的增益。重要的是,RAF训练的BigVGAN基础在感知质量方面优于LSGAN训练的BigVGAN,仅使用12%的参数。比较研究进一步证实了RAF作为GAN声码器训练框架的有效性。
摘要:We propose Relativistic Adversarial Feedback (RAF), a novel training objective for GAN vocoders that improves in-domain fidelity and generalization to unseen scenarios. Although modern GAN vocoders employ advanced architectures, their training objectives often fail to promote generalizable representations. RAF addresses this problem by leveraging speech self-supervised learning models to assist discriminators in evaluating sample quality, encouraging the generator to learn richer representations. Furthermore, we utilize relativistic pairing for real and fake waveforms to improve the modeling of the training data distribution. Experiments across multiple datasets show consistent gains in both objective and subjective metrics on GAN-based vocoders. Importantly, the RAF-trained BigVGAN-base outperforms the LSGAN-trained BigVGAN in perceptual quality using only 12\% of the parameters. Comparative studies further confirm the effectiveness of RAF as a training framework for GAN vocoders.
【8】SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
标题:SEmamba++:利用全局、局部和周期性频谱模式的通用语音恢复框架
链接:https://arxiv.org/abs/2603.11669
备注:Submitted to Interspeech 2026
摘要:一般的语音恢复需要能够在各种失真下解释复杂语音结构的技术。虽然像SEAmba这样的状态空间模型在语音去噪方面取得了先进的进展,但它们并没有针对关键语音特征进行固有优化,例如频谱周期性或多分辨率频率分析。在这项工作中,我们引入了一个架构,将语音特定的功能作为归纳偏见。特别是,我们提出了频率GLP,频率特征提取块,有效地和高效地利用频率箱的属性。然后,我们设计了一个多分辨率并行时频双处理模块来捕获不同的频谱模式,并设计了一个可学习的映射来进一步提高模型的性能。结合我们所有的想法,提出的SEMamba++在保持计算效率的同时,在多个基线模型中实现了最佳性能。
摘要:General speech restoration demands techniques that can interpret complex speech structures under various distortions. While State-Space Models like SEMamba have advanced the state-of-the-art in speech denoising, they are not inherently optimized for critical speech characteristics, such as spectral periodicity or multi-resolution frequency analysis. In this work, we introduce an architecture tailored to incorporate speech-specific features as inductive biases. In particular, we propose Frequency GLP, a frequency feature extraction block that effectively and efficiently leverages the properties of frequency bins. Then, we design a multi-resolution parallel time-frequency dual-processing block to capture diverse spectral patterns, and a learnable mapping to further enhance model performance. With all our ideas combined, the proposed SEMamba++ achieves the best performance among multiple baseline models while remaining computationally efficient.
【9】Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
标题:使用CSC编码器草案对基于LLM的ASB进行自我推测解码
链接:https://arxiv.org/abs/2603.11243
摘要:我们提出了自推测解码的语音感知LLM使用CTC编码器作为草案模型,以加速自回归(AR)推理和提高ASR的准确性。我们的三步过程如下工作:(1)如果CTC输出分布的帧熵低于阈值,则贪婪CTC假设被接受为最终的;(2)否则,使用基于令牌似然的宽松接受标准在单个LLM前向传递中验证CTC假设;(3)如果验证失败,则AR解码从接受的CTC前缀恢复。在九个语料库和五种语言上的实验表明,该方法可以同时加速解码和减少WER。在具有1B参数LLM和440 M参数CTC编码器的HuggingFace Open ASR基准测试中,我们实现了创纪录的5.58% WER,并将逆实时因子提高了4.4倍,相对于AR搜索仅增加了12%的WER。代码和模型权重在许可证下公开可用。
摘要:We propose self-speculative decoding for speech-aware LLMs by using the CTC encoder as a draft model to accelerate auto-regressive (AR) inference and improve ASR accuracy. Our three-step procedure works as follows: (1) if the frame entropies of the CTC output distributions are below a threshold, the greedy CTC hypothesis is accepted as final; (2) otherwise, the CTC hypothesis is verified in a single LLM forward pass using a relaxed acceptance criterion based on token likelihoods; (3) if verification fails, AR decoding resumes from the accepted CTC prefix. Experiments on nine corpora and five languages show that this approach can simultaneously accelerate decoding and reduce WER. On the HuggingFace Open ASR benchmark with a 1B parameter LLM and 440M parameter CTC encoder, we achieve a record 5.58% WER and improve the inverse real time factor by a factor of 4.4 with only a 12% relative WER increase over AR search. Code and model weights are publicly available under a permissive license.
【10】Cough activity detection for automatic tuberculosis screening
标题:自动结核病筛查的咳嗽活动检测
链接:https://arxiv.org/abs/2603.11241
摘要:通过确定开始点和结束点来自动识别音频中的咳嗽片段对于在用于肺部相关疾病的健康技术中构建可扩展的筛查工具至关重要。我们提出了两个当前的预训练架构的应用程序的咳嗽活动检测的任务。采用了一个记录数据集,该数据集包含来自南非和乌干达社区一级护理中心的结核病(TB)患者的咳嗽症状。当使用XLS-R确定自动开始和结束点时,测试集的平均精度为0.96,接收器操作特性下的面积为0.99。我们表明,最好的平均精度是通过只利用前三层的网络,这具有减少计算和内存需求的双重好处,基于智能手机的应用程序的关键。该XLS-R配置被示出在测试集平均精度方面分别以9%和27%的绝对值优于音频频谱图Transformer(AST)以及逻辑回归基线。此外,使用由XLS-R自动隔离的咳嗽训练的下游TB分类模型轻松地优于在由AST隔离的咳嗽上训练的模型,并且仅勉强优于在地面真实咳嗽上训练的分类器。我们得出结论,应用大型预训练的Transformer模型是识别咳嗽终点的有效方法,并且将这种模型集成到筛选工具中是可行的。
摘要:The automatic identification of cough segments in audio through the determination of start and end points is pivotal to building scalable screening tools in health technologies for pulmonary related diseases. We propose the application of two current pre-trained architectures to the task of cough activity detection. A dataset of recordings containing cough from patients symptomatic for tuberculosis (TB) who self-present at community-level care centres in South Africa and Uganda is employed. When automatic start and end points are determined using XLS-R, an average precision of 0.96 and an area under the receiver-operating characteristic of 0.99 are achieved for the test set. We show that best average precision is achieved by utilising only the first three layers of the network, which has the dual benefits of reduced computational and memory requirements, pivotal for smartphone-based applications. This XLS-R configuration is shown to outperform an audio spectrogram transformer (AST) as well as a logistic regression baseline by 9% and 27% absolute in test set average precision respectively. Furthermore, a downstream TB classification model trained using the coughs automatically isolated by XLS-R comfortably outperforms a model trained on the coughs isolated by AST, and is only narrowly outperformed by a classifier trained on the ground truth coughs. We conclude that the application of large pre-trained transformer models is an effective approach to identifying cough end-points and that the integration of such a model into a screening tool is feasible.
【11】Can LLMs Help Localize Fake Words in Partially Fake Speech?
标题:LLM能否帮助本地化部分虚假言语中的虚假词语?
链接:https://arxiv.org/abs/2603.11205
备注:Submitted to Interspeech 2026; put on arxiv based on requirement from Interspeech: "Interspeech no longer enforces an anonymity period for submissions." and "For authors that prefer to upload their paper online, a note indicating that the paper was submitted for review to Interspeech should be included in the posting."
摘要:在大规模文本上训练的大型语言模型(LLM)最近因其在许多任务中的出色表现而引起了人们的极大关注。出于这一动机,我们研究了文本训练的LLM是否可以帮助定位部分假语音中的假单词,其中只有语音中的特定单词被编辑。我们建立了一个语音LLM通过下一个令牌预测执行假词定位。在AV-Deepfake 1 M和PartialEdit上的实验和分析表明,该模型经常利用从训练数据中学习到的编辑风格模式,特别是我们讨论的这两个数据库的单词级极性替换,作为定位假单词的线索。尽管这些特定模式在域内场景中提供了有用的信息,但如何避免过度依赖这些特定模式并改进对不可见编辑样式的泛化仍然是一个悬而未决的问题。
摘要:Large language models (LLMs), trained on large-scale text, have recently attracted significant attention for their strong performance across many tasks. Motivated by this, we investigate whether a text-trained LLM can help localize fake words in partially fake speech, where only specific words within a speech are edited. We build a speech LLM to perform fake word localization via next token prediction. Experiments and analyses on AV-Deepfake1M and PartialEdit indicates that the model frequently leverages editing-style pattern learned from the training data, particularly word-level polarity substitutions for those two databases we discussed, as cues for localizing fake words. Although such particular patterns provide useful information in an in-domain scenario, how to avoid over-reliance on such particular pattern and improve generalization to unseen editing styles remains an open question.
【12】Resurfacing Paralinguistic Awareness in Large Audio Language Models
标题:大型音频语言模型中重塑副语言意识
链接:https://arxiv.org/abs/2603.11947
备注:Submitted to Interspeech 2026
摘要:大型音频语言模型(LALM)已经将与人类的交互扩展到语音模态,这引入了巨大的交互潜力,由于隐含地指示用户上下文的非语言线索。然而,建立在当前的内容为中心的范例,LALM通常忽略这样的非语言线索和响应仅基于查询内容。在这项工作中,重新出现在LALMs的语言意识,我们引入了五个不同的逐层分析,共同确定语言层和语义理解层。基于这些见解,我们提出了一个语言增强的微调(PE-FT)协议,相应地配备LALM与语言感知的能力,包括(1)选择层微调,和(2)一个辅助的双层分类头。我们的实验表明,PE-FT协议有效地,有效地重现了语言意识,甚至超过了性能的全层微调策略。
摘要:Large Audio Language Models (LALMs) have expanded the interaction with human to speech modality, which introduces great interactive potential, due to the paralinguistic cues implicitly indicating the user context. However, building on the current content-centred paradigm, LALMs usually neglect such paralinguistic cues and respond solely based on query content. In this work, to resurface the paralinguistic awareness in LALMs, we introduce five diverse layer-wise analyses to jointly identify paralinguistic layers and semantic understanding layers. Based on these insights, we propose a paralinguistic-enhanced fine-tuning (PE-FT) protocol accordingly to equip LALMs with paralinguistic-aware capabilities, including (1) selective-layer fine-tuning, and (2) an auxiliary dual-level classification head. Our experiments demonstrate that PE-FT protocol efficiently and effectively resurfaces the paralinguistic awareness, even surpassing the performance of the all-layer fine-tuning strategy.
【13】AnimeScore: A Preference-Based Dataset and Framework for Evaluating Anime-Like Speech Style
标题:AnimeScore:用于评估类似动画的语音风格的基于偏好的数据集和框架
链接:https://arxiv.org/abs/2603.11482
摘要:目前,评估“动画般”的声音依赖于昂贵的主观判断,但没有标准化的客观指标存在。一个关键的挑战是,与自然不同,动画相似性缺乏共享的绝对尺度,使得传统的平均意见评分(MOS)协议不可靠。为了解决这个差距,我们提出了AnimeScore,一个基于偏好的框架,通过成对排名自动评估动画相似性。我们收集了15,000成对的判断,从187个评价与自由形式的描述,声学分析表明,感知动画相似性是由受控的共振成形,韵律连续性和故意清晰度,而不是简单的发音,如高音。我们发现,手工制作的声学特征达到了69.3%的AUC上限,而基于SSL的排名模型达到了90.8%的AUC,提供了一个实用的指标,也可以作为基于偏好的生成语音模型优化的奖励信号。
摘要:Evaluating 'anime-like' voices currently relies on costly subjective judgments, yet no standardized objective metric exists. A key challenge is that anime-likeness, unlike naturalness, lacks a shared absolute scale, making conventional Mean Opinion Score (MOS) protocols unreliable. To address this gap, we propose AnimeScore, a preference-based framework for automatic anime-likeness evaluation via pairwise ranking. We collect 15,000 pairwise judgments from 187 evaluators with free-form descriptions, and acoustic analysis reveals that perceived anime-likeness is driven by controlled resonance shaping, prosodic continuity, and deliberate articulation rather than simple heuristics such as high pitch. We show that handcrafted acoustic features reach a 69.3% AUC ceiling, while SSL-based ranking models achieve up to 90.8% AUC, providing a practical metric that can also serve as a reward signal for preference-based optimization of generative speech models.
【14】Continued Pretraining for Low-Resource Swahili ASR: Achieving State-of-the-Art Performance with Minimal Labeled Data
标题:低资源斯瓦希里语SVR的持续预训练:利用最少的标签数据实现最先进的性能
链接:https://arxiv.org/abs/2603.11378
摘要:我们调查持续预训练(CPT)适应wav 2 vec 2-BERT-2.0斯瓦希里语自动语音识别(ASR)。我们的方法通过伪标记CPT将未标记的音频与有限的标记数据结合起来,然后进行监督微调。有20,000个标记的样本,我们实现了3.24%的WER的共同语音斯瓦希里语-一个82%的相对改善基线。这一结果超过了以前报道的最好的学术系统(8.3% WER从XLS-R)61%的相对改善。我们提供了具体的数据要求和适用于其他低资源语言的可复制方法。
摘要:We investigate continued pretraining (CPT) for adapting wav2vec2-bert-2.0 to Swahili automatic speech recognition (ASR). Our approach combines unlabeled audio with limited labeled data through pseudo-labeled CPT followed by supervised finetuning. With 20,000 labeled samples, we achieve 3.24% WER on Common Voice Swahili-an 82% relative improvement over the baseline. This result surpasses the best previously reported academic system (8.3% WER from XLS-R) by 61% relative improvement. We provide concrete data requirements and a replicable methodology applicable to other low-resource languages.
【15】Fair-Gate: Fairness-Aware Interpretable Risk Gating for Sex-Fair Voice Biometrics
标题:Fair-Gate:性别公平语音生物识别技术的公平性可解释风险门控
链接:https://arxiv.org/abs/2603.11360
摘要:语音生物识别系统可能会表现出与性别相关的性能差距,即使整体验证精度很高。我们将这些差距归因于两个实际机制:(i)人口统计学捷径学习,其中说话人分类训练利用性别和说话人身份之间的虚假相关性,以及(ii)特征纠缠,其中与性别相关的声学变化与身份线索重叠,并且在不降低说话人歧视的情况下无法去除。我们提出了公平门,一个公平意识和可解释的风险门控框架,在一个单一的管道中解决这两种机制。公平门应用风险外推,以减少代理性别群体的扬声器分类风险的变化,并引入了一个本地的互补门,路由中间功能到一个身份分支和性别分支。该门通过产生一个明确的路由掩码来提供可解释性,可以检查该掩码以了解哪些特征被分配给身份与性别相关的路径。在VoxCeleb 1上的实验表明,Fair-Gate改进了效用-公平性权衡,在具有挑战性的评估条件下产生更多的性别公平ASV性能。
摘要:Voice biometric systems can exhibit sex-related performance gaps even when overall verification accuracy is strong. We attribute these gaps to two practical mechanisms: (i) demographic shortcut learning, where speaker classification training exploits spurious correlations between sex and speaker identity, and (ii) feature entanglement, where sex-linked acoustic variation overlaps with identity cues and cannot be removed without degrading speaker discrimination. We propose Fair-Gate, a fairness-aware and interpretable risk-gating framework that addresses both mechanisms in a single pipeline. Fair-Gate applies risk extrapolation to reduce variation in speaker-classification risk across proxy sex groups, and introduces a local complementary gate that routes intermediate features into an identity branch and a sex branch. The gate provides interpretability by producing an explicit routing mask that can be inspected to understand which features are allocated to identity versus sex-related pathways. Experiments on VoxCeleb1 show that Fair-Gate improves the utility--fairness trade-off, yielding more sex-fair ASV performance under challenging evaluation conditions.
【16】V2A-DPO: Omni-Preference Optimization for Video-to-Audio Generation
标题:V2 A-DPO:视频到音频生成的全偏好优化
链接:https://arxiv.org/abs/2603.11089
备注:Accepted at ICASSP2026
摘要:本文介绍了V2 A-DPO,一种新的直接偏好优化(DPO)框架,专为基于流的视频到音频生成(V2 A)模型量身定制,结合关键的调整,以有效地将生成的音频与人类的偏好相匹配。我们的方法结合了三个核心创新:(1)AudioScore-一个全面的人类偏好对齐评分系统,用于评估合成音频的语义一致性,时间对齐和感知质量;(2)自动化AudioScore驱动的管道,用于生成用于DPO优化的大规模偏好对数据;(3)课程学习授权的DPO优化策略,专门为基于流的生成模型量身定制。在基准VGGSound数据集上的实验表明,使用V2 A-DPO的人类偏好对齐的Frieren和MMAudio优于使用去噪扩散策略优化(DDPO)以及预训练基线优化的同行。此外,我们的DPO优化MMAudio在多个指标上实现了最先进的性能,超过了已发布的V2 A型号。
摘要:This paper introduces V2A-DPO, a novel Direct Preference Optimization (DPO) framework tailored for flow-based video-to-audio generation (V2A) models, incorporating key adaptations to effectively align generated audio with human preferences. Our approach incorporates three core innovations: (1) AudioScore-a comprehensive human preference-aligned scoring system for assessing semantic consistency, temporal alignment, and perceptual quality of synthesized audio; (2) an automated AudioScore-driven pipeline for generating large-scale preference pair data for DPO optimization; (3) a curriculum learning-empowered DPO optimization strategy specifically tailored for flow-based generative models. Experiments on benchmark VGGSound dataset demonstrate that human-preference aligned Frieren and MMAudio using V2A-DPO outperform their counterparts optimized using Denoising Diffusion Policy Optimization (DDPO) as well as pre-trained baselines. Furthermore, our DPO-optimized MMAudio achieves state-of-the-art performance across multiple metrics, surpassing published V2A models.
机器翻译由腾讯交互翻译提供,仅供参考
