今日论文合集:cs.SD语音3篇,eess.AS音频处理2篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Zero-Shot to Zero-Lies: Detecting Bengali Deepfake Audio through Transfer Learning
标题:Zero-Shot to Zero-Lies:通过迁移学习检测孟加拉语Deepfake音频
链接:https://arxiv.org/abs/2512.21702

作者:Most. Sharmin Sultana Samu,Md. Rakibul Islam,Md. Zahid Hossain,Md. Kamrozzaman Bhuiyan,Farhad Uz Zaman
备注:Accepted for publication in 2025 28th International Conference on Computer and Information Technology (ICCIT)
摘要:语音合成和语音转换系统的快速增长使Deepfake音频成为一个主要的安全问题。孟加拉语的deepfake检测在很大程度上仍未被探索。在这项工作中,我们研究了使用BanglaFake数据集自动检测孟加拉语音频deepfake。我们使用几个预训练模型来评估zeroshoot推断。其中包括Wav 2 Vec 2-XLSR-53、Whisper、PANNsCNN 14、WavLM和音频频谱图Transformer。Zero-shot结果显示有限的检测能力。最佳模型Wav 2 Vec 2-XLSR-53的准确率为53.80%,AUC为56.60%,EER为46.20%。然后,我们对孟加拉语deepfake检测的多个架构进行微调。其中包括Wav 2 Vec 2-Base、LCNN、LCNN-Attention、ResNet 18、ViT-B16和CNN-BiLSTM。经过微调的模型显示出强大的性能增益。ResNet 18的准确率最高,为79.17%,F1评分为79.12%,AUC为84.37%,EER为24.35%。实验结果证实,微调显着提高性能超过zero-shot推理。这项研究提供了孟加拉语deepfake音频检测的第一个系统基准。它强调了微调深度学习模型对这种低资源语言的有效性。
摘要:The rapid growth of speech synthesis and voice conversion systems has made deepfake audio a major security concern. Bengali deepfake detection remains largely unexplored. In this work, we study automatic detection of Bengali audio deepfakes using the BanglaFake dataset. We evaluate zeroshot inference with several pretrained models. These include Wav2Vec2-XLSR-53, Whisper, PANNsCNN14, WavLM and Audio Spectrogram Transformer. Zero-shot results show limited detection ability. The best model, Wav2Vec2-XLSR-53, achieves 53.80% accuracy, 56.60% AUC and 46.20% EER. We then f ine-tune multiple architectures for Bengali deepfake detection. These include Wav2Vec2-Base, LCNN, LCNN-Attention, ResNet18, ViT-B16 and CNN-BiLSTM. Fine-tuned models show strong performance gains. ResNet18 achieves the highest accuracy of 79.17%, F1 score of 79.12%, AUC of 84.37% and EER of 24.35%. Experimental results confirm that fine-tuning significantly improves performance over zero-shot inference. This study provides the first systematic benchmark of Bengali deepfake audio detection. It highlights the effectiveness of f ine-tuned deep learning models for this low-resource language.


【2】Semantic Codebooks as Effective Priors for Neural Speech Compression
标题:语义码本作为神经语音压缩的有效先验
链接:https://arxiv.org/abs/2512.21653

作者:Liuyang Bai,Weiyi Lu,Li Guo
摘要:语音编解码器传统上是针对波形保真度进行优化的,分配比特以保留声学细节,即使其中大部分可以从语言结构中推断出来。这导致了低效率的压缩和下游识别任务的次优性能。我们提出SemDAC,语义感知神经音频编解码器,利用语义码本作为有效的语音压缩的先验。在SemDAC中,残差矢量量化(RVQ)堆栈中的第一个量化器从HuBERT特征中提取,以产生捕获语音内容的语义令牌,而后续量化器则对残差声学进行建模。FiLM条件解码器重构基于语义标记的音频,提高了声学码本的使用效率。尽管其简单性,但这种设计被证明是非常有效的:SemDAC在感知指标上优于DAC,并且在重构语音上运行Whisper时实现了较低的WER,所有这些都是在基本上较低的比特率下操作的(例如,0.95 kbps与DAC的2.5 kbps)。这些结果表明,语义码本提供了一个有效的神经语音压缩的归纳偏见,产生紧凑,但简化友好的表示。
摘要:Speech codecs are traditionally optimized for waveform fidelity, allocating bits to preserve acoustic detail even when much of it can be inferred from linguistic structure. This leads to inefficient compression and suboptimal performance on downstream recognition tasks. We propose SemDAC, a semantic-aware neural audio codec that leverages semantic codebooks as effective priors for speech compression. In SemDAC, the first quantizer in a residual vector quantization (RVQ) stack is distilled from HuBERT features to produce semantic tokens that capture phonetic content, while subsequent quantizers model residual acoustics. A FiLM-conditioned decoder reconstructs audio conditioned on the semantic tokens, improving efficiency in the use of acoustic codebooks. Despite its simplicity, this design proves highly effective: SemDAC outperforms DAC across perceptual metrics and achieves lower WER when running Whisper on reconstructed speech, all while operating at substantially lower bitrates (e.g., 0.95 kbps vs. 2.5 kbps for DAC). These results demonstrate that semantic codebooks provide an effective inductive bias for neural speech compression, producing compact yet recognition-friendly representations.


【3】Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
标题:无需语音模型中任务载体微调的稀有词识别和翻译
链接:https://arxiv.org/abs/2512.21894

作者:Ruihao Jing,Cheng Gong,Yu Jiang,Boyu Zhu,Shansong Liu,Chi Zhang,Xiao-Lei Zhang,Xuelong Li
摘要:生僻词仍然是语音到文本系统的关键瓶颈。虽然直接微调提高了目标词的识别,但它通常会导致高成本,灾难性的遗忘和有限的可扩展性。为了应对这些挑战,我们提出了一种基于任务向量的免训练范式,用于稀有词识别和翻译。通过将任务向量定义为参数差,并引入词级任务向量算法,实现了稀有词能力的灵活组合,极大地提高了可扩展性和可重用性。跨多个领域的广泛实验表明,该方法匹配或超越目标词的微调模型,提高约5 BLEU的一般性能,并减轻灾难性遗忘。
摘要:Rare words remain a critical bottleneck for speech-to-text systems. While direct fine-tuning improves recognition of target words, it often incurs high cost, catastrophic forgetting, and limited scalability. To address these challenges, we propose a training-free paradigm based on task vectors for rare word recognition and translation. By defining task vectors as parameter differences and introducing word-level task vector arithmetic, our approach enables flexible composition of rare-word capabilities, greatly enhancing scalability and reusability. Extensive experiments across multiple domains show that the proposed method matches or surpasses fine-tuned models on target words, improves general performance by about 5 BLEU, and mitigates catastrophic forgetting.


eess.AS音频处理


【1】Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
标题:无需语音模型中任务载体微调的稀有词识别和翻译
链接:https://arxiv.org/abs/2512.21894

作者:Ruihao Jing,Cheng Gong,Yu Jiang,Boyu Zhu,Shansong Liu,Chi Zhang,Xiao-Lei Zhang,Xuelong Li
摘要:生僻词仍然是语音到文本系统的关键瓶颈。虽然直接微调提高了目标词的识别,但它通常会导致高成本,灾难性的遗忘和有限的可扩展性。为了应对这些挑战,我们提出了一种基于任务向量的免训练范式,用于稀有词识别和翻译。通过将任务向量定义为参数差,并引入词级任务向量算法,实现了稀有词能力的灵活组合,极大地提高了可扩展性和可重用性。跨多个领域的大量实验表明,该方法匹配或超越目标词的微调模型,提高约5 BLEU的一般性能,并减轻灾难性遗忘。
摘要:Rare words remain a critical bottleneck for speech-to-text systems. While direct fine-tuning improves recognition of target words, it often incurs high cost, catastrophic forgetting, and limited scalability. To address these challenges, we propose a training-free paradigm based on task vectors for rare word recognition and translation. By defining task vectors as parameter differences and introducing word-level task vector arithmetic, our approach enables flexible composition of rare-word capabilities, greatly enhancing scalability and reusability. Extensive experiments across multiple domains show that the proposed method matches or surpasses fine-tuned models on target words, improves general performance by about 5 BLEU, and mitigates catastrophic forgetting.


【2】Contextual Biasing for LLM-Based ASR with Hotword Retrieval and Reinforcement Learning
标题:具有Hotword检索和强化学习的基于LLM的ASB的上下文偏置
链接:https://arxiv.org/abs/2512.21828

作者:YuXiang Kong,JunFeng Hou,Jian Tang,Bingqing Zhu,Jicheng Zhang,Shaofei Xue
摘要:基于大语言模型(LLM)的自动语音识别(ASR)最近在不同的任务中取得了很好的性能,但在大词汇表下命名实体和热词的上下文偏置仍然具有挑战性。在这项工作中,我们提出了一个可扩展的两阶段框架,集成热词检索LLM-ASR适应。首先,我们扩展了全局-局部对比存储-音频预训练模型(GLCLAP),通过鲁棒性感知数据增强和模糊匹配从大词汇表中检索紧凑的top-k热词候选集。其次,我们将检索到的候选项作为文本提示注入LLM-ASR模型,并使用基于生成拒绝的策略优化(GRPO)对其进行微调,使用任务驱动的奖励,共同优化热词识别和整体转录准确性。热词为重点的测试集上的实验表明,大量的关键字错误率(KER)的减少,同时保持句子的准确性一般ASR基准,证明了所提出的框架的有效性,大词汇量的上下文偏置。
摘要:Large language model (LLM)-based automatic speech recognition (ASR) has recently achieved strong performance across diverse tasks, yet contextual biasing for named entities and hotwords under large vocabularies remains challenging. In this work, we propose a scalable two-stage framework that integrates hotword retrieval with LLM-ASR adaptation. First, we extend the Global-Local Contrastive Language-Audio pre-trained model (GLCLAP) to retrieve a compact top-k set of hotword candidates from a large vocabulary via robustness-aware data augmentation and fuzzy matching. Second, we inject the retrieved candidates as textual prompts into an LLM-ASR model and fine-tune it with Generative Rejection-Based Policy Optimization (GRPO), using a task-driven reward that jointly optimizes hotword recognition and overall transcription accuracy. Experiments on hotword-focused test sets show substantial keyword error rate (KER) reductions while maintaining sentence accuracy on general ASR benchmarks, demonstrating the effectiveness of the proposed framework for large-vocabulary contextual biasing.


机器翻译由腾讯交互翻译提供,仅供参考