今日论文合集:cs.SD语音4篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】AudioKV: KV Cache Eviction in Efficient Large Audio Language Models
标题:AudioKV:高效大型音频语言模型中的KV缓存驱逐
链接:https://arxiv.org/abs/2604.06694

作者:Yuxuan Wang,Peize He,Xiyan Gui,Xiaoqian Liu,Junhao He,Xuyang Liu,Zichen Wen,Xuming Hu,Linfeng Zhang
摘要:大型音频语言模型(LALM)已经在语音处理中建立了新的基准,但它们的部署受到长上下文推理期间键值(KV)缓存的内存占用的阻碍。虽然一般KV高速缓存压缩技术在LLM中表现出色,但它们通常由于忽略了声学信号的固有时间连续性而在音频域中失败。为了弥合这一差距,我们提出了AudioKV,一个新的框架,通过硬件友好的语义声学对齐机制,强大的优先级音频关键注意头。具体来说,我们确定这些模态专用头通过分析注意力分数在ASR任务和动态分配KV缓存预算优先给他们。此外,我们还引入了频谱分数平滑(SSS),这是一种基于FFT的全局过滤策略,旨在抑制高频噪声并从重要性分数中恢复平滑的全局趋势,从而确保以前所未有的精度进行更平衡的令牌选择。对多个LALM(包括Qwen和Gemma系列)的广泛评估表明,AudioKV在提高计算效率的同时显著优于基线。值得注意的是,在40%的压缩比下,AudioKV在Qwen 3-Omni-30 B上保持了接近全精度,只有0.45%的下降,而传统方法遭受灾难性的性能下降和重复。我们的代码将在验收后发布。
摘要:Large Audio-Language Models (LALMs) have set new benchmarks in speech processing, yet their deployment is hindered by the memory footprint of the Key-Value (KV) cache during long-context inference. While general KV cache compression techniques excel in LLMs, they often fail in the audio domain by overlooking the intrinsic temporal continuity of acoustic signals. To bridge this gap, we propose AudioKV, a novel framework that robustly prioritizes audio-critical attention heads through a hardware-friendly semantic-acoustic alignment mechanism. Specifically, we identify these modality-specialized heads by analyzing attention scores in ASR tasks and dynamically allocate KV cache budgets preferentially to them. Furthermore, we introduce Spectral Score Smoothing (SSS), an FFT-based global filtering strategy designed to suppress high-frequency noise and recover smooth global trends from importance scores, ensuring more balanced token selection with unprecedented precision. Extensive evaluations across multiple LALMs, including Qwen and Gemma series, demonstrate that AudioKV significantly outperforms baselines while enhancing computational efficiency. Notably, at a 40% compression ratio, AudioKV maintains near-full accuracy on Qwen3-Omni-30B with only a 0.45% drop, whereas traditional methods suffer from catastrophic performance degradation and repetition. Our code will be released after acceptance.


【2】A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
标题:合成语音中说话人漂移检测的新型自动框架
链接:https://arxiv.org/abs/2604.06327

作者:Jia-Hong Huang,Seulgi Kim,Yi Chieh Liu,Yixian Shen,Hongyi Zhu,Prayag Tiwari,Stevan Rudinac,Evangelos Kanoulas
备注:The paper has been accepted by the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
摘要:最近的扩散为基础的文本到语音(TTS)模型实现了高的自然度和表现力,但往往遭受扬声器漂移,一个微妙的,逐渐转变的感知扬声器身份在一个单一的话语。这种未被充分研究的现象破坏了合成语音的连贯性,特别是在长格式或交互式环境中。我们介绍了第一个自动检测扬声器漂移的框架,制定它作为一个二进制分类任务的话语级扬声器的一致性。我们的方法计算合成语音重叠段之间的余弦相似性,并提示大型语言模型(LLM)与结构化表示评估漂移。我们提供了理论保证余弦为基础的漂移检测,并证明扬声器嵌入表现出有意义的几何聚类的单位球。为了支持评估,我们构建了一个高质量的合成基准与人类验证的扬声器漂移注释。多个国家的最先进的LLM的实验证实了这种嵌入到推理流水线的可行性。我们的工作建立扬声器漂移作为一个独立的研究问题和桥梁的几何信号分析与LLM为基础的感知推理在现代TTS。
摘要:Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker identity within a single utterance. This underexplored phenomenon undermines the coherence of synthetic speech, especially in long-form or interactive settings. We introduce the first automatic framework for detecting speaker drift by formulating it as a binary classification task over utterance-level speaker consistency. Our method computes cosine similarity across overlapping segments of synthesized speech and prompts large language models (LLMs) with structured representations to assess drift. We provide theoretical guarantees for cosine-based drift detection and demonstrate that speaker embeddings exhibit meaningful geometric clustering on the unit sphere. To support evaluation, we construct a high-quality synthetic benchmark with human-validated speaker drift annotations. Experiments with multiple state-of-the-art LLMs confirm the viability of this embedding-to-reasoning pipeline. Our work establishes speaker drift as a standalone research problem and bridges geometric signal analysis with LLM-based perceptual reasoning in modern TTS.


【3】Development of ML model for triboelectric nanogenerator based sign language detection system
标题:基于摩擦电纳米发电机的手语检测系统ML模型的开发
链接:https://arxiv.org/abs/2604.06220

作者:Meshv Patel,Bikash Baro,Sayan Bayan,Mohendra Roy
备注:This paper has been accepted at the IEEE GCON 2026 (https://gcon2026.in/) Conference, organized by IIT Guwahati
摘要:手语识别(SLR)对于弥合聋人和听力社区之间的沟通差距至关重要。基于视觉的方法受到遮挡、计算成本和物理约束的影响。这项工作提出了一种基于自定义摩擦电纳米发电机(TENG)的传感器手套的机器学习(ML)和深度学习模型的比较。该研究利用来自五个Flex传感器的多变量时间序列数据,对传统ML算法、前馈神经网络、基于LSTM的时间模型以及跨11个符号类(数字1-5,字母A-F)的多传感器MFCC CNN-LSTM架构进行基准测试。所提出的MFCC CNN-LSTM架构在融合之前通过独立的卷积分支处理来自每个传感器的频域特征。它实现了93.33%的准确率和95.56%的精度,比最好的ML算法(随机森林:70.38%)提高了23个点。消融研究表明,50时间步长窗口提供了时间背景和训练数据量之间的权衡,与100时间步长窗口的58.06%相比,准确率为84.13%。MFCC特征提取将时间变化映射到执行速度不变的频谱表示,并且数据增强方法(时间扭曲,噪声注入)对于泛化是必不可少的。结果表明,频域特征表示与并行多传感器处理架构相结合,为基于传感器的可穿戴手势识别提供了优于经典算法和时域深度学习的增强。这有助于辅助技术的发展。
摘要:Sign language recognition (SLR) is vital for bridging communication gaps between deaf and hearing communities. Vision-based approaches suffer from occlusion, computational costs, and physical constraints. This work presents a comparison of machine learning (ML) and deep learning models for a custom triboelectric nanogenerator (TENG)-based sensor glove. Utilizing multivariate time-series data from five flex sensors, the study benchmarks traditional ML algorithms, feedforward neural networks, LSTM-based temporal models, and a multi-sensor MFCC CNN-LSTM architecture across 11 sign classes (digits 1-5, letters A-F). The proposed MFCC CNN-LSTM architecture processes frequency-domain features from each sensor through independent convolutional branches before fusion. It achieves 93.33% accuracy and 95.56% precision, a 23-point improvement over the best ML algorithm (Random Forest: 70.38%). Ablation studies reveal 50-timestep windows offer a tradeoff between temporal context and training data volume, yielding 84.13% accuracy compared to 58.06% with 100-timestep windows. MFCC feature extraction maps temporal variations to execution-speed-invariant spectral representations, and data augmentation methods (time warping, noise injection) are essential for generalization. Results demonstrate that frequency-domain feature representations combined with parallel multi-sensor processing architectures offer enhancement over classical algorithms and time-domain deep learning for wearable sensor-based gesture recognition. This aids assistive technology development.


【4】Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
标题:半言语:阿拉伯音素级言语评估的临床一致框架
链接:https://arxiv.org/abs/2604.06191

作者:Asif Azad,MD Sadik Hossain Shanto,Mohammad Sadat Hossain,Bdour Alwuqaysi,Sabri Boughorbel,Yahya Bokhari,Abdulrhman Aljouie,Ayah Othman Sindi,Ehsan Hoque
摘要:自动音素级发音评估对于可扩展的语音治疗和语言学习至关重要,但阿拉伯语的有效工具仍然很少。我们提出了哈夫语音,一个模块化的系统评分阿拉伯语发音在音素水平上的临床规模。它结合了一个MSA发音器,微调语音音素模型,Levenshtein对齐,并使用最长的共同子序列和编辑距离度量混合评分。我们微调三个ASR架构的阿拉伯语音素数据和基准测试他们与zero-shot多模态模型;最好的,OmniASR-CTC-1B-v2,达到8.92%的音素错误率。三位经过认证的言语语言病理学家独立地对40个话语进行了临床验证。Harf-Speech与平均专家评分的Pearson相关系数为0.791,ICC(2,1)为0.659,优于现有的端到端评估框架。这些结果表明,Harf-Speech产生的临床一致、可解释的评分与评分者间专家一致性相当。
摘要:Automated phoneme-level pronunciation assessment is vital for scalable speech therapy and language learning, yet validated tools for Arabic remain scarce. We present Harf-Speech, a modular system scoring Arabic pronunciation at the phoneme level on a clinical scale. It combines an MSA phonetizer, a fine-tuned speech-to-phoneme model, Levenshtein alignment, and a blended scorer using longest common subsequence and edit-distance metrics. We fine-tune three ASR architectures on Arabic phoneme data and benchmark them with zero-shot multimodal models; the best, OmniASR-CTC-1B-v2, achieves 8.92\% phoneme error rate. Three certified speech-language pathologists independently scored 40 utterances for clinical validation. Harf-Speech attains a Pearson correlation of 0.791 and ICC(2,1) of 0.659 with mean expert scores, outperforming existing end-to-end assessment frameworks. These results show Harf-Speech yields clinically aligned, interpretable scores comparable to inter-rater expert agreement.


eess.AS音频处理


【1】EvoTSE: Evolving Enrollment for Target Speaker Extraction
标题:EvoPSE:不断发展的目标说话人提取注册
链接:https://arxiv.org/abs/2604.06810

作者:Zikai Liu,Ziqian Wang,Xingchen Li,Yike Zhu,Shuai Wang,Longshuai Xiao,Lei Xie
摘要:目标说话人提取(TSE)的目的是分离出一个特定的说话人的声音从混合,由预先录制的注册指导。虽然TSE绕过了盲源分离的全局置换模糊性,但它仍然容易受到说话人混淆的影响,其中模型错误地提取了干扰说话人。此外,常规TSE依赖于静态推理流水线,其中性能受到固定登记的质量的限制。为了克服这些限制,我们提出了EvoTSE,一个不断发展的TSE框架,在该框架中,通过可靠性过滤检索高置信度的历史估计,不断更新注册。这种机制减少了说话人混淆,并放宽了对预先录制的注册的质量要求,而不依赖于额外的注释数据。跨多个基准测试的实验表明,EvoTSE实现了一致的改进,特别是在域外(OOD)场景下进行评估时。我们的代码和检查点都是可用的。
摘要:Target Speaker Extraction (TSE) aims to isolate a specific speaker's voice from a mixture, guided by a pre-recorded enrollment. While TSE bypasses the global permutation ambiguity of blind source separation, it remains vulnerable to speaker confusion, where models mistakenly extract the interfering speaker. Furthermore, conventional TSE relies on static inference pipeline, where performance is limited by the quality of the fixed enrollment. To overcome these limitations, we propose EvoTSE, an evolving TSE framework in which the enrollment is continuously updated through reliability-filtered retrieval over high-confidence historical estimates. This mechanism reduces speaker confusion and relaxes the quality requirements for pre-recorded enrollment without relying on additional annotated data. Experiments across multiple benchmarks demonstrate that EvoTSE achieves consistent improvements, especially when evaluated on out-of-domain (OOD) scenarios. Our code and checkpoints are available.


【2】DAT-CFTNet: Speech Enhancement for Cochlear Implant Recipients using Attention-based Dual-Path Recurrent Neural Network
标题:DAT-CFTNet:使用基于注意力的双路径递归神经网络对耳蜗植入者进行语音增强
链接:https://arxiv.org/abs/2604.06744

作者:Nursadul Mamun,John H. L. Hansen
备注:5 pages
摘要:人类听觉系统具有选择性地聚焦于音频流中的关键语音元素的能力,同时对背景内的诸如噪声或失真之类的不太相关的区域给予次要关注,从而随时间动态地调整其注意力。受近年来注意力模型研究的启发,本文在并发语音增强网络的瓶颈层引入了双路径注意力模型。我们的研究提出了一种基于注意力的双路径RNN(DAT-RNN),当它与改进的复值频率变换网络(CFTNet)相结合时,形成了DAT-CFTNet。这种注意力机制允许在频谱图的时频(T-F)区域中精确区分语音和噪声,优化CFTNet中的局部和全局上下文信息处理。我们的实验表明,DAT-CFTNet在语音清晰度和质量方面,与现有模型(包括CFTNet和DCCRN)相比,性能得到了持续改善。此外,所提出的模型在增强人工耳蜗(CI)接受者的语音可懂度方面表现出优异的性能,已知人工耳蜗接受者具有严重有限的T-F听力恢复(例如,>10%)表明,该方法能够有效抑制非平稳噪声,避免了传统语音增强方法中常见的音乐伪影。拟议模式的实施情况将向公众公布。
摘要:The human auditory system has the ability to selectively focus on key speech elements in an audio stream while giving secondary attention to less relevant areas such as noise or distortion within the background, dynamically adjusting its attention over time. Inspired by the recent success of attention models, this study introduces a dual-path attention module in the bottleneck layer of a concurrent speech enhancement network. Our study proposes an attention-based dual-path RNN (DAT-RNN), which, when combined with the modified complex-valued frequency transformation network (CFTNet), forms the DAT-CFTNet. This attention mechanism allows for precise differentiation between speech and noise in time-frequency (T-F) regions of spectrograms, optimizing both local and global context information processing in the CFTNet. Our experiments suggest that the DAT-CFTNet leads to consistently improved performance over the existing models, including CFTNet and DCCRN, in terms of speech intelligibility and quality. Moreover, the proposed model exhibits superior performance in enhancing speech intelligibility for cochlear implant (CI) recipients, who are known to have severely limited T-F hearing restoration (e.g., >10%) in CI listener studies in noisy settings show the proposed solution is capable of suppressing non-stationary noise, avoiding the musical artifacts often seen in traditional speech enhancement methods. The implementation of the proposed model will be publicly available.


【3】ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals
标题:ULTRAS --音频和语音信号转换器表示的统一学习
链接:https://arxiv.org/abs/2604.06702

作者:Ameenudeen P E,Charumathi Narayanan,Sriram Ganapathy
摘要:自监督学习(SSL)通过采用时域预测目标,在语音处理方面取得了令人印象深刻的进步,而音频表示学习框架则在时间-频率频谱图上运行。为一种范式优化的模式很难转移到另一种范式,这突出表明需要一个联合框架。我们提出了音频和语音Transformer表示的统一学习(ULTRAS),其中在长数据块上执行掩蔽和预测建模。该模型基于Transformer架构,对对数梅尔谱图特征的谱块进行编码。掩蔽段的预测建模是使用组合损失函数对频谱和时间目标进行的,迫使表示对时间和频率特性进行编码。在各种语音和音频任务上进行实验,我们说明了ULTRAS框架在其他已建立的基线上实现了更好的性能。
摘要:Self-supervised learning (SSL) has driven impressive advances in speech processing by adopting time-domain prediction objectives, while audio representation learning frameworks operate on time-frequency spectrograms. Models optimized for one paradigm struggle to transfer to the other, highlighting the need for a joint framework. We propose Unified Learning of Transformer Representations for Audio and Speech (ULTRAS), where the masking and predictive modeling is performed over long patches of the data. The model, based on the transformer architecture, encodes spectral-patches of log-mel spectrogram features. The predictive modeling of masked segments is performed on spectral and temporal targets using a combined loss-function, forcing the representations to encode time and frequency traits. Experiments are performed on a variety of speech and audio tasks, where we illustrate that the ULTRAS framework achieves improved performance over other established baselines.


【4】Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
标题:半言语:阿拉伯音素级言语评估的临床一致框架
链接:https://arxiv.org/abs/2604.06191

作者:Asif Azad,MD Sadik Hossain Shanto,Mohammad Sadat Hossain,Bdour Alwuqaysi,Sabri Boughorbel,Yahya Bokhari,Abdulrhman Aljouie,Ayah Othman Sindi,Ehsan Hoque
摘要:自动音素级发音评估对于可扩展的语音治疗和语言学习至关重要,但阿拉伯语的有效工具仍然很少。我们提出了哈夫语音,一个模块化的系统评分阿拉伯语发音在音素水平上的临床规模。它结合了一个MSA发音器,微调语音音素模型,Levenshtein对齐,并使用最长的共同子序列和编辑距离度量混合评分。我们微调三个ASR架构的阿拉伯语音素数据和基准测试他们与zero-shot多模态模型;最好的,OmniASR-CTC-1B-v2,达到8.92%的音素错误率。三位经过认证的言语语言病理学家独立地对40个话语进行了临床验证。Harf-Speech与平均专家评分的Pearson相关系数为0.791,ICC(2,1)为0.659,优于现有的端到端评估框架。这些结果表明,Harf-Speech产生的临床一致、可解释的评分与评分者间专家一致性相当。
摘要:Automated phoneme-level pronunciation assessment is vital for scalable speech therapy and language learning, yet validated tools for Arabic remain scarce. We present Harf-Speech, a modular system scoring Arabic pronunciation at the phoneme level on a clinical scale. It combines an MSA phonetizer, a fine-tuned speech-to-phoneme model, Levenshtein alignment, and a blended scorer using longest common subsequence and edit-distance metrics. We fine-tune three ASR architectures on Arabic phoneme data and benchmark them with zero-shot multimodal models; the best, OmniASR-CTC-1B-v2, achieves 8.92\% phoneme error rate. Three certified speech-language pathologists independently scored 40 utterances for clinical validation. Harf-Speech attains a Pearson correlation of 0.791 and ICC(2,1) of 0.659 with mean expert scores, outperforming existing end-to-end assessment frameworks. These results show Harf-Speech yields clinically aligned, interpretable scores comparable to inter-rater expert agreement.


机器翻译由腾讯交互翻译提供,仅供参考