今日论文合集:cs.SD语音11篇,eess.AS音频处理12篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Multi-level SSL Feature Gating for Audio Deepfake Detection
标题:用于音频Deepfake检测的多层SSL功能门控
链接:https://arxiv.org/abs/2509.03409

作者:Hoan My Tran, Damien Lolive, Aghilas Sini, Arnaud Delhay, Pierre-François Marteau, David Guennec
备注:This paper has been accepted by ACM MM 2025
摘要:生成人工智能的最新进展,特别是在语音合成方面,已经能够生成高度自然的合成语音,这些语音非常接近人类的声音。虽然这些创新为辅助技术等应用带来了希望,但它们也带来了重大风险,包括滥用欺诈活动,身份盗窃和安全威胁。目前对欺骗检测对策的研究仍然局限于对看不见的deepfake攻击和语言的推广。为了解决这个问题,我们提出了一个门控机制提取相关功能的语音基础XLS-R模型作为前端特征提取器。对于下游后端分类器,我们采用多核门控卷积(MultiConv)来捕获局部和全局语音伪影。此外,我们引入了中心核对齐(CKA)作为相似性度量,以加强不同MultiConv层中学习特征的多样性。通过将CKA与我们的门控机制相结合,我们假设每个组件都有助于改善不同合成语音模式的学习。实验结果表明,我们的方法实现了最先进的性能在域内的基准,而鲁棒地推广到域外的数据集,包括多语言语音样本。这凸显了它作为检测不断发展的语音Deepfake威胁的通用解决方案的潜力。
摘要:Recent advancements in generative AI, particularly in speech synthesis, have enabled the generation of highly natural-sounding synthetic speech that closely mimics human voices. While these innovations hold promise for applications like assistive technologies, they also pose significant risks, including misuse for fraudulent activities, identity theft, and security threats. Current research on spoofing detection countermeasures remains limited by generalization to unseen deepfake attacks and languages. To address this, we propose a gating mechanism extracting relevant feature from the speech foundation XLS-R model as a front-end feature extractor. For downstream back-end classifier, we employ Multi-kernel gated Convolution (MultiConv) to capture both local and global speech artifacts. Additionally, we introduce Centered Kernel Alignment (CKA) as a similarity metric to enforce diversity in learned features across different MultiConv layers. By integrating CKA with our gating mechanism, we hypothesize that each component helps improving the learning of distinct synthetic speech patterns. Experimental results demonstrate that our approach achieves state-of-the-art performance on in-domain benchmarks while generalizing robustly to out-of-domain datasets, including multilingual speech samples. This underscores its potential as a versatile solution for detecting evolving speech deepfake threats.


【2】Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge
标题:NOCASA 2025挑战赛的端到端语音评估模型比较
链接:https://arxiv.org/abs/2509.03256

作者:Aleksei Zavoronkov, Tanel Alumäe
备注:Published at IEEE MLSP 2025
摘要:本文分析了为NOCASA 2025挑战赛开发的三个端到端模型,旨在为学习挪威语作为第二语言的儿童进行单词级发音自动评估。我们的模型包括一个编码器-解码器连体架构(E2 E-R),一个前缀调整的直接分类模型,利用预训练的wav2vec2.0表示,以及一个新的模型,集成了通过CTC计算的无干扰的良好发音(GOP)功能。我们引入了一个加权序数交叉熵损失量身定制的优化指标,如未加权平均召回率和平均绝对误差。在探索的方法中,我们基于GOP-CTC的模型实现了最高的性能,大大超过了挑战基线,并获得了最高的排行榜分数。
摘要:This paper presents an analysis of three end-to-end models developed for the NOCASA 2025 Challenge, aimed at automatic word-level pronunciation assessment for children learning Norwegian as a second language. Our models include an encoder-decoder Siamese architecture (E2E-R), a prefix-tuned direct classification model leveraging pretrained wav2vec2.0 representations, and a novel model integrating alignment-free goodness-of-pronunciation (GOP) features computed via CTC. We introduce a weighted ordinal cross-entropy loss tailored for optimizing metrics such as unweighted average recall and mean absolute error. Among the explored methods, our GOP-CTC-based model achieved the highest performance, substantially surpassing challenge baselines and attaining top leaderboard scores.


【3】Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
标题:Speech DF Arena:Speech DeepFake检测模型的排行榜
链接:https://arxiv.org/abs/2509.02859

作者:Sandipana Dowerah, Atharva Kulkarni, Ajinkya Kulkarni, Hoan My Tran, Joonas Kalda, Artem Fedorchenko, Benoit Fauve, Damien Lolive, Tanel Alumäe, Matthew Magimai Doss
摘要:与高级deepfake音频生成的发展并行,音频deepfake检测也取得了重大进展。然而,仍然缺少一个标准化和全面的基准。为了解决这个问题,我们引入了Speech DeepFake(DF)Arena,这是音频deepfake检测的第一个综合基准。Speech DF Arena提供了一个工具包来统一评估检测系统,目前涵盖14个不同的数据集和攻击场景,标准化的评估指标和协议,以实现可重复性和透明度。它还包括一个排行榜来比较和排名系统,以帮助研究人员和开发人员提高其可靠性和鲁棒性。我们包括14个评估集,12个最先进的开源和3个专有检测系统。我们的研究提出了许多系统表现出高EER域外的情况下,强调需要广泛的跨域评估。排行榜托管在Huggingface1上,GitHub上提供了一个工具包,用于在所列数据集上复制结果。
摘要:Parallel to the development of advanced deepfake audio generation, audio deepfake detection has also seen significant progress. However, a standardized and comprehensive benchmark is still missing. To address this, we introduce Speech DeepFake (DF) Arena, the first comprehensive benchmark for audio deepfake detection. Speech DF Arena provides a toolkit to uniformly evaluate detection systems, currently across 14 diverse datasets and attack scenarios, standardized evaluation metrics and protocols for reproducibility and transparency. It also includes a leaderboard to compare and rank the systems to help researchers and developers enhance their reliability and robustness. We include 14 evaluation sets, 12 state-of-the-art open-source and 3 proprietary detection systems. Our study presents many systems exhibiting high EER in out-of-domain scenarios, highlighting the need for extensive cross-domain evaluation. The leaderboard is hosted on Huggingface1 and a toolkit for reproducing results across the listed datasets is available on GitHub.


【4】Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission
标题:神经音频编解码器传输的说话人验证性能权衡分析
链接:https://arxiv.org/abs/2509.02771

作者:Nirmalya Mallick Thakur, Jia Qi Yip, Eng Siong Chng
备注:Accepted by APSIPA ASC 2025
摘要:神经音频编解码器(NAC)近年来取得了重大进展,并迅速被许多音频处理管道采用。然而,它们会引入音频失真,从而降低说话人验证(SV)性能。这项研究调查了传统和神经音频编解码器在不同比特率下对VoxCeleb 1数据集上评估的三种最先进的SV模型的影响。我们的研究结果显示,随着比特率的降低,所有模型和编解码器的SV性能都出现了一致的下降。值得注意的是,与传统编解码器相比,NAC不会从根本上破坏SV性能。它们在低比特率(< 12 kbps)下的表现优于Opus 6-8%,在高比特率(约24 kbps)下仍略落后,EER仅增加0.4- 0.7%。较高比特率下的差异可能是由于NAC对感知质量的主要优化,这可能会无意中丢弃关键的说话者区分特征,而Opus则旨在保留声音特征。我们的研究表明,NAC是一个可行的替代传统的编解码器,特别是在带宽限制。为了弥补更高比特率的差距,未来的工作应该集中在开发说话者感知的NAC或重新训练和适应SV模型上。
摘要:Neural audio codecs (NACs) have made significant advancements in recent years and are rapidly being adopted in many audio processing pipelines. However, they can introduce audio distortions which degrade speaker verification (SV) performance. This study investigates the impact of both traditional and neural audio codecs at varying bitrates on three state of-the-art SV models evaluated on the VoxCeleb1 dataset. Our findings reveal a consistent degradation in SV performance across all models and codecs as bitrates decrease. Notably, NACs do not fundamentally break SV performance when compared to traditional codecs. They outperform Opus by 6-8% at low-bitrates (< 12 kbps) and remain marginally behind at higher bitrates ($\approx$ 24 kbps), with an EER increase of only 0.4-0.7%. The disparity at higher bitrates is likely due to the primary optimization of NACs for perceptual quality, which can inadvertently discard critical speaker-discriminative features, unlike Opus which was designed to preserve vocal characteristics. Our investigation suggests that NACs are a feasible alternative to traditional codecs, especially under bandwidth limitations. To bridge the gap at higher bitrates, future work should focus on developing speaker-aware NACs or retraining and adapting SV models.


【5】An Effective Strategy for Modeling Score Ordinality and Non-uniform Intervals in Automated Speaking Assessment
标题:自动口语评估中得分有序度和非均匀间隔建模的有效策略
链接:https://arxiv.org/abs/2509.03372

作者:Tien-Hong Lo, Szu-Yu Chen, Yao-Ting Sung, Berlin Chen
备注:Accepted at ASRU 2025
摘要:最近一系列关于自动口语评估(ASA)的研究受益于自监督学习(SSL)表示,它可以在没有特征策展的基础假设的情况下捕获非母语语音中丰富的声学和语言模式。然而,基于语音的SSL模型捕获声学相关的特征,但忽略了语言内容,而基于文本的SSL模型依赖于ASR输出,无法编码韵律的细微差别。此外,大多数现有技术将熟练程度水平视为名义类别,忽略了它们的顺序结构和熟练程度标签之间的非均匀间隔。为了解决这些限制,我们提出了一个有效的ASA方法相结合的SSL手工制作的指标功能,通过一个新的建模范例。我们进一步引入了一个多利润率序数损失,共同模型的得分序数和非均匀间隔的熟练度标签。在TEEMI语料库上进行的大量实验表明,我们的方法始终优于强基线,并很好地推广到看不见的提示。
摘要:A recent line of research on automated speaking assessment (ASA) has benefited from self-supervised learning (SSL) representations, which capture rich acoustic and linguistic patterns in non-native speech without underlying assumptions of feature curation. However, speech-based SSL models capture acoustic-related traits but overlook linguistic content, while text-based SSL models rely on ASR output and fail to encode prosodic nuances. Moreover, most prior arts treat proficiency levels as nominal classes, ignoring their ordinal structure and non-uniform intervals between proficiency labels. To address these limitations, we propose an effective ASA approach combining SSL with handcrafted indicator features via a novel modeling paradigm. We further introduce a multi-margin ordinal loss that jointly models both the score ordinality and non-uniform intervals of proficiency labels. Extensive experiments on the TEEMI corpus show that our method consistently outperforms strong baselines and generalizes well to unseen prompts.


【6】Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
标题:通过三重缺失和自我监督嵌入改善知觉音频审美评估
链接:https://arxiv.org/abs/2509.03292

作者:Dyah A. M. G. Wisnu. Wisnu, Ryandhimas E. Zezario, Stefano Rini, Hsin-Min Wang, Yu Tsao
备注:Accepted by IEEE Automatic Speech Recognition and Understanding Workshop(ASRU), 2025
摘要:我们提出了一个用于生成音频的自动多轴感知质量预测的系统,该系统是为AudioMOS Challenge 2025的Track 2开发的。该任务是预测四个音频美学分数-生产质量,生产复杂性,内容享受和内容丰富性-由文本到语音(TTS),文本到音频(TTA)和文本到音乐(TTM)系统生成的音频。一个主要的挑战是自然训练数据和综合评估数据之间的域转移。为了解决这个问题,我们将BEAT(一种预训练的基于变换的音频表示模型)与多分支长短期记忆(LSTM)预测器相结合,并使用基于缓冲区的采样的三重丢失来通过感知相似性来构建嵌入空间。我们的研究结果表明,这提高了嵌入的可辨别性和泛化性,使域鲁棒的音频质量评估没有合成的训练数据。
摘要:We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness--for audio generated by text-to-speech (TTS), text-to-audio (TTA), and text-to-music (TTM) systems. A main challenge is the domain shift between natural training data and synthetic evaluation data. To address this, we combine BEATs, a pretrained transformer-based audio representation model, with a multi-branch long short-term memory (LSTM) predictor and use a triplet loss with buffer-based sampling to structure the embedding space by perceptual similarity. Our results show that this improves embedding discriminability and generalization, enabling domain-robust audio quality assessment without synthetic training data.


【7】A Study on Zero-Shot Non-Intrusive Speech Intelligibility for Hearing Aids Using Large Language Models
标题:使用大语言模型的助听器Zero-Shot非侵入性语音可理解性研究
链接:https://arxiv.org/abs/2509.03021

作者:Ryandhimas E. Zezario, Dyah A.M.G. Wisnu, Hsin-Min Wang, Yu Tsao
备注:Accepted to IEEE ICCE-TW 2025
摘要:本工作的重点是zero-shot非侵入性的语音评估助听器(HA)使用大语言模型(LLM)。具体来说,我们介绍了GPT-Whisper-HA,GPT-Whisper的扩展,zero-shot非侵入性语音评估模型的基础上LLM。GPT-Whisper-HA是专为HA的语音评估而设计的,它结合了MSBG听力损失和NAL-R模拟,以根据每个人的听力图处理音频输入,两个自动语音识别(ASR)模块用于音频到文本的表示,GPT-4 o用于预测两个相应的分数,然后对最终估计分数进行分数平均。实验结果表明,GPT-Whisper-HA实现了2.59%的相对均方根误差(RMSE)的改善GPT-Whisper,确认潜在的LLM的zero-shot语音评估预测HA用户的主观可懂度。
摘要:This work focuses on zero-shot non-intrusive speech assessment for hearing aids (HA) using large language models (LLMs). Specifically, we introduce GPT-Whisper-HA, an extension of GPT-Whisper, a zero-shot non-intrusive speech assessment model based on LLMs. GPT-Whisper-HA is designed for speech assessment for HA, incorporating MSBG hearing loss and NAL-R simulations to process audio input based on each individual's audiogram, two automatic speech recognition (ASR) modules for audio-to-text representation, and GPT-4o to predict two corresponding scores, followed by score averaging for the final estimated score. Experimental results indicate that GPT-Whisper-HA achieves a 2.59% relative root mean square error (RMSE) improvement over GPT-Whisper, confirming the potential of LLMs for zero-shot speech assessment in predicting subjective intelligibility for HA users.


【8】Non-Intrusive Intelligibility Prediction for Hearing Aids: Recent Advances, Trends, and Challenges
标题:助听器的非侵入性可认知度预测:最新进展、趋势和挑战
链接:https://arxiv.org/abs/2509.03017

作者:Ryandhimas E. Zezario
备注:APSIPA ASC 2025 perspective paper
摘要:本文综述了助听器(HA)的非侵入性语言清晰度预测的最新进展。我们总结了强大的声学特征提取,听力损失建模,并使用新兴的架构长序列处理的发展。还讨论了特定于听众的适应策略和领域泛化方法,旨在提高在看不见的声学环境中的鲁棒性。认识到仍然存在的挑战,例如需要大规模、多样化的数据集和可靠的跨剖面归纳。我们的目标是提供对当前趋势,持续的挑战,以及未来可能的方向,以实用和可靠的HA为导向的可懂度预测系统的观点。
摘要:This paper provides an overview of recent progress in non-intrusive speech intelligibility prediction for hearing aids (HA). We summarize developments in robust acoustic feature extraction, hearing loss modeling, and the use of emerging architectures for long-sequence processing. Listener-specific adaptation strategies and domain generalization approaches that aim to improve robustness in unseen acoustic environments are also discussed. Remaining challenges, such as the need for large-scale, diverse datasets and reliable cross-profile generalization, are acknowledged. Our goal is to offer a perspective on current trends, ongoing challenges, and possible future directions toward practical and reliable HA-oriented intelligibility prediction systems.


【9】Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
标题:使用不确定性Whisper嵌入和sLSTM的语音可理解性评估
链接:https://arxiv.org/abs/2509.03013

作者:Ryandhimas E. Zezario, Dyah A.M.G. Wisnu, Hsin-Min Wang, Yu Tsao
备注:Accepted to APSIPA ASC 2025
摘要:由于说话者、噪声条件和主观感知的变化,非侵入式语音可懂度预测仍然具有挑战性。我们提出了一种不确定性感知方法,该方法利用Whisper嵌入与统计特征相结合,特别是在嵌入维度上计算的平均值,标准差和熵。通过softmax在特征维度上计算的熵作为不确定性的代理,补充了由平均值和标准差捕获的全局信息。为了对语音的序列结构进行建模,我们采用了标量长短期记忆(sLSTM)网络,它可以有效地捕获长程依赖关系。在此基础上,我们提出了iMTI-Net,这是一种改进的多目标可懂度预测网络,它在多任务学习框架内集成了卷积神经网络(CNN)和sLSTM组件。它联合预测来自Google ASR和Whisper的人类可懂度分数和基于机器的单词错误率(WER)。实验结果表明,iMTI-Net在多个评估指标上优于原始MTI-Net,证明了结合不确定性感知特征和CNN-sLSTM架构的有效性。
摘要:Non-intrusive speech intelligibility prediction remains challenging due to variability in speakers, noise conditions, and subjective perception. We propose an uncertainty-aware approach that leverages Whisper embeddings in combination with statistical features, specifically the mean, standard deviation, and entropy computed across the embedding dimensions. The entropy, computed via a softmax over the feature dimension, serves as a proxy for uncertainty, complementing global information captured by the mean and standard deviation. To model the sequential structure of speech, we adopt a scalar long short-term memory (sLSTM) network, which efficiently captures long-range dependencies. Building on this foundation, we propose iMTI-Net, an improved multi-target intelligibility prediction network that integrates convolutional neural network (CNN) and sLSTM components within a multitask learning framework. It jointly predicts human intelligibility scores and machine-based word error rates (WER) from Google ASR and Whisper. Experimental results show that iMTI-Net outperforms the original MTI-Net across multiple evaluation metrics, demonstrating the effectiveness of incorporating uncertainty-aware features and the CNN-sLSTM architecture.


【10】IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
标题:IS${}' 3 $:通用脉冲--使用深度过滤在声学场景中进行静态声音分离
链接:https://arxiv.org/abs/2509.02622

作者:Clementine Berger , Paraskevas Stamadiatis, Roland Badeau, Slim Essid
备注:None
摘要:我们感兴趣的音频系统能够执行一个不同的处理静态背景和孤立的声学场景内的声学事件,无论是应用特定的处理方法,每个部分或专注于一个而忽略其他。这样的系统在现实世界场景中具有应用,包括鲁棒自适应音频渲染系统(例如,EQ或压缩)、语音混合中的爆破音衰减、噪声抑制或减少、稳健的声学事件分类或甚至生物声学。为此,我们引入了IS${}^3$,一个为脉冲-平稳声音分离而设计的神经网络,它使用深度滤波方法将脉冲声事件从平稳背景中分离出来,可以作为上述任务的预处理阶段。为了确保最佳的训练,我们提出了一个复杂的数据生成管道,可以为这项任务管理和调整现有的数据集。我们证明了一种基于学习的方法,建立在一个相对轻量级的神经体系结构上,并用精心设计和不同的数据进行训练,在这个以前未解决的任务中是成功的,优于谐波-打击乐声音分离掩蔽方法,适用于音乐信号处理研究,以及小波滤波对客观分离指标。
摘要:We are interested in audio systems capable of performing a differentiated processing of stationary backgrounds and isolated acoustic events within an acoustic scene, whether for applying specific processing methods to each part or for focusing solely on one while ignoring the other. Such systems have applications in real-world scenarios, including robust adaptive audio rendering systems (e.g., EQ or compression), plosive attenuation in voice mixing, noise suppression or reduction, robust acoustic event classification or even bioacoustics. To this end, we introduce IS${}^3$, a neural network designed for Impulsive--Stationary Sound Separation, that isolates impulsive acoustic events from the stationary background using a deep filtering approach, that can act as a pre-processing stage for the above-mentioned tasks. To ensure optimal training, we propose a sophisticated data generation pipeline that curates and adapts existing datasets for this task. We demonstrate that a learning-based approach, build on a relatively lightweight neural architecture and trained with well-designed and varied data, is successful in this previously unaddressed task, outperforming the Harmonic--Percussive Sound Separation masking method, adapted from music signal processing research, and wavelet filtering on objective separation metrics.


【11】Gaussian Process Regression of Steering Vectors With Physics-Aware Deep Composite Kernels for Augmented Listening
标题:具有物理感知深度复合核的引导载体的高斯过程回归以增强听力
链接:https://arxiv.org/abs/2509.02571

作者:Diego Di Carlo (RIKEN AIP), Koyama Shoichi (UTokyo), Nugraha Aditya Arie (RIKEN AIP), Fontaine Mathieu (LTCI, S2A), Bando Yoshiaki (AIST), Yoshii Kazuyoshi (RIKEN AIP)
摘要:本文研究了用于增强收听的导向向量在麦克风和源的频率和位置上的连续表示(例如,空间滤波和双耳再现),并精确控制用户感知的声场。导向矢量通常用于将声场的空间特性表示为收听位置的函数。假设理想环境的导向矢量的基本代数表示不能处理声场的散射效应。因此,可以收集在专用设施中测量的实际导向矢量的离散集合,并且超分辨(即,upsample)。最近,物理感知的深度学习方法已有效地用于此目的。然而,这种确定性的超分辨率,遭受过拟合问题,由于测量空间上的非均匀的不确定性。为了解决这个问题,我们将基于神经场(NF)的表达表示集成到基于高斯过程(GP)的原则概率框架中。具体来说,我们提出了一个物理感知的复合内核,模型的方向传入波和随后的散射效果。通过综合对比实验,验证了该方法在数据不足情况下的有效性.在下游任务中,如语音增强和双耳渲染,使用SPECTRA挑战的模拟数据,预言性能达到不到十倍的测量。
摘要:This paper investigates continuous representations of steering vectors over frequency and position of microphone and source for augmented listening (e.g., spatial filtering and binaural rendering) with precise control of the sound field perceived by the user. Steering vectors have typically been used for representing the spatial characteristics of the sound field as a function of the listening position. The basic algebraic representation of steering vectors assuming an idealized environment cannot deal with the scattering effect of the sound field. One may thus collect a discrete set of real steering vectors measured in dedicated facilities and super-resolve (i.e., upsample) them. Recently, physics-aware deep learning methods have been effectively used for this purpose. Such deterministic super-resolution, however, suffers from the overfitting problem due to the non-uniform uncertainty over the measurement space. To solve this problem, we integrate an expressive representation based on the neural field (NF) into the principled probabilistic framework based on the Gaussian process (GP). Specifically, we propose a physics-aware composite kernel that model the directional incoming waves and the subsequent scattering effect. Our comprehensive comparative experiment showed the effectiveness of the proposed method under data insufficiency conditions. In downstream tasks such as speech enhancement and binaural rendering using the simulated data of the SPEAR challenge, the oracle performances were attained with less than ten times fewer measurements.


eess.AS音频处理


【1】An Effective Strategy for Modeling Score Ordinality and Non-uniform Intervals in Automated Speaking Assessment
标题:自动口语评估中得分有序度和非均匀间隔建模的有效策略
链接:https://arxiv.org/abs/2509.03372

作者: Tien-Hong Lo, Szu-Yu Chen, Yao-Ting Sung, Berlin Chen
备注:Accepted at ASRU 2025
摘要:最近一系列关于自动口语评估(ASA)的研究受益于自监督学习(SSL)表示,它可以在没有特征策展的基础假设的情况下捕获非母语语音中丰富的声学和语言模式。然而,基于语音的SSL模型捕获声学相关的特征,但忽略了语言内容,而基于文本的SSL模型依赖于ASR输出,无法编码韵律的细微差别。此外,大多数现有技术将熟练程度水平视为名义类别,忽略了它们的顺序结构和熟练程度标签之间的非均匀间隔。为了解决这些限制,我们提出了一个有效的ASA方法相结合的SSL手工制作的指标功能,通过一个新的建模范例。我们进一步引入了一个多利润率序数损失,共同模型的得分序数和非均匀间隔的熟练度标签。在TEEMI语料库上进行的大量实验表明,我们的方法始终优于强基线,并很好地推广到看不见的提示。
摘要:A recent line of research on automated speaking assessment (ASA) has benefited from self-supervised learning (SSL) representations, which capture rich acoustic and linguistic patterns in non-native speech without underlying assumptions of feature curation. However, speech-based SSL models capture acoustic-related traits but overlook linguistic content, while text-based SSL models rely on ASR output and fail to encode prosodic nuances. Moreover, most prior arts treat proficiency levels as nominal classes, ignoring their ordinal structure and non-uniform intervals between proficiency labels. To address these limitations, we propose an effective ASA approach combining SSL with handcrafted indicator features via a novel modeling paradigm. We further introduce a multi-margin ordinal loss that jointly models both the score ordinality and non-uniform intervals of proficiency labels. Extensive experiments on the TEEMI corpus show that our method consistently outperforms strong baselines and generalizes well to unseen prompts.


【2】Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
标题:通过三重缺失和自我监督嵌入改善知觉音频审美评估
链接:https://arxiv.org/abs/2509.03292

作者:Dyah A. M. G. Wisnu, Ryandhimas E. Zezario, Stefano Rini, Hsin-Min Wang, Yu Tsao
备注:Accepted by IEEE Automatic Speech Recognition and Understanding Workshop(ASRU), 2025
摘要:我们提出了一个用于生成音频的自动多轴感知质量预测的系统,该系统是为AudioMOS Challenge 2025的Track 2开发的。该任务是预测四个音频美学分数-生产质量,生产复杂性,内容享受和内容丰富性-由文本到语音(TTS),文本到音频(TTA)和文本到音乐(TTM)系统生成的音频。一个主要的挑战是自然训练数据和综合评估数据之间的域转移。为了解决这个问题,我们将BEAT(一种预训练的基于变换的音频表示模型)与多分支长短期记忆(LSTM)预测器相结合,并使用基于缓冲区的采样的三重丢失来通过感知相似性来构建嵌入空间。我们的研究结果表明,这提高了嵌入的可辨别性和泛化性,使域鲁棒的音频质量评估没有合成的训练数据。
摘要:We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness--for audio generated by text-to-speech (TTS), text-to-audio (TTA), and text-to-music (TTM) systems. A main challenge is the domain shift between natural training data and synthetic evaluation data. To address this, we combine BEATs, a pretrained transformer-based audio representation model, with a multi-branch long short-term memory (LSTM) predictor and use a triplet loss with buffer-based sampling to structure the embedding space by perceptual similarity. Our results show that this improves embedding discriminability and generalization, enabling domain-robust audio quality assessment without synthetic training data.


【3】A Study on Zero-Shot Non-Intrusive Speech Intelligibility for Hearing Aids Using Large Language Models
标题:使用大语言模型的助听器Zero-Shot非侵入性语音可理解性研究
链接:https://arxiv.org/abs/2509.03021

作者:Ryandhimas E. Zezario, Dyah A.M.G. Wisnu, Hsin-Min Wang, Yu Tsao
备注:Accepted to IEEE ICCE-TW 2025
摘要:本工作的重点是zero-shot非侵入性的语音评估助听器(HA)使用大语言模型(LLM)。具体来说,我们介绍了GPT-Whisper-HA,GPT-Whisper的扩展,zero-shot非侵入性语音评估模型的基础上LLM。GPT-Whisper-HA是专为HA的语音评估而设计的,它结合了MSBG听力损失和NAL-R模拟,以根据每个人的听力图处理音频输入,两个自动语音识别(ASR)模块用于音频到文本的表示,GPT-4 o用于预测两个相应的分数,然后对最终估计分数进行分数平均。实验结果表明,GPT-Whisper-HA实现了2.59%的相对均方根误差(RMSE)的改善GPT-Whisper,确认潜在的LLM的zero-shot语音评估预测HA用户的主观可懂度。
摘要:This work focuses on zero-shot non-intrusive speech assessment for hearing aids (HA) using large language models (LLMs). Specifically, we introduce GPT-Whisper-HA, an extension of GPT-Whisper, a zero-shot non-intrusive speech assessment model based on LLMs. GPT-Whisper-HA is designed for speech assessment for HA, incorporating MSBG hearing loss and NAL-R simulations to process audio input based on each individual's audiogram, two automatic speech recognition (ASR) modules for audio-to-text representation, and GPT-4o to predict two corresponding scores, followed by score averaging for the final estimated score. Experimental results indicate that GPT-Whisper-HA achieves a 2.59% relative root mean square error (RMSE) improvement over GPT-Whisper, confirming the potential of LLMs for zero-shot speech assessment in predicting subjective intelligibility for HA users.


【4】Non-Intrusive Intelligibility Prediction for Hearing Aids: Recent Advances, Trends, and Challenges
标题:助听器的非侵入性可认知度预测:最新进展、趋势和挑战
链接:https://arxiv.org/abs/2509.03017

作者:Ryandhimas E. Zezario
备注:APSIPA ASC 2025 perspective paper
摘要:本文综述了助听器(HA)的非侵入性语言清晰度预测的最新进展。我们总结了强大的声学特征提取,听力损失建模,并使用新兴的架构长序列处理的发展。还讨论了特定于听众的适应策略和领域泛化方法,旨在提高在看不见的声学环境中的鲁棒性。认识到仍然存在的挑战,例如需要大规模、多样化的数据集和可靠的跨剖面归纳。我们的目标是提供对当前趋势,持续的挑战,以及未来可能的方向,以实用和可靠的HA为导向的可懂度预测系统的观点。
摘要:This paper provides an overview of recent progress in non-intrusive speech intelligibility prediction for hearing aids (HA). We summarize developments in robust acoustic feature extraction, hearing loss modeling, and the use of emerging architectures for long-sequence processing. Listener-specific adaptation strategies and domain generalization approaches that aim to improve robustness in unseen acoustic environments are also discussed. Remaining challenges, such as the need for large-scale, diverse datasets and reliable cross-profile generalization, are acknowledged. Our goal is to offer a perspective on current trends, ongoing challenges, and possible future directions toward practical and reliable HA-oriented intelligibility prediction systems.


【5】Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
标题:使用不确定性Whisper嵌入和sLSTM的语音可理解性评估
链接:https://arxiv.org/abs/2509.03013

作者:Ryandhimas E. Zezario, Dyah A.M.G. Wisnu, Hsin-Min Wang, Yu Tsao
备注:Accepted to APSIPA ASC 2025
摘要:由于说话者、噪声条件和主观感知的变化,非侵入式语音可懂度预测仍然具有挑战性。我们提出了一种不确定性感知方法,该方法利用Whisper嵌入与统计特征相结合,特别是在嵌入维度上计算的平均值,标准差和熵。通过softmax在特征维度上计算的熵作为不确定性的代理,补充了由平均值和标准差捕获的全局信息。为了对语音的序列结构进行建模,我们采用了标量长短期记忆(sLSTM)网络,它可以有效地捕获长程依赖关系。在此基础上,我们提出了iMTI-Net,这是一种改进的多目标可懂度预测网络,它在多任务学习框架内集成了卷积神经网络(CNN)和sLSTM组件。它联合预测来自Google ASR和Whisper的人类可懂度分数和基于机器的单词错误率(WER)。实验结果表明,iMTI-Net在多个评估指标上的性能优于原始MTI-Net,证明了结合不确定性感知特征和CNN-sLSTM架构的有效性。
摘要:Non-intrusive speech intelligibility prediction remains challenging due to variability in speakers, noise conditions, and subjective perception. We propose an uncertainty-aware approach that leverages Whisper embeddings in combination with statistical features, specifically the mean, standard deviation, and entropy computed across the embedding dimensions. The entropy, computed via a softmax over the feature dimension, serves as a proxy for uncertainty, complementing global information captured by the mean and standard deviation. To model the sequential structure of speech, we adopt a scalar long short-term memory (sLSTM) network, which efficiently captures long-range dependencies. Building on this foundation, we propose iMTI-Net, an improved multi-target intelligibility prediction network that integrates convolutional neural network (CNN) and sLSTM components within a multitask learning framework. It jointly predicts human intelligibility scores and machine-based word error rates (WER) from Google ASR and Whisper. Experimental results show that iMTI-Net outperforms the original MTI-Net across multiple evaluation metrics, demonstrating the effectiveness of incorporating uncertainty-aware features and the CNN-sLSTM architecture.


【6】IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
标题:IS${}' 3 $:通用脉冲--使用深度过滤在声学场景中进行静态声音分离
链接:https://arxiv.org/abs/2509.02622

作者:Clementine Berger (IDS, S2A), Stamadiatis Paraskevas (IDS, S2A), Badeau Roland (IDS, S2A), Essid Slim (IDS, S2A)
备注:None
摘要:我们感兴趣的音频系统能够执行一个不同的处理静态背景和孤立的声学场景内的声学事件,无论是应用特定的处理方法,每个部分或专注于一个而忽略其他。这样的系统在现实世界场景中具有应用,包括鲁棒自适应音频渲染系统(例如,EQ或压缩)、语音混合中的爆破音衰减、噪声抑制或减少、稳健的声学事件分类或甚至生物声学。为此,我们引入了IS${}^3$,一个为脉冲-平稳声音分离而设计的神经网络,它使用深度滤波方法将脉冲声事件从平稳背景中分离出来,可以作为上述任务的预处理阶段。为了确保最佳的训练,我们提出了一个复杂的数据生成管道,可以为这项任务管理和调整现有的数据集。我们证明了一种基于学习的方法,建立在一个相对轻量级的神经体系结构上,并用精心设计和不同的数据进行训练,在这个以前未解决的任务中是成功的,优于谐波-打击乐声音分离掩蔽方法,适用于音乐信号处理研究,以及小波滤波对客观分离指标。
摘要:We are interested in audio systems capable of performing a differentiated processing of stationary backgrounds and isolated acoustic events within an acoustic scene, whether for applying specific processing methods to each part or for focusing solely on one while ignoring the other. Such systems have applications in real-world scenarios, including robust adaptive audio rendering systems (e.g., EQ or compression), plosive attenuation in voice mixing, noise suppression or reduction, robust acoustic event classification or even bioacoustics. To this end, we introduce IS${}^3$, a neural network designed for Impulsive--Stationary Sound Separation, that isolates impulsive acoustic events from the stationary background using a deep filtering approach, that can act as a pre-processing stage for the above-mentioned tasks. To ensure optimal training, we propose a sophisticated data generation pipeline that curates and adapts existing datasets for this task. We demonstrate that a learning-based approach, build on a relatively lightweight neural architecture and trained with well-designed and varied data, is successful in this previously unaddressed task, outperforming the Harmonic--Percussive Sound Separation masking method, adapted from music signal processing research, and wavelet filtering on objective separation metrics.


【7】Gaussian Process Regression of Steering Vectors With Physics-Aware Deep Composite Kernels for Augmented Listening
标题:具有物理感知深度复合核的引导载体的高斯过程回归以增强听力
链接:https://arxiv.org/abs/2509.02571

作者:Diego Di Carlo (RIKEN AIP), Koyama Shoichi (UTokyo), Nugraha Aditya Arie (RIKEN AIP), Fontaine Mathieu (LTCI, S2A), Bando Yoshiaki (AIST), Yoshii Kazuyoshi (RIKEN AIP)
摘要:本文研究了用于增强收听的导向向量在麦克风和源的频率和位置上的连续表示(例如,空间滤波和双耳再现),并精确控制用户感知的声场。导向矢量通常用于将声场的空间特征表示为收听位置的函数。假设理想环境的导向矢量的基本代数表示不能处理声场的散射效应。因此,可以收集在专用设施中测量的实际导向矢量的离散集合,并且超分辨(即,upsample)。最近,物理感知的深度学习方法已有效地用于此目的。然而,这种确定性的超分辨率,遭受过拟合问题,由于测量空间上的非均匀的不确定性。为了解决这个问题,我们将基于神经场(NF)的表达表示集成到基于高斯过程(GP)的原则概率框架中。具体来说,我们提出了一个物理感知的复合内核,模型的方向传入波和随后的散射效果。通过综合对比实验,验证了该方法在数据不足情况下的有效性.在下游任务中,如语音增强和双耳渲染,使用SPECTRA挑战的模拟数据,预言性能达到不到十倍的测量。
摘要:This paper investigates continuous representations of steering vectors over frequency and position of microphone and source for augmented listening (e.g., spatial filtering and binaural rendering) with precise control of the sound field perceived by the user. Steering vectors have typically been used for representing the spatial characteristics of the sound field as a function of the listening position. The basic algebraic representation of steering vectors assuming an idealized environment cannot deal with the scattering effect of the sound field. One may thus collect a discrete set of real steering vectors measured in dedicated facilities and super-resolve (i.e., upsample) them. Recently, physics-aware deep learning methods have been effectively used for this purpose. Such deterministic super-resolution, however, suffers from the overfitting problem due to the non-uniform uncertainty over the measurement space. To solve this problem, we integrate an expressive representation based on the neural field (NF) into the principled probabilistic framework based on the Gaussian process (GP). Specifically, we propose a physics-aware composite kernel that model the directional incoming waves and the subsequent scattering effect. Our comprehensive comparative experiment showed the effectiveness of the proposed method under data insufficiency conditions. In downstream tasks such as speech enhancement and binaural rendering using the simulated data of the SPEAR challenge, the oracle performances were attained with less than ten times fewer measurements.


【8】Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge
标题:NOCASA 2025挑战赛的端到端语音评估模型比较
链接:https://arxiv.org/abs/2509.03256

作者:Aleksei Zavoronkov, Tanel Alumäe
备注:Published at IEEE MLSP 2025
摘要:本文分析了为NOCASA 2025挑战赛开发的三个端到端模型,旨在为学习挪威语作为第二语言的儿童进行单词级发音自动评估。我们的模型包括一个编码器-解码器连体架构(E2 E-R),一个前缀调整的直接分类模型,利用预训练的wav2vec2.0表示,以及一个新的模型,集成了通过CTC计算的无干扰的良好发音(GOP)功能。我们引入了一种加权有序交叉熵损失,专门用于优化未加权平均召回率和平均绝对误差等指标。在探索的方法中,我们基于GOP-CTC的模型实现了最高的性能,大大超过了挑战基线,并获得了最高的排行榜分数。
摘要:This paper presents an analysis of three end-to-end models developed for the NOCASA 2025 Challenge, aimed at automatic word-level pronunciation assessment for children learning Norwegian as a second language. Our models include an encoder-decoder Siamese architecture (E2E-R), a prefix-tuned direct classification model leveraging pretrained wav2vec2.0 representations, and a novel model integrating alignment-free goodness-of-pronunciation (GOP) features computed via CTC. We introduce a weighted ordinal cross-entropy loss tailored for optimizing metrics such as unweighted average recall and mean absolute error. Among the explored methods, our GOP-CTC-based model achieved the highest performance, substantially surpassing challenge baselines and attaining top leaderboard scores.


【9】Mitigating Data Imbalance in Automated Speaking Assessment
标题:缓解自动演讲评估中的数据失衡
链接:https://arxiv.org/abs/2509.03010

作者:Fong-Chun Tsai, Kuan-Tang Huang, Bi-Cheng Yan, Tien-Hong Lo, Berlin Chen
备注:Submitted to APSIPA 2025
摘要:自动口语评估在评估二语学习者的语言水平方面起着至关重要的作用。然而,ASA模型经常遭受类不平衡,导致有偏见的预测。为了解决这个问题,我们引入了一个新的目标来训练ASA模型,称为平衡Logit变异(BLV)损失,它会干扰模型预测,以改善少数类的特征表示,而无需修改数据集。对ICNALE基准数据集的评估表明,将BLV损失集成到著名的基于文本的(BERT)模型中显着提高了分类准确性和公平性,使自动语音评估对不同的学习者更加强大。
摘要:Automated Speaking Assessment (ASA) plays a crucial role in evaluating second-language (L2) learners proficiency. However, ASA models often suffer from class imbalance, leading to biased predictions. To address this, we introduce a novel objective for training ASA models, dubbed the Balancing Logit Variation (BLV) loss, which perturbs model predictions to improve feature representation for minority classes without modifying the dataset. Evaluations on the ICNALE benchmark dataset show that integrating the BLV loss into a celebrated text-based (BERT) model significantly enhances classification accuracy and fairness, making automated speech evaluation more robust for diverse learners.


【10】Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
标题:Speech DF Arena:Speech DeepFake检测模型的排行榜
链接:https://arxiv.org/abs/2509.02859

作者:Sandipana Dowerah, Atharva Kulkarni, Ajinkya Kulkarni, Hoan My Tran, Joonas Kalda, Artem Fedorchenko, Benoit Fauve, Damien Lolive, Tanel Alumäe, Matthew Magimai Doss
摘要:与高级deepfake音频生成的发展并行,音频deepfake检测也取得了重大进展。然而,仍然缺少一个标准化和全面的基准。为了解决这个问题,我们引入了Speech DeepFake(DF)Arena,这是音频deepfake检测的第一个综合基准。Speech DF Arena提供了一个工具包来统一评估检测系统,目前涵盖14个不同的数据集和攻击场景,标准化的评估指标和协议,以实现可重复性和透明度。它还包括一个排行榜来比较和排名系统,以帮助研究人员和开发人员提高其可靠性和鲁棒性。我们包括14个评估集,12个最先进的开源和3个专有检测系统。我们的研究提出了许多系统表现出高EER域外的情况下,强调需要广泛的跨域评估。排行榜托管在Huggingface1上,GitHub上提供了一个工具包,用于在所列数据集上复制结果。
摘要:Parallel to the development of advanced deepfake audio generation, audio deepfake detection has also seen significant progress. However, a standardized and comprehensive benchmark is still missing. To address this, we introduce Speech DeepFake (DF) Arena, the first comprehensive benchmark for audio deepfake detection. Speech DF Arena provides a toolkit to uniformly evaluate detection systems, currently across 14 diverse datasets and attack scenarios, standardized evaluation metrics and protocols for reproducibility and transparency. It also includes a leaderboard to compare and rank the systems to help researchers and developers enhance their reliability and robustness. We include 14 evaluation sets, 12 state-of-the-art open-source and 3 proprietary detection systems. Our study presents many systems exhibiting high EER in out-of-domain scenarios, highlighting the need for extensive cross-domain evaluation. The leaderboard is hosted on Huggingface1 and a toolkit for reproducing results across the listed datasets is available on GitHub.


【11】SSVD: Structured SVD for Parameter-Efficient Fine-Tuning and Benchmarking under Domain Shift in ASR
标题:SSVD:结构化MVD,用于ASB中域转移下的参数高效微调和基准测试
链接:https://arxiv.org/abs/2509.02830

作者:Pu Wang, Shinji Watanabe, Hugo Van hamme
备注:Accepted by IEEE ASRU 2025
摘要:参数高效微调(PEFT)已成为适应大型基础模型的可扩展解决方案。虽然低秩自适应(LoRA)广泛用于语音应用中,但其最先进的变体,例如,VeRA、DoRA、PiSSA和SVFT主要是为语言和视觉任务开发的,在语音方面的验证有限。这项工作提出了第一个全面的集成和ESPnet内的这些PEFT方法的基准。我们进一步引入结构化SVD引导(SSVD)微调,它选择性地旋转输入相关的右奇异向量,同时保持输出相关向量固定,以保持语义映射。这种设计能够以最小的可训练参数和提高的效率实现鲁棒的域自适应。我们评估了域转移语音识别任务的所有方法,包括儿童语音和方言变化,模型规模从0.1B到2B。所有实现都在ESPnet中发布,以支持可重复性和未来的工作。
摘要:Parameter-efficient fine-tuning (PEFT) has emerged as a scalable solution for adapting large foundation models. While low-rank adaptation (LoRA) is widely used in speech applications, its state-of-the-art variants, e.g., VeRA, DoRA, PiSSA, and SVFT, are developed mainly for language and vision tasks, with limited validation in speech. This work presents the first comprehensive integration and benchmarking of these PEFT methods within ESPnet. We further introduce structured SVD-guided (SSVD) fine-tuning, which selectively rotates input-associated right singular vectors while keeping output-associated vectors fixed to preserve semantic mappings. This design enables robust domain adaptation with minimal trainable parameters and improved efficiency. We evaluate all methods on domain-shifted speech recognition tasks, including child speech and dialectal variation, across model scales from 0.1B to 2B. All implementations are released in ESPnet to support reproducibility and future work.


【12】Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission
标题:神经音频编解码器传输的说话人验证性能权衡分析
链接:https://arxiv.org/abs/2509.02771

作者:Nirmalya Mallick Thakur, Jia Qi Yip, Eng Siong Chng
备注:Accepted by APSIPA ASC 2025
摘要:神经音频编解码器(NAC)近年来取得了重大进展,并迅速被许多音频处理管道采用。然而,它们会引入音频失真,从而降低说话人验证(SV)性能。这项研究调查了传统和神经音频编解码器在不同比特率下对VoxCeleb 1数据集上评估的三种最先进的SV模型的影响。我们的研究结果显示,随着比特率的降低,所有模型和编解码器的SV性能都出现了一致的下降。值得注意的是,与传统编解码器相比,NAC不会从根本上破坏SV性能。它们在低比特率(< 12 kbps)下的表现优于Opus 6-8%,在高比特率(约24 kbps)下仍略落后,EER仅增加0.4- 0.7%。较高比特率下的差异可能是由于NAC对感知质量的主要优化,这可能会无意中丢弃关键的说话者区分特征,而Opus则旨在保留声音特征。我们的研究表明,NAC是一个可行的替代传统的编解码器,特别是在带宽限制。为了弥补更高比特率的差距,未来的工作应该集中在开发说话者感知的NAC或重新训练和适应SV模型上。
摘要:Neural audio codecs (NACs) have made significant advancements in recent years and are rapidly being adopted in many audio processing pipelines. However, they can introduce audio distortions which degrade speaker verification (SV) performance. This study investigates the impact of both traditional and neural audio codecs at varying bitrates on three state of-the-art SV models evaluated on the VoxCeleb1 dataset. Our findings reveal a consistent degradation in SV performance across all models and codecs as bitrates decrease. Notably, NACs do not fundamentally break SV performance when compared to traditional codecs. They outperform Opus by 6-8% at low-bitrates (< 12 kbps) and remain marginally behind at higher bitrates ($\approx$ 24 kbps), with an EER increase of only 0.4-0.7%. The disparity at higher bitrates is likely due to the primary optimization of NACs for perceptual quality, which can inadvertently discard critical speaker-discriminative features, unlike Opus which was designed to preserve vocal characteristics. Our investigation suggests that NACs are a feasible alternative to traditional codecs, especially under bandwidth limitations. To bridge the gap at higher bitrates, future work should focus on developing speaker-aware NACs or retraining and adapting SV models.


机器翻译由腾讯交互翻译提供,仅供参考