微信公众号:arXiv_Daily
cs.SD语音
标题:在具有反向和不匹配的语音文本方向的端到端TTC中探索人类关节表达限制
链接:https://arxiv.org/abs/2602.14664
备注:A shorter version of this paper appeared in ACPR 2025
摘要:端到端(e2 e)文本到语音(TTS)系统是一种深度架构,它学习将文本字符串与来自精选数据集的声学语音模式相关联。预计与语音产生相关联的所有方面,例如电话持续时间、扬声器特征和语调等,都将在训练的TTS模型中捕获,以使合成语音能够自然和可理解。人类语音是复杂的,涉及发音配置(AC)之间的平滑过渡。由于解剖结构的限制,一些AC很难模拟或转换。在本文中,我们的实验研究,如果人体解剖学所施加的限制有一个e2 e-TTS系统的培训的影响。我们实验两个e2 e-TTS架构,即Tacotron-2的自回归模型和VITS-TTS的非自回归模型。在这项研究中,我们使用(a)前向文本,前向语音(传统,e2 e-TTS),(b)反向文本,反向语音(r-e2 e-TTS),(c)反向文本,前向语音(rtfs-e2 e-TTS)构建TTS系统。实验表明,e2 e-TTS系统是纯数据驱动的。有趣的是,由r-e2 e-TTS系统生成的语音表现出更好的保真度、更好的感知可懂度和更好的自然度
摘要:An end-to-end (e2e) text-to-speech (TTS) system is a deep architecture that learns to associate a text string with acoustic speech patterns from a curated dataset. It is expected that all aspects associated with speech production, such as phone duration, speaker characteristics, and intonation among other things are captured in the trained TTS model to enable the synthesized speech to be natural and intelligible. Human speech is complex, involving smooth transitions between articulatory configurations (ACs). Due to anatomical constraints, some ACs are challenging to mimic or transition between. In this paper, we experimentally study if the constraints imposed by human anatomy have an implication on training an e2e-TTS systems. We experiment with two e2e-TTS architectures, namely, Tacotron-2 an autoregressive model and VITS-TTS a non-autoregressive model. In this study, we build TTS systems using (a) forward text, forward speech (conventional, e2e-TTS), (b) reverse text, reverse speech (r-e2e-TTS), and (c) reverse text, forward speech (rtfs-e2e-TTS). Experiments demonstrate that e2e-TTS systems are purely data-driven. Interestingly, the generated speech by r-e2e-TTS systems exhibits better fidelity, better perceptual intelligibility, and better naturalness
【2】Bengali-Loop: Community Benchmarks for Long-Form Bangla ASR and Speaker Diarization
标题:Bengali-Loop:长格式孟加拉语ASB和扬声器拨号的社区基准链接:https://arxiv.org/abs/2602.14291
摘要:孟加拉语(孟加拉语)尽管被广泛使用,但在长格式语音技术方面仍然资源不足。我们提出了孟加拉循环,两个社区基准来解决这个差距:(1)191个录音的长格式ASR语料库(158.6小时,79.2万字)来自11个YouTube频道,通过可复制的字幕提取管道和人工参与的转录验证收集;以及(2)24个录音(22小时,5,744个注释片段)的说话者日记语料库,其具有CSV格式的完全手动的说话者转向标签。这两个基准都针对真实的多扬声器、长持续时间内容(例如,Bangla drama/natok).我们建立基线(Tugstugi:34.07% WER; pyannote.audio:40.08% DER),并提供标准化的评估协议(WER/CER,DER),注释规则和数据格式,以支持孟加拉语长格式ASR和日记的可重复基准测试和未来模型开发。
摘要:Bengali (Bangla) remains under-resourced in long-form speech technology despite its wide use. We present Bengali-Loop, two community benchmarks to address this gap: (1) a long-form ASR corpus of 191 recordings (158.6 hours, 792k words) from 11 YouTube channels, collected via a reproducible subtitle-extraction pipeline and human-in-the-loop transcript verification; and (2) a speaker diarization corpus of 24 recordings (22 hours, 5,744 annotated segments) with fully manual speaker-turn labels in CSV format. Both benchmarks target realistic multi-speaker, long-duration content (e.g., Bangla drama/natok). We establish baselines (Tugstugi: 34.07% WER; pyannote.audio: 40.08% DER) and provide standardized evaluation protocols (WER/CER, DER), annotation rules, and data formats to support reproducible benchmarking and future model development for Bangla long-form ASR and diarization.
【3】The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents
标题:Interspeech 2026音频推理挑战:评估音频推理模型和代理的推理过程质量链接:https://arxiv.org/abs/2602.14224
备注:The official website of the Audio Reasoning Challenge: https://audio-reasoning-challenge.github.io
摘要:最近的大型音频语言模型(LALM)擅长理解,但往往缺乏透明的推理。为了解决这个“黑箱”限制,我们在Interspeech 2026上组织了音频推理挑战赛,这是第一个致力于评估音频领域中的思想链(CoT)质量的共享任务。该挑战引入了MMAR-Rubrics,这是一种新颖的实例级协议,用于评估推理链的真实性和逻辑。比赛以单一模式和代理人为特色,吸引了来自18个国家和地区的156支队伍。结果表明,代理系统目前领先的推理质量,利用迭代工具编排和跨模态分析。此外,单一模型通过强化学习和复杂的数据管道快速发展。我们详细介绍了挑战设计,方法和对最先进系统的全面分析,为可解释的音频智能提供了新的见解。
摘要:Recent Large Audio Language Models (LALMs) excel in understanding but often lack transparent reasoning. To address this "black-box" limitation, we organized the Audio Reasoning Challenge at Interspeech 2026, the first shared task dedicated to evaluating Chain-of-Thought (CoT) quality in the audio domain. The challenge introduced MMAR-Rubrics, a novel instance-level protocol assessing the factuality and logic of reasoning chains. Featured Single Model and Agent tracks, the competition attracting 156 teams from 18 countries and regions. Results show agent systems currently lead in reasoning quality, utilizing iterative tool orchestration and cross-modal analysis. Besides, single models are rapidly advancing via reinforcement learning and sophisticated data pipeline. We details the challenge design, methodology, and a comprehensive analysis of state-of-the-art systems, providing new insights for explainable audio intelligence.
【4】Investigation for Relative Voice Impression Estimation
标题:相对语音印象估计的研究链接:https://arxiv.org/abs/2602.14172
备注:5 pages,3 figures, Accepted to Speech Prosody 2026
摘要:言语的副语言和非语言方面强烈影响听者的印象。虽然大多数研究都集中在绝对印象评分,本研究调查相对语音印象估计(RIE),一个框架,用于预测来自同一扬声器的两个话语之间的感知差异。估计目标是从主观评估导出的低维向量,其量化第二话语相对于第一话语沿着反义轴(例如,"Dark-“)。为了分离表达和韵律的变化,我们使用了专业演讲者以各种风格阅读文本的录音。我们比较了三种建模方法:经典的语音情感识别,自我监督的语音表示,和多模态大语言模型(MLLM)常用的声学特征。我们的研究结果表明,使用自监督表示的模型优于具有经典声学特征的方法,特别是在捕获复杂和动态印象(例如,“冷--暖”),经典特征失败的地方。相比之下,目前的MLLM证明这种细粒度的成对任务是不可靠的。这项研究提供了RIE的第一个系统的调查,并证明了自我监督的语音模型在捕捉微妙的感知变化的力量。
摘要:Paralinguistic and non-linguistic aspects of speech strongly influence listener impressions. While most research focuses on absolute impression scoring, this study investigates relative voice impression estimation (RIE), a framework for predicting the perceptual difference between two utterances from the same speaker. The estimation target is a low-dimensional vector derived from subjective evaluations, quantifying the perceptual shift of the second utterance relative to the first along an antonymic axis (e.g., ``Dark--Bright''). To isolate expressive and prosodic variation, we used recordings of a professional speaker reading a text in various styles. We compare three modeling approaches: classical acoustic features commonly used for speech emotion recognition, self-supervised speech representations, and multimodal large language models (MLLMs). Our results demonstrate that models using self-supervised representations outperform methods with classical acoustic features, particularly in capturing complex and dynamic impressions (e.g., ``Cold--Warm'') where classical features fail. In contrast, current MLLMs prove unreliable for this fine-grained pairwise task. This study provides the first systematic investigation of RIE and demonstrates the strength of self-supervised speech models in capturing subtle perceptual variations.
【5】MUKA: Multi Kernel Audio Adaptation Of Audio-Language Models
标题:MUKA:音频语言模型的多核心音频适应链接:https://arxiv.org/abs/2602.14127
摘要:多模态基础模型已经展示了令人印象深刻的泛化能力,但在Few-Shot设置中有效地使其适应新任务仍然是一个关键的挑战。在这项工作中,我们研究了Few-Shot适应大型音频语言模型(ALMs)通过基于训练和训练的方法。我们介绍了MUKA,这是一个多内核自适应框架,它结合了Pengi等基于自适应调整的模型的细粒度、上下文相关的表示和CLAP等对比预训练方法的全局语义表示。通过构建一个将局部相似性与全局语义相结合的产品内核,MUKA增强了表示能力,同时保留了内核方法的理论保证并避免了额外的训练。在11个不同的音频数据集上进行的广泛实验表明,MUKA在无训练方法中实现了最先进的性能,甚至在几种情况下超过了基于训练的适配器,在适应性和效率之间实现了令人信服的平衡。
摘要:Multimodal foundation models have demonstrated impressive generalization capabilities, yet efficiently adapting them to new tasks in a few-shot setting remains a critical challenge. In this work, we investigate the few-shot adaptation of Large Audio-Language Models (ALMs) through both training-based and training-free approaches. We introduce MUKA, a multi-kernel adaptation framework that combines the fine-grained, context-dependent representations of instruction-tuning based models like Pengi with the global semantic representations of contrastive pretraining methods like CLAP. By constructing a product kernel that aligns local similarity with global semantics, MUKA enhances representational power while preserving the theoretical guarantees of kernel methods and avoiding additional training. Extensive experiments across 11 diverse audio datasets demonstrate that MUKA achieves state-of-the-art performance among training-free methods and even surpasses training-based adapters in several scenarios, offering a compelling balance between adaptability and efficiency.
【6】From Scarcity to Scale: A Release-Level Analysis of the Pashto Common Voice Dataset
标题:从稀缺到规模:普什图语常见语音数据集的发布级分析链接:https://arxiv.org/abs/2602.14062
摘要:大型、开放许可的语音数据集对于构建自动语音识别(ASR)系统至关重要,但许多广泛使用的语言在公共资源中的代表性仍然不足。普什图语有超过6000万人使用,历史上一直缺乏适合现代ASR开发的大规模公开许可的语音数据。 本文对Mozilla Common Voice语料库的普什图语组件进行了发布级分析,重点关注24.0版(2025年12月),并对主要版本的趋势进行了上下文分析。我们记录了从2023年年中的1. 49小时记录到2025年的2,768. 7小时的快速增长,其中包括975. 89小时可用于监督ASR培训的验证小时。 除了规模,我们分析验证吞吐量,贡献者参与不平等,人口统计元数据的完整性,并在验证的子集中的企业级浓度。我们发现,参与是非常集中的(基尼系数= 0.941),年龄代表性是强烈倾向于年轻人,和41.97%的剪辑缺乏自我报告的性别标签,限制了基于元数据的亚组审计。在文本层面上,提示重用是温和的:35.88%的独特的句子占50%的验证剪辑,这表明结构集中主要是由不均匀的贡献者活动,而不是一个小的提示集的优势。 这些结果提供了一个快速扩展的低资源语音语料库的定量审计,并突出了提高数据集成熟度的实际优先事项,包括扩大验证能力和更广泛的人口参与。
摘要:Large, openly licensed speech datasets are essential for building automatic speech recognition (ASR) systems, yet many widely spoken languages remain underrepresented in public resources. Pashto, spoken by more than 60 million people, has historically lacked large-scale openly licensed speech data suitable for modern ASR development. This paper presents a release-level analysis of the Pashto component of the Mozilla Common Voice corpus, focusing on version 24.0 (December 2025) and contextualizing trends across major releases. We document rapid growth from 1.49 recorded hours in mid-2023 to 2,768.7 total hours in 2025, including 975.89 validated hours available for supervised ASR training. Beyond scale, we analyze validation throughput, contributor participation inequality, demographic metadata completeness, and sentence-level concentration in the validated subset. We find that participation is extremely concentrated (Gini = 0.941), age representation is strongly skewed toward young adults, and 41.97\% of clips lack self-reported gender labels, limiting subgroup auditing based on metadata. At the textual level, prompt reuse is moderate: 35.88\% of unique sentences account for 50\% of validated clips, suggesting that structural concentration is driven primarily by uneven contributor activity rather than dominance of a small prompt set. These results provide a quantitative audit of a rapidly scaling low-resource speech corpus and highlight practical priorities for improving dataset maturity, including expanded validation capacity and broader demographic participation.
【7】Eureka-Audio: Triggering Audio Intelligence in Compact Language Models
标题:尤里卡音频:在紧凑语言模型中触发音频智能链接:https://arxiv.org/abs/2602.13954
备注:23 pages, 4 figures
摘要:Eureka-Audio是一种紧凑而高性能的音频语言模型,在广泛的音频理解基准测试中,它与4到18倍的模型相比具有竞争力的性能。尽管仅包含1.7B参数,但Eureka-Audio在自动语音识别(ASR)、音频理解和密集音频字幕方面表现出强大的性能,匹配或超越了多个7 B至30 B音频和全模态基线。该模型采用统一的端到端架构,由轻量级语言主干、基于Whisper的音频编码器和稀疏激活的Mixture-of-Experts(MoE)适配器组成,该适配器明确考虑了音频异构性,并在有限容量下消除了跨模态优化冲突。为了进一步增强语言推理,我们引入了DataFlux,这是一个闭环音频指令数据合成和验证管道,可以从原始音频中构建高质量,逻辑一致的监督。在ASR、知识推理、安全性、指令遵循和语言基准方面的广泛评估表明,Eureka-Audio在计算成本和性能之间实现了有效的平衡。这些结果使Eureka Audio成为轻量级音频理解模型的强大而实用的基线。
摘要:We present Eureka-Audio, a compact yet high-performance audio language model that achieves competitive performance against models that are 4 to 18 times larger across a broad range of audio understanding benchmarks. Despite containing only 1.7B parameters, Eureka-Audio demonstrates strong performance on automatic speech recognition (ASR), audio understanding, and dense audio captioning, matching or surpassing multiple 7B to 30B audio and omni-modal baselines. The model adopts a unified end-to-end architecture composed of a lightweight language backbone, a Whisper-based audio encoder, and a sparsely activated Mixture-of-Experts (MoE) adapter that explicitly accounts for audio heterogeneity and alleviates cross-modal optimization conflicts under limited capacity. To further enhance paralinguistic reasoning, we introduce DataFlux, a closed loop audio instruction data synthesis and verification pipeline that constructs high quality, logically consistent supervision from raw audio. Extensive evaluations across ASR, knowledge reasoning, safety, instruction following, and paralinguistic benchmarks, demonstrate that Eureka-Audio achieves an efficient balance between computational cost and performance. These results establish Eureka Audio as a strong and practical baseline for lightweight audio understanding models.
【8】voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
标题:voice2 mode:使用自我监督语音模型进行歌唱中的发音模式分类链接:https://arxiv.org/abs/2602.13928
备注:Accepted to the Speech, Music and Mind (SMM26) workshop at the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026). This is the preprint version of the paper to appear in the proceedings
摘要:我们提出了voice 2 mode,一种使用从大型自监督语音模型中提取的嵌入对四种歌唱发声模式(呼吸,中性(模态),流动和按压)进行分类的方法。以前的工作歌唱发声依赖于手工制作的信号功能或特定任务的神经网络,这项工作评估的可移植性语音基础模型歌唱发声分类。voice 2 mode从HuBERT和两个wav 2 vec 2变体中提取分层表示,应用全局时间池化,并使用轻量级分类器(SVM,XGBoost)对池化嵌入进行分类。公开的女高音数据集(763持续元音录音,四个标签)的实验表明,基础模型功能大大优于传统的频谱基线(频谱图,梅尔频谱图,MFCC)。从早期层获得的HuBERT嵌入产生了最好的结果(SVM的准确率约为95.7%),比最好的传统基线绝对提高了约12-15%。我们还展示了分层行为:保留声学/语音细节的较低层比专门用于自动语音识别(ASR)的顶层更有效。
摘要:We present voice2mode, a method for classification of four singing phonation modes (breathy, neutral (modal), flow, and pressed) using embeddings extracted from large self-supervised speech models. Prior work on singing phonation has relied on handcrafted signal features or task-specific neural nets; this work evaluates the transferability of speech foundation models to singing phonation classification. voice2mode extracts layer-wise representations from HuBERT and two wav2vec2 variants, applies global temporal pooling, and classifies the pooled embeddings with lightweight classifiers (SVM, XGBoost). Experiments on a publicly available soprano dataset (763 sustained vowel recordings, four labels) show that foundation-model features substantially outperform conventional spectral baselines (spectrogram, mel-spectrogram, MFCC). HuBERT embeddings obtained from early layers yield the best result (~95.7% accuracy with SVM), an absolute improvement of ~12-15% over the best traditional baseline. We also show layer-wise behaviour: lower layers, which retain acoustic/phonetic detail, are more effective than top layers specialized for Automatic Speech Recognition (ASR).
【9】GSRM: Generative Speech Reward Model for Speech RLHF
标题:GSOM:语音RL HF的生成语音奖励模型链接:https://arxiv.org/abs/2602.13891
摘要:语音语言模型的最新进展,如GPT-4 o语音模式和Gemini Live,已经证明了有前途的语音生成能力。然而,合成音频的美学自然度仍然落后于人类语音。提高生成质量需要可靠的语音自然度评估器。然而,现有的自然度评估器通常将原始音频回归到标量分数,提供有限的评估可解释性,而且无法推广到不同分类的语音。受生成奖励模型最新进展的启发,我们提出了生成语音奖励模型(GSRM),这是一种为语音量身定制的以推理为中心的奖励模型。GSRM经过训练,将语音自然度评估分解为可解释的声学特征提取阶段,然后进行基于特征的思维链推理,从而实现可解释的判断。为了实现这一目标,我们策划了一个大规模的人类反馈数据集,其中包括31 k个专家评级和一个真实世界用户辅助语音交互的域外基准。实验表明,GSRM大大优于现有的语音自然度预测,实现模型的人的相关性的自然度分数预测,接近人类评分员之间的一致性。我们进一步展示了如何GSRM可以提高自然的语音LLM代作为一个有效的验证器在线RLHF。
摘要:Recent advances in speech language models, such as GPT-4o Voice Mode and Gemini Live, have demonstrated promising speech generation capabilities. Nevertheless, the aesthetic naturalness of the synthesized audio still lags behind that of human speech. Enhancing generation quality requires a reliable evaluator of speech naturalness. However, existing naturalness evaluators typically regress raw audio to scalar scores, offering limited interpretability of the evaluation and moreover fail to generalize to speech across different taxonomies. Inspired by recent advances in generative reward modeling, we propose the Generative Speech Reward Model (GSRM), a reasoning-centric reward model tailored for speech. The GSRM is trained to decompose speech naturalness evaluation into an interpretable acoustic feature extraction stage followed by feature-grounded chain-of-thought reasoning, enabling explainable judgments. To achieve this, we curated a large-scale human feedback dataset comprising 31k expert ratings and an out-of-domain benchmark of real-world user-assistant speech interactions. Experiments show that GSRM substantially outperforms existing speech naturalness predictors, achieving model-human correlation of naturalness score prediction that approaches human inter-rater consistency. We further show how GSRM can improve the naturalness of speech LLM generations by serving as an effective verifier for online RLHF.
【10】Audiocards: Structured Metadata Improves Audio Language Models For Sound Design
标题:音频卡:结构化元数据改进声音设计的音频语言模型链接:https://arxiv.org/abs/2602.13835
备注:Accepted at ICASSP 2026
摘要:声音设计师使用声音类或视觉上下文等方面在大型声音效果库中搜索声音。然而,这样的搜索所需的元数据往往是缺失或不完整的,并需要大量的手动工作来添加。现有的解决方案,以自动化这一任务,通过生成元数据,即字幕,并使用学习嵌入搜索,即文本音频检索,没有训练元数据的结构和信息有关的声音设计。为此,我们提出了声卡,结构化的元数据接地声学属性和声音描述符,通过利用世界知识的LLM。我们表明,培训的音频卡提高下游的文本音频检索,描述性字幕,元数据生成专业的音效库。此外,音频卡还提高了性能的一般音频字幕和检索的基线单句字幕的方法。我们发布了一个精心策划的音效声卡数据集,以邀请进一步研究声音设计的音频语言建模。
摘要:Sound designers search for sounds in large sound effects libraries using aspects such as sound class or visual context. However, the metadata needed for such search is often missing or incomplete, and requires significant manual effort to add. Existing solutions to automate this task by generating metadata, i.e. captioning, and search using learned embeddings, i.e. text-audio retrieval, are not trained on metadata with the structure and information pertinent to sound design. To this end we propose audiocards, structured metadata grounded in acoustic attributes and sonic descriptors, by exploiting the world knowledge of LLMs. We show that training on audiocards improves downstream text-audio retrieval, descriptive captioning, and metadata generation on professional sound effects libraries. Moreover, audiocards also improve performance on general audio captioning and retrieval over the baseline single-sentence captioning approach. We release a curated dataset of sound effects audiocards to invite further research in audio language modeling for sound design.
【11】Learning Vocal-Tract Area and Radiation with a Physics-Informed Webster Model
标题:使用了解物理情况的韦伯斯特模型学习气道区域和辐射链接:https://arxiv.org/abs/2602.13834
备注:Accepted at IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026
摘要:我们提出了一个物理通知有声后端渲染器唱歌的声音合成。给定合成单声道音频和基频轨迹,我们训练时域韦伯斯特模型作为物理信息的神经网络来估计可解释的声道面积函数和开口辐射系数。训练强制执行偏微分方程和边界一致性;轻量级DDSP路径仅用于稳定学习,而推理纯粹基于物理。在持续元音(/a/,/i/,/u/),由独立的时域有限差分韦伯斯特求解器呈现的参数与紧凑的DDSP基线竞争性地再现频谱包络,并且在离散化、适度的源变化和大约百分之十的音高偏移的变化下保持稳定。图内波形保持比参考呼吸,激励在未来的工作中意识到目标和明确的声门先验。
摘要:We present a physics-informed voiced backend renderer for singing-voice synthesis. Given synthetic single-channel audio and a fund-amental--frequency trajectory, we train a time-domain Webster model as a physics-informed neural network to estimate an interpretable vocal-tract area function and an open-end radiation coefficient. Training enforces partial differential equation and boundary consistency; a lightweight DDSP path is used only to stabilize learning, while inference is purely physics-based. On sustained vowels (/a/, /i/, /u/), parameters rendered by an independent finite-difference time-domain Webster solver reproduce spectral envelopes competitively with a compact DDSP baseline and remain stable under changes in discretization, moderate source variations, and about ten percent pitch shifts. The in-graph waveform remains breathier than the reference, motivating periodicity-aware objectives and explicit glottal priors in future work.
【12】Enhancing spatial hearing with cochlear implants: exploring the role of AI, multimodal interaction and perceptual training
标题:通过人工智能增强空间听力:探索人工智能、多模式交互和感知训练的作用链接:https://arxiv.org/abs/2602.13787
摘要:髋关节植入物(CI)已经发展到可以恢复大部分患者的听力和语言理解的程度。虽然空间听觉是控制和引导注意力以及在嘈杂环境中实现言语理解的核心,但在过去它在很大程度上被忽视了。在这里,我们提出了一个多学科的研究框架,医生,心理学家和工程师合作,以提高CI用户的空间听力。
摘要:Cochlear implants (CIs) have been developed to the point where they can restore hearing and speech understanding in a large proportion of patients. Although spatial hearing is central to controlling and directing attention and to enabling speech understanding in noisy environments, it has been largely neglected in the past. We propose here a multi-disciplinary research framework in which physicians, psychologists and engineers collaborate to improve spatial hearing for CI users.
【13】AuTAgent: A Reinforcement Learning Framework for Tool-Augmented Audio Reasoning
标题:AutAGent:用于工具增强音频推理的强化学习框架链接:https://arxiv.org/abs/2602.13685
摘要:大型音频语言模型(LALM)擅长感知,但难以进行需要精确声学测量的复杂推理。虽然外部工具可以提取精确的节奏或音高等细粒度特征,但有效的集成仍然具有挑战性:天真地使用所有工具会导致信息过载,而基于上下文的选择无法评估依赖于上下文的效用。为了解决这个问题,我们提出了AutAgent(音频工具代理),这是一个强化学习框架,可以学习何时调用哪些工具。通过采用稀疏反馈训练策略和一种新的差分奖励机制,智能体学会过滤掉不相关的工具,只有当它产生的净性能增益超过基本模型时才调用外部援助。实验结果证实,AuTAgent通过提供可验证的声学证据补充了LALM的表示瓶颈。在MMAU Test-mini和MMAR基准测试中,开源和闭源主干的准确率分别提高了4.20% / 6.20%和9.80% / 8.00%。此外,进一步的实验证明了特殊的可转移性。我们强调了外部工具在增强音频模型推理中的补充作用。
摘要:Large Audio Language Models (LALMs) excel at perception but struggle with complex reasoning requiring precise acoustic measurements. While external tools can extract fine-grained features like exact tempo or pitch, effective integration remains challenging: naively using all tools causes information overload, while prompt-based selection fails to assess context-dependent utility. To address this, we propose AuTAgent (Audio Tool Agent), a reinforcement learning framework that learns when and which tools to invoke. By employing a sparse-feedback training strategy with a novel Differential Reward mechanism, the agent learns to filter out irrelevant tools and invokes external assistance only when it yields a net performance gain over the base model. Experimental results confirm that AuTAgent complements the representation bottleneck of LALMs by providing verifiable acoustic evidence. It improves accuracy by 4.20% / 6.20% and 9.80% / 8.00% for open-source and closed-source backbones on the MMAU Test-mini and the MMAR benchmarks, respectively. In addition, further experiments demonstrate exceptional transferability. We highlight the complementary role of external tools in augmenting audio model reasoning.
【14】BreathNet: Generalizable Audio Deepfake Detection via Breath-Cue-Guided Feature Refinement
标题:BreathNet:通过呼吸提示引导功能细化的可推广音频Deepfake检测链接:https://arxiv.org/abs/2602.13596
备注:Under Review
摘要:随着deepfake音频变得更加真实和多样化,开发可推广的对策系统变得至关重要。现有的检测方法主要依赖于XLS-R前端特征来提高泛化能力。尽管如此,它们的性能仍然有限,部分原因是对细粒度信息(如生理线索或频域特征)的关注不够。在本文中,我们提出了BreathNet,这是一种新型的音频深度伪造检测框架,它集成了细粒度的呼吸信息以提高泛化能力。具体来说,我们设计BreathFilm,一个功能明智的线性调制机制,选择性地放大基于呼吸声的存在的时间表示。BreathFiLM与XLS-R提取器联合训练,进而鼓励提取器学习呼吸相关线索并将其编码到时间特征中。然后,我们使用频率前端提取频谱特征,然后与时间特征融合,以提供由声码器或压缩伪影引入的补充信息。此外,我们提出了一组特征损失,包括仅正监督对比损失(PSCL),中心损失和对比度损失。这些损失共同增强了区分能力,鼓励模型在特征空间中更有效地分离真实和深度伪造样本。在五个基准数据集上进行的大量实验证明了最先进的(SOTA)性能。使用ASVspoof 2019 LA训练集,我们的方法在四个相关的评估基准中达到了1.99%的平均EER,在野外数据集上的性能尤其强劲,达到了4.70%的EER。此外,在ASVspoof 5评估协议下,我们的方法在这个最新的基准测试中实现了4.94%的EER。
摘要:As deepfake audio becomes more realistic and diverse, developing generalizable countermeasure systems has become crucial. Existing detection methods primarily depend on XLS-R front-end features to improve generalization. Nonetheless, their performance remains limited, partly due to insufficient attention to fine-grained information, such as physiological cues or frequency-domain features. In this paper, we propose BreathNet, a novel audio deepfake detection framework that integrates fine-grained breath information to improve generalization. Specifically, we design BreathFiLM, a feature-wise linear modulation mechanism that selectively amplifies temporal representations based on the presence of breathing sounds. BreathFiLM is trained jointly with the XLS-R extractor, in turn encouraging the extractor to learn and encode breath-related cues into the temporal features. Then, we use the frequency front-end to extract spectral features, which are then fused with temporal features to provide complementary information introduced by vocoders or compression artifacts. Additionally, we propose a group of feature losses comprising Positive-only Supervised Contrastive Loss (PSCL), center loss, and contrast loss. These losses jointly enhance the discriminative ability, encouraging the model to separate bona fide and deepfake samples more effectively in the feature space. Extensive experiments on five benchmark datasets demonstrate state-of-the-art (SOTA) performance. Using the ASVspoof 2019 LA training set, our method attains 1.99% average EER across four related eval benchmarks, with particularly strong performance on the In-the-Wild dataset, where it achieves 4.70% EER. Moreover, under the ASVspoof5 evaluation protocol, our method achieves an EER of 4.94% on this latest benchmark.
【15】Multimodal Consistency-Guided Reference-Free Data Selection for ASR Accent Adaptation
标题:用于ASB口音适应的多模式一致性引导的无参考数据选择链接:https://arxiv.org/abs/2602.13263
摘要:自动语音识别(ASR)系统通常会在带口音的语音上降级,因为声学语音和韵律移位会导致训练数据不匹配,从而使标记的口音适应成本高昂。然而,常见的伪标签选择算法在很大程度上是以文本为中心的(例如,复杂度(PPL)滤波),并且可能偏好流畅但声学上不匹配的假设,从而在微调时导致误差放大。为了解决这个问题,我们引入了一个多模式的一致性指导,无参考的数据选择管道的ASR口音适应下的转导,无标签协议。该流水线从基于子模块互信息的目标感知预选步骤开始,以提高查询相关性并减少下游计算。然后,它通过基于扰动的解码生成每个话语的多个伪transmits,并使用两个无参考信号对每个假设进行评分:共享嵌入空间中的语音-文本对齐和预测单词错误率(WER)。一个简单的基于语音的选择规则保留可靠的伪标签进行微调,同时丢弃嘈杂的话语。在域内设置中,从30 k池中选择约1.5k话语实现了10.91%的WER,接近使用30 k监督标签获得的10.45%。在具有不匹配的候选池的跨域设置中,一致性过滤的子集避免了在强重音移位下由未过滤的伪标签引起的退化,并且在更强的ASR骨干上的匹配小时实验进一步确认了随机采样和最近选择基线的增益。
摘要:Automatic speech recognition (ASR) systems often degrade on accented speech because acoustic-phonetic and prosodic shifts induce a mismatch to training data, making labeled accent adaptation costly. However, common pseudo-label selection heuristics are largely text-centric (e.g., perplexity (PPL) filtering) and can prefer fluent yet acoustically mismatched hypotheses, leading to error amplification when fine-tuning. To address this, we introduce a multimodal consistency-guided, reference-free data selection pipeline for ASR accent adaptation under a transductive, label-free protocol. The pipeline starts with a target-aware preselection step based on submodular mutual information to improve query relevance and reduce downstream computation. It then generates multiple pseudo-transcriptions per utterance via perturbation-based decoding and scores each hypothesis using two reference-free signals: speech--text alignment in a shared embedding space and predicted word error rate (WER). A simple percentile-based selection rule retains reliable pseudo-labels for fine-tuning while discarding noisy utterances. In an in-domain setting, selecting ~1.5k utterances from a 30k pool achieves 10.91% WER, close to 10.45% obtained using 30k supervised labels. In a cross-domain setting with a mismatched candidate pool, consistency-filtered subsets avoid the degradation caused by unfiltered pseudo-labels under strong accent shift, and matched-hour experiments on a stronger ASR backbone further confirm gains over random sampling and recent selection baselines.
【16】Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
标题:学习基于生理学的声乐光谱时间表示用于语音情感识别链接:https://arxiv.org/abs/2602.13259
备注:13 pages, 5 figures
摘要:语音情感识别(SER)对于社交机器人交互和机器人心理诊断等类人机器人任务至关重要,其中可解释和有效的模型对于安全性和性能至关重要。在大型数据集上训练的现有深度模型在很大程度上仍然无法解释,通常无法充分建模潜在的情感声学信号,并且无法捕获和分析情感发声行为的核心生理学。对人声的生理研究表明,人声的幅度和相位的动态变化通过声道滤波器和声门源与情绪相关。然而,大多数现有的深度模型仅涉及幅度,但未能耦合幅度和相位的生理特征以及幅度和相位之间的生理特征。在这里,我们提出了PhysioSER,一个生理学知情的声乐spectrotemporal表示学习方法,以解决这些问题与紧凑,即插即用的设计。PhysioSER构建了基于语音解剖学和生理学(VAP)的幅度和相位视图,以补充用于SER的SSL模型。该VAP通知框架包含两个并行工作流程:基于VAP的语音特征表示分支,用于分解语音信号,将它们嵌入到四元数场中,并使用Hamilton结构的四元数卷积来建模它们的动态交互;以及基于冻结SSL主干的潜在表示分支。然后,从两个工作流程的话语级特征对齐的对比投影和对齐框架,其次是一个浅注意力融合头SER分类。通过对14个数据集、10种语言和6个主干的广泛评估,PhysioSER对SER来说是可解释的和有效的,其实际功效通过在人形机器人平台上的实时部署得到了验证。
摘要:Speech emotion recognition (SER) is essential for humanoid robot tasks such as social robotic interactions and robotic psychological diagnosis, where interpretable and efficient models are critical for safety and performance. Existing deep models trained on large datasets remain largely uninterpretable, often insufficiently modeling underlying emotional acoustic signals and failing to capture and analyze the core physiology of emotional vocal behaviors. Physiological research on human voices shows that the dynamics of vocal amplitude and phase correlate with emotions through the vocal tract filter and the glottal source. However, most existing deep models solely involve amplitude but fail to couple the physiological features of and between amplitude and phase. Here, we propose PhysioSER, a physiology-informed vocal spectrotemporal representation learning method, to address these issues with a compact, plug-and-play design. PhysioSER constructs amplitude and phase views informed by voice anatomy and physiology (VAP) to complement SSL models for SER. This VAP-informed framework incorporates two parallel workflows: a vocal feature representation branch to decompose vocal signals based on VAP, embed them into a quaternion field, and use Hamilton-structured quaternion convolutions for modeling their dynamic interactions; and a latent representation branch based on a frozen SSL backbone. Then, utterance-level features from both workflows are aligned by a Contrastive Projection and Alignment framework, followed by a shallow attention fusion head for SER classification. PhysioSER is shown to be interpretable and efficient for SER through extensive evaluations across 14 datasets, 10 languages, and 6 backbones, and its practical efficacy is validated by real-time deployment on a humanoid robotic platform.
【17】CLAP-Based Automatic Word Naming Recognition in Post-Stroke Aphasia
标题:基于CLAP的中风后失语症自动词表识别链接:https://arxiv.org/abs/2602.14584
备注:Submitted to EUSIPCO 2026
摘要:传统的自动单词命名识别系统很难识别中风后失语症患者的单词,因为不流利和发音错误,限制了该人群的可靠自动评估。在本文中,我们提出了一种基于对比存储音频预训练(CLAP)的自动单词命名识别方法,通过利用文本音频对齐来解决这一挑战。我们的方法将单词命名识别视为音频文本匹配问题,将语音信号和文本提示投射到共享的嵌入空间中,即使在具有挑战性的录音中也能识别出预期的单词。两个语音数据集的法国中风后失语症患者进行评估,我们的方法达到了90%的准确率,优于现有的基于分类和基于自动语音识别的基线。
摘要:Conventional automatic word-naming recognition systems struggle to recognize words from post-stroke patients with aphasia because of disfluencies and mispronunciations, limiting reliable automated assessment in this population. In this paper, we propose a Contrastive Language-Audio Pretraining (CLAP) based approach for automatic word-naming recognition to address this challenge by leveraging text-audio alignment. Our approach treats word-naming recognition as an audio-text matching problem, projecting speech signals and textual prompts into a shared embedding space to identify intended words even in challenging recordings. Evaluated on two speech datasets of French post-stroke patients with aphasia, our approach achieves up to 90% accuracy, outperforming existing classification-based and automatic speech recognition-based baselines.
【18】Preliminary sonification of ENSO using traditional Javanese gamelan scales
标题:使用传统爪哇加美兰音阶对ENSO进行初步超声处理链接:https://arxiv.org/abs/2602.14560
备注:16 pages, 7 figures
摘要:声音化(将数据映射为非语音音频)为表示复杂的动态系统提供了一个尚未开发的渠道。我们对待厄尔尼诺-南方涛动(ENSO),低维气候混乱的典型例子,作为一个测试案例,通过复杂的系统诊断评估的文化定位声化。使用参数映射sonification的Niño 3.4海面温度异常指数(1870- 2024),我们编码ENSO变化到两个传统的爪哇加美兰五声系统(pelog和slendro)在四个组成策略,然后分析产生的音频作为轨迹在二维声学相空间。基于递归的诊断,凸壳几何和耦合分析表明,声化管道保留了关键的动力学特征:交替模式产生最高的轨迹递归率,呼应ENSO的准周期性;分层复调模式探索最广泛的相空间区域;这两个标度族在光谱亮度和能量之间诱导了定性上不同的耦合机制--主要是反-pelog中的相位,但在slendro中接近独立。相空间轨迹分析提供了一个严格的几何框架,在一个复杂的系统背景下比较声化设计。感知验证仍然是必要的,我们有助于动态系统的方法来评估这样的映射。
摘要:Sonification -- the mapping of data to non-speech audio -- offers an underexplored channel for representing complex dynamical systems. We treat El Niño-Southern Oscillation (ENSO), a canonical example of low-dimensional climate chaos, as a test case for culturally-situated sonification evaluated through complex systems diagnostics. Using parameter-mapping sonification of the Niño 3.4 sea surface temperature anomaly index (1870--2024), we encode ENSO variability into two traditional Javanese gamelan pentatonic systems (pelog and slendro) across four composition strategies, then analyze the resulting audio as trajectories in a two-dimensional acoustic phase space. Recurrence-based diagnostics, convex hull geometry, and coupling analysis reveal that the sonification pipeline preserves key dynamical signatures: alternating modes produce the highest trajectory recurrence rates, echoing ENSO's quasi-periodicity; layered polyphonic modes explore the broadest phase space regions; and the two scale families induce qualitatively distinct coupling regimes between spectral brightness and energy -- predominantly anti-phase in pelog but near-independent in slendro. Phase space trajectory analysis provides a rigorous geometric framework for comparing sonification designs within a complex systems context. Perceptual validation remains necessary; we contribute the dynamical systems methodology for evaluating such mappings.
标题:SA-SSL-MOS:具有频谱增强的自监督学习MOS预测,用于广义多速率语音评估
链接:https://arxiv.org/abs/2602.14785
备注:Accepted at ICASSP 2026
摘要:设计一个语音质量评估(SQA)系统来估计具有不同采样频率(16-48 kHz)的多速率语音的平均意见得分(MOS)是一项具有挑战性的任务。由于包括多速率语音样本的MOS标记的训练数据集的有限可用性,出现了挑战。虽然自监督学习(SSL)模型已被广泛用于SQA以提高性能,但一个关键的限制是它们是在16 kHz语音上进行预训练的,因此会丢弃较高采样率下的高频信息。为了解决这个问题,我们提出了一个频谱增强的SSL方法,通过并行分支架构,采用高频功能(高达48 kHz的采样率)。我们进一步介绍了一个两步训练方案:首先在一个大型48 kHz数据集上对模型进行预训练,然后在一个较小的多速率数据集上进行微调。实验结果表明,利用SSL特征忽略的高频信息对于准确的多速率SQA至关重要,并且当多速率数据有限时,所提出的两步训练大大提高了泛化能力。
摘要:Designing a speech quality assessment (SQA) system for estimating mean-opinion-score (MOS) of multi-rate speech with varying sampling frequency (16-48 kHz) is a challenging task. The challenge arises due to the limited availability of a MOS-labeled training dataset comprising multi-rate speech samples. While self-supervised learning (SSL) models have been widely adopted in SQA to boost performance, a key limitation is that they are pretrained on 16 kHz speech and therefore discard high-frequency information present in higher sampling rates. To address this issue, we propose a spectrogram-augmented SSL method that incorporates high-frequency features (up to 48 kHz sampling rate) through a parallel-branch architecture. We further introduce a two-step training scheme: the model is first pre-trained on a large 48 kHz dataset and then fine-tuned on a smaller multi-rate dataset. Experimental results show that leveraging high-frequency information overlooked by SSL features is crucial for accurate multi-rate SQA, and that the proposed two-step training substantially improves generalization when multi-rate data is limited.
【2】Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
标题:语音合成中理清音调和裂缝以保持说话者身份链接:https://arxiv.org/abs/2602.14686
摘要:我们介绍了一个系统,能够忠实地修改的感知语音质量的吱吱声,同时保留扬声器的感知身份。虽然众所周知,高吱吱声概率通常与低音调相关,但重要的是要注意,这是在说话者群体上观察到的属性,但不一定在所有情况下都成立。基于条件连续规范化流的说话人操作模块,通过对语音合成系统的训练数据集进行扩充,实现了基音与吱吱声的分离。实验表明,大大提高了说话人确认性能的范围内吱吱操纵强度。
摘要:We introduce a system capable of faithfully modifying the perceptual voice quality of creak while preserving the speaker's perceived identity. While it is well known that high creak probability is typically correlated with low pitch, it is important to note that this is a property observed on a population of speakers but does not necessarily hold across all situations. Disentanglement of pitch from creak is achieved by augmentation of the training dataset of a speech synthesis system with a speaker manipulation block based on conditional continuous normalizing flow. The experiments show greatly improved speaker verification performance over a range of creak manipulation strengths.
【3】Data Augmentation for Pathological Speech Enhancement
标题:病理性语音增强的数据增强链接:https://arxiv.org/abs/2602.14671
摘要:由于非典型的声学特性和有限的数据可用性,最先进的语音增强(SE)模型的性能大大降低病理语音。本文系统地研究了数据增强(DA)策略,以提高SE性能的病理扬声器,评估预测和生成SE模型。我们研究了三个DA类别,即,变革性、生成性和噪声增强,通过客观的SE指标评估其影响。实验结果表明,噪声增强始终提供最大和最鲁棒的增益,变换增强提供适度的改进,而生成增强产生的好处有限,并且随着合成数据量的增加会损害性能。此外,我们发现DA的有效性取决于SE模型,DA更有利于预测SE模型。虽然我们的研究结果表明,DA提高SE性能的病理扬声器,神经典型和病理语音之间的性能差距仍然存在,突出了未来的研究需要有针对性的DA策略病理语音。
摘要:The performance of state-of-the-art speech enhancement (SE) models considerably degrades for pathological speech due to atypical acoustic characteristics and limited data availability. This paper systematically investigates data augmentation (DA) strategies to improve SE performance for pathological speakers, evaluating both predictive and generative SE models. We examine three DA categories, i.e., transformative, generative, and noise augmentation, assessing their impact with objective SE metrics. Experimental results show that noise augmentation consistently delivers the largest and most robust gains, transformative augmentations provide moderate improvements, while generative augmentation yields limited benefits and can harm performance as the amount of synthetic data increases. Furthermore, we show that the effectiveness of DA varies depending on the SE model, with DA being more beneficial for predictive SE models. While our results demonstrate that DA improves SE performance for pathological speakers, a performance gap between neurotypical and pathological speech persists, highlighting the need for future research on targeted DA strategies for pathological speech.
【4】LongAudio-RAG: Event-Grounded Question Answering over Multi-Hour Long Audio
标题:LongAudio-RAG:通过多小时长音频进行基于事件的问题解答链接:https://arxiv.org/abs/2602.14612
摘要:长时间的音频在工业和消费环境中越来越普遍,但回顾多小时的录音是不切实际的,激励系统以精确的时间基础和最小的幻觉回答自然语言查询。现有的音频语言模型显示出了希望,但由于上下文长度的限制,长音频问题的回答仍然很困难。我们介绍LongAudio-RAG(LA-RAG),这是一个混合框架,它将大语言模型(LLM)输出基于检索到的带时间戳的声学事件检测,而不是原始音频。多小时流被转换为存储在SQL数据库中的结构化事件记录,在推理时,系统解析自然语言时间参考,对意图进行分类,仅检索相关事件,并使用此约束证据生成答案。为了评估性能,我们构建了一个合成的长音频基准,通过连接记录与保存的时间戳,并生成基于模板的问答对检测,计数和总结任务。最后,我们通过将其部署在混合边缘云环境中来展示我们方法的实用性,其中音频接地模型在物联网级硬件上的设备上运行,而LLM托管在GPU支持的服务器上。这种架构可以在边缘实现低延迟事件提取,并在云中实现高质量的语言推理。实验表明,结构化的,事件级检索显着提高准确性相比香草检索增强生成(RAG)或文本到SQL的方法。
摘要:Long-duration audio is increasingly common in industrial and consumer settings, yet reviewing multi-hour recordings is impractical, motivating systems that answer natural-language queries with precise temporal grounding and minimal hallucination. Existing audio-language models show promise, but long-audio question answering remains difficult due to context-length limits. We introduce LongAudio-RAG (LA-RAG), a hybrid framework that grounds Large Language Model (LLM) outputs in retrieved, timestamped acoustic event detections rather than raw audio. Multi-hour streams are converted into structured event records stored in an SQL database, and at inference time the system resolves natural-language time references, classifies intent, retrieves only the relevant events, and generates answers using this constrained evidence. To evaluate performance, we construct a synthetic long-audio benchmark by concatenating recordings with preserved timestamps and generating template-based question-answer pairs for detection, counting, and summarization tasks. Finally, we demonstrate the practicality of our approach by deploying it in a hybrid edge-cloud environment, where the audio grounding model runs on-device on IoT-class hardware while the LLM is hosted on a GPU-backed server. This architecture enables low-latency event extraction at the edge and high-quality language reasoning in the cloud. Experiments show that structured, event-level retrieval significantly improves accuracy compared to vanilla Retrieval-Augmented Generation (RAG) or text-to-SQL approaches.
【5】CLAP-Based Automatic Word Naming Recognition in Post-Stroke Aphasia
标题:基于CLAP的中风后失语症自动词表识别链接:https://arxiv.org/abs/2602.14584
备注:Submitted to EUSIPCO 2026
摘要:传统的自动单词命名识别系统很难识别中风后失语症患者的单词,因为不流利和发音错误,限制了该人群的可靠自动评估。在本文中,我们提出了一种基于对比存储音频预训练(CLAP)的自动单词命名识别方法,通过利用文本音频对齐来解决这一挑战。我们的方法将单词命名识别视为音频文本匹配问题,将语音信号和文本提示投射到共享的嵌入空间中,即使在具有挑战性的录音中也能识别出预期的单词。两个语音数据集的法国中风后失语症患者进行评估,我们的方法达到了90%的准确率,优于现有的基于分类和基于自动语音识别的基线。
摘要:Conventional automatic word-naming recognition systems struggle to recognize words from post-stroke patients with aphasia because of disfluencies and mispronunciations, limiting reliable automated assessment in this population. In this paper, we propose a Contrastive Language-Audio Pretraining (CLAP) based approach for automatic word-naming recognition to address this challenge by leveraging text-audio alignment. Our approach treats word-naming recognition as an audio-text matching problem, projecting speech signals and textual prompts into a shared embedding space to identify intended words even in challenging recordings. Evaluated on two speech datasets of French post-stroke patients with aphasia, our approach achieves up to 90% accuracy, outperforming existing classification-based and automatic speech recognition-based baselines.
【6】ELEAT-SAGA: Early & Late Integration with Evading Alternating Training for Spoof-Robust Speaker Verification
标题:ELEAT-SAGA:早期和晚期集成与避免交替训练,以实现欺骗稳健的说话人验证链接:https://arxiv.org/abs/2602.13761
摘要:欺骗鲁棒自动说话人验证(SASV)旨在建立自动说话人验证系统,该系统对零努力冒名顶替者攻击和复杂的欺骗技术(如语音转换(VC)和文本到语音(TTS))都具有鲁棒性。在这项工作中,我们提出了一种新的SASV架构,引入分数感知门控注意力(SAGA),SASV-SAGA,使动态调制的扬声器嵌入的基础上对策(CM)分数。通过分别从预训练的ECAPA-TDNN和AASIST模型中整合说话人嵌入和CM分数,我们探索了几种整合策略,包括早期,晚期和完全整合。我们进一步介绍了多模块交替训练(ATMM)和一个改进的变体,回避交替训练(EAT)。ASVspoof 2019 Logical Access(LA)和Spoofceleb数据集的实验结果表明,与基线相比有了显着改进,在ASVspoof 2019评估集上实现了1.22%的欺骗感知说话人验证等误率(SASV-EER)和0.0304的最小归一化不可知检测成本函数(min a-DCF)。这些结果证实了分数感知注意机制和交替训练策略在增强SASV系统鲁棒性方面的有效性。
摘要:Spoofing-robust automatic speaker verification (SASV) seeks to build automatic speaker verification systems that are robust against both zero-effort impostor attacks and sophisticated spoofing techniques such as voice conversion (VC) and text-to-speech (TTS). In this work, we propose a novel SASV architecture that introduces score-aware gated attention (SAGA), SASV-SAGA, enabling dynamic modulation of speaker embeddings based on countermeasure (CM) scores. By integrating speaker embeddings and CM scores from pre-trained ECAPA-TDNN and AASIST models respectively, we explore several integration strategies including early, late, and full integration. We further introduce alternating training for multi-module (ATMM) and a refined variant, evading alternating training (EAT). Experimental results on the ASVspoof 2019 Logical Access (LA) and Spoofceleb datasets demonstrate significant improvements over baselines, achieving a spoofing aware speaker verification equal error rate (SASV-EER) of 1.22% and minimum normalized agnostic detection cost function (min a-DCF) of 0.0304 on the ASVspoof 2019 evaluation set. These results confirm the effectiveness of score-aware attention mechanisms and alternating training strategies in enhancing the robustness of SASV systems.
【7】Bengali-Loop: Community Benchmarks for Long-Form Bangla ASR and Speaker Diarization
标题:Bengali-Loop:长格式孟加拉语ASB和扬声器拨号的社区基准链接:https://arxiv.org/abs/2602.14291
摘要:孟加拉语(孟加拉语)尽管被广泛使用,但在长格式语音技术方面仍然资源不足。我们提出了孟加拉循环,两个社区基准来解决这个差距:(1)191个录音的长格式ASR语料库(158.6小时,79.2万字)来自11个YouTube频道,通过可复制的字幕提取管道和人工参与的转录验证收集;以及(2)24个录音(22小时,5,744个注释片段)的说话者日记语料库,其具有CSV格式的完全手动的说话者转向标签。这两个基准都针对真实的多扬声器、长持续时间内容(例如,Bangla drama/natok).我们建立基线(Tugstugi:34.07% WER; pyannote.audio:40.08% DER),并提供标准化的评估协议(WER/CER,DER),注释规则和数据格式,以支持孟加拉语长格式ASR和日记的可重复基准测试和未来模型开发。
摘要:Bengali (Bangla) remains under-resourced in long-form speech technology despite its wide use. We present Bengali-Loop, two community benchmarks to address this gap: (1) a long-form ASR corpus of 191 recordings (158.6 hours, 792k words) from 11 YouTube channels, collected via a reproducible subtitle-extraction pipeline and human-in-the-loop transcript verification; and (2) a speaker diarization corpus of 24 recordings (22 hours, 5,744 annotated segments) with fully manual speaker-turn labels in CSV format. Both benchmarks target realistic multi-speaker, long-duration content (e.g., Bangla drama/natok). We establish baselines (Tugstugi: 34.07% WER; pyannote.audio: 40.08% DER) and provide standardized evaluation protocols (WER/CER, DER), annotation rules, and data formats to support reproducible benchmarking and future model development for Bangla long-form ASR and diarization.
【8】Learning Vocal-Tract Area and Radiation with a Physics-Informed Webster Model
标题:使用了解物理情况的韦伯斯特模型学习气道区域和辐射链接:https://arxiv.org/abs/2602.13834
备注:Accepted at IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026
摘要:我们提出了一个物理通知有声后端渲染器唱歌的声音合成。给定合成单声道音频和基频轨迹,我们训练时域韦伯斯特模型作为物理信息的神经网络来估计可解释的声道面积函数和开口辐射系数。训练强制执行偏微分方程和边界一致性;轻量级DDSP路径仅用于稳定学习,而推理纯粹基于物理。在持续元音(/a/,/i/,/u/),由独立的时域有限差分韦伯斯特求解器呈现的参数与紧凑的DDSP基线竞争性地再现频谱包络,并且在离散化、适度的源变化和大约百分之十的音高偏移的变化下保持稳定。图内波形保持比参考呼吸,激励在未来的工作中意识到目标和明确的声门先验。
摘要:We present a physics-informed voiced backend renderer for singing-voice synthesis. Given synthetic single-channel audio and a fund-amental--frequency trajectory, we train a time-domain Webster model as a physics-informed neural network to estimate an interpretable vocal-tract area function and an open-end radiation coefficient. Training enforces partial differential equation and boundary consistency; a lightweight DDSP path is used only to stabilize learning, while inference is purely physics-based. On sustained vowels (/a/, /i/, /u/), parameters rendered by an independent finite-difference time-domain Webster solver reproduce spectral envelopes competitively with a compact DDSP baseline and remain stable under changes in discretization, moderate source variations, and about ten percent pitch shifts. The in-graph waveform remains breathier than the reference, motivating periodicity-aware objectives and explicit glottal priors in future work.
【9】Enhancing spatial hearing with cochlear implants: exploring the role of AI, multimodal interaction and perceptual training
标题:通过人工智能增强空间听力:探索人工智能、多模式交互和感知训练的作用链接:https://arxiv.org/abs/2602.13787
摘要:髋关节植入物(CI)已经发展到可以恢复大部分患者的听力和语言理解的程度。虽然空间听觉是控制和引导注意力以及在嘈杂环境中实现言语理解的核心,但在过去它在很大程度上被忽视了。在这里,我们提出了一个多学科的研究框架,医生,心理学家和工程师合作,以提高CI用户的空间听力。
摘要:Cochlear implants (CIs) have been developed to the point where they can restore hearing and speech understanding in a large proportion of patients. Although spatial hearing is central to controlling and directing attention and to enabling speech understanding in noisy environments, it has been largely neglected in the past. We propose here a multi-disciplinary research framework in which physicians, psychologists and engineers collaborate to improve spatial hearing for CI users.
【10】BreathNet: Generalizable Audio Deepfake Detection via Breath-Cue-Guided Feature Refinement
标题:BreathNet:通过呼吸提示引导功能细化的可推广音频Deepfake检测链接:https://arxiv.org/abs/2602.13596
备注:Under Review
摘要:随着deepfake音频变得更加真实和多样化,开发可推广的对策系统变得至关重要。现有的检测方法主要依赖于XLS-R前端特征来提高泛化能力。尽管如此,它们的性能仍然有限,部分原因是对细粒度信息(如生理线索或频域特征)的关注不够。在本文中,我们提出了BreathNet,这是一种新型的音频深度伪造检测框架,它集成了细粒度的呼吸信息以提高泛化能力。具体来说,我们设计BreathFilm,一个功能明智的线性调制机制,选择性地放大基于呼吸声的存在的时间表示。BreathFiLM与XLS-R提取器联合训练,进而鼓励提取器学习呼吸相关线索并将其编码到时间特征中。然后,我们使用频率前端提取频谱特征,然后与时间特征融合,以提供由声码器或压缩伪影引入的补充信息。此外,我们提出了一组特征损失,包括仅正监督对比损失(PSCL),中心损失和对比度损失。这些损失共同增强了区分能力,鼓励模型在特征空间中更有效地分离真实和深度伪造样本。在五个基准数据集上进行的大量实验证明了最先进的(SOTA)性能。使用ASVspoof 2019 LA训练集,我们的方法在四个相关的评估基准中达到了1.99%的平均EER,在野外数据集上的性能尤其强劲,达到了4.70%的EER。此外,在ASVspoof 5评估协议下,我们的方法在这个最新的基准测试中实现了4.94%的EER。
摘要:As deepfake audio becomes more realistic and diverse, developing generalizable countermeasure systems has become crucial. Existing detection methods primarily depend on XLS-R front-end features to improve generalization. Nonetheless, their performance remains limited, partly due to insufficient attention to fine-grained information, such as physiological cues or frequency-domain features. In this paper, we propose BreathNet, a novel audio deepfake detection framework that integrates fine-grained breath information to improve generalization. Specifically, we design BreathFiLM, a feature-wise linear modulation mechanism that selectively amplifies temporal representations based on the presence of breathing sounds. BreathFiLM is trained jointly with the XLS-R extractor, in turn encouraging the extractor to learn and encode breath-related cues into the temporal features. Then, we use the frequency front-end to extract spectral features, which are then fused with temporal features to provide complementary information introduced by vocoders or compression artifacts. Additionally, we propose a group of feature losses comprising Positive-only Supervised Contrastive Loss (PSCL), center loss, and contrast loss. These losses jointly enhance the discriminative ability, encouraging the model to separate bona fide and deepfake samples more effectively in the feature space. Extensive experiments on five benchmark datasets demonstrate state-of-the-art (SOTA) performance. Using the ASVspoof 2019 LA training set, our method attains 1.99% average EER across four related eval benchmarks, with particularly strong performance on the In-the-Wild dataset, where it achieves 4.70% EER. Moreover, under the ASVspoof5 evaluation protocol, our method achieves an EER of 4.94% on this latest benchmark.
【11】Fast Swap-Based Element Selection for Multiplication-Free Dimension Reduction
标题:基于快速交换的元素选择以实现无乘降维链接:https://arxiv.org/abs/2602.13532
备注:11 pages, 4 figures
摘要:在本文中,我们提出了一个快速算法的元素选择,乘法自由形式的降维,产生一个降维向量,只需从输入中选择一个子集的元素。降维是减少不必要的模型参数、减轻过拟合以及加速训练和推理的基本技术。一种标准的方法是主成分分析(PCA),但PCA依赖于矩阵乘法;在资源受限的系统中,乘法计数本身可能成为瓶颈。元素选择消除了这种成本,因为减少只包括选择元素,因此关键的挑战是确定哪些元素应该被保留。我们通过线性回归的最小均方误差来评估候选子集,该线性回归从所选元素预测目标向量,其中目标可以是例如分类中的独热标签向量。当显式目标不可用时,输入本身可以用作目标,从而产生基于重构的标准。由此产生的优化是组合的,穷举搜索是不切实际的。为了解决这个问题,我们推导出一个有效的公式交换所选择的和一个非线性元素所造成的客观变化,使用矩阵求逆引理,我们执行一个基于交换的局部搜索,反复应用目标减少交换,直到没有进一步的改进是可能的。在MNIST手写数字图像上的实验证明了该方法的有效性。
摘要:In this paper, we propose a fast algorithm for element selection, a multiplication-free form of dimension reduction that produces a dimension-reduced vector by simply selecting a subset of elements from the input. Dimension reduction is a fundamental technique for reducing unnecessary model parameters, mitigating overfitting, and accelerating training and inference. A standard approach is principal component analysis (PCA), but PCA relies on matrix multiplications; on resource-constrained systems, the multiplication count itself can become a bottleneck. Element selection eliminates this cost because the reduction consists only of selecting elements, and thus the key challenge is to determine which elements should be retained. We evaluate a candidate subset through the minimum mean-squared error of linear regression that predicts a target vector from the selected elements, where the target may be, for example, a one-hot label vector in classification. When an explicit target is unavailable, the input itself can be used as the target, yielding a reconstruction-based criterion. The resulting optimization is combinatorial, and exhaustive search is impractical. To address this, we derive an efficient formula for the objective change caused by swapping a selected and an unselected element, using the matrix inversion lemma, and we perform a swap-based local search that repeatedly applies objective-decreasing swaps until no further improvement is possible. Experiments on MNIST handwritten-digit images demonstrate the effectiveness of the proposed method.
【12】Multimodal Consistency-Guided Reference-Free Data Selection for ASR Accent Adaptation
标题:用于ASB口音适应的多模式一致性引导的无参考数据选择链接:https://arxiv.org/abs/2602.13263
摘要:自动语音识别(ASR)系统通常会在带口音的语音上降级,因为声学语音和韵律移位会导致训练数据不匹配,从而使标记的口音适应成本高昂。然而,常见的伪标签选择算法在很大程度上是以文本为中心的(例如,复杂度(PPL)滤波),并且可能偏好流畅但声学上不匹配的假设,从而在微调时导致误差放大。为了解决这个问题,我们引入了一个多模式的一致性指导,无参考的数据选择管道的ASR口音适应下的转导,无标签协议。该流水线从基于子模块互信息的目标感知预选步骤开始,以提高查询相关性并减少下游计算。然后,它通过基于扰动的解码生成每个话语的多个伪transmits,并使用两个无参考信号对每个假设进行评分:共享嵌入空间中的语音-文本对齐和预测单词错误率(WER)。一个简单的基于语音的选择规则保留可靠的伪标签进行微调,同时丢弃嘈杂的话语。在域内设置中,从30 k池中选择约1.5k话语实现了10.91%的WER,接近使用30 k监督标签获得的10.45%。在具有不匹配的候选池的跨域设置中,一致性过滤的子集避免了在强重音移位下由未过滤的伪标签引起的退化,并且在更强的ASR骨干上的匹配小时实验进一步确认了随机采样和最近选择基线的增益。
摘要:Automatic speech recognition (ASR) systems often degrade on accented speech because acoustic-phonetic and prosodic shifts induce a mismatch to training data, making labeled accent adaptation costly. However, common pseudo-label selection heuristics are largely text-centric (e.g., perplexity (PPL) filtering) and can prefer fluent yet acoustically mismatched hypotheses, leading to error amplification when fine-tuning. To address this, we introduce a multimodal consistency-guided, reference-free data selection pipeline for ASR accent adaptation under a transductive, label-free protocol. The pipeline starts with a target-aware preselection step based on submodular mutual information to improve query relevance and reduce downstream computation. It then generates multiple pseudo-transcriptions per utterance via perturbation-based decoding and scores each hypothesis using two reference-free signals: speech--text alignment in a shared embedding space and predicted word error rate (WER). A simple percentile-based selection rule retains reliable pseudo-labels for fine-tuning while discarding noisy utterances. In an in-domain setting, selecting ~1.5k utterances from a 30k pool achieves 10.91% WER, close to 10.45% obtained using 30k supervised labels. In a cross-domain setting with a mismatched candidate pool, consistency-filtered subsets avoid the degradation caused by unfiltered pseudo-labels under strong accent shift, and matched-hour experiments on a stronger ASR backbone further confirm gains over random sampling and recent selection baselines.
【13】Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
标题:学习基于生理学的声乐光谱时间表示用于语音情感识别链接:https://arxiv.org/abs/2602.13259
备注:13 pages, 5 figures
摘要:语音情感识别(SER)对于社交机器人交互和机器人心理诊断等类人机器人任务至关重要,其中可解释和有效的模型对于安全性和性能至关重要。在大型数据集上训练的现有深度模型在很大程度上仍然无法解释,通常无法充分建模潜在的情感声学信号,并且无法捕获和分析情感发声行为的核心生理学。对人声的生理研究表明,人声的幅度和相位的动态变化通过声道滤波器和声门源与情绪相关。然而,大多数现有的深度模型仅涉及幅度,但未能耦合幅度和相位的生理特征以及幅度和相位之间的生理特征。在这里,我们提出了PhysioSER,一个生理学知情的声乐spectrotemporal表示学习方法,以解决这些问题与紧凑,即插即用的设计。PhysioSER构建了基于语音解剖学和生理学(VAP)的幅度和相位视图,以补充用于SER的SSL模型。该VAP通知框架包含两个并行工作流程:基于VAP的语音特征表示分支,用于分解语音信号,将它们嵌入到四元数场中,并使用Hamilton结构的四元数卷积来建模它们的动态交互;以及基于冻结SSL主干的潜在表示分支。然后,从两个工作流程的话语级特征对齐的对比投影和对齐框架,其次是一个浅注意力融合头SER分类。通过对14个数据集、10种语言和6个主干的广泛评估,PhysioSER对SER来说是可解释的和有效的,其实际功效通过在人形机器人平台上的实时部署得到了验证。
摘要:Speech emotion recognition (SER) is essential for humanoid robot tasks such as social robotic interactions and robotic psychological diagnosis, where interpretable and efficient models are critical for safety and performance. Existing deep models trained on large datasets remain largely uninterpretable, often insufficiently modeling underlying emotional acoustic signals and failing to capture and analyze the core physiology of emotional vocal behaviors. Physiological research on human voices shows that the dynamics of vocal amplitude and phase correlate with emotions through the vocal tract filter and the glottal source. However, most existing deep models solely involve amplitude but fail to couple the physiological features of and between amplitude and phase. Here, we propose PhysioSER, a physiology-informed vocal spectrotemporal representation learning method, to address these issues with a compact, plug-and-play design. PhysioSER constructs amplitude and phase views informed by voice anatomy and physiology (VAP) to complement SSL models for SER. This VAP-informed framework incorporates two parallel workflows: a vocal feature representation branch to decompose vocal signals based on VAP, embed them into a quaternion field, and use Hamilton-structured quaternion convolutions for modeling their dynamic interactions; and a latent representation branch based on a frozen SSL backbone. Then, utterance-level features from both workflows are aligned by a Contrastive Projection and Alignment framework, followed by a shallow attention fusion head for SER classification. PhysioSER is shown to be interpretable and efficient for SER through extensive evaluations across 14 datasets, 10 languages, and 6 backbones, and its practical efficacy is validated by real-time deployment on a humanoid robotic platform.
机器翻译由腾讯交互翻译提供,仅供参考
