今日论文合集:cs.SD语音9篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
标题:AudioMarathon:音频LLM中长上下文音频理解和效率的综合基准
链接:https://arxiv.org/abs/2510.07293

作者:Peize He, Zichen Wen, Yubo Wang, Yuxuan Wang, Xiaoqian Liu, Jiajie Huang, Zehui Lei, Zhuangcheng Gu, Xiangqi Jin, Jiabing Yang, Kai Li, Zhifei Liu, Weijia Li, Cunxiang Wang, Conghui He, Linfeng Zhang
备注:26 pages, 23 figures, the code is available at \url{this https URL}
摘要:处理长格式音频是大型音频语言模型(LALM)的主要挑战。这些模型与注意力的二次成本($O(N^2)$)和建模长期时间依赖性作斗争。现有的音频基准大多是从短片段构建的,并且不评估现实长上下文设置中的模型。为了解决这一差距,我们引入AudioMarathon,这是一个旨在评估长格式音频的理解和推理效率的基准测试。AudioMarathon提供了基于三大支柱的各种任务:持续时间从90.0秒到300.0秒的长上下文音频输入,分别对应于2,250到7,500个音频令牌的编码序列,语音,声音和音乐的全域覆盖,以及需要多跳推理的复杂推理。我们评估了最先进的LALM,并观察到随着音频长度的增加,性能明显下降。我们还研究了加速技术,并分析了令牌修剪和KV缓存驱逐的权衡。结果表明,目前的LALM之间存在很大的差距,并强调需要更好的时间推理和内存效率的架构。我们相信AudioMarathon将推动音频和多模态研究社区开发能够解决复杂音频任务的更先进的音频理解模型。
摘要:Processing long-form audio is a major challenge for Large Audio Language models (LALMs). These models struggle with the quadratic cost of attention ($O(N^2)$) and with modeling long-range temporal dependencies. Existing audio benchmarks are built mostly from short clips and do not evaluate models in realistic long context settings. To address this gap, we introduce AudioMarathon, a benchmark designed to evaluate both understanding and inference efficiency on long-form audio. AudioMarathon provides a diverse set of tasks built upon three pillars: long-context audio inputs with durations ranging from 90.0 to 300.0 seconds, which correspond to encoded sequences of 2,250 to 7,500 audio tokens, respectively, full domain coverage across speech, sound, and music, and complex reasoning that requires multi-hop inference. We evaluate state-of-the-art LALMs and observe clear performance drops as audio length grows. We also study acceleration techniques and analyze the trade-offs of token pruning and KV cache eviction. The results show large gaps across current LALMs and highlight the need for better temporal reasoning and memory-efficient architectures. We believe AudioMarathon will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.


【2】Making Machines Sound Sarcastic: LLM-Enhanced and Retrieval-Guided Sarcastic Speech Synthesis
标题:让机器听起来讽刺:LLM增强和检索引导讽刺语音合成
链接:https://arxiv.org/abs/2510.07096

作者:Zhu Li, Yuqing Zhang, Xiyuan Gao, Shekhar Nayak, Matt Coler
摘要:讽刺是一种微妙的非字面语言形式,由于其依赖于微妙的语义,上下文和韵律线索,对语音合成构成了重大挑战。虽然现有的语音合成研究主要集中在广泛的情感类别上,但讽刺在很大程度上仍未得到探索。在本文中,我们提出了一个大语言模型(LLM)增强检索增强框架的语音感知语音合成。我们的方法结合了(1)语义嵌入从LoRA微调LLaMA 3,捕捉语用不一致和话语层面的讽刺线索,(2)韵律样本检索通过检索增强生成(RAG)模块,提供表达参考模式的讽刺交付。集成在VITS主干中,这种双重条件作用使讽刺言语更自然,更适合上下文。实验表明,我们的方法优于基线的客观措施和主观评价,提高语音自然度,讽刺的表现力,和下游的讽刺检测。
摘要:Sarcasm is a subtle form of non-literal language that poses significant challenges for speech synthesis due to its reliance on nuanced semantic, contextual, and prosodic cues. While existing speech synthesis research has focused primarily on broad emotional categories, sarcasm remains largely unexplored. In this paper, we propose a Large Language Model (LLM)-enhanced Retrieval-Augmented framework for sarcasm-aware speech synthesis. Our approach combines (1) semantic embeddings from a LoRA-fine-tuned LLaMA 3, which capture pragmatic incongruity and discourse-level cues of sarcasm, and (2) prosodic exemplars retrieved via a Retrieval Augmented Generation (RAG) module, which provide expressive reference patterns of sarcastic delivery. Integrated within a VITS backbone, this dual conditioning enables more natural and contextually appropriate sarcastic speech. Experiments demonstrate that our method outperforms baselines in both objective measures and subjective evaluations, yielding improvements in speech naturalness, sarcastic expressivity, and downstream sarcasm detection.


【3】Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
标题:开放的ASB排行榜:迈向可重复和透明的多语言和长形式语音识别评估
链接:https://arxiv.org/abs/2510.06961

作者:Vaibhav Srivastav, Steven Zheng, Eric Bezzam, Eustache Le Bihan, Nithin Koluguri, Piotr Żelasko, Somshubra Majumdar, Adel Moumen, Sanchit Gandhi
备注:Submitted to ICASSP 2026; Leaderboard: this https URL Code: this https URL
摘要:尽管进展迅速,但ASR评估仍然充斥着简短的英语,效率很少报告。我们推出了开放式ASR排行榜,这是一个完全可复制的基准和交互式排行榜,比较了11个数据集的60多个开源和专有系统,包括专用的多语言和长格式跟踪。我们标准化文本规范化并报告单词错误率(WER)和反向实时因子(RTFx),从而实现公平的准确性-效率比较。对于英语转录,与LLM解码器配对的Conformer编码器实现了最佳的平均WER,但速度较慢,而CTC和TDT解码器提供了更好的RTFx,使其对长格式和离线使用具有吸引力。耳语衍生的编码器针对英语进行了微调,提高了准确性,但往往会牺牲多语言覆盖率。所有代码和数据集加载器都是开源的,以支持透明的,可扩展的评估。
摘要:Despite rapid progress, ASR evaluation remains saturated with short-form English, and efficiency is rarely reported. We present the Open ASR Leaderboard, a fully reproducible benchmark and interactive leaderboard comparing 60+ open-source and proprietary systems across 11 datasets, including dedicated multilingual and long-form tracks. We standardize text normalization and report both word error rate (WER) and inverse real-time factor (RTFx), enabling fair accuracy-efficiency comparisons. For English transcription, Conformer encoders paired with LLM decoders achieve the best average WER but are slower, while CTC and TDT decoders deliver much better RTFx, making them attractive for long-form and offline use. Whisper-derived encoders fine-tuned for English improve accuracy but often trade off multilingual coverage. All code and dataset loaders are open-sourced to support transparent, extensible evaluation.


【4】XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
标题:XLSR-Kanformer:合成语音检测的KAN集成模型
链接:https://arxiv.org/abs/2510.06706

作者:Phuong Tuan Dat, Tran Huy Dat
备注:Accepted to 2025 IEEE International Conference on Advanced Video and Signal-Based Surveillance
摘要:语音合成技术的最新发展导致了越来越复杂的欺骗攻击,对自动说话人确认系统提出了重大挑战。虽然基于自监督学习(SSL)模型的系统,特别是XLSR-Conformer架构,在合成语音检测方面表现出了卓越的性能,但仍有架构改进的空间。在本文中,我们提出了一种新的方法,取代了传统的多层感知器(MLP)的XLSR-构象模型与Kolmogorov-Arnold网络(KAN),一个强大的通用近似的基础上Kolmogorov-Arnold表示定理。我们在ASVspoof 2021上的实验结果表明,KAN到XLSR-Conformer模型的集成可以在等错误率(EER)LA和DF集上相对提高60.55%的性能,在21 LA集上进一步达到0.70%的EER。此外,所提出的替代方案对各种SSL架构也具有鲁棒性。这些研究结果表明,将KAN到基于SSL的模型是一个很有前途的方向,在合成语音检测的进步。
摘要:Recent advancements in speech synthesis technologies have led to increasingly sophisticated spoofing attacks, posing significant challenges for automatic speaker verification systems. While systems based on self-supervised learning (SSL) models, particularly the XLSR-Conformer architecture, have demonstrated remarkable performance in synthetic speech detection, there remains room for architectural improvements. In this paper, we propose a novel approach that replaces the traditional Multi-Layer Perceptron (MLP) in the XLSR-Conformer model with a Kolmogorov-Arnold Network (KAN), a powerful universal approximator based on the Kolmogorov-Arnold representation theorem. Our experimental results on ASVspoof2021 demonstrate that the integration of KAN to XLSR-Conformer model can improve the performance by 60.55% relatively in Equal Error Rate (EER) LA and DF sets, further achieving 0.70% EER on the 21LA set. Besides, the proposed replacement is also robust to various SSL architectures. These findings suggest that incorporating KAN into SSL-based models is a promising direction for advances in synthetic speech detection.


【5】Pitch Estimation With Mean Averaging Smoothed Product Spectrum And Musical Consonance Evaluation Using MASP
标题:平均平均平滑产物谱的音调估计和使用MASP的音乐协和评估
链接:https://arxiv.org/abs/2510.06625

作者:Murat Yasar Baskin
摘要:这项研究介绍了平均平滑产品(MASP)频谱,这是一个修改版本的谐波产品频谱,旨在提高音调估计的许多算法明智的欺骗性频谱,仍然导致清晰的音高,谐波和非谐波的情况下。通过引入基于全局均值的光谱平滑,MASP算法减少了HPS对缺失分光的光谱的不必要的敏感性。该方法表现出强大的音高估计与感性的期望。基于和谐和周期性之间的强相关性,相同的算法被扩展,并且随着和谐性度量(H)的提出,用于评估两个和三个音调的音乐和谐;产生与音乐理论的感知和实践相一致的和谐层次。这些发现表明,音高和协和音的感知可能有着相似的基础机制,依赖于频谱。
摘要:This study introduces Mean Averaging Smoothed Product (MASP) Spectrum, which is a modified version of the Harmonic Product Spectrum, designed to enhance pitch estimation for many algorithm-wise deceptive frequency spectra that still lead clear pitches, for both harmonic and inharmonic cases. By introducing a global mean based smoothing for spectrum, the MASP algorithm diminishes the unwanted sensitivity of HPS for spectra with missing partials. The method exhibited robust pitch estimations consistent with perceptual expectations. Motivated upon the strong correlation between consonance and periodicity, the same algorithm is extended and, with the proposition of a harmonicity measure (H), used to evaluate musical consonance for two and three tones; yielding consonance hierarchies that align with perception and practice of music theory. These findings suggest that perception of pitch and consonance may share a similar underlying mechanism that depend on spectrum.


【6】Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race
标题:在假语音一代军备竞赛中对假语音检测进行基准测试
链接:https://arxiv.org/abs/2510.06544

作者:Xutao Mao, Ke Li, Cameron Baird, Ezra Xuanru Tao, Dan Lin
摘要:随着合成语音生成技术的加速发展,越来越多的假语音生成器出现,产生的音频通常与真实的人类语音无法区分。这一演变对录音作为关键证据的各个部门构成了新的严重威胁。虽然假语音检测器也在进步,但假语音生成和检测之间的军备竞赛变得更加激烈和复杂。在这项工作中,我们提出了第一个大规模的,跨域的评估假语音检测器,基准8个国家的最先进的模型对数据集合成的20个不同的假语音生成系统。据我们所知,这是迄今为止进行的最全面的跨领域评估。我们的研究揭示了当前虚假语音检测系统中存在的大量安全漏洞,强调了其在现实世界鲁棒性方面的关键差距。为了推进这一领域,我们提出了一个统一而有效的指标,巩固了以前在不同研究中使用的多样化且往往不一致的评价标准。该指标可以对虚假语音检测器的鲁棒性进行标准化、直接的比较。最后,我们为构建更具弹性的虚假语音检测技术提供了可行的建议,其更广泛的目标是加强人工智能安全性和可信度的基础。
摘要:As advances in synthetic voice generation accelerate, an increasing variety of fake voice generators have emerged, producing audio that is often indistinguishable from real human speech. This evolution poses new and serious threats across sectors where audio recordings serve as critical evidence. Although fake voice detectors are also advancing, the arms race between fake voice generation and detection has become more intense and complex. In this work, we present the first large-scale, cross-domain evaluation of fake voice detectors, benchmarking 8 state-of-the-art models against datasets synthesized by 20 different fake voice generation systems. To the best of our knowledge, this is the most comprehensive cross-domain assessment conducted to date. Our study reveals substantial security vulnerabilities in current fake voice detection systems, underscoring critical gaps in their real-world robustness. To advance the field, we propose a unified and effective metric that consolidates the diverse and often inconsistent evaluation criteria previously used across different studies. This metric enables standardized, straightforward comparisons of the robustness of fake voice detectors. We conclude by offering actionable recommendations for building more resilient fake voice detection technologies, with the broader goal of reinforcing the foundations of AI security and trustworthiness.


【7】BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music
标题:BACHI:通过流行音乐和古典音乐的掩蔽迭代解码实现边界意识符号和弦识别
链接:https://arxiv.org/abs/2510.06528

作者:Mingyang Yao, Ke Chen, Shlomo Dubnov, Taylor Berg-Kirkpatrick
备注:Under review
摘要:通过深度学习模型进行的自动和弦识别(ACR)已经逐渐实现了有希望的识别精度,但仍然存在两个关键挑战。首先,先前的工作主要集中在音频域ACR上,而符号音乐(例如,由于数据稀缺,ACR受到的关注有限。其次,现有的方法仍然忽略了与人类音乐分析实践相一致的策略。为了应对这些挑战,我们做了两个贡献:(1)我们引入了POP 909-CL,它是POP 909数据集的增强版本,具有与节奏一致的内容和人类校正的和弦,节拍,基调和时间签名标签;(2)提出了BACHI符号和弦识别模型,该模型将和弦识别任务分解为不同的决策步骤,即边界检测和和弦根、质量低音(倒置)。这种机制反映了人类的听力训练实践。实验表明,BACHI在古典音乐和流行音乐基准上都达到了最先进的和弦识别性能,消融研究验证了每个模块的有效性。
摘要:Automatic chord recognition (ACR) via deep learning models has gradually achieved promising recognition accuracy, yet two key challenges remain. First, prior work has primarily focused on audio-domain ACR, while symbolic music (e.g., score) ACR has received limited attention due to data scarcity. Second, existing methods still overlook strategies that are aligned with human music analytical practices. To address these challenges, we make two contributions: (1) we introduce POP909-CL, an enhanced version of POP909 dataset with tempo-aligned content and human-corrected labels of chords, beats, keys, and time signatures; and (2) We propose BACHI, a symbolic chord recognition model that decomposes the task into different decision steps, namely boundary detection and iterative ranking of chord root, quality, and bass (inversion). This mechanism mirrors the human ear-training practices. Experiments demonstrate that BACHI achieves state-of-the-art chord recognition performance on both classical and pop music benchmarks, with ablation studies validating the effectiveness of each module.


【8】Comparison of Speech Tasks in Human Expert and Machine Detection of Parkinson's Disease
标题:帕金森病人类专家和机器检测中语音任务的比较
链接:https://arxiv.org/abs/2510.07299

作者:Peter Plantinga, Roozbeh Sattari, Karine Marcotte, Carla Di Gironimo, Madeleine Sharp, Liziane Bouvier, Maiya Geddes, Ingrid Verduyckt, Étienne de Villers-Sidani, Mirco Ravanelli, Denise Klein
备注:Accepted to SMASH 2025
摘要:帕金森氏病(PD)患者的语言已被证明是关于疾病存在和进展的重要线索。我们调查的因素的基础上,人类专家作出判断的存在疾病的语音样本超过五个不同的语音任务:发音,句子重复,阅读,回忆和图片描述。我们通过进行听力测试来进行比较,以确定临床医生仅从音频中识别PD体征的准确性,并且我们使用机器学习系统进行基于Whisper的检测实验。在所有任务中,当只有音频可用时,Whisper的表现与人类专家相当或更好,特别是在具有挑战性但重要的数据亚组:年轻患者,轻度病例和女性患者。Whisper在困难情况下识别声学线索的能力补充了人类专家的多模态和上下文优势。
摘要:The speech of people with Parkinson's Disease (PD) has been shown to hold important clues about the presence and progression of the disease. We investigate the factors based on which humans experts make judgments of the presence of disease in speech samples over five different speech tasks: phonations, sentence repetition, reading, recall, and picture description. We make comparisons by conducting listening tests to determine clinicians accuracy at recognizing signs of PD from audio alone, and we conduct experiments with a machine learning system for detection based on Whisper. Across tasks, Whisper performs on par or better than human experts when only audio is available, especially on challenging but important subgroups of the data: younger patients, mild cases, and female patients. Whisper's ability to recognize acoustic cues in difficult cases complements the multimodal and contextual strengths of human experts.


【9】Moises-Light: Resource-efficient Band-split U-Net For Music Source Separation
标题:Moises-Light:资源高效的带宽分离U-Net,用于音乐源分离
链接:https://arxiv.org/abs/2510.06785

作者:Yun-Ning (Amy)Hung, Igor Pereira, Filip Korzeniowski
摘要:近年来,音乐源分离取得了重大进展,双通道建模、频带分离模块或Transformer层等模型架构取得了非常好的效果。然而,这些模型通常包含大量参数,在训练和实际应用方面对计算资源有限的设备构成挑战。虽然已经引入了一些轻量级模型,但与较大的模型相比,它们的性能通常较差。在本文中,我们从这些最新的进展,以改善轻量级模型的灵感。我们证明,通过精心设计,轻量级模型可以实现与具有多达13倍参数的模型相当的SDR。我们提出的模型,Moises-Light,在MUSDB-HQ基准数据集上分离四个音乐干时取得了有竞争力的结果。当使用MoisesDB作为额外的训练数据时,所提出的模型还展示了具有竞争力的可扩展性。
摘要:In recent years, significant advances have been made in music source separation, with model architectures such as dual-path modeling, band-split modules, or transformer layers achieving comparably good results. However, these models often contain a significant number of parameters, posing challenges to devices with limited computational resources in terms of training and practical application. While some lightweight models have been introduced, they generally perform worse compared to their larger counterparts. In this paper, we take inspiration from these recent advances to improve a lightweight model. We demonstrate that with careful design, a lightweight model can achieve comparable SDRs to models with up to 13 times more parameters. Our proposed model, Moises-Light, achieves competitive results in separating four musical stems on the MUSDB-HQ benchmark dataset. The proposed model also demonstrates competitive scalability when using MoisesDB as additional training data.


eess.AS音频处理


【1】Comparison of Speech Tasks in Human Expert and Machine Detection of Parkinson's Disease
标题:帕金森病人类专家和机器检测中语音任务的比较
链接:https://arxiv.org/abs/2510.07299

作者:Peter Plantinga, Roozbeh Sattari, Karine Marcotte, Carla Di Gironimo, Madeleine Sharp, Liziane Bouvier, Maiya Geddes, Ingrid Verduyckt, Étienne de Villers-Sidani, Mirco Ravanelli, Denise Klein
备注:Accepted to SMASH 2025
摘要:帕金森氏病(PD)患者的语言已被证明是关于疾病存在和进展的重要线索。我们调查的因素的基础上,人类专家作出判断的存在疾病的语音样本超过五个不同的语音任务:发音,句子重复,阅读,回忆和图片描述。我们通过进行听力测试来进行比较,以确定临床医生仅从音频中识别PD体征的准确性,并且我们使用机器学习系统进行基于Whisper的检测实验。在所有任务中,当只有音频可用时,Whisper的表现与人类专家相当或更好,特别是在具有挑战性但重要的数据亚组:年轻患者,轻度病例和女性患者。Whisper在困难情况下识别声学线索的能力补充了人类专家的多模态和上下文优势。
摘要:The speech of people with Parkinson's Disease (PD) has been shown to hold important clues about the presence and progression of the disease. We investigate the factors based on which humans experts make judgments of the presence of disease in speech samples over five different speech tasks: phonations, sentence repetition, reading, recall, and picture description. We make comparisons by conducting listening tests to determine clinicians accuracy at recognizing signs of PD from audio alone, and we conduct experiments with a machine learning system for detection based on Whisper. Across tasks, Whisper performs on par or better than human experts when only audio is available, especially on challenging but important subgroups of the data: younger patients, mild cases, and female patients. Whisper's ability to recognize acoustic cues in difficult cases complements the multimodal and contextual strengths of human experts.


【2】Towards Responsible Evaluation for Text-to-Speech
标题:迈向负责任的文本转语音评估
链接:https://arxiv.org/abs/2510.06927

作者:Yifan Yang, Hui Wang, Bing Han, Shujie Liu, Jinyu Li, Yong Qin, Xie Chen
摘要:文本到语音(TTS)技术的最新进展使系统能够产生人类无法区分的语音,从而在可访问性、内容创建和人机交互方面带来好处。然而,目前的评价做法越来越不足以全面反映能力、局限性和社会影响。本文介绍了负责任评价的概念,认为负责任评价是下一阶段文语转换系统发展的必要和紧迫的任务,它分为三个层次:(1)确保真实准确地反映模型的真实能力,采用更稳健、更有区别和更全面的客观和主观评分方法;(2)通过标准化基准、透明的报告和可转移的评估指标实现可比性、标准化和可转移性;(3)评估和减轻与伪造、滥用、侵犯隐私和安全漏洞相关的道德风险。通过这一概念,我们批判性地审视当前的评估实践,找出系统性的缺陷,并提出可操作的建议。我们希望这种负责任评估的概念能够促进更加值得信赖和可靠的TTS技术,并引导其朝着符合道德规范和对社会有益的应用方向发展。
摘要:Recent advances in text-to-speech (TTS) technology have enabled systems to produce human-indistinguishable speech, bringing benefits across accessibility, content creation, and human-computer interaction. However, current evaluation practices are increasingly inadequate for capturing the full range of capabilities, limitations, and societal implications. This position paper introduces the concept of Responsible Evaluation and argues that it is essential and urgent for the next phase of TTS development, structured through three progressive levels: (1) ensuring the faithful and accurate reflection of a model's true capabilities, with more robust, discriminative, and comprehensive objective and subjective scoring methodologies; (2) enabling comparability, standardization, and transferability through standardized benchmarks, transparent reporting, and transferable evaluation metrics; and (3) assessing and mitigating ethical risks associated with forgery, misuse, privacy violations, and security vulnerabilities. Through this concept, we critically examine current evaluation practices, identify systemic shortcomings, and propose actionable recommendations. We hope this concept of Responsible Evaluation will foster more trustworthy and reliable TTS technology and guide its development toward ethically sound and societally beneficial applications.


【3】Moises-Light: Resource-efficient Band-split U-Net For Music Source Separation
标题:Moises-Light:资源高效的带宽分离U-Net,用于音乐源分离
链接:https://arxiv.org/abs/2510.06785

作者:Yun-Ning (Amy)Hung, Igor Pereira, Filip Korzeniowski
摘要:近年来,音乐源分离取得了重大进展,双通道建模、频带分离模块或Transformer层等模型架构取得了非常好的效果。然而,这些模型通常包含大量参数,在训练和实际应用方面对计算资源有限的设备构成挑战。虽然已经引入了一些轻量级模型,但与较大的模型相比,它们的性能通常较差。在本文中,我们从这些最新的进展,以改善轻量级模型的灵感。我们证明,通过精心设计,轻量级模型可以实现与具有多达13倍参数的模型相当的SDR。我们提出的模型,Moises-Light,在MUSDB-HQ基准数据集上分离四个音乐干时取得了有竞争力的结果。当使用MoisesDB作为额外的训练数据时,所提出的模型还展示了具有竞争力的可扩展性。
摘要:In recent years, significant advances have been made in music source separation, with model architectures such as dual-path modeling, band-split modules, or transformer layers achieving comparably good results. However, these models often contain a significant number of parameters, posing challenges to devices with limited computational resources in terms of training and practical application. While some lightweight models have been introduced, they generally perform worse compared to their larger counterparts. In this paper, we take inspiration from these recent advances to improve a lightweight model. We demonstrate that with careful design, a lightweight model can achieve comparable SDRs to models with up to 13 times more parameters. Our proposed model, Moises-Light, achieves competitive results in separating four musical stems on the MUSDB-HQ benchmark dataset. The proposed model also demonstrates competitive scalability when using MoisesDB as additional training data.


【4】AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
标题:AudioMarathon:音频LLM中长上下文音频理解和效率的综合基准
链接:https://arxiv.org/abs/2510.07293

作者:Peize He, Zichen Wen, Yubo Wang, Yuxuan Wang, Xiaoqian Liu, Jiajie Huang, Zehui Lei, Zhuangcheng Gu, Xiangqi Jin, Jiabing Yang, Kai Li, Zhifei Liu, Weijia Li, Cunxiang Wang, Conghui He, Linfeng Zhang
备注:26 pages, 23 figures, the code is available at \url{this https URL}
摘要:处理长格式音频是大型音频语言模型(LALM)的主要挑战。这些模型与注意力的二次成本($O(N^2)$)和建模长期时间依赖性作斗争。现有的音频基准大多是从短片段构建的,并且不评估现实长上下文设置中的模型。为了解决这一差距,我们引入AudioMarathon,这是一个旨在评估长格式音频的理解和推理效率的基准测试。AudioMarathon提供了基于三大支柱的各种任务:持续时间从90.0秒到300.0秒的长上下文音频输入,分别对应于2,250到7,500个音频令牌的编码序列,语音,声音和音乐的全域覆盖,以及需要多跳推理的复杂推理。我们评估了最先进的LALM,并观察到随着音频长度的增加,性能明显下降。我们还研究了加速技术,并分析了令牌修剪和KV缓存驱逐的权衡。结果表明,目前的LALM之间存在很大的差距,并强调需要更好的时间推理和内存效率的架构。我们相信AudioMarathon将推动音频和多模态研究社区开发能够解决复杂音频任务的更先进的音频理解模型。
摘要:Processing long-form audio is a major challenge for Large Audio Language models (LALMs). These models struggle with the quadratic cost of attention ($O(N^2)$) and with modeling long-range temporal dependencies. Existing audio benchmarks are built mostly from short clips and do not evaluate models in realistic long context settings. To address this gap, we introduce AudioMarathon, a benchmark designed to evaluate both understanding and inference efficiency on long-form audio. AudioMarathon provides a diverse set of tasks built upon three pillars: long-context audio inputs with durations ranging from 90.0 to 300.0 seconds, which correspond to encoded sequences of 2,250 to 7,500 audio tokens, respectively, full domain coverage across speech, sound, and music, and complex reasoning that requires multi-hop inference. We evaluate state-of-the-art LALMs and observe clear performance drops as audio length grows. We also study acceleration techniques and analyze the trade-offs of token pruning and KV cache eviction. The results show large gaps across current LALMs and highlight the need for better temporal reasoning and memory-efficient architectures. We believe AudioMarathon will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.


【5】Making Machines Sound Sarcastic: LLM-Enhanced and Retrieval-Guided Sarcastic Speech Synthesis
标题:让机器听起来讽刺:LLM增强和检索引导讽刺语音合成
链接:https://arxiv.org/abs/2510.07096

作者:Zhu Li, Yuqing Zhang, Xiyuan Gao, Shekhar Nayak, Matt Coler
摘要:讽刺是一种微妙的非字面语言形式,由于其依赖于微妙的语义,上下文和韵律线索,对语音合成构成了重大挑战。虽然现有的语音合成研究主要集中在广泛的情感类别上,但讽刺在很大程度上仍未得到探索。在本文中,我们提出了一个大语言模型(LLM)增强检索增强框架的语音感知语音合成。我们的方法结合了(1)语义嵌入从LoRA微调LLaMA 3,捕捉语用不一致和话语层面的讽刺线索,(2)韵律样本检索通过检索增强生成(RAG)模块,提供表达参考模式的讽刺交付。集成在VITS主干中,这种双重条件作用使讽刺言语更自然,更适合上下文。实验表明,我们的方法优于基线的客观措施和主观评价,提高语音自然度,讽刺的表现力,和下游的讽刺检测。
摘要:Sarcasm is a subtle form of non-literal language that poses significant challenges for speech synthesis due to its reliance on nuanced semantic, contextual, and prosodic cues. While existing speech synthesis research has focused primarily on broad emotional categories, sarcasm remains largely unexplored. In this paper, we propose a Large Language Model (LLM)-enhanced Retrieval-Augmented framework for sarcasm-aware speech synthesis. Our approach combines (1) semantic embeddings from a LoRA-fine-tuned LLaMA 3, which capture pragmatic incongruity and discourse-level cues of sarcasm, and (2) prosodic exemplars retrieved via a Retrieval Augmented Generation (RAG) module, which provide expressive reference patterns of sarcastic delivery. Integrated within a VITS backbone, this dual conditioning enables more natural and contextually appropriate sarcastic speech. Experiments demonstrate that our method outperforms baselines in both objective measures and subjective evaluations, yielding improvements in speech naturalness, sarcastic expressivity, and downstream sarcasm detection.


【6】Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
标题:开放的ASB排行榜:迈向可重复和透明的多语言和长形式语音识别评估
链接:https://arxiv.org/abs/2510.06961

作者:Vaibhav Srivastav, Steven Zheng, Eric Bezzam, Eustache Le Bihan, Nithin Koluguri, Piotr Żelasko, Somshubra Majumdar, Adel Moumen, Sanchit Gandhi
备注:Submitted to ICASSP 2026; Leaderboard: this https URL Code: this https URL
摘要:尽管进展迅速,但ASR评估仍然充斥着简短的英语,效率很少报告。我们推出了开放式ASR排行榜,这是一个完全可复制的基准和交互式排行榜,比较了11个数据集的60多个开源和专有系统,包括专用的多语言和长格式跟踪。我们标准化文本规范化并报告单词错误率(WER)和反向实时因子(RTFx),从而实现公平的准确性-效率比较。对于英语转录,与LLM解码器配对的Conformer编码器实现了最佳的平均WER,但速度较慢,而CTC和TDT解码器提供了更好的RTFx,使其对长格式和离线使用具有吸引力。耳语衍生的编码器针对英语进行了微调,提高了准确性,但往往会牺牲多语言覆盖率。所有代码和数据集加载器都是开源的,以支持透明的,可扩展的评估。
摘要:Despite rapid progress, ASR evaluation remains saturated with short-form English, and efficiency is rarely reported. We present the Open ASR Leaderboard, a fully reproducible benchmark and interactive leaderboard comparing 60+ open-source and proprietary systems across 11 datasets, including dedicated multilingual and long-form tracks. We standardize text normalization and report both word error rate (WER) and inverse real-time factor (RTFx), enabling fair accuracy-efficiency comparisons. For English transcription, Conformer encoders paired with LLM decoders achieve the best average WER but are slower, while CTC and TDT decoders deliver much better RTFx, making them attractive for long-form and offline use. Whisper-derived encoders fine-tuned for English improve accuracy but often trade off multilingual coverage. All code and dataset loaders are open-sourced to support transparent, extensible evaluation.


【7】SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
标题:SHANKS:口语模型的同时听力和思维
链接:https://arxiv.org/abs/2510.06917

作者:Cheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin, Kevin Lin, Shujie Liu, Zhendong Wang, Zhengyuan Yang, Hung-yi Lee, Lijuan Wang
备注:Work in progress
摘要:当前的大型语言模型(LLM)和口语模型(SLM)只有在用户完成他们的回合之后才开始思考和采取行动。这阻止了模型在用户的回合期间进行交互,并且在它等待思考时可能导致高响应延迟。因此,在接收到完整输入后进行思考不适合语音到语音交互,其中实时,低延迟交换很重要。我们通过注意到人类自然地“边听边思考”来解决这个问题。“在本文中,我们提出了SHANKS,一个通用的推理框架,使SLM生成潜台词的思维链推理,同时听取用户的输入。SHANKS将输入语音以固定持续时间的组块进行流式传输,并且一旦接收到组块,就基于所有先前的语音和推理生成未说出的推理,同时用户继续说话。SHANKS使用这种不言而喻的推理来决定是否打断用户并进行工具调用以完成任务。我们证明了SHANKS在两种情况下增强了用户与SLM的实时交互:(1)当用户提出一个数学问题的逐步解决方案时,SHANKS可以倾听,推理,并在用户犯错误时打断,比不加思考地打断的基线高出37.1%;(2)在工具增强对话中,SHANKS可以在用户完成其回合之前完成56.9%的工具调用。总的来说,SHANKS倾向于在整个对话过程中不断思考的模式,而不仅仅是在一个回合结束后。Shanks的动画插图可以在https://d223302.github.io/SHANKS/上找到
摘要:Current large language models (LLMs) and spoken language models (SLMs) begin thinking and taking actions only after the user has finished their turn. This prevents the model from interacting during the user's turn and can lead to high response latency while it waits to think. Consequently, thinking after receiving the full input is not suitable for speech-to-speech interaction, where real-time, low-latency exchange is important. We address this by noting that humans naturally "think while listening." In this paper, we propose SHANKS, a general inference framework that enables SLMs to generate unspoken chain-of-thought reasoning while listening to the user input. SHANKS streams the input speech in fixed-duration chunks and, as soon as a chunk is received, generates unspoken reasoning based on all previous speech and reasoning, while the user continues speaking. SHANKS uses this unspoken reasoning to decide whether to interrupt the user and to make tool calls to complete the task. We demonstrate that SHANKS enhances real-time user-SLM interaction in two scenarios: (1) when the user is presenting a step-by-step solution to a math problem, SHANKS can listen, reason, and interrupt when the user makes a mistake, achieving 37.1% higher interruption accuracy than a baseline that interrupts without thinking; and (2) in a tool-augmented dialogue, SHANKS can complete 56.9% of the tool calls before the user finishes their turn. Overall, SHANKS moves toward models that keep thinking throughout the conversation, not only after a turn ends. Animated illustrations of Shanks can be found at https://d223302.github.io/SHANKS/


【8】XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
标题:XLSR-Kanformer:合成语音检测的KAN集成模型
链接:https://arxiv.org/abs/2510.06706

作者:Phuong Tuan Dat, Tran Huy Dat
备注:Accepted to 2025 IEEE International Conference on Advanced Video and Signal-Based Surveillance
摘要:语音合成技术的最新发展导致了越来越复杂的欺骗攻击,对自动说话人确认系统提出了重大挑战。虽然基于自监督学习(SSL)模型的系统,特别是XLSR-Conformer架构,在合成语音检测方面表现出了卓越的性能,但仍有架构改进的空间。在本文中,我们提出了一种新的方法,取代了传统的多层感知器(MLP)的XLSR-构象模型与Kolmogorov-Arnold网络(KAN),一个强大的通用近似的基础上Kolmogorov-Arnold表示定理。我们在ASVspoof 2021上的实验结果表明,KAN到XLSR-Conformer模型的集成可以在等错误率(EER)LA和DF集上相对提高60.55%的性能,在21 LA集上进一步达到0.70%的EER。此外,所提出的替代方案对各种SSL架构也具有鲁棒性。这些研究结果表明,将KAN到基于SSL的模型是一个很有前途的方向,在合成语音检测的进步。
摘要:Recent advancements in speech synthesis technologies have led to increasingly sophisticated spoofing attacks, posing significant challenges for automatic speaker verification systems. While systems based on self-supervised learning (SSL) models, particularly the XLSR-Conformer architecture, have demonstrated remarkable performance in synthetic speech detection, there remains room for architectural improvements. In this paper, we propose a novel approach that replaces the traditional Multi-Layer Perceptron (MLP) in the XLSR-Conformer model with a Kolmogorov-Arnold Network (KAN), a powerful universal approximator based on the Kolmogorov-Arnold representation theorem. Our experimental results on ASVspoof2021 demonstrate that the integration of KAN to XLSR-Conformer model can improve the performance by 60.55% relatively in Equal Error Rate (EER) LA and DF sets, further achieving 0.70% EER on the 21LA set. Besides, the proposed replacement is also robust to various SSL architectures. These findings suggest that incorporating KAN into SSL-based models is a promising direction for advances in synthetic speech detection.


【9】Learning to Rewrite Prompts for Bootstrapping LLMs on Downstream Tasks
标题:学习重写脚本以在下游任务上引导LLM
链接:https://arxiv.org/abs/2510.06695

作者:Qinhao Zhou, Xiang Xiang, Kun He, John E. Hopcroft
摘要:近年来,人们对大型语言模型(LLM)的兴趣日益浓厚,这极大地促进了快速工程,从手动设计过渡到基于模型的优化。LLM的指令通常包括两个组件:定义任务或目标的\textit{instruction}和针对指令类型定制的\textit{input}。在机器翻译等自然语言生成(NLG)任务中,\textit{input}组件尤为关键,而\textit{instruction}组件则趋于简洁。现有的快速工程方法主要集中在优化一般任务的\textit{instruction}组件,通常需要大参数LLM作为辅助工具。然而,这些方法对机器翻译等任务的适用性有限,其中\textit{input}组件起着更关键的作用。为了解决这个问题,本文介绍了一种新的提示优化方法,专门为机器翻译任务设计。所提出的方法采用了一个使用基于反向翻译的策略训练的小参数模型,显著降低了单任务优化的训练开销,同时提供了高效的性能。经过一定的调整,这种方法也可以扩展到其他下游任务。
摘要:In recent years, the growing interest in Large Language Models (LLMs) has significantly advanced prompt engineering, transitioning from manual design to model-based optimization. Prompts for LLMs generally comprise two components: the \textit{instruction}, which defines the task or objective, and the \textit{input}, which is tailored to the instruction type. In natural language generation (NLG) tasks such as machine translation, the \textit{input} component is particularly critical, while the \textit{instruction} component tends to be concise. Existing prompt engineering methods primarily focus on optimizing the \textit{instruction} component for general tasks, often requiring large-parameter LLMs as auxiliary tools. However, these approaches exhibit limited applicability for tasks like machine translation, where the \textit{input} component plays a more pivotal role. To address this limitation, this paper introduces a novel prompt optimization method specifically designed for machine translation tasks. The proposed approach employs a small-parameter model trained using a back-translation-based strategy, significantly reducing training overhead for single-task optimization while delivering highly effective performance. With certain adaptations, this method can also be extended to other downstream tasks.


【10】Utilizing Information Theoretic Approach to Study Cochlear Neural Degeneration
标题:利用信息论方法研究皮质神经退行性变
链接:https://arxiv.org/abs/2510.06671

作者:Ahsan J. Cheema, Sunil Puria
摘要:隐性听力损失,或耳蜗神经变性(CND),破坏阈上听觉编码,而不影响临床阈值,使其难以诊断。我们提出了一个信息理论的框架来评估语音刺激,最大限度地揭示CND通过量化内毛细胞(IHC)的受体电位和听觉神经纤维(ANF)的反应,声学输入和ANF反应之间的互信息(MI)损失。使用现象学的听觉模型,我们模拟了50个CVC单词的反应,在不同的呈现水平下,在干净,时间压缩,混响和组合条件下,低,中,高自发率纤维的生存率有系统的变化。在IHC和ANF响应之间按通道计算MI,并在特征频率之间积分。相对于正常听力基线定义信息损失。结果表明,随着CND的增加,最明显的时间压缩语音进行性MI损失,而混响产生相对较小的影响。这些研究结果确定快速,时间密集的语音作为CND的最佳探针,为客观临床诊断的设计提供信息,同时揭示与混响作为探针相关的问题。
摘要:Hidden hearing loss, or cochlear neural degeneration (CND), disrupts suprathreshold auditory coding without affecting clinical thresholds, making it difficult to diagnose. We present an information-theoretic framework to evaluate speech stimuli that maximally reveal CND by quantifying mutual information (MI) loss between inner hair cell (IHC) receptor potentials and auditory nerve fiber (ANF) responses and acoustic input and ANF responses. Using a phenomenological auditory model, we simulated responses to 50 CVC words under clean, time-compressed, reverberant, and combined conditions across different presentation levels, with systematically varied survival of low-, medium-, and high-spontaneous-rate fibers. MI was computed channel-wise between IHC and ANF responses and integrated across characteristic frequencies. Information loss was defined relative to a normal-hearing baseline. Results demonstrate progressive MI loss with increasing CND, most pronounced for time-compressed speech, while reverberation produced comparatively smaller effects. These findings identify rapid, temporally dense speech as optimal probes for CND, informing the design of objective clinical diagnostics while revealing problems associated with reverberation as a probe.


【11】Pitch Estimation With Mean Averaging Smoothed Product Spectrum And Musical Consonance Evaluation Using MASP
标题:平均平均平滑产物谱的音调估计和使用MASP的音乐协和评估
链接:https://arxiv.org/abs/2510.06625

作者:Murat Yasar Baskin
摘要:这项研究介绍了平均平滑产品(MASP)频谱,这是一个修改版本的谐波产品频谱,旨在提高音调估计的许多算法明智的欺骗性频谱,仍然导致清晰的音高,谐波和非谐波的情况下。通过引入基于全局均值的光谱平滑,MASP算法减少了HPS对缺失分光的光谱的不必要的敏感性。该方法表现出强大的音高估计与感性的期望。基于和谐和周期性之间的强相关性,相同的算法被扩展,并且随着和谐性度量(H)的提出,用于评估两个和三个音调的音乐和谐;产生与音乐理论的感知和实践相一致的和谐层次。这些发现表明,音高和协和音的感知可能有着相似的基础机制,依赖于频谱。
摘要:This study introduces Mean Averaging Smoothed Product (MASP) Spectrum, which is a modified version of the Harmonic Product Spectrum, designed to enhance pitch estimation for many algorithm-wise deceptive frequency spectra that still lead clear pitches, for both harmonic and inharmonic cases. By introducing a global mean based smoothing for spectrum, the MASP algorithm diminishes the unwanted sensitivity of HPS for spectra with missing partials. The method exhibited robust pitch estimations consistent with perceptual expectations. Motivated upon the strong correlation between consonance and periodicity, the same algorithm is extended and, with the proposition of a harmonicity measure (H), used to evaluate musical consonance for two and three tones; yielding consonance hierarchies that align with perception and practice of music theory. These findings suggest that perception of pitch and consonance may share a similar underlying mechanism that depend on spectrum.


【12】Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race
标题:在假语音一代军备竞赛中对假语音检测进行基准测试
链接:https://arxiv.org/abs/2510.06544

作者:Xutao Mao, Ke Li, Cameron Baird, Ezra Xuanru Tao, Dan Lin
摘要:随着合成语音生成技术的加速发展,越来越多的假语音生成器出现,产生的音频通常与真实的人类语音无法区分。这一演变对录音作为关键证据的各个部门构成了新的严重威胁。虽然假语音检测器也在进步,但假语音生成和检测之间的军备竞赛变得更加激烈和复杂。在这项工作中,我们提出了第一个大规模的,跨域的评估假语音检测器,基准8个国家的最先进的模型对数据集合成的20个不同的假语音生成系统。据我们所知,这是迄今为止进行的最全面的跨领域评估。我们的研究揭示了当前虚假语音检测系统中的重大安全漏洞,强调了其在现实世界中的鲁棒性方面的关键差距。为了推进这一领域,我们提出了一个统一而有效的指标,巩固了以前在不同研究中使用的多样化且往往不一致的评价标准。该指标可以对虚假语音检测器的鲁棒性进行标准化、直接的比较。最后,我们为构建更具弹性的虚假语音检测技术提供了可行的建议,其更广泛的目标是加强人工智能安全性和可信度的基础。
摘要:As advances in synthetic voice generation accelerate, an increasing variety of fake voice generators have emerged, producing audio that is often indistinguishable from real human speech. This evolution poses new and serious threats across sectors where audio recordings serve as critical evidence. Although fake voice detectors are also advancing, the arms race between fake voice generation and detection has become more intense and complex. In this work, we present the first large-scale, cross-domain evaluation of fake voice detectors, benchmarking 8 state-of-the-art models against datasets synthesized by 20 different fake voice generation systems. To the best of our knowledge, this is the most comprehensive cross-domain assessment conducted to date. Our study reveals substantial security vulnerabilities in current fake voice detection systems, underscoring critical gaps in their real-world robustness. To advance the field, we propose a unified and effective metric that consolidates the diverse and often inconsistent evaluation criteria previously used across different studies. This metric enables standardized, straightforward comparisons of the robustness of fake voice detectors. We conclude by offering actionable recommendations for building more resilient fake voice detection technologies, with the broader goal of reinforcing the foundations of AI security and trustworthiness.


【13】BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music
标题:BACHI:通过流行音乐和古典音乐的掩蔽迭代解码实现边界意识符号和弦识别
链接:https://arxiv.org/abs/2510.06528

作者:Mingyang Yao, Ke Chen, Shlomo Dubnov, Taylor Berg-Kirkpatrick
备注:Under review
摘要:通过深度学习模型进行的自动和弦识别(ACR)已经逐渐实现了有希望的识别精度,但仍然存在两个关键挑战。首先,先前的工作主要集中在音频域ACR上,而符号音乐(例如,由于数据稀缺,ACR受到的关注有限。其次,现有的方法仍然忽略了与人类音乐分析实践相一致的策略。为了应对这些挑战,我们做了两个贡献:(1)我们引入了POP 909-CL,它是POP 909数据集的增强版本,具有与节奏一致的内容和人类校正的和弦,节拍,基调和时间签名标签;(2)提出了BACHI符号和弦识别模型,该模型将和弦识别任务分解为不同的决策步骤,即边界检测和和弦根、质量低音(倒置)。这种机制反映了人类的听力训练实践。实验表明,BACHI在古典音乐和流行音乐基准上都达到了最先进的和弦识别性能,消融研究验证了每个模块的有效性。
摘要:Automatic chord recognition (ACR) via deep learning models has gradually achieved promising recognition accuracy, yet two key challenges remain. First, prior work has primarily focused on audio-domain ACR, while symbolic music (e.g., score) ACR has received limited attention due to data scarcity. Second, existing methods still overlook strategies that are aligned with human music analytical practices. To address these challenges, we make two contributions: (1) we introduce POP909-CL, an enhanced version of POP909 dataset with tempo-aligned content and human-corrected labels of chords, beats, keys, and time signatures; and (2) We propose BACHI, a symbolic chord recognition model that decomposes the task into different decision steps, namely boundary detection and iterative ranking of chord root, quality, and bass (inversion). This mechanism mirrors the human ear-training practices. Experiments demonstrate that BACHI achieves state-of-the-art chord recognition performance on both classical and pop music benchmarks, with ablation studies validating the effectiveness of each module.


机器翻译由腾讯交互翻译提供,仅供参考