今日论文合集:cs.SD语音12篇,eess.AS音频处理8篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Spectral and Rhythm Feature Performance Evaluation for Category and Class Level Audio Classification with Deep Convolutional Neural Networks
标题:使用深度卷积神经网络进行类别和类别级音频分类的频谱和节奏特征性能评估
链接:https://arxiv.org/abs/2509.07756

作者: Friedrich Wolf-Monheim
摘要:除了决策树和k近邻算法之外,深度卷积神经网络(CNN)被广泛用于对音乐、语音或环境声音等许多领域的音频数据进行分类。为了训练特定的CNN,可以将各种频谱和节奏特征(如梅尔缩放频谱图、梅尔频率倒谱系数(MFCC)、循环温度图、短时傅里叶变换(STFT)色度图、恒定Q变换(CQT)色度图和色度能量归一化统计(CENS)色度图)用作神经网络的数字图像输入数据。使用深度CNN和ESC-50数据集,使用端到端深度学习管道,详细研究了这些频谱和节奏特征在音频类别级别以及音频类级别分类中的性能,其中ESC-50数据集包含2,000个标记的环境音频记录。多类分类的评估指标准确度、精确度、召回率和F1得分清楚地表明,梅尔缩放频谱图和梅尔频率倒谱系数(MFCC)的表现明显优于本研究中使用深度CNN进行音频分类任务所研究的其他频谱和节奏特征。
摘要:Next to decision tree and k-nearest neighbours algorithms deep convolutional neural networks (CNNs) are widely used to classify audio data in many domains like music, speech or environmental sounds. To train a specific CNN various spectral and rhythm features like mel-scaled spectrograms, mel-frequency cepstral coefficients (MFCC), cyclic tempograms, short-time Fourier transform (STFT) chromagrams, constant-Q transform (CQT) chromagrams and chroma energy normalized statistics (CENS) chromagrams can be used as digital image input data for the neural network. The performance of these spectral and rhythm features for audio category level as well as audio class level classification is investigated in detail with a deep CNN and the ESC-50 dataset with 2,000 labeled environmental audio recordings using an end-to-end deep learning pipeline. The evaluated metrics accuracy, precision, recall and F1 score for multiclass classification clearly show that the mel-scaled spectrograms and the mel-frequency cepstral coefficients (MFCC) perform significantly better then the other spectral and rhythm features investigated in this research for audio classification tasks using deep CNNs.


【2】Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
标题:频谱掩蔽和内插攻击(SMIA):针对语音认证和反欺骗系统的黑匣子对抗攻击
链接:https://arxiv.org/abs/2509.07677

作者:Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal
摘要:语音认证系统(VAS)使用独特的声音特征进行验证。他们越来越多地融入银行和医疗保健等高安全性行业。尽管他们使用深度学习进行了改进,但他们面临着来自深度伪造和对抗性攻击等复杂威胁的严重漏洞。真实声音克隆的出现使检测复杂化,因为系统很难区分真实的合成音频。虽然存在反欺骗对策(CM)来减轻这些风险,但许多依赖于静态检测模型,这些模型可以被新的对抗方法绕过,从而留下关键的安全漏洞。为了证明这种脆弱性,我们提出了频谱掩蔽和插值攻击(SMIA),这是一种新的方法,可以策略性地操纵人工智能生成的音频的听不见的频率区域。通过改变人耳难以察觉的区域中的声音,SMIA创建了听起来真实的对抗样本,同时欺骗CM。我们在模拟的真实世界条件下,对跨多个任务的最先进(SOTA)模型进行了全面的评估。SMIA针对组合VAS/CM系统实现了至少82%的强大攻击成功率(ASR),针对独立说话人验证系统至少97.5%,针对对策至少100%。这些发现最终表明,当前的安全态势不足以应对自适应对抗性攻击。这项工作强调了迫切需要向下一代防御模式转变,这些防御采用能够随着威胁形势而发展的动态上下文感知框架。
摘要:Voice Authentication Systems (VAS) use unique vocal characteristics for verification. They are increasingly integrated into high-security sectors such as banking and healthcare. Despite their improvements using deep learning, they face severe vulnerabilities from sophisticated threats like deepfakes and adversarial attacks. The emergence of realistic voice cloning complicates detection, as systems struggle to distinguish authentic from synthetic audio. While anti-spoofing countermeasures (CMs) exist to mitigate these risks, many rely on static detection models that can be bypassed by novel adversarial methods, leaving a critical security gap. To demonstrate this vulnerability, we propose the Spectral Masking and Interpolation Attack (SMIA), a novel method that strategically manipulates inaudible frequency regions of AI-generated audio. By altering the voice in imperceptible zones to the human ear, SMIA creates adversarial samples that sound authentic while deceiving CMs. We conducted a comprehensive evaluation of our attack against state-of-the-art (SOTA) models across multiple tasks, under simulated real-world conditions. SMIA achieved a strong attack success rate (ASR) of at least 82% against combined VAS/CM systems, at least 97.5% against standalone speaker verification systems, and 100% against countermeasures. These findings conclusively demonstrate that current security postures are insufficient against adaptive adversarial attacks. This work highlights the urgent need for a paradigm shift toward next-generation defenses that employ dynamic, context-aware frameworks capable of evolving with the threat landscape.


【3】Neural Proxies for Sound Synthesizers: Learning Perceptually Informed Preset Representations
标题:声音合成器的神经代理:学习感知知情的预设表示
链接:https://arxiv.org/abs/2509.07635

作者:Paolo Combes, Stefan Weinzierl, Klaus Obermayer
备注:17 pages, 4 figures, published in the Journal of the Audio   Engineering Society
摘要:深度学习似乎是自动合成器编程(ASP)的一个有吸引力的解决方案,旨在帮助音乐家和声音设计师对声音合成器进行编程。然而,由于其潜在的不可微性,将软件合成器集成到训练管道中具有挑战性。这项工作通过引入一种方法来近似任意合成器来应对这一挑战。具体来说,我们训练一个神经网络,以将合成器映射到从预训练模型导出的音频嵌入空间。这有助于定义产生紧凑而有效的表示的神经代理,从而能够将音频嵌入损失集成到黑盒合成器的基于神经的ASP系统中。我们在基于神经的nASP的背景下评估了各种预训练音频模型的表示,并评估了几种神经网络架构的有效性,包括前馈,递归和基于transformer的模型,在定义神经代理。我们评估所提出的方法使用合成和手工制作的合成器从三个流行的软件合成器,并评估其性能的合成器的声音匹配下游任务。虽然学习的代表性的好处是细微差别的资源需求,令人鼓舞的结果,获得了所有的合成器,为未来的研究铺平了道路,为基于神经的ASP系统的合成器代理的应用。
摘要:Deep learning appears as an appealing solution for Automatic Synthesizer Programming (ASP), which aims to assist musicians and sound designers in programming sound synthesizers. However, integrating software synthesizers into training pipelines is challenging due to their potential non-differentiability. This work tackles this challenge by introducing a method to approximate arbitrary synthesizers. Specifically, we train a neural network to map synthesizer presets onto an audio embedding space derived from a pretrained model. This facilitates the definition of a neural proxy that produces compact yet effective representations, thereby enabling the integration of audio embedding loss into neural-based ASP systems for black-box synthesizers. We evaluate the representations derived by various pretrained audio models in the context of neural-based nASP and assess the effectiveness of several neural network architectures, including feedforward, recurrent, and transformer-based models, in defining neural proxies. We evaluate the proposed method using both synthetic and hand-crafted presets from three popular software synthesizers and assess its performance in a synthesizer sound matching downstream task. While the benefits of the learned representation are nuanced by resource requirements, encouraging results were obtained for all synthesizers, paving the way for future research into the application of synthesizer proxies for neural-based ASP systems.


【4】Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
标题:在公共数据上进行数据有效单阶段训练的竞争性音频语言模型
链接:https://arxiv.org/abs/2509.07526

作者:Gokul Karthik Kumar, Rishabh Saraf, Ludovick Lepauloux, Abdul Muneer, Billel Mokeddem, Hakim Hacid
备注:Accepted at ASRU 2025
摘要:大型语言模型(LLM)已经改变了NLP,但它们与音频的集成仍然没有得到充分的探索-尽管音频是人类交流的中心。我们介绍Falcon 3-Audio,这是一个基于音频调优LLM和Whisper编码器构建的音频语言模型(ALM)家族。使用非常少量的公共音频数据-不到30 K小时(5 K唯一)-Falcon 3-Audio-7 B在MMAU基准测试中与开放重量模型中报告的最佳性能相匹配,得分为64.14,与R1-AQA相匹配,同时通过卓越的数据和参数效率,单阶段训练和透明度而脱颖而出。值得注意的是,我们最小的1B模型与从2B到13 B参数的较大开放模型相比仍然具有竞争力。通过广泛的消融,我们发现常见的复杂性-例如课程学习,多个音频编码器和复杂的交叉注意连接器-即使与超过50万小时的数据训练模型相比,也不需要强大的性能。
摘要:Large language models (LLMs) have transformed NLP, yet their integration with audio remains underexplored -- despite audio's centrality to human communication. We introduce Falcon3-Audio, a family of Audio-Language Models (ALMs) built on instruction-tuned LLMs and Whisper encoders. Using a remarkably small amount of public audio data -- less than 30K hours (5K unique) -- Falcon3-Audio-7B matches the best reported performance among open-weight models on the MMAU benchmark, with a score of 64.14, matching R1-AQA, while distinguishing itself through superior data and parameter efficiency, single-stage training, and transparency. Notably, our smallest 1B model remains competitive with larger open models ranging from 2B to 13B parameters. Through extensive ablations, we find that common complexities -- such as curriculum learning, multiple audio encoders, and intricate cross-attention connectors -- are not required for strong performance, even compared to models trained on over 500K hours of data.


【5】Target matching based generative model for speech enhancement
标题:基于目标匹配的语音增强生成模型
链接:https://arxiv.org/abs/2509.07521

作者:Taihui Wang, Rilin Chen, Tong Lei, Andong Li, Jinzheng Zhao, Meng Yu, Dong Yu
备注:12 pages, 5 figures
摘要:扰动信号的均值和方差时间表的设计是生成模型中的一个基本挑战。基于分数和基于Schr“odinger桥的模型需要仔细选择随机微分方程来推导相应的时间表,而基于流的模型通过向量场匹配来解决这个问题。然而,由于向量场中可能包含随机分量,这种策略通常会导致幻觉伪影和低效的训练和推理过程。此外,广泛采用的扩散骨干,NCSN++,遭受高计算复杂度。为了克服这些限制,我们提出了一种新的基于目标的生成框架,提高了均值/方差时间表设计的灵活性和训练和推理过程的效率。具体来说,我们通过将生成式语音增强任务重新定义为目标信号估计问题来消除训练损失中的随机成分,从而导致更稳定和有效的训练和推理过程。此外,我们采用了逻辑平均时间表和桥方差时间表,产生更有利的信噪比轨迹相比,几个广泛使用的时间表,从而导致更有效的扰动策略。此外,我们提出了一种新的音频扩散骨干,通过显式建模长期帧相关性和跨频带依赖性,显着提高了NCSN++的效率。
摘要:The design of mean and variance schedules for the perturbed signal is a fundamental challenge in generative models. While score-based and Schr\"odinger bridge-based models require careful selection of the stochastic differential equation to derive the corresponding schedules, flow-based models address this issue via vector field matching. However, this strategy often leads to hallucination artifacts and inefficient training and inference processes due to the potential inclusion of stochastic components in the vector field. Additionally, the widely adopted diffusion backbone, NCSN++, suffers from high computational complexity. To overcome these limitations, we propose a novel target-based generative framework that enhances both the flexibility of mean/variance schedule design and the efficiency of training and inference processes. Specifically, we eliminate the stochastic components in the training loss by reformulating the generative speech enhancement task as a target signal estimation problem, which therefore leads to more stable and efficient training and inference processes. In addition, we employ a logistic mean schedule and a bridge variance schedule, which yield a more favorable signal-to-noise ratio trajectory compared to several widely used schedules and thus leads to a more efficient perturbation strategy. Furthermore, we propose a new diffusion backbone for audio, which significantly improves the efficiency over NCSN++ by explicitly modeling long-term frame correlations and cross-band dependencies.


【6】Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis
标题:基于双边属性增强的渐进面部粒度聚合用于面部到语音合成
链接:https://arxiv.org/abs/2509.07376

作者:Yejin Jeon, Youngjae Kim, Jihyun Lee, Hyounghun Kim, Gary Geunbae Lee
备注:EMNLP Findings
摘要:对于经历过中风等创伤事件的人来说,语言可能不再是一种可行的沟通方式。虽然文本到语音(TTS)可以用作通信辅助,因为它生成合成语音,但它无法保留用户自己的声音。因此,从面部图像中导出相应语音的面部到语音(FTV)合成提供了一种有前途的替代方案。然而,现有的方法依赖于预先训练的视觉编码器,并对其进行微调以与语音嵌入保持一致,这从性别或种族等面部输入中剥离了细粒度信息,尽管它们与声音特征存在已知的相关性。此外,这些管道是多级的,这需要单独训练多个组件,从而导致训练效率低下。为了解决这些局限性,我们利用细粒度的面部属性建模,将面部图像分解成不重叠的片段,并逐步将它们集成到一个多粒度的表示。这种表示是进一步完善,通过多任务学习的扬声器属性,如性别和种族在视觉和声学领域。此外,为了提高对齐的鲁棒性,我们采用了多视图训练策略,通过配对不同角度和照明条件下的扬声器的各种视觉视角,与相同的语音记录。广泛的主观和客观的评价证实,我们的方法大大提高了脸的声音一致性和合成稳定性。
摘要:For individuals who have experienced traumatic events such as strokes, speech may no longer be a viable means of communication. While text-to-speech (TTS) can be used as a communication aid since it generates synthetic speech, it fails to preserve the user's own voice. As such, face-to-voice (FTV) synthesis, which derives corresponding voices from facial images, provides a promising alternative. However, existing methods rely on pre-trained visual encoders, and finetune them to align with speech embeddings, which strips fine-grained information from facial inputs such as gender or ethnicity, despite their known correlation with vocal traits. Moreover, these pipelines are multi-stage, which requires separate training of multiple components, thus leading to training inefficiency. To address these limitations, we utilize fine-grained facial attribute modeling by decomposing facial images into non-overlapping segments and progressively integrating them into a multi-granular representation. This representation is further refined through multi-task learning of speaker attributes such as gender and ethnicity at both the visual and acoustic domains. Moreover, to improve alignment robustness, we adopt a multi-view training strategy by pairing various visual perspectives of a speaker in terms of different angles and lighting conditions, with identical speech recordings. Extensive subjective and objective evaluations confirm that our approach substantially enhances face-voice congruence and synthesis stability.


【7】When Fine-Tuning is Not Enough: Lessons from HSAD on Hybrid and Adversarial Audio Spoof Detection
标题:当微调还不够时:HSAD关于混合和对抗性音频欺骗检测的教训
链接:https://arxiv.org/abs/2509.07323

作者:Bin Hu, Kunyang Huang, Daehan Kwak, Meng Xu, Kuan Huang
备注:13 pages, 11 this http URL work has been submitted to the IEEE for possible publication
摘要:人工智能的快速发展使高度逼真的语音合成和语音克隆成为可能,对语音认证、智能助手和电信安全构成严重风险。虽然大多数先前的工作框架将欺骗检测作为二元任务,但现实世界的攻击通常涉及混合真实语音和合成语音的混合话语,使得检测更具挑战性。为了解决这一差距,我们引入了混合欺骗音频数据集(HSAD),这是一个基准测试,包含1,248个干净的和41,044个降级的话语,分为四类:人类,克隆,zero-shot AI生成和混合音频。每个样本都用欺骗方法、说话者身份和降级元数据进行注释,以实现细粒度分析。我们评估了六个基于变换器的模型,包括频谱图编码器(MIT-AST,MattyB 95-AST)和自监督波形模型(Wav 2 Vec 2,HuBERT)。结果揭示了关键的教训:预训练的模型在混合条件下过度泛化和崩溃;欺骗特定的微调提高了可分性,但与看不见的成分斗争; HSAD上的特定于HSAD的自适应产生了很大的性能增益(AST大于97%,F1得分约为99%),尽管复杂混合的残余错误仍然存在。这些研究结果表明,仅进行微调并不可靠-像HSAD这样的强大的混合感知基准对于暴露校准失败,模型偏差以及对抗环境中影响欺骗检测的因素至关重要。因此,HSAD提供了一个数据集和一个分析框架,用于构建弹性和可信的语音认证系统。
摘要:The rapid advancement of AI has enabled highly realistic speech synthesis and voice cloning, posing serious risks to voice authentication, smart assistants, and telecom security. While most prior work frames spoof detection as a binary task, real-world attacks often involve hybrid utterances that mix genuine and synthetic speech, making detection substantially more challenging. To address this gap, we introduce the Hybrid Spoofed Audio Dataset (HSAD), a benchmark containing 1,248 clean and 41,044 degraded utterances across four classes: human, cloned, zero-shot AI-generated, and hybrid audio. Each sample is annotated with spoofing method, speaker identity, and degradation metadata to enable fine-grained analysis. We evaluate six transformer-based models, including spectrogram encoders (MIT-AST, MattyB95-AST) and self-supervised waveform models (Wav2Vec2, HuBERT). Results reveal critical lessons: pretrained models overgeneralize and collapse under hybrid conditions; spoof-specific fine-tuning improves separability but struggles with unseen compositions; and dataset-specific adaptation on HSAD yields large performance gains (AST greater than 97 percent and F1 score is approximately 99 percent), though residual errors persist for complex hybrids. These findings demonstrate that fine-tuning alone is not sufficient-robust hybrid-aware benchmarks like HSAD are essential to expose calibration failures, model biases, and factors affecting spoof detection in adversarial environments. HSAD thus provides both a dataset and an analytic framework for building resilient and trustworthy voice authentication systems.


【8】Adversarial Attacks on Audio Deepfake Detection: A Benchmark and Comparative Study
标题:音频Deepfake检测的对抗攻击:基准和比较研究
链接:https://arxiv.org/abs/2509.07132

作者:Kutub Uddin, Muhammad Umar Farooq, Awais Khan, Khalid Mahmood Malik
摘要:生成式人工智能的广泛使用在制作高度逼真的deepfake方面取得了显着的成功,对各种语音生物识别应用构成了严重威胁,包括说话人验证,语音生物识别,音频会议和刑事调查。为了解决这个问题,已经提出了几种最先进的(SoTA)音频深度伪造检测(ADD)方法来识别生成AI签名,以区分真实和深度伪造音频。然而,这些方法的有效性被隐藏生成签名的反取证(AF)攻击严重破坏。这些AF攻击涵盖了广泛的技术,包括统计修改(例如,音调移位、滤波、噪声添加和量化)和基于优化的攻击(例如,FGSM、PGD、C \& W和DeepFool)。在本文中,我们研究了SoTA ADD方法,并提供了一个比较分析,以突出它们在暴露deepfake签名方面的有效性,以及它们在对抗条件下的漏洞。我们使用两类方法在五个deepfake基准数据集上对ADD方法进行了广泛的评估:原始方法和基于频谱图的方法。这种比较分析使人们能够更深入地了解SoTA ADD方法对抗各种AF攻击的优势和局限性。它不仅突出了ADD方法的弱点,而且还为现实世界的语音生物特征识别设计更强大和更通用的检测器提供了信息。它将进一步指导未来的研究,开发适应性防御策略,可以有效地对抗不断发展的AF技术。
摘要:The widespread use of generative AI has shown remarkable success in producing highly realistic deepfakes, posing a serious threat to various voice biometric applications, including speaker verification, voice biometrics, audio conferencing, and criminal investigations. To counteract this, several state-of-the-art (SoTA) audio deepfake detection (ADD) methods have been proposed to identify generative AI signatures to distinguish between real and deepfake audio. However, the effectiveness of these methods is severely undermined by anti-forensic (AF) attacks that conceal generative signatures. These AF attacks span a wide range of techniques, including statistical modifications (e.g., pitch shifting, filtering, noise addition, and quantization) and optimization-based attacks (e.g., FGSM, PGD, C \& W, and DeepFool). In this paper, we investigate the SoTA ADD methods and provide a comparative analysis to highlight their effectiveness in exposing deepfake signatures, as well as their vulnerabilities under adversarial conditions. We conducted an extensive evaluation of ADD methods on five deepfake benchmark datasets using two categories: raw and spectrogram-based approaches. This comparative analysis enables a deeper understanding of the strengths and limitations of SoTA ADD methods against diverse AF attacks. It does not only highlight vulnerabilities of ADD methods, but also informs the design of more robust and generalized detectors for real-world voice biometrics. It will further guide future research in developing adaptive defense strategies that can effectively counter evolving AF techniques.


【9】End-to-End Efficiency in Keyword Spotting: A System-Level Approach for Embedded Microcontrollers
标题:关键词发现的端到端效率:嵌入式微控制器的系统级方法
链接:https://arxiv.org/abs/2509.07051

作者:Pietro Bartoli, Tommaso Bondini, Christian Veronesi, Andrea Giudici, Niccolò Antonello, Franco Zappa
备注:4 pages, 2 figures, 1 table. Accepted for publication in IEEE Sensors 2025. \c{opyright} 2025 IEEE. Personal use permitted. Permission from IEEE required for all other uses
摘要:关键字识别(KWS)是嵌入式和物联网设备中免提交互的关键支持技术,其中严格的内存和能源限制对支持AI的设备的部署构成了挑战。在这项工作中,我们系统地评估和比较了几种最先进的轻量级神经网络架构,包括DS-CNN,LiCoNet和TENet,以及我们提出的基于MobileNet的Typman-KWS(TKWS)架构,专为微控制器单元(MCU)上的高效KWS而设计。与之前仅关注模型推理的研究不同,我们的分析涵盖了从梅尔频率倒谱系数(MFCC)特征提取到神经推理的整个处理流程,并在三个STM 32平台(N6,H7和U 5)上进行了基准测试。我们的研究结果表明,TKWS与三个残留块实现高达92.4%的F1分数,只有14.4k的参数,减少内存占用,而不影响准确性。此外,具有集成神经加速功能的N6 MCU实现了最佳的能量延迟积(EDP),即使在高分辨率特性下也能实现高效、低延迟的操作。我们的研究结果强调了模型的准确性本身并不能决定现实世界的有效性;相反,最佳的关键字定位部署需要仔细考虑特征提取参数和特定于硬件的优化。
摘要:Keyword spotting (KWS) is a key enabling technology for hands-free interaction in embedded and IoT devices, where stringent memory and energy constraints challenge the deployment of AI-enabeld devices. In this work, we systematically evaluate and compare several state-of-the-art lightweight neural network architectures, including DS-CNN, LiCoNet, and TENet, alongside our proposed Typman-KWS (TKWS) architecture built upon MobileNet, specifically designed for efficient KWS on microcontroller units (MCUs). Unlike prior studies focused solely on model inference, our analysis encompasses the entire processing pipeline, from Mel-Frequency Cepstral Coefficient (MFCC) feature extraction to neural inference, and is benchmarked across three STM32 platforms (N6, H7, and U5). Our results show that TKWS with three residual blocks achieves up to 92.4% F1-score with only 14.4k parameters, reducing memory footprint without compromising the accuracy. Moreover, the N6 MCU with integrated neural acceleration achieves the best energy-delay product (EDP), enabling efficient, low-latency operation even with high-resolution features. Our findings highlight the model accuracy alone does not determine real-world effectiveness; rather, optimal keyword spotting deployments require careful consideration of feature extraction parameters and hardware-specific optimization.


【10】Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence
标题:利用音素级能量序列的可控歌唱声音合成
链接:https://arxiv.org/abs/2509.07038

作者:Yerin Ryu, Inseop Shin, Chanwoo Kim
备注:Accepted to ASRU 2025
摘要:可控歌声合成(SVS)旨在生成反映用户意图的富有表现力的歌声。虽然最近的SVS系统实现了高音频质量,但大多数依赖于概率建模,限制了对动态等属性的精确控制。我们通过专注于动态控制来解决这个问题-时间响度变化对音乐表现力至关重要-并明确地将SVS模型置于从地面实况频谱图中提取的能量序列上,从而降低注释成本并提高可控性。我们还提出了一个音素级的能量序列,用户友好的控制。据我们所知,这是第一次尝试在SVS中启用用户驱动的动态控制。实验表明,我们的方法实现了超过50%的减少,平均绝对误差的音素级输入的能量序列相比,基线和能量预测模型,而不影响合成质量。
摘要:Controllable Singing Voice Synthesis (SVS) aims to generate expressive singing voices reflecting user intent. While recent SVS systems achieve high audio quality, most rely on probabilistic modeling, limiting precise control over attributes such as dynamics. We address this by focusing on dynamic control--temporal loudness variation essential for musical expressiveness--and explicitly condition the SVS model on energy sequences extracted from ground-truth spectrograms, reducing annotation costs and improving controllability. We also propose a phoneme-level energy sequence for user-friendly control. To the best of our knowledge, this is the first attempt enabling user-driven dynamics control in SVS. Experiments show our method achieves over 50% reduction in mean absolute error of energy sequences for phoneme-level inputs compared to baseline and energy-predictor models, without compromising synthesis quality.


【11】Prototype: A Keyword Spotting-Based Intelligent Audio SoC for IoT
标题:原型:基于关键字点识别的物联网智能音频SOC
链接:https://arxiv.org/abs/2509.06964

作者:Huihong Liang, Dongxuan Jia, Youquan Wang, Longtao Huang, Shida Zhong, Luping Xiang, Lei Huang, Tao Yuan
摘要:在此演示中,我们展示了一款集成关键词识别加速器的紧凑型智能音频片上系统(SoC),可在物联网(IoT)设备中实现超低延迟、低功耗和低成本的语音交互。通过算法-硬件协同设计,系统的能量效率最大化。我们通过基于FPGA的实时原型展示了该系统的功能,展示了边缘智能应用的稳定性能和实时语音交互。
摘要:In this demo, we present a compact intelligent audio system-on-chip (SoC) integrated with a keyword spotting accelerator, enabling ultra-low latency, low-power, and low-cost voice interaction in Internet of Things (IoT) devices. Through algorithm-hardware co-design, the system's energy efficiency is maximized. We demonstrate the system's capabilities through a live FPGA-based prototype, showcasing stable performance and real-time voice interaction for edge intelligence applications.


【12】Exploring System Adaptations For Minimum Latency Real-Time Piano Transcription
标题:探索最小延迟实时钢琴抄写的系统适应
链接:https://arxiv.org/abs/2509.07586

作者:Patricia Hu, Silvan David Peter, Jan Schlüter, Gerhard Widmer
备注:to be published in Proceedings of the 26th International Society for Music Information Retrieval (ISMIR) Conference 2025, Daejeon, South Korea
摘要:神经网络设计的进步和大规模标记数据集的可用性推动了钢琴转录的重大改进。现有的方法针对离线应用程序,没有限制的计算需求,或在线转录,与延迟128-320毫秒。然而,大多数实时音乐应用程序需要延迟低于30毫秒。在这项工作中,我们调查是否以及如何目前最先进的在线转录模型可以适应实时钢琴转录。具体来说,我们消除了所有非因果处理,并通过跨核心模型组件和模型大小变化的共享计算来减少计算负载。此外,我们还探索了不同的预处理和后处理策略以及相关的标签编码方案,并讨论了它们对实时转录的适用性。评估MAESTRO数据集上的自适应,我们发现由于严格的因果处理以及预处理延迟和预测准确性之间的权衡,转录准确性下降。我们发布我们的系统作为基线,以支持研究人员设计最小延迟实时转录的模型。
摘要:Advances in neural network design and the availability of large-scale labeled datasets have driven major improvements in piano transcription. Existing approaches target either offline applications, with no restrictions on computational demands, or online transcription, with delays of 128-320 ms. However, most real-time musical applications require latencies below 30 ms. In this work, we investigate whether and how the current state-of-the-art online transcription model can be adapted for real-time piano transcription. Specifically, we eliminate all non-causal processing, and reduce computational load through shared computations across core model components and variations in model size. Additionally, we explore different pre- and postprocessing strategies, and related label encoding schemes, and discuss their suitability for real-time transcription. Evaluating the adaptions on the MAESTRO dataset, we find a drop in transcription accuracy due to strictly causal processing as well as a tradeoff between the preprocessing latency and prediction accuracy. We release our system as a baseline to support researchers in designing models towards minimum latency real-time transcription.


eess.AS音频处理


【1】Exploring System Adaptations For Minimum Latency Real-Time Piano Transcription
标题:探索最小延迟实时钢琴抄写的系统适应
链接:https://arxiv.org/abs/2509.07586

作者:Patricia Hu, Silvan David Peter, Jan Schlüter, Gerhard Widmer
备注:to be published in Proceedings of the 26th International Society for Music Information Retrieval (ISMIR) Conference 2025, Daejeon, South Korea
摘要:神经网络设计的进步和大规模标记数据集的可用性推动了钢琴转录的重大改进。现有的方法针对离线应用程序,没有限制的计算需求,或在线转录,与延迟128-320毫秒。然而,大多数实时音乐应用程序需要延迟低于30毫秒。在这项工作中,我们调查是否以及如何目前最先进的在线转录模型可以适应实时钢琴转录。具体来说,我们消除了所有非因果处理,并通过跨核心模型组件和模型大小变化的共享计算来减少计算负载。此外,我们探讨不同的预处理和后处理策略,以及相关的标签编码方案,并讨论其适用于实时转录。评估MAESTRO数据集上的自适应,我们发现由于严格的因果处理以及预处理延迟和预测准确性之间的权衡,转录准确性下降。我们发布我们的系统作为基线,以支持研究人员设计最小延迟实时转录的模型。
摘要:Advances in neural network design and the availability of large-scale labeled datasets have driven major improvements in piano transcription. Existing approaches target either offline applications, with no restrictions on computational demands, or online transcription, with delays of 128-320 ms. However, most real-time musical applications require latencies below 30 ms. In this work, we investigate whether and how the current state-of-the-art online transcription model can be adapted for real-time piano transcription. Specifically, we eliminate all non-causal processing, and reduce computational load through shared computations across core model components and variations in model size. Additionally, we explore different pre- and postprocessing strategies, and related label encoding schemes, and discuss their suitability for real-time transcription. Evaluating the adaptions on the MAESTRO dataset, we find a drop in transcription accuracy due to strictly causal processing as well as a tradeoff between the preprocessing latency and prediction accuracy. We release our system as a baseline to support researchers in designing models towards minimum latency real-time transcription.


【2】Affine Modulation-based Audiogram Fusion Network for Joint Noise Reduction and Hearing Loss Compensation
标题:基于仿射调制的联合降噪和听力损失补偿的声图融合网络
链接:https://arxiv.org/abs/2509.07341

作者:Ye Ni, Ruiyu Liang, Xiaoshuai Hao, Jiaming Cheng, Qingyun Wang, Chengwei Huang, Cairong Zou, Wei Zhou, Weiping Ding, Björn W. Schuller
摘要:助听器(HA)被广泛用于提供个性化的语音增强(PSE)服务,从而改善听力损失患者的生活质量。然而,HA性能在嘈杂的环境中显著下降,因为它将降噪(NR)和听力损失补偿(HLC)视为单独的任务。这种分离导致缺乏系统优化,忽略了这两个关键任务之间的相互作用,并增加了系统的复杂性。为了解决这些挑战,我们提出了一种新的听力图融合网络,命名为AFN-HearNet,它通过融合跨域听力图和频谱特征同时处理NR和HLC任务。我们提出了一个音频特定的编码器,将稀疏的听力曲线转换为一个深的表示,解决融合前的跨域特征的对齐问题。为了结合NR和HLC任务之间的相互作用,我们提出了基于仿射调制的听力图融合频率-时间一致性,自适应地将这两个特征融合到一个统一的深度表示中,用于语音重建。此外,我们引入了语音活动检测辅助训练任务,将语音和非语音模式隐式地嵌入到统一的深度表示中。我们在多个数据集上进行了全面的实验,以验证每个模块的有效性。结果表明,AFN-HearNet在关键指标(如HASQI和PESQ)方面明显优于最先进的上下文融合联合模型,在性能和效率之间实现了相当大的权衡。源代码和数据将在https://github.com/deepnetni/AFN-HearNet上发布。
摘要:Hearing aids (HAs) are widely used to provide personalized speech enhancement (PSE) services, improving the quality of life for individuals with hearing loss. However, HA performance significantly declines in noisy environments as it treats noise reduction (NR) and hearing loss compensation (HLC) as separate tasks. This separation leads to a lack of systematic optimization, overlooking the interactions between these two critical tasks, and increases the system complexity. To address these challenges, we propose a novel audiogram fusion network, named AFN-HearNet, which simultaneously tackles the NR and HLC tasks by fusing cross-domain audiogram and spectrum features. We propose an audiogram-specific encoder that transforms the sparse audiogram profile into a deep representation, addressing the alignment problem of cross-domain features prior to fusion. To incorporate the interactions between NR and HLC tasks, we propose the affine modulation-based audiogram fusion frequency-temporal Conformer that adaptively fuses these two features into a unified deep representation for speech reconstruction. Furthermore, we introduce a voice activity detection auxiliary training task to embed speech and non-speech patterns into the unified deep representation implicitly. We conduct comprehensive experiments across multiple datasets to validate the effectiveness of each proposed module. The results indicate that the AFN-HearNet significantly outperforms state-of-the-art in-context fusion joint models regarding key metrics such as HASQI and PESQ, achieving a considerable trade-off between performance and efficiency. The source code and data will be released at https://github.com/deepnetni/AFN-HearNet.


【3】Identifying and Calibrating Overconfidence in Noisy Speech Recognition
标题:识别和校准有噪语音识别中的过度自信
链接:https://arxiv.org/abs/2509.07195

作者:Mingyue Huo, Yuheng Zhang, Yan Tang
备注:Accepted to ASRU2025
摘要:像Whisper这样的现代端到端自动语音识别(ASR)模型不仅在噪声中识别准确率降低,而且表现出过度自信-将高置信度分配给错误的预测。我们对Whisper在加性噪声条件下的行为进行了系统的分析,发现在低信噪比下,过度自信的错误会急剧增加,10-20%的令牌在置信度高于0.7的情况下被错误地预测。为了缓解这一问题,我们提出了一个轻量级的事后校准框架,该框架可以检测潜在的过度自信,并选择性地对这些令牌应用温度缩放,而不会改变底层的ASR模型。R-SPIN数据集上的评估表明,在低信噪比范围(-18至-5 dB),我们的方法将预期校准误差(ECE)降低了58%,并将归一化交叉熵(NCE)提高了两倍,在严重的噪声条件下产生更可靠的置信度估计。
摘要:Modern end-to-end automatic speech recognition (ASR) models like Whisper not only suffer from reduced recognition accuracy in noise, but also exhibit overconfidence - assigning high confidence to wrong predictions. We conduct a systematic analysis of Whisper's behavior in additive noise conditions and find that overconfident errors increase dramatically at low signal-to-noise ratios, with 10-20% of tokens incorrectly predicted with confidence above 0.7. To mitigate this, we propose a lightweight, post-hoc calibration framework that detects potential overconfidence and applies temperature scaling selectively to those tokens, without altering the underlying ASR model. Evaluations on the R-SPIN dataset demonstrate that, in the low signal-to-noise ratio range (-18 to -5 dB), our method reduces the expected calibration error (ECE) by 58% and triples the normalized cross entropy (NCE), yielding more reliable confidence estimates under severe noise conditions.


【4】Spectral and Rhythm Feature Performance Evaluation for Category and Class Level Audio Classification with Deep Convolutional Neural Networks
标题:使用深度卷积神经网络进行类别和类别级音频分类的频谱和节奏特征性能评估
链接:https://arxiv.org/abs/2509.07756

作者: Friedrich Wolf-Monheim
摘要:除了决策树和k近邻算法之外,深度卷积神经网络(CNN)被广泛用于对音乐、语音或环境声音等许多领域的音频数据进行分类。为了训练特定的CNN,可以将各种频谱和节奏特征(如梅尔缩放频谱图、梅尔频率倒谱系数(MFCC)、循环温度图、短时傅里叶变换(STFT)色度图、恒定Q变换(CQT)色度图和色度能量归一化统计(CENS)色度图)用作神经网络的数字图像输入数据。使用深度CNN和ESC-50数据集,使用端到端深度学习管道,详细研究了这些频谱和节奏特征在音频类别级别以及音频类级别分类中的性能,其中ESC-50数据集包含2,000个标记的环境音频记录。多类分类的评估指标准确度、精确度、召回率和F1得分清楚地表明,梅尔缩放频谱图和梅尔频率倒谱系数(MFCC)的表现明显优于本研究中使用深度CNN进行音频分类任务所研究的其他频谱和节奏特征。
摘要:Next to decision tree and k-nearest neighbours algorithms deep convolutional neural networks (CNNs) are widely used to classify audio data in many domains like music, speech or environmental sounds. To train a specific CNN various spectral and rhythm features like mel-scaled spectrograms, mel-frequency cepstral coefficients (MFCC), cyclic tempograms, short-time Fourier transform (STFT) chromagrams, constant-Q transform (CQT) chromagrams and chroma energy normalized statistics (CENS) chromagrams can be used as digital image input data for the neural network. The performance of these spectral and rhythm features for audio category level as well as audio class level classification is investigated in detail with a deep CNN and the ESC-50 dataset with 2,000 labeled environmental audio recordings using an end-to-end deep learning pipeline. The evaluated metrics accuracy, precision, recall and F1 score for multiclass classification clearly show that the mel-scaled spectrograms and the mel-frequency cepstral coefficients (MFCC) perform significantly better then the other spectral and rhythm features investigated in this research for audio classification tasks using deep CNNs.


【5】Neural Proxies for Sound Synthesizers: Learning Perceptually Informed Preset Representations
标题:声音合成器的神经代理:学习感知知情的预设表示
链接:https://arxiv.org/abs/2509.07635

作者:Paolo Combes, Stefan Weinzierl, Klaus Obermayer
备注:17 pages, 4 figures, published in the Journal of the Audio   Engineering Society
摘要:深度学习似乎是自动合成器编程(ASP)的一个有吸引力的解决方案,旨在帮助音乐家和声音设计师对声音合成器进行编程。然而,由于其潜在的不可微性,将软件合成器集成到训练管道中具有挑战性。这项工作通过引入一种方法来近似任意合成器来应对这一挑战。具体来说,我们训练一个神经网络,以将合成器映射到从预训练模型导出的音频嵌入空间。这有助于定义产生紧凑而有效的表示的神经代理,从而能够将音频嵌入损失集成到黑盒合成器的基于神经的ASP系统中。我们在基于神经的nASP的背景下评估了各种预训练音频模型的表示,并评估了几种神经网络架构的有效性,包括前馈,递归和基于transformer的模型,在定义神经代理。我们评估所提出的方法使用合成和手工制作的合成器从三个流行的软件合成器,并评估其性能的合成器的声音匹配下游任务。虽然学习的代表性的好处是细微差别的资源需求,令人鼓舞的结果,获得了所有的合成器,为未来的研究铺平了道路,为基于神经的ASP系统的合成器代理的应用。
摘要:Deep learning appears as an appealing solution for Automatic Synthesizer Programming (ASP), which aims to assist musicians and sound designers in programming sound synthesizers. However, integrating software synthesizers into training pipelines is challenging due to their potential non-differentiability. This work tackles this challenge by introducing a method to approximate arbitrary synthesizers. Specifically, we train a neural network to map synthesizer presets onto an audio embedding space derived from a pretrained model. This facilitates the definition of a neural proxy that produces compact yet effective representations, thereby enabling the integration of audio embedding loss into neural-based ASP systems for black-box synthesizers. We evaluate the representations derived by various pretrained audio models in the context of neural-based nASP and assess the effectiveness of several neural network architectures, including feedforward, recurrent, and transformer-based models, in defining neural proxies. We evaluate the proposed method using both synthetic and hand-crafted presets from three popular software synthesizers and assess its performance in a synthesizer sound matching downstream task. While the benefits of the learned representation are nuanced by resource requirements, encouraging results were obtained for all synthesizers, paving the way for future research into the application of synthesizer proxies for neural-based ASP systems.


【6】The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
标题:ML-SURB 2.0挑战:迈向所有语言变体的包容性ASB基准测试
链接:https://arxiv.org/abs/2509.07139

作者:William Chen, Chutong Meng, Jiatong Shi, Martijn Bartelds, Shih-Heng Wang, Hsiu-Hsuan Wang, Rafael Mosquera, Sara Hincapie, Dan Jurafsky, Antonis Anastasopoulos, Hung-yi Lee, Karen Livescu, Shinji Watanabe
备注:Interspeech 2025
摘要:多语言ASR的最新改进并没有在语言和语言种类之间均匀分布。为了推进最先进的(SOTA)ASR模型,我们推出了Interspeech 2025 ML-SUPERB 2.0挑战赛。我们构建了一个新的测试套件,其中包含来自200多种语言,口音和方言的数据,以评估SOTA多语言语音模型。挑战赛还引入了基于DynaBench的在线评估服务器,允许参与者灵活地进行模型设计和架构。挑战赛收到了来自3个团队的5份参赛作品,所有参赛作品都超过了我们的基线。与一般多语言测试集的最佳基线相比,表现最好的提交在LID准确性方面实现了23%的绝对提高,CER减少了18%。在口音和方言数据上,最好的提交获得了30.2%的CER和15.7%的LID准确性,这表明社区挑战在使语音技术更具包容性方面的重要性。
摘要:Recent improvements in multilingual ASR have not been equally distributed across languages and language varieties. To advance state-of-the-art (SOTA) ASR models, we present the Interspeech 2025 ML-SUPERB 2.0 Challenge. We construct a new test suite that consists of data from 200+ languages, accents, and dialects to evaluate SOTA multilingual speech models. The challenge also introduces an online evaluation server based on DynaBench, allowing for flexibility in model design and architecture for participants. The challenge received 5 submissions from 3 teams, all of which outperformed our baselines. The best-performing submission achieved an absolute improvement in LID accuracy of 23% and a reduction in CER of 18% when compared to the best baseline on a general multilingual test set. On accented and dialectal data, the best submission obtained 30.2% lower CER and 15.7% higher LID accuracy, showing the importance of community challenges in making speech technologies more inclusive.


【7】Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence
标题:利用音素级能量序列的可控歌唱声音合成
链接:https://arxiv.org/abs/2509.07038

作者:Yerin Ryu, Inseop Shin, Chanwoo Kim
备注:Accepted to ASRU 2025
摘要:可控歌声合成(SVS)旨在生成反映用户意图的富有表现力的歌声。虽然最近的SVS系统实现了高音频质量,但大多数依赖于概率建模,限制了对动态等属性的精确控制。我们通过专注于动态控制来解决这个问题-时间响度变化对音乐表现力至关重要-并明确地将SVS模型置于从地面实况频谱图中提取的能量序列上,从而降低注释成本并提高可控性。我们还提出了一个音素级的能量序列,用户友好的控制。据我们所知,这是第一次尝试在SVS中启用用户驱动的动态控制。实验表明,我们的方法实现了超过50%的减少,平均绝对误差的音素级输入的能量序列相比,基线和能量预测模型,而不影响合成质量。
摘要:Controllable Singing Voice Synthesis (SVS) aims to generate expressive singing voices reflecting user intent. While recent SVS systems achieve high audio quality, most rely on probabilistic modeling, limiting precise control over attributes such as dynamics. We address this by focusing on dynamic control--temporal loudness variation essential for musical expressiveness--and explicitly condition the SVS model on energy sequences extracted from ground-truth spectrograms, reducing annotation costs and improving controllability. We also propose a phoneme-level energy sequence for user-friendly control. To the best of our knowledge, this is the first attempt enabling user-driven dynamics control in SVS. Experiments show our method achieves over 50% reduction in mean absolute error of energy sequences for phoneme-level inputs compared to baseline and energy-predictor models, without compromising synthesis quality.


【8】Prototype: A Keyword Spotting-Based Intelligent Audio SoC for IoT
标题:原型:基于关键字点识别的物联网智能音频SOC
链接:https://arxiv.org/abs/2509.06964

作者:Huihong Liang, Dongxuan Jia, Youquan Wang, Longtao Huang, Shida Zhong, Luping Xiang, Lei Huang, Tao Yuan
摘要:在此演示中,我们展示了一款集成关键词识别加速器的紧凑型智能音频片上系统(SoC),可在物联网(IoT)设备中实现超低延迟、低功耗和低成本的语音交互。通过算法-硬件协同设计,系统的能量效率最大化。我们通过现场基于FPGA的原型展示了该系统的功能,展示了边缘智能应用的稳定性能和实时语音交互。
摘要:In this demo, we present a compact intelligent audio system-on-chip (SoC) integrated with a keyword spotting accelerator, enabling ultra-low latency, low-power, and low-cost voice interaction in Internet of Things (IoT) devices. Through algorithm-hardware co-design, the system's energy efficiency is maximized. We demonstrate the system's capabilities through a live FPGA-based prototype, showcasing stable performance and real-time voice interaction for edge intelligence applications.


机器翻译由腾讯交互翻译提供,仅供参考