今日论文合集:cs.SD语音13篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
标题:使用交织的语义、声学和文本令牌扩展开放离散音频基础模型
链接:https://arxiv.org/abs/2602.16687

作者:Potsawee Manakul,Woody Haosheng Gan,Martijn Bartelds,Guangzhi Sun,William Held,Diyi Yang
摘要:当前的音频语言模型主要是文本优先的,要么扩展预训练的文本LLM主干,要么依赖于仅语义的音频令牌,限制了一般的音频建模。本文提出了一个系统的实证研究的原生音频基础模型,应用下一个令牌预测音频的规模,联合建模语义内容,声学细节和文本,以支持通用音频生成和跨模态的能力。我们为构建这样的模型提供了全面的经验见解:(1)我们系统地研究了设计选择-数据源,文本混合比例和令牌组成-建立了一个经过验证的训练配方。(2)我们通过IsoFLOP分析对离散音频模型进行了第一次标度律研究,对64个模型进行了3 {\times}10^{18}$到3 {\times}10^{20}$FLOP分析,发现最佳数据的增长速度比最佳模型大小快1.6\times $。(3)我们将这些经验教训应用于训练SODA(Scaling Open Discrete Audio,扩展开放离散音频),这是一套在500 B令牌上从135 M到4 B参数的模型,与我们的扩展预测和现有模型进行比较。SODA是各种音频/文本任务的灵活骨干--我们通过使用相同的统一架构对保留语音的语音到语音翻译进行微调来证明这一点。
摘要:Current audio language models are predominantly text-first, either extending pre-trained text LLM backbones or relying on semantic-only audio tokens, limiting general audio modeling. This paper presents a systematic empirical study of native audio foundation models that apply next-token prediction to audio at scale, jointly modeling semantic content, acoustic details, and text to support both general audio generation and cross-modal capabilities. We provide comprehensive empirical insights for building such models: (1) We systematically investigate design choices -- data sources, text mixture ratios, and token composition -- establishing a validated training recipe. (2) We conduct the first scaling law study for discrete audio models via IsoFLOP analysis on 64 models spanning $3{\times}10^{18}$ to $3{\times}10^{20}$ FLOPs, finding that optimal data grows 1.6$\times$ faster than optimal model size. (3) We apply these lessons to train SODA (Scaling Open Discrete Audio), a suite of models from 135M to 4B parameters on 500B tokens, comparing against our scaling predictions and existing models. SODA serves as a flexible backbone for diverse audio/text tasks -- we demonstrate this by fine-tuning for voice-preserving speech-to-speech translation, using the same unified architecture.


【2】Hardware-accelerated graph neural networks: an alternative approach for neuromorphic event-based audio classification and keyword spotting on SoC FPGA
标题:硬件加速图神经网络:在SoCFPG上进行基于神经形态事件的音频分类和关键词识别的替代方法
链接:https://arxiv.org/abs/2602.16442

作者:Kamil Jeziorek,Piotr Wzorek,Krzysztof Blachut,Hiroshi Nakano,Manon Dampfhoffer,Thomas Mesquida,Hiroaki Nishi,Thomas Dalgaty,Tomasz Kryjak
备注:Under revision in TRETS Journal
摘要:随着嵌入式边缘传感器记录的数据量的增加,特别是来自产生离散事件流的神经形态设备的数据量的增加,对硬件感知神经架构的需求越来越大,这些架构能够实现高效,低延迟和节能的本地处理。我们提出了一个FPGA实现的事件图神经网络的音频处理。我们利用人工耳蜗将时间序列信号转换为稀疏事件数据,从而降低内存和计算成本。我们的架构在SoC FPGA上实现,并在两个开源数据集上进行了评估。对于分类任务,我们的基线浮点模型在SHD数据集上实现了92.7%的准确率-仅比最先进水平低2.4%-同时需要的参数减少了10倍和67倍以上。在SSC上,我们的模型实现了66.9-71.0%的准确度。与基于FPGA的脉冲神经网络相比,我们的量化模型达到了92.3%的准确率,比它们高出19.3%,同时减少了资源使用和延迟。对于SSC,我们报告了第一个硬件加速评估。我们进一步展示了事件音频关键字定位的第一个端到端FPGA实现,将图卷积层与递归序列建模相结合。该系统实现了高达95%的字端检测准确率,仅10.53微秒的延迟和1.18 W的功耗,为高能效的事件驱动KWS建立了强大的基准。
摘要:As the volume of data recorded by embedded edge sensors increases, particularly from neuromorphic devices producing discrete event streams, there is a growing need for hardware-aware neural architectures that enable efficient, low-latency, and energy-conscious local processing. We present an FPGA implementation of event-graph neural networks for audio processing. We utilise an artificial cochlea that converts time-series signals into sparse event data, reducing memory and computation costs. Our architecture was implemented on a SoC FPGA and evaluated on two open-source datasets. For classification task, our baseline floating-point model achieves 92.7% accuracy on SHD dataset - only 2.4% below the state of the art - while requiring over 10x and 67x fewer parameters. On SSC, our models achieve 66.9-71.0% accuracy. Compared to FPGA-based spiking neural networks, our quantised model reaches 92.3% accuracy, outperforming them by up to 19.3% while reducing resource usage and latency. For SSC, we report the first hardware-accelerated evaluation. We further demonstrate the first end-to-end FPGA implementation of event-audio keyword spotting, combining graph convolutional layers with recurrent sequence modelling. The system achieves up to 95% word-end detection accuracy, with only 10.53 microsecond latency and 1.18 W power consumption, establishing a strong benchmark for energy-efficient event-driven KWS.


【3】How to Label Resynthesized Audio: The Dual Role of Neural Audio Codecs in Audio Deepfake Detection
标题:如何标记重新合成的音频:神经音频编解码器在音频Deepfake检测中的双重角色
链接:https://arxiv.org/abs/2602.16343

作者:Yixuan Xiao,Florian Lux,Alejandro Pérez-González-de-Martos,Ngoc Thang Vu
备注:Accepted to ICASSP 2026
摘要:由于文本到语音系统通常不直接产生波形,最近的欺骗检测研究使用来自声码器和神经音频编解码器的重新合成波形来模拟攻击者。与专为语音合成而设计的声码器不同,神经音频编解码器最初是为了压缩音频以进行存储和传输而开发的。然而,他们离散语音的能力也引发了人们对基于语言建模的语音合成的兴趣。由于这种双重功能,编解码器重新合成的数据可以被标记为真实或欺骗。到目前为止,很少有研究涉及这个问题。在这项研究中,我们提出了一个具有挑战性的扩展ASVspoof 5数据集为此目的而构建。我们研究不同的标签选择如何影响检测性能,并提供标签策略的见解。
摘要:Since Text-to-Speech systems typically don't produce waveforms directly, recent spoof detection studies use resynthesized waveforms from vocoders and neural audio codecs to simulate an attacker. Unlike vocoders, which are specifically designed for speech synthesis, neural audio codecs were originally developed for compressing audio for storage and transmission. However, their ability to discretize speech also sparked interest in language-modeling-based speech synthesis. Owing to this dual functionality, codec resynthesized data may be labeled as either bonafide or spoof. So far, very little research has addressed this issue. In this study, we present a challenging extension of the ASVspoof 5 dataset constructed for this purpose. We examine how different labeling choices affect detection performance and provide insights into labeling strategies.


【4】Spatial Audio Question Answering and Reasoning on Dynamic Source Movements
标题:动态源运动的空间音频问题回答和推理
链接:https://arxiv.org/abs/2602.16334

作者:Arvind Krishna Sridhar,Yinyi Guo,Erik Visser
摘要:空间音频理解旨在使机器能够解释复杂的听觉场景,特别是当声源随时间移动时。在这项工作中,我们研究了空间音频问题推理(空间AQA),重点是运动推理,其中模型必须直接从立体声音频中推断对象的运动,位置和方向变化。首先,我们介绍了一个以运动为中心的空间音频增强框架,该框架可以从孤立的单声道音频事件中合成不同的运动模式,从而实现受控和可扩展的训练数据生成。其次,我们提出了一种具有思维模式的端到端多模态微调方法,该方法允许音频语言模型在预测答案之前产生明确的中间推理步骤。第三,我们调查查询条件源分离作为预处理阶段的影响,并比较三个推理制度:无掩蔽,音频接地模型(AGM),地面真相面具。我们的研究结果表明,推理放大了源分离的好处,当问题中存在单个事件时,思维模式显示出+5.1%的显着改善。这些发现突出了运动建模,推理和分离质量之间的相互作用,为推进空间音频理解提供了新的见解。
摘要:Spatial audio understanding aims to enable machines to interpret complex auditory scenes, particularly when sound sources move over time. In this work, we study Spatial Audio Question Answering (Spatial AQA) with a focus on movement reasoning, where a model must infer object motion, position, and directional changes directly from stereo audio. First, we introduce a movement-centric spatial audio augmentation framework that synthesizes diverse motion patterns from isolated mono audio events, enabling controlled and scalable training data generation. Second, we propose an end-to-end multimodal finetuning approach with a thinking mode, which allows audio-language models to produce explicit intermediate reasoning steps before predicting an answer. Third, we investigate the impact of query-conditioned source separation as a preprocessing stage and compare three inference regimes: no masking, an audio grounding model (AGM), and ground-truth masks. Our results show that reasoning amplifies the benefits of source separation, with thinking mode showing significant improvement of +5.1% when a single event is present in the question. These findings highlight the interplay between movement modeling, reasoning, and separation quality, offering new insights for advancing spatial audio understanding.


【5】BAT: Better Audio Transformer Guided by Convex Gated Probing
标题:BAT:由凸面门控探测引导的更好的音频Transformer
链接:https://arxiv.org/abs/2602.16305

作者:Houtan Ghaffari,Lukas Rauch,Christoph Scholz,Paul Devos
摘要:探测在计算机视觉中被广泛采用,以忠实地评估自监督学习(SSL)嵌入,因为微调可能会歪曲其内在质量。相比之下,音频SSL模型仍然依赖于微调,因为简单的探测无法释放其全部潜力,并在AudioSet上竞争SOTA时改变其排名。因此,需要一个强大而有效的探测机制来引导音频SSL的轨迹走向可靠和可重复的方法。我们介绍了凸门探测(CGP),这是一种基于原型的方法,可以大大缩小音频中微调和探测之间的差距。CGP通过门控机制有效地利用所有冻结层,并暴露潜在的任务相关信息的位置。在CGP的指导下,我们重新构建了当前SOTA音频模型的整个SSL管道,这些模型使用了以前SSL方法的遗留实现。通过改进数据预处理、模型架构和预训练配方,我们引入了Better Audio Transformer(BAT),并在音频基准上建立了新的SOTA。
摘要:Probing is widely adopted in computer vision to faithfully evaluate self-supervised learning (SSL) embeddings, as fine-tuning may misrepresent their inherent quality. In contrast, audio SSL models still rely on fine-tuning because simple probing fails to unlock their full potential and alters their rankings when competing for SOTA on AudioSet. Hence, a robust and efficient probing mechanism is required to guide the trajectory of audio SSL towards reliable and reproducible methods. We introduce Convex Gated Probing (CGP), a prototype-based method that drastically closes the gap between fine-tuning and probing in audio. CGP efficiently utilizes all frozen layers via a gating mechanism and exposes the location of latent task-relevant information. Guided by CGP, we rework the entire SSL pipeline of current SOTA audio models that use legacy implementations of prior SSL methods. By refining data preprocessing, model architecture, and pre-training recipe, we introduce Better Audio Transformer (BAT), and establish new SOTA on audio benchmarks.


【6】MAEB: Massive Audio Embedding Benchmark
标题:MAEB:海量音频嵌入基准
链接:https://arxiv.org/abs/2602.16008

作者:Adnan El Assadi,Isaac Chung,Chenghao Xiao,Roman Solomatin,Animesh Jha,Rahul Chand,Silky Singh,Kaitlyn Wang,Ali Sartaz Khan,Marc Moussa Nasser,Sufen Fong,Pengfei He,Alan Xiao,Ayush Sunil Munot,Aditya Shrivastava,Artem Gazizov,Niklas Muennighoff,Kenneth Enevoldsen
摘要:我们介绍了Massive Audio Embedding Benchmark(MAEB),这是一个大规模的基准测试,涵盖了100多种语言的语音,音乐,环境声音和跨模态音频文本推理的30个任务。我们评估了50多个模型,发现没有一个模型在所有任务中占主导地位:对比音频文本模型在环境声音分类方面表现出色(例如,ESC 50),但在多语言语音任务中得分接近随机(例如,SIB-FLEURS),而语音预训练模型则显示出相反的模式。聚类对所有模型来说都是一个挑战,即使是性能最好的模型也只能获得适度的结果。我们观察到,擅长声学理解的模型通常在语言任务上表现不佳,反之亦然。我们还表明,音频编码器在MAEB上的性能与它们在音频大语言模型中使用时的性能高度相关。MAEB是由MAEB+衍生而来的,MAEB+是98个任务的集合。MAEB旨在保持任务多样性,同时降低评估成本,并集成到MTEB生态系统中,以实现跨文本、图像和音频模式的统一评估。我们在https://github.com/embeddings-benchmark/mteb上发布了MAEB和所有98个任务以及代码和排行榜。
摘要:We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasoning in 100+ languages. We evaluate 50+ models and find that no single model dominates across all tasks: contrastive audio-text models excel at environmental sound classification (e.g., ESC50) but score near random on multilingual speech tasks (e.g., SIB-FLEURS), while speech-pretrained models show the opposite pattern. Clustering remains challenging for all models, with even the best-performing model achieving only modest results. We observe that models excelling on acoustic understanding often perform poorly on linguistic tasks, and vice versa. We also show that the performance of audio encoders on MAEB correlates highly with their performance when used in audio large language models. MAEB is derived from MAEB+, a collection of 98 tasks. MAEB is designed to maintain task diversity while reducing evaluation cost, and it integrates into the MTEB ecosystem for unified evaluation across text, image, and audio modalities. We release MAEB and all 98 tasks along with code and a leaderboard at https://github.com/embeddings-benchmark/mteb.


【7】SELEBI: Percussion-aware Time Stretching via Selective Magnitude Spectrogram Compression by Nonstationary Gabor Transform
标题:SEN BI:通过非平稳Gabor变换的选择性幅度谱图压缩进行打击感知时间扩展
链接:https://arxiv.org/abs/2602.16421

作者:Natsuki Akaishi,Nicki Holighaus,Kohei Yatabe
备注:This work has been submitted to the IEEE for possible publication
摘要:基于相位声码器的时间拉伸是用于音频信号的时间尺度修改的广泛使用的技术。然而,传统的实现方式遭受“撞击拖尾”,这是一种众所周知的伪像,其显著降低了撞击组件的质量。我们将此伪影归因于时间上模糊的幅度谱图和本地化的新生成的相位之间的基本时间尺度失配。为了解决这个问题,我们提出了SELEBI,信号自适应相位声码器算法,显着减少打击拖影,同时保持稳定性和完美的重建性能。与依赖于启发式处理或分量分离的传统方法不同,我们的方法利用了非平稳Gabor变换。通过动态调整分析窗口长度,将短窗口分配给包含与离散分量相关的显著能量的间隔,我们直接从时域信号计算时间局部幅度谱图。这种方法确保了幅度和相位的时间结构之间的更大一致性。此外,非平稳Gabor变换的完美重构特性保证了稳定,高保真的信号合成,与以前的启发式方法相比。实验结果表明,该方法有效地减轻了打击拖影,并产生自然的声音质量。
摘要:Phase vocoder-based time-stretching is a widely used technique for the time-scale modification of audio signals. However, conventional implementations suffer from ``percussion smearing,'' a well-known artifact that significantly degrades the quality of percussive components. We attribute this artifact to a fundamental time-scale mismatch between the temporally smeared magnitude spectrogram and the localized, newly generated phase. To address this, we propose SELEBI, a signal-adaptive phase vocoder algorithm that significantly reduces percussion smearing while preserving stability and the perfect reconstruction property. Unlike conventional methods that rely on heuristic processing or component separation, our approach leverages the nonstationary Gabor transform. By dynamically adapting analysis window lengths to assign short windows to intervals containing significant energy associated with percussive components, we directly compute a temporally localized magnitude spectrogram from the time-domain signal. This approach ensures greater consistency between the temporal structures of the magnitude and phase. Furthermore, the perfect reconstruction property of the nonstationary Gabor transform guarantees stable, high-fidelity signal synthesis, in contrast to previous heuristic approaches. Experimental results demonstrate that the proposed method effectively mitigates percussion smearing and yields natural sound quality.


【8】Online Single-Channel Audio-Based Sound Speed Estimation for Robust Multi-Channel Audio Control
标题:基于单通道音频的在线音速估计,用于鲁棒的多通道音频控制
链接:https://arxiv.org/abs/2602.16416

作者:Andreas Jonas Fuglsig,Mads Græsbøll Christensen,Jesper Rindom Jensen
备注:Preprint submitted to EUSIPCO 2026, under review
摘要:鲁棒的空间音频控制依赖于准确的声学传播模型,但环境变化,特别是声速的变化,会导致系统性不匹配,从而降低性能。现有的方法要么假设已知的声速,需要多个麦克风,要么依赖于单独的校准,这使得它们对于具有最小感测的系统来说是不切实际的。我们提出了一个在线的声速估计器,在一般多通道音频播放,只需要一个单一的观察麦克风。该方法利用再现信号的声速的结构化效果,并通过最小化测量的音频和参数声学模型之间的失配来估计它。仿真结果表明,准确的跟踪不同的输入信号的声速和改进的空间控制性能时,估计用于补偿传播误差的声音区域控制框架。
摘要:Robust spatial audio control relies on accurate acoustic propagation models, yet environmental variations, especially changes in the speed of sound, cause systematic mismatches that degrade performance. Existing methods either assume known sound speed, require multiple microphones, or rely on separate calibration, making them impractical for systems with minimal sensing. We propose an online sound speed estimator that operates during general multichannel audio playback and requires only a single observation microphone. The method exploits the structured effect of sound speed on the reproduced signal and estimates it by minimizing the mismatch between the measured audio and a parametric acoustic model. Simulations show accurate tracking of sound speed for diverse input signals and improved spatial control performance when the estimates are used to compensate propagation errors in a sound zone control framework.


【9】Multi-Channel Replay Speech Detection using Acoustic Maps
标题:使用声学地图的多通道回放语音检测
链接:https://arxiv.org/abs/2602.16399

作者:Michael Neri,Tuomas Virtanen
备注:Submitted to EUSIPCO 2026
摘要:重放攻击仍然是自动说话人验证系统的一个关键漏洞,特别是在实时语音助理应用程序中。在这项工作中,我们提出了声学地图作为一种新的空间特征表示重放语音检测从多通道录音。从离散方位角和仰角网格上的经典波束形成导出,声学地图对反映人类语音辐射和基于语音扬声器的重放之间的物理差异的定向能量分布进行编码。一个轻量级的卷积神经网络被设计来对这种表示进行操作,在具有大约6k个可训练参数的ReMASC数据集上实现有竞争力的性能。实验结果表明,声学地图提供了一个紧凑的和物理上可解释的特征空间,在不同的设备和声学环境的重放攻击检测。
摘要:Replay attacks remain a critical vulnerability for automatic speaker verification systems, particularly in real-time voice assistant applications. In this work, we propose acoustic maps as a novel spatial feature representation for replay speech detection from multi-channel recordings. Derived from classical beamforming over discrete azimuth and elevation grids, acoustic maps encode directional energy distributions that reflect physical differences between human speech radiation and loudspeaker-based replay. A lightweight convolutional neural network is designed to operate on this representation, achieving competitive performance on the ReMASC dataset with approximately 6k trainable parameters. Experimental results show that acoustic maps provide a compact and physically interpretable feature space for replay attack detection across different devices and acoustic environments.


【10】Color-based Emotion Representation for Speech Emotion Recognition
标题:语音情感识别中基于颜色的情感表示
链接:https://arxiv.org/abs/2602.16256

作者:Ryotaro Nagase,Ryoichi Takashima,Yoichi Yamashita
备注:Submitted to EUSIPCO2026
摘要:语音情感识别(SER)传统上依赖于分类或维度标签。然而,这种技术在表现情感的多样性和可解释性方面是有限的。为了克服这一限制,我们专注于颜色属性,如色调,饱和度和价值,以表示情绪作为连续和可解释的分数。我们通过众包的情感语音语料库的颜色属性进行标注和分析。此外,我们使用机器学习和深度学习建立了SER中颜色属性的回归模型,并探索了颜色属性回归和情感分类的多任务学习。结果,我们证明了颜色属性和情感之间的关系,在语音,并成功地开发了颜色属性回归模型SER。我们还表明,多任务学习提高了每个任务的性能。
摘要:Speech emotion recognition (SER) has traditionally relied on categorical or dimensional labels. However, this technique is limited in representing both the diversity and interpretability of emotions. To overcome this limitation, we focus on color attributes, such as hue, saturation, and value, to represent emotions as continuous and interpretable scores. We annotated an emotional speech corpus with color attributes via crowdsourcing and analyzed them. Moreover, we built regression models for color attributes in SER using machine learning and deep learning, and explored the multitask learning of color attribute regression and emotion classification. As a result, we demonstrated the relationship between color attributes and emotions in speech, and successfully developed color attribute regression models for SER. We also showed that multitask learning improved the performance of each task.


【11】How Much Does Machine Identity Matter in Anomalous Sound Detection at Test Time?
标题:机器身份在测试时异常声音检测中有多重要?
链接:https://arxiv.org/abs/2602.16253

作者:Kevin Wilkinghoff,Keisuke Imoto,Zheng-Hua Tan
摘要:异常声音检测(ASD)基准通常假设被监控机器的身份在测试时是已知的,并且记录以机器方式进行评估。然而,在多个已知机器同时操作的现实监控场景中,测试记录可能无法可靠地归因于特定机器,并且要求机器身份强加了部署约束,例如每台机器专用传感器。为了揭示隐藏在标准机器评估下的鲁棒性的性能下降和方法特定的差异,我们考虑对ASD评估协议进行最小修改,其中来自多台机器的测试记录被合并并联合评估,而无需在推理时访问机器身份。训练数据和评估指标保持不变,机器标识标签仅用于事后评估。具有代表性的ASD方法的实验表明,放松这一假设揭示了隐藏在标准机器评估下的性能退化和方法特定的鲁棒性差异,并且这些退化与隐式机器识别精度密切相关。
摘要:Anomalous sound detection (ASD) benchmarks typically assume that the identity of the monitored machine is known at test time and that recordings are evaluated in a machine-wise manner. However, in realistic monitoring scenarios with multiple known machines operating concurrently, test recordings may not be reliably attributable to a specific machine, and requiring machine identity imposes deployment constraints such as dedicated sensors per machine. To reveal performance degradations and method-specific differences in robustness that are hidden under standard machine-wise evaluation, we consider a minimal modification of the ASD evaluation protocol in which test recordings from multiple machines are merged and evaluated jointly without access to machine identity at inference time. Training data and evaluation metrics remain unchanged, and machine identity labels are used only for post hoc evaluation. Experiments with representative ASD methods show that relaxing this assumption reveals performance degradations and method-specific differences in robustness that are hidden under standard machine-wise evaluation, and that these degradations are strongly related to implicit machine identification accuracy.


【12】Real time fault detection in 3D printers using Convolutional Neural Networks and acoustic signals
标题:使用卷积神经网络和声信号对3D打印机进行实时故障检测
链接:https://arxiv.org/abs/2602.16118

作者:Muhammad Fasih Waheed,Shonda Bernadin
备注:6 pages
摘要:3D打印过程的可靠性和质量关键取决于机械故障的及时检测。传统的监测方法通常依赖于视觉检查和硬件传感器,这可能既昂贵又有限。本文探讨了一种可扩展的非接触式方法,用于实时音频信号分析,以检测3D打印机中的机械故障。通过在打印过程中捕获和分类声发射,我们的目标是识别常见故障,如喷嘴堵塞,断丝,跳轮和各种其他机械故障。利用卷积神经网络,我们实现了能够实时音频分类的算法,以及时检测这些故障。我们的方法包括进行一系列受控实验来收集音频数据,然后应用先进的机器学习模型进行故障检测。此外,我们还回顾了现有的关于制造业和3D打印中基于音频的故障检测的文献,以使我们的研究在更广泛的领域内得到应用。初步结果表明,当使用机器学习技术进行分析时,音频信号提供了增强实时故障检测的可靠且具有成本效益的方法。
摘要:The reliability and quality of 3D printing processes are critically dependent on the timely detection of mechanical faults. Traditional monitoring methods often rely on visual inspection and hardware sensors, which can be both costly and limited in scope. This paper explores a scalable and contactless method for the use of real-time audio signal analysis for detecting mechanical faults in 3D printers. By capturing and classifying acoustic emissions during the printing process, we aim to identify common faults such as nozzle clogging, filament breakage, pully skipping and various other mechanical faults. Utilizing Convolutional neural networks, we implement algorithms capable of real-time audio classification to detect these faults promptly. Our methodology involves conducting a series of controlled experiments to gather audio data, followed by the application of advanced machine learning models for fault detection. Additionally, we review existing literature on audio-based fault detection in manufacturing and 3D printing to contextualize our research within the broader field. Preliminary results demonstrate that audio signals, when analyzed with machine learning techniques, provide a reliable and cost-effective means of enhancing real-time fault detection.


【13】Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis
标题:Resp-Agent:一种基于代理的多模式呼吸声生成和疾病诊断系统
链接:https://arxiv.org/abs/2602.15909

作者:Pengfei Zhang,Tianxin Xie,Minghao Yang,Li Liu
备注:24 pages, 3 figures. Published as a conference paper at ICLR 2026. Code and data available at https://github.com/zpforlove/Resp-Agent
摘要:基于深度学习的呼吸听诊目前受到两个基本挑战的阻碍:(i)固有的信息丢失,因为将信号转换为频谱图会丢弃瞬态声学事件和临床背景;(ii)数据可用性有限,严重的类别不平衡加剧了这种情况。为了弥合这些差距,我们提出了Resp-Agent,一个自主的多模态系统编排的一种新的主动对抗课程代理(思想家-A $^2$CA)。与静态管道不同,Thinker-A$^2$CA作为一个中央控制器,可以主动识别诊断弱点,并在闭环中安排目标合成。为了解决代表性的差距,我们引入了一个模态编织诊断器,通过战略全球注意力和稀疏的音频锚将EHR数据与音频令牌编织在一起,捕获远程临床背景和毫秒级瞬态。为了解决数据缺口,我们设计了一个流匹配生成器,它通过模态注入来适应纯文本的大语言模型(LLM),将病理内容与声学风格解耦,以合成难以诊断的样本。作为这些努力的基础,我们介绍了Resp-229 k,这是一个229 k记录的基准语料库,与LLM提炼的临床叙述配对。大量的实验表明,Resp-Agent在不同的评估设置中始终优于先前的方法,提高了数据稀缺和长尾类不平衡下的诊断鲁棒性。我们的代码和数据可在https://github.com/zpforlove/Resp-Agent上获得。
摘要:Deep learning-based respiratory auscultation is currently hindered by two fundamental challenges: (i) inherent information loss, as converting signals into spectrograms discards transient acoustic events and clinical context; (ii) limited data availability, exacerbated by severe class imbalance. To bridge these gaps, we present Resp-Agent, an autonomous multimodal system orchestrated by a novel Active Adversarial Curriculum Agent (Thinker-A$^2$CA). Unlike static pipelines, Thinker-A$^2$CA serves as a central controller that actively identifies diagnostic weaknesses and schedules targeted synthesis in a closed loop. To address the representation gap, we introduce a Modality-Weaving Diagnoser that weaves EHR data with audio tokens via Strategic Global Attention and sparse audio anchors, capturing both long-range clinical context and millisecond-level transients. To address the data gap, we design a Flow Matching Generator that adapts a text-only Large Language Model (LLM) via modality injection, decoupling pathological content from acoustic style to synthesize hard-to-diagnose samples. As a foundation for these efforts, we introduce Resp-229k, a benchmark corpus of 229k recordings paired with LLM-distilled clinical narratives. Extensive experiments demonstrate that Resp-Agent consistently outperforms prior approaches across diverse evaluation settings, improving diagnostic robustness under data scarcity and long-tailed class imbalance. Our code and data are available at https://github.com/zpforlove/Resp-Agent.


eess.AS音频处理


【1】SELEBI: Percussion-aware Time Stretching via Selective Magnitude Spectrogram Compression by Nonstationary Gabor Transform
标题:SEN BI:通过非平稳Gabor变换的选择性幅度谱图压缩进行打击感知时间扩展
链接:https://arxiv.org/abs/2602.16421

作者:Natsuki Akaishi,Nicki Holighaus,Kohei Yatabe
备注:This work has been submitted to the IEEE for possible publication
摘要:基于相位声码器的时间拉伸是用于音频信号的时间尺度修改的广泛使用的技术。然而,传统的实现方式遭受“撞击拖尾”,这是一种众所周知的伪像,其显著降低了撞击组件的质量。我们将此伪影归因于时间上模糊的幅度谱图和本地化的新生成的相位之间的基本时间尺度失配。为了解决这个问题,我们提出了SELEBI,信号自适应相位声码器算法,显着减少打击拖影,同时保持稳定性和完美的重建性能。与依赖于启发式处理或分量分离的传统方法不同,我们的方法利用了非平稳Gabor变换。通过动态调整分析窗口长度,将短窗口分配给包含与离散分量相关的显著能量的间隔,我们直接从时域信号计算时间局部幅度谱图。这种方法确保了幅度和相位的时间结构之间的更大一致性。此外,非平稳Gabor变换的完美重构特性保证了稳定,高保真的信号合成,与以前的启发式方法相比。实验结果表明,该方法有效地减轻了打击拖影,并产生自然的声音质量。
摘要:Phase vocoder-based time-stretching is a widely used technique for the time-scale modification of audio signals. However, conventional implementations suffer from ``percussion smearing,'' a well-known artifact that significantly degrades the quality of percussive components. We attribute this artifact to a fundamental time-scale mismatch between the temporally smeared magnitude spectrogram and the localized, newly generated phase. To address this, we propose SELEBI, a signal-adaptive phase vocoder algorithm that significantly reduces percussion smearing while preserving stability and the perfect reconstruction property. Unlike conventional methods that rely on heuristic processing or component separation, our approach leverages the nonstationary Gabor transform. By dynamically adapting analysis window lengths to assign short windows to intervals containing significant energy associated with percussive components, we directly compute a temporally localized magnitude spectrogram from the time-domain signal. This approach ensures greater consistency between the temporal structures of the magnitude and phase. Furthermore, the perfect reconstruction property of the nonstationary Gabor transform guarantees stable, high-fidelity signal synthesis, in contrast to previous heuristic approaches. Experimental results demonstrate that the proposed method effectively mitigates percussion smearing and yields natural sound quality.


【2】Online Single-Channel Audio-Based Sound Speed Estimation for Robust Multi-Channel Audio Control
标题:基于单通道音频的在线音速估计,用于鲁棒的多通道音频控制
链接:https://arxiv.org/abs/2602.16416

作者:Andreas Jonas Fuglsig,Mads Græsbøll Christensen,Jesper Rindom Jensen
备注:Preprint submitted to EUSIPCO 2026, under review
摘要:鲁棒的空间音频控制依赖于准确的声学传播模型,但环境变化,特别是声速的变化,会导致系统性不匹配,从而降低性能。现有的方法要么假设已知的声速,需要多个麦克风,要么依赖于单独的校准,这使得它们对于具有最小感测的系统来说是不切实际的。我们提出了一个在线的声速估计器,在一般多通道音频播放,只需要一个单一的观察麦克风。该方法利用再现信号的声速的结构化效果,并通过最小化测量的音频和参数声学模型之间的失配来估计它。仿真结果表明,准确的跟踪不同的输入信号的声速和改进的空间控制性能时,估计用于补偿传播误差的声音区域控制框架。
摘要:Robust spatial audio control relies on accurate acoustic propagation models, yet environmental variations, especially changes in the speed of sound, cause systematic mismatches that degrade performance. Existing methods either assume known sound speed, require multiple microphones, or rely on separate calibration, making them impractical for systems with minimal sensing. We propose an online sound speed estimator that operates during general multichannel audio playback and requires only a single observation microphone. The method exploits the structured effect of sound speed on the reproduced signal and estimates it by minimizing the mismatch between the measured audio and a parametric acoustic model. Simulations show accurate tracking of sound speed for diverse input signals and improved spatial control performance when the estimates are used to compensate propagation errors in a sound zone control framework.


【3】Multi-Channel Replay Speech Detection using Acoustic Maps
标题:使用声学地图的多通道回放语音检测
链接:https://arxiv.org/abs/2602.16399

作者:Michael Neri,Tuomas Virtanen
备注:Submitted to EUSIPCO 2026
摘要:重放攻击仍然是自动说话人验证系统的一个关键漏洞,特别是在实时语音助理应用程序中。在这项工作中,我们提出了声学地图作为一种新的空间特征表示重放语音检测从多通道录音。从离散方位角和仰角网格上的经典波束形成导出,声学地图对反映人类语音辐射和基于语音扬声器的重放之间的物理差异的定向能量分布进行编码。一个轻量级的卷积神经网络被设计来对这种表示进行操作,在具有大约6k个可训练参数的ReMASC数据集上实现有竞争力的性能。实验结果表明,声学地图提供了一个紧凑的和物理上可解释的特征空间,在不同的设备和声学环境的重放攻击检测。
摘要:Replay attacks remain a critical vulnerability for automatic speaker verification systems, particularly in real-time voice assistant applications. In this work, we propose acoustic maps as a novel spatial feature representation for replay speech detection from multi-channel recordings. Derived from classical beamforming over discrete azimuth and elevation grids, acoustic maps encode directional energy distributions that reflect physical differences between human speech radiation and loudspeaker-based replay. A lightweight convolutional neural network is designed to operate on this representation, achieving competitive performance on the ReMASC dataset with approximately 6k trainable parameters. Experimental results show that acoustic maps provide a compact and physically interpretable feature space for replay attack detection across different devices and acoustic environments.


【4】Color-based Emotion Representation for Speech Emotion Recognition
标题:语音情感识别中基于颜色的情感表示
链接:https://arxiv.org/abs/2602.16256

作者:Ryotaro Nagase,Ryoichi Takashima,Yoichi Yamashita
备注:Submitted to EUSIPCO2026
摘要:语音情感识别(SER)传统上依赖于分类或维度标签。然而,这种技术在表现情感的多样性和可解释性方面是有限的。为了克服这一限制,我们专注于颜色属性,如色调,饱和度和价值,以表示情绪作为连续和可解释的分数。我们通过众包的情感语音语料库的颜色属性进行标注和分析。此外,我们使用机器学习和深度学习建立了SER中颜色属性的回归模型,并探索了颜色属性回归和情感分类的多任务学习。结果,我们证明了颜色属性和情感之间的关系,在语音,并成功地开发了颜色属性回归模型SER。我们还表明,多任务学习提高了每个任务的性能。
摘要:Speech emotion recognition (SER) has traditionally relied on categorical or dimensional labels. However, this technique is limited in representing both the diversity and interpretability of emotions. To overcome this limitation, we focus on color attributes, such as hue, saturation, and value, to represent emotions as continuous and interpretable scores. We annotated an emotional speech corpus with color attributes via crowdsourcing and analyzed them. Moreover, we built regression models for color attributes in SER using machine learning and deep learning, and explored the multitask learning of color attribute regression and emotion classification. As a result, we demonstrated the relationship between color attributes and emotions in speech, and successfully developed color attribute regression models for SER. We also showed that multitask learning improved the performance of each task.


【5】How Much Does Machine Identity Matter in Anomalous Sound Detection at Test Time?
标题:机器身份在测试时异常声音检测中有多重要?
链接:https://arxiv.org/abs/2602.16253

作者:Kevin Wilkinghoff,Keisuke Imoto,Zheng-Hua Tan
摘要:异常声音检测(ASD)基准通常假设被监控机器的身份在测试时是已知的,并且记录以机器方式进行评估。然而,在多个已知机器同时操作的现实监控场景中,测试记录可能无法可靠地归因于特定机器,并且要求机器身份强加了部署约束,例如每台机器专用传感器。为了揭示隐藏在标准机器评估下的鲁棒性的性能下降和方法特定的差异,我们考虑对ASD评估协议进行最小修改,其中来自多台机器的测试记录被合并并联合评估,而无需在推理时访问机器身份。训练数据和评估指标保持不变,机器标识标签仅用于事后评估。具有代表性的ASD方法的实验表明,放松这一假设揭示了隐藏在标准机器评估下的性能退化和方法特定的鲁棒性差异,并且这些退化与隐式机器识别精度密切相关。
摘要:Anomalous sound detection (ASD) benchmarks typically assume that the identity of the monitored machine is known at test time and that recordings are evaluated in a machine-wise manner. However, in realistic monitoring scenarios with multiple known machines operating concurrently, test recordings may not be reliably attributable to a specific machine, and requiring machine identity imposes deployment constraints such as dedicated sensors per machine. To reveal performance degradations and method-specific differences in robustness that are hidden under standard machine-wise evaluation, we consider a minimal modification of the ASD evaluation protocol in which test recordings from multiple machines are merged and evaluated jointly without access to machine identity at inference time. Training data and evaluation metrics remain unchanged, and machine identity labels are used only for post hoc evaluation. Experiments with representative ASD methods show that relaxing this assumption reveals performance degradations and method-specific differences in robustness that are hidden under standard machine-wise evaluation, and that these degradations are strongly related to implicit machine identification accuracy.


【6】Real time fault detection in 3D printers using Convolutional Neural Networks and acoustic signals
标题:使用卷积神经网络和声信号对3D打印机进行实时故障检测
链接:https://arxiv.org/abs/2602.16118

作者:Muhammad Fasih Waheed,Shonda Bernadin
备注:6 pages
摘要:3D打印过程的可靠性和质量关键取决于机械故障的及时检测。传统的监测方法通常依赖于视觉检查和硬件传感器,这可能既昂贵又有限。本文探讨了一种可扩展的非接触式方法,用于实时音频信号分析,以检测3D打印机中的机械故障。通过在打印过程中捕获和分类声发射,我们的目标是识别常见故障,如喷嘴堵塞,断丝,跳轮和各种其他机械故障。利用卷积神经网络,我们实现了能够实时音频分类的算法,以及时检测这些故障。我们的方法包括进行一系列受控实验来收集音频数据,然后应用先进的机器学习模型进行故障检测。此外,我们还回顾了现有的关于制造业和3D打印中基于音频的故障检测的文献,以使我们的研究在更广泛的领域内得到应用。初步结果表明,当使用机器学习技术进行分析时,音频信号提供了增强实时故障检测的可靠且具有成本效益的方法。
摘要:The reliability and quality of 3D printing processes are critically dependent on the timely detection of mechanical faults. Traditional monitoring methods often rely on visual inspection and hardware sensors, which can be both costly and limited in scope. This paper explores a scalable and contactless method for the use of real-time audio signal analysis for detecting mechanical faults in 3D printers. By capturing and classifying acoustic emissions during the printing process, we aim to identify common faults such as nozzle clogging, filament breakage, pully skipping and various other mechanical faults. Utilizing Convolutional neural networks, we implement algorithms capable of real-time audio classification to detect these faults promptly. Our methodology involves conducting a series of controlled experiments to gather audio data, followed by the application of advanced machine learning models for fault detection. Additionally, we review existing literature on audio-based fault detection in manufacturing and 3D printing to contextualize our research within the broader field. Preliminary results demonstrate that audio signals, when analyzed with machine learning techniques, provide a reliable and cost-effective means of enhancing real-time fault detection.


【7】Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis
标题:Resp-Agent:一种基于代理的多模式呼吸声生成和疾病诊断系统
链接:https://arxiv.org/abs/2602.15909

作者:Pengfei Zhang,Tianxin Xie,Minghao Yang,Li Liu
备注:24 pages, 3 figures. Published as a conference paper at ICLR 2026. Code and data available at https://github.com/zpforlove/Resp-Agent
摘要:基于深度学习的呼吸听诊目前受到两个基本挑战的阻碍:(i)固有的信息丢失,因为将信号转换为频谱图会丢弃瞬态声学事件和临床背景;(ii)数据可用性有限,严重的类别不平衡加剧了这种情况。为了弥合这些差距,我们提出了Resp-Agent,一个自主的多模态系统编排的一种新的主动对抗课程代理(思想家-A $^2$CA)。与静态管道不同,Thinker-A$^2$CA作为一个中央控制器,可以主动识别诊断弱点,并在闭环中安排目标合成。为了解决代表性的差距,我们引入了一个模态编织诊断器,通过战略全球注意力和稀疏的音频锚将EHR数据与音频令牌编织在一起,捕获远程临床背景和毫秒级瞬态。为了解决数据缺口,我们设计了一个流匹配生成器,它通过模态注入来适应纯文本的大语言模型(LLM),将病理内容与声学风格解耦,以合成难以诊断的样本。作为这些努力的基础,我们介绍了Resp-229 k,这是一个229 k记录的基准语料库,与LLM提炼的临床叙述配对。大量的实验表明,Resp-Agent在不同的评估设置中始终优于先前的方法,提高了数据稀缺和长尾类不平衡下的诊断鲁棒性。我们的代码和数据可在https://github.com/zpforlove/Resp-Agent上获得。
摘要:Deep learning-based respiratory auscultation is currently hindered by two fundamental challenges: (i) inherent information loss, as converting signals into spectrograms discards transient acoustic events and clinical context; (ii) limited data availability, exacerbated by severe class imbalance. To bridge these gaps, we present Resp-Agent, an autonomous multimodal system orchestrated by a novel Active Adversarial Curriculum Agent (Thinker-A$^2$CA). Unlike static pipelines, Thinker-A$^2$CA serves as a central controller that actively identifies diagnostic weaknesses and schedules targeted synthesis in a closed loop. To address the representation gap, we introduce a Modality-Weaving Diagnoser that weaves EHR data with audio tokens via Strategic Global Attention and sparse audio anchors, capturing both long-range clinical context and millisecond-level transients. To address the data gap, we design a Flow Matching Generator that adapts a text-only Large Language Model (LLM) via modality injection, decoupling pathological content from acoustic style to synthesize hard-to-diagnose samples. As a foundation for these efforts, we introduce Resp-229k, a benchmark corpus of 229k recordings paired with LLM-distilled clinical narratives. Extensive experiments demonstrate that Resp-Agent consistently outperforms prior approaches across diverse evaluation settings, improving diagnostic robustness under data scarcity and long-tailed class imbalance. Our code and data are available at https://github.com/zpforlove/Resp-Agent.


【8】Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
标题:使用交织的语义、声学和文本令牌扩展开放离散音频基础模型
链接:https://arxiv.org/abs/2602.16687

作者:Potsawee Manakul,Woody Haosheng Gan,Martijn Bartelds,Guangzhi Sun,William Held,Diyi Yang
摘要:当前的音频语言模型主要是文本优先的,要么扩展预训练的文本LLM主干,要么依赖于仅语义的音频令牌,限制了一般的音频建模。本文提出了一个系统的实证研究的原生音频基础模型,应用下一个令牌预测音频的规模,联合建模语义内容,声学细节和文本,以支持通用音频生成和跨模态的能力。我们为构建此类模型提供了全面的经验见解:(1)我们系统地研究设计选择--数据源、文本混合比和令牌组成--建立经过验证的训练配方。(2)我们通过IsoFLOP分析对离散音频模型进行了第一次标度律研究,对64个模型进行了3 {\times}10^{18}$到3 {\times}10^{20}$FLOP分析,发现最佳数据的增长速度比最佳模型大小快1.6\times $。(3)我们将这些经验教训应用于训练SODA(Scaling Open Discrete Audio,扩展开放离散音频),这是一套在500 B令牌上从135 M到4 B参数的模型,与我们的扩展预测和现有模型进行比较。SODA是各种音频/文本任务的灵活骨干--我们通过使用相同的统一架构对保留语音的语音到语音翻译进行微调来证明这一点。
摘要:Current audio language models are predominantly text-first, either extending pre-trained text LLM backbones or relying on semantic-only audio tokens, limiting general audio modeling. This paper presents a systematic empirical study of native audio foundation models that apply next-token prediction to audio at scale, jointly modeling semantic content, acoustic details, and text to support both general audio generation and cross-modal capabilities. We provide comprehensive empirical insights for building such models: (1) We systematically investigate design choices -- data sources, text mixture ratios, and token composition -- establishing a validated training recipe. (2) We conduct the first scaling law study for discrete audio models via IsoFLOP analysis on 64 models spanning $3{\times}10^{18}$ to $3{\times}10^{20}$ FLOPs, finding that optimal data grows 1.6$\times$ faster than optimal model size. (3) We apply these lessons to train SODA (Scaling Open Discrete Audio), a suite of models from 135M to 4B parameters on 500B tokens, comparing against our scaling predictions and existing models. SODA serves as a flexible backbone for diverse audio/text tasks -- we demonstrate this by fine-tuning for voice-preserving speech-to-speech translation, using the same unified architecture.


【9】Hardware-accelerated graph neural networks: an alternative approach for neuromorphic event-based audio classification and keyword spotting on SoC FPGA
标题:硬件加速图神经网络:在SoCFPG上进行基于神经形态事件的音频分类和关键词识别的替代方法
链接:https://arxiv.org/abs/2602.16442

作者:Kamil Jeziorek,Piotr Wzorek,Krzysztof Blachut,Hiroshi Nakano,Manon Dampfhoffer,Thomas Mesquida,Hiroaki Nishi,Thomas Dalgaty,Tomasz Kryjak
备注:Under revision in TRETS Journal
摘要:随着嵌入式边缘传感器记录的数据量的增加,特别是来自产生离散事件流的神经形态设备的数据量的增加,对硬件感知神经架构的需求越来越大,这些架构能够实现高效,低延迟和节能的本地处理。我们提出了一个FPGA实现的事件图神经网络的音频处理。我们利用人工耳蜗将时间序列信号转换为稀疏事件数据,从而降低内存和计算成本。我们的架构在SoC FPGA上实现,并在两个开源数据集上进行了评估。对于分类任务,我们的基线浮点模型在SHD数据集上实现了92.7%的准确率-仅比最先进水平低2.4%-同时需要的参数减少了10倍和67倍以上。在SSC上,我们的模型实现了66.9-71.0%的准确度。与基于FPGA的脉冲神经网络相比,我们的量化模型达到了92.3%的准确率,比它们高出19.3%,同时减少了资源使用和延迟。对于SSC,我们报告了第一个硬件加速评估。我们进一步展示了事件音频关键字定位的第一个端到端FPGA实现,将图卷积层与递归序列建模相结合。该系统实现了高达95%的字端检测准确率,仅10.53微秒的延迟和1.18 W的功耗,为高能效的事件驱动KWS建立了强大的基准。
摘要:As the volume of data recorded by embedded edge sensors increases, particularly from neuromorphic devices producing discrete event streams, there is a growing need for hardware-aware neural architectures that enable efficient, low-latency, and energy-conscious local processing. We present an FPGA implementation of event-graph neural networks for audio processing. We utilise an artificial cochlea that converts time-series signals into sparse event data, reducing memory and computation costs. Our architecture was implemented on a SoC FPGA and evaluated on two open-source datasets. For classification task, our baseline floating-point model achieves 92.7% accuracy on SHD dataset - only 2.4% below the state of the art - while requiring over 10x and 67x fewer parameters. On SSC, our models achieve 66.9-71.0% accuracy. Compared to FPGA-based spiking neural networks, our quantised model reaches 92.3% accuracy, outperforming them by up to 19.3% while reducing resource usage and latency. For SSC, we report the first hardware-accelerated evaluation. We further demonstrate the first end-to-end FPGA implementation of event-audio keyword spotting, combining graph convolutional layers with recurrent sequence modelling. The system achieves up to 95% word-end detection accuracy, with only 10.53 microsecond latency and 1.18 W power consumption, establishing a strong benchmark for energy-efficient event-driven KWS.


机器翻译由腾讯交互翻译提供,仅供参考