微信公众号:arXiv_Daily
cs.SD语音
【1】Improving Out-of-Domain Audio Deepfake Detection via Layer Selection and Fusion of SSL-Based Countermeasures
标题:通过基于SSL的对策的层选择和融合改进域外音频深度伪造检测
链接:https://arxiv.org/abs/2509.12003
摘要:基于冻结预训练自监督学习(SSL)编码器的音频深度伪造检测系统在与层加权池化方法(如多头因子化注意池化(MHFA))相结合时表现出高水平的性能。然而,它们仍然难以推广到域外(OOD)条件。我们通过在四个不同的测试语料库上研究六个不同的预训练SSL的行为来解决这个问题。我们执行逐层分析,以确定哪些层贡献最大。接下来,我们研究了池头,比较了基于单层的策略和通过MHFA自动选择的策略。我们观察到,选择最好的层可以获得非常好的结果,同时将系统参数减少高达80%。还观察到作为测试语料库和SSL模型的函数的性能的广泛变化,表明编码器的预训练策略起作用。最后,分数级融合的几个编码器提高了泛化到OOD攻击。
摘要:Audio deepfake detection systems based on frozen pre-trained self-supervised learning (SSL) encoders show a high level of performance when combined with layer-weighted pooling methods, such as multi-head factorized attentive pooling (MHFA). However, they still struggle to generalize to out-of-domain (OOD) conditions. We tackle this problem by studying the behavior of six different pre-trained SSLs, on four different test corpora. We perform a layer-by-layer analysis to determine which layers contribute most. Next, we study the pooling head, comparing a strategy based on a single layer with automatic selection via MHFA. We observed that selecting the best layer gave very good results, while reducing system parameters by up to 80%. A wide variation in performance as a function of test corpus and SSL model is also observed, showing that the pre-training strategy of the encoder plays a role. Finally, score-level fusion of several encoders improved generalization to OOD attacks.
【2】PoolingVQ: A VQVAE Variant for Reducing Audio Redundancy and Boosting Multi-Modal Fusion in Music Emotion Analysis
标题:PoolingVQ:一种VQVAE变体,用于减少音频冗余并促进音乐情感分析中的多模式融合
链接:https://arxiv.org/abs/2509.11976
摘要:多模态音乐情感分析利用音频和音频模态来增强性能。虽然主流方法专注于复杂的特征提取网络,但我们认为,缩短音频序列特征的长度以减轻冗余,特别是与ESTA的紧凑表示相比,可以有效地提高任务性能。为了实现这一点,我们开发了PoolingVQ相结合的矢量量化变分自编码器(VQVAE)和空间池,它直接压缩音频特征序列通过局部聚合,以减少冗余,然后设计了一个两阶段的共同注意力的方法来融合音频和视频信息。在公共数据集EMOPIA和VGQQ上的实验结果表明,我们的多模态框架实现了最先进的整体性能,PoolingVQ产生了一些改进。
摘要:Multimodal music emotion analysis leverages audio and MIDI modalities to enhance performance. While mainstream approaches focus on complex feature extraction networks, we posit that shortening the length of audio sequence features to mitigate redundancy, especially in contrast to MIDI's compact representation, may effectively boost task performance. To achieve this, we developed PoolingVQ by combining Vector Quantized Variational Autoencoder (VQVAE) with spatial pooling, which directly compresses audio feature sequences through local aggregation to reduce redundancy, then devised a two-stage co-attention approach to fuse audio and MIDI information. Experimental results on the public datasets EMOPIA and VGMIDI demonstrate that our multimodal framework achieves state-of-the-art overall performance, with PoolingVQ yielding some improvement.
【3】MusicSwarm: Biologically Inspired Intelligence for Music Composition
标题:MusicSwarm:生物启发的音乐创作智能
链接:https://arxiv.org/abs/2509.11973
摘要:我们表明,连贯的,长形式的音乐作品可以出现在一个分散的群相同的,冻结的基础模型,通过stigmergic,点对点的信号协调,没有任何权重更新。我们比较了一个集中的多智能体系统与一个全球性的评论家,一个完全分散的群体中,酒吧明智的代理商的感觉和存款和谐,节奏和结构的线索,适应短期记忆,并达成共识。通过符号、音频和图论分析,群体产生了卓越的质量,同时提供了更大的多样性和结构多样性,并在创造力指标上领先。动态合同走向一个稳定的配置互补的角色,和自相似性网络揭示了一个小世界的架构,有效的远程连接和专门的桥接图案,澄清如何当地的新奇巩固到全球音乐形式。通过将专业化从参数更新转移到交互规则,共享内存和动态共识,MusicSwarm提供了一种计算和数据高效的长期创意结构,可以立即从音乐转移到协作写作,设计和科学发现。
摘要:We show that coherent, long-form musical composition can emerge from a decentralized swarm of identical, frozen foundation models that coordinate via stigmergic, peer-to-peer signals, without any weight updates. We compare a centralized multi-agent system with a global critic to a fully decentralized swarm in which bar-wise agents sense and deposit harmonic, rhythmic, and structural cues, adapt short-term memory, and reach consensus. Across symbolic, audio, and graph-theoretic analyses, the swarm yields superior quality while delivering greater diversity and structural variety and leads across creativity metrics. The dynamics contract toward a stable configuration of complementary roles, and self-similarity networks reveal a small-world architecture with efficient long-range connectivity and specialized bridging motifs, clarifying how local novelties consolidate into global musical form. By shifting specialization from parameter updates to interaction rules, shared memory, and dynamic consensus, MusicSwarm provides a compute- and data-efficient route to long-horizon creative structure that is immediately transferable beyond music to collaborative writing, design, and scientific discovery.
【4】Data-Driven Analysis of Text-Conditioned AI-Generated Music: A Case Study with Suno and Udio
标题:文本条件人工智能生成音乐的数据驱动分析:Suno和Udio的案例研究
链接:https://arxiv.org/abs/2509.11824
备注:Submitted for review to TISMIR Digital Musicology special issue
摘要:通过文本提示创建音乐的在线人工智能平台(人工智能音乐),如Suno和Udio,现在正被数十万用户使用。一些人工智能音乐出现在多个国家的广告甚至图表中。这些平台是如何使用的?哪些主题能激励用户?本文使用2024年5月至10月期间这些平台用户生成的大量歌曲,为Suno和Udio回答了这些问题。使用最先进的文本嵌入模型,降维和聚类方法的组合,我们分析提示,标签和歌词,并自动注释和显示处理后的数据在交互式图。我们的研究结果揭示了突出的主题,歌词,语言偏好,提示策略,以及特殊的尝试,通过使用元标记的转向模型。为了促进对人工智能生成音乐的文化实践的音乐学研究,我们分享了我们的代码和资源。
摘要:Online AI platforms for creating music from text prompts (AI music), such as Suno and Udio, are now being used by hundreds of thousands of users. Some AI music is appearing in advertising, and even charting, in multiple countries. How are these platforms being used? What subjects are inspiring their users? This article answers these questions for Suno and Udio using a large collection of songs generated by users of these platforms from May to October 2024. Using a combination of state-of-the-art text embedding models, dimensionality reduction and clustering methods, we analyze the prompts, tags and lyrics, and automatically annotate and display the processed data in interactive plots. Our results reveal prominent themes in lyrics, language preference, prompting strategies, as well as peculiar attempts at steering models through the use of metatags. To promote the musicological study of the developing cultural practice of AI-generated music we share our code and resources.
【5】Neural Audio Codecs for Prompt-Driven Universal Source Separation
标题:用于预算驱动通用源分离的神经音频编解码器
链接:https://arxiv.org/abs/2509.11717
备注:21 pages, 1 figure, pre-print, under review
摘要:文本引导的源分离支持跨媒体和辅助应用程序的灵活音频编辑,但AudioSep等现有模型对于边缘部署来说计算量太大。神经音频编解码器(NAC)模型,如CodecFormer和SDCodec,计算效率高,但限于固定类别分离。我们介绍CodecSep,这是第一个基于NAC的设备通用文本驱动分离模型。CodecSep将DAC压缩与Transformer掩蔽相结合,后者由CLAP派生的FilLM参数调制。在匹配训练/提示协议下的六个开放域基准测试中,\textbf{CodecSep}在分离保真度(SI-SDR)方面超过\textbf{AudioSep},同时在感知质量(ViSQOL)和匹配或超过固定词干基线(TDANet、CodecFormer、SDCodec)方面保持竞争力。在代码流部署中,它只需要1.35~ GMAC的端到端--比像AudioSep这样的频谱域分离器大约少54\times $的计算(仅架构)--同时保持完全的比特流兼容。
摘要:Text-guided source separation supports flexible audio editing across media and assistive applications, but existing models like AudioSep are too compute-heavy for edge deployment. Neural audio codec (NAC) models such as CodecFormer and SDCodec are compute-efficient but limited to fixed-class separation. We introduce CodecSep, the first NAC-based model for on-device universal, text-driven separation. CodecSep combines DAC compression with a Transformer masker modulated by CLAP-derived FiLM parameters. Across six open-domain benchmarks under matched training/prompt protocols, \textbf{CodecSep} surpasses \textbf{AudioSep} in separation fidelity (SI-SDR) while remaining competitive in perceptual quality (ViSQOL) and matching or exceeding fixed-stem baselines (TDANet, CodecFormer, SDCodec). In code-stream deployments, it needs just 1.35~GMACs end-to-end -- approximately $54\times$ less compute ($25\times$ architecture-only) than spectrogram-domain separators like AudioSep -- while remaining fully bitstream-compatible.
【6】Scaling to Multimodal and Multichannel Heart Sound Classification: Fine-Tuning Wav2Vec 2.0 with Synthetic and Augmented Biosignals
标题:扩展到多模式和多通道心脏声音分类:利用合成和增强生物信号进行微调Wav2Vec 2.0
链接:https://arxiv.org/abs/2509.11606
备注:35 pages, 37 figures, 19 tables
摘要:心血管疾病(CVD)是全球死亡的主要原因,每年约有1790万人死亡。早期检测至关重要,这就需要准确和廉价的预筛查方法。深度学习最近已被应用于使用同步心音图(PCG)和心电图(ECG)信号以及多通道PCG(mPCG)对指示CVD的异常心音进行分类。然而,由于同步和多通道数据集的可用性有限,最先进的架构仍然没有得到充分利用。增强的数据集和预训练的模型提供了克服这些限制的途径,使基于transformer的架构能够得到有效的训练。这项工作将传统的信号处理与去噪扩散模型WaveGrad和DiffWave相结合,以创建一个增强的数据集,从而在多模态和多通道心音数据集上微调基于Wav 2 Vec 2.0的分类器。该方法实现了最先进的性能。在2016年CinC单通道PCG数据集上,准确率、未加权平均召回率(UAR)、敏感度、特异度和Matthew相关系数(MCC)分别达到92.48%、93.05%、93.63%、92.48%、94.93%和0.8283。使用来自CinC的训练-a数据集的同步PCG和ECG信号,准确性、UAR、灵敏度、特异性和MCC分别达到93.14%、92.21%、94.35%、90.10%、95.12%和0.8380。使用由mPCG数据组成的可穿戴背心数据集,该模型的准确率为77.13%,UAR为74.25%,灵敏度为86.47%,特异度为62.04%,MCC为0.5082。这些结果证明了在增强数据集的支持下,基于变压器的CVD检测模型的有效性,突出了其推进多模态和多通道心音分类的潜力。
摘要:Cardiovascular diseases (CVDs) are the leading cause of death worldwide, accounting for approximately 17.9 million deaths each year. Early detection is critical, creating a demand for accurate and inexpensive pre-screening methods. Deep learning has recently been applied to classify abnormal heart sounds indicative of CVDs using synchronised phonocardiogram (PCG) and electrocardiogram (ECG) signals, as well as multichannel PCG (mPCG). However, state-of-the-art architectures remain underutilised due to the limited availability of synchronised and multichannel datasets. Augmented datasets and pre-trained models provide a pathway to overcome these limitations, enabling transformer-based architectures to be trained effectively. This work combines traditional signal processing with denoising diffusion models, WaveGrad and DiffWave, to create an augmented dataset to fine-tune a Wav2Vec 2.0-based classifier on multimodal and multichannel heart sound datasets. The approach achieves state-of-the-art performance. On the Computing in Cardiology (CinC) 2016 dataset of single channel PCG, accuracy, unweighted average recall (UAR), sensitivity, specificity and Matthew's correlation coefficient (MCC) reach 92.48\%, 93.05\%, 93.63\%, 92.48\%, 94.93\% and 0.8283, respectively. Using the synchronised PCG and ECG signals of the training-a dataset from CinC, 93.14\%, 92.21\%, 94.35\%, 90.10\%, 95.12\% and 0.8380 are achieved for accuracy, UAR, sensitivity, specificity and MCC, respectively. Using a wearable vest dataset consisting of mPCG data, the model achieves 77.13\% accuracy, 74.25\% UAR, 86.47\% sensitivity, 62.04\% specificity, and 0.5082 MCC. These results demonstrate the effectiveness of transformer-based models for CVD detection when supported by augmented datasets, highlighting their potential to advance multimodal and multichannel heart sound classification.
【7】Acoustic Overspecification in Electronic Dance Music Taxonomy
标题:电子舞曲音乐分类中的声学过度规范
链接:https://arxiv.org/abs/2509.11474
备注:5 pages, 3 figures, conference paper
摘要:电子舞曲(EDM)分类通常依赖于行业定义的分类法,其中包含许多子流派,但这些区分的声学基础仍不清楚。目前的方法使用监督学习与规定的体裁标签,假设他们的有效性没有系统的评估。在本文中,我们提出了一个无监督的方法来发现的自然声学结构的EDM独立的商业标签。我们的方法结合了新颖的tempogram-based功能捕捉EDM的分层节奏模式与多标准的功能选择。为了验证我们的研究结果反映了真正的声学结构而不是方法上的伪影,我们将我们的结果与最先进的预训练音频嵌入(MERT和CLAP)进行了比较。我们的特征空间和嵌入表示都收敛到19-23个自然声学家族,而规定的是35个,这提供了一致的证据,表明当前EDM分类法中的显著过度指定约为三分之一。
摘要:Electronic Dance Music (EDM) classification typically relies on industry-defined taxonomies with numerous subgenres, yet the acoustic basis for these distinctions remains unclear. Current approaches use supervised learning with prescribed genre labels, assuming their validity without systematic evaluation. In this paper, we propose an unsupervised approach to discover the natural acoustic structure of EDM independent of commercial labels. Our method combines novel tempogram-based features capturing EDM's layered rhythmic patterns with multi-criteria feature selection. To validate that our findings reflect genuine acoustic structure rather than methodological artifacts, we compare our results against state-of-the-art pre-trained audio embeddings (MERT and CLAP). Both our feature space and embedding representations converge to 19-23 natural acoustic families compared to the prescribed 35, providing consistent evidence of significant overspecification in current EDM taxonomy by approximately one-third.
【8】FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
标题:FuseCodec:神经编解码器的语义上下文融合和监督
链接:https://arxiv.org/abs/2509.11425
摘要:语音标记化使离散表示和促进语音语言建模。然而,现有的神经编解码器捕获低级别的声学特征,忽略了人类语音固有的语义和上下文线索。虽然最近的努力引入了来自自监督语音模型的语义表示或结合了来自预训练语言模型的上下文表示,但在对齐和统一语义和上下文表示方面仍然存在挑战。我们引入了FuseCodec,它通过强大的跨模态对齐和全局信息监督来统一声学,语义和上下文表示。我们提出了三种互补技术:(i)潜在表示融合,将语义和上下文特征直接集成到编码器潜在空间中,以实现鲁棒和统一的表示学习;(ii)全局语义上下文监督,用全局池化和广播表示监督离散令牌,以增强时间一致性和跨模态对齐;以及(iii)时间对齐的上下文监督,通过在本地窗口内动态匹配上下文和语音令牌来加强对齐,以进行细粒度的令牌级监督。我们进一步介绍FuseCodec-TTS,证明我们的方法的适用性zero-shot语音合成。从经验上看,FuseCodec在LibriSpeech中实现了最先进的性能,在转录准确性、感知质量、可理解性和扬声器相似性方面超过了EnCodec、SpeechTokenizer和DAC。结果突出了上下文和语义引导的语音标记和下游任务的标记化的有效性。代码和预训练模型可在https://github.com/mubtasimahasan/FuseCodec上获得。
摘要:Speech tokenization enables discrete representation and facilitates speech language modeling. However, existing neural codecs capture low-level acoustic features, overlooking the semantic and contextual cues inherent to human speech. While recent efforts introduced semantic representations from self-supervised speech models or incorporated contextual representations from pre-trained language models, challenges remain in aligning and unifying the semantic and contextual representations. We introduce FuseCodec, which unifies acoustic, semantic, and contextual representations through strong cross-modal alignment and globally informed supervision. We propose three complementary techniques: (i) Latent Representation Fusion, integrating semantic and contextual features directly into the encoder latent space for robust and unified representation learning; (ii) Global Semantic-Contextual Supervision, supervising discrete tokens with globally pooled and broadcasted representations to enhance temporal consistency and cross-modal alignment; and (iii) Temporally Aligned Contextual Supervision, strengthening alignment by dynamically matching contextual and speech tokens within a local window for fine-grained token-level supervision. We further introduce FuseCodec-TTS, demonstrating our methodology's applicability to zero-shot speech synthesis. Empirically, FuseCodec achieves state-of-the-art performance in LibriSpeech, surpassing EnCodec, SpeechTokenizer, and DAC in transcription accuracy, perceptual quality, intelligibility, and speaker similarity. Results highlight the effectiveness of contextually and semantically guided tokenization for speech tokenization and downstream tasks. Code and pretrained models are available at https://github.com/mubtasimahasan/FuseCodec.
【9】Revisiting Meter Tracking in Carnatic Music using Deep Learning Approaches
标题:使用深度学习方法重新审视狂欢音乐中的电表跟踪
链接:https://arxiv.org/abs/2509.11241
摘要:节拍和强拍跟踪,统称为节拍跟踪,是音乐信息检索(MIR)中的一项基本任务。在这一领域,深度学习模型已经远远超过了传统的信号处理和经典的机器学习方法,特别是对于西方(欧洲遗传学)流派来说,大型注释数据集广泛存在。然而,这些系统在代表性不足的音乐传统上表现得不太可靠。卡纳蒂克音乐是印度次大陆的丰富传统,以其复杂的节奏和独特的韵律结构而闻名(唉)。在此背景下,关于仪表跟踪的最值得注意的先前工作采用了概率动态贝叶斯网络(DBN)。然而,最先进的(SOTA)深度学习模型在Carnatic音乐上的性能在很大程度上仍未得到探索。 在这项研究中,我们评估了两个模型的节拍跟踪在卡纳蒂克音乐:时间卷积网络(TCN),一个轻量级的架构,已成功地适应拉丁节奏,和击败这个!,一个基于transformer的模型,设计用于广泛的风格覆盖,而不需要后期处理。在Carnatic音乐节奏(CMR$_f$)数据集上复制DBN基线的实验设置,我们系统地评估了这些模型在直接可比设置中的性能。我们进一步研究了适应策略,包括微调Carnatic数据模型和使用音乐信息参数。结果表明,虽然现成的模型并不总是优于DBN,但通过迁移学习,它们的性能大幅提高,匹配或超过基线。这些发现表明,SOTA深度学习模型可以有效地适应代表性不足的传统,为更具包容性和广泛适用的仪表跟踪系统铺平道路。
摘要:Beat and downbeat tracking, jointly referred to as Meter Tracking, is a fundamental task in Music Information Retrieval (MIR). Deep learning models have far surpassed traditional signal processing and classical machine learning approaches in this domain, particularly for Western (Eurogenetic) genres, where large annotated datasets are widely available. These systems, however, perform less reliably on underrepresented musical traditions. Carnatic music, a rich tradition from the Indian subcontinent, is renowned for its rhythmic intricacy and unique metrical structures (t\=alas). The most notable prior work on meter tracking in this context employed probabilistic Dynamic Bayesian Networks (DBNs). The performance of state-of-the-art (SOTA) deep learning models on Carnatic music, however, remains largely unexplored. In this study, we evaluate two models for meter tracking in Carnatic music: the Temporal Convolutional Network (TCN), a lightweight architecture that has been successfully adapted for Latin rhythms, and Beat This!, a transformer-based model designed for broad stylistic coverage without the need for post-processing. Replicating the experimental setup of the DBN baseline on the Carnatic Music Rhythm (CMR$_f$) dataset, we systematically assess the performance of these models in a directly comparable setting. We further investigate adaptation strategies, including fine-tuning the models on Carnatic data and the use of musically informed parameters. Results show that while off-the-shelf models do not always outperform the DBN, their performance improves substantially with transfer learning, matching or surpassing the baseline. These findings indicate that SOTA deep learning models can be effectively adapted to underrepresented traditions, paving the way for more inclusive and broadly applicable meter tracking systems.
【10】WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation
标题:WeaveMuse:一个用于多模式音乐理解和生成的开放式统计系统
链接:https://arxiv.org/abs/2509.11183
备注:Accepted at Large Language Models for Music & Audio Workshop (LLM4MA) 2025
摘要:人工智能已经在工业中标准化,作为协调专业模型和工具以解决复杂多模式任务的实用范例。在这项工作中,我们提出了WeaveMuse,一个多智能体系统的音乐理解,符号组成和音频合成。每个专家代理解释用户请求,导出机器可操作的要求(模态,格式,约束),并验证自己的输出,而管理员代理选择和排序工具,调解用户交互,并保持跨回合的状态。该系统可以在本地扩展和部署,使用量化和推理策略来适应不同的硬件预算,或者通过HFApi来保留对开放模型的免费社区访问。除了开箱即用之外,该系统还通过约束模式、结构化解码、基于策略的推理和参数高效适配器或为MIR任务定制模型的提炼变体来强调可控性和适应性。一个中心的设计目标是促进跨文本,符号符号和可视化,和音频,使分析-合成-渲染循环和解决跨格式的约束多式联运的互动。该框架旨在通过支持各种大小的可互换开源模型、灵活的内存管理和可重现的部署路径,使MIR工具民主化、实现和可访问。
摘要:Agentic AI has been standardized in industry as a practical paradigm for coordinating specialized models and tools to solve complex multimodal tasks. In this work, we present WeaveMuse, a multi-agent system for music understanding, symbolic composition, and audio synthesis. Each specialist agent interprets user requests, derives machine-actionable requirements (modalities, formats, constraints), and validates its own outputs, while a manager agent selects and sequences tools, mediates user interaction, and maintains state across turns. The system is extendable and deployable either locally, using quantization and inference strategies to fit diverse hardware budgets, or via the HFApi to preserve free community access to open models. Beyond out-of-the-box use, the system emphasizes controllability and adaptation through constraint schemas, structured decoding, policy-based inference, and parameter-efficient adapters or distilled variants that tailor models to MIR tasks. A central design goal is to facilitate intermodal interaction across text, symbolic notation and visualization, and audio, enabling analysis-synthesis-render loops and addressing cross-format constraints. The framework aims to democratize, implement, and make accessible MIR tools by supporting interchangeable open-source models of various sizes, flexible memory management, and reproducible deployment paths.
【11】An Entropy-Guided Curriculum Learning Strategy for Data-Efficient Acoustic Scene Classification under Domain Shift
标题:领域转移下数据高效的声学场景分类的信息引导课程学习策略
链接:https://arxiv.org/abs/2509.11168
备注:Accepted at the Detection and Classification of Acoustic Scenes and Events (DCASE) Workshop 2025
摘要:声学场景分类(ASC)面临着在记录设备之间推广的挑战,特别是当标记数据有限时。DCASE 2024挑战任务1通过要求模型从记录在几个设备上的小标记子集中学习来强调这个问题。然后,这些模型需要在严格的复杂性约束下推广到以前看不见的设备的记录。虽然数据增强和使用预先训练的模型等技术已经很好地用于提高模型泛化,但优化训练策略代表了一种补充但较少探索的路径,不会引入额外的架构复杂性或推理开销。在各种培训策略中,课程学习提供了一个很有前途的范例,通过从简单到困难的例子来构建学习过程。在这项工作中,我们提出了一个熵引导的课程学习策略,以解决数据高效ASC域转移问题。具体来说,我们量化的不确定性,设备域预测每个训练样本通过计算香农熵的设备后验概率估计的辅助域分类。使用熵作为域不变性的代理,该课程从高熵样本开始,并逐渐结合低熵,特定于域的样本,以促进可推广表示的学习。在多个DCASE 2024 ASC基线上的实验结果表明,我们的策略有效地减轻了域偏移,特别是在有限的标记数据条件下。我们的策略与架构无关,并且不引入额外的推理成本,使其可以轻松集成到现有的ASC基线中,并为域转移提供实用的解决方案。
摘要:Acoustic Scene Classification (ASC) faces challenges in generalizing across recording devices, particularly when labeled data is limited. The DCASE 2024 Challenge Task 1 highlights this issue by requiring models to learn from small labeled subsets recorded on a few devices. These models need to then generalize to recordings from previously unseen devices under strict complexity constraints. While techniques such as data augmentation and the use of pre-trained models are well-established for improving model generalization, optimizing the training strategy represents a complementary yet less-explored path that introduces no additional architectural complexity or inference overhead. Among various training strategies, curriculum learning offers a promising paradigm by structuring the learning process from easier to harder examples. In this work, we propose an entropy-guided curriculum learning strategy to address the domain shift problem in data-efficient ASC. Specifically, we quantify the uncertainty of device domain predictions for each training sample by computing the Shannon entropy of the device posterior probabilities estimated by an auxiliary domain classifier. Using entropy as a proxy for domain invariance, the curriculum begins with high-entropy samples and gradually incorporates low-entropy, domain-specific ones to facilitate the learning of generalizable representations. Experimental results on multiple DCASE 2024 ASC baselines demonstrate that our strategy effectively mitigates domain shift, particularly under limited labeled data conditions. Our strategy is architecture-agnostic and introduces no additional inference cost, making it easily integrable into existing ASC baselines and offering a practical solution to domain shift.
【12】ENJ: Optimizing Noise with Genetic Algorithms to Jailbreak LSMs
标题:ENJ:用遗传算法优化噪音以越狱LSM
链接:https://arxiv.org/abs/2509.11128
摘要:大型语音模型的广泛应用使得其安全隐患日益突出。传统的语音对抗攻击方法面临着平衡有效性和隐蔽性的挑战。本文提出了进化噪声越狱(ENJ),它利用遗传算法将环境噪声从被动干扰转化为可主动优化的攻击载体,用于越狱LSM。该方法通过种群初始化、交叉融合和概率变异等操作,迭代进化出一系列音频样本,将恶意指令与背景噪声融合。这些样本对人类来说听起来像是无害的噪音,但可能会导致模型解析和执行有害的命令。在多个主流语音模型上的实验表明,ENJ的攻击效果明显优于现有的基线方法。该研究揭示了噪声在语音安全中的双重作用,为复杂声学环境下的模型安全防御提供了新的重要见解。
摘要:The widespread application of Large Speech Models (LSMs) has made their security risks increasingly prominent. Traditional speech adversarial attack methods face challenges in balancing effectiveness and stealth. This paper proposes Evolutionary Noise Jailbreak (ENJ), which utilizes a genetic algorithm to transform environmental noise from a passive interference into an actively optimizable attack carrier for jailbreaking LSMs. Through operations such as population initialization, crossover fusion, and probabilistic mutation, this method iteratively evolves a series of audio samples that fuse malicious instructions with background noise. These samples sound like harmless noise to humans but can induce the model to parse and execute harmful commands. Extensive experiments on multiple mainstream speech models show that ENJ's attack effectiveness is significantly superior to existing baseline methods. This research reveals the dual role of noise in speech security and provides new critical insights for model security defense in complex acoustic environments.
【13】STASE: A spatialized text-to-audio synthesis engine for music generation
标题:STASE:用于音乐生成的空间化文本到音频合成引擎
链接:https://arxiv.org/abs/2509.11124
备注:Accepted to LLM4Music @ ISMIR 2025
摘要:虽然许多文本到音频系统产生单声道或固定立体声输出,但生成具有用户定义的空间属性的音频仍然是一个挑战。现有的基于深度学习的空间化方法通常依赖于潜在空间操作,这可能会限制对空间感知至关重要的心理声学参数的直接控制。为了解决这个问题,我们引入STASE,一个利用大型语言模型(LLM)作为代理来解释文本中的空间线索的系统。STASE的一个关键特征是将语义解释从一个单独的、基于物理的空间渲染引擎中分离出来,这有利于可解释和用户可控的空间推理。LLM过程通过两个主要途径提示:(i)描述符,用于直接映射显式空间信息(例如,“将主音吉他放置在45{\deg}方位角,10米距离”),以及(ii)抽象图,其中检索增强生成(RAG)模块检索相关的空间模板以通知渲染。本文详细介绍了STASE的工作流程,讨论了实现的考虑因素,并强调了当前在评估生成空间音频的挑战。
摘要:While many text-to-audio systems produce monophonic or fixed-stereo outputs, generating audio with user-defined spatial properties remains a challenge. Existing deep learning-based spatialization methods often rely on latent-space manipulations, which can limit direct control over psychoacoustic parameters critical to spatial perception. To address this, we introduce STASE, a system that leverages a Large Language Model (LLM) as an agent to interpret spatial cues from text. A key feature of STASE is the decoupling of semantic interpretation from a separate, physics-based spatial rendering engine, which facilitates interpretable and user-controllable spatial reasoning. The LLM processes prompts through two main pathways: (i) Description Prompts, for direct mapping of explicit spatial information (e.g., "place the lead guitar at 45{\deg} azimuth, 10 m distance"), and (ii) Abstract Prompts, where a Retrieval-Augmented Generation (RAG) module retrieves relevant spatial templates to inform the rendering. This paper details the STASE workflow, discusses implementation considerations, and highlights current challenges in evaluating generative spatial audio.
【14】Emoanti: audio anti-deepfake with refined emotion-guided representations
标题:Deliveranti:具有精致情感引导表示的音频反深度伪造
链接:https://arxiv.org/abs/2509.10781
摘要:音频deepfake是如此复杂,缺乏有效的检测方法是致命的。虽然大多数检测系统主要依赖于低级声学特征或预训练的语音表示,但它们经常忽略高级情感线索,这些线索可以提供补充和潜在的反深度伪造信息,以增强泛化能力。在这项工作中,我们提出了一种新的音频反深度伪造系统,该系统通过利用预先训练的Wav 2 Vec 2(W2 V2)模型来利用情感特征(EST-Anti),该模型在情感识别任务上进行了微调,该模型导出了情感引导的表示,然后设计了一个基于卷积层的专用特征提取器,该卷积层具有剩余连接,以有效地捕获和细化Transformer层输出的情感特征。实验结果表明,我们提出的架构在ASVspoof 2019 LA和ASVspoof 2021 LA基准测试中都达到了最先进的性能,并在ASVspoof 2021 DF数据集上表现出很强的泛化能力。我们提出的方法的代码可以在Anonymous GitHub 1上找到。
摘要:Audio deepfake is so sophisticated that the lack of effective detection methods is fatal. While most detection systems primarily rely on low-level acoustic features or pretrained speech representations, they frequently neglect high-level emotional cues, which can offer complementary and potentially anti-deepfake information to enhance generalization. In this work, we propose a novel audio anti-deepfake system that utilizes emotional features (EmoAnti) by exploiting a pretrained Wav2Vec2 (W2V2) model fine-tuned on emotion recognition tasks, which derives emotion-guided representations, then designing a dedicated feature extractor based on convolutional layers with residual connections to effectively capture and refine emotional characteristics from the transformer layers outputs. Experimental results show that our proposed architecture achieves state-of-the-art performance on both the ASVspoof2019LA and ASVspoof2021LA benchmarks, and demonstrates strong generalization on the ASVspoof2021DF dataset. Our proposed approach's code is available at Anonymous GitHub1.
【15】Combining Audio and Non-Audio Inputs in Evolved Neural Networks for Ovenbird
标题:在Ovenbird的进化神经网络中结合音频和非音频输入
链接:https://arxiv.org/abs/2509.10566
摘要:在过去的几年里,神经网络作为从数字数据中自动进行物种分类的工具的使用有所增加。这部分是由于通过卷积神经网络(CNN)的图像分类的高分类精度。在音频数据的情况下,基于CNN的识别器用于通过使用来自声音可视化的信息(即,光谱图)。这些识别器通常使用声谱图作为其唯一输入。然而,研究人员有其他非音频数据,如物种的栖息地偏好,物候和范围信息,可以改善物种分类。在本文中,我们提出了如何单物种识别神经网络的准确性可以提高使用非音频数据作为输入,除了频谱图信息。我们还分析了这些改进是否仅仅是具有更多参数的神经网络的结果,而不是结合两个输入。我们发现,使用两种不同输入的网络比仅使用其中一种输入的类似规模的网络具有更高的分类准确性。
摘要:In the last several years the use of neural networks as tools to automate species classification from digital data has increased. This has been due in part to the high classification accuracy of image classification through Convolutional Neural Networks (CNN). In the case of audio data CNN based recognizers are used to automate the classification of species in audio recordings by using information from sound visualization (i.e., spectrograms). It is common for these recognizers to use the spectrogram as their sole input. However, researchers have other non-audio data, such as habitat preferences of a species, phenology, and range information, available that could improve species classification. In this paper we present how a single-species recognizer neural network's accuracy can be improved by using non-audio data as inputs in addition to spectrogram information. We also analyze if the improvements are merely a result of having a neural network with a higher number of parameters instead of combining the two inputs. We find that networks that use the two different inputs have a higher classification accuracy than networks of similar size that use only one of the inputs.
【16】Evaluating Automatic Speech Recognition Systems for Korean Meteorological Experts
标题:韩国气象专家评估自动语音识别系统
链接:https://arxiv.org/abs/2410.18444
备注:EMNLP 2025 Findings
摘要:本文探讨了将自动语音识别(ASR)集成到自然语言查询系统中,以提高韩国气象学家的天气预报效率。我们解决了为韩国天气领域开发ASR系统的挑战,特别是专业词汇和韩国语言的复杂性。为了解决这些问题,我们构建了一个评估数据集的口语查询记录的母语为韩国人。使用该数据集,我们评估了多语言ASR模型系列的各种配置,确定了与特定领域术语相关的性能限制。然后,我们实现了一个简单的基于文本到语音的数据增强方法,该方法在保持一般域性能的同时提高了对专业术语的识别。我们的贡献包括创建特定领域的数据集,全面的ASR模型评估和有效的增强技术。我们相信,我们的工作为韩国天气预报领域ASR的未来发展奠定了基础。
摘要:This paper explores integrating Automatic Speech Recognition (ASR) into natural language query systems to improve weather forecasting efficiency for Korean meteorologists. We address challenges in developing ASR systems for the Korean weather domain, specifically specialized vocabulary and Korean linguistic intricacies. To tackle these issues, we constructed an evaluation dataset of spoken queries recorded by native Korean speakers. Using this dataset, we assessed various configurations of a multilingual ASR model family, identifying performance limitations related to domain-specific terminology. We then implemented a simple text-to-speech-based data augmentation method, which improved the recognition of specialized terms while maintaining general-domain performance. Our contributions include creating a domain-specific dataset, comprehensive ASR model evaluations, and an effective augmentation technique. We believe our work provides a foundation for future advancements in ASR for the Korean weather forecasting domain.
【17】Spectral and Rhythm Features for Audio Classification with Deep Convolutional Neural Networks
标题:利用深度卷积神经网络进行音频分类的频谱和节奏特征
链接:https://arxiv.org/abs/2410.06927
摘要:卷积神经网络(CNN)广泛应用于计算机视觉。它们不仅可以用于传统的数字图像材料来识别模式,而且还可以用于从数字图像中提取特征,这些特征表示从时域数字音频信号中提取的频谱和节奏特征,用于声音的声学分类。不同的频谱和节奏特征表示,如梅尔缩放频谱图,梅尔频率倒谱系数(MFCC),循环tempogram,短时傅里叶变换(STFT)chromagram,恒定Q变换(CQT)chromagram和色度能量归一化统计(CENS)chromagram的音频分类性能使用深度卷积神经网络进行了研究。可以清楚地表明,对于使用深度CNN的音频分类任务,梅尔缩放谱图和梅尔频率倒谱系数(MFCC)的性能明显优于本研究中研究的其他谱和节奏特征。实验是在ESC-50数据集的帮助下进行的,该数据集包含2,000个标记的环境音频记录。
摘要:Convolutional neural networks (CNNs) are widely used in computer vision. They can be used not only for conventional digital image material to recognize patterns, but also for feature extraction from digital imagery representing spectral and rhythm features extracted from time-domain digital audio signals for the acoustic classification of sounds. Different spectral and rhythm feature representations like mel-scaled spectrograms, mel-frequency cepstral coefficients (MFCCs), cyclic tempograms, short-time Fourier transform (STFT) chromagrams, constant-Q transform (CQT) chromagrams and chroma energy normalized statistics (CENS) chromagrams are investigated in terms of the audio classification performance using a deep convolutional neural network. It can be clearly shown that the mel-scaled spectrograms and the mel-frequency cepstral coefficients (MFCCs) perform significantly better than the other spectral and rhythm features investigated in this research for audio classification tasks using deep CNNs. The experiments were carried out with the aid of the ESC-50 dataset with 2,000 labeled environmental audio recordings.
【18】Length-Aware Rotary Position Embedding for Text-Speech Alignment
标题:用于文本语音对齐的长度感知旋转位置嵌入
链接:https://arxiv.org/abs/2509.11084
备注:5 pages, 3 figures, preprint
摘要:许多最近的文本到语音(TTS)系统是建立在Transformer架构,并采用交叉注意机制的文本语音对齐。在这些系统中,旋转位置嵌入(RoPE)通常用于在文本和语音表示中编码位置信息。在这项工作中,我们引入了长度感知的RoPE(LARoPE),一个简单而有效的扩展RoPE,提高文本语音对齐。与依赖于绝对索引的RoPE不同,LARoPE使用长度规范化索引计算查询和关键位置之间的相对距离。实验结果表明,LARoPE始终优于RoPE,提供更快的损失收敛,更准确的文本语音对齐,更高的整体TTS质量。此外,LARoPE表现出更大的弹性话语持续时间的变化,并保持稳定的性能,在延长语音生成高达30秒,而RoPE遭受显着的退化。值得注意的是,我们的方法实现了最先进的字错误率的标准zero-shot TTS基准。
摘要:Many recent text-to-speech (TTS) systems are built on transformer architectures and employ cross-attention mechanisms for text-speech alignment. Within these systems, rotary position embedding (RoPE) is commonly used to encode positional information in text and speech representations. In this work, we introduce length-aware RoPE (LARoPE), a simple yet effective extension of RoPE that improves text-speech alignment. Unlike RoPE, which relies on absolute indices, LARoPE computes relative distances between query and key positions using length-normalized indices. Experimental results show that LARoPE consistently outperforms RoPE, offering faster loss convergence, more accurate text-speech alignment, and higher overall TTS quality. Furthermore, LARoPE demonstrates greater resilience to variations in utterance duration and maintains stable performance in extended speech generation up to 30 seconds, whereas RoPE suffers from notable degradation. Notably, our method achieves a state-of-the-art word error rate on a standard zero-shot TTS benchmark.
【19】Local Density-Based Anomaly Score Normalization for Domain Generalization
标题:基于局部密度的异常分数规范化领域概括
链接:https://arxiv.org/abs/2509.10951
摘要:域移位条件下的最先进的异常声音检测(ASD)系统依赖于将音频信号投影到嵌入空间中并使用基于距离的离群值检测来计算异常分数。要克服的主要困难之一是在声学上和在所提供的训练数据的量方面不同的源域和目标域的异常分数分布之间的所谓的域失配。对于一个域最优的决策阈值对于另一个域可能是高度次优的,反之亦然。当仅使用单个决策阈值时,这显著降低了性能,这是在泛化到在训练期间可能不可见的多个数据域同时仍然使用与源域中相同的经训练的ASD系统时所需要的。为了减少域之间的这种不匹配,我们提出了一个简单的基于局部密度的异常分数归一化方案。在几个ASD数据集上进行的实验中,我们表明,所提出的归一化方案始终提高了各种类型的基于嵌入的ASD系统的性能,并产生比现有的异常分数归一化方法更好的结果。
摘要:State-of-the-art anomalous sound detection (ASD) systems in domain-shifted conditions rely on projecting audio signals into an embedding space and using distance-based outlier detection to compute anomaly scores. One of the major difficulties to overcome is the so-called domain mismatch between the anomaly score distributions of a source domain and a target domain that differ acoustically and in terms of the amount of training data provided. A decision threshold that is optimal for one domain may be highly sub-optimal for the other domain and vice versa. This significantly degrades the performance when only using a single decision threshold, as is required when generalizing to multiple data domains that are possibly unseen during training while still using the same trained ASD system as in the source domain. To reduce this mismatch between the domains, we propose a simple local-density-based anomaly score normalization scheme. In experiments conducted on several ASD datasets, we show that the proposed normalization scheme consistently improves performance for various types of embedding-based ASD systems and yields better results than existing anomaly score normalization approaches.
【20】Sound Matching an Analogue Levelling Amplifier Using the Newton-Raphson Method
标题:使用Newton-Raphson方法对模拟均衡放大器进行声音匹配
链接:https://arxiv.org/abs/2509.10706
备注:Published at 2025 AES International Conference on Artificial Intelligence and Machine Learning for Audio (https://aes2.org/publications/elibrary-page/?id=22991)
摘要:通过数字信号处理算法进行虚拟模拟建模的自动微分最近得到了普及。这些算法通常比依赖于密集矩阵乘法的黑盒神经网络在计算上更有效。由于它们的可微性,它们可以与神经网络集成,并使用梯度下降算法进行联合训练,从而产生更有效的系统。此外,信号处理算法具有比神经网络少得多的参数,允许应用牛顿-拉夫森方法。该方法比梯度下降法具有更快的收敛速度和更强的鲁棒性,但代价是二次存储。本文提出了一种方法来模拟电平放大器使用前馈数字压缩器的参数优化,通过牛顿-拉夫逊法。我们证明,数字压缩机可以成功地近似我们的目标单位,Teletronix LA-2A的行为。不同的策略计算海森矩阵的基准。我们利用递归滤波器的并行算法在现代GPU上实现高效训练。生成的模型被制作成VST插件,并在https://github.com/aim-qmul/4a2a上开源。
摘要:Automatic differentiation through digital signal processing algorithms for virtual analogue modelling has recently gained popularity. These algorithms are typically more computationally efficient than black-box neural networks that rely on dense matrix multiplications. Due to their differentiable nature, they can be integrated with neural networks and jointly trained using gradient descent algorithms, resulting in more efficient systems. Furthermore, signal processing algorithms have significantly fewer parameters than neural networks, allowing the application of the Newton-Raphson method. This method offers faster and more robust convergence than gradient descent at the cost of quadratic storage. This paper presents a method to emulate analogue levelling amplifiers using a feed-forward digital compressor with parameters optimised via the Newton-Raphson method. We demonstrate that a digital compressor can successfully approximate the behaviour of our target unit, the Teletronix LA-2A. Different strategies for computing the Hessian matrix are benchmarked. We leverage parallel algorithms for recursive filters to achieve efficient training on modern GPUs. The resulting model is made into a VST plugin and is open-sourced at https://github.com/aim-qmul/4a2a.
【21】Spectral Bottleneck in Deep Neural Networks: Noise is All You Need
标题:深度神经网络中的频谱瓶颈:噪音就是你所需要的一切
链接:https://arxiv.org/abs/2509.09719
摘要:众所周知,深度神经网络表现出频谱学习偏差,其中低频分量在训练早期学习,而高频模式在后期逐渐出现。然而,当目标信号缺乏低频分量并且由宽带高频占主导地位时,训练会遭受“频谱检查”,并且模型无法重建整个信号,包括位于网络表示能力内的频率分量。我们研究这样的情况下,隐式神经表示(INR)与正弦表示网络(SIREN)的背景下,专注于拟合高频占主导地位的信号,容易受到频谱瓶颈的挑战。为了有效地适应任何目标信号,无论它的频率内容,我们提出了一个广义的目标感知的“权重扰动方案”(WINNER -权重初始化与噪声的神经表示)的网络初始化。该方案用高斯噪声扰动均匀初始化的权值,其中噪声尺度由目标信号的谱质心自适应确定。我们表明,噪声尺度可以提供控制网络激活的频谱和经验神经正切内核的本征基。这种方法不仅解决了频谱瓶颈,而且收敛速度更快,表示精度更高,在音频拟合方面优于最先进的方法,并在图像拟合和去噪任务中取得了显着的收益。除了信号重建之外,我们的方法还为计算机视觉和科学机器学习中的自适应权重初始化策略开辟了新的方向。
摘要:Deep neural networks are known to exhibit a spectral learning bias, wherein low-frequency components are learned early in training, while high-frequency modes emerge more gradually in later epochs. However, when the target signal lacks low-frequency components and is dominated by broadband high frequencies, training suffers from a 'spectral bottleneck', and the model fails to reconstruct the entire signal, including the frequency components that lie within the network's representational capacity. We examine such a scenario in the context of implicit neural representations (INRs) with sinusoidal representation networks (SIRENs), focusing on the challenge of fitting high-frequency-dominant signals that are susceptible to spectral bottleneck. To effectively fit any target signal irrespective of it's frequency content, we propose a generalized target-aware 'weight perturbation scheme' (WINNER - weight initialization with noise for neural representations) for network initialization. The scheme perturbs uniformly initialized weights with Gaussian noise, where the noise scales are adaptively determined by the spectral centroid of the target signal. We show that the noise scales can provide control over the spectra of network activations and the eigenbasis of the empirical neural tangent kernel. This method not only addresses the spectral bottleneck but also yields faster convergence and with improved representation accuracy, outperforming state-of-the-art approaches in audio fitting and achieving notable gains in image fitting and denoising tasks. Beyond signal reconstruction, our approach opens new directions for adaptive weight initialization strategies in computer vision and scientific machine learning.
【1】EEND-SAA: Enrollment-Less Main Speaker Voice Activity Detection Using Self-Attention Attractors
标题:EEND-SBA:使用自我注意力吸引器的少注册主要说话者语音活动检测
链接:https://arxiv.org/abs/2509.11957
摘要:语音活动检测(VAD)是基于语音的系统中必不可少的,但传统的方法只检测语音存在,而不识别说话人。目标说话人VAD(TS-VAD)通过使用短登记话语检测已知说话人的语音来扩展这一点,但是这种假设在开放域场景中失败,例如会议或客户服务呼叫,其中主要说话人是未知的。我们提出了EEND-SAA,一个注册少,流兼容的框架,主扬声器VAD,它确定了主扬声器没有先验知识。与TS-VAD不同,我们的方法基于语音连续性和音量将主要发言者确定为说话更稳定和清晰的人。我们在EEND上使用Transformer中的两个自注意吸引子构建模型,并应用因果掩蔽进行实时使用。多扬声器LibriSpeech混合的实验表明,EEND-SAA将主扬声器DER从6.63%降低到3.61%,并将F1从SA-EEND基线的0.9667提高到0.9818,在涉及扬声器重叠和噪声的条件下实现了最先进的性能。
摘要:Voice activity detection (VAD) is essential in speech-based systems, but traditional methods detect only speech presence without identifying speakers. Target-speaker VAD (TS-VAD) extends this by detecting the speech of a known speaker using a short enrollment utterance, but this assumption fails in open-domain scenarios such as meetings or customer service calls, where the main speaker is unknown. We propose EEND-SAA, an enrollment-less, streaming-compatible framework for main-speaker VAD, which identifies the primary speaker without prior knowledge. Unlike TS-VAD, our method determines the main speaker as the one who talks more steadily and clearly, based on speech continuity and volume. We build our model on EEND using two self-attention attractors in a Transformer and apply causal masking for real-time use. Experiments on multi-speaker LibriSpeech mixtures show that EEND-SAA reduces main-speaker DER from 6.63% to 3.61% and improves F1 from 0.9667 to 0.9818 over the SA-EEND baseline, achieving state-of-the-art performance under conditions involving speaker overlap and noise.
【2】Length-Aware Rotary Position Embedding for Text-Speech Alignment
标题:用于文本语音对齐的长度感知旋转位置嵌入
链接:https://arxiv.org/abs/2509.11084
备注:5 pages, 3 figures, preprint
摘要:许多最近的文本到语音(TTS)系统是建立在Transformer架构,并采用交叉注意机制的文本语音对齐。在这些系统中,旋转位置嵌入(RoPE)通常用于在文本和语音表示中编码位置信息。在这项工作中,我们引入了长度感知的RoPE(LARoPE),一个简单而有效的扩展RoPE,提高文本语音对齐。与依赖于绝对索引的RoPE不同,LARoPE使用长度规范化索引计算查询和关键位置之间的相对距离。实验结果表明,LARoPE始终优于RoPE,提供更快的损失收敛,更准确的文本语音对齐,更高的整体TTS质量。此外,LARoPE表现出更大的弹性话语持续时间的变化,并保持稳定的性能,在延长语音生成高达30秒,而RoPE遭受显着的退化。值得注意的是,我们的方法实现了最先进的字错误率的标准zero-shot TTS基准。
摘要:Many recent text-to-speech (TTS) systems are built on transformer architectures and employ cross-attention mechanisms for text-speech alignment. Within these systems, rotary position embedding (RoPE) is commonly used to encode positional information in text and speech representations. In this work, we introduce length-aware RoPE (LARoPE), a simple yet effective extension of RoPE that improves text-speech alignment. Unlike RoPE, which relies on absolute indices, LARoPE computes relative distances between query and key positions using length-normalized indices. Experimental results show that LARoPE consistently outperforms RoPE, offering faster loss convergence, more accurate text-speech alignment, and higher overall TTS quality. Furthermore, LARoPE demonstrates greater resilience to variations in utterance duration and maintains stable performance in extended speech generation up to 30 seconds, whereas RoPE suffers from notable degradation. Notably, our method achieves a state-of-the-art word error rate on a standard zero-shot TTS benchmark.
【3】Local Density-Based Anomaly Score Normalization for Domain Generalization
标题:基于局部密度的异常分数规范化领域概括
链接:https://arxiv.org/abs/2509.10951
摘要:域移位条件下的最先进的异常声音检测(ASD)系统依赖于将音频信号投影到嵌入空间中并使用基于距离的离群值检测来计算异常分数。要克服的主要困难之一是在声学上和在所提供的训练数据的量方面不同的源域和目标域的异常分数分布之间的所谓的域失配。对于一个域最优的决策阈值对于另一个域可能是高度次优的,反之亦然。当仅使用单个决策阈值时,这显著降低了性能,这是在泛化到在训练期间可能不可见的多个数据域同时仍然使用与源域中相同的经训练的ASD系统时所需要的。为了减少域之间的这种不匹配,我们提出了一个简单的基于局部密度的异常分数归一化方案。在几个ASD数据集上进行的实验中,我们表明,所提出的归一化方案始终提高了各种类型的基于嵌入的ASD系统的性能,并产生比现有的异常分数归一化方法更好的结果。
摘要:State-of-the-art anomalous sound detection (ASD) systems in domain-shifted conditions rely on projecting audio signals into an embedding space and using distance-based outlier detection to compute anomaly scores. One of the major difficulties to overcome is the so-called domain mismatch between the anomaly score distributions of a source domain and a target domain that differ acoustically and in terms of the amount of training data provided. A decision threshold that is optimal for one domain may be highly sub-optimal for the other domain and vice versa. This significantly degrades the performance when only using a single decision threshold, as is required when generalizing to multiple data domains that are possibly unseen during training while still using the same trained ASD system as in the source domain. To reduce this mismatch between the domains, we propose a simple local-density-based anomaly score normalization scheme. In experiments conducted on several ASD datasets, we show that the proposed normalization scheme consistently improves performance for various types of embedding-based ASD systems and yields better results than existing anomaly score normalization approaches.
【4】Sound Matching an Analogue Levelling Amplifier Using the Newton-Raphson Method
标题:使用Newton-Raphson方法对模拟均衡放大器进行声音匹配
链接:https://arxiv.org/abs/2509.10706
备注:Published at 2025 AES International Conference on Artificial Intelligence and Machine Learning for Audio (https://aes2.org/publications/elibrary-page/?id=22991)
摘要:通过数字信号处理算法进行虚拟模拟建模的自动微分最近得到了普及。这些算法通常比依赖于密集矩阵乘法的黑盒神经网络在计算上更有效。由于它们的可微性,它们可以与神经网络集成,并使用梯度下降算法进行联合训练,从而产生更有效的系统。此外,信号处理算法具有比神经网络少得多的参数,允许应用牛顿-拉夫森方法。该方法比梯度下降法具有更快的收敛速度和更强的鲁棒性,但代价是二次存储。本文提出了一种方法来模拟电平放大器使用前馈数字压缩器的参数优化,通过牛顿-拉夫逊法。我们证明,数字压缩机可以成功地近似我们的目标单位,Teletronix LA-2A的行为。不同的策略计算海森矩阵的基准。我们利用递归滤波器的并行算法在现代GPU上实现高效训练。生成的模型被制作成VST插件,并在https://github.com/aim-qmul/4a2a上开源。
摘要:Automatic differentiation through digital signal processing algorithms for virtual analogue modelling has recently gained popularity. These algorithms are typically more computationally efficient than black-box neural networks that rely on dense matrix multiplications. Due to their differentiable nature, they can be integrated with neural networks and jointly trained using gradient descent algorithms, resulting in more efficient systems. Furthermore, signal processing algorithms have significantly fewer parameters than neural networks, allowing the application of the Newton-Raphson method. This method offers faster and more robust convergence than gradient descent at the cost of quadratic storage. This paper presents a method to emulate analogue levelling amplifiers using a feed-forward digital compressor with parameters optimised via the Newton-Raphson method. We demonstrate that a digital compressor can successfully approximate the behaviour of our target unit, the Teletronix LA-2A. Different strategies for computing the Hessian matrix are benchmarked. We leverage parallel algorithms for recursive filters to achieve efficient training on modern GPUs. The resulting model is made into a VST plugin and is open-sourced at https://github.com/aim-qmul/4a2a.
【5】Room acoustics affect communicative success in hybrid meeting spaces: a pilot study
标题:房间声学影响混合会议空间中的沟通成功:一项试点研究
链接:https://arxiv.org/abs/2509.11709
摘要:自2020年COVID-19疫情以来,大学和公司越来越多地将混合功能融入其会议空间,甚至为此目的创建专用房间。虽然快速稳定的互联网连接的重要性往往被优先考虑,但会议室的声学设计经常被忽视。声学效果不佳,尤其是过度的混响,会导致误解、语音清晰度降低或认知和声音疲劳等问题。这项试点研究调查是否在格拉茨科技大学的研讨室的房间声学干预支持更好的沟通混合会议。为此,我们对两组人进行了两次录音,一次是在改善房间声学效果之前,一次是在改善房间声学效果之后。我们的研究结果-尽管没有达到统计学意义,由于样本量小-清楚地表明,我们的空间干预提高混合会议的沟通成功。为了使读者也可以从语音通信社区的文件,我们解释房间声学背景,相关的解释我们的结果。
摘要:Since the COVID-19 pandemic in 2020, universities and companies have increasingly integrated hybrid features into their meeting spaces, or even created dedicated rooms for this purpose. While the importance of a fast and stable internet connection is often prioritized, the acoustic design of seminar rooms is frequently overlooked. Poor acoustics, particularly excessive reverberation, can lead to issues such as misunderstandings, reduced speech intelligibility or cognitive and vocal fatigue. This pilot study investigates whether room acoustic interventions in a seminar room at Graz University of Technology support better communication in hybrid meetings. For this purpose, we recorded two groups of persons twice, once before and once after improving the acoustics of the room. Our findings -- despite not reaching statistical significance due to the small sample size - indicate clearly that our spatial interventions improve communicative success in hybrid meetings. To make the paper accessible also for readers from the speech communication community, we explain room acoustics background, relevant for the interpretation of our results.
【6】FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
标题:FuseCodec:神经编解码器的语义上下文融合和监督
链接:https://arxiv.org/abs/2509.11425
摘要:语音标记化使离散表示和促进语音语言建模。然而,现有的神经编解码器捕获低级别的声学特征,忽略了人类语音固有的语义和上下文线索。虽然最近的努力引入了来自自监督语音模型的语义表示或结合了来自预训练语言模型的上下文表示,但在对齐和统一语义和上下文表示方面仍然存在挑战。我们引入了FuseCodec,它通过强大的跨模态对齐和全局信息监督来统一声学,语义和上下文表示。我们提出了三种互补技术:(i)潜在表示融合,将语义和上下文特征直接集成到编码器潜在空间中,以实现鲁棒和统一的表示学习;(ii)全局语义上下文监督,用全局池化和广播表示监督离散令牌,以增强时间一致性和跨模态对齐;以及(iii)时间对齐的上下文监督,通过在本地窗口内动态匹配上下文和语音令牌来加强对齐,以进行细粒度的令牌级监督。我们进一步介绍FuseCodec-TTS,证明我们的方法的适用性zero-shot语音合成。从经验上看,FuseCodec在LibriSpeech中实现了最先进的性能,在转录准确性、感知质量、可理解性和扬声器相似性方面超过了EnCodec、SpeechTokenizer和DAC。结果突出了上下文和语义引导的语音标记和下游任务的标记化的有效性。代码和预训练模型可在https://github.com/mubtasimahasan/FuseCodec上获得。
摘要:Speech tokenization enables discrete representation and facilitates speech language modeling. However, existing neural codecs capture low-level acoustic features, overlooking the semantic and contextual cues inherent to human speech. While recent efforts introduced semantic representations from self-supervised speech models or incorporated contextual representations from pre-trained language models, challenges remain in aligning and unifying the semantic and contextual representations. We introduce FuseCodec, which unifies acoustic, semantic, and contextual representations through strong cross-modal alignment and globally informed supervision. We propose three complementary techniques: (i) Latent Representation Fusion, integrating semantic and contextual features directly into the encoder latent space for robust and unified representation learning; (ii) Global Semantic-Contextual Supervision, supervising discrete tokens with globally pooled and broadcasted representations to enhance temporal consistency and cross-modal alignment; and (iii) Temporally Aligned Contextual Supervision, strengthening alignment by dynamically matching contextual and speech tokens within a local window for fine-grained token-level supervision. We further introduce FuseCodec-TTS, demonstrating our methodology's applicability to zero-shot speech synthesis. Empirically, FuseCodec achieves state-of-the-art performance in LibriSpeech, surpassing EnCodec, SpeechTokenizer, and DAC in transcription accuracy, perceptual quality, intelligibility, and speaker similarity. Results highlight the effectiveness of contextually and semantically guided tokenization for speech tokenization and downstream tasks. Code and pretrained models are available at https://github.com/mubtasimahasan/FuseCodec.
【7】Revisiting Meter Tracking in Carnatic Music using Deep Learning Approaches
标题:使用深度学习方法重新审视狂欢音乐中的电表跟踪
链接:https://arxiv.org/abs/2509.11241
摘要:节拍和强拍跟踪,统称为节拍跟踪,是音乐信息检索(MIR)中的一项基本任务。在这一领域,深度学习模型已经远远超过了传统的信号处理和经典的机器学习方法,特别是对于西方(欧洲遗传学)流派来说,大型注释数据集广泛存在。然而,这些系统在代表性不足的音乐传统上表现得不太可靠。卡纳提克音乐是印度次大陆的一种丰富的传统音乐,以其复杂的节奏和独特的韵律结构而闻名。在此背景下,关于仪表跟踪的最值得注意的先前工作采用了概率动态贝叶斯网络(DBN)。然而,最先进的(SOTA)深度学习模型在Carnatic音乐上的性能在很大程度上仍未得到探索。 在这项研究中,我们评估了两个模型的节拍跟踪在卡纳蒂克音乐:时间卷积网络(TCN),一个轻量级的架构,已成功地适应拉丁节奏,和击败这个!,一个基于transformer的模型,设计用于广泛的风格覆盖,而不需要后期处理。在Carnatic音乐节奏(CMR$_f$)数据集上复制DBN基线的实验设置,我们系统地评估了这些模型在直接可比设置中的性能。我们进一步研究了适应策略,包括微调Carnatic数据模型和使用音乐信息参数。结果表明,虽然现成的模型并不总是优于DBN,但通过迁移学习,它们的性能大幅提高,匹配或超过基线。这些发现表明,SOTA深度学习模型可以有效地适应代表性不足的传统,为更具包容性和广泛适用的仪表跟踪系统铺平道路。
摘要:Beat and downbeat tracking, jointly referred to as Meter Tracking, is a fundamental task in Music Information Retrieval (MIR). Deep learning models have far surpassed traditional signal processing and classical machine learning approaches in this domain, particularly for Western (Eurogenetic) genres, where large annotated datasets are widely available. These systems, however, perform less reliably on underrepresented musical traditions. Carnatic music, a rich tradition from the Indian subcontinent, is renowned for its rhythmic intricacy and unique metrical structures (t\=alas). The most notable prior work on meter tracking in this context employed probabilistic Dynamic Bayesian Networks (DBNs). The performance of state-of-the-art (SOTA) deep learning models on Carnatic music, however, remains largely unexplored. In this study, we evaluate two models for meter tracking in Carnatic music: the Temporal Convolutional Network (TCN), a lightweight architecture that has been successfully adapted for Latin rhythms, and Beat This!, a transformer-based model designed for broad stylistic coverage without the need for post-processing. Replicating the experimental setup of the DBN baseline on the Carnatic Music Rhythm (CMR$_f$) dataset, we systematically assess the performance of these models in a directly comparable setting. We further investigate adaptation strategies, including fine-tuning the models on Carnatic data and the use of musically informed parameters. Results show that while off-the-shelf models do not always outperform the DBN, their performance improves substantially with transfer learning, matching or surpassing the baseline. These findings indicate that SOTA deep learning models can be effectively adapted to underrepresented traditions, paving the way for more inclusive and broadly applicable meter tracking systems.
【8】WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation
标题:WeaveMuse:一个用于多模式音乐理解和生成的开放式统计系统
链接:https://arxiv.org/abs/2509.11183
备注:Accepted at Large Language Models for Music & Audio Workshop (LLM4MA) 2025
摘要:人工智能已经在工业中标准化,作为协调专业模型和工具以解决复杂多模式任务的实用范例。在这项工作中,我们提出了WeaveMuse,一个多智能体系统的音乐理解,符号组成和音频合成。每个专家代理解释用户请求,导出机器可操作的要求(模态,格式,约束),并验证自己的输出,而管理员代理选择和排序工具,调解用户交互,并保持跨回合的状态。该系统可以在本地扩展和部署,使用量化和推理策略来适应不同的硬件预算,或者通过HFApi来保留对开放模型的免费社区访问。除了开箱即用之外,该系统还通过约束模式、结构化解码、基于策略的推理和参数高效适配器或为MIR任务定制模型的提炼变体来强调可控性和适应性。一个中心的设计目标是促进跨文本,符号符号和可视化,和音频,使分析-合成-渲染循环和解决跨格式的约束多式联运的互动。该框架旨在通过支持各种大小的可互换开源模型、灵活的内存管理和可重现的部署路径,使MIR工具民主化、实现和可访问。
摘要:Agentic AI has been standardized in industry as a practical paradigm for coordinating specialized models and tools to solve complex multimodal tasks. In this work, we present WeaveMuse, a multi-agent system for music understanding, symbolic composition, and audio synthesis. Each specialist agent interprets user requests, derives machine-actionable requirements (modalities, formats, constraints), and validates its own outputs, while a manager agent selects and sequences tools, mediates user interaction, and maintains state across turns. The system is extendable and deployable either locally, using quantization and inference strategies to fit diverse hardware budgets, or via the HFApi to preserve free community access to open models. Beyond out-of-the-box use, the system emphasizes controllability and adaptation through constraint schemas, structured decoding, policy-based inference, and parameter-efficient adapters or distilled variants that tailor models to MIR tasks. A central design goal is to facilitate intermodal interaction across text, symbolic notation and visualization, and audio, enabling analysis-synthesis-render loops and addressing cross-format constraints. The framework aims to democratize, implement, and make accessible MIR tools by supporting interchangeable open-source models of various sizes, flexible memory management, and reproducible deployment paths.
【9】Combining Audio and Non-Audio Inputs in Evolved Neural Networks for Ovenbird
标题:在Ovenbird的进化神经网络中结合音频和非音频输入
链接:https://arxiv.org/abs/2509.10566
摘要:在过去的几年里,神经网络作为从数字数据中自动进行物种分类的工具的使用有所增加。这部分是由于通过卷积神经网络(CNN)的图像分类的高分类精度。在音频数据的情况下,基于CNN的识别器用于通过使用来自声音可视化的信息(即,光谱图)。这些识别器通常使用声谱图作为其唯一输入。然而,研究人员有其他非音频数据,如物种的栖息地偏好,物候和范围信息,可以改善物种分类。在本文中,我们提出了如何单物种识别神经网络的准确性可以提高使用非音频数据作为输入,除了频谱图信息。我们还分析了这些改进是否仅仅是具有更多参数的神经网络的结果,而不是结合两个输入。我们发现,使用两种不同输入的网络比仅使用其中一种输入的类似规模的网络具有更高的分类准确性。
摘要:In the last several years the use of neural networks as tools to automate species classification from digital data has increased. This has been due in part to the high classification accuracy of image classification through Convolutional Neural Networks (CNN). In the case of audio data CNN based recognizers are used to automate the classification of species in audio recordings by using information from sound visualization (i.e., spectrograms). It is common for these recognizers to use the spectrogram as their sole input. However, researchers have other non-audio data, such as habitat preferences of a species, phenology, and range information, available that could improve species classification. In this paper we present how a single-species recognizer neural network's accuracy can be improved by using non-audio data as inputs in addition to spectrogram information. We also analyze if the improvements are merely a result of having a neural network with a higher number of parameters instead of combining the two inputs. We find that networks that use the two different inputs have a higher classification accuracy than networks of similar size that use only one of the inputs.
【10】Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
标题:用于ATCO命令任务组建模和任务组预测的多模式深度学习
链接:https://arxiv.org/abs/2509.10522
摘要:空中交通管制员(ATCO)在密集的空域中发出高强度的语音命令,准确的工作负载建模对于安全和效率至关重要。本文提出了一种多模态深度学习框架,该框架集成了结构化数据、轨迹序列和图像特征,以估计ATCO命令生命周期中的两个关键参数:命令与飞机机动之间的时间偏移,以及命令持续时间。构建了高质量的数据集,并使用滑动窗口和基于直方图的方法检测机动点。CNN-Transformer系综模型用于准确、可推广和可解释的预测。通过将轨迹与语音命令联系起来,这项工作提供了第一个支持智能命令生成的模型,并为工作量评估、人员配备和调度提供了实用价值。
摘要:Air traffic controllers (ATCOs) issue high-intensity voice commands in dense airspace, where accurate workload modeling is critical for safety and efficiency. This paper proposes a multimodal deep learning framework that integrates structured data, trajectory sequences, and image features to estimate two key parameters in the ATCO command lifecycle: the time offset between a command and the resulting aircraft maneuver, and the command duration. A high-quality dataset was constructed, with maneuver points detected using sliding window and histogram-based methods. A CNN-Transformer ensemble model was developed for accurate, generalizable, and interpretable predictions. By linking trajectories to voice commands, this work offers the first model of its kind to support intelligent command generation and provides practical value for workload assessment, staffing, and scheduling.
【11】Evaluating Automatic Speech Recognition Systems for Korean Meteorological Experts
标题:韩国气象专家评估自动语音识别系统
链接:https://arxiv.org/abs/2410.18444
备注:EMNLP 2025 Findings
摘要:本文探讨了将自动语音识别(ASR)集成到自然语言查询系统中,以提高韩国气象学家的天气预报效率。我们解决了为韩国天气领域开发ASR系统的挑战,特别是专业词汇和韩国语言的复杂性。为了解决这些问题,我们构建了一个评估数据集的口语查询记录的母语为韩国人。使用该数据集,我们评估了多语言ASR模型系列的各种配置,确定了与特定领域术语相关的性能限制。然后,我们实现了一种简单的基于文本到语音的数据增强方法,该方法提高了专业术语的识别能力,同时保持了通用领域的性能。我们的贡献包括创建特定领域的数据集,全面的ASR模型评估和有效的增强技术。我们相信,我们的工作为韩国天气预报领域ASR的未来发展奠定了基础。
摘要:This paper explores integrating Automatic Speech Recognition (ASR) into natural language query systems to improve weather forecasting efficiency for Korean meteorologists. We address challenges in developing ASR systems for the Korean weather domain, specifically specialized vocabulary and Korean linguistic intricacies. To tackle these issues, we constructed an evaluation dataset of spoken queries recorded by native Korean speakers. Using this dataset, we assessed various configurations of a multilingual ASR model family, identifying performance limitations related to domain-specific terminology. We then implemented a simple text-to-speech-based data augmentation method, which improved the recognition of specialized terms while maintaining general-domain performance. Our contributions include creating a domain-specific dataset, comprehensive ASR model evaluations, and an effective augmentation technique. We believe our work provides a foundation for future advancements in ASR for the Korean weather forecasting domain.
机器翻译由腾讯交互翻译提供,仅供参考
