今日论文合集:cs.SD语音12篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Parallel Delayed Memory Units for Enhanced Temporal Modeling in Biomedical and Bioacoustic Signal Analysis
标题:用于生物医学和生物声学信号分析中增强时间建模的并行延迟存储单元
链接:https://arxiv.org/abs/2512.01626

作者:Pengfei Sun,Wenyu Jiang,Paul Devos,Dick Botteldooren
备注:Accepted for publication in IEEE Transactions on Audio, Speech and Language Processing, 2025
摘要:先进的深度学习架构,特别是递归神经网络(RNN),已广泛应用于音频、生物声学和生物医学信号分析,特别是在数据稀缺的环境中。虽然门控RNN仍然有效,但它们在某些情况下可能相对过度参数化并且训练效率较低,而线性RNN往往无法捕捉生物信号中固有的复杂性。为了解决这些挑战,我们提出了并行延迟存储器单元(PDMU),一个{延迟门控状态空间模块短期时间信用分配}针对音频和生物声信号,增强短期时间状态的相互作用和内存效率通过门控延迟线机制。与先前将时间动态嵌入延迟线架构的延迟存储单元(DMU)不同,PDMU进一步使用勒让德存储单元(LMU)将时间信息压缩成向量表示。这种设计作为因果注意的一种形式,允许模型动态调整其对过去状态的依赖,并提高实时学习性能。值得注意的是,在低信息场景中,门控机制的行为类似于通过绕过状态衰减和保留早期表示来跳过连接,从而促进长期记忆保持。PDMU是模块化的,支持并行训练和顺序推理,并且可以轻松集成到现有的线性RNN框架中。此外,我们还引入了该架构的双向、高效和尖峰变体,每种变体都能在性能或能效方面提供额外的收益。不同的音频和生物医学基准的实验结果表明,PDMU显着提高了内存容量和整体模型的性能。
摘要:Advanced deep learning architectures, particularly recurrent neural networks (RNNs), have been widely applied in audio, bioacoustic, and biomedical signal analysis, especially in data-scarce environments. While gated RNNs remain effective, they can be relatively over-parameterised and less training-efficient in some regimes, while linear RNNs tend to fall short in capturing the complexity inherent in bio-signals. To address these challenges, we propose the Parallel Delayed Memory Unit (PDMU), a {delay-gated state-space module for short-term temporal credit assignment} targeting audio and bioacoustic signals, which enhances short-term temporal state interactions and memory efficiency via a gated delay-line mechanism. Unlike previous Delayed Memory Units (DMU) that embed temporal dynamics into the delay-line architecture, the PDMU further compresses temporal information into vector representations using Legendre Memory Units (LMU). This design serves as a form of causal attention, allowing the model to dynamically adjust its reliance on past states and improve real-time learning performance. Notably, in low-information scenarios, the gating mechanism behaves similarly to skip connections by bypassing state decay and preserving early representations, thereby facilitating long-term memory retention. The PDMU is modular, supporting parallel training and sequential inference, and can be easily integrated into existing linear RNN frameworks. Furthermore, we introduce bidirectional, efficient, and spiking variants of the architecture, each offering additional gains in performance or energy efficiency. Experimental results on diverse audio and biomedical benchmarks demonstrate that the PDMU significantly enhances both memory capacity and overall model performance.


【2】LLM2Fx-Tools: Tool Calling For Music Post-Production
标题:LLM 2FX-Tools:音乐后期制作的工具
链接:https://arxiv.org/abs/2512.01559

作者:Seungheon Doh,Junghyun Koo,Marco A. Martínez-Ramírez,Woosung Choi,Wei-Hsiang Liao,Qiyu Wu,Juhan Nam,Yuki Mitsufuji
摘要:本文介绍了LLM 2Fx-Tools,一个多模态的工具调用框架,生成可执行序列的音频效果(FX链)的音乐后期制作。LLM 2Fx-Tools使用大型语言模型(LLM)来理解音频输入,选择音频效果类型,确定它们的顺序,并在思想链(CoT)规划的指导下估计参数。我们还提出了LP-Fx,一个新的结构化CoT注释和音频效果模块的工具调用遵循的数据集。实验表明,LLM 2Fx-Tools可以通过自回归序列建模、工具调用和CoT推理,从未处理和已处理的音频对中推断Fx链及其参数。我们进一步验证了系统的风格转移设置,其中音频效果信息从参考源转移并应用于新的内容。最后,法学硕士作为一个法官的评价表明,我们的方法产生适当的CoT推理和音乐制作查询的反应。据我们所知,这是第一个将基于LLM的工具调用应用于音频效果模块的工作,从而实现可解释和可控的音乐制作。
摘要:This paper introduces LLM2Fx-Tools, a multimodal tool-calling framework that generates executable sequences of audio effects (Fx-chain) for music post-production. LLM2Fx-Tools uses a large language model (LLM) to understand audio inputs, select audio effects types, determine their order, and estimate parameters, guided by chain-of-thought (CoT) planning. We also present LP-Fx, a new instruction-following dataset with structured CoT annotations and tool calls for audio effects modules. Experiments show that LLM2Fx-Tools can infer an Fx-chain and its parameters from pairs of unprocessed and processed audio, enabled by autoregressive sequence modeling, tool calling, and CoT reasoning. We further validate the system in a style transfer setting, where audio effects information is transferred from a reference source and applied to new content. Finally, LLM-as-a-judge evaluation demonstrates that our approach generates appropriate CoT reasoning and responses for music production queries. To our knowledge, this is the first work to apply LLM-based tool calling to audio effects modules, enabling interpretable and controllable music production.


【3】Q2D2: A Geometry-Aware Audio Codec Leveraging Two-Dimensional Quantization
标题:Q2 D2:利用二维量化的几何感知音频编解码器
链接:https://arxiv.org/abs/2512.01537

作者:Tal Shuster,Eliya Nachmani
摘要:最近的神经音频编解码器已经实现了令人印象深刻的重建质量,通常依赖于诸如残差矢量量化(RVQ)、矢量量化(VQ)和有限标量量化(FSQ)之类的量化方法。然而,这些量化技术限制了潜在空间的几何结构,使得难以捕获特征之间的相关性,从而导致表示学习、码本利用率和令牌速率的低效率。在本文中,我们介绍了二维量化(Q2D2),一种量化方案,其中特征对被投影到结构化的二维网格,如六边形,菱形或矩形平铺和量化到最近的网格值,产生一个隐式码本定义的产品的网格级别,与传统的方法相比,码本的大小。尽管其几何公式简单,但Q2 D2提高了音频压缩效率,具有低令牌速率和高码本利用率,同时保持最先进的重建质量。具体而言,Q2D2在各种客观和主观重建度量方面实现了具有竞争力的卓越性能,与最先进的模型相比,在语音域中进行了广泛的实验。全面的消融研究进一步证实了我们设计选择的有效性。
摘要:Recent neural audio codecs have achieved impressive reconstruction quality, typically relying on quantization methods such as Residual Vector Quantization (RVQ), Vector Quantization (VQ) and Finite Scalar Quantization (FSQ). However, these quantization techniques limit the geometric structure of the latent space, make it harder to capture correlations between features leading to inefficiency in representation learning, codebook utilization and token rate. In this paper we introduce Two Dimensional Quantization (Q2D2), a quantization scheme in which feature pairs are projected onto structured 2D grids such as hexagonal, rhombic, or rectangular tiling and quantized to the nearest grid values, yielding an implicit codebook defined by the product of grid levels, with codebook sizes comparable to conventional methods. Despite its simple geometric formulation, Q2D2 improves audio compression efficiency, with low token rates and high codebook utilization while maintaining state of the art reconstruction quality. Specifically, Q2D2 achieves competitive to superior performance in various objective and subjective reconstruction metrics, across extensive experiments in speech domain compared to state of the art models. Comprehensive ablation studies further confirm the effectiveness of our design choices.


【4】MEGConformer: Conformer-Based MEG Decoder for Robust Speech and Phoneme Classification
标题:MEGConformer:基于Conformer的MEG解码器,用于稳健的语音和音素分类
链接:https://arxiv.org/abs/2512.01443

作者:Xabier de Zuazo,Ibon Saratxaga,Eva Navas
备注:10 pages, 5 figures, 4 tables, LibriBrain Workshop, NeurIPS 2025
摘要:我们为LibriBrain 2025 PNPL竞赛提供了基于Conformer的解码器,目标是两个基本的MEG任务:语音检测和音素分类。我们的方法将紧凑型Conformer适配于原始306通道MEG信号,具有轻量级卷积投影层和特定任务的头。对于语音检测,面向MEG的SpecAugment提供了对MEG特定增强的第一次探索。对于音素分类,我们使用反平方根类加权和动态分组加载器来处理100个样本平均的示例。此外,一个简单的实例级规范化被证明是至关重要的,以减轻分布变化的坚持分裂。使用官方的标准音轨分割和F1宏进行模型选择,我们最好的系统在排行榜上获得了88.9%(语音)和65.8%(音素),超过了竞争基线,并在两项任务中排名前10。有关进一步的实现细节,请访问https://github.com/neural2speech/libribrain-experiments获取技术文档、源代码和检查点。
摘要:We present Conformer-based decoders for the LibriBrain 2025 PNPL competition, targeting two foundational MEG tasks: Speech Detection and Phoneme Classification. Our approach adapts a compact Conformer to raw 306-channel MEG signals, with a lightweight convolutional projection layer and task-specific heads. For Speech Detection, a MEG-oriented SpecAugment provided a first exploration of MEG-specific augmentation. For Phoneme Classification, we used inverse-square-root class weighting and a dynamic grouping loader to handle 100-sample averaged examples. In addition, a simple instance-level normalization proved critical to mitigate distribution shifts on the holdout split. Using the official Standard track splits and F1-macro for model selection, our best systems achieved 88.9% (Speech) and 65.8% (Phoneme) on the leaderboard, surpassing the competition baselines and ranking within the top-10 in both tasks. For further implementation details, the technical documentation, source code, and checkpoints are available at https://github.com/neural2speech/libribrain-experiments.


【5】ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
标题:ZO-ASB:无需反向传播的语音基础模型零阶微调
链接:https://arxiv.org/abs/2512.01267

作者:Yuezhang Peng,Yuxin Liu,Yao Li,Sheng Wang,Fei Wen,Xie Chen
备注:2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
摘要:微调预训练的语音基础模型用于自动语音识别(ASR)是普遍的,但受到大量GPU内存需求的限制。我们介绍了ZO-ASR,一种记忆高效的零阶(ZO)方法,通过向前传递估计梯度来避免反向传播(BP)和激活记忆。当与SGD优化器结合使用时,ZO-ASR-SGD仅使用推理内存来微调ASR模型。我们的评估涵盖监督和无监督任务。对于Whisper-Large-V3上的监督域自适应,ZO-ASR的多查询机制增强了鲁棒性,并在zero-shot基线上实现了高达18.9\ %的相对字错误率降低,优于现有的ZO方法。对于Wav2Vec2-Base上的无监督测试时自适应,ZO-ASR与一阶优化器Adam相比表现出较低的性能。我们的BP自由的方法提供了一个可行的解决方案,在计算资源受限或梯度不可访问的情况下微调ASR模型。
摘要:Fine-tuning pre-trained speech foundation models for Automatic Speech Recognition (ASR) is prevalent, yet constrained by substantial GPU memory requirements. We introduce ZO-ASR, a memory-efficient Zeroth-Order (ZO) method that avoids Back-Propagation (BP) and activation memory by estimating gradients via forward passes. When combined with SGD optimizer, ZO-ASR-SGD fine-tunes ASR models using only inference memory. Our evaluation spans supervised and unsupervised tasks. For Supervised Domain Adaptation on Whisper-Large-V3, ZO-ASR's multiple query mechanism enhances robustness and achieves up to an 18.9\% relative Word Error Rate reduction over zero-shot baselines, outperforming existing ZO methods. For unsupervised Test-Time Adaptation on Wav2Vec2-Base, ZO-ASR exhibits moderately lower performance compared to first-order optimizer Adam. Our BP-free approach provides a viable solution for fine-tuning ASR models in computationally resource-constrained or gradient-inaccessible scenarios.


【6】Audio-Visual World Models: Towards Multisensory Imagination in Sight and Sound
标题:视听世界模型:走向视觉和声音的多感官想象
链接:https://arxiv.org/abs/2512.00883

作者:Jiahua Wang,Shannan Yan,Leqi Zheng,Jialong Wu,Yaoxin Mao
摘要:世界模型模拟环境动态,使智能体能够计划和推理未来的状态。虽然现有的方法主要集中在视觉观察,但现实世界的感知本质上涉及多种感官形式。音频提供了关键的空间和时间线索,如声源定位和声学场景属性,但其融入世界模型仍然在很大程度上未被探索。没有以前的工作已经正式定义了什么构成了一个视听世界模型,或如何联合捕捉双耳空间音频和视觉动态下精确的动作控制与任务奖励预测。这项工作提出了视听世界模型(AVWM)的第一个正式框架,将多模式环境模拟制定为一个部分可观察的马尔可夫决策过程,具有同步的视听观察、细粒度的动作和任务奖励。为了解决缺乏合适的训练数据的问题,我们构建了AVW-4k,这是一个包含30小时双耳视听轨迹的数据集,其中包含76个室内环境中的动作注释和奖励信号。我们提出了AV-CDiT,视听条件扩散Transformer与新的模态专家架构,平衡视觉和听觉学习,优化通过三阶段的训练策略,有效的多模态整合。大量的实验表明,AV-CDiT实现了高保真的多模态预测跨视觉和听觉模式与奖励。此外,我们验证了它的实际效用在连续的视听导航任务,AVWM显着提高代理的性能。
摘要:World models simulate environmental dynamics to enable agents to plan and reason about future states. While existing approaches have primarily focused on visual observations, real-world perception inherently involves multiple sensory modalities. Audio provides crucial spatial and temporal cues such as sound source localization and acoustic scene properties, yet its integration into world models remains largely unexplored. No prior work has formally defined what constitutes an audio-visual world model or how to jointly capture binaural spatial audio and visual dynamics under precise action control with task reward prediction. This work presents the first formal framework for Audio-Visual World Models (AVWM), formulating multimodal environment simulation as a partially observable Markov decision process with synchronized audio-visual observations, fine-grained actions, and task rewards. To address the lack of suitable training data, we construct AVW-4k, a dataset comprising 30 hours of binaural audio-visual trajectories with action annotations and reward signals across 76 indoor environments. We propose AV-CDiT, an Audio-Visual Conditional Diffusion Transformer with a novel modality expert architecture that balances visual and auditory learning, optimized through a three-stage training strategy for effective multimodal integration. Extensive experiments demonstrate that AV-CDiT achieves high-fidelity multimodal prediction across visual and auditory modalities with reward. Furthermore, we validate its practical utility in continuous audio-visual navigation tasks, where AVWM significantly enhances the agent's performance.


【7】Melody or Machine: Detecting Synthetic Music with Dual-Stream Contrastive Learning
标题:旋律还是机器:通过双流对比学习检测合成音乐
链接:https://arxiv.org/abs/2512.00621

作者:Arnesh Batra,Dev Sharma,Krish Thukral,Ruhani Bhatia,Naman Batra,Aditya Gautam
备注:Accepted at Transactions on Machine Learning Research (TMLR)
摘要:端到端人工智能音乐生成的快速发展对艺术真实性和版权构成了不断升级的威胁,需要能够跟上步伐的检测方法。虽然SpectTTTra等现有的基础模型在面对新生成器的多样化和快速发展的生态系统时会出现动摇,但在分发外(OOD)内容上表现出显着的性能下降。这种泛化失败突出了一个关键的差距:需要更具挑战性的基准和更强大的检测架构。为了解决这个问题,我们首先介绍Melody or Machine(MoM),这是一个新的大规模基准,包含超过130,000首歌曲(6,665小时)。MoM是迄今为止最多样化的数据集,由开放和闭源模型的混合构建,以及专门为促进真正可推广的检测器的开发而设计的精心策划的OOD测试集。除了这个基准,我们介绍CLAM,一种新的双流检测架构。我们假设,微妙的,机器引起的声乐和器乐元素之间的不一致,往往难以察觉的混合信号,提供了一个强有力的综合迹象。CLAM旨在通过采用两个不同的预训练音频编码器(MERT和Wave2Vec2)来创建音频的并行表示来测试这一假设。这些表示由一个可学习的交叉聚合模块融合,该模块对它们的相互依赖性进行建模。该模型使用双重损失目标进行训练:用于分类的标准二进制交叉熵损失,辅以对比性三重损失,该损失训练模型区分相干和人工不匹配的流配对,增强其对合成伪影的敏感性,而无需假设简单的特征对齐。CLAM在合成音乐取证方面建立了一个新的最先进的技术。它在我们具有挑战性的MoM基准测试中获得了0.925的F1分数。
摘要:The rapid evolution of end-to-end AI music generation poses an escalating threat to artistic authenticity and copyright, demanding detection methods that can keep pace. While foundational, existing models like SpecTTTra falter when faced with the diverse and rapidly advancing ecosystem of new generators, exhibiting significant performance drops on out-of-distribution (OOD) content. This generalization failure highlights a critical gap: the need for more challenging benchmarks and more robust detection architectures. To address this, we first introduce Melody or Machine (MoM), a new large-scale benchmark of over 130,000 songs (6,665 hours). MoM is the most diverse dataset to date, built with a mix of open and closed-source models and a curated OOD test set designed specifically to foster the development of truly generalizable detectors. Alongside this benchmark, we introduce CLAM, a novel dual-stream detection architecture. We hypothesize that subtle, machine-induced inconsistencies between vocal and instrumental elements, often imperceptible in a mixed signal, offer a powerful tell-tale sign of synthesis. CLAM is designed to test this hypothesis by employing two distinct pre-trained audio encoders (MERT and Wave2Vec2) to create parallel representations of the audio. These representations are fused by a learnable cross-aggregation module that models their inter-dependencies. The model is trained with a dual-loss objective: a standard binary cross-entropy loss for classification, complemented by a contrastive triplet loss which trains the model to distinguish between coherent and artificially mismatched stream pairings, enhancing its sensitivity to synthetic artifacts without presuming a simple feature alignment. CLAM establishes a new state-of-the-art in synthetic music forensics. It achieves an F1 score of 0.925 on our challenging MoM benchmark.


【8】Explainable Multi-Modal Deep Learning for Automatic Detection of Lung Diseases from Respiratory Audio Signals
标题:可解释的多模式深度学习用于从呼吸音频信号自动检测肺部疾病
链接:https://arxiv.org/abs/2512.00563

作者:S M Asiful Islam Saky,Md Rashidul Islam,Md Saiful Arefin,Shahaba Alam
摘要:呼吸系统疾病仍然是全球主要的健康挑战,传统的听诊往往受到主观性,环境噪声和临床医生之间的差异的限制。这项研究提出了一个可解释的多模态深度学习框架,用于使用呼吸音频信号自动检测肺部疾病。所提出的系统集成了两种互补的表示:基于CNN-BiLSTM Attention架构的频谱-时间编码器,以及捕获生理上有意义的描述符(如MFCC,频谱质心,频谱带宽和过零率)的手工声学特征编码器。这些分支通过后期融合结合在一起,以利用数据驱动的学习和领域信息的声学线索。该模型在哮喘检测数据集版本2上进行训练和评估,使用严格的预处理,包括恢复,归一化,噪声过滤,数据增强和患者级别分层分区。该研究实现了较强的泛化能力,准确率为91.21%,宏观F1评分为0.899,宏观ROC-AUC为0.9866,优于所有消融变体。消融研究证实了时间建模,注意力机制和多模态融合的重要性。该框架结合了Grad-CAM,集成的声学标记和SHAP,生成可解释的频谱,时间和特征级解释,与已知的声学生物标志物对齐,以建立临床透明度。研究结果证明了该框架在远程医疗、即时诊断和现实世界呼吸筛查方面的潜力。
摘要:Respiratory diseases remain major global health challenges, and traditional auscultation is often limited by subjectivity, environmental noise, and inter-clinician variability. This study presents an explainable multimodal deep learning framework for automatic lung-disease detection using respiratory audio signals. The proposed system integrates two complementary representations: a spectral-temporal encoder based on a CNN-BiLSTM Attention architecture, and a handcrafted acoustic-feature encoder capturing physiologically meaningful descriptors such as MFCCs, spectral centroid, spectral bandwidth, and zero-crossing rate. These branches are combined through late-stage fusion to leverage both data-driven learning and domain-informed acoustic cues. The model is trained and evaluated on the Asthma Detection Dataset Version 2 using rigorous preprocessing, including resampling, normalization, noise filtering, data augmentation, and patient-level stratified partitioning. The study achieved strong generalization with 91.21% accuracy, 0.899 macro F1-score, and 0.9866 macro ROC-AUC, outperforming all ablated variants. An ablation study confirms the importance of temporal modeling, attention mechanisms, and multimodal fusion. The framework incorporates Grad-CAM, Integrated Gradients, and SHAP, generating interpretable spectral, temporal, and feature-level explanations aligned with known acoustic biomarkers to build clinical transparency. The findings demonstrate the framework's potential for telemedicine, point-of-care diagnostics, and real-world respiratory screening.


【9】STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
标题:STCTS:通过显式文本-韵律-音色分解实现超低比特率语音的生成式语义压缩
链接:https://arxiv.org/abs/2512.00451

作者:Siyu Wang,Haitao Li
备注:The complete source code and online speech reconstruction demo is publicly available at https://github.com/dywsy21/STCTS
摘要:在带宽受限的环境中--海上、卫星和战术网络--语音通信仍然非常昂贵。传统的编解码器在低于1 kbps的情况下挣扎,而现有的语义方法(STT-TTS)牺牲了韵律和说话者身份。我们提出了STCTS,生成语义压缩框架,使自然语音通信约80 bps。STCTS显式地将语音分解为语言内容,韵律表达和扬声器音色,应用定制的压缩:上下文感知文本编码(约70 bps),通过TTS插值的稀疏韵律传输(在0.1-1 Hz时小于14 bps)和摊销扬声器嵌入。   对LibriSpeech的评估表明,与Opus(6 kbps)相比,比特率降低了75倍,与EnCodec(1 kbps)相比,比特率降低了12倍,同时保持了感知质量(NISQA MOS大于4.26)。我们还发现了一个双峰质量分布与韵律采样率:稀疏和密集的更新都实现了高质量,而中档率下降,由于感知的不连续性-指导最佳配置设计。除了效率之外,我们的模块化架构还支持隐私保护加密、人类可理解的传输以及在边缘设备上的灵活部署,为超低带宽场景提供了强大的解决方案。
摘要:Voice communication in bandwidth-constrained environments--maritime, satellite, and tactical networks--remains prohibitively expensive. Traditional codecs struggle below 1 kbps, while existing semantic approaches (STT-TTS) sacrifice prosody and speaker identity. We present STCTS, a generative semantic compression framework enabling natural voice communication at approximately 80 bps. STCTS explicitly decomposes speech into linguistic content, prosodic expression, and speaker timbre, applying tailored compression: context-aware text encoding (approximately 70 bps), sparse prosody transmission via TTS interpolation (less than 14 bps at 0.1-1 Hz), and amortized speaker embedding.   Evaluations on LibriSpeech demonstrate a 75x bitrate reduction versus Opus (6 kbps) and 12x versus EnCodec (1 kbps), while maintaining perceptual quality (NISQA MOS greater than 4.26). We also discover a bimodal quality distribution with prosody sampling rate: sparse and dense updates both achieve high quality, while mid-range rates degrade due to perceptual discontinuities--guiding optimal configuration design. Beyond efficiency, our modular architecture supports privacy-preserving encryption, human-interpretable transmission, and flexible deployment on edge devices, offering a robust solution for ultra-low bandwidth scenarios.


【10】Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment
标题:Art 2 Music:通过多模式感觉对齐为艺术图像生成音乐
链接:https://arxiv.org/abs/2512.00120

作者:Jiaying Hong,Ting Zhu,Thanet Markchom,Huizhi Liang
摘要:随着人工智能生成内容(AIGC)的兴起,从多模态输入中生成感知自然和情感一致的音乐已成为一个核心挑战。现有的方法通常依赖于需要昂贵注释的明确情感标签,强调需要更灵活的情感对齐方法。为了支持多模态音乐生成,我们构建了ArtiCaps,这是一个伪感觉对齐的图像-音乐-文本数据集,它是通过从ArtEmis和MusicCaps语义匹配描述创建的。我们进一步提出了Art 2 Music,一个轻量级的跨模态框架,从艺术图像和用户评论合成音乐。在第一阶段,图像和文本使用OpenCLIP编码,并使用门控残差模块进行融合;融合后的表示由双向LSTM解码为具有频率加权L1损失的Mel频谱图,以增强高频保真度。在第二阶段,微调的HiFi-GAN声码器重建高质量的音频波形。在ArtiCaps上的实验表明,梅尔倒谱系数失真,弗雷歇音频距离,对数频谱距离和余弦相似性有明显的改善。一个小型的基于LLM的评级研究进一步验证了一致的跨模态感觉对齐,并提供了跨模态匹配和不匹配的可解释的解释。这些结果证明了改进的感知自然度、光谱保真度和语义一致性。Art 2 Music还仅使用5万个训练样本就保持了强大的性能,为交互式艺术、个性化音景和数字艺术展览中的情感一致的创意音频生成提供了可扩展的解决方案。
摘要:With the rise of AI-generated content (AIGC), generating perceptually natural and feeling-aligned music from multimodal inputs has become a central challenge. Existing approaches often rely on explicit emotion labels that require costly annotation, underscoring the need for more flexible feeling-aligned methods. To support multimodal music generation, we construct ArtiCaps, a pseudo feeling-aligned image-music-text dataset created by semantically matching descriptions from ArtEmis and MusicCaps. We further propose Art2Music, a lightweight cross-modal framework that synthesizes music from artistic images and user comments. In the first stage, images and text are encoded with OpenCLIP and fused using a gated residual module; the fused representation is decoded by a bidirectional LSTM into Mel-spectrograms with a frequency-weighted L1 loss to enhance high-frequency fidelity. In the second stage, a fine-tuned HiFi-GAN vocoder reconstructs high-quality audio waveforms. Experiments on ArtiCaps show clear improvements in Mel-Cepstral Distortion, Frechet Audio Distance, Log-Spectral Distance, and cosine similarity. A small LLM-based rating study further verifies consistent cross-modal feeling alignment and offers interpretable explanations of matches and mismatches across modalities. These results demonstrate improved perceptual naturalness, spectral fidelity, and semantic consistency. Art2Music also maintains robust performance with only 50k training samples, providing a scalable solution for feeling-aligned creative audio generation in interactive art, personalized soundscapes, and digital art exhibitions.


【11】MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
标题:MoLT:分层代币混合,实现高效视听学习
链接:https://arxiv.org/abs/2512.00115

作者:Kyeongha Rho,Hyeongkeun Lee,Jae Won Cho,Joon Son Chung
备注:10 pages, 5 figures
摘要:在本文中,我们提出了分层令牌混合(MoLT),这是一种用于视听学习的参数和内存高效的自适应框架。MoLT的关键思想是在每个Transformer层用一个并行的、轻量级的方案来取代传统的、计算量大的顺序自适应,该方案仅从后期层提取和融合逐层令牌。我们采用了两种类型的适配器提取特定于模态的信息和跨模态的交互到紧凑的潜在令牌在一个逐层的方式。然后,令牌融合模块通过考虑它们的相对重要性来动态地融合这些逐层令牌。为了防止潜在标记的冗余,我们在训练期间在潜在标记之间应用正交正则化。通过系统地分析预训练的Transformers中自适应的位置,我们只从Transformers的后期层中提取潜在标记。这种策略性的自适应方法避免了来自易失性早期层特征的错误传播,从而在保持参数和存储器效率的同时最大化自适应性能。通过大量的实验,我们证明了MoLT在各种视听基准上优于现有方法,包括视听问题分类,视听分割和视听事件定位。
摘要:In this paper, we propose Mixture of Layer-Wise Tokens (MoLT), a parameter- and memory-efficient adaptation framework for audio-visual learning. The key idea of MoLT is to replace conventional, computationally heavy sequential adaptation at every transformer layer with a parallel, lightweight scheme that extracts and fuses layer-wise tokens only from the late layers. We adopt two types of adapters to distill modality-specific information and cross-modal interaction into compact latent tokens in a layer-wise manner. A token fusion module then dynamically fuses these layer-wise tokens by taking into account their relative significance. To prevent the redundancy of latent tokens, we apply an orthogonality regularization between latent tokens during training. Through the systematic analysis of the position of adaptation in the pre-trained transformers, we extract latent tokens only from the late layers of the transformers. This strategic adaptation approach avoids error propagation from the volatile early-layer features, thereby maximizing the adaptation performance while maintaining parameter and memory efficiency. Through extensive experiments, we demonstrate that MoLT outperforms existing methods on diverse audio-visual benchmarks, including Audio-Visual Question Answering, Audio-Visual Segmentation, and Audio-Visual Event Localization.


【12】Masked Symbol Modeling for Demodulation of Oversampled Baseband Communication Signals in Impulsive Noise-Dominated Channels
标题:脉冲噪音主导通道中过采样基带通信信号解调的掩蔽符号建模
链接:https://arxiv.org/abs/2512.01428

作者:Oguz Bedir,Nurullah Sevim,Mostafa Ibrahim,Sabit Ekin
备注:Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on AI and ML for Next-Generation Wireless Communications and Networking (AI4NextG), non-archival
摘要:自然语言处理领域的最新突破表明,通过屏蔽标记预测训练的Transformer网络中的注意力机制使模型能够捕获标记的语义上下文并内化语言的语法。虽然Transformers在通信系统中的应用是一个新兴的领域,但物理波形中的上下文概念仍未得到充分探索。本文通过重新研究脉冲整形重叠引起的符号间贡献(ISC)来解决这一差距。而不是把ISC作为一个滋扰,我们认为它是一个确定性的上下文信息源嵌入过采样复基带信号。我们提出了屏蔽符号建模(MSM),物理(PHY)层的框架启发双向编码器表示从Transformers方法。在MSM中,符号对齐样本的子集被随机掩蔽,并且Transformer使用周围的“中间”样本来预测丢失的符号标识符。通过这个目标,该模型学习复杂基带波形的潜在语法。我们说明MSM的潜力,通过将其应用到解调脉冲噪声破坏的信号的任务,其中模型推断损坏段利用学习的上下文。我们的研究结果表明,接收器的解释,而不仅仅是检测通信信号的路径,为上下文感知的物理层设计开辟了新的途径。
摘要:Recent breakthroughs in natural language processing show that attention mechanism in Transformer networks, trained via masked-token prediction, enables models to capture the semantic context of the tokens and internalize the grammar of language. While the application of Transformers to communication systems is a burgeoning field, the notion of context within physical waveforms remains under-explored. This paper addresses that gap by re-examining inter-symbol contribution (ISC) caused by pulse-shaping overlap. Rather than treating ISC as a nuisance, we view it as a deterministic source of contextual information embedded in oversampled complex baseband signals. We propose Masked Symbol Modeling (MSM), a framework for the physical (PHY) layer inspired by Bidirectional Encoder Representations from Transformers methodology. In MSM, a subset of symbol aligned samples is randomly masked, and a Transformer predicts the missing symbol identifiers using the surrounding "in-between" samples. Through this objective, the model learns the latent syntax of complex baseband waveforms. We illustrate MSM's potential by applying it to the task of demodulating signals corrupted by impulsive noise, where the model infers corrupted segments by leveraging the learned context. Our results suggest a path toward receivers that interpret, rather than merely detect communication signals, opening new avenues for context-aware PHY layer design.


eess.AS音频处理


【1】Identifiability Conditions for Acoustic Feedback Cancellation with the Two-Channel Adaptive Feedback Canceller Algorithm
标题:双通道自适应反馈抵消器算法声反馈抵消的可识别条件
链接:https://arxiv.org/abs/2512.01466

作者:Arnout Roebben,Toon van Waterschoot,Jan Wouters,Marc Moonen
备注:Accepted for publication in IEEE Open Journal of Signal Processing (OJSP)
摘要:在同一声学环境中具有麦克风和扬声器的音频信号处理应用中,扬声器信号可以反馈到麦克风中,从而创建可能导致系统不稳定的闭环系统。为了消除这种声学耦合,预测误差法(PEM)反馈消除算法旨在通过假设输入信号可以借助于自回归(AR)模型来建模来识别扬声器与麦克风之间的反馈路径。先前已经表明,该PEM框架和所得算法可以在从麦克风到扬声器的前向路径充分时变或非线性的情况下,或者当前向路径延迟等于或超过AR模型的阶数时,正确地识别反馈路径。在本文中,它表明,这种基于延迟的条件可以概括为一个特定的PEM为基础的算法,所谓的双通道自适应反馈消除器(2CH-AFC),一个可逆性为基础的条件,它表明,可识别性时,可以实现前向路径前馈滤波器的顺序超过了AR模型的顺序。此外,在2ch-AFC算法中使用的相关矩阵的求逆的条件数可以用作用于监测可识别性的测量。
摘要:In audio signal processing applications with a microphone and a loudspeaker within the same acoustic environment, the loudspeaker signals can feed back into the microphone, thereby creating a closed-loop system that potentially leads to system instability. To remove this acoustic coupling, prediction error method (PEM) feedback cancellation algorithms aim to identify the feedback path between the loudspeaker and the microphone by assuming that the input signal can be modelled by means of an autoregressive (AR) model. It has previously been shown that this PEM framework and resulting algorithms can identify the feedback path correctly in cases where the forward path from microphone to loudspeaker is sufficiently time-varying or non-linear, or when the forward path delay equals or exceeds the order of the AR model. In this paper, it is shown that this delay-based condition can be generalised for one particular PEM-based algorithm, the so-called two-channel adaptive feedback canceller (2ch-AFC), to an invertibility-based condition, for which it is shown that identifiability can be achieved when the order of the forward path feedforward filter exceeds the order of the AR model. Additionally, the condition number of inversion of the correlation matrix as used in the 2ch-AFC algorithm can serve as a measure for monitoring the identifiability.


【2】Arabic TTS with FastPitch: Reproducible Baselines, Adversarial Training, and Oversmoothing Analysis
标题:带FastPitch的阿拉伯语TTC:可重复基线、对抗训练和过度平滑分析
链接:https://arxiv.org/abs/2512.00937

作者:Lars Nippert
摘要:由于资源有限和复杂的语音模式,阿拉伯文的文语转换(TTS)仍然具有挑战性。我们提出了可重复的基线,阿拉伯文TTS的FastPitch架构的基础上,并引入倒谱域指标分析梅尔频谱预测的过平滑。虽然传统的Lp重建损失产生平滑但过平均的输出,但所提出的度量在整个训练过程中揭示了它们的时间和频谱效应。为了解决这个问题,我们引入了一个轻量级的对抗谱图损失,它可以稳定地训练,并大大减少过度平滑。我们进一步探索多扬声器阿拉伯文TTS增强FastPitch与使用XTTSv2生成的合成语音,从而提高韵律的多样性,而不会失去稳定性。代码、预训练模型和训练方法可在https://github.com/nipponjo/tts-arabic-pytorch上公开获取。
摘要:Arabic text-to-speech (TTS) remains challenging due to limited resources and complex phonological patterns. We present reproducible baselines for Arabic TTS built on the FastPitch architecture and introduce cepstral-domain metrics for analyzing oversmoothing in mel-spectrogram prediction. While traditional Lp reconstruction losses yield smooth but over-averaged outputs, the proposed metrics reveal their temporal and spectral effects throughout training. To address this, we incorporate a lightweight adversarial spectrogram loss, which trains stably and substantially reduces oversmoothing. We further explore multi-speaker Arabic TTS by augmenting FastPitch with synthetic voices generated using XTTSv2, resulting in improved prosodic diversity without loss of stability. The code, pretrained models, and training recipes are publicly available at: https://github.com/nipponjo/tts-arabic-pytorch.


【3】A Low-Complexity Speech Codec Using Parametric Dithering for ASR
标题:一种采用参数抖动的低复杂度语音编解码器
链接:https://arxiv.org/abs/2512.00511

作者:Ellison Murray,Morriel Kasher,Predrag Spasojevic
备注:10 pages, 8 figures, Accepted 2026 Data Compression Conference
摘要:抖动是一种通常用于改善有损数据压缩的感知质量的技术。在这项工作中,我们分析和实验证明使用抖动的ASR输入压缩。我们形式化的最佳ASR性能的有损输入压缩下的理解,并利用这一点,提出了一个低复杂度的语音压缩管道的参数抖动技术。该方法在1位分辨率下表现良好,表现出25%的相对CER改善,同时还表现出在2位和3位分辨率下分别改善了32.4%和33.5%,其中我们的第二抖动选择产生降低的数据速率。所提出的编解码器是自适应的,以满足性能目标或留在熵约束。
摘要:Dithering is a technique commonly used to improve the perceptual quality of lossy data compression. In this work, we analytically and experimentally justify the use of dithering for ASR input compression. We formalize an understanding of optimal ASR performance under lossy input compression and leverage this to propose a parametric dithering technique for a low-complexity speech compression pipeline. The method performs well at 1-bit resolution, showing a 25\% relative CER improvement, while also demonstrating improvements of 32.4\% and 33.5\% at 2- and 3-bit resolution, respectively, with our second dither choice yielding a reduced data rate. The proposed codec is adaptable to meet performance targets or stay within entropy constraints.


【4】Beyond Performance: Probing Representation Dynamics In Speech Enhancement Models
标题:超越性能:探索语音增强模型中的表示动态
链接:https://arxiv.org/abs/2512.00482

作者:Yair Amar,Amir Ivry,Israel Cohen
摘要:我们探讨内部表示的语音增强(SE)模型在噪声条件下。使用在VoiceBank DEMAND上训练的变换卷积模型MUSE,我们分析了编码器,潜在,解码器和细化模块中的激活,同时将输入信噪比(SNR)从-10到30 dB进行扫描。我们使用中心核对齐(CKA)来测量逐点表示相似性和扩散距离,以捕获SNR之间的分布变化。结果表明,编码器的CKA在噪声和干净输入之间保持稳定和潜在的,解码器的CKA急剧下降,随着信噪比的降低。CKA与SNR的线性拟合揭示了深度依赖的鲁棒性-灵敏度权衡。扩散距离随着SNR在每个层内递增地变化,但在层间差异很大,特别是在低SNR下。总之,这些研究结果表明,噪声水平差异激活模型区域,并诱导不同的层间动态,激励SNR意识的条件反射和细化策略SE。
摘要:We probe internal representations of a speech enhancement (SE) model across noise conditions. Using MUSE, a transformer-convolutional model trained on VoiceBank DEMAND, we analyze activations in encoder, latent, decoder, and refinement blocks while sweeping input signal-to-noise-ratios (SNRs) from -10 to 30 dB. We use Centered Kernel Alignment (CKA) to measure point-wise representation similarity and diffusion distance to capture distributional shifts across SNRs. Results show that the encoder CKA between noisy and clean inputs remains stable and latent and decoder CKA drop sharply as SNR decreases. Linear fits of CKA versus SNR reveal a depth-dependent robustness-sensitivity trade-off. The diffusion distance varies incrementally with SNR within each layer but differs strongly across layers, especially at low SNRs. Together, these findings indicate that noise levels differentially activate model regions and induce distinct inter-layer dynamics, motivating SNR-aware conditioning and refinement strategies for SE.


机器翻译由腾讯交互翻译提供,仅供参考