今日论文合集:cs.SD语音5篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning
标题:Omni-AutoThink:通过强化学习进行自适应多模式推理
链接:https://arxiv.org/abs/2512.03783

作者:Dongchao Yang,Songxiang Liu,Disong Wang,Yuanyuan Wang,Guanglu Wan,Helen Meng
摘要:Omni模型的最新进展使统一的多模态感知和生成成为可能。然而,大多数现有的系统仍然表现出僵化的推理行为,要么过度思考简单的问题,要么在必要时无法推理。为了解决这个问题,我们提出了Omni-AutoThink,一种新的自适应推理框架,根据任务难度动态调整模型的推理深度。我们的框架包括两个阶段:(1)自适应监督微调(自适应SFT)阶段,使用大规模推理增强数据赋予Omni模型基本推理能力,以及(2)自适应强化学习(自适应GRPO)阶段,基于任务复杂性和奖励反馈优化推理行为。我们进一步构建了一个全面的自适应推理基准,跨越纯文本,文本音频,文本视觉和文本音频视觉模式,为多模态推理评估提供训练和评估。实验结果表明,我们提出的框架显着提高自适应推理性能相比,以前的基线。所有基准测试数据和代码都将公开发布。
摘要:Recent advances in Omni models have enabled unified multimodal perception and generation. However, most existing systems still exhibit rigid reasoning behaviors, either overthinking simple problems or failing to reason when necessary. To address this limitation, we propose Omni-AutoThink, a novel adaptive reasoning framework that dynamically adjusts the model's reasoning depth according to task difficulty. Our framework comprises two stages: (1) an Adaptive Supervised Fine-Tuning (Adaptive SFT) stage, which endows the Omni model with fundamental reasoning capability using large-scale reasoning-augmented data, and (2) an Adaptive Reinforcement Learning (Adaptive GRPO) stage, which optimizes reasoning behaviors based on task complexity and reward feedback. We further construct a comprehensive adaptive reasoning benchmark that spans text-only, text-audio, text-visual, and text-audio-visual modalities, providing both training and evaluation splits for multimodal reasoning assessment. Experimental results demonstrate that our proposed framework significantly improves adaptive reasoning performance compared to previous baselines. All benchmark data and code will be publicly released.


【2】AaPE: Aliasing-aware Patch Embedding for Self-Supervised Audio Representation Learning
标题:AaPE:用于自我监督音频表示学习的混淆感知补丁嵌入
链接:https://arxiv.org/abs/2512.03637

作者:Kohei Yamamoto,Kosuke Okusa
备注:11 pages, 4 figures
摘要:基于transformer的音频SSL(自监督学习)模型通常将频谱图视为图像,应用卷积补丁化和大量时间下采样。这会降低有效奈奎斯特频率并引入混叠,而简单的低通滤波会删除与任务相关的高频提示。在这项研究中,我们提出了混叠感知补丁嵌入(AaPE),一个下拉式补丁干,减轻混叠,同时保留高频信息。AaPE增强了标准补丁令牌,其特征由带限复正弦内核使用动态针对易混叠频带的双侧指数窗口产生。从输入估计内核的频率和衰减参数,使并行,自适应子带分析,其输出与标准补丁令牌融合。AaPE无缝集成到蒙面师生自我监督学习中。此外,我们将多掩码策略与对比目标相结合,以加强不同掩码模式的一致性,稳定训练。在AudioSet上进行预培训,然后在不同的下游基准上进行微调评估,这些基准涵盖环境声音和其他常见音频领域等类别。这种方法在一部分任务上产生了最先进的性能,在其余任务上产生了有竞争力的结果。互补的线性探测评估反映了这种模式,在几个基准测试中获得了明显的收益,在其他地方也有很好的表现。对这些结果的集体分析表明,AaPE有助于减轻混叠的影响,而不会丢弃信息丰富的高频内容。
摘要:Transformer-based audio SSL (self-supervised learning) models often treat spectrograms as images, applying convolutional patchification with heavy temporal downsampling. This lowers the effective Nyquist frequency and introduces aliasing, while naïve low-pass filtering removes task-relevant high-frequency cues. In this study, we present Aliasing-aware Patch Embedding (AaPE), a drop-in patch stem that mitigates aliasing while preserving high-frequency information. AaPE augments standard patch tokens with features produced by a band-limited complex sinusoidal kernel using a two-sided exponential window that dynamically targets alias-prone bands. Frequency and decay parameters of the kernel are estimated from the input, enabling parallel, adaptive subband analysis whose outputs are fused with the standard patch tokens. AaPE integrates seamlessly into the masked teacher-student self-supervised learning. In addition, we combine a multi-mask strategy with a contrastive objective to enforce consistency across diverse mask patterns, stabilizing training. Pre-training on AudioSet followed by fine-tuning evaluation across diverse downstream benchmarks, which spanned categories, such as environmental sounds and other common audio domains. This approach yields state-of-the-art performance on a subset of tasks and competitive results across the remainder. Complementary linear probing evaluation mirrors this pattern, yielding clear gains on several benchmarks and strong performance elsewhere. The collective analysis of these results indicates that AaPE serves to mitigate the effects of aliasing without discarding of informative high-frequency content.


【3】Head, posture, and full-body gestures in interactive communication
标题:互动交流中的头部、姿势和全身手势
链接:https://arxiv.org/abs/2512.03636

作者:Ľuboš Hládek,Bernhard U. Seeber
备注:7 figures, 10 tables, 30 pages
摘要:当面对面的沟通由于背景噪音或干扰谈话者而变得费力时,视觉线索的作用对于沟通的成功变得越来越重要。虽然以前的研究有选择地检查头部或手部的运动,在这里,我们探讨在声学不利条件下的整个身体的运动。我们假设,在对话中增加背景噪音会导致增加手势频率的手,头部,躯干,腿部运动的典型对话。增加手部动作的使用应该支持说话者的角色,而增加头部和躯干的运动可能会帮助听者。我们进行了一个自由的二元对话实验正常听力参与者(n=8)在虚拟声学环境。使用新开发的典型会话动作的标记系统描述会话动作,并分析各个类型的频率。此外,我们通过评估手语音同步性来分析手势质量,假设更高水平的背景噪声会导致根据交互耦合模型的同步性丧失。更高的噪音水平导致说话和倾听过程中手势复杂性增加,头部上下运动更明显,与预期相反,倾听过程中的头部运动通常相对于说话减少。同步性和峰值速度不受噪声影响,而手势质量仅适度缩放。研究结果支持了之前关于手势频率的发现,但我们只发现了有限的证据来证明语音手势同步的变化。这项工作揭示了整个身体的沟通模式,并说明了多模态适应的沟通需求。
摘要:When face-to-face communication becomes effortful due to background noise or interfering talkers, the role of visual cues becomes increasingly important for communication success. While previous research has selectively examined head or hand movements, here we explore movements of the whole body in acoustically adverse conditions. We hypothesized that increasing background noise in conversations would lead to increased gesture frequency in hand, head, trunk, and leg movements typical of conversation. Increased use of hand movements should support the speaker's role, while increased head and trunk movements may help the listener. We conducted a free dyadic conversation experiment with normal-hearing participants (n=8) in a virtual acoustic environment. Conversational movements were described with a newly developed labeling system for typical conversational actions, and the frequency of individual types was analyzed. In addition, we analyzed gesture quality by assessing hand-speech synchrony, with the hypothesis that higher levels of background noise would lead to a loss of synchrony according to an interactive coupling model. Higher noise levels led to increased hand-gesture complexity during speaking and listening, more pronounced up-down head movements, and contrary to expectations, head movements during listening generally decreased relative to speaking. Synchrony and peak velocity were unaffected by noise, while gesture quality scaled only modestly. The results support previous findings regarding gesturing frequency, but we found only limited evidence for changes in speech-gesture synchrony. This work reveals communication patterns of the whole body and illustrates multimodal adaptation to communication demands.


【4】State Space Models for Bioacoustics: A comparative Evaluation with Transformers
标题:生物声学的状态空间模型:与Transformer的比较评估
链接:https://arxiv.org/abs/2512.03563

作者:Chengyu Tang,Sanjeev Baskiyar
摘要:在这项研究中,我们评估的曼巴模型在生物声学领域的功效。我们首先使用自监督学习在大型音频数据语料库上预训练基于Mamba的音频大语言模型(LLM)。我们在BEANS基准上对BioMamba进行了微调和评估,BEANS基准是一系列不同的生物声学任务,包括分类和检测,并将其性能和效率与多个基线模型进行了比较,包括AVES,一种最先进的基于Transformer的模型。结果表明,BioMamba实现了与AVES相当的性能,同时消耗的VRAM显著减少,证明了其在该领域的潜力。
摘要:In this study, we evaluate the efficacy of the Mamba model in the field of bioacoustics. We first pretrain a Mamba-based audio large language model (LLM) on a large corpus of audio data using self-supervised learning. We fine-tune and evaluate BioMamba on the BEANS benchmark, a collection of diverse bioacoustic tasks including classification and detection, and compare its performance and efficiency with multiple baseline models, including AVES, a state-of-the-art Transformer-based model. The results show that BioMamba achieves comparable performance with AVES while consumption significantly less VRAM, demonstrating its potential in this domain.


【5】A Convolutional Framework for Mapping Imagined Auditory MEG into Listened Brain Responses
标题:将想象的听觉MEG映射到聆听的大脑反应的卷积框架
链接:https://arxiv.org/abs/2512.03458

作者:Maryam Maghsoudi,Mohsen Rezaeizadeh,Shihab Shamma
摘要:解码想象的语音涉及复杂的神经过程,由于时间的不确定性和有限的可用性,这些神经过程很难解释。在这项研究中,我们提出了一个脑磁图(MEG)数据集收集训练有素的音乐家,因为他们想象和听音乐和诗歌刺激。我们发现,想象和感知的大脑反应包含一致的,特定条件的信息。使用滑动窗口岭回归模型,我们第一次映射想象的反应,听取反应在单一主题的水平,但发现有限的泛化跨主题。在组级别,我们开发了一个编码器-解码器卷积神经网络,它具有特定于主题的校准层,可以产生稳定和可推广的映射。CNN的表现一直优于空模型,几乎所有被试的预测反应和真实反应之间的相关性都显著更高。我们的研究结果表明,想象的神经活动可以转化为类似感知的反应,为未来涉及想象语音和音乐的脑机接口应用提供了基础。
摘要:Decoding imagined speech engages complex neural processes that are difficult to interpret due to uncertainty in timing and the limited availability of imagined-response datasets. In this study, we present a Magnetoencephalography (MEG) dataset collected from trained musicians as they imagined and listened to musical and poetic stimuli. We show that both imagined and perceived brain responses contain consistent, condition-specific information. Using a sliding-window ridge regression model, we first mapped imagined responses to listened responses at the single-subject level, but found limited generalization across subjects. At the group level, we developed an encoder-decoder convolutional neural network with a subject-specific calibration layer that produced stable and generalizable mappings. The CNN consistently outperformed the null model, yielding significantly higher correlations between predicted and true listened responses for nearly all held-out subjects. Our findings demonstrate that imagined neural activity can be transformed into perception-like responses, providing a foundation for future brain-computer interface applications involving imagined speech and music.


eess.AS音频处理


【1】A Universal Harmonic Discriminator for High-quality GAN-based Vocoder
标题:用于高质量基于GAN的声码器的通用调和鉴别器
链接:https://arxiv.org/abs/2512.03486

作者:Nan Xu,Zhaolong Huang,Xiao Zeng
备注:Accepted by ASRU2025
摘要:随着基于GAN的声码器的出现,作为关键部件的声码器最近也得到了发展。在我们的工作中,我们专注于改进基于时频的OFDM。特别地,短时傅立叶变换(STFT)表示通常用作基于时频的OFDM的输入。然而,STFT谱图在不同的频率点具有相同的频率分辨率,这导致性能较差,特别是对于歌唱声音。受此启发,我们提出了一个通用的谐波分析仪的动态频率分辨率建模和谐波跟踪。具体来说,我们设计了一个谐波滤波器与可学习的三角带通滤波器组,其中每个频率箱有一个灵活的带宽。此外,我们增加了一个半谐波捕捉细粒度的谐波关系在低频带。在语音和歌唱数据集上的实验验证了该方法在主观和客观指标上的有效性。
摘要:With the emergence of GAN-based vocoders, the discriminator, as a crucial component, has been developed recently. In our work, we focus on improving the time-frequency based discriminator. Particularly, Short-Time Fourier Transform (STFT) representation is usually used as input of time-frequency based discriminator. However, the STFT spectrogram has the same frequency resolution at different frequency bins, which results in an inferior performance, especially for singing voices. Motivated by this, we propose a universal harmonic discriminator for dynamic frequency resolution modeling and harmonic tracking. Specifically, we design a harmonic filter with learnable triangular band-pass filter banks, where each frequency bin has a flexible bandwidth. Additionally, we add a half-harmonic to capture fine-grained harmonic relationships at low-frequency band. Experiments on speech and singing datasets validate the effectiveness of the proposed discriminator on both subjective and objective metrics.


【2】A Convolutional Framework for Mapping Imagined Auditory MEG into Listened Brain Responses
标题:将想象的听觉MEG映射到聆听的大脑反应的卷积框架
链接:https://arxiv.org/abs/2512.03458

作者:Maryam Maghsoudi,Mohsen Rezaeizadeh,Shihab Shamma
摘要:解码想象的语音涉及复杂的神经过程,由于时间的不确定性和有限的可用性,这些神经过程很难解释。在这项研究中,我们提出了一个脑磁图(MEG)数据集收集训练有素的音乐家,因为他们想象和听音乐和诗歌刺激。我们发现,想象和感知的大脑反应包含一致的,特定条件的信息。使用滑动窗口岭回归模型,我们第一次映射想象的反应,听取反应在单一主题的水平,但发现有限的泛化跨主题。在组级别,我们开发了一个编码器-解码器卷积神经网络,它具有特定于主题的校准层,可以产生稳定和可推广的映射。CNN的表现一直优于空模型,几乎所有被试的预测反应和真实反应之间的相关性都显著更高。我们的研究结果表明,想象的神经活动可以转化为类似感知的反应,为未来涉及想象语音和音乐的脑机接口应用提供了基础。
摘要:Decoding imagined speech engages complex neural processes that are difficult to interpret due to uncertainty in timing and the limited availability of imagined-response datasets. In this study, we present a Magnetoencephalography (MEG) dataset collected from trained musicians as they imagined and listened to musical and poetic stimuli. We show that both imagined and perceived brain responses contain consistent, condition-specific information. Using a sliding-window ridge regression model, we first mapped imagined responses to listened responses at the single-subject level, but found limited generalization across subjects. At the group level, we developed an encoder-decoder convolutional neural network with a subject-specific calibration layer that produced stable and generalizable mappings. The CNN consistently outperformed the null model, yielding significantly higher correlations between predicted and true listened responses for nearly all held-out subjects. Our findings demonstrate that imagined neural activity can be transformed into perception-like responses, providing a foundation for future brain-computer interface applications involving imagined speech and music.


【3】Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
标题:比较无监督和有监督的语义语音标记:儿童ASB的案例研究
链接:https://arxiv.org/abs/2512.03301

作者:Mohan Shi,Natarajan Balaji Shankar,Kaiyuan Zhang,Zilai Wang,Abeer Alwan
备注:ASRU-AI4CSL
摘要:离散语音标记因其存储效率和与大型语言模型(LLM)的集成而受到关注。它们通常分为声学和语义标记,后者对自动语音识别(ASR)更有利。传统上,无监督K均值聚类已被用于从语音基础模型(SFM)中提取语义语音标记。最近,出现了用于语音生成的监督方法,例如用ASR损失训练的有限标量量化(FSQ)。这两种方法都利用了预先训练的SFM,有利于低资源任务,如儿童ASR。   本文系统地比较了有监督和无监督的语义语音令牌的儿童ASR。结果表明,有监督的方法不仅优于无监督的方法,甚至出乎意料地超过连续表示,即使在超低比特率设置下也表现良好。这些发现突出了监督语义标记的优势,并为改进离散语音标记化提供了见解。
摘要:Discrete speech tokens have gained attention for their storage efficiency and integration with Large Language Models (LLMs). They are commonly categorized into acoustic and semantic tokens, with the latter being more advantageous for Automatic Speech Recognition (ASR). Traditionally, unsupervised K-means clustering has been used to extract semantic speech tokens from Speech Foundation Models (SFMs). Recently, supervised methods, such as finite scalar quantization (FSQ) trained with ASR loss, have emerged for speech generation. Both approaches leverage pre-trained SFMs, benefiting low-resource tasks such as child ASR.   This paper systematically compares supervised and unsupervised semantic speech tokens for child ASR. Results show that supervised methods not only outperform unsupervised ones but even unexpectedly surpass continuous representations, and they perform well even in ultra-low bitrate settings. These findings highlight the advantages of supervised semantic tokens and offer insights for improving discrete speech tokenization.


【4】Head, posture, and full-body gestures in interactive communication
标题:互动交流中的头部、姿势和全身手势
链接:https://arxiv.org/abs/2512.03636

作者:Ľuboš Hládek,Bernhard U. Seeber
备注:7 figures, 10 tables, 30 pages
摘要:当面对面的沟通由于背景噪音或干扰谈话者而变得费力时,视觉线索的作用对于沟通的成功变得越来越重要。虽然以前的研究有选择地检查头部或手部的运动,在这里,我们探讨在声学不利条件下的整个身体的运动。我们假设,在对话中增加背景噪音会导致增加手势频率的手,头部,躯干,腿部运动的典型对话。增加手部动作的使用应该支持说话者的角色,而增加头部和躯干的运动可能会帮助听者。我们进行了一个自由的二元对话实验正常听力参与者(n=8)在虚拟声学环境。使用新开发的典型会话动作的标记系统描述会话动作,并分析各个类型的频率。此外,我们通过评估手语音同步性来分析手势质量,假设更高水平的背景噪声会导致根据交互耦合模型的同步性丧失。更高的噪音水平导致说话和倾听过程中手势复杂性增加,头部上下运动更明显,与预期相反,倾听过程中的头部运动通常相对于说话减少。同步性和峰值速度不受噪声影响,而手势质量仅适度缩放。研究结果支持了之前关于手势频率的发现,但我们只发现了有限的证据来证明语音手势同步的变化。这项工作揭示了整个身体的沟通模式,并说明了多模态适应的沟通需求。
摘要:When face-to-face communication becomes effortful due to background noise or interfering talkers, the role of visual cues becomes increasingly important for communication success. While previous research has selectively examined head or hand movements, here we explore movements of the whole body in acoustically adverse conditions. We hypothesized that increasing background noise in conversations would lead to increased gesture frequency in hand, head, trunk, and leg movements typical of conversation. Increased use of hand movements should support the speaker's role, while increased head and trunk movements may help the listener. We conducted a free dyadic conversation experiment with normal-hearing participants (n=8) in a virtual acoustic environment. Conversational movements were described with a newly developed labeling system for typical conversational actions, and the frequency of individual types was analyzed. In addition, we analyzed gesture quality by assessing hand-speech synchrony, with the hypothesis that higher levels of background noise would lead to a loss of synchrony according to an interactive coupling model. Higher noise levels led to increased hand-gesture complexity during speaking and listening, more pronounced up-down head movements, and contrary to expectations, head movements during listening generally decreased relative to speaking. Synchrony and peak velocity were unaffected by noise, while gesture quality scaled only modestly. The results support previous findings regarding gesturing frequency, but we found only limited evidence for changes in speech-gesture synchrony. This work reveals communication patterns of the whole body and illustrates multimodal adaptation to communication demands.


机器翻译由腾讯交互翻译提供,仅供参考