今日论文合集:cs.SD语音16篇,eess.AS音频处理17篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Modulation Discovery with Differentiable Digital Signal Processing
标题:利用差异数字信号处理实现调制发现
链接:https://arxiv.org/abs/2510.06204

作者:Christopher Mitcheltree, Hao Hao Tan, Joshua D. Reiss
备注:Accepted to WASPAA 2025 (best paper award candidate). Code, audio samples, and plugins can be found at this https URL
摘要:调制是声音设计和音乐制作的关键部分,可以创建复杂和不断发展的音频。现代频率合成器提供包络、低频振荡器(LFO)和更多的参数自动化工具,使用户能够轻松调制输出。然而,确定用于创建声音的调制信号是困难的,并且现有的声音匹配/参数估计系统通常是不可解释的黑箱或预测高维逐帧参数值,而不考虑底层调制曲线的形状、结构和路由。我们提出了一种神经声音匹配的方法,利用调制提取,约束控制信号参数化,和微分数字信号处理(DDSP),发现调制存在于一个声音。我们证明了我们的方法对高度调制的合成和真实音频样本的有效性,它适用于不同的DDSP合成器架构,并调查它在可解释性和声音匹配精度之间的权衡。我们提供代码和音频样本,并在VST插件中提供经过训练的DDSP合成器。
摘要:Modulations are a critical part of sound design and music production, enabling the creation of complex and evolving audio. Modern synthesizers provide envelopes, low frequency oscillators (LFOs), and more parameter automation tools that allow users to modulate the output with ease. However, determining the modulation signals used to create a sound is difficult, and existing sound-matching / parameter estimation systems are often uninterpretable black boxes or predict high-dimensional framewise parameter values without considering the shape, structure, and routing of the underlying modulation curves. We propose a neural sound-matching approach that leverages modulation extraction, constrained control signal parameterizations, and differentiable digital signal processing (DDSP) to discover the modulations present in a sound. We demonstrate the effectiveness of our approach on highly modulated synthetic and real audio samples, its applicability to different DDSP synth architectures, and investigate the trade-off it incurs between interpretability and sound-matching accuracy. We make our code and audio samples available and provide the trained DDSP synths in a VST plugin.


【2】EmoHRNet: High-Resolution Neural Network Based Speech Emotion Recognition
标题:CLARHRNet:基于高分辨率神经网络的语音情感识别
链接:https://arxiv.org/abs/2510.06072

作者:Akshay Muppidi, Martin Radfar
备注:None
摘要:语音情感识别(SER)是增强人机交互的关键。本文介绍了一种新的适应高分辨率网络(HRNet)为SER。HRNet结构的设计,以保持高分辨率表示从初始到最终层。通过将音频样本转换为频谱图,RISHHRNet利用HRNet架构来提取高级特征。CNOHRNet的独特架构始终保持高分辨率表示,从语音信号中捕获粒度和总体情感线索。该模型优于领先的模型,在RAVDESS上达到92.45%的准确率,在IEMOCAP上达到80.06%,在EMOVO上达到92.77%。因此,我们表明,ARMHRNet在SER域中设置了一个新的基准。
摘要:Speech emotion recognition (SER) is pivotal for enhancing human-machine interactions. This paper introduces "EmoHRNet", a novel adaptation of High-Resolution Networks (HRNet) tailored for SER. The HRNet structure is designed to maintain high-resolution representations from the initial to the final layers. By transforming audio samples into spectrograms, EmoHRNet leverages the HRNet architecture to extract high-level features. EmoHRNet's unique architecture maintains high-resolution representations throughout, capturing both granular and overarching emotional cues from speech signals. The model outperforms leading models, achieving accuracies of 92.45% on RAVDESS, 80.06% on IEMOCAP, and 92.77% on EMOVO. Thus, we show that EmoHRNet sets a new benchmark in the SER domain.


【3】ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
标题:ECTSpeech:通过简单的一致性调整增强高效的语音合成
链接:https://arxiv.org/abs/2510.05984

作者:Tao Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun, Wendong Zheng
备注:Accepted for publication by Proceedings of the 2025 ACM Multimedia Asia Conference(MMAsia '25)
摘要:扩散模型在语音合成中表现出显著的性能,但通常需要多步采样,导致推理效率低。最近的研究通过将扩散模型蒸馏成一致性模型来解决这个问题,从而实现高效的一步生成。然而,这些方法引入了额外的培训成本,并严重依赖于预先培训的教师模型的性能。在本文中,我们提出了ECTSpeech,一个简单而有效的一步语音合成框架,第一次,将容易的一致性调整(ECT)策略到语音合成。通过逐步收紧对预训练扩散模型的一致性约束,ECTSpeech实现了高质量的一步生成,同时显着降低了训练复杂性。此外,我们设计了一个多尺度门模块(MSGate),以提高去噪器的能力,融合功能在不同的尺度。LJSpeech数据集上的实验结果表明,ECTSpeech在单步采样下实现了与最先进方法相当的音频质量,同时大大降低了模型的训练成本和复杂性。
摘要:Diffusion models have demonstrated remarkable performance in speech synthesis, but typically require multi-step sampling, resulting in low inference efficiency. Recent studies address this issue by distilling diffusion models into consistency models, enabling efficient one-step generation. However, these approaches introduce additional training costs and rely heavily on the performance of pre-trained teacher models. In this paper, we propose ECTSpeech, a simple and effective one-step speech synthesis framework that, for the first time, incorporates the Easy Consistency Tuning (ECT) strategy into speech synthesis. By progressively tightening consistency constraints on a pre-trained diffusion model, ECTSpeech achieves high-quality one-step generation while significantly reducing training complexity. In addition, we design a multi-scale gate module (MSGate) to enhance the denoiser's ability to fuse features at different scales. Experimental results on the LJSpeech dataset demonstrate that ECTSpeech achieves audio quality comparable to state-of-the-art methods under single-step sampling, while substantially reducing the model's training cost and complexity.


【4】Segment-Factorized Full-Song Generation on Symbolic Piano Music
标题:象征性钢琴音乐的分段分解全曲生成
链接:https://arxiv.org/abs/2510.05881

作者:Ping-Yi Chen, Chih-Pin Tan, Yi-Hsuan Yang
备注:Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: AI for Music
摘要:我们提出了分段全歌模型(SFS)的符号全歌生成。该模型接受用户提供的歌曲结构和可选的短种子片段,该片段锚定歌曲开发的主要思想。通过将歌曲分解为片段并通过选择性注意相关片段来生成每个片段,该模型与先前的工作相比实现了更高的质量和效率。为了证明它对人类与人工智能交互的适用性,我们进一步将SFS包装到一个Web应用程序中,使用户能够以可定制的结构和灵活的顺序在钢琴卷上迭代地共同创作音乐。
摘要:We propose the Segmented Full-Song Model (SFS) for symbolic full-song generation. The model accepts a user-provided song structure and an optional short seed segment that anchors the main idea around which the song is developed. By factorizing a song into segments and generating each one through selective attention to related segments, the model achieves higher quality and efficiency compared to prior work. To demonstrate its suitability for human-AI interaction, we further wrap SFS into a web application that enables users to iteratively co-create music on a piano roll with customizable structures and flexible ordering.


【5】LARA-Gen: Enabling Continuous Emotion Control for Music Generation Models via Latent Affective Representation Alignment
标题:LARA-Gen:通过潜在情感表示对齐为音乐生成模型实现连续情感控制
链接:https://arxiv.org/abs/2510.05875

作者:Jiahao Mei, Xuenan Xu, Zeyu Xie, Zihao Zheng, Ye Tao, Yue Ding, Mengyue Wu
摘要:文本到音乐模型的最新进展已经能够从文本提示中生成连贯的音乐,但细粒度的情感控制仍然没有解决。我们介绍LARA-Gen,这是一个持续情感控制的框架,它通过潜在情感表征对齐(LARA)将内部隐藏状态与外部音乐理解模型对齐,从而实现有效的训练。此外,我们设计了一个基于连续效价唤醒空间的情感控制模块,将情感属性从文本内容中分离出来,绕过了基于文本提示的瓶颈。此外,我们通过精心策划的测试集和强大的情感预测器建立了基准,促进了对音乐生成中情感可控性的客观评估。大量的实验表明,LARA-Gen实现了对情感的连续、细粒度控制,并且在情感坚持和音乐质量方面都显著优于基线。生成的示例可在https://nieeim.github.io/LARA-Gen/上获得。
摘要:Recent advances in text-to-music models have enabled coherent music generation from text prompts, yet fine-grained emotional control remains unresolved. We introduce LARA-Gen, a framework for continuous emotion control that aligns the internal hidden states with an external music understanding model through Latent Affective Representation Alignment (LARA), enabling effective training. In addition, we design an emotion control module based on a continuous valence-arousal space, disentangling emotional attributes from textual content and bypassing the bottlenecks of text-based prompting. Furthermore, we establish a benchmark with a curated test set and a robust Emotion Predictor, facilitating objective evaluation of emotional controllability in music generation. Extensive experiments demonstrate that LARA-Gen achieves continuous, fine-grained control of emotion and significantly outperforms baselines in both emotion adherence and music quality. Generated samples are available at https://nieeim.github.io/LARA-Gen/.


【6】FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
标题:FoleyFinder:使用与GRAM对齐的多模式编码器的视频到音频生成
链接:https://arxiv.org/abs/2510.05829

作者:Riccardo Fosco Gramaccioni, Christian Marinoni, Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello
备注:Acepted at IJCNN 2025
摘要:在这项工作中,我们提出了Foleystrike,一种新的方法来视频到音频的生成,强调通过使用对齐的多模态编码器的语义条件。基于视频到音频生成的先前进步,Foleywalk利用Gramian表示对齐度量(RMM)来对齐视频,文本和音频模态的嵌入,从而实现对音频生成过程的精确语义控制。Foleystrike的核心是一个基于扩散的音频合成模型,以GRAM对齐的嵌入和波形包络为条件,确保语义丰富性和与相应输入视频的时间对齐。我们在Greatest Hits数据集上评估了Foleywatch,这是视频到音频模型的标准基准。我们的实验表明,对齐多模式编码器使用的MPEG增强了系统的能力,语义上对齐生成的音频与视频内容,推进了视频到音频合成的最新技术。
摘要:In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM leverages the Gramian Representation Alignment Measure (GRAM) to align embeddings across video, text, and audio modalities, enabling precise semantic control over the audio generation process. The core of FoleyGRAM is a diffusion-based audio synthesis model conditioned on GRAM-aligned embeddings and waveform envelopes, ensuring both semantic richness and temporal alignment with the corresponding input video. We evaluate FoleyGRAM on the Greatest Hits dataset, a standard benchmark for video-to-audio models. Our experiments demonstrate that aligning multimodal encoders using GRAM enhances the system's ability to semantically align generated audio with video content, advancing the state of the art in video-to-audio synthesis.


【7】StereoSync: Spatially-Aware Stereo Audio Generation from Video
标题:StereoLock:从视频中生成空间感知的立体声音频
链接:https://arxiv.org/abs/2510.05828

作者:Christian Marinoni, Riccardo Fosco Gramaccioni, Kazuki Shimada, Takashi Shibuya, Yuki Mitsufuji, Danilo Comminiello
备注:Accepted at IJCNN 2025
摘要:虽然近年来音频生成已被广泛研究,但视频对齐的音频生成仍然是一个相对未开发的前沿。为了解决这一差距,我们引入StereoSync,一种新颖而高效的模型,旨在生成与参考视频在时间上同步并与其视觉上下文在空间上对齐的音频。此外,StereoSync还通过利用预训练的基础模型来实现效率,减少了对大量训练的需求,同时保持高质量的合成。与主要关注时间同步的现有方法不同,StereoSync通过将空间感知纳入视频对齐的音频生成来引入显著的进步。事实上,给定输入视频,我们的方法从深度图和边界框中提取空间线索,将它们用作基于扩散的音频生成模型中的交叉注意调节。这种方法允许StereoSync超越简单的同步,产生动态适应视频场景的空间结构和移动的立体声音频。我们评估了StereoSync上的Walking The Maps,这是一个策展数据集,包括来自视频游戏的视频,这些视频游戏的特点是动画角色在不同的环境中行走。实验结果证明了StereoSync能够实现时间和空间对齐,推进了视频到音频生成的最新技术水平,并带来了更加身临其境和逼真的音频体验。
摘要:Although audio generation has been widely studied over recent years, video-aligned audio generation still remains a relatively unexplored frontier. To address this gap, we introduce StereoSync, a novel and efficient model designed to generate audio that is both temporally synchronized with a reference video and spatially aligned with its visual context. Moreover, StereoSync also achieves efficiency by leveraging pretrained foundation models, reducing the need for extensive training while maintaining high-quality synthesis. Unlike existing methods that primarily focus on temporal synchronization, StereoSync introduces a significant advancement by incorporating spatial awareness into video-aligned audio generation. Indeed, given an input video, our approach extracts spatial cues from depth maps and bounding boxes, using them as cross-attention conditioning in a diffusion-based audio generation model. Such an approach allows StereoSync to go beyond simple synchronization, producing stereo audio that dynamically adapts to the spatial structure and movement of a video scene. We evaluate StereoSync on Walking The Maps, a curated dataset comprising videos from video games that feature animated characters walking through diverse environments. Experimental results demonstrate the ability of StereoSync to achieve both temporal and spatial alignment, advancing the state of the art in video-to-audio generation and resulting in a significantly more immersive and realistic audio experience.


【8】Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech
标题:基于LLM的文本转语音的数据高效有针对性的令牌级偏好优化
链接:https://arxiv.org/abs/2510.05799

作者:Rikuto Kotoge, Yuichi Sasaki
摘要:通过偏好优化将文本到语音(TTS)系统输出与人类反馈对齐,已被证明可以有效地提高基于语言模型的TTS模型的鲁棒性和自然度。目前的方法主要需要在话语水平配对的期望和不期望的样本。然而,这样的对通常在TTS输出数据中是有限的,并且话语级公式化防止精确发音对齐所需的细粒度令牌级优化。在这项研究中,我们提出了TKTO,它消除了对配对数据的需求,实现了更有效的数据训练范式,并直接针对令牌级单元,自动提供细粒度的对齐信号,而无需令牌级注释。TKTO将具有挑战性的日语TTS准确性提高了39%,并将CER降低了54%,自动为目标代币分配12.8倍的奖励。
摘要:Aligning text-to-speech (TTS) system outputs with human feedback through preference optimization has been shown to effectively improve the robustness and naturalness of language model-based TTS models. Current approaches primarily require paired desirable and undesirable samples at the utterance level. However, such pairs are often limited in TTS output data, and utterance-level formulation prevents fine-grained token-level optimization needed for accurate pronunciation alignment. In this study, we propose TKTO that eliminates the need for paired data, enabling a more data-efficient training paradigm, and directly targets token-level units, automatically providing fine-grained alignment signals without token-level annotations. TKTO improves the challenging Japanese TTS accuracy by 39% and reduces CER by 54%, automatically assigning 12.8 times stronger reward to targeted tokens.


【9】EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTS
标题:EMOLL-TTC:基于LLM的TTC中用于细粒度情绪控制的强化学习
链接:https://arxiv.org/abs/2510.05758

作者:Haoxun Li, Yu Liu, Yuqing Sun, Hanlei Shi, Leyuan Qu, Taihao Li
备注:Under review for ICASSP 2026
摘要:目前基于LLM的TTS系统具有很强的质量和zero-shot能力,但由于依赖于离散的语音标记,缺乏细粒度的情感控制。现有的方法要么将情感限制在分类标签上,要么不能推广到基于LLM的架构。我们提出了EMORL-TTS(Fine-grained prediction-controllable TTS with Reinforcement Learning),这是一个将VAD空间中的全局强度控制与局部强调调节相结合的框架。我们的方法将监督微调与强化学习相结合,由针对情绪类别,强度和重点的任务特定奖励指导。此外,我们进一步研究如何强调放置调制细粒度的情绪强度。实验表明,EMORL-TTS提高了情感的准确性,强度差异和重点清晰度,同时保持合成质量与基于LLM的基线相当。
摘要:Recent LLM-based TTS systems achieve strong quality and zero-shot ability, but lack fine-grained emotional control due to their reliance on discrete speech tokens. Existing approaches either limit emotions to categorical labels or cannot generalize to LLM-based architectures. We propose EMORL-TTS (Fine-grained Emotion-controllable TTS with Reinforcement Learning), a framework that unifies global intensity control in the VAD space with local emphasis regulation. Our method combines supervised fine-tuning with reinforcement learning guided by task-specific rewards for emotion category, intensity, and emphasis. Moreover, we further investigate how emphasis placement modulates fine-grained emotion intensity. Experiments show that EMORL-TTS improves emotion accuracy, intensity differentiation, and emphasis clarity, while preserving synthesis quality comparable to strong LLM-based baselines.


【10】Transcribing Rhythmic Patterns of the Guitar Track in Polyphonic Music
标题:复调音乐中吉他音轨节奏模式的转录
链接:https://arxiv.org/abs/2510.05756

作者:Aleksandr Lukoianov, Anssi Klapuri
备注:Accepted to WASPAA 2025
摘要:然而,在过去的几十年里,和弦转录受到了相当大的关注,很少有人致力于转录和编码歌曲中出现的节奏模式。这一主题与节奏吉他等乐器尤其相关,节奏吉他通常是通过重复和随时间变化的节奏模式来演奏的。然而,在许多情况下,人们不能客观地定义一个单一的“正确”的节奏模式,为一个给定的歌曲部分。为了创建一个具有明确定义的地面实况标签的数据集,我们请专业音乐家转录410首流行歌曲和唱片封面版本中的节奏模式,其中吉他曲目跟随这些转录。为了记录弦乐器及其相应的节奏模式,我们提出了一个三步框架。首先,我们执行近似干分离提取吉他部分从复调混合。其次,我们使用预训练的基础模型(MERT)作为骨干,在分离的吉他音频中检测单个琴弦。最后,我们进行了一个模式解码过程中,吉他弹奏的转录序列是由来自专家策划的词汇表的模式表示。我们表明,它是可以转录的节奏模式的吉他音轨在复调音乐具有相当高的准确性,产生一个表示,是人类可读的,包括自动检测的酒吧线和时间签名标记。我们进行消融研究和错误分析,并提出了一套评估指标,以评估预测的节奏模式序列的准确性和可读性。
摘要:Whereas chord transcription has received considerable attention during the past couple of decades, far less work has been devoted to transcribing and encoding the rhythmic patterns that occur in a song. The topic is especially relevant for instruments such as the rhythm guitar, which is typically played by strumming rhythmic patterns that repeat and vary over time. However, in many cases one cannot objectively define a single "right" rhythmic pattern for a given song section. To create a dataset with well-defined ground-truth labels, we asked expert musicians to transcribe the rhythmic patterns in 410 popular songs and record cover versions where the guitar tracks followed those transcriptions. To transcribe the strums and their corresponding rhythmic patterns, we propose a three-step framework. Firstly, we perform approximate stem separation to extract the guitar part from the polyphonic mixture. Secondly, we detect individual strums within the separated guitar audio, using a pre-trained foundation model (MERT) as a backbone. Finally, we carry out a pattern-decoding process in which the transcribed sequence of guitar strums is represented by patterns drawn from an expert-curated vocabulary. We show that it is possible to transcribe the rhythmic patterns of the guitar track in polyphonic music with quite high accuracy, producing a representation that is human-readable and includes automatically detected bar lines and time signature markers. We perform ablation studies and error analysis and propose a set of evaluation metrics to assess the accuracy and readability of the predicted rhythmic pattern sequence.


【11】MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition
标题:MSF-BER:用多粒度语义丰富声学建模以实现语音情感识别
链接:https://arxiv.org/abs/2510.05749

作者:Haoxun Li, Yuqing Sun, Hanlei Shi, Yu Liu, Leyuan Qu, Taihao Li
备注:Under review for ICASSP 2026
摘要:连续维度语音情感识别捕获沿着效价、唤醒和支配的情感变化,提供比分类方法更细粒度的表示。然而,大多数多模态方法仅依赖于全局转录,导致两个局限性:(1)所有单词都被平等对待,忽略了句子不同部分的强调可以改变情感意义;(2)只表示表面词汇内容,缺乏更高层次的解释线索。为了克服这些问题,我们提出了MSF-SER(多粒度语义融合语音情感识别),它增强了声学特征与文本语义的三个互补层次-局部强调语义(LES),全局语义(GS),和扩展语义(ES)。这些通过模态内门控融合和跨模态FiLM调制的轻量级专家混合(FM-MOE)进行集成。在MSP-Podcast和IEMOCAP上的实验表明,MSF-SER一致地提高了维度预测,证明了增强语义融合的有效性。
摘要:Continuous dimensional speech emotion recognition captures affective variation along valence, arousal, and dominance, providing finer-grained representations than categorical approaches. Yet most multimodal methods rely solely on global transcripts, leading to two limitations: (1) all words are treated equally, overlooking that emphasis on different parts of a sentence can shift emotional meaning; (2) only surface lexical content is represented, lacking higher-level interpretive cues. To overcome these issues, we propose MSF-SER (Multi-granularity Semantic Fusion for Speech Emotion Recognition), which augments acoustic features with three complementary levels of textual semantics--Local Emphasized Semantics (LES), Global Semantics (GS), and Extended Semantics (ES). These are integrated via an intra-modal gated fusion and a cross-modal FiLM-modulated lightweight Mixture-of-Experts (FM-MOE). Experiments on MSP-Podcast and IEMOCAP show that MSF-SER consistently improves dimensional prediction, demonstrating the effectiveness of enriched semantic fusion for SER.


【12】Sparse deepfake detection promotes better disentanglement
标题:稀疏深度伪造检测促进更好的解纠缠
链接:https://arxiv.org/abs/2510.05696

作者:Antoine Teissier, Marie Tahon, Nicolas Dugué, Aghilas Sini
摘要:由于语音合成的快速发展,deepfake检测已成为语音处理社区的主要关注点。因为这是一项关键任务,系统不仅必须高效和健壮,而且还必须提供可解释的解释。在可解释性的不同方法中,我们专注于潜在表征的解释。在这篇论文中,我们重点介绍了AASIST的最后一层嵌入,这是一种deepfake检测架构。我们在这一层上使用受SAE启发的TopK激活来获得用于决策过程的稀疏表示。我们证明了稀疏深度伪造检测可以提高检测性能,在ASVSpoof5测试集上的EER为23.36%,稀疏度为95%。然后,我们表明,这些表示提供更好的解纠缠,使用基于互信息的完整性和模块化度量。值得注意的是,一些攻击直接编码在潜在空间中。
摘要:Due to the rapid progress of speech synthesis, deepfake detection has become a major concern in the speech processing community. Because it is a critical task, systems must not only be efficient and robust, but also provide interpretable explanations. Among the different approaches for explainability, we focus on the interpretation of latent representations. In such paper, we focus on the last layer of embeddings of AASIST, a deepfake detection architecture. We use a TopK activation inspired by SAEs on this layer to obtain sparse representations which are used in the decision process. We demonstrate that sparse deepfake detection can improve detection performance, with an EER of 23.36% on ASVSpoof5 test set, with 95% of sparsity. We then show that these representations provide better disentanglement, using completeness and modularity metrics based on mutual information. Notably, some attacks are directly encoded in the latent space.


【13】Sci-Phi: A Large Language Model Spatial Audio Descriptor
标题:Sci-Phi:一种大型语言模型空间音频描述符
链接:https://arxiv.org/abs/2510.05542

作者:Xilin Jiang, Hannes Gamper, Sebastian Braun
摘要:声学场景感知包括描述声音的类型,它们的时间,它们的方向和距离,以及它们的响度和混响。虽然音频语言模型在声音识别方面表现出色,但单通道输入从根本上限制了空间理解。这项工作提出了Sci-Phi,这是一种具有双重空间和频谱编码器的空间音频大语言模型,可以为所有声源和周围环境估计完整的参数集。Sci-Phi从超过4,000小时的合成一阶Ambisonics录音(包括元数据)中学习,一次性列举并描述了多达四个定向声源,以及非定向背景声音和房间特征。我们使用置换不变协议和15个涵盖内容,位置,定时,响度和混响的度量来评估模型,并分析其在源计数,信噪比,混响水平以及具有挑战性的声学,空间或时间相似源的混合物中的鲁棒性。值得注意的是,Sci-Phi推广到实际房间脉冲响应,只有轻微的性能下降。总的来说,这项工作建立了第一个音频LLM能够完整的空间场景描述,具有强大的潜力,为现实世界的部署。演示:https://sci-phi-audio.github.io/demo
摘要:Acoustic scene perception involves describing the type of sounds, their timing, their direction and distance, as well as their loudness and reverberation. While audio language models excel in sound recognition, single-channel input fundamentally limits spatial understanding. This work presents Sci-Phi, a spatial audio large language model with dual spatial and spectral encoders that estimates a complete parameter set for all sound sources and the surrounding environment. Learning from over 4,000 hours of synthetic first-order Ambisonics recordings including metadata, Sci-Phi enumerates and describes up to four directional sound sources in one pass, alongside non-directional background sounds and room characteristics. We evaluate the model with a permutation-invariant protocol and 15 metrics covering content, location, timing, loudness, and reverberation, and analyze its robustness across source counts, signal-to-noise ratios, reverberation levels, and challenging mixtures of acoustically, spatially, or temporally similar sources. Notably, Sci-Phi generalizes to real room impulse responses with only minor performance degradation. Overall, this work establishes the first audio LLM capable of full spatial-scene description, with strong potential for real-world deployment. Demo: https://sci-phi-audio.github.io/demo


【14】AUREXA-SE: Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement
标题:AURREXA-SE:具有交叉注意力和挤压器的视听统一表示交换架构,用于语音增强
链接:https://arxiv.org/abs/2510.05295

作者:M. Sajid, Deepanshu Gupta, Yash Modi, Sanskriti Jain, Harshith Jai Surya Ganji, A. Rahaman, Harshvardhan Choudhary, Nasir Saleem, Amir Hussain, M. Tanveer
摘要:在本文中,我们提出了AUREXA-SE(视听统一表示交换架构与交叉注意力和挤压语音增强),一个渐进的双峰框架视听语音增强(AVSE)量身定制。AUREXA-SE通过采用基于U-Net的1D卷积编码器和Swin Transformer V2来实现高效和富有表现力的视觉特征提取,从而联合利用原始音频波形和视觉线索。该架构的核心是一种新颖的双向交叉注意机制,它有助于模态之间的深度上下文融合,从而实现丰富和互补的表征学习。为了捕获融合嵌入中的时间依赖性,引入了一堆结合卷积和注意力模块的轻量级Squeezeformer块。增强的嵌入然后通过U-Net风格的解码器进行解码,以进行直接波形重建,确保感知一致和可理解的语音输出。实验评估证明了AUREXA-SE的有效性,在噪声基线上实现了显着的性能改善,STOI为0.516,PESQ为1.323,SI-SDR为-4.322 dB。AUREXA-SE的源代码可在https://github.com/mtanveer1/AVSEC-4-Challenge-2025上获得。
摘要:In this paper, we propose AUREXA-SE (Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement), a progressive bimodal framework tailored for audio-visual speech enhancement (AVSE). AUREXA-SE jointly leverages raw audio waveforms and visual cues by employing a U-Net-based 1D convolutional encoder for audio and a Swin Transformer V2 for efficient and expressive visual feature extraction. Central to the architecture is a novel bidirectional cross-attention mechanism, which facilitates deep contextual fusion between modalities, enabling rich and complementary representation learning. To capture temporal dependencies within the fused embeddings, a stack of lightweight Squeezeformer blocks combining convolutional and attention modules is introduced. The enhanced embeddings are then decoded via a U-Net-style decoder for direct waveform reconstruction, ensuring perceptually consistent and intelligible speech output. Experimental evaluations demonstrate the effectiveness of AUREXA-SE, achieving significant performance improvements over noisy baselines, with STOI of 0.516, PESQ of 1.323, and SI-SDR of -4.322 dB. The source code of AUREXA-SE is available at https://github.com/mtanveer1/AVSEC-4-Challenge-2025.


【15】Provable Speech Attributes Conversion via Latent Independence
标题:通过潜在独立性的可证明语音属性转换
链接:https://arxiv.org/abs/2510.05191

作者:Jonathan Svirsky, Ofir Lindenbaum, Uri Shaham
摘要:虽然信号转换和分解表示学习已经显示出在音频、图像和多模态生成等领域操纵数据属性的前景,但现有方法,特别是语音风格转换方法,在很大程度上是经验性的,缺乏严格的理论基础来保证可靠和可解释的控制。在这项工作中,我们提出了一个语音属性转换的一般框架,伴随着理论分析和合理的假设下的保证。我们的框架建立在一个非概率自动编码器架构与预测的潜在变量和目标可控变量之间的独立约束。这种设计确保了一致的信号转换,以观察到的样式变量为条件,同时保留原始内容并修改所需的属性。我们进一步证明了我们的方法的多功能性,通过评估它的讲话风格,包括扬声器的身份和情感。定量评价证实了所提出的方法的有效性和普遍性。
摘要:While signal conversion and disentangled representation learning have shown promise for manipulating data attributes across domains such as audio, image, and multimodal generation, existing approaches, especially for speech style conversion, are largely empirical and lack rigorous theoretical foundations to guarantee reliable and interpretable control. In this work, we propose a general framework for speech attribute conversion, accompanied by theoretical analysis and guarantees under reasonable assumptions. Our framework builds on a non-probabilistic autoencoder architecture with an independence constraint between the predicted latent variable and the target controllable variable. This design ensures a consistent signal transformation, conditioned on an observed style variable, while preserving the original content and modifying the desired attribute. We further demonstrate the versatility of our method by evaluating it on speech styles, including speaker identity and emotion. Quantitative evaluations confirm the effectiveness and generality of the proposed approach.


【16】TokenChain: A Discrete Speech Chain via Semantic Token Modeling
标题:TokenChain:通过语义代币建模的离散语音链
链接:https://arxiv.org/abs/2510.06201

作者:Mingxuan Wang, Satoshi Nakamura
备注:5 pages, 3 figures. Submitted to IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026
摘要:机器语音链,模拟人类的感知-生产循环,证明有效地联合改善ASR和TTS。我们提出了TokenChain,这是一个完全离散的语音链,它将语义令牌ASR与两阶段TTS耦合在一起:一个与ASR共同训练的自回归文本到语义模型和一个仅用于合成的掩蔽生成语义到声学模型。整个文本界面的端到端反馈通过直通argmax/Gumbel-Softmax实现,并通过动态权重平均与监督ASR平衡。烧蚀检查最佳的温度时间表内和跨域转移。评估显示,TokenChain在LibriSpeech上的T2 S稳定,比基线精度提前2-6个epoch,等时误差降低5-13%,在TED-LIUM上的相对ASR WER降低56%,T2 S WER降低31%,遗忘最小,表明链学习对令牌接口和模型仍然有效。
摘要:Machine Speech Chain, simulating the human perception-production loop, proves effective in jointly improving ASR and TTS. We propose TokenChain, a fully discrete speech chain coupling semantic-token ASR with a two-stage TTS: an autoregressive text-to-semantic model co-trained with ASR and a masked-generative semantic-to-acoustic model for synthesis only. End-to-end feedback across the text interface is enabled with straight-through argmax/Gumbel-Softmax and balanced with supervised ASR via dynamic weight averaging. Ablations examine optimal temperature schedules for in- and cross-domain transfer. Evaluation reveals TokenChain surpasses baseline accuracy 2-6 epochs earlier and yields 5-13% lower equal-epoch error with stable T2S on LibriSpeech, and reduces relative ASR WER by 56% and T2S WER by 31% on TED-LIUM with minimal forgetting, showing that chain learning remains effective with token interfaces and models.


eess.AS音频处理


【1】TokenChain: A Discrete Speech Chain via Semantic Token Modeling
标题:TokenChain:通过语义代币建模的离散语音链
链接:https://arxiv.org/abs/2510.06201

作者:Mingxuan Wang, Satoshi Nakamura
备注:5 pages, 3 figures. Submitted to IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026
摘要:机器语音链,模拟人类的感知-生产循环,证明有效地联合改善ASR和TTS。我们提出了TokenChain,这是一个完全离散的语音链,它将语义令牌ASR与两阶段TTS耦合在一起:一个与ASR共同训练的自回归文本到语义模型和一个仅用于合成的掩蔽生成语义到声学模型。整个文本界面的端到端反馈通过直通argmax/Gumbel-Softmax实现,并通过动态权重平均与监督ASR平衡。烧蚀检查最佳的温度时间表内和跨域转移。评估显示,TokenChain在LibriSpeech上的T2 S稳定,比基线精度提前2-6个epoch,等时误差降低5-13%,在TED-LIUM上的相对ASR WER降低56%,T2 S WER降低31%,遗忘最小,表明链学习对令牌接口和模型仍然有效。
摘要:Machine Speech Chain, simulating the human perception-production loop, proves effective in jointly improving ASR and TTS. We propose TokenChain, a fully discrete speech chain coupling semantic-token ASR with a two-stage TTS: an autoregressive text-to-semantic model co-trained with ASR and a masked-generative semantic-to-acoustic model for synthesis only. End-to-end feedback across the text interface is enabled with straight-through argmax/Gumbel-Softmax and balanced with supervised ASR via dynamic weight averaging. Ablations examine optimal temperature schedules for in- and cross-domain transfer. Evaluation reveals TokenChain surpasses baseline accuracy 2-6 epochs earlier and yields 5-13% lower equal-epoch error with stable T2S on LibriSpeech, and reduces relative ASR WER by 56% and T2S WER by 31% on TED-LIUM with minimal forgetting, showing that chain learning remains effective with token interfaces and models.


【2】Revisiting Modeling and Evaluation Approaches in Speech Emotion Recognition: Considering Subjectivity of Annotators and Ambiguity of Emotions
标题:重新审视语音情感识别中的建模和评估方法:考虑注释者的主观性和情感的模糊性
链接:https://arxiv.org/abs/2510.05934

作者:Huang-Cheng Chou, Chi-Chun Lee
备注:PhD Thesis; ACLCLP Doctoral Dissertation Award -- Honorable Mention
摘要:在过去的二十年里,语音情感识别(SER)受到了越来越多的关注。为了训练SER系统,研究人员收集了由众包或内部评分员从预定义的类别中选择情绪进行注释的情绪语音数据库。然而,评分员之间的分歧是常见的。传统的方法将这些分歧视为噪音,将标签聚合成单个共识目标。虽然这将SER简化为单标签任务,但它忽略了人类情感感知的固有主观性。本文对这些假设提出了挑战,并提出了以下问题:(1)少数民族的情感评级应该被抛弃吗?(2)SER系统应该只从少数人的感知中学习吗?(3)SER系统是否应该对每个样本只预测一种情绪?   心理学研究表明,情绪感知具有主观性和模糊性,具有重叠的情绪边界。我们提出了新的建模和评估观点:(1)保留所有的情感评级,并表示它们与软标签分布。模型在单个注释者评级上进行训练,并与标准SER系统联合优化,提高了共识标记测试的性能。(2)重新定义SER评估,包括所有情绪数据,并允许共同出现的情绪(例如,悲伤和愤怒)。我们提出了一个“包罗万象的规则”,聚合所有的评级,以最大限度地提高标签表示的多样性。在四个英语情感数据库上的实验表明,大多数和复数标记的性能优越。(3)构建一个惩罚矩阵,以阻止训练过程中不太可能的情绪组合。将其集成到损失函数中进一步提高了性能。总体而言,包含少数评级,多个注释者和多情感预测会产生更强大和更人性化的SER系统。
摘要:Over the past two decades, speech emotion recognition (SER) has received growing attention. To train SER systems, researchers collect emotional speech databases annotated by crowdsourced or in-house raters who select emotions from predefined categories. However, disagreements among raters are common. Conventional methods treat these disagreements as noise, aggregating labels into a single consensus target. While this simplifies SER as a single-label task, it ignores the inherent subjectivity of human emotion perception. This dissertation challenges such assumptions and asks: (1) Should minority emotional ratings be discarded? (2) Should SER systems learn from only a few individuals' perceptions? (3) Should SER systems predict only one emotion per sample?   Psychological studies show that emotion perception is subjective and ambiguous, with overlapping emotional boundaries. We propose new modeling and evaluation perspectives: (1) Retain all emotional ratings and represent them with soft-label distributions. Models trained on individual annotator ratings and jointly optimized with standard SER systems improve performance on consensus-labeled tests. (2) Redefine SER evaluation by including all emotional data and allowing co-occurring emotions (e.g., sad and angry). We propose an ``all-inclusive rule'' that aggregates all ratings to maximize diversity in label representation. Experiments on four English emotion databases show superior performance over majority and plurality labeling. (3) Construct a penalization matrix to discourage unlikely emotion combinations during training. Integrating it into loss functions further improves performance. Overall, embracing minority ratings, multiple annotators, and multi-emotion predictions yields more robust and human-aligned SER systems.


【3】Revisiting MFCCs: Evidence for Spectral-Prosodic Coupling
标题:重温MFCC:频谱-韵律耦合的证据
链接:https://arxiv.org/abs/2510.05922

作者:Vitor Magno de O. S. Bezerra, Gabriel F. A. Bastos, Jugurta Montalvão
备注:5 pages, 3 figures, ISCMI 2025
摘要:梅尔倒谱系数(MFCC)是语音处理中的一个重要特征。更深入地了解它们的属性可以有助于经典和深度学习模型的工作。本研究挑战了长期以来的假设,即MFCC缺乏相关的时间信息,通过调查它们与语音韵律的关系。使用一个零假设的显着性检验框架,系统的评估MFCC和三个韵律特征:能量,基频(F0),和清化之间的统计独立性。结果表明,这是统计上令人难以置信的MFCC是独立的这三个韵律特征。这一发现表明,MFCC固有地携带有价值的韵律信息,这可以为未来的语音分析和识别模型的设计提供信息。
摘要:Mel-frequency cepstral coefficients (MFCCs) are an important feature in speech processing. A deeper understanding of their properties can contribute to the work that is being done with both classical and deep learning models. This study challenges the long-held assumption that MFCCs lack relevant temporal information by investigating their relationship with speech prosody. Using a null hypothesis significance testing framework, a systematic assessment is made about the statistical independence between MFCCs and the three prosodic features: energy, fundamental frequency (F0), and voicing. The results demonstrate that it is statistically implausible that the MFCCs are independent of any of these three prosodic features. This finding suggests that MFCCs inherently carry valuable prosodic information, which can inform the design of future models in speech analysis and recognition.


【4】Neural Forward Filtering for Speaker-Image Separation
标题:语音图像分离的神经前向过滤
链接:https://arxiv.org/abs/2510.05757

作者:Jingqi Sun, Shulin He, Ruizhe Pang, Zhong-Qiu Wang
备注:in submission
摘要:我们解决了混响条件下的单声道多扬声器图像分离,旨在分离混合扬声器,但保留每个扬声器的混响。一个简单的方法是直接训练端到端DNN系统,以预测每个扬声器的混响语音的基础上输入的混合。虽然有效,但该方法没有明确地利用混响语音可以通过将直接路径信号与线性滤波器卷积来再现的物理约束。为了解决这个问题,我们提出了CxNet,这是一个两个DNN系统,中间有一个神经前向滤波模块。第一DNN被训练为联合预测直接路径信号和混响语音。基于直接路径估计,神经前向滤波模块估计线性滤波器,然后将估计的滤波器与直接路径估计卷积以获得混响语音的另一估计,该估计被用作区分特征以帮助第二DNN更好地估计混响语音。通过对线性滤波器进行显式建模,CxNet可以利用直达路径信号和混响语音之间的物理约束来捕获有关混响尾部的关键信息。在SMS-WSJ数据集上的测试结果表明了算法的有效性。
摘要:We address monaural multi-speaker-image separation in reverberant conditions, aiming at separating mixed speakers but preserving the reverberation of each speaker. A straightforward approach for this task is to directly train end-to-end DNN systems to predict the reverberant speech of each speaker based on the input mixture. Although effective, this approach does not explicitly exploit the physical constraint that reverberant speech can be reproduced by convolving the direct-path signal with a linear filter. To address this, we propose CxNet, a two-DNN system with a neural forward filtering module in between. The first DNN is trained to jointly predict the direct-path signal and reverberant speech. Based on the direct-path estimate, the neural forward filtering module estimates the linear filter, and the estimated filter is then convolved with the direct-path estimate to obtain another estimate of reverberant speech, which is utilized as a discriminative feature to help the second DNN better estimate the reverberant speech. By explicitly modeling the linear filter, CxNet could leverage the physical constraint between the direct-path signal and reverberant speech to capture crucial information about reverberation tails. Evaluation results on the SMS-WSJ dataset show the effectiveness of the proposed algorithms.


【5】Investigation of perception inconsistency in speaker embedding for asynchronous voice anonymization
标题:用于同步语音匿名化的说话人嵌入中感知不一致性的研究
链接:https://arxiv.org/abs/2510.05718

作者:Rui Wang, Liping Chen, Kong Aik Lee, Zhengpeng Zha, Zhenhua Ling
摘要:给定用嵌入向量表示说话人属性的语音生成框架,可以通过修改从原始语音导出的说话人嵌入来实现异步语音匿名化。然而,机器和人类之间的不一致的感知扬声器属性内的扬声器嵌入仍然未被探索,限制其性能在异步语音匿名化。为此,本研究通过修改语音生成过程中的扬声器嵌入调查这种不一致。在FACodec和Diff-HierVC语音生成模型上进行的实验发现了一个子空间,该子空间的去除改变了机器感知,同时保留了所生成语音中说话者属性的人类感知。利用这些发现,开发了异步语音匿名化,实现了100%的人类感知保留率,同时模糊了机器感知。音频样本可以在https://voiceprivacy.github.io/speaker-embedding-eigen-decomposition/上找到。
摘要:Given the speech generation framework that represents the speaker attribute with an embedding vector, asynchronous voice anonymization can be achieved by modifying the speaker embedding derived from the original speech. However, the inconsistency between machine and human perceptions of the speaker attribute within the speaker embedding remains unexplored, limiting its performance in asynchronous voice anonymization. To this end, this study investigates this inconsistency via modifications to speaker embedding in the speech generation process. Experiments conducted on the FACodec and Diff-HierVC speech generation models discover a subspace whose removal alters machine perception while preserving its human perception of the speaker attribute in the generated speech. With these findings, an asynchronous voice anonymization is developed, achieving 100% human perception preservation rate while obscuring the machine perception. Audio samples can be found in https://voiceprivacy.github.io/speaker-embedding-eigen-decomposition/.


【6】Teaching Machines to Speak Using Articulatory Control
标题:使用发音控制教机器说话
链接:https://arxiv.org/abs/2510.05619

作者:Akshay Anand, Chenxu Guo, Cheol Jun Cho, Jiachen Lian, Gopala Anumanchipalli
摘要:当前的语音产生系统主要依赖于作为黑盒操作的大型Transformer模型,在人类语音的物理机制中提供很少的可解释性或基础。我们提出了一个新的框架:通过明确的发音控制语音生成解决这个限制。这将语音重新定义为类似于机器人操作的运动控制任务。我们的方法使用强化学习来训练一种直接控制声道发音器官(如舌头、嘴唇和下巴)运动的策略,以产生音节级别的语音。具体来说,我们采用近端策略优化算法来学习最佳发音运动的基础上提供的声音反馈,我们的音频感知器,Sylber。由此产生的发音轨迹使用预先训练的发音到语音解码器(ARM2.0)解码成音频。我们在六个目标音节上训练这个框架,它展示了成功的收敛,策略生成的音频和目标音节之间的相似性得分超过0.85。人类对诸如“please”、“loot”和“cat”等音节的音频的准确转录证明了该框架的可理解性。
摘要:Current speech production systems predominantly rely on large transformer models that operate as black boxes, providing little interpretability or grounding in the physical mechanisms of human speech. We address this limitation by proposing a new framework: speech generation through explicit articulatory control. This reframes speech as a motor control task similar to robotic manipulation. Our approach uses reinforcement learning to train a policy that directly controls the movements of vocal tract articulators, such as the tongue, lips, and jaw, to produce syllable-level speech. Specifically, we employ the Proximal Policy Optimization algorithm to learn optimal articulatory movements based on acoustic feedback provided by our audio perceiver, Sylber. The resulting articulatory trajectories are decoded into audio using SPARC, a pre-trained articulatory-to-speech decoder. We train this framework on six target syllables, and it demonstrates successful convergence, with similarity scores between the policy-generated audio and the target syllables exceeding 0.85. Accurate human transcription of the audio for syllables such as "please", "loot", and "cat" demonstrates the intelligibility of this framework.


【7】AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning
标题:AQA-TTRL:通过测试时强化学习实现音频问题回答的自适应
链接:https://arxiv.org/abs/2510.05478

作者:Haoyu Zhang, Jiaxian Guo, Yusuke Iwasawa, Yutaka Matsuo
备注:5 pages, 4 figures, Submitted to ICASSP 2026
摘要:大型音频语言模型(LALM)展示了令人印象深刻的一般音频理解,但一旦部署,它们是静态的,无法通过新的真实世界音频数据进行改进。由于传统的监督微调是昂贵的,我们介绍了一种新的框架,测试时的音频理解,AQA-TTRL,其中LALM演变的飞行只使用未标记的测试数据。它首先通过多数投票从预测中生成伪标签,然后通过强化学习优化模型。为了处理这些自生成标签中的固有噪声,我们引入了一种基于置信度的加权方法来调整训练信号。此外,多次尝试采样操作减轻了优势崩溃并稳定了训练。在MMAU(test-mini/test)、MMAR和MMSU基准测试中,AQA-TTRL在Qwen2.5-Omni 7 B模型和3B模型中分别实现了4.42%和11.04%的显著平均改进。值得注意的是,适配的3B模型始终优于未适配的7 B模型的直接推断,突出了先前未探索的测试时间适配在音频理解中的有效性。
摘要:Large Audio Language Models (LALMs) demonstrate impressive general audio understanding, but once deployed, they are static and fail to improve with new real-world audio data. As traditional supervised fine-tuning is costly, we introduce a novel framework for test-time audio understanding, AQA-TTRL, where an LALM evolves on-the-fly using only unlabeled test data. It first generates pseudo-labels from the prediction via majority voting, then optimizes the model via reinforcement learning. To handle the inherent noise in these self-generated labels, we introduce a confidence-based weighting method to adjust training signals. Furthermore, a multiple-attempt sampling operation mitigates advantage collapse and stabilizes training. On the MMAU (test-mini/test), MMAR, and MMSU benchmarks, AQA-TTRL achieves significant average improvements of 4.42% for the Qwen2.5-Omni 7B model and 11.04% for the 3B model. Notably, the adapted 3B model consistently outperforms the direct inference of the unadapted 7B model, highlighting the effectiveness of previously unexplored test-time adaptations in audio understanding.


【8】WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection
标题:WaveSP-Net:用于语音深度伪造检测的可学习波域稀疏提示调整
链接:https://arxiv.org/abs/2510.05305

作者:Xi Xuan, Xuechen Liu, Wenxin Zhang, Yi-Cheng Lin, Xiaojian Lin, Tomi Kinnunen
备注:Submitted to ICASSP 2026
摘要:语音深度伪造检测的现代前端设计依赖于对XLSR等大型预训练模型的全面微调。然而,这种方法不是参数有效的,可能会导致次优的泛化到现实的,在野生数据类型。为了解决这些局限性,我们引入了一个新的家庭的参数高效的前端,融合了传统的信号处理变换的调谐。这些包括使用傅立叶变换的FourierPT-XLSR,以及基于小波变换的两个变体:WSPT-XLSR和Partial-WSPT-XLSR。我们进一步提出了WaveSP-Net,这是一种结合了Partial-WSPT-XLSR前端和基于双向Mamba的后端的新型架构。该设计将多分辨率特征注入提示嵌入,这增强了细微合成伪影的定位,而不改变冻结的XLSR参数。实验结果表明,WaveSP-Net在两个新的具有挑战性的基准测试Deepfake-Eval-2024和SpoofCeleb上的性能优于几种最先进的模型,具有较低的可训练参数和显着的性能提升。代码和模型可在https://github.com/xxuan-acoustics/WaveSP-Net上获得。
摘要:Modern front-end design for speech deepfake detection relies on full fine-tuning of large pre-trained models like XLSR. However, this approach is not parameter-efficient and may lead to suboptimal generalization to realistic, in-the-wild data types. To address these limitations, we introduce a new family of parameter-efficient front-ends that fuse prompt-tuning with classical signal processing transforms. These include FourierPT-XLSR, which uses the Fourier Transform, and two variants based on the Wavelet Transform: WSPT-XLSR and Partial-WSPT-XLSR. We further propose WaveSP-Net, a novel architecture combining a Partial-WSPT-XLSR front-end and a bidirectional Mamba-based back-end. This design injects multi-resolution features into the prompt embeddings, which enhances the localization of subtle synthetic artifacts without altering the frozen XLSR parameters. Experimental results demonstrate that WaveSP-Net outperforms several state-of-the-art models on two new and challenging benchmarks, Deepfake-Eval-2024 and SpoofCeleb, with low trainable parameters and notable performance gains. The code and models are available at https://github.com/xxuan-acoustics/WaveSP-Net.


【9】Modulation Discovery with Differentiable Digital Signal Processing
标题:利用差异数字信号处理实现调制发现
链接:https://arxiv.org/abs/2510.06204

作者:Christopher Mitcheltree, Hao Hao Tan, Joshua D. Reiss
备注:Accepted to WASPAA 2025 (best paper award candidate). Code, audio samples, and plugins can be found at this https URL
摘要:调制是声音设计和音乐制作的关键部分,可以创建复杂和不断发展的音频。现代频率合成器提供包络、低频振荡器(LFO)和更多的参数自动化工具,使用户能够轻松调制输出。然而,确定用于创建声音的调制信号是困难的,并且现有的声音匹配/参数估计系统通常是不可解释的黑箱或预测高维逐帧参数值,而不考虑底层调制曲线的形状、结构和路由。我们提出了一种神经声音匹配的方法,利用调制提取,约束控制信号参数化,和微分数字信号处理(DDSP),发现调制存在于一个声音。我们证明了我们的方法对高度调制的合成和真实音频样本的有效性,它适用于不同的DDSP合成器架构,并调查它在可解释性和声音匹配精度之间的权衡。我们提供代码和音频样本,并在VST插件中提供经过训练的DDSP合成器。
摘要:Modulations are a critical part of sound design and music production, enabling the creation of complex and evolving audio. Modern synthesizers provide envelopes, low frequency oscillators (LFOs), and more parameter automation tools that allow users to modulate the output with ease. However, determining the modulation signals used to create a sound is difficult, and existing sound-matching / parameter estimation systems are often uninterpretable black boxes or predict high-dimensional framewise parameter values without considering the shape, structure, and routing of the underlying modulation curves. We propose a neural sound-matching approach that leverages modulation extraction, constrained control signal parameterizations, and differentiable digital signal processing (DDSP) to discover the modulations present in a sound. We demonstrate the effectiveness of our approach on highly modulated synthetic and real audio samples, its applicability to different DDSP synth architectures, and investigate the trade-off it incurs between interpretability and sound-matching accuracy. We make our code and audio samples available and provide the trained DDSP synths in a VST plugin.


【10】Latent Speech-Text Transformer
标题:潜在的语音-文本转换器Transformer
链接:https://arxiv.org/abs/2510.06195

作者:Yen-Ju Lu, Yashesh Gaur, Wei Zhou, Benjamin Muller, Jesus Villalba, Najim Dehak, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Srinivasan Iyer, Duc Le
备注:16 pages, 13 figures
摘要:自回归语音-文本模型通常在大量的文本令牌的交织序列上进行预训练,并且使用矢量量化将原始语音编码为语音令牌。这些模型在语音到语音理解和生成基准方面表现出了最先进的性能,以及有前途的缩放律,主要是通过文本和语音之间的代表性对齐实现的。然而,他们遭受的缺点,部分原因是不成比例的较长序列的语音令牌相比,文本令牌。这导致在预训练期间以及在推理期间模态之间的大的计算不平衡,以及有效地对齐语音和文本的潜在障碍,最终转化为几个数量级的较慢的缩放律。我们引入了潜在的语音文本Transformer(LST),它通过动态和廉价地将语音令牌聚合到潜在的语音补丁中,使预训练的语音文本模型更具数据效率。这些补丁作为更高级别的单元,可以与相应的文本单元对齐以帮助能力转移,甚至可以封装常见的语音序列,如沉默,以提高计算效率。我们表明,LST优于香草的方法在语音到语音以及文本到文本的基准在数据和计算机控制的设置,前者表明更有效的代表性对齐和后者表明更陡峭的缩放律语音文本模型。在HellaSwag故事完成时,LST在计算机控制的训练下实现了6.5%的语音准确率绝对增益,在数据控制的训练下实现了5.3%的语音准确率绝对增益,同时还提高了文本性能。我们将发布我们的模型,代码和评估数据,以方便进一步的研究。
摘要:Auto-regressive speech-text models are typically pre-trained on a large number of interleaved sequences of text tokens and raw speech encoded as speech tokens using vector quantization. These models have demonstrated state-of-the-art performance in speech-to-speech understanding and generation benchmarks, together with promising scaling laws, primarily enabled by the representational alignment between text and speech. Nevertheless, they suffer from shortcomings, partly owing to the disproportionately longer sequences of speech tokens in contrast to textual tokens. This results in a large compute imbalance between modalities during pre-training as well as during inference, and a potential hindrance to effectively aligning speech and text, ultimately translating to several orders of magnitude slower scaling laws. We introduce the Latent Speech-Text Transformer (LST), which makes pre-training speech-text models more data-efficient by dynamically and inexpensively aggregating speech tokens into latent speech patches. These patches serve as higher-level units that can either align with corresponding textual units to aid capability transfer or even encapsulate common speech sequences like silences to be more compute-efficient. We show that LST outperforms vanilla approaches on speech-to-speech as well as text-to-text benchmarks in both data- and compute-controlled settings, the former indicating more effective representational alignment and the latter indicating steeper scaling laws for speech-text models. On HellaSwag story completion, LST achieves 6.5% absolute gain in speech accuracy under compute-controlled training and 5.3% under data-controlled training, while also improving text performance. We will release our models, code, and the evaluation data to facilitate further research.


【11】ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
标题:ECTSpeech:通过简单的一致性调整增强高效的语音合成
链接:https://arxiv.org/abs/2510.05984

作者:Tao Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun, Wendong Zheng
备注:Accepted for publication by Proceedings of the 2025 ACM Multimedia Asia Conference(MMAsia '25)
摘要:扩散模型在语音合成中表现出显著的性能,但通常需要多步采样,导致推理效率低。最近的研究通过将扩散模型蒸馏成一致性模型来解决这个问题,从而实现高效的一步生成。然而,这些方法引入了额外的培训成本,并严重依赖于预先培训的教师模型的性能。在本文中,我们提出了ECTSpeech,一个简单而有效的一步语音合成框架,第一次,将容易的一致性调整(ECT)策略到语音合成。通过逐步收紧对预训练扩散模型的一致性约束,ECTSpeech实现了高质量的一步生成,同时显着降低了训练复杂性。此外,我们设计了一个多尺度门模块(MSGate),以提高去噪器的能力,融合功能在不同的尺度。LJSpeech数据集上的实验结果表明,ECTSpeech在单步采样下实现了与最先进方法相当的音频质量,同时大大降低了模型的训练成本和复杂性。
摘要:Diffusion models have demonstrated remarkable performance in speech synthesis, but typically require multi-step sampling, resulting in low inference efficiency. Recent studies address this issue by distilling diffusion models into consistency models, enabling efficient one-step generation. However, these approaches introduce additional training costs and rely heavily on the performance of pre-trained teacher models. In this paper, we propose ECTSpeech, a simple and effective one-step speech synthesis framework that, for the first time, incorporates the Easy Consistency Tuning (ECT) strategy into speech synthesis. By progressively tightening consistency constraints on a pre-trained diffusion model, ECTSpeech achieves high-quality one-step generation while significantly reducing training complexity. In addition, we design a multi-scale gate module (MSGate) to enhance the denoiser's ability to fuse features at different scales. Experimental results on the LJSpeech dataset demonstrate that ECTSpeech achieves audio quality comparable to state-of-the-art methods under single-step sampling, while substantially reducing the model's training cost and complexity.


【12】Segment-Factorized Full-Song Generation on Symbolic Piano Music
标题:象征性钢琴音乐的分段分解全曲生成
链接:https://arxiv.org/abs/2510.05881

作者:Ping-Yi Chen, Chih-Pin Tan, Yi-Hsuan Yang
备注:Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: AI for Music
摘要:我们提出了分段全歌模型(SFS)的符号全歌生成。该模型接受用户提供的歌曲结构和可选的短种子片段,该片段锚定歌曲开发的主要思想。通过将歌曲分解为片段并通过选择性注意相关片段来生成每个片段,该模型与先前的工作相比实现了更高的质量和效率。为了证明它对人类与人工智能交互的适用性,我们进一步将SFS包装到一个Web应用程序中,使用户能够以可定制的结构和灵活的顺序在钢琴卷上迭代地共同创作音乐。
摘要:We propose the Segmented Full-Song Model (SFS) for symbolic full-song generation. The model accepts a user-provided song structure and an optional short seed segment that anchors the main idea around which the song is developed. By factorizing a song into segments and generating each one through selective attention to related segments, the model achieves higher quality and efficiency compared to prior work. To demonstrate its suitability for human-AI interaction, we further wrap SFS into a web application that enables users to iteratively co-create music on a piano roll with customizable structures and flexible ordering.


【13】FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
标题:FoleyFinder:使用与GRAM对齐的多模式编码器的视频到音频生成
链接:https://arxiv.org/abs/2510.05829

作者:Riccardo Fosco Gramaccioni, Christian Marinoni, Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello
备注:Acepted at IJCNN 2025
摘要:在这项工作中,我们提出了Foleystrike,一种新的方法来视频到音频的生成,强调通过使用对齐的多模态编码器的语义条件。基于视频到音频生成的先前进步,Foleywalk利用Gramian表示对齐度量(RMM)来对齐视频,文本和音频模态的嵌入,从而实现对音频生成过程的精确语义控制。Foleystrike的核心是一个基于扩散的音频合成模型,以GRAM对齐的嵌入和波形包络为条件,确保语义丰富性和与相应输入视频的时间对齐。我们在Greatest Hits数据集上评估了Foleywatch,这是视频到音频模型的标准基准。我们的实验表明,对齐多模式编码器使用的MPEG增强了系统的能力,语义上对齐生成的音频与视频内容,推进了视频到音频合成的最新技术。
摘要:In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM leverages the Gramian Representation Alignment Measure (GRAM) to align embeddings across video, text, and audio modalities, enabling precise semantic control over the audio generation process. The core of FoleyGRAM is a diffusion-based audio synthesis model conditioned on GRAM-aligned embeddings and waveform envelopes, ensuring both semantic richness and temporal alignment with the corresponding input video. We evaluate FoleyGRAM on the Greatest Hits dataset, a standard benchmark for video-to-audio models. Our experiments demonstrate that aligning multimodal encoders using GRAM enhances the system's ability to semantically align generated audio with video content, advancing the state of the art in video-to-audio synthesis.


【14】StereoSync: Spatially-Aware Stereo Audio Generation from Video
标题:StereoLock:从视频中生成空间感知的立体声音频
链接:https://arxiv.org/abs/2510.05828

作者:Christian Marinoni, Riccardo Fosco Gramaccioni, Kazuki Shimada, Takashi Shibuya, Yuki Mitsufuji, Danilo Comminiello
备注:Accepted at IJCNN 2025
摘要:虽然近年来音频生成已被广泛研究,但视频对齐的音频生成仍然是一个相对未开发的前沿。为了解决这一差距,我们引入StereoSync,一种新颖而高效的模型,旨在生成与参考视频在时间上同步并与其视觉上下文在空间上对齐的音频。此外,StereoSync还通过利用预训练的基础模型来实现效率,减少了对大量训练的需求,同时保持高质量的合成。与主要关注时间同步的现有方法不同,StereoSync通过将空间感知纳入视频对齐的音频生成来引入显著的进步。事实上,给定输入视频,我们的方法从深度图和边界框中提取空间线索,将它们用作基于扩散的音频生成模型中的交叉注意调节。这种方法允许StereoSync超越简单的同步,产生动态适应视频场景的空间结构和移动的立体声音频。我们评估了StereoSync上的Walking The Maps,这是一个策展数据集,包括来自视频游戏的视频,这些视频游戏的特点是动画角色在不同的环境中行走。实验结果证明了StereoSync能够实现时间和空间对齐,推进了视频到音频生成的最新技术水平,并带来了更加身临其境和逼真的音频体验。
摘要:Although audio generation has been widely studied over recent years, video-aligned audio generation still remains a relatively unexplored frontier. To address this gap, we introduce StereoSync, a novel and efficient model designed to generate audio that is both temporally synchronized with a reference video and spatially aligned with its visual context. Moreover, StereoSync also achieves efficiency by leveraging pretrained foundation models, reducing the need for extensive training while maintaining high-quality synthesis. Unlike existing methods that primarily focus on temporal synchronization, StereoSync introduces a significant advancement by incorporating spatial awareness into video-aligned audio generation. Indeed, given an input video, our approach extracts spatial cues from depth maps and bounding boxes, using them as cross-attention conditioning in a diffusion-based audio generation model. Such an approach allows StereoSync to go beyond simple synchronization, producing stereo audio that dynamically adapts to the spatial structure and movement of a video scene. We evaluate StereoSync on Walking The Maps, a curated dataset comprising videos from video games that feature animated characters walking through diverse environments. Experimental results demonstrate the ability of StereoSync to achieve both temporal and spatial alignment, advancing the state of the art in video-to-audio generation and resulting in a significantly more immersive and realistic audio experience.


【15】Transcribing Rhythmic Patterns of the Guitar Track in Polyphonic Music
标题:复调音乐中吉他音轨节奏模式的转录
链接:https://arxiv.org/abs/2510.05756

作者:Aleksandr Lukoianov, Anssi Klapuri
备注:Accepted to WASPAA 2025
摘要:然而,在过去的几十年里,和弦转录受到了相当大的关注,很少有人致力于转录和编码歌曲中出现的节奏模式。这一主题与节奏吉他等乐器尤其相关,节奏吉他通常是通过重复和随时间变化的节奏模式来演奏的。然而,在许多情况下,人们不能客观地定义一个单一的“正确”的节奏模式,为一个给定的歌曲部分。为了创建一个具有明确定义的地面实况标签的数据集,我们请专业音乐家转录410首流行歌曲和唱片封面版本中的节奏模式,其中吉他曲目遵循这些转录。为了记录弦乐器及其相应的节奏模式,我们提出了一个三步框架。首先,我们执行近似干分离提取吉他部分从复调混合。其次,我们使用预训练的基础模型(MERT)作为骨干,在分离的吉他音频中检测单个琴弦。最后,我们进行了一个模式解码过程中,吉他弹奏的转录序列是由来自专家策划的词汇表的模式表示。我们表明,它是可以转录的节奏模式的吉他音轨在复调音乐具有相当高的准确性,产生一个表示,是人类可读的,包括自动检测的酒吧线和时间签名标记。我们进行消融研究和错误分析,并提出了一套评估指标,以评估预测的节奏模式序列的准确性和可读性。
摘要:Whereas chord transcription has received considerable attention during the past couple of decades, far less work has been devoted to transcribing and encoding the rhythmic patterns that occur in a song. The topic is especially relevant for instruments such as the rhythm guitar, which is typically played by strumming rhythmic patterns that repeat and vary over time. However, in many cases one cannot objectively define a single "right" rhythmic pattern for a given song section. To create a dataset with well-defined ground-truth labels, we asked expert musicians to transcribe the rhythmic patterns in 410 popular songs and record cover versions where the guitar tracks followed those transcriptions. To transcribe the strums and their corresponding rhythmic patterns, we propose a three-step framework. Firstly, we perform approximate stem separation to extract the guitar part from the polyphonic mixture. Secondly, we detect individual strums within the separated guitar audio, using a pre-trained foundation model (MERT) as a backbone. Finally, we carry out a pattern-decoding process in which the transcribed sequence of guitar strums is represented by patterns drawn from an expert-curated vocabulary. We show that it is possible to transcribe the rhythmic patterns of the guitar track in polyphonic music with quite high accuracy, producing a representation that is human-readable and includes automatically detected bar lines and time signature markers. We perform ablation studies and error analysis and propose a set of evaluation metrics to assess the accuracy and readability of the predicted rhythmic pattern sequence.


【16】Sci-Phi: A Large Language Model Spatial Audio Descriptor
标题:Sci-Phi:一种大型语言模型空间音频描述符
链接:https://arxiv.org/abs/2510.05542

作者:Xilin Jiang, Hannes Gamper, Sebastian Braun
摘要:声学场景感知包括描述声音的类型,它们的时间,它们的方向和距离,以及它们的响度和混响。虽然音频语言模型在声音识别方面表现出色,但单通道输入从根本上限制了空间理解。这项工作提出了Sci-Phi,这是一种具有双重空间和频谱编码器的空间音频大语言模型,可以为所有声源和周围环境估计完整的参数集。Sci-Phi从超过4,000小时的合成一阶Ambisonics录音(包括元数据)中学习,一次性列举并描述了多达四个定向声源,以及非定向背景声音和房间特征。我们使用置换不变协议和15个涵盖内容,位置,定时,响度和混响的度量来评估模型,并分析其在源计数,信噪比,混响水平以及具有挑战性的声学,空间或时间相似源的混合物中的鲁棒性。值得注意的是,Sci-Phi推广到实际房间脉冲响应,只有轻微的性能下降。总的来说,这项工作建立了第一个音频LLM能够完整的空间场景描述,具有强大的潜力,为现实世界的部署。演示:https://sci-phi-audio.github.io/demo
摘要:Acoustic scene perception involves describing the type of sounds, their timing, their direction and distance, as well as their loudness and reverberation. While audio language models excel in sound recognition, single-channel input fundamentally limits spatial understanding. This work presents Sci-Phi, a spatial audio large language model with dual spatial and spectral encoders that estimates a complete parameter set for all sound sources and the surrounding environment. Learning from over 4,000 hours of synthetic first-order Ambisonics recordings including metadata, Sci-Phi enumerates and describes up to four directional sound sources in one pass, alongside non-directional background sounds and room characteristics. We evaluate the model with a permutation-invariant protocol and 15 metrics covering content, location, timing, loudness, and reverberation, and analyze its robustness across source counts, signal-to-noise ratios, reverberation levels, and challenging mixtures of acoustically, spatially, or temporally similar sources. Notably, Sci-Phi generalizes to real room impulse responses with only minor performance degradation. Overall, this work establishes the first audio LLM capable of full spatial-scene description, with strong potential for real-world deployment. Demo: https://sci-phi-audio.github.io/demo


【17】Advancing Automated Spatio-Semantic Analysis in Picture Description Using Language Models
标题:使用语言模型推进图片描述中的自动空间语义分析
链接:https://arxiv.org/abs/2510.05128

作者:Si-Ioi Ng, Pranav S. Ambadi, Kimberly D. Mueller, Julie Liss, Visar Berisha
摘要:目前通过图片描述自动评估认知语言障碍的方法往往忽略了视觉叙事路径-说话者在图片中描述的元素的顺序和位置。空间语义特征的分析使用内容信息单元(CIU)捕获该路径,但是手动标记或基于字典的映射是劳动密集型的。本研究提出了一个基于BERT的管道,微调二进制交叉熵和成对排名损失,自动CIU提取和排序的Cookie盗窃图片描述。通过5倍交叉验证评估,它在CIU检测中达到93%的中位数精度,96%的中位数召回率和24%的序列错误率。所提出的方法提取的功能,表现出强烈的皮尔逊相关性与地面真相,超越了基于字典的基线在外部验证。这些特征还对通过ANCOVA评估组差异时从手动注释中获得的特征进行了验证。该管道被证明可以有效地表征认知障碍评估的视觉叙事路径,其实现和模型向公众开放。
摘要:Current methods for automated assessment of cognitive-linguistic impairment via picture description often neglect the visual narrative path - the sequence and locations of elements a speaker described in the picture. Analyses of spatio-semantic features capture this path using content information units (CIUs), but manual tagging or dictionary-based mapping is labor-intensive. This study proposes a BERT-based pipeline, fine tuned with binary cross-entropy and pairwise ranking loss, for automated CIU extraction and ordering from the Cookie Theft picture description. Evaluated by 5-fold cross-validation, it achieves 93% median precision, 96% median recall in CIU detection, and 24% sequence error rates. The proposed method extracts features that exhibit strong Pearson correlations with ground truth, surpassing the dictionary-based baseline in external validation. These features also perform comparably to those derived from manual annotations in evaluating group differences via ANCOVA. The pipeline is shown to effectively characterize visual narrative paths for cognitive impairment assessment, with the implementation and models open-sourced to public.


机器翻译由腾讯交互翻译提供,仅供参考