今日论文合集:cs.SD语音9篇,eess.AS音频处理7篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Content Anonymization for Privacy in Long-form Audio
标题:长篇音频中的隐私内容匿名化
链接:https://arxiv.org/abs/2510.12780

作者:Cristina Aggazzotti, Ashi Garg, Zexin Cai, Nicholas Andrews
摘要:语音匿名化技术已被发现成功地掩盖了说话者的声音身份在短,孤立的话语在基准,如语音隐私的挑战。然而,在实践中,话语很少孤立地发生:长格式音频在采访,电话和会议等领域中很常见。在这些情况下,来自同一个说话者的许多话语都是可用的,这带来了更大的隐私风险:给定来自同一个说话者的多个话语,攻击者可以利用个人的词汇,语法和短语来重新识别它们,即使他们的声音完全伪装。为了解决这一风险,我们提出了新的内容匿名化方法。我们的方法在ASR-TTS管道中执行转录的上下文重写,以消除特定于说话者的风格,同时保留意义。我们目前的结果在一个长形式的电话交谈设置演示基于内容的攻击语音匿名语音的有效性。然后,我们将展示如何提出基于内容的匿名化方法可以减轻这种风险,同时保持语音效用。总的来说,我们发现释义是一种有效的防御基于内容的攻击,并建议利益相关者采取这一步骤,以确保长格式音频的匿名性。
摘要:Voice anonymization techniques have been found to successfully obscure a speaker's acoustic identity in short, isolated utterances in benchmarks such as the VoicePrivacy Challenge. In practice, however, utterances seldom occur in isolation: long-form audio is commonplace in domains such as interviews, phone calls, and meetings. In these cases, many utterances from the same speaker are available, which pose a significantly greater privacy risk: given multiple utterances from the same speaker, an attacker could exploit an individual's vocabulary, syntax, and turns of phrase to re-identify them, even when their voice is completely disguised. To address this risk, we propose new content anonymization approaches. Our approach performs a contextual rewriting of the transcripts in an ASR-TTS pipeline to eliminate speaker-specific style while preserving meaning. We present results in a long-form telephone conversation setting demonstrating the effectiveness of a content-based attack on voice-anonymized speech. Then we show how the proposed content-based anonymization methods can mitigate this risk while preserving speech utility. Overall, we find that paraphrasing is an effective defense against content-based attacks and recommend that stakeholders adopt this step to ensure anonymity in long-form audio.


【2】Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
标题:Omni-Captioner:Omni详细感知的数据管道、模型和基准
链接:https://arxiv.org/abs/2510.12720

作者:Ziyang Ma, Ruiyang Xu, Zhenghao Xing, Yunfei Chu, Yuxuan Wang, Jinzheng He, Jin Xu, Pheng-Ann Heng, Kai Yu, Junyang Lin, Eng Siong Chng, Xie Chen
备注:his https URL
摘要:对多模态信息的细粒度感知对于推进人机交互至关重要。随着视听技术的最新进展,能够并行处理音频和视频信号的Omni Language Models(OLM)已经成为实现更丰富理解和推理的有前途的范例。然而,它们捕获和描述细粒度细节的能力仍然有限。在这项工作中,我们提出了一个系统的和全面的调查全方位的详细感知的数据管道,模型和基准的角度。我们首先确定了一个内在的“共同成长”的细节和幻觉在当前的OLMs。为了解决这个问题,我们提出了Omni-Detective,一个集成工具调用的代理数据生成管道,以自主生成高度详细但最小化幻觉的多模态数据。基于Omni-Detective生成的数据,我们训练了两个字幕模型:Audio-Captioner用于仅音频的详细感知,Omni-Captioner用于视听的详细感知。在级联评估协议下,Audio-Captioner在所有开源模型中实现了MMAU和MMAR上的最佳性能,超过了Gemini 2.5 Flash,并提供了与Gemini 2.5 Pro相当的性能。在现有的详细字幕基准上,Omni-Captioner在VDC上设置了一个新的最先进的水平,并在video-SALMONN 2测试集上实现了细节和幻觉之间的最佳平衡。由于缺乏一个专门的基准,全方位的细节感知,我们设计了全方位完形填空,一种新颖的完形填空风格的评估详细的音频,视频和视听字幕,确保稳定,高效,可靠的评估。实验结果和分析表明,Omni-Detective在生成高质量的详细字幕方面是有效的,Omni-Cloze在评价此类详细字幕方面具有优越性。
摘要:Fine-grained perception of multimodal information is critical for advancing human-AI interaction. With recent progress in audio-visual technologies, Omni Language Models (OLMs), capable of processing audio and video signals in parallel, have emerged as a promising paradigm for achieving richer understanding and reasoning. However, their capacity to capture and describe fine-grained details remains limited explored. In this work, we present a systematic and comprehensive investigation of omni detailed perception from the perspectives of the data pipeline, models, and benchmark. We first identify an inherent "co-growth" between detail and hallucination in current OLMs. To address this, we propose Omni-Detective, an agentic data generation pipeline integrating tool-calling, to autonomously produce highly detailed yet minimally hallucinatory multimodal data. Based on the data generated with Omni-Detective, we train two captioning models: Audio-Captioner for audio-only detailed perception, and Omni-Captioner for audio-visual detailed perception. Under the cascade evaluation protocol, Audio-Captioner achieves the best performance on MMAU and MMAR among all open-source models, surpassing Gemini 2.5 Flash and delivering performance comparable to Gemini 2.5 Pro. On existing detailed captioning benchmarks, Omni-Captioner sets a new state-of-the-art on VDC and achieves the best trade-off between detail and hallucination on the video-SALMONN 2 testset. Given the absence of a dedicated benchmark for omni detailed perception, we design Omni-Cloze, a novel cloze-style evaluation for detailed audio, visual, and audio-visual captioning that ensures stable, efficient, and reliable assessment. Experimental results and analysis demonstrate the effectiveness of Omni-Detective in generating high-quality detailed captions, as well as the superiority of Omni-Cloze in evaluating such detailed captions.


【3】TFGA-Net: Temporal-Frequency Graph Attention Network for Brain-Controlled Speaker Extraction
标题:TFGA-Net:用于脑控说话人提取的时频图注意力网络
链接:https://arxiv.org/abs/2510.12275

作者:Youhao Si, Yuan Liao, Qiushi Han, Yuhang Yang, Rui Dai, Liya Huang
备注:5 pages, 3 figures
摘要:基于脑电信号的听觉注意解码技术的迅速发展为脑电信号驱动的目标说话人提取提供了可能。然而,如何有效地利用脑电和语音之间的目标说话人共同信息仍然是一个悬而未决的问题。在本文中,我们提出了一个脑控制的说话人提取模型,它利用从听者记录的EEG提取目标语音。为了有效地从EEG信号中提取信息,我们推导出多尺度的时间-频率特征,并进一步将在任务期间选择性地参与的皮层拓扑结构。此外,为了有效地利用脑电信号的非欧结构和捕捉其全局特征,在脑电信号编码器中使用了图卷积网络和自注意机制。此外,为了充分利用融合的EEG和语音特征,保持全局上下文并捕获语音节奏和韵律,我们引入了MossFormer和RNN-Free Recurrent相结合的MossFormer 2作为分离器。本文在公开的Cocktail Party和KUL数据集上的实验结果表明,我们的TFGA-Net模型在某些客观评价指标上明显优于最先进的方法。源代码可从以下网址获得:https://github.com/LaoDa-X/TFGA-NET。
摘要:The rapid development of auditory attention decoding (AAD) based on electroencephalography (EEG) signals offers the possibility EEG-driven target speaker extraction. However, how to effectively utilize the target-speaker common information between EEG and speech remains an unresolved problem. In this paper, we propose a model for brain-controlled speaker extraction, which utilizes the EEG recorded from the listener to extract the target speech. In order to effectively extract information from EEG signals, we derive multi-scale time--frequency features and further incorporate cortical topological structures that are selectively engaged during the task. Moreover, to effectively exploit the non-Euclidean structure of EEG signals and capture their global features, the graph convolutional networks and self-attention mechanism are used in the EEG encoder. In addition, to make full use of the fused EEG and speech feature and preserve global context and capture speech rhythm and prosody, we introduce MossFormer2 which combines MossFormer and RNN-Free Recurrent as separator. Experimental results on both the public Cocktail Party and KUL dataset in this paper show that our TFGA-Net model significantly outper-forms the state-of-the-art method in certain objective evaluation metrics. The source code is available at: https://github.com/LaoDa-X/TFGA-NET.


【4】Not in Sync: Unveiling Temporal Bias in Audio Chat Models
标题:不同步:揭露音频聊天模型中的时间偏差
链接:https://arxiv.org/abs/2510.12185

作者:Jiayu Yao, Shenghua Liu, Yiwei Wang, Rundong Cheng, Lingrui Mei, Baolong Bi, Zhen Xiong, Xueqi Cheng
摘要:大型音频语言模型(LALM)越来越多地应用于音频理解和多模态推理,但它们在事件发生时的定位能力仍然有待研究。我们提出了第一个系统的研究LALM的时间偏差,揭示了他们的时间戳预测的一个关键限制。例如,当被问到“讲师在哪一秒引入关键公式?“,模型通常预测始终早于或晚于地面实况的时间戳。通过对时间戳数据集的对照实验,我们发现时间偏差(i)在数据集和模型中普遍存在,(ii)随着音频长度的增加而增加-甚至在延长的录音中积累到数十秒,(iii)在事件类型和位置之间变化。我们用时间偏差指数(TBI)来量化这种影响,测量预测事件时间的系统性偏差,并用可视化框架来补充它。我们的研究结果突出了当前LALM的根本局限性,并呼吁开发时间上鲁棒的架构。
摘要:Large Audio Language Models (LALMs) are increasingly applied to audio understanding and multimodal reasoning, yet their ability to locate when events occur remains underexplored. We present the first systematic study of temporal bias in LALMs, revealing a key limitation in their timestamp prediction. For example, when asked "At which second does the lecturer introduce the key formula?", models often predict timestamps that are consistently earlier or later than the ground truth. Through controlled experiments on timestamped datasets, we find that temporal bias (i) is prevalent across datasets and models, (ii) increases with audio length - even accumulating to tens of seconds in extended recordings, and (iii) varies across event types and positions. We quantify this effect with the Temporal Bias Index (TBI), measuring systematic misalignment in predicted event timings, and complement it with a visualization framework. Our findings highlight a fundamental limitation in current LALMs and call for the development of temporally robust architectures.


【5】Audio Palette: A Diffusion Transformer with Multi-Signal Conditioning for Controllable Foley Synthesis
标题:音频转换器:具有多信号调节的扩散Transformer,用于可控弗利合成
链接:https://arxiv.org/abs/2510.12175

作者:Junnuo Wang
备注:Accepted for publication in the Journal of Artificial Intelligence Research (JAIR), Vol. 3 No. 2, December 2025
摘要:基于扩散的生成模型的最新进展已经实现了高质量的文本到音频合成,但细粒度的声学控制仍然是开源研究中的一个重大挑战。我们提出了音频转换器,扩散Transformer(DiT)为基础的模型,扩展了稳定的音频开放架构,以解决这个“控制差距”在可控的音频生成。与先前仅依赖于语义条件反射的方法不同,音频控制引入了四个时变控制信号:响度、音高、频谱质心和音色,以精确和可解释地操纵声学特征。该模型在AudioSet的精选子集上使用低秩自适应(LoRA)有效地适应了福利合成的细微差别,只需要训练原始参数的0.85%。实验表明,音频编码器实现了对声音属性的细粒度、可解释的控制。至关重要的是,它实现了这种新颖的可控性,同时保持了高音频质量和与文本提示的强语义对齐,并且在Frechet音频距离(FAD)和LAION-CLAP评分等标准指标上的性能与原始基线模型相当。我们为音频研究提供了一个可扩展的模块化管道,强调基于序列的调节,内存效率和三尺度分类器的指导机制,用于细致入微的推理时间控制。这项工作为开源环境中的可控声音设计和表演性音频合成奠定了坚实的基础,实现了更加以艺术家为中心的工作流程。
摘要:Recent advances in diffusion-based generative models have enabled high-quality text-to-audio synthesis, but fine-grained acoustic control remains a significant challenge in open-source research. We present Audio Palette, a diffusion transformer (DiT) based model that extends the Stable Audio Open architecture to address this "control gap" in controllable audio generation. Unlike prior approaches that rely solely on semantic conditioning, Audio Palette introduces four time-varying control signals: loudness, pitch, spectral centroid, and timbre, for precise and interpretable manipulation of acoustic features. The model is efficiently adapted for the nuanced domain of Foley synthesis using Low-Rank Adaptation (LoRA) on a curated subset of AudioSet, requiring only 0.85 percent of the original parameters to be trained. Experiments demonstrate that Audio Palette achieves fine-grained, interpretable control of sound attributes. Crucially, it accomplishes this novel controllability while maintaining high audio quality and strong semantic alignment to text prompts, with performance on standard metrics such as Frechet Audio Distance (FAD) and LAION-CLAP scores remaining comparable to the original baseline model. We provide a scalable, modular pipeline for audio research, emphasizing sequence-based conditioning, memory efficiency, and a three-scale classifier-free guidance mechanism for nuanced inference-time control. This work establishes a robust foundation for controllable sound design and performative audio synthesis in open-source settings, enabling a more artist-centric workflow.


【6】UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
标题:UALM:用于理解、生成和推理的统一音频语言模型
链接:https://arxiv.org/abs/2510.12000

作者:Jinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh, Arushi Goel, Chao-Han Huck Yang, Wenliang Dai, Zihan Liu, Hanrong Ye, Shinji Watanabe, Mohammad Shoeybi, Bryan Catanzaro, Rafael Valle, Wei Ping
摘要:音频语言建模(ALM)领域的最新进展将音频理解和文本到音频生成作为单独的任务来处理。很少有研究试图将这些任务统一起来--这是迈向高级多模态推理的重要一步。本文介绍了通用音频语言模型(UALM),它旨在将音频理解、文本到音频生成和多模态推理统一在一个模型中。为了实现这一目标,我们首先提出了UALM-Gen,这是一种文本到音频的语言模型,可以直接预测音频标记,并且可以与最先进的基于扩散的模型相媲美。然后,我们使用适当的数据混合,训练配方和推理技术证明,我们的单一UALM模型在音频理解,文本到音频生成和文本推理方面与最先进的专业模型的质量相匹配。此外,我们提出了UALM-Reason,一种多模态推理模型,它在中间思维步骤中利用文本和音频来促进复杂的生成任务。据我们所知,这是音频研究中跨模态生成推理的第一次演示,其有效性得到了主观评价的证实。
摘要:Recent advances in the audio language modeling (ALM) domain tackle audio understanding and text-to-audio generation as separate tasks. Very few studies attempt to unify these tasks -- an essential step toward advanced multimodal reasoning. This paper introduces U}nified Audio Language Model (UALM), which aims to unify audio understanding, text-to-audio generation, and multimodal reasoning in a single model. To achieve this goal, we first present UALM-Gen, a text-to-audio language model that directly predicts audio tokens and is comparable to state-of-the-art diffusion-based models. We then demonstrate, using proper data blending, training recipes, and inference techniques, that our single UALM model matches the quality of state-of-the-art specialized models in audio understanding, text-to-audio generation, and text reasoning. Furthermore, we present UALM-Reason, a multimodal reasoning model that utilizes both text and audio in the intermediate thinking steps to facilitate complex generation tasks. To our knowledge, this is the first demonstration in audio research of cross-modal generative reasoning, with its effectiveness confirmed by subjective evaluations.


【7】Audio-Guided Visual Perception for Audio-Visual Navigation
标题:用于视听导航的音频引导视觉感知
链接:https://arxiv.org/abs/2510.11760

作者:Yi Wang, Yinfeng Yu, Fuchun Sun, Liejun Wang, Wendong Zheng
备注:Main paper (6 pages). Accepted for publication by International Conference on Virtual Reality and Visualization 2025 (ICVRV 2025)
摘要:视听导航的目的是使代理自主导航到声源在未知的3D环境中使用听觉线索。虽然目前的AVN方法在分布声源上表现出色,但它们表现出较差的跨源泛化能力:当代理遇到听不到的声音或看不见的环境时,导航成功率骤降,搜索路径变得过长。这种限制源于听觉信号和相应视觉区域之间缺乏明确的对齐机制。策略倾向于在训练过程中记住虚假的声音指纹相关性,从而导致在暴露于新声源时的盲目探索。为了解决这个问题,我们提出了AGVP框架,它将声音从政策难忘的声学指纹线索转化为空间指导。该框架首先通过音频自注意提取全局听觉上下文,然后使用此上下文作为查询来引导视觉特征注意,在特征级突出声源相关区域。然后执行后续的时间建模和策略优化。这种设计以可解释的跨模态对齐和区域重新加权为中心,减少了对特定声学指纹的依赖。实验结果表明,AGVP提高了导航效率和鲁棒性,同时实现了优越的跨场景概括以前闻所未闻的声音。
摘要:Audio-Visual Embodied Navigation aims to enable agents to autonomously navigate to sound sources in unknown 3D environments using auditory cues. While current AVN methods excel on in-distribution sound sources, they exhibit poor cross-source generalization: navigation success rates plummet and search paths become excessively long when agents encounter unheard sounds or unseen environments. This limitation stems from the lack of explicit alignment mechanisms between auditory signals and corresponding visual regions. Policies tend to memorize spurious \enquote{acoustic fingerprint-scenario} correlations during training, leading to blind exploration when exposed to novel sound sources. To address this, we propose the AGVP framework, which transforms sound from policy-memorable acoustic fingerprint cues into spatial guidance. The framework first extracts global auditory context via audio self-attention, then uses this context as queries to guide visual feature attention, highlighting sound-source-related regions at the feature level. Subsequent temporal modeling and policy optimization are then performed. This design, centered on interpretable cross-modal alignment and region reweighting, reduces dependency on specific acoustic fingerprints. Experimental results demonstrate that AGVP improves both navigation efficiency and robustness while achieving superior cross-scenario generalization on previously unheard sounds.


【8】SeeingSounds: Learning Audio-to-Visual Alignment via Text
标题:SeeingSounds:通过文本学习视听对齐
链接:https://arxiv.org/abs/2510.11738

作者:Simone Carnemolla, Matteo Pennisi, Chiara Russo, Simone Palazzo, Daniela Giordano, Concetto Spampinato
备注:accepted to ACM Multimedia Asia 2025
摘要:我们介绍SeeingSounds,一个轻量级的模块化框架,用于音频到图像的生成,利用音频,语言和视觉之间的相互作用,而不需要任何配对的视听数据或视觉生成模型的训练。我们的方法不是将音频视为文本的替代品或仅仅依赖于音频到文本的映射,而是执行双重对齐:音频通过冻结的语言编码器投射到语义语言空间中,并且使用视觉语言模型根据上下文接地到视觉域中。这种方法受到认知神经科学的启发,反映了在人类感知中观察到的自然跨模态关联。该模型在冻结的扩散骨干上运行,只训练轻量级适配器,从而实现高效和可扩展的学习。此外,它通过过程文本提示生成支持细粒度和可解释的控制,其中音频转换(例如,音量或音高变化)转换成描述性提示(例如,“遥远的雷声”)引导视觉输出。在标准基准测试中进行的大量实验证实,SeeingSounds在zero-shot和监督设置中的性能优于现有方法,在可控的音频到视频生成中建立了一个新的艺术状态。
摘要:We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on visual generative models. Rather than treating audio as a substitute for text or relying solely on audio-to-text mappings, our method performs dual alignment: audio is projected into a semantic language space via a frozen language encoder, and, contextually grounded into the visual domain using a vision-language model. This approach, inspired by cognitive neuroscience, reflects the natural cross-modal associations observed in human perception. The model operates on frozen diffusion backbones and trains only lightweight adapters, enabling efficient and scalable learning. Moreover, it supports fine-grained and interpretable control through procedural text prompt generation, where audio transformations (e.g., volume or pitch shifts) translate into descriptive prompts (e.g., "a distant thunder") that guide visual outputs. Extensive experiments across standard benchmarks confirm that SeeingSounds outperforms existing methods in both zero-shot and supervised settings, establishing a new state of the art in controllable audio-to-visual generation.


【9】Serial-Parallel Dual-Path Architecture for Speaking Style Recognition
标题:用于说话风格识别的串并行双路径架构
链接:https://arxiv.org/abs/2510.11732

作者:Guojian Li, Qijie Shao, Zhixian Zhao, Shuiyuan Wang, Zhonghua Fu, Lei Xie
备注:Accepted by NCMMSC2025
摘要:说话风格识别是从语音中识别说话人的说话风格特征。现有的风格识别方法主要依赖于语言信息,对声学信息的整合有限,这限制了识别精度的提高。声学和语言模态的融合提供了显著的潜力,以提高识别性能。在本文中,我们提出了一种新的串行-并行双路径架构SSR,利用声学语言双峰信息。串行路径遵循ASR+STYLE串行范式,反映了顺序的时间依赖性,而并行路径集成了我们设计的声学语言相似性模块(ALSM),以促进跨模态交互与时间的一致性。与现有的SSR基线-OSUM模型相比,我们的方法减少了88.4%的参数大小,并实现了30.3%的提高SSR的准确性为测试集上的八种风格。
摘要:Speaking Style Recognition (SSR) identifies a speaker's speaking style characteristics from speech. Existing style recognition approaches primarily rely on linguistic information, with limited integration of acoustic information, which restricts recognition accuracy improvements. The fusion of acoustic and linguistic modalities offers significant potential to enhance recognition performance. In this paper, we propose a novel serial-parallel dual-path architecture for SSR that leverages acoustic-linguistic bimodal information. The serial path follows the ASR+STYLE serial paradigm, reflecting a sequential temporal dependency, while the parallel path integrates our designed Acoustic-Linguistic Similarity Module (ALSM) to facilitate cross-modal interaction with temporal simultaneity. Compared to the existing SSR baseline -- the OSUM model, our approach reduces parameter size by 88.4% and achieves a 30.3% improvement in SSR accuracy for eight styles on the test set.


eess.AS音频处理


【1】I-DCCRN-VAE: An Improved Deep Representation Learning Framework for Complex VAE-based Single-channel Speech Enhancement
标题:I-DCCRN-VAE:用于基于复杂VAE的单通道语音增强的改进深度表示学习框架
链接:https://arxiv.org/abs/2510.12485

作者:Jiatong Li, Simon Doclo
摘要:最近,一个复杂的变分自动编码器(VAE)的基础上DCCRN架构的单通道语音增强系统已经提出。在该系统中,噪声抑制VAE(NSVAE)学习使用预先训练的干净语音和具有跳过连接的噪声VAE从有噪语音中提取干净语音表示。在本文中,我们通过合并三个关键修改来改进DCCRN-VAE:1)删除预训练VAE中的跳过连接,以鼓励更多信息的语音和噪声潜在表示; 2)在预训练中使用$\beta$-VAE,以更好地平衡重建和潜在空间正则化; 3)NSVAE生成语音和噪声潜在表示。实验表明,该系统在匹配的DNS 3数据集上实现了与DCCRN和DCCRN-VAE基线相当的性能,但在不匹配的数据集(WSJ 0-QUT,Voicebank-DEMEND)上优于基线,表现出更好的泛化能力。此外,一项消融研究表明,使用经典微调而不是对抗性训练可以实现类似的性能,从而使训练管道更简单。
摘要:Recently, a complex variational autoencoder (VAE)-based single-channel speech enhancement system based on the DCCRN architecture has been proposed. In this system, a noise suppression VAE (NSVAE) learns to extract clean speech representations from noisy speech using pretrained clean speech and noise VAEs with skip connections. In this paper, we improve DCCRN-VAE by incorporating three key modifications: 1) removing the skip connections in the pretrained VAEs to encourage more informative speech and noise latent representations; 2) using $\beta$-VAE in pretraining to better balance reconstruction and latent space regularization; and 3) a NSVAE generating both speech and noise latent representations. Experiments show that the proposed system achieves comparable performance as the DCCRN and DCCRN-VAE baselines on the matched DNS3 dataset but outperforms the baselines on mismatched datasets (WSJ0-QUT, Voicebank-DEMEND), demonstrating improved generalization ability. In addition, an ablation study shows that a similar performance can be achieved with classical fine-tuning instead of adversarial training, resulting in a simpler training pipeline.


【2】A Phase Synthesizer for Decorrelation to Improve Acoustic Feedback Cancellation
标题:一种用于去相关改善声反馈抵消的相合成器
链接:https://arxiv.org/abs/2510.12377

作者:Klaus Linhard, Philipp Bulling
摘要:不期望的声学反馈是通信系统中的已知问题,例如语音车内通信、公共广播系统或助听器。如果没有额外的预防措施,自适应滤波器(旨在消除反馈路径)也会抑制部分所需信号的风险很高。一种解决方案是去相关扬声器和麦克风信号。在这项工作中,我们结合了两种去相关方法的频率偏移和相位调制在一个统一的框架:一个所谓的\textit{相位合成器},实现在离散傅里叶变换(DFT)滤波器组。此外,我们扩展了相位调制技术,使用可变延迟线,从颤音和合唱效果。我们证明了所提出的相位合成器的好处,从语音车内通信的一个例子,采用自适应频域卡尔曼滤波器。提出了在系统稳定性、语音质量感知评价(PESQ)等方面的改进.
摘要:Undesired acoustic feedback is a known issue in communication systems, such as speech in-car communication, public address systems, or hearing aids. Without additional precautions, there is a high risk that the adaptive filter - intended to cancel the feedback path - also suppresses parts of the desired signal. One solution is to decorrelate the loudspeaker and microphone signals. In this work, we combine the two decorrelation approaches frequency shifting and phase modulation in a unified framework: a so-called \textit{phase synthesizer}, implemented in a discrete Fourier transform (DFT) filter bank. Furthermore, we extend the phase modulation technique using variable delay lines, as known from vibrato and chorus effects. We demonstrate the benefits of the proposed phase synthesizer using an example from speech in-car communication, employing an adaptive frequency-domain Kalman filter. Improvements in system stability, speech quality measured by perceptual evaluation of speech quality (PESQ) are presented.


【3】DeePAQ: A Perceptual Audio Quality Metric Based On Foundational Models and Weakly Supervised Learning
标题:DeePAQ:基于基础模型和弱监督学习的感知音频质量指标
链接:https://arxiv.org/abs/2510.12326

作者:Guanxin Jiang, Andreas Brendel, Pablo M. Delgado, Jürgen Herre
备注:5 pages, 2 figures
摘要:本文提出了用于评估一般音频质量的基于深度学习的感知音频质量度量(DeePAQ)。我们的方法利用度量学习与音乐基础模型MERT一起,由代理标签指导,构建一个嵌入空间,捕获一般音频中的失真强度。据我们所知,DeePAQ是通用音频质量领域中第一个利用弱监督标签和度量学习来微调具有低秩自适应(LoRA)的音乐基础模型的方法,这是其他最先进方法尚未探索的方向。我们基准测试所提出的模型对国家的最先进的客观的音频质量指标,跨越音频编码和源分离的听力测试。结果表明,我们的方法在检测编码伪影方面超越了现有的指标,并且很好地推广到了看不见的失真,如源分离,突出了其鲁棒性和通用性。
摘要:This paper presents the Deep learning-based Perceptual Audio Quality metric (DeePAQ) for evaluating general audio quality. Our approach leverages metric learning together with the music foundation model MERT, guided by surrogate labels, to construct an embedding space that captures distortion intensity in general audio. To the best of our knowledge, DeePAQ is the first in the general audio quality domain to leverage weakly supervised labels and metric learning for fine-tuning a music foundation model with Low-Rank Adaptation (LoRA), a direction not yet explored by other state-of-the-art methods. We benchmark the proposed model against state-of-the-art objective audio quality metrics across listening tests spanning audio coding and source separation. Results show that our method surpasses existing metrics in detecting coding artifacts and generalizes well to unseen distortions such as source separation, highlighting its robustness and versatility.


【4】DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
标题:DiSTAR:语音生成的可扩展令牌自回归表示上的扩散
链接:https://arxiv.org/abs/2510.12210

作者:Yakun Song, Xiaobin Zhuang, Jiawei Chen, Zhikang Niu, Guanrou Yang, Chenpeng Du, Zhuo Chen, Yuping Wang, Yuxuan Wang, Xie Chen
摘要:最近的尝试交错自回归(AR)草图与扩散为基础的精炼连续语音表示已经显示出希望,但他们仍然脆弱的分布变化下,并提供有限的杠杆可控性。我们介绍DISTAR,一个zero-shot文本到语音的框架,完全在离散残差矢量量化(RVQ)代码空间和紧密耦合的AR语言模型与掩蔽扩散模型,没有强制对齐或持续时间预测。具体地说,DISTAR用AR语言模型起草块级RVQ令牌,然后根据草案执行并行掩蔽扩散填充以完成下一个块,从而产生具有分块并行性的长形式合成,同时减轻经典的AR曝光偏差。离散代码空间在推理时提供显式控制:DISTAR使用无分类器指导在贪婪和基于样本的解码下产生高质量音频,支持鲁棒性和多样性之间的权衡,并在测试时通过RVQ层修剪实现可变比特率和可控计算。大量的实验和消融表明,DISTAR超越国家的最先进的zero-shot TTS系统的鲁棒性,自然性和扬声器/风格的一致性,同时保持丰富的输出多样性。音频样本在https://anonymous.4open.science/w/DiSTAR_demo上提供。
摘要:Recent attempts to interleave autoregressive (AR) sketchers with diffusion-based refiners over continuous speech representations have shown promise, but they remain brittle under distribution shift and offer limited levers for controllability. We introduce DISTAR, a zero-shot text-to-speech framework that operates entirely in a discrete residual vector quantization (RVQ) code space and tightly couples an AR language model with a masked diffusion model, without forced alignment or a duration predictor. Concretely, DISTAR drafts block-level RVQ tokens with an AR language model and then performs parallel masked-diffusion infilling conditioned on the draft to complete the next block, yielding long-form synthesis with blockwise parallelism while mitigating classic AR exposure bias. The discrete code space affords explicit control at inference: DISTAR produces high-quality audio under both greedy and sample-based decoding using classifier-free guidance, supports trade-offs between robustness and diversity, and enables variable bit-rate and controllable computation via RVQ layer pruning at test time. Extensive experiments and ablations demonstrate that DISTAR surpasses state-of-the-art zero-shot TTS systems in robustness, naturalness, and speaker/style consistency, while maintaining rich output diversity. Audio samples are provided on https://anonymous.4open.science/w/DiSTAR_demo.


【5】FakeMark: Deepfake Speech Attribution With Watermarked Artifacts
标题:FakeMark:带有水印文物的Deepfake语音归因
链接:https://arxiv.org/abs/2510.12042

作者:Wanying Ge, Xin Wang, Junichi Yamagishi
摘要:Deepfake语音归因对于现有解决方案来说仍然具有挑战性。基于分类器的解决方案通常不能推广到域移位的样本,并且基于水印的解决方案容易受到诸如编解码器压缩或恶意移除攻击之类的失真的损害。为了解决这些问题,我们提出了FakeMark,这是一种新的水印框架,它注入了与deepfake系统相关的伪影相关水印,而不是预先分配的位串消息。这种设计允许检测器通过利用注入的水印和固有的deepfake伪影来归属源系统,即使这些线索之一难以捉摸或被移除也保持有效。实验结果表明,FakeMark提高了对跨数据集样本的泛化能力,而传统的基于分类器的方法在处理不同的失真时仍能保持较高的准确率。
摘要:Deepfake speech attribution remains challenging for existing solutions. Classifier-based solutions often fail to generalize to domain-shifted samples, and watermarking-based solutions are easily compromised by distortions like codec compression or malicious removal attacks. To address these issues, we propose FakeMark, a novel watermarking framework that injects artifact-correlated watermarks associated with deepfake systems rather than pre-assigned bitstring messages. This design allows a detector to attribute the source system by leveraging both injected watermark and intrinsic deepfake artifacts, remaining effective even if one of these cues is elusive or removed. Experimental results show that FakeMark improves generalization to cross-dataset samples where classifier-based solutions struggle and maintains high accuracy under various distortions where conventional watermarking-based solutions fail.


【6】Audio Palette: A Diffusion Transformer with Multi-Signal Conditioning for Controllable Foley Synthesis
标题:音频转换器:具有多信号调节的扩散Transformer,用于可控弗利合成
链接:https://arxiv.org/abs/2510.12175

作者:Junnuo Wang
备注:Accepted for publication in the Journal of Artificial Intelligence Research (JAIR), Vol. 3 No. 2, December 2025
摘要:基于扩散的生成模型的最新进展已经实现了高质量的文本到音频合成,但细粒度的声学控制仍然是开源研究中的一个重大挑战。我们提出了音频转换器,扩散Transformer(DiT)为基础的模型,扩展了稳定的音频开放架构,以解决这个“控制差距”在可控的音频生成。与先前仅依赖于语义条件反射的方法不同,音频控制引入了四个时变控制信号:响度、音高、频谱质心和音色,以精确和可解释地操纵声学特征。该模型在AudioSet的精选子集上使用低秩自适应(LoRA)有效地适应了福利合成的细微差别,只需要训练原始参数的0.85%。实验表明,音频编码器实现了对声音属性的细粒度、可解释的控制。至关重要的是,它实现了这种新颖的可控性,同时保持了高音频质量和与文本提示的强语义对齐,并且在Frechet音频距离(FAD)和LAION-CLAP评分等标准指标上的性能与原始基线模型相当。我们为音频研究提供了一个可扩展的模块化管道,强调基于序列的调节,内存效率和三尺度分类器的指导机制,用于细致入微的推理时间控制。这项工作为开源环境中的可控声音设计和表演性音频合成奠定了坚实的基础,实现了更加以艺术家为中心的工作流程。
摘要:Recent advances in diffusion-based generative models have enabled high-quality text-to-audio synthesis, but fine-grained acoustic control remains a significant challenge in open-source research. We present Audio Palette, a diffusion transformer (DiT) based model that extends the Stable Audio Open architecture to address this "control gap" in controllable audio generation. Unlike prior approaches that rely solely on semantic conditioning, Audio Palette introduces four time-varying control signals: loudness, pitch, spectral centroid, and timbre, for precise and interpretable manipulation of acoustic features. The model is efficiently adapted for the nuanced domain of Foley synthesis using Low-Rank Adaptation (LoRA) on a curated subset of AudioSet, requiring only 0.85 percent of the original parameters to be trained. Experiments demonstrate that Audio Palette achieves fine-grained, interpretable control of sound attributes. Crucially, it accomplishes this novel controllability while maintaining high audio quality and strong semantic alignment to text prompts, with performance on standard metrics such as Frechet Audio Distance (FAD) and LAION-CLAP scores remaining comparable to the original baseline model. We provide a scalable, modular pipeline for audio research, emphasizing sequence-based conditioning, memory efficiency, and a three-scale classifier-free guidance mechanism for nuanced inference-time control. This work establishes a robust foundation for controllable sound design and performative audio synthesis in open-source settings, enabling a more artist-centric workflow.


【7】Serial-Parallel Dual-Path Architecture for Speaking Style Recognition
标题:用于说话风格识别的串并行双路径架构
链接:https://arxiv.org/abs/2510.11732

作者:Guojian Li, Qijie Shao, Zhixian Zhao, Shuiyuan Wang, Zhonghua Fu, Lei Xie
备注:Accepted by NCMMSC2025
摘要:说话风格识别是从语音中识别说话人的说话风格特征。现有的风格识别方法主要依赖于语言信息,对声学信息的整合有限,这限制了识别精度的提高。声学和语言模态的融合提供了显著的潜力,以提高识别性能。在本文中,我们提出了一种新的串行-并行双路径架构SSR,利用声学语言双峰信息。串行路径遵循ASR+STYLE串行范式,反映了顺序的时间依赖性,而并行路径集成了我们设计的声学语言相似性模块(ALSM),以促进跨模态交互与时间的一致性。与现有的SSR基线-OSUM模型相比,我们的方法减少了88.4%的参数大小,并实现了30.3%的提高SSR的准确性为测试集上的八种风格。
摘要:Speaking Style Recognition (SSR) identifies a speaker's speaking style characteristics from speech. Existing style recognition approaches primarily rely on linguistic information, with limited integration of acoustic information, which restricts recognition accuracy improvements. The fusion of acoustic and linguistic modalities offers significant potential to enhance recognition performance. In this paper, we propose a novel serial-parallel dual-path architecture for SSR that leverages acoustic-linguistic bimodal information. The serial path follows the ASR+STYLE serial paradigm, reflecting a sequential temporal dependency, while the parallel path integrates our designed Acoustic-Linguistic Similarity Module (ALSM) to facilitate cross-modal interaction with temporal simultaneity. Compared to the existing SSR baseline -- the OSUM model, our approach reduces parameter size by 88.4% and achieves a 30.3% improvement in SSR accuracy for eight styles on the test set.


机器翻译由腾讯交互翻译提供,仅供参考