今日论文合集:cs.SD语音11篇,eess.AS音频处理7篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training
标题:联合训练实现分离优先多流音频水印
链接:https://arxiv.org/abs/2603.16805

作者:Houmin Sun,Zi Hu,Linxi Li,Yechen Wang,Liwei Jin,Ming Li
摘要:现代音频是通过混合来自不同来源的茎来创建的,这就提出了一个问题:我们能否独立地对每个茎添加水印,并在分离后恢复所有水印?我们研究了一个分离优先的多流水印框架,使用唯一的密钥但共享的结构,混合,分离和解码从每个输出嵌入不同的信息到茎。朴素流水线(鲁棒水印+现成的分离)产生差的比特恢复,示出对通用失真的鲁棒性并不能确保对分离伪像的鲁棒性。为了实现这一点,我们联合训练水印系统和分离器在端到端的方式,鼓励分离器保留水印线索,同时适应嵌入分离特定的失真。语音+音乐和声乐+伴奏混合物的实验表明,分离后的恢复,同时保持感知质量的实质性收益。
摘要:Modern audio is created by mixing stems from different sources, raising the question: can we independently watermark each stem and recover all watermarks after separation? We study a separation-first, multi-stream watermarking framework-embedding distinct information into stems using unique keys but a shared structure, mixing, separating, and decoding from each output. A naive pipeline (robust watermarking + off-the-shelf separation) yields poor bit recovery, showing robustness to generic distortions does not ensure robustness to separation artifacts. To enable this, we jointly train the watermark system and the separator in an end-to-end manner, encouraging the separator to preserve watermark cues while adapting embedding to separation-specific distortions. Experiments on speech+music and vocal+accompaniment mixtures show substantial gains in post-separation recovery while maintaining perceptual quality.


【2】Evaluating Latent Space Structure in Timbre VAEs: A Comparative Study of Unsupervised, Descriptor-Conditioned, and Perceptual Feature-Conditioned Models
标题:评估音色VAE中的潜在空间结构:无监督、描述符条件和知觉条件模型的比较研究
链接:https://arxiv.org/abs/2603.16713

作者:Joseph Cameron,Alan Blackwell
备注:5 pages, 1 figure, 1 table
摘要:我们提出了一个潜在的空间组织在三个变分自动编码器(VAE)的音乐音色生成的比较评估:一个无监督的VAE,一个有条件的VAE,和一个VAE条件上的连续感知功能的AudioCommons音色模型。使用一个精心策划的电吉他声音数据集,标记了四个强度级别的19个语义描述符,我们用一套聚类和可解释性指标来评估每个模型的潜在结构。这些包括轮廓分数、音色描述符紧凑性、音高条件分离、轨迹线性和跨音高一致性。我们的研究结果表明,对感知特征的调节产生了一个更紧凑,更有鉴别力和音高不变的潜在空间,优于无监督和离散的干扰器条件模型。这项工作突出了一热语义条件反射的局限性,并提供了评估音色潜在空间的方法工具,有助于开发更可控和可解释的生成音频模型。
摘要:We present a comparative evaluation of latent space organization in three Variational Autoencoders (VAEs) for musical timbre generation: an unsupervised VAE, a descriptor-conditioned VAE, and a VAE conditioned on continuous perceptual features from the AudioCommons timbral models. Using a curated dataset of electric guitar sounds labeled with 19 semantic descriptors across four intensity levels, we assess each model's latent structure with a suite of clustering and interpretability metrics. These include silhouette scores, timbre descriptor compactness, pitch-conditional separation, trajectory linearity, and cross-pitch consistency. Our findings show that conditioning on perceptual features yields a more compact, discriminative, and pitch-invariant latent space, outperforming both the unsupervised and discrete descriptor-conditioned models. This work highlights the limitations of one-hot semantic conditioning and provides methodological tools for evaluating timbre latent spaces, contributing to the development of more controllable and interpretable generative audio models.


【3】A Semantic Timbre Dataset for the Electric Guitar
标题:电吉他的语义音色数据集
链接:https://arxiv.org/abs/2603.16682

作者:Joseph Cameron,Alan Blackwell
备注:5 pages, 7 figures, 2 tables
摘要:理解和操纵音色是音频合成的核心,但由于缺乏将感知音色维度与语义描述符联系起来的注释数据集,这在机器学习中仍然没有得到充分的探索。我们提出了语义音色数据集,一个精心策划的单声道电吉他声音的集合,每个都标有19个语义音色描述符和相应的幅度之一。这些描述符来自物理和虚拟吉他效果单元的定性分析,并系统地应用于干净的吉他音调。该数据集连接感知音色和机器学习表示,支持音色控制和语义音频生成的学习。我们通过在其潜在空间上训练变分自动编码器(VAE)并使用人类感知判断和描述符分类器对其进行评估来验证数据集。结果表明,VAE捕捉音色结构,并使平滑插值描述符。我们发布了数据集、代码和评估协议,以支持音色感知的生成式AI研究。
摘要:Understanding and manipulating timbre is central to audio synthesis, yet this remains under-explored in machine learning due to a lack of annotated datasets linking perceptual timbre dimensions to semantic descriptors. We present the Semantic Timbre Dataset, a curated collection of monophonic electric guitar sounds, each labeled with one of 19 semantic timbre descriptors and corresponding magnitudes. These descriptors were derived from a qualitative analysis of physical and virtual guitar effect units and applied systematically to clean guitar tones. The dataset bridges perceptual timbre and machine learning representations, supporting learning for timbre control and semantic audio generation. We validate the dataset by training a variational autoencoder (VAE) on its latent space and evaluating it using human perceptual judgments and descriptor classifiers. Results show that the VAE captures timbral structure and enables smooth interpolation across descriptors. We release the dataset, code, and evaluation protocols to support timbre-aware generative AI research.


【4】CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
标题:CAST-DTS:一个简单的交叉注意框架,用于DTS中统一音色控制
链接:https://arxiv.org/abs/2603.16280

作者:Zihao Zheng,Wen Wu,Chao Zhang,Mengyue Wu,Xuenan Xu
备注:Submitted to Interspeech 2026
摘要:当前的文本到语音(TTS)系统通常使用单独的模型来进行语音提示和文本提示的音色控制。虽然将两个控制信号统一到单个模型中是可取的,但跨模态对齐的挑战通常会导致过于复杂的架构和训练目标。为了应对这一挑战,我们提出了CAST-TTS,一个简单而有效的框架,统一的音色控制。使用预先训练的编码器从语音提示和文本提示中提取特征。多阶段训练策略在共享嵌入空间内有效地对齐语音和投影文本表示。然后,一个单一的交叉注意机制允许模型使用这些表示中的任何一个来控制音色。大量实验验证了统一的交叉注意机制对于实现高质量合成至关重要。CAST-TTS在统一架构内运行时,可实现与专用单输入模型相当的性能。演示页面可以在https://HiRookie9.github.io/CAST-TTS-Page上访问。
摘要:Current Text-to-Speech (TTS) systems typically use separate models for speech-prompted and text-prompted timbre control. While unifying both control signals into a single model is desirable, the challenge of cross-modal alignment often results in overly complex architectures and training objective. To address this challenge, we propose CAST-TTS, a simple yet effective framework for unified timbre control. Features are extracted from speech prompts and text prompts using pre-trained encoders. The multi-stage training strategy efficiently aligns the speech and projected text representations within a shared embedding space. A single cross-attention mechanism then allows the model to use either of these representations to control the timbre. Extensive experiments validate that the unified cross-attention mechanism is critical for achieving high-quality synthesis. CAST-TTS achieves performance comparable to specialized single-input models while operating within a unified architecture. The demo page can be accessed at https://HiRookie9.github.io/CAST-TTS-Page.


【5】Diffusion Models for Joint Audio-Video Generation
标题:音频-视频联合生成的扩散模型
链接:https://arxiv.org/abs/2603.16093

作者:Alejandro Paredes La Torre
摘要:多模态生成模型在单模态视频和音频合成方面取得了显着进展,但真正的联合音视频生成仍然是一个开放的挑战。在本文中,我探讨了推进这一领域的四个关键贡献。首先,我发布了两个高质量的配对音频视频数据集。这些数据集包括13小时的视频游戏片段和64小时的音乐会表演,每个片段都分割成一致的34秒样本,以促进可重复的研究。其次,我训练我们的数据集从头开始的MM扩散架构,展示了它的能力,产生语义一致的音频-视频对,并定量评估对齐快速行动和音乐线索。第三,我通过利用预训练的视频和音频编码器-解码器来研究联合潜在扩散,揭示多模式解码阶段的挑战和不一致性。最后,我提出了一个顺序的两步文本到音频视频生成流水线:首先生成视频,然后调节视频输出和原始提示合成时间同步的音频。我的实验表明,这种模块化的方法产生高保真的音频视频生成。
摘要:Multimodal generative models have shown remarkable progress in single-modality video and audio synthesis, yet truly joint audio-video generation remains an open challenge. In this paper, I explore four key contributions to advance this field. First, I release two high-quality, paired audio-video datasets. The datasets consisting on 13 hours of video-game clips and 64 hours of concert performances, each segmented into consistent 34-second samples to facilitate reproducible research. Second, I train the MM-Diffusion architecture from scratch on our datasets, demonstrating its ability to produce semantically coherent audio-video pairs and quantitatively evaluating alignment on rapid actions and musical cues. Third, I investigate joint latent diffusion by leveraging pretrained video and audio encoder-decoders, uncovering challenges and inconsistencies in the multimodal decoding stage. Finally, I propose a sequential two-step text-to-audio-video generation pipeline: first generating video, then conditioning on both the video output and the original prompt to synthesize temporally synchronized audio. My experiments show that this modular approach yields high-fidelity generations of audio video generation.


【6】Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
标题:迈向视觉-声音-创伤-动作范式:以声音为中心操纵的HEAR框架
链接:https://arxiv.org/abs/2603.16086

作者:Chang Nie,Tianchen Deng,Guangming Wang,Zhe Liu,Hesheng Wang
摘要:虽然最近的视觉-语言-动作(VLA)模型已经开始包含音频,但它们通常将声音视为静态的执行前提示或仅关注人类语音。这在实时、以声音为中心的操作中留下了显著的差距,其中短暂的环境声学在任务执行期间提供关键状态验证。因此,由于低频更新或系统延迟,很容易错过关键声音。开环执行的动作分块会加剧这个问题,这会创建一个盲执行间隔,其中声学事件在离散的音频观察窗口之间丢失。认识到连续听觉意识的必要性,我们正式视觉-声音-听觉-动作(VSLA)作为一个连续的控制范式条件下的视觉,音频流,语言和本体感受延迟决策循环。作为一个实例,我们介绍了HEAR,一个集成了四个组件的VSLA框架:(i)一个流Historizer,用于在执行间隙中保持紧凑的因果音频上下文;(ii)一个Envisionary,适用于全方位基础模型,用于对多感官输入进行推理;(iii)一个Advancer,被公式化为音频世界模型,通过预测不久的将来的音频代码来学习时间动态;以及(iv)流匹配实现器策略以生成平滑动作块。为了解决VSLA预训练数据和评估的稀缺性,我们构建了OpenX-Sound用于预训练,以及HEAR-Bench,这是第一个以声音为中心的操作基准,具有严格的因果时序规则。我们的研究结果表明,强大的声音为中心的操纵需要因果的持久性和明确的时间学习。这个框架提供了一个实际的一步多感官的基础模型体现代理,使机器人感知和动态环境中进行交互。代码和视频可在https://hear.irmv.top上获得。
摘要:While recent Vision-Language-Action (VLA) models have begun to incorporate audio, they typically treat sound as static pre-execution prompts or focus exclusively on human speech. This leaves a significant gap in real-time, sound-centric manipulation where fleeting environmental acoustics provide critical state verification during task execution. Consequently, key sounds are easily missed due to low-frequency updates or system latency. This problem is exacerbated by action chunking with open-loop execution, which creates a Blind Execution Interval where acoustic events are lost between discrete audio observation windows. Recognizing the necessity of continuous auditory awareness, we formalize Vision-Sound-Language-Action (VSLA) as a continuous control paradigm conditioned on vision, streaming audio, language, and proprioception under delayed decision loops. As an instantiation, we introduce HEAR, a VSLA framework integrating four components: (i) a streaming Historizer to maintain a compact, causal audio context across execution gaps; (ii) an Envisioner adapted from omni foundation models to reason over multi-sensory inputs; (iii) an Advancer, formulated as an audio world model, to learn temporal dynamics by predicting near-future audio codes; and (iv) a flow-matching Realizer policy to generate smooth action chunks. To address the scarcity of pretraining data and evaluations for VSLA, we construct OpenX-Sound for pretraining, alongside HEAR-Bench, the first sound-centric manipulation benchmark with strict causal timing rules. Our results suggest that robust sound-centric manipulation necessitates causal persistence and explicit temporal learning. This framework provides a practical step toward multi-sensory foundation models for embodied agents, enabling robots to perceive and interact with dynamic environments. Code and videos are available at https://hear.irmv.top.


【7】INSTRUMENTAL: Automatic Synthesizer Parameter Recovery from Audio via Evolutionary Optimization
标题:INSTRUMENTAL:通过进化优化从音频中自动恢复合成器参数
链接:https://arxiv.org/abs/2603.15905

作者:Philipp Bogdan
备注:5 pages
摘要:现有的音频转换工具提取音符,但丢弃了定义乐器身份的音色特征。我们提出了仪器,一个系统,恢复连续合成器参数从音频耦合微分28参数减法合成器与CMA-ES,衍生物免费进化优化。我们优化了一个复合感知损失结合梅尔缩放STFT,频谱质心,MFCC发散,实现了2.09的匹配损失的真实记录的音频。我们系统地评估了八个假设,以提高收敛性,发现只有参数EQ提升产生有意义的改善。我们的研究结果表明,CMA-ES在这个非凸景观上优于梯度下降,更多的参数不会单调地改善匹配,并且谱分析初始化加速了随机开始的收敛。
摘要:Existing audio-to-MIDI tools extract notes but discard the timbral characteristics that define an instrument's identity. We present Instrumental, a system that recovers continuous synthesizer parameters from audio by coupling a differentiable 28-parameter subtractive synthesizer with CMA-ES, a derivative-free evolutionary optimizer. We optimize a composite perceptual loss combining mel-scaled STFT, spectral centroid, and MFCC divergence, achieving a matching loss of 2.09 on real recorded audio. We systematically evaluate eight hypotheses for improving convergence and find that only parametric EQ boosting yields meaningful improvement. Our results show that CMA-ES outperforms gradient descent on this non-convex landscape, that more parameters do not monotonically improve matching, and that spectral analysis initialization accelerates convergence over random starts.


【8】PulmoVec: A Two-Stage Stacking Meta-Learning Architecture Built on the HeAR Foundation Model for Multi-Task Classification of Pediatric Respiratory Sounds
标题:PulmoVec:基于HeAR基金会模型构建的两阶段堆叠元学习架构,用于儿科呼吸音的多任务分类
链接:https://arxiv.org/abs/2603.15688

作者:Izzet Turkalp Akbasli,Oguzhan Serin
备注:14 pages, 2 figures, 4 tables; supplementary material included (4 tables, 3 multi-panel figures)
摘要:背景资料:呼吸道疾病是儿童发病和死亡的主要原因,但肺部听诊仍然具有主观性,并且受到听众间差异的限制,特别是在儿科人群中。现有的人工智能方法进一步受到小数据集和单任务设计的限制。我们开发了PulmoVec,这是一个基于健康声学表示(HeAR)基础模型的多任务框架,用于儿科呼吸音的分类。方法:在SPRSound数据库的回顾性分析中,分析了来自1,652名儿科患者的24,808个事件级注释片段。三个特定任务的分类器进行了训练,用于筛选,声音模式识别和疾病组预测。他们的出折概率输出与LightGBM堆叠元模型中的人口统计学元数据相结合,并使用集合投票将事件级预测聚合到患者级。结果如下:在事件水平,筛选模型的ROC-AUC为0.96(95% CI,0.95-0.97),声音模式识别模型的宏观ROC-AUC为0.96(95% CI,0.96-0.97),疾病组预测模型的宏观ROC-AUC为0.94(95% CI,0.93-0.94)。在患者水平,疾病组分类的准确性为0.74(95% CI,0.71-0.77),加权F1评分为0.73,宏观ROC-AUC为0.91(95% CI,0.90-0.93)。与单独的基本模型相比,堆叠提高了所有任务的性能。结论:PulmoVec将事件级声学表型与患者级临床分类联系起来,支持基于基础模型的数字听诊在儿科呼吸医学中的潜力。跨设备和真实世界条件的多中心外部验证仍然至关重要。
摘要:Background: Respiratory diseases are a leading cause of childhood morbidity and mortality, yet lung auscultation remains subjective and limited by inter-listener variability, particularly in pediatric populations. Existing AI approaches are further constrained by small datasets and single-task designs. We developed PulmoVec, a multi-task framework built on the Health Acoustic Representations (HeAR) foundation model for classification of pediatric respiratory sounds. Methods: In this retrospective analysis of the SPRSound database, 24,808 event-level annotated segments from 1,652 pediatric patients were analyzed. Three task-specific classifiers were trained for screening, sound-pattern recognition, and disease-group prediction. Their out-of-fold probability outputs were combined with demographic metadata in a LightGBM stacking meta-model, and event-level predictions were aggregated to the patient level using ensemble voting. Results: At the event level, the screening model achieved an ROC-AUC of 0.96 (95% CI, 0.95-0.97), the sound-pattern recognition model a macro ROC-AUC of 0.96 (95% CI, 0.96-0.97), and the disease-group prediction model a macro ROC-AUC of 0.94 (95% CI, 0.93-0.94). At the patient level, disease-group classification yielded an accuracy of 0.74 (95% CI, 0.71-0.77), a weighted F1-score of 0.73, and a macro ROC-AUC of 0.91 (95% CI, 0.90-0.93). Stacking improved performance across all tasks compared with base models alone. Conclusions: PulmoVec links event-level acoustic phenotyping with patient-level clinical classification, supporting the potential of foundation-model-based digital auscultation in pediatric respiratory medicine. Multi-center external validation across devices and real-world conditions remains essential.


【9】DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
标题:DASH:动态音频驱动的语义分块,用于高效的全模式令牌压缩
链接:https://arxiv.org/abs/2603.15685

作者:Bingzhou Li,Tao Huang
摘要:全模态大型语言模型(OmniLLM)联合处理音频和视频流,但由此产生的长多模态令牌序列使推理变得非常昂贵。现有的压缩方法通常依赖于固定的窗口分割和基于注意力的修剪,忽略了分段的语义结构的视听信号,并成为下积极的令牌减少脆弱。我们提出了动态音频驱动的语义cHunking(DASH),这是一个无需训练的框架,将令牌压缩与语义结构相结合。DASH将音频嵌入视为语义锚点,并通过余弦相似性不连续来检测边界候选,从而产生动态的可变长度片段,这些片段近似于序列的底层分段连贯组织。这些边界被投影到视频标记上,以建立显式的跨模态分割。在每个片段内,令牌保留由三信号重要性估计器确定,该估计器融合了结构边界线索、代表性独特性和基于注意力的显著性,减轻了仅注意力选择的稀疏性偏差。这种结构感知的分配保留了转换关键令牌,同时减少了冗余区域。在AVUT、VideoMME和WorldSense上进行的大量实验表明,与现有方法相比,DASH保持了卓越的准确性,同时实现了更高的压缩比。代码可从以下网址获得:https://github.com/laychou666/DASH。
摘要:Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make inference prohibitively expensive. Existing compression methods typically rely on fixed window partitioning and attention-based pruning, which overlook the piecewise semantic structure of audio-visual signals and become fragile under aggressive token reduction. We propose Dynamic Audio-driven Semantic cHunking (DASH), a training-free framework that aligns token compression with semantic structure. DASH treats audio embeddings as a semantic anchor and detects boundary candidates via cosine-similarity discontinuities, inducing dynamic, variable-length segments that approximate the underlying piecewise-coherent organization of the sequence. These boundaries are projected onto video tokens to establish explicit cross-modal segmentation. Within each segment, token retention is determined by a tri-signal importance estimator that fuses structural boundary cues, representational distinctiveness, and attention-based salience, mitigating the sparsity bias of attention-only selection. This structure-aware allocation preserves transition-critical tokens while reducing redundant regions. Extensive experiments on AVUT, VideoMME, and WorldSense demonstrate that DASH maintains superior accuracy while achieving higher compression ratios compared to prior methods. Code is available at: https://github.com/laychou666/DASH.


【10】HRTF-guided Binaural Target Speaker Extraction with Real-World Validation
标题:HRTF引导的双耳目标说话人提取及真实世界验证
链接:https://arxiv.org/abs/2603.16668

作者:Yoav Ellinson,Sharon Gannot
备注:Submitted to Interspeech 2026
摘要:提出了一种基于头相关传递函数(HRTF)的双耳目标说话人提取(TSE)方法。与传统的TSE方法的基础上到达方向(DOA)估计或登记信号,这往往扭曲感知的空间位置,所提出的方法利用听众的HRTF作为一个明确的空间先验。所提出的框架是建立在多通道深度盲源分离骨干,适应双耳TSE设置。它是在来自不同人群的测量HRTF上训练的,能够实现跨听众的泛化,而不是特定于主题的调谐。通过调节提取HRTF派生的空间信息,该方法保留双耳线索,同时提高语音质量和可懂度。所提出的框架的性能进行了验证,通过模拟和真实的录音从头部和躯干模拟器(HATS)。
摘要:This paper presents a Head-Related Transfer Function (HRTF)-guided framework for binaural Target Speaker Extraction (TSE) from mixtures of concurrent sources. Unlike conventional TSE methods based on Direction of Arrival (DOA) estimation or enrollment signals, which often distort perceived spatial location, the proposed approach leverages the listener's HRTF as an explicit spatial prior. The proposed framework is built upon a multi-channel deep blind source separation backbone, adapted to the binaural TSE setting. It is trained on measured HRTFs from a diverse population, enabling cross-listener generalization rather than subject-specific tuning. By conditioning the extraction on HRTF-derived spatial information, the method preserves binaural cues while enhancing speech quality and intelligibility. The performance of the proposed framework is validated through simulations and real recordings obtained from a head and torso simulator (HATS).


【11】Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
标题:鲁棒的生成音频质量评估:从虚假相关中分离质量
链接:https://arxiv.org/abs/2603.16201

作者:Kuan-Tang Huang,Chien-Chun Wang,Cheng-Yeh Yang,Hung-Shin Lee,Hsin-Min Wang,Berlin Chen
备注:Accepted to IEEE ICME 2026
摘要:人工智能生成内容(AIGC)的快速增长需要强大的感知质量评估指标。然而,自动平均意见得分(MOS)预测模型往往受到数据稀缺的影响,使它们容易学习虚假的相关性-例如特定于网络的声学特征-而不是广义的质量特征。为了解决这个问题,我们利用领域对抗训练(DAT)来从这些讨厌的因素中分离出真正的质量感知。不同于以往的作品,依赖于静态域先验,我们系统地研究域定义策略,从显式元数据驱动的标签,隐式数据驱动的集群。我们的研究结果表明,没有“一刀切”的域定义,相反,最佳策略是高度依赖于具体的MOS方面进行评估。实验结果表明,我们的特定方面的域策略有效地减轻了声学偏见,显着提高与人类评级的相关性,并实现了对看不见的生成场景的卓越泛化。
摘要:The rapid proliferation of AI-Generated Content (AIGC) has necessitated robust metrics for perceptual quality assessment. However, automatic Mean Opinion Score (MOS) prediction models are often compromised by data scarcity, predisposing them to learn spurious correlations-- such as dataset-specific acoustic signatures-- rather than generalized quality features. To address this, we leverage domain adversarial training (DAT) to disentangle true quality perception from these nuisance factors. Unlike prior works that rely on static domain priors, we systematically investigate domain definition strategies ranging from explicit metadata-driven labels to implicit data-driven clusters. Our findings reveal that there is no "one-size-fits-all" domain definition; instead, the optimal strategy is highly dependent on the specific MOS aspect being evaluated. Experimental results demonstrate that our aspect-specific domain strategy effectively mitigates acoustic biases, significantly improving correlation with human ratings and achieving superior generalization on unseen generative scenarios.


eess.AS音频处理


【1】HRTF-guided Binaural Target Speaker Extraction with Real-World Validation
标题:HRTF引导的双耳目标说话人提取及真实世界验证
链接:https://arxiv.org/abs/2603.16668

作者:Yoav Ellinson,Sharon Gannot
备注:Submitted to Interspeech 2026
摘要:提出了一种基于头相关传递函数(HRTF)的双耳目标说话人提取(TSE)方法。与传统的TSE方法的基础上到达方向(DOA)估计或登记信号,这往往扭曲感知的空间位置,所提出的方法利用听众的HRTF作为一个明确的空间先验。所提出的框架是建立在多通道深度盲源分离骨干,适应双耳TSE设置。它是在来自不同人群的测量HRTF上训练的,能够实现跨听众的泛化,而不是特定于主题的调谐。通过调节提取HRTF派生的空间信息,该方法保留双耳线索,同时提高语音质量和可懂度。所提出的框架的性能进行了验证,通过模拟和真实的录音从头部和躯干模拟器(HATS)。
摘要:This paper presents a Head-Related Transfer Function (HRTF)-guided framework for binaural Target Speaker Extraction (TSE) from mixtures of concurrent sources. Unlike conventional TSE methods based on Direction of Arrival (DOA) estimation or enrollment signals, which often distort perceived spatial location, the proposed approach leverages the listener's HRTF as an explicit spatial prior. The proposed framework is built upon a multi-channel deep blind source separation backbone, adapted to the binaural TSE setting. It is trained on measured HRTFs from a diverse population, enabling cross-listener generalization rather than subject-specific tuning. By conditioning the extraction on HRTF-derived spatial information, the method preserves binaural cues while enhancing speech quality and intelligibility. The performance of the proposed framework is validated through simulations and real recordings obtained from a head and torso simulator (HATS).


【2】Speakers Localization Using Batch EM In Unfolding Neural Network
标题:在展开神经网络中使用批量EM进行扬声器定位
链接:https://arxiv.org/abs/2603.16278

作者:Rina Veler,Sharon Gannot
备注:3 pages, 1 figure, ICSEE 2026
摘要:我们提出了一个可解释的Batch-EM展开网络,用于鲁棒的说话人定位。通过在编码器-EM-解码器架构内嵌入迭代EM过程,该方法减轻了初始化敏感性并提高了收敛性。实验表明,在混响条件下,优于经典的批量EM的准确性和鲁棒性。
摘要:We propose an interpretable Batch-EM Unfolded Network for robust speaker localization. By embedding the iterative EM procedure within an encoder-EM-decoder architecture, the method mitigates initialization sensitivity and improves convergence. Experiments show superior accuracy and robustness over the classical Batch-EM in reverberant conditions.


【3】Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
标题:鲁棒的生成音频质量评估:从虚假相关中分离质量
链接:https://arxiv.org/abs/2603.16201

作者:Kuan-Tang Huang,Chien-Chun Wang,Cheng-Yeh Yang,Hung-Shin Lee,Hsin-Min Wang,Berlin Chen
备注:Accepted to IEEE ICME 2026
摘要:人工智能生成内容(AIGC)的快速增长需要强大的感知质量评估指标。然而,自动平均意见得分(MOS)预测模型往往受到数据稀缺的影响,使它们容易学习虚假的相关性-例如特定于网络的声学特征-而不是广义的质量特征。为了解决这个问题,我们利用领域对抗训练(DAT)来从这些讨厌的因素中分离出真正的质量感知。不同于以往的作品,依赖于静态域先验,我们系统地研究域定义策略,从显式元数据驱动的标签,隐式数据驱动的集群。我们的研究结果表明,没有“一刀切”的域定义,相反,最佳策略是高度依赖于具体的MOS方面进行评估。实验结果表明,我们的特定方面的域策略有效地减轻了声学偏见,显着提高与人类评级的相关性,并实现了对看不见的生成场景的卓越泛化。
摘要:The rapid proliferation of AI-Generated Content (AIGC) has necessitated robust metrics for perceptual quality assessment. However, automatic Mean Opinion Score (MOS) prediction models are often compromised by data scarcity, predisposing them to learn spurious correlations-- such as dataset-specific acoustic signatures-- rather than generalized quality features. To address this, we leverage domain adversarial training (DAT) to disentangle true quality perception from these nuisance factors. Unlike prior works that rely on static domain priors, we systematically investigate domain definition strategies ranging from explicit metadata-driven labels to implicit data-driven clusters. Our findings reveal that there is no "one-size-fits-all" domain definition; instead, the optimal strategy is highly dependent on the specific MOS aspect being evaluated. Experimental results demonstrate that our aspect-specific domain strategy effectively mitigates acoustic biases, significantly improving correlation with human ratings and achieving superior generalization on unseen generative scenarios.


【4】AILive Mixer: A Deep Learning based Zero Latency Automatic Music Mixer for Live Music Performances
标题:AILive Mixer:一款基于深度学习的零延迟自动音乐混音器,用于现场音乐表演
链接:https://arxiv.org/abs/2603.15995

作者:Devansh Zurale,Iris Lorente,Michael Lester,Alex Mitchell
备注:5 pages, 4 figures, accepted to ICASSP 2026
摘要:在这项工作中,我们提出了一个基于深度学习的自动多轨音乐混合系统,以满足现场表演的需要。在现场表演中,通道经常被位于同一位置的乐器的声泄漏破坏。此外,视听同步至关重要,因此对音频延迟施加了严格的约束。在这项工作中,我们主要解决这两个挑战,处理输入通道中的出血,以产生零延迟的音乐混合。虽然最近在自动音乐混音领域有了一些发展,但大多数或所有以前的作品都集中在隔离乐器信号的离线制作上,据我们所知,这是第一个为现场音乐表演开发的端到端深度学习系统。我们提出的系统目前预测单声道增益的多轨输入,但它的设计以及在过去的作品中的先例,可以很容易地适应未来的工作,预测其他相关的音乐混合参数。
摘要:In this work, we present a deep learning-based automatic multitrack music mixing system catered towards live performances. In a live performance, channels are often corrupted with acoustic bleeds of co-located instruments. Moreover, audio-visual synchronization is of critical importance thus putting a tight constraint on the audio latency. In this work we primarily tackle these two challenges of handling bleeds in the input channels to produce the music mix with zero latency. Although there have been several developments in the field of automatic music mixing in recent times, most or all previous works focus on offline production for isolated instrument signals and to the best of our knowledge, this is the first end-to-end deep learning system developed for live music performances. Our proposed system currently predicts mono gains for a multitrack input, but its design along with the precedent set in past works, allows for easy adaptation to future work of predicting other relevant music mixing parameters.


【5】Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech
标题:从无到有:数据增强,用于稳健的发音障碍严重程度估计
链接:https://arxiv.org/abs/2603.15988

作者:Jaesung Bae,Xiuwen Zheng,Minje Kim,Chang D. Yoo,Mark Hasegawa-Johnson
备注:Submitted to Interspeech 2026
摘要:构音障碍语音质量评估(DSQA)对于临床诊断和包容性语音技术至关重要。然而,主观评估是昂贵的,难以规模化,标记数据的稀缺性限制了鲁棒的客观建模。为了解决这个问题,我们提出了一个三阶段的框架,利用未标记的构音障碍语音和大规模的典型语音数据集来扩展训练。教师模型首先为未标记的样本生成伪标签,然后使用标签感知对比学习策略进行弱监督预训练,该策略将模型暴露于不同的扬声器和声学条件。然后,针对下游DSQA任务对预训练的模型进行微调。五个看不见的数据集,跨越多种病因和语言的实验证明了我们的方法的鲁棒性。我们基于Whisper的基线显著优于SpICE等SOTA DSQA预测器,整个框架在未见过的测试数据集上实现了0.761的平均SRCC。
摘要:Dysarthric speech quality assessment (DSQA) is critical for clinical diagnostics and inclusive speech technologies. However, subjective evaluation is costly and difficult to scale, and the scarcity of labeled data limits robust objective modeling. To address this, we propose a three-stage framework that leverages unlabeled dysarthric speech and large-scale typical speech datasets to scale training. A teacher model first generates pseudo-labels for unlabeled samples, followed by weakly supervised pretraining using a label-aware contrastive learning strategy that exposes the model to diverse speakers and acoustic conditions. The pretrained model is then fine-tuned for the downstream DSQA task. Experiments on five unseen datasets spanning multiple etiologies and languages demonstrate the robustness of our approach. Our Whisper-based baseline significantly outperforms SOTA DSQA predictors such as SpICE, and the full framework achieves an average SRCC of 0.761 across unseen test datasets.


【6】RECOVER: Robust Entity Correction via agentic Orchestration of hypothesis Variants for Evidence-based Recovery
标题:RECVERER:通过基于证据的恢复的假设变体的代理描述进行稳健的实体纠正
链接:https://arxiv.org/abs/2603.16411

作者:Abhishek Kumar,Aashraya Sachdeva
备注:Under review. Submitted to Interspeech 2026
摘要:自动语音识别(ASR)中的实体识别对于稀有和特定领域的术语具有挑战性。在金融、医药和空中交通管制等领域,这些错误代价高昂。如果实体完全不存在于ASR输出中,则ASR后校正变得困难。为了解决这个问题,我们引入了RECOVER,这是一个代理纠正框架,充当工具使用代理。它利用多个假设作为ASR的证据,检索相关实体,并在约束条件下应用大型语言模型(LLM)校正。假设使用不同的策略,即1-最佳,智能感知选择,识别器输出投票错误减少(ROVER)Entrance和LLM-Select。在五个不同的数据集上进行评估,它实现了实体短语单词错误率(E-WER)相对降低8-46%,并将召回率提高了22个百分点。LLM-Select在实体校正方面实现了最佳的整体性能,同时保持了整体WER。
摘要:Entity recognition in Automatic Speech Recognition (ASR) is challenging for rare and domain-specific terms. In domains such as finance, medicine, and air traffic control, these errors are costly. If the entities are entirely absent from the ASR output, post-ASR correction becomes difficult. To address this, we introduce RECOVER, an agentic correction framework that serves as a tool-using agent. It leverages multiple hypotheses as evidence from ASR, retrieves relevant entities, and applies Large Language Model (LLM) correction under constraints. The hypotheses are used using different strategies, namely, 1-Best, Entity-Aware Select, Recognizer Output Voting Error Reduction (ROVER) Ensemble, and LLM-Select. Evaluated across five diverse datasets, it achieves 8-46% relative reductions in entity-phrase word error rate (E-WER) and increases recall by up to 22 percentage points. The LLM-Select achieves the best overall performance in entity correction while maintaining overall WER.


【7】CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
标题:CAST-DTS:一个简单的交叉注意框架,用于DTS中统一音色控制
链接:https://arxiv.org/abs/2603.16280

作者:Zihao Zheng,Wen Wu,Chao Zhang,Mengyue Wu,Xuenan Xu
备注:Submitted to Interspeech 2026
摘要:当前的文本到语音(TTS)系统通常使用单独的模型来进行语音提示和文本提示的音色控制。虽然将两个控制信号统一到单个模型中是可取的,但跨模态对齐的挑战通常会导致过于复杂的架构和训练目标。为了应对这一挑战,我们提出了CAST-TTS,一个简单而有效的框架,统一的音色控制。使用预先训练的编码器从语音提示和文本提示中提取特征。多阶段训练策略在共享嵌入空间内有效地对齐语音和投影文本表示。然后,一个单一的交叉注意机制允许模型使用这些表示中的任何一个来控制音色。大量的实验验证了统一的交叉注意机制是实现高质量合成的关键。CAST-TTS在统一架构内运行时,可实现与专用单输入模型相当的性能。演示页面可以在https://HiRookie9.github.io/CAST-TTS-Page上访问。
摘要:Current Text-to-Speech (TTS) systems typically use separate models for speech-prompted and text-prompted timbre control. While unifying both control signals into a single model is desirable, the challenge of cross-modal alignment often results in overly complex architectures and training objective. To address this challenge, we propose CAST-TTS, a simple yet effective framework for unified timbre control. Features are extracted from speech prompts and text prompts using pre-trained encoders. The multi-stage training strategy efficiently aligns the speech and projected text representations within a shared embedding space. A single cross-attention mechanism then allows the model to use either of these representations to control the timbre. Extensive experiments validate that the unified cross-attention mechanism is critical for achieving high-quality synthesis. CAST-TTS achieves performance comparable to specialized single-input models while operating within a unified architecture. The demo page can be accessed at https://HiRookie9.github.io/CAST-TTS-Page.


机器翻译由腾讯交互翻译提供,仅供参考