今日论文合集:cs.SD语音14篇,eess.AS音频处理8篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
标题:对话西顿:从野外对话音频中恢复全速对话曲目
链接:https://arxiv.org/abs/2604.09344

作者:Wataru Nakata,Yuki Saito,Kazuki Yamauchi,Emiru Tsunoo,Hiroshi Saruwatari
备注:12 pages, 2 figures
摘要:全双工对话音频,其中每个扬声器被记录在单独的轨道上,是口语对话研究的重要资源,但难以大规模收集。大多数野外的两个扬声器对话只能作为退化的单声道混合信号,这使得它不适合需要清晰的扬声器信号的系统。我们提出了DialogueSidon,一个模型,用于联合恢复和分离退化的单声道两个扬声器对话音频。DialogueSidon结合了变分自动编码器(VAE)对语音自监督学习(SSL)模型特征进行操作,该特征将SSL模型特征压缩到紧凑的潜在空间中,并结合了基于扩散的潜在预测器,该预测器从降级的混合物中恢复说话者的潜在表示。在英语、多语言和野外对话数据集上的实验表明,DialogueSidon大大提高了基线的可理解性和分离质量,同时也实现了更快的推理。
摘要:Full-duplex dialogue audio, in which each speaker is recorded on a separate track, is an important resource for spoken dialogue research, but is difficult to collect at scale. Most in-the-wild two-speaker dialogue is available only as degraded monaural mixtures, making it unsuitable for systems requiring clean speaker-wise signals. We propose DialogueSidon, a model for joint restoration and separation of degraded monaural two-speaker dialogue audio. DialogueSidon combines a variational autoencoder (VAE) operates on the speech self-supervised learning (SSL) model feature, which compresses SSL model features into a compact latent space, with a diffusion-based latent predictor that recovers speaker-wise latent representations from the degraded mixture. Experiments on English, multilingual, and in-the-wild dialogue datasets show that DialogueSidon substantially improves intelligibility and separation quality over a baseline, while also achieving much faster inference.


【2】DDSP-QbE++: Improving Speech Quality for Speech Anonymisation for Atypical Speech
标题:DDSP-QbE++:通过非典型语音的语音匿名化提高语音质量
链接:https://arxiv.org/abs/2604.09246

作者:Suhita Ghosh,Yamini Sinha,Sebastian Stober
备注:accepted in CHI workshop (Speech AI For All) 2026
摘要:用于语音转换的可微分数字信号处理(DDSP)流水线依赖于减法合成,其中周期性激励信号由学习的频谱包络整形以重构目标语音。在DDSP-QbE中,激励通过相位累积产生,产生锯齿状波形,其突然的不连续性引入混叠伪像,其在感知上表现为混叠和频谱失真,特别是在较高的基频处。我们提出了两个有针对性的改进DDSP-QbE减法合成器的激励阶段。首先,我们采用显式的浊音检测门谐波激励,抑制无声区域的周期性成分,并将其替换为过滤后的噪声,从而避免混叠谐波内容,它是最感知破坏性。其次,我们应用多项式带限步(PolyBLEP)校正的相位累积振荡器,取代硬波形的不连续性在每个相位缠绕一个平滑的多项式残差,取消混叠生成组件,而无需过采样或频谱截断。总之,这些修改产生更干净的谐波滚降,减少高频伪影,并改善感知自然度,如MOS所测量的。所提出的方法是轻量级的,可区分的,并且无缝集成到现有的DDSP-QbE训练管道中,无需额外的可学习参数。
摘要:Differentiable Digital Signal Processing (DDSP) pipelines for voice conversion rely on subtractive synthesis, where a periodic excitation signal is shaped by a learned spectral envelope to reconstruct the target voice. In DDSP-QbE, the excitation is generated via phase accumulation, producing a sawtooth-like waveform whose abrupt discontinuities introduce aliasing artefacts that manifest perceptually as buzziness and spectral distortion, particularly at higher fundamental frequencies. We propose two targeted improvements to the excitation stage of the DDSP-QbE subtractive synthesizer. First, we incorporate explicit voicing detection to gate the harmonic excitation, suppressing the periodic component in unvoiced regions and replacing it with filtered noise, thereby avoiding aliased harmonic content where it is most perceptually disruptive. Second, we apply Polynomial Band-Limited Step (PolyBLEP) correction to the phase-accumulated oscillator, substituting the hard waveform discontinuity at each phase wrap with a smooth polynomial residual that cancels alias-generating components without oversampling or spectral truncation. Together, these modifications yield a cleaner harmonic roll-off, reduced high-frequency artefacts, and improved perceptual naturalness, as measured by MOS. The proposed approach is lightweight, differentiable, and integrates seamlessly into the existing DDSP-QbE training pipeline with no additional learnable parameters.


【3】GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking
标题:GRM:通过攻击比率掩蔽对音频LLM进行实用性感知越狱攻击
链接:https://arxiv.org/abs/2604.09222

作者:Yunqiang Wang,Hengyuan Na,Di Wu,Miao Hu,Guocong Quan
备注:Under Review
摘要:音频大语言模型(ALLM)支持丰富的语音-文本交互,但它们也在音频模态中引入了越狱漏洞。现有的音频越狱方法主要是优化越狱成功,而忽略了效用保存,如转录质量和问题回答性能。在实践中,更强大的攻击往往以降低效用为代价。为了研究这种权衡,我们重新审视现有的攻击,通过改变它们在频域中的扰动覆盖范围,从部分频带到全频带,并发现更广泛的频率覆盖范围并不一定能提高越狱性能,而实用性不断恶化。这表明,集中扰动的一个子集的波段可以产生一个更好的攻击效用权衡比不分青红皂白的全波段覆盖。基于这一见解,我们提出了GRM,一个实用意识的频率选择性越狱框架。它排名梅尔乐队的攻击贡献相对于实用程序的敏感性,扰动只有一个选定的子集的频带,并学习一个可重用的通用扰动下的语义保存目标。在四个有代表性的ALLM上的实验表明,GRM实现了88.46%的平均越狱成功率(JSR),同时提供了比代表性基线更好的攻击效用权衡。这些结果突出了频率选择性扰动在音频越狱中更好地平衡攻击有效性和实用性保护的潜力。内容警告:本文包含有害查询示例和不安全的模型响应。
摘要:Audio large language models (ALLMs) enable rich speech-text interaction, but they also introduce jailbreak vulnerabilities in the audio modality. Existing audio jailbreak methods mainly optimize jailbreak success while overlooking utility preservation, as reflected in transcription quality and question answering performance. In practice, stronger attacks often come at the cost of degraded utility. To study this trade-off, we revisit existing attacks by varying their perturbation coverage in the frequency domain, from partial-band to full-band, and find that broader frequency coverage does not necessarily improve jailbreak performance, while utility consistently deteriorates. This suggests that concentrating perturbation on a subset of bands can yield a better attack-utility trade-off than indiscriminate full-band coverage. Based on this insight, we propose GRM, a utility-aware frequency-selective jailbreak framework. It ranks Mel bands by their attack contribution relative to utility sensitivity, perturbs only a selected subset of bands, and learns a reusable universal perturbation under a semantic-preservation objective. Experiments on four representative ALLMs show that GRM achieves an average Jailbreak Success Rate (JSR) of 88.46% while providing a better attack-utility trade-off than representative baselines. These results highlight the potential of frequency-selective perturbation for better balancing attack effectiveness and utility preservation in audio jailbreak. Content Warning: This paper includes harmful query examples and unsafe model responses.


【4】LatentFlowSR: High-Fidelity Audio Super-Resolution via Noise-Robust Latent Flow Matching
标题:LatentFlowSR:通过噪音稳健的潜伏流匹配实现高保真音频超分辨率
链接:https://arxiv.org/abs/2604.09188

作者:Fei Liu,Yang Ai,Hui-Peng Du,Yu-Fei Shi,Zhen-Hua Ling
摘要:音频超分辨率旨在从带宽有限的低分辨率音频中恢复丢失的高频细节,从而提高重建信号的自然度和感知质量。然而,大多数现有的方法直接在波形或时频域中操作,这不仅涉及高维生成空间,而且在很大程度上限于语音任务,从而为更复杂的音频类型(例如音效和音乐)留下了很大的改进空间。为了减轻这些限制,我们引入了LatentFlowSR,这是一种新的音频超分辨率方法,它利用了潜在表示空间内的条件流匹配(CFM)。具体来说,我们首先训练一个噪声鲁棒的自编码器,它将低分辨率音频编码到一个连续的潜在空间。在低分辨率潜在表示的条件下,CFM机制逐步从高斯先验利用一步常微分方程(ODE)求解器生成相应的高分辨率潜在表示。然后,由预训练的自动编码器对所得到的高分辨率潜在表示进行解码,以重建高分辨率音频。实验结果表明,LatentFlowSR在各种音频类型和超分辨率设置中始终优于基线方法。实验结果表明,该方法具有较强的高频重构能力和鲁棒的泛化能力,为潜空间建模在音频超分辨率中的有效性提供了有力的证据。所有相关代码将在完成论文评审过程后公开提供。
摘要:Audio super-resolution aims to recover missing high-frequency details from bandwidth-limited low-resolution audio, thereby improving the naturalness and perceptual quality of the reconstructed signal. However, most existing methods directly operate in the waveform or time-frequency domain, which not only involves high-dimensional generation spaces but is also largely limited to speech tasks, leaving substantial room for improvement on more complex audio types such as sound effects and music. To mitigate these limitations, we introduce LatentFlowSR, a new audio super-resolution approach that leverages conditional flow matching (CFM) within a latent representation space. Specifically, we first train a noise-robust autoencoder, which encodes low-resolution audio into a continuous latent space. Conditioned on the low-resolution latent representation, a CFM mechanism progressively generates the corresponding high-resolution latent representation from a Gaussian prior with a one-step ordinary differential equation (ODE) solver. The resulting high-resolution latent representation is then decoded by the pretrained autoencoder to reconstruct the high-resolution audio. Experimental results demonstrate that LatentFlowSR consistently outperforms baseline methods across various audio types and super-resolution settings. These results indicate that the proposed method possesses strong high-frequency reconstruction capability and robust generalization performance, providing compelling evidence for the effectiveness of latent-space modeling in audio super-resolution. All relevant code will be made publicly available upon completion of the paper review process.


【5】Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
标题:交互式ASB:面向语音识别的类人交互和语义一致性评估
链接:https://arxiv.org/abs/2604.09121

作者:Peng Wang,Yanqiao Zhu,Zixuan Jiang,Qinyuan Chen,Xingjian Zhao,Xipeng Qiu,Wupeng Wang,Zhifu Gao,Xiangang Li,Kai Yu,Xie Chen
摘要:近年来,在模型架构和大规模训练数据的推动下,自动语音识别(ASR)取得了显着进展。然而,有两个重要方面仍然没有得到充分探讨。首先,词错误率(WER),几十年来占主导地位的评估指标,平等地对待所有的单词,往往不能反映句子层次上的话语的语义正确性。第二,交互式校正-人类沟通的重要组成部分-很少被系统地研究在ASR的研究。在本文中,我们将这两个角度下的互动ASR的代理框架。我们建议利用LLM-as-a-Judge作为语义感知的评估指标来评估识别质量,超越标记级别的准确性。此外,我们设计了一个LLM驱动的代理框架来模拟类人的多回合交互,通过语义反馈实现识别输出的迭代细化。在标准基准测试上进行了广泛的实验,包括GigaSpeech(英语),WenetSpeech(中文),ASRU 2019代码转换测试集。客观和主观的评价表明,所提出的框架在提高语义保真度和交互校正能力的有效性。我们将发布代码,以促进互动和代理ASR的未来研究。
摘要:Recent years have witnessed remarkable progress in automatic speech recognition (ASR), driven by advances in model architectures and large-scale training data. However, two important aspects remain underexplored. First, Word Error Rate (WER), the dominant evaluation metric for decades, treats all words equally and often fails to reflect the semantic correctness of an utterance at the sentence level. Second, interactive correction-an essential component of human communication-has rarely been systematically studied in ASR research. In this paper, we integrate these two perspectives under an agentic framework for interactive ASR. We propose leveraging LLM-as-a-Judge as a semantic-aware evaluation metric to assess recognition quality beyond token-level accuracy. Furthermore, we design an LLM-driven agent framework to simulate human-like multi-turn interaction, enabling iterative refinement of recognition outputs through semantic feedback. Extensive experiments are conducted on standard benchmarks, including GigaSpeech (English), WenetSpeech (Chinese), the ASRU 2019 code-switching test set. Both objective and subjective evaluations demonstrate the effectiveness of the proposed framework in improving semantic fidelity and interactive correction capability. We will release the code to facilitate future research in interactive and agentic ASR.


【6】Few-Shot Contrastive Adaptation for Audio Abuse Detection in Low-Resource Indic Languages
标题:用于低资源印度语音频滥用检测的Few-Shot对比自适应
链接:https://arxiv.org/abs/2604.09094

作者:Aditya Narayan Sankaran,Reza Farahbakhsh,Noel Crespi
备注:14 pages, preprint under review
摘要:随着社交媒体转向基于语音的交互,滥用语音检测变得越来越重要,特别是在多语言和低资源环境中。目前大多数系统依赖于自动语音识别(ASR),然后是基于文本的仇恨语音分类,但这种管道容易出现转录错误,并丢弃语音中携带的韵律信息。我们研究对比存储音频预训练(CLAP)是否可以支持直接从音频中检测辱骂性语音。使用ADIMA数据集,我们评估了基于CLAP的表示下,Few-Shot监督对比适应跨语言和leave-one-language-out设置,与zero-shot提示包括作为辅助分析。我们的研究结果表明,CLAP在10种印度语言中产生了强大的跨语言音频表示,并且轻量级的仅投影自适应在完整的训练数据上训练的完全监督系统方面具有竞争力的性能。然而,Few-Shot自适应的好处是依赖于语言的,并且不随镜头大小而单调。这些研究结果表明,对比音频文本模型提供了一个很有前途的基础,跨语言音频滥用检测在低资源环境中,同时也表明,转让仍然不完整和语言特定的重要方式。
摘要:Abusive speech detection is becoming increasingly important as social media shifts towards voice-based interaction, particularly in multilingual and low-resource settings. Most current systems rely on automatic speech recognition (ASR) followed by text-based hate speech classification, but this pipeline is vulnerable to transcription errors and discards prosodic information carried in speech. We investigate whether Contrastive Language-Audio Pre-training (CLAP) can support abusive speech detection directly from audio. Using the ADIMA dataset, we evaluate CLAP-based representations under few-shot supervised contrastive adaptation in cross-lingual and leave-one-language-out settings, with zero-shot prompting included as an auxiliary analysis. Our results show that CLAP yields strong cross-lingual audio representations across ten Indic languages, and that lightweight projection-only adaptation achieves competitive performance with respect to fully supervised systems trained on complete training data. However, the benefits of few-shot adaptation are language-dependent and not monotonic with shot size. These findings suggest that contrastive audio-text models provide a promising basis for cross-lingual audio abuse detection in low-resource settings, while also indicating that transfer remains incomplete and language-specific in important ways.


【7】Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
标题:Tora 3:具有物理一致性的轨迹引导音频视频生成
链接:https://arxiv.org/abs/2604.09057

作者:Junchao Liao,Zhenghao Zhang,Xiangyu Meng,Litao Li,Ziying Zhang,Siyu Zhu,Long Qin,Weizhi Wang
摘要:音视频(AV)生成最近在感知质量和多模态一致性方面取得了很大进展,但生成具有合理运动-声音关系的内容仍然具有挑战性。现有方法通常产生视觉上不稳定的对象运动和仅与显著运动或接触事件松散对齐的声音,这主要是因为它们缺乏由视频和音频生成共享的明确的运动感知结构。我们提出了Tora 3,一个自动引导的AV生成框架,通过使用对象轨迹作为共享的运动学先验来提高物理一致性。Tora 3没有将轨迹视为仅限视频的控制信号,而是使用它们来联合引导视觉运动和声学事件。具体而言,我们设计了一个基于运动学的视频运动表示,一个基于运动学的二阶运动学状态驱动的运动学-音频对齐模块,以及一个混合流匹配方案,该方案在保持局部相干性的同时,在基于运动学的条件区域中保持轨迹保真度。我们进一步策划了PAV,这是一个大规模的AV数据集,强调运动相关模式,并自动提取运动注释。大量的实验表明,Tora 3在强大的开源基线上提高了运动真实感,运动声音同步和整体AV生成质量。
摘要:Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible motion-sound relations remains challenging. Existing methods often produce object motions that are visually unstable and sounds that are only loosely aligned with salient motion or contact events, largely because they lack an explicit motion-aware structure shared by video and audio generation. We present Tora3, a trajectory-guided AV generation framework that improves physical coherence by using object trajectories as a shared kinematic prior. Rather than treating trajectories as a video-only control signal, Tora3 uses them to jointly guide visual motion and acoustic events. Specifically, we design a trajectory-aligned motion representation for video, a kinematic-audio alignment module driven by trajectory-derived second-order kinematic states, and a hybrid flow matching scheme that preserves trajectory fidelity in trajectory-conditioned regions while maintaining local coherence elsewhere. We further curate PAV, a large-scale AV dataset emphasizing motion-relevant patterns with automatically extracted motion annotations. Extensive experiments show that Tora3 improves motion realism, motion-sound synchronization, and overall AV generation quality over strong open-source baselines.


【8】AccompGen: Hierarchical Autoregressive Vocal Accompaniment Generation with Dual-Rate Codec Tokenization
标题:AccomGen:采用双速率编解码器令牌化的分层自回归声乐伴奏生成
链接:https://arxiv.org/abs/2604.09054

作者:Jian Zhu,Jianwei Cui,Shihao Chen,Yubang Zhang,Cheng Luo
摘要:我们提出AccompGen,一个系统,产生器乐音频陪同输入人声。给定孤立的歌声,AccompGen产生一个连贯的乐器伴奏,可以直接与输入混合,以创建完整的音乐。我们提出了三个关键的创新在以前的工作:(1)一个双速率编解码器令牌化方案,使用HuBERT语义令牌在50,Hz的人声和EnCodec声学令牌在75,Hz的乐器,使时间对齐但速率无关的建模;(2)一个三阶段的分层自回归架构(语义到粗声学到细声学),具有交织多码本预测和无分类器引导;以及(3)现代Transformer设计选择,包括QK范数、GEGLU激活、RMS范数和T5风格的相对位置偏置,用于改进训练稳定性和序列泛化。
摘要:We present AccompGen, a system that generates instrumental music audio to accompany input vocals. Given isolated singing voice, AccompGen produces a coherent instrumental accompaniment that can be directly mixed with the input to create complete music. We propose three key innovations over prior work: (1) a dual-rate codec tokenization scheme using HuBERT semantic tokens at 50,Hz for vocals and EnCodec acoustic tokens at 75,Hz for instrumentals, enabling time-aligned yet rate-independent modeling; (2) a three-stage hierarchical autoregressive architecture (semantic to coarse acoustic to fine acoustic) with interleaved multi-codebook prediction and classifier-free guidance; and (3) modern Transformer design choices including QK-norm, GEGLU activations, RMSNorm, and T5-style relative position bias for improved training stability and sequence generalization.


【9】Noise-Aware In-Context Learning for Hallucination Mitigation in ALLMs
标题:噪音感知的上下文学习以缓解ALLM中的幻觉
链接:https://arxiv.org/abs/2604.09021

作者:Qixuan Huang,Khalid Zaman,Masashi Unoki
摘要:听觉大语言模型(ALLM)在音频理解和推理任务中表现出强大的通用能力。然而,他们的可靠性仍然受到幻觉问题的影响。现有的幻觉评估方法被制定为二进制分类任务,这是不足以表征更复杂的幻觉模式中出现的生成任务。此外,目前的幻觉缓解策略依赖于微调,导致高计算成本。为了解决上述限制,我们提出了一种即插即用的噪声感知上下文学习(NAICL)方法。具体而言,我们构建了一个噪声先验库,检索与输入音频相关的噪声示例,并将其作为上下文先验,从而指导模型在声学证据不足时减少推测性关联,并采用更保守的生成策略。此外,我们还建立了音频字幕任务的幻听基准,包括Clotho-1 K多事件基准数据集的构建、四种幻听类型的定义,以及引入幻听类型分布等指标来支持细粒度分析。实验结果表明,所有评估ALLM表现出相同的幻觉行为。此外,提出的NAICL方法将总体幻觉率从26.53%降低到16.98%。
摘要:Auditory large language models (ALLMs) have demonstrated strong general capabilities in audio understanding and reasoning tasks. However, their reliability is still undermined by hallucination issues. Existing hallucination evaluation methods are formulated as binary classification tasks, which are insufficient to characterize the more complex hallucination patterns that arise in generative tasks. Moreover, current hallucination mitigation strategies rely on fine-tuning, resulting in high computational costs. To address the above limitations, we propose a plug-and-play Noise-Aware In-Context Learning (NAICL) method. Specifically, we construct a noise prior library, retrieve noise examples relevant to the input audio, and incorporate them as contextual priors, thereby guiding the model to reduce speculative associations when acoustic evidence is insufficient and to adopt a more conservative generation strategy. In addition, we establish a hallucination benchmark for audio caption tasks including the construction of the Clotho-1K multi-event benchmark dataset, the definition of four types of auditory hallucinations, and the introduction of metrics such as hallucination type distribution to support fine-grained analysis. Experimental results show that all evaluated ALLMs exhibit same hallucination behaviors. Moreover, the proposed NAICL method reduces the overall hallucination rate from 26.53% to 16.98%.


【10】Accessible Fine-grained Data Representation via Spatial Audio
标题:通过空间音频的可访问细粒度数据表示
链接:https://arxiv.org/abs/2604.08979

作者:Can Liu,Wenjie Jiang,Shaolun Ruan,Kotaro Hara,Yong Wang
备注:Accepted by IEEE Computer Graphics and Applications (IEEE CG&A)
摘要:定量数据的基于音高的声音化增加了盲人和低视力(BLV)个体无法获得的数据可视化的可访问性。我们认为,虽然音高表示可以揭示数据的粗粒度信息,如数据趋势和值比较,但它们不能有效地传达细粒度的细节,如单个数据点的符号和确切值。根据现有的声音感知研究,我们提出了一种基于空间音频的方法,通过将数据值表示为方位平面中的声音方向,以实现可访问的细粒度数据表示。我们对26名参与者(包括10名BLV参与者)进行了四项数据感知任务的用户研究。结果表明,我们的方法在细粒度数据感知任务(如识别数据符号和精确值)上的性能明显优于音高表示,并且在数据趋势识别上的性能相似,尽管其在数据值比较上的准确性较差。
摘要:Pitch-based sonification of quantitative data increases the accessibility of data visualizations that are otherwise inaccessible for blind and low-vision (BLV) individuals. We argue that, although pitch representations can reveal the coarse-grained information of data, such as data trend and value comparison, they cannot effectively convey the fine-grained details like the sign and exact value of individual data points. Informed by existing sound perception research, we propose a spatial audio-based approach by representing data values as the sound direction in the azimuth plane to achieve accessible fine-grained data representation. We conducted a user study with 26 participants (including 10 BLV participants) on four data perception tasks. The results show our approach significantly outperforms pitch representation on fine-grained data perception tasks like recognizing data signs and exact values, and performs similarly on data trend identification, despite its inferior accuracy on data value comparison.


【11】AudioGS: Spectrogram-Based Audio Gaussian Splatting for Sound Field Reconstruction
标题:AudioGS:基于谱图的音频高斯飞溅用于声学场重建
链接:https://arxiv.org/abs/2604.08967

作者:Chunhao Bi,Houqiang Zhong,Zhixin Xu,Li Song,Zhengxue Cheng
摘要:空间音频是沉浸式虚拟体验的基础,但从稀疏的观察合成高保真双耳音频仍然是一个重大挑战。现有的方法通常依赖于以视觉先验为条件的隐式神经表示,这通常难以捕获细粒度的声学结构。受3D高斯飞溅(3DGS)的启发,我们引入了AudioGS,这是一种新颖的无视觉框架,它基于频谱图将声场显式编码为一组音频高斯。AudioGS将每个时间频率仓与配备有双球谐(SH)系数和衰减系数的音频高斯相关联。对于目标姿态,我们通过评估SH场来捕获方向性,结合几何引导的距离衰减和相位校正,并重建波形来渲染双耳音频。在Replay-NVAS数据集上的实验表明,AudioGS成功地捕获了复杂的空间线索,并优于最先进的视觉依赖基线。具体而言,与性能最佳的视觉引导方法相比,AudioGS将幅度重建误差(MAG)降低了14%以上,并将感知质量度量(DPAM)降低了约25%。
摘要:Spatial audio is fundamental to immersive virtual experiences, yet synthesizing high-fidelity binaural audio from sparse observations remains a significant challenge. Existing methods typically rely on implicit neural representations conditioned on visual priors, which often struggle to capture fine-grained acoustic structures. Inspired by 3D Gaussian Splatting (3DGS), we introduce AudioGS, a novel visual-free framework that explicitly encodes the sound field as a set of Audio Gaussians based on spectrograms. AudioGS associates each time-frequency bin with an Audio Gaussian equipped with dual Spherical Harmonic (SH) coefficients and a decay coefficient. For a target pose, we render binaural audio by evaluating the SH field to capture directionality, incorporating geometry-guided distance attenuation and phase correction, and reconstructing the waveform. Experiments on the Replay-NVAS dataset demonstrate that AudioGS successfully captures complex spatial cues and outperforms state-of-the-art visual-dependent baselines. Specifically, AudioGS reduces the magnitude reconstruction error (MAG) by over 14% and reduces the perceptual quality metric (DPAM) by approximately 25% compared to the best performing visual-guided method.


【12】AudioGuard: Toward Comprehensive Audio Safety Protection Across Diverse Threat Models
标题:AudioGuard:针对各种威胁模型提供全面的音频安全保护
链接:https://arxiv.org/abs/2604.08867

作者:Mintong Kang,Chen Fang,Bo Li
摘要:音频已迅速成为基础模型的主要接口,为实时语音助手提供支持。确保音频系统的安全性本质上比“大声说出的不安全文本”更复杂:现实世界的风险可能取决于音频原生有害声音事件、扬声器属性(例如,童声)、模仿/语音克隆滥用以及语音内容合成危害,例如童声加上性内容。音频的性质使得针对这种独特的风险环境制定全面的基准或护栏具有挑战性。为了缩小这一差距,我们对音频系统进行了大规模的红色团队,系统地发现音频中的漏洞,并开发了一个全面的、基于政策的音频风险分类和AudioSafetyBench,这是第一个基于政策的音频安全基准,适用于各种威胁模型。AudioSafetyBench支持多种语言,可疑的声音(例如,名人/模仿和儿童语音)、有风险的语音内容组合以及非语音声音事件。为了抵御这些威胁,我们提出了AudioGuard,这是一个统一的护栏,包括1)用于波形级音频原生检测的SoundGuard和2)用于基于策略的语义保护的ContentGuard。在AudioSafetyBench和四个补充基准测试上进行的广泛实验表明,AudioGuard在基于强大音频LLM的基线上持续提高了护栏精度,并且延迟大大降低。
摘要:Audio has rapidly become a primary interface for foundation models, powering real-time voice assistants. Ensuring safety in audio systems is inherently more complex than just "unsafe text spoken aloud": real-world risks can hinge on audio-native harmful sound events, speaker attributes (e.g., child voice), impersonation/voice-cloning misuse, and voice-content compositional harms, such as child voice plus sexual content. The nature of audio makes it challenging to develop comprehensive benchmarks or guardrails against this unique risk landscape. To close this gap, we conduct large-scale red teaming on audio systems, systematically uncover vulnerabilities in audio, and develop a comprehensive, policy-grounded audio risk taxonomy and AudioSafetyBench, the first policy-based audio safety benchmark across diverse threat models. AudioSafetyBench supports diverse languages, suspicious voices (e.g., celebrity/impersonation and child voice), risky voice-content combinations, and non-speech sound events. To defend against these threats, we propose AudioGuard, a unified guardrail consisting of 1) SoundGuard for waveform-level audio-native detection and 2) ContentGuard for policy-grounded semantic protection. Extensive experiments on AudioSafetyBench and four complementary benchmarks show that AudioGuard consistently improves guardrail accuracy over strong audio-LLM-based baselines with substantially lower latency.


【13】Script Collapse in Multilingual ASR: Defining and Measuring Script Fidelity Rate
标题:多语言ASB中的脚本崩溃:定义和测量脚本保真度
链接:https://arxiv.org/abs/2604.08786

作者:Hanif Rahman
摘要:单词错误率(WER)是自动语音识别的主要指标,但它无法检测系统故障模式:在错误的书写系统中产生流畅输出的模型。我们定义了脚本保真度率(SFR),即目标脚本块中假设字符的比例,无需参考转录即可计算,并报告了跨四种书写系统(普什图语、乌尔都语、印地语、孟加拉语、马拉雅拉姆语、索马里语)的六种语言的脚本崩溃的首次系统测量。)和FLEURS测试集上的九个ASR模型。在53个评估的模型-语言对中,18个(34%; 95% Wilson CI:23-47%)表现出脚本崩溃(SFR < 10%); MMS-1B和M4 T-v2在评估的每种语言上保持SFR高于99%,证实SFR正确识别高保真度。我们确定了三种不同的崩溃模式:拉丁语音替代(较小的耳语对印度语言),阿拉伯语替代索马里的拉丁字母正字法,和梵文替代较大的耳语模型将所有印度语音频作为印地语,即使在耳语大V3中也存在失败。
摘要:Word error rate (WER) is the dominant metric for automatic speech recognition, yet it cannot detect a systematic failure mode: models that produce fluent output in the wrong writing system. We define Script Fidelity Rate (SFR), the fraction of hypothesis characters in the target script block, computable without reference transcriptions, and report the first systematic measurement of script collapse across six languages spanning four writing systems (Pashto, Urdu, Hindi, Bengali, Malayalam, Somali) and nine ASR models on FLEURS test sets. Across 53 evaluated model-language pairs, 18 (34%; 95% Wilson CI: 23-47%) exhibit script collapse (SFR < 10%); MMS-1B and SeamlessM4T-v2 maintain SFR above 99% on every language evaluated, confirming that SFR correctly identifies high fidelity where it is present. We identify three distinct collapse patterns: Latin phonetic substitution (smaller Whisper on Indic languages), Arabic substitution for Somali's Latin-script orthography, and Devanagari substitution where larger Whisper models treat all Indic audio as Hindi, a failure present even in Whisper large-v3.


【14】Neural networks for Text-to-Speech evaluation
标题:用于文本到语音评估的神经网络
链接:https://arxiv.org/abs/2604.08562

作者:Ilya Trofimenko,David Kocharyan,Aleksandr Zaitsev,Pavel Repnikov,Mark Levin,Nikita Shevtsov
摘要:确保文本到语音(TTS)系统提供大规模的人类感知质量是现代语音技术的核心挑战。人类主观评价协议,如平均意见评分(MOS)和并排(SBS)比较仍然是事实上的黄金标准,但它们是昂贵的,缓慢的,并对普遍的评估偏见敏感。本研究通过制定和实施一套新颖的神经模型来解决这些障碍,这些模型旨在近似相对(SBS)和绝对(MOS)设置中的专家判断。对于相对评估,我们提出了NeuralSBS,这是一个HuBERT支持的模型,准确率达到73.7%(在SOMOS数据集上)。对于绝对评估,我们引入了使用自定义序列长度优化的MOSNet增强功能,以及WhisperBert,这是一种多模态堆叠集成,通过弱学习器将Whisper音频功能和BERT文本嵌入相结合。我们最好的MOS模型实现了~0.40的均方根误差(RMSE),显著优于0.62的人类评分员间RMSE基线。此外,我们的消融研究表明,通过交叉注意天真地融合文本会降低性能,突出了基于集成的堆叠比直接潜在融合的有效性。我们还报告了基于SpeechLM的架构和zero-shot LLM评估器(Qwen 2-Audio,Gemini 2.5 flash预览)的负面结果,加强了专用度量学习框架的必要性。
摘要:Ensuring that Text-to-Speech (TTS) systems deliver human-perceived quality at scale is a central challenge for modern speech technologies. Human subjective evaluation protocols such as Mean Opinion Score (MOS) and Side-by-Side (SBS) comparisons remain the de facto gold standards, yet they are expensive, slow, and sensitive to pervasive assessor biases. This study addresses these barriers by formulating, and implementing a suite of novel neural models designed to approximate expert judgments in both relative (SBS) and absolute (MOS) settings. For relative assessment, we propose NeuralSBS, a HuBERT-backed model achieving 73.7% accuracy (on SOMOS dataset). For absolute assessment, we introduce enhancements to MOSNet using custom sequence-length batching, as well as WhisperBert, a multimodal stacking ensemble that combines Whisper audio features and BERT textual embeddings via weak learners. Our best MOS models achieve a Root Mean Square Error (RMSE) of ~0.40, significantly outperforming the human inter-rater RMSE baseline of 0.62. Furthermore, our ablation studies reveal that naively fusing text via cross-attention can degrade performance, highlighting the effectiveness of ensemble-based stacking over direct latent fusion. We additionally report negative results with SpeechLM-based architectures and zero-shot LLM evaluators (Qwen2-Audio, Gemini 2.5 flash preview), reinforcing the necessity of dedicated metric learning frameworks.


eess.AS音频处理


【1】Data Selection Effects on Self-Supervised Learning of Audio Representations for French Audiovisual Broadcasts
标题:数据选择对法国视听广播音频表示的自我监督学习的影响
链接:https://arxiv.org/abs/2604.09472

作者:Valentin Pelloin,Lina Bekkali,Reda Dehak,David Doukhan
备注:To be published in the Fifteenth International Conference on Language Resources and Evaluation (LREC 2026)
摘要:音频和语音自监督编码器模型现在广泛用于许多不同的任务。这些模型中的许多通常是在干净的分段语音内容上训练的,例如LibriSpeech。在本文中,我们研究了这种SSL(自我监督学习)模型的预训练数据集如何影响其下游结果。我们建立了一个大型的预训练语料库的高度多样化的电视和广播音频内容,我们描述与自动工具。我们使用这些注释来构建更小的子集,我们使用这些子集来训练音频SSL模型。然后,我们评估多个下游任务的模型,如自动语音识别,语音活动和音乐检测,或说话人识别。结果显示了在不同音频内容上预训练SSL模型的潜力,而不限于语音。我们还执行了一个成员推理攻击,以评估编码器的能力,记住他们的训练数据集,这突出了重复数据删除的重要性。这种统一的训练可以连接语音和音乐机器学习社区。
摘要:Audio and speech self-supervised encoder models are now widely used for a lot of different tasks. Many of these models are often trained on clean segmented speech content such as LibriSpeech. In this paper, we look into how the pretraining datasets of such SSL (Self-Supervised Learning) models impact their downstream results. We build a large pretraining corpus of highly diverse TV and Radio broadcast audio content, which we describe with automatic tools. We use these annotations to build smaller subsets, which we use to train audio SSL models. Then, we evaluate the models on multiple downstream tasks such as automatic speech recognition, voice activity and music detection, or speaker recognition. The results show the potential of pretraining SSL models on diverse audio content without restricting it to speech. We also perform a membership inference attack to evaluate the encoder ability to memorize their training datasets, which highlight the importance of data deduplication. This unified training could bridge speech and music machine learning communities.


【2】Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models
标题:使用语言模型实现多干音乐源分离的离散令牌建模
链接:https://arxiv.org/abs/2604.09371

作者:Pengbo Lyu,Xiangyu Zhao,Chengwei Liu,Haoyin Yan,Xiaotao Liang,Hongyu Wang,Shaofei Xue
备注:5 pages, 2 figures, 3 tables. Submitted to INTERSPEECH 2026
摘要:我们提出了一个生成框架多轨音乐源分离(MSS),重新制定的任务作为条件离散令牌生成。与直接在时域或频域中估计连续信号的传统方法不同,我们的方法结合了基于Conformer的条件编码器,双路径神经音频编解码器(HCodec)和仅解码器语言模型,以自回归方式为四个目标轨道生成音频令牌。生成的令牌通过编解码器解码回波形。在MUSDB 18-HQ基准测试中的评估表明,我们的生成方法实现了接近最先进的判别方法的感知质量,同时在人声轨道上获得了最高的NISQA分数。消融研究证实了可学习的Conformer编码器的有效性和顺序交叉轨迹生成的好处。
摘要:We propose a generative framework for multi-track music source separation (MSS) that reformulates the task as conditional discrete token generation. Unlike conventional approaches that directly estimate continuous signals in the time or frequency domain, our method combines a Conformer-based conditional encoder, a dual-path neural audio codec (HCodec), and a decoder-only language model to autoregressively generate audio tokens for four target tracks. The generated tokens are decoded back to waveforms through the codec decoder. Evaluation on the MUSDB18-HQ benchmark shows that our generative approach achieves perceptual quality approaching state-of-the-art discriminative methods, while attaining the highest NISQA score on the vocals track. Ablation studies confirm the effectiveness of the learnable Conformer encoder and the benefit of sequential cross-track generation.


【3】Phonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASR
标题:音素与投影仪:基于LLM的ASB语音语言接口的调查
链接:https://arxiv.org/abs/2604.09332

作者:Ziwei Li,Lukuang Dong,Saierdaer Yusuyin,Xianyu Zhao,Zhijian Ou
备注:Update after INTERSPEECH2026 submission
摘要:将预训练的语音编码器与大型语言模型(LLM)集成在一起对于ASR来说是有希望的,但性能和数据效率取决于语音语言接口。一个常见的选择是一个学习投影仪,映射编码器的功能到LLM嵌入空间,而另一种选择是暴露离散音素序列的LLM。使用相同的编码器和LLM骨干,我们比较基于音素和香草投影仪为基础的接口在高资源英语和低资源鞑靼语。我们还提出了一个BPE音素接口组频繁的本地音素模式,同时保留显式的字边界线索音素到字素生成。在LibriSpeech上,基于音素的接口与vanilla projector竞争,BPE音素接口产生进一步的收益。在鞑靼语上,基于音素的界面大大优于普通投影仪。我们进一步发现,音素监督产生一个音素知情的混合接口,比香草投影仪更强。
摘要:Integrating pretrained speech encoders with large language models (LLMs) is promising for ASR, but performance and data efficiency depend on the speech-language interface. A common choice is a learned projector that maps encoder features into the LLM embedding space, whereas an alternative is to expose discrete phoneme sequences to the LLM. Using the same encoder and LLM backbones, we compare phoneme-based and vanilla projector-based interfaces in high-resource English and low-resource Tatar. We also propose a BPE-phoneme interface that groups frequent local phoneme patterns while preserving explicit word-boundary cues for phoneme-to-grapheme generation. On LibriSpeech, the phoneme-based interface is competitive with the vanilla projector, and the BPE-phoneme interface yields further gains. On Tatar, the phoneme-based interface substantially outperforms the vanilla projector. We further find that phoneme supervision yields a phoneme-informed hybrid interface that is stronger than the vanilla projector.


【4】PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing
标题:PS-TTC:文本到语音的语音同步,以实现自然自动配音
链接:https://arxiv.org/abs/2604.09111

作者:Changi Hong,Yoonah Song,Hwayoung Park,Chaewoon Bang,Dayeon Gu,Do Hyun Lee,Hong Kook Kim
备注:Accepted to ICPR 2026
摘要:最近,基于人工智能的配音技术已经发展,使得自动配音(AD)能够将视频的源语音转换为不同语言的目标语音。然而,自然广告仍然面临同步挑战,例如持续时间和嘴唇同步(嘴唇同步),这对于保持观众体验至关重要。因此,本文提出了一种同步方法的AD过程,释义翻译文本,包括两个步骤:等时的时间约束和语音同步(PS),以保持唇同步。首先,我们通过用语言模型对翻译文本进行释义来实现等时性,确保目标语音的时长与源语音的时长相匹配。其次,我们引入PS,它采用动态时间规整(DTW)与本地成本的元音距离从训练数据中测量,使目标文本组成元音发音类似源元音。第三,我们将这种方法扩展到PSComet,它联合考虑语义和语音的相似性,以更好地保留意义。所提出的方法被纳入到文本到语音系统,PS-TTS和PS-Comet TTS。使用韩语和英语唇读数据集和配音演员配音数据集的性能评估表明,这两个系统在几个客观指标上优于没有PS的TTS,并且在韩语到英语和英语到韩语配音中优于配音演员。我们将实验扩展到法语,测试这些语言之间的所有对,以评估跨语言的适用性。在所有的语言对中,PS-Comet表现最好,平衡了唇音同步的准确性和语义保留,证实了PS-Comet比单独的PS实现了更准确的唇音同步和语义保留。
摘要:Recently, artificial intelligence-based dubbing technology has advanced, enabling automated dubbing (AD) to convert the source speech of a video into target speech in different languages. However, natural AD still faces synchronization challenges such as duration and lip-synchronization (lip-sync), which are crucial for preserving the viewer experience. Therefore, this paper proposes a synchronization method for AD processes that paraphrases translated text, comprising two steps: isochrony for timing constraints and phonetic synchronization (PS) to preserve lip-sync. First, we achieve isochrony by paraphrasing the translated text with a language model, ensuring the target speech duration matches that of the source speech. Second, we introduce PS, which employs dynamic time warping (DTW) with local costs of vowel distances measured from training data so that the target text composes vowels with pronunciations similar to source vowels. Third, we extend this approach to PSComet, which jointly considers semantic and phonetic similarity to preserve meaning better. The proposed methods are incorporated into text-to-speech systems, PS-TTS and PS-Comet TTS. The performance evaluation using Korean and English lip-reading datasets and a voice-actor dubbing dataset demonstrates that both systems outperform TTS without PS on several objective metrics and outperform voice actors in Korean-to-English and English-to-Korean dubbing. We extend the experiments to French, testing all pairs among these languages to evaluate cross-linguistic applicability. Across all language pairs, PS-Comet performed best, balancing lip-sync accuracy with semantic preservation, confirming that PS-Comet achieves more accurate lip-sync with semantic preservation than PS alone.


【5】Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning
标题:通过级联预算和基于ICL的在线强化学习增强对话TTC
链接:https://arxiv.org/abs/2604.08709

作者:Zhicheng Ouyang,Seong-Gyun Leem,Bach Viet Do,Haibin Wu,Ariya Rastrow,Yuzong Liu,Florian Metze
摘要:会话人工智能已经取得了重大进展,但生成具有表达力和可控性的文本到语音(TTS)仍然具有挑战性。具体来说,控制细粒度的语音风格和情感是非常困难的,通常需要大量的大量注释训练数据。为了克服这个数据瓶颈,我们提出了一个可扩展的,数据高效的级联框架,将文本风格的令牌与人工策划的高质量音频提示配对。这种方法使单镜头适应细粒度的说话风格和人物的声音。在TTS的上下文中,这种音频提示充当上下文学习(ICL),指导模型的韵律和音色,而不需要大量的参数更新或大规模的重新训练。为了进一步提高生成质量并减轻幻觉,我们引入了一种新的基于ICL的在线强化学习(RL)策略。该策略使用主观审美奖励直接优化自回归韵律模型,同时受到联结主义时间分类(CTC)对齐的约束,以保持可理解性。综合的人类感知评估表明,在合成语音的自然度和表现力方面都有显着改善,建立了我们基于ICL的在线RL方法的有效性。
摘要:Conversational AI has made significant progress, yet generating expressive and controllable text-to-speech (TTS) remains challenging. Specifically, controlling fine-grained voice styles and emotions is notoriously difficult and typically requires massive amounts of heavily annotated training data. To overcome this data bottleneck, we present a scalable, data-efficient cascaded framework that pairs textual style tokens with human-curated, high-quality audio prompts. This approach enables single-shot adaptation to fine-grained speaking styles and character voices. In the context of TTS, this audio prompting acts as In-Context Learning (ICL), guiding the model's prosody and timbre without requiring massive parameter updates or large-scale retraining. To further enhance generation quality and mitigate hallucinations, we introduce a novel ICL-based online reinforcement learning (RL) strategy. This strategy directly optimizes the autoregressive prosody model using subjective aesthetic rewards while being constrained by Connectionist Temporal Classification (CTC) alignment to preserve intelligibility. Comprehensive human perception evaluations demonstrate significant improvements in both the naturalness and expressivity of the synthesized speech, establishing the efficacy of our ICL-based online RL approach.


【6】DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
标题:对话西顿:从野外对话音频中恢复全速对话曲目
链接:https://arxiv.org/abs/2604.09344

作者:Wataru Nakata,Yuki Saito,Kazuki Yamauchi,Emiru Tsunoo,Hiroshi Saruwatari
备注:12 pages, 2 figures
摘要:全双工对话音频,其中每个扬声器被记录在单独的轨道上,是口语对话研究的重要资源,但难以大规模收集。大多数野外的两个扬声器对话只能作为退化的单声道混合信号,这使得它不适合需要清晰的扬声器信号的系统。我们提出了DialogueSidon,一个模型,用于联合恢复和分离退化的单声道两个扬声器对话音频。DialogueSidon结合了变分自动编码器(VAE)对语音自监督学习(SSL)模型特征进行操作,该特征将SSL模型特征压缩到紧凑的潜在空间中,并结合了基于扩散的潜在预测器,该预测器从降级的混合物中恢复说话者的潜在表示。在英语、多语言和野外对话数据集上的实验表明,DialogueSidon大大提高了基线的可理解性和分离质量,同时也实现了更快的推理。
摘要:Full-duplex dialogue audio, in which each speaker is recorded on a separate track, is an important resource for spoken dialogue research, but is difficult to collect at scale. Most in-the-wild two-speaker dialogue is available only as degraded monaural mixtures, making it unsuitable for systems requiring clean speaker-wise signals. We propose DialogueSidon, a model for joint restoration and separation of degraded monaural two-speaker dialogue audio. DialogueSidon combines a variational autoencoder (VAE) operates on the speech self-supervised learning (SSL) model feature, which compresses SSL model features into a compact latent space, with a diffusion-based latent predictor that recovers speaker-wise latent representations from the degraded mixture. Experiments on English, multilingual, and in-the-wild dialogue datasets show that DialogueSidon substantially improves intelligibility and separation quality over a baseline, while also achieving much faster inference.


【7】Script Collapse in Multilingual ASR: Defining and Measuring Script Fidelity Rate
标题:多语言ASB中的脚本崩溃:定义和测量脚本保真度
链接:https://arxiv.org/abs/2604.08786

作者:Hanif Rahman
摘要:单词错误率(WER)是自动语音识别的主要指标,但它无法检测系统故障模式:在错误的书写系统中产生流畅输出的模型。我们定义脚本保真度率(SFR),假设字符在目标脚本块中的比例,可计算的参考transmittance,并报告了第一个系统的测量跨六种语言的脚本崩溃跨越四个书写系统(普什图语,乌尔都语,印地语,孟加拉语,马拉雅拉姆语,索马里语)和九个ASR模型的FLEURS测试集。在53个评估的模型-语言对中,18个(34%; 95% Wilson CI:23-47%)表现出脚本崩溃(SFR < 10%); MMS-1B和M4 T-v2在评估的每种语言上保持SFR高于99%,证实SFR正确识别高保真度。我们确定了三种不同的崩溃模式:拉丁语音替代(较小的耳语对印度语言),阿拉伯语替代索马里的拉丁字母正字法,以及梵文替代较大的耳语模型将所有印度语音频视为印地语,即使在耳语大V3中也存在失败。
摘要:Word error rate (WER) is the dominant metric for automatic speech recognition, yet it cannot detect a systematic failure mode: models that produce fluent output in the wrong writing system. We define Script Fidelity Rate (SFR), the fraction of hypothesis characters in the target script block, computable without reference transcriptions, and report the first systematic measurement of script collapse across six languages spanning four writing systems (Pashto, Urdu, Hindi, Bengali, Malayalam, Somali) and nine ASR models on FLEURS test sets. Across 53 evaluated model-language pairs, 18 (34%; 95% Wilson CI: 23-47%) exhibit script collapse (SFR < 10%); MMS-1B and SeamlessM4T-v2 maintain SFR above 99% on every language evaluated, confirming that SFR correctly identifies high fidelity where it is present. We identify three distinct collapse patterns: Latin phonetic substitution (smaller Whisper on Indic languages), Arabic substitution for Somali's Latin-script orthography, and Devanagari substitution where larger Whisper models treat all Indic audio as Hindi, a failure present even in Whisper large-v3.


【8】Neural networks for Text-to-Speech evaluation
标题:用于文本到语音评估的神经网络
链接:https://arxiv.org/abs/2604.08562

作者:Ilya Trofimenko,David Kocharyan,Aleksandr Zaitsev,Pavel Repnikov,Mark Levin,Nikita Shevtsov
摘要:确保文本到语音(TTS)系统提供大规模的人类感知质量是现代语音技术的核心挑战。人类主观评价协议,如平均意见评分(MOS)和并排(SBS)比较仍然是事实上的黄金标准,但它们是昂贵的,缓慢的,并对普遍的评估偏见敏感。本研究通过制定和实施一套新颖的神经模型来解决这些障碍,这些模型旨在近似相对(SBS)和绝对(MOS)设置中的专家判断。对于相对评估,我们提出了NeuralSBS,这是一个HuBERT支持的模型,准确率达到73.7%(在SOMOS数据集上)。对于绝对评估,我们引入了使用自定义序列长度优化的MOSNet增强功能,以及WhisperBert,这是一种多模态堆叠集成,通过弱学习器将Whisper音频功能和BERT文本嵌入相结合。我们最好的MOS模型实现了~0.40的均方根误差(RMSE),显著优于0.62的人类评分员间RMSE基线。此外,我们的消融研究表明,通过交叉注意天真地融合文本会降低性能,突出了基于集成的堆叠比直接潜在融合的有效性。我们还报告了基于SpeechLM的架构和zero-shot LLM评估器(Qwen 2-Audio,Gemini 2.5 flash预览)的负面结果,加强了专用度量学习框架的必要性。
摘要:Ensuring that Text-to-Speech (TTS) systems deliver human-perceived quality at scale is a central challenge for modern speech technologies. Human subjective evaluation protocols such as Mean Opinion Score (MOS) and Side-by-Side (SBS) comparisons remain the de facto gold standards, yet they are expensive, slow, and sensitive to pervasive assessor biases. This study addresses these barriers by formulating, and implementing a suite of novel neural models designed to approximate expert judgments in both relative (SBS) and absolute (MOS) settings. For relative assessment, we propose NeuralSBS, a HuBERT-backed model achieving 73.7% accuracy (on SOMOS dataset). For absolute assessment, we introduce enhancements to MOSNet using custom sequence-length batching, as well as WhisperBert, a multimodal stacking ensemble that combines Whisper audio features and BERT textual embeddings via weak learners. Our best MOS models achieve a Root Mean Square Error (RMSE) of ~0.40, significantly outperforming the human inter-rater RMSE baseline of 0.62. Furthermore, our ablation studies reveal that naively fusing text via cross-attention can degrade performance, highlighting the effectiveness of ensemble-based stacking over direct latent fusion. We additionally report negative results with SpeechLM-based architectures and zero-shot LLM evaluators (Qwen2-Audio, Gemini 2.5 flash preview), reinforcing the necessity of dedicated metric learning frameworks.


机器翻译由腾讯交互翻译提供,仅供参考