微信公众号:arXiv_Daily
cs.SD语音
【1】UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
标题:UniSS:与您的声音进行统一的表达性语音翻译
链接:https://arxiv.org/abs/2509.21144
摘要:表达性语音到语音翻译(S2ST)的最终目标是准确地翻译口语内容,同时保留说话者身份和情感风格。然而,这一领域的进展在很大程度上受到三个关键挑战的阻碍:保留表达风格的成对语音数据的稀缺性,多阶段处理管道的复杂性,以及从大型语言模型(LLM)转移翻译能力的有限性。在这项工作中,我们通过引入UniSS,一种新颖的单阶段表达S2ST框架来解决这些挑战。我们的方法具有精心设计的语音语义和风格建模,能够与现有的基于文本的LLM框架无缝集成,以开发统一的文本语音语言模型。为了将翻译能力从文本转移到语音,我们提出了一个跨模态的思想链提示过程,该过程逐步将音频语义与文本对齐,并确保解码结果中的风格保留。此外,我们构建并发布了一个大规模,高质量的表达S2ST数据集,UniST,包括44.8k小时的数据。实验结果表明,UniSS显着优于以前的方法在翻译保真度和语音质量,同时保持语音,情感和持续时间的一致性。我们的工作建立了一个更简单,更有效的模式,建立下一代的表达S2ST系统。音频样本可在https://cmots.github.io/uniss-demo上获得。
摘要:The ultimate goal of expressive speech-to-speech translation (S2ST) is to accurately translate spoken content while preserving the speaker identity and emotional style. However, progress in this field is largely hindered by three key challenges: the scarcity of paired speech data that retains expressive styles, the complexity of multi-stage processing pipelines, and the limited transfer of translation capabilities from large language models (LLMs). In this work, we address these challenges by introducing UniSS, a novel single-stage framework for expressive S2ST. Our approach features carefully designed speech semantic and style modeling, enabling seamless integration with existing text-based LLM frameworks to develop a unified text-speech language model. To transfer translation capabilities from text to speech, we propose a cross-modal chain-of-thought prompting process that progressively aligns audio semantics with text and ensures style preservation in the decoded results. Furthermore, we construct and release a large-scale, high-quality expressive S2ST dataset, UniST, comprising 44.8k hours of data. Experimental results show that UniSS significantly outperforms previous methods in translation fidelity and speech quality while preserving voice, emotion, and duration consistency. Our work establishes a simpler and more effective paradigm for building the next generation of expressive S2ST systems. Audio samples are available at https://cmots.github.io/uniss-demo.
【2】SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization
标题:SupCLAP:利用支持载体正规化控制音频文本对比学习中的优化轨迹漂移
链接:https://arxiv.org/abs/2509.21033
摘要:对比语言音频预训练旨在将多模态表示统一在共享嵌入空间中,是构建从跨模态检索到尖端多模态大型语言模型等广泛应用的基石。然而,我们发现对比学习中负样本推力的垂直分量是一把双刃剑:它包含了来自负样本的丰富补充信息,但其不受约束的性质导致优化轨迹漂移和训练不稳定。为了解决这个问题,我们提出了支持向量正则化(SVR),这是一种引入辅助支持向量来控制这个垂直分量的方法,旨在利用其丰富的信息,同时减轻相关的轨迹漂移。SVR的有效性主要取决于其语义半径,为此,我们探索了两种无监督建模策略:直接参数化和自适应半径预测模块,通过约束来提高其预测精度。大量的实验结果表明,我们的方法超越了广泛使用的基线,如InfoNCE和SigLIP损失跨分类,单语检索,多语言检索标准的音频文本数据集。理论分析和优化轨迹漂移的实验结果验证了该方法的正确性和有效性。
摘要:Contrastive language-audio pretraining, which aims to unify multimodal representations in a shared embedding space, serves as a cornerstone for building a wide range of applications, from cross-modal retrieval to cutting-edge multimodal large language models. However, we find that the perpendicular component of the pushing force from negative samples in contrastive learning is a double-edged sword: it contains rich supplementary information from negative samples, yet its unconstrained nature causes optimization trajectory drift and training instability. To address this, we propose Support Vector Regularization (SVR), a method that introduces an auxiliary support vector to control this perpendicular component, aiming to harness its rich information while mitigating the associated trajectory drift. The efficacy of SVR is critically governed by its semantic radius, for which we explore two unsupervised modeling strategies: direct parameterization and an adaptive radius predictor module enhanced with constraints to improve its predicting accuracy. Extensive experimental results demonstrate that our method surpasses widely used baselines like InfoNCE and SigLIP loss across classification, monolingual retrieval, and multilingual retrieval on standard audio-text datasets. Both the theoretical analysis and the experimental results on optimizing trajectory drift validate the correctness and effectiveness of our SVR method.
【3】i-LAVA: Insights on Low Latency Voice-2-Voice Architecture for Agents
标题:i-LAVA:低延迟语音见解-2-代理语音架构
链接:https://arxiv.org/abs/2509.20971
备注:This paper analyzes a low-latency, end-to-end voice-to-voice (V-2-V) architecture, identifying that the Text-to-Speech (TTS) component has the highest impact on real-time performance. By reducing the number of Residual Vector Quantization (RVQ) iterations in the TTS model, latency can be effectively halved, creating a direct trade-off between conversational speed and audio quality
摘要:我们实验了一个低延迟的端到端语音到语音通信模型,以优化它的实时会话应用程序。通过分析语音到语音(V-2-V)系统的关键组件,即自动语音识别(ASR),文本到语音(TTS)和对话管理,我们的工作分析了如何减少处理时间,同时保持高质量的交互,以确定优化V-2-V系统的杠杆。我们的工作发现,TTS组件生成逼真的语音,充满情感,包括自然停顿和感叹,对实时因素(RTF)影响最大。实验的V-2-V架构利用CSM 1b具有理解音调以及会话的上下文的能力,其通过将先前交换的音频和文本两者结合以生成上下文准确的语音。我们探索了TTS解码器对残差矢量量化(RVQ)迭代的优化,这是以降低生成的语音质量为代价的。我们的实验评估还表明,对于基于CSM的V-2-V实现,最重要的优化可以通过减少RVQ迭代次数以及Mimi中使用的码本来实现。
摘要:We experiment with a low-latency, end-to-end voice-to-voice communication model to optimize it for real-time conversational applications. By analyzing components essential to voice to voice (V-2-V) system viz. automatic speech recognition (ASR), text-to-speech (TTS), and dialog management, our work analyzes how to reduce processing time while maintaining high-quality interactions to identify the levers for optimizing V-2-V system. Our work identifies that TTS component which generates life-like voice, full of emotions including natural pauses and exclamations has highest impact on Real time factor (RTF). The experimented V-2-V architecture utilizes CSM1b has the capability to understand tone as well as context of conversation by ingesting both audio and text of prior exchanges to generate contextually accurate speech. We explored optimization of Residual Vector Quantization (RVQ) iterations by the TTS decoder which come at a cost of decrease in the quality of voice generated. Our experimental evaluations also demonstrate that for V-2-V implementations based on CSM most important optimizations can be brought by reducing the number of RVQ Iterations along with the codebooks used in Mimi.
【4】SingVERSE: A Diverse, Real-World Benchmark for Singing Voice Enhancement
标题:SingVERSE:歌唱声音增强的多元化、现实世界基准
链接:https://arxiv.org/abs/2509.20969
备注:Demopage: this https URL, Dataset: this https URL
摘要:本文提出了一个歌唱声增强的基准。由于缺乏真实的评价数据,限制了歌唱嗓音增强技术的发展。为了解决这一差距,本文介绍了SingVERSE,这是第一个真实世界的歌唱声音增强基准,涵盖了各种声学场景,并提供了成对的,录音室质量的干净参考。利用SingVERSE,我们对最先进的模型进行全面评估,并发现感知质量和可理解性之间的一致权衡。最后,我们表明,在域内歌唱数据的训练大大提高了增强性能,而不会降低语音能力,建立一个简单而有效的路径前进。这项工作为社区提供了一个基本的基准以及关键的见解,以指导这个尚未开发的领域的未来发展。Demopage:https://singverse.github.io
摘要:This paper presents a benchmark for singing voice enhancement. The development of singing voice enhancement is limited by the lack of realistic evaluation data. To address this gap, this paper introduces SingVERSE, the first real-world benchmark for singing voice enhancement, covering diverse acoustic scenarios and providing paired, studio-quality clean references. Leveraging SingVERSE, we conduct a comprehensive evaluation of state-of-the-art models and uncover a consistent trade-off between perceptual quality and intelligibility. Finally, we show that training on in-domain singing data substantially improves enhancement performance without degrading speech capabilities, establishing a simple yet effective path forward. This work offers the community a foundational benchmark together with critical insights to guide future advances in this underexplored domain. Demopage: https://singverse.github.io
【5】AIBA: Attention-based Instrument Band Alignment for Text-to-Audio Diffusion
标题:AIBA:基于注意力的文本到音频扩散的乐器频带对齐
链接:https://arxiv.org/abs/2509.20891
备注:NeurIPS 2025 AI for Music Workshop
摘要:我们提出了AIBA(带内注意力对齐),这是一个轻量级的,无需训练的管道,用于量化文本到音频扩散模型在时频(T-F)平面上的位置。AIBA(i)在推理时钩住交叉注意力,以记录注意力概率,而不修改权重;(ii)将它们投射到与音频能量直接可比的固定大小的梅尔网格;(iii)通过可解释的度量(T-F IoU/AP,频率配置文件相关性和指向游戏)与仪器波段地面实况达成一致。在具有AudioLDM 2主干的Slakh 2100上,AIBA揭示了一致的仪器依赖性趋势(例如,低音有利于低波段),并实现高精度与适度的召回。
摘要:We present AIBA (Attention-In-Band Alignment), a lightweight, training-free pipeline to quantify where text-to-audio diffusion models attend on the time-frequency (T-F) plane. AIBA (i) hooks cross-attention at inference to record attention probabilities without modifying weights; (ii) projects them to fixed-size mel grids that are directly comparable to audio energy; and (iii) scores agreement with instrument-band ground truth via interpretable metrics (T-F IoU/AP, frequency-profile correlation, and a pointing game). On Slakh2100 with an AudioLDM2 backbone, AIBA reveals consistent instrument-dependent trends (e.g., bass favoring low bands) and achieves high precision with moderate recall.
【6】AuthGlass: Enhancing Voice Authentication on Smart Glasses via Air-Bone Acoustic Features
标题:AuthGlass:通过Air-Bone声学功能增强智能眼镜上的语音认证
链接:https://arxiv.org/abs/2509.20799
备注:24 pages, 12 figures, submitted to CHI'26
摘要:随着智能眼镜的快速发展,语音交互由于其自然和方便而得到广泛部署。然而,它的实用性往往受到欺骗攻击和周围声音干扰的脆弱性的破坏,使得无缝语音认证对于智能眼镜的使用至关重要。为了应对这一挑战,我们提出了AuthGlass,一种利用空气和骨传导语音特征来提高准确性和活性检测的语音认证方法。为了获得与语音相关的声学和振动特征的全面知识,我们构建了一个具有冗余同步麦克风的智能眼镜原型:14个空气传导麦克风和2个骨传导单元。在一项有42名参与者的研究中,我们验证了将声场和振动特征相结合可以显着提高身份验证的鲁棒性和抗攻击性。此外,实验表明,AuthGlass即使在各种实际场景下也保持了具有竞争力的准确性,突出了其在实际部署中的适用性和可扩展性。
摘要:With the rapid advancement of smart glasses, voice interaction has become widely deployed due to its naturalness and convenience. However, its practicality is often undermined by the vulnerability to spoofing attacks and interference from surrounding sounds, making seamless voice authentication crucial for smart glasses usage. To address this challenge, we propose AuthGlass, a voice authentication approach that leverages both air- and bone-conducted speech features to enhance accuracy and liveness detection. Aiming to gain comprehensive knowledge on speech-related acoustic and vibration features, we built a smart glasses prototype with redundant synchronized microphones: 14 air-conductive microphones and 2 bone-conductive units. In a study with 42 participants, we validated that combining sound-field and vibration features significantly improves authentication robustness and attack resistance. Furthermore, experiments demonstrated that AuthGlass maintains competitive accuracy even under various practical scenarios, highlighting its applicability and scalability for real-world deployment.
【7】MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model
标题:MI-SYS:利用闭源大音频语言模型进行无监督领域自适应的标签融合
链接:https://arxiv.org/abs/2509.20706
备注:5 pages, 2 figures, 2 tables
摘要:大型音频语言模型(LALM)在语音任务上表现出很强的zero-shot能力,为语音情感识别(SER)提供了希望。然而,SER在实际部署中经常在域不匹配的情况下失败,其中源数据不可用,并且只能通过API访问强大的LALM。我们要问:如果只给出未标记的目标域音频和仅API的LALM,学生模型是否可以在目标域中超越LALM?为此,我们提出了MI的,去噪标签融合框架,补充了LALM与源域训练的SER分类器作为辅助教师。该框架从两个教师中提取多个随机预测,通过基于互信息的不确定性来加权其平均分布,并使用指数移动平均教师来稳定训练。在三个公共情绪数据集和六个跨域传输的实验显示出一致的收益,学生超过了LALM,并超过了最强的基线3.9%。这种方法在不共享源数据的情况下增强了情感感知语音系统,从而实现了现实的适应。
摘要:Large audio-language models (LALMs) show strong zero-shot ability on speech tasks, suggesting promise for speech emotion recognition (SER). However, SER in real-world deployments often fails under domain mismatch, where source data are unavailable and powerful LALMs are accessible only through an API. We ask: given only unlabeled target-domain audio and an API-only LALM, can a student model be adapted to outperform the LALM in the target domain? To this end, we propose MI-Fuse, a denoised label fusion framework that supplements the LALM with a source-domain trained SER classifier as an auxiliary teacher. The framework draws multiple stochastic predictions from both teachers, weights their mean distributions by mutual-information-based uncertainty, and stabilizes training with an exponential moving average teacher. Experiments across three public emotion datasets and six cross-domain transfers show consistent gains, with the student surpassing the LALM and outperforming the strongest baseline by 3.9%. This approach strengthens emotion-aware speech systems without sharing source data, enabling realistic adaptation.
【8】Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection
标题:解决数据增强训练中的梯度失调以实现稳健的语音深度伪造检测
链接:https://arxiv.org/abs/2509.20682
备注:5 pages, 4 figures
摘要:在语音深度伪造检测(SDD)中,数据增强(DA)通常用于改善不同语音条件和欺骗攻击的模型泛化。然而,在训练期间,来自原始输入和增强输入的反向传播梯度可能不对齐,这可能导致冲突的参数更新。这些冲突可能会阻碍收敛,并将模型推向次优解决方案,从而减少DA的好处。为了研究和解决这个问题,我们设计了一个双路径数据增强(DPDA)训练框架,并为SDD进行梯度对齐。在我们的框架中,每个训练话语通过两个输入路径进行处理:一个使用原始语音,另一个使用其增强版本。这种设计允许我们比较和对齐它们的反向传播梯度方向,以减少优化冲突。我们的分析表明,当使用RawBoost增强时,大约25%的训练迭代在原始输入和它们的增强对应物之间表现出梯度冲突。通过使用梯度对齐解决这些冲突,我们的方法通过减少训练时期的数量来加速收敛,并且与基线相比,在野外数据集上实现了高达18.69%的等错误率相对降低。
摘要:In speech deepfake detection (SDD), data augmentation (DA) is commonly used to improve model generalization across varied speech conditions and spoofing attacks. However, during training, the backpropagated gradients from original and augmented inputs may misalign, which can result in conflicting parameter updates. These conflicts could hinder convergence and push the model toward suboptimal solutions, thereby reducing the benefits of DA. To investigate and address this issue, we design a dual-path data-augmented (DPDA) training framework with gradient alignment for SDD. In our framework, each training utterance is processed through two input paths: one using the original speech and the other with its augmented version. This design allows us to compare and align their backpropagated gradient directions to reduce optimization conflicts. Our analysis shows that approximately 25% of training iterations exhibit gradient conflicts between the original inputs and their augmented counterparts when using RawBoost augmentation. By resolving these conflicts with gradient alignment, our method accelerates convergence by reducing the number of training epochs and achieves up to an 18.69% relative reduction in Equal Error Rate on the In-the-Wild dataset compared to the baseline.
【9】QAMO: Quality-aware Multi-centroid One-class Learning For Speech Deepfake Detection
标题:QAMO:语音深度伪造检测的质量感知多中心一级学习
链接:https://arxiv.org/abs/2509.20679
备注:5 pages, 4 figures
摘要:最近的工作表明,单类学习可以通过对围绕单个质心的真实语音的紧凑分布进行建模来检测看不见的deepfake攻击。然而,单质心假设可能会过度简化真实的语音表示,并忽略有用的线索,如语音质量,这反映了语音的自然性。语音质量可以很容易地获得使用现有的语音质量评估模型,估计它通过平均意见得分。在本文中,我们提出了QAMO:用于语音深度伪造检测的质量感知多中心一类学习。QAMO通过引入多个质量感知质心扩展了传统的单类学习。在QAMO中,每个质心被优化以表示不同的语音质量子空间,从而能够更好地建模真实语音中的类内变化。此外,QAMO还支持多质心集成评分策略,该策略可提高决策阈值并减少推理过程中对质量标签的需求。使用两个质心来表示高质量和低质量的语音,我们提出的QAMO在野外数据集中实现了5.09%的相等错误率,优于以前的一类和质量感知系统。
摘要:Recent work shows that one-class learning can detect unseen deepfake attacks by modeling a compact distribution of bona fide speech around a single centroid. However, the single-centroid assumption can oversimplify the bona fide speech representation and overlook useful cues, such as speech quality, which reflects the naturalness of the speech. Speech quality can be easily obtained using existing speech quality assessment models that estimate it through Mean Opinion Score. In this paper, we propose QAMO: Quality-Aware Multi-Centroid One-Class Learning for speech deepfake detection. QAMO extends conventional one-class learning by introducing multiple quality-aware centroids. In QAMO, each centroid is optimized to represent a distinct speech quality subspaces, enabling better modeling of intra-class variability in bona fide speech. In addition, QAMO supports a multi-centroid ensemble scoring strategy, which improves decision thresholding and reduces the need for quality labels during inference. With two centroids to represent high- and low-quality speech, our proposed QAMO achieves an equal error rate of 5.09% in In-the-Wild dataset, outperforming previous one-class and quality-aware systems.
【10】Building Tailored Speech Recognizers for Japanese Speaking Assessment
标题:构建面向日语口语评估的定制语音识别器
链接:https://arxiv.org/abs/2509.20655
摘要:本文介绍了构建适合日语口语评估任务的语音识别器的方法。具体来说,我们建立了一个语音识别器,输出带有重音标记的音素标签。虽然日语资源丰富,但只有少量数据用于训练模型以产生包括重音标记的准确音素transmits。我们提出了两种方法来减轻数据稀疏。首先,多任务训练方案引入辅助损失函数来估计输入信号的正字法文本标签和音调模式,以便在训练中可以利用仅具有正字法注释的话语。第二个融合两个估计,一个在语音字母串,和其他文本令牌序列。为了结合这些估计,我们开发了一种基于有限状态换能器框架的算法。我们的研究结果表明,使用多任务学习和融合是有效的建立一个准确的音素识别器。我们表明,这种方法是有利的相比,使用通用的多语言识别器。比较了各种方法的相对优势。我们提出的方法降低了平均mora-label错误率从12.3%到7.1%的CSJ核心评估集。
摘要:This paper presents methods for building speech recognizers tailored for Japanese speaking assessment tasks. Specifically, we build a speech recognizer that outputs phonemic labels with accent markers. Although Japanese is resource-rich, there is only a small amount of data for training models to produce accurate phonemic transcriptions that include accent marks. We propose two methods to mitigate data sparsity. First, a multitask training scheme introduces auxiliary loss functions to estimate orthographic text labels and pitch patterns of the input signal, so that utterances with only orthographic annotations can be leveraged in training. The second fuses two estimators, one over phonetic alphabet strings, and the other over text token sequences. To combine these estimates we develop an algorithm based on the finite-state transducer framework. Our results indicate that the use of multitask learning and fusion is effective for building an accurate phonemic recognizer. We show that this approach is advantageous compared to the use of generic multilingual recognizers. The relative advantages of the proposed methods were also compared. Our proposed methods reduced the average of mora-label error rates from 12.3% to 7.1% over the CSJ core evaluation sets.
【11】Investigating Modality Contribution in Audio LLMs for Music
标题:调查音乐音频LLM中的情态贡献
链接:https://arxiv.org/abs/2509.20641
摘要:音频大语言模型(Audio LLM)可以实现关于音乐的类似人类的对话,但目前还不清楚它们是真的在听音频还是只是使用文本推理,正如最近的基准测试所表明的那样。本文通过量化每种模态对模型输出的贡献来研究这个问题。我们采用MM-SHAP框架,这是一个基于Shapley值的性能不可知分数,它量化了每个模态对模型预测的相对贡献。我们在MuChoMusic基准测试中评估了两个模型,发现准确率更高的模型更多地依赖于文本来回答问题,但进一步的检查表明,即使整体音频贡献很低,模型也可以成功地定位关键的声音事件,这表明音频并没有完全被忽略。我们的研究是MM-SHAP在音频LLM中的首次应用,我们希望它能成为未来可解释AI和音频研究的基础。
摘要:Audio Large Language Models (Audio LLMs) enable human-like conversation about music, yet it is unclear if they are truly listening to the audio or just using textual reasoning, as recent benchmarks suggest. This paper investigates this issue by quantifying the contribution of each modality to a model's output. We adapt the MM-SHAP framework, a performance-agnostic score based on Shapley values that quantifies the relative contribution of each modality to a model's prediction. We evaluate two models on the MuChoMusic benchmark and find that the model with higher accuracy relies more on text to answer questions, but further inspection shows that even if the overall audio contribution is low, models can successfully localize key sound events, suggesting that audio is not entirely ignored. Our study is the first application of MM-SHAP to Audio LLMs and we hope it will serve as a foundational step for future research in explainable AI and audio.
【12】Why Speech Deepfake Detectors Won't Generalize: The Limits of Detection in an Open World
标题:为什么语音Deepfake检测器不会泛化:开放世界中检测的极限
链接:https://arxiv.org/abs/2509.20405
摘要:语音Deepfake检测器通常在干净的基准测试条件下进行评估,但部署发生在移动设备、采样率、编解码器、环境和攻击家族的开放世界中。这为基于人工智能的检测器创造了“覆盖债务”:每一个新的条件都会与现有的条件相乘,从而产生数据盲点,其增长速度超过了数据收集的速度。由于攻击者可以针对这些未覆盖的区域,因此最坏情况下的性能(而不是平均基准得分)决定了安全性。为了证明覆盖债务问题的影响,我们分析了最近的交叉测试框架的结果。在真正的领域和欺骗发布年的性能方面,出现了两种模式:新的合成器消除了传统的工件检测器所依赖的,而会话语音领域(电话会议,采访,社交媒体)一直是最难保护的。这些发现表明,在做出高风险决策时,不应仅仅依靠检测。检测器应该被视为分层防御中的辅助信号,包括出处、人格凭证和策略保障。
摘要:Speech deepfake detectors are often evaluated on clean, benchmark-style conditions, but deployment occurs in an open world of shifting devices, sampling rates, codecs, environments, and attack families. This creates a ``coverage debt" for AI-based detectors: every new condition multiplies with existing ones, producing data blind spots that grow faster than data can be collected. Because attackers can target these uncovered regions, worst-case performance (not average benchmark scores) determines security. To demonstrate the impact of the coverage debt problem, we analyze results from a recent cross-testing framework. Grouping performance by bona fide domain and spoof release year, two patterns emerge: newer synthesizers erase the legacy artifacts detectors rely on, and conversational speech domains (teleconferencing, interviews, social media) are consistently the hardest to secure. These findings show that detection alone should not be relied upon for high-stakes decisions. Detectors should be treated as auxiliary signals within layered defenses that include provenance, personhood credentials, and policy safeguards.
【13】Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?
标题:现代语音增强系统容易受到对抗攻击吗?
链接:https://arxiv.org/abs/2509.21087
备注:Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:用于语音增强的机器学习方法正变得越来越有表现力,能够对输入信号进行更强大的修改。在本文中,我们证明了这种表现力引入了一个漏洞:先进的语音增强模型可能容易受到对抗性攻击。具体来说,我们表明,敌对的噪音,精心制作和心理声学掩盖的原始输入,可以注入这样的增强语音输出传达一个完全不同的语义含义。我们通过实验验证,当代预测语音增强模型确实可以以这种方式进行操作。此外,我们强调,具有随机采样器的扩散模型通过设计表现出对这种对抗性攻击的固有鲁棒性。
摘要:Machine learning approaches for speech enhancement are becoming increasingly expressive, enabling ever more powerful modifications of input signals. In this paper, we demonstrate that this expressiveness introduces a vulnerability: advanced speech enhancement models can be susceptible to adversarial attacks. Specifically, we show that adversarial noise, carefully crafted and psychoacoustically masked by the original input, can be injected such that the enhanced speech output conveys an entirely different semantic meaning. We experimentally verify that contemporary predictive speech enhancement models can indeed be manipulated in this way. Furthermore, we highlight that diffusion models with stochastic samplers exhibit inherent robustness to such adversarial attacks by design.
【14】SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
标题:SPADE:结构化修剪和自适应蒸馏,实现高效LLM-TTC
链接:https://arxiv.org/abs/2509.20802
备注:Submitted to ICASSP 2026
摘要:本文的目标是介绍SPADE,一个用于高效的基于大语言模型的文本到语音(LLM-TTS)的结构化修剪和自适应蒸馏的框架。近年来的LLM-TTS系统实现了较强的可控性和zero-shot泛化能力,但其参数数量多、延迟高,限制了实际应用。SPADE通过结合(i)基于单词错误率的层重要性指数指导的修剪步骤来去除非必要的Transformer层,与(ii)多级知识蒸馏来恢复自回归一致性来解决这个问题。在zero-shot基准测试中,SPADE保留了接近奇偶校验的感知质量,同时将Transformer深度减半,将VRAM使用量减少了20%,并在原始训练数据不到5%的情况下实现了高达1.7倍的实时因子。这些结果表明,紧凑的LLM-TTS模型可以保持自然度和说话人相似性,同时使实际的实时语音生成。音频样本可在https://mm.kaist.ac.kr/projects/SPADE/上获得。
摘要:The goal of this paper is to introduce SPADE, a framework for Structured Pruning and Adaptive Distillation for Efficient Large Language Model-based text-to-speech (LLM-TTS). Recent LLM-TTS systems achieve strong controllability and zero-shot generalization, but their large parameter counts and high latency limit real-world deployment. SPADE addresses this by combining (i) a pruning step guided by a word-error-rate-based layer importance index to remove non-essential Transformer layers, with (ii) multi-level knowledge distillation to restore autoregressive coherence. On zero-shot benchmarks, SPADE preserves near-parity perceptual quality while halving Transformer depth, reducing VRAM usage by up to 20%, and achieving up to 1.7x faster real-time factor with less than 5% of the original training data. These results show that compact LLM-TTS models can maintain naturalness and speaker similarity while enabling practical real-time speech generation. Audio samples are available at https://mm.kaist.ac.kr/projects/SPADE/.
【15】Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
标题:通过离散令牌的条件预测客观评估语音合成中的韵律和可理解度
链接:https://arxiv.org/abs/2509.20485
备注:Under review for IEEE OJSP
摘要:合成语音的客观评估对于推进语音生成系统至关重要,但现有的可懂度和韵律指标范围仍然有限,并且与人类感知的相关性较弱。词错误率(WER)只提供了一个粗略的基于文本的可理解性的测量,而F0-RMSE和相关的基于音高的度量提供了一个狭窄的,依赖于参考的韵律视图。为了解决这些限制,我们提出了TTScore,一个有针对性的和无参考的评估框架的基础上的离散语音标记的条件预测。TTScore采用两个以输入文本为条件的序列到序列预测器:TTScore-int,它通过内容令牌测量可理解性,TTScore-pro,它通过韵律令牌评估韵律。对于每个合成的话语,预测器计算相应的令牌序列的可能性,产生可解释的分数,捕获与预期的语言内容和韵律结构的对齐。SOMOS,VoiceMOS和TTSArena基准测试的实验表明,TTScore-int和TTScore-pro提供了可靠的,特定方面的评估,并实现了更强的相关性与人类判断的整体质量比现有的可理解性和韵律为重点的指标。
摘要:Objective evaluation of synthesized speech is critical for advancing speech generation systems, yet existing metrics for intelligibility and prosody remain limited in scope and weakly correlated with human perception. Word Error Rate (WER) provides only a coarse text-based measure of intelligibility, while F0-RMSE and related pitch-based metrics offer a narrow, reference-dependent view of prosody. To address these limitations, we propose TTScore, a targeted and reference-free evaluation framework based on conditional prediction of discrete speech tokens. TTScore employs two sequence-to-sequence predictors conditioned on input text: TTScore-int, which measures intelligibility through content tokens, and TTScore-pro, which evaluates prosody through prosody tokens. For each synthesized utterance, the predictors compute the likelihood of the corresponding token sequences, yielding interpretable scores that capture alignment with intended linguistic content and prosodic structure. Experiments on the SOMOS, VoiceMOS, and TTSArena benchmarks demonstrate that TTScore-int and TTScore-pro provide reliable, aspect-specific evaluation and achieve stronger correlations with human judgments of overall quality than existing intelligibility and prosody-focused metrics.
【16】Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
标题:Phoenix-VAD:用于全双工语音交互的流语义端点检测
链接:https://arxiv.org/abs/2509.20410
摘要:口语对话模型具有显着先进的智能人机交互,但它们缺乏用于语义端点检测的即插即用的“textendash”和“textendash play full”textendash双工预测模块,阻碍了无缝音频交互。在本文中,我们介绍凤凰\textendashVAD,LLM\textendash为基础的模型,使流语义端点检测。具体来说,Phoenix\textendash VAD利用LLM的语义理解能力和滑动窗口训练策略来实现可靠的语义端点检测,同时支持流推理。在语义完全和不完全语音场景下的实验表明,Phoenix\textendash VAD具有优异的性能。此外,该设计使全文本双工预测模块能够独立于对话模型进行优化,为下一代文本人机交互提供更可靠、更灵活的支持。
摘要:Spoken dialogue models have significantly advanced intelligent human\textendash computer interaction, yet they lack a plug\textendash and\textendash play full\textendash duplex prediction module for semantic endpoint detection, hindering seamless audio interactions. In this paper, we introduce Phoenix\textendashVAD, an LLM\textendash based model that enables streaming semantic endpoint detection. Specifically, Phoenix\textendash VAD leverages the semantic comprehension capability of the LLM and a sliding window training strategy to achieve reliable semantic endpoint detection while supporting streaming inference. Experiments on both semantically complete and incomplete speech scenarios indicate that Phoenix\textendash VAD achieves excellent and competitive performance. Furthermore, this design enables the full\textendash duplex prediction module to be optimized independently of the dialogue model, providing more reliable and flexible support for next\textendash generation human\textendash computer interaction.
【17】Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
标题:基于不确定性音素难度分数的非规范语音ASR个性化
链接:https://arxiv.org/abs/2509.20396
摘要:自动语音识别(ASR)系统与来自患有脑瘫或结构异常等疾病的个人的非规范性语音作斗争。高声学可变性和训练数据的稀缺性严重降低了模型的性能。这项工作介绍了一种数据高效的个性化方法,量化音素级的不确定性,以指导微调。我们利用Monte Carlo Dropout来估计模型发现最困难的音素,并将这些估计用于有针对性的过采样策略。我们验证我们的方法在英语和德语数据集。至关重要的是,我们证明了我们的模型推导的不确定性与专家临床logopedic报告中确定为具有挑战性的音素密切相关,据我们所知,这是第一次成功地将模型不确定性与专家评估的语音难度相结合。我们的研究结果表明,这种经过临床验证的、不确定性引导的抽样显著提高了ASR的准确性,为个性化和包容性的ASR提供了一个实用的框架。
摘要:Automatic speech recognition (ASR) systems struggle with non-normative speech from individuals with impairments caused by conditions like cerebral palsy or structural anomalies. The high acoustic variability and scarcity of training data severely degrade model performance. This work introduces a data-efficient personalization method that quantifies phoneme-level uncertainty to guide fine-tuning. We leverage Monte Carlo Dropout to estimate which phonemes a model finds most difficult and use these estimates for a targeted oversampling strategy. We validate our method on English and German datasets. Crucially, we demonstrate that our model-derived uncertainty strongly correlates with phonemes identified as challenging in an expert clinical logopedic report, marking, to our knowledge, the first work to successfully align model uncertainty with expert assessment of speech difficulty. Our results show that this clinically-validated, uncertainty-guided sampling significantly improves ASR accuracy, delivering a practical framework for personalized and inclusive ASR.
【1】MeanSE: Efficient Generative Speech Enhancement with Mean Flows
标题:MeanSE:使用Mean Flow的高效生成语音增强
链接:https://arxiv.org/abs/2509.21214
备注:Submitted to ICASSP 2026
摘要:语音增强(SE)提高了退化语音的质量,其中生成模型(如流匹配)因其出色的感知质量而受到关注。然而,基于流的模型需要多个数量的函数评估(NFE),以实现稳定和令人满意的性能,导致高计算负载和差的1-NFE性能。在本文中,我们提出了MeanSE,一个有效的生成语音增强模型,使用平均流,模型的平均速度场,以实现高质量的1-NFE增强。实验结果表明,我们提出的MeanSE显着优于流匹配基线与一个单一的NFE,表现出非常好的域外泛化能力。
摘要:Speech enhancement (SE) improves degraded speech's quality, with generative models like flow matching gaining attention for their outstanding perceptual quality. However, the flow-based model requires multiple numbers of function evaluations (NFEs) to achieve stable and satisfactory performance, leading to high computational load and poor 1-NFE performance. In this paper, we propose MeanSE, an efficient generative speech enhancement model using mean flows, which models the average velocity field to achieve high-quality 1-NFE enhancement. Experimental results demonstrate that our proposed MeanSE significantly outperforms the flow matching baseline with a single NFE, exhibiting extremely better out-of-domain generalization capabilities.
【2】Hybrid Real- And Complex-Valued Neural Network Concept For Low-Complexity Phase-Aware Speech Enhancement
标题:用于低复杂度相感知语音增强的混合实值和复值神经网络概念
链接:https://arxiv.org/abs/2509.21185
摘要:在本文中,我们提出了混合实值和复值神经网络的语音增强。实值或复值模型要么效率低,要么复杂度高。我们设计了一个简单的设计方法,将实值网络扩展为混合网络。基于语音可懂度和质量指标,我们比较了卷积和卷积递归架构的真实,复杂和混合版本。混合网络始终优于具有相同数量参数的同行。此外,混合模型在乘-累加运算方面的复杂度大大低于其对应模型。
摘要:In this paper, we propose hybrid real- and complex-valued neural networks for speech enhancement. Real- or complex-valued models are either inefficient or present high complexity. We devise a straightforward design method for extending a real-valued network into its hybrid counterpart. Based on speech intelligibility and quality metrics, we compare the real, complex, and hybrid versions of a convolutional and a convolutional-recurrent architecture. The hybrid network consistently outperforms its counterparts with the same number of parameters. Additionally, the hybrid models' complexity in terms of multiply-accumulate operations is substantially lower than that of their counterparts.
【3】Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?
标题:现代语音增强系统容易受到对抗攻击吗?
链接:https://arxiv.org/abs/2509.21087
备注:Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:用于语音增强的机器学习方法正变得越来越有表现力,能够对输入信号进行更强大的修改。在本文中,我们证明了这种表现力引入了一个漏洞:先进的语音增强模型可能容易受到对抗性攻击。具体来说,我们表明,敌对的噪音,精心制作和心理声学掩盖的原始输入,可以注入这样的增强语音输出传达一个完全不同的语义含义。我们通过实验验证,当代预测语音增强模型确实可以以这种方式进行操作。此外,我们强调,具有随机采样器的扩散模型通过设计表现出对这种对抗性攻击的固有鲁棒性。
【4】Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
标题:衡量音频对正确性的影响:大型音频语言模型的音频贡献感知训练后
链接:https://arxiv.org/abs/2509.21060
摘要:大型音频语言模型(LALM)代表了多模态AI的一个重要前沿,可以解决各种音频任务。最近,后训练的LALM已受到越来越多的关注,由于显着的性能改进的基础模型。虽然单阶段后训练(如强化学习(RL))已经证明了有希望的结果,但多阶段方法(如监督微调(SFT))以及随后的RL仍然是次优的。跨多个训练阶段分配数据以最大化LALM能力的问题尚未得到充分探索,也缺乏用于此类研究的大规模、高质量数据集。为了解决这些问题,我们首先提出了AudioMCQ,一个全面的音频多项选择题数据集,包括571 k个样本,具有两种思维链注释。其次,我们研究了LALM中普遍存在的零音频贡献现象,其中模型仅从文本信息中获得正确答案,而不处理音频内容。我们提出了音频贡献过滤来将数据划分为弱音频贡献子集和强音频贡献子集。基于这些见解,我们开发了两种有效的后训练范式:弱到强(弱音频贡献数据的SFT,然后是强音频贡献数据的RL)和混合到强(混合音频贡献数据的SFT,然后是强音频贡献数据的RL)。我们通过使用AudioMCQ在DCASE 2025音频提问挑战赛中获得第一名。此外,利用我们的数据集和不同的训练策略,我们在MMAU-test-mini上实现了78.2%,在MMAU上实现了75.6%,在MMAR上实现了67.1%,在MMSU上实现了70.7%,在这些基准测试中建立了新的最先进的性能。
【5】TF-Restormer: Complex Spectral Prediction for Speech Restoration
标题:TF-Restormer:语音恢复的复谱预测
链接:https://arxiv.org/abs/2509.21003
备注:Preprint. Under review
摘要:由于诸如削波、带通滤波、数字伪影、噪声和混响等复合失真以及低采样率,真实世界条件下的语音恢复具有挑战性。现有的系统,包括基于声码器的方法,往往牺牲信号保真度,而扩散模型仍然不切实际的流。此外,大多数假设一个固定的目标采样率,需要外部复位,导致冗余计算。我们提出了TF-Restormer,一个编码器-解码器架构,集中分析输入带宽与时间-频率双路径编码器和重建丢失的高频带通过光解码器与频率扩展查询。它可以在任意输入输出速率下进行有效和通用的恢复,而无需冗余恢复。为了支持不同速率的对抗训练,我们引入了一个共享的采样频率无关(SFI)STFT。TF-Restormer还支持带有因果时间模块的流,并通过将频谱感应偏置注入频率模块来提高极端退化情况下的鲁棒性。最后,我们提出了一个缩放的对数谱损失,稳定优化在恶劣的条件下,同时强调良好的预测光谱的细节。作为跨采样率的单一模型,TF-Restormer始终优于之前的系统,在信号保真度和感知质量方面实现了平衡增益,同时其流模式在实时应用中保持了竞争力。代码和演示可在https://tf-restormer.github.io/demo上获得。
【6】PAS-SE: Personalized Auxiliary-Sensor Speech Enhancement for Voice Pickup in Hearables
标题:PAS-SE:用于可听语音拾取的个性化辅助传感器语音增强
链接:https://arxiv.org/abs/2509.20875
备注:Submitted to ICASSP 2026
摘要:语音增强是指通过抑制噪声和干扰说话者来改善用户的语音,同时保持自身的语音质量。对于单通道方法,在没有附加上下文的情况下将目标与干扰讲话者区分开是特别具有挑战性的。在本文中,我们比较了两种策略来解决这种模糊性:个性化语音增强(PSE),它使用注册话语来代表目标,和内置传感器语音增强(AS-SE),它使用入耳式麦克风作为额外的输入。我们在两个公共数据集上评估了这些策略,采用不同的辅助传感器阵列,以研究它们的跨数据集泛化。我们提出了训练时间增强,以促进AS-SE系统的跨数据集泛化。我们还表明,PSE和AS-SE(PAS-SE)相结合,提供了互补的性能优势,特别是当注册语音与入耳式麦克风记录。我们进一步证明,PAS-SE个性化与嘈杂的入耳式注册保持性能优势的AS-SE系统。
【7】SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
标题:SPADE:结构化修剪和自适应蒸馏,实现高效LLM-TTC
链接:https://arxiv.org/abs/2509.20802
备注:Submitted to ICASSP 2026
摘要:本文的目标是介绍SPADE,一个用于高效的基于大语言模型的文本到语音(LLM-TTS)的结构化修剪和自适应蒸馏的框架。近年来的LLM-TTS系统实现了较强的可控性和zero-shot泛化能力,但其参数数量多、延迟高,限制了实际应用。SPADE通过结合(i)基于单词错误率的层重要性指数指导的修剪步骤来去除非必要的Transformer层,与(ii)多级知识蒸馏来恢复自回归一致性来解决这个问题。在zero-shot基准测试中,SPADE保留了接近奇偶校验的感知质量,同时将Transformer深度减半,将VRAM使用量减少了20%,并在原始训练数据不到5%的情况下实现了高达1.7倍的实时因子。这些结果表明,紧凑的LLM-TTS模型可以保持自然度和说话人相似性,同时使实际的实时语音生成。音频样本可在https://mm.kaist.ac.kr/projects/SPADE/上获得。
【8】Real-Time System for Audio-Visual Target Speech Enhancement
标题:视听目标语音增强实时系统
链接:https://arxiv.org/abs/2509.20741
备注:Accepted into WASPAA 2025 demo session
摘要:我们提出了一个现场演示的RAVEN,实时视听语音增强系统设计完全运行在一个CPU上。在单通道、纯音频环境中,语音增强传统上被视为从环境噪声中提取干净语音的任务。最近的工作探索了使用视觉线索,如嘴唇运动,以提高鲁棒性,特别是在干扰扬声器的存在。然而,据我们所知,没有先前的工作已经证明了一个交互式系统的实时视听语音增强CPU硬件上运行。RAVEN通过使用来自视听语音识别模型的预训练视觉嵌入来编码嘴唇运动信息,填补了这一空白。该系统可以概括环境噪声、干扰扬声器、瞬态声音甚至歌声。在此演示中,与会者将能够使用麦克风和网络摄像头设置体验现场视听目标语音增强,并通过耳机播放清晰的语音。
【9】Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
标题:通过离散令牌的条件预测客观评估语音合成中的韵律和可理解度
链接:https://arxiv.org/abs/2509.20485
备注:Under review for IEEE OJSP
摘要:合成语音的客观评估对于推进语音生成系统至关重要,但现有的可懂度和韵律指标范围仍然有限,并且与人类感知的相关性较弱。词错误率(WER)只提供了一个粗略的基于文本的可理解性的测量,而F0-RMSE和相关的基于音高的度量提供了一个狭窄的,依赖于参考的韵律视图。为了解决这些限制,我们提出了TTScore,一个有针对性的和无参考的评估框架的基础上的离散语音标记的条件预测。TTScore采用两个以输入文本为条件的序列到序列预测器:TTScore-int,它通过内容令牌测量可理解性,TTScore-pro,它通过韵律令牌评估韵律。对于每个合成的话语,预测器计算相应的令牌序列的可能性,产生可解释的分数,捕获与预期的语言内容和韵律结构的对齐。SOMOS,VoiceMOS和TTSArena基准测试的实验表明,TTScore-int和TTScore-pro提供了可靠的,特定方面的评估,并实现了更强的相关性与人类判断的整体质量比现有的可理解性和韵律为重点的指标。
【10】Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
标题:Phoenix-VAD:用于全双工语音交互的流语义端点检测
链接:https://arxiv.org/abs/2509.20410
摘要:口语对话模型具有显着先进的智能人机交互,但它们缺乏用于语义端点检测的即插即用的“textendash”和“textendash play full”textendash双工预测模块,阻碍了无缝音频交互。在本文中,我们介绍凤凰\textendashVAD,LLM\textendash为基础的模型,使流语义端点检测。具体来说,Phoenix\textendash VAD利用LLM的语义理解能力和滑动窗口训练策略来实现可靠的语义端点检测,同时支持流推理。在语义完全和不完全语音场景下的实验表明,Phoenix\textendash VAD具有优异的性能。此外,该设计使全文本双工预测模块能够独立于对话模型进行优化,为下一代文本人机交互提供更可靠、更灵活的支持。
【11】Variational Low-Rank Adaptation for Personalized Impaired Speech Recognition
标题:基于变分低秩自适应的个性化受损语音识别
链接:https://arxiv.org/abs/2509.20397
摘要:由先天性疾病(例如脑瘫、唐氏综合征或Apert综合征)以及由中风、创伤性事故或肿瘤引起的获得性脑损伤引起的言语损伤对自动语音识别(ASR)系统提出了主要挑战。尽管最近取得了进展,但由于训练数据有限和声学可变性高,像Whisper这样的最先进的ASR模型仍然难以处理非规范性语音。此外,收集和注释非规范性言语是繁重的:对于许多受影响的个体来说,说话是费力的,而费力的注释通常需要熟悉说话者的护理人员。本文介绍了一种基于贝叶斯低秩自适应的ASR个性化方法,用于数据高效的微调。我们验证我们的方法上的英语UA语音数据集和新收集的德语语音数据集,BF-Sprache,从一个孩子的结构性语音障碍。该数据集和方法旨在反映低资源环境的挑战,包括有语言障碍的个人。我们的方法显着提高了受损语音的ASR准确性,同时保持数据和注释效率,为实现包容性ASR提供了一条实用的途径。
【12】Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
标题:基于不确定性音素难度分数的非规范语音ASR个性化
链接:https://arxiv.org/abs/2509.20396
摘要:自动语音识别(ASR)系统与来自患有脑瘫或结构异常等疾病的个人的非规范性语音作斗争。高声学可变性和训练数据的稀缺性严重降低了模型的性能。这项工作介绍了一种数据高效的个性化方法,量化音素级的不确定性,以指导微调。我们利用Monte Carlo Dropout来估计模型发现最困难的音素,并将这些估计用于有针对性的过采样策略。我们验证我们的方法在英语和德语数据集。至关重要的是,我们证明了我们的模型推导的不确定性与专家临床logopedic报告中确定为具有挑战性的音素密切相关,据我们所知,这是第一次成功地将模型不确定性与专家评估的语音难度相结合。我们的研究结果表明,这种经过临床验证的、不确定性引导的抽样显著提高了ASR的准确性,为个性化和包容性的ASR提供了一个实用的框架。
【13】SingVERSE: A Diverse, Real-World Benchmark for Singing Voice Enhancement
标题:SingVERSE:歌唱声音增强的多元化、现实世界基准
链接:https://arxiv.org/abs/2509.20969
备注:Demopage: this https URL, Dataset: this https URL
摘要:本文提出了一个歌唱声增强的基准。由于缺乏真实的评价数据,限制了歌唱嗓音增强技术的发展。为了解决这一差距,本文介绍了SingVERSE,这是第一个真实世界的歌唱声音增强基准,涵盖了各种声学场景,并提供了成对的,录音室质量的干净参考。利用SingVERSE,我们对最先进的模型进行全面评估,并发现感知质量和可理解性之间的一致权衡。最后,我们表明,在域内歌唱数据的训练大大提高了增强性能,而不会降低语音能力,建立一个简单而有效的路径前进。这项工作为社区提供了一个基本的基准以及关键的见解,以指导这个尚未开发的领域的未来发展。Demopage:https://singverse.github.io
【14】MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model
标题:MI-SYS:利用闭源大音频语言模型进行无监督领域自适应的标签融合
链接:https://arxiv.org/abs/2509.20706
备注:5 pages, 2 figures, 2 tables
摘要:大型音频语言模型(LALM)在语音任务上表现出很强的zero-shot能力,为语音情感识别(SER)提供了希望。然而,SER在实际部署中经常在域不匹配的情况下失败,其中源数据不可用,并且只能通过API访问强大的LALM。我们要问:如果只给出未标记的目标域音频和仅API的LALM,学生模型是否可以在目标域中超越LALM?为此,我们提出了MI的,去噪标签融合框架,补充了LALM与源域训练的SER分类器作为辅助教师。该框架从两个教师中提取多个随机预测,通过基于互信息的不确定性来加权其平均分布,并使用指数移动平均教师来稳定训练。在三个公共情绪数据集和六个跨域传输的实验显示出一致的收益,学生超过了LALM,并超过了最强的基线3.9%。这种方法在不共享源数据的情况下增强了情感感知语音系统,从而实现了现实的适应。
【15】Building Tailored Speech Recognizers for Japanese Speaking Assessment
标题:构建面向日语口语评估的定制语音识别器
链接:https://arxiv.org/abs/2509.20655
摘要:本文介绍了构建适合日语口语评估任务的语音识别器的方法。具体来说,我们建立了一个语音识别器,输出带有重音标记的音素标签。虽然日语资源丰富,但只有少量数据用于训练模型以产生包括重音标记的准确音素transmits。我们提出了两种方法来减轻数据稀疏。首先,多任务训练方案引入辅助损失函数来估计输入信号的正字法文本标签和音高模式,使得在训练中可以利用仅具有正字法注释的话语。第二个融合两个估计,一个在语音字母串,和其他文本令牌序列。为了结合这些估计,我们开发了一种基于有限状态换能器框架的算法。我们的研究结果表明,使用多任务学习和融合是有效的建立一个准确的音素识别器。我们表明,这种方法是有利的相比,使用通用的多语言识别器。比较了各种方法的相对优势。我们提出的方法降低了平均mora-label错误率从12.3%到7.1%的CSJ核心评估集。
【16】Why Speech Deepfake Detectors Won't Generalize: The Limits of Detection in an Open World
标题:为什么语音Deepfake检测器不会泛化:开放世界中检测的极限
链接:https://arxiv.org/abs/2509.20405
摘要:语音Deepfake检测器通常在干净的基准测试条件下进行评估,但部署发生在移动设备、采样率、编解码器、环境和攻击家族的开放世界中。这为基于人工智能的检测器创造了“覆盖债务”:每一个新的条件都会与现有的条件相乘,从而产生数据盲点,其增长速度超过了数据收集的速度。由于攻击者可以针对这些未覆盖的区域,因此最坏情况下的性能(而不是平均基准得分)决定了安全性。为了证明覆盖债务问题的影响,我们分析了最近的交叉测试框架的结果。在真正的领域和欺骗发布年的性能方面,出现了两种模式:新的合成器消除了传统的工件检测器所依赖的,而会话语音领域(电话会议,采访,社交媒体)一直是最难保护的。这些发现表明,在做出高风险决策时,不应仅仅依靠检测。检测器应该被视为分层防御中的辅助信号,包括出处、人格凭证和策略保障。
机器翻译由腾讯交互翻译提供,仅供参考
