微信公众号:arXiv_Daily
cs.SD语音
标题:超越声学情感识别:使用基于LLM和声学情感模型进行政治演讲中的多模式病态分析
链接:https://arxiv.org/abs/2605.22732
备注:13 pages, 1 figure
摘要:
摘要:
【2】Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
标题:现场音乐扩散模型:交互式扩散音乐生成器的高效微调和后训练链接:https://arxiv.org/abs/2605.22717
摘要:
摘要:
【3】Automatic Contextual Audio Denoising
标题:自动上下文音频去噪链接:https://arxiv.org/abs/2605.22262
摘要:音频上下文确定哪些声音分量和源是相关的,哪些可以被听众感知为不相关的(噪声)。例如,交通噪音在城市监控中是有用的,但在同一地点打电话时噪音是有用的。大多数当前的音频去噪系统应用固定的目标噪声定义,通常在一个上下文中去除有用的分量,而无法抑制不相关的分量。为了解决这个问题,我们引入了自动上下文音频去噪(ACAD)的概念,它根据推断的上下文定义目标和噪声。在这项工作中,我们限制上下文与声学场景类。我们将场景类(噪声)的事件分布之外的声音事件标记为上下文外(OC),将该场景的典型事件标记为上下文内(IC)。我们实现了一种深度学习方法,可以自动推断音频信号的上下文并删除OC组件,并将其与变体进行基准测试:没有上下文推断,使用Oracle上下文,以及单独提供的无信息上下文。在不同背景下配对的干净/嘈杂数据上,其中一个背景下的OC组件可能是另一个背景下的IC,我们提出的方法在标准客观指标上优于其他方法,这表明该模型可以推断上下文,并且上下文相关处理可以增强去噪。
摘要:Audio context determines which sound components and sources are relevant and which can be perceived as irrelevant (noise) by listeners. For example, traffic noise is informative in urban surveillance but noise for a phone call at the same location. Most current audio denoising systems apply fixed target-noise definitions, often removing useful components in one context while failing to suppress irrelevant components. To address this, we introduce the concept automatic contextual audio denoising (ACAD) which defines target and noise based on the inferred context. In this work, we restrict context to be associated with an acoustic scene class. We label sound events outside the event distribution of a scene class (noise) as out-of-context (OC) and events typical for that scene as in-context (IC). We implement a deep learning method that automatically infers the context of the audio signal and removes OC components, and benchmark it against variants: without context inference, with oracle context, and with separately provided uninformative context. On paired clean/noisy data across diverse contexts, where OC components in one context may be IC in another, our proposed method outperforms other approaches across standard objective metrics, indicating that the model can infer context and context-dependent processing can enhance denoising.
【4】RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
标题:RobustSpeechFlow:通过基于增强的对比流匹配学习稳健的文本到语音轨迹链接:https://arxiv.org/abs/2605.22083
备注:Submitted to INTERSPEECH 2026
摘要:
摘要:
【5】Real-time, EDM-inspired sonfication of the activity of a supercomputer
标题:超级计算机活动的实时、受EDM启发的超声处理链接:https://arxiv.org/abs/2605.21874
备注:7 pages, 2 figures, accepted conference paper
摘要:本文所描述的项目探讨了从超级计算机实时接收的数据的信息化。这些数据捕获计算机的所有节点中的当前活动,因此,它们的发音功能作为对节点行为的连续监视的一种形式,并且通过扩展来监视整个系统。由于这种监听在理论上是无休止的,因此产生的发声必须在音乐上能够通过声音以一种既可理解又能长时间吸引人的方式传达信息。我们并没有将预先定义的音乐风格强加给数据,而是试图确定数据本身可以支持的音乐风格。从一小部分候选者中,我们选择了EDM,因为它是一个流派家族,其结构和时间特征与连续的数据驱动过程和长期聆听保持一致。通过这种基于风格的方法,这项研究建立在计算机数据发音的悠久传统之上,同时独特地结合了三个很少一起解决的元素:监控(而不是调试)作为主要目标,实时(而不是事后)数据解释,以及生成几乎无限和风格连贯(而不是不协调)的音乐。
摘要:The project described in this paper explores the informative sonification of data received in real time from a supercomputer. These data capture the current activities in all the nodes of the computer, therefore, their sonification functions as a form of continuous monitoring of the nodes' behavior and, by extension, of the system as a whole. Because such monitoring is theoretically unending, the resulting sonification must be musically capable of conveying information through sound in a way that remains both intelligible and engaging over long durations. Rather than imposing a predefined musical style onto the data, we sought to identify one which the data themselves could plausibly support. From a small set of candidates, we selected EDM because it is a family of genres whose structural and temporal characteristics align well with continuous, data-driven processes and long-term listening. Through this style-based approach, this research builds on the long tradition of computer data sonification while uniquely combining three elements rarely addressed together: monitoring (rather than debugging) as the primary goal, real-time (rather than post-mortem) data interpretation, and generation of virtually infinite and stylistically coherent (rather than incongruous) music.
【6】Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods
标题:学术文本到音乐大挑战:数据集、基线和评估方法链接:https://arxiv.org/abs/2605.21538
备注:Accepted to IEEE ICME 2026 Grand Challenge Paper
摘要:本文介绍了ICME 2026学术文本到音乐生成(ATTM)大挑战的概述和技术框架。尽管文本到音乐生成(TTM)系统取得了快速进展,但该领域目前仍由在具有工业规模计算资源的大规模专有数据集上训练的模型主导,这为学术研究造成了重大障碍。为了解决这个问题,ATTM挑战赛建立了一个公平竞争的基准,要求参与者使用MTG-Jamendo数据集的标准化、CC许可的子集(仅包含器乐),严格从头开始训练生成模型。挑战赛分为两个赛道:效率赛道(限制500 M参数)和性能赛道(无参数限制)。提交的作品通过多阶段的过程进行评估,涉及客观指标,包括Frechet音频距离,CLAP评分和新的概念覆盖评分(CCS),然后进行主观听力测试。通过提供开源基线、预处理管道、参考说明和用于计算FAD和CLAP的公共评估代码,这项挑战旨在促进和促进学术环境中的TTM研究。
摘要:This paper presents an overview and the technical framework of the ICME 2026 Grand Challenge on Academic Text-to-Music Generation (ATTM). Despite the rapid progress in text-to-music generation (TTM) systems, the field is currently dominated by models trained on massive proprietary datasets with industrial-scale computational resources, creating a significant barrier for academic research. To address this, the ATTM Challenge establishes a fair-play benchmark that requires participants to train generative models strictly from scratch using a standardized, CC-licensed subset of the MTG-Jamendo dataset containing only instrumental music. The challenge is divided into two tracks: the Efficiency Track (limited to 500M parameters) and the Performance Track (no parameter limit). Submissions are evaluated through a multi-stage process involving objective metrics, including Frechet Audio Distance, CLAP score, and a novel Concept Coverage Score (CCS), followed by a subjective listening test. By providing open-source baselines, preprocessing pipelines, reference captions, and public evaluation code for computing FAD and CLAP, this challenge aims to facilitate and promote TTM research in academic contexts.
【7】Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation
标题:通过双阶段匹配、多模式注册和连续适应实现有效的用户定义关键词定位链接:https://arxiv.org/abs/2605.22120
备注:14 pages, 13 figures, 12 tables. Accepted by TASLP
摘要:用户定义的关键词定位(KWS)是个性化语音交互的关键,但现有的方法面临着几个挑战:(1)混淆词之间的可辨别性不足,(2)不同发音的说话人之间的性能不一致,以及(3)高数据成本,以确保可靠的唤醒词性能。在本文中,我们介绍DMA-KWS,一个有效的和强大的框架,用户定义的关键字发现。首先,它采用了一个双阶段匹配流水线:CTC解码与流音素搜索定位候选段,其次是QbyT与音素匹配器进行细粒度验证,使其能够更好地区分易混淆的单词。接下来,多模式注册将特定于用户的语音与文本嵌入融合,以进一步提高注册用户的准确性。最后,一个参数高效的持续适应机制使用合成和真实数据执行轻量级更新。大量的实验证明了DMA-KWS的优越性能。在LibriPhrase Hard子集上,它达到了97.85%的AUC和6.13%的EER,达到了最先进的性能。在依赖于说话者的设置中,DMA-KWS始终优于纯文本注册,表现出显着的性能提升。此外,所提出的参数高效的微调机制仅用187 k更新参数来适应DMA-KWS,进一步增强KWS性能,同时确保设备上部署的适用性。
摘要:User-defined keyword spotting (KWS) is crucial for personalized voice interaction, yet existing methods face several challenges: (1) insufficient discriminability among confusable words, (2) performance inconsistency across speakers with varying pronunciations, and (3) high data cost to ensure reliable wake-word performance. In this paper, we introduce DMA-KWS, an efficient and robust framework for user-defined keyword spotting. First, it adopts a dual-stage matching pipeline: CTC decoding with streaming phoneme search to locate candidate segments, followed by QbyT with a phoneme matcher for fine-grained verification, enabling it to better distinguish confusable words. Next, multi-modal enrollment fuses user-specific speech with text embeddings to further improve accuracy for registered users. Finally, a parameter-efficient continual adaptation mechanism performs lightweight updates using synthetic and real data. Extensive experiments demonstrate the superior performance of DMA-KWS. On the LibriPhrase Hard subset, it achieves 97.85% AUC and 6.13% EER, reaching state-of-the-art performance. In speaker-dependent settings, DMA-KWS consistently outperforms text-only enrollment, demonstrating significant performance gains. Moreover, the proposed parameter-efficient fine-tuning mechanism adapts DMA-KWS with only 187k updated parameters, further enhancing KWS performance while ensuring suitability for on-device deployment.
标题:通过双阶段匹配、多模式注册和连续适应实现有效的用户定义关键词定位
链接:https://arxiv.org/abs/2605.22120
备注:14 pages, 13 figures, 12 tables. Accepted by TASLP
摘要:用户定义的关键词定位(KWS)是个性化语音交互的关键,但现有的方法面临着几个挑战:(1)混淆词之间的可辨别性不足,(2)不同发音的说话人之间的性能不一致,以及(3)高数据成本,以确保可靠的唤醒词性能。在本文中,我们介绍DMA-KWS,一个有效的和强大的框架,用户定义的关键字发现。首先,它采用了一个双阶段匹配流水线:CTC解码与流音素搜索定位候选段,其次是QbyT与音素匹配器进行细粒度验证,使其能够更好地区分易混淆的单词。接下来,多模式注册将特定于用户的语音与文本嵌入融合,以进一步提高注册用户的准确性。最后,一个参数高效的持续适应机制使用合成和真实数据执行轻量级更新。大量的实验证明了DMA-KWS的优越性能。在LibriPhrase Hard子集上,它达到了97.85%的AUC和6.13%的EER,达到了最先进的性能。在依赖于说话者的设置中,DMA-KWS始终优于纯文本注册,表现出显着的性能提升。此外,所提出的参数高效的微调机制仅用187 k更新参数来适应DMA-KWS,进一步增强KWS性能,同时确保设备上部署的适用性。
摘要:User-defined keyword spotting (KWS) is crucial for personalized voice interaction, yet existing methods face several challenges: (1) insufficient discriminability among confusable words, (2) performance inconsistency across speakers with varying pronunciations, and (3) high data cost to ensure reliable wake-word performance. In this paper, we introduce DMA-KWS, an efficient and robust framework for user-defined keyword spotting. First, it adopts a dual-stage matching pipeline: CTC decoding with streaming phoneme search to locate candidate segments, followed by QbyT with a phoneme matcher for fine-grained verification, enabling it to better distinguish confusable words. Next, multi-modal enrollment fuses user-specific speech with text embeddings to further improve accuracy for registered users. Finally, a parameter-efficient continual adaptation mechanism performs lightweight updates using synthetic and real data. Extensive experiments demonstrate the superior performance of DMA-KWS. On the LibriPhrase Hard subset, it achieves 97.85% AUC and 6.13% EER, reaching state-of-the-art performance. In speaker-dependent settings, DMA-KWS consistently outperforms text-only enrollment, demonstrating significant performance gains. Moreover, the proposed parameter-efficient fine-tuning mechanism adapts DMA-KWS with only 187k updated parameters, further enhancing KWS performance while ensuring suitability for on-device deployment.
【2】Neighbor-Consistent Neural Filters for Robust Personal Sound Zones Under Localization Uncertainty
标题:定位不确定性下鲁棒个人声区的邻居一致神经过滤器链接:https://arxiv.org/abs/2605.21891
摘要:坐标调节神经网络可以实时生成头部跟踪个人声音区域(PSZ)扬声器滤波器,但它们对定位不确定性很敏感。由光学失真、临时遮挡或跟踪抖动引起的估计的收听者坐标的小波动可能产生大的滤波器变化,即使当收听者物理上静止时也是如此。本文提出了邻域一致性神经滤波器,通过在训练期间惩罚随机扰动的相邻坐标处的滤波器差异来正则化坐标到滤波器的映射。为了评估对跟踪噪声的鲁棒性,我们引入了一个解耦的协议,将声学传递函数固定在一个物理锚点,同时只扰动用于滤波器生成的坐标输入。隔离质量和局部稳定性进行评估,使用邻域中位数和下尾统计的区域间和程序间隔离,连同空间变化率,量化度量灵敏度内的坐标邻域。在使用分离频段低音-高音扬声器系统和25个随机采样锚点位置的仿真中,邻居一致性将低音频段的均方根(RMS)变化率降低了55.9%,高音频段降低了30.3%,同时在很大程度上保持了隔离质量并提高了下尾鲁棒性。在使用24个驱动器阵列和两个固定的头部和躯干模拟器的现场测量中,所提出的正则化将最坏情况下的邻域隔离提高了16.9%,并将空间变化率降低了61.8%。这些结果表明,邻域一致性正则化有效地稳定PSZ渲染的本地化不确定性。
摘要:Coordinate-conditioned neural networks can generate head-tracked personal sound zone (PSZ) loudspeaker filters in real time, but they are sensitive to localization uncertainty. Small fluctuations in estimated listener coordinates, caused by optical distortion, temporary occlusions, or tracking jitter, may produce large filter changes even when listeners are physically stationary. This paper proposes neighbor-consistent neural filters that regularize the coordinate-to-filter mapping by penalizing filter differences at randomly perturbed neighboring coordinates during training. To evaluate robustness against tracking noise, we introduce a decoupled protocol that fixes the acoustic transfer functions at a physical anchor while perturbing only the coordinate inputs used for filter generation. Isolation quality and local stability are evaluated using neighborhood median and lower-tail statistics of inter-zone and inter-program isolation, together with spatial variation rates that quantify metric sensitivity within a coordinate neighborhood. In simulation with a split-band woofer-tweeter system and 25 randomly sampled anchor positions, neighbor consistency reduces the root-mean-square (RMS) variation rate by up to 55.9% in the woofer band and 30.3% in the tweeter band while largely preserving isolation quality and improving lower-tail robustness. In in-situ measurements using a 24-driver array and two stationary head-and-torso simulators, the proposed regularization improves worst-case neighborhood isolation by up to 16.9% and reduces spatial variation rates by up to 61.8%. These results demonstrate that neighbor-consistency regularization effectively stabilizes PSZ rendering under localization uncertainty.
【3】Plug-in Losses for Evidential Deep Learning: A Simplified Framework for Uncertainty Estimation that Includes the Softmax Classifier
标题:证据深度学习的插件损失:包括Softmax分类器的不确定性估计简化框架链接:https://arxiv.org/abs/2605.22746
摘要:现实世界中基于传感器的学习系统需要可靠且计算效率高的不确定性估计。证据深度学习(EDL)通过Dirichlet分布对类概率进行建模,提供单遍不确定性估计,其中Dirichlet参数由学习的神经网络映射预测。然而,这种方法可能会导致计算挑战,因为Dirichlet预期目标比标准的监督学习损失更复杂,使其分析和实现复杂化。我们解决这个问题,通过近似的一阶经验风险最小化问题的目标引起的EDL与插件的损失评估的Dirichlet平均值,并表明,在温和的假设下,近似误差衰减与越来越多的证据广泛的一类损失函数,包括均方误差和交叉熵损失。作为一个特殊的情况下,我们的分析提供了使用softmax的不确定性估计的上下文中的理由,因为在一个特定的证据狄利克雷映射,我们的框架包括标准softmax分类。我们在Google Speech Commands数据集上验证了所提出的简化目标,并表明它们实现了与经典EDL相当的预测准确性和选择性预测性能,同时使用标准深度学习损失和训练管道更容易实现。据我们所知,这种实证分析是第一次通过EDL获得语音识别任务的覆盖率-准确性权衡。
摘要:Real-world sensor-based learning systems require uncertainty estimation that is both reliable and computationally efficient. Evidential Deep Learning (EDL) provides single-pass uncertainty estimation by modeling the class probabilities via Dirichlet distributions, where the Dirichlet parameters are predicted by a learned neural network mapping. However, this approach can lead to computational challenges, as Dirichlet expected objectives are more complex than standard supervised learning losses, complicating their analysis and implementation. We address this issue by approximating the objective of the first-order empirical risk minimization problem induced by EDL with a plug-in loss evaluated at the Dirichlet mean and show that, under mild assumptions, the approximation error decays with growing evidence for a broad class of loss functions, including mean-squared error and cross-entropy loss. As a special case, our analysis provides justification for the use of softmax in the context of uncertainty estimation, since under a particular evidence-to-Dirichlet mapping, our framework includes the standard softmax classifier. We validate the proposed simplified objectives on the Google Speech Commands dataset and show that they achieve predictive accuracy and selective prediction performance comparable to classical EDL, while being simpler to implement using standard deep learning losses and training pipelines. To the best of our knowledge, this empirical analysis is the first to obtain coverage-accuracy trade-offs for speech recognition tasks through EDL.
【4】Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
标题:超越声学情感识别:使用基于LLM和声学情感模型进行政治演讲中的多模式病态分析链接:https://arxiv.org/abs/2605.22732
备注:13 pages, 1 figure
摘要:我们调查声学情感识别模型是否可以作为代理的悲情维度在政治演讲分析,可操作的信任多智能体大语言模型(LLM)管道。费利克斯·巴纳斯扎克(Felix Banaszak)在德国联邦议院全体会议上的讲话(51段,245秒)作为案例研究,我们比较了三种分析模式:(1)emotion2vec_plus_large,一种声学语音情感识别(SER)模型,其连续唤醒值和效价值通过事后Russell环投影导出;(2)双子座2.5闪光灯,法学硕士分析完整的语音音频连同其成绩单在一个开放式的,上下文感知的方式;和(3)TRUST-Pathos分数从三个主张法学硕士主管合奏。Spearman等级相关显示,Gemini效价与TRUST-Pathos强烈相关(rho =+0.664,p < 0.001),而emotion 2 vec效价则不相关(rho =+0.097,p = 0.499)。我们进一步表明,通过系统的质量评估的情感语音的柏林数据库(EMO-DB)使用双子座在一个开放式的注释范例,标准SER基准语料库遭受表演的讲话,文化偏见,和类别不兼容。我们的研究结果表明,基于LLM的多模态分析捕捉语义定义的政治情绪大大优于单独的声学模型,而声学特征仍然为低水平的唤醒估计提供信息。未来的工作将扩展这种方法,以视频为基础的分析,结合面部表情和凝视。
摘要:We investigate whether acoustic emotion recognition models can serve as proxies for the Pathos dimension in political speech analysis, as operationalised by the TRUST multi-agent large language model (LLM) pipeline. Using a Bundestag plenary speech by Felix Banaszak (51 segments, 245 s) as a case study, we compare three analysis modalities: (1) emotion2vec_plus_large, an acoustic speech emotion recognition (SER) model whose continuous Arousal and Valence values are derived via post-hoc Russell Circumplex projection; (2) Gemini 2.5 Flash, an LLM analysing the full speech audio together with its transcript in an open-ended, context-aware fashion; and (3) TRUST-Pathos scores from a three-advocate LLM supervisor ensemble. Spearman rank correlations reveal that Gemini Valence correlates strongly with TRUST-Pathos (rho = +0.664, p < 0.001), whereas emotion2vec Valence does not (rho = +0.097, p = 0.499). We further demonstrate, via a systematic quality evaluation of the Berlin Database of Emotional Speech (EMO-DB) using Gemini in an open-ended annotation paradigm, that standard SER benchmark corpora suffer from acted speech, cultural bias, and category incompatibility. Our results suggest that LLM-based multimodal analysis captures semantically defined political emotion substantially better than acoustic models alone, while acoustic features remain informative for low-level Arousal estimation. Future work will extend this approach to video-based analysis incorporating facial expression and gaze.
【5】Automatic Contextual Audio Denoising
标题:自动上下文音频去噪链接:https://arxiv.org/abs/2605.22262
摘要:音频上下文确定哪些声音分量和源是相关的,哪些可以被听众感知为不相关的(噪声)。例如,交通噪音在城市监控中是有用的,但在同一地点打电话时噪音是有用的。大多数当前的音频去噪系统应用固定的目标噪声定义,通常在一个上下文中去除有用的分量,而无法抑制不相关的分量。为了解决这个问题,我们引入了自动上下文音频去噪(ACAD)的概念,它根据推断的上下文定义目标和噪声。在这项工作中,我们限制上下文与声学场景类。我们将场景类(噪声)的事件分布之外的声音事件标记为上下文外(OC),将该场景的典型事件标记为上下文内(IC)。我们实现了一种深度学习方法,可以自动推断音频信号的上下文并删除OC组件,并将其与变体进行基准测试:没有上下文推断,使用Oracle上下文,以及单独提供的无信息上下文。在不同背景下配对的干净/嘈杂数据上,其中一个背景下的OC组件可能是另一个背景下的IC,我们提出的方法在标准客观指标上优于其他方法,这表明该模型可以推断上下文,并且上下文相关处理可以增强去噪。
摘要:Audio context determines which sound components and sources are relevant and which can be perceived as irrelevant (noise) by listeners. For example, traffic noise is informative in urban surveillance but noise for a phone call at the same location. Most current audio denoising systems apply fixed target-noise definitions, often removing useful components in one context while failing to suppress irrelevant components. To address this, we introduce the concept automatic contextual audio denoising (ACAD) which defines target and noise based on the inferred context. In this work, we restrict context to be associated with an acoustic scene class. We label sound events outside the event distribution of a scene class (noise) as out-of-context (OC) and events typical for that scene as in-context (IC). We implement a deep learning method that automatically infers the context of the audio signal and removes OC components, and benchmark it against variants: without context inference, with oracle context, and with separately provided uninformative context. On paired clean/noisy data across diverse contexts, where OC components in one context may be IC in another, our proposed method outperforms other approaches across standard objective metrics, indicating that the model can infer context and context-dependent processing can enhance denoising.
【6】RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
标题:RobustSpeechFlow:通过基于增强的对比流匹配学习稳健的文本到语音轨迹链接:https://arxiv.org/abs/2605.22083
备注:Submitted to INTERSPEECH 2026
摘要:虽然流匹配文本到语音(TTS)实现了强的zero-shot说话者相似性和自然度,但它仍然容易受到内容保真度问题的影响,特别是来自不完美对齐的跳过和重复错误。我们提出了RobustSpeechFlow,这是一种训练策略,通过扩展具有长度保持重复和跳过潜在增强的对比流匹配来提高对齐鲁棒性。不需要外部校准器或偏好数据,我们的方法直接惩罚现实的故障模式,并很容易集成到现有的管道。在Seed-TTS-eval上,它仅使用0.06B参数就将字错误率(WER)从1.44降低到1.38。在我们的ZERO 500基准测试中,它在不同的说话者和韵律条件下提供了一致的可懂度改善;在NFE=24时,它将英语字符错误率(CER)从0.48\%降低到0.35\%,将韩语CER从0.81\%降低到0.57\%。音频示例:https://robustspeechflow.github.io/
摘要:While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repeat errors from imperfect alignment. We propose RobustSpeechFlow, a training strategy that improves alignment robustness by extending contrastive flow matching with length-preserving repeat and skip latent augmentations. Requiring no external aligners or preference data, our method directly penalizes realistic failure modes and readily integrates into existing pipelines. On Seed-TTS-eval, it reduces the word error rate (WER) from 1.44 to 1.38 using only 0.06B parameters. On our ZERO500 benchmark, it delivers consistent intelligibility improvements across diverse speaker and prosody conditions; at NFE=24, it reduces English character error rate (CER) from 0.48\% to 0.35\% and Korean CER from 0.81\% to 0.57\%. Audio samples: https://robustspeechflow.github.io/
机器翻译由腾讯交互翻译提供,仅供参考
