微信公众号:arXiv_Daily
cs.SD语音
【1】Covertly improving intelligibility with data-driven adaptations of speech timing
标题:通过数据驱动的语音计时调整隐性提高清晰度
链接:https://arxiv.org/abs/2603.30032
摘要:人类说话者经常通过全局放慢他们的讲话来解决具有语言理解挑战的听众,例如听力困难或非母语成年人。然而,目前还不清楚这种策略是否真的使语音更容易理解。在这里,我们利用机器生成语音的最新进展,允许更精确地控制语速,以便系统地研究有针对性的语速调整如何提高理解力。我们首先使用反向相关实验表明,语音速率的时间影响之前,目标元音对比度(例如。紧张-放松的区别)实际上表现为一种剪刀状的模式,在早期和晚期的语境窗口中具有相反的效果;这种模式在个体内以及在母语为法语、汉语和日语的母语为英语的母语为英语的听众和母语为英语的母语为英语的听众中都非常稳定。其次,我们表明,这种语音速率结构不仅有利于L2听众的目标元音对比度的理解,但本机听众也依赖于这种模式在具有挑战性的声学条件。最后,我们建立了一个数据驱动的文本到语音的算法,复制新的语音序列的时间结构。在不同的句子和元音对比中,听众仍然没有意识到这种有针对性的放慢可以提高单词的理解能力。引人注目的是,参与者反而认为全球经济放缓的共同战略更加清晰,尽管它实际上增加了理解错误。总而言之,这些结果表明,对语音速率的有针对性的调整在具有挑战性的条件下显着提高了清晰度,同时往往被忽视。更一般地说,本文提供了一种数据驱动的方法,以提高机器生成的语音,可以扩展到语音理解的其他方面和各种各样的听众和环境的可访问性。
摘要:Human talkers often address listeners with language-comprehension challenges, such as hard-of-hearing or non-native adults, by globally slowing down their speech. However, it remains unclear whether this strategy actually makes speech more intelligible. Here, we take advantage of recent advancements in machine-generated speech allowing more precise control of speech rate in order to systematically examine how targeted speech-rate adjustments may improve comprehension. We first use reverse-correlation experiments to show that the temporal influence of speech rate prior to a target vowel contrast (ex. the tense-lax distinction) in fact manifests in a scissor-like pattern, with opposite effects in early versus late context windows; this pattern is remarkably stable both within individuals and across native L1-English listeners and L2-English listeners with French, Mandarin, and Japanese L1s. Second, we show that this speech rate structure not only facilitates L2 listeners' comprehension of the target vowel contrast, but that native listeners also rely on this pattern in challenging acoustic conditions. Finally, we build a data-driven text-to-speech algorithm that replicates this temporal structure on novel speech sequences. Across a variety of sentences and vowel contrasts, listeners remained unaware that such targeted slowing improved word comprehension. Strikingly, participants instead judged the common strategy of global slowing as clearer, even though it actually increased comprehension errors. Together, these results show that targeted adjustments to speech rate significantly aid intelligibility under challenging conditions, while often going unnoticed. More generally, this paper provides a data-driven methodology to improve the accessibility of machine-generated speech which can be extended to other aspects of speech comprehension and a wide variety of listeners and environments.
【2】SIREN: Spatially-Informed Reconstruction of Binaural Audio with Vision
标题:SIREN:立体声立体声视觉重建
链接:https://arxiv.org/abs/2603.29820
备注:5 pages, 1 figure, to appear in ICASSP 2026
摘要:双耳音频提供了沉浸感所必需的空间线索,但由于捕获限制,大多数消费者视频都是单声道的。我们介绍SIREN,一个视觉引导的单声道到双耳框架,明确预测左声道和右声道。基于ViT的编码器学习双头自我注意力,以产生共享的场景映射和端到端的L/R注意力,取代手工制作的遮罩。一个软的,退火的空间先验轻轻地偏置早期的L/R接地,和两个阶段,置信加权波形域融合(由单声道重建和耳间相位一致性引导)抑制串扰时,聚合多作物和重叠窗口。通过FAIR-Play和MUSIC-Stereo评估,SIREN在时频和相位敏感指标上产生一致的增益,并具有竞争力的SNR。该设计是模块化和通用的,不需要特定于任务的注释,并与标准的视听管道集成。
摘要:Binaural audio delivers spatial cues essential for immersion, yet most consumer videos are monaural due to capture constraints. We introduce SIREN, a visually guided mono to binaural framework that explicitly predicts left and right channels. A ViT-based encoder learns dual-head self-attention to produce a shared scene map and end-to-end L/R attention, replacing hand-crafted masks. A soft, annealed spatial prior gently biases early L/R grounding, and a two-stage, confidence-weighted waveform-domain fusion (guided by mono reconstruction and interaural phase consistency) suppresses crosstalk when aggregating multi-crop and overlapping windows. Evaluated on FAIR-Play and MUSIC-Stereo, SIREN yields consistent gains on time-frequency and phase-sensitive metrics with competitive SNR. The design is modular and generic, requires no task-specific annotations, and integrates with standard audio-visual pipelines.
【3】A Comprehensive Corpus of Biomechanically Constrained Piano Chords: Generation, Analysis, and Implications for Voicing and Psychoacoustics
标题:生物力学约束钢琴和弦的综合数据库:发声和心理声学的生成、分析及其含义
链接:https://arxiv.org/abs/2603.29710
备注:10 pages, 3 figures
摘要:我介绍了已知最大的可演奏钢琴和弦开源语料库(约1930万个条目)的生成和分析。这个数据集列举了双手搜索空间受到生物力学约束(双手,每只手1.5倍频程)到前所未有的程度。为了证明语料库的效用,发声形状和心理声学目标之间的关系进行了建模。和声性被证明是音高类别特性的内在特征:发声统计增加了可忽略的方差($ΔR^2 \approx 0.014\%$,$p \approx 0.13$)。相反,发声显着预测不和谐音($ΔR^2 \approx 6.75\%$,$p \approx 0.0008$)。重要的是,在预测粗糙度方面,偏度($β\approx +0.145 $)比扩散($β\approx-0.025 $)有效约5.8倍。分析挑战教学强调“传播”:偏度是一个更强的预测不和谐比传播。这表明,在“开放式”清晰度驱动的宽度比负偏度少;实现较低的登记清除放置在底部宽的差距,并允许在高音更紧密的集群。结果表明,语料库的能力,使未来的研究,特别是在生成建模,语音领导的拓扑结构和心理声学分析等领域。
摘要:I present the generation and analysis of the largest known open-source corpus of playable piano chords (approximately 19.3 million entries). This dataset enumerates the two-handed search space subject to biomechanical constraints (two hands, each with 1.5 octave reach) to an unprecedented extent. To demonstrate the corpus's utility, the relationship between voicing shape and psychoacoustic targets was modeled. Harmonicity proved intrinsic to pitch-class identity: voicing statistics added negligible variance ($ΔR^2 \approx 0.014\%$, $p \approx 0.13$). Conversely, voicing significantly predicted dissonance ($ΔR^2 \approx 6.75\%$, $p \approx 0.0008$). Crucially, skewness ($β\approx +0.145$) was approximately 5.8$\times$ more effective than spread ($β\approx -0.025$) at predicting roughness. The analysis challenges the pedagogical emphasis on ``spread'': skewness is a stronger predictor of dissonance than spread. This suggests that clarity in ``open voicings'' is driven less by width than by negative skewness; achieving lower-register clearance by placing wide gaps at the bottom and allowing tighter clustering in the treble. The results demonstrate the corpus's ability to enable future research, especially in areas such as generative modeling, voice-leading topology, and psychoacoustic analysis.
【4】LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
标题:LongCat-AudioDiT:波形潜在空间中的高保真扩散文本到语音
链接:https://arxiv.org/abs/2603.29339
备注:Code and model weights are available at https://github.com/meituan-longcat/LongCat-AudioDiT
摘要:我们提出了LongCat-AudioDiT,一种新的,基于非自回归扩散的文本到语音(TTS)模型,实现了最先进的(SOTA)性能。与之前依赖于中间声学表示(如梅尔频谱图)的方法不同,LongCat-AudioDiT的核心创新在于直接在波形潜在空间内操作。这种方法有效地减轻了复合误差,并大大简化了TTS管道,只需要一个波形变分自动编码器(Wav-VAE)和扩散骨干。此外,我们对推理过程进行了两项关键改进:首先,我们识别并纠正了长期存在的训练-推理不匹配;其次,我们用自适应投影指导取代了传统的无分类器指导,以提高生成质量。实验结果表明,尽管缺乏复杂的多阶段训练管道或高质量的人类注释数据集,LongCat-AudioDiT在种子基准上实现了SOTA zero-shot语音克隆性能,同时保持了竞争力的可懂度。具体来说,我们最大的变体LongCat-AudioDiT-3.5B的性能优于之前的SOTA模型(Seed-TTS),将Seed-ZH上的说话人相似性(SIM)分数从0.809提高到0.818,Seed-Hard上从0.776提高到0.797。最后,通过全面的消融研究和系统的分析,我们验证了我们提出的模块的有效性。值得注意的是,我们研究了Wav-VAE和TTS骨干之间的相互作用,揭示了违反直觉的发现,即Wav-VAE中的卓越重建保真度并不一定会导致更好的整体TTS性能。发布代码和模型权重,以促进语音社区内的进一步研究。
摘要:We present LongCat-AudioDiT, a novel, non-autoregressive diffusion-based text-to-speech (TTS) model that achieves state-of-the-art (SOTA) performance. Unlike previous methods that rely on intermediate acoustic representations such as mel-spectrograms, the core innovation of LongCat-AudioDiT lies in operating directly within the waveform latent space. This approach effectively mitigates compounding errors and drastically simplifies the TTS pipeline, requiring only a waveform variational autoencoder (Wav-VAE) and a diffusion backbone. Furthermore, we introduce two critical improvements to the inference process: first, we identify and rectify a long-standing training-inference mismatch; second, we replace traditional classifier-free guidance with adaptive projection guidance to elevate generation quality. Experimental results demonstrate that, despite the absence of complex multi-stage training pipelines or high-quality human-annotated datasets, LongCat-AudioDiT achieves SOTA zero-shot voice cloning performance on the Seed benchmark while maintaining competitive intelligibility. Specifically, our largest variant, LongCat-AudioDiT-3.5B, outperforms the previous SOTA model (Seed-TTS), improving the speaker similarity (SIM) scores from 0.809 to 0.818 on Seed-ZH, and from 0.776 to 0.797 on Seed-Hard. Finally, through comprehensive ablation studies and systematic analysis, we validate the effectiveness of our proposed modules. Notably, we investigate the interplay between the Wav-VAE and the TTS backbone, revealing the counterintuitive finding that superior reconstruction fidelity in the Wav-VAE does not necessarily lead to better overall TTS performance. Code and model weights are released to foster further research within the speech community.
【5】Real-Time Band-Grouped Vocal Denoising Using Sigmoid-Driven Ideal Ratio Masking
标题:基于Sigmoid驱动的理想比率掩蔽的实时带分组人声去噪
链接:https://arxiv.org/abs/2603.29326
摘要:基于深度学习的实时语音去噪在过去几年中取得了重大进展,证明了人工智能在保持语音自然度的同时提高信噪比(SNR)的能力。然而,许多深度学习方法具有很高的延迟,并且需要很长的上下文帧,这使得它们难以为实时应用程序进行配置。为了解决这些挑战,我们提出了一个S形驱动的理想比率掩模训练的频谱损失,以鼓励增加SNR和最大化的感知质量的声音。该模型使用了频带分组的编码器-解码器架构与频率的注意力,并实现了小于10,ms的总延迟,与PESQ-WB的改进0.21上的平稳噪声和0.12上的非平稳噪声。
摘要:Real-time, deep learning-based vocal denoising has seen significant progress over the past few years, demonstrating the capability of artificial intelligence in preserving the naturalness of the voice while increasing the signal-to-noise ratio (SNR). However, many deep learning approaches have high amounts of latency and require long frames of context, making them difficult to configure for live applications. To address these challenges, we propose a sigmoid-driven ideal ratio mask trained with a spectral loss to encourage an increased SNR and maximized perceptual quality of the voice. The proposed model uses a band-grouped encoder-decoder architecture with frequency attention and achieves a total latency of less than 10,ms, with PESQ-WB improvements of 0.21 on stationary noise and 0.12 on nonstationary noise.
【6】Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
标题:音频幻觉攻击:探索大型音频语言模型的可靠性
链接:https://arxiv.org/abs/2603.29263
摘要:大型音频语言模型(LALM)在音频语言任务上实现了强大的性能;然而,它们在现实世界中的可靠性仍然有待探索。我们介绍了音频幻觉攻击(AHA),一种称为AHA-Eval的攻击套件,包括6.5K QA对,旨在测试LALM是否真正将其响应置于音频输入中。AHA针对两个攻击面:(i)基于查询的攻击,其利用问题结构来诱导关于不存在的声音的幻觉,以及(ii)基于音频的攻击,其将描述不存在的事件的合成语音注入音频流。评估最先进的LALM,包括Audio Flamingo 3和Gemini 3 Pro,我们观察到高攻击成功率分别为95.35%和79.65%,揭示了标准基准性能隐藏的可靠性差距。为了缓解这一问题,我们提出了一个120 K QA后对齐数据集AHA-Guard,它成功地将攻击成功率降低了49%。
摘要:Large Audio Language Models (LALMs) achieve strong performance on audio-language tasks; however, their reliability in real-world settings remains underexplored. We introduce Audio Hallucination Attacks (AHA), an attack suite called AHA-Eval, comprising 6.5K QA pairs designed to test whether LALMs genuinely ground their responses in the audio input. AHA targets two attack surfaces: (i) query-based attacks, which exploit question structure to induce hallucinations about absent sounds, and (ii) audio-based attacks, which inject synthetic speech describing non-existent events into the audio stream. Evaluating state-of-the-art LALMs, including Audio Flamingo 3 and Gemini 3 Pro, we observe high attack success rates of 95.35% and 79.65%, respectively, revealing a reliability gap that is hidden by standard benchmark performance. To mitigate this, we propose a 120K QA post-alignment dataset, AHA-Guard, which successfully reduces attack success rates by up to 49%.
【7】IQRA 2026: Interspeech Challenge on Automatic Assessment Pronunciation for Modern Standard Arabic (MSA)
标题:MQRA 2026:现代标准阿拉伯语(GMA)自动评估发音的过渡挑战
链接:https://arxiv.org/abs/2603.29087
备注:5 pages paper
摘要:我们提出了第二版的IQRA语际挑战,对现代标准阿拉伯语(MSA)的自动发音错误检测和诊断(MDD)的挑战的结果。在上一版的基础上,本次迭代引入了\textbf{Iqra\_Extra\_IS26},这是一个新的真实人类发音错误的语音数据集,补充了现有的训练和评估资源。提交的系统采用了多种方法,包括基于CTC的自监督学习模型、两阶段微调策略以及使用大型音频语言模型。与第一版相比,我们观察到\textbf{0.28 in F1-score}的大幅跃升,这既归功于参与者提出的新颖架构和建模策略,也归功于提供的额外真实的发音错误数据。这些结果表明,阿拉伯语MDD研究的日益成熟,并建立了一个更强大的基础,为今后的工作,在阿拉伯语发音评估。
摘要:We present the findings of the second edition of the IQRA Interspeech Challenge, a challenge on automatic Mispronunciation Detection and Diagnosis (MDD) for Modern Standard Arabic (MSA). Building on the previous edition, this iteration introduces \textbf{Iqra\_Extra\_IS26}, a new dataset of authentic human mispronounced speech, complementing the existing training and evaluation resources. Submitted systems employed a diverse range of approaches, spanning CTC-based self-supervised learning models, two-stage fine-tuning strategies, and using large audio-language models. Compared to the first edition, we observe a substantial jump of \textbf{0.28 in F1-score}, attributable both to novel architectures and modeling strategies proposed by participants and to the additional authentic mispronunciation data made available. These results demonstrate the growing maturity of Arabic MDD research and establish a stronger foundation for future work in Arabic pronunciation assessment.
【8】Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
标题:推进基于LLM的音素到字素以实现多语言语音识别
链接:https://arxiv.org/abs/2603.29217
备注:Update after INTERSPEECH2026 submission
摘要:基于音素的ASR将识别分解为语音到音素(S2 P)和音素到字素(P2 G),从而实现跨语言声学共享,同时将特定语言的正字法保持在单独的模块中。虽然大型语言模型(LLM)是有希望的P2 G,多语言P2 G仍然具有挑战性,由于语言感知的生成和严重的跨语言数据不平衡。我们研究了基于多语言LLM的P2 G十语言CV-Lang 10基准。我们研究了考虑S2 P不确定性的鲁棒性策略,包括DANP和简化SKM(S-SKM)。S-SKM是一种蒙特卡洛近似,避免了P2 G训练中基于CTC的S2 P概率加权。强大的训练和低资源过采样将平均WER从10.56%降低到7.66%。
摘要:Phoneme-based ASR factorizes recognition into speech-to-phoneme (S2P) and phoneme-to-grapheme (P2G), enabling cross-lingual acoustic sharing while keeping language-specific orthography in a separate module. While large language models (LLMs) are promising for P2G, multilingual P2G remains challenging due to language-aware generation and severe cross-language data imbalance. We study multilingual LLM-based P2G on the ten-language CV-Lang10 benchmark. We examine robustness strategies that account for S2P uncertainty, including DANP and Simplified SKM (S-SKM). S-SKM is a Monte Carlo approximation that avoids CTC-based S2P probability weighting in P2G training. Robust training and low-resource oversampling reduce the average WER from 10.56% to 7.66%.
【9】Asymmetric Encoder-Decoder Based on Time-Frequency Correlation for Speech Separation
标题:基于时频相关的非对称编解码器语音分离
链接:https://arxiv.org/abs/2603.29097
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing (T-ASLP)
摘要:在现实的声学环境中的语音分离仍然具有挑战性,因为重叠的扬声器,背景噪声和混响必须同时解决。虽然最近的时频域模型表现出了很强的性能,但大多数仍然依赖于后期分裂架构,其中扬声器解纠缠被推迟到最后阶段,从而在不利条件下产生了信息瓶颈并削弱了可辨别性。为了解决这个问题,我们提出了SR-CorrNet,一个非对称的编码器-解码器框架,引入分离-重构(SepRe)策略到TF双路径骨干。编码器从混合观测中执行粗略分离,而权重共享解码器通过跨说话者交互逐步重建说话者区分特征,从而实现逐阶段细化。为了补充这种架构,我们将语音分离公式化为结构化的相关滤波器问题:从观察结果计算的空间-频谱-时间相关性用作输入特征,并估计相应的深度滤波器以恢复目标信号。我们还结合了一个基于吸引子的动态分割模块,以适应实际扬声器配置的输出流的数量。WSJ0-2/3/4/5Mix,WHAMR!,和LibriCSS在单通道和多通道设置中的消声、噪声混响和真实记录条件下表现出一致的改进,突出了TF域SepRe与基于相关性的滤波器估计用于语音分离的有效性。
摘要:Speech separation in realistic acoustic environments remains challenging because overlapping speakers, background noise, and reverberation must be resolved simultaneously. Although recent time-frequency (TF) domain models have shown strong performance, most still rely on late-split architectures, where speaker disentanglement is deferred to the final stage, creating an information bottleneck and weakening discriminability under adverse conditions. To address this issue, we propose SR-CorrNet, an asymmetric encoder-decoder framework that introduces the separation-reconstruction (SepRe) strategy into a TF dual-path backbone. The encoder performs coarse separation from mixture observations, while the weight-shared decoder progressively reconstructs speaker-discriminative features with cross-speaker interaction, enabling stage-wise refinement. To complement this architecture, we formulate speech separation as a structured correlation-to-filter problem: spatio-spectro-temporal correlations computed from the observations are used as input features, and the corresponding deep filters are estimated to recover target signals. We further incorporate an attractor-based dynamic split module to adapt the number of output streams to the actual speaker configuration. Experimental results on WSJ0-2/3/4/5Mix, WHAMR!, and LibriCSS demonstrate consistent improvements across anechoic, noisy-reverberant, and real-recorded conditions in both single- and multi-channel settings, highlighting the effectiveness of TF-domain SepRe with correlation-based filter estimation for speech separation.
【1】An Information-Theoretic Method for Dynamic System Identification With Output-Only Damping Estimation
标题:基于仅输出衰减估计的动态系统识别的信息论方法
链接:https://arxiv.org/abs/2603.29956
备注:18 pages, 16 figures, 4 tables. Published in Journal of Dynamic Systems, Measurement, and Control (ASME), 2026. Licensed under CC BY 4.0
摘要:一种新的信息理论方法的系统识别能力在这里进行检查。具体来说,这项工作使用信息理论的度量和振动为基础的测量,以提高机械系统的阻尼估计精度。该方法涉及系统识别、信号处理、监控和警报系统中的关键限制。这些系统集成了各种组件,包括传感器、数据采集设备和警报机制。它们被设计用于在计算关键参数的环境中运行,例如峰值加速度和高加速度值的持续时间。然而,当前的操作模态识别方法由于其经验性质而受到与获得差的阻尼估计相关的限制。这对警报系统有重大影响。这发生在它们的持续时间被错误估计时;特别是,当使用振动幅度作为损坏或异常检测场景中的监视系统的危险警报的指示时。为此,提出了基于香农熵和Kullback-Leibler发散概念的方法。其主要目的是近实时监控振动水平,并在超过预定义阈值时立即发出警报。在考虑所提出的方法,新的现实世界的数据从多轴仿真表在巴斯大学,以及基准的国际结构控制协会-美国土木工程师协会(IASC-ASCE)的结构健康监测问题。重要的是,该方法被证明是选择最佳模型,准确地捕捉正确的警报持续时间,系统识别和监控提供了一个强大的工具。
摘要:The system identification capabilities of a novel information-theoretic method are examined here. Specifically, this work uses information-theoretic metrics and vibration-based measurements to enhance damping estimation accuracy in mechanical systems. The method refers to a key limitation in system identification, signal processing, monitoring, and alert systems. These systems integrate various components, including sensors, data acquisition devices, and alert mechanisms. They are designed to operate in an environment to calculate key parameters such as peak accelerations and duration of high acceleration values. The current operational modal identification methods, though, suffer from limitations related to obtaining poor damping estimates due to their empirical nature. This has a significant impact on alert warning systems. This occurs when their duration is misestimated; specifically, when using the vibration amplitudes as an indicator of danger alerts for monitoring systems in damage or anomaly detection scenarios. To this end, approaches based on the Shannon entropy and the Kullback-Leibler divergence concept are proposed. The primary objective is to monitor the vibration levels in near real-time and provide immediate alerts when predefined thresholds are exceeded. In considering the proposed approach, both new real-world data from the multi-axis simulation table at the University of Bath, as well as the benchmark International Association for Structural Control-American Society of Civil Engineers (IASC-ASCE) structural health monitoring problem are considered. Importantly, the approach is shown to select the optimal model, which accurately captures the correct alert duration, providing a powerful tool for system identification and monitoring.
【2】Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
标题:推进基于LLM的音素到字素以实现多语言语音识别
链接:https://arxiv.org/abs/2603.29217
备注:Update after INTERSPEECH2026 submission
摘要:基于音素的ASR将识别分解为语音到音素(S2 P)和音素到字素(P2 G),从而实现跨语言声学共享,同时将特定语言的正字法保持在单独的模块中。虽然大型语言模型(LLM)是有希望的P2 G,多语言P2 G仍然具有挑战性,由于语言感知的生成和严重的跨语言数据不平衡。我们研究了基于多语言LLM的P2 G十语言CV-Lang 10基准。我们研究了考虑S2 P不确定性的鲁棒性策略,包括DANP和简化SKM(S-SKM)。S-SKM是一种蒙特卡洛近似,避免了P2 G训练中基于CTC的S2 P概率加权。强大的训练和低资源过采样将平均WER从10.56%降低到7.66%。
摘要:Phoneme-based ASR factorizes recognition into speech-to-phoneme (S2P) and phoneme-to-grapheme (P2G), enabling cross-lingual acoustic sharing while keeping language-specific orthography in a separate module. While large language models (LLMs) are promising for P2G, multilingual P2G remains challenging due to language-aware generation and severe cross-language data imbalance. We study multilingual LLM-based P2G on the ten-language CV-Lang10 benchmark. We examine robustness strategies that account for S2P uncertainty, including DANP and Simplified SKM (S-SKM). S-SKM is a Monte Carlo approximation that avoids CTC-based S2P probability weighting in P2G training. Robust training and low-resource oversampling reduce the average WER from 10.56% to 7.66%.
【3】Asymmetric Encoder-Decoder Based on Time-Frequency Correlation for Speech Separation
标题:基于时频相关的非对称编解码器语音分离
链接:https://arxiv.org/abs/2603.29097
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing (T-ASLP)
摘要:在现实的声学环境中的语音分离仍然具有挑战性,因为重叠的扬声器,背景噪声和混响必须同时解决。虽然最近的时频域模型表现出了很强的性能,但大多数仍然依赖于后期分裂架构,其中扬声器解纠缠被推迟到最后阶段,从而在不利条件下产生了信息瓶颈并削弱了可辨别性。为了解决这个问题,我们提出了SR-CorrNet,一个非对称的编码器-解码器框架,引入分离-重构(SepRe)策略到TF双路径骨干。编码器从混合观测中执行粗略分离,而权重共享解码器通过跨说话者交互逐步重建说话者区分特征,从而实现逐阶段细化。为了补充这种架构,我们将语音分离公式化为结构化的相关滤波器问题:从观察结果计算的空间-频谱-时间相关性用作输入特征,并估计相应的深度滤波器以恢复目标信号。我们还结合了一个基于吸引子的动态分割模块,以适应实际扬声器配置的输出流的数量。WSJ0-2/3/4/5Mix,WHAMR!,和LibriCSS在单通道和多通道设置中的消声、噪声混响和真实记录条件下表现出一致的改进,突出了TF域SepRe与基于相关性的滤波器估计用于语音分离的有效性。
摘要:Speech separation in realistic acoustic environments remains challenging because overlapping speakers, background noise, and reverberation must be resolved simultaneously. Although recent time-frequency (TF) domain models have shown strong performance, most still rely on late-split architectures, where speaker disentanglement is deferred to the final stage, creating an information bottleneck and weakening discriminability under adverse conditions. To address this issue, we propose SR-CorrNet, an asymmetric encoder-decoder framework that introduces the separation-reconstruction (SepRe) strategy into a TF dual-path backbone. The encoder performs coarse separation from mixture observations, while the weight-shared decoder progressively reconstructs speaker-discriminative features with cross-speaker interaction, enabling stage-wise refinement. To complement this architecture, we formulate speech separation as a structured correlation-to-filter problem: spatio-spectro-temporal correlations computed from the observations are used as input features, and the corresponding deep filters are estimated to recover target signals. We further incorporate an attractor-based dynamic split module to adapt the number of output streams to the actual speaker configuration. Experimental results on WSJ0-2/3/4/5Mix, WHAMR!, and LibriCSS demonstrate consistent improvements across anechoic, noisy-reverberant, and real-recorded conditions in both single- and multi-channel settings, highlighting the effectiveness of TF-domain SepRe with correlation-based filter estimation for speech separation.
【4】A Comprehensive Corpus of Biomechanically Constrained Piano Chords: Generation, Analysis, and Implications for Voicing and Psychoacoustics
标题:生物力学约束钢琴和弦的综合数据库:发声和心理声学的生成、分析及其含义
链接:https://arxiv.org/abs/2603.29710
备注:10 pages, 3 figures
摘要:我介绍了已知最大的可演奏钢琴和弦开源语料库(约1930万个条目)的生成和分析。这个数据集列举了双手搜索空间受到生物力学约束(双手,每只手1.5倍频程)到前所未有的程度。为了证明语料库的效用,发声形状和心理声学目标之间的关系进行了建模。和声性被证明是音高类别特性的内在特征:发声统计增加了可忽略的方差($ΔR^2 \approx 0.014\%$,$p \approx 0.13$)。相反,发声显着预测不和谐音($ΔR^2 \approx 6.75\%$,$p \approx 0.0008$)。重要的是,在预测粗糙度方面,偏度($β\approx +0.145 $)比扩散($β\approx-0.025 $)有效约5.8倍。分析挑战教学强调“传播”:偏度是一个更强的预测不和谐比传播。这表明,在“开放式”清晰度驱动的宽度比负偏度少;实现较低的登记清除放置在底部宽的差距,并允许在高音更紧密的集群。结果表明,语料库的能力,使未来的研究,特别是在生成建模,语音领导的拓扑结构和心理声学分析等领域。
摘要:I present the generation and analysis of the largest known open-source corpus of playable piano chords (approximately 19.3 million entries). This dataset enumerates the two-handed search space subject to biomechanical constraints (two hands, each with 1.5 octave reach) to an unprecedented extent. To demonstrate the corpus's utility, the relationship between voicing shape and psychoacoustic targets was modeled. Harmonicity proved intrinsic to pitch-class identity: voicing statistics added negligible variance ($ΔR^2 \approx 0.014\%$, $p \approx 0.13$). Conversely, voicing significantly predicted dissonance ($ΔR^2 \approx 6.75\%$, $p \approx 0.0008$). Crucially, skewness ($β\approx +0.145$) was approximately 5.8$\times$ more effective than spread ($β\approx -0.025$) at predicting roughness. The analysis challenges the pedagogical emphasis on ``spread'': skewness is a stronger predictor of dissonance than spread. This suggests that clarity in ``open voicings'' is driven less by width than by negative skewness; achieving lower-register clearance by placing wide gaps at the bottom and allowing tighter clustering in the treble. The results demonstrate the corpus's ability to enable future research, especially in areas such as generative modeling, voice-leading topology, and psychoacoustic analysis.
【5】LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
标题:LongCat-AudioDiT:波形潜在空间中的高保真扩散文本到语音
链接:https://arxiv.org/abs/2603.29339
备注:Code and model weights are available at https://github.com/meituan-longcat/LongCat-AudioDiT
摘要:我们提出了LongCat-AudioDiT,一种新的,基于非自回归扩散的文本到语音(TTS)模型,实现了最先进的(SOTA)性能。与之前依赖于中间声学表示(如梅尔频谱图)的方法不同,LongCat-AudioDiT的核心创新在于直接在波形潜在空间内操作。这种方法有效地减轻了复合误差,并大大简化了TTS管道,只需要一个波形变分自动编码器(Wav-VAE)和扩散骨干。此外,我们对推理过程进行了两项关键改进:首先,我们识别并纠正了长期存在的训练-推理不匹配;其次,我们用自适应投影指导取代了传统的无分类器指导,以提高生成质量。实验结果表明,尽管缺乏复杂的多阶段训练管道或高质量的人类注释数据集,LongCat-AudioDiT在种子基准上实现了SOTA zero-shot语音克隆性能,同时保持了竞争力的可懂度。具体来说,我们最大的变体LongCat-AudioDiT-3.5B的性能优于之前的SOTA模型(Seed-TTS),将Seed-ZH上的说话人相似性(SIM)分数从0.809提高到0.818,Seed-Hard上从0.776提高到0.797。最后,通过全面的消融研究和系统的分析,我们验证了我们提出的模块的有效性。值得注意的是,我们研究了Wav-VAE和TTS骨干之间的相互作用,揭示了违反直觉的发现,即Wav-VAE中的卓越重建保真度并不一定会导致更好的整体TTS性能。发布代码和模型权重,以促进语音社区内的进一步研究。
摘要:We present LongCat-AudioDiT, a novel, non-autoregressive diffusion-based text-to-speech (TTS) model that achieves state-of-the-art (SOTA) performance. Unlike previous methods that rely on intermediate acoustic representations such as mel-spectrograms, the core innovation of LongCat-AudioDiT lies in operating directly within the waveform latent space. This approach effectively mitigates compounding errors and drastically simplifies the TTS pipeline, requiring only a waveform variational autoencoder (Wav-VAE) and a diffusion backbone. Furthermore, we introduce two critical improvements to the inference process: first, we identify and rectify a long-standing training-inference mismatch; second, we replace traditional classifier-free guidance with adaptive projection guidance to elevate generation quality. Experimental results demonstrate that, despite the absence of complex multi-stage training pipelines or high-quality human-annotated datasets, LongCat-AudioDiT achieves SOTA zero-shot voice cloning performance on the Seed benchmark while maintaining competitive intelligibility. Specifically, our largest variant, LongCat-AudioDiT-3.5B, outperforms the previous SOTA model (Seed-TTS), improving the speaker similarity (SIM) scores from 0.809 to 0.818 on Seed-ZH, and from 0.776 to 0.797 on Seed-Hard. Finally, through comprehensive ablation studies and systematic analysis, we validate the effectiveness of our proposed modules. Notably, we investigate the interplay between the Wav-VAE and the TTS backbone, revealing the counterintuitive finding that superior reconstruction fidelity in the Wav-VAE does not necessarily lead to better overall TTS performance. Code and model weights are released to foster further research within the speech community.
【6】IQRA 2026: Interspeech Challenge on Automatic Assessment Pronunciation for Modern Standard Arabic (MSA)
标题:MQRA 2026:现代标准阿拉伯语(GMA)自动评估发音的过渡挑战
链接:https://arxiv.org/abs/2603.29087
备注:5 pages paper
摘要:我们提出了第二版的IQRA语际挑战,对现代标准阿拉伯语(MSA)的自动发音错误检测和诊断(MDD)的挑战的结果。在上一版的基础上,本次迭代引入了\textbf{Iqra\_Extra\_IS26},这是一个新的真实人类发音错误的语音数据集,补充了现有的训练和评估资源。提交的系统采用了多种方法,包括基于CTC的自监督学习模型、两阶段微调策略以及使用大型音频语言模型。与第一版相比,我们观察到\textbf{0.28 in F1-score}的大幅跃升,这既归功于参与者提出的新颖架构和建模策略,也归功于提供的额外真实的发音错误数据。这些结果表明,阿拉伯语MDD研究的日益成熟,并建立了一个更强大的基础,为今后的工作,在阿拉伯语发音评估。
摘要:We present the findings of the second edition of the IQRA Interspeech Challenge, a challenge on automatic Mispronunciation Detection and Diagnosis (MDD) for Modern Standard Arabic (MSA). Building on the previous edition, this iteration introduces \textbf{Iqra\_Extra\_IS26}, a new dataset of authentic human mispronounced speech, complementing the existing training and evaluation resources. Submitted systems employed a diverse range of approaches, spanning CTC-based self-supervised learning models, two-stage fine-tuning strategies, and using large audio-language models. Compared to the first edition, we observe a substantial jump of \textbf{0.28 in F1-score}, attributable both to novel architectures and modeling strategies proposed by participants and to the additional authentic mispronunciation data made available. These results demonstrate the growing maturity of Arabic MDD research and establish a stronger foundation for future work in Arabic pronunciation assessment.
机器翻译由腾讯交互翻译提供,仅供参考
