今日论文合集:cs.SD语音8篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
标题:韵律ABX:一种测量语音表达中韵律对比度的模糊不可知方法
链接:https://arxiv.org/abs/2604.02102

作者:Haitong Sun,Stephen McIntosh,Kwanghee Choi,Eunjung Yeo,Daisuke Saito,Nobuaki Minematsu
备注:Submitted to Interspeech 2026; 6 pages, 4 figures
摘要:众所周知,自监督语音模型(S3 M)的语音表示对音素对比敏感,但它们对韵律对比的敏感性尚未直接测量。ABX歧视任务已被用来衡量音位对比S3M表示通过最小对。我们引入韵律ABX,这个框架的扩展,以评估韵律对比度只有少数的例子,没有明确的标签。此外,我们构建并发布了英语和日语最小对的数据集,并将其与普通话数据集一起使用,以评估英语重音,日语音高口音和普通话音调的对比。最后,我们表明,模型和层的排名往往保留在几个实验条件下,使其实用的低资源设置。
摘要:Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly measured. The ABX discrimination task has been used to measure phonemic contrast in S3M representations via minimal pairs. We introduce prosodic ABX, an extension of this framework to evaluate prosodic contrast with only a handful of examples and no explicit labels. Also, we build and release a dataset of English and Japanese minimal pairs and use it along with a Mandarin dataset to evaluate contrast in English stress, Japanese pitch accent, and Mandarin tone. Finally, we show that model and layer rankings are often preserved across several experimental conditions, making it practical for low-resource settings.


【2】Woosh: A Sound Effects Foundation Model
标题:Woosh:音效基础模型
链接:https://arxiv.org/abs/2604.01929

作者:Gaëtan Hadjeres,Marc Ferras,Khaled Koutini,Benno Weck,Alexandre Bittar,Thomas Hummel,Zineb Lahrici,Hakim Missoum,Joan Serrà,Yuki Mitsufuji
摘要:音频研究社区依赖于开放生成模型作为构建新方法和建立基线的基础工具。在本报告中,我们介绍了索尼AI公开发布的音效基础模型Woosh,详细介绍了其架构,训练过程以及对其他流行开放模型的评估。针对声音效果进行了优化,我们提供了(1)高质量的音频编码器/解码器模型和(2)用于调节的文本-音频对齐模型,以及(3)文本到音频和(4)视频到音频生成模型。该版本还包括经过提炼的文本到音频和视频到音频模型,允许低资源操作和快速推理。我们对公共和私人数据的评估显示,与现有的开放替代方案(如StableAudio-Open和TangoFlux)相比,每个模块都具有竞争力或更好的性能。推断代码和模型权重可在https://github.com/SonyResearch/Woosh上获得。演示示例可以在https://sonyresearch.github.io/Woosh/上找到。
摘要:The audio research community depends on open generative models as foundational tools for building novel approaches and establishing baselines. In this report, we present Woosh, Sony AI's publicly released sound effect foundation model, detailing its architecture, training process, and an evaluation against other popular open models. Being optimized for sound effects, we provide (1) a high-quality audio encoder/decoder model and (2) a text-audio alignment model for conditioning, together with (3) text-to-audio and (4) video-to-audio generative models. Distilled text-to-audio and video-to-audio models are also included in the release, allowing for low-resource operation and fast inference. Our evaluation on both public and private data shows competitive or better performance for each module when compared to existing open alternatives like StableAudio-Open and TangoFlux. Inference code and model weights are available at https://github.com/SonyResearch/Woosh. Demo samples can be found at https://sonyresearch.github.io/Woosh/.


【3】FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection
标题:FastTurn:统一声学和流语义线索,实现低延迟和稳健的转弯检测
链接:https://arxiv.org/abs/2604.01897

作者:Chengyou Wang,Hongfei Xue,Chunjiang He,Jingbin Hu,Shuiyuan Wang,Bo Wu,Yuyu Ji,Jimeng Zheng,Ruofei Chen,Zhou Zhu,Lei Xie
备注:5 pages, 2 figures
摘要:AudioLLM的最新进展使口语对话系统能够超越基于回合的交互,转向实时全双工通信,其中代理必须在用户仍在说话时决定何时说话,屈服或中断。现有的全双工方法要么依赖于缺乏语义理解的语音活动线索,要么依赖于基于ASR的模块,这会引入延迟并在重叠的语音和噪声下降级。此外,现有的数据集很少捕捉现实的互动动态,限制了评估和部署。为了缓解这个问题,我们提出了\textbf{FastTurn},一个低延迟和鲁棒的转弯检测的统一框架。为了在保持性能的同时提高延迟,FastTurn将流式CTC解码与声学特征相结合,从而在保留语义线索的同时从部分观察中实现早期决策。我们还发布了一个基于真实人类对话的测试集,捕捉真实的转折过渡,重叠语音,反向通道,停顿,音高变化和环境噪声。实验表明,FastTurn实现了更高的决策准确性与较低的中断延迟比代表性的基线,并保持强大的挑战性的声学条件下,证明了其有效性的实际全双工对话系统。
摘要:Recent advances in AudioLLMs have enabled spoken dialogue systems to move beyond turn-based interaction toward real-time full-duplex communication, where the agent must decide when to speak, yield, or interrupt while the user is still talking. Existing full-duplex approaches either rely on voice activity cues, which lack semantic understanding, or on ASR-based modules, which introduce latency and degrade under overlapping speech and noise. Moreover, available datasets rarely capture realistic interaction dynamics, limiting evaluation and deployment. To mitigate the problem, we propose \textbf{FastTurn}, a unified framework for low-latency and robust turn detection. To advance latency while maintaining performance, FastTurn combines streaming CTC decoding with acoustic features, enabling early decisions from partial observations while preserving semantic cues. We also release a test set based on real human dialogue, capturing authentic turn transitions, overlapping speech, backchannels, pauses, pitch variation, and environmental noise. Experiments show FastTurn achieves higher decision accuracy with lower interruption latency than representative baselines and remains robust under challenging acoustic conditions, demonstrating its effectiveness for practical full-duplex dialogue systems.


【4】Acoustic and perceptual differences between standard and accented Chinese speech and their voice clones
标题:标准和重读中文语音及其语音克隆之间的声学和感知差异
链接:https://arxiv.org/abs/2604.01562

作者:Tianle Yang,Chengzhe Sun,Phil Rose,Siwei Lyu
摘要:语音克隆通常是根据整体质量进行评估,但对口音保留及其感知后果知之甚少。我们比较标准和重口音的普通话语音和他们的声音克隆使用相结合的计算和感知设计。嵌入为基础的分析表明,没有可靠的强调标准的差异,跨系统的原始克隆的距离。在感知研究中,克隆被评为更类似于他们的原始标准比口音的扬声器,和可理解性增加从原来的克隆,与口音的语音更大的增益。这些结果表明,口音的变化可以塑造感知身份匹配和语音克隆的可懂度,即使它不是反映在现成的扬声器嵌入距离,他们激励评估扬声器身份保护和口音保护作为可分离的维度。
摘要:Voice cloning is often evaluated in terms of overall quality, but less is known about accent preservation and its perceptual consequences. We compare standard and heavily accented Mandarin speech and their voice clones using a combined computational and perceptual design. Embedding-based analyses show no reliable accented-standard difference in original-clone distances across systems. In the perception study, clones are rated as more similar to their originals for standard than for accented speakers, and intelligibility increases from original to clone, with a larger gain for accented speech. These results show that accent variation can shape perceived identity match and intelligibility in voice cloning even when it is not reflected in an off-the-shelf speaker-embedding distance, and they motivate evaluating speaker identity preservation and accent preservation as separable dimensions.


【5】Evolutionary Multi-Objective Fusion of Deepfake Speech Detectors
标题:Deepfake语音检测器的进化多目标融合
链接:https://arxiv.org/abs/2604.01330

作者:Vojtěch Staněk,Martin Perešíni,Lukáš Sekanina,Anton Firc,Kamil Malinka
备注:Accepted to WCCI CEC 2026
摘要:虽然建立在大型自监督学习(SSL)模型上的Deepfake语音检测器实现了高精度,但采用标准集成融合来进一步增强鲁棒性往往会导致系统规模过大,收益递减。为了解决这个问题,我们提出了一个进化的多目标得分融合框架,共同最大限度地减少检测错误和系统的复杂性。我们探讨了两种编码优化的NSGA-II:二进制编码的检测器选择得分平均和一个实值的计划,优化检测器的加权和的权重。在ASVspoof 5数据集上使用36个基于SSL的检测器进行的实验表明,所获得的Pareto前沿优于简单平均和逻辑回归基线。实值变体实现了2.37%的EER(0.0684 minDCF),并确定了与最先进性能相匹配的配置,同时显著降低了系统复杂性,仅需要一半的参数。我们的方法还提供了一套不同的权衡解决方案,使部署选择,平衡精度和计算成本。
摘要:While deepfake speech detectors built on large self-supervised learning (SSL) models achieve high accuracy, employing standard ensemble fusion to further enhance robustness often results in oversized systems with diminishing returns. To address this, we propose an evolutionary multi-objective score fusion framework that jointly minimizes detection error and system complexity. We explore two encodings optimized by NSGA-II: binary-coded detector selection for score averaging and a real-valued scheme that optimizes detector weights for a weighted sum. Experiments on the ASVspoof 5 dataset with 36 SSL-based detectors show that the obtained Pareto fronts outperform simple averaging and logistic regression baselines. The real-valued variant achieves 2.37% EER (0.0684 minDCF) and identifies configurations that match state-of-the-art performance while significantly reducing system complexity, requiring only half the parameters. Our method also provides a diverse set of trade-off solutions, enabling deployment choices that balance accuracy and computational cost.


【6】Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
标题:结合掩蔽语言建模和跨模式对比学习实现韵律感知的TTC
链接:https://arxiv.org/abs/2604.01247

作者:Kirill Borodin,Vasiliy Kudryavtsev,Maxim Maslov,Nikita Vasiliev,Mikhail Gorodnichev,Grach Mkrtchian
备注:This paper has been submitted to Interspeech 2026 for review
摘要:我们研究了基于扩散的TTS中韵律建模的多阶段预训练。说话者条件双流编码器是用掩蔽语言建模训练的,然后使用混合音素批次进行SigLIP风格的跨模态对比学习,并单独研究了额外的相同音素细化阶段。我们评估内在的文本音频检索和下游合成梯度TTS和潜在的扩散TTS系统。两阶段课程(MLM +混合音素对比学习)在可懂度、说话人相似性和感知测量方面实现了最佳的整体合成质量。虽然相同的音素细化提高韵律检索,它减少了音素歧视和降低合成。这些研究结果表明,嵌入空间指标的改善并不一定转化为更好的生成性能,并强调需要平衡音素歧视和韵律敏感性TTS预训练。
摘要:We investigate multi-stage pretraining for prosody modeling in diffusion-based TTS. A speaker-conditioned dual-stream encoder is trained with masked language modeling followed by SigLIP-style cross-modal contrastive learning using mixed-phoneme batches, with an additional same-phoneme refinement stage studied separately. We evaluate intrinsic text-audio retrieval and downstream synthesis in Grad-TTS and a latent diffusion TTS system. The two-stage curriculum (MLM + mixed-phoneme contrastive learning) achieves the best overall synthesis quality in terms of intelligibility, speaker similarity, and perceptual measures. Although same-phoneme refinement improves prosodic retrieval, it reduces phoneme discrimination and degrades synthesis. These findings indicate that improvements in embedding-space metrics do not necessarily translate to better generative performance and highlight the need to balance phoneme discrimination and prosodic sensitivity in TTS pretraining.


【7】GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
标题:GAP-URGENet:通用语音增强的生成预测融合框架
链接:https://arxiv.org/abs/2604.01832

作者:Xiaobin Rong,Yushi Wang,Zheng Wang,Jing Lu
备注:Awarded 1st place in the URGENT 2026 Challenge (objective phase), accepted by ICASSP 2026
摘要:我们介绍GAP-URGENet,这是一个为ICASSP 2026 URGENT挑战赛的轨道1开发的生成预测融合框架。该系统集成了一个生成分支,该分支在自监督表示域中执行全栈语音恢复,并通过神经声码器重建波形,以及一个预测分支,该分支执行频谱域增强,提供互补线索。来自两个分支的输出由后处理模块融合,后处理模块还执行带宽扩展以生成48 kHz的增强波形,随后下采样到原始采样率。这种生成预测融合提高了鲁棒性和感知质量,在盲测阶段实现了最佳性能,并在客观评估中排名第一。音频示例可在https://xiaobin-rong.github.io/gap-urgenet_demo上获得。
摘要:We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system integrates a generative branch, which performs full-stack speech restoration in a self-supervised representation domain and reconstructs the waveform via a neural vocoder, along with a predictive branch that performs spectrogram-domain enhancement, providing complementary cues. Outputs from both branches are fused by a post-processing module, which also performs bandwidth extension to generate the enhanced waveform at 48 kHz, later downsampled to the original sampling rate. This generative-predictive fusion improves robustness and perceptual quality, achieving top performance in the blind-test phase and ranking 1st in the objective evaluation. Audio examples are available at https://xiaobin-rong.github.io/gap-urgenet_demo.


【8】PhiNet: Speaker Verification with Phonetic Interpretability
标题:PhiNet:具有语音解释性的说话者验证
链接:https://arxiv.org/abs/2604.01590

作者:Yi Ma,Shuai Wang,Tianchi Liu,Haizhou Li
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing. Codes: https://github.com/mmmmayi/PhiNet
摘要:尽管取得了显著的进展,但自动说话人验证(ASV)系统通常缺乏高问责制应用程序所需的透明度。受人类专家如何进行法医说话人比较(FSC)的启发,我们提出了一个具有语音可解释性的说话人验证网络PhiNet,旨在通过利用语音证据进行决策来增强本地和全球的可解释性。对于用户来说,PhiNet提供了详细的语音级别比较,可以手动检查特定于说话者的功能,并促进对验证结果进行更严格的评估。对于开发人员来说,它提供了验证决策背后的显式推理,简化了错误跟踪并为超参数选择提供了信息。在我们的实验中,我们用实际的例子证明了PhiNet的可解释性,包括它在分析不同超参数的影响方面的应用。我们对所提出的可解释性方法进行了定性和定量评估,并评估了多个基准数据集(包括VoxCeleb,SITW和LibriSpeech)的说话人验证性能。结果表明,PhiNet实现了与传统黑盒ASV模型相当的性能,同时为其决策提供了有意义的,可解释的解释,弥合了ASV和法医分析之间的差距。
摘要:Despite remarkable progress, automatic speaker verification (ASV) systems typically lack the transparency required for high-accountability applications. Motivated by how human experts perform forensic speaker comparison (FSC), we propose a speaker verification network with phonetic interpretability, PhiNet, designed to enhance both local and global interpretability by leveraging phonetic evidence in decision-making. For users, PhiNet provides detailed phonetic-level comparisons that enable manual inspection of speaker-specific features and facilitate a more critical evaluation of verification outcomes. For developers, it offers explicit reasoning behind verification decisions, simplifying error tracing and informing hyperparameter selection. In our experiments, we demonstrate PhiNet's interpretability with practical examples, including its application in analyzing the impact of different hyperparameters. We conduct both qualitative and quantitative evaluations of the proposed interpretability methods and assess speaker verification performance across multiple benchmark datasets, including VoxCeleb, SITW, and LibriSpeech. Results show that PhiNet achieves performance comparable to traditional black-box ASV models while offering meaningful, interpretable explanations for its decisions, bridging the gap between ASV and forensic analysis.


eess.AS音频处理


【1】GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
标题:GAP-URGENet:通用语音增强的生成预测融合框架
链接:https://arxiv.org/abs/2604.01832

作者:Xiaobin Rong,Yushi Wang,Zheng Wang,Jing Lu
备注:Awarded 1st place in the URGENT 2026 Challenge (objective phase), accepted by ICASSP 2026
摘要:我们介绍GAP-URGENet,这是一个为ICASSP 2026 URGENT挑战赛的轨道1开发的生成预测融合框架。该系统集成了一个生成分支,该分支在自监督表示域中执行全栈语音恢复,并通过神经声码器重建波形,以及一个预测分支,该分支执行频谱域增强,提供互补线索。来自两个分支的输出由后处理模块融合,后处理模块还执行带宽扩展以生成48 kHz的增强波形,随后下采样到原始采样率。这种生成预测融合提高了鲁棒性和感知质量,在盲测阶段实现了最佳性能,并在客观评估中排名第一。音频示例可在https://xiaobin-rong.github.io/gap-urgenet_demo上获得。
摘要:We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system integrates a generative branch, which performs full-stack speech restoration in a self-supervised representation domain and reconstructs the waveform via a neural vocoder, along with a predictive branch that performs spectrogram-domain enhancement, providing complementary cues. Outputs from both branches are fused by a post-processing module, which also performs bandwidth extension to generate the enhanced waveform at 48 kHz, later downsampled to the original sampling rate. This generative-predictive fusion improves robustness and perceptual quality, achieving top performance in the blind-test phase and ranking 1st in the objective evaluation. Audio examples are available at https://xiaobin-rong.github.io/gap-urgenet_demo.


【2】T5Gemma-TTS Technical Report
标题:T5 Gemma-TTC技术报告
链接:https://arxiv.org/abs/2604.01760

作者:Chihiro Arata,Kiyoshi Kurihara
摘要:自回归神经编解码器语言模型已经显示出强大的zero-shot语音克隆能力,但是仅解码器架构将输入文本视为与不断增长的音频序列竞争位置容量的前缀,从而削弱了长话语上的文本调节。我们提出了T5 Gemma-TTS,编码器-解码器编解码器语言模型,通过在每个解码器层的交叉注意路由双向文本表示来保持持久的文本调节。它建立在T5 Gemma预训练的编码器-解码器骨干(2B编码器+ 2B解码器; 4 B参数)上,继承了丰富的语言知识,无需音素转换,并直接在子字级别处理文本。为了改善持续时间控制,我们在所有26个交叉注意层中引入了进度监控旋转位置嵌入(PM-RoPE),注入归一化的进度信号,帮助解码器跟踪目标语音长度。经过170,000小时的英语、汉语和日语多语言语音训练,T5 Gemma-TTS在日语上的说话者相似度在统计上显著高于XTTSv 2(0.677对0.622;非重叠的95%置信区间)和最高的韩语说话者相似性数值(0.747),尽管韩国人没有被纳入培训,尽管超过XTTSv 2(0.741)的这一幅度在统计学上并不确定。在五个基线中,它也达到了最低的日语字符错误率(0.126),尽管这个排名应该谨慎解释,因为部分置信区间与Kokoro重叠。LibriSpeech上的英语结果应该被视为一个上限估计,因为LibriHeavy是LibriSpeech的超集。使用相同的检查点,在推理时禁用PM-RoPE会导致几乎完全的合成失败:CER从0.129下降到0.982,持续时间准确度从79%下降到46%。代码和重量可在https://github.com/Aratako/T5Gemma-TTS上获得。
摘要:Autoregressive neural codec language models have shown strong zero-shot voice cloning ability, but decoder-only architectures treat input text as a prefix that competes with the growing audio sequence for positional capacity, weakening text conditioning over long utterances. We present T5Gemma-TTS, an encoder-decoder codec language model that maintains persistent text conditioning by routing bidirectional text representations through cross-attention at every decoder layer. Built on the T5Gemma pretrained encoder-decoder backbone (2B encoder + 2B decoder; 4B parameters), it inherits rich linguistic knowledge without phoneme conversion and processes text directly at the subword level. To improve duration control, we introduce Progress-Monitoring Rotary Position Embedding (PM-RoPE) in all 26 cross-attention layers, injecting normalized progress signals that help the decoder track target speech length. Trained on 170,000 hours of multilingual speech in English, Chinese, and Japanese, T5Gemma-TTS achieves a statistically significant speaker-similarity gain on Japanese over XTTSv2 (0.677 vs. 0.622; non-overlapping 95% confidence intervals) and the highest numerical Korean speaker similarity (0.747) despite Korean not being included in training, although this margin over XTTSv2 (0.741) is not statistically conclusive. It also attains the lowest numerical Japanese character error rate among five baselines (0.126), though this ranking should be interpreted cautiously because of partial confidence-interval overlap with Kokoro. English results on LibriSpeech should be viewed as an upper-bound estimate because LibriHeavy is a superset of LibriSpeech. Using the same checkpoint, disabling PM-RoPE at inference causes near-complete synthesis failure: CER degrades from 0.129 to 0.982 and duration accuracy drops from 79% to 46%. Code and weights are available at https://github.com/Aratako/T5Gemma-TTS.


【3】PhiNet: Speaker Verification with Phonetic Interpretability
标题:PhiNet:具有语音解释性的说话者验证
链接:https://arxiv.org/abs/2604.01590

作者:Yi Ma,Shuai Wang,Tianchi Liu,Haizhou Li
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing. Codes: https://github.com/mmmmayi/PhiNet
摘要:尽管取得了显著的进展,但自动说话人验证(ASV)系统通常缺乏高问责制应用程序所需的透明度。受人类专家如何进行法医说话人比较(FSC)的启发,我们提出了一个具有语音可解释性的说话人验证网络PhiNet,旨在通过利用语音证据进行决策来增强本地和全球的可解释性。对于用户来说,PhiNet提供了详细的语音级别比较,可以手动检查特定于说话者的功能,并促进对验证结果进行更严格的评估。对于开发人员来说,它提供了验证决策背后的显式推理,简化了错误跟踪并为超参数选择提供了信息。在我们的实验中,我们用实际的例子证明了PhiNet的可解释性,包括它在分析不同超参数的影响方面的应用。我们对所提出的可解释性方法进行了定性和定量评估,并评估了多个基准数据集(包括VoxCeleb,SITW和LibriSpeech)的说话人验证性能。结果表明,PhiNet实现了与传统黑盒ASV模型相当的性能,同时为其决策提供了有意义的,可解释的解释,弥合了ASV和法医分析之间的差距。
摘要:Despite remarkable progress, automatic speaker verification (ASV) systems typically lack the transparency required for high-accountability applications. Motivated by how human experts perform forensic speaker comparison (FSC), we propose a speaker verification network with phonetic interpretability, PhiNet, designed to enhance both local and global interpretability by leveraging phonetic evidence in decision-making. For users, PhiNet provides detailed phonetic-level comparisons that enable manual inspection of speaker-specific features and facilitate a more critical evaluation of verification outcomes. For developers, it offers explicit reasoning behind verification decisions, simplifying error tracing and informing hyperparameter selection. In our experiments, we demonstrate PhiNet's interpretability with practical examples, including its application in analyzing the impact of different hyperparameters. We conduct both qualitative and quantitative evaluations of the proposed interpretability methods and assess speaker verification performance across multiple benchmark datasets, including VoxCeleb, SITW, and LibriSpeech. Results show that PhiNet achieves performance comparable to traditional black-box ASV models while offering meaningful, interpretable explanations for its decisions, bridging the gap between ASV and forensic analysis.


【4】Robust Pitch Estimation and Tracking for Speakers Based on Subband Encoding and the Generalized Labeled Multi-Bernoulli Filter
标题:基于子带编码和广义标记多伯努里过滤器的稳健扬声器音调估计和跟踪
链接:https://arxiv.org/abs/2604.01541

作者:Shoufeng Lin
摘要:本文提出了一种新的基音估计器和基音跟踪器。我们首先使用听觉滤波器组将声音信号分解成子带,假设人类语音的时频稀疏性。而不是直接选择的子带的数量根据经验,我们提出了一种新的频率覆盖度量来获得的子带的数量和滤波器组的中心频率。然后,子带信号编码的启发下的计算听觉场景分析(CASA)的方法,和归一化的自相关计算的基音估计。为了抑制虚假错误和跟踪说话人身份,利用时间连续性约束和广义标记多伯努利(GLMB)滤波器适应基音跟踪,其中我们使用一种新的基于Ornstein-Uhlenbeck过程的基音状态转换模型,和测量驱动的出生模型自适应新出生的音高目标。各种加性噪声的实验评估表明,所提出的方法相比,在大多数研究的情况下,几个国家的最先进的基音周期估计方法取得了更好的准确性。在混响室中使用真实录音的测试也表明,所提出的方法对混响是鲁棒的。
摘要:This paper proposes a new pitch estimator and a novel pitch tracker for speakers. We first decompose the sound signal into subbands using an auditory filterbank, assuming time-frequency sparsity of human speech. Instead of directly selecting the number of subbands according to experience, we propose a novel frequency coverage metric to derive the number of subbands and the center frequencies of the filterbank. The subband signals are then encoded inspired by the computational auditory scene analysis (CASA) approach, and the normalized autocorrelations are calculated for pitch estimation. To suppress spurious errors and track the speaker identity, the temporal continuity constraint is exploited and a Generalized Labeled Multi-Bernoulli (GLMB) filter is adapted for pitch tracking, where we use a novel pitch state transition model based on the Ornstein-Uhlenbeck process, and the measurement driven birth model for adaptive new births of pitch targets. Experimental evaluations with various additive noises demonstrate that the proposed methods have achieved better accuracy compared with several state-of-the-art pitch estimation methods in most studied scenarios. Tests using real recordings in a reverberant room also show that the proposed method is robust against reverberation.


【5】Validating Computational Markers of Depressive Behavior: Cross-Linguistic Speech-Based Depression Detection with Neurophysiological Validation
标题:验证抑郁行为的计算标记:基于跨语言言语的抑郁检测与神经生理学验证
链接:https://arxiv.org/abs/2604.01533

作者:Fuxiang Tao,Dongwei Li,Shuning Tang,Xuri Ge,Wei Ma,Anna Esposito,Alessandro Vinciarelli
备注:12 pages, 6 figures
摘要:基于语音的抑郁症检测已显示出作为客观诊断工具的前景,但声学标记及其神经生物学基础的跨语言鲁棒性仍有待探索。本研究扩展了最初在意大利语上验证的跨数据多级注意力(CDMA)框架,使用中国普通话数据集与脑电图(EEG)记录来研究这些维度。我们系统地融合阅读语音与自发语音在不同的情绪效价(积极,中性,消极),以调查是否情绪唤醒是一个更关键的因素比效价极性在提高检测性能的讲话。此外,我们建立了第一个基于语音的抑郁症模型的神经生理学验证相关的预测与神经振荡模式在情绪化的面孔处理。我们的研究结果表明,CDMA框架具有很强的跨语言通用性,在中国数据集上实现了最先进的性能(F1得分高达89.6%),与之前的意大利验证相当。关键的是,情绪化的言语(包括积极和消极的)明显优于中性言语。积极和消极任务之间的这种可比性支持了情绪唤起假说。最重要的是,EEG分析揭示了该模型的语音衍生抑郁估计与神经振荡模式(θ和α波段)之间的显著相关性,表明与抑郁症情绪失调的既定神经标志物一致。这种对齐,结合该模型的跨语言的鲁棒性,不仅支持CDMA框架的方法是一个普遍适用的和神经生物学验证的策略,但也建立了一个新的范式计算心理健康模型的神经生理学验证。
摘要:Speech-based depression detection has shown promise as an objective diagnostic tool, yet the cross-linguistic robustness of acoustic markers and their neurobiological underpinnings remain underexplored. This study extends Cross-Data Multilevel Attention (CDMA) framework, initially validated on Italian, to investigate these dimensions using a Chinese Mandarin dataset with Electroencephalography (EEG) recordings. We systematically fuse read speech with spontaneous speech across different emotional valences (positive, neutral, negative) to investigate whether emotional arousal is a more critical factor than valence polarity in enhancing detection performance in speech. Additionally, we establish the first neurophysiological validation for a speech-based depression model by correlating its predictions with neural oscillatory patterns during emotional face processing. Our results demonstrate strong cross-linguistic generalizability of the CDMA framework, achieving state-of-the-art performance (F1-score up to 89.6%) on the Chinese dataset, which is comparable to the previous Italian validation. Critically, emotionally valenced speech (both positive and negative) significantly outperformed neutral speech. This comparable performance between positive and negative tasks supports the emotional arousal hypothesis. Most importantly, EEG analysis revealed significant correlations between the model's speech-derived depression estimates and neural oscillatory patterns (theta and alpha bands), demonstrating alignment with established neural markers of emotional dysregulation in depression. This alignment, combined with the model's cross-linguistic robustness, not only supports that the CDMA framework's approach is a universally applicable and neurobiologically validated strategy but also establishes a novel paradigm for the neurophysiological validation of computational mental health models.


【6】Reverberation-Robust Localization of Speakers Using Distinct Speech Onsets and Multi-channel Cross-Correlations
标题:利用不同语音起始点和多通道互相关的混响鲁棒说话人定位
链接:https://arxiv.org/abs/2604.01524

作者:Shoufeng Lin
摘要:许多说话人定位方法可以在文献中找到。然而,强混响环境下的说话人定位仍然是现实应用中的一个主要挑战。本文提出了两种算法定位扬声器使用麦克风阵列记录的混响声音。为了分离并发的扬声器,第一种算法通过听觉滤波器组将麦克风信号谱时分解成子带。为了抑制混响,我们提出了一种新的语音起始点检测方法来自语音信号和脉冲响应模型,并进一步提出制定在每个子带中的编码语音起始点的多通道互相关系数(MCCC)。子带的结果相结合,以估计的到达方向(DOA)的扬声器。第二种算法扩展了广义互相关相位变换(GCC-PHAT)方法,利用多个麦克风的冗余信息来解决混响问题。所提出的方法已评估不利条件下,不仅使用模拟信号(混响时间$T_{60}$的高达$1$s),但也记录在一个真正的混响室($T_{60} \approximately 0.65$s)。实验结果表明,与现有的定位方法相比,本文提出的方法在混响存在的情况下,能够可靠地定位静止和运动的说话人。
摘要:Many speaker localization methods can be found in the literature. However, speaker localization under strong reverberation still remains a major challenge in the real-world applications. This paper proposes two algorithms for localizing speakers using microphone array recordings of reverberated sounds. To separate concurrent speakers, the first algorithm decomposes microphone signals spectrotemporally into subbands via an auditory filterbank. To suppress reverberation, we propose a novel speech onset detection approach derived from the speech signal and impulse response models, and further propose to formulate the multi-channel cross-correlation coefficient (MCCC) of encoded speech onsets in each subband. The subband results are combined to estimate the directions-of-arrival (DOAs) of speakers. The second algorithm extends the generalized cross-correlation - phase transform (GCC-PHAT) method by using redundant information of multiple microphones to address the reverberation problem. The proposed methods have been evaluated under adverse conditions using not only simulated signals (reverberation time $T_{60}$ of up to $1$s) but also recordings in a real reverberant room ($T_{60} \approx 0.65$s). Comparing with some state-of-the-art localization methods, experimental results confirm that the proposed methods can reliably locate static and moving speakers, in presence of reverberation.


【7】Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
标题:韵律ABX:一种测量语音表达中韵律对比度的模糊不可知方法
链接:https://arxiv.org/abs/2604.02102

作者:Haitong Sun,Stephen McIntosh,Kwanghee Choi,Eunjung Yeo,Daisuke Saito,Nobuaki Minematsu
备注:Submitted to Interspeech 2026; 6 pages, 4 figures
摘要:自监督语音模型(S3M)的语音表示是已知的音素对比敏感,但他们的韵律对比的敏感性还没有被直接测量。ABX歧视任务已被用来衡量音位对比S3M表示通过最小对。我们引入韵律ABX,这个框架的扩展,以评估韵律对比度只有少数的例子,没有明确的标签。此外,我们构建并发布了英语和日语最小对的数据集,并将其与普通话数据集一起使用,以评估英语重音,日语音高口音和普通话音调的对比。最后,我们表明,模型和层的排名往往保留在几个实验条件下,使其实用的低资源设置。
摘要:Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly measured. The ABX discrimination task has been used to measure phonemic contrast in S3M representations via minimal pairs. We introduce prosodic ABX, an extension of this framework to evaluate prosodic contrast with only a handful of examples and no explicit labels. Also, we build and release a dataset of English and Japanese minimal pairs and use it along with a Mandarin dataset to evaluate contrast in English stress, Japanese pitch accent, and Mandarin tone. Finally, we show that model and layer rankings are often preserved across several experimental conditions, making it practical for low-resource settings.


【8】Tracking the emergence of linguistic structure in self-supervised models learning from speech
标题:跟踪语音学习的自我监督模型中语言结构的出现
链接:https://arxiv.org/abs/2604.02043

作者:Marianne de Heer Kloots,Martijn Bentum,Hosein Mohebbi,Charlotte Pouw,Gaofei Shen,Willem Zuidema
摘要:自监督语音模型学习口语的有效表示,这已被证明反映了语言结构的各个方面。但是,这种结构何时出现在模型训练中呢?我们研究了广泛的语言结构的编码,跨层和中间检查点的六个Wav2Vec2和HuBERT模型的荷兰语口语训练。我们发现,不同层次的语言结构显示出明显不同的分层模式以及学习轨迹,这可以部分地解释为他们的程度的差异,从声学信号的抽象和时间尺度的信息从输入集成。此外,我们发现,预训练目标的定义水平强烈影响了语言结构的分层组织和学习轨迹,高阶预测任务(即迭代细化的伪标签)引起了更大的并行性。
摘要:Self-supervised speech models learn effective representations of spoken language, which have been shown to reflect various aspects of linguistic structure. But when does such structure emerge in model training? We study the encoding of a wide range of linguistic structures, across layers and intermediate checkpoints of six Wav2Vec2 and HuBERT models trained on spoken Dutch. We find that different levels of linguistic structure show notably distinct layerwise patterns as well as learning trajectories, which can partially be explained by differences in their degree of abstraction from the acoustic signal and the timescale at which information from the input is integrated. Moreover, we find that the level at which pre-training objectives are defined strongly affects both the layerwise organization and the learning trajectories of linguistic structures, with greater parallelism induced by higher-order prediction tasks (i.e. iteratively refined pseudo-labels).


【9】Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
标题:结合掩蔽语言建模和跨模式对比学习实现韵律感知的TTC
链接:https://arxiv.org/abs/2604.01247

作者:Kirill Borodin,Vasiliy Kudryavtsev,Maxim Maslov,Nikita Vasiliev,Mikhail Gorodnichev,Grach Mkrtchian
备注:This paper has been submitted to Interspeech 2026 for review
摘要:我们研究了基于扩散的TTS中韵律建模的多阶段预训练。说话者条件双流编码器是用掩蔽语言建模训练的,然后使用混合音素批次进行SigLIP风格的跨模态对比学习,并单独研究了额外的相同音素细化阶段。我们评估内在的文本音频检索和下游合成梯度TTS和潜在的扩散TTS系统。两阶段课程(MLM +混合音素对比学习)在可懂度、说话人相似性和感知测量方面实现了最佳的整体合成质量。虽然相同的音素细化提高韵律检索,它减少了音素歧视和降低合成。这些研究结果表明,嵌入空间指标的改善并不一定转化为更好的生成性能,并强调需要平衡音素歧视和韵律敏感性TTS预训练。
摘要:We investigate multi-stage pretraining for prosody modeling in diffusion-based TTS. A speaker-conditioned dual-stream encoder is trained with masked language modeling followed by SigLIP-style cross-modal contrastive learning using mixed-phoneme batches, with an additional same-phoneme refinement stage studied separately. We evaluate intrinsic text-audio retrieval and downstream synthesis in Grad-TTS and a latent diffusion TTS system. The two-stage curriculum (MLM + mixed-phoneme contrastive learning) achieves the best overall synthesis quality in terms of intelligibility, speaker similarity, and perceptual measures. Although same-phoneme refinement improves prosodic retrieval, it reduces phoneme discrimination and degrades synthesis. These findings indicate that improvements in embedding-space metrics do not necessarily translate to better generative performance and highlight the need to balance phoneme discrimination and prosodic sensitivity in TTS pretraining.


机器翻译由腾讯交互翻译提供,仅供参考