今日论文合集:cs.SD语音10篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Low-Resource Guidance for Controllable Latent Audio Diffusion
标题:用于可控潜在音频扩散的低资源指导
链接:https://arxiv.org/abs/2603.04366

作者:Zachary Novack,Zack Zukowski,CJ Carr,Julian Parker,Zach Evans,Josiah Taylor,Taylor Berg-Kirkpatrick,Julian McAuley,Jordi Pons
备注:Accepted at ICASSP 2026
摘要:生成音频需要细粒度的可控输出,但大多数现有方法需要对特定控件或推理时间控件进行模型再训练(\textit{e.g.},引导),这也可能是计算上的要求。通过研究现有的基于指导的控制的瓶颈,特别是由于解码器反向传播的每步成本高,我们引入了一个基于指导的方法,通过选择性TFG和潜在控制头(LatCH),这使得控制潜在的音频扩散模型具有较低的计算开销。LatCH直接在潜在空间中操作,避免了昂贵的解码器步骤,并且需要最少的训练资源(7 M参数和大约4小时的训练)。Stable Audio Open的实验证明了对强度,音高和节拍(以及它们的组合)的有效控制,同时保持生成质量。我们的方法平衡了精度和音频保真度,计算成本远低于标准的端到端指导。演示示例可以在https://zacharynovack.github.io/latch/latch.html上找到。
摘要:Generative audio requires fine-grained controllable outputs, yet most existing methods require model retraining on specific controls or inference-time controls (\textit{e.g.}, guidance) that can also be computationally demanding. By examining the bottlenecks of existing guidance-based controls, in particular their high cost-per-step due to decoder backpropagation, we introduce a guidance-based approach through selective TFG and Latent-Control Heads (LatCHs), which enables controlling latent audio diffusion models with low computational overhead. LatCHs operate directly in latent space, avoiding the expensive decoder step, and requiring minimal training resources (7M parameters and $\approx$ 4 hours of training). Experiments with Stable Audio Open demonstrate effective control over intensity, pitch, and beats (and a combination of those) while maintaining generation quality. Our method balances precision and audio fidelity with far lower computational costs than standard end-to-end guidance. Demo examples can be found at https://zacharynovack.github.io/latch/latch.html.


【2】LabelBuddy: An Open Source Music and Audio Language Annotation Tagging Tool Using AI Assistance
标题:LabelBuddy:使用人工智能协助的开源音乐和音频语言注释标记工具
链接:https://arxiv.org/abs/2603.04293

作者:Ioannis Prokopiou,Ioannis Sina,Agisilaos Kounelis,Pantelis Vikatos,Themos Stafylakis
备注:Accepted at NLP4MusA 2026 (4th Workshop on NLP for Music and Audio)
摘要:机器学习(ML),大型音频语言模型(LALM)和音乐信息检索(MIR)中的自主AI代理的进步需要从静态标记转向丰富的,与人类对齐的表示学习。然而,能够捕捉音频注释的主观细微差别的开源基础设施的稀缺仍然是一个关键的瓶颈。本文介绍了\textbf{LabelBuddy},这是一个开源的协作自动标记音频注释工具,旨在弥合人类意图和机器理解之间的差距。与静态工具不同,它通过容器化后端从推理中提取接口,允许用户插入自定义模型以进行AI辅助的预注释。我们描述了系统架构,它支持多用户共识,容器化的模型隔离,以及扩展代理和LALM的路线图。代码可在https://github.com/GiannisProkopiou/gsoc2022-Label-buddy上获得。
摘要:The advancement of Machine learning (ML), Large Audio Language Models (LALMs), and autonomous AI agents in Music Information Retrieval (MIR) necessitates a shift from static tagging to rich, human-aligned representation learning. However, the scarcity of open-source infrastructure capable of capturing the subjective nuances of audio annotation remains a critical bottleneck. This paper introduces \textbf{LabelBuddy}, an open-source collaborative auto-tagging audio annotation tool designed to bridge the gap between human intent and machine understanding. Unlike static tools, it decouples the interface from inference via containerized backends, allowing users to plug in custom models for AI-assisted pre-annotation. We describe the system architecture, which supports multi-user consensus, containerized model isolation, and a roadmap for extending agents and LALMs. Code available at https://github.com/GiannisProkopiou/gsoc2022-Label-buddy.


【3】ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
标题:ZeSTA:采用域条件训练的Zero-ShotTTC增强,实现数据高效的个性化语音合成
链接:https://arxiv.org/abs/2603.04219

作者:Youngwon Choi,Jinwoo Oh,Hwayeon Kim,Hyeonyu Kim
备注:6 pages, submitted to INTERSPEECH 2026
摘要:我们调查使用zero-shot文本到语音(TTS)作为低资源个性化语音合成的数据增强源。虽然合成增强可以提供语言丰富和语音多样的语音,但天真地将大量合成语音与有限的真实录音混合通常会导致微调期间扬声器相似性降低。为了解决这个问题,我们提出了ZeSTA,一个简单的域条件训练框架,通过轻量级域嵌入区分真实和合成语音,结合真实数据过采样,在极其有限的目标数据下稳定自适应,而无需修改基础架构。LibriTTS和内部数据集上的实验表明,我们的方法在保持可懂度和感知质量的同时,提高了说话人的相似性。
摘要:We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide linguistically rich and phonetically diverse speech, naively mixing large amounts of synthetic speech with limited real recordings often leads to speaker similarity degradation during fine-tuning. To address this issue, we propose ZeSTA, a simple domain-conditioned training framework that distinguishes real and synthetic speech via a lightweight domain embedding, combined with real-data oversampling to stabilize adaptation under extremely limited target data, without modifying the base architecture. Experiments on LibriTTS and an in-house dataset with two ZS-TTS sources demonstrate that our approach improves speaker similarity over naive synthetic augmentation while preserving intelligibility and perceptual quality.


【4】FastWave: Optimized Diffusion Model for Audio Super-Resolution
标题:FastWave:音频超分辨率的优化扩散模型
链接:https://arxiv.org/abs/2603.04122

作者:Nikita Kuznetsov,Maksim Kaledin
摘要:音频超分辨率是一组旨在对给定信号进行高质量估计的技术,就好像它将以更高的采样率进行采样一样。在建议的方法中,有扩散和流动模型(被认为是较慢的),生成对抗网络(被认为是较快的),但这两种方法目前都是由高参数网络提出的,需要很高的训练和推理计算成本。我们通过重新考虑扩散模型训练的最新进展并将其应用于从任何到48 kHz采样率的超分辨率,提出了这两个问题的解决方案。我们的方法显示出比NU-Wave 2更好的结果,并且与最先进的模型相当。我们的FastWave模型具有大约50 GFLOPs的计算复杂度和130万个参数,可以用更少的资源进行训练,并且比大多数最近提出的基于扩散和流的解决方案要快得多。该守则已公开提供。
摘要:Audio Super-Resolution is a set of techniques aimed at high-quality estimation of the given signal as if it would be sampled with higher sample rate. Among suggested methods there are diffusion and flow models (which are considered slower), generative adversarial networks (which are considered faster), however both approaches are currently presented by high-parametric networks, requiring high computational costs both for training and inference. We propose a solution to both these problems by re-considering the recent advances in the training of diffusion models and applying them to super-resolution from any to 48 kHz sample rate. Our approach shows better results than NU-Wave 2 and is comparable to state-of-the-art models. Our model called FastWave has around 50 GFLOPs of computational complexity and 1.3 M parameters and can be trained with less resources and significantly faster than the majority of recently proposed diffusion- and flow-based solutions. The code has been made publicly available.


【5】Multi-Stage Music Source Restoration with BandSplit-RoFormer Separation and HiFi++ GAN
标题:利用BandSplit-RoFormer Separation和HiFi++ GAN进行多阶段音乐源恢复
链接:https://arxiv.org/abs/2603.04032

作者:Tobias Morocutti,Emmanouil Karystinaios,Jonathan Greif,Gerhard Widmer
备注:ICASSP 2026 Music Source Restoration (MSR) Challenge
摘要:音乐源恢复(MSR)的目标是恢复原始的,未经处理的乐器源于完全混合和掌握的音频,其中生产效果和分布文物违反常见的线性混合假设。本技术报告介绍了CP-JKU团队为2025年MSR ICASSP挑战赛设计的系统。我们的方法分解MSR分离和恢复。首先,单个BandSplit-RoFormer分离器预测八个阀杆加上一个辅助阀杆,并通过三阶段课程进行培训,从4阀杆热启动微调(使用LoRA)到通过头部扩展的8阀杆扩展。其次,我们应用HiFi++ GAN波形恢复器,该波形恢复器被训练为通才,然后被专业化为八个特定于乐器的专家。
摘要:Music Source Restoration (MSR) targets recovery of original, unprocessed instrument stems from fully mixed and mastered audio, where production effects and distribution artifacts violate common linear-mixture assumptions. This technical report presents the CP-JKU team's system for the MSR ICASSP Challenge 2025. Our approach decomposes MSR into separation and restoration. First, a single BandSplit-RoFormer separator predicts eight stems plus an auxiliary other stem, and is trained with a three-stage curriculum that progresses from 4-stem warm-start fine-tuning (with LoRA) to 8-stem extension via head expansion. Second, we apply a HiFi++ GAN waveform restorer trained as a generalist and then specialized into eight instrument-specific experts.


【6】A Sensitivity Analysis of Multi-Event Audio Grounding in Audio LLMs
标题:音频LLM中多事件音频接地的灵敏度分析
链接:https://arxiv.org/abs/2603.03855

作者:Taehan Lee,Jaehan Jung,Hyukjun Lee
备注:6 pages, Submitted to Interspeech 2026
摘要:音频LLM已经显示出很强的理解音频样本的能力,但它们在复杂声学场景中的可靠性仍然没有得到充分的探索。与以前的工作局限于小规模或较少控制的查询建设,我们提出了一个大规模的评估事件接地和假警报的听觉场景的复杂性增加。使用71 K AudioCapsV 2剪辑,我们提取归一化(源,属性)事件并构建两种查询类型:用于地面实况检测的当前事件查询和用于探测幻觉的缺席事件查询,在音频对齐的文本嵌入空间中使用相似性过滤负采样。我们评估了四个SOTA音频LLM,每个模型有12个提示变量超过500 K是/否查询。在所有模型中,增加事件计数始终会降低真阳性率并提高假阳性率,而提示则会在两者之间产生强烈的权衡。我们的置信度分析表明,模型在多事件音频上变得更加不确定,显示出改进的空间。
摘要:Audio LLMs have shown a strong ability to understand audio samples, yet their reliability in complex acoustic scenes remains under-explored. Unlike prior work limited to small scale or less controlled query construction, we present a large-scale evaluation of event grounding and false alarms as auditory scene complexity increases. Using 71K AudioCapsV2 clips, we extract normalized (source, attribute) events and build two query types: present-event queries for ground-truth detection and absent-event queries to probe hallucinations, using similarity-filtered negative sampling in an audio-aligned text embedding space. We evaluate four SOTA Audio LLMs with 12 prompt variants over 500K yes/no queries per model. Across models, increasing event count consistently lowers true-positive rate and raises false-positive rate, while prompts induce a strong trade-off between the two. Our confidence analysis shows that models become more uncertain on multi-event audio, revealing room for improvement.


【7】Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
标题:具有稀疏模式对齐和视觉单元引导细化的鲁棒的基于LLM的视听语音识别
链接:https://arxiv.org/abs/2603.03811

作者:Fei Su,Cancan Li,Juan Liu,Wei Ju,Hongbin Suo,Ming Li
备注:submitted to Interspeech 2026
摘要:视听语音识别(AVSR)集成了声学和视觉信息,以增强在不利声学条件下的鲁棒性。大语言模型(LLM)的最新进展已经产生了有竞争力的自动语音识别性能,并显示出AVSR的有效性。然而,先前的方法独立地投射音频和视觉特征或应用浅融合,限制了跨模态对准和互补交换,同时增加了LLM的计算负荷。为了解决这个问题,我们提出了AVUR-LLM,一个基于LLM的视听语音识别通过稀疏模态对齐和视觉单元引导的细化。LRS 3上的实验证明了AVSR的最新结果。在0 dB SNR的加性噪声条件下,它比基线系统实现了37%的相对改善。
摘要:Audio-Visual Speech Recognition (AVSR) integrates acoustic and visual information to enhance robustness in adverse acoustic conditions. Recent advances in Large Language Models (LLMs) have yielded competitive automatic speech recognition performance and shown effectiveness for AVSR. However, prior approaches project audio and visual features independently or apply shallow fusion, limiting cross-modal alignment and complementary exchange while increasing the LLM's computational load. To address this, we propose AVUR-LLM, an LLM-based Audio-Visual Speech Recognition via Sparse Modality Alignment and Visual Unit-Guided Refinement. Experiments on LRS3 demonstrate state-of-the-art results for AVSR. Under additive-noise conditions at 0 dB SNR, it achieves 37% relative improvement over the baseline system.


【8】ACES: Accent Subspaces for Coupling, Explanations, and Stress-Testing in Automatic Speech Recognition
标题:ACES:自动语音识别中用于耦合、简化和压力测试的口音子空间
链接:https://arxiv.org/abs/2603.03359

作者:Swapnil Parekh
摘要:ASR系统在不同口音之间表现出持续的性能差异,但这些差异背后的内部机制仍然知之甚少。我们介绍ACES,一个代表为中心的审计,提取口音歧视子空间,并使用它们来探测模型的脆弱性和差距。分析Wav 2 Vec 2-base的五个英语口音,我们发现口音信息集中在一个低维的早期层子空间(层3,k=8)。投影幅度与每话语WER(r=0.26),至关重要的是,子空间约束扰动产生更强的耦合之间的表示移位和退化(r=0.32)比随机子空间控制(r=0.15)。最后,这个子空间的线性衰减,但并没有减少差距,并略有abrases it. We的研究结果表明,口音相关的功能深深纠缠在一起的发音关键线索,定位口音子空间作为重要的诊断工具,而不是简单的“擦除”公平杠杆。
摘要:ASR systems exhibit persistent performance disparities across accents, yet the internal mechanisms underlying these gaps remain poorly understood. We introduce ACES, a representation-centric audit that extracts accent-discriminative subspaces and uses them to probe model fragility and disparity. Analyzing Wav2Vec2-base with five English accents, we find that accent information concentrates in a low-dimensional early-layer subspace (layer 3, k=8). Projection magnitude correlates with per-utterance WER (r=0.26), and crucially, subspace-constrained perturbations yield stronger coupling between representation shift and degradation (r=0.32) than random-subspace controls (r=0.15). Finally, linear attenuation of this subspace however does not reduce disparity and slightly worsens it. Our findings suggest that accent-relevant features are deeply entangled with recognition-critical cues, positioning accent subspaces as vital diagnostic tools rather than simple "erasure" levers for fairness.


【9】FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
标题:FlowW2N:通过流匹配的耳语到正常语音转换
链接:https://arxiv.org/abs/2603.04296

作者:Fabian Ritter-Gutierrez,Md Asif Jalal,Pablo Peso Parada,Karthikeyan Saravanan,Yusun Shul,Minseung Kim,Gun-Woo Lee,Han-Gil Moon
备注:Submitted to Interspeech 2026
摘要:耳语到正常(W2 N)语音转换旨在从耳语输入重建丢失的发声,同时保留内容和说话人身份。这项任务是具有挑战性的,由于耳语和有声录音之间的时间错位和配对数据的缺乏。我们提出了FlowW 2N,这是一种条件流匹配方法,它只在合成的、时间对齐的耳语正常对和域不变特征的条件上进行训练。我们利用高级ASR嵌入,在合成语音和真实耳语之间表现出很强的不变性,尽管在训练过程中从未观察到,但仍然可以推广到真实耳语。我们验证了这种不变性跨ASR层,并提出了一个选择标准,优化内容的信息性和跨域的不变性。我们的方法在CHAINS和wTIMIT数据集上实现了SOTA可理解性,相对于先前的工作,将单词错误率降低了26-46%,同时仅使用10个推理步骤,并且不需要真正的配对数据。
摘要:Whispered-to-normal (W2N) speech conversion aims to reconstruct missing phonation from whispered input while preserving content and speaker identity. This task is challenging due to temporal misalignment between whisper and voiced recordings and lack of paired data. We propose FlowW2N, a conditional flow matching approach that trains exclusively on synthetic, time-aligned whisper-normal pairs and conditions on domain-invariant features. We exploit high-level ASR embeddings that exhibits strong invariance between synthetic and real whispered speech, enabling generalization to real whispers despite never observing it during training. We verify this invariance across ASR layers and propose a selection criterion optimizing content informativeness and cross-domain invariance. Our method achieves SOTA intelligibility on the CHAINS and wTIMIT datasets, reducing Word Error Rate by 26-46% relative to prior work while using only 10 steps at inference and requiring no real paired data.


【10】Automated Measurement of Geniohyoid Muscle Thickness During Speech Using Deep Learning and Ultrasound
标题:使用深度学习和超声自动测量语音期间的膝舌骨肌厚度
链接:https://arxiv.org/abs/2603.03350

作者:Alisher Myrgyyassov,Bruce Xiao Wang,Yu Sun,Shuming Huang,Zhen Song,Min Ney Wong,Yongping Zheng
备注:6 pages, including references and acknowledgements. Submitted to Interspeech 2026
摘要:在讲话期间从超声手动测量肌肉形态是耗时的,并且限制了大规模研究。我们提出了SMMA,这是一个完全自动化的框架,它将深度学习分割与基于神经元的厚度量化相结合,以分析颏舌骨(GH)肌肉动力学。验证证明了接近人类水平的准确度(Dice = 0.9037,MAE = 0.53 mm,r = 0.901)。应用于粤语元音产生(N = 11)揭示了系统模式:/a:/显示显著更大的GH厚度(7.29 mm)比/i:/(5.95 mm,p < 0.001,Cohen's d > 1.3),表明在产生/a:/比/i:/期间更大的GH激活,与其在下颌凹陷中的作用一致。性别差异(男性高5-8%)反映了解剖尺度。SMMA实现了专家验证的准确性,同时消除了手动注释的需要,从而能够对言语运动控制进行可扩展的调查,并对言语和吞咽障碍进行客观评估。
摘要:Manual measurement of muscle morphology from ultrasound during speech is time-consuming and limits large-scale studies. We present SMMA, a fully automated framework that combines deep-learning segmentation with skeleton-based thickness quantification to analyze geniohyoid (GH) muscle dynamics. Validation demonstrates near-human-level accuracy (Dice = 0.9037, MAE = 0.53 mm, r = 0.901). Application to Cantonese vowel production (N = 11) reveals systematic patterns: /a:/ shows significantly greater GH thickness (7.29 mm) than /i:/ (5.95 mm, p < 0.001, Cohen's d > 1.3), suggesting greater GH activation during production of /a:/ than /i:/, consistent with its role in mandibular depression. Sex differences (5-8% greater in males) reflect anatomical scaling. SMMA achieves expert-validated accuracy while eliminating the need for manual annotation, enabling scalable investigations of speech motor control and objective assessment of speech and swallowing disorders.


eess.AS音频处理


【1】FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
标题:FlowW2N:通过流匹配的耳语到正常语音转换
链接:https://arxiv.org/abs/2603.04296

作者:Fabian Ritter-Gutierrez,Md Asif Jalal,Pablo Peso Parada,Karthikeyan Saravanan,Yusun Shul,Minseung Kim,Gun-Woo Lee,Han-Gil Moon
备注:Submitted to Interspeech 2026
摘要:耳语到正常(W2 N)语音转换旨在从耳语输入重建丢失的发声,同时保留内容和说话人身份。这项任务是具有挑战性的,由于耳语和有声录音之间的时间错位和配对数据的缺乏。我们提出了FlowW 2N,这是一种条件流匹配方法,它只在合成的、时间对齐的耳语正常对和域不变特征的条件上进行训练。我们利用高级ASR嵌入,在合成语音和真实耳语之间表现出很强的不变性,尽管在训练过程中从未观察到,但仍然可以推广到真实耳语。我们验证了这种不变性跨ASR层,并提出了一个选择标准,优化内容的信息性和跨域的不变性。我们的方法在CHAINS和wTIMIT数据集上实现了SOTA可理解性,相对于先前的工作,将单词错误率降低了26-46%,同时仅使用10个推理步骤,并且不需要真正的配对数据。
摘要:Whispered-to-normal (W2N) speech conversion aims to reconstruct missing phonation from whispered input while preserving content and speaker identity. This task is challenging due to temporal misalignment between whisper and voiced recordings and lack of paired data. We propose FlowW2N, a conditional flow matching approach that trains exclusively on synthetic, time-aligned whisper-normal pairs and conditions on domain-invariant features. We exploit high-level ASR embeddings that exhibits strong invariance between synthetic and real whispered speech, enabling generalization to real whispers despite never observing it during training. We verify this invariance across ASR layers and propose a selection criterion optimizing content informativeness and cross-domain invariance. Our method achieves SOTA intelligibility on the CHAINS and wTIMIT datasets, reducing Word Error Rate by 26-46% relative to prior work while using only 10 steps at inference and requiring no real paired data.


【2】Cyclostationarity Analysis as a Complement to Self-Supervised Representations for Speech Deepfake Detection
标题:循环平稳性分析作为语音深度伪造检测的自我监督表示的补充
链接:https://arxiv.org/abs/2603.03921

作者:Cemal Hanilçi,Md Sahidullah,Tomi Kinnunen
备注:submitted to IEEE Transactions on Audio, Speech and Language Processing
摘要:语音深度伪造检测(SDD)对于保持对语音驱动技术和数字媒体的信任至关重要。尽管最近的SDD系统越来越依赖于捕获丰富上下文信息的自监督学习(SSL)表示,但互补信号驱动的声学特征对于建模语音的细粒度结构属性仍然很重要。大多数现有的声学前端都基于时频表示,没有充分利用语音信号固有的高阶谱依赖性。我们介绍了一个循环平稳启发的声音特征提取框架SDD的基础上的谱相关密度(SCD)。所提出的特征通过捕获频率分量之间的频谱相关性来对语音中的周期性统计结构进行建模。特别是,我们提出了时间结构的SCD功能,其特征在于随着时间的推移频谱和循环频率分量的演变。使用多个对策架构,包括卷积神经网络,基于SSL的嵌入系统,和混合融合模型的有效性和互补性的建议功能进行评估。在ASVspoof 2019 LA、ASVspoof 2021 DF和ASVspoof 5上的实验表明,基于SCD的特征为SSL嵌入和传统声学表示提供了互补的判别信息。特别是,SSL和SCD嵌入的融合将ASVspoof 2019 LA的相等错误率从8.28美元降低到0.98美元,并在具有挑战性的ASVspoof 5数据集上产生一致的改进。结果突出了循环平稳信号分析作为语音深度伪造检测的理论基础和有效前端。
摘要:Speech deepfake detection (SDD) is essential for maintaining trust in voice-driven technologies and digital media. Although recent SDD systems increasingly rely on self-supervised learning (SSL) representations that capture rich contextual information, complementary signal-driven acoustic features remain important for modeling fine-grained structural properties of speech. Most existing acoustic front ends are based on time-frequency representations, which do not fully exploit higher-order spectral dependencies inherent in speech signals. We introduce a cyclostationarity-inspired acoustic feature extraction framework for SDD based on spectral correlation density (SCD). The proposed features model periodic statistical structures in speech by capturing spectral correlations between frequency components. In particular, we propose temporally structured SCD features that characterize the evolution of spectral and cyclic-frequency components over time. The effectiveness and complementarity of the proposed features are evaluated using multiple countermeasure architectures, including convolutional neural networks, SSL-based embedding systems, and hybrid fusion models. Experiments on ASVspoof 2019 LA, ASVspoof 2021 DF, and ASVspoof 5 demonstrate that SCD-based features provide complementary discriminative information to SSL embeddings and conventional acoustic representations. In particular, fusion of SSL and SCD embeddings reduces the equal error rate on ASVspoof 2019 LA from $8.28\%$ to $0.98\%$, and yields consistent improvements on the challenging ASVspoof 5 dataset. The results highlight cyclostationary signal analysis as a theoretically grounded and effective front end for speech deepfake detection.


【3】The PARLO Dementia Corpus: A German Multi-Center Resource for Alzheimer's Disease
标题:PARLO痴呆症谱系:德国阿尔茨海默病多中心资源
链接:https://arxiv.org/abs/2603.03471

作者:Franziska Braun,Christopher Witzl,Florian Hönig,Elmar Nöth,Tobias Bocklet,Korbinian Riedhammer
备注:Accepted at LREC 2026
摘要:阿尔茨海默病(AD)的早期和可获得的检测仍然是一个重大挑战,因为目前的诊断方法往往依赖于昂贵的和侵入性的生物标志物。语音和语言分析已经成为一种有前途的非侵入性和可扩展的方法来检测认知障碍,但这一领域的研究受到缺乏公开数据集的阻碍,特别是对于英语以外的语言。本文介绍了PARLO痴呆语料库(PDC),一个新的多中心,临床验证的德国资源AD收集在德国的九个学术记忆诊所。该数据集包括AD相关轻度认知障碍和轻度至中度痴呆患者以及认知健康对照的语音记录。使用标准化测试电池的八个神经心理任务,包括对抗命名,言语流畅性,单词重复,图片描述,故事阅读,回忆任务。除了音频记录外,数据集还包括手动验证的转录和详细的人口统计学,临床和生物标志物元数据。ASR基准测试,自动测试评估和基于LLM的分类的基线实验说明了自动的,基于语音的认知评估的可行性,并突出了召回驱动的语音生产的诊断价值。因此,PDC建立了第一个公开的德国神经退行性疾病多模式和跨语言研究基准。
摘要:Early and accessible detection of Alzheimer's disease (AD) remains a major challenge, as current diagnostic methods often rely on costly and invasive biomarkers. Speech and language analysis has emerged as a promising non-invasive and scalable approach to detecting cognitive impairment, but research in this area is hindered by the lack of publicly available datasets, especially for languages other than English. This paper introduces the PARLO Dementia Corpus (PDC), a new multi-center, clinically validated German resource for AD collected across nine academic memory clinics in Germany. The dataset comprises speech recordings from individuals with AD-related mild cognitive impairment and mild to moderate dementia, as well as cognitively healthy controls. Speech was elicited using a standardized test battery of eight neuropsychological tasks, including confrontation naming, verbal fluency, word repetition, picture description, story reading, and recall tasks. In addition to audio recordings, the dataset includes manually verified transcriptions and detailed demographic, clinical, and biomarker metadata. Baseline experiments on ASR benchmarking, automated test evaluation, and LLM-based classification illustrate the feasibility of automatic, speech-based cognitive assessment and highlight the diagnostic value of recall-driven speech production. The PDC thus establishes the first publicly available German benchmark for multi-modal and cross-lingual research on neurodegenerative diseases.


【4】ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
标题:ZeSTA:采用域条件训练的Zero-ShotTTC增强,实现数据高效的个性化语音合成
链接:https://arxiv.org/abs/2603.04219

作者:Youngwon Choi,Jinwoo Oh,Hwayeon Kim,Hyeonyu Kim
备注:6 pages, submitted to INTERSPEECH 2026
摘要:我们调查使用zero-shot文本到语音(TTS)作为低资源个性化语音合成的数据增强源。虽然合成增强可以提供语言丰富和语音多样的语音,但天真地将大量合成语音与有限的真实录音混合通常会导致微调期间扬声器相似性降低。为了解决这个问题,我们提出了ZeSTA,一个简单的域条件训练框架,通过轻量级域嵌入区分真实和合成语音,结合真实数据过采样,在极其有限的目标数据下稳定自适应,而无需修改基础架构。LibriTTS和内部数据集上的实验表明,我们的方法在保持可懂度和感知质量的同时,提高了说话人的相似性。
摘要:We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide linguistically rich and phonetically diverse speech, naively mixing large amounts of synthetic speech with limited real recordings often leads to speaker similarity degradation during fine-tuning. To address this issue, we propose ZeSTA, a simple domain-conditioned training framework that distinguishes real and synthetic speech via a lightweight domain embedding, combined with real-data oversampling to stabilize adaptation under extremely limited target data, without modifying the base architecture. Experiments on LibriTTS and an in-house dataset with two ZS-TTS sources demonstrate that our approach improves speaker similarity over naive synthetic augmentation while preserving intelligibility and perceptual quality.


【5】Multi-Stage Music Source Restoration with BandSplit-RoFormer Separation and HiFi++ GAN
标题:利用BandSplit-RoFormer Separation和HiFi++ GAN进行多阶段音乐源恢复
链接:https://arxiv.org/abs/2603.04032

作者:Tobias Morocutti,Emmanouil Karystinaios,Jonathan Greif,Gerhard Widmer
备注:ICASSP 2026 Music Source Restoration (MSR) Challenge
摘要:音乐源恢复(MSR)的目标是恢复原始的,未经处理的乐器源于完全混合和掌握的音频,其中生产效果和分布文物违反常见的线性混合假设。本技术报告介绍了CP-JKU团队为2025年MSR ICASSP挑战赛设计的系统。我们的方法分解MSR分离和恢复。首先,单个BandSplit-RoFormer分离器预测八个阀杆加上一个辅助阀杆,并通过三阶段课程进行培训,从4阀杆热启动微调(使用LoRA)到通过头部扩展的8阀杆扩展。其次,我们应用HiFi++ GAN波形恢复器,该波形恢复器被训练为通才,然后被专业化为八个特定于乐器的专家。
摘要:Music Source Restoration (MSR) targets recovery of original, unprocessed instrument stems from fully mixed and mastered audio, where production effects and distribution artifacts violate common linear-mixture assumptions. This technical report presents the CP-JKU team's system for the MSR ICASSP Challenge 2025. Our approach decomposes MSR into separation and restoration. First, a single BandSplit-RoFormer separator predicts eight stems plus an auxiliary other stem, and is trained with a three-stage curriculum that progresses from 4-stem warm-start fine-tuning (with LoRA) to 8-stem extension via head expansion. Second, we apply a HiFi++ GAN waveform restorer trained as a generalist and then specialized into eight instrument-specific experts.


【6】Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
标题:具有稀疏模式对齐和视觉单元引导细化的鲁棒的基于LLM的视听语音识别
链接:https://arxiv.org/abs/2603.03811

作者:Fei Su,Cancan Li,Juan Liu,Wei Ju,Hongbin Suo,Ming Li
备注:submitted to Interspeech 2026
摘要:视听语音识别(AVSR)集成了声学和视觉信息,以增强在不利声学条件下的鲁棒性。大语言模型(LLM)的最新进展已经产生了有竞争力的自动语音识别性能,并显示出AVSR的有效性。然而,先前的方法独立地投射音频和视觉特征或应用浅融合,限制了跨模态对准和互补交换,同时增加了LLM的计算负荷。为了解决这个问题,我们提出了AVUR-LLM,一个基于LLM的视听语音识别通过稀疏模态对齐和视觉单元引导的细化。LRS 3上的实验证明了AVSR的最新结果。在0 dB SNR的加性噪声条件下,它比基线系统实现了37%的相对改善。
摘要:Audio-Visual Speech Recognition (AVSR) integrates acoustic and visual information to enhance robustness in adverse acoustic conditions. Recent advances in Large Language Models (LLMs) have yielded competitive automatic speech recognition performance and shown effectiveness for AVSR. However, prior approaches project audio and visual features independently or apply shallow fusion, limiting cross-modal alignment and complementary exchange while increasing the LLM's computational load. To address this, we propose AVUR-LLM, an LLM-based Audio-Visual Speech Recognition via Sparse Modality Alignment and Visual Unit-Guided Refinement. Experiments on LRS3 demonstrate state-of-the-art results for AVSR. Under additive-noise conditions at 0 dB SNR, it achieves 37% relative improvement over the baseline system.


【7】ACES: Accent Subspaces for Coupling, Explanations, and Stress-Testing in Automatic Speech Recognition
标题:ACES:自动语音识别中用于耦合、简化和压力测试的口音子空间
链接:https://arxiv.org/abs/2603.03359

作者:Swapnil Parekh
摘要:ASR系统在不同口音之间表现出持续的性能差异,但这些差异背后的内部机制仍然知之甚少。我们介绍ACES,一个代表为中心的审计,提取口音歧视子空间,并使用它们来探测模型的脆弱性和差距。分析Wav 2 Vec 2-base的五个英语口音,我们发现口音信息集中在一个低维的早期层子空间(层3,k=8)。投影幅度与每话语WER(r=0.26),至关重要的是,子空间约束扰动产生更强的耦合之间的表示移位和退化(r=0.32)比随机子空间控制(r=0.15)。最后,这个子空间的线性衰减,但并没有减少差距,并略有abrases it. We的研究结果表明,口音相关的功能深深纠缠在一起的发音关键线索,定位口音子空间作为重要的诊断工具,而不是简单的“擦除”公平杠杆。
摘要:ASR systems exhibit persistent performance disparities across accents, yet the internal mechanisms underlying these gaps remain poorly understood. We introduce ACES, a representation-centric audit that extracts accent-discriminative subspaces and uses them to probe model fragility and disparity. Analyzing Wav2Vec2-base with five English accents, we find that accent information concentrates in a low-dimensional early-layer subspace (layer 3, k=8). Projection magnitude correlates with per-utterance WER (r=0.26), and crucially, subspace-constrained perturbations yield stronger coupling between representation shift and degradation (r=0.32) than random-subspace controls (r=0.15). Finally, linear attenuation of this subspace however does not reduce disparity and slightly worsens it. Our findings suggest that accent-relevant features are deeply entangled with recognition-critical cues, positioning accent subspaces as vital diagnostic tools rather than simple "erasure" levers for fairness.


【8】Automated Measurement of Geniohyoid Muscle Thickness During Speech Using Deep Learning and Ultrasound
标题:使用深度学习和超声自动测量语音期间的膝舌骨肌厚度
链接:https://arxiv.org/abs/2603.03350

作者:Alisher Myrgyyassov,Bruce Xiao Wang,Yu Sun,Shuming Huang,Zhen Song,Min Ney Wong,Yongping Zheng
备注:6 pages, including references and acknowledgements. Submitted to Interspeech 2026
摘要:在讲话期间从超声手动测量肌肉形态是耗时的,并且限制了大规模研究。我们提出了SMMA,这是一个完全自动化的框架,它将深度学习分割与基于神经元的厚度量化相结合,以分析颏舌骨(GH)肌肉动力学。验证证明了接近人类水平的准确度(Dice = 0.9037,MAE = 0.53 mm,r = 0.901)。应用于粤语元音产生(N = 11)揭示了系统模式:/a:/显示显著更大的GH厚度(7.29 mm)比/i:/(5.95 mm,p < 0.001,Cohen's d > 1.3),表明在产生/a:/比/i:/期间更大的GH激活,与其在下颌凹陷中的作用一致。性别差异(男性高5-8%)反映了解剖尺度。SMMA实现了专家验证的准确性,同时消除了手动注释的需要,从而能够对言语运动控制进行可扩展的调查,并对言语和吞咽障碍进行客观评估。
摘要:Manual measurement of muscle morphology from ultrasound during speech is time-consuming and limits large-scale studies. We present SMMA, a fully automated framework that combines deep-learning segmentation with skeleton-based thickness quantification to analyze geniohyoid (GH) muscle dynamics. Validation demonstrates near-human-level accuracy (Dice = 0.9037, MAE = 0.53 mm, r = 0.901). Application to Cantonese vowel production (N = 11) reveals systematic patterns: /a:/ shows significantly greater GH thickness (7.29 mm) than /i:/ (5.95 mm, p < 0.001, Cohen's d > 1.3), suggesting greater GH activation during production of /a:/ than /i:/, consistent with its role in mandibular depression. Sex differences (5-8% greater in males) reflect anatomical scaling. SMMA achieves expert-validated accuracy while eliminating the need for manual annotation, enabling scalable investigations of speech motor control and objective assessment of speech and swallowing disorders.


【9】Escaping the BLEU Trap: A Signal-Grounded Framework with Decoupled Semantic Guidance for EEG-to-Text Decoding
标题:摆脱BLEU陷阱:用于脑电到文本解码的信号接地框架,具有去耦合语义指导
链接:https://arxiv.org/abs/2603.03312

作者:Yuchen Wang,Haonan Wang,Yu Guo,Honglong Yang,Xiaomeng Li
摘要:从非侵入性脑电信号中解码自然语言是一项有前途但又具有挑战性的任务。然而,目前最先进的模型仍然受到三个基本限制的限制:语义偏差(模式崩溃为通用模板),信号神经元(基于语言先验而不是神经输入的幻觉)和BLEU陷阱,其中评估指标被高频停用词人为夸大,掩盖了缺乏真正的语义保真度。为了解决这些挑战,我们提出了SemKey,一个新的多阶段框架,通过四个解耦的语义目标:情感,主题,长度和语义,强制信号接地生成。我们重新设计了神经编码器和大型语言模型(LLM)之间的交互,通过注入语义提示作为Key-Value对和EEG嵌入,严格迫使模型关注神经输入。此外,我们超越了标准的翻译指标,采用N向检索准确性和Fréchet距离来严格评估多样性和一致性。大量的实验表明,我们的方法有效地消除了幻觉的噪声输入,并实现SOTA性能这些强大的协议。代码将在https://github.com/xmed-lab/SemKey接受后发布。
摘要:Decoding natural language from non-invasive EEG signals is a promising yet challenging task. However, current state-of-the-art models remain constrained by three fundamental limitations: Semantic Bias (mode collapse into generic templates), Signal Neglect (hallucination based on linguistic priors rather than neural inputs), and the BLEU Trap, where evaluation metrics are artificially inflated by high-frequency stopwords, masking a lack of true semantic fidelity. To address these challenges, we propose SemKey, a novel multi-stage framework that enforces signal-grounded generation through four decoupled semantic objectives: sentiment, topic, length, and surprisal. We redesign the interaction between the neural encoder and the Large Language Model (LLM) by injecting semantic prompts as Queries and EEG embeddings as Key-Value pairs, strictly forcing the model to attend to neural inputs. Furthermore, we move beyond standard translation metrics by adopting N-way Retrieval Accuracy and Fréchet Distance to rigorously assess diversity and alignment. Extensive experiments demonstrate that our approach effectively eliminates hallucinations on noise inputs and achieves SOTA performance on these robust protocols. Code will be released upon acceptance at https://github.com/xmed-lab/SemKey.


机器翻译由腾讯交互翻译提供,仅供参考