微信公众号:arXiv_Daily
cs.SD语音
【1】Adapting Language Balance in Code-Switching Speech
标题:码转换语音中的语言平衡调整
链接:https://arxiv.org/abs/2510.18724
备注:Submitted to ICASSP 2026
摘要:尽管在标准基准测试中取得了令人印象深刻的结果,但大型基础模型仍然难以应对代码切换测试用例。当数据稀缺不能被用作表现不佳的通常理由时,原因可能在于语码转换时刻很少发生,第二语言的嵌入微妙地出现。与其期望模型自己学习这种不频繁的情况,不如为训练过程提供标签。评估模型性能的代码转换数据需要仔细定位的代码转换点,识别错误是最重要的,使分析强调在这些时刻发生的错误。基于这一观察,我们利用嵌入式语言和主语言之间的差异来突出这些代码转换点,从而强调在这些位置的学习。这种简单而有效的可区分替代物减轻了生成过程中的上下文偏差-代码转换的核心挑战-从而提高了模型的鲁棒性。我们对阿拉伯语和汉英双语的实验表明,该模型能够更正确地预测转换位置,这反映在减少的替换错误上。
摘要:Despite achieving impressive results on standard benchmarks, large foundational models still struggle against code-switching test cases. When data scarcity cannot be used as the usual justification for poor performance, the reason may lie in the infrequent occurrence of code-switched moments, where the embedding of the second language appears subtly. Instead of expecting the models to learn this infrequency on their own, it might be beneficial to provide the training process with labels. Evaluating model performance on code-switching data requires careful localization of code-switching points where recognition errors are most consequential, so that the analysis emphasizes mistakes occurring at those moments. Building on this observation, we leverage the difference between the embedded and the main language to highlight those code-switching points and thereby emphasize learning at those locations. This simple yet effective differentiable surrogate mitigates context bias during generation -- the central challenge in code-switching -- thereby improving the model's robustness. Our experiments with Arabic and Chinese-English showed that the models are able to predict the switching places more correctly, reflected by the reduced substitution error.
【2】Bayesian Low-Rank Factorization for Robust Model Adaptation
标题:用于鲁棒模型自适应的Bayesian低等级因子分解
链接:https://arxiv.org/abs/2510.18723
备注:Submitted to ICASSP 2026
摘要:大型语音基础模型在许多领域都具有很强的性能,但它们通常需要适应以处理本地需求,例如语码转换,即说话者在同一话语中混合语言。直接对这些模型进行微调会冒着过度拟合目标域的风险,并破坏基础模型的广泛功能。为了解决这一挑战,我们探索贝叶斯分解适配器的语音基础模型,将先验接近零,以实现稀疏的适应矩阵,从而保持一般性能,同时适应特定的领域。我们将我们的方法应用到Whisper模型中,并对不同的多语言代码切换场景进行评估。我们的研究结果表明,只有最小的适应损失,同时显着减少灾难性遗忘的基础模型。与LoRA相比,我们的方法实现了54%的后向增益,在新域上仅下降4%。这些发现突出了贝叶斯适应的有效性,微调语音基础模型,而不牺牲泛化。
摘要:Large speech foundation models achieve strong performance across many domains, but they often require adaptation to handle local needs such as code-switching, where speakers mix languages within the same utterance. Direct fine-tuning of these models risks overfitting to the target domain and overwriting the broad capabilities of the base model. To address this challenge, we explore Bayesian factorized adapters for speech foundation models, which place priors near zero to achieve sparser adaptation matrices and thereby retain general performance while adapting to specific domains. We apply our approach to the Whisper model and evaluate on different multilingual code-switching scenarios. Our results show only minimal adaptation loss while significantly reducing catastrophic forgetting of the base model. Compared to LoRA, our method achieves a backward gain of 54% with only a 4% drop on the new domain. These findings highlight the effectiveness of Bayesian adaptation for fine-tuning speech foundation models without sacrificing generalization.
【3】MLMA: Towards Multilingual with Mamba Based Architectures
标题:MLMA:利用基于曼巴的架构实现多语言
链接:https://arxiv.org/abs/2510.18684
备注:The paper is under review at ICASSP 2026
摘要:多语言自动语音识别(ASR)仍然是一项具有挑战性的任务,特别是在平衡高资源和低资源语言的性能时。序列建模的最新进展表明,超越Transformers的体系结构可以提供更好的可伸缩性和效率。在这项工作中,我们介绍了MLMA(多语言语言建模与Mamba的ASR),一种新的方法,利用Mamba架构-一个有效的状态空间模型优化的长上下文序列处理-多语言ASR。使用Mamba,MLMA隐式地结合了语言感知条件和共享表示,以支持跨多种语言的鲁棒识别。标准的多语言基准测试的实验表明,MLMA实现竞争力的性能相比,基于transformer的架构。这些结果突出了Mamba作为可扩展,高效和准确的多语言语音识别的强大骨干的潜力。
摘要:Multilingual automatic speech recognition (ASR) remains a challenging task, especially when balancing performance across high- and low-resource languages. Recent advances in sequence modeling suggest that architectures beyond Transformers may offer better scalability and efficiency. In this work, we introduce MLMA (Multilingual Language Modeling with Mamba for ASR), a new approach that leverages the Mamba architecture--an efficient state-space model optimized for long-context sequence processing--for multilingual ASR. Using Mamba, MLMA implicitly incorporates language-aware conditioning and shared representations to support robust recognition across diverse languages. Experiments on standard multilingual benchmarks show that MLMA achieves competitive performance compared to Transformer-based architectures. These results highlight Mamba's potential as a strong backbone for scalable, efficient, and accurate multilingual speech recognition.
【4】Noise-Conditioned Mixture-of-Experts Framework for Robust Speaker Verification
标题:用于鲁棒说话人验证的噪音条件专家混合框架
链接:https://arxiv.org/abs/2510.18533
摘要:噪声条件下的鲁棒说话人确认仍然是一个公开的挑战。传统的深度学习方法可以针对不同的背景噪声学习鲁棒的统一说话人表示空间,并取得显着的改进。与此相反,本文提出了一个噪声条件下的混合专家框架,将特征空间分解成专门的噪声感知子空间的说话人验证。具体来说,我们提出了一个噪声条件下的专家路由机制,一个通用的模型为基础的专家专业化策略,和SNR衰减的课程学习协议,共同提高模型的鲁棒性和泛化在不同的噪声条件下。所提出的方法可以自动路由输入到专家网络的基础上从输入,其中每个专家的目标不同的噪声特性,同时保留扬声器的身份信息的噪声信息。全面的实验表明,一致优于基线,确认显式噪声相关的特征建模显着提高鲁棒性,而不牺牲验证精度。
摘要:Robust speaker verification under noisy conditions remains an open challenge. Conventional deep learning methods learn a robust unified speaker representation space against diverse background noise and achieve significant improvement. In contrast, this paper presents a noise-conditioned mixture-ofexperts framework that decomposes the feature space into specialized noise-aware subspaces for speaker verification. Specifically, we propose a noise-conditioned expert routing mechanism, a universal model based expert specialization strategy, and an SNR-decaying curriculum learning protocol, collectively improving model robustness and generalization under diverse noise conditions. The proposed method can automatically route inputs to expert networks based on noise information derived from the inputs, where each expert targets distinct noise characteristics while preserving speaker identity information. Comprehensive experiments demonstrate consistent superiority over baselines, confirming that explicit noise-dependent feature modeling significantly enhances robustness without sacrificing verification accuracy.
【5】A Stage-Wise Learning Strategy with Fixed Anchors for Robust Speaker Verification
标题:用于鲁棒说话人验证的固定条件分阶段学习策略
链接:https://arxiv.org/abs/2510.18530
摘要:在噪声条件下学习鲁棒的说话人表示提出了重大挑战,这需要仔细处理的歧视性和噪声不变的属性。在这项工作中,我们提出了一个基于锚点的阶段式学习策略,用于鲁棒的说话人表示学习。具体来说,我们的方法首先训练一个基础模型来建立区分性的说话人边界,然后从这个模型中提取锚嵌入作为稳定的参考。最后,基础模型的副本在有噪声的输入上进行微调,通过强制接近其相应的固定锚嵌入来正则化,以在失真下保留说话人身份。实验结果表明,这种策略提供了传统的联合优化的优势,特别是在保持歧视,同时提高噪声鲁棒性。所提出的方法在各种噪声条件下表现出一致的改进,可能是由于其能够单独处理边界稳定和变化抑制。
摘要:Learning robust speaker representations under noisy conditions presents significant challenges, which requires careful handling of both discriminative and noise-invariant properties. In this work, we proposed an anchor-based stage-wise learning strategy for robust speaker representation learning. Specifically, our approach begins by training a base model to establish discriminative speaker boundaries, and then extract anchor embeddings from this model as stable references. Finally, a copy of the base model is fine-tuned on noisy inputs, regularized by enforcing proximity to their corresponding fixed anchor embeddings to preserve speaker identity under distortion. Experimental results suggest that this strategy offers advantages over conventional joint optimization, particularly in maintaining discrimination while improving noise robustness. The proposed method demonstrates consistent improvements across various noise conditions, potentially due to its ability to handle boundary stabilization and variation suppression separately.
【6】SegTune: Structured and Fine-Grained Control for Song Generation
标题:SegButton:歌曲生成的结构化和细粒度控制
链接:https://arxiv.org/abs/2510.18416
摘要:歌曲生成的最新进展已经显示出从歌词和/或全局文本提示生成歌曲的有希望的结果。然而,大多数现有的系统缺乏对歌曲的随时间变化的属性进行建模的能力,限制了对音乐结构和动态的细粒度控制。在本文中,我们提出了SegTune,结构化和可控的歌曲生成的非自回归框架。SegTune允许用户或大型语言模型指定与歌曲片段对齐的本地音乐描述,从而实现片段级控制。片段提示通过临时广播到相应的时间窗口来注入模型,而全局提示则影响整首歌曲,以确保风格一致性。为了获得准确的片段持续时间并实现精确的歌词到音乐对齐,我们引入了一个基于LLM的持续时间预测器,该预测器以LRC格式自回归地生成具有时间戳的歌词。我们进一步构建了一个大规模的数据管道,用于收集具有对齐歌词和提示的高质量歌曲,并提出了新的评估指标来评估片段级对齐和声乐属性一致性。实验结果表明,与现有的基线相比,SegTune实现了更好的可控性和音乐连贯性。请访问https://cai525.github.io/SegTune_demo查看我们工作的演示。
摘要:Recent advancements in song generation have shown promising results in generating songs from lyrics and/or global text prompts. However, most existing systems lack the ability to model the temporally varying attributes of songs, limiting fine-grained control over musical structure and dynamics. In this paper, we propose SegTune, a non-autoregressive framework for structured and controllable song generation. SegTune enables segment-level control by allowing users or large language models to specify local musical descriptions aligned to song sections.The segmental prompts are injected into the model by temporally broadcasting them to corresponding time windows, while global prompts influence the whole song to ensure stylistic coherence. To obtain accurate segment durations and enable precise lyric-to-music alignment, we introduce an LLM-based duration predictor that autoregressively generates sentence-level timestamped lyrics in LRC format. We further construct a large-scale data pipeline for collecting high-quality songs with aligned lyrics and prompts, and propose new evaluation metrics to assess segment-level alignment and vocal attribute consistency. Experimental results show that SegTune achieves superior controllability and musical coherence compared to existing baselines. See https://cai525.github.io/SegTune_demo for demos of our work.
【7】ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
标题:ParaStyleTTC:实现高效且稳健的副语言风格控制,以实现表达性文本转语音生成
链接:https://arxiv.org/abs/2510.18308
摘要:在文语转换(TTS)系统中控制说话风格已经成为学术界和工业界日益关注的焦点。虽然许多现有的方法依赖于参考音频来指导样式生成,但是由于隐私问题和有限的可访问性,这样的方法通常是不切实际的。最近,大型语言模型(LLM)已被用于控制说话风格,通过自然语言提示,然而,其高计算成本,缺乏可解释性,并提示措辞的敏感性限制了其在实时和资源受限的环境中的适用性。在这项工作中,我们提出了ParaStyleTTS,一个轻量级的和可解释的TTS框架,使表达风格控制从文本提示。ParaStyleTTS的特点是一种新颖的两级风格适应架构,分离韵律和非语言学的语音风格建模。它允许对情绪,性别和年龄等因素进行细粒度和鲁棒的控制。与基于LLM的方法不同,ParaStyleTTS在不同的提示公式中保持一致的风格实现,非常适合现实世界的应用程序,包括设备上和低资源部署。实验结果表明,ParaStyleTTS生成高质量的语音,其性能与最先进的基于LLM的系统相当,同时速度快30倍,使用的参数少8倍,需要的CUDA内存少2.5倍。此外,ParaStyleTTS表现出优越的鲁棒性和可控性,超过非语言的说话风格,提供了一个实用和有效的解决方案,风格可控的文本到语音生成。演示可以在https://parastyletts.github.io/ParaStyleTTS_Demo/上找到。代码可以在https://github.com/haoweilou/ParaStyleTTS上找到。
摘要:Controlling speaking style in text-to-speech (TTS) systems has become a growing focus in both academia and industry. While many existing approaches rely on reference audio to guide style generation, such methods are often impractical due to privacy concerns and limited accessibility. More recently, large language models (LLMs) have been used to control speaking style through natural language prompts; however, their high computational cost, lack of interpretability, and sensitivity to prompt phrasing limit their applicability in real-time and resource-constrained environments. In this work, we propose ParaStyleTTS, a lightweight and interpretable TTS framework that enables expressive style control from text prompts alone. ParaStyleTTS features a novel two-level style adaptation architecture that separates prosodic and paralinguistic speech style modeling. It allows fine-grained and robust control over factors such as emotion, gender, and age. Unlike LLM-based methods, ParaStyleTTS maintains consistent style realization across varied prompt formulations and is well-suited for real-world applications, including on-device and low-resource deployment. Experimental results show that ParaStyleTTS generates high-quality speech with performance comparable to state-of-the-art LLM-based systems while being 30x faster, using 8x fewer parameters, and requiring 2.5x less CUDA memory. Moreover, ParaStyleTTS exhibits superior robustness and controllability over paralinguistic speaking styles, providing a practical and efficient solution for style-controllable text-to-speech generation. Demo can be found at https://parastyletts.github.io/ParaStyleTTS_Demo/. Code can be found at https://github.com/haoweilou/ParaStyleTTS.
【8】Transformer Redesign for Late Fusion of Audio-Text Features on Ultra-Low-Power Edge Hardware
标题:Transformer重新设计,以在超低功耗边缘硬件上实现音频文本功能的后期融合
链接:https://arxiv.org/abs/2510.18036
摘要:在设备必须小巧、低功耗和私密的现实环境中部署情感识别系统仍然是一个重大挑战。这对于诸如紧张监控、冲突降级和响应式可穿戴设备等应用尤其重要,在这些应用中,基于云的解决方案不切实际。多模态情感识别通过深度学习取得了进展,但大多数系统仍然不适合部署在超受限的边缘设备上。以前的工作通常依赖于强大的硬件,缺乏实时性能,或使用单峰输入。本文通过提出一种硬件感知的情感识别系统来解决这一差距,该系统使用针对Edge TPU优化的后期融合架构将声学和语言特征相结合。该设计将基于量化变换器的声学模型与DSResNet-SE网络中的冻结关键字嵌入集成在一起,从而在1.8MB内存预算和21- 23 ms延迟内实现实时推理。该管道使用MicroFrontend和MLTK确保训练和部署之间的频谱图对齐。对通过Coral Dev Board Micro麦克风捕获的重新记录的分段IEMOCAP样本进行评估,结果显示与单峰基线相比,宏观F1改善了6.3%。这项工作表明,通过特定于任务的融合和硬件引导的模型设计,可以在微处理器级边缘平台上实现准确,实时的多模态情感推理。
摘要:Deploying emotion recognition systems in real-world environments where devices must be small, low-power, and private remains a significant challenge. This is especially relevant for applications such as tension monitoring, conflict de-escalation, and responsive wearables, where cloud-based solutions are impractical. Multimodal emotion recognition has advanced through deep learning, but most systems remain unsuitable for deployment on ultra-constrained edge devices. Prior work typically relies on powerful hardware, lacks real-time performance, or uses unimodal input. This paper addresses that gap by presenting a hardware-aware emotion recognition system that combines acoustic and linguistic features using a late-fusion architecture optimised for Edge TPU. The design integrates a quantised transformer-based acoustic model with frozen keyword embeddings from a DSResNet-SE network, enabling real-time inference within a 1.8MB memory budget and 21-23ms latency. The pipeline ensures spectrogram alignment between training and deployment using MicroFrontend and MLTK. Evaluation on re-recorded, segmented IEMOCAP samples captured through the Coral Dev Board Micro microphone shows a 6.3% macro F1 improvement over unimodal baselines. This work demonstrates that accurate, real-time multimodal emotion inference is achievable on microcontroller-class edge platforms through task-specific fusion and hardware-guided model design.
【9】Diffusion Buffer for Online Generative Speech Enhancement
标题:用于在线生成语音增强的扩散缓冲区
链接:https://arxiv.org/abs/2510.18744
摘要:在线语音增强主要用于预测模型。这些模型的一个主要优点是,对于来自数据流的传入信号帧,模型仅被调用一次以进行增强。相比之下,生成式语音增强模型通常需要多次调用,从而导致对于许多在线语音增强应用而言太高的计算复杂度。这项工作提出了扩散缓冲区,这是一种基于生成扩散的语音增强模型,它只需要从数据流中每个传入信号帧调用一个神经网络,并在消费级GPU上以在线方式执行增强。扩散缓冲区的关键思想是将物理时间与扩散时间步长对齐。该方法通过物理时间逐步对帧进行降噪,其中过去的帧具有更多的噪声被去除。因此,增强帧以扩散缓冲器定义的延迟输出给收听者,并且输出帧具有相应的前瞻。在这项工作中,我们通过仔细设计一个2D卷积UNet架构来扩展我们以前的工作,该架构专门与扩散缓冲区的前瞻性保持一致。我们观察到,建议UNet提高性能,特别是当算法延迟低。此外,我们表明,使用数据预测损失而不是去噪分数匹配损失,可以灵活地控制推理过程中算法延迟和质量之间的权衡。扩展的扩散缓冲区配备了一个新的神经网络和损失函数,大大减少了算法延迟从320 - 960毫秒到32 - 176毫秒,甚至提高了性能。虽然之前已经表明离线生成扩散模型在看不见的噪声语音数据中优于预测方法,但我们确认在线扩散缓冲区在看不见的噪声语音数据上也优于其预测对应物。
摘要:Online Speech Enhancement was mainly reserved for predictive models. A key advantage of these models is that for an incoming signal frame from a stream of data, the model is called only once for enhancement. In contrast, generative Speech Enhancement models often require multiple calls, resulting in a computational complexity that is too high for many online speech enhancement applications. This work presents the Diffusion Buffer, a generative diffusion-based Speech Enhancement model which only requires one neural network call per incoming signal frame from a stream of data and performs enhancement in an online fashion on a consumer-grade GPU. The key idea of the Diffusion Buffer is to align physical time with Diffusion time-steps. The approach progressively denoises frames through physical time, where past frames have more noise removed. Consequently, an enhanced frame is output to the listener with a delay defined by the Diffusion Buffer, and the output frame has a corresponding look-ahead. In this work, we extend upon our previous work by carefully designing a 2D convolutional UNet architecture that specifically aligns with the Diffusion Buffer's look-ahead. We observe that the proposed UNet improves performance, particularly when the algorithmic latency is low. Moreover, we show that using a Data Prediction loss instead of Denoising Score Matching loss enables flexible control over the trade-off between algorithmic latency and quality during inference. The extended Diffusion Buffer equipped with a novel NN and loss function drastically reduces the algorithmic latency from 320 - 960 ms to 32 - 176 ms with an even increased performance. While it has been shown before that offline generative diffusion models outperform predictive approaches in unseen noisy speech data, we confirm that the online Diffusion Buffer also outperforms its predictive counterpart on unseen noisy speech data.
【10】ProLAP: Probabilistic Language-Audio Pre-Training
标题:Propriate:概率放大-音频预训练
链接:https://arxiv.org/abs/2510.18423
备注:Under review
摘要:音频-音频联合表示学习框架通常依赖于确定性嵌入,假设音频和文本之间存在一一对应关系。然而,在现实世界中,语言-音频关系本质上是多对多的:一个音频片段可以由多个字幕描述,反之亦然。为了解决这个问题,我们提出了概率音频预训练(Probability-Audio Pre-training,Probability-Audio Pre-training),它将多重性建模为联合语言音频嵌入空间中概率分布的分布。为了有效地训练模态内层次关系,我们还引入了两个目标:(i)层次包含损失,以促进对输入的语义层次理解;(ii)屏蔽排斥损失,以在优化层次包含损失时提高学习效率。通过这种训练策略,我们的模型甚至可以从小数据集中学习数据中固有的层次结构,这与依赖于大规模数据集的先前概率方法相反。在我们的实验中,Proximity优于现有的确定性方法的音频文本检索任务。此外,通过本文中介绍的音频遍历任务的实验,我们证明了Proximity捕获的似是而非的语义层次结构。
摘要:Language-audio joint representation learning frameworks typically depend on deterministic embeddings, assuming a one-to-one correspondence between audio and text. In real-world settings, however, the language-audio relationship is inherently many-to-many: one audio segment can be described by multiple captions and vice versa. To address this, we propose Probabilistic Language-Audio Pre-training (ProLAP), which models multiplicity as the spread of probability distributions in a joint language-audio embedding space. To train the intra-modal hierarchical relationship effectively, we also introduce two objectives: (i) hierarchical inclusion loss to promote semantic hierarchical understanding of inputs and (ii) mask repulsive loss to improve the efficiency of learning when optimizing the hierarchical inclusion loss. With this training strategy, our model can learn the hierarchical structure inherent in the data even from small datasets, in contrast to prior probabilistic approaches that rely on large-scale datasets. In our experiments, ProLAP outperforms existing deterministic approaches on audio-text retrieval tasks. Moreover, through experiments on the audio traversal task introduced in this paper, we demonstrate that ProLAP captures the plausible semantic hierarchy.
【11】MVDR Beamforming for Cyclostationary Processes
标题:循环平稳过程的MVDR束形成
链接:https://arxiv.org/abs/2510.18391
备注:Under review for publication from September 2025
摘要:传统的声学波束形成器假设噪声在短时间帧内是静止的。这种假设阻止了他们利用几乎周期性噪声源(如乐器,风扇和发动机)频率之间的相关性。这些信号表现出周期性变化的统计,更好地建模为循环平稳过程。本文介绍了循环MVDR(cMVDR)波束形成器,传统的MVDR的扩展,利用空间和频谱相关性,以提高降噪,特别是在低信噪比的情况下。该方法建立在频移(FRESH)滤波的基础上,其中输入的移位版本被组合以衰减或放大跨频率相干的分量。为了解决谐波偏频偏离基频的精确整数倍的不和谐性,我们提出了一种数据驱动的策略,该策略通过周期图分析估计谐振频率,并从其间距计算频移。分析和实验结果表明,性能提高光谱相关性。在实际录音中,cMVDR在标度不变信号失真比(SI-SDR)方面比MVDR增益高达5 dB,即使使用单个麦克风也能保持有效。代码可在https://github.com/Screeen/cMVDR上获得。
摘要:Conventional acoustic beamformers assume that noise is stationary within short time frames. This assumption prevents them from exploiting correlations between frequencies in almost-periodic noise sources such as musical instruments, fans, and engines. These signals exhibit periodically varying statistics and are better modeled as cyclostationary processes. This paper introduces the cyclic MVDR (cMVDR) beamformer, an extension of the conventional MVDR that leverages both spatial and spectral correlations to improve noise reduction, particularly in low-SNR scenarios. The method builds on frequency-shifted (FRESH) filtering, where shifted versions of the input are combined to attenuate or amplify components that are coherent across frequency. To address inharmonicity, where harmonic partials deviate from exact integer multiples of the fundamental frequency, we propose a data-driven strategy that estimates resonant frequencies via periodogram analysis and computes the frequency shifts from their spacing. Analytical and experimental results demonstrate that performance improves with increasing spectral correlation. On real recordings, the cMVDR achieves up to 5 dB gain in scale-invariant signal-to-distortion ratio (SI-SDR) over the MVDR and remains effective even with a single microphone. Code is available at https://github.com/Screeen/cMVDR.
【12】Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing
标题:用于鲁棒音频信号处理的自适应单通道能量归一化前端
链接:https://arxiv.org/abs/2510.18206
备注:Submitted to ICASSP2026
摘要:在音频信号处理中,可学习的前端通过优化特定于任务的表示,在不同的任务中表现出强大的性能。然而,它们的参数在训练后保持固定,在推理过程中缺乏灵活性,并且在动态复杂声学环境下限制了鲁棒性。在本文中,我们介绍了一种新的自适应模式的音频前端,取代静态参数化与闭环神经控制器。具体来说,我们简化了可学习的前端LEAF架构,并通过动态调整每通道能量归一化来集成神经控制器进行自适应表示。神经控制器利用当前和缓冲过去的子带能量,使输入相关的自适应推理过程中。多个音频分类任务的实验结果表明,所提出的自适应前端始终优于以前的固定和可学习的前端在干净和复杂的声学条件下。这些结果突出了神经适应性作为下一代音频前端的一个有前途的方向。
摘要:In audio signal processing, learnable front-ends have shown strong performance across diverse tasks by optimizing task-specific representation. However, their parameters remain fixed once trained, lacking flexibility during inference and limiting robustness under dynamic complex acoustic environments. In this paper, we introduce a novel adaptive paradigm for audio front-ends that replaces static parameterization with a closed-loop neural controller. Specifically, we simplify the learnable front-end LEAF architecture and integrate a neural controller for adaptive representation via dynamically tuning Per-Channel Energy Normalization. The neural controller leverages both the current and the buffered past subband energies to enable input-dependent adaptation during inference. Experimental results on multiple audio classification tasks demonstrate that the proposed adaptive front-end consistently outperforms prior fixed and learnable front-ends under both clean and complex acoustic conditions. These results highlight neural adaptability as a promising direction for the next generation of audio front-ends.
【13】Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network
标题:基于多任务多尺度网络的钢琴动力学与韵律结构联合估计
链接:https://arxiv.org/abs/2510.18190
备注:Paper submitted to ICASSP2026
摘要:从音频记录中估计钢琴动态是计算音乐分析中的一个基本挑战。在本文中,我们提出了一个高效的多任务网络,它可以从共享的潜在表示中联合预测动态水平,变化点,节拍和下拍。这四个目标构成了乐谱中力度的韵律结构。受近年来声乐动力学研究的启发,我们采用多尺度网络作为主干,以树皮尺度特定响度作为输入特征。与作为输入的log-Mel相比,这将模型大小从14.7 M减少到0.5 M,从而实现长序列输入。我们在音频分割中使用60秒的音频长度,这是通常使用的节拍跟踪长度的两倍。在公共MazurkaBL数据集上进行评估,我们的模型在所有任务中都获得了最先进的结果。这项工作为钢琴动态估计设定了一个新的基准,并提供了一个强大而紧凑的工具,为大规模,资源有效的音乐表达分析铺平了道路。
摘要:Estimating piano dynamic from audio recordings is a fundamental challenge in computational music analysis. In this paper, we propose an efficient multi-task network that jointly predicts dynamic levels, change points, beats, and downbeats from a shared latent representation. These four targets form the metrical structure of dynamics in the music score. Inspired by recent vocal dynamic research, we use a multi-scale network as the backbone, which takes Bark-scale specific loudness as the input feature. Compared to log-Mel as input, this reduces model size from 14.7 M to 0.5 M, enabling long sequential input. We use a 60-second audio length in audio segmentation, which doubled the length of beat tracking commonly used. Evaluated on the public MazurkaBL dataset, our model achieves state-of-the-art results across all tasks. This work sets a new benchmark for piano dynamic estimation and delivers a powerful and compact tool, paving the way for large-scale, resource-efficient analysis of musical expression.
【14】Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction
标题:家庭医疗保健中的听力健康:利用LLM进行疾病评分,利用ILM进行声乐生物标志物提取
链接:https://arxiv.org/abs/2510.18169
备注:The Second Workshop on GenAI for Health at NeurIPS 2025
摘要:对家庭医疗保健的需求不断增长,需要能够支持医疗服务的工具。在这项研究中,我们使用真实世界的家庭护理访问数据,利用其包含的各种患者信息,探索语音自动健康评估。首先,我们利用大型语言模型(LLM)将来自非结构化音频转录和结构化生命体征的主观,客观,评估和计划(SOAP)笔记整合到反映患者整体健康状况的整体疾病评分中。这种紧凑的表示有利于跨访视健康状况比较和下游分析。接下来,我们设计了一个多阶段的预处理管道,从家庭护理录音中的目标扬声器中提取短语音片段进行声学分析。然后,我们采用音频语言模型(ALM)来产生声音生物标志物的简单语言描述,并检查它们与个人健康状况的关联。我们的实验结果在估计疾病评分方面对商业和开源LLM进行了基准测试,证明了它们与实际临床结果的一致性,并揭示了SOAP笔记比生命体征提供的信息要多得多。在疾病评分的基础上,我们提供了第一个证据,即ALMs可以从家庭护理记录中识别与健康相关的声学模式,并以人类可读的形式呈现它们。总之,这些研究结果突出了LLM和ALM利用异构家庭访视数据进行更好的患者监测和护理的潜力。
摘要:The growing demand for home healthcare calls for tools that can support care delivery. In this study, we explore automatic health assessment from voice using real-world home care visit data, leveraging the diverse patient information it contains. First, we utilize Large Language Models (LLMs) to integrate Subjective, Objective, Assessment, and Plan (SOAP) notes derived from unstructured audio transcripts and structured vital signs into a holistic illness score that reflects a patient's overall health. This compact representation facilitates cross-visit health status comparisons and downstream analysis. Next, we design a multi-stage preprocessing pipeline to extract short speech segments from target speakers in home care recordings for acoustic analysis. We then employ an Audio Language Model (ALM) to produce plain-language descriptions of vocal biomarkers and examine their association with individuals' health status. Our experimental results benchmark both commercial and open-source LLMs in estimating illness scores, demonstrating their alignment with actual clinical outcomes, and revealing that SOAP notes are substantially more informative than vital signs. Building on the illness scores, we provide the first evidence that ALMs can identify health-related acoustic patterns from home care recordings and present them in a human-readable form. Together, these findings highlight the potential of LLMs and ALMs to harness heterogeneous in-home visit data for better patient monitoring and care.
【15】SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
标题:SAC:语义-声学双流量化神经语音编解码器
链接:https://arxiv.org/abs/2510.16841
摘要:将连续语音信号转换为离散标记的语音编解码器已成为语音语言模型(SLM)的基本组成部分。然而,现有的编解码器难以平衡高质量的重建与语义丰富的表示,限制了它们在生成和理解任务中的有效性。在这项工作中,我们提出了SAC,神经语音编解码器与语义声学双流量化。通过将语义和声学建模分解为两个专用流,SAC使每个流都能够针对其各自的角色进行优化。综合评估表明,SAC在干净和嘈杂的条件下,在不同的比特率下都实现了强大的重建性能,UTMOS和WER的得分特别高,表现出卓越的感知质量和可懂度。此外,SAC在语义表示方面远远优于最先进的编解码器,达到了与自监督学习(SSL)连续嵌入相当的水平。最后,我们的语音解纠缠的分析突出了双流设计的有效性,可控语音应用提供了新的潜力。
摘要:Speech codecs that convert continuous speech signals into discrete tokens have become essential for speech language models (SLMs). However, existing codecs struggle to balance high-quality reconstruction with semantically rich representations, limiting their effectiveness in both generative and understanding tasks. In this work, we propose SAC, a neural speech codec with semantic-acoustic dual-stream quantization. By disentangling semantic and acoustic modeling into two dedicated streams, SAC enables each to be optimized for its respective role. Comprehensive evaluations show that SAC achieves strong reconstruction performance across diverse bitrates under both clean and noisy conditions, with particularly high scores on UTMOS and WER, demonstrating superior perceptual quality and intelligibility. Moreover, SAC substantially outperforms state-of-the-art codecs in semantic representation, achieving a level comparable to that of self-supervised learning (SSL) continuous embeddings. Finally, our analysis of speech disentanglement highlights the effectiveness of the dual-stream design, offering new potential for controllable speech applications.
【1】Diffusion Buffer for Online Generative Speech Enhancement
标题:用于在线生成语音增强的扩散缓冲区
链接:https://arxiv.org/abs/2510.18744
摘要:在线语音增强主要用于预测模型。这些模型的一个主要优点是,对于来自数据流的传入信号帧,模型仅被调用一次以进行增强。相比之下,生成式语音增强模型通常需要多次调用,从而导致对于许多在线语音增强应用而言太高的计算复杂度。这项工作提出了扩散缓冲区,这是一种基于生成扩散的语音增强模型,它只需要从数据流中每个传入信号帧调用一个神经网络,并在消费级GPU上以在线方式执行增强。扩散缓冲区的关键思想是将物理时间与扩散时间步长对齐。该方法通过物理时间逐步对帧进行降噪,其中过去的帧具有更多的噪声被去除。因此,增强帧以扩散缓冲器定义的延迟输出给收听者,并且输出帧具有相应的前瞻。在这项工作中,我们通过仔细设计一个2D卷积UNet架构来扩展我们以前的工作,该架构专门与扩散缓冲区的前瞻性保持一致。我们观察到,建议UNet提高性能,特别是当算法延迟低。此外,我们表明,使用数据预测损失而不是去噪分数匹配损失,可以灵活地控制推理过程中算法延迟和质量之间的权衡。扩展的扩散缓冲区配备了一个新的神经网络和损失函数,大大减少了算法延迟从320 - 960毫秒到32 - 176毫秒,甚至提高了性能。虽然之前已经表明离线生成扩散模型在看不见的噪声语音数据中优于预测方法,但我们确认在线扩散缓冲区在看不见的噪声语音数据上也优于其预测对应物。
摘要:Online Speech Enhancement was mainly reserved for predictive models. A key advantage of these models is that for an incoming signal frame from a stream of data, the model is called only once for enhancement. In contrast, generative Speech Enhancement models often require multiple calls, resulting in a computational complexity that is too high for many online speech enhancement applications. This work presents the Diffusion Buffer, a generative diffusion-based Speech Enhancement model which only requires one neural network call per incoming signal frame from a stream of data and performs enhancement in an online fashion on a consumer-grade GPU. The key idea of the Diffusion Buffer is to align physical time with Diffusion time-steps. The approach progressively denoises frames through physical time, where past frames have more noise removed. Consequently, an enhanced frame is output to the listener with a delay defined by the Diffusion Buffer, and the output frame has a corresponding look-ahead. In this work, we extend upon our previous work by carefully designing a 2D convolutional UNet architecture that specifically aligns with the Diffusion Buffer's look-ahead. We observe that the proposed UNet improves performance, particularly when the algorithmic latency is low. Moreover, we show that using a Data Prediction loss instead of Denoising Score Matching loss enables flexible control over the trade-off between algorithmic latency and quality during inference. The extended Diffusion Buffer equipped with a novel NN and loss function drastically reduces the algorithmic latency from 320 - 960 ms to 32 - 176 ms with an even increased performance. While it has been shown before that offline generative diffusion models outperform predictive approaches in unseen noisy speech data, we confirm that the online Diffusion Buffer also outperforms its predictive counterpart on unseen noisy speech data.
【2】ProLAP: Probabilistic Language-Audio Pre-Training
标题:Propriate:概率放大-音频预训练
链接:https://arxiv.org/abs/2510.18423
备注:Under review
摘要:语言-音频联合表示学习框架通常依赖于确定性嵌入,假设音频和文本之间存在一一对应关系。然而,在现实世界中,语言-音频关系本质上是多对多的:一个音频片段可以由多个字幕描述,反之亦然。为了解决这个问题,我们提出了概率音频预训练(Probability-Audio Pre-training,Probability-Audio Pre-training),它将多重性建模为联合语言音频嵌入空间中概率分布的分布。为了有效地训练模态内层次关系,我们还引入了两个目标:(i)层次包含损失,以促进对输入的语义层次理解;(ii)屏蔽排斥损失,以在优化层次包含损失时提高学习效率。通过这种训练策略,我们的模型甚至可以从小数据集中学习数据中固有的层次结构,这与依赖于大规模数据集的先前概率方法相反。在我们的实验中,Proximity优于现有的确定性方法的音频文本检索任务。此外,通过本文中介绍的音频遍历任务的实验,我们证明了Proximity捕获的似是而非的语义层次结构。
摘要:Language-audio joint representation learning frameworks typically depend on deterministic embeddings, assuming a one-to-one correspondence between audio and text. In real-world settings, however, the language-audio relationship is inherently many-to-many: one audio segment can be described by multiple captions and vice versa. To address this, we propose Probabilistic Language-Audio Pre-training (ProLAP), which models multiplicity as the spread of probability distributions in a joint language-audio embedding space. To train the intra-modal hierarchical relationship effectively, we also introduce two objectives: (i) hierarchical inclusion loss to promote semantic hierarchical understanding of inputs and (ii) mask repulsive loss to improve the efficiency of learning when optimizing the hierarchical inclusion loss. With this training strategy, our model can learn the hierarchical structure inherent in the data even from small datasets, in contrast to prior probabilistic approaches that rely on large-scale datasets. In our experiments, ProLAP outperforms existing deterministic approaches on audio-text retrieval tasks. Moreover, through experiments on the audio traversal task introduced in this paper, we demonstrate that ProLAP captures the plausible semantic hierarchy.
【3】MVDR Beamforming for Cyclostationary Processes
标题:循环平稳过程的MVDR束形成
链接:https://arxiv.org/abs/2510.18391
备注:Under review for publication from September 2025
摘要:传统的声学波束形成器假设噪声在短时间帧内是静止的。这种假设阻止了他们利用几乎周期性噪声源(如乐器,风扇和发动机)频率之间的相关性。这些信号表现出周期性变化的统计,更好地建模为循环平稳过程。本文介绍了循环MVDR(cMVDR)波束形成器,传统的MVDR的扩展,利用空间和频谱相关性,以提高降噪,特别是在低信噪比的情况下。该方法建立在频移(FRESH)滤波的基础上,其中输入的移位版本被组合以衰减或放大跨频率相干的分量。为了解决谐波偏频偏离基频的精确整数倍的不和谐性,我们提出了一种数据驱动的策略,该策略通过周期图分析估计谐振频率,并从其间距计算频移。分析和实验结果表明,性能提高光谱相关性。在实际录音中,cMVDR在标度不变信号失真比(SI-SDR)方面比MVDR增益高达5 dB,即使使用单个麦克风也能保持有效。代码可在https://github.com/Screeen/cMVDR上获得。
摘要:Conventional acoustic beamformers assume that noise is stationary within short time frames. This assumption prevents them from exploiting correlations between frequencies in almost-periodic noise sources such as musical instruments, fans, and engines. These signals exhibit periodically varying statistics and are better modeled as cyclostationary processes. This paper introduces the cyclic MVDR (cMVDR) beamformer, an extension of the conventional MVDR that leverages both spatial and spectral correlations to improve noise reduction, particularly in low-SNR scenarios. The method builds on frequency-shifted (FRESH) filtering, where shifted versions of the input are combined to attenuate or amplify components that are coherent across frequency. To address inharmonicity, where harmonic partials deviate from exact integer multiples of the fundamental frequency, we propose a data-driven strategy that estimates resonant frequencies via periodogram analysis and computes the frequency shifts from their spacing. Analytical and experimental results demonstrate that performance improves with increasing spectral correlation. On real recordings, the cMVDR achieves up to 5 dB gain in scale-invariant signal-to-distortion ratio (SI-SDR) over the MVDR and remains effective even with a single microphone. Code is available at https://github.com/Screeen/cMVDR.
【4】Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing
标题:用于鲁棒音频信号处理的自适应单通道能量归一化前端
链接:https://arxiv.org/abs/2510.18206
备注:Submitted to ICASSP2026
摘要:在音频信号处理中,可学习的前端通过优化特定于任务的表示,在不同的任务中表现出强大的性能。然而,它们的参数在训练后保持固定,在推理过程中缺乏灵活性,并且在动态复杂声学环境下限制了鲁棒性。在本文中,我们介绍了一种新的自适应模式的音频前端,取代静态参数化与闭环神经控制器。具体来说,我们简化了可学习的前端LEAF架构,并通过动态调整每通道能量归一化来集成神经控制器进行自适应表示。神经控制器利用当前和缓冲过去的子带能量,使输入相关的自适应推理过程中。多个音频分类任务的实验结果表明,所提出的自适应前端始终优于以前的固定和可学习的前端在干净和复杂的声学条件下。这些结果突出了神经适应性作为下一代音频前端的一个有前途的方向。
摘要:In audio signal processing, learnable front-ends have shown strong performance across diverse tasks by optimizing task-specific representation. However, their parameters remain fixed once trained, lacking flexibility during inference and limiting robustness under dynamic complex acoustic environments. In this paper, we introduce a novel adaptive paradigm for audio front-ends that replaces static parameterization with a closed-loop neural controller. Specifically, we simplify the learnable front-end LEAF architecture and integrate a neural controller for adaptive representation via dynamically tuning Per-Channel Energy Normalization. The neural controller leverages both the current and the buffered past subband energies to enable input-dependent adaptation during inference. Experimental results on multiple audio classification tasks demonstrate that the proposed adaptive front-end consistently outperforms prior fixed and learnable front-ends under both clean and complex acoustic conditions. These results highlight neural adaptability as a promising direction for the next generation of audio front-ends.
【5】Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network
标题:基于多任务多尺度网络的钢琴动力学与韵律结构联合估计
链接:https://arxiv.org/abs/2510.18190
备注:Paper submitted to ICASSP2026
摘要:从音频记录中估计钢琴动态是计算音乐分析中的一个基本挑战。在本文中,我们提出了一个高效的多任务网络,可以从共享的潜在表示中联合预测动态水平、变化点、节拍和强拍。这四个目标构成了乐谱中力度的韵律结构。受近年来声乐动力学研究的启发,我们采用多尺度网络作为主干,以树皮尺度特定响度作为输入特征。与作为输入的log-Mel相比,这将模型大小从14.7 M减少到0.5 M,从而实现长序列输入。我们在音频分割中使用60秒的音频长度,这是通常使用的节拍跟踪长度的两倍。在公共MazurkaBL数据集上进行评估,我们的模型在所有任务中都获得了最先进的结果。这项工作为钢琴动态估计设定了一个新的基准,并提供了一个强大而紧凑的工具,为大规模,资源有效的音乐表达分析铺平了道路。
摘要:Estimating piano dynamic from audio recordings is a fundamental challenge in computational music analysis. In this paper, we propose an efficient multi-task network that jointly predicts dynamic levels, change points, beats, and downbeats from a shared latent representation. These four targets form the metrical structure of dynamics in the music score. Inspired by recent vocal dynamic research, we use a multi-scale network as the backbone, which takes Bark-scale specific loudness as the input feature. Compared to log-Mel as input, this reduces model size from 14.7 M to 0.5 M, enabling long sequential input. We use a 60-second audio length in audio segmentation, which doubled the length of beat tracking commonly used. Evaluated on the public MazurkaBL dataset, our model achieves state-of-the-art results across all tasks. This work sets a new benchmark for piano dynamic estimation and delivers a powerful and compact tool, paving the way for large-scale, resource-efficient analysis of musical expression.
【6】Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction
标题:家庭医疗保健中的听力健康:利用LLM进行疾病评分,利用ILM进行声乐生物标志物提取
链接:https://arxiv.org/abs/2510.18169
备注:The Second Workshop on GenAI for Health at NeurIPS 2025
摘要:对家庭医疗保健的需求不断增长,需要能够支持医疗服务的工具。在这项研究中,我们使用真实世界的家庭护理访问数据,利用其包含的各种患者信息,探索语音自动健康评估。首先,我们利用大型语言模型(LLM)将来自非结构化音频转录和结构化生命体征的主观,客观,评估和计划(SOAP)笔记整合到反映患者整体健康状况的整体疾病评分中。这种紧凑的表示有利于跨访视健康状况比较和下游分析。接下来,我们设计了一个多阶段的预处理管道,从家庭护理录音中的目标扬声器中提取短语音片段进行声学分析。然后,我们采用音频语言模型(ALM)来产生声音生物标志物的简单语言描述,并检查它们与个人健康状况的关联。我们的实验结果在估计疾病评分方面对商业和开源LLM进行了基准测试,证明了它们与实际临床结果的一致性,并揭示了SOAP笔记比生命体征信息量更大。在疾病评分的基础上,我们提供了第一个证据,即ALMs可以从家庭护理记录中识别与健康相关的声学模式,并以人类可读的形式呈现它们。总之,这些研究结果突出了LLM和ALM利用异构家庭访视数据进行更好的患者监测和护理的潜力。
摘要:The growing demand for home healthcare calls for tools that can support care delivery. In this study, we explore automatic health assessment from voice using real-world home care visit data, leveraging the diverse patient information it contains. First, we utilize Large Language Models (LLMs) to integrate Subjective, Objective, Assessment, and Plan (SOAP) notes derived from unstructured audio transcripts and structured vital signs into a holistic illness score that reflects a patient's overall health. This compact representation facilitates cross-visit health status comparisons and downstream analysis. Next, we design a multi-stage preprocessing pipeline to extract short speech segments from target speakers in home care recordings for acoustic analysis. We then employ an Audio Language Model (ALM) to produce plain-language descriptions of vocal biomarkers and examine their association with individuals' health status. Our experimental results benchmark both commercial and open-source LLMs in estimating illness scores, demonstrating their alignment with actual clinical outcomes, and revealing that SOAP notes are substantially more informative than vital signs. Building on the illness scores, we provide the first evidence that ALMs can identify health-related acoustic patterns from home care recordings and present them in a human-readable form. Together, these findings highlight the potential of LLMs and ALMs to harness heterogeneous in-home visit data for better patient monitoring and care.
【7】Adapting Language Balance in Code-Switching Speech
标题:码转换语音中的语言平衡调整
链接:https://arxiv.org/abs/2510.18724
备注:Submitted to ICASSP 2026
摘要:尽管在标准基准测试中取得了令人印象深刻的结果,但大型基础模型仍然难以应对代码切换测试用例。当数据稀缺不能被用作表现不佳的通常理由时,原因可能在于语码转换时刻很少发生,第二语言的嵌入微妙地出现。与其期望模型自己学习这种不频繁的情况,不如为训练过程提供标签。评估模型性能的代码转换数据需要仔细定位的代码转换点,识别错误是最重要的,使分析强调在这些时刻发生的错误。基于这一观察,我们利用嵌入式语言和主语言之间的差异来突出这些代码转换点,从而强调在这些位置的学习。这种简单而有效的可区分替代物减轻了生成过程中的上下文偏差-代码转换的核心挑战-从而提高了模型的鲁棒性。我们对阿拉伯语和汉英双语的实验表明,该模型能够更正确地预测转换位置,这反映在减少的替换错误上。
摘要:Despite achieving impressive results on standard benchmarks, large foundational models still struggle against code-switching test cases. When data scarcity cannot be used as the usual justification for poor performance, the reason may lie in the infrequent occurrence of code-switched moments, where the embedding of the second language appears subtly. Instead of expecting the models to learn this infrequency on their own, it might be beneficial to provide the training process with labels. Evaluating model performance on code-switching data requires careful localization of code-switching points where recognition errors are most consequential, so that the analysis emphasizes mistakes occurring at those moments. Building on this observation, we leverage the difference between the embedded and the main language to highlight those code-switching points and thereby emphasize learning at those locations. This simple yet effective differentiable surrogate mitigates context bias during generation -- the central challenge in code-switching -- thereby improving the model's robustness. Our experiments with Arabic and Chinese-English showed that the models are able to predict the switching places more correctly, reflected by the reduced substitution error.
【8】Bayesian Low-Rank Factorization for Robust Model Adaptation
标题:用于鲁棒模型自适应的Bayesian低等级因子分解
链接:https://arxiv.org/abs/2510.18723
备注:Submitted to ICASSP 2026
摘要:大型语音基础模型在许多领域都具有很强的性能,但它们通常需要适应以处理本地需求,例如语码转换,即说话者在同一话语中混合语言。直接对这些模型进行微调会冒着过度拟合目标域的风险,并破坏基础模型的广泛功能。为了解决这一挑战,我们探索贝叶斯分解适配器的语音基础模型,将先验接近零,以实现稀疏的适应矩阵,从而保持一般性能,同时适应特定的领域。我们将我们的方法应用到Whisper模型中,并对不同的多语言代码切换场景进行评估。我们的研究结果表明,只有最小的适应损失,同时显着减少灾难性遗忘的基础模型。与LoRA相比,我们的方法实现了54%的后向增益,在新域上仅下降4%。这些发现突出了贝叶斯适应的有效性,微调语音基础模型,而不牺牲泛化。
摘要:Large speech foundation models achieve strong performance across many domains, but they often require adaptation to handle local needs such as code-switching, where speakers mix languages within the same utterance. Direct fine-tuning of these models risks overfitting to the target domain and overwriting the broad capabilities of the base model. To address this challenge, we explore Bayesian factorized adapters for speech foundation models, which place priors near zero to achieve sparser adaptation matrices and thereby retain general performance while adapting to specific domains. We apply our approach to the Whisper model and evaluate on different multilingual code-switching scenarios. Our results show only minimal adaptation loss while significantly reducing catastrophic forgetting of the base model. Compared to LoRA, our method achieves a backward gain of 54% with only a 4% drop on the new domain. These findings highlight the effectiveness of Bayesian adaptation for fine-tuning speech foundation models without sacrificing generalization.
【9】Noise-Conditioned Mixture-of-Experts Framework for Robust Speaker Verification
标题:用于鲁棒说话人验证的噪音条件专家混合框架
链接:https://arxiv.org/abs/2510.18533
摘要:噪声条件下的鲁棒说话人确认仍然是一个公开的挑战。传统的深度学习方法可以针对不同的背景噪声学习鲁棒的统一说话人表示空间,并取得显着的改进。与此相反,本文提出了一个噪声条件下的混合专家框架,将特征空间分解成专门的噪声感知子空间的说话人验证。具体来说,我们提出了一个噪声条件下的专家路由机制,一个通用的模型为基础的专家专业化策略,和SNR衰减的课程学习协议,共同提高模型的鲁棒性和泛化在不同的噪声条件下。所提出的方法可以自动路由输入到专家网络的基础上从输入,其中每个专家的目标不同的噪声特性,同时保留扬声器的身份信息的噪声信息。全面的实验表明,一致优于基线,确认显式噪声相关的特征建模显着提高鲁棒性,而不牺牲验证精度。
摘要:Robust speaker verification under noisy conditions remains an open challenge. Conventional deep learning methods learn a robust unified speaker representation space against diverse background noise and achieve significant improvement. In contrast, this paper presents a noise-conditioned mixture-ofexperts framework that decomposes the feature space into specialized noise-aware subspaces for speaker verification. Specifically, we propose a noise-conditioned expert routing mechanism, a universal model based expert specialization strategy, and an SNR-decaying curriculum learning protocol, collectively improving model robustness and generalization under diverse noise conditions. The proposed method can automatically route inputs to expert networks based on noise information derived from the inputs, where each expert targets distinct noise characteristics while preserving speaker identity information. Comprehensive experiments demonstrate consistent superiority over baselines, confirming that explicit noise-dependent feature modeling significantly enhances robustness without sacrificing verification accuracy.
【10】A Stage-Wise Learning Strategy with Fixed Anchors for Robust Speaker Verification
标题:用于鲁棒说话人验证的固定条件分阶段学习策略
链接:https://arxiv.org/abs/2510.18530
摘要:在噪声条件下学习鲁棒的说话人表示提出了重大挑战,这需要仔细处理的歧视性和噪声不变的属性。在这项工作中,我们提出了一个基于锚点的阶段式学习策略,用于鲁棒的说话人表示学习。具体来说,我们的方法首先训练一个基础模型来建立区分性的说话人边界,然后从这个模型中提取锚嵌入作为稳定的参考。最后,基础模型的副本在有噪声的输入上进行微调,通过强制接近其相应的固定锚嵌入来正则化,以在失真下保留说话人身份。实验结果表明,这种策略提供了传统的联合优化的优势,特别是在保持歧视,同时提高噪声鲁棒性。所提出的方法在各种噪声条件下表现出一致的改进,可能是由于其能够单独处理边界稳定和变化抑制。
摘要:Learning robust speaker representations under noisy conditions presents significant challenges, which requires careful handling of both discriminative and noise-invariant properties. In this work, we proposed an anchor-based stage-wise learning strategy for robust speaker representation learning. Specifically, our approach begins by training a base model to establish discriminative speaker boundaries, and then extract anchor embeddings from this model as stable references. Finally, a copy of the base model is fine-tuned on noisy inputs, regularized by enforcing proximity to their corresponding fixed anchor embeddings to preserve speaker identity under distortion. Experimental results suggest that this strategy offers advantages over conventional joint optimization, particularly in maintaining discrimination while improving noise robustness. The proposed method demonstrates consistent improvements across various noise conditions, potentially due to its ability to handle boundary stabilization and variation suppression separately.
【11】ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
标题:ParaStyleTTC:实现高效且稳健的副语言风格控制,以实现表达性文本转语音生成
链接:https://arxiv.org/abs/2510.18308
摘要:在文语转换(TTS)系统中控制说话风格已经成为学术界和工业界日益关注的焦点。虽然许多现有的方法依赖于参考音频来指导样式生成,但是由于隐私问题和有限的可访问性,这样的方法通常是不切实际的。最近,大型语言模型(LLM)已被用于控制说话风格,通过自然语言提示,然而,其高计算成本,缺乏可解释性,并提示措辞的敏感性限制了其在实时和资源受限的环境中的适用性。在这项工作中,我们提出了ParaStyleTTS,一个轻量级的和可解释的TTS框架,使表达风格控制从文本提示。ParaStyleTTS的特点是一种新颖的两级风格适应架构,分离韵律和非语言学的语音风格建模。它允许对情绪,性别和年龄等因素进行细粒度和鲁棒的控制。与基于LLM的方法不同,ParaStyleTTS在不同的提示公式中保持一致的风格实现,非常适合现实世界的应用程序,包括设备上和低资源部署。实验结果表明,ParaStyleTTS生成高质量的语音,其性能与最先进的基于LLM的系统相当,同时速度快30倍,使用的参数少8倍,需要的CUDA内存少2.5倍。此外,ParaStyleTTS表现出优越的鲁棒性和可控性,超过非语言的说话风格,提供了一个实用和有效的解决方案,风格可控的文本到语音生成。演示可以在https://parastyletts.github.io/ParaStyleTTS_Demo/上找到。代码可在https://github.com/haoweilou/ParaStyleTTS上找到。
摘要:Controlling speaking style in text-to-speech (TTS) systems has become a growing focus in both academia and industry. While many existing approaches rely on reference audio to guide style generation, such methods are often impractical due to privacy concerns and limited accessibility. More recently, large language models (LLMs) have been used to control speaking style through natural language prompts; however, their high computational cost, lack of interpretability, and sensitivity to prompt phrasing limit their applicability in real-time and resource-constrained environments. In this work, we propose ParaStyleTTS, a lightweight and interpretable TTS framework that enables expressive style control from text prompts alone. ParaStyleTTS features a novel two-level style adaptation architecture that separates prosodic and paralinguistic speech style modeling. It allows fine-grained and robust control over factors such as emotion, gender, and age. Unlike LLM-based methods, ParaStyleTTS maintains consistent style realization across varied prompt formulations and is well-suited for real-world applications, including on-device and low-resource deployment. Experimental results show that ParaStyleTTS generates high-quality speech with performance comparable to state-of-the-art LLM-based systems while being 30x faster, using 8x fewer parameters, and requiring 2.5x less CUDA memory. Moreover, ParaStyleTTS exhibits superior robustness and controllability over paralinguistic speaking styles, providing a practical and efficient solution for style-controllable text-to-speech generation. Demo can be found at https://parastyletts.github.io/ParaStyleTTS_Demo/. Code can be found at https://github.com/haoweilou/ParaStyleTTS.
【12】Transformer Redesign for Late Fusion of Audio-Text Features on Ultra-Low-Power Edge Hardware
标题:Transformer重新设计,以在超低功耗边缘硬件上实现音频文本功能的后期融合
链接:https://arxiv.org/abs/2510.18036
摘要:在设备必须小巧、低功耗和私密的现实环境中部署情感识别系统仍然是一个重大挑战。这对于诸如紧张监控、冲突降级和响应式可穿戴设备等应用尤其重要,在这些应用中,基于云的解决方案不切实际。多模态情感识别通过深度学习取得了进展,但大多数系统仍然不适合部署在超受限的边缘设备上。以前的工作通常依赖于强大的硬件,缺乏实时性能,或使用单峰输入。本文通过提出一种硬件感知的情感识别系统来解决这一差距,该系统使用针对Edge TPU优化的后期融合架构将声学和语言特征相结合。该设计将基于量化变换器的声学模型与DSResNet-SE网络中的冻结关键字嵌入集成在一起,从而在1.8MB内存预算和21- 23 ms延迟内实现实时推理。该管道使用MicroFrontend和MLTK确保训练和部署之间的频谱图对齐。对通过Coral Dev Board Micro麦克风捕获的重新记录的分段IEMOCAP样本进行评估,结果显示与单峰基线相比,宏观F1改善了6.3%。这项工作表明,通过特定于任务的融合和硬件引导的模型设计,可以在微处理器级边缘平台上实现准确,实时的多模态情感推理。
摘要:Deploying emotion recognition systems in real-world environments where devices must be small, low-power, and private remains a significant challenge. This is especially relevant for applications such as tension monitoring, conflict de-escalation, and responsive wearables, where cloud-based solutions are impractical. Multimodal emotion recognition has advanced through deep learning, but most systems remain unsuitable for deployment on ultra-constrained edge devices. Prior work typically relies on powerful hardware, lacks real-time performance, or uses unimodal input. This paper addresses that gap by presenting a hardware-aware emotion recognition system that combines acoustic and linguistic features using a late-fusion architecture optimised for Edge TPU. The design integrates a quantised transformer-based acoustic model with frozen keyword embeddings from a DSResNet-SE network, enabling real-time inference within a 1.8MB memory budget and 21-23ms latency. The pipeline ensures spectrogram alignment between training and deployment using MicroFrontend and MLTK. Evaluation on re-recorded, segmented IEMOCAP samples captured through the Coral Dev Board Micro microphone shows a 6.3% macro F1 improvement over unimodal baselines. This work demonstrates that accurate, real-time multimodal emotion inference is achievable on microcontroller-class edge platforms through task-specific fusion and hardware-guided model design.
机器翻译由腾讯交互翻译提供,仅供参考
