微信公众号:arXiv_Daily
cs.SD语音
【1】Towards explainable reference-free speech intelligibility evaluation of people with pathological speech
标题:对病态言语患者进行可解释的无参考言语清晰度评估
链接:https://arxiv.org/abs/2602.12723
备注:Plan to be submited to Interspeech 2026
摘要:客观评估言语,反映有意义的变化,沟通是至关重要的临床决策和可重复的研究。虽然现有的客观评估,特别是基于参考的方法,可以捕捉可理解性的变化,他们往往是缺乏可解释性和劳动密集型的手动transmittance的需要阻碍。为了解决这些问题,这项工作提出了无参考,可解释的ASR不一致性分数。我们评估这种方法在荷兰语,西班牙语和英语的病理语音,并比较其性能的参考为基础的单词错误率(WER)的基线。我们的研究结果表明,ASR不一致性分数实现了与专家感知评级的高度相关性,性能密切匹配,并在一种情况下超过了标准的基于参考的单词错误率(WER)基线。
摘要:Objective assessment of speech that reflects meaningful changes in communication is crucial for clinical decision making and reproducible research. While existing objective assessments, particularly reference-based approaches, can capture intelligibility changes, they are often hindered by lack of explainability and the need for labor-intensive manual transcriptions. To address these issues, this work proposes the reference-free, explainable ASR Inconsistency Score. We evaluate this method on pathological speech in Dutch, Spanish and English, and compare its performance to a reference-based Word Error Rate (WER) baseline. Our results demonstrate that the ASR Inconsistency Score achieves a high correlation with expert perceptual ratings, with performance closely matching, and in one case exceeding, a standard reference-based Word Error Rate (WER) baseline.
【2】DisSR: Disentangling Speech Representation for Degradation-Prior Guided Cross-Domain Speech Restoration
标题:DisSR:分解语音表示以实现降级优先引导的跨域语音恢复
链接:https://arxiv.org/abs/2602.12701
备注:Accepted to 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)
摘要:以往的语音恢复(SR)主要集中在单任务语音恢复(SSR),它不能解决一般的语音恢复问题。针对不同失真训练特定SSR模型是耗时的并且缺乏通用性。此外,大多数研究忽略了模型在未知领域的泛化问题。为了克服这些局限性,我们提出了DisSR,一个基于解纠缠语音表示的通用语音恢复模型,具有两个特性:1)退化先验指导,它提取说话人不变的退化表示来指导基于扩散的语音恢复模型。2)领域自适应,我们设计了跨领域对齐训练,以增强模型对跨领域数据的适应性和泛化能力。实验结果表明,我们的方法可以在各种失真条件下产生高质量的恢复语音。音频样本可以在https://itspsp.github.io/DisSR上找到。
摘要:Previous speech restoration (SR) primarily focuses on single-task speech restoration (SSR), which cannot address general speech restoration problems. Training specific SSR models for different distortions is time-consuming and lacks generality. In addition, most studies ignore the problem of model generalization across unseen domains. To overcome those limitations, we propose DisSR, a Disentangling Speech Representation based general speech restoration model with two properties: 1) Degradation-prior guidance, which extracts speaker-invariant degradation representation to guide the diffusion-based speech restoration model. 2) Domain adaptation, where we design cross-domain alignment training to enhance the model's adaptability and generalization on cross-domain data, respectively. Experimental results demonstrate that our method can produce high-quality restored speech under various distortion conditions. Audio samples can be found at https://itspsp.github.io/DisSR.
【3】OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
标题:OmniCustoms:通过联合音视频生成模型同步音视频定制
链接:https://arxiv.org/abs/2602.12304
备注:16 pages
摘要:现有主流的视频定制方法集中于基于给定的参考图像和文本提示生成身份一致的视频。得益于音视频联合生成技术的快速发展,本文提出了一个更加引人注目的新课题:音视频同步定制,旨在同步定制视频标识和音频音色。具体来说,给定参考图像$I^{r}$和参考音频$A^{r}$,这个新颖的任务需要生成保持参考图像的身份同时模仿参考音频的音色的视频,其中口语内容可以通过用户自由指定提供的文本提示。为此,我们提出了OmniCustom,一个功能强大的基于DiT的音视频定制框架,可以合成一个视频以下的参考图像身份,音频音色,和文本提示一次在一个zero-shot的方式。我们的框架建立在三个关键贡献之上。首先,身份和音频音色控制是通过独立的参考身份和音频LoRA模块实现的,这些模块通过基础音频-视频生成模型内的自注意层进行操作。其次,我们引入了一个对比学习目标旁边的标准流匹配目标。它使用的参考输入条件下的预测流作为正面的例子和那些没有参考条件作为负面的例子,从而提高模型的能力,以保持身份和音色。第三,我们训练OmniCustom在我们构建的大规模,高质量的视听人类数据集。大量的实验表明,OmniCustom在生成具有一致身份和音色保真度的音频视频内容方面优于现有方法。
摘要:Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, this paper proposes a more compelling new task: sync audio-video customization, which aims to synchronously customize both video identity and audio timbre. Specifically, given a reference image $I^{r}$ and a reference audio $A^{r}$, this novel task requires generating videos that maintain the identity of the reference image while imitating the timbre of the reference audio, with spoken content freely specifiable through user-provided textual prompts. To this end, we propose OmniCustom, a powerful DiT-based audio-video customization framework that can synthesize a video following reference image identity, audio timbre, and text prompts all at once in a zero-shot manner. Our framework is built on three key contributions. First, identity and audio timbre control are achieved through separate reference identity and audio LoRA modules that operate through self-attention layers within the base audio-video generation model. Second, we introduce a contrastive learning objective alongside the standard flow matching objective. It uses predicted flows conditioned on reference inputs as positive examples and those without reference conditions as negative examples, thereby enhancing the model ability to preserve identity and timbre. Third, we train OmniCustom on our constructed large-scale, high-quality audio-visual human dataset. Extensive experiments demonstrate that OmniCustom outperforms existing methods in generating audio-video content with consistent identity and timbre fidelity.
【4】Beyond Musical Descriptors: Extracting Preference-Bearing Intent in Music Queries
标题:超越音乐描述符:提取音乐预设中的偏好意图
链接:https://arxiv.org/abs/2602.12301
备注:Accepted at NLP4MusA 2026 (4th Workshop on NLP for Music and Audio)
摘要:虽然用于用户查询的带注释的音乐描述符数据集越来越常见,但很少考虑这些描述符背后的用户意图,这对于有效满足他们的需求至关重要。我们介绍了MusicallyIntent,这是一个包含2,291个Reddit音乐请求的手动注释语料库,将音乐描述符标记为七个类别,具有积极,消极或参考偏好承载角色。然后,我们研究了大型语言模型(LLM)如何可靠地提取这些音乐描述符,发现它们确实捕获了明确的描述符,但与上下文相关的描述符有关。这项工作可以进一步作为用户意图的细粒度建模的基准,并获得改善基于LLM的音乐理解系统的见解。
摘要:Although annotated music descriptor datasets for user queries are increasingly common, few consider the user's intent behind these descriptors, which is essential for effectively meeting their needs. We introduce MusicRecoIntent, a manually annotated corpus of 2,291 Reddit music requests, labeling musical descriptors across seven categories with positive, negative, or referential preference-bearing roles. We then investigate how reliably large language models (LLMs) can extract these music descriptors, finding that they do capture explicit descriptors but struggle with context-dependent ones. This work can further serve as a benchmark for fine-grained modeling of user intent and for gaining insights into improving LLM-based music understanding systems.
【5】A two-step approach for speech enhancement in low-SNR scenarios using cyclostationary beamforming and DNNs
标题:使用循环平稳射束和DNN在低SNR场景中进行语音增强的两步方法
链接:https://arxiv.org/abs/2602.12986
备注:Submitted version
摘要:深度神经网络(DNN)通常难以在低信噪比(SNR)下抑制噪声。本文讨论了在谐波噪声占主导地位的情况下的语音增强,并提出了一个框架,该框架集成了循环平稳性感知预处理与轻量级的基于DNN的去噪。循环最小功率无失真响应(CMPDR)频谱波束形成器被用作预处理块。它利用循环平稳噪声的频谱相关性来抑制谐波分量,然后进行基于学习的增强,并且不需要修改DNN架构。建议的管道使用两种DNN架构在单通道设置中进行评估:一种是简单而轻量级的卷积递归神经网络(CRNN),另一种是最先进的模型,即超低复杂度网络(ULCNet)。对旋转机械噪声主导的合成数据和真实世界记录的实验表明,端到端DNN基线的一致改进,特别是在低SNR下。值得注意的是,具有cMPDR预处理的参数高效CRNN超过了在原始或维纳滤波输入上操作的较大ULCNet的性能。这些结果表明,显式地将循环平稳性作为信号先验比单独增加模型容量更有效地抑制谐波干扰。
摘要:Deep Neural Networks (DNNs) often struggle to suppress noise at low signal-to-noise ratios (SNRs). This paper addresses speech enhancement in scenarios dominated by harmonic noise and proposes a framework that integrates cyclostationarity-aware preprocessing with lightweight DNN-based denoising. A cyclic minimum power distortionless response (cMPDR) spectral beamformer is used as a preprocessing block. It exploits the spectral correlations of cyclostationary noise to suppress harmonic components prior to learning-based enhancement and does not require modifications to the DNN architecture. The proposed pipeline is evaluated in a single-channel setting using two DNN architectures: a simple and lightweight convolutional recurrent neural network (CRNN), and a state-of-the-art model, namely ultra-low complexity network (ULCNet). Experiments on synthetic data and real-world recordings dominated by rotating machinery noise demonstrate consistent improvements over end-to-end DNN baselines, particularly at low SNRs. Remarkably, a parameter-efficient CRNN with cMPDR preprocessing surpasses the performance of the larger ULCNet operating on raw or Wiener-filtered inputs. These results indicate that explicitly incorporating cyclostationarity as a signal prior is more effective than increasing model capacity alone for suppressing harmonic interference.
【6】A Wavefield Correlation Approach to Improve Sound Speed Estimation in Ultrasound Autofocusing
标题:一种改进超声自聚焦声速估计的波场相关方法
链接:https://arxiv.org/abs/2602.12805
摘要:当波束形成不考虑波前失真时,像差通常会降低超声图像质量。在过去的十年中,局部声速估计器已被开发用于整个介质中的分布畸变校正。最近,迭代声速优化方法已经实现了比早期方法更准确的估计,但这些较新的方法仍然与具有混响杂波和大声速变化的媒体的降低的准确性作斗争。为了解决这些挑战,我们建议使用波场相关(WFC)波束成形方法时,执行声速优化。WFC使模拟的前向传播的发射波场和后向传播的接收波场相关,以便形成图像。该过程更准确地模拟了波在非均匀介质中的传播,并且由于其时空匹配滤波效应可以减少漫射杂波。该波束形成器使用自动微分软件来实现,然后在波束形成期间使用的局部声速图上使用总变差正则化的公共中点相位聚焦度量损失来执行梯度下降优化。这种方法相比,使用延迟和总和(DAS)与直射线时间延迟计算在相同的声速优化方法上的各种模拟,幻影,并在体内数据与大的声速变化和杂波。结果表明,使用WFC减小声速估计误差,并使用估计的像差校正提高图像的分辨率和对比度。这些有希望的结果有可能改善具有挑战性的临床场景的脉冲回波成像。
摘要:Aberration often degrades ultrasound image quality when beamforming does not account for wavefront distortions. In the past decade, local sound speed estimators have been developed for distributed aberration correction throughout a medium. Recently, iterative sound speed optimization approaches have achieved more accurate estimates than earlier approaches, but these newer methods still struggle with decreased accuracy for media with reverberation clutter and large sound speed changes. To address these challenges, we propose using a wavefield correlation (WFC) beamforming approach when performing sound speed optimization. WFC correlates simulated forward-propagated transmit wavefields and backwards-propagated receive wavefields in order to form images. This process more accurately models wave propagation in heterogeneous media and can decrease diffuse clutter due to its spatiotemporal matched filtering effect. This beamformer is implemented using auto-differentiation software to then perform gradient descent optimization, using a total-variation regularized common midpoint phase focus metric loss, on the local sound speed map used during beamforming. This approach is compared to using delay and sum (DAS) with straight-ray time delay calculations in the same sound speed optimization approach on a variety of simulated, phantom, and in vivo data with large sound speed changes and clutter. Results show that using WFC decreases sound speed estimation error, and using the estimates for aberration correction improves image resolution and contrast. These promising results have potential to improve pulse-echo imaging for challenging clinical scenarios.
【7】Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
标题:具有模式感知的ASB专家稀疏混合的仅解码器一致器
链接:https://arxiv.org/abs/2602.12546
备注:Accepted to ICASSP 2026
摘要:我们提出了一种用于自动语音识别(ASR)的仅解码器的Conformer,该Conformer在单个堆栈中处理语音和文本,而无需外部语音编码器或预训练的大型语言模型(LLM)。该模型使用模态感知稀疏混合专家(MoE):不相交的专家池的语音和文本硬路由和前1选择,嵌入在混合因果一致性块(双向语音,因果文本)。训练将语音位置上的CTC与用于文本生成的标签平滑交叉熵相结合。我们的113 M参数模型在Librispeech的139 M AED基线上持续改善WER(2.8% vs. 3.2%测试-清洁; 5.6% vs. 6.0%测试-其他)。在Common Voice 16.1上,使用跨五种语言的单一多语言模型,我们的方法将平均WER从12.2%降低到10.6%。据我们所知,这是第一个随机初始化的仅解码器ASR,其通过模态感知路由和稀疏MoE超过强AED基线,以更少的活动参数实现更好的准确性,并且没有对齐/自适应模块。
摘要:We present a decoder-only Conformer for automatic speech recognition (ASR) that processes speech and text in a single stack without external speech encoders or pretrained large language models (LLM). The model uses a modality-aware sparse mixture of experts (MoE): disjoint expert pools for speech and text with hard routing and top-1 selection, embedded in hybrid-causality Conformer blocks (bidirectional for speech, causal for text). Training combines CTC on speech positions with label-smoothed cross-entropy for text generation. Our 113M-parameter model consistently improves WER over a 139M AED baseline on Librispeech (2.8% vs. 3.2% test-clean; 5.6% vs. 6.0% test-other). On Common Voice 16.1 with a single multilingual model across five languages, our approach reduces average WER from 12.2% to 10.6%. To our knowledge, this is the first randomly initialized decoder-only ASR that surpasses strong AED baselines via modality-aware routing and sparse MoE, achieving better accuracy with fewer active parameters and without alignment/adaptation modules.
【8】Acoustivision Pro: An Open-Source Interactive Platform for Room Impulse Response Analysis and Acoustic Characterization
标题:Acoustivision Pro:一个用于房间脉冲响应分析和声学特性的开源交互平台
链接:https://arxiv.org/abs/2602.12299
摘要:室内声学分析在建筑设计、音频工程、语音清晰度评估和听力研究中起着核心作用。尽管混响时间、清晰度和语音传输指数等标准化指标可用,但将严格的信号处理与直观的可视化相结合的可用工具仍然很少。本文介绍了AcoustiVision Pro,一个开源的基于Web的平台,全面的房间脉冲响应(RIR)分析。该系统从上传或上传的RIR中计算出12个不同的声学参数,提供早期反射的交互式3D可视化,通过瀑布图生成频率相关的衰减特性,并检查是否符合ANSI S12.60和ISO 3382等国际标准。我们介绍了随附的RIRMega和RIRMega Speech数据集,这些数据集托管在Hugging Face上,包含数千个模拟的房间脉冲响应和完整的元数据。该平台通过基于FFT的卷积支持实时可听化,导出适合工程文档的详细PDF报告,并提供CSV数据导出以供进一步分析。我们描述了每个声学指标的数学基础,详细介绍了系统架构,并提出了初步的案例研究,展示了该平台在不同应用领域的实用性,包括教室声学,医疗设施设计和录音室评估。
摘要:Room acoustics analysis plays a central role in architectural design, audio engineering, speech intelligibility assessment, and hearing research. Despite the availability of standardized metrics such as reverberation time, clarity, and speech transmission index, accessible tools that combine rigorous signal processing with intuitive visualization remain scarce. This paper presents AcoustiVision Pro, an open-source web-based platform for comprehensive room impulse response (RIR) analysis. The system computes twelve distinct acoustic parameters from uploaded or dataset-sourced RIRs, provides interactive 3D visualizations of early reflections, generates frequency-dependent decay characteristics through waterfall plots, and checks compliance against international standards including ANSI S12.60 and ISO 3382. We introduce the accompanying RIRMega and RIRMega Speech datasets hosted on Hugging Face, containing thousands of simulated room impulse responses with full metadata. The platform supports real-time auralization through FFT-based convolution, exports detailed PDF reports suitable for engineering documentation, and provides CSV data export for further analysis. We describe the mathematical foundations underlying each acoustic metric, detail the system architecture, and present preliminary case studies demonstrating the platform's utility across diverse application domains including classroom acoustics, healthcare facility design, and recording studio evaluation.
【1】A two-step approach for speech enhancement in low-SNR scenarios using cyclostationary beamforming and DNNs
标题:使用循环平稳射束和DNN在低SNR场景中进行语音增强的两步方法
链接:https://arxiv.org/abs/2602.12986
备注:Submitted version
摘要:深度神经网络(DNN)通常难以在低信噪比(SNR)下抑制噪声。本文讨论了在谐波噪声占主导地位的情况下的语音增强,并提出了一个框架,该框架集成了循环平稳性感知预处理与轻量级的基于DNN的去噪。循环最小功率无失真响应(CMPDR)频谱波束形成器被用作预处理块。它利用循环平稳噪声的频谱相关性来抑制谐波分量,然后进行基于学习的增强,并且不需要修改DNN架构。建议的管道使用两种DNN架构在单通道设置中进行评估:一种是简单而轻量级的卷积递归神经网络(CRNN),另一种是最先进的模型,即超低复杂度网络(ULCNet)。对以旋转机械噪音为主的合成数据和真实世界录音进行的实验表明,与端到端DNN基线相比,尤其是在低SNR下,取得了一致的改进。值得注意的是,具有cMPDR预处理的参数高效CRNN超过了在原始或维纳滤波输入上操作的较大ULCNet的性能。这些结果表明,显式地将循环平稳性作为信号先验比单独增加模型容量更有效地抑制谐波干扰。
摘要:Deep Neural Networks (DNNs) often struggle to suppress noise at low signal-to-noise ratios (SNRs). This paper addresses speech enhancement in scenarios dominated by harmonic noise and proposes a framework that integrates cyclostationarity-aware preprocessing with lightweight DNN-based denoising. A cyclic minimum power distortionless response (cMPDR) spectral beamformer is used as a preprocessing block. It exploits the spectral correlations of cyclostationary noise to suppress harmonic components prior to learning-based enhancement and does not require modifications to the DNN architecture. The proposed pipeline is evaluated in a single-channel setting using two DNN architectures: a simple and lightweight convolutional recurrent neural network (CRNN), and a state-of-the-art model, namely ultra-low complexity network (ULCNet). Experiments on synthetic data and real-world recordings dominated by rotating machinery noise demonstrate consistent improvements over end-to-end DNN baselines, particularly at low SNRs. Remarkably, a parameter-efficient CRNN with cMPDR preprocessing surpasses the performance of the larger ULCNet operating on raw or Wiener-filtered inputs. These results indicate that explicitly incorporating cyclostationarity as a signal prior is more effective than increasing model capacity alone for suppressing harmonic interference.
【2】Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
标题:具有模式感知的ASB专家稀疏混合的仅解码器一致器
链接:https://arxiv.org/abs/2602.12546
备注:Accepted to ICASSP 2026
摘要:我们提出了一种用于自动语音识别(ASR)的仅解码器的Conformer,该Conformer在单个堆栈中处理语音和文本,而无需外部语音编码器或预训练的大型语言模型(LLM)。该模型使用模态感知稀疏混合专家(MoE):不相交的专家池的语音和文本硬路由和前1选择,嵌入在混合因果一致性块(双向语音,因果文本)。训练将语音位置上的CTC与用于文本生成的标签平滑交叉熵相结合。我们的113 M参数模型在Librispeech的139 M AED基线上持续改善WER(2.8% vs. 3.2%测试-清洁; 5.6% vs. 6.0%测试-其他)。在Common Voice 16.1上,使用跨五种语言的单一多语言模型,我们的方法将平均WER从12.2%降低到10.6%。据我们所知,这是第一个随机初始化的仅解码器ASR,其通过模态感知路由和稀疏MoE超过强AED基线,以更少的活动参数实现更好的准确性,并且没有对齐/自适应模块。
摘要:We present a decoder-only Conformer for automatic speech recognition (ASR) that processes speech and text in a single stack without external speech encoders or pretrained large language models (LLM). The model uses a modality-aware sparse mixture of experts (MoE): disjoint expert pools for speech and text with hard routing and top-1 selection, embedded in hybrid-causality Conformer blocks (bidirectional for speech, causal for text). Training combines CTC on speech positions with label-smoothed cross-entropy for text generation. Our 113M-parameter model consistently improves WER over a 139M AED baseline on Librispeech (2.8% vs. 3.2% test-clean; 5.6% vs. 6.0% test-other). On Common Voice 16.1 with a single multilingual model across five languages, our approach reduces average WER from 12.2% to 10.6%. To our knowledge, this is the first randomly initialized decoder-only ASR that surpasses strong AED baselines via modality-aware routing and sparse MoE, achieving better accuracy with fewer active parameters and without alignment/adaptation modules.
【3】Acoustivision Pro: An Open-Source Interactive Platform for Room Impulse Response Analysis and Acoustic Characterization
标题:Acoustivision Pro:一个用于房间脉冲响应分析和声学特性的开源交互平台
链接:https://arxiv.org/abs/2602.12299
摘要:房间声学分析在建筑设计、音频工程、语音清晰度评估和听力研究中发挥着核心作用。尽管混响时间、清晰度和语音传输指数等标准化指标可用,但将严格的信号处理与直观的可视化相结合的可用工具仍然很少。本文介绍了AcoustiVision Pro,一个开源的基于Web的平台,全面的房间脉冲响应(RIR)分析。该系统从上传或上传的RIR中计算出12个不同的声学参数,提供早期反射的交互式3D可视化,通过瀑布图生成频率相关的衰减特性,并检查是否符合ANSI S12.60和ISO 3382等国际标准。我们介绍了随附的RIRMega和RIRMega Speech数据集,这些数据集托管在Hugging Face上,包含数千个模拟的房间脉冲响应和完整的元数据。该平台通过基于FFT的卷积支持实时可听化,导出适合工程文档的详细PDF报告,并提供CSV数据导出以供进一步分析。我们描述了每个声学指标的数学基础,详细介绍了系统架构,并提出了初步的案例研究,展示了该平台在不同应用领域的实用性,包括教室声学,医疗设施设计和录音室评估。
摘要:Room acoustics analysis plays a central role in architectural design, audio engineering, speech intelligibility assessment, and hearing research. Despite the availability of standardized metrics such as reverberation time, clarity, and speech transmission index, accessible tools that combine rigorous signal processing with intuitive visualization remain scarce. This paper presents AcoustiVision Pro, an open-source web-based platform for comprehensive room impulse response (RIR) analysis. The system computes twelve distinct acoustic parameters from uploaded or dataset-sourced RIRs, provides interactive 3D visualizations of early reflections, generates frequency-dependent decay characteristics through waterfall plots, and checks compliance against international standards including ANSI S12.60 and ISO 3382. We introduce the accompanying RIRMega and RIRMega Speech datasets hosted on Hugging Face, containing thousands of simulated room impulse responses with full metadata. The platform supports real-time auralization through FFT-based convolution, exports detailed PDF reports suitable for engineering documentation, and provides CSV data export for further analysis. We describe the mathematical foundations underlying each acoustic metric, detail the system architecture, and present preliminary case studies demonstrating the platform's utility across diverse application domains including classroom acoustics, healthcare facility design, and recording studio evaluation.
【4】Lamer-SSL: Layer-aware Mixture of LoRA Experts for Continual Multilingual Expansion of Self-supervised Models without Forgetting
标题:Lamer-SSL:分层感知的LoRA专家混合,用于自我监督模型的持续多语言扩展,而不会忘记
链接:https://arxiv.org/abs/2602.12746
备注:Accepted by ICASSP 2026
摘要:尽管自监督语音模型的性能令人印象深刻,但它们往往难以推广到新的语言,并且在持续训练期间往往会忘记以前获得的知识。为了解决这个问题,我们提出了Lamer-SSL,这是一个参数高效的框架,它将LoRA专家(Lamer)模块的层感知混合与重放策略集成在一起。Lamer模块实现了共享和特定语言表示之间的灵活平衡,而层感知专家分配将更多专家分配到语义信息更丰富的更深层。同时,重放策略使用最少的数据保留先验知识,减少在连续训练过程中的遗忘。自动语音识别(ASR)和语言识别(LID)的实验表明,Lamer-SSL有效地将自监督模型扩展到新的语言,同时保持对先前学习的语言的强大性能,只有2.14%的参数是可训练的。
摘要:Despite their impressive performance, self-supervised speech models often struggle to generalize to new languages and tend to forget previously acquired knowledge during continual training. To address this, we propose Lamer-SSL, a parameter-efficient framework that integrates a Layer-Aware MixturE of LoRA Experts (Lamer) module with a replay strategy. The Lamer module enables flexible balancing between shared and language-specific representations, while layer-aware expert allocation assigns more experts to deeper layers where semantic information is richer. Meanwhile, the replay strategy retains prior knowledge using minimal data, mitigating forgetting during continual training. Experiments on automatic speech recognition (ASR) and language identification (LID) demonstrate that Lamer-SSL extends self-supervised models to new languages effectively while maintaining strong performance on previously learned languages with only 2.14% parameters being trainable.
【5】OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
标题:OmniCustoms:通过联合音视频生成模型同步音视频定制
链接:https://arxiv.org/abs/2602.12304
备注:16 pages
摘要:现有主流的视频定制方法集中于基于给定的参考图像和文本提示生成身份一致的视频。得益于音视频联合生成技术的快速发展,本文提出了一个更加引人注目的新课题:音视频同步定制,旨在同步定制视频标识和音频音色。具体地,给定参考图像$I^{r}$和参考音频$A^{r}$,该新颖的任务需要生成保持参考图像的身份同时模仿参考音频的音色的视频,其中通过用户提供的文本提示可自由指定口语内容。为此,我们提出了OmniCustom,一个功能强大的基于DiT的音视频定制框架,可以合成一个视频以下的参考图像身份,音频音色,和文本提示一次在一个zero-shot的方式。我们的框架建立在三个关键贡献之上。首先,身份和音频音色控制是通过独立的参考身份和音频LoRA模块实现的,这些模块通过基础音频-视频生成模型内的自注意层进行操作。其次,我们引入了一个对比学习目标旁边的标准流匹配目标。它使用的参考输入条件下的预测流作为正面的例子和那些没有参考条件作为负面的例子,从而提高模型的能力,以保持身份和音色。第三,我们训练OmniCustom在我们构建的大规模,高质量的视听人类数据集。大量实验表明,OmniCustom在生成具有一致身份和音色保真度的音频视频内容方面优于现有方法。
摘要:Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, this paper proposes a more compelling new task: sync audio-video customization, which aims to synchronously customize both video identity and audio timbre. Specifically, given a reference image $I^{r}$ and a reference audio $A^{r}$, this novel task requires generating videos that maintain the identity of the reference image while imitating the timbre of the reference audio, with spoken content freely specifiable through user-provided textual prompts. To this end, we propose OmniCustom, a powerful DiT-based audio-video customization framework that can synthesize a video following reference image identity, audio timbre, and text prompts all at once in a zero-shot manner. Our framework is built on three key contributions. First, identity and audio timbre control are achieved through separate reference identity and audio LoRA modules that operate through self-attention layers within the base audio-video generation model. Second, we introduce a contrastive learning objective alongside the standard flow matching objective. It uses predicted flows conditioned on reference inputs as positive examples and those without reference conditions as negative examples, thereby enhancing the model ability to preserve identity and timbre. Third, we train OmniCustom on our constructed large-scale, high-quality audio-visual human dataset. Extensive experiments demonstrate that OmniCustom outperforms existing methods in generating audio-video content with consistent identity and timbre fidelity.
【6】Beyond Musical Descriptors: Extracting Preference-Bearing Intent in Music Queries
标题:超越音乐描述符:提取音乐预设中的偏好意图
链接:https://arxiv.org/abs/2602.12301
备注:Accepted at NLP4MusA 2026 (4th Workshop on NLP for Music and Audio)
摘要:虽然用于用户查询的带注释的音乐描述符数据集越来越常见,但很少考虑这些描述符背后的用户意图,这对于有效满足他们的需求至关重要。我们介绍了MusicallyIntent,这是一个包含2,291个Reddit音乐请求的手动注释语料库,将音乐描述符标记为七个类别,具有积极,消极或参考偏好承载角色。然后,我们研究了大型语言模型(LLM)如何可靠地提取这些音乐描述符,发现它们确实捕获了明确的描述符,但与上下文相关的描述符有关。这项工作可以进一步作为用户意图的细粒度建模的基准,并获得改善基于LLM的音乐理解系统的见解。
摘要:Although annotated music descriptor datasets for user queries are increasingly common, few consider the user's intent behind these descriptors, which is essential for effectively meeting their needs. We introduce MusicRecoIntent, a manually annotated corpus of 2,291 Reddit music requests, labeling musical descriptors across seven categories with positive, negative, or referential preference-bearing roles. We then investigate how reliably large language models (LLMs) can extract these music descriptors, finding that they do capture explicit descriptors but struggle with context-dependent ones. This work can further serve as a benchmark for fine-grained modeling of user intent and for gaining insights into improving LLM-based music understanding systems.
【7】Retrieval-Augmented Self-Taught Reasoning Model with Adaptive Chain-of-Thought for ASR Named Entity Correction
标题:具有自适应思想链的用于SVR命名实体纠正的检索增强自学推理模型
链接:https://arxiv.org/abs/2602.12287
摘要:端到端自动语音识别(ASR)系统经常会错误识别领域特定的短语,如命名实体,这可能会导致下游任务的灾难性失败。最近出现了一种新的基于大型语言模型(LLM)的命名实体校正方法。然而,这些方法尚未充分利用LLM固有的复杂推理能力。为了弥补这一差距,我们提出了一种新的检索增强生成框架,用于纠正命名实体错误ASR。我们的方法包括两个关键部分:(1)用于命名实体识别的改写语言模型(RLM),然后使用语音级编辑距离进行候选检索;以及(2)一种新的自学推理模型,具有自适应思维链(A-STAR),可根据任务难度动态调整推理深度。在AISHELL-1和Homophone数据集上的实验证明了该方法的有效性,与强基线相比,命名实体字符错误率分别降低了17.96%和34.42%.
摘要:End-to-end automatic speech recognition (ASR) systems frequently misrecognize domain-specific phrases like named entities, which can cause catastrophic failures in downstream tasks. A new family of named entity correction methods based on large language models (LLMs) has recently emerged. However, these approaches have yet to fully exploit the sophisticated reasoning capabilities inherent to LLMs. To bridge this gap, we propose a novel retrieval-augmented generation framework for correcting named entity errors in ASR. Our approach consists of two key components: (1) a rephrasing language model (RLM) for named entity recognition, followed by candidate retrieval using a phonetic-level edit distance; and (2) a novel self-taught reasoning model with adaptive chain-of-thought (A-STAR) that dynamically adjusts the depth of its reasoning based on task difficulty. Experiments on the AISHELL-1 and Homophone datasets demonstrate the effectiveness of our method, which achieves relative reductions in the named entity character error rate of 17.96\% and 34.42\%, respectively, compared to a strong baseline.
机器翻译由腾讯交互翻译提供,仅供参考
