今日论文合集:cs.SD语音5篇,eess.AS音频处理5篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline
标题:作曲歌曲管道中的MIDI知情歌唱伴奏一代
链接:https://arxiv.org/abs/2602.22029

作者:Fang-Duo Tsai,Yi-An Lai,Fei-Yueh Chen,Hsueh-Wei Fu,Li Chai,Wei-Jaw Lee,Hao-Chung Cheng,Yi-Hsuan Yang
摘要:歌曲生成的目标是从歌词和文本描述中生成带有人声和伴奏的完整歌曲,但端到端模型仍然是数据和计算密集型的,并且提供有限的可编辑性。我们提倡一种作曲的替代方案,将任务分解为旋律创作,歌唱声音合成和歌唱伴奏生成。我们的方法的核心是MIDI知情的歌唱伴奏生成(MIDI-SAG),它的条件伴奏的象征性的声乐旋律旋律,以提高歌唱和乐器之间的节奏和和声对齐。此外,除了传统的SAG设置,假设连续演唱的人声,作曲歌曲生成功能间歇性的人声,我们解决这个问题,通过结合明确的节奏/谐波控制与音频延续,以保持整个声乐和非声乐区域的背景曲目一致。凭借轻量级的新训练组件,在单个RTX 3090上仅需要2.5k小时的音频,我们的管道在几个指标上接近最近开源端到端基线的感知质量。我们提供音频演示,并将在https://composerflow.github.io/web/上开源我们的模型。
摘要:Song generation aims to produce full songs with vocals and accompaniment from lyrics and text descriptions, yet end-to-end models remain data- and compute-intensive and provide limited editability. We advocate a compositional alternative that decomposes the task into melody composition, singing voice synthesis, and singing accompaniment generation. Central to our approach is MIDI-informed singing accompaniment generation (MIDI-SAG), which conditions accompaniment on the symbolic vocal-melody MIDI to improve rhythmic and harmonic alignment between singing and instrumentation. Moreover, beyond conventional SAG settings that assume continuously sung vocals, compositional song generation features intermittent vocals; we address this by combining explicit rhythmic/harmonic controls with audio continuation to keep the backing track consistent across vocal and non-vocal regions. With lightweight newly trained components requiring only 2.5k hours of audio on a single RTX 3090, our pipeline approaches the perceptual quality of recent open-source end-to-end baselines in several metrics. We provide audio demos and will open-source our model at https://composerflow.github.io/web/.


【2】EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
标题:Omni:在Omni-Modal LLM中架起情感理解和表达的桥梁
链接:https://arxiv.org/abs/2602.21900

作者:Wenjie Tian,Zhixian Zhao,Jingbin Hu,Huakang Chen,Haohe Liu,Binshen Mu,Lei Xie
摘要:全模态大语言模型(Omni-LLM)的发展彻底改变了人机交互,使视听感知和语音响应实现了统一。然而,现有的Omni-LLM在复杂的现实世界场景中苦苦挣扎,往往导致肤浅的理解和上下文不匹配的情绪反应。Omni-LLM的Thinker-Talker架构进一步加剧了这个问题,这些架构通过隐藏状态隐式连接,导致情感细节的丢失。在这项工作中,我们提出了一个统一的框架,准确的理解和表达多模态情感对话。在其核心,我们引入了情感的思维链~(E-CoT),它强制执行从细粒度的多模态感知到文本响应的推理。此外,我们明确地将E-CoT视为指导谈话者的高级情感指令,从而实现准确的情感表达。作为对模型的补充,我们构建了多模态情感对话数据库来获取真实世界的标注对话数据,并建立了一个多模态情感对话任务的系统评估基准--多模态情感对话评测。实验表明,在相同的说话人下,Qwen 3 Omni-7 B与Qwen 3 Omni-30 B-A3 B-Thinking的性能相当。
摘要:The evolution of Omni-Modal Large Language Models~(Omni-LLMs) has revolutionized human--computer interaction, enabling unified audio-visual perception and speech response. However, existing Omni-LLMs struggle with complex real-world scenarios, often leading to superficial understanding and contextually mismatched emotional responses. This issue is further intensified by Omni-LLM's Thinker-Talker architectures, which are implicitly connected through hidden states, leading to the loss of emotional details. In this work, we present EmoOmni, a unified framework for accurate understanding and expression in multimodal emotional dialogue. At its core, we introduce the emotional Chain-of-Thought~(E-CoT), which enforces a reasoning from fine-grained multimodal perception to textual response. Moreover, we explicitly treat E-CoT as high-level emotional instructions that guide the talker, enabling accurate emotional expression. Complementing the model, we construct EmoOmniPipe to obtain the real-world annotated dialogue data and establish a benchmark, EmoOmniEval, to facilitate systematic assessment of multimodal emotional dialogue task. Experiments show that EmoOmni-7B achieves comparable performance with Qwen3Omni-30B-A3B-Thinking under the same talker.


【3】UniWhisper: Efficient Continual Multi-task Training for Robust Universal Audio Representation
标题:UniWhisper:高效的连续多任务训练,以实现稳健的通用音频表示
链接:https://arxiv.org/abs/2602.21772

作者:Yuxuan Chen,Peize He,Haoyuan Xu,Junzi Zhang
摘要:通用的音频表示应该在单个编码器中捕获环境声音和音乐的细粒度语音线索和高级语义。现有的编码器通常在一个领域表现出色,但在其他领域却有所下降。我们提出了UniWhisper,一个高效的连续多任务训练框架,将异构的音频任务转换为统一的指令和答案格式。这使得标准的下一个令牌训练没有特定于任务的头和损失。我们在38 k小时的公共音频上对它进行训练,并使用浅层MLP探针和k-最近邻(kNN)对涵盖语音,环境声音和音乐的20个任务进行评估。UniWhisper使用MLP探针达到归一化加权平均值0.81,使用kNN达到归一化加权平均值0.61,而Whisper为0.64和0.46,同时保持强大的语音性能。
摘要:A universal audio representation should capture fine-grained speech cues and high-level semantics for environmental sounds and music in a single encoder. Existing encoders often excel in one domain but degrade in others. We propose UniWhisper, an efficient continual multi-task training framework that casts heterogeneous audio tasks into a unified instruction and answer format. This enables standard next-token training without task-specific heads and losses. We train it on 38k hours of public audio and assess the encoder using shallow MLP probes and k-nearest neighbors (kNN) on 20 tasks spanning speech, environmental sound, and music. UniWhisper reaches normalized weighted averages of 0.81 with MLP probes and 0.61 with kNN, compared to 0.64 and 0.46 for Whisper, while retaining strong speech performance.


【4】Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
标题:鲁棒的长形式孟加拉语语音处理:自动语音识别和说话人分区化
链接:https://arxiv.org/abs/2602.21741

作者:MD. Sagor Chowdhury,Adiba Fairooz Chowdhury
备注:6 pages, 5 figures, 3 tables; system paper submitted to DL Sprint 4.0 (Kaggle)
摘要:我们描述了我们的端到端系统孟加拉语长格式语音识别(ASR)和扬声器日记提交给DL Sprint 4.0的竞争Kaggle。孟加拉语提出了实质性的挑战,这两项任务:一个大的音素库存,显着的方言变异,频繁的代码混合与英语,以及相对稀缺的大规模标记语料库。对于ASR,我们实现了0.37738的最佳私有字错误率(WER)和0.36137的公共WER,将BengaliAI微调的Whisper媒体模型与Demucs源分离相结合,用于语音隔离,沉默边界分块和精心调整的生成超参数。对于说话人日记,我们通过将pyannote.audio管道内的默认分割模型替换为孟加拉语微调变体,将其与wespeaker-voxceleb-resnet 34-LM嵌入和基于质心的凝聚聚类配对,达到了0.27671的最佳私有日记错误率(DER)和0.20936的公共DER。我们的实验表明,特定领域的微调分割组件,声音源分离,自然的沉默意识分块是三个最有效的设计选择低资源孟加拉语语音处理。
摘要:We describe our end-to-end system for Bengali long-form speech recognition (ASR) and speaker diarization submitted to the DL Sprint 4.0 competition on Kaggle. Bengali presents substantial challenges for both tasks: a large phoneme inventory, significant dialectal variation, frequent code-mixing with English, and a relative scarcity of large-scale labelled corpora. For ASR we achieve a best private Word Error Rate (WER) of 0.37738 and public WER of 0.36137, combining a BengaliAI fine-tuned Whisper medium model with Demucs source separation for vocal isolation, silence-boundary chunking, and carefully tuned generation hyperparameters. For speaker diarization we reach a best private Diarization Error Rate (DER) of 0.27671 and public DER of 0.20936 by replacing the default segmentation model inside the pyannote.audio pipeline with a Bengali-fine-tuned variant, pairing it with wespeaker-voxceleb-resnet34-LM embeddings and centroid-based agglomerative clustering. Our experiments demonstrate that domain-specific fine-tuning of the segmentation component, vocal source separation, and natural silence-aware chunking are the three most impactful design choices for low-resource Bengali speech processing.


【5】TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
标题:TG-ASB:具有并行门控交叉注意力的翻译引导学习,用于低资源自动语音识别
链接:https://arxiv.org/abs/2602.22039

作者:Cheng-Yeh Yang,Chien-Chun Wang,Li-Wei Chen,Hung-Shin Lee,Hsin-Min Wang,Berlin Chen
备注:Accepted to LREC 2026
摘要:低资源自动语音识别(ASR)继续构成重大挑战,主要是由于许多语言的转录数据有限。虽然在电视剧和网络视频中可以看到大量的口语内容,但台湾闽南语却解决了这一问题,因为字幕通常很少,而且大多数可用的字幕都只有普通话。为了解决这一不足,我们引入了TG-ASR台湾闽南语戏剧语音识别,一个引导的ASR框架,利用多语言翻译嵌入,以提高识别性能在低资源环境中。该框架围绕并行门控交叉注意(PGCA)机制,自适应地集成嵌入从各种辅助语言到ASR解码器。这种机制有助于强大的跨语言语义指导,同时确保稳定的优化和最大限度地减少语言之间的干扰。为了支持正在进行的研究计划,我们提出了YT-THDC,一个30小时的语料库台湾闽南语戏剧讲话对齐普通话字幕和手动验证的台湾闽南语transmittance。综合的实验和分析确定了最有效地提高ASR性能的辅助语言,实现了14.77%的字符错误率相对降低,并在实际应用中证明了对代表性不足的语言进行预防引导学习的有效性。
摘要:Low-resource automatic speech recognition (ASR) continues to pose significant challenges, primarily due to the limited availability of transcribed data for numerous languages. While a wealth of spoken content is accessible in television dramas and online videos, Taiwanese Hokkien exemplifies this issue, with transcriptions often being scarce and the majority of available subtitles provided only in Mandarin. To address this deficiency, we introduce TG-ASR for Taiwanese Hokkien drama speech recognition, a translation-guided ASR framework that utilizes multilingual translation embeddings to enhance recognition performance in low-resource environments. The framework is centered around the parallel gated cross-attention (PGCA) mechanism, which adaptively integrates embeddings from various auxiliary languages into the ASR decoder. This mechanism facilitates robust cross-linguistic semantic guidance while ensuring stable optimization and minimizing interference between languages. To support ongoing research initiatives, we present YT-THDC, a 30-hour corpus of Taiwanese Hokkien drama speech with aligned Mandarin subtitles and manually verified Taiwanese Hokkien transcriptions. Comprehensive experiments and analyses identify the auxiliary languages that most effectively enhance ASR performance, achieving a 14.77% relative reduction in character error rate and demonstrating the efficacy of translation-guided learning for underrepresented languages in practical applications.


eess.AS音频处理


【1】TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
标题:TG-ASB:具有并行门控交叉注意力的翻译引导学习,用于低资源自动语音识别
链接:https://arxiv.org/abs/2602.22039

作者:Cheng-Yeh Yang,Chien-Chun Wang,Li-Wei Chen,Hung-Shin Lee,Hsin-Min Wang,Berlin Chen
备注:Accepted to LREC 2026
摘要:低资源自动语音识别(ASR)继续构成重大挑战,主要是由于许多语言的转录数据有限。虽然在电视剧和网络视频中可以看到大量的口语内容,但台湾闽南语却解决了这一问题,因为字幕通常很少,而且大多数可用的字幕都只有普通话。为了解决这一不足,我们引入了TG-ASR台湾闽南语戏剧语音识别,一个引导的ASR框架,利用多语言翻译嵌入,以提高识别性能在低资源环境中。该框架围绕并行门控交叉注意(PGCA)机制,自适应地集成嵌入从各种辅助语言到ASR解码器。这种机制有助于强大的跨语言语义指导,同时确保稳定的优化和最大限度地减少语言之间的干扰。为了支持正在进行的研究计划,我们提出了YT-THDC,一个30小时的语料库台湾闽南语戏剧讲话对齐普通话字幕和手动验证的台湾闽南语transmittance。综合的实验和分析确定了最有效地提高ASR性能的辅助语言,实现了14.77%的字符错误率相对降低,并在实际应用中证明了对代表性不足的语言进行预防引导学习的有效性。
摘要:Low-resource automatic speech recognition (ASR) continues to pose significant challenges, primarily due to the limited availability of transcribed data for numerous languages. While a wealth of spoken content is accessible in television dramas and online videos, Taiwanese Hokkien exemplifies this issue, with transcriptions often being scarce and the majority of available subtitles provided only in Mandarin. To address this deficiency, we introduce TG-ASR for Taiwanese Hokkien drama speech recognition, a translation-guided ASR framework that utilizes multilingual translation embeddings to enhance recognition performance in low-resource environments. The framework is centered around the parallel gated cross-attention (PGCA) mechanism, which adaptively integrates embeddings from various auxiliary languages into the ASR decoder. This mechanism facilitates robust cross-linguistic semantic guidance while ensuring stable optimization and minimizing interference between languages. To support ongoing research initiatives, we present YT-THDC, a 30-hour corpus of Taiwanese Hokkien drama speech with aligned Mandarin subtitles and manually verified Taiwanese Hokkien transcriptions. Comprehensive experiments and analyses identify the auxiliary languages that most effectively enhance ASR performance, achieving a 14.77% relative reduction in character error rate and demonstrating the efficacy of translation-guided learning for underrepresented languages in practical applications.


【2】A Knowledge-Driven Approach to Music Segmentation, Music Source Separation and Cinematic Audio Source Separation
标题:知识驱动的音乐分割、音乐源分离和电影音频源分离方法
链接:https://arxiv.org/abs/2602.21476

作者:Chun-wei Ho,Sabato Marco Siniscalchi,Kai Li,Chin-Hui Lee
摘要:我们提出了一个知识驱动的,基于模型的方法来分割音频成单类和混合类块源分离的应用程序。这里的“知识”表示与数据相关联的信息,例如乐谱。这里的“模型”是指可以用于音频分割和识别的工具,例如隐马尔可夫模型。与通常依赖于具有给定片段类别及其相应边界的注释数据来指导学习过程的传统学习相比,所提出的框架不依赖于任何预先分段的训练数据,并且直接从输入音频及其相关知识源学习以自主地构建所有必要的模型。仿真数据评估表明,分数引导学习取得了非常好的音乐分割和分离效果。对电影音轨数据进行的电影音频源分离测试也表明,利用声音类别知识比不使用此类信息的数据驱动技术获得的分离结果更好。
摘要:We propose a knowledge-driven, model-based approach to segmenting audio into single-category and mixed-category chunks with applications to source separation. "Knowledge" here denotes information associated with the data, such as music scores. "Model" here refers to tool that can be used for audio segmentation and recognition, such as hidden Markov models. In contrast to conventional learning that often relies on annotated data with given segment categories and their corresponding boundaries to guide the learning process, the proposed framework does not depend on any pre-segmented training data and learns directly from the input audio and its related knowledge sources to build all necessary models autonomously. Evaluation on simulation data shows that score-guided learning achieves very good music segmentation and separation results. Tested on movie track data for cinematic audio source separation also shows that utilizing sound category knowledge achieves better separation results than those obtained with data-driven techniques without using such information.


【3】iMiGUE-Speech: A Spontaneous Speech Dataset for Affective Analysis
标题:iMiGUE-Speech:用于情感分析的自发言语数据集
链接:https://arxiv.org/abs/2602.21464

作者:Sofoklis Kakouros,Fang Kang,Haoyu Chen
备注:Accepted to Speech Prosody 2026
摘要:这项工作提出了iMiGUE语音,iMiGUE数据集的扩展,提供了一个自发的情感语料库研究情绪和情感状态。新版本侧重于语音,并通过额外的元数据丰富了原始数据集,包括语音转录,面试官和受访者之间的演讲者角色分离以及单词级强制对齐。与现有的依赖于行为或实验室引发的情感的情感语音数据集不同,iMiGUE-Speech捕获了从真实比赛结果中自然产生的自发影响。为了展示数据集的实用性并建立初始基准,我们引入了两个评估任务进行比较评估:语音情感识别和基于成绩单的情感分析。这些任务利用最先进的预训练表示来评估数据集从声学和语言模态捕获自发情感状态的能力。iMiGUE语音还可以与来自原始iMiGUE数据集的微手势注释同步配对,形成用于研究语音-手势情感动态的独特多模态资源。扩展数据集可在https://github.com/CV-AC/imigue-speech上获得。
摘要:This work presents iMiGUE-Speech, an extension of the iMiGUE dataset that provides a spontaneous affective corpus for studying emotional and affective states. The new release focuses on speech and enriches the original dataset with additional metadata, including speech transcripts, speaker-role separation between interviewer and interviewee, and word-level forced alignments. Unlike existing emotional speech datasets that rely on acted or laboratory-elicited emotions, iMiGUE-Speech captures spontaneous affect arising naturally from real match outcomes. To demonstrate the utility of the dataset and establish initial benchmarks, we introduce two evaluation tasks for comparative assessment: speech emotion recognition and transcript-based sentiment analysis. These tasks leverage state-of-the-art pre-trained representations to assess the dataset's ability to capture spontaneous affective states from both acoustic and linguistic modalities. iMiGUE-Speech can also be synchronously paired with micro-gesture annotations from the original iMiGUE dataset, forming a uniquely multimodal resource for studying speech-gesture affective dynamics. The extended dataset is available at https://github.com/CV-AC/imigue-speech.


【4】MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline
标题:作曲歌曲管道中的MIDI知情歌唱伴奏一代
链接:https://arxiv.org/abs/2602.22029

作者:Fang-Duo Tsai,Yi-An Lai,Fei-Yueh Chen,Hsueh-Wei Fu,Li Chai,Wei-Jaw Lee,Hao-Chung Cheng,Yi-Hsuan Yang
摘要:歌曲生成的目标是从歌词和文本描述中生成带有人声和伴奏的完整歌曲,但端到端模型仍然是数据和计算密集型的,并且提供有限的可编辑性。我们提倡一种作曲的替代方案,将任务分解为旋律创作,歌唱声音合成和歌唱伴奏生成。我们的方法的核心是MIDI知情的歌唱伴奏生成(MIDI-SAG),它的条件伴奏的象征性的声乐旋律旋律,以提高歌唱和乐器之间的节奏和和声对齐。此外,除了传统的SAG设置,假设连续演唱的人声,作曲歌曲生成功能间歇性的人声,我们解决这个问题,通过结合明确的节奏/谐波控制与音频延续,以保持整个声乐和非声乐区域的背景曲目一致。凭借轻量级的新训练组件,在单个RTX 3090上仅需要2.5k小时的音频,我们的管道在几个指标上接近最近开源端到端基线的感知质量。我们提供音频演示,并将在https://composerflow.github.io/web/上开源我们的模型。
摘要:Song generation aims to produce full songs with vocals and accompaniment from lyrics and text descriptions, yet end-to-end models remain data- and compute-intensive and provide limited editability. We advocate a compositional alternative that decomposes the task into melody composition, singing voice synthesis, and singing accompaniment generation. Central to our approach is MIDI-informed singing accompaniment generation (MIDI-SAG), which conditions accompaniment on the symbolic vocal-melody MIDI to improve rhythmic and harmonic alignment between singing and instrumentation. Moreover, beyond conventional SAG settings that assume continuously sung vocals, compositional song generation features intermittent vocals; we address this by combining explicit rhythmic/harmonic controls with audio continuation to keep the backing track consistent across vocal and non-vocal regions. With lightweight newly trained components requiring only 2.5k hours of audio on a single RTX 3090, our pipeline approaches the perceptual quality of recent open-source end-to-end baselines in several metrics. We provide audio demos and will open-source our model at https://composerflow.github.io/web/.


【5】EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
标题:Omni:在Omni-Modal LLM中架起情感理解和表达的桥梁
链接:https://arxiv.org/abs/2602.21900

作者:Wenjie Tian,Zhixian Zhao,Jingbin Hu,Huakang Chen,Haohe Liu,Binshen Mu,Lei Xie
摘要:全模态大语言模型(Omni-LLM)的发展彻底改变了人机交互,使视听感知和语音响应实现了统一。然而,现有的Omni-LLM在复杂的现实世界场景中苦苦挣扎,往往导致肤浅的理解和上下文不匹配的情绪反应。Omni-LLM的Thinker-Talker架构进一步加剧了这个问题,这些架构通过隐藏状态隐式连接,导致情感细节的丢失。在这项工作中,我们提出了一个统一的框架,准确的理解和表达多模态情感对话。在其核心,我们引入了情感的思维链~(E-CoT),它强制执行从细粒度的多模态感知到文本响应的推理。此外,我们明确地将E-CoT视为指导谈话者的高级情感指令,从而实现准确的情感表达。作为对模型的补充,我们构建了多模态情感对话数据库来获取真实世界的标注对话数据,并建立了一个多模态情感对话任务的系统评估基准--多模态情感对话评测。实验表明,在相同的说话人下,Qwen 3 Omni-7 B与Qwen 3 Omni-30 B-A3 B-Thinking的性能相当。
摘要:The evolution of Omni-Modal Large Language Models~(Omni-LLMs) has revolutionized human--computer interaction, enabling unified audio-visual perception and speech response. However, existing Omni-LLMs struggle with complex real-world scenarios, often leading to superficial understanding and contextually mismatched emotional responses. This issue is further intensified by Omni-LLM's Thinker-Talker architectures, which are implicitly connected through hidden states, leading to the loss of emotional details. In this work, we present EmoOmni, a unified framework for accurate understanding and expression in multimodal emotional dialogue. At its core, we introduce the emotional Chain-of-Thought~(E-CoT), which enforces a reasoning from fine-grained multimodal perception to textual response. Moreover, we explicitly treat E-CoT as high-level emotional instructions that guide the talker, enabling accurate emotional expression. Complementing the model, we construct EmoOmniPipe to obtain the real-world annotated dialogue data and establish a benchmark, EmoOmniEval, to facilitate systematic assessment of multimodal emotional dialogue task. Experiments show that EmoOmni-7B achieves comparable performance with Qwen3Omni-30B-A3B-Thinking under the same talker.


机器翻译由腾讯交互翻译提供,仅供参考