今日论文合集:cs.SD语音4篇,eess.AS音频处理3篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】ARIA: A Diagnostic Framework for Music Training Data Attribution
标题:ARIA:音乐训练数据归因的诊断框架
链接:https://arxiv.org/abs/2605.16181
作者:Changheon Han,Ashkan Panahi,Kıvanç Tatar
备注:Working Paper
摘要:用于音乐生成的训练数据属性(TDA)必须回答版权分析所需的两个问题,即哪些训练歌曲影响生成的输出以及这种影响在哪些音乐方面起作用。现有的方法将影响减少到单个标量,而没有揭示哪些音乐方面在该影响中占主导地位。我们提出了ARIA,一个框架,分解归因沿音乐方面(五个符号音乐,三个音频)和配对的分解与可靠性诊断计算的段级得分矩阵。它测量前K个属性轨道与从训练池中提取的随机参考组之间的组内相似性,并通过其奇异值分解和列统计来诊断得分矩阵。在一个符号音乐模型中,归因地面真相是通过反事实再训练,可靠性诊断排名四归因方法相同的地面真相。在音频音乐生成模型上,ARIA揭示了TDA方法之间差异很大的归因行为,标记其检索到的曲目在查询中几乎相同而不是反映每个查询归因的得分矩阵,并通过每个编码器表面的音乐方面表征嵌入相似性检索基线。总之,ARIA产生每个方面的归属证据与版权分析中的思想表达区别下考虑的音乐方面相一致。
摘要:Training data attribution (TDA) for music generation must answer two questions that copyright analysis requires, namely which training songs influence a generated output and along which musical aspects the influence operates. Existing methods reduce influence to a single scalar, without revealing which musical aspects are dominant in that influence. We propose ARIA, a framework that decomposes attribution along musical aspects (five for symbolic music, three for audio) and pairs the decomposition with reliability diagnostics computed from the segment-level score matrix. It measures within-group similarity among the top-K attributed tracks against random reference groups drawn from the training pool, and diagnoses the score matrix through its singular value decomposition and column statistics. On a symbolic-music model where attribution ground truth is available through counterfactual retraining, the reliability diagnostics rank four attribution methods identically to that ground truth. On an audio music generation model, ARIA reveals attribution behaviors that vary substantially across TDA methods, flags score matrices whose retrieved tracks are nearly identical across queries rather than reflecting per-query attribution, and characterizes embedding-similarity retrieval baselines by the musical aspect each encoder surfaces. Together, ARIA produces per-aspect attribution evidence aligned with the musical aspects considered under the idea-expression distinction in copyright analysis.


【2】Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues

标题:超越内容:综合语音毒性数据集和检测框架,体现副语言线索
链接:https://arxiv.org/abs/2605.15984
作者:Zhongjie Ba, Liang Yi, Peng Cheng, Qingcao Li, Qinglong Wang, Li Lu
摘要:
摘要:


【3】Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation

标题:将音乐建模为时频图像:音乐生成的2D代币器
链接:https://arxiv.org/abs/2605.15831
作者:Yuqing Cheng,Xingyu Ma,Guochen Yu,Xiaotao Gu
摘要:自回归音乐生成强烈依赖于音频标记器。现有的高保真编解码器通常使用残差多码本量化,这保留了重建质量,但使序列平坦化后的语言建模复杂化,因为残差层次结构强加了强的顺序依赖性,并且可以放大错误累积。我们提出了BandTok,一个面向生成的2D梅尔频谱图标记器,表示每个帧梅尔频带令牌从一个单一的共享码本。这种设计产生了一个物理上可解释的时间-频率令牌网格,具有更独立的令牌结构,使其更适合自回归建模。BandTok通过多尺度PatchGAN目标和EMA码本更新改进了重建。我们进一步引入了一个自回归语言模型与二维旋转位置嵌入(2D RoPE),以保持时间和频带结构在生成过程中。实验表明,BandTok改进了剩余码本标记器,并在数据受限的情况下取得了很好的效果。这项工作的源代码和生成演示是公开的。
摘要:Autoregressive music generation depends strongly on the audio tokenizer. Existing high-fidelity codecs often use residual multi-codebook quantization, which preserves reconstruction quality but complicates language modeling after sequence flattening, as the residual hierarchy imposes strong sequential dependencies and can amplify error accumulation. We propose BandTok, a generation-oriented 2D Mel-spectrogram tokenizer that represents each frame with Mel-frequency band tokens from a single shared codebook. This design yields a physically interpretable time-frequency token grid with a more independent token structure, making it better suited for autoregressive modeling. BandTok improves reconstruction with a multi-scale PatchGAN objective and EMA codebook updates. We further introduce an autoregressive language model with 2D Rotary Position Embedding (2D RoPE) to preserve temporal and frequency-band structure during generation. Experiments show that BandTok improves over residual-codebook tokenizers and achieves strong results in a data-limited setting. The source code and generation demos for this work are publicly available.


【4】Sound Sparks Motion: Audio and Text Tuning for Video Editing

标题:声音火花运动:用于视频编辑的音频和文本调整
链接:https://arxiv.org/abs/2605.15307
作者:AmirHossein Naghi Razlighi,Aryan Mikaeili,Ali Mahdavi-Amiri,Daniel Cohen-Or,Yiorgos Chrysanthou
备注:Project Page: https://amirhossein-razlighi.github.io/Sound_Sparks_Motion
摘要:以运动为中心的视频编辑对于大型生成视频模型仍然很困难,这些模型通常对外观变化响应良好,但难以在现有剪辑中产生特定的局部动作或状态转换。我们介绍了声音火花运动,一个无训练的框架,使运动编辑的视听视频生成模型,通过调整其内部多模态条件信号在测试时。而不是修改模型的权重,我们的方法只调整两个轻量级的变量:一个音频潜伏源视频和残余扰动的文本调节。我们发现,这种组合可以鼓励运动编辑,底层模型往往很难实现下,只有控制。由于没有直接的方法来评估文本和运动之间的时间对齐,我们使用视觉语言模型来指导调整过程,该模型提供反馈,指示预期的运动是否出现在生成的视频中。这种简单的监督为运动编辑提供了有效的语义目标,而正则化和感知时间约束有助于保持内容和视觉质量。除了每个视频的调整,我们表明,学习的潜在控制是跨视频可转移的,这表明他们捕捉可重用的运动编辑方向,而不是过拟合到一个单一的例子。   我们的研究结果突出了多模态条件调整,特别是通过音频通路,作为运动感知视频编辑的一个有前途的方向,并建议测试时间调整可以作为一个轻量级的探测机制,有助于揭示嵌入在模型的多模态条件下的潜在运动控制。代码和数据可通过我们的项目页面获得:https://amirhossein-razlighi.github.io/Sound_Sparks_Motion/
摘要:Motion-centric video editing remains difficult for large generative video models, which often respond well to appearance changes but struggle to produce specific, localized actions or state transitions in an existing clip. We introduce Sound Sparks Motion, a training-free framework that enables motion editing in an audio-visual video generation model by tuning its internal multimodal conditioning signals at test time. Rather than modifying model weights, our method tunes only two lightweight variables: an audio latent derived from the source video and a residual perturbation in the text-conditioning. We find that this combination can encourage motion edits that the underlying model often struggles to realize under prompt-only control. Since there is no direct way to evaluate temporal alignment between text and motion, we guide the tuning process using a vision-language model that provides feedback indicating whether the intended motion appears in the generated video. This simple supervision yields an effective semantic objective for motion editing, while regularization and perceptual-temporal constraints help preserve content and visual quality. Beyond per-video tuning, we show that the learned latent controls are transferable across videos, suggesting that they capture reusable motion-edit directions rather than overfitting to a single example.   Our results highlight multimodal conditioning tuning, particularly through the audio pathway, as a promising direction for motion-aware video editing, and suggest that test-time tuning can serve as a lightweight probing mechanism that helps reveal latent motion controls embedded in the model's multimodal conditioning. Code and data are available via our project page: https://amirhossein-razlighi.github.io/Sound_Sparks_Motion/


eess.AS音频处理


【1】Real-time Speech Restoration using Data Prediction Mean Flows
标题:使用数据预测平均流的实时语音恢复
链接:https://arxiv.org/abs/2605.16251
作者:Sebastian Braun
摘要:生成模型能够解决非唯一解决方案的难题,如带宽扩展和间隙填充,从编解码器中去除高度非线性的伪影,剪切和失真,而不是去除噪声和混响等线性添加剂成分。虽然大型离线处理模型已经显示出令人印象深刻的结果,但这些任务还没有通过具有低延迟和计算的实时模型来解决。我们提出了一个几步流匹配模型,使用数据预测平均流结合适当的新的低延迟架构,使流匹配模型在这些约束下的一个有吸引力的选择。与现有技术相比,我们提出的平均流模型使用的计算量减少了120倍,并且除了STFT之外没有引入算法延迟,同时实现了类似的音频质量。
摘要:Generative models are capable to address difficult problems with non-unique solutions like bandwidth extension and gap filling, removing highly non-linear artifacts from codecs, clipping and distortion, as opposed to removing linear additive components like noise and reverb. While large offline processing models have shown impressive results, these tasks have not been solved with real-time capable models with low latency and compute. We propose a few-step flow matching model using Data Prediction Mean Flows in combination with suitable novel low-latency architecture to make flow matching models an attractive choice under theses constraints. Compared to state-of-the-art, our proposed mean flow model uses 120x less compute and introduces no algorithmic latency other than the STFT, while achieving similar audio quality.


【2】Improving Automatic Speech Recognition for Speakers Treated for Oral Cancer using Data Augmentation and LLM Error Correction

标题:使用数据增强和LLM错误纠正改善接受口腔癌治疗的说话者的自动语音识别
链接:https://arxiv.org/abs/2605.15854
作者:Hidde Folkertsma,Thomas Tienkamp,Sebastiaan de Visscher,Max Witjes,Rob van Son,Jiapan Guo,Bence Mark Halpern
备注:7 pages, 3 tables. Accepted by EMBC 2026
摘要:近年来,自动语音识别(ASR)系统的性能取得了长足的进步。不幸的是,对于有语言障碍的人,例如接受口腔癌(OC)治疗的人,ASR性能仍然落后。OC语音数据的稀缺性和可变性使得针对这种类型的语音开发ASR模型变得困难。在这项工作中,我们使用数据增强和大型语言模型(LLM)纠错来缓解这个问题。我们应用各种增强技术在荷兰口腔癌语音语料库上创建合成数据,并评估其对ASR性能的影响。我们为每种增强技术微调Whisper和大规模多语言语音(MMS)模型,并观察到,当包括使用文本到语音(TTS)创建的数据时,平均而言,单词错误率(WER)相对降低8%。当采用LLM进行误差校正时,我们看到微调ASR模型的WER进一步相对降低21.4-26.2%,非微调模型的WER相对降低10.0%。总体而言,我们实现了40%的相对WER降低耳语和50%的相对WER降低MMS,这表明数据增强和LLM校正的组合是一个可行的策略识别OC语音。
摘要:In recent years, the performance of automatic speech recognition (ASR) systems has made considerable progress. Unfortunately, for people with speech impairments, such as people treated for oral cancer (OC), ASR performance is still lagging behind. The scarcity and variability of OC speech data makes development of ASR models for this type of speech difficult. In this work, we use data augmentation and large language model (LLM) error correction to mitigate this problem. We apply various augmentation techniques on a corpus of Dutch oral cancer speech to create synthetic data, and evaluate their effect on ASR performance. We finetune Whisper and Massively Multilingual Speech (MMS) models for each augmentation technique and observe, on average, an 8% relative decrease in Word Error Rate (WER) when including data created using text-to-speech (TTS). When employing LLMs for error correction, we see a further 21.4-26.2% relative decrease in WER for finetuned ASR models and a 10.0% relative decrease for non-finetuned models. Overall, we achieve a 40% relative WER decrease for Whisper and a 50% relative WER decrease for MMS, indicating that a combination of data augmentation and LLM correction is a viable strategy for the recognition of OC speech.


【3】Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization

标题:注意差距:合成对话数据对多说话者ASB和说话者拨号的影响
链接:https://arxiv.org/abs/2605.15442
作者:Alexander Polok,Ivan Medennikov,Jan Černocký,Shinji Watanabe,Lukáš Burget,Samuele Cornell
备注:Submitted to INTERSPEECH 2026
摘要:最近在多说话者ASR(MT-ASR)和说话者日记(SD)方面的突破依赖于合成数据来缓解大规模对话录音的稀缺性,但对特定模拟选择的影响仍然知之甚少。考虑到模拟混合物和现实世界的相互作用之间的差距,我们提出了一个领先的MT-ASR(DiCoW)和SD(Sortformer)系统的合成数据生成的研究。通过介绍FastMSS,一个高效的开源模拟器,我们分析了轮流动态,源域,声学增强和数据混合策略。我们的研究结果表明,最佳的模拟配方是高度依赖于任务:增加语音重叠的好处ASR,但降低日记。此外,广泛的源多样性始终优于精确的域匹配。最终,仅合成训练接近真实数据基线,并且将模拟数据与真实记录相结合,在这两项任务中都比仅真实训练获得了实质性的收益。
摘要:Recent breakthroughs in multi-talker ASR (MT-ASR) and speaker diarization (SD) rely on synthetic data to mitigate the scarcity of large-scale conversational recordings, yet the impact of specific simulation choices remains poorly understood. To mind the gap between simulated mixtures and real-world interactions, we present a study of synthetic data generation for leading MT-ASR (DiCoW) and SD (Sortformer) systems. By introducing FastMSS, a highly efficient open-source simulator, we analyze turn-taking dynamics, source domain, acoustic augmentation, and data mixing strategies. Our findings reveal that optimal simulation recipes are highly task-dependent: increasing speech overlap benefits ASR but degrades diarization. Furthermore, broad source diversity consistently outperforms exact domain matching. Ultimately, synthetic-only training approaches real-data baselines, and combining simulated data with real recordings yields substantial gains over real-only training across both tasks.


机器翻译由腾讯交互翻译提供,仅供参考