微信公众号:arXiv_Daily
cs.SD语音
标题:通过从未发生过的对话进行高效的ASB训练
链接:https://arxiv.org/abs/2606.03957
摘要:低资源语言和利基领域的会话ASR受到缺乏领域匹配的多说话人训练数据的限制。我们提出了一个增强管道,生成与参与者的元数据,扬声器属性映射到TTS语音配置文件,合成的话语到扬声器感知的模拟对话。我们评估了五个LLM家庭在单一发电机,固定预算的混合,并使用相同的FastConformer-Large训练配方的每个设置。我们对匈牙利BEA-Dialogue基准语料库进行了全面的评估,该方法本身适用于任何语言,只要每个组件都有资源。结果表明,合成对话一致提高语音识别性能,但发电机的选择和数据组成强烈影响收益。我们最大的训练配置仅使用67小时的真实对话和636小时的模拟数据,在评估基准上取得了比在2700小时匈牙利语语音上训练的zero-shot模型更好的性能。这些研究结果表明,LLM生成的会话数据合成的TTS是一个实际的补充,真正的会话语料库的语音模型训练。
摘要:Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations. We evaluated five LLM families under single-generator, fixed-budget mixture, and scale-up settings using the same FastConformer-Large training recipe for each one. We ran comprehensive evaluations on the Hungarian BEA-Dialogue benchmark corpus, with the method itself being applicable to any language given the resources for each component. The results show that synthetic conversations consistently improve speech recognition performance, but generator choice and data composition strongly affect the gains. Our largest training configuration, using only 67 hours of real conversations and 636 hours of simulated data, achieves better performance on the evaluation benchmark than a zero-shot model trained on 2700 hours of Hungarian speech. These findings indicate that LLM-generated conversational data synthesized with TTS is a practical complement to real conversational corpora for speech model training.
【2】LiveBand: Live Accompaniment Generation in the Audio Domain
标题:LiveBand:音频领域的现场伴奏一代链接:https://arxiv.org/abs/2606.03803
摘要:我们提出了LiveBand,一个实时系统,生成高保真的音乐伴奏,以现场音频输入,尊重严格的因果约束。我们的方法在预先训练的因果音频自动编码器的连续潜在空间中训练因果Transformer生成器,使用来自编码器的对抗序列级监督。在每个时步,生成器仅接收因果可用的混合上下文和高斯噪声,并预测伴奏潜伏期,而不访问未来的混合帧或地面实况目标潜伏期。训练在因果掩蔽下以单个并行向前传递进行,而流推理则以滚动注意状态进行自回归。该模型的训练和推理计算通过设计进行匹配,消除了教师强迫和相关的暴露偏差。在多乐器音乐伴奏基准测试中,LiveBand在音频质量、节拍对齐和混音一致性的客观测量方面优于之前的工作,同时在消费者硬件上实现实时流媒体生成,而无需展望未来。
摘要:We present LiveBand, a real-time system that generates high-fidelity music accompaniments to live audio input, respecting strict causal constraints. Our method trains a causal transformer generator in the continuous latent space of a pre-trained causal audio autoencoder, using adversarial sequence-level supervision from a discriminator. At each timestep, the generator receives only the causally available mix context and Gaussian noise, and predicts accompaniment latents without access to future mix frames or ground-truth target latents. Training is performed in a single parallel forward pass under causal masking, while streaming inference proceeds autoregressively with a rolling attention state. The model's training and inference computations are matched by design, eliminating teacher forcing and the associated exposure bias. On a multi-instrument music accompaniment benchmark, LiveBand improves over prior work on objective measures of audio quality, beat alignment, and mix adherence, while enabling real-time streaming generation without lookahead into the future on consumer hardware.
【3】Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation
标题:Foley-Omni:从任务级音频合成到完整视频原声生成的统一多模式生成模型链接:https://arxiv.org/abs/2606.03672
摘要:
摘要:
【4】Tonal parsimony in chord-sequence analysis: combining modulation cost and tonal vocabulary
标题:和弦序列分析中的音调简约:结合调制成本和音调词汇链接:https://arxiv.org/abs/2606.03459
备注:20 pages, 1 figure
摘要:
摘要:
【5】Speech Emotion Recognition using Attention-based LSTM-Network with Residual Connection
标题:使用具有剩余连接的基于注意力的LSTM网络的语音情感识别链接:https://arxiv.org/abs/2606.03359
备注:6 pages, 5 figures, DSPA 2026
摘要:
摘要:
【6】Inference-Time Scaling for Joint Audio-Video Generation
标题:用于联合音频-视频生成的推理时间缩放链接:https://arxiv.org/abs/2606.03183
备注:Accepted by Transactions on Machine Learning Research (TMLR). Project page: this https URL
摘要:
摘要:
【7】SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling
标题:SketchSong:具有草图规划和细粒度多轨建模的分层歌曲生成链接:https://arxiv.org/abs/2606.03169
摘要:
摘要:
【8】Audio Spotforming via Post-Filtering Using Cross-Array Non-target Estimates
标题:通过使用交叉阵列非目标估计的后过滤的音频点形成链接:https://arxiv.org/abs/2606.03028
备注:Accepted for EUSIPCO 2026
摘要:音频点形成是一种利用多个麦克风阵列从噪声混合中提取目标语音的技术。常规方法使用低秩近似从由每个阵列获得的线性分离的信号估计共享目标语音分量,并且基于该估计的低秩表示来应用后滤波(PF)。然而,由于低秩模型之间的不匹配和语音信号的复杂结构,直接依赖于低秩近似PF会降低语音提取性能。在这项研究中,我们利用的观察,从一个阵列的角度来看,位于目标语音方向的非目标组件可以在空间上分离时,从其他阵列。这一见解激发了一种新的点形成方法,用于使用跨阵列的非目标估计而不是依赖于低秩近似的高效滤波后估计。实验表明,该方法优于传统的斑点形成方法。
摘要:Audio spotforming is a technique for extracting target speech from noisy mixtures by utilizing multiple microphone arrays. Conventional methods estimate a shared target speech component from linearly separated signals obtained by each array using low-rank approximations and apply post filtering (PF) based on this estimated low-rank representation. However, owing to the mismatch between low-rank models and the complex structure of speech signals, directly relying on low-rank approximations for PF can degrade the speech extraction performance. In this study, we leverage the observation that non-target components located in the target speech direction from the perspective of one array can be spatially separated when viewed from other arrays. This insight motivates a new spotforming method for efficient post-filter estimation using non-target estimates across arrays instead of relying on low-rank approximations. Experiments demonstrate that the proposed method outperforms conventional spotforming methods.
【9】A Training-Efficient Transformer-Based Anti-Spoofing Network for Logical Access in ASVspoof 5
标题:ASVspoof 5中用于逻辑访问的训练高效的基于转换器的反欺骗网络链接:https://arxiv.org/abs/2606.02980
备注:11 pages, 2 figures
摘要:合成语音和人工语音会降低自动说话人确认系统的可靠性,因此反欺骗方法需要在训练和推理方面既准确又有效。本文重点关注ASVspoof 5 Track 1封闭条件,其中标准交叉熵训练可能没有对硬试验给予足够的关注,并且没有直接与基于排名和阈值的评估指标保持一致。我们提出了TFPARN,一个基于transformer的焦点成对关注排名网络。该系统从语音中提取对数梅尔特征,使用Transformer编码器来建模帧级信息,应用注意力池来获得话语级表示,并使用焦点分类损失和成对排序损失的组合进行训练。在训练期间使用RawBoost增强,在评估期间应用测试时增强以提高鲁棒性。与相同协议下重新实现的AASIST和RawNet2基线相比,TFPARN获得了最佳结果,minDCF为0.2430,EER为12.52%。消融实验进一步表明,成对损失,焦点损失和注意力池都提高了性能。TFPARN还使用了比较系统中最低的推理内存,为1.4 GB,每次发声运行约0.79 ms,并且在比AASIST更少的训练时间内达到其最佳检查点。这些结果表明,TFPARN提供了一个很好的平衡之间的检测精度和计算成本的逻辑访问反欺骗。
摘要:Synthetic and manipulated speech can reduce the reliability of automatic speaker verification systems, so anti-spoofing methods need to be both accurate and efficient in training and inference. This paper focuses on the ASVspoof 5 Track 1 closed condition, where standard cross-entropy training may not give enough attention to hard trials and is not directly aligned with ranking- and threshold-based evaluation metrics. We propose TFPARN, a Transformer-based focal-pairwise attentive ranking network. The system extracts log-Mel features from speech, uses a Transformer encoder to model frame-level information, applies attention pooling to obtain utterance-level representations, and is trained with a combination of focal classification loss and pairwise ranking loss. RawBoost augmentation is used during training, and test-time augmentation is applied during evaluation to improve robustness. Compared with re-implemented AASIST and RawNet2 baselines under the same protocol, TFPARN achieves the best results, with a minDCF of 0.2430 and an EER of 12.52%. Ablation experiments further show that the pairwise loss, focal loss, and attention pooling all improve performance. TFPARN also uses the lowest inference memory among the compared systems, at 1.4 GB, runs at about 0.79 ms per utterance, and reaches its best checkpoint in less training time than AASIST. These results show that TFPARN provides a good balance between detection accuracy and computational cost for logical access anti-spoofing.
【10】EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement
标题:EntangleCodec:通过语义-声学纠缠的统一离散音频令牌器链接:https://arxiv.org/abs/2606.02739
备注:17 pages, 10 figures
摘要:音频分词器作为连续音频和音频语言模型(ALM)之间的离散接口,但现有的分词器往往难以支持理解和生成。面向重构的编解码器保持声学保真度,但缺乏丰富的语义,而语义感知的标记器通常依赖于单独的语义和声学流,从而引入冗余或不对齐。 我们提出了\textbf{EntangleCodec},一个统一的离散音频标记器,在量化之前学习标题对齐的语义声学表示。通过将音频与丰富的字幕而不是ASR转录对齐,EntangleCodec在紧凑的令牌流中捕获语言内容,说话者身份,情感,韵律和声学场景。流匹配扩散解码器进一步实现跨语音、音乐和一般音频的高质量重构。 EntangleCodec实现了与专业编解码器竞争的重建质量,在MMAR上的音频理解方面优于所有基于编解码器的基线,最高可达\textbf{+7.4\%},并在统一的框架中支持TTS和TTA生成。此外,基于EntangleCodec的音频语言模型表现出强大的缩放行为:即使在\textit{0.6B}参数下,该模型在三个基准测试中使用更少的参数,超过了具有超过\textit{13 B}参数的专用连续表示LLM;缩放到\textit{8B}进一步建立了关于MMAR的新的最先进结果,强调了表示质量与音频语言建模中的模型比例一样重要。代码和型号重量可在https://github.com/luckyerr/EntangleCodec上获得。
摘要:Audio tokenizers serve as the discrete interface between continuous audio and Audio Language Models (ALMs), but existing tokenizers often struggle to support both understanding and generation. Reconstruction-oriented codecs preserve acoustic fidelity but lack rich semantics, while semantic-aware tokenizers typically rely on separate semantic and acoustic streams, introducing redundancy or misalignment. We propose \textbf{EntangleCodec}, a unified discrete audio tokenizer that learns caption-aligned semantic-acoustic representations before quantization. By aligning audio with rich captions rather than ASR transcripts, EntangleCodec captures linguistic content, speaker identity, emotion, prosody, and acoustic scenes within a compact token stream. A flow-matching diffusion decoder further enables high-quality reconstruction across speech, music, and general audio. EntangleCodec achieves reconstruction quality competitive with specialized codecs, outperforms all codec-based baselines on audio understanding by up to \textbf{+7.4\%} on MMAR, and supports both TTS and TTA generation in a unified framework. Furthermore, EntangleCodec-based audio language models demonstrate strong scaling behavior: even at \textit{0.6B} parameters, the model surpasses specialized continuous-representation LLMs with over \textit{13B} parameters across three benchmarks using \textbf{22$\times$} fewer parameters; scaling to \textit{8B} further establishes new state-of-the-art results on MMAR, highlighting that representation quality is as critical as model scale in audio language modeling. Code and model weights are available at https://github.com/luckyerr/EntangleCodec.
【11】Before Fusion, Ask What to Keep: Contextual Calibration of Multimodal Signals
标题:Fusion之前,询问要保留什么:多模式信号的上下文校准链接:https://arxiv.org/abs/2606.02679
备注:11 pages, 7 figures, 9 tables
摘要:多模态系统通常受益于跨语言,声音和视觉流的信息组合,但这种好处并不能保证。对一种输入有用的模态可能会分散另一种输入的注意力,并且同一模态内的局部特征响应可能与来自其他来源的证据不一致。这项工作研究了如何调整多模态表示之前,他们合并的下游预测。我们开发了一个紧凑的校准模块,比较每个模态与其他的汇总水平,提取线索的跨源支持和冲突,并将这些线索转换成实例和维度的调制信号。校准应用于原始模态特征,而不是已经融合的表示,使模型能够抑制误导性成分,保留弱但有用的证据,并强调当前多模态上下文更好地支持的响应。该模块被设计为一个插件组件,可以连接到不同的融合骨干,而无需改变他们的预测头。在涵盖情感理解、动作识别、视听事件检测和视听情感分类的五个基准测试中,所提出的预组合校准策略提高了基于序列和卷积融合设置下的性能。在模态移除、合成腐败、训练动态和特征级可视化下的其他分析表明,在融合之前校准信号可以减少不可靠模态的干扰,并产生更稳定的多模态优化。
摘要:Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed. A modality that is useful for one input may become distracting for another, and local feature responses within the same modality can disagree with evidence from other sources. This work investigates how to adjust multimodal representations before they are merged by a downstream predictor. We develop a compact calibration module that compares each modality with the others at the summary level, extracts cues of cross-source support and conflict, and converts these cues into instance-wise and dimension-wise modulation signals. The calibration is applied to the original modality features rather than to already fused representations, enabling the model to suppress misleading components, preserve weak but useful evidence, and emphasize responses that are better supported by the current multimodal context. The module is designed as a plug-in component and can be attached to different fusion backbones without changing their prediction heads. Across five benchmarks covering sentiment understanding, action recognition, audio-visual event detection, and audio-visual emotion classification, the proposed pre-combination calibration strategy improves performance under both sequence-based and convolutional fusion settings. Additional analyses under modality removal, synthetic corruption, training dynamics, and feature-level visualization show that calibrating signals before fusion can reduce interference from unreliable modalities and produce more stable multimodal optimization.
【12】SegTune: Structured and Fine-Grained Control for Song Generation
标题:SegButton:歌曲生成的结构化和细粒度控制链接:https://arxiv.org/abs/2606.02638
备注:This paper has been accepted to ACL 2026 as an oral presentation and has been nominated for the Best Paper Award. This work is a revised and extended version of an earlier technical report (arXiv:2510.18416). arXiv admin note: text overlap with arXiv:2510.18416
摘要:神经歌曲生成的最新进展使歌词和全局文本提示的高质量合成成为可能。然而,大多数系统无法对歌曲的时间变化属性进行建模,严重限制了对音乐结构和动态的细粒度控制。为了解决这个问题,我们提出了SegTune,一个基于扩散变换器的框架,通过允许用户或大型语言模型(LLM)指定与歌曲片段对齐的本地音乐描述,实现结构化和细粒度的可控性。这些片段提示在时间上广播到对应的时间窗口,而全局提示确保风格一致性。为了支持精确的歌词到音乐的对齐,我们引入了一个基于LLM的持续时间预测器,该预测器以LyRiCs格式自回归生成音乐级时间戳。我们进一步构建了一个大规模的数据管道,用于高质量的歌曲收集与对齐的歌词和提示,并提出了新的指标来评估段对齐和声乐一致性。实验表明,SegTune优于现有的基线在音乐性和可控性。请访问我们的项目页面(https://github.com/KlingAIResearch/SegTune)获取代码和更多生成的歌曲。
摘要:Recent advances in neural song generation have enabled high-quality synthesis from lyrics and global textual prompts. However, most systems fail to model temporally varying attributes of songs, severely limiting fine-grained control over musical structure and dynamics. To address this, we propose SegTune, a Diffusion Transformer-based framework enabling structured and fine-grained controllability by allowing users or large language models (LLMs) to specify local musical descriptions aligned to song segments. These segment prompts are temporally broadcast to corresponding time windows, while global prompts ensure stylistic coherence. To support precise lyric-to-music alignment, we introduce an LLM-based duration predictor that autoregressively generates sentence-level timestamps in LyRiCs format. We further construct a large-scale data pipeline for high-quality song collection with aligned lyrics and prompts, and propose new metrics to evaluate segment alignment and vocal consistency. Experiments demonstrate that SegTune outperforms existing baselines in both musicality and controllability. Visit our project page (https://github.com/KlingAIResearch/SegTune) for codes and more generated songs.
【13】WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling
标题:WavTTC:通过直接原始波建模实现高质量Zero-ShotTTC链接:https://arxiv.org/abs/2606.03455
摘要:近年来,基于VAE潜伏期或Mel谱图的扩散模型已经成为zero-shot TTS的主要范例。虽然这些压缩表示提高了生成效率,但它们不可避免地遭受信息丢失和非端到端训练。从理论上讲,直接对原始波形进行建模可以避免这些问题;然而,由于音频信号的序列长度非常长,这个方向仍然没有得到充分的探索,并且通常被认为是困难的。为了克服这一点,我们提出了WavTTS,第一个原始波形生成TTS模型,大大缩小了与潜在空间生成模型的差距。WavTTS建立在与Diffusion Transformer(DiT)的流匹配的基础上,通过简单的拼接策略直接对语音波形进行建模,同时集成多尺度梅尔频谱图监督,以在训练期间提供感知指导。此外,我们调查的预测目标和噪声调度波形扩散的影响,并开发一个有效的时间表设计,以提高发电质量。对开源基准的评估表明,WavTTS接近当前最先进的潜在生成zero-shot TTS模型的性能,同时大大优于以前的端到端语音生成模型。我们的研究结果表明直接在波形空间中缩放基于扩散的TTS的可行性,为端到端语音生成开辟了新的方向。
摘要:Recently, diffusion models operating on VAE latents or mel-spectrograms have become the dominant paradigm for zero-shot TTS. Although these compressed representations improve generation efficiency, they inevitably suffer from information loss and non-end-to-end training. Theoretically, directly modeling raw waveforms circumvents these issues; however, this direction remains underexplored and is often deemed difficult due to the extremely long sequence length of audio signals. To overcome this, we propose WavTTS, the first raw waveform generative TTS model that substantially narrows the gap with latent-space generative models. Built upon the flow matching with Diffusion Transformer (DiT), WavTTS directly models speech waveforms via a simple patchification strategy, while integrating multi-scale mel-spectrogram supervision to provide perceptual guidance during training. Furthermore, we investigate the impact of prediction targets and noise scheduling in waveform diffusion, and develop an effective schedule design to improve generation quality. Evaluations on open-source benchmarks demonstrate that WavTTS closely approaches the performance of current state-of-the-art latent generative zero-shot TTS models, while substantially outperforming previous end-to-end speech generation models. Our findings demonstrate the feasibility of scaling diffusion-based TTS directly in the waveform space, opening a new direction for end-to-end speech generation.
【14】SpeakerCard-1M: An Evidence-Grounded Speaker Card Corpus for In-the-Wild Speaker Verification
标题:SpeakerCard-1 M:一个基于证据的说话人卡片语料库链接:https://arxiv.org/abs/2606.03283
备注:Corpus and protocols at https://junyipeng00.github.io/SpeakerCard-1M-page
摘要:现代说话人确认(SV)系统依赖于有效但难以用自然语言解释或查询的说话人嵌入。大多数现有的语音文本语料库的目标可控合成或话语级字幕,并提供有限的说话人级的监督,在野生说话人识别。本文介绍了SpeakerCard-1 M,一个以说话者为中心的双语资源,以证据为基础的SV,来自VoxCeleb 1/2和CN-Celeb 1/2,其中“-1M”后缀是指1.78 M的话语级字幕包含在释放。我们采用了一个工具第一,LLM最后的方法:10声探头产生字段级的证据,证据被聚合到扬声器配置文件下的模式,分离相对稳定的特征从话语级状态,和双语扬声器卡呈现由一个受约束的LLM,只看到结构化的字段。该版本包括超过10.2K扬声器的56.7K扬声器卡记录,1.78M话语级字幕和扬声器ID不相交的硬否定三元组。我们进一步定义了两个面向SV的跨模态协议,双向说话人文本检索(T2 S-R/S2 T-R)和属性条件验证(AC-验证),并比较了双编码器基线对最近的音频语言模型下的zero-shot强制选择设置。联合音频文本训练使VoxCeleb 1-O EER比仅音频基线增加了0.31%。在风格对称的LLM生成的反事实协议下,八个最近的音频语言模型(7 B-30 B+参数,开源和闭源)在双向强制选择下的音高级别AC验证得分为49-77%,而我们的双编码器达到了88.66%。
摘要:Modern speaker verification (SV) systems rely on speaker embeddings that are effective but difficult to interpret or query in natural language. Most existing speech-text corpora target controllable synthesis or utterance-level captioning, and provide limited speaker-level supervision for in-the-wild speaker recognition. This paper introduces SpeakerCard-1M, a bilingual speaker-centric resource for evidence-grounded SV, derived from VoxCeleb1/2 and CN-Celeb1/2, where the "-1M" suffix refers to the 1.78M utterance-level captions contained in the release. We adopt a tool-first, LLM-last approach: ten acoustic probes produce field-level evidence, the evidence is aggregated into speaker profiles under a schema that separates relatively stable traits from utterance-level states, and bilingual Speaker Cards are rendered by a constrained LLM that sees only the structured fields. The release includes 56.7K Speaker Card records over 10.2K speakers, 1.78M utterance-level captions, and speaker-ID-disjoint hard-negative triplets. We further define two SV-oriented cross-modal protocols, bidirectional Speaker-Text Retrieval (T2S-R / S2T-R) and Attribute-Conditioned Verification (AC-Verify), and compare a dual-encoder baseline against recent audio language models under a zero-shot forced-choice setting. Joint audio-text training increases VoxCeleb1-O EER by 0.31% absolute over the audio-only baseline. Under a style-symmetric LLM-generated counterfactual protocol, eight recent audio language models (7B-30B+ parameters, both open- and closed-source) score 49-77% on pitch-level AC-Verify under two-way forced choice, compared with 88.66% reached by our dual encoder.
【15】AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following
标题:AnyAudio-Judge:一个基于规则的动态音频教学跟踪基准和评估器链接:https://arxiv.org/abs/2606.03116
摘要:导航音频生成的快速发展突出了对鲁棒对准评估的迫切需求。目前的自动评估方法严重依赖于通用大型语言模型的整体评分,这些模型难以解耦复杂的指令,缺乏可解释性,并且无法捕获细粒度的属性不匹配。为了解决这个问题,我们引入了一种新的动态的基于规则的评价范式,自适应复杂的音频字幕分解成一个独立的,可验证的二进制规则项的可变数量。为了严格地对这种能力进行基准测试,我们提出了AnyAudio-Judge Bench,这是一个全面的双语基准测试,包括四个不同音频领域(语音,声音,音乐和混合)的7,920个精心策划的样本,具有故意构建的硬底片。此外,我们构建了一个大规模的语料库105 K的样本与明确的思想链(CoT)的理由训练我们的专用评估,AnyAudio-Judge模型。通过采用结合监督微调(SFT)和组相对策略优化(GRPO)的训练管道,我们的模型成功地将其推理路径与基于规则的评分机制相结合。大量的实验表明,AnyAudio-Judge不仅显著增强了与最先进的基线相比的zero-shot对齐检测,而且还提供了精确和可解释的奖励信号,大大改善了下游强化学习中的指令对齐,用于音频生成。
摘要:The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current automated evaluation methods heavily rely on holistic scoring from general-purpose large language models, which struggle to decouple complex instructions, lack interpretability, and fail to capture fine-grained attribute mismatches. To address this, we introduce a novel dynamic rubric-based evaluation paradigm that adaptively decomposes complex audio captions into a variable number of independent, verifiable binary rubric items. To rigorously benchmark this capability, we propose the AnyAudio-Judge Bench, a comprehensive, bilingual benchmark comprising 7,920 meticulously curated samples across four diverse audio domains (speech, sound, music, and mixed), featuring deliberately constructed hard negatives. Furthermore, we construct a large-scale corpus of 105K samples with explicit Chain-of-Thought (CoT) rationales to train our dedicated evaluator, the AnyAudio-Judge model. By employing a training pipeline that combines Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO), our model successfully aligns its reasoning paths with the rubric-based scoring mechanism. Extensive experiments demonstrate that AnyAudio-Judge not only significantly enhances zero-shot alignment detection compared to state-of-the-art baselines, but also provides precise and interpretable reward signals that substantially improve instruction alignment in downstream reinforcement learning for audio generation.
【16】A Comparison of Generative and Discriminative Methods for Speech Enhancement: Robustness, Complexity, and Hallucination
标题:语音增强生成方法和区分方法的比较:稳健性、复杂性和幻觉链接:https://arxiv.org/abs/2606.02913
摘要:在这项研究中,我们对基于生成式和判别式深度学习的语音增强方法进行了全面的比较分析,特别是在降噪任务中。我们的调查重点是评估他们的有效性,在高和低信噪比的条件下,考虑匹配和不匹配的训练方案。我们进一步研究了训练数据量,模型收敛速度的影响,并解释了所考虑的训练范式的客观结果方面的性能差异。此外,我们比较了这些方法的复杂性和性能的权衡和实际可行性。为了进一步加强评价,我们研究了生成方法的幻觉特征,包括单词错误率和音素相似性。从这项研究中得出的见解提供了经验证据,以帮助研究人员和从业人员了解不同方法的感知增益是否证明其在实际应用中的计算成本是合理的。
摘要:In this study, we conduct a comprehensive comparative analysis of generative and discriminative deep learning-based speech enhancement methods, specifically in noise reduction tasks. Our investigation focuses on evaluating their effectiveness under high and low signal-to-noise ratio conditions, considering both matched and mismatched training scenarios. We further investigate the impact of training data volume, model convergence speed, and interpret the performance differences in terms of objective results for the considered training paradigms. Additionally, we compare the complexity-performance trade-off and the practical viability of these approaches. To further strengthen the evaluation, we study the hallucination characteristics of generative approaches in terms of word error rate and phoneme similarity. The insights derived from this study provide empirical evidence to assist researchers and practitioners in understanding whether the perceptual gains of different approaches justify their computational cost in practical applications.
【17】SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models
标题:SVHalluc:视听大型语言模型中的言语视觉幻觉基准链接:https://arxiv.org/abs/2606.02642
备注:Accepted at CVPR 2026
摘要:尽管视听大语言模型(LLM)取得了成功,但它们可以产生看似合理但没有根据的输出,称为幻觉。现有基准侧重于环境声音(例如,狗吠)以指示事件发生。相比之下,人类语音具有根本不同的、丰富的语义和时间结构,但目前的模型是否能准确地将语音内容与相应的视觉信号对齐,仍有待探索。在这项工作中,我们表明,语音内容可以诱导视听LLM的幻觉。为了系统地研究这一点,我们介绍了SVHalluc,这是第一个用于评估视听LLM中的语音-视觉幻觉的综合基准。我们的基准诊断言语视觉幻觉从两个关键和互补的方面:语义和时间。实验结果表明,最先进的开源视听LLM努力将语音内容与相应的视觉信号对齐,在多个任务中具有近乎随机的准确性。相比之下,Gemini 2.5 Pro明显优于开源模型。我们的分析表明,他们的失败源于有限的能力,在跨通道的理解,尽管在单通道的知觉表现强劲。我们的工作揭示了当前视听LLM的一个新的和根本的限制,并强调了对基于语音的视频理解的需要。项目页面:https://chenshuang-zhang.github.io/projects/svhalluc/。
摘要:Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmental sounds (e.g., dog barking) to indicate event occurrence. In contrast, human speech carries fundamentally different, rich semantics and temporal structures, yet it remains unexplored whether current models can accurately align speech content with corresponding visual signals. In this work, we show that speech content can induce hallucinations in audio-visual LLMs. To systematically study this, we introduce SVHalluc, the first comprehensive benchmark for evaluating speech-vision hallucination in audio-visual LLMs. Our benchmark diagnoses speech-vision hallucinations from two critical and complementary aspects: semantic and temporal. Experimental results demonstrate that state-of-the-art open-source audio-visual LLMs struggle with aligning speech content with corresponding visual signals, with a near-random accuracy on multiple tasks. In contrast, Gemini 2.5 Pro significantly outperforms the open-source models. Our analysis suggests that their failures stem from limited ability in cross-modality understanding, despite strong performance in single-modality perception. Our work uncovers a new and fundamental limitation of current audio-visual LLMs and highlights the need for speech-grounded video comprehension. Project page: https://chenshuang-zhang.github.io/projects/svhalluc/.
【18】Wavelet as Tokenizer: Preliminary Results on a Shared Wavelet Token Schema for Natural Signals
标题:作为令牌化器的子波:自然信号共享子波令牌模式的初步结果链接:https://arxiv.org/abs/2606.02631
备注:12 pages, 3 figures
摘要:本文研究音频,图像和视频是否可以共享一个共同的小波令牌模式,而不是依赖于单独的模态特定的潜在网格。它介绍了一个初步的连续令牌模型,该模型围绕一级Haar DWT/IDWT前端、共享系数令牌布局、可选的结构元数据、轻量级模态值适配器和共享令牌式编码器-解码器主干构建。在语音命令、EuroSAT RGB和DAVIS 2017数据上,密集共享模型可达到39.92 dB音频、29.37 dB图像和23.93 dB视频PSNR。连续潜在标量预算下的匹配率扫描表明,视觉增益并不能仅仅由潜在容量来解释,同时也表明添加剂元数据嵌入并不是一个普遍的改进来源。最后,固定速率的能量选择提供了一个强大的非参数基线:energy_global在压缩保持率下将音频的平均PSNR提高了16.73 dB,图像提高了16.90 dB,视频提高了15.86 dB。屏蔽稀疏训练达到34.45 dB的视频PSNR与50%的密集令牌。结果支持一个统一的小波令牌模式和稀疏令牌接口,而停止建立一个通用的离散词汇表。
摘要:This paper studies whether audio, images, and video can share a common wavelet token schema rather than relying on separate modality-specific latent grids. It introduces a preliminary continuous-token model built around a one-level Haar DWT/IDWT frontend, a shared coefficient-token layout, optional structural metadata, lightweight modality value adapters, and a shared token-wise encoder-decoder trunk. On Speech Commands, EuroSAT RGB, and DAVIS 2017 data, a dense shared model reaches 39.92 dB audio, 29.37 dB image, and 23.93 dB video PSNR. A matched-rate sweep under continuous latent scalar budgets indicates that the visual gains are not explained solely by latent capacity, while also showing that additive metadata embeddings are not a universal source of improvement. Finally, fixed-rate energy selection provides a strong non-parametric baseline: energy_global improves average PSNR over uniform selection by 16.73 dB for audio, 16.90 dB for images, and 15.86 dB for video under compressed keep ratios. Masked sparse training reaches 34.45 dB video PSNR with 50% of dense tokens. The results support a unified wavelet token schema and sparse token interface, while stopping short of establishing a universal discrete vocabulary.
【19】FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations
标题:FSA-GRPO:教听觉LLM使用Few-Shot演示链接:https://arxiv.org/abs/2606.02615
摘要:Few-Shot提示提供了一种有效的方法来使听觉大型语言模型适应低资源任务,例如儿童语音识别。然而,大多数听觉大语言模型没有被明确地训练来以这种演示条件格式执行推理,这限制了它们可以从Few-Shot提示中受益的程度。为了解决这一限制,我们引入了Few-Shot Aware GRPO(FSA-GRPO),这是一种基于RL的后训练配方,使用专门设计的奖励来鼓励模型利用Few-Shot演示,从而增强其Few-Shot适应能力。值得注意的是,仅使用高资源成人ASR数据进行训练可以提高模型的一般Few-Shot适应能力,不仅在儿童语音识别方面,而且在语音翻译和音频理解方面都有收获。我们进一步研究了数据选择和辅助奖励加权,以确定有效的训练配方。我们的实验表明,当域内数据不可用或不能用于训练时,FSA-GRPO比直接调整相关的域外数据更有效。
摘要:Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognition. However, most auditory large language models are not explicitly trained to perform inference in this demonstration-conditioned format, limiting the extent to which they can benefit from few-shot prompting. To address this limitation, we introduce Few-Shot Aware GRPO (FSA-GRPO), an RL-based post-training recipe that uses a specially designed reward to encourage the model to leverage few-shot demonstrations, thereby strengthening its few-shot adaptation ability. Notably, training with only high-resource adult ASR data improves the model's general few-shot adaptation ability, yielding gains not only in children's speech recognition but also in speech translation and audio understanding. We further study data selection and auxiliary reward weighting to identify an effective training recipe. Our experiments show that when in-domain data are unavailable or cannot be used for training, FSA-GRPO is more effective than direct tuning on related out-of-domain data.
标题:助听器深度反馈消除的环内训练
链接:https://arxiv.org/abs/2606.03832
摘要:声反馈限制了助听器的最大增益。除了几种基于自适应滤波的方法之外,最近还提出了一种基于深度神经网络的反馈消除(DFC)方法,该方法通过开环框架进行训练。由于开环训练的DFC(DFC-OL)在高增益的推理过程中可能会变得不稳定,因此在本文中,我们提出了一种在环训练的DFC(DFC-IL),它将DFC直接集成到优化循环中。这允许模型在训练期间暴露于不稳定的条件。一个两阶段的训练策略,包括对稳定系统的预训练和对更宽增益范围的微调,使DFC-IL能够学习鲁棒的啸叫减少。实验结果表明,在小增益的情况下,所提出的DFC-IL执行类似的DFC-OL,都超过了自适应滤波器的性能。在具有高放大增益的场景中,DFC-IL通过保持系统稳定性而明显优于DFC-OL。
摘要:Acoustic feedback limits the maximum gain in hearing aids. In addition to several approaches based on adaptive filtering, recently a deep-neural-network-based feedback cancellation (DFC) approach has been proposed, which is trained via an open-loop framework. Since open-loop-trained DFC (DFC-OL) can become unstable during inference at high gains, in this paper we propose an in-the-loop-trained DFC (DFC-IL) that integrates the DFC directly into the optimisation loop. This allows the model to be exposed to unstable conditions during training. A two-stage training strategy involving pre-training on stable systems and fine-tuning on a wider gain range enables DFC-IL to learn robust howling reduction. Experimental results on measured feedback paths demonstrate that in scenarios with small gains, the proposed DFC-IL performs similarly to DFC-OL, and both exceed the performance of adaptive filters. In scenarios with high amplification gains, DFC-IL clearly outperforms DFC-OL by maintaining system stability.
【2】Stable Hybrid Cross-Attention Fusion for Audio-Visual Event Recognition
标题:用于视听事件识别的稳定混合交叉注意融合链接:https://arxiv.org/abs/2606.03747
备注:6 pages, 4 Figures
摘要:视听事件识别(AVER)对于智能城市监控系统至关重要,因为这些系统需要对复杂环境进行强大的多模态理解。本文提出了一种稳定的混合交叉注意融合框架,用于智能城市环境中的视听事件识别。该架构结合了预训练的视频掩蔽自动编码器(VideoMAE)和音频频谱图Transformer(AST)表示与基于FiLM的音频调节,双向交叉注意融合,多模态Transformer编码和模态时间注意。为了提高计算效率和训练的稳定性,冻结预训练的骨干和缓存的特征提取。在AVE数据集上进行的大量实验表明,所提出的框架在多个评估指标的评估单峰和多峰基线中实现了最高的平均性能,在五次独立运行中获得了91.74%的最佳验证准确度和83.85 ± 1.40%的测试准确度。结果表明,所提出的混合融合策略有效地捕捉互补的视听信息,并提供了强大的多模态表示学习具有挑战性的现实世界的城市监控场景。
摘要:Audio-Visual Event Recognition (AVER) is essential for intelligent urban monitoring systems, where robust multimodal understanding of complex environments is required. This paper proposes a stable hybrid cross-attention fusion framework for audio-visual event recognition in smart urban environments. The proposed architecture combines pretrained Video Masked Autoencoder (VideoMAE) and Audio Spectrogram Transformer (AST) representations with FiLM-based audio conditioning, bidirectional cross-attention fusion, multimodal Transformer encoding, and modality-temporal attention. To improve computational efficiency and training stability, frozen pretrained backbones and cached feature extraction are employed. Extensive experiments on the AVE dataset show that the proposed framework achieves the highest average performance among the evaluated unimodal and multimodal baselines across multiple evaluation metrics, obtaining a best validation accuracy of 91.74% and a test accuracy of 83.85 plus/minus 1.40% over five independent runs. The results indicate that the proposed hybrid fusion strategy effectively captures complementary audio-visual information and provides robust multimodal representation learning for challenging realworld urban monitoring scenarios.
【3】WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling
标题:WavTTC:通过直接原始波建模实现高质量Zero-ShotTTC链接:https://arxiv.org/abs/2606.03455
摘要:近年来,基于VAE潜伏期或Mel谱图的扩散模型已经成为zero-shot TTS的主要范例。虽然这些压缩表示提高了生成效率,但它们不可避免地遭受信息丢失和非端到端训练。从理论上讲,直接对原始波形进行建模可以避免这些问题;然而,由于音频信号的序列长度非常长,这个方向仍然没有得到充分的探索,并且通常被认为是困难的。为了克服这一点,我们提出了WavTTS,第一个原始波形生成TTS模型,大大缩小了与潜在空间生成模型的差距。WavTTS建立在与Diffusion Transformer(DiT)的流匹配的基础上,通过简单的拼接策略直接对语音波形进行建模,同时集成多尺度梅尔频谱图监督,以在训练期间提供感知指导。此外,我们调查的预测目标和噪声调度波形扩散的影响,并开发一个有效的时间表设计,以提高发电质量。对开源基准的评估表明,WavTTS接近当前最先进的潜在生成zero-shot TTS模型的性能,同时大大优于以前的端到端语音生成模型。我们的研究结果表明直接在波形空间中缩放基于扩散的TTS的可行性,为端到端语音生成开辟了新的方向。
摘要:Recently, diffusion models operating on VAE latents or mel-spectrograms have become the dominant paradigm for zero-shot TTS. Although these compressed representations improve generation efficiency, they inevitably suffer from information loss and non-end-to-end training. Theoretically, directly modeling raw waveforms circumvents these issues; however, this direction remains underexplored and is often deemed difficult due to the extremely long sequence length of audio signals. To overcome this, we propose WavTTS, the first raw waveform generative TTS model that substantially narrows the gap with latent-space generative models. Built upon the flow matching with Diffusion Transformer (DiT), WavTTS directly models speech waveforms via a simple patchification strategy, while integrating multi-scale mel-spectrogram supervision to provide perceptual guidance during training. Furthermore, we investigate the impact of prediction targets and noise scheduling in waveform diffusion, and develop an effective schedule design to improve generation quality. Evaluations on open-source benchmarks demonstrate that WavTTS closely approaches the performance of current state-of-the-art latent generative zero-shot TTS models, while substantially outperforming previous end-to-end speech generation models. Our findings demonstrate the feasibility of scaling diffusion-based TTS directly in the waveform space, opening a new direction for end-to-end speech generation.
【4】SpeakerCard-1M: An Evidence-Grounded Speaker Card Corpus for In-the-Wild Speaker Verification
标题:SpeakerCard-1 M:一个基于证据的说话人卡片语料库链接:https://arxiv.org/abs/2606.03283
备注:Corpus and protocols at https://junyipeng00.github.io/SpeakerCard-1M-page
摘要:现代说话人确认(SV)系统依赖于有效但难以用自然语言解释或查询的说话人嵌入。大多数现有的语音文本语料库的目标可控合成或话语级字幕,并提供有限的说话人级的监督,在野生说话人识别。本文介绍了SpeakerCard-1 M,一个以说话者为中心的双语资源,以证据为基础的SV,来自VoxCeleb 1/2和CN-Celeb 1/2,其中“-1M”后缀是指1.78 M的话语级字幕包含在释放。我们采用了一个工具第一,LLM最后的方法:10声探头产生字段级的证据,证据被聚合到扬声器配置文件下的模式,分离相对稳定的特征从话语级状态,和双语扬声器卡呈现由一个受约束的LLM,只看到结构化的字段。该版本包括超过10.2K扬声器的56.7K扬声器卡记录,1.78M话语级字幕和扬声器ID不相交的硬否定三元组。我们进一步定义了两个面向SV的跨模态协议,双向说话人文本检索(T2 S-R/S2 T-R)和属性条件验证(AC-验证),并比较了双编码器基线对最近的音频语言模型下的zero-shot强制选择设置。联合音频文本训练使VoxCeleb 1-O EER比仅音频基线增加了0.31%。在风格对称的LLM生成的反事实协议下,八个最近的音频语言模型(7 B-30 B+参数,开源和闭源)在双向强制选择下的音高级别AC验证得分为49-77%,而我们的双编码器达到了88.66%。
摘要:Modern speaker verification (SV) systems rely on speaker embeddings that are effective but difficult to interpret or query in natural language. Most existing speech-text corpora target controllable synthesis or utterance-level captioning, and provide limited speaker-level supervision for in-the-wild speaker recognition. This paper introduces SpeakerCard-1M, a bilingual speaker-centric resource for evidence-grounded SV, derived from VoxCeleb1/2 and CN-Celeb1/2, where the "-1M" suffix refers to the 1.78M utterance-level captions contained in the release. We adopt a tool-first, LLM-last approach: ten acoustic probes produce field-level evidence, the evidence is aggregated into speaker profiles under a schema that separates relatively stable traits from utterance-level states, and bilingual Speaker Cards are rendered by a constrained LLM that sees only the structured fields. The release includes 56.7K Speaker Card records over 10.2K speakers, 1.78M utterance-level captions, and speaker-ID-disjoint hard-negative triplets. We further define two SV-oriented cross-modal protocols, bidirectional Speaker-Text Retrieval (T2S-R / S2T-R) and Attribute-Conditioned Verification (AC-Verify), and compare a dual-encoder baseline against recent audio language models under a zero-shot forced-choice setting. Joint audio-text training increases VoxCeleb1-O EER by 0.31% absolute over the audio-only baseline. Under a style-symmetric LLM-generated counterfactual protocol, eight recent audio language models (7B-30B+ parameters, both open- and closed-source) score 49-77% on pitch-level AC-Verify under two-way forced choice, compared with 88.66% reached by our dual encoder.
【5】AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following
标题:AnyAudio-Judge:一个基于规则的动态音频教学跟踪基准和评估器链接:https://arxiv.org/abs/2606.03116
摘要:导航音频生成的快速发展突出了对鲁棒对准评估的迫切需求。目前的自动评估方法严重依赖于通用大型语言模型的整体评分,这些模型难以解耦复杂的指令,缺乏可解释性,并且无法捕获细粒度的属性不匹配。为了解决这个问题,我们引入了一种新的动态的基于规则的评价范式,自适应复杂的音频字幕分解成一个独立的,可验证的二进制规则项的可变数量。为了严格地对这种能力进行基准测试,我们提出了AnyAudio-Judge Bench,这是一个全面的双语基准测试,包括四个不同音频领域(语音,声音,音乐和混合)的7,920个精心策划的样本,具有故意构建的硬底片。此外,我们构建了一个大规模的语料库105 K的样本与明确的思想链(CoT)的理由训练我们的专用评估,AnyAudio-Judge模型。通过采用结合监督微调(SFT)和组相对策略优化(GRPO)的训练管道,我们的模型成功地将其推理路径与基于规则的评分机制相结合。大量的实验表明,AnyAudio-Judge不仅显著增强了与最先进的基线相比的zero-shot对齐检测,而且还提供了精确和可解释的奖励信号,大大改善了下游强化学习中的指令对齐,用于音频生成。
摘要:The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current automated evaluation methods heavily rely on holistic scoring from general-purpose large language models, which struggle to decouple complex instructions, lack interpretability, and fail to capture fine-grained attribute mismatches. To address this, we introduce a novel dynamic rubric-based evaluation paradigm that adaptively decomposes complex audio captions into a variable number of independent, verifiable binary rubric items. To rigorously benchmark this capability, we propose the AnyAudio-Judge Bench, a comprehensive, bilingual benchmark comprising 7,920 meticulously curated samples across four diverse audio domains (speech, sound, music, and mixed), featuring deliberately constructed hard negatives. Furthermore, we construct a large-scale corpus of 105K samples with explicit Chain-of-Thought (CoT) rationales to train our dedicated evaluator, the AnyAudio-Judge model. By employing a training pipeline that combines Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO), our model successfully aligns its reasoning paths with the rubric-based scoring mechanism. Extensive experiments demonstrate that AnyAudio-Judge not only significantly enhances zero-shot alignment detection compared to state-of-the-art baselines, but also provides precise and interpretable reward signals that substantially improve instruction alignment in downstream reinforcement learning for audio generation.
【6】A Comparison of Generative and Discriminative Methods for Speech Enhancement: Robustness, Complexity, and Hallucination
标题:语音增强生成方法和区分方法的比较:稳健性、复杂性和幻觉链接:https://arxiv.org/abs/2606.02913
摘要:在这项研究中,我们对基于生成式和判别式深度学习的语音增强方法进行了全面的比较分析,特别是在降噪任务中。我们的调查重点是评估他们的有效性,在高和低信噪比的条件下,考虑匹配和不匹配的训练方案。我们进一步研究了训练数据量,模型收敛速度的影响,并解释了所考虑的训练范式的客观结果方面的性能差异。此外,我们比较了这些方法的复杂性和性能的权衡和实际可行性。为了进一步加强评价,我们研究了生成方法的幻觉特征,包括单词错误率和音素相似性。从这项研究中得出的见解提供了经验证据,以帮助研究人员和从业人员了解不同方法的感知增益是否证明其在实际应用中的计算成本是合理的。
摘要:In this study, we conduct a comprehensive comparative analysis of generative and discriminative deep learning-based speech enhancement methods, specifically in noise reduction tasks. Our investigation focuses on evaluating their effectiveness under high and low signal-to-noise ratio conditions, considering both matched and mismatched training scenarios. We further investigate the impact of training data volume, model convergence speed, and interpret the performance differences in terms of objective results for the considered training paradigms. Additionally, we compare the complexity-performance trade-off and the practical viability of these approaches. To further strengthen the evaluation, we study the hallucination characteristics of generative approaches in terms of word error rate and phoneme similarity. The insights derived from this study provide empirical evidence to assist researchers and practitioners in understanding whether the perceptual gains of different approaches justify their computational cost in practical applications.
【7】SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models
标题:SVHalluc:视听大型语言模型中的言语视觉幻觉基准链接:https://arxiv.org/abs/2606.02642
备注:Accepted at CVPR 2026
摘要:尽管视听大语言模型(LLM)取得了成功,但它们可以产生看似合理但没有根据的输出,称为幻觉。现有基准侧重于环境声音(例如,狗吠)以指示事件发生。相比之下,人类语音具有根本不同的、丰富的语义和时间结构,但目前的模型是否能准确地将语音内容与相应的视觉信号对齐,仍有待探索。在这项工作中,我们表明,语音内容可以诱导视听LLM的幻觉。为了系统地研究这一点,我们介绍了SVHalluc,这是第一个用于评估视听LLM中的语音-视觉幻觉的综合基准。我们的基准诊断言语视觉幻觉从两个关键和互补的方面:语义和时间。实验结果表明,最先进的开源视听LLM努力将语音内容与相应的视觉信号对齐,在多个任务中具有近乎随机的准确性。相比之下,Gemini 2.5 Pro明显优于开源模型。我们的分析表明,他们的失败源于有限的能力,在跨通道的理解,尽管在单通道的知觉表现强劲。我们的工作揭示了当前视听LLM的一个新的和根本的限制,并强调了对基于语音的视频理解的需要。项目页面:https://chenshuang-zhang.github.io/projects/svhalluc/。
摘要:Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmental sounds (e.g., dog barking) to indicate event occurrence. In contrast, human speech carries fundamentally different, rich semantics and temporal structures, yet it remains unexplored whether current models can accurately align speech content with corresponding visual signals. In this work, we show that speech content can induce hallucinations in audio-visual LLMs. To systematically study this, we introduce SVHalluc, the first comprehensive benchmark for evaluating speech-vision hallucination in audio-visual LLMs. Our benchmark diagnoses speech-vision hallucinations from two critical and complementary aspects: semantic and temporal. Experimental results demonstrate that state-of-the-art open-source audio-visual LLMs struggle with aligning speech content with corresponding visual signals, with a near-random accuracy on multiple tasks. In contrast, Gemini 2.5 Pro significantly outperforms the open-source models. Our analysis suggests that their failures stem from limited ability in cross-modality understanding, despite strong performance in single-modality perception. Our work uncovers a new and fundamental limitation of current audio-visual LLMs and highlights the need for speech-grounded video comprehension. Project page: https://chenshuang-zhang.github.io/projects/svhalluc/.
【8】Wavelet as Tokenizer: Preliminary Results on a Shared Wavelet Token Schema for Natural Signals
标题:作为令牌化器的子波:自然信号共享子波令牌模式的初步结果链接:https://arxiv.org/abs/2606.02631
备注:12 pages, 3 figures
摘要:本文研究音频,图像和视频是否可以共享一个共同的小波令牌模式,而不是依赖于单独的模态特定的潜在网格。它介绍了一个初步的连续令牌模型,该模型围绕一级Haar DWT/IDWT前端、共享系数令牌布局、可选的结构元数据、轻量级模态值适配器和共享令牌式编码器-解码器主干构建。在语音命令、EuroSAT RGB和DAVIS 2017数据上,密集共享模型可达到39.92 dB音频、29.37 dB图像和23.93 dB视频PSNR。连续潜在标量预算下的匹配率扫描表明,视觉增益并不能仅仅由潜在容量来解释,同时也表明添加剂元数据嵌入并不是一个普遍的改进来源。最后,固定速率的能量选择提供了一个强大的非参数基线:energy_global在压缩保持率下将音频的平均PSNR提高了16.73 dB,图像提高了16.90 dB,视频提高了15.86 dB。屏蔽稀疏训练达到34.45 dB的视频PSNR与50%的密集令牌。结果支持一个统一的小波令牌模式和稀疏令牌接口,而停止建立一个通用的离散词汇表。
摘要:This paper studies whether audio, images, and video can share a common wavelet token schema rather than relying on separate modality-specific latent grids. It introduces a preliminary continuous-token model built around a one-level Haar DWT/IDWT frontend, a shared coefficient-token layout, optional structural metadata, lightweight modality value adapters, and a shared token-wise encoder-decoder trunk. On Speech Commands, EuroSAT RGB, and DAVIS 2017 data, a dense shared model reaches 39.92 dB audio, 29.37 dB image, and 23.93 dB video PSNR. A matched-rate sweep under continuous latent scalar budgets indicates that the visual gains are not explained solely by latent capacity, while also showing that additive metadata embeddings are not a universal source of improvement. Finally, fixed-rate energy selection provides a strong non-parametric baseline: energy_global improves average PSNR over uniform selection by 16.73 dB for audio, 16.90 dB for images, and 15.86 dB for video under compressed keep ratios. Masked sparse training reaches 34.45 dB video PSNR with 50% of dense tokens. The results support a unified wavelet token schema and sparse token interface, while stopping short of establishing a universal discrete vocabulary.
【9】FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations
标题:FSA-GRPO:教听觉LLM使用Few-Shot演示链接:https://arxiv.org/abs/2606.02615
摘要:Few-Shot提示提供了一种有效的方法来使听觉大型语言模型适应低资源任务,例如儿童语音识别。然而,大多数听觉大语言模型没有被明确地训练来以这种演示条件格式执行推理,这限制了它们可以从Few-Shot提示中受益的程度。为了解决这一限制,我们引入了Few-Shot Aware GRPO(FSA-GRPO),这是一种基于RL的后训练配方,使用专门设计的奖励来鼓励模型利用Few-Shot演示,从而增强其Few-Shot适应能力。值得注意的是,仅使用高资源成人ASR数据进行训练可以提高模型的一般Few-Shot适应能力,不仅在儿童语音识别方面,而且在语音翻译和音频理解方面都有收获。我们进一步研究了数据选择和辅助奖励加权,以确定有效的训练配方。我们的实验表明,当域内数据不可用或不能用于训练时,FSA-GRPO比直接调整相关的域外数据更有效。
摘要:Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognition. However, most auditory large language models are not explicitly trained to perform inference in this demonstration-conditioned format, limiting the extent to which they can benefit from few-shot prompting. To address this limitation, we introduce Few-Shot Aware GRPO (FSA-GRPO), an RL-based post-training recipe that uses a specially designed reward to encourage the model to leverage few-shot demonstrations, thereby strengthening its few-shot adaptation ability. Notably, training with only high-resource adult ASR data improves the model's general few-shot adaptation ability, yielding gains not only in children's speech recognition but also in speech translation and audio understanding. We further study data selection and auxiliary reward weighting to identify an effective training recipe. Our experiments show that when in-domain data are unavailable or cannot be used for training, FSA-GRPO is more effective than direct tuning on related out-of-domain data.
【10】Efficient ASR Training with Conversations that Never Happened
标题:通过从未发生过的对话进行高效的ASB训练链接:https://arxiv.org/abs/2606.03957
摘要:低资源语言和利基领域的会话ASR受到缺乏领域匹配的多说话人训练数据的限制。我们提出了一个增强管道,生成与参与者的元数据,扬声器属性映射到TTS语音配置文件,合成的话语到扬声器感知的模拟对话。我们评估了五个LLM家庭在单一发电机,固定预算的混合,并使用相同的FastConformer-Large训练配方的每个设置。我们对匈牙利BEA-Dialogue基准语料库进行了全面的评估,该方法本身适用于任何语言,只要每个组件都有资源。结果表明,合成对话一致提高语音识别性能,但发电机的选择和数据组成强烈影响收益。我们最大的训练配置仅使用67小时的真实对话和636小时的模拟数据,在评估基准上取得了比在2700小时匈牙利语语音上训练的zero-shot模型更好的性能。这些研究结果表明,LLM生成的会话数据合成的TTS是一个实际的补充,真正的会话语料库的语音模型训练。
摘要:Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations. We evaluated five LLM families under single-generator, fixed-budget mixture, and scale-up settings using the same FastConformer-Large training recipe for each one. We ran comprehensive evaluations on the Hungarian BEA-Dialogue benchmark corpus, with the method itself being applicable to any language given the resources for each component. The results show that synthetic conversations consistently improve speech recognition performance, but generator choice and data composition strongly affect the gains. Our largest training configuration, using only 67 hours of real conversations and 636 hours of simulated data, achieves better performance on the evaluation benchmark than a zero-shot model trained on 2700 hours of Hungarian speech. These findings indicate that LLM-generated conversational data synthesized with TTS is a practical complement to real conversational corpora for speech model training.
【11】LiveBand: Live Accompaniment Generation in the Audio Domain
标题:LiveBand:音频领域的现场伴奏一代链接:https://arxiv.org/abs/2606.03803
摘要:我们提出了LiveBand,一个实时系统,生成高保真的音乐伴奏,以现场音频输入,尊重严格的因果约束。我们的方法在预先训练的因果音频自动编码器的连续潜在空间中训练因果Transformer生成器,使用来自编码器的对抗序列级监督。在每个时步,生成器仅接收因果可用的混合上下文和高斯噪声,并预测伴奏潜伏期,而不访问未来的混合帧或地面实况目标潜伏期。训练在因果掩蔽下以单个并行向前传递进行,而流推理则以滚动注意状态进行自回归。该模型的训练和推理计算通过设计进行匹配,消除了教师强迫和相关的暴露偏差。在多乐器音乐伴奏基准测试中,LiveBand在音频质量、节拍对齐和混音一致性的客观测量方面优于之前的工作,同时在消费者硬件上实现实时流媒体生成,而无需展望未来。
摘要:We present LiveBand, a real-time system that generates high-fidelity music accompaniments to live audio input, respecting strict causal constraints. Our method trains a causal transformer generator in the continuous latent space of a pre-trained causal audio autoencoder, using adversarial sequence-level supervision from a discriminator. At each timestep, the generator receives only the causally available mix context and Gaussian noise, and predicts accompaniment latents without access to future mix frames or ground-truth target latents. Training is performed in a single parallel forward pass under causal masking, while streaming inference proceeds autoregressively with a rolling attention state. The model's training and inference computations are matched by design, eliminating teacher forcing and the associated exposure bias. On a multi-instrument music accompaniment benchmark, LiveBand improves over prior work on objective measures of audio quality, beat alignment, and mix adherence, while enabling real-time streaming generation without lookahead into the future on consumer hardware.
【12】Benchmarking Speech-to-Speech Translation Models
标题:语音翻译模型基准链接:https://arxiv.org/abs/2606.03241
备注:Paper under submission
摘要:语音到语音翻译(S2 ST)发展迅速,但离线评估缺乏统一的协议:研究报告了不重叠的指标子集,无法进行直接比较。我们介绍了COMPASS,一个统一的和可重复的基准测试框架,集成了八个维度的46个指标,并将其部署在来自FLEURS和CVSS的1,248个模型语言配置上,跨越了十种语言对的级联和端到端架构。架构表现出互补的优势:最好与最差的差距超过30%的自然和扬声器保存,但保持在几个点的翻译质量,所以单一的指标排名系统性地歪曲系统质量。相关滤波将46个度量减少到每个方向10个,其中三个轴需要跨X$\到$EN和EN$\到$X的不同度量(例如,TER/UTMOS vs. ChrF++/NISQA-MOS);这些子集保持排名(斯皮尔曼的$ρ>0.80$),同时减少评估时间约2.5倍。配音,播客和医疗领域的人类验证显示,独立的MOS预测器无法预测听众的偏好,而顶级领域特定的指标与人类判断相关($ρ\geq 0.90$)。我们发布COMPASS作为域感知S2 ST评估的基础。
摘要:Speech-to-speech translation (S2ST) has advanced rapidly, but offline evaluation lacks a unified protocol: studies report non-overlapping metric subsets, preventing direct comparisons. We introduce COMPASS, a unified and reproducible benchmarking framework integrating 46 metrics across eight dimensions, and deploy it on 1,248 model-language configurations from FLEURS and CVSS, spanning cascaded and end-to-end architectures over ten language pairs. Architectures exhibit complementary strengths: best-vs-worst gaps exceed 30\% on naturalness and speaker preservation but remain within a few points on translation quality, so single-metric rankings systematically misrepresent system quality. Correlation filtering reduces 46 metrics to 10 per direction, with three axes requiring different metrics across X$\to$EN and EN$\to$X (e.g., TER/UTMOS vs. ChrF++/NISQA-MOS); these subsets preserve rankings (Spearman's $ρ>0.80$) while cutting evaluation time by $\approx 2.5\times$. Human validation across dubbing, podcasts, and medical domains shows standalone MOS predictors fail to predict listener preference, while top domain-specific metrics correlate with human judgment ($ρ\geq 0.90$). We release COMPASS as a foundation for domain-aware S2ST evaluation.
【13】Inference-Time Scaling for Joint Audio-Video Generation
标题:用于联合音频-视频生成的推理时间缩放链接:https://arxiv.org/abs/2606.03183
备注:Accepted by Transactions on Machine Learning Research (TMLR). Project page: https://jung-jaemin.github.io/ITS-AVGen-Proj/
摘要:联合音视频生成的目标是合成逼真的音视频对,既符合文本提示的语义和精确同步。虽然现有的联合音视频生成模型通常需要大量的训练资源来提高保真度,但推理时间缩放(ITS)最近已成为单模态领域中一种有前途的无训练替代方案。然而,扩展ITS从一个单一的模态到多模态域是不平凡的,因为它需要平衡多个异构的目标。在本文中,我们提出了第一个全面的研究ITS联合音视频生成。我们首先证明,多验证框架是必不可少的,以解决单目标指导的局限性,包括不对称的性能权衡和验证黑客。通过系统的分析,我们确定了一个最佳的多验证器组合,在所有质量维度上产生平衡的改进。最后,为了有效地聚合不同的奖励信号,我们提出了自适应奖励加权(ARW),一种新的测试时间优化算法。ARW将奖励聚合视为在线优化问题,利用可学习的参数来校准奖励方差,而不需要奖励分布的先验知识,从而确保鲁棒的多目标选择。VGGSound和JavisBench-mini基准测试的实验结果表明,我们的框架显着提高了语义对齐,感知质量和生成的输出的视听同步。在项目页面https://jung-jaemin.github.io/ITS-AVGen-Proj上可以找到合成的示例和代码。
摘要:Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchronized. While existing joint audio-video generation models often require substantial training resources to improve fidelity, Inference-Time Scaling (ITS) has recently emerged as a promising training-free alternative in single-modality domains. However, extending ITS from a single modality to multimodal domains is non-trivial, as it requires balancing multiple heterogeneous objectives. In this paper, we present the first comprehensive study of ITS for joint audio-video generation. We first demonstrate that a multi-verifier framework is essential to address the limitations of single-objective guidance, including asymmetric performance trade-offs and verifier hacking. Through systematic analysis, we then identify an optimal multi-verifier combination that yields balanced improvements across all quality dimensions. Finally, to effectively aggregate diverse reward signals, we propose Adaptive Reward Weighting (ARW), a novel test-time optimization algorithm. ARW treats reward aggregation as an online optimization problem, utilizing learnable parameters to calibrate reward variances without requiring prior knowledge of reward distributions, thereby ensuring robust multi-objective selection. Experimental results on VGGSound and JavisBench-mini benchmarks demonstrate that our framework significantly enhances semantic alignment, perceptual quality, and audio-visual synchronization of generated outputs. Synthesized samples and code are available on the project page: https://jung-jaemin.github.io/ITS-AVGen-Proj.
【14】CoughSense: Five-Class Respiratory Disease Classification via Whisper Encoder Fine-Tuning and Dual-Encoder Cross-Attention Fusion with Balanced Contrastive Learning
标题:CoughSense:通过Whisper编码器微调和双编码器交叉注意融合以及平衡对比学习进行五级呼吸道疾病分类链接:https://arxiv.org/abs/2606.02998
备注:26 pages, 3 figures
摘要:自动咳嗽分析为低成本呼吸道筛查提供了一条途径,但大多数现有工作都停留在二元COVID-19检测上。一个实用的工具需要从消费者智能手机上的一个咳嗽记录中区分出几种呼吸状况。我们提出CoughSense,一个将咳嗽记录分为五类的系统。这些是健康的,COVID-19,哮喘或呼吸道疾病,支气管炎和肺炎。我们汇总了来自四个公共数据集(Coswara,CoughVID,Virufy和华西医院儿科咳嗽数据集)的18,301个记录,并使用OpenAI Whisper编码器作为咳嗽疾病分类的预训练骨干。主要贡献是活动帧QKV注意力池,它将注意力限制在1500个编码器令牌中的前200个。这避免了沉默稀释问题,因为3秒的咳嗽只填充了Whisper 30秒输入窗口的150个标记。其他训练部分处理19比1的类不平衡和四个数据集的域转移。这些包括WeightedRandomSampler,SpecAugment,带有强制少数配对的平衡混合,监督对比辅助损失,Film症状调节和梯度反转域适应。双编码器模型通过交叉注意将Whisper与OPERA-CT呼吸基础模型融合。CoughSense(Whisper-tiny,8.6M参数)在五重交叉验证中达到82.3%的平衡准确度(macro-F1为0.817,AUC为0.941)。它以11.1分的优势击败了ImageNet预训练的EfficientNet-B2,以29.6分的优势击败了从头开始训练的ViT。所有五个班都通过了74%的回忆,五个班中有四个通过了80%。双编码器模型达到了85.4%的平衡精度。活动帧合并是所有消融组件中最大的单一贡献者,为5.1分,这应该有助于使用Whisper作为主干的任何短音频任务。
摘要:Automated cough analysis offers a path to low-cost respiratory screening, but most existing work stops at binary COVID-19 detection. A practical tool needs to tell apart several respiratory conditions from one cough recording on a consumer smartphone. We present CoughSense, a system that sorts cough recordings into five classes. These are healthy, COVID-19, asthma or respiratory condition, bronchitis, and pneumonia. We aggregated 18,301 recordings from four public datasets (Coswara, CoughVID, Virufy, and the West China Hospital Pediatric Cough Dataset) and used the OpenAI Whisper encoder as a pretrained backbone for cough disease classification. The main contribution is active-frame QKV attention pooling, which restricts attention to the first 200 of 1500 encoder tokens. This avoids the silence-dilution problem that arises because a 3-second cough fills only 150 tokens of Whisper's 30-second input window. Other training parts handle the 19 to 1 class imbalance and the four-dataset domain shift. These include WeightedRandomSampler, SpecAugment, Balanced Mixup with forced minority pairing, a supervised contrastive auxiliary loss, FiLM symptom conditioning, and gradient-reversal domain adaptation. A dual-encoder model fuses Whisper with the OPERA-CT respiratory foundation model through cross-attention. CoughSense (Whisper-tiny, 8.6M parameters) reached 82.3 percent balanced accuracy on five-fold cross-validation (macro-F1 of 0.817, AUC of 0.941). It beat an ImageNet-pretrained EfficientNet-B2 by 11.1 points and a ViT trained from scratch by 29.6 points. All five classes passed 74 percent recall and four of five passed 80 percent. The dual-encoder model reached 85.4 percent balanced accuracy. Active-frame pooling is the largest single contributor across all ablation components at 5.1 points, which should help any short-audio task using Whisper as a backbone.
【15】EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement
标题:EntangleCodec:通过语义-声学纠缠的统一离散音频令牌器链接:https://arxiv.org/abs/2606.02739
备注:17 pages, 10 figures
摘要:音频分词器作为连续音频和音频语言模型(ALM)之间的离散接口,但现有的分词器往往难以支持理解和生成。面向重构的编解码器保持声学保真度,但缺乏丰富的语义,而语义感知的标记器通常依赖于单独的语义和声学流,从而引入冗余或不对齐。 我们提出了\textbf{EntangleCodec},一个统一的离散音频标记器,在量化之前学习标题对齐的语义声学表示。通过将音频与丰富的字幕而不是ASR转录对齐,EntangleCodec在紧凑的令牌流中捕获语言内容,说话者身份,情感,韵律和声学场景。流匹配扩散解码器进一步实现跨语音、音乐和一般音频的高质量重构。 EntangleCodec实现了与专业编解码器竞争的重建质量,在MMAR上的音频理解方面优于所有基于编解码器的基线,最高可达\textbf{+7.4\%},并在统一的框架中支持TTS和TTA生成。此外,基于EntangleCodec的音频语言模型表现出强大的缩放行为:即使在\textit{0.6B}参数下,该模型在三个基准测试中使用更少的参数,超过了具有超过\textit{13 B}参数的专用连续表示LLM;缩放到\textit{8B}进一步建立了关于MMAR的新的最先进结果,强调了表示质量与音频语言建模中的模型比例一样重要。代码和型号重量可在https://github.com/luckyerr/EntangleCodec上获得。
摘要:Audio tokenizers serve as the discrete interface between continuous audio and Audio Language Models (ALMs), but existing tokenizers often struggle to support both understanding and generation. Reconstruction-oriented codecs preserve acoustic fidelity but lack rich semantics, while semantic-aware tokenizers typically rely on separate semantic and acoustic streams, introducing redundancy or misalignment. We propose \textbf{EntangleCodec}, a unified discrete audio tokenizer that learns caption-aligned semantic-acoustic representations before quantization. By aligning audio with rich captions rather than ASR transcripts, EntangleCodec captures linguistic content, speaker identity, emotion, prosody, and acoustic scenes within a compact token stream. A flow-matching diffusion decoder further enables high-quality reconstruction across speech, music, and general audio. EntangleCodec achieves reconstruction quality competitive with specialized codecs, outperforms all codec-based baselines on audio understanding by up to \textbf{+7.4\%} on MMAR, and supports both TTS and TTA generation in a unified framework. Furthermore, EntangleCodec-based audio language models demonstrate strong scaling behavior: even at \textit{0.6B} parameters, the model surpasses specialized continuous-representation LLMs with over \textit{13B} parameters across three benchmarks using \textbf{22$\times$} fewer parameters; scaling to \textit{8B} further establishes new state-of-the-art results on MMAR, highlighting that representation quality is as critical as model scale in audio language modeling. Code and model weights are available at https://github.com/luckyerr/EntangleCodec.
【16】Before Fusion, Ask What to Keep: Contextual Calibration of Multimodal Signals
标题:Fusion之前,询问要保留什么:多模式信号的上下文校准链接:https://arxiv.org/abs/2606.02679
备注:11 pages, 7 figures, 9 tables
摘要:多模态系统通常受益于跨语言,声音和视觉流的信息组合,但这种好处并不能保证。对一种输入有用的模态可能会分散另一种输入的注意力,并且同一模态内的局部特征响应可能与来自其他来源的证据不一致。这项工作研究了如何调整多模态表示之前,他们合并的下游预测。我们开发了一个紧凑的校准模块,比较每个模态与其他的汇总水平,提取线索的跨源支持和冲突,并将这些线索转换成实例和维度的调制信号。校准应用于原始模态特征,而不是已经融合的表示,使模型能够抑制误导性成分,保留弱但有用的证据,并强调当前多模态上下文更好地支持的响应。该模块被设计为一个插件组件,可以连接到不同的融合骨干,而无需改变他们的预测头。在涵盖情感理解、动作识别、视听事件检测和视听情感分类的五个基准测试中,所提出的预组合校准策略提高了基于序列和卷积融合设置下的性能。在模态移除、合成腐败、训练动态和特征级可视化下的其他分析表明,在融合之前校准信号可以减少不可靠模态的干扰,并产生更稳定的多模态优化。
摘要:Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed. A modality that is useful for one input may become distracting for another, and local feature responses within the same modality can disagree with evidence from other sources. This work investigates how to adjust multimodal representations before they are merged by a downstream predictor. We develop a compact calibration module that compares each modality with the others at the summary level, extracts cues of cross-source support and conflict, and converts these cues into instance-wise and dimension-wise modulation signals. The calibration is applied to the original modality features rather than to already fused representations, enabling the model to suppress misleading components, preserve weak but useful evidence, and emphasize responses that are better supported by the current multimodal context. The module is designed as a plug-in component and can be attached to different fusion backbones without changing their prediction heads. Across five benchmarks covering sentiment understanding, action recognition, audio-visual event detection, and audio-visual emotion classification, the proposed pre-combination calibration strategy improves performance under both sequence-based and convolutional fusion settings. Additional analyses under modality removal, synthetic corruption, training dynamics, and feature-level visualization show that calibrating signals before fusion can reduce interference from unreliable modalities and produce more stable multimodal optimization.
【17】SegTune: Structured and Fine-Grained Control for Song Generation
标题:SegButton:歌曲生成的结构化和细粒度控制链接:https://arxiv.org/abs/2606.02638
备注:This paper has been accepted to ACL 2026 as an oral presentation and has been nominated for the Best Paper Award. This work is a revised and extended version of an earlier technical report (arXiv:2510.18416). arXiv admin note: text overlap with arXiv:2510.18416
摘要:神经歌曲生成的最新进展使歌词和全局文本提示的高质量合成成为可能。然而,大多数系统无法对歌曲的时间变化属性进行建模,严重限制了对音乐结构和动态的细粒度控制。为了解决这个问题,我们提出了SegTune,一个基于扩散变换器的框架,通过允许用户或大型语言模型(LLM)指定与歌曲片段对齐的本地音乐描述,实现结构化和细粒度的可控性。这些片段提示在时间上广播到对应的时间窗口,而全局提示确保风格一致性。为了支持精确的歌词到音乐的对齐,我们引入了一个基于LLM的持续时间预测器,该预测器以LyRiCs格式自回归生成音乐级时间戳。我们进一步构建了一个大规模的数据管道,用于高质量的歌曲收集与对齐的歌词和提示,并提出了新的指标来评估段对齐和声乐一致性。实验表明,SegTune优于现有的基线在音乐性和可控性。请访问我们的项目页面(https://github.com/KlingAIResearch/SegTune)获取代码和更多生成的歌曲。
摘要:Recent advances in neural song generation have enabled high-quality synthesis from lyrics and global textual prompts. However, most systems fail to model temporally varying attributes of songs, severely limiting fine-grained control over musical structure and dynamics. To address this, we propose SegTune, a Diffusion Transformer-based framework enabling structured and fine-grained controllability by allowing users or large language models (LLMs) to specify local musical descriptions aligned to song segments. These segment prompts are temporally broadcast to corresponding time windows, while global prompts ensure stylistic coherence. To support precise lyric-to-music alignment, we introduce an LLM-based duration predictor that autoregressively generates sentence-level timestamps in LyRiCs format. We further construct a large-scale data pipeline for high-quality song collection with aligned lyrics and prompts, and propose new metrics to evaluate segment alignment and vocal consistency. Experiments demonstrate that SegTune outperforms existing baselines in both musicality and controllability. Visit our project page (https://github.com/KlingAIResearch/SegTune) for codes and more generated songs.
机器翻译由腾讯交互翻译提供,仅供参考
