微信公众号:arXiv_Daily
cs.SD语音
【1】Joint Fullband-Subband Modeling for High-Resolution SingFake Detection
标题:用于高分辨率SingFake检测的全带-子带联合建模
链接:https://arxiv.org/abs/2604.04841
备注:Submitted to INTERSPEECH 2026
摘要:歌声合成的快速发展增加了未经授权的模仿风险,迫切需要更好的歌声Deepfake(SingFake)检测,也称为SVDD。与语音不同,歌唱包含复杂的音高,宽动态范围和音色变化。传统的16 kHz采样探测器被证明是不够的,因为它们丢弃了重要的高频信息。这项研究提出了第一个系统的分析高分辨率(44.1 kHz采样率)音频SVDD。我们提出了一个联合全波段子带建模框架:全波段捕捉全球范围内,而特定的子带专家隔离细粒度的合成文物不均匀分布在整个频谱。WildSVDD数据集上的实验表明,高频子带提供了必要的补充线索。我们的框架显著优于16 kHz采样模型,证明了高分辨率音频和战略子带集成对于稳健的野外检测至关重要。
摘要:Rapid advances in singing voice synthesis have increased unauthorized imitation risks, creating an urgent need for better Singing Voice Deepfake (SingFake) Detection, also known as SVDD. Unlike speech, singing contains complex pitch, wide dynamic range, and timbral variations. Conventional 16 kHz-sampled detectors prove inadequate, as they discard vital high-frequency information. This study presents the first systematic analysis of high-resolution (44.1 kHz sampling rate) audio for SVDD. We propose a joint fullband-subband modeling framework: the fullband captures global context, while subband-specific experts isolate fine-grained synthesis artifacts unevenly distributed across the spectrum. Experiments on the WildSVDD dataset demonstrate that high-frequency subbands provide essential complementary cues. Our framework significantly outperforms 16 kHz-sampled models, proving that high-resolution audio and strategic subband integration are critical for robust in-the-wild detection.
【2】OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
标题:OmniSonic:从视频和文本实现普遍和整体的音频生成
链接:https://arxiv.org/abs/2604.04348
备注:CVPR 2026
摘要:在本文中,我们提出了通用整体音频生成(UniHAGen),这是一项用于合成综合听觉场景的任务,包括跨不同领域的屏幕上和屏幕外声音(例如,环境事件、乐器和人类语言)。现有的视频调节音频生成模型通常集中于产生对应于可见发声事件的屏幕上环境声音,而忽略屏幕外听觉事件。虽然最近的整体联合文本视频到音频生成模型旨在产生具有屏幕上和屏幕外声音的听觉场景,但它们仅限于非语音声音,缺乏生成或整合人类语音的能力。为了克服这些限制,我们引入了OmniSonic,一个基于流匹配的扩散框架,它以视频和文本为条件。它采用TriAttn-DiT架构,执行三个交叉注意操作,同时处理屏幕上的环境声音,屏幕外的环境声音和语音条件,并采用混合专家(MoE)门控机制,在生成过程中自适应地平衡它们的贡献。此外,我们构建了UniHAGen-Bench,一个新的基准测试,超过一千个样本,涵盖三个代表性的屏幕上/屏幕外的语音环境场景。大量实验表明,OmniSonic在客观指标和人工评估方面始终优于最先进的方法,为通用和整体音频生成建立了强大的基线。项目页面:https://weiguopian.github.io/OmniSonic_webpage/
摘要:In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on-screen and off-screen sounds across diverse domains (e.g., ambient events, musical instruments, and human speech). Prior video-conditioned audio generation models typically focus on producing on-screen environmental sounds that correspond to visible sounding events, neglecting off-screen auditory events. While recent holistic joint text-video-to-audio generation models aim to produce auditory scenes with both on- and off-screen sound but they are limited to non-speech sounds, lacking the ability to generate or integrate human speech. To overcome these limitations, we introduce OmniSonic, a flow-matching-based diffusion framework jointly conditioned on video and text. It features a TriAttn-DiT architecture that performs three cross-attention operations to process on-screen environmental sound, off-screen environmental sound, and speech conditions simultaneously, with a Mixture-of-Experts (MoE) gating mechanism that adaptively balances their contributions during generation. Furthermore, we construct UniHAGen-Bench, a new benchmark with over one thousand samples covering three representative on/off-screen speech-environment scenarios. Extensive experiments show that OmniSonic consistently outperforms state-of-the-art approaches on both objective metrics and human evaluations, establishing a strong baseline for universal and holistic audio generation. Project page: https://weiguopian.github.io/OmniSonic_webpage/
【3】Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning
标题:用于无监督视听表示学习的分层语义相关性感知掩蔽自动编码器
链接:https://arxiv.org/abs/2604.04229
备注:6 pages, 2 tables, 4 figures. Accepted by IEEE ICME 2026
摘要:从弱配对、无标签语料库中学习对齐的多模态嵌入是一项挑战:管道通常只提供预提取的特征,剪辑包含多个事件,以及虚假的同现。我们提出HSC-MAE(Hierarchical Semantic Correlation-Aware Masked Autoencoder),一种双路径教师-学生框架,其在三个互补的表示级别上强制语义一致性-从粗到细:(i)经由DCCA的全局级别规范几何相关,其在共享模态不变子空间内对齐音频和视觉嵌入;(ii)通过教师挖掘的软top-k亲和度进行局部邻域-语义相关,保持语义相似实例之间的多正关系结构;以及(iii)通过掩码自编码的样本级条件充分性相关,其确保个体嵌入在部分观察下保留有区别的语义内容。具体地说,学生MAE路径是用掩码特征重建和仿射加权软top-k InfoNCE训练的; EMA教师通过CCA路径对未掩码输入进行操作,提供稳定的规范几何和软阳性。可学习的多任务权重协调竞争目标,并且可选的蒸馏损失将教师的几何形状转移到学生身上。在AVE和VEGAS上的实验表明,与强无监督基线相比,mAP有了实质性的改进,验证了HSC-MAE产生了鲁棒且结构良好的视听表示。
摘要:Learning aligned multimodal embeddings from weakly paired, label-free corpora is challenging: pipelines often provide only pre-extracted features, clips contain multiple events, and spurious co-occurrences. We propose HSC-MAE (Hierarchical Semantic Correlation-Aware Masked Autoencoder), a dual-path teacher-student framework that enforces semantic consistency across three complementary levels of representation - from coarse to fine: (i) global-level canonical-geometry correlation via DCCA, which aligns audio and visual embeddings within a shared modality-invariant subspace; (ii) local-level neighborhood-semantics correlation via teacher-mined soft top-k affinities, which preserves multi-positive relational structure among semantically similar instances; and (iii) sample-level conditional-sufficiency correlation via masked autoencoding, which ensures individual embeddings retain discriminative semantic content under partial observation. Concretely, a student MAE path is trained with masked feature reconstruction and affinity-weighted soft top-k InfoNCE; an EMA teacher operating on unmasked inputs via the CCA path supplies stable canonical geometry and soft positives. Learnable multi-task weights reconcile competing objectives, and an optional distillation loss transfers teacher geometry into the student. Experiments on AVE and VEGAS demonstrate substantial mAP improvements over strong unsupervised baselines, validating that HSC-MAE yields robust and well-structured audio-visual representations.
【4】Measuring Robustness of Speech Recognition from MEG Signals Under Distribution Shift
标题:分布漂移下MEG信号语音识别的鲁棒性测量
链接:https://arxiv.org/abs/2604.04129
备注:17 pages, 6 figures, LibriBrain Competition @NeurIPS2025
摘要:本研究使用2025年PNPL竞赛的LibriBrain音素分类基准,研究了非侵入性MEG信号的鲁棒语音相关解码。我们比较了残差卷积神经网络(CNN),基于STFT的CNN和CNN-Transformer混合,同时还检查了组平均,标签平衡,重复分组,归一化策略和数据增强的影响。在我们的内部实现中,预处理和数据配置的选择比额外的架构复杂性更重要,其中实例规范化是最有影响力的泛化修改。我们自己的模型中最强的一个,具有组平均、标签平衡、重复分组和实例归一化的CNN,在测试分割上实现了60.95%的F1-macro,而普通CNN基线的F1-macro为39.53%。然而,我们的大多数模型,没有实例规范化,显示出大量的验证测试退化,表明不同的规范化统计引起的分布偏移是我们实验中泛化的主要障碍。相比之下,MEGConformer在验证和测试中都保持了64.09%的F1-macro,显着性图分析与这种对比在定性上是一致的:较弱的模型在分裂中表现出更集中或重复的音素敏感模式,而MEGConformer似乎更分散。总的来说,结果表明,提高非侵入性音素解码的可靠性可能需要更好地处理归一化相关的分布偏移,同时也解决了单次尝试解码的挑战。
摘要:This study investigates robust speech-related decoding from non-invasive MEG signals using the LibriBrain phoneme-classification benchmark from the 2025 PNPL competition. We compare residual convolutional neural networks (CNNs), an STFT-based CNN, and a CNN--Transformer hybrid, while also examining the effects of group averaging, label balancing, repeated grouping, normalization strategies, and data augmentation. Across our in-house implementations, preprocessing and data-configuration choices matter more than additional architectural complexity, among which instance normalization emerges as the most influential modification for generalization. The strongest of our own models, a CNN with group averaging, label balancing, repeated grouping, and instance normalization, achieves 60.95% F1-macro on the test split, compared with 39.53% for the plain CNN baseline. However, most of our models, without instance normalization, show substantial validation-to-test degradation, indicating that distribution shift induced by different normalization statistics is a major obstacle to generalization in our experiments. By contrast, MEGConformer maintains 64.09% F1-macro on both validation and test, and saliency-map analysis is qualitatively consistent with this contrast: weaker models exhibit more concentrated or repetitive phoneme-sensitive patterns across splits, whereas MEGConformer appears more distributed. Overall, the results suggest that improving the reliability of non-invasive phoneme decoding will likely require better handling of normalization-related distribution shift while also addressing the challenge of single-trial decoding.
【5】A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
标题:视听推理跨模式印刷攻击的系统研究
链接:https://arxiv.org/abs/2604.03995
摘要:随着视听多模态大型语言模型(MLLM)越来越多地部署在安全关键型应用程序中,了解其漏洞至关重要。为此,我们介绍了多模态排版,一个系统的研究,探讨如何在多个模态的排版攻击产生不利影响MLLM。虽然以前的工作主要集中在单峰攻击,我们暴露了跨模态的脆弱性MLLM。我们分析了音频、视频和文本扰动之间的相互作用,并揭示了协同多模态攻击比单模态攻击产生的威胁更大(攻击成功率= 83.43\%$ vs 34.93\%$)。我们在多个前沿MLLM,任务,并且常识推理和内容调节基准将多模态排版建立为多模态推理中的关键且未充分探索的攻击策略。代码和数据将公开。
摘要:As audio-visual multi-modal large language models (MLLMs) are increasingly deployed in safety-critical applications, understanding their vulnerabilities is crucial. To this end, we introduce Multi-Modal Typography, a systematic study examining how typographic attacks across multiple modalities adversely influence MLLMs. While prior work focuses narrowly on unimodal attacks, we expose the cross-modal fragility of MLLMs. We analyze the interactions between audio, visual, and text perturbations and reveal that coordinated multi-modal attack creates a significantly more potent threat than single-modality attacks (attack success rate = $83.43\%$ vs $34.93\%$).Our findings across multiple frontier MLLMs, tasks, and common-sense reasoning and content moderation benchmarks establishes multi-modal typography as a critical and underexplored attack strategy in multi-modal reasoning. Code and data will be publicly available.
【6】FlueBricks: A Construction Kit of Flute-like Instruments for Acoustic Reasoning
标题:节日砖:用于声学推理的长笛状仪器的构建套件
链接:https://arxiv.org/abs/2604.03636
备注:Accepted to CHI 2026
摘要:我们提出了FestivalBricks,一个通过构建和定制长笛类乐器进行声学推理的构建工具包。通过组装体现各种航空声学特性的发生器、谐振器和连接器模块,用户可以通过动手实验更深入地了解气孔、管长度和音孔位置如何改变起始、音高和音色。这就形成了一个设计者-演奏者的循环,通过配置和演奏来形成、测试和完善声学行为-声学推理-将声学乐器从静态工件转变为动态系统。为了了解用户如何参与这个系统,我们进行了一项探索性研究,有12名参与者,从新手到专业音乐家。在他们的探索过程中,我们观察到参与者流利地在设计师和演奏者角色之间切换,从熟悉的乐器中构建设计,形成和完善他们对长度,音孔和发生器几何形状的声学理解,重新解释超出其预期功能的模块,并将他们的创作用于表演行为,如教学展示和音乐表达。这些共同展示了FestivalBricks作为具体声学推理的教学工具的潜力。
摘要:We present FlueBricks, a construction kit for acoustic reasoning via building and customizing flute-like instruments. By assembling generator, resonator, and connector modules that embody various aeroacoustic properties, users gain deeper understanding of how blowhole, tube length, and tone-hole placement alter onset, pitch, and timbre through hands-on experimentation. This forms a designer-player loop of configuring and playing to form, test, and refine acoustic behaviors-acoustic reasoning-shifting acoustic instruments from static artifacts to dynamic systems. To understand how users engage with this system, we conducted an exploratory study with 12 participants ranging from novices to professional musicians. During their explorations, we observed participants fluently switching between designer and player roles, scaffolding designs from familiar instruments, forming and refining their acoustic understanding of length, tone holes, and generator geometry, reinterpreting modules beyond their intended functions, and using their creations for performative acts such as pedagogical showing and musical expression. These collectively demonstrated FlueBricks's potential as a pedagogical tool for embodied acoustic reasoning.
【7】Composer Vector: Style-steering Symbolic Music Generation in a Latent Space
标题:作曲家Vector:潜在空间中的风格引导象征音乐生成
链接:https://arxiv.org/abs/2604.03333
摘要:符号音乐生成已经取得了重大进展,但实现对作曲家风格的细粒度和灵活控制仍然具有挑战性。现有的基于训练的作曲家风格调节方法依赖于大型标记数据集。此外,这些方法通常一次只支持单个作曲家的生成,限制了它们对更具创造性或混合场景的适用性。在这项工作中,我们提出了作曲家向量,一个推理时间转向方法,直接在模型的潜在空间控制作曲家的风格,而无需重新训练。通过对多个符号音乐生成模型的实验,我们表明Composer Vector可以有效地引导一代人走向目标作曲家风格,通过连续的转向系数实现平滑和可解释的控制。它还可以在统一的潜在空间框架内无缝融合多种风格。总的来说,我们的工作表明,简单的潜在空间转向提供了一个实用和通用的机制,可控的符号音乐生成,使更灵活和互动的创作工作流程。代码和演示可在此处获得:https://github.com/JiangXunyi/Composer-Vector和https://jiangxunyi.github.io/composervector.github.io/
摘要:Symbolic music generation has made significant progress, yet achieving fine-grained and flexible control over composer style remains challenging. Existing training-based methods for composer style conditioning depend on large labeled datasets. Besides, these methods typically support only single-composer generation at a time, limiting their applicability to more creative or blended scenarios. In this work, we propose Composer Vector, an inference-time steering method that operates directly in the model's latent space to control composer style without retraining. Through experiments on multiple symbolic music generation models, we show that Composer Vector effectively guides generations toward target composer styles, enabling smooth and interpretable control through a continuous steering coefficient. It also enables seamless fusion of multiple styles within a unified latent space framework. Overall, our work demonstrates that simple latent space steering provides a practical and general mechanism for controllable symbolic music generation, enabling more flexible and interactive creative workflows. Code and Demo are available here: https://github.com/JiangXunyi/Composer-Vector and https://jiangxunyi.github.io/composervector.github.io/
【8】CoLoRSMamba: Conditional LoRA-Steered Mamba for Supervised Multimodal Violence Detection
标题:CoLoRSMamba:有条件LoRA引导的曼巴,用于监督多模式暴力检测
链接:https://arxiv.org/abs/2604.03329
摘要:暴力检测受益于音频,但现实世界的音景可能是嘈杂的或与可见场景弱相关的。我们提出了CoLoRSMamba,一个定向的视频到音频多模式架构,通过CLS引导的条件LoRA耦合VideoMamba和AudioMamba。在每一层,VideoMamba CLS令牌产生一个通道调制矢量和一个稳定门,用于调整AudioMamba投影,负责选择性状态空间参数(Delta,B,C),包括步长路径,产生场景感知的音频动态,而没有令牌级的交叉注意。训练将二进制分类与对称的AV-InfoNCE目标相结合,该目标将剪辑级音频和视频嵌入对齐。为了支持公平的多模态评估,我们从时间注释中策划NTU-CCTV和DVD数据集的音频过滤剪辑级别子集,仅保留具有可用音频的剪辑。在这些子集上,CoLoRSMAamba优于代表性的仅音频,仅视频和多模式基线,在NTU-CCTV上实现88.63%的准确性/86.24%F1-V,在DVD上实现75.77%的准确性/72.94%F1-V。它还提供了一个有利的准确性和效率的权衡,超过了几个更大的模型,更少的参数和FLOP。
摘要:Violence detection benefits from audio, but real-world soundscapes can be noisy or weakly related to the visible scene. We present CoLoRSMamba, a directional Video to Audio multimodal architecture that couples VideoMamba and AudioMamba through CLS-guided conditional LoRA. At each layer, the VideoMamba CLS token produces a channel-wise modulation vector and a stabilization gate that adapt the AudioMamba projections responsible for the selective state-space parameters (Delta, B, C), including the step-size pathway, yielding scene-aware audio dynamics without token-level cross-attention. Training combines binary classification with a symmetric AV-InfoNCE objective that aligns clip-level audio and video embeddings. To support fair multimodal evaluation, we curate audio-filtered clip level subsets of the NTU-CCTV and DVD datasets from temporal annotations, retaining only clips with available audio. On these subsets, CoLoRSMamba outperforms representative audio-only, video-only, and multimodal baselines, achieving 88.63% accuracy / 86.24% F1-V on NTU-CCTV and 75.77% accuracy / 72.94% F1-V on DVD. It further offers a favorable accuracy-efficiency tradeoff, surpassing several larger models with fewer parameters and FLOPs.
【9】AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
标题:AffectSpeech:具有细粒度文本描述的大规模情感语音数据集,用于语音情感字幕和合成
链接:https://arxiv.org/abs/2604.04160
备注:Submitted to IEEE Transactions
摘要:情感是口语交际中必不可少的一部分,但现有的语音情感建模框架大多依赖于预定义的类别或低维连续属性,表达能力有限。语音情感字幕和合成的最新进展表明,文本描述提供了一个更灵活和可解释的替代表示语音中的情感特征。然而,由于缺乏与可靠和细粒度的自然语言注释相匹配的情感语音数据集,这一方向的进展受到阻碍。为了解决这个问题,我们引入了AffectSpeech,这是一个大规模的人类记录语音语料库,其中包含了用于细粒度情感分析和生成的结构化描述。每个话语的特征在于六个互补的维度,包括情感极性,开放词汇情感标题,强度水平,韵律属性,突出段和语义内容,使语音表达的多粒度建模。为了平衡注释质量和可扩展性,我们采用了一个人-LLM协作注释管道,该管道集成了算法预标记,多LLM描述生成和人在环验证。此外,这些注释被重新表述为不同的描述风格,以提高语言的多样性和减少下游建模的风格偏差。语音情感字幕和合成的实验结果表明,在AffectSpeech上训练的模型在多个评估设置中始终获得优异的性能。
摘要:Emotion is essential in spoken communication, yet most existing frameworks in speech emotion modeling rely on predefined categories or low-dimensional continuous attributes, which offer limited expressive capacity. Recent advances in speech emotion captioning and synthesis have shown that textual descriptions provide a more flexible and interpretable alternative for representing affective characteristics in speech. However, progress in this direction is hindered by the lack of an emotional speech dataset aligned with reliable and fine-grained natural language annotations. To tackle this, we introduce AffectSpeech, a large-scale corpus of human-recorded speech enriched with structured descriptions for fine-grained emotion analysis and generation. Each utterance is characterized across six complementary dimensions, including sentiment polarity, open-vocabulary emotion captions, intensity level, prosodic attributes, prominent segments, and semantic content, enabling multi-granular modeling of vocal expression. To balance annotation quality and scalability, we adopt a human-LLM collaborative annotation pipeline that integrates algorithmic pre-labeling, multi-LLM description generation, and human-in-the-loop verification. Furthermore, these annotations are reformulated into diverse descriptive styles to enhance linguistic diversity and reduce stylistic bias in downstream modeling. Experimental results on speech emotion captioning and synthesis demonstrate that models trained on AffectSpeech consistently achieve superior performance across multiple evaluation settings.
【10】Neurological Plausibility of AI-Generated Music for Commercial Environments: An In-Silico Cortical Investigation Using Wubble and TRIBE v2
标题:人工智能生成的音乐在商业环境中的神经学合理性:使用Wubble和TRUTE v2的In-Silico皮质研究
链接:https://arxiv.org/abs/2604.04025
备注:IEEE-style preprint; 4 figures; 4 tables
摘要:背景音乐塑造了商业环境中的注意力、影响力和接近行为,但人工智能生成的音乐在这种环境中的神经可接受性仍然很差。我们提出了一个在silico试点研究,结合Wubble,生成音乐系统,与TRIBE V2,公开发布的全脑编码模型,估计皮质反应配置文件的无条件零售音乐。生成五个完全仪器化的轨道,以跨越低到高的唤醒,稀疏到密集的安排,和中性到积极的效价提示,然后分析与音频唯一的TRIBE v2推理的一致性标准化波形。分析集中在对听觉、上颞叶、颞顶叶和下额叶HCP包裹总结的fsaverage 5皮质预测。快速明亮的大流行条件产生了最大的全皮质平均激活(0.0402)、最强的前额叶感兴趣区复合反应(0.0704)以及IFJa(0.1102)、IFJp(0.0995)、A5(0.0188)和45区(0.0015)的最高包裹平均值。成对的空间相关性范围从0.787到0.974,表明提示变化调制预测的皮质状态,而不是产生一个单一的未分化的响应配置文件。预测皮层表面地图进一步揭示了视觉上不同的空间组织之间的低唤醒和高唤醒条件。这些结果支持了一个谨慎的说法,即皮层神经系统的可解释性:条件反射的人工智能音乐可以系统地改变与显著性和评价相关的预测皮层-颞叶-前额叶模式。虽然这项研究没有建立皮层下的奖励参与或消费者行为,但它为商业音乐生成的神经预筛选和预优化提供了一个可重复的框架,以对抗生物学上知情的皮层代理。
摘要:Background music shapes attention, affect, and approach behavior in commercial environments, yet the neural plausibility of AI-generated music for such settings remains poorly characterized. We present an in-silico pilot study that combines Wubble, a generative music system, with TRIBE v2, a publicly released whole-brain encoding model, to estimate cortical response profiles for prompt-conditioned retail music. Five fully instrumental tracks were generated to span low-to-high arousal, sparse-to-dense arrangement, and neutral-to-positive valence prompts, then analyzed with audio-only TRIBE v2 inference on loudness-normalized waveforms. Analysis focused on fsaverage5 cortical predictions summarized over auditory, superior temporal, temporo-parietal, and inferior frontal HCP parcels. The fast bright major-pop condition produced the largest whole-cortex mean activation (0.0402), the strongest prefrontal ROI composite response (0.0704), and the highest parcel means in IFJa (0.1102), IFJp (0.0995), A5 (0.0188), and area 45 (0.0015). Pairwise spatial correlations ranged from 0.787 to 0.974, indicating that prompt variation modulated predicted cortical states rather than yielding a single undifferentiated response profile. Predicted cortical surface maps further revealed visually distinct spatial organization between low-arousal and high-arousal conditions. These results support a cautious claim of cortical neurological plausibility: prompt-conditioned AI music can systematically shift predicted auditory-temporal-prefrontal patterns relevant to salience and valuation. Although the study does not establish subcortical reward engagement or consumer behavior, it provides a reproducible framework for neural pre-screening and pre-optimization of commercial music generation against biologically informed cortical proxies.
【11】Rewriting TTS Inference Economics: Lightning V2 on Tenstorrent Achieves 4x Lower Cost Than NVIDIA L40S
标题:重写TTC推理经济学:Tenstorrent上的Lightning V2成本比NVIDIA L40 S低4倍
链接:https://arxiv.org/abs/2604.03279
摘要:文本到语音(TTS)模型在数值上比大型语言模型(LLM)脆弱得多,这是由于它们的连续波形生成和对小数值扰动的感知敏感性。虽然诸如BlockFloat8(BFP8)和低保真度(LoFi)计算等积极的精度降低技术已被广泛用于语言模型中,但将类似的策略应用于TTS系统通常会导致可听伪影,相位不稳定和频谱失真。 在这项工作中,我们提出了Lightning V2,这是一种针对Tenstorrent硬件协同优化的生产级TTS模型。通过精确感知的架构设计和软硬件协同优化,我们实现了超过95%的LoFi计算保真度和超过80%的BlockFloat8部署,而没有可测量的音频质量下降。利用Tenstorrent的片上网络(NoC)、分布式SRAM和确定性执行模型,我们减少了内存移动和冗余权重提取,从而实现了高效的低精度推理。 与NVIDIA L40S基准相比,Lightning V2在同等吞吐量下实现了约4倍的本地加速器成本,同时保持了生产音频保真度。我们的研究结果表明,精确协同设计,结合硬件感知优化,可以从根本上重塑实时语音推理的经济性。
摘要:Text-to-Speech (TTS) models are significantly more numerically fragile than Large Language Models (LLMs) due to their continuous waveform generation and perceptual sensitivity to small numerical perturbations. While aggressive precision reduction techniques such as BlockFloat8 (BFP8) and low-fidelity (LoFi) compute have been widely adopted in language models, applying similar strategies to TTS systems often results in audible artifacts, phase instability, and spectral distortion. In this work, we present Lightning V2, a production-grade TTS model co-optimized for Tenstorrent hardware. Through precision-aware architectural design and hardware-software co-optimization, we achieve over 95% LoFi computational fidelity and more than 80% BlockFloat8 deployment without measurable degradation in audio quality. Leveraging Tenstorrent's Network-on-Chip (NoC), distributed SRAM, and deterministic execution model, we reduce memory movement and redundant weight fetches, enabling efficient low-precision inference. Compared to an NVIDIA L40S baseline, Lightning V2 achieves approximately 4x lower on-prem accelerator cost at equivalent throughput, while maintaining production audio fidelity. Our results demonstrate that precision co-design, combined with hardware-aware optimization, can fundamentally reshape the economics of real-time speech inference.
【1】Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
标题:Full-Duplex-Bench-v3:在现实世界不流利情况下用于Full-Duplex语音代理的基准工具
链接:https://arxiv.org/abs/2604.04847
备注:Work in progress. Demo at https://daniellin94144.github.io/FDB-v3-demo
摘要:我们介绍Full-Duplex-Bench-v3(FDB-v3),这是一个在自然语音条件和多步工具使用下评估口语模型的基准。与之前的工作不同,我们的数据集完全由五个不流利类别的真实人类音频注释组成,与需要跨四个任务域的链式API调用的场景配对。我们评估了六种模型配置-GPT实时,Gemini Live 2.5,Gemini Live 3.1,Grok,Ultravox v0.7和传统的级联管道(Whisper$\rightarrow$GPT-4o$\rightarrow$TTS)-在准确性,延迟和转向方面。GPT-Realtime在Pass@1(0.600)和中断避免(13.5%)方面领先; Gemini Live 3.1实现了最快的延迟(4.25s),但最低的话轮率(78.0%);级联基线尽管有完美的话轮率,但延迟最高(10.12s)。在所有系统中,自我纠正处理和硬场景下的多步推理仍然是最一致的故障模式。
摘要:We introduce Full-Duplex-Bench-v3 (FDB-v3), a benchmark for evaluating spoken language models under naturalistic speech conditions and multi-step tool use. Unlike prior work, our dataset consists entirely of real human audio annotated for five disfluency categories, paired with scenarios requiring chained API calls across four task domains. We evaluate six model configurations -- GPT-Realtime, Gemini Live 2.5, Gemini Live 3.1, Grok, Ultravox v0.7, and a traditional Cascaded pipeline (Whisper$\rightarrow$GPT-4o$\rightarrow$TTS) -- across accuracy, latency, and turn-taking dimensions. GPT-Realtime leads on Pass@1 (0.600) and interruption avoidance (13.5\%); Gemini Live 3.1 achieves the fastest latency (4.25~s) but the lowest turn-take rate (78.0\%); and the Cascaded baseline, despite a perfect turn-take rate, incurs the highest latency (10.12~s). Across all systems, self-correction handling and multi-step reasoning under hard scenarios remain the most consistent failure modes.
【2】AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
标题:AffectSpeech:具有细粒度文本描述的大规模情感语音数据集,用于语音情感字幕和合成
链接:https://arxiv.org/abs/2604.04160
备注:Submitted to IEEE Transactions
摘要:情感是口语交际中必不可少的一部分,但现有的语音情感建模框架大多依赖于预定义的类别或低维连续属性,表达能力有限。语音情感字幕和合成的最新进展表明,文本描述提供了一个更灵活和可解释的替代表示语音中的情感特征。然而,由于缺乏与可靠和细粒度的自然语言注释相匹配的情感语音数据集,这一方向的进展受到阻碍。为了解决这个问题,我们引入了AffectSpeech,这是一个大规模的人类记录语音语料库,其中包含了用于细粒度情感分析和生成的结构化描述。每个话语的特征在于六个互补的维度,包括情感极性,开放词汇情感标题,强度水平,韵律属性,突出段和语义内容,使语音表达的多粒度建模。为了平衡注释质量和可扩展性,我们采用了一个人-LLM协作注释管道,该管道集成了算法预标记,多LLM描述生成和人在环验证。此外,这些注释被重新表述为不同的描述风格,以提高语言的多样性和减少下游建模的风格偏差。语音情感字幕和合成的实验结果表明,在AffectSpeech上训练的模型在多个评估设置中始终获得优异的性能。
摘要:Emotion is essential in spoken communication, yet most existing frameworks in speech emotion modeling rely on predefined categories or low-dimensional continuous attributes, which offer limited expressive capacity. Recent advances in speech emotion captioning and synthesis have shown that textual descriptions provide a more flexible and interpretable alternative for representing affective characteristics in speech. However, progress in this direction is hindered by the lack of an emotional speech dataset aligned with reliable and fine-grained natural language annotations. To tackle this, we introduce AffectSpeech, a large-scale corpus of human-recorded speech enriched with structured descriptions for fine-grained emotion analysis and generation. Each utterance is characterized across six complementary dimensions, including sentiment polarity, open-vocabulary emotion captions, intensity level, prosodic attributes, prominent segments, and semantic content, enabling multi-granular modeling of vocal expression. To balance annotation quality and scalability, we adopt a human-LLM collaborative annotation pipeline that integrates algorithmic pre-labeling, multi-LLM description generation, and human-in-the-loop verification. Furthermore, these annotations are reformulated into diverse descriptive styles to enhance linguistic diversity and reduce stylistic bias in downstream modeling. Experimental results on speech emotion captioning and synthesis demonstrate that models trained on AffectSpeech consistently achieve superior performance across multiple evaluation settings.
【3】MALEFA: Multi-grAnularity Learning and Effective False Alarm Suppression for Zero-shot Keyword Spotting
标题:MALEFA:多群体学习和有效的虚警抑制,用于零触发关键词发现
链接:https://arxiv.org/abs/2604.03689
备注:Accepted by ICASSP 2026. 5 pages, 4 figures
摘要:用户定义的关键字定位(KWS),而不诉诸特定领域的预标记的训练数据是建立适应性和个性化的语音接口的根本重要性。然而,这样的系统仍然面临着艰巨的挑战,包括有限的计算资源和有限的注释训练数据。现有的方法也很难区分声学上相似的关键字,这通常会导致现实部署中令人讨厌的误报率(FAR)。为了减轻这些限制,我们提出了MALEFA,一个新的轻量级的zero-shot KWS框架,共同学习话语和音素级对齐通过交叉注意和多粒度对比学习目标。在四个公共基准数据集上的评估表明,MALEFA实现了90%的高准确率,在AMI数据集上将FAR显著降低到0.007%。除了强大的性能外,MALEFA还具有高计算效率,可以随时支持在资源受限的设备上进行实时部署。
摘要:User-defined keyword spotting (KWS) without resorting to domain-specific pre-labeled training data is of fundamental importance in building adaptable and personalized voice interfaces. However, such systems are still faced with arduous challenges, including constrained computational resources and limited annotated training data. Existing methods also struggle to distinguish acoustically similar keywords, often leading to a pesky false alarm rate (FAR) in real-world deployments. To mitigate these limitations, we put forward MALEFA, a novel lightweight zero-shot KWS framework that jointly learns utterance- and phoneme-level alignments via cross-attention and a multi-granularity contrastive learning objective. Evaluations on four public benchmark datasets show that MALEFA achieves a high accuracy of 90%, significantly reducing FAR to 0.007% on the AMI dataset. Beyond its strong performance, MALEFA demonstrates high computational efficiency and can readily support real-time deployment on resource-constrained devices.
【4】Rewriting TTS Inference Economics: Lightning V2 on Tenstorrent Achieves 4x Lower Cost Than NVIDIA L40S
标题:重写TTC推理经济学:Tenstorrent上的Lightning V2成本比NVIDIA L40 S低4倍
链接:https://arxiv.org/abs/2604.03279
摘要:文本到语音(TTS)模型在数值上比大型语言模型(LLM)脆弱得多,这是由于它们的连续波形生成和对小数值扰动的感知敏感性。虽然诸如BlockFloat8(BFP8)和低保真度(LoFi)计算等积极的精度降低技术已被广泛用于语言模型中,但将类似的策略应用于TTS系统通常会导致可听伪影,相位不稳定和频谱失真。 在这项工作中,我们提出了Lightning V2,这是一种针对Tenstorrent硬件协同优化的生产级TTS模型。通过精确感知的架构设计和软硬件协同优化,我们实现了超过95%的LoFi计算保真度和超过80%的BlockFloat8部署,而没有可测量的音频质量下降。利用Tenstorrent的片上网络(NoC)、分布式SRAM和确定性执行模型,我们减少了内存移动和冗余权重获取,从而实现高效的低精度推理。 与NVIDIA L40S基准相比,Lightning V2在同等吞吐量下实现了约4倍的本地加速器成本,同时保持了生产音频保真度。我们的研究结果表明,精确协同设计,结合硬件感知优化,可以从根本上重塑实时语音推理的经济性。
摘要:Text-to-Speech (TTS) models are significantly more numerically fragile than Large Language Models (LLMs) due to their continuous waveform generation and perceptual sensitivity to small numerical perturbations. While aggressive precision reduction techniques such as BlockFloat8 (BFP8) and low-fidelity (LoFi) compute have been widely adopted in language models, applying similar strategies to TTS systems often results in audible artifacts, phase instability, and spectral distortion. In this work, we present Lightning V2, a production-grade TTS model co-optimized for Tenstorrent hardware. Through precision-aware architectural design and hardware-software co-optimization, we achieve over 95% LoFi computational fidelity and more than 80% BlockFloat8 deployment without measurable degradation in audio quality. Leveraging Tenstorrent's Network-on-Chip (NoC), distributed SRAM, and deterministic execution model, we reduce memory movement and redundant weight fetches, enabling efficient low-precision inference. Compared to an NVIDIA L40S baseline, Lightning V2 achieves approximately 4x lower on-prem accelerator cost at equivalent throughput, while maintaining production audio fidelity. Our results demonstrate that precision co-design, combined with hardware-aware optimization, can fundamentally reshape the economics of real-time speech inference.
【5】Joint Fullband-Subband Modeling for High-Resolution SingFake Detection
标题:用于高分辨率SingFake检测的全带-子带联合建模
链接:https://arxiv.org/abs/2604.04841
备注:Submitted to INTERSPEECH 2026
摘要:歌声合成的快速发展增加了未经授权的模仿风险,迫切需要更好的歌声Deepfake(SingFake)检测,也称为SVDD。与语音不同,歌唱包含复杂的音高,宽动态范围和音色变化。传统的16 kHz采样探测器被证明是不够的,因为它们丢弃了重要的高频信息。这项研究提出了第一个系统的分析高分辨率(44.1 kHz采样率)音频SVDD。我们提出了一个联合全波段子带建模框架:全波段捕捉全球范围内,而特定的子带专家隔离细粒度的合成文物不均匀分布在整个频谱。WildSVDD数据集上的实验表明,高频子带提供了必要的补充线索。我们的框架显著优于16 kHz采样模型,证明了高分辨率音频和战略子带集成对于稳健的野外检测至关重要。
摘要:Rapid advances in singing voice synthesis have increased unauthorized imitation risks, creating an urgent need for better Singing Voice Deepfake (SingFake) Detection, also known as SVDD. Unlike speech, singing contains complex pitch, wide dynamic range, and timbral variations. Conventional 16 kHz-sampled detectors prove inadequate, as they discard vital high-frequency information. This study presents the first systematic analysis of high-resolution (44.1 kHz sampling rate) audio for SVDD. We propose a joint fullband-subband modeling framework: the fullband captures global context, while subband-specific experts isolate fine-grained synthesis artifacts unevenly distributed across the spectrum. Experiments on the WildSVDD dataset demonstrate that high-frequency subbands provide essential complementary cues. Our framework significantly outperforms 16 kHz-sampled models, proving that high-resolution audio and strategic subband integration are critical for robust in-the-wild detection.
【6】DHFP-PE: Dual-Precision Hybrid Floating Point Processing Element for AI Acceleration
标题:DHFP-PE:用于人工智能加速的双精度混合浮点处理元件
链接:https://arxiv.org/abs/2604.04507
备注:Accepted in ANRF-sponsored 2nd International Conference on Next Generation Electronics (NEleX-2026)
摘要:低精度算法在人工智能和边缘计算中的快速采用,对节能和灵活的浮点乘法累加(MAC)单元产生了强烈的需求。本文提出了一种全流水线双精度浮点MAC处理引擎,支持FP8格式(E4M3,E5M2)和FP4格式(E2M1,E1M2),专门针对低功耗和高吞吐量AI工作负载进行了优化。所提出的架构采用了一种新的位分区技术,使一个单一的4位单位乘法器操作,无论是作为一个标准的4x4乘法器FP8或作为两个并行的2x2乘法器2位操作数,实现100%的硬件利用率,而无需重复逻辑。该处理引擎采用28 nm工艺实现,工作频率为1.94 GHz,面积为0.00396 mm^2,功耗为2.13 mW,与最先进的设计相比,面积减少了60.4%,功耗节省了86.6%。
摘要:The rapid adoption of low-precision arithmetic in artificial intelligence and edge computing has created a strong demand for energy-efficient and flexible floating-point multiply-accumulate (MAC) units. This paper presents a fully pipelined dual-precision floating-point MAC processing engine supporting FP8 formats (E4M3, E5M2) and FP4 formats (E2M1, E1M2), specifically optimized for low-power and high-throughput AI workloads. The proposed architecture employs a novel bit-partitioning technique that enables a single 4-bit unit multiplier to operate either as a standard 4x4 multiplier for FP8 or as two parallel 2x2 multipliers for 2-bit operands, achieving 100 percent hardware utilization without duplicating logic. Implemented in 28 nm technology, the proposed processing engine achieves an operating frequency of 1.94 GHz with an area of 0.00396 mm^2 and power consumption of 2.13 mW, resulting in up to 60.4 percent area reduction and 86.6 percent power savings compared to state-of-the-art designs.
【7】FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection
标题:FastTurn:统一声学和流语义线索,实现低延迟和稳健的转弯检测
链接:https://arxiv.org/abs/2604.01897
备注:5 pages, 2 figures
摘要:AudioLLM的最新进展使口语对话系统能够超越基于回合的交互,转向实时全双工通信,其中代理必须在用户仍在说话时决定何时说话,屈服或中断。现有的全双工方法要么依赖于缺乏语义理解的语音活动线索,要么依赖于基于ASR的模块,这会引入延迟并在重叠的语音和噪声下降级。此外,现有的数据集很少捕捉现实的互动动态,限制了评估和部署。为了缓解这个问题,我们提出了\textbf{FastTurn},一个低延迟和鲁棒的转弯检测的统一框架。为了在保持性能的同时提高延迟,FastTurn将流式CTC解码与声学特征相结合,从而在保留语义线索的同时从部分观察中实现早期决策。我们还发布了一个基于真实人类对话的测试集,捕捉真实的转折过渡,重叠语音,反向通道,停顿,音高变化和环境噪声。实验表明,与代表性基线相比,FastTurn实现了更高的决策准确性和更低的中断延迟,并且在具有挑战性的声学条件下仍然保持稳健,证明了其对于实际全双工对话系统的有效性。
摘要:Recent advances in AudioLLMs have enabled spoken dialogue systems to move beyond turn-based interaction toward real-time full-duplex communication, where the agent must decide when to speak, yield, or interrupt while the user is still talking. Existing full-duplex approaches either rely on voice activity cues, which lack semantic understanding, or on ASR-based modules, which introduce latency and degrade under overlapping speech and noise. Moreover, available datasets rarely capture realistic interaction dynamics, limiting evaluation and deployment. To mitigate the problem, we propose \textbf{FastTurn}, a unified framework for low-latency and robust turn detection. To advance latency while maintaining performance, FastTurn combines streaming CTC decoding with acoustic features, enabling early decisions from partial observations while preserving semantic cues. We also release a test set based on real human dialogue, capturing authentic turn transitions, overlapping speech, backchannels, pauses, pitch variation, and environmental noise. Experiments show FastTurn achieves higher decision accuracy with lower interruption latency than representative baselines and remains robust under challenging acoustic conditions, demonstrating its effectiveness for practical full-duplex dialogue systems.
机器翻译由腾讯交互翻译提供,仅供参考
