今日论文合集:cs.SD语音8篇,eess.AS音频处理6篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Audio Avatar Fingerprinting: An Approach for Authorized Use of Voice Cloning in the Era of Synthetic Audio
标题:音频化身指纹:合成音频时代授权使用语音克隆的方法
链接:https://arxiv.org/abs/2603.20165

作者:Candice R. Gerstner
摘要:随着人工智能语音合成的进步,在目标语音中生成逼真的音频比以往任何时候都更容易。一个人只需要几秒钟的参考音频从目标,相当字面上把话在目标的人的嘴。这给基于语音的认证系统、视频会议和视听广播平台带来了一系列新的取证相关挑战,我们希望在这些平台上检测合成语音。与此同时,利用人工智能语音合成可以通过低带宽通信和音频增强等功能增强不同的通信模式,从而导致合成音频的合法用例不断增加。在这种情况下,我们想要验证合成的语音是否真的是用户说的。这将需要一种机制来验证给定的合成音频是否由授权身份驱动。我们把这个任务称为音频化身指纹识别。作为在这些新的和新兴的情况下对音频取证的一步,我们分析和扩展了现成的扬声器验证模型,该模型是在取证背景之外开发的,用于虚假语音检测和音频化身指纹识别的任务,这是同类实验中的第一个。此外,我们观察到,没有现有的数据集允许验证合成音频的授权使用的新任务-我们通过引入新的语音取证数据集来解决这个新任务的限制。
摘要:With the advancements in AI speech synthesis, it is easier than ever before to generate realistic audio in a target voice. One only needs a few seconds of reference audio from the target, quite literally putting words in the target person's mouth. This imposes a new set of forensics-related challenges on speech-based authentication systems, videoconferencing, and audio-visual broadcasting platforms, where we want to detect synthetic speech. At the same time, leveraging AI speech synthesis can enhance the different modes of communication through features such as low-bandwidth communication and audio enhancements - leading to ever-increasing legitimate use-cases of synthetic audio. In this case, we want to verify if the synthesized voice is actually spoken by the user. This will require a mechanism to verify whether a given synthetic audio is driven by an authorized identity, or not. We term this task audio avatar fingerprinting. As a step towards audio forensics in these new and emerging situations, we analyze and extend an off-the-shelf speaker verification model developed outside of forensics context for the task of fake speech detection and audio avatar fingerprinting, the first experimentation of its kind. Furthermore, we observe that no existing dataset allows for the novel task of verifying the authorized use of synthetic audio - a limitation which we address by introducing a new speech forensics dataset for this novel task.


【2】FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
标题:FoleyDirector:通过结构化收件箱实现视频到音频生成的细粒度时间引导
链接:https://arxiv.org/abs/2603.19857

作者:You Li,Dewei Zhou,Fan Ma,Fu Li,Dongliang He,Yi Yang
备注:Accepted at IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026, 18 pages
摘要:最近的视频到音频(V2A)方法已经取得了显着的进步,使现实的,高质量的音频合成。然而,他们在多事件场景中或当视觉提示不足时(例如小区域,屏幕外的声音或遮挡或部分可见的对象),难以进行细粒度的时间控制。在本文中,我们提出了FoleyDirector框架,该框架首次在基于DiT的V2A生成中实现了精确的时间指导,同时保留了基础模型的音频质量,并允许在V2A生成和时间控制合成之间进行无缝切换。FoleyDirector引入了结构化时间段(STS),一组对应于短时间段的字幕,以提供更丰富的时间信息。这些功能集成通过脚本引导的时间融合模块,采用时间脚本注意力融合STS功能连贯。为了处理复杂的多事件场景,我们进一步提出了双帧声音合成,实现了并行的帧内和帧外音频生成,并提高了可控性。为了支持训练和评估,我们构建了DirectorSound数据集,并引入了VGGSoundDirector和DirectorBench。实验表明,FoleyDirector大大提高了时间的可控性,同时保持高音频保真度,使用户能够充当Foley导演和推进V2A向更有表现力和可控的生成。
摘要:Recent Video-to-Audio (V2A) methods have achieved remarkable progress, enabling the synthesis of realistic, high-quality audio. However, they struggle with fine-grained temporal control in multi-event scenarios or when visual cues are insufficient, such as small regions, off-screen sounds, or occluded or partially visible objects. In this paper, we propose FoleyDirector, a framework that, for the first time, enables precise temporal guidance in DiT-based V2A generation while preserving the base model's audio quality and allowing seamless switching between V2A generation and temporally controlled synthesis. FoleyDirector introduces Structured Temporal Scripts (STS), a set of captions corresponding to short temporal segments, to provide richer temporal information. These features are integrated via the Script-Guided Temporal Fusion Module, which employs Temporal Script Attention to fuse STS features coherently. To handle complex multi-event scenarios, we further propose Bi-Frame Sound Synthesis, enabling parallel in-frame and out-of-frame audio generation and improving controllability. To support training and evaluation, we construct the DirectorSound dataset and introduce VGGSoundDirector and DirectorBench. Experiments demonstrate that FoleyDirector substantially enhances temporal controllability while maintaining high audio fidelity, empowering users to act as Foley directors and advancing V2A toward more expressive and controllable generation.


【3】Borderless Long Speech Synthesis
标题:无边界长语音合成
链接:https://arxiv.org/abs/2603.19798

作者:Xingchen Song,Di Wu,Dinghao Zhou,Pengyu Cheng,Hongwu Ding,Yunchao He,Jie Wang,Shengfan Shen,Sixiang Lv,Lichun Fan,Hang Su,Yifeng Wang,Shuai Wang,Meng Meng,Jian Luan
摘要:大多数现有的文本到语音(TTS)系统或者逐句合成语音并将结果缝合在一起,或者单独驱动纯文本对话的合成。这两种方法都使模型对全局上下文或非语言线索的理解很少,因此很难捕捉真实世界的现象,例如多说话者交互(中断,重叠语音),不断变化的情感弧和各种声学环境。我们引入了无边界长语音合成框架,用于以代理为中心的无边界长音频合成。而不是针对一个单一的狭窄的任务,该系统被设计为一个统一的能力集跨越VoiceDesigner,多扬声器合成,指令TTS,和长格式文本合成。在数据方面,我们提出了一个“标签过滤/清洗”的策略,并设计了一个自顶向下的,多层次的注释模式,我们称之为全局语句令牌。在模型方面,我们采用了一个连续的tokenizer骨干,并添加了思想链(CoT)推理与维度丢弃,这两个显着提高在复杂条件下的指令遵循。我们进一步表明,该系统是原生的设计:层次注释双打作为LLM代理和合成引擎之间的结构化语义接口,创建一个分层的控制协议栈,从场景语义到语音细节。因此,文本成为一个信息完整的宽带控制通道,使前端LLM转换成结构化的生成命令的任何形式的输入,从Text 2Speech扩展到无边界的长语音合成的范例。
摘要:Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both approaches leave models with little understanding of global context or paralinguistic cues, making it hard to capture real-world phenomena such as multi-speaker interactions (interruptions, overlapping speech), evolving emotional arcs, and varied acoustic environments. We introduce the Borderless Long Speech Synthesis framework for agent-centric, borderless long audio synthesis. Rather than targeting a single narrow task, the system is designed as a unified capability set spanning VoiceDesigner, multi-speaker synthesis, Instruct TTS, and long-form text synthesis. On the data side, we propose a "Labeling over filtering/cleaning" strategy and design a top-down, multi-level annotation schema we call Global-Sentence-Token. On the model side, we adopt a backbone with a continuous tokenizer and add Chain-of-Thought (CoT) reasoning together with Dimension Dropout, both of which markedly improve instruction following under complex conditions. We further show that the system is Native Agentic by design: the hierarchical annotation doubles as a Structured Semantic Interface between the LLM Agent and the synthesis engine, creating a layered control protocol stack that spans from scene semantics down to phonetic detail. Text thereby becomes an information-complete, wide-band control channel, enabling a front-end LLM to convert inputs of any modality into structured generation commands, extending the paradigm from Text2Speech to borderless long speech synthesis.


【4】MOSS-TTSD: Text to Spoken Dialogue Generation
标题:MOSS-TTSD:文本到口语对话的生成
链接:https://arxiv.org/abs/2603.19739

作者:Yuqian Zhang,Donghua Yu,Zhengyuan Lin,Botian Jiang,Mingshu Chen,Yaozhou Jiang,Yiwei Zhao,Yiyang Zhang,Yucheng Yuan,Hanfu Chen,Kexin Huang,Jun Zhan,Cheng Chang,Zhaoye Fei,Shimin Li,Xiaogui Yang,Qinyuan Cheng,Xipeng Qiu
摘要:口语对话生成对于播客、动态评论和娱乐内容等应用至关重要,但与单话语文本到语音(TTS)相比,它带来了重大挑战。关键的要求包括准确的话轮转换,跨话轮声学一致性和长形式的稳定性,目前的模型往往无法解决由于缺乏对话上下文建模。为了弥合这一差距,我们提出了MOSS-TTSD,一个口语对话合成模型,旨在表达,跨多种语言的多方对话语音。通过增强的长上下文建模,MOSS-TTSD从具有显式扬声器标签的对话脚本生成长形式的口语对话,支持长达60分钟的单通道合成,最多5个扬声器的多方对话,以及从短参考音频片段中克隆zero-shot语音。该模型支持多种主流语言,包括英语和中文,并适应多种长格式场景。此外,为了解决现有评估方法的局限性,我们提出了TTSD-eval,一个客观的评估框架,基于强制对齐,测量说话人属性的准确性和说话人相似性,而不依赖于说话人日记工具。客观和主观的评估结果表明,MOSS-TTSD超越强大的开源和专有的对话合成基线。
摘要:Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance text-to-speech (TTS). Key requirements include accurate turn-taking, cross-turn acoustic consistency, and long-form stability, which current models often fail to address due to a lack of dialogue context modeling. To bridge this gap, we present MOSS-TTSD, a spoken dialogue synthesis model designed for expressive, multi-party conversational speech across multiple languages. With enhanced long-context modeling, MOSS-TTSD generates long-form spoken conversations from dialogue scripts with explicit speaker tags, supporting up to 60 minutes of single-pass synthesis, multi-party dialogue with up to 5 speakers, and zero-shot voice cloning from a short reference audio clip. The model supports various mainstream languages, including English and Chinese, and is adapted to several long-form scenarios. Additionally, to address limitations of existing evaluation methods, we propose TTSD-eval, an objective evaluation framework based on forced alignment that measures speaker attribution accuracy and speaker similarity without relying on speaker diarization tools. Both objective and subjective evaluation results show that MOSS-TTSD surpasses strong open-source and proprietary baselines in dialogue synthesis.


【5】CAF-Score: Calibrating CLAP with LALMs for Reference-free Audio Captioning Evaluation
标题:CAF-Score:使用LALM校准CLAP,以进行无参考音频字幕评估
链接:https://arxiv.org/abs/2603.19615

作者:Insung Lee,Taeyoung Jeong,Haejun Yoo,Du-Seong Chang,Myoung-Wan Koo
备注:A condensed version of this work has been submitted to Interspeech 2026. Section 10 is an extended analysis added in this version
摘要:虽然大型音频语言模型(LALM)具有先进的音频字幕,但稳健的评估仍然很困难。基于参考的度量是昂贵的,并且通常无法评估声学保真度,而基于对比存储音频预训练(CLAP)的方法经常忽略语法错误和细粒度的细节。我们提出CAF-Score,一个无参考的度量标准,校准CLAP的粗粒度的语义对齐与细粒度的理解和语法意识的LALM。通过将对比音频文本嵌入与LALM推理相结合,CAF-Score可以有效地检测句法不一致和微妙的幻觉。在BRACE基准测试上的实验表明,我们的方法与人类判断的相关性最高,甚至在具有挑战性的场景中优于基于参考的基线。这些结果突出了CAF-Score用于无参考音频字幕评估的有效性。代码和结果可在https://github.com/inseong00/CAF-Score上获得。
摘要:While Large Audio-Language Models (LALMs) have advanced audio captioning, robust evaluation remains difficult. Reference-based metrics are expensive and often fail to assess acoustic fidelity, while Contrastive Language-Audio Pretraining (CLAP)-based approaches frequently overlook syntactic errors and fine-grained details. We propose CAF-Score, a reference-free metric that calibrates CLAP's coarse-grained semantic alignment with the fine-grained comprehension and syntactic awareness of LALMs. By combining contrastive audio-text embeddings with LALM reasoning, CAF-Score effectively detects syntactic inconsistencies and subtle hallucinations. Experiments on the BRACE benchmark demonstrate that our approach achieves the highest correlation with human judgments, even outperforming reference-based baselines in challenging scenarios. These results highlight the efficacy of CAF-Score for reference-free audio captioning evaluation. Code and results are available at https://github.com/inseong00/CAF-Score.


【6】Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
标题:先听,然后回答:基于时间戳的语音推理
链接:https://arxiv.org/abs/2603.19468

作者:Jihoon Jeong,Pooneh Mousavi,Mirco Ravanelli,Cem Subakan
备注:Submitted to Interspeech 2026
摘要:大型音频语言模型(LALM)可以为其预测生成推理链,但目前尚不清楚这些推理链是否仍然基于输入音频。在本文中,我们提出了一个基于RL的策略,理由的LALM推理输出明确的时间戳注释指的是音频信号的相关部分。我们的分析表明,时间戳接地导致模型在推理生成过程中更强烈地关注音频令牌。在四个基于语音的基准数据集上的实验表明,与zero-shot推理和没有时间戳基础的微调相比,我们的方法提高了性能。此外,接地放大理想的推理行为,如区域探索,听力学验证和一致性,强调接地机制的重要性,忠实的多模态推理。
摘要:Large audio-language models (LALMs) can generate reasoning chains for their predictions, but it remains unclear whether these reasoning chains remain grounded in the input audio. In this paper, we propose an RL-based strategy that grounds the reasoning outputs of LALMs with explicit timestamp annotations referring to relevant segments of the audio signal. Our analysis shows that timestamp grounding leads the model to attend more strongly to audio tokens during reasoning generation. Experiments on four speech-based benchmark datasets demonstrate that our approach improves performance compared to both zero-shot reasoning and fine-tuning without timestamp grounding. Additionally, grounding amplifies desirable reasoning behaviors, such as region exploration, audiology verification, and consistency, underscoring the importance of grounding mechanisms for faithful multimodal reasoning.


【7】BioDCASE 2026 Challenge Baseline for Cross-Domain Mosquito Species Classification
标题:BioDUSE 2026跨领域蚊子物种分类挑战基线
链接:https://arxiv.org/abs/2603.20118

作者:Yuanbo Hou,Vanja Zdravkovic,Marianne Sinka,Yunpeng Li,Wenwu Wang,Mark D. Plumbley,Kathy Willis,Stephen Roberts
备注:BioDCASE 2026 CD-MSC Baseline, source code and models: https://github.com/Yuanbo2020/CD-MSC
摘要:蚊子传播的疾病每年影响超过10亿人,并造成近100万人死亡。传统的监测方法依赖于陷阱和人工识别,速度慢,劳动密集,难以扩展。基于音频的蚊子监测为基于陷阱的监测提供了一种非破坏性、低成本和更可扩展的补充,但在真实世界的记录条件下,可靠的物种分类仍然很困难。蚊子的飞行音调是窄带的,通常信噪比很低,很容易被背景噪音掩盖,几个流行病学相关物种的记录仍然有限,造成明显的阶级不平衡。设备、环境和收集协议之间的差异进一步增加了稳健分类的难度。这种变化可能会导致模型依赖于特定领域的记录文物,而不是物种相关的声学线索,这使得转移到新的采集设置困难。BioDCASE 2026跨域苔藓物种分类(CD-MSC)挑战赛是围绕这一部署问题设计的,通过评估可见和不可见域的性能。本文介绍了官方基线系统和评估管道,作为CD-MSC挑战任务的简单,完全可重复的参考。基线使用对数梅尔特征和具有物种和辅助域输出的多时间分辨率卷积神经网络(MTRCNN),以及完整的训练和测试脚本。基线系统表现强劲的领域,但看不见的领域显着下降,表明跨域概括,而不是域内识别,是从多源生物声学记录的实际蚊子物种分类的核心挑战。
摘要:Mosquito-borne diseases affect more than one billion people each year and cause close to one million deaths. Traditional surveillance methods rely on traps and manual identification that are slow, labor-intensive, and difficult to scale. Audio-based mosquito monitoring offers a non-destructive, lower-cost, and more scalable complement to trap-based surveillance, but reliable species classification remains difficult under real-world recording conditions. Mosquito flight tones are narrow-band, often low in signal-to-noise ratio, and easily masked by background noise, and recordings for several epidemiologically relevant species remain limited, creating pronounced class imbalance. Variation across devices, environments, and collection protocols further increases the difficulty of robust classification. Such variation can cause models to rely on domain-specific recording artefacts rather than species-relevant acoustic cues, which makes transfer to new acquisition settings difficult. The BioDCASE 2026 Cross-Domain Mosquito Species Classification (CD-MSC) challenge is designed around this deployment problem by evaluating performance on both seen and unseen domains. This paper presents the official baseline system and evaluation pipeline as a simple, fully reproducible reference for the CD-MSC challenge task. The baseline uses log-mel features and a multitemporal resolution convolutional neural network (MTRCNN) with species and auxiliary domain outputs, together with complete training and test scripts. The baseline system performs strongly on seen domains but degrades markedly on unseen domains, showing that cross-domain generalisation, rather than within-domain recognition, is the central challenge for practical mosquito species classification from multi-source bioacoustic recordings.


【8】Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
标题:即插即用:视听目标说话人提取中的分离和选择脱钩
链接:https://arxiv.org/abs/2603.19697

作者:Doyeop Kwak,Suyeon Lee,Joon Son Chung
备注:Submitted to Interspeech 2026; demo available https://plugandsteer.github.io
摘要:本文的目的是提供一个新的视角视听目标说话人提取(AV-TSE)分离和目标选择。传统的AV-TSE系统通常将音频和视觉特征深度集成以重新学习整个分离过程,由于野外视听数据集的噪声性质,这可以充当保真度天花板。为了解决这个问题,我们提出了插件和转向,它分配高保真分离到一个冻结的音频只骨干和限制的作用,视觉模态严格的目标选择。我们引入了潜在的转向矩阵(LSM),一个最低限度的线性变换,重新路由的骨干内的潜在功能锚定的目标扬声器到指定的通道。在四个有代表性的架构的实验表明,我们的方法有效地保留了不同的骨干的声学先验,实现与原始骨干的感知质量。音频样本可在以下网址获得:https://plugandsteer.github.io
摘要:The goal of this paper is to provide a new perspective on audio-visual target speaker extraction (AV-TSE) by decoupling the separation and target selection. Conventional AV-TSE systems typically integrate audio and visual features deeply to re-learn the entire separation process, which can act as a fidelity ceiling due to the noisy nature of in-the-wild audio-visual datasets. To address this, we propose Plug-and-Steer, which assigns high-fidelity separation to a frozen audio-only backbone and limits the role of visual modality strictly to target selection. We introduce the Latent Steering Matrix (LSM), a minimalist linear transformation that re-routes latent features within the backbone to anchor the target speaker to a designated channel. Experiments across four representative architectures show that our method effectively preserves the acoustic priors of diverse backbones, achieving perceptual quality comparable to the original backbones. Audio samples are available at: https://plugandsteer.github.io


eess.AS音频处理


【1】BioDCASE 2026 Challenge Baseline for Cross-Domain Mosquito Species Classification
标题:BioDUSE 2026跨领域蚊子物种分类挑战基线
链接:https://arxiv.org/abs/2603.20118

作者:Yuanbo Hou,Vanja Zdravkovic,Marianne Sinka,Yunpeng Li,Wenwu Wang,Mark D. Plumbley,Kathy Willis,Stephen Roberts
备注:BioDCASE 2026 CD-MSC Baseline, source code and models: https://github.com/Yuanbo2020/CD-MSC
摘要:蚊子传播的疾病每年影响超过10亿人,并造成近100万人死亡。传统的监测方法依赖于陷阱和人工识别,速度慢,劳动密集,难以扩展。基于音频的蚊子监测为基于陷阱的监测提供了一种非破坏性、低成本和更可扩展的补充,但在真实世界的记录条件下,可靠的物种分类仍然很困难。蚊子的飞行音调是窄带的,通常信噪比很低,很容易被背景噪音掩盖,几个流行病学相关物种的记录仍然有限,造成明显的阶级不平衡。设备、环境和收集协议之间的差异进一步增加了稳健分类的难度。这种变化可能会导致模型依赖于特定领域的记录文物,而不是物种相关的声学线索,这使得转移到新的采集设置困难。BioDCASE 2026跨域苔藓物种分类(CD-MSC)挑战赛是围绕这一部署问题设计的,通过评估可见和不可见域的性能。本文介绍了官方基线系统和评估管道,作为CD-MSC挑战任务的简单,完全可重复的参考。基线使用对数梅尔特征和具有物种和辅助域输出的多时间分辨率卷积神经网络(MTRCNN),以及完整的训练和测试脚本。基线系统表现强劲的领域,但看不见的领域显着下降,表明跨域概括,而不是域内识别,是从多源生物声学记录的实际蚊子物种分类的核心挑战。
摘要:Mosquito-borne diseases affect more than one billion people each year and cause close to one million deaths. Traditional surveillance methods rely on traps and manual identification that are slow, labor-intensive, and difficult to scale. Audio-based mosquito monitoring offers a non-destructive, lower-cost, and more scalable complement to trap-based surveillance, but reliable species classification remains difficult under real-world recording conditions. Mosquito flight tones are narrow-band, often low in signal-to-noise ratio, and easily masked by background noise, and recordings for several epidemiologically relevant species remain limited, creating pronounced class imbalance. Variation across devices, environments, and collection protocols further increases the difficulty of robust classification. Such variation can cause models to rely on domain-specific recording artefacts rather than species-relevant acoustic cues, which makes transfer to new acquisition settings difficult. The BioDCASE 2026 Cross-Domain Mosquito Species Classification (CD-MSC) challenge is designed around this deployment problem by evaluating performance on both seen and unseen domains. This paper presents the official baseline system and evaluation pipeline as a simple, fully reproducible reference for the CD-MSC challenge task. The baseline uses log-mel features and a multitemporal resolution convolutional neural network (MTRCNN) with species and auxiliary domain outputs, together with complete training and test scripts. The baseline system performs strongly on seen domains but degrades markedly on unseen domains, showing that cross-domain generalisation, rather than within-domain recognition, is the central challenge for practical mosquito species classification from multi-source bioacoustic recordings.


【2】Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?
标题:Gesture 2 Speech:手部动作能在多大程度上塑造表达性演讲?
链接:https://arxiv.org/abs/2603.19831

作者:Lokesh Kumar,Nirmesh Shah,Ashishkumar P. Gudmalwar,Pankaj Wasnik
备注:Accepted at The 2nd International Workshop on Bodily Expressed Emotion Understanding (BEEU) at AAAI 2026 [non-archival]
摘要:人类交流无缝地集成了语音和身体动作,手势自然地补充了语音韵律来表达意图,情感和重点。虽然最近的文本到语音(TTS)系统已经开始纳入多模态线索,如面部表情或嘴唇运动,手势在塑造韵律的作用仍然在很大程度上未被探索。我们提出了一种新的多模态TTS框架,Gesture2Speech,利用视觉手势线索来调制合成语音中的韵律。出于观察,自信和富有表现力的扬声器协调手势与语音韵律,我们引入了一个多模态混合专家(MoE)架构,动态融合语言内容和手势功能在一个专用的风格提取模块。融合的表示条件的LLM为基础的语音解码器,使韵律调制,是时间上对齐的手的运动。我们进一步设计了一个手势语音对齐损失,明确建模其时间对应,以确保手势和韵律轮廓之间的细粒度同步。对PATS数据集的评估表明,Gesture2Speech在语音自然度和手势语音同步方面都优于最先进的基线。据我们所知,这是第一个工作,利用手势线索的韵律控制神经语音合成。演示示例可在https://research.sri-media-analysis.com/aaai26-beeu-gesture2speech/上获得
摘要:Human communication seamlessly integrates speech and bodily motion, where hand gestures naturally complement vocal prosody to express intent, emotion, and emphasis. While recent text-to-speech (TTS) systems have begun incorporating multimodal cues such as facial expressions or lip movements, the role of hand gestures in shaping prosody remains largely underexplored. We propose a novel multimodal TTS framework, Gesture2Speech, that leverages visual gesture cues to modulate prosody in synthesized speech. Motivated by the observation that confident and expressive speakers coordinate gestures with vocal prosody, we introduce a multimodal Mixture-of-Experts (MoE) architecture that dynamically fuses linguistic content and gesture features within a dedicated style extraction module. The fused representation conditions an LLM-based speech decoder, enabling prosodic modulation that is temporally aligned with hand movements. We further design a gesture-speech alignment loss that explicitly models their temporal correspondence to ensure fine-grained synchrony between gestures and prosodic contours. Evaluations on the PATS dataset show that Gesture2Speech outperforms state-of-the-art baselines in both speech naturalness and gesture-speech synchrony. To the best of our knowledge, this is the first work to utilize hand gesture cues for prosody control in neural speech synthesis. Demo samples are available at https://research.sri-media-analysis.com/aaai26-beeu-gesture2speech/


【3】Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
标题:即插即用:视听目标说话人提取中的分离和选择脱钩
链接:https://arxiv.org/abs/2603.19697

作者:Doyeop Kwak,Suyeon Lee,Joon Son Chung
备注:Submitted to Interspeech 2026; demo available https://plugandsteer.github.io
摘要:本文的目的是提供一个新的视角视听目标说话人提取(AV-TSE)分离和目标选择。传统的AV-TSE系统通常将音频和视觉特征深度集成以重新学习整个分离过程,由于野外视听数据集的噪声性质,这可以充当保真度天花板。为了解决这个问题,我们提出了插件和转向,它分配高保真分离到一个冻结的音频只骨干和限制的作用,视觉模态严格的目标选择。我们引入了潜在的转向矩阵(LSM),一个最低限度的线性变换,重新路由的骨干内的潜在功能锚定的目标扬声器到指定的通道。在四个有代表性的架构的实验表明,我们的方法有效地保留了不同的骨干的声学先验,实现与原始骨干的感知质量。音频样本可在以下网址获得:https://plugandsteer.github.io
摘要:The goal of this paper is to provide a new perspective on audio-visual target speaker extraction (AV-TSE) by decoupling the separation and target selection. Conventional AV-TSE systems typically integrate audio and visual features deeply to re-learn the entire separation process, which can act as a fidelity ceiling due to the noisy nature of in-the-wild audio-visual datasets. To address this, we propose Plug-and-Steer, which assigns high-fidelity separation to a frozen audio-only backbone and limits the role of visual modality strictly to target selection. We introduce the Latent Steering Matrix (LSM), a minimalist linear transformation that re-routes latent features within the backbone to anchor the target speaker to a designated channel. Experiments across four representative architectures show that our method effectively preserves the acoustic priors of diverse backbones, achieving perceptual quality comparable to the original backbones. Audio samples are available at: https://plugandsteer.github.io


【4】Audio Avatar Fingerprinting: An Approach for Authorized Use of Voice Cloning in the Era of Synthetic Audio
标题:音频化身指纹:合成音频时代授权使用语音克隆的方法
链接:https://arxiv.org/abs/2603.20165

作者:Candice R. Gerstner
摘要:随着人工智能语音合成的进步,在目标语音中生成逼真的音频比以往任何时候都更容易。一个人只需要几秒钟的参考音频从目标,相当字面上把话在目标的人的嘴。这给基于语音的认证系统、视频会议和视听广播平台带来了一系列新的取证相关挑战,我们希望在这些平台上检测合成语音。与此同时,利用人工智能语音合成可以通过低带宽通信和音频增强等功能增强不同的通信模式,从而导致合成音频的合法用例不断增加。在这种情况下,我们想要验证合成的语音是否真的是用户说的。这将需要一种机制来验证给定的合成音频是否由授权身份驱动。我们把这个任务称为音频化身指纹识别。作为在这些新的和新兴的情况下对音频取证的一步,我们分析和扩展了现成的扬声器验证模型,该模型是在取证背景之外开发的,用于虚假语音检测和音频化身指纹识别的任务,这是同类实验中的第一个。此外,我们观察到,没有现有的数据集允许验证合成音频的授权使用的新任务-我们通过引入新的语音取证数据集来解决这个新任务的限制。
摘要:With the advancements in AI speech synthesis, it is easier than ever before to generate realistic audio in a target voice. One only needs a few seconds of reference audio from the target, quite literally putting words in the target person's mouth. This imposes a new set of forensics-related challenges on speech-based authentication systems, videoconferencing, and audio-visual broadcasting platforms, where we want to detect synthetic speech. At the same time, leveraging AI speech synthesis can enhance the different modes of communication through features such as low-bandwidth communication and audio enhancements - leading to ever-increasing legitimate use-cases of synthetic audio. In this case, we want to verify if the synthesized voice is actually spoken by the user. This will require a mechanism to verify whether a given synthetic audio is driven by an authorized identity, or not. We term this task audio avatar fingerprinting. As a step towards audio forensics in these new and emerging situations, we analyze and extend an off-the-shelf speaker verification model developed outside of forensics context for the task of fake speech detection and audio avatar fingerprinting, the first experimentation of its kind. Furthermore, we observe that no existing dataset allows for the novel task of verifying the authorized use of synthetic audio - a limitation which we address by introducing a new speech forensics dataset for this novel task.


【5】Borderless Long Speech Synthesis
标题:无边界长语音合成
链接:https://arxiv.org/abs/2603.19798

作者:Xingchen Song,Di Wu,Dinghao Zhou,Pengyu Cheng,Hongwu Ding,Yunchao He,Jie Wang,Shengfan Shen,Sixiang Lv,Lichun Fan,Hang Su,Yifeng Wang,Shuai Wang,Meng Meng,Jian Luan
摘要:大多数现有的文本到语音(TTS)系统或者逐句合成语音并将结果缝合在一起,或者单独驱动纯文本对话的合成。这两种方法都使模型对全局上下文或非语言线索的理解很少,因此很难捕捉真实世界的现象,例如多说话者交互(中断,重叠语音),不断变化的情感弧和各种声学环境。我们引入了无边界长语音合成框架,用于以代理为中心的无边界长音频合成。而不是针对一个单一的狭窄的任务,该系统被设计为一个统一的能力集跨越VoiceDesigner,多扬声器合成,指令TTS,和长格式文本合成。在数据方面,我们提出了一个“标签过滤/清洗”的策略,并设计了一个自顶向下的,多层次的注释模式,我们称之为全局语句令牌。在模型方面,我们采用了一个连续的tokenizer骨干,并添加了思想链(CoT)推理与维度丢弃,这两个显着提高在复杂条件下的指令遵循。我们进一步表明,该系统是原生的设计:层次注释双打作为LLM代理和合成引擎之间的结构化语义接口,创建一个分层的控制协议栈,从场景语义到语音细节。因此,文本成为信息完整的宽带控制通道,使前端LLM能够将任何模态的输入转换为结构化生成命令,将范式从Text 2Speech扩展到无边界长语音合成。
摘要:Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both approaches leave models with little understanding of global context or paralinguistic cues, making it hard to capture real-world phenomena such as multi-speaker interactions (interruptions, overlapping speech), evolving emotional arcs, and varied acoustic environments. We introduce the Borderless Long Speech Synthesis framework for agent-centric, borderless long audio synthesis. Rather than targeting a single narrow task, the system is designed as a unified capability set spanning VoiceDesigner, multi-speaker synthesis, Instruct TTS, and long-form text synthesis. On the data side, we propose a "Labeling over filtering/cleaning" strategy and design a top-down, multi-level annotation schema we call Global-Sentence-Token. On the model side, we adopt a backbone with a continuous tokenizer and add Chain-of-Thought (CoT) reasoning together with Dimension Dropout, both of which markedly improve instruction following under complex conditions. We further show that the system is Native Agentic by design: the hierarchical annotation doubles as a Structured Semantic Interface between the LLM Agent and the synthesis engine, creating a layered control protocol stack that spans from scene semantics down to phonetic detail. Text thereby becomes an information-complete, wide-band control channel, enabling a front-end LLM to convert inputs of any modality into structured generation commands, extending the paradigm from Text2Speech to borderless long speech synthesis.


【6】Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
标题:先听,然后回答:基于时间戳的语音推理
链接:https://arxiv.org/abs/2603.19468

作者:Jihoon Jeong,Pooneh Mousavi,Mirco Ravanelli,Cem Subakan
备注:Submitted to Interspeech 2026
摘要:大型音频语言模型(LALM)可以为其预测生成推理链,但目前尚不清楚这些推理链是否仍然基于输入音频。在本文中,我们提出了一个基于RL的策略,理由的LALM推理输出明确的时间戳注释指的是音频信号的相关部分。我们的分析表明,时间戳接地导致模型在推理生成过程中更强烈地关注音频令牌。在四个基于语音的基准数据集上的实验表明,与zero-shot推理和没有时间戳基础的微调相比,我们的方法提高了性能。此外,接地放大理想的推理行为,如区域探索,听力学验证和一致性,强调接地机制的重要性,忠实的多模态推理。
摘要:Large audio-language models (LALMs) can generate reasoning chains for their predictions, but it remains unclear whether these reasoning chains remain grounded in the input audio. In this paper, we propose an RL-based strategy that grounds the reasoning outputs of LALMs with explicit timestamp annotations referring to relevant segments of the audio signal. Our analysis shows that timestamp grounding leads the model to attend more strongly to audio tokens during reasoning generation. Experiments on four speech-based benchmark datasets demonstrate that our approach improves performance compared to both zero-shot reasoning and fine-tuning without timestamp grounding. Additionally, grounding amplifies desirable reasoning behaviors, such as region exploration, audiology verification, and consistency, underscoring the importance of grounding mechanisms for faithful multimodal reasoning.


机器翻译由腾讯交互翻译提供,仅供参考