微信公众号:arXiv_Daily
cs.SD语音
【1】V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
标题:V2 M-Zero:零对时间对齐的视频到音乐一代
链接:https://arxiv.org/abs/2603.11042
备注:Project page: https://genjib.github.io/v2m_zero/
摘要:生成在时间上与视频事件对齐的音乐对于现有的文本到音乐模型是具有挑战性的,这些模型缺乏细粒度的时间控制。我们介绍了V2 M-Zero,一种零对视频到音乐生成方法,可以为视频输出时间对齐的音乐。我们的方法的动机是一个关键的观察:时间同步需要匹配的时间和多少变化发生,而不是什么变化。虽然音乐和视觉事件在语义上不同,但它们表现出共享的时间结构,可以在每种模态中独立捕获。我们通过使用预先训练的音乐和视频编码器从模态内相似性计算的事件曲线来捕获这种结构。通过独立测量每个模态内的时间变化,这些曲线提供了跨模态的可比表示。这实现了一个简单的训练策略:在音乐事件曲线上微调文本到音乐的模型,然后在没有交叉模式训练或配对数据的情况下在推理时替换视频事件曲线。在OES-Pub、MovieGenBench-Music和AIST++中,V2 M-Zero在配对数据基线上实现了显著的收益:音频质量提高了5-21%,语义对齐提高了13-15%,时间同步提高了21-52%,舞蹈视频的节拍对齐提高了28%。我们通过一个大型的众包主观听力测试发现了类似的结果。总的来说,我们的研究结果验证了时间对齐通过内模态功能,而不是成对的跨模态监督,是有效的视频到音乐的生成。结果见https://genjib.github.io/v2m_zero/
摘要:Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-Zero, a zero-pair video-to-music generation approach that outputs time-aligned music for video. Our method is motivated by a key observation: temporal synchronization requires matching when and how much change occurs, not what changes. While musical and visual events differ semantically, they exhibit shared temporal structure that can be captured independently within each modality. We capture this structure through event curves computed from intra-modal similarity using pretrained music and video encoders. By measuring temporal change within each modality independently, these curves provide comparable representations across modalities. This enables a simple training strategy: fine-tune a text-to-music model on music-event curves, then substitute video-event curves at inference without cross-modal training or paired data. Across OES-Pub, MovieGenBench-Music, and AIST++, V2M-Zero achieves substantial gains over paired-data baselines: 5-21% higher audio quality, 13-15% better semantic alignment, 21-52% improved temporal synchronization, and 28% higher beat alignment on dance videos. We find similar results via a large crowd-source subjective listening test. Overall, our results validate that temporal alignment through within-modality features, rather than paired cross-modal supervision, is effective for video-to-music generation. Results are available at https://genjib.github.io/v2m_zero/
【2】Training-Free Multi-Step Inference for Target Speaker Extraction
标题:目标说话人提取的免训练多步推理
链接:https://arxiv.org/abs/2603.10921
摘要:目标说话人提取(TSE)的目的是使用参考话语作为线索从混合语音中恢复目标说话人的语音。大多数TSE系统采用具有一步推理的条件自动编码器架构。受测试时间缩放的启发,我们提出了一种无需训练的多步推理方法,该方法可以使用冻结的预训练模型进行迭代细化。在每一步中,通过对原始混合和先前估计进行插值来生成新的候选,并且选择最佳候选以进一步细化直到收敛。实验表明,当地面实况目标语音可用时,优化侵入性度量(SI-SDRi)在多个评估度量中产生一致的增益。在没有地面实况的情况下,优化非侵入性指标(UTMOS或SpkSim)可以改进相应的指标,但可能会损害其他指标。因此,我们引入联合度量优化来平衡这些目标,使可控的提取偏好的实际部署。
摘要:Target speaker extraction (TSE) aims to recover a target speaker's speech from a mixture using a reference utterance as a cue. Most TSE systems adopt conditional auto-encoder architectures with one-step inference. Inspired by test-time scaling, we propose a training-free multi-step inference method that enables iterative refinement with a frozen pretrained model. At each step, new candidates are generated by interpolating the original mixture and the previous estimate, and the best candidate is selected for further refinement until convergence. Experiments show that, when ground-truth target speech is available, optimizing an intrusive metric (SI-SDRi) yields consistent gains across multiple evaluation metrics. Without ground truth, optimizing non-intrusive metrics (UTMOS or SpkSim) improves the corresponding metric but may hurt others. We therefore introduce joint metric optimization to balance these objectives, enabling controllable extraction preferences for practical deployment.
【3】When Fine-Tuning Fails and when it Generalises: Role of Data Diversity and Mixed Training in LLM-based TTS
标题:当微调失败时,当它泛化时:数据多样性和混合训练在基于LLM的TTS中的作用
链接:https://arxiv.org/abs/2603.10904
备注:We finetune the Qwen 0.5B backbone in an LLM TTS with LoRA to raise MOS speaker similarity and SNR. It works best with diverse training audio with uniform data it can amplify noise so tune decoding and use GGUF quantization for low latency stable quality
摘要:大型语言模型越来越多地被用作神经文本到语音系统的语义骨干。然而,冻结LLM表示不足以建模扬声器特定的声学和感知特性。我们的实验涉及微调的TTS的语言模型骨干显示出改善语音克隆任务中的语音一致性和信噪比SNR的承诺。在多个扬声器中,LoRA微调在语音质量的三个互补维度上始终优于非微调的基础Qwen-0.5B模型。首先,对于训练数据表现出足够的声学可变性的扬声器,感知质量显著提高,DNS-MOS增益高达0.42点。其次,扬声器保真度提高所有评估的扬声器与语音相似性的一致增加,表明LoRA有效地适应扬声器身份表示,而不会降低语言建模。第三,在大多数情况下,信号电平质量得到改善,信噪比提高了34%。至关重要的是,这些改进在很大程度上取决于训练数据的特征。具有高可变性的声能和感知质量的扬声器实现DNS-MOS语音相似性和SNR的同时增益。总的来说,这项工作建立了LoRA微调不仅是一个参数有效的优化技术,但在紧凑的基于LLM的TTS系统更好的扬声器电平自适应的有效机制。当有足够多样化的训练数据支持时,LoRA适配的Qwen-0.5B在感知质量扬声器相似性方面始终优于其冻结的基础模型,并且使用以量化形式托管的GGUF模型具有低延迟。
摘要:Large language models are increasingly adopted as semantic backbones for neural text-to-speech systems. However, frozen LLM representations are insufficient for modeling speaker specific acoustic and perceptual characteristics. Our experiments involving fine tuning of the Language Model backbone of TTS show promise in improving the voice consistency and Signal to Noise ratio SNR in voice cloning task. Across multiple speakers LoRA finetuning consistently outperforms the non-finetuned base Qwen-0.5B model across three complementary dimensions of speech quality. First, perceptual quality improves significantly with DNS-MOS gains of up to 0.42 points for speakers whose training data exhibits sufficient acoustic variability. Second, speaker fidelity improves for all evaluated speakers with consistent increases in voice similarity indicating that LoRA effectively adapts speaker identity representations without degrading linguistic modeling. Third, signal level quality improves in most cases with signal to noise ratio increasing by as much as 34 percent. Crucially these improvements are strongly governed by the characteristics of the training data. Speakers with high variability in acoustic energy and perceptual quality achieve simultaneous gains in DNS-MOS voice similarity and SNR. Overall this work establishes that LoRA finetuning is not merely a parameter efficient optimization technique but an effective mechanism for better speaker level adaptation in compact LLM-based TTS systems. When supported by sufficiently diverse training data LoRA adapted Qwen-0.5B consistently surpasses its frozen base model in perceptual quality speaker similarity with low latency using GGUF model hosted in quantized form.
【4】VoxCare: Studying Natural Communication Behaviors of Hospital Caregivers through Wearable Sensing of Egocentric Audio
标题:VoxCare:通过对自我中心音频的可穿戴传感研究医院护理人员的自然沟通行为
链接:https://arxiv.org/abs/2603.10888
摘要:医疗保健专业人员在复杂、高风险的环境中工作,有效的沟通对医疗服务、团队协调和个人福祉至关重要。然而,在日常临床环境中的沟通活动仍然具有挑战性的测量,在人类行为研究中基本上未被探索。我们提出了VoxCare,一个可扩展的以自我为中心的可穿戴音频传感和计算系统,它可以在不存储原始音频的情况下,在现实世界中捕获医院专业人员的自然通信行为。VoxCare执行实时、设备上的声学特征提取,并应用语音基础模型引导的师生框架来识别前景语音活动。从这些功能中,VoxCare推导出通信频率,持续时间和声音唤醒的可解释行为测量。我们的分析揭示了临床医生如何,何时以及多久在不同的班次和工作单位进行沟通,并表明沟通活动反映了潜在的工作量和压力。通过对日常环境中的沟通模式进行持续评估,这项研究提供了数据驱动的方法来了解医疗服务提供者的行为,并最终改善医疗服务。
摘要:Healthcare professionals work in complex, high-stakes environments where effective communication is critical for care delivery, team coordination, and individual well-being. However, communication activity in everyday clinical settings remains challenging to measure and largely unexplored in human behavioral research. We present VoxCare, a scalable egocentric wearable audio sensing and computing system that captures natural communication behaviors of hospital professionals in real-world settings without storing raw audio. VoxCare performs real-time, on-device acoustic feature extraction and applies a speech foundation model-guided teacher-student framework to identify foreground speech activity. From these features, VoxCare derives interpretable behavioral measures of communication frequency, duration, and vocal arousal. Our analyses reveal how, when, and how often clinicians communicate across different shifts and working units, and suggest that communication activity reflects underlying workload and stress. By enabling continuous assessment of communication patterns in everyday contexts, this study provides data-driven approaches to understand the behaviors of healthcare providers and ultimately improve healthcare delivery.
【5】OSUM-Pangu: An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs
标题:OSUM-Pangu:基于Ascend NPU上OpenPangu的开源多维语音理解基础模型
链接:https://arxiv.org/abs/2603.10862
备注:5 pages, 2 figures
摘要:语音大语言模型的最新进展显着增强了多维语音理解。然而,大多数高性能框架主要针对以GPU为中心的生态系统和专有骨干进行了优化,这为非CUDA计算基础设施的部署带来了巨大的差距。在本文中,我们介绍了OSUM-Pangu,这是一个完全开源的语音理解基础模型,在完全非CUDA的软件和硬件堆栈上开发。通过将音频编码器与openPangu-7 B LLM主干集成,我们成功地在Ascend NPU平台上实现了整个训练和推理管道。为了在非CUDA资源限制下促进有效的任务对齐,我们采用了一个实际的训练过程,该过程顺序地桥接语音感知和用户意图识别。实验结果表明,OSUM-Pangu在保持强大的自然语言交互能力的同时,实现了与主流基于GPU的模型相当的任务准确性。我们的工作为开源语音社区提供了一个可复制的非CUDA基线,促进了多模态智能的独立发展。
摘要:Recent advancements in Speech Large Language Models have significantly enhanced multi-dimensional speech understanding. However, the majority of high-performance frameworks are predominantly optimized for GPU centric ecosystems and proprietary backbones, creating a significant gap for deployment on non-CUDA computing infrastructures. In this paper, we present OSUM-Pangu, a fully open-source speech understanding foundation model developed on a completely non-CUDA software and hardware stack. By integrating an audio encoder with the openPangu-7B LLM backbone, we successfully implement the entire training and inference pipeline on the Ascend NPU platform. To facilitate efficient task alignment under non-CUDA resource constraints, we adopt a practical training process that sequentially bridges speech perception and user intent recognition. Experimental results demonstrate that OSUM-Pangu achieves task accuracy comparable to mainstream GPU-based models while maintaining robust natural language interaction capabilities. Our work provides a reproducible, non-CUDA baseline for the open-source speech community, promoting the independent evolution of multimodal intelligence.
【6】Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
标题:使用语音感知LLM进行说话者验证:评估和增强
链接:https://arxiv.org/abs/2603.10827
备注:3 Tables, 1 Figure, Under review
摘要:语音感知的大型语言模型(LLM)可以接受语音输入,但它们的训练目标主要强调语言内容或特定领域,如情感或说话者的性别,因此不清楚它们是否编码说话者身份。首先,我们提出了一个模型不可知的评分协议,使用来自是/否令牌概率的置信度分数或对数似然比,为仅API和开放权重模型生成连续的验证分数。使用该协议,我们对最近的语音感知LLM进行了基准测试,并观察到弱说话者区分(VoxCeleb 1上的EER高于20%)。其次,我们引入了一个轻量级的增强,通过学习投影和仅训练LoRA适配器注入冻结的ECAPA-TDNN扬声器嵌入,为LLM配备ASV功能。在TinyLLaMA-1.1B上,ECAPA-LLM在VoxCeleb 1-E上实现了1.03%的EER,接近专用的说话人验证系统,同时保留了自然语言界面。
摘要:Speech-aware large language models (LLMs) can accept speech inputs, yet their training objectives largely emphasize linguistic content or specific fields such as emotions or the speaker's gender, leaving it unclear whether they encode speaker identity. First, we propose a model-agnostic scoring protocol that produces continuous verification scores for both API-only and open-weight models, using confidence scores or log-likelihood ratios from the Yes/No token probabilities. Using this protocol, we benchmark recent speech-aware LLMs and observe weak speaker discrimination (EERs above 20% on VoxCeleb1). Second, we introduce a lightweight augmentation that equips an LLM with ASV capability by injecting frozen ECAPA-TDNN speaker embeddings through a learned projection and training only LoRA adapters. On TinyLLaMA-1.1B, the resulting ECAPA-LLM achieves 1.03% EER on VoxCeleb1-E, approaching a dedicated speaker verification system while preserving a natural-language interface.
【7】Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
标题:通过人类推理实现稳健的语音深度伪造检测
链接:https://arxiv.org/abs/2603.10725
摘要:现代生成音频模型可以被对手以非法的方式使用,特别是模仿其他人来获取私人信息。为了缓解这个问题,语音深度伪造检测(SDD)方法开始发展。不幸的是,目前的SDD方法通常缺乏对新的音频域和生成器的推广。更重要的是,它们缺乏可解释性,特别是类似人类的推理,可以自然地解释给定音频的真实或欺骗类的属性,并提供人类可感知的线索。在本文中,我们提出了HIR-SDD,这是一种新的SDD框架,它结合了大型音频语言模型(LALM)的优势和来自新提出的人类注释数据集的思维链推理。实验评估表明,所提出的方法的有效性和它的能力,提供合理的理由预测。
摘要:The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information. To mitigate this issue, speech deepfake detection (SDD) methods started to evolve. Unfortunately, current SDD methods generally suffer from the lack of generalization to new audio domains and generators. More than that, they lack interpretability, especially human-like reasoning that would naturally explain the attribution of a given audio to the bona fide or spoof class and provide human-perceptible cues. In this paper, we propose HIR-SDD, a novel SDD framework that combines the strengths of Large Audio Language Models (LALMs) with the chain-of-thought reasoning derived from the novel proposed human-annotated dataset. Experimental evaluation demonstrates both the effectiveness of the proposed method and its ability to provide reasonable justifications for predictions.
【8】Probabilistic Verification of Voice Anti-Spoofing Models
标题:语音反欺骗模型的概率验证
链接:https://arxiv.org/abs/2603.10713
摘要:生成模型的最新进展放大了恶意滥用语音合成技术的风险,使对手能够模仿目标说话者并访问敏感资源。虽然语音深度伪造检测进展迅速,但大多数现有对策缺乏正式的鲁棒性保证,或未能推广到不可见的生成技术。我们提出了PV-VASM,一个概率框架,用于验证语音反欺骗模型(VASMs)的鲁棒性。PV-VASM估计文本到语音(TTS)、语音克隆(VC)和参数信号转换下的误分类概率。该方法是模型不可知的,并使鲁棒性验证对看不见的语音合成技术和输入扰动。我们推导出一个理论上的错误概率上限,并验证了该方法在不同的实验设置,证明其有效性作为一个实用的鲁棒性验证工具。
摘要:Recent advances in generative models have amplified the risk of malicious misuse of speech synthesis technologies, enabling adversaries to impersonate target speakers and access sensitive resources. Although speech deepfake detection has progressed rapidly, most existing countermeasures lack formal robustness guarantees or fail to generalize to unseen generation techniques. We propose PV-VASM, a probabilistic framework for verifying the robustness of voice anti-spoofing models (VASMs). PV-VASM estimates the probability of misclassification under text-to-speech (TTS), voice cloning (VC), and parametric signal transformations. The approach is model-agnostic and enables robustness verification against unseen speech synthesis techniques and input perturbations. We derive a theoretical upper bound on the error probability and validate the method across diverse experimental settings, demonstrating its effectiveness as a practical robustness verification tool.
【9】AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow
标题:AlphaFlowPSE:通过条件AlphaFlow一步生成目标说话人提取
链接:https://arxiv.org/abs/2603.10701
备注:Submitted to Interspeech 2026 for review
摘要:在目标说话人提取(TSE)中,我们的目标是从多说话人的混合使用短注册话语作为参考恢复目标语音。最近的研究扩散和流匹配发生器提高了目标语音保真度。然而,多步采样会增加延迟,而一步解决方案通常依赖于混合依赖的时间坐标,这对于现实世界的对话来说可能是不可靠的。我们提出了AlphaFlowTSE,这是一个使用无雅可比向量积(JVP)AlphaFlow目标训练的一步条件生成模型。AlphaFlowTSE从观察到的混合物开始沿着混合物到目标轨迹学习平均速度传输,消除辅助混合比预测,并通过将流量匹配与间隔一致性教师-学生目标相结合来稳定训练。Libri 2 Mix和REAL-T上的实验证实,AlphaFlowTSE提高了目标说话人相似度和真实混合物的泛化能力,用于下游自动语音识别(ASR)。
摘要:In target speaker extraction (TSE), we aim to recover target speech from a multi-talker mixture using a short enrollment utterance as reference. Recent studies on diffusion and flow-matching generators have improved target-speech fidelity. However, multi-step sampling increases latency, and one-step solutions often rely on a mixture-dependent time coordinate that can be unreliable for real-world conversations. We present AlphaFlowTSE, a one-step conditional generative model trained with a Jacobian-vector product (JVP)-free AlphaFlow objective. AlphaFlowTSE learns mean-velocity transport along a mixture-to-target trajectory starting from the observed mixture, eliminating auxiliary mixing-ratio prediction, and stabilizes training by combining flow matching with an interval-consistency teacher-student target. Experiments on Libri2Mix and REAL-T confirm that AlphaFlowTSE improves target-speaker similarity and real-mixture generalization for downstream automatic speech recognition (ASR).
【10】Distilling LLM Semantic Priors into Encoder-Only Multi-Talker ASR with Talker-Count Routing
标题:将LLM语义先验提炼为具有说话者计数路由的仅编码器多说话者ASB
链接:https://arxiv.org/abs/2603.10587
摘要:大型语言模型(LLM)提供了强大的语义先验,可以改善多说话者自动语音识别(MT-ASR),但使用LLM作为自回归解码器在计算上是昂贵的,并且在严重重叠的情况下仍然很脆弱。在本文中,我们提出了一个仅编码器的MT-ASR框架,该框架使LLM适应多说话者条件反射,并在训练期间将其语义指导提取到编码器中,同时在推理时保留快速CTC风格的解码。我们的模型采用了一个编码器后分离器与序列化的CTC,以产生说话者排序的成绩单,并利用一个适应的基于LLM的SOT目标作为多说话者感知的教师信号,以显式地正则化混合语音表示。为了进一步支持可变数量的扬声器,我们引入了一个扬声器计数头,预测扬声器计数和动态选择适当的解码分支。LibriMix上的实验表明,所提出的仅编码器模型在两个说话者条件下实现了与基于LLM的系统相当的性能,同时在具有显著小RTF的三个说话者条件下实现了显著的改进。
摘要:Large language models (LLMs) provide strong semantic priors that can improve multi-talker automatic speech recognition (MT-ASR), but using an LLM as an autoregressive decoder is computationally expensive and remains fragile under heavy overlap. In this paper, we propose an encoder-only MT-ASR framework that adapts an LLM to multi-talker conditioning and distills its semantic guidance into the encoder during training, while retaining fast CTC-style decoding at inference. Our model employs a post-encoder separator with serialized CTC to produce talker-ordered transcripts, and leverages an adapted LLM-based SOT objective as a multi-talker-aware teacher signal to explicitly regularize mixed-speech representations. To further support variable numbers of talkers, we introduce a Talker-Count Head that predicts the talker count and dynamically selects the appropriate decoding branch. Experiments on LibriMix show that the proposed encoder-only model achieves comparable performance to LLM-based systems in the two-talker condition, while delivering significant improvements in the three-talker condition with significant small RTF.
【11】MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
标题:MoXaRT:XR的视听对象引导声音交互
链接:https://arxiv.org/abs/2603.10465
摘要:在延展实境(XR)中,复杂的声学环境常常使用户不知所措,由于纠缠的声源而损害场景感知和社交参与。我们介绍MoXaRt,这是一个实时XR系统,它使用视听线索来分离这些来源,并实现细粒度的声音交互。MoXaRt的核心是级联架构,其与源的视觉检测并行地执行粗略的仅音频分离(例如,面孔,工具)。然后,这些视觉锚点引导细化网络隔离各个源,分离多达5个并发源的复杂混合(例如,2个声音+3个乐器),处理延迟约为2秒。我们通过对30个一分钟录音的新数据集进行技术评估来验证MoXaRt,这些录音具有并发语音和音乐,以及22名参与者的用户研究。实验结果表明,我们的系统显著提高了语音清晰度,在对抗性声学环境中的听力理解能力提高了36.2%(p < 0.01),同时大大降低了认知负荷(p < 0.001),从而为更具感知力和社交能力的XR体验铺平了道路。
摘要:In Extended Reality (XR), complex acoustic environments often overwhelm users, compromising both scene awareness and social engagement due to entangled sound sources. We introduce MoXaRt, a real-time XR system that uses audio-visual cues to separate these sources and enable fine-grained sound interaction. MoXaRt's core is a cascaded architecture that performs coarse, audio-only separation in parallel with visual detection of sources (e.g., faces, instruments). These visual anchors then guide refinement networks to isolate individual sources, separating complex mixes of up to 5 concurrent sources (e.g., 2 voices + 3 instruments) with ~2 second processing latency. We validate MoXaRt through a technical evaluation on a new dataset of 30 one-minute recordings featuring concurrent speech and music, and a 22-participant user study. Empirical results indicate that our system significantly enhances speech intelligibility, yielding a 36.2% (p < 0.01) increase in listening comprehension within adversarial acoustic environments while substantially reducing cognitive load (p < 0.001), thereby paving the way for more perceptive and socially adept XR experiences.
【12】NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
标题:NasoVoce:一种鼻式低可听语音接口,用于始终可用的语音交互
链接:https://arxiv.org/abs/2603.10324
备注:ACM CHI 2026 paper
摘要:无声和低声的语音为始终可用的人工智能语音交互提供了希望,但现有的方法很难平衡词汇量、可穿戴性、无声和噪声鲁棒性。我们提出NasoVoce,鼻桥安装接口,集成了麦克风和振动传感器。它位于智能眼镜的鼻垫处,可以不引人注目地捕获声学和振动信号。靠近嘴部的鼻梁允许访问骨骼和皮肤传导的语音,并且能够可靠地捕获诸如低声语音的低音量话语。虽然麦克风可以捕捉高质量的音频,但它对环境噪音非常敏感。相反,振动传感器对噪声是鲁棒的,但产生较低的信号质量。通过融合这些互补的输入,NasoVoce可以生成抗干扰的高质量语音。评估与耳语大v2,PESQ,STOI,和MUSHRA评级确认提高识别和质量. NasoVoce展示了一个实用界面的可行性,用于始终可用,连续和谨慎的人工智能语音对话。
摘要:Silent and whispered speech offer promise for always-available voice interaction with AI, yet existing methods struggle to balance vocabulary size, wearability, silence, and noise robustness. We present NasoVoce, a nose-bridge-mounted interface that integrates a microphone and a vibration sensor. Positioned at the nasal pads of smart glasses, it unobtrusively captures both acoustic and vibration signals. The nasal bridge, close to the mouth, allows access to bone- and skin-conducted speech and enables reliable capture of low-volume utterances such as whispered speech. While the microphone captures high-quality audio, it is highly sensitive to environmental noise. Conversely, the vibration sensor is robust to noise but yields lower signal quality. By fusing these complementary inputs, NasoVoce generates high-quality speech robust against interference. Evaluation with Whisper Large-v2, PESQ, STOI, and MUSHRA ratings confirms improved recognition and quality. NasoVoce demonstrates the feasibility of a practical interface for always-available, continuous, and discreet AI voice conversations.
【13】PRoADS: Provably Secure and Robust Audio Diffusion Steganography with latent optimization and backward Euler Inversion
标题:PRoADS:具有潜在优化和向后欧拉倒置的可证明安全且鲁棒的音频扩散隐写术
链接:https://arxiv.org/abs/2603.10314
备注:This paper has been accepted for presentation at the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)
摘要:提出了一种基于音频扩散模型的可证明安全的鲁棒音频隐写框架PRoADS。作为一种生成式隐写方案,PRoADS通过正交矩阵投影将秘密信息嵌入到扩散模型的初始噪声中。为了解决扩散反演中导致高误码率(BER)的重建误差,我们引入了潜在优化和向后欧拉反演,以最大限度地减少潜在的重建和扩散反演误差。实验结果表明,在64 kbps的MP3压缩下,该方案的误码率为0.15%,明显优于现有方法,具有较强的鲁棒性。
摘要:This paper proposes PRoADS, a provably secure and robust audio steganographic framework based on audio diffusion models. As a generative steganography scheme, PRoADS embeds secret messages into the initial noise of diffusion models via orthogonal matrix projection. To address the reconstruction errors in diffusion inversion that cause high bit error rates (BER), we introduce Latent Optimization and Backward Euler Inversion to minimize the latent reconstruction and diffusion inversion errors. Comprehensive experiments demonstrate that our scheme sustains a remarkably low BER of 0.15\% under 64 kbps MP3 compression, significantly outperforming existing methods and exhibiting strong robustness.
【14】ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA
标题:ID-LoRA:通过上下文LoRA实现身份驱动的音频视频个性化
链接:https://arxiv.org/abs/2603.10256
摘要:现有的视频个性化方法保持视觉相似性,但将视频和音频分开处理。如果不能访问视觉场景,音频模型就无法将声音与屏幕上的动作同步;由于经典的语音克隆模型只以参考录音为条件,文本提示无法重定向说话风格或声学环境。我们提出了ID-LoRA(身份驱动的上下文LoRA),它在单个模型中联合生成主体的外观和声音,让文本提示,参考图像和简短的音频剪辑共同管理这两种模式。ID-LoRA通过参数高效的In-Context LoRA适应LTX-2联合音频-视频扩散骨干,据我们所知,这是第一种在单个生成通道中个性化视觉外观和语音的方法。出现了两个挑战。参考和生成标记共享相同的位置编码空间,使得它们难以区分;我们用负时间位置来解决这个问题,将参考标记放置在不相交的RoPE区域中,同时保留其内部时间结构。说话者的特征也往往在去噪过程中被稀释;我们引入了身份指导,一种无分类器的指导变体,通过对比有和没有参考信号的预测来放大说话者特定的特征。在人类偏好研究中,73%的注释者更喜欢ID-LoRA而不是Kling 2.6 Pro,因为语音相似性和65%的说话风格。在跨环境设置中,说话人相似性比Kling提高了24%,随着条件的不同,差距会扩大。初步的用户研究进一步表明,联合发电提供了一个有用的电感偏置物理接地的声音合成。ID-LoRA在单个GPU上仅使用约3 K个训练对即可实现这些结果。代码、模型和数据将被发布。
摘要:Existing video personalization methods preserve visual likeness but treat video and audio separately. Without access to the visual scene, audio models cannot synchronize sounds with on-screen actions; and because classical voice-cloning models condition only on a reference recording, a text prompt cannot redirect speaking style or acoustic environment. We propose ID-LoRA (Identity-Driven In-Context LoRA), which jointly generates a subject's appearance and voice in a single model, letting a text prompt, a reference image, and a short audio clip govern both modalities together. ID-LoRA adapts the LTX-2 joint audio-video diffusion backbone via parameter-efficient In-Context LoRA and, to our knowledge, is the first method to personalize visual appearance and voice in a single generative pass. Two challenges arise. Reference and generation tokens share the same positional-encoding space, making them hard to distinguish; we address this with negative temporal positions, placing reference tokens in a disjoint RoPE region while preserving their internal temporal structure. Speaker characteristics also tend to be diluted during denoising; we introduce identity guidance, a classifier-free guidance variant that amplifies speaker-specific features by contrasting predictions with and without the reference signal. In human preference studies, ID-LoRA is preferred over Kling 2.6 Pro by 73% of annotators for voice similarity and 65% for speaking style. On cross-environment settings, speaker similarity improves by 24% over Kling, with the gap widening as conditions diverge. A preliminary user study further suggests that joint generation provides a useful inductive bias for physically grounded sound synthesis. ID-LoRA achieves these results with only ~3K training pairs on a single GPU. Code, models, and data will be released.
【15】nlm: Real-Time Non-linear Modal Synthesis in Max
标题:nlm:Max中的实时非线性模式合成
链接:https://arxiv.org/abs/2603.10240
备注:accepted to PdMaxCon25~ (https://music.illinois.edu/pd-max-con/)
摘要:我们提出了\texttt{nlm},一组最大外部,使有效的实时非线性模态合成字符串,膜和板。外部,在C++实现,提供物理参数的交互式控制,允许自定义模态数据的加载,并提供多通道输出。通过将交互式物理建模功能集成到熟悉的环境中,降低了作曲家、表演者和声音设计师探索非线性模态合成的表达潜力的障碍。这些外部软件可以在https://github.com/rodrigodzf/nlm上以开源软件的形式获得。
摘要:We present \texttt{nlm}, a set of Max externals that enable efficient real-time non-linear modal synthesis for strings, membranes, and plates. The externals, implemented in C++, offer interactive control of physical parameters, allow the loading of custom modal data, and provide multichannel output. By integrating interactive physical-modelling capabilities into a familiar environment, \texttt{nlm} lowers the barrier for composers, performers, and sound designers to explore the expressive potential of non-linear modal synthesis. The externals are available as open-source software at https://github.com/rodrigodzf/nlm.
【16】AMB-DSGDN: Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network for Multimodal Emotion Recognition
标题:AMB-DSGDN:用于多模式情感识别的自适应模式平衡动态语义图差异网络
链接:https://arxiv.org/abs/2603.10043
备注:18 pages
摘要:多模态对话情感识别通过融合文本、视觉和音频模态来捕获情感线索。然而,现有的方法仍然遭受显着的局限性,在建模情感依赖和学习多模态表示。一方面,它们无法有效地过滤掉多模态特征中的冗余或噪声信号,这阻碍了准确捕捉说话者之间和说话者内部的情绪状态的动态演变。另一方面,在多模态特征学习过程中,主导模态往往会压倒融合过程,从而抑制了语音和视觉等非主导模态的互补贡献,最终限制了整体识别性能。为了应对这些挑战,我们提出了一个自适应模态平衡动态语义图差分网络(AMB-DSGDN)。具体地说,我们首先为文本,语音和视觉构建特定于模态的子图,其中每个模态包含说话者内和说话者间的图,以捕获自我连续性和跨说话者的情感依赖。在这些子图之上,我们引入了差分图注意力机制,该机制计算两组注意力图之间的差异。通过明确地对比这些注意力分布,该机制消除了共享的噪声模式,同时保留了特定于模态和上下文相关的信号,从而产生更纯粹和更具区分力的情感表征。此外,我们还设计了一个自适应的模态平衡机制,它根据每个模态在情感建模中的相对贡献来估计其丢失概率。
摘要:Multimodal dialogue emotion recognition captures emotional cues by fusing text, visual, and audio modalities. However, existing approaches still suffer from notable limitations in modeling emotional dependencies and learning multimodal representations. On the one hand, they are unable to effectively filter out redundant or noisy signals within multimodal features, which hinders the accurate capture of the dynamic evolution of emotional states across and within speakers. On the other hand, during multimodal feature learning, dominant modalities tend to overwhelm the fusion process, thereby suppressing the complementary contributions of non-dominant modalities such as speech and vision, ultimately constraining the overall recognition performance. To address these challenges, we propose an Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network (AMB-DSGDN). Concretely, we first construct modality-specific subgraphs for text, speech, and vision, where each modality contains intra-speaker and inter-speaker graphs to capture both self-continuity and cross-speaker emotional dependencies. On top of these subgraphs, we introduce a differential graph attention mechanism, which computes the discrepancy between two sets of attention maps. By explicitly contrasting these attention distributions, the mechanism cancels out shared noise patterns while retaining modality-specific and context-relevant signals, thereby yielding purer and more discriminative emotional representations. In addition, we design an adaptive modality balancing mechanism, which estimates a dropout probability for each modality according to its relative contribution in emotion modeling.
【17】Geo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context
标题:Geo-ATBench:一个基于地理语义上下文的地理空间音频标注基准
链接:https://arxiv.org/abs/2603.10623
摘要:在计算听觉场景分析(CASA)中,环境声音理解通常被表述为一个纯音频识别问题。这个公式在多标签音频标记(AT)中留下了一个持久的缺点:声学相似性可能使某些事件难以单独从波形中分离出来。在这种情况下,消除歧义的线索往往位于波形之外。从地理信息系统数据导出的地理空间语义上下文(GSC),例如,兴趣点(POI)提供了与位置相关的环境先验,可以帮助减少这种模糊性。通过建议的地理空间音频标记(Geo-AT)任务,使这个方向的一个系统的研究,其中的条件下,多标签的声音事件标记GSC旁边的音频。为了对Geo-AT进行基准测试,Geo-ATBench被引入作为具有地理注释的复调音频基准,包含28个事件类别的10.71小时音频;每个片段与来自11个语义上下文类别的GSC表示配对。GeoFusion-AT被提出作为一个统一的地理音频融合框架,评估特征,表示和决策级融合的代表性音频骨干,音频和GSC的基线。结果表明,将GSC提高AT的性能,特别是在声学混淆的标签,表明地理空间语义提供有效的先验超越音频单独。一项由10名参与者对579个样本进行的众包听力研究表明,Geo-ATBench标签和聚合人类标签上的模型之间的性能没有显着差异,支持Geo-ATBench作为人类对齐的基准。Geo-AT任务、基准Geo-ATBench和可再现的地理音频融合框架GeoFusion-AT为CASA社区内研究具有地理空间语义背景的AT提供了基础。数据集,代码,模型都在主页上(https://github.com/WuYanru2002/Geo-ATBench)。
摘要:Environmental sound understanding in computational auditory scene analysis (CASA) is often formulated as an audio-only recognition problem. This formulation leaves a persistent drawback in multi-label audio tagging (AT): acoustic similarity can make certain events difficult to separate from waveforms alone. In such cases, disambiguating cues often lie outside the waveform. Geospatial semantic context (GSC), derived from geographic information system data, e.g., points of interest (POI), provides location-tied environmental priors that can help reduce this ambiguity. A systematic study of this direction is enabled through the proposed geospatial audio tagging (Geo-AT) task, which conditions multi-label sound event tagging on GSC alongside audio. To benchmark Geo-AT, Geo-ATBench is introduced as a polyphonic audio benchmark with geographical annotations, containing 10.71 hours of audio across 28 event categories; each clip is paired with a GSC representation from 11 semantic context categories. GeoFusion-AT is proposed as a unified geo-audio fusion framework that evaluates feature-, representation-, and decision-level fusion on representative audio backbones, with audio- and GSC-only baselines. Results show that incorporating GSC improves AT performance, especially on acoustically confounded labels, indicating geospatial semantics provide effective priors beyond audio alone. A crowdsourced listening study with 10 participants on 579 samples shows that there is no significant difference in performance between models on Geo-ATBench labels and aggregated human labels, supporting Geo-ATBench as a human-aligned benchmark. The Geo-AT task, benchmark Geo-ATBench, and reproducible geo-audio fusion framework GeoFusion-AT provide a foundation for studying AT with geospatial semantic context within the CASA community. Dataset, code, models are on homepage (https://github.com/WuYanru2002/Geo-ATBench).
【18】G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
标题:G-STAR:端到端全球说话者跟踪归因识别
链接:https://arxiv.org/abs/2603.10468
备注:submitted to Interspeech 2026
摘要:我们研究了时间戳的扬声器归因ASR的长形式,多方语音重叠,其中块明智的推理必须保持会议级扬声器身份的一致性,同时产生时间戳,扬声器标记的成绩单。以前的Speech-LLM系统倾向于优先考虑局部日志化或全局标记,但通常缺乏捕获细粒度时间边界或鲁棒的跨块身份链接的能力。我们提出了G-STAR,一个端到端的系统,耦合一个时间感知说话人跟踪模块与Speech-LLM转录骨干。该跟踪器提供结构化的扬声器提示与时间接地,和LLM生成归因于这些提示的条件下的文本。G-STAR支持组件优化和联合端到端培训,从而在异构监督和领域转移下实现灵活的学习。实验分析线索融合,本地与长期的上下文权衡和分层目标。
摘要:We study timestamped speaker-attributed ASR for long-form, multi-party speech with overlap, where chunk-wise inference must preserve meeting-level speaker identity consistency while producing time-stamped, speaker-labeled transcripts. Previous Speech-LLM systems tend to prioritize either local diarization or global labeling, but often lack the ability to capture fine-grained temporal boundaries or robust cross-chunk identity linking. We propose G-STAR, an end-to-end system that couples a time-aware speaker-tracking module with a Speech-LLM transcription backbone. The tracker provides structured speaker cues with temporal grounding, and the LLM generates attributed text conditioned on these cues. G-STAR supports both component-wise optimization and joint end-to-end training, enabling flexible learning under heterogeneous supervision and domain shift. Experiments analyze cue fusion, local versus long-context trade-offs and hierarchical objectives.
【19】FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System
标题:FireRedASB 2S:最先进的工业级一体化自动语音识别系统
链接:https://arxiv.org/abs/2603.10420
摘要:我们介绍FireRedASR 2S,一个最先进的工业级一体化自动语音识别(ASR)系统。它在一个统一的管道中集成了四个模块:ASR,语音活动检测(VAD),口语识别(LID)和标点预测(Punc)。FireRedASR 2:一个ASR模块,有两个变体,FireRedASR 2-LLM(8B+参数)和FireRedASR 2-AED(1B+参数),支持普通话,中国方言和口音,英语和代码切换的语音和唱歌转录。与FireRedASR相比,FireRedASR 2提供了更高的识别准确性和更广泛的方言和口音覆盖范围。FireRedASR 2-LLM在4个公共普通话基准测试中的平均CER为2.89%,在19个公共汉语方言和口音基准测试中的平均CER为11.55%,优于竞争对手的基准测试,包括Doubao-ASR,Qwen 3-ASR和Fun-ASR。FireRedVAD:基于深度前馈顺序存储网络(DFSMN)的超轻量模块(0.6M参数),支持流式VAD、非流式VAD和多标签VAD(mVAD)。在FLEURS-VAD-102基准测试中,它实现了97.57%的帧级F1和99.60%的AUC-ROC,优于Silero-VAD,TEN-VAD,FunASR-VAD和WebRTC-VAD。FireRedLID:一个编码器-解码器LID模块,支持100多种语言和20多种中国方言和口音。在FLEURS(82种语言)上,它达到了97.18%的话语级准确率,优于Whisper和SpeechBrain。FireRedPunc:一个BERT风格的中文和英文标点预测模块。在多域基准测试中,它达到了78.90%的平均F1,优于FunASR-Punc(62.77%)。为了推进语音处理的研究,我们在https://github.com/FireRedTeam/FireRedASR2S上发布了模型权重和代码。
摘要:We present FireRedASR2S, a state-of-the-art industrial-grade all-in-one automatic speech recognition (ASR) system. It integrates four modules in a unified pipeline: ASR, Voice Activity Detection (VAD), Spoken Language Identification (LID), and Punctuation Prediction (Punc). All modules achieve SOTA performance on the evaluated benchmarks: FireRedASR2: An ASR module with two variants, FireRedASR2-LLM (8B+ parameters) and FireRedASR2-AED (1B+ parameters), supporting speech and singing transcription for Mandarin, Chinese dialects and accents, English, and code-switching. Compared to FireRedASR, FireRedASR2 delivers improved recognition accuracy and broader dialect and accent coverage. FireRedASR2-LLM achieves 2.89% average CER on 4 public Mandarin benchmarks and 11.55% on 19 public Chinese dialects and accents benchmarks, outperforming competitive baselines including Doubao-ASR, Qwen3-ASR, and Fun-ASR. FireRedVAD: An ultra-lightweight module (0.6M parameters) based on the Deep Feedforward Sequential Memory Network (DFSMN), supporting streaming VAD, non-streaming VAD, and multi-label VAD (mVAD). On the FLEURS-VAD-102 benchmark, it achieves 97.57% frame-level F1 and 99.60% AUC-ROC, outperforming Silero-VAD, TEN-VAD, FunASR-VAD, and WebRTC-VAD. FireRedLID: An Encoder-Decoder LID module supporting 100+ languages and 20+ Chinese dialects and accents. On FLEURS (82 languages), it achieves 97.18% utterance-level accuracy, outperforming Whisper and SpeechBrain. FireRedPunc: A BERT-style punctuation prediction module for Chinese and English. On multi-domain benchmarks, it achieves 78.90% average F1, outperforming FunASR-Punc (62.77%). To advance research in speech processing, we release model weights and code at https://github.com/FireRedTeam/FireRedASR2S.
【1】MOS-Bias: From Hidden Gender Bias to Gender-Aware Speech Quality Assessment
标题:MOS偏见:从隐藏的性别偏见到性别意识的言语质量评估
链接:https://arxiv.org/abs/2603.10723
备注:Submitted to Interspeech 2026
摘要:平均意见得分(MOS)作为语音质量评估的标准度量,但人类注释中的偏见仍然没有得到充分研究。我们对MOS中的性别偏见进行了首次系统分析,揭示了男性听众始终比女性听众给予更高的分数--这一差距在低质量语音中最为明显,并随着质量的提高而逐渐缩小。这种依赖质量的结构证明难以通过简单的校准来消除。我们进一步证明,在聚合标签上训练的自动MOS模型显示出偏向男性感知标准的预测。为了解决这个问题,我们提出了一个性别感知模型,通过抽象二进制组嵌入来学习特定性别的评分模式,从而提高整体和特定性别的预测准确性。这项研究建立了MOS中的性别偏见构成了一个系统的,可学习的模式,要求在公平的言论评价的关注。
摘要:The Mean Opinion Score (MOS) serves as the standard metric for speech quality assessment, yet biases in human annotations remain underexplored. We conduct the first systematic analysis of gender bias in MOS, revealing that male listeners consistently assign higher scores than female listeners--a gap that is most pronounced in low-quality speech and gradually diminishes as quality improves. This quality-dependent structure proves difficult to eliminate through simple calibration. We further demonstrate that automated MOS models trained on aggregated labels exhibit predictions skewed toward male standards of perception. To address this, we propose a gender-aware model that learns gender-specific scoring patterns through abstracting binary group embeddings, thereby improving overall and gender-specific prediction accuracy. This study establishes that gender bias in MOS constitutes a systematic, learnable pattern demanding attention in equitable speech evaluation.
【2】Geo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context
标题:Geo-ATBench:一个基于地理语义上下文的地理空间音频标注基准
链接:https://arxiv.org/abs/2603.10623
摘要:在计算听觉场景分析(CASA)中,环境声音理解通常被表述为一个纯音频识别问题。这个公式在多标签音频标记(AT)中留下了一个持久的缺点:声学相似性可能使某些事件难以单独从波形中分离出来。在这种情况下,消除歧义的线索往往位于波形之外。从地理信息系统数据导出的地理空间语义上下文(GSC),例如,兴趣点(POI)提供了与位置相关的环境先验,可以帮助减少这种模糊性。通过建议的地理空间音频标记(Geo-AT)任务,使这个方向的一个系统的研究,其中的条件下,多标签的声音事件标记GSC旁边的音频。为了对Geo-AT进行基准测试,Geo-ATBench被引入作为具有地理注释的复调音频基准,包含28个事件类别的10.71小时音频;每个片段与来自11个语义上下文类别的GSC表示配对。GeoFusion-AT被提出作为一个统一的地理音频融合框架,评估特征,表示和决策级融合的代表性音频骨干,音频和GSC的基线。结果表明,将GSC提高AT的性能,特别是在声学混淆的标签,表明地理空间语义提供有效的先验超越音频单独。一项由10名参与者对579个样本进行的众包听力研究表明,Geo-ATBench标签和聚合人类标签上的模型之间的性能没有显着差异,支持Geo-ATBench作为人类对齐的基准。Geo-AT任务、基准Geo-ATBench和可再现的地理音频融合框架GeoFusion-AT为CASA社区内研究具有地理空间语义背景的AT提供了基础。数据集,代码,模型都在主页上(https://github.com/WuYanru2002/Geo-ATBench)。
摘要:Environmental sound understanding in computational auditory scene analysis (CASA) is often formulated as an audio-only recognition problem. This formulation leaves a persistent drawback in multi-label audio tagging (AT): acoustic similarity can make certain events difficult to separate from waveforms alone. In such cases, disambiguating cues often lie outside the waveform. Geospatial semantic context (GSC), derived from geographic information system data, e.g., points of interest (POI), provides location-tied environmental priors that can help reduce this ambiguity. A systematic study of this direction is enabled through the proposed geospatial audio tagging (Geo-AT) task, which conditions multi-label sound event tagging on GSC alongside audio. To benchmark Geo-AT, Geo-ATBench is introduced as a polyphonic audio benchmark with geographical annotations, containing 10.71 hours of audio across 28 event categories; each clip is paired with a GSC representation from 11 semantic context categories. GeoFusion-AT is proposed as a unified geo-audio fusion framework that evaluates feature-, representation-, and decision-level fusion on representative audio backbones, with audio- and GSC-only baselines. Results show that incorporating GSC improves AT performance, especially on acoustically confounded labels, indicating geospatial semantics provide effective priors beyond audio alone. A crowdsourced listening study with 10 participants on 579 samples shows that there is no significant difference in performance between models on Geo-ATBench labels and aggregated human labels, supporting Geo-ATBench as a human-aligned benchmark. The Geo-AT task, benchmark Geo-ATBench, and reproducible geo-audio fusion framework GeoFusion-AT provide a foundation for studying AT with geospatial semantic context within the CASA community. Dataset, code, models are on homepage (https://github.com/WuYanru2002/Geo-ATBench).
【3】G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
标题:G-STAR:端到端全球说话者跟踪归因识别
链接:https://arxiv.org/abs/2603.10468
备注:submitted to Interspeech 2026
摘要:我们研究了时间戳的扬声器归因ASR的长形式,多方语音重叠,其中块明智的推理必须保持会议级扬声器身份的一致性,同时产生时间戳,扬声器标记的成绩单。以前的Speech-LLM系统倾向于优先考虑局部日志化或全局标记,但通常缺乏捕获细粒度时间边界或鲁棒的跨块身份链接的能力。我们提出了G-STAR,一个端到端的系统,耦合一个时间感知说话人跟踪模块与Speech-LLM转录骨干。该跟踪器提供结构化的扬声器提示与时间接地,和LLM生成归因于这些提示的条件下的文本。G-STAR支持组件优化和联合端到端培训,从而在异构监督和领域转移下实现灵活的学习。实验分析线索融合,本地与长期的上下文权衡和分层目标。
摘要:We study timestamped speaker-attributed ASR for long-form, multi-party speech with overlap, where chunk-wise inference must preserve meeting-level speaker identity consistency while producing time-stamped, speaker-labeled transcripts. Previous Speech-LLM systems tend to prioritize either local diarization or global labeling, but often lack the ability to capture fine-grained temporal boundaries or robust cross-chunk identity linking. We propose G-STAR, an end-to-end system that couples a time-aware speaker-tracking module with a Speech-LLM transcription backbone. The tracker provides structured speaker cues with temporal grounding, and the LLM generates attributed text conditioned on these cues. G-STAR supports both component-wise optimization and joint end-to-end training, enabling flexible learning under heterogeneous supervision and domain shift. Experiments analyze cue fusion, local versus long-context trade-offs and hierarchical objectives.
【4】FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System
标题:FireRedASB 2S:最先进的工业级一体化自动语音识别系统
链接:https://arxiv.org/abs/2603.10420
摘要:我们介绍FireRedASR 2S,一个最先进的工业级一体化自动语音识别(ASR)系统。它在一个统一的管道中集成了四个模块:ASR,语音活动检测(VAD),口语识别(LID)和标点预测(Punc)。FireRedASR 2:一个ASR模块,有两个变体,FireRedASR 2-LLM(8B+参数)和FireRedASR 2-AED(1B+参数),支持普通话,中国方言和口音,英语和代码切换的语音和唱歌转录。与FireRedASR相比,FireRedASR 2提供了更高的识别准确性和更广泛的方言和口音覆盖范围。FireRedASR 2-LLM在4个公共普通话基准测试中的平均CER为2.89%,在19个公共汉语方言和口音基准测试中的平均CER为11.55%,优于竞争对手的基准测试,包括Doubao-ASR,Qwen 3-ASR和Fun-ASR。FireRedVAD:基于深度前馈顺序存储网络(DFSMN)的超轻量模块(0.6M参数),支持流式VAD、非流式VAD和多标签VAD(mVAD)。在FLEURS-VAD-102基准测试中,它实现了97.57%的帧级F1和99.60%的AUC-ROC,优于Silero-VAD,TEN-VAD,FunASR-VAD和WebRTC-VAD。FireRedLID:一个编码器-解码器LID模块,支持100多种语言和20多种中国方言和口音。在FLEURS(82种语言)上,它达到了97.18%的话语级准确率,优于Whisper和SpeechBrain。FireRedPunc:一个BERT风格的中文和英文标点预测模块。在多域基准测试中,它达到了78.90%的平均F1,优于FunASR-Punc(62.77%)。为了推进语音处理的研究,我们在https://github.com/FireRedTeam/FireRedASR2S上发布了模型权重和代码。
摘要:We present FireRedASR2S, a state-of-the-art industrial-grade all-in-one automatic speech recognition (ASR) system. It integrates four modules in a unified pipeline: ASR, Voice Activity Detection (VAD), Spoken Language Identification (LID), and Punctuation Prediction (Punc). All modules achieve SOTA performance on the evaluated benchmarks: FireRedASR2: An ASR module with two variants, FireRedASR2-LLM (8B+ parameters) and FireRedASR2-AED (1B+ parameters), supporting speech and singing transcription for Mandarin, Chinese dialects and accents, English, and code-switching. Compared to FireRedASR, FireRedASR2 delivers improved recognition accuracy and broader dialect and accent coverage. FireRedASR2-LLM achieves 2.89% average CER on 4 public Mandarin benchmarks and 11.55% on 19 public Chinese dialects and accents benchmarks, outperforming competitive baselines including Doubao-ASR, Qwen3-ASR, and Fun-ASR. FireRedVAD: An ultra-lightweight module (0.6M parameters) based on the Deep Feedforward Sequential Memory Network (DFSMN), supporting streaming VAD, non-streaming VAD, and multi-label VAD (mVAD). On the FLEURS-VAD-102 benchmark, it achieves 97.57% frame-level F1 and 99.60% AUC-ROC, outperforming Silero-VAD, TEN-VAD, FunASR-VAD, and WebRTC-VAD. FireRedLID: An Encoder-Decoder LID module supporting 100+ languages and 20+ Chinese dialects and accents. On FLEURS (82 languages), it achieves 97.18% utterance-level accuracy, outperforming Whisper and SpeechBrain. FireRedPunc: A BERT-style punctuation prediction module for Chinese and English. On multi-domain benchmarks, it achieves 78.90% average F1, outperforming FunASR-Punc (62.77%). To advance research in speech processing, we release model weights and code at https://github.com/FireRedTeam/FireRedASR2S.
【5】Speech Codec Probing from Semantic and Phonetic Perspectives
标题:从语义和语音角度探索语音编解码器
链接:https://arxiv.org/abs/2603.10371
摘要:在多模态系统中,语音标记器对于将语音连接到大型语言模型(LLM)是必不可少的。这些标记器预计将保留语义和声学信息,以供下游理解和生成。然而,新出现的证据表明,什么是所谓的“语义”的语音表示不符合文本派生的语义:不匹配,可以降低多模态LLM性能。在本文中,我们系统地分析了几个广泛使用的语音标记器编码的信息,通过词级探测任务,分层表示分析和跨模态对齐度量(如CKA)来解开它们的语义和语音内容。我们的研究结果表明,目前的标记器主要捕捉语音,而不是词汇语义结构,我们得到的下一代语音标记方法的设计的实际影响。
摘要:Speech tokenizers are essential for connecting speech to large language models (LLMs) in multimodal systems. These tokenizers are expected to preserve both semantic and acoustic information for downstream understanding and generation. However, emerging evidence suggests that what is termed "semantic" in speech representations does not align with text-derived semantics: a mismatch that can degrade multimodal LLM performance. In this paper, we systematically analyze the information encoded by several widely used speech tokenizers, disentangling their semantic and phonetic content through word-level probing tasks, layerwise representation analysis, and cross-modal alignment metrics such as CKA. Our results show that current tokenizers primarily capture phonetic rather than lexical-semantic structure, and we derive practical implications for the design of next-generation speech tokenization methods.
【6】Calibration-Reasoning Framework for Descriptive Speech Quality Assessment
标题:描述性语音质量评估的校准推理框架
链接:https://arxiv.org/abs/2603.10175
备注:Submitted to Interspeech 2026
摘要:可解释的语音质量评估需要超越平均意见得分(MOS)来分析潜在的感知维度。为了解决这个问题,我们引入了一种新的后训练方法,该方法为音频伪影的多维推理,检测和分类定制了基础音频大语言模型。首先,校准阶段对齐模型以预测预定义的感知维度。其次,强化学习阶段利用具有维度特定奖励的组相对策略优化(GRPO)来大大提高描述的准确性和质量问题的时间定位。通过这种方法,我们达到了最先进的结果0.71平均PCC得分的多维语音基准和13%的改善MOS预测驱动的基于RL的推理。此外,我们的细粒度GRPO奖励大大提高了模型及时查明和分类音频伪影的能力。
摘要:Explainable speech quality assessment requires moving beyond Mean Opinion Scores (MOS) to analyze underlying perceptual dimensions. To address this, we introduce a novel post-training method that tailors the foundational Audio Large Language Model for multidimensional reasoning, detection and classification of audio artifacts. First, a calibration stage aligns the model to predict predefined perceptual dimensions. Second, a reinforcement learning stage leverages Group Relative Policy Optimization (GRPO) with dimension-specific rewards to heavily enhance accuracy of descriptions and temporal localization of quality issues. With this approach we reach state-of-the-art results of 0.71 mean PCC score on the multidimensional QualiSpeech benchmark and 13% improvement in MOS prediction driven by RL-based reasoning. Furthermore, our fine-grained GRPO rewards substantially advance the model's ability to pinpoint and classify audio artifacts in time.
【7】nlm: Real-Time Non-linear Modal Synthesis in Max
标题:nlm:Max中的实时非线性模式合成
链接:https://arxiv.org/abs/2603.10240
备注:accepted to PdMaxCon25~ (https://music.illinois.edu/pd-max-con/)
摘要:我们提出了\texttt{nlm},一组最大外部,使有效的实时非线性模态合成字符串,膜和板。外部,在C++实现,提供物理参数的交互式控制,允许自定义模态数据的加载,并提供多通道输出。通过将交互式物理建模功能集成到熟悉的环境中,降低了作曲家、表演者和声音设计师探索非线性模态合成的表达潜力的障碍。这些外部软件可以在https://github.com/rodrigodzf/nlm上以开源软件的形式获得。
摘要:We present \texttt{nlm}, a set of Max externals that enable efficient real-time non-linear modal synthesis for strings, membranes, and plates. The externals, implemented in C++, offer interactive control of physical parameters, allow the loading of custom modal data, and provide multichannel output. By integrating interactive physical-modelling capabilities into a familiar environment, \texttt{nlm} lowers the barrier for composers, performers, and sound designers to explore the expressive potential of non-linear modal synthesis. The externals are available as open-source software at https://github.com/rodrigodzf/nlm.
机器翻译由腾讯交互翻译提供,仅供参考
