微信公众号:arXiv_Daily
cs.SD语音
【1】VoiceAgengRAG: Solving the RAG Latency Bottleneck in Real-Time Voice Agents Using Dual-Agent Architectures
标题:VoiceDelivergRAG:使用双代理架构解决实时语音代理中的RAG延迟瓶颈
链接:https://arxiv.org/abs/2603.02206
摘要:我们提出了VoiceAgentRAG,一个开源的双代理内存路由器,从响应生成检索。后台的Slow Thinker代理持续监控会话流,使用LLM预测可能的后续主题,并将相关文档块预取到FAISS支持的语义缓存中。前台FastTalker代理仅从该亚毫秒级缓存读取,在缓存命中时完全绕过矢量数据库。
摘要:We present VoiceAgentRAG, an open-source dual-agent memory router that decouples retrieval from response generation. A background Slow Thinker agent continuously monitors the conversation stream, predicts likely follow-up topics using an LLM, and pre-fetches relevant document chunks into a FAISS-backed semantic cache. A foreground Fast Talker agent reads only from this sub-millisecond cache, bypassing the vector database entirely on cache hits.
【2】Analytical Exploration of Spatial Audio Cues: A Differentiable Multi-Sphere Scattering Model
标题:空间音频线索的分析探索:可区分的多球散射模型
链接:https://arxiv.org/abs/2603.02205
摘要:开发合成空间听觉系统(特别是水下)的主要挑战是精确建模声音散射。生物有机体通过利用其身体的声音散射来产生位置相关的耳间水平和时间差(ITD/ILD)来实现3D空间听觉。虽然基于刚性散射的头部相关传递函数(HRTF)模型足以满足陆地人类,但由于水和软组织之间的近阻抗匹配,它们在水下环境中失败。出于水下动物的声学解剖,我们介绍了一种新的,解析推导,封闭形式的前向散射模型从一个半透明的球体包含两个刚性的球形散射体。该模型准确地将源方向、频率和材料特性映射到压力场,捕捉分层、可穿透结构的复杂物理特性。重要的是,我们的模型是在完全可微的设置中实现的,使其能够与机器学习算法集成,以优化主动定位的成本函数。我们展示了增强的收敛性,在噪声下使用物理信息的频率加权方案的本地化,并提出了准确的移动源跟踪通过扩展卡尔曼滤波器(EKF)与分析计算的雅可比矩阵。我们的工作表明,分层的刚性和透明的几何形状的散射微分模型提供了一个很有前途的新的基础麦克风阵列,利用基于散射的空间线索在传统的波束形成,适用于陆地和水下应用。我们的模型将是开源的。
摘要:A primary challenge in developing synthetic spatial hearing systems, particularly underwater, is accurately modeling sound scattering. Biological organisms achieve 3D spatial hearing by exploiting sound scattering off their bodies to generate location-dependent interaural level and time differences (ITD/ILD). While Head-Related Transfer Function (HRTF) models based on rigid scattering suffice for terrestrial humans, they fail in underwater environments due to the near-impedance match between water and soft tissue. Motivated by the acoustic anatomy of underwater animals, we introduce a novel, analytically derived, closed-form forward model for scattering from a semi-transparent sphere containing two rigid spherical scatterers. This model accurately maps source direction, frequency, and material properties to the pressure field, capturing the complex physics of layered, penetrable structures. Critically, our model is implemented in a fully differentiable setting, enabling its integration with a machine learning algorithm to optimize a cost function for active localization. We demonstrate enhanced convergence for localization under noise using a physics-informed frequency weighting scheme, and present accurate moving-source tracking via an Extended Kalman Filter (EKF) with analytically computed Jacobians. Our work suggests that differentiable models of scattering from layered rigid and transparent geometries offer a promising new foundation for microphone arrays that leverage scattering-based spatial cues over conventional beamforming, applicable to both terrestrial and underwater applications. Our model will be made open source.
【3】CodecFlow: Efficient Bandwidth Extension via Conditional Flow Matching in Neural Codec Latent Space
标题:CodecFlow:通过神经编解码器潜在空间中的条件流匹配高效带宽扩展
链接:https://arxiv.org/abs/2603.02022
备注:7 pages, 7 figures
摘要:语音带宽扩展通过为低带宽语音恢复/推断适当的高频内容来提高清晰度和可懂度。现有的方法通常依赖于频谱图或波形建模,这会导致较高的计算成本并且具有有限的高频保真度。神经音频编解码器提供了紧凑的潜在表示,可以更好地保留声学细节,但由于表示不匹配,准确恢复高分辨率潜在信息仍然具有挑战性。我们提出了CodecFlow,一个基于神经编解码器的BWE框架,在紧凑的潜在空间中执行有效的语音重建。CodecFlow在连续编解码器嵌入和结构约束的残差矢量量化器上采用了一个具有感知能力的条件流转换器,以提高潜在对齐稳定性。经过端到端优化的CodecFlow在8 kHz至16 kHz和44.1 kHz语音BWE任务中实现了强大的频谱保真度和增强的感知质量。
摘要:Speech Bandwidth Extension improves clarity and intelligibility by restoring/inferring appropriate high-frequency content for low-bandwidth speech. Existing methods often rely on spectrogram or waveform modeling, which can incur higher computational cost and have limited high-frequency fidelity. Neural audio codecs offer compact latent representations that better preserve acoustic detail, yet accurately recovering high-resolution latent information remains challenging due to representation mismatch. We present CodecFlow, a neural codec-based BWE framework that performs efficient speech reconstruction in a compact latent space. CodecFlow employs a voicing-aware conditional flow converter on continuous codec embeddings and a structure-constrained residual vector quantizer to improve latent alignment stability. Optimized end-to-end, CodecFlow achieves strong spectral fidelity and enhanced perceptual quality on 8 kHz to 16 kHz and 44.1 kHz speech BWE tasks.
【4】ViTex: Visual Texture Control for Multi-Track Symbolic Music Generation via Discrete Diffusion Models
标题:ViTex:通过离散扩散模型实现多轨符号音乐生成的视觉纹理控制
链接:https://arxiv.org/abs/2603.01984
摘要:在自动音乐生成中,一个核心挑战是设计能够实现有意义的人机交互的控件。现有的系统通常依赖于外部输入,如文本提示或元数据,这不允许人类直接塑造组合。虽然以前的工作已经探索了内在的控制,如和弦或层次结构,这些方法主要解决钢琴或声乐伴奏设置,留下多轨象征性的音乐在很大程度上探索不足。我们确定仪器,仪器的选择和它们的作用,作为一个自然的多轨道组成的控制维度,并提出ViTex,一个视觉表示的工具纹理。在ViTex中,颜色编码乐器选择,空间位置表示音高和时间,笔划属性捕获局部纹理。在此表示的基础上,我们开发了一个离散扩散模型条件ViTex和和弦进行生成8测量多轨道符号音乐,使明确的纹理级控制,同时保持强大的无条件生成质量。演示页面和代码可在https://vitex2025.github.io/上获得。
摘要:In automatic music generation, a central challenge is to design controls that enable meaningful human-machine interaction. Existing systems often rely on extrinsic inputs such as text prompts or metadata, which do not allow humans to directly shape the composition. While prior work has explored intrinsic controls such as chords or hierarchical structure, these approaches mainly address piano or vocal-accompaniment settings, leaving multitrack symbolic music largely underexplored. We identify instrumentation, the choice of instruments and their roles, as a natural dimension of control in multi-track composition, and propose ViTex, a visual representation of instrumental texture. In ViTex, color encodes instrument choice, spatial position represents pitch and time, and stroke properties capture local textures. Building on this representation, we develop a discrete diffusion model conditioned on ViTex and chord progressions to generate 8-measure multi-track symbolic music, enabling explicit texture-level control while maintaining strong unconditional generation quality. The demo page and code are avaliable at https://vitex2025.github.io/.
【5】VietSuperSpeech: A Large-Scale Vietnamese Conversational Speech Dataset for ASR Fine-Tuning in Chatbot, Customer Support, and Call Center Applications
标题:VietSuperSpeech:一个大规模越南对话语音数据集,用于Chatbot、客户支持和呼叫中心应用程序中的ASB微调
链接:https://arxiv.org/abs/2603.01894
摘要:我们介绍VietSuperSpeech,这是一个大规模的越南语自动语音识别(ASR)数据集,包含52,023个音频文本对,总计267.39小时,特别关注休闲会话语音。与现有的越南语ASR语料库主要以阅读演讲,新闻叙述或有声读物内容为特色不同,VietSuperSpeech来自四个可公开访问的YouTube频道,涵盖日常对话,个人vlogging,海外越南社区对话和非正式评论-现实世界聊天机器人,客户支持,呼叫中心和热线部署中遇到的演讲风格。所有音频标准化为16 kHz单声道PCM WAV,并分段为3-30秒的话语。Transcript是使用Zipformer-30 M-RNNT-6, 000 h模型(Nguyen,2025)通过伪标记生成的,该模型通过Sherpa-ONNX部署,并在6,000小时的越南语语音上进行了预训练。经过质量过滤后,数据集被分成46,822个训练样本(240.67小时)和5,201个开发/测试样本(26.72小时),并具有固定的随机种子。文本平均每个话语266个字符,总计1380万个完全变音标记的越南语字符。我们证明了VietSuperSpeech填补了越南ASR生态系统中的一个关键空白:虽然VLSP 2020,VIET_BUD500,VietSpeech,FLEURS,VietMed,Sub-GigaSpeech 2-Vi,viVoice和Sub-PhoAudioBook等语料库提供了正式和阅读语音的广泛覆盖,但没有一个专门针对会话AI应用程序不可或缺的随意,自发注册。VietSuperSpeech在https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech公开发布。
摘要:We introduce VietSuperSpeech, a large-scale Vietnamese automatic speech recognition (ASR) dataset of 52,023 audio-text pairs totaling 267.39 hours, with a distinctive focus on casual conversational speech. Unlike existing Vietnamese ASR corpora that predominantly feature read speech, news narration, or audiobook content, VietSuperSpeech is sourced from four publicly accessible YouTube channels spanning everyday conversation, personal vlogging, overseas Vietnamese community dialogue, and informal commentary - the very speech styles encountered in real-world chatbot, customer support, call center, and hotline deployments. All audio is standardized to 16 kHz mono PCM WAV and segmented into 3-30 second utterances. Transcriptions are generated via pseudo-labeling using the Zipformer-30M-RNNT-6000h model (Nguyen, 2025) deployed through Sherpa-ONNX, pre-trained on 6,000 hours of Vietnamese speech. After quality filtering, the dataset is split into 46,822 training samples (240.67 hours) and 5,201 development/test samples (26.72 hours) with a fixed random seed. The text averages 266 characters per utterance, totaling 13.8 million fully diacritically marked Vietnamese characters. We demonstrate that VietSuperSpeech fills a critical gap in the Vietnamese ASR ecosystem: while corpora such as VLSP2020, VIET_BUD500, VietSpeech, FLEURS, VietMed, Sub-GigaSpeech2-Vi, viVoice, and Sub-PhoAudioBook provide broad coverage of formal and read speech, none specifically targets the casual, spontaneous register indispensable for conversational AI applications. VietSuperSpeech is publicly released at https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.
【6】TQCodec: Towards neural audio codec for high-fidelity music streaming
标题:TQCodec:迈向高保真音乐流媒体的神经音频编解码器
链接:https://arxiv.org/abs/2603.01592
摘要:我们提出了TQCodec,一种神经音频编解码器,专为高比特率,高保真音乐流而设计。与现有的主要针对超低比特率(<= 16 kbps)的神经编解码器不同,TQCodec以44.1 kHz运行,支持32 kbps至128 kbps的比特率,符合现代音乐流媒体平台的标准质量。该模型采用了基于SEANet的编码器-解码器架构,以实现高效的设备上计算,并引入了几项增强功能:用于以低开销提高质量的不平衡网络设计,用于中频细节保留的SimVQ,以及相位感知波形丢失。此外,我们引入了一个感知驱动的波段明智的比特分配策略,优先考虑感知关键的低频。对不同音乐数据集的评估表明,TQCodec在目标比特率下实现了卓越的音频质量,使其非常适合高质量的音频应用。
摘要:We propose TQCodec, a neural audio codec designed for high-bitrate, high-fidelity music streaming. Unlike existing neural codecs that primarily target ultra-low bitrates (<= 16kbps), TQCodec operates at 44.1 kHz and supports bitrates from 32 kbps to 128 kbps, aligning with the standard quality of modern music streaming platforms. The model adopts an encoder-decoder architecture based on SEANet for efficient on-device computation and introduces several enhancements: an imbalanced network design for improved quality with low overhead, SimVQ for mid-frequency detail preservation, and a phase-aware waveform loss. Additionally, we introduce a perception-driven band-wise bit allocation strategy to prioritize perceptually critical lower frequencies. Evaluations on diverse music datasets demonstrate that TQCodec achieves superior audio quality at target bitrates, making it well-suited for high-quality audio applications.
【7】UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
标题:UniTalking:用于说话肖像生成的统一音视频框架
链接:https://arxiv.org/abs/2603.01418
备注:Accepted at CVPR 2026 (Findings Track)
摘要:虽然Veo3和Sora2等最先进的音频视频生成模型展示了卓越的功能,但它们的闭源特性使其架构和训练范式无法访问。为了弥合可访问性和性能方面的差距,我们引入了UniTalking,这是一个统一的端到端扩散框架,用于生成高保真语音和嘴唇同步视频。在其核心,我们的框架采用多模态Transformer块显式地模拟细粒度的时间之间的对应关系的音频和视频潜在令牌通过共享的自我注意力机制。通过利用来自预训练视频生成模型的强大先验,我们的框架确保了最先进的视觉保真度,同时实现了高效的训练。此外,UniTalking还集成了个性化的语音克隆功能,允许从简短的音频参考中生成目标风格的语音。定性和定量的结果表明,我们的方法产生高度逼真的说话肖像,实现优于现有的开源方法在唇同步的准确性,音频自然度和整体感知质量。
摘要:While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and performance, we introduce UniTalking, a unified, end-to-end diffusion framework for generating high-fidelity speech and lip-synchronized video. At its core, our framework employs Multi-Modal Transformer Blocks to explicitly model the fine-grained temporal correspondence between audio and video latent tokens via a shared self-attention mechanism. By leveraging powerful priors from a pre-trained video generation model, our framework ensures state-of-the-art visual fidelity while enabling efficient training. Furthermore, UniTalking incorporates a personalized voice cloning capability, allowing the generation of speech in a target style from a brief audio reference. Qualitative and quantitative results demonstrate that our method produces highly realistic talking portraits, achieving superior performance over existing open-source approaches in lip-sync accuracy, audio naturalness, and overall perceptual quality.
【8】End-to-End Simultaneous Dysarthric Speech Reconstruction with Frame-Level Adaptor and Multiple Wait-k Knowledge Distillation
标题:使用帧级适配器和多重Wait-k知识提炼的端到端同时发音障碍语音重建
链接:https://arxiv.org/abs/2603.01382
备注:Submitted to 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
摘要:构音障碍语音重建(DSR)通常采用级联系统,该级联系统将自动语音识别(ASR)和语音级文本到语音(TTS)相结合,以将构音障碍语音转换为正常韵律的语音。然而,构音障碍的个体通常说话更慢,导致这种系统的响应时间过长,使得它们在长时间讲话的情况下不切实际。基于流式ASR和增量TTS的级联DSR系统可以帮助减少延迟。然而,具有不同构音障碍严重程度的患者对于相同文本表现出显著的发音变异性,导致ASR的鲁棒性差并且限制了重建语音的可懂度。此外,增量TTS遭受差的韵律特征预测由于有限的感受野。在本研究中,我们提出了一个端到端的同步DSR系统,其中有两个关键的创新:1)引入了一个帧级适配器模块来桥接ASR和TTS。通过采用显隐语义信息融合和联合模块训练,提高了TTS对ASR输出的容错能力。2)设计了一个多等待k自回归TTS模块,通过多视角知识提取来减轻韵律退化。我们的系统在Tesla A100上的平均响应时间为1.03秒,平均实时因子(RTF)为0.71。在UASpeech数据集上,它的平均意见得分(MOS)为4.67,与最先进的技术相比,单词错误率(WER)相对降低了54.25%。我们的演示可在以下网站获得:https://wflrz123.github.io/
摘要:Dysarthric speech reconstruction (DSR) typically employs a cascaded system that combines automatic speech recognition (ASR) and sentence-level text-to-speech (TTS) to convert dysarthric speech into normally-prosodied speech. However, dysarthric individuals often speak more slowly, leading to excessively long response times in such systems, rendering them impractical in long-speech scenarios. Cascaded DSR systems based on streaming ASR and incremental TTS can help reduce latency. However, patients with differing dysarthria severity exhibit substantial pronunciation variability for the same text, resulting in poor robustness of ASR and limiting the intelligibility of reconstructed speech. In addition, incremental TTS suffers from poor prosodic feature prediction due to a limited receptive field. In this study, we propose an end-to-end simultaneous DSR system with two key innovations: 1) A frame-level adaptor module is introduced to bridge ASR and TTS. By employing explicit-implicit semantic information fusion and joint module training, it enhances the error tolerance of TTS to ASR outputs. 2) A multiple wait-k autoregressive TTS module is designed to mitigate prosodic degradation via multi-view knowledge distillation. Our system has an average response time of 1.03 seconds on Tesla A100, with an average real-time factor (RTF) of 0.71. On the UASpeech dataset, it attains a mean opinion score (MOS) of 4.67 and demonstrates a 54.25% relative reduction in word error rate (WER) compared to the state-of-the-art. Our demo is available at: https://wflrz123.github.io/
【9】DARS: Dysarthria-Aware Rhythm-Style Synthesis for ASR Enhancement
标题:DARS:用于ASB增强的构音障碍意识节奏风格合成
链接:https://arxiv.org/abs/2603.01369
备注:Submitted to 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
摘要:构音障碍语音表现出异常的韵律和显著的说话者变异性,给自动语音识别(ASR)带来了持续的挑战。虽然基于文本到语音(TTS)的数据增强已显示出潜力,但现有的方法往往无法准确地模拟构音障碍语音的病理节奏和声学风格。为了解决这个问题,我们提出了DARS,一个发音困难的节奏风格的合成框架的基础上的Matcha-TTS架构。DARS结合了一个多阶段的节奏预测器,通过正常和构音障碍语音之间的对比偏好进行优化,以及构音障碍风格的条件流匹配机制,共同增强时间节奏重建和病理声学风格模拟。在TORGO数据集上的实验表明,DARS实现了4.29的平均倒频谱失真(MCD),非常接近真实的构音障碍语音。与最先进的方法相比,使用来自DARS的合成构音障碍语音适配基于Whisper的ASR系统实现了54.22%的字错误率(WER)相对降低,证明了该框架在提高识别性能方面的有效性。
摘要:Dysarthric speech exhibits abnormal prosody and significant speaker variability, presenting persistent challenges for automatic speech recognition (ASR). While text-to-speech (TTS)-based data augmentation has shown potential, existing methods often fail to accurately model the pathological rhythm and acoustic style of dysarthric speech. To address this, we propose DARS, a dysarthria-aware rhythm-style synthesis framework based on the Matcha-TTS architecture. DARS incorporates a multi-stage rhythm predictor optimized by contrastive preferences between normal and dysarthric speech, along with a dysarthric-style conditional flow matching mechanism, jointly enhancing temporal rhythm reconstruction and pathological acoustic style simulation. Experiments on the TORGO dataset demonstrate that DARS achieves a Mean Cepstral Distortion (MCD) of 4.29, closely approximating real dysarthric speech. Adapting a Whisper-based ASR system with synthetic dysarthric speech from DARS achieves a 54.22% relative reduction in word error rate (WER) compared to state-of-the-art methods, demonstrating the framework's effectiveness in enhancing recognition performance.
【10】SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music Generation
标题:SyncTrack:多轨音乐生成中的节奏稳定性和同步性
链接:https://arxiv.org/abs/2603.01101
备注:Accepted by ICLR 2026
摘要:多音轨音乐生成由于其精确的混音和再混音能力而获得了重大的研究兴趣。然而,现有的模型往往忽略了基本属性,如节奏的稳定性和同步,导致重点放在轨道之间的差异,而不是其固有的属性。在本文中,我们介绍了同步跟踪,同步多轨波形音乐生成模型,旨在捕捉多轨音乐的独特特性。SyncTrack采用了一种新颖的架构,其中包括轨道共享模块,以在所有轨道和轨道特定模块之间建立共同的节奏,以适应不同的音色和音高范围。每个音轨共享模块采用两个跨音轨注意机制来同步节奏信息,而每个音轨特定模块利用可学习的乐器先验来更好地表示音色和其他独特特征。此外,我们通过三个新的指标引入节奏一致性来增强对多轨音乐质量的评估:内轨节奏稳定性(IRS)、跨轨节拍同步(CBS)和跨轨节拍分散(CBD)。实验结果表明,SyncTrack通过增强节奏的一致性,显著提高了多音轨音乐的质量。
摘要:Multi-track music generation has garnered significant research interest due to its precise mixing and remixing capabilities. However, existing models often overlook essential attributes such as rhythmic stability and synchronization, leading to a focus on differences between tracks rather than their inherent properties. In this paper, we introduce SyncTrack, a synchronous multi-track waveform music generation model designed to capture the unique characteristics of multi-track music. SyncTrack features a novel architecture that includes track-shared modules to establish a common rhythm across all tracks and track-specific modules to accommodate diverse timbres and pitch ranges. Each track-shared module employs two cross-track attention mechanisms to synchronize rhythmic information, while each track-specific module utilizes learnable instrument priors to better represent timbre and other unique features. Additionally, we enhance the evaluation of multi-track music quality by introducing rhythmic consistency through three novel metrics: Inner-track Rhythmic Stability (IRS), Cross-track Beat Synchronization (CBS), and Cross-track Beat Dispersion (CBD). Experiments demonstrate that SyncTrack significantly improves the multi-track music quality by enhancing rhythmic consistency.
【11】AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
标题:AG-REPA:音频流匹配中用于表示对齐的因果层选择
链接:https://arxiv.org/abs/2603.01006
备注:13 pages, 4 figures, 4 tables
摘要:REPresentation对齐(REPA)通过将中间隐藏状态与预训练的教师特征对齐来改进生成流模型的训练,但其在令牌条件音频流匹配中的有效性关键取决于监督层的选择,这通常是基于深度的启发式的。在这项工作中,我们介绍了属性引导REPresentation对齐(AG-REPA),一种新的因果层选择策略表示对齐音频流匹配。首先,我们发现最好地存储语义/声学信息(高教师空间相似性)的层不一定是对驱动生成的速度场贡献最大的层,我们称之为存储贡献分离(SCD)。为了将这种洞察力转化为可操作的训练指导,我们提出了一种仅向前的门消融(FoG-A),通过预测速度场的诱导变化来量化每个层的因果贡献,从而实现稀疏层选择和自适应加权对齐。在不同令牌条件拓扑下的统一语音和通用音频训练(LibriSpeech + AudioSet)中,AG-REPA始终优于REPA基线。总的来说,我们的研究结果表明,对齐是最有效的,当应用到因果关系占主导地位的层,驱动速度场,而不是层的代表性丰富,但功能被动。
摘要:REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth. In this work, we introduce Attribution-Guided REPresentation Alignment (AG-REPA), a novel causal layer selection strategy for representation alignment in audio Flow Matching. Firstly, we find that layers that best store semantic/acoustic information (high teacher-space similarity) are not necessarily the layers that contribute most to the velocity field that drives generation, and we call it Store-Contribute Dissociation (SCD). To turn this insight into an actionable training guidance, we propose a forward-only gate ablation (FoG-A) that quantifies each layer's causal contribution via the induced change in the predicted velocity field, enabling sparse layer selection and adaptive weighting for alignment. Across unified speech and general-audio training (LibriSpeech + AudioSet) under different token-conditioning topologies, AG-REPA consistently outperforms REPA baselines. Overall, our results show that alignment is most effective when applied to the causally dominant layers that drive the velocity field, rather than to layers that are representationally rich but functionally passive.
【12】Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages
标题:对印度语言语音识别系统进行正字知情评估
链接:https://arxiv.org/abs/2603.00941
备注:Accepted in ICASSP 2026
摘要:评估印度语言的ASR系统是具有挑战性的,因为拼写变化,后缀分裂的灵活性,以及代码混合单词的非标准拼写。传统的字错误率(WER)通常呈现出比人类用户感知的更黯淡的系统性能。更好地将评估与真实世界的表现相结合需要捕捉允许的正字法变化,这对于资源不足的印度语言来说是极具挑战性的。利用LLM的最新进展,我们提出了一个框架,用于创建捕获允许变化的基准。通过大量的实验,我们证明了OIWER通过考虑正交变化,降低了悲观错误率(平均提高6.3分),缩小了膨胀的模型差距(例如,Gemini-Canary性能差异从18.1下降到11.5分),并且比WER-SN等现有方法更接近人类感知4.9分。
摘要:Evaluating ASR systems for Indian languages is challenging due to spelling variations, suffix splitting flexibility, and non-standard spellings in code-mixed words. Traditional Word Error Rate (WER) often presents a bleaker picture of system performance than what human users perceive. Better aligning evaluation with real-world performance requires capturing permissible orthographic variations, which is extremely challenging for under-resourced Indian languages. Leveraging recent advances in LLMs, we propose a framework for creating benchmarks that capture permissible variations. Through extensive experiments, we demonstrate that OIWER, by accounting for orthographic variations, reduces pessimistic error rates (an average improvement of 6.3 points), narrows inflated model gaps (e.g., Gemini-Canary performance difference drops from 18.1 to 11.5 points), and aligns more closely with human perception than prior methods like WER-SN by 4.9 points.
【13】SpectroFusion-ViT: A Lightweight Transformer for Speech Emotion Recognition Using Harmonic Mel-Chroma Fusion
标题:SpectroFusion-ViT:使用Harmonic Mel-Chroma融合的语音情感识别轻量级Transformer
链接:https://arxiv.org/abs/2603.00746
摘要:言语是表达情感的自然手段,是理解和表达人类情感的有效方法。可靠的语音情感识别(SER)是人机交互、医疗保健、教育和客户服务等应用的核心。然而,大多数SER方法依赖于沉重的主干模型或手工制作的功能,无法平衡准确性和效率,特别是对于像Bangla这样的低资源语言。在这项工作中,我们提出了SpectroFusion-ViT,一个轻量级的SER框架,利用EfficientViT-b 0,一个紧凑的Vision Transformer架构,配备了自注意力,以捕获长距离的时间和光谱模式。该模型仅包含2.04 M参数,需要0.1 GFLOP,可在资源受限的环境中进行部署,而不会影响准确性。我们的管道首先对原始音频进行预处理和增强,然后提取色度和梅尔频率倒谱系数(MFCC)特征。这些表示融合成一个互补的时间-频率描述符,保留了细粒度的频谱细节和更广泛的谐波结构。使用迁移学习,EfficientViT-b 0针对多类情感分类进行了微调。我们评估系统的两个基准孟加拉语情感语音数据集,SUBESCO和BanglaSER,不同的扬声器的多样性,录音条件和声学特性。该方法在SUBESCO上实现了92.56%的准确率,在BanglaSER上实现了82.19%的准确率,超过了现有的最先进的方法。这些研究结果表明,轻量级的Transformer架构可以提供强大的SER性能,同时保持计算效率的现实世界的部署。
摘要:Speech is a natural means of conveying emotions, making it an effective method for understanding and representing human feelings. Reliable speech emotion recognition (SER) is central to applications in human-computer interaction, healthcare, education, and customer service. However, most SER methods depend on heavy backbone models or hand-crafted features that fail to balance accuracy and efficiency, particularly for low-resource languages like Bangla. In this work, we present SpectroFusion-ViT, a lightweight SER framework built utilizing EfficientViT-b0, a compact Vision Transformer architecture equipped with self-attention to capture long-range temporal and spectral patterns. The model contains only 2.04M parameters and requires 0.1 GFLOPs, enabling deployment in resource-constrained settings without compromising accuracy. Our pipeline first performs preprocessing and augmentation on raw audio, then extracts Chroma and Mel-frequency cepstral coefficient (MFCC) features. These representations are fused into a complementary time-frequency descriptor that preserves both fine-grained spectral detail and broader harmonic structure. Using transfer learning, EfficientViT-b0 is fine-tuned for multi-class emotion classification. We evaluate the system on two benchmark Bangla emotional speech datasets, SUBESCO and BanglaSER, which vary in speaker diversity, recording conditions, and acoustic characteristics. The proposed approach achieves 92.56% accuracy on SUBESCO and 82.19% on BanglaSER, surpassing existing state-of-the-art methods. These findings demonstrate that lightweight transformer architectures can deliver robust SER performance while remaining computationally efficient for real-world deployment.
【14】CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
标题:CMI-RewardBench:使用合成多模式教学评估音乐奖励模型
链接:https://arxiv.org/abs/2603.00610
摘要:虽然音乐生成模型已经发展到可以处理混合了文本、歌词和参考音频的复杂多模态输入,但评估机制却落后了。在本文中,我们通过建立一个综合的生态系统,在作曲多模态教学(CMI)下的音乐奖励建模,其中生成的音乐可能会以文本描述,歌词和音频提示为条件,来弥合这一关键差距。我们首先介绍CMI-Pref-Pseudo,一个包含110 k伪标记样本的大规模偏好数据集,以及CMI-Pref,一个为细粒度对齐任务量身定制的高质量人工注释语料库。为了统一评估环境,我们提出了CMI-RewardBench,这是一个统一的基准,可以在音乐性,文本音乐对齐和作曲指令对齐的异构样本上评估音乐奖励模型。利用这些资源,我们开发了CMI奖励模型(CMI-RM),一个参数有效的奖励模型家族能够处理异构输入。我们评估了它们与人类对音乐性的判断分数的相关性,以及与以前的数据集在CMI-Pref上的对齐。进一步的实验表明,CMI-RM不仅与人类的判断密切相关,而且还可以通过top-k滤波实现有效的推理时间缩放。必要的训练数据、基准和奖励模型是公开的。
摘要:While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive ecosystem for music reward modeling under Compositional Multimodal Instruction (CMI), where the generated music may be conditioned on text descriptions, lyrics, and audio prompts. We first introduce CMI-Pref-Pseudo, a large-scale preference dataset comprising 110k pseudo-labeled samples, and CMI-Pref, a high-quality, human-annotated corpus tailored for fine-grained alignment tasks. To unify the evaluation landscape, we propose CMI-RewardBench, a unified benchmark that evaluates music reward models on heterogeneous samples across musicality, text-music alignment, and compositional instruction alignment. Leveraging these resources, we develop CMI reward models (CMI-RMs), a parameter-efficient reward model family capable of processing heterogeneous inputs. We evaluate their correlation with human judgments scores on musicality and alignment on CMI-Pref along with previous datasets. Further experiments demonstrate that CMI-RM not only correlates strongly with human judgments, but also enables effective inference-time scaling via top-k filtering. The necessary training data, benchmarks, and reward models are publicly available.
【15】Efficient Long-Sequence Diffusion Modeling for Symbolic Music Generation
标题:符号音乐生成的有效长序列扩散模型
链接:https://arxiv.org/abs/2603.00576
备注:17 pages, 5 figures
摘要:符号音乐生成是多媒体生成中的一项具有挑战性的任务,涉及具有层次时间结构的长序列,长范围依赖性和细粒度局部细节。虽然最近的基于扩散的模型产生高质量的生成,但由于迭代去噪和序列长度相关的成本,它们往往会遭受长符号序列的高训练和推理成本。针对这一问题,本文提出了一种有效的全局结构构建和轻度局部细化相结合的扩散策略SMDIM。SMDIM使用结构化的状态空间模型,以接近线性的成本捕捉远程音乐上下文,并通过混合细化方案选择性地细化局部音乐细节。在包含各种西方古典音乐、流行音乐和传统民间音乐的符号音乐数据集上进行的实验表明,SMDIM模型在生成质量和计算效率方面优于其他最先进的方法,并且对未充分探索的音乐风格具有鲁棒的泛化能力。这些结果表明,SMDIM提供了一个原则性的解决方案,长序列的符号音乐生成,包括伴随序列的相关属性。我们在https://3328702107.github.io/smdim-music/上提供了一个包含音频示例和补充材料的项目网页。
摘要:Symbolic music generation is a challenging task in multimedia generation, involving long sequences with hierarchical temporal structures, long-range dependencies, and fine-grained local details. Though recent diffusion-based models produce high quality generations, they tend to suffer from high training and inference costs with long symbolic sequences due to iterative denoising and sequence-length-related costs. To deal with such problem, we put forth a diffusing strategy named SMDIM to combine efficient global structure construction and light local refinement. SMDIM uses structured state space models to capture long range musical context at near linear cost, and selectively refines local musical details via a hybrid refinement scheme. Experiments performed on a wide range of symbolic music datasets which encompass various Western classical music, popular music and traditional folk music show that the SMDIM model outperforms the other state-of-the-art approaches on both the generation quality and the computational efficiency, and it has robust generalization to underexplored musical styles. These results show that SMDIM offers a principled solution for long-sequence symbolic music generation, including associated attributes that accompany the sequences. We provide a project webpage with audio examples and supplementary materials at https://3328702107.github.io/smdim-music/.
【16】Whisper-MLA: Reducing GPU Memory Consumption of ASR Models based on MHA2MLA Conversion
标题:Whisper-MLA:基于MHA 2 MLA转换降低ASB型号的图形内存消耗
链接:https://arxiv.org/abs/2603.00563
备注:5 pages, 3 figures, accepted at ICASSP 2026
摘要:基于transformer的Whisper模型在自动语音识别(ASR)中已经达到了最先进的性能。然而,由于线性增长的键值(KV)缓存使用率,其多头注意力(MHA)机制会导致显著的GPU内存消耗,这对于许多应用程序来说是个问题,特别是对于长格式音频。为了解决这个问题,我们引入了Whisper-MLA,这是一种将多头潜在注意力(MLA)纳入Whisper模型的新型架构。具体来说,我们适应MLA耳语的绝对位置嵌入和系统地研究其应用程序在编码器的自我注意,解码器的自我注意,和交叉注意模块。经验结果表明,专门应用MLA解码器的自我注意力产生的性能和内存效率之间的理想平衡。我们提出的方法允许将预训练的Whisper模型转换为Whisper-MLA,只需最小的微调。LibriSpeech基准测试的大量实验验证了这种转换的有效性,表明Whisper-MLA将KV缓存大小减少了87.5%,同时保持了具有竞争力的准确性。
摘要:The Transformer-based Whisper model has achieved state-of-the-art performance in Automatic Speech Recognition (ASR). However, its Multi-Head Attention (MHA) mechanism results in significant GPU memory consumption due to the linearly growing Key-Value (KV) cache usage, which is problematic for many applications especially with long-form audio. To address this, we introduce Whisper-MLA, a novel architecture that incorporates Multi-Head Latent Attention (MLA) into the Whisper model. Specifically, we adapt MLA for Whisper's absolute positional embeddings and systematically investigate its application across encoder self-attention, decoder self-attention, and cross-attention modules. Empirical results indicate that applying MLA exclusively to decoder self-attention yields the desired balance between performance and memory efficiency. Our proposed approach allows conversion of a pretrained Whisper model to Whisper-MLA with minimal fine-tuning. Extensive experiments on the LibriSpeech benchmark validate the effectiveness of this conversion, demonstrating that Whisper-MLA reduces the KV cache size by up to 87.5% while maintaining competitive accuracy.
【17】Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
标题:文明之声:全球音乐理解的多语言QA基准
链接:https://arxiv.org/abs/2603.00533
备注:2 pages, 2 figures, 1 table, accepted by ISMIR 2025 LBD
摘要:我们介绍了文明的声音,第一个多语言QA基准评估音频LLM的文化理解全长音乐录音。覆盖38种语言的380首曲目,我们的自动化管道通过四个阶段产生1,190个多项选择题-每个阶段都需要人工验证:1)编制代表性音乐列表; 2)通过LLM为音乐列表中的每个样本生成文化背景文档; 3)从这些文档中提取关键属性;(4)设计多项选择题,探讨语言、地域联想、语气和主题内容。我们在四种条件下评估模型,并报告每种语言的准确性。我们的研究结果表明,即使是最先进的音频LLM也很难在没有丰富文本背景的情况下捕捉微妙的文化细微差别,并在解释来自不同文化传统的音乐时表现出系统性偏见。该数据集在Hugging Face上公开,以促进文化包容性的音乐理解研究。
摘要:We introduce Voices of Civilizations, the first multilingual QA benchmark for evaluating audio LLMs' cultural comprehension on full-length music recordings. Covering 380 tracks across 38 languages, our automated pipeline yields 1,190 multiple-choice questions through four stages - each followed by manual verification: 1) compiling a representative music list; 2) generating cultural-background documents for each sample in the music list via LLMs; 3) extracting key attributes from those documents; and 4) constructing multiple-choice questions probing language, region associations, mood, and thematic content. We evaluate models under four conditions and report per-language accuracy. Our findings demonstrate that even state-of-the-art audio LLMs struggle to capture subtle cultural nuances without rich textual context and exhibit systematic biases in interpreting music from different cultural traditions. The dataset is publicly available on Hugging Face to foster culturally inclusive music understanding research.
【18】Aurchestra: Fine-Grained, Real-Time Soundscape Control on Resource-Constrained Hearables
标题:Aurchestra:对资源受限的听力进行细粒度、实时音景控制
链接:https://arxiv.org/abs/2603.00395
备注:15 pages, 11 figures, 4 tables, submitted to ACM MobiSys 2026
摘要:可听设备正变得无处不在,但它们的声音控制仍然很迟钝:用户可以启用全局噪声抑制或专注于单个目标声音。然而,真实世界的声学场景包含许多用户可能想要独立调整的同时源。我们介绍了Aurchestra,第一个系统提供细粒度的,实时的声音景观控制资源受限的听觉。我们的系统有两个关键组件:(1)一个动态界面,只显示活动的声音类别;(2)一个实时的设备上多输出提取网络,为每个选定的类别生成单独的流,为多达5个重叠的目标声音实现强大的性能,并让用户通过自定义每个类别的音量来混合他们的环境,就像音频工程师混合曲目一样。我们优化了多个计算受限平台的模型架构,并在6 ms流式音频块上展示了实时性能。在以前看不见的室内和室外场景中的真实环境中,我们的系统能够实现富有表现力的每类声音控制,并在目标类增强和干扰抑制方面实现了实质性的改进。我们的研究结果表明,世界不需要作为一个单一的,无差别的流听到:与Aurchestra,声景变得真正可编程。
摘要:Hearables are becoming ubiquitous, yet their sound controls remain blunt: users can either enable global noise suppression or focus on a single target sound. Real-world acoustic scenes, however, contain many simultaneous sources that users may want to adjust independently. We introduce Aurchestra, the first system to provide fine-grained, real-time soundscape control on resource-constrained hearables. Our system has two key components: (1) a dynamic interface that surfaces only active sound classes and (2) a real-time, on-device multi-output extraction network that generates separate streams for each selected class, achieving robust performance for upto 5 overlapping target sounds, and letting users mix their environment by customizing per-class volumes, much like an audio engineer mixes tracks. We optimize the model architecture for multiple compute-limited platforms and demonstrate real-time performance on 6 ms streaming audio chunks. Across real-world environments in previously unseen indoor and outdoor scenarios, our system enables expressive per-class sound control and achieves substantial improvements in target-class enhancement and interference suppression. Our results show that the world need not be heard as a single, undifferentiated stream: with Aurchestra, the soundscape becomes truly programmable.
【19】StethoLM: Audio Language Model for Cardiopulmonary Analysis Across Clinical Tasks
标题:StethoLM:跨临床任务心肺分析的音频语言模型
链接:https://arxiv.org/abs/2603.00355
备注:To be published in TMLR
摘要:听心脏和肺的声音-听诊-是临床检查中最基本的步骤之一。尽管快速且无创,但它需要多年的经验来解释微妙的音频提示。最近的深度学习方法在自动化心肺音分析方面取得了进展,但大多数都局限于简单的分类,几乎没有提供临床可解释性或决策支持。我们提出StethoLM,第一个专门用于心肺听诊的音频语言模型,能够在整个听诊分析范围内执行预防驱动的临床任务。StethoLM将音频编码与医学语言模型主干集成在一起,并在StethoBench上进行训练,StethoBench是一个全面的基准测试,包括从16,125个标记的心肺记录中合成的77,027个预防-反应对,涵盖七个临床任务类别:二元分类,检测,报告,推理,鉴别诊断,比较和基于位置的分析。通过结合监督微调和直接偏好优化的多阶段训练,StethoLM在非分布数据上实现了性能和鲁棒性的大幅提升。我们的工作为临床听诊中的预防跟踪AI系统奠定了基础。
摘要:Listening to heart and lung sounds - auscultation - is one of the first and most fundamental steps in a clinical examination. Despite being fast and non-invasive, it demands years of experience to interpret subtle audio cues. Recent deep learning methods have made progress in automating cardiopulmonary sound analysis, yet most are restricted to simple classification and offer little clinical interpretability or decision support. We present StethoLM, the first audio-language model specialized for cardiopulmonary auscultation, capable of performing instruction-driven clinical tasks across the full spectrum of auscultation analysis. StethoLM integrates audio encoding with a medical language model backbone and is trained on StethoBench, a comprehensive benchmark comprising 77,027 instruction-response pairs synthesized from 16,125 labeled cardiopulmonary recordings spanning seven clinical task categories: binary classification, detection, reporting, reasoning, differential diagnosis, comparison, and location-based analysis. Through multi-stage training that combines supervised fine-tuning and direct preference optimization, StethoLM achieves substantial gains in performance and robustness on out-of-distribution data. Our work establishes a foundation for instruction-following AI systems in clinical auscultation.
【20】Acoustic Sensing for Universal Jamming Grippers
标题:通用干扰钳的声学传感
链接:https://arxiv.org/abs/2603.00351
备注:Accepted at ICRA 2026, supplementary material under https://rbo.gitlab-pages.tu-berlin.de/papers/acoustic-jamming-icra26/
摘要:通用干扰夹持器擅长抓住未知的物体,因为他们的顺应机构。传统的触觉传感器可能会损害这种顺应性,降低抓取性能。我们目前的声学传感作为一种形式的形态感测,其中夹具的柔软的身体本身成为传感器。扬声器和麦克风被放置在夹具腔内,远离可变形膜,完全保持顺应性。声音通过夹具和物体传播,编码物体属性,然后通过机器学习重建。我们的传感器在检测物体尺寸(2.6 mm误差)和方向(0.6度误差)方面实现了高空间分辨率,对80 dBA的外部噪声水平保持稳健,并区分物体材料(高达100%准确度)和16种日常物体(85.6%准确度)。我们验证了传感器在一个现实的触觉物体分类任务,实现了53分钟的不间断的抓取和传感,确认保存抓取性能。最后,我们证明,解开声学表示可以学习,提高鲁棒性无关的声学变化。
摘要:Universal jamming grippers excel at grasping unknown objects due to their compliant bodies. Traditional tactile sensors can compromise this compliance, reducing grasping performance. We present acoustic sensing as a form of morphological sensing, where the gripper's soft body itself becomes the sensor. A speaker and microphone are placed inside the gripper cavity, away from the deformable membrane, fully preserving compliance. Sound propagates through the gripper and object, encoding object properties, which are then reconstructed via machine learning. Our sensor achieves high spatial resolution in sensing object size (2.6 mm error) and orientation (0.6 deg error), remains robust to external noise levels of 80 dBA, and discriminates object materials (up to 100% accuracy) and 16 everyday objects (85.6% accuracy). We validate the sensor in a realistic tactile object sorting task, achieving 53 minutes of uninterrupted grasping and sensing, confirming the preserved grasping performance. Finally, we demonstrate that disentangled acoustic representations can be learned, improving robustness to irrelevant acoustic variations.
【21】FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
标题:FlowPortrait:用于音频驱动肖像视频生成的强化学习
链接:https://arxiv.org/abs/2603.00159
摘要:生成逼真的讲话头部视频仍然具有挑战性,这是由于持续存在的问题,如不完美的嘴唇同步,不自然的运动,以及与人类感知相关性差的评估指标。我们提出了FlowPortrait,一个基于自回归音频到视频生成的多模态主干的音频驱动的肖像动画的学习框架。FlowPortrait引入了一个基于多模态大型语言模型(MLLM)的人类对齐评估系统,以评估唇同步的准确性,表现力和运动质量。这些信号与感知和时间一致性正则化器相结合,形成稳定的复合奖励,用于通过组相对策略优化(GRPO)对生成器进行后训练。包括自动评估和人类偏好研究在内的大量实验表明,FlowPortrait始终能够生成更高质量的说话头部视频,突出了强化学习对肖像动画的有效性。
摘要:Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate poorly with human perception. We propose FlowPortrait, a reinforcement-learning framework for audio-driven portrait animation built on a multimodal backbone for autoregressive audio-to-video generation. FlowPortrait introduces a human-aligned evaluation system based on Multimodal Large Language Models (MLLMs) to assess lip-sync accuracy, expressiveness, and motion quality. These signals are combined with perceptual and temporal consistency regularizers to form a stable composite reward, which is used to post-train the generator via Group Relative Policy Optimization (GRPO). Extensive experiments, including both automatic evaluations and human preference studies, demonstrate that FlowPortrait consistently produces higher-quality talking-head videos, highlighting the effectiveness of reinforcement learning for portrait animation.
【22】Iterative LLM-based improvement for French Clinical Interview Transcription and Speaker Diarization
标题:基于LLM的法语临床访谈转录和演讲者拨号的迭代改进
链接:https://arxiv.org/abs/2603.00086
摘要:法语医学对话的自动语音识别仍然具有挑战性,自发临床语音的单词错误率通常超过30%。本研究提出了一种多通道LLM后处理架构,在说话人识别和单词识别通道之间交替,以提高转录准确性和说话人属性。消融研究两个法国临床数据集(自杀预防电话咨询和术前清醒神经外科咨询)调查四个设计选择:模型选择,提示策略,通过排序,迭代深度。使用Qwen 3-Next-80 B,Wilcoxon符号秩检验证实了自杀预防对话的WDER显著降低(p < 0.05,n=18),同时保持清醒神经外科会诊的稳定性(n=10),零输出故障和可接受的计算成本(RTF 0.32),表明离线临床部署的可行性。
摘要:Automatic speech recognition for French medical conversations remains challenging, with word error rates often exceeding 30% in spontaneous clinical speech. This study proposes a multi-pass LLM post-processing architecture alternating between Speaker Recognition and Word Recognition passes to improve transcription accuracy and speaker attribution. Ablation studies on two French clinical datasets (suicide prevention telephone counseling and preoperative awake neurosurgery consultations) investigate four design choices: model selection, prompting strategy, pass ordering, and iteration depth. Using Qwen3-Next-80B, Wilcoxon signed-rank tests confirm significant WDER reductions on suicide prevention conversations (p < 0.05, n=18), while maintaining stability on awake neurosurgery consultations (n=10), with zero output failures and acceptable computational cost (RTF 0.32), suggesting feasibility for offline clinical deployment.
【23】Investigating Group Relative Policy Optimization for Diffusion Transformer based Text-to-Audio Generation
标题:调查小组基于扩散Transformer的文本到音频生成的相对政策优化
链接:https://arxiv.org/abs/2603.01565
摘要:文本到音频(T2 A)生成近年来已经有了相当大的进步,然而现有方法在准确地呈现复杂的文本提示(特别是涉及复杂音频效果的那些文本提示)以及实现精确的文本-音频对齐方面继续面临挑战。虽然先前的方法已经探索了数据增强、显式定时调节和强化学习,但整体合成质量仍然受到限制。在这项工作中,我们实验强化学习,以进一步提高T2 A生成质量,建立在扩散Transformer(DiT)为基础的架构。我们的方法首先采用了一个大的语言模型(LLM)来生成高保真,丰富详细的音频字幕,大大提高了文本音频语义对齐,特别是对于模糊或指定不足的提示。然后,我们应用组相对策略优化(GRPO),最近推出的强化学习算法,微调T2 A模型。通过对不同奖励函数(包括CLAP、KL、FAD及其组合)的系统实验,我们确定了音频合成中有效RL的关键驱动因素,并分析了奖励设计如何影响最终音频质量。实验结果表明,基于GRPO的微调产生大量的收益,在合成保真度和及时遵守。
摘要:Text-to-audio (T2A) generation has advanced considerably in recent years, yet existing methods continue to face challenges in accurately rendering complex text prompts, particularly those involving intricate audio effects, and achieving precise text-audio alignment. While prior approaches have explored data augmentation, explicit timing conditioning, and reinforcement learning, overall synthesis quality remains constrained. In this work, we experiment with reinforcement learning to further enhance T2A generation quality, building on diffusion transformer (DiT)-based architectures. Our method first employs a large language model (LLM) to generate high-fidelity, richly detailed audio captions, substantially improving text-audio semantic alignment, especially for ambiguous or underspecified prompts. We then apply Group Relative Policy Optimization (GRPO), a recently introduced reinforcement learning algorithm, to fine-tune the T2A model. Through systematic experimentation with diverse reward functions (including CLAP, KL, FAD, and their combinations), we identify the key drivers of effective RL in audio synthesis and analyze how reward design impacts final audio quality. Experimental results demonstrate that GRPO-based fine-tuning yield substantial gains in synthesis fidelity and prompt adherence.
【24】VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling
标题:VoxKnesset:用于老化说话人建模的大规模纵向希伯来语语音数据集
链接:https://arxiv.org/abs/2603.01270
备注:4 pages, 5 figures, 2 tables
摘要:语音处理系统面临着一个根本性的挑战:人类的声音随着年龄的增长而变化,但很少有数据集支持严格的纵向评估。我们介绍VoxKnesset,这是一个开放获取的数据集,包含了2009-2025年约2,300小时的希伯来议会演讲,包括393位演讲者,记录跨度长达15年。每一部分都包括来自正式会议记录的经调整的记录誊本和经核实的人口统计元数据。我们基准现代语音嵌入(WavLM大,ECAPA-TDNN,Wav 2 Vec 2-XLSR-1B)的年龄预测和说话人验证纵向条件下。最强模型的说话人确认EER在15年内从2.15%上升到4.58%,并且横截面训练的年龄回归器未能捕获说话人内老化,而纵向训练的模型恢复了有意义的时间信号。我们公开发布数据集和管道,以支持老化鲁棒的语音系统和希伯来语语音处理。
摘要:Speech processing systems face a fundamental challenge: the human voice changes with age, yet few datasets support rigorous longitudinal evaluation. We introduce VoxKnesset, an open-access dataset of ~2,300 hours of Hebrew parliamentary speech spanning 2009-2025, comprising 393 speakers with recording spans of up to 15 years. Each segment includes aligned transcripts and verified demographic metadata from official parliamentary records. We benchmark modern speech embeddings (WavLM-Large, ECAPA-TDNN, Wav2Vec2-XLSR-1B) on age prediction and speaker verification under longitudinal conditions. Speaker verification EER rises from 2.15\% to 4.58\% over 15 years for the strongest model, and cross-sectionally trained age regressors fail to capture within-speaker aging, while longitudinally trained models recover a meaningful temporal signal. We publicly release the dataset and pipeline to support aging-robust speech systems and Hebrew speech processing.
【1】TCG CREST System Description for the DISPLACE-M Challenge
标题:DISPLACE-M挑战赛的TCG CREST系统描述
链接:https://arxiv.org/abs/2603.02030
备注:Report submitted for the DISPLACE-M challenge
摘要:本报告介绍了TCG CREST系统描述的轨道1(扬声器日记)的位移-M的挑战,专注于自然主义的医疗对话在嘈杂的农村医疗保健方案。我们的研究评估了各种语音活动检测(VAD)方法和先进的聚类算法对整体扬声器日志(SD)性能的影响。我们比较和分析了两个SD框架:一个是利用SpeechBrain和ECAPA-TDNN嵌入的模块化管道,另一个是最先进的(SOTA)混合端到端神经日志化系统Diarizen,它建立在预训练的WavLM之上。有了这些框架,我们探索了不同的聚类技术,包括凝聚层次聚类(AHC),和多种新的谱聚类变体,如SC-适应,SC-PNA和SC-MK。实验结果表明,Diarizen系统提供了一个约39\%$的相对改善的日志化错误率(DER)的后评估分析阶段~I相比,SpeechBrain基线。我们提交的性能最好的系统采用Diarizen基线,AHC采用中值滤波,背景窗口更大,为29美元,在开发和评估集上分别实现了10.37%和9.21%的DER。经过第一阶段的评估,我们队在11个参赛队中排名第六。
摘要:This report presents the TCG CREST system description for Track 1 (Speaker Diarization) of the DISPLACE-M challenge, focusing on naturalistic medical conversations in noisy rural-healthcare scenarios. Our study evaluates the impact of various voice activity detection (VAD) methods and advanced clustering algorithms on overall speaker diarization (SD) performance. We compare and analyze two SD frameworks: a modular pipeline utilizing SpeechBrain with ECAPA-TDNN embeddings, and a state-of-the-art (SOTA) hybrid end-to-end neural diarization system, Diarizen, built on top of a pre-trained WavLM. With these frameworks, we explore diverse clustering techniques, including agglomerative hierarchical clustering (AHC), and multiple novel variants of spectral clustering, such as SC-adapt, SC-PNA, and SC-MK. Experimental results demonstrate that the Diarizen system provides an approximate $39\%$ relative improvement in the diarization error rate (DER) on the post-evaluation analysis of Phase~I compared to the SpeechBrain baseline. Our best-performing submitted system employing the Diarizen baseline with AHC employing a median filtering with a larger context window of $29$ achieved a DER of 10.37\% on the development and 9.21\% on the evaluation sets, respectively. Our team ranked sixth out of the 11 participating teams after the Phase~I evaluation.
【2】Investigating Group Relative Policy Optimization for Diffusion Transformer based Text-to-Audio Generation
标题:调查小组基于扩散Transformer的文本到音频生成的相对政策优化
链接:https://arxiv.org/abs/2603.01565
摘要:文本到音频(T2 A)生成近年来已经有了相当大的进步,然而现有方法在准确地呈现复杂的文本提示(特别是涉及复杂音频效果的那些文本提示)以及实现精确的文本-音频对齐方面继续面临挑战。虽然先前的方法已经探索了数据增强、显式定时调节和强化学习,但整体合成质量仍然受到限制。在这项工作中,我们实验强化学习,以进一步提高T2 A生成质量,建立在扩散Transformer(DiT)为基础的架构。我们的方法首先采用了一个大的语言模型(LLM)来生成高保真,丰富详细的音频字幕,大大提高了文本音频语义对齐,特别是对于模糊或指定不足的提示。然后,我们应用组相对策略优化(GRPO),最近推出的强化学习算法,微调T2 A模型。通过对不同奖励函数(包括CLAP、KL、FAD及其组合)的系统实验,我们确定了音频合成中有效RL的关键驱动因素,并分析了奖励设计如何影响最终音频质量。实验结果表明,基于GRPO的微调产生大量的收益,在合成保真度和及时遵守。
摘要:Text-to-audio (T2A) generation has advanced considerably in recent years, yet existing methods continue to face challenges in accurately rendering complex text prompts, particularly those involving intricate audio effects, and achieving precise text-audio alignment. While prior approaches have explored data augmentation, explicit timing conditioning, and reinforcement learning, overall synthesis quality remains constrained. In this work, we experiment with reinforcement learning to further enhance T2A generation quality, building on diffusion transformer (DiT)-based architectures. Our method first employs a large language model (LLM) to generate high-fidelity, richly detailed audio captions, substantially improving text-audio semantic alignment, especially for ambiguous or underspecified prompts. We then apply Group Relative Policy Optimization (GRPO), a recently introduced reinforcement learning algorithm, to fine-tune the T2A model. Through systematic experimentation with diverse reward functions (including CLAP, KL, FAD, and their combinations), we identify the key drivers of effective RL in audio synthesis and analyze how reward design impacts final audio quality. Experimental results demonstrate that GRPO-based fine-tuning yield substantial gains in synthesis fidelity and prompt adherence.
【3】A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection
标题:用于音频深度伪造检测的自监督语音模型的SURB式基准
链接:https://arxiv.org/abs/2603.01482
备注:Accepted at ICASSP
摘要:自我监督学习(SSL)已经改变了语音处理,SUPERB等基准测试在不同的下游任务之间建立了公平的比较。尽管它的安全至关重要,但音频deepfake检测仍然在这些努力之外。在这项工作中,我们介绍了Spoof-SUPERB,这是一个用于音频深度伪造检测的基准,系统地评估了20个SSL模型,这些模型跨越了生成,判别和基于频谱图的架构。我们在多个域内和域外数据集上评估了这些模型。我们的研究结果表明,XLS-R、UniSpeech-SAT和WavLM Large等大规模判别模型的性能始终优于其他模型,这得益于多语言预训练、说话者感知目标和模型规模。我们进一步分析了这些模型在声学退化下的鲁棒性,表明生成方法急剧退化,而判别模型仍然具有弹性。该基准建立了一个可重复的基线,并提供了实用的见解,即SSL表示对于保护语音系统免受音频deepfake影响最可靠。
摘要:Self-supervised learning (SSL) has transformed speech processing, with benchmarks such as SUPERB establishing fair comparisons across diverse downstream tasks. Despite it's security-critical importance, Audio deepfake detection has remained outside these efforts. In this work, we introduce Spoof-SUPERB, a benchmark for audio deepfake detection that systematically evaluates 20 SSL models spanning generative, discriminative, and spectrogram-based architectures. We evaluated these models on multiple in-domain and out-of-domain datasets. Our results reveal that large-scale discriminative models such as XLS-R, UniSpeech-SAT, and WavLM Large consistently outperform other models, benefiting from multilingual pretraining, speaker-aware objectives, and model scale. We further analyze the robustness of these models under acoustic degradations, showing that generative approaches degrade sharply, while discriminative models remain resilient. This benchmark establishes a reproducible baseline and provides practical insights into which SSL representations are most reliable for securing speech systems against audio deepfakes.
【4】Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
标题:用于超低比特率神经语音编解码器的信息引导GRVQ
链接:https://arxiv.org/abs/2603.01476
摘要:神经音频编解码器(NAC)是重建高质量语音信号和为下游语音语言模型生成离散表示的关键。然而,确保准确的语义建模,同时保持高保真度重建超低比特率的约束下仍然具有挑战性。我们提出了一种熵引导的组残差矢量量化(EG-GRVQ)的超低比特率神经语音编解码器,它保留了语义分支的语言信息,并在声学分支中采用了熵引导的分组策略。假设信道激活近似遵循高斯统计,则每个信道的方差可以用作其信息内容的原则代理。基于这一假设,我们对编码器输出进行分区,使得每个组承载相等份额的总信息。这种平衡分配提高了码本效率并减少了冗余。在LibriTTS和VCTK上训练后,我们的模型在超低比特率条件下显示出感知质量和可懂度指标的改善,重点是面向通信场景的编解码器级保真度。
摘要:Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensuring accurate semantic modeling while maintaining high-fidelity reconstruction under ultra-low bitrate constraints remains challenging. We propose an entropy-guided group residual vector quantization (EG-GRVQ) for an ultra-low bitrate neural speech codec, which retains a semantic branch for linguistic information and incorporates an entropy-guided grouping strategy in the acoustic branch. Assuming that channel activations follow approximately Gaussian statistics, the variance of each channel can serve as a principled proxy for its information content. Based on this assumption, we partition the encoder output such that each group carries an equal share of the total information. This balanced allocation improves codebook efficiency and reduces redundancy. Trained on LibriTTS and VCTK, our model shows improvements in perceptual quality and intelligibility metrics under ultra-low bitrate conditions, with a focus on codec-level fidelity for communication-oriented scenarios.
【5】Conversational Speech Naturalness Predictor
标题:会话语音自然度预测器
链接:https://arxiv.org/abs/2603.01467
备注:Under review for Interspeech 2026
摘要:会话自然度评价是开发类人语音智能体的关键。然而,现有的语音自然度预测器通常被设计为评估来自单个说话者的话语,未能捕获会话级的自然度质量。在本文中,我们提出了一个自动自然度预测两个扬声器,多轮对话的框架。我们首先表明,现有的自然度估计有低,有时甚至是负的,与会话的自然度的相关性,根据会话记录注释与人类的评级。然后,我们提出了一个双通道自然度估计器,其中我们研究了多个预训练编码器和数据增强。我们提出的模型实现了更高的相关性与人类的判断相比,现有的自然度预测域内和域外的条件。
摘要:Evaluation of conversational naturalness is essential for developing human-like speech agents. However, existing speech naturalness predictors are often designed to assess utterances from a single speaker, failing to capture conversation-level naturalness qualities. In this paper, we present a framework for an automatic naturalness predictor for two-speaker, multi-turn conversations. We first show that existing naturalness estimators have low, or sometimes even negative, correlations with conversational naturalness, based on conversational recordings annotated with human ratings. We then propose a dual-channel naturalness estimator, in which we investigate multiple pre-trained encoders with data augmentation. Our proposed model achieves substantially higher correlation with human judgments compared to existing naturalness predictors for both in-domain and out-of-domain conditions.
【6】The USTC-NERCSLIP Systems for the CHiME-9 MCoRec Challenge
标题:CHiME-9 MCoRec挑战赛的USTC-NERCSLIP系统
链接:https://arxiv.org/abs/2603.01415
摘要:这份报告详细介绍了我们提交给CHiME-9 MCoRec挑战赛的关于在室内社交环境中识别和聚类多个并发自然对话的信息。与传统的以单一主题为中心的会议不同,这种场景包含多个并行对话-最多八个发言者,最多四个同时进行的对话-语音重叠率超过90%。为了解决这个问题,我们提出了一个多模式级联系统,利用每个扬声器的视觉流提取同步的360度视频与单声道音频。我们的系统通过利用增强的视听预训练模型改进了管道的三个组件:主动说话人检测(ASD),视听目标语音提取(AVTSE)和视听语音识别(AVSR)。AVSR模块还结合了Whisper和LLM技术,以提高转录准确性。我们最好的单级联系统在开发集上实现了32.44%的扬声器字错误率(WER)。通过进一步应用ROVER融合来自不同前端和后端变体的输出,我们将扬声器WER降低到31.40%。值得注意的是,我们基于LLM的zero-shot会话聚类实现了1.0的说话者聚类F1得分,产生了15.70%的最终联合ASR聚类错误率(JACER)。
摘要:This report details our submission to the CHiME-9 MCoRec Challenge on recognizing and clustering multiple concurrent natural conversations within indoor social settings. Unlike conventional meetings centered on a single shared topic, this scenario contains multiple parallel dialogues--up to eight speakers across up to four simultaneous conversations--with a speech overlap rate exceeding 90%. To tackle this, we propose a multimodal cascaded system that leverages per-speaker visual streams extracted from synchronized 360 degree video together with single-channel audio. Our system improves three components of the pipeline by leveraging enhanced audio-visual pretrained models: Active Speaker Detection (ASD), Audio-Visual Target Speech Extraction (AVTSE), and Audio-Visual Speech Recognition (AVSR). The AVSR module further incorporates Whisper and LLM techniques to boost transcription accuracy. Our best single cascaded system achieves a Speaker Word Error Rate (WER) of 32.44% on the development set. By further applying ROVER to fuse outputs from diverse front-end and back-end variants, we reduce Speaker WER to 31.40%. Notably, our LLM-based zero-shot conversational clustering achieves a speaker clustering F1 score of 1.0, yielding a final Joint ASR-Clustering Error Rate (JACER) of 15.70%.
【7】Inter-Speaker Relative Cues for Two-Stage Text-Guided Target Speech Extraction
标题:两阶段文本引导目标语音提取的说话者间相对线索
链接:https://arxiv.org/abs/2603.01316
摘要:本文研究了在基于文本的目标语音提取(TSE)中使用相关线索。我们首先从人类感知和标签量化的角度为相对线索提供了理论依据,表明相对线索保留了在绝对分类表示中经常丢失的细粒度区分。在此分析的基础上,我们提出了一个两阶段的TSE框架,其中语音分离模型生成候选源,然后由文本引导的分类器,选择基于嵌入相似性的目标扬声器。使用这个框架,我们训练两个独立的分类模型,以评估的优势,相对于独立的线索在分类精度和TSE性能。实验结果表明,(i)与独立线索相比,相对线索实现了更高的整体分类准确率和改进的TSE性能,(ii)两阶段框架在信号级和客观感知度量上都大大优于单阶段文本条件提取方法,以及(iii)某些相对线索(语言、性别、响度、距离、时间顺序、说话持续时间、随机线索和所有线索)的性能可以超过基于音频的TSE系统的性能。进一步的分析揭示了显着的差异,在不同的线索类型的区分能力,提供不同的相对线索TSE的有效性的见解。
摘要:This paper investigates the use of relative cues for text-based target speech extraction (TSE). We first provide a theoretical justification for relative cues from the perspectives of human perception and label quantization, showing that relative cues preserve fine-grained distinctions often lost in absolute categorical representations. Building on this analysis, we propose a two-stage TSE framework, in which a speech separation model generates candidate sources, followed by a text-guided classifier that selects the target speaker based on embedding similarity. Using this framework, we train two separate classification models to evaluate the advantages of relative cues over independent cues in terms of both classification accuracy and TSE performance. Experimental results demonstrate that (i) relative cues achieve higher overall classification accuracy and improved TSE performance compared with independent cues, (ii) the two-stage framework substantially outperforms single-stage text-conditioned extraction methods on both signal-level and objective perceptual metrics, and (iii) certain relative cues (language, gender, loudness, distance, temporal order, speaking duration, random cue and all cue) can surpass the performance of an audio-based TSE system. Further analysis reveals notable differences in discriminative power across cue types, providing insights into the effectiveness of different relative cues for TSE.
【8】VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling
标题:VoxKnesset:用于老化说话人建模的大规模纵向希伯来语语音数据集
链接:https://arxiv.org/abs/2603.01270
备注:4 pages, 5 figures, 2 tables
摘要:语音处理系统面临着一个根本性的挑战:人类的声音随着年龄的增长而变化,但很少有数据集支持严格的纵向评估。我们介绍VoxKnesset,这是一个开放获取的数据集,包含了2009-2025年约2,300小时的希伯来议会演讲,包括393位演讲者,记录跨度长达15年。每一部分都包括来自正式会议记录的经调整的记录誊本和经核实的人口统计元数据。我们基准现代语音嵌入(WavLM大,ECAPA-TDNN,Wav 2 Vec 2-XLSR-1B)的年龄预测和说话人验证纵向条件下。最强模型的说话人确认EER在15年内从2.15%上升到4.58%,并且横截面训练的年龄回归器未能捕获说话人内老化,而纵向训练的模型恢复了有意义的时间信号。我们公开发布数据集和管道,以支持老化鲁棒的语音系统和希伯来语语音处理。
摘要:Speech processing systems face a fundamental challenge: the human voice changes with age, yet few datasets support rigorous longitudinal evaluation. We introduce VoxKnesset, an open-access dataset of ~2,300 hours of Hebrew parliamentary speech spanning 2009-2025, comprising 393 speakers with recording spans of up to 15 years. Each segment includes aligned transcripts and verified demographic metadata from official parliamentary records. We benchmark modern speech embeddings (WavLM-Large, ECAPA-TDNN, Wav2Vec2-XLSR-1B) on age prediction and speaker verification under longitudinal conditions. Speaker verification EER rises from 2.15\% to 4.58\% over 15 years for the strongest model, and cross-sectionally trained age regressors fail to capture within-speaker aging, while longitudinally trained models recover a meaningful temporal signal. We publicly release the dataset and pipeline to support aging-robust speech systems and Hebrew speech processing.
【9】Using Songs to Improve Kazakh Automatic Speech Recognition
标题:利用歌曲提高哈萨克语自动语音识别
链接:https://arxiv.org/abs/2603.00961
备注:9 pages, 7 tables, to appear in Proceedings of the 2026 Language Resources and Evaluation Conference
摘要:开发低资源语言的自动语音识别(ASR)系统受到转录语料库稀缺的阻碍。这项概念验证研究探讨了歌曲作为哈萨克语ASR的非传统但有前途的数据源。我们从36位艺术家的195首歌曲中挑选了3,013个音频文本对(约4.5小时)的数据集,在歌词行级别进行分段。使用Whisper作为基本识别器,我们在涉及歌曲,通用语音语料库(CVC)和FLEURS的七种训练场景下微调模型,并在三个基准上进行评估:CVC,FLEURS和哈萨克语语音语料库2(KSC 2)。结果表明,基于歌曲的微调提高了性能超过zero-shot基线。例如,在歌曲,CVC和FLEURS的混合物上训练的Whisper Large-V3 Turbo在CVC上实现了27.6%的归一化WER,在FLEURS上实现了11.8%的归一化WER,同时相对于zero-shot模型,KSC 2上的误差减半(39.3% vs. 81.2%)。尽管这些增益仍然低于在1,100小时KSC 2语料库上训练的模型,但它们表明,即使是适度的歌曲-语音混合也可以在低资源ASR中产生有意义的适应性改进。该数据集在Hugging Face上发布,用于研究目的,根据非商业许可证。
摘要:Developing automatic speech recognition (ASR) systems for low-resource languages is hindered by the scarcity of transcribed corpora. This proof-of-concept study explores songs as an unconventional yet promising data source for Kazakh ASR. We curate a dataset of 3,013 audio-text pairs (about 4.5 hours) from 195 songs by 36 artists, segmented at the lyric-line level. Using Whisper as the base recogniser, we fine-tune models under seven training scenarios involving Songs, Common Voice Corpus (CVC), and FLEURS, and evaluate them on three benchmarks: CVC, FLEURS, and Kazakh Speech Corpus 2 (KSC2). Results show that song-based fine-tuning improves performance over zero-shot baselines. For instance, Whisper Large-V3 Turbo trained on a mixture of Songs, CVC, and FLEURS achieves 27.6% normalised WER on CVC and 11.8% on FLEURS, while halving the error on KSC2 (39.3% vs. 81.2%) relative to the zero-shot model. Although these gains remain below those of models trained on the 1,100-hour KSC2 corpus, they demonstrate that even modest song-speech mixtures can yield meaningful adaptation improvements in low-resource ASR. The dataset is released on Hugging Face for research purposes under a gated, non-commercial licence.
【10】Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
标题:情态差距剖析:剖析端到端言语LLM的内部状态
链接:https://arxiv.org/abs/2603.01502
摘要:大型语音语言模型的最新进展大大弥合了声学信号和语言理解之间的差距。然而,一个持久的性能差距仍然在基于语音的输入任务相比,直接文本推理。在本文中,我们调查的动态根源,这种模态差距超出静态几何对齐,分析如何语音和文本表示演变逐层。我们在SpeechMMLU和VoiceBench BBH上评估了四个开放权重的端到端模型。使用跨层CKA分析与语音文本标记对齐,我们发现,语音表示表现出广泛的跨层对齐带,由于语音的冗余性,语义内容跨越多个帧。我们表明,这些对齐模式在不同的分析配置结构稳定。至关重要的是,简单的统计校准是不够的,并且当应用于输入层时可能是有害的,这表明模态间隙不仅仅是分布偏移。总的来说,我们的研究结果表明,瓶颈在于将冗余语音压缩成稳定的后期层决策,激励未来的解决方案在令牌或时间粒度上操作,而不是特征级匹配。
摘要:Recent advancements in Large Speech-Language Models have significantly bridged the gap between acoustic signals and linguistic understanding. However, a persistent performance disparity remains in speech-based input tasks compared to direct text inference. In this paper, we investigate the dynamic roots of this modality gap beyond static geometric alignment, analyzing how speech and text representations evolve layer-by-layer. We evaluate four open-weight end-to-end models on SpeechMMLU and VoiceBench BBH. Using cross-layer CKA analysis with speech-text token alignment, we find that speech representations exhibit a broad cross-layer alignment band, attributable to the redundant nature of speech where semantic content spans multiple frames. We show that these alignment patterns are structurally stable across different analysis configurations. Crucially, simple statistical calibration is insufficient and can be detrimental when applied at the input layer, indicating that the modality gap is not a mere distribution shift. Overall, our results suggest that the bottleneck lies in condensing redundant speech into stable late-layer decisions, motivating future solutions that operate at the token or temporal granularity instead of feature-level matching.
【11】CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
标题:CMI-RewardBench:使用合成多模式教学评估音乐奖励模型
链接:https://arxiv.org/abs/2603.00610
摘要:虽然音乐生成模型已经发展到可以处理混合了文本、歌词和参考音频的复杂多模态输入,但评估机制却落后了。在本文中,我们通过建立一个综合的生态系统,在作曲多模态教学(CMI)下的音乐奖励建模,其中生成的音乐可能会以文本描述,歌词和音频提示为条件,来弥合这一关键差距。我们首先介绍CMI-Pref-Pseudo,一个包含110 k伪标记样本的大规模偏好数据集,以及CMI-Pref,一个为细粒度对齐任务量身定制的高质量人工注释语料库。为了统一评估环境,我们提出了CMI-RewardBench,这是一个统一的基准,可以在音乐性,文本音乐对齐和作曲指令对齐的异构样本上评估音乐奖励模型。利用这些资源,我们开发了CMI奖励模型(CMI-RM),一个参数有效的奖励模型家族能够处理异构输入。我们评估了它们与人类对音乐性的判断分数的相关性,以及与以前的数据集在CMI-Pref上的对齐。进一步的实验表明,CMI-RM不仅与人类的判断密切相关,而且还可以通过top-k滤波实现有效的推理时间缩放。必要的训练数据、基准和奖励模型是公开的。
摘要:While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive ecosystem for music reward modeling under Compositional Multimodal Instruction (CMI), where the generated music may be conditioned on text descriptions, lyrics, and audio prompts. We first introduce CMI-Pref-Pseudo, a large-scale preference dataset comprising 110k pseudo-labeled samples, and CMI-Pref, a high-quality, human-annotated corpus tailored for fine-grained alignment tasks. To unify the evaluation landscape, we propose CMI-RewardBench, a unified benchmark that evaluates music reward models on heterogeneous samples across musicality, text-music alignment, and compositional instruction alignment. Leveraging these resources, we develop CMI reward models (CMI-RMs), a parameter-efficient reward model family capable of processing heterogeneous inputs. We evaluate their correlation with human judgments scores on musicality and alignment on CMI-Pref along with previous datasets. Further experiments demonstrate that CMI-RM not only correlates strongly with human judgments, but also enables effective inference-time scaling via top-k filtering. The necessary training data, benchmarks, and reward models are publicly available.
【12】Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
标题:文明之声:全球音乐理解的多语言QA基准
链接:https://arxiv.org/abs/2603.00533
备注:2 pages, 2 figures, 1 table, accepted by ISMIR 2025 LBD
摘要:我们介绍了文明的声音,第一个多语言QA基准评估音频LLM的文化理解全长音乐录音。覆盖38种语言的380首曲目,我们的自动化管道通过四个阶段产生1,190个多项选择题-每个阶段都需要人工验证:1)编制代表性音乐列表; 2)通过LLM为音乐列表中的每个样本生成文化背景文档; 3)从这些文档中提取关键属性;(4)设计多项选择题,探讨语言、地域联想、语气和主题内容。我们在四种条件下评估模型,并报告每种语言的准确性。我们的研究结果表明,即使是最先进的音频LLM也很难在没有丰富文本背景的情况下捕捉微妙的文化细微差别,并在解释来自不同文化传统的音乐时表现出系统性偏见。该数据集在Hugging Face上公开,以促进文化包容性的音乐理解研究。
摘要:We introduce Voices of Civilizations, the first multilingual QA benchmark for evaluating audio LLMs' cultural comprehension on full-length music recordings. Covering 380 tracks across 38 languages, our automated pipeline yields 1,190 multiple-choice questions through four stages - each followed by manual verification: 1) compiling a representative music list; 2) generating cultural-background documents for each sample in the music list via LLMs; 3) extracting key attributes from those documents; and 4) constructing multiple-choice questions probing language, region associations, mood, and thematic content. We evaluate models under four conditions and report per-language accuracy. Our findings demonstrate that even state-of-the-art audio LLMs struggle to capture subtle cultural nuances without rich textual context and exhibit systematic biases in interpreting music from different cultural traditions. The dataset is publicly available on Hugging Face to foster culturally inclusive music understanding research.
【13】Aurchestra: Fine-Grained, Real-Time Soundscape Control on Resource-Constrained Hearables
标题:Aurchestra:对资源受限的听力进行细粒度、实时音景控制
链接:https://arxiv.org/abs/2603.00395
备注:15 pages, 11 figures, 4 tables, submitted to ACM MobiSys 2026
摘要:可听设备正变得无处不在,但它们的声音控制仍然很迟钝:用户可以启用全局噪声抑制或专注于单个目标声音。然而,真实世界的声学场景包含许多用户可能想要独立调整的同时源。我们介绍了Aurchestra,第一个系统提供细粒度的,实时的声音景观控制资源受限的听觉。我们的系统有两个关键组件:(1)一个动态界面,只显示活动的声音类别;(2)一个实时的设备上多输出提取网络,为每个选定的类别生成单独的流,为多达5个重叠的目标声音实现强大的性能,并让用户通过自定义每个类别的音量来混合他们的环境,就像音频工程师混合曲目一样。我们优化了多个计算受限平台的模型架构,并在6 ms流式音频块上展示了实时性能。在以前看不见的室内和室外场景中的真实环境中,我们的系统能够实现富有表现力的每类声音控制,并在目标类增强和干扰抑制方面实现了实质性的改进。我们的研究结果表明,世界不需要作为一个单一的,无差别的流听到:与Aurchestra,声景变得真正可编程。
摘要:Hearables are becoming ubiquitous, yet their sound controls remain blunt: users can either enable global noise suppression or focus on a single target sound. Real-world acoustic scenes, however, contain many simultaneous sources that users may want to adjust independently. We introduce Aurchestra, the first system to provide fine-grained, real-time soundscape control on resource-constrained hearables. Our system has two key components: (1) a dynamic interface that surfaces only active sound classes and (2) a real-time, on-device multi-output extraction network that generates separate streams for each selected class, achieving robust performance for upto 5 overlapping target sounds, and letting users mix their environment by customizing per-class volumes, much like an audio engineer mixes tracks. We optimize the model architecture for multiple compute-limited platforms and demonstrate real-time performance on 6 ms streaming audio chunks. Across real-world environments in previously unseen indoor and outdoor scenarios, our system enables expressive per-class sound control and achieves substantial improvements in target-class enhancement and interference suppression. Our results show that the world need not be heard as a single, undifferentiated stream: with Aurchestra, the soundscape becomes truly programmable.
【14】StethoLM: Audio Language Model for Cardiopulmonary Analysis Across Clinical Tasks
标题:StethoLM:跨临床任务心肺分析的音频语言模型
链接:https://arxiv.org/abs/2603.00355
备注:To be published in TMLR
摘要:听心脏和肺的声音-听诊-是临床检查中最基本的步骤之一。尽管快速且无创,但它需要多年的经验来解释微妙的音频提示。最近的深度学习方法在自动化心肺音分析方面取得了进展,但大多数都局限于简单的分类,几乎没有提供临床可解释性或决策支持。我们提出StethoLM,第一个专门用于心肺听诊的音频语言模型,能够在整个听诊分析范围内执行预防驱动的临床任务。StethoLM将音频编码与医学语言模型主干集成在一起,并在StethoBench上进行训练,StethoBench是一个全面的基准测试,包括从16,125个标记的心肺记录中合成的77,027个预防-反应对,涵盖七个临床任务类别:二元分类,检测,报告,推理,鉴别诊断,比较和基于位置的分析。通过结合监督微调和直接偏好优化的多阶段训练,StethoLM在非分布数据上实现了性能和鲁棒性的大幅提升。我们的工作为临床听诊中的预防跟踪AI系统奠定了基础。
摘要:Listening to heart and lung sounds - auscultation - is one of the first and most fundamental steps in a clinical examination. Despite being fast and non-invasive, it demands years of experience to interpret subtle audio cues. Recent deep learning methods have made progress in automating cardiopulmonary sound analysis, yet most are restricted to simple classification and offer little clinical interpretability or decision support. We present StethoLM, the first audio-language model specialized for cardiopulmonary auscultation, capable of performing instruction-driven clinical tasks across the full spectrum of auscultation analysis. StethoLM integrates audio encoding with a medical language model backbone and is trained on StethoBench, a comprehensive benchmark comprising 77,027 instruction-response pairs synthesized from 16,125 labeled cardiopulmonary recordings spanning seven clinical task categories: binary classification, detection, reporting, reasoning, differential diagnosis, comparison, and location-based analysis. Through multi-stage training that combines supervised fine-tuning and direct preference optimization, StethoLM achieves substantial gains in performance and robustness on out-of-distribution data. Our work establishes a foundation for instruction-following AI systems in clinical auscultation.
【15】Iterative LLM-based improvement for French Clinical Interview Transcription and Speaker Diarization
标题:基于LLM的法语临床访谈转录和演讲者拨号的迭代改进
链接:https://arxiv.org/abs/2603.00086
摘要:法语医学对话的自动语音识别仍然具有挑战性,自发临床语音的单词错误率通常超过30%。本研究提出了一种多通道LLM后处理架构,在说话人识别和单词识别通道之间交替,以提高转录准确性和说话人属性。消融研究两个法国临床数据集(自杀预防电话咨询和术前清醒神经外科咨询)调查四个设计选择:模型选择,提示策略,通过排序,迭代深度。使用Qwen 3-Next-80 B,Wilcoxon符号秩检验证实了自杀预防对话的WDER显著降低(p < 0.05,n=18),同时保持清醒神经外科会诊的稳定性(n=10),零输出故障和可接受的计算成本(RTF 0.32),表明离线临床部署的可行性。
摘要:Automatic speech recognition for French medical conversations remains challenging, with word error rates often exceeding 30% in spontaneous clinical speech. This study proposes a multi-pass LLM post-processing architecture alternating between Speaker Recognition and Word Recognition passes to improve transcription accuracy and speaker attribution. Ablation studies on two French clinical datasets (suicide prevention telephone counseling and preoperative awake neurosurgery consultations) investigate four design choices: model selection, prompting strategy, pass ordering, and iteration depth. Using Qwen3-Next-80B, Wilcoxon signed-rank tests confirm significant WDER reductions on suicide prevention conversations (p < 0.05, n=18), while maintaining stability on awake neurosurgery consultations (n=10), with zero output failures and acceptable computational cost (RTF 0.32), suggesting feasibility for offline clinical deployment.
机器翻译由腾讯交互翻译提供,仅供参考
