微信公众号:arXiv_Daily
cs.SD语音
【1】SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
标题:SpeakerLLM:一种用于说话人理解和验证推理的演讲人专业音频LLM
链接:https://arxiv.org/pdf/2605.15044v1
摘要:随着音频优先代理在物理AI、会话机器人和无屏可穿戴设备中越来越普遍,音频大型语言模型(audio-LLM)必须集成特定于说话者的理解,以支持用户授权、个性化和上下文感知交互。这需要建模谁在说话,声音听起来如何,以及录音条件如何影响扬声器提示。传统的说话人验证系统提供了强大的标量分数,但很少的语言证据,而目前的音频LLM和说话人感知语言模型有有限的能力,组织说话人信息超出二进制标签或描述性配置文件。我们提出了SpeakerLLM,一个扬声器专用的音频LLM框架,它统一了单话语扬声器分析,记录条件理解,话语对扬声器比较和自然语言界面内的证据组织验证推理。我们构建了验证推理目标和决策组合策略,该策略将配置文件级别的证据与最终相同或不同的决策分开,并将记录条件、配置文件证据和决策组织成结构化的轨迹。在其核心,SpeakerLLM使用分层扬声器标记器,旨在捕获扬声器证据的多个粒度。话语级的扬声器嵌入总结身份和配置文件级别的线索,而帧级的扬声器功能保留细粒度的声学描述符。实验表明,SpeakerLLM-Base比一般音频LLM提高了说话者个人资料和录音条件的理解,而SpeakerLLM-VR保留了较强的生成判决准确性,并产生基于监督验证推理模式的决策轨迹。我们将发布元数据丰富的监督数据集和目标构建代码,以实现可重复性。摘要:As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-specific understanding to support user authorization, personalization, and context-aware interaction. This requires modeling who is speaking, how the voice sounds, and how recording conditions affect speaker cues. Conventional speaker verification systems provide strong scalar scores but little linguistic evidence, while current audio-LLMs and speaker-aware language models have limited ability to organize speaker information beyond binary labels or descriptive profiles. We present SpeakerLLM, a speaker-specialized audio-LLM framework that unifies single-utterance speaker profiling, recording-condition understanding, utterance-pair speaker comparison, and evidence-organized verification reasoning within a natural-language interface. We construct verification-reasoning targets and a decision-composition policy that separate profile-level evidence from the final same-or-different decision and organize recording condition, profile evidence, and the decision into a structured trace. At its core, SpeakerLLM uses a hierarchical speaker tokenizer designed to capture multiple granularities of speaker evidence. Utterance-level speaker embeddings summarize identity and profile-level cues, whereas frame-level speaker features preserve fine-grained acoustic descriptors. Experiments show that SpeakerLLM-Base improves speaker-profile and recording-condition understanding over general audio-LLMs, while SpeakerLLM-VR preserves strong generated-verdict accuracy and produces decision traces grounded in the supervised verification reasoning schema. We will release the metadata-enriched supervision dataset and target-construction code for reproducibility.
【2】Text-Dependent Speaker Verification (TdSV) Challenge 2024: Team Naive System Report
标题:2024年文本相关说话者验证(TdSV)挑战:Team Naive系统报告
链接:https://arxiv.org/pdf/2605.14896v1
摘要:本文介绍了2024年文本相关说话人验证(TdSV)挑战赛的系统。该系统实现了0.0461的最小检测成本函数(MinDCF)和1.3%的等错误率(EER)。我们的方法专注于适应现有的最先进的神经网络,ResNet-TDNN和NeXt-TDNN,最初在VoxCeleb数据集上训练。选择该策略是因为挑战持续时间有限,并且当时可用资源有限。此外,我们还设计了一个轻量级和资源高效的模型EfficientNet-A0,专门针对挑战数据集进行训练,以提高适应性并加强集成方法。我们的系统结合了先进的神经架构,广泛的数据增强和优化的超参数。这些组件有助于在文本相关的说话人确认中实现强大的性能。实验结果也证明了多模型集成学习在说话人和短语确认中的有效性。摘要:This paper presents a system for the 2024 Text-Dependent Speaker Verification (TdSV) Challenge. The system achieved a Minimum Detection Cost Function (MinDCF) of 0.0461 and an Equal Error Rate (EER) of 1.3 %. Our approach focused on adapting existing state-of-the-art neural networks, ResNet-TDNN and NeXt-TDNN, originally trained on the VoxCeleb dataset. This strategy was chosen because of the limited challenge duration and the available resources at the time. In addition, we designed a lightweight and resource-efficient model, EfficientNet-A0, trained specifically on the challenge dataset to improve adaptation and strengthen the ensemble approach. Our system combines advanced neural architectures, extensive data augmentation, and optimised hyperparameters. These components helped achieve strong performance in text-dependent speaker verification. The results also demonstrate the effectiveness of multi-model ensemble learning for both speaker and phrase verification.
【3】PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection
标题:Process-2:用于早期认知障碍检测的基准言语库
链接:https://arxiv.org/pdf/2605.14888v1
摘要:基于语音的分析为检测认知能力下降提供了一种可扩展的非侵入性方法,但进展受到现实条件下收集的临床验证数据集有限的限制。我们介绍PROCESS-2,一个大规模的语音数据集,旨在支持研究自动评估认知障碍的自发和面向任务的语音。该数据集包括使用CognoMemory数字评估平台收集的200名健康对照,150名轻度认知障碍和50名痴呆症诊断的记录。每个参与者都完成了一次评估,包括图片描述和语言流畅性任务,并附有手动验证的成绩单和参与者级别的元数据。PROCESS-2包含约21小时的语音音频,具有预定义的训练 测试分区。全面的技术验证评估了人口统计学平衡、临床一致性、记录稳定性、嵌入空间结构和可重现的基线建模性能,证明了具有临床意义的组分离和建模方法的稳定性能,同时保留了真实世界的会话变异性。PROCESS-2通过Hugging Face在受控访问下发布,以实现负责任的重用,同时保护参与者隐私,为基于语音的认知评估研究提供可复制的基准资源。摘要:Speech-based analysis offers a scalable and non-invasive approach for detecting cognitive decline, yet progress has been constrained by the limited availability of clinically validated datasets collected under realistic conditions. We introduce PROCESS-2, a large-scale speech dataset designed to support research on automatic assessment of cognitive impairment from spontaneous and task-oriented speech. The dataset comprises recordings from 200 healthy controls, 150 mild cognitive impairment, and 50 dementia diagnoses collected using the CognoMemory digital assessment platform. Each participant completed a single assessment session, including picture description and verbal fluency tasks, accompanied by manually verified transcripts and participant-level metadata. PROCESS-2 contains approximately 21 hours of speech audio with predefined train test partitions. Comprehensive technical validation evaluated demographic balance, clinical consistency, recording stability, embedding-space structure, and reproducible baseline modelling performance, demonstrating clinically meaningful group separation and stable performance across modelling approaches while preserving real-world conversational variability. PROCESS-2 is released under controlled access via Hugging Face to enable responsible reuse while protecting participant privacy, providing a reproducible benchmark resource for speech-based cognitive assessment research.
【4】Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music
标题:波斯音乐世代:波斯音乐的大规模数据集和文化感知生成模型
链接:https://arxiv.org/pdf/2605.14765v1
备注:9 pages, 2 figures, 3 tables
摘要:波斯音乐以其独特的音调、调式系统(Dastgah)和节奏结构,对主要接受西方音乐训练的音乐生成模型提出了重大挑战。我们通过策划第一个大规模的波斯歌曲数据集来解决这一差距,该数据集包括900多个小时的高质量音频样本,包括流行音乐,传统和当代风格。该数据集捕捉了波斯音乐丰富的旋律和文化多样性,并作为微调MusicGen的基础,MusicGen是最先进的生成音乐模型。我们适应MusicGen这个领域,并利用主观和客观的指标来评估其性能。为了评估生成的音乐和预期的风格标签之间的语义对齐,我们报告了在生成的输出中准确反映的相关标签的比例。我们的研究结果表明,微调模型产生的组合物,更符合波斯文体惯例。这项工作介绍了生成音乐研究的新资源,并说明了音乐生成模型的适应性不足的文化和语言环境。摘要:Persian music, with its unique tonalities, modal systems (Dastgah), and rhythmic structures, presents significant challenges for music generation models trained primarily on Western music. We address this gap by curating the first large-scale dataset of Persian songs, comprising over 900 hours high-quality audio samples across diverse sub-genres, including pop, traditional, and contemporary styles. This dataset captures the rich melodic and cultural diversity of Persian music and serves as the foundation for fine-tuning MusicGen, a state-of-the-art generative music model. We adapt MusicGen to this domain and evaluate its performance by utilizing subjective and objective metrics. To assess the semantic alignment between generated music and intended style tags, we report the proportion of relevant tags accurately reflected in the generated outputs. Our results demonstrate that the fine-tuned model produces compositions that more align with Persian stylistic conventions. This work introduces a new resource for generative music research and illustrates the adaptability of music generation models to underrepresented cultural and linguistic contexts.
【5】IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments
标题:IsoNet:复杂声学环境中的空间感知视听目标语音提取
链接:https://arxiv.org/pdf/2605.14736v1
备注:8 pages
摘要:目标语音提取仍然很难紧凑的设备,因为单耳神经模型缺乏空间的证据和经典的波束形成器失去了解决能力时,麦克风孔径只有几厘米。我们提出了IsoNet,一个用户可选择的视听目标语音提取系统的紧凑型4麦克风阵列。IsoNet将复杂的多通道STFT特征、GCC-PHAT空间线索、面部条件视觉嵌入和辅助到达方向监督结合在U-Net掩码估计网络中。三个课程变体在25,000个模拟VoxCeleb混合物上进行了训练,SNR制度逐渐困难。在-1至10 dB SNR的硬测试集上,IsoNet-CL 1实现了9.31 dB SI-SDR,比混合物提高了4.85 dB,PESQ为2.13,STOI为0.84。Oracle延迟求和和MVDR波束形成器分别将相同的混合物降低4.82 dB和6.08 dB SI-SDRi,表明所提出的学习多模态调节解决了传统空间滤波无效的情况。消融研究表明,一致的收益从视觉条件反射,GCC-PHAT功能,和扩展的延迟箱编码。结果建立了一个紧凑的阵列,面部可选的语音提取基线控制下的模拟,并确定剩余的障碍,真正的部署,特别是相位重建,多重混合物,和模拟到真实的传输。摘要:Target speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. We present IsoNet, a user-selectable audio-visual target speech extraction system for a compact 4-microphone array. IsoNet combines complex multi-channel STFT features, GCC-PHAT spatial cues, face-conditioned visual embeddings, and auxiliary direction-of-arrival supervision inside a U-Net mask estimation network. Three curriculum variants were trained on 25,000 simulated VoxCeleb mixtures with progressively difficult SNR regimes. On a hard test set spanning -1 to 10 dB SNR, IsoNet-CL1 achieves 9.31 dB SI-SDR, a 4.85 dB improvement over the mixture, with PESQ 2.13 and STOI 0.84. Oracle delay-and-sum and MVDR beamformers degrade the same mixtures by 4.82 dB and 6.08 dB SI-SDRi, respectively, showing that the proposed learned multimodal conditioning solves a regime where conventional spatial filtering is ineffective. Ablation studies show consistent gains from visual conditioning, GCC-PHAT features, and extended delay-bin encoding. The results establish a compact-array, face-selectable speech extraction baseline under controlled simulation and identify the remaining barriers to real deployment, especially phase reconstruction, multi-interferer mixtures, and simulation-to-real transfer.
【6】UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars
标题:Umo:实时同声语音化身的统一稀疏运动建模
链接:https://arxiv.org/pdf/2605.14731v1
摘要:语音驱动的手势和面部动画是游戏、虚拟制作和交互式媒体中富有表现力的数字化身的基础。然而,现有的方法要么局限于单一模态的音频运动对齐,未能充分利用海量人体运动数据的潜力,或受到多模态模型的表示能力和吞吐量的限制,这使得难以实现高质量的运动生成或实时性能。我们提出了UMO,一个统一的稀疏运动建模架构的实时语音头像,它处理文本,音频和运动令牌在一个统一的配方。利用空间稀疏的Mixture-of-Experts框架和时间稀疏的以关键帧为中心的设计,UMo有效地执行实时密集重建,从而为面部表情和手势生成时间连贯和高保真的动画。此外,我们实施了一个多阶段的训练策略,有针对性的音频增强,以提高声学多样性和语义一致性。因此,即使在严格的延迟约束下,UMo也能保持细粒度的语音运动对齐。广泛的定量和定性评估表明,Umo在低延迟和实时性能约束下实现了更好的输出质量,为高保真实时语音化身提供了一个实用的解决方案。摘要:Speech-driven gestures and facial animations are fundamental to expressive digital avatars in games, virtual production, and interactive media. However, existing methods are either limited to a single modality for audio motion alignment, failing to fully utilize the potential of massive human motion data, or are constrained by the representation ability and throughput of multimodal models, which makes it difficult to achieve high-quality motion generation or real-time performance. We present UMo, a unified sparse motion modeling architecture for real-time co-speech avatars, which processes text, audio, and motion tokens within a unified formulation. Leveraging a spatially sparse Mixture-of-Experts framework and a temporally sparse, keyframe-centric design, UMo efficiently performs real-time dense reconstruction, enabling temporally coherent and high-fidelity animation generation for both facial expressions and gestures. Furthermore, we implement a multi-stage training strategy with targeted audio augmentation to enhance acoustic diversity and semantic consistency. Consequently, UMo preserves fine-grained speech-motion alignment even under strict latency constraints. Extensive quantitative and qualitative evaluations show that UMo achieves better output quality under low latency and real-time performance constraints, offering a practical solution for high-fidelity real-time co-speech avatars.
【7】Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis
标题:打破节奏!可控的MIDI到鼓音频合成
链接:https://arxiv.org/pdf/2605.14555v1
摘要:用于在数字音乐制作中创建鼓循环音频的当前方法(诸如使用单次采样或重采样)通常需要创作者的不平凡的努力。虽然最近的生成模型实现了高保真度并坚持文本,但它们缺乏此类任务所需的特定控制。现有的符号到音频的研究往往集中在单一的,音调的文书,离开复调的挑战,鼓合成未解决。我们通过引入“打破节拍!”来解决这一差距,“一个能够以参考音频的音色呈现鼓的模型。它是通过使用我们提出的内容编码器和有效的混合调节机制微调预训练的文本到音频模型而构建的。为了实现这一点,我们从现有的鼓音频数据集构造成对的目标参考鼓音频的新数据集。实验表明,我们的模型生成高质量的鼓音频,遵循高分辨率鼓的声音,实现了强大的性能指标的音频质量,节奏对齐,和节拍连续性。这为生产者提供了一个新的,可控的创造性生产工具。演示页面:https: ik4sumii.github.io break-the-beat 摘要:Current methods for creating drum loop audio in digital music production, such as using one-shot samples or resampling, often demand non-trivial efforts of creators. While recent generative models achieve high fidelity and adhere to text, they lack the specific control needed for such a task. Existing symbolic-to-audio research often focuses on single, tonal instruments, leaving the challenge of polyphonic, percussive drum synthesis unaddressed. We address this gap by introducing Break-the-Beat!,'' a model capable of rendering a drum MIDI with the timbre of a reference audio. It is built by fine-tuning a pre-trained text-to-audio model with our proposed content encoder and a effective hybrid conditioning mechanism. To enable this, we construct a new dataset of paired target-reference drum audio from existing drum audio datasets. Experiments demonstrate that our model generates high-quality drum audio that follows high-resolution drum MIDI, achieving strong performance across metrics of audio quality, rhythmic alignment, and beat continuity. This offer producers a new, controllable tool for creative production. Demo page: https: ik4sumii.github.io break-the-beat
【8】Physics-Based iOCT Sonification for Real-time Interaction Awareness in Subretinal Injection
标题:基于物理的iOptical Sonification用于视网膜下注射中的实时交互感知
链接:https://arxiv.org/pdf/2605.14500v1
摘要:视网膜下注射是一种精细的玻璃体视网膜手术,需要在视网膜下腔内精确放置针头,同时避免视网膜色素上皮(RPE)穿孔,RPE是直接位于目标下方的一层,再生能力极其有限。为了增强插管推进过程中的深度感知,术中光学相干断层扫描(iOCT)提供了针-组织相互作用的高分辨率横截面可视化;然而,解释这些图像需要在正面显微镜视图的同时保持视觉注意力,从而增加了关键阶段的认知负荷,并对外科医生的本体感受控制提出了额外要求。在本文中,我们提出了一个结构化的,实时的发音框架,旨在可扩展的映射iOCT派生的解剖特征到感知听觉反馈。该方法采用由来自iOCT B扫描流的分割的视网膜层驱动的物理学启发的声学模型,其中针运动和注射诱导的视网膜层位移用作声音模型的激励输入,从而实现工具位置和视网膜变形的感知。在一项对照用户研究(n=34)中,所提出的超声处理实现了高视网膜层识别准确性和视网膜变形相关事件的稳健检测,在总体事件识别中显著优于最新基线(83.4% vs. 60.6%,p < 0.001),主要通过增强对注射诱导的视网膜变形的检测来驱动增益。专家评价(n=4)证实了该方法的临床相关性和潜在术中适用性。这些结果确立了结构化iOCT超声作为视网膜下注射实时手术指导的可行补充模式。摘要:Subretinal injection is a delicate vitreoretinal procedure requiring precise needle placement within the subretinal space while avoiding perforation of the retinal pigment epithelium (RPE), a layer directly beneath the target with extremely limited regenerative capacity. To enhance depth perception during cannula advancement, intraoperative optical coherence tomography (iOCT) offers high-resolution cross-sectional visualization of needle-tissue interaction; however, interpreting these images requires sustained visual attention alongside the en face microscope view, thereby increasing cognitive load during critical phases and placing additional demands on the surgeon's proprioceptive control. In this paper, we propose a structured, real-time sonification framework designed for extensible mapping of iOCT-derived anatomical features into perceptual auditory feedback. The method employs a physics-inspired acoustic model driven by segmented retinal layers from a stream of iOCT B-scans, with needle motion and injection-induced retinal layer displacements serving as excitation inputs to the sound model, enabling perception of tool position and retinal deformation. In a controlled user study (n=34), the proposed sonification achieved high retinal layer identification accuracy and robust detection of retinal deformation-related events, significantly outperforming a state-of-the-art baseline in overall event identification (83.4% vs. 60.6%, p < 0.001), with gains driven primarily by enhanced detection of injection-induced retinal deformation. Evaluation by experts (n=4) confirmed the clinical relevance and potential intraoperative applicability of the method. These results establish structured iOCT sonification as a viable complementary modality for real-time surgical guidance in subretinal injection.
【9】A Calculus-Based Framework for Determining Vocabulary Size in End-to-End ASR
标题:用于确定端到端ASB中词汇量的基于微积分的框架
链接:https://arxiv.org/pdf/2605.14427v1
备注:8 pages, is an extension of the paper S. K. Kopparapu and A. Panda, A cost minimization approach to fix the vocabulary size in a tokenizer for an end-to-end ASR system, in Proceedings of the 2024 International Conference on Pattern Recognition, Kolkata, India, 2024
摘要:在混合自动语音识别(ASR)系统中,词汇量是明确的,通常由语言中存在的音素、双音素或三音素的数量确定。相比之下,端到端ASR系统从用于训练的文本语料库中获取词汇,通常称为令牌。词汇表的选择以及更重要的是词汇表的大小是训练端到端ASR系统的关键超参数。诸如字节对编码(BPE)、WordPiece和Unigram语言模型(ULM)之类的令牌化算法使用词汇大小作为输入超参数来生成在ASR训练期间采用的子词。像ESPNet这样的流行工具包在其训练食谱中提供了固定的词汇量,但文献中几乎没有关于如何确定这些值的文档或讨论。最近的工作[1]已经正式确定了一种方法来确定最适合端到端ASR的词汇量,引入了一个成本函数框架,将标记化过程视为黑盒。在本文中,我们建立在这个基础上,通过曲线拟合的训练数据,并使用微积分中的一阶和二阶导数测试的原则,正式估计词汇量超参数。我们证明了我们的方法的实用性和有用性,通过将其应用于标准的Librisepeech语料库,并表明词汇量超参数的最佳选择提高了ASR的性能。本文的主要贡献是形式化一种方法,以确定最适合训练端到端ASR系统的词汇量。摘要:In hybrid automatic speech recognition (ASR) systems, the vocabulary size is unambiguous, typically determined by the number of phones, bi-phones, or tri-phones present in the language. In contrast, end-to-end ASR systems derive their vocabulary, often referred to as tokens from the text corpus used for training. The choice and, more importantly, the size of this vocabulary is a critical hyper-parameter in training end-to-end ASR systems. Tokenization algorithms such as Byte Pair Encoding (BPE), WordPiece, and Unigram Language Model (ULM) use the vocabulary size as an input hyper-parameter to generate the sub-words employed during ASR training. Popular toolkits like ESPNet provide a fixed vocabulary size in their training recipes, but there is little documentation or discussion in the literature regarding how these values are determined. Recent work [1] has formalized an approach to identify the vocabulary size best suited for end-to-end ASR, introducing a cost function framework that treats the tokenization process as a black box. In this paper, we build upon that foundation by curve fitting the training data and using the principle of first and second derivative tests in calculus to formally estimate the vocabulary size hyper-parameter. We demonstrate the utility and usefulness of our approach by applying it on a standard Librispeech corpus and show that the optimal choice of vocabulary size hyper-parameter improves the performance of the ASR. The main contribution of this paper in formalizing an approach to identify the vocabulary size best suited for training an end-to-end ASR system.【10】Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR
标题:在基于LLM的ASB中,通过语音-文本对齐来细化伪音频脚本,以实现纯文本域自适应
链接:https://arxiv.org/pdf/2605.14340v1
备注:Submitted to Interspeech 2026
摘要:基于LLM的自动语音识别模型通过连接音频编码器和LLM表现出强大的性能。然而,配对语音和转录的数据稀缺往往阻碍了他们适应新的领域,使纯文本域适应至关重要。现有的方法通常依赖于单独微调LLM或采用伪音频提示。前者忽略了必要的声学背景,而后者要么遭受有限的可扩展性,在数据稀缺的条件下,或产生缺乏表现力的提示,仅利用文本功能,忽略音频模态。为了解决这个问题,我们提出了一个增强的框架,显式模型语音文本对齐。我们的方法有效地生成高度表达的伪音频提示,弥合模态差距,实现有效的目标域适应。实验表明,我们的方法优于现有的纯文本方法,提高了整体错误率和词汇覆盖率。摘要:LLM-based automatic speech recognition models demonstrate strong performance by connecting audio encoders and LLMs. However, data scarcity of paired speech and transcription often hinders their adaptation to new domains, making text-only domain adaptation crucial. Existing methods typically rely on either fine-tuning the LLM alone or employing pseudo-audio prompts. The former neglects essential acoustic context, while the latter either suffers from limited scalability in data-scarce conditions, or yields inexpressive prompts by leveraging only textual features, ignoring audio modality. To address this, we propose an enhanced framework that explicitly models speech-text alignment. Our method efficiently generates highly expressive pseudo-audio prompts that bridges the modality gap, enabling effective target-domain adaptation. Experiments demonstrate that our approach outperforms existing text-only methods, improving both overall error rates and out-of-vocabulary coverage.【11】AudioMosaic: Contrastive Masked Audio Representation Learning
标题:AudioMosaic:对比掩蔽音频表示学习
链接:https://arxiv.org/pdf/2605.14231v1
备注:ICML2026
摘要:音频自监督学习(SSL)旨在从大规模未标记的音频数据中学习通用表示。虽然最近的进展主要是由生成重建目标驱动的,但对比方法仍然较少探索,部分原因是设计有效的音频增强的难度以及对比预训练所需的大批量。我们介绍 textbf{AudioMosaic},一个基于对比学习的音频编码器,用于一般音频理解。在预训练期间,AudioMosaic通过对频谱图补丁应用结构化时频掩蔽来构建正对,从而减少内存使用并实现高效的大批量训练。与生成方法相比,AudioMosaic编码器学习更具鉴别力的话语级表示,这些表示在数据集,域和声学条件之间表现出强大的可转移性。大量的实验表明,AudioMosaic在线性探测和微调的情况下,在几个标准音频基准测试中达到了最先进的性能。我们进一步表明,将预训练的AudioMosaic编码器集成到音频语言模型中可以提高音频语言任务的性能。该代码在我们的 href{https: github.com HanxunH AudioMosaic}{GitHub存储库}中公开提供。摘要:Audio self-supervised learning (SSL) aims to learn general-purpose representations from large-scale unlabeled audio data. While recent advances have been driven mainly by generative reconstruction objectives, contrastive approaches remain less explored, partly due to the difficulty of designing effective audio augmentations and the large batch sizes required for contrastive pre-training. We introduce textbf{AudioMosaic}, a contrastive learning-based audio encoder for general audio understanding. During pre-training, AudioMosaic constructs positive pairs by applying structured time-frequency masking to spectrogram patches, which reduces memory usage and enables efficient large-batch training. Compared with generative approaches, the AudioMosaic encoder learns more discriminative utterance-level representations that demonstrate strong transferability across datasets, domains, and acoustic conditions. Extensive experiments show that AudioMosaic achieves state-of-the-art performance on several standard audio benchmarks under both linear probing and fine-tuning. We further show that integrating the pretrained AudioMosaic encoder into audio-language models improves performance on audio-language tasks. The code is publicly available in our href{https: github.com HanxunH AudioMosaic}{GitHub repository}.
【12】Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study
标题:数据有限的掩蔽自动编码器:它有效吗?细粒度生物声学案例研究
链接:https://arxiv.org/pdf/2605.14031
备注:Workshop on Fine-Grained Visual Categorization (FGVC) at CVPR 2026. 8 pages, 6 figures
摘要:
摘要:
【13】Case Studies and Reflections on Agentic Software Engineering for Rapid Development of Digital Music Instruments
标题:数字乐器快速开发的动态软件工程案例研究与思考
链接:https://arxiv.org/pdf/2605.14016
摘要:
摘要:
【14】A Benchmark for Early-stage Parkinson's Disease Detection from Speech
标题:通过言语检测早期帕金森病的基准
链接:https://arxiv.org/pdf/2605.14066v1
备注:Submitted to Interspeech2026
摘要:从语音中检测早期帕金森病(EarlyPD)具有临床意义,但尚未充分探索,并且由于研究在数据集,语言,任务,评估协议和EarlyPD定义方面存在差异,因此很难比较已发表的结果。为了解决这个问题,我们提出了基于语音的EarlyPD检测的第一个基准,该基准具有独立于说话者的分割,旨在对研究人员可访问的数据集进行公平和可复制的跨方法评估。该基准涵盖了三种常见的语音任务,并在不同的训练资源设置下评估方法。我们还按数据集、聚合水平、性别和疾病阶段提供了多维评估细分,以支持细粒度比较和临床采用。我们的研究结果提供了一个可复制的参考和可操作的见解,鼓励采用这种公开的基准,以推进强大的和临床上有意义的早期PD检测从语音。摘要:Early-stage Parkinson's disease (EarlyPD) detection from speech is clinically meaningful yet underexplored, and published results are hard to compare because studies differ in datasets, languages, tasks, evaluation protocols, and EarlyPD definitions. To address this issue, we propose the first benchmark for speech-based EarlyPD detection, with a speaker-independent split designed for fair and replicable cross-method evaluation on researcher-accessible datasets. The benchmark covers three common speech tasks and evaluates methods under different training-resource settings. We also present multi-dimensional evaluation breakdowns by dataset, aggregation level, gender, and disease stage to support fine-grained comparisons and clinical adoption. Our results provide a replicable reference and actionable insights, encouraging the adoption of this publicly available benchmark to advance robust and clinically meaningful EarlyPD detection from speech.
【1】A Benchmark for Early-stage Parkinson's Disease Detection from Speech
标题:通过言语检测早期帕金森病的基准
链接:https://arxiv.org/pdf/2605.14066v1
【2】FSD50K-Solo: Automated Curation of Single-Source Sound Events
标题:FSD 50 K-Solo:单源声音事件的自动化处理
链接:https://arxiv.org/pdf/2605.13931v1
备注:Accepted to EUSIPCO 2026. 5 pages, 3 figures
摘要:高质量的训练数据集对于神经网络的性能至关重要。然而,音频领域仍然缺乏大规模、强标记和单源声音事件数据集。尽管FSD 50 K数据集相对较大且开放,但它包含相当一部分多源样本,其中背景干扰或重叠事件可能会限制数据的有用性。为了应对这一挑战,我们引入了一个数据策展框架,专为大型开放音频语料库。我们的方法利用生成扩散模型来合成干净的单类事件,以构建用于监督的受控噪声混合物。随后,我们采用一个预先训练的音频编码器,再加上一个判别分类器,自动识别和过滤出多源样本。实验表明,我们的框架在人类专家策划的测试集上实现了强大的性能。最后,我们发布了FSD 50 K-Solo,这是FSD 50 K的一个模型精选子集,包含通过我们的方法识别的单源音频样本。除了FSD 50 K,我们的方法建立了一个可扩展的模式,策划开源音频语料库。摘要:High-quality training datasets are essential for the performance of neural networks. However, the audio domain still lacks a large-scale, strongly-labeled, and single-source sound event dataset. The FSD50K dataset, despite being relatively large and open, contains a considerable fraction of multi-source samples where background interference or overlapping events could limit the usefulness of the data. To address this challenge, we introduce a data curation framework designed for large-scale open audio corpora. Our approach leverages a generative diffusion model to synthesize clean single-class events to construct controlled noisy mixtures for supervision. We subsequently employ a pre-trained audio encoder coupled with a discriminative classifier to automatically identify and filter out multi-source samples. Experiments show that our framework achieves strong performance on a human expert-curated test set. Finally, we release FSD50K-Solo, a model-curated subset of FSD50K containing single-source audio samples identified by our method. Beyond FSD50K, our method establishes a scalable paradigm for curating open source audio corpora.
【3】SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
标题:SpeakerLLM:一种用于说话人理解和验证推理的演讲人专业音频LLM
链接:https://arxiv.org/pdf/2605.15044v1
【4】Streaming Speech-to-Text Translation with a SpeechLLM
标题:使用SpeechLLM流媒体语音转文本翻译
链接:https://arxiv.org/pdf/2605.14766v1
备注:9 pages of main text; 24 pages in total
摘要:通常,将语音翻译成文本的系统由用于语音识别和文本到文本翻译的单独模块组成。将这些任务组合成SpeechLLM有望利用语音中的非语言信息并减少级联错误。但是现有的SpeechLLM系统速度很慢,因为它们不能以真正的流媒体方式工作:它们在输出翻译之前等待完整的音频话语,或者以固定的时间间隔输出令牌,这不适合实际应用。这项工作提出了一个基于LLM的架构,真正的流语音到文本翻译。LLM不仅学习发出输出令牌,而且还决定它是否已经看到足够的音频来这样做。该系统使用输入语音和输出文本的自动对齐来训练。在不同语言对的实验中,该系统实现了接近非流基线的翻译质量,但延迟仅为1-2秒。摘要:Normally, a system that translates speech into text consists of separate modules for speech recognition and text-to-text translation. Combining those tasks into a SpeechLLM promises to exploit paralinguistic information in the speech and to reduce cascaded errors. But existing SpeechLLM systems are slow since they do not work in a real streaming fashion: they wait for a complete utterance of audio before outputting a translation, or output tokens at fixed intervals, which is not suitable for real applications. This work proposes an LLM-based architecture for real streaming speech-to-text translation. The LLM learns not just to emit output tokens, but also to decide whether it has seen enough audio to do so. The system is trained using automatic alignments of the input speech and the output text. In experiments on different language pairs, the system achieves a translation quality close to the non-streaming baseline, but with a latency of only 1-2 seconds.
机器翻译由腾讯交互翻译提供,仅供参考
