微信公众号:arXiv_Daily
cs.SD语音
【1】Beyond Fixed Frames: Dynamic Character-Aligned Speech Tokenization
标题:超越固定帧:动态字符串对齐语音令牌化
链接:https://arxiv.org/abs/2601.23174
备注:18 pages, 3 figures
摘要:神经音频编解码器是现代会话语音技术的核心,将连续语音转换为可以由LLM处理的离散令牌序列。然而,现有的编解码器通常以固定的帧速率操作,在时间上均匀地分配令牌并产生不必要的长序列。在这项工作中,我们介绍了DyCAST,一个动态的字符对齐的语音标记,使可变帧速率标记,通过软字符级对齐和明确的持续时间建模。DyCAST在训练过程中学习将令牌与字符级语言单元相关联,并在解码时直接控制令牌持续时间,从而支持无干扰推理。为了在低帧速率下提高语音再合成质量,我们进一步引入了一种检索增强解码机制,在不增加比特率的情况下提高重建保真度。实验表明,DyCAST实现竞争力的语音再合成质量和下游性能,同时使用显着更少的令牌比固定帧速率编解码器。
摘要:Neural audio codecs are at the core of modern conversational speech technologies, converting continuous speech into sequences of discrete tokens that can be processed by LLMs. However, existing codecs typically operate at fixed frame rates, allocating tokens uniformly in time and producing unnecessarily long sequences. In this work, we introduce DyCAST, a Dynamic Character-Aligned Speech Tokenizer that enables variable-frame-rate tokenization through soft character-level alignment and explicit duration modeling. DyCAST learns to associate tokens with character-level linguistic units during training and supports alignment-free inference with direct control over token durations at decoding time. To improve speech resynthesis quality at low frame rates, we further introduce a retrieval-augmented decoding mechanism that enhances reconstruction fidelity without increasing bitrate. Experiments show that DyCAST achieves competitive speech resynthesis quality and downstream performance while using significantly fewer tokens than fixed-frame-rate codecs.
【2】DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
标题:DIFFA-2:通用音频理解的实用扩散大型语言模型
链接:https://arxiv.org/abs/2601.23161
摘要:自回归(AR)大型音频语言模型(LALM),如Qwen-2.5-Omni,在音频理解和交互方面取得了很好的性能,但扩展它们仍然需要大量的数据和计算,并且严格的顺序解码限制了推理效率。扩散大语言模型(dLLM)最近已被证明可以有效地利用有限的训练数据,并且DIFFA的先前工作表明,用扩散对应物替换AR主干可以在匹配的设置下大大提高音频理解,尽管在概念验证的规模上没有大规模的指令调整,偏好对齐或实际的解码方案。我们介绍DIFFA-2,一个实用的基于扩散的LALM一般音频理解。DIFFA-2升级了语音编码器,采用双重语义和声学适配器,并使用四阶段课程进行训练,该课程结合了语义和声学对齐,大规模监督微调和方差减少偏好优化,仅使用完全开源的语料库。在MMSU,MMAU和MMAR上的实验表明,DIFFA-2在实际训练预算下始终优于DIFFA,并且与强AR LALM竞争,支持基于扩散的建模是大规模音频理解的可行支柱。我们的代码可在https://github.com/NKU-HLT/DIFFA.git上获得。
摘要:Autoregressive (AR) large audio language models (LALMs) such as Qwen-2.5-Omni have achieved strong performance on audio understanding and interaction, but scaling them remains costly in data and computation, and strictly sequential decoding limits inference efficiency. Diffusion large language models (dLLMs) have recently been shown to make effective use of limited training data, and prior work on DIFFA indicates that replacing an AR backbone with a diffusion counterpart can substantially improve audio understanding under matched settings, albeit at a proof-of-concept scale without large-scale instruction tuning, preference alignment, or practical decoding schemes. We introduce DIFFA-2, a practical diffusion-based LALM for general audio understanding. DIFFA-2 upgrades the speech encoder, employs dual semantic and acoustic adapters, and is trained with a four-stage curriculum that combines semantic and acoustic alignment, large-scale supervised fine-tuning, and variance-reduced preference optimization, using only fully open-source corpora. Experiments on MMSU, MMAU, and MMAR show that DIFFA-2 consistently improves over DIFFA and is competitive to strong AR LALMs under practical training budgets, supporting diffusion-based modeling is a viable backbone for large-scale audio understanding. Our code is available at https://github.com/NKU-HLT/DIFFA.git.
【3】Hearing is Believing? Evaluating and Analyzing Audio Language Model Sycophancy with SYAUDIO
标题:听到就是相信?利用SYAUDIO评估和分析音频语言模型的谄媚性
链接:https://arxiv.org/abs/2601.23149
摘要:音频语言模型(ALM)最近在语音,声音和自然语言的统一推理方面表现出强大的能力;然而,它们继承了在大型语言模型中观察到的行为问题,包括奉承-即使与客观证据相矛盾,也倾向于同意用户断言。虽然奉承已经在文本和视觉语言模型中得到了广泛的研究,但其在音频条件推理中的表现在很大程度上仍未被探索,尽管ALM需要依赖于听觉线索,如声学事件,扬声器特性和语音速率。为了解决这一差距,我们引入了SYAUDIO,这是第一个专门用于评估ALM中奉承行为的基准,由4,319个音频问题组成,涵盖音频感知,音频推理,音频数学和音频道德。SYAUDIO建立在既定的音频基准之上,并通过TTS生成的算术和道德推理任务进行了增强,它能够通过仔细验证的数据质量,对多个领域和奉承类型进行系统评估。此外,我们分析了在涉及噪声和速率的现实条件下的音频特定的阿谀奉承,并证明了用思想链数据进行监督微调是减少ALM中阿谀奉承行为的有效缓解策略。
摘要:Audio Language Models (ALMs) have recently shown strong capabilities in unified reasoning over speech, sound, and natural language; yet they inherit behavioral issues observed in Large Language Models, including sycophancy--the tendency to agree with user assertions even when they contradict objective evidence. While sycophancy has been extensively studied in text and vision-language models, its manifestation in audio-conditioned reasoning remains largely unexplored, despite the need for ALMs to rely on auditory cues such as acoustic events, speaker characteristics, and speech rate. To address this gap, we introduce SYAUDIO, the first benchmark dedicated to evaluating sycophancy in ALMs, consisting of 4,319 audio questions spanning Audio Perception, Audio Reasoning, Audio Math, and Audio Ethics. Built upon established audio benchmarks and augmented with TTS-generated arithmetic and moral reasoning tasks, SYAUDIO enables systematic evaluation across multiple domains and sycophancy types with carefully verified data quality. Furthermore, we analyze audio-specific sycophancy under realistic conditions involving noise and rate, and demonstrate that supervised fine-tuning with chain-of-thought data is an effective mitigation strategy for reducing sycophantic behavior in ALMs.
【4】Towards Explicit Acoustic Evidence Perception in Audio LLMs for Speech Deepfake Detection
标题:在音频LLM中实现显式声学证据感知以实现语音深度伪造检测
链接:https://arxiv.org/abs/2601.23066
备注:9 pages, 4 figures
摘要:语音深度伪造检测(SDD)专注于识别给定的语音信号是真实的还是合成的。现有的基于音频大语言模型(LLM)的方法在内容理解方面表现出色;然而,它们的预测往往偏向于语义相关的线索,这导致细粒度的声学伪影在决策过程中被忽视。因此,具有自然语义的假语音可以绕过检测器,尽管隐藏着微妙的声学异常;这表明,挑战不是源于声学数据的缺乏,而是当语义主导推理盛行时,其可访问性不足。为了解决这个问题,我们调查SDD内的音频LLM范式,并引入SDD与听觉感知增强音频大语言模型(SDD-APALLM),声学增强的框架,旨在明确地暴露细粒度的时间-频率证据作为可访问的声学线索。通过将原始音频与结构化频谱图相结合,所提出的框架使音频LLM能够更有效地捕获细微的声学不一致,而不会影响其语义理解。实验结果表明,一致的增益检测精度和鲁棒性,特别是在语义线索是误导的情况下。进一步的分析表明,这些改进源于语义和声学信息的协调利用,而不是简单的模态聚合。
摘要:Speech deepfake detection (SDD) focuses on identifying whether a given speech signal is genuine or has been synthetically generated. Existing audio large language model (LLM)-based methods excel in content understanding; however, their predictions are often biased toward semantically correlated cues, which results in fine-grained acoustic artifacts being overlooked during the decisionmaking process. Consequently, fake speech with natural semantics can bypass detectors despite harboring subtle acoustic anomalies; this suggests that the challenge stems not from the absence of acoustic data, but from its inadequate accessibility when semantic-dominant reasoning prevails. To address this issue, we investigate SDD within the audio LLM paradigm and introduce SDD with Auditory Perception-enhanced Audio Large Language Model (SDD-APALLM), an acoustically enhanced framework designed to explicitly expose fine-grained time-frequency evidence as accessible acoustic cues. By combining raw audio with structured spectrograms, the proposed framework empowers audio LLMs to more effectively capture subtle acoustic inconsistencies without compromising their semantic understanding. Experimental results indicate consistent gains in detection accuracy and robustness, especially in cases where semantic cues are misleading. Further analysis reveals that these improvements stem from a coordinated utilization of semantic and acoustic information, as opposed to simple modality aggregation.
【5】DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
标题:扩散言语:无声的思想,通过统一的言语-文本传播说出的答案
链接:https://arxiv.org/abs/2601.22889
摘要:当前的语音语言模型直接生成响应而没有明确的推理,导致一旦产生音频就无法纠正的错误。我们介绍\textbf{``沉默的思想,口语的批评者'}-一个范例,其中语音LLM生成内部文本推理旁边的口头反应,与思维痕迹通知语音质量。为了实现这一点,我们提出了\method{},第一个基于扩散的语音-文本语言模型,支持理解和生成,统一离散文本和标记语音下一个单一的蒙面扩散框架。与自回归方法不同,\method{}通过迭代去噪联合生成推理轨迹和语音标记,并使用特定于模态的掩蔽时间表。我们还构建了\dataset{},这是第一个具有成对文本推理痕迹的语音QA数据集,包含26 K样本,总计319小时。实验表明,\method{}实现了最先进的语音到语音QA的准确性,超过最佳基线高达9分,同时在生成模型中获得最佳TTS质量(6.2\% WER)并保持语言理解(66.2\% MMLU)。烧蚀证实,扩散架构和思维痕迹有助于这些收益。
摘要:Current speech language models generate responses directly without explicit reasoning, leading to errors that cannot be corrected once audio is produced. We introduce \textbf{``Silent Thought, Spoken Answer''} -- a paradigm where speech LLMs generate internal text reasoning alongside spoken responses, with thinking traces informing speech quality. To realize this, we present \method{}, the first diffusion-based speech-text language model supporting both understanding and generation, unifying discrete text and tokenized speech under a single masked diffusion framework. Unlike autoregressive approaches, \method{} jointly generates reasoning traces and speech tokens through iterative denoising, with modality-specific masking schedules. We also construct \dataset{}, the first speech QA dataset with paired text reasoning traces, containing 26K samples totaling 319 hours. Experiments show \method{} achieves state-of-the-art speech-to-speech QA accuracy, outperforming the best baseline by up to 9 points, while attaining the best TTS quality among generative models (6.2\% WER) and preserving language understanding (66.2\% MMLU). Ablations confirm that both the diffusion architecture and thinking traces contribute to these gains.
【6】Compact Hypercube Embeddings for Fast Text-based Wildlife Observation Retrieval
标题:用于快速基于文本的野生动物观察检索的紧凑超立方体嵌入
链接:https://arxiv.org/abs/2601.22783
摘要:大规模生物多样性监测平台越来越依赖于多模式野生动物观测。虽然最近的基础模型支持跨视觉、音频和语言的丰富语义表示,但由于高维相似性搜索的计算成本,从大量档案中检索相关观察结果仍然具有挑战性。在这项工作中,我们引入紧凑的超立方体嵌入快速基于文本的野生动物观察检索,一个框架,使高效的基于文本的搜索大规模的野生动物图像和音频数据库使用紧凑的二进制表示。基于跨视图代码对齐哈希框架,我们将轻量级哈希扩展到单模态设置之外,以将自然语言描述与共享汉明空间中的视觉或声学观察对齐。我们的方法利用了预训练的野生动物基础模型,包括BioCLIP和BioLingual,并使用参数高效的微调有效地调整它们以进行散列。我们在大规模基准测试中评估了我们的方法,包括用于文本到图像检索的iNaturalist 2024和用于文本到音频检索的iNatSounds2024,以及多个音景数据集,以评估域转移下的鲁棒性。结果表明,检索使用离散超立方体嵌入实现竞争力,并在某些情况下优越,性能相比,连续嵌入,同时大大降低了内存和搜索成本。此外,我们观察到散列目标一致地改进了底层编码器表示,从而导致更强的检索和zero-shot泛化。这些结果表明,二进制,基于语言的检索,使可扩展的和有效的搜索大型野生动物档案的生物多样性监测系统。
摘要:Large-scale biodiversity monitoring platforms increasingly rely on multimodal wildlife observations. While recent foundation models enable rich semantic representations across vision, audio, and language, retrieving relevant observations from massive archives remains challenging due to the computational cost of high-dimensional similarity search. In this work, we introduce compact hypercube embeddings for fast text-based wildlife observation retrieval, a framework that enables efficient text-based search over large-scale wildlife image and audio databases using compact binary representations. Building on the cross-view code alignment hashing framework, we extend lightweight hashing beyond a single-modality setup to align natural language descriptions with visual or acoustic observations in a shared Hamming space. Our approach leverages pretrained wildlife foundation models, including BioCLIP and BioLingual, and adapts them efficiently for hashing using parameter-efficient fine-tuning. We evaluate our method on large-scale benchmarks, including iNaturalist2024 for text-to-image retrieval and iNatSounds2024 for text-to-audio retrieval, as well as multiple soundscape datasets to assess robustness under domain shift. Results show that retrieval using discrete hypercube embeddings achieves competitive, and in several cases superior, performance compared to continuous embeddings, while drastically reducing memory and search cost. Moreover, we observe that the hashing objective consistently improves the underlying encoder representations, leading to stronger retrieval and zero-shot generalization. These results demonstrate that binary, language-based retrieval enables scalable and efficient search over large wildlife archives for biodiversity monitoring systems.
【7】How Far Can Pretrained LLMs Go in Symbolic Music? Controlled Comparisons of Supervised and Preference-based Adaptation
标题:经过预训练的法学硕士在象征音乐方面能走多远?监督适应和基于偏好的适应的对照比较
链接:https://arxiv.org/abs/2601.22764
备注:Accepted at NLP4MusA 2026
摘要:音乐通常与语言有着显著的相似之处,这促使人们使用预训练的大型语言模型(LLM)来理解和生成符号音乐。尽管越来越多的兴趣,适应的象征性音乐调整LLMs的实际效果仍然不够的特点。我们提出了一个基于ABC的生成和理解的微调策略的对照比较研究,比较了现成的调整调整骨干域适应的变体和音乐专业的LLM基线。在多个符号音乐语料库和评价信号,我们提供了一些见解适应选择的符号音乐应用程序。我们强调域适应与~保留先前的信息折衷以及用于测量符号音乐的域适应的度量的不同行为。
摘要:Music often shares notable parallels with language, motivating the use of pretrained large language models (LLMs) for symbolic music understanding and generation. Despite growing interest, the practical effectiveness of adapting instruction-tuned LLMs to symbolic music remains insufficiently characterized. We present a controlled comparative study of finetuning strategies for ABC-based generation and understanding, comparing an off-the-shelf instruction-tuned backbone to domain-adapted variants and a music-specialized LLM baseline. Across multiple symbolic music corpora and evaluation signals, we provide some insights into adaptation choices for symbolic music applications. We highlight the domain adaptation vs.~preserving prior information tradeoff as well as the distinct behaviour of metrics used to measure the domain adaptation for symbolic music.
【8】Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability
标题:通过平均延续日志概率评估和奖励具有表达性角色扮演的TTC的LALM
链接:https://arxiv.org/abs/2601.22661
摘要:大型音频语言模型(LALM)的最新进展已经将文本到语音(TTS)扩展到交互式角色扮演场景,这需要高表现力和严格遵守角色扮演指令。然而,现有的模型很难在多回合对话中保持与角色轮廓和场景描述的风格一致性。一个关键的瓶颈是缺乏量化说话风格的客观指标。为了弥补这一差距,我们提出了平均连续对数概率(MCLP)作为评估指标和奖励信号,验证基于LALM的角色扮演TTS(RP-TTS)任务。关键是,我们利用预先训练的LALM的上下文学习能力,通过连续对数概率预测来制定MCLP。该度量通过测量以生成的语音为条件的地面实况语音的可能性来量化风格一致性。此外,我们采用MCLP作为强化学习奖励,以增强生成的语音和角色扮演指令之间的风格对齐。为了便于评估,我们构建了一个RP-TTS数据集,具有丰富的场景和字符注释。实验结果表明,我们的方法显着优于强LALM基线的客观和主观指标。
摘要:Recent advances in Large Audio Language Models (LALMs) have extended Text-to-Speech (TTS) to interactive role-play scenarios, which demand high expressiveness and strict adherence to role-play instructions. However, existing models struggle to maintain stylistic consistency with character profiles and scene descriptions across multi-turn dialogues. A critical bottleneck is the lack of objective metrics for quantifying speaking style. To bridge this gap, we propose Mean Continuation Log-Probability (MCLP) as both an evaluation metric and a reward signal, validated on LALM-based Role-Play TTS (RP-TTS) tasks. Critically, we leverage the In-Context Learning capability of pre-trained LALMs to formulate MCLP via a continuation log-probability prediction. This metric quantifies stylistic consistency by measuring the likelihood of the ground-truth speech conditioned on the generated speech. Furthermore, we employ MCLP as a reinforcement learning reward to enhance the style alignment between generated speech and Role-Play instructions. To facilitate evaluation, we construct an RP-TTS dataset with rich scene and character annotations. Experimental results demonstrate that our method significantly outperforms strong LALM baselines on both objective and subjective metrics.
【9】A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
标题:用于数据高效的基于查询的通用声音分离的语义一致数据集
链接:https://arxiv.org/abs/2601.22599
备注:Technical Report
摘要:基于查询的通用声音分离是智能听觉系统的基础,旨在从混合物中分离出特定的源。尽管最近的进展,现有的方法继续遭受复杂的声学场景中的残留干扰。这种性能限制主要源于数据瓶颈:野外数据集包含弱标签和严重的事件共现。这些缺陷导致模型学习背景噪声和目标类别之间的虚假相关性,而不是鲁棒的声学特征。为了解决这个问题,我们提出了一个自动化的管道,通过语义一致的合成协议从野外数据集中挖掘高纯度的单事件片段来消除事件的同现。利用这个管道,我们构建了Hive,这是一个高质量的合成数据集,包含2.4k小时的原始音频。实验结果表明,与在比Hive大500倍的巨大数据集上训练的最先进的模型SAM-Audio相比,在Hive上训练的某些开源模型实现了具有竞争力的分离精度和感知质量。此外,这些模型表现出显着的zero-shot泛化分布外的评价基准。这些发现强调,优先考虑监督信号的纯度可以提高数据效率,为训练鲁棒的听觉基础模型提供了一种新的范式,同时降低了计算成本。代码和数据集可在https://shandaai.github.io/Hive上获得。
摘要:Query-based universal sound separation is fundamental to intelligent auditory systems, aiming to isolate specific sources from mixtures. Despite recent advances, existing methods continue to suffer from residual interference in complex acoustic scenes. This performance limitation stems largely from a data bottleneck: in-the-wild datasets contain weak labels and severe co-occurrence of events. These flaws induce models to learn spurious correlations between background noise and target categories instead of robust acoustic features. To address this, we propose an automated pipeline that eliminates co-occurrence of events by mining high-purity single-event segments from in-the-wild datasets via a semantically consistent synthesis protocol. Utilizing this pipeline, we constructed Hive, a high-quality synthetic dataset comprising 2.4k hours of raw audio. Experimental results demonstrate that, compared with the state-of-the-art model SAM-Audio which was trained on a huge dataset $\sim$500 times larger than Hive, certain open-source models trained on Hive achieve competitive separation accuracy and perceptual quality. Moreover, these models exhibited remarkable zero-shot generalization on out-of-distribution evaluation benchmarks. These findings highlight that prioritizing purity of supervised signals enables significant data efficiency, offering a new paradigm for training robust auditory foundation models with reduced computational costs. Code and dataset are available at https://shandaai.github.io/Hive.
【10】MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control
标题:镜报:通过解开风格和分层运动控制打造个性化化身
链接:https://arxiv.org/abs/2601.22501
备注:Accepted to 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)
摘要:合成个性化的说话脸,坚持和突出发言者的独特风格,同时保持唇同步的准确性仍然是一个重大的挑战。现有方法的一个主要限制是说话者特定的谈话风格和面部运动内的语义内容的内在混淆,这阻止了说话者的独特人物的忠实转移到任意语音。在本文中,我们提出了一个基于条件扩散模型的生成框架,它结合了一个语义分解风格编码器(SDSE),可以从一个简短的参考视频中提取纯风格表示。为了有效地利用这种表示,我们进一步引入了一个分层的调制策略内的扩散过程。这种机制通过动态平衡不同面部区域的音频和风格特征的贡献来指导合成,确保精确的唇同步准确性和富有表现力的全脸动态。大量的实验表明,在唇同步准确性和个性化保留方面,与最先进的方法相比,TumblTalk实现了显着的改进。
摘要:Synthesizing personalized talking faces that uphold and highlight a speaker's unique style while maintaining lip-sync accuracy remains a significant challenge. A primary limitation of existing approaches is the intrinsic confounding of speaker-specific talking style and semantic content within facial motions, which prevents the faithful transfer of a speaker's unique persona to arbitrary speech. In this paper, we propose MirrorTalk, a generative framework based on a conditional diffusion model, combined with a Semantically-Disentangled Style Encoder (SDSE) that can distill pure style representations from a brief reference video. To effectively utilize this representation, we further introduce a hierarchical modulation strategy within the diffusion process. This mechanism guides the synthesis by dynamically balancing the contributions of audio and style features across distinct facial regions, ensuring both precise lip-sync accuracy and expressive full-face dynamics. Extensive experiments demonstrate that MirrorTalk achieves significant improvements over state-of-the-art methods in terms of lip-sync accuracy and personalization preservation.
【11】Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
标题:重新思考语音增强中的语音表示聚合:语音互信息的角度
链接:https://arxiv.org/abs/2601.22480
备注:Accepted to ICASSP 2026
摘要:最近的语音增强(SE)模型越来越多地利用自监督学习(SSL)表示其丰富的语义信息。通常,中间特征通过轻量级适配模块聚合到单个表示中。然而,大多数SSL模型没有经过噪声鲁棒性训练,这可能导致语义表示损坏。此外,自适应模块与SE模型联合训练,潜在地将声学细节优先于语义信息,这与最初的目的相矛盾。为了解决这个问题,我们首先从信息论的角度分析了SSL模型对噪声语音的行为。具体来说,我们测量损坏的SSL表示和相应的音素标签之间的互信息(MI),专注于保存的语言内容。在此分析的基础上,我们引入了语言聚合层,该层经过预训练,以使用音素标签(可选动态聚合)最大化MI,然后在SE训练期间冻结。实验表明,这种解耦的方法提高了单词错误率(WER)在联合优化的基线,展示了显式对齐的适应模块与语言内容的好处。
摘要:Recent speech enhancement (SE) models increasingly leverage self-supervised learning (SSL) representations for their rich semantic information. Typically, intermediate features are aggregated into a single representation via a lightweight adaptation module. However, most SSL models are not trained for noise robustness, which can lead to corrupted semantic representations. Moreover, the adaptation module is trained jointly with the SE model, potentially prioritizing acoustic details over semantic information, contradicting the original purpose. To address this issue, we first analyze the behavior of SSL models on noisy speech from an information-theoretic perspective. Specifically, we measure the mutual information (MI) between the corrupted SSL representations and the corresponding phoneme labels, focusing on preservation of linguistic contents. Building upon this analysis, we introduce the linguistic aggregation layer, which is pre-trained to maximize MI with phoneme labels (with optional dynamic aggregation) and then frozen during SE training. Experiments show that this decoupled approach improves Word Error Rate (WER) over jointly optimized baselines, demonstrating the benefit of explicitly aligning the adaptation module with linguistic contents.
【12】An Effective Energy Mask-based Adversarial Evasion Attacks against Misclassification in Speaker Recognition Systems
标题:基于有效能量面具的对抗规避攻击说话人识别系统中的误分类
链接:https://arxiv.org/abs/2601.22390
摘要:规避攻击对人工智能系统构成重大威胁,利用机器学习模型中的漏洞来绕过检测机制。目前,在有前途的未来行业中,语音数据(包括deepfake)的广泛使用受到法律框架不足的阻碍。对抗性攻击方法已成为防止滥用此类数据的最有效对策。本文介绍了一种利用功率谱对原始语音数据进行能量掩蔽的新方法-掩蔽能量扰动法。MEP在生成对抗性扰动之前对频域中的小能量区域应用掩蔽,目标是人类听觉模型不太明显的区域。该研究主要采用先进的说话人识别模型,包括ECAPA-TDNN和ResNet 34,这些模型在说话人验证任务中表现出出色的性能。所提出的MEP方法在音频质量和规避有效性方面表现出很强的性能。能量掩蔽方法有效地最大限度地减少了语音质量的感知评估(PESQ)的退化,这表明尽管有对抗性扰动,但人类收听者的感知失真最小。具体而言,在PESQ评估中,与快速梯度符号法(FGSM)和迭代FGSM相比,MEP方法的相对性能为26.68%。
摘要:Evasion attacks pose significant threats to AI systems, exploiting vulnerabilities in machine learning models to bypass detection mechanisms. The widespread use of voice data, including deepfakes, in promising future industries is currently hindered by insufficient legal frameworks. Adversarial attack methods have emerged as the most effective countermeasure against the indiscriminate use of such data. This research introduces masked energy perturbation (MEP), a novel approach using power spectrum for energy masking of original voice data. MEP applies masking to small energy regions in the frequency domain before generating adversarial perturbations, targeting areas less noticeable to the human auditory model. The study primarily employs advanced speaker recognition models, including ECAPA-TDNN and ResNet34, which have shown remarkable performance in speaker verification tasks. The proposed MEP method demonstrated strong performance in both audio quality and evasion effectiveness. The energy masking approach effectively minimizes the perceptual evaluation of speech quality (PESQ) degradation, indicating that minimal perceptual distortion occurs to the human listener despite the adversarial perturbations. Specifically, in the PESQ evaluation, the relative performance of the MEP method was 26.68% when compared to the fast gradient sign method (FGSM) and iterative FGSM.
【13】Attention Isn't All You Need for Emotion Recognition:Domain Features Outperform Transformers on the EAV Dataset
标题:情感识别需要的不仅仅是注意力:领域功能在EAV数据集中优于Transformer
链接:https://arxiv.org/abs/2601.22161
摘要:我们提出了一个系统的研究多模态情感识别使用EAV数据集,调查是否复杂的注意力机制提高性能的小数据集。我们实现了三个模型类别:基线Transformers(M1),新型因子化注意力机制(M2)和改进的CNN基线(M3)。我们的实验表明,复杂的注意力机制在小数据集上表现不佳。由于过度拟合和预训练特征的破坏,M2模型比基线低5到13个百分点。相比之下,简单的域适当的修改被证明是有效的:将增量MFCC添加到音频CNN中将准确率从61.9\%提高到\textbf{65.56\%}(+3.66pp),而EEG的频域特征达到\textbf{67.62\%}(比论文基线高出+7.62pp)。我们的Vision Transformer基线(M1)达到了\textbf{75.30\%},超过了论文的ViViT结果(74.5\%),通过特定领域的预训练,视觉增量特征达到了\textbf{72.68\%}(比论文CNN增加了1.28 pp)。这些发现表明,对于小规模的情感识别,领域知识和适当的实现优于架构的复杂性。
摘要:We present a systematic study of multimodal emotion recognition using the EAV dataset, investigating whether complex attention mechanisms improve performance on small datasets. We implement three model categories: baseline transformers (M1), novel factorized attention mechanisms (M2), and improved CNN baselines (M3). Our experiments show that sophisticated attention mechanisms consistently underperform on small datasets. M2 models achieved 5 to 13 percentage points below baselines due to overfitting and destruction of pretrained features. In contrast, simple domain-appropriate modifications proved effective: adding delta MFCCs to the audio CNN improved accuracy from 61.9\% to \textbf{65.56\%} (+3.66pp), while frequency-domain features for EEG achieved \textbf{67.62\%} (+7.62pp over the paper baseline). Our vision transformer baseline (M1) reached \textbf{75.30\%}, exceeding the paper's ViViT result (74.5\%) through domain-specific pretraining, and vision delta features achieved \textbf{72.68\%} (+1.28pp over the paper CNN). These findings demonstrate that for small-scale emotion recognition, domain knowledge and proper implementation outperform architectural complexity.
【14】EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
标题:MIDI Change:轻量级激活引导,用于增强的语音感知合成
链接:https://arxiv.org/abs/2601.22873
备注:Activation Steering; Emotion-Aware TTS; Speech Synthesis; Accepted by ICASSP 2026
摘要:在文语转换(TTS)合成中,实现精确可控的情感表达对于产生自然和上下文合适的语音至关重要。然而,许多情感感知的TTS系统,包括基于大型语言模型(LLM)的设计,依赖于缩放固定的情感嵌入或外部指导,限制了它们对特定于情感的潜在特征进行建模的能力。为了解决这一差距,我们提出了一个轻量级的激活转向框架,它包含一个激活转向层,它为输出嵌入空间中的每个目标情感学习一个转向向量,以捕获其潜在的偏移量,并在话语和类别中保持稳定,适当的表达。可训练参数仅为10 M,不足全微调的1/30,在客观和主观评测中均优于zero-shot和全微调基线,在保持自然度和说话人相似度的同时,增强情感表现力。进一步的分析证实了所提出的控制器Steer层的有效性,并揭示了其在语音合成中可控情感强度的潜力。
摘要:Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language model (LLM)-based designs, rely on scaling fixed emotion embeddings or external guidance, limiting their ability to model emotion-specific latent characteristics. To address this gap, we present EmoShift, a lightweight activation-steering framework incorporating a EmoSteer layer, which learns a steering vector for each target emotion in the output embedding space to capture its latent offset and maintain stable, appropriate expression across utterances and categories. With only 10M trainable parameters,less than 1/30 of full fine-tuning, EmoShift outperforms zero-shot and fully fine-tuned baselines in objective and subjective evaluations, enhancing emotional expressiveness while preserving naturalness and speaker similarity. Further analysis confirms the proposed EmoSteer layer's effectiveness and reveals its potential for controllable emotional intensity in speech synthesis.
【15】CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
标题:CALM:用于多说话者ASB个性化的联合上下文声学语言建模
链接:https://arxiv.org/abs/2601.22792
备注:Accepted to IEEE ICASSP 2026
摘要:我们提出了CALM,一个联合上下文声学语言建模框架的多说话人自动语音识别(ASR)。在个性化的人工智能场景中,声学和语言线索的联合可用性自然会激发目标说话者条件反射与重叠对话中的上下文偏置的整合。CALM通过说话人嵌入驱动的目标说话人提取和基于词汇的动态上下文偏置在端到端框架中实现了这种集成。我们评估CALM模拟英语(LibriSpeechMix)和日语(语料库的自发日语混合物,CSJMix)。在两个扬声器的混合,CALM减少了偏见的单词错误率(B-WER)从12.7到4.7 LibriSpeech 2 Mix和偏见的字符错误率(B-CER)从16.6到8.4 CSJMix 2(eval 3),证明了跨语言的联合声学语言建模的有效性。我们还报告了AMI语料库(IHM混合条件)的结果,以验证标准化语音混合物的性能。
摘要:We present CALM, a joint Contextual Acoustic-Linguistic Modeling framework for multi-speaker automatic speech recognition (ASR). In personalized AI scenarios, the joint availability of acoustic and linguistic cues naturally motivates the integration of target-speaker conditioning with contextual biasing in overlapping conversations. CALM implements this integration in an end-to-end framework through speaker embedding-driven target-speaker extraction and dynamic vocabulary-based contextual biasing. We evaluate CALM on simulated English (LibriSpeechMix) and Japanese (Corpus of Spontaneous Japanese mixtures, CSJMix). On two-speaker mixtures, CALM reduces biased word error rate (B-WER) from 12.7 to 4.7 on LibriSpeech2Mix and biased character error rate (B-CER) from 16.6 to 8.4 on CSJMix2 (eval3), demonstrating the effectiveness of joint acoustic-linguistic modeling across languages. We additionally report results on the AMI corpus (IHM-mix condition) to validate performance on standardized speech mixtures.
【16】Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
标题:使用仅解码器大语言模型和延迟优化的流语音识别
链接:https://arxiv.org/abs/2601.22779
备注:accepted to ICASSP 2026
摘要:最近的进展表明,解码器只有大语言模型(LLM)的自动语音识别(ASR)的潜力。然而,在这个框架内实现流识别仍然是一个挑战。在这项工作中,我们提出了一种新的流ASR方法,集成了读/写策略网络与单调分块注意力(MoChA)动态分段语音嵌入。这些片段在训练期间与标签序列交错,从而实现与LLM的无缝集成。在推断期间,音频流被缓冲,直到MoChA模块触发读取信号,此时缓冲的片段与先前的令牌一起被馈送到LLM中用于下一个令牌预测。我们还引入了最小延迟训练目标,以引导策略网络实现准确的分割边界。此外,我们采用了一种联合训练策略,其中非流式LLM-ASR模型和我们的流式模型共享参数。在AISHELL-1和AISHELL-2普通话基准测试上的实验表明,我们的方法始终优于最近的流ASR基线,分别实现了5.1%和5.5%的字符错误率。延迟优化导致平均令牌生成延迟减少62.5%,对识别准确性的影响可以忽略不计
摘要:Recent advances have demonstrated the potential of decoderonly large language models (LLMs) for automatic speech recognition (ASR). However, enabling streaming recognition within this framework remains a challenge. In this work, we propose a novel streaming ASR approach that integrates a read/write policy network with monotonic chunkwise attention (MoChA) to dynamically segment speech embeddings. These segments are interleaved with label sequences during training, enabling seamless integration with the LLM. During inference, the audio stream is buffered until the MoChA module triggers a read signal, at which point the buffered segment together with the previous token is fed into the LLM for the next token prediction. We also introduce a minimal-latency training objective to guide the policy network toward accurate segmentation boundaries. Furthermore, we adopt a joint training strategy in which a non-streaming LLM-ASR model and our streaming model share parameters. Experiments on the AISHELL-1 and AISHELL-2 Mandarin benchmarks demonstrate that our method consistently outperforms recent streaming ASR baselines, achieving character error rates of 5.1% and 5.5%, respectively. The latency optimization results in a 62.5% reduction in average token generation delay with negligible impact on recognition accuracy
【17】Proliferating series by Jean Barraqué: a study and classification in mathematical terms
标题:让·巴拉奎(Jean Barraqué)的沸腾系列:数学术语的研究和分类
链接:https://arxiv.org/abs/2601.22176
备注:28 pages, 8 figures
摘要:Barraqué的增殖系列通过在构建系列时创建一个新的不变量,对经典序列主义的概念进行了有趣的转变:在给定基本序列的增殖的构造期间保持不变的不是连续音符之间的间隔,而是发生在两个连续序列之间的音符的排列,也就是说,系列中音符顺序的转换。这为对序列方法感兴趣的作曲家提供了新的可能性,因为通过这种方法获得的间隔的多样性远远大于经典的序列主义。在这份手稿中,我们将从数学的角度研究激增的系列提供的一些未探索的可能性,这将使作曲家对它们更加熟悉,并可能导致创作出将序列主义提升到下一个层次的作品。
摘要:Barraqué's proliferating series give an interesting turn on the concept of classic serialism by creating a new invariant when it comes to constructing the series: rather than the intervals between consecutive notes, what remains unaltered during the construction of the proliferations of the given base series is the permutation of the notes which happens between two consecutive series, that is to say, the transformation of the order of the notes in the series. This presents new possibilities for composers interested in the serial method, given the fact that the variety of intervals obtained by this method is far greater than that of classic serialism. In this manuscript, we will study some unexplored possibilities that the proliferating series offer from a mathematical point of view, which will allow composers to gain much more familiarity with them and potentially result in the creation of pieces that take serialism to the next level.
【1】Beyond Omnidirectional: Neural Ambisonics Encoding for Arbitrary Microphone Directivity Patterns using Cross-Attention
标题:超越全方位:使用交叉注意力对任意麦克风方向性模式进行神经立体声编码
链接:https://arxiv.org/abs/2601.23196
备注:Accepted to ICASSP 2026
摘要:我们提出了一种用于将麦克风阵列信号编码为高保真度立体声的深度神经网络方法,该方法可推广到具有固定麦克风数量但位置和频率相关方向特性不同的任意麦克风阵列配置。与以前的方法,只依赖于阵列几何作为元数据,我们的方法使用定向阵列传递函数,使真实世界的阵列的准确表征。所提出的架构采用单独的编码器的音频和方向性的反应,结合它们通过交叉注意机制,以产生阵列独立的空间音频表示。我们评估的方法在两种设置的模拟数据:一个手机与复杂的身体散射,和自由场条件下,都与不同数量的声源在混响环境中。评估表明,我们的方法优于传统的基于数字信号处理的方法和现有的深度神经网络解决方案。此外,使用阵列传递函数而不是几何结构作为元数据输入提高了真实阵列的准确性。
摘要:We present a deep neural network approach for encoding microphone array signals into Ambisonics that generalizes to arbitrary microphone array configurations with fixed microphone count but varying locations and frequency-dependent directional characteristics. Unlike previous methods that rely only on array geometry as metadata, our approach uses directional array transfer functions, enabling accurate characterization of real-world arrays. The proposed architecture employs separate encoders for audio and directional responses, combining them through cross-attention mechanisms to generate array-independent spatial audio representations. We evaluate the method on simulated data in two settings: a mobile phone with complex body scattering, and a free-field condition, both with varying numbers of sound sources in reverberant environments. Evaluations demonstrate that our approach outperforms both conventional digital signal processing-based methods and existing deep neural network solutions. Furthermore, using array transfer functions instead of geometry as metadata input improves accuracy on realistic arrays.
【2】Layer-Aware Early Fusion of Acoustic and Linguistic Embeddings for Cognitive Status Classification
标题:用于认知状态分类的声学和语言嵌入的层感知早期融合
链接:https://arxiv.org/abs/2601.23004
备注:5 pages, 3 figures, paper accepted for ICASSP 2026 conference
摘要:语音包含反映认知衰退的声学和语言模式,因此仅描述一个域的模型无法完全捕获这种复杂性。本研究探讨如何早期融合(EF)的语音及其相应的转录文本嵌入,注意编码器层的深度,可以提高认知状态分类。使用来自DementiaBank的录音集合(1,629名扬声器;认知正常对照$\unicode{x2013}$CN,轻度认知障碍$\unicode{x2013}$MCI,阿尔茨海默病和相关痴呆症$\unicode {x2013}$ADRD),我们从wav 2 vec 2.0或Whisper的不同内部层中提取了与DistilBERT或RoBERTa相结合的帧对齐嵌入。用Transformer分类器训练单峰、EF和后期融合(LF)模型,对其进行优化,然后在10个种子上进行评估。性能在中间编码器层($\sim$8$\unicode{x2013}$10)中始终达到峰值,其中Whisper + RoBERTa层9处的单个最佳F1和Whisper + DistilBERT层10处的最佳对数损失。纯声学模型始终优于纯文本变体。EF增强了对真正声学嵌入的区分,而LF提高了概率校准。层的选择决定性地塑造了临床多模式协同。
摘要:Speech contains both acoustic and linguistic patterns that reflect cognitive decline, and therefore models describing only one domain cannot fully capture such complexity. This study investigates how early fusion (EF) of speech and its corresponding transcription text embeddings, with attention to encoder layer depth, can improve cognitive status classification. Using a DementiaBank-derived collection of recordings (1,629 speakers; cognitively normal controls$\unicode{x2013}$CN, Mild Cognitive Impairment$\unicode{x2013}$MCI, and Alzheimer's Disease and Related Dementias$\unicode{x2013}$ADRD), we extracted frame-aligned embeddings from different internal layers of wav2vec 2.0 or Whisper combined with DistilBERT or RoBERTa. Unimodal, EF and late fusion (LF) models were trained with a transformer classifier, optimized, and then evaluated across 10 seeds. Performance consistently peaked in mid encoder layers ($\sim$8$\unicode{x2013}$10), with the single best F1 at Whisper + RoBERTa layer 9 and the best log loss at Whisper + DistilBERT layer 10. Acoustic-only models consistently outperformed text-only variants. EF boosts discrimination for genuinely acoustic embeddings, whereas LF improves probability calibration. Layer choice critically shapes clinical multimodal synergy.
【3】EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
标题:MIDI Change:轻量级激活引导,用于增强的语音感知合成
链接:https://arxiv.org/abs/2601.22873
备注:Activation Steering; Emotion-Aware TTS; Speech Synthesis; Accepted by ICASSP 2026
摘要:在文语转换(TTS)合成中,实现精确可控的情感表达对于产生自然和上下文合适的语音至关重要。然而,许多情感感知的TTS系统,包括基于大型语言模型(LLM)的设计,依赖于缩放固定的情感嵌入或外部指导,限制了它们对特定于情感的潜在特征进行建模的能力。为了解决这一差距,我们提出了一个轻量级的激活转向框架,它包含一个激活转向层,它为输出嵌入空间中的每个目标情感学习一个转向向量,以捕获其潜在的偏移量,并在话语和类别中保持稳定,适当的表达。可训练参数仅为10 M,不足全微调的1/30,在客观和主观评测中均优于zero-shot和全微调基线,在保持自然度和说话人相似度的同时,增强情感表现力。进一步的分析证实了所提出的控制器Steer层的有效性,并揭示了其在语音合成中可控情感强度的潜力。
摘要:Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language model (LLM)-based designs, rely on scaling fixed emotion embeddings or external guidance, limiting their ability to model emotion-specific latent characteristics. To address this gap, we present EmoShift, a lightweight activation-steering framework incorporating a EmoSteer layer, which learns a steering vector for each target emotion in the output embedding space to capture its latent offset and maintain stable, appropriate expression across utterances and categories. With only 10M trainable parameters,less than 1/30 of full fine-tuning, EmoShift outperforms zero-shot and fully fine-tuned baselines in objective and subjective evaluations, enhancing emotional expressiveness while preserving naturalness and speaker similarity. Further analysis confirms the proposed EmoSteer layer's effectiveness and reveals its potential for controllable emotional intensity in speech synthesis.
【4】CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
标题:CALM:用于多说话者ASB个性化的联合上下文声学语言建模
链接:https://arxiv.org/abs/2601.22792
备注:Accepted to IEEE ICASSP 2026
摘要:我们提出了CALM,一个联合上下文声学语言建模框架的多说话人自动语音识别(ASR)。在个性化的人工智能场景中,声学和语言线索的联合可用性自然会激发目标说话者条件反射与重叠对话中的上下文偏置的整合。CALM通过说话人嵌入驱动的目标说话人提取和基于词汇的动态上下文偏置在端到端框架中实现了这种集成。我们评估CALM模拟英语(LibriSpeechMix)和日语(语料库的自发日语混合物,CSJMix)。在两个扬声器的混合,CALM减少了偏见的单词错误率(B-WER)从12.7到4.7 LibriSpeech 2 Mix和偏见的字符错误率(B-CER)从16.6到8.4 CSJMix 2(eval 3),证明了跨语言的联合声学语言建模的有效性。我们还报告了AMI语料库(IHM混合条件)的结果,以验证标准化语音混合物的性能。
摘要:We present CALM, a joint Contextual Acoustic-Linguistic Modeling framework for multi-speaker automatic speech recognition (ASR). In personalized AI scenarios, the joint availability of acoustic and linguistic cues naturally motivates the integration of target-speaker conditioning with contextual biasing in overlapping conversations. CALM implements this integration in an end-to-end framework through speaker embedding-driven target-speaker extraction and dynamic vocabulary-based contextual biasing. We evaluate CALM on simulated English (LibriSpeechMix) and Japanese (Corpus of Spontaneous Japanese mixtures, CSJMix). On two-speaker mixtures, CALM reduces biased word error rate (B-WER) from 12.7 to 4.7 on LibriSpeech2Mix and biased character error rate (B-CER) from 16.6 to 8.4 on CSJMix2 (eval3), demonstrating the effectiveness of joint acoustic-linguistic modeling across languages. We additionally report results on the AMI corpus (IHM-mix condition) to validate performance on standardized speech mixtures.
【5】Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
标题:使用仅解码器大语言模型和延迟优化的流语音识别
链接:https://arxiv.org/abs/2601.22779
备注:accepted to ICASSP 2026
摘要:最近的进展表明,解码器只有大语言模型(LLM)的自动语音识别(ASR)的潜力。然而,在这个框架内实现流识别仍然是一个挑战。在这项工作中,我们提出了一种新的流ASR方法,集成了读/写策略网络与单调分块注意力(MoChA)动态分段语音嵌入。这些片段在训练期间与标签序列交错,从而实现与LLM的无缝集成。在推断期间,音频流被缓冲,直到MoChA模块触发读取信号,此时缓冲的片段与先前的令牌一起被馈送到LLM中用于下一个令牌预测。我们还引入了最小延迟训练目标,以引导策略网络实现准确的分割边界。此外,我们采用了一种联合训练策略,其中非流式LLM-ASR模型和我们的流式模型共享参数。在AISHELL-1和AISHELL-2普通话基准测试上的实验表明,我们的方法始终优于最近的流ASR基线,分别实现了5.1%和5.5%的字符错误率。延迟优化导致平均令牌生成延迟减少62.5%,对识别准确性的影响可以忽略不计
摘要:Recent advances have demonstrated the potential of decoderonly large language models (LLMs) for automatic speech recognition (ASR). However, enabling streaming recognition within this framework remains a challenge. In this work, we propose a novel streaming ASR approach that integrates a read/write policy network with monotonic chunkwise attention (MoChA) to dynamically segment speech embeddings. These segments are interleaved with label sequences during training, enabling seamless integration with the LLM. During inference, the audio stream is buffered until the MoChA module triggers a read signal, at which point the buffered segment together with the previous token is fed into the LLM for the next token prediction. We also introduce a minimal-latency training objective to guide the policy network toward accurate segmentation boundaries. Furthermore, we adopt a joint training strategy in which a non-streaming LLM-ASR model and our streaming model share parameters. Experiments on the AISHELL-1 and AISHELL-2 Mandarin benchmarks demonstrate that our method consistently outperforms recent streaming ASR baselines, achieving character error rates of 5.1% and 5.5%, respectively. The latency optimization results in a 62.5% reduction in average token generation delay with negligible impact on recognition accuracy
【6】Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources
标题:类感知置换不变信失真比用于具有同类源的声音场景的语义分割
链接:https://arxiv.org/abs/2601.22504
备注:Accepted by ICASSP 2026
摘要:为了推进沉浸式通信,声学场景和事件的检测和分类(DCASE)2025挑战赛最近推出了关于声音场景的空间语义分割(S5)的任务4。S5系统将多声道音频混合作为输入,并输出单声道干源及其相应的类别标签。虽然DCASE 2025挑战通过将每个混合物中的类别标签约束为互斥来简化任务,但现实世界的混合物通常包含来自同一类别的多个来源。重复标签的存在会显著降低标签查询源分离(LQSS)模型的性能,该模型是许多现有S5系统的关键组件,并且还可以限制DCASE 2025任务4的官方评估指标的有效性。为了解决这些问题,我们提出了一个类感知的置换不变损失函数,使LQSS模型处理涉及重复标签的查询。此外,我们重新设计了S5评估指标,以消除这些同类源造成的歧义。为了在S5系统中评估所提出的方法,我们扩展了标签预测模型以支持同类标签。实验结果表明,所提出的方法的有效性和鲁棒性的新度量的混合有和没有同类源。
摘要:To advance immersive communication, the Detection and Classification of Acoustic Scenes and Events (DCASE) 2025 Challenge recently introduced Task 4 on Spatial Semantic Segmentation of Sound Scenes (S5). An S5 system takes a multi-channel audio mixture as input and outputs single-channel dry sources along with their corresponding class labels. Although the DCASE 2025 Challenge simplifies the task by constraining class labels in each mixture to be mutually exclusive, real-world mixtures frequently contain multiple sources from the same class. The presence of duplicated labels can significantly degrade the performance of the label-queried source separation (LQSS) model, which is the key component of many existing S5 systems, and can also limit the validity of the official evaluation metric of DCASE 2025 Task 4. To address these issues, we propose a class-aware permutation-invariant loss function that enables the LQSS model to handle queries involving duplicated labels. In addition, we redesign the S5 evaluation metric to eliminate ambiguities caused by these same-class sources. To evaluate the proposed method within the S5 system, we extend the label prediction model to support same-class labels. Experimental results demonstrate the effectiveness of the proposed methods and the robustness of the new metric on mixtures both with and without same-class sources.
【7】Optimizing Domain-Adaptive Self-Supervised Learning for Clinical Voice-Based Disease Classification
标题:优化领域自适应自我监督学习以实现临床语音疾病分类
链接:https://arxiv.org/abs/2601.22319
备注:Accepted at IEEE ICASSP 2026
摘要:人类声音是一种很有前途的非侵入性数字生物标志物,但基于声音的健康分析的深度学习受到数据稀缺和域不匹配的阻碍,其中在一般音频上预训练的模型无法捕捉临床声音数据的微妙病理特征。为了解决这些挑战,我们研究了具有掩蔽自动编码器(MAE)的域自适应自监督学习(SSL),并证明了标准配置对于健康相关音频来说是次优的。使用Bridge 2AI-Voice数据集,一个多机构收集的病理声音,我们系统地检查了三个性能关键因素:重建损失(平均绝对误差与均方误差),归一化(分片与全局)和掩蔽(随机与内容感知)。我们的优化设计结合了平均绝对误差(MA-Error)损失,分片归一化和内容感知掩蔽,实现了0.688\pm 0.009$的宏F1(超过10次微调运行),优于在大规模通用音频上预训练的强大域外SSL基线,其宏F1为0.663\pm 0.011$。结果表明,MA错误损失提高了鲁棒性,内容感知掩蔽通过强调信息丰富的区域来提高性能。这些发现突出了组件级优化在依赖音频数据的数据受限医疗应用中的重要性。
摘要:The human voice is a promising non-invasive digital biomarker, yet deep learning for voice-based health analysis is hindered by data scarcity and domain mismatch, where models pre-trained on general audio fail to capture the subtle pathological features characteristic of clinical voice data. To address these challenges, we investigate domain-adaptive self-supervised learning (SSL) with Masked Autoencoders (MAE) and demonstrate that standard configurations are suboptimal for health-related audio. Using the Bridge2AI-Voice dataset, a multi-institutional collection of pathological voices, we systematically examine three performance-critical factors: reconstruction loss (Mean Absolute Error vs. Mean Squared Error), normalization (patch-wise vs. global), and masking (random vs. content-aware). Our optimized design, which combines Mean Absolute Error (MA-Error) loss, patch-wise normalization, and content-aware masking, achieves a Macro F1 of $0.688 \pm 0.009$ (over 10 fine-tuning runs), outperforming a strong out-of-domain SSL baseline pre-trained on large-scale general audio, which has a Macro F1 of $0.663 \pm 0.011$. The results show that MA-Error loss improves robustness and content-aware masking boosts performance by emphasizing information-rich regions. These findings highlight the importance of component-level optimization in data-constrained medical applications that rely on audio data.
【8】Sylber 2.0: A Universal Syllable Embedding
标题:Sylber 2.0:通用音节嵌入
链接:https://arxiv.org/abs/2601.22306
摘要:扩展口语建模需要高效且通用的语音令牌。最近的工作提出了音节作为有前途的语音令牌在低时间分辨率,但现有的模型被限制到英语,并未能捕捉到足够的声学细节。为了解决这一差距,我们提出了Sylber 2.0,一个自我监督的框架,用于在音节水平上对语音进行编码,从而实现有效的时间压缩和高保真重建。Sylber 2.0实现了约5 Hz的极低令牌频率,同时保留了多种语言和表达风格的语言和声学细节。实验表明,它与以前的模型在高频基线上运行的性能相当。此外,Sylber 2.0支持高效的TTS建模,仅使用72M参数即可生成具有竞争力的可懂度和质量的语音。此外,Sylber 2.0的通用性为低资源ASR提供了比以前的语音编码框架更有效的功能。总之,我们建立了一个有效的音节级抽象一般口语。
摘要:Scaling spoken language modeling requires speech tokens that are both efficient and universal. Recent work has proposed syllables as promising speech tokens at low temporal resolution, but existing models are constrained to English and fail to capture sufficient acoustic detail. To address this gap, we present Sylber 2.0, a self-supervised framework for coding speech at the syllable level that enables efficient temporal compression and high-fidelity reconstruction. Sylber 2.0 achieves a very low token frequency around 5 Hz, while retaining both linguistic and acoustic detail across multiple languages and expressive styles. Experiments show that it performs on par with previous models operating on high-frequency baselines. Furthermore, Sylber 2.0 enables efficient TTS modeling which can generate speech with competitive intelligibility and quality with SOTA models using only 72M parameters. Moreover, the universality of Sylber 2.0 provides more effective features for low resource ASR than previous speech coding frameworks. In sum, we establish an effective syllable-level abstraction for general spoken language.
【9】Brain-Informed Speech Separation for Cochlear Implants
标题:人工晶状体植入物的大脑知情言语分离
链接:https://arxiv.org/abs/2601.22260
摘要:我们提出了一个人工耳蜗植入(CI),使用脑电图(EEG)派生的注意力线索,引导增强对出席扬声器的大脑知情的语音分离方法。注意力引导网络通过轻量级融合层将音频混合物与EEG特征融合,产生用于CI刺激的关注源电描记图,同时解决仅音频分离器的标签排列模糊性。通过在训练期间改变线索质量的混合课程,可以提高对退化注意力线索的鲁棒性,即使在脑电与语音相关性中等的情况下也能产生稳定的收益。在多说话者条件下,该模型实现了比仅音频电描记图基线更高的信干比改善,同时保持略小(167k与171k参数)。与2毫秒的算法延迟和可比的成本,该方法突出了耦合的听觉和神经线索的认知自适应CI处理的承诺。
摘要:We propose a brain-informed speech separation method for cochlear implants (CIs) that uses electroencephalography (EEG)-derived attention cues to guide enhancement toward the attended speaker. An attention-guided network fuses audio mixtures with EEG features through a lightweight fusion layer, producing attended-source electrodograms for CI stimulation while resolving the label-permutation ambiguity of audio-only separators. Robustness to degraded attention cues is improved with a mixed curriculum that varies cue quality during training, yielding stable gains even when EEG-speech correlation is moderate. In multi-talker conditions, the model achieves higher signal-to-interference ratio improvements than an audio-only electrodogram baseline while remaining slightly smaller (167k vs. 171k parameters). With 2 ms algorithmic latency and comparable cost, the approach highlights the promise of coupling auditory and neural cues for cognitively adaptive CI processing.
【10】Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
标题:重新思考语音增强中的语音表示聚合:语音互信息的角度
链接:https://arxiv.org/abs/2601.22480
备注:Accepted to ICASSP 2026
摘要:最近的语音增强(SE)模型越来越多地利用自监督学习(SSL)表示其丰富的语义信息。通常,中间特征通过轻量级适配模块聚合到单个表示中。然而,大多数SSL模型没有经过噪声鲁棒性训练,这可能导致语义表示损坏。此外,自适应模块与SE模型联合训练,潜在地将声学细节优先于语义信息,这与最初的目的相矛盾。为了解决这个问题,我们首先从信息论的角度分析了SSL模型对噪声语音的行为。具体来说,我们测量损坏的SSL表示和相应的音素标签之间的互信息(MI),专注于保存的语言内容。在此分析的基础上,我们引入了语言聚合层,该层经过预训练,以使用音素标签(可选动态聚合)最大化MI,然后在SE训练期间冻结。实验表明,这种解耦的方法提高了单词错误率(WER)在联合优化的基线,展示了显式对齐的适应模块与语言内容的好处。
摘要:Recent speech enhancement (SE) models increasingly leverage self-supervised learning (SSL) representations for their rich semantic information. Typically, intermediate features are aggregated into a single representation via a lightweight adaptation module. However, most SSL models are not trained for noise robustness, which can lead to corrupted semantic representations. Moreover, the adaptation module is trained jointly with the SE model, potentially prioritizing acoustic details over semantic information, contradicting the original purpose. To address this issue, we first analyze the behavior of SSL models on noisy speech from an information-theoretic perspective. Specifically, we measure the mutual information (MI) between the corrupted SSL representations and the corresponding phoneme labels, focusing on preservation of linguistic contents. Building upon this analysis, we introduce the linguistic aggregation layer, which is pre-trained to maximize MI with phoneme labels (with optional dynamic aggregation) and then frozen during SE training. Experiments show that this decoupled approach improves Word Error Rate (WER) over jointly optimized baselines, demonstrating the benefit of explicitly aligning the adaptation module with linguistic contents.
【11】An Effective Energy Mask-based Adversarial Evasion Attacks against Misclassification in Speaker Recognition Systems
标题:基于有效能量面具的对抗规避攻击说话人识别系统中的误分类
链接:https://arxiv.org/abs/2601.22390
摘要:规避攻击对人工智能系统构成重大威胁,利用机器学习模型中的漏洞绕过检测机制。目前,在有前途的未来行业中,语音数据(包括deepfake)的广泛使用受到法律框架不足的阻碍。对抗性攻击方法已成为防止滥用此类数据的最有效对策。本文介绍了一种利用功率谱对原始语音数据进行能量掩蔽的新方法-掩蔽能量扰动法。MEP在生成对抗性扰动之前对频域中的小能量区域应用掩蔽,目标是人类听觉模型不太明显的区域。该研究主要采用先进的说话人识别模型,包括ECAPA-TDNN和ResNet 34,这些模型在说话人验证任务中表现出出色的性能。所提出的MEP方法在音频质量和规避有效性方面表现出很强的性能。能量掩蔽方法有效地最大限度地减少了语音质量的感知评估(PESQ)的退化,这表明尽管有对抗性扰动,但人类收听者的感知失真最小。具体而言,在PESQ评估中,与快速梯度符号法(FGSM)和迭代FGSM相比,MEP方法的相对性能为26.68%。
摘要:Evasion attacks pose significant threats to AI systems, exploiting vulnerabilities in machine learning models to bypass detection mechanisms. The widespread use of voice data, including deepfakes, in promising future industries is currently hindered by insufficient legal frameworks. Adversarial attack methods have emerged as the most effective countermeasure against the indiscriminate use of such data. This research introduces masked energy perturbation (MEP), a novel approach using power spectrum for energy masking of original voice data. MEP applies masking to small energy regions in the frequency domain before generating adversarial perturbations, targeting areas less noticeable to the human auditory model. The study primarily employs advanced speaker recognition models, including ECAPA-TDNN and ResNet34, which have shown remarkable performance in speaker verification tasks. The proposed MEP method demonstrated strong performance in both audio quality and evasion effectiveness. The energy masking approach effectively minimizes the perceptual evaluation of speech quality (PESQ) degradation, indicating that minimal perceptual distortion occurs to the human listener despite the adversarial perturbations. Specifically, in the PESQ evaluation, the relative performance of the MEP method was 26.68% when compared to the fast gradient sign method (FGSM) and iterative FGSM.
【12】PersonaCite: VoC-Grounded Interviewable Agentic Synthetic AI Personas for Verifiable User and Design Research
标题:PersonaCite:基于VoC的可采访抽象合成人工智能角色,用于可验证的用户和设计研究
链接:https://arxiv.org/abs/2601.22288
摘要:基于LLM和基于代理的合成人物角色越来越多地用于设计和产品决策,但先前的工作表明,基于LLM的人物角色通常会产生有说服力但无法验证的响应,从而掩盖其证据基础。我们提出了PersonaCite,这是一个代理系统,通过检索增强交互将AI人物角色重新构建为证据约束的研究工具。与依赖于基于角色扮演的先前方法不同,PersonaCite在每个会话回合期间检索实际的客户之声工件,约束对检索到的证据的响应,在证据缺失时明确弃权,并提供响应级别的源归因。通过对14位行业专家的半结构化访谈和部署研究,我们确定了关于感知利益、有效性问题和设计紧张局势的初步发现,并提出了Persona起源卡作为在以人为本的设计工作流中使用负责任的AI角色的文档模式。
摘要:LLM-based and agent-based synthetic personas are increasingly used in design and product decision-making, yet prior work shows that prompt-based personas often produce persuasive but unverifiable responses that obscure their evidentiary basis. We present PersonaCite, an agentic system that reframes AI personas as evidence-bounded research instruments through retrieval-augmented interaction. Unlike prior approaches that rely on prompt-based roleplaying, PersonaCite retrieves actual voice-of-customer artifacts during each conversation turn, constrains responses to retrieved evidence, explicitly abstains when evidence is missing, and provides response-level source attribution. Through semi-structured interviews and deployment study with 14 industry experts, we identify preliminary findings on perceived benefits, validity concerns, and design tensions, and propose Persona Provenance Cards as a documentation pattern for responsible AI persona use in human-centered design workflows.
【13】Proliferating series by Jean Barraqué: a study and classification in mathematical terms
标题:让·巴拉奎(Jean Barraqué)的沸腾系列:数学术语的研究和分类
链接:https://arxiv.org/abs/2601.22176
备注:28 pages, 8 figures
摘要:Barraqué的增殖系列通过在构建系列时创建一个新的不变量,对经典序列主义的概念进行了有趣的转变:在给定基本序列的增殖的构造期间保持不变的不是连续音符之间的间隔,而是发生在两个连续序列之间的音符的排列,也就是说,系列中音符顺序的转换。这为对序列方法感兴趣的作曲家提供了新的可能性,因为通过这种方法获得的间隔的多样性远远大于经典的序列主义。在这份手稿中,我们将从数学的角度研究激增的系列提供的一些未探索的可能性,这将使作曲家对它们更加熟悉,并可能导致创作出将序列主义提升到下一个层次的作品。
摘要:Barraqué's proliferating series give an interesting turn on the concept of classic serialism by creating a new invariant when it comes to constructing the series: rather than the intervals between consecutive notes, what remains unaltered during the construction of the proliferations of the given base series is the permutation of the notes which happens between two consecutive series, that is to say, the transformation of the order of the notes in the series. This presents new possibilities for composers interested in the serial method, given the fact that the variety of intervals obtained by this method is far greater than that of classic serialism. In this manuscript, we will study some unexplored possibilities that the proliferating series offer from a mathematical point of view, which will allow composers to gain much more familiarity with them and potentially result in the creation of pieces that take serialism to the next level.
【14】Attention Isn't All You Need for Emotion Recognition:Domain Features Outperform Transformers on the EAV Dataset
标题:情感识别需要的不仅仅是注意力:领域功能在EAV数据集中优于Transformer
链接:https://arxiv.org/abs/2601.22161
摘要:我们提出了一个系统的研究多模态情感识别使用EAV数据集,调查是否复杂的注意力机制提高性能的小数据集。我们实现了三个模型类别:基线Transformers(M1),新型因子化注意力机制(M2)和改进的CNN基线(M3)。我们的实验表明,复杂的注意力机制在小数据集上表现不佳。由于过度拟合和预训练特征的破坏,M2模型比基线低5到13个百分点。相比之下,简单的域适当的修改被证明是有效的:将增量MFCC添加到音频CNN中将准确率从61.9\%提高到\textbf{65.56\%}(+3.66pp),而EEG的频域特征达到\textbf{67.62\%}(比论文基线高出+7.62pp)。我们的Vision Transformer基线(M1)达到了\textbf{75.30\%},超过了论文的ViViT结果(74.5\%),通过特定领域的预训练,视觉增量特征达到了\textbf{72.68\%}(比论文CNN增加了1.28 pp)。这些发现表明,对于小规模的情感识别,领域知识和适当的实现优于架构的复杂性。
摘要:We present a systematic study of multimodal emotion recognition using the EAV dataset, investigating whether complex attention mechanisms improve performance on small datasets. We implement three model categories: baseline transformers (M1), novel factorized attention mechanisms (M2), and improved CNN baselines (M3). Our experiments show that sophisticated attention mechanisms consistently underperform on small datasets. M2 models achieved 5 to 13 percentage points below baselines due to overfitting and destruction of pretrained features. In contrast, simple domain-appropriate modifications proved effective: adding delta MFCCs to the audio CNN improved accuracy from 61.9\% to \textbf{65.56\%} (+3.66pp), while frequency-domain features for EEG achieved \textbf{67.62\%} (+7.62pp over the paper baseline). Our vision transformer baseline (M1) reached \textbf{75.30\%}, exceeding the paper's ViViT result (74.5\%) through domain-specific pretraining, and vision delta features achieved \textbf{72.68\%} (+1.28pp over the paper CNN). These findings demonstrate that for small-scale emotion recognition, domain knowledge and proper implementation outperform architectural complexity.
机器翻译由腾讯交互翻译提供,仅供参考
