今日论文合集:cs.SD语音16篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Gen-SER: When the generative model meets speech emotion recognition
标题:Gen-BER:当生成模型遇到语音情感识别时
链接:https://arxiv.org/abs/2601.20573

作者:Taihui Wang,Jinzheng Zhao,Rilin Chen,Tong Lei,Wenwu Wang,Dong Yu
备注:Accepted to IEEE ICASSP 2026
摘要:语音情感识别是语音理解和生成的关键。大多数方法基于分类模型或大型语言模型。与以往的方法不同,我们提出了Gen-SER,一种新的方法,通过生成模型将SER重新表述为分布偏移问题。我们建议将离散的类别标签投影到一个连续的空间中,并通过正弦分类编码获得终端分布。采用基于目标匹配的生成模型,有效地将初始分布转化为最终分布。通过计算生成的终端分布和地面实况终端分布的相似性来实现分类。实验结果证实了所提出的方法的有效性,证明了其可扩展性的各种语音理解任务,并建议其潜在的适用范围更广的分类任务。
摘要:Speech emotion recognition (SER) is crucial in speech understanding and generation. Most approaches are based on either classification models or large language models. Different from previous methods, we propose Gen-SER, a novel approach that reformulates SER as a distribution shift problem via generative models. We propose to project discrete class labels into a continuous space, and obtain the terminal distribution via sinusoidal taxonomy encoding. The target-matching-based generative model is adopted to transform the initial distribution into the terminal distribution efficiently. The classification is achieved by calculating the similarity of the generated terminal distribution and ground truth terminal distribution. The experimental results confirm the efficacy of the proposed method, demonstrating its extensibility to various speech-understanding tasks and suggesting its potential applicability to a broader range of classification tasks.


【2】Audio Deepfake Detection in the Age of Advanced Text-to-Speech models
标题:高级文本到语音模型时代的音频深度伪造检测
链接:https://arxiv.org/abs/2601.20510

作者:Robin Singh,Aditya Yogesh Nair,Fabio Palumbo,Florian Barbaro,Anna Dyka,Lohith Rachakonda
备注:This work was performed using HPC resources from GENCI-IDRIS (Grant 2025- AD011016076)
摘要:文本到语音(TTS)系统的最新进展大大提高了合成语音的真实性,为音频深度伪造检测提出了新的挑战。这项工作提出了一个比较评估的三个国家的最先进的TTS模型-Dia 2,Maya 1和MeloTTS-代表流,基于LLM,和非自回归架构。使用Daily-Dialog数据集生成了12,000个合成音频样本的语料库,并针对四个检测框架进行了评估,包括语义,结构和信号级方法。结果显示,在整个生成机制的检测器性能的显着变化:对一个TTS架构有效的模型可能会失败对其他人,特别是基于LLM的合成。相比之下,结合互补分析水平的多视图检测方法在所有评估的模型中表现出稳健的性能。这些发现强调了单一范式检测器的局限性,并强调了集成检测策略的必要性,以应对不断变化的音频deepfake威胁。
摘要:Recent advances in Text-to-Speech (TTS) systems have substantially increased the realism of synthetic speech, raising new challenges for audio deepfake detection. This work presents a comparative evaluation of three state-of-the-art TTS models--Dia2, Maya1, and MeloTTS--representing streaming, LLM-based, and non-autoregressive architectures. A corpus of 12,000 synthetic audio samples was generated using the Daily-Dialog dataset and evaluated against four detection frameworks, including semantic, structural, and signal-level approaches. The results reveal significant variability in detector performance across generative mechanisms: models effective against one TTS architecture may fail against others, particularly LLM-based synthesis. In contrast, a multi-view detection approach combining complementary analysis levels demonstrates robust performance across all evaluated models. These findings highlight the limitations of single-paradigm detectors and emphasize the necessity of integrated detection strategies to address the evolving landscape of audio deepfake threats.


【3】On Every Note a Griff: Looking for a Useful Representation of Basso Continuo Performance Style
标题:每个音符上都有一个格里夫:寻找巴索连续演奏风格的有用表现
链接:https://arxiv.org/abs/2601.20478

作者:Adam Štefunko,Carlos Eduardo Cancino-Chacón,Jan Hajič
备注:6 pages, 5 figures, accepted to the Music Encoding Conference (MEC) 2026
摘要:通奏低音是一种巴洛克风格的即兴伴奏风格,它涉及在大键琴或管风琴上即兴演奏乐谱中给定低音线以上的多个部分。连续低音不仅仅是一个历史问题;此外,它是一个受历史启发的生活实践,而The Aligned Continuo Dataset(ACoRD)记录了第一个现代连续低音在符号领域演奏的样本。这个数据集包含了7位演奏者演奏的5首连续低音乐谱的175段录音,让我们可以开始观察和分析连续低音即兴演奏带来的多样性。最近提出的连续低音演奏与乐谱对齐系统提供了一种将即兴演奏音符映射到乐谱音符的方法。为了研究对齐的连续低音演奏,我们需要一个适当的特征表示。我们提出格里夫,一个代表性的灵感来自历史的连续低音论文。它使我们能够编码的音高内容和结构的连续低音实现的转置不变的方式。Griffs是直接从对齐的连续低音演奏中提取的,通过将与同一乐谱音符对齐的演奏音符以起始时间排序的方式分组在一起,它们提供了有意义的标记,形成了一个特征空间,我们可以在其中分析连续低音演奏风格。我们统计描述griffs提取的ACoRD数据集记录,并显示在两个实验中griffs如何可以用于统计分析不同的球员的连续低音演奏风格的个性。最后,我们提出了一个论点,为什么它是可取的,以保持连续低音即兴的结构,以进行个人的连续低音练习者的个人演奏风格的精细分析,以及为什么griffs可以提供一个有意义的历史知情的功能空间值得更强大的经验验证。
摘要:Basso continuo is a baroque improvisatory accompaniment style which involves improvising multiple parts above a given bass line in a musical score on a harpsichord or organ. Basso continuo is not merely a matter of history; moreover, it is a historically inspired living practice, and The Aligned Continuo Dataset (ACoRD) records the first sample of modern-day basso continuo playing in the symbolic domain. This dataset, containing 175 MIDI recordings of 5 basso continuo scores performed by 7 players, allows us to start observing and analyzing the variety that basso continuo improvisation brings. A recently proposed basso continuo performance-to-score alignment system provides a way of mapping improvised performance notes to score notes. In order to study aligned basso continuo performances, we need an appropriate feature representation. We propose griff, a representation inspired by historical basso continuo treatises. It enables us to encode both pitch content and structure of a basso continuo realization in a transposition-invariant way. Griffs are directly extracted from aligned basso continuo performances by grouping together performance notes aligned to the same score note in a onset-time ordered way, and they provide meaningful tokens that form a feature space in which we can analyze basso continuo performance styles. We statistically describe griffs extracted from the ACoRD dataset recordings, and show in two experiments how griffs can be used for statistical analysis of individuality of different players' basso continuo performance styles. We finally present an argument why it is desirable to preserve the structure of a basso continuo improvisation in order to conduct a refined analysis of personal performance styles of individual basso continuo practitioners, and why griffs can provide a meaningful historically informed feature space worthy of a more robust empirical validation.


【4】Self Voice Conversion as an Attack against Neural Audio Watermarking
标题:自语音转换作为对神经音频水印的攻击
链接:https://arxiv.org/abs/2601.20432

作者:Yigitcan Özer,Wanying Ge,Zhe Zhang,Xin Wang,Junichi Yamagishi
备注:7 pages; 2 figures; 2 tables; accepted at IEICE, SP/SLP 2026
摘要:音频水印技术在语音中嵌入辅助信息,同时保持说话人身份、语言内容和感知质量。虽然神经和数字信号处理为基础的水印方法的最新进展,提高了不可感知性和嵌入能力,鲁棒性仍然主要评估对传统的失真,如压缩,加性噪声,和resstrike。然而,基于深度学习的攻击的兴起给水印安全带来了新的重大威胁。在这项工作中,我们调查自己的声音转换作为一个普遍的,内容保持攻击音频水印系统。自语音转换将说话者的语音重新映射到相同的身份,同时通过语音转换模型改变声学特性。我们证明,这种攻击严重降低了国家的最先进的水印方法的可靠性,并强调其对现代音频水印技术的安全性的影响。
摘要:Audio watermarking embeds auxiliary information into speech while maintaining speaker identity, linguistic content, and perceptual quality. Although recent advances in neural and digital signal processing-based watermarking methods have improved imperceptibility and embedding capacity, robustness is still primarily assessed against conventional distortions such as compression, additive noise, and resampling. However, the rise of deep learning-based attacks introduces novel and significant threats to watermark security. In this work, we investigate self voice conversion as a universal, content-preserving attack against audio watermarking systems. Self voice conversion remaps a speaker's voice to the same identity while altering acoustic characteristics through a voice conversion model. We demonstrate that this attack severely degrades the reliability of state-of-the-art watermarking approaches and highlight its implications for the security of modern audio watermarking techniques.


【5】Mix2Morph: Learning Sound Morphing from Noisy Mixes
标题:Mix 2 Morph:从Noisy Mixes中学习声音变形
链接:https://arxiv.org/abs/2601.20426

作者:Annie Chu,Hugo Flores García,Oriol Nieto,Justin Salamon,Bryan Pardo,Prem Seetharaman
备注:Accepted into ICASSP 2026
摘要:我们引入了Mix2Morph,这是一个文本到音频的扩散模型,经过微调,可以在没有专用变形数据集的情况下执行声音变形。通过在更高的扩散时间步长上对有噪声的替代混合进行微调,Mix2Morph产生稳定的、感知上连贯的变形,令人信服地整合了两个源的质量。我们专门针对声音注入,一个实际和感知动机的变形子类,其中一个声音作为主导的主要来源,提供整体的时间和结构行为,而第二个声音被注入整个,丰富其音色和纹理质量。客观的评估和听力测试表明,Mix2Morph的表现优于之前的基线,并在不同的类别中产生高质量的声音注入,代表着朝着更可控和概念驱动的声音设计工具迈出了一步。在https://anniejchu.github.io/mix2morph上可以找到正确的例子。
摘要:We introduce Mix2Morph, a text-to-audio diffusion model fine-tuned to perform sound morphing without a dedicated dataset of morphs. By finetuning on noisy surrogate mixes at higher diffusion timesteps, Mix2Morph yields stable, perceptually coherent morphs that convincingly integrate qualities of both sources. We specifically target sound infusions, a practically and perceptually motivated subclass of morphing in which one sound acts as the dominant primary source, providing overall temporal and structural behavior, while a secondary sound is infused throughout, enriching its timbral and textural qualities. Objective evaluations and listening tests show that Mix2Morph outperforms prior baselines and produces high-quality sound infusions across diverse categories, representing a step toward more controllable and concept-driven tools for sound design. Sound examples are available at https://anniejchu.github.io/mix2morph .


【6】Switchcodec: Adaptive residual-expert sparse quantization for high-fidelity neural audio coding
标题:Switch Codec:用于高保真神经音频编码的自适应剩余专家稀疏量化
链接:https://arxiv.org/abs/2601.20362

作者:Xiangbo Wang,Wenbin Jiang,Jin Wang,Yubo You,Sheng Fang,Fei Wen
备注:4page,3figure,Accepted by ICASSP 2026,We would like to express our sincere gratitude to Senior Fellow Jing Wang for his continuous support and assistance. He has made an indelible and significant contribution to this work
摘要:最近的神经音频压缩模型往往依赖于残差矢量量化高保真编码,但使用固定数量的每帧码本是次优的音频内容的广泛变化,特别是对于信号是非常简单或非常复杂。为了解决这一限制,我们提出了SwitchCodec,这是一种基于残差专家矢量量化(REVQ)的神经音频编解码器。REVQ将共享量化器与根据输入音频激活的动态路由专家量化器相结合,将比特率与码本容量解耦,并提高压缩效率。这种设计确保了每个量化器的充分训练和利用。此外,可变比特率机制在推理时调整活动专家量化器的数量,从而实现多比特率操作而无需再训练。实验表明,SwitchCodec在客观指标和主观听力测试上都超过了现有的基线。
摘要:Recent neural audio compression models often rely on residual vector quantization for high-fidelity coding, but using a fixed number of per-frame codebooks is suboptimal for the wide variability of audio content-especially for signals that are either very simple or highly complex. To address this limitation, we propose SwitchCodec, a neural audio codec based on Residual Experts Vector Quantization (REVQ). REVQ combines a shared quantizer with dynamically routed expert quantizers that are activated according to the input audio, decoupling bitrate from codebook capacity and improving compression efficiency. This design ensures full training and utilization of each quantizer. In addition, a variable-bitrate mechanism adjusts the number of active expert quantizers at inference, enabling multi-bitrate operation without retraining. Experiments demonstrate that SwitchCodec surpasses existing baselines on both objective metrics and subjective listening tests.


【7】Improving X-Codec-2.0 for Multi-Lingual Speech: 25 Hz Latent Rate and 24 kHz Sampling
标题:改进多语言语音的X-Codec-2.0:25 Hz潜在速率和24 GHz采样
链接:https://arxiv.org/abs/2601.20185

作者:Husein Zolkepli
摘要:X-Codec-2.0在神经音频压缩和多语言语音建模方面表现出了强大的性能,使用冻结的HuBERT功能以50 Hz的潜伏速率和16 kHz的采样率运行。虽然有效,但是这种配置限制了时间效率和音频保真度。在这项工作中,我们探索了一个简单而有效的修改,通过引入额外的池化和增加解码器跳数。这将潜在速率从50 Hz降低到25 Hz,同时将输出采样速率从16 kHz提高到24 kHz,在不改变内核架构的情况下提高了效率和感知质量。在多语言Common Voice 17测试集上进行了评估,所提出的配置比基于UTMOSv 2的原始X-Codec-2.0基线提高了0.29 MOS,并在所有以25 Hz工作的编解码器中获得了最佳性能。源代码、检查点和生成比较在\href{https://huggingface.co/Scicom-intl/xcodec2-25TPS-24k}{https://huggingface.co/Scicom-intl/xcodec2-25TPS-24k}上发布。
摘要:X-Codec-2.0 has shown strong performance in neural audio compression and multilingual speech modeling, operating at a 50 Hz latent rate and a 16 kHz sampling rate using frozen HuBERT features. While effective, this configuration limits temporal efficiency and audio fidelity. In this work, we explore a simple and effective modification by introducing additional pooling and increasing the decoder hop size. This reduces the latent rate from 50 Hz to 25 Hz and simultaneously raises the output sampling rate from 16 kHz to 24 kHz, improving efficiency and perceptual quality without altering the core architecture. Evaluated on the multilingual Common Voice 17 test set, the proposed configuration achieves a 0.29 MOS improvement over the original X-Codec-2.0 baseline based on UTMOSv2, and attains the best reported performance among all codecs operating at 25 Hz. The source code, checkpoints, and generation comparisons are released at \href{https://huggingface.co/Scicom-intl/xcodec2-25TPS-24k}{https://huggingface.co/Scicom-intl/xcodec2-25TPS-24k}.


【8】Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR
标题:注意转变:使用Delta SSL嵌入式增强儿童ASB
链接:https://arxiv.org/abs/2601.20142

作者:Zilai Wang,Natarajan Balaji Shankar,Kaiyuan Zhang,Zihan Wang,Abeer Alwan
备注:ICASSP 2026
摘要:自监督学习(SSL)模型在许多语音任务中取得了令人印象深刻的结果,但由于数据有限和预训练域不匹配,儿童自动语音识别(ASR)仍然具有挑战性。微调SSL模型的儿童语音诱导的代表性空间的变化。我们假设delta SSL嵌入(定义为来自微调模型的嵌入与来自其预训练对应模型的嵌入之间的差异)编码特定于任务的信息,以补充来自另一个SSL模型的微调特征。我们使用不同的模型在MyST儿童语料库上评估多种融合策略。结果表明,与微调嵌入融合相比,WavLM的delta嵌入融合对HuBERT产生了高达10%的相对WER降低,对W2V2产生了4.4%的降低。值得注意的是,融合WavLM与delta W2V2嵌入实现了9.64的WER,在MyST语料库上的SSL模型中树立了新的艺术水平。这些发现证明了增量嵌入和突出特征融合的有效性,作为推进儿童ASR的一个有前途的方向。
摘要:Self-supervised learning (SSL) models have achieved impressive results across many speech tasks, yet child automatic speech recognition (ASR) remains challenging due to limited data and pretraining domain mismatch. Fine-tuning SSL models on child speech induces shifts in the representation space. We hypothesize that delta SSL embeddings, defined as the differences between embeddings from a finetuned model and those from its pretrained counterpart, encode task-specific information that complements finetuned features from another SSL model. We evaluate multiple fusion strategies on the MyST childrens corpus using different models. Results show that delta embedding fusion with WavLM yields up to a 10 percent relative WER reduction for HuBERT and a 4.4 percent reduction for W2V2, compared to finetuned embedding fusion. Notably, fusing WavLM with delta W2V2 embeddings achieves a WER of 9.64, setting a new state of the art among SSL models on the MyST corpus. These findings demonstrate the effectiveness of delta embeddings and highlight feature fusion as a promising direction for advancing child ASR.


【9】LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning
标题:LTS-VoiceAgent:一个通过语义触发和增量推理进行高效流媒体语音交互的听-想-说框架
链接:https://arxiv.org/abs/2601.19952

作者:Wenhao Zou,Yuwei Miao,Zhanyu Ma,Jun Xu,Jiuchong Gao,Jinghua Hao,Renqing He,Jingwen Xu
摘要:实时语音代理面临着一个困境:端到端模型通常缺乏深度推理,而级联管道通过严格按顺序执行ASR,LLM推理和TTS而导致高延迟,这与人类对话不同,在人类对话中,听众通常在说话者完成之前开始思考。由于级联架构仍然是复杂任务的主要选择,因此现有的级联流传输策略试图通过机械分段(例如,固定块、基于VAD的分割)或推测性生成,但它们经常要么破坏语义单元,要么在必须回滚的预测上浪费计算。为了解决这些挑战,我们提出了LTS-语音代理,一个听,想,说框架,明确地分离时,从如何逐步推理思考。它具有一个动态语义触发器来检测有意义的前缀,以及一个双角色流推理器,它协调后台思想者(用于状态维护)和前台发言者(用于推测性解决)。这种并行设计使“边思考边说话”不会阻碍响应。我们还介绍了一个暂停和修复基准包含自然的不流畅压力测试流鲁棒性。VERA,Spoken-MQA,BigBenchAudio和我们的基准测试的实验表明,LTS-VoiceAgent实现了比串行级联基线和现有流媒体策略更强的准确性-延迟-效率权衡。
摘要:Real-time voice agents face a dilemma: end-to-end models often lack deep reasoning, while cascaded pipelines incur high latency by executing ASR, LLM reasoning, and TTS strictly in sequence, unlike human conversation where listeners often start thinking before the speaker finishes. Since cascaded architectures remain the dominant choice for complex tasks, existing cascaded streaming strategies attempt to reduce this latency via mechanical segmentation (e.g., fixed chunks, VAD-based splitting) or speculative generation, but they frequently either break semantic units or waste computation on predictions that must be rolled back. To address these challenges, we propose LTS-VoiceAgent, a Listen-Think-Speak framework that explicitly separates when to think from how to reason incrementally. It features a Dynamic Semantic Trigger to detect meaningful prefixes, and a Dual-Role Stream Orchestrator that coordinates a background Thinker (for state maintenance) and a foreground Speaker (for speculative solving). This parallel design enables "thinking while speaking" without blocking responses. We also introduce a Pause-and-Repair benchmark containing natural disfluencies to stress-test streaming robustness. Experiments across VERA, Spoken-MQA, BigBenchAudio, and our benchmark show that LTS-VoiceAgent achieves a stronger accuracy-latency-efficiency trade-off than serial cascaded baselines and existing streaming strategies.


【10】Pianoroll-Event: A Novel Score Representation for Symbolic Music
标题:钢琴演奏事件:一种新的象征性音乐乐谱表现形式
链接:https://arxiv.org/abs/2601.19951

作者:Lekai Qian,Haoyu Gu,Dehan Li,Boyu Cao,Qi Liu
摘要:音乐符号表示是计算音乐学的一个基本挑战。虽然基于网格的表示有效地保持了基音时间的空间对应性,但其固有的数据稀疏性导致编码效率低。离散事件表示实现了紧凑的编码,但未能充分捕捉结构不变性和空间局部性。为了解决这些互补的限制,我们提出了Pianoroll事件,一种新的编码方案,通过事件描述pianoroll表示,结合结构特性与编码效率,同时保持时间依赖性和局部空间模式。具体来说,我们设计了四个互补的事件类型:帧事件的时间边界,间隙事件的稀疏区域,模式事件的注意模式,音乐元数据的音乐结构事件。Pianoroll-Event在序列长度和词汇量之间取得了有效的平衡,编码效率比典型的离散序列方法提高了1.36 ~ 7.16倍。跨多个自回归架构的实验表明,使用我们的表示的模型在定量和人工评估中始终优于基线。
摘要:Symbolic music representation is a fundamental challenge in computational musicology. While grid-based representations effectively preserve pitch-time spatial correspondence, their inherent data sparsity leads to low encoding efficiency. Discrete-event representations achieve compact encoding but fail to adequately capture structural invariance and spatial locality. To address these complementary limitations, we propose Pianoroll-Event, a novel encoding scheme that describes pianoroll representations through events, combining structural properties with encoding efficiency while maintaining temporal dependencies and local spatial patterns. Specifically, we design four complementary event types: Frame Events for temporal boundaries, Gap Events for sparse regions, Pattern Events for note patterns, and Musical Structure Events for musical metadata. Pianoroll-Event strikes an effective balance between sequence length and vocabulary size, improving encoding efficiency by 1.36\times to 7.16\times over representative discrete sequence methods. Experiments across multiple autoregressive architectures show models using our representation consistently outperform baselines in both quantitative and human evaluations.


【11】FastWhisper: Adaptive Self-knowledge Distillation for Real-time Automatic Speech Recognition
标题:FastWhisper:用于实时自动语音识别的自适应自我知识提取
链接:https://arxiv.org/abs/2601.19919

作者:Junseok Lee,Nahoon Kim,Sangyong Lee,Chang-Jae Chun
摘要:知识提取是模型压缩的有效方法之一。以往的研究主要集中在学生模型有效训练教师模型的预测分布上。然而,在训练过程中,学生模型可能继承教师模型的缺点,这可能导致泛化能力下降。为了缓解这个问题,我们提出了自适应自我知识蒸馏(ASKD),它动态地减少了教师模型的依赖性,以提高自我训练能力,并执行自我知识蒸馏方法,以提高学生模型的泛化能力。我们进一步将Whisper模型提炼成一个更小的变体,称为FastWhisper。在我们的训练后设置中,FastWhisper的单词错误率比教师模型Whisper低1.07%,其相对推理时间快5倍。
摘要:Knowledge distillation is one of the most effective methods for model compression. Previous studies have focused on the student model effectively training the predictive distribution of the teacher model. However, during training, the student model may inherit the shortcomings of the teacher model, which can lead to a decline in generalization capacity. To mitigate this issue, we propose adaptive self-knowledge distillation (ASKD), which dynamically reduces the dependence of the teacher model to improve the self-training capacity, and performs the self-knowledge distillation method to improve the generalization capacity of the student model. We further distill the Whisper model into a smaller variant, called FastWhisper. In our post-training setting, FastWhisper achieved a word error rate of 1.07% lower than the teacher model Whisper, and its relative inference time was 5 times faster.


【12】Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
标题:在听到之前消除你的声音:免训练的说话者放弃Zero-Shot文本到语音的学习
链接:https://arxiv.org/abs/2601.20481

作者:Myungjin Lee,Eunji Shin,Jiyoung Lee
备注:ICASSP'2026
摘要:现代zero-shot文本到语音(TTS)模型提供了前所未有的表达能力,但也带来了严重的犯罪风险,因为它们可以合成从未同意过的个人的声音。在这种情况下,说话人遗忘的目的是防止产生特定的发言人身份的请求。现有的方法,依赖于再培训,是昂贵的,仅限于在培训中看到的扬声器。我们提出了TruS,一个免训练的说话人遗忘框架,将范式从数据删除转移到推理时间控制。TruS引导身份特定的隐藏激活以抑制目标说话者,同时保留其他属性(例如,韵律和情感)。实验结果表明,TruS有效地防止语音生成的可见和不可见的选择退出扬声器,建立一个可扩展的语音合成的保障。演示和代码可在http://mmai.ewha.ac.kr/trus上获取。
摘要:Modern zero-shot text-to-speech (TTS) models offer unprecedented expressivity but also pose serious crime risks, as they can synthesize voices of individuals who never consented. In this context, speaker unlearning aims to prevent the generation of specific speaker identities upon request. Existing approaches, reliant on retraining, are costly and limited to speakers seen in the training set. We present TruS, a training-free speaker unlearning framework that shifts the paradigm from data deletion to inference-time control. TruS steers identity-specific hidden activations to suppress target speakers while preserving other attributes (e.g., prosody and emotion). Experimental results show that TruS effectively prevents voice generation on both seen and unseen opt-out speakers, establishing a scalable safeguard for speech synthesis. The demo and code are available on http://mmai.ewha.ac.kr/trus.


【13】Do we really need Self-Attention for Streaming Automatic Speech Recognition?
标题:我们真的需要自我关注来实现流媒体自动语音识别吗?
链接:https://arxiv.org/abs/2601.19960

作者:Youness Dkhissi,Valentin Vielzeuf,Elys Allesiardo,Anthony Larcher
摘要:基于transformer的架构是许多深度学习领域中最常用的架构,如自然语言处理,计算机视觉或语音处理。它可以鼓励在受约束的任务中直接使用Transformers,而不质疑它是否会产生与标准任务相同的好处。  给定特定的约束条件,必须评估Transformer模型的相关性。这项工作质疑Transformers对特定领域的适用性。我们认为,与这些模型相关的高计算要求和延迟问题与流媒体应用程序不一致。我们的研究促进了寻找替代策略,以提高效率,而不牺牲性能。  根据这一观察,我们的论文批判性地探讨了Transformer架构在这种受限环境中的实用性。作为第一次尝试,我们表明,流自动语音识别(ASR)的计算成本可以减少使用变形卷积,而不是自我注意。此外,我们表明,自我注意机制可以完全删除,而不是取代,没有观察到显着的字错误率下降。
摘要:Transformer-based architectures are the most used architectures in many deep learning fields like Natural Language Processing, Computer Vision or Speech processing. It may encourage the direct use of Transformers in the constrained tasks, without questioning whether it will yield the same benefits as in standard tasks.  Given specific constraints, it is essential to evaluate the relevance of transformer models. This work questions the suitability of transformers for specific domains. We argue that the high computational requirements and latency issues associated with these models do not align well with streaming applications. Our study promotes the search for alternative strategies to improve efficiency without sacrificing performance.  In light of this observation, our paper critically examines the usefulness of transformer architecture in such constrained environments. As a first attempt, we show that the computational cost for Streaming Automatic Speech Recognition (ASR) can be reduced using deformable convolution instead of Self-Attention. Furthermore, we show that Self-Attention mechanisms can be entirely removed and not replaced, without observing significant degradation in the Word Error Rate.


【14】VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
标题:VoxPrivacy:评估语音语言模型交互隐私的基准
链接:https://arxiv.org/abs/2601.19956

作者:Yuxiang Wang,Hongyu Liu,Dekun Chen,Xueyao Zhang,Zhizheng Wu
摘要:随着语音语言模型(SLM)从个人设备过渡到智能家居等共享的多用户环境,一个新的挑战出现了:该模型预计将区分用户以适当地管理信息流。如果没有这种能力,SLM可能会将一个用户的机密日程泄露给另一个用户,这是一种隐私失败,我们称之为互操作隐私。因此,生成说话者感知响应的能力对于SLM安全部署至关重要。目前的SLM基准测试对话能力,但忽视了发言者的身份。多说话者基准检查谁说了什么,而不评估SLM是否适应他们的反应。隐私基准侧重于全球敏感数据(例如,银行密码)而忽略上下文隐私敏感信息(例如,用户的私人约会)。为了解决这一差距,我们引入VoxPrivacy,第一个基准设计来评估在SLM的隐私。VoxPrivacy跨越了三个层次,从遵循直接的保密命令到主动保护隐私。我们在32小时双语数据集上对9个SLM进行的评估揭示了一个普遍存在的漏洞:大多数开源模型在有条件的隐私决策上的表现接近随机机会(约50%的准确率),而即使是强大的闭源系统也无法进行主动隐私推断。我们进一步验证了这些发现的真实VoxPrivacy,一个人类记录的子集,确认在合成数据上观察到的失败在真实语音中持续存在。最后,我们展示了一条可行的前进道路:通过对新的4,000小时训练集进行微调,我们在保持鲁棒性的同时提高了隐私保护能力。为了支持未来的工作,我们发布了VoxPrivacy基准测试、大规模训练集和微调模型,以促进更安全、更上下文感知的SLM的开发。
摘要:As Speech Language Models (SLMs) transition from personal devices to shared, multi-user environments such as smart homes, a new challenge emerges: the model is expected to distinguish between users to manage information flow appropriately. Without this capability, an SLM could reveal one user's confidential schedule to another, a privacy failure we term interactional privacy. Thus, the ability to generate speaker-aware responses becomes essential for SLM safe deployment. Current SLM benchmarks test dialogue ability but overlook speaker identity. Multi-speaker benchmarks check who said what without assessing whether SLMs adapt their responses. Privacy benchmarks focus on globally sensitive data (e.g., bank passwords) while neglecting contextual privacy-sensitive information (e.g., a user's private appointment). To address this gap, we introduce VoxPrivacy, the first benchmark designed to evaluate interactional privacy in SLMs. VoxPrivacy spans three tiers of increasing difficulty, from following direct secrecy commands to proactively protecting privacy. Our evaluation of nine SLMs on a 32-hour bilingual dataset reveals a widespread vulnerability: most open-source models perform close to random chance (around 50% accuracy) on conditional privacy decisions, while even strong closed-source systems fall short on proactive privacy inference. We further validate these findings on Real-VoxPrivacy, a human-recorded subset, confirming that failures observed on synthetic data persist in real speech. Finally, we demonstrate a viable path forward: by fine-tuning on a new 4,000-hour training set, we improve privacy-preserving abilities while maintaining robustness. To support future work, we release the VoxPrivacy benchmark, the large-scale training set, and the fine-tuned model to foster the development of safer and more context-aware SLMs.


【15】RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation
标题:RIR-Mega-Speech:具有全面声学元数据和可重复评估的回响语音库
链接:https://arxiv.org/abs/2601.19949

作者:Mandip Goswami
摘要:尽管对混响语音进行了数十年的研究,但比较方法仍然很困难,因为大多数语料库缺乏每个文件的声学注释或提供有限的复制文档。我们提出了RIR-Mega-Speech,一个大约117.5小时的语料库,通过卷积LibriSpeech话语与来自RIR-Mega集合的大约5,000个模拟房间脉冲响应创建。每个文件包括RT 60,直接混响比(DRR),和清晰度指数($C_{50}$)计算从源RIR使用明确定义的,可重复的程序。我们还提供脚本来重建数据集并重现所有评估结果。   使用Whisper small对1,500对配对的话语,我们测量了5.20%的WER(95%CI:4.69- 5.78)对干净的语音和7.70%(7.04- 8.35)对混响版本,对应于2.50个百分点(2.06- 2.98)的配对增加。这表示48%的相对降解。WER随RT 60单调增加,随DRR降低,与先前的感知研究一致。虽然混响损害识别的核心发现已经确立,但我们的目标是为社区提供一个标准化的资源,其中声学条件是透明的,结果可以独立验证。该存储库包含适用于Windows和Linux环境的单命令重建说明。
摘要:Despite decades of research on reverberant speech, comparing methods remains difficult because most corpora lack per-file acoustic annotations or provide limited documentation for reproduction. We present RIR-Mega-Speech, a corpus of approximately 117.5 hours created by convolving LibriSpeech utterances with roughly 5,000 simulated room impulse responses from the RIR-Mega collection. Every file includes RT60, direct-to-reverberant ratio (DRR), and clarity index ($C_{50}$) computed from the source RIR using clearly defined, reproducible procedures. We also provide scripts to rebuild the dataset and reproduce all evaluation results.   Using Whisper small on 1,500 paired utterances, we measure 5.20% WER (95% CI: 4.69--5.78) on clean speech and 7.70% (7.04--8.35) on reverberant versions, corresponding to a paired increase of 2.50 percentage points (2.06--2.98). This represents a 48% relative degradation. WER increases monotonically with RT60 and decreases with DRR, consistent with prior perceptual studies. While the core finding that reverberation harms recognition is well established, we aim to provide the community with a standardized resource where acoustic conditions are transparent and results can be verified independently. The repository includes one-command rebuild instructions for both Windows and Linux environments.


【16】MK-SGC-SC: Multiple Kernel guided Sparse Graph Construction in Spectral Clustering for Unsupervised Speaker Diarization
标题:MK-SRC-SC:用于无监督说话者二元化的谱簇中的多核引导稀疏图构建
链接:https://arxiv.org/abs/2601.19946

作者:Nikhil Raghav,Avisek Gupta,Swagatam Das,Md Sahidullah
备注:5 pages
摘要:说话人日志化的目的是将音频记录分割成与各个说话人相对应的区域。虽然无监督的说话人日志化具有内在的挑战性,但在没有预训练或弱监督的情况下识别说话人区域的前景激发了对聚类技术的研究。在这项工作中,我们分享了一个值得注意的观察结果,即测量说话人嵌入的多个内核相似性,然后以原则性的方式为谱聚类制作一个稀疏图,足以在完全无监督的环境中实现最先进的性能。具体来说,我们考虑四个多项式内核和一个度反余弦内核来衡量扬声器嵌入的相似性,使用稀疏图构建的原则性的方式来强调本地的相似性。实验表明,该方法优于在DIHARD-III,AMI,和VoxConverse语料库的各种具有挑战性的环境中的无监督扬声器日记。为了鼓励进一步的研究,我们的实现可以在https://github.com/nikhilraghav29/MK-SGC-SC上获得。
摘要:Speaker diarization aims to segment audio recordings into regions corresponding to individual speakers. Although unsupervised speaker diarization is inherently challenging, the prospect of identifying speaker regions without pretraining or weak supervision motivates research on clustering techniques. In this work, we share the notable observation that measuring multiple kernel similarities of speaker embeddings to thereafter craft a sparse graph for spectral clustering in a principled manner is sufficient to achieve state-of-the-art performances in a fully unsupervised setting. Specifically, we consider four polynomial kernels and a degree one arccosine kernel to measure similarities in speaker embeddings, using which sparse graphs are constructed in a principled manner to emphasize local similarities. Experiments show the proposed approach excels in unsupervised speaker diarization over a variety of challenging environments in the DIHARD-III, AMI, and VoxConverse corpora. To encourage further research, our implementations are available at https://github.com/nikhilraghav29/MK-SGC-SC.


eess.AS音频处理


【1】Decoding Speech Envelopes from Electroencephalogram with a Contrastive Pearson Correlation Coefficient Loss
标题:具有对比Pearson相关系数损失的脑电波解码语音信封
链接:https://arxiv.org/abs/2601.20542

作者:Yayun Liang,Yuanming Zhang,Fei Chen,Jing Lu,Zhibin Lin
摘要:最近的进展,从脑电信号(EEG)重建语音包络,使连续听觉注意解码(AAD)在多说话人环境中。大多数基于深度神经网络(DNN)的包络重建模型都经过训练,以最大化参与包络和重建包络(参与PCC)之间的皮尔逊相关系数(PCC)。虽然有人注意PCC和无人注意PCC之间的差异在听觉注意解码中起着至关重要的作用,但现有方法通常集中于最大化有人注意PCC。因此,我们提出了一个对比PCC的损失,它代表了参加PCC和无人值守PCC之间的差异。所提出的方法进行了评估三个公共EEG AAD数据集使用四个DNN架构。在许多设置中,所提出的目标提高了包络可分性和AAD准确性,同时还揭示了数据集和架构相关的故障情况。
摘要:Recent advances in reconstructing speech envelopes from Electroencephalogram (EEG) signals have enabled continuous auditory attention decoding (AAD) in multi-speaker environments. Most Deep Neural Network (DNN)-based envelope reconstruction models are trained to maximize the Pearson correlation coefficients (PCC) between the attended envelope and the reconstructed envelope (attended PCC). While the difference between the attended PCC and the unattended PCC plays an essential role in auditory attention decoding, existing methods often focus on maximizing the attended PCC. We therefore propose a contrastive PCC loss which represents the difference between the attended PCC and the unattended PCC. The proposed approach is evaluated on three public EEG AAD datasets using four DNN architectures. Across many settings, the proposed objective improves envelope separability and AAD accuracy, while also revealing dataset- and architecture-dependent failure cases.


【2】Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
标题:在听到之前消除你的声音:免训练的说话者放弃Zero-Shot文本到语音的学习
链接:https://arxiv.org/abs/2601.20481

作者:Myungjin Lee,Eunji Shin,Jiyoung Lee
备注:ICASSP'2026
摘要:现代zero-shot文本到语音(TTS)模型提供了前所未有的表达能力,但也带来了严重的犯罪风险,因为它们可以合成从未同意过的个人的声音。在这种情况下,说话人遗忘的目的是防止产生特定的发言人身份的请求。现有的方法,依赖于再培训,是昂贵的,仅限于在培训中看到的扬声器。我们提出了TruS,一个免训练的说话人遗忘框架,将范式从数据删除转移到推理时间控制。TruS引导身份特定的隐藏激活以抑制目标说话者,同时保留其他属性(例如,韵律和情感)。实验结果表明,TruS有效地防止语音生成的可见和不可见的选择退出扬声器,建立一个可扩展的语音合成的保障。演示和代码可在http://mmai.ewha.ac.kr/trus上获得。
摘要:Modern zero-shot text-to-speech (TTS) models offer unprecedented expressivity but also pose serious crime risks, as they can synthesize voices of individuals who never consented. In this context, speaker unlearning aims to prevent the generation of specific speaker identities upon request. Existing approaches, reliant on retraining, are costly and limited to speakers seen in the training set. We present TruS, a training-free speaker unlearning framework that shifts the paradigm from data deletion to inference-time control. TruS steers identity-specific hidden activations to suppress target speakers while preserving other attributes (e.g., prosody and emotion). Experimental results show that TruS effectively prevents voice generation on both seen and unseen opt-out speakers, establishing a scalable safeguard for speech synthesis. The demo and code are available on http://mmai.ewha.ac.kr/trus.


【3】ASR for Affective Speech: Investigating Impact of Emotion and Speech Generative Strategy
标题:情感言语的ASB:调查情感和言语生成策略的影响
链接:https://arxiv.org/abs/2601.20319

作者:Ya-Tse Wu,Chi-Chun Lee
备注:Accepted for publication at IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) 2025
摘要:本研究探讨了情感言语和生成策略如何影响ASR绩效。我们分析了从三种情感TTS模型合成的语音,发现替代错误占主导地位,情感表达能力在不同的模型。基于这些见解,我们引入了两种生成策略:一种使用转录正确性,另一种使用情感显著性,以构建微调子集。结果表明,WER在真实情感数据集上得到了一致的改善,而在干净的LibriSpeech话语上没有明显的退化。组合策略实现了最强的增益,特别是表达性语音。这些发现强调了有针对性的增强对于构建情感感知ASR系统的重要性。
摘要:This work investigates how emotional speech and generative strategies affect ASR performance. We analyze speech synthesized from three emotional TTS models and find that substitution errors dominate, with emotional expressiveness varying across models. Based on these insights, we introduce two generative strategies: one using transcription correctness and another using emotional salience, to construct fine-tuning subsets. Results show consistent WER improvements on real emotional datasets without noticeable degradation on clean LibriSpeech utterances. The combined strategy achieves the strongest gains, particularly for expressive speech. These findings highlight the importance of targeted augmentation for building emotion-aware ASR systems.


【4】T-Mimi: A Transformer-based Mimi Decoder for Real-Time On-Phone TTS
标题:T-Mimi:一个基于转换器的Mimi解码器,用于实时电话TTC
链接:https://arxiv.org/abs/2601.20094

作者:Haibin Wu,Bach Viet Do,Naveen Suda,Julian Chan,Madhavan C R,Gene-Ping Yang,Yi-Chiao Wu,Naoyuki Kanda,Yossef Adi,Xin Lei,Yue Liu,Florian Metze,Yuzong Liu
备注:Accepted by ICASSP 2026
摘要:神经音频编解码器为语音合成提供了有前途的声学特征,其中代表性的流式编解码器(如Mimi)为实时文本到语音(TTS)应用提供了高质量的声学特征。然而,采用混合Transformer和卷积架构的Mimi解码器在边缘设备上引入了显著的延迟瓶颈,这是由于解卷积层的计算密集型性质,其对于移动CPU不友好,例如最具代表性的框架XNNPACK。本文介绍了T-Mimi,一种新的修改Mimi编解码器的解码器,取代其卷积组件与一个纯粹的基于变换的解码器,灵感来自TS 3编解码器架构。这一变化将设备上的TTS延迟从42.1ms大幅降低到仅4.4ms。此外,我们进行量化感知训练,并得出一个重要的发现:最后两个Transformer层和解码器的最后线性层,它们接近波形,对量化高度敏感,必须以全精度保留以保持音频质量。
摘要:Neural audio codecs provide promising acoustic features for speech synthesis, with representative streaming codecs like Mimi providing high-quality acoustic features for real-time Text-to-Speech (TTS) applications. However, Mimi's decoder, which employs a hybrid transformer and convolution architecture, introduces significant latency bottlenecks on edge devices due to the the compute intensive nature of deconvolution layers which are not friendly for mobile-CPUs, such as the most representative framework XNNPACK. This paper introduces T-Mimi, a novel modification of the Mimi codec decoder that replaces its convolutional components with a purely transformer-based decoder, inspired by the TS3-Codec architecture. This change dramatically reduces on-device TTS latency from 42.1ms to just 4.4ms. Furthermore, we conduct quantization aware training and derive a crucial finding: the final two transformer layers and the concluding linear layers of the decoder, which are close to the waveform, are highly sensitive to quantization and must be preserved at full precision to maintain audio quality.


【5】Do we really need Self-Attention for Streaming Automatic Speech Recognition?
标题:我们真的需要自我关注来实现流媒体自动语音识别吗?
链接:https://arxiv.org/abs/2601.19960

作者:Youness Dkhissi,Valentin Vielzeuf,Elys Allesiardo,Anthony Larcher
摘要:基于transformer的架构是许多深度学习领域中最常用的架构,如自然语言处理,计算机视觉或语音处理。它可以鼓励在受约束的任务中直接使用Transformers,而不质疑它是否会产生与标准任务相同的好处。  给定特定的约束条件,必须评估Transformer模型的相关性。这项工作质疑Transformers对特定领域的适用性。我们认为,与这些模型相关的高计算要求和延迟问题与流媒体应用程序不一致。我们的研究促进了寻找替代策略,以提高效率,而不牺牲性能。  根据这一观察,我们的论文批判性地探讨了Transformer架构在这种受限环境中的实用性。作为第一次尝试,我们表明,流自动语音识别(ASR)的计算成本可以减少使用变形卷积,而不是自我注意。此外,我们表明,自我注意机制可以完全删除,而不是取代,没有观察到显着的字错误率下降。
摘要:Transformer-based architectures are the most used architectures in many deep learning fields like Natural Language Processing, Computer Vision or Speech processing. It may encourage the direct use of Transformers in the constrained tasks, without questioning whether it will yield the same benefits as in standard tasks.  Given specific constraints, it is essential to evaluate the relevance of transformer models. This work questions the suitability of transformers for specific domains. We argue that the high computational requirements and latency issues associated with these models do not align well with streaming applications. Our study promotes the search for alternative strategies to improve efficiency without sacrificing performance.  In light of this observation, our paper critically examines the usefulness of transformer architecture in such constrained environments. As a first attempt, we show that the computational cost for Streaming Automatic Speech Recognition (ASR) can be reduced using deformable convolution instead of Self-Attention. Furthermore, we show that Self-Attention mechanisms can be entirely removed and not replaced, without observing significant degradation in the Word Error Rate.


【6】VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
标题:VoxPrivacy:评估语音语言模型交互隐私的基准
链接:https://arxiv.org/abs/2601.19956

作者:Yuxiang Wang,Hongyu Liu,Dekun Chen,Xueyao Zhang,Zhizheng Wu
摘要:随着语音语言模型(SLM)从个人设备过渡到智能家居等共享的多用户环境,一个新的挑战出现了:该模型预计将区分用户以适当地管理信息流。如果没有这种能力,SLM可能会将一个用户的机密日程泄露给另一个用户,这是一种隐私失败,我们称之为互操作隐私。因此,生成说话者感知响应的能力对于SLM安全部署至关重要。目前的SLM基准测试对话能力,但忽视了发言者的身份。多说话者基准检查谁说了什么,而不评估SLM是否适应他们的反应。隐私基准侧重于全球敏感数据(例如,银行密码)而忽略上下文隐私敏感信息(例如,用户的私人约会)。为了解决这一差距,我们引入VoxPrivacy,第一个基准设计来评估在SLM的隐私。VoxPrivacy跨越了三个层次,从遵循直接的保密命令到主动保护隐私。我们在32小时双语数据集上对9个SLM进行的评估揭示了一个普遍存在的漏洞:大多数开源模型在有条件的隐私决策上的表现接近随机机会(约50%的准确率),而即使是强大的闭源系统也无法进行主动隐私推断。我们进一步验证了这些发现的真实VoxPrivacy,一个人类记录的子集,确认在合成数据上观察到的失败在真实语音中持续存在。最后,我们展示了一条可行的前进道路:通过对新的4,000小时训练集进行微调,我们在保持鲁棒性的同时提高了隐私保护能力。为了支持未来的工作,我们发布了VoxPrivacy基准测试、大规模训练集和微调模型,以促进更安全、更上下文感知的SLM的开发。
摘要:As Speech Language Models (SLMs) transition from personal devices to shared, multi-user environments such as smart homes, a new challenge emerges: the model is expected to distinguish between users to manage information flow appropriately. Without this capability, an SLM could reveal one user's confidential schedule to another, a privacy failure we term interactional privacy. Thus, the ability to generate speaker-aware responses becomes essential for SLM safe deployment. Current SLM benchmarks test dialogue ability but overlook speaker identity. Multi-speaker benchmarks check who said what without assessing whether SLMs adapt their responses. Privacy benchmarks focus on globally sensitive data (e.g., bank passwords) while neglecting contextual privacy-sensitive information (e.g., a user's private appointment). To address this gap, we introduce VoxPrivacy, the first benchmark designed to evaluate interactional privacy in SLMs. VoxPrivacy spans three tiers of increasing difficulty, from following direct secrecy commands to proactively protecting privacy. Our evaluation of nine SLMs on a 32-hour bilingual dataset reveals a widespread vulnerability: most open-source models perform close to random chance (around 50% accuracy) on conditional privacy decisions, while even strong closed-source systems fall short on proactive privacy inference. We further validate these findings on Real-VoxPrivacy, a human-recorded subset, confirming that failures observed on synthetic data persist in real speech. Finally, we demonstrate a viable path forward: by fine-tuning on a new 4,000-hour training set, we improve privacy-preserving abilities while maintaining robustness. To support future work, we release the VoxPrivacy benchmark, the large-scale training set, and the fine-tuned model to foster the development of safer and more context-aware SLMs.


【7】RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation
标题:RIR-Mega-Speech:具有全面声学元数据和可重复评估的回响语音库
链接:https://arxiv.org/abs/2601.19949

作者:Mandip Goswami
摘要:尽管对混响语音进行了数十年的研究,但比较方法仍然很困难,因为大多数语料库缺乏每个文件的声学注释或提供有限的复制文档。我们提出了RIR-Mega-Speech,一个大约117.5小时的语料库,通过卷积LibriSpeech话语与来自RIR-Mega集合的大约5,000个模拟房间脉冲响应创建。每个文件包括RT 60,直接混响比(DRR),和清晰度指数($C_{50}$)计算从源RIR使用明确定义的,可重复的程序。我们还提供脚本来重建数据集并重现所有评估结果。   使用Whisper small对1,500对配对的话语,我们测量了5.20%的WER(95%CI:4.69- 5.78)对干净的语音和7.70%(7.04- 8.35)对混响版本,对应于2.50个百分点(2.06- 2.98)的配对增加。这表示48%的相对降解。WER随RT 60单调增加,随DRR降低,与先前的感知研究一致。虽然混响损害识别的核心发现已经确立,但我们的目标是为社区提供一个标准化的资源,其中声学条件是透明的,结果可以独立验证。该存储库包含适用于Windows和Linux环境的单命令重建说明。
摘要:Despite decades of research on reverberant speech, comparing methods remains difficult because most corpora lack per-file acoustic annotations or provide limited documentation for reproduction. We present RIR-Mega-Speech, a corpus of approximately 117.5 hours created by convolving LibriSpeech utterances with roughly 5,000 simulated room impulse responses from the RIR-Mega collection. Every file includes RT60, direct-to-reverberant ratio (DRR), and clarity index ($C_{50}$) computed from the source RIR using clearly defined, reproducible procedures. We also provide scripts to rebuild the dataset and reproduce all evaluation results.   Using Whisper small on 1,500 paired utterances, we measure 5.20% WER (95% CI: 4.69--5.78) on clean speech and 7.70% (7.04--8.35) on reverberant versions, corresponding to a paired increase of 2.50 percentage points (2.06--2.98). This represents a 48% relative degradation. WER increases monotonically with RT60 and decreases with DRR, consistent with prior perceptual studies. While the core finding that reverberation harms recognition is well established, we aim to provide the community with a standardized resource where acoustic conditions are transparent and results can be verified independently. The repository includes one-command rebuild instructions for both Windows and Linux environments.


【8】MK-SGC-SC: Multiple Kernel guided Sparse Graph Construction in Spectral Clustering for Unsupervised Speaker Diarization
标题:MK-SRC-SC:用于无监督说话者二元化的谱簇中的多核引导稀疏图构建
链接:https://arxiv.org/abs/2601.19946

作者:Nikhil Raghav,Avisek Gupta,Swagatam Das,Md Sahidullah
备注:5 pages
摘要:说话人日志化的目的是将音频记录分割成与各个说话人相对应的区域。虽然无监督的说话人日志化具有内在的挑战性,但在没有预训练或弱监督的情况下识别说话人区域的前景激发了对聚类技术的研究。在这项工作中,我们分享了一个值得注意的观察结果,即测量说话人嵌入的多个内核相似性,然后以原则性的方式为谱聚类制作一个稀疏图,足以在完全无监督的环境中实现最先进的性能。具体来说,我们考虑四个多项式内核和一个度反余弦内核来衡量扬声器嵌入的相似性,使用稀疏图构建的原则性的方式来强调本地的相似性。实验表明,该方法优于在DIHARD-III,AMI,和VoxConverse语料库的各种具有挑战性的环境中的无监督扬声器日记。为了鼓励进一步的研究,我们的实现可以在https://github.com/nikhilraghav29/MK-SGC-SC上获得。
摘要:Speaker diarization aims to segment audio recordings into regions corresponding to individual speakers. Although unsupervised speaker diarization is inherently challenging, the prospect of identifying speaker regions without pretraining or weak supervision motivates research on clustering techniques. In this work, we share the notable observation that measuring multiple kernel similarities of speaker embeddings to thereafter craft a sparse graph for spectral clustering in a principled manner is sufficient to achieve state-of-the-art performances in a fully unsupervised setting. Specifically, we consider four polynomial kernels and a degree one arccosine kernel to measure similarities in speaker embeddings, using which sparse graphs are constructed in a principled manner to emphasize local similarities. Experiments show the proposed approach excels in unsupervised speaker diarization over a variety of challenging environments in the DIHARD-III, AMI, and VoxConverse corpora. To encourage further research, our implementations are available at https://github.com/nikhilraghav29/MK-SGC-SC.


【9】Audio Deepfake Detection in the Age of Advanced Text-to-Speech models
标题:高级文本到语音模型时代的音频深度伪造检测
链接:https://arxiv.org/abs/2601.20510

作者:Robin Singh,Aditya Yogesh Nair,Fabio Palumbo,Florian Barbaro,Anna Dyka,Lohith Rachakonda
备注:This work was performed using HPC resources from GENCI-IDRIS (Grant 2025- AD011016076)
摘要:文本到语音(TTS)系统的最新进展大大提高了合成语音的真实性,为音频深度伪造检测提出了新的挑战。这项工作提出了一个比较评估的三个国家的最先进的TTS模型-Dia 2,Maya 1和MeloTTS-代表流,基于LLM,和非自回归架构。使用Daily-Dialog数据集生成了12,000个合成音频样本的语料库,并针对四个检测框架进行了评估,包括语义,结构和信号级方法。结果显示,在整个生成机制的检测器性能的显着变化:对一个TTS架构有效的模型可能会失败对其他人,特别是基于LLM的合成。相比之下,结合互补分析水平的多视图检测方法在所有评估的模型中表现出稳健的性能。这些发现强调了单一范式检测器的局限性,并强调了集成检测策略的必要性,以应对不断变化的音频deepfake威胁。
摘要:Recent advances in Text-to-Speech (TTS) systems have substantially increased the realism of synthetic speech, raising new challenges for audio deepfake detection. This work presents a comparative evaluation of three state-of-the-art TTS models--Dia2, Maya1, and MeloTTS--representing streaming, LLM-based, and non-autoregressive architectures. A corpus of 12,000 synthetic audio samples was generated using the Daily-Dialog dataset and evaluated against four detection frameworks, including semantic, structural, and signal-level approaches. The results reveal significant variability in detector performance across generative mechanisms: models effective against one TTS architecture may fail against others, particularly LLM-based synthesis. In contrast, a multi-view detection approach combining complementary analysis levels demonstrates robust performance across all evaluated models. These findings highlight the limitations of single-paradigm detectors and emphasize the necessity of integrated detection strategies to address the evolving landscape of audio deepfake threats.


【10】MiLorE-SSL: Scaling Multilingual Capabilities in Self-Supervised Models without Forgetting
标题:MiLorE-SSL:在自我监督模型中扩展多语言功能而不会忘记
链接:https://arxiv.org/abs/2601.20300

作者:Jing Xu,Minglin Wu,Xueyuan Chen,Xixin Wu,Helen Meng
备注:Accepted by ICASSP2026
摘要:自监督学习(SSL)极大地推进了语音表示学习,但多语言SSL模型仍然局限于预训练期间遇到的语言。从头开始重新训练以融入新语言在计算上是昂贵的,而没有迁移策略的顺序训练通常会导致灾难性的遗忘。为了解决这个问题,我们提出了MiLorE-SSL,一个轻量级的框架,结合了LoRA模块与软混合专家(MoE)机制,有效的持续多语言培训。LoRA提供高效的低秩自适应,而软MoE促进跨语言的灵活专家共享,减少跨语言干扰。为了进一步减轻遗忘,我们从现有语言中引入有限的重放数据,避免依赖大型历史语料库。在ML-SUPERB上的实验表明,MiLorE-SSL在新语言中获得了很好的性能,并且仅用2.14%的可训练参数就提高了现有语言的能力。
摘要:Self-supervised learning (SSL) has greatly advanced speech representation learning, but multilingual SSL models remain constrained to languages encountered during pretraining. Retraining from scratch to incorporate new languages is computationally expensive, while sequential training without migitation strategies often leads to catastrophic forgetting. To address this, we propose MiLorE-SSL, a lightweight framework that combines LoRA modules with a soft mixture-of-experts (MoE) mechanism for efficient continual multilingual training. LoRA provides efficient low-rank adaptation, while soft MoE promotes flexible expert sharing across languages, reducing cross-lingual interference. To further mitigate forgetting, we introduce limited replay data from existing languages, avoiding reliance on large historical corpora. Experiments on ML-SUPERB demonstrate that MiLorE-SSL achieves strong performance in new languages and improves the ability in existing ones with only 2.14% trainable parameters.


【11】Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR
标题:注意转变:使用Delta SSL嵌入式增强儿童ASB
链接:https://arxiv.org/abs/2601.20142

作者:Zilai Wang,Natarajan Balaji Shankar,Kaiyuan Zhang,Zihan Wang,Abeer Alwan
备注:ICASSP 2026
摘要:自监督学习(SSL)模型在许多语音任务中取得了令人印象深刻的结果,但由于数据有限和预训练域不匹配,儿童自动语音识别(ASR)仍然具有挑战性。微调SSL模型的儿童语音诱导的代表性空间的变化。我们假设delta SSL嵌入(定义为来自微调模型的嵌入与来自其预训练对应模型的嵌入之间的差异)编码特定于任务的信息,以补充来自另一个SSL模型的微调特征。我们使用不同的模型在MyST儿童语料库上评估多种融合策略。结果表明,与微调嵌入融合相比,WavLM的delta嵌入融合对HuBERT产生了高达10%的相对WER降低,对W2V2产生了4.4%的降低。值得注意的是,融合WavLM与delta W2V2嵌入实现了9.64的WER,在MyST语料库上的SSL模型中树立了新的艺术水平。这些发现证明了增量嵌入和突出特征融合的有效性,作为推进儿童ASR的一个有前途的方向。
摘要:Self-supervised learning (SSL) models have achieved impressive results across many speech tasks, yet child automatic speech recognition (ASR) remains challenging due to limited data and pretraining domain mismatch. Fine-tuning SSL models on child speech induces shifts in the representation space. We hypothesize that delta SSL embeddings, defined as the differences between embeddings from a finetuned model and those from its pretrained counterpart, encode task-specific information that complements finetuned features from another SSL model. We evaluate multiple fusion strategies on the MyST childrens corpus using different models. Results show that delta embedding fusion with WavLM yields up to a 10 percent relative WER reduction for HuBERT and a 4.4 percent reduction for W2V2, compared to finetuned embedding fusion. Notably, fusing WavLM with delta W2V2 embeddings achieves a WER of 9.64, setting a new state of the art among SSL models on the MyST corpus. These findings demonstrate the effectiveness of delta embeddings and highlight feature fusion as a promising direction for advancing child ASR.


【12】LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning
标题:LTS-VoiceAgent:一个通过语义触发和增量推理进行高效流媒体语音交互的听-想-说框架
链接:https://arxiv.org/abs/2601.19952

作者:Wenhao Zou,Yuwei Miao,Zhanyu Ma,Jun Xu,Jiuchong Gao,Jinghua Hao,Renqing He,Jingwen Xu
摘要:实时语音代理面临着一个困境:端到端模型通常缺乏深度推理,而级联管道通过严格按顺序执行ASR,LLM推理和TTS而导致高延迟,这与人类对话不同,在人类对话中,听众通常在说话者完成之前开始思考。由于级联架构仍然是复杂任务的主要选择,因此现有的级联流传输策略试图通过机械分段(例如,固定块、基于VAD的分割)或推测性生成,但它们经常要么破坏语义单元,要么在必须回滚的预测上浪费计算。为了解决这些挑战,我们提出了LTS-语音代理,一个听,想,说框架,明确地分离时,从如何逐步推理思考。它具有一个动态语义触发器来检测有意义的前缀,以及一个双角色流推理器,它协调后台思想者(用于状态维护)和前台发言者(用于推测性解决)。这种并行设计使“边思考边说话”不会阻碍响应。我们还介绍了一个暂停和修复基准包含自然的不流畅压力测试流鲁棒性。VERA,Spoken-MQA,BigBenchAudio和我们的基准测试的实验表明,LTS-VoiceAgent实现了比串行级联基线和现有流媒体策略更强的准确性-延迟-效率权衡。
摘要:Real-time voice agents face a dilemma: end-to-end models often lack deep reasoning, while cascaded pipelines incur high latency by executing ASR, LLM reasoning, and TTS strictly in sequence, unlike human conversation where listeners often start thinking before the speaker finishes. Since cascaded architectures remain the dominant choice for complex tasks, existing cascaded streaming strategies attempt to reduce this latency via mechanical segmentation (e.g., fixed chunks, VAD-based splitting) or speculative generation, but they frequently either break semantic units or waste computation on predictions that must be rolled back. To address these challenges, we propose LTS-VoiceAgent, a Listen-Think-Speak framework that explicitly separates when to think from how to reason incrementally. It features a Dynamic Semantic Trigger to detect meaningful prefixes, and a Dual-Role Stream Orchestrator that coordinates a background Thinker (for state maintenance) and a foreground Speaker (for speculative solving). This parallel design enables "thinking while speaking" without blocking responses. We also introduce a Pause-and-Repair benchmark containing natural disfluencies to stress-test streaming robustness. Experiments across VERA, Spoken-MQA, BigBenchAudio, and our benchmark show that LTS-VoiceAgent achieves a stronger accuracy-latency-efficiency trade-off than serial cascaded baselines and existing streaming strategies.


【13】Pianoroll-Event: A Novel Score Representation for Symbolic Music
标题:钢琴演奏事件:一种新的象征性音乐乐谱表现形式
链接:https://arxiv.org/abs/2601.19951

作者:Lekai Qian,Haoyu Gu,Dehan Li,Boyu Cao,Qi Liu
摘要:音乐符号表示是计算音乐学的一个基本挑战。虽然基于网格的表示有效地保持了基音时间的空间对应性,但其固有的数据稀疏性导致编码效率低。离散事件表示实现了紧凑的编码,但未能充分捕捉结构不变性和空间局部性。为了解决这些互补的限制,我们提出了Pianoroll事件,一种新的编码方案,通过事件描述pianoroll表示,结合结构特性与编码效率,同时保持时间依赖性和局部空间模式。具体来说,我们设计了四个互补的事件类型:帧事件的时间边界,间隙事件的稀疏区域,模式事件的注意模式,音乐元数据的音乐结构事件。Pianoroll-Event在序列长度和词汇量之间取得了有效的平衡,编码效率比典型的离散序列方法提高了1.36 ~ 7.16倍。跨多个自回归架构的实验表明,使用我们的表示的模型在定量和人工评估中始终优于基线。
摘要:Symbolic music representation is a fundamental challenge in computational musicology. While grid-based representations effectively preserve pitch-time spatial correspondence, their inherent data sparsity leads to low encoding efficiency. Discrete-event representations achieve compact encoding but fail to adequately capture structural invariance and spatial locality. To address these complementary limitations, we propose Pianoroll-Event, a novel encoding scheme that describes pianoroll representations through events, combining structural properties with encoding efficiency while maintaining temporal dependencies and local spatial patterns. Specifically, we design four complementary event types: Frame Events for temporal boundaries, Gap Events for sparse regions, Pattern Events for note patterns, and Musical Structure Events for musical metadata. Pianoroll-Event strikes an effective balance between sequence length and vocabulary size, improving encoding efficiency by 1.36\times to 7.16\times over representative discrete sequence methods. Experiments across multiple autoregressive architectures show models using our representation consistently outperform baselines in both quantitative and human evaluations.


机器翻译由腾讯交互翻译提供,仅供参考