今日论文合集:cs.SD语音8篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Beyond Rules: Towards Basso Continuo Personal Style Identification
标题:超越规则:走向巴索继续个人风格认同
链接:https://arxiv.org/abs/2604.21822
作者:Adam Štefunko,Jan Hajič
备注:8 pages, 4 figures, accepted to the 13th International Conference on Digital Libraries for Musicology (DLfM)
摘要:当代历史知情实践运动的一个核心部分是连续低音,这是一种即兴伴奏流派,其传统起源于巴洛克时代,现在许多键盘手都在积极练习。虽然计算音乐学已经研究了连续低音的理论基础,包括和声和声音引导的规则和约束,但由于缺乏合适的演奏数据可以进行实证分析,连续低音作为一种活跃的表演艺术的特点在很大程度上被忽视了。随着Aligned Continuo Realization Dataset(ACoRD)和Basso Continuo Realization-to-Score对齐的引入,这种情况已经改变。连续低音演奏是由来自历史论文的风格传统塑造的,但它也可能为展示其从业者的个人演奏风格提供空间。在本文中,我们试图探讨的问题存在的个人风格的连续低音实现的球员在ACoRD数据集。我们使用一个历史上知情的结构化表示的连续低音表现球场的内容称为griffs和支持向量机,看看是否有可能根据他们的表现对球员进行分类。结果表明,我们可以识别球员从他们的表现。除了球员的分类问题,我们讨论的元素,使球员的个人风格。
摘要:A central part of the contemporary Historically Informed Practice movement is basso continuo, an improvised accompaniment genre with its traditions originating in the baroque era and actively practiced by many keyboard players nowadays. Although computational musicology has studied the theoretical foundations of basso continuo expressed by harmonic and voice-leading rules and constraints, characteristics of basso continuo as an active performing art have been largely overlooked mostly due to a lack of suitable performance data that could be empirically analyzed. This has changed with the introduction of The Aligned Continuo Realization Dataset (ACoRD) and the basso continuo realization-to-score alignment. Basso continuo playing is shaped by stylistic traditions coming from historical treatises, but it also may provide space for showcasing individual performance styles of its practitioners. In this paper, we attempt to explore the question of the presence of personal styles in the basso continuo realizations of players in the ACoRD dataset. We use a historically informed structured representation of basso continuo performance pitch content called griffs and Support Vector Machines to see whether it is possible to classify players based on their performances. The results show that we can identify players from their performances. In addition to the player classification problem, we discuss the elements that make up the individual styles of the players.


【2】Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0

标题:时间与层:在wav2vec 2.0中定位结构障碍语音描述符的预测线索
链接:https://arxiv.org/abs/2604.21628
作者:Natalie Engert,Dominik Wagner,Korbinian Riedhammer,Tobias Bocklet
备注:Accepted to IEEE ICASSP 2026
摘要:Wav2vec 2.0(W2V2)通过有效地捕捉非典型语音的特征,在病理语音分析中表现出强大的性能。尽管它取得了成功,但仍不清楚其学习表征的哪些组成部分对特定的下游任务最有用。在这项研究中,我们解决了这个问题,通过调查使用注释从语音无障碍项目数据集的构音障碍语音描述符的回归。我们专注于五个描述符,每一个都涉及语音或声音产生的不同方面:可理解性,不精确的辅音,不适当的沉默,刺耳的声音和单响。语音表示来自一个基于W2V2的特征提取器,我们系统地比较层和时间的聚合策略,使用细心的统计池。我们的研究结果表明,清晰度是最好的捕获通过逐层表示,而不精确的辅音,刺耳的声音和单响度受益于时间明智的建模。对于不适当的沉默,两种方法都没有明显的优势。
摘要:Wav2vec 2.0 (W2V2) has shown strong performance in pathological speech analysis by effectively capturing the characteristics of atypical speech. Despite its success, it remains unclear which components of its learned representations are most informative for specific downstream tasks. In this study, we address this question by investigating the regression of dysarthric speech descriptors using annotations from the Speech Accessibility Project dataset. We focus on five descriptors, each addressing a different aspect of speech or voice production: intelligibility, imprecise consonants, inappropriate silences, harsh voice and monoloudness. Speech representations are derived from a W2V2-based feature extractor, and we systematically compare layer-wise and time-wise aggregation strategies using attentive statistics pooling. Our results show that intelligibility is best captured through layer-wise representations, whereas imprecise consonants, harsh voice and monoloudness benefit from time-wise modeling. For inappropriate silences, no clear advantage could be observed for either approach.


【3】Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition

标题:LLM解码器是否公平倾听?对语言模型先验如何影响语音识别中的偏差进行基准测试
链接:https://arxiv.org/abs/2604.21276
作者:Srishti Ginjala,Eric Fosler-Lussier,Christopher W. Myers,Srinivasan Parthasarathy
摘要:随着预训练的大型语言模型在语音识别中取代特定任务的解码器,一个关键问题出现了:它们的文本衍生先验是否使识别在人口统计群体中更公平或更有偏见?我们使用Common Voice 24和Meta的Fair-Speech(一种消除词汇混淆的受控提示数据集)评估了跨越三代架构的九个模型(没有语言模型的CTC,具有隐式LM的编码器-解码器,以及基于LLM的显式预训练解码器),涉及五个人口统计轴(种族,口音,性别,年龄,第一语言)的约43,000个话语。在干净的音频上,有三个发现挑战了假设:LLM解码器不会放大种族偏见(Granite-8B具有最好的种族公平性,最大/最小WER = 2.28); Whisper在印度口音语音上表现出病态幻觉,在大v3下非单调插入率飙升至9.62%;音频压缩预测口音公平性超过LLM规模。然后,我们在两个数据集的12个声学退化条件(噪声,混响,沉默注入,块掩蔽)下对这些发现进行了压力测试,总共216次推理运行。严重的退化矛盾地压缩了公平差距,因为所有群体都收敛到高WER,但沉默注入通过触发人口统计选择性幻觉将Whisper的口音偏见放大了4.64倍。在掩蔽下,Whisper进入灾难性的重复循环(51,797次插入中的86%),而显式LLM解码器在接近零重复的情况下产生的插入减少了38倍;高压缩音频编码(Q-former)甚至在LLM解码器中重新引入重复病理。这些结果表明,音频编码器的设计,而不是LLM缩放,是公平和强大的语音识别的主要杠杆。
摘要:As pretrained large language models replace task-specific decoders in speech recognition, a critical question arises: do their text-derived priors make recognition fairer or more biased across demographic groups? We evaluate nine models spanning three architectural generations (CTC with no language model, encoder-decoder with an implicit LM, and LLM-based with an explicit pretrained decoder) on about 43,000 utterances across five demographic axes (ethnicity, accent, gender, age, first language) using Common Voice 24 and Meta's Fair-Speech, a controlled-prompt dataset that eliminates vocabulary confounds. On clean audio, three findings challenge assumptions: LLM decoders do not amplify racial bias (Granite-8B has the best ethnicity fairness, max/min WER = 2.28); Whisper exhibits pathological hallucination on Indian-accented speech with a non-monotonic insertion-rate spike to 9.62% at large-v3; and audio compression predicts accent fairness more than LLM scale. We then stress-test these findings under 12 acoustic degradation conditions (noise, reverberation, silence injection, chunk masking) across both datasets, totaling 216 inference runs. Severe degradation paradoxically compresses fairness gaps as all groups converge to high WER, but silence injection amplifies Whisper's accent bias up to 4.64x by triggering demographic-selective hallucination. Under masking, Whisper enters catastrophic repetition loops (86% of 51,797 insertions) while explicit-LLM decoders produce 38x fewer insertions with near-zero repetition; high-compression audio encoding (Q-former) reintroduces repetition pathology even in LLM decoders. These results suggest that audio encoder design, not LLM scaling, is the primary lever for equitable and robust speech recognition.


【4】MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

标题:MAGIC-TTS:具有显式局部持续时间和重复控制的细粒度可控语音合成
链接:https://arxiv.org/abs/2604.21164
作者:Jialong Mai,Xiaofen Xing,Xiangmin Xu
摘要:细粒度的本地定时控制仍然缺乏现代文本到语音系统:现有的方法通常只提供话语级持续时间或全局说话速率控制,而精确的令牌级定时操作仍然不可用。据我们所知,MAGIC-TTS是第一个对令牌级内容持续时间和暂停进行显式本地定时控制的TTS模型。MAGIC-TTS通过明确的令牌级持续时间条件,精心准备的高置信度持续时间监督以及纠正零值偏差并使模型对缺失的局部控制具有鲁棒性的训练机制来实现。在我们的时间控制基准测试中,MAGIC-TTS大大改善了自发合成之后的令牌级持续时间和暂停。即使不提供时序控制,MAGIC-TTS也能保持自然的高品质合成。我们进一步评估了实用的本地编辑与一个基于JavaScript的基准涵盖导航指导,指导阅读,并面向可访问性的代码阅读。在这种设置下,MAGIC-TTS实现了可再现的均匀定时基线,然后以低平均偏差将编辑区域移向所需的局部目标。这些结果表明,显式细粒度可控性可以有效地实现在一个高质量的TTS系统,可以支持现实的本地定时编辑应用程序。
摘要:Fine-grained local timing control is still absent from modern text-to-speech systems: existing approaches typically provide only utterance-level duration or global speaking-rate control, while precise token-level timing manipulation remains unavailable. To the best of our knowledge, MAGIC-TTS is the first TTS model with explicit local timing control over token-level content duration and pause. MAGIC-TTS is enabled by explicit token-level duration conditioning, carefully prepared high-confidence duration supervision, and training mechanisms that correct zero-value bias and make the model robust to missing local controls. On our timing-control benchmark, MAGIC-TTS substantially improves token-level duration and pause following over spontaneous synthesis. Even when no timing control is provided, MAGIC-TTS maintains natural high-quality synthesis. We further evaluate practical local editing with a scenario-based benchmark covering navigation guidance, guided reading, and accessibility-oriented code reading. In this setting, MAGIC-TTS realizes a reproducible uniform-timing baseline and then moves the edited regions toward the requested local targets with low mean bias. These results show that explicit fine-grained controllability can be implemented effectively in a high-quality TTS system and can support realistic local timing-editing applications.


【5】Materialistic RIR: Material Conditioned Realistic RIR Generation

标题:唯物主义RIR:物质条件现实主义RIR生成
链接:https://arxiv.org/abs/2604.21119
作者:Mahnoor Fatima Saad,Sagnik Majumder,Kristen Grauman,Ziad Al-Halah
备注:Accepted to CVPR 2026 Findings. Project page: https://mahnoor-fatima-saad.github.io/MatRIR.html
摘要:戒指像金子,重击像木头!我们在场景中听到的声音不仅受环境的空间布局的影响,还受其中物体和表面的材料的影响。例如,一个有木墙的房间会产生与空间布局相同但有混凝土墙的房间不同的声学体验。精确地建模这些效果对于虚拟现实、机器人、建筑设计和音频工程等应用至关重要。然而,现有的声学建模方法往往在相关表示中纠缠空间和材料影响,这限制了用户控制并降低了所生成的声学的真实性。在这项工作中,我们提出了一种新的方法,材料控制的房间脉冲响应(RIR)的生成,明确解开空间和材料线索在场景中的影响。我们的方法模型的RIR使用两个模块:一个空间模块,捕捉场景的空间布局的影响,和一个材料模块,根据用户指定的材料配置调制这个空间RIR。这种明确的设计使用户能够轻松地修改场景的材料配置,并观察其对声学的影响,而不会改变空间结构或场景内容。我们的模型在基于声学的指标(RTE上高达+16%)和基于材料的指标(高达+70%)上都比以前的方法有了显著的改进。此外,通过人类感知研究,我们证明了与最强基线相比,我们模型的真实性和材料敏感性有所改善。
摘要:Rings like gold, thuds like wood! The sound we hear in a scene is shaped not only by the spatial layout of the environment but also by the materials of the objects and surfaces within it. For instance, a room with wooden walls will produce a different acoustic experience from a room with the same spatial layout but concrete walls. Accurately modeling these effects is essential for applications such as virtual reality, robotics, architectural design, and audio engineering. Yet, existing methods for acoustic modeling often entangle spatial and material influences in correlated representations, which limits user control and reduces the realism of the generated acoustics. In this work, we present a novel approach for material-controlled Room Impulse Response (RIR) generation that explicitly disentangles the effects of spatial and material cues in a scene. Our approach models the RIR using two modules: a spatial module that captures the influence of the spatial layout of the scene, and a material module that modulates this spatial RIR according to a user-specified material configuration. This explicitly disentangled design allows users to easily modify the material configuration of a scene and observe its impact on acoustics without altering the spatial structure or scene content. Our model provides significant improvements over prior approaches on both acoustic-based metrics (up to +16% on RTE) and material-based metrics (up to +70%). Furthermore, through a human perceptual study, we demonstrate the improved realism and material sensitivity of our model compared to the strongest baselines.


【6】Sema: Semantic Transport for Real-Time Multimodal Agents

标题:Sema:实时多模式代理的语义传输
链接:https://arxiv.org/abs/2604.20940
作者:Jiaying Meng,Bojie Li
摘要:实时多模式代理使用为人类接收器设计的网络堆栈传输原始音频和屏幕截图,这些堆栈优化了感知保真度和流畅播放。然而,代理模型充当事件驱动的处理器,没有固有的物理时间感,消耗任务相关的语义,而不是实时重建信号。这种根本差异将传输目标从信号保真度的技术问题(香农-韦弗A级)转移到含义保留的语义问题(B级)。这种不匹配带来了显著的开销。在可视化管道中,屏幕截图上传占受限上行链路上端到端动作延迟的60%以上,而在语音管道中,传统传输具有大量冗余,发送的数据比保持任务准确性所需的多43- 64倍。我们提出了Sema,一个语义传输系统,它结合了离散音频tokenizer与混合屏幕表示(无损可访问性树或OCR文本,加上紧凑的视觉令牌)和突发令牌交付,消除抖动缓冲区。在模拟WAN条件下的模拟中,Sema将音频的上行链路带宽减少了64倍,将屏幕截图的上行链路带宽减少了130- 210倍,同时将任务精度保持在原始基线的0.7个百分点以内。
摘要:Real-time multimodal agents transport raw audio and screenshots using networking stacks designed for human receivers, which optimize for perceptual fidelity and smooth playout. Yet agent models act as event-driven processors with no inherent sense of physical time, consuming task-relevant semantics rather than reconstructing signals in real time. This fundamental difference shifts the transport goal from the technical problem of signal fidelity (Shannon-Weaver Level A) to the semantic problem of meaning preservation (Level B). This mismatch imposes significant overhead. In visual pipelines, screenshot upload accounts for over 60% of end-to-end action latency on constrained uplinks, and in voice pipelines, conventional transport carries massive redundancy, sending 43-64x more data than needed to maintain task accuracy. We present Sema, a semantic transport system that combines discrete audio tokenizers with a hybrid screen representation (lossless accessibility-tree or OCR text, plus compact visual tokens) and bursty token delivery that eliminates jitter buffers. In simulations under emulated WAN conditions, Sema reduces uplink bandwidth by 64x for audio and 130-210x for screenshots while preserving task accuracy within 0.7 percentage points of the raw baseline.


【7】DiariZen Explained: A Tutorial for the Open Source State-of-the-Art Speaker Diarization Pipeline

标题:DiariZen解释:开源最先进的扬声器Dializer管道的开发
链接:https://arxiv.org/abs/2604.21507
作者:Nikhil Raghav
备注:13 pages, 7 figures, 2 tables. Code available at https://github.com/nikhilraghav29/diarizen-tutorial
摘要:说话人日志(SD)是在多说话人音频流中回答“谁在什么时候说话”的任务。典型地,SD系统聚类属于单个说话者的身份的语音片段。近年来,通过端到端神经日志(EEND)方法,SD取得了实质性进展。DiariZen是一种混合SD管道,构建在结构修剪的WavLM-Large编码器,具有powerset分类的Conformer后端和VBx集群上,代表了在编写多个基准测试时领先的开源技术水平。尽管DiariZen架构具有强大的性能,但它跨越了多个存储库和框架,使得研究人员和从业人员难以理解、复制或扩展整个系统。本教程提供了完整的DiariZen管道的独立,逐块解释,将其分解为七个阶段:(1)音频加载和滑动窗口分割,(2)具有学习层加权的WavLM特征提取,(3)Conformer后端和幂集分类,(4)经由叠加的分割聚合,(5)具有重叠排除的说话人嵌入提取,(6)VBx聚类与PLDA评分,以及(7)重建和RTTM输出。对于每个块,我们提供了概念动机,源代码参考,中间张量形状,并从AMI会议语料库中摘录了30年代的实际输出的注释可视化。该实现可在https://github.com/nikhilraghav29/diarizen-tutorial上获得,其中包括每个块的独立可执行脚本和一个端到端运行完整管道的Applyter Notebook。
摘要:Speaker diarization (SD) is the task of answering "who spoke when" in a multi-speaker audio stream. Classically, an SD system clusters segments of speech belonging to an individual speaker's identity. Recent years have seen substantial progress in SD through end-to-end neural diarization (EEND) approaches. DiariZen, a hybrid SD pipeline built upon a structurally pruned WavLM-Large encoder, a Conformer backend with powerset classification, and VBx clustering, represents the leading open-source state of the art at the time of writing across multiple benchmarks. Despite its strong performance, the DiariZen architecture spans several repositories and frameworks, making it difficult for researchers and practitioners to understand, reproduce, or extend the system as a whole. This tutorial paper provides a self-contained, block-by-block explanation of the complete DiariZen pipeline, decomposing it into seven stages: (1) audio loading and sliding window segmentation, (2) WavLM feature extraction with learned layer weighting, (3) Conformer backend and powerset classification, (4) segmentation aggregation via overlap-add, (5) speaker embedding extraction with overlap exclusion, (6) VBx clustering with PLDA scoring, and (7) reconstruction and RTTM output. For each block, we provide the conceptual motivation, source code references, intermediate tensor shapes, and annotated visualizations of the actual outputs on a 30s excerpt from the AMI Meeting Corpus. The implementation is available at https://github.com/nikhilraghav29/diarizen-tutorial, which includes standalone executable scripts for each block and a Jupyter notebook that runs the complete pipeline end-to-end.


【8】HHL with a Coherent Fourier Oracle: A Proof-of-Concept Quantum Architecture for Joint Melody-Harmony Generation

标题:具有连贯傅里叶Oracle的HHL:用于联合旋律和声生成的概念验证量子架构
链接:https://arxiv.org/abs/2604.20882
作者:Alexis Kirke
摘要:量子算法与经典计算的理论加速比是罕见的。其中最突出的是用于求解稀疏线性系统的Harrow-Hassidim-Lloyd(HHL)算法。这里,HHL被应用于编码旋律偏好:系统矩阵编码Narmour隐含实现和Krumhansl-Kessler音调稳定性,因此其解向量是音乐认知加权的音符对分布。HHL的关键约束是经典地读取其输出取消了量子加速;解决方案必须被相干地消耗。这激发了一个连贯的傅立叶谐波预言:一个将和弦过渡权重直接应用于HHL振幅向量的酉,使得单个测量联合选择旋律音符和两个和弦进行。   两个音符/两个弦(2/2)块用于包含关节状态空间的指数增长,否则将使较大块的经典模拟不可行。对于较长段落的演示,块被经典地链接-每个块的折叠输出条件下一个-作为临时解决方案,直到容错硬件允许更大的单片电路。四块链产生8个音符超过8个和弦,在每个块边界处具有语法上有效的过渡。   独立的基于规则的和声验证证实,97%的生成的和弦进行被评为强或可接受的。主要的动机是,HHL进行了经典的线性求解器的指数加速;这项工作表明,一个连贯的HHL+Oracle管道-在音乐环境中实现加速的先决条件-是机械实现的。代表性输出的音频实现可供在线收听。
摘要:Quantum algorithms with a proven theoretical speedup over classical computation are rare. Among the most prominent is the Harrow-Hassidim-Lloyd (HHL) algorithm for solving sparse linear systems. Here, HHL is applied to encode melodic preference: the system matrix encodes Narmour implication-realisation and Krumhansl-Kessler tonal stability, so its solution vector is a music-cognition-weighted note-pair distribution. The key constraint of HHL is that reading its output classically cancels the quantum speedup; the solution must be consumed coherently. This motivates a coherent Fourier harmonic oracle: a unitary that applies chord-transition weights directly to the HHL amplitude vector, so that a single measurement jointly selects both melody notes and a two-chord progression.   A two-note/two-chord (2/2) block is used to contain the exponential growth of the joint state space that would otherwise make classical simulation of larger blocks infeasible. For demonstrations of longer passages, blocks are chained classically - each block's collapsed output conditions the next -- as a temporary workaround until fault-tolerant hardware permits larger monolithic circuits. A four-block chain produces 8 notes over 8 chords with grammatically valid transitions at every block boundary.   Independent rule-based harmony validation confirms that 97% of generated chord progressions are rated strong or acceptable. The primary motivation is that HHL carries a proven exponential speedup over classical linear solvers; this work demonstrates that a coherent HHL+oracle pipeline - the prerequisite for that speedup to be realised in a musical setting - is mechanically achievable. Audio realisations of representative outputs are made available for listening online.


eess.AS音频处理


【1】PHOTON: Non-Invasive Optical Tracking of Key-Lever Motion in Historical Keyboard Instruments
标题:PHOTON:历史键盘乐器中键杆运动的非侵入式光学跟踪
链接:https://arxiv.org/abs/2604.21682
作者:Noah Jaffe,John Ashley Burgoyne
备注:NIME 2026
摘要:本文介绍了PHOTON(PHysical Optical Tracking of Notes),一种用于测量键盘乐器键杆运动的非侵入式光学传感系统。PHOTON跟踪键杆本身的垂直位移,捕获由表演者输入和乐器的机械施加的随时间变化的负载形成的运动。安装在每个杠杆远端下方的反射式光学传感器提供连续的位移、定时和关节连接数据,而不会干扰动作。与为现代钢琴设计的现有光学系统不同,PHOTON可适应大键琴、击弦琴和早期fortepianos的不同几何形状、有限间隙和非标准布局。其模块化、低轮廓架构可实现跨多个手册和可变按键数的高分辨率、低延迟感测。除了性能捕获,PHOTON还提供实时手势输出,并支持表达性手势,人机交互以及使用真实历史机制构建特定于乐器的手势语料库的实证研究。完整的系统以开源硬件和软件的形式发布,从KiCad开发的原理图和PCB布局到CircuitPython编写的固件,降低了采用,复制和扩展的障碍。
摘要:This paper introduces PHOTON (PHysical Optical Tracking of Notes), a non-invasive optical sensing system for measuring key-lever motion in historical keyboard instruments. PHOTON tracks the vertical displacement of the key lever itself, capturing motion shaped by both performer input and the instrument's mechanically imposed, time-varying load. Reflective optical sensors mounted beneath the distal end of each lever provide continuous displacement, timing, and articulation data without interfering with the action. Unlike existing optical systems designed for modern pianos, PHOTON accommodates the diverse geometries, limited clearances, and non-standard layouts of harpsichords, clavichords, and early fortepianos. Its modular, low-profile architecture enables high-resolution, low-latency sensing across multiple manuals and variable key counts. Beyond performance capture, PHOTON provides real-time MIDI output and supports empirical study of expressive gesture, human-instrument interaction, and the construction of instrument-specific MIDI corpora using real historical mechanisms. The complete system is released as open-source hardware and software, from schematics and PCB layouts developed in KiCad to firmware written in CircuitPython, lowering the barrier to adoption, replication, and extension.


【2】DiariZen Explained: A Tutorial for the Open Source State-of-the-Art Speaker Diarization Pipeline

标题:DiariZen解释:开源最先进的扬声器Dializer管道的开发
链接:https://arxiv.org/abs/2604.21507
作者:Nikhil Raghav
备注:13 pages, 7 figures, 2 tables. Code available at https://github.com/nikhilraghav29/diarizen-tutorial
摘要:说话人日志(SD)是在多说话人音频流中回答“谁在什么时候说话”的任务。典型地,SD系统聚类属于单个说话者的身份的语音片段。近年来,通过端到端神经日志(EEND)方法,SD取得了实质性进展。DiariZen是一种混合SD管道,构建在结构修剪的WavLM-Large编码器,具有powerset分类的Conformer后端和VBx集群上,代表了在编写多个基准测试时领先的开源技术水平。尽管DiariZen架构具有强大的性能,但它跨越了多个存储库和框架,使得研究人员和从业人员难以理解、复制或扩展整个系统。本教程提供了完整的DiariZen管道的独立,逐块解释,将其分解为七个阶段:(1)音频加载和滑动窗口分割,(2)具有学习层加权的WavLM特征提取,(3)Conformer后端和幂集分类,(4)经由叠加的分割聚合,(5)具有重叠排除的说话人嵌入提取,(6)VBx聚类与PLDA评分,以及(7)重建和RTTM输出。对于每个块,我们提供了概念动机,源代码参考,中间张量形状,并从AMI会议语料库中摘录了30年代的实际输出的注释可视化。该实现可在https://github.com/nikhilraghav29/diarizen-tutorial上获得,其中包括每个块的独立可执行脚本和一个端到端运行完整管道的Applyter Notebook。
摘要:Speaker diarization (SD) is the task of answering "who spoke when" in a multi-speaker audio stream. Classically, an SD system clusters segments of speech belonging to an individual speaker's identity. Recent years have seen substantial progress in SD through end-to-end neural diarization (EEND) approaches. DiariZen, a hybrid SD pipeline built upon a structurally pruned WavLM-Large encoder, a Conformer backend with powerset classification, and VBx clustering, represents the leading open-source state of the art at the time of writing across multiple benchmarks. Despite its strong performance, the DiariZen architecture spans several repositories and frameworks, making it difficult for researchers and practitioners to understand, reproduce, or extend the system as a whole. This tutorial paper provides a self-contained, block-by-block explanation of the complete DiariZen pipeline, decomposing it into seven stages: (1) audio loading and sliding window segmentation, (2) WavLM feature extraction with learned layer weighting, (3) Conformer backend and powerset classification, (4) segmentation aggregation via overlap-add, (5) speaker embedding extraction with overlap exclusion, (6) VBx clustering with PLDA scoring, and (7) reconstruction and RTTM output. For each block, we provide the conceptual motivation, source code references, intermediate tensor shapes, and annotated visualizations of the actual outputs on a 30s excerpt from the AMI Meeting Corpus. The implementation is available at https://github.com/nikhilraghav29/diarizen-tutorial, which includes standalone executable scripts for each block and a Jupyter notebook that runs the complete pipeline end-to-end.


【3】Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge

标题:口语对话系统中的全场交互:ICASP 2026 HumDial挑战赛的综合研究
链接:https://arxiv.org/abs/2604.21406
作者:Chengyou Wang,Hongfei Yue,Guojian Li,Zhixian Zhao,Shuiyuan Wang,Shuai Wang,Xin Xu,Hui Bu,Lei Xie
备注:5 pages, 1 figures
摘要:全双工交互,即说话者和听众同时交谈,是传统口语对话系统中经常缺失的人际交流的关键要素。这些系统基于严格的话轮转换模式,很难在动态对话中自然地做出反应。ICASSP 2026类人口语对话系统挑战赛(HumDial Challenge)的全双工交互轨道旨在通过提供处理实时中断,语音重叠和动态转向协商的框架来推进全双工系统的评估。我们介绍了一个全面的基准全双工口语对话系统,从HumDial挑战。我们发布了一个高质量的双通道数据集,包含真实的人类记录对话,捕捉中断,重叠语音和反馈机制。该数据集构成了HumDial-FDBench基准的基础,该基准评估了系统在保持会话流的同时处理中断的能力。此外,我们还创建了一个公共排行榜来比较开源和专有模型的性能,促进透明,可重复的评估。这些资源支持开发更具响应性、适应性和人性化的对话系统。
摘要:Full-duplex interaction, where speakers and listeners converse simultaneously, is a key element of human communication often missing from traditional spoken dialogue systems. These systems, based on rigid turn-taking paradigms, struggle to respond naturally in dynamic conversations. The Full-Duplex Interaction Track of ICASSP 2026 Human-like Spoken Dialogue Systems Challenge (HumDial Challenge) aims to advance the evaluation of full-duplex systems by offering a framework for handling real-time interruptions, speech overlap, and dynamic turn negotiation. We introduce a comprehensive benchmark for full-duplex spoken dialogue systems, built from the HumDial Challenge. We release a high-quality dual-channel dataset of real human-recorded conversations, capturing interruptions, overlapping speech, and feedback mechanisms. This dataset forms the basis for the HumDial-FDBench benchmark, which assesses a system's ability to handle interruptions while maintaining conversational flow. Additionally, we create a public leaderboard to compare the performance of open-source and proprietary models, promoting transparent, reproducible evaluation. These resources support the development of more responsive, adaptive, and human-like dialogue systems.


【4】Dilated CNNs for Periodic Signal Processing: A Low-Complexity Approach

标题:用于周期性信号处理的扩展CNN:一种低复杂性方法
链接:https://arxiv.org/abs/2604.21651
作者:Eli Gildish,Michael Grebshtein,Igor Makienko
备注:16 pages, 8 figures, the use of deep learning in IoT devices
摘要:周期信号的去噪和精确的波形估计是许多信号处理领域的核心任务,包括语音、音乐、医疗诊断、无线电和声纳。尽管深度学习方法最近显示出比经典方法更好的性能,但它们需要大量的计算资源,并且通常针对每个信号观测单独进行训练。这项研究提出了一种基于DCNN和重采样的计算效率高的方法,称为R-DCNN,专为在严格的功率和资源限制下运行而设计。该方法针对具有不同基频的信号,并且只需要一次观测来进行训练。它通过一个轻量级的重新配置步骤来推广到其他信号,该步骤将信号中的时间尺度与不同频率对齐,以重新使用相同的网络权重。尽管R-DCNN的计算复杂度较低,但其性能可与最先进的经典方法(如基于自回归(AR)的技术)以及为每个观察单独训练的传统DCNN相媲美。这种效率和性能的结合使得所提出的方法特别适合在资源受限的环境中部署,而不会牺牲去噪或估计精度。
摘要:Denoising of periodic signals and accurate waveform estimation are core tasks across many signal processing domains, including speech, music, medical diagnostics, radio, and sonar. Although deep learning methods have recently shown performance improvements over classical approaches, they require substantial computational resources and are usually trained separately for each signal observation. This study proposes a computationally efficient method based on DCNN and Re-sampling, termed R-DCNN, designed for operation under strict power and resource constraints. The approach targets signals with varying fundamental frequencies and requires only a single observation for training. It generalizes to additional signals via a lightweight resampling step that aligns time scales in signals with different frequencies to re-use the same network weights. Despite its low computational complexity, R-DCNN achieves performance comparable to state-of-the-art classical methods, such as autoregressive (AR)-based techniques, as well as conventional DCNNs trained individually for each observation. This combination of efficiency and performance makes the proposed method particularly well suited for deployment in resource-constrained environments without sacrificing denoising or estimation accuracy.


机器翻译由腾讯交互翻译提供,仅供参考