【1】Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition标题:无源跨语料语音情感识别中的自适应对比网络链接:https://arxiv.org/abs/2401.12925作者:Yan Zhao,Jincen Wang,Cheng Lu,Sunan Li,Björn Schuller,Yuan Zong,Wenming Zheng备注:Accepted by ICASSP 2024摘要:跨语料库语音情感识别的目的是将情感知识从有标签的语料库迁移到无标签的语料库。然而,现有方法需要在适应期间访问源数据,由于数据隐私保护问题,这在现实生活场景中是无法实现的。本文解决了一个更实际的任务,即无源跨语料库SER,其中预先训练的源模型适用于目标域,而无需访问源数据。为了解决这个问题,我们提出了一种新的方法称为情感感知对比适应网络(ECAN)。其核心思想是在考虑全局类级自适应的同时,捕捉样本间的局部邻域信息。具体来说,我们提出了一个最近邻对比学习,以促进高度相似的样本的特征之间的局部情感一致性。此外,仅仅依赖最近的邻域可能导致聚类之间的边界模糊。因此,我们结合了监督对比学习,以鼓励代表不同情绪的聚类之间的更大分离,从而促进改进类级适应。大量的实验表明,我们提出的ECAN显着优于国家的最先进的方法下的无源跨语料库SER设置几个语音情感语料库。摘要:Cross-corpus speech emotion recognition (SER) aims to transfer emotional knowledge from a labeled source corpus to an unlabeled corpus. However, prior methods require access to source data during adaptation, which is unattainable in real-life scenarios due to data privacy protection concerns. This paper tackles a more practical task, namely source-free cross-corpus SER, where a pre-trained source model is adapted to the target domain without access to source data. To address the problem, we propose a novel method called emotion-aware contrastive adaptation network (ECAN). The core idea is to capture local neighborhood information between samples while considering the global class-level adaptation. Specifically, we propose a nearest neighbor contrastive learning to promote local emotion consistency among features of highly similar samples. Furthermore, relying solely on nearest neighborhoods may lead to ambiguous boundaries between clusters. Thus, we incorporate supervised contrastive learning to encourage greater separation between clusters representing different emotions, thereby facilitating improved class-level adaptation. Extensive experiments indicate that our proposed ECAN significantly outperforms state-of-the-art methods under the source-free cross-corpus SER setting on several speech emotion corpora.
【2】 Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study标题:多语言、完全非自回归ASR与大语言模型融合的综合研究链接:https://arxiv.org/abs/2401.12789作者:W. Ronny Huang,Cyril Allauzen,Tongzhou Chen,Kilol Gupta,Ke Hu,James Qin,Yu Zhang,Yongqiang Wang,Shuo-Yiin Chang,Tara N. Sainath备注:ICASSP 2024摘要:在大型模型时代,解码的自回归性质通常导致延迟成为一个重要的瓶颈。我们提出了一个非自回归LM融合ASR系统,有效地利用加速器硬件的并行化能力。我们的方法结合了通用语音模型(USM)和PaLM 2语言模型在每段评分模式,实现了平均相对WER改善所有语言的10.8%的FLEURS和3.6%的YouTube字幕。此外,我们的综合消融研究分析了关键参数,如LLM大小,上下文长度,词汇量,融合方法。例如,我们探讨了从128 M到340 B的LLM大小参数对ASR性能的影响。这项研究提供了有价值的见解的影响因素,实际的大规模LM融合语音识别系统的有效性。摘要:In the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system that effectively leverages the parallelization capabilities of accelerator hardware. Our approach combines the Universal Speech Model (USM) and the PaLM 2 language model in per-segment scoring mode, achieving an average relative WER improvement across all languages of 10.8% on FLEURS and 3.6% on YouTube captioning. Furthermore, our comprehensive ablation study analyzes key parameters such as LLM size, context length, vocabulary size, fusion methodology. For instance, we explore the impact of LLM size ranging from 128M to 340B parameters on ASR performance. This study provides valuable insights into the factors influencing the effectiveness of practical large-scale LM-fused speech recognition systems. 【3】 MoodLoopGP: Generating Emotion-Conditioned Loop Tablature Music with Multi-Granular Features标题:MoodLoopGP:生成具有多粒度特征的情感条件循环曲谱音乐链接:https://arxiv.org/abs/2401.12656作者:Wenqian Cui,Pedro Sarmento,Mathieu Barthet备注:This preprint is licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0). The Version of Record of this contribution is published in Proceedings of EvoMUSART: International Conference on Computational Intelligence in Music, Sound, Art and Design (Part of EvoStar) 2024摘要:可循环音乐生成系统支持多种应用,但它们通常缺乏可控性和定制能力。我们认为,提高可控性可以丰富这些模型,情感表达是创作者和听众的一个重要方面。因此,基于LooperGP,一个可循环的指谱生成模型,本文探讨了赋予系统对所传达的情感的控制。为了实现这样的条件生成,我们建议在模型训练和推理过程中利用多粒度语义和音乐特征来整合音乐知识。具体来说,我们将歌曲级别的功能(情感标签,节奏和模式)和酒吧级别的功能(音调张力)一起指导情感表达。通过算法和人类的评估,我们证明了这种方法在制作传达两种截然不同的目标情感(幸福和悲伤)的音乐方面的有效性。还进行了消融研究,以澄清我们的方法的结果背后的影响因素。摘要:Loopable music generation systems enable diverse applications, but they often lack controllability and customization capabilities. We argue that enhancing controllability can enrich these models, with emotional expression being a crucial aspect for both creators and listeners. Hence, building upon LooperGP, a loopable tablature generation model, this paper explores endowing systems with control over conveyed emotions. To enable such conditional generation, we propose integrating musical knowledge by utilizing multi-granular semantic and musical features during model training and inference. Specifically, we incorporate song-level features (Emotion Labels, Tempo, and Mode) and bar-level features (Tonal Tension) together to guide emotional expression. Through algorithmic and human evaluations, we demonstrate the approach's effectiveness in producing music conveying two contrasting target emotions, happiness and sadness. An ablation study is also conducted to clarify the contributing factors behind our approach's results.
【4】 EEND-M2F: Masked-attention mask transformers for speaker diarization标题:EEND-M2F:用于扬声器二值化的屏蔽注意掩码转换器链接:https://arxiv.org/abs/2401.12600作者:Marc Härkönen,Samuel J. Broughton,Lahiru Samarakoon备注:14 pages, 2 figures摘要:在本文中,我们明确的图像分割方法和端到端的日记化方法之间的联系。从这些见解,我们提出了一种新的,完全端到端的日记模型,EEND-M2 F,基于Mask 2Former架构。使用Transformer解码器的堆栈并行计算扬声器表示,其中使用来自先前层的预测明确地从交叉关注中屏蔽不相关的帧。EEND-M2 F是轻量级的,高效的,真正的端到端,因为它不需要运行任何额外的日志,说话人验证或分割模型,也不需要运行任何聚类算法。我们的模型在几个公共数据集上实现了最先进的性能,如AMI,AliMeeting和RAMC。最值得注意的是,我们在DIHARD-III上的16.07%的DER是挑战获胜系统的第一个重大改进。摘要:In this paper, we make the explicit connection between image segmentation methods and end-to-end diarization methods. From these insights, we propose a novel, fully end-to-end diarization model, EEND-M2F, based on the Mask2Former architecture. Speaker representations are computed in parallel using a stack of transformer decoders, in which irrelevant frames are explicitly masked from the cross attention using predictions from previous layers. EEND-M2F is lightweight, efficient, and truly end-to-end, as it does not require any additional diarization, speaker verification, or segmentation models to run, nor does it require running any clustering algorithms. Our model achieves state-of-the-art performance on several public datasets, such as AMI, AliMeeting and RAMC. Most notably our DER of 16.07% on DIHARD-III is the first major improvement upon the challenge winning system. 【5】 An Exploratory Study of Multimodal Physiological Data in Jazz Improvisation Using Basic Machine Learning Techniques标题:基于基本机器学习技术的爵士乐即兴演奏多通道生理数据的探索性研究链接:https://arxiv.org/abs/2401.12266作者:Yawen Zhang备注:Master's thesis摘要:我们的研究深入到“即兴音乐数据集”,探索即兴音乐表演过程中生理和心理维度之间的相互交织的关系和相关性。主要目的是确定这些状态之间存在明确的因果关系或相关关系,并理解它们在音乐作品中的表现。这个丰富的数据集提供了一个关于音乐家如何在实时即兴表演场景中协调他们的身体与声音事件的视角,强调了“即兴音乐”的概念。“摘要:Our study delves into the "Embodied Musicking Dataset," exploring the intertwined relationships and correlations between physiological and psychological dimensions during improvisational music performances. The primary objective is to ascertain the presence of a definitive causal or correlational relationship between these states and comprehend their manifestation in musical compositions. This rich dataset provides a perspective on how musicians coordinate their physicality with sonic events in real-time improvisational scenarios, emphasizing the concept of "Embodied Musicking." 【6】 Overlap-aware End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization标题:基于重叠感知的端到端监督层次图聚类的说话人日志化链接:https://arxiv.org/abs/2401.12850作者:Prachi Singh,Sriram Ganapathy备注:10 pages摘要:说话人日志化是基于说话人身份对音频记录进行分段的任务,它构成了一些下游应用的重要语音预处理步骤。传统的日志化方法涉及嵌入提取和聚类的多个步骤,这些步骤通常以孤立的方式进行优化。虽然端到端日志化系统尝试学习用于任务的单个模型,但它们通常训练起来很麻烦,并且需要大的监督数据集。在本文中,我们提出了一个端到端的监督层次聚类算法的基础上图神经网络(GNN),称为端到端的监督层次聚类(E-SHARC)。E-SHARC方法使用前端mel滤波器组特征作为输入,并联合学习嵌入提取器和GNN聚类模块,执行表示学习,度量学习和端到端优化的聚类。此外,利用来自外部重叠检测器的附加输入,E-SHARC方法能够预测重叠语音区域中的说话者。在AMI、VoxConverse和DISPLACE等几个基准数据集上的实验结果表明,E-SHARC框架比现有的日志系统有了显著的改进。摘要:Speaker diarization, the task of segmenting an audio recording based on speaker identity, constitutes an important speech pre-processing step for several downstream applications. The conventional approach to diarization involves multiple steps of embedding extraction and clustering, which are often optimized in an isolated fashion. While end-to-end diarization systems attempt to learn a single model for the task, they are often cumbersome to train and require large supervised datasets. In this paper, we propose an end-to-end supervised hierarchical clustering algorithm based on graph neural networks (GNN), called End-to-end Supervised HierARchical Clustering (E-SHARC). The E-SHARC approach uses front-end mel-filterbank features as input and jointly learns an embedding extractor and the GNN clustering module, performing representation learning, metric learning, and clustering with end-to-end optimization. Further, with additional inputs from an external overlap detector, the E-SHARC approach is capable of predicting the speakers in the overlapping speech regions. The experimental evaluation on several benchmark datasets like AMI, VoxConverse and DISPLACE, illustrates that the proposed E-SHARC framework improves significantly over the state-of-art diarization systems. 【7】 DiffMoog: a Differentiable Modular Synthesizer for Sound Matching标题:DiffMoog:一种用于声音匹配的可差分模块合成器链接:https://arxiv.org/abs/2401.12570作者:Noy Uzrad,Oren Barkan,Almog Elharar,Shlomi Shvartzman,Moshe Laufer,Lior Wolf,Noam Koenigstein备注:5 pages, 7 figures, 1 table, Our code is released at this https URL摘要:本文介绍了DiffMoog -一个微分模块化合成器,具有一套全面的模块,通常在商业仪器中找到。由于是可区分的,它允许集成到神经网络中,从而实现自动声音匹配,以复制给定的音频输入。值得注意的是,DiffMoog促进了调制功能(FM/AM),低频振荡器(LFO),滤波器,包络整形器以及用户创建自定义信号链的能力。我们介绍了一个开源平台,包括DiffMoog和端到端的声音匹配框架。该框架利用了一种新的信号链损失和编码器网络,自编程的输出预测DiffMoogs参数的基础上,用户定义的模块化架构。此外,我们提供的见解和经验教训,对声音匹配使用微分合成。DiffMoog将强大的声音功能与整体平台相结合,是加快音频合成和机器学习研究的首要资产。摘要:This paper presents DiffMoog - a differentiable modular synthesizer with a comprehensive set of modules typically found in commercial instruments. Being differentiable, it allows integration into neural networks, enabling automated sound matching, to replicate a given audio input. Notably, DiffMoog facilitates modulation capabilities (FM/AM), low-frequency oscillators (LFOs), filters, envelope shapers, and the ability for users to create custom signal chains. We introduce an open-source platform that comprises DiffMoog and an end-to-end sound matching framework. This framework utilizes a novel signal-chain loss and an encoder network that self-programs its outputs to predict DiffMoogs parameters based on the user-defined modular architecture. Moreover, we provide insights and lessons learned towards sound matching using differentiable synthesis. Combining robust sound capabilities with a holistic platform, DiffMoog stands as a premier asset for expediting research in audio synthesis and machine learning. 【8】 Boosting Unknown-number Speaker Separation with Transformer Decoder-based Attractor标题:基于Transformer解码器的吸引子增强未知数说话人分离链接:https://arxiv.org/abs/2401.12473作者:Younglo Lee,Shukjae Choi,Byeong-Yeol Kim,Zhong-Qiu Wang,Shinji Watanabe备注:5 pages, 4 figures, accepted by ICASSP 2024摘要:我们提出了一种新的语音分离模型,旨在分离混合物与未知数量的扬声器。所提出的模型堆叠1)可以对频谱-时间模式进行建模的双路径处理块,2)可以处理未知数量的扬声器的基于Transformer解码器的吸引子(TDA)计算模块,以及3)可以对扬声器间关系进行建模的三路径处理块。给定一个固定的,小的一组学习的扬声器查询和混合嵌入的双路径块,TDA推断这些查询的关系,并为每个扬声器生成一个吸引子向量。然后,通过特征线性调制条件将估计的吸引子与混合嵌入相结合,从而创建扬声器维度。以TDA产生的扬声器信息为条件的混合嵌入被馈送到最终的三路径块,该三路径块用专用于扬声器间处理的附加路径来增强双路径块。所提出的方法优于文献中先前报道的最佳方法,在WSJ 0 -2和3 mix上分别实现24.0和23.7 dB SI-SDR改进(SI-SDRi),单个模型被训练用于分离2和3扬声器混合。所提出的模型也表现出很强的性能和泛化能力,在计数源和分离混合与多达5个扬声器。摘要:We propose a novel speech separation model designed to separate mixtures with an unknown number of speakers. The proposed model stacks 1) a dual-path processing block that can model spectro-temporal patterns, 2) a transformer decoder-based attractor (TDA) calculation module that can deal with an unknown number of speakers, and 3) triple-path processing blocks that can model inter-speaker relations. Given a fixed, small set of learned speaker queries and the mixture embedding produced by the dual-path blocks, TDA infers the relations of these queries and generates an attractor vector for each speaker. The estimated attractors are then combined with the mixture embedding by feature-wise linear modulation conditioning, creating a speaker dimension. The mixture embedding, conditioned with speaker information produced by TDA, is fed to the final triple-path blocks, which augment the dual-path blocks with an additional pathway dedicated to inter-speaker processing. The proposed approach outperforms the previous best reported in the literature, achieving 24.0 and 23.7 dB SI-SDR improvement (SI-SDRi) on WSJ0-2 and 3mix respectively, with a single model trained to separate 2- and 3-speaker mixtures. The proposed model also exhibits strong performance and generalizability at counting sources and separating mixtures with up to 5 speakers. 【9】 Post-Training Embedding Alignment for Decoupling Enrollment and Runtime Speaker Recognition Models标题:训练后嵌入对齐用于解耦注册和运行时说话人识别模型链接:https://arxiv.org/abs/2401.12440作者:Chenyang Gao,Brecht Desplanques,Chelsea J. -T. Ju,Aman Chadha,Andreas Stolcke备注:Accepted to ICASSP 2024摘要:自动说话人识别(SID)是实现各种语音服务个性化的关键一步。典型的SID系统使用具有单个模型的对称登记验证框架来离线地导出从登记话语提取的语音简档的嵌入,以及在线地从运行时话语导出嵌入。由于注册和运行时的不同情况,例如不同的计算和延迟约束,一些应用程序将受益于使用不同模型进行注册和运行时嵌入生成的非对称注册验证框架。为了支持这种非对称SID,其中两个模型中的每一个都可以独立更新,我们建议使用轻量级神经网络将两个独立模型的嵌入映射到共享的说话人嵌入空间。我们的研究结果表明,这种方法显着优于余弦评分在共享扬声器logit空间中的模型,这些模型是在具有许多扬声器身份的大型数据集上进行对比损失训练的。这种建议的神经嵌入扬声器空间对齐(NESSA)结合只有一个模型的不对称更新提供了至少60%的性能增益,通过更新两个模型在标准的对称SID方法。摘要:Automated speaker identification (SID) is a crucial step for the personalization of a wide range of speech-enabled services. Typical SID systems use a symmetric enrollment-verification framework with a single model to derive embeddings both offline for voice profiles extracted from enrollment utterances, and online from runtime utterances. Due to the distinct circumstances of enrollment and runtime, such as different computation and latency constraints, several applications would benefit from an asymmetric enrollment-verification framework that uses different models for enrollment and runtime embedding generation. To support this asymmetric SID where each of the two models can be updated independently, we propose using a lightweight neural network to map the embeddings from the two independent models to a shared speaker embedding space. Our results show that this approach significantly outperforms cosine scoring in a shared speaker logit space for models that were trained with a contrastive loss on large datasets with many speaker identities. This proposed Neural Embedding Speaker Space Alignment (NESSA) combined with an asymmetric update of only one of the models delivers at least 60% of the performance gain achieved by updating both models in the standard symmetric SID approach. 【10】 CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing标题:CoAVT:一种面向多模式处理的认知启发式统一视听文本预训练模型链接:https://arxiv.org/abs/2401.12264作者:Xianghu Yue,Xiaohai Tian,Malu Zhang,Zhizheng Wu,Haizhou Li摘要:长期以来,人们一直在寻求一种统一的视听文本模型,以实现各种多模态理解任务,该任务模仿人类的听,看和阅读过程。人类倾向于使用两个独立的系统来表示知识:一个用于表示语言(文本)信息,另一个用于表示非语言(视觉和听觉)信息。这两个系统可以独立运行,但也可以相互作用。出于对人类认知的这种理解,本文引入了CoAVT --一种新颖的认知启发的相关视听文本预训练模型来连接这三种模态。它包含一个联合视听编码器,学习将视听同步信息与非语言信息的视听内容一起编码,以及一个文本编码器,用于处理语言信息的文本输入。为了弥合模态之间的差距,CoAVT采用查询编码器,其中包含一组可学习的查询嵌入,并提取相应文本的最丰富的视听功能。此外,为了利用音频和视觉分别与语言之间的对应关系,我们还在基础视听文本三模态对齐的基础上建立了视听文本和视觉文本双模态对齐,以增强多模态表征学习。最后,我们联合优化CoAVT模型的三个多模态目标:对比损失,匹配损失和语言建模损失。大量的实验表明,CoAVT可以学习强多模态相关性,并推广到各种下游任务。CoAVT为AudioCaps上的文本视频检索任务建立了新的最先进的性能,用于AudioSet和VGGSound上的zero-shot和微调设置、视听事件分类和视听检索任务。摘要:There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represent knowledge using two separate systems: one for representing verbal (textual) information and one for representing non-verbal (visual and auditory) information. These two systems can operate independently but can also interact with each other. Motivated by this understanding of human cognition, in this paper, we introduce CoAVT -- a novel cognition-inspired Correlated Audio-Visual-Text pre-training model to connect the three modalities. It contains a joint audio-visual encoder that learns to encode audio-visual synchronization information together with the audio and visual content for non-verbal information, and a text encoder to handle textual input for verbal information. To bridge the gap between modalities, CoAVT employs a query encoder, which contains a set of learnable query embeddings, and extracts the most informative audiovisual features of the corresponding text. Additionally, to leverage the correspondences between audio and vision with language respectively, we also establish the audio-text and visual-text bi-modal alignments upon the foundational audiovisual-text tri-modal alignment to enhance the multimodal representation learning. Finally, we jointly optimize CoAVT model with three multimodal objectives: contrastive loss, matching loss and language modeling loss. Extensive experiments show that CoAVT can learn strong multimodal correlations and be generalized to various downstream tasks. CoAVT establishes new state-of-the-art performance on text-video retrieval task on AudioCaps for both zero-shot and fine-tuning settings, audio-visual event classification and audio-visual retrieval tasks on AudioSet and VGGSound. 【11】 Spatial Scaper: A Library to Simulate and Augment Soundscapes for Sound Event Localization and Detection in Realistic Rooms标题:Space Sfaer:模拟和增强真实感房间声事件定位和检测的声景库链接:https://arxiv.org/abs/2401.12238作者:Iran R. Roman,Christopher Ick,Sivan Ding,Adrian S. Roman,Brian McFee,Juan P. Bello备注:5 pages, 4 figures, 1 table, to be presented at ICASSP 2024 in Seoul, South Korea摘要:声音事件定位与检测是机器听觉中的一项重要任务。主要的进步依赖于模拟数据与声音事件在特定的房间和强大的时空标签。SELD数据是通过将空间定位的房间脉冲响应(RIR)与声音波形进行卷积来模拟的,以将声音事件放置在声景中。然而,RIR需要在特定房间进行人工采集。我们提出SpatialScaper,SELD数据模拟和增强库。与现有工具相比,SpatialScaper通过尺寸和墙壁吸收等参数模拟虚拟房间。这允许前景和背景声源的参数化放置(包括移动)。SpatialScaper还包括可应用于现有SELD数据的数据增强管道。作为一个案例研究,我们使用SpatialScaper向DCASE SELD数据添加房间。用我们的数据训练模型导致渐进性能的改善,这是声学多样性的直接函数。这些结果表明,SpatialScaper对训练鲁棒SELD模型是有价值的.摘要:Sound event localization and detection (SELD) is an important task in machine listening. Major advancements rely on simulated data with sound events in specific rooms and strong spatio-temporal labels. SELD data is simulated by convolving spatialy-localized room impulse responses (RIRs) with sound waveforms to place sound events in a soundscape. However, RIRs require manual collection in specific rooms. We present SpatialScaper, a library for SELD data simulation and augmentation. Compared to existing tools, SpatialScaper emulates virtual rooms via parameters such as size and wall absorption. This allows for parameterized placement (including movement) of foreground and background sound sources. SpatialScaper also includes data augmentation pipelines that can be applied to existing SELD data. As a case study, we use SpatialScaper to add rooms to the DCASE SELD data. Training a model with our data led to progressive performance improves as a direct function of acoustic diversity. These results show that SpatialScaper is valuable to train robust SELD models.
eess.AS音频处理【1】 Overlap-aware End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization标题:基于重叠感知的端到端监督层次图聚类的说话人日志化链接:https://arxiv.org/abs/2401.12850作者:Prachi Singh,Sriram Ganapathy备注:10 pages摘要:说话人日志化是基于说话人身份对音频记录进行分段的任务,它构成了一些下游应用的重要语音预处理步骤。传统的日志化方法涉及嵌入提取和聚类的多个步骤,这些步骤通常以孤立的方式进行优化。虽然端到端日志化系统尝试学习用于任务的单个模型,但它们通常训练起来很麻烦,并且需要大的监督数据集。在本文中,我们提出了一个端到端的监督层次聚类算法的基础上图神经网络(GNN),称为端到端的监督层次聚类(E-SHARC)。E-SHARC方法使用前端mel滤波器组特征作为输入,并联合学习嵌入提取器和GNN聚类模块,执行表示学习,度量学习和端到端优化的聚类。此外,利用来自外部重叠检测器的附加输入,E-SHARC方法能够预测重叠语音区域中的说话者。在AMI、VoxConverse和DISPLACE等几个基准数据集上的实验结果表明,E-SHARC框架比现有的日志系统有了显著的改进。摘要:Speaker diarization, the task of segmenting an audio recording based on speaker identity, constitutes an important speech pre-processing step for several downstream applications. The conventional approach to diarization involves multiple steps of embedding extraction and clustering, which are often optimized in an isolated fashion. While end-to-end diarization systems attempt to learn a single model for the task, they are often cumbersome to train and require large supervised datasets. In this paper, we propose an end-to-end supervised hierarchical clustering algorithm based on graph neural networks (GNN), called End-to-end Supervised HierARchical Clustering (E-SHARC). The E-SHARC approach uses front-end mel-filterbank features as input and jointly learns an embedding extractor and the GNN clustering module, performing representation learning, metric learning, and clustering with end-to-end optimization. Further, with additional inputs from an external overlap detector, the E-SHARC approach is capable of predicting the speakers in the overlapping speech regions. The experimental evaluation on several benchmark datasets like AMI, VoxConverse and DISPLACE, illustrates that the proposed E-SHARC framework improves significantly over the state-of-art diarization systems.
【2】 DiffMoog: a Differentiable Modular Synthesizer for Sound Matching标题:DiffMoog:一种用于声音匹配的可微分模块化合成器链接:https://arxiv.org/abs/2401.12570作者:Noy Uzrad,Oren Barkan,Almog Elharar,Shlomi Shvartzman,Moshe Laufer,Lior Wolf,Noam Koenigstein备注:5 pages, 7 figures, 1 table, Our code is released at this https URL摘要:本文介绍了DiffMoog -一个微分模块化合成器,具有一套全面的模块,通常在商业仪器中找到。由于是可区分的,它允许集成到神经网络中,从而实现自动声音匹配,以复制给定的音频输入。值得注意的是,DiffMoog促进了调制功能(FM/AM),低频振荡器(LFO),滤波器,包络整形器以及用户创建自定义信号链的能力。我们介绍了一个开源平台,包括DiffMoog和端到端的声音匹配框架。该框架利用了一种新的信号链损失和编码器网络,自编程的输出预测DiffMoogs参数的基础上,用户定义的模块化架构。此外,我们提供的见解和经验教训,对声音匹配使用微分合成。DiffMoog将强大的声音功能与整体平台相结合,是加快音频合成和机器学习研究的首要资产。摘要:This paper presents DiffMoog - a differentiable modular synthesizer with a comprehensive set of modules typically found in commercial instruments. Being differentiable, it allows integration into neural networks, enabling automated sound matching, to replicate a given audio input. Notably, DiffMoog facilitates modulation capabilities (FM/AM), low-frequency oscillators (LFOs), filters, envelope shapers, and the ability for users to create custom signal chains. We introduce an open-source platform that comprises DiffMoog and an end-to-end sound matching framework. This framework utilizes a novel signal-chain loss and an encoder network that self-programs its outputs to predict DiffMoogs parameters based on the user-defined modular architecture. Moreover, we provide insights and lessons learned towards sound matching using differentiable synthesis. Combining robust sound capabilities with a holistic platform, DiffMoog stands as a premier asset for expediting research in audio synthesis and machine learning. 【3】 Boosting Unknown-number Speaker Separation with Transformer Decoder-based Attractor标题:基于Transformer解码器的吸引子增强未知数说话人分离链接:https://arxiv.org/abs/2401.12473作者:Younglo Lee,Shukjae Choi,Byeong-Yeol Kim,Zhong-Qiu Wang,Shinji Watanabe备注:5 pages, 4 figures, accepted by ICASSP 2024摘要:我们提出了一种新的语音分离模型,旨在分离混合物与未知数量的扬声器。所提出的模型堆叠1)可以对频谱-时间模式进行建模的双路径处理块,2)可以处理未知数量的扬声器的基于Transformer解码器的吸引子(TDA)计算模块,以及3)可以对扬声器间关系进行建模的三路径处理块。给定一个固定的,小的一组学习的扬声器查询和混合嵌入的双路径块,TDA推断这些查询的关系,并为每个扬声器生成一个吸引子向量。然后,通过特征线性调制条件将估计的吸引子与混合嵌入相结合,从而创建扬声器维度。以TDA产生的扬声器信息为条件的混合嵌入被馈送到最终的三路径块,该三路径块用专用于扬声器间处理的附加路径来增强双路径块。所提出的方法优于文献中先前报道的最佳方法,在WSJ 0 -2和3 mix上分别实现24.0和23.7 dB SI-SDR改进(SI-SDRi),单个模型被训练用于分离2和3扬声器混合。所提出的模型也表现出很强的性能和泛化能力,在计数源和分离混合与多达5个扬声器。摘要:We propose a novel speech separation model designed to separate mixtures with an unknown number of speakers. The proposed model stacks 1) a dual-path processing block that can model spectro-temporal patterns, 2) a transformer decoder-based attractor (TDA) calculation module that can deal with an unknown number of speakers, and 3) triple-path processing blocks that can model inter-speaker relations. Given a fixed, small set of learned speaker queries and the mixture embedding produced by the dual-path blocks, TDA infers the relations of these queries and generates an attractor vector for each speaker. The estimated attractors are then combined with the mixture embedding by feature-wise linear modulation conditioning, creating a speaker dimension. The mixture embedding, conditioned with speaker information produced by TDA, is fed to the final triple-path blocks, which augment the dual-path blocks with an additional pathway dedicated to inter-speaker processing. The proposed approach outperforms the previous best reported in the literature, achieving 24.0 and 23.7 dB SI-SDR improvement (SI-SDRi) on WSJ0-2 and 3mix respectively, with a single model trained to separate 2- and 3-speaker mixtures. The proposed model also exhibits strong performance and generalizability at counting sources and separating mixtures with up to 5 speakers. 【4】 Post-Training Embedding Alignment for Decoupling Enrollment and Runtime Speaker Recognition Models标题:解耦登记和分离说话人识别模型的训练后嵌入对齐链接:https://arxiv.org/abs/2401.12440作者:Chenyang Gao,Brecht Desplanques,Chelsea J. -T. Ju,Aman Chadha,Andreas Stolcke备注:Accepted to ICASSP 2024摘要:自动说话人识别(SID)是实现各种语音服务个性化的关键一步。典型的SID系统使用具有单个模型的对称登记验证框架来离线地导出从登记话语提取的语音简档的嵌入,以及在线地从运行时话语导出嵌入。由于注册和运行时的不同情况,例如不同的计算和延迟约束,一些应用程序将受益于使用不同模型进行注册和运行时嵌入生成的非对称注册验证框架。为了支持这种非对称SID,其中两个模型中的每一个都可以独立更新,我们建议使用轻量级神经网络将两个独立模型的嵌入映射到共享的说话人嵌入空间。我们的研究结果表明,这种方法显着优于余弦评分在共享扬声器logit空间中的模型,这些模型是在具有许多扬声器身份的大型数据集上进行对比损失训练的。这种建议的神经嵌入扬声器空间对齐(NESSA)结合只有一个模型的不对称更新提供了至少60%的性能增益,通过更新两个模型在标准的对称SID方法。摘要:Automated speaker identification (SID) is a crucial step for the personalization of a wide range of speech-enabled services. Typical SID systems use a symmetric enrollment-verification framework with a single model to derive embeddings both offline for voice profiles extracted from enrollment utterances, and online from runtime utterances. Due to the distinct circumstances of enrollment and runtime, such as different computation and latency constraints, several applications would benefit from an asymmetric enrollment-verification framework that uses different models for enrollment and runtime embedding generation. To support this asymmetric SID where each of the two models can be updated independently, we propose using a lightweight neural network to map the embeddings from the two independent models to a shared speaker embedding space. Our results show that this approach significantly outperforms cosine scoring in a shared speaker logit space for models that were trained with a contrastive loss on large datasets with many speaker identities. This proposed Neural Embedding Speaker Space Alignment (NESSA) combined with an asymmetric update of only one of the models delivers at least 60% of the performance gain achieved by updating both models in the standard symmetric SID approach.
【5】 CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing标题:CoAVT:一种基于认知启发的多模态视听文本统一预训练模型链接:https://arxiv.org/abs/2401.12264作者:Xianghu Yue,Xiaohai Tian,Malu Zhang,Zhizheng Wu,Haizhou Li摘要:长期以来,人们一直在寻求一种统一的视听文本模型,以实现各种多模态理解任务,该任务模仿人类的听,看和阅读过程。人类倾向于使用两个独立的系统来表示知识:一个用于表示语言(文本)信息,另一个用于表示非语言(视觉和听觉)信息。这两个系统可以独立运行,但也可以相互作用。出于对人类认知的这种理解,本文引入了CoAVT --一种新颖的认知启发的相关视听文本预训练模型来连接这三种模态。它包含一个联合视听编码器,学习将视听同步信息与非语言信息的视听内容一起编码,以及一个文本编码器,用于处理语言信息的文本输入。为了弥合模态之间的差距,CoAVT采用查询编码器,其中包含一组可学习的查询嵌入,并提取相应文本的最丰富的视听功能。此外,为了利用音频和视觉分别与语言之间的对应关系,我们还在基础视听文本三模态对齐的基础上建立了视听文本和视觉文本双模态对齐,以增强多模态表征学习。最后,我们联合优化CoAVT模型的三个多模态目标:对比损失,匹配损失和语言建模损失。大量的实验表明,CoAVT可以学习强多模态相关性,并推广到各种下游任务。CoAVT为AudioCaps上的文本视频检索任务建立了新的最先进的性能,用于AudioSet和VGGSound上的zero-shot和微调设置、视听事件分类和视听检索任务。摘要:There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represent knowledge using two separate systems: one for representing verbal (textual) information and one for representing non-verbal (visual and auditory) information. These two systems can operate independently but can also interact with each other. Motivated by this understanding of human cognition, in this paper, we introduce CoAVT -- a novel cognition-inspired Correlated Audio-Visual-Text pre-training model to connect the three modalities. It contains a joint audio-visual encoder that learns to encode audio-visual synchronization information together with the audio and visual content for non-verbal information, and a text encoder to handle textual input for verbal information. To bridge the gap between modalities, CoAVT employs a query encoder, which contains a set of learnable query embeddings, and extracts the most informative audiovisual features of the corresponding text. Additionally, to leverage the correspondences between audio and vision with language respectively, we also establish the audio-text and visual-text bi-modal alignments upon the foundational audiovisual-text tri-modal alignment to enhance the multimodal representation learning. Finally, we jointly optimize CoAVT model with three multimodal objectives: contrastive loss, matching loss and language modeling loss. Extensive experiments show that CoAVT can learn strong multimodal correlations and be generalized to various downstream tasks. CoAVT establishes new state-of-the-art performance on text-video retrieval task on AudioCaps for both zero-shot and fine-tuning settings, audio-visual event classification and audio-visual retrieval tasks on AudioSet and VGGSound. 【6】 Spatial Scaper: A Library to Simulate and Augment Soundscapes for Sound Event Localization and Detection in Realistic Rooms标题:Spatial Scaper:模拟和增强真实房间中声音事件定位和检测的音景库链接:https://arxiv.org/abs/2401.12238作者:Iran R. Roman,Christopher Ick,Sivan Ding,Adrian S. Roman,Brian McFee,Juan P. Bello备注:5 pages, 4 figures, 1 table, to be presented at ICASSP 2024 in Seoul, South Korea摘要:声音事件定位与检测是机器听觉中的一项重要任务。主要的进步依赖于模拟数据与声音事件在特定的房间和强大的时空标签。SELD数据是通过将空间定位的房间脉冲响应(RIR)与声音波形进行卷积来模拟的,以将声音事件放置在声景中。然而,RIR需要在特定房间进行人工采集。我们提出SpatialScaper,SELD数据模拟和增强库。与现有工具相比,SpatialScaper通过尺寸和墙壁吸收等参数模拟虚拟房间。这允许前景和背景声源的参数化放置(包括移动)。SpatialScaper还包括可应用于现有SELD数据的数据增强管道。作为一个案例研究,我们使用SpatialScaper向DCASE SELD数据添加房间。用我们的数据训练模型导致渐进性能的改善,这是声学多样性的直接函数。这些结果表明,SpatialScaper对训练鲁棒SELD模型是有价值的.摘要:Sound event localization and detection (SELD) is an important task in machine listening. Major advancements rely on simulated data with sound events in specific rooms and strong spatio-temporal labels. SELD data is simulated by convolving spatialy-localized room impulse responses (RIRs) with sound waveforms to place sound events in a soundscape. However, RIRs require manual collection in specific rooms. We present SpatialScaper, a library for SELD data simulation and augmentation. Compared to existing tools, SpatialScaper emulates virtual rooms via parameters such as size and wall absorption. This allows for parameterized placement (including movement) of foreground and background sound sources. SpatialScaper also includes data augmentation pipelines that can be applied to existing SELD data. As a case study, we use SpatialScaper to add rooms to the DCASE SELD data. Training a model with our data led to progressive performance improves as a direct function of acoustic diversity. These results show that SpatialScaper is valuable to train robust SELD models. 【7】 Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition标题:跨语料库无源语音情感识别的情感感知对比适应网络链接:https://arxiv.org/abs/2401.12925作者:Yan Zhao,Jincen Wang,Cheng Lu,Sunan Li,Björn Schuller,Yuan Zong,Wenming Zheng备注:Accepted by ICASSP 2024摘要:跨语料库语音情感识别的目的是将情感知识从有标签的语料库迁移到无标签的语料库。然而,现有方法需要在适应期间访问源数据,由于数据隐私保护问题,这在现实生活场景中是无法实现的。本文解决了一个更实际的任务,即无源跨语料库SER,其中预先训练的源模型适用于目标域,而无需访问源数据。为了解决这个问题,我们提出了一种新的方法称为情感感知对比适应网络(ECAN)。其核心思想是在考虑全局类级自适应的同时,捕捉样本间的局部邻域信息。具体来说,我们提出了一个最近邻对比学习,以促进高度相似的样本的特征之间的局部情感一致性。此外,仅仅依赖最近的邻域可能导致聚类之间的边界模糊。因此,我们结合了监督对比学习,以鼓励代表不同情绪的聚类之间的更大分离,从而促进改进类级适应。大量的实验表明,我们提出的ECAN显着优于国家的最先进的方法下的无源跨语料库SER设置几个语音情感语料库。摘要:Cross-corpus speech emotion recognition (SER) aims to transfer emotional knowledge from a labeled source corpus to an unlabeled corpus. However, prior methods require access to source data during adaptation, which is unattainable in real-life scenarios due to data privacy protection concerns. This paper tackles a more practical task, namely source-free cross-corpus SER, where a pre-trained source model is adapted to the target domain without access to source data. To address the problem, we propose a novel method called emotion-aware contrastive adaptation network (ECAN). The core idea is to capture local neighborhood information between samples while considering the global class-level adaptation. Specifically, we propose a nearest neighbor contrastive learning to promote local emotion consistency among features of highly similar samples. Furthermore, relying solely on nearest neighborhoods may lead to ambiguous boundaries between clusters. Thus, we incorporate supervised contrastive learning to encourage greater separation between clusters representing different emotions, thereby facilitating improved class-level adaptation. Extensive experiments indicate that our proposed ECAN significantly outperforms state-of-the-art methods under the source-free cross-corpus SER setting on several speech emotion corpora. 【8】 Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study标题:多语言、完全非自回归ASR与大语言模型融合的综合研究链接:https://arxiv.org/abs/2401.12789作者:W. Ronny Huang,Cyril Allauzen,Tongzhou Chen,Kilol Gupta,Ke Hu,James Qin,Yu Zhang,Yongqiang Wang,Shuo-Yiin Chang,Tara N. Sainath备注:ICASSP 2024摘要:在大型模型时代,解码的自回归性质通常导致延迟成为一个重要的瓶颈。我们提出了一个非自回归LM融合ASR系统,有效地利用加速器硬件的并行化能力。我们的方法结合了通用语音模型(USM)和PaLM 2语言模型在每段评分模式,实现了平均相对WER改善所有语言的10.8%的FLEURS和3.6%的YouTube字幕。此外,我们的综合消融研究分析了关键参数,如LLM大小,上下文长度,词汇量,融合方法。例如,我们探讨了从128 M到340 B的LLM大小参数对ASR性能的影响。这项研究提供了有价值的见解的影响因素,实际的大规模LM融合语音识别系统的有效性。摘要:In the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system that effectively leverages the parallelization capabilities of accelerator hardware. Our approach combines the Universal Speech Model (USM) and the PaLM 2 language model in per-segment scoring mode, achieving an average relative WER improvement across all languages of 10.8% on FLEURS and 3.6% on YouTube captioning. Furthermore, our comprehensive ablation study analyzes key parameters such as LLM size, context length, vocabulary size, fusion methodology. For instance, we explore the impact of LLM size ranging from 128M to 340B parameters on ASR performance. This study provides valuable insights into the factors influencing the effectiveness of practical large-scale LM-fused speech recognition systems. 【9】 MoodLoopGP: Generating Emotion-Conditioned Loop Tablature Music with Multi-Granular Features标题:MoodLoopGP:生成具有多粒度特征的情感条件循环曲谱音乐链接:https://arxiv.org/abs/2401.12656作者:Wenqian Cui,Pedro Sarmento,Mathieu Barthet备注:This preprint is licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0). The Version of Record of this contribution is published in Proceedings of EvoMUSART: International Conference on Computational Intelligence in Music, Sound, Art and Design (Part of EvoStar) 2024摘要:可循环音乐生成系统支持多种应用,但它们通常缺乏可控性和定制能力。我们认为,提高可控性可以丰富这些模型,情感表达是创作者和听众的一个重要方面。因此,基于LooperGP,一个可循环的指谱生成模型,本文探讨了赋予系统对所传达的情感的控制。为了实现这样的条件生成,我们建议在模型训练和推理过程中利用多粒度语义和音乐特征来整合音乐知识。具体来说,我们将歌曲级别的功能(情感标签,节奏和模式)和酒吧级别的功能(音调张力)一起指导情感表达。通过算法和人类的评估,我们证明了这种方法在制作传达两种截然不同的目标情感(幸福和悲伤)的音乐方面的有效性。还进行了消融研究,以澄清我们的方法的结果背后的影响因素。摘要:Loopable music generation systems enable diverse applications, but they often lack controllability and customization capabilities. We argue that enhancing controllability can enrich these models, with emotional expression being a crucial aspect for both creators and listeners. Hence, building upon LooperGP, a loopable tablature generation model, this paper explores endowing systems with control over conveyed emotions. To enable such conditional generation, we propose integrating musical knowledge by utilizing multi-granular semantic and musical features during model training and inference. Specifically, we incorporate song-level features (Emotion Labels, Tempo, and Mode) and bar-level features (Tonal Tension) together to guide emotional expression. Through algorithmic and human evaluations, we demonstrate the approach's effectiveness in producing music conveying two contrasting target emotions, happiness and sadness. An ablation study is also conducted to clarify the contributing factors behind our approach's results.
【10】 EEND-M2F: Masked-attention mask transformers for speaker diarization标题:EEND-M2F:用于扬声器二值化的屏蔽注意掩码转换器链接:https://arxiv.org/abs/2401.12600作者:Marc Härkönen,Samuel J. Broughton,Lahiru Samarakoon备注:14 pages, 2 figures摘要:在本文中,我们明确的图像分割方法和端到端的日记化方法之间的联系。从这些见解,我们提出了一种新的,完全端到端的日记模型,EEND-M2 F,基于Mask 2Former架构。使用Transformer解码器的堆栈并行计算扬声器表示,其中使用来自先前层的预测明确地从交叉关注中屏蔽不相关的帧。EEND-M2 F是轻量级的,高效的,真正的端到端,因为它不需要运行任何额外的日志,说话人验证或分割模型,也不需要运行任何聚类算法。我们的模型在几个公共数据集上实现了最先进的性能,如AMI,AliMeeting和RAMC。最值得注意的是,我们在DIHARD-III上的16.07%的DER是挑战获胜系统的第一个重大改进。摘要:In this paper, we make the explicit connection between image segmentation methods and end-to-end diarization methods. From these insights, we propose a novel, fully end-to-end diarization model, EEND-M2F, based on the Mask2Former architecture. Speaker representations are computed in parallel using a stack of transformer decoders, in which irrelevant frames are explicitly masked from the cross attention using predictions from previous layers. EEND-M2F is lightweight, efficient, and truly end-to-end, as it does not require any additional diarization, speaker verification, or segmentation models to run, nor does it require running any clustering algorithms. Our model achieves state-of-the-art performance on several public datasets, such as AMI, AliMeeting and RAMC. Most notably our DER of 16.07% on DIHARD-III is the first major improvement upon the challenge winning system. 【11】 An Exploratory Study of Multimodal Physiological Data in Jazz Improvisation Using Basic Machine Learning Techniques标题:基于基本机器学习技术的爵士乐即兴演奏多通道生理数据的探索性研究链接:https://arxiv.org/abs/2401.12266作者:Yawen Zhang备注:Master's thesis摘要:我们的研究深入到“即兴音乐数据集”,探索即兴音乐表演过程中生理和心理维度之间的相互交织的关系和相关性。主要目的是确定这些状态之间存在明确的因果关系或相关关系,并理解它们在音乐作品中的表现。这个丰富的数据集提供了一个关于音乐家如何在实时即兴表演场景中协调他们的身体与声音事件的视角,强调了“即兴音乐”的概念。“摘要:Our study delves into the "Embodied Musicking Dataset," exploring the intertwined relationships and correlations between physiological and psychological dimensions during improvisational music performances. The primary objective is to ascertain the presence of a definitive causal or correlational relationship between these states and comprehend their manifestation in musical compositions. This rich dataset provides a perspective on how musicians coordinate their physicality with sonic events in real-time improvisational scenarios, emphasizing the concept of "Embodied Musicking." 机器翻译由腾讯交互翻译提供,仅供参考