今日论文合集:cs.SD语音18篇,eess.AS音频处理22篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
标题:SA-SOT:多说话人ASR的说话人感知串行化输出训练
链接:https://arxiv.org/abs/2403.02010
作者:Zhiyun Fan,Linhao Dong,Jun Zhang,Lu Lu,Zejun Ma
摘要:多说话者自动语音识别在涉及多方交互的场景中起着至关重要的作用,例如会议和对话。由于其固有的复杂性,这项任务越来越受到重视。值得注意的是,序列化输出训练(SOT)由于其简单的架构和出色的性能而在各种方法中脱颖而出。然而,在令牌级SOT(t-SOT)的频繁的扬声器的变化提出了挑战的自回归解码器在有效地利用上下文来预测输出序列。为了解决这个问题,我们引入了一个掩蔽的t-SOT标签,它作为辅助训练损失的基石。此外,我们利用说话人相似度矩阵来改进解码器的自注意机制。这种策略性的调整增强了同一说话人的话语标记内的上下文关系,同时最大限度地减少了不同说话人话语标记之间的交互。我们将我们的方法称为说话者感知SOT(SA-SOT)。在Librispeech数据集上的实验表明,我们的SA-SOT在多人测试集上获得了12.75%到22.03%的相对cpWER降低。此外,通过更广泛的训练,我们的方法实现了令人印象深刻的3.41%的cpWER,在LibrispeechMix数据集上建立了一个新的最先进的结果。
摘要:Multi-talker automatic speech recognition plays a crucial role in scenarios involving multi-party interactions, such as meetings and conversations. Due to its inherent complexity, this task has been receiving increasing attention. Notably, the serialized output training (SOT) stands out among various approaches because of its simplistic architecture and exceptional performance. However, the frequent speaker changes in token-level SOT (t-SOT) present challenges for the autoregressive decoder in effectively utilizing context to predict output sequences. To address this issue, we introduce a masked t-SOT label, which serves as the cornerstone of an auxiliary training loss. Additionally, we utilize a speaker similarity matrix to refine the self-attention mechanism of the decoder. This strategic adjustment enhances contextual relationships within the same speaker's tokens while minimizing interactions between different speakers' tokens. We denote our method as speaker-aware SOT (SA-SOT). Experiments on the Librispeech datasets demonstrate that our SA-SOT obtains a relative cpWER reduction ranging from 12.75% to 22.03% on the multi-talker test sets. Furthermore, with more extensive training, our method achieves an impressive cpWER of 3.41%, establishing a new state-of-the-art result on the LibrispeechMix dataset.


【2】 Fine-Grained Quantitative Emotion Editing for Speech Generation
标题:面向语音生成的细粒度量化情感编辑
链接:https://arxiv.org/abs/2403.02002
作者:Sho Inoue,Kun Zhou,Shuai Wang,Haizhou Li
备注:This paper is submitted to IEEE Signal Processing Letters
摘要:如何定量地控制语音情感的表现力是语音生成中的一个重要挑战。在这项工作中,我们提出了一种新的方法来操纵渲染的情绪语音生成。我们提出了一个分层的情感分布提取器,即分层ED,量化的情绪强度在不同级别的粒度。支持向量机(SVM)被用来排名情感强度,从而产生一个层次化的情感嵌入。分层ED随后被集成到FastSpeech2框架中,指导模型在音素,单词和话语水平上学习情感强度。在合成过程中,用户可以手动编辑生成的语音的情感强度。客观和主观的评价表明,所提出的网络在细粒度的定量情感编辑方面的有效性。
摘要:It remains a significant challenge how to quantitatively control the expressiveness of speech emotion in speech generation. In this work, we present a novel approach for manipulating the rendering of emotions for speech generation. We propose a hierarchical emotion distribution extractor, i.e. Hierarchical ED, that quantifies the intensity of emotions at different levels of granularity. Support vector machines (SVMs) are employed to rank emotion intensity, resulting in a hierarchical emotional embedding. Hierarchical ED is subsequently integrated into the FastSpeech2 framework, guiding the model to learn emotion intensity at phoneme, word, and utterance levels. During synthesis, users can manually edit the emotional intensity of the generated voices. Both objective and subjective evaluations demonstrate the effectiveness of the proposed network in terms of fine-grained quantitative emotion editing.

【3】 A robust audio deepfake detection system via multi-view feature
标题:一种基于多视角特征的鲁棒音频Deepfake检测系统
链接:https://arxiv.org/abs/2403.01960
作者:Yujie Yang,Haochen Qin,Hang Zhou,Chengcheng Wang,Tianyu Guo,Kai Han,Yunhe Wang
备注:5 pages, 2 figures
摘要:随着生成建模技术的进步,合成的人类语音变得越来越难以与真实语音区分,这给音频深度伪造检测(ADD)系统带来了棘手的挑战。在本文中,我们利用音频功能,以提高ADD系统的泛化能力。ADD任务性能的调查进行了广泛的音频功能,包括各种手工制作的功能和学习为基础的功能。实验表明,在大量数据上预训练的基于学习的音频特征在域外场景上比手工制作的特征更好地泛化。随后,我们进一步提高了推广的ADD系统使用建议的多功能的方法,将互补的信息,从不同的观点的功能。在ASV 2019数据上训练的模型在In-the-Wild数据集上实现了24.27%的相等错误率。
摘要:With the advancement of generative modeling techniques, synthetic human speech becomes increasingly indistinguishable from real, and tricky challenges are elicited for the audio deepfake detection (ADD) system. In this paper, we exploit audio features to improve the generalizability of ADD systems. Investigation of the ADD task performance is conducted over a broad range of audio features, including various handcrafted features and learning-based features. Experiments show that learning-based audio features pretrained on a large amount of data generalize better than hand-crafted features on out-of-domain scenarios. Subsequently, we further improve the generalizability of the ADD system using proposed multi-feature approaches to incorporate complimentary information from features of different views. The model trained on ASV2019 data achieves an equal error rate of 24.27\% on the In-the-Wild dataset.

【4】 ConSep: a Noise- and Reverberation-Robust Speech Separation Framework by  Magnitude Conditioning
标题:Consep:一种基于幅度条件的抗噪混响语音分离框架
链接:https://arxiv.org/abs/2403.01792
作者:Kuan-Hsun Ho,Jeih-weih Hung,Berlin Chen
摘要:由于时域方法中使用的细粒度视觉,语音分离最近取得了重大进展。然而,一些研究表明,采用短时傅立叶变换(STFT)进行特征提取在遇到更苛刻的条件(如噪声或混响)时可能是有益的。因此,我们提出了一个幅度条件的时域框架,ConSep,继承的有益特性。实验表明,ConSep促进性能在消声,嘈杂,混响的设置相比,两个著名的方法,SepFormer和Bi-Sep。此外,我们可视化ConSep的组件,以加强优势和一致性的现状,我们已经发现在初步研究。
摘要:Speech separation has recently made significant progress thanks to the fine-grained vision used in time-domain methods. However, several studies have shown that adopting Short-Time Fourier Transform (STFT) for feature extraction could be beneficial when encountering harsher conditions, such as noise or reverberation. Therefore, we propose a magnitude-conditioned time-domain framework, ConSep, to inherit the beneficial characteristics. The experiment shows that ConSep promotes performance in anechoic, noisy, and reverberant settings compared to two celebrated methods, SepFormer and Bi-Sep. Furthermore, we visualize the components of ConSep to strengthen the advantages and cohere with the actualities we have found in preliminary studies.

【5】 What do neural networks listen to? Exploring the crucial bands in Speech  Enhancement using Sinc-convolution
标题:神经网络听什么?正弦卷积在语音增强中关键频段的探索
链接:https://arxiv.org/abs/2403.01785
作者:Kuan-Hsun Ho,Jeih-weih Hung,Berlin Chen
摘要:本研究介绍了一种改进的Sinc-convolution(Sincconv)框架,该框架专为深度网络的编码器组件定制,用于语音增强(SE)。改革后的Sincconv,参数化sinc函数作为带通滤波器的基础上,提供了显着的优势,在训练效率,滤波器的多样性,和可解释性。改革后的Sinc-conv与各种SE模型一起进行评估,展示了其提高SE性能的能力。此外,改进的Sincconv提供了对在SE场景中优先化的特定频率分量的有价值的见解。这开辟了SE研究的新方向,并提高了我们对其运行动态的了解。
摘要:This study introduces a reformed Sinc-convolution (Sincconv) framework tailored for the encoder component of deep networks for speech enhancement (SE). The reformed Sincconv, based on parametrized sinc functions as band-pass filters, offers notable advantages in terms of training efficiency, filter diversity, and interpretability. The reformed Sinc-conv is evaluated in conjunction with various SE models, showcasing its ability to boost SE performance. Furthermore, the reformed Sincconv provides valuable insights into the specific frequency components that are prioritized in an SE scenario. This opens up a new direction of SE research and improving our knowledge of their operating dynamics.

【6】 Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based  Audio-Visual Conformer
标题:基于视听同步器的帧级跨通道注意的健壮唤醒词检测
链接:https://arxiv.org/abs/2403.01700
作者:Haoxu Wang,Ming Cheng,Qiang Fu,Ming Li
备注:Accepted by ICASSP 2024
摘要:近年来,基于神经网络的Wake Word Spotting在干净的音频样本上取得了良好的性能,但在嘈杂的环境中表现不佳。听觉-视觉唤醒词识别技术(AVWWS)由于不受复杂声音环境的影响而受到广泛关注。以往的作品通常使用简单的加法或拼接进行多模态融合。模式间的相关性仍然相对不足的探索。在本文中,我们提出了一种新的模块称为帧级跨模态注意(FLCMA),以提高AVWWS系统的性能。该模块可以通过同步的嘴唇运动和语音信号在帧级帮助建模多模态信息。我们训练基于端到端FLCMA的视听一致性,并通过微调预训练的单模态模型来进一步提高性能。所提出的系统实现了一个新的国家的最先进的结果(4.57% WWS分数)的远场MISP数据集。
摘要:In recent years, neural network-based Wake Word Spotting achieves good performance on clean audio samples but struggles in noisy environments. Audio-Visual Wake Word Spotting (AVWWS) receives lots of attention because visual lip movement information is not affected by complex acoustic scenes. Previous works usually use simple addition or concatenation for multi-modal fusion. The inter-modal correlation remains relatively under-explored. In this paper, we propose a novel module called Frame-Level Cross-Modal Attention (FLCMA) to improve the performance of AVWWS systems. This module can help model multi-modal information at the frame-level through synchronous lip movements and speech signals. We train the end-to-end FLCMA based Audio-Visual Conformer and further improve the performance by fine-tuning pre-trained uni-modal models for the AVWWS task. The proposed system achieves a new state-of-the-art result (4.57% WWS score) on the far-field MISP dataset.

【7】 Brilla AI: AI Contestant for the National Science and Maths Quiz
标题:Brilla AI:美国国家科学和数学竞赛的AI选手
链接:https://arxiv.org/abs/2403.01699
作者:George Boateng,Jonathan Abrefah Mensah,Kevin Takyi Yeboah,William Edor,Andrew Kojo Mensah-Onumah,Naafi Dasana Ibrahim,Nana Sam Yeboah
备注:13 pages. Under review at the 25th International Conference on AI in Education (AIED 2024)
摘要:非洲大陆缺乏足够的合格教师,这妨碍了提供适当的学习支助。人工智能可能会增加数量有限的教师的努力,从而带来更好的学习效果。为此,这项工作描述和评估了NATURAL AI Grand Challenge的第一个关键输出,它为这样的AI提出了一个强大的现实世界基准:“构建一个AI来参加加纳的国家科学和数学测验(NATURAL)比赛并获胜-在比赛的所有回合和阶段中表现得比最好的参赛者更好。在加纳,每年一度的高中生科学和数学竞赛,由两名学生组成的三个团队通过回答生物,化学,物理和数学的问题进行竞争,在5轮中进行5个渐进阶段,直到获胜的团队获得该年度的冠军。在这项工作中,我们构建了Brilla AI,这是一个人工智能参赛者,我们部署它来进行非正式的远程比赛,并参加2023年Numbers总决赛的谜语回合,这是该比赛30年历史上的第一次。Brilla AI目前是一款网络应用程序,可以直播比赛的谜语回合,并运行4个机器学习系统:(1)语音到文本(2)问题提取(3)问题回答和(4)文本到语音,这些系统实时协同工作,快速准确地提供答案,然后用加纳口音说出来。在首次亮相时,我们的人工智能回答了4个谜语中的一个,领先于3个人类参赛队,非正式地排名第二(并列)。这种人工智能的改进和扩展可能被部署为向学生提供科学辅导,并最终使非洲数百万人能够进行一对一的学习互动,使科学教育民主化。
摘要:The African continent lacks enough qualified teachers which hampers the provision of adequate learning support. An AI could potentially augment the efforts of the limited number of teachers, leading to better learning outcomes. Towards that end, this work describes and evaluates the first key output for the NSMQ AI Grand Challenge, which proposes a robust, real-world benchmark for such an AI: "Build an AI to compete live in Ghana's National Science and Maths Quiz (NSMQ) competition and win - performing better than the best contestants in all rounds and stages of the competition". The NSMQ is an annual live science and mathematics competition for senior secondary school students in Ghana in which 3 teams of 2 students compete by answering questions across biology, chemistry, physics, and math in 5 rounds over 5 progressive stages until a winning team is crowned for that year. In this work, we built Brilla AI, an AI contestant that we deployed to unofficially compete remotely and live in the Riddles round of the 2023 NSMQ Grand Finale, the first of its kind in the 30-year history of the competition. Brilla AI is currently available as a web app that livestreams the Riddles round of the contest, and runs 4 machine learning systems: (1) speech to text (2) question extraction (3) question answering and (4) text to speech that work together in real-time to quickly and accurately provide an answer, and then say it with a Ghanaian accent. In its debut, our AI answered one of the 4 riddles ahead of the 3 human contesting teams, unofficially placing second (tied). Improvements and extensions of this AI could potentially be deployed to offer science tutoring to students and eventually enable millions across Africa to have one-on-one learning interactions, democratizing science education.

【8】 Enhancing Audio Generation Diversity with Visual Information
标题:利用视觉信息增强音频生成多样性
链接:https://arxiv.org/abs/2403.01278
作者:Zeyu Xie,Baihan Li,Xuenan Xu,Mengyue Wu,Kai Yu
摘要:近年来,音频和声音生成获得了极大的关注,主要集中在提高生成的音频的质量上。然而,关于增强所生成的音频的多样性的研究有限,特别是当涉及到特定类别内的音频生成时。当前模型倾向于在类别内产生同质音频样本。这项工作的目的是解决这个限制,通过改善生成的音频与视觉信息的多样性。我们提出了一种基于聚类的方法,利用视觉信息来指导模型在每个类别中生成不同的音频内容。七个类别的结果表明,额外的视觉输入可以大大提高音频生成的多样性。音频样本可在https://zeyuxie29.github.io/DiverseAudioGeneration上获得。
摘要:Audio and sound generation has garnered significant attention in recent years, with a primary focus on improving the quality of generated audios. However, there has been limited research on enhancing the diversity of generated audio, particularly when it comes to audio generation within specific categories. Current models tend to produce homogeneous audio samples within a category. This work aims to address this limitation by improving the diversity of generated audio with visual information. We propose a clustering-based method, leveraging visual information to guide the model in generating distinct audio content within each category. Results on seven categories indicate that extra visual input can largely enhance audio generation diversity. Audio samples are available at https://zeyuxie29.github.io/DiverseAudioGeneration.


【9】 Automatic Speech Recognition using Advanced Deep Learning Approaches: A  survey
标题:使用高级深度学习方法的自动语音识别:综述
链接:https://arxiv.org/abs/2403.01255
作者:Hamza Kheddar,Mustapha Hemis,Yassine Himeur
摘要:深度学习(DL)的最新进展对自动语音识别(ASR)提出了重大挑战。ASR依赖于广泛的训练数据集,包括机密数据集,并需要大量的计算和存储资源。启用自适应系统可提高动态环境中的ASR性能。DL技术假设训练和测试数据来自同一个域,但这并不总是正确的。深度迁移学习(DTL)、联邦学习(FL)和强化学习(RL)等高级DL技术可以解决这些问题。DTL允许使用小而相关的数据集进行高性能模型,FL允许在不拥有数据集的情况下对机密数据进行训练,RL优化动态环境中的决策,降低计算成本。该调查对基于DTL、FL和RL的ASR框架进行了全面的回顾,旨在深入了解最新发展,并帮助研究人员和专业人士了解当前的挑战。此外,Transformers,这是先进的DL技术,大量使用在拟议的ASR框架中,被认为是在这个调查中,他们的能力,以捕捉广泛的依赖关系的输入ASR序列。本文首先介绍了DTL、FL、RL和Transformers的背景,然后采用了一个精心设计的分类法来概述最先进的方法。随后,进行了批判性分析,以确定每个框架的长处和短处。此外,还进行了比较研究,以突出现有的挑战,为未来的研究机会铺平道路。
摘要:Recent advancements in deep learning (DL) have posed a significant challenge for automatic speech recognition (ASR). ASR relies on extensive training datasets, including confidential ones, and demands substantial computational and storage resources. Enabling adaptive systems improves ASR performance in dynamic environments. DL techniques assume training and testing data originate from the same domain, which is not always true. Advanced DL techniques like deep transfer learning (DTL), federated learning (FL), and reinforcement learning (RL) address these issues. DTL allows high-performance models using small yet related datasets, FL enables training on confidential data without dataset possession, and RL optimizes decision-making in dynamic environments, reducing computation costs. This survey offers a comprehensive review of DTL, FL, and RL-based ASR frameworks, aiming to provide insights into the latest developments and aid researchers and professionals in understanding the current challenges. Additionally, transformers, which are advanced DL techniques heavily used in proposed ASR frameworks, are considered in this survey for their ability to capture extensive dependencies in the input ASR sequence. The paper starts by presenting the background of DTL, FL, RL, and Transformers and then adopts a well-designed taxonomy to outline the state-of-the-art approaches. Subsequently, a critical analysis is conducted to identify the strengths and weaknesses of each framework. Additionally, a comparative study is presented to highlight the existing challenges, paving the way for future research opportunities.

【10】 MPIPN: A Multi Physics-Informed PointNet for solving parametric  acoustic-structure systems
标题:MPIPN:求解参数声学结构系统的多物理信息点网络
链接:https://arxiv.org/abs/2403.01132
作者:Chu Wang,Jinhong Wu,Yanzhi Wang,Zhijian Zha,Qi Zhou
备注:The number of figures is 16. The number of tables is 5. The number of words is 9717
摘要:机器学习用于求解由一般非线性偏微分方程(PDE)控制的物理系统。然而,复杂的多物理场系统,如声-结构耦合,往往是由一系列的偏微分方程,其中包括可变的物理量,这被称为参数系统。对于包含显式和隐式量的偏微分方程所控制的参数系统,缺乏求解策略。本文提出了一种基于深度学习的多物理信息点网络(MPIPN),用于求解参数化声学-结构系统。首先,MPIPN诱导一个增强的点云架构,包括明确的物理量和几何特征的计算域。然后,MPIPN提取局部和全局特征的重建点云的参数系统的解决标准的一部分,分别。此外,隐式物理量作为求解准则的另一部分通过编码技术嵌入。最后,所有的解决标准,表征参数系统的合并,形成独特的序列作为输入的MPIPN,其输出是系统的解决方案。所提出的框架是由相应的计算域的自适应物理信息损失函数训练的。该框架被推广到处理新的参数条件的系统。通过求解Helmholtz方程控制的稳态参数声固耦合系统,验证了MPIPN的有效性。消融实验已经实施,以证明与少数监督数据的物理通知的影响的功效。所提出的方法产生合理的精度在所有的计算域恒定的参数条件下和可变的组合的参数条件下的声学结构系统。
摘要:Machine learning is employed for solving physical systems governed by general nonlinear partial differential equations (PDEs). However, complex multi-physics systems such as acoustic-structure coupling are often described by a series of PDEs that incorporate variable physical quantities, which are referred to as parametric systems. There are lack of strategies for solving parametric systems governed by PDEs that involve explicit and implicit quantities. In this paper, a deep learning-based Multi Physics-Informed PointNet (MPIPN) is proposed for solving parametric acoustic-structure systems. First, the MPIPN induces an enhanced point-cloud architecture that encompasses explicit physical quantities and geometric features of computational domains. Then, the MPIPN extracts local and global features of the reconstructed point-cloud as parts of solving criteria of parametric systems, respectively. Besides, implicit physical quantities are embedded by encoding techniques as another part of solving criteria. Finally, all solving criteria that characterize parametric systems are amalgamated to form distinctive sequences as the input of the MPIPN, whose outputs are solutions of systems. The proposed framework is trained by adaptive physics-informed loss functions for corresponding computational domains. The framework is generalized to deal with new parametric conditions of systems. The effectiveness of the MPIPN is validated by applying it to solve steady parametric acoustic-structure coupling systems governed by the Helmholtz equations. An ablation experiment has been implemented to demonstrate the efficacy of physics-informed impact with a minority of supervised data. The proposed method yields reasonable precision across all computational domains under constant parametric conditions and changeable combinations of parametric conditions for acoustic-structure systems.

【11】 Towards Accurate Lip-to-Speech Synthesis in-the-Wild
标题:在野外实现精确的唇语合成
链接:https://arxiv.org/abs/2403.01087
作者:Sindhu Hegde,Rudrabha Mukhopadhyay,C. V. Jawahar,Vinay Namboodiri
备注:None
摘要:在本文中,我们介绍了一种新的方法来解决的任务合成语音从无声的视频中的任何在野生扬声器完全基于嘴唇运动。传统的从嘴唇视频直接生成语音的方法面临着无法单独从语音中学习鲁棒的语言模型的挑战,导致结果不令人满意。为了克服这个问题,我们建议使用最先进的唇到文本网络将噪声文本监督纳入我们的模型中。噪声文本是使用预先训练的唇到文本模型生成的,使我们的方法在推理过程中无需文本注释即可工作。我们设计了一个可视化的文本到语音的网络,利用视觉流,以产生准确的语音,这是同步的无声输入视频。我们进行了广泛的实验和消融研究,证明了我们的方法在各种基准数据集上优于当前最先进的方法。此外,我们证明了我们的方法在辅助技术中的一个重要的实际应用,通过为ALS患者谁失去了声音,但可以使嘴部运动产生语音。可以在\url{http://cvit.iiit.ac.in/research/projects/cvit-projects/ms-l2s-itw}找到我们的演示视频、代码和其他详细信息。
摘要:In this paper, we introduce a novel approach to address the task of synthesizing speech from silent videos of any in-the-wild speaker solely based on lip movements. The traditional approach of directly generating speech from lip videos faces the challenge of not being able to learn a robust language model from speech alone, resulting in unsatisfactory outcomes. To overcome this issue, we propose incorporating noisy text supervision using a state-of-the-art lip-to-text network that instills language information into our model. The noisy text is generated using a pre-trained lip-to-text model, enabling our approach to work without text annotations during inference. We design a visual text-to-speech network that utilizes the visual stream to generate accurate speech, which is in-sync with the silent input video. We perform extensive experiments and ablation studies, demonstrating our approach's superiority over the current state-of-the-art methods on various benchmark datasets. Further, we demonstrate an essential practical application of our method in assistive technology by generating speech for an ALS patient who has lost the voice but can make mouth movements. Our demo video, code, and additional details can be found at \url{http://cvit.iiit.ac.in/research/projects/cvit-projects/ms-l2s-itw}.

【12】 Scaling Up Adaptive Filter Optimizers
标题:放大自适应滤波优化器
链接:https://arxiv.org/abs/2403.00977
作者:Jonah Casebeer,Nicholas J. Bryan,Paris Smaragdis
摘要:我们介绍了一种新的在线自适应滤波方法称为监督多步自适应滤波器(SMS-AF)。我们的方法使用神经网络来控制或优化线性多延迟或多通道频域滤波器,并且可以以增加计算为代价灵活地扩展性能-这是AF文献中很少涉及的属性,但对许多应用至关重要。为了做到这一点,我们扩展了最近的工作与一组改进,包括功能修剪,监督损失,和多个优化步骤,每个时间框架。这些改进以一种内聚的方式工作,以解锁缩放。此外,我们展示了我们的方法如何与卡尔曼滤波和元自适应滤波相关,使其无缝适用于各种AF任务。我们评估我们的方法声学回声消除(AEC)和多通道语音增强任务,并与标准的合成和真实世界的数据集上的几个基线进行比较。结果表明,我们的方法的性能尺度与推理成本和模型容量,产生多dB的性能增益为两个任务,并在一个单一的CPU核心上是实时的能力。
摘要:We introduce a new online adaptive filtering method called supervised multi-step adaptive filters (SMS-AF). Our method uses neural networks to control or optimize linear multi-delay or multi-channel frequency-domain filters and can flexibly scale-up performance at the cost of increased compute -- a property rarely addressed in the AF literature, but critical for many applications. To do so, we extend recent work with a set of improvements including feature pruning, a supervised loss, and multiple optimization steps per time-frame. These improvements work in a cohesive manner to unlock scaling. Furthermore, we show how our method relates to Kalman filtering and meta-adaptive filtering, making it seamlessly applicable to a diverse set of AF tasks. We evaluate our method on acoustic echo cancellation (AEC) and multi-channel speech enhancement tasks and compare against several baselines on standard synthetic and real-world datasets. Results show our method performance scales with inference cost and model capacity, yields multi-dB performance gains for both tasks, and is real-time capable on a single CPU core.


【13】 Structuring Concept Space with the Musical Circle of Fifths by Utilizing  Music Grammar Based Activations
标题:利用基于音乐语法的激活构建与五度音乐圈的概念空间
链接:https://arxiv.org/abs/2403.00790
作者:Tofara Moyo
备注:3 pages
摘要:在本文中,我们探索了离散神经网络(如尖峰网络)的结构与钢琴作品之间的有趣相似性。虽然两者都涉及顺序或并行激活的节点或音符,但后者受益于丰富的音乐理论来指导有意义的组合。我们提出了一种新的方法,利用音乐语法来调节尖峰神经网络中的激活,允许将符号表示为吸引子。通过应用音乐理论中的和弦进行规则,我们展示了某些激活是如何自然地跟随其他激活的,类似于吸引力的概念。此外,我们引入了调制键的概念,以导航网络内的不同流域的吸引力。最终,我们展示了我们模型中的概念图是由五度音乐圈构成的,突出了在深度学习算法中利用音乐理论原理的潜力。
摘要:In this paper, we explore the intriguing similarities between the structure of a discrete neural network, such as a spiking network, and the composition of a piano piece. While both involve nodes or notes that are activated sequentially or in parallel, the latter benefits from the rich body of music theory to guide meaningful combinations. We propose a novel approach that leverages musical grammar to regulate activations in a spiking neural network, allowing for the representation of symbols as attractors. By applying rules for chord progressions from music theory, we demonstrate how certain activations naturally follow others, akin to the concept of attraction. Furthermore, we introduce the concept of modulating keys to navigate different basins of attraction within the network. Ultimately, we show that the map of concepts in our model is structured by the musical circle of fifths, highlighting the potential for leveraging music theory principles in deep learning algorithms.

【14】 Speech emotion recognition from voice messages recorded in the wild
标题:野外录音语音信息的语音情感识别
链接:https://arxiv.org/abs/2403.02167
作者:Lucía Gómez-Zaragozá,Óscar Valls,Rocío del Amor,María José Castro-Bleda,Valery Naranjo,Mariano Alcañiz Raya,Javier Marín-Morales
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
摘要:用于语音情感识别(SER)的情感数据集通常包含行为或引发的语音,限制了其在现实世界场景中的适用性。在这项工作中,我们使用了情绪语音信息(EMOVOME)数据库,包括来自100名西班牙语使用者在消息应用程序上的对话的自发语音信息,由专家和非专家注释者以连续和离散的情绪进行标记。我们使用eGeMAPS功能,基于transformer的模型及其组合创建了与说话者无关的SER模型。我们将结果与参考数据库进行了比较,并分析了注释者和性别公平性的影响。预训练的Unispeech-L模型及其与eGeMAPS的组合取得了最高的结果,3类效价和唤醒预测的未加权准确率(UA)分别为61.64%和55.57%,比基线模型提高了10%。对于情绪类别,获得42.58%的UA。EMOVOME的性能低于作用的RAVDESS数据库。诱发IEMOCAP数据库也优于EMOVOME的情绪类别的预测,而在效价和唤醒得到类似的结果。此外,EMOVOME结果因注释者标签而异,当结合专家和非专家注释时,显示出更好的结果和更好的公平性。这项研究显着有助于评估SER模型在现实生活中的情况下,推进应用程序的发展,分析自发的语音信息。
摘要:Emotion datasets used for Speech Emotion Recognition (SER) often contain acted or elicited speech, limiting their applicability in real-world scenarios. In this work, we used the Emotional Voice Messages (EMOVOME) database, including spontaneous voice messages from conversations of 100 Spanish speakers on a messaging app, labeled in continuous and discrete emotions by expert and non-expert annotators. We created speaker-independent SER models using the eGeMAPS features, transformer-based models and their combination. We compared the results with reference databases and analyzed the influence of annotators and gender fairness. The pre-trained Unispeech-L model and its combination with eGeMAPS achieved the highest results, with 61.64% and 55.57% Unweighted Accuracy (UA) for 3-class valence and arousal prediction respectively, a 10% improvement over baseline models. For the emotion categories, 42.58% UA was obtained. EMOVOME performed lower than the acted RAVDESS database. The elicited IEMOCAP database also outperformed EMOVOME in the prediction of emotion categories, while similar results were obtained in valence and arousal. Additionally, EMOVOME outcomes varied with annotator labels, showing superior results and better fairness when combining expert and non-expert annotations. This study significantly contributes to the evaluation of SER models in real-life situations, advancing in the development of applications for analyzing spontaneous voice messages.


【15】 6DoF SELD: Sound Event Localization and Detection Using Microphones and  Motion Tracking Sensors on self-motioning human
标题:6DoF SELD:利用麦克风和运动跟踪传感器对自主运动的人进行声音事件定位和检测
链接:https://arxiv.org/abs/2403.01670
作者:Masahiro Yasuda,Shoichiro Saito,Akira Nakayama,Noboru Harada
备注:ICASSP2024 accepted
摘要:我们的目标是执行声音事件定位和检测(SELD)使用可穿戴设备的移动的人,如行人。传统的SELD任务只处理位于静态位置的麦克风阵列。然而,对于可穿戴麦克风阵列,应考虑具有三个旋转和三个平移自由度(6DoF)的自运动。仅使用固定位置的麦克风阵列用数据集训练的系统将无法适应与自运动相关联的声音事件的快速相对运动,从而导致SELD性能的下降。为了解决这个问题,我们为可穿戴系统设计了6DoF SELD数据集,这是第一个考虑麦克风自运动的SELD数据集。此外,我们提出了一个多模态SELD系统,联合利用音频和运动跟踪传感器信号。这些传感器信号被期望帮助系统基于当前的自运动状态找到用于SELD的有用的声学线索。在我们的数据集上的实验结果表明,该方法有效地提高了SELD的性能与机制,以提取由传感器信号条件下的声学特征。
摘要:We aim to perform sound event localization and detection (SELD) using wearable equipment for a moving human, such as a pedestrian. Conventional SELD tasks have dealt only with microphone arrays located in static positions. However, self-motion with three rotational and three translational degrees of freedom (6DoF) shall be considered for wearable microphone arrays. A system trained only with a dataset using microphone arrays in a fixed position would be unable to adapt to the fast relative motion of sound events associated with self-motion, resulting in the degradation of SELD performance. To address this, we designed 6DoF SELD Dataset for wearable systems, the first SELD dataset considering the self-motion of microphones. Furthermore, we proposed a multi-modal SELD system that jointly utilizes audio and motion tracking sensor signals. These sensor signals are expected to help the system find useful acoustic cues for SELD on the basis of the current self-motion state. Experimental results on our dataset show that the proposed method effectively improves SELD performance with a mechanism to extract acoustic features conditioned by sensor signals.


【16】 PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice  Conversion
标题:PAVITS:探索韵律感知的VITS用于端到端情感语音转换
链接:https://arxiv.org/abs/2403.01494
作者:Tianhua Qi,Wenming Zheng,Cheng Lu,Yuan Zong,Hailun Lian
备注:Accepted to ICASSP2024
摘要:在本文中,我们提出了韵律感知的语音转换(PAVITS)的情感语音转换(EVC),旨在实现EVC的两个主要目标:高内容的自然性和高情感的自然性,这是至关重要的,以满足人类的感知需求。为了提高转换音频的内容自然度,我们开发了一种端到端EVC架构,其灵感来自于VITS的高音频质量。通过无缝集成声学转换器和声码器,我们有效地解决了情感韵律训练和运行时转换之间的不匹配,这是在现有的EVC模型普遍存在的共同问题。为了进一步增强语音的情感自然度,我们引入了一个情感描述子来描述不同语音情感的细微韵律变化。此外,我们提出了一个韵律预测器,它预测的基础上提供的情感标签的文本的韵律特征。值得注意的是,我们引入了韵律对齐损失来建立两种不同模态的潜在韵律特征之间的联系,以确保有效的训练。实验结果表明,PAVITS的性能优于最先进的EVC方法。语音样本可在https://jeremychee4.github.io/pavits4EVC/上获得。
摘要:In this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for meeting the demands of human perception. To improve the content naturalness of converted audio, we have developed an end-to-end EVC architecture inspired by the high audio quality of VITS. By seamlessly integrating an acoustic converter and vocoder, we effectively address the common issue of mismatch between emotional prosody training and run-time conversion that is prevalent in existing EVC models. To further enhance the emotional naturalness, we introduce an emotion descriptor to model the subtle prosody variations of different speech emotions. Additionally, we propose a prosody predictor, which predicts prosody features from text based on the provided emotion label. Notably, we introduce a prosody alignment loss to establish a connection between latent prosody features from two distinct modalities, ensuring effective training. Experimental results show that the performance of PAVITS is superior to the state-of-the-art EVC methods. Speech Samples are available at https://jeremychee4.github.io/pavits4EVC/ .


【17】 SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in  Speech
标题:SEGAA:预测言语中年龄、性别和情感的统一方法
链接:https://arxiv.org/abs/2403.00887
作者:Aron R,Indra Sigicharla,Chirag Periwal,Mohanaprasad K,Nithya Darisini P S,Sourabh Tiwari,Shivani Arora
摘要:人类声音的解释在各种应用中都很重要。这项研究冒险预测年龄,性别和情绪的声音线索,一个领域具有广泛的应用。语音分析技术的进步跨越了各个领域,从改善客户互动到增强医疗保健和零售体验。辨别情绪有助于心理健康,而年龄和性别检测在各种情况下都至关重要。探索这些预测的深度学习模型涉及比较本文中强调的单输出,多输出和顺序模型。寻找合适的数据带来了挑战,导致CREMA-D和EMO-DB数据集的合并。先前的工作表明,在个别预测的承诺,但有限的研究同时考虑所有三个变量。本文指出了个体模型方法的缺陷,并倡导我们新的多输出学习架构基于语音的情感性别和年龄分析(SEGAA)模型。实验表明,多输出模型对单个模型进行了改进,有效地捕捉了变量和语音输入之间的复杂关系,同时提高了运行时间。
摘要:The interpretation of human voices holds importance across various applications. This study ventures into predicting age, gender, and emotion from vocal cues, a field with vast applications. Voice analysis tech advancements span domains, from improving customer interactions to enhancing healthcare and retail experiences. Discerning emotions aids mental health, while age and gender detection are vital in various contexts. Exploring deep learning models for these predictions involves comparing single, multi-output, and sequential models highlighted in this paper. Sourcing suitable data posed challenges, resulting in the amalgamation of the CREMA-D and EMO-DB datasets. Prior work showed promise in individual predictions, but limited research considered all three variables simultaneously. This paper identifies flaws in an individual model approach and advocates for our novel multi-output learning architecture Speech-based Emotion Gender and Age Analysis (SEGAA) model. The experiments suggest that Multi-output models perform comparably to individual models, efficiently capturing the intricate relationships between variables and speech inputs, all while achieving improved runtime.

【18】 Speaker-Independent Dysarthria Severity Classification using  Self-Supervised Transformers and Multi-Task Learning
标题:使用自我监督转换器和多任务学习的非说话者依赖性构音障碍严重程度分类
链接:https://arxiv.org/abs/2403.00854
作者:Lauren Stumpf,Balasundaram Kadirvelu,Sigourney Waibel,A. Aldo Faisal
备注:17 pages, 2 tables, 4 main figures, 2 supplemental figures, prepared for journal submission
摘要:构音障碍是一种由于神经系统疾病导致的言语肌肉控制受损而导致的疾病,严重影响了患者的沟通和生活质量。这种情况的复杂性,人为评分和各种各样的表现使其评估和管理具有挑战性。这项研究提出了一个基于transformer的框架,自动评估构音障碍的严重程度,从原始语音数据。它可以提供一个客观的,可重复的,可访问的,标准化的和具有成本效益的,与传统的方法相比,需要人类专家评估。我们开发了一个Transformer框架,称为Speaker-Agnostic Latent Regularisation(SALR),将多任务学习目标和对比学习用于独立于说话者的多类构音障碍严重程度分类。多任务框架的设计,以减少对特定于扬声器的特征的依赖,并解决构音障碍的语音的内在类内变异。我们使用leave-one-speaker-out交叉验证对Universal Access Speech数据集进行了评估,我们的模型表现出优于传统机器学习方法的性能,准确率为70.48美元,F1得分为59.23美元。我们的SALR模型也超过了之前使用支持向量机的基于AI的分类基准,达到了16.58美元。我们通过可视化潜在空间来打开模型的黑盒子,在那里我们可以观察模型如何大大减少说话者特定的线索并放大特定于任务的线索,从而显示其鲁棒性。总之,SALR使用生成式AI在说话者独立的多类构音障碍严重程度分类中建立了一个新的基准。我们的研究结果对自动构音障碍严重程度评估的更广泛临床应用的潜在影响。
摘要:Dysarthria, a condition resulting from impaired control of the speech muscles due to neurological disorders, significantly impacts the communication and quality of life of patients. The condition's complexity, human scoring and varied presentations make its assessment and management challenging. This study presents a transformer-based framework for automatically assessing dysarthria severity from raw speech data. It can offer an objective, repeatable, accessible, standardised and cost-effective and compared to traditional methods requiring human expert assessors. We develop a transformer framework, called Speaker-Agnostic Latent Regularisation (SALR), incorporating a multi-task learning objective and contrastive learning for speaker-independent multi-class dysarthria severity classification. The multi-task framework is designed to reduce reliance on speaker-specific characteristics and address the intrinsic intra-class variability of dysarthric speech. We evaluated on the Universal Access Speech dataset using leave-one-speaker-out cross-validation, our model demonstrated superior performance over traditional machine learning approaches, with an accuracy of $70.48\%$ and an F1 score of $59.23\%$. Our SALR model also exceeded the previous benchmark for AI-based classification, which used support vector machines, by $16.58\%$. We open the black box of our model by visualising the latent space where we can observe how the model substantially reduces speaker-specific cues and amplifies task-specific ones, thereby showing its robustness. In conclusion, SALR establishes a new benchmark in speaker-independent multi-class dysarthria severity classification using generative AI. The potential implications of our findings for broader clinical applications in automated dysarthria severity assessments.


eess.AS音频处理
【1】 PixIT: Joint Training of Speaker Diarization and Speech Separation from  Real-world Multi-speaker Recordings
标题:PixIT:真实多人录音中说话人对分和语音分离的联合训练
链接:https://arxiv.org/abs/2403.02288
作者:Joonas Kalda,Clément Pagés,Ricard Marxer,Tanel Alumäe,Hervé Bredin
备注:submitted to Speaker Odyssey 2024
摘要:监督语音分离(SSep)系统的一个主要缺点是它们依赖于合成数据,导致真实世界的泛化能力差。混合不变训练(MixIT)是作为一种无监督的替代方案提出的,它使用真实的录音,但要克服过度分离和适应长格式音频。我们介绍PixIT,这是一种联合方法,它结合了用于扬声器diarization(SD)的置换不变训练(PIT)和用于SSep的MixIT。通过需要SD标签的小的额外要求,它解决了过度分离的问题,并允许利用基于聚类的神经SD的现有工作来拼接本地分离的源。我们通过应用自动语音识别(ASR)系统来测量分离的源的质量。PixIT提升了两个会议语料库中各种ASR系统的性能,无论是在说话者属性还是基于话语的单词错误率方面,都不需要任何微调。
摘要:A major drawback of supervised speech separation (SSep) systems is their reliance on synthetic data, leading to poor real-world generalization. Mixture invariant training (MixIT) was proposed as an unsupervised alternative that uses real recordings, yet struggles with overseparation and adapting to long-form audio. We introduce PixIT, a joint approach that combines permutation invariant training (PIT) for speaker diarization (SD) and MixIT for SSep. With a small extra requirement of needing SD labels, it solves the problem of overseparation and allows stitching local separated sources leveraging existing work on clustering-based neural SD. We measure the quality of the separated sources via applying automatic speech recognition (ASR) systems to them. PixIT boosts the performance of various ASR systems across two meeting corpora both in terms of the speaker-attributed and utterance-based word error rates while not requiring any fine-tuning.


【2】 Speech emotion recognition from voice messages recorded in the wild
标题:野外录音语音信息的语音情感识别
链接:https://arxiv.org/abs/2403.02167
作者:Lucía Gómez-Zaragozá,Óscar Valls,Rocío del Amor,María José Castro-Bleda,Valery Naranjo,Mariano Alcañiz Raya,Javier Marín-Morales
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
摘要:用于语音情感识别(SER)的情感数据集通常包含行为或引发的语音,限制了其在现实世界场景中的适用性。在这项工作中,我们使用了情绪语音信息(EMOVOME)数据库,包括来自100名西班牙语使用者在消息应用程序上的对话的自发语音信息,由专家和非专家注释者以连续和离散的情绪进行标记。我们使用eGeMAPS功能,基于transformer的模型及其组合创建了与说话者无关的SER模型。我们将结果与参考数据库进行了比较,并分析了注释者和性别公平性的影响。预训练的Unispeech-L模型及其与eGeMAPS的组合取得了最高的结果,3类效价和唤醒预测的未加权准确率(UA)分别为61.64%和55.57%,比基线模型提高了10%。对于情绪类别,获得42.58%的UA。EMOVOME的性能低于作用的RAVDESS数据库。诱发IEMOCAP数据库也优于EMOVOME的情绪类别的预测,而在效价和唤醒得到类似的结果。此外,EMOVOME结果因注释者标签而异,当结合专家和非专家注释时,显示出更好的结果和更好的公平性。这项研究显着有助于评估SER模型在现实生活中的情况下,推进应用程序的发展,分析自发的语音信息。
摘要:Emotion datasets used for Speech Emotion Recognition (SER) often contain acted or elicited speech, limiting their applicability in real-world scenarios. In this work, we used the Emotional Voice Messages (EMOVOME) database, including spontaneous voice messages from conversations of 100 Spanish speakers on a messaging app, labeled in continuous and discrete emotions by expert and non-expert annotators. We created speaker-independent SER models using the eGeMAPS features, transformer-based models and their combination. We compared the results with reference databases and analyzed the influence of annotators and gender fairness. The pre-trained Unispeech-L model and its combination with eGeMAPS achieved the highest results, with 61.64% and 55.57% Unweighted Accuracy (UA) for 3-class valence and arousal prediction respectively, a 10% improvement over baseline models. For the emotion categories, 42.58% UA was obtained. EMOVOME performed lower than the acted RAVDESS database. The elicited IEMOCAP database also outperformed EMOVOME in the prediction of emotion categories, while similar results were obtained in valence and arousal. Additionally, EMOVOME outcomes varied with annotator labels, showing superior results and better fairness when combining expert and non-expert annotations. This study significantly contributes to the evaluation of SER models in real-life situations, advancing in the development of applications for analyzing spontaneous voice messages.

【3】 6DoF SELD: Sound Event Localization and Detection Using Microphones and  Motion Tracking Sensors on self-motioning human
标题:6DoF SELD:使用麦克风和运动跟踪传感器对自运动人体进行声音事件定位和检测
链接:https://arxiv.org/abs/2403.01670
作者:Masahiro Yasuda,Shoichiro Saito,Akira Nakayama,Noboru Harada
备注:ICASSP2024 accepted
摘要:我们的目标是执行声音事件定位和检测(SELD)使用可穿戴设备的移动的人,如行人。传统的SELD任务只处理位于静态位置的麦克风阵列。然而,对于可穿戴麦克风阵列,应考虑具有三个旋转和三个平移自由度(6DoF)的自运动。仅使用固定位置的麦克风阵列用数据集训练的系统将无法适应与自运动相关联的声音事件的快速相对运动,从而导致SELD性能的下降。为了解决这个问题,我们为可穿戴系统设计了6DoF SELD数据集,这是第一个考虑麦克风自运动的SELD数据集。此外,我们提出了一个多模态SELD系统,联合利用音频和运动跟踪传感器信号。这些传感器信号被期望帮助系统基于当前的自运动状态找到用于SELD的有用的声学线索。在我们的数据集上的实验结果表明,该方法有效地提高了SELD的性能与机制,以提取由传感器信号条件下的声学特征。
摘要:We aim to perform sound event localization and detection (SELD) using wearable equipment for a moving human, such as a pedestrian. Conventional SELD tasks have dealt only with microphone arrays located in static positions. However, self-motion with three rotational and three translational degrees of freedom (6DoF) shall be considered for wearable microphone arrays. A system trained only with a dataset using microphone arrays in a fixed position would be unable to adapt to the fast relative motion of sound events associated with self-motion, resulting in the degradation of SELD performance. To address this, we designed 6DoF SELD Dataset for wearable systems, the first SELD dataset considering the self-motion of microphones. Furthermore, we proposed a multi-modal SELD system that jointly utilizes audio and motion tracking sensor signals. These sensor signals are expected to help the system find useful acoustic cues for SELD on the basis of the current self-motion state. Experimental results on our dataset show that the proposed method effectively improves SELD performance with a mechanism to extract acoustic features conditioned by sensor signals.

【4】 PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice  Conversion
标题:PAVITS:探索端到端情感语音转换的韵律感知VITS
链接:https://arxiv.org/abs/2403.01494
作者:Tianhua Qi,Wenming Zheng,Cheng Lu,Yuan Zong,Hailun Lian
备注:Accepted to ICASSP2024
摘要:在本文中,我们提出了韵律感知的语音转换(PAVITS)的情感语音转换(EVC),旨在实现EVC的两个主要目标:高内容的自然性和高情感的自然性,这是至关重要的,以满足人类的感知需求。为了提高转换音频的内容自然度,我们开发了一种端到端EVC架构,其灵感来自于VITS的高音频质量。通过无缝集成声学转换器和声码器,我们有效地解决了情感韵律训练和运行时转换之间的不匹配,这是在现有的EVC模型普遍存在的共同问题。为了进一步增强语音的情感自然度,我们引入了一个情感描述子来描述不同语音情感的细微韵律变化。此外,我们提出了一个韵律预测器,它预测的基础上提供的情感标签的文本的韵律特征。值得注意的是,我们引入了韵律对齐损失来建立两种不同模态的潜在韵律特征之间的联系,以确保有效的训练。实验结果表明,PAVITS的性能优于最先进的EVC方法。语音样本可在https://jeremychee4.github.io/pavits4EVC/上获得。
摘要:In this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for meeting the demands of human perception. To improve the content naturalness of converted audio, we have developed an end-to-end EVC architecture inspired by the high audio quality of VITS. By seamlessly integrating an acoustic converter and vocoder, we effectively address the common issue of mismatch between emotional prosody training and run-time conversion that is prevalent in existing EVC models. To further enhance the emotional naturalness, we introduce an emotion descriptor to model the subtle prosody variations of different speech emotions. Additionally, we propose a prosody predictor, which predicts prosody features from text based on the provided emotion label. Notably, we introduce a prosody alignment loss to establish a connection between latent prosody features from two distinct modalities, ensuring effective training. Experimental results show that the performance of PAVITS is superior to the state-of-the-art EVC methods. Speech Samples are available at https://jeremychee4.github.io/pavits4EVC/ .

【5】 A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech  Enhancement标题:深入了解用于设备上单通道语音增强的Wav2Vec2嵌入
链接:https://arxiv.org/abs/2403.01369
作者:Ravi Shankar,Ke Tan,Buye Xu,Anurag Kumar
备注:8 pages; Shorter form accepted in ICASSP 2024
摘要:已经发现自监督学习模型对于某些语音任务非常有效,例如自动语音识别,说话人识别,关键字定位等。虽然这些特征在语音识别和相关任务中是非常有用的,但它们在语音增强系统中的实用性还没有被牢固地建立,并且可能没有被正确地理解。在本文中,我们调查的SSL表示单通道语音增强在具有挑战性的条件下,发现它们添加很少的增强任务的价值。我们的约束是围绕设备上的实时语音增强设计的-模型是因果关系的,计算占用空间小。此外,我们专注于低信噪比条件下,这些模型的斗争,以提供良好的增强。为了系统地研究SSL表示如何影响这种增强模型的性能,我们提出了各种技术来利用这些嵌入,其中包括不同形式的知识提取和预训练。
摘要:Self-supervised learned models have been found to be very effective for certain speech tasks such as automatic speech recognition, speaker identification, keyword spotting and others. While the features are undeniably useful in speech recognition and associated tasks, their utility in speech enhancement systems is yet to be firmly established, and perhaps not properly understood. In this paper, we investigate the uses of SSL representations for single-channel speech enhancement in challenging conditions and find that they add very little value for the enhancement task. Our constraints are designed around on-device real-time speech enhancement -- model is causal, the compute footprint is small. Additionally, we focus on low SNR conditions where such models struggle to provide good enhancement. In order to systematically examine how SSL representations impact performance of such enhancement models, we propose a variety of techniques to utilize these embeddings which include different forms of knowledge-distillation and pre-training.


【6】 a-DCF: an architecture agnostic metric with application to  spoofing-robust speaker verification
标题:A-DCF:一种与体系结构无关的度量方法及其在欺骗性说话人验证中的应用
链接:https://arxiv.org/abs/2403.01355
作者:Hye-jin Shim,Jee-weon Jung,Tomi Kinnunen,Nicholas Evans,Jean-Francois Bonastre,Itshak Lapidot
备注:8 pages, submitted to Speaker Odyssey 2024
摘要:欺骗检测是当今的主流研究课题。标准度量可以应用于评估孤立欺骗检测解决方案的性能,并且已经提出了其他度量来支持它们与扬声器检测相结合时的评估。这些要么具有众所周知的缺陷,要么限制了将扬声器和欺骗检测器相结合的架构方法。在本文中,我们提出了一个架构无关的检测成本函数(a-DCF)。a-DCF是广泛用于自动说话人确认(ASV)评估的原始DCF的概括,旨在评估欺骗鲁棒ASV。与DCF一样,a-DCF反映了贝叶斯风险意义上的决策成本,具有明确定义的类先验和检测成本模型。我们证明的优点的a-DCF通过基准测试评估架构异构欺骗强大的ASV解决方案。
摘要:Spoofing detection is today a mainstream research topic. Standard metrics can be applied to evaluate the performance of isolated spoofing detection solutions and others have been proposed to support their evaluation when they are combined with speaker detection. These either have well-known deficiencies or restrict the architectural approach to combine speaker and spoof detectors. In this paper, we propose an architecture-agnostic detection cost function (a-DCF). A generalisation of the original DCF used widely for the assessment of automatic speaker verification (ASV), the a-DCF is designed for the evaluation of spoofing-robust ASV. Like the DCF, the a-DCF reflects the cost of decisions in a Bayes risk sense, with explicitly defined class priors and detection cost model. We demonstrate the merit of the a-DCF through the benchmarking evaluation of architecturally-heterogeneous spoofing-robust ASV solutions.


【7】 Arbitrary Discrete Fourier Analysis and Its Application in Replayed  Speech Detection
标题:任意离散傅里叶分析及其在重放语音检测中的应用
链接:https://arxiv.org/abs/2403.01130
作者:Shih-Kuang Lee
备注:this https URL
摘要:在本文中,信号分析的概念时,推导出一个特定的频率成分的频谱分析中的傅立叶分析。三种信号分析方法,然后开发的基础上衍生的概念,即任意离散傅立叶分析(ADFA),梅尔尺度离散傅立叶分析(MDFA),和常数Q分析(CQA)。我验证了这三种信号分析方法的有效性,通过测试他们的性能重放语音检测基准(即,ASVspoof 2019物理访问)以及最先进的模型。实验结果表明,这三种信号分析方法的性能与最好的报告系统。同时,本文提出的CQA方法的计算时间比常用的常Q变换方法要短得多,常Q变换方法在语音检测和音乐处理中有着广泛的应用。
摘要:In this paper, a signal analysis concept is derived when revisiting how a specific frequency component in spectrum is analyzed in Fourier analysis. Three signal analysis methods are then developed based on the derived concept, namely Arbitrary Discrete Fourier Analysis (ADFA), Mel-scale Discrete Fourier Analysis (MDFA), and constant Q Analysis (CQA). I validate the effectiveness of these three signal analysis methods by testing their performance on a replayed speech detection benchmark (i.e., the ASVspoof 2019 Physical Access) along with a state-of-the-art model. Experimental results show that the performance of these three signal analysis methods is comparable to the best reported systems. At the same time, it is show that the computation time of the developed method CQA is much shorter than the convention method constant Q Transform, which is commonly used in spoofed and fake speech detection and music processing.

【8】 SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in  Speech
标题:SEGAA:预测语音中的年龄、性别和情绪的统一方法
链接:https://arxiv.org/abs/2403.00887
作者:Aron R,Indra Sigicharla,Chirag Periwal,Mohanaprasad K,Nithya Darisini P S,Sourabh Tiwari,Shivani Arora
摘要:人类声音的解释在各种应用中都很重要。这项研究冒险预测年龄,性别和情绪的声音线索,一个领域具有广泛的应用。语音分析技术的进步跨越了各个领域,从改善客户互动到增强医疗保健和零售体验。辨别情绪有助于心理健康,而年龄和性别检测在各种情况下都至关重要。探索这些预测的深度学习模型涉及比较本文中强调的单输出,多输出和顺序模型。寻找合适的数据带来了挑战,导致CREMA-D和EMO-DB数据集的合并。先前的工作表明,在个别预测的承诺,但有限的研究同时考虑所有三个变量。本文指出了个体模型方法的缺陷,并倡导我们新的多输出学习架构基于语音的情感性别和年龄分析(SEGAA)模型。实验表明,多输出模型对单个模型进行了改进,有效地捕捉了变量和语音输入之间的复杂关系,同时提高了运行时间。
摘要:The interpretation of human voices holds importance across various applications. This study ventures into predicting age, gender, and emotion from vocal cues, a field with vast applications. Voice analysis tech advancements span domains, from improving customer interactions to enhancing healthcare and retail experiences. Discerning emotions aids mental health, while age and gender detection are vital in various contexts. Exploring deep learning models for these predictions involves comparing single, multi-output, and sequential models highlighted in this paper. Sourcing suitable data posed challenges, resulting in the amalgamation of the CREMA-D and EMO-DB datasets. Prior work showed promise in individual predictions, but limited research considered all three variables simultaneously. This paper identifies flaws in an individual model approach and advocates for our novel multi-output learning architecture Speech-based Emotion Gender and Age Analysis (SEGAA) model. The experiments suggest that Multi-output models perform comparably to individual models, efficiently capturing the intricate relationships between variables and speech inputs, all while achieving improved runtime.


【9】 SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
标题:SA-SOT:面向多说话者ASR的说话者感知串行输出训练
链接:https://arxiv.org/abs/2403.02010
作者:Zhiyun Fan,Linhao Dong,Jun Zhang,Lu Lu,Zejun Ma
摘要:多说话者自动语音识别在涉及多方交互的场景中起着至关重要的作用,例如会议和对话。由于其固有的复杂性,这项任务越来越受到重视。值得注意的是,序列化输出训练(SOT)由于其简单的架构和出色的性能而在各种方法中脱颖而出。然而,在令牌级SOT(t-SOT)的频繁的扬声器的变化提出了挑战的自回归解码器在有效地利用上下文来预测输出序列。为了解决这个问题,我们引入了一个掩蔽的t-SOT标签,它作为辅助训练损失的基石。此外,我们利用说话人相似度矩阵来改进解码器的自注意机制。这种策略性的调整增强了同一说话人的话语标记内的上下文关系,同时最大限度地减少了不同说话人话语标记之间的交互。我们将我们的方法称为说话者感知SOT(SA-SOT)。在Librispeech数据集上的实验表明,我们的SA-SOT在多人测试集上获得了12.75%到22.03%的相对cpWER降低。此外,通过更广泛的训练,我们的方法实现了令人印象深刻的3.41%的cpWER,在LibrispeechMix数据集上建立了一个新的最先进的结果。
摘要:Multi-talker automatic speech recognition plays a crucial role in scenarios involving multi-party interactions, such as meetings and conversations. Due to its inherent complexity, this task has been receiving increasing attention. Notably, the serialized output training (SOT) stands out among various approaches because of its simplistic architecture and exceptional performance. However, the frequent speaker changes in token-level SOT (t-SOT) present challenges for the autoregressive decoder in effectively utilizing context to predict output sequences. To address this issue, we introduce a masked t-SOT label, which serves as the cornerstone of an auxiliary training loss. Additionally, we utilize a speaker similarity matrix to refine the self-attention mechanism of the decoder. This strategic adjustment enhances contextual relationships within the same speaker's tokens while minimizing interactions between different speakers' tokens. We denote our method as speaker-aware SOT (SA-SOT). Experiments on the Librispeech datasets demonstrate that our SA-SOT obtains a relative cpWER reduction ranging from 12.75% to 22.03% on the multi-talker test sets. Furthermore, with more extensive training, our method achieves an impressive cpWER of 3.41%, establishing a new state-of-the-art result on the LibrispeechMix dataset.

【10】 Fine-Grained Quantitative Emotion Editing for Speech Generation
标题:面向语音生成的细粒度量化情感编辑
链接:https://arxiv.org/abs/2403.02002
作者:Sho Inoue,Kun Zhou,Shuai Wang,Haizhou Li
备注:This paper is submitted to IEEE Signal Processing Letters
摘要:如何定量地控制语音情感的表现力是语音生成中的一个重要挑战。在这项工作中,我们提出了一种新的方法来操纵渲染的情绪语音生成。我们提出了一个分层的情感分布提取器,即分层ED,量化的情绪强度在不同级别的粒度。支持向量机(SVM)被用来排名情感强度,从而产生一个层次化的情感嵌入。分层ED随后被集成到FastSpeech2框架中,指导模型在音素,单词和话语水平上学习情感强度。在合成过程中,用户可以手动编辑生成的语音的情感强度。客观和主观的评价表明,所提出的网络在细粒度的定量情感编辑方面的有效性。
摘要:It remains a significant challenge how to quantitatively control the expressiveness of speech emotion in speech generation. In this work, we present a novel approach for manipulating the rendering of emotions for speech generation. We propose a hierarchical emotion distribution extractor, i.e. Hierarchical ED, that quantifies the intensity of emotions at different levels of granularity. Support vector machines (SVMs) are employed to rank emotion intensity, resulting in a hierarchical emotional embedding. Hierarchical ED is subsequently integrated into the FastSpeech2 framework, guiding the model to learn emotion intensity at phoneme, word, and utterance levels. During synthesis, users can manually edit the emotional intensity of the generated voices. Both objective and subjective evaluations demonstrate the effectiveness of the proposed network in terms of fine-grained quantitative emotion editing.

【11】 A robust audio deepfake detection system via multi-view feature
标题:一种基于多视角特征的健壮音频深度伪码检测系统
链接:https://arxiv.org/abs/2403.01960
作者:Yujie Yang,Haochen Qin,Hang Zhou,Chengcheng Wang,Tianyu Guo,Kai Han,Yunhe Wang
备注:5 pages, 2 figures
摘要:随着生成建模技术的进步,合成的人类语音变得越来越难以与真实语音区分,这给音频深度伪造检测(ADD)系统带来了棘手的挑战。在本文中,我们利用音频功能,以提高ADD系统的泛化能力。ADD任务性能的调查进行了广泛的音频功能,包括各种手工制作的功能和学习为基础的功能。实验表明,在大量数据上预训练的基于学习的音频特征在域外场景上比手工制作的特征更好地泛化。随后,我们进一步提高了推广的ADD系统使用建议的多功能的方法,将互补的信息,从不同的观点的功能。在ASV 2019数据上训练的模型在In-the-Wild数据集上实现了24.27%的相等错误率。
摘要:With the advancement of generative modeling techniques, synthetic human speech becomes increasingly indistinguishable from real, and tricky challenges are elicited for the audio deepfake detection (ADD) system. In this paper, we exploit audio features to improve the generalizability of ADD systems. Investigation of the ADD task performance is conducted over a broad range of audio features, including various handcrafted features and learning-based features. Experiments show that learning-based audio features pretrained on a large amount of data generalize better than hand-crafted features on out-of-domain scenarios. Subsequently, we further improve the generalizability of the ADD system using proposed multi-feature approaches to incorporate complimentary information from features of different views. The model trained on ASV2019 data achieves an equal error rate of 24.27\% on the In-the-Wild dataset.


【12】 ConSep: a Noise- and Reverberation-Robust Speech Separation Framework by  Magnitude Conditioning
标题:Consep:一种基于幅度条件的抗噪混响语音分离框架
链接:https://arxiv.org/abs/2403.01792
作者:Kuan-Hsun Ho,Jeih-weih Hung,Berlin Chen
摘要:由于时域方法中使用的细粒度视觉,语音分离最近取得了重大进展。然而,一些研究表明,采用短时傅立叶变换(STFT)进行特征提取在遇到更苛刻的条件(如噪声或混响)时可能是有益的。因此,我们提出了一个幅度条件的时域框架,ConSep,继承的有益特性。实验表明,ConSep促进性能在消声,嘈杂,混响的设置相比,两个著名的方法,SepFormer和Bi-Sep。此外,我们可视化ConSep的组件,以加强优势和一致性的现状,我们已经发现在初步研究。
摘要:Speech separation has recently made significant progress thanks to the fine-grained vision used in time-domain methods. However, several studies have shown that adopting Short-Time Fourier Transform (STFT) for feature extraction could be beneficial when encountering harsher conditions, such as noise or reverberation. Therefore, we propose a magnitude-conditioned time-domain framework, ConSep, to inherit the beneficial characteristics. The experiment shows that ConSep promotes performance in anechoic, noisy, and reverberant settings compared to two celebrated methods, SepFormer and Bi-Sep. Furthermore, we visualize the components of ConSep to strengthen the advantages and cohere with the actualities we have found in preliminary studies.

【13】 What do neural networks listen to? Exploring the crucial bands in Speech  Enhancement using Sinc-convolution
标题:神经网络听什么?正弦卷积在语音增强中关键频段的探索
链接:https://arxiv.org/abs/2403.01785
作者:Kuan-Hsun Ho,Jeih-weih Hung,Berlin Chen
摘要:本研究介绍了一种改进的Sinc-convolution(Sincconv)框架,该框架专为深度网络的编码器组件定制,用于语音增强(SE)。改革后的Sincconv,参数化sinc函数作为带通滤波器的基础上,提供了显着的优势,在训练效率,滤波器的多样性,和可解释性。改革后的Sinc-conv与各种SE模型一起进行评估,展示了其提高SE性能的能力。此外,改进的Sincconv提供了对在SE场景中优先化的特定频率分量的有价值的见解。这开辟了SE研究的新方向,并提高了我们对其运行动态的了解。
摘要:This study introduces a reformed Sinc-convolution (Sincconv) framework tailored for the encoder component of deep networks for speech enhancement (SE). The reformed Sincconv, based on parametrized sinc functions as band-pass filters, offers notable advantages in terms of training efficiency, filter diversity, and interpretability. The reformed Sinc-conv is evaluated in conjunction with various SE models, showcasing its ability to boost SE performance. Furthermore, the reformed Sincconv provides valuable insights into the specific frequency components that are prioritized in an SE scenario. This opens up a new direction of SE research and improving our knowledge of their operating dynamics.


【14】 Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based  Audio-Visual Conformer
标题:基于视听同步器的帧级跨通道注意的健壮唤醒词检测
链接:https://arxiv.org/abs/2403.01700
作者:Haoxu Wang,Ming Cheng,Qiang Fu,Ming Li
备注:Accepted by ICASSP 2024
摘要:近年来,基于神经网络的Wake Word Spotting在干净的音频样本上取得了良好的性能,但在嘈杂的环境中表现不佳。听觉-视觉唤醒词识别技术(AVWWS)由于不受复杂声音环境的影响而受到广泛关注。以往的作品通常使用简单的加法或拼接进行多模态融合。模式间的相关性仍然相对不足的探索。在本文中,我们提出了一种新的模块称为帧级跨模态注意(FLCMA),以提高AVWWS系统的性能。该模块可以通过同步的嘴唇运动和语音信号在帧级帮助建模多模态信息。我们训练基于端到端FLCMA的视听一致性,并通过微调预训练的单模态模型来进一步提高性能。所提出的系统实现了一个新的国家的最先进的结果(4.57% WWS分数)的远场MISP数据集。
摘要:In recent years, neural network-based Wake Word Spotting achieves good performance on clean audio samples but struggles in noisy environments. Audio-Visual Wake Word Spotting (AVWWS) receives lots of attention because visual lip movement information is not affected by complex acoustic scenes. Previous works usually use simple addition or concatenation for multi-modal fusion. The inter-modal correlation remains relatively under-explored. In this paper, we propose a novel module called Frame-Level Cross-Modal Attention (FLCMA) to improve the performance of AVWWS systems. This module can help model multi-modal information at the frame-level through synchronous lip movements and speech signals. We train the end-to-end FLCMA based Audio-Visual Conformer and further improve the performance by fine-tuning pre-trained uni-modal models for the AVWWS task. The proposed system achieves a new state-of-the-art result (4.57% WWS score) on the far-field MISP dataset.


【15】 Brilla AI: AI Contestant for the National Science and Maths Quiz
标题:Brilla AI:美国国家科学和数学竞赛的AI选手
链接:https://arxiv.org/abs/2403.01699
作者:George Boateng,Jonathan Abrefah Mensah,Kevin Takyi Yeboah,William Edor,Andrew Kojo Mensah-Onumah,Naafi Dasana Ibrahim,Nana Sam Yeboah
备注:13 pages. Under review at the 25th International Conference on AI in Education (AIED 2024)
摘要:非洲大陆缺乏足够的合格教师,这妨碍了提供适当的学习支助。人工智能可能会增加数量有限的教师的努力,从而带来更好的学习效果。为此,这项工作描述和评估了NATURAL AI Grand Challenge的第一个关键输出,它为这样的AI提出了一个强大的现实世界基准:“构建一个AI来参加加纳的国家科学和数学测验(NATURAL)比赛并获胜-在比赛的所有回合和阶段中表现得比最好的参赛者更好。在加纳,每年一度的高中生科学和数学竞赛,由两名学生组成的三个团队通过回答生物,化学,物理和数学的问题进行竞争,在5轮中进行5个渐进阶段,直到获胜的团队获得该年度的冠军。在这项工作中,我们构建了Brilla AI,这是一个人工智能参赛者,我们部署它来进行非正式的远程比赛,并参加2023年Numbers总决赛的谜语回合,这是该比赛30年历史上的第一次。Brilla AI目前是一款网络应用程序,可以直播比赛的谜语回合,并运行4个机器学习系统:(1)语音到文本(2)问题提取(3)问题回答和(4)文本到语音,这些系统实时协同工作,快速准确地提供答案,然后用加纳口音说出来。在首次亮相时,我们的人工智能回答了4个谜语中的一个,领先于3个人类参赛队,非正式地排名第二(并列)。这种人工智能的改进和扩展可能被部署为向学生提供科学辅导,并最终使非洲数百万人能够进行一对一的学习互动,使科学教育民主化。
摘要:The African continent lacks enough qualified teachers which hampers the provision of adequate learning support. An AI could potentially augment the efforts of the limited number of teachers, leading to better learning outcomes. Towards that end, this work describes and evaluates the first key output for the NSMQ AI Grand Challenge, which proposes a robust, real-world benchmark for such an AI: "Build an AI to compete live in Ghana's National Science and Maths Quiz (NSMQ) competition and win - performing better than the best contestants in all rounds and stages of the competition". The NSMQ is an annual live science and mathematics competition for senior secondary school students in Ghana in which 3 teams of 2 students compete by answering questions across biology, chemistry, physics, and math in 5 rounds over 5 progressive stages until a winning team is crowned for that year. In this work, we built Brilla AI, an AI contestant that we deployed to unofficially compete remotely and live in the Riddles round of the 2023 NSMQ Grand Finale, the first of its kind in the 30-year history of the competition. Brilla AI is currently available as a web app that livestreams the Riddles round of the contest, and runs 4 machine learning systems: (1) speech to text (2) question extraction (3) question answering and (4) text to speech that work together in real-time to quickly and accurately provide an answer, and then say it with a Ghanaian accent. In its debut, our AI answered one of the 4 riddles ahead of the 3 human contesting teams, unofficially placing second (tied). Improvements and extensions of this AI could potentially be deployed to offer science tutoring to students and eventually enable millions across Africa to have one-on-one learning interactions, democratizing science education.


【16】 Enhancing Audio Generation Diversity with Visual Information
标题:利用视觉信息增强音频生成多样性
链接:https://arxiv.org/abs/2403.01278
作者:Zeyu Xie,Baihan Li,Xuenan Xu,Mengyue Wu,Kai Yu
摘要:近年来,音频和声音生成获得了极大的关注,主要集中在提高生成的音频的质量上。然而,关于增强所生成的音频的多样性的研究有限,特别是当涉及到特定类别内的音频生成时。当前模型倾向于在类别内产生同质音频样本。这项工作的目的是解决这个限制,通过改善生成的音频与视觉信息的多样性。我们提出了一种基于聚类的方法,利用视觉信息来指导模型在每个类别中生成不同的音频内容。七个类别的结果表明,额外的视觉输入可以大大提高音频生成的多样性。音频样本可在https://zeyuxie29.github.io/DiverseAudioGeneration上获得。
摘要:Audio and sound generation has garnered significant attention in recent years, with a primary focus on improving the quality of generated audios. However, there has been limited research on enhancing the diversity of generated audio, particularly when it comes to audio generation within specific categories. Current models tend to produce homogeneous audio samples within a category. This work aims to address this limitation by improving the diversity of generated audio with visual information. We propose a clustering-based method, leveraging visual information to guide the model in generating distinct audio content within each category. Results on seven categories indicate that extra visual input can largely enhance audio generation diversity. Audio samples are available at https://zeyuxie29.github.io/DiverseAudioGeneration.


【17】 Automatic Speech Recognition using Advanced Deep Learning Approaches: A  survey
标题:使用高级深度学习方法的自动语音识别:综述
链接:https://arxiv.org/abs/2403.01255
作者:Hamza Kheddar,Mustapha Hemis,Yassine Himeur
摘要:深度学习(DL)的最新进展对自动语音识别(ASR)提出了重大挑战。ASR依赖于广泛的训练数据集,包括机密数据集,并需要大量的计算和存储资源。启用自适应系统可提高动态环境中的ASR性能。DL技术假设训练和测试数据来自同一个域,但这并不总是正确的。深度迁移学习(DTL)、联邦学习(FL)和强化学习(RL)等高级DL技术可以解决这些问题。DTL允许使用小而相关的数据集进行高性能模型,FL允许在不拥有数据集的情况下对机密数据进行训练,RL优化动态环境中的决策,降低计算成本。该调查对基于DTL、FL和RL的ASR框架进行了全面的回顾,旨在深入了解最新发展,并帮助研究人员和专业人士了解当前的挑战。此外,Transformers,这是先进的DL技术,大量使用在拟议的ASR框架中,被认为是在这个调查中,他们的能力,以捕捉广泛的依赖关系的输入ASR序列。本文首先介绍了DTL、FL、RL和Transformers的背景,然后采用了一个精心设计的分类法来概述最先进的方法。随后,进行了批判性分析,以确定每个框架的长处和短处。此外,还进行了比较研究,以突出现有的挑战,为未来的研究机会铺平道路。
摘要:Recent advancements in deep learning (DL) have posed a significant challenge for automatic speech recognition (ASR). ASR relies on extensive training datasets, including confidential ones, and demands substantial computational and storage resources. Enabling adaptive systems improves ASR performance in dynamic environments. DL techniques assume training and testing data originate from the same domain, which is not always true. Advanced DL techniques like deep transfer learning (DTL), federated learning (FL), and reinforcement learning (RL) address these issues. DTL allows high-performance models using small yet related datasets, FL enables training on confidential data without dataset possession, and RL optimizes decision-making in dynamic environments, reducing computation costs. This survey offers a comprehensive review of DTL, FL, and RL-based ASR frameworks, aiming to provide insights into the latest developments and aid researchers and professionals in understanding the current challenges. Additionally, transformers, which are advanced DL techniques heavily used in proposed ASR frameworks, are considered in this survey for their ability to capture extensive dependencies in the input ASR sequence. The paper starts by presenting the background of DTL, FL, RL, and Transformers and then adopts a well-designed taxonomy to outline the state-of-the-art approaches. Subsequently, a critical analysis is conducted to identify the strengths and weaknesses of each framework. Additionally, a comparative study is presented to highlight the existing challenges, paving the way for future research opportunities.


【18】 MPIPN: A Multi Physics-Informed PointNet for solving parametric  acoustic-structure systems
标题:MPIPN:求解参数声学结构系统的多物理信息点网络
链接:https://arxiv.org/abs/2403.01132
作者:Chu Wang,Jinhong Wu,Yanzhi Wang,Zhijian Zha,Qi Zhou
备注:The number of figures is 16. The number of tables is 5. The number of words is 9717
摘要:机器学习用于求解由一般非线性偏微分方程(PDE)控制的物理系统。然而,复杂的多物理场系统,如声-结构耦合,往往是由一系列的偏微分方程,其中包括可变的物理量,这被称为参数系统。对于包含显式和隐式量的偏微分方程所控制的参数系统,缺乏求解策略。本文提出了一种基于深度学习的多物理信息点网络(MPIPN),用于求解参数化声学-结构系统。首先,MPIPN诱导一个增强的点云架构,包括明确的物理量和几何特征的计算域。然后,MPIPN提取局部和全局特征的重建点云的参数系统的解决标准的一部分,分别。此外,隐式物理量作为求解准则的另一部分通过编码技术嵌入。最后,所有的解决标准,表征参数系统的合并,形成独特的序列作为输入的MPIPN,其输出是系统的解决方案。所提出的框架是由相应的计算域的自适应物理信息损失函数训练的。该框架被推广到处理新的参数条件的系统。通过求解Helmholtz方程控制的稳态参数声固耦合系统,验证了MPIPN的有效性。消融实验已经实施,以证明与少数监督数据的物理通知的影响的功效。所提出的方法产生合理的精度在所有的计算域恒定的参数条件下和可变的组合的参数条件下的声学结构系统。
摘要:Machine learning is employed for solving physical systems governed by general nonlinear partial differential equations (PDEs). However, complex multi-physics systems such as acoustic-structure coupling are often described by a series of PDEs that incorporate variable physical quantities, which are referred to as parametric systems. There are lack of strategies for solving parametric systems governed by PDEs that involve explicit and implicit quantities. In this paper, a deep learning-based Multi Physics-Informed PointNet (MPIPN) is proposed for solving parametric acoustic-structure systems. First, the MPIPN induces an enhanced point-cloud architecture that encompasses explicit physical quantities and geometric features of computational domains. Then, the MPIPN extracts local and global features of the reconstructed point-cloud as parts of solving criteria of parametric systems, respectively. Besides, implicit physical quantities are embedded by encoding techniques as another part of solving criteria. Finally, all solving criteria that characterize parametric systems are amalgamated to form distinctive sequences as the input of the MPIPN, whose outputs are solutions of systems. The proposed framework is trained by adaptive physics-informed loss functions for corresponding computational domains. The framework is generalized to deal with new parametric conditions of systems. The effectiveness of the MPIPN is validated by applying it to solve steady parametric acoustic-structure coupling systems governed by the Helmholtz equations. An ablation experiment has been implemented to demonstrate the efficacy of physics-informed impact with a minority of supervised data. The proposed method yields reasonable precision across all computational domains under constant parametric conditions and changeable combinations of parametric conditions for acoustic-structure systems.


【19】 Towards Accurate Lip-to-Speech Synthesis in-the-Wild
标题:走向野外精确的唇语合成
链接:https://arxiv.org/abs/2403.01087
作者:Sindhu Hegde,Rudrabha Mukhopadhyay,C. V. Jawahar,Vinay Namboodiri
备注:None
摘要:在本文中,我们介绍了一种新的方法来解决的任务合成语音从无声的视频中的任何在野生扬声器完全基于嘴唇运动。传统的从嘴唇视频直接生成语音的方法面临着无法单独从语音中学习鲁棒的语言模型的挑战,导致结果不令人满意。为了克服这个问题,我们建议使用最先进的唇到文本网络将噪声文本监督纳入我们的模型中。噪声文本是使用预先训练的唇到文本模型生成的,使我们的方法在推理过程中无需文本注释即可工作。我们设计了一个可视化的文本到语音的网络,利用视觉流,以产生准确的语音,这是同步的无声输入视频。我们进行了广泛的实验和消融研究,证明了我们的方法在各种基准数据集上优于当前最先进的方法。此外,我们证明了我们的方法在辅助技术中的一个重要的实际应用,通过为ALS患者谁失去了声音,但可以使嘴部运动产生语音。可以在\url{http://cvit.iiit.ac.in/research/projects/cvit-projects/ms-l2s-itw}找到我们的演示视频、代码和其他详细信息。
摘要:In this paper, we introduce a novel approach to address the task of synthesizing speech from silent videos of any in-the-wild speaker solely based on lip movements. The traditional approach of directly generating speech from lip videos faces the challenge of not being able to learn a robust language model from speech alone, resulting in unsatisfactory outcomes. To overcome this issue, we propose incorporating noisy text supervision using a state-of-the-art lip-to-text network that instills language information into our model. The noisy text is generated using a pre-trained lip-to-text model, enabling our approach to work without text annotations during inference. We design a visual text-to-speech network that utilizes the visual stream to generate accurate speech, which is in-sync with the silent input video. We perform extensive experiments and ablation studies, demonstrating our approach's superiority over the current state-of-the-art methods on various benchmark datasets. Further, we demonstrate an essential practical application of our method in assistive technology by generating speech for an ALS patient who has lost the voice but can make mouth movements. Our demo video, code, and additional details can be found at \url{http://cvit.iiit.ac.in/research/projects/cvit-projects/ms-l2s-itw}.

【20】 Scaling Up Adaptive Filter Optimizers
标题:放大自适应滤波优化器
链接:https://arxiv.org/abs/2403.00977
作者:Jonah Casebeer,Nicholas J. Bryan,Paris Smaragdis
摘要:我们介绍了一种新的在线自适应滤波方法称为监督多步自适应滤波器(SMS-AF)。我们的方法使用神经网络来控制或优化线性多延迟或多通道频域滤波器,并且可以以增加计算为代价灵活地扩展性能-这是AF文献中很少涉及的属性,但对许多应用至关重要。为了做到这一点,我们扩展了最近的工作与一组改进,包括功能修剪,监督损失,和多个优化步骤,每个时间框架。这些改进以一种内聚的方式工作,以解锁缩放。此外,我们展示了我们的方法如何与卡尔曼滤波和元自适应滤波相关,使其无缝适用于各种AF任务。我们评估我们的方法声学回声消除(AEC)和多通道语音增强任务,并与标准的合成和真实世界的数据集上的几个基线进行比较。结果表明,我们的方法的性能尺度与推理成本和模型容量,产生多dB的性能增益为两个任务,并在一个单一的CPU核心上是实时的能力。
摘要:We introduce a new online adaptive filtering method called supervised multi-step adaptive filters (SMS-AF). Our method uses neural networks to control or optimize linear multi-delay or multi-channel frequency-domain filters and can flexibly scale-up performance at the cost of increased compute -- a property rarely addressed in the AF literature, but critical for many applications. To do so, we extend recent work with a set of improvements including feature pruning, a supervised loss, and multiple optimization steps per time-frame. These improvements work in a cohesive manner to unlock scaling. Furthermore, we show how our method relates to Kalman filtering and meta-adaptive filtering, making it seamlessly applicable to a diverse set of AF tasks. We evaluate our method on acoustic echo cancellation (AEC) and multi-channel speech enhancement tasks and compare against several baselines on standard synthetic and real-world datasets. Results show our method performance scales with inference cost and model capacity, yields multi-dB performance gains for both tasks, and is real-time capable on a single CPU core.

【21】 Speaker-Independent Dysarthria Severity Classification using  Self-Supervised Transformers and Multi-Task Learning
标题:使用自我监督转换器和多任务学习的非说话者依赖性构音障碍严重程度分类
链接:https://arxiv.org/abs/2403.00854
作者:Lauren Stumpf,Balasundaram Kadirvelu,Sigourney Waibel,A. Aldo Faisal
备注:17 pages, 2 tables, 4 main figures, 2 supplemental figures, prepared for journal submission
摘要:构音障碍是一种由于神经系统疾病导致的言语肌肉控制受损而导致的疾病,严重影响了患者的沟通和生活质量。这种情况的复杂性,人为评分和各种各样的表现使其评估和管理具有挑战性。这项研究提出了一个基于transformer的框架,自动评估构音障碍的严重程度,从原始语音数据。它可以提供一个客观的,可重复的,可访问的,标准化的和具有成本效益的,与传统的方法相比,需要人类专家评估。我们开发了一个Transformer框架,称为Speaker-Agnostic Latent Regularisation(SALR),将多任务学习目标和对比学习用于独立于说话者的多类构音障碍严重程度分类。多任务框架的设计,以减少对特定于扬声器的特征的依赖,并解决构音障碍的语音的内在类内变异。我们使用leave-one-speaker-out交叉验证对Universal Access Speech数据集进行了评估,我们的模型表现出优于传统机器学习方法的性能,准确率为70.48美元,F1得分为59.23美元。我们的SALR模型也超过了之前使用支持向量机的基于AI的分类基准,达到了16.58美元。我们通过可视化潜在空间来打开模型的黑盒子,在那里我们可以观察模型如何大大减少说话者特定的线索并放大特定于任务的线索,从而显示其鲁棒性。总之,SALR使用生成式AI在说话者独立的多类构音障碍严重程度分类中建立了一个新的基准。我们的研究结果对自动构音障碍严重程度评估的更广泛临床应用的潜在影响。
摘要:Dysarthria, a condition resulting from impaired control of the speech muscles due to neurological disorders, significantly impacts the communication and quality of life of patients. The condition's complexity, human scoring and varied presentations make its assessment and management challenging. This study presents a transformer-based framework for automatically assessing dysarthria severity from raw speech data. It can offer an objective, repeatable, accessible, standardised and cost-effective and compared to traditional methods requiring human expert assessors. We develop a transformer framework, called Speaker-Agnostic Latent Regularisation (SALR), incorporating a multi-task learning objective and contrastive learning for speaker-independent multi-class dysarthria severity classification. The multi-task framework is designed to reduce reliance on speaker-specific characteristics and address the intrinsic intra-class variability of dysarthric speech. We evaluated on the Universal Access Speech dataset using leave-one-speaker-out cross-validation, our model demonstrated superior performance over traditional machine learning approaches, with an accuracy of $70.48\%$ and an F1 score of $59.23\%$. Our SALR model also exceeded the previous benchmark for AI-based classification, which used support vector machines, by $16.58\%$. We open the black box of our model by visualising the latent space where we can observe how the model substantially reduces speaker-specific cues and amplifies task-specific ones, thereby showing its robustness. In conclusion, SALR establishes a new benchmark in speaker-independent multi-class dysarthria severity classification using generative AI. The potential implications of our findings for broader clinical applications in automated dysarthria severity assessments.


【22】 Structuring Concept Space with the Musical Circle of Fifths by Utilizing  Music Grammar Based Activations
标题:利用基于音乐语法的激活构建与五度音乐圈的概念空间
链接:https://arxiv.org/abs/2403.00790
作者:Tofara Moyo
备注:3 pages
摘要:在本文中,我们探索了离散神经网络(如尖峰网络)的结构与钢琴作品之间的有趣相似性。虽然两者都涉及顺序或并行激活的节点或音符,但后者受益于丰富的音乐理论来指导有意义的组合。我们提出了一种新的方法,利用音乐语法来调节尖峰神经网络中的激活,允许将符号表示为吸引子。通过应用音乐理论中的和弦进行规则,我们展示了某些激活是如何自然地跟随其他激活的,类似于吸引力的概念。此外,我们引入了调制键的概念,以导航网络内的不同流域的吸引力。最终,我们展示了我们模型中的概念图是由五度音乐圈构成的,突出了在深度学习算法中利用音乐理论原理的潜力。
摘要:In this paper, we explore the intriguing similarities between the structure of a discrete neural network, such as a spiking network, and the composition of a piano piece. While both involve nodes or notes that are activated sequentially or in parallel, the latter benefits from the rich body of music theory to guide meaningful combinations. We propose a novel approach that leverages musical grammar to regulate activations in a spiking neural network, allowing for the representation of symbols as attractors. By applying rules for chord progressions from music theory, we demonstrate how certain activations naturally follow others, akin to the concept of attraction. Furthermore, we introduce the concept of modulating keys to navigate different basins of attraction within the network. Ultimately, we show that the map of concepts in our model is structured by the musical circle of fifths, highlighting the potential for leveraging music theory principles in deep learning algorithms.


机器翻译由腾讯交互翻译提供,仅供参考