本文经arXiv每日学术速递授权转载
标题: GAMA:一个具有高级音频理解和复杂推理能力的大型音频语言模型
作者:Sreyan Ghosh,Sonal Kumar,Ashish Seth,Chandra Kiran Reddy Evuru,Utkarsh Tyagi,S Sakshi,Oriol Nieto,Ramani Duraiswami,Dinesh Manocha
备注:Project Website: this https URL
链接:点击下载PDF文件
摘要:感知和理解非言语声音和非言语言语对于帮助我们与周围环境互动的决策至关重要。在本文中,我们提出了GAMA,一种新的通用大型音频语言模型(LALM)与高级音频理解和复杂的推理能力。我们通过将LLM与多种类型的音频表示集成来构建GAMA,包括来自自定义Audio Q-Former的功能,这是一个多层聚合器,可以聚合来自音频编码器多个层的功能。我们在一个大规模的音频语言数据集上对GAMA进行了微调,从而增强了音频理解能力。接下来,我们提出了CompA-R(复杂音频推理的指令调整),这是一个合成生成的指令调整(IT)数据集,其指令要求模型对输入音频执行复杂推理。我们使用CompA-R对GAMA进行了调整,以赋予其复杂的推理能力,其中我们进一步通过利用输入音频的事件标签来添加软提示作为具有高级语义证据的输入。最后,我们还提出了CompA-R-test,一个人类标记的评估数据集,用于评估LALM在需要复杂推理的开放式音频问答上的能力。通过自动化和专家人工评估,我们表明GAMA在各种音频理解任务上的表现优于文献中的所有其他LALM,幅度为1%-84%。此外,CompA-R上的GAMA IT版证明在其复杂推理和指令遵循能力方面具有优势。摘要:Perceiving and understanding non-speech sounds and non-verbal speech is essential to making decisions that help us interact with our surroundings. In this paper, we propose GAMA, a novel General-purpose Large Audio-Language Model (LALM) with Advanced Audio Understanding and Complex Reasoning Abilities. We build GAMA by integrating an LLM with multiple types of audio representations, including features from a custom Audio Q-Former, a multi-layer aggregator that aggregates features from multiple layers of an audio encoder. We fine-tune GAMA on a large-scale audio-language dataset, which augments it with audio understanding capabilities. Next, we propose CompA-R (Instruction-Tuning for Complex Audio Reasoning), a synthetically generated instruction-tuning (IT) dataset with instructions that require the model to perform complex reasoning on the input audio. We instruction-tune GAMA with CompA-R to endow it with complex reasoning abilities, where we further add a soft prompt as input with high-level semantic evidence by leveraging event tags of the input audio. Finally, we also propose CompA-R-test, a human-labeled evaluation dataset for evaluating the capabilities of LALMs on open-ended audio question-answering that requires complex reasoning. Through automated and expert human evaluations, we show that GAMA outperforms all other LALMs in literature on diverse audio understanding tasks by margins of 1%-84%. Further, GAMA IT-ed on CompA-R proves to be superior in its complex reasoning and instruction following capabilities.
【2】 Towards an End-to-End Framework for Invasive Brain Signal Decoding with Large Language Models
标题: 采用大型语言模型建立有创大脑信号解码的端到端框架
作者:Sheng Feng,Heyang Liu,Yu Wang,Yanfeng Wang
链接:点击下载PDF文件
摘要:在本文中,我们介绍了一个突破性的端到端(E2E)框架,用于解码侵入性大脑信号,标志着语音神经假体领域的重大进展。我们的方法利用大型语言模型(LLM)的综合推理能力来促进直接解码。通过完全集成LLM,我们实现了与最先进的级联模型相当的结果。我们的研究结果强调了E2E框架在语音神经假体中的巨大潜力,特别是随着脑机接口(BCI)背后的技术和相关数据集的可用性不断发展。这项工作不仅展示了将LLM与E2E解码相结合用于增强语音神经假体的功效,还为BCI应用的未来研究确定了新的方向,强调了LLM在解码复杂神经信号以恢复通信方面的影响。代码将在https: github.com FsFrancis15 BrainLLM上提供。摘要:In this paper, we introduce a groundbreaking end-to-end (E2E) framework for decoding invasive brain signals, marking a significant advancement in the field of speech neuroprosthesis. Our methodology leverages the comprehensive reasoning abilities of large language models (LLMs) to facilitate direct decoding. By fully integrating LLMs, we achieve results comparable to the state-of-the-art cascade models. Our findings underscore the immense potential of E2E frameworks in speech neuroprosthesis, particularly as the technology behind brain-computer interfaces (BCIs) and the availability of relevant datasets continue to evolve. This work not only showcases the efficacy of combining LLMs with E2E decoding for enhancing speech neuroprosthesis but also sets a new direction for future research in BCI applications, underscoring the impact of LLMs in decoding complex neural signals for communication restoration. Code will be made available at https: github.com FsFrancis15 BrainLLM.
【3】 MusicScore: A Dataset for Music Score Modeling and Generation
标题: MusicScore:乐谱建模和生成的数据集
作者:Yuheng Lin,Zheqi Dai,Qiuqiang Kong
备注:Dataset paper, dataset link: this https URL
链接:点击下载PDF文件
摘要:乐谱是音乐的书面表示,包含有关音乐成分的丰富信息。乐谱上的视觉信息包括音符、休止符、五线谱、谱号、力度和发音。乐谱中的视觉信息比音乐的音频和符号表示包含更多的语义信息。以前的乐谱数据集大小有限,主要用于光学音乐识别(OMR)。目前缺乏关于创建用于音乐建模和生成的大规模基准数据集的研究。在这项工作中,我们提出了MusicScore,一个大规模的乐谱数据集收集和处理的国际乐谱库项目(IMSLP)。MusicScore由图像-文本对组成,其中图像是乐谱的页面,文本是音乐的元数据。MusicScore的元数据取自IMSLP页面的一般信息部分。元数据包括关于音乐作品的作曲家、乐器、作品风格和流派的丰富信息。MusicScore被分别策划成400、14 k和200 k大小的图像-文本对,并具有不同的多样性。我们构建了一个基于UNet扩散模型的乐谱生成系统,以生成视觉可读的乐谱,并以文本描述为条件,以MusicScore数据集为基准进行乐谱生成。MusicScore在https: huggingface.co datasets ZheqiDAI MusicScore上向公众发布。摘要:Music scores are written representations of music and contain rich information about musical components. The visual information on music scores includes notes, rests, staff lines, clefs, dynamics, and articulations. This visual information in music scores contains more semantic information than audio and symbolic representations of music. Previous music score datasets have limited sizes and are mainly designed for optical music recognition (OMR). There is a lack of research on creating a large-scale benchmark dataset for music modeling and generation. In this work, we propose MusicScore, a large-scale music score dataset collected and processed from the International Music Score Library Project (IMSLP). MusicScore consists of image-text pairs, where the image is a page of a music score and the text is the metadata of the music. The metadata of MusicScore is extracted from the general information section of the IMSLP pages. The metadata includes rich information about the composer, instrument, piece style, and genre of the music pieces. MusicScore is curated into small, medium, and large scales of 400, 14k, and 200k image-text pairs with varying diversity, respectively. We build a score generation system based on a UNet diffusion model to generate visually readable music scores conditioned on text descriptions to benchmark the MusicScore dataset for music score generation. MusicScore is released to the public at https: huggingface.co datasets ZheqiDAI MusicScore.
【4】 AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
标题: Anopatch:在机器异常声音检测中实现更好的一致性
作者:Anbai Jiang,Bing Han,Zhiqiang Lv,Yufeng Deng,Wei-Qiang Zhang,Xie Chen,Yanmin Qian,Jia Liu,Pingyi Fan
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:大型预训练模型在多个领域表现出了主导性的性能,其中预训练和微调之间的一致性是成功的关键。然而,很少有工程报告的机器异常声音检测(ASD)任务的预训练模型的令人满意的结果。这可能是由于预训练模型的不一致性和机器音频的归纳偏差,导致数据和架构不一致。因此,我们提出了AnoPatch,它利用在AudioSet上预先训练的ViT骨干,并在机器音频上对其进行微调。人们认为,机器音频与音频数据集比语音数据集更相关,并且从补丁级别对其建模适合机器音频的稀疏性。因此,AnoPatch在DCASE 2020 ASD数据集和DCASE 2023 ASD数据集上展示了最先进的(SOTA)性能。我们还比较了多个预先训练的模型,并通过经验证明了更好的一致性会带来相当大的改进。摘要:Large pre-trained models have demonstrated dominant performances in multiple areas, where the consistency between pre-training and fine-tuning is the key to success. However, few works reported satisfactory results of pre-trained models for the machine anomalous sound detection (ASD) task. This may be caused by the inconsistency of the pre-trained model and the inductive bias of machine audio, resulting in inconsistency in data and architecture. Thus, we propose AnoPatch which utilizes a ViT backbone pre-trained on AudioSet and fine-tunes it on machine audio. It is believed that machine audio is more related to audio datasets than speech datasets, and modeling it from patch level suits the sparsity of machine audio. As a result, AnoPatch showcases state-of-the-art (SOTA) performances on the DCASE 2020 ASD dataset and the DCASE 2023 ASD dataset. We also compare multiple pre-trained models and empirically demonstrate that better consistency yields considerable improvement.
【5】 SMRU: Split-and-Merge Recurrent-based UNet for Acoustic Echo Cancellation and Noise Suppression
标题: SMRU:基于分离合并的回归UNet,用于声学回声消除和噪音抑制
作者:Zhihang Sun,Andong Li,Rilin Chen,Hao Zhang,Meng Yu,Yi Zhou,Dong Yu
链接:点击下载PDF文件
摘要:深度神经网络的激增催生了声学回声消除和噪声抑制的快速发展,并且已经提出了大量现有技术,这些技术产生了有前途的性能。然而,他们很少考虑不同处理场景(如边缘设备和云处理)中的部署通用性。为此,本文提出了一个通用的模型,称为SMRU,以涵盖不同的应用场景。新颖之处在于双重性。首先,提出了多尺度频带分离层和频带合并层,以有效地融合局部频带,从而降低建模复杂度。此外,通过模拟经典UNet结构的多分辨率特征建模特性,设计了一种新的递归支配UNet结构。它由多个可变帧速率块组成,每个块都涉及具有不同压缩比的因果时间下 上采样层以及用于带间和带内建模的双路径结构。该模型配置为从50 M s到6.8 G s的MAC,实验结果表明,所提出的方法产生的竞争力,甚至更好的性能超过现有的基线,并有充分的潜力,以适应更一般的情况下,不同的复杂性要求。摘要:The proliferation of deep neural networks has spawned the rapid development of acoustic echo cancellation and noise suppression, and plenty of prior arts have been proposed, which yield promising performance. Nevertheless, they rarely consider the deployment generality in different processing scenarios, such as edge devices, and cloud processing. To this end, this paper proposes a general model, termed SMRU, to cover different application scenarios. The novelty lies in two-fold. First, a multi-scale band split layer and band merge layer are proposed to effectively fuse local frequency bands for lower complexity modeling. Besides, by simulating the multi-resolution feature modeling characteristic of the classical UNet structure, a novel recurrent-dominated UNet is devised. It consists of multiple variable frame rate blocks, each of which involves the causal time down- up-sampling layer with varying compression ratios and the dual-path structure for inter- and intra-band modeling. The model is configured from 50 M s to 6.8 G s in terms of MACs, and the experimental results show that the proposed approach yields competitive or even better performance over existing baselines, and has the full potential to adapt to more general scenarios with varying complexity requirements.
【6】 Identification of Physical Properties in Acoustic Tubes Using Physics-Informed Neural Networks
标题: 使用物理信息神经网络识别声管的物理性能
作者:Kazuya Yokota,Masataka Ogura,Masajiro Abe
备注:10 pages, 7 figures, The following article has been submitted to Mechanical Engineering Journal. After it is published, it will be found at this https URL
链接:点击下载PDF文件
摘要:物理信息神经网络(Physics-informed Neural Networks,PINN)是一种数值模拟方法,它将与控制方程对应的损失函数合并到神经网络中。虽然PINN已经被探索用于逆分析,但它们在声学分析中的应用仍然有限。本文提出了一种利用PINNs识别声管内损耗参数的方法。我们将损失参数分为两类:一类依赖于管道直径,另一类是与管道直径无关的常数,后者被设置为神经网络的可训练参数。将损耗参数的确定问题转化为一个优化问题,通过这个过程确定材料的物理性质。所采用的神经网络架构是基于我们以前提出的ResoNet,这是专为分析声学共振。所提出的方法的有效性进行评估,通过正向和反向分析,特别是通过识别的损失参数。研究结果表明,它是可行的,以准确地识别参数,显着影响下的声场分析。通过改变损失函数中的控制方程,该方法可以适用于各种声场,具有广泛的应用前景。摘要:Physics-informed Neural Networks (PINNs) is a method for numerical simulation that incorporates a loss function corresponding to the governing equations into a neural network. While PINNs have been explored for their utility in inverse analysis, their application in acoustic analysis remains limited. This study presents a method to identify loss parameters in acoustic tubes using PINNs. We categorized the loss parameters into two groups: one dependent on the tube's diameter and another constant, independent of it. The latter were set as the trainable parameters of the neural network. The problem of identifying the loss parameter was formulated as an optimization problem, with the physical properties being determined through this process. The neural network architecture employed was based on our previously proposed ResoNet, which is designed for analyzing acoustic resonance. The efficacy of the proposed method is assessed through both forward and inverse analysis, specifically through the identification of loss parameters. The findings demonstrate that it is feasible to accurately identify parameters that significantly impact the sound field under analysis. By merely altering the governing equations in the loss function, this method could be adapted to various sound fields, suggesting its potential for broad application.
【7】 NAST: Noise Aware Speech Tokenization for Speech Language Models
标题: NAST:语音语言模型的噪音感知语音令牌化
作者:Shoval Messica,Yossi Adi
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:语音标记化是将语音信号表示为离散单元序列的任务。这样的表示可以用于各种下游任务,包括自动语音识别,文本到语音等更相关的这项研究,这样的表示作为语音语言模型的基础。在这项工作中,我们解决了在噪声环境下的语音标记化任务,并提出了NAST:噪声感知语音标记化的语音语言模型。NAST由三个主要组件组成:(i)预测器;(ii)残差编码器;以及(iii)解码器。我们评估NAST的效率,考虑几个口语建模任务,并表明NAST优于所有设置的评估基线。最后,我们分析了NAST,并显示其解纠缠特性和鲁棒性的信号变化的形式的噪声,混响,音高移位,和时间拉伸。代码和预训练模型可在https: github.com ShovalMessica NAST上获得。摘要:Speech tokenization is the task of representing speech signals as a sequence of discrete units. Such representations can be later used for various downstream tasks including automatic speech recognition, text-to-speech, etc. More relevant to this study, such representation serves as the basis of Speech Language Models. In this work, we tackle the task of speech tokenization under the noisy setup and present NAST: Noise Aware Speech Tokenization for Speech Language Models. NAST is composed of three main components: (i) a predictor; (ii) a residual encoder; and (iii) a decoder. We evaluate the efficiency of NAST considering several spoken language modeling tasks and show that NAST is superior to the evaluated baselines across all setups. Lastly, we analyze NAST and show its disentanglement properties and robustness to signal variations in the form of noise, reverberation, pitch-shift, and time-stretch. Code and pre-trained models are available at https: github.com ShovalMessica NAST.
【8】 Large Language Models for Dysfluency Detection in Stuttered Speech
标题: 用于口吃语音流畅性检测的大型语言模型
作者:Dominik Wagner,Sebastian P. Bayerl,Ilja Baumann,Korbinian Riedhammer,Elmar Nöth,Tobias Bocklet
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:准确检测口语中的不流利可以帮助提高自动语音和语言处理组件的性能,并支持开发更具包容性的语音和语言技术。受最近部署大型语言模型(LLM)作为非词汇输入(如音频和视频)的通用学习器和处理器的趋势的启发,我们将多标签不流利检测任务作为语言建模问题。我们提出的假设候选人产生的自动语音识别系统和声学表示提取的音频编码器模型的LLM,微调系统预测不流利的标签上的三个数据集包含英语和德语口吃的语音。实验结果表明,我们的系统有效地结合了声学和词汇信息,并取得了多标签口吃检测任务的竞争力的结果。摘要:Accurately detecting dysfluencies in spoken language can help to improve the performance of automatic speech and language processing components and support the development of more inclusive speech and language technologies. Inspired by the recent trend towards the deployment of large language models (LLMs) as universal learners and processors of non-lexical inputs, such as audio and video, we approach the task of multi-label dysfluency detection as a language modeling problem. We present hypotheses candidates generated with an automatic speech recognition system and acoustic representations extracted from an audio encoder model to an LLM, and finetune the system to predict dysfluency labels on three datasets containing English and German stuttered speech. The experimental results show that our system effectively combines acoustic and lexical information and achieves competitive results on the multi-label stuttering detection task.
【9】 Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
标题: 利用门控注意力减少离群值以改进大型序列到序列语音基础模型中的训练后量化
作者:Dominik Wagner,Ilja Baumann,Korbinian Riedhammer,Tobias Bocklet
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:本文探讨了在Whisper语音基础模型族中进行知识蒸馏后的后训练量化(PTQ)的改进。我们解决了权重和激活张量中离群值的挑战,已知这些离群值会阻碍基于变换的语言和视觉模型中的量化质量。将这一观察结果扩展到Whisper,我们证明了当基于transformer的模型被训练来执行自动语音识别时,这些离群值也存在,从而需要PTQ的缓解策略。我们表明,离群值可以减少最近提出的门控机制,在学生模型的注意力块,使有效的8位量化,和较低的字错误率相比,学生模型没有门控机制到位。摘要:This paper explores the improvement of post-training quantization (PTQ) after knowledge distillation in the Whisper speech foundation model family. We address the challenge of outliers in weights and activation tensors, known to impede quantization quality in transformer-based language and vision models. Extending this observation to Whisper, we demonstrate that these outliers are also present when transformer-based models are trained to perform automatic speech recognition, necessitating mitigation strategies for PTQ. We show that outliers can be reduced by a recently proposed gating mechanism in the attention blocks of the student model, enabling effective 8-bit quantization, and lower word error rates compared to student models without the gating mechanism in place.
【10】 SPEAR: Receiver-to-Receiver Acoustic Neural Warping Field
标题: SPSYS:接收器到接收器的声学神经扭曲场
作者:Yuhang He,Shitong Xu,Jia-Xing Zhong,Sangyun Shin,Niki Trigoni,Andrew Markham
备注:9 pages, 5 figures in main paper
链接:点击下载PDF文件
摘要:我们提出了一个连续的接收器到接收器的声学神经扭曲场,用于在具有单个固定音频源的声学3D空间中进行空间声学效果预测。与传统的源到接收器建模方法,需要先前的空间声学特性知识,严格地模拟音频传播从源到接收器,我们建议预测通过扭曲的空间声学效果从一个参考接收器位置到另一个目标接收器位置,使扭曲的音频基本上容纳所有的空间声学效果属于目标位置。SPARK可以以一种更容易获得数据的方式进行训练,我们只需让两个机器人在不同的位置独立地记录空间音频。我们进一步从理论上证明了翘曲场的普遍存在当且仅当一个音频源存在。三个物理原则被纳入到指导SPRINT网络设计,导致学习翘曲场物理意义。我们在合成的、照片般逼真的和真实世界的数据集上展示了SPARTS的优越性,展示了SPARTS在各种下游机器人任务中的巨大潜力。摘要:We present SPEAR, a continuous receiver-to-receiver acoustic neural warping field for spatial acoustic effects prediction in an acoustic 3D space with a single stationary audio source. Unlike traditional source-to-receiver modelling methods that require prior space acoustic properties knowledge to rigorously model audio propagation from source to receiver, we propose to predict by warping the spatial acoustic effects from one reference receiver position to another target receiver position, so that the warped audio essentially accommodates all spatial acoustic effects belonging to the target position. SPEAR can be trained in a data much more readily accessible manner, in which we simply ask two robots to independently record spatial audio at different positions. We further theoretically prove the universal existence of the warping field if and only if one audio source presents. Three physical principles are incorporated to guide SPEAR network design, leading to the learned warping field physically meaningful. We demonstrate SPEAR superiority on both synthetic, photo-realistic and real-world dataset, showing the huge potential of SPEAR to various down-stream robotic tasks.
【11】 CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving
标题: CoSTA:使用对齐语音文本交织的代码交换语音翻译
作者:Bhavani Shankar,Preethi Jyothi,Pushpak Bhattacharyya
链接:点击下载PDF文件
摘要:语码转换是印度等多语言社会中普遍存在的语言现象。由于数据集的可用性有限,为代码切换语音构建语音到文本模型具有挑战性。在这项工作中,我们专注于口语翻译(ST)的问题,在印度语言的代码转换语音到英语文本。我们提出了一个新的端到端模型架构COSTA,它以预训练的自动语音识别(ASR)和机器翻译(MT)模块(适用于许多语言)为基础。语音和ASR文本表示使用对齐的交织方案进行融合,并进一步作为输入馈送到预训练的MT模块;然后使用合成创建的ST数据对整个管道进行端到端的口语翻译训练。我们还发布了一个新的评估基准代码切换孟加拉语英语,印地语英语,马拉地语英语和泰卢固语英语语音到英语文本。COSTA显著优于许多具有竞争力的级联和端到端多模式基线,最高可达3.5个BLEU点。摘要:Code-switching is a widely prevalent linguistic phenomenon in multilingual societies like India. Building speech-to-text models for code-switched speech is challenging due to limited availability of datasets. In this work, we focus on the problem of spoken translation (ST) of code-switched speech in Indian languages to English text. We present a new end-to-end model architecture COSTA that scaffolds on pretrained automatic speech recognition (ASR) and machine translation (MT) modules (that are more widely available for many languages). Speech and ASR text representations are fused using an aligned interleaving scheme and are fed further as input to a pretrained MT module; the whole pipeline is then trained end-to-end for spoken translation using synthetically created ST data. We also release a new evaluation benchmark for code-switched Bengali-English, Hindi-English, Marathi-English and Telugu- English speech to English text. COSTA significantly outperforms many competitive cascaded and end-to-end multimodal baselines by up to 3.5 BLEU points.
【12】 Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
标题: 用于时间控制文本到音乐生成的联合音频和符号条件处理
作者:Or Tal,Alon Ziv,Itai Gat,Felix Kreuk,Yossi Adi
链接:点击下载PDF文件
摘要:我们提出了JASCO,一个时间控制的文本到音乐生成模型,利用符号和基于音频的条件。JASCO可以生成高质量的音乐样本,条件是全局文本描述以及细粒度的本地控件。JASCO是基于流匹配建模范式与一种新的空调方法。这允许本地控制的音乐生成(例如,和弦)和全局(文本描述)。具体来说,我们应用信息瓶颈层结合时间模糊提取相关信息的特定控件。这允许在相同的文本到音乐模型中结合符号和基于音频的条件。我们用各种符号控制信号进行实验(例如,和弦,旋律),以及音频表示(例如,分离的鼓轨道,全混合)。我们评估JASCO同时考虑发电质量和条件的坚持,使用客观指标和人体研究。结果表明,JASCO是可比的,考虑到生成质量的评估基线,同时允许显着更好,更灵活的控制所生成的音乐。样品可在我们的演示页面https: pages.cs.huji.ac.il adiyoss-lab JASCO上获得。摘要:We present JASCO, a temporally controlled text-to-music generation model utilizing both symbolic and audio-based conditions. JASCO can generate high-quality music samples conditioned on global text descriptions along with fine-grained local controls. JASCO is based on the Flow Matching modeling paradigm together with a novel conditioning method. This allows music generation controlled both locally (e.g., chords) and globally (text description). Specifically, we apply information bottleneck layers in conjunction with temporal blurring to extract relevant information with respect to specific controls. This allows the incorporation of both symbolic and audio-based conditions in the same text-to-music model. We experiment with various symbolic control signals (e.g., chords, melody), as well as with audio representations (e.g., separated drum tracks, full-mix). We evaluate JASCO considering both generation quality and condition adherence, using both objective metrics and human studies. Results suggest that JASCO is comparable to the evaluated baselines considering generation quality while allowing significantly better and more versatile controls over the generated music. Samples are available on our demo page https: pages.cs.huji.ac.il adiyoss-lab JASCO.
【13】 Robust Channel Learning for Large-Scale Radio Speaker Verification
标题: 用于大规模无线电扬声器验证的稳健通道学习
作者:Wenhao Yang,Jianguo Wei,Wenhuan Lu,Lei Li,Xugang Lu
备注:12 pages, 11 figures
链接:点击下载PDF文件
摘要:说话人确认的研究越来越多地集中在具有挑战性的信道条件和噪声环境下实现鲁棒和可靠的识别。在无线电通信中识别说话者是特别困难的,这是由于固有的限制,如有限的带宽和普遍的噪声干扰。为了解决这个问题,我们提出了一个通道鲁棒说话人学习(CRSL)框架,提高了当前说话人验证管道的鲁棒性,考虑到数据源,数据增强和模型传输过程的效率。我们的框架引入了一个增强模块,通过操纵训练输入的带宽来减轻无线电语音数据集的带宽变化。它还通过在流形空间内引入噪声来解决未知噪声。此外,我们提出了一种有效的微调方法,减少了对大量额外训练时间和大量数据的需求。此外,我们开发了一个工具包,用于组装一个大规模的无线电语音语料库,并建立了一个专门为无线电场景说话人验证研究量身定制的基准。实验结果表明,我们提出的方法有效地提高了性能,并减轻在说话人确认任务中无线电传输所造成的退化。代码将在Github上提供。摘要:Recent research in speaker verification has increasingly focused on achieving robust and reliable recognition under challenging channel conditions and noisy environments. Identifying speakers in radio communications is particularly difficult due to inherent limitations such as constrained bandwidth and pervasive noise interference. To address this issue, we present a Channel Robust Speaker Learning (CRSL) framework that enhances the robustness of the current speaker verification pipeline, considering data source, data augmentation, and the efficiency of model transfer processes. Our framework introduces an augmentation module that mitigates bandwidth variations in radio speech datasets by manipulating the bandwidth of training inputs. It also addresses unknown noise by introducing noise within the manifold space. Additionally, we propose an efficient fine-tuning method that reduces the need for extensive additional training time and large amounts of data. Moreover, we develop a toolkit for assembling a large-scale radio speech corpus and establish a benchmark specifically tailored for radio scenario speaker verification studies. Experimental results demonstrate that our proposed methodology effectively enhances performance and mitigates degradation caused by radio transmission in speaker verification tasks. The code will be available on Github.
【14】 Imperceptible Rhythm Backdoor Attacks: Exploring Rhythm Transformation for Embedding Undetectable Vulnerabilities on Speech Recognition
标题: 不可感知的节奏后门攻击:探索节奏转换以在语音识别中嵌入不可检测的漏洞
作者:Wenhan Yao,Jiangkun Yang,Yongqiang He,Jia Liu,Weiping Wen
链接:点击下载PDF文件
摘要:语音识别是人机交互的重要起点,最近,深度学习模型在这项任务中取得了巨大的成功。然而,当模型训练和私有数据提供者总是分离时,一些使深度神经网络(DNN)异常的安全威胁值得研究。近年来,语音识别系统中的典型后门攻击已成为研究热点。现有的后门方法是基于数据中毒。攻击者将一些合并的变化添加到良性语音频谱图或改变语音成分,如音高和音色。因此,中毒数据可以通过人类听觉或自动深度算法检测到。为了提高数据中毒的隐蔽性,本文提出了一种非神经网络的快速算法--随机谱图节奏变换(RSRT)。该算法结合了四个步骤来生成隐形有毒话语。从节奏成分转换的角度来看,我们提出的触发拉伸或挤压梅尔频谱图,并恢复他们回到信号。该操作保持音色和内容不变,具有良好的隐蔽性。我们的实验是在两种语音识别任务上进行的,包括通过说话人确认和自动语音识别来测试中毒样本的隐蔽性。实验结果表明,该方法具有良好的有效性和隐蔽性。节奏触发需要低中毒率,并获得非常高的攻击成功率。摘要:Speech recognition is an essential start ring of human-computer interaction, and recently, deep learning models have achieved excellent success in this task. However, when the model training and private data provider are always separated, some security threats that make deep neural networks (DNNs) abnormal deserve to be researched. In recent years, the typical backdoor attacks have been researched in speech recognition systems. The existing backdoor methods are based on data poisoning. The attacker adds some incorporated changes to benign speech spectrograms or changes the speech components, such as pitch and timbre. As a result, the poisoned data can be detected by human hearing or automatic deep algorithms. To improve the stealthiness of data poisoning, we propose a non-neural and fast algorithm called Random Spectrogram Rhythm Transformation (RSRT) in this paper. The algorithm combines four steps to generate stealthy poisoned utterances. From the perspective of rhythm component transformation, our proposed trigger stretches or squeezes the mel spectrograms and recovers them back to signals. The operation keeps timbre and content unchanged for good stealthiness. Our experiments are conducted on two kinds of speech recognition tasks, including testing the stealthiness of poisoned samples by speaker verification and automatic speech recognition. The results show that our method has excellent effectiveness and stealthiness. The rhythm trigger needs a low poisoning rate and gets a very high attack success rate.
【15】 SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
标题: SingMOS:用于MOS预测的广泛开源歌唱声音数据集
作者:Yuxun Tang,Jiatong Shi,Yuning Wu,Qin Jin
链接:点击下载PDF文件
摘要:在语音生成任务中,人的主观评分,通常被称为意见分数,被认为是语音质量评估的“金标准”,平均意见分数(MOS)作为主要的评估指标。由于人工注释的高成本,在语音领域出现了几种MOS预测系统,表现出良好的性能。这些MOS预测模型使用来自先前语音相关挑战的注释进行训练。然而,与语音领域相比,歌唱领域面临着数据稀缺和更严格的版权保护,导致缺乏高质量的MOS注释歌唱数据集。为了解决这个问题,我们提出了SingMOS,这是一个高质量和多样化的MOS歌唱数据集,涵盖了一系列中国和日本的数据集。这些合成的人声是使用歌唱合成、转换或再合成任务中最先进的模型生成的,并由专业注释者与真实人声一起进行评级。数据分析证明了我们数据集的多样性和可靠性。此外,我们还对SingMOS进行了进一步的探索,为SingMOS的预测提供了见解,并为SingMOS的持续扩展提供了指导。摘要:In speech generation tasks, human subjective ratings, usually referred to as the opinion score, are considered the "gold standard" for speech quality evaluation, with the mean opinion score (MOS) serving as the primary evaluation metric. Due to the high cost of human annotation, several MOS prediction systems have emerged in the speech domain, demonstrating good performance. These MOS prediction models are trained using annotations from previous speech-related challenges. However, compared to the speech domain, the singing domain faces data scarcity and stricter copyright protections, leading to a lack of high-quality MOS-annotated datasets for singing. To address this, we propose SingMOS, a high-quality and diverse MOS dataset for singing, covering a range of Chinese and Japanese datasets. These synthesized vocals are generated using state-of-the-art models in singing synthesis, conversion, or resynthesis tasks and are rated by professional annotators alongside real vocals. Data analysis demonstrates the diversity and reliability of our dataset. Additionally, we conduct further exploration on SingMOS, providing insights for singing MOS prediction and guidance for the continued expansion of SingMOS.
【16】 Optimizing Automatic Speech Assessment: W-RankSim Regularization and Hybrid Feature Fusion Strategies
标题: 优化自动语音评估:W-RankSim正规化和混合特征融合策略
作者:Chung-Wen Wu,Berlin Chen
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:在最近的研究中,自动语音评估(ASA)随着自监督特征(SSL)的利用而取得了显着的进步。然而,ASA的一个关键挑战在于数据的不平衡分布,特别是在英语测试数据集中。为了解决这一挑战,我们将ASA作为一个有序分类任务,引入加权向量排序相似性(W-RankSim)作为一种新的正则化技术。W-RankSim鼓励输出层中相似类的加权向量更接近,这意味着具有相似标签的特征向量将随着它们向相应的加权向量收敛而逐渐相互靠近。广泛的实验评估证实了我们的方法在提高ASA的顺序分类性能的有效性。此外,我们提出了一个混合模型,结合SSL和手工制作的功能,展示如何包含手工制作的功能,提高性能的ASA系统。摘要:Automatic Speech Assessment (ASA) has seen notable advancements with the utilization of self-supervised features (SSL) in recent research. However, a key challenge in ASA lies in the imbalanced distribution of data, particularly evident in English test datasets. To address this challenge, we approach ASA as an ordinal classification task, introducing Weighted Vectors Ranking Similarity (W-RankSim) as a novel regularization technique. W-RankSim encourages closer proximity of weighted vectors in the output layer for similar classes, implying that feature vectors with similar labels would be gradually nudged closer to each other as they converge towards corresponding weighted vectors. Extensive experimental evaluations confirm the effectiveness of our approach in improving ordinal classification performance for ASA. Furthermore, we propose a hybrid model that combines SSL and handcrafted features, showcasing how the inclusion of handcrafted features enhances performance in an ASA system.
【17】 Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
标题: 使用CNN的语音情感识别及其在数字医疗保健中的用例
作者:Nishargo Nigar
备注:Master's Thesis at Hamburg University of Technology
链接:点击下载PDF文件
摘要:从语音中识别人类情感和情感状态的过程被称为语音情感识别(SER)。这是基于这样的观察,即声音中的音调和音高经常传达潜在的情感。语音识别包括识别情感的能力,这变得越来越流行,需求量也越来越大。在适当因素的帮助下(如模式,情绪,强度,重复等)基于数据中发现的情感,我的研究试图使用卷积神经网络(CNN)来区分音频记录中的情感,并根据不同情感的范围对其进行标记。我开发了一个机器学习模型,可以借助机器学习方法从提供的音频文件中识别情绪。评估主要集中在精度,召回率和F1分数,这些都是常见的机器学习指标。为了正确地建立和训练机器学习框架,主要目标是研究所有输入和输出参数的影响和相互关系。为了提高识别意图的能力,这是沟通的一个关键条件,我使用我的专业机器学习算法通过语音来评估情绪,该算法将在数字医疗保健的帮助下从语音中解决情绪状态,弥合人类和人工智能(AI)之间的差距。摘要:The process of identifying human emotion and affective states from speech is known as speech emotion recognition (SER). This is based on the observation that tone and pitch in the voice frequently convey underlying emotion. Speech recognition includes the ability to recognize emotions, which is becoming increasingly popular and in high demand. With the help of appropriate factors (such modalities, emotions, intensities, repetitions, etc.) found in the data, my research seeks to use the Convolutional Neural Network (CNN) to distinguish emotions from audio recordings and label them in accordance with the range of different emotions. I have developed a machine learning model to identify emotions from supplied audio files with the aid of machine learning methods. The evaluation is mostly focused on precision, recall, and F1 score, which are common machine learning metrics. To properly set up and train the machine learning framework, the main objective is to investigate the influence and cross-relation of all input and output parameters. To improve the ability to recognize intentions, a key condition for communication, I have evaluated emotions using my specialized machine learning algorithm via voice that would address the emotional state from voice with the help of digital healthcare, bridging the gap between human and artificial intelligence (AI).
【18】 How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
标题: 我们应该如何从自我监督模型中提取离散音频令牌?
作者:Pooneh Mousavi,Jarod Duret,Salah Zaiem,Luca Della Libera,Artem Ploujnikov,Cem Subakan,Mirco Ravanelli
备注:4 pages, 2 figures, 2 tables, Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:离散音频令牌最近因其在弥合音频和语言处理之间的差距方面的潜力而受到关注。理想的音频令牌必须保留内容、非语言元素、说话者身份和许多其他音频细节。当前的音频标记化方法分为两类:通过自监督学习(SSL)模型的量化获得的语义标记,以及基于神经压缩的标记(编解码器)。虽然以前的研究已经对编解码器模型进行了基准测试以确定最佳配置,但量化预训练SSL模型的理想设置仍然不清楚。本文探讨了在区分和生成任务中语义标记的最佳配置。我们提出了一个可扩展的解决方案,训练跨多个SSL层的通用声码器。此外,注意力机制被用来识别特定于任务的影响层,增强了语义令牌在不同音频应用中的适应性和性能。摘要:Discrete audio tokens have recently gained attention for their potential to bridge the gap between audio and language processing. Ideal audio tokens must preserve content, paralinguistic elements, speaker identity, and many other audio details. Current audio tokenization methods fall into two categories: Semantic tokens, acquired through quantization of Self-Supervised Learning (SSL) models, and Neural compression-based tokens (codecs). Although previous studies have benchmarked codec models to identify optimal configurations, the ideal setup for quantizing pretrained SSL models remains unclear. This paper explores the optimal configuration of semantic tokens across discriminative and generative tasks. We propose a scalable solution to train a universal vocoder across multiple SSL layers. Furthermore, an attention mechanism is employed to identify task-specific influential layers, enhancing the adaptability and performance of semantic tokens in diverse audio applications.
【19】 Improving child speech recognition with augmented child-like speech
标题: 通过增强的儿童语音来提高儿童语音识别
作者:Yuanyuan Zhang,Zhengjun Yue,Tanvina Patel,Odette Scharenborg
备注:5 pages, 1 figure Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:最先进的ASR显示出儿童语音的次优性能。儿童语言的缺乏限制了儿童语音识别的发展。因此,我们研究了儿童到儿童的语音转换(VC),从现有的儿童扬声器的数据集和额外的(新的)儿童扬声器通过单语和跨语言(荷兰语到德语)VC,分别。结果表明,跨语言儿童对儿童VC显着提高儿童ASR性能。关于儿童对儿童跨语言VC生成的数据量对微调(FT)ASR模型的影响的实验给出了最佳结果,其中我们的FT-Conformer模型和FT-Whisper模型的两倍增强与基线相比减少了约3%的绝对WER,并且从头开始训练的模型的六倍增强提高了绝对3.6% WER。此外,使用少量的“高质量”VC生成的数据,获得了与我们最好的FT模型相似的结果。摘要:State-of-the-art ASRs show suboptimal performance for child speech. The scarcity of child speech limits the development of child speech recognition (CSR). Therefore, we studied child-to-child voice conversion (VC) from existing child speakers in the dataset and additional (new) child speakers via monolingual and cross-lingual (Dutch-to-German) VC, respectively. The results showed that cross-lingual child-to-child VC significantly improved child ASR performance. Experiments on the impact of the quantity of child-to-child cross-lingual VC-generated data on fine-tuning (FT) ASR models gave the best results with two-fold augmentation for our FT-Conformer model and FT-Whisper model which reduced WERs with ~3% absolute compared to the baseline, and with six-fold augmentation for the model trained from scratch, which improved by an absolute 3.6% WER. Moreover, using a small amount of "high-quality" VC-generated data achieved similar results to those of our best-FT models.
【20】 Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
标题: 精心合并预训练语音模型中的隐藏嵌入以实现反欺骗检测
作者:Zihan Pan,Tianchi Liu,Hardik B. Sailor,Qiongqiong Wang
链接:点击下载PDF文件
摘要:在大型语音语料库上训练的自监督学习(SSL)语音表示模型已经证明了通过多个Transformer层提取分层语音嵌入的有效性。然而,这些嵌入在特定任务中的行为仍然不确定。本文研究了WavLM模型在反欺骗中的多层行为,并提出了一种注意的合并方法来利用层次隐藏嵌入。结果证明了微调WavLM的可行性,以在ASVspoof 2019LA,2021LA和2021DF评估集上分别实现0.65%,3.50%和3.19%的最佳等误率(EER)。值得注意的是,我们发现WavLM大模型的早期隐藏Transformer层对反欺骗任务有显著贡献,通过利用部分预训练模型实现计算效率。摘要:Self-supervised learning (SSL) speech representation models, trained on large speech corpora, have demonstrated effectiveness in extracting hierarchical speech embeddings through multiple transformer layers. However, the behavior of these embeddings in specific tasks remains uncertain. This paper investigates the multi-layer behavior of the WavLM model in anti-spoofing and proposes an attentive merging method to leverage the hierarchical hidden embeddings. Results demonstrate the feasibility of fine-tuning WavLM to achieve the best equal error rate (EER) of 0.65%, 3.50%, and 3.19% on the ASVspoof 2019LA, 2021LA, and 2021DF evaluation sets, respectively. Notably, We find that the early hidden transformer layers of the WavLM large model contribute significantly to anti-spoofing task, enabling computational efficiency by utilizing a partial pre-trained model.
【21】 Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
标题: 模糊不可知多对一端到端语音翻译的软语言识别
作者:Peidong Wang,Jian Xue,Jinyu Li,Junkun Chen,Aswin Shanmugam Subramanian
链接:点击下载PDF文件
摘要:语音不可知的多对一端到端语音翻译模型可以将来自不同源语言的音频信号转换为目标语言的文本。这些模型不需要源语言识别,这提高了用户体验。在某些情况下,输入语言可以被给定或估计。我们的目标是使用这些额外的语言信息,同时保持其他语言的质量。我们通过引入一个简单有效的线性输入网络来实现这一点。线性输入网络被初始化为单位矩阵,这确保模型可以与原始模型一样好,甚至更好。实验结果表明,该方法可以成功地增强指定的语言,同时保持语言无关的能力的多对一ST模型。摘要:Language-agnostic many-to-one end-to-end speech translation models can convert audio signals from different source languages into text in a target language. These models do not need source language identification, which improves user experience. In some cases, the input language can be given or estimated. Our goal is to use this additional language information while preserving the quality of the other languages. We accomplish this by introducing a simple and effective linear input network. The linear input network is initialized as an identity matrix, which ensures that the model can perform as well as, or better than, the original model. Experimental results show that the proposed method can successfully enhance the specified language, while keeping the language-agnostic ability of the many-to-one ST models.
【22】 Connected Speech-Based Cognitive Assessment in Chinese and English
标题: 中文和英语的连接言语认知评估
作者:aturnino Luz,Sofia De La Fuente Garcia,Fasih Haider,Davida Fromm,Brian MacWhinney,Alyssa Lanzi,Ya-Ning Chang,Chia-Ju Chou,Yi-Chien Liu
备注:To appear in Proceedings of Interspeech 2024
链接:点击下载PDF文件
摘要:我们提出了一种新的基准数据集和预测任务,用于研究通过分析连接语音来评估认知功能的方法。该数据集包括具有不同程度认知障碍的汉语普通话和英语使用者以及具有正常认知的个体的语音样本和临床信息。这些数据已通过倾向评分分析按年龄和性别仔细匹配,以确保模型训练的平衡性和代表性。预测任务包括轻度认知障碍诊断和认知测试分数预测。该框架旨在鼓励开发基于语音的认知评估方法,这些方法可以概括各种语言。我们通过提出基线预测模型来说明这一点,这些模型采用与语言无关的和可比较的特征来进行诊断和认知测试分数预测。模型在诊断中的平均召回率为59.2%,评分预测的均方根误差为2.89。摘要:We present a novel benchmark dataset and prediction tasks for investigating approaches to assess cognitive function through analysis of connected speech. The dataset consists of speech samples and clinical information for speakers of Mandarin Chinese and English with different levels of cognitive impairment as well as individuals with normal cognition. These data have been carefully matched by age and sex by propensity score analysis to ensure balance and representativity in model training. The prediction tasks encompass mild cognitive impairment diagnosis and cognitive test score prediction. This framework was designed to encourage the development of approaches to speech-based cognitive assessment which generalise across languages. We illustrate it by presenting baseline prediction models that employ language-agnostic and comparable features for diagnosis and cognitive test score prediction. The models achieved unweighted average recall was 59.2% in diagnosis, and root mean squared error of 2.89 in score prediction.
【23】 Towards Signal Processing In Large Language Models
标题: 走向大型语言模型中的信号处理
作者:Prateek Verma,Mert Pilanci
备注:12 pages, 3 figures
链接:点击下载PDF文件
摘要:本文介绍了在大型语言模型(LLM)中应用信号处理的思想。随着最近生成式人工智能的爆发,我们的工作可以帮助将两个领域连接在一起,即信号处理领域和大型语言模型。我们绘制了经典的傅里叶变换和傅里叶变换式的可学习的时间-频率表示的LLM的每一个中间激活信号之间的平行。一旦我们将令牌上的每个激活信号分解为时频表示,我们就可以学习如何过滤和重建它们,从头开始学习所有组件,以预测给定先前上下文的下一个令牌。我们表明,对于类似GPT的架构,我们的工作实现了更快的收敛,并通过在相同时期的训练中添加少量的额外参数来显着提高性能。我们希望这项工作为算法探索LLM等神经架构中发现的信号内部的信号处理铺平道路。摘要:This paper introduces the idea of applying signal processing inside a Large Language Model (LLM). With the recent explosion of generative AI, our work can help bridge two fields together, namely the field of signal processing and large language models. We draw parallels between classical Fourier-Transforms and Fourier Transform-like learnable time-frequency representations for every intermediate activation signal of an LLM. Once we decompose every activation signal across tokens into a time-frequency representation, we learn how to filter and reconstruct them, with all components learned from scratch, to predict the next token given the previous context. We show that for GPT-like architectures, our work achieves faster convergence and significantly increases performance by adding a minuscule number of extra parameters when trained for the same epochs. We hope this work paves the way for algorithms exploring signal processing inside the signals found in neural architectures like LLMs and beyond.
【24】 GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
标题: GigaSpeech 2:一个不断发展的、大规模和多领域的SVR数据库,用于低资源语言,具有自动抓取、转录和细化
作者:Yifan Yang,Zheshu Song,Jianheng Zhuo,Mingyu Cui,Jinpeng Li,Bo Yang,Yexing Du,Ziyang Ma,Xunying Liu,Ziyuan Wang,Ke Li,Shuai Fan,Kai Yu,Wei-Qiang Zhang,Guoguo Chen,Xie Chen
备注:Under review
链接:点击下载PDF文件
摘要:语音技术的发展受到数据集规模快速增长的推动。传统的语音模型通常依赖于大量的标记训练数据,这对于低资源语言来说是稀缺的。本文介绍了GigaSpeech 2,一个大规模,多领域,多语种的语音识别语料库。它是为低资源语言设计的,不依赖于成对的语音和文本数据。GigaSpeech 2包含约30,000小时的自动转录语音,包括泰国语,印度尼西亚语和越南语,这些语音来自未标记的YouTube视频。我们还引入了一个自动化的管道,用于数据抓取,转录和标签优化。具体来说,该管道使用Whisper进行初始转录,使用TorchAudio进行强制对齐,并结合多维过滤来保证数据质量。改进的Noisy Student训练方法进一步迭代地改进有缺陷的伪标签,从而提高模型性能。我们的人工转录的评价集和两个公共测试集,从普通的声音和FLEURS的实验结果证实了我们的语料库的高质量和广泛的适用性。值得注意的是,在GigaSpeech 2上训练的ASR模型可以在我们具有挑战性和现实性的YouTube测试集上将泰国语,印度尼西亚语和越南语的单词错误率降低25%至40%,而Whisper大v3模型只有10%的模型参数。此外,与商业服务相比,我们在Gigaspeech 2上训练的ASR模型具有卓越的性能。我们相信,我们新引入的语料库和管道将为低资源语音识别开辟一条新的途径,并大大促进这一领域的研究。摘要:The evolution of speech technology has been spurred by the rapid increase in dataset sizes. Traditional speech models generally depend on a large amount of labeled training data, which is scarce for low-resource languages. This paper presents GigaSpeech 2, a large-scale, multi-domain, multilingual speech recognition corpus. It is designed for low-resource languages and does not rely on paired speech and text data. GigaSpeech 2 comprises about 30,000 hours of automatically transcribed speech, including Thai, Indonesian, and Vietnamese, gathered from unlabeled YouTube videos. We also introduce an automated pipeline for data crawling, transcription, and label refinement. Specifically, this pipeline uses Whisper for initial transcription and TorchAudio for forced alignment, combined with multi-dimensional filtering for data quality assurance. A modified Noisy Student Training is developed to further refine flawed pseudo labels iteratively, thus enhancing model performance. Experimental results on our manually transcribed evaluation set and two public test sets from Common Voice and FLEURS confirm our corpus's high quality and broad applicability. Notably, ASR models trained on GigaSpeech 2 can reduce the word error rate for Thai, Indonesian, and Vietnamese on our challenging and realistic YouTube test set by 25% to 40% compared to the Whisper large-v3 model, with merely 10% model parameters. Furthermore, our ASR models trained on Gigaspeech 2 yield superior performance compared to commercial services. We believe that our newly introduced corpus and pipeline will open a new avenue for low-resource speech recognition and significantly facilitate research in this area.
【25】 DiTTo-TTS: Efficient and Scalable Zero-Shot Text-to-Speech with Diffusion Transformer
标题: DiTTo-TTC:具有扩散Transformer的高效且可扩展的Zero-Shot文本到语音
作者:Keon Lee,Dong Won Kim,Jaehyeon Kim,Jaewoong Cho
链接:点击下载PDF文件
摘要:大规模扩散模型在包括图像、视频和音频在内的多种模态中表现出出色的生成能力。然而,文本到语音(TTS)系统通常涉及域特定建模因素(例如,音素和音素级持续时间),以确保文本和语音之间的精确时间对齐,这阻碍了TTS扩散模型的效率和可扩展性。在这项工作中,我们提出了一个有效的和可扩展的扩散Transformer(DiT),利用现成的预先训练的文本和语音编码器。我们的方法解决了文本语音对齐的挑战,通过交叉注意机制与语音表示的总长度的预测。为了实现这一目标,我们增强了DiT架构,以适应TTS和提高对齐,将语义指导到语音的潜在空间。我们将训练数据集和模型大小分别扩展到82K小时和790M参数。我们广泛的实验表明,大规模的扩散模型的TTS没有特定领域的建模,不仅简化了训练管道,但也产生优越的或可比的zero-shot性能,以国家的最先进的TTS模型的自然度,可理解性和说话人相似性。我们的语音样本可以在https: ditto-tts.github.io上找到。摘要:Large-scale diffusion models have shown outstanding generative abilities across multiple modalities including images, videos, and audio. However, text-to-speech (TTS) systems typically involve domain-specific modeling factors (e.g., phonemes and phoneme-level durations) to ensure precise temporal alignments between text and speech, which hinders the efficiency and scalability of diffusion models for TTS. In this work, we present an efficient and scalable Diffusion Transformer (DiT) that utilizes off-the-shelf pre-trained text and speech encoders. Our approach addresses the challenge of text-speech alignment via cross-attention mechanisms with the prediction of the total length of speech representations. To achieve this, we enhance the DiT architecture to suit TTS and improve the alignment by incorporating semantic guidance into the latent space of speech. We scale the training dataset and the model size to 82K hours and 790M parameters, respectively. Our extensive experiments demonstrate that the large-scale diffusion model for TTS without domain-specific modeling not only simplifies the training pipeline but also yields superior or comparable zero-shot performance to state-of-the-art TTS models in terms of naturalness, intelligibility, and speaker similarity. Our speech samples are available at https: ditto-tts.github.io.
【26】 Performance Improvement of Language-Queried Audio Source Separation Based on Caption Augmentation From Large Language Models for DCASE Challenge 2024 Task 9
标题: 基于来自大型语言模型的字幕增强的数字查询音频源分离的性能改进DUSE Challenge 2024任务9
作者:Do Hyun Lee,Yoonah Song,Hong Kook Kim
备注:DCASE 2024 Challenge Task 9, 4 pages
链接:点击下载PDF文件
摘要:我们提出了一个基于工程的文本增强方法应用于语言查询的音频源分离(LASS)任务。为了提高LASS的性能,所提出的方法利用大语言模型(LLM)来生成对应于训练数据集的每个句子的多个字幕。为此,我们首先进行实验,以确定最有效的提示字幕扩增与较少的标题。使用这些增强字幕训练的LASS模型在DCASE 2024任务9验证集上表现出与未经增强训练的LASS模型相比的改进性能。这项研究强调了基于LLM的字幕增强在推进语言查询音频源分离方面的有效性。摘要:We present a prompt-engineering-based text-augmentation approach applied to a language-queried audio source separation (LASS) task. To enhance the performance of LASS, the proposed approach utilizes large language models (LLMs) to generate multiple captions corresponding to each sentence of the training dataset. To this end, we first perform experiments to identify the most effective prompts for caption augmentation with a smaller number of captions. A LASS model trained with these augmented captions demonstrates improved performance on the DCASE 2024 Task 9 validation set compared to that trained without augmentation. This study highlights the effectiveness of LLM-based caption augmentation in advancing language-queried audio source separation.
【27】 Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
标题: 自蒸馏原型网络:在没有监督的情况下学习稳健的说话者表示
作者:Yafeng Chen,Siqi Zheng,Hui Wang,Luyao Cheng,Qian Chen,Shiliang Zhang,Wen Wang
链接:点击下载PDF文件
摘要:在没有明确说话人标签的情况下训练说话人区分和鲁棒的说话人确认系统仍然是一个持续的挑战。在本文中,我们提出了一种新的自监督说话人确认方法,自蒸馏原型网络(SDPN),它有效地促进了自监督说话人表示学习。SDPN将话语的增强视图的表示分配给与原始视图的表示相同的原型,从而实现视图之间的有效知识转移。最初,由于SDPN训练过程中缺少负对,网络往往会在嵌入空间中非常紧密地对齐正对,这种现象称为模型崩溃。为了缓解这个问题,我们在SDPN中的嵌入中引入了多样性正则化项。在VoxCeleb数据集上的综合实验证明了SDPN在自监督说话人确认中的优越性。SDPN在VoxCeleb 1说话者验证评估基准上设定了一个新的最先进水平,在训练中不使用任何说话者标签的情况下,VoxCeleb 1-O、VoxCeleb 1-E和VoxCeleb 1-H的试验分别实现了1.80%、1.99%和3.62%的等错误率。摘要:Training speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persisting challenge. In this paper, we propose a new self-supervised speaker verification approach, Self-Distillation Prototypes Network (SDPN), which effectively facilitates self-supervised speaker representation learning. SDPN assigns the representation of the augmented views of an utterance to the same prototypes as the representation of the original view, thereby enabling effective knowledge transfer between the views. Originally, due to the lack of negative pairs in the SDPN training process, the network tends to align positive pairs very closely in the embedding space, a phenomenon known as model collapse. To alleviate this problem, we introduce a diversity regularization term to embeddings in SDPN. Comprehensive experiments on the VoxCeleb datasets demonstrate the superiority of SDPN in self-supervised speaker verification. SDPN sets a new state-of-the-art on the VoxCeleb1 speaker verification evaluation benchmark, achieving Equal Error Rate 1.80%, 1.99%, and 3.62% for trial VoxCeleb1-O, VoxCeleb1-E and VoxCeleb1-H respectively, without using any speaker labels in training.
【28】 Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
标题: 含噪语音端到端语音识别的连续测试时间自适应
作者:Guan-Ting Lin,Wei-Ping Huang,Hung-yi Lee
备注:13 pages
链接:点击下载PDF文件
摘要:基于深度学习的端到端自动语音识别(ASR)已经取得了重大进展,但由于现实场景中的域转移,在域外(OOD)样本上的性能仍然很差。测试时自适应(TTA)方法通过在推理时使用测试样本来适应模型来解决这个问题。然而,目前的ASR TTA方法主要集中在非连续TTA,这限制了跨样本知识学习相比,连续TTA。在这项工作中,我们提出了一个快-慢TTA框架的ASR,它利用了连续和非连续TTA的优势。在这个框架内,我们介绍了动态SUTA(DSUTA),一个基于熵最小化的连续TTA方法的ASR。为了增强DSUTA对时变数据的鲁棒性,我们提出了一种动态重置策略,可以自动检测域偏移并重置模型,使其在处理多域数据时更有效。我们的方法在各种嘈杂的ASR数据集上表现出卓越的性能,优于非连续和连续的TTA基线,同时保持对域变化的鲁棒性,而不需要域边界信息。摘要:Deep learning-based end-to-end automatic speech recognition (ASR) has made significant strides but still struggles with performance on out-of-domain (OOD) samples due to domain shifts in real-world scenarios. Test-Time Adaptation (TTA) methods address this issue by adapting models using test samples at inference time. However, current ASR TTA methods have largely focused on non-continual TTA, which limits cross-sample knowledge learning compared to continual TTA. In this work, we propose a Fast-slow TTA framework for ASR, which leverages the advantage of continual and non-continual TTA. Within this framework, we introduce Dynamic SUTA (DSUTA), an entropy-minimization-based continual TTA method for ASR. To enhance DSUTA's robustness on time-varying data, we propose a dynamic reset strategy that automatically detects domain shifts and resets the model, making it more effective at handling multi-domain data. Our method demonstrates superior performance on various noisy ASR datasets, outperforming both non-continual and continual TTA baselines while maintaining robustness to domain changes without requiring domain boundary information.
【29】 Multi-Scale Accent Modeling with Disentangling for Multi-Speaker Multi-Accent TTS Synthesis
标题: 多说话人多口音TTC合成的多尺度口音建模
作者:Xuehao Zhou,Mingyang Zhang,Yi Zhou,Zhizheng Wu,Haizhou Li
链接:点击下载PDF文件
摘要:在保留说话者身份的同时,跨不同口音合成语音对于各种现实世界的客户应用至关重要。然而,由于口音变化的复杂性以及口音和说话人身份之间的内在纠缠,在文本到语音(TTS)系统中对口音和说话人的个体和准确建模是具有挑战性的。在本文中,我们提出了一种新的方法,多说话人多口音的TTS合成,其目的是合成多个说话人的声音,每个不同的口音。我们提出的方法采用了多尺度口音建模策略,以解决口音的变化在不同的水平。具体来说,我们引入全球(话语级)和本地(音素级)口音建模,监督个别口音分类器,以捕捉重音话语和细粒度的音素之间的变化,分别在整体上的变化。为了分别控制口音和说话人,说话人无关的口音建模是必要的,这是通过使用说话人分类器进行对抗训练来实现的,以在多尺度口音建模中解开说话人身份。因此,我们获得了说话人无关和口音歧视的多尺度嵌入作为综合口音特征。此外,我们提出了一个本地口音预测模型,允许直接从音素输入生成口音语音。在带口音的英语语音语料库上进行了大量的实验。客观和主观的评价表明,我们提出的系统相比,基线系统的优越性。详细的成分分析表明,全局和局部口音建模,说话人解纠缠的多说话人多口音语音合成的有效性。摘要:Synthesizing speech across different accents while preserving the speaker identity is essential for various real-world customer applications. However, the individual and accurate modeling of accents and speakers in a text-to-speech (TTS) system is challenging due to the complexity of accent variations and the intrinsic entanglement between the accent and speaker identity. In this paper, we present a novel approach for multi-speaker multi-accent TTS synthesis, which aims to synthesize voices of multiple speakers, each with various accents. Our proposed approach employs a multi-scale accent modeling strategy to address accent variations at different levels. Specifically, we introduce both global (utterance level) and local (phoneme level) accent modeling, supervised by individual accent classifiers to capture the overall variation within accented utterances and fine-grained variations between phonemes, respectively. To control accents and speakers separately, speaker-independent accent modeling is necessary, which is achieved by adversarial training with speaker classifiers to disentangle speaker identity within the multi-scale accent modeling. Consequently, we obtain speaker-independent and accent-discriminative multi-scale embeddings as comprehensive accent features. Additionally, we propose a local accent prediction model that allows to generate accented speech directly from phoneme inputs. Extensive experiments are conducted on an accented English speech corpus. Both objective and subjective evaluations show the superiority of our proposed system compared to baselines systems. Detailed component analysis demonstrates the effectiveness of global and local accent modeling, and speaker disentanglement on multi-speaker multi-accent speech synthesis.
【30】 Revisiting and Improving Scoring Fusion for Spoofing-aware Speaker Verification Using Compositional Data Analysis
标题: 使用合成数据分析重新审视和改进具有欺骗意识的说话人验证的评分融合
作者:Xin Wang,Tomi Kinnunen,Kong Aik Lee,Paul-Gauthier Noé,Junichi Yamagishi
备注:Interspeech 2024 Accepted. this https URL
链接:点击下载PDF文件
摘要:融合自动说话人验证(ASV)和欺骗对策(CM)的输出,预计将使一个集成的系统强大的零努力冒名顶替者和合成欺骗攻击。已经提出了许多分数级融合方法,但许多仍然是启发式的。本文回顾了得分级融合决策理论的工具,并提出了三个主要的发现。首先,通过对ASV和CM分数求和的融合可以在成分数据分析的基础上进行解释,并且融合之前的分数校准是必不可少的。其次,解释导致一个改进的融合方法,线性结合ASV和CM的对数似然比。然而,正如第三个发现所揭示的那样,这种线性组合在做出最优决策方面不如非线性组合。这些发现的结果,即融合前的评分校准、改进的线性融合和更好的非线性融合,被认为对SASV挑战数据库有效。摘要:Fusing outputs from automatic speaker verification (ASV) and spoofing countermeasure (CM) is expected to make an integrated system robust to zero-effort imposters and synthesized spoofing attacks. Many score-level fusion methods have been proposed, but many remain heuristic. This paper revisits score-level fusion using tools from decision theory and presents three main findings. First, fusion by summing the ASV and CM scores can be interpreted on the basis of compositional data analysis, and score calibration before fusion is essential. Second, the interpretation leads to an improved fusion method that linearly combines the log-likelihood ratios of ASV and CM. However, as the third finding reveals, this linear combination is inferior to a non-linear one in making optimal decisions. The outcomes of these findings, namely, the score calibration before fusion, improved linear fusion, and better non-linear fusion, were found to be effective on the SASV challenge database.
【31】 Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
标题: 奥德赛2024年语音情感识别挑战赛的双多头注意力多模式系统
作者:Federico Costa,Miquel India,Javier Hernando
备注:Odyssey 2024: The Speaker and Language Recognition Workshop
链接:点击下载PDF文件
摘要:随着计算机应用越来越多地融入我们的日常生活,语音情感识别(SER)的重要性显着增加。通过SER的创新方法促进研究,Odyssey 2024语音情感识别挑战赛是Odyssey 2024扬声器和语言识别研讨会的一部分。在本文中,我们描述了双多头注意力多模式系统开发的这一挑战。预训练的自监督模型用于提取信息丰富的声学和文本特征。采用早期融合策略,其中多头注意层将这些混合特征转换为互补的上下文表示。然后应用第二注意机制将这些表示汇集到话语级向量中。我们提出的系统在分类任务排名中以34.41%的Macro-F1得分获得第三名,共有31个团队参与。摘要:As computer-based applications are becoming more integrated into our daily lives, the importance of Speech Emotion Recognition (SER) has increased significantly. Promoting research with innovative approaches in SER, the Odyssey 2024 Speech Emotion Recognition Challenge was organized as part of the Odyssey 2024 Speaker and Language Recognition Workshop. In this paper we describe the Double Multi-Head Attention Multimodal System developed for this challenge. Pre-trained self-supervised models were used to extract informative acoustic and text features. An early fusion strategy was adopted, where a Multi-Head Attention layer transforms these mixed features into complementary contextualized representations. A second attention mechanism is then applied to pool these representations into an utterance-level vector. Our proposed system achieved the third position in the categorical task ranking with a 34.41% Macro-F1 score, where 31 teams participated in total.
【32】 MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation
标题: MNT:用于Foley音频内容规划和生成的多模式图像和叙事文本配音数据集
作者:Ruibo Fu,Shuchen Shi,Hongming Guo,Tao Wang,Chunyu Qiang,Zhengqi Wen,Jianhua Tao,Xin Qi,Yi Lu,Xiaopeng Wang,Zhiyong Wang,Yukun Liu,Xuefei Liu,Shuai Zhang,Guanjun Li
链接:点击下载PDF文件
摘要:Foley音频对于增强多媒体内容的沉浸式体验至关重要,但在AI生成内容(AIGC)领域面临着重大挑战。尽管AIGC技术在文本和图像生成方面取得了进步,但由于跨模态场景匹配和内容相关性方面的困难,Foley音频配音仍然是基本的。目前的文本到音频技术,它依赖于详细的和声学相关的文本描述,在实际的视频配音应用中不足。现有的数据集,如AudioSet,AudioCaps,Clotho,Sound-of-Story和WavCaps并不能完全满足真实世界的Foley音频配音任务的要求。为了解决这个问题,我们引入了多模态图像和叙事文本配音数据集(MINT),旨在增强主流配音任务,如文学故事有声读物配音,图像 无声视频配音。此外,为了解决现有TTA技术在理解和规划复杂提示方面的局限性,提出了一种Foley音频内容规划、生成和对齐(CPGA)框架,该框架包括利用大型语言模型进行复杂多模态提示理解的内容规划模块。此外,使用基于邻近策略优化的强化学习优化了训练过程,显著提高了生成的Foley音频的对齐和听觉真实性。实验结果表明,我们的方法显着推进领域的福利音频配音,提供了强大的解决方案的挑战,多模态配音。即使在使用相对轻量级的GPT-2模型时,我们的框架也优于LLaVA,DeepSeek-VL和Moondream 2等开源多模式大型模型。该数据集可在https: github.com borisfrb MINT上获得。摘要:Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape. Despite advancements in AIGC technologies for text and image generation, the foley audio dubbing remains rudimentary due to difficulties in cross-modal scene matching and content correlation. Current text-to-audio technology, which relies on detailed and acoustically relevant textual descriptions, falls short in practical video dubbing applications. Existing datasets like AudioSet, AudioCaps, Clotho, Sound-of-Story, and WavCaps do not fully meet the requirements for real-world foley audio dubbing task. To address this, we introduce the Multi-modal Image and Narrative Text Dubbing Dataset (MINT), designed to enhance mainstream dubbing tasks such as literary story audiobooks dubbing, image silent video dubbing. Besides, to address the limitations of existing TTA technology in understanding and planning complex prompts, a Foley Audio Content Planning, Generation, and Alignment (CPGA) framework is proposed, which includes a content planning module leveraging large language models for complex multi-modal prompts comprehension. Additionally, the training process is optimized using Proximal Policy Optimization based reinforcement learning, significantly improving the alignment and auditory realism of generated foley audio. Experimental results demonstrate that our approach significantly advances the field of foley audio dubbing, providing robust solutions for the challenges of multi-modal dubbing. Even when utilizing the relatively lightweight GPT-2 model, our framework outperforms open-source multimodal large models such as LLaVA, DeepSeek-VL, and Moondream2. The dataset is available at https: github.com borisfrb MINT .
【33】 Lightweight Audio Segmentation for Long-form Speech Translation
标题: 用于长格式语音翻译的轻量级音频分割
作者:Jaesong Lee,Soyoon Kim,Hanbyul Kim,Joon Son Chung
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:语音分割是语音翻译系统的重要组成部分。由于大多数ST模型被设计为处理语音片段,因此在翻译之前必须将长格式音频划分为较短的片段。最近,数据驱动的语音分割任务的方法已经开发出来。虽然这些方法提高了整体翻译质量,但由于模型和ST系统之间的不匹配,存在性能差距。此外,现有的工作需要大型的自监督语音模型,这消耗了大量的计算资源。在这项工作中,我们提出了一个分割模型,实现更好的语音翻译质量与一个小的模型大小。我们提出了一个带有标点符号的ASR任务作为分割模型的有效预训练策略。我们还表明,适当的语音分割模型集成到底层ST系统是至关重要的,以提高整体翻译质量在推理时间。摘要:Speech segmentation is an essential part of speech translation (ST) systems in real-world scenarios. Since most ST models are designed to process speech segments, long-form audio must be partitioned into shorter segments before translation. Recently, data-driven approaches for the speech segmentation task have been developed. Although the approaches improve overall translation quality, a performance gap exists due to a mismatch between the models and ST systems. In addition, the prior works require large self-supervised speech models, which consume significant computational resources. In this work, we propose a segmentation model that achieves better speech translation quality with a small model size. We propose an ASR-with-punctuation task as an effective pre-training strategy for the segmentation model. We also show that proper integration of the speech segmentation model into the underlying ST system is critical to improve overall translation quality at inference time.
【34】 Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
标题: 关节发音学知情的可控表达性语音合成
作者:Zehua Kcriss Li,Meiying Melissa Chen,Yi Zhong,Pinxin Liu,Zhiyao Duan
链接:点击下载PDF文件
摘要:表达性语音合成旨在生成捕捉广泛的准语言特征的语音,包括情感和发音,尽管目前的研究主要强调情感方面,而不是专业配音演员掌握的细微差别的发音特征。受此启发,我们通过发音语音学的镜头探索表达性语音合成。具体来说,我们定义了一个框架与三个维度:声门化,紧张,共鸣(GTR),指导合成的声音生产水平。在这个框架下,我们记录了一个名为GTR-Voice的高质量语音数据集,该数据集包含专业配音演员在125个不同GTR组合中表达的20个中文句子。我们通过自动分类和听力测试验证了框架和GTR注释,并在两个微调的表达TTS模型上沿着GTR维度展示了精确的可控性。我们开源了数据集和TTS模型。摘要:Expressive speech synthesis aims to generate speech that captures a wide range of para-linguistic features, including emotion and articulation, though current research primarily emphasizes emotional aspects over the nuanced articulatory features mastered by professional voice actors. Inspired by this, we explore expressive speech synthesis through the lens of articulatory phonetics. Specifically, we define a framework with three dimensions: Glottalization, Tenseness, and Resonance (GTR), to guide the synthesis at the voice production level. With this framework, we record a high-quality speech dataset named GTR-Voice, featuring 20 Chinese sentences articulated by a professional voice actor across 125 distinct GTR combinations. We verify the framework and GTR annotations through automatic classification and listening tests, and demonstrate precise controllability along the GTR dimensions on two fine-tuned expressive TTS models. We open-source the dataset and TTS models.
【35】 SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR
标题: SOC:通过低资源ASB的纯语音自适应减少SSL管道中的域不匹配
作者:Natarajan Balaji Shankar,Ruchao Fan,Abeer Alwan
备注:Accepted to ICASSP 2024 SASB Workshop
链接:点击下载PDF文件
摘要:最近,语音基础模型由于其在微调下游ASR任务方面的优势而受到欢迎。然而,在某些领域进行微调的模型,如LibriSpeech(成人阅读语音),在其他领域(儿童或嘈杂语音)表现不佳。一个解决方案可能是收集尽可能多的标记和不同的数据,以便在各个领域进行联合微调。然而,收集目标域语音文本配对数据和重新训练模型通常是昂贵的和计算昂贵的。在本文中,我们介绍了一种简单而有效的方法,语音只适应(SOA),语音基础模型(Wav2vec 2.0)的基础上,它只需要从目标域的语音输入数据。具体来说,Wav2vec 2.0特征编码器在域自适应的源和目标域数据上连续预训练Wav2vec 2.0损失,而上下文编码器被冻结。与在训练过程中冻结特征编码器的源域微调模型相比,我们发现,用适应的特征编码器替换冻结的特征编码器可以显著改善目标域的WER,同时保持源域的性能。SOA的有效性进行了检查各种低资源或域不匹配的ASR设置,包括成人儿童和清洁嘈杂的语音。摘要:Recently, speech foundation models have gained popularity due to their superiority in finetuning downstream ASR tasks. However, models finetuned on certain domains, such as LibriSpeech (adult read speech), behave poorly on other domains (child or noisy speech). One solution could be collecting as much labeled and diverse data as possible for joint finetuning on various domains. However, collecting target domain speech-text paired data and retraining the model is often costly and computationally expensive. In this paper, we introduce a simple yet effective method, speech only adaptation (SOA), based on speech foundation models (Wav2vec 2.0), which requires only speech input data from the target domain. Specifically, the Wav2vec 2.0 feature encoder is continually pretrained with the Wav2vec 2.0 loss on both the source and target domain data for domain adaptation, while the contextual encoder is frozen. Compared to a source domain finetuned model with the feature encoder being frozen during training, we find that replacing the frozen feature encoder with the adapted one provides significant WER improvements to the target domain while preserving the performance of the source domain. The effectiveness of SOA is examined on various low resource or domain mismatched ASR settings, including adult-child and clean-noisy speech.
【36】 Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
标题: 使用监督和自我监督言语基础模型对儿童ASB进行基准测试
作者:Ruchao Fan,Natarajan Balaji Shankar,Abeer Alwan
备注:To appear in Interspeech 2024
链接:点击下载PDF文件
摘要:语音基础模型(SFM)已经在监督(例如Whisper)或自监督系统(例如WavLM)中的各种语音任务中取得了最先进的结果。然而,儿童ASR的SFM的性能还没有得到系统的研究。此外,没有标准评估儿童ASR的基准,这使得比较新颖的想法变得困难。在本文中,我们发起并提出了一个全面的基准几个孩子的语音数据库的基础上,各种SFM(耳语,Wav2vec2.0,HuBERT,和WavLM)。此外,我们研究微调策略,通过比较各种数据增强和参数有效的微调(PEFT)方法。我们观察到,这些方法的行为是不同的,当模型的大小增加。例如,PEFT匹配大模型的完全微调性能,但小模型的性能更差。为了稳定使用增强数据的微调,我们提出了一个扰动不变微调(PIF)损失作为正则化。摘要:Speech foundation models (SFMs) have achieved state-of-the-art results for various speech tasks in supervised (e.g. Whisper) or self-supervised systems (e.g. WavLM). However, the performance of SFMs for child ASR has not been systematically studied. In addition, there is no benchmark for child ASR with standard evaluations, making the comparisons of novel ideas difficult. In this paper, we initiate and present a comprehensive benchmark on several child speech databases based on various SFMs (Whisper, Wav2vec2.0, HuBERT, and WavLM). Moreover, we investigate finetuning strategies by comparing various data augmentation and parameter-efficient finetuning (PEFT) methods. We observe that the behaviors of these methods are different when the model size increases. For example, PEFT matches the performance of full finetuning for large models but worse for small models. To stabilize finetuning using augmented data, we propose a perturbation invariant finetuning (PIF) loss as a regularization.
【37】 AVR: Synergizing Foundation Models for Audio-Visual Humor Detection
标题: AVR:协同视听幽默检测的基础模型
作者:Sarthak Sharma,Orchid Chetia Phukan,Drishti Singh,Arun Balaji Buduru,Rajesh Sharma
备注:Accepted to INTERSPEECH 2024 Show & Tell Demonstrations
链接:点击下载PDF文件
摘要:在这项工作中,我们提出,AVR应用视听幽默检测。虽然幽默检测传统上以文本分析为中心,但最近的进展突出了多模态方法。然而,这些方法依赖于文本线索作为模态,需要使用ASR系统来转录音频数据。这种对ASR准确性的严重依赖可能会在实际应用中带来挑战。为了解决这个瓶颈,我们提出了一个创新的视听幽默检测系统,绕过文本依赖,消除了对ASR模型的需要。相反,所提出的方法取决于音频和视觉内容之间复杂的相互作用,以实现有效的幽默检测。摘要:In this work, we present, AVR application for audio-visual humor detection. While humor detection has traditionally centered around textual analysis, recent advancements have spotlighted multimodal approaches. However, these methods lean on textual cues as a modality, necessitating the use of ASR systems for transcribing the audio-data. This heavy reliance on ASR accuracy can pose challenges in real-world applications. To address this bottleneck, we propose an innovative audio-visual humor detection system that circumvents textual reliance, eliminating the need for ASR models. Instead, the proposed approach hinges on the intricate interplay between audio and visual content for effective humor detection.
【38】 Phoneme Discretized Saliency Maps for Explainable Detection of AI-Generated Voice
标题: 音素离散显着图用于人工智能生成语音的可解释检测
作者:Shubham Gupta,Mirco Ravanelli,Pascal Germain,Cem Subakan
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了音素离散显着性图(PDSM),这是一种显着性图的离散化算法,它利用音素边界对AI生成的语音进行可解释的检测。我们用两种不同的文本到语音系统(即,Tacotron 2和Fastspeech 2),该算法产生的显着图,导致更忠实的解释相比,标准的事后解释方法。此外,通过将显着性图与音素表示相关联,这种方法生成的解释往往比幅度谱图上的标准显着性图更容易理解。摘要:In this paper, we propose Phoneme Discretized Saliency Maps (PDSM), a discretization algorithm for saliency maps that takes advantage of phoneme boundaries for explainable detection of AI-generated voice. We experimentally show with two different Text-to-Speech systems (i.e., Tacotron2 and Fastspeech2) that the proposed algorithm produces saliency maps that result in more faithful explanations compared to standard posthoc explanation methods. Moreover, by associating the saliency maps to the phoneme representations, this methodology generates explanations that tend to be more understandable than standard saliency maps on magnitude spectrograms.
【39】 Evaluating Speaker Identity Coding in Self-supervised Models and Humans
标题: 评估自我监督模型和人类中的说话者身份编码
作者:Gasser Elbanna
备注:Masters Thesis
链接:点击下载PDF文件
摘要:说话人身份在人类交流中发挥着重要作用,并且越来越多地用于社会应用,其中许多是通过机器学习的进步。说话人身份感知是一种重要的认知现象,可以概括为两个主要任务:识别声音或区分声音。一些研究试图确定身份感知的声学相关性,以确定这样一个任务的显着参数。与其他沟通性的社会信号不同,大多数努力都产生了无效的结论。此外,目前的语音识别处理的神经认知模型认为感知的基础是声学维度,如基频,谐波噪声比和共振峰分散。然而,这些研究结果并没有考虑到自然主义的讲话和说话人内的变化。当前自监督模型的表征空间在各种语音相关任务中表现出显着的性能。在这项工作中,我们证明了来自不同家庭的自我监督表示(例如,生成、对比和预测模型)对于说话人识别明显优于声学表示。我们还表明,这样的说话人识别任务可以用来更好地理解这些强大的网络的不同层中的声学信息表示的性质。通过评估声学,音素,韵律和语言变体的说话人识别精度,我们报告模型性能和人类身份感知之间的相似性。我们进一步研究这些相似之处并列的编码空间的模型和人类和挑战使用的距离度量作为一个代理扬声器接近。最后,我们证明了一些模型可以预测自然刺激过程中听觉和语言区域的大脑反应。摘要:Speaker identity plays a significant role in human communication and is being increasingly used in societal applications, many through advances in machine learning. Speaker identity perception is an essential cognitive phenomenon that can be broadly reduced to two main tasks: recognizing a voice or discriminating between voices. Several studies have attempted to identify acoustic correlates of identity perception to pinpoint salient parameters for such a task. Unlike other communicative social signals, most efforts have yielded inefficacious conclusions. Furthermore, current neurocognitive models of voice identity processing consider the bases of perception as acoustic dimensions such as fundamental frequency, harmonics-to-noise ratio, and formant dispersion. However, these findings do not account for naturalistic speech and within-speaker variability. Representational spaces of current self-supervised models have shown significant performance in various speech-related tasks. In this work, we demonstrate that self-supervised representations from different families (e.g., generative, contrastive, and predictive models) are significantly better for speaker identification over acoustic representations. We also show that such a speaker identification task can be used to better understand the nature of acoustic information representation in different layers of these powerful networks. By evaluating speaker identification accuracy across acoustic, phonemic, prosodic, and linguistic variants, we report similarity between model performance and human identity perception. We further examine these similarities by juxtaposing the encoding spaces of models and humans and challenging the use of distance metrics as a proxy for speaker proximity. Lastly, we show that some models can predict brain responses in Auditory and Language regions during naturalistic stimuli.
【40】 Gender Representation in TV and Radio: Automatic Information Extraction methods versus Manual Analyses
标题: 电视和广播中的性别表现:自动信息提取方法与手动分析
作者:David Doukhan,Lena Dodson,Manon Conan,Valentin Pelloin,Aurélien Clamouse,Mélina Lepape,Géraldine Van Hille,Cécile Méadel,Marlène Coulomb-Gully
备注:keywords : Gender representation, computational humanities, TV, Radio, face classification, speaker traits, ASR, media, SLU. Accepted to InterSpeech 2024, Kos Island, Greece, september 2024
链接:点击下载PDF文件
摘要:本研究探讨自动信息提取描述符和人工分析之间的关系,以描述在电视和广播的性别代表性差异。自动描述符,包括语音时间,面部分类和语音transmittance进行了比较,从2023年起,在一个巨大的32,000小时的法语广播语料库的频道报告。调查结果揭示了系统性的性别不平衡,在所有描述中,与男子相比,妇女的代表性不足。值得注意的是,人工渠道报告显示妇女的出席率高于自动估计,提及妇女的时间低于其发言时间。描述符在高和低观众,战争报道,或私人与公共频道之间有着共同的动态。虽然在法国电视中,妇女的形象比声音更明显,但在新闻中,这一趋势却相反,看不见的记者描绘了男性主角。统计检验表明,影响女性参考的主要因素有三个:节目类别、频道和说话人性别。摘要:This study investigates the relationship between automatic information extraction descriptors and manual analyses to describe gender representation disparities in TV and Radio. Automatic descriptors, including speech time, facial categorization and speech transcriptions are compared with channel reports on a vast 32,000-hour corpus of French broadcasts from 2023. Findings reveal systemic gender imbalances, with women underrepresented compared to men across all descriptors. Notably, manual channel reports show higher women's presence than automatic estimates and references to women are lower than their speech time. Descriptors share common dynamics during high and low audiences, war coverage, or private versus public channels. While women are more visible than audible in French TV, this trend is inverted in news with unseen journalists depicting male protagonists. A statistical test shows 3 main effects influencing references to women: program category, channel and speaker gender.
eess.AS音频处理
【1】 1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis标题: 1000个非洲之声:推进包容性多说话人多口音语音合成
作者:Sewade Ogun,Abraham T. Owodunni,Tobi Olatunji,Eniola Alese,Babatunde Oladimeji,Tejumade Afonja,Kayode Olaleye,Naome A. Etori,Tosin Adewumi
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:语音合成的最新进展使许多有用的应用程序成为可能,如谷歌地图中的音频方向,屏幕阅读器以及TikTok等平台上的自动内容生成。然而,这些系统大多由来自数据丰富的地理位置的声音主导,其中人物角色代表其源数据。虽然世界上有3000种语言在非洲使用,但非洲的声音和人物在这些系统中的代表性不足。随着语音合成变得越来越民主化,人们希望增加非洲英语口音的代表性。我们提出了Afro-TTS,第一个泛非洲口音英语语音合成系统,能够生成86种非洲口音的语音,1000个人物代表了整个非洲大陆丰富的语音多样性,用于教育,公共卫生和自动内容创建的下游应用。扬声器插值保持自然和重音,使新的声音的创建。摘要:Recent advances in speech synthesis have enabled many useful applications like audio directions in Google Maps, screen readers, and automated content generation on platforms like TikTok. However, these systems are mostly dominated by voices sourced from data-rich geographies with personas representative of their source data. Although 3000 of the world's languages are domiciled in Africa, African voices and personas are under-represented in these systems. As speech synthesis becomes increasingly democratized, it is desirable to increase the representation of African English accents. We present Afro-TTS, the first pan-African accented English speech synthesis system able to generate speech in 86 African accents, with 1000 personas representing the rich phonological diversity across the continent for downstream application in Education, Public Health, and Automated Content Creation. Speaker interpolation retains naturalness and accentedness, enabling the creation of new voices.
【2】 AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling
标题: AV-CrossNet:一种通过利用窄带和跨带建模实现语音分离的视听复谱映射网络
作者:Vahid Ahmadi Kalkhorani,Cheng Yu,Anurag Kumar,Ke Tan,Buye Xu,DeLiang Wang
备注:10 pages, 4 Figures, and 4 Tables
链接:点击下载PDF文件
摘要:在基于音频的语音分离中加入视觉线索可以提高分离性能。本文介绍了AV-CrossNet,这是一个用于语音增强、目标说话人提取和多说话人说话人分离的AV系统。AV-CrossNet是从CrossNet架构扩展而来的,CrossNet架构是最近提出的一种网络,通过利用全局注意力和位置编码来执行复杂的频谱映射以进行语音分离。为了有效地利用视觉线索,所提出的系统结合了预提取的视觉嵌入,并采用包括时间卷积层的视觉编码器。音频和视觉功能在馈送到AV-CrossNet块之前在早期融合层中融合。我们在多个数据集上评估了AV-CrossNet,包括LRS,VoxCeleb和COG-MHEAR挑战。评估结果表明,AV-CrossNet提高了所有视听任务的最先进性能,即使是在未经训练和不匹配的数据集上。摘要:Adding visual cues to audio-based speech separation can improve separation performance. This paper introduces AV-CrossNet, an gls{av} system for speech enhancement, target speaker extraction, and multi-talker speaker separation. AV-CrossNet is extended from the CrossNet architecture, which is a recently proposed network that performs complex spectral mapping for speech separation by leveraging global attention and positional encoding. To effectively utilize visual cues, the proposed system incorporates pre-extracted visual embeddings and employs a visual encoder comprising temporal convolutional layers. Audio and visual features are fused in an early fusion layer before feeding to AV-CrossNet blocks. We evaluate AV-CrossNet on multiple datasets, including LRS, VoxCeleb, and COG-MHEAR challenge. Evaluation results demonstrate that AV-CrossNet advances the state-of-the-art performance in all audiovisual tasks, even on untrained and mismatched datasets.
【3】 GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
标题: GigaSpeech 2:一个不断发展的、大规模和多领域的SVR数据库,用于低资源语言,具有自动抓取、转录和细化
作者:Yifan Yang,Zheshu Song,Jianheng Zhuo,Mingyu Cui,Jinpeng Li,Bo Yang,Yexing Du,Ziyang Ma,Xunying Liu,Ziyuan Wang,Ke Li,Shuai Fan,Kai Yu,Wei-Qiang Zhang,Guoguo Chen,Xie Chen
备注:Under review
链接:点击下载PDF文件
摘要:语音技术的发展受到数据集规模快速增长的推动。传统的语音模型通常依赖于大量的标记训练数据,这对于低资源语言来说是稀缺的。本文介绍了GigaSpeech 2,一个大规模,多领域,多语种的语音识别语料库。它是为低资源语言设计的,不依赖于成对的语音和文本数据。GigaSpeech 2包含约30,000小时的自动转录语音,包括泰国语,印度尼西亚语和越南语,这些语音来自未标记的YouTube视频。我们还引入了一个自动化的管道,用于数据抓取,转录和标签优化。具体来说,该管道使用Whisper进行初始转录,使用TorchAudio进行强制对齐,并结合多维过滤来保证数据质量。改进的Noisy Student训练方法进一步迭代地改进有缺陷的伪标签,从而提高模型性能。我们的人工转录的评价集和两个公共测试集,从普通的声音和FLEURS的实验结果证实了我们的语料库的高质量和广泛的适用性。值得注意的是,在GigaSpeech 2上训练的ASR模型可以在我们具有挑战性和现实性的YouTube测试集上将泰国语,印度尼西亚语和越南语的单词错误率降低25%至40%,而Whisper大v3模型只有10%的模型参数。此外,与商业服务相比,我们在Gigaspeech 2上训练的ASR模型具有卓越的性能。我们相信,我们新引入的语料库和管道将为低资源语音识别开辟一条新的途径,并大大促进这一领域的研究。摘要:The evolution of speech technology has been spurred by the rapid increase in dataset sizes. Traditional speech models generally depend on a large amount of labeled training data, which is scarce for low-resource languages. This paper presents GigaSpeech 2, a large-scale, multi-domain, multilingual speech recognition corpus. It is designed for low-resource languages and does not rely on paired speech and text data. GigaSpeech 2 comprises about 30,000 hours of automatically transcribed speech, including Thai, Indonesian, and Vietnamese, gathered from unlabeled YouTube videos. We also introduce an automated pipeline for data crawling, transcription, and label refinement. Specifically, this pipeline uses Whisper for initial transcription and TorchAudio for forced alignment, combined with multi-dimensional filtering for data quality assurance. A modified Noisy Student Training is developed to further refine flawed pseudo labels iteratively, thus enhancing model performance. Experimental results on our manually transcribed evaluation set and two public test sets from Common Voice and FLEURS confirm our corpus's high quality and broad applicability. Notably, ASR models trained on GigaSpeech 2 can reduce the word error rate for Thai, Indonesian, and Vietnamese on our challenging and realistic YouTube test set by 25% to 40% compared to the Whisper large-v3 model, with merely 10% model parameters. Furthermore, our ASR models trained on Gigaspeech 2 yield superior performance compared to commercial services. We believe that our newly introduced corpus and pipeline will open a new avenue for low-resource speech recognition and significantly facilitate research in this area.
【4】 DiTTo-TTS: Efficient and Scalable Zero-Shot Text-to-Speech with Diffusion Transformer
标题: DiTTo-TTC:具有扩散Transformer的高效且可扩展的Zero-Shot文本到语音
作者:Keon Lee,Dong Won Kim,Jaehyeon Kim,Jaewoong Cho
链接:点击下载PDF文件
摘要:大规模扩散模型在包括图像、视频和音频在内的多种模态中表现出出色的生成能力。然而,文本到语音(TTS)系统通常涉及域特定建模因素(例如,音素和音素级持续时间),以确保文本和语音之间的精确时间对齐,这阻碍了TTS扩散模型的效率和可扩展性。在这项工作中,我们提出了一个有效的和可扩展的扩散Transformer(DiT),利用现成的预先训练的文本和语音编码器。我们的方法解决了文本语音对齐的挑战,通过交叉注意机制与语音表示的总长度的预测。为了实现这一目标,我们增强了DiT架构,以适应TTS和提高对齐,将语义指导到语音的潜在空间。我们将训练数据集和模型大小分别扩展到82K小时和790M参数。我们广泛的实验表明,大规模的扩散模型的TTS没有特定领域的建模,不仅简化了训练管道,但也产生优越的或可比的zero-shot性能,以国家的最先进的TTS模型的自然度,可理解性和说话人相似性。我们的语音样本可以在https: ditto-tts.github.io上找到。摘要:Large-scale diffusion models have shown outstanding generative abilities across multiple modalities including images, videos, and audio. However, text-to-speech (TTS) systems typically involve domain-specific modeling factors (e.g., phonemes and phoneme-level durations) to ensure precise temporal alignments between text and speech, which hinders the efficiency and scalability of diffusion models for TTS. In this work, we present an efficient and scalable Diffusion Transformer (DiT) that utilizes off-the-shelf pre-trained text and speech encoders. Our approach addresses the challenge of text-speech alignment via cross-attention mechanisms with the prediction of the total length of speech representations. To achieve this, we enhance the DiT architecture to suit TTS and improve the alignment by incorporating semantic guidance into the latent space of speech. We scale the training dataset and the model size to 82K hours and 790M parameters, respectively. Our extensive experiments demonstrate that the large-scale diffusion model for TTS without domain-specific modeling not only simplifies the training pipeline but also yields superior or comparable zero-shot performance to state-of-the-art TTS models in terms of naturalness, intelligibility, and speaker similarity. Our speech samples are available at https: ditto-tts.github.io.
【5】 An Exploration of Length Generalization in Transformer-Based Speech Enhancement
标题: 基于变换器的语音增强中长度概括的探索
作者:Qiquan Zhang,Hongxu Zhu,Xinyuan Qian,Eliathamby Ambikairajah,Haizhou Li
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:Transformer架构的使用促进了语音增强的显著进展。使用相当长的语音话语来训练Transformers通常是不可行的,因为自我注意力受到二次复杂度的影响。对于基于transformer的语音增强模型来说,从短语音中学习并推广到长语音是一个关键的和未探索的挑战。本文对Transformer语音增强中的长度泛化问题进行了综合实验研究。我们的研究结果首先建立了位置嵌入提供了一个有效的工具,以减轻话语长度的影响,基于transformer的语音增强。具体来说,我们探索了四种不同的位置嵌入方案,使长度泛化。结果证实了相对位置嵌入在长度推广方面优于绝对位置嵌入。摘要:The use of Transformer architectures has facilitated remarkable progress in speech enhancement. Training Transformers using substantially long speech utterances is often infeasible as self-attention suffers from quadratic complexity. It is a critical and unexplored challenge for a Transformer-based speech enhancement model to learn from short speech utterances and generalize to longer ones. In this paper, we conduct comprehensive experiments to explore the length generalization problem in speech enhancement with Transformer. Our findings first establish that position embedding provides an effective instrument to alleviate the impact of utterance length on Transformer-based speech enhancement. Specifically, we explore four different position embedding schemes to enable length generalization. The results confirm the superiority of relative position embeddings (RPEs) over absolute PE (APEs) in length generalization.
【6】 Spatially constrained vs. unconstrained filtering in neural spatiospectral filters for multichannel speech enhancement
标题: 用于多通道语音增强的神经空间谱过滤器中的空间约束与无约束过滤
作者:Annika Briegleb,Walter Kellermann
备注:Accepted to the 32nd European Signal Processing Conference (EUSIPCO 2024), Lyon, France. 5 pages, 4 figures
链接:点击下载PDF文件
摘要:当使用人工神经网络进行多通道语音增强时,滤波通常通过估计应用于输入信号的所有或一个参考通道的复值掩码来实现。该掩模的估计是基于有噪声的多通道信号,因此,可以同时利用空间和频谱线索。虽然已经表明,联合利用空间和频谱线索有利于语音增强结果,但神经网络内部两者相互作用的机制在很大程度上仍然是未知的。在这篇文章中,我们研究了两个概念上不同的神经空间谱滤波器(NSSF)如何根据训练目标信号利用空间线索,并表明,虽然一个NSSF总是执行空间滤波,但另一个NSSF在利用空间信息方面是有选择性的,这取决于手头的任务。这些见解提供了对NSSF用于进行预测的信息的更好理解,从而可以就其设计和部署做出明智的决策。摘要:When using artificial neural networks for multichannel speech enhancement, filtering is often achieved by estimating a complex-valued mask that is applied to all or one reference channel of the input signal. The estimation of this mask is based on the noisy multichannel signal and, hence, can exploit spatial and spectral cues simultaneously. While it has been shown that exploiting spatial and spectral cues jointly is beneficial for the speech enhancement result, the mechanics of the interplay of the two inside the neural network are still largely unknown. In this contribution, we investigate how two conceptually different neural spatiospectral filters (NSSFs) exploit spatial cues depending on the training target signal and show that, while one NSSF always performs spatial filtering, the other one is selective in leveraging spatial information depending on the task at hand. These insights provide better understanding of the information the NSSFs use to make their prediction and, thus, allow to make informed decisions regarding their design and deployment.
【7】 Performance Improvement of Language-Queried Audio Source Separation Based on Caption Augmentation From Large Language Models for DCASE Challenge 2024 Task 9
标题: 基于来自大型语言模型的字幕增强的数字查询音频源分离的性能改进DUSE Challenge 2024任务9
作者:Do Hyun Lee,Yoonah Song,Hong Kook Kim
备注:DCASE 2024 Challenge Task 9, 4 pages
链接:点击下载PDF文件
摘要:我们提出了一个基于工程的文本增强方法应用于语言查询的音频源分离(LASS)任务。为了提高LASS的性能,所提出的方法利用大语言模型(LLM)来生成对应于训练数据集的每个句子的多个字幕。为此,我们首先进行实验,以确定最有效的提示字幕扩增与较少的标题。使用这些增强字幕训练的LASS模型在DCASE 2024任务9验证集上表现出与未经增强训练的LASS模型相比的改进性能。这项研究强调了基于LLM的字幕增强在推进语言查询音频源分离方面的有效性。摘要:We present a prompt-engineering-based text-augmentation approach applied to a language-queried audio source separation (LASS) task. To enhance the performance of LASS, the proposed approach utilizes large language models (LLMs) to generate multiple captions corresponding to each sentence of the training dataset. To this end, we first perform experiments to identify the most effective prompts for caption augmentation with a smaller number of captions. A LASS model trained with these augmented captions demonstrates improved performance on the DCASE 2024 Task 9 validation set compared to that trained without augmentation. This study highlights the effectiveness of LLM-based caption augmentation in advancing language-queried audio source separation.
【8】 Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
标题: 自蒸馏原型网络:在没有监督的情况下学习稳健的说话者表示
作者:Yafeng Chen,Siqi Zheng,Hui Wang,Luyao Cheng,Qian Chen,Shiliang Zhang,Wen Wang
链接:点击下载PDF文件
摘要:在没有明确说话人标签的情况下训练说话人区分和鲁棒的说话人确认系统仍然是一个持续的挑战。在本文中,我们提出了一种新的自监督说话人确认方法,自蒸馏原型网络(SDPN),它有效地促进了自监督说话人表示学习。SDPN将话语的增强视图的表示分配给与原始视图的表示相同的原型,从而实现视图之间的有效知识转移。最初,由于SDPN训练过程中缺少负对,网络往往会在嵌入空间中非常紧密地对齐正对,这种现象称为模型崩溃。为了缓解这个问题,我们在SDPN中的嵌入中引入了多样性正则化项。在VoxCeleb数据集上的综合实验证明了SDPN在自监督说话人确认中的优越性。SDPN在VoxCeleb 1说话者验证评估基准上设定了一个新的最先进水平,在训练中不使用任何说话者标签的情况下,VoxCeleb 1-O、VoxCeleb 1-E和VoxCeleb 1-H的试验分别实现了1.80%、1.99%和3.62%的等错误率。摘要:Training speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persisting challenge. In this paper, we propose a new self-supervised speaker verification approach, Self-Distillation Prototypes Network (SDPN), which effectively facilitates self-supervised speaker representation learning. SDPN assigns the representation of the augmented views of an utterance to the same prototypes as the representation of the original view, thereby enabling effective knowledge transfer between the views. Originally, due to the lack of negative pairs in the SDPN training process, the network tends to align positive pairs very closely in the embedding space, a phenomenon known as model collapse. To alleviate this problem, we introduce a diversity regularization term to embeddings in SDPN. Comprehensive experiments on the VoxCeleb datasets demonstrate the superiority of SDPN in self-supervised speaker verification. SDPN sets a new state-of-the-art on the VoxCeleb1 speaker verification evaluation benchmark, achieving Equal Error Rate 1.80%, 1.99%, and 3.62% for trial VoxCeleb1-O, VoxCeleb1-E and VoxCeleb1-H respectively, without using any speaker labels in training.
【9】 Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
标题: 含噪语音端到端语音识别的连续测试时间自适应
作者:Guan-Ting Lin,Wei-Ping Huang,Hung-yi Lee
备注:13 pages
链接:点击下载PDF文件
摘要:基于深度学习的端到端自动语音识别(ASR)已经取得了重大进展,但由于现实场景中的域转移,在域外(OOD)样本上的性能仍然很差。测试时自适应(TTA)方法通过在推理时使用测试样本来适应模型来解决这个问题。然而,目前的ASR TTA方法主要集中在非连续TTA,这限制了跨样本知识学习相比,连续TTA。在这项工作中,我们提出了一个快-慢TTA框架的ASR,它利用了连续和非连续TTA的优势。在这个框架内,我们介绍了动态SUTA(DSUTA),一个基于熵最小化的连续TTA方法的ASR。为了增强DSUTA对时变数据的鲁棒性,我们提出了一种动态重置策略,可以自动检测域偏移并重置模型,使其在处理多域数据时更有效。我们的方法在各种嘈杂的ASR数据集上表现出卓越的性能,优于非连续和连续的TTA基线,同时保持对域变化的鲁棒性,而不需要域边界信息。摘要:Deep learning-based end-to-end automatic speech recognition (ASR) has made significant strides but still struggles with performance on out-of-domain (OOD) samples due to domain shifts in real-world scenarios. Test-Time Adaptation (TTA) methods address this issue by adapting models using test samples at inference time. However, current ASR TTA methods have largely focused on non-continual TTA, which limits cross-sample knowledge learning compared to continual TTA. In this work, we propose a Fast-slow TTA framework for ASR, which leverages the advantage of continual and non-continual TTA. Within this framework, we introduce Dynamic SUTA (DSUTA), an entropy-minimization-based continual TTA method for ASR. To enhance DSUTA's robustness on time-varying data, we propose a dynamic reset strategy that automatically detects domain shifts and resets the model, making it more effective at handling multi-domain data. Our method demonstrates superior performance on various noisy ASR datasets, outperforming both non-continual and continual TTA baselines while maintaining robustness to domain changes without requiring domain boundary information.
【10】 Multi-Scale Accent Modeling with Disentangling for Multi-Speaker Multi-Accent TTS Synthesis
标题: 多说话人多口音TTC合成的多尺度口音建模
作者:Xuehao Zhou,Mingyang Zhang,Yi Zhou,Zhizheng Wu,Haizhou Li
链接:点击下载PDF文件
摘要:在保留说话者身份的同时,跨不同口音合成语音对于各种现实世界的客户应用至关重要。然而,由于口音变化的复杂性以及口音和说话人身份之间的内在纠缠,在文本到语音(TTS)系统中对口音和说话人的个体和准确建模是具有挑战性的。在本文中,我们提出了一种新的方法,多说话人多口音的TTS合成,其目的是合成多个说话人的声音,每个不同的口音。我们提出的方法采用了多尺度口音建模策略,以解决口音的变化在不同的水平。具体来说,我们引入全球(话语级)和本地(音素级)口音建模,监督个别口音分类器,以捕捉重音话语和细粒度的音素之间的变化,分别在整体上的变化。为了分别控制口音和说话人,说话人无关的口音建模是必要的,这是通过使用说话人分类器进行对抗训练来实现的,以在多尺度口音建模中解开说话人身份。因此,我们获得了说话人无关和口音歧视的多尺度嵌入作为综合口音特征。此外,我们提出了一个本地口音预测模型,允许直接从音素输入生成口音语音。在带口音的英语语音语料库上进行了大量的实验。客观和主观的评价表明,我们提出的系统相比,基线系统的优越性。详细的成分分析表明,全局和局部口音建模,说话人解纠缠的多说话人多口音语音合成的有效性。摘要:Synthesizing speech across different accents while preserving the speaker identity is essential for various real-world customer applications. However, the individual and accurate modeling of accents and speakers in a text-to-speech (TTS) system is challenging due to the complexity of accent variations and the intrinsic entanglement between the accent and speaker identity. In this paper, we present a novel approach for multi-speaker multi-accent TTS synthesis, which aims to synthesize voices of multiple speakers, each with various accents. Our proposed approach employs a multi-scale accent modeling strategy to address accent variations at different levels. Specifically, we introduce both global (utterance level) and local (phoneme level) accent modeling, supervised by individual accent classifiers to capture the overall variation within accented utterances and fine-grained variations between phonemes, respectively. To control accents and speakers separately, speaker-independent accent modeling is necessary, which is achieved by adversarial training with speaker classifiers to disentangle speaker identity within the multi-scale accent modeling. Consequently, we obtain speaker-independent and accent-discriminative multi-scale embeddings as comprehensive accent features. Additionally, we propose a local accent prediction model that allows to generate accented speech directly from phoneme inputs. Extensive experiments are conducted on an accented English speech corpus. Both objective and subjective evaluations show the superiority of our proposed system compared to baselines systems. Detailed component analysis demonstrates the effectiveness of global and local accent modeling, and speaker disentanglement on multi-speaker multi-accent speech synthesis.
【11】 Revisiting and Improving Scoring Fusion for Spoofing-aware Speaker Verification Using Compositional Data Analysis
标题: 使用合成数据分析重新审视和改进具有欺骗意识的说话人验证的评分融合
作者:Xin Wang,Tomi Kinnunen,Kong Aik Lee,Paul-Gauthier Noé,Junichi Yamagishi
备注:Interspeech 2024 Accepted. this https URL
链接:点击下载PDF文件
摘要:融合自动说话人验证(ASV)和欺骗对策(CM)的输出,预计将使一个集成的系统强大的零努力冒名顶替者和合成欺骗攻击。已经提出了许多分数级融合方法,但许多仍然是启发式的。本文回顾了得分级融合决策理论的工具,并提出了三个主要的发现。首先,通过对ASV和CM分数求和的融合可以在成分数据分析的基础上进行解释,并且融合之前的分数校准是必不可少的。其次,解释导致一个改进的融合方法,线性结合ASV和CM的对数似然比。然而,正如第三个发现所揭示的那样,这种线性组合在做出最优决策方面不如非线性组合。这些发现的结果,即融合前的评分校准、改进的线性融合和更好的非线性融合,被认为对SASV挑战数据库有效。摘要:Fusing outputs from automatic speaker verification (ASV) and spoofing countermeasure (CM) is expected to make an integrated system robust to zero-effort imposters and synthesized spoofing attacks. Many score-level fusion methods have been proposed, but many remain heuristic. This paper revisits score-level fusion using tools from decision theory and presents three main findings. First, fusion by summing the ASV and CM scores can be interpreted on the basis of compositional data analysis, and score calibration before fusion is essential. Second, the interpretation leads to an improved fusion method that linearly combines the log-likelihood ratios of ASV and CM. However, as the third finding reveals, this linear combination is inferior to a non-linear one in making optimal decisions. The outcomes of these findings, namely, the score calibration before fusion, improved linear fusion, and better non-linear fusion, were found to be effective on the SASV challenge database.
【12】 Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
标题: 奥德赛2024年语音情感识别挑战赛的双多头注意力多模式系统
作者:Federico Costa,Miquel India,Javier Hernando
备注:Odyssey 2024: The Speaker and Language Recognition Workshop
链接:点击下载PDF文件
摘要:随着计算机应用越来越多地融入我们的日常生活,语音情感识别(SER)的重要性显着增加。通过SER的创新方法促进研究,Odyssey 2024语音情感识别挑战赛是Odyssey 2024扬声器和语言识别研讨会的一部分。在本文中,我们描述了双多头注意力多模式系统开发的这一挑战。预训练的自监督模型用于提取信息丰富的声学和文本特征。采用早期融合策略,其中多头注意层将这些混合特征转换为互补的上下文表示。然后应用第二注意机制将这些表示汇集到话语级向量中。我们提出的系统在分类任务排名中以34.41%的Macro-F1得分获得第三名,共有31个团队参与。摘要:As computer-based applications are becoming more integrated into our daily lives, the importance of Speech Emotion Recognition (SER) has increased significantly. Promoting research with innovative approaches in SER, the Odyssey 2024 Speech Emotion Recognition Challenge was organized as part of the Odyssey 2024 Speaker and Language Recognition Workshop. In this paper we describe the Double Multi-Head Attention Multimodal System developed for this challenge. Pre-trained self-supervised models were used to extract informative acoustic and text features. An early fusion strategy was adopted, where a Multi-Head Attention layer transforms these mixed features into complementary contextualized representations. A second attention mechanism is then applied to pool these representations into an utterance-level vector. Our proposed system achieved the third position in the categorical task ranking with a 34.41% Macro-F1 score, where 31 teams participated in total.
【13】 MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation
标题: MNT:用于Foley音频内容规划和生成的多模式图像和叙事文本配音数据集
作者:Ruibo Fu,Shuchen Shi,Hongming Guo,Tao Wang,Chunyu Qiang,Zhengqi Wen,Jianhua Tao,Xin Qi,Yi Lu,Xiaopeng Wang,Zhiyong Wang,Yukun Liu,Xuefei Liu,Shuai Zhang,Guanjun Li
链接:点击下载PDF文件
摘要:Foley音频对于增强多媒体内容的沉浸式体验至关重要,但在AI生成内容(AIGC)领域面临着重大挑战。尽管AIGC技术在文本和图像生成方面取得了进步,但由于跨模态场景匹配和内容相关性方面的困难,Foley音频配音仍然是基本的。目前的文本到音频技术,它依赖于详细的和声学相关的文本描述,在实际的视频配音应用中不足。现有的数据集,如AudioSet,AudioCaps,Clotho,Sound-of-Story和WavCaps并不能完全满足真实世界的Foley音频配音任务的要求。为了解决这个问题,我们引入了多模态图像和叙事文本配音数据集(MINT),旨在增强主流配音任务,如文学故事有声读物配音,图像 无声视频配音。此外,为了解决现有TTA技术在理解和规划复杂提示方面的局限性,提出了一种Foley音频内容规划、生成和对齐(CPGA)框架,该框架包括利用大型语言模型进行复杂多模态提示理解的内容规划模块。此外,使用基于邻近策略优化的强化学习优化了训练过程,显著提高了生成的Foley音频的对齐和听觉真实性。实验结果表明,我们的方法显着推进领域的福利音频配音,提供了强大的解决方案的挑战,多模态配音。即使在使用相对轻量级的GPT-2模型时,我们的框架也优于LLaVA,DeepSeek-VL和Moondream 2等开源多模式大型模型。该数据集可在https: github.com borisfrb MINT上获得。摘要:Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape. Despite advancements in AIGC technologies for text and image generation, the foley audio dubbing remains rudimentary due to difficulties in cross-modal scene matching and content correlation. Current text-to-audio technology, which relies on detailed and acoustically relevant textual descriptions, falls short in practical video dubbing applications. Existing datasets like AudioSet, AudioCaps, Clotho, Sound-of-Story, and WavCaps do not fully meet the requirements for real-world foley audio dubbing task. To address this, we introduce the Multi-modal Image and Narrative Text Dubbing Dataset (MINT), designed to enhance mainstream dubbing tasks such as literary story audiobooks dubbing, image silent video dubbing. Besides, to address the limitations of existing TTA technology in understanding and planning complex prompts, a Foley Audio Content Planning, Generation, and Alignment (CPGA) framework is proposed, which includes a content planning module leveraging large language models for complex multi-modal prompts comprehension. Additionally, the training process is optimized using Proximal Policy Optimization based reinforcement learning, significantly improving the alignment and auditory realism of generated foley audio. Experimental results demonstrate that our approach significantly advances the field of foley audio dubbing, providing robust solutions for the challenges of multi-modal dubbing. Even when utilizing the relatively lightweight GPT-2 model, our framework outperforms open-source multimodal large models such as LLaVA, DeepSeek-VL, and Moondream2. The dataset is available at https: github.com borisfrb MINT .
【14】 Lightweight Audio Segmentation for Long-form Speech Translation
标题: 用于长格式语音翻译的轻量级音频分割
作者:Jaesong Lee,Soyoon Kim,Hanbyul Kim,Joon Son Chung
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:语音分割是语音翻译系统的重要组成部分。由于大多数ST模型被设计为处理语音片段,因此在翻译之前必须将长格式音频划分为较短的片段。最近,数据驱动的语音分割任务的方法已经开发出来。虽然这些方法提高了整体翻译质量,但由于模型和ST系统之间的不匹配,存在性能差距。此外,现有的工作需要大型的自监督语音模型,这消耗了大量的计算资源。在这项工作中,我们提出了一个分割模型,实现更好的语音翻译质量与一个小的模型大小。我们提出了一个带有标点符号的ASR任务作为分割模型的有效预训练策略。我们还表明,适当的语音分割模型集成到底层ST系统是至关重要的,以提高整体翻译质量在推理时间。摘要:Speech segmentation is an essential part of speech translation (ST) systems in real-world scenarios. Since most ST models are designed to process speech segments, long-form audio must be partitioned into shorter segments before translation. Recently, data-driven approaches for the speech segmentation task have been developed. Although the approaches improve overall translation quality, a performance gap exists due to a mismatch between the models and ST systems. In addition, the prior works require large self-supervised speech models, which consume significant computational resources. In this work, we propose a segmentation model that achieves better speech translation quality with a small model size. We propose an ASR-with-punctuation task as an effective pre-training strategy for the segmentation model. We also show that proper integration of the speech segmentation model into the underlying ST system is critical to improve overall translation quality at inference time.
【15】 Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
标题: 关节发音学知情的可控表达性语音合成
作者:Zehua Kcriss Li,Meiying Melissa Chen,Yi Zhong,Pinxin Liu,Zhiyao Duan
链接:点击下载PDF文件
摘要:表达性语音合成旨在生成捕捉广泛的准语言特征的语音,包括情感和发音,尽管目前的研究主要强调情感方面,而不是专业配音演员掌握的细微差别的发音特征。受此启发,我们通过发音语音学的镜头探索表达性语音合成。具体来说,我们定义了一个框架与三个维度:声门化,紧张,共鸣(GTR),指导合成的声音生产水平。在这个框架下,我们记录了一个名为GTR-Voice的高质量语音数据集,该数据集包含专业配音演员在125个不同GTR组合中表达的20个中文句子。我们通过自动分类和听力测试验证了框架和GTR注释,并在两个微调的表达TTS模型上沿着GTR维度展示了精确的可控性。我们开源了数据集和TTS模型。摘要:Expressive speech synthesis aims to generate speech that captures a wide range of para-linguistic features, including emotion and articulation, though current research primarily emphasizes emotional aspects over the nuanced articulatory features mastered by professional voice actors. Inspired by this, we explore expressive speech synthesis through the lens of articulatory phonetics. Specifically, we define a framework with three dimensions: Glottalization, Tenseness, and Resonance (GTR), to guide the synthesis at the voice production level. With this framework, we record a high-quality speech dataset named GTR-Voice, featuring 20 Chinese sentences articulated by a professional voice actor across 125 distinct GTR combinations. We verify the framework and GTR annotations through automatic classification and listening tests, and demonstrate precise controllability along the GTR dimensions on two fine-tuned expressive TTS models. We open-source the dataset and TTS models.
【16】 SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR
标题: SOC:通过低资源ASB的纯语音自适应减少SSL管道中的域不匹配
作者:Natarajan Balaji Shankar,Ruchao Fan,Abeer Alwan
备注:Accepted to ICASSP 2024 SASB Workshop
链接:点击下载PDF文件
摘要:最近,语音基础模型由于其在微调下游ASR任务方面的优势而受到欢迎。然而,在某些领域进行微调的模型,如LibriSpeech(成人阅读语音),在其他领域(儿童或嘈杂语音)表现不佳。一个解决方案可能是收集尽可能多的标记和不同的数据,以便在各个领域进行联合微调。然而,收集目标域语音文本配对数据和重新训练模型通常是昂贵的和计算昂贵的。在本文中,我们介绍了一种简单而有效的方法,语音只适应(SOA),语音基础模型(Wav2vec 2.0)的基础上,它只需要从目标域的语音输入数据。具体来说,Wav2vec 2.0特征编码器在域自适应的源和目标域数据上连续预训练Wav2vec 2.0损失,而上下文编码器被冻结。与在训练过程中冻结特征编码器的源域微调模型相比,我们发现,用适应的特征编码器替换冻结的特征编码器可以显著改善目标域的WER,同时保持源域的性能。SOA的有效性进行了检查各种低资源或域不匹配的ASR设置,包括成人儿童和清洁嘈杂的语音。摘要:Recently, speech foundation models have gained popularity due to their superiority in finetuning downstream ASR tasks. However, models finetuned on certain domains, such as LibriSpeech (adult read speech), behave poorly on other domains (child or noisy speech). One solution could be collecting as much labeled and diverse data as possible for joint finetuning on various domains. However, collecting target domain speech-text paired data and retraining the model is often costly and computationally expensive. In this paper, we introduce a simple yet effective method, speech only adaptation (SOA), based on speech foundation models (Wav2vec 2.0), which requires only speech input data from the target domain. Specifically, the Wav2vec 2.0 feature encoder is continually pretrained with the Wav2vec 2.0 loss on both the source and target domain data for domain adaptation, while the contextual encoder is frozen. Compared to a source domain finetuned model with the feature encoder being frozen during training, we find that replacing the frozen feature encoder with the adapted one provides significant WER improvements to the target domain while preserving the performance of the source domain. The effectiveness of SOA is examined on various low resource or domain mismatched ASR settings, including adult-child and clean-noisy speech.
【17】 Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
标题: 使用监督和自我监督言语基础模型对儿童ASB进行基准测试
作者:Ruchao Fan,Natarajan Balaji Shankar,Abeer Alwan
备注:To appear in Interspeech 2024
链接:点击下载PDF文件
摘要:语音基础模型(SFM)已经在监督(例如Whisper)或自监督系统(例如WavLM)中的各种语音任务中取得了最先进的结果。然而,儿童ASR的SFM的性能还没有得到系统的研究。此外,没有标准评估儿童ASR的基准,这使得比较新颖的想法变得困难。在本文中,我们发起并提出了一个全面的基准几个孩子的语音数据库的基础上,各种SFM(耳语,Wav2vec2.0,HuBERT,和WavLM)。此外,我们研究微调策略,通过比较各种数据增强和参数有效的微调(PEFT)方法。我们观察到,这些方法的行为是不同的,当模型的大小增加。例如,PEFT匹配大模型的完全微调性能,但小模型的性能更差。为了稳定使用增强数据的微调,我们提出了一个扰动不变微调(PIF)损失作为正则化。摘要:Speech foundation models (SFMs) have achieved state-of-the-art results for various speech tasks in supervised (e.g. Whisper) or self-supervised systems (e.g. WavLM). However, the performance of SFMs for child ASR has not been systematically studied. In addition, there is no benchmark for child ASR with standard evaluations, making the comparisons of novel ideas difficult. In this paper, we initiate and present a comprehensive benchmark on several child speech databases based on various SFMs (Whisper, Wav2vec2.0, HuBERT, and WavLM). Moreover, we investigate finetuning strategies by comparing various data augmentation and parameter-efficient finetuning (PEFT) methods. We observe that the behaviors of these methods are different when the model size increases. For example, PEFT matches the performance of full finetuning for large models but worse for small models. To stabilize finetuning using augmented data, we propose a perturbation invariant finetuning (PIF) loss as a regularization.
【18】 AVR: Synergizing Foundation Models for Audio-Visual Humor Detection
标题: AVR:协同视听幽默检测的基础模型
作者:Sarthak Sharma,Orchid Chetia Phukan,Drishti Singh,Arun Balaji Buduru,Rajesh Sharma
备注:Accepted to INTERSPEECH 2024 Show & Tell Demonstrations
链接:点击下载PDF文件
摘要:在这项工作中,我们提出,AVR应用视听幽默检测。虽然幽默检测传统上以文本分析为中心,但最近的进展突出了多模态方法。然而,这些方法依赖于文本线索作为模态,需要使用ASR系统来转录音频数据。这种对ASR准确性的严重依赖可能会在实际应用中带来挑战。为了解决这个瓶颈,我们提出了一个创新的视听幽默检测系统,绕过文本依赖,消除了对ASR模型的需要。相反,所提出的方法取决于音频和视觉内容之间复杂的相互作用,以实现有效的幽默检测。摘要:In this work, we present, AVR application for audio-visual humor detection. While humor detection has traditionally centered around textual analysis, recent advancements have spotlighted multimodal approaches. However, these methods lean on textual cues as a modality, necessitating the use of ASR systems for transcribing the audio-data. This heavy reliance on ASR accuracy can pose challenges in real-world applications. To address this bottleneck, we propose an innovative audio-visual humor detection system that circumvents textual reliance, eliminating the need for ASR models. Instead, the proposed approach hinges on the intricate interplay between audio and visual content for effective humor detection.
【19】 Phoneme Discretized Saliency Maps for Explainable Detection of AI-Generated Voice
标题: 音素离散显着图用于人工智能生成语音的可解释检测
作者:Shubham Gupta,Mirco Ravanelli,Pascal Germain,Cem Subakan
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了音素离散显着性图(PDSM),这是一种显着性图的离散化算法,它利用音素边界对AI生成的语音进行可解释的检测。我们用两种不同的文本到语音系统(即,Tacotron 2和Fastspeech 2),该算法产生的显着图,导致更忠实的解释相比,标准的事后解释方法。此外,通过将显着性图与音素表示相关联,这种方法生成的解释往往比幅度谱图上的标准显着性图更容易理解。摘要:In this paper, we propose Phoneme Discretized Saliency Maps (PDSM), a discretization algorithm for saliency maps that takes advantage of phoneme boundaries for explainable detection of AI-generated voice. We experimentally show with two different Text-to-Speech systems (i.e., Tacotron2 and Fastspeech2) that the proposed algorithm produces saliency maps that result in more faithful explanations compared to standard posthoc explanation methods. Moreover, by associating the saliency maps to the phoneme representations, this methodology generates explanations that tend to be more understandable than standard saliency maps on magnitude spectrograms.
【20】 Evaluating Speaker Identity Coding in Self-supervised Models and Humans
标题: 评估自我监督模型和人类中的说话者身份编码
作者:Gasser Elbanna
备注:Masters Thesis
链接:点击下载PDF文件
摘要:说话人身份在人类交流中发挥着重要作用,并且越来越多地用于社会应用,其中许多是通过机器学习的进步。说话人身份感知是一种重要的认知现象,可以概括为两个主要任务:识别声音或区分声音。一些研究试图确定身份感知的声学相关性,以确定这样一个任务的显着参数。与其他沟通性的社会信号不同,大多数努力都产生了无效的结论。此外,目前的语音识别处理的神经认知模型认为感知的基础是声学维度,如基频,谐波噪声比和共振峰分散。然而,这些研究结果并没有考虑到自然主义的讲话和说话人内的变化。当前自监督模型的表征空间在各种语音相关任务中表现出显着的性能。在这项工作中,我们证明了来自不同家庭的自我监督表示(例如,生成、对比和预测模型)对于说话人识别明显优于声学表示。我们还表明,这样的说话人识别任务可以用来更好地理解这些强大的网络的不同层中的声学信息表示的性质。通过评估声学,音素,韵律和语言变体的说话人识别精度,我们报告模型性能和人类身份感知之间的相似性。我们进一步研究这些相似之处并列的编码空间的模型和人类和挑战使用的距离度量作为一个代理扬声器接近。最后,我们证明了一些模型可以预测自然刺激过程中听觉和语言区域的大脑反应。摘要:Speaker identity plays a significant role in human communication and is being increasingly used in societal applications, many through advances in machine learning. Speaker identity perception is an essential cognitive phenomenon that can be broadly reduced to two main tasks: recognizing a voice or discriminating between voices. Several studies have attempted to identify acoustic correlates of identity perception to pinpoint salient parameters for such a task. Unlike other communicative social signals, most efforts have yielded inefficacious conclusions. Furthermore, current neurocognitive models of voice identity processing consider the bases of perception as acoustic dimensions such as fundamental frequency, harmonics-to-noise ratio, and formant dispersion. However, these findings do not account for naturalistic speech and within-speaker variability. Representational spaces of current self-supervised models have shown significant performance in various speech-related tasks. In this work, we demonstrate that self-supervised representations from different families (e.g., generative, contrastive, and predictive models) are significantly better for speaker identification over acoustic representations. We also show that such a speaker identification task can be used to better understand the nature of acoustic information representation in different layers of these powerful networks. By evaluating speaker identification accuracy across acoustic, phonemic, prosodic, and linguistic variants, we report similarity between model performance and human identity perception. We further examine these similarities by juxtaposing the encoding spaces of models and humans and challenging the use of distance metrics as a proxy for speaker proximity. Lastly, we show that some models can predict brain responses in Auditory and Language regions during naturalistic stimuli.
【21】 Gender Representation in TV and Radio: Automatic Information Extraction methods versus Manual Analyses
标题: 电视和广播中的性别表现:自动信息提取方法与手动分析
作者:David Doukhan,Lena Dodson,Manon Conan,Valentin Pelloin,Aurélien Clamouse,Mélina Lepape,Géraldine Van Hille,Cécile Méadel,Marlène Coulomb-Gully
备注:keywords : Gender representation, computational humanities, TV, Radio, face classification, speaker traits, ASR, media, SLU. Accepted to InterSpeech 2024, Kos Island, Greece, september 2024
链接:点击下载PDF文件
摘要:本研究探讨自动信息提取描述符和人工分析之间的关系,以描述在电视和广播的性别代表性差异。自动描述符,包括语音时间,面部分类和语音transmittance进行了比较,从2023年起,在一个巨大的32,000小时的法语广播语料库的频道报告。调查结果揭示了系统性的性别不平衡,在所有描述中,与男子相比,妇女的代表性不足。值得注意的是,人工渠道报告显示妇女的出席率高于自动估计,提及妇女的时间低于其发言时间。描述符在高和低观众,战争报道,或私人与公共频道之间有着共同的动态。虽然在法国电视中,妇女的形象比声音更明显,但在新闻中,这一趋势却相反,看不见的记者描绘了男性主角。统计检验表明,影响女性参考的主要因素有三个:节目类别、频道和说话人性别。摘要:This study investigates the relationship between automatic information extraction descriptors and manual analyses to describe gender representation disparities in TV and Radio. Automatic descriptors, including speech time, facial categorization and speech transcriptions are compared with channel reports on a vast 32,000-hour corpus of French broadcasts from 2023. Findings reveal systemic gender imbalances, with women underrepresented compared to men across all descriptors. Notably, manual channel reports show higher women's presence than automatic estimates and references to women are lower than their speech time. Descriptors share common dynamics during high and low audiences, war coverage, or private versus public channels. While women are more visible than audible in French TV, this trend is inverted in news with unseen journalists depicting male protagonists. A statistical test shows 3 main effects influencing references to women: program category, channel and speaker gender.
【22】 GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
标题: GAMA:一个具有高级音频理解和复杂推理能力的大型音频语言模型
作者:Sreyan Ghosh,Sonal Kumar,Ashish Seth,Chandra Kiran Reddy Evuru,Utkarsh Tyagi,S Sakshi,Oriol Nieto,Ramani Duraiswami,Dinesh Manocha
备注:Project Website: this https URL
链接:点击下载PDF文件
摘要:感知和理解非言语声音和非言语言语对于帮助我们与周围环境互动的决策至关重要。在本文中,我们提出了GAMA,一种新的通用大型音频语言模型(LALM)与高级音频理解和复杂的推理能力。我们通过将LLM与多种类型的音频表示集成来构建GAMA,包括来自自定义Audio Q-Former的功能,这是一个多层聚合器,可以聚合来自音频编码器多个层的功能。我们在一个大规模的音频语言数据集上对GAMA进行了微调,从而增强了音频理解能力。接下来,我们提出了CompA-R(复杂音频推理的指令调整),这是一个合成生成的指令调整(IT)数据集,其指令要求模型对输入音频执行复杂推理。我们使用CompA-R对GAMA进行了调整,以赋予其复杂的推理能力,其中我们进一步通过利用输入音频的事件标签来添加软提示作为具有高级语义证据的输入。最后,我们还提出了CompA-R-test,一个人类标记的评估数据集,用于评估LALM在需要复杂推理的开放式音频问答上的能力。通过自动化和专家人工评估,我们表明GAMA在各种音频理解任务上的表现优于文献中的所有其他LALM,幅度为1%-84%。此外,CompA-R上的GAMA IT版证明在其复杂推理和指令遵循能力方面具有优势。摘要:Perceiving and understanding non-speech sounds and non-verbal speech is essential to making decisions that help us interact with our surroundings. In this paper, we propose GAMA, a novel General-purpose Large Audio-Language Model (LALM) with Advanced Audio Understanding and Complex Reasoning Abilities. We build GAMA by integrating an LLM with multiple types of audio representations, including features from a custom Audio Q-Former, a multi-layer aggregator that aggregates features from multiple layers of an audio encoder. We fine-tune GAMA on a large-scale audio-language dataset, which augments it with audio understanding capabilities. Next, we propose CompA-R (Instruction-Tuning for Complex Audio Reasoning), a synthetically generated instruction-tuning (IT) dataset with instructions that require the model to perform complex reasoning on the input audio. We instruction-tune GAMA with CompA-R to endow it with complex reasoning abilities, where we further add a soft prompt as input with high-level semantic evidence by leveraging event tags of the input audio. Finally, we also propose CompA-R-test, a human-labeled evaluation dataset for evaluating the capabilities of LALMs on open-ended audio question-answering that requires complex reasoning. Through automated and expert human evaluations, we show that GAMA outperforms all other LALMs in literature on diverse audio understanding tasks by margins of 1%-84%. Further, GAMA IT-ed on CompA-R proves to be superior in its complex reasoning and instruction following capabilities.
【23】 Towards an End-to-End Framework for Invasive Brain Signal Decoding with Large Language Models
标题: 采用大型语言模型建立有创大脑信号解码的端到端框架
作者:Sheng Feng,Heyang Liu,Yu Wang,Yanfeng Wang
链接:点击下载PDF文件
摘要:在本文中,我们介绍了一个突破性的端到端(E2E)框架,用于解码侵入性大脑信号,标志着语音神经假体领域的重大进展。我们的方法利用大型语言模型(LLM)的综合推理能力来促进直接解码。通过完全集成LLM,我们实现了与最先进的级联模型相当的结果。我们的研究结果强调了E2E框架在语音神经假体中的巨大潜力,特别是随着脑机接口(BCI)背后的技术和相关数据集的可用性不断发展。这项工作不仅展示了将LLM与E2E解码相结合用于增强语音神经假体的功效,还为BCI应用的未来研究确定了新的方向,强调了LLM在解码复杂神经信号以恢复通信方面的影响。代码将在https: github.com FsFrancis15 BrainLLM上提供。摘要:In this paper, we introduce a groundbreaking end-to-end (E2E) framework for decoding invasive brain signals, marking a significant advancement in the field of speech neuroprosthesis. Our methodology leverages the comprehensive reasoning abilities of large language models (LLMs) to facilitate direct decoding. By fully integrating LLMs, we achieve results comparable to the state-of-the-art cascade models. Our findings underscore the immense potential of E2E frameworks in speech neuroprosthesis, particularly as the technology behind brain-computer interfaces (BCIs) and the availability of relevant datasets continue to evolve. This work not only showcases the efficacy of combining LLMs with E2E decoding for enhancing speech neuroprosthesis but also sets a new direction for future research in BCI applications, underscoring the impact of LLMs in decoding complex neural signals for communication restoration. Code will be made available at https: github.com FsFrancis15 BrainLLM.
【24】 MusicScore: A Dataset for Music Score Modeling and Generation
标题: MusicScore:乐谱建模和生成的数据集
作者:Yuheng Lin,Zheqi Dai,Qiuqiang Kong
备注:Dataset paper, dataset link: this https URL
链接:点击下载PDF文件
摘要:乐谱是音乐的书面表示,包含有关音乐成分的丰富信息。乐谱上的视觉信息包括音符、休止符、五线谱、谱号、力度和发音。乐谱中的视觉信息比音乐的音频和符号表示包含更多的语义信息。以前的乐谱数据集大小有限,主要用于光学音乐识别(OMR)。目前缺乏关于创建用于音乐建模和生成的大规模基准数据集的研究。在这项工作中,我们提出了MusicScore,一个大规模的乐谱数据集收集和处理的国际乐谱库项目(IMSLP)。MusicScore由图像-文本对组成,其中图像是乐谱的页面,文本是音乐的元数据。MusicScore的元数据取自IMSLP页面的一般信息部分。元数据包括关于音乐作品的作曲家、乐器、作品风格和流派的丰富信息。MusicScore被分别策划成400、14 k和200 k大小的图像-文本对,并具有不同的多样性。我们构建了一个基于UNet扩散模型的乐谱生成系统,以生成视觉可读的乐谱,并以文本描述为条件,以MusicScore数据集为基准进行乐谱生成。MusicScore在https: huggingface.co datasets ZheqiDAI MusicScore上向公众发布。摘要:Music scores are written representations of music and contain rich information about musical components. The visual information on music scores includes notes, rests, staff lines, clefs, dynamics, and articulations. This visual information in music scores contains more semantic information than audio and symbolic representations of music. Previous music score datasets have limited sizes and are mainly designed for optical music recognition (OMR). There is a lack of research on creating a large-scale benchmark dataset for music modeling and generation. In this work, we propose MusicScore, a large-scale music score dataset collected and processed from the International Music Score Library Project (IMSLP). MusicScore consists of image-text pairs, where the image is a page of a music score and the text is the metadata of the music. The metadata of MusicScore is extracted from the general information section of the IMSLP pages. The metadata includes rich information about the composer, instrument, piece style, and genre of the music pieces. MusicScore is curated into small, medium, and large scales of 400, 14k, and 200k image-text pairs with varying diversity, respectively. We build a score generation system based on a UNet diffusion model to generate visually readable music scores conditioned on text descriptions to benchmark the MusicScore dataset for music score generation. MusicScore is released to the public at https: huggingface.co datasets ZheqiDAI MusicScore.
【25】 AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
标题: Anopatch:在机器异常声音检测中实现更好的一致性
作者:Anbai Jiang,Bing Han,Zhiqiang Lv,Yufeng Deng,Wei-Qiang Zhang,Xie Chen,Yanmin Qian,Jia Liu,Pingyi Fan
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:大型预训练模型在多个领域表现出了主导性的性能,其中预训练和微调之间的一致性是成功的关键。然而,很少有工程报告的机器异常声音检测(ASD)任务的预训练模型的令人满意的结果。这可能是由于预训练模型的不一致性和机器音频的归纳偏差,导致数据和架构不一致。因此,我们提出了AnoPatch,它利用在AudioSet上预先训练的ViT骨干,并在机器音频上对其进行微调。人们认为,机器音频与音频数据集比语音数据集更相关,并且从补丁级别对其建模适合机器音频的稀疏性。因此,AnoPatch在DCASE 2020 ASD数据集和DCASE 2023 ASD数据集上展示了最先进的(SOTA)性能。我们还比较了多个预先训练的模型,并通过经验证明了更好的一致性会带来相当大的改进。摘要:Large pre-trained models have demonstrated dominant performances in multiple areas, where the consistency between pre-training and fine-tuning is the key to success. However, few works reported satisfactory results of pre-trained models for the machine anomalous sound detection (ASD) task. This may be caused by the inconsistency of the pre-trained model and the inductive bias of machine audio, resulting in inconsistency in data and architecture. Thus, we propose AnoPatch which utilizes a ViT backbone pre-trained on AudioSet and fine-tunes it on machine audio. It is believed that machine audio is more related to audio datasets than speech datasets, and modeling it from patch level suits the sparsity of machine audio. As a result, AnoPatch showcases state-of-the-art (SOTA) performances on the DCASE 2020 ASD dataset and the DCASE 2023 ASD dataset. We also compare multiple pre-trained models and empirically demonstrate that better consistency yields considerable improvement.
【26】 SMRU: Split-and-Merge Recurrent-based UNet for Acoustic Echo Cancellation and Noise Suppression
标题: SMRU:基于分离合并的回归UNet,用于声学回声消除和噪音抑制
作者:Zhihang Sun,Andong Li,Rilin Chen,Hao Zhang,Meng Yu,Yi Zhou,Dong Yu
链接:点击下载PDF文件
摘要:深度神经网络的激增催生了声学回声消除和噪声抑制的快速发展,并且已经提出了大量现有技术,这些技术产生了有前途的性能。然而,他们很少考虑不同处理场景(如边缘设备和云处理)中的部署通用性。为此,本文提出了一个通用的模型,称为SMRU,以涵盖不同的应用场景。新颖之处在于双重性。首先,提出了多尺度频带分离层和频带合并层,以有效地融合局部频带,从而降低建模复杂度。此外,通过模拟经典UNet结构的多分辨率特征建模特性,设计了一种新的递归支配UNet结构。它由多个可变帧速率块组成,每个块都涉及具有不同压缩比的因果时间下 上采样层以及用于带间和带内建模的双路径结构。该模型配置为从50 M s到6.8 G s的MAC,实验结果表明,所提出的方法产生的竞争力,甚至更好的性能超过现有的基线,并有充分的潜力,以适应更一般的情况下,不同的复杂性要求。摘要:The proliferation of deep neural networks has spawned the rapid development of acoustic echo cancellation and noise suppression, and plenty of prior arts have been proposed, which yield promising performance. Nevertheless, they rarely consider the deployment generality in different processing scenarios, such as edge devices, and cloud processing. To this end, this paper proposes a general model, termed SMRU, to cover different application scenarios. The novelty lies in two-fold. First, a multi-scale band split layer and band merge layer are proposed to effectively fuse local frequency bands for lower complexity modeling. Besides, by simulating the multi-resolution feature modeling characteristic of the classical UNet structure, a novel recurrent-dominated UNet is devised. It consists of multiple variable frame rate blocks, each of which involves the causal time down- up-sampling layer with varying compression ratios and the dual-path structure for inter- and intra-band modeling. The model is configured from 50 M s to 6.8 G s in terms of MACs, and the experimental results show that the proposed approach yields competitive or even better performance over existing baselines, and has the full potential to adapt to more general scenarios with varying complexity requirements.
【27】 Identification of Physical Properties in Acoustic Tubes Using Physics-Informed Neural Networks
标题: 使用物理信息神经网络识别声管的物理性能
作者:Kazuya Yokota,Masataka Ogura,Masajiro Abe
备注:10 pages, 7 figures, The following article has been submitted to Mechanical Engineering Journal. After it is published, it will be found at this https URL
链接:点击下载PDF文件
摘要:物理信息神经网络(Physics-informed Neural Networks,PINN)是一种数值模拟方法,它将与控制方程对应的损失函数合并到神经网络中。虽然PINN已经被探索用于逆分析,但它们在声学分析中的应用仍然有限。本文提出了一种利用PINNs识别声管内损耗参数的方法。我们将损失参数分为两类:一类依赖于管道直径,另一类是与管道直径无关的常数,后者被设置为神经网络的可训练参数。将损耗参数的确定问题转化为一个优化问题,通过这个过程确定材料的物理性质。所采用的神经网络架构是基于我们以前提出的ResoNet,这是专为分析声学共振。所提出的方法的有效性进行评估,通过正向和反向分析,特别是通过识别的损失参数。研究结果表明,它是可行的,以准确地识别参数,显着影响下的声场分析。通过改变损失函数中的控制方程,该方法可以适用于各种声场,具有广泛的应用前景。摘要:Physics-informed Neural Networks (PINNs) is a method for numerical simulation that incorporates a loss function corresponding to the governing equations into a neural network. While PINNs have been explored for their utility in inverse analysis, their application in acoustic analysis remains limited. This study presents a method to identify loss parameters in acoustic tubes using PINNs. We categorized the loss parameters into two groups: one dependent on the tube's diameter and another constant, independent of it. The latter were set as the trainable parameters of the neural network. The problem of identifying the loss parameter was formulated as an optimization problem, with the physical properties being determined through this process. The neural network architecture employed was based on our previously proposed ResoNet, which is designed for analyzing acoustic resonance. The efficacy of the proposed method is assessed through both forward and inverse analysis, specifically through the identification of loss parameters. The findings demonstrate that it is feasible to accurately identify parameters that significantly impact the sound field under analysis. By merely altering the governing equations in the loss function, this method could be adapted to various sound fields, suggesting its potential for broad application.
【28】 NAST: Noise Aware Speech Tokenization for Speech Language Models
标题: NAST:语音语言模型的噪音感知语音令牌化
作者:Shoval Messica,Yossi Adi
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:语音标记化是将语音信号表示为离散单元序列的任务。这样的表示可以用于各种下游任务,包括自动语音识别,文本到语音等更相关的这项研究,这样的表示作为语音语言模型的基础。在这项工作中,我们解决了在噪声环境下的语音标记化任务,并提出了NAST:噪声感知语音标记化的语音语言模型。NAST由三个主要组件组成:(i)预测器;(ii)残差编码器;以及(iii)解码器。我们评估NAST的效率,考虑几个口语建模任务,并表明NAST优于所有设置的评估基线。最后,我们分析了NAST,并显示其解纠缠特性和鲁棒性的信号变化的形式的噪声,混响,音高移位,和时间拉伸。代码和预训练模型可在https: github.com ShovalMessica NAST上获得。摘要:Speech tokenization is the task of representing speech signals as a sequence of discrete units. Such representations can be later used for various downstream tasks including automatic speech recognition, text-to-speech, etc. More relevant to this study, such representation serves as the basis of Speech Language Models. In this work, we tackle the task of speech tokenization under the noisy setup and present NAST: Noise Aware Speech Tokenization for Speech Language Models. NAST is composed of three main components: (i) a predictor; (ii) a residual encoder; and (iii) a decoder. We evaluate the efficiency of NAST considering several spoken language modeling tasks and show that NAST is superior to the evaluated baselines across all setups. Lastly, we analyze NAST and show its disentanglement properties and robustness to signal variations in the form of noise, reverberation, pitch-shift, and time-stretch. Code and pre-trained models are available at https: github.com ShovalMessica NAST.
【29】 Large Language Models for Dysfluency Detection in Stuttered Speech
标题: 用于口吃语音流畅性检测的大型语言模型
作者:Dominik Wagner,Sebastian P. Bayerl,Ilja Baumann,Korbinian Riedhammer,Elmar Nöth,Tobias Bocklet
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:准确检测口语中的不流利可以帮助提高自动语音和语言处理组件的性能,并支持开发更具包容性的语音和语言技术。受最近部署大型语言模型(LLM)作为非词汇输入(如音频和视频)的通用学习器和处理器的趋势的启发,我们将多标签不流利检测任务作为语言建模问题。我们提出的假设候选人产生的自动语音识别系统和声学表示提取的音频编码器模型的LLM,微调系统预测不流利的标签上的三个数据集包含英语和德语口吃的语音。实验结果表明,我们的系统有效地结合了声学和词汇信息,并取得了多标签口吃检测任务的竞争力的结果。摘要:Accurately detecting dysfluencies in spoken language can help to improve the performance of automatic speech and language processing components and support the development of more inclusive speech and language technologies. Inspired by the recent trend towards the deployment of large language models (LLMs) as universal learners and processors of non-lexical inputs, such as audio and video, we approach the task of multi-label dysfluency detection as a language modeling problem. We present hypotheses candidates generated with an automatic speech recognition system and acoustic representations extracted from an audio encoder model to an LLM, and finetune the system to predict dysfluency labels on three datasets containing English and German stuttered speech. The experimental results show that our system effectively combines acoustic and lexical information and achieves competitive results on the multi-label stuttering detection task.
【30】 Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
标题: 利用门控注意力减少离群值以改进大型序列到序列语音基础模型中的训练后量化
作者:Dominik Wagner,Ilja Baumann,Korbinian Riedhammer,Tobias Bocklet
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:本文探讨了在Whisper语音基础模型族中进行知识蒸馏后的后训练量化(PTQ)的改进。我们解决了权重和激活张量中离群值的挑战,已知这些离群值会阻碍基于变换的语言和视觉模型中的量化质量。将这一观察结果扩展到Whisper,我们证明了当基于transformer的模型被训练来执行自动语音识别时,这些离群值也存在,从而需要PTQ的缓解策略。我们表明,离群值可以减少最近提出的门控机制,在学生模型的注意力块,使有效的8位量化,和较低的字错误率相比,学生模型没有门控机制到位。摘要:This paper explores the improvement of post-training quantization (PTQ) after knowledge distillation in the Whisper speech foundation model family. We address the challenge of outliers in weights and activation tensors, known to impede quantization quality in transformer-based language and vision models. Extending this observation to Whisper, we demonstrate that these outliers are also present when transformer-based models are trained to perform automatic speech recognition, necessitating mitigation strategies for PTQ. We show that outliers can be reduced by a recently proposed gating mechanism in the attention blocks of the student model, enabling effective 8-bit quantization, and lower word error rates compared to student models without the gating mechanism in place.
【31】 SPEAR: Receiver-to-Receiver Acoustic Neural Warping Field
标题: SPSYS:接收器到接收器的声学神经扭曲场
作者:Yuhang He,Shitong Xu,Jia-Xing Zhong,Sangyun Shin,Niki Trigoni,Andrew Markham
备注:9 pages, 5 figures in main paper
链接:点击下载PDF文件
摘要:我们提出了一个连续的接收器到接收器的声学神经扭曲场,用于在具有单个固定音频源的声学3D空间中进行空间声学效果预测。与传统的源到接收器建模方法,需要先前的空间声学特性知识,严格地模拟音频传播从源到接收器,我们建议预测通过扭曲的空间声学效果从一个参考接收器位置到另一个目标接收器位置,使扭曲的音频基本上容纳所有的空间声学效果属于目标位置。SPARK可以以一种更容易获得数据的方式进行训练,我们只需让两个机器人在不同的位置独立地记录空间音频。我们进一步从理论上证明了翘曲场的普遍存在当且仅当一个音频源存在。三个物理原则被纳入到指导SPRINT网络设计,导致学习翘曲场物理意义。我们在合成的、照片般逼真的和真实世界的数据集上展示了SPARTS的优越性,展示了SPARTS在各种下游机器人任务中的巨大潜力。摘要:We present SPEAR, a continuous receiver-to-receiver acoustic neural warping field for spatial acoustic effects prediction in an acoustic 3D space with a single stationary audio source. Unlike traditional source-to-receiver modelling methods that require prior space acoustic properties knowledge to rigorously model audio propagation from source to receiver, we propose to predict by warping the spatial acoustic effects from one reference receiver position to another target receiver position, so that the warped audio essentially accommodates all spatial acoustic effects belonging to the target position. SPEAR can be trained in a data much more readily accessible manner, in which we simply ask two robots to independently record spatial audio at different positions. We further theoretically prove the universal existence of the warping field if and only if one audio source presents. Three physical principles are incorporated to guide SPEAR network design, leading to the learned warping field physically meaningful. We demonstrate SPEAR superiority on both synthetic, photo-realistic and real-world dataset, showing the huge potential of SPEAR to various down-stream robotic tasks.
【32】 CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving
标题: CoSTA:使用对齐语音文本交织的代码交换语音翻译
作者:Bhavani Shankar,Preethi Jyothi,Pushpak Bhattacharyya
链接:点击下载PDF文件
摘要:语码转换是印度等多语言社会中普遍存在的语言现象。由于数据集的可用性有限,为代码切换语音构建语音到文本模型具有挑战性。在这项工作中,我们专注于口语翻译(ST)的问题,在印度语言的代码转换语音到英语文本。我们提出了一个新的端到端模型架构COSTA,它以预训练的自动语音识别(ASR)和机器翻译(MT)模块(适用于许多语言)为基础。语音和ASR文本表示使用对齐的交织方案进行融合,并进一步作为输入馈送到预训练的MT模块;然后使用合成创建的ST数据对整个管道进行端到端的口语翻译训练。我们还发布了一个新的评估基准代码切换孟加拉语英语,印地语英语,马拉地语英语和泰卢固语英语语音到英语文本。COSTA显著优于许多具有竞争力的级联和端到端多模式基线,最高可达3.5个BLEU点。摘要:Code-switching is a widely prevalent linguistic phenomenon in multilingual societies like India. Building speech-to-text models for code-switched speech is challenging due to limited availability of datasets. In this work, we focus on the problem of spoken translation (ST) of code-switched speech in Indian languages to English text. We present a new end-to-end model architecture COSTA that scaffolds on pretrained automatic speech recognition (ASR) and machine translation (MT) modules (that are more widely available for many languages). Speech and ASR text representations are fused using an aligned interleaving scheme and are fed further as input to a pretrained MT module; the whole pipeline is then trained end-to-end for spoken translation using synthetically created ST data. We also release a new evaluation benchmark for code-switched Bengali-English, Hindi-English, Marathi-English and Telugu- English speech to English text. COSTA significantly outperforms many competitive cascaded and end-to-end multimodal baselines by up to 3.5 BLEU points.
【33】 Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
标题: 用于时间控制文本到音乐生成的联合音频和符号条件处理
作者:Or Tal,Alon Ziv,Itai Gat,Felix Kreuk,Yossi Adi
链接:点击下载PDF文件
摘要:我们提出了JASCO,一个时间控制的文本到音乐生成模型,利用符号和基于音频的条件。JASCO可以生成高质量的音乐样本,条件是全局文本描述以及细粒度的本地控件。JASCO是基于流匹配建模范式与一种新的空调方法。这允许本地控制的音乐生成(例如,和弦)和全局(文本描述)。具体来说,我们应用信息瓶颈层结合时间模糊提取相关信息的特定控件。这允许在相同的文本到音乐模型中结合符号和基于音频的条件。我们用各种符号控制信号进行实验(例如,和弦,旋律),以及音频表示(例如,分离的鼓轨道,全混合)。我们评估JASCO同时考虑发电质量和条件的坚持,使用客观指标和人体研究。结果表明,JASCO是可比的,考虑到生成质量的评估基线,同时允许显着更好,更灵活的控制所生成的音乐。示例可在我们的演示页面https: pages.cs.huji.ac.il adiyoss-lab JASCO上获取。摘要:We present JASCO, a temporally controlled text-to-music generation model utilizing both symbolic and audio-based conditions. JASCO can generate high-quality music samples conditioned on global text descriptions along with fine-grained local controls. JASCO is based on the Flow Matching modeling paradigm together with a novel conditioning method. This allows music generation controlled both locally (e.g., chords) and globally (text description). Specifically, we apply information bottleneck layers in conjunction with temporal blurring to extract relevant information with respect to specific controls. This allows the incorporation of both symbolic and audio-based conditions in the same text-to-music model. We experiment with various symbolic control signals (e.g., chords, melody), as well as with audio representations (e.g., separated drum tracks, full-mix). We evaluate JASCO considering both generation quality and condition adherence, using both objective metrics and human studies. Results suggest that JASCO is comparable to the evaluated baselines considering generation quality while allowing significantly better and more versatile controls over the generated music. Samples are available on our demo page https: pages.cs.huji.ac.il adiyoss-lab JASCO.
【34】 Robust Channel Learning for Large-Scale Radio Speaker Verification
标题: 用于大规模无线电扬声器验证的稳健通道学习
作者:Wenhao Yang,Jianguo Wei,Wenhuan Lu,Lei Li,Xugang Lu
备注:12 pages, 11 figures
链接:点击下载PDF文件
摘要:说话人确认的研究越来越多地集中在具有挑战性的信道条件和噪声环境下实现鲁棒和可靠的识别。在无线电通信中识别说话者是特别困难的,这是由于固有的限制,如有限的带宽和普遍的噪声干扰。为了解决这个问题,我们提出了一个通道鲁棒说话人学习(CRSL)框架,提高了当前说话人验证管道的鲁棒性,考虑到数据源,数据增强和模型传输过程的效率。我们的框架引入了一个增强模块,通过操纵训练输入的带宽来减轻无线电语音数据集的带宽变化。它还通过在流形空间内引入噪声来解决未知噪声。此外,我们提出了一种有效的微调方法,减少了对大量额外训练时间和大量数据的需求。此外,我们开发了一个工具包,用于组装一个大规模的无线电语音语料库,并建立了一个专门为无线电场景说话人验证研究量身定制的基准。实验结果表明,我们提出的方法有效地提高了性能,并减轻在说话人确认任务中无线电传输所造成的退化。代码将在Github上提供。摘要:Recent research in speaker verification has increasingly focused on achieving robust and reliable recognition under challenging channel conditions and noisy environments. Identifying speakers in radio communications is particularly difficult due to inherent limitations such as constrained bandwidth and pervasive noise interference. To address this issue, we present a Channel Robust Speaker Learning (CRSL) framework that enhances the robustness of the current speaker verification pipeline, considering data source, data augmentation, and the efficiency of model transfer processes. Our framework introduces an augmentation module that mitigates bandwidth variations in radio speech datasets by manipulating the bandwidth of training inputs. It also addresses unknown noise by introducing noise within the manifold space. Additionally, we propose an efficient fine-tuning method that reduces the need for extensive additional training time and large amounts of data. Moreover, we develop a toolkit for assembling a large-scale radio speech corpus and establish a benchmark specifically tailored for radio scenario speaker verification studies. Experimental results demonstrate that our proposed methodology effectively enhances performance and mitigates degradation caused by radio transmission in speaker verification tasks. The code will be available on Github.
【35】 Imperceptible Rhythm Backdoor Attacks: Exploring Rhythm Transformation for Embedding Undetectable Vulnerabilities on Speech Recognition
标题: 不可感知的节奏后门攻击:探索节奏转换以在语音识别中嵌入不可检测的漏洞
作者:Wenhan Yao,Jiangkun Yang,Yongqiang He,Jia Liu,Weiping Wen
链接:点击下载PDF文件
摘要:语音识别是人机交互的重要起点,最近,深度学习模型在这项任务中取得了巨大的成功。然而,当模型训练和私有数据提供者总是分离时,一些使深度神经网络(DNN)异常的安全威胁值得研究。近年来,语音识别系统中的典型后门攻击已成为研究热点。现有的后门方法是基于数据中毒。攻击者将一些合并的变化添加到良性语音频谱图或改变语音成分,如音高和音色。因此,中毒数据可以通过人类听觉或自动深度算法检测到。为了提高数据中毒的隐蔽性,本文提出了一种非神经网络的快速算法--随机谱图节奏变换(RSRT)。该算法结合了四个步骤来生成隐形有毒话语。从节奏成分转换的角度来看,我们提出的触发拉伸或挤压梅尔频谱图,并恢复他们回到信号。该操作保持音色和内容不变,具有良好的隐蔽性。我们的实验是在两种语音识别任务上进行的,包括通过说话人确认和自动语音识别来测试中毒样本的隐蔽性。实验结果表明,该方法具有良好的有效性和隐蔽性。节奏触发需要低中毒率,并获得非常高的攻击成功率。摘要:Speech recognition is an essential start ring of human-computer interaction, and recently, deep learning models have achieved excellent success in this task. However, when the model training and private data provider are always separated, some security threats that make deep neural networks (DNNs) abnormal deserve to be researched. In recent years, the typical backdoor attacks have been researched in speech recognition systems. The existing backdoor methods are based on data poisoning. The attacker adds some incorporated changes to benign speech spectrograms or changes the speech components, such as pitch and timbre. As a result, the poisoned data can be detected by human hearing or automatic deep algorithms. To improve the stealthiness of data poisoning, we propose a non-neural and fast algorithm called Random Spectrogram Rhythm Transformation (RSRT) in this paper. The algorithm combines four steps to generate stealthy poisoned utterances. From the perspective of rhythm component transformation, our proposed trigger stretches or squeezes the mel spectrograms and recovers them back to signals. The operation keeps timbre and content unchanged for good stealthiness. Our experiments are conducted on two kinds of speech recognition tasks, including testing the stealthiness of poisoned samples by speaker verification and automatic speech recognition. The results show that our method has excellent effectiveness and stealthiness. The rhythm trigger needs a low poisoning rate and gets a very high attack success rate.
【36】 SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
标题: SingMOS:用于MOS预测的广泛开源歌唱声音数据集
作者:Yuxun Tang,Jiatong Shi,Yuning Wu,Qin Jin
链接:点击下载PDF文件
摘要:在语音生成任务中,人的主观评分,通常被称为意见分数,被认为是语音质量评估的“金标准”,平均意见分数(MOS)作为主要的评估指标。由于人工注释的高成本,在语音领域出现了几种MOS预测系统,表现出良好的性能。这些MOS预测模型使用来自先前语音相关挑战的注释进行训练。然而,与语音领域相比,歌唱领域面临着数据稀缺和更严格的版权保护,导致缺乏高质量的MOS注释歌唱数据集。为了解决这个问题,我们提出了SingMOS,这是一个高质量和多样化的MOS歌唱数据集,涵盖了一系列中国和日本的数据集。这些合成的人声是使用歌唱合成、转换或再合成任务中最先进的模型生成的,并由专业注释者与真实人声一起进行评级。数据分析证明了我们数据集的多样性和可靠性。此外,我们还对SingMOS进行了进一步的探索,为SingMOS的预测提供了见解,并为SingMOS的持续扩展提供了指导。摘要:In speech generation tasks, human subjective ratings, usually referred to as the opinion score, are considered the "gold standard" for speech quality evaluation, with the mean opinion score (MOS) serving as the primary evaluation metric. Due to the high cost of human annotation, several MOS prediction systems have emerged in the speech domain, demonstrating good performance. These MOS prediction models are trained using annotations from previous speech-related challenges. However, compared to the speech domain, the singing domain faces data scarcity and stricter copyright protections, leading to a lack of high-quality MOS-annotated datasets for singing. To address this, we propose SingMOS, a high-quality and diverse MOS dataset for singing, covering a range of Chinese and Japanese datasets. These synthesized vocals are generated using state-of-the-art models in singing synthesis, conversion, or resynthesis tasks and are rated by professional annotators alongside real vocals. Data analysis demonstrates the diversity and reliability of our dataset. Additionally, we conduct further exploration on SingMOS, providing insights for singing MOS prediction and guidance for the continued expansion of SingMOS.
【37】 Optimizing Automatic Speech Assessment: W-RankSim Regularization and Hybrid Feature Fusion Strategies
标题: 优化自动语音评估:W-RankSim正规化和混合特征融合策略
作者:Chung-Wen Wu,Berlin Chen
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:在最近的研究中,自动语音评估(ASA)随着自监督特征(SSL)的利用而取得了显着的进步。然而,ASA的一个关键挑战在于数据的不平衡分布,特别是在英语测试数据集中。为了解决这一挑战,我们将ASA作为一个有序分类任务,引入加权向量排序相似性(W-RankSim)作为一种新的正则化技术。W-RankSim鼓励输出层中相似类的加权向量更接近,这意味着具有相似标签的特征向量将随着它们向相应的加权向量收敛而逐渐相互靠近。广泛的实验评估证实了我们的方法在提高ASA的顺序分类性能的有效性。此外,我们提出了一个混合模型,结合SSL和手工制作的功能,展示如何包含手工制作的功能,提高性能的ASA系统。摘要:Automatic Speech Assessment (ASA) has seen notable advancements with the utilization of self-supervised features (SSL) in recent research. However, a key challenge in ASA lies in the imbalanced distribution of data, particularly evident in English test datasets. To address this challenge, we approach ASA as an ordinal classification task, introducing Weighted Vectors Ranking Similarity (W-RankSim) as a novel regularization technique. W-RankSim encourages closer proximity of weighted vectors in the output layer for similar classes, implying that feature vectors with similar labels would be gradually nudged closer to each other as they converge towards corresponding weighted vectors. Extensive experimental evaluations confirm the effectiveness of our approach in improving ordinal classification performance for ASA. Furthermore, we propose a hybrid model that combines SSL and handcrafted features, showcasing how the inclusion of handcrafted features enhances performance in an ASA system.
【38】 Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
标题: 使用CNN的语音情感识别及其在数字医疗保健中的用例
作者:Nishargo Nigar
备注:Master's Thesis at Hamburg University of Technology
链接:点击下载PDF文件
摘要:从语音中识别人类情感和情感状态的过程被称为语音情感识别(SER)。这是基于这样的观察,即声音中的音调和音高经常传达潜在的情感。语音识别包括识别情感的能力,这变得越来越流行,需求量也越来越大。在适当因素的帮助下(如模式,情绪,强度,重复等)基于数据中发现的情感,我的研究试图使用卷积神经网络(CNN)来区分音频记录中的情感,并根据不同情感的范围对其进行标记。我开发了一个机器学习模型,可以借助机器学习方法从提供的音频文件中识别情绪。评估主要集中在精度,召回率和F1分数,这些都是常见的机器学习指标。为了正确地建立和训练机器学习框架,主要目标是研究所有输入和输出参数的影响和相互关系。为了提高识别意图的能力,这是沟通的一个关键条件,我使用我的专业机器学习算法通过语音来评估情绪,该算法将在数字医疗保健的帮助下从语音中解决情绪状态,弥合人类和人工智能(AI)之间的差距。摘要:The process of identifying human emotion and affective states from speech is known as speech emotion recognition (SER). This is based on the observation that tone and pitch in the voice frequently convey underlying emotion. Speech recognition includes the ability to recognize emotions, which is becoming increasingly popular and in high demand. With the help of appropriate factors (such modalities, emotions, intensities, repetitions, etc.) found in the data, my research seeks to use the Convolutional Neural Network (CNN) to distinguish emotions from audio recordings and label them in accordance with the range of different emotions. I have developed a machine learning model to identify emotions from supplied audio files with the aid of machine learning methods. The evaluation is mostly focused on precision, recall, and F1 score, which are common machine learning metrics. To properly set up and train the machine learning framework, the main objective is to investigate the influence and cross-relation of all input and output parameters. To improve the ability to recognize intentions, a key condition for communication, I have evaluated emotions using my specialized machine learning algorithm via voice that would address the emotional state from voice with the help of digital healthcare, bridging the gap between human and artificial intelligence (AI).
【39】 How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
标题: 我们应该如何从自我监督模型中提取离散音频令牌?
作者:Pooneh Mousavi,Jarod Duret,Salah Zaiem,Luca Della Libera,Artem Ploujnikov,Cem Subakan,Mirco Ravanelli
备注:4 pages, 2 figures, 2 tables, Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:离散音频令牌最近因其在弥合音频和语言处理之间的差距方面的潜力而受到关注。理想的音频令牌必须保留内容、非语言元素、说话者身份和许多其他音频细节。当前的音频标记化方法分为两类:通过自监督学习(SSL)模型的量化获得的语义标记,以及基于神经压缩的标记(编解码器)。虽然以前的研究已经对编解码器模型进行了基准测试以确定最佳配置,但量化预训练SSL模型的理想设置仍然不清楚。本文探讨了在区分和生成任务中语义标记的最佳配置。我们提出了一个可扩展的解决方案,训练跨多个SSL层的通用声码器。此外,注意力机制被用来识别特定于任务的影响层,增强了语义令牌在不同音频应用中的适应性和性能。摘要:Discrete audio tokens have recently gained attention for their potential to bridge the gap between audio and language processing. Ideal audio tokens must preserve content, paralinguistic elements, speaker identity, and many other audio details. Current audio tokenization methods fall into two categories: Semantic tokens, acquired through quantization of Self-Supervised Learning (SSL) models, and Neural compression-based tokens (codecs). Although previous studies have benchmarked codec models to identify optimal configurations, the ideal setup for quantizing pretrained SSL models remains unclear. This paper explores the optimal configuration of semantic tokens across discriminative and generative tasks. We propose a scalable solution to train a universal vocoder across multiple SSL layers. Furthermore, an attention mechanism is employed to identify task-specific influential layers, enhancing the adaptability and performance of semantic tokens in diverse audio applications.
【40】 Enhancing Multilingual Voice Toxicity Detection with Speech-Text Alignment
标题: 通过语音文本对齐增强多语言语音毒性检测
作者:Joseph Liu,Mahesh Kumar Nandwana,Janne Pylkkönen,Hannes Heikinheimo,Morgan McGuire
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:语音的毒性分类严重依赖于语音的语义内容。我们提出了一个新的框架,利用跨模态学习集成到一个多标签语音毒性分类器在训练过程中的文本的语义嵌入。这使我们能够在训练过程中结合文本信息,同时在推理过程中仍然只需要音频。我们在具有真实世界特征的大规模数据集上对该分类器进行了评估,以验证该框架的有效性。通过消融研究,我们证明了通用语义文本嵌入是丰富的,并与语音毒性分类的目的。通过在多种语言中进行大规模的实验,我们发现五种语言和不同毒性类别的语音毒性分类有所改善。摘要:Toxicity classification for voice heavily relies on the semantic content of speech. We propose a novel framework that utilizes cross-modal learning to integrate the semantic embedding of text into a multilabel speech toxicity classifier during training. This enables us to incorporate textual information during training while still requiring only audio during inference. We evaluate this classifier on large-scale datasets with real-world characteristics to validate the effectiveness of this framework. Through ablation studies, we demonstrate that general-purpose semantic text embeddings are rich and aligned with speech for toxicity classification purposes. Conducting experiments across multiple languages at scale, we show improvements in voice toxicity classification across five languages and different toxicity categories.
【41】 Improving child speech recognition with augmented child-like speech
标题: 通过增强的儿童语音来提高儿童语音识别
作者:Yuanyuan Zhang,Zhengjun Yue,Tanvina Patel,Odette Scharenborg
备注:5 pages, 1 figure Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:最先进的ASR显示出儿童语音的次优性能。儿童语言的缺乏限制了儿童语音识别的发展。因此,我们研究了儿童到儿童的语音转换(VC),从现有的儿童扬声器的数据集和额外的(新的)儿童扬声器通过单语和跨语言(荷兰语到德语)VC,分别。结果表明,跨语言儿童对儿童VC显着提高儿童ASR性能。关于儿童对儿童跨语言VC生成的数据量对微调(FT)ASR模型的影响的实验给出了最佳结果,其中我们的FT-Conformer模型和FT-Whisper模型的两倍增强与基线相比减少了约3%的绝对WER,并且从头开始训练的模型的六倍增强提高了绝对3.6% WER。此外,使用少量的“高质量”VC生成的数据,获得了与我们最好的FT模型相似的结果。摘要:State-of-the-art ASRs show suboptimal performance for child speech. The scarcity of child speech limits the development of child speech recognition (CSR). Therefore, we studied child-to-child voice conversion (VC) from existing child speakers in the dataset and additional (new) child speakers via monolingual and cross-lingual (Dutch-to-German) VC, respectively. The results showed that cross-lingual child-to-child VC significantly improved child ASR performance. Experiments on the impact of the quantity of child-to-child cross-lingual VC-generated data on fine-tuning (FT) ASR models gave the best results with two-fold augmentation for our FT-Conformer model and FT-Whisper model which reduced WERs with ~3% absolute compared to the baseline, and with six-fold augmentation for the model trained from scratch, which improved by an absolute 3.6% WER. Moreover, using a small amount of "high-quality" VC-generated data achieved similar results to those of our best-FT models.
【42】 Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
标题: 精心合并预训练语音模型中的隐藏嵌入以实现反欺骗检测
作者:Zihan Pan,Tianchi Liu,Hardik B. Sailor,Qiongqiong Wang
链接:点击下载PDF文件
摘要:在大型语音语料库上训练的自监督学习(SSL)语音表示模型已经证明了通过多个Transformer层提取分层语音嵌入的有效性。然而,这些嵌入在特定任务中的行为仍然不确定。本文研究了WavLM模型在反欺骗中的多层行为,并提出了一种注意的合并方法来利用层次隐藏嵌入。结果证明了微调WavLM的可行性,以在ASVspoof 2019LA,2021LA和2021DF评估集上分别实现0.65%,3.50%和3.19%的最佳等误率(EER)。值得注意的是,我们发现WavLM大模型的早期隐藏Transformer层对反欺骗任务有显著贡献,通过利用部分预训练模型实现计算效率。摘要:Self-supervised learning (SSL) speech representation models, trained on large speech corpora, have demonstrated effectiveness in extracting hierarchical speech embeddings through multiple transformer layers. However, the behavior of these embeddings in specific tasks remains uncertain. This paper investigates the multi-layer behavior of the WavLM model in anti-spoofing and proposes an attentive merging method to leverage the hierarchical hidden embeddings. Results demonstrate the feasibility of fine-tuning WavLM to achieve the best equal error rate (EER) of 0.65%, 3.50%, and 3.19% on the ASVspoof 2019LA, 2021LA, and 2021DF evaluation sets, respectively. Notably, We find that the early hidden transformer layers of the WavLM large model contribute significantly to anti-spoofing task, enabling computational efficiency by utilizing a partial pre-trained model.
【43】 Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
标题: 模糊不可知多对一端到端语音翻译的软语言识别
作者:Peidong Wang,Jian Xue,Jinyu Li,Junkun Chen,Aswin Shanmugam Subramanian
链接:点击下载PDF文件
摘要:语音不可知的多对一端到端语音翻译模型可以将来自不同源语言的音频信号转换为目标语言的文本。这些模型不需要源语言识别,这提高了用户体验。在某些情况下,输入语言可以被给定或估计。我们的目标是使用这些额外的语言信息,同时保持其他语言的质量。我们通过引入一个简单有效的线性输入网络来实现这一点。线性输入网络被初始化为单位矩阵,这确保模型可以与原始模型一样好,甚至更好。实验结果表明,该方法可以成功地增强指定的语言,同时保持语言无关的能力的多对一ST模型。摘要:Language-agnostic many-to-one end-to-end speech translation models can convert audio signals from different source languages into text in a target language. These models do not need source language identification, which improves user experience. In some cases, the input language can be given or estimated. Our goal is to use this additional language information while preserving the quality of the other languages. We accomplish this by introducing a simple and effective linear input network. The linear input network is initialized as an identity matrix, which ensures that the model can perform as well as, or better than, the original model. Experimental results show that the proposed method can successfully enhance the specified language, while keeping the language-agnostic ability of the many-to-one ST models.
【44】 Connected Speech-Based Cognitive Assessment in Chinese and English
标题: 中文和英语的连接言语认知评估
作者:aturnino Luz,Sofia De La Fuente Garcia,Fasih Haider,Davida Fromm,Brian MacWhinney,Alyssa Lanzi,Ya-Ning Chang,Chia-Ju Chou,Yi-Chien Liu
备注:To appear in Proceedings of Interspeech 2024
链接:点击下载PDF文件
摘要:我们提出了一种新的基准数据集和预测任务,用于研究通过分析连接语音来评估认知功能的方法。该数据集包括具有不同程度认知障碍的汉语普通话和英语使用者以及具有正常认知的个体的语音样本和临床信息。这些数据已通过倾向评分分析按年龄和性别仔细匹配,以确保模型训练的平衡性和代表性。预测任务包括轻度认知障碍诊断和认知测试分数预测。该框架旨在鼓励开发基于语音的认知评估方法,这些方法可以概括各种语言。我们通过提出基线预测模型来说明这一点,这些模型采用与语言无关的和可比较的特征来进行诊断和认知测试分数预测。模型在诊断中的平均召回率为59.2%,评分预测的均方根误差为2.89。摘要:We present a novel benchmark dataset and prediction tasks for investigating approaches to assess cognitive function through analysis of connected speech. The dataset consists of speech samples and clinical information for speakers of Mandarin Chinese and English with different levels of cognitive impairment as well as individuals with normal cognition. These data have been carefully matched by age and sex by propensity score analysis to ensure balance and representativity in model training. The prediction tasks encompass mild cognitive impairment diagnosis and cognitive test score prediction. This framework was designed to encourage the development of approaches to speech-based cognitive assessment which generalise across languages. We illustrate it by presenting baseline prediction models that employ language-agnostic and comparable features for diagnosis and cognitive test score prediction. The models achieved unweighted average recall was 59.2% in diagnosis, and root mean squared error of 2.89 in score prediction.
【45】 Towards Signal Processing In Large Language Models
标题: 走向大型语言模型中的信号处理
作者:Prateek Verma,Mert Pilanci
备注:12 pages, 3 figures
链接:点击下载PDF文件
摘要:本文介绍了在大型语言模型(LLM)中应用信号处理的思想。随着最近生成式人工智能的爆发,我们的工作可以帮助将两个领域连接在一起,即信号处理领域和大型语言模型。我们绘制了经典的傅里叶变换和傅里叶变换式的可学习的时间-频率表示的LLM的每一个中间激活信号之间的平行。一旦我们将令牌上的每个激活信号分解为时频表示,我们就可以学习如何过滤和重建它们,从头开始学习所有组件,以预测给定先前上下文的下一个令牌。我们表明,对于类似GPT的架构,我们的工作实现了更快的收敛,并通过在相同时期的训练中添加少量的额外参数来显着提高性能。我们希望这项工作为算法探索LLM等神经架构中发现的信号内部的信号处理铺平道路。摘要:This paper introduces the idea of applying signal processing inside a Large Language Model (LLM). With the recent explosion of generative AI, our work can help bridge two fields together, namely the field of signal processing and large language models. We draw parallels between classical Fourier-Transforms and Fourier Transform-like learnable time-frequency representations for every intermediate activation signal of an LLM. Once we decompose every activation signal across tokens into a time-frequency representation, we learn how to filter and reconstruct them, with all components learned from scratch, to predict the next token given the previous context. We show that for GPT-like architectures, our work achieves faster convergence and significantly increases performance by adding a minuscule number of extra parameters when trained for the same epochs. We hope this work paves the way for algorithms exploring signal processing inside the signals found in neural architectures like LLMs and beyond.
机器翻译,仅供参考
