今日论文合集:cs.SD语音14篇,eess.AS音频处理15篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 MuDiT & MuSiT: Alignment with Colloquial Expression in Description-to-Song Generation
标题: MuDiT和MuSiT:在描述到歌曲生成中与口语表达的一致
作者:Zihao Wang,Haoxuan Liu,Jiaxing Yu,Tao Zhang,Yan Liu,Kejun Zhang
备注:19 pages, 5 figures
链接:点击下载PDF文件
摘要:在生成人工智能和人类艺术过程的交叉中,这项研究探讨了以人为中心的自动歌曲创作中关键但较少探索的对齐领域。我们提出了一个新的任务,口语描述到歌曲生成,其重点是对齐生成的内容与口语的人类表达。这项任务旨在弥合人工智能模型中口语理解和听觉表达之间的差距,最终目标是创建准确满足人类听觉期望并在结构上符合音乐规范的歌曲。目前的数据集是有限的,由于其狭窄的描述范围,语义差距和不准确。为了克服这一领域的数据稀缺性,我们提出了蔡冲音乐数据集(CaiMD)。CaiMD由专业音乐家和业余爱好者手动注释,提供不同的视角和对口语描述的全面理解。与现有的数据集不同,这些数据集预先设置了专家注释或具有固有偏见的自动生成的数据集,CaiMD更充分地满足了我们将人工智能生成的音乐与广泛的用户期望结果相匹配的目的。此外,我们提出了一个创新的单阶段框架,称为MuDiT MuSiT,使有效的人机对齐歌曲创作。该框架不仅实现了口语和听觉音乐感知之间的跨模态理解,而且还确保生成的歌曲与用户期望的结果一致。MuDiT MuSiT采用一个DiT SiT模型,用于端到端生成音乐组件,如旋律,和声,节奏,人声和乐器。该方法确保了所有生成的音乐成分之间的和谐声音凝聚力,促进了与人类听觉期望的更好共鸣。摘要:Amid the rising intersection of generative AI and human artistic processes, this study probes the critical yet less-explored terrain of alignment in human-centric automatic song composition. We propose a novel task of Colloquial Description-to-Song Generation, which focuses on aligning the generated content with colloquial human expressions. This task is aimed at bridging the gap between colloquial language understanding and auditory expression within an AI model, with the ultimate goal of creating songs that accurately satisfy human auditory expectations and structurally align with musical norms. Current datasets are limited due to their narrow descriptive scope, semantic gaps and inaccuracies. To overcome data scarcity in this domain, we present the Caichong Music Dataset (CaiMD). CaiMD is manually annotated by both professional musicians and amateurs, offering diverse perspectives and a comprehensive understanding of colloquial descriptions. Unlike existing datasets pre-set with expert annotations or auto-generated ones with inherent biases, CaiMD caters more sufficiently to our purpose of aligning AI-generated music with widespread user-desired results. Moreover, we propose an innovative single-stage framework called MuDiT MuSiT for enabling effective human-machine alignment in song creation. This framework not only achieves cross-modal comprehension between colloquial language and auditory music perceptions but also ensures generated songs align with user-desired results. MuDiT MuSiT employs one DiT SiT model for end-to-end generation of musical components like melody, harmony, rhythm, vocals, and instrumentation. The approach ensures harmonious sonic cohesiveness amongst all generated musical components, facilitating better resonance with human auditory expectations.

【2】 Investigating Decoder-only Large Language Models for Speech-to-text Translation
标题: 研究语音到文本翻译的纯解码器大型语言模型
作者:Chao-Wei Huang,Hui Lu,Hongyu Gong,Hirofumi Inaguma,Ilia Kulikov,Ruslan Mavlyutov,Sravya Popuri
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:大型语言模型(LLM)以其卓越的推理能力,泛化能力和跨不同领域的流畅性而闻名,为增强语音相关任务提供了一条有前途的途径。在本文中,我们专注于将仅解码器LLM集成到语音到文本翻译(S2TT)的任务中。我们提出了一个解码器的架构,使LLM直接消费的编码语音表示和生成的文本翻译。此外,我们研究了不同的参数有效的微调技术和任务制定的影响。我们的模型在CoVoST 2和FLEURS上实现了最先进的性能,在没有专有数据的情况下训练模型。我们还进行了分析,以验证我们提出的模型的设计选择,并为LLM与S2TT的集成带来见解。摘要:Large language models (LLMs), known for their exceptional reasoning capabilities, generalizability, and fluency across diverse domains, present a promising avenue for enhancing speech-related tasks. In this paper, we focus on integrating decoder-only LLMs to the task of speech-to-text translation (S2TT). We propose a decoder-only architecture that enables the LLM to directly consume the encoded speech representation and generate the text translation. Additionally, we investigate the effects of different parameter-efficient fine-tuning techniques and task formulation. Our model achieves state-of-the-art performance on CoVoST 2 and FLEURS among models trained without proprietary data. We also conduct analyses to validate the design choices of our proposed model and bring insights to the integration of LLMs to S2TT.

【3】 GMM-ResNext: Combining Generative and Discriminative Models for Speaker Verification
标题: GMM-ResNext:结合生成模型和区分模型进行说话人验证
作者:Hui Yan,Zhenchun Lei,Changhong Liu,Yong Zhou
链接:点击下载PDF文件
摘要:随着深度学习技术的发展,人们在说话人确认中探索了许多不同的网络架构。然而,大多数网络架构依赖于单一的深度学习架构,而在ASV任务中,结合不同架构的混合网络研究很少。在本文中,我们提出了GMM-ResNext模型的说话人确认。传统的高斯混合模型没有考虑每个帧特征在所有高斯分量上的分数分布,忽略了相邻语音帧之间的关系。因此,我们在原始声学特征的基础上提取对数高斯概率特征,并使用基于ResNext的网络作为骨干来提取说话人嵌入。GMM-ResNext结合了生成模型和判别模型,以提高深度学习模型的泛化能力,并允许人们更容易地指定模型参数的有意义的先验。提出了一种基于两个性别相关的GMM-ResNext模型。实验结果表明,在VoxCeleb 1-O测试集上,与ResNet 34和ECAPA-TDNN相比,GMM-ResNext的EER分别提高了48.1%和11.3%。摘要:With the development of deep learning, many different network architectures have been explored in speaker verification. However, most network architectures rely on a single deep learning architecture, and hybrid networks combining different architectures have been little studied in ASV tasks. In this paper, we propose the GMM-ResNext model for speaker verification. Conventional GMM does not consider the score distribution of each frame feature over all Gaussian components and ignores the relationship between neighboring speech frames. So, we extract the log Gaussian probability features based on the raw acoustic features and use ResNext-based network as the backbone to extract the speaker embedding. GMM-ResNext combines Generative and Discriminative Models to improve the generalization ability of deep learning models and allows one to more easily specify meaningful priors on model parameters. A two-path GMM-ResNext model based on two gender-related GMMs has also been proposed. The Experimental results show that the proposed GMM-ResNext achieves relative improvements of 48.1 % and 11.3 % in EER compared with ResNet34 and ECAPA-TDNN on VoxCeleb1-O test set.

【4】 Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
标题: 言语中的发音运动和音素对齐的与说话者和文本无关的估计
作者:Tobias Weise,Philipp Klumpp,Kubilay Can Demir,Paula Andrea Pérez-Toro,Maria Schuster,Elmar Noeth,Bjoern Heismann,Andreas Maier,Seung Hee Yang
备注:to be published in Interspeech 2024 proceedings
链接:点击下载PDF文件
摘要:本文介绍了之前单独处理的两项任务的新颖组合:声学到发音语音倒置(AAI)和音素到发音(PTA)运动估计。我们将这种联合任务称为声学音素到发音语音倒置(APTAI),并探索两种不同的方法,在推理过程中独立于说话者和文本工作。我们使用多任务学习设置,端到端的目标是将原始语音作为输入并估计相应的发音运动、音素序列和音素对齐。虽然这两种提出的方法都有相同的要求,但它们实现音素相关预测的方式不同:一种是基于帧分类,另一种是基于两阶段训练过程和强制对齐。我们达到0.73平均相关的AAI任务的竞争力表现,并实现高达约87%的帧重叠相比,一个国家的最先进的文本相关的音素力对齐。摘要:This paper introduces a novel combination of two tasks, previously treated separately: acoustic-to-articulatory speech inversion (AAI) and phoneme-to-articulatory (PTA) motion estimation. We refer to this joint task as acoustic phoneme-to-articulatory speech inversion (APTAI) and explore two different approaches, both working speaker- and text-independently during inference. We use a multi-task learning setup, with the end-to-end goal of taking raw speech as input and estimating the corresponding articulatory movements, phoneme sequence, and phoneme alignment. While both proposed approaches share these same requirements, they differ in their way of achieving phoneme-related predictions: one is based on frame classification, the other on a two-staged training procedure and forced alignment. We reach competitive performance of 0.73 mean correlation for the AAI task and achieve up to approximately 87% frame overlap compared to a state-of-the-art text-dependent phoneme force aligner.

【5】 A Toolchain for Comprehensive AudioVideo Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)
标题: 使用基于深度学习的多模式方法进行全面音频视频分析的工具链(骚乱或暴力上下文检测的用例)
作者:Lam Pham,Phat Lam,Tin Nguyen,Hieu Tang,Alexander Schindler
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个工具链,通过利用基于深度学习的多模态方法进行全面的音频 视频分析。为此,语音到文本(S2 T),声学场景分类(ASC),声学事件检测(AED),视觉对象检测(VOD),图像字幕(IC)和视频字幕(VC)的不同特定任务被执行并集成到工具链中。通过组合各个任务并分析从输入视频中提取的音频和视频数据,工具链提供了各种基于音频 视频的应用程序:音频 视频聚类的两个一般应用程序,综合音频 视频摘要和骚乱或暴力上下文检测的特定应用程序。此外,该工具链提供了一个灵活且适应性强的架构,可以有效地为进一步的基于音频 视频的应用程序集成新模型。摘要:In this paper, we present a toolchain for a comprehensive audio video analysis by leveraging deep learning based multimodal approach. To this end, different specific tasks of Speech to Text (S2T), Acoustic Scene Classification (ASC), Acoustic Event Detection (AED), Visual Object Detection (VOD), Image Captioning (IC), and Video Captioning (VC) are conducted and integrated into the toolchain. By combining individual tasks and analyzing both audio & visual data extracted from input video, the toolchain offers various audio video-based applications: Two general applications of audio video clustering, comprehensive audio video summary and a specific application of riot or violent context detection. Furthermore, the toolchain presents a flexible and adaptable architecture that is effective to integrate new models for further audio video-based applications.

【6】 Qifusion-Net: Layer-adapted StreamNon-stream Model for End-to-End Multi-Accent Speech Recognition
标题: Qifusion-Net:用于端到端多口音语音识别的分层自适应流非流模型
作者:Jinming Chen,Jingyi Fang,Yuanzhong Zheng,Yaoxuan Wang,Haojun Fei
备注:accpeted by interspeech 2014, 5 pages, 1 figure
链接:点击下载PDF文件
摘要:目前,端到端(E2 E)语音识别方法已经取得了令人满意的性能。然而,自动语音识别(ASR)模型在准确识别多口音语音方面仍然面临挑战。我们提出了一个层自适应融合(LAF)模型,称为Qifusion-Net,它不需要任何先验知识的目标口音。该方法基于动态组块策略,实现了流解码,并能提取帧级声学特征,有利于细粒度的信息融合。实验结果表明,我们提出的方法优于基线的字符错误率(CER)的相对减少22.1$ %$和17.2$ %$在KeSpeech和MagicData-RMAC多口音测试数据集。摘要:Currently, end-to-end (E2E) speech recognition methods have achieved promising performance. However, auto speech recognition (ASR) models still face challenges in recognizing multi-accent speech accurately. We propose a layer-adapted fusion (LAF) model, called Qifusion-Net, which does not require any prior knowledge about the target accent. Based on dynamic chunk strategy, our approach enables streaming decoding and can extract frame-level acoustic feature, facilitating fine-grained information fusion. Experiment results demonstrate that our proposed methods outperform the baseline with relative reductions of 22.1$ %$ and 17.2$ %$ in character error rate (CER) across multi accent test datasets on KeSpeech and MagicData-RMAC.

【7】 Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
标题: 神经语音模型中的类人语言偏见:Wav2Vec2.0中的语音分类和音素约束
作者:Marianne de Heer Kloots,Willem Zuidema
Journal-ref:Proc. INTERSPEECH 2024
链接:点击下载PDF文件
摘要:深度神经语音模型对语音学了解多少?现有的工作已经研究了编码的个别语言单位,如音素在这些模型。在这里,我们研究单位之间的相互作用。受人类语音感知的经典实验的启发,我们研究了Wav 2 Vec 2如何解决语音定位约束。我们在 l 和 r 之间的声学连续体上合成声音,并将它们嵌入到英语中只有 l ,只有 r 或两者都没有出现的受控上下文中。像人类一样,Wav 2 Vec 2模型在处理这种模棱两可的声音时也表现出对语音可接受类别的偏见。使用简单的措施来分析模型内部的个人刺激的水平上,我们发现,这种偏见出现在模型的Transformer模块的早期层。这种效果被ASR微调放大,但也存在于完全自我监督的模型中。我们的方法演示了如何控制刺激设计可以帮助本地化特定的语言知识在神经语音模型。摘要:What do deep neural speech models know about phonology? Existing work has examined the encoding of individual linguistic units such as phonemes in these models. Here we investigate interactions between units. Inspired by classic experiments on human speech perception, we study how Wav2Vec2 resolves phonotactic constraints. We synthesize sounds on an acoustic continuum between l and r and embed them in controlled contexts where only l , only r , or neither occur in English. Like humans, Wav2Vec2 models show a bias towards the phonotactically admissable category in processing such ambiguous sounds. Using simple measures to analyze model internals on the level of individual stimuli, we find that this bias emerges in early layers of the model's Transformer module. This effect is amplified by ASR finetuning but also present in fully self-supervised models. Our approach demonstrates how controlled stimulus designs can help localize specific linguistic knowledge in neural speech models.

【8】 Probing the Feasibility of Multilingual Speaker Anonymization
标题: 探讨多语言说话人语音化的可行性
作者:Sarina Meyer,Florian Lux,Ngoc Thang Vu
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:在说话者匿名化中,语音记录以说话者的身份保持隐藏的方式进行修改。虽然这项技术可以帮助保护全球个人的隐私,但目前的研究几乎只关注英语数据,从而限制了这一点。在这项研究中,我们将一个最先进的匿名系统扩展到九种语言,将依赖语言的组件转换为多语言组件。实验测试的匿名语音对隐私攻击和语音恶化的鲁棒性表明,该系统的所有语言的整体成功。结果表明,在英语数据上训练的说话人嵌入可以应用于各种语言,并且一种语言的匿名化性能主要受其使用的语音合成组件的质量影响。摘要:In speaker anonymization, speech recordings are modified in a way that the identity of the speaker remains hidden. While this technology could help to protect the privacy of individuals around the globe, current research restricts this by focusing almost exclusively on English data. In this study, we extend a state-of-the-art anonymization system to nine languages by transforming language-dependent components to their multilingual counterparts. Experiments testing the robustness of the anonymized speech against privacy attacks and speech deterioration show an overall success of this system for all languages. The results suggest that speaker embeddings trained on English data can be applied across languages, and that the anonymization performance for a language is mainly affected by the quality of the speech synthesis component used for it.

【9】 PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
标题: PicoAudio:在文本转音频生成中实现音频事件的精确时间戳和频率可控性
作者:Zeyu Xie,Xuenan Xu,Zhizheng Wu,Mengyue Wu
链接:点击下载PDF文件
摘要:最近,音频生成任务吸引了相当大的研究兴趣。精确的时间可控性对于将音频生成与实际应用集成至关重要。在这项工作中,我们提出了一个时间控制的音频生成框架,PicoAudio。PicoAudio集成了时间信息,通过量身定制的模型设计来指导音频生成。它利用数据抓取,分割,过滤和模拟细粒度的时间对齐的音频文本数据。主观和客观评估都表明,PicoAudio在时间戳和发生频率可控性方面大大超过了当前最先进的生成模型。生成的示例可在演示网站https: PicoAudio.github.io上获得。摘要:Recently, audio generation tasks have attracted considerable research interests. Precise temporal controllability is essential to integrate audio generation with real applications. In this work, we propose a temporal controlled audio generation framework, PicoAudio. PicoAudio integrates temporal information to guide audio generation through tailored model design. It leverages data crawling, segmentation, filtering, and simulation of fine-grained temporally-aligned audio-text data. Both subjective and objective evaluations demonstrate that PicoAudio dramantically surpasses current state-of-the-art generation models in terms of timestamp and occurrence frequency controllability. The generated samples are available on the demo website https: PicoAudio.github.io.

【10】 AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
标题: AudioTime:时间对齐的音频文本基准数据集
作者:Zeyu Xie,Xuenan Xu,Zhizheng Wu,Mengyue Wu
链接:点击下载PDF文件
摘要:音频生成的最新进展已经使得能够从自由形式的文本描述创建高保真音频剪辑。然而,时间关系,音频内容的一个重要特征,目前在主流模型中代表性不足,导致不精确的时间可控性。具体来说,用户无法使用自由格式的文本精确地控制声音事件的时间戳。我们承认,一个重要的因素是缺乏高质量的,时间对齐的音频文本数据集,这对于训练具有时间控制的模型至关重要。注释的时间对齐越多,模型就越能理解音频输出和时间文本提示之间的精确关系。因此,我们提出了一个强对齐的音频文本数据集AudioTime。它提供了丰富的时间信息,如时间戳,持续时间,频率和顺序的文本注释,几乎涵盖了时间控制的所有方面。此外,我们提供了一个全面的测试集和评估指标,以评估各种模型的时间控制性能。示例可在https: zeyuxie29.github.io AudioTime 上获得摘要:Recent advancements in audio generation have enabled the creation of high-fidelity audio clips from free-form textual descriptions. However, temporal relationships, a critical feature for audio content, are currently underrepresented in mainstream models, resulting in an imprecise temporal controllability. Specifically, users cannot accurately control the timestamps of sound events using free-form text. We acknowledge that a significant factor is the absence of high-quality, temporally-aligned audio-text datasets, which are essential for training models with temporal control. The more temporally-aligned the annotations, the better the models can understand the precise relationship between audio outputs and temporal textual prompts. Therefore, we present a strongly aligned audio-text dataset, AudioTime. It provides text annotations rich in temporal information such as timestamps, duration, frequency, and ordering, covering almost all aspects of temporal control. Additionally, we offer a comprehensive test set and evaluation metric to assess the temporal control performance of various models. Examples are available on the https: zeyuxie29.github.io AudioTime

【11】 Nollywood: Let's Go to the Movies!
标题: 尼莱坞:我们去看电影吧!
作者:John E. Ortega,Ibrahim Said Ahmad,William Chen
备注:8 pages, 4 figures, 2 tables
链接:点击下载PDF文件
摘要:尼莱坞,基于印度宝莱坞的想法,是一系列源于尼日利亚的优秀电影。不幸的是,虽然电影是英语的,但由于英语方言,许多母语人士很难理解。在这篇文章中,我们实现了两个目标:(1)创建一个语音字幕模型,能够将尼日利亚英语语音翻译为美国英语;(2)使用最先进的毒性检测器来发现语音的毒性。我们的目的是突出这些视频中的文字,这些文字往往因为缺乏方言理解而被忽视,因为尼日利亚的许多人在家里说豪萨语等母语。摘要:Nollywood, based on the idea of Bollywood from India, is a series of outstanding movies that originate from Nigeria. Unfortunately, while the movies are in English, they are hard to understand for many native speakers due to the dialect of English that is spoken. In this article, we accomplish two goals: (1) create a phonetic sub-title model that is able to translate Nigerian English speech to American English and (2) use the most advanced toxicity detectors to discover how toxic the speech is. Our aim is to highlight the text in these videos which is often times ignored for lack of dialectal understanding due the fact that many people in Nigeria speak a native language like Hausa at home.

【12】 Towards the Next Frontier in Speech Representation Learning Using Disentanglement
标题: 使用解纠缠迈向语音表示学习的下一个前沿
作者:Varun Krishna,Sriram Ganapathy
链接:点击下载PDF文件
摘要:语音表示的自监督学习的流行框架主要集中在语音区域的帧级掩蔽预测上。虽然这已经显示出语音识别和相关任务的有希望的下游任务性能,但这在很大程度上忽略了在较粗糙级别编码的语音因素,例如在整个语音话语中保持一致的扬声器或通道的特性。在这项工作中,我们提出了一个框架,学习解开自我监督(称为Learn2Diss)表示的语音,其中包括帧级和话语级编码器模块。这两个编码器最初是独立学习的,其中帧级模型在很大程度上受到现有自我监督技术的启发,从而学习伪音素表示,而话语级编码器受到池化嵌入的启发,从而学习伪说话者表示。这两个模块的联合学习包括使用基于互信息的标准解开两个编码器。通过几个下游评估实验,我们证明了所提出的Learn2Diss在各种任务上都取得了最先进的结果,其中帧级编码器表示改善了语义任务,而话语级表示改善了非语义任务。摘要:The popular frameworks for self-supervised learning of speech representations have largely focused on frame-level masked prediction of speech regions. While this has shown promising downstream task performance for speech recognition and related tasks, this has largely ignored factors of speech that are encoded at coarser level, like characteristics of the speaker or channel that remain consistent through-out a speech utterance. In this work, we propose a framework for Learning Disentangled Self Supervised (termed as Learn2Diss) representations of speech, which consists of frame-level and an utterance-level encoder modules. The two encoders are initially learned independently, where the frame-level model is largely inspired by existing self supervision techniques, thereby learning pseudo-phonemic representations, while the utterance-level encoder is inspired by constrastive learning of pooled embeddings, thereby learning pseudo-speaker representations. The joint learning of these two modules consists of disentangling the two encoders using a mutual information based criterion. With several downstream evaluation experiments, we show that the proposed Learn2Diss achieves state-of-the-art results on a variety of tasks, with the frame-level encoder representations improving semantic tasks, while the utterance-level representations improve non-semantic tasks.

【13】 VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features
标题: 使用梯度软化和SSL声学特征的基于VAE的音素对齐
作者:Tomoki Koriyama
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:本文提出了一种精确的音素对齐模型,旨在语音分析和视频内容创建。我们提出了一个变分自动编码器(VAE)为基础的对齐模型,其中一个可能的路径搜索使用编码的声学和语言嵌入在一个无监督的方式。我们提出的模型是基于一个TTS对齐(OTA)和扩展,以获得音素边界。具体来说,我们采用VAE架构来保持嵌入和输入之间的一致性,应用梯度退火来避免训练过程中的局部最优,并引入基于自监督学习(SSL)的声学特征输入和状态级语言单元来利用丰富而详细的信息。实验结果表明,与传统的OTA模型、基于CTC的分割模型和广泛使用的MFA工具相比,该模型生成的音素边界更接近标注的音素边界。摘要:This paper presents an accurate phoneme alignment model that aims for speech analysis and video content creation. We propose a variational autoencoder (VAE)-based alignment model in which a probable path is searched using encoded acoustic and linguistic embeddings in an unsupervised manner. Our proposed model is based on one TTS alignment (OTA) and extended to obtain phoneme boundaries. Specifically, we incorporate a VAE architecture to maintain consistency between the embedding and input, apply gradient annealing to avoid local optimum during training, and introduce a self-supervised learning (SSL)-based acoustic-feature input and state-level linguistic unit to utilize rich and detailed information. Experimental results show that the proposed model generated phoneme boundaries closer to annotated ones compared with the conventional OTA model, the CTC-based segmentation model, and the widely-used tool MFA.

【14】 Livestock feeding behaviour: A review on automated systems for ruminant monitoring
标题: 牲畜进食行为:反刍动物监测自动化系统回顾
作者:José Chelotti,Luciano Martinez-Rau,Mariano Ferrero,Leandro Vignolo,Julio Galli,Alejandra Planisich,H. Leonardo Rufiner,Leonardo Giovanini
备注:Preprint submitted to the journal biosystems engineering
链接:点击下载PDF文件
摘要:家畜饲养行为是畜牧业和农业领域的一个有影响力的研究领域。近年来,人们对用于监测反刍动物行为的自动化系统越来越感兴趣。尽管在过去十年中取得了进展,但在测量和分析牲畜饲养行为的方法方面仍有许多工作要做。自动监控系统主要使用运动、声学和图像传感器来收集动物行为数据。现有方法的性能评价是一项复杂的任务,研究之间的直接比较是困难的。从实验中使用的数据和性能指标的多样性开始,有几个因素阻止了直接比较。据我们所知,这项工作代表了第一次辅导式的反刍动物的摄食行为的分析,强调传感方法,信号处理和计算智能方法之间的关系。它评估了主要的传感方法(即基于运动,声音,图像 视频和压力)以及测量和分析与进食行为相关的信号的主要技术,评估了它们在不同环境和情况下的使用。它还强调了自动监测系统的潜力,以提供有价值的信息,提高我们对牲畜饲养行为的理解。由于这些系统对生产系统和研究的影响,它们的相关性越来越重要。最后,本文最后讨论了未来的挑战和机遇,在牲畜饲养行为监测。摘要:Livestock feeding behaviour is an influential research area for those involved in animal husbandry and agriculture. In recent years, there has been a growing interest in automated systems for monitoring the behaviour of ruminants. Despite the developments accomplished in the last decade, there is still much to do and learn about the methods for measuring and analysing livestock feeding behaviour. Automated monitoring systems mainly use motion, acoustic, and image sensors to collect animal behavioural data. The performance evaluation of existing methods is a complex task and direct comparisons between studies are difficult. Several factors prevent a direct comparison, starting from the diversity of data and performance metrics used in the experiments. To the best of our knowledge, this work represents the first tutorial-style review on the analysis of the feeding behaviour of ruminants, emphasising the relationship between sensing methodologies, signal processing, and computational intelligence methods. It assesses the main sensing methodologies (i.e. based on movement, sound, images videos, and pressure) and the main techniques to measure and analyse the signals associated with feeding behaviour, evaluating their use in different settings and situations. It also highlights the potentiality of automated monitoring systems to provide valuable information that improves our understanding of livestock feeding behaviour. The relevance of these systems is increasingly important due to their impact on production systems and research. Finally, the paper closes by discussing future challenges and opportunities in livestock feeding behaviour monitoring.


eess.AS音频处理
【1】 SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
标题: SA-WavLM:混合语音的说话者感知自我监督预训练
作者:Jingru Lin,Meng Ge,Junyi Ao,Liqun Deng,Haizhou Li
备注:InterSpeech 2024
链接:点击下载PDF文件
摘要:结果表明,具有自监督学习(SSL)技术的预训练模型在各种下游语音任务中是有效的。然而,大多数这样的模型都是在单说话人语音数据上训练的,这限制了它们在混合语音中的有效性。这促使我们探索混合语音的预训练。本文提出了一种新的混合语音预训练模型SA-WavLM。具体来说,SA-WavLM遵循“提取-合并-预测”流水线,其中输入混合中每个扬声器的表示首先单独提取,然后在最终预测之前合并。在这个流水线中,SA-WavLM执行说话者通知提取,考虑不同说话者之间的交互。此外,提出了一种说话人重排策略,以提高对说话人缺失的鲁棒性。实验表明,SA-WavLM匹配或改进了最先进的预训练模型。摘要:It was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks. However, most such models are trained on single-speaker speech data, limiting their effectiveness in mixture speech. This motivates us to explore pre-training on mixture speech. This work presents SA-WavLM, a novel pre-trained model for mixture speech. Specifically, SA-WavLM follows an "extract-merge-predict" pipeline in which the representations of each speaker in the input mixture are first extracted individually and then merged before the final prediction. In this pipeline, SA-WavLM performs speaker-informed extractions with the consideration of the interactions between different speakers. Furthermore, a speaker shuffling strategy is proposed to enhance the robustness towards the speaker absence. Experiments show that SA-WavLM either matches or improves upon the state-of-the-art pre-trained models.

【2】 VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features
标题: 使用梯度软化和SSL声学特征的基于VAE的音素对齐
作者:Tomoki Koriyama
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:本文提出了一种精确的音素对齐模型,旨在语音分析和视频内容创建。我们提出了一个变分自动编码器(VAE)为基础的对齐模型,其中一个可能的路径搜索使用编码的声学和语言嵌入在一个无监督的方式。我们提出的模型是基于一个TTS对齐(OTA)和扩展,以获得音素边界。具体来说,我们采用VAE架构来保持嵌入和输入之间的一致性,应用梯度退火来避免训练过程中的局部最优,并引入基于自监督学习(SSL)的声学特征输入和状态级语言单元来利用丰富而详细的信息。实验结果表明,与传统的OTA模型、基于CTC的分割模型和广泛使用的MFA工具相比,该模型生成的音素边界更接近标注的音素边界。摘要:This paper presents an accurate phoneme alignment model that aims for speech analysis and video content creation. We propose a variational autoencoder (VAE)-based alignment model in which a probable path is searched using encoded acoustic and linguistic embeddings in an unsupervised manner. Our proposed model is based on one TTS alignment (OTA) and extended to obtain phoneme boundaries. Specifically, we incorporate a VAE architecture to maintain consistency between the embedding and input, apply gradient annealing to avoid local optimum during training, and introduce a self-supervised learning (SSL)-based acoustic-feature input and state-level linguistic unit to utilize rich and detailed information. Experimental results show that the proposed model generated phoneme boundaries closer to annotated ones compared with the conventional OTA model, the CTC-based segmentation model, and the widely-used tool MFA.

【3】 Zero-Bit Transmission of Adaptive Pre- and De-emphasis Filters for Speech and Audio Coding
标题: 用于语音和音频编码的自适应预加重和去加重过滤器的零比特传输
作者:Niloofar Omidi Piralideh,Philippe Gournay,Roch Lefebvre
备注:This paper has been accepted by the 47th International Conference on Telecommunications and Signal Processing (TSP 2024)
链接:点击下载PDF文件
摘要:本文介绍了一种新的自适应一阶预加重和去加重滤波器,在许多语音和音频编解码器,以提高编码效率和感知质量的重要工具的方法。所提出的零比特自适应方法与经典的前向和后向自适应方法的不同之处在于,在接收机处从解码的预加重信号估计去加重系数。这消除了对传输由前向自适应产生的信息以及后向自适应中固有的信号滤波器滞后的需要。评估结果表明,去加重系数可以准确地估计从解码的预强调信号和建议的零比特自适应方法提供了可比的主观改善前向适应。摘要:This paper introduces a novel adaptation approach for first-order pre- and de-emphasis filters, an essential tool in many speech and audio codecs to increase coding efficiency and perceived quality. The proposed zero-bit self-adaptation approach differs from classical forward and backward adaptation approaches in that the de-emphasis coefficient is estimated at the receiver, from the decoded pre-emphasized signal. This eliminates the need to transmit information that arises from forward adaptation as well as the signal-filter lag that is inherent in backward adaptation. Evaluation results show that the de-emphasis coefficient can be estimated accurately from the decoded pre-emphasized signal and that the proposed zero-bit self-adaptation approach provides comparable subjective improvement to forward adaptation.

【4】 MuDiT & MuSiT: Alignment with Colloquial Expression in Description-to-Song Generation
标题: MuDiT和MuSiT:在描述到歌曲生成中与口语表达的一致
作者:Zihao Wang,Haoxuan Liu,Jiaxing Yu,Tao Zhang,Yan Liu,Kejun Zhang
备注:19 pages, 5 figures
链接:点击下载PDF文件
摘要:在生成人工智能和人类艺术过程的交叉中,这项研究探讨了以人为中心的自动歌曲创作中关键但较少探索的对齐领域。我们提出了一个新的任务,口语描述到歌曲生成,其重点是对齐生成的内容与口语的人类表达。这项任务旨在弥合人工智能模型中口语理解和听觉表达之间的差距,最终目标是创建准确满足人类听觉期望并在结构上符合音乐规范的歌曲。目前的数据集是有限的,由于其狭窄的描述范围,语义差距和不准确。为了克服这一领域的数据稀缺性,我们提出了蔡冲音乐数据集(CaiMD)。CaiMD由专业音乐家和业余爱好者手动注释,提供不同的视角和对口语描述的全面理解。与现有的数据集不同,这些数据集预先设置了专家注释或具有固有偏见的自动生成的数据集,CaiMD更充分地满足了我们将人工智能生成的音乐与广泛的用户期望结果相匹配的目的。此外,我们提出了一个创新的单阶段框架,称为MuDiT MuSiT,使有效的人机对齐歌曲创作。该框架不仅实现了口语和听觉音乐感知之间的跨模态理解,而且还确保生成的歌曲与用户期望的结果一致。MuDiT MuSiT采用一个DiT SiT模型,用于端到端生成音乐组件,如旋律,和声,节奏,人声和乐器。该方法确保了所有生成的音乐成分之间的和谐声音凝聚力,促进了与人类听觉期望的更好共鸣。摘要:Amid the rising intersection of generative AI and human artistic processes, this study probes the critical yet less-explored terrain of alignment in human-centric automatic song composition. We propose a novel task of Colloquial Description-to-Song Generation, which focuses on aligning the generated content with colloquial human expressions. This task is aimed at bridging the gap between colloquial language understanding and auditory expression within an AI model, with the ultimate goal of creating songs that accurately satisfy human auditory expectations and structurally align with musical norms. Current datasets are limited due to their narrow descriptive scope, semantic gaps and inaccuracies. To overcome data scarcity in this domain, we present the Caichong Music Dataset (CaiMD). CaiMD is manually annotated by both professional musicians and amateurs, offering diverse perspectives and a comprehensive understanding of colloquial descriptions. Unlike existing datasets pre-set with expert annotations or auto-generated ones with inherent biases, CaiMD caters more sufficiently to our purpose of aligning AI-generated music with widespread user-desired results. Moreover, we propose an innovative single-stage framework called MuDiT MuSiT for enabling effective human-machine alignment in song creation. This framework not only achieves cross-modal comprehension between colloquial language and auditory music perceptions but also ensures generated songs align with user-desired results. MuDiT MuSiT employs one DiT SiT model for end-to-end generation of musical components like melody, harmony, rhythm, vocals, and instrumentation. The approach ensures harmonious sonic cohesiveness amongst all generated musical components, facilitating better resonance with human auditory expectations.

【5】 Investigating Decoder-only Large Language Models for Speech-to-text Translation
标题: 研究语音到文本翻译的纯解码器大型语言模型
作者:Chao-Wei Huang,Hui Lu,Hongyu Gong,Hirofumi Inaguma,Ilia Kulikov,Ruslan Mavlyutov,Sravya Popuri
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:大型语言模型(LLM)以其卓越的推理能力,泛化能力和跨不同领域的流畅性而闻名,为增强语音相关任务提供了一条有前途的途径。在本文中,我们专注于将仅解码器LLM集成到语音到文本翻译(S2TT)的任务中。我们提出了一个解码器的架构,使LLM直接消费的编码语音表示和生成的文本翻译。此外,我们研究了不同的参数有效的微调技术和任务制定的影响。我们的模型在CoVoST 2和FLEURS上实现了最先进的性能,在没有专有数据的情况下训练模型。我们还进行了分析,以验证我们提出的模型的设计选择,并为LLM与S2TT的集成带来见解。摘要:Large language models (LLMs), known for their exceptional reasoning capabilities, generalizability, and fluency across diverse domains, present a promising avenue for enhancing speech-related tasks. In this paper, we focus on integrating decoder-only LLMs to the task of speech-to-text translation (S2TT). We propose a decoder-only architecture that enables the LLM to directly consume the encoded speech representation and generate the text translation. Additionally, we investigate the effects of different parameter-efficient fine-tuning techniques and task formulation. Our model achieves state-of-the-art performance on CoVoST 2 and FLEURS among models trained without proprietary data. We also conduct analyses to validate the design choices of our proposed model and bring insights to the integration of LLMs to S2TT.

【6】 GMM-ResNext: Combining Generative and Discriminative Models for Speaker Verification
标题: GMM-ResNext:结合生成模型和区分模型进行说话人验证
作者:Hui Yan,Zhenchun Lei,Changhong Liu,Yong Zhou
链接:点击下载PDF文件
摘要:随着深度学习技术的发展,人们在说话人确认中探索了许多不同的网络架构。然而,大多数网络架构依赖于单一的深度学习架构,而在ASV任务中,结合不同架构的混合网络研究很少。在本文中,我们提出了GMM-ResNext模型的说话人确认。传统的高斯混合模型没有考虑每个帧特征在所有高斯分量上的分数分布,忽略了相邻语音帧之间的关系。因此,我们在原始声学特征的基础上提取对数高斯概率特征,并使用基于ResNext的网络作为骨干来提取说话人嵌入。GMM-ResNext结合了生成模型和判别模型,以提高深度学习模型的泛化能力,并允许人们更容易地指定模型参数的有意义的先验。提出了一种基于两个性别相关的GMM-ResNext模型。实验结果表明,在VoxCeleb 1-O测试集上,与ResNet 34和ECAPA-TDNN相比,GMM-ResNext的EER分别提高了48.1%和11.3%。摘要:With the development of deep learning, many different network architectures have been explored in speaker verification. However, most network architectures rely on a single deep learning architecture, and hybrid networks combining different architectures have been little studied in ASV tasks. In this paper, we propose the GMM-ResNext model for speaker verification. Conventional GMM does not consider the score distribution of each frame feature over all Gaussian components and ignores the relationship between neighboring speech frames. So, we extract the log Gaussian probability features based on the raw acoustic features and use ResNext-based network as the backbone to extract the speaker embedding. GMM-ResNext combines Generative and Discriminative Models to improve the generalization ability of deep learning models and allows one to more easily specify meaningful priors on model parameters. A two-path GMM-ResNext model based on two gender-related GMMs has also been proposed. The Experimental results show that the proposed GMM-ResNext achieves relative improvements of 48.1 % and 11.3 % in EER compared with ResNet34 and ECAPA-TDNN on VoxCeleb1-O test set.

【7】 Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
标题: 言语中的发音运动和音素对齐的与说话者和文本无关的估计
作者:Tobias Weise,Philipp Klumpp,Kubilay Can Demir,Paula Andrea Pérez-Toro,Maria Schuster,Elmar Noeth,Bjoern Heismann,Andreas Maier,Seung Hee Yang
备注:to be published in Interspeech 2024 proceedings
链接:点击下载PDF文件
摘要:本文介绍了一种新的两个任务的组合,以前分别处理:声学发音语音反转(AAI)和音素发音(PTA)运动估计。我们把这个联合任务称为声学音素发音语音反转(APTAI),并探讨两种不同的方法,无论是工作扬声器和文本独立推理。我们使用多任务学习设置,以原始语音作为输入,并估计相应的发音运动,音素序列和音素对齐的端到端的目标。虽然这两种方法都有相同的要求,但它们在实现音素相关预测的方式上有所不同:一种是基于帧分类,另一种是基于两阶段训练过程和强制对齐。我们达到0.73平均相关的AAI任务的竞争力表现,并实现高达约87%的帧重叠相比,一个国家的最先进的文本相关的音素力对齐。摘要:This paper introduces a novel combination of two tasks, previously treated separately: acoustic-to-articulatory speech inversion (AAI) and phoneme-to-articulatory (PTA) motion estimation. We refer to this joint task as acoustic phoneme-to-articulatory speech inversion (APTAI) and explore two different approaches, both working speaker- and text-independently during inference. We use a multi-task learning setup, with the end-to-end goal of taking raw speech as input and estimating the corresponding articulatory movements, phoneme sequence, and phoneme alignment. While both proposed approaches share these same requirements, they differ in their way of achieving phoneme-related predictions: one is based on frame classification, the other on a two-staged training procedure and forced alignment. We reach competitive performance of 0.73 mean correlation for the AAI task and achieve up to approximately 87% frame overlap compared to a state-of-the-art text-dependent phoneme force aligner.

【8】 A Toolchain for Comprehensive AudioVideo Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)
标题: 使用基于深度学习的多模式方法进行全面音频视频分析的工具链(骚乱或暴力上下文检测的用例)
作者:Lam Pham,Phat Lam,Tin Nguyen,Hieu Tang,Alexander Schindler
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个工具链,通过利用基于深度学习的多模态方法进行全面的音频 视频分析。为此,语音到文本(S2 T),声学场景分类(ASC),声学事件检测(AED),视觉对象检测(VOD),图像字幕(IC)和视频字幕(VC)的不同特定任务被执行并集成到工具链中。通过组合各个任务并分析从输入视频中提取的音频和视频数据,工具链提供了各种基于音频 视频的应用程序:音频 视频聚类的两个一般应用程序,综合音频 视频摘要和骚乱或暴力上下文检测的特定应用程序。此外,该工具链提供了一个灵活且适应性强的架构,可以有效地为进一步的基于音频 视频的应用程序集成新模型。摘要:In this paper, we present a toolchain for a comprehensive audio video analysis by leveraging deep learning based multimodal approach. To this end, different specific tasks of Speech to Text (S2T), Acoustic Scene Classification (ASC), Acoustic Event Detection (AED), Visual Object Detection (VOD), Image Captioning (IC), and Video Captioning (VC) are conducted and integrated into the toolchain. By combining individual tasks and analyzing both audio & visual data extracted from input video, the toolchain offers various audio video-based applications: Two general applications of audio video clustering, comprehensive audio video summary and a specific application of riot or violent context detection. Furthermore, the toolchain presents a flexible and adaptable architecture that is effective to integrate new models for further audio video-based applications.

【9】 Qifusion-Net: Layer-adapted StreamNon-stream Model for End-to-End Multi-Accent Speech Recognition
标题: Qifusion-Net:用于端到端多口音语音识别的分层自适应流非流模型
作者:Jinming Chen,Jingyi Fang,Yuanzhong Zheng,Yaoxuan Wang,Haojun Fei
备注:accpeted by interspeech 2014, 5 pages, 1 figure
链接:点击下载PDF文件
摘要:目前,端到端(E2 E)语音识别方法已经取得了令人满意的性能。然而,自动语音识别(ASR)模型在准确识别多口音语音方面仍然面临挑战。我们提出了一个层自适应融合(LAF)模型,称为Qifusion-Net,它不需要任何先验知识的目标口音。该方法基于动态组块策略,实现了流解码,并能提取帧级声学特征,有利于细粒度的信息融合。实验结果表明,我们提出的方法优于基线的字符错误率(CER)的相对减少22.1$ %$和17.2$ %$在KeSpeech和MagicData-RMAC多口音测试数据集。摘要:Currently, end-to-end (E2E) speech recognition methods have achieved promising performance. However, auto speech recognition (ASR) models still face challenges in recognizing multi-accent speech accurately. We propose a layer-adapted fusion (LAF) model, called Qifusion-Net, which does not require any prior knowledge about the target accent. Based on dynamic chunk strategy, our approach enables streaming decoding and can extract frame-level acoustic feature, facilitating fine-grained information fusion. Experiment results demonstrate that our proposed methods outperform the baseline with relative reductions of 22.1$ %$ and 17.2$ %$ in character error rate (CER) across multi accent test datasets on KeSpeech and MagicData-RMAC.

【10】 Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
标题: 神经语音模型中的类人语言偏见:Wav2Vec2.0中的语音分类和音素约束
作者:Marianne de Heer Kloots,Willem Zuidema
Journal-ref:Proc. INTERSPEECH 2024
链接:点击下载PDF文件
摘要:深度神经语音模型对语音学了解多少?现有的工作已经研究了编码的个别语言单位,如音素在这些模型。在这里,我们研究单位之间的相互作用。受人类语音感知的经典实验的启发,我们研究了Wav 2 Vec 2如何解决语音定位约束。我们在 l 和 r 之间的声学连续体上合成声音,并将它们嵌入到英语中只有 l ,只有 r 或两者都没有出现的受控上下文中。像人类一样,Wav 2 Vec 2模型在处理这种模棱两可的声音时也表现出对语音可接受类别的偏见。使用简单的措施来分析模型内部的个人刺激的水平上,我们发现,这种偏见出现在模型的Transformer模块的早期层。这种效果被ASR微调放大,但也存在于完全自我监督的模型中。我们的方法演示了如何控制刺激设计可以帮助本地化特定的语言知识在神经语音模型。摘要:What do deep neural speech models know about phonology? Existing work has examined the encoding of individual linguistic units such as phonemes in these models. Here we investigate interactions between units. Inspired by classic experiments on human speech perception, we study how Wav2Vec2 resolves phonotactic constraints. We synthesize sounds on an acoustic continuum between l and r and embed them in controlled contexts where only l , only r , or neither occur in English. Like humans, Wav2Vec2 models show a bias towards the phonotactically admissable category in processing such ambiguous sounds. Using simple measures to analyze model internals on the level of individual stimuli, we find that this bias emerges in early layers of the model's Transformer module. This effect is amplified by ASR finetuning but also present in fully self-supervised models. Our approach demonstrates how controlled stimulus designs can help localize specific linguistic knowledge in neural speech models.

【11】 Probing the Feasibility of Multilingual Speaker Anonymization
标题: 探讨多语言说话人语音化的可行性
作者:Sarina Meyer,Florian Lux,Ngoc Thang Vu
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:在说话者匿名化中,语音记录以说话者的身份保持隐藏的方式进行修改。虽然这项技术可以帮助保护全球个人的隐私,但目前的研究几乎只关注英语数据,从而限制了这一点。在这项研究中,我们将一个最先进的匿名系统扩展到九种语言,将依赖语言的组件转换为多语言组件。实验测试的匿名语音对隐私攻击和语音恶化的鲁棒性表明,该系统的所有语言的整体成功。结果表明,在英语数据上训练的说话人嵌入可以应用于各种语言,并且一种语言的匿名化性能主要受其使用的语音合成组件的质量影响。摘要:In speaker anonymization, speech recordings are modified in a way that the identity of the speaker remains hidden. While this technology could help to protect the privacy of individuals around the globe, current research restricts this by focusing almost exclusively on English data. In this study, we extend a state-of-the-art anonymization system to nine languages by transforming language-dependent components to their multilingual counterparts. Experiments testing the robustness of the anonymized speech against privacy attacks and speech deterioration show an overall success of this system for all languages. The results suggest that speaker embeddings trained on English data can be applied across languages, and that the anonymization performance for a language is mainly affected by the quality of the speech synthesis component used for it.

【12】 PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
标题: PicoAudio:在文本转音频生成中实现音频事件的精确时间戳和频率可控性
作者:Zeyu Xie,Xuenan Xu,Zhizheng Wu,Mengyue Wu
链接:点击下载PDF文件
摘要:最近,音频生成任务吸引了相当大的研究兴趣。精确的时间可控性对于将音频生成与实际应用集成至关重要。在这项工作中,我们提出了一个时间控制的音频生成框架,PicoAudio。PicoAudio集成了时间信息,通过量身定制的模型设计来指导音频生成。它利用数据抓取,分割,过滤和模拟细粒度的时间对齐的音频文本数据。主观和客观评估都表明,PicoAudio在时间戳和发生频率可控性方面大大超过了当前最先进的生成模型。生成的示例可在演示网站https: PicoAudio.github.io上获得。摘要:Recently, audio generation tasks have attracted considerable research interests. Precise temporal controllability is essential to integrate audio generation with real applications. In this work, we propose a temporal controlled audio generation framework, PicoAudio. PicoAudio integrates temporal information to guide audio generation through tailored model design. It leverages data crawling, segmentation, filtering, and simulation of fine-grained temporally-aligned audio-text data. Both subjective and objective evaluations demonstrate that PicoAudio dramantically surpasses current state-of-the-art generation models in terms of timestamp and occurrence frequency controllability. The generated samples are available on the demo website https: PicoAudio.github.io.

【13】 AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
标题: AudioTime:时间对齐的音频文本基准数据集
作者:Zeyu Xie,Xuenan Xu,Zhizheng Wu,Mengyue Wu
链接:点击下载PDF文件
摘要:音频生成的最新进展已经使得能够从自由形式的文本描述创建高保真音频剪辑。然而,时间关系,音频内容的一个重要特征,目前在主流模型中代表性不足,导致不精确的时间可控性。具体来说,用户无法使用自由格式的文本精确地控制声音事件的时间戳。我们承认,一个重要的因素是缺乏高质量的,时间对齐的音频文本数据集,这对于训练具有时间控制的模型至关重要。注释的时间对齐越多,模型就越能理解音频输出和时间文本提示之间的精确关系。因此,我们提出了一个强对齐的音频文本数据集AudioTime。它提供了丰富的时间信息,如时间戳,持续时间,频率和顺序的文本注释,几乎涵盖了时间控制的所有方面。此外,我们提供了一个全面的测试集和评估指标,以评估各种模型的时间控制性能。示例可在https: zeyuxie29.github.io AudioTime 上获得摘要:Recent advancements in audio generation have enabled the creation of high-fidelity audio clips from free-form textual descriptions. However, temporal relationships, a critical feature for audio content, are currently underrepresented in mainstream models, resulting in an imprecise temporal controllability. Specifically, users cannot accurately control the timestamps of sound events using free-form text. We acknowledge that a significant factor is the absence of high-quality, temporally-aligned audio-text datasets, which are essential for training models with temporal control. The more temporally-aligned the annotations, the better the models can understand the precise relationship between audio outputs and temporal textual prompts. Therefore, we present a strongly aligned audio-text dataset, AudioTime. It provides text annotations rich in temporal information such as timestamps, duration, frequency, and ordering, covering almost all aspects of temporal control. Additionally, we offer a comprehensive test set and evaluation metric to assess the temporal control performance of various models. Examples are available on the https: zeyuxie29.github.io AudioTime

【14】 Nollywood: Let's Go to the Movies!
标题: 尼莱坞:我们去看电影吧!
作者:John E. Ortega,Ibrahim Said Ahmad,William Chen
备注:8 pages, 4 figures, 2 tables
链接:点击下载PDF文件
摘要:尼莱坞,基于印度宝莱坞的想法,是一系列源于尼日利亚的优秀电影。不幸的是,虽然电影是英语的,但由于英语方言,许多母语人士很难理解。在这篇文章中,我们实现了两个目标:(1)创建一个语音字幕模型,能够将尼日利亚英语语音翻译为美国英语;(2)使用最先进的毒性检测器来发现语音的毒性。我们的目的是突出这些视频中的文字,这些文字往往因为缺乏方言理解而被忽视,因为尼日利亚的许多人在家里说豪萨语等母语。摘要:Nollywood, based on the idea of Bollywood from India, is a series of outstanding movies that originate from Nigeria. Unfortunately, while the movies are in English, they are hard to understand for many native speakers due to the dialect of English that is spoken. In this article, we accomplish two goals: (1) create a phonetic sub-title model that is able to translate Nigerian English speech to American English and (2) use the most advanced toxicity detectors to discover how toxic the speech is. Our aim is to highlight the text in these videos which is often times ignored for lack of dialectal understanding due the fact that many people in Nigeria speak a native language like Hausa at home.

【15】 Towards the Next Frontier in Speech Representation Learning Using Disentanglement
标题: 使用解纠缠迈向语音表示学习的下一个前沿
作者:Varun Krishna,Sriram Ganapathy
链接:点击下载PDF文件
摘要:语音表示的自监督学习的流行框架主要集中在语音区域的帧级掩蔽预测上。虽然这已经显示出语音识别和相关任务的有希望的下游任务性能,但这在很大程度上忽略了在较粗糙级别编码的语音因素,例如在整个语音话语中保持一致的扬声器或通道的特性。在这项工作中,我们提出了一个框架,学习解开自我监督(称为Learn2Diss)表示的语音,其中包括帧级和话语级编码器模块。这两个编码器最初是独立学习的,其中帧级模型在很大程度上受到现有自我监督技术的启发,从而学习伪音素表示,而话语级编码器受到池化嵌入的启发,从而学习伪说话者表示。这两个模块的联合学习包括使用基于互信息的标准解开两个编码器。通过几个下游评估实验,我们证明了所提出的Learn2Diss在各种任务上都取得了最先进的结果,其中帧级编码器表示改善了语义任务,而话语级表示改善了非语义任务。摘要:The popular frameworks for self-supervised learning of speech representations have largely focused on frame-level masked prediction of speech regions. While this has shown promising downstream task performance for speech recognition and related tasks, this has largely ignored factors of speech that are encoded at coarser level, like characteristics of the speaker or channel that remain consistent through-out a speech utterance. In this work, we propose a framework for Learning Disentangled Self Supervised (termed as Learn2Diss) representations of speech, which consists of frame-level and an utterance-level encoder modules. The two encoders are initially learned independently, where the frame-level model is largely inspired by existing self supervision techniques, thereby learning pseudo-phonemic representations, while the utterance-level encoder is inspired by constrastive learning of pooled embeddings, thereby learning pseudo-speaker representations. The joint learning of these two modules consists of disentangling the two encoders using a mutual information based criterion. With several downstream evaluation experiments, we show that the proposed Learn2Diss achieves state-of-the-art results on a variety of tasks, with the frame-level encoder representations improving semantic tasks, while the utterance-level representations improve non-semantic tasks.


机器翻译,仅供参考