今日论文合集:cs.SD语音5篇,eess.AS音频处理7篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音

【1】Video-Guided Foley Sound Generation with Multimodal Controls

标题:具有多模式控制的视频引导Foley声音生成
链接:https://arxiv.org/abs/2411.17698
作者:Ziyang Chen,  Prem Seetharaman,  Bryan Russell,  Oriol Nieto,  David Bourgin,  Andrew Owens,  Justin Salamon
备注:Project site: this https URL
摘要:为视频生成声音效果通常需要创建与现实生活来源显著不同的艺术声音效果以及声音设计中的灵活控制。为了解决这个问题,我们引入MultiFoley,一个模型,设计用于视频引导的声音生成,支持多模态条件通过文本,音频和视频。给定无声视频和文本提示,MultiFoley允许用户创建干净的声音(例如,没有风噪声的旋转的滑板轮)或更古怪的声音(例如,使狮子的吼声听起来像猫的叫声)。MultiFoley还允许用户从声音效果(SFX)库或部分视频中选择参考音频进行调节。我们模型的一个关键新颖之处在于它在具有低质量音频和专业SFX录音的互联网视频数据集上进行联合训练,从而实现高质量,全带宽(48kHz)音频生成。通过自动评估和人类研究,我们证明了MultiFoley成功地在各种条件输入中生成同步的高质量声音,并优于现有方法。请参阅我们的项目页面视频结果:https://ificl.github.io/MultiFoley/
摘要:Generating sound effects for videos often requires creating artistic soundeffects that diverge significantly from real-life sources and flexible controlin the sound design. To address this problem, we introduce MultiFoley, a modeldesigned for video-guided sound generation that supports multimodalconditioning through text, audio, and video. Given a silent video and a textprompt, MultiFoley allows users to create clean sounds (e.g., skateboard wheelsspinning without wind noise) or more whimsical sounds (e.g., making a lion'sroar sound like a cat's meow). MultiFoley also allows users to choose referenceaudio from sound effects (SFX) libraries or partial videos for conditioning. Akey novelty of our model lies in its joint training on both internet videodatasets with low-quality audio and professional SFX recordings, enablinghigh-quality, full-bandwidth (48kHz) audio generation. Through automatedevaluations and human studies, we demonstrate that MultiFoley successfullygenerates synchronized high-quality sounds across varied conditional inputs andoutperforms existing methods. Please see our project page for video results:https://ificl.github.io/MultiFoley/

【2】 Visatronic: A Multimodal Decoder-Only Model for Speech Synthesis
标题:Visatronic:用于语音合成的多模式纯解码器模型
链接:https://arxiv.org/abs/2411.17690
作者:Akshita Gupta,  Tatiana Likhomanenko,  Karren Dai Yang,  Richard He Bai,  Zakaria Aldeneh,  Navdeep Jaitly
摘要:在本文中,我们提出了一个新的任务-从视频的人和他们的成绩单(VTTS)生成语音-激励多模态语音生成的新技术。该任务概括了从裁剪的嘴唇视频生成语音的任务,并且也比生成通用音频剪辑的任务更复杂(例如,狗叫声)从视频和文本。多语言版本的任务可能会导致跨语言配音的新技术。我们还提出了一个解码器的多模态模型,我们称之为Visatronic。该模型将视觉、文本和语音直接嵌入到Transformer模型的公共子空间中,并使用自回归损失来学习以扬声器视频和其语音转录为条件的离散化梅尔频谱图的生成模型。通过将所有模态嵌入到一个公共子空间中,Visatronic可以实现比仅使用文本或视频作为输入的模型更好的结果。此外,它提出了一个更简单的方法,多模态语音生成相比,目前的方法,依赖于唇检测器和复杂的架构融合的方式,同时产生更好的结果。由于该模型足够灵活,可以适应将输入排序为序列的不同方式,因此我们仔细探索不同的策略,以更好地了解将信息传播到生成步骤的最佳方式。为了促进对VTTS的进一步研究,我们将发布(i)我们的代码,(ii)大规模VoxCeleb 2数据集的干净transmittance,以及(iii)结合客观和主观指标的VTTS标准化评估协议。
摘要:In this paper, we propose a new task -- generating speech from videos ofpeople and their transcripts (VTTS) -- to motivate new techniques formultimodal speech generation. This task generalizes the task of generatingspeech from cropped lip videos, and is also more complicated than the task ofgenerating generic audio clips (e.g., dog barking) from videos and text.Multilingual versions of the task could lead to new techniques forcross-lingual dubbing. We also present a decoder-only multimodal model for thistask, which we call Visatronic. This model embeds vision, text and speechdirectly into the common subspace of a transformer model and uses anautoregressive loss to learn a generative model of discretized mel-spectrogramsconditioned on speaker videos and transcripts of their speech. By embedding allmodalities into a common subspace, Visatronic can achieve improved results overmodels that use only text or video as input. Further, it presents a muchsimpler approach for multimodal speech generation compared to prevailingapproaches which rely on lip-detectors and complicated architectures to fusemodalities while producing better results. Since the model is flexible enoughto accommodate different ways of ordering inputs as a sequence, we carefullyexplore different strategies to better understand the best way to propagateinformation to the generative steps. To facilitate further research on VTTS, wewill release (i) our code, (ii) clean transcriptions for the large-scaleVoxCeleb2 dataset, and (iii) a standardized evaluation protocol for VTTSincorporating both objective and subjective metrics.

【3】 Scaling Speech-Text Pre-training with Synthetic Interleaved Data
标题:使用合成交织数据扩展语音文本预训练
链接:https://arxiv.org/abs/2411.17607
作者:Aohan Zeng,  Zhengxiao Du,  Mingdao Liu,  Lei Zhang,  Shengmin Jiang,  Yuxiao Dong,  Jie Tang
摘要:语音语言模型(SpeechLM)接受语音输入并产生语音输出,与基于文本的大型语言模型(LLM)相比,允许更自然的人机交互。开发SpeechLM的传统方法受到无监督语音数据和并行语音文本数据的有限可用性的限制,这些数据比文本预训练数据少得多,从而限制了它们作为LLM的可扩展性。我们提出了一种新颖的方法,通过利用源自文本语料库的大规模合成交织数据来扩展语音文本预训练,从而消除了对并行语音文本数据集的需求。我们的方法有效地构建语音文本交错数据采样文本跨度从现有的文本语料库和合成相应的语音跨度使用文本到令牌模型,绕过需要生成实际的语音。我们还采用了一个监督的语音标记来自自动语音识别(ASR)模型,通过将矢量量化的瓶颈到编码器。这种有监督的训练方法导致即使在较低的采样率(例如12.5Hz)下也具有强语义保留的离散语音令牌,同时仍然保持语音重建质量。从一个预训练的语言模型开始,并将我们的预训练扩展到1万亿个令牌(600 B合成交错语音文本数据),我们在语音语言建模和口语问答方面实现了最先进的性能,将口语问答任务的性能从之前的13%(Moshi)提高到31%。我们进一步证明,通过用语音对话数据微调预训练模型,我们可以开发一个端到端的口语聊天机器人,它在会话能力和语音质量方面都达到了与现有基线相当的竞争力,甚至只在语音领域运行。
摘要:Speech language models (SpeechLMs) accept speech input and produce speechoutput, allowing for more natural human-computer interaction compared totext-based large language models (LLMs). Traditional approaches for developingSpeechLMs are constrained by the limited availability of unsupervised speechdata and parallel speech-text data, which are significantly less abundant thantext pre-training data, thereby limiting their scalability as LLMs. We proposea novel approach to scaling speech-text pre-training by leveraging large-scalesynthetic interleaved data derived from text corpora, eliminating the need forparallel speech-text datasets. Our method efficiently constructs speech-textinterleaved data by sampling text spans from existing text corpora andsynthesizing corresponding speech spans using a text-to-token model, bypassingthe need to generate actual speech. We also employ a supervised speechtokenizer derived from an automatic speech recognition (ASR) model byincorporating a vector-quantized bottleneck into the encoder. This supervisedtraining approach results in discrete speech tokens with strong semanticpreservation even at lower sampling rates (e.g. 12.5Hz), while stillmaintaining speech reconstruction quality. Starting from a pre-trained languagemodel and scaling our pre-training to 1 trillion tokens (with 600B syntheticinterleaved speech-text data), we achieve state-of-the-art performance inspeech language modeling and spoken question answering, improving performanceon spoken questions tasks from the previous SOTA of 13% (Moshi) to 31%. Wefurther demonstrate that by fine-tuning the pre-trained model with speechdialogue data, we can develop an end-to-end spoken chatbot that achievescompetitive performance comparable to existing baselines in both conversationalabilities and speech quality, even operating exclusively in the speech domain.

【4】 Comparative Analysis of ASR Methods for Speech Deepfake Detection
标题:语音深度伪造检测的ASB方法比较分析
链接:https://arxiv.org/abs/2411.17349
作者:Davide Salvi,  Amit Kumar Singh Yadav,  Kratika Bhagtani,  Viola Negroni,  Paolo Bestagini,  Edward J. Delp
备注:Published at Asilomar Conference on Signals, Systems, and Computers 2024
摘要:最近的语音deepfake检测技术通常依赖于预先训练的自监督模型。这些系统最初是为自动语音识别(ASR)开发的,已经证明了它们能够提供有意义的语音信号表示,这可以使各种任务受益,包括deepfake检测。在这种情况下,预先训练的模型充当特征提取器,并用于从输入语音中提取嵌入,然后将其馈送到二进制语音deepfake检测器。通过这种方法实现的显著准确性强调了ASR和语音深度伪造检测之间的潜在关系。然而,这种联系还不完全清楚,我们不知道ASR的性能改善是否与更高的语音深度伪造检测能力相对应。在本文中,我们通过系统的分析来解决这个问题。我们考虑了两种不同的预训练自监督ASR模型Whisper和Wav2Vec 2.0,并将其用于语音深度伪造检测任务。这些模型已经发布了多个版本,参数数量不断增加,ASR性能得到增强。我们调查了ASR的性能改善是否与语音deepfake检测的改善相关。我们的研究结果为这两项任务之间的关系提供了见解,并为开发更有效的语音深度伪造检测器提供了有价值的指导。
摘要:Recent techniques for speech deepfake detection often rely on pre-trainedself-supervised models. These systems, initially developed for Automatic SpeechRecognition (ASR), have proved their ability to offer a meaningfulrepresentation of speech signals, which can benefit various tasks, includingdeepfake detection. In this context, pre-trained models serve as featureextractors and are used to extract embeddings from input speech, which are thenfed to a binary speech deepfake detector. The remarkable accuracy achievedthrough this approach underscores a potential relationship between ASR andspeech deepfake detection. However, this connection is not yet entirely clear,and we do not know whether improved performance in ASR corresponds to higherspeech deepfake detection capabilities. In this paper, we address this questionthrough a systematic analysis. We consider two different pre-trainedself-supervised ASR models, Whisper and Wav2Vec 2.0, and adapt them for thespeech deepfake detection task. These models have been released in multipleversions, with increasing number of parameters and enhanced ASR performance. Weinvestigate whether performance improvements in ASR correlate with improvementsin speech deepfake detection. Our results provide insights into therelationship between these two tasks and offer valuable guidance for thedevelopment of more effective speech deepfake detectors.

【5】 DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer  Normalization Mamba-2
标题:DiM-GSYS:采用自适应层规范化Mamba-2的联合语音手势生成
链接:https://arxiv.org/abs/2411.16729
作者:Fan Zhang,  Siyuan Zhao,  Naye Ji,  Zhaohan Wang,  Jingmei Wu,  Fuxing Gao,  Zhenqing Ye,  Leyao Yan,  Lanxin Dai,  Weidong Geng,  Xin Lyu,  Bozuo Zhao,  Dingguo Yu,  Hui Du,  Bin Hu
备注:13 pages, 11 figures
摘要:使用基于变换器的生成模型的语音驱动的手势生成代表了虚拟人创建中的快速发展领域。然而,现有的模型面临着巨大的挑战,由于其平方的时间和空间的复杂性,限制了可扩展性和效率。为了解决这些局限性,我们引入了DiM-GARCH,这是一种利用Mamba-2架构的创新端到端生成模型。DiM-Ghosts具有双组件框架:(1)模糊特征提取器和(2)语音到手势映射模块,两者都构建在Mamba-2上。模糊特征提取器,结合汉语预训练模型和Mamba-2,自主提取隐含的,连续的语音特征。这些特征被合成为统一的潜在表示,然后由语音到手势映射模块进行处理。该模块采用自适应层规范化(AdaLN)增强的Mamba-2机制,在所有序列标记上统一应用转换。这使得能够精确建模语音特征和手势动态之间的微妙相互作用。我们利用扩散模型来训练和推断不同的手势输出。在最新发布的中文协同语音手势数据集上进行的大量主观和客观评估证实了我们提出的模型的有效性。与基于transformer的架构相比,评估表明,我们的方法提供了有竞争力的结果,并显着减少了内存使用量,约为2.4倍,并将推理速度提高了2到4倍。此外,我们还发布了CCG数据集,这是一个中文协同语音手势数据集,包括由专业中国电视广播公司执行的15.97小时(五种场景中的六种风格)的3D全身骨架手势运动。
摘要:Speech-driven gesture generation using transformer-based generative modelsrepresents a rapidly advancing area within virtual human creation. However,existing models face significant challenges due to their quadratic time andspace complexities, limiting scalability and efficiency. To address theselimitations, we introduce DiM-Gestor, an innovative end-to-end generative modelleveraging the Mamba-2 architecture. DiM-Gestor features a dual-componentframework: (1) a fuzzy feature extractor and (2) a speech-to-gesture mappingmodule, both built on the Mamba-2. The fuzzy feature extractor, integrated witha Chinese Pre-trained Model and Mamba-2, autonomously extracts implicit,continuous speech features. These features are synthesized into a unifiedlatent representation and then processed by the speech-to-gesture mappingmodule. This module employs an Adaptive Layer Normalization (AdaLN)-enhancedMamba-2 mechanism to uniformly apply transformations across all sequencetokens. This enables precise modeling of the nuanced interplay between speechfeatures and gesture dynamics. We utilize a diffusion model to train and inferdiverse gesture outputs. Extensive subjective and objective evaluationsconducted on the newly released Chinese Co-Speech Gestures dataset corroboratethe efficacy of our proposed model. Compared with Transformer-basedarchitecture, the assessments reveal that our approach delivers competitiveresults and significantly reduces memory usage, approximately 2.4 times, andenhances inference speeds by 2 to 4 times. Additionally, we released the CCGdataset, a Chinese Co-Speech Gestures dataset, comprising 15.97 hours (sixstyles across five scenarios) of 3D full-body skeleton gesture motion performedby professional Chinese TV broadcasters.

eess.AS音频处理

【1】 Towards Maximum Likelihood Training for Transducer-based Streaming  Speech Recognition
标题:基于传感器的流语音识别的最大可能性训练
链接:https://arxiv.org/abs/2411.17537
作者:Hyeonseung Lee,  Ji Won Yoon,  Sungsoo Kim,  Nam Soo Kim
备注:5 pages, 1 figure, 1 table
摘要:传感器神经网络已成为流式自动语音识别(ASR)的主流方法,在平衡准确性和延迟方面提供了最先进的性能。在传统的框架中,流式换能器模型被训练为基于非流式递归规则来最大化似然函数。然而,这种方法会导致训练和推理之间的不匹配,从而导致变形的可能性问题,并因此导致ASR准确度次优。我们引入了一个数学量化的差距之间的实际可能性和变形的可能性,即前向变量因果补偿(FoCC)。我们还提出了它的估计,FoCCE,作为一个解决方案,以估计确切的可能性。通过在LibriSpeech数据集上的实验,我们表明FoCCE训练提高了流式换能器的准确性。
摘要:Transducer neural networks have emerged as the mainstream approach forstreaming automatic speech recognition (ASR), offering state-of-the-artperformance in balancing accuracy and latency. In the conventional framework,streaming transducer models are trained to maximize the likelihood functionbased on non-streaming recursion rules. However, this approach leads to amismatch between training and inference, resulting in the issue of deformedlikelihood and consequently suboptimal ASR accuracy. We introduce amathematical quantification of the gap between the actual likelihood and thedeformed likelihood, namely forward variable causal compensation (FoCC). Wealso present its estimator, FoCCE, as a solution to estimate the exactlikelihood. Through experiments on the LibriSpeech dataset, we show that FoCCEtraining improves the accuracy of the streaming transducers.

【2】 Typical vs. Atypical Disfluency Classification: Introducing the  IIITH-TISA Corpus and Temporal Context-Based Feature Representations
标题:典型与非典型不流利分类:介绍IIII-TISA数据库和基于时间上下文的特征表示
链接:https://arxiv.org/abs/2411.17149
作者:Priyanka Kommagouni,  Vamshiraghusimha Narasinga,  Purva Barche,  Sai Akarsh C,  Anil Vuppala
摘要:自发交际中的言语不流利可以分为典型和非典型两类。典型的不流利,如犹豫和重复,是日常讲话中的自然现象,而非典型的不流利则表明有病理性疾病,如口吃。区分这些类别对于改善口吃者(PWS)的语音助理(VA)至关重要,因为口吃者经常由于错误识别语音终止而面临过早的切断。准确的分类也有助于早期发现儿童口吃,防止误诊为语言发展不流利。本研究介绍了III-TISA数据集,第一个印度英语口吃语料库,捕捉非典型的不流利。此外,我们扩展了III-IED数据集与典型的不流利的详细注释。我们提出了感知增强零时间窗倒谱系数(PE-ZTWCC)结合移位三角倒谱系数(SDC)作为输入功能的浅时间延迟神经网络(TDNN)分类器,捕捉本地和更广泛的时间背景。我们的方法实现了平均F1得分为85.01%的不流利分类,优于传统的功能。
摘要:Speech disfluencies in spontaneous communication can be categorized as eithertypical or atypical. Typical disfluencies, such as hesitations and repetitions,are natural occurrences in everyday speech, while atypical disfluencies areindicative of pathological disorders like stuttering. Distinguishing betweenthese categories is crucial for improving voice assistants (VAs) for PersonsWho Stutter (PWS), who often face premature cutoffs due to misidentification ofspeech termination. Accurate classification also aids in detecting stutteringearly in children, preventing misdiagnosis as language development disfluency.This research introduces the IIITH-TISA dataset, the first Indian Englishstammer corpus, capturing atypical disfluencies. Additionally, we extend theIIITH-IED dataset with detailed annotations for typical disfluencies. Wepropose Perceptually Enhanced Zero-Time Windowed Cepstral Coefficients(PE-ZTWCC) combined with Shifted Delta Cepstra (SDC) as input features to ashallow Time Delay Neural Network (TDNN) classifier, capturing both local andwider temporal contexts. Our method achieves an average F1 score of 85.01% fordisfluency classification, outperforming traditional features.

【3】 k2SSL: A Faster and Better Framework for Self-Supervised Speech  Representation Learning
标题:k2 SSL:更快、更好的自我监督语音表示学习框架
链接:https://arxiv.org/abs/2411.17100
作者:Yifan Yang,  Jianheng Zhuo,  Zengrui Jin,  Ziyang Ma,  Xiaoyu Yang,  Zengwei Yao,  Liyong Guo,  Wei Kang,  Fangjun Kuang,  Long Lin,  Daniel Povey,  Xie Chen
备注:Submitted to ICASSP 2025
摘要:在语音编码器架构的进步和数据集的扩展的推动下,自监督学习(SSL)在语音相关任务中取得了巨大成功。虽然Transformer和Conformer架构已经主导了SSL主干,但像Zipformer这样在自动语音识别(ASR)方面表现出色的编码器在SSL中仍然没有得到开发。同时,现有SSL培训框架(如fairseq)中的数据处理效率低下,给管理不断增长的培训数据带来了挑战。为了解决这些问题,我们提出了k2 SSL,这是一个开源框架,它提供了更快,更高效,更好的自我监督语音表示学习,重点是下游ASR任务。优化的HuBERT和提出的基于Zipformer的SSL系统在SSL训练期间的训练时间和内存使用量都大幅减少。在LibriSpeech和Libri-Light上的实验表明,基于Zipformer的SSL系统的性能明显优于可比的HuBERT和WavLM系统,在监督微调后,与HuBERT Base相比,dev-other/test-other的相对WER降低高达34.8%/32.4%,以及总GPU时间的3.5倍预训练加速。
摘要:Self-supervised learning (SSL) has achieved great success in speech-relatedtasks, driven by advancements in speech encoder architectures and the expansionof datasets. While Transformer and Conformer architectures have dominated SSLbackbones, encoders like Zipformer, which excel in automatic speech recognition(ASR), remain unexplored in SSL. Concurrently, inefficiencies in dataprocessing within existing SSL training frameworks, such as fairseq, posechallenges in managing the growing volumes of training data. To address theseissues, we propose k2SSL, an open-source framework that offers faster, morememory-efficient, and better-performing self-supervised speech representationlearning, with a focus on downstream ASR tasks. The optimized HuBERT andproposed Zipformer-based SSL systems exhibit substantial reductions in bothtraining time and memory usage during SSL training. Experiments on LibriSpeechand Libri-Light demonstrate that Zipformer-based SSL systems significantlyoutperform comparable HuBERT and WavLM systems, achieving a relative WERreduction on dev-other/test-other of up to 34.8%/32.4% compared to HuBERT Baseafter supervised fine-tuning, along with a 3.5x pre-training speedup in totalGPU hours.

【4】 Video-Guided Foley Sound Generation with Multimodal Controls
标题:具有多模式控制的视频引导Foley声音生成
链接:https://arxiv.org/abs/2411.17698
作者:Ziyang Chen,  Prem Seetharaman,  Bryan Russell,  Oriol Nieto,  David Bourgin,  Andrew Owens,  Justin Salamon
备注:Project site: this https URL
摘要:为视频生成声音效果通常需要创建与现实生活来源显著不同的艺术声音效果以及声音设计中的灵活控制。为了解决这个问题,我们引入MultiFoley,一个模型,设计用于视频引导的声音生成,支持多模态条件通过文本,音频和视频。给定无声视频和文本提示,MultiFoley允许用户创建干净的声音(例如,没有风噪声的旋转的滑板轮)或更古怪的声音(例如,使狮子的吼声听起来像猫的叫声)。MultiFoley还允许用户从声音效果(SFX)库或部分视频中选择参考音频进行调节。我们模型的一个关键新颖之处在于它在具有低质量音频和专业SFX录音的互联网视频数据集上进行联合训练,从而实现高质量,全带宽(48kHz)音频生成。通过自动评估和人类研究,我们证明了MultiFoley成功地在各种条件输入中生成同步的高质量声音,并优于现有方法。请参阅我们的项目页面视频结果:https://ificl.github.io/MultiFoley/
摘要:Generating sound effects for videos often requires creating artistic soundeffects that diverge significantly from real-life sources and flexible controlin the sound design. To address this problem, we introduce MultiFoley, a modeldesigned for video-guided sound generation that supports multimodalconditioning through text, audio, and video. Given a silent video and a textprompt, MultiFoley allows users to create clean sounds (e.g., skateboard wheelsspinning without wind noise) or more whimsical sounds (e.g., making a lion'sroar sound like a cat's meow). MultiFoley also allows users to choose referenceaudio from sound effects (SFX) libraries or partial videos for conditioning. Akey novelty of our model lies in its joint training on both internet videodatasets with low-quality audio and professional SFX recordings, enablinghigh-quality, full-bandwidth (48kHz) audio generation. Through automatedevaluations and human studies, we demonstrate that MultiFoley successfullygenerates synchronized high-quality sounds across varied conditional inputs andoutperforms existing methods. Please see our project page for video results:https://ificl.github.io/MultiFoley/

【5】 Visatronic: A Multimodal Decoder-Only Model for Speech Synthesis
标题:Visatronic:用于语音合成的多模式纯解码器模型
链接:https://arxiv.org/abs/2411.17690
作者:Akshita Gupta,  Tatiana Likhomanenko,  Karren Dai Yang,  Richard He Bai,  Zakaria Aldeneh,  Navdeep Jaitly
摘要:在本文中,我们提出了一个新的任务-从视频的人和他们的成绩单(VTTS)生成语音-激励多模态语音生成的新技术。该任务概括了从裁剪的嘴唇视频生成语音的任务,并且也比生成通用音频剪辑的任务更复杂(例如,狗叫声)从视频和文本。多语言版本的任务可能会导致跨语言配音的新技术。我们还提出了一个解码器的多模态模型,我们称之为Visatronic。该模型将视觉、文本和语音直接嵌入到Transformer模型的公共子空间中,并使用自回归损失来学习以扬声器视频和其语音转录为条件的离散化梅尔频谱图的生成模型。通过将所有模态嵌入到一个公共子空间中,Visatronic可以实现比仅使用文本或视频作为输入的模型更好的结果。此外,它提出了一种更简单的方法,多模态语音生成相比,目前的方法,依赖于唇检测器和复杂的架构融合的方式,同时产生更好的结果。由于该模型足够灵活,可以适应将输入排序为序列的不同方式,因此我们仔细探索不同的策略,以更好地了解将信息传播到生成步骤的最佳方式。为了促进对VTTS的进一步研究,我们将发布(i)我们的代码,(ii)大规模VoxCeleb 2数据集的干净transmittance,以及(iii)结合客观和主观指标的VTTS标准化评估协议。
摘要:In this paper, we propose a new task -- generating speech from videos ofpeople and their transcripts (VTTS) -- to motivate new techniques formultimodal speech generation. This task generalizes the task of generatingspeech from cropped lip videos, and is also more complicated than the task ofgenerating generic audio clips (e.g., dog barking) from videos and text.Multilingual versions of the task could lead to new techniques forcross-lingual dubbing. We also present a decoder-only multimodal model for thistask, which we call Visatronic. This model embeds vision, text and speechdirectly into the common subspace of a transformer model and uses anautoregressive loss to learn a generative model of discretized mel-spectrogramsconditioned on speaker videos and transcripts of their speech. By embedding allmodalities into a common subspace, Visatronic can achieve improved results overmodels that use only text or video as input. Further, it presents a muchsimpler approach for multimodal speech generation compared to prevailingapproaches which rely on lip-detectors and complicated architectures to fusemodalities while producing better results. Since the model is flexible enoughto accommodate different ways of ordering inputs as a sequence, we carefullyexplore different strategies to better understand the best way to propagateinformation to the generative steps. To facilitate further research on VTTS, wewill release (i) our code, (ii) clean transcriptions for the large-scaleVoxCeleb2 dataset, and (iii) a standardized evaluation protocol for VTTSincorporating both objective and subjective metrics.

【6】 Scaling Speech-Text Pre-training with Synthetic Interleaved Data
标题:使用合成交织数据扩展语音文本预训练
链接:https://arxiv.org/abs/2411.17607
作者:Aohan Zeng,  Zhengxiao Du,  Mingdao Liu,  Lei Zhang,  Shengmin Jiang,  Yuxiao Dong,  Jie Tang
摘要:语音语言模型(SpeechLM)接受语音输入并产生语音输出,与基于文本的大型语言模型(LLM)相比,允许更自然的人机交互。开发SpeechLM的传统方法受到无监督语音数据和并行语音文本数据的有限可用性的限制,这些数据比文本预训练数据少得多,从而限制了它们作为LLM的可扩展性。我们提出了一种新颖的方法,通过利用源自文本语料库的大规模合成交织数据来扩展语音文本预训练,从而消除了对并行语音文本数据集的需求。我们的方法有效地构建语音文本交错数据采样文本跨度从现有的文本语料库和合成相应的语音跨度使用文本到令牌模型,绕过需要生成实际的语音。我们还采用了一个监督的语音标记来自自动语音识别(ASR)模型,通过将矢量量化的瓶颈到编码器。这种有监督的训练方法导致即使在较低的采样率(例如12.5Hz)下也具有强语义保留的离散语音令牌,同时仍然保持语音重建质量。从一个预训练的语言模型开始,并将我们的预训练扩展到1万亿个令牌(600 B合成交错语音文本数据),我们在语音语言建模和口语问答方面实现了最先进的性能,将口语问答任务的性能从之前的13%(Moshi)提高到31%。我们进一步证明,通过用语音对话数据微调预训练模型,我们可以开发一个端到端的口语聊天机器人,它在会话能力和语音质量方面都达到了与现有基线相当的竞争力,甚至只在语音领域运行。
摘要:Speech language models (SpeechLMs) accept speech input and produce speechoutput, allowing for more natural human-computer interaction compared totext-based large language models (LLMs). Traditional approaches for developingSpeechLMs are constrained by the limited availability of unsupervised speechdata and parallel speech-text data, which are significantly less abundant thantext pre-training data, thereby limiting their scalability as LLMs. We proposea novel approach to scaling speech-text pre-training by leveraging large-scalesynthetic interleaved data derived from text corpora, eliminating the need forparallel speech-text datasets. Our method efficiently constructs speech-textinterleaved data by sampling text spans from existing text corpora andsynthesizing corresponding speech spans using a text-to-token model, bypassingthe need to generate actual speech. We also employ a supervised speechtokenizer derived from an automatic speech recognition (ASR) model byincorporating a vector-quantized bottleneck into the encoder. This supervisedtraining approach results in discrete speech tokens with strong semanticpreservation even at lower sampling rates (e.g. 12.5Hz), while stillmaintaining speech reconstruction quality. Starting from a pre-trained languagemodel and scaling our pre-training to 1 trillion tokens (with 600B syntheticinterleaved speech-text data), we achieve state-of-the-art performance inspeech language modeling and spoken question answering, improving performanceon spoken questions tasks from the previous SOTA of 13% (Moshi) to 31%. Wefurther demonstrate that by fine-tuning the pre-trained model with speechdialogue data, we can develop an end-to-end spoken chatbot that achievescompetitive performance comparable to existing baselines in both conversationalabilities and speech quality, even operating exclusively in the speech domain.

【7】 Automatic Album Sequencing
标题:自动专辑排序
链接:https://arxiv.org/abs/2411.07772
作者:Vincent Herrmann,  Dylan R. Ashley,  Jürgen Schmidhuber
备注:presented as a late breaking demo in the 25th International Society for Music Information Retrieval Conference; 3 pages in main text + 1 page of references, 3 figures in main text; source code available at this https URL
摘要:专辑排序是专辑制作过程中至关重要的一部分。最近,提出了一种数据驱动的方法,通过提取集合中项目的叙事本质来对独立媒体的一般集合进行排序。虽然这种方法意味着一种专辑排序技术,但它并不广泛适用于技术含量较低的受众,需要先进的机器学习技术知识才能使用。为了解决这个问题,我们引入了一个新的用户友好的基于Web的工具,允许技术含量较低的观众上传音乐曲目,只需单击一下即可执行此技术,随后将结果以清晰的可视化方式呈现给用户。为了增加用户可用的模板数量并解决以前工作的缺点,我们还引入了一种新的直接基于transformer的相册排序方法。我们发现,我们更直接的方法优于随机基线,但没有达到相同的性能作为叙事本质的方法。这两种方法都包含在我们基于Web的用户界面中,并且可以在https://github.com/dylanashley/automatic-album-sequencing上公开获得该界面以及我们实现的完整副本
摘要:Album sequencing is a critical part of the album production process.Recently, a data-driven approach was proposed that sequences generalcollections of independent media by extracting the narrative essence of theitems in the collections. While this approach implies an album sequencingtechnique, it is not widely accessible to a less technical audience, requiringadvanced knowledge of machine learning techniques to use. To address this, weintroduce a new user-friendly web-based tool that allows a less technicalaudience to upload music tracks, execute this technique in one click, andsubsequently presents the result in a clean visualization to the user. To bothincrease the number of templates available to the user and address shortcomingsof previous work, we also introduce a new direct transformer-based albumsequencing method. We find that our more direct method outperforms a randombaseline but does not reach the same performance as the narrative essenceapproach. Both methods are included in our web-based user interface, and this-- alongside a full copy of our implementation -- is publicly available athttps://github.com/dylanashley/automatic-album-sequencing

机器翻译由腾讯交互翻译提供,仅供参考