今日论文合集:cs.SD语音32篇,eess.AS音频处理36篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Presto! Distilling Steps and Layers for Accelerating Music Generation
标题: 快点!提炼步骤和层次以加速音乐生成
作者: Zachary Novack, Ge Zhu, Jonah Casebeer, Julian McAuley, Taylor Berg-Kirkpatrick, Nicholas J. Bryan
链接:点击下载PDF文件
摘要:尽管基于扩散的文本到音乐(TTM)方法取得了进展,但高效、高质量的生成仍然是一个挑战。我们介绍Presto!,一种通过减少采样步骤和每步成本来加速基于分数的扩散Transformers的推理的方法。为了减少步骤,我们开发了一种新的基于分数的分布匹配蒸馏(DMD)方法,用于EDM系列扩散模型,这是第一种基于GAN的TTM蒸馏方法。为了降低每一步的成本,我们开发了一个简单的,但强大的改进最近层蒸馏方法,通过更好地保留隐藏状态方差来提高学习。最后,我们结合我们的步骤和层蒸馏方法在一起的一个双方面的方法。我们独立评估我们的阶梯和分层蒸馏方法,并展示每种方法的最佳性能。我们的组合蒸馏方法可以产生高质量的输出,提高多样性,将我们的基础模型加速10- 18倍(32秒单声道 立体声44.1kHz的230 435 ms延迟,比同类SOTA快15倍)-据我们所知,这是最快的高质量TTM。可以在https: presto-music.github.io web 上找到合理的例子。摘要:Despite advances in diffusion-based text-to-music (TTM) methods, efficient, high-quality generation remains a challenge. We introduce Presto!, an approach to inference acceleration for score-based diffusion transformers via reducing both sampling steps and cost per step. To reduce steps, we develop a new score-based distribution matching distillation (DMD) method for the EDM-family of diffusion models, the first GAN-based distillation method for TTM. To reduce the cost per step, we develop a simple, but powerful improvement to a recent layer distillation method that improves learning via better preserving hidden state variance. Finally, we combine our step and layer distillation methods together for a dual-faceted approach. We evaluate our step and layer distillation methods independently and show each yield best-in-class performance. Our combined distillation method can generate high-quality outputs with improved diversity, accelerating our base model by 10-18x (230 435ms latency for 32 second mono stereo 44.1kHz, 15x faster than comparable SOTA) -- the fastest high-quality TTM to our knowledge. Sound examples can be found at https: presto-music.github.io web .

【2】 Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
标题: 使用多尺度特征的对比损失改进说话者表示
作者: Satvik Dixit, Massa Baali, Rita Singh, Bhiksha Raj
链接:点击下载PDF文件
摘要:随着多尺度特征聚合(MFA)架构的引入,说话人确认系统已经取得了重大进展,例如MFA-Conformer和ECAPA-TDNN。这些模型通过在池化层和投影层之前连接中间特征图来利用来自不同网络深度的信息,表明即使是较浅的特征图也会编码有价值的说话者特定信息。在此基础上,我们提出了一个多尺度特征对比(MFCon)损失,直接提高这些中间表示的质量。我们的MFCon损失将对比学习应用于网络中的所有特征图,鼓励模型在中间阶段学习更多的判别表示。通过执行更好的特征映射学习,我们表明,由此产生的扬声器嵌入表现出更高的区分能力。我们的方法实现了9.05%的改善,在等错误率(EER)相比,标准的MFA-构象的VoxCeleb-1 O测试集。摘要:Speaker verification systems have seen significant advancements with the introduction of Multi-scale Feature Aggregation (MFA) architectures, such as MFA-Conformer and ECAPA-TDNN. These models leverage information from various network depths by concatenating intermediate feature maps before the pooling and projection layers, demonstrating that even shallower feature maps encode valuable speaker-specific information. Building upon this foundation, we propose a Multi-scale Feature Contrastive (MFCon) loss that directly enhances the quality of these intermediate representations. Our MFCon loss applies contrastive learning to all feature maps within the network, encouraging the model to learn more discriminative representations at the intermediate stage itself. By enforcing better feature map learning, we show that the resulting speaker embeddings exhibit increased discriminative power. Our method achieves a 9.05% improvement in equal error rate (EER) compared to the standard MFA-Conformer on the VoxCeleb-1O test set.

【3】 RelUNet: Relative Channel Fusion U-Net for Multichannel Speech Enhancement
标题: RelUNet:用于多通道语音增强的相对通道融合U-Net
作者: Ibrahim Aldarmaki, Thamar Solorio, Bhiksha Raj, Hanan Aldarmaki
链接:点击下载PDF文件
摘要:神经多通道语音增强模型,特别是那些基于U-Net架构,表现出良好的性能和推广潜力。这些模型通常独立编码输入通道,并在网络的后期阶段集成通道。在本文中,我们提出了一种新的修改这些模型,从一开始就将相关信息,其中每个通道的处理与参考通道通过堆叠。该输入策略利用比较差异自适应地融合通道之间的信息,从而捕获关键的空间信息并提高整体性能。在CHiME-3数据集上进行的实验证明了各种架构的语音增强指标的改进。摘要:Neural multi-channel speech enhancement models, in particular those based on the U-Net architecture, demonstrate promising performance and generalization potential. These models typically encode input channels independently, and integrate the channels during later stages of the network. In this paper, we propose a novel modification of these models by incorporating relative information from the outset, where each channel is processed in conjunction with a reference channel through stacking. This input strategy exploits comparative differences to adaptively fuse information between channels, thereby capturing crucial spatial information and enhancing the overall performance. The experiments conducted on the CHiME-3 dataset demonstrate improvements in speech enhancement metrics across various architectures.

【4】 Stage-Wise and Prior-Aware Neural Speech Phase Prediction
标题: 分阶段和优先感知神经语音阶段预测
作者: Fei Liu, Yang Ai, Hui-Peng Du, Ye-Xin Lu, Rui-Chen Zheng, Zhen-Hua Ling
备注:Accepted by SLT2024
链接:点击下载PDF文件
摘要:本文提出了一种新的逐段和先验感知的神经语音相位预测(SP-NSPP)模型,该模型通过两段神经网络从输入的幅度谱中预测相位谱。在初始先验构造阶段,我们从振幅谱中初步预测出一个粗略的先验相位谱。随后的细化阶段将幅度谱变换成以先前相位为条件的细化的高质量相位谱。这两个阶段的网络都使用ConvNeXt v2块作为主干,并通过创新地引入相位谱鉴别器(PSD)来采用对抗训练。为了进一步提高细化相位的连续性,我们还在细化阶段引入了时频积分差分(TFID)损耗。实验结果表明,与基于神经网络的无先验相位预测方法相比,SP-NSPP由于引入了粗相位先验和多样化的训练准则,获得了更高的相位预测精度。与迭代相位估计算法相比,我们提出的SP-NSPP不需要多轮分阶段迭代,从而产生更高的效率。摘要:This paper proposes a novel Stage-wise and Prior-aware Neural Speech Phase Prediction (SP-NSPP) model, which predicts the phase spectrum from input amplitude spectrum by two-stage neural networks. In the initial prior-construction stage, we preliminarily predict a rough prior phase spectrum from the amplitude spectrum. The subsequent refinement stage transforms the amplitude spectrum into a refined high-quality phase spectrum conditioned on the prior phase. Networks in both stages use ConvNeXt v2 blocks as the backbone and adopt adversarial training by innovatively introducing a phase spectrum discriminator (PSD). To further improve the continuity of the refined phase, we also incorporate a time-frequency integrated difference (TFID) loss in the refinement stage. Experimental results confirm that, compared to neural network-based no-prior phase prediction methods, the proposed SP-NSPP achieves higher phase prediction accuracy, thanks to introducing the coarse phase priors and diverse training criteria. Compared to iterative phase estimation algorithms, our proposed SP-NSPP does not require multiple rounds of staged iterations, resulting in higher generation efficiency.

【5】 Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation
标题: Art 2 Mus:通过跨模式一代架起视觉艺术与音乐的桥梁
作者: Ivan Rinaldi, Nicola Fanelli, Giovanna Castellano, Gennaro Vessio
备注:Presented at the AI for Visual Arts (AI4VA) workshop at ECCV 2024
链接:点击下载PDF文件
摘要:Artificial Intelligence and generative models have revolutionized music creation, with many models leveraging textual or visual prompts for guidance. However, existing image-to-music models are limited to simple images, lacking the capability to generate music from complex digitized artworks. To address this gap, we introduce $ mathcal{A} textit{rt2} mathcal{M} textit{us}$, a novel model designed to create music from digitized artworks or text inputs. $ mathcal{A} textit{rt2} mathcal{M} textit{us}$ extends the AudioLDM~2 architecture, a text-to-audio model, and employs our newly curated datasets, created via ImageBind, which pair digitized artworks with music. Experimental results demonstrate that $ mathcal{A} textit{rt2} mathcal{M} textit{us}$ can generate music that resonates with the input stimuli. These findings suggest promising applications in multimedia art, interactive installations, and AI-driven creative tools.摘要:Artificial Intelligence and generative models have revolutionized music creation, with many models leveraging textual or visual prompts for guidance. However, existing image-to-music models are limited to simple images, lacking the capability to generate music from complex digitized artworks. To address this gap, we introduce $ mathcal{A} textit{rt2} mathcal{M} textit{us}$, a novel model designed to create music from digitized artworks or text inputs. $ mathcal{A} textit{rt2} mathcal{M} textit{us}$ extends the AudioLDM~2 architecture, a text-to-audio model, and employs our newly curated datasets, created via ImageBind, which pair digitized artworks with music. Experimental results demonstrate that $ mathcal{A} textit{rt2} mathcal{M} textit{us}$ can generate music that resonates with the input stimuli. These findings suggest promising applications in multimedia art, interactive installations, and AI-driven creative tools.

【6】 Attentive-based Multi-level Feature Fusion for Voice Disorder Diagnosis
标题: 基于注意力的多层特征融合用于语音障碍诊断
作者: Lipeng Shen, Yifan Xiong, Dongyue Guo, Wei Mo, Lingyu Yu, Hui Yang, Yi Lin
链接:点击下载PDF文件
摘要:语音障碍以各种方式对日常生活质量产生负面影响。然而,由于数据集有限,从原始音频中准确识别病理特征的类别仍然是一个相当大的挑战。一个很有前途的方法来处理这个问题是提取多层次的病理信息,在语音的综合方式融合特征的潜在空间。本文设计了一种新的框架,探索高质量的特征融合的方式,有效的和广义的检测性能。具体而言,该模型采用两阶段训练模式:(1)采用在多个领域都表现出显著效果的ECAPA-TDNN和Wav 2 vec 2.0,从原始音频中学习普遍的病理信息;(2)专门设计了一个注意融合模块,用于建立EcapTdnn和Wav 2 vec 2.0分别投影的病理特征之间的相互作用,并引导多个特征的融合。层融合,整个模型由自动语音病理检测任务从预训练的特征联合微调。最后,在FEMH和SVD数据集上的综合实验表明,该框架优于竞争基线,达到90.51%和87.68%的准确率。摘要:Voice disorders negatively impact the quality of daily life in various ways. However, accurately recognizing the category of pathological features from raw audio remains a considerable challenge due to the limited dataset. A promising method to handle this issue is extracting multi-level pathological information from speech in a comprehensive manner by fusing features in the latent space. In this paper, a novel framework is designed to explore the way of high-quality feature fusion for effective and generalized detection performance. Specifically, the proposed model follows a two-stage training paradigm: (1) ECAPA-TDNN and Wav2vec 2.0 which have shown remarkable effectiveness in various domains are employed to learn the universal pathological information from raw audio; (2) An attentive fusion module is dedicatedly designed to establish the interaction between pathological features projected by EcapTdnn and Wav2vec 2.0 respectively and guide the multi-layer fusion, the entire model is jointly fine-tuned from pre-trained features by the automatic voice pathology detection task. Finally, comprehensive experiments on the FEMH and SVD datasets demonstrate that the proposed framework outperforms the competitive baselines, and achieves the accuracy of 90.51% and 87.68%.

【7】 Modeling and Estimation of Vocal Tract and Glottal Source Parameters Using ARMAX-LF Model
标题: 基于ARMAX-LF模型的声道和喉舌源参数建模与估计
作者: Kai Lia, Masato Akagia, Yongwei Lib, Masashi Unokia
链接:点击下载PDF文件
摘要:来自原始语音的元音的声道和声门源参数的建模和估计通常可以通过使用具有外源输入的自回归(ARX)模型和具有基于迭代的估计方法的Liljencrants-Fant(LF)模型来完成。然而,声道滤波器建模中的全极点自回归模型不能提供反共振峰(零)的位置,这增加了某些类别的语音声音(诸如鼻音、摩擦音和塞音)的估计误差。在本文中,我们提出了自回归移动平均线eXogenous LF(ARMAX-LF)模型扩展的ARX-LF模型,以更广泛的语音,包括元音和鼻音辅音。LF模型将声门源导数表示为参数化时域模型,并且ARMAX模型将声道表示为具有额外的外源LF激励作为输入的极零滤波器。为了以更少的误差估计多个参数,我们首先利用深度神经网络(DNN)强大的非线性拟合能力,从提取的声门源导数或语音波形到相应的LF参数建立映射。然后,声门源和声道参数可以估计较少的估计误差和没有任何迭代的分析合成策略。使用线性源滤波器模型的合成语音、使用物理模型的合成语音和真实语音信号的实验结果表明,所提出的ARMAX-LF模型与基于DNN的估计方法可以估计元音和鼻音的参数,具有更少的误差和估计时间。摘要:Modeling and estimation of the vocal tract and glottal source parameters of vowels from raw speech can be typically done by using the Auto-Regressive with eXogenous input (ARX) model and Liljencrants-Fant (LF) model with an iteration-based estimation approach. However, the all-pole autoregressive model in the modeling of vocal tract filters cannot provide the locations of anti-formants (zeros), which increases the estimation errors in certain classes of speech sounds, such as nasal, fricative, and stop consonants. In this paper, we propose the Auto-Regressive Moving Average eXogenous with LF (ARMAX-LF) model to extend the ARX-LF model to a wider variety of speech sounds, including vowels and nasalized consonants. The LF model represents the glottal source derivative as a parametrized time-domain model, and the ARMAX model represents the vocal tract as a pole-zero filter with an additional exogenous LF excitation as input. To estimate multiple parameters with fewer errors, we first utilize the powerful nonlinear fitting ability of deep neural networks (DNNs) to build a mapping from extracted glottal source derivatives or speech waveforms to corresponding LF parameters. Then, glottal source and vocal tract parameters can be estimated with fewer estimation errors and without any iterations as in the analysis-by-synthesis strategy. Experimental results with synthesized speech using the linear source-filter model, synthesized speech using the physical model, and real speech signals showed that the proposed ARMAX-LF model with a DNN-based estimation method can estimate the parameters of both vowels and nasalized sounds with fewer errors and estimation time.

【8】 Demo of Zero-Shot Guitar Amplifier Modelling: Enhancing Modeling with Hyper Neural Networks
标题: Zero-Shot吉他放大器建模演示:用超神经网络增强建模
作者: Yu-Hua Chen, Yuan-Chiao Cheng, Yen-Tung Yeh, Jui-Te Wu, Yu-Hsiang Ho, Jyh-Shing Roger Jang, Yi-Hsuan Yang
备注:demo of the ISMIR paper
链接:点击下载PDF文件
摘要:电吉他音调建模通常集中于从干净音频到增强器渲染音频的非线性变换。传统方法依赖于一对一映射,将设备参数纳入神经模型以复制特定的放大器。然而,这些方法受到特定训练数据需求的限制。在本文中,我们适应了一个模型的基础上,以前的工作,它利用了音调嵌入编码器和功能明智的线性调制(薄膜)条件方法。在这项工作中,我们使用基于超网络的门控卷积网络(GCN)来改变条件反射方法,以生成将干净输入与参考音频的音调特征混合的音频。通过扩展训练数据以覆盖更广泛的放大器音调,我们的模型能够捕获更广泛的音调。此外,我们还开发了一个实时插件来演示系统的实际应用,让用户可以交互式地体验其性能。我们的研究结果表明,该系统实现了优越的音调建模的通用性相比,传统的方法。摘要:Electric guitar tone modeling typically focuses on the non-linear transformation from clean to amplifier-rendered audio. Traditional methods rely on one-to-one mappings, incorporating device parameters into neural models to replicate specific amplifiers. However, these methods are limited by the need for specific training data. In this paper, we adapt a model based on the previous work, which leverages a tone embedding encoder and a feature wise linear modulation (FiLM) condition method. In this work, we altered conditioning method using a hypernetwork-based gated convolutional network (GCN) to generate audio that blends clean input with the tone characteristics of reference audio. By extending the training data to cover a wider variety of amplifier tones, our model is able to capture a broader range of tones. Additionally, we developed a real-time plugin to demonstrate the system's practical application, allowing users to experience its performance interactively. Our results indicate that the proposed system achieves superior tone modeling versatility compared to traditional methods.

【9】 UniMuMo: Unified Text, Music and Motion Generation
标题: UniMuMo:统一的文本、音乐和动作生成
作者: Han Yang, Kun Su, Yutong Zhang, Jiaben Chen, Kaizhi Qian, Gaowen Liu, Chuang Gan
链接:点击下载PDF文件
摘要:我们介绍了UniMuMo,一个统一的多模态模型,能够将任意文本,音乐和运动数据作为输入条件,以生成所有三种模态的输出。为了解决缺乏时间同步数据的问题,我们根据节奏模式对齐未配对的音乐和运动数据,以利用现有的大规模纯音乐和纯运动数据集。通过将音乐、运动和文本转换为基于标记的表示,我们的模型通过统一的编码器-解码器Transformer架构来桥接这些模态。为了在一个框架内支持多个生成任务,我们引入了几个体系结构的改进。我们建议用音乐码本对运动进行编码,将运动映射到与音乐相同的特征空间。我们引入了一个音乐运动并行生成方案,统一到一个单一的Transformer解码器架构的音乐运动联合生成的单一训练任务的所有音乐和运动生成任务。此外,该模型是通过微调现有预训练的单模态模型来设计的,从而显着降低了计算需求。大量的实验表明,UniMuMo实现了竞争力的结果,所有单向生成基准跨越音乐,运动和文本形式。定量结果可在 href{https: hanyangclarence.github.io unimumo_demo }{project page}中找到。摘要:We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data, we align unpaired music and motion data based on rhythmic patterns to leverage existing large-scale music-only and motion-only datasets. By converting music, motion, and text into token-based representation, our model bridges these modalities through a unified encoder-decoder transformer architecture. To support multiple generation tasks within a single framework, we introduce several architectural improvements. We propose encoding motion with a music codebook, mapping motion into the same feature space as music. We introduce a music-motion parallel generation scheme that unifies all music and motion generation tasks into a single transformer decoder architecture with a single training task of music-motion joint generation. Moreover, the model is designed by fine-tuning existing pre-trained single-modality models, significantly reducing computational demands. Extensive experiments demonstrate that UniMuMo achieves competitive results on all unidirectional generation benchmarks across music, motion, and text modalities. Quantitative results are available in the href{https: hanyangclarence.github.io unimumo_demo }{project page}.

【10】 Configurable Multilingual ASR with Speech Summary Representations
标题: 具有语音摘要表示的可配置多语言ASB
作者: Harrison Zhu, Ivan Fung, Yingke Zhu, Lahiru Samarakoon
备注:A preprint
链接:点击下载PDF文件
摘要:世界上大约有一半的人口是多语言的,这使得多语言ASR(MASR)至关重要。当地面实况语言事先未知时,部署多个单语模型具有挑战性。这激发了对可配置的多语言MASR模型的研究工作,这些模型可以手动提示或自动调整以识别特定语言。在本文中,我们提出了可配置的MASR模型与摘要向量(csvMASR),一种新的架构,旨在提高可配置性。我们的方法利用适配器,并引入语音摘要向量表示,灵感来自语音日记中的会话摘要表示,在话语级别结合特定语言组件的输出。我们还将一个辅助语言分类损失,以提高可配置性。使用多语言Libripeech(MLS)数据集中的7种语言的数据,csvMASR优于现有的MASR模型,并将单词错误率(WER)从10.33%降低到9.95%。此外,csvMASR在语言分类和提示任务中表现出卓越的性能。摘要:Approximately half of the world's population is multilingual, making multilingual ASR (MASR) essential. Deploying multiple monolingual models is challenging when the ground-truth language is unknown in advance. This motivates research efforts on configurable multilingual MASR models that can be prompted manually or adapted automatically to recognise specific languages. In this paper, we present the Configurable MASR model with Summary Vector (csvMASR), a novel architecture designed to enhance configurability. Our approach leverages adapters and introduces speech summary vector representations, inspired by conversational summary representations in speech diarization, to combine outputs from language-specific components at the utterance level. We also incorporate an auxiliary language classification loss to enhance configurability. Using data from 7 languages in the Multilingual Librispeech (MLS) dataset, csvMASR outperforms existing MASR models and reduces the word error rate (WER) from 10.33 % to 9.95 % when compared with the baseline. Additionally, csvMASR demonstrates superior performance in language classification and prompting tasks.

【11】 SONAR: A Synthetic AI-Audio Detection Framework~and Benchmark
标题: SONAR:合成人工智能音频检测框架~和基准
作者: Xiang Li, Pin-Yu Chen, Wenqi Wei
链接:点击下载PDF文件
摘要:使用生成人工智能(AI)技术的文本到语音(TTS)和语音转换(VC)的最新进展使得生成高质量和逼真的类人音频成为可能。这给区分人工智能合成语音与真实人类语音带来了重大挑战,并可能引发潜在的恶意滥用问题,如模仿和欺诈、传播错误信息、深度伪造和诈骗。然而,现有的人工智能合成音频检测技术并没有跟上步伐,并且在不同的数据集上往往表现出很差的泛化能力。在本文中,我们介绍了SONAR,一个人工智能音频检测框架和基准,旨在为区分尖端的人工智能合成的听觉内容提供全面的评估。SONAR包括一个来自9个不同音频合成平台的新型评估数据集,包括领先的TTS提供商和最先进的TTS模型。它是第一个在传统和基于基础模型的deepfake检测系统中统一基准测试AI音频检测的框架。通过大量的实验,我们揭示了现有检测方法的泛化局限性,并证明基础模型具有更强的泛化能力,这可以归因于它们的模型大小以及预训练数据的规模和质量。此外,我们探讨了Few-Shot微调在提高泛化能力方面的有效性和效率,突出了其在定制应用中的潜力,例如针对特定实体或个人的个性化检测系统。代码和数据集可在https: github.com Jessegator SONAR上获得。摘要:Recent advances in Text-to-Speech (TTS) and Voice-Conversion (VC) using generative Artificial Intelligence (AI) technology have made it possible to generate high-quality and realistic human-like audio. This introduces significant challenges to distinguishing AI-synthesized speech from the authentic human voice and could raise potential issues of misuse for malicious purposes such as impersonation and fraud, spreading misinformation, deepfakes, and scams. However, existing detection techniques for AI-synthesized audio have not kept pace and often exhibit poor generalization across diverse datasets. In this paper, we introduce SONAR, a synthetic AI-Audio Detection Framework and Benchmark, aiming to provide a comprehensive evaluation for distinguishing cutting-edge AI-synthesized auditory content. SONAR includes a novel evaluation dataset sourced from 9 diverse audio synthesis platforms, including leading TTS providers and state-of-the-art TTS models. It is the first framework to uniformly benchmark AI-audio detection across both traditional and foundation model-based deepfake detection systems. Through extensive experiments, we reveal the generalization limitations of existing detection methods and demonstrate that foundation models exhibit stronger generalization capabilities, which can be attributed to their model size and the scale and quality of pretraining data. Additionally, we explore the effectiveness and efficiency of few-shot fine-tuning in improving generalization, highlighting its potential for tailored applications, such as personalized detection systems for specific entities or individuals. Code and dataset are available at https: github.com Jessegator SONAR.

【12】 Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
标题: 使用混合H3-Conformer实现高效、稳健的长形式语音识别
作者: Tomoki Honda, Shinsuke Sakai, Tatsuya Kawahara
备注:Submitted to InterSpeech2024, Sample code is available at this https URL
链接:点击下载PDF文件
摘要:最近,Conformer在许多语音识别任务中取得了最先进的性能。然而,基于transformer的模型对于长形式的语音(例如演讲)表现出显著的恶化,因为自我注意机制随着输入长度的平方阶的计算而变得不可靠。为了解决这个问题,我们引入了一种状态空间模型,饥饿的饥饿河马(H3),以取代或补充多头自我注意(MHSA)。H3允许使用线性阶计算对长形式序列进行有效建模。在使用CSJ和LibriSpeech两个数据集的实验中,我们提出的H3-Conformer模型对长格式语音进行了高效和鲁棒的识别。此外,我们提出了一个混合的H3和MHSA,并表明,使用H3在较高层和MHSA在较低层提供了显着的改善在线识别。我们还研究了在所有层中并行使用H3和MHSA,从而获得最佳性能。摘要:Recently, Conformer has achieved state-of-the-art performance in many speech recognition tasks. However, the Transformer-based models show significant deterioration for long-form speech, such as lectures, because the self-attention mechanism becomes unreliable with the computation of the square order of the input length. To solve the problem, we incorporate a kind of state-space model, Hungry Hungry Hippos (H3), to replace or complement the multi-head self-attention (MHSA). H3 allows for efficient modeling of long-form sequences with a linear-order computation. In experiments using two datasets of CSJ and LibriSpeech, our proposed H3-Conformer model performs efficient and robust recognition of long-form speech. Moreover, we propose a hybrid of H3 and MHSA and show that using H3 in higher layers and MHSA in lower layers provides significant improvement in online recognition. We also investigate a parallel use of H3 and MHSA in all layers, resulting in the best performance.

【13】 The OCON model: an old but green solution for distributable supervised classification for acoustic monitoring in smart cities
标题: OCON模型:一种古老但绿色的解决方案,用于智能城市声学监测的分布式监督分类
作者: Stefano Giacomelli, Marco Giordano, Claudia Rinaldi
备注:Accepted at "IEEE 5th International Symposium on the Internet of Sounds, 30 Sep 2 Oct 2024, Erlangen, Germany"
Journal-ref:in Proceedings of the 5th IEEE International Symposium on the Internet of Sounds (IEEE IS2 2024, https:internetofsounds.netis2_2024)
链接:点击下载PDF文件
摘要:本文探讨了一个结构化的应用程序的一类方法和一类一网络模型的监督分类任务,侧重于元音音素分类和说话人识别的自动语音识别(ASR)域。在我们的案例研究中,ASR模型运行在专有的传感和闪电系统上,用于监测城市街道上的声学和空气污染。我们正式组合的伪神经架构搜索和超参数调整实验,使用一个明智的网格搜索方法,以实现分类精度相媲美,现在最复杂的架构,深入到说话人识别和能源效率方面。尽管它的简单性,我们的模型建议有一个很好的机会来概括的语言和扬声器性别的背景下,广泛适用于计算约束的情况下,证明了相关的统计和性能指标。我们的实验代码可以在我们的GitHub上公开访问。摘要:This paper explores a structured application of the One-Class approach and the One-Class-One-Network model for supervised classification tasks, focusing on vowel phonemes classification and speakers recognition for the Automatic Speech Recognition (ASR) domain. For our case-study, the ASR model runs on a proprietary sensing and lightning system, exploited to monitor acoustic and air pollution on urban streets. We formalize combinations of pseudo-Neural Architecture Search and Hyper-Parameters Tuning experiments, using an informed grid-search methodology, to achieve classification accuracy comparable to nowadays most complex architectures, delving into the speaker recognition and energy efficiency aspects. Despite its simplicity, our model proposal has a very good chance to generalize the language and speaker genders context for widespread applicability in computational constrained contexts, proved by relevant statistical and performance metrics. Our experiments code is openly accessible on our GitHub.

【14】 Cross-Lingual Query-by-Example Spoken Term Detection: A Transformer-Based Approach
标题: 跨语言逐例查询口语检测:基于转换器的方法
作者: Allahdadi Fatemeh, Mahdian Toroghi Rahil, Zareian Hassan
链接:点击下载PDF文件
摘要:通过实例查询的口语术语检测(QbE-STD)通常受到转录数据稀缺性和语言特异性的约束。本文介绍了一种新的,语言无关的QbE-STD模型,利用图像处理技术和Transformer架构。通过采用预训练的XLSR-53网络进行特征提取和Hough变换进行检测,我们的模型可以有效地搜索任何音频文件中的用户定义的口语术语。四种语言的实验结果表明,与基于CNN的基线相比,性能有显着提高(19-54%)。虽然与DTW相比处理时间有所改善,但准确性仍然较差。值得注意的是,我们的模型提供了准确计算目标音频内的查询词重复的优点。摘要:Query-by-example spoken term detection (QbE-STD) is typically constrained by transcribed data scarcity and language specificity. This paper introduces a novel, language-agnostic QbE-STD model leveraging image processing techniques and transformer architecture. By employing a pre-trained XLSR-53 network for feature extraction and a Hough transform for detection, our model effectively searches for user-defined spoken terms within any audio file. Experimental results across four languages demonstrate significant performance gains (19-54%) over a CNN-based baseline. While processing time is improved compared to DTW, accuracy remains inferior. Notably, our model offers the advantage of accurately counting query term repetitions within the target audio.

【15】 Reverb: Open-Source ASR and Diarization from Rev
标题: Reverb:开源ASB和Rev的日记
作者: Nishchal Bhandari, Danny Chen, Miguel Ángel del Río Fernández, Natalie Delworth, Jennifer Drexler Fox, Migüel Jetté, Quinten McNamara, Corey Miller, Ondřej Novotný, Ján Profant, Nan Qin, Martin Ratajczak, Jean-Philippe Robichaud
链接:点击下载PDF文件
摘要:今天,我们正在开源我们的核心语音识别和日志模型,用于非商业用途。我们正在为开发人员发布一个完整的生产管道,以及用于实验的精简研究模型。Rev希望这些版本将刺激快速发展的语音技术领域的研究和创新。今天发布的语音识别模型在各种长格式语音识别领域的性能优于所有现有的开源语音识别模型。摘要:Today, we are open-sourcing our core speech recognition and diarization models for non-commercial use. We are releasing both a full production pipeline for developers as well as pared-down research models for experimentation. Rev hopes that these releases will spur research and innovation in the fast-moving domain of voice technology. The speech recognition models released today outperform all existing open source speech recognition models across a variety of long-form speech recognition domains.

【16】 Did You Hear That? Introducing AADG: A Framework for Generating Benchmark Data in Audio Anomaly Detection
标题: 你听到了吗?介绍AADG:音频异常检测中生成基准数据的框架
作者: Ksheeraja Raghavan, Samiran Gode, Ankit Shah, Surabhi Raghavan, Wolfram Burgard, Bhiksha Raj, Rita Singh
备注:9 pages, under review
链接:点击下载PDF文件
摘要:我们介绍了一种新的,通用的音频生成框架,专门设计用于异常检测和定位。与主要关注工业和机器相关声音的现有数据集不同,我们的框架关注更广泛的环境,特别是在只有音频数据可用的现实世界场景中,例如视频衍生或电话音频。为了生成这样的数据,我们提出了一种受LLM-Modulo框架启发的新方法,该框架利用大型语言模型(LLM)作为世界模型来模拟真实世界的场景。该工具是模块化的,允许即插即用的方法。它首先使用LLM来预测合理的现实世界场景。LLM进一步提取组成声音,顺序和方式,这些应该合并,以创建连贯的整体。与LLM-Modulo框架非常相似,我们对每个输出阶段都进行了严格的验证,以确保生成数据的可靠性。使用该框架产生的数据作为异常检测应用程序的基准,可能会提高在音频数据上训练的模型的性能,特别是在处理分发外的情况下。因此,我们的贡献填补了音频异常检测资源的关键空白,并提供了一个可扩展的工具,用于生成多样化的,逼真的音频数据。摘要:We introduce a novel, general-purpose audio generation framework specifically designed for anomaly detection and localization. Unlike existing datasets that predominantly focus on industrial and machine-related sounds, our framework focuses a broader range of environments, particularly useful in real-world scenarios where only audio data are available, such as in video-derived or telephonic audio. To generate such data, we propose a new method inspired by the LLM-Modulo framework, which leverages large language models(LLMs) as world models to simulate such real-world scenarios. This tool is modular allowing a plug-and-play approach. It operates by first using LLMs to predict plausible real-world scenarios. An LLM further extracts the constituent sounds, the order and the way in which these should be merged to create coherent wholes. Much like the LLM-Modulo framework, we include rigorous verification of each output stage, ensuring the reliability of the generated data. The data produced using the framework serves as a benchmark for anomaly detection applications, potentially enhancing the performance of models trained on audio data, particularly in handling out-of-distribution cases. Our contributions thus fill a critical void in audio anomaly detection resources and provide a scalable tool for generating diverse, realistic audio data.

【17】 SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
标题: SONIQUE:使用未配对视听数据生成视频背景音乐
作者: Liqian Zhang, Magdalena Fuentes
链接:点击下载PDF文件
摘要:我们提出了SONIQUE,一个模型,用于生成定制的视频内容的背景音乐。与传统的视频到音乐生成方法不同,它严重依赖于配对的视听数据集,SONIQUE利用未配对的数据,将免版税音乐和独立的视频源相结合。通过利用大型语言模型(LLM)进行视频理解并将视觉描述转换为音乐标签,以及基于U-Net的条件扩散模型,SONIQUE实现了可定制的音乐生成。用户可以控制音乐的特定方面,如乐器、流派、节奏和旋律,确保生成的输出符合他们的创意愿景。SONIQUE是开源的,有一个在线演示。摘要:We present SONIQUE, a model for generating background music tailored to video content. Unlike traditional video-to-music generation approaches, which rely heavily on paired audio-visual datasets, SONIQUE leverages unpaired data, combining royalty-free music and independent video sources. By utilizing large language models (LLMs) for video understanding and converting visual descriptions into musical tags, alongside a U-Net-based conditional diffusion model, SONIQUE enables customizable music generation. Users can control specific aspects of the music, such as instruments, genres, tempo, and melodies, ensuring the generated output fits their creative vision. SONIQUE is open-source, with a demo available online.

【18】 SOI: Scaling Down Computational Complexity by Estimating Partial States of the Model
标题: SIM:通过估计模型的部分状态来降低计算复杂性
作者: Grzegorz Stefański, Paweł Daniluk, Artur Szumaczuk, Jakub Tkaczuk
备注:NeurIPS 2024
链接:点击下载PDF文件
摘要:消费电子产品过去一直遵循摩尔定律所描述的小型化趋势。尽管微控制器单元(MCU)的处理能力有所提高,但用于最小家电的MCU仍然无法运行即使是中等规模的最先进的人工神经网络(ANN),特别是在时间敏感的情况下。在这项工作中,我们提出了一种新的方法称为分散在线推理(SOI),旨在降低人工神经网络的计算复杂度。SOI利用时间序列数据和模型预测的连续性和季节性,实现外推以提高处理速度,特别是在更深层。通过应用压缩,SOI生成ANN的更一般的内部部分状态,允许在每次推理时跳过完整的模型重新计算。摘要:Consumer electronics used to follow the miniaturization trend described by Moore's Law. Despite increased processing power in Microcontroller Units (MCUs), MCUs used in the smallest appliances are still not capable of running even moderately big, state-of-the-art artificial neural networks (ANNs) especially in time-sensitive scenarios. In this work, we present a novel method called Scattered Online Inference (SOI) that aims to reduce the computational complexity of ANNs. SOI leverages the continuity and seasonality of time-series data and model predictions, enabling extrapolation for processing speed improvements, particularly in deeper layers. By applying compression, SOI generates more general inner partial states of ANN, allowing skipping full model recalculation at each inference.

【19】 Self-Powered LLM Modality Expansion for Large Speech-Text Models
标题: 针对大型语音文本模型的自供电LLM情态扩展
作者: Tengfei Yu, Xuebo Liu, Zhiyi Hou, Liang Ding, Dacheng Tao, Min Zhang
备注:Accepted to EMNLP 2024
链接:点击下载PDF文件
摘要:大型语言模型(LLM)在不同的任务中表现出显着的性能,这表明它们有潜力通过集成语音功能扩展到大型语音-文本模型(LSM)。虽然统一的语音文本预训练和多模态数据预调整提供了相当大的好处,但这些方法通常需要大量的资源需求,并且往往会过度适应特定的任务。本研究旨在通过解决vanilla指令调整的局限性来改进LSM训练中语音数据集的使用。我们探讨了LSMs内的解释,以下动态,确定一个关键的问题,称为语音锚偏置LSMs过度依赖语音输入的倾向,错误地解释整个语音模态的指令,从而忽略了文本的指示。为了抵消这种偏见,我们引入了一个自供电的LSM,利用模型本身生成的增强自动语音识别数据进行更有效的指令调整。我们在一系列基于语音的任务中的实验表明,自供电LSM减轻了语音锚偏置,提高了LSM中语音和文本模态的融合。数据、代码和脚本可在https: github.com ytf-philp Self-powered-LSM上免费获得。摘要:Large language models (LLMs) exhibit remarkable performance across diverse tasks, indicating their potential for expansion into large speech-text models (LSMs) by integrating speech capabilities. Although unified speech-text pre-training and multimodal data instruction-tuning offer considerable benefits, these methods generally entail significant resource demands and tend to overfit specific tasks. This study aims to refine the use of speech datasets for LSM training by addressing the limitations of vanilla instruction tuning. We explore the instruction-following dynamics within LSMs, identifying a critical issue termed speech anchor bias-a tendency for LSMs to over-rely on speech inputs, mistakenly interpreting the entire speech modality as directives, thereby neglecting textual instructions. To counteract this bias, we introduce a self-powered LSM that leverages augmented automatic speech recognition data generated by the model itself for more effective instruction tuning. Our experiments across a range of speech-based tasks demonstrate that self-powered LSM mitigates speech anchor bias and improves the fusion of speech and text modalities in LSMs. Data, code and scripts are freely available at https: github.com ytf-philp Self-powered-LSM.

【20】 People are poorly equipped to detect AI-powered voice clones
标题: 人们检测人工智能语音克隆的能力很差
作者: Sarah Barrington, Hany Farid
链接:点击下载PDF文件
摘要:随着生成式人工智能继续其弹道轨迹,从文本到音频,图像和视频生成的一切都在模仿人类生成的内容方面继续改进。通过一系列的感知研究,我们报告了AI生成的声音在身份匹配和自然性方面的真实性。我们发现人类参与者无法可靠地识别人工智能生成的语音的短录音(不到20秒)。具体来说,参与者在80%的时间里将人工智能语音的身份误认为是其真实的对应物,而只有60%的时间将语音正确识别为人工智能生成的。在所有情况下,性能是独立的人口统计的发言者或听众。摘要:As generative AI continues its ballistic trajectory, everything from text to audio, image, and video generation continues to improve in mimicking human-generated content. Through a series of perceptual studies, we report on the realism of AI-generated voices in terms of identity matching and naturalness. We find human participants cannot reliably identify short recordings (less than 20 seconds) of AI-generated voices. Specifically, participants mistook the identity of an AI-voice for its real counterpart 80% of the time, and correctly identified a voice as AI-generated only 60% of the time. In all cases, performance is independent of the demographics of the speaker or listener.

【21】 Efficient Streaming LLM for Speech Recognition
标题: 用于语音识别的高效流媒体LLM
作者: Junteng Jia, Gil Keren, Wei Zhou, Egor Lakomkin, Xiaohui Zhang, Chunyang Wu, Frank Seide, Jay Mahadeokar, Ozlem Kalinli
链接:点击下载PDF文件
摘要:最近的工作表明,用音频编码提示大型语言模型可以解锁语音识别功能。然而,现有的技术不能有效地扩展,特别是在处理长格式流音频输入时--它们不仅在训练期间看到的音频长度之外外推得很差,而且由于注意力的二次成本,它们在计算上效率低下。 在这项工作中,我们介绍了SpeechLLM-XL,一个线性缩放解码器的流语音识别模型。我们使用有限的注意力窗口来处理可配置块中的音频以减少计算,并且每个音频块的文本令牌自回归地生成,直到预测EOS。在训练期间,使用从编码器输出估计的CTC强制对齐将转录本分割成块。具有1.28秒块大小的SpeechLLM-XL在LibriSpeech测试clean other上实现了2.7% 6.7%的WER,并且它在比训练话语长10倍的长形式话语上没有显示出质量下降。摘要:Recent works have shown that prompting large language models with audio encodings can unlock speech recognition capabilities. However, existing techniques do not scale efficiently, especially while handling long form streaming audio inputs -- not only do they extrapolate poorly beyond the audio length seen during training, but they are also computationally inefficient due to the quadratic cost of attention. In this work, we introduce SpeechLLM-XL, a linear scaling decoder-only model for streaming speech recognition. We process audios in configurable chunks using limited attention window for reduced computation, and the text tokens for each audio chunk are generated auto-regressively until an EOS is predicted. During training, the transcript is segmented into chunks, using a CTC forced alignment estimated from encoder output. SpeechLLM-XL with 1.28 seconds chunk size achieves 2.7% 6.7% WER on LibriSpeech test clean other, and it shows no quality degradation on long form utterances 10x longer than the training utterances.

【22】 Recent Advances in Speech Language Models: A Survey
标题: 言语语言模型的最新进展:调查
作者: Wenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng, Guangyan Zhang, Qichao Wang, Yiwen Guo, Irwin King
备注:Work in progress
链接:点击下载PDF文件
摘要:大型语言模型(LLM)最近获得了极大的关注,主要是因为它们在基于文本的交互中的能力。然而,自然的人类交互通常依赖于语音,因此需要转向基于语音的模型。实现这一目标的一种简单方法涉及“自动语音识别(ASR)+ LLM +文本到语音(TTS)”的管道,其中输入语音被转录为文本,由LLM处理,然后转换回语音。尽管是直接的,这种方法遭受固有的局限性,如模态转换过程中的信息丢失和三个阶段的误差积累。为了解决这些问题,语音语言模型(SpeechLM)-生成语音而不转换文本的端到端模型-已成为一种有前途的替代方案。这篇调查论文首次全面概述了构建SpeechLM的最新方法,详细介绍了其架构的关键组件以及对其开发不可或缺的各种培训食谱。此外,我们系统地调查的各种能力的SpeechLMs,分类SpeechLMs的评价指标,并讨论在这个快速发展的领域的挑战和未来的研究方向。摘要:Large Language Models (LLMs) have recently garnered significant attention, primarily for their capabilities in text-based interactions. However, natural human interaction often relies on speech, necessitating a shift towards voice-based models. A straightforward approach to achieve this involves a pipeline of Automatic Speech Recognition (ASR) + LLM + Text-to-Speech (TTS)", where input speech is transcribed to text, processed by an LLM, and then converted back to speech. Despite being straightforward, this method suffers from inherent limitations, such as information loss during modality conversion and error accumulation across the three stages. To address these issues, Speech Language Models (SpeechLMs) -- end-to-end models that generate speech without converting from text -- have emerged as a promising alternative. This survey paper provides the first comprehensive overview of recent methodologies for constructing SpeechLMs, detailing the key components of their architecture and the various training recipes integral to their development. Additionally, we systematically survey the various capabilities of SpeechLMs, categorize the evaluation metrics for SpeechLMs, and discuss the challenges and future research directions in this rapidly evolving field.

【23】 Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
标题: 使用离散单元与从可控口音TTC合成的并行数据进行口音转换
作者: Tuan Nam Nguyen, Ngoc Quan Pham, Alexander Waibel
备注:Accepted at Syndata4genAI
链接:点击下载PDF文件
摘要:口音转换(AC)的目标是转换语音口音,同时保留内容和说话人身份。以前的方法要么在推理过程中需要参考话语,不能很好地保留说话者身份,要么使用只能针对每个非母语口音进行训练的一对一系统。本文提出了一个很有前途的AC模型,可以将许多口音转换为本地,以克服这些问题。我们的方法利用从对母语的自监督表示进行聚类中得出的离散单元作为口音转换的中间目标。利用多说话人文本到语音合成,它将这些离散表示转换回本地语音,同时保留说话人身份。此外,我们开发了一种有效的数据增强方法来训练系统,而不需要大量的非本地资源。我们的系统被证明可以提高非母语人士的流利性,听起来像一个本地口音,并保持原来的发言人身份很好。摘要:The goal of accent conversion (AC) is to convert speech accents while preserving content and speaker identity. Previous methods either required reference utterances during inference, did not preserve speaker identity well, or used one-to-one systems that could only be trained for each non-native accent. This paper presents a promising AC model that can convert many accents into native to overcome these issues. Our approach utilizes discrete units, derived from clustering self-supervised representations of native speech, as an intermediary target for accent conversion. Leveraging multi-speaker text-to-speech synthesis, it transforms these discrete representations back into native speech while retaining the speaker identity. Additionally, we develop an efficient data augmentation method to train the system without demanding a lot of non-native resources. Our system is proved to improve non-native speaker fluency, sound like a native accent, and preserve original speaker identity well.

【24】 FluentEditor+: Text-based Speech Editing by Modeling Local Hierarchical Acoustic Smoothness and Global Prosody Consistency
标题: FluentEditor+:通过建模局部分层声学平滑度和全局韵律一致性来进行基于文本的语音编辑
作者: Rui Liu, Jiatian Xi, Ziyue Jiang, Haizhou Li
备注:Work in progress
链接:点击下载PDF文件
摘要:基于文本的语音编辑(TSE)允许用户通过编辑相应的文本并执行剪切、复制和粘贴等操作来修改语音,以生成更新的音频,而无需直接更改原始录音。基于文本的语音编辑(TSE)允许用户通过编辑相应的文本并执行剪切、复制和粘贴等操作来修改语音,以生成更新的音频,而无需直接更改原始录音。虽然目前的TSE技术专注于最大限度地减少编辑片段内生成的语音和参考目标之间的差异,但它们往往忽视了在原始话语的上下文中保持局部和全局流畅性的重要性。此外,将编辑的片段与音频的未更改部分无缝集成仍然具有挑战性,通常需要文本到语音(TTS)系统的支持。本文介绍了一种新的方法,FluentEditor$ tiny +$,旨在克服这些限制。FluentEditor$ tiny +$采用高级特征提取技术来捕获声学和韵律特征,确保编辑区域和未编辑区域之间的流畅过渡。该模型确保了分段声学平滑性和全局韵律一致性,允许语音的无缝拼接,同时保持输出的连贯性和自然性。在VCTK和LibriTTS数据集上进行的大量实验表明,FluentEditor$ tiny +$在流畅性和韵律方面都优于现有的基于TTS的方法,包括Editspeech,Campnet,$A^3T$ FluentSpeech和Fluenteditor。消融研究进一步强调了每个模块对系统整体有效性的贡献。摘要:Text-based speech editing (TSE) allows users to modify speech by editing the corresponding text and performing operations such as cutting, copying, and pasting to generate updated audio without altering the original recording directly. Text-based speech editing (TSE) allows users to modify speech by editing the corresponding text and performing operations such as cutting, copying, and pasting to generate updated audio without altering the original recording directly. While current TSE techniques focus on minimizing discrepancies between generated speech and reference targets within edited segments, they often neglect the importance of maintaining both local and global fluency in the context of the original discourse. Additionally, seamlessly integrating edited segments with unaltered portions of the audio remains challenging, typically requiring support from text-to-speech (TTS) systems. This paper introduces a novel approach, FluentEditor$ tiny +$, designed to overcome these limitations. FluentEditor$ tiny +$ employs advanced feature extraction techniques to capture both acoustic and prosodic characteristics, ensuring fluent transitions between edited and unedited regions. The model ensures segmental acoustic smoothness and global prosody consistency, allowing seamless splicing of speech while preserving the coherence and naturalness of the output. Extensive experiments on the VCTK and LibriTTS datasets show that FluentEditor$ tiny +$ surpasses existing TTS-based methods, including Editspeech, Campnet, $A^3T$ FluentSpeech, and Fluenteditor, in both fluency and prosody. Ablation studies further highlight the contributions of each module to the overall effectiveness of the system.

【25】 A quest through interconnected datasets: lessons from highly-cited ICASSP papers
标题: 探索相互关联的数据集:备受引用的ICASP论文的教训
作者: Cynthia C. S. Liem, Doğa Taşcılar, Andrew M. Demetriou
备注:in Proceedings of the 21st International Conference on Content-based Multimedia Indexing, September 18-20 2024, Reykjavik, Iceland
链接:点击下载PDF文件
摘要:随着音频机器学习成果部署在具有社会影响力的应用程序中,重要的是要了解所使用数据的质量和来源。注意到在应用机器学习领域的学术出版中,明确这种意义并不是微不足道的奖励,也不包括在典型的应用机器学习课程中,我们在声学,语音和信号处理国际会议(ICASSP)上提出了一项与前5名引用论文相关的数据集使用研究。在这方面,我们对使用的数据集的来源进行了彻底的深度优先分析,通常导致搜索必须超出官方论文中报道的内容,并最终导致不清楚或纠缠的来源。特别是在当前对更大的、可能是生成性的人工智能模型的需求中,人们越来越意识到需要对数据来源进行问责。因此,我们呼吁社区不仅要专注于工程更大的模型,而且要为明确构建这些模型的基础创造更多的空间和奖励。摘要:As audio machine learning outcomes are deployed in societally impactful applications, it is important to have a sense of the quality and origins of the data used. Noticing that being explicit about this sense is not trivially rewarded in academic publishing in applied machine learning domains, and neither is included in typical applied machine learning curricula, we present a study into dataset usage connected to the top-5 cited papers at the International Conference on Acoustics, Speech, and Signal Processing (ICASSP). In this, we conduct thorough depth-first analyses towards origins of used datasets, often leading to searches that had to go beyond what was reported in official papers, and ending into unclear or entangled origins. Especially in the current pull towards larger, and possibly generative AI models, awareness of the need for accountability on data provenance is increasing. With this, we call on the community to not only focus on engineering larger models, but create more room and reward for explicitizing the foundations on which such models should be built.

【26】 Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
标题: 使用旋律和文本编辑音乐:使用Control Net进行扩散Transformer
作者: Siyuan Hou, Shansong Liu, Ruibin Yuan, Wei Xue, Ying Shan, Mangsuo Zhao, Chao Zhang
备注:5 pages, 1 figure
链接:点击下载PDF文件
摘要:尽管在可控音乐生成和编辑方面取得了重大进展,但由于使用Mel频谱图表示和基于UNet的模型结构,所生成音乐的质量和长度仍然存在挑战。为了解决这些限制,我们提出了一种新的方法,使用扩散Transformer(DiT)增加了一个额外的控制分支,使用ControlNet。这允许通过文本和旋律提示控制长格式和可变长度的音乐生成和编辑。为了更精确和细粒度的旋律控制,我们引入了一种新颖的top-$k$ constant-Q Transform表示作为旋律提示,与以前的表示相比减少了模糊性(例如,色度),特别是对于具有多个音轨或宽范围的音高值的音乐。为了有效地平衡文本和旋律提示的控制信号,我们采用了一种课程学习策略,逐步掩盖旋律提示,从而使训练过程更加稳定。实验已经进行了文本到音乐的生成和音乐风格的转移任务,使用开源的乐器录音数据。结果表明,通过扩展StableAudio(一种预先训练的文本控制DiT模型),我们的方法可以实现卓越的旋律控制编辑,同时保持良好的文本到音乐生成性能。这些结果在基于文本的生成和用于编辑的旋律保存方面都优于强大的MusicGen基线。音频示例可以在https: stable-audio-control.github.io web 上找到。摘要:Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structures. To address these limitations, we propose a novel approach using a Diffusion Transformer (DiT) augmented with an additional control branch using ControlNet. This allows for long-form and variable-length music generation and editing controlled by text and melody prompts. For more precise and fine-grained melody control, we introduce a novel top-$k$ constant-Q Transform representation as the melody prompt, reducing ambiguity compared to previous representations (e.g., chroma), particularly for music with multiple tracks or a wide range of pitch values. To effectively balance the control signals from text and melody prompts, we adopt a curriculum learning strategy that progressively masks the melody prompt, resulting in a more stable training process. Experiments have been performed on text-to-music generation and music-style transfer tasks using open-source instrumental recording data. The results demonstrate that by extending StableAudio, a pre-trained text-controlled DiT model, our approach enables superior melody-controlled editing while retaining good text-to-music generation performance. These results outperform a strong MusicGen baseline in terms of both text-based generation and melody preservation for editing. Audio examples can be found at https: stable-audio-control.github.io web .

【27】 CR-CTC: Consistency regularization on CTC for improved speech recognition
标题: CR-ctc:对CTC进行一致性规范化,以改进语音识别
作者: Zengwei Yao, Wei Kang, Xiaoyu Yang, Fangjun Kuang, Liyong Guo, Han Zhu, Zengrui Jin, Zhaoqing Li, Long Lin, Daniel Povey
链接:点击下载PDF文件
摘要:连接主义时态分类(CTC)是一种广泛应用于自动语音识别(ASR)的方法,以其简单性和计算效率而闻名。然而,它往往低于识别性能相比,传感器或系统结合CTC和基于注意力的编码器-解码器(CTC AED)。在这项工作中,我们提出了一致性正则化CTC(CR-CTC),它强制执行从输入语音梅尔频谱图的不同增强视图获得的两个CTC分布之间的一致性。我们从三个方面深入研究了它的基本行为:1)它在处理不同增强视图的随机子模型对之间进行自蒸馏; 2)它通过对时间掩蔽区域内的位置进行掩蔽预测来学习上下文表示,特别是当我们增加时间掩蔽量时; 3)抑制了CTC分布的极端峰值,从而减少了过拟合,提高了泛化能力。在LibriSpeech、Aishell-1和GigaSpeech数据集上进行的大量实验证明了我们的CR-CTC的有效性,其性能与换能器和CTC AED相当,甚至略好于换能器和CTC AED。摘要:Connectionist Temporal Classification (CTC) is a widely used method for automatic speech recognition (ASR), renowned for its simplicity and computational efficiency. However, it often falls short in recognition performance compared to transducer or systems combining CTC and attention-based encoder-decoder (CTC AED). In this work, we propose the Consistency-Regularized CTC (CR-CTC), which enforces consistency between two CTC distributions obtained from different augmented views of the input speech mel-spectrogram. We provide in-depth insights into its essential behaviors from three perspectives: 1) it conducts self-distillation between random pairs of sub-models that process different augmented views; 2) it learns contextual representation through masked prediction for positions within time-masked regions, especially when we increase the amount of time masking; 3) it suppresses the extremely peaky CTC distributions, thereby reducing overfitting and improving the generalization ability. Extensive experiments on LibriSpeech, Aishell-1, and GigaSpeech datasets demonstrate the effectiveness of our CR-CTC, which achieves performance comparable to, or even slightly better than, that of transducer and CTC AED.

【28】 A decade of DCASE: Achievements, practices, evaluations and future challenges
标题: DUSE十年:成就、实践、评估和未来挑战
作者: Annamaria Mesaros, Romain Serizel, Toni Heittola, Tuomas Virtanen, Mark D. Plumbley
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:本文简要介绍了声场景和事件的检测与分类(DCASE)挑战赛的历史和发展,研讨会,研究领域和研究团体。DCASE创建于2013年,作为一项数据评估挑战,已成为音频和声学信号处理领域的主要研究课题。它的成功来自于多种因素:挑战提供了每年更新的各种任务;讲习班提供了传播相关工作的渠道,吸引了年轻和充满活力的社区。与此同时,DCASE也面临着自己的挑战,不断发展并扩展到不同的领域。DCASE的核心原则之一是开放科学和可重复性:公开可用的数据集,基线系统,技术报告和研讨会出版物。虽然DCASE挑战和研讨会独立于IEEE SPS,但该挑战每年都会得到AASP TC的认可,DCASE社区为ICASSP旗舰会议和SPS在其许多活动中的成功做出了重大贡献。摘要:This paper introduces briefly the history and growth of the Detection and Classification of Acoustic Scenes and Events (DCASE) challenge, workshop, research area and research community. Created in 2013 as a data evaluation challenge, DCASE has become a major research topic in the Audio and Acoustic Signal Processing area. Its success comes from a combination of factors: the challenge offers a large variety of tasks that are renewed each year; and the workshop offers a channel for dissemination of related work, engaging a young and dynamic community. At the same time, DCASE faces its own challenges, growing and expanding to different areas. One of the core principles of DCASE is open science and reproducibility: publicly available datasets, baseline systems, technical reports and workshop publications. While the DCASE challenge and workshop are independent of IEEE SPS, the challenge receives annual endorsement from the AASP TC, and the DCASE community contributes significantly to the ICASSP flagship conference and the success of SPS in many of its activities.

【29】 Towards Ultra-Low-Power Neuromorphic Speech Enhancement with Spiking-FullSubNet
标题: 利用尖峰全SubNet实现超低功耗神经形态语音增强
作者: Xiang Hao, Chenxiang Ma, Qu Yang, Jibin Wu, Kay Chen Tan
备注:under review
链接:点击下载PDF文件
摘要:语音增强对于提高各种音频设备中的语音清晰度和质量至关重要。近年来,基于深度学习的方法显着提高了语音增强性能,但它们通常具有很高的计算成本,这对于耳机和助听器等大量边缘设备来说是令人望而却步的。本文提出了一种基于脑激励脉冲神经网络(SNN)的超低功耗语音增强系统,称为Spiking-FullSubNet。Spiking-FullSubNet采用全波段和子波段融合的方法,有效地捕获全局和局部光谱信息。为了提高计算昂贵的子带建模的效率,我们引入了一个频率划分方法的灵感来自人类外周听觉系统的灵敏度分布。此外,我们引入了一种新的尖峰神经元模型,可以动态控制输入信息的整合和遗忘,增强SNN的多尺度时间处理能力,这是语音去噪的关键。在最近的Intel Neuromorphic Deep Noise Suppression(N-DNS)Challenge数据集上进行的实验表明,Spiking-FullSubNet在语音质量和能效指标方面都大大超过了最先进的方法。值得一提的是,我们的系统赢得了英特尔N-DNS挑战赛(英语:Intel N-DNS Challenge)的冠军,为边缘的超低功耗语音增强提供了无数机会。我们的源代码和模型检查点可在https: github.com haoxiangsnr spiking-fullsubnet上公开获取。摘要:Speech enhancement is critical for improving speech intelligibility and quality in various audio devices. In recent years, deep learning-based methods have significantly improved speech enhancement performance, but they often come with a high computational cost, which is prohibitive for a large number of edge devices, such as headsets and hearing aids. This work proposes an ultra-low-power speech enhancement system based on the brain-inspired spiking neural network (SNN) called Spiking-FullSubNet. Spiking-FullSubNet follows a full-band and sub-band fusioned approach to effectively capture both global and local spectral information. To enhance the efficiency of computationally expensive sub-band modeling, we introduce a frequency partitioning method inspired by the sensitivity profile of the human peripheral auditory system. Furthermore, we introduce a novel spiking neuron model that can dynamically control the input information integration and forgetting, enhancing the multi-scale temporal processing capability of SNN, which is critical for speech denoising. Experiments conducted on the recent Intel Neuromorphic Deep Noise Suppression (N-DNS) Challenge dataset show that the Spiking-FullSubNet surpasses state-of-the-art methods by large margins in terms of both speech quality and energy efficiency metrics. Notably, our system won the championship of the Intel N-DNS Challenge (Algorithmic Track), opening up a myriad of opportunities for ultra-low-power speech enhancement at the edge. Our source code and model checkpoints are publicly available at https: github.com haoxiangsnr spiking-fullsubnet.

【30】 Example-Based Framework for Perceptually Guided Audio Texture Generation
标题: 基于示例的感知引导音频纹理生成框架
作者: Purnima Kamath, Chitralekha Gupta, Lonce Wyse, Suranga Nanayakkara
备注:Accepted for publication at IEEE Transactions on Audio, Speech and Language Processing
链接:点击下载PDF文件
摘要:使用StyleGAN的可控生成通常是通过使用标记数据训练模型来实现的。然而,对于音频纹理,目前缺乏大型语义标记数据集。因此,为了控制生成,我们开发了一种方法,用于在没有此类标记数据集的情况下对无条件训练的StyleGAN进行语义控制。在本文中,我们提出了一个基于实例的框架,以确定指导向量的音频纹理生成的基础上,用户定义的语义属性。我们的方法利用了无条件训练的StyleGAN的语义分解潜在空间。通过使用一些合成的例子来指示语义属性的存在或不存在,我们推断StyleGAN的潜在空间中的指导向量,以在生成过程中控制该属性。我们的研究结果表明,我们的框架可以找到用户定义的和感知相关的指导矢量可控生成音频纹理。此外,我们展示了我们的框架的应用程序的其他任务,如选择性语义属性转移。摘要:Controllable generation using StyleGANs is usually achieved by training the model using labeled data. For audio textures, however, there is currently a lack of large semantically labeled datasets. Therefore, to control generation, we develop a method for semantic control over an unconditionally trained StyleGAN in the absence of such labeled datasets. In this paper, we propose an example-based framework to determine guidance vectors for audio texture generation based on user-defined semantic attributes. Our approach leverages the semantically disentangled latent space of an unconditionally trained StyleGAN. By using a few synthetic examples to indicate the presence or absence of a semantic attribute, we infer the guidance vectors in the latent space of the StyleGAN to control that attribute during generation. Our results show that our framework can find user-defined and perceptually relevant guidance vectors for controllable generation for audio textures. Furthermore, we demonstrate an application of our framework to other tasks, such as selective semantic attribute transfer.

【31】 Towards Controllable Audio Texture Morphing
标题: 迈向可控音频纹理变形
作者: Chitralekha Gupta, Purnima Kamath, Yize Wei, Zhuoyao Li, Suranga Nanayakkara, Lonce Wyse
备注:accepted to ICASSP 2023
链接:点击下载PDF文件
摘要:在本文中,我们提出了一种数据驱动的方法来训练一个生成对抗网络(GAN)的“软标签”的条件下,从音频分类器的倒数第二层的音频纹理类的目标集训练。我们证明,这样的条件或控制向量之间的插值提供平滑变形之间生成的音频纹理,并显示类似或更好的音频纹理变形能力相比,国家的最先进的方法。所提出的方法导致在一个组织良好的潜在空间,产生新的音频输出,同时保持与语义的条件参数一致。这是朝着设计具有定制控件的生成音频模型的通用数据驱动方法迈出的一步,该定制控件能够遍历分布外区域以进行新颖的声音合成。摘要:In this paper, we propose a data-driven approach to train a Generative Adversarial Network (GAN) conditioned on "soft-labels" distilled from the penultimate layer of an audio classifier trained on a target set of audio texture classes. We demonstrate that interpolation between such conditions or control vectors provides smooth morphing between the generated audio textures, and shows similar or better audio texture morphing capability compared to the state-of-the-art methods. The proposed approach results in a well-organized latent space that generates novel audio outputs while remaining consistent with the semantics of the conditioning parameters. This is a step towards a general data-driven approach to designing generative audio models with customized controls capable of traversing out-of-distribution regions for novel sound synthesis.

【32】 Jointly Fine-Tuning "BERT-like" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
作者: Shamane Siriwardhana, Andrew Reis, Rivindu Weerasekera, Suranga Nanayakkara
备注:Accepted to INTERSPEECH 2020
链接:点击下载PDF文件
摘要:语音多模态情感识别是情感计算的一个重要研究领域。融合多个数据模态和学习具有有限数量的标记数据的表示是一项具有挑战性的任务。在本文中,我们将探讨使用特定模态的“BERT样”预训练的自监督学习(SSL)架构来表示语音和文本模态的多模态语音情感识别任务。通过对三个公开可用的数据集(IEMOCAP,CMU-MOSEI和CMU-MOSI)进行实验,我们表明,联合微调“类BERT”SSL架构实现了最先进的(SOTA)结果。我们还评估了两种方法的融合语音和文本模态,并表明一个简单的融合机制可以优于更复杂的SSL模型时,具有类似的建筑属性BERT。摘要:Multimodal emotion recognition from speech is an important area in affective computing. Fusing multiple data modalities and learning representations with limited amounts of labeled data is a challenging task. In this paper, we explore the use of modality-specific "BERT-like" pretrained Self Supervised Learning (SSL) architectures to represent both speech and text modalities for the task of multimodal speech emotion recognition. By conducting experiments on three publicly available datasets (IEMOCAP, CMU-MOSEI, and CMU-MOSI), we show that jointly fine-tuning "BERT-like" SSL architectures achieve state-of-the-art (SOTA) results. We also evaluate two methods of fusing speech and text modalities and show that a simple fusion mechanism can outperform more complex ones when using SSL models that have similar architectural properties to BERT.


eess.AS音频处理
【1】 Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
标题: 使用旋律和文本编辑音乐:使用Control Net进行扩散Transformer
作者: Siyuan Hou, Shansong Liu, Ruibin Yuan, Wei Xue, Ying Shan, Mangsuo Zhao, Chao Zhang
备注:5 pages, 1 figure
链接:点击下载PDF文件
摘要:尽管在可控音乐生成和编辑方面取得了重大进展,但由于使用Mel频谱图表示和基于UNet的模型结构,所生成音乐的质量和长度仍然存在挑战。为了解决这些限制,我们提出了一种新的方法,使用扩散Transformer(DiT)增加了一个额外的控制分支,使用ControlNet。这允许通过文本和旋律提示控制长格式和可变长度的音乐生成和编辑。为了更精确和细粒度的旋律控制,我们引入了一种新颖的top-$k$ constant-Q Transform表示作为旋律提示,与以前的表示相比减少了模糊性(例如,色度),特别是对于具有多个音轨或宽范围的音高值的音乐。为了有效地平衡文本和旋律提示的控制信号,我们采用了一种课程学习策略,逐步掩盖旋律提示,从而使训练过程更加稳定。实验已经进行了文本到音乐的生成和音乐风格的转移任务,使用开源的乐器录音数据。结果表明,通过扩展StableAudio(一种预先训练的文本控制DiT模型),我们的方法可以实现卓越的旋律控制编辑,同时保持良好的文本到音乐生成性能。这些结果在基于文本的生成和用于编辑的旋律保存方面都优于强大的MusicGen基线。音频示例可以在https: stable-audio-control.github.io web 上找到。摘要:Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structures. To address these limitations, we propose a novel approach using a Diffusion Transformer (DiT) augmented with an additional control branch using ControlNet. This allows for long-form and variable-length music generation and editing controlled by text and melody prompts. For more precise and fine-grained melody control, we introduce a novel top-$k$ constant-Q Transform representation as the melody prompt, reducing ambiguity compared to previous representations (e.g., chroma), particularly for music with multiple tracks or a wide range of pitch values. To effectively balance the control signals from text and melody prompts, we adopt a curriculum learning strategy that progressively masks the melody prompt, resulting in a more stable training process. Experiments have been performed on text-to-music generation and music-style transfer tasks using open-source instrumental recording data. The results demonstrate that by extending StableAudio, a pre-trained text-controlled DiT model, our approach enables superior melody-controlled editing while retaining good text-to-music generation performance. These results outperform a strong MusicGen baseline in terms of both text-based generation and melody preservation for editing. Audio examples can be found at https: stable-audio-control.github.io web .

【2】 CR-CTC: Consistency regularization on CTC for improved speech recognition
标题: CR-ctc:对CTC进行一致性规范化,以改进语音识别
作者: Zengwei Yao, Wei Kang, Xiaoyu Yang, Fangjun Kuang, Liyong Guo, Han Zhu, Zengrui Jin, Zhaoqing Li, Long Lin, Daniel Povey
链接:点击下载PDF文件
摘要:连接主义时态分类(CTC)是一种广泛应用于自动语音识别(ASR)的方法,以其简单性和计算效率而闻名。然而,它往往低于识别性能相比,传感器或系统结合CTC和基于注意力的编码器-解码器(CTC AED)。在这项工作中,我们提出了一致性正则化CTC(CR-CTC),它强制执行从输入语音梅尔频谱图的不同增强视图获得的两个CTC分布之间的一致性。我们从三个方面深入研究了它的基本行为:1)它在处理不同增强视图的随机子模型对之间进行自蒸馏; 2)它通过对时间掩蔽区域内的位置进行掩蔽预测来学习上下文表示,特别是当我们增加时间掩蔽量时; 3)抑制极峰值的CTC分布,从而减少过拟合并提高泛化能力。在LibriSpeech、Aishell-1和GigaSpeech数据集上进行的大量实验证明了我们的CR-CTC的有效性,其性能与换能器和CTC AED相当,甚至略好于换能器和CTC AED。摘要:Connectionist Temporal Classification (CTC) is a widely used method for automatic speech recognition (ASR), renowned for its simplicity and computational efficiency. However, it often falls short in recognition performance compared to transducer or systems combining CTC and attention-based encoder-decoder (CTC AED). In this work, we propose the Consistency-Regularized CTC (CR-CTC), which enforces consistency between two CTC distributions obtained from different augmented views of the input speech mel-spectrogram. We provide in-depth insights into its essential behaviors from three perspectives: 1) it conducts self-distillation between random pairs of sub-models that process different augmented views; 2) it learns contextual representation through masked prediction for positions within time-masked regions, especially when we increase the amount of time masking; 3) it suppresses the extremely peaky CTC distributions, thereby reducing overfitting and improving the generalization ability. Extensive experiments on LibriSpeech, Aishell-1, and GigaSpeech datasets demonstrate the effectiveness of our CR-CTC, which achieves performance comparable to, or even slightly better than, that of transducer and CTC AED.

【3】 A decade of DCASE: Achievements, practices, evaluations and future challenges
标题: DUSE十年:成就、实践、评估和未来挑战
作者: Annamaria Mesaros, Romain Serizel, Toni Heittola, Tuomas Virtanen, Mark D. Plumbley
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:本文简要介绍了声场景和事件的检测与分类(DCASE)挑战赛的历史和发展,研讨会,研究领域和研究团体。DCASE创建于2013年,作为一项数据评估挑战,已成为音频和声学信号处理领域的主要研究课题。它的成功来自于多种因素:挑战提供了每年更新的各种任务;讲习班提供了传播相关工作的渠道,吸引了年轻和充满活力的社区。与此同时,DCASE也面临着自己的挑战,不断发展并扩展到不同的领域。DCASE的核心原则之一是开放科学和可重复性:公开可用的数据集,基线系统,技术报告和研讨会出版物。虽然DCASE挑战和研讨会独立于IEEE SPS,但该挑战每年都会得到AASP TC的认可,DCASE社区为ICASSP旗舰会议和SPS在其许多活动中的成功做出了重大贡献。摘要:This paper introduces briefly the history and growth of the Detection and Classification of Acoustic Scenes and Events (DCASE) challenge, workshop, research area and research community. Created in 2013 as a data evaluation challenge, DCASE has become a major research topic in the Audio and Acoustic Signal Processing area. Its success comes from a combination of factors: the challenge offers a large variety of tasks that are renewed each year; and the workshop offers a channel for dissemination of related work, engaging a young and dynamic community. At the same time, DCASE faces its own challenges, growing and expanding to different areas. One of the core principles of DCASE is open science and reproducibility: publicly available datasets, baseline systems, technical reports and workshop publications. While the DCASE challenge and workshop are independent of IEEE SPS, the challenge receives annual endorsement from the AASP TC, and the DCASE community contributes significantly to the ICASSP flagship conference and the success of SPS in many of its activities.

【4】 Towards Ultra-Low-Power Neuromorphic Speech Enhancement with Spiking-FullSubNet
标题: 利用尖峰全SubNet实现超低功耗神经形态语音增强
作者: Xiang Hao, Chenxiang Ma, Qu Yang, Jibin Wu, Kay Chen Tan
备注:under review
链接:点击下载PDF文件
摘要:语音增强对于提高各种音频设备中的语音清晰度和质量至关重要。近年来,基于深度学习的方法显着提高了语音增强性能,但它们通常具有很高的计算成本,这对于耳机和助听器等大量边缘设备来说是令人望而却步的。本文提出了一种基于脑激励脉冲神经网络(SNN)的超低功耗语音增强系统,称为Spiking-FullSubNet。Spiking-FullSubNet采用全波段和子波段融合的方法,有效地捕获全局和局部光谱信息。为了提高计算昂贵的子带建模的效率,我们引入了一个频率划分方法的灵感来自人类外周听觉系统的灵敏度分布。此外,我们引入了一种新的尖峰神经元模型,可以动态控制输入信息的整合和遗忘,增强SNN的多尺度时间处理能力,这是语音去噪的关键。在最近的Intel Neuromorphic Deep Noise Suppression(N-DNS)Challenge数据集上进行的实验表明,Spiking-FullSubNet在语音质量和能效指标方面都大大超过了最先进的方法。值得注意的是,我们的系统赢得了英特尔N-DNS挑战赛(英语:Intel N-DNS Challenge)的冠军,为边缘的超低功耗语音增强提供了无数机会。我们的源代码和模型检查点可在https: github.com haoxiangsnr spiking-fullsubnet上公开获取。摘要:Speech enhancement is critical for improving speech intelligibility and quality in various audio devices. In recent years, deep learning-based methods have significantly improved speech enhancement performance, but they often come with a high computational cost, which is prohibitive for a large number of edge devices, such as headsets and hearing aids. This work proposes an ultra-low-power speech enhancement system based on the brain-inspired spiking neural network (SNN) called Spiking-FullSubNet. Spiking-FullSubNet follows a full-band and sub-band fusioned approach to effectively capture both global and local spectral information. To enhance the efficiency of computationally expensive sub-band modeling, we introduce a frequency partitioning method inspired by the sensitivity profile of the human peripheral auditory system. Furthermore, we introduce a novel spiking neuron model that can dynamically control the input information integration and forgetting, enhancing the multi-scale temporal processing capability of SNN, which is critical for speech denoising. Experiments conducted on the recent Intel Neuromorphic Deep Noise Suppression (N-DNS) Challenge dataset show that the Spiking-FullSubNet surpasses state-of-the-art methods by large margins in terms of both speech quality and energy efficiency metrics. Notably, our system won the championship of the Intel N-DNS Challenge (Algorithmic Track), opening up a myriad of opportunities for ultra-low-power speech enhancement at the edge. Our source code and model checkpoints are publicly available at https: github.com haoxiangsnr spiking-fullsubnet.

【5】 SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
标题: SegTIN:神经文本到语音中用于序列对齐的逐段隐式神经表示
作者: Minchan Kim, Myeonghun Jeong, Joun Yeop Lee, Nam Soo Kim
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:我们提出了SegINR,一种新的神经文本到语音(TTS)的方法,它解决了序列对齐问题,而不依赖于辅助持续时间预测器和复杂的自回归(AR)或非自回归(NAR)帧级序列建模。SegINR通过将文本序列直接转换为帧级特征来简化该过程。它利用最佳文本编码器来提取嵌入,使用条件隐式神经表示(INR)将每个嵌入转换为一段帧级特征。这种方法被称为分段INR(SegINR),它对每个分段内的时间动态进行建模,并自主定义分段边界,从而降低计算成本。我们将SegINR集成到一个两阶段的TTS框架中,使用它进行语义标记预测。我们在zero-shot自适应TTS场景中的实验表明,SegINR在语音质量和计算效率方面优于传统方法。摘要:We present SegINR, a novel approach to neural Text-to-Speech (TTS) that addresses sequence alignment without relying on an auxiliary duration predictor and complex autoregressive (AR) or non-autoregressive (NAR) frame-level sequence modeling. SegINR simplifies the process by converting text sequences directly into frame-level features. It leverages an optimal text encoder to extract embeddings, transforming each into a segment of frame-level features using a conditional implicit neural representation (INR). This method, named segment-wise INR (SegINR), models temporal dynamics within each segment and autonomously defines segment boundaries, reducing computational costs. We integrate SegINR into a two-stage TTS framework, using it for semantic token prediction. Our experiments in zero-shot adaptive TTS scenarios demonstrate that SegINR outperforms conventional methods in speech quality with computational efficiency.

【6】 HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
标题: HALL-E:用于分钟长Zero-Shot文本到语音合成的分层神经编解码语言模型
作者: Yuto Nishimura, Takumi Hirose, Masanari Ohi, Hideki Nakayama, Nakamasa Inoue
链接:点击下载PDF文件
摘要:最近,基于将自然语言文本翻译成离散音频令牌序列的大语言模型(LLM)的文本到语音(TTS)模型获得了极大的研究关注,其中神经音频编解码器(NAC)模型使用残差矢量量化(RVQ)的进展。然而,由于高帧速率,长形式语音合成仍然是一个重大挑战,这增加了音频令牌的长度,并且使得自回归语言模型难以为甚至一分钟的语音生成音频令牌。为了应对这一挑战,本文介绍了两种新的后训练方法:1)多分辨率重新量化(MReQ)和2)HALL-E。MReQ是一个降低预训练NAC模型帧速率的框架。具体来说,它采用了多分辨率残差矢量量化(MRVQ)模块,分层重组离散的音频令牌,通过师生蒸馏。HALL-E是一种基于LLM的TTS模型,旨在预测MReQ的分层令牌。具体来说,它结合了使用MRVQ子模块的技术,并从预先训练的基于LLM的TTS模型继续训练。此外,为了促进TTS研究,我们创建了MinutesSpeech,这是一个新的基准数据集,由4万小时的过滤语音数据组成,用于训练和评估从3s到180 s的语音合成。在实验中,我们通过将我们的后训练框架应用于VALL-E来证明我们方法的有效性。我们实现了低至8 Hz的帧速率,从而在单个推理步骤中实现了稳定的小时长语音合成。音频样本、数据集、代码和预训练模型可在https: yutonishimura-v2.github.io HALL-E_DEMO 上获得。摘要:Recently, Text-to-speech (TTS) models based on large language models (LLMs) that translate natural language text into sequences of discrete audio tokens have gained great research attention, with advances in neural audio codec (NAC) models using residual vector quantization (RVQ). However, long-form speech synthesis remains a significant challenge due to the high frame rate, which increases the length of audio tokens and makes it difficult for autoregressive language models to generate audio tokens for even a minute of speech. To address this challenge, this paper introduces two novel post-training approaches: 1) Multi-Resolution Requantization (MReQ) and 2) HALL-E. MReQ is a framework to reduce the frame rate of pre-trained NAC models. Specifically, it incorporates multi-resolution residual vector quantization (MRVQ) module that hierarchically reorganizes discrete audio tokens through teacher-student distillation. HALL-E is an LLM-based TTS model designed to predict hierarchical tokens of MReQ. Specifically, it incorporates the technique of using MRVQ sub-modules and continues training from a pre-trained LLM-based TTS model. Furthermore, to promote TTS research, we create MinutesSpeech, a new benchmark dataset consisting of 40k hours of filtered speech data for training and evaluating speech synthesis ranging from 3s up to 180s. In experiments, we demonstrated the effectiveness of our approaches by applying our post-training framework to VALL-E. We achieved the frame rate down to as low as 8 Hz, enabling the stable minitue-long speech synthesis in a single inference step. Audio samples, dataset, codes and pre-trained models are available at https: yutonishimura-v2.github.io HALL-E_DEMO .

【7】 DJ Mix Transcription with Multi-Pass Non-Negative Matrix Factorization
标题: 具有多遍非负矩阵分解的DJ混音转录
作者: Étienne Paul André, Dominique Fourer, Diemo Schwarz
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:DJ混音转录是DJ混音逆向工程的关键一步,它估计应用于一组现有曲目的参数集和音频效果,以产生表演性的DJ混音。我们介绍了一种新的方法的基础上,多通NMF算法的字典矩阵对应于一组的频谱图切片的源轨道中存在的混合。 多遍策略的动机是由于使用大的NMF字典导致的高计算成本。该方法使用通间滤波,有利于时间连续性和稀疏性,并在公开可用的数据集上进行评估。 我们的比较结果考虑基于动态时间规整(DTW)的基线方法是有前途的,为未来基于NMF的应用铺平了道路。摘要:DJ mix transcription is a crucial step towards DJ mix reverse engineering, which estimates the set of parameters and audio effects applied to a set of existing tracks to produce a performative DJ mix. We introduce a new approach based on a multi-pass NMF algorithm where the dictionary matrix corresponds to a set of spectrogram slices of the source tracks present in the mix. The multi-pass strategy is motivated by the high computational cost resulting from the use of a large NMF dictionary. The proposed method uses inter-pass filtering to favor temporal continuity and sparseness and is evaluated on a publicly available dataset. Our comparative results considering a baseline method based on dynamic time warping (DTW) are promising and pave the way of future NMF-based applications.

【8】 Enhancement of Dysarthric Speech Reconstruction by Contrastive Learning
标题: 通过对比学习增强结构障碍性言语重建
作者: Keshvari Fatemeh, Mahdian Toroghi Rahil, Zareian Hassan
链接:点击下载PDF文件
摘要:构音障碍的语音重建是具有挑战性的,由于其病理性的声音模式。保护说话人的身份,特别是在没有正常语音的情况下,是一个关键的挑战。我们提出的方法使用对比学习提取说话人嵌入重建,而采用XLS-R表示,而不是滤波器组。结果表明,改善语音质量,自然度,可懂度,说话人身份的保护,和性别的一致性为女性发言人。重建的语音表现出1.51和2.12 MOS分数的改善,减少25.45%和32.1%的错误率为中度和中度-重度构音障碍的发言者使用Jasper语音识别系统,分别。这种方法为构音障碍言语重建提供了有希望的进展。摘要:Dysarthric speech reconstruction is challenging due to its pathological sound patterns. Preserving speaker identity, especially without access to normal speech, is a key challenge. Our proposed approach uses contrastive learning to extract speaker embedding for reconstruction, while employing XLS-R representations instead of filter banks. The results show improved speech quality, naturalness, intelligibility, speaker identity preservation, and gender consistency for female speakers. Reconstructed speech exhibits 1.51 and 2.12 MOS score improvements and reduces word error rates by 25.45% and 32.1% for moderate and moderate-severe dysarthria speakers using Jasper speech recognition system, respectively. This approach offers promising advancements in dysarthric speech reconstruction.

【9】 Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech System
标题: 基于说话人嵌入的Zero-Shot文本到语音系统中的对抗攻击和鲁棒防御
作者: Ze Li, Yao Shi, Yunfei Xu, Ming Li
链接:点击下载PDF文件
摘要:基于说话人嵌入的zero-shot文本到语音(TTS)系统使得能够使用最少的数据为看不见的说话人进行高质量的语音合成。然而,这些系统很容易受到对抗性攻击,攻击者会对原始说话者的音频波形引入难以察觉的扰动,导致合成语音听起来像另一个人。此漏洞会带来重大的安全风险,包括扬声器身份欺骗和未经授权的语音操作。本文研究了两种主要的防御策略来解决这些威胁:对抗性训练和对抗性净化。对抗性训练通过在训练过程中整合对抗性示例来增强模型的鲁棒性,从而提高对此类攻击的抵抗力。另一方面,对抗性净化采用扩散概率模型将受对抗性干扰的音频恢复为其干净的形式。实验结果表明,这些防御机制可以有效降低对抗干扰的影响,提高基于说话人嵌入的zero-shot TTS系统在对抗环境中的安全性和可靠性。摘要:Speaker embedding based zero-shot Text-to-Speech (TTS) systems enable high-quality speech synthesis for unseen speakers using minimal data. However, these systems are vulnerable to adversarial attacks, where an attacker introduces imperceptible perturbations to the original speaker's audio waveform, leading to synthesized speech sounds like another person. This vulnerability poses significant security risks, including speaker identity spoofing and unauthorized voice manipulation. This paper investigates two primary defense strategies to address these threats: adversarial training and adversarial purification. Adversarial training enhances the model's robustness by integrating adversarial examples during the training process, thereby improving resistance to such attacks. Adversarial purification, on the other hand, employs diffusion probabilistic models to revert adversarially perturbed audio to its clean form. Experimental results demonstrate that these defense mechanisms can significantly reduce the impact of adversarial perturbations, enhancing the security and reliability of speaker embedding based zero-shot TTS systems in adversarial environments.

【10】 Presto! Distilling Steps and Layers for Accelerating Music Generation
标题: 快点!提炼步骤和层次以加速音乐生成
作者: Zachary Novack, Ge Zhu, Jonah Casebeer, Julian McAuley, Taylor Berg-Kirkpatrick, Nicholas J. Bryan
链接:点击下载PDF文件
摘要:尽管基于扩散的文本到音乐(TTM)方法取得了进展,但高效、高质量的生成仍然是一个挑战。我们介绍Presto!,一种通过减少采样步骤和每步成本来加速基于分数的扩散Transformers的推理的方法。为了减少步骤,我们开发了一种新的基于分数的分布匹配蒸馏(DMD)方法,用于EDM系列扩散模型,这是第一种基于GAN的TTM蒸馏方法。为了降低每一步的成本,我们开发了一个简单的,但强大的改进最近层蒸馏方法,通过更好地保留隐藏状态方差来提高学习。最后,我们结合我们的步骤和层蒸馏方法在一起的一个双方面的方法。我们独立评估我们的阶梯和分层蒸馏方法,并展示每种方法的最佳性能。我们的组合蒸馏方法可以产生高质量的输出,提高多样性,将我们的基础模型加速10- 18倍(32秒单声道 立体声44.1kHz的230 435 ms延迟,比同类SOTA快15倍)-据我们所知,这是最快的高质量TTM。可以在https: presto-music.github.io web 上找到合理的例子。摘要:Despite advances in diffusion-based text-to-music (TTM) methods, efficient, high-quality generation remains a challenge. We introduce Presto!, an approach to inference acceleration for score-based diffusion transformers via reducing both sampling steps and cost per step. To reduce steps, we develop a new score-based distribution matching distillation (DMD) method for the EDM-family of diffusion models, the first GAN-based distillation method for TTM. To reduce the cost per step, we develop a simple, but powerful improvement to a recent layer distillation method that improves learning via better preserving hidden state variance. Finally, we combine our step and layer distillation methods together for a dual-faceted approach. We evaluate our step and layer distillation methods independently and show each yield best-in-class performance. Our combined distillation method can generate high-quality outputs with improved diversity, accelerating our base model by 10-18x (230 435ms latency for 32 second mono stereo 44.1kHz, 15x faster than comparable SOTA) -- the fastest high-quality TTM to our knowledge. Sound examples can be found at https: presto-music.github.io web .

【11】 CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
标题: CTC-GMM:用于快速准确的流语音翻译的CTC-GMM引导模式匹配
作者: Rui Zhao, Jinyu Li, Ruchao Fan, Matt Post
备注:Accepted by IEEE Spoken Language Technology Workshop (SLT 2024)
链接:点击下载PDF文件
摘要:流语音翻译(ST)的模型可以实现高准确性和低延迟,如果它们是用源语言的大量配对音频和目标语言的书面文本开发的。然而,由于手动ST数据标记的高昂成本,这些用于目标语言的文本标签通常是伪标签。在本文中,我们介绍了一种名为连接主义时间分类引导模态匹配(CTC-GMM)的方法,该方法通过利用广泛的机器翻译(MT)文本数据来增强流ST模型。该技术采用CTC将语音序列压缩成与相应文本序列相匹配的紧凑嵌入序列,允许我们利用来自MT语料库的匹配的{源-目标}语言文本对来进一步细化流ST模型。我们对FLEURS和CoVoST 2的评估表明,CTC-GMM方法可以分别将翻译准确率提高13.9%和6.4%,同时还可以将GPU上的解码速度提高59.7%。摘要:Models for streaming speech translation (ST) can achieve high accuracy and low latency if they're developed with vast amounts of paired audio in the source language and written text in the target language. Yet, these text labels for the target language are often pseudo labels due to the prohibitive cost of manual ST data labeling. In this paper, we introduce a methodology named Connectionist Temporal Classification guided modality matching (CTC-GMM) that enhances the streaming ST model by leveraging extensive machine translation (MT) text data. This technique employs CTC to compress the speech sequence into a compact embedding sequence that matches the corresponding text sequence, allowing us to utilize matched {source-target} language text pairs from the MT corpora to refine the streaming ST model further. Our evaluations with FLEURS and CoVoST2 show that the CTC-GMM approach can increase translation accuracy relatively by 13.9% and 6.4% respectively, while also boosting decoding speed by 59.7% on GPU.

【12】 Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
标题: 使用多尺度特征的对比损失改进说话者表示
作者: Satvik Dixit, Massa Baali, Rita Singh, Bhiksha Raj
链接:点击下载PDF文件
摘要:随着多尺度特征聚合(MFA)架构的引入,说话人确认系统已经取得了重大进展,例如MFA-Conformer和ECAPA-TDNN。这些模型通过在池化层和投影层之前连接中间特征图来利用来自不同网络深度的信息,表明即使是较浅的特征图也会编码有价值的说话者特定信息。在此基础上,我们提出了一个多尺度特征对比(MFCon)损失,直接提高这些中间表示的质量。我们的MFCon损失将对比学习应用于网络中的所有特征图,鼓励模型在中间阶段学习更多的判别表示。通过执行更好的特征映射学习,我们表明,由此产生的扬声器嵌入表现出更高的区分能力。我们的方法实现了9.05%的改善,在等错误率(EER)相比,标准的MFA-构象的VoxCeleb-1 O测试集。摘要:Speaker verification systems have seen significant advancements with the introduction of Multi-scale Feature Aggregation (MFA) architectures, such as MFA-Conformer and ECAPA-TDNN. These models leverage information from various network depths by concatenating intermediate feature maps before the pooling and projection layers, demonstrating that even shallower feature maps encode valuable speaker-specific information. Building upon this foundation, we propose a Multi-scale Feature Contrastive (MFCon) loss that directly enhances the quality of these intermediate representations. Our MFCon loss applies contrastive learning to all feature maps within the network, encouraging the model to learn more discriminative representations at the intermediate stage itself. By enforcing better feature map learning, we show that the resulting speaker embeddings exhibit increased discriminative power. Our method achieves a 9.05% improvement in equal error rate (EER) compared to the standard MFA-Conformer on the VoxCeleb-1O test set.

【13】 RelUNet: Relative Channel Fusion U-Net for Multichannel Speech Enhancement
标题: RelUNet:用于多通道语音增强的相对通道融合U-Net
作者: Ibrahim Aldarmaki, Thamar Solorio, Bhiksha Raj, Hanan Aldarmaki
链接:点击下载PDF文件
摘要:神经多通道语音增强模型,特别是那些基于U-Net架构,表现出良好的性能和推广潜力。这些模型通常独立地对输入通道进行编码,并在网络的后期阶段集成通道。在本文中,我们提出了一种新的修改这些模型,从一开始就将相关信息,其中每个通道的处理与参考通道通过堆叠。该输入策略利用比较差异自适应地融合通道之间的信息,从而捕获关键的空间信息并提高整体性能。在CHiME-3数据集上进行的实验证明了各种架构的语音增强指标的改进。摘要:Neural multi-channel speech enhancement models, in particular those based on the U-Net architecture, demonstrate promising performance and generalization potential. These models typically encode input channels independently, and integrate the channels during later stages of the network. In this paper, we propose a novel modification of these models by incorporating relative information from the outset, where each channel is processed in conjunction with a reference channel through stacking. This input strategy exploits comparative differences to adaptively fuse information between channels, thereby capturing crucial spatial information and enhancing the overall performance. The experiments conducted on the CHiME-3 dataset demonstrate improvements in speech enhancement metrics across various architectures.

【14】 Stage-Wise and Prior-Aware Neural Speech Phase Prediction
标题: 分阶段和优先感知神经语音阶段预测
作者: Fei Liu, Yang Ai, Hui-Peng Du, Ye-Xin Lu, Rui-Chen Zheng, Zhen-Hua Ling
备注:Accepted by SLT2024
链接:点击下载PDF文件
摘要:本文提出了一种新的逐段和先验感知的神经语音相位预测(SP-NSPP)模型,该模型通过两段神经网络从输入的幅度谱中预测相位谱。在初始先验构造阶段,我们从振幅谱中初步预测出一个粗略的先验相位谱。随后的细化阶段将幅度谱变换成以先前相位为条件的细化的高质量相位谱。这两个阶段的网络都使用ConvNeXt v2块作为骨干,并通过创新性地引入相位谱图(PSD)来采用对抗训练。为了进一步提高细化相位的连续性,我们还在细化阶段引入了时频积分差分(TFID)损耗。实验结果表明,与基于神经网络的无先验相位预测方法相比,SP-NSPP由于引入了粗相位先验和多样化的训练准则,获得了更高的相位预测精度。与迭代相位估计算法相比,我们提出的SP-NSPP不需要多轮分阶段迭代,从而产生更高的效率。摘要:This paper proposes a novel Stage-wise and Prior-aware Neural Speech Phase Prediction (SP-NSPP) model, which predicts the phase spectrum from input amplitude spectrum by two-stage neural networks. In the initial prior-construction stage, we preliminarily predict a rough prior phase spectrum from the amplitude spectrum. The subsequent refinement stage transforms the amplitude spectrum into a refined high-quality phase spectrum conditioned on the prior phase. Networks in both stages use ConvNeXt v2 blocks as the backbone and adopt adversarial training by innovatively introducing a phase spectrum discriminator (PSD). To further improve the continuity of the refined phase, we also incorporate a time-frequency integrated difference (TFID) loss in the refinement stage. Experimental results confirm that, compared to neural network-based no-prior phase prediction methods, the proposed SP-NSPP achieves higher phase prediction accuracy, thanks to introducing the coarse phase priors and diverse training criteria. Compared to iterative phase estimation algorithms, our proposed SP-NSPP does not require multiple rounds of staged iterations, resulting in higher generation efficiency.

【15】 Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation
标题: Art 2 Mus:通过跨模式一代架起视觉艺术与音乐的桥梁
作者: Ivan Rinaldi, Nicola Fanelli, Giovanna Castellano, Gennaro Vessio
备注:Presented at the AI for Visual Arts (AI4VA) workshop at ECCV 2024
链接:点击下载PDF文件
摘要:Artificial Intelligence and generative models have revolutionized music creation, with many models leveraging textual or visual prompts for guidance. However, existing image-to-music models are limited to simple images, lacking the capability to generate music from complex digitized artworks. To address this gap, we introduce $ mathcal{A} textit{rt2} mathcal{M} textit{us}$, a novel model designed to create music from digitized artworks or text inputs. $ mathcal{A} textit{rt2} mathcal{M} textit{us}$ extends the AudioLDM~2 architecture, a text-to-audio model, and employs our newly curated datasets, created via ImageBind, which pair digitized artworks with music. Experimental results demonstrate that $ mathcal{A} textit{rt2} mathcal{M} textit{us}$ can generate music that resonates with the input stimuli. These findings suggest promising applications in multimedia art, interactive installations, and AI-driven creative tools.摘要:Artificial Intelligence and generative models have revolutionized music creation, with many models leveraging textual or visual prompts for guidance. However, existing image-to-music models are limited to simple images, lacking the capability to generate music from complex digitized artworks. To address this gap, we introduce $ mathcal{A} textit{rt2} mathcal{M} textit{us}$, a novel model designed to create music from digitized artworks or text inputs. $ mathcal{A} textit{rt2} mathcal{M} textit{us}$ extends the AudioLDM~2 architecture, a text-to-audio model, and employs our newly curated datasets, created via ImageBind, which pair digitized artworks with music. Experimental results demonstrate that $ mathcal{A} textit{rt2} mathcal{M} textit{us}$ can generate music that resonates with the input stimuli. These findings suggest promising applications in multimedia art, interactive installations, and AI-driven creative tools.

【16】 Attentive-based Multi-level Feature Fusion for Voice Disorder Diagnosis
标题: 基于注意力的多层特征融合用于语音障碍诊断
作者: Lipeng Shen, Yifan Xiong, Dongyue Guo, Wei Mo, Lingyu Yu, Hui Yang, Yi Lin
链接:点击下载PDF文件
摘要:语音障碍以各种方式对日常生活质量产生负面影响。然而,由于数据集有限,从原始音频中准确识别病理特征的类别仍然是一个相当大的挑战。一个很有前途的方法来处理这个问题是提取多层次的病理信息,在语音的综合方式融合特征的潜在空间。本文设计了一种新的框架,探索高质量的特征融合的方式,有效的和广义的检测性能。具体而言,该模型采用两阶段训练模式:(1)采用在多个领域都表现出显著效果的ECAPA-TDNN和Wav 2 vec 2.0,从原始音频中学习普遍的病理信息;(2)专门设计了一个注意融合模块,用于建立EcapTdnn和Wav 2 vec 2.0分别投影的病理特征之间的相互作用,并引导多个特征的融合。层融合,整个模型由自动语音病理检测任务从预训练的特征联合微调。最后,在FEMH和SVD数据集上的综合实验表明,该框架优于竞争基线,达到90.51%和87.68%的准确率。摘要:Voice disorders negatively impact the quality of daily life in various ways. However, accurately recognizing the category of pathological features from raw audio remains a considerable challenge due to the limited dataset. A promising method to handle this issue is extracting multi-level pathological information from speech in a comprehensive manner by fusing features in the latent space. In this paper, a novel framework is designed to explore the way of high-quality feature fusion for effective and generalized detection performance. Specifically, the proposed model follows a two-stage training paradigm: (1) ECAPA-TDNN and Wav2vec 2.0 which have shown remarkable effectiveness in various domains are employed to learn the universal pathological information from raw audio; (2) An attentive fusion module is dedicatedly designed to establish the interaction between pathological features projected by EcapTdnn and Wav2vec 2.0 respectively and guide the multi-layer fusion, the entire model is jointly fine-tuned from pre-trained features by the automatic voice pathology detection task. Finally, comprehensive experiments on the FEMH and SVD datasets demonstrate that the proposed framework outperforms the competitive baselines, and achieves the accuracy of 90.51% and 87.68%.

【17】 Modeling and Estimation of Vocal Tract and Glottal Source Parameters Using ARMAX-LF Model
标题: 基于ARMAX-LF模型的声道和喉舌源参数建模与估计
作者: Kai Lia, Masato Akagia, Yongwei Lib, Masashi Unokia
链接:点击下载PDF文件
摘要:来自原始语音的元音的声道和声门源参数的建模和估计通常可以通过使用具有外源输入的自回归(ARX)模型和具有基于迭代的估计方法的Liljencrants-Fant(LF)模型来完成。然而,声道滤波器建模中的全极点自回归模型不能提供反共振峰(零)的位置,这增加了某些类别的语音声音(诸如鼻音、摩擦音和塞音)的估计误差。在本文中,我们提出了自回归移动平均线eXogenous LF(ARMAX-LF)模型扩展的ARX-LF模型,以更广泛的语音,包括元音和鼻音辅音。LF模型将声门源导数表示为参数化时域模型,并且ARMAX模型将声道表示为具有额外的外源LF激励作为输入的极零滤波器。为了以更少的误差估计多个参数,我们首先利用深度神经网络(DNN)强大的非线性拟合能力,从提取的声门源导数或语音波形到相应的LF参数建立映射。然后,声门源和声道参数可以估计较少的估计误差和没有任何迭代的分析合成策略。使用线性源滤波器模型的合成语音、使用物理模型的合成语音和真实语音信号的实验结果表明,所提出的ARMAX-LF模型与基于DNN的估计方法可以估计元音和鼻音的参数,具有更少的误差和估计时间。摘要:Modeling and estimation of the vocal tract and glottal source parameters of vowels from raw speech can be typically done by using the Auto-Regressive with eXogenous input (ARX) model and Liljencrants-Fant (LF) model with an iteration-based estimation approach. However, the all-pole autoregressive model in the modeling of vocal tract filters cannot provide the locations of anti-formants (zeros), which increases the estimation errors in certain classes of speech sounds, such as nasal, fricative, and stop consonants. In this paper, we propose the Auto-Regressive Moving Average eXogenous with LF (ARMAX-LF) model to extend the ARX-LF model to a wider variety of speech sounds, including vowels and nasalized consonants. The LF model represents the glottal source derivative as a parametrized time-domain model, and the ARMAX model represents the vocal tract as a pole-zero filter with an additional exogenous LF excitation as input. To estimate multiple parameters with fewer errors, we first utilize the powerful nonlinear fitting ability of deep neural networks (DNNs) to build a mapping from extracted glottal source derivatives or speech waveforms to corresponding LF parameters. Then, glottal source and vocal tract parameters can be estimated with fewer estimation errors and without any iterations as in the analysis-by-synthesis strategy. Experimental results with synthesized speech using the linear source-filter model, synthesized speech using the physical model, and real speech signals showed that the proposed ARMAX-LF model with a DNN-based estimation method can estimate the parameters of both vowels and nasalized sounds with fewer errors and estimation time.

【18】 Demo of Zero-Shot Guitar Amplifier Modelling: Enhancing Modeling with Hyper Neural Networks
标题: Zero-Shot吉他放大器建模演示:用超神经网络增强建模
作者: Yu-Hua Chen, Yuan-Chiao Cheng, Yen-Tung Yeh, Jui-Te Wu, Yu-Hsiang Ho, Jyh-Shing Roger Jang, Yi-Hsuan Yang
备注:demo of the ISMIR paper
链接:点击下载PDF文件
摘要:电吉他音调建模通常集中于从干净音频到增强器渲染音频的非线性变换。传统方法依赖于一对一映射,将设备参数纳入神经模型以复制特定的放大器。然而,这些方法受到特定训练数据需求的限制。在本文中,我们适应了一个模型的基础上,以前的工作,它利用了音调嵌入编码器和功能明智的线性调制(薄膜)条件方法。在这项工作中,我们使用基于超网络的门控卷积网络(GCN)来改变条件反射方法,以生成将干净输入与参考音频的音调特征混合的音频。通过扩展训练数据以覆盖更广泛的放大器音调,我们的模型能够捕获更广泛的音调。此外,我们还开发了一个实时插件来演示系统的实际应用,让用户可以交互式地体验其性能。我们的研究结果表明,该系统实现了优越的音调建模的通用性相比,传统的方法。摘要:Electric guitar tone modeling typically focuses on the non-linear transformation from clean to amplifier-rendered audio. Traditional methods rely on one-to-one mappings, incorporating device parameters into neural models to replicate specific amplifiers. However, these methods are limited by the need for specific training data. In this paper, we adapt a model based on the previous work, which leverages a tone embedding encoder and a feature wise linear modulation (FiLM) condition method. In this work, we altered conditioning method using a hypernetwork-based gated convolutional network (GCN) to generate audio that blends clean input with the tone characteristics of reference audio. By extending the training data to cover a wider variety of amplifier tones, our model is able to capture a broader range of tones. Additionally, we developed a real-time plugin to demonstrate the system's practical application, allowing users to experience its performance interactively. Our results indicate that the proposed system achieves superior tone modeling versatility compared to traditional methods.

【19】 UniMuMo: Unified Text, Music and Motion Generation
标题: UniMuMo:统一的文本、音乐和动作生成
作者: Han Yang, Kun Su, Yutong Zhang, Jiaben Chen, Kaizhi Qian, Gaowen Liu, Chuang Gan
链接:点击下载PDF文件
摘要:我们介绍了UniMuMo,一个统一的多模态模型,能够将任意文本,音乐和运动数据作为输入条件,以生成所有三种模态的输出。为了解决缺乏时间同步数据的问题,我们根据节奏模式对齐未配对的音乐和运动数据,以利用现有的大规模纯音乐和纯运动数据集。通过将音乐、运动和文本转换为基于标记的表示,我们的模型通过统一的编码器-解码器Transformer架构来桥接这些模态。为了在一个框架内支持多个生成任务,我们引入了几个体系结构的改进。我们建议用音乐码本对运动进行编码,将运动映射到与音乐相同的特征空间。我们引入了一个音乐运动并行生成方案,统一到一个单一的Transformer解码器架构的音乐运动联合生成的单一训练任务的所有音乐和运动生成任务。此外,该模型是通过微调现有的预训练单模态模型来设计的,大大减少了计算需求。大量的实验表明,UniMuMo实现了竞争力的结果,所有单向生成基准跨越音乐,运动和文本形式。定量结果可在 href{https: hanyangclarence.github.io unimumo_demo }{project page}中找到。摘要:We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data, we align unpaired music and motion data based on rhythmic patterns to leverage existing large-scale music-only and motion-only datasets. By converting music, motion, and text into token-based representation, our model bridges these modalities through a unified encoder-decoder transformer architecture. To support multiple generation tasks within a single framework, we introduce several architectural improvements. We propose encoding motion with a music codebook, mapping motion into the same feature space as music. We introduce a music-motion parallel generation scheme that unifies all music and motion generation tasks into a single transformer decoder architecture with a single training task of music-motion joint generation. Moreover, the model is designed by fine-tuning existing pre-trained single-modality models, significantly reducing computational demands. Extensive experiments demonstrate that UniMuMo achieves competitive results on all unidirectional generation benchmarks across music, motion, and text modalities. Quantitative results are available in the href{https: hanyangclarence.github.io unimumo_demo }{project page}.

【20】 Configurable Multilingual ASR with Speech Summary Representations
标题: 具有语音摘要表示的可配置多语言ASB
作者: Harrison Zhu, Ivan Fung, Yingke Zhu, Lahiru Samarakoon
备注:A preprint
链接:点击下载PDF文件
摘要:世界上大约有一半的人口是多语言的,这使得多语言ASR(MASR)至关重要。当地面实况语言事先未知时,部署多个单语模型具有挑战性。这激发了对可配置的多语言MASR模型的研究工作,这些模型可以手动提示或自动调整以识别特定语言。在本文中,我们提出了可配置的MASR模型与摘要向量(csvMASR),一种新的架构,旨在提高可配置性。我们的方法利用适配器,并引入语音摘要向量表示,灵感来自语音日记中的会话摘要表示,在话语级别结合特定语言组件的输出。我们还将一个辅助语言分类损失,以提高可配置性。使用多语言Libripeech(MLS)数据集中的7种语言的数据,csvMASR优于现有的MASR模型,并将单词错误率(WER)从10.33%降低到9.95%。此外,csvMASR在语言分类和提示任务中表现出卓越的性能。摘要:Approximately half of the world's population is multilingual, making multilingual ASR (MASR) essential. Deploying multiple monolingual models is challenging when the ground-truth language is unknown in advance. This motivates research efforts on configurable multilingual MASR models that can be prompted manually or adapted automatically to recognise specific languages. In this paper, we present the Configurable MASR model with Summary Vector (csvMASR), a novel architecture designed to enhance configurability. Our approach leverages adapters and introduces speech summary vector representations, inspired by conversational summary representations in speech diarization, to combine outputs from language-specific components at the utterance level. We also incorporate an auxiliary language classification loss to enhance configurability. Using data from 7 languages in the Multilingual Librispeech (MLS) dataset, csvMASR outperforms existing MASR models and reduces the word error rate (WER) from 10.33 % to 9.95 % when compared with the baseline. Additionally, csvMASR demonstrates superior performance in language classification and prompting tasks.

【21】 SONAR: A Synthetic AI-Audio Detection Framework~and Benchmark
标题: SONAR:合成人工智能音频检测框架~和基准
作者: Xiang Li, Pin-Yu Chen, Wenqi Wei
链接:点击下载PDF文件
摘要:使用生成人工智能(AI)技术的文本到语音(TTS)和语音转换(VC)的最新进展使得生成高质量和逼真的类人音频成为可能。这给区分人工智能合成语音与真实人类语音带来了重大挑战,并可能引发潜在的恶意滥用问题,如模仿和欺诈、传播错误信息、深度伪造和诈骗。然而,现有的人工智能合成音频检测技术并没有跟上步伐,并且在不同的数据集上往往表现出很差的泛化能力。在本文中,我们介绍了SONAR,一个人工智能音频检测框架和基准,旨在为区分尖端的人工智能合成的听觉内容提供全面的评估。SONAR包括一个来自9个不同音频合成平台的新型评估数据集,包括领先的TTS提供商和最先进的TTS模型。它是第一个在传统和基于基础模型的deepfake检测系统中统一基准测试AI音频检测的框架。通过大量的实验,我们揭示了现有检测方法的泛化局限性,并证明基础模型具有更强的泛化能力,这可以归因于它们的模型大小以及预训练数据的规模和质量。此外,我们探讨了Few-Shot微调在提高泛化能力方面的有效性和效率,突出了其在定制应用中的潜力,例如针对特定实体或个人的个性化检测系统。代码和数据集可在https: github.com Jessegator SONAR上获得。摘要:Recent advances in Text-to-Speech (TTS) and Voice-Conversion (VC) using generative Artificial Intelligence (AI) technology have made it possible to generate high-quality and realistic human-like audio. This introduces significant challenges to distinguishing AI-synthesized speech from the authentic human voice and could raise potential issues of misuse for malicious purposes such as impersonation and fraud, spreading misinformation, deepfakes, and scams. However, existing detection techniques for AI-synthesized audio have not kept pace and often exhibit poor generalization across diverse datasets. In this paper, we introduce SONAR, a synthetic AI-Audio Detection Framework and Benchmark, aiming to provide a comprehensive evaluation for distinguishing cutting-edge AI-synthesized auditory content. SONAR includes a novel evaluation dataset sourced from 9 diverse audio synthesis platforms, including leading TTS providers and state-of-the-art TTS models. It is the first framework to uniformly benchmark AI-audio detection across both traditional and foundation model-based deepfake detection systems. Through extensive experiments, we reveal the generalization limitations of existing detection methods and demonstrate that foundation models exhibit stronger generalization capabilities, which can be attributed to their model size and the scale and quality of pretraining data. Additionally, we explore the effectiveness and efficiency of few-shot fine-tuning in improving generalization, highlighting its potential for tailored applications, such as personalized detection systems for specific entities or individuals. Code and dataset are available at https: github.com Jessegator SONAR.

【22】 Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
标题: 使用混合H3-conformer实现高效、稳健的长形式语音识别
作者: Tomoki Honda, Shinsuke Sakai, Tatsuya Kawahara
备注:Submitted to InterSpeech2024, Sample code is available at this https URL
链接:点击下载PDF文件
摘要:最近,Conformer在许多语音识别任务中取得了最先进的性能。然而,基于transformer的模型对于长形式的语音(例如演讲)表现出显著的恶化,因为自我注意机制随着输入长度的平方阶的计算而变得不可靠。为了解决这个问题,我们引入了一种状态空间模型,饥饿的饥饿河马(H3),以取代或补充多头自我注意(MHSA)。H3允许使用线性阶计算对长形式序列进行有效建模。在使用CSJ和LibriSpeech两个数据集的实验中,我们提出的H3-Conformer模型对长格式语音进行了高效和鲁棒的识别。此外,我们提出了一个混合的H3和MHSA,并表明,使用H3在较高层和MHSA在较低层提供了显着的改善在线识别。我们还研究了在所有层中并行使用H3和MHSA,从而获得最佳性能。摘要:Recently, Conformer has achieved state-of-the-art performance in many speech recognition tasks. However, the Transformer-based models show significant deterioration for long-form speech, such as lectures, because the self-attention mechanism becomes unreliable with the computation of the square order of the input length. To solve the problem, we incorporate a kind of state-space model, Hungry Hungry Hippos (H3), to replace or complement the multi-head self-attention (MHSA). H3 allows for efficient modeling of long-form sequences with a linear-order computation. In experiments using two datasets of CSJ and LibriSpeech, our proposed H3-Conformer model performs efficient and robust recognition of long-form speech. Moreover, we propose a hybrid of H3 and MHSA and show that using H3 in higher layers and MHSA in lower layers provides significant improvement in online recognition. We also investigate a parallel use of H3 and MHSA in all layers, resulting in the best performance.

【23】 The OCON model: an old but green solution for distributable supervised classification for acoustic monitoring in smart cities
标题: OCON模型:一种古老但绿色的解决方案,用于智能城市声学监测的分布式监督分类
作者: Stefano Giacomelli, Marco Giordano, Claudia Rinaldi
备注:Accepted at "IEEE 5th International Symposium on the Internet of Sounds, 30 Sep 2 Oct 2024, Erlangen, Germany"
Journal-ref:in Proceedings of the 5th IEEE International Symposium on the Internet of Sounds (IEEE IS2 2024, https:internetofsounds.netis2_2024)
链接:点击下载PDF文件
摘要:本文探讨了一个结构化的应用程序的一类方法和一类一网络模型的监督分类任务,侧重于元音音素分类和说话人识别的自动语音识别(ASR)域。在我们的案例研究中,ASR模型运行在专有的传感和闪电系统上,用于监测城市街道上的声学和空气污染。我们正式组合的伪神经架构搜索和超参数调整实验,使用一个明智的网格搜索方法,以实现分类精度相媲美,现在最复杂的架构,深入到说话人识别和能源效率方面。尽管它的简单性,我们的模型建议有一个很好的机会来概括的语言和扬声器性别的背景下,广泛适用于计算约束的情况下,证明了相关的统计和性能指标。我们的实验代码可以在我们的GitHub上公开访问。摘要:This paper explores a structured application of the One-Class approach and the One-Class-One-Network model for supervised classification tasks, focusing on vowel phonemes classification and speakers recognition for the Automatic Speech Recognition (ASR) domain. For our case-study, the ASR model runs on a proprietary sensing and lightning system, exploited to monitor acoustic and air pollution on urban streets. We formalize combinations of pseudo-Neural Architecture Search and Hyper-Parameters Tuning experiments, using an informed grid-search methodology, to achieve classification accuracy comparable to nowadays most complex architectures, delving into the speaker recognition and energy efficiency aspects. Despite its simplicity, our model proposal has a very good chance to generalize the language and speaker genders context for widespread applicability in computational constrained contexts, proved by relevant statistical and performance metrics. Our experiments code is openly accessible on our GitHub.

【24】 Cross-Lingual Query-by-Example Spoken Term Detection: A Transformer-Based Approach
标题: 跨语言逐例查询口语检测:基于转换器的方法
作者: Allahdadi Fatemeh, Mahdian Toroghi Rahil, Zareian Hassan
链接:点击下载PDF文件
摘要:通过实例查询的口语术语检测(QbE-STD)通常受到转录数据稀缺性和语言特异性的约束。本文介绍了一种新的,语言无关的QbE-STD模型,利用图像处理技术和Transformer架构。通过采用预训练的XLSR-53网络进行特征提取和Hough变换进行检测,我们的模型可以有效地搜索任何音频文件中的用户定义的口语术语。四种语言的实验结果表明,与基于CNN的基线相比,性能有显着提高(19-54%)。虽然与DTW相比处理时间有所改善,但准确性仍然较差。值得注意的是,我们的模型提供了准确计算目标音频内的查询词重复的优点。摘要:Query-by-example spoken term detection (QbE-STD) is typically constrained by transcribed data scarcity and language specificity. This paper introduces a novel, language-agnostic QbE-STD model leveraging image processing techniques and transformer architecture. By employing a pre-trained XLSR-53 network for feature extraction and a Hough transform for detection, our model effectively searches for user-defined spoken terms within any audio file. Experimental results across four languages demonstrate significant performance gains (19-54%) over a CNN-based baseline. While processing time is improved compared to DTW, accuracy remains inferior. Notably, our model offers the advantage of accurately counting query term repetitions within the target audio.

【25】 SyllableLM: Learning Coarse Semantic Units for Speech Language Models
标题: SyllableLM:学习语音语言模型的粗语义单元
作者: Alan Baade, Puyuan Peng, David Harwath
备注:10 pages, 2 figures
链接:点击下载PDF文件
摘要:语言模型需要标记化的输入。然而,音频和视觉等连续数据的标记化策略通常基于简单的算法,例如固定大小的卷积或离散聚类,这些算法不一定与数据的语义结构一致。特别是对于语音,波形的高分辨率(16,000个样本 秒或更高)提出了一个重大挑战,因为基于语音的语言模型必须使用比基于文本的语言模型多几倍的每个单词的令牌。在这项工作中,我们引入了一个可控的自我监督技术合并语音表示成粗糙的音节状单位,同时仍然保留语义信息。我们通过以下方式做到这一点:1)通过分析预训练编码器损失中的相关性来提取噪声边界; 2)使用一种新的蒸馏技术迭代地改进模型表示。我们的方法产生可重复的速率语义单位在低至5 Hz和60 bps,并达到SotA在音节分割和聚类。使用这些粗令牌,我们成功地训练了SyllableLM,这是一种语音语言模型(SpeechLM),在一系列口语建模任务上与当前SotA SpeechLM相匹配或优于当前SotA SpeechLM。SyllableLM还实现了效率的显着提高,训练计算减少了30倍,推理加速了4倍。摘要:Language models require tokenized inputs. However, tokenization strategies for continuous data like audio and vision are often based on simple heuristics such as fixed sized convolutions or discrete clustering, which do not necessarily align with the semantic structure of the data. For speech in particular, the high resolution of waveforms (16,000 samples second or more) presents a significant challenge as speech-based language models have had to use several times more tokens per word than text-based language models. In this work, we introduce a controllable self-supervised technique to merge speech representations into coarser syllable-like units while still preserving semantic information. We do this by 1) extracting noisy boundaries through analyzing correlations in pretrained encoder losses and 2) iteratively improving model representations with a novel distillation technique. Our method produces controllable-rate semantic units at as low as 5Hz and 60bps and achieves SotA in syllabic segmentation and clustering. Using these coarse tokens, we successfully train SyllableLM, a Speech Language Model (SpeechLM) that matches or outperforms current SotA SpeechLMs on a range of spoken language modeling tasks. SyllableLM also achieves significant improvements in efficiency with a 30x reduction in training compute and a 4x wall-clock inference speedup.

【26】 Reverb: Open-Source ASR and Diarization from Rev
标题: Reverb:开源ASB和Rev的日记
作者: Nishchal Bhandari, Danny Chen, Miguel Ángel del Río Fernández, Natalie Delworth, Jennifer Drexler Fox, Migüel Jetté, Quinten McNamara, Corey Miller, Ondřej Novotný, Ján Profant, Nan Qin, Martin Ratajczak, Jean-Philippe Robichaud
链接:点击下载PDF文件
摘要:今天,我们正在开源我们的核心语音识别和日志模型,用于非商业用途。我们正在为开发人员发布一个完整的生产管道,以及用于实验的精简研究模型。Rev希望这些版本将刺激快速发展的语音技术领域的研究和创新。今天发布的语音识别模型在各种长格式语音识别领域的性能优于所有现有的开源语音识别模型。摘要:Today, we are open-sourcing our core speech recognition and diarization models for non-commercial use. We are releasing both a full production pipeline for developers as well as pared-down research models for experimentation. Rev hopes that these releases will spur research and innovation in the fast-moving domain of voice technology. The speech recognition models released today outperform all existing open source speech recognition models across a variety of long-form speech recognition domains.

【27】 Did You Hear That? Introducing AADG: A Framework for Generating Benchmark Data in Audio Anomaly Detection
标题: 你听到了吗?介绍AADG:音频异常检测中生成基准数据的框架
作者: Ksheeraja Raghavan, Samiran Gode, Ankit Shah, Surabhi Raghavan, Wolfram Burgard, Bhiksha Raj, Rita Singh
备注:9 pages, under review
链接:点击下载PDF文件
摘要:我们介绍了一种新的,通用的音频生成框架,专门设计用于异常检测和定位。与主要关注工业和机器相关声音的现有数据集不同,我们的框架关注更广泛的环境,特别是在只有音频数据可用的现实世界场景中,例如视频衍生或电话音频。为了生成这样的数据,我们提出了一种受LLM-Modulo框架启发的新方法,该框架利用大型语言模型(LLM)作为世界模型来模拟真实世界的场景。该工具是模块化的,允许即插即用的方法。它首先使用LLM来预测合理的现实世界场景。LLM进一步提取组成声音,顺序和方式,这些应该合并,以创建连贯的整体。与LLM-Modulo框架非常相似,我们对每个输出阶段都进行了严格的验证,以确保生成数据的可靠性。使用该框架产生的数据作为异常检测应用程序的基准,可能会提高在音频数据上训练的模型的性能,特别是在处理分发外的情况下。因此,我们的贡献填补了音频异常检测资源的关键空白,并提供了一个可扩展的工具,用于生成多样化的,逼真的音频数据。摘要:We introduce a novel, general-purpose audio generation framework specifically designed for anomaly detection and localization. Unlike existing datasets that predominantly focus on industrial and machine-related sounds, our framework focuses a broader range of environments, particularly useful in real-world scenarios where only audio data are available, such as in video-derived or telephonic audio. To generate such data, we propose a new method inspired by the LLM-Modulo framework, which leverages large language models(LLMs) as world models to simulate such real-world scenarios. This tool is modular allowing a plug-and-play approach. It operates by first using LLMs to predict plausible real-world scenarios. An LLM further extracts the constituent sounds, the order and the way in which these should be merged to create coherent wholes. Much like the LLM-Modulo framework, we include rigorous verification of each output stage, ensuring the reliability of the generated data. The data produced using the framework serves as a benchmark for anomaly detection applications, potentially enhancing the performance of models trained on audio data, particularly in handling out-of-distribution cases. Our contributions thus fill a critical void in audio anomaly detection resources and provide a scalable tool for generating diverse, realistic audio data.

【28】 SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
标题: SONIQUE:使用未配对视听数据生成视频背景音乐
作者: Liqian Zhang, Magdalena Fuentes
链接:点击下载PDF文件
摘要:我们提出了SONIQUE,一个模型,用于生成定制的视频内容的背景音乐。与传统的视频到音乐生成方法不同,它严重依赖于配对的视听数据集,SONIQUE利用未配对的数据,将免版税音乐和独立的视频源相结合。通过利用大型语言模型(LLM)进行视频理解并将视觉描述转换为音乐标签,以及基于U-Net的条件扩散模型,SONIQUE可以实现可定制的音乐生成。用户可以控制音乐的特定方面,如乐器、流派、节奏和旋律,确保生成的输出符合他们的创意愿景。SONIQUE是开源的,有一个在线演示。摘要:We present SONIQUE, a model for generating background music tailored to video content. Unlike traditional video-to-music generation approaches, which rely heavily on paired audio-visual datasets, SONIQUE leverages unpaired data, combining royalty-free music and independent video sources. By utilizing large language models (LLMs) for video understanding and converting visual descriptions into musical tags, alongside a U-Net-based conditional diffusion model, SONIQUE enables customizable music generation. Users can control specific aspects of the music, such as instruments, genres, tempo, and melodies, ensuring the generated output fits their creative vision. SONIQUE is open-source, with a demo available online.

【29】 SOI: Scaling Down Computational Complexity by Estimating Partial States of the Model
标题: SIM:通过估计模型的部分状态来降低计算复杂性
作者: Grzegorz Stefański, Paweł Daniluk, Artur Szumaczuk, Jakub Tkaczuk
备注:NeurIPS 2024
链接:点击下载PDF文件
摘要:消费电子产品过去一直遵循摩尔定律所描述的小型化趋势。尽管微控制器单元(MCU)的处理能力有所提高,但用于最小家电的MCU仍然无法运行即使是中等规模的最先进的人工神经网络(ANN),特别是在时间敏感的情况下。在这项工作中,我们提出了一种新的方法称为分散在线推理(SOI),旨在降低人工神经网络的计算复杂度。SOI利用时间序列数据和模型预测的连续性和季节性,实现外推以提高处理速度,特别是在更深层。通过应用压缩,SOI生成ANN的更一般的内部部分状态,允许在每次推理时跳过完整的模型重新计算。摘要:Consumer electronics used to follow the miniaturization trend described by Moore's Law. Despite increased processing power in Microcontroller Units (MCUs), MCUs used in the smallest appliances are still not capable of running even moderately big, state-of-the-art artificial neural networks (ANNs) especially in time-sensitive scenarios. In this work, we present a novel method called Scattered Online Inference (SOI) that aims to reduce the computational complexity of ANNs. SOI leverages the continuity and seasonality of time-series data and model predictions, enabling extrapolation for processing speed improvements, particularly in deeper layers. By applying compression, SOI generates more general inner partial states of ANN, allowing skipping full model recalculation at each inference.

【30】 Self-Powered LLM Modality Expansion for Large Speech-Text Models
标题: 针对大型语音文本模型的自供电LLM情态扩展
作者: Tengfei Yu, Xuebo Liu, Zhiyi Hou, Liang Ding, Dacheng Tao, Min Zhang
备注:Accepted to EMNLP 2024
链接:点击下载PDF文件
摘要:大型语言模型(LLM)在不同的任务中表现出显着的性能,这表明它们有潜力通过集成语音功能扩展到大型语音-文本模型(LSM)。虽然统一的语音文本预训练和多模态数据预调整提供了相当大的好处,但这些方法通常需要大量的资源需求,并且往往会过度适应特定的任务。本研究旨在通过解决vanilla指令调整的局限性来改进LSM训练中语音数据集的使用。我们探讨了LSMs内的解释,以下动态,确定一个关键的问题,称为语音锚偏置LSMs过度依赖语音输入的倾向,错误地解释整个语音模态的指令,从而忽略了文本的指示。为了抵消这种偏见,我们引入了一个自供电的LSM,利用模型本身生成的增强自动语音识别数据进行更有效的指令调整。我们在一系列基于语音的任务中的实验表明,自供电LSM减轻了语音锚偏置,提高了LSM中语音和文本模态的融合。数据、代码和脚本可在https: github.com ytf-philp Self-powered-LSM上免费获得。摘要:Large language models (LLMs) exhibit remarkable performance across diverse tasks, indicating their potential for expansion into large speech-text models (LSMs) by integrating speech capabilities. Although unified speech-text pre-training and multimodal data instruction-tuning offer considerable benefits, these methods generally entail significant resource demands and tend to overfit specific tasks. This study aims to refine the use of speech datasets for LSM training by addressing the limitations of vanilla instruction tuning. We explore the instruction-following dynamics within LSMs, identifying a critical issue termed speech anchor bias-a tendency for LSMs to over-rely on speech inputs, mistakenly interpreting the entire speech modality as directives, thereby neglecting textual instructions. To counteract this bias, we introduce a self-powered LSM that leverages augmented automatic speech recognition data generated by the model itself for more effective instruction tuning. Our experiments across a range of speech-based tasks demonstrate that self-powered LSM mitigates speech anchor bias and improves the fusion of speech and text modalities in LSMs. Data, code and scripts are freely available at https: github.com ytf-philp Self-powered-LSM.

【31】 People are poorly equipped to detect AI-powered voice clones
标题: 人们检测人工智能语音克隆的能力很差
作者: Sarah Barrington, Hany Farid
链接:点击下载PDF文件
摘要:随着生成式人工智能继续其弹道轨迹,从文本到音频,图像和视频生成的一切都在模仿人类生成的内容方面继续改进。通过一系列的感知研究,我们报告了AI生成的声音在身份匹配和自然性方面的真实性。我们发现人类参与者无法可靠地识别人工智能生成的语音的短录音(不到20秒)。具体来说,参与者在80%的时间里将人工智能语音的身份误认为是其真实的对应物,而只有60%的时间将语音正确识别为人工智能生成的。在所有情况下,性能是独立的人口统计的发言者或听众。摘要:As generative AI continues its ballistic trajectory, everything from text to audio, image, and video generation continues to improve in mimicking human-generated content. Through a series of perceptual studies, we report on the realism of AI-generated voices in terms of identity matching and naturalness. We find human participants cannot reliably identify short recordings (less than 20 seconds) of AI-generated voices. Specifically, participants mistook the identity of an AI-voice for its real counterpart 80% of the time, and correctly identified a voice as AI-generated only 60% of the time. In all cases, performance is independent of the demographics of the speaker or listener.

【32】 Efficient Streaming LLM for Speech Recognition
标题: 用于语音识别的高效流媒体LLM
作者: Junteng Jia, Gil Keren, Wei Zhou, Egor Lakomkin, Xiaohui Zhang, Chunyang Wu, Frank Seide, Jay Mahadeokar, Ozlem Kalinli
链接:点击下载PDF文件
摘要:最近的工作表明,用音频编码提示大型语言模型可以解锁语音识别功能。然而,现有的技术不能有效地扩展,特别是在处理长格式流音频输入时--它们不仅在训练期间看到的音频长度之外外推得很差,而且由于注意力的二次成本,它们在计算上效率低下。 在这项工作中,我们介绍了SpeechLLM-XL,一个线性缩放解码器的流语音识别模型。我们使用有限的注意力窗口来处理可配置块中的音频以减少计算,并且每个音频块的文本令牌自回归地生成,直到预测EOS。在训练期间,使用从编码器输出估计的CTC强制对齐将转录本分割成块。具有1.28秒块大小的SpeechLLM-XL在LibriSpeech测试clean other上实现了2.7% 6.7%的WER,并且它在比训练话语长10倍的长形式话语上没有显示出质量下降。摘要:Recent works have shown that prompting large language models with audio encodings can unlock speech recognition capabilities. However, existing techniques do not scale efficiently, especially while handling long form streaming audio inputs -- not only do they extrapolate poorly beyond the audio length seen during training, but they are also computationally inefficient due to the quadratic cost of attention. In this work, we introduce SpeechLLM-XL, a linear scaling decoder-only model for streaming speech recognition. We process audios in configurable chunks using limited attention window for reduced computation, and the text tokens for each audio chunk are generated auto-regressively until an EOS is predicted. During training, the transcript is segmented into chunks, using a CTC forced alignment estimated from encoder output. SpeechLLM-XL with 1.28 seconds chunk size achieves 2.7% 6.7% WER on LibriSpeech test clean other, and it shows no quality degradation on long form utterances 10x longer than the training utterances.

【33】 Recent Advances in Speech Language Models: A Survey
标题: 言语语言模型的最新进展:调查
作者: Wenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng, Guangyan Zhang, Qichao Wang, Yiwen Guo, Irwin King
备注:Work in progress
链接:点击下载PDF文件
摘要:大型语言模型(LLM)最近获得了极大的关注,主要是因为它们在基于文本的交互中的能力。然而,自然的人类交互通常依赖于语音,因此需要转向基于语音的模型。实现这一目标的一种简单方法涉及“自动语音识别(ASR)+ LLM +文本到语音(TTS)”的管道,其中输入语音被转录为文本,由LLM处理,然后转换回语音。尽管是直接的,这种方法遭受固有的局限性,如模态转换过程中的信息丢失和三个阶段的误差积累。为了解决这些问题,语音语言模型(SpeechLM)-生成语音而不转换文本的端到端模型-已成为一种有前途的替代方案。本调查报告首次全面概述了构建SpeechLM的最新方法,详细介绍了其架构的关键组件以及对其开发不可或缺的各种培训配方。此外,我们系统地调查的各种能力的SpeechLMs,分类SpeechLMs的评价指标,并讨论在这个快速发展的领域的挑战和未来的研究方向。摘要:Large Language Models (LLMs) have recently garnered significant attention, primarily for their capabilities in text-based interactions. However, natural human interaction often relies on speech, necessitating a shift towards voice-based models. A straightforward approach to achieve this involves a pipeline of Automatic Speech Recognition (ASR) + LLM + Text-to-Speech (TTS)", where input speech is transcribed to text, processed by an LLM, and then converted back to speech. Despite being straightforward, this method suffers from inherent limitations, such as information loss during modality conversion and error accumulation across the three stages. To address these issues, Speech Language Models (SpeechLMs) -- end-to-end models that generate speech without converting from text -- have emerged as a promising alternative. This survey paper provides the first comprehensive overview of recent methodologies for constructing SpeechLMs, detailing the key components of their architecture and the various training recipes integral to their development. Additionally, we systematically survey the various capabilities of SpeechLMs, categorize the evaluation metrics for SpeechLMs, and discuss the challenges and future research directions in this rapidly evolving field.

【34】 Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
标题: 使用离散单元与从可控口音TTC合成的并行数据进行口音转换
作者: Tuan Nam Nguyen, Ngoc Quan Pham, Alexander Waibel
备注:Accepted at Syndata4genAI
链接:点击下载PDF文件
摘要:口音转换(AC)的目标是转换语音口音,同时保留内容和说话人身份。以前的方法要么在推理过程中需要参考话语,不能很好地保留说话者身份,要么使用只能针对每个非母语口音进行训练的一对一系统。本文提出了一个很有前途的AC模型,可以将许多口音转换为本地,以克服这些问题。我们的方法利用离散的单位,来自集群的自监督表示的母语,作为中介目标口音转换。利用多说话人文本到语音合成,它将这些离散表示转换回本地语音,同时保留说话人身份。此外,我们开发了一种有效的数据增强方法来训练系统,而不需要大量的非本地资源。我们的系统被证明可以提高非母语人士的流利性,听起来像一个本地口音,并保持原来的发言人身份很好。摘要:The goal of accent conversion (AC) is to convert speech accents while preserving content and speaker identity. Previous methods either required reference utterances during inference, did not preserve speaker identity well, or used one-to-one systems that could only be trained for each non-native accent. This paper presents a promising AC model that can convert many accents into native to overcome these issues. Our approach utilizes discrete units, derived from clustering self-supervised representations of native speech, as an intermediary target for accent conversion. Leveraging multi-speaker text-to-speech synthesis, it transforms these discrete representations back into native speech while retaining the speaker identity. Additionally, we develop an efficient data augmentation method to train the system without demanding a lot of non-native resources. Our system is proved to improve non-native speaker fluency, sound like a native accent, and preserve original speaker identity well.

【35】 FluentEditor+: Text-based Speech Editing by Modeling Local Hierarchical Acoustic Smoothness and Global Prosody Consistency
标题: FluentEditor+:通过建模局部分层声学平滑度和全局韵律一致性来进行基于文本的语音编辑
作者: Rui Liu, Jiatian Xi, Ziyue Jiang, Haizhou Li
备注:Work in progress
链接:点击下载PDF文件
摘要:基于文本的语音编辑(TSE)允许用户通过编辑相应的文本并执行剪切、复制和粘贴等操作来修改语音,以生成更新的音频,而无需直接更改原始录音。基于文本的语音编辑(TSE)允许用户通过编辑相应的文本并执行剪切、复制和粘贴等操作来修改语音,以生成更新的音频,而无需直接更改原始录音。虽然目前的TSE技术专注于最大限度地减少编辑片段内生成的语音和参考目标之间的差异,但它们往往忽视了在原始话语的上下文中保持局部和全局流畅性的重要性。此外,将编辑的片段与音频的未更改部分无缝集成仍然具有挑战性,通常需要文本到语音(TTS)系统的支持。本文介绍了一种新的方法,FluentEditor$ tiny +$,旨在克服这些限制。FluentEditor$ tiny +$采用高级特征提取技术来捕获声学和韵律特征,确保编辑和未编辑区域之间的流畅过渡。该模型确保了分段声学平滑性和全局韵律一致性,允许语音的无缝拼接,同时保持输出的连贯性和自然性。在VCTK和LibriTTS数据集上进行的大量实验表明,FluentEditor$ tiny +$在流畅性和韵律方面都优于现有的基于TTS的方法,包括Editspeech,Campnet,$A^3T$ FluentSpeech和Fluenteditor。消融研究进一步强调了每个模块对系统整体有效性的贡献。摘要:Text-based speech editing (TSE) allows users to modify speech by editing the corresponding text and performing operations such as cutting, copying, and pasting to generate updated audio without altering the original recording directly. Text-based speech editing (TSE) allows users to modify speech by editing the corresponding text and performing operations such as cutting, copying, and pasting to generate updated audio without altering the original recording directly. While current TSE techniques focus on minimizing discrepancies between generated speech and reference targets within edited segments, they often neglect the importance of maintaining both local and global fluency in the context of the original discourse. Additionally, seamlessly integrating edited segments with unaltered portions of the audio remains challenging, typically requiring support from text-to-speech (TTS) systems. This paper introduces a novel approach, FluentEditor$ tiny +$, designed to overcome these limitations. FluentEditor$ tiny +$ employs advanced feature extraction techniques to capture both acoustic and prosodic characteristics, ensuring fluent transitions between edited and unedited regions. The model ensures segmental acoustic smoothness and global prosody consistency, allowing seamless splicing of speech while preserving the coherence and naturalness of the output. Extensive experiments on the VCTK and LibriTTS datasets show that FluentEditor$ tiny +$ surpasses existing TTS-based methods, including Editspeech, Campnet, $A^3T$ FluentSpeech, and Fluenteditor, in both fluency and prosody. Ablation studies further highlight the contributions of each module to the overall effectiveness of the system.

【36】 A quest through interconnected datasets: lessons from highly-cited ICASSP papers
标题: 探索相互关联的数据集:备受引用的ICASP论文的教训
作者: Cynthia C. S. Liem, Doğa Taşcılar, Andrew M. Demetriou
备注:in Proceedings of the 21st International Conference on Content-based Multimedia Indexing, September 18-20 2024, Reykjavik, Iceland
链接:点击下载PDF文件
摘要:随着音频机器学习成果部署在具有社会影响力的应用程序中,重要的是要了解所使用数据的质量和来源。注意到在应用机器学习领域的学术出版中,明确这种意义并不是微不足道的奖励,也不包括在典型的应用机器学习课程中,我们在声学,语音和信号处理国际会议(ICASSP)上提出了一项与前5名引用论文相关的数据集使用研究。在这方面,我们对使用的数据集的来源进行了彻底的深度优先分析,通常导致搜索必须超出官方论文中报道的内容,并最终导致不清楚或纠缠的来源。特别是在当前对更大的、可能是生成性的人工智能模型的需求中,人们越来越意识到需要对数据来源进行问责。因此,我们呼吁社区不仅要专注于工程更大的模型,而且要为明确构建这些模型的基础创造更多的空间和奖励。摘要:As audio machine learning outcomes are deployed in societally impactful applications, it is important to have a sense of the quality and origins of the data used. Noticing that being explicit about this sense is not trivially rewarded in academic publishing in applied machine learning domains, and neither is included in typical applied machine learning curricula, we present a study into dataset usage connected to the top-5 cited papers at the International Conference on Acoustics, Speech, and Signal Processing (ICASSP). In this, we conduct thorough depth-first analyses towards origins of used datasets, often leading to searches that had to go beyond what was reported in official papers, and ending into unclear or entangled origins. Especially in the current pull towards larger, and possibly generative AI models, awareness of the need for accountability on data provenance is increasing. With this, we call on the community to not only focus on engineering larger models, but create more room and reward for explicitizing the foundations on which such models should be built.


机器翻译,仅供参考