今日论文合集:cs.SD语音9篇,eess.AS音频处理14篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 HeightCeleb -- an enrichment of VoxCeleb dataset with speaker height information
标题: HeightCeleb --包含说话者身高信息的VoxCeleb数据集的丰富内容
作者: Stanisław Kacprzak, Konrad Kowalczyk
备注:Accepted at IEEE SLT 2024
链接:点击下载PDF文件
摘要:说话人身高的预测对于语音取证、监控和自动说话人分析都有重要意义。到目前为止,TIMIT一直是训练和评估身高估计方法的最受欢迎的数据集。在本文中,我们介绍了HeightCeleb,VoxCeleb的扩展,VoxCeleb是说话人识别任务中常用的数据集。这种丰富包括添加有关VoxCeleb中所有1251个扬声器的高度的信息,这些信息是用自动化方法从公开来源提取的。这样的注释数据将使研究界能够利用免费提供的扬声器嵌入提取器,在VoxCeleb上进行预训练,以构建更有效的扬声器高度估计器。在这项工作中,我们描述了HeightCeleb数据集的创建,并表明使用它可以通过使用简单的统计回归方法和使用流行的扬声器模型获得的嵌入(无需任何额外的微调)在TIMIT测试集上实现最先进的结果。摘要:Prediction of speaker's height is of interest for voice forensics, surveillance, and automatic speaker profiling. Until now, TIMIT has been the most popular dataset for training and evaluation of the height estimation methods. In this paper, we introduce HeightCeleb, an extension to VoxCeleb, which is the dataset commonly used in speaker recognition tasks. This enrichment consists in adding information about the height of all 1251 speakers from VoxCeleb that has been extracted with an automated method from publicly available sources. Such annotated data will enable the research community to utilize freely available speaker embedding extractors, pre-trained on VoxCeleb, to build more efficient speaker height estimators. In this work, we describe the creation of the HeightCeleb dataset and show that using it enables to achieve state-of-the-art results on the TIMIT test set by using simple statistical regression methods and embeddings obtained with a popular speaker model (without any additional fine-tuning).

【2】 Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
标题: 通过自我监督学习特征的分段平均池增强语音情感识别
作者: Jonghwan Hyeon, Yung-Hwan Oh, Ho-Jin Choi
链接:点击下载PDF文件
摘要:语音情感识别(SER)分析通过语音表达的人类情感。自监督学习(SSL)通过从大量未标记的音频数据中学习有意义的表示,为SER提供了一种很有前途的方法。然而,现有的基于SSL的方法依赖于全局平均池(GAP)来表示音频信号,平等地对待语音和非语音段。这可能导致信息丰富的语音特征被不相关的非语音信息冲淡。为了解决这个问题,本文提出了分段平均池(SAP),它选择性地专注于信息语音段,而忽略非语音段。通过将GAP和SAP应用于SSL功能,我们的方法利用了GAP的整体语音信号信息和SAP的特定信息,从而提高了SER性能。实验表明,IEMOCAP的最先进的结果为英语和卓越的性能KEMDy19韩国数据集在未加权和加权的准确性。摘要:Speech Emotion Recognition (SER) analyzes human emotions expressed through speech. Self-supervised learning (SSL) offers a promising approach to SER by learning meaningful representations from a large amount of unlabeled audio data. However, existing SSL-based methods rely on Global Average Pooling (GAP) to represent audio signals, treating speech and non-speech segments equally. This can lead to dilution of informative speech features by irrelevant non-speech information. To address this, the paper proposes Segmental Average Pooling (SAP), which selectively focuses on informative speech segments while ignoring non-speech segments. By applying both GAP and SAP to SSL features, our approach utilizes overall speech signal information from GAP and specific information from SAP, leading to improved SER performance. Experiments show state-of-the-art results on the IEMOCAP for English and superior performance on KEMDy19 for Korean datasets in both unweighted and weighted accuracies.

【3】 SF-Speech: Straightened Flow for Zero-Shot Voice Clone on Small-Scale Dataset
标题: SF-Speech:小规模数据集中Zero-Shot语音克隆的简化流程
作者: Xuyuan Li, Zengqiang Shang, Hua Hua, Peiyang Shi, Chen Yang, Li Wang, Pengyuan Zhang
备注:Submitted to TASLP
链接:点击下载PDF文件
摘要:大规模语音生成模型在依赖于大规模数据集的zero-shot语音克隆任务中取得了令人瞩目的性能。然而,探索如何在小规模数据集上实现zero-shot语音克隆也是必要的。本文提出了SF语音,一种新的国家的最先进的语音克隆模型的基础上常微分方程和上下文学习。与以往的工作不同,SF-Speech采用多阶段生成策略来获得粗糙的声学特征,并利用该特征来拉直通过流匹配训练常微分方程模型而产生的弯曲反向轨迹。此外,我们发现了不同类型的声学特征的局部相关性之间的差异,并证明了2D卷积在建模梅尔频谱图特征中的潜在作用。在使用不到1000小时的语音进行训练后,SF-Speech的性能明显优于基于全局说话人嵌入或自回归大语言模型的方法。特别是,在相似的参数尺度下,SF Speech在语音可懂度(误词率相对降低22.4%)和音色相似度(余弦距离相对提高5.6%)方面也明显优于性能最好的常微分方程模型VoiceBox,当VoiceBox的参数增加到三倍时,SF Speech甚至保持了微弱的优势。摘要:Large-scale speech generation models have achieved impressive performance in the zero-shot voice clone tasks relying on large-scale datasets. However, exploring how to achieve zero-shot voice clone with small-scale datasets is also essential. This paper proposes SF-Speech, a novel state-of-the-art voice clone model based on ordinary differential equations and contextual learning. Unlike the previous works, SF-Speech employs a multi-stage generation strategy to obtain the coarse acoustic feature and utilizes this feature to straighten the curved reverse trajectories caused by training the ordinary differential equation model with flow matching. In addition, we find the difference between the local correlations of different types of acoustic features and demonstrate the potential role of 2D convolution in modeling mel-spectrogram features. After training with less than 1000 hours of speech, SF-Speech significantly outperforms those methods based on global speaker embedding or autoregressive large language models. In particular, SF-Speech also shows a significant advantage over VoiceBox, the best-performing ordinary differential equation model, in speech intelligibility (a relative decrease of 22.4 % on word error rate) and timbre similarity (a relative improvement of 5.6 % on cosine distance) at a similar scale of parameters, and even keep a slight advantage when the parameters of VoiceBox are tripled.

【4】 Learning to rumble: Automated elephant call classification, detection and endpointing using deep architectures
标题: 学习隆隆声:使用深度架构自动化大象呼叫分类、检测和端点定位
作者: Christiaan M. Geldenhuys, Thomas R. Niesler
链接:点击下载PDF文件
摘要:我们考虑的问题,检测,隔离和分类大象呼叫连续记录的音频。这种自动呼叫表征可以帮助保护工作,并为环境管理策略提供信息。在以前的工作中,呼叫检测是在一个段级进行,我们执行呼叫检测在帧级,这也隐含地允许呼叫端点,在一个较长的记录中的呼叫隔离。对于实验,我们采用两个注释的数据集,一个包含亚洲和其他非洲大象发声。我们评估了几个浅和深分类模型,并表明,目前最好的性能可以通过使用音频频谱图Transformer(AST),神经架构,尚未用于此目的之前,我们已经配置在一个新的序列到序列的方式。我们还表明,通过预训练使用迁移学习可以进一步提高计算复杂度和性能。最后,我们考虑子调用分类使用一个公认的分类的调用类型,一个任务,以前没有被考虑。我们还表明,在这种情况下,Transformer架构提供了最佳的性能。我们最好的分类器实现了平均精度(AP)为0.962帧的二进制呼叫分类,和0.957和0.979下的接收器操作特性(AUC)的呼叫分类与5类和子呼叫分类与7类分别的面积。所有这些都代表了新的基准(子呼叫分类)或对以前最好的系统的改进。我们的结论是,一个全自动的大象呼叫检测和子呼叫分类系统是触手可及的。这样一个系统将提供有关象群的行为和状况的宝贵信息,以便进行养护和管理。摘要:We consider the problem of detecting, isolating and classifying elephant calls in continuously recorded audio. Such automatic call characterisation can assist conservation efforts and inform environmental management strategies. In contrast to previous work in which call detection was performed at a segment level, we perform call detection at a frame level which implicitly also allows call endpointing, the isolation of a call in a longer recording. For experimentation, we employ two annotated datasets, one containing Asian and the other African elephant vocalisations. We evaluate several shallow and deep classifier models, and show that the current best performance can be improved by using an audio spectrogram transformer (AST), a neural architecture which has not been used for this purpose before, and which we have configured in a novel sequence-to-sequence manner. We also show that using transfer learning by pre-training leads to further improvements both in terms of computational complexity and performance. Finally, we consider sub-call classification using an accepted taxonomy of call types, a task which has not previously been considered. We show that also in this case the transformer architectures provide the best performance. Our best classifiers achieve an average precision (AP) of 0.962 for framewise binary call classification, and an area under the receiver operating characteristic (AUC) of 0.957 and 0.979 for call classification with 5 classes and sub-call classification with 7 classes respectively. All of these represent either new benchmarks (sub-call classifications) or improvements on previously best systems. We conclude that a fully-automated elephant call detection and subcall classification system is within reach. Such a system would provide valuable information on the behaviour and state of elephant herds for the purposes of conservation and management.

【5】 EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation
标题: 描述Caps:通过描述增强数据生成增强音频字幕
作者: Mithun Manivannan (1), Vignesh Nethrapalli (1), Mark Cartwright (1) ((1) New Jersey Institute of Technology)
链接:点击下载PDF文件
摘要:音频语言建模的最新进展,如自动音频字幕,得益于在大语言模型的帮助下生成的合成数据的训练。然而,这种用于环境声音字幕的方法主要集中在音频事件标签上,并且还没有探索利用可能存在于记录中的情感信息。在这项工作中,我们探索的好处,生成情感增强的合成音频字幕数据,指示ChatGPT与额外的声学信息的形式估计的音景情感。为此,我们引入了一个音频字幕数据集,它由大约120,000个音频片段组成,这些音频片段具有丰富的音景情感识别(SER)信息的成对合成描述。我们假设,这些额外的信息将产生更高质量的字幕,与音频记录的情感基调相匹配,这反过来又会提高使用这些数据训练的字幕模型的性能。我们通过客观和主观评估来测试这一假设,将使用AppropriationCaps数据集训练的模型与多个基线模型进行比较。我们的研究结果挑战了当前的字幕方法,并为开发和评估字幕模型提出了新的方向。摘要:Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have primarily focused on audio event tags and have not explored leveraging emotional information that may be present in recordings. In this work, we explore the benefit of generating emotion-augmented synthetic audio caption data by instructing ChatGPT with additional acoustic information in the form of estimated soundscape emotion. To do so, we introduce EmotionCaps, an audio captioning dataset comprised of approximately 120,000 audio clips with paired synthetic descriptions enriched with soundscape emotion recognition (SER) information. We hypothesize that this additional information will result in higher-quality captions that match the emotional tone of the audio recording, which will, in turn, improve the performance of captioning models trained with this data. We test this hypothesis through both objective and subjective evaluation, comparing models trained with the EmotionCaps dataset to multiple baseline models. Our findings challenge current approaches to captioning and suggest new directions for developing and assessing captioning models.

【6】 SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
标题: SeQuiFi:通过顺序类微调减轻语音情感识别中的灾难性遗忘
作者: Sarthak Jain, Orchid Chetia Phukan, Swarup Ranjan Behera, Arun Balaji Buduru, Rajesh Sharma
链接:点击下载PDF文件
摘要:在这项工作中,我们介绍了SeQuiFi,一种用于减轻语音情感识别(SER)中的灾难性遗忘(CF)的新方法。SeQuiFi采用了顺序类微调策略,其中模型每次在一个情感类上进行增量微调,保留并增强每个类的保留。虽然已经提出了各种最先进的(SOTA)方法,例如基于正则化的,基于内存的和加权平均技术来解决CF,但它仍然是一个挑战,特别是对于多样化和多语言数据集。通过大量的实验,我们证明了SeQuiFi在多个基准SER数据集(包括CREMA-D,RAVDESS,MESD-DB,MESD和SHEMO)的准确性和F1得分方面显著优于香草微调和SOTA持续学习技术,涵盖不同的语言。摘要:In this work, we introduce SeQuiFi, a novel approach for mitigating catastrophic forgetting (CF) in speech emotion recognition (SER). SeQuiFi adopts a sequential class-finetuning strategy, where the model is fine-tuned incrementally on one emotion class at a time, preserving and enhancing retention for each class. While various state-of-the-art (SOTA) methods, such as regularization-based, memory-based, and weight-averaging techniques, have been proposed to address CF, it still remains a challenge, particularly with diverse and multilingual datasets. Through extensive experiments, we demonstrate that SeQuiFi significantly outperforms both vanilla fine-tuning and SOTA continual learning techniques in terms of accuracy and F1 scores on multiple benchmark SER datasets, including CREMA-D, RAVDESS, Emo-DB, MESD, and SHEMO, covering different languages.

【7】 SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
标题: SiFiSinger:基于源过滤器模型的高保真端到端歌唱语音合成器
作者: Jianwei Cui, Yu Gu, Chao Weng, Jie Zhang, Liping Chen, Lirong Dai
备注:Accepted by ICASSP 2024, Synthesized audio samples are available at: this https URL
链接:点击下载PDF文件
摘要:本文提出了一种先进的端到端的歌唱声音合成(SVS)系统的基础上,源过滤器机制,直接翻译抒情和旋律线索到富有表现力和高保真的人类一样的歌唱。与VISinger 2类似,所提出的系统也利用了从VITS演变而来的训练范例,并结合了基本音高(F0)预测器和波形生成解码器等元素。为了解决这个问题,梅尔频谱图功能与F0信息的耦合可能会在F0预测过程中引入错误,我们考虑两种策略。首先,我们利用梅尔倒谱(mcep)功能,以解耦相互交织的梅尔频谱图和F0特性。其次,受神经源滤波器模型的启发,我们在SVS系统中引入源激励信号作为F0的表示,旨在更准确地捕捉基音细微差别。同时,采用可微mcep和F0损失作为波形解码器监督,以增强生成语音中语音包络和音调的预测准确性。在Opencpop数据集上的实验证明了该模型在合成质量和语调准确性方面的有效性。摘要:This paper presents an advanced end-to-end singing voice synthesis (SVS) system based on the source-filter mechanism that directly translates lyrical and melodic cues into expressive and high-fidelity human-like singing. Similarly to VISinger 2, the proposed system also utilizes training paradigms evolved from VITS and incorporates elements like the fundamental pitch (F0) predictor and waveform generation decoder. To address the issue that the coupling of mel-spectrogram features with F0 information may introduce errors during F0 prediction, we consider two strategies. Firstly, we leverage mel-cepstrum (mcep) features to decouple the intertwined mel-spectrogram and F0 characteristics. Secondly, inspired by the neural source-filter models, we introduce source excitation signals as the representation of F0 in the SVS system, aiming to capture pitch nuances more accurately. Meanwhile, differentiable mcep and F0 losses are employed as the waveform decoder supervision to fortify the prediction accuracy of speech envelope and pitch in the generated speech. Experiments on the Opencpop dataset demonstrate efficacy of the proposed model in synthesis quality and intonation accuracy.

【8】 FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
标题: Flash音频:用于快速、高保真文本到音频生成的纠正流
作者: Huadai Liu, Jialei Wang, Rongjie Huang, Yang Liu, Heng Lu, Wei Xue, Zhou Zhao
链接:点击下载PDF文件
摘要:潜在扩散模型(LDMs)的最新进展显着增强了文本到音频的生成,但其迭代采样过程施加了大量的计算需求,限制了实际部署。虽然最近的方法利用基于一致性的蒸馏旨在实现几步或单步推理,但它们的单步性能受到弯曲轨迹的限制,使它们无法超越传统的扩散模型。在这项工作中,我们介绍了FlashAudio整流流学习直接流快速模拟。为了改善低效率的时间步长分配和次优的噪声分布,FlashAudio使用双焦点采样器优化整流流的时间分布,并提出不混溶流以最小化批处理过孔分配中的数据-噪声对的总距离。此外,为了解决由无分类器制导(CFG)引起的放大累积误差,我们提出了锚定优化,通过将其锚定到参考轨迹来细化制导尺度。文本到音频生成的实验结果表明,FlashAudio的一步生成性能超过了基于扩散的模型,具有数百个采样步骤的音频质量,并使采样速度比单个NVIDIA 4090 Ti GPU上的实时速度快400倍。摘要:Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While recent methods utilizing consistency-based distillation aim to achieve few-step or single-step inference, their one-step performance is constrained by curved trajectories, preventing them from surpassing traditional diffusion models. In this work, we introduce FlashAudio with rectified flows to learn straight flow for fast simulation. To alleviate the inefficient timesteps allocation and suboptimal distribution of noise, FlashAudio optimizes the time distribution of rectified flow with Bifocal Samplers and proposes immiscible flow to minimize the total distance of data-noise pairs in a batch vias assignment. Furthermore, to address the amplified accumulation error caused by the classifier-free guidance (CFG), we propose Anchored Optimization, which refines the guidance scale by anchoring it to a reference trajectory. Experimental results on text-to-audio generation demonstrate that FlashAudio's one-step generation performance surpasses the diffusion-based models with hundreds of sampling steps on audio quality and enables a sampling speed of 400x faster than real-time on a single NVIDIA 4090Ti GPU.

【9】 Guided Speaker Embedding
标题: 引导演讲者嵌入
作者: Shota Horiguchi, Takafumi Moriya, Atsushi Ando, Takanori Ashihara, Hiroshi Sato, Naohiro Tawara, Marc Delcroix
链接:点击下载PDF文件
摘要:本文提出了一种引导式说话人嵌入提取系统,该系统以目标说话人和干扰说话人的语音活动为线索提取目标说话人的说话人嵌入。用于长格式重叠多扬声器音频处理的若干方法通常是两阶段的:i)分段级处理和ii)分段间扬声器匹配。扬声器嵌入通常用于后一目的。典型的说话人嵌入提取方法只使用单个说话人间隔,以避免嵌入与干扰说话人的语音破坏。然而,这通常使得说话人嵌入不可能提取,因为足够长的非重叠间隔并不总是可用的。在本文中,我们提出使用扬声器活动作为线索,直接从重叠语音中提取感兴趣的扬声器的嵌入。具体来说,我们将目标和非目标说话者的活动与声学特征连接起来,然后再输入模型。我们还条件的注意力权重用于池,使目标扬声器是不活跃的时间间隔的注意力权重为零。在说话人确认和说话人日志化中的实验结果表明了该方法的有效性。摘要:This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped multi-speaker audio processing are typically two-staged: i) segment-level processing and ii) inter-segment speaker matching. Speaker embeddings are often used for the latter purpose. Typical speaker embedding extraction approaches only use single-speaker intervals to avoid corrupting the embeddings with speech from interference speakers. However, this often makes speaker embeddings impossible to extract because sufficiently long non-overlapping intervals are not always available. In this paper, we propose using speaker activities as clues to extract the embedding of the speaker-of-interest directly from overlapping speech. Specifically, we concatenate the activity of target and non-target speakers to acoustic features before being fed to the model. We also condition the attention weights used for pooling so that the attention weights of the intervals in which the target speaker is inactive are zero. The effectiveness of the proposed method is demonstrated in speaker verification and speaker diarization.


eess.AS音频处理
【1】 SWIM: An Attention-Only Model for Speech Quality Assessment Under Subjective Variance
标题: SWIM:主观方差下的仅注意力语音质量评估模型
作者: Imran E Kibria, Donald S. Williamson
链接:点击下载PDF文件
摘要:语音质量最好通过使用平均意见分数(MOS)的人类反馈来评估。然而,收听者之间的评级的变化可能在话语的真实质量标签中引入噪声。目前,深度学习网络(包括卷积、递归和基于注意力的架构)已被用于质量估计。本文提出了一个专门的注意力为基础的模型,涉及一个Swin Transformer MOS估计(SWIM)。我们的网络捕获反映话语声学特性的局部和全局依赖性。为了抵消MOS标签中的主观方差,我们提出了一个正常的基于距离的目标,占每个标签的标准差,我们利用多级自学策略,以进一步提高泛化。我们的模型比现有的基于注意力的网络更紧凑,用于质量估计。最后,我们在Samsung Open Mean Opinion Score(SOMOS)数据集上的实验显示,在从头开始训练时,现有的基线模型有所改进。摘要:Speech quality is best evaluated by human feedback using mean opinion scores (MOS). However, variance in ratings between listeners can introduce noise in the true quality label of an utterance. Currently, deep learning networks including convolutional, recurrent, and attention-based architectures have been explored for quality estimation. This paper proposes an exclusively attention-based model involving a Swin Transformer for MOS estimation (SWIM). Our network captures local and global dependencies that reflect the acoustic properties of an utterance. To counteract subjective variance in MOS labels, we propose a normal distance-based objective that accounts for standard deviation in each label, and we avail a multistage self-teaching strategy to improve generalization further. Our model is significantly more compact than existing attention-based networks for quality estimation. Finally, our experiments on the Samsung Open Mean Opinion Score (SOMOS) dataset show improvement over existing baseline models when trained from scratch.

【2】 Beyond Speech and More: Investigating the Emergent Ability of Speech Foundation Models for Classifying Physiological Time-Series Signals
标题: 超越语音及更多:研究语音基础模型对生理时间序列信号进行分类的紧急能力
作者: Orchid Chetia Phukan, Swarup Ranjan Behera, Girish, Mohd Mujtaba Akhtar, Arun Balaji Buduru, Rajesh Sharma
链接:点击下载PDF文件
摘要:尽管只在语音数据上进行训练,但像Whisper这样的语音基础模型(SFM)在音频分类等非语音任务中表现出令人印象深刻的性能。这部分是因为语音与音频有一些共同的特征,使SFM能够有效地传输。在这项研究中,我们通过评估SFM对一个更具挑战性的域外(OOD)任务的边界:分类生理时间序列信号。我们测试两个关键假设:第一,SFM可以通过捕获共享的时间模式来概括生理信号;第二,多语言SFM将优于其他SFM,因为它们在预训练期间暴露于更大的可变性,从而导致更鲁棒的,概括的表示。我们使用ECG(心电图),EMG(肌电图)和EDA(皮肤电活动)信号进行压力识别的实验表明,在SFM衍生表示上训练的模型优于在原始生理信号上训练的模型。在所有的模型中,多语言的SFM实现了最高的准确性,支持我们的假设,并展示了他们的OOD能力。这项工作的立场SFMs作为有前途的工具,为新的未知领域以外的讲话。摘要:Despite being trained exclusively on speech data, speech foundation models (SFMs) like Whisper have shown impressive performance in non-speech tasks such as audio classification. This is partly because speech shares some common traits with audio, enabling SFMs to transfer effectively. In this study, we push the boundaries by evaluating SFMs on a more challenging out-of-domain (OOD) task: classifying physiological time-series signals. We test two key hypotheses: first, that SFMs can generalize to physiological signals by capturing shared temporal patterns; second, that multilingual SFMs will outperform others due to their exposure to greater variability during pre-training, leading to more robust, generalized representations. Our experiments, conducted for stress recognition using ECG (Electrocardiogram), EMG (Electromyography), and EDA (Electrodermal Activity) signals, reveal that models trained on SFM-derived representations outperform those trained on raw physiological signals. Among all models, multilingual SFMs achieve the highest accuracy, supporting our hypothesis and demonstrating their OOD capabilities. This work positions SFMs as promising tools for new uncharted domains beyond speech.

【3】 SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
标题: SeQuiFi:通过顺序类微调减轻语音情感识别中的灾难性遗忘
作者: Sarthak Jain, Orchid Chetia Phukan, Swarup Ranjan Behera, Arun Balaji Buduru, Rajesh Sharma
链接:点击下载PDF文件
摘要:在这项工作中,我们介绍了SeQuiFi,一种用于减轻语音情感识别(SER)中的灾难性遗忘(CF)的新方法。SeQuiFi采用了顺序类微调策略,其中模型每次在一个情感类上进行增量微调,保留并增强每个类的保留。虽然已经提出了各种最先进的(SOTA)方法,例如基于正则化的,基于内存的和加权平均技术来解决CF,但它仍然是一个挑战,特别是对于多样化和多语言数据集。通过大量的实验,我们证明了SeQuiFi在多个基准SER数据集(包括CREMA-D,RAVDESS,MESD-DB,MESD和SHEMO)的准确性和F1得分方面显著优于香草微调和SOTA持续学习技术,涵盖不同的语言。摘要:In this work, we introduce SeQuiFi, a novel approach for mitigating catastrophic forgetting (CF) in speech emotion recognition (SER). SeQuiFi adopts a sequential class-finetuning strategy, where the model is fine-tuned incrementally on one emotion class at a time, preserving and enhancing retention for each class. While various state-of-the-art (SOTA) methods, such as regularization-based, memory-based, and weight-averaging techniques, have been proposed to address CF, it still remains a challenge, particularly with diverse and multilingual datasets. Through extensive experiments, we demonstrate that SeQuiFi significantly outperforms both vanilla fine-tuning and SOTA continual learning techniques in terms of accuracy and F1 scores on multiple benchmark SER datasets, including CREMA-D, RAVDESS, Emo-DB, MESD, and SHEMO, covering different languages.

【4】 SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
标题: SiFiSinger:基于源过滤器模型的高保真端到端歌唱语音合成器
作者: Jianwei Cui, Yu Gu, Chao Weng, Jie Zhang, Liping Chen, Lirong Dai
备注:Accepted by ICASSP 2024, Synthesized audio samples are available at: this https URL
链接:点击下载PDF文件
摘要:本文提出了一种先进的端到端的歌唱声音合成(SVS)系统的基础上,源过滤器机制,直接翻译抒情和旋律线索到富有表现力和高保真的人类一样的歌唱。与VISinger 2类似,所提出的系统也利用了从VITS演变而来的训练范例,并结合了基本音高(F0)预测器和波形生成解码器等元素。为了解决这个问题,梅尔频谱图功能与F0信息的耦合可能会在F0预测过程中引入错误,我们考虑两种策略。首先,我们利用梅尔倒谱(mcep)功能,以解耦相互交织的梅尔频谱图和F0特性。其次,受神经源滤波器模型的启发,我们在SVS系统中引入源激励信号作为F0的表示,旨在更准确地捕捉基音细微差别。同时,采用可微mcep和F0损失作为波形解码器监督,以增强生成语音中语音包络和音调的预测准确性。在Opencpop数据集上的实验证明了该模型在合成质量和语调准确性方面的有效性。摘要:This paper presents an advanced end-to-end singing voice synthesis (SVS) system based on the source-filter mechanism that directly translates lyrical and melodic cues into expressive and high-fidelity human-like singing. Similarly to VISinger 2, the proposed system also utilizes training paradigms evolved from VITS and incorporates elements like the fundamental pitch (F0) predictor and waveform generation decoder. To address the issue that the coupling of mel-spectrogram features with F0 information may introduce errors during F0 prediction, we consider two strategies. Firstly, we leverage mel-cepstrum (mcep) features to decouple the intertwined mel-spectrogram and F0 characteristics. Secondly, inspired by the neural source-filter models, we introduce source excitation signals as the representation of F0 in the SVS system, aiming to capture pitch nuances more accurately. Meanwhile, differentiable mcep and F0 losses are employed as the waveform decoder supervision to fortify the prediction accuracy of speech envelope and pitch in the generated speech. Experiments on the Opencpop dataset demonstrate efficacy of the proposed model in synthesis quality and intonation accuracy.

【5】 ERVQ: Enhanced Residual Vector Quantization with Intra-and-Inter-Codebook Optimization for Neural Audio Codecs
标题: ERVQ:神经音频编解码器的增强残留量量化,采用码本内和码本间优化
作者: Rui-Chen Zheng, Hui-Peng Du, Xiao-Hang Jiang, Yang Ai, Zhen-Hua Ling
链接:点击下载PDF文件
摘要:当前的神经音频编解码器通常使用残差矢量量化(RVQ)来离散化语音信号。然而,它们经常经历码本崩溃,这会减少有效码本大小并导致性能不佳。为了解决这个问题,我们引入了ERVQ,增强的残差矢量量化,一种新的增强策略,用于神经音频编解码器中的RVQ框架。ERVQ通过码本内和码本间优化来减轻码本崩溃并提高编解码器性能。码本内优化结合了在线聚类策略和码平衡损失,以确保均衡和有效的码本利用率。码本间优化通过最小化连续量化之间的相似性来提高量化特征的多样性。我们的实验表明,ERVQ显着提高音频编解码器的性能在不同的模型,采样率和比特率,实现卓越的质量和泛化能力。它还在最先进的神经音频编解码器上实现了100%的码本利用率。进一步的实验表明,通过ERVQ策略改进的音频编解码器可以改进统一的语音和文本大语言模型(LLM)。具体地,在下游zero-shot文本到语音任务中生成的语音的自然度有显著的改善。音频样本可在此处获取。摘要:Current neural audio codecs typically use residual vector quantization (RVQ) to discretize speech signals. However, they often experience codebook collapse, which reduces the effective codebook size and leads to suboptimal performance. To address this problem, we introduce ERVQ, Enhanced Residual Vector Quantization, a novel enhancement strategy for the RVQ framework in neural audio codecs. ERVQ mitigates codebook collapse and boosts codec performance through both intra- and inter-codebook optimization. Intra-codebook optimization incorporates an online clustering strategy and a code balancing loss to ensure balanced and efficient codebook utilization. Inter-codebook optimization improves the diversity of quantized features by minimizing the similarity between successive quantizations. Our experiments show that ERVQ significantly enhances audio codec performance across different models, sampling rates, and bitrates, achieving superior quality and generalization capabilities. It also achieves 100% codebook utilization on one of the most advanced neural audio codecs. Further experiments indicate that audio codecs improved by the ERVQ strategy can improve unified speech-and-text large language models (LLMs). Specifically, there is a notable improvement in the naturalness of generated speech in downstream zero-shot text-to-speech tasks. Audio samples are available here.

【6】 Beyond Oversmoothing: Evaluating DDPM and MSE for Scalable Speech Synthesis in ASR
标题: 超越过度平滑:评估DDPM和SSE以实现ASB中的可扩展语音合成
作者: Christoph Minixhofer, Ondrej Klejch, Peter Bell
备注:Under review at ICASSP 2025
链接:点击下载PDF文件
摘要:人工合成的语音已经迅速接近人类的自然水平。然而,悖论仍然存在,即ASR系统在被人类判断为自然的TTS输出上进行训练时,在真实语音上的表现仍然很差。在这项工作中,我们探讨这种现象是否是由于TTS中常用的模型的过度平滑行为,特别关注TTS对ASR的行为,因为TTS训练数据的数量增加了。当用于ASR模型训练时,我们系统地比较了去噪扩散概率模型(DDPM)和基于均方误差(MSE)的TTS模型。我们测试的可扩展性的两种方法,不同的小时数,和不同的扬声器的数量。我们发现,对于一个给定的模型大小,DDPM可以更好地利用更多的数据,更多样化的扬声器,比MSE模型。我们实现了迄今为止真实和合成语音WER之间的最佳报告比率(1.46),但也发现仍然存在很大的差距。摘要:Synthetically generated speech has rapidly approached human levels of naturalness. However, the paradox remains that ASR systems, when trained on TTS output that is judged as natural by humans, continue to perform badly on real speech. In this work, we explore whether this phenomenon is due to the oversmoothing behaviour of models commonly used in TTS, with a particular focus on the behaviour of TTS-for-ASR as the amount of TTS training data is scaled up. We systematically compare Denoising Diffusion Probabilistic Models (DDPM) to Mean Squared Error (MSE) based models for TTS, when used for ASR model training. We test the scalability of the two approaches, varying both the number hours, and the number of different speakers. We find that for a given model size, DDPM can make better use of more data, and a more diverse set of speakers, than MSE models. We achieve the best reported ratio between real and synthetic speech WER to date (1.46), but also find that a large gap remains.

【7】 FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
标题: Flash音频:用于快速、高保真文本到音频生成的纠正流
作者: Huadai Liu, Jialei Wang, Rongjie Huang, Yang Liu, Heng Lu, Wei Xue, Zhou Zhao
链接:点击下载PDF文件
摘要:潜在扩散模型(LDMs)的最新进展显着增强了文本到音频的生成,但其迭代采样过程施加了大量的计算需求,限制了实际部署。虽然最近的方法利用基于一致性的蒸馏旨在实现几步或单步推理,但它们的单步性能受到弯曲轨迹的限制,使它们无法超越传统的扩散模型。在这项工作中,我们介绍了FlashAudio整流流学习直接流快速模拟。为了改善低效率的时间步长分配和次优的噪声分布,FlashAudio使用双焦点采样器优化整流流的时间分布,并提出不混溶流以最小化批处理过孔分配中的数据-噪声对的总距离。此外,为了解决由无分类器制导(CFG)引起的放大累积误差,我们提出了锚定优化,通过将其锚定到参考轨迹来细化制导尺度。文本到音频生成的实验结果表明,FlashAudio的一步生成性能超过了基于扩散的模型,具有数百个采样步骤的音频质量,并使采样速度比单个NVIDIA 4090 Ti GPU上的实时速度快400倍。摘要:Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While recent methods utilizing consistency-based distillation aim to achieve few-step or single-step inference, their one-step performance is constrained by curved trajectories, preventing them from surpassing traditional diffusion models. In this work, we introduce FlashAudio with rectified flows to learn straight flow for fast simulation. To alleviate the inefficient timesteps allocation and suboptimal distribution of noise, FlashAudio optimizes the time distribution of rectified flow with Bifocal Samplers and proposes immiscible flow to minimize the total distance of data-noise pairs in a batch vias assignment. Furthermore, to address the amplified accumulation error caused by the classifier-free guidance (CFG), we propose Anchored Optimization, which refines the guidance scale by anchoring it to a reference trajectory. Experimental results on text-to-audio generation demonstrate that FlashAudio's one-step generation performance surpasses the diffusion-based models with hundreds of sampling steps on audio quality and enables a sampling speed of 400x faster than real-time on a single NVIDIA 4090Ti GPU.

【8】 Guided Speaker Embedding
标题: 引导演讲者嵌入
作者: Shota Horiguchi, Takafumi Moriya, Atsushi Ando, Takanori Ashihara, Hiroshi Sato, Naohiro Tawara, Marc Delcroix
链接:点击下载PDF文件
摘要:本文提出了一种引导式说话人嵌入提取系统,该系统以目标说话人和干扰说话人的语音活动为线索提取目标说话人的说话人嵌入。用于长格式重叠多扬声器音频处理的若干方法通常是两阶段的:i)分段级处理和ii)分段间扬声器匹配。扬声器嵌入通常用于后一目的。典型的说话人嵌入提取方法只使用单个说话人间隔,以避免嵌入与干扰说话人的语音破坏。然而,这通常使得说话人嵌入不可能提取,因为足够长的非重叠间隔并不总是可用的。在本文中,我们提出使用扬声器活动作为线索,直接从重叠语音中提取感兴趣的扬声器的嵌入。具体来说,我们将目标和非目标说话者的活动与声学特征连接起来,然后再输入模型。我们还条件的注意力权重用于池,使目标扬声器是不活跃的时间间隔的注意力权重为零。在说话人确认和说话人日志化中的实验结果表明了该方法的有效性。摘要:This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped multi-speaker audio processing are typically two-staged: i) segment-level processing and ii) inter-segment speaker matching. Speaker embeddings are often used for the latter purpose. Typical speaker embedding extraction approaches only use single-speaker intervals to avoid corrupting the embeddings with speech from interference speakers. However, this often makes speaker embeddings impossible to extract because sufficiently long non-overlapping intervals are not always available. In this paper, we propose using speaker activities as clues to extract the embedding of the speaker-of-interest directly from overlapping speech. Specifically, we concatenate the activity of target and non-target speakers to acoustic features before being fed to the model. We also condition the attention weights used for pooling so that the attention weights of the intervals in which the target speaker is inactive are zero. The effectiveness of the proposed method is demonstrated in speaker verification and speaker diarization.

【9】 Automatic Screening for Children with Speech Disorder using Automatic Speech Recognition: Opportunities and Challenges
标题: 使用自动语音识别自动筛查言语障碍儿童:机遇与挑战
作者: Dancheng Liu, Jason Yang, Ishan Albrecht-Buehler, Helen Qin, Sophie Li, Yuting Hu, Amir Nassereldine, Jinjun Xiong
备注:AAAI-FSS 24
链接:点击下载PDF文件
摘要:言语是人类生活的一个基本方面,不仅对交流至关重要,而且对认知,社会和学术发展也至关重要。言语障碍儿童面临着重大挑战,如果不加以解决,可能会造成持久的负面影响。传统上,语音和语言评估(SLA)由熟练的语音语言病理学家(SLP)进行,但越来越需要由人工智能驱动的高效和可扩展的SLA方法。本立场文件对适用于自动化SLA管道的现有技术进行了调查,重点是针对儿童语音调整自动语音识别(ASR)模型,概述了当前SLA及其自动化对应物,以证明AI增强的SLA管道的可行性,并讨论了与AI驱动的SLA部署相关的实际考虑因素,包括可访问性和隐私问题。摘要:Speech is a fundamental aspect of human life, crucial not only for communication but also for cognitive, social, and academic development. Children with speech disorders (SD) face significant challenges that, if unaddressed, can result in lasting negative impacts. Traditionally, speech and language assessments (SLA) have been conducted by skilled speech-language pathologists (SLPs), but there is a growing need for efficient and scalable SLA methods powered by artificial intelligence. This position paper presents a survey of existing techniques suitable for automating SLA pipelines, with an emphasis on adapting automatic speech recognition (ASR) models for children's speech, an overview of current SLAs and their automated counterparts to demonstrate the feasibility of AI-enhanced SLA pipelines, and a discussion of practical considerations, including accessibility and privacy concerns, associated with the deployment of AI-powered SLAs.

【10】 HeightCeleb -- an enrichment of VoxCeleb dataset with speaker height information
标题: HeightCeleb --包含说话者身高信息的VoxCeleb数据集的丰富内容
作者: Stanisław Kacprzak, Konrad Kowalczyk
备注:Accepted at IEEE SLT 2024
链接:点击下载PDF文件
摘要:说话人身高的预测对于语音取证、监控和自动说话人分析都有重要意义。到目前为止,TIMIT一直是训练和评估身高估计方法的最受欢迎的数据集。在本文中,我们介绍了HeightCeleb,VoxCeleb的扩展,VoxCeleb是说话人识别任务中常用的数据集。这种丰富包括添加有关VoxCeleb中所有1251个扬声器的高度的信息,这些信息是用自动化方法从公开来源提取的。这样的注释数据将使研究界能够利用免费提供的扬声器嵌入提取器,在VoxCeleb上进行预训练,以构建更有效的扬声器高度估计器。在这项工作中,我们描述了HeightCeleb数据集的创建,并表明使用它可以通过使用简单的统计回归方法和使用流行的扬声器模型获得的嵌入(无需任何额外的微调)在TIMIT测试集上实现最先进的结果。摘要:Prediction of speaker's height is of interest for voice forensics, surveillance, and automatic speaker profiling. Until now, TIMIT has been the most popular dataset for training and evaluation of the height estimation methods. In this paper, we introduce HeightCeleb, an extension to VoxCeleb, which is the dataset commonly used in speaker recognition tasks. This enrichment consists in adding information about the height of all 1251 speakers from VoxCeleb that has been extracted with an automated method from publicly available sources. Such annotated data will enable the research community to utilize freely available speaker embedding extractors, pre-trained on VoxCeleb, to build more efficient speaker height estimators. In this work, we describe the creation of the HeightCeleb dataset and show that using it enables to achieve state-of-the-art results on the TIMIT test set by using simple statistical regression methods and embeddings obtained with a popular speaker model (without any additional fine-tuning).

【11】 Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
标题: 通过自我监督学习特征的分段平均池增强语音情感识别
作者: Jonghwan Hyeon, Yung-Hwan Oh, Ho-Jin Choi
链接:点击下载PDF文件
摘要:语音情感识别(SER)分析通过语音表达的人类情感。自监督学习(SSL)通过从大量未标记的音频数据中学习有意义的表示,为SER提供了一种很有前途的方法。然而,现有的基于SSL的方法依赖于全局平均池(GAP)来表示音频信号,平等地对待语音和非语音段。这可能导致信息丰富的语音特征被不相关的非语音信息冲淡。为了解决这个问题,本文提出了分段平均池(SAP),它选择性地专注于信息语音段,而忽略非语音段。通过将GAP和SAP应用于SSL功能,我们的方法利用了GAP的整体语音信号信息和SAP的特定信息,从而提高了SER性能。实验表明,IEMOCAP的最先进的结果为英语和卓越的性能KEMDy19韩国数据集在未加权和加权的准确性。摘要:Speech Emotion Recognition (SER) analyzes human emotions expressed through speech. Self-supervised learning (SSL) offers a promising approach to SER by learning meaningful representations from a large amount of unlabeled audio data. However, existing SSL-based methods rely on Global Average Pooling (GAP) to represent audio signals, treating speech and non-speech segments equally. This can lead to dilution of informative speech features by irrelevant non-speech information. To address this, the paper proposes Segmental Average Pooling (SAP), which selectively focuses on informative speech segments while ignoring non-speech segments. By applying both GAP and SAP to SSL features, our approach utilizes overall speech signal information from GAP and specific information from SAP, leading to improved SER performance. Experiments show state-of-the-art results on the IEMOCAP for English and superior performance on KEMDy19 for Korean datasets in both unweighted and weighted accuracies.

【12】 SF-Speech: Straightened Flow for Zero-Shot Voice Clone on Small-Scale Dataset
标题: SF-Speech:小规模数据集中Zero-Shot语音克隆的简化流程
作者: Xuyuan Li, Zengqiang Shang, Hua Hua, Peiyang Shi, Chen Yang, Li Wang, Pengyuan Zhang
备注:Submitted to TASLP
链接:点击下载PDF文件
摘要:大规模语音生成模型在依赖于大规模数据集的zero-shot语音克隆任务中取得了令人瞩目的性能。然而,探索如何在小规模数据集上实现zero-shot语音克隆也是必要的。本文提出了SF语音,一种新的国家的最先进的语音克隆模型的基础上常微分方程和上下文学习。与以前的作品不同,SF-Speech采用多阶段生成策略来获得粗糙的声学特征,并利用该特征来拉直由流匹配训练常微分方程模型引起的弯曲的反向轨迹。此外,我们发现了不同类型的声学特征的局部相关性之间的差异,并证明了2D卷积在建模梅尔频谱图特征中的潜在作用。在使用不到1000小时的语音进行训练后,SF-Speech的性能明显优于基于全局说话人嵌入或自回归大语言模型的方法。特别是,在相似的参数尺度下,SF Speech在语音可懂度(误词率相对降低22.4%)和音色相似度(余弦距离相对提高5.6%)方面也明显优于性能最好的常微分方程模型VoiceBox,当VoiceBox的参数增加到三倍时,SF Speech甚至保持了微弱的优势。摘要:Large-scale speech generation models have achieved impressive performance in the zero-shot voice clone tasks relying on large-scale datasets. However, exploring how to achieve zero-shot voice clone with small-scale datasets is also essential. This paper proposes SF-Speech, a novel state-of-the-art voice clone model based on ordinary differential equations and contextual learning. Unlike the previous works, SF-Speech employs a multi-stage generation strategy to obtain the coarse acoustic feature and utilizes this feature to straighten the curved reverse trajectories caused by training the ordinary differential equation model with flow matching. In addition, we find the difference between the local correlations of different types of acoustic features and demonstrate the potential role of 2D convolution in modeling mel-spectrogram features. After training with less than 1000 hours of speech, SF-Speech significantly outperforms those methods based on global speaker embedding or autoregressive large language models. In particular, SF-Speech also shows a significant advantage over VoiceBox, the best-performing ordinary differential equation model, in speech intelligibility (a relative decrease of 22.4 % on word error rate) and timbre similarity (a relative improvement of 5.6 % on cosine distance) at a similar scale of parameters, and even keep a slight advantage when the parameters of VoiceBox are tripled.

【13】 Learning to rumble: Automated elephant call classification, detection and endpointing using deep architectures
标题: 学习隆隆声:使用深度架构自动化大象呼叫分类、检测和端点定位
作者: Christiaan M. Geldenhuys, Thomas R. Niesler
链接:点击下载PDF文件
摘要:我们考虑的问题,检测,隔离和分类大象呼叫连续记录的音频。这种自动呼叫表征可以帮助保护工作,并为环境管理策略提供信息。在以前的工作中,呼叫检测是在一个段级进行,我们执行呼叫检测在帧级,这也隐含地允许呼叫端点,在一个较长的记录中的呼叫隔离。对于实验,我们采用两个注释的数据集,一个包含亚洲和其他非洲大象发声。我们评估了几个浅和深分类模型,并表明,目前最好的性能可以通过使用音频频谱图Transformer(AST),神经架构,尚未用于此目的之前,我们已经配置在一个新的序列到序列的方式。我们还表明,通过预训练使用迁移学习可以进一步提高计算复杂度和性能。最后,我们考虑子呼叫分类使用一个公认的分类呼叫类型,一个任务,以前没有被考虑。我们还表明,在这种情况下,Transformer架构提供了最佳的性能。我们最好的分类器实现了平均精度(AP)为0.962帧的二进制呼叫分类,和0.957和0.979下的接收器操作特性(AUC)的呼叫分类与5类和子呼叫分类与7类分别的面积。所有这些都代表了新的基准(子呼叫分类)或对以前最好的系统的改进。我们的结论是,一个全自动的大象呼叫检测和子呼叫分类系统是触手可及的。这样一个系统将提供有关象群的行为和状况的宝贵信息,以便进行养护和管理。摘要:We consider the problem of detecting, isolating and classifying elephant calls in continuously recorded audio. Such automatic call characterisation can assist conservation efforts and inform environmental management strategies. In contrast to previous work in which call detection was performed at a segment level, we perform call detection at a frame level which implicitly also allows call endpointing, the isolation of a call in a longer recording. For experimentation, we employ two annotated datasets, one containing Asian and the other African elephant vocalisations. We evaluate several shallow and deep classifier models, and show that the current best performance can be improved by using an audio spectrogram transformer (AST), a neural architecture which has not been used for this purpose before, and which we have configured in a novel sequence-to-sequence manner. We also show that using transfer learning by pre-training leads to further improvements both in terms of computational complexity and performance. Finally, we consider sub-call classification using an accepted taxonomy of call types, a task which has not previously been considered. We show that also in this case the transformer architectures provide the best performance. Our best classifiers achieve an average precision (AP) of 0.962 for framewise binary call classification, and an area under the receiver operating characteristic (AUC) of 0.957 and 0.979 for call classification with 5 classes and sub-call classification with 7 classes respectively. All of these represent either new benchmarks (sub-call classifications) or improvements on previously best systems. We conclude that a fully-automated elephant call detection and subcall classification system is within reach. Such a system would provide valuable information on the behaviour and state of elephant herds for the purposes of conservation and management.

【14】 EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation
标题: 描述Caps:通过描述增强数据生成增强音频字幕
作者: Mithun Manivannan (1), Vignesh Nethrapalli (1), Mark Cartwright (1) ((1) New Jersey Institute of Technology)
链接:点击下载PDF文件
摘要:音频语言建模的最新进展,如自动音频字幕,得益于在大语言模型的帮助下生成的合成数据的训练。然而,这种用于环境声音字幕的方法主要集中在音频事件标签上,并且还没有探索利用可能存在于记录中的情感信息。在这项工作中,我们探索的好处,生成情感增强的合成音频字幕数据,指示ChatGPT与额外的声学信息的形式估计的音景情感。为此,我们引入了一个音频字幕数据集,它由大约120,000个音频片段组成,这些音频片段具有丰富的音景情感识别(SER)信息的成对合成描述。我们假设,这些额外的信息将产生更高质量的字幕,与音频记录的情感基调相匹配,这反过来又会提高使用这些数据训练的字幕模型的性能。我们通过客观和主观评估来测试这一假设,将使用AppropriationCaps数据集训练的模型与多个基线模型进行比较。我们的研究结果挑战了当前的字幕方法,并为开发和评估字幕模型提出了新的方向。摘要:Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have primarily focused on audio event tags and have not explored leveraging emotional information that may be present in recordings. In this work, we explore the benefit of generating emotion-augmented synthetic audio caption data by instructing ChatGPT with additional acoustic information in the form of estimated soundscape emotion. To do so, we introduce EmotionCaps, an audio captioning dataset comprised of approximately 120,000 audio clips with paired synthetic descriptions enriched with soundscape emotion recognition (SER) information. We hypothesize that this additional information will result in higher-quality captions that match the emotional tone of the audio recording, which will, in turn, improve the performance of captioning models trained with this data. We test this hypothesis through both objective and subjective evaluation, comparing models trained with the EmotionCaps dataset to multiple baseline models. Our findings challenge current approaches to captioning and suggest new directions for developing and assessing captioning models.


机器翻译,仅供参考