本文经arXiv每日学术速递授权转载
标题: EMVD数据集:重金属中使用的极端声乐失真技术数据集
作者:Modan Tailleur,Julien Pinquier,Laurent Millot,Corsin Vogel,Mathieu Lagrange
Journal-ref:21st International Conference on Content-based Multimedia Indexing (CBMI), Gylfi {TH}{'o}r Gu{dh}mundsson; Laurent Amsaleg; Omar Shahbaz Khan; Ralph Gasser; Shin'ichi Satoh; Maria Pegia; Aladine Chetouani; Bj{"o}rn {TH}{'o}r J{'o}nsson; Claudio Gennaro; Ewa Kijak; Ilias Gialampoukidis; Liting Zhou; Jenny Benois-Pineau; Stevan Rudinac, Sep 2024, Reykjavik, Iceland
链接:点击下载PDF文件
摘要:在本文中,我们介绍了极端金属声乐数据集,其中包括一个极端的声乐技术在重金属音乐领域内执行的录音集合。该数据集由760段1秒至30秒长的音频摘录组成,总计约100分钟的音频材料,大致由60分钟的失真语音和40分钟的清晰语音录音组成。这些声乐录音来自27个不同的歌手,没有附带的乐器或后处理效果。该数据集中的失真分类包括四种不同的失真技术和三种声音效果,都在不同的音高范围内进行。针对与声乐技术相关的两种不同分类任务评估了最先进的深度学习模型的性能,展示了该资源在音频处理社区的潜力。摘要:In this paper, we introduce the Extreme Metal Vocals Dataset, which comprises a collection of recordings of extreme vocal techniques performed within the realm of heavy metal music. The dataset consists of 760 audio excerpts of 1 second to 30 seconds long, totaling about 100 min of audio material, roughly composed of 60 minutes of distorted voices and 40 minutes of clear voice recordings. These vocal recordings are from 27 different singers and are provided without accompanying musical instruments or post-processing effects. The distortion taxonomy within this dataset encompasses four distinct distortion techniques and three vocal effects, all performed in different pitch ranges. Performance of a state-of-the-art deep learning model is evaluated for two different classification tasks related to vocal techniques, demonstrating the potential of this resource for the audio processing community.
【2】 Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
标题: 空间语音转换:保留空间信息和非目标信号的语音转换
作者:Kentaro Seki,Shinnosuke Takamichi,Norihiro Takamune,Yuki Saito,Kanami Imamura,Hiroshi Saruwatari
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的任务称为空间语音转换,其目的是转换的目标语音,同时保留空间信息和非目标信号。传统的语音转换方法主要关注单声道波形,忽略了人耳听觉固有的立体声听觉体验。我们的基线方法通过集成盲源分离(BSS)、语音转换(VC)和空间混合来处理多通道波形,从而解决了这一差距。通过实验评估,我们组织并确定了这项任务中固有的关键挑战,例如保持音频质量和准确保留空间信息。我们的研究结果突出了平衡这些方面的根本困难,为未来的空间语音转换研究提供了一个基准。所提出的方法的代码是公开的,以鼓励在这一领域的进一步探索。摘要:This paper proposes a new task called spatial voice conversion, which aims to convert a target voice while preserving spatial information and non-target signals. Traditional voice conversion methods focus on single-channel waveforms, ignoring the stereo listening experience inherent in human hearing. Our baseline approach addresses this gap by integrating blind source separation (BSS), voice conversion (VC), and spatial mixing to handle multi-channel waveforms. Through experimental evaluations, we organize and identify the key challenges inherent in this task, such as maintaining audio quality and accurately preserving spatial information. Our results highlight the fundamental difficulties in balancing these aspects, providing a benchmark for future research in spatial voice conversion. The proposed method's code is publicly available to encourage further exploration in this domain.
【3】 SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
标题: SpecMaskGIT:音频频谱图的掩蔽生成建模,以实现高效音频合成及超越
作者:Marco Comunita,Zhi Zhong,Akira Takahashi,Shiqi Yang,Mengjie Zhao,Koichi Saito,Yukara Ikemiya. Takashi Shibuya,Shusuke Takahashi,Yuki Mitsufuji
备注:6 pages, 8 figures, 8 tables
链接:点击下载PDF文件
摘要:迭代合成音频片段的生成模型的最新进展为文本到音频合成(TTA)带来了巨大的成功,但合成速度慢,计算量大。虽然已经尝试加速迭代过程,但由于在推理阶段需要数百次迭代和大量的模型参数,高质量的TTA系统仍然效率低下。为了应对这些挑战,我们提出了SpecMaskGIT,这是一种轻量级的,高效的TTA模型,基于频谱图的掩蔽生成建模。首先,SpecMaskGIT通过不到16次迭代合成了一个逼真的10秒音频片段,比以前的迭代TTA方法少了一个数量级。作为一个离散模型,SpecMaskGIT在TTA基准测试中优于更大的VQ-Diffusion和自回归模型,同时仅需4个CPU内核即可实现实时,甚至在GPU上速度快30倍。接下来,建立在Mel频谱图的潜在空间上,SpecMaskGIT具有更广泛的应用(例如,zero-shot带宽扩展)。此外,我们解释SpecMaskGIT作为一个生成扩展到以前的歧视性音频掩蔽Transformers,并阐明其音频表示学习潜力。我们希望我们的工作能够激发对掩蔽音频建模的探索,以进一步多样化的场景。摘要:Recent advances in generative models that iteratively synthesize audio clips sparked great success to text-to-audio synthesis (TTA), but with the cost of slow synthesis speed and heavy computation. Although there have been attempts to accelerate the iterative procedure, high-quality TTA systems remain inefficient due to hundreds of iterations required in the inference phase and large amount of model parameters. To address the challenges, we propose SpecMaskGIT, a light-weighted, efficient yet effective TTA model based on the masked generative modeling of spectrograms. First, SpecMaskGIT synthesizes a realistic 10s audio clip by less than 16 iterations, an order-of-magnitude less than previous iterative TTA methods.As a discrete model, SpecMaskGIT outperforms larger VQ-Diffusion and auto-regressive models in the TTA benchmark, while being real-time with only 4 CPU cores or even 30x faster with a GPU. Next, built upon a latent space of Mel-spectrogram, SpecMaskGIT has a wider range of applications (e.g., the zero-shot bandwidth extension) than similar methods built on the latent wave domain. Moreover, we interpret SpecMaskGIT as a generative extension to previous discriminative audio masked Transformers, and shed light on its audio representation learning potential. We hope our work inspires the exploration of masked audio modeling toward further diverse scenarios.
【4】 This Paper Had the Smartest Reviewers -- Flattery Detection Utilising an Audio-Textual Transformer-Based Approach
标题: 本文拥有最聪明的评论者--利用基于音频文本转换器的方法进行奉承检测
作者:Lukas Christ,Shahin Amiriparian,Friederike Hawighorst,Ann-Kathrin Schill,Angelo Boutalikakis,Lorenz Graf-Vlachy,Andreas König,Björn W. Schuller
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:奉承是人类交流的一个重要方面,通过策略性的赞美和赞扬促进社会联系,塑造感知,影响行为,利用言语的力量有效地建立融洽关系。因此,它的自动检测可以增强人类与AI交互的自然性。为了满足这一需求,我们提出了一种新的音频文本数据集,包括20小时的语音和训练机器学习模型自动奉承检测。特别是,我们采用预先训练的AST,Wav2Vec2和Whisper模型的语音模态,和Whisper TTS模型结合的文本模态的RoberTa文本分类器。随后,我们建立了一个多模态分类器,结合文本和音频表示。对未见过的测试数据的评估表明了有希望的结果,在纯音频实验中,未加权平均召回分数达到82.46%,纯文本实验中为85.97%,使用多模态方法为87.16%。摘要:Flattery is an important aspect of human communication that facilitates social bonding, shapes perceptions, and influences behavior through strategic compliments and praise, leveraging the power of speech to build rapport effectively. Its automatic detection can thus enhance the naturalness of human-AI interactions. To meet this need, we present a novel audio textual dataset comprising 20 hours of speech and train machine learning models for automatic flattery detection. In particular, we employ pretrained AST, Wav2Vec2, and Whisper models for the speech modality, and Whisper TTS models combined with a RoBERTa text classifier for the textual modality. Subsequently, we build a multimodal classifier by combining text and audio representations. Evaluation on unseen test data demonstrates promising results, with Unweighted Average Recall scores reaching 82.46% in audio-only experiments, 85.97% in text-only experiments, and 87.16% using a multimodal approach.
【5】 Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights
标题: 探索大型多模式模型中特定于语音的风险:分类学、基准和见解
作者:Hao Yang,Lizhen Qu,Ehsan Shareghi,Gholamreza Haffari
链接:点击下载PDF文件
摘要:大型多模态模型(Large Multimodal Models,简称LMFM)在理解多模态信息和与人类用户交互方面取得了巨大的成功。尽管取得了进展,但在多模式环境中,特别是在语音模式中检测高风险互动的挑战在很大程度上仍未得到探索。关于语音模态风险的传统研究主要强调内容(例如,被捕获为转录的内容)。然而,在基于语音的交互中,音频中的非语言线索可以显著地改变话语背后的意图。在这项工作中,我们提出了一个语音特定的风险分类,涵盖了敌意(恶意讽刺和威胁),恶意模仿(年龄,性别,种族)和刻板偏见(年龄,性别,种族)下的8个风险类别。基于分类法,我们创建了一个小规模的数据集,用于评估当前LIFE在检测这些类别的风险方面的能力。我们观察到,即使是最新的模型在检测语音中的各种语言特异性风险(例如,Gemini 1.5 Pro的性能仅略高于随机基线)。警告:本文包含有偏见和攻击性的例子。摘要:Large Multimodal Models (LMMs) have achieved great success recently, demonstrating a strong capability to understand multimodal information and to interact with human users. Despite the progress made, the challenge of detecting high-risk interactions in multimodal settings, and in particular in speech modality, remains largely unexplored. Conventional research on risk for speech modality primarily emphasises the content (e.g., what is captured as transcription). However, in speech-based interactions, paralinguistic cues in audio can significantly alter the intended meaning behind utterances. In this work, we propose a speech-specific risk taxonomy, covering 8 risk categories under hostility (malicious sarcasm and threats), malicious imitation (age, gender, ethnicity), and stereotypical biases (age, gender, ethnicity). Based on the taxonomy, we create a small-scale dataset for evaluating current LMMs capability in detecting these categories of risk. We observe even the latest models remain ineffective to detect various paralinguistic-specific risks in speech (e.g., Gemini 1.5 Pro is performing only slightly above random baseline). Warning: this paper contains biased and offensive examples.
【6】 Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
标题: 用于合成语音检测的多头自注意中的时间通道建模
作者:Duc-Tuan Truong,Ruijie Tao,Tuan Nguyen,Hieu-Thi Luong,Kong Aik Lee,Eng Siong Chng
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:利用Transformer模型的最近的合成语音检测器与卷积神经网络对应物相比具有优越的性能。这种改进可能是由于Transformer模型中多头自注意力(MHSA)的强大建模能力,它学习每个输入令牌的时间关系。然而,人工合成语音可以位于特定区域的频率通道和时间段,而MHSA忽略了这种时间通道依赖性的输入序列。在这项工作中,我们提出了一个时间通道建模(TCM)模块,以提高MHSA的能力,捕捉时间通道的依赖关系。在ASVspoof 2021上的实验结果表明,仅需0.03M额外参数,TCM模块就可以在EER方面超过最先进的系统9.25%。进一步的消融研究表明,同时利用时间和信道信息产生最大的改善检测合成语音。摘要:Recent synthetic speech detectors leveraging the Transformer model have superior performance compared to the convolutional neural network counterparts. This improvement could be due to the powerful modeling ability of the multi-head self-attention (MHSA) in the Transformer model, which learns the temporal relationship of each input token. However, artifacts of synthetic speech can be located in specific regions of both frequency channels and temporal segments, while MHSA neglects this temporal-channel dependency of the input sequence. In this work, we proposed a Temporal-Channel Modeling (TCM) module to enhance MHSA's capability for capturing temporal-channel dependencies. Experimental results on the ASVspoof 2021 show that with only 0.03M additional parameters, the TCM module can outperform the state-of-the-art system by 9.25% in EER. Further ablation study reveals that utilizing both temporal and channel information yields the most improvement for detecting synthetic speech.
【7】 Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
标题: 利用合成音频数据进行端到端低资源语音翻译
作者:Yasmin Moslem
备注:IWSLT 2024
链接:点击下载PDF文件
摘要:本文介绍了我们的系统提交给国际会议上的口语翻译(IWITH2024)爱尔兰语到英语的语音翻译。我们基于Whisper构建了端到端系统,并采用了许多数据增强技术,如语音回译和噪声增强。我们研究了使用合成音频数据的效果,并讨论了几种方法来丰富信号的多样性。摘要:This paper describes our system submission to the International Conference on Spoken Language Translation (IWSLT 2024) for Irish-to-English speech translation. We built end-to-end systems based on Whisper, and employed a number of data augmentation techniques, such as speech back-translation and noise augmentation. We investigate the effect of using synthetic audio data and discuss several methods for enriching signal diversity.
【8】 Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
标题: 利用参数高效的迁移学习进行多语言文本到语音适应
作者:Yingting Li,Ambuj Mehrish,Bryan Chew,Bo Cheng,Soujanya Poria
链接:点击下载PDF文件
摘要:不同的语言有不同的语音系统,并在他们的韵律特征不同,这使得它具有挑战性的开发文本到语音(TTS)模型,可以有效地合成语音在多语言环境。此外,TTS架构需要足够高效,以捕捉多种语言的细微差别,并足够高效,以便于部署。标准方法是构建基于Transformer的模型,如SpeechT5,并在大型多语言数据集上对其进行训练。随着这些模型的大小的增长,用于适应这些模型的常规微调由于沉重的计算成本而变得不切实际。在本文中,我们提出了参数有效的迁移学习(PETL)方法,如适配器和超网络与TTS架构的多语言语音合成。值得注意的是,在我们的实验中,PETL方法能够实现与完全微调相比甚至更好的性能,只需2.5 %的可调参数。代码和示例可在:https: anonymous.4open.science r multilingualTTS-BA4C。摘要:Different languages have distinct phonetic systems and vary in their prosodic features making it challenging to develop a Text-to-Speech (TTS) model that can effectively synthesise speech in multilingual settings. Furthermore, TTS architecture needs to be both efficient enough to capture nuances in multiple languages and efficient enough to be practical for deployment. The standard approach is to build transformer based model such as SpeechT5 and train it on large multilingual dataset. As the size of these models grow the conventional fine-tuning for adapting these model becomes impractical due to heavy computational cost. In this paper, we proposes to integrate parameter-efficient transfer learning (PETL) methods such as adapters and hypernetwork with TTS architecture for multilingual speech synthesis. Notably, in our experiments PETL methods able to achieve comparable or even better performance compared to full fine-tuning with only $ sim$2.5 % tunable parameters.The code and samples are available at: https: anonymous.4open.science r multilingualTTS-BA4C.
【9】 Beyond Silence: Bias Analysis through Loss and Asymmetric Approach in Audio Anti-Spoofing
标题: 超越沉默:通过音频反欺骗中的丢失和不对称方法进行偏见分析
作者:Hye-jin Shim,Md Sahidullah,Jee-weon Jung,Shinji Watanabe,Tomi Kinnunen
备注:5 pages, 1 figure, 5 tables
链接:点击下载PDF文件
摘要:音频反欺骗检测研究的当前趋势是通过学习识别各种欺骗伪像来提高模型在看不见的攻击中的泛化能力。这个重点主要集中在恶搞类。最近,一些研究指出,沉默的分布在这两个阶层之间是不同的,这可以作为一个捷径。在本文中,我们扩展类明智的解释超越沉默。我们采用损失分析和非对称方法,从传统的以攻击为中心和以结果为导向的评估转向对模型行为的更深入检查。我们的调查突出了两个类之间的训练动态的显着差异,强调未来的研究需要专注于鲁棒建模的真正的类。摘要:Current trends in audio anti-spoofing detection research strive to improve models' ability to generalize across unseen attacks by learning to identify a variety of spoofing artifacts. This emphasis has primarily focused on the spoof class. Recently, several studies have noted that the distribution of silence differs between the two classes, which can serve as a shortcut. In this paper, we extend class-wise interpretations beyond silence. We employ loss analysis and asymmetric methodologies to move away from traditional attack-focused and result-oriented evaluations towards a deeper examination of model behaviors. Our investigations highlight the significant differences in training dynamics between the two classes, emphasizing the need for future research to focus on robust modeling of the bonafide class.
【10】 Self-Supervised Embeddings for Detecting Individual Symptoms of Depression
标题: 检测抑郁症个体症状的自我监督嵌入
作者:Sri Harsha Dumpala,Katerina Dikaios,Abraham Nunes,Frank Rudzicz,Rudolf Uher,Sageev Oore
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:抑郁症是一种影响全球数百万人的普遍心理健康疾病,需要可靠的评估系统。与以前的研究不同,我们的研究只关注检测抑郁症或预测其严重程度,我们的工作识别了抑郁症的个体症状,同时还使用语音输入预测其严重程度。我们利用基于自我监督学习(SSL)的语音模型来更好地利用在此任务中经常遇到的小规模数据集。我们的研究表明,显着的性能改善,利用SSL嵌入相比,传统的语音功能。我们比较了各种类型的SSL预训练模型,以阐明在识别不同症状方面贡献最大的语音信息类型(语义,说话者或韵律)。此外,我们还评估了结合多个SSL嵌入对性能的影响。此外,我们显示了多任务学习的意义,有效地识别抑郁症状。摘要:Depression, a prevalent mental health disorder impacting millions globally, demands reliable assessment systems. Unlike previous studies that focus solely on either detecting depression or predicting its severity, our work identifies individual symptoms of depression while also predicting its severity using speech input. We leverage self-supervised learning (SSL)-based speech models to better utilize the small-sized datasets that are frequently encountered in this task. Our study demonstrates notable performance improvements by utilizing SSL embeddings compared to conventional speech features. We compare various types of SSL pretrained models to elucidate the type of speech information (semantic, speaker, or prosodic) that contributes the most in identifying different symptoms. Additionally, we evaluate the impact of combining multiple SSL embeddings on performance. Furthermore, we show the significance of multi-task learning for identifying depressive symptoms effectively.
【11】 Sound Tagging in Infant-centric Home Soundscapes
标题: 以婴儿为中心的家庭声音场景中的声音标记
作者:Mohammad Nur Hossain Khan,Jialu Li,Nancy L. McElwain,Mark Hasegawa-Johnson,Bashima Islam
备注:Accepted in IEEEACM CHASE 2024
链接:点击下载PDF文件
摘要:某些环境噪音与婴幼儿的负面发育结果有关。虽然对家庭环境中的声音事件进行分类或标记是一个活跃的研究领域,但以前的研究主要集中在从放置在环境中的非固定麦克风或从成人的角度收集的数据。此外,这些工作中的许多忽略了环境中的婴儿或幼儿,或者仅从单个家庭收集数据,其中来自固定声源的噪声在婴儿的位置处可能是中等的,反之亦然。因此,尽管最近大型预训练模型在噪声事件检测方面取得了成功,但这些模型在家庭中以婴儿为中心的噪声音景上的性能还有待探索。为了弥补这一差距,我们以一种不引人注目的方式收集并标记了22个家庭的家庭音景中的噪音,这些数据是通过婴儿佩戴的录音设备收集的。在本文中,我们探索了一个大型预训练模型(音频频谱图Transformer [AST])在我们的以婴儿为中心的环境数据以及公开的家庭环境数据集上的性能。利用不同的训练策略,如重新训练,利用公共数据集,混合公共和以婴儿为中心的训练集,以及使用噪声和掩蔽的数据增强,我们评估了一个大型预训练模型在稀疏和不平衡的以婴儿为中心的数据上的性能。我们的研究结果表明,通过将我们收集的数据集与公共数据集相结合来微调大型预训练模型,(公共数据集)和0.76(收集的数据集)至0.84(组合数据集)和Cohen的Kappa值从0.013(公共数据集)和0.77(收集的数据集)到0.83(组合数据集),分别与仅使用公共或收集的数据集进行训练相比。摘要:Certain environmental noises have been associated with negative developmental outcomes for infants and young children. Though classifying or tagging sound events in a domestic environment is an active research area, previous studies focused on data collected from a non-stationary microphone placed in the environment or from the perspective of adults. Further, many of these works ignore infants or young children in the environment or have data collected from only a single family where noise from the fixed sound source can be moderate at the infant's position or vice versa. Thus, despite the recent success of large pre-trained models for noise event detection, the performance of these models on infant-centric noise soundscapes in the home is yet to be explored. To bridge this gap, we have collected and labeled noises in home soundscapes from 22 families in an unobtrusive manner, where the data are collected through an infant-worn recording device. In this paper, we explore the performance of a large pre-trained model (Audio Spectrogram Transformer [AST]) on our noise-conditioned infant-centric environmental data as well as publicly available home environmental datasets. Utilizing different training strategies such as resampling, utilizing public datasets, mixing public and infant-centric training sets, and data augmentation using noise and masking, we evaluate the performance of a large pre-trained model on sparse and imbalanced infant-centric data. Our results show that fine-tuning the large pre-trained model by combining our collected dataset with public datasets increases the F1-score from 0.11 (public datasets) and 0.76 (collected datasets) to 0.84 (combined datasets) and Cohen's Kappa from 0.013 (public datasets) and 0.77 (collected datasets) to 0.83 (combined datasets) compared to only training with public or collected datasets, respectively.
【12】 Investigating Confidence Estimation Measures for Speaker Diarization
标题: 研究说话人数字化的置信度估计措施
作者:Anurag Chowdhury,Abhinav Misra,Mark C. Fuhs,Monika Woszczyna
备注:Accepted in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:说话人日记系统根据说话人的身份对会话记录进行分段。这样的系统可能由于诸如语音模式变化、背景噪声和重叠语音之类的各种因素而对音频的一部分的说话者进行错误分类。这些错误会传播到依赖说话者身份的下游系统(如说话者自适应语音识别),并对其产生不利影响。减轻这些错误的方法之一是向下游系统提供片段级日志化置信度分数。在这项工作中,我们研究了多种方法来生成日记置信分数,包括那些来自原始日记系统和那些来自外部模型。我们在多个数据集和日志化系统上的实验表明,最具竞争力的置信度方法可以隔离出约30%的日志化错误,其中置信度最低的约为10%。摘要:Speaker diarization systems segment a conversation recording based on the speakers' identity. Such systems can misclassify the speaker of a portion of audio due to a variety of factors, such as speech pattern variation, background noise, and overlapping speech. These errors propagate to, and can adversely affect, downstream systems that rely on the speaker's identity, such as speaker-adapted speech recognition. One of the ways to mitigate these errors is to provide segment-level diarization confidence scores to downstream systems. In this work, we investigate multiple methods for generating diarization confidence scores, including those derived from the original diarization system and those derived from an external model. Our experiments across multiple datasets and diarization systems demonstrate that the most competitive confidence score methods can isolate ~30% of the diarization errors within segments with the lowest ~10% of confidence scores.
【13】 Sound Field Synthesis with Acoustic Waves
标题: 利用声波合成的声学场
作者:Mohamed F. Mansour
备注:16 pages, double-spaced, 5 figures
链接:点击下载PDF文件
摘要:基于声传播的物理学原理,提出了一种在小刚性表面上合成宽带声场的实用框架。声场被生成为两个组件的合成图:房间组件和设备组件;其中声学平面波作为核心元素。房间和设备组件的这种解耦显著降低了问题的复杂性,并提供了声场的准确呈现。我们详细描述了所提出的框架的理论基础,有效的程序,和工程应用。所提出的框架的有效性是通过在不同的环境设置下进行严格的验证。摘要:We propose a practical framework to synthesize broadband sound-field on a small rigid surface based on the physics of sound propagation. The sound-field is generated as a composite map of two components: room component and device component; with acoustic plane waves as core element. This decoupling of room and device components significantly reduces the problem complexity and provide accurate rendering of the sound-field. We describe in details the theoretical foundations, efficient procedures, and engineering applications of the proposed framework. The effectiveness of the proposed framework is established through rigorous validation under different environment setups.
【14】 Maximum Likelihood Estimation of the Direction of Sound In A Reverberant Noisy Environment
标题: 回响噪音环境中声音方向的最大似然估计
作者:Mohamed F. Mansour
备注:5 pages, 2 figures, conference
链接:点击下载PDF文件
摘要:我们描述了一种新的方法,估计在混响环境中的声音传播的基本原理的方向。该方法利用信噪比自适应的功能,从时间延迟和能量的方向分量后,声波分解的观测声场估计的视线方向在噪声和混响条件下。该方法的有效性建立与不同的麦克风阵列配置在各种使用场景下的实际数据。摘要:We describe a new method for estimating the direction of sound in a reverberant environment from basic principles of sound propagation. The method utilizes SNR-adaptive features from time-delay and energy of the directional components after acoustic wave decomposition of the observed sound field to estimate the line-of-sight direction under noisy and reverberant conditions. The effectiveness of the approach is established with real-data of different microphone array configurations under various usage scenarios.
【15】 AND: Audio Network Dissection for Interpreting Deep Acoustic
标题: AND:用于解释深层声学的音频网络剖析
作者:Tung-Yu Wu,Yu-Xiang Lin,Tsui-Wei Weng
备注:Accepted by ICML'24
链接:点击下载PDF文件
摘要:神经元层面的解释旨在通过研究神经元对特定感知或结构输入模式的反应来解释网络的行为和属性。虽然在视觉和语言领域有一些新兴的工作,但没有一个是针对声学模型的。为了弥补这一差距,我们引入了$ texttit {AND}$,这是第一个$ textbf{A}$udio $ textbf{N}$etwork $ textbf{D}$issection框架,它可以基于高响应音频自动建立声学神经元的自然语言解释。$ textit{AND}$的特点是使用LLM来总结音频之间的相互声学特征和身份。进行了大量的实验,以验证$ textit{AND}$的精确和翔实的描述。此外,我们展示了一个潜在的使用$ textit{AND}$音频机器unlearning进行概念特定的修剪生成的描述的基础上。最后,通过$ textit{AND}$的分析,我们强调了两个声学模型的行为:(i)模型通过基本声学特征的组合而不是高级抽象概念来区分音频;(ii)训练策略影响模型行为和神经元可解释性--监督训练引导神经元逐渐缩小注意力,而自我监督学习鼓励神经元成为多语义的,以探索高级特征。摘要:Neuron-level interpretations aim to explain network behaviors and properties by investigating neurons responsive to specific perceptual or structural input patterns. Although there is emerging work in the vision and language domains, none is explored for acoustic models. To bridge the gap, we introduce $ textit{AND}$, the first $ textbf{A}$udio $ textbf{N}$etwork $ textbf{D}$issection framework that automatically establishes natural language explanations of acoustic neurons based on highly-responsive audio. $ textit{AND}$ features the use of LLMs to summarize mutual acoustic features and identities among audio. Extensive experiments are conducted to verify $ textit{AND}$'s precise and informative descriptions. In addition, we demonstrate a potential use of $ textit{AND}$ for audio machine unlearning by conducting concept-specific pruning based on the generated descriptions. Finally, we highlight two acoustic model behaviors with analysis by $ textit{AND}$: (i) models discriminate audio with a combination of basic acoustic features rather than high-level abstract concepts; (ii) training strategies affect model behaviors and neuron interpretability -- supervised training guides neurons to gradually narrow their attention, while self-supervised learning encourages neurons to be polysemantic for exploring high-level features.
【16】 Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
标题: 建立端到端多语言自动歌词转录模型
作者:Jiawen Huang,Emmanouil Benetos
备注:Accepted at EUSIPCO 2024
链接:点击下载PDF文件
摘要:与多语言自动语音识别相比,多语言自动歌词转录(ALT)是一项具有挑战性的任务,这是由于标记数据的可用性有限以及唱歌带来的挑战。虽然最近发布了一些多语言歌唱数据集,但英语仍然在这些数据集中占主导地位。由于数据规模和注释质量的原因,多语言ALT仍然未得到充分探索。在本文中,我们的目标是创建一个多语言ALT系统与可用的数据集。受到已被证明有效的英语ALT架构的启发,我们通过扩展目标词汇集来适应这些技术的多语言场景。然后,我们评估的多语言模型的性能相比,其单语同行。此外,我们还探索了各种条件反射方法,将语言信息纳入模型。我们采用按语言分析,并将其与语言分类性能相结合。我们的研究结果表明,多语言模型的性能始终优于在语言子集上训练的单语模型。此外,我们证明,将语言信息显着提高性能。摘要:Multilingual automatic lyrics transcription (ALT) is a challenging task due to the limited availability of labelled data and the challenges introduced by singing, compared to multilingual automatic speech recognition. Although some multilingual singing datasets have been released recently, English continues to dominate these collections. Multilingual ALT remains underexplored due to the scale of data and annotation quality. In this paper, we aim to create a multilingual ALT system with available datasets. Inspired by architectures that have been proven effective for English ALT, we adapt these techniques to the multilingual scenario by expanding the target vocabulary set. We then evaluate the performance of the multilingual model in comparison to its monolingual counterparts. Additionally, we explore various conditioning methods to incorporate language information into the model. We apply analysis by language and combine it with the language classification performance. Our findings reveal that the multilingual model performs consistently better than the monolingual models trained on the language subsets. Furthermore, we demonstrate that incorporating language information significantly enhances performance.
【17】 Speaker-Independent Acoustic-to-Articulatory Inversion through Multi-Channel Attention Discriminator
标题: 通过多通道注意力鉴别器实现与扬声器无关的声学-关节翻转
作者:Woo-Jin Chung,Hong-Goo Kang
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:我们提出了一种新的扬声器独立的声学发音反转(AAI)模型,克服了传统的AAI模型,依赖于来自有限的数据集的声学特征中观察到的局限性。为了解决这些挑战,我们利用来自预训练的自监督学习(SSL)模型的表示,在AAI过程中更有效地估计电磁关节成像(EMA)信号中的全局,局部和运动模式信息。我们使用对抗方法训练我们的模型,并引入了一个基于注意力的多持续时间音素识别(MDPD),旨在充分捕捉多通道发音信号之间的复杂关系。我们的方法实现了皮尔逊相关系数为0.847,标志着最先进的性能在说话人独立的AAI模型。实现细节和代码可以在网上找到。摘要:We present a novel speaker-independent acoustic-to-articulatory inversion (AAI) model, overcoming the limitations observed in conventional AAI models that rely on acoustic features derived from restricted datasets. To address these challenges, we leverage representations from a pre-trained self-supervised learning (SSL) model to more effectively estimate the global, local, and kinematic pattern information in Electromagnetic Articulography (EMA) signals during the AAI process. We train our model using an adversarial approach and introduce an attention-based Multi-duration phoneme discriminator (MDPD) designed to fully capture the intricate relationship among multi-channel articulatory signals. Our method achieves a Pearson correlation coefficient of 0.847, marking state-of-the-art performance in speaker-independent AAI models. The implementation details and code can be found online.
【18】 Exploring compressibility of transformer based text-to-music (TTM) models
标题: 探索基于Transformer的文本到音乐(TTM)模型的可压缩性
作者:Vasileios Moschopoulos,Thanasis Kotsiopoulos,Pablo Peso Parada,Konstantinos Nikiforidis,Alexandros Stergiadis,Gerasimos Papakostas,Md Asif Jalal,Jisi Zhang,Anastasios Drosou,Karthikeyan Saravanan
备注:Proceedings of INTERSPEECH 2024
链接:点击下载PDF文件
摘要:最先进的文本到音乐(TTM)生成AI模型很大,需要桌面或服务器级计算,这使得它们无法在手机上部署。本文提出了一个模型压缩和生成性能的TTM模型之间的权衡分析。我们通过知识蒸馏和特定的修改来研究压缩,这些修改使TTM模型的各个组件(编码器,生成模型和解码器)都具有适用性。利用这些方法,我们创建了TinyTTM(89.2M参数),它在MusicBench数据集上实现了3.66的FAD和1.32的KL,优于MusicGen-Small(557.6M参数),但不低于MusicBench上微调的MusicGen-small。摘要:State-of-the art Text-To-Music (TTM) generative AI models are large and require desktop or server class compute, making them infeasible for deployment on mobile phones. This paper presents an analysis of trade-offs between model compression and generation performance of TTM models. We study compression through knowledge distillation and specific modifications that enable applicability over the various components of the TTM model (encoder, generative model and the decoder). Leveraging these methods we create TinyTTM (89.2M params) that achieves a FAD of 3.66 and KL of 1.32 on MusicBench dataset, better than MusicGen-Small (557.6M params) but not lower than MusicGen-small fine-tuned on MusicBench.
【19】 Rational-Exponent Filters with Applications to Generalized Auditory Filterbanks
标题: 一般指数过滤器及其在广义听觉过滤器组中的应用
作者:Samiya A Alkhairy
备注:14 pages, 9 figures, 2 tables, 32 equations. Submitted to IEEE TCAS-I
链接:点击下载PDF文件
摘要:我们提出了过滤器的有理指数,以提供一个连续的过滤器的行为是经典的实现。我们讨论了它们的稳定性,它们提供的灵活性,以及用于分析,设计和实现的各种表示。我们这样做是为了推广二阶滤波器,我们称之为有理指数广义听觉滤波器 滤波器组(GAF),适用于各种应用。我们提出了在时域和频域中的合理阶GAF的等效表示:传递函数,脉冲响应和积分表达式-其中最后一个允许有效的实时处理,而无需预处理要求。指数滤波器使滤波器特性能够处于连续体上,而不是将其限制为离散值,从而导致这些滤波器的行为具有更大的灵活性。在GAF的情况下,这允许具有任意连续而非离散的滤波器特性值,诸如(1)3dB品质因数与最大群延迟的比率-对于同时具有对频率选择性和同步的要求的滤波器组特别重要;以及(2)3dB与15 dB品质因数的比率,其指示频率响应幅度的形状。摘要:We present filters with rational exponents in order to provide a continuum of filter behavior not classically achievable. We discuss their stability, the flexibility they afford, and various representations useful for analysis, design and implementations. We do this for a generalization of second order filters which we refer to as rational-exponent Generalized Auditory Filters Filterbanks (GAFs) that are useful for a diverse array of applications. We present equivalent representations for rational-order GAFs in the time and frequency domains: transfer functions, impulse responses, and integral expressions - the last of which allows for efficient real-time processing without preprocessing requirements. Rational-exponent filters enable filter characteristics to be on a continuum rather than limiting them to discrete values thereby resulting in greater flexibility in the behavior of these filters. In the case of GAFs, this allows for having arbitrary continuous rather than discrete values for filter characteristics such as (1) the ratio of 3dB quality factor to maximum group delay - particularly important for filterbanks which have simultaneous requirements on frequency selectivity and synchronization; and (2) the ratio of 3dB to 15dB quality factors that dictates the shape of the frequency response magnitude.
eess.AS音频处理
【1】 Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model标题: 建立端到端多语言自动歌词转录模型
作者:Jiawen Huang,Emmanouil Benetos
备注:Accepted at EUSIPCO 2024
链接:点击下载PDF文件
摘要:与多语言自动语音识别相比,多语言自动歌词转录(ALT)是一项具有挑战性的任务,这是由于标记数据的可用性有限以及唱歌带来的挑战。虽然最近发布了一些多语言歌唱数据集,但英语仍然在这些数据集中占主导地位。由于数据规模和注释质量的原因,多语言ALT仍然未得到充分探索。在本文中,我们的目标是创建一个多语言ALT系统与可用的数据集。受到已被证明有效的英语ALT架构的启发,我们通过扩展目标词汇集来适应这些技术的多语言场景。然后,我们评估的多语言模型的性能相比,其单语同行。此外,我们还探索了各种条件反射方法,将语言信息纳入模型。我们采用按语言分析,并将其与语言分类性能相结合。我们的研究结果表明,多语言模型的性能始终优于在语言子集上训练的单语模型。此外,我们证明,将语言信息显着提高性能。摘要:Multilingual automatic lyrics transcription (ALT) is a challenging task due to the limited availability of labelled data and the challenges introduced by singing, compared to multilingual automatic speech recognition. Although some multilingual singing datasets have been released recently, English continues to dominate these collections. Multilingual ALT remains underexplored due to the scale of data and annotation quality. In this paper, we aim to create a multilingual ALT system with available datasets. Inspired by architectures that have been proven effective for English ALT, we adapt these techniques to the multilingual scenario by expanding the target vocabulary set. We then evaluate the performance of the multilingual model in comparison to its monolingual counterparts. Additionally, we explore various conditioning methods to incorporate language information into the model. We apply analysis by language and combine it with the language classification performance. Our findings reveal that the multilingual model performs consistently better than the monolingual models trained on the language subsets. Furthermore, we demonstrate that incorporating language information significantly enhances performance.
【2】 Speaker-Independent Acoustic-to-Articulatory Inversion through Multi-Channel Attention Discriminator
标题: 通过多通道注意力鉴别器实现与扬声器无关的声学-关节翻转
作者:Woo-Jin Chung,Hong-Goo Kang
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:我们提出了一种新的扬声器独立的声学发音反转(AAI)模型,克服了传统的AAI模型,依赖于来自有限的数据集的声学特征中观察到的局限性。为了解决这些挑战,我们利用来自预训练的自监督学习(SSL)模型的表示,在AAI过程中更有效地估计电磁关节成像(EMA)信号中的全局,局部和运动模式信息。我们使用对抗方法训练我们的模型,并引入了一个基于注意力的多持续时间音素识别(MDPD),旨在充分捕捉多通道发音信号之间的复杂关系。我们的方法实现了皮尔逊相关系数为0.847,标志着最先进的性能在说话人独立的AAI模型。实现细节和代码可以在网上找到。摘要:We present a novel speaker-independent acoustic-to-articulatory inversion (AAI) model, overcoming the limitations observed in conventional AAI models that rely on acoustic features derived from restricted datasets. To address these challenges, we leverage representations from a pre-trained self-supervised learning (SSL) model to more effectively estimate the global, local, and kinematic pattern information in Electromagnetic Articulography (EMA) signals during the AAI process. We train our model using an adversarial approach and introduce an attention-based Multi-duration phoneme discriminator (MDPD) designed to fully capture the intricate relationship among multi-channel articulatory signals. Our method achieves a Pearson correlation coefficient of 0.847, marking state-of-the-art performance in speaker-independent AAI models. The implementation details and code can be found online.
【3】 High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
标题: 使用令牌转换器和群屏蔽语言模型通过离散令牌实现高保真文本到语音
作者:Joun Yeop Lee,Myeonghun Jeong,Minchan Kim,Ji-Hyun Lee,Hoon-Young Cho,Nam Soo Kim
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
摘要:我们提出了一种新的两阶段的文本到语音(TTS)框架与两种类型的离散令牌,即,语义和声学标记,用于高保真语音合成。它有两个核心组件:解释模块,它将文本和语音提示处理成语义标记,专注于语言内容和对齐,以及说话模块,它捕获目标语音的音色,从语义标记生成声学标记,丰富语音重建。解释阶段采用转换器,因为它在将文本与语音对齐时具有鲁棒性。相比之下,Speaking阶段利用基于Conformer的架构与分组掩码语言模型(G-MLM)集成,以提高计算效率。实验结果表明,该结构在zero-shot场景下,语音质量和说话人相似度都优于传统模型。摘要:We propose a novel two-stage text-to-speech (TTS) framework with two types of discrete tokens, i.e., semantic and acoustic tokens, for high-fidelity speech synthesis. It features two core components: the Interpreting module, which processes text and a speech prompt into semantic tokens focusing on linguistic contents and alignment, and the Speaking module, which captures the timbre of the target voice to generate acoustic tokens from semantic tokens, enriching speech reconstruction. The Interpreting stage employs a transducer for its robustness in aligning text to speech. In contrast, the Speaking stage utilizes a Conformer-based architecture integrated with a Grouped Masked Language Model (G-MLM) to boost computational efficiency. Our experiments verify that this innovative structure surpasses the conventional models in the zero-shot scenario in terms of speech quality and speaker similarity.
【4】 AG-LSEC: Audio Grounded Lexical Speaker Error Correction
标题: AG-LSEC:音频接地词汇说话者错误纠正
作者:Rohit Paturi,Xiang Li,Sundararajan Srinivasan
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:说话者日记(SD)系统通常是基于音频的,并且独立于传统语音转录流水线中的ASR系统操作,并且可能由于SD和 或ASR协调而具有说话者错误,特别是在说话者转向和语音重叠区域周围。为了减少这些错误,词汇说话人错误校正(LSEC),其中外部语言模型提供词汇信息来纠正说话人错误,最近被提出。虽然该方法实现了很好的字拨号错误率(WDER)的改善,它不使用任何额外的声学信息,是容易出现错误校正。在本文中,我们提出了增强和声学地面的LSEC系统与扬声器的分数直接来自现有的SD管道。这种方法在基于音频的SD、ASR系统上实现了25-40%范围内的显著相对WDER降低,并且在RT 03-CTS、Callhome美式英语和Fisher数据集上相对于LSEC系统击败了15-25%。摘要:Speaker Diarization (SD) systems are typically audio-based and operate independently of the ASR system in traditional speech transcription pipelines and can have speaker errors due to SD and or ASR reconciliation, especially around speaker turns and regions of speech overlap. To reduce these errors, a Lexical Speaker Error Correction (LSEC), in which an external language model provides lexical information to correct the speaker errors, was recently proposed. Though the approach achieves good Word Diarization error rate (WDER) improvements, it does not use any additional acoustic information and is prone to miscorrections. In this paper, we propose to enhance and acoustically ground the LSEC system with speaker scores directly derived from the existing SD pipeline. This approach achieves significant relative WDER reductions in the range of 25-40% over the audio-based SD, ASR system and beats the LSEC system by 15-25% relative on RT03-CTS, Callhome American English and Fisher datasets.
【5】 Exploring compressibility of transformer based text-to-music (TTM) models
标题: 探索基于Transformer的文本到音乐(TTM)模型的可压缩性
作者:Vasileios Moschopoulos,Thanasis Kotsiopoulos,Pablo Peso Parada,Konstantinos Nikiforidis,Alexandros Stergiadis,Gerasimos Papakostas,Md Asif Jalal,Jisi Zhang,Anastasios Drosou,Karthikeyan Saravanan
备注:Proceedings of INTERSPEECH 2024
链接:点击下载PDF文件
摘要:最先进的文本到音乐(TTM)生成AI模型很大,需要桌面或服务器级计算,这使得它们无法在手机上部署。本文提出了一个模型压缩和生成性能的TTM模型之间的权衡分析。我们通过知识蒸馏和特定的修改来研究压缩,这些修改使TTM模型的各个组件(编码器,生成模型和解码器)都具有适用性。利用这些方法,我们创建了TinyTTM(89.2M参数),它在MusicBench数据集上实现了3.66的FAD和1.32的KL,优于MusicGen-Small(557.6M参数),但不低于MusicBench上微调的MusicGen-small。摘要:State-of-the art Text-To-Music (TTM) generative AI models are large and require desktop or server class compute, making them infeasible for deployment on mobile phones. This paper presents an analysis of trade-offs between model compression and generation performance of TTM models. We study compression through knowledge distillation and specific modifications that enable applicability over the various components of the TTM model (encoder, generative model and the decoder). Leveraging these methods we create TinyTTM (89.2M params) that achieves a FAD of 3.66 and KL of 1.32 on MusicBench dataset, better than MusicGen-Small (557.6M params) but not lower than MusicGen-small fine-tuned on MusicBench.
【6】 Rational-Exponent Filters with Applications to Generalized Auditory Filterbanks
标题: 一般指数过滤器及其在广义听觉过滤器组中的应用
作者:Samiya A Alkhairy
备注:14 pages, 9 figures, 2 tables, 32 equations. Submitted to IEEE TCAS-I
链接:点击下载PDF文件
摘要:我们提出了过滤器的有理指数,以提供一个连续的过滤器的行为是经典的实现。我们讨论了它们的稳定性,它们提供的灵活性,以及用于分析,设计和实现的各种表示。我们这样做是为了推广二阶滤波器,我们称之为有理指数广义听觉滤波器 滤波器组(GAF),适用于各种应用。我们提出了在时域和频域中的合理阶GAF的等效表示:传递函数,脉冲响应和积分表达式-其中最后一个允许有效的实时处理,而无需预处理要求。指数滤波器使滤波器特性能够处于连续体上,而不是将其限制为离散值,从而导致这些滤波器的行为具有更大的灵活性。在GAF的情况下,这允许具有任意连续而非离散的滤波器特性值,诸如(1)3dB品质因数与最大群延迟的比率-对于同时具有对频率选择性和同步的要求的滤波器组特别重要;以及(2)3dB与15 dB品质因数的比率,其指示频率响应幅度的形状。摘要:We present filters with rational exponents in order to provide a continuum of filter behavior not classically achievable. We discuss their stability, the flexibility they afford, and various representations useful for analysis, design and implementations. We do this for a generalization of second order filters which we refer to as rational-exponent Generalized Auditory Filters Filterbanks (GAFs) that are useful for a diverse array of applications. We present equivalent representations for rational-order GAFs in the time and frequency domains: transfer functions, impulse responses, and integral expressions - the last of which allows for efficient real-time processing without preprocessing requirements. Rational-exponent filters enable filter characteristics to be on a continuum rather than limiting them to discrete values thereby resulting in greater flexibility in the behavior of these filters. In the case of GAFs, this allows for having arbitrary continuous rather than discrete values for filter characteristics such as (1) the ratio of 3dB quality factor to maximum group delay - particularly important for filterbanks which have simultaneous requirements on frequency selectivity and synchronization; and (2) the ratio of 3dB to 15dB quality factors that dictates the shape of the frequency response magnitude.
【7】 Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
标题: 空间语音转换:保留空间信息和非目标信号的语音转换
作者:Kentaro Seki,Shinnosuke Takamichi,Norihiro Takamune,Yuki Saito,Kanami Imamura,Hiroshi Saruwatari
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的任务称为空间语音转换,其目的是转换的目标语音,同时保留空间信息和非目标信号。传统的语音转换方法主要关注单声道波形,忽略了人耳听觉固有的立体声听觉体验。我们的基线方法通过集成盲源分离(BSS)、语音转换(VC)和空间混合来处理多通道波形,从而解决了这一差距。通过实验评估,我们组织并确定了这项任务中固有的关键挑战,例如保持音频质量和准确保留空间信息。我们的研究结果突出了平衡这些方面的根本困难,为未来的空间语音转换研究提供了一个基准。所提出的方法的代码是公开的,以鼓励在这一领域的进一步探索。摘要:This paper proposes a new task called spatial voice conversion, which aims to convert a target voice while preserving spatial information and non-target signals. Traditional voice conversion methods focus on single-channel waveforms, ignoring the stereo listening experience inherent in human hearing. Our baseline approach addresses this gap by integrating blind source separation (BSS), voice conversion (VC), and spatial mixing to handle multi-channel waveforms. Through experimental evaluations, we organize and identify the key challenges inherent in this task, such as maintaining audio quality and accurately preserving spatial information. Our results highlight the fundamental difficulties in balancing these aspects, providing a benchmark for future research in spatial voice conversion. The proposed method's code is publicly available to encourage further exploration in this domain.
【8】 SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
标题: SpecMaskGIT:音频频谱图的掩蔽生成建模,以实现高效音频合成及超越
作者:Marco Comunita,Zhi Zhong,Akira Takahashi,Shiqi Yang,Mengjie Zhao,Koichi Saito,Yukara Ikemiya. Takashi Shibuya,Shusuke Takahashi,Yuki Mitsufuji
备注:6 pages, 8 figures, 8 tables
链接:点击下载PDF文件
摘要:迭代合成音频片段的生成模型的最新进展为文本到音频合成(TTA)带来了巨大的成功,但合成速度慢,计算量大。虽然已经尝试加速迭代过程,但由于在推理阶段需要数百次迭代和大量的模型参数,高质量的TTA系统仍然效率低下。为了应对这些挑战,我们提出了SpecMaskGIT,这是一种轻量级的,高效的TTA模型,基于频谱图的掩蔽生成建模。首先,SpecMaskGIT通过不到16次迭代合成了一个逼真的10秒音频片段,比以前的迭代TTA方法少了一个数量级。作为一个离散模型,SpecMaskGIT在TTA基准测试中优于更大的VQ-Diffusion和自回归模型,同时仅需4个CPU内核即可实现实时,甚至在GPU上速度快30倍。接下来,建立在Mel频谱图的潜在空间上,SpecMaskGIT具有更广泛的应用(例如,zero-shot带宽扩展)。此外,我们解释SpecMaskGIT作为一个生成扩展到以前的歧视性音频掩蔽Transformers,并阐明其音频表示学习潜力。我们希望我们的工作能够激发对掩蔽音频建模的探索,以进一步多样化的场景。摘要:Recent advances in generative models that iteratively synthesize audio clips sparked great success to text-to-audio synthesis (TTA), but with the cost of slow synthesis speed and heavy computation. Although there have been attempts to accelerate the iterative procedure, high-quality TTA systems remain inefficient due to hundreds of iterations required in the inference phase and large amount of model parameters. To address the challenges, we propose SpecMaskGIT, a light-weighted, efficient yet effective TTA model based on the masked generative modeling of spectrograms. First, SpecMaskGIT synthesizes a realistic 10s audio clip by less than 16 iterations, an order-of-magnitude less than previous iterative TTA methods.As a discrete model, SpecMaskGIT outperforms larger VQ-Diffusion and auto-regressive models in the TTA benchmark, while being real-time with only 4 CPU cores or even 30x faster with a GPU. Next, built upon a latent space of Mel-spectrogram, SpecMaskGIT has a wider range of applications (e.g., the zero-shot bandwidth extension) than similar methods built on the latent wave domain. Moreover, we interpret SpecMaskGIT as a generative extension to previous discriminative audio masked Transformers, and shed light on its audio representation learning potential. We hope our work inspires the exploration of masked audio modeling toward further diverse scenarios.
【9】 This Paper Had the Smartest Reviewers -- Flattery Detection Utilising an Audio-Textual Transformer-Based Approach
标题: 本文拥有最聪明的评论者--利用基于音频文本转换器的方法进行奉承检测
作者:Lukas Christ,Shahin Amiriparian,Friederike Hawighorst,Ann-Kathrin Schill,Angelo Boutalikakis,Lorenz Graf-Vlachy,Andreas König,Björn W. Schuller
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:奉承是人类交流的一个重要方面,通过策略性的赞美和赞扬促进社会联系,塑造感知,影响行为,利用言语的力量有效地建立融洽关系。因此,它的自动检测可以增强人类与AI交互的自然性。为了满足这一需求,我们提出了一种新的音频文本数据集,包括20小时的语音和训练机器学习模型自动奉承检测。特别是,我们采用预先训练的AST,Wav2Vec2和Whisper模型的语音模态,和Whisper TTS模型结合的文本模态的RoberTa文本分类器。随后,我们建立了一个多模态分类器,结合文本和音频表示。对未见过的测试数据的评估表明了有希望的结果,在纯音频实验中,未加权平均召回分数达到82.46%,纯文本实验中为85.97%,使用多模态方法为87.16%。摘要:Flattery is an important aspect of human communication that facilitates social bonding, shapes perceptions, and influences behavior through strategic compliments and praise, leveraging the power of speech to build rapport effectively. Its automatic detection can thus enhance the naturalness of human-AI interactions. To meet this need, we present a novel audio textual dataset comprising 20 hours of speech and train machine learning models for automatic flattery detection. In particular, we employ pretrained AST, Wav2Vec2, and Whisper models for the speech modality, and Whisper TTS models combined with a RoBERTa text classifier for the textual modality. Subsequently, we build a multimodal classifier by combining text and audio representations. Evaluation on unseen test data demonstrates promising results, with Unweighted Average Recall scores reaching 82.46% in audio-only experiments, 85.97% in text-only experiments, and 87.16% using a multimodal approach.
【10】 Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights
标题: 探索大型多模式模型中特定于语音的风险:分类学、基准和见解
作者:Hao Yang,Lizhen Qu,Ehsan Shareghi,Gholamreza Haffari
链接:点击下载PDF文件
摘要:大型多模态模型(Large Multimodal Models,简称LMFM)在理解多模态信息和与人类用户交互方面取得了巨大的成功。尽管取得了进展,但在多模式环境中,特别是在语音模式中检测高风险互动的挑战在很大程度上仍未得到探索。关于语音模态风险的传统研究主要强调内容(例如,被捕获为转录的内容)。然而,在基于语音的交互中,音频中的非语言线索可以显著地改变话语背后的意图。在这项工作中,我们提出了一个语音特定的风险分类,涵盖了敌意(恶意讽刺和威胁),恶意模仿(年龄,性别,种族)和刻板偏见(年龄,性别,种族)下的8个风险类别。基于分类法,我们创建了一个小规模的数据集,用于评估当前LIFE在检测这些类别的风险方面的能力。我们观察到,即使是最新的模型在检测语音中的各种语言特异性风险(例如,Gemini 1.5 Pro的性能仅略高于随机基线)。警告:本文包含有偏见和攻击性的例子。摘要:Large Multimodal Models (LMMs) have achieved great success recently, demonstrating a strong capability to understand multimodal information and to interact with human users. Despite the progress made, the challenge of detecting high-risk interactions in multimodal settings, and in particular in speech modality, remains largely unexplored. Conventional research on risk for speech modality primarily emphasises the content (e.g., what is captured as transcription). However, in speech-based interactions, paralinguistic cues in audio can significantly alter the intended meaning behind utterances. In this work, we propose a speech-specific risk taxonomy, covering 8 risk categories under hostility (malicious sarcasm and threats), malicious imitation (age, gender, ethnicity), and stereotypical biases (age, gender, ethnicity). Based on the taxonomy, we create a small-scale dataset for evaluating current LMMs capability in detecting these categories of risk. We observe even the latest models remain ineffective to detect various paralinguistic-specific risks in speech (e.g., Gemini 1.5 Pro is performing only slightly above random baseline). Warning: this paper contains biased and offensive examples.
【11】 Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
标题: 用于合成语音检测的多头自注意中的时间通道建模
作者:Duc-Tuan Truong,Ruijie Tao,Tuan Nguyen,Hieu-Thi Luong,Kong Aik Lee,Eng Siong Chng
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:利用Transformer模型的最近的合成语音检测器与卷积神经网络对应物相比具有优越的性能。这种改进可能是由于Transformer模型中多头自注意力(MHSA)的强大建模能力,它学习每个输入令牌的时间关系。然而,人工合成语音可以位于特定区域的频率通道和时间段,而MHSA忽略了这种时间通道依赖性的输入序列。在这项工作中,我们提出了一个时间通道建模(TCM)模块,以提高MHSA的能力,捕捉时间通道的依赖关系。在ASVspoof 2021上的实验结果表明,仅需0.03M额外参数,TCM模块就可以在EER方面超过最先进的系统9.25%。进一步的消融研究表明,同时利用时间和信道信息产生最大的改善检测合成语音。摘要:Recent synthetic speech detectors leveraging the Transformer model have superior performance compared to the convolutional neural network counterparts. This improvement could be due to the powerful modeling ability of the multi-head self-attention (MHSA) in the Transformer model, which learns the temporal relationship of each input token. However, artifacts of synthetic speech can be located in specific regions of both frequency channels and temporal segments, while MHSA neglects this temporal-channel dependency of the input sequence. In this work, we proposed a Temporal-Channel Modeling (TCM) module to enhance MHSA's capability for capturing temporal-channel dependencies. Experimental results on the ASVspoof 2021 show that with only 0.03M additional parameters, the TCM module can outperform the state-of-the-art system by 9.25% in EER. Further ablation study reveals that utilizing both temporal and channel information yields the most improvement for detecting synthetic speech.
【12】 Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
标题: 利用合成音频数据进行端到端低资源语音翻译
作者:Yasmin Moslem
备注:IWSLT 2024
链接:点击下载PDF文件
摘要:本文介绍了我们的系统提交给国际会议上的口语翻译(IWITH2024)爱尔兰语到英语的语音翻译。我们基于Whisper构建了端到端系统,并采用了许多数据增强技术,如语音回译和噪声增强。我们研究了使用合成音频数据的效果,并讨论了几种方法来丰富信号的多样性。摘要:This paper describes our system submission to the International Conference on Spoken Language Translation (IWSLT 2024) for Irish-to-English speech translation. We built end-to-end systems based on Whisper, and employed a number of data augmentation techniques, such as speech back-translation and noise augmentation. We investigate the effect of using synthetic audio data and discuss several methods for enriching signal diversity.
【13】 Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
标题: 利用参数高效的迁移学习进行多语言文本到语音适应
作者:Yingting Li,Ambuj Mehrish,Bryan Chew,Bo Cheng,Soujanya Poria
链接:点击下载PDF文件
摘要:不同的语言有不同的语音系统,并在他们的韵律特征不同,这使得它具有挑战性的开发文本到语音(TTS)模型,可以有效地合成语音在多语言环境。此外,TTS架构需要足够高效,以捕捉多种语言的细微差别,并足够高效,以便于部署。标准方法是构建基于Transformer的模型,如SpeechT5,并在大型多语言数据集上对其进行训练。随着这些模型的大小的增长,用于适应这些模型的常规微调由于沉重的计算成本而变得不切实际。在本文中,我们提出了参数有效的迁移学习(PETL)方法,如适配器和超网络与TTS架构的多语言语音合成。值得注意的是,在我们的实验中,PETL方法能够实现与完全微调相比甚至更好的性能,只需2.5 %的可调参数。代码和示例可在:https: anonymous.4open.science r multilingualTTS-BA4C。摘要:Different languages have distinct phonetic systems and vary in their prosodic features making it challenging to develop a Text-to-Speech (TTS) model that can effectively synthesise speech in multilingual settings. Furthermore, TTS architecture needs to be both efficient enough to capture nuances in multiple languages and efficient enough to be practical for deployment. The standard approach is to build transformer based model such as SpeechT5 and train it on large multilingual dataset. As the size of these models grow the conventional fine-tuning for adapting these model becomes impractical due to heavy computational cost. In this paper, we proposes to integrate parameter-efficient transfer learning (PETL) methods such as adapters and hypernetwork with TTS architecture for multilingual speech synthesis. Notably, in our experiments PETL methods able to achieve comparable or even better performance compared to full fine-tuning with only $ sim$2.5 % tunable parameters.The code and samples are available at: https: anonymous.4open.science r multilingualTTS-BA4C.
【14】 Beyond Silence: Bias Analysis through Loss and Asymmetric Approach in Audio Anti-Spoofing
标题: 超越沉默:通过音频反欺骗中的丢失和不对称方法进行偏见分析
作者:Hye-jin Shim,Md Sahidullah,Jee-weon Jung,Shinji Watanabe,Tomi Kinnunen
备注:5 pages, 1 figure, 5 tables
链接:点击下载PDF文件
摘要:音频反欺骗检测研究的当前趋势是通过学习识别各种欺骗伪像来提高模型在看不见的攻击中的泛化能力。这个重点主要集中在恶搞类。最近,一些研究指出,沉默的分布在这两个阶层之间是不同的,这可以作为一个捷径。在本文中,我们扩展类明智的解释超越沉默。我们采用损失分析和非对称方法,从传统的以攻击为中心和以结果为导向的评估转向对模型行为的更深入检查。我们的调查突出了两个类之间的训练动态的显着差异,强调未来的研究需要专注于鲁棒建模的真正的类。摘要:Current trends in audio anti-spoofing detection research strive to improve models' ability to generalize across unseen attacks by learning to identify a variety of spoofing artifacts. This emphasis has primarily focused on the spoof class. Recently, several studies have noted that the distribution of silence differs between the two classes, which can serve as a shortcut. In this paper, we extend class-wise interpretations beyond silence. We employ loss analysis and asymmetric methodologies to move away from traditional attack-focused and result-oriented evaluations towards a deeper examination of model behaviors. Our investigations highlight the significant differences in training dynamics between the two classes, emphasizing the need for future research to focus on robust modeling of the bonafide class.
【15】 Self-Supervised Embeddings for Detecting Individual Symptoms of Depression
标题: 检测抑郁症个体症状的自我监督嵌入
作者:Sri Harsha Dumpala,Katerina Dikaios,Abraham Nunes,Frank Rudzicz,Rudolf Uher,Sageev Oore
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:抑郁症是一种影响全球数百万人的普遍心理健康疾病,需要可靠的评估系统。与以前的研究不同,我们的研究只关注检测抑郁症或预测其严重程度,我们的工作识别了抑郁症的个体症状,同时还使用语音输入预测其严重程度。我们利用基于自我监督学习(SSL)的语音模型来更好地利用在此任务中经常遇到的小规模数据集。我们的研究表明,显着的性能改善,利用SSL嵌入相比,传统的语音功能。我们比较了各种类型的SSL预训练模型,以阐明在识别不同症状方面贡献最大的语音信息类型(语义,说话者或韵律)。此外,我们还评估了结合多个SSL嵌入对性能的影响。此外,我们显示了多任务学习的意义,有效地识别抑郁症状。摘要:Depression, a prevalent mental health disorder impacting millions globally, demands reliable assessment systems. Unlike previous studies that focus solely on either detecting depression or predicting its severity, our work identifies individual symptoms of depression while also predicting its severity using speech input. We leverage self-supervised learning (SSL)-based speech models to better utilize the small-sized datasets that are frequently encountered in this task. Our study demonstrates notable performance improvements by utilizing SSL embeddings compared to conventional speech features. We compare various types of SSL pretrained models to elucidate the type of speech information (semantic, speaker, or prosodic) that contributes the most in identifying different symptoms. Additionally, we evaluate the impact of combining multiple SSL embeddings on performance. Furthermore, we show the significance of multi-task learning for identifying depressive symptoms effectively.
【16】 Sound Tagging in Infant-centric Home Soundscapes
标题: 以婴儿为中心的家庭声音场景中的声音标记
作者:Mohammad Nur Hossain Khan,Jialu Li,Nancy L. McElwain,Mark Hasegawa-Johnson,Bashima Islam
备注:Accepted in IEEEACM CHASE 2024
链接:点击下载PDF文件
摘要:某些环境噪音与婴幼儿的负面发育结果有关。虽然对家庭环境中的声音事件进行分类或标记是一个活跃的研究领域,但以前的研究主要集中在从放置在环境中的非固定麦克风或从成人的角度收集的数据。此外,这些工作中的许多忽略了环境中的婴儿或幼儿,或者仅从单个家庭收集数据,其中来自固定声源的噪声在婴儿的位置处可能是中等的,反之亦然。因此,尽管最近大型预训练模型在噪声事件检测方面取得了成功,但这些模型在家庭中以婴儿为中心的噪声音景上的性能还有待探索。为了弥补这一差距,我们以一种不引人注目的方式收集并标记了22个家庭的家庭音景中的噪音,这些数据是通过婴儿佩戴的录音设备收集的。在本文中,我们探索了一个大型预训练模型(音频频谱图Transformer [AST])在我们的以婴儿为中心的环境数据以及公开的家庭环境数据集上的性能。利用不同的训练策略,如重新训练,利用公共数据集,混合公共和以婴儿为中心的训练集,以及使用噪声和掩蔽的数据增强,我们评估了一个大型预训练模型在稀疏和不平衡的以婴儿为中心的数据上的性能。我们的研究结果表明,通过将我们收集的数据集与公共数据集相结合来微调大型预训练模型,(公共数据集)和0.76(收集的数据集)至0.84(组合数据集)和Cohen的Kappa值从0.013(公共数据集)和0.77(收集的数据集)到0.83(组合数据集),分别与仅使用公共或收集的数据集进行训练相比。摘要:Certain environmental noises have been associated with negative developmental outcomes for infants and young children. Though classifying or tagging sound events in a domestic environment is an active research area, previous studies focused on data collected from a non-stationary microphone placed in the environment or from the perspective of adults. Further, many of these works ignore infants or young children in the environment or have data collected from only a single family where noise from the fixed sound source can be moderate at the infant's position or vice versa. Thus, despite the recent success of large pre-trained models for noise event detection, the performance of these models on infant-centric noise soundscapes in the home is yet to be explored. To bridge this gap, we have collected and labeled noises in home soundscapes from 22 families in an unobtrusive manner, where the data are collected through an infant-worn recording device. In this paper, we explore the performance of a large pre-trained model (Audio Spectrogram Transformer [AST]) on our noise-conditioned infant-centric environmental data as well as publicly available home environmental datasets. Utilizing different training strategies such as resampling, utilizing public datasets, mixing public and infant-centric training sets, and data augmentation using noise and masking, we evaluate the performance of a large pre-trained model on sparse and imbalanced infant-centric data. Our results show that fine-tuning the large pre-trained model by combining our collected dataset with public datasets increases the F1-score from 0.11 (public datasets) and 0.76 (collected datasets) to 0.84 (combined datasets) and Cohen's Kappa from 0.013 (public datasets) and 0.77 (collected datasets) to 0.83 (combined datasets) compared to only training with public or collected datasets, respectively.
【17】 Investigating Confidence Estimation Measures for Speaker Diarization
标题: 研究说话人数字化的置信度估计措施
作者:Anurag Chowdhury,Abhinav Misra,Mark C. Fuhs,Monika Woszczyna
备注:Accepted in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:说话人日记系统根据说话人的身份对会话记录进行分段。这样的系统可能由于诸如语音模式变化、背景噪声和重叠语音之类的各种因素而对音频的一部分的说话者进行错误分类。这些错误会传播到依赖说话者身份的下游系统(如说话者自适应语音识别),并对其产生不利影响。减轻这些错误的方法之一是向下游系统提供片段级日志化置信度分数。在这项工作中,我们研究了多种方法来生成日记置信分数,包括那些来自原始日记系统和那些来自外部模型。我们在多个数据集和日志化系统上的实验表明,最具竞争力的置信度方法可以隔离出约30%的日志化错误,其中置信度最低的约为10%。摘要:Speaker diarization systems segment a conversation recording based on the speakers' identity. Such systems can misclassify the speaker of a portion of audio due to a variety of factors, such as speech pattern variation, background noise, and overlapping speech. These errors propagate to, and can adversely affect, downstream systems that rely on the speaker's identity, such as speaker-adapted speech recognition. One of the ways to mitigate these errors is to provide segment-level diarization confidence scores to downstream systems. In this work, we investigate multiple methods for generating diarization confidence scores, including those derived from the original diarization system and those derived from an external model. Our experiments across multiple datasets and diarization systems demonstrate that the most competitive confidence score methods can isolate ~30% of the diarization errors within segments with the lowest ~10% of confidence scores.
【18】 Sound Field Synthesis with Acoustic Waves
标题: 利用声波合成的声学场
作者:Mohamed F. Mansour
备注:16 pages, double-spaced, 5 figures
链接:点击下载PDF文件
摘要:基于声传播的物理学原理,提出了一种在小刚性表面上合成宽带声场的实用框架。声场被生成为两个组件的合成图:房间组件和设备组件;其中声学平面波作为核心元素。房间和设备组件的这种解耦显著降低了问题的复杂性,并提供了声场的准确呈现。我们详细描述了所提出的框架的理论基础,有效的程序,和工程应用。所提出的框架的有效性是通过在不同的环境设置下进行严格的验证。摘要:We propose a practical framework to synthesize broadband sound-field on a small rigid surface based on the physics of sound propagation. The sound-field is generated as a composite map of two components: room component and device component; with acoustic plane waves as core element. This decoupling of room and device components significantly reduces the problem complexity and provide accurate rendering of the sound-field. We describe in details the theoretical foundations, efficient procedures, and engineering applications of the proposed framework. The effectiveness of the proposed framework is established through rigorous validation under different environment setups.
【19】 Maximum Likelihood Estimation of the Direction of Sound In A Reverberant Noisy Environment
标题: 回响噪音环境中声音方向的最大似然估计
作者:Mohamed F. Mansour
备注:5 pages, 2 figures, conference
链接:点击下载PDF文件
摘要:我们描述了一种新的方法,估计在混响环境中的声音传播的基本原理的方向。该方法利用信噪比自适应的功能,从时间延迟和能量的方向分量后,声波分解的观测声场估计的视线方向在噪声和混响条件下。该方法的有效性建立与不同的麦克风阵列配置在各种使用场景下的实际数据。摘要:We describe a new method for estimating the direction of sound in a reverberant environment from basic principles of sound propagation. The method utilizes SNR-adaptive features from time-delay and energy of the directional components after acoustic wave decomposition of the observed sound field to estimate the line-of-sight direction under noisy and reverberant conditions. The effectiveness of the approach is established with real-data of different microphone array configurations under various usage scenarios.
【20】 AND: Audio Network Dissection for Interpreting Deep Acoustic
标题: AND:用于解释深层声学的音频网络剖析
作者:Tung-Yu Wu,Yu-Xiang Lin,Tsui-Wei Weng
备注:Accepted by ICML'24
链接:点击下载PDF文件
摘要:神经元层面的解释旨在通过研究神经元对特定感知或结构输入模式的反应来解释网络的行为和属性。虽然在视觉和语言领域有一些新兴的工作,但没有一个是针对声学模型的。为了弥补这一差距,我们引入了$ texttit {AND}$,这是第一个$ textbf{A}$udio $ textbf{N}$etwork $ textbf{D}$issection框架,它可以基于高响应音频自动建立声学神经元的自然语言解释。$ textit{AND}$的特点是使用LLM来总结音频之间的相互声学特征和身份。进行了大量的实验,以验证$ textit{AND}$的精确和翔实的描述。此外,我们展示了一个潜在的使用$ textit{AND}$音频机器unlearning进行概念特定的修剪生成的描述的基础上。最后,通过$ textit{AND}$的分析,我们强调了两个声学模型的行为:(i)模型通过基本声学特征的组合而不是高级抽象概念来区分音频;(ii)训练策略影响模型行为和神经元可解释性--监督训练引导神经元逐渐缩小注意力,而自我监督学习鼓励神经元成为多语义的,以探索高级特征。摘要:Neuron-level interpretations aim to explain network behaviors and properties by investigating neurons responsive to specific perceptual or structural input patterns. Although there is emerging work in the vision and language domains, none is explored for acoustic models. To bridge the gap, we introduce $ textit{AND}$, the first $ textbf{A}$udio $ textbf{N}$etwork $ textbf{D}$issection framework that automatically establishes natural language explanations of acoustic neurons based on highly-responsive audio. $ textit{AND}$ features the use of LLMs to summarize mutual acoustic features and identities among audio. Extensive experiments are conducted to verify $ textit{AND}$'s precise and informative descriptions. In addition, we demonstrate a potential use of $ textit{AND}$ for audio machine unlearning by conducting concept-specific pruning based on the generated descriptions. Finally, we highlight two acoustic model behaviors with analysis by $ textit{AND}$: (i) models discriminate audio with a combination of basic acoustic features rather than high-level abstract concepts; (ii) training strategies affect model behaviors and neuron interpretability -- supervised training guides neurons to gradually narrow their attention, while self-supervised learning encourages neurons to be polysemantic for exploring high-level features.
机器翻译,仅供参考
