今天跟大家分享一篇语音相关的论文合集:cs.SD语音12篇,eess.AS音频处理15篇。

cs.SD语音

【1】 How to hide your voice: Noise-cancelling bird photography blind

标题:如何隐藏你的声音:消除噪音的鸟类摄影盲人

链接:https://arxiv.org/abs/2206.12340

作者:C. Baydur,B. Pu,X. Xu

备注:22 pages, 11 figures

摘要:在野生动物摄影中,接近鸟类是一个巨大的挑战。鸟类摄影盲板可能是最有效、干扰最小的方法。如果设计得当,这些基本结构可以让摄影师在视觉上和听觉上远离栖息地。然而,百叶窗的声学设计被忽视了。在此,我们提出了噪声消除盲板,允许近距离拍摄鸟类。首先,我们在位于中国云南的生态旅游中心进行了问卷调查。因此,我们确定了观鸟者对室内声环境的期望。然后,我们确定了四个变量,以检查建筑和声学决策对噪声传播的影响。在Comsol MultiPhysics的声学模块中进行了数值模拟。在建筑设计过程中,减小结构尺寸和规划封闭窗户的建筑是减少噪音的正确决策。吸声材料可以降低室内的声能,从而降低室外噪声。隔音材料有助于消除室内到室外的声音传输。同时使用吸音和隔音材料是将室内外噪音降至最低的最佳方法。我们的研究表明,为了人类和鸟类的健康,摄影百叶窗需要强大而彻底的声学设计。

摘要:Getting close to birds is a great challenge in wildlife photography. Bird photography blinds may be the most effective and least intrusive way. These essential structures can allow to visually and audibly conceal photographers from the habitat if properly designed. However, the acoustic design of the blinds has been overlooked. Herein, we present noise-cancelling blinds which allow photographing birds at close range. Firstly, we conduct a questionnaire in the eco-tourism centre located in Yunnan, China. Thus, we determine the birders' expectations of the indoor sound environment. We then identify four variables to examine the impact of architectural and acoustic decisions on noise propagation. The numerical simulations are performed in the acoustic module of Comsol MultiPhysics. Minimizing the structural size and planning the building with closed windows is a proper decision to reduce noise in the architectural design process. Sound-absorbing materials reduce the acoustic energy indoors, thus decreasing the outdoor noise. Sound-proofing materials help to cancel the acoustic transmission indoors to outdoors. Using sound-absorbing and proofing materials together is the best way to minimize noise both indoors and outdoors. Our study demonstrated that photography blinds require a strong and thorough acoustic design for both human and bird well-being.


【2】 PoCaP Corpus: A Multimodal Dataset for Smart Operating Room Speech  Assistant using Interventional Radiology Workflow Analysis

标题:PoCaP语料库:用于介入放射工作流分析的智能手术室语音助手的多模式数据集

链接:https://arxiv.org/abs/2206.12320

作者:Kubilay Can Demir,Matthias May,Axel Schmid,Michael Uder,Katharina Breininger,Tobias Weise,Andreas Maier,Seung Hee Yang
备注:8 pages, 4 figures, Text, Speech and Dialogue 2022 Conference
摘要:本文提出了一种新的多模式介入放射学数据集,称为PoCaP(端口导管放置)语料库。该语料库包括德语语音和音频信号、X射线图像和系统命令,这些信号由6名外科医生从31次PoCaP干预中收集,平均持续时间为81.4$\pm$41.0分钟。该语料库旨在为在手术室开发智能语音助手提供资源。尤其是,它可用于开发一种语音控制系统,使外科医生能够控制手术参数,如C形臂移动和手术台位置。为了记录数据集,我们获得了爱尔兰大学医院机构审查委员会和工人委员会以及患者对数据隐私的同意。我们描述了记录设置、数据结构、工作流程和预处理步骤,并使用预训练模型报告了第一个PoCaP语料库语音识别分析结果,错误率为11.52$\%$。研究结果表明,这些数据有可能构建一个强大的命令识别系统,并将允许开发一种新型的干预支持系统,该系统在医学领域使用语音和图像处理。
摘要:This paper presents a new multimodal interventional radiology dataset, called PoCaP (Port Catheter Placement) Corpus. This corpus consists of speech and audio signals in German, X-ray images, and system commands collected from 31 PoCaP interventions by six surgeons with average duration of 81.4 $\pm$ 41.0 minutes. The corpus aims to provide a resource for developing a smart speech assistant in operating rooms. In particular, it may be used to develop a speech controlled system that enables surgeons to control the operation parameters such as C-arm movements and table positions. In order to record the dataset, we acquired consent by the institutional review board and workers council in the University Hospital Erlangen and by the patients for data privacy. We describe the recording set-up, data structure, workflow and preprocessing steps, and report the first PoCaP Corpus speech recognition analysis results with 11.52 $\%$ word error rate using pretrained models. The findings suggest that the data has the potential to build a robust command recognition system and will allow the development of a novel intervention support systems using speech and image processing in the medical domain.


【3】 Deformable CNN and Imbalance-Aware Feature Learning for Singing  Technique Classification

标题:用于歌唱技巧分类的可变形CNN和不平衡感知特征学习

链接:https://arxiv.org/abs/2206.12230

作者:Yuya Yamamoto,Juhan Nam,Hiroko Terasawa
备注:Accepted to INTERSPEECH2022
摘要:歌唱技巧是利用音色、音高和其他声音成分的时间波动来进行富有表现力的声乐表演。他们的分类是一项具有挑战性的任务,因为主要有两个因素:1)歌唱技巧的波动具有广泛的多样性,并受到许多因素的影响;2)现有数据集不平衡。为了解决这些问题,我们开发了一种新的基于变形卷积的音频特征学习方法,该方法使用类加权损失函数对特征抽取器和分类器进行解耦训练。实验结果表明:1)可变形卷积改进了分类结果,尤其是将其应用于最后两个卷积层时,以及2)重新训练分类器和通过平滑的逆频率加权交叉熵损失函数都提高了分类性能。
摘要:Singing techniques are used for expressive vocal performances by employing temporal fluctuations of the timbre, the pitch, and other components of the voice. Their classification is a challenging task, because of mainly two factors: 1) the fluctuations in singing techniques have a wide variety and are affected by many factors and 2) existing datasets are imbalanced. To deal with these problems, we developed a novel audio feature learning method based on deformable convolution with decoupled training of the feature extractor and the classifier using a class-weighted loss function. The experimental results show the following: 1) the deformable convolution improves the classification results, particularly when it is applied to the last two convolutional layers, and 2) both re-training the classifier and weighting the cross-entropy loss function by a smoothed inverse frequency enhance the classification performance.


【4】 Prosody Cloning in Zero-Shot Multispeaker Text-to-Speech

标题:零发声多人文语转换中的韵律克隆

链接:https://arxiv.org/abs/2206.12229

作者:Florian Lux,Julia Koch,Ngoc Thang Vu
摘要:使用未翻译的参考样本克隆说话人的声音是现代神经文语转换(TTS)方法的重大进步之一。最近也提出了模仿转录参考音频韵律的方法。在这项工作中,我们首次通过话语级规范化和话语级说话人嵌入将这两项任务结合在一起。我们进一步介绍了一种用于提取细粒度韵律特征的轻量级对齐器,可以在几秒钟内对单个样本进行微调。我们的客观评估和人类研究表明,可以独立克隆说话人的声音以及口语参考的韵律,而不会降低质量,并且与原始声音和韵律高度相似。我们的所有代码和经过训练的模型都可用,还有静态和交互式演示。
摘要:The cloning of a speaker's voice using an untranscribed reference sample is one of the great advances of modern neural text-to-speech (TTS) methods. Approaches for mimicking the prosody of a transcribed reference audio have also been proposed recently. In this work, we bring these two tasks together for the first time through utterance level normalization in conjunction with an utterance level speaker embedding. We further introduce a lightweight aligner for extracting fine-grained prosodic features, that can be finetuned on individual samples within seconds. We show that it is possible to clone the voice of a speaker as well as the prosody of a spoken reference independently without any degradation in quality and high similarity to both original voice and prosody, as our objective evaluation and human study show. All of our code and trained models are available, alongside static and interactive demos.


【5】 BYOL-S: Learning Self-supervised Speech Representations by Bootstrapping

标题:BYOL-S:通过自举学习自监督语音表示

链接:https://arxiv.org/abs/2206.12038

作者:Gasser Elbanna,Neil Scheidwasser-Clow,Mikolaj Kegler,Pierre Beckmann,Karl El Hajal,Milos Cernak

摘要:自几十年前频谱分析的开创性工作以来,音频和语音特征的提取方法一直在研究中。最近的努力以开发通用音频表示的雄心壮志为指导。例如,如果对大型音频数据集进行训练,深度神经网络可以提取最佳嵌入。这项工作扩展了现有的基于自引导自监督学习的方法,提出了各种编码器体系结构,并探讨了使用不同预训练数据集的效果。最后,我们提出了一个新的训练框架,提出了一种混合音频表示,它结合了手工制作和数据驱动的学习音频特征。在听觉场景分类和时间戳检测任务的听觉神经末梢2021挑战中,对所有提议的表示进行了评估。我们的结果表明,以卷积Transformer作为编码器的混合模型在大多数听力挑战任务中具有优异的性能。

摘要:Methods for extracting audio and speech features have been studied since pioneering work on spectrum analysis decades ago. Recent efforts are guided by the ambition to develop general-purpose audio representations. For example, deep neural networks can extract optimal embeddings if they are trained on large audio datasets. This work extends existing methods based on self-supervised learning by bootstrapping, proposes various encoder architectures, and explores the effects of using different pre-training datasets. Lastly, we present a novel training framework to come up with a hybrid audio representation, which combines handcrafted and data-driven learned audio features. All the proposed representations were evaluated within the HEAR NeurIPS 2021 challenge for auditory scene classification and timestamp detection tasks. Our results indicate that the hybrid model with a convolutional transformer as the encoder yields superior performance in most HEAR challenge tasks.


【6】 Comparing supervised and self-supervised embedding for ExVo Multi-Task  learning track

标题:ExVo多任务学习路径的监督嵌入和自监督嵌入的比较

链接:https://arxiv.org/abs/2206.11968

作者:Tilak Purohit,Imen Ben Mahmoud,Bogdan Vlasenko,Mathew Magimai. -Doss
摘要:ICML表达性发声(ExVo)2022年多任务挑战赛侧重于理解非语言发声(发声爆发(VB))的情感方面。这项挑战的目标是预测VB的情绪强度,这是一项多任务挑战,还需要预测说话者的年龄和母语。针对这一挑战,我们研究并比较了两种不同的嵌入空间,即基于自监督学习(SSL)的嵌入和基于任务特定监督学习的嵌入。为此,我们研究了从几个预先训练的SSL神经网络和任务特定监督分类神经网络获得的特征表示。我们的研究表明,混合方法可以获得最佳性能,其中使用了通过SSL和任务特定监督学习得出的预测。我们在测试集上的最佳系统超过了比较基线(所有子任务得分的调和平均值,即$S\u{MTL}$),相对优势为13\%$。
摘要:The ICML Expressive Vocalizations (ExVo) Multi-task challenge 2022, focuses on understanding the emotional facets of the non-linguistic vocalizations (vocal bursts (VB)). The objective of this challenge is to predict emotional intensities for VB, being a multi-task challenge it also requires to predict speakers' age and native-country. For this challenge we study and compare two distinct embedding spaces namely, self-supervised learning (SSL) based embeddings and task-specific supervised learning based embeddings. Towards that, we investigate feature representations obtained from several pre-trained SSL neural networks and task-specific supervised classification neural networks. Our studies show that the best performance is obtained with a hybrid approach, where predictions derived via both SSL and task-specific supervised learning are used. Our best system on test-set surpasses the ComPARE baseline (harmonic mean of all sub-task scores i.e., $S_{MTL}$) by a relative $13\%$ margin.


【7】 SAQAM: Spatial Audio Quality Assessment Metric

标题:SAQAM:空间音频质量评估指标

链接:https://arxiv.org/abs/2206.12297

作者:Pranay Manocha,Anurag Kumar,Buye Xu,Anjali Menon,Israel D. Gebru,Vamsi K. Ithapu,Paul Calamia
备注:To Appear, Interspeech 2022
摘要:音频质量评估对于评估声音的感知真实性至关重要。然而,获得“金标准”人类判断的时间和费用限制了此类数据的可用性。对于AR和VR,良好的感知音质和声源的本地化是确保用户完全沉浸其中的关键因素。我们的工作介绍了SAQAM,它使用多任务学习框架来评估任何给定双耳信号对之间的听力质量(LQ)和空间化质量(SQ),而不使用任何主观数据。我们通过在三重人体判断的模拟数据集上进行训练来建立LQ模型,并通过利用来自经过波达方向(DOA)估计训练的网络的激活水平距离来建立SQ模型。我们表明,SAQAM与四个不同数据集的人类反应有很好的相关性。由于它是一个深度网络,度量是可微的,因此它适合作为其他任务的损失函数。例如,简单地用我们的度量替换现有的损耗,可以提高语音增强网络的性能。
摘要:Audio quality assessment is critical for assessing the perceptual realism of sounds. However, the time and expense of obtaining ''gold standard'' human judgments limit the availability of such data. For AR&VR, good perceived sound quality and localizability of sources are among the key elements to ensure complete immersion of the user. Our work introduces SAQAM which uses a multi-task learning framework to assess listening quality (LQ) and spatialization quality (SQ) between any given pair of binaural signals without using any subjective data. We model LQ by training on a simulated dataset of triplet human judgments, and SQ by utilizing activation-level distances from networks trained for direction of arrival (DOA) estimation. We show that SAQAM correlates well with human responses across four diverse datasets. Since it is a deep network, the metric is differentiable, making it suitable as a loss function for other tasks. For example, simply replacing an existing loss with our metric yields improvement in a speech-enhancement network.


【8】 Speech Quality Assessment through MOS using Non-Matching References

标题:基于非匹配参考的MOS语音质量评估

链接:https://arxiv.org/abs/2206.12285

作者:Pranay Manocha,Anurag Kumar

备注:To Appear, Interspeech 2022

摘要:通过平均意见得分(MOS)获得的人类判断是评估语音信号质量最可靠的方法。然而,最近一些使用深度学习方法自动估计MOS的尝试缺乏稳健性和泛化能力,限制了它们在实际应用中的使用。在这项工作中,我们提出了一种新的框架,NORESQA-MOS,用于估计语音信号的MOS。与之前的工作不同,我们的方法使用非匹配参考作为条件,通过神经网络为MOS估计奠定基础。我们表明,NORESQA-MOS比以前的最新方法(如DNSMOS和NISQA)具有更好的泛化能力和更稳健的MOS估计,即使我们使用的训练集较小。此外,我们还表明,我们的通用框架可以与其他学习方法(如自监督学习)相结合,并可以进一步补充这些方法的优点。

摘要:Human judgments obtained through Mean Opinion Scores (MOS) are the most reliable way to assess the quality of speech signals. However, several recent attempts to automatically estimate MOS using deep learning approaches lack robustness and generalization capabilities, limiting their use in real-world applications. In this work, we present a novel framework, NORESQA-MOS, for estimating the MOS of a speech signal. Unlike prior works, our approach uses non-matching references as a form of conditioning to ground the MOS estimation by neural networks. We show that NORESQA-MOS provides better generalization and more robust MOS estimation than previous state-of-the-art methods such as DNSMOS and NISQA, even though we use a smaller training set. Moreover, we also show that our generic framework can be combined with other learning methods such as self-supervised learning and can further supplement the benefits from these methods.


【9】 Open-source objective-oriented framework for head-related transfer  function

标题:开源面向目标的头部相关传递函数框架

链接:https://arxiv.org/abs/2206.12283

作者:Adam Szwajcowski
备注:Not submitted anywhere in the current form
摘要:在过去的30年中,已经开发了许多与头部相关的传递函数(HRTF)模型,还有更多的模型。本文描述了一个基于面向目标编程范式的框架,其中每个HRTF表示方法都可以作为一个单独的类来实现。它的模块化结构允许源代码在研究人员之间方便地共享,而公共接口提供了对数据的轻松访问,而不管类的内部结构如何。本文讨论了设计该框架、保持其灵活性和找到每个可能的方向性表示的共同特征之间的平衡的困难。包括并解释了示例性用例。采用该框架将提高各种HRTF模型之间准确性比较的可能性,从而改进对当前和未来表示方法的评估。该框架以MATLAB工具箱的形式开发,不仅可以处理HRTF,还可以处理其他类型的空间数据,例如声源方向性、麦克风方向性等。
摘要:Throughout last 30 years, numerous head-related transfer function (HRTF) models have been developed and there are more to come. This paper describes a framework based on objective-oriented programming paradigm, in which each HRTF representation method can be implemented as a separate class. Its modular structure allows the source code to be conveniently shared between researchers, while common interface provides easy access to data regardless of the internal structure of the classes. The paper discusses difficulties of designing the framework, maintaining the balance between its flexibility and finding common features of every possible directivity representation. Exemplary use cases are included and explained. Adoption of the framework will enhance possibilities of accuracy comparison between various HRTF models, thus improving the evaluation of current and future representation methods. The framework, developed in the form of a MATLAB toolbox, is designed to handle not only HRTFs but also other types of spatial data, such as e.g. sound source directivity, microphone directivity, etc.


【10】 Data Augmentation and Squeeze-and-Excitation Network on Multiple  Dimension for Sound Event Localization and Detection in Real Scenes

标题:用于真实场景声事件定位和检测的多维数据增强和挤压激励网络

链接:https://arxiv.org/abs/2206.12059

作者:Byeong-Yun Ko,Hyeonuk Nam,Seong-Hu Kim,Deokki Min,Seung-Deok Choi,Yong-Hwa Park
备注:Technical Report submitted for DCASE2022 Challenge Task3
摘要:真实场景中声音事件定位与检测(SELD)的性能受到SELD数据集较小的限制,因为难以获得足够数量的真实多通道音频数据记录和准确的标签。我们使用两种主要策略来解决小型真实SELD数据集产生的问题。首先,我们在所有数据维度上应用了各种数据增强方法:信道、频率和时间。我们还提出了一种称为适度混合的原始数据增强方法,以模拟存在噪声地板或干扰事件的情况。其次,我们在通道和频率维度上应用挤压和激励块来有效地提取特征。我们在STARSS22测试数据集上训练的模型的结果达到了最佳ER、F1、LE和LR,分别为0.53、49.8%、16.0度。,和56.2%。
摘要:Performance of sound event localization and detection (SELD) in real scenes is limited by small size of SELD dataset, due to difficulty in obtaining sufficient amount of realistic multi-channel audio data recordings with accurate label. We used two main strategies to solve problems arising from the small real SELD dataset. First, we applied various data augmentation methods on all data dimensions: channel, frequency and time. We also propose original data augmentation method named Moderate Mixup in order to simulate situations where noise floor or interfering events exist. Second, we applied Squeeze-and-Excitation block on channel and frequency dimensions to efficiently extract feature characteristics. Result of our trained models on the STARSS22 test dataset achieved the best ER, F1, LE, and LR of 0.53, 49.8%, 16.0deg., and 56.2% respectively.


【11】 Confidence Score Based Conformer Speaker Adaptation for Speech  Recognition

标题:基于置信度的语音识别整形说话人自适应

链接:https://arxiv.org/abs/2206.12045

作者:Jiajun Deng,Xurong Xie,Tianzi Wang,Mingyu Cui,Boyang Xue,Zengrui Jin,Mengzhe Geng,Guinan Li,Xunying Liu,Helen Meng

备注:It's accepted to INTERSPEECH 2022. arXiv admin note: text overlap with arXiv:2206.11596

摘要:自动语音识别(ASR)系统面临的一个关键挑战是对说话人级别的可变性进行建模。本文采用紧凑的说话人相关学习隐藏单元贡献(LHUC)来促进基于一致性的端到端ASR系统的说话人自适应训练(SAT)和测试时无监督说话人自适应。使用基于可信度得分的选择说话人特定数据中更“可信”的子集,可以降低自适应过程中对监督错误率的敏感性。置信度估计模块用于在用作置信度得分之前平滑过度置信度一致性解码器的输出概率。利用LHUC参数的贝叶斯估计,解决了说话人级数据选择导致的数据稀疏性增加的问题。在300小时的交换机语料库上进行的实验表明,在NIST Hub5'00、RT02、,和RT03评估集。在使用外部Transformer和LSTM语言模型进行重新排序后,保持了一致的性能改进。

摘要:A key challenge for automatic speech recognition (ASR) systems is to model the speaker level variability. In this paper, compact speaker dependent learning hidden unit contributions (LHUC) are used to facilitate both speaker adaptive training (SAT) and test time unsupervised speaker adaptation for state-of-the-art Conformer based end-to-end ASR systems. The sensitivity during adaptation to supervision error rate is reduced using confidence score based selection of the more "trustworthy" subset of speaker specific data. A confidence estimation module is used to smooth the over-confident Conformer decoder output probabilities before serving as confidence scores. The increased data sparsity due to speaker level data selection is addressed using Bayesian estimation of LHUC parameters. Experiments on the 300-hour Switchboard corpus suggest that the proposed LHUC-SAT Conformer with confidence score based test time unsupervised adaptation outperformed the baseline speaker independent and i-vector adapted Conformer systems by up to 1.0%, 1.0%, and 1.2% absolute (9.0%, 7.9%, and 8.9% relative) word error rate (WER) reductions on the NIST Hub5'00, RT02, and RT03 evaluation sets respectively. Consistent performance improvements were retained after external Transformer and LSTM language models were used for rescoring.


【12】 End-to-End Text-to-Speech Based on Latent Representation of Speaking  Styles Using Spontaneous Dialogue

标题:基于自发对话的潜在语体表征的端到端文语转换

链接:https://arxiv.org/abs/2206.12040

作者:Kentaro Mitsui,Tianyu Zhao,Kei Sawada,Yukiya Hono,Yoshihiko Nankaku,Keiichi Tokuda

备注:5 pages, 3 figures, accepted for INTERSPEECH 2022. Audio samples: this https URL

摘要:最近的文语转换(TTS)已经达到了与人类相当的质量;然而,它在口语对话中的应用还没有得到广泛的研究。本研究旨在实现一种与人类对话非常相似的TTS。首先,我们记录并转录真实的自发对话。然后,将所提出的对话TTS分为两个阶段进行训练:第一阶段,训练变分自动编码器(VAE)-VITS或高斯混合变分自动编码器(GMVAE)-VITS,该阶段通过端到端文本到语音(VITS)的对抗式学习将话语级潜变量引入到变分推理中,VITS是最近提出的端到端TTS模型。从语音中提取潜在说话风格表示的风格编码器与TTS联合训练。在第二阶段,训练一个风格预测因子来预测从对话历史中综合出来的说话风格。在推理过程中,通过将风格预测器预测的说话风格表示传递给VAE/GMVAE-VITS,可以以适合对话上下文的风格合成语音。主观评价结果表明,该方法在对话级自然度方面优于原VITS。

摘要:The recent text-to-speech (TTS) has achieved quality comparable to that of humans; however, its application in spoken dialogue has not been widely studied. This study aims to realize a TTS that closely resembles human dialogue. First, we record and transcribe actual spontaneous dialogues. Then, the proposed dialogue TTS is trained in two stages: first stage, variational autoencoder (VAE)-VITS or Gaussian mixture variational autoencoder (GMVAE)-VITS is trained, which introduces an utterance-level latent variable into variational inference with adversarial learning for end-to-end text-to-speech (VITS), a recently proposed end-to-end TTS model. A style encoder that extracts a latent speaking style representation from speech is trained jointly with TTS. In the second stage, a style predictor is trained to predict the speaking style to be synthesized from dialogue history. During inference, by passing the speaking style representation predicted by the style predictor to VAE/GMVAE-VITS, speech can be synthesized in a style appropriate to the context of the dialogue. Subjective evaluation results demonstrate that the proposed method outperforms the original VITS in terms of dialogue-level naturalness.


eess.AS音频处理

【1】 Analyzing the impact of SARS-CoV-2 variants on respiratory sound signals

标题:SARS-CoV-2变异株对呼吸音信号的影响分析

链接:https://arxiv.org/abs/2206.12309

作者:Debarpan Bhattacharya,Debottam Dutta,Neeraj Kumar Sharma,Srikanth Raj Chetupalli,Pravin Mote,Sriram Ganapathy,Chandrakiran C,Sahiti Nori,Suhail K K,Sadhana Gonuguntla,Murali Alagesan

摘要:2019冠状病毒疾病疫情导致了多波与不同SARS-CoV-2变体相关的感染。研究报告了变异对患者呼吸健康的不同影响。我们探讨了从2019冠状病毒疾病受试者身上采集的声学信号是否显示出计算上可区分的声学模式,这表明有可能预测潜在的病毒变体。我们分析了从三个受试者库中收集的Coswara数据集,即:i)健康,ii)delta变异显性期记录的2019冠状病毒疾病受试者,以及iii)来自omicron激增期间记录的2019冠状病毒疾病受试者的数据。我们的研究结果表明,当将2019冠状病毒疾病受试者与omicron和delta变异体进行比较时,咳嗽、呼吸和言语等多种声音类别显示出显著的声学特征差异。曲线下的分类区域明显高于区分奥米克罗感染者和delta感染者的可能性。使用多个声音类别的评分融合,我们在95%的特异性下获得了89%和52.4%的曲线下面积。此外,采用分层三级方法将声学数据分为健康和2019冠状病毒疾病阳性,并进一步将2019冠状病毒疾病受试者分为delta和omicron变体,提供高水平的三级分类准确性。这些结果为设计基于声音的2019冠状病毒疾病诊断方法提供了新方法。

摘要:The COVID-19 outbreak resulted in multiple waves of infections that have been associated with different SARS-CoV-2 variants. Studies have reported differential impact of the variants on respiratory health of patients. We explore whether acoustic signals, collected from COVID-19 subjects, show computationally distinguishable acoustic patterns suggesting a possibility to predict the underlying virus variant. We analyze the Coswara dataset which is collected from three subject pools, namely, i) healthy, ii) COVID-19 subjects recorded during the delta variant dominant period, and iii) data from COVID-19 subjects recorded during the omicron surge. Our findings suggest that multiple sound categories, such as cough, breathing, and speech, indicate significant acoustic feature differences when comparing COVID-19 subjects with omicron and delta variants. The classification areas-under-the-curve are significantly above chance for differentiating subjects infected by omicron from those infected by delta. Using a score fusion from multiple sound categories, we obtained an area-under-the-curve of 89% and 52.4% sensitivity at 95% specificity. Additionally, a hierarchical three class approach was used to classify the acoustic data into healthy and COVID-19 positive, and further COVID-19 subjects into delta and omicron variants providing high level of 3-class classification accuracy. These results suggest new ways for designing sound based COVID-19 diagnosis approaches.


【2】 SAQAM: Spatial Audio Quality Assessment Metric

标题:SAQAM:空间音频质量评估指标

链接:https://arxiv.org/abs/2206.12297

* 与cs.SD语音【7】为同一篇

作者:Pranay Manocha,Anurag Kumar,Buye Xu,Anjali Menon,Israel D. Gebru,Vamsi K. Ithapu,Paul Calamia

备注:To Appear, Interspeech 2022

摘要:音频质量评估对于评估声音的感知真实性至关重要。然而,获得“金标准”人类判断的时间和费用限制了此类数据的可用性。对于AR和VR,良好的感知音质和声源的本地化是确保用户完全沉浸其中的关键因素。我们的工作介绍了SAQAM,它使用多任务学习框架来评估任何给定双耳信号对之间的听力质量(LQ)和空间化质量(SQ),而不使用任何主观数据。我们通过在三重人体判断的模拟数据集上进行训练来建立LQ模型,并通过利用来自经过波达方向(DOA)估计训练的网络的激活水平距离来建立SQ模型。我们表明,SAQAM与四个不同数据集的人类反应有很好的相关性。由于它是一个深度网络,度量是可微的,因此它适合作为其他任务的损失函数。例如,简单地用我们的度量替换现有的损耗,可以提高语音增强网络的性能。

摘要:Audio quality assessment is critical for assessing the perceptual realism of sounds. However, the time and expense of obtaining ''gold standard'' human judgments limit the availability of such data. For AR&VR, good perceived sound quality and localizability of sources are among the key elements to ensure complete immersion of the user. Our work introduces SAQAM which uses a multi-task learning framework to assess listening quality (LQ) and spatialization quality (SQ) between any given pair of binaural signals without using any subjective data. We model LQ by training on a simulated dataset of triplet human judgments, and SQ by utilizing activation-level distances from networks trained for direction of arrival (DOA) estimation. We show that SAQAM correlates well with human responses across four diverse datasets. Since it is a deep network, the metric is differentiable, making it suitable as a loss function for other tasks. For example, simply replacing an existing loss with our metric yields improvement in a speech-enhancement network.


【3】 Speech Quality Assessment through MOS using Non-Matching References

标题:基于非匹配参考的MOS语音质量评估

链接:https://arxiv.org/abs/2206.12285

* 与cs.SD语音【8】为同一篇

作者:Pranay Manocha,Anurag Kumar
备注:To Appear, Interspeech 2022
摘要:通过平均意见得分(MOS)获得的人类判断是评估语音信号质量最可靠的方法。然而,最近一些使用深度学习方法自动估计MOS的尝试缺乏稳健性和泛化能力,限制了它们在实际应用中的使用。在这项工作中,我们提出了一种新的框架,NORESQA-MOS,用于估计语音信号的MOS。与之前的工作不同,我们的方法使用非匹配参考作为条件,通过神经网络为MOS估计奠定基础。我们表明,NORESQA-MOS比以前的最新方法(如DNSMOS和NISQA)具有更好的泛化能力和更稳健的MOS估计,即使我们使用的训练集较小。此外,我们还表明,我们的通用框架可以与其他学习方法(如自监督学习)相结合,并可以进一步补充这些方法的优点。
摘要:Human judgments obtained through Mean Opinion Scores (MOS) are the most reliable way to assess the quality of speech signals. However, several recent attempts to automatically estimate MOS using deep learning approaches lack robustness and generalization capabilities, limiting their use in real-world applications. In this work, we present a novel framework, NORESQA-MOS, for estimating the MOS of a speech signal. Unlike prior works, our approach uses non-matching references as a form of conditioning to ground the MOS estimation by neural networks. We show that NORESQA-MOS provides better generalization and more robust MOS estimation than previous state-of-the-art methods such as DNSMOS and NISQA, even though we use a smaller training set. Moreover, we also show that our generic framework can be combined with other learning methods such as self-supervised learning and can further supplement the benefits from these methods.


【4】 Open-source objective-oriented framework for head-related transfer  function

标题:开源面向目标的头部相关传递函数框架

链接:https://arxiv.org/abs/2206.12283

* 与cs.SD语音【9】为同一篇

作者:Adam Szwajcowski

备注:Not submitted anywhere in the current form

摘要:在过去的30年中,已经开发了许多与头部相关的传递函数(HRTF)模型,还有更多的模型。本文描述了一个基于面向目标编程范式的框架,其中每个HRTF表示方法都可以作为一个单独的类来实现。它的模块化结构允许源代码在研究人员之间方便地共享,而公共接口提供了对数据的轻松访问,而不管类的内部结构如何。本文讨论了设计该框架、保持其灵活性和找到每个可能的方向性表示的共同特征之间的平衡的困难。包括并解释了示例性用例。采用该框架将提高各种HRTF模型之间准确性比较的可能性,从而改进对当前和未来表示方法的评估。该框架以MATLAB工具箱的形式开发,不仅可以处理HRTF,还可以处理其他类型的空间数据,例如声源方向性、麦克风方向性等。

摘要:Throughout last 30 years, numerous head-related transfer function (HRTF) models have been developed and there are more to come. This paper describes a framework based on objective-oriented programming paradigm, in which each HRTF representation method can be implemented as a separate class. Its modular structure allows the source code to be conveniently shared between researchers, while common interface provides easy access to data regardless of the internal structure of the classes. The paper discusses difficulties of designing the framework, maintaining the balance between its flexibility and finding common features of every possible directivity representation. Exemplary use cases are included and explained. Adoption of the framework will enhance possibilities of accuracy comparison between various HRTF models, thus improving the evaluation of current and future representation methods. The framework, developed in the form of a MATLAB toolbox, is designed to handle not only HRTFs but also other types of spatial data, such as e.g. sound source directivity, microphone directivity, etc.


【5】 Iterative Sound Source Localization for Unknown Number of Sources

标题:未知声源个数的迭代声源定位

链接:https://arxiv.org/abs/2206.12273

作者:Yanjie Fu,Meng Ge,Haoran Yin,Xinyuan Qian,Longbiao Wang,Gaoyan Zhang,Jianwu Dang
备注:Accepted by Interspeech 2022
摘要:声源定位的目的是从观测到的多通道音频中寻找所有声源的到达方向(DOA)。对于信源数目未知的实际问题,现有的定位算法试图预测基于似然的编码(即空间谱),并使用预先确定的阈值来检测信源数目和相应的DOA值。然而,这些基于阈值的算法由于受到阈值选择的限制而不稳定。为了解决这个问题,我们提出了一种迭代声源定位方法ISSL,该方法可以迭代地提取每个声源的DOA,无需阈值,直到满足终止条件。与基于阈值的算法不同,ISSL设计了一个基于二进制分类器的有源源检测网络,以接受剩余空间谱并决定是否停止迭代。通过这样做,我们的ISSL可以处理任意数量的源,甚至比训练阶段看到的源数量还要多。实验结果表明,与现有的基于阈值的算法相比,我们的ISSL在DOA估计和信源数目检测方面都取得了显著的性能改进。
摘要:Sound source localization aims to seek the direction of arrival (DOA) of all sound sources from the observed multi-channel audio. For the practical problem of unknown number of sources, existing localization algorithms attempt to predict a likelihood-based coding (i.e., spatial spectrum) and employ a pre-determined threshold to detect the source number and corresponding DOA value. However, these threshold-based algorithms are not stable since they are limited by the careful choice of threshold. To address this problem, we propose an iterative sound source localization approach called ISSL, which can iteratively extract each source's DOA without threshold until the termination criterion is met. Unlike threshold-based algorithms, ISSL designs an active source detector network based on binary classifier to accept residual spatial spectrum and decide whether to stop the iteration. By doing so, our ISSL can deal with an arbitrary number of sources, even more than the number of sources seen during the training stage. The experimental results show that our ISSL achieves significant performance improvements in both DOA estimation and source number detection compared with the existing threshold-based algorithms.


【6】 SANE-TTS: Stable And Natural End-to-End Multilingual Text-to-Speech

标题:Sane-TTS:稳定自然的端到端多语言文语转换

链接:https://arxiv.org/abs/2206.12132

作者:Hyunjae Cho,Wonbin Jung,Junhyeok Lee,Sang Hoon Woo
备注:Accepted to Interspeech 2022
摘要:在本文中,我们提出了SANE-TTS,一种稳定自然的端到端多语言TTS模型。由于特定说话人很难获得多语语料库,用单语语料库训练多语TTS模型是不可避免的。我们引入了说话人正则化丢失,在跨语言合成过程中提高了语音的自然度,并引入了域对抗训练,这在其他多语言TTS模型中也得到了应用。此外,通过增加说话人正则化损失,在持续时间预测器中用零向量代替说话人嵌入,稳定了跨语言推理。通过这种替换,我们的模型可以生成具有中等节奏的演讲,而不考虑跨语言合成中的源说话人。在MOS评估中,SANE-TTS在跨语言和语内合成中的自然度得分均达到3.80以上,其中基本真实度得分为3.99。此外,即使在跨语言推理中,SANE-TTS也能保持说话人与背景真实的相似性。音频样本可在我们的网页上找到。
摘要:In this paper, we present SANE-TTS, a stable and natural end-to-end multilingual TTS model. By the difficulty of obtaining multilingual corpus for given speaker, training multilingual TTS model with monolingual corpora is unavoidable. We introduce speaker regularization loss that improves speech naturalness during cross-lingual synthesis as well as domain adversarial training, which is applied in other multilingual TTS models. Furthermore, by adding speaker regularization loss, replacing speaker embedding with zero vector in duration predictor stabilizes cross-lingual inference. With this replacement, our model generates speeches with moderate rhythm regardless of source speaker in cross-lingual synthesis. In MOS evaluation, SANE-TTS achieves naturalness score above 3.80 both in cross-lingual and intralingual synthesis, where the ground truth score is 3.99. Also, SANE-TTS maintains speaker similarity close to that of ground truth even in cross-lingual inference. Audio samples are available on our web page.


【7】 Data Augmentation and Squeeze-and-Excitation Network on Multiple  Dimension for Sound Event Localization and Detection in Real Scenes

标题:用于真实场景声事件定位和检测的多维数据增强和挤压激励网络

链接:https://arxiv.org/abs/2206.12059

* 与cs.SD语音【10】为同一篇

作者:Byeong-Yun Ko,Hyeonuk Nam,Seong-Hu Kim,Deokki Min,Seung-Deok Choi,Yong-Hwa Park

备注:Technical Report submitted for DCASE2022 Challenge Task3

摘要:真实场景中声音事件定位与检测(SELD)的性能受到SELD数据集较小的限制,因为难以获得足够数量的真实多通道音频数据记录和准确的标签。我们使用两种主要策略来解决小型真实SELD数据集产生的问题。首先,我们在所有数据维度上应用了各种数据增强方法:信道、频率和时间。我们还提出了一种称为适度混合的原始数据增强方法,以模拟存在噪声地板或干扰事件的情况。其次,我们在通道和频率维度上应用挤压和激励块来有效地提取特征。我们在STARSS22测试数据集上训练的模型的结果达到了最佳ER、F1、LE和LR,分别为0.53、49.8%、16.0度。,和56.2%。

摘要:Performance of sound event localization and detection (SELD) in real scenes is limited by small size of SELD dataset, due to difficulty in obtaining sufficient amount of realistic multi-channel audio data recordings with accurate label. We used two main strategies to solve problems arising from the small real SELD dataset. First, we applied various data augmentation methods on all data dimensions: channel, frequency and time. We also propose original data augmentation method named Moderate Mixup in order to simulate situations where noise floor or interfering events exist. Second, we applied Squeeze-and-Excitation block on channel and frequency dimensions to efficiently extract feature characteristics. Result of our trained models on the STARSS22 test dataset achieved the best ER, F1, LE, and LR of 0.53, 49.8%, 16.0deg., and 56.2% respectively.


【8】 Confidence Score Based Conformer Speaker Adaptation for Speech  Recognition

标题:基于置信度的语音识别整形说话人自适应

链接:https://arxiv.org/abs/2206.12045

* 与cs.SD语音【11】为同一篇

作者:Jiajun Deng,Xurong Xie,Tianzi Wang,Mingyu Cui,Boyang Xue,Zengrui Jin,Mengzhe Geng,Guinan Li,Xunying Liu,Helen Meng
备注:It's accepted to INTERSPEECH 2022. arXiv admin note: text overlap with arXiv:2206.11596
摘要:自动语音识别(ASR)系统面临的一个关键挑战是对说话人级别的可变性进行建模。本文采用紧凑的说话人相关学习隐藏单元贡献(LHUC)来促进基于一致性的端到端ASR系统的说话人自适应训练(SAT)和测试时无监督说话人自适应。使用基于可信度得分的选择说话人特定数据中更“可信”的子集,可以降低自适应过程中对监督错误率的敏感性。置信度估计模块用于在用作置信度得分之前平滑过度置信度一致性解码器的输出概率。利用LHUC参数的贝叶斯估计,解决了说话人级数据选择导致的数据稀疏性增加的问题。在300小时的交换机语料库上进行的实验表明,在NIST Hub5'00、RT02、,和RT03评估集。在使用外部Transformer和LSTM语言模型进行重新排序后,保持了一致的性能改进。
摘要:A key challenge for automatic speech recognition (ASR) systems is to model the speaker level variability. In this paper, compact speaker dependent learning hidden unit contributions (LHUC) are used to facilitate both speaker adaptive training (SAT) and test time unsupervised speaker adaptation for state-of-the-art Conformer based end-to-end ASR systems. The sensitivity during adaptation to supervision error rate is reduced using confidence score based selection of the more "trustworthy" subset of speaker specific data. A confidence estimation module is used to smooth the over-confident Conformer decoder output probabilities before serving as confidence scores. The increased data sparsity due to speaker level data selection is addressed using Bayesian estimation of LHUC parameters. Experiments on the 300-hour Switchboard corpus suggest that the proposed LHUC-SAT Conformer with confidence score based test time unsupervised adaptation outperformed the baseline speaker independent and i-vector adapted Conformer systems by up to 1.0%, 1.0%, and 1.2% absolute (9.0%, 7.9%, and 8.9% relative) word error rate (WER) reductions on the NIST Hub5'00, RT02, and RT03 evaluation sets respectively. Consistent performance improvements were retained after external Transformer and LSTM language models were used for rescoring.


【9】 End-to-End Text-to-Speech Based on Latent Representation of Speaking  Styles Using Spontaneous Dialogue

标题:基于自发对话的潜在语体表征的端到端文语转换

链接:https://arxiv.org/abs/2206.12040

* 与cs.SD语音【12】为同一篇

作者:Kentaro Mitsui,Tianyu Zhao,Kei Sawada,Yukiya Hono,Yoshihiko Nankaku,Keiichi Tokuda
备注:5 pages, 3 figures, accepted for INTERSPEECH 2022. Audio samples: this https URL
摘要:最近的文语转换(TTS)已经达到了与人类相当的质量;然而,它在口语对话中的应用还没有得到广泛的研究。本研究旨在实现一种与人类对话非常相似的TTS。首先,我们记录并转录真实的自发对话。然后,将所提出的对话TTS分为两个阶段进行训练:第一阶段,训练变分自动编码器(VAE)-VITS或高斯混合变分自动编码器(GMVAE)-VITS,该阶段通过端到端文本到语音(VITS)的对抗式学习将话语级潜变量引入到变分推理中,VITS是最近提出的端到端TTS模型。从语音中提取潜在说话风格表示的风格编码器与TTS联合训练。在第二阶段,训练一个风格预测因子来预测从对话历史中综合出来的说话风格。在推理过程中,通过将风格预测器预测的说话风格表示传递给VAE/GMVAE-VITS,可以以适合对话上下文的风格合成语音。主观评价结果表明,该方法在对话级自然度方面优于原VITS。
摘要:The recent text-to-speech (TTS) has achieved quality comparable to that of humans; however, its application in spoken dialogue has not been widely studied. This study aims to realize a TTS that closely resembles human dialogue. First, we record and transcribe actual spontaneous dialogues. Then, the proposed dialogue TTS is trained in two stages: first stage, variational autoencoder (VAE)-VITS or Gaussian mixture variational autoencoder (GMVAE)-VITS is trained, which introduces an utterance-level latent variable into variational inference with adversarial learning for end-to-end text-to-speech (VITS), a recently proposed end-to-end TTS model. A style encoder that extracts a latent speaking style representation from speech is trained jointly with TTS. In the second stage, a style predictor is trained to predict the speaking style to be synthesized from dialogue history. During inference, by passing the speaking style representation predicted by the style predictor to VAE/GMVAE-VITS, speech can be synthesized in a style appropriate to the context of the dialogue. Subjective evaluation results demonstrate that the proposed method outperforms the original VITS in terms of dialogue-level naturalness.


【10】 How to hide your voice: Noise-cancelling bird photography blind

标题:如何隐藏你的声音:消除噪音的鸟类摄影盲人

链接:https://arxiv.org/abs/2206.12340

* 与cs.SD语音【1】为同一篇

作者:C. Baydur,B. Pu,X. Xu
备注:22 pages, 11 figures
摘要:在野生动物摄影中,接近鸟类是一个巨大的挑战。鸟类摄影盲板可能是最有效、干扰最小的方法。如果设计得当,这些基本结构可以让摄影师在视觉上和听觉上远离栖息地。然而,百叶窗的声学设计被忽视了。在此,我们提出了噪声消除盲板,允许近距离拍摄鸟类。首先,我们在位于中国云南的生态旅游中心进行了问卷调查。因此,我们确定了观鸟者对室内声环境的期望。然后,我们确定了四个变量,以检查建筑和声学决策对噪声传播的影响。在Comsol MultiPhysics的声学模块中进行了数值模拟。在建筑设计过程中,减小结构尺寸和规划封闭窗户的建筑是减少噪音的正确决策。吸声材料可以降低室内的声能,从而降低室外噪声。隔音材料有助于消除室内到室外的声音传输。同时使用吸音和隔音材料是将室内外噪音降至最低的最佳方法。我们的研究表明,为了人类和鸟类的健康,摄影百叶窗需要强大而彻底的声学设计。
摘要:Getting close to birds is a great challenge in wildlife photography. Bird photography blinds may be the most effective and least intrusive way. These essential structures can allow to visually and audibly conceal photographers from the habitat if properly designed. However, the acoustic design of the blinds has been overlooked. Herein, we present noise-cancelling blinds which allow photographing birds at close range. Firstly, we conduct a questionnaire in the eco-tourism centre located in Yunnan, China. Thus, we determine the birders' expectations of the indoor sound environment. We then identify four variables to examine the impact of architectural and acoustic decisions on noise propagation. The numerical simulations are performed in the acoustic module of Comsol MultiPhysics. Minimizing the structural size and planning the building with closed windows is a proper decision to reduce noise in the architectural design process. Sound-absorbing materials reduce the acoustic energy indoors, thus decreasing the outdoor noise. Sound-proofing materials help to cancel the acoustic transmission indoors to outdoors. Using sound-absorbing and proofing materials together is the best way to minimize noise both indoors and outdoors. Our study demonstrated that photography blinds require a strong and thorough acoustic design for both human and bird well-being.


【11】 PoCaP Corpus: A Multimodal Dataset for Smart Operating Room Speech  Assistant using Interventional Radiology Workflow Analysis

标题:PoCaP语料库:用于介入放射工作流分析的智能手术室语音助手的多模式数据集

链接:https://arxiv.org/abs/2206.12320

* 与cs.SD语音【2】为同一篇

作者:Kubilay Can Demir,Matthias May,Axel Schmid,Michael Uder,Katharina Breininger,Tobias Weise,Andreas Maier,Seung Hee Yang
备注:8 pages, 4 figures, Text, Speech and Dialogue 2022 Conference
摘要:本文提出了一种新的多模式介入放射学数据集,称为PoCaP(端口导管放置)语料库。该语料库包括德语语音和音频信号、X射线图像和系统命令,这些信号由6名外科医生从31次PoCaP干预中收集,平均持续时间为81.4$\pm$41.0分钟。该语料库旨在为在手术室开发智能语音助手提供资源。尤其是,它可用于开发一种语音控制系统,使外科医生能够控制手术参数,如C形臂移动和手术台位置。为了记录数据集,我们获得了爱尔兰大学医院机构审查委员会和工人委员会以及患者对数据隐私的同意。我们描述了记录设置、数据结构、工作流程和预处理步骤,并使用预训练模型报告了第一个PoCaP语料库语音识别分析结果,错误率为11.52$\%$。研究结果表明,这些数据有可能构建一个强大的命令识别系统,并将允许开发一种新型的干预支持系统,该系统在医学领域使用语音和图像处理。
摘要:This paper presents a new multimodal interventional radiology dataset, called PoCaP (Port Catheter Placement) Corpus. This corpus consists of speech and audio signals in German, X-ray images, and system commands collected from 31 PoCaP interventions by six surgeons with average duration of 81.4 $\pm$ 41.0 minutes. The corpus aims to provide a resource for developing a smart speech assistant in operating rooms. In particular, it may be used to develop a speech controlled system that enables surgeons to control the operation parameters such as C-arm movements and table positions. In order to record the dataset, we acquired consent by the institutional review board and workers council in the University Hospital Erlangen and by the patients for data privacy. We describe the recording set-up, data structure, workflow and preprocessing steps, and report the first PoCaP Corpus speech recognition analysis results with 11.52 $\%$ word error rate using pretrained models. The findings suggest that the data has the potential to build a robust command recognition system and will allow the development of a novel intervention support systems using speech and image processing in the medical domain.


【12】 Deformable CNN and Imbalance-Aware Feature Learning for Singing  Technique Classification

标题:用于歌唱技巧分类的可变形CNN和不平衡感知特征学习

链接:https://arxiv.org/abs/2206.12230

* 与cs.SD语音【3】为同一篇

作者:Yuya Yamamoto,Juhan Nam,Hiroko Terasawa
备注:Accepted to INTERSPEECH2022
摘要:歌唱技巧是利用音色、音高和其他声音成分的时间波动来进行富有表现力的声乐表演。他们的分类是一项具有挑战性的任务,因为主要有两个因素:1)歌唱技巧的波动具有广泛的多样性,并受到许多因素的影响;2)现有数据集不平衡。为了解决这些问题,我们开发了一种新的基于变形卷积的音频特征学习方法,该方法使用类加权损失函数对特征抽取器和分类器进行解耦训练。实验结果表明:1)可变形卷积改进了分类结果,尤其是将其应用于最后两个卷积层时,以及2)重新训练分类器和通过平滑的逆频率加权交叉熵损失函数都提高了分类性能。
摘要:Singing techniques are used for expressive vocal performances by employing temporal fluctuations of the timbre, the pitch, and other components of the voice. Their classification is a challenging task, because of mainly two factors: 1) the fluctuations in singing techniques have a wide variety and are affected by many factors and 2) existing datasets are imbalanced. To deal with these problems, we developed a novel audio feature learning method based on deformable convolution with decoupled training of the feature extractor and the classifier using a class-weighted loss function. The experimental results show the following: 1) the deformable convolution improves the classification results, particularly when it is applied to the last two convolutional layers, and 2) both re-training the classifier and weighting the cross-entropy loss function by a smoothed inverse frequency enhance the classification performance.


【13】 Prosody Cloning in Zero-Shot Multispeaker Text-to-Speech

标题:零发声多人文语转换中的韵律克隆

链接:https://arxiv.org/abs/2206.12229

* 与cs.SD语音【4】为同一篇

作者:Florian Lux,Julia Koch,Ngoc Thang Vu

摘要:使用未翻译的参考样本克隆说话人的声音是现代神经文语转换(TTS)方法的重大进步之一。最近也提出了模仿转录参考音频韵律的方法。在这项工作中,我们首次通过话语级规范化和话语级说话人嵌入将这两项任务结合在一起。我们进一步介绍了一种用于提取细粒度韵律特征的轻量级对齐器,可以在几秒钟内对单个样本进行微调。我们的客观评估和人类研究表明,可以独立克隆说话人的声音以及口语参考的韵律,而不会降低质量,并且与原始声音和韵律高度相似。我们的所有代码和经过训练的模型都可用,还有静态和交互式演示。

摘要:The cloning of a speaker's voice using an untranscribed reference sample is one of the great advances of modern neural text-to-speech (TTS) methods. Approaches for mimicking the prosody of a transcribed reference audio have also been proposed recently. In this work, we bring these two tasks together for the first time through utterance level normalization in conjunction with an utterance level speaker embedding. We further introduce a lightweight aligner for extracting fine-grained prosodic features, that can be finetuned on individual samples within seconds. We show that it is possible to clone the voice of a speaker as well as the prosody of a spoken reference independently without any degradation in quality and high similarity to both original voice and prosody, as our objective evaluation and human study show. All of our code and trained models are available, alongside static and interactive demos.


【14】 BYOL-S: Learning Self-supervised Speech Representations by Bootstrapping

标题:BYOL-S:通过自举学习自监督语音表示

链接:https://arxiv.org/abs/2206.12038

* 与cs.SD语音【5】为同一篇

作者:Gasser Elbanna,Neil Scheidwasser-Clow,Mikolaj Kegler,Pierre Beckmann,Karl El Hajal,Milos Cernak

摘要:自几十年前频谱分析的开创性工作以来,音频和语音特征的提取方法一直在研究中。最近的努力以开发通用音频表示的雄心壮志为指导。例如,如果对大型音频数据集进行训练,深度神经网络可以提取最佳嵌入。这项工作扩展了现有的基于自引导自监督学习的方法,提出了各种编码器体系结构,并探讨了使用不同预训练数据集的效果。最后,我们提出了一个新的训练框架,提出了一种混合音频表示,它结合了手工制作和数据驱动的学习音频特征。在听觉场景分类和时间戳检测任务的听觉神经末梢2021挑战中,对所有提议的表示进行了评估。我们的结果表明,以卷积Transformer作为编码器的混合模型在大多数听力挑战任务中具有优异的性能。

摘要:Methods for extracting audio and speech features have been studied since pioneering work on spectrum analysis decades ago. Recent efforts are guided by the ambition to develop general-purpose audio representations. For example, deep neural networks can extract optimal embeddings if they are trained on large audio datasets. This work extends existing methods based on self-supervised learning by bootstrapping, proposes various encoder architectures, and explores the effects of using different pre-training datasets. Lastly, we present a novel training framework to come up with a hybrid audio representation, which combines handcrafted and data-driven learned audio features. All the proposed representations were evaluated within the HEAR NeurIPS 2021 challenge for auditory scene classification and timestamp detection tasks. Our results indicate that the hybrid model with a convolutional transformer as the encoder yields superior performance in most HEAR challenge tasks.


【15】 Comparing supervised and self-supervised embedding for ExVo Multi-Task  learning track

标题:ExVo多任务学习路径的监督嵌入和自监督嵌入的比较

链接:https://arxiv.org/abs/2206.11968

* 与cs.SD语音【6】为同一篇

作者:Tilak Purohit,Imen Ben Mahmoud,Bogdan Vlasenko,Mathew Magimai. -Doss
摘要:ICML表达性发声(ExVo)2022年多任务挑战赛侧重于理解非语言发声(发声爆发(VB))的情感方面。这项挑战的目标是预测VB的情绪强度,这是一项多任务挑战,还需要预测说话者的年龄和母语。针对这一挑战,我们研究并比较了两种不同的嵌入空间,即基于自监督学习(SSL)的嵌入和基于任务特定监督学习的嵌入。为此,我们研究了从几个预先训练的SSL神经网络和任务特定监督分类神经网络获得的特征表示。我们的研究表明,混合方法可以获得最佳性能,其中使用了通过SSL和任务特定监督学习得出的预测。我们在测试集上的最佳系统超过了比较基线(所有子任务得分的调和平均值,即$S\u{MTL}$),相对优势为13\%$。
摘要:The ICML Expressive Vocalizations (ExVo) Multi-task challenge 2022, focuses on understanding the emotional facets of the non-linguistic vocalizations (vocal bursts (VB)). The objective of this challenge is to predict emotional intensities for VB, being a multi-task challenge it also requires to predict speakers' age and native-country. For this challenge we study and compare two distinct embedding spaces namely, self-supervised learning (SSL) based embeddings and task-specific supervised learning based embeddings. Towards that, we investigate feature representations obtained from several pre-trained SSL neural networks and task-specific supervised classification neural networks. Our studies show that the best performance is obtained with a hybrid approach, where predictions derived via both SSL and task-specific supervised learning are used. Our best system on test-set surpasses the ComPARE baseline (harmonic mean of all sub-task scores i.e., $S_{MTL}$) by a relative $13\%$ margin.