今日论文合集:cs.SD语音12篇,eess.AS音频处理17篇。本文经arXiv每日学术速递授权转载
【1】Towards Environmental Preference Based Speech Enhancement For Individualised Multi-Modal Hearing Aids
标题:基于环境偏好的个性化多模态助听器语音增强
链接:https://arxiv.org/abs/2402.16757
作者:Jasper Kirton-Wingate,Shafique Ahmed,Adeel Hussain,Mandar Gogate,Kia Dashtipour,Jen-Cheng Hou,Tassadaq Hussain,Yu Tsao,Amir Hussain
备注:This has been submitted to the Trends in Hearing journal
摘要:自从深度学习(DL)出现以来,语音增强(SE)模型在各种噪声条件下都表现良好。然而,这样的系统仍然可能引入声音伪像、不自然的声音,并且限制用户听到可能是重要的环境声音的能力。助听器(HA)用户可能希望定制他们的SE系统,以适应他们的个人喜好和日常生活方式。在本文中,我们引入了一个基于偏好学习的SE(PLSE)模型,用于未来的多模态HA,可以根据用户的偏好,根据上下文利用音频信息来提高听觉舒适度。所提出的系统估计的信噪比(SNR)作为一个基本的客观语音质量的措施,量化的相对量的背景噪声存在于语音,并直接相关的信号的可懂度。此外,为了提供上下文信息,我们预测用户所处的声学场景。这些任务是通过多任务DL模型来实现的,该模型通过联合利用共享的编码特征空间来超越单独推断声学场景或SNR的性能。这些环境的推断中利用的偏好启发框架,线性学习一组预测功能,以确定AV(视听)SE系统的目标SNR。通过在具有挑战性的收听条件下大大降低噪声,并通过新颖地缩放SE模型的输出,我们能够为HA用户提供上下文个性化的SE。初步结果表明,在一些参与者中,非个体化基线模型有所改善。
摘要:Since the advent of Deep Learning (DL), Speech Enhancement (SE) models have performed well under a variety of noise conditions. However, such systems may still introduce sonic artefacts, sound unnatural, and restrict the ability for a user to hear ambient sound which may be of importance. Hearing Aid (HA) users may wish to customise their SE systems to suit their personal preferences and day-to-day lifestyle. In this paper, we introduce a preference learning based SE (PLSE) model for future multi-modal HAs that can contextually exploit audio information to improve listening comfort, based upon the preferences of the user. The proposed system estimates the Signal-to-noise ratio (SNR) as a basic objective speech quality measure which quantifies the relative amount of background noise present in speech, and directly correlates to the intelligibility of the signal. Additionally, to provide contextual information we predict the acoustic scene in which the user is situated. These tasks are achieved via a multi-task DL model, which surpasses the performance of inferring the acoustic scene or SNR separately, by jointly leveraging a shared encoded feature space. These environmental inferences are exploited in a preference elicitation framework, which linearly learns a set of predictive functions to determine the target SNR of an AV (Audio-Visual) SE system. By greatly reducing noise in challenging listening conditions, and by novelly scaling the output of the SE model, we are able to provide HA users with contextually individualised SE. Preliminary results suggest an improvement over the non-individualised baseline model in some participants.
【2】 Open Your Ears to Take a Look: A State-of-the-Art Report on the Integration of Sonification and Visualization标题:打开耳朵看一看:可听化与可视化融合的最新报告作者:Kajetan Enge,Elias Elmquist,Valentina Caiola,Niklas Rönnberg,Alexander Rind,Michael Iber,Sara Lenzi,Fangfei Lan,Robert Höldrich,Wolfgang Aigner备注:27 pages, 10 figures, submitted to EuroVis 2024 conference摘要:研究数据显示和分析的可视化和声音化的研究社区有着非常相似的目标,基本上使任何类型的数据都可以被人类解释。一个社区通过使用数据的视觉表示来这样做,另一个社区通过使用数据的听觉(非语音)表示来这样做。虽然这两个社区有很多共同点,但在过去几十年中,它们大多是平行发展的。通过这个STAR,我们讨论了一系列跨越两个社区边界的作品,因此,一系列旨在将这两种技术整合到一种视听展示形式中的作品,我们认为这是“超过两者的总和”。“我们引入并激励适用于此类视听显示的分类系统,并将2011年至2023年期间出现的57种学术出版物的语料库按阅读水平,数据集类型或评估系统等类别进行分类。该语料库还可以对该领域进行元分析,包括定期出现的设计模式,如可视化和声音处理技术的类型,或视觉和听觉通道的使用,以及对该领域的合著者网络的分析,该网络显示了没有太多相互联系的各个团队。本STAR涵盖的工作主体还涉及三个相邻的主题:视听监控,可访问性和视听数据艺术。除了本研究的系统进行部分外,还单独讨论了这三个主题。本报告的研究结果可供这两个领域的研究人员使用,以了解这种集成设计的潜力和挑战,同时激励他们未来与其他领域的专家合作。摘要:The research communities studying visualization and sonification for data display and analysis share exceptionally similar goals, essentially making data of any kind interpretable to humans. One community does so by using visual representations of data, the other community does so by employing auditory (non-speech) representations of data. While the two communities have a lot in common, they developed mostly in parallel over the course of the last few decades. With this STAR, we discuss a collection of work that bridges the borders of the two communities, hence a collection of work that aims to integrate the two techniques to one form of audiovisual display, which we argue to be "more than the sum of the two." We introduce and motivate a classification system applicable to such audiovisual displays and categorize a corpus of 57 academic publications that appeared between 2011 and 2023 in categories such as reading level, dataset type, or evaluation system, to mention a few. The corpus also enables a meta-analysis of the field, including regularly occurring design patterns such as type of visualization and sonification techniques, or the use of visual and auditory channels, and the analysis of a co-author network of the field which shows individual teams without much interconnection. The body of work covered in this STAR also relates to three adjacent topics: audiovisual monitoring, accessibility, and audiovisual data art. These three topics are discussed individually in addition to the systematically conducted part of this research. The findings of this report may be used by researchers from both fields to understand the potentials and challenges of such integrated designs, while inspiring them for future collaboration with experts from the respective other field.【3】 Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech作者:Szu-Wei Fu,Kuo-Hsuan Hung,Yu Tsao,Yu-Chiang Frank Wang备注:Published as a conference paper at ICLR 2024摘要:语音质量估计最近经历了从人类听觉专家设计到机器学习模型的范式转变。然而,目前的模型主要依赖于监督学习,这是耗时和昂贵的标签收集。为了解决这个问题,我们提出了VQScore,一种基于矢量量化变分自编码器(VQ-VAE)的量化误差的自监督语音评估度量。VQ-VAE的训练依赖于干净的语音;因此,当语音失真时,可以预期大的量化误差。为了进一步提高与真实质量分数的相关性,将语音处理的领域知识并入模型设计中。我们发现,矢量量化机制也可以用于自监督语音增强(SE)模型训练。为了提高编码器对SE的鲁棒性,引入了一种新的结合对抗训练的自蒸馏机制。总之,所提出的语音质量估计方法和增强模型只需要干净的语音进行训练,而没有任何标签要求。实验结果表明,所提出的VQScore和增强模型与监督基线相比具有竞争力。代码将在发布后发布。摘要:Speech quality estimation has recently undergone a paradigm shift from human-hearing expert designs to machine-learning models. However, current models rely mainly on supervised learning, which is time-consuming and expensive for label collection. To solve this problem, we propose VQScore, a self-supervised metric for evaluating speech based on the quantization error of a vector-quantized-variational autoencoder (VQ-VAE). The training of VQ-VAE relies on clean speech; hence, large quantization errors can be expected when the speech is distorted. To further improve correlation with real quality scores, domain knowledge of speech processing is incorporated into the model design. We found that the vector quantization mechanism could also be used for self-supervised speech enhancement (SE) model training. To improve the robustness of the encoder for SE, a novel self-distillation mechanism combined with adversarial training is introduced. In summary, the proposed speech quality estimation method and enhancement models require only clean speech for training without any label requirements. Experimental results show that the proposed VQScore and enhancement model are competitive with supervised baselines. The code will be released after publication.【4】 ChatMusician: Understanding and Generating Music Intrinsically with LLM标题:ChatMusic:用LLM内在地理解和生成音乐作者:Ruibin Yuan,Hanfeng Lin,Yi Wang,Zeyue Tian,Shangda Wu,Tianhao Shen,Ge Zhang,Yuhang Wu,Cong Liu,Ziya Zhou,Ziyang Ma,Liumeng Xue,Ziyu Wang,Qin Liu,Tianyu Zheng,Yizhi Li,Yinghao Ma,Yiming Liang,Xiaowei Chi,Ruibo Liu,Zili Wang,Pengfei Li,Jingcheng Wu,Chenghua Lin,Qifeng Liu,Tao Jiang,Wenhao Huang,Wenhu Chen,Emmanouil Benetos,Jie Fu,Gus Xia,Roger Dannenberg,Wei Xue,Shiyin Kang,Yike Guo备注:GitHub: this https URL摘要:虽然大型语言模型(LLM)在文本生成方面表现出令人印象深刻的能力,但我们发现它们的能力尚未推广到音乐,人类的创造性语言。我们介绍ChatMusician,一个开源的LLM,集成了内在的音乐能力。它基于对文本兼容的音乐表示(ABC记谱法)的持续预训练和微调LLaMA 2,音乐被视为第二语言。ChatMusician可以使用纯文本标记器来理解和生成音乐,而无需任何外部多模态神经结构或标记器。有趣的是,赋予音乐能力并不会损害语言能力,甚至可以获得略高的MMLU分数。我们的模型能够创作结构良好的全长音乐,以文本,和弦,旋律,主题,音乐形式等为条件,超过GPT-4基线。在我们精心策划的大学水平的音乐理解基准,MusicTheoryBench,ChatMusician超过LLaMA 2和GPT-3.5的zero-shot设置的显着保证金。我们的工作表明,LLM可以是一个很好的音乐压缩机,但仍然有显着的领域有待征服。我们在GitHub中发布了我们的4 B代币音乐语言语料库MusicPile,收集的MusicTheoryBench,代码,模型和演示。摘要:While Large Language Models (LLMs) demonstrate impressive capabilities in text generation, we find that their ability has yet to be generalized to music, humanity's creative language. We introduce ChatMusician, an open-source LLM that integrates intrinsic musical abilities. It is based on continual pre-training and finetuning LLaMA2 on a text-compatible music representation, ABC notation, and the music is treated as a second language. ChatMusician can understand and generate music with a pure text tokenizer without any external multi-modal neural structures or tokenizers. Interestingly, endowing musical abilities does not harm language abilities, even achieving a slightly higher MMLU score. Our model is capable of composing well-structured, full-length music, conditioned on texts, chords, melodies, motifs, musical forms, etc, surpassing GPT-4 baseline. On our meticulously curated college-level music understanding benchmark, MusicTheoryBench, ChatMusician surpasses LLaMA2 and GPT-3.5 on zero-shot setting by a noticeable margin. Our work reveals that LLMs can be an excellent compressor for music, but there remains significant territory to be conquered. We release our 4B token music-language corpora MusicPile, the collected MusicTheoryBench, code, model and demo in GitHub.【5】 Phonetic and Lexical Discovery of a Canine Language using HuBERT作者:Xingyuan Li,Sinong Wang,Zeyu Xie,Mengyue Wu,Kenny Q. Zhu摘要:本文深入探讨了狗发声中潜在的通信模式的开拓性探索,并超越了传统的语言分析障碍,这种障碍严重依赖于人类对有限数据集的先验知识来寻找狗发声中的声音单元。我们提出了一个自我监督的方法与休伯特,使准确的音素标签分类和识别的声音模式,建议狗发声的基本词汇。我们的研究结果表明,在这些确定的犬词汇,涵盖了整个观察到的狗发声序列的声学一致性。我们进一步开发了一个基于网络的狗叫声标注系统。该系统可以突出显示用户上传的狗音频中存在于词汇表中的音素n-grams。摘要:This paper delves into the pioneering exploration of potential communication patterns within dog vocalizations and transcends traditional linguistic analysis barriers, which heavily relies on human priori knowledge on limited datasets to find sound units in dog vocalization. We present a self-supervised approach with HuBERT, enabling the accurate classification of phoneme labels and the identification of vocal patterns that suggest a rudimentary vocabulary within dog vocalizations. Our findings indicate a significant acoustic consistency in these identified canine vocabulary, covering the entirety of observed dog vocalization sequences. We further develop a web-based dog vocalization labeling system. This system can highlight phoneme n-grams, present in the vocabulary, in the dog audio uploaded by users.【6】 Direct Punjabi to English speech translation using discrete units作者:Prabhjot Kaur,L. Andrew M. Bush,Weisong Shi摘要:语音到语音翻译尚未达到与文本到文本翻译系统相同的覆盖水平。目前的语音技术在覆盖全世界7000多种语言方面非常有限,使一半以上的人口被剥夺了这种技术和共享经验。随着语音辅助技术(如社交机器人和语音转文本应用程序)和听觉内容(如播客和讲座)的兴起,确保所有人都能使用这项技术比以往任何时候都更加重要。语音翻译可以在缩小技术差距和创造更具包容性的社会方面发挥至关重要的作用。为了促进低资源语言的语音翻译研究,我们的工作提出了一个直接的语音到语音翻译模型的印度语言之一,称为旁遮普语到英语。此外,我们探讨了使用称为离散声学单元的离散语音表示作为基于transformer的翻译模型的输入的性能。该模型,简称为单元到单元翻译(U2UT),采用源语言(被翻译的语言)的离散单元序列,并输出目标语言(被翻译的语言)的离散单元序列。我们的研究结果表明,U2UT模型的性能优于语音到单元翻译(S2UT)模型,BLEU得分为3.69。摘要:Speech-to-speech translation is yet to reach the same level of coverage as text-to-text translation systems. The current speech technology is highly limited in its coverage of over 7000 languages spoken worldwide, leaving more than half of the population deprived of such technology and shared experiences. With voice-assisted technology (such as social robots and speech-to-text apps) and auditory content (such as podcasts and lectures) on the rise, ensuring that the technology is available for all is more important than ever. Speech translation can play a vital role in mitigating technological disparity and creating a more inclusive society. With a motive to contribute towards speech translation research for low-resource languages, our work presents a direct speech-to-speech translation model for one of the Indic languages called Punjabi to English. Additionally, we explore the performance of using a discrete representation of speech called discrete acoustic units as input to the Transformer-based translation model. The model, abbreviated as Unit-to-Unit Translation (U2UT), takes a sequence of discrete units of the source language (the language being translated from) and outputs a sequence of discrete units of the target language (the language being translated to). Our results show that the U2UT model performs better than the Speech-to-Unit Translation (S2UT) model by a 3.69 BLEU score.
【7】 Alternating Weak Triphone/BPE Alignment Supervision from Hybrid Model Improves End-to-End ASR标题:基于混合模型的交替弱三音素/BPE对齐监控改善端到端ASR作者:Jintao Jiang,Yingbo Gao,Mohammad Zeineldeen,Zoltan Tuske备注:5 pages, 1 figure, 3 tables摘要:本文提出了交替弱三音子/BPE对齐监督来改进端到端模型训练。为此,三音子和BPE对齐提取使用预先存在的混合ASR系统。然后,通过在用于三音子对准的编码器的中间层表示处和用于BPE对准的编码器处对这样的对准计算基于交叉熵的中间辅助损失来获得正则化效果。弱监督通过参数为0.5的强标签平滑来实现。TED-LIUM 2上的实验结果表明,无论是基于三音子或BPE对齐的弱监督提高ASR性能超过标准CTC辅助损失。此外,它们的组合进一步降低了字错误率。我们还研究了模型训练过程中两个辅助任务的交替,并观察到额外的性能增益。总的来说,所提出的技术导致超过10%的相对错误率降低超过CTC正则化基线系统。摘要:In this paper, alternating weak triphone/BPE alignment supervision is proposed to improve end-to-end model training. Towards this end, triphone and BPE alignments are extracted using a pre-existing hybrid ASR system. Then, regularization effect is obtained by cross-entropy based intermediate auxiliary losses computed on such alignments at a mid-layer representation of the encoder for triphone alignments and at the encoder for BPE alignments. Weak supervision is achieved through strong label smoothing with parameter of 0.5. Experimental results on TED-LIUM 2 indicate that either triphone or BPE alignment based weak supervision improves ASR performance over standard CTC auxiliary loss. Moreover, their combination lowers the word error rate further. We also investigate the alternation of the two auxiliary tasks during model training, and additional performance gain is observed. Overall, the proposed techniques result in over 10% relative error rate reduction over a CTC-regularized baseline system.
【8】 GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model标题:GLA-Grad:一种扩展的Griffin-Lim波形产生扩散模型作者:Haocheng Liu,Teysir Baoueb,Mathieu Fontaine,Jonathan Le Roux,Gael Richard摘要:扩散模型在语音或音乐合成等各种信号生成任务中受到越来越多的关注。例如,WaveGrad是一个成功的扩散模型,它有条件地使用梅尔频谱图来指导用于生成高保真音频的扩散过程。然而,这些模型面临着关于训练和推理的噪声扩散过程的重要挑战,并且它们难以为训练期间未看到的说话者生成高质量的语音。为了最大限度地减少条件误差和提高噪声扩散过程的效率,本文提出了一种新的方案,称为GLA-Grad,它包括在定期扩散过程的每一步引入一个相位恢复算法,如Griffin-Lim算法(GLA)。此外,它可以直接应用于已经训练过的波形生成模型,而无需额外的训练或微调。我们表明,我们的算法优于国家的最先进的扩散模型的语音生成,特别是当生成语音为以前看不见的目标扬声器。摘要:Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectrogram to guide a diffusion process for the generation of high-fidelity audio. However, such models face important challenges concerning the noise diffusion process for training and inference, and they have difficulty generating high-quality speech for speakers that were not seen during training. With the aim of minimizing the conditioning error and increasing the efficiency of the noise diffusion process, we propose in this paper a new scheme called GLA-Grad, which consists in introducing a phase recovery algorithm such as the Griffin-Lim algorithm (GLA) at each step of the regular diffusion process. Furthermore, it can be directly applied to an already-trained waveform generation model, without additional training or fine-tuning. We show that our algorithm outperforms state-of-the-art diffusion models for speech generation, especially when generating speech for a previously unseen target speaker.【9】 SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning作者:Luca Zampierin,Ghouthi Boukli Hacene,Bac Nguyen,Mirco Ravanelli备注:Accepted at the Self-supervision in Audio, Speech and Beyond (SASB) Workshop at ICASSP 2024摘要:自监督学习(SSL)在各种语音处理任务中取得了显着的成功。为了提高其效率,以前的作品往往利用压缩技术的使用。最近一个值得注意的尝试是DPHuBERT,它应用联合知识蒸馏(KD)和结构化修剪来学习一个明显较小的SSL模型。在本文中,我们通过引入SKILL,这是一种新的方法,它在教师网络中进行跨层组的蒸馏,而不是蒸馏单个任意选择的层,从而为这一研究领域做出贡献。通过应用于层相似性度量的分层聚类过程来实现要提取的层的识别。大量的实验表明,我们的WavLM Base+的简化版本不仅优于DPHuBERT,而且在几个SUPERB任务中的30M参数模型类中获得了最先进的结果。摘要:Self-supervised learning (SSL) has achieved remarkable success across various speech-processing tasks. To enhance its efficiency, previous works often leverage the use of compression techniques. A notable recent attempt is DPHuBERT, which applies joint knowledge distillation (KD) and structured pruning to learn a significantly smaller SSL model. In this paper, we contribute to this research domain by introducing SKILL, a novel method that conducts distillation across groups of layers instead of distilling individual arbitrarily selected layers within the teacher network. The identification of the layers to distill is achieved through a hierarchical clustering procedure applied to layer similarity measures. Extensive experiments demonstrate that our distilled version of WavLM Base+ not only outperforms DPHuBERT but also achieves state-of-the-art results in the 30M parameters model class across several SUPERB tasks.
【10】 Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues标题:基于情感的语境线索在噪声环境下的视听语音增强作者:Tassadaq Hussain,Kia Dashtipour,Yu Tsao,Amir Hussain摘要:在现实世界的环境中,背景噪声显着降低人类语音的可懂度和清晰度。视听语音增强(AVSE)试图恢复语音质量,但现有的方法往往达不到,特别是在动态噪声条件下。本研究调查包括情绪作为一种新的上下文线索内AVSE,假设,将情绪理解可以提高语音增强性能。我们提出了一种新的情感感知AVSE系统,利用听觉和视觉信息。它从说话人的面部标志中提取情感特征,并将其与相应的音频和视觉模态融合。这种丰富的数据作为基于深度UNet的编码器-解码器网络的输入,该网络专门设计用于协调情感增强的多模态信息的融合。该网络通过编码器-解码器架构迭代地改进增强的语音表示,由感知启发的损失函数指导,用于联合学习和优化。我们训练和评估CMU多模态意见情绪和情绪强度(CMU-MOSEI)数据集,一个丰富的音频视频记录与注释的情绪库的模型。我们的综合评估证明了情绪作为AVSE的上下文线索的有效性。通过整合情感特征,所提出的系统实现了显着改善语音质量和可懂度的客观和主观评估,特别是在具有挑战性的噪声环境。与基线AVSE和纯音频语音增强系统相比,我们的方法在PESQ和STOI方面表现出明显的增加,表明更高的感知质量和可懂度。大规模的听力测试证实了这些发现,表明人类对增强语音的理解有所改善。摘要:In real-world environments, background noise significantly degrades the intelligibility and clarity of human speech. Audio-visual speech enhancement (AVSE) attempts to restore speech quality, but existing methods often fall short, particularly in dynamic noise conditions. This study investigates the inclusion of emotion as a novel contextual cue within AVSE, hypothesizing that incorporating emotional understanding can improve speech enhancement performance. We propose a novel emotion-aware AVSE system that leverages both auditory and visual information. It extracts emotional features from the facial landmarks of the speaker and fuses them with corresponding audio and visual modalities. This enriched data serves as input to a deep UNet-based encoder-decoder network, specifically designed to orchestrate the fusion of multimodal information enhanced with emotion. The network iteratively refines the enhanced speech representation through an encoder-decoder architecture, guided by perceptually-inspired loss functions for joint learning and optimization. We train and evaluate the model on the CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) dataset, a rich repository of audio-visual recordings with annotated emotions. Our comprehensive evaluation demonstrates the effectiveness of emotion as a contextual cue for AVSE. By integrating emotional features, the proposed system achieves significant improvements in both objective and subjective assessments of speech quality and intelligibility, especially in challenging noise environments. Compared to baseline AVSE and audio-only speech enhancement systems, our approach exhibits a noticeable increase in PESQ and STOI, indicating higher perceptual quality and intelligibility. Large-scale listening tests corroborate these findings, suggesting improved human understanding of enhanced speech.
【11】 A circular microphone array with virtual microphones based on acoustics-informed neural networks标题:基于声学信息神经网络的虚拟麦克风圆形麦克风阵列备注:Submitted to JASA on 24/02/2024摘要:声波束形成的目的是将声信号聚焦到特定方向,并抑制来自其他方向的不期望的干扰。尽管具有灵活性和可操纵性,但使用圆形麦克风阵列的波束成形在对应于贝塞尔函数的零点的频率处遭受显著的性能降级。为了克服这一限制,已经研究了挡板式或同心圆形麦克风阵列;然而,前者需要干扰原始声场的笨重挡板,而后者需要更多的麦克风,这增加了复杂性和成本,这两者在实际应用中都是不期望的。为了解决这个问题,本文提出了一种圆形麦克风阵列配备虚拟麦克风,它解决了性能下降通常与圆形麦克风阵列,而不诉诸物理修改。基于声学信息神经网络从由物理麦克风测量的声压预测虚拟麦克风处的声压,并且然后将由物理麦克风测量的声压和在虚拟麦克风处预测的声压集成以设计波束形成器。实验结果表明,该方法不仅消除了性能下降,但也抑制了在高频空间混叠,从而强调其有前途的潜力。摘要:Acoustic beamforming aims to focus acoustic signals to a specific direction and suppress undesirable interferences from other directions. Despite its flexibility and steerability, beamforming with circular microphone arrays suffers from significant performance degradation at frequencies corresponding to zeros of the Bessel functions. To conquer this constraint, baffled or concentric circular microphone arrays have been studied; however, the former needs a bulky baffle that interferes with the original sound field whereas the latter requires more microphones that increase the complexity and cost, both of which are undesirable in practical applications. To tackle this challenge, this paper proposes a circular microphone array equipped with virtual microphones, which resolves the performance degradation commonly associated with circular microphone arrays without resorting to physical modifications. The sound pressures at the virtual microphones are predicted from those measured by the physical microphones based on an acoustics-informed neural network, and then the sound pressures measured by the physical microphones and those predicted at the virtual microphones are integrated to design the beamformer. Experimental results demonstrate that the proposed approach not only eliminates the performance degradation but also suppresses spatial aliasing at high frequencies, thereby underscoring its promising potential.
【12】 Toward Fully Self-Supervised Multi-Pitch Estimation作者:Frank Cwitkowitz,Zhiyao Duan摘要:多音高估计是一个长达数十年的研究问题,涉及检测与多乐器混合物中的并发音乐事件相关联的音高活动。监督学习技术在更窄的任务特征上表现出了良好的性能,但受到缺乏具有多音高注释的大规模和多样化复调音乐数据集的限制。我们提出了一套用于多音高估计的自监督学习目标,其鼓励围绕谐波、音色变换的不变性和几何变换的等变性的集中支持。这些目标足以训练完全卷积的自动编码器,以直接产生多音调显著图,而无需任何微调。尽管只在合成单音符音频样本的集合上进行训练,但我们的完全自监督框架可以推广到复调音乐混合,并实现了与在传统多音高数据集上训练的监督模型相当的性能。摘要:Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supervised learning techniques have demonstrated solid performance on more narrow characterizations of the task, but suffer from limitations concerning the shortage of large-scale and diverse polyphonic music datasets with multi-pitch annotations. We present a suite of self-supervised learning objectives for multi-pitch estimation, which encourage the concentration of support around harmonics, invariance to timbral transformations, and equivariance to geometric transformations. These objectives are sufficient to train an entirely convolutional autoencoder to produce multi-pitch salience-grams directly, without any fine-tuning. Despite training exclusively on a collection of synthetic single-note audio samples, our fully self-supervised framework generalizes to polyphonic music mixtures, and achieves performance comparable to supervised models trained on conventional multi-pitch datasets.
【1】 SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning作者:Luca Zampierin,Ghouthi Boukli Hacene,Bac Nguyen,Mirco Ravanelli备注:Accepted at the Self-supervision in Audio, Speech and Beyond (SASB) Workshop at ICASSP 2024摘要:自监督学习(SSL)在各种语音处理任务中取得了显着的成功。为了提高其效率,以前的作品往往利用压缩技术的使用。最近一个值得注意的尝试是DPHuBERT,它应用联合知识蒸馏(KD)和结构化修剪来学习一个明显较小的SSL模型。在本文中,我们通过引入SKILL,这是一种新的方法,它在教师网络中进行跨层组的蒸馏,而不是蒸馏单个任意选择的层,从而为这一研究领域做出贡献。通过应用于层相似性度量的分层聚类过程来实现要提取的层的识别。大量的实验表明,我们的WavLM Base+的简化版本不仅优于DPHuBERT,而且在几个SUPERB任务中的30M参数模型类中获得了最先进的结果。摘要:Self-supervised learning (SSL) has achieved remarkable success across various speech-processing tasks. To enhance its efficiency, previous works often leverage the use of compression techniques. A notable recent attempt is DPHuBERT, which applies joint knowledge distillation (KD) and structured pruning to learn a significantly smaller SSL model. In this paper, we contribute to this research domain by introducing SKILL, a novel method that conducts distillation across groups of layers instead of distilling individual arbitrarily selected layers within the teacher network. The identification of the layers to distill is achieved through a hierarchical clustering procedure applied to layer similarity measures. Extensive experiments demonstrate that our distilled version of WavLM Base+ not only outperforms DPHuBERT but also achieves state-of-the-art results in the 30M parameters model class across several SUPERB tasks.
【2】 Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues标题:基于情感的语境线索在噪声环境下的视听语音增强作者:Tassadaq Hussain,Kia Dashtipour,Yu Tsao,Amir Hussain摘要:在现实世界的环境中,背景噪声显着降低人类语音的可懂度和清晰度。视听语音增强(AVSE)试图恢复语音质量,但现有的方法往往达不到,特别是在动态噪声条件下。本研究调查包括情绪作为一种新的上下文线索内AVSE,假设,将情绪理解可以提高语音增强性能。我们提出了一种新的情感感知AVSE系统,利用听觉和视觉信息。它从说话人的面部标志中提取情感特征,并将其与相应的音频和视觉模态融合。这种丰富的数据作为基于深度UNet的编码器-解码器网络的输入,该网络专门设计用于协调情感增强的多模态信息的融合。该网络通过编码器-解码器架构迭代地改进增强的语音表示,由感知启发的损失函数指导,用于联合学习和优化。我们训练和评估CMU多模态意见情绪和情绪强度(CMU-MOSEI)数据集,一个丰富的音频视频记录与注释的情绪库的模型。我们的综合评估证明了情绪作为AVSE的上下文线索的有效性。通过整合情感特征,所提出的系统实现了显着改善语音质量和可懂度的客观和主观评估,特别是在具有挑战性的噪声环境。与基线AVSE和纯音频语音增强系统相比,我们的方法在PESQ和STOI方面表现出明显的增加,表明更高的感知质量和可懂度。大规模的听力测试证实了这些发现,表明人类对增强语音的理解有所改善。摘要:In real-world environments, background noise significantly degrades the intelligibility and clarity of human speech. Audio-visual speech enhancement (AVSE) attempts to restore speech quality, but existing methods often fall short, particularly in dynamic noise conditions. This study investigates the inclusion of emotion as a novel contextual cue within AVSE, hypothesizing that incorporating emotional understanding can improve speech enhancement performance. We propose a novel emotion-aware AVSE system that leverages both auditory and visual information. It extracts emotional features from the facial landmarks of the speaker and fuses them with corresponding audio and visual modalities. This enriched data serves as input to a deep UNet-based encoder-decoder network, specifically designed to orchestrate the fusion of multimodal information enhanced with emotion. The network iteratively refines the enhanced speech representation through an encoder-decoder architecture, guided by perceptually-inspired loss functions for joint learning and optimization. We train and evaluate the model on the CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) dataset, a rich repository of audio-visual recordings with annotated emotions. Our comprehensive evaluation demonstrates the effectiveness of emotion as a contextual cue for AVSE. By integrating emotional features, the proposed system achieves significant improvements in both objective and subjective assessments of speech quality and intelligibility, especially in challenging noise environments. Compared to baseline AVSE and audio-only speech enhancement systems, our approach exhibits a noticeable increase in PESQ and STOI, indicating higher perceptual quality and intelligibility. Large-scale listening tests corroborate these findings, suggesting improved human understanding of enhanced speech.
【3】 An Automated End-to-End Open-Source Software for High-Quality Text-to-Speech Dataset Generation标题:用于高质量文语转换数据集生成的自动化端到端开源软件作者:Ahmet Gunduz,Kamer Ali Yuksel,Kareem Darwish,Golara Javadi,Fabio Minazzi,Nicola Sobieski,Sebastien Bratieres备注:9 Pages, 6 Figures, 4 Tables, LREC-COLING 2024摘要:数据可用性对于推进人工智能应用(包括基于语音的技术)至关重要。随着内容创作,特别是社交媒体中的内容创作需求不断增加,翻译和文本到语音(TTS)技术已成为必不可少的工具。值得注意的是,这些TTS技术的性能高度依赖于训练数据的质量,强调数据可用性和技术进步的相互依赖性。本文介绍了一种端到端工具,用于为文本到语音(TTS)模型生成高质量的数据集,以满足对高质量数据的这一关键需求。这项工作的贡献是多方面的,包括:将特定语言的音素分布整合到样本选择中,录音过程的自动化,录音的自动化和人在回路的质量保证,以及处理录音以满足特定的格式。拟议的应用程序旨在通过这些功能简化TTS模型的数据集创建过程,从而促进基于语音的技术的进步。摘要:Data availability is crucial for advancing artificial intelligence applications, including voice-based technologies. As content creation, particularly in social media, experiences increasing demand, translation and text-to-speech (TTS) technologies have become essential tools. Notably, the performance of these TTS technologies is highly dependent on the quality of the training data, emphasizing the mutual dependence of data availability and technological progress. This paper introduces an end-to-end tool to generate high-quality datasets for text-to-speech (TTS) models to address this critical need for high-quality data. The contributions of this work are manifold and include: the integration of language-specific phoneme distribution into sample selection, automation of the recording process, automated and human-in-the-loop quality assurance of recordings, and processing of recordings to meet specified formats. The proposed application aims to streamline the dataset creation process for TTS models through these features, thereby facilitating advancements in voice-based technologies.【4】 Exploring the Power of Pure Attention Mechanisms in Blind Room Parameter Estimation作者:Chunxi Wang,Maoshen Jia,Meiran Li,Changchun Bao,Wenyu Jin备注:27 pages, 9 figures, submitted to EURASIP Journal On Audio Speech And Music Processing摘要:声环境的动态参数化在音频处理领域引起了广泛的关注。在为各种音频渲染应用设计音频滤波器时,精确表示局部房间声学特性至关重要。在这种情况下的关键参数包括混响时间(RT 60)和几何房间体积。近年来,神经网络在盲室参数估计中得到了广泛的应用。然而,仍然存在一个问题,即纯注意机制是否可以在这项任务中取得优异的成绩。为了解决这个问题,本研究采用盲室参数估计的基础上单声道噪声语音信号。各种模型架构进行了研究,包括建议的注意力为基础的模型。该模型是一个无卷积的音频频谱图Transformer,利用补丁分裂,注意力机制和来自预训练的Vision Transformer的跨模态迁移学习。实验结果表明,所提出的基于注意力机制的模型,纯粹依赖于注意力机制而不使用卷积,在各种房间参数估计任务中表现出显着提高的性能,特别是在专用预训练和数据增强方案的帮助下。此外,与现有方法相比,该模型在处理可变长度音频输入时表现出更有利的适应性和鲁棒性。摘要:Dynamic parameterization of acoustic environments has drawn widespread attention in the field of audio processing. Precise representation of local room acoustic characteristics is crucial when designing audio filters for various audio rendering applications. Key parameters in this context include reverberation time (RT60) and geometric room volume. In recent years, neural networks have been extensively applied in the task of blind room parameter estimation. However, there remains a question of whether pure attention mechanisms can achieve superior performance in this task. To address this issue, this study employs blind room parameter estimation based on monaural noisy speech signals. Various model architectures are investigated, including a proposed attention-based model. This model is a convolution-free Audio Spectrogram Transformer, utilizing patch splitting, attention mechanisms, and cross-modality transfer learning from a pretrained Vision Transformer. Experimental results suggest that the proposed attention mechanism-based model, relying purely on attention mechanisms without using convolution, exhibits significantly improved performance across various room parameter estimation tasks, especially with the help of dedicated pretraining and data augmentation schemes. Additionally, the model demonstrates more advantageous adaptability and robustness when handling variable-length audio inputs compared to existing methods.【5】 A circular microphone array with virtual microphones based on acoustics-informed neural networks标题:基于声学信息神经网络的虚拟麦克风圆形麦克风阵列备注:Submitted to JASA on 24/02/2024摘要:声波束形成的目的是将声信号聚焦到特定方向,并抑制来自其他方向的不期望的干扰。尽管具有灵活性和可操纵性,但使用圆形麦克风阵列的波束成形在对应于贝塞尔函数的零点的频率处遭受显著的性能降级。为了克服这一限制,已经研究了挡板式或同心圆形麦克风阵列;然而,前者需要干扰原始声场的笨重挡板,而后者需要更多的麦克风,这增加了复杂性和成本,这两者在实际应用中都是不期望的。为了解决这个问题,本文提出了一种圆形麦克风阵列配备虚拟麦克风,它解决了性能下降通常与圆形麦克风阵列,而不诉诸物理修改。基于声学信息神经网络从由物理麦克风测量的声压预测虚拟麦克风处的声压,并且然后将由物理麦克风测量的声压和在虚拟麦克风处预测的声压集成以设计波束形成器。实验结果表明,该方法不仅消除了性能下降,但也抑制了在高频空间混叠,从而强调其有前途的潜力。摘要:Acoustic beamforming aims to focus acoustic signals to a specific direction and suppress undesirable interferences from other directions. Despite its flexibility and steerability, beamforming with circular microphone arrays suffers from significant performance degradation at frequencies corresponding to zeros of the Bessel functions. To conquer this constraint, baffled or concentric circular microphone arrays have been studied; however, the former needs a bulky baffle that interferes with the original sound field whereas the latter requires more microphones that increase the complexity and cost, both of which are undesirable in practical applications. To tackle this challenge, this paper proposes a circular microphone array equipped with virtual microphones, which resolves the performance degradation commonly associated with circular microphone arrays without resorting to physical modifications. The sound pressures at the virtual microphones are predicted from those measured by the physical microphones based on an acoustics-informed neural network, and then the sound pressures measured by the physical microphones and those predicted at the virtual microphones are integrated to design the beamformer. Experimental results demonstrate that the proposed approach not only eliminates the performance degradation but also suppresses spatial aliasing at high frequencies, thereby underscoring its promising potential.
【6】 Text-guided HuBERT: Self-Supervised Speech Pre-training via Generative Adversarial Networks标题:文本引导的HuBERT:通过生成对抗网络进行自监督语音预训练作者:Duo Ma,Xianghu Yue,Junyi Ao,Xiaoxue Gao,Haizhou Li备注:5 pages, 1 figures,5 tables, submit to IEEE Signal Processing Letters(SPL)摘要:人类语言可以用书面或口头形式表达,即文本或语音。人类可以从文本中获取知识,以提高口语和听力。然而,对语音预训练模型的探索才刚刚开始。在本文中,我们研究了一种预训练这种联合语音-文本模型的新方法,以学习增强的语音表示并使各种语音相关的下游任务受益。具体来说,我们提出了一种新的预训练方法,文本引导的HuBERT或T-HuBERT,它对语音进行自监督学习,以获得类似音素的离散表示。这些音素类伪标签序列首先通过生成对抗网络(GAN)从语音中导出,以在统计上与来自额外的未配对文本数据的伪标签序列相似。通过这种方式,我们以无监督的方式在未配对的语音和文本之间建立了一座桥梁。大量的实验表明,我们提出的方法在各种强基线上具有显着的优越性,在LibriSpeech数据集上实现了高达15.3%的相对字错误率(WER)降低。摘要:Human language can be expressed in either written or spoken form, i.e. text or speech. Humans can acquire knowledge from text to improve speaking and listening. However, the quest for speech pre-trained models to leverage unpaired text has just started. In this paper, we investigate a new way to pre-train such a joint speech-text model to learn enhanced speech representations and benefit various speech-related downstream tasks. Specifically, we propose a novel pre-training method, text-guided HuBERT, or T-HuBERT, which performs self-supervised learning over speech to derive phoneme-like discrete representations. And these phoneme-like pseudo-label sequences are firstly derived from speech via the generative adversarial networks (GAN) to be statistically similar to those from additional unpaired textual data. In this way, we build a bridge between unpaired speech and text in an unsupervised manner. Extensive experiments demonstrate the significant superiority of our proposed method over various strong baselines, which achieves up to 15.3% relative Word Error Rate (WER) reduction on the LibriSpeech dataset.
【7】 Toward Fully Self-Supervised Multi-Pitch Estimation作者:Frank Cwitkowitz,Zhiyao Duan摘要:多音高估计是一个长达数十年的研究问题,涉及检测与多乐器混合物中的并发音乐事件相关联的音高活动。监督学习技术在更窄的任务特征上表现出了良好的性能,但受到缺乏具有多音高注释的大规模和多样化复调音乐数据集的限制。我们提出了一套用于多音高估计的自监督学习目标,其鼓励围绕谐波、音色变换的不变性和几何变换的等变性的集中支持。这些目标足以训练完全卷积的自动编码器,以直接产生多音调显著图,而无需任何微调。尽管只在合成单音符音频样本的集合上进行训练,但我们的完全自监督框架可以推广到复调音乐混合,并实现了与在传统多音高数据集上训练的监督模型相当的性能。摘要:Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supervised learning techniques have demonstrated solid performance on more narrow characterizations of the task, but suffer from limitations concerning the shortage of large-scale and diverse polyphonic music datasets with multi-pitch annotations. We present a suite of self-supervised learning objectives for multi-pitch estimation, which encourage the concentration of support around harmonics, invariance to timbral transformations, and equivariance to geometric transformations. These objectives are sufficient to train an entirely convolutional autoencoder to produce multi-pitch salience-grams directly, without any fine-tuning. Despite training exclusively on a collection of synthetic single-note audio samples, our fully self-supervised framework generalizes to polyphonic music mixtures, and achieves performance comparable to supervised models trained on conventional multi-pitch datasets.
【8】 Speech Corpus for Korean Children with Autism Spectrum Disorder: Towards Automatic Assessment Systems标题:韩国自闭症谱系障碍儿童语音语料库:自动评估系统作者:Seonwoo Lee,Jihyun Mun,Sunhee Kim,Minhwa Chung备注:11 pages, Accepted for LREC-COLING 2024摘要:尽管对自闭症谱系障碍(ASD)儿童的数字治疗需求不断增长,但目前还没有针对韩国ASD儿童的语音语料库。本文介绍了一个专门为韩国ASD儿童设计的语音语料库,旨在提高语音技术,如发音和严重程度评估。从演讲和语言评估会议的演讲录音转录,并注释发音和语言特征。三位言语和语言病理学家使用3分制李克特量表对这些录音进行社交严重性(SCS)和发音熟练度(PP)评分。参与者总数将为300名ASD儿童和50名典型发育(TD)儿童。本文还分析了从73名ASD儿童和9名TD儿童收集并完成注释的语音数据中提取的声学和语言特征,以调查ASD儿童的特征,并确定与临床评分相关的显著特征。结果揭示了ASD儿童的一些言语和语言特征,这些特征不同于TD儿童或按临床评分分类的另一个ASD亚组,表明了开发SCS和PP自动评估系统的潜力。摘要:Despite the growing demand for digital therapeutics for children with Autism Spectrum Disorder (ASD), there is currently no speech corpus available for Korean children with ASD. This paper introduces a speech corpus specifically designed for Korean children with ASD, aiming to advance speech technologies such as pronunciation and severity evaluation. Speech recordings from speech and language evaluation sessions were transcribed, and annotated for articulatory and linguistic characteristics. Three speech and language pathologists rated these recordings for social communication severity (SCS) and pronunciation proficiency (PP) using a 3-point Likert scale. The total number of participants will be 300 for children with ASD and 50 for typically developing (TD) children. The paper also analyzes acoustic and linguistic features extracted from speech data collected and completed for annotation from 73 children with ASD and 9 TD children to investigate the characteristics of children with ASD and identify significant features that correlate with the clinical scores. The results reveal some speech and linguistic characteristics in children with ASD that differ from those in TD children or another subgroup of ASD categorized by clinical scores, demonstrating the potential for developing automatic assessment systems for SCS and PP.
【9】 Towards Environmental Preference Based Speech Enhancement For Individualised Multi-Modal Hearing Aids作者:Jasper Kirton-Wingate,Shafique Ahmed,Adeel Hussain,Mandar Gogate,Kia Dashtipour,Jen-Cheng Hou,Tassadaq Hussain,Yu Tsao,Amir Hussain备注:This has been submitted to the Trends in Hearing journal摘要:自从深度学习(DL)出现以来,语音增强(SE)模型在各种噪声条件下都表现良好。然而,这样的系统仍然可能引入声音伪像、不自然的声音,并且限制用户听到可能是重要的环境声音的能力。助听器(HA)用户可能希望定制他们的SE系统,以适应他们的个人喜好和日常生活方式。在本文中,我们引入了一个基于偏好学习的SE(PLSE)模型,用于未来的多模态HA,可以根据用户的偏好,根据上下文利用音频信息来提高听觉舒适度。所提出的系统估计的信噪比(SNR)作为一个基本的客观语音质量的措施,量化的相对量的背景噪声存在于语音,并直接相关的信号的可懂度。此外,为了提供上下文信息,我们预测用户所处的声学场景。这些任务是通过多任务DL模型来实现的,该模型通过联合利用共享的编码特征空间来超越单独推断声学场景或SNR的性能。这些环境的推断中利用的偏好启发框架,线性学习一组预测功能,以确定AV(视听)SE系统的目标SNR。通过在具有挑战性的收听条件下大大降低噪声,并通过新颖地缩放SE模型的输出,我们能够为HA用户提供上下文个性化的SE。初步结果表明,在一些参与者中,非个体化基线模型有所改善。摘要:Since the advent of Deep Learning (DL), Speech Enhancement (SE) models have performed well under a variety of noise conditions. However, such systems may still introduce sonic artefacts, sound unnatural, and restrict the ability for a user to hear ambient sound which may be of importance. Hearing Aid (HA) users may wish to customise their SE systems to suit their personal preferences and day-to-day lifestyle. In this paper, we introduce a preference learning based SE (PLSE) model for future multi-modal HAs that can contextually exploit audio information to improve listening comfort, based upon the preferences of the user. The proposed system estimates the Signal-to-noise ratio (SNR) as a basic objective speech quality measure which quantifies the relative amount of background noise present in speech, and directly correlates to the intelligibility of the signal. Additionally, to provide contextual information we predict the acoustic scene in which the user is situated. These tasks are achieved via a multi-task DL model, which surpasses the performance of inferring the acoustic scene or SNR separately, by jointly leveraging a shared encoded feature space. These environmental inferences are exploited in a preference elicitation framework, which linearly learns a set of predictive functions to determine the target SNR of an AV (Audio-Visual) SE system. By greatly reducing noise in challenging listening conditions, and by novelly scaling the output of the SE model, we are able to provide HA users with contextually individualised SE. Preliminary results suggest an improvement over the non-individualised baseline model in some participants.
【10】 Open Your Ears to Take a Look: A State-of-the-Art Report on the Integration of Sonification and Visualization标题:打开耳朵看一看:可听化与可视化融合的最新报告作者:Kajetan Enge,Elias Elmquist,Valentina Caiola,Niklas Rönnberg,Alexander Rind,Michael Iber,Sara Lenzi,Fangfei Lan,Robert Höldrich,Wolfgang Aigner备注:27 pages, 10 figures, submitted to EuroVis 2024 conference摘要:研究数据显示和分析的可视化和声音化的研究社区有着非常相似的目标,基本上使任何类型的数据都可以被人类解释。一个社区通过使用数据的视觉表示来这样做,另一个社区通过使用数据的听觉(非语音)表示来这样做。虽然这两个社区有很多共同点,但在过去几十年中,它们大多是平行发展的。通过这个STAR,我们讨论了一系列跨越两个社区边界的作品,因此,一系列旨在将这两种技术整合到一种视听展示形式中的作品,我们认为这是“超过两者的总和”。“我们引入并激励适用于此类视听显示的分类系统,并将2011年至2023年期间出现的57种学术出版物的语料库按阅读水平,数据集类型或评估系统等类别进行分类。该语料库还可以对该领域进行元分析,包括定期出现的设计模式,如可视化和声音处理技术的类型,或视觉和听觉通道的使用,以及对该领域的合著者网络的分析,该网络显示了没有太多相互联系的各个团队。本STAR涵盖的工作主体还涉及三个相邻的主题:视听监控,可访问性和视听数据艺术。除了本研究的系统进行部分外,还单独讨论了这三个主题。本报告的研究结果可供这两个领域的研究人员使用,以了解这种集成设计的潜力和挑战,同时激励他们未来与其他领域的专家合作。摘要:The research communities studying visualization and sonification for data display and analysis share exceptionally similar goals, essentially making data of any kind interpretable to humans. One community does so by using visual representations of data, the other community does so by employing auditory (non-speech) representations of data. While the two communities have a lot in common, they developed mostly in parallel over the course of the last few decades. With this STAR, we discuss a collection of work that bridges the borders of the two communities, hence a collection of work that aims to integrate the two techniques to one form of audiovisual display, which we argue to be "more than the sum of the two." We introduce and motivate a classification system applicable to such audiovisual displays and categorize a corpus of 57 academic publications that appeared between 2011 and 2023 in categories such as reading level, dataset type, or evaluation system, to mention a few. The corpus also enables a meta-analysis of the field, including regularly occurring design patterns such as type of visualization and sonification techniques, or the use of visual and auditory channels, and the analysis of a co-author network of the field which shows individual teams without much interconnection. The body of work covered in this STAR also relates to three adjacent topics: audiovisual monitoring, accessibility, and audiovisual data art. These three topics are discussed individually in addition to the systematically conducted part of this research. The findings of this report may be used by researchers from both fields to understand the potentials and challenges of such integrated designs, while inspiring them for future collaboration with experts from the respective other field.
【11】 Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech作者:Szu-Wei Fu,Kuo-Hsuan Hung,Yu Tsao,Yu-Chiang Frank Wang备注:Published as a conference paper at ICLR 2024摘要:语音质量估计最近经历了从人类听觉专家设计到机器学习模型的范式转变。然而,目前的模型主要依赖于监督学习,这是耗时和昂贵的标签收集。为了解决这个问题,我们提出了VQScore,一种基于矢量量化变分自编码器(VQ-VAE)的量化误差的自监督语音评估度量。VQ-VAE的训练依赖于干净的语音;因此,当语音失真时,可以预期大的量化误差。为了进一步提高与真实质量分数的相关性,将语音处理的领域知识并入模型设计中。我们发现,矢量量化机制也可以用于自监督语音增强(SE)模型训练。为了提高编码器对SE的鲁棒性,引入了一种新的结合对抗训练的自蒸馏机制。总之,所提出的语音质量估计方法和增强模型只需要干净的语音进行训练,而没有任何标签要求。实验结果表明,所提出的VQScore和增强模型与监督基线相比具有竞争力。代码将在发布后发布。摘要:Speech quality estimation has recently undergone a paradigm shift from human-hearing expert designs to machine-learning models. However, current models rely mainly on supervised learning, which is time-consuming and expensive for label collection. To solve this problem, we propose VQScore, a self-supervised metric for evaluating speech based on the quantization error of a vector-quantized-variational autoencoder (VQ-VAE). The training of VQ-VAE relies on clean speech; hence, large quantization errors can be expected when the speech is distorted. To further improve correlation with real quality scores, domain knowledge of speech processing is incorporated into the model design. We found that the vector quantization mechanism could also be used for self-supervised speech enhancement (SE) model training. To improve the robustness of the encoder for SE, a novel self-distillation mechanism combined with adversarial training is introduced. In summary, the proposed speech quality estimation method and enhancement models require only clean speech for training without any label requirements. Experimental results show that the proposed VQScore and enhancement model are competitive with supervised baselines. The code will be released after publication.
【12】 ChatMusician: Understanding and Generating Music Intrinsically with LLM标题:ChatMusic:用LLM内在地理解和生成音乐作者:Ruibin Yuan,Hanfeng Lin,Yi Wang,Zeyue Tian,Shangda Wu,Tianhao Shen,Ge Zhang,Yuhang Wu,Cong Liu,Ziya Zhou,Ziyang Ma,Liumeng Xue,Ziyu Wang,Qin Liu,Tianyu Zheng,Yizhi Li,Yinghao Ma,Yiming Liang,Xiaowei Chi,Ruibo Liu,Zili Wang,Pengfei Li,Jingcheng Wu,Chenghua Lin,Qifeng Liu,Tao Jiang,Wenhao Huang,Wenhu Chen,Emmanouil Benetos,Jie Fu,Gus Xia,Roger Dannenberg,Wei Xue,Shiyin Kang,Yike Guo备注:GitHub: this https URL摘要:虽然大型语言模型(LLM)在文本生成方面表现出令人印象深刻的能力,但我们发现它们的能力尚未推广到音乐,人类的创造性语言。我们介绍ChatMusician,一个开源的LLM,集成了内在的音乐能力。它基于对文本兼容的音乐表示(ABC记谱法)的持续预训练和微调LLaMA 2,音乐被视为第二语言。ChatMusician可以使用纯文本标记器来理解和生成音乐,而无需任何外部多模态神经结构或标记器。有趣的是,赋予音乐能力并不会损害语言能力,甚至可以获得略高的MMLU分数。我们的模型能够创作结构良好的全长音乐,以文本,和弦,旋律,主题,音乐形式等为条件,超过GPT-4基线。在我们精心策划的大学水平的音乐理解基准,MusicTheoryBench,ChatMusician超过LLaMA 2和GPT-3.5的zero-shot设置的显着保证金。我们的工作表明,LLM可以是一个很好的音乐压缩机,但仍然有显着的领域有待征服。我们在GitHub中发布了我们的4 B代币音乐语言语料库MusicPile,收集的MusicTheoryBench,代码,模型和演示。摘要:While Large Language Models (LLMs) demonstrate impressive capabilities in text generation, we find that their ability has yet to be generalized to music, humanity's creative language. We introduce ChatMusician, an open-source LLM that integrates intrinsic musical abilities. It is based on continual pre-training and finetuning LLaMA2 on a text-compatible music representation, ABC notation, and the music is treated as a second language. ChatMusician can understand and generate music with a pure text tokenizer without any external multi-modal neural structures or tokenizers. Interestingly, endowing musical abilities does not harm language abilities, even achieving a slightly higher MMLU score. Our model is capable of composing well-structured, full-length music, conditioned on texts, chords, melodies, motifs, musical forms, etc, surpassing GPT-4 baseline. On our meticulously curated college-level music understanding benchmark, MusicTheoryBench, ChatMusician surpasses LLaMA2 and GPT-3.5 on zero-shot setting by a noticeable margin. Our work reveals that LLMs can be an excellent compressor for music, but there remains significant territory to be conquered. We release our 4B token music-language corpora MusicPile, the collected MusicTheoryBench, code, model and demo in GitHub.【13】 TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different Languages标题:TMT:通过将不同的模态处理为不同的语言,实现语音、图像和文本之间的三模态翻译作者:Minsu Kim,Jee-weon Jung,Hyeongseop Rha,Soumi Maiti,Siddhant Arora,Xuankai Chang,Shinji Watanabe,Yong Man Ro摘要:联合处理多模态信息的能力正在成为一项重要任务。然而,有限的成对多模态数据的数量和多模态学习的大计算需求阻碍了发展。我们提出了一种新的三模态翻译(TMT)模型,翻译之间的任意模态跨越语音,图像和文本。我们引入了一个新的观点,我们解释不同的语言不同的模态,并把多模态翻译作为一个完善的机器翻译问题。为此,我们将语音和图像数据标记为离散标记,这提供了跨模态的统一接口,并显着降低了计算成本。在所提出的TMT中,多模态编码器-解码器进行核心翻译,而特定于模态的处理仅在标记化和去标记化阶段内进行。我们评估建议的TMT对所有六个模态翻译任务。TMT的表现一直优于单一模型的同行,这表明统一任务不仅有利于实用性,而且有利于性能。摘要:The capability to jointly process multi-modal information is becoming an essential task. However, the limited number of paired multi-modal data and the large computational requirements in multi-modal learning hinder the development. We propose a novel Tri-Modal Translation (TMT) model that translates between arbitrary modalities spanning speech, image, and text. We introduce a novel viewpoint, where we interpret different modalities as different languages, and treat multi-modal translation as a well-established machine translation problem. To this end, we tokenize speech and image data into discrete tokens, which provide a unified interface across modalities and significantly decrease the computational cost. In the proposed TMT, a multi-modal encoder-decoder conducts the core translation, whereas modality-specific processing is conducted only within the tokenization and detokenization stages. We evaluate the proposed TMT on all six modality translation tasks. TMT outperforms single model counterparts consistently, demonstrating that unifying tasks is beneficial not only for practicality but also for performance.【14】 Phonetic and Lexical Discovery of a Canine Language using HuBERT作者:Xingyuan Li,Sinong Wang,Zeyu Xie,Mengyue Wu,Kenny Q. Zhu摘要:本文深入探讨了狗发声中潜在的通信模式的开拓性探索,并超越了传统的语言分析障碍,这种障碍严重依赖于人类对有限数据集的先验知识来寻找狗发声中的声音单元。我们提出了一个自我监督的方法与休伯特,使准确的音素标签分类和识别的声音模式,建议狗发声的基本词汇。我们的研究结果表明,在这些确定的犬词汇,涵盖了整个观察到的狗发声序列的声学一致性。我们进一步开发了一个基于网络的狗叫声标注系统。该系统可以突出显示用户上传的狗音频中存在于词汇表中的音素n-grams。摘要:This paper delves into the pioneering exploration of potential communication patterns within dog vocalizations and transcends traditional linguistic analysis barriers, which heavily relies on human priori knowledge on limited datasets to find sound units in dog vocalization. We present a self-supervised approach with HuBERT, enabling the accurate classification of phoneme labels and the identification of vocal patterns that suggest a rudimentary vocabulary within dog vocalizations. Our findings indicate a significant acoustic consistency in these identified canine vocabulary, covering the entirety of observed dog vocalization sequences. We further develop a web-based dog vocalization labeling system. This system can highlight phoneme n-grams, present in the vocabulary, in the dog audio uploaded by users.【15】 Direct Punjabi to English speech translation using discrete units作者:Prabhjot Kaur,L. Andrew M. Bush,Weisong Shi摘要:语音到语音翻译尚未达到与文本到文本翻译系统相同的覆盖水平。目前的语音技术在覆盖全世界7000多种语言方面非常有限,使一半以上的人口被剥夺了这种技术和共享经验。随着语音辅助技术(如社交机器人和语音转文本应用程序)和听觉内容(如播客和讲座)的兴起,确保所有人都能使用这项技术比以往任何时候都更加重要。语音翻译可以在缩小技术差距和创造更具包容性的社会方面发挥至关重要的作用。为了促进低资源语言的语音翻译研究,我们的工作提出了一个直接的语音到语音翻译模型的印度语言之一,称为旁遮普语到英语。此外,我们探讨了使用称为离散声学单元的离散语音表示作为基于transformer的翻译模型的输入的性能。该模型,简称为单元到单元翻译(U2UT),采用源语言(被翻译的语言)的离散单元序列,并输出目标语言(被翻译的语言)的离散单元序列。我们的研究结果表明,U2UT模型的性能优于语音到单元翻译(S2UT)模型,BLEU得分为3.69。摘要:Speech-to-speech translation is yet to reach the same level of coverage as text-to-text translation systems. The current speech technology is highly limited in its coverage of over 7000 languages spoken worldwide, leaving more than half of the population deprived of such technology and shared experiences. With voice-assisted technology (such as social robots and speech-to-text apps) and auditory content (such as podcasts and lectures) on the rise, ensuring that the technology is available for all is more important than ever. Speech translation can play a vital role in mitigating technological disparity and creating a more inclusive society. With a motive to contribute towards speech translation research for low-resource languages, our work presents a direct speech-to-speech translation model for one of the Indic languages called Punjabi to English. Additionally, we explore the performance of using a discrete representation of speech called discrete acoustic units as input to the Transformer-based translation model. The model, abbreviated as Unit-to-Unit Translation (U2UT), takes a sequence of discrete units of the source language (the language being translated from) and outputs a sequence of discrete units of the target language (the language being translated to). Our results show that the U2UT model performs better than the Speech-to-Unit Translation (S2UT) model by a 3.69 BLEU score.
【16】 Alternating Weak Triphone/BPE Alignment Supervision from Hybrid Model Improves End-to-End ASR标题:基于混合模型的交替弱三音素/BPE对齐监控改善端到端ASR作者:Jintao Jiang,Yingbo Gao,Mohammad Zeineldeen,Zoltan Tuske备注:5 pages, 1 figure, 3 tables摘要:本文提出了交替弱三音子/BPE对齐监督来改进端到端模型训练。为此,三音子和BPE对齐提取使用预先存在的混合ASR系统。然后,通过在用于三音子对准的编码器的中间层表示处和用于BPE对准的编码器处对这样的对准计算基于交叉熵的中间辅助损失来获得正则化效果。弱监督通过参数为0.5的强标签平滑来实现。TED-LIUM 2上的实验结果表明,无论是基于三音子或BPE对齐的弱监督提高ASR性能超过标准CTC辅助损失。此外,它们的组合进一步降低了字错误率。我们还研究了模型训练过程中两个辅助任务的交替,并观察到额外的性能增益。总的来说,所提出的技术导致超过10%的相对错误率降低超过CTC正则化基线系统。摘要:In this paper, alternating weak triphone/BPE alignment supervision is proposed to improve end-to-end model training. Towards this end, triphone and BPE alignments are extracted using a pre-existing hybrid ASR system. Then, regularization effect is obtained by cross-entropy based intermediate auxiliary losses computed on such alignments at a mid-layer representation of the encoder for triphone alignments and at the encoder for BPE alignments. Weak supervision is achieved through strong label smoothing with parameter of 0.5. Experimental results on TED-LIUM 2 indicate that either triphone or BPE alignment based weak supervision improves ASR performance over standard CTC auxiliary loss. Moreover, their combination lowers the word error rate further. We also investigate the alternation of the two auxiliary tasks during model training, and additional performance gain is observed. Overall, the proposed techniques result in over 10% relative error rate reduction over a CTC-regularized baseline system.【17】 GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model标题:GLA-Grad:一种扩展的Griffin-Lim波形产生扩散模型作者:Haocheng Liu,Teysir Baoueb,Mathieu Fontaine,Jonathan Le Roux,Gael Richard摘要:扩散模型在语音或音乐合成等各种信号生成任务中受到越来越多的关注。例如,WaveGrad是一个成功的扩散模型,它有条件地使用梅尔频谱图来指导用于生成高保真音频的扩散过程。然而,这些模型面临着关于训练和推理的噪声扩散过程的重要挑战,并且它们难以为训练期间未看到的说话者生成高质量的语音。为了最大限度地减少条件误差和提高噪声扩散过程的效率,本文提出了一种新的方案,称为GLA-Grad,它包括在定期扩散过程的每一步引入一个相位恢复算法,如Griffin-Lim算法(GLA)。此外,它可以直接应用于已经训练过的波形生成模型,而无需额外的训练或微调。我们表明,我们的算法优于国家的最先进的扩散模型的语音生成,特别是当生成语音为以前看不见的目标扬声器。摘要:Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectrogram to guide a diffusion process for the generation of high-fidelity audio. However, such models face important challenges concerning the noise diffusion process for training and inference, and they have difficulty generating high-quality speech for speakers that were not seen during training. With the aim of minimizing the conditioning error and increasing the efficiency of the noise diffusion process, we propose in this paper a new scheme called GLA-Grad, which consists in introducing a phase recovery algorithm such as the Griffin-Lim algorithm (GLA) at each step of the regular diffusion process. Furthermore, it can be directly applied to an already-trained waveform generation model, without additional training or fine-tuning. We show that our algorithm outperforms state-of-the-art diffusion models for speech generation, especially when generating speech for a previously unseen target speaker.