今日论文合集:cs.SD语音13篇,eess.AS音频处理16篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos
标题: VMAS:通过网络音乐视频中的语义对齐生成视频到音乐
作者:Yan-Bo Lin,Yu Tian,Linjie Yang,Gedas Bertasius,Heng Wang
备注:Project Page: this https URL
链接:点击下载PDF文件
摘要:我们提出了一个框架,学习从视频输入生成背景音乐。与现有的作品,依赖于象征性的音乐注释,这是有限的数量和多样性,我们的方法利用大规模的网络视频伴随着背景音乐。这使我们的模型能够学习生成逼真和多样化的音乐。为了实现这一目标,我们开发了一个生成的视频音乐Transformer与一个新的语义视频音乐对齐计划。我们的模型使用联合自回归和对比学习目标,鼓励生成与高级视频内容一致的音乐。我们还引入了一种新的视频节拍对齐方案,以匹配生成的音乐节拍与视频中的低级别运动。最后,为了捕获逼真背景音乐生成所需的视频中的细粒度视觉线索,我们引入了一种新的时间视频编码器架构,使我们能够有效地处理由许多密集采样帧组成的视频。我们在新策划的DISCO-MV数据集上训练我们的框架,该数据集由220万个视频音乐样本组成,比之前用于视频音乐生成的任何数据集都大几个数量级。根据各种音乐生成评估指标,包括人类评估,我们的方法在DISCO-MV和MusicCaps数据集上的性能优于现有方法。结果见https: genjib.github.io project_page VMAs index.html摘要:We present a framework for learning to generate background music from video inputs. Unlike existing works that rely on symbolic musical annotations, which are limited in quantity and diversity, our method leverages large-scale web videos accompanied by background music. This enables our model to learn to generate realistic and diverse music. To accomplish this goal, we develop a generative video-music Transformer with a novel semantic video-music alignment scheme. Our model uses a joint autoregressive and contrastive learning objective, which encourages the generation of music aligned with high-level video content. We also introduce a novel video-beat alignment scheme to match the generated music beats with the low-level motions in the video. Lastly, to capture fine-grained visual cues in a video needed for realistic background music generation, we introduce a new temporal video encoder architecture, allowing us to efficiently process videos consisting of many densely sampled frames. We train our framework on our newly curated DISCO-MV dataset, consisting of 2.2M video-music samples, which is orders of magnitude larger than any prior datasets used for video music generation. Our method outperforms existing approaches on the DISCO-MV and MusicCaps datasets according to various music generation evaluation metrics, including human evaluation. Results are available at https: genjib.github.io project_page VMAs index.html

【2】 A Suite for Acoustic Language Model Evaluation
标题: 声学语言模型评估套件
作者:Gallil Maimon,Amit Roth,Yossi Adi
链接:点击下载PDF文件
摘要:近年来,语音语言模型作为通用的语音处理系统显示出巨大的潜力。这样的模型有能力建模丰富的声学信息存在于音频信号中,除了口语内容,如情感,背景噪声等,尽管如此,评估基准,评估意识到广泛的声学方面,是缺乏的。为了帮助弥合这一差距,我们介绍了SALMon,一种新颖的评估套件,包括背景噪声,情感,扬声器身份和房间脉冲响应。所提出的基准测试既评估了被检查元素的一致性,又评估了它与口语文本的匹配程度。我们遵循基于建模的方法,测量模型是否给正确的样本比不正确的更高的分数。这种方法使基准测试快速计算,即使是大型模型。我们在SALMon上评估了几种语音语言模型,从而突出了每种评估方法的优点和缺点。代码和数据可在https: pages.cs.huji.ac.il adiyoss-lab salmon 上公开获取。摘要:Speech language models have recently demonstrated great potential as universal speech processing systems. Such models have the ability to model the rich acoustic information existing in audio signals, beyond spoken content, such as emotion, background noise, etc. Despite this, evaluation benchmarks which evaluate awareness to a wide range of acoustic aspects, are lacking. To help bridge this gap, we introduce SALMon, a novel evaluation suite encompassing background noise, emotion, speaker identity and room impulse response. The proposed benchmarks both evaluate the consistency of the inspected element and how much it matches the spoken text. We follow a modelling based approach, measuring whether a model gives correct samples higher scores than incorrect ones. This approach makes the benchmark fast to compute even for large models. We evaluated several speech language models on SALMon, thus highlighting the strengths and weaknesses of each evaluated method. Code and data are publicly available at https: pages.cs.huji.ac.il adiyoss-lab salmon .

【3】 D-CAPTCHA++: A Study of Resilience of Deepfake CAPTCHA under Transferable Imperceptible Adversarial Attack
标题: D-CAPTCHA++:Deepfake CAPTCHA在可转移不可感知对抗攻击下的弹性研究
作者:Hong-Hanh Nguyen-Le,Van-Tuan Tran,Dinh-Thuc Nguyen,Nhien-An Le-Khac
备注:14 pages
链接:点击下载PDF文件
摘要:生成式人工智能的进步使得音频合成模型得以改进,包括文本到语音和语音转换。这引起了人们对它在社会操纵和政治干预中可能被滥用的担忧,因为合成语音已经无法与自然人类语音区分开来。一些语音生成程序被用于恶意目的,特别是通过电话冒充个人。因此,检测虚假音频对于维护社会安全和保护信息完整性至关重要。最近的研究提出了一种基于挑战-响应协议的D-CAPTCHA系统,以区分虚假电话和真实电话。在这项工作中,我们研究了这个系统的弹性,并引入了一个更强大的版本,D-CAPTCHA++,以抵御虚假呼叫。具体来说,我们首先暴露了D-CAPTCHA系统在可转移的不可感知对抗攻击下的脆弱性。其次,我们通过在D-CAPTCHA deepfake检测器和任务分类器中使用对抗训练来提高系统的鲁棒性,从而减轻这种脆弱性。摘要:The advancements in generative AI have enabled the improvement of audio synthesis models, including text-to-speech and voice conversion. This raises concerns about its potential misuse in social manipulation and political interference, as synthetic speech has become indistinguishable from natural human speech. Several speech-generation programs are utilized for malicious purposes, especially impersonating individuals through phone calls. Therefore, detecting fake audio is crucial to maintain social security and safeguard the integrity of information. Recent research has proposed a D-CAPTCHA system based on the challenge-response protocol to differentiate fake phone calls from real ones. In this work, we study the resilience of this system and introduce a more robust version, D-CAPTCHA++, to defend against fake calls. Specifically, we first expose the vulnerability of the D-CAPTCHA system under transferable imperceptible adversarial attack. Secondly, we mitigate such vulnerability by improving the robustness of the system by using adversarial training in D-CAPTCHA deepfake detectors and task classifiers.

【4】 Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
标题: 音调口音语言中的跨方言文本转语音涉及多方言音素级BERT
作者:Kazuki Yamauchi,Yuki Saito,Hiroshi Saruwatari
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
摘要:我们探索跨方言的文本到语音(CD-TTS),一个任务,合成学习扬声器的声音在非母语方言,特别是在音高口音的语言。CD-TTS对于开发能够自然地与跨地区的人进行通信的语音代理非常重要。我们提出了一种新的TTS模型,包括三个子模块,以执行这项任务的竞争力。我们首先训练一个主干TTS模型,以根据参考编码器从语音中提取的音素级口音潜在变量(ALV)从文本中合成方言语音。然后,我们训练一个ALV预测器,利用我们新颖的多方言音素级BERT,从输入文本中预测针对目标方言定制的ALV。我们进行多方言TTS实验,并通过比较它与来自传统方言TTS方法的基线,我们的模型的有效性进行评估。实验结果表明,该模型提高了CD-TTS合成语音的方言自然度。摘要:We explore cross-dialect text-to-speech (CD-TTS), a task to synthesize learned speakers' voices in non-native dialects, especially in pitch-accent languages. CD-TTS is important for developing voice agents that naturally communicate with people across regions. We present a novel TTS model comprising three sub-modules to perform competitively at this task. We first train a backbone TTS model to synthesize dialect speech from a text conditioned on phoneme-level accent latent variables (ALVs) extracted from speech by a reference encoder. Then, we train an ALV predictor to predict ALVs tailored to a target dialect from input text leveraging our novel multi-dialect phoneme-level BERT. We conduct multi-dialect TTS experiments and evaluate the effectiveness of our model by comparing it with a baseline derived from conventional dialect TTS methods. The results show that our model improves the dialectal naturalness of synthetic speech in CD-TTS.

【5】 ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
标题: ManaRTS波斯语:为低级资源语言创建RTS数据集的秘诀
作者:Mahta Fetrat Qharabagh,Zahra Dehghanian,Hamid R. Rabiee
备注:33 pages, 12 figures
链接:点击下载PDF文件
摘要:在这项研究中,我们介绍ManaTTS,最广泛的公开访问的单扬声器波斯语语料库,和一个全面的框架,收集波斯语转录语音数据集。ManaTTS在开放CC-0许可下发布,包括大约86小时的音频,采样率为44.1 kHz。除了ManaTTS,我们还生成了VirgoolInformal数据集,以评估用于强制对齐的波斯语语音识别模型,扩展了5个小时的音频。这些数据集由完全透明的、麻省理工学院许可的管道支持,这证明了该领域的创新。它包括用于句子标记化的独特工具,有界音频分割和一种新颖的强制对齐方法。这种对齐技术是专门为低资源语言设计的,解决了该领域的关键需求。利用这个数据集,我们训练了一个基于Tacotron 2的TTS模型,实现了3.76的平均意见得分(MOS),这非常接近由相同声码器和自然频谱图生成的话语的MOS 3.86,以及自然波形的MOS 4.01,证明了语料库的卓越质量和有效性。摘要:In this study, we introduce ManaTTS, the most extensive publicly accessible single-speaker Persian corpus, and a comprehensive framework for collecting transcribed speech datasets for the Persian language. ManaTTS, released under the open CC-0 license, comprises approximately 86 hours of audio with a sampling rate of 44.1 kHz. Alongside ManaTTS, we also generated the VirgoolInformal dataset to evaluate Persian speech recognition models used for forced alignment, extending over 5 hours of audio. The datasets are supported by a fully transparent, MIT-licensed pipeline, a testament to innovation in the field. It includes unique tools for sentence tokenization, bounded audio segmentation, and a novel forced alignment method. This alignment technique is specifically designed for low-resource languages, addressing a crucial need in the field. With this dataset, we trained a Tacotron2-based TTS model, achieving a Mean Opinion Score (MOS) of 3.76, which is remarkably close to the MOS of 3.86 for the utterances generated by the same vocoder and natural spectrogram, and the MOS of 4.01 for the natural waveform, demonstrating the exceptional quality and effectiveness of the corpus.

【6】 Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
标题: Muskits-ESPnet:新范式中歌唱声音合成的综合工具包
作者:Yuning Wu,Jiatong Shi,Yifeng Yu,Yuxun Tang,Tao Qian,Yueqian Lin,Jionghao Han,Xinyi Bai,Shinji Watanabe,Qin Jin
备注:Accepted by ACMMM 2024 demo track
链接:点击下载PDF文件
摘要:这项研究提出了Muskits-ESPnet,这是一个多功能的工具包,通过在连续和离散方法中应用预训练的音频模型,为歌唱语音合成(SVS)引入了新的范例。具体而言,我们探索了从SSL模型和音频编解码器中衍生的离散表示,并在多功能性和智能性方面提供了显着优势,支持多种格式的输入和各种SVS模型的自适应数据处理工作流。该工具包具有自动乐谱错误检测和校正,以及一个感知自动评估模块,以模仿人类的主观评估分数。Muskits-ESPnet可以在 url{https: github.com espnet espnet}上找到。摘要:This research presents Muskits-ESPnet, a versatile toolkit that introduces new paradigms to Singing Voice Synthesis (SVS) through the application of pretrained audio models in both continuous and discrete approaches. Specifically, we explore discrete representations derived from SSL models and audio codecs and offer significant advantages in versatility and intelligence, supporting multi-format inputs and adaptable data processing workflows for various SVS models. The toolkit features automatic music score error detection and correction, as well as a perception auto-evaluation module to imitate human subjective evaluating scores. Muskits-ESPnet is available at url{https: github.com espnet espnet}.

【7】 Analytic Class Incremental Learning for Sound Source Localization with Privacy Protection
标题: 分析类增量学习,实现具有隐私保护的声音源本地化
作者:Xinyuan Qian,Xianghu Yue,Jiadong Wang,Huiping Zhuang,Haizhou Li
链接:点击下载PDF文件
摘要:声源定位(SSL)技术可用于监控和机器人等应用。虽然传统的基于信号处理(SP)的SSL方法在特定的信号和噪声假设下提供了分析解决方案,但最近基于深度学习(DL)的方法已经大大优于它们。然而,它们的成功取决于大量的训练数据和大量的计算资源。此外,它们通常依赖于大规模的注释空间数据,并且在适应不断变化的声音类别时可能会遇到困难。为了缓解这些挑战,我们提出了一种新的类增量学习(CIL)方法,称为SSL-CIL,它避免了严重的准确性下降,由于灾难性的遗忘通过增量更新基于DL的SSL模型通过一个封闭的形式的解析解。特别是,由于学习过程不会重新访问任何历史数据(无样本),因此确保了数据隐私,这更适合智能家居场景。在公开的SSLR数据集上的实证结果表明,我们的建议具有优越的性能,实现了90.9%的定位准确率,超过了其他竞争对手的方法。摘要:Sound Source Localization (SSL) enabling technology for applications such as surveillance and robotics. While traditional Signal Processing (SP)-based SSL methods provide analytic solutions under specific signal and noise assumptions, recent Deep Learning (DL)-based methods have significantly outperformed them. However, their success depends on extensive training data and substantial computational resources. Moreover, they often rely on large-scale annotated spatial data and may struggle when adapting to evolving sound classes. To mitigate these challenges, we propose a novel Class Incremental Learning (CIL) approach, termed SSL-CIL, which avoids serious accuracy degradation due to catastrophic forgetting by incrementally updating the DL-based SSL model through a closed-form analytic solution. In particular, data privacy is ensured since the learning process does not revisit any historical data (exemplar-free), which is more suitable for smart home scenarios. Empirical results in the public SSLR dataset demonstrate the superior performance of our proposal, achieving a localization accuracy of 90.9%, surpassing other competitive methods.

【8】 Enhancing CTC-Based Visual Speech Recognition
标题: 增强基于ATC的视觉语音识别
作者:Hendrik Laux,Anke Schmeink
链接:点击下载PDF文件
摘要:本文介绍了LiteVSR 2,我们以前介绍的有效的视觉语音识别(VSR)的方法的增强版本。基于我们的知识蒸馏框架,从一个预先训练的自动语音识别(ASR)模型,我们引入了两个关键的改进:一个稳定的视频预处理技术和特征归一化的蒸馏过程。这些改进在LRS2和LRS3基准测试中产生了实质性的性能提升,将LiteVSR 2定位为当前最佳的基于CTC的VSR模型,而不会增加训练数据量或使用的计算资源。此外,我们通过检查不同模型复杂性和训练数据量的性能指标来探索我们方法的可扩展性。LiteVSR 2保持了其前身的效率,同时显著提高了准确性,从而展示了VSR技术在资源效率方面的进步潜力。摘要:This paper presents LiteVSR2, an enhanced version of our previously introduced efficient approach to Visual Speech Recognition (VSR). Building upon our knowledge distillation framework from a pre-trained Automatic Speech Recognition (ASR) model, we introduce two key improvements: a stabilized video preprocessing technique and feature normalization in the distillation process. These improvements yield substantial performance gains on the LRS2 and LRS3 benchmarks, positioning LiteVSR2 as the current best CTC-based VSR model without increasing the volume of training data or computational resources utilized. Furthermore, we explore the scalability of our approach by examining performance metrics across varying model complexities and training data volumes. LiteVSR2 maintains the efficiency of its predecessor while significantly enhancing accuracy, thereby demonstrating the potential for resource-efficient advancements in VSR technology.

【9】 Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition
标题: 用于流语音识别的带有SummaryMixing的线性时间复杂度一致器
作者:Titouan Parcollet,Rogier van Dalen,Shucong Zhang,Sourav Batthacharya
链接:点击下载PDF文件
摘要:自动语音识别(ASR)的编码器配备有自我注意,无论是流或非流,需要的时间在语音话语的长度的平方。这减慢了训练和解码,增加了它们的成本,并且限制了ASR在受限设备中的部署。SummaryMixing是一种很有前途的线性时间复杂度替代自注意的非流语音识别方法,首次保留或优于自注意模型的准确性。不幸的是,SummaryMixing的原始定义不适合流式语音识别。因此,这项工作将SummaryMixing扩展到在流式传输和离线模式下工作的Conformer换能器。结果表明,这种新的线性时间复杂度的语音编码器优于自注意在这两种情况下,同时需要更少的计算和内存在训练和解码。摘要:Automatic speech recognition (ASR) with an encoder equipped with self-attention, whether streaming or non-streaming, takes quadratic time in the length of the speech utterance. This slows down training and decoding, increase their cost, and limit the deployment of the ASR in constrained devices. SummaryMixing is a promising linear-time complexity alternative to self-attention for non-streaming speech recognition that, for the first time, preserves or outperforms the accuracy of self-attention models. Unfortunately, the original definition of SummaryMixing is not suited to streaming speech recognition. Hence, this work extends SummaryMixing to a Conformer Transducer that works in both a streaming and an offline mode. It shows that this new linear-time complexity speech encoder outperforms self-attention in both scenarios while requiring less compute and memory during training and decoding.

【10】 Developing a Framework for Sonifying Variational Quantum Algorithms: Implications for Music Composition
标题: 开发发声变分量子算法的框架:对音乐创作的影响
作者:Paulo Vitor Itaboraí,Peter Thomas,Arianna Crippa,Karl Jansen,Tim Schwägerl,María Aguado Yáñez
备注:This is a non-edited pre-publication version of a chapter to appear in the book Advances in Quantum Computer Music, World Scientific, editor E. R. Miranda, 2024. ISBN:978-981-98-0017-9
链接:点击下载PDF文件
摘要:本章介绍了变分量子调和器,一个软件工具和音乐接口,重点是变分量子算法(VQA)的最小化步骤的声音化问题,用于模拟量子系统的属性和量子硬件辅助的优化问题。特别是,它详细的二次无约束二元优化(QUBO)问题的发音使用VQA。灵活的设计使其未来的应用既可以作为科学研究中听觉显示的发声工具,也可以作为艺术创作的混合量子数字乐器。反过来,声化可以帮助研究人员更好地理解复杂系统,并可以为量子物理和量子计算的培训服务。VQH结构,包括它的软件实现,控制机制,和发音映射进行了详细说明。此外,它指导VQH中QUBO代价函数作为音乐合成对象的设计。讨论扩展到应用量子计算机辅助组成和现场编码性能的量子辅助模拟的影响。艺术作品《Hundreonal Chambers》(Thomas and Itabora,2023年)展示了这一点。摘要:This chapter examines the Variational Quantum Harmonizer, a software tool and musical interface that focuses on the problem of sonification of the minimization steps of Variational Quantum Algorithms (VQA), used for simulating properties of quantum systems and optimization problems assisted by quantum hardware. Particularly, it details the sonification of Quadratic Unconstrained Binary Optimization (QUBO) problems using VQA. A flexible design enables its future applications both as a sonification tool for auditory displays in scientific investigation, and as a hybrid quantum-digital musical instrument for artistic endeavours. In turn, sonification can help researchers understand complex systems better and can serve for the training of quantum physics and quantum computing. The VQH structure, including its software implementation, control mechanisms, and sonification mappings are detailed. Moreover, it guides the design of QUBO cost functions in VQH as a music compositional object. The discussion is extended to the implications of applying quantum-assisted simulation in quantum-computer aided composition and live-coding performances. An artistic output is showcased by the piece textit{Hexagonal Chambers} (Thomas and Itabora 'i, 2023).

【11】 Improving Anomalous Sound Detection via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio Models
标题: 通过预训练音频模型的低等级自适应微调来改进异常声音检测
作者:Xinhu Zheng,Anbai Jiang,Bing Han,Yanmin Qian,Pingyi Fan,Jia Liu,Wei-Qiang Zhang
链接:点击下载PDF文件
摘要:异常声音检测(ASD)通过各种人工智能(AI)技术在工业环境中的应用获得了极大的兴趣。虽然ASD系统具有很大的潜力,但由于数据收集的困难和环境因素的复杂性而导致的泛化问题,ASD系统很难在实际生产现场部署。本文介绍了一种利用音频预训练模型的鲁棒ASD模型。具体来说,我们使用机器操作数据微调这些模型,采用SpecAug作为数据增强策略。此外,我们研究了利用低秩自适应(LoRA)调整而不是完全微调的影响,以解决微调数据有限的问题。我们在DCASE2023任务2数据集上的实验在评估集上建立了77.75%的新基准,与之前的最先进(SOTA)模型相比,显着提高了6.48%,包括顶级传统卷积网络和语音预训练模型,这证明了具有LoRA调整的音频预训练模型的有效性。还进行了消融研究,以展示所提出的方案的有效性。摘要:Anomalous Sound Detection (ASD) has gained significant interest through the application of various Artificial Intelligence (AI) technologies in industrial settings. Though possessing great potential, ASD systems can hardly be readily deployed in real production sites due to the generalization problem, which is primarily caused by the difficulty of data collection and the complexity of environmental factors. This paper introduces a robust ASD model that leverages audio pre-trained models. Specifically, we fine-tune these models using machine operation data, employing SpecAug as a data augmentation strategy. Additionally, we investigate the impact of utilizing Low-Rank Adaptation (LoRA) tuning instead of full fine-tuning to address the problem of limited data for fine-tuning. Our experiments on the DCASE2023 Task 2 dataset establish a new benchmark of 77.75% on the evaluation set, with a significant improvement of 6.48% compared with previous state-of-the-art (SOTA) models, including top-tier traditional convolutional networks and speech pre-trained models, which demonstrates the effectiveness of audio pre-trained models with LoRA tuning. Ablation studies are also conducted to showcase the efficacy of the proposed scheme.

【12】 The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
标题: VoiceMOS 2024挑战:超越语音质量预测
作者:Wen-Chin Huang,Szu-Wei Fu,Erica Cooper,Ryandhimas E. Zezario,Tomoki Toda,Hsin-Min Wang,Junichi Yamagishi,Yu Tsao
备注:Accepted to SLT2024
链接:点击下载PDF文件
摘要:我们介绍了VoiceMOS挑战赛的第三版,这是一项旨在推进人类语音评级自动预测研究的科学举措。有三条轨道。第一个轨道是预测“放大”的语音合成系统的高质量样本的质量。第二个轨道是预测评级的样本从歌唱声音合成和声音转换与各种各样的系统,听众和语言。第三个轨道是半监督质量预测,用于噪声,干净和增强的语音,其中提供了非常少量的标记训练数据。在来自学术界和工业界的八个团队中,我们发现许多团队的性能都超过了基准系统。成功的技术包括基于检索的方法和使用非自我监督的表示,如频谱图和音高直方图。这些结果表明,这一挑战推动了主观语音评级预测领域的发展。摘要:We present the third edition of the VoiceMOS Challenge, a scientific initiative designed to advance research into automatic prediction of human speech ratings. There were three tracks. The first track was on predicting the quality of zoomed-in'' high-quality samples from speech synthesis systems. The second track was to predict ratings of samples from singing voice synthesis and voice conversion with a large variety of systems, listeners, and languages. The third track was semi-supervised quality prediction for noisy, clean, and enhanced speech, where a very small amount of labeled training data was provided. Among the eight teams from both academia and industry, we found that many were able to outperform the baseline systems. Successful techniques included retrieval-based methods and the use of non-self-supervised representations like spectrograms and pitch histograms. These results showed that the challenge has advanced the field of subjective speech rating prediction.

【13】 Unveiling Visual Biases in Audio-Visual Localization Benchmarks
标题: 揭露视听本地化基准中的视觉偏见
作者:Liangyu Chen,Zihao Yue,Boshen Xu,Qin Jin
备注:Accepted by ECCV24 AVGenL Workshop
链接:点击下载PDF文件
摘要:视听源定位(AVSL)旨在定位视频中的声源。在本文中,我们确定了一个重要的问题,在现有的基准:发声对象往往很容易识别的基础上,仅仅是视觉线索,我们称之为视觉偏见。这种偏见阻碍了这些基准有效地评估AVSL模型。为了进一步验证我们关于视觉偏差的假设,我们研究了两个代表性的AVSL基准,VGG-SS和EpicSounding-Object,其中仅视觉模型优于所有视听基线。我们的研究结果表明,现有的AVSL基准需要进一步完善,以促进视听学习。摘要:Audio-Visual Source Localization (AVSL) aims to localize the source of sound within a video. In this paper, we identify a significant issue in existing benchmarks: the sounding objects are often easily recognized based solely on visual cues, which we refer to as visual bias. Such biases hinder these benchmarks from effectively evaluating AVSL models. To further validate our hypothesis regarding visual biases, we examine two representative AVSL benchmarks, VGG-SS and EpicSounding-Object, where the vision-only models outperform all audiovisual baselines. Our findings suggest that existing AVSL benchmarks need further refinement to facilitate audio-visual learning.


eess.AS音频处理
【1】 Rethinking Mamba in Speech Processing by Self-Supervised Models
标题: 通过自我监督模型重新思考语音处理中的曼巴语
作者:Xiangyu Zhang,Jianbo Ma,Mostafa Shahin,Beena Ahmed,Julien Epps
链接:点击下载PDF文件
摘要:基于Mamba的模型在计算机视觉、自然语言处理和语音处理方面表现出了出色的性能。然而,在语音处理领域,基于Mamba的模型的性能在不同的任务中会有所不同。例如,在语音增强和频谱重建等任务中,Mamba模型在独立使用时表现良好。然而,对于像语音识别这样的任务,需要额外的模块来超越基于注意力的模型的性能。我们提出的假设,曼巴为基础的模型擅长在语音处理的“重建”任务。然而,对于“分类任务”(如语音识别),需要额外的模块来完成“重建”步骤。为了验证我们的假设,我们从信息论的角度分析了以前的基于曼巴的语音模型。此外,我们在研究中利用了HuBERT的特性。我们训练了一个基于Mamba的HuBERT模型,互信息模式以及模型的性能指标证实了我们的假设。摘要:The Mamba-based model has demonstrated outstanding performance across tasks in computer vision, natural language processing, and speech processing. However, in the realm of speech processing, the Mamba-based model's performance varies across different tasks. For instance, in tasks such as speech enhancement and spectrum reconstruction, the Mamba model performs well when used independently. However, for tasks like speech recognition, additional modules are required to surpass the performance of attention-based models. We propose the hypothesis that the Mamba-based model excels in "reconstruction" tasks within speech processing. However, for "classification tasks" such as Speech Recognition, additional modules are necessary to accomplish the "reconstruction" step. To validate our hypothesis, we analyze the previous Mamba-based Speech Models from an information theory perspective. Furthermore, we leveraged the properties of HuBERT in our study. We trained a Mamba-based HuBERT model, and the mutual information patterns, along with the model's performance metrics, confirmed our assumptions.

【2】 Zero-Shot Text-to-Speech as Golden Speech Generator: A Systematic Framework and its Applicability in Automatic Pronunciation Assessment
标题: Zero-Shot文本到语音作为黄金语音生成器:一个系统框架及其在自动发音评估中的适用性
作者:Tien-Hong Lo,Meng-Ting Tsai,Berlin Chen
备注:11 pages, 4 figures, 4 tables
链接:点击下载PDF文件
摘要:第二语言学习者可以通过模仿黄金语音来提高自己的语音水平,尤其是模仿符合自己语音特点的语音。本研究探讨了zero-shot文语转换(TTS)技术产生的学习者特定的黄金语音可以作为衡量二语学习者发音水平的有效指标这一假设。基于这种探索,本研究的贡献至少有两个方面:1)设计和开发了一个系统的框架,用于评估合成模型生成黄金语音的能力,以及2)在自动发音评估(APA)中使用黄金语音的有效性的深入调查。在L2-ARCTIC和Speechocean 762基准数据集上进行的综合实验表明,我们提出的建模可以在与一些现有技术相关的各种评估指标方面产生显着的性能改进。据我们所知,这项研究是第一个探讨黄金语音的作用,在两个TTS和APA,提供了一个有前途的制度,计算机辅助发音训练(CAPT)。摘要:Second language (L2) learners can improve their pronunciation by imitating golden speech, especially when the speech that aligns with their respective speech characteristics. This study explores the hypothesis that learner-specific golden speech generated with zero-shot text-to-speech (ZS-TTS) techniques can be harnessed as an effective metric for measuring the pronunciation proficiency of L2 learners. Building on this exploration, the contributions of this study are at least two-fold: 1) design and development of a systematic framework for assessing the ability of a synthesis model to generate golden speech, and 2) in-depth investigations of the effectiveness of using golden speech in automatic pronunciation assessment (APA). Comprehensive experiments conducted on the L2-ARCTIC and Speechocean762 benchmark datasets suggest that our proposed modeling can yield significant performance improvements with respect to various assessment metrics in relation to some prior arts. To our knowledge, this study is the first to explore the role of golden speech in both ZS-TTS and APA, offering a promising regime for computer-assisted pronunciation training (CAPT).

【3】 Neural Ambisonic Encoding For Multi-Speaker Scenarios Using A Circular Microphone Array
标题: 使用圆形麦克风阵列的多扬声器场景的神经立体声编码
作者:Yue Qiao,Vinay Kothapally,Meng Yu,Dong Yu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:像高保真立体声这样的空间音频格式与回放设备布局无关,非常适合电话会议和虚拟现实等应用。传统的高保真度立体声编码方法通常依赖于球形麦克风阵列来进行有效的声场捕获,这限制了它们在实际场景中的灵活性。我们提出了一种基于深度学习(DL)的方法,利用两级网络架构将圆形麦克风阵列信号编码为多扬声器环境中的二阶高保真度立体声(SOA)。此外,我们还介绍:(i)基于空间功率图的新颖损失函数以规则化高保真度立体声信号的声道间相关性,以及(ii)声道置换技术以解决使用水平圆形阵列编码垂直信息的模糊性。模拟语音和噪声数据集的评估表明,我们的方法始终优于传统的信号处理(SP)和DL为基础的方法,提供显着更好的音色和空间质量和更高的源定位精度。具有可视化效果的双耳音频演示可在https: bridgoon97.github.io NeuralAmbisonicEncoding 上获得。摘要:Spatial audio formats like Ambisonics are playback device layout-agnostic and well-suited for applications such as teleconferencing and virtual reality. Conventional Ambisonic encoding methods often rely on spherical microphone arrays for efficient sound field capture, which limits their flexibility in practical scenarios. We propose a deep learning (DL)-based approach, leveraging a two-stage network architecture for encoding circular microphone array signals into second-order Ambisonics (SOA) in multi-speaker environments. In addition, we introduce: (i) a novel loss function based on spatial power maps to regularize inter-channel correlations of the Ambisonic signals, and (ii) a channel permutation technique to resolve the ambiguity of encoding vertical information using a horizontal circular array. Evaluation on simulated speech and noise datasets shows that our approach consistently outperforms traditional signal processing (SP) and DL-based methods, providing significantly better timbral and spatial quality and higher source localization accuracy. Binaural audio demos with visualizations are available at https: bridgoon97.github.io NeuralAmbisonicEncoding .

【4】 VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos
标题: VMAS:通过网络音乐视频中的语义对齐生成视频到音乐
作者:Yan-Bo Lin,Yu Tian,Linjie Yang,Gedas Bertasius,Heng Wang
备注:Project Page: this https URL
链接:点击下载PDF文件
摘要:我们提出了一个框架,学习从视频输入生成背景音乐。与现有的作品,依赖于象征性的音乐注释,这是有限的数量和多样性,我们的方法利用大规模的网络视频伴随着背景音乐。这使我们的模型能够学习生成逼真和多样化的音乐。为了实现这一目标,我们开发了一个生成的视频音乐Transformer与一个新的语义视频音乐对齐计划。我们的模型使用联合自回归和对比学习目标,鼓励生成与高级视频内容一致的音乐。我们还引入了一种新的视频节拍对齐方案,以匹配生成的音乐节拍与视频中的低级别运动。最后,为了捕获逼真背景音乐生成所需的视频中的细粒度视觉线索,我们引入了一种新的时间视频编码器架构,使我们能够有效地处理由许多密集采样帧组成的视频。我们在新策划的DISCO-MV数据集上训练我们的框架,该数据集由220万个视频音乐样本组成,比之前用于视频音乐生成的任何数据集都大几个数量级。根据各种音乐生成评估指标,包括人类评估,我们的方法在DISCO-MV和MusicCaps数据集上的性能优于现有方法。结果见https: genjib.github.io project_page VMAs index.html摘要:We present a framework for learning to generate background music from video inputs. Unlike existing works that rely on symbolic musical annotations, which are limited in quantity and diversity, our method leverages large-scale web videos accompanied by background music. This enables our model to learn to generate realistic and diverse music. To accomplish this goal, we develop a generative video-music Transformer with a novel semantic video-music alignment scheme. Our model uses a joint autoregressive and contrastive learning objective, which encourages the generation of music aligned with high-level video content. We also introduce a novel video-beat alignment scheme to match the generated music beats with the low-level motions in the video. Lastly, to capture fine-grained visual cues in a video needed for realistic background music generation, we introduce a new temporal video encoder architecture, allowing us to efficiently process videos consisting of many densely sampled frames. We train our framework on our newly curated DISCO-MV dataset, consisting of 2.2M video-music samples, which is orders of magnitude larger than any prior datasets used for video music generation. Our method outperforms existing approaches on the DISCO-MV and MusicCaps datasets according to various music generation evaluation metrics, including human evaluation. Results are available at https: genjib.github.io project_page VMAs index.html

【5】 A Suite for Acoustic Language Model Evaluation
标题: 声学语言模型评估套件
作者:Gallil Maimon,Amit Roth,Yossi Adi
链接:点击下载PDF文件
摘要:近年来,语音语言模型作为通用的语音处理系统显示出巨大的潜力。这样的模型有能力建模丰富的声学信息存在于音频信号中,除了口语内容,如情感,背景噪声等,尽管如此,评估基准,评估意识到广泛的声学方面,是缺乏的。为了帮助弥合这一差距,我们介绍了SALMon,一种新颖的评估套件,包括背景噪声,情感,扬声器身份和房间脉冲响应。所提出的基准测试既评估了被检查元素的一致性,又评估了它与口语文本的匹配程度。我们遵循基于建模的方法,测量模型是否给正确的样本比不正确的更高的分数。这种方法使基准测试快速计算,即使是大型模型。我们在SALMon上评估了几种语音语言模型,从而突出了每种评估方法的优点和缺点。代码和数据可在https: pages.cs.huji.ac.il adiyoss-lab salmon 上公开获取。摘要:Speech language models have recently demonstrated great potential as universal speech processing systems. Such models have the ability to model the rich acoustic information existing in audio signals, beyond spoken content, such as emotion, background noise, etc. Despite this, evaluation benchmarks which evaluate awareness to a wide range of acoustic aspects, are lacking. To help bridge this gap, we introduce SALMon, a novel evaluation suite encompassing background noise, emotion, speaker identity and room impulse response. The proposed benchmarks both evaluate the consistency of the inspected element and how much it matches the spoken text. We follow a modelling based approach, measuring whether a model gives correct samples higher scores than incorrect ones. This approach makes the benchmark fast to compute even for large models. We evaluated several speech language models on SALMon, thus highlighting the strengths and weaknesses of each evaluated method. Code and data are publicly available at https: pages.cs.huji.ac.il adiyoss-lab salmon .

【6】 D-CAPTCHA++: A Study of Resilience of Deepfake CAPTCHA under Transferable Imperceptible Adversarial Attack
标题: D-CAPTCHA++:Deepfake CAPTCHA在可转移不可感知对抗攻击下的弹性研究
作者:Hong-Hanh Nguyen-Le,Van-Tuan Tran,Dinh-Thuc Nguyen,Nhien-An Le-Khac
备注:14 pages
链接:点击下载PDF文件
摘要:生成式人工智能的进步使得音频合成模型得以改进,包括文本到语音和语音转换。这引起了人们对它在社会操纵和政治干预中可能被滥用的担忧,因为合成语音已经无法与自然人类语音区分开来。一些语音生成程序被用于恶意目的,特别是通过电话冒充个人。因此,检测虚假音频对于维护社会安全和保护信息完整性至关重要。最近的研究提出了一种基于挑战-响应协议的D-CAPTCHA系统,以区分虚假电话和真实电话。在这项工作中,我们研究了这个系统的弹性,并引入了一个更强大的版本,D-CAPTCHA++,以抵御虚假呼叫。具体来说,我们首先暴露了D-CAPTCHA系统在可转移的不可感知对抗攻击下的脆弱性。其次,我们通过在D-CAPTCHA deepfake检测器和任务分类器中使用对抗训练来提高系统的鲁棒性,从而减轻这种脆弱性。摘要:The advancements in generative AI have enabled the improvement of audio synthesis models, including text-to-speech and voice conversion. This raises concerns about its potential misuse in social manipulation and political interference, as synthetic speech has become indistinguishable from natural human speech. Several speech-generation programs are utilized for malicious purposes, especially impersonating individuals through phone calls. Therefore, detecting fake audio is crucial to maintain social security and safeguard the integrity of information. Recent research has proposed a D-CAPTCHA system based on the challenge-response protocol to differentiate fake phone calls from real ones. In this work, we study the resilience of this system and introduce a more robust version, D-CAPTCHA++, to defend against fake calls. Specifically, we first expose the vulnerability of the D-CAPTCHA system under transferable imperceptible adversarial attack. Secondly, we mitigate such vulnerability by improving the robustness of the system by using adversarial training in D-CAPTCHA deepfake detectors and task classifiers.

【7】 Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
标题: 音调口音语言中的跨方言文本转语音涉及多方言音素级BERT
作者:Kazuki Yamauchi,Yuki Saito,Hiroshi Saruwatari
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
摘要:我们探索跨方言的文本到语音(CD-TTS),一个任务,合成学习扬声器的声音在非母语方言,特别是在音高口音的语言。CD-TTS对于开发能够自然地与跨地区的人进行通信的语音代理非常重要。我们提出了一种新的TTS模型,包括三个子模块,以执行这项任务的竞争力。首先,我们训练一个骨干TTS模型来合成方言语音从一个文本的音素级口音的潜变量(ALV)从语音提取的参考编码器的条件。然后,我们训练一个ALV预测器,利用我们新颖的多方言音素级BERT,从输入文本中预测针对目标方言定制的ALV。我们进行多方言TTS实验,并通过比较它与来自传统方言TTS方法的基线,我们的模型的有效性进行评估。实验结果表明,该模型提高了CD-TTS合成语音的方言自然度。摘要:We explore cross-dialect text-to-speech (CD-TTS), a task to synthesize learned speakers' voices in non-native dialects, especially in pitch-accent languages. CD-TTS is important for developing voice agents that naturally communicate with people across regions. We present a novel TTS model comprising three sub-modules to perform competitively at this task. We first train a backbone TTS model to synthesize dialect speech from a text conditioned on phoneme-level accent latent variables (ALVs) extracted from speech by a reference encoder. Then, we train an ALV predictor to predict ALVs tailored to a target dialect from input text leveraging our novel multi-dialect phoneme-level BERT. We conduct multi-dialect TTS experiments and evaluate the effectiveness of our model by comparing it with a baseline derived from conventional dialect TTS methods. The results show that our model improves the dialectal naturalness of synthetic speech in CD-TTS.

【8】 ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
标题: ManaRTS波斯语:为低级资源语言创建RTS数据集的秘诀
作者:Mahta Fetrat Qharabagh,Zahra Dehghanian,Hamid R. Rabiee
备注:33 pages, 12 figures
链接:点击下载PDF文件
摘要:在这项研究中,我们介绍ManaTTS,最广泛的公开访问的单扬声器波斯语语料库,和一个全面的框架,收集波斯语转录语音数据集。ManaTTS在开放CC-0许可下发布,包括大约86小时的音频,采样率为44.1 kHz。除了ManaTTS,我们还生成了VirgoolInformal数据集,以评估用于强制对齐的波斯语语音识别模型,扩展了5个小时的音频。这些数据集由一个完全透明的、麻省理工学院许可的管道提供支持,这证明了该领域的创新。它包括独特的句子标记化工具、有界音频分割和新颖的强制对齐方法。这种对齐技术是专门为低资源语言设计的,解决了该领域的关键需求。使用该数据集,我们训练了一个基于Tacotron 2的TTS模型,实现了3.76的平均意见得分(MOS),这非常接近由相同声码器和自然频谱图生成的话语的MOS 3.86,以及自然波形的MOS 4.01,证明了语料库的卓越质量和有效性。摘要:In this study, we introduce ManaTTS, the most extensive publicly accessible single-speaker Persian corpus, and a comprehensive framework for collecting transcribed speech datasets for the Persian language. ManaTTS, released under the open CC-0 license, comprises approximately 86 hours of audio with a sampling rate of 44.1 kHz. Alongside ManaTTS, we also generated the VirgoolInformal dataset to evaluate Persian speech recognition models used for forced alignment, extending over 5 hours of audio. The datasets are supported by a fully transparent, MIT-licensed pipeline, a testament to innovation in the field. It includes unique tools for sentence tokenization, bounded audio segmentation, and a novel forced alignment method. This alignment technique is specifically designed for low-resource languages, addressing a crucial need in the field. With this dataset, we trained a Tacotron2-based TTS model, achieving a Mean Opinion Score (MOS) of 3.76, which is remarkably close to the MOS of 3.86 for the utterances generated by the same vocoder and natural spectrogram, and the MOS of 4.01 for the natural waveform, demonstrating the exceptional quality and effectiveness of the corpus.

【9】 Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
标题: Muskits-ESPnet:新范式中歌唱声音合成的综合工具包
作者:Yuning Wu,Jiatong Shi,Yifeng Yu,Yuxun Tang,Tao Qian,Yueqian Lin,Jionghao Han,Xinyi Bai,Shinji Watanabe,Qin Jin
备注:Accepted by ACMMM 2024 demo track
链接:点击下载PDF文件
摘要:这项研究提出了Muskits-ESPnet,这是一个多功能的工具包,通过在连续和离散方法中应用预训练的音频模型,为歌唱语音合成(SVS)引入了新的范例。具体而言,我们探索了从SSL模型和音频编解码器中衍生的离散表示,并在多功能性和智能性方面提供了显着优势,支持多种格式的输入和各种SVS模型的自适应数据处理工作流。该工具包具有自动乐谱错误检测和校正,以及一个感知自动评估模块,以模仿人类的主观评估分数。Muskits-ESPnet可以在 url{https: github.com espnet espnet}上找到。摘要:This research presents Muskits-ESPnet, a versatile toolkit that introduces new paradigms to Singing Voice Synthesis (SVS) through the application of pretrained audio models in both continuous and discrete approaches. Specifically, we explore discrete representations derived from SSL models and audio codecs and offer significant advantages in versatility and intelligence, supporting multi-format inputs and adaptable data processing workflows for various SVS models. The toolkit features automatic music score error detection and correction, as well as a perception auto-evaluation module to imitate human subjective evaluating scores. Muskits-ESPnet is available at url{https: github.com espnet espnet}.

【10】 Analytic Class Incremental Learning for Sound Source Localization with Privacy Protection
标题: 分析类增量学习,实现具有隐私保护的声音源本地化
作者:Xinyuan Qian,Xianghu Yue,Jiadong Wang,Huiping Zhuang,Haizhou Li
链接:点击下载PDF文件
摘要:声源定位(SSL)技术可用于监控和机器人等应用。虽然传统的基于信号处理(SP)的SSL方法在特定的信号和噪声假设下提供了分析解决方案,但最近基于深度学习(DL)的方法已经大大优于它们。然而,它们的成功取决于大量的训练数据和大量的计算资源。此外,它们通常依赖于大规模的注释空间数据,并且在适应不断变化的声音类别时可能会遇到困难。为了缓解这些挑战,我们提出了一种新的类增量学习(CIL)方法,称为SSL-CIL,它避免了严重的准确性下降,由于灾难性的遗忘通过增量更新基于DL的SSL模型通过一个封闭的形式的解析解。特别是,由于学习过程不会重新访问任何历史数据(无样本),因此确保了数据隐私,这更适合智能家居场景。在公开的SSLR数据集上的实证结果表明,我们的建议具有优越的性能,实现了90.9%的定位准确率,超过了其他竞争对手的方法。摘要:Sound Source Localization (SSL) enabling technology for applications such as surveillance and robotics. While traditional Signal Processing (SP)-based SSL methods provide analytic solutions under specific signal and noise assumptions, recent Deep Learning (DL)-based methods have significantly outperformed them. However, their success depends on extensive training data and substantial computational resources. Moreover, they often rely on large-scale annotated spatial data and may struggle when adapting to evolving sound classes. To mitigate these challenges, we propose a novel Class Incremental Learning (CIL) approach, termed SSL-CIL, which avoids serious accuracy degradation due to catastrophic forgetting by incrementally updating the DL-based SSL model through a closed-form analytic solution. In particular, data privacy is ensured since the learning process does not revisit any historical data (exemplar-free), which is more suitable for smart home scenarios. Empirical results in the public SSLR dataset demonstrate the superior performance of our proposal, achieving a localization accuracy of 90.9%, surpassing other competitive methods.

【11】 Enhancing CTC-Based Visual Speech Recognition
标题: 增强基于ATC的视觉语音识别
作者:Hendrik Laux,Anke Schmeink
链接:点击下载PDF文件
摘要:本文介绍了LiteVSR 2,我们以前介绍的有效的视觉语音识别(VSR)的方法的增强版本。基于我们的知识蒸馏框架,从一个预先训练的自动语音识别(ASR)模型,我们引入了两个关键的改进:一个稳定的视频预处理技术和特征归一化的蒸馏过程。这些改进在LRS2和LRS3基准测试中产生了实质性的性能提升,将LiteVSR 2定位为当前最佳的基于CTC的VSR模型,而不会增加训练数据量或使用的计算资源。此外,我们通过检查不同模型复杂性和训练数据量的性能指标来探索我们方法的可扩展性。LiteVSR 2保持了其前身的效率,同时显著提高了准确性,从而展示了VSR技术在资源效率方面的进步潜力。摘要:This paper presents LiteVSR2, an enhanced version of our previously introduced efficient approach to Visual Speech Recognition (VSR). Building upon our knowledge distillation framework from a pre-trained Automatic Speech Recognition (ASR) model, we introduce two key improvements: a stabilized video preprocessing technique and feature normalization in the distillation process. These improvements yield substantial performance gains on the LRS2 and LRS3 benchmarks, positioning LiteVSR2 as the current best CTC-based VSR model without increasing the volume of training data or computational resources utilized. Furthermore, we explore the scalability of our approach by examining performance metrics across varying model complexities and training data volumes. LiteVSR2 maintains the efficiency of its predecessor while significantly enhancing accuracy, thereby demonstrating the potential for resource-efficient advancements in VSR technology.

【12】 Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition
标题: 用于流语音识别的带有SummaryMixing的线性时间复杂度一致器
作者:Titouan Parcollet,Rogier van Dalen,Shucong Zhang,Sourav Batthacharya
链接:点击下载PDF文件
摘要:自动语音识别(ASR)的编码器配备有自我注意,无论是流或非流,需要二次时间的语音话语的长度。这减慢了训练和解码,增加了它们的成本,并且限制了ASR在受限设备中的部署。SummaryMixing是一种很有前途的线性时间复杂度替代自注意的非流语音识别方法,首次保留或优于自注意模型的准确性。不幸的是,SummaryMixing的原始定义不适合流式语音识别。因此,这项工作将SummaryMixing扩展到在流式传输和离线模式下工作的Conformer换能器。结果表明,这种新的线性时间复杂度的语音编码器优于自注意在这两种情况下,同时需要更少的计算和内存在训练和解码。摘要:Automatic speech recognition (ASR) with an encoder equipped with self-attention, whether streaming or non-streaming, takes quadratic time in the length of the speech utterance. This slows down training and decoding, increase their cost, and limit the deployment of the ASR in constrained devices. SummaryMixing is a promising linear-time complexity alternative to self-attention for non-streaming speech recognition that, for the first time, preserves or outperforms the accuracy of self-attention models. Unfortunately, the original definition of SummaryMixing is not suited to streaming speech recognition. Hence, this work extends SummaryMixing to a Conformer Transducer that works in both a streaming and an offline mode. It shows that this new linear-time complexity speech encoder outperforms self-attention in both scenarios while requiring less compute and memory during training and decoding.

【13】 Developing a Framework for Sonifying Variational Quantum Algorithms: Implications for Music Composition
标题: 开发发声变分量子算法的框架:对音乐创作的影响
作者:Paulo Vitor Itaboraí,Peter Thomas,Arianna Crippa,Karl Jansen,Tim Schwägerl,María Aguado Yáñez
备注:This is a non-edited pre-publication version of a chapter to appear in the book Advances in Quantum Computer Music, World Scientific, editor E. R. Miranda, 2024. ISBN:978-981-98-0017-9
链接:点击下载PDF文件
摘要:本章介绍了变分量子调和器,一个软件工具和音乐接口,重点是变分量子算法(VQA)的最小化步骤的声音化问题,用于模拟量子系统的属性和量子硬件辅助的优化问题。特别是,它详细介绍了使用VQA对二次无约束二元优化(QUBO)问题的可听化。灵活的设计使其未来的应用既可以作为科学研究中听觉显示的发声工具,也可以作为艺术创作的混合量子数字乐器。反过来,声化可以帮助研究人员更好地理解复杂系统,并可以为量子物理和量子计算的培训服务。VQH结构,包括它的软件实现,控制机制,和发音映射进行了详细说明。此外,它指导VQH中QUBO代价函数作为音乐合成对象的设计。讨论扩展到应用量子计算机辅助组成和现场编码性能的量子辅助模拟的影响。艺术作品《Hundreonal Chambers》(Thomas and Itabora,2023年)展示了这一点。摘要:This chapter examines the Variational Quantum Harmonizer, a software tool and musical interface that focuses on the problem of sonification of the minimization steps of Variational Quantum Algorithms (VQA), used for simulating properties of quantum systems and optimization problems assisted by quantum hardware. Particularly, it details the sonification of Quadratic Unconstrained Binary Optimization (QUBO) problems using VQA. A flexible design enables its future applications both as a sonification tool for auditory displays in scientific investigation, and as a hybrid quantum-digital musical instrument for artistic endeavours. In turn, sonification can help researchers understand complex systems better and can serve for the training of quantum physics and quantum computing. The VQH structure, including its software implementation, control mechanisms, and sonification mappings are detailed. Moreover, it guides the design of QUBO cost functions in VQH as a music compositional object. The discussion is extended to the implications of applying quantum-assisted simulation in quantum-computer aided composition and live-coding performances. An artistic output is showcased by the piece textit{Hexagonal Chambers} (Thomas and Itabora 'i, 2023).

【14】 Improving Anomalous Sound Detection via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio Models
标题: 通过预训练音频模型的低等级自适应微调来改进异常声音检测
作者:Xinhu Zheng,Anbai Jiang,Bing Han,Yanmin Qian,Pingyi Fan,Jia Liu,Wei-Qiang Zhang
链接:点击下载PDF文件
摘要:异常声音检测(ASD)通过各种人工智能(AI)技术在工业环境中的应用获得了极大的兴趣。虽然ASD系统具有很大的潜力,但由于数据收集的困难和环境因素的复杂性而导致的泛化问题,ASD系统很难在实际生产现场部署。本文介绍了一种利用音频预训练模型的鲁棒ASD模型。具体来说,我们使用机器操作数据微调这些模型,采用SpecAug作为数据增强策略。此外,我们研究了利用低秩自适应(LoRA)调整而不是完全微调的影响,以解决微调数据有限的问题。我们在DCASE2023任务2数据集上的实验在评估集上建立了77.75%的新基准,与之前的最先进(SOTA)模型相比,显着提高了6.48%,包括顶级传统卷积网络和语音预训练模型,这证明了具有LoRA调整的音频预训练模型的有效性。还进行了消融研究,以展示所提出的方案的有效性。摘要:Anomalous Sound Detection (ASD) has gained significant interest through the application of various Artificial Intelligence (AI) technologies in industrial settings. Though possessing great potential, ASD systems can hardly be readily deployed in real production sites due to the generalization problem, which is primarily caused by the difficulty of data collection and the complexity of environmental factors. This paper introduces a robust ASD model that leverages audio pre-trained models. Specifically, we fine-tune these models using machine operation data, employing SpecAug as a data augmentation strategy. Additionally, we investigate the impact of utilizing Low-Rank Adaptation (LoRA) tuning instead of full fine-tuning to address the problem of limited data for fine-tuning. Our experiments on the DCASE2023 Task 2 dataset establish a new benchmark of 77.75% on the evaluation set, with a significant improvement of 6.48% compared with previous state-of-the-art (SOTA) models, including top-tier traditional convolutional networks and speech pre-trained models, which demonstrates the effectiveness of audio pre-trained models with LoRA tuning. Ablation studies are also conducted to showcase the efficacy of the proposed scheme.

【15】 The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
标题: VoiceMOS 2024挑战:超越语音质量预测
作者:Wen-Chin Huang,Szu-Wei Fu,Erica Cooper,Ryandhimas E. Zezario,Tomoki Toda,Hsin-Min Wang,Junichi Yamagishi,Yu Tsao
备注:Accepted to SLT2024
链接:点击下载PDF文件
摘要:我们介绍了VoiceMOS挑战赛的第三版,这是一项旨在推进人类语音评级自动预测研究的科学举措。有三条轨道。第一个轨道是预测“放大”的语音合成系统的高质量样本的质量。第二个轨道是预测评级的样本从歌唱声音合成和声音转换与各种各样的系统,听众和语言。第三个轨道是半监督质量预测,用于噪声,干净和增强的语音,其中提供了非常少量的标记训练数据。在来自学术界和工业界的八个团队中,我们发现许多团队的表现都优于基线系统。成功的技术包括基于检索的方法和使用非自我监督的表示,如频谱图和音高直方图。这些结果表明,这一挑战推动了主观语音评级预测领域的发展。摘要:We present the third edition of the VoiceMOS Challenge, a scientific initiative designed to advance research into automatic prediction of human speech ratings. There were three tracks. The first track was on predicting the quality of zoomed-in'' high-quality samples from speech synthesis systems. The second track was to predict ratings of samples from singing voice synthesis and voice conversion with a large variety of systems, listeners, and languages. The third track was semi-supervised quality prediction for noisy, clean, and enhanced speech, where a very small amount of labeled training data was provided. Among the eight teams from both academia and industry, we found that many were able to outperform the baseline systems. Successful techniques included retrieval-based methods and the use of non-self-supervised representations like spectrograms and pitch histograms. These results showed that the challenge has advanced the field of subjective speech rating prediction.

【16】 Unveiling Visual Biases in Audio-Visual Localization Benchmarks
标题: 揭露视听本地化基准中的视觉偏见
作者:Liangyu Chen,Zihao Yue,Boshen Xu,Qin Jin
备注:Accepted by ECCV24 AVGenL Workshop
链接:点击下载PDF文件
摘要:视听源定位(AVSL)旨在定位视频中的声源。在本文中,我们确定了一个重要的问题,在现有的基准:发声对象往往很容易识别的基础上,仅仅是视觉线索,我们称之为视觉偏见。这种偏见阻碍了这些基准有效地评估AVSL模型。为了进一步验证我们关于视觉偏差的假设,我们研究了两个代表性的AVSL基准,VGG-SS和EpicSounding-Object,其中仅视觉模型优于所有视听基线。我们的研究结果表明,现有的AVSL基准需要进一步完善,以促进视听学习。摘要:Audio-Visual Source Localization (AVSL) aims to localize the source of sound within a video. In this paper, we identify a significant issue in existing benchmarks: the sounding objects are often easily recognized based solely on visual cues, which we refer to as visual bias. Such biases hinder these benchmarks from effectively evaluating AVSL models. To further validate our hypothesis regarding visual biases, we examine two representative AVSL benchmarks, VGG-SS and EpicSounding-Object, where the vision-only models outperform all audiovisual baselines. Our findings suggest that existing AVSL benchmarks need further refinement to facilitate audio-visual learning.


机器翻译,仅供参考