今日论文合集:cs.SD语音8篇,eess.AS音频处理8篇。本文经arXiv每日学术速递授权转载
【1】Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization
标题:双重Mean-Teacher:一种无偏半监督视听资源定位框架
链接:https://arxiv.org/abs/2403.03145
作者:Yuxin Guo,Shijie Ma,Hu Su,Zhiqing Wang,Yuhao Zhao,Wei Zou,Siyang Sun,Yun Zheng
备注:Accepted to NeurIPS2023
摘要:视听源定位(AVSL)的目的是在给定配对音频片段的视频帧内定位发声对象。现有的方法主要依赖于视听对应的自监督对比学习。如果没有任何边界框注释,它们很难实现精确的定位,特别是对于小对象,并且遭受模糊边界和误报。此外,朴素半监督方法在充分利用大量未标记数据的信息方面很差。在本文中,我们提出了一种新的半监督学习框架AVSL,即双均值教师(DMT),包括两个师生结构,以规避确认偏差问题。具体来说,两个老师,在有限的标记数据上进行预训练,通过他们的预测之间的共识过滤掉噪音样本,然后通过交叉他们的置信度图生成高质量的伪标签。对标记和未标记数据的充分利用以及所提出的无偏框架使DMT能够大幅优于当前最先进的方法,Flickr-SoundNet和VGG-Sound Source上的CIoU分别为90.4%和48.8%,分别比自监督和半监督方法提高了8.9%,9.6%和4.6%,6.4%,只有3%的位置标注。我们还将我们的框架扩展到一些现有的AVSL方法,并不断提高其性能。
摘要:Audio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised contrastive learning of audio-visual correspondence. Without any bounding-box annotations, they struggle to achieve precise localization, especially for small objects, and suffer from blurry boundaries and false positives. Moreover, the naive semi-supervised method is poor in fully leveraging the information of abundant unlabeled data. In this paper, we propose a novel semi-supervised learning framework for AVSL, namely Dual Mean-Teacher (DMT), comprising two teacher-student structures to circumvent the confirmation bias issue. Specifically, two teachers, pre-trained on limited labeled data, are employed to filter out noisy samples via the consensus between their predictions, and then generate high-quality pseudo-labels by intersecting their confidence maps. The sufficient utilization of both labeled and unlabeled data and the proposed unbiased framework enable DMT to outperform current state-of-the-art methods by a large margin, with CIoU of 90.4% and 48.8% on Flickr-SoundNet and VGG-Sound Source, obtaining 8.9%, 9.6% and 4.6%, 6.4% improvements over self- and semi-supervised methods respectively, given only 3% positional-annotations. We also extend our framework to some existing AVSL methods and consistently boost their performance.
【2】 Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization作者:Yuxin Guo,Shijie Ma,Yuhao Zhao,Hu Su,Wei Zou备注:Accepted To ICASSP2024摘要:视听源定位(AVSL)是在给定音频线索的场景中识别特定发声对象的任务。在我们的工作中,我们专注于半监督AVSL与伪标记。为了解决香草硬伪标签的问题,包括偏差积累,噪声敏感性和不稳定性,我们提出了一种新的方法称为交叉伪标签(XPL),其中两个模型相互学习交叉细化机制,以避免偏差积累。我们为XPL配备了两个有效的组件。首先,具有锐化和伪标签指数移动平均机制的软伪标签使模型能够实现渐进的自我改进,并确保稳定的训练。其次,课程数据选择模块在训练期间自适应地选择具有高质量的伪标签,以减轻潜在的偏差。实验结果表明,XPL显著优于现有方法,实现了最先进的性能,同时有效地减轻了确认偏差并确保了训练稳定性。摘要:Audio-Visual Source Localization (AVSL) is the task of identifying specific sounding objects in the scene given audio cues. In our work, we focus on semi-supervised AVSL with pseudo-labeling. To address the issues with vanilla hard pseudo-labels including bias accumulation, noise sensitivity, and instability, we propose a novel method named Cross Pseudo-Labeling (XPL), wherein two models learn from each other with the cross-refine mechanism to avoid bias accumulation. We equip XPL with two effective components. Firstly, the soft pseudo-labels with sharpening and pseudo-label exponential moving average mechanisms enable models to achieve gradual self-improvement and ensure stable training. Secondly, the curriculum data selection module adaptively selects pseudo-labels with high quality during training to mitigate potential bias. Experimental results demonstrate that XPL significantly outperforms existing methods, achieving state-of-the-art performance while effectively mitigating confirmation bias and ensuring training stability.【3】 AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models标题:AIX速度:使用语音识别模型的听力理解优化播放速度作者:Kazuki Kawamura,Jun Rekimoto摘要:由于人类可以以比实际观察到的更快的速度收听音频和观看视频,因此我们经常以更高的播放速度收听或观看这些内容,以提高内容理解的时间效率。为了进一步利用这种能力,已经开发了根据用户的状况和内容的类型自动调整回放速度以帮助更有效地理解时间序列内容的系统。然而,这些系统仍然有空间通过生成具有针对更精细的时间单位优化的回放速度的语音并将其提供给人类来进一步扩展人类的速度收听能力。在这项研究中,我们确定人类是否可以听到优化的语音,并提出了一个系统,自动调整播放速度的单位小音素,同时确保语音清晰度。该系统使用语音识别器分数作为人类能够多好地听到某个语音单元的代理,并且将语音回放速度最大化到人类能够听到的程度。这种方法可以用来产生快速但可理解的语音。在评估实验中,我们比较了语音播放在一个恒定的快速和灵活的速度加快所提出的方法产生的语音在盲测试,并确认所提出的方法产生的语音更容易听。摘要:Since humans can listen to audio and watch videos at faster speeds than actually observed, we often listen to or watch these pieces of content at higher playback speeds to increase the time efficiency of content comprehension. To further utilize this capability, systems that automatically adjust the playback speed according to the user's condition and the type of content to assist in more efficient comprehension of time-series content have been developed. However, there is still room for these systems to further extend human speed-listening ability by generating speech with playback speed optimized for even finer time units and providing it to humans. In this study, we determine whether humans can hear the optimized speech and propose a system that automatically adjusts playback speed at units as small as phonemes while ensuring speech intelligibility. The system uses the speech recognizer score as a proxy for how well a human can hear a certain unit of speech and maximizes the speech playback speed to the extent that a human can hear. This method can be used to produce fast but intelligible speech. In the evaluation experiment, we compared the speech played back at a constant fast speed and the flexibly speed-up speech generated by the proposed method in a blind test and confirmed that the proposed method produced speech that was easier to listen to.【4】 Single-Channel Robot Ego-Speech Filtering during Human-Robot Interaction作者:Yue Li,Koen V Hindriks,Florian Kunneman备注:Accepted by ACM Technological Advances in Human-Robot Interaction. 9 pages摘要:在本文中,我们研究如何以及人类语音可以自动过滤时,这与社会机器人,胡椒的声音和风扇噪音重叠。我们最终的目标是HRI场景,其中麦克风可以在机器人说话时保持打开,从而实现更自然的回合转换方案,其中人类可以打断机器人。为了做出适当的反应,机器人需要理解对话者在语音的重叠部分说了什么,这可以通过目标语音提取(TSE)来完成。为了研究在流行的社交机器人Pepper的背景下如何很好地完成TSE,我们着手制造一个数据库,该数据库由Pepper本身的录音语音、它的风扇噪音(靠近麦克风)和Pepper麦克风记录的人类语音组成,在一个低混响和高混响的房间里。比较信号处理方法,有和没有后滤波,和卷积递归神经网络(CRNN)的方法,一个国家的最先进的说话人识别为基础的TSE模型,我们发现,没有后滤波的信号处理方法产生的字错误率方面的最佳性能的重叠语音信号低混响,而CRNN方法是更强大的混响。这些结果表明,估计重叠语音与机器人的人的声音是可能的,在现实生活中的应用,提供的房间混响是低的,人的语音具有高音量或高音调。摘要:In this paper, we study how well human speech can automatically be filtered when this overlaps with the voice and fan noise of a social robot, Pepper. We ultimately aim for an HRI scenario where the microphone can remain open when the robot is speaking, enabling a more natural turn-taking scheme where the human can interrupt the robot. To respond appropriately, the robot would need to understand what the interlocutor said in the overlapping part of the speech, which can be accomplished by target speech extraction (TSE). To investigate how well TSE can be accomplished in the context of the popular social robot Pepper, we set out to manufacture a datase composed of a mixture of recorded speech of Pepper itself, its fan noise (which is close to the microphones), and human speech as recorded by the Pepper microphone, in a room with low reverberation and high reverberation. Comparing a signal processing approach, with and without post-filtering, and a convolutional recurrent neural network (CRNN) approach to a state-of-the-art speaker identification-based TSE model, we found that the signal processing approach without post-filtering yielded the best performance in terms of Word Error Rate on the overlapping speech signals with low reverberation, while the CRNN approach is more robust for reverberation. These results show that estimating the human voice in overlapping speech with a robot is possible in real-life application, provided that the room reverberation is low and the human speech has a high volume or high pitch.
【5】 Fighting Game Adaptive Background Music for Improved Gameplay作者:Ibrahim Khan,Thai Van Nguyen,Chollakorn Nimpattanavong,Ruck Thawonmas备注:This is an updated version of our IEEE CoG 2023 paper (this https URL). This version has revised the description of the association between the distance between the two players (PD) and the instrument's volume on page 2. arXiv admin note: substantial text overlap with arXiv:2303.15734摘要:本文介绍了我们的工作,以提高背景音乐(BGM)在DareFightingICE通过添加自适应功能。自适应BGM由三种不同类别的乐器组成,演奏2022年DareFightingICE比赛获奖声音设计的BGM。BGM通过改变每种乐器的音量进行调整。每个类别都连接到游戏的不同元素。然后,我们通过使用仅使用音频作为输入的深度强化学习AI代理(盲DL AI)来运行实验以评估自适应BGM。结果表明,与没有自适应BGM的播放相比,在播放自适应BGM时,盲DL AI的性能有所改善。摘要:This paper presents our work to enhance the background music (BGM) in DareFightingICE by adding adaptive features. The adaptive BGM consists of three different categories of instruments playing the BGM of the winner sound design from the 2022 DareFightingICE Competition. The BGM adapts by changing the volume of each category of instruments. Each category is connected to a different element of the game. We then run experiments to evaluate the adaptive BGM by using a deep reinforcement learning AI agent that only uses audio as input (Blind DL AI). The results show that the performance of the Blind DL AI improves while playing with the adaptive BGM as compared to playing without the adaptive BGM.【6】 Enhanced DareFightingICE Competitions: Sound Design and AI Competitions标题:增强的DareFightingICE比赛:声音设计和人工智能比赛作者:Ibrahim Khan,Chollakorn Nimpattanavong,Thai Van Nguyen,Kantinan Plupattanakit,Ruck Thawonmas摘要:本文介绍了一个新的和改进的DareFightingICE平台,一个专注于视障玩家(VIP)的格斗游戏平台,在Unity游戏引擎中。它还介绍了DareFightingICE比赛分为两个独立的比赛,称为DareFightingICE声音设计比赛和DareFightingICE AI比赛-在2024年IEEE游戏会议(CoG)上-其中将使用新平台。这个新平台是旧的DareFightingICE平台的增强版本,具有更好的音频系统来传达3D声音,以及更好的方式将音频数据发送给AI代理。通过这种增强和利用Unity,新的DareFightingICE平台在为VIP和未来的音频研究添加新功能方面更容易访问。本文还改进了声音设计比赛中声音设计的评价方法,以确保在未来的CoG中继续举办这项比赛,为VIP提供更好的声音设计。据我们所知,我们的两个比赛都是同类比赛中的第一个,比赛之间的联系是随着时间的推移相互提高参赛作品的质量,这使得这些比赛成为代表更广泛的游戏社区中经常被忽视的部分的重要组成部分,VIP。摘要:This paper presents a new and improved DareFightingICE platform, a fighting game platform with a focus on visually impaired players (VIPs), in the Unity game engine. It also introduces the separation of the DareFightingICE Competition into two standalone competitions called DareFightingICE Sound Design Competition and DareFightingICE AI Competition--at the 2024 IEEE Conference on Games (CoG)--in which a new platform will be used. This new platform is an enhanced version of the old DareFightingICE platform, having a better audio system to convey 3D sound and a better way to send audio data to AI agents. With this enhancement and by utilizing Unity, the new DareFightingICE platform is more accessible in terms of adding new features for VIPs and future audio research. This paper also improves the evaluation method for evaluating sound designs in the Sound Design Competition which will ensure a better sound design for VIPs as this competition continues to run at future CoG. To the best of our knowledge, both of our competitions are first of their kind, and the connection between the competitions to mutually improve the entries' quality with time makes these competitions an important part of representing an often overlooked segment within the broader gaming community, VIPs.
【7】 NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models标题:NaturalSpeech 3:基于因子分解编解码器和扩散模型的零发声语音合成作者:Zeqian Ju,Yuancheng Wang,Kai Shen,Xu Tan,Detai Xin,Dongchao Yang,Yanqing Liu,Yichong Leng,Kaitao Song,Siliang Tang,Zhizheng Wu,Tao Qin,Xiang-Yang Li,Wei Ye,Shikun Zhang,Jiang Bian,Lei He,Jinyu Li,Sheng Zhao备注:22 pages, 15 tables, 3 figures摘要:虽然最近的大规模文本到语音(TTS)模型已经取得了显着的进展,他们仍然在语音质量,相似性和韵律不足。考虑到语音复杂地包含各种属性(例如,内容、韵律、音色和声学细节)对生成构成重大挑战,自然的想法是将语音分解成表示不同属性的各个子空间并单独生成它们。受此启发,我们提出了NaturalSpeech 3,一个TTS系统,采用新颖的因子分解扩散模型以zero-shot方式生成自然语音。具体而言,1)我们设计了一个神经编解码器与因子分解矢量量化(FVQ),以解开语音波形到子空间的内容,韵律,音色,和声学细节; 2)我们提出了一个因子分解的扩散模型,以产生属性在每个子空间后,其相应的提示。通过这种分解设计,NaturalSpeech 3可以以分而治之的方式有效地建模具有分解子空间的复杂语音。实验表明,NaturalSpeech 3在质量、相似度、韵律和可懂度方面优于最先进的TTS系统。此外,我们通过扩展到1B参数和20万小时的训练数据来实现更好的性能。摘要:While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intricately encompasses various attributes (e.g., content, prosody, timbre, and acoustic details) that pose significant challenges for generation, a natural idea is to factorize speech into individual subspaces representing different attributes and generate them individually. Motivated by it, we propose NaturalSpeech 3, a TTS system with novel factorized diffusion models to generate natural speech in a zero-shot way. Specifically, 1) we design a neural codec with factorized vector quantization (FVQ) to disentangle speech waveform into subspaces of content, prosody, timbre, and acoustic details; 2) we propose a factorized diffusion model to generate attributes in each subspace following its corresponding prompt. With this factorization design, NaturalSpeech 3 can effectively and efficiently model the intricate speech with disentangled subspaces in a divide-and-conquer way. Experiments show that NaturalSpeech 3 outperforms the state-of-the-art TTS systems on quality, similarity, prosody, and intelligibility. Furthermore, we achieve better performance by scaling to 1B parameters and 200K hours of training data.
【8】 NeuroVoz: a Castillian Spanish corpus of parkinsonian speech标题:NeuroVoz:帕金森症演讲的卡斯蒂利亚西班牙语语料库作者:Janaína Mendes-Laureano,Jorge A. Gómez-García,Alejandro Guerrero-López,Elisa Luque-Buzo,Julián D. Arias-Londoño,Francisco J. Grandas-Pérez,Juan I. Godino-Llorente备注:Preprint version摘要:通过语音分析诊断帕金森病(PD)的进展受到明显缺乏公开可用的多样化语言数据集的阻碍,限制了现有研究的可重复性和进一步探索。 为了应对这一差距,我们引入了一个全面的语料库,来自108名西班牙语卡斯蒂利亚本地人,包括55名健康对照和53名诊断为PD的个体,所有这些人都接受了药物治疗,并在药物优化状态下记录。这个独特的数据集具有广泛的语音任务,包括五个西班牙元音的持续发声,diadochokinetic测试,16个重复和重复的话语,以及自由独白。该数据集通过专家手动翻译重复任务强调准确性和可靠性,并利用Whisper进行自动独白翻译,使其成为最完整的帕金森氏症语音公共语料库,也是卡斯蒂利亚西班牙语中的第一个。 NeuroVoz由2,903个音频记录组成,平均每位参与者的录音价格为26.88美元,为PD对语音的影响的科学探索提供了大量资源。该数据集已经支持了几项研究,在PD语音模式识别中达到了89%的基准准确率,表明PD导致了显著的语音改变。尽管有这些进展,更广泛的挑战进行语言不可知论,跨语料库分析帕金森氏症的语音模式仍然是一个开放的领域,为未来的研究。这一贡献不仅填补了PD语音分析资源的关键空白,而且为全球研究界利用语音作为神经退行性疾病的诊断工具设定了新标准。摘要:The advancement of Parkinson's Disease (PD) diagnosis through speech analysis is hindered by a notable lack of publicly available, diverse language datasets, limiting the reproducibility and further exploration of existing research. In response to this gap, we introduce a comprehensive corpus from 108 native Castilian Spanish speakers, comprising 55 healthy controls and 53 individuals diagnosed with PD, all of whom were under pharmacological treatment and recorded in their medication-optimized state. This unique dataset features a wide array of speech tasks, including sustained phonation of the five Spanish vowels, diadochokinetic tests, 16 listen-and-repeat utterances, and free monologues. The dataset emphasizes accuracy and reliability through specialist manual transcriptions of the listen-and-repeat tasks and utilizes Whisper for automated monologue transcriptions, making it the most complete public corpus of Parkinsonian speech, and the first in Castillian Spanish. NeuroVoz is composed by 2,903 audio recordings averaging $26.88 \pm 3.35$ recordings per participant, offering a substantial resource for the scientific exploration of PD's impact on speech. This dataset has already underpinned several studies, achieving a benchmark accuracy of 89% in PD speech pattern identification, indicating marked speech alterations attributable to PD. Despite these advances, the broader challenge of conducting a language-agnostic, cross-corpora analysis of Parkinsonian speech patterns remains an open area for future research. This contribution not only fills a critical void in PD speech analysis resources but also sets a new standard for the global research community in leveraging speech as a diagnostic tool for neurodegenerative diseases.
【1】 NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models标题:NaturalSpeech 3:基于因子分解编解码器和扩散模型的零发声语音合成作者:Zeqian Ju,Yuancheng Wang,Kai Shen,Xu Tan,Detai Xin,Dongchao Yang,Yanqing Liu,Yichong Leng,Kaitao Song,Siliang Tang,Zhizheng Wu,Tao Qin,Xiang-Yang Li,Wei Ye,Shikun Zhang,Jiang Bian,Lei He,Jinyu Li,Sheng Zhao备注:22 pages, 15 tables, 3 figures摘要:虽然最近的大规模文本到语音(TTS)模型已经取得了显着的进展,他们仍然在语音质量,相似性和韵律不足。考虑到语音复杂地包含各种属性(例如,内容、韵律、音色和声学细节)对生成构成重大挑战,自然的想法是将语音分解成表示不同属性的各个子空间并单独生成它们。受此启发,我们提出了NaturalSpeech 3,一个TTS系统,采用新颖的因子分解扩散模型以zero-shot方式生成自然语音。具体而言,1)我们设计了一个神经编解码器与因子分解矢量量化(FVQ),以解开语音波形到子空间的内容,韵律,音色,和声学细节; 2)我们提出了一个因子分解的扩散模型,以产生属性在每个子空间后,其相应的提示。通过这种分解设计,NaturalSpeech 3可以以分而治之的方式有效地建模具有分解子空间的复杂语音。实验表明,NaturalSpeech 3在质量、相似度、韵律和可懂度方面优于最先进的TTS系统。此外,我们通过扩展到1B参数和20万小时的训练数据来实现更好的性能。摘要:While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intricately encompasses various attributes (e.g., content, prosody, timbre, and acoustic details) that pose significant challenges for generation, a natural idea is to factorize speech into individual subspaces representing different attributes and generate them individually. Motivated by it, we propose NaturalSpeech 3, a TTS system with novel factorized diffusion models to generate natural speech in a zero-shot way. Specifically, 1) we design a neural codec with factorized vector quantization (FVQ) to disentangle speech waveform into subspaces of content, prosody, timbre, and acoustic details; 2) we propose a factorized diffusion model to generate attributes in each subspace following its corresponding prompt. With this factorization design, NaturalSpeech 3 can effectively and efficiently model the intricate speech with disentangled subspaces in a divide-and-conquer way. Experiments show that NaturalSpeech 3 outperforms the state-of-the-art TTS systems on quality, similarity, prosody, and intelligibility. Furthermore, we achieve better performance by scaling to 1B parameters and 200K hours of training data.【2】 NeuroVoz: a Castillian Spanish corpus of parkinsonian speech标题:NeuroVoz:帕金森症演讲的卡斯蒂利亚西班牙语语料库作者:Janaína Mendes-Laureano,Jorge A. Gómez-García,Alejandro Guerrero-López,Elisa Luque-Buzo,Julián D. Arias-Londoño,Francisco J. Grandas-Pérez,Juan I. Godino-Llorente摘要:通过语音分析诊断帕金森病(PD)的进展受到明显缺乏公开可用的多样化语言数据集的阻碍,限制了现有研究的可重复性和进一步探索。 为了应对这一差距,我们引入了一个全面的语料库,来自108名西班牙语卡斯蒂利亚本地人,包括55名健康对照和53名诊断为PD的个体,所有这些人都接受了药物治疗,并在药物优化状态下记录。这个独特的数据集具有广泛的语音任务,包括五个西班牙元音的持续发声,diadochokinetic测试,16个重复和重复的话语,以及自由独白。该数据集通过专家手动翻译重复任务强调准确性和可靠性,并利用Whisper进行自动独白翻译,使其成为最完整的帕金森氏症语音公共语料库,也是卡斯蒂利亚西班牙语中的第一个。 NeuroVoz由2,903个音频记录组成,平均每位参与者的录音价格为26.88美元,为PD对语音的影响的科学探索提供了大量资源。该数据集已经支持了几项研究,在PD语音模式识别中达到了89%的基准准确率,表明PD导致了显著的语音改变。尽管有这些进展,更广泛的挑战进行语言不可知论,跨语料库分析帕金森氏症的语音模式仍然是一个开放的领域,为未来的研究。这一贡献不仅填补了PD语音分析资源的关键空白,而且为全球研究界利用语音作为神经退行性疾病的诊断工具设定了新标准。摘要:The advancement of Parkinson's Disease (PD) diagnosis through speech analysis is hindered by a notable lack of publicly available, diverse language datasets, limiting the reproducibility and further exploration of existing research. In response to this gap, we introduce a comprehensive corpus from 108 native Castilian Spanish speakers, comprising 55 healthy controls and 53 individuals diagnosed with PD, all of whom were under pharmacological treatment and recorded in their medication-optimized state. This unique dataset features a wide array of speech tasks, including sustained phonation of the five Spanish vowels, diadochokinetic tests, 16 listen-and-repeat utterances, and free monologues. The dataset emphasizes accuracy and reliability through specialist manual transcriptions of the listen-and-repeat tasks and utilizes Whisper for automated monologue transcriptions, making it the most complete public corpus of Parkinsonian speech, and the first in Castillian Spanish. NeuroVoz is composed by 2,903 audio recordings averaging $26.88 \pm 3.35$ recordings per participant, offering a substantial resource for the scientific exploration of PD's impact on speech. This dataset has already underpinned several studies, achieving a benchmark accuracy of 89% in PD speech pattern identification, indicating marked speech alterations attributable to PD. Despite these advances, the broader challenge of conducting a language-agnostic, cross-corpora analysis of Parkinsonian speech patterns remains an open area for future research. This contribution not only fills a critical void in PD speech analysis resources but also sets a new standard for the global research community in leveraging speech as a diagnostic tool for neurodegenerative diseases.【3】 Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization标题:双重Mean-Teacher:一种无偏半监督视听资源定位框架作者:Yuxin Guo,Shijie Ma,Hu Su,Zhiqing Wang,Yuhao Zhao,Wei Zou,Siyang Sun,Yun Zheng备注:Accepted to NeurIPS2023摘要:视听源定位(AVSL)的目的是在给定配对音频片段的视频帧内定位发声对象。现有的方法主要依赖于视听对应的自监督对比学习。如果没有任何边界框注释,它们很难实现精确的定位,特别是对于小对象,并且遭受模糊边界和误报。此外,朴素半监督方法在充分利用大量未标记数据的信息方面很差。在本文中,我们提出了一种新的半监督学习框架AVSL,即双均值教师(DMT),包括两个师生结构,以规避确认偏差问题。具体来说,两个老师,在有限的标记数据上进行预训练,通过他们的预测之间的共识过滤掉噪音样本,然后通过交叉他们的置信度图生成高质量的伪标签。对标记和未标记数据的充分利用以及所提出的无偏框架使DMT能够大幅优于当前最先进的方法,Flickr-SoundNet和VGG-Sound Source上的CIoU分别为90.4%和48.8%,分别比自监督和半监督方法提高了8.9%,9.6%和4.6%,6.4%,只有3%的位置标注。我们还将我们的框架扩展到一些现有的AVSL方法,并不断提高其性能。摘要:Audio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised contrastive learning of audio-visual correspondence. Without any bounding-box annotations, they struggle to achieve precise localization, especially for small objects, and suffer from blurry boundaries and false positives. Moreover, the naive semi-supervised method is poor in fully leveraging the information of abundant unlabeled data. In this paper, we propose a novel semi-supervised learning framework for AVSL, namely Dual Mean-Teacher (DMT), comprising two teacher-student structures to circumvent the confirmation bias issue. Specifically, two teachers, pre-trained on limited labeled data, are employed to filter out noisy samples via the consensus between their predictions, and then generate high-quality pseudo-labels by intersecting their confidence maps. The sufficient utilization of both labeled and unlabeled data and the proposed unbiased framework enable DMT to outperform current state-of-the-art methods by a large margin, with CIoU of 90.4% and 48.8% on Flickr-SoundNet and VGG-Sound Source, obtaining 8.9%, 9.6% and 4.6%, 6.4% improvements over self- and semi-supervised methods respectively, given only 3% positional-annotations. We also extend our framework to some existing AVSL methods and consistently boost their performance.【4】 Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization作者:Yuxin Guo,Shijie Ma,Yuhao Zhao,Hu Su,Wei Zou备注:Accepted To ICASSP2024摘要:视听源定位(AVSL)是在给定音频线索的场景中识别特定发声对象的任务。在我们的工作中,我们专注于半监督AVSL与伪标记。为了解决香草硬伪标签的问题,包括偏差积累,噪声敏感性和不稳定性,我们提出了一种新的方法称为交叉伪标签(XPL),其中两个模型相互学习交叉细化机制,以避免偏差积累。我们为XPL配备了两个有效的组件。首先,具有锐化和伪标签指数移动平均机制的软伪标签使模型能够实现渐进的自我改进,并确保稳定的训练。其次,课程数据选择模块在训练期间自适应地选择具有高质量的伪标签,以减轻潜在的偏差。实验结果表明,XPL显著优于现有方法,实现了最先进的性能,同时有效地减轻了确认偏差并确保了训练稳定性。摘要:Audio-Visual Source Localization (AVSL) is the task of identifying specific sounding objects in the scene given audio cues. In our work, we focus on semi-supervised AVSL with pseudo-labeling. To address the issues with vanilla hard pseudo-labels including bias accumulation, noise sensitivity, and instability, we propose a novel method named Cross Pseudo-Labeling (XPL), wherein two models learn from each other with the cross-refine mechanism to avoid bias accumulation. We equip XPL with two effective components. Firstly, the soft pseudo-labels with sharpening and pseudo-label exponential moving average mechanisms enable models to achieve gradual self-improvement and ensure stable training. Secondly, the curriculum data selection module adaptively selects pseudo-labels with high quality during training to mitigate potential bias. Experimental results demonstrate that XPL significantly outperforms existing methods, achieving state-of-the-art performance while effectively mitigating confirmation bias and ensuring training stability.【5】 AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models标题:AIX速度:使用语音识别模型的听力理解优化播放速度作者:Kazuki Kawamura,Jun Rekimoto摘要:由于人类可以以比实际观察到的更快的速度收听音频和观看视频,因此我们经常以更高的播放速度收听或观看这些内容,以提高内容理解的时间效率。为了进一步利用这种能力,已经开发了根据用户的状况和内容的类型自动调整回放速度以帮助更有效地理解时间序列内容的系统。然而,这些系统仍然有空间通过生成具有针对更精细的时间单位优化的回放速度的语音并将其提供给人类来进一步扩展人类的速度收听能力。在这项研究中,我们确定人类是否可以听到优化的语音,并提出了一个系统,自动调整播放速度的单位小音素,同时确保语音清晰度。该系统使用语音识别器分数作为人类能够多好地听到某个语音单元的代理,并且将语音回放速度最大化到人类能够听到的程度。这种方法可以用来产生快速但可理解的语音。在评估实验中,我们比较了语音播放在一个恒定的快速和灵活的速度加快所提出的方法产生的语音在盲测试,并确认所提出的方法产生的语音更容易听。摘要:Since humans can listen to audio and watch videos at faster speeds than actually observed, we often listen to or watch these pieces of content at higher playback speeds to increase the time efficiency of content comprehension. To further utilize this capability, systems that automatically adjust the playback speed according to the user's condition and the type of content to assist in more efficient comprehension of time-series content have been developed. However, there is still room for these systems to further extend human speed-listening ability by generating speech with playback speed optimized for even finer time units and providing it to humans. In this study, we determine whether humans can hear the optimized speech and propose a system that automatically adjusts playback speed at units as small as phonemes while ensuring speech intelligibility. The system uses the speech recognizer score as a proxy for how well a human can hear a certain unit of speech and maximizes the speech playback speed to the extent that a human can hear. This method can be used to produce fast but intelligible speech. In the evaluation experiment, we compared the speech played back at a constant fast speed and the flexibly speed-up speech generated by the proposed method in a blind test and confirmed that the proposed method produced speech that was easier to listen to.【6】 Single-Channel Robot Ego-Speech Filtering during Human-Robot Interaction作者:Yue Li,Koen V Hindriks,Florian Kunneman备注:Accepted by ACM Technological Advances in Human-Robot Interaction. 9 pages摘要:在本文中,我们研究如何以及人类语音可以自动过滤时,这与社会机器人,胡椒的声音和风扇噪音重叠。我们最终的目标是HRI场景,其中麦克风可以在机器人说话时保持打开,从而实现更自然的回合转换方案,其中人类可以打断机器人。为了做出适当的反应,机器人需要理解对话者在语音的重叠部分说了什么,这可以通过目标语音提取(TSE)来完成。为了研究在流行的社交机器人Pepper的背景下如何很好地完成TSE,我们着手制造一个数据库,该数据库由Pepper本身的录音语音、它的风扇噪音(靠近麦克风)和Pepper麦克风记录的人类语音组成,在一个低混响和高混响的房间里。比较信号处理方法,有和没有后滤波,和卷积递归神经网络(CRNN)的方法,一个国家的最先进的说话人识别为基础的TSE模型,我们发现,没有后滤波的信号处理方法产生的字错误率方面的最佳性能的重叠语音信号低混响,而CRNN方法是更强大的混响。这些结果表明,估计重叠语音与机器人的人的声音是可能的,在现实生活中的应用,提供的房间混响是低的,人的语音具有高音量或高音调。摘要:In this paper, we study how well human speech can automatically be filtered when this overlaps with the voice and fan noise of a social robot, Pepper. We ultimately aim for an HRI scenario where the microphone can remain open when the robot is speaking, enabling a more natural turn-taking scheme where the human can interrupt the robot. To respond appropriately, the robot would need to understand what the interlocutor said in the overlapping part of the speech, which can be accomplished by target speech extraction (TSE). To investigate how well TSE can be accomplished in the context of the popular social robot Pepper, we set out to manufacture a datase composed of a mixture of recorded speech of Pepper itself, its fan noise (which is close to the microphones), and human speech as recorded by the Pepper microphone, in a room with low reverberation and high reverberation. Comparing a signal processing approach, with and without post-filtering, and a convolutional recurrent neural network (CRNN) approach to a state-of-the-art speaker identification-based TSE model, we found that the signal processing approach without post-filtering yielded the best performance in terms of Word Error Rate on the overlapping speech signals with low reverberation, while the CRNN approach is more robust for reverberation. These results show that estimating the human voice in overlapping speech with a robot is possible in real-life application, provided that the room reverberation is low and the human speech has a high volume or high pitch.【7】 Fighting Game Adaptive Background Music for Improved Gameplay作者:Ibrahim Khan,Thai Van Nguyen,Chollakorn Nimpattanavong,Ruck Thawonmas备注:This is an updated version of our IEEE CoG 2023 paper (this https URL). This version has revised the description of the association between the distance between the two players (PD) and the instrument's volume on page 2. arXiv admin note: substantial text overlap with arXiv:2303.15734摘要:本文介绍了我们的工作,以提高背景音乐(BGM)在DareFightingICE通过添加自适应功能。自适应BGM由三种不同类别的乐器组成,演奏2022年DareFightingICE比赛获奖声音设计的BGM。BGM通过改变每种乐器的音量进行调整。每个类别都连接到游戏的不同元素。然后,我们通过使用仅使用音频作为输入的深度强化学习AI代理(盲DL AI)来运行实验以评估自适应BGM。结果表明,与没有自适应BGM的播放相比,在播放自适应BGM时,盲DL AI的性能有所改善。摘要:This paper presents our work to enhance the background music (BGM) in DareFightingICE by adding adaptive features. The adaptive BGM consists of three different categories of instruments playing the BGM of the winner sound design from the 2022 DareFightingICE Competition. The BGM adapts by changing the volume of each category of instruments. Each category is connected to a different element of the game. We then run experiments to evaluate the adaptive BGM by using a deep reinforcement learning AI agent that only uses audio as input (Blind DL AI). The results show that the performance of the Blind DL AI improves while playing with the adaptive BGM as compared to playing without the adaptive BGM.
【8】 Enhanced DareFightingICE Competitions: Sound Design and AI Competitions标题:增强的DareFightingICE比赛:声音设计和人工智能比赛作者:Ibrahim Khan,Chollakorn Nimpattanavong,Thai Van Nguyen,Kantinan Plupattanakit,Ruck Thawonmas摘要:本文介绍了一个新的和改进的DareFightingICE平台,一个专注于视障玩家(VIP)的格斗游戏平台,在Unity游戏引擎中。它还介绍了DareFightingICE比赛分为两个独立的比赛,称为DareFightingICE声音设计比赛和DareFightingICE AI比赛-在2024年IEEE游戏会议(CoG)上-其中将使用新平台。这个新平台是旧的DareFightingICE平台的增强版本,具有更好的音频系统来传达3D声音,以及更好的方式将音频数据发送给AI代理。通过这种增强和利用Unity,新的DareFightingICE平台在为VIP和未来的音频研究添加新功能方面更容易访问。本文还改进了声音设计比赛中声音设计的评价方法,以确保在未来的CoG中继续举办这项比赛,为VIP提供更好的声音设计。据我们所知,我们的两个比赛都是同类比赛中的第一个,比赛之间的联系是随着时间的推移相互提高参赛作品的质量,这使得这些比赛成为代表更广泛的游戏社区中经常被忽视的部分的重要组成部分,VIP。摘要:This paper presents a new and improved DareFightingICE platform, a fighting game platform with a focus on visually impaired players (VIPs), in the Unity game engine. It also introduces the separation of the DareFightingICE Competition into two standalone competitions called DareFightingICE Sound Design Competition and DareFightingICE AI Competition--at the 2024 IEEE Conference on Games (CoG)--in which a new platform will be used. This new platform is an enhanced version of the old DareFightingICE platform, having a better audio system to convey 3D sound and a better way to send audio data to AI agents. With this enhancement and by utilizing Unity, the new DareFightingICE platform is more accessible in terms of adding new features for VIPs and future audio research. This paper also improves the evaluation method for evaluating sound designs in the Sound Design Competition which will ensure a better sound design for VIPs as this competition continues to run at future CoG. To the best of our knowledge, both of our competitions are first of their kind, and the connection between the competitions to mutually improve the entries' quality with time makes these competitions an important part of representing an often overlooked segment within the broader gaming community, VIPs.