本文经arXiv每日学术速递授权转载
【1】 Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation
链接:https://arxiv.org/abs/2411.12719
备注:19 pages, 12 Figures
摘要:尽管TTS模型发展迅速,但仍然缺乏一致和强大的人类评估框架。例如,MOS测试无法区分相似的模型,CMOS的成对比较是时间密集型的。MUSHRA测试是同时评估多个TTS系统的一个有前途的替代方案,但在这项工作中,我们表明,它对匹配人类参考语音的依赖过度地损害了现代TTS系统的得分,这些系统的语音质量可能超过人类语音质量。更具体地说,我们对MUSHRA测试进行了全面的评估,重点关注其对评分员变异性,听众疲劳和参考偏差等因素的敏感性。基于我们对471名印地语和泰米尔语的人类听众的广泛评估,我们发现了两个主要缺点:(i)参考匹配偏差,其中评分员受到人类参考的过度影响,以及(ii)判断模糊性,由于缺乏明确的细粒度指导方针。为了解决这些问题,我们提出了MUSHRA测试的两个改进版本。第一种变体能够对超过人类参考质量的合成样本进行更公平的评级。第二种变体减少了模糊性,如评分员之间相对较低的方差所示。通过结合这些方法,我们实现了更可靠和更细粒度的评估。我们还发布了MANGO,这是一个包含47,100个人类评分的大规模数据集,是印度语言的首个同类集合,有助于分析人类偏好并开发用于评估TTS系统的自动指标。
摘要:Despite rapid advancements in TTS models, a consistent and robust humanevaluation framework is still lacking. For example, MOS tests fail todifferentiate between similar models, and CMOS's pairwise comparisons aretime-intensive. The MUSHRA test is a promising alternative for evaluatingmultiple TTS systems simultaneously, but in this work we show that its relianceon matching human reference speech unduly penalises the scores of modern TTSsystems that can exceed human speech quality. More specifically, we conduct acomprehensive assessment of the MUSHRA test, focusing on its sensitivity tofactors such as rater variability, listener fatigue, and reference bias. Basedon our extensive evaluation involving 471 human listeners across Hindi andTamil we identify two primary shortcomings: (i) reference-matching bias, whereraters are unduly influenced by the human reference, and (ii) judgementambiguity, arising from a lack of clear fine-grained guidelines. To addressthese issues, we propose two refined variants of the MUSHRA test. The firstvariant enables fairer ratings for synthesized samples that surpass humanreference quality. The second variant reduces ambiguity, as indicated by therelatively lower variance across raters. By combining these approaches, weachieve both more reliable and more fine-grained assessments. We also releaseMANGO, a massive dataset of 47,100 human ratings, the first-of-its-kindcollection for Indian languages, aiding in analyzing human preferences anddeveloping automatic metrics for evaluating TTS systems.
标题:提高预训练文本到音乐生成模型的可控性和可编辑性
链接:https://arxiv.org/abs/2411.12641
备注:PhD Thesis
摘要:人工智能辅助音乐创作领域已经取得了重大进展,但现有系统往往难以满足迭代和细致入微的音乐制作需求。这些挑战包括对生成的内容提供足够的控制,并允许灵活,精确的编辑。本文通过引入一系列相互促进的进步来解决这些问题,增强了文本到音乐生成模型的可控性和可编辑性。 首先,我们介绍Loop Copilot,这是一个试图解决音乐创作中迭代优化需求的系统。Loop Copilot利用大型语言模型(LLM)来协调多个专业的AI模型,使用户能够通过对话界面交互式地生成和优化音乐。该系统的核心是全局属性表,它记录并维护整个迭代过程中的关键音乐属性,确保任何阶段的修改都能保持音乐的整体连贯性。虽然Loop Copilot在编排音乐创作过程方面表现出色,但它并不能直接解决对生成内容进行详细编辑的需求。 为了克服这一限制,MusicMagus被作为编辑AI生成的音乐的进一步解决方案。MusicMagus引入了一种zero-shot文本到音乐编辑方法,允许修改特定的音乐属性,如流派,情绪和乐器,而无需重新训练。通过在预先训练的扩散模型中操纵潜在空间,MusicMagus确保这些编辑在风格上是一致的,并且非目标属性保持不变。该系统在编辑期间保持音乐的结构完整性方面特别有效,但它在更复杂和真实的音频场景中遇到了挑战。 ...
摘要:The field of AI-assisted music creation has made significant strides, yetexisting systems often struggle to meet the demands of iterative and nuancedmusic production. These challenges include providing sufficient control overthe generated content and allowing for flexible, precise edits. This thesistackles these issues by introducing a series of advancements that progressivelybuild upon each other, enhancing the controllability and editability oftext-to-music generation models. First, we introduce Loop Copilot, a system that tries to address the need foriterative refinement in music creation. Loop Copilot leverages a large languagemodel (LLM) to coordinate multiple specialised AI models, enabling users togenerate and refine music interactively through a conversational interface.Central to this system is the Global Attribute Table, which records andmaintains key musical attributes throughout the iterative process, ensuringthat modifications at any stage preserve the overall coherence of the music.While Loop Copilot excels in orchestrating the music creation process, it doesnot directly address the need for detailed edits to the generated content. To overcome this limitation, MusicMagus is presented as a further solutionfor editing AI-generated music. MusicMagus introduces a zero-shot text-to-musicediting approach that allows for the modification of specific musicalattributes, such as genre, mood, and instrumentation, without the need forretraining. By manipulating the latent space within pre-trained diffusionmodels, MusicMagus ensures that these edits are stylistically coherent and thatnon-targeted attributes remain unchanged. This system is particularly effectivein maintaining the structural integrity of the music during edits, but itencounters challenges with more complex and real-world audio scenarios. ...
标题:DGSBA:基于预算的动态生成基于场景的噪音添加方法
链接:https://arxiv.org/abs/2411.12363
摘要:本文讨论了准确列举和描述场景的挑战,以及使用非生成方法复制声学环境所需的劳动密集型过程。本文提出了基于场景的动态生成场景加噪方法(DGSNA),该方法创新性地将场景信息动态生成(DGSI)和基于场景的音频加噪(SNAA)相结合。DGSI组件采用在背景示例任务(BET)提示框架内构造的生成式聊天模型,促进了针对特定声学环境定制的场景信息(SI)的动态合成。此外,SNAA组件利用房间脉冲响应(RIR)滤波器和文本到音频(TTA)系统来生成逼真的、基于场景的噪声,可以适应室内和室外环境。通过全面的实验,证明了DGSNA在不同生成式聊天模型中的适应性。通过客观和主观评估的结果表明,DGSNA提供了强大的性能,动态生成精确的SI和有效地提高基于场景的噪声添加能力,从而提供显着的改进,在声学场景模拟的传统方法。我们的实施和演示可在https://dgsna.github.io上获得。
摘要:This paper addresses the challenges of accurately enumerating and describingscenes and the labor-intensive process required to replicate acousticenvironments using non-generative methods. We introduce the prompt-basedDynamic Generative Sce-ne-based Noise Addition method (DGSNA), whichinnovatively combines the Dynamic Generation of Scene Information (DGSI) withScene-based Noise Addition for Audio (SNAA). Employing generative chat modelsstructured within the Back-ground-Examples-Task (BET) prompt framework, DGSIcom-ponent facilitates the dynamic synthesis of tailored Scene Infor-mation(SI) for specific acoustic environments. Additionally, the SNAA componentleverages Room Impulse Response (RIR) fil-ters and Text-To-Audio (TTA) systemsto generate realistic, scene-based noise that can be adapted for both indoorand out-door environments. Through comprehensive experiments, the adaptabilityof DGSNA across different generative chat models was demonstrated. The results,assessed through both objective and subjective evaluations, show that DGSNAprovides robust performance in dynamically generating precise SI andeffectively enhancing scene-based noise addition capabilities, thus offeringsignificant improvements over traditional methods in acoustic scene simulation.Our implementation and demos are available at https://dgsna.github.io.
标题:从音乐发现对话中预测用户意图和音乐属性
链接:https://arxiv.org/abs/2411.12254
备注:8 pages, 4 figures
摘要:意图分类是一个文本理解任务,它从输入文本查询中识别用户需求。虽然意图分类已经在各个领域得到了广泛的研究,但它在音乐领域并没有得到太多的关注。在本文中,我们研究了音乐发现会话的意图分类模型,重点是预训练的语言模型。我们不仅预测功能需求:意图分类,还包括对音乐需求进行分类的任务:音乐属性分类。此外,我们提出了一种方法,将以前的聊天历史与输入文本中的单轮用户查询连接起来,使模型能够更好地理解整个对话上下文。我们提出的模型显著提高了用户意图和音乐属性分类的F1得分,并超过了预训练的Llama 3模型的zero-shot和Few-Shot性能。
摘要:Intent classification is a text understanding task that identifies user needsfrom input text queries. While intent classification has been extensivelystudied in various domains, it has not received much attention in the musicdomain. In this paper, we investigate intent classification models for musicdiscovery conversation, focusing on pre-trained language models. Rather thanonly predicting functional needs: intent classification, we also include a taskfor classifying musical needs: musical attribute classification. Additionally,we propose a method of concatenating previous chat history with justsingle-turn user queries in the input text, allowing the model to understandthe overall conversation context better. Our proposed model significantlyimproves the F1 score for both user intent and musical attributeclassification, and surpasses the zero-shot and few-shot performance of thepretrained Llama 3 model.
标题:Zero-Shot板条挖掘:使用语音活动、音乐结构和CLAP嵌入的DJ工具检索
链接:https://arxiv.org/abs/2411.12209
摘要:在Hip-Hop、RnB、Reggae、Dancehall和几乎每一种电子/舞蹈/俱乐部风格中,DJ工具是一套特殊的音频文件,旨在提高DJ的音乐表现和创造性的混音选择。在这项工作中,我们展示了一种在个人音乐收藏中发现DJ工具的方法。利用开放源代码库的语音/音乐活动,音乐边界分析和对比存储音频预训练(CLAP)模型的zero-shot音频分类,我们展示了一种新颖的系统,旨在检索(或重新发现)引人注目的DJ工具,用于现场或在演播室。
摘要:In genres like Hip-Hop, RnB, Reggae, Dancehall and just about everyElectronic/Dance/Club style, DJ tools are a special set of audio files curatedto heighten the DJ's musical performance and creative mixing choices. In thiswork we demonstrate an approach to discovering DJ tools in personal musiccollections. Leveraging open-source libraries for speech/music activity, musicboundary analysis and a Contrastive Language-Audio Pretraining (CLAP) model forzero-shot audio classification, we demonstrate a novel system designed toretrieve (or rediscover) compelling DJ tools for use live or in the studio.
标题:视觉语言模型是Few-Shot音频谱图分类器
链接:https://arxiv.org/abs/2411.12058
摘要:我们证明了视觉语言模型(VLM)能够识别音频记录中的内容时,给出相应的频谱图图像。具体地,我们指示VLM在Few-Shot设置中执行音频分类任务,这是通过提示它们对每个类别的给定示例频谱图图像的频谱图图像进行分类来实现的。通过仔细设计的声谱图图像表示和选择良好的Few-Shot的例子,我们表明,GPT-4 o可以实现ESC-10环境声音分类数据集上的交叉验证的准确率为59.00%。此外,我们证明了VLM目前在等效音频分类任务上的表现优于唯一具有音频理解能力的商业音频语言模型(Gemini-1.5)(59.00% vs. 49.62%),甚至在视觉频谱图分类上的表现略好于人类专家(第一次折叠时为73.75% vs. 72.50%)。我们为这些发现设想了两个潜在的用例:(1)将VLM的声谱图和语言理解能力结合起来用于音频字幕增强,以及(2)将视觉声谱图分类作为VLM的挑战任务。
摘要:We demonstrate that vision language models (VLMs) are capable of recognizingthe content in audio recordings when given corresponding spectrogram images.Specifically, we instruct VLMs to perform audio classification tasks in afew-shot setting by prompting them to classify a spectrogram image givenexample spectrogram images of each class. By carefully designing thespectrogram image representation and selecting good few-shot examples, we showthat GPT-4o can achieve 59.00% cross-validated accuracy on the ESC-10environmental sound classification dataset. Moreover, we demonstrate that VLMscurrently outperform the only available commercial audio language model withaudio understanding capabilities (Gemini-1.5) on the equivalent audioclassification task (59.00% vs. 49.62%), and even perform slightly better thanhuman experts on visual spectrogram classification (73.75% vs. 72.50% on firstfold). We envision two potential use cases for these findings: (1) combiningthe spectrogram and language understanding capabilities of VLMs for audiocaption augmentation, and (2) posing visual spectrogram classification as achallenge task for VLMs.
标题:用多通道RVQGAN压缩高级立体声
链接:https://arxiv.org/abs/2411.12008
摘要:提出了一种对RVQGAN神经编码方法的多通道扩展,并实现了三阶高保真度立体声音频的数据驱动压缩。修改生成器模型的输入层和输出层,以便在不增加模型比特率的情况下接受多(16)个通道。我们还提出了一个损失函数,用于解释沉浸式再现中的空间感知,以及从单通道模型中迁移学习。7.1.4沉浸式播放的听力测试结果表明,所提出的扩展适用于编码基于场景的16声道高保真度立体声内容,在16 kbit/s的质量良好。
摘要:A multichannel extension to the RVQGAN neural coding method is proposed, andrealized for data-driven compression of third-order Ambisonics audio. Theinput- and output layers of the generator and discriminator models are modifiedto accept multiple (16) channels without increasing the model bitrate. We alsopropose a loss function for accounting for spatial perception in immersivereproduction, and transfer learning from single-channel models. Listening testresults with 7.1.4 immersive playback show that the proposed extension issuitable for coding scene-based, 16-channel Ambisonics content with goodquality at 16 kbit/s.
标题:基于幅度和相预测从有噪的Mel谱图中生成干净波形的神经去噪声码器
链接:https://arxiv.org/abs/2411.12268
备注:Accepted by NCMMSC2024
摘要:本文提出了一种新的神经去噪声码器,可以产生干净的语音波形从噪声梅尔频谱图。所提出的神经去噪声码器由两个组件组成,即,频谱预测器和增强模块。频谱预测器首先从输入的有噪梅尔频谱图预测出有噪的幅度谱和相位谱,然后增强模块从有噪的幅度谱和相位谱中恢复出干净的幅度谱和相位谱。最后,通过逆短时傅立叶变换(iSTFT)重建干净的语音波形。所有操作都在帧级谱域执行,APNet声码器和MP-SENet语音增强模型分别用作两个分量的主干。实验结果表明,我们提出的神经去噪声码器实现了最先进的性能相比,现有的神经声码器的VoiceBank+DEMAND数据集。此外,尽管缺乏相位信息和部分幅度信息的输入梅尔频谱图,所提出的神经去噪声码器仍然实现了与几个先进的语音增强方法相当的性能。
摘要:This paper proposes a novel neural denoising vocoder that can generate cleanspeech waveforms from noisy mel-spectrograms. The proposed neural denoisingvocoder consists of two components, i.e., a spectrum predictor and aenhancement module. The spectrum predictor first predicts the noisy amplitudeand phase spectra from the input noisy mel-spectrogram, and subsequently theenhancement module recovers the clean amplitude and phase spectrum from noisyones. Finally, clean speech waveforms are reconstructed through inverseshort-time Fourier transform (iSTFT). All operations are performed at theframe-level spectral domain, with the APNet vocoder and MP-SENet speechenhancement model used as the backbones for the two components, respectively.Experimental results demonstrate that our proposed neural denoising vocoderachieves state-of-the-art performance compared to existing neural vocoders onthe VoiceBank+DEMAND dataset. Additionally, despite the lack of phaseinformation and partial amplitude information in the input mel-spectrogram, theproposed neural denoising vocoder still achieves comparable performance withthe serveral advanced speech enhancement methods.
标题:重新思考MUHRA:应对文本到语音评估的现代挑战
链接:https://arxiv.org/abs/2411.12719
备注:19 pages, 12 Figures
摘要:尽管TTS模型发展迅速,但仍然缺乏一致和强大的人类评估框架。例如,MOS测试无法区分相似的模型,CMOS的成对比较是时间密集型的。MUSHRA测试是同时评估多个TTS系统的一个有前途的替代方案,但在这项工作中,我们表明,它对匹配人类参考语音的依赖过度地损害了现代TTS系统的得分,这些系统的语音质量可能超过人类语音质量。更具体地说,我们对MUSHRA测试进行了全面的评估,重点关注其对评分员变异性,听众疲劳和参考偏差等因素的敏感性。基于我们对471名印地语和泰米尔语的人类听众的广泛评估,我们发现了两个主要缺点:(i)参考匹配偏差,其中评分员受到人类参考的过度影响,以及(ii)判断模糊性,由于缺乏明确的细粒度指导方针。为了解决这些问题,我们提出了MUSHRA测试的两个改进版本。第一种变体能够对超过人类参考质量的合成样本进行更公平的评级。第二种变体减少了模糊性,正如评分者之间相对较低的方差所表明的那样。通过结合这些方法,我们实现了更可靠和更细粒度的评估。我们还发布了MANGO,这是一个包含47,100个人类评分的大规模数据集,是印度语言的首个同类集合,有助于分析人类偏好并开发用于评估TTS系统的自动指标。
摘要:Despite rapid advancements in TTS models, a consistent and robust humanevaluation framework is still lacking. For example, MOS tests fail todifferentiate between similar models, and CMOS's pairwise comparisons aretime-intensive. The MUSHRA test is a promising alternative for evaluatingmultiple TTS systems simultaneously, but in this work we show that its relianceon matching human reference speech unduly penalises the scores of modern TTSsystems that can exceed human speech quality. More specifically, we conduct acomprehensive assessment of the MUSHRA test, focusing on its sensitivity tofactors such as rater variability, listener fatigue, and reference bias. Basedon our extensive evaluation involving 471 human listeners across Hindi andTamil we identify two primary shortcomings: (i) reference-matching bias, whereraters are unduly influenced by the human reference, and (ii) judgementambiguity, arising from a lack of clear fine-grained guidelines. To addressthese issues, we propose two refined variants of the MUSHRA test. The firstvariant enables fairer ratings for synthesized samples that surpass humanreference quality. The second variant reduces ambiguity, as indicated by therelatively lower variance across raters. By combining these approaches, weachieve both more reliable and more fine-grained assessments. We also releaseMANGO, a massive dataset of 47,100 human ratings, the first-of-its-kindcollection for Indian languages, aiding in analyzing human preferences anddeveloping automatic metrics for evaluating TTS systems.
标题:提高预训练文本到音乐生成模型的可控性和可编辑性
链接:https://arxiv.org/abs/2411.12641
备注:PhD Thesis
摘要:人工智能辅助音乐创作领域已经取得了重大进展,但现有系统往往难以满足迭代和细致入微的音乐制作需求。这些挑战包括对生成的内容提供足够的控制,并允许灵活,精确的编辑。本文通过引入一系列相互促进的进步来解决这些问题,增强了文本到音乐生成模型的可控性和可编辑性。 首先,我们介绍Loop Copilot,这是一个试图解决音乐创作中迭代优化需求的系统。Loop Copilot利用大型语言模型(LLM)来协调多个专业的AI模型,使用户能够通过对话界面交互式地生成和优化音乐。该系统的核心是全局属性表,它记录并维护整个迭代过程中的关键音乐属性,确保任何阶段的修改都能保持音乐的整体连贯性。虽然Loop Copilot在编排音乐创作过程方面表现出色,但它并不能直接解决对生成内容进行详细编辑的需求。 为了克服这一限制,MusicMagus被作为编辑AI生成的音乐的进一步解决方案。MusicMagus引入了一种zero-shot文本到音乐编辑方法,允许修改特定的音乐属性,如流派,情绪和乐器,而无需重新训练。通过在预先训练的扩散模型中操纵潜在空间,MusicMagus确保这些编辑在风格上是一致的,并且非目标属性保持不变。该系统在编辑期间保持音乐的结构完整性方面特别有效,但它在更复杂和真实的音频场景中遇到了挑战。 ...
摘要:The field of AI-assisted music creation has made significant strides, yetexisting systems often struggle to meet the demands of iterative and nuancedmusic production. These challenges include providing sufficient control overthe generated content and allowing for flexible, precise edits. This thesistackles these issues by introducing a series of advancements that progressivelybuild upon each other, enhancing the controllability and editability oftext-to-music generation models. First, we introduce Loop Copilot, a system that tries to address the need foriterative refinement in music creation. Loop Copilot leverages a large languagemodel (LLM) to coordinate multiple specialised AI models, enabling users togenerate and refine music interactively through a conversational interface.Central to this system is the Global Attribute Table, which records andmaintains key musical attributes throughout the iterative process, ensuringthat modifications at any stage preserve the overall coherence of the music.While Loop Copilot excels in orchestrating the music creation process, it doesnot directly address the need for detailed edits to the generated content. To overcome this limitation, MusicMagus is presented as a further solutionfor editing AI-generated music. MusicMagus introduces a zero-shot text-to-musicediting approach that allows for the modification of specific musicalattributes, such as genre, mood, and instrumentation, without the need forretraining. By manipulating the latent space within pre-trained diffusionmodels, MusicMagus ensures that these edits are stylistically coherent and thatnon-targeted attributes remain unchanged. This system is particularly effectivein maintaining the structural integrity of the music during edits, but itencounters challenges with more complex and real-world audio scenarios. ...
标题:DGSBA:基于预算的动态生成基于场景的噪音添加方法
链接:https://arxiv.org/abs/2411.12363
摘要:本文讨论了准确列举和描述场景的挑战,以及使用非生成方法复制声学环境所需的劳动密集型过程。本文提出了基于场景的动态生成场景加噪方法(DGSNA),该方法创新性地将场景信息动态生成(DGSI)和基于场景的音频加噪(SNAA)相结合。DGSI组件采用在背景示例任务(BET)提示框架内构造的生成式聊天模型,促进了针对特定声学环境定制的场景信息(SI)的动态合成。此外,SNAA组件利用房间脉冲响应(RIR)滤波器和文本到音频(TTA)系统来生成逼真的、基于场景的噪声,可以适应室内和室外环境。通过综合实验,验证了DGSNA在不同生成式聊天模型间的适应性.通过客观和主观评估的结果表明,DGSNA提供了强大的性能,动态生成精确的SI和有效地提高基于场景的噪声添加能力,从而提供显着的改进,在声学场景模拟的传统方法。我们的实现和演示可以在https://dgsna.github.io上找到。
摘要:This paper addresses the challenges of accurately enumerating and describingscenes and the labor-intensive process required to replicate acousticenvironments using non-generative methods. We introduce the prompt-basedDynamic Generative Sce-ne-based Noise Addition method (DGSNA), whichinnovatively combines the Dynamic Generation of Scene Information (DGSI) withScene-based Noise Addition for Audio (SNAA). Employing generative chat modelsstructured within the Back-ground-Examples-Task (BET) prompt framework, DGSIcom-ponent facilitates the dynamic synthesis of tailored Scene Infor-mation(SI) for specific acoustic environments. Additionally, the SNAA componentleverages Room Impulse Response (RIR) fil-ters and Text-To-Audio (TTA) systemsto generate realistic, scene-based noise that can be adapted for both indoorand out-door environments. Through comprehensive experiments, the adaptabilityof DGSNA across different generative chat models was demonstrated. The results,assessed through both objective and subjective evaluations, show that DGSNAprovides robust performance in dynamically generating precise SI andeffectively enhancing scene-based noise addition capabilities, thus offeringsignificant improvements over traditional methods in acoustic scene simulation.Our implementation and demos are available at https://dgsna.github.io.
标题:从音乐发现对话中预测用户意图和音乐属性
链接:https://arxiv.org/abs/2411.12254
备注:8 pages, 4 figures
摘要:意图分类是一个文本理解任务,它从输入文本查询中识别用户需求。虽然意图分类已经在各个领域得到了广泛的研究,但它在音乐领域并没有得到太多的关注。在本文中,我们研究了音乐发现会话的意图分类模型,重点是预训练的语言模型。我们不仅预测功能需求:意图分类,还包括对音乐需求进行分类的任务:音乐属性分类。此外,我们提出了一种方法,将以前的聊天历史与输入文本中的单轮用户查询连接起来,使模型能够更好地理解整个对话上下文。我们提出的模型显著提高了用户意图和音乐属性分类的F1得分,并超过了预训练的Llama 3模型的zero-shot和Few-Shot性能。
摘要:Intent classification is a text understanding task that identifies user needsfrom input text queries. While intent classification has been extensivelystudied in various domains, it has not received much attention in the musicdomain. In this paper, we investigate intent classification models for musicdiscovery conversation, focusing on pre-trained language models. Rather thanonly predicting functional needs: intent classification, we also include a taskfor classifying musical needs: musical attribute classification. Additionally,we propose a method of concatenating previous chat history with justsingle-turn user queries in the input text, allowing the model to understandthe overall conversation context better. Our proposed model significantlyimproves the F1 score for both user intent and musical attributeclassification, and surpasses the zero-shot and few-shot performance of thepretrained Llama 3 model.
标题:Zero-Shot板条挖掘:使用语音活动、音乐结构和CLAP嵌入的DJ工具检索
链接:https://arxiv.org/abs/2411.12209
摘要:在Hip-Hop、RnB、Reggae、Dancehall和几乎每一种电子/舞蹈/俱乐部风格中,DJ工具是一套特殊的音频文件,旨在提高DJ的音乐表现和创造性的混音选择。在这项工作中,我们展示了一种方法来发现个人音乐收藏中的DJ工具。利用开放源代码库的语音/音乐活动,音乐边界分析和对比存储音频预训练(CLAP)模型的zero-shot音频分类,我们展示了一种新颖的系统,旨在检索(或重新发现)引人注目的DJ工具,用于现场或在演播室。
摘要:In genres like Hip-Hop, RnB, Reggae, Dancehall and just about everyElectronic/Dance/Club style, DJ tools are a special set of audio files curatedto heighten the DJ's musical performance and creative mixing choices. In thiswork we demonstrate an approach to discovering DJ tools in personal musiccollections. Leveraging open-source libraries for speech/music activity, musicboundary analysis and a Contrastive Language-Audio Pretraining (CLAP) model forzero-shot audio classification, we demonstrate a novel system designed toretrieve (or rediscover) compelling DJ tools for use live or in the studio.
标题:视觉语言模型是Few-Shot音频谱图分类器
链接:https://arxiv.org/abs/2411.12058
摘要:我们证明了视觉语言模型(VLM)能够识别音频记录中的内容时,给出相应的频谱图图像。具体地,我们指示VLM在Few-Shot设置中执行音频分类任务,这是通过提示它们对每个类别的给定示例频谱图图像的频谱图图像进行分类来实现的。通过仔细设计的声谱图图像表示和选择良好的Few-Shot的例子,我们表明,GPT-4 o可以实现ESC-10环境声音分类数据集上的交叉验证的准确率为59.00%。此外,我们证明了VLM目前在等效音频分类任务上的表现优于唯一具有音频理解能力的商业音频语言模型(Gemini-1.5)(59.00% vs. 49.62%),甚至在视觉频谱图分类上的表现略好于人类专家(第一次折叠时为73.75% vs. 72.50%)。我们为这些发现设想了两个潜在的用例:(1)将VLM的声谱图和语言理解能力结合起来用于音频字幕增强,以及(2)将视觉声谱图分类作为VLM的挑战任务。
摘要:We demonstrate that vision language models (VLMs) are capable of recognizingthe content in audio recordings when given corresponding spectrogram images.Specifically, we instruct VLMs to perform audio classification tasks in afew-shot setting by prompting them to classify a spectrogram image givenexample spectrogram images of each class. By carefully designing thespectrogram image representation and selecting good few-shot examples, we showthat GPT-4o can achieve 59.00% cross-validated accuracy on the ESC-10environmental sound classification dataset. Moreover, we demonstrate that VLMscurrently outperform the only available commercial audio language model withaudio understanding capabilities (Gemini-1.5) on the equivalent audioclassification task (59.00% vs. 49.62%), and even perform slightly better thanhuman experts on visual spectrogram classification (73.75% vs. 72.50% on firstfold). We envision two potential use cases for these findings: (1) combiningthe spectrogram and language understanding capabilities of VLMs for audiocaption augmentation, and (2) posing visual spectrogram classification as achallenge task for VLMs.
标题:用多通道RVQGAN压缩高级立体声
链接:https://arxiv.org/abs/2411.12008
摘要:提出了一种对RVQGAN神经编码方法的多通道扩展,并实现了三阶高保真度立体声音频的数据驱动压缩。修改生成器模型的输入层和输出层,以便在不增加模型比特率的情况下接受多(16)个通道。我们还提出了一个损失函数,用于解释沉浸式再现中的空间感知,以及从单通道模型中迁移学习。7.1.4沉浸式播放的听力测试结果表明,所提出的扩展适用于编码基于场景的16声道高保真度立体声内容,在16 kbit/s的质量良好。
摘要:A multichannel extension to the RVQGAN neural coding method is proposed, andrealized for data-driven compression of third-order Ambisonics audio. Theinput- and output layers of the generator and discriminator models are modifiedto accept multiple (16) channels without increasing the model bitrate. We alsopropose a loss function for accounting for spatial perception in immersivereproduction, and transfer learning from single-channel models. Listening testresults with 7.1.4 immersive playback show that the proposed extension issuitable for coding scene-based, 16-channel Ambisonics content with goodquality at 16 kbit/s.
