【1】Advancing Audio Fingerprinting Accuracy Addressing Background Noise and Distortion Challenges标题:提高音频指纹识别精度解决背景噪声和失真挑战链接:https://arxiv.org/abs/2402.13957作者:Navin Kamuni,Sathishkumar Chintala,Naveen Kunchakuri,Jyothi Swaroop Arlagadda Narasimharaju,Venkat Kumar摘要:以Shazam等先驱为例的音频指纹识别已经改变了数字音频识别。然而,现有的系统在具有挑战性的条件下难以准确,限制了广泛的适用性。本研究提出一种AI和ML整合的音频指纹算法,以提高准确性。该研究建立在Dejavu项目的基础上,强调具有不同背景噪音和失真的真实场景模拟。信号处理是Dejavu模型的核心,包括快速傅立叶变换、频谱图和峰值提取。“星座”概念和指纹散列实现了独特的歌曲识别。性能评估证明,在5秒的音频输入内,100%的准确性,系统展示了可预测的匹配速度,以提高效率。存储分析强调了实际实施的关键空间-速度权衡。这项研究提高了音频指纹的适应性,解决了各种环境和应用中的挑战。摘要:Audio fingerprinting, exemplified by pioneers like Shazam, has transformed digital audio recognition. However, existing systems struggle with accuracy in challenging conditions, limiting broad applicability. This research proposes an AI and ML integrated audio fingerprinting algorithm to enhance accuracy. Built on the Dejavu Project's foundations, the study emphasizes real-world scenario simulations with diverse background noises and distortions. Signal processing, central to Dejavu's model, includes the Fast Fourier Transform, spectrograms, and peak extraction. The "constellation" concept and fingerprint hashing enable unique song identification. Performance evaluation attests to 100% accuracy within a 5-second audio input, with a system showcasing predictable matching speed for efficiency. Storage analysis highlights the critical space-speed trade-off for practical implementation. This research advances audio fingerprinting's adaptability, addressing challenges in varied environments and applications.
【2】 Voice-Driven Mortality Prediction in Hospitalized Heart Failure Patients: A Machine Learning Approach Enhanced with Diagnostic Biomarkers标题:心力衰竭住院患者的语音驱动死亡率预测:一种通过诊断生物标志物增强的机器学习方法链接:https://arxiv.org/abs/2402.13812作者:Nihat Ahmadli,Mehmet Ali Sarsil,Berk Mizrak,Kurtulus Karauzum,Ata Shaker,Erol Tulumen,Didar Mirzamidinov,Dilek Ural,Onur Ergen备注:11 pages, 6 figures, 5 tables. THe first 2 authors have contributed equally摘要:解决心力衰竭(HF)作为一个普遍的全球健康问题带来了困难,实施创新的方法,以加强病人的护理。特别是预测HF患者的死亡率是困难的,但也是关键的,需要个性化护理,积极主动的管理,并使教育决策,以提高结果。最近,声音生物标志物与机器学习(ML)相结合的重要性激增,表现出显着的功效,特别是在预测心力衰竭方面。语音分析和ML算法的协同作用提供了一种非侵入性且易于访问的方法来评估患者的健康状况。然而,缺乏语音生物标志物来预测标准化语音协议的心力衰竭患者的死亡率。在这里,我们展示了一个强大而有效的ML模型,通过利用语音生物标志物来预测住院HF患者的死亡率。通过将语音生物标志物无缝集成到常规患者监测中,该策略有可能改善患者结局,优化资源分配,并推进以患者为中心的HF管理。在这项研究中,一个机器学习系统,特别是一个逻辑回归模型,被训练来预测患者的5年死亡率,使用他们的语音作为输入。交叉验证和统计方法证明,该模型的表现令人钦佩且一致(p值< 0.001)。此外,整合NT-proBNP,HF中的诊断生物标志物,大大提高了模型的预测准确性。摘要:Addressing heart failure (HF) as a prevalent global health concern poses difficulties in implementing innovative approaches for enhanced patient care. Predicting mortality rates in HF patients, in particular, is difficult yet critical, necessitating individualized care, proactive management, and enabling educated decision-making to enhance outcomes. Recently, the significance of voice biomarkers coupled with Machine Learning (ML) has surged, demonstrating remarkable efficacy, particularly in predicting heart failure. The synergy of voice analysis and ML algorithms provides a non-invasive and easily accessible means to evaluate patients' health. However, there is a lack of voice biomarkers for predicting mortality rates among heart failure patients with standardized speech protocols. Here, we demonstrate a powerful and effective ML model for predicting mortality rates in hospitalized HF patients through the utilization of voice biomarkers. By seamlessly integrating voice biomarkers into routine patient monitoring, this strategy has the potential to improve patient outcomes, optimize resource allocation, and advance patient-centered HF management. In this study, a Machine Learning system, specifically a logistic regression model, is trained to predict patients' 5-year mortality rates using their speech as input. The model performs admirably and consistently, as demonstrated by cross-validation and statistical approaches (p-value < 0.001). Furthermore, integrating NT-proBNP, a diagnostic biomarker in HF, improves the model's predictive accuracy substantially. 【3】 Music Style Transfer with Time-Varying Inversion of Diffusion Models标题:基于扩散模型时变逆的音乐风格转换链接:https://arxiv.org/abs/2402.13763作者:Sifei Li,Yuxin Zhang,Fan Tang,Chongyang Ma,Weiming dong,Changsheng Xu备注:7 pages, 4 figures, AAAI 2024摘要:随着扩散模型的发展,文本引导的图像风格转换已经显示出高质量的可控合成结果。然而,利用文本进行不同的音乐风格转移带来了重大挑战,主要是由于匹配的音频文本数据集的可用性有限。音乐是一种抽象而复杂的艺术形式,即使在同一流派中也会表现出变化和复杂性,从而使准确的文本描述具有挑战性。本文提出了一种音乐风格转移的方法,有效地捕捉音乐属性,使用最少的数据。我们引入了一种新的时变文本反转模块,以精确地捕获不同级别的梅尔频谱图特征。在推理过程中,我们提出了一个减少偏见的风格化技术,以获得稳定的结果。实验结果表明,我们的方法可以转移特定乐器的风格,以及纳入自然的声音组成的旋律。示例和源代码可在https://lsfhuihuiff.github.io/MusicTI/上获得。摘要:With the development of diffusion models, text-guided image style transfer has demonstrated high-quality controllable synthesis results. However, the utilization of text for diverse music style transfer poses significant challenges, primarily due to the limited availability of matched audio-text datasets. Music, being an abstract and complex art form, exhibits variations and intricacies even within the same genre, thereby making accurate textual descriptions challenging. This paper presents a music style transfer approach that effectively captures musical attributes using minimal data. We introduce a novel time-varying textual inversion module to precisely capture mel-spectrogram features at different levels. During inference, we propose a bias-reduced stylization technique to obtain stable results. Experimental results demonstrate that our method can transfer the style of specific instruments, as well as incorporate natural sounds to compose melodies. Samples and source code are available at https://lsfhuihuiff.github.io/MusicTI/.
【4】 The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning标题:批量大小对对比自监督语音表征学习的影响链接:https://arxiv.org/abs/2402.13723作者:Nik Vaessen,David A. van Leeuwen摘要:语音中的基础模型通常使用许多GPU进行训练,这隐含地导致了大的有效批量。在本文中,我们研究了批量大小对预训练的影响,包括在训练过程中可以监控的统计数据,以及对下游微调任务性能的影响。通过使用从87.5秒到80分钟不等的批量大小,我们表明,对于固定的迭代次数,较大的批量大小会产生更好的预训练模型。然而,稳定性有下限,有效性有上限。然后,我们表明,预训练模型的质量主要取决于训练期间看到的语音数据量,即,批量大小和迭代次数的乘积。所有结果都是通过wav2vec 2.0架构的独立实现产生的,该架构在很大程度上再现了原始工作的结果(arXiv:2006.11477)。我们的扩展可以帮助研究人员在研究语音中的自监督学习时选择有效的操作条件,并提示使用固定数量的可见数据对自监督进行基准测试。代码和模型检查点可以在https://github.com/nikvaessen/w2v2-batch-size上找到。摘要:Foundation models in speech are often trained using many GPUs, which implicitly leads to large effective batch sizes. In this paper we study the effect of batch size on pre-training, both in terms of statistics that can be monitored during training, and in the effect on the performance of a downstream fine-tuning task. By using batch sizes varying from 87.5 seconds to 80 minutes of speech we show that, for a fixed amount of iterations, larger batch sizes result in better pre-trained models. However, there is lower limit for stability, and an upper limit for effectiveness. We then show that the quality of the pre-trained model depends mainly on the amount of speech data seen during training, i.e., on the product of batch size and number of iterations. All results are produced with an independent implementation of the wav2vec 2.0 architecture, which to a large extent reproduces the results of the original work (arXiv:2006.11477). Our extensions can help researchers choose effective operating conditions when studying self-supervised learning in speech, and hints towards benchmarking self-supervision with a fixed amount of seen data. Code and model checkpoints are available at https://github.com/nikvaessen/w2v2-batch-size.
【5】 Structure-informed Positional Encoding for Music Generation标题:用于音乐生成的结构信息位置编码链接:https://arxiv.org/abs/2402.13301作者:Manvi Agarwal,Changhong Wang,Gaël Richard备注:None摘要:深度学习方法生成的音乐通常缺乏连贯性和长期组织。然而,多尺度层次结构是音乐信号的显著特征。为了利用这些信息,我们提出了一个结构知情的位置编码框架与Transformers的音乐生成。我们设计了三个变种的绝对,相对和非静止的位置信息。我们全面测试他们的两个象征性的音乐生成任务:下一个时间步预测和伴奏生成。作为比较,我们从文献中选择多个基线,并使用几个音乐动机的评价指标来证明我们的方法的优点。特别是,我们的方法提高了生成的作品的旋律和结构的一致性。摘要:Music generated by deep learning methods often suffers from a lack of coherence and long-term organization. Yet, multi-scale hierarchical structure is a distinctive feature of music signals. To leverage this information, we propose a structure-informed positional encoding framework for music generation with Transformers. We design three variants in terms of absolute, relative and non-stationary positional information. We comprehensively test them on two symbolic music generation tasks: next-timestep prediction and accompaniment generation. As a comparison, we choose multiple baselines from the literature and demonstrate the merits of our methods using several musically-motivated evaluation metrics. In particular, our methods improve the melodic and structural consistency of the generated pieces. 【6】 When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection标题:当LLMS遇到声学标志时:一种有效的将语音集成到大型语言模型中进行抑郁检测的方法链接:https://arxiv.org/abs/2402.13276作者:Xiangyu Zhang,Hexin Liu,Kaishuai Xu,Qiquan Zhang,Daijiao Liu,Beena Ahmed,Julien Epps摘要:抑郁症是全球心理健康的一个重要问题,促使人们对基于人工智能的检测方法进行广泛研究。在各种人工智能技术中,大型语言模型(LLM)因其在心理健康应用中的多功能性而脱颖而出。然而,它们的主要局限性来自于它们对文本输入的排他性依赖,这限制了它们的整体能力。此外,利用LLM识别和分析抑郁状态仍然是相对未开发的。在本文中,我们提出了一种创新的方法,将声学语音信息集成到LLM框架多模态抑郁症检测。我们研究了一种有效的方法,抑郁症检测集成语音信号到LLM利用声学地标。通过结合声学地标,这是特定于口语单词的发音,我们的方法增加了关键尺寸的文本成绩单。这种整合还提供了对个人独特的言语模式的见解,揭示了个人的潜在心理状态。在DAIC-WOZ数据集上对所提出的方法进行的评估显示,与现有的音频文本基线相比,该方法具有最先进的结果。此外,这种方法不仅是有价值的抑郁症的检测,但也代表了一个新的视角,在提高LLM理解和处理语音信号的能力。摘要:Depression is a critical concern in global mental health, prompting extensive research into AI-based detection methods. Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in mental healthcare applications. However, their primary limitation arises from their exclusive dependence on textual input, which constrains their overall capabilities. Furthermore, the utilization of LLMs in identifying and analyzing depressive states is still relatively untapped. In this paper, we present an innovative approach to integrating acoustic speech information into the LLMs framework for multimodal depression detection. We investigate an efficient method for depression detection by integrating speech signals into LLMs utilizing Acoustic Landmarks. By incorporating acoustic landmarks, which are specific to the pronunciation of spoken words, our method adds critical dimensions to text transcripts. This integration also provides insights into the unique speech patterns of individuals, revealing the potential mental states of individuals. Evaluations of the proposed approach on the DAIC-WOZ dataset reveal state-of-the-art results when compared with existing Audio-Text baselines. In addition, this approach is not only valuable for the detection of depression but also represents a new perspective in enhancing the ability of LLMs to comprehend and process speech signals.
eess.AS音频处理【1】 HOMULA-RIR: A Room Impulse Response Dataset for Teleconferencing and Spatial Audio Applications Acquired Through Higher-Order Microphones and Uniform Linear Microphone Arrays标题:HOMULA-RIR:通过高阶麦克风和均匀线性麦克风阵列获取的用于电话会议和空间音频应用的房间脉冲响应数据集链接:https://arxiv.org/abs/2402.13896作者:Federico Miotello,Paolo Ostan,Mirco Pezzoli,Luca Comanducci,Alberto Bernardini,Fabio Antonacci,Augusto Sarti备注:Accepted for publication at ICASSP 2024 - HSCMA Workshop摘要:在本文中,我们提出了HOMULA-RIR,一个数据集的房间脉冲响应(RIR)获得使用高阶麦克风(HOM)和均匀线性阵列(OLS),为了模拟远程出席电话会议的情况。具体来说,测量是在一个研讨会室中进行的,其中64个麦克风的麦克风被用作扬声器附近的多通道音频采集系统,而HOM被用来模拟实际存在于研讨会室中的25名与会者。HOM覆盖了房间的大面积,使得数据集也适合虚拟声学的应用。通过混响时间和清晰度指数的测量,以及源定位和分离等示例应用,我们证明了HOMULA-RIR数据集的有效性。摘要:In this paper, we present HOMULA-RIR, a dataset of room impulse responses (RIRs) acquired using both higher-order microphones (HOMs) and a uniform linear array (ULA), in order to model a remote attendance teleconferencing scenario. Specifically, measurements were performed in a seminar room, where a 64-microphone ULA was used as a multichannel audio acquisition system in the proximity of the speakers, while HOMs were used to model 25 attendees actually present in the seminar room. The HOMs cover a wide area of the room, making the dataset suitable also for applications of virtual acoustics. Through the measurement of the reverberation time and clarity index, and sample applications such as source localization and separation, we demonstrate the effectiveness of the HOMULA-RIR dataset. 【2】 Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR标题:Mel-FullSubNet:同时提高语音质量和ASR的Mel谱图增强链接:https://arxiv.org/abs/2402.13511作者:Rui Zhou,Xian Li,Ying Fang,Xiaofei Li摘要:在这项工作中,我们提出了Mel-FullSubNet,这是一个单通道Mel频谱图去噪和去混响网络,用于提高语音质量和自动语音识别(ASR)性能。Mel-FullSubNet将噪声和混响的Mel频谱图作为输入,并预测相应的干净Mel频谱图。增强后的Mel谱图既可以用神经声码器转换成语音波形,也可以直接用于ASR。Mel-FullSubNet封装了交错的全频带和子频带网络,分别用于学习信号的全频带频谱模式和信号的子频带/窄带特性。与线性频域或时域语音增强相比,Mel谱图增强的主要优点是Mel频率以更紧凑的方式呈现语音,因此更容易学习,这将有利于语音质量和ASR。实验结果表明,该模型在语音质量和ASR性能方面都有显着的改善。摘要:In this work, we propose Mel-FullSubNet, a single-channel Mel-spectrogram denoising and dereverberation network for improving both speech quality and automatic speech recognition (ASR) performance. Mel-FullSubNet takes as input the noisy and reverberant Mel-spectrogram and predicts the corresponding clean Mel-spectrogram. The enhanced Mel-spectrogram can be either transformed to speech waveform with a neural vocoder or directly used for ASR. Mel-FullSubNet encapsulates interleaved full-band and sub-band networks, for learning the full-band spectral pattern of signals and the sub-band/narrow-band properties of signals, respectively. Compared to linear-frequency domain or time-domain speech enhancement, the major advantage of Mel-spectrogram enhancement is that Mel-frequency presents speech in a more compact way and thus is easier to learn, which will benefit both speech quality and ASR. Experimental results demonstrate a significant improvement in both speech quality and ASR performance achieved by the proposed model.
【3】 When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection标题:当LLMS遇到声学标志时:一种有效的将语音集成到大型语言模型中进行抑郁检测的方法链接:https://arxiv.org/abs/2402.13276作者:Xiangyu Zhang,Hexin Liu,Kaishuai Xu,Qiquan Zhang,Daijiao Liu,Beena Ahmed,Julien Epps摘要:抑郁症是全球心理健康的一个重要问题,促使人们对基于人工智能的检测方法进行广泛研究。在各种人工智能技术中,大型语言模型(LLM)因其在心理健康应用中的多功能性而脱颖而出。然而,它们的主要局限性来自于它们对文本输入的排他性依赖,这限制了它们的整体能力。此外,利用LLM识别和分析抑郁状态仍然是相对未开发的。在本文中,我们提出了一种创新的方法,将声学语音信息集成到LLM框架多模态抑郁症检测。我们研究了一种有效的方法,抑郁症检测集成语音信号到LLM利用声学地标。通过结合声学地标,这是特定于口语单词的发音,我们的方法增加了关键尺寸的文本成绩单。这种整合还提供了对个人独特的言语模式的见解,揭示了个人的潜在心理状态。在DAIC-WOZ数据集上对所提出的方法进行的评估显示,与现有的音频文本基线相比,该方法具有最先进的结果。此外,这种方法不仅是有价值的抑郁症的检测,但也代表了一个新的视角,在提高LLM理解和处理语音信号的能力。摘要:Depression is a critical concern in global mental health, prompting extensive research into AI-based detection methods. Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in mental healthcare applications. However, their primary limitation arises from their exclusive dependence on textual input, which constrains their overall capabilities. Furthermore, the utilization of LLMs in identifying and analyzing depressive states is still relatively untapped. In this paper, we present an innovative approach to integrating acoustic speech information into the LLMs framework for multimodal depression detection. We investigate an efficient method for depression detection by integrating speech signals into LLMs utilizing Acoustic Landmarks. By incorporating acoustic landmarks, which are specific to the pronunciation of spoken words, our method adds critical dimensions to text transcripts. This integration also provides insights into the unique speech patterns of individuals, revealing the potential mental states of individuals. Evaluations of the proposed approach on the DAIC-WOZ dataset reveal state-of-the-art results when compared with existing Audio-Text baselines. In addition, this approach is not only valuable for the detection of depression but also represents a new perspective in enhancing the ability of LLMs to comprehend and process speech signals. 【4】 Advancing Audio Fingerprinting Accuracy Addressing Background Noise and Distortion Challenges标题:提高音频指纹识别精度解决背景噪声和失真挑战链接:https://arxiv.org/abs/2402.13957作者:Navin Kamuni,Sathishkumar Chintala,Naveen Kunchakuri,Jyothi Swaroop Arlagadda Narasimharaju,Venkat Kumar摘要:以Shazam等先驱为例的音频指纹识别已经改变了数字音频识别。然而,现有的系统在具有挑战性的条件下难以准确,限制了广泛的适用性。本研究提出一种AI和ML整合的音频指纹算法,以提高准确性。该研究建立在Dejavu项目的基础上,强调具有不同背景噪音和失真的真实场景模拟。信号处理是Dejavu模型的核心,包括快速傅立叶变换、频谱图和峰值提取。“星座”概念和指纹散列实现了独特的歌曲识别。性能评估证明,在5秒的音频输入内,100%的准确性,系统展示了可预测的匹配速度,以提高效率。存储分析强调了实际实施的关键空间-速度权衡。这项研究提高了音频指纹的适应性,解决了各种环境和应用中的挑战。摘要:Audio fingerprinting, exemplified by pioneers like Shazam, has transformed digital audio recognition. However, existing systems struggle with accuracy in challenging conditions, limiting broad applicability. This research proposes an AI and ML integrated audio fingerprinting algorithm to enhance accuracy. Built on the Dejavu Project's foundations, the study emphasizes real-world scenario simulations with diverse background noises and distortions. Signal processing, central to Dejavu's model, includes the Fast Fourier Transform, spectrograms, and peak extraction. The "constellation" concept and fingerprint hashing enable unique song identification. Performance evaluation attests to 100% accuracy within a 5-second audio input, with a system showcasing predictable matching speed for efficiency. Storage analysis highlights the critical space-speed trade-off for practical implementation. This research advances audio fingerprinting's adaptability, addressing challenges in varied environments and applications. 【5】 Voice-Driven Mortality Prediction in Hospitalized Heart Failure Patients: A Machine Learning Approach Enhanced with Diagnostic Biomarkers标题:心力衰竭住院患者的语音驱动死亡率预测:一种通过诊断生物标志物增强的机器学习方法链接:https://arxiv.org/abs/2402.13812作者:Nihat Ahmadli,Mehmet Ali Sarsil,Berk Mizrak,Kurtulus Karauzum,Ata Shaker,Erol Tulumen,Didar Mirzamidinov,Dilek Ural,Onur Ergen备注:11 pages, 6 figures, 5 tables. THe first 2 authors have contributed equally摘要:解决心力衰竭(HF)作为一个普遍的全球健康问题带来了困难,实施创新的方法,以加强病人的护理。特别是预测HF患者的死亡率是困难的,但也是关键的,需要个性化护理,积极主动的管理,并使教育决策,以提高结果。最近,声音生物标志物与机器学习(ML)相结合的重要性激增,表现出显着的功效,特别是在预测心力衰竭方面。语音分析和ML算法的协同作用提供了一种非侵入性且易于访问的方法来评估患者的健康状况。然而,缺乏语音生物标志物来预测标准化语音协议的心力衰竭患者的死亡率。在这里,我们展示了一个强大而有效的ML模型,通过利用语音生物标志物来预测住院HF患者的死亡率。通过将语音生物标志物无缝集成到常规患者监测中,该策略有可能改善患者结局,优化资源分配,并推进以患者为中心的HF管理。在这项研究中,一个机器学习系统,特别是一个逻辑回归模型,被训练来预测患者的5年死亡率,使用他们的语音作为输入。交叉验证和统计方法证明,该模型的表现令人钦佩且一致(p值< 0.001)。此外,整合NT-proBNP,HF中的诊断生物标志物,大大提高了模型的预测准确性。摘要:Addressing heart failure (HF) as a prevalent global health concern poses difficulties in implementing innovative approaches for enhanced patient care. Predicting mortality rates in HF patients, in particular, is difficult yet critical, necessitating individualized care, proactive management, and enabling educated decision-making to enhance outcomes. Recently, the significance of voice biomarkers coupled with Machine Learning (ML) has surged, demonstrating remarkable efficacy, particularly in predicting heart failure. The synergy of voice analysis and ML algorithms provides a non-invasive and easily accessible means to evaluate patients' health. However, there is a lack of voice biomarkers for predicting mortality rates among heart failure patients with standardized speech protocols. Here, we demonstrate a powerful and effective ML model for predicting mortality rates in hospitalized HF patients through the utilization of voice biomarkers. By seamlessly integrating voice biomarkers into routine patient monitoring, this strategy has the potential to improve patient outcomes, optimize resource allocation, and advance patient-centered HF management. In this study, a Machine Learning system, specifically a logistic regression model, is trained to predict patients' 5-year mortality rates using their speech as input. The model performs admirably and consistently, as demonstrated by cross-validation and statistical approaches (p-value < 0.001). Furthermore, integrating NT-proBNP, a diagnostic biomarker in HF, improves the model's predictive accuracy substantially.
【6】 Music Style Transfer with Time-Varying Inversion of Diffusion Models标题:基于扩散模型时变逆的音乐风格转换链接:https://arxiv.org/abs/2402.13763作者:Sifei Li,Yuxin Zhang,Fan Tang,Chongyang Ma,Weiming dong,Changsheng Xu备注:7 pages, 4 figures, AAAI 2024摘要:随着扩散模型的发展,文本引导的图像风格转换已经显示出高质量的可控合成结果。然而,利用文本进行不同的音乐风格转移带来了重大挑战,主要是由于匹配的音频文本数据集的可用性有限。音乐是一种抽象而复杂的艺术形式,即使在同一流派中也会表现出变化和复杂性,从而使准确的文本描述具有挑战性。本文提出了一种音乐风格转移的方法,有效地捕捉音乐属性,使用最少的数据。我们引入了一种新的时变文本反转模块,以精确地捕获不同级别的梅尔频谱图特征。在推理过程中,我们提出了一个减少偏见的风格化技术,以获得稳定的结果。实验结果表明,我们的方法可以转移特定乐器的风格,以及纳入自然的声音组成的旋律。示例和源代码可在https://lsfhuihuiff.github.io/MusicTI/上获得。摘要:With the development of diffusion models, text-guided image style transfer has demonstrated high-quality controllable synthesis results. However, the utilization of text for diverse music style transfer poses significant challenges, primarily due to the limited availability of matched audio-text datasets. Music, being an abstract and complex art form, exhibits variations and intricacies even within the same genre, thereby making accurate textual descriptions challenging. This paper presents a music style transfer approach that effectively captures musical attributes using minimal data. We introduce a novel time-varying textual inversion module to precisely capture mel-spectrogram features at different levels. During inference, we propose a bias-reduced stylization technique to obtain stable results. Experimental results demonstrate that our method can transfer the style of specific instruments, as well as incorporate natural sounds to compose melodies. Samples and source code are available at https://lsfhuihuiff.github.io/MusicTI/.
【7】 The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning标题:批次大小对对比性自监督语音表征学习的影响链接:https://arxiv.org/abs/2402.13723作者:Nik Vaessen,David A. van Leeuwen摘要:语音中的基础模型通常使用许多GPU进行训练,这隐含地导致了大的有效批量。在本文中,我们研究了批量大小对预训练的影响,包括在训练过程中可以监控的统计数据,以及对下游微调任务性能的影响。通过使用从87.5秒到80分钟不等的批量大小,我们表明,对于固定的迭代次数,较大的批量大小会产生更好的预训练模型。然而,稳定性有下限,有效性有上限。然后,我们表明,预训练模型的质量主要取决于训练期间看到的语音数据量,即,批量大小和迭代次数的乘积。所有结果都是通过wav2vec 2.0架构的独立实现产生的,该架构在很大程度上再现了原始工作的结果(arXiv:2006.11477)。我们的扩展可以帮助研究人员在研究语音中的自监督学习时选择有效的操作条件,并提示使用固定数量的可见数据对自监督进行基准测试。代码和模型检查点可以在https://github.com/nikvaessen/w2v2-batch-size上找到。摘要:Foundation models in speech are often trained using many GPUs, which implicitly leads to large effective batch sizes. In this paper we study the effect of batch size on pre-training, both in terms of statistics that can be monitored during training, and in the effect on the performance of a downstream fine-tuning task. By using batch sizes varying from 87.5 seconds to 80 minutes of speech we show that, for a fixed amount of iterations, larger batch sizes result in better pre-trained models. However, there is lower limit for stability, and an upper limit for effectiveness. We then show that the quality of the pre-trained model depends mainly on the amount of speech data seen during training, i.e., on the product of batch size and number of iterations. All results are produced with an independent implementation of the wav2vec 2.0 architecture, which to a large extent reproduces the results of the original work (arXiv:2006.11477). Our extensions can help researchers choose effective operating conditions when studying self-supervised learning in speech, and hints towards benchmarking self-supervision with a fixed amount of seen data. Code and model checkpoints are available at https://github.com/nikvaessen/w2v2-batch-size. 【8】 Structure-informed Positional Encoding for Music Generation标题:用于音乐生成的结构信息位置编码链接:https://arxiv.org/abs/2402.13301作者:Manvi Agarwal,Changhong Wang,Gaël Richard备注:None摘要:深度学习方法生成的音乐通常缺乏连贯性和长期组织。然而,多尺度层次结构是音乐信号的显著特征。为了利用这些信息,我们提出了一个结构知情的位置编码框架与Transformers的音乐生成。我们设计了三个变种的绝对,相对和非静止的位置信息。我们全面测试他们的两个象征性的音乐生成任务:下一个时间步预测和伴奏生成。作为比较,我们从文献中选择多个基线,并使用几个音乐动机的评价指标来证明我们的方法的优点。特别是,我们的方法提高了生成的作品的旋律和结构的一致性。摘要:Music generated by deep learning methods often suffers from a lack of coherence and long-term organization. Yet, multi-scale hierarchical structure is a distinctive feature of music signals. To leverage this information, we propose a structure-informed positional encoding framework for music generation with Transformers. We design three variants in terms of absolute, relative and non-stationary positional information. We comprehensively test them on two symbolic music generation tasks: next-timestep prediction and accompaniment generation. As a comparison, we choose multiple baselines from the literature and demonstrate the merits of our methods using several musically-motivated evaluation metrics. In particular, our methods improve the melodic and structural consistency of the generated pieces.