微信公众号:arXiv_Daily
cs.SD语音
标题: 潜表示与Cosine距离对齐的语音去噪知识提炼
链接:https://arxiv.org/abs/2505.03442
摘要:语音去噪是一项普遍采用且有影响力的任务,出现在许多常见和日常生活中的用例中。虽然有非常强大的方法公布,其中大多数是太复杂的部署在日常和低资源的计算环境,如手持设备,智能眼镜,助听器等知识蒸馏(KD)是一个突出的方式,以减轻这种复杂性不匹配,是基于转移/蒸馏的知识从一个预先训练的复杂模型,教师,另一个不太复杂的,那个学生现有的KD语音去噪方法的基础上的过程,可能会阻碍KD的学生的学习,分布,信息排序,和教师学习的特征维度。在本文中,我们提出并评估了一种方法,试图处理这个问题,通过利用著名的去噪自动编码器框架,线性倒瓶颈,和余弦相似性的属性。我们使用一个公共数据集,并在教师和学生之间的不同不匹配场景下进行重复实验,报告我们的方法和另一种用作基线的最先进方法的指标的平均值和标准差。我们的研究结果表明,与所提出的方法相比,学生可以表现得更好,也可以保留更大的不匹配条件的老师。
摘要:Speech denoising is a generally adopted and impactful task, appearing in many common and everyday-life use cases. Although there are very powerful methods published, most of those are too complex for deployment in everyday and low-resources computational environments, like hand-held devices, intelligent glasses, hearing aids, etc. Knowledge distillation (KD) is a prominent way for alleviating this complexity mismatch and is based on the transferring/distilling of knowledge from a pre-trained complex model, the teacher, to another less complex one, the student. Existing KD methods for speech denoising are based on processes that potentially hamper the KD by bounding the learning of the student to the distribution, information ordering, and feature dimensionality learned by the teacher. In this paper, we present and assess a method that tries to treat this issue, by exploiting the well-known denoising-autoencoder framework, the linear inverted bottlenecks, and the properties of the cosine similarity. We use a public dataset and conduct repeated experiments with different mismatching scenarios between the teacher and the student, reporting the mean and standard deviation of the metrics of our method and another, state-of-the-art method that is used as a baseline. Our results show that with the proposed method, the student can perform better and can also retain greater mismatching conditions compared to the teacher.
【2】 The Inverse Drum Machine: Source Separation Through Joint Transcription and Analysis-by-Synthesis
标题: 倒置鼓机:通过联合转录和综合分析进行源分离链接:https://arxiv.org/abs/2505.03337
摘要:我们介绍了反向鼓机(IDM),这是一种将综合分析与深度学习相结合的鼓源分离新方法。与最近的监督方法,依赖于孤立的茎,IDM只需要转录注释。它在端到端框架中联合优化了自动鼓转录和一次性鼓样本合成。通过将合成的一次性样本与估计的起始时间进行卷积-模仿鼓机-IDM重建单个鼓茎并训练神经网络以匹配原始混合物。对StemGMD数据集的评估表明,IDM实现了与最先进的监督方法相当的分离性能,同时大大优于矩阵分解基线。
摘要:We introduce the Inverse Drum Machine (IDM), a novel approach to drum source separation that combines analysis-by-synthesis with deep learning. Unlike recent supervised methods that rely on isolated stems, IDM requires only transcription annotations. It jointly optimizes automatic drum transcription and one-shot drum sample synthesis in an end-to-end framework. By convolving synthesized one-shot samples with estimated onsets-mimicking a drum machine-IDM reconstructs individual drum stems and trains a neural network to match the original mixture. Evaluations on the StemGMD dataset show that IDM achieves separation performance on par with state-of-the-art supervised methods, while substantially outperforming matrix decomposition baselines.
【3】 Mamba-Diffusion Model with Learnable Wavelet for Controllable Symbolic Music Generation
标题: 用于可控符号音乐生成的可学习子波曼巴扩散模型链接:https://arxiv.org/abs/2505.03314
摘要:最近激增的普及扩散模型的图像合成吸引了新的注意力,他们的潜力,在其他领域的生成任务。然而,它们在符号音乐生成中的应用在很大程度上仍然未被探索,因为符号音乐通常表示为离散事件序列,标准扩散模型不适合离散数据。我们代表象征性的音乐形象般的钢琴,促进使用扩散模型的象征性音乐的产生。此外,本研究还引入了一种新的扩散模型,该模型结合了我们提出的Transformer-Mamba块和可学习的小波变换。无分类器的指导被用来生成具有目标和弦的符号音乐。我们的评估表明,我们的方法在音乐质量和可控性方面取得了令人信服的结果,在钢琴演奏生成方面优于强大的基线。我们的代码可在https://github.com/jinchengzhanggg/proffusion上获得。
摘要:The recent surge in the popularity of diffusion models for image synthesis has attracted new attention to their potential for generation tasks in other domains. However, their applications to symbolic music generation remain largely under-explored because symbolic music is typically represented as sequences of discrete events and standard diffusion models are not well-suited for discrete data. We represent symbolic music as image-like pianorolls, facilitating the use of diffusion models for the generation of symbolic music. Moreover, this study introduces a novel diffusion model that incorporates our proposed Transformer-Mamba block and learnable wavelet transform. Classifier-free guidance is utilised to generate symbolic music with target chords. Our evaluation shows that our method achieves compelling results in terms of music quality and controllability, outperforming the strong baseline in pianoroll generation. Our code is available at https://github.com/jinchengzhanggg/proffusion.
【4】 SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation
标题: SepILM:音频语言模型是鲁棒语音分离的错误纠正器链接:https://arxiv.org/abs/2505.03273
备注:Appears in IJCAI 2025
摘要:虽然当代语音分离技术熟练地处理冗长的混合音频波形,但它们经常受到现实世界环境的复杂性的挑战,包括嘈杂和混响设置,这可能导致分离语音中的伪像或失真。为了克服这些局限性,我们引入SepALM,一种开创性的方法,采用音频语言模型(ALMs)来纠正和重新合成语音的文本域初步分离后。SepALM包括四个核心组件:分离器、校正器、合成器和对准器。通过集成基于ALM的端到端纠错机制,我们减轻了错误积累的风险,并规避了传统方法中通常遇到的优化障碍,这些方法将自动语音识别(ASR)与大型语言模型(LLM)相结合。此外,我们还开发了思想链(CoT)提示和知识蒸馏技术,以促进ALM的推理和培训过程。实验结果表明,SepALM不仅提高了语音分离的精度,而且在新的声学环境中也显着增强了适应性。
摘要:While contemporary speech separation technologies adeptly process lengthy mixed audio waveforms, they are frequently challenged by the intricacies of real-world environments, including noisy and reverberant settings, which can result in artifacts or distortions in the separated speech. To overcome these limitations, we introduce SepALM, a pioneering approach that employs audio language models (ALMs) to rectify and re-synthesize speech within the text domain following preliminary separation. SepALM comprises four core components: a separator, a corrector, a synthesizer, and an aligner. By integrating an ALM-based end-to-end error correction mechanism, we mitigate the risk of error accumulation and circumvent the optimization hurdles typically encountered in conventional methods that amalgamate automatic speech recognition (ASR) with large language models (LLMs). Additionally, we have developed Chain-of-Thought (CoT) prompting and knowledge distillation techniques to facilitate the reasoning and training processes of the ALM. Our experiments substantiate that SepALM not only elevates the precision of speech separation but also markedly bolsters adaptability in novel acoustic environments.
【5】 SonicRAG : High Fidelity Sound Effects Synthesis Based on Retrival Augmented Generation
标题: SonicRAG:基于检索增强生成的高保真音效合成链接:https://arxiv.org/abs/2505.03244
备注:8 pages, 5 figures
摘要:大型语言模型(LLM)在自然语言处理(NLP)和多模态学习方面表现出了卓越的能力,在文本生成和语音合成方面的成功应用,使人们能够更深入地理解和生成多模态内容。在声音效果(SFX)生成领域,LLM已经被利用来编排用于音频合成的多个模型。然而,由于标注数据集的稀缺性,以及时间建模的复杂性。当前的SFX生成技术在实现高保真音频方面仍然不足。为了解决这些限制,本文介绍了一种新的框架,集成LLM与现有的音效数据库,允许检索,重组和合成的音频根据用户的要求。通过利用这种方法,我们提高了生成的声音效果的多样性和质量,同时消除了额外的录音成本的需要,为声音设计和应用提供了灵活而高效的解决方案。
摘要:Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language processing (NLP) and multimodal learning, with successful applications in text generation and speech synthesis, enabling a deeper understanding and generation of multimodal content. In the field of sound effects (SFX) generation, LLMs have been leveraged to orchestrate multiple models for audio synthesis. However, due to the scarcity of annotated datasets, and the complexity of temproal modeling. current SFX generation techniques still fall short in achieving high-fidelity audio. To address these limitations, this paper introduces a novel framework that integrates LLMs with existing sound effect databases, allowing for the retrieval, recombination, and synthesis of audio based on user requirements. By leveraging this approach, we enhance the diversity and quality of generated sound effects while eliminating the need for additional recording costs, offering a flexible and efficient solution for sound design and application.
【6】 MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification
标题: MGFF-TDNN:一种具有深度可分离模块的多粒度特征融合TDNN模型,用于说话人验证链接:https://arxiv.org/abs/2505.03228
摘要:在说话人确认中,传统的模型往往强调建模长期的上下文特征,以捕获全局说话人特征。然而,这种方法可能会忽略细粒度的声纹信息,其中包含了鲁棒的说话人嵌入所必需的高度判别特征。提出了一种基于多粒度特征融合的模型结构MGFF-TDNN。MGFF-TDNN利用二维深度可分离卷积模块,通过局部特征建模进行增强,作为前端特征提取器,以有效地捕获时频域特征。为了实现全面的多粒度特征融合,我们提出了M-TDNN结构,该结构通过结合时延神经网络和音素级特征池,将全局上下文建模与细粒度特征提取相结合。在VoxCeleb数据集上的实验表明,MGFF-TDNN在说话人确认方面取得了出色的性能,同时在参数和计算资源方面保持了高效。
摘要:In speaker verification, traditional models often emphasize modeling long-term contextual features to capture global speaker characteristics. However, this approach can neglect fine-grained voiceprint information, which contains highly discriminative features essential for robust speaker embeddings. This paper introduces a novel model architecture, termed MGFF-TDNN, based on multi-granularity feature fusion. The MGFF-TDNN leverages a two-dimensional depth-wise separable convolution module, enhanced with local feature modeling, as a front-end feature extractor to effectively capture time-frequency domain features. To achieve comprehensive multi-granularity feature fusion, we propose the M-TDNN structure, which integrates global contextual modeling with fine-grained feature extraction by combining time-delay neural networks and phoneme-level feature pooling. Experiments on the VoxCeleb dataset demonstrate that the MGFF-TDNN achieves outstanding performance in speaker verification while remaining efficient in terms of parameters and computational resources.
【7】 A study on audio synchronous steganography detection and distributed guide inference model based on sliding spectral features and intelligent inference drive
标题: 基于滑动谱特征和智能推理驱动的音频同步隐写检测和分布式引导推理模型研究链接:https://arxiv.org/abs/2505.03193
备注:This paper proposes a novel framework for detecting steganographic content in short video audio streams using sliding spectral features and distributed inference models, combining STFT analysis, entropy-based synchronization, and deep learning-driven decoding strategies
摘要:随着短视频平台在全球传播中的兴起,在音频同步流中嵌入隐写数据成为一种新的隐蔽通信方法。针对传统技术在同步隐写检测上的局限性,本文提出了一种基于中国南海舰队在TikTok上发布的短视频“玉盘”样本的检测和分布式引导重构模型。该方法集成了滑动谱特征提取和智能推理机制。使用具有短时傅立叶变换(STFT)的25 ms滑动窗口来提取主频率轨迹并构造同步帧检测模型(M1),识别帧标志“FFFFFFFFFFFFFF80”。随后的32字节有效载荷由结构化模型(M2)解码以推断分布式制导命令。分析揭示了在具有高度集中的频谱能量的36至45秒音频段中的低熵重复字节序列,从而确认了同步帧的存在。虽然未恢复明文语义,但命令字段布局的一致性表明了军事通信协议的特征。多段拼接模型进一步展示了跨视频嵌入和集中解码能力。该框架验证了滑动谱特征在同步隐写检测中的有效性,并为开放平台上的隐蔽通信分析和战术制导仿真建立了可扩展的推理模型。
摘要:With the rise of short video platforms in global communication, embedding steganographic data in audio synchronization streams has emerged as a new covert communication method. To address the limitations of traditional techniques in detecting synchronized steganography, this paper proposes a detection and distributed guidance reconstruction model based on short video "Yupan" samples released by China's South Sea Fleet on TikTok. The method integrates sliding spectrum feature extraction and intelligent inference mechanisms. A 25 ms sliding window with short-time Fourier transform (STFT) is used to extract the main frequency trajectory and construct the synchronization frame detection model (M1), identifying a frame flag "FFFFFFFFFFFFFFFFFF80". The subsequent 32-byte payload is decoded by a structured model (M2) to infer distributed guidance commands. Analysis reveals a low-entropy, repetitive byte sequence in the 36 to 45 second audio segment with highly concentrated spectral energy, confirming the presence of synchronization frames. Although plaintext semantics are not restored, the consistency in command field layout suggests features of military communication protocols. The multi-segment splicing model further shows cross-video embedding and centralized decoding capabilities. The proposed framework validates the effectiveness of sliding spectral features for synchronized steganography detection and builds an extensible inference model for covert communication analysis and tactical guidance simulation on open platforms.
【8】 CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
标题: CoGenAV:通过对比生成同步的多功能视听表示学习链接:https://arxiv.org/abs/2505.03186
摘要:说话者的嘴唇运动、语音和潜在的语言内容之间的固有同步为改进语音处理任务提供了丰富的信息源,特别是在传统的仅音频系统不稳定的挑战性条件下。我们介绍CoGenAV,一个强大的和数据高效的模型,旨在学习适用于各种语音和视听任务的多功能视听表示。CoGenAV通过优化来自自然视听同步、对比特征对齐和生成文本预测的双重目标进行训练,仅使用来自LRS2数据集的223小时标记数据。这种对比生成同步策略有效地捕获了基本的跨模态相关性。我们展示了多个基准的学习CoGenAV表示的有效性和多功能性。当用于LRS2上的视听语音识别(AVSR)时,这些表示有助于实现1.27的最先进的字错误率(WER)。它们还在视觉语音识别(VSR)方面实现了强大的性能,在LRS2上的WER为22.0,并将嘈杂环境中的性能显著提高了70%以上。此外,CoGenAV表示有利于语音重建任务,提高语音增强和分离的性能,并在主动说话人检测(ASD)等视听同步任务中取得有竞争力的结果。我们的模型将是开源的,以促进学术界和工业界的进一步发展和合作。
摘要:The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions where traditional audio-only systems falter. We introduce CoGenAV, a powerful and data-efficient model designed to learn versatile audio-visual representations applicable across a wide range of speech and audio-visual tasks. CoGenAV is trained by optimizing a dual objective derived from natural audio-visual synchrony, contrastive feature alignment and generative text prediction, using only 223 hours of labeled data from the LRS2 dataset. This contrastive-generative synchronization strategy effectively captures fundamental cross-modal correlations. We showcase the effectiveness and versatility of the learned CoGenAV representations on multiple benchmarks. When utilized for Audio-Visual Speech Recognition (AVSR) on LRS2, these representations contribute to achieving a state-of-the-art Word Error Rate (WER) of 1.27. They also enable strong performance in Visual Speech Recognition (VSR) with a WER of 22.0 on LRS2, and significantly improve performance in noisy environments by over 70%. Furthermore, CoGenAV representations benefit speech reconstruction tasks, boosting performance in Speech Enhancement and Separation, and achieve competitive results in audio-visual synchronization tasks like Active Speaker Detection (ASD). Our model will be open-sourced to facilitate further development and collaboration within both academia and industry.
【9】 Coupling the Heart to Musical Machines
标题: 将心脏与音乐机器相结合链接:https://arxiv.org/abs/2505.03073
摘要:生物反馈最近被用作人机接口(HCI)的一般控制范例。虽然生物反馈,特别是来自呼吸的生物反馈,已经越来越多地被用作新型音乐界面的控制器,新的音乐表达界面(NIME),但社区并没有对心脏给予太多的关注。心脏和呼吸一样,是音乐的一个重要组成部分,有人认为心脏决定了我们对时间的感知,从而间接地决定了我们对音乐的感知。受此启发,我演示了一个基于光电容积描记图(PPG)的NIME控制器,使用心率作为1D控制参数,通过蓝牙无线HCI实时转换声音的质量。我将时间缩放应用于“扭曲”音频缓冲区到声卡,并将这些转换后的音频缓冲区播放给佩戴PPG传感器的听众,创建一个假设的感知生物反馈回路:声音的变化改变心率,改变PPG测量值,从而改变声音。我讨论了声音-心脏-PPG生物反馈回路如何通过1D控制器提供更大的控制和/或各种运动,如何通过生物反馈控制声音回放的空间和/或时间尺度,从而实现性能氛围的可能性,我简要讨论了生成潜在空间作为扩展1D PPG控制空间的可能方式。
摘要:Biofeedback is being used more recently as a general control paradigm for human-computer interfaces (HCIs). While biofeedback especially from breath has seen increasing uptake as a controller for novel musical interfaces, new interfaces for musical expression (NIMEs), the community has not given as much attention to the heart. The heart is just as intimate a part of music as breath and it is argued that the heart determines our perception of time and so indirectly our perception of music. Inspired by this I demonstrate a photoplethysmogram (PPG)-based NIME controller using heart rate as a 1D control parameter to transform the qualities of sounds in real-time over a Bluetooth wireless HCI. I apply time scaling to "warp" audio buffers inbound to the sound card, and play these transformed audio buffers back to the listener wearing the PPG sensor, creating a hypothetical perceptual biofeedback loop: changes in sound change heart rate to change PPG measurements to change sound. I discuss how a sound-heart-PPG biofeedback loop possibly affords greater control and/or variety of movements with a 1D controller, how controlling the space and/or time scale of sound playback with biofeedback makes for possibilities in performance ambience, and I briefly discuss generative latent spaces as a possible way to extend a 1D PPG control space.
【10】 BLAB: Brutally Long Audio Bench
标题: BLAB:残酷的长音频长凳链接:https://arxiv.org/abs/2505.03054
摘要:开发能够理解不同口语交互的大型音频语言模型(LM)对于适应人类交流的多模态性质至关重要,并且可以增加语言技术在不同用户群体中的可访问性。最近关于音频LM的工作主要评估了它们在短音频片段上的性能,通常在30秒以下,对长形式对话语音片段的探索有限,这些片段更接近地反映了用户与这些模型的自然交互。我们介绍了Brutally Long Audio Bench(BLAB),这是一个具有挑战性的长格式音频基准测试,使用平均长度为51分钟的音频片段评估音频LM的本地化,持续时间估计,情感和计数任务。BLAB由833小时以上的各种全长音频片段组成,每个片段都配有人工注释的基于文本的自然语言问题和答案。我们的音频数据是从许可的来源收集的,并经过人工辅助过滤过程,以确保任务合规性。我们评估了BLAB上的六个开源和专有音频LM,发现所有这些,包括Gemini 2.0 Pro和GPT-4 o等高级型号,都在BLAB中的任务中挣扎。我们的全面分析揭示了任务难度和音频持续时间之间权衡的关键见解。总的来说,我们发现音频LM在长形式的语音中挣扎,随着持续时间的增加,性能下降。他们在定位、时间推理、计数方面表现不佳,并且很难理解非音素信息,更多地依赖提示而不是音频内容。BLAB是一个具有挑战性的评估框架,用于开发具有强大的长格式音频理解能力的音频LM。
摘要:Developing large audio language models (LMs) capable of understanding diverse spoken interactions is essential for accommodating the multimodal nature of human communication and can increase the accessibility of language technologies across different user populations. Recent work on audio LMs has primarily evaluated their performance on short audio segments, typically under 30 seconds, with limited exploration of long-form conversational speech segments that more closely reflect natural user interactions with these models. We introduce Brutally Long Audio Bench (BLAB), a challenging long-form audio benchmark that evaluates audio LMs on localization, duration estimation, emotion, and counting tasks using audio segments averaging 51 minutes in length. BLAB consists of 833+ hours of diverse, full-length audio clips, each paired with human-annotated, text-based natural language questions and answers. Our audio data were collected from permissively licensed sources and underwent a human-assisted filtering process to ensure task compliance. We evaluate six open-source and proprietary audio LMs on BLAB and find that all of them, including advanced models such as Gemini 2.0 Pro and GPT-4o, struggle with the tasks in BLAB. Our comprehensive analysis reveals key insights into the trade-offs between task difficulty and audio duration. In general, we find that audio LMs struggle with long-form speech, with performance declining as duration increases. They perform poorly on localization, temporal reasoning, counting, and struggle to understand non-phonemic information, relying more on prompts than audio content. BLAB serves as a challenging evaluation framework to develop audio LMs with robust long-form audio understanding capabilities.
标题: 唇裂语音自动语音识别的公平性
链接:https://arxiv.org/abs/2505.03697
备注:Submitted to Digital Signal Processing
摘要:唇腭裂(CLP)患者产生的语音通常由于结构异常而高度鼻音化和呼吸,导致共振峰结构的变化,影响自动语音识别(ASR)的性能和公平性。这项研究假设,公开可用的ASR系统表现出降低公平性CLP语音,并通过实验证实了这一点。尽管共振峰中断,轻度和中度CLP语音保留一些频谱时间对齐与正常语音,激励增强策略,以提高公平性。该研究系统地探讨了增强CLP语音与正常语音的严重程度,并评估其对ASR公平性的影响。在AIISH和NMCPC数据集上测试了三种ASR模型GMM-HMM,Whisper和XLS-R。结果表明,正常语音的训练和混合数据的测试提高了字错误率(WER)。值得注意的是,WER从22.64美元降至18.76美元(GMM-HMM,AIISH),从28.45美元降至18.89美元(Whisper,NMCPC)。GMM-HMM在AIISH上的优异性能可能是由于其适用于卡纳达语儿童的语音,这对XLS-R和Whisper等基础模型来说是一个挑战。为了评估公平性,引入了公平性评分,揭示了增加后的17.89美元(AIISH)和47.50美元(NMCPC)的改善。
摘要:Speech produced by individuals with cleft lip and palate (CLP) is often highly nasalized and breathy due to structural anomalies, causing shifts in formant structure that affect automatic speech recognition (ASR) performance and fairness. This study hypothesizes that publicly available ASR systems exhibit reduced fairness for CLP speech and confirms this through experiments. Despite formant disruptions, mild and moderate CLP speech retains some spectro-temporal alignment with normal speech, motivating augmentation strategies to enhance fairness. The study systematically explores augmenting CLP speech with normal speech across severity levels and evaluates its impact on ASR fairness. Three ASR models-GMM-HMM, Whisper, and XLS-R-were tested on AIISH and NMCPC datasets. Results indicate that training with normal speech and testing on mixed data improves word error rate (WER). Notably, WER decreased from $22.64\%$ to $18.76\%$ (GMM-HMM, AIISH) and $28.45\%$ to $18.89\%$ (Whisper, NMCPC). The superior performance of GMM-HMM on AIISH may be due to its suitability for Kannada children's speech, a challenge for foundation models like XLS-R and Whisper. To assess fairness, a fairness score was introduced, revealing improvements of $17.89\%$ (AIISH) and $47.50\%$ (NMCPC) with augmentation.
【2】 The Search for Squawk: Agile Modeling in Bioacoustics
标题: 寻找Squawk:生物声学中的敏捷建模链接:https://arxiv.org/abs/2505.03071
摘要:被动声监测(PAM)在帮助生态学家了解动物种群和生态系统的健康方面显示出巨大的潜力。然而,从数百万小时的音频记录中提取见解需要开发专门的识别器。这通常是一项具有挑战性的任务,需要大量的训练数据和机器学习专业知识。在这项工作中,我们介绍了一个通用的,可扩展的和数据高效的系统,用于在一个小时内开发新的生物声学问题的识别器。我们的系统由几个关键组件组成,可以解决以前生物声学工作流程中的问题:1)预先训练用于鸟鸣分类的高度概括的声学嵌入,最大限度地减少数据饥饿; 2)索引音频搜索允许有效创建分类器训练数据集,以及3)嵌入的预计算实现了有效的主动学习循环,以最小的等待时间迭代地提高分类器质量。生态学家在三个新的案例研究中使用了我们的系统:通过不明声音分析珊瑚礁的健康状况;识别夏威夷幼鸟的叫声,以量化繁殖成功率并改善濒危物种监测;圣诞岛鸟类居住建模。我们通过模拟实验来增加案例研究,这些实验以结构化的方式探索设计决策的范围,并帮助建立最佳实践。总之,这些实验展示了我们系统的可扩展性,效率和通用性,使科学家能够快速应对新的生物声学挑战。
摘要:Passive acoustic monitoring (PAM) has shown great promise in helping ecologists understand the health of animal populations and ecosystems. However, extracting insights from millions of hours of audio recordings requires the development of specialized recognizers. This is typically a challenging task, necessitating large amounts of training data and machine learning expertise. In this work, we introduce a general, scalable and data-efficient system for developing recognizers for novel bioacoustic problems in under an hour. Our system consists of several key components that tackle problems in previous bioacoustic workflows: 1) highly generalizable acoustic embeddings pre-trained for birdsong classification minimize data hunger; 2) indexed audio search allows the efficient creation of classifier training datasets, and 3) precomputation of embeddings enables an efficient active learning loop, improving classifier quality iteratively with minimal wait time. Ecologists employed our system in three novel case studies: analyzing coral reef health through unidentified sounds; identifying juvenile Hawaiian bird calls to quantify breeding success and improve endangered species monitoring; and Christmas Island bird occupancy modeling. We augment the case studies with simulated experiments which explore the range of design decisions in a structured way and help establish best practices. Altogether these experiments showcase our system's scalability, efficiency, and generalizability, enabling scientists to quickly address new bioacoustic challenges.
【3】 Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance
标题: 潜表示与Cosine距离对齐的语音去噪知识提炼链接:https://arxiv.org/abs/2505.03442
摘要:语音去噪是一项普遍采用且有影响力的任务,出现在许多常见的日常生活用例中。虽然有非常强大的方法公布,其中大多数是太复杂的部署在日常和低资源的计算环境,如手持设备,智能眼镜,助听器等知识蒸馏(KD)是一个突出的方式,以减轻这种复杂性不匹配,是基于转移/蒸馏的知识从一个预先训练的复杂模型,教师,另一个不太复杂的,那个学生现有的KD语音去噪方法的基础上的过程,可能会阻碍KD的学生的学习,分布,信息排序,和教师学习的特征维度。在本文中,我们提出并评估了一种方法,试图处理这个问题,通过利用著名的去噪自动编码器框架,线性倒瓶颈,和余弦相似性的属性。我们使用一个公共数据集,并在教师和学生之间的不同不匹配场景下进行重复实验,报告我们的方法和另一种用作基线的最先进方法的指标的平均值和标准差。我们的研究结果表明,与所提出的方法相比,学生可以表现得更好,也可以保留更大的不匹配条件的老师。
摘要:Speech denoising is a generally adopted and impactful task, appearing in many common and everyday-life use cases. Although there are very powerful methods published, most of those are too complex for deployment in everyday and low-resources computational environments, like hand-held devices, intelligent glasses, hearing aids, etc. Knowledge distillation (KD) is a prominent way for alleviating this complexity mismatch and is based on the transferring/distilling of knowledge from a pre-trained complex model, the teacher, to another less complex one, the student. Existing KD methods for speech denoising are based on processes that potentially hamper the KD by bounding the learning of the student to the distribution, information ordering, and feature dimensionality learned by the teacher. In this paper, we present and assess a method that tries to treat this issue, by exploiting the well-known denoising-autoencoder framework, the linear inverted bottlenecks, and the properties of the cosine similarity. We use a public dataset and conduct repeated experiments with different mismatching scenarios between the teacher and the student, reporting the mean and standard deviation of the metrics of our method and another, state-of-the-art method that is used as a baseline. Our results show that with the proposed method, the student can perform better and can also retain greater mismatching conditions compared to the teacher.
【4】 The Inverse Drum Machine: Source Separation Through Joint Transcription and Analysis-by-Synthesis
标题: 倒置鼓机:通过联合转录和综合分析进行源分离链接:https://arxiv.org/abs/2505.03337
摘要:我们介绍了反向鼓机(IDM),这是一种将综合分析与深度学习相结合的鼓源分离新方法。与最近的监督方法,依赖于孤立的茎,IDM只需要转录注释。它在端到端框架中联合优化了自动鼓转录和一次性鼓样本合成。通过将合成的一次性样本与估计的起始时间进行卷积-模仿鼓机-IDM重建单个鼓茎并训练神经网络以匹配原始混合物。对StemGMD数据集的评估表明,IDM实现了与最先进的监督方法相当的分离性能,同时大大优于矩阵分解基线。
摘要:We introduce the Inverse Drum Machine (IDM), a novel approach to drum source separation that combines analysis-by-synthesis with deep learning. Unlike recent supervised methods that rely on isolated stems, IDM requires only transcription annotations. It jointly optimizes automatic drum transcription and one-shot drum sample synthesis in an end-to-end framework. By convolving synthesized one-shot samples with estimated onsets-mimicking a drum machine-IDM reconstructs individual drum stems and trains a neural network to match the original mixture. Evaluations on the StemGMD dataset show that IDM achieves separation performance on par with state-of-the-art supervised methods, while substantially outperforming matrix decomposition baselines.
【5】 Mamba-Diffusion Model with Learnable Wavelet for Controllable Symbolic Music Generation
标题: 用于可控符号音乐生成的可学习子波曼巴扩散模型链接:https://arxiv.org/abs/2505.03314
摘要:最近激增的普及扩散模型的图像合成吸引了新的注意力,他们的潜力,在其他领域的生成任务。然而,它们在符号音乐生成中的应用在很大程度上仍然未被探索,因为符号音乐通常表示为离散事件序列,标准扩散模型不适合离散数据。我们代表象征性的音乐形象般的钢琴,促进使用扩散模型的象征性音乐的产生。此外,本研究还引入了一种新的扩散模型,该模型结合了我们提出的Transformer-Mamba块和可学习的小波变换。无分类器的指导被用来生成具有目标和弦的符号音乐。我们的评估表明,我们的方法在音乐质量和可控性方面取得了令人信服的结果,在钢琴演奏生成方面优于强大的基线。我们的代码可在https://github.com/jinchengzhanggg/proffusion上获得。
摘要:The recent surge in the popularity of diffusion models for image synthesis has attracted new attention to their potential for generation tasks in other domains. However, their applications to symbolic music generation remain largely under-explored because symbolic music is typically represented as sequences of discrete events and standard diffusion models are not well-suited for discrete data. We represent symbolic music as image-like pianorolls, facilitating the use of diffusion models for the generation of symbolic music. Moreover, this study introduces a novel diffusion model that incorporates our proposed Transformer-Mamba block and learnable wavelet transform. Classifier-free guidance is utilised to generate symbolic music with target chords. Our evaluation shows that our method achieves compelling results in terms of music quality and controllability, outperforming the strong baseline in pianoroll generation. Our code is available at https://github.com/jinchengzhanggg/proffusion.
【6】 SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation
标题: SepILM:音频语言模型是鲁棒语音分离的错误纠正器链接:https://arxiv.org/abs/2505.03273
备注:Appears in IJCAI 2025
摘要:虽然当代语音分离技术熟练地处理冗长的混合音频波形,但它们经常受到现实世界环境的复杂性的挑战,包括嘈杂和混响设置,这可能导致分离语音中的伪像或失真。为了克服这些局限性,我们引入SepALM,一种开创性的方法,采用音频语言模型(ALMs)来纠正和重新合成语音的文本域初步分离后。SepALM包括四个核心组件:分离器、校正器、合成器和对准器。通过集成基于ALM的端到端纠错机制,我们减轻了错误积累的风险,并规避了传统方法中通常遇到的优化障碍,这些方法将自动语音识别(ASR)与大型语言模型(LLM)相结合。此外,我们还开发了思想链(CoT)提示和知识蒸馏技术,以促进ALM的推理和培训过程。实验结果表明,SepALM不仅提高了语音分离的精度,而且在新的声学环境中也显着增强了适应性。
摘要:While contemporary speech separation technologies adeptly process lengthy mixed audio waveforms, they are frequently challenged by the intricacies of real-world environments, including noisy and reverberant settings, which can result in artifacts or distortions in the separated speech. To overcome these limitations, we introduce SepALM, a pioneering approach that employs audio language models (ALMs) to rectify and re-synthesize speech within the text domain following preliminary separation. SepALM comprises four core components: a separator, a corrector, a synthesizer, and an aligner. By integrating an ALM-based end-to-end error correction mechanism, we mitigate the risk of error accumulation and circumvent the optimization hurdles typically encountered in conventional methods that amalgamate automatic speech recognition (ASR) with large language models (LLMs). Additionally, we have developed Chain-of-Thought (CoT) prompting and knowledge distillation techniques to facilitate the reasoning and training processes of the ALM. Our experiments substantiate that SepALM not only elevates the precision of speech separation but also markedly bolsters adaptability in novel acoustic environments.
【7】 SonicRAG : High Fidelity Sound Effects Synthesis Based on Retrival Augmented Generation
标题: SonicRAG:基于检索增强生成的高保真音效合成链接:https://arxiv.org/abs/2505.03244
备注:8 pages, 5 figures
摘要:大型语言模型(LLM)在自然语言处理(NLP)和多模态学习方面表现出了卓越的能力,在文本生成和语音合成方面的成功应用,使人们能够更深入地理解和生成多模态内容。在声音效果(SFX)生成领域,LLM已经被利用来编排用于音频合成的多个模型。然而,由于标注数据集的稀缺性,以及时间建模的复杂性。当前的SFX生成技术在实现高保真音频方面仍然不足。为了解决这些限制,本文介绍了一种新的框架,集成LLM与现有的音效数据库,允许检索,重组和合成的音频根据用户的要求。通过利用这种方法,我们提高了生成的声音效果的多样性和质量,同时消除了额外的录音成本的需要,为声音设计和应用提供了灵活而高效的解决方案。
摘要:Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language processing (NLP) and multimodal learning, with successful applications in text generation and speech synthesis, enabling a deeper understanding and generation of multimodal content. In the field of sound effects (SFX) generation, LLMs have been leveraged to orchestrate multiple models for audio synthesis. However, due to the scarcity of annotated datasets, and the complexity of temproal modeling. current SFX generation techniques still fall short in achieving high-fidelity audio. To address these limitations, this paper introduces a novel framework that integrates LLMs with existing sound effect databases, allowing for the retrieval, recombination, and synthesis of audio based on user requirements. By leveraging this approach, we enhance the diversity and quality of generated sound effects while eliminating the need for additional recording costs, offering a flexible and efficient solution for sound design and application.
【8】 MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification
标题: MGFF-TDNN:一种具有深度可分离模块的多粒度特征融合TDNN模型,用于说话人验证链接:https://arxiv.org/abs/2505.03228
摘要:在说话人确认中,传统的模型往往强调建模长期的上下文特征,以捕获全局说话人特征。然而,这种方法可能会忽略细粒度的声纹信息,其中包含了鲁棒的说话人嵌入所必需的高度判别特征。提出了一种基于多粒度特征融合的模型结构MGFF-TDNN。MGFF-TDNN利用二维深度可分离卷积模块,通过局部特征建模进行增强,作为前端特征提取器,以有效地捕获时频域特征。为了实现全面的多粒度特征融合,我们提出了M-TDNN结构,该结构通过结合时延神经网络和音素级特征池,将全局上下文建模与细粒度特征提取相结合。在VoxCeleb数据集上的实验表明,MGFF-TDNN在说话人确认方面取得了出色的性能,同时在参数和计算资源方面保持了高效。
摘要:In speaker verification, traditional models often emphasize modeling long-term contextual features to capture global speaker characteristics. However, this approach can neglect fine-grained voiceprint information, which contains highly discriminative features essential for robust speaker embeddings. This paper introduces a novel model architecture, termed MGFF-TDNN, based on multi-granularity feature fusion. The MGFF-TDNN leverages a two-dimensional depth-wise separable convolution module, enhanced with local feature modeling, as a front-end feature extractor to effectively capture time-frequency domain features. To achieve comprehensive multi-granularity feature fusion, we propose the M-TDNN structure, which integrates global contextual modeling with fine-grained feature extraction by combining time-delay neural networks and phoneme-level feature pooling. Experiments on the VoxCeleb dataset demonstrate that the MGFF-TDNN achieves outstanding performance in speaker verification while remaining efficient in terms of parameters and computational resources.
【9】 A study on audio synchronous steganography detection and distributed guide inference model based on sliding spectral features and intelligent inference drive
标题: 基于滑动谱特征和智能推理驱动的音频同步隐写检测和分布式引导推理模型研究链接:https://arxiv.org/abs/2505.03193
备注:This paper proposes a novel framework for detecting steganographic content in short video audio streams using sliding spectral features and distributed inference models, combining STFT analysis, entropy-based synchronization, and deep learning-driven decoding strategies
摘要:随着短视频平台在全球传播中的兴起,在音频同步流中嵌入隐写数据成为一种新的隐蔽通信方法。针对传统技术在同步隐写检测上的局限性,本文提出了一种基于中国南海舰队在TikTok上发布的短视频“玉盘”样本的检测和分布式引导重构模型。该方法集成了滑动谱特征提取和智能推理机制。使用具有短时傅立叶变换(STFT)的25 ms滑动窗口来提取主频率轨迹并构造同步帧检测模型(M1),识别帧标志“FFFFFFFFFFFFFF80”。随后的32字节有效载荷由结构化模型(M2)解码以推断分布式制导命令。分析揭示了在具有高度集中的频谱能量的36至45秒音频段中的低熵重复字节序列,从而确认了同步帧的存在。虽然未恢复明文语义,但命令字段布局的一致性表明了军事通信协议的特征。多段拼接模型进一步展示了跨视频嵌入和集中解码能力。该框架验证了滑动谱特征在同步隐写检测中的有效性,并为开放平台上的隐蔽通信分析和战术制导仿真建立了可扩展的推理模型。
摘要:With the rise of short video platforms in global communication, embedding steganographic data in audio synchronization streams has emerged as a new covert communication method. To address the limitations of traditional techniques in detecting synchronized steganography, this paper proposes a detection and distributed guidance reconstruction model based on short video "Yupan" samples released by China's South Sea Fleet on TikTok. The method integrates sliding spectrum feature extraction and intelligent inference mechanisms. A 25 ms sliding window with short-time Fourier transform (STFT) is used to extract the main frequency trajectory and construct the synchronization frame detection model (M1), identifying a frame flag "FFFFFFFFFFFFFFFFFF80". The subsequent 32-byte payload is decoded by a structured model (M2) to infer distributed guidance commands. Analysis reveals a low-entropy, repetitive byte sequence in the 36 to 45 second audio segment with highly concentrated spectral energy, confirming the presence of synchronization frames. Although plaintext semantics are not restored, the consistency in command field layout suggests features of military communication protocols. The multi-segment splicing model further shows cross-video embedding and centralized decoding capabilities. The proposed framework validates the effectiveness of sliding spectral features for synchronized steganography detection and builds an extensible inference model for covert communication analysis and tactical guidance simulation on open platforms.
【10】 CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
标题: CoGenAV:通过对比生成同步的多功能视听表示学习链接:https://arxiv.org/abs/2505.03186
摘要:说话者的嘴唇运动、语音和潜在的语言内容之间的固有同步为改进语音处理任务提供了丰富的信息源,特别是在传统的仅音频系统不稳定的挑战性条件下。我们介绍CoGenAV,一个强大的和数据高效的模型,旨在学习适用于各种语音和视听任务的多功能视听表示。CoGenAV通过优化来自自然视听同步、对比特征对齐和生成文本预测的双重目标进行训练,仅使用来自LRS2数据集的223小时标记数据。这种对比生成同步策略有效地捕获了基本的跨模态相关性。我们展示了多个基准的学习CoGenAV表示的有效性和多功能性。当用于LRS2上的视听语音识别(AVSR)时,这些表示有助于实现1.27的最先进的字错误率(WER)。它们还在视觉语音识别(VSR)方面实现了强大的性能,在LRS2上的WER为22.0,并将嘈杂环境中的性能显著提高了70%以上。此外,CoGenAV表示有利于语音重建任务,提高语音增强和分离的性能,并在主动说话人检测(ASD)等视听同步任务中取得有竞争力的结果。我们的模型将是开源的,以促进学术界和工业界的进一步发展和合作。
摘要:The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions where traditional audio-only systems falter. We introduce CoGenAV, a powerful and data-efficient model designed to learn versatile audio-visual representations applicable across a wide range of speech and audio-visual tasks. CoGenAV is trained by optimizing a dual objective derived from natural audio-visual synchrony, contrastive feature alignment and generative text prediction, using only 223 hours of labeled data from the LRS2 dataset. This contrastive-generative synchronization strategy effectively captures fundamental cross-modal correlations. We showcase the effectiveness and versatility of the learned CoGenAV representations on multiple benchmarks. When utilized for Audio-Visual Speech Recognition (AVSR) on LRS2, these representations contribute to achieving a state-of-the-art Word Error Rate (WER) of 1.27. They also enable strong performance in Visual Speech Recognition (VSR) with a WER of 22.0 on LRS2, and significantly improve performance in noisy environments by over 70%. Furthermore, CoGenAV representations benefit speech reconstruction tasks, boosting performance in Speech Enhancement and Separation, and achieve competitive results in audio-visual synchronization tasks like Active Speaker Detection (ASD). Our model will be open-sourced to facilitate further development and collaboration within both academia and industry.
【11】 Coupling the Heart to Musical Machines
标题: 将心脏与音乐机器相结合链接:https://arxiv.org/abs/2505.03073
摘要:生物反馈最近被用作人机接口(HCI)的一般控制范例。虽然生物反馈,特别是来自呼吸的生物反馈,已经越来越多地被用作新型音乐界面的控制器,新的音乐表达界面(NIME),但社区并没有对心脏给予太多的关注。心脏和呼吸一样,是音乐的一个重要组成部分,有人认为心脏决定了我们对时间的感知,从而间接地决定了我们对音乐的感知。受此启发,我演示了一个基于光电容积描记图(PPG)的NIME控制器,使用心率作为1D控制参数,通过蓝牙无线HCI实时转换声音的质量。我将时间缩放应用于“扭曲”音频缓冲区到声卡,并将这些转换后的音频缓冲区播放给佩戴PPG传感器的听众,创建一个假设的感知生物反馈回路:声音的变化改变心率,改变PPG测量值,从而改变声音。我讨论了声音-心脏-PPG生物反馈回路如何通过1D控制器提供更大的控制和/或各种运动,如何通过生物反馈控制声音回放的空间和/或时间尺度,从而实现性能氛围的可能性,我简要讨论了生成潜在空间作为扩展1D PPG控制空间的可能方式。
摘要:Biofeedback is being used more recently as a general control paradigm for human-computer interfaces (HCIs). While biofeedback especially from breath has seen increasing uptake as a controller for novel musical interfaces, new interfaces for musical expression (NIMEs), the community has not given as much attention to the heart. The heart is just as intimate a part of music as breath and it is argued that the heart determines our perception of time and so indirectly our perception of music. Inspired by this I demonstrate a photoplethysmogram (PPG)-based NIME controller using heart rate as a 1D control parameter to transform the qualities of sounds in real-time over a Bluetooth wireless HCI. I apply time scaling to "warp" audio buffers inbound to the sound card, and play these transformed audio buffers back to the listener wearing the PPG sensor, creating a hypothetical perceptual biofeedback loop: changes in sound change heart rate to change PPG measurements to change sound. I discuss how a sound-heart-PPG biofeedback loop possibly affords greater control and/or variety of movements with a 1D controller, how controlling the space and/or time scale of sound playback with biofeedback makes for possibilities in performance ambience, and I briefly discuss generative latent spaces as a possible way to extend a 1D PPG control space.
【12】 BLAB: Brutally Long Audio Bench
标题: BLAB:残酷的长音频长凳链接:https://arxiv.org/abs/2505.03054
摘要:None
摘要:Developing large audio language models (LMs) capable of understanding diverse spoken interactions is essential for accommodating the multimodal nature of human communication and can increase the accessibility of language technologies across different user populations. Recent work on audio LMs has primarily evaluated their performance on short audio segments, typically under 30 seconds, with limited exploration of long-form conversational speech segments that more closely reflect natural user interactions with these models. We introduce Brutally Long Audio Bench (BLAB), a challenging long-form audio benchmark that evaluates audio LMs on localization, duration estimation, emotion, and counting tasks using audio segments averaging 51 minutes in length. BLAB consists of 833+ hours of diverse, full-length audio clips, each paired with human-annotated, text-based natural language questions and answers. Our audio data were collected from permissively licensed sources and underwent a human-assisted filtering process to ensure task compliance. We evaluate six open-source and proprietary audio LMs on BLAB and find that all of them, including advanced models such as Gemini 2.0 Pro and GPT-4o, struggle with the tasks in BLAB. Our comprehensive analysis reveals key insights into the trade-offs between task difficulty and audio duration. In general, we find that audio LMs struggle with long-form speech, with performance declining as duration increases. They perform poorly on localization, temporal reasoning, counting, and struggle to understand non-phonemic information, relying more on prompts than audio content. BLAB serves as a challenging evaluation framework to develop audio LMs with robust long-form audio understanding capabilities.
机器翻译由腾讯交互翻译提供,仅供参考
