今日论文合集:cs.SD语音21篇,eess.AS音频处理27篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】Music Information Retrieval on Representative Mexican Folk Vocal  Melodies Through MIDI Feature Extraction
标题: 通过MIDI特征提取实现墨西哥代表性民间声乐歌曲的音乐信息检索
链接:https://arxiv.org/abs/2503.24243
作者: Mario Alberto Vallejo Reyes 
备注:10 pages, 5 figures, 2 tables
摘要:本研究分析了具有代表性的墨西哥民间声乐旋律,使用小波特征提取,检查ambitus,音高类熵,和间隔分布。它还探讨了这些功能与歌曲流行度之间的关系,如Spotify播放量所衡量的。本研究使用MATLAB与统计分析软体进行音乐特徴的撷取与统计分析。研究结果显示,在ambitus的显着变化,从8到27的范围内的值,表明不同的作曲风格和声乐需求的体裁。音高类熵的分析展示了广泛的旋律复杂性,阿曼多·曼萨内罗的《Somos Novios》显示了最高的熵,表明了多样而复杂的旋律结构,而像《La Bamba》这样的传统作品表现出较低的熵,表明了更简单,更重复的模式。音程分布主要以主音程(P1)、大、小三度音程(M2、m2)为特征,这表明作曲者偏爱紧密、连续的音程,这有助于旋律的可及性和吸引力。统计分析并没有在范围或熵与Spotify播放次数之间建立显着的相关性。
摘要:This study analyzes representative Mexican folk vocal melodies using MIDI feature extraction, examining ambitus, pitch-class entropy, and interval distribution. It also explores the relationship between these features and song popularity, as measured by Spotify plays. The study employs MATLAB and the MIDI Toolbox for extracting musical features and performing statistical analysis. The findings reveal a significant variation in ambitus, with values ranging from 8 to 27 semitones, indicating a diverse compositional style and vocal demand across the genre. The analysis of pitch-class entropy showcases a broad spectrum of melodic complexity, with Armando Manzanero's `Somos Novios' displaying the highest entropy, suggesting varied and complex melodic structures, while traditional pieces like `La Bamba' exhibit lower entropy, indicating simpler, more repetitive patterns. The interval distribution predominantly features prime intervals (P1), major and minor seconds (M2, m2), pointing to a compositional preference for close, contiguous intervals that contribute to the melodies' accessibility and appeal. Statistical analysis do not establish a significant correlation between the ambitus or entropy and the number of Spotify plays.


【2】 UniSep: Universal Target Audio Separation with Language Models at Scale
标题: UniSep:采用大规模语言模型的通用目标音频分离
链接:https://arxiv.org/abs/2503.23762

作者: Yuanyuan Wang,  Hangting Chen,  Dongchao Yang,  Weiqin Li,  Dan Luo,  Guangzhi Li,  Shan Yang,  Zhiyong Wu,  Helen Meng,  Xixin Wu 
备注:Accepted by ICME 2025
摘要:我们提出了通用目标音频分离(UniSep),解决了不同类型音频的任意混合的分离任务。区别于以往的研究,UniSep进行无限的源域和无限的源号码。我们制定的分离任务作为一个序列到序列的问题,和一个大的语言模型(LLM)是用来模拟的离散潜在空间中的音频序列,利用LLM在处理复杂的混合音频与大规模数据的权力。此外,提出了一种新的预训练策略,利用音频的数据,这减少了大规模的数据模拟的努力,并提高了LLM的能力,以了解音频序列中的信息的一致性和相关性。我们还证明了在音频分离任务中缩放数据集的有效性:我们使用大规模数据(36.5k小时),包括语音,音乐和声音,训练一个通用的目标音频分离模型,不限于特定的域。实验表明,与单任务模型相比,UniSep获得了具有竞争力的主观和客观评价结果。
摘要:We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep is performed on unlimited source domains and unlimited source numbers. We formulate the separation task as a sequence-to-sequence problem, and a large language model (LLM) is used to model the audio sequence in the discrete latent space, leveraging the power of LLM in handling complex mixture audios with large-scale data. Moreover, a novel pre-training strategy is proposed to utilize audio-only data, which reduces the efforts of large-scale data simulation and enhances the ability of LLMs to understand the consistency and correlation of information within audio sequences. We also demonstrate the effectiveness of scaling datasets in an audio separation task: we use large-scale data (36.5k hours), including speech, music, and sound, to train a universal target audio separation model that is not limited to a specific domain. Experiments show that UniSep achieves competitive subjective and objective evaluation results compared with single-task models.


【3】 Evaluation of the Pronunciation of Tajweed Rules Based on DNN as a Step  Towards Interactive Recitation Learning
标题: 基于DNN的Tajweed规则发音评估作为交互式背诵学习的一步
链接:https://arxiv.org/abs/2503.23470

作者: Dim Shaiakhmetov,  Gulnaz Gimaletdinova,  Selcuk Cankurt,  Kadyrmamat Momunov 
摘要:正确的古兰经背诵,坚持Tajweed的规则,是至关重要的,以防止在背诵过程中的错误,并需要大量的努力掌握。传统的教授这些规则的方法受到合格教师的可用性和时间限制的限制。背诵的自动评估可以通过提供及时的反馈和支持独立练习来解决这些挑战。这项研究的重点是开发一个深度学习模型来分类三个Tajweed规则-分离拉伸(Al Mad),紧中午(Ghunnah)和隐藏(Ikhfaa)-使用公开的QDAT数据集,其中包含超过1,500个音频记录。输入数据由来自该数据集的音频记录组成,转换为归一化的梅尔频谱图。对于分类,使用了EfficientNet-B 0架构,并通过挤压和激发注意力机制进行了增强。开发的模型实现了95.35%,99.34%和97.01%的准确率为各自的规则。对学习曲线的分析证实了该模型的鲁棒性和不存在过拟合。所提出的方法具有很高的效率,并为开发交互式教育系统的Tajweed研究铺平了道路。
摘要:Proper recitation of the Quran, adhering to the rules of Tajweed, is crucial for preventing mistakes during recitation and requires significant effort to master. Traditional methods of teaching these rules are limited by the availability of qualified instructors and time constraints. Automatic evaluation of recitation can address these challenges by providing prompt feedback and supporting independent practice. This study focuses on developing a deep learning model to classify three Tajweed rules - separate stretching (Al Mad), tight noon (Ghunnah), and hide (Ikhfaa) - using the publicly available QDAT dataset, which contains over 1,500 audio recordings. The input data consisted of audio recordings from this dataset, transformed into normalized mel-spectrograms. For classification, the EfficientNet-B0 architecture was used, enhanced with a Squeeze-and-Excitation attention mechanism. The developed model achieved accuracy rates of 95.35%, 99.34%, and 97.01% for the respective rules. An analysis of the learning curves confirmed the model's robustness and absence of overfitting. The proposed approach demonstrates high efficiency and paves the way for developing interactive educational systems for Tajweed study.


【4】 Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
标题: 用于高效语音聊天机器人助手的推测性结束转向检测器
链接:https://arxiv.org/abs/2503.23439

作者: Hyunjong Ok,  Suho Yoo,  Jaeho Lee 
备注:Preprint
摘要:由大型语言模型驱动的口语对话系统在理解人类语音和生成适当的口语响应方面表现出了卓越的能力。然而,这些系统与转弯结束检测(ETD)-区分用户转弯完成和犹豫的能力相矛盾。这种限制往往会导致过早或延迟的反应,中断口语对话的流动。在本文中,我们介绍了ETD数据集,这是第一个用于末端转弯检测的公共数据集。ETD数据集由文本到语音模型生成的合成语音数据和从网络源收集的真实语音数据组成。我们还提出了SpeculativeETD,一种新的协作推理框架,平衡效率和准确性,以提高实时ETD在资源受限的环境。我们的方法联合采用了一个轻量级的基于GRU的模型,该模型可以在本地设备上实时快速检测非说话单元,以及一个在服务器上运行的高性能基于Wav 2 vec的模型,以进行更具挑战性的分类,区分转弯结束和仅仅是停顿。实验表明,所提出的SpeculativeETD显着提高ETD的准确性,同时保持所需的计算量低。数据集和代码将在审查后提供。
摘要:Spoken dialogue systems powered by large language models have demonstrated remarkable abilities in understanding human speech and generating appropriate spoken responses. However, these systems struggle with end-turn detection (ETD) -- the ability to distinguish between user turn completion and hesitation. This limitation often leads to premature or delayed responses, disrupting the flow of spoken conversations. In this paper, we introduce the ETD Dataset, the first public dataset for end-turn detection. The ETD dataset consists of both synthetic speech data generated with text-to-speech models and real-world speech data collected from web sources. We also propose SpeculativeETD, a novel collaborative inference framework that balances efficiency and accuracy to improve real-time ETD in resource-constrained environments. Our approach jointly employs a lightweight GRU-based model, which rapidly detects the non-speaking units in real-time on local devices, and a high-performance Wav2vec-based model running on the server to make a more challenging classification of distinguishing turn ends from mere pauses. Experiments demonstrate that the proposed SpeculativeETD significantly improves ETD accuracy while keeping the required computations low. Datasets and code will be available after the review.


【5】 Scaling Auditory Cognition via Test-Time Compute in Audio Language  Models
标题: 通过音频语言模型中的测试时间计算缩放听觉认知
链接:https://arxiv.org/abs/2503.23395

作者: Ting Dang,  Yan Gao,  Hong Jia 
摘要:大型语言模型(LLM)在自然语言处理中表现出了非凡的多功能性,促使最近的努力通过开发音频大型语言模型(Audio LLM)将其多模态功能扩展到语音处理。虽然音频LLM在语音识别和合成等任务中表现出色,但目前尚不清楚它们在面对现实环境带来的听觉认知挑战时的表现,例如音频理解和听力回忆,特别是在存在背景噪声或重叠语音的情况下。与基于文本的LLM不同,基于文本的LLM可以访问大量的文本数据进行预训练,由于模拟真实世界听觉认知场景的数据集有限以及获取听觉认知标签进行训练的挑战,使用不同的听觉认知场景重新训练音频LLM是困难的。虽然测试时间计算(TTC)方法已被证明可以增强基于文本的LLM在推理过程中的能力,但一个关键的挑战在于设计这些TTC方法来提高音频LLM的听觉能力。本研究旨在解决这两个研究空白:i)探索音频LLM的听觉认知能力,以及ii)使用TTC方法增强其能力。我们已经调查了五种不同的音频LLM听觉认知使用的\textit{自我收集}数据库,并提出了五个TTC方法,以提高听觉认知能力的推理。我们的研究结果表明,音频LLM性能下降更具挑战性的听觉认知任务。所提出的TTC方法显著增强了认知听觉能力,推动了更具适应性和弹性的音频LLM的开发,用于辅助听力设备,基于语音的AI助手和通信技术等实际应用。
摘要:Large language models (LLMs) have shown exceptional versatility in natural language processing, prompting recent efforts to extend their multimodal capabilities to speech processing through the development of audio large language models (Audio LLMs). While Audio LLMs excel in tasks such as speech recognition and synthesis, it remains unclear how they perform when faced with the auditory cognitive challenges posed by real-world environments, such as audio comprehension and listening recall, particularly in the presence of background noise or overlapping speech. Unlike text-based LLMs, which have access to vast amounts of text data for pre-training, retraining Audio LLMs with diverse auditory cognitive scenes is difficult due to the limited datasets that simulate real-world auditory cognitive scenarios and the challenge of acquiring auditory cognitive labels for training. While test-time compute (TTC) methods have been shown to enhance the capabilities of text-based LLMs during inference, a key challenge lies in designing these TTC methods to improve the auditory capabilities of Audio LLMs. This study aims to address these two research gaps by: i) exploring the auditory cognitive capabilities of Audio LLMs, and ii) enhancing their capabilities using TTC approaches. We have investigated five different Audio LLMs for auditory cognition using a \textit{self-collected} database and have proposed five TTC approaches to enhance auditory cognitive capabilities during inference. Our findings reveal that Audio LLMs performance decreases in more challenging auditory cognitive tasks. The proposed TTC approaches significantly enhance cognitive auditory capabilities, advancing the development of more adaptable and resilient Audio LLMs for practical applications such as assistive listening devices, voice-based AI assistants, and communication technologies.


【6】 D3-Guard: Acoustic-based Drowsy Driving Detection Using Smartphones
标题: D3-Guard:使用智能手机基于声学的昏昏欲驾驶检测
链接:https://arxiv.org/abs/2503.23393

作者: Yadong Xie,  Fan Li,  Yue Wu,  Song Yang,  Yu Wang 
备注:IEEE INFOCOM 2019-IEEE Conference on Computer Communications
摘要:近年来,随着汽车数量的快速增长,行车安全越来越受到人们的关注。疲劳驾驶是对行车安全的最大威胁之一。因此,一个简单但强大的系统,可以检测困倦驾驶与商业现成的设备(如智能手机)是非常必要的。出于这个动机,我们探索纯粹使用嵌入在智能手机中的声学传感器来检测困倦驾驶的可行性。本文首先研究了驾驶员在驾驶过程中的疲劳行为,发现了驾驶员在驾驶过程中的三种典型疲劳行为,即点头、打哈欠和操作方向盘时所引起的多普勒频移的一些独特模式。然后,我们通过对从真实驾驶环境中收集的驾驶数据进行实证分析来验证我们的重要发现。我们进一步提出了一种基于嵌入在智能手机中的音频设备的实时昏昏欲睡驾驶检测系统(D3-Guard)。为了提高系统的性能,我们采用了一种基于欠采样技术和FFT的有效特征提取方法,并精心设计了一种基于LSTM网络的高精度疲劳驾驶早期检测器。通过在真实驾驶环境中对5名志愿者驾驶员的大量实验,我们的系统可以实时区分困倦驾驶行为,平均总准确率为93.31%。超过80%的困倦驾驶动作可以在动作持续时间的前70%内检测到。
摘要:Since the number of cars has grown rapidly in recent years, driving safety draws more and more public attention. Drowsy driving is one of the biggest threatens to driving safety. Therefore, a simple but robust system that can detect drowsy driving with commercial off-the-shelf devices (such as smartphones) is very necessary. With this motivation, we explore the feasibility of purely using acoustic sensors embedded in smartphones to detect drowsy driving. We first study characteristics of drowsy driving, and find some unique patterns of Doppler shift caused by three typical drowsy behaviors, i.e. nodding, yawning and operating steering wheel. We then validate our important findings through empirical analysis of the driving data collected from real driving environments. We further propose a real-time Drowsy Driving Detection system (D3-Guard) based on audio devices embedded in smartphones. In order to improve the performance of our system, we adopt an effective feature extraction method based on undersampling technique and FFT, and carefully design a high-accuracy detector based on LSTM networks for the early detection of drowsy driving. Through extensive experiments with 5 volunteer drivers in real driving environments, our system can distinguish drowsy driving actions with an average total accuracy of 93.31% in real-time. Over 80% drowsy driving actions can be detected within first 70% of action duration.


【7】 HearSmoking: Smoking Detection in Driving Environment via Acoustic  Sensing on Smartphones
标题: HearSmoking:通过智能手机上的声学传感检测驾驶环境中的吸烟
链接:https://arxiv.org/abs/2503.23391

作者: Yadong Xie,  Fan Li,  Yue Wu,  Song Yang,  Yu Wang 
备注:IEEE Transactions on Mobile Computing ( Volume: 21, Issue: 8, 01 August 2022)
摘要:近年来,由于汽车数量的快速增长,驾驶安全引起了公众的广泛关注。吸烟是对驾驶安全的威胁之一,但往往被司机忽视。现有的吸烟检测工作要么以接触方式工作,要么需要额外的设备。这促使我们探索使用智能手机来检测驾驶环境中的吸烟事件的实用性。在本文中,我们提出了一种名为HearSmoking的吸烟检测系统,该系统仅使用智能手机上的声学传感器来提高驾驶安全性。在调查了驾驶员的典型吸烟习惯(包括手部运动和胸部波动)后,我们设计了一个由扬声器发出并由麦克风接收的声学信号。我们计算接收信号的相对相关系数,以获得手和胸部的运动模式。处理后的数据被发送到一个训练好的卷积神经网络进行分类的手运动。我们还设计了一种同时检测呼吸的方法。为了提高系统性能,我们进一步分析了复合吸烟运动的周期性。通过在真实驾驶环境中的大量实验,HearSmoking实时检测吸烟事件的平均总准确率为93.44%。
摘要:Driving safety has drawn much public attention in recent years due to the fast-growing number of cars. Smoking is one of the threats to driving safety but is often ignored by drivers. Existing works on smoking detection either work in contact manner or need additional devices. This motivates us to explore the practicability of using smartphones to detect smoking events in driving environment. In this paper, we propose a cigarette smoking detection system, named HearSmoking, which only uses acoustic sensors on smartphones to improve driving safety. After investigating typical smoking habits of drivers, including hand movement and chest fluctuation, we design an acoustic signal to be emitted by the speaker and received by the microphone. We calculate Relative Correlation Coefficient of received signals to obtain movement patterns of hands and chest. The processed data is sent into a trained Convolutional Neural Network for classification of hand movement. We also design a method to detect respiration at the same time. To improve system performance, we further analyse the periodicity of the composite smoking motion. Through extensive experiments in real driving environments, HearSmoking detects smoking events with an average total accuracy of 93.44 percent in real-time.


【8】 HearFit+: Personalized Fitness Monitoring via Audio Signals on Smart  Speakers
标题: HearFit+:通过智能扬声器上的音频信号进行个性化健身监测
链接:https://arxiv.org/abs/2503.23387

作者: Yadong Xie,  Fan Li,  Yue Wu,  Yu Wang 
备注:IEEE Transactions on Mobile Computing ( Volume: 22, Issue: 5, 01 May 2023)
摘要:健身有助于强健肌肉,增强抗病能力,改善体形。现在,很多人选择在家里/办公室锻炼,而不是在健身房,因为缺乏时间。但是,如果没有专业的指导,他们很难获得良好的健身效果。出于这一动机,我们提出了第一个个性化的健身监测系统,HearFit+,在家庭/办公室使用智能扬声器。我们探讨了利用声学传感监测健身的可行性。设计了一种基于多普勒频移的适应度检测方法,并采用短时能量对适应度动作进行分割。基于深度学习,HearFit+可以同时进行健身分类和用户识别。结合增量学习,用户可以轻松添加新操作。我们设计了4个评估指标(即,持续时间、强度、连续性和平滑度),以帮助用户提高健身效果。通过对12名志愿者10种健身类型的9,000多个动作的大量实验,HearFit+在健身分类上的平均准确率达到96.13%,在用户识别上的准确率达到91%。所有志愿者都确认HearFit+可以帮助改善各种环境下的健身效果。
摘要:Fitness can help to strengthen muscles, increase resistance to diseases, and improve body shape. Nowadays, a great number of people choose to exercise at home/office rather than at the gym due to lack of time. However, it is difficult for them to get good fitness effects without professional guidance. Motivated by this, we propose the first personalized fitness monitoring system, HearFit+, using smart speakers at home/office. We explore the feasibility of using acoustic sensing to monitor fitness. We design a fitness detection method based on Doppler shift and adopt the short time energy to segment fitness actions. Based on deep learning, HearFit+ can perform fitness classification and user identification at the same time. Combined with incremental learning, users can easily add new actions. We design 4 evaluation metrics (i.e., duration, intensity, continuity, and smoothness) to help users to improve fitness effects. Through extensive experiments including over 9,000 actions of 10 types of fitness from 12 volunteers, HearFit+ can achieve an average accuracy of 96.13% on fitness classification and 91% accuracy for user identification. All volunteers confirm that HearFit+ can help improve the fitness effect in various environments.


【9】 JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical  Spatio-Temporal Prior Synchronization
标题: JavisDiT:具有分层时空先验同步的联合音频-视频扩散Transformer
链接:https://arxiv.org/abs/2503.23377

作者: Kai Liu,  Wei Li,  Lai Chen,  Shengqiong Wu,  Yanhao Zheng,  Jiayi Ji,  Fan Zhou,  Rongxin Jiang,  Jiebo Luo,  Hao Fei,  Tat-Seng Chua 
备注:Work in progress. Homepage: this https URL
摘要:本文介绍了一种用于同步音视频生成的新型联合音视频扩散Transformer--JavisDiT。建立在强大的扩散Transformer(DiT)架构,JavisDiT是能够产生高品质的音频和视频内容,同时从开放式的用户提示。为了确保最佳的同步,我们引入了一个细粒度的时空对齐机制,通过分层时空同步先验(HiST-Sypo)估计。该模块提取全局和细粒度的时空先验,指导视觉和听觉组件之间的同步。此外,我们提出了一个新的基准,JavisBench,由10,140个高质量的文本字幕声音视频组成,跨越不同的场景和复杂的现实世界场景。此外,我们专门设计了一个强大的度量标准,用于评估在现实世界的复杂内容中生成的音频-视频对之间的同步。实验结果表明,JavisDiT通过确保高质量的生成和精确的同步,显着优于现有的方法,为JAVG任务设定了新的标准。我们的代码、模型和数据集将在https://javisdit.github.io/上公开。
摘要:This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Transformer designed for synchronized audio-video generation (JAVG). Built upon the powerful Diffusion Transformer (DiT) architecture, JavisDiT is able to generate high-quality audio and video content simultaneously from open-ended user prompts. To ensure optimal synchronization, we introduce a fine-grained spatio-temporal alignment mechanism through a Hierarchical Spatial-Temporal Synchronized Prior (HiST-Sypo) Estimator. This module extracts both global and fine-grained spatio-temporal priors, guiding the synchronization between the visual and auditory components. Furthermore, we propose a new benchmark, JavisBench, consisting of 10,140 high-quality text-captioned sounding videos spanning diverse scenes and complex real-world scenarios. Further, we specifically devise a robust metric for evaluating the synchronization between generated audio-video pairs in real-world complex content. Experimental results demonstrate that JavisDiT significantly outperforms existing methods by ensuring both high-quality generation and precise synchronization, setting a new standard for JAVG tasks. Our code, model, and dataset will be made publicly available at https://javisdit.github.io/.


【10】 Joint Source-Environment Adaptation for Deep Learning-Based Underwater  Acoustic Source Ranging
标题: 基于深度学习的水下声波源距离联合源环境自适应
链接:https://arxiv.org/abs/2503.23262

作者: Dariush Kari,  Andrew C. Singer 
摘要:在本文中,我们提出了一种方法,使预先训练的基于深度学习的水下声学定位模型适应新的环境。我们使用无监督域自适应来提高模型的泛化性能,即,使用无监督损失,微调预训练的网络参数,而无需访问目标环境的任何标签或用于预训练模型的任何数据。该方法通过将预训练模型预测与基于接收信号能量(取决于源)的几乎独立的估计相耦合来改进预训练模型预测。我们展示了这种方法在类似于SWellEx-96实验的环境中对Bellhop生成的数据的有效性,该实验被来自KAM 11实验的真实海洋噪声污染。
摘要:In this paper, we propose a method to adapt a pre-trained deep-learning-based model for underwater acoustic localization to a new environment. We use unsupervised domain adaptation to improve the generalization performance of the model, i.e., using an unsupervised loss, fine-tune the pre-trained network parameters without access to any labels of the target environment or any data used to pre-train the model. This method improves the pre-trained model prediction by coupling that with an almost independent estimation based on the received signal energy (that depends on the source). We show the effectiveness of this approach on Bellhop generated data in an environment similar to that of the SWellEx-96 experiment contaminated with real ocean noise from the KAM11 experiment.


【11】 Mismatch-Robust Underwater Acoustic Localization Using A Differentiable  Modular Forward Model
标题: 使用可微模块正演模型的匹配鲁棒性水下声学定位
链接:https://arxiv.org/abs/2503.23260

作者: Dariush Kari,  Yongjie Zhuang,  Andrew C. Singer 
摘要:本文研究了环境失配条件下的水声定位问题。特别是,我们在基于梯度的优化框架中利用预先训练的声波传播神经网络来估计源位置。为了减轻训练数据和测试数据之间的不匹配的影响,我们同时优化网络的权重在推理时,并提供条件下,该方法是有效的。此外,我们在前向模型中引入了物理启发的模块化,使我们能够以端到端的训练方式学习多路径结构的路径长度,而无需访问特定的路径标签。我们调查的有效性的假设在一个简单而说明性的环境模型。
摘要:In this paper, we study the underwater acoustic localization in the presence of environmental mismatch. Especially, we exploit a pre-trained neural network for the acoustic wave propagation in a gradient-based optimization framework to estimate the source location. To alleviate the effect of mismatch between the training data and the test data, we simultaneously optimize over the network weights at the inference time, and provide conditions under which this method is effective. Moreover, we introduce a physics-inspired modularity in the forward model that enables us to learn the path lengths of the multipath structure in an end-to-end training manner without access to the specific path labels. We investigate the validity of the assumptions in a simple yet illustrative environment model.


【12】 Joint Source-Environment Adaptation of Data-Driven Underwater Acoustic  Source Ranging Based on Model Uncertainty
标题: 基于模型不确定性的数据驱动水下声源联合自适应测距
链接:https://arxiv.org/abs/2503.23258

作者: Dariush Kari,  Hari Vishnu,  Andrew C. Singer 
摘要:使预训练的深度学习模型适应新的未知环境是水声定位中的一项艰巨挑战。我们发现,尽管预训练模型的性能受到训练数据和测试数据之间不匹配的影响,但在不匹配较多的环境中,它们通常表现出较高的“隐含不确定性”。利用隐含不确定性的概念,我们将测试样本划分为更确定和更不确定的集合,并使用确定的样本来实现估计方法,以改善不确定样本的标记,这有助于适应模型。我们使用一种有效的方法来量化模型预测的不确定性,并使用一种创新的方法来使预训练的模型在测试时适应看不见的水下环境。这消除了对来自目标环境或原始训练数据的标记数据的需要。通过基于接收信号能量对独立估计进行积分来增强这种自适应。我们使用真实的实验数据以及由模型生成的信号与真实海洋噪声组成的合成数据来验证该方法。结果表明,模型预测准确性得到了显着提高,强调了该方法在多样化、嘈杂和未知环境中增强水下声学定位的潜力。
摘要:Adapting pre-trained deep learning models to new and unknown environments is a difficult challenge in underwater acoustic localization. We show that although pre-trained models have performance that suffers from mismatch between the training and test data, they generally exhibit a higher ``implied uncertainty'' in environments where there is more mismatch. Leveraging this notion of implied uncertainty, we partition the test samples into more certain and less certain sets, and implement an estimation method using the certain samples to improve the labeling for uncertain samples, which helps to adapt the model. We use an efficient method to quantify model prediction uncertainty, and an innovative approach to adapt a pre-trained model to unseen underwater environments at test time. This eliminates the need for labeled data from the target environment or the original training data. This adaptation is enhanced by integrating an independent estimate based on the received signal energy. We validate the approach extensively using real experimental data, as well as synthetic data consisting of model-generated signals with real ocean noise. The results demonstrate significant improvements in model prediction accuracy, underscoring the potential of the method to enhance underwater acoustic localization in diverse, noisy, and unknown environments.


【13】 CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with  LLM-Powered Text Description Sourcing and Mining
标题: CrossMuSim:一个基于LLM-Powered文本描述源和挖掘的跨模态音乐相似性检索框架
链接:https://arxiv.org/abs/2503.23128

作者: Tristan Tsoi,  Jiajun Deng,  Yaolong Ju,  Benno Weck,  Holger Kirchhoff,  Simon Lui 
备注:Accepted by ICME2025
摘要:音乐相似性检索是管理和探索流媒体平台中大型收藏中相关内容的基础。本文提出了一种新的跨模态对比学习框架,利用文本描述的开放性来指导音乐相似性建模,解决了传统单模态方法在捕捉复杂音乐关系方面的局限性。为了克服高质量的文本-音乐配对数据的稀缺性,本文介绍了一种结合在线抓取和基于LLM的提示的双源数据获取方法,其中精心设计的提示利用LLM的综合音乐知识来生成上下文丰富的描述。在华为Music流媒体平台上,通过客观指标、主观评价和真实A/B测试,验证了该框架在性能上的显著提升。
摘要:Music similarity retrieval is fundamental for managing and exploring relevant content from large collections in streaming platforms. This paper presents a novel cross-modal contrastive learning framework that leverages the open-ended nature of text descriptions to guide music similarity modeling, addressing the limitations of traditional uni-modal approaches in capturing complex musical relationships. To overcome the scarcity of high-quality text-music paired data, this paper introduces a dual-source data acquisition approach combining online scraping and LLM-based prompting, where carefully designed prompts leverage LLMs' comprehensive music knowledge to generate contextually rich descriptions. Exten1sive experiments demonstrate that the proposed framework achieves significant performance improvements over existing benchmarks through objective metrics, subjective evaluations, and real-world A/B testing on the Huawei Music streaming platform.


【14】 Teaching LLMs Music Theory with In-Context Learning and Chain-of-Thought  Prompting: Pedagogical Strategies for Machines
标题: 教学LLM音乐理论与上下文学习和思想链的修正:机器的教学策略
链接:https://arxiv.org/abs/2503.22853

作者: Liam Pond,  Ichiro Fujinaga 
备注:11 pages, 4 figures, 3 tables. Published in Volume 1 of the Proceedings of the 17th International Conference on Computer Supported Music Education (CSME 2025). Presented on 3 April 2025 in Porto, Portugal
摘要:这项研究评估了ChatGPT、Claude和Gemini等大型语言模型(LLM)通过上下文学习和思维链提示来学习音乐理论中的概念的基本能力。使用精心设计的提示(上下文学习)和一步一步的工作示例(思想链提示),我们探讨了LLM如何教授越来越复杂的材料,以及人类学习者的教学策略如何转化为教育机器。演奏使用加拿大皇家音乐学院(RCM)官方6级考试的问题进行评估,该考试涵盖了广泛的主题,包括音程和和弦识别,关键检测,节奏分类和韵律分析。此外,我们评估了各种音乐编码格式对这些任务的适用性(ABC,Humbrum,MEI,MusicXML)。所有实验都在有和没有上下文提示的情况下运行。结果表明,在没有上下文的情况下,具有MEI的ChatGPT表现最好,为52%,而在有上下文的情况下,具有MEI的Claude表现最好,为75%。未来的工作将进一步完善提示,并扩大到涵盖更先进的音乐理论概念。这项研究有助于更广泛地理解教学LLM,并为教育工作者,学生和AI音乐工具开发人员提供应用程序。
摘要:This study evaluates the baseline capabilities of Large Language Models (LLMs) like ChatGPT, Claude, and Gemini to learn concepts in music theory through in-context learning and chain-of-thought prompting. Using carefully designed prompts (in-context learning) and step-by-step worked examples (chain-of-thought prompting), we explore how LLMs can be taught increasingly complex material and how pedagogical strategies for human learners translate to educating machines. Performance is evaluated using questions from an official Canadian Royal Conservatory of Music (RCM) Level 6 examination, which covers a comprehensive range of topics, including interval and chord identification, key detection, cadence classification, and metrical analysis. Additionally, we evaluate the suitability of various music encoding formats for these tasks (ABC, Humdrum, MEI, MusicXML). All experiments were run both with and without contextual prompts. Results indicate that without context, ChatGPT with MEI performs the best at 52%, while with context, Claude with MEI performs the best at 75%. Future work will further refine prompts and expand to cover more advanced music theory concepts. This research contributes to the broader understanding of teaching LLMs and has applications for educators, students, and developers of AI music tools alike.


【15】 Dual Audio-Centric Modality Coupling for Talking Head Generation
标题: 用于说话的头部生成的双音频中心模式耦合
链接:https://arxiv.org/abs/2503.22728

作者: Ao Fu,  Ziqi Ni,  Yi Zhou 
备注:9 pages, 4 figures
摘要:音频驱动的说话头部视频的生成是计算机视觉和图形学中的一个关键挑战,其应用于虚拟化身和数字媒体。传统方法通常难以捕捉音频和面部动态之间的复杂交互,导致嘴唇同步和视觉质量问题。在本文中,我们提出了一种新的NeRF为基础的框架,双音频中心的模态耦合(DAMC),它有效地集成了音频输入的内容和动态功能。通过利用双编码器结构,DAMC通过内容感知编码器捕获语义内容,并通过动态同步编码器确保精确的视觉同步。这些功能融合使用交叉同步融合模块(CSFM),增强内容表示和唇同步。大量的实验表明,我们的方法优于现有的最先进的方法在关键指标,如唇同步精度和图像质量,表现出强大的泛化在各种音频输入,包括合成语音从文本到语音(TTS)系统。我们的研究结果提供了一个有前途的解决方案,高品质,音频驱动的说话头生成,并提出了一个可扩展的方法来创建逼真的说话头。
摘要:The generation of audio-driven talking head videos is a key challenge in computer vision and graphics, with applications in virtual avatars and digital media. Traditional approaches often struggle with capturing the complex interaction between audio and facial dynamics, leading to lip synchronization and visual quality issues. In this paper, we propose a novel NeRF-based framework, Dual Audio-Centric Modality Coupling (DAMC), which effectively integrates content and dynamic features from audio inputs. By leveraging a dual encoder structure, DAMC captures semantic content through the Content-Aware Encoder and ensures precise visual synchronization through the Dynamic-Sync Encoder. These features are fused using a Cross-Synchronized Fusion Module (CSFM), enhancing content representation and lip synchronization. Extensive experiments show that our method outperforms existing state-of-the-art approaches in key metrics such as lip synchronization accuracy and image quality, demonstrating robust generalization across various audio inputs, including synthetic speech from text-to-speech (TTS) systems. Our results provide a promising solution for high-quality, audio-driven talking head generation and present a scalable approach for creating realistic talking heads.


【16】 Risk-Calibrated Affective Speech Recognition via Conformal Coverage  Guarantees: A Stochastic Calibrative Framework for Emergent Uncertainty  Quantification
标题: 通过保形覆盖保证进行风险校准的情感语音识别:紧急不确定性量化的随机校准框架
链接:https://arxiv.org/abs/2503.22712

作者: Zijun Jia 
摘要:极端驾驶员情绪引发的交通安全挑战凸显了对可靠情绪识别系统的迫切需求。语音情感识别中的传统深度学习方法存在过拟合和校准不佳的置信度估计。我们提出了一个整合共形预测(CP)和风险控制的框架,使用通过预训练的卷积神经网络处理的Mel谱图特征。我们的主要创新是开发了一个不一致性评分,该评分可以精确地衡量分类器的预测与给定输入的一致程度。通过校准样本,我们计算这个分数,并根据用户指定的风险水平$\alpha$推导出一个统计上严格的阈值,构建具有可证明覆盖保证($\geq 1-\alpha$)的预测集。风险控制框架通过可定制的损失函数实现特定于任务的自适应,动态调整预测集大小,同时保持覆盖率保证。IEMOCAP和TESS的跨数据集实验表明:1)严格的覆盖保证,2)平均预测集大小(APSS)和$\alpha$之间显着的负相关性,揭示了在高风险条件下降低模型的不确定性。我们进一步提出APSS作为一种新的度量分类不确定性。该方法提高了语音情感识别的可靠性,可直接应用于智能交通系统和实时情感监控。
摘要:Traffic safety challenges arising from extreme driver emotions highlight the urgent need for reliable emotion recognition systems. Traditional deep learning approaches in speech emotion recognition suffer from overfitting and poorly calibrated confidence estimates. We propose a framework integrating Conformal Prediction (CP) and Risk Control,using Mel-spectrogram features processed through a pre-trained convolutional neural network. Our key innovation is the development of a nonconformity score that heuristically measures how closely a classifier's predictions align with given inputs. Through calibration samples, we compute this score and derive a statistically rigorous threshold based on user-specified risk level $\alpha$, constructing prediction sets with provable coverage guarantees ($\geq 1-\alpha$). The Risk Control framework enables task-specific adaptation through customizable loss functions, dynamically adjusting prediction set sizes while maintaining coverage guarantees. Cross-dataset experiments on IEMOCAP and TESS demonstrate: 1) Strict coverage guarantee, 2) Significant negative correlation between Average Prediction Set Size (APSS) and $\alpha$, revealing reduced model uncertainty under high-risk conditions. We further propose APSS as a novel metric for evaluating classification uncertainty. This approach enhances speech emotion recognition reliability, with direct applications in intelligent transportation systems and real-time emotion monitoring.


【17】 Modeling speech emotion with label variance and analyzing performance  across speakers and unseen acoustic conditions
标题: 利用标签方差建模语音情感并分析说话者和看不见的声学条件之间的性能
链接:https://arxiv.org/abs/2503.22711

作者: Vikramjit Mitra,  Amrit Romana,  Dung T. Tran,  Erdrin Azemi 
备注:11 pages, 5 figures
摘要:自发语音情感数据通常包含感知等级,其中分级者在听完语音文件后分配情感分数。这样的感知等级由于分级者意见变化而在标签中引入不确定性。评分员的差异是通过使用共识评分作为地面实况来解决的,其中具有最高投票的情感被选择。共识等级没有考虑模糊的情况下,语音样本可能包含多种情绪,通过分级意见的不确定性捕获。我们证明,使用的概率密度函数的情绪等级作为目标,而不是常用的共识等级,提供更好的性能基准评估集相比,在文献中报道的结果。我们发现,显着性驱动的基础模型(FM)表示选择有助于训练一个国家的最先进的语音情感模型的维度和类别的情感识别。比较从不同FM获得的表示,我们观察到,专注于整体测试集性能可能具有欺骗性,因为它无法揭示模型在说话者和性别之间的泛化能力。我们表明,跨多个测试集的性能评估和跨性别和扬声器的性能分析是有用的评估情感模型的有用性。最后,我们证明了标签的不确定性和数据倾斜对模型评估提出了挑战,而不是使用最佳假设,考虑2或3个最佳假设是有用的。
摘要:Spontaneous speech emotion data usually contain perceptual grades where graders assign emotion score after listening to the speech files. Such perceptual grades introduce uncertainty in labels due to grader opinion variation. Grader variation is addressed by using consensus grades as groundtruth, where the emotion with the highest vote is selected. Consensus grades fail to consider ambiguous instances where a speech sample may contain multiple emotions, as captured through grader opinion uncertainty. We demonstrate that using the probability density function of the emotion grades as targets instead of the commonly used consensus grades, provide better performance on benchmark evaluation sets compared to results reported in the literature. We show that a saliency driven foundation model (FM) representation selection helps to train a state-of-the-art speech emotion model for both dimensional and categorical emotion recognition. Comparing representations obtained from different FMs, we observed that focusing on overall test-set performance can be deceiving, as it fails to reveal the models generalization capacity across speakers and gender. We demonstrate that performance evaluation across multiple test-sets and performance analysis across gender and speakers are useful in assessing usefulness of emotion models. Finally, we demonstrate that label uncertainty and data-skew pose a challenge to model evaluation, where instead of using the best hypothesis, it is useful to consider the 2- or 3-best hypotheses.


【18】 Exploring In-Context Learning Capabilities of ChatGPT for Pathological  Speech Detection
标题: 探索ChatGPT用于病理性语音检测的上下文学习能力
链接:https://arxiv.org/abs/2503.23873

作者: Mahdi Amiri,  Hatef Otroshi Shahreza,  Ina Kodrasi 
备注:submitted to EUSIPCO 2025
摘要:自动病理语音检测方法已经显示出有希望的结果,作为潜在的诊断工具与昂贵的传统方法一起受到关注。虽然这些方法可以实现高准确性,但它们缺乏可解释性,限制了它们在临床实践中的适用性。在本文中,我们研究了使用多模态大语言模型(LLM),特别是ChatGPT-4 o,在一个Few-Shot的上下文中学习设置自动病理语音检测。实验结果表明,这种方法不仅提供了有前途的性能,但也提供了解释其决策,增强模型的可解释性。为了进一步了解其有效性,我们进行了一项消融研究,以分析不同因素(如输入类型和系统提示)对最终结果的影响。我们的研究结果突出了多模态LLM在自动病理语音检测中的进一步探索和进步的潜力。
摘要:Automatic pathological speech detection approaches have shown promising results, gaining attention as potential diagnostic tools alongside costly traditional methods. While these approaches can achieve high accuracy, their lack of interpretability limits their applicability in clinical practice. In this paper, we investigate the use of multimodal Large Language Models (LLMs), specifically ChatGPT-4o, for automatic pathological speech detection in a few-shot in-context learning setting. Experimental results show that this approach not only delivers promising performance but also provides explanations for its decisions, enhancing model interpretability. To further understand its effectiveness, we conduct an ablation study to analyze the impact of different factors, such as input type and system prompts, on the final results. Our findings highlight the potential of multimodal LLMs for further exploration and advancement in automatic pathological speech detection.


【19】 SupertonicTTS: Towards Highly Scalable and Efficient Text-to-Speech  System
标题: SupertonicTTC:迈向高度可扩展和高效的文本到语音系统
链接:https://arxiv.org/abs/2503.23108

作者: Hyeongju Kim,  Jinhyeok Yang,  Yechan Yu,  Seunghun Ji,  Jacob Morton,  Frederik Bous,  Joon Byun,  Juheon Lee 
备注:19 pages, preprint
摘要:我们提出了一种新的文本到语音(TTS)系统,即SupertonicTTS,提高语音合成的可扩展性和效率。SupertonicTTS由三个组件组成:用于连续潜在表示的语音自动编码器,利用流匹配进行文本到潜在映射的文本到潜在模块,以及话语级持续时间预测器。为了实现轻量级架构,我们采用了低维的潜在空间,潜在的时间压缩和ConvNeXt块。我们通过直接在原始字符级文本上操作并采用交叉注意进行文本语音对齐来进一步简化TTS管道,从而消除了对字素到音素(G2P)模块和外部对齐器的需要。此外,我们引入了上下文共享批量扩展,加速损失收敛和稳定的文本语音对齐。实验结果表明,SupertonicTTS实现竞争力的性能,同时显着降低建筑的复杂性和计算开销相比,当代TTS模型。演示SupertonicTTS功能的音频样本可在https://supertonictts.github.io/上获得。
摘要:We present a novel text-to-speech (TTS) system, namely SupertonicTTS, for improved scalability and efficiency in speech synthesis. SupertonicTTS is comprised of three components: a speech autoencoder for continuous latent representation, a text-to-latent module leveraging flow-matching for text-to-latent mapping, and an utterance-level duration predictor. To enable a lightweight architecture, we employ a low-dimensional latent space, temporal compression of latents, and ConvNeXt blocks. We further simplify the TTS pipeline by operating directly on raw character-level text and employing cross-attention for text-speech alignment, thus eliminating the need for grapheme-to-phoneme (G2P) modules and external aligners. In addition, we introduce context-sharing batch expansion that accelerates loss convergence and stabilizes text-speech alignment. Experimental results demonstrate that SupertonicTTS achieves competitive performance while significantly reducing architectural complexity and computational overhead compared to contemporary TTS models. Audio samples demonstrating the capabilities of SupertonicTTS are available at: https://supertonictts.github.io/.


【20】 Chirp Localization via Fine-Tuned Transformer Model: A Proof-of-Concept  Study
标题: 基于精细调谐Transformer模型的啁啾定位:概念验证研究
链接:https://arxiv.org/abs/2503.22713

作者: Nooshin Bahador,  Milad Lankarany 
备注:19 pages, 8 figures
摘要:频谱图是时频信号分析的关键,广泛应用于音频处理和计算神经科学。脑电图(EEG)频谱图(由线性或指数频率扫描标记)中的Chirp样模式是癫痫发作动力学的关键生物标志物,但缺乏用于其检测、定位和特征提取的自动化工具。这项研究通过对合成频谱图的Vision Transformer(ViT)模型进行微调来弥合这一差距,并通过低秩自适应(LoRA)来增强适应性。我们生成了100000合成光谱与啁啾参数,创造了第一个大规模的基准啁啾定位。这些频谱图使用线性或指数频率扫描、高斯噪声和平滑来模拟神经啁啾。一个ViT模型,适用于回归,预测啁啾参数。LoRA微调了注意力层,使预训练的骨干能够有效更新。训练使用MSE损失和AdamW优化器,具有学习率调度器和早期停止以抑制过度拟合。仅针对三个特征:Chirp开始时间(开始时间)、Chirp开始频率(开始频率)和Chirp结束频率(偏移频率)。通过预测和实际标签之间的Pearson相关性评估性能。结果显示出强烈的一致性:线性调频脉冲开始时间的相关性为0.9841,具有稳定的推断时间(137至140秒)和误差分布的最小偏差。这种方法提供了一个工具,啁啾分析脑电时频表示,填补了一个关键的方法空白。
摘要:Spectrograms are pivotal in time-frequency signal analysis, widely used in audio processing and computational neuroscience. Chirp-like patterns in electroencephalogram (EEG) spectrograms (marked by linear or exponential frequency sweep) are key biomarkers for seizure dynamics, but automated tools for their detection, localization, and feature extraction are lacking. This study bridges this gap by fine-tuning a Vision Transformer (ViT) model on synthetic spectrograms, augmented with Low-Rank Adaptation (LoRA) to boost adaptability. We generated 100000 synthetic spectrograms with chirp parameters, creating the first large-scale benchmark for chirp localization. These spectrograms mimic neural chirps using linear or exponential frequency sweep, Gaussian noise, and smoothing. A ViT model, adapted for regression, predicted chirp parameters. LoRA fine-tuned the attention layers, enabling efficient updates to the pre-trained backbone. Training used MSE loss and the AdamW optimizer, with a learning rate scheduler and early stopping to curb overfitting. Only three features were targeted: Chirp Start Time (Onset Time), Chirp Start Frequency (Onset Frequency), and Chirp End Frequency (Offset Frequency). Performance was evaluated via Pearson correlation between predicted and actual labels. Results showed strong alignment: 0.9841 correlation for chirp start time, with stable inference times (137 to 140s) and minimal bias in error distributions. This approach offers a tool for chirp analysis in EEG time-frequency representation, filling a critical methodological void.


【21】 Audio Compression using Periodic Gabor with Biorthogonal Exchange:  Implementation Using the Zak Transform
标题: 使用周期Gabor和双正交交换的音频压缩:使用Zak变换的实现
链接:https://arxiv.org/abs/2503.22703

作者: Roger Alimi,  David J. Tannor 
摘要:提出了一种基于Gabor基变换的信号压缩新方法。在Shimshovitz和Tannor的早期工作之后,我们将传统的Gabor函数与Dirichlet函数卷积以获得周期Gabor基集(PG)。PG基对于周期性带限的连续函数是精确的。使用狄利克雷函数的正交性,PG系数的计算变得平凡且数值稳定,但其表示不允许压缩。通过将PG基与其双正交基交换,从而使用局部PG基来计算系数(PGB),来实现大的压缩因子。在这里,我们实现了PGB形式主义使用快速扎克变换,并获得非常高的效率与CPU和内存。我们比较的方法与国家的最先进的短时傅立叶变换(STFT)和离散小波变换(DWT)的方法对各种音频文件,包括音乐和语音样本。在所有的测试情况下,我们的计划远远超过STFT,在大多数情况下优于DWT。
摘要:An efficient new approach to signal compression is presented based of a novel variation on the Gabor basis set. Following earlier work by Shimshovitz and Tannor, we convolve the conventional Gabor functions with Dirichlet functions to obtain a Periodic Gabor basis set (PG). The PG basis is exact for continuous functions that are periodic band-limited. Using the orthonormality of the Dirichlet functions, the calculation of the PG coefficients becomes trivial and numerically stable, but its representation does not allow compression. Large compression factors are achieved by exchanging the PG basis with its biorthogonal basis, thereby using the localized PG basis to calculate the coefficients (PGB). Here we implement the PGB formalism using the Fast Zak Transform and obtain very high efficiency with respect to both CPU and memory. We compare the method with the state of the art Short-Time Fourier Transform (STFT) and Discrete Wavelet Transform (DWT) methods on a variety of audio files, including music and speech samples. In all cases tested our scheme surpasses the STFT by far and in most cases outperforms DWT.


eess.AS音频处理


【1】 Exploring In-Context Learning Capabilities of ChatGPT for Pathological  Speech Detection
标题: 探索ChatGPT用于病理性语音检测的上下文学习能力
链接:https://arxiv.org/abs/2503.23873

作者: Mahdi Amiri,  Hatef Otroshi Shahreza,  Ina Kodrasi 
备注:submitted to EUSIPCO 2025
摘要:自动病理语音检测方法已经显示出有希望的结果,作为潜在的诊断工具与昂贵的传统方法一起受到关注。虽然这些方法可以实现高准确性,但它们缺乏可解释性,限制了它们在临床实践中的适用性。在本文中,我们研究了使用多模态大语言模型(LLM),特别是ChatGPT-4 o,在一个Few-Shot的上下文中学习设置自动病理语音检测。实验结果表明,这种方法不仅提供了有前途的性能,但也提供了解释其决策,增强模型的可解释性。为了进一步了解其有效性,我们进行了一项消融研究,以分析不同因素(如输入类型和系统提示)对最终结果的影响。我们的研究结果突出了多模态LLM在自动病理语音检测中的进一步探索和进步的潜力。
摘要:Automatic pathological speech detection approaches have shown promising results, gaining attention as potential diagnostic tools alongside costly traditional methods. While these approaches can achieve high accuracy, their lack of interpretability limits their applicability in clinical practice. In this paper, we investigate the use of multimodal Large Language Models (LLMs), specifically ChatGPT-4o, for automatic pathological speech detection in a few-shot in-context learning setting. Experimental results show that this approach not only delivers promising performance but also provides explanations for its decisions, enhancing model interpretability. To further understand its effectiveness, we conduct an ablation study to analyze the impact of different factors, such as input type and system prompts, on the final results. Our findings highlight the potential of multimodal LLMs for further exploration and advancement in automatic pathological speech detection.


【2】 Aud-Sur: An Audio Analyzer Assistant for Audio Surveillance Applications
标题: Aud-Sur:音频监控应用的音频分析仪助手
链接:https://arxiv.org/abs/2503.23827

作者: Phat Lam,  Lam Pham,  Dat Tran,  Alexander Schindler,  Silvia Poletti,  Marcel Hasenbalg,  David Fischinger,  Martin Boyer 
备注:A preprint for conference paper, 8 pages, 9 figures
摘要:在本文中,我们提出了一个音频分析器辅助工具,专为各种基于音频的监控应用程序(这项工作是我们的DEFAME FAKES和EUCINF项目的一部分)。建议的工具,简称为Aud-Sur,包括两个主要阶段音频分析和音频检索,分别。在第一阶段,利用多个开源音频模型从用户上传的输入音频记录中提取信息。在第二阶段,用户通过自然的问答方式与Aud-Sur工具进行交互,由大型语言模型(LLM)提供支持,以检索从处理后的音频文件中提取的信息。Aud-Sur工具使用Docker部署在基于微服务的架构设计上。通过利用开源音频模型进行信息提取,LLM用于音频信息检索,以及基于微服务的部署方法,拟议的Aud-Sur工具提供了一个高度可扩展和适应性强的框架,可以集成更多的音频任务,并在音频社区中广泛共享以供进一步开发。
摘要:In this paper, we present an audio analyzer assistant tool designed for a wide range of audio-based surveillance applications (This work is a part of our DEFAME FAKES and EUCINF projects). The proposed tool, refered to as Aud-Sur, comprises two main phases Audio Analysis and Audio Retrieval, respectively. In the first phase, multiple open-source audio models are leveraged to extract information from input audio recording uploaded by a user. In the second phase, users interact with the Aud-Sur tool via a natural question-and-answer manner, powered by a large language model (LLM), to retrieve the information extracted from the processed audio file. The Aud-Sur tool was deployed using Docker on a microservices-based architecture design. By leveraging open-source audio models for information extraction, LLM for audio information retrieval, and a microservices-based deployment approach, the proposed Aud-Sur tool offers a highly extensible and adaptable framework that can integrate more audio tasks, and be widely shared within the audio community for further development.


【3】 A first-order DirAC-based parametric Ambisonic coder for immersive  communications
标题: 一种用于沉浸式通信的一阶DirAC参数化立体混响编码器
链接:https://arxiv.org/abs/2503.23586

作者: Guillaume Fuchs,  Florin Ghido,  Dominik Weckbecker,  Oliver Thiergart 
备注:Accepted at ICASSP'25
摘要:定向音频编码(DirAC)是一种经过验证的方法,用于以B格式参数化地表示3D音频场景,并且能够在任意扬声器布局上再现它。虽然这样的方法似乎很适合低比特率的高保真度立体声传输,很少的工作已经做了一个真正的系统上it.In本文中,我们提出了一个基于DirAC的编码高阶高保真度立体声(HOA),开发的标准化工作的一部分,扩展3GPP EVS编解码器沉浸式通信的可行性。从一阶DirAC模型开始,我们展示了如何通过在球谐域中进行全合成来降低算法延迟、参数所需的比特率和复杂度。对所提出的用于以从32到128 kbps的比特率对3阶高保真度立体声响复制进行编码的技术的评估示出了与现有解决方案相比的参数化方法的相关性。
摘要:Directional Audio Coding (DirAC) is a proven method for parametrically representing a 3D audio scene in B-format and is capable of reproducing it on arbitrary loudspeaker layouts. Although such a method seems well suited for low bitrate Ambisonic transmission, little work has been done on the feasibility of building a real system upon it. In this paper, we present a DirAC-based coding for Higher-Order Ambisonics (HOA), developed as part of a standardisation effort to extend the 3GPP EVS codec to immersive communications. Starting from the first-order DirAC model, we show how to reduce algorithmic delay, the bitrate required for the parameters and complexity by bringing the full synthesis in the spherical harmonic domain. The evaluation of the proposed technique for coding 3\textsuperscript{rd} order Ambisonics at bitrates from 32 to 128 kbps shows the relevance of the parametric approach compared with existing solutions.


【4】 Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs
标题: Aurelia:视听LLM中的测试时推理蒸馏
链接:https://arxiv.org/abs/2503.23219

作者: Sanjoy Chowdhury,  Hanan Gani,  Nishit Anand,  Sayan Nag,  Ruohan Gao,  Mohamed Elhoseiny,  Salman Khan,  Dinesh Manocha 
摘要:推理优化的最新进展极大地提高了大型语言模型(LLM)的性能。然而,现有的工作未能解决视听场景的复杂性,强调需要进一步研究。在本文中,我们介绍了AURELIA,一种新的基于演员-评论家的视听(AV)推理框架,它在测试时将结构化的,逐步的推理提取到AVLLM中,提高了它们处理复杂多模态输入的能力,而无需额外的训练或微调。为了进一步提高AVLLM推理技能,我们提出了AVReasonBench,这是一个具有挑战性的基准测试,包括4500个视听问题,每个问题都配有详细的分步推理。我们的基准测试涵盖了六个不同的任务,包括AV-GeoIQ,它评估了AV推理与地理和文化知识的结合。在AVReasonBench上对18个AVLLM进行评估,发现它们的多模态推理能力存在明显的局限性。使用AURELIA,我们实现了高达100%的相对改善,证明了其有效性。这一性能提升凸显了推理增强数据生成在现实应用中推进AVLLM的潜力。我们的代码和数据将在以下网址公开发布:https://github.com/schowdhury671/aurelia。
摘要:Recent advancements in reasoning optimization have greatly enhanced the performance of large language models (LLMs). However, existing work fails to address the complexities of audio-visual scenarios, underscoring the need for further research. In this paper, we introduce AURELIA, a novel actor-critic based audio-visual (AV) reasoning framework that distills structured, step-by-step reasoning into AVLLMs at test time, improving their ability to process complex multi-modal inputs without additional training or fine-tuning. To further advance AVLLM reasoning skills, we present AVReasonBench, a challenging benchmark comprising 4500 audio-visual questions, each paired with detailed step-by-step reasoning. Our benchmark spans six distinct tasks, including AV-GeoIQ, which evaluates AV reasoning combined with geographical and cultural knowledge. Evaluating 18 AVLLMs on AVReasonBench reveals significant limitations in their multi-modal reasoning capabilities. Using AURELIA, we achieve up to a 100% relative improvement, demonstrating its effectiveness. This performance gain highlights the potential of reasoning-enhanced data generation for advancing AVLLMs in real-world applications. Our code and data will be publicly released at: https: //github.com/schowdhury671/aurelia.


【5】 SupertonicTTS: Towards Highly Scalable and Efficient Text-to-Speech  System
标题: SupertonicTTC:迈向高度可扩展和高效的文本到语音系统
链接:https://arxiv.org/abs/2503.23108

作者: Hyeongju Kim,  Jinhyeok Yang,  Yechan Yu,  Seunghun Ji,  Jacob Morton,  Frederik Bous,  Joon Byun,  Juheon Lee 
备注:19 pages, preprint
摘要:我们提出了一种新的文本到语音(TTS)系统,即SupertonicTTS,提高语音合成的可扩展性和效率。SupertonicTTS由三个组件组成:用于连续潜在表示的语音自动编码器,利用流匹配进行文本到潜在映射的文本到潜在模块,以及话语级持续时间预测器。为了实现轻量级架构,我们采用了低维的潜在空间,潜在的时间压缩和ConvNeXt块。我们通过直接在原始字符级文本上操作并采用交叉注意进行文本语音对齐来进一步简化TTS管道,从而消除了对字素到音素(G2P)模块和外部对齐器的需要。此外,我们引入了上下文共享批量扩展,加速损失收敛和稳定的文本语音对齐。实验结果表明,SupertonicTTS实现竞争力的性能,同时显着降低建筑的复杂性和计算开销相比,当代TTS模型。演示SupertonicTTS功能的音频样本可在https://supertonictts.github.io/上获得。
摘要:We present a novel text-to-speech (TTS) system, namely SupertonicTTS, for improved scalability and efficiency in speech synthesis. SupertonicTTS is comprised of three components: a speech autoencoder for continuous latent representation, a text-to-latent module leveraging flow-matching for text-to-latent mapping, and an utterance-level duration predictor. To enable a lightweight architecture, we employ a low-dimensional latent space, temporal compression of latents, and ConvNeXt blocks. We further simplify the TTS pipeline by operating directly on raw character-level text and employing cross-attention for text-speech alignment, thus eliminating the need for grapheme-to-phoneme (G2P) modules and external aligners. In addition, we introduce context-sharing batch expansion that accelerates loss convergence and stabilizes text-speech alignment. Experimental results demonstrate that SupertonicTTS achieves competitive performance while significantly reducing architectural complexity and computational overhead compared to contemporary TTS models. Audio samples demonstrating the capabilities of SupertonicTTS are available at: https://supertonictts.github.io/.


【6】 The trajectoRIR Database: Room Acoustic Recordings Along a Trajectory of  Moving Microphones
标题: AtlantoRIR数据库:沿着移动麦克风轨迹的房间声学记录
链接:https://arxiv.org/abs/2503.23004

作者: Stefano Damiano,  Kathleen MacWilliam,  Valerio Lorenzoni,  Thomas Dietzen,  Toon van Waterschoot 
备注:15 pages, 7 figures
摘要:数据可用性对于开发声学信号处理算法至关重要,尤其是对于需要大量且多样化训练数据集的数据驱动方法。出于这个原因,近年来发布了越来越多的数据库,包括房间脉冲响应(RIR)或移动音频的记录。在本文中,我们介绍了一个广泛的,多阵列收集的动态和静态的声音记录沿一个受控的轨迹在一个房间里的ARATORIR数据库。具体来说,该数据库的特点是使用移动麦克风和固定RIR沿着L形,3.74米长的轨迹对房间声学进行空间采样。这种组合使得OCToRIR在从声源定位和跟踪到空间动态声场重建和系统识别的各种任务中具有独特性和适用性。录音室具有0.5秒的混响时间,并且所采用的三种不同的麦克风配置包括具有位于耳朵旁边的附加参考麦克风的假人头部、3个一阶高保真度立体声响复制麦克风、16和4声道的两个圆形阵列以及12声道线性阵列。麦克风的运动是使用机器人推车以三种速度穿过轨道实现的:[0.2,0.4,0.8] m/s。使用两个固定扬声器再现音频信号。收集的数据库功能8648静止RIR,以及完美的扫描,语音,音乐和运动过程中记录的静止噪音。MATLAB和Python脚本包括访问录制的音频以及检索几何信息。
摘要:Data availability is essential to develop acoustic signal processing algorithms, especially when it comes to data-driven approaches that demand large and diverse training datasets. For this reason, an increasing number of databases have been published in recent years, including either room impulse responses (RIRs) or recordings of moving audio. In this paper we introduce the trajectoRIR database, an extensive, multi-array collection of both dynamic and stationary acoustic recordings along a controlled trajectory in a room. Specifically, the database features recordings using moving microphones and stationary RIRs spatially sampling the room acoustics along an L-shaped, 3.74-meter-long trajectory. This combination makes trajectoRIR unique and applicable in various tasks ranging from sound source localization and tracking to spatially dynamic sound field reconstruction and system identification. The recording room has a reverberation time of 0.5 seconds, and the three different microphone configurations employed include a dummy head, with additional reference microphones located next to the ears, 3 first-order Ambisonics microphones, two circular arrays of 16 and 4 channels, and a 12-channel linear array. The motion of the microphones was achieved using a robotic cart traversing a rail at three speeds: [0.2,0.4,0.8] m/s. Audio signals were reproduced using two stationary loudspeakers. The collected database features 8648 stationary RIRs, as well as perfect sweeps, speech, music, and stationary noise recorded during motion. MATLAB and Python scripts are included to access the recorded audio as well as to retrieve geometrical information.


【7】 Congenital Heart Disease Classification Using Phonocardiograms: A  Scalable Screening Tool for Diverse Environments
标题: 使用听诊器进行先天性心脏病分类:适用于不同环境的可扩展筛查工具
链接:https://arxiv.org/abs/2503.22773

作者: Abdul Jabbar,  Ethan Grooby,  Jack Crozier,  Alexander Gallon,  Vivian Pham,  Khawza I Ahmad,  Md Hassanuzzaman,  Raqibul Mostafa,  Ahsan H. Khandoker,  Faezeh Marzbanrad 
备注:12 pages, 6 figures
摘要:先天性心脏病(CHD)是一种需要早期发现的严重疾病,特别是在婴儿和儿童时期。这项研究提出了一种深度学习模型,旨在使用心音图(PCG)信号检测CHD,重点关注其在全球健康中的应用。我们在几个数据集上评估了我们的模型,包括来自孟加拉国的主要数据集,实现了94.1%的高准确性,92.7%的灵敏度,96.3%的特异性。该模型还在公共PhysioNet Challenge 2022和2016数据集上表现出强大的性能,强调了其对不同人群和数据源的通用性。我们评估了该算法在胸部单个和多个听诊部位的性能,表明即使使用单个位置,该模型也能保持85%以上的准确性。此外,我们的算法能够在心脏病专家认为非诊断性的低质量记录上实现80%的准确性。这项研究表明,人工智能驱动的数字听诊器可以在资源有限的环境中作为CHD的成本效益筛选工具,增强临床决策支持并最终改善患者预后。
摘要:Congenital heart disease (CHD) is a critical condition that demands early detection, particularly in infancy and childhood. This study presents a deep learning model designed to detect CHD using phonocardiogram (PCG) signals, with a focus on its application in global health. We evaluated our model on several datasets, including the primary dataset from Bangladesh, achieving a high accuracy of 94.1%, sensitivity of 92.7%, specificity of 96.3%. The model also demonstrated robust performance on the public PhysioNet Challenge 2022 and 2016 datasets, underscoring its generalizability to diverse populations and data sources. We assessed the performance of the algorithm for single and multiple auscultation sites on the chest, demonstrating that the model maintains over 85% accuracy even when using a single location. Furthermore, our algorithm was able to achieve an accuracy of 80% on low-quality recordings, which cardiologists deemed non-diagnostic. This research suggests that an AI- driven digital stethoscope could serve as a cost-effective screening tool for CHD in resource-limited settings, enhancing clinical decision support and ultimately improving patient outcomes.


【8】 Chirp Localization via Fine-Tuned Transformer Model: A Proof-of-Concept  Study
标题: 基于精细调谐Transformer模型的啁啾定位:概念验证研究
链接:https://arxiv.org/abs/2503.22713

作者: Nooshin Bahador,  Milad Lankarany 
备注:19 pages, 8 figures
摘要:频谱图是时频信号分析的关键,广泛应用于音频处理和计算神经科学。脑电图(EEG)频谱图(由线性或指数频率扫描标记)中的Chirp样模式是癫痫发作动力学的关键生物标志物,但缺乏用于其检测、定位和特征提取的自动化工具。这项研究通过对合成频谱图的Vision Transformer(ViT)模型进行微调来弥合这一差距,并通过低秩自适应(LoRA)来增强适应性。我们生成了100000合成光谱与啁啾参数,创造了第一个大规模的基准啁啾定位。这些频谱图使用线性或指数频率扫描、高斯噪声和平滑来模拟神经啁啾。一个ViT模型,适用于回归,预测啁啾参数。LoRA微调了注意力层,使预训练的骨干能够有效更新。训练使用MSE损失和AdamW优化器,具有学习率调度器和早期停止以抑制过度拟合。仅针对三个特征:Chirp开始时间(开始时间)、Chirp开始频率(开始频率)和Chirp结束频率(偏移频率)。通过预测和实际标签之间的Pearson相关性评估性能。结果显示出强烈的一致性:0.9841线性调频脉冲开始时间的相关性,具有稳定的推断时间(137至140秒)和误差分布的最小偏差。这种方法提供了一个工具,啁啾分析脑电时频表示,填补了一个关键的方法空白。
摘要:Spectrograms are pivotal in time-frequency signal analysis, widely used in audio processing and computational neuroscience. Chirp-like patterns in electroencephalogram (EEG) spectrograms (marked by linear or exponential frequency sweep) are key biomarkers for seizure dynamics, but automated tools for their detection, localization, and feature extraction are lacking. This study bridges this gap by fine-tuning a Vision Transformer (ViT) model on synthetic spectrograms, augmented with Low-Rank Adaptation (LoRA) to boost adaptability. We generated 100000 synthetic spectrograms with chirp parameters, creating the first large-scale benchmark for chirp localization. These spectrograms mimic neural chirps using linear or exponential frequency sweep, Gaussian noise, and smoothing. A ViT model, adapted for regression, predicted chirp parameters. LoRA fine-tuned the attention layers, enabling efficient updates to the pre-trained backbone. Training used MSE loss and the AdamW optimizer, with a learning rate scheduler and early stopping to curb overfitting. Only three features were targeted: Chirp Start Time (Onset Time), Chirp Start Frequency (Onset Frequency), and Chirp End Frequency (Offset Frequency). Performance was evaluated via Pearson correlation between predicted and actual labels. Results showed strong alignment: 0.9841 correlation for chirp start time, with stable inference times (137 to 140s) and minimal bias in error distributions. This approach offers a tool for chirp analysis in EEG time-frequency representation, filling a critical methodological void.


【9】 Enhancing nonnative speech perception and production through an  AI-powered application
标题: 通过人工智能驱动的应用程序增强非母语语音感知和生成
链接:https://arxiv.org/abs/2503.22705

作者: Georgios P. Georgiou 
摘要:虽然通过各种应用程序使用人工智能(AI)来增强外语发音的研究正在扩大,但它主要集中在可理解性和可理解性等方面,在很大程度上忽视了在感知和生产方面对个人语音的改善。这项研究旨在通过研究使用人工智能驱动的移动应用程序进行培训对非母语声音感知和产生的影响来解决这一差距。参与者完成了一个前测,评估他们的能力,区分第二语言英语注意隐藏的对比,并产生这些元音的句子上下文。干预措施包括使用Speakometer移动应用程序进行培训,其中包括以英语元音为特色的记录任务,以及发音反馈和练习。后测反映了前测,以衡量表现的变化。结果显示,干预后,辨别准确性和目标对比度的产生都有显着改善。然而,参与者并没有达到像本地人一样的能力。这些发现突出了人工智能驱动的应用程序在促进语音获取方面的有效性,并支持它们在课堂之外用于个性化,交互式发音训练的潜在用途。
摘要:While research on using Artificial Intelligence (AI) through various applications to enhance foreign language pronunciation is expanding, it has primarily focused on aspects such as comprehensibility and intelligibility, largely neglecting the improvement of individual speech sounds in both perception and production. This study seeks to address this gap by examining the impact of training with an AI-powered mobile application on nonnative sound perception and production. Participants completed a pretest assessing their ability to discriminate the second language English heed-hid contrast and produce these vowels in sentence contexts. The intervention involved training with the Speakometer mobile application, which incorporated recording tasks featuring the English vowels, along with pronunciation feedback and practice. The posttest mirrored the pretest to measure changes in performance. The results revealed significant improvements in both discrimination accuracy and production of the target contrast following the intervention. However, participants did not achieve native-like competence. These findings highlight the effectiveness of AI-powered applications in facilitating speech acquisition and support their potential use for personalized, interactive pronunciation training beyond the classroom.


【10】 Audio Compression using Periodic Gabor with Biorthogonal Exchange:  Implementation Using the Zak Transform
标题: 使用周期Gabor和双正交交换的音频压缩:使用Zak变换的实现
链接:https://arxiv.org/abs/2503.22703

作者: Roger Alimi,  David J. Tannor 
摘要:提出了一种基于Gabor基变换的信号压缩新方法。在Shimshovitz和Tannor的早期工作之后,我们将传统的Gabor函数与Dirichlet函数卷积以获得周期Gabor基集(PG)。PG基对于周期性带限的连续函数是精确的。使用狄利克雷函数的正交性,PG系数的计算变得平凡且数值稳定,但其表示不允许压缩。通过将PG基与其双正交基交换,从而使用局部PG基来计算系数(PGB),来实现大的压缩因子。在这里,我们实现了PGB形式主义使用快速扎克变换,并获得非常高的效率与CPU和内存。我们比较的方法与国家的最先进的短时傅立叶变换(STFT)和离散小波变换(DWT)的方法对各种音频文件,包括音乐和语音样本。在所有的测试情况下,我们的计划远远超过STFT,在大多数情况下优于DWT。
摘要:An efficient new approach to signal compression is presented based of a novel variation on the Gabor basis set. Following earlier work by Shimshovitz and Tannor, we convolve the conventional Gabor functions with Dirichlet functions to obtain a Periodic Gabor basis set (PG). The PG basis is exact for continuous functions that are periodic band-limited. Using the orthonormality of the Dirichlet functions, the calculation of the PG coefficients becomes trivial and numerically stable, but its representation does not allow compression. Large compression factors are achieved by exchanging the PG basis with its biorthogonal basis, thereby using the localized PG basis to calculate the coefficients (PGB). Here we implement the PGB formalism using the Fast Zak Transform and obtain very high efficiency with respect to both CPU and memory. We compare the method with the state of the art Short-Time Fourier Transform (STFT) and Discrete Wavelet Transform (DWT) methods on a variety of audio files, including music and speech samples. In all cases tested our scheme surpasses the STFT by far and in most cases outperforms DWT.


【11】 Enhancing Aviation Communication Transcription: Fine-Tuning  Distil-Whisper with LoRA
标题: 增强航空通信转录:用LoRA微调Distil-Whisper
链接:https://arxiv.org/abs/2503.22692

作者: Shokoufeh Mirzaei,  Jesse Arzate,  Yukti Vijay 
备注:14 pages, 4 Figures, 4 Tables, Under review by Journal of Aerospace Information Systems
摘要:航空通信的转录有几个应用,从协助空中交通管制员识别回读错误的准确性到搜索和救援行动。人工智能的最新进展为改善航空通信转录任务提供了前所未有的机会。OpenAI的Whisper是领先的自动语音识别模型。然而,微调耳语航空通信转录是计算效率不高。因此,本文旨在使用一种称为低秩自适应的参数高效微调方法来微调计算效率更高的Whisper版本Distil-Whisper。为了进行微调,我们使用了来自语言数据联盟的空中交通管制语料库数据集,其中包含美国三个主要机场附近大约70小时的管制员和飞行员传输。其目的是减少字的错误率,以提高航空通信记录的准确性。首先,从LoRA的一组初始超参数(Alpha = 64,Rank = 32)开始,我们进行了网格搜索。我们应用了5重交叉验证来找到蒸馏-Whisper超参数的最佳组合。然后,我们对LoRA超参数模型进行了微调,在五倍范围内实现了令人印象深刻的平均单词错误率3.86%。这一结果突出了该模型在驾驶舱中使用的潜力。
摘要:Transcription of aviation communications has several applications, from assisting air traffic controllers in identifying the accuracy of read-back errors to search and rescue operations. Recent advances in artificial intelligence have provided unprecedented opportunities for improving aviation communication transcription tasks. OpenAI's Whisper is one of the leading automatic speech recognition models. However, fine-tuning Whisper for aviation communication transcription is not computationally efficient. Thus, this paper aims to use a Parameter-Efficient Fine-tuning method called Low-Rank Adaptation to fine-tune a more computationally efficient version of Whisper, distil-Whisper. To perform the fine-tuning, we used the Air Traffic Control Corpus dataset from the Linguistic Data Consortium, which contains approximately 70 hours of controller and pilot transmissions near three major airports in the US. The objective was to reduce the word error rate to enhance accuracy in the transcription of aviation communication. First, starting with an initial set of hyperparameters for LoRA (Alpha = 64 and Rank = 32), we performed a grid search. We applied a 5-fold cross-validation to find the best combination of distil-Whisper hyperparameters. Then, we fine-tuned the model for LoRA hyperparameters, achieving an impressive average word error rate of 3.86% across five folds. This result highlights the model's potential for use in the cockpit.


【12】 Qieemo: Speech Is All You Need in the Emotion Recognition in  Conversations
标题: 切莫:言语就是对话中情感识别的全部
链接:https://arxiv.org/abs/2503.22687

作者: Jinming Chen,  Jingyi Fang,  Yuanzhong Zheng,  Yaoxuan Wang,  Haojun Fei 
摘要:情感识别在智能人机交互系统中起着举足轻重的作用。多模态方法受益于多种模态的融合,从而提高识别准确性。然而,缺乏高质量的多模式数据和实现不同模式之间的最佳对齐的挑战,大大限制了改进多模式方法的潜力。在本文中,所提出的Qieemo框架有效地利用预训练的自动语音识别(ASR)模型骨干,其中包含自然帧对齐的文本和情感特征,以实现精确的情感分类仅基于音频模态。此外,我们设计了多模态融合(MMF)模块和跨模态注意(CMA)模块,以融合语音后验图(PPG)和情感特征提取的ASR编码器,以提高识别精度。在IEMOCAP数据集上的实验结果表明,Qieemo优于基准单峰,多峰和自监督模型,分别有3.0%,1.2%和1.9%的绝对改进。
摘要:Emotion recognition plays a pivotal role in intelligent human-machine interaction systems. Multimodal approaches benefit from the fusion of diverse modalities, thereby improving the recognition accuracy. However, the lack of high-quality multimodal data and the challenge of achieving optimal alignment between different modalities significantly limit the potential for improvement in multimodal approaches. In this paper, the proposed Qieemo framework effectively utilizes the pretrained automatic speech recognition (ASR) model backbone which contains naturally frame aligned textual and emotional features, to achieve precise emotion classification solely based on the audio modality. Furthermore, we design the multimodal fusion (MMF) module and cross-modal attention (CMA) module in order to fuse the phonetic posteriorgram (PPG) and emotional features extracted by the ASR encoder for improving recognition accuracy. The experimental results on the IEMOCAP dataset demonstrate that Qieemo outperforms the benchmark unimodal, multimodal, and self-supervised models with absolute improvements of 3.0%, 1.2%, and 1.9% respectively.


【13】 UniSep: Universal Target Audio Separation with Language Models at Scale
标题: UniSep:采用大规模语言模型的通用目标音频分离
链接:https://arxiv.org/abs/2503.23762

作者: Yuanyuan Wang,  Hangting Chen,  Dongchao Yang,  Weiqin Li,  Dan Luo,  Guangzhi Li,  Shan Yang,  Zhiyong Wu,  Helen Meng,  Xixin Wu 
备注:Accepted by ICME 2025
摘要:我们提出了通用目标音频分离(UniSep),解决了不同类型音频的任意混合的分离任务。区别于以往的研究,UniSep进行无限的源域和无限的源号码。我们制定的分离任务作为一个序列到序列的问题,和一个大的语言模型(LLM)是用来模拟的离散潜在空间中的音频序列,利用LLM在处理复杂的混合音频与大规模数据的权力。此外,提出了一种新的预训练策略,利用音频的数据,这减少了大规模的数据模拟的努力,并提高了LLM的能力,以了解音频序列中的信息的一致性和相关性。我们还证明了在音频分离任务中缩放数据集的有效性:我们使用大规模数据(36.5k小时),包括语音,音乐和声音,训练一个通用的目标音频分离模型,不限于特定的域。实验表明,与单任务模型相比,UniSep获得了具有竞争力的主观和客观评价结果。
摘要:We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep is performed on unlimited source domains and unlimited source numbers. We formulate the separation task as a sequence-to-sequence problem, and a large language model (LLM) is used to model the audio sequence in the discrete latent space, leveraging the power of LLM in handling complex mixture audios with large-scale data. Moreover, a novel pre-training strategy is proposed to utilize audio-only data, which reduces the efforts of large-scale data simulation and enhances the ability of LLMs to understand the consistency and correlation of information within audio sequences. We also demonstrate the effectiveness of scaling datasets in an audio separation task: we use large-scale data (36.5k hours), including speech, music, and sound, to train a universal target audio separation model that is not limited to a specific domain. Experiments show that UniSep achieves competitive subjective and objective evaluation results compared with single-task models.


【14】 Evaluation of the Pronunciation of Tajweed Rules Based on DNN as a Step  Towards Interactive Recitation Learning
标题: 基于DNN的Tajweed规则发音评估作为交互式背诵学习的一步
链接:https://arxiv.org/abs/2503.23470

作者: Dim Shaiakhmetov,  Gulnaz Gimaletdinova,  Selcuk Cankurt,  Kadyrmamat Momunov 
摘要:正确的古兰经背诵,坚持Tajweed的规则,是至关重要的,以防止在背诵过程中的错误,并需要大量的努力掌握。传统的教授这些规则的方法受到合格教师的可用性和时间限制的限制。背诵的自动评估可以通过提供及时的反馈和支持独立练习来解决这些挑战。这项研究的重点是开发一个深度学习模型来分类三个Tajweed规则-分离拉伸(Al Mad),紧中午(Ghunnah)和隐藏(Ikhfaa)-使用公开的QDAT数据集,其中包含超过1,500个音频记录。输入数据由来自该数据集的音频记录组成,转换为归一化的梅尔频谱图。对于分类,使用了EfficientNet-B 0架构,并通过挤压和激发注意力机制进行了增强。开发的模型实现了95.35%,99.34%和97.01%的准确率为各自的规则。对学习曲线的分析证实了该模型的鲁棒性和不存在过拟合。所提出的方法具有很高的效率,并为开发交互式教育系统的Tajweed研究铺平了道路。
摘要:Proper recitation of the Quran, adhering to the rules of Tajweed, is crucial for preventing mistakes during recitation and requires significant effort to master. Traditional methods of teaching these rules are limited by the availability of qualified instructors and time constraints. Automatic evaluation of recitation can address these challenges by providing prompt feedback and supporting independent practice. This study focuses on developing a deep learning model to classify three Tajweed rules - separate stretching (Al Mad), tight noon (Ghunnah), and hide (Ikhfaa) - using the publicly available QDAT dataset, which contains over 1,500 audio recordings. The input data consisted of audio recordings from this dataset, transformed into normalized mel-spectrograms. For classification, the EfficientNet-B0 architecture was used, enhanced with a Squeeze-and-Excitation attention mechanism. The developed model achieved accuracy rates of 95.35%, 99.34%, and 97.01% for the respective rules. An analysis of the learning curves confirmed the model's robustness and absence of overfitting. The proposed approach demonstrates high efficiency and paves the way for developing interactive educational systems for Tajweed study.


【15】 Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
标题: 用于高效语音聊天机器人助手的推测性结束转向检测器
链接:https://arxiv.org/abs/2503.23439

作者: Hyunjong Ok,  Suho Yoo,  Jaeho Lee 
备注:Preprint
摘要:由大型语言模型驱动的口语对话系统在理解人类语音和生成适当的口语响应方面表现出了卓越的能力。然而,这些系统与转弯结束检测(ETD)-区分用户转弯完成和犹豫的能力相矛盾。这种限制往往会导致过早或延迟的反应,扰乱口语对话的流动。在本文中,我们介绍了ETD数据集,这是第一个用于末端转弯检测的公共数据集。ETD数据集由文本到语音模型生成的合成语音数据和从网络源收集的真实语音数据组成。我们还提出了SpeculativeETD,一种新的协作推理框架,平衡效率和准确性,以提高实时ETD在资源受限的环境。我们的方法联合采用了一个轻量级的基于GRU的模型,该模型可以在本地设备上实时快速检测非说话单元,以及一个在服务器上运行的高性能基于Wav 2 vec的模型,以进行更具挑战性的分类,区分转弯结束和仅仅是停顿。实验表明,所提出的SpeculativeETD显着提高ETD的准确性,同时保持所需的计算量低。数据集和代码将在审查后提供。
摘要:Spoken dialogue systems powered by large language models have demonstrated remarkable abilities in understanding human speech and generating appropriate spoken responses. However, these systems struggle with end-turn detection (ETD) -- the ability to distinguish between user turn completion and hesitation. This limitation often leads to premature or delayed responses, disrupting the flow of spoken conversations. In this paper, we introduce the ETD Dataset, the first public dataset for end-turn detection. The ETD dataset consists of both synthetic speech data generated with text-to-speech models and real-world speech data collected from web sources. We also propose SpeculativeETD, a novel collaborative inference framework that balances efficiency and accuracy to improve real-time ETD in resource-constrained environments. Our approach jointly employs a lightweight GRU-based model, which rapidly detects the non-speaking units in real-time on local devices, and a high-performance Wav2vec-based model running on the server to make a more challenging classification of distinguishing turn ends from mere pauses. Experiments demonstrate that the proposed SpeculativeETD significantly improves ETD accuracy while keeping the required computations low. Datasets and code will be available after the review.


【16】 Scaling Auditory Cognition via Test-Time Compute in Audio Language  Models
标题: 通过音频语言模型中的测试时间计算缩放听觉认知
链接:https://arxiv.org/abs/2503.23395

作者: Ting Dang,  Yan Gao,  Hong Jia 
摘要:大型语言模型(LLM)在自然语言处理中表现出了非凡的多功能性,促使最近的努力通过开发音频大型语言模型(Audio LLM)将其多模态功能扩展到语音处理。虽然音频LLM在语音识别和合成等任务中表现出色,但目前尚不清楚它们在面对现实环境带来的听觉认知挑战时的表现,例如音频理解和听力回忆,特别是在存在背景噪声或重叠语音的情况下。与基于文本的LLM不同,基于文本的LLM可以访问大量的文本数据进行预训练,由于模拟真实世界听觉认知场景的数据集有限以及获取听觉认知标签进行训练的挑战,使用不同的听觉认知场景重新训练音频LLM是困难的。虽然测试时间计算(TTC)方法已被证明可以增强基于文本的LLM在推理过程中的能力,但一个关键的挑战在于设计这些TTC方法来提高音频LLM的听觉能力。本研究旨在解决这两个研究空白:i)探索音频LLM的听觉认知能力,以及ii)使用TTC方法增强其能力。我们已经调查了五种不同的音频LLM听觉认知使用的\textit{自我收集}数据库,并提出了五个TTC方法,以提高听觉认知能力的推理。我们的研究结果表明,音频LLM性能下降更具挑战性的听觉认知任务。所提出的TTC方法显著增强了认知听觉能力,推动了更具适应性和弹性的音频LLM的开发,用于辅助听力设备,基于语音的AI助手和通信技术等实际应用。
摘要:Large language models (LLMs) have shown exceptional versatility in natural language processing, prompting recent efforts to extend their multimodal capabilities to speech processing through the development of audio large language models (Audio LLMs). While Audio LLMs excel in tasks such as speech recognition and synthesis, it remains unclear how they perform when faced with the auditory cognitive challenges posed by real-world environments, such as audio comprehension and listening recall, particularly in the presence of background noise or overlapping speech. Unlike text-based LLMs, which have access to vast amounts of text data for pre-training, retraining Audio LLMs with diverse auditory cognitive scenes is difficult due to the limited datasets that simulate real-world auditory cognitive scenarios and the challenge of acquiring auditory cognitive labels for training. While test-time compute (TTC) methods have been shown to enhance the capabilities of text-based LLMs during inference, a key challenge lies in designing these TTC methods to improve the auditory capabilities of Audio LLMs. This study aims to address these two research gaps by: i) exploring the auditory cognitive capabilities of Audio LLMs, and ii) enhancing their capabilities using TTC approaches. We have investigated five different Audio LLMs for auditory cognition using a \textit{self-collected} database and have proposed five TTC approaches to enhance auditory cognitive capabilities during inference. Our findings reveal that Audio LLMs performance decreases in more challenging auditory cognitive tasks. The proposed TTC approaches significantly enhance cognitive auditory capabilities, advancing the development of more adaptable and resilient Audio LLMs for practical applications such as assistive listening devices, voice-based AI assistants, and communication technologies.


【17】 D3-Guard: Acoustic-based Drowsy Driving Detection Using Smartphones
标题: D3-Guard:使用智能手机基于声学的昏昏欲驾驶检测
链接:https://arxiv.org/abs/2503.23393

作者: Yadong Xie,  Fan Li,  Yue Wu,  Song Yang,  Yu Wang 
备注:IEEE INFOCOM 2019-IEEE Conference on Computer Communications
摘要:近年来,随着汽车数量的快速增长,行车安全越来越受到人们的关注。疲劳驾驶是对行车安全的最大威胁之一。因此,一个简单但强大的系统,可以检测困倦驾驶与商业现成的设备(如智能手机)是非常必要的。出于这个动机,我们探索纯粹使用嵌入在智能手机中的声学传感器来检测困倦驾驶的可行性。本文首先研究了驾驶员在驾驶过程中的疲劳行为,发现了驾驶员在驾驶过程中的三种典型疲劳行为,即点头、打哈欠和操作方向盘时所引起的多普勒频移的一些独特模式。然后,我们通过对从真实驾驶环境中收集的驾驶数据进行实证分析来验证我们的重要发现。我们进一步提出了一种基于嵌入在智能手机中的音频设备的实时昏昏欲睡驾驶检测系统(D3-Guard)。为了提高系统的性能,我们采用了一种基于欠采样技术和FFT的有效特征提取方法,并精心设计了一种基于LSTM网络的高精度疲劳驾驶早期检测器。通过在真实驾驶环境中对5名志愿者驾驶员的大量实验,我们的系统可以实时区分困倦驾驶行为,平均总准确率为93.31%。超过80%的困倦驾驶动作可以在动作持续时间的前70%内检测到。
摘要:Since the number of cars has grown rapidly in recent years, driving safety draws more and more public attention. Drowsy driving is one of the biggest threatens to driving safety. Therefore, a simple but robust system that can detect drowsy driving with commercial off-the-shelf devices (such as smartphones) is very necessary. With this motivation, we explore the feasibility of purely using acoustic sensors embedded in smartphones to detect drowsy driving. We first study characteristics of drowsy driving, and find some unique patterns of Doppler shift caused by three typical drowsy behaviors, i.e. nodding, yawning and operating steering wheel. We then validate our important findings through empirical analysis of the driving data collected from real driving environments. We further propose a real-time Drowsy Driving Detection system (D3-Guard) based on audio devices embedded in smartphones. In order to improve the performance of our system, we adopt an effective feature extraction method based on undersampling technique and FFT, and carefully design a high-accuracy detector based on LSTM networks for the early detection of drowsy driving. Through extensive experiments with 5 volunteer drivers in real driving environments, our system can distinguish drowsy driving actions with an average total accuracy of 93.31% in real-time. Over 80% drowsy driving actions can be detected within first 70% of action duration.


【18】 HearSmoking: Smoking Detection in Driving Environment via Acoustic  Sensing on Smartphones
标题: HearSmoking:通过智能手机上的声学传感检测驾驶环境中的吸烟
链接:https://arxiv.org/abs/2503.23391

作者: Yadong Xie,  Fan Li,  Yue Wu,  Song Yang,  Yu Wang 
备注:IEEE Transactions on Mobile Computing ( Volume: 21, Issue: 8, 01 August 2022)
摘要:近年来,由于汽车数量的快速增长,驾驶安全引起了公众的广泛关注。吸烟是对驾驶安全的威胁之一,但往往被司机忽视。现有的吸烟检测工作要么以接触方式工作,要么需要额外的设备。这促使我们探索使用智能手机来检测驾驶环境中的吸烟事件的实用性。在本文中,我们提出了一种名为HearSmoking的吸烟检测系统,该系统仅使用智能手机上的声学传感器来提高驾驶安全性。在调查了驾驶员的典型吸烟习惯(包括手部运动和胸部波动)后,我们设计了一个由扬声器发出并由麦克风接收的声学信号。我们计算接收信号的相对相关系数,以获得手和胸部的运动模式。处理后的数据被发送到一个训练好的卷积神经网络进行分类的手运动。我们还设计了一种同时检测呼吸的方法。为了提高系统性能,我们进一步分析了复合吸烟运动的周期性。通过在真实驾驶环境中的大量实验,HearSmoking实时检测吸烟事件的平均总准确率为93.44%。
摘要:Driving safety has drawn much public attention in recent years due to the fast-growing number of cars. Smoking is one of the threats to driving safety but is often ignored by drivers. Existing works on smoking detection either work in contact manner or need additional devices. This motivates us to explore the practicability of using smartphones to detect smoking events in driving environment. In this paper, we propose a cigarette smoking detection system, named HearSmoking, which only uses acoustic sensors on smartphones to improve driving safety. After investigating typical smoking habits of drivers, including hand movement and chest fluctuation, we design an acoustic signal to be emitted by the speaker and received by the microphone. We calculate Relative Correlation Coefficient of received signals to obtain movement patterns of hands and chest. The processed data is sent into a trained Convolutional Neural Network for classification of hand movement. We also design a method to detect respiration at the same time. To improve system performance, we further analyse the periodicity of the composite smoking motion. Through extensive experiments in real driving environments, HearSmoking detects smoking events with an average total accuracy of 93.44 percent in real-time.


【19】 HearFit+: Personalized Fitness Monitoring via Audio Signals on Smart  Speakers
标题: HearFit+:通过智能扬声器上的音频信号进行个性化健身监测
链接:https://arxiv.org/abs/2503.23387

作者: Yadong Xie,  Fan Li,  Yue Wu,  Yu Wang 
备注:IEEE Transactions on Mobile Computing ( Volume: 22, Issue: 5, 01 May 2023)
摘要:健身有助于强健肌肉,增强抗病能力,改善体形。现在,很多人选择在家里/办公室锻炼,而不是在健身房,因为缺乏时间。但是,如果没有专业的指导,他们很难获得良好的健身效果。出于这一动机,我们提出了第一个个性化的健身监测系统,HearFit+,在家庭/办公室使用智能扬声器。我们探讨了利用声学传感监测健身的可行性。设计了一种基于多普勒频移的适应度检测方法,并采用短时能量对适应度动作进行分割。基于深度学习,HearFit+可以同时进行健身分类和用户识别。结合增量学习,用户可以轻松添加新操作。我们设计了4个评估指标(即,持续时间、强度、连续性和平滑度),以帮助用户提高健身效果。通过对12名志愿者10种健身类型的9,000多个动作的大量实验,HearFit+在健身分类上的平均准确率达到96.13%,在用户识别上的准确率达到91%。所有志愿者都确认HearFit+可以帮助改善各种环境下的健身效果。
摘要:Fitness can help to strengthen muscles, increase resistance to diseases, and improve body shape. Nowadays, a great number of people choose to exercise at home/office rather than at the gym due to lack of time. However, it is difficult for them to get good fitness effects without professional guidance. Motivated by this, we propose the first personalized fitness monitoring system, HearFit+, using smart speakers at home/office. We explore the feasibility of using acoustic sensing to monitor fitness. We design a fitness detection method based on Doppler shift and adopt the short time energy to segment fitness actions. Based on deep learning, HearFit+ can perform fitness classification and user identification at the same time. Combined with incremental learning, users can easily add new actions. We design 4 evaluation metrics (i.e., duration, intensity, continuity, and smoothness) to help users to improve fitness effects. Through extensive experiments including over 9,000 actions of 10 types of fitness from 12 volunteers, HearFit+ can achieve an average accuracy of 96.13% on fitness classification and 91% accuracy for user identification. All volunteers confirm that HearFit+ can help improve the fitness effect in various environments.


【20】 JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical  Spatio-Temporal Prior Synchronization
标题: JavisDiT:具有分层时空先验同步的联合音频-视频扩散Transformer
链接:https://arxiv.org/abs/2503.23377

作者: Kai Liu,  Wei Li,  Lai Chen,  Shengqiong Wu,  Yanhao Zheng,  Jiayi Ji,  Fan Zhou,  Rongxin Jiang,  Jiebo Luo,  Hao Fei,  Tat-Seng Chua 
备注:Work in progress. Homepage: this https URL
摘要:本文介绍了一种用于同步音视频生成的新型联合音视频扩散Transformer--JavisDiT。建立在强大的扩散Transformer(DiT)架构,JavisDiT是能够产生高品质的音频和视频内容,同时从开放式的用户提示。为了确保最佳的同步,我们引入了一个细粒度的时空对齐机制,通过分层时空同步先验(HiST-Sypo)估计。该模块提取全局和细粒度的时空先验,指导视觉和听觉组件之间的同步。此外,我们提出了一个新的基准,JavisBench,由10,140个高质量的文本字幕声音视频组成,跨越不同的场景和复杂的现实世界场景。此外,我们专门设计了一个强大的度量标准,用于评估在现实世界的复杂内容中生成的音频-视频对之间的同步。实验结果表明,JavisDiT通过确保高质量的生成和精确的同步,显着优于现有的方法,为JAVG任务设定了新的标准。我们的代码、模型和数据集将在www.example.com上公开。
摘要:This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Transformer designed for synchronized audio-video generation (JAVG). Built upon the powerful Diffusion Transformer (DiT) architecture, JavisDiT is able to generate high-quality audio and video content simultaneously from open-ended user prompts. To ensure optimal synchronization, we introduce a fine-grained spatio-temporal alignment mechanism through a Hierarchical Spatial-Temporal Synchronized Prior (HiST-Sypo) Estimator. This module extracts both global and fine-grained spatio-temporal priors, guiding the synchronization between the visual and auditory components. Furthermore, we propose a new benchmark, JavisBench, consisting of 10,140 high-quality text-captioned sounding videos spanning diverse scenes and complex real-world scenarios. Further, we specifically devise a robust metric for evaluating the synchronization between generated audio-video pairs in real-world complex content. Experimental results demonstrate that JavisDiT significantly outperforms existing methods by ensuring both high-quality generation and precise synchronization, setting a new standard for JAVG tasks. Our code, model, and dataset will be made publicly available at https://javisdit.github.io/.


【21】 Joint Source-Environment Adaptation for Deep Learning-Based Underwater  Acoustic Source Ranging
标题: 基于深度学习的水下声波源距离联合源环境自适应
链接:https://arxiv.org/abs/2503.23262

作者: Dariush Kari,  Andrew C. Singer 
摘要:在本文中,我们提出了一种方法,使预先训练的基于深度学习的水下声学定位模型适应新的环境。我们使用无监督域自适应来提高模型的泛化性能,即,使用无监督损失,微调预训练的网络参数,而无需访问目标环境的任何标签或用于预训练模型的任何数据。该方法通过将预训练模型预测与基于接收信号能量(取决于源)的几乎独立的估计相耦合来改进预训练模型预测。我们展示了这种方法在类似于SWellEx-96实验的环境中对Bellhop生成的数据的有效性,该实验被来自KAM 11实验的真实海洋噪声污染。
摘要:In this paper, we propose a method to adapt a pre-trained deep-learning-based model for underwater acoustic localization to a new environment. We use unsupervised domain adaptation to improve the generalization performance of the model, i.e., using an unsupervised loss, fine-tune the pre-trained network parameters without access to any labels of the target environment or any data used to pre-train the model. This method improves the pre-trained model prediction by coupling that with an almost independent estimation based on the received signal energy (that depends on the source). We show the effectiveness of this approach on Bellhop generated data in an environment similar to that of the SWellEx-96 experiment contaminated with real ocean noise from the KAM11 experiment.


【22】 Mismatch-Robust Underwater Acoustic Localization Using A Differentiable  Modular Forward Model
标题: 使用可微模块正演模型的匹配鲁棒性水下声学定位
链接:https://arxiv.org/abs/2503.23260

作者: Dariush Kari,  Yongjie Zhuang,  Andrew C. Singer 
摘要:本文研究了环境失配条件下的水声定位问题。特别是,我们利用一个预先训练的神经网络的声波传播在一个基于梯度的优化框架来估计源位置。为了减轻训练数据和测试数据之间的不匹配的影响,我们同时优化网络的权重在推理时,并提供条件下,该方法是有效的。此外,我们在前向模型中引入了物理启发的模块化,使我们能够以端到端的训练方式学习多路径结构的路径长度,而无需访问特定的路径标签。我们调查的有效性的假设在一个简单而说明性的环境模型。
摘要:In this paper, we study the underwater acoustic localization in the presence of environmental mismatch. Especially, we exploit a pre-trained neural network for the acoustic wave propagation in a gradient-based optimization framework to estimate the source location. To alleviate the effect of mismatch between the training data and the test data, we simultaneously optimize over the network weights at the inference time, and provide conditions under which this method is effective. Moreover, we introduce a physics-inspired modularity in the forward model that enables us to learn the path lengths of the multipath structure in an end-to-end training manner without access to the specific path labels. We investigate the validity of the assumptions in a simple yet illustrative environment model.


【23】 Joint Source-Environment Adaptation of Data-Driven Underwater Acoustic  Source Ranging Based on Model Uncertainty
标题: 基于模型不确定性的数据驱动水下声源联合自适应测距
链接:https://arxiv.org/abs/2503.23258

作者: Dariush Kari,  Hari Vishnu,  Andrew C. Singer 
摘要:使预训练的深度学习模型适应新的未知环境是水声定位中的一项艰巨挑战。我们发现,尽管预训练模型的性能受到训练数据和测试数据之间不匹配的影响,但在不匹配较多的环境中,它们通常表现出较高的“隐含不确定性”。利用隐含不确定性的概念,我们将测试样本划分为更确定和更不确定的集合,并使用确定的样本来实现估计方法,以改善不确定样本的标记,这有助于适应模型。我们使用一种有效的方法来量化模型预测的不确定性,并使用一种创新的方法来使预训练的模型在测试时适应看不见的水下环境。这消除了对来自目标环境或原始训练数据的标记数据的需要。通过基于接收信号能量对独立估计进行积分来增强这种自适应。我们使用真实的实验数据以及由模型生成的信号与真实海洋噪声组成的合成数据来验证该方法。结果表明,模型预测准确性得到了显着提高,强调了该方法在多样化、嘈杂和未知环境中增强水下声学定位的潜力。
摘要:Adapting pre-trained deep learning models to new and unknown environments is a difficult challenge in underwater acoustic localization. We show that although pre-trained models have performance that suffers from mismatch between the training and test data, they generally exhibit a higher ``implied uncertainty'' in environments where there is more mismatch. Leveraging this notion of implied uncertainty, we partition the test samples into more certain and less certain sets, and implement an estimation method using the certain samples to improve the labeling for uncertain samples, which helps to adapt the model. We use an efficient method to quantify model prediction uncertainty, and an innovative approach to adapt a pre-trained model to unseen underwater environments at test time. This eliminates the need for labeled data from the target environment or the original training data. This adaptation is enhanced by integrating an independent estimate based on the received signal energy. We validate the approach extensively using real experimental data, as well as synthetic data consisting of model-generated signals with real ocean noise. The results demonstrate significant improvements in model prediction accuracy, underscoring the potential of the method to enhance underwater acoustic localization in diverse, noisy, and unknown environments.


【24】 CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with  LLM-Powered Text Description Sourcing and Mining
标题: CrossMuSim:一个基于LLM-Powered文本描述源和挖掘的跨模态音乐相似性检索框架
链接:https://arxiv.org/abs/2503.23128

作者: Tristan Tsoi,  Jiajun Deng,  Yaolong Ju,  Benno Weck,  Holger Kirchhoff,  Simon Lui 
备注:Accepted by ICME2025
摘要:音乐相似性检索是管理和探索流媒体平台中大型收藏中相关内容的基础。本文提出了一种新的跨模态对比学习框架,利用文本描述的开放性来指导音乐相似性建模,解决了传统单模态方法在捕捉复杂音乐关系方面的局限性。为了克服高质量的文本-音乐配对数据的稀缺性,本文介绍了一种结合在线抓取和基于LLM的提示的双源数据获取方法,其中精心设计的提示利用LLM的综合音乐知识来生成上下文丰富的描述。在华为Music流媒体平台上,通过客观指标、主观评价和真实A/B测试,验证了该框架在性能上的显著提升。
摘要:Music similarity retrieval is fundamental for managing and exploring relevant content from large collections in streaming platforms. This paper presents a novel cross-modal contrastive learning framework that leverages the open-ended nature of text descriptions to guide music similarity modeling, addressing the limitations of traditional uni-modal approaches in capturing complex musical relationships. To overcome the scarcity of high-quality text-music paired data, this paper introduces a dual-source data acquisition approach combining online scraping and LLM-based prompting, where carefully designed prompts leverage LLMs' comprehensive music knowledge to generate contextually rich descriptions. Exten1sive experiments demonstrate that the proposed framework achieves significant performance improvements over existing benchmarks through objective metrics, subjective evaluations, and real-world A/B testing on the Huawei Music streaming platform.


【25】 Dual Audio-Centric Modality Coupling for Talking Head Generation
标题: 用于说话的头部生成的双音频中心模式耦合
链接:https://arxiv.org/abs/2503.22728

作者: Ao Fu,  Ziqi Ni,  Yi Zhou 
备注:9 pages, 4 figures
摘要:音频驱动的说话头部视频的生成是计算机视觉和图形学中的一个关键挑战,其应用于虚拟化身和数字媒体。传统方法通常难以捕捉音频和面部动态之间的复杂交互,导致嘴唇同步和视觉质量问题。在本文中,我们提出了一种新的NeRF为基础的框架,双音频中心的模态耦合(DAMC),它有效地集成了音频输入的内容和动态功能。通过利用双编码器结构,DAMC通过内容感知编码器捕获语义内容,并通过动态同步编码器确保精确的视觉同步。这些功能融合使用交叉同步融合模块(CSFM),增强内容表示和唇同步。大量的实验表明,我们的方法优于现有的最先进的方法在关键指标,如唇同步精度和图像质量,表现出强大的泛化在各种音频输入,包括合成语音从文本到语音(TTS)系统。我们的研究结果提供了一个有前途的解决方案,高品质,音频驱动的说话头生成,并提出了一个可扩展的方法来创建逼真的说话头。
摘要:The generation of audio-driven talking head videos is a key challenge in computer vision and graphics, with applications in virtual avatars and digital media. Traditional approaches often struggle with capturing the complex interaction between audio and facial dynamics, leading to lip synchronization and visual quality issues. In this paper, we propose a novel NeRF-based framework, Dual Audio-Centric Modality Coupling (DAMC), which effectively integrates content and dynamic features from audio inputs. By leveraging a dual encoder structure, DAMC captures semantic content through the Content-Aware Encoder and ensures precise visual synchronization through the Dynamic-Sync Encoder. These features are fused using a Cross-Synchronized Fusion Module (CSFM), enhancing content representation and lip synchronization. Extensive experiments show that our method outperforms existing state-of-the-art approaches in key metrics such as lip synchronization accuracy and image quality, demonstrating robust generalization across various audio inputs, including synthetic speech from text-to-speech (TTS) systems. Our results provide a promising solution for high-quality, audio-driven talking head generation and present a scalable approach for creating realistic talking heads.


【26】 Risk-Calibrated Affective Speech Recognition via Conformal Coverage  Guarantees: A Stochastic Calibrative Framework for Emergent Uncertainty  Quantification
标题: 通过保形覆盖保证进行风险校准的情感语音识别:紧急不确定性量化的随机校准框架
链接:https://arxiv.org/abs/2503.22712

作者: Zijun Jia 
摘要:驾驶员极端情绪引发的交通安全挑战凸显了对可靠情绪识别系统的迫切需求。语音情感识别中的传统深度学习方法存在过拟合和校准不佳的置信度估计。我们提出了一个整合共形预测(CP)和风险控制的框架,使用通过预训练的卷积神经网络处理的Mel谱图特征。我们的主要创新是开发了一个不一致性评分,该评分可以精确地衡量分类器的预测与给定输入的一致程度。通过校准样本,我们计算这个分数,并根据用户指定的风险水平$\alpha$推导出一个统计上严格的阈值,构建具有可证明覆盖保证($\geq 1-\alpha$)的预测集。风险控制框架通过可定制的损失函数实现特定于任务的自适应,动态调整预测集大小,同时保持覆盖率保证。IEMOCAP和TESS的跨数据集实验表明:1)严格的覆盖保证,2)平均预测集大小(APSS)和$\alpha$之间显着的负相关性,揭示了在高风险条件下降低模型的不确定性。我们进一步提出APSS作为一种新的度量分类不确定性。该方法提高了语音情感识别的可靠性,可直接应用于智能交通系统和实时情感监控。
摘要:Traffic safety challenges arising from extreme driver emotions highlight the urgent need for reliable emotion recognition systems. Traditional deep learning approaches in speech emotion recognition suffer from overfitting and poorly calibrated confidence estimates. We propose a framework integrating Conformal Prediction (CP) and Risk Control,using Mel-spectrogram features processed through a pre-trained convolutional neural network. Our key innovation is the development of a nonconformity score that heuristically measures how closely a classifier's predictions align with given inputs. Through calibration samples, we compute this score and derive a statistically rigorous threshold based on user-specified risk level $\alpha$, constructing prediction sets with provable coverage guarantees ($\geq 1-\alpha$). The Risk Control framework enables task-specific adaptation through customizable loss functions, dynamically adjusting prediction set sizes while maintaining coverage guarantees. Cross-dataset experiments on IEMOCAP and TESS demonstrate: 1) Strict coverage guarantee, 2) Significant negative correlation between Average Prediction Set Size (APSS) and $\alpha$, revealing reduced model uncertainty under high-risk conditions. We further propose APSS as a novel metric for evaluating classification uncertainty. This approach enhances speech emotion recognition reliability, with direct applications in intelligent transportation systems and real-time emotion monitoring.


【27】 Modeling speech emotion with label variance and analyzing performance  across speakers and unseen acoustic conditions
标题: 利用标签方差建模语音情感并分析说话者和看不见的声学条件之间的性能
链接:https://arxiv.org/abs/2503.22711

作者: Vikramjit Mitra,  Amrit Romana,  Dung T. Tran,  Erdrin Azemi 
备注:11 pages, 5 figures
摘要:自发语音情感数据通常包含感知等级,其中分级者在听完语音文件后分配情感分数。这样的感知等级由于分级者意见变化而在标签中引入不确定性。评分员的差异是通过使用共识评分作为地面实况来解决的,其中具有最高投票的情感被选择。共识等级没有考虑模糊的情况下,语音样本可能包含多种情绪,通过分级意见的不确定性捕获。我们证明,使用的概率密度函数的情绪等级作为目标,而不是常用的共识等级,提供更好的性能基准评估集相比,在文献中报道的结果。我们发现,显着性驱动的基础模型(FM)表示选择有助于训练一个国家的最先进的语音情感模型的维度和类别的情感识别。比较从不同FM获得的表示,我们观察到,专注于整体测试集性能可能具有欺骗性,因为它无法揭示模型在说话者和性别之间的泛化能力。我们表明,跨多个测试集的性能评估和跨性别和扬声器的性能分析是有用的评估情感模型的有用性。最后,我们证明了标签的不确定性和数据倾斜对模型评估提出了挑战,而不是使用最佳假设,考虑2或3个最佳假设是有用的。
摘要:Spontaneous speech emotion data usually contain perceptual grades where graders assign emotion score after listening to the speech files. Such perceptual grades introduce uncertainty in labels due to grader opinion variation. Grader variation is addressed by using consensus grades as groundtruth, where the emotion with the highest vote is selected. Consensus grades fail to consider ambiguous instances where a speech sample may contain multiple emotions, as captured through grader opinion uncertainty. We demonstrate that using the probability density function of the emotion grades as targets instead of the commonly used consensus grades, provide better performance on benchmark evaluation sets compared to results reported in the literature. We show that a saliency driven foundation model (FM) representation selection helps to train a state-of-the-art speech emotion model for both dimensional and categorical emotion recognition. Comparing representations obtained from different FMs, we observed that focusing on overall test-set performance can be deceiving, as it fails to reveal the models generalization capacity across speakers and gender. We demonstrate that performance evaluation across multiple test-sets and performance analysis across gender and speakers are useful in assessing usefulness of emotion models. Finally, we demonstrate that label uncertainty and data-skew pose a challenge to model evaluation, where instead of using the best hypothesis, it is useful to consider the 2- or 3-best hypotheses.


机器翻译由腾讯交互翻译提供,仅供参考