今日论文合集:cs.SD语音43篇,eess.AS音频处理46篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
标题:Moonshine的味道:用于边缘设备的微型专用ASR模型
链接:https://arxiv.org/abs/2509.02523

摘要:我们提出了Moonshine的味道,一套专门用于一系列代表性不足的语言的微型自动语音识别(ASR)模型。主流观点认为,多语言ASR模型通过利用跨语言语音相似性而优于单语言模型。我们对这一假设提出了挑战,表明对于足够小的模型(27M参数),在高质量的人类标记,伪标记和合成数据的精心平衡组合上训练单语系统会产生显著的优越性能。平均而言,我们的模型实现的错误率比10倍大小的Whisper Tiny模型低48%,优于9倍大的Whisper Small模型,并且在大多数情况下匹配或优于28倍大的Whisper Medium模型。这些结果推进了这种规模模型的最新技术水平,为以前支持有限的语言实现了准确的设备上ASR。我们发布阿拉伯语,中文,日语,韩语,乌克兰语和越南语的Moonshine模型在一个宽松的开源许可证下。


【2】FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
标题:FLM-音频:自然独白通过双重训练改进原生的全复式聊天机器人
链接:https://arxiv.org/abs/2509.02521

摘要:全双工对话模型旨在同时收听和说话,并对快速变化的用户输入做出快速响应。在现有的方法中,本地全双工模型在单个时间步长中合并不同的信道(例如,听和说),克服了时分复用时分复用(TDM)替代方案固有的高响应延迟。然而,一个关键的挑战仍然存在:将文本独白与以不同比特率运行的音频流对齐。主流的解决方案依赖于单词级对齐,但这可能会降低大型预训练模型的语言能力。此外,它需要为每个令牌提供高度准确的时间戳,这会引入级联错误并增加预处理成本。在本文中,我们提出了连续令牌序列的文本独白,即“自然”独白,模仿人类的认知行为的对话。对于时间对齐,我们在不同的训练阶段交替自然独白的位置-领先或落后的音频。这种“双重”训练模式在构建FLM音频方面非常有效,我们的7 B口语对话模型展示了卓越的响应能力,准确性和聊天体验,实验结果证实了这一点。


【3】ESTM: An Enhanced Dual-Branch Spectral-Temporal Mamba for Anomalous Sound Detection
标题:ESTM:用于异常声音检测的增强型双分支频谱-时间曼巴
链接:https://arxiv.org/abs/2509.02471

备注:Accepted in IEEE Signal Processing Letters 2025
摘要:工业设备异常声检测的核心挑战在于对声学特征的时频耦合特性进行建模。现有的建模方法受到局部感受野的限制,难以捕捉机器声学特征中的长范围时间模式和跨频带动态耦合效应。在本文中,我们提出了一种新的框架,ESTM,这是基于一个双路径的Mamba架构与时间-频率解耦建模,并利用选择性状态空间模型(SSM)的远程序列建模。ESTM通过融合增强的Mel频谱图和原始音频特征,从不同的时间段和频带提取丰富的特征表示,同时通过TriStat-Gating(TSG)模块进一步提高对异常模式的灵敏度。我们的实验表明,ESTM提高了DCASE 2020任务2数据集上的异常检测性能,进一步验证了所提出的方法的有效性。


【4】TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
标题:TTA-Bench:一个用于评估文本到音频模型的综合基准
链接:https://arxiv.org/abs/2509.02398

摘要:文本到音频(TTA)的生成已经取得了迅速的进展,但目前的评估方法仍然狭窄,主要集中在感知质量,而忽略了鲁棒性,泛化和道德问题。我们提出了TTA-Bench,这是一个全面的基准,用于评估TTA模型的功能性能,可靠性和社会责任。它涵盖了准确性、鲁棒性、公平性和毒性等七个维度,包括通过自动化和手动方法生成的2,999个不同提示。我们引入了一个统一的评估协议,该协议将客观指标与来自专家和普通用户的118,000多个人工注释相结合。在此框架下,对10种最先进的模型进行了基准测试,详细了解了它们的优势和局限性。TTA-Bench为TTA系统的全面和负责任的评估建立了新的标准。数据集和评估工具在https://nku-hlt.github.io/tta-bench/上开源。


【5】AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
标题:AudioCodecBench:音频编解码器评估的全面基准
链接:https://arxiv.org/abs/2509.02349

摘要:多模态大语言模型(MLLM)在语音和音乐领域有着广泛的应用。这种趋势导致了对大型模型(LM)的音频标记化的关注。与仅语义的文本令牌不同,音频令牌必须捕获全局语义内容并保留细粒度的声学细节。此外,它们提供了一种用于语音和音乐的离散方法,可以有效地集成到MLLM中。然而,现有的研究是不合适的语义令牌和声学令牌的定义。此外,对不同编解码器的评估通常集中在特定的领域或任务上,例如重建或自动语音识别(ASR)任务,这妨碍了公平和全面的比较。为了解决这些问题,本文提供了合适的语义和声学令牌的定义,并介绍了一个系统的评估框架。该框架允许对编解码器的能力进行全面评估,其跨四个维度进行评估:音频重构度量、码本索引(ID)稳定性、仅解码器Transformer困惑以及下游探测任务的性能。我们的研究结果表明,所提供的合适的定义和重建度量,码本ID稳定性,下游探测任务和困惑之间的相关性的正确性。


【6】Speech transformer models for extracting information from baby cries
标题:用于从婴儿哭声中提取信息的语音Transformer模型
链接:https://arxiv.org/abs/2509.02259

备注:Accepted to WOCCI2025 (interspeech2025 workshop)
摘要:使用来自预训练语音模型的潜在表示的迁移学习在标记数据稀缺的任务中实现了出色的性能。然而,它们对非语音数据的适用性以及在这些表示中编码的特定声学特性在很大程度上仍未被探索。在这项研究中,我们调查了这两个方面。我们在8个婴儿哭声数据集上评估了5个预训练的语音模型,其中包括来自960个婴儿的115小时音频。对于每个数据集,我们在所有可用的分类任务中评估每个模型的潜在表示。我们的研究结果表明,这些模型的潜在表示可以有效地分类人类婴儿的哭声和编码的关键信息相关的声源不稳定性和身份的哭泣的婴儿。此外,这些模型的架构和训练策略的比较为设计未来针对类似任务(如情感检测)的模型提供了有价值的见解。


【7】Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
标题:频谱图补丁编解码器:用于神经语音编码的2D块量化VQ-VAE和HiFi-GAN
链接:https://arxiv.org/abs/2509.02244

摘要:我们提出了一种神经语音编解码器,通过引入一种更简单的单级量化方法来挑战对复杂残差矢量量化(RVQ)堆栈的需求。我们的方法直接对梅尔频谱图进行操作,将其视为2D数据,并将非重叠的4x 4补丁量化为单个共享码本。这种拼接式设计简化了架构,实现了低延迟流,并产生了离散的潜在网格。为了确保高保真度的合成,我们采用了后期对抗微调的VQ-VAE和训练一个HiFi-GAN声码器从头开始的编解码器的重建频谱图。在大约7.5 kbits/s的16 kHz的语音操作,我们的系统进行了评估,对几个国家的最先进的神经编解码器使用客观的指标,如STOI,PESQ,MCD,和ViSQOL。结果表明,我们的简化,无残留的架构实现了有竞争力的感知质量和可理解性,验证它作为一个有效的和开放的基础,为未来的低延迟编解码器设计。


【8】AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
标题:AudioRWKV:用于音频模式识别的高效稳定的双向RWKV
链接:https://arxiv.org/abs/2509.02167

备注:6 pages, 3 figures
摘要:最近,Transformers(例如,音频频谱图Transformers,AST)和状态空间模型(例如,Audio Mamba,AuM)在音频建模方面取得了显著的进展。然而,Transformer架构的O(L^2)计算复杂度阻碍了有效的长序列处理,而Mamba架构在缩放参数和数据时往往变得不稳定。为了应对这些挑战,本文提出了AudioRWKV(A-RWKV),一个高效,稳定的音频建模架构。具体来说,我们继承了RWKV 7的稳定和高效的递归公式,并将其1D标记移位操作替换为2D dependency可分离卷积,以更好地捕获局部频谱-时间模式。此外,我们将原始的因果WKV内核调整为双向WKV内核(Bi-WKV),从而在整个音频序列上实现全局上下文建模,同时保持线性计算复杂度。得益于RWKV 7基础的固有稳定性,A-RWKV可无缝扩展到更大的模型尺寸。实验结果表明,在相同的线性模型机制下,A-RWKV-S(22 M)实现了与AuM-B(92 M)相同的性能,同时表现出比AST更稳定的吞吐量;对于长格式音频(约5分28秒),WKV 7在处理方面实现了高达13.3倍的加速。


【9】NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
标题:NADI 2025:首个多方言阿拉伯语语音处理共享任务
链接:https://arxiv.org/abs/2509.02038

摘要:我们提出了第六个细微差别的阿拉伯语方言识别(NADI 2025)共享任务的研究结果,该任务侧重于三个子任务的阿拉伯语语音方言处理:口语方言识别(子任务1),语音识别(子任务2)和口语方言的变音符号恢复(子任务3)。共有44个团队注册,在测试阶段,从8个独特的团队收到了100份有效的提交材料。分配如下:子任务1“五个团队”提交了34份,子任务2“六个团队”提交了47份,子任务3“两个团队”提交了19份。性能最好的系统在子任务1上实现了79.8%的准确率,在子任务2上实现了35.68/12.20 WER/CER(总体平均值),在子任务3上实现了55/13 WER/CER。这些结果突出了阿拉伯方言语音处理,特别是在方言识别,识别和变音符号恢复的持续挑战。我们还总结了参与团队所采用的方法,并简要概述了NADI未来版本的方向。


【10】FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
标题:FireRedRTS-2:播客和Chatbot的长对话语音生成
链接:https://arxiv.org/abs/2509.02020

摘要:目前的对话生成方法通常需要完整的对话文本合成之前,并产生一个单一的,不可分割的语音包含所有的声音,使他们不适合交互式聊天,此外,他们遭受不稳定的合成,不准确的扬声器过渡,和不连贯的韵律。在这项工作中,我们提出了FireRedTTS-2,一个长格式的流TTS系统,多扬声器对话生成,提供稳定,自然的语音与可靠的扬声器切换和上下文感知韵律。一个新的12.5Hz流语音标记器加速了训练和推理,延长了最大对话长度,编码了更丰富的语义以稳定文本到令牌的建模,并支持实时应用的高保真流生成。我们采用了文本语音交错格式,连接扬声器标记的文本与对齐的语音令牌按时间顺序排列,并建模与双Transformer:一个大的解码器只Transformer预测令牌在第一层,和一个较小的完成后续层。实验结果表明,FireRedTTS-2与聊天框架无缝集成,并以最小的微调,产生情感表达的语音由隐式上下文线索的指导。在播客生成中,它超越了现有的系统,包括MoonCast,Zipvoice-Dialogue和MOSS-TTSD在客观的可懂度,说话者转向的可靠性,和感知自然与上下文一致的韵律。我们的演示可在https://fireredteam.github.io/demos/firered_tts_2上获得。


【11】Music Genre Classification Using Machine Learning Techniques
标题:使用机器学习技术的音乐流派分类
链接:https://arxiv.org/abs/2509.01762

备注:10 pages, 20 figures. Submitted in partial fulfillment of the requirements for the Bachelor of Technology (this http URL) degree in Artificial Intelligence and Data Science
摘要:本文提出了一个比较分析的机器学习方法自动音乐流派分类。我们评估了经典分类器的性能,包括支持向量机(SVM)和集成方法,训练一套全面的手工制作的音频功能,对卷积神经网络(CNN)梅尔频谱图操作。该研究在广泛使用的GTZAN数据集上进行。我们的研究结果证明了一个值得注意的结果:与端到端CNN模型相比,利用特定领域特征工程的SVM实现了更高的分类准确性。我们将这一结果归因于基准数据集的数据约束性质,其中工程特征的强归纳偏差提供了正则化效应,减轻了高容量深度学习模型中固有的过度拟合风险。这项工作强调了传统特征提取在实际音频处理任务中的持久相关性,并为深度学习的普遍适用性提供了一个重要的视角,特别是对于中等规模的数据集。


【12】From Discord to Harmony: Decomposed Consonance-based Training for Improved Audio Chord Estimation
标题:从不和谐到和谐:用于改进音频和弦估计的分解基于协和的训练
链接:https://arxiv.org/abs/2509.01588

备注:9 pages, 3 figures, 3 tables
摘要:音频和弦估计(ACE)在音乐信息研究中发挥着关键作用,由于其与音乐转录和分析的相关性,二十多年来一直受到关注。尽管取得了显著的进步,但任务中仍然存在挑战,特别是关于谐波含量的独特特性,这导致现有系统的性能达到玻璃天花板。这些挑战包括注释者的主观性,其中注释者之间的不同解释导致不一致,以及和弦数据集内的类不平衡,其中某些和弦类与其他和弦类相比被过度表示,这给模型训练和评估带来了困难。作为第一个贡献,本文提出了一个评价的注释者之间的协议,在和弦注释,使用的指标,超越传统的二进制措施。此外,我们提出了一个和谐的距离度量,反映了和谐注释之间的感知相似性。我们的分析表明,基于和谐的距离度量更有效地捕捉音乐有意义的协议之间的注释。扩展这些发现,我们引入了一种新的ACE一致性为基础的模型,整合到模型中的概念,通过基于辅音的标签平滑的辅音。该模型还通过分别估计根音、低音和所有音符激活来解决类不平衡问题,从而能够从分解的输出中重建和弦标签。


【13】ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition
标题:Arabspel Net:一个轻量级混合2D CNN-BiLSTM模型,注重稳健的阿拉伯语语音情感识别
链接:https://arxiv.org/abs/2509.01401

备注:Accepted (The Third Arabic Natural Language Processing Conference)
摘要:语音情感识别对于人机交互至关重要,特别是对于像阿拉伯语这样的低资源语言,由于有限的数据和研究而面临挑战。我们引入ArabnetNet,这是一种轻量级架构,旨在克服这些限制并提供最先进的性能。与之前依赖于离散MFCC特征和1D卷积的系统不同,这些系统错过了细微的频谱-时间模式,ArabaNet使用通过2D卷积处理的Mel频谱图,保留了传统方法中经常丢失的关键情感线索。   虽然最近的模型倾向于拥有数百万个参数的大规模架构,但ArabnetNet仅用100万个参数就取得了优异的结果,比HuBERT base小90倍,比Whisper小74倍。这种效率使其成为资源受限环境的理想选择。阿拉伯语网络推进阿拉伯语语音情感识别,为现实世界的应用程序提供卓越的性能和可访问性。


【14】CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays
标题:CabinSep:基于IR增强屏蔽的MVDR,通过分布式异类阵列实现实时车内语音分离
链接:https://arxiv.org/abs/2509.01399

备注:Accepted by Interspeech 2025
摘要:从多个说话者中分离重叠语音对于有效的人车交互至关重要。本文提出了CabinSep,一种轻量级的基于神经掩码的最小方差无失真响应(MVDR)语音分离方法,以减少后端自动语音识别(ASR)模型中的语音识别错误。我们的贡献有三个方面:首先,我们利用信道信息来提取空间特征,这提高了语音和噪声掩模的估计。其次,我们在推理过程中使用MVDR,减少语音失真,使其对ASR更友好。第三,我们介绍了一种数据增强方法相结合的模拟和真实记录的脉冲响应(IR),提高扬声器定位在区域边界,进一步减少语音识别错误。CabinSep的计算复杂度仅为0.4 GMAC,与最先进的DualSep模型相比,在真实记录的数据集中,语音识别错误率相对降低了17.5%。演示可在https://cabinsep.github.io/cabinsep/上获得。


【15】The AudioMOS Challenge 2025
标题:2025年AudioMOS挑战赛
链接:https://arxiv.org/abs/2509.01336

备注:IEEE ASRU 2025
摘要:这是AudioMOS Challenge 2025的总结论文,这是合成音频自动主观质量预测的第一个挑战。挑战包括三条轨道。第一条轨道旨在评估文本到音乐样本的整体质量和文本对齐。第二个轨道是基于Meta Audiobox Aesthetics的四个评估维度,测试集由文本到语音,文本到音频和文本到音乐样本组成。第三部分是不同采样率下的合成语音质量评价。这项挑战吸引了来自学术界和工业界的24个独特团队,并确认了基线的改进。这一挑战的结果预计将促进音频生成系统自动评估领域的发展和进步。


【16】SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation
标题:Simulega:MoE路由器是同步语音翻译的高级政策制定者
链接:https://arxiv.org/abs/2509.01200

摘要:同步语音翻译(SimulST)通过在严格的延迟限制下联合优化语音识别和机器翻译,实现实时跨语言通信。现有系统难以平衡翻译质量、延迟和语义一致性,特别是在多语言多对多场景中,不同的读写策略阻碍了统一的策略学习。在本文中,我们提出了SimulMEGA(通过混合专家门控同时生成),这是一个无监督的策略学习框架,它将基于前缀的训练与混合专家细化器相结合,以隐式方式学习有效的读写决策,而不会增加推理时间开销。我们的设计只需要对标准Transformer架构进行最小限度的修改,并且可以通用于语音转文本和文本转语音流传输任务。通过对六种语言对的综合评估,我们的500 M参数语音到文本模型优于Seamless基线,在1.5秒的平均延迟下实现了低于7%的BLEU降级,在3秒内实现了低于3%的BLEU降级。我们进一步证明了SimulMEGA的多功能性,通过将其扩展到具有单向骨干的流式TTS,从而获得卓越的延迟质量权衡。


【17】EZhouNet:A framework based on graph neural network and anchor interval for the respiratory sound event detection
标题:EZhouNet:基于图神经网络和锚点区间的呼吸音事件检测框架
链接:https://arxiv.org/abs/2509.01153

摘要:听诊是呼吸系统和肺部疾病早期诊断的关键方法,依赖于熟练的医疗保健专业人员。然而,这一过程往往是主观的,专家之间存在差异。因此,出现了许多基于深度学习的自动分类方法,其中大多数侧重于呼吸声分类。相比之下,对呼吸声事件检测的研究仍然有限。现有的声音事件检测方法通常依赖于帧级预测,然后进行后处理以生成事件级输出,这使得直接学习间隔边界具有挑战性。此外,许多方法只能处理固定长度的音频,限制了它们对可变长度呼吸声的适用性。此外,呼吸声位置信息对检测性能的影响尚未得到广泛研究。为了解决这些问题,我们提出了一个基于图神经网络的框架与锚定间隔,能够处理可变长度的音频,并提供更精确的时间定位异常呼吸声音事件。我们的方法提高了呼吸音检测的灵活性和适用性。在SPRSound 2024和HF Lung V1数据集上的实验证明了该方法的有效性,并结合呼吸位置信息增强了异常声音之间的区分。


【18】A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR
标题:自我监督孟加拉方言ASB的统一去噪和适应框架
链接:https://arxiv.org/abs/2509.00988

摘要:孟加拉语是世界第五大语言,其自动语音识别(ASR)仍然是一个重大挑战,严重阻碍了超过2.7亿人的技术可访问性。这一挑战由于两个持续存在且相互交织的因素而变得更加复杂:语言的巨大方言多样性和现实环境中普遍存在的噪音。虽然最先进的自监督学习(SSL)模型为低资源语言提供了先进的ASR,但它们通常缺乏明确的机制来处理预训练期间的环境噪声或针对孟加拉语方言中复杂语音和词汇变化的专门适应策略。本文介绍了一种新颖的,统一的框架,旨在同时解决这些双重挑战。我们的方法建立在WavLM模型的基础上,WavLM模型是用掩蔽语音去噪目标进行唯一预训练的,使其对声学失真具有固有的鲁棒性。我们提出了一种专门的多阶段微调策略,首先将模型适应通用领域标准孟加拉语,以建立强大的语言基础,然后通过有针对性的数据增强将其专门用于噪声鲁棒性方言识别。该框架在广泛的模拟噪声条件下,从干净的音频到低信噪比(SNR)水平,在包括多种孟加拉方言的综合基准上进行了严格的评估。   实验结果表明,该框架显着优于强基线,包括标准微调wav2vec 2.0和大规模多语言Whisper模型。这项工作为这项任务建立了一个新的最先进的技术,并为全球其他低资源,高变化的语言开发实用的ASR系统提供了一个可扩展的,有效的蓝图。


【19】IoT-based Noise Monitoring using Mobile Nodes for Smart Cities
标题:使用移动节点实现智能城市基于物联网的噪音监控
链接:https://arxiv.org/abs/2509.00979

摘要:城市噪声污染对公众健康构成重大威胁,但现有的监测基础设施提供有限的空间覆盖范围和适应性。本文提出了一种可扩展的、低成本的、基于物联网的、使用移动节点(移动车辆上的传感器节点)的实时环境噪声监测解决方案。该系统利用一个低成本的声音传感器与GPS功能模块集成,以一秒的间隔收集地理标记的噪音数据。声音节点在实验室环境中根据参考声级计进行校准,以使用各种机器学习(ML)算法确保准确性,例如简单线性回归(SLR)、多元线性回归(MLR)、多项式回归(PR)、分段回归(SR)、支持向量回归(SVR)、决策树(DT)和随机森林回归(RFR)。虽然实验室校准证明了高精度,它表明,节点的性能下降,在移动车辆中的数据收集过程中。为了解决这个问题,证明了必须基于在移动环境中与参考设备一起收集的数据在基于IoT的节点上执行校准。在所采用的ML模型中,RFR实现了最佳性能,R2为0.937,RMSE为1.09。该系统部署在印度海得拉巴,通过27天的三次测量活动,捕获了436,420个数据点。结果突出了工作日,周末和排灯节期间的时间和空间噪声变化。将车辆速度纳入校准中显著提高了准确性。该系统展示了在智慧城市中广泛部署基于物联网的噪声传感网络的潜力,从而实现有效的噪声污染管理和城市规划。


【20】TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
标题:TinyMusician:基于知识蒸馏和混合精确量化的设备上音乐生成
链接:https://arxiv.org/abs/2509.00914

备注:12 pages for main context, 5 figures
摘要:生成模式的成功在音乐生成领域引起了前所未有的关注。基于Transformer的架构为模型性能设定了新的基准。然而,它们的实际采用受到一些关键挑战的阻碍:由于它们的大量参数,需要大量的计算资源和推理时间。这些障碍使它们无法部署在计算资源有限的边缘设备(例如智能手机和可穿戴设备)上。在这项工作中,我们提出了TinyMusician,一个轻量级的音乐生成模型从MusicGen(一个国家的最先进的音乐生成模型)蒸馏。TinyMusician集成了两项创新:(i)阶段混合双向和倾斜KL发散和(ii)自适应混合精度量化。实验结果表明,TinyMusician保留93%的MusicGen-Small性能,模型大小减少55%。TinyMusician是第一个可移动部署的音乐生成模型,消除了对云的依赖,同时保持了高音频保真度和高效的资源使用率


【21】Speech Command Recognition Using LogNNet Reservoir Computing for Embedded Systems
标题:嵌入式系统中使用LogNNet水库计算的语音命令识别
链接:https://arxiv.org/abs/2509.00862

备注:20 pages, 6 figures
摘要:本文提出了一种低资源的语音命令识别器,结合基于能量的语音活动检测(VAD),优化的梅尔频率倒谱系数(MFCC)管道,和LogNNet算法计算分类器。使用四个命令从语音命令da-taset下采样到8 kHz,我们评估四个MFCC聚合方案,并发现自适应分仓(64维特征向量)提供了最好的精度紧凑性权衡。架构为64:33:9:4的LogNNet分类器在说话人独立评估下达到了92.04%的准确率,同时需要的参数比传统的深度学习模型少得多。在Arduino Nano 33 IoT(ARM Cor-tex-M0+,48 MHz,32 KB RAM)上的硬件实现验证了实际可行性,实现了约90%的实时识别准确率,同时仅消耗18 KB RAM(55%利用率)。因此,完整的管道(VAD -> MFCC -> LogNNet)可以在严格的内存和计算限制下实现可靠的设备上语音命令识别,使其适用于电池供电的物联网节点、无线传感器网络和免提控制接口。


【22】Adaptive Vehicle Speed Classification via BMCNN with Reinforcement Learning-Enhanced Acoustic Processing
标题:通过BMCNN和强化学习增强声学处理的自适应车辆速度分类
链接:https://arxiv.org/abs/2509.00839

摘要:交通拥堵仍然是一个紧迫的城市挑战,需要智能交通系统进行实时管理。我们提出了一个混合框架,结合了深度学习和强化学习的声学车辆速度分类。双分支BMCNN处理MFCC和小波特征以捕获互补频率模式。注意力增强的DQN自适应地选择最小数量的音频帧,并在达到置信度阈值时触发早期决策。对IDMT-Traffic和我们的SZUR-Acoustic(苏州)数据集的评估显示,准确率分别为95.99%和92.3%,通过提前终止,平均处理速度提高了1.63倍。与A3 C、DDDQN、SA 2C、PPO和TD 3相比,该方法提供了优越的精度-效率权衡,适合于异构城市环境中的实时ITS部署。


【23】AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation
标题:Aimocilles:评估文本到音乐生成中情感传递的基准
链接:https://arxiv.org/abs/2509.00813

备注:to be published in HCMIR25: 3rd Workshop on Human-Centric Music Information Research
摘要:文本到音乐(TTM)生成方面的最新进展使使用自然语言提示进行可控且富有表现力的音乐创作成为可能。然而,与人类偏好或文本对齐相比,TTM系统的情感保真度在很大程度上仍然未被探索。在这项研究中,我们介绍了AImoclips,这是一个用于评估TTM系统如何向人类听众传达预期情感的基准,涵盖了开源和商业模式。我们选择了12个情绪意图,跨越了效价-唤醒空间的四个象限,并使用了六个最先进的TTM系统来生成超过1,000个音乐片段。共有111名参与者在9点Likert量表上对每个剪辑的感知效价和唤醒进行了评分。我们的研究结果表明,商业系统往往会产生比预期更令人愉快的音乐,而开源系统往往会表现相反。在所有模型中,情绪在高唤醒条件下都能更准确地传达。此外,所有的系统都表现出对情感中立的偏见,突出了情感可控性的关键限制。该基准测试为模型特定的情感渲染特性提供了有价值的见解,并支持情感对齐TTM系统的未来发展。


【24】PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
标题:PicoAudio2:自然语言描述的时间可控文本到音频生成
链接:https://arxiv.org/abs/2509.00683

备注:Demo page: this https URL
摘要:可控文本到音频生成(TTA)最近引起了人们的广泛关注。虽然现有的作品可以实现细粒度的可控性的基础上的时间戳信息,声音事件类别被限制在一个固定的集合。此外,由于仅使用模拟数据进行训练,因此所生成的音频质量和对真实数据的泛化性能受到限制。为了解决这个问题,我们提出了PicoAudio2,通过新的数据处理管道和模型架构来改进时间可控的TTA。具体来说,我们使用接地模型来注释真实音频文本数据集的事件时间戳,以管理时间上强的真实数据,以及现有作品的模拟数据。该模型是在真实和模拟数据的组合上训练的。此外,在PicoAudio之后,我们将时间戳信息编码到时间戳矩阵中,以在粗粒度文本描述之上为模型提供额外的细粒度时间对齐信息。实验表明,PicoAudio2在时间可控性和音频质量方面表现出优越的性能。


【25】The Name-Free Gap: Policy-Aware Stylistic Control in Music Generation
标题:无名差距:音乐世代中的政策意识风格控制
链接:https://arxiv.org/abs/2509.00654

备注:10 pages, 2 figures
摘要:文本到音乐模型捕捉广泛的属性,如乐器或情绪,但细粒度的风格控制仍然是一个开放的挑战。现有的风格化方法通常需要重新训练或专门的条件反射,这使得再现性变得复杂,并且在艺术家姓名受到限制时限制了政策合规性。我们研究是否轻量级的,人类可读的修饰符采样从一个大型的语言模型可以提供一个策略强大的替代风格控制。使用MusicGen-small,我们评估两个艺术家:Billie Eilish(声乐流行)和Ludovico Einaudi(器乐钢琴)。对于每个艺术家,我们使用15个参考摘录,并在三个条件下评估匹配的种子:基线提示,艺术家姓名提示和五个描述符集。所有提示都是使用大型语言模型生成的。评估使用VGGish和CLAP嵌入以及分布和每个剪辑的相似性度量,包括新的最小距离归因度量。结果表明,艺术家的名字是最强的控制信号,在这两个艺术家,而无名称的描述符恢复这种效果。这突出表明,现有的保障措施,如在音乐生成提示中限制艺术家姓名,可能无法完全防止风格模仿。跨艺术家转移减少对齐,表明描述符编码有针对性的风格线索。我们还提出了一个描述符表在十个当代艺术家,以说明令牌的广度。这些发现共同定义了无名称的差距,艺术家的名字提示和符合政策的描述符之间的可控性差异,通过一个可重复的评估协议,为艺术家级别的可控性。


【26】Real-Time Piano Note Frequency Detection Using FPGA and FFT Core
标题:利用现场可编程逻辑器件和快速傅里叶变换核实现钢琴音符频率检测
链接:https://arxiv.org/abs/2509.00589

备注:20 pages, 11 Figures
摘要:钢琴等乐器的实时频率分析是电子调谐器、音乐可视化器和现场声音监控等领域的重要功能。传统方法通常依赖于基于软件的数字信号处理(DSP),这可能会引入延迟并需要大量的计算能力。相比之下,诸如FPGA(现场可编程门阵列)的硬件平台由于其并行处理能力而提供了以更快的速度和确定性执行此类分析的能力。该项目的主要目标是使用基于FPGA的实时快速傅立叶变换(FFT)系统分析来自数字钢琴的模拟音频信号。


【27】SaD: A Scenario-Aware Discriminator for Speech Enhancement
标题:SaD:语音增强的场景感知识别器
链接:https://arxiv.org/abs/2509.00405

备注:5 pages, 2 this http URL by InterSpeech2025
摘要:基于生成对抗网络的模型在语音增强领域表现出了卓越的性能。然而,这些模型的当前优化策略主要集中在改进生成器的架构或提高质量评估指标的可重用性。这种方法往往忽略了不同场景中固有的丰富上下文信息。在本文中,我们提出了一个感知的语音识别,捕捉场景的特定功能,并执行频域划分,从而使生成器生成的增强语音的质量评估更准确。我们使用两个公开的数据集对三个代表性模型进行了全面的实验。结果表明,我们的方法可以有效地适应各种生成器架构,而不会改变它们的结构,从而在不同的场景中解锁语音增强的进一步性能增益。


【28】Towards High-Fidelity and Controllable Bioacoustic Generation via Enhanced Diffusion Learning
标题:通过增强的扩散学习实现高保真和可控的生物声学生成
链接:https://arxiv.org/abs/2509.00318

摘要:生成建模为生物声学提供了新的机会,可以合成逼真的动物发声,从而支持生物监测工作并补充濒危物种的稀缺数据。然而,直接从嘈杂的现场录音中产生鸟鸣波形仍然是一个重大挑战。   我们提出了BirdDiff,一个生成框架,旨在从12种野生鸟类的嘈杂数据集中合成鸟的叫声。该模型采用了一个“零层”阶段的多尺度自适应鸟叫增强,其次是一个基于扩散的发电机条件下的三种模式:梅尔频率倒谱系数,物种标签,和文字描述。增强阶段提高了信噪比(SNR),同时最大限度地减少频谱失真,与三种广泛使用的非训练增强方法相比,实现了最高的SNR增益(+10.45 dB)和最低的Itakura-Saito距离(0.54)。   我们根据基线生成模型DiffWave评估BirdDiff。我们的方法在生成质量度量方面产生了实质性的改进:Fr 'echet音频距离(0.590至0.213),Jensen-Shannon发散度(0.259至0.226)和统计上不同的Bins数量(7.33至5.58)。为了评估物种特异性细节保留,我们使用在原始数据集上训练的ResNet 50分类器来识别生成的样本。分类准确率从35.9%(DiffWave)提高到70.1%(BirdDiff),12个物种中有8个超过70%的准确率。   这些结果表明,BirdDiff能够直接从嘈杂的现场录音中产生高保真、可控的鸟鸣。


【29】Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
标题:评估Wav 2 Vec 2.0、XLS-R和Whisper中Transformer层对说话人识别任务的有效性
链接:https://arxiv.org/abs/2509.00230

摘要:本研究评估了三种先进的语音编码器模型,Wav 2 Vec 2.0,XLS-R,和耳语,在说话人识别任务的性能。通过微调这些模型并使用SVCCA、k-means聚类和t-SNE可视化分析其分层表示,我们发现Wav 2 Vec 2.0和XLS-R在其早期层中有效地捕获了特定于说话者的特征,并通过微调提高了稳定性和性能。Whisper在更深的层中显示出更好的性能。此外,我们确定了每个模型的Transformer层的最佳数量时,微调说话人识别任务。


【30】Generalizable Audio Spoofing Detection using Non-Semantic Representations
标题:使用非语义表示的可推广音频欺骗检测
链接:https://arxiv.org/abs/2509.00186

备注:None
摘要:生成建模的快速发展使得合成音频生成变得容易,使得基于语音的服务容易受到欺骗攻击。因此,现在比以往任何时候都迫切需要强有力的反措施。现有的deepfake检测解决方案经常被批评缺乏通用性,并且在应用于真实世界数据时会严重失败。本研究提出一种利用非语义通用音频表示的通用欺骗检测新方法。大量的实验已经进行了使用TRILL和TRILLsson模型找到合适的非语义特征。结果表明,所提出的方法在域内测试集上实现了相当的性能,同时在域外测试集上显著优于最先进的方法。值得注意的是,它在公共领域数据上表现出了卓越的泛化能力,超越了基于手工制作的功能,语义嵌入和端到端架构的方法。


【31】CoComposer: LLM Multi-agent Collaborative Music Composition
标题:联合作曲家:LLM多主体协作音乐作曲
链接:https://arxiv.org/abs/2509.00132

摘要:现有的人工智能音乐创作工具在生成时间、音乐质量和可控性方面受到限制。我们介绍CoComposer,一个多代理系统,由五个合作代理,每个任务的基础上,传统的音乐创作工作流程。使用AudioBox美学系统,我们实验评估CoComposer的四个组成标准。我们使用三个LLM(GPT-4 o,DeepSeek-V3-0324,Gemini-2.5-Flash)进行测试,发现(1)CoComposer在音乐质量方面优于现有的基于多智能体LLM的系统,(2)与单智能体系统相比,在生产复杂性方面。与非LLM MusicLM相比,CoComposer具有更好的可解释性和可编辑性,尽管MusicLM仍然产生更好的音乐。


【32】Algorithms for Collaborative Harmonization
标题:协作协调算法
链接:https://arxiv.org/abs/2509.00120

备注:Presented at the 15th Multidisciplinary Workshop on Advances in Preference Handling M-PREF 2024, Santiago de Compostela, Oct 20, 2024
摘要:我们认为,在音乐和谐领域的文本聚合的一个特定的场景。音乐和声与文本聚合有相似之处,但和声语言比一般文本更有结构性。具体地说,给定一个给定的音乐旋律的一组和声建议,我们的兴趣在于设计聚合算法,产生一个和声序列,满足以下两个关键标准:(1)集体建议的有效表示;(2)音乐上连贯的和声。我们提出了不同的算法聚合的谐波由一组代理商,并分析其复杂性。结果表明,Kemeny和基于复数的算法是最有效的评估代表性和保持音乐的连贯性。


【33】A Survey on Evaluation Metrics for Music Generation
标题:音乐生成评估指标调查
链接:https://arxiv.org/abs/2509.00051

备注:19 pages, 2 figures
摘要:尽管音乐生成系统取得了显著的进步,但由于音乐的复杂性质,用于评估所生成的音乐的方法没有如预期的那样发展,其中包括结构、连贯性、创造性和情感表现力等方面。在本文中,我们揭示了这一研究差距,介绍了一个详细的分类评估指标的音频和符号音乐表示。我们包括一个批判性的审查,确定目前的评估方法,其中包括客观指标和人类感知之间的相关性差,跨文化偏见,缺乏标准化,阻碍跨模型比较的主要局限性。针对这些差距,我们进一步提出了未来的研究方向,建立一个全面的评价框架,音乐生成的评价。


【34】From Sound to Sight: Towards AI-authored Music Videos
标题:从声音到视觉:走向人工智能创作的音乐视频
链接:https://arxiv.org/abs/2509.00029

备注:1st Workshop on Generative AI for Storytelling (AISTORY), 2025
摘要:传统的音乐可视化系统依赖于手工制作的形状和颜色的临时变换,这些变换仅提供有限的表现力。我们提出了两个新的管道,用于使用现成的深度学习模型从任何用户指定的声乐或器乐歌曲自动生成音乐视频。受音乐视频制作人手动工作流程的启发,我们实验了基于潜在特征的技术如何分析音频以检测音乐品质,如情感线索和乐器模式,并使用语言模型将其转换为文本场景描述。接下来,我们采用生成模型来生成相应的视频剪辑。为了评估生成的视频,我们确定了几个关键方面,并设计和进行了初步的用户评估,展示了讲故事的潜力,视觉连贯性和情感与音乐的一致性。我们的研究结果强调了潜在特征技术和深度生成模型在传统方法之外扩展音乐可视化的潜力。


【35】Prospects for acoustically monitoring ecosystem tipping points
标题:声学监测生态系统临界点的前景
链接:https://arxiv.org/abs/2509.02201

备注:44 pages (including Supporting Information), 1 figure. Review article submitted to Global Change Biology
摘要:许多生态系统可能会发生重要的质的变化,包括突然过渡到其他稳定状态,以应对扰动或条件的增加。在这种“临界点”之前,生态系统的复原力,即从扰动中恢复的能力,往往会下降,从而留下各种空间和时间特征。这些所谓的“早期预警信号”已被用于预测不同真实系统中的过渡,但许多正在改变生态的高通量自主监测技术尚未完全利用。特别是声学监测是量化生物多样性、跟踪生态系统健康和促进保护的有力工具。通过在不同的环境中部署声学记录器,研究人员从单个物种的叫声和行为中获得了更高层次的声景特征,这些特征描述了栖息地质量,甚至预测了物种的发生。在这里,我们借鉴理论和实践,倡导使用声学来探测生态系统的弹性,并确定新兴的和既定的预警信号的临界点。重点放在务实的考虑,我们强调,尽管限制临界点理论和目前的规模和数据的可转移性,声学可以在理解不同的生态系统和规模的弹性和小费的潜力。


【36】AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
标题:AHAMaskk:无指令的大型音频语言模型的可靠任务规范
链接:https://arxiv.org/abs/2509.01787

备注:15 pages, 7 tables, 6 figures
摘要:虽然目前的大型音频语言模型(LALM)扩展了文本大型语言模型(LLM)与通用的声学理解能力,他们通常遭受指令敏感性,其中相同意图的不同指令可以产生截然不同的结果。在这项工作中,我们提出了AHAMask,在那里我们只是简单地屏蔽了一些注意头的解码器,只有LLM骨干的LALM,触发特定的声学任务功能,没有指令。这些掩码通过在LALM上训练而有效地获得,其中可训练参数的数量等于其LLM骨干中的注意力头数。我们通过实验表明,应用这种选择性注意头罩实现可比的,甚至更好的性能比使用指令,无论是在单一或复合任务。除了实现可靠的声学任务规范的LALM,这也表明,LALM表现出一定的“功能通路”在他们的注意头


【37】Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
标题:基于三阶段训练的噪音去纠缠实现抗噪语音识别
链接:https://arxiv.org/abs/2509.01087

备注:11 pages,4 figures
摘要:为了提高端到端(E2 E)语音识别系统在噪声或低信噪比(SNR)条件下的性能,本文介绍了NoisyD-CT,一种新的三阶段训练框架,建立在Conformer-Transducer架构上。NoisyD-CT的核心是一个专门设计的紧凑型噪声解纠缠(NoisyD)模块(仅添加1.71 M参数),集成在Conformer模块和换能器解码器之间,可在具有挑战性的声学噪声环境中执行深度噪声抑制并提高ASR鲁棒性。为了充分利用NoisyD-CT的噪声抑制能力,我们进一步提出了一个干净的表示一致性损失对齐高层次的表示来自嘈杂的语音与那些从相应的干净的语音。与噪声重建损失一起,这种一致性对齐使NoisyD模块能够有效地抑制噪声,同时在干净和嘈杂的条件下保持基本的声学和语言特征一致,从而产生更清晰的内部表示,增强ASR性能。此外,我们的三阶段训练策略旨在在整个模型训练过程中充分利用噪声解纠缠和语音识别模块的功能,最终在噪声条件下最大限度地提高性能。我们的实验是在LibriSpeech和CHiME-4数据集上进行的,大量的结果表明,我们提出的NoisyD-CT显著优于竞争对手的Conformer-Transducer基线,在模拟和真实噪声测试集上分别实现了高达25.7%和10.6%的相对单词错误率降低,同时保持甚至提高了干净语音测试集的性能。源代码、模型检查点和数据模拟脚本将在https://github.com/litchimo/NoisyD-CT上提供。
摘要:To enhance the performance of end-to-end (E2E) speech recognition systems in noisy or low signal-to-noise ratio (SNR) conditions, this paper introduces NoisyD-CT, a novel tri-stage training framework built on the Conformer-Transducer architecture. The core of NoisyD-CT is a especially designed compact noisy disentanglement (NoisyD) module (adding only 1.71M parameters), integrated between the Conformer blocks and Transducer Decoder to perform deep noise suppression and improve ASR robustness in challenging acoustic noise environments. To fully exploit the noise suppression capability of the NoisyD-CT, we further propose a clean representation consistency loss to align high-level representations derived from noisy speech with those obtained from corresponding clean speech. Together with a noisy reconstruction loss, this consistency alignment enables the NoisyD module to effectively suppress noise while preserving essential acoustic and linguistic features consistent across both clean and noisy conditions, thereby producing cleaner internal representations that enhance ASR performance. Moreover, our tri-stage training strategy is designed to fully leverage the functionalities of both the noisy disentanglement and speech recognition modules throughout the model training process, ultimately maximizing performance gains under noisy conditions. Our experiments are performed on the LibriSpeech and CHiME-4 datasets, extensive results demonstrate that our proposed NoisyD-CT significantly outperforms the competitive Conformer-Transducer baseline, achieving up to 25.7% and 10.6% relative word error rate reductions on simulated and real-world noisy test sets, respectively, while maintaining or even improving performance on clean speech test sets. The source code, model checkpoint and data simulation scripts will be available at https://github.com/litchimo/NoisyD-CT.


【38】MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
标题:MPO:基于语言模型的文本到语音的多维偏好优化
链接:https://arxiv.org/abs/2509.00685

备注:Accepted by NCMMSC2025
摘要:近年来,文本到语音(TTS)通过大规模语言模型取得了令人印象深刻的进步,实现了人类水平的语音质量。整合人的反馈已被证明是有效的,以提高这些系统的鲁棒性。然而,目前的方法面临的挑战,优化TTS与偏好数据在多个维度上,往往遭受性能下降,由于过度自信的奖励。我们提出了多维偏好优化(MPO),以更好地调整TTS系统与人类的喜好。MPO引入了一个偏好集,简化了多维偏好优化的数据构建,实现了多维度对齐。此外,我们在训练过程中加入正则化,以解决基于DPO的方法中的典型退化问题。我们的实验证明MPO的有效性,在可懂度,说话人相似性和韵律相比,基线系统显着改善。
摘要:In recent years, text-to-speech (TTS) has seen impressive advancements through large-scale language models, achieving human-level speech quality. Integrating human feedback has proven effective for enhancing robustness in these systems. However, current approaches face challenges in optimizing TTS with preference data across multiple dimensions and often suffer from performance degradation due to overconfidence in rewards. We propose Multidimensional Preference Optimization (MPO) to better align TTS systems with human preferences. MPO introduces a preference set that streamlines the construction of data for multidimensional preference optimization, enabling alignment with multiple dimensions. Additionally, we incorporate regularization during training to address the typical degradation issues in DPO-based approaches. Our experiments demonstrate MPO's effectiveness, showing significant improvements in intelligibility, speaker similarity, and prosody compared to baseline systems.


【39】Deep Learning for Personalized Binaural Audio Reproduction
标题:个性化双耳音频复制的深度学习
链接:https://arxiv.org/abs/2509.00400

摘要:个性化双耳音频再现是逼真的空间定位、声音外化和沉浸式聆听的基础,直接塑造用户体验和聆听效果。本调查回顾了深度学习在这一任务中的最新进展,并将其按生成机制分为两种范式:显式个性化过滤和端到端渲染。显式方法从稀疏测量、形态学特征或环境线索预测个性化的头部相关传递函数(HRTF),然后在传统的渲染管道中使用它们。端到端方法将源信号直接映射到双耳信号,并辅以其他输入,如视觉、文本或参数指导,并在模型中学习个性化。我们还总结了该领域的主要数据集和评估指标,以支持公平和可重复的比较。最后,我们讨论了这些技术实现的关键应用,当前的技术限制以及基于深度学习的空间音频系统的潜在研究方向。
摘要:Personalized binaural audio reproduction is the basis of realistic spatial localization, sound externalization, and immersive listening, directly shaping user experience and listening effort. This survey reviews recent advances in deep learning for this task and organizes them by generation mechanism into two paradigms: explicit personalized filtering and end-to-end rendering. Explicit methods predict personalized head-related transfer functions (HRTFs) from sparse measurements, morphological features, or environmental cues, and then use them in the conventional rendering pipeline. End-to-end methods map source signals directly to binaural signals, aided by other inputs such as visual, textual, or parametric guidance, and they learn personalization within the model. We also summarize the field's main datasets and evaluation metrics to support fair and repeatable comparison. Finally, we conclude with a discussion of key applications enabled by these technologies, current technical limitations, and potential research directions for deep learning-based spatial audio systems.


【40】Quantum-Enhanced Analysis and Grading of Vocal Performance
标题:声乐表演的量子增强分析与评分
链接:https://arxiv.org/abs/2509.00106

备注:4 pages, 5 figures. Hybrid quantum - classical feasibility study; simulator - only results
摘要:我们提出了QuantumMelody,一个混合量子经典的客观歌唱评估方法。分组的声音特征(音高稳定性,动态,音色)被编码到一个小型的模拟量子电路中;所有九个量子比特都用每个量子比特上的Hadamard初始化,然后接收Rx,Ry和Rz旋转,并进行组内和跨组纠缠。将电路测量概率与频谱图Transformer嵌入融合,以估计标签2-5上的等级并使技术级反馈浮出水面。在168个标记的20秒摘录中,混合体与专家评分员的一致率达到74.29%,比经典特征基线增加了+12.86分。在笔记本电脑级的Qiskit模拟器上,每个记录的处理时间为亚分钟;我们不要求硬件加速。这是一个可行的一步,可解释的,客观的歌唱评估应用音频信号处理。
摘要:We present QuantumMelody, a hybrid quantum-classical method for objective singing assessment. Grouped vocal features (pitch stability, dynamics, timbre) are encoded into a small simulated quantum circuit; all nine qubits are initialized with a Hadamard on each qubit and then receive Rx, Ry, and Rz rotations, with intra- and cross-group entanglement. The circuit measurement probabilities are fused with spectrogram transformer embeddings to estimate a grade on labels 2-5 and to surface technique-level feedback. On 168 labeled 20 second excerpts, the hybrid reaches 74.29% agreement with expert graders, a +12.86 point gain over a classical-features baseline. Processing is sub-minute per recording on a laptop-class Qiskit simulator; we do not claim hardware speedups. This is a feasibility step toward interpretable, objective singing assessment in applied audio signal processing.


【41】Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
标题:使用深度学习自动检测和纠正《古兰经》学习者的发音错误
链接:https://arxiv.org/abs/2509.00094

摘要:评估口语是一项挑战,而量化机器学习模型的发音指标就更难了。然而,对于《古兰经》来说,穆斯林学者制定的严格的背诵规则(tajweed)简化了这项任务,使评估变得非常有效。尽管有这一优势,但缺乏高质量的注释数据仍然是一个重大障碍。   在这项工作中,我们通过引入:(1)一个98%自动化的管道来产生高质量的古兰经数据集-包括:从专家背诵者那里收集背诵,使用我们微调的wav2vec2-BERT模型在暂停点(waqf)进行分割,片段的转录,通过我们新颖的Tasmeea算法进行转录验证;(2)850+小时的音频(~300K注释的话语);(3)一种新的基于ASR的发音错误检测方法,利用我们的自定义古兰经音标(QPS)编码Tajweed规则(不同于现代标准阿拉伯语的IPA标准)。QPS使用两级脚本:(音素级):用短/长元音编码阿拉伯字母。(Sifa级别):编码每个音素的发音特征。我们还包括全面的建模与我们的新的多级CTC模型,达到0.16%的平均音素错误率(PER)的测试集。我们以开源方式发布所有代码、数据和模型:www.example.com
摘要:Assessing spoken language is challenging, and quantifying pronunciation metrics for machine learning models is even harder. However, for the Holy Quran, this task is simplified by the rigorous recitation rules (tajweed) established by Muslim scholars, enabling highly effective assessment. Despite this advantage, the scarcity of high-quality annotated data remains a significant barrier.   In this work, we bridge these gaps by introducing: (1) A 98% automated pipeline to produce high-quality Quranic datasets -- encompassing: Collection of recitations from expert reciters, Segmentation at pause points (waqf) using our fine-tuned wav2vec2-BERT model, Transcription of segments, Transcript verification via our novel Tasmeea algorithm; (2) 850+ hours of audio (~300K annotated utterances); (3) A novel ASR-based approach for pronunciation error detection, utilizing our custom Quran Phonetic Script (QPS) to encode Tajweed rules (unlike the IPA standard for Modern Standard Arabic). QPS uses a two-level script: (Phoneme level): Encodes Arabic letters with short/long vowels. (Sifa level): Encodes articulation characteristics of every phoneme. We further include comprehensive modeling with our novel multi-level CTC Model which achieved 0.16% average Phoneme Error Rate (PER) on the testset. We release all code, data, and models as open-source: https://obadx.github.io/prepare-quran-dataset/


【42】ChipChat: Low-Latency Cascaded Conversational Agent in MLX
标题:ChipChat:MLX中的低延迟级联对话代理
链接:https://arxiv.org/abs/2509.00078

备注:ASRU 2025
摘要:大型语言模型(LLM)的出现已经改变了口语对话系统,但实时设备上语音代理的最佳架构仍然是一个悬而未决的问题。虽然端到端方法在理论上有优势,但级联系统(CS)在语言理解任务中继续优于它们,尽管受到顺序处理延迟的限制。在这项工作中,我们介绍了ChipChat,一种新型的低延迟CS,通过架构创新和流优化克服了传统的瓶颈。我们的系统集成了流(a)对话语音识别与混合专家,(b)状态动作增强LLM,(c)文本到语音合成,(d)神经声码器,(e)扬声器建模。ChipChat使用MLX实现,在没有专用GPU的Mac Studio上实现亚秒级响应延迟,同时通过完整的设备上处理保护用户隐私。我们的工作表明,从战略上重新设计的CS可以克服其历史延迟限制,为实际的基于语音的AI代理提供了一条有前途的道路。
摘要:The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages, cascaded systems (CSs) continue to outperform them in language understanding tasks, despite being constrained by sequential processing latency. In this work, we introduce ChipChat, a novel low-latency CS that overcomes traditional bottlenecks through architectural innovations and streaming optimizations. Our system integrates streaming (a) conversational speech recognition with mixture-of-experts, (b) state-action augmented LLM, (c) text-to-speech synthesis, (d) neural vocoder, and (e) speaker modeling. Implemented using MLX, ChipChat achieves sub-second response latency on a Mac Studio without dedicated GPUs, while preserving user privacy through complete on-device processing. Our work shows that strategically redesigned CSs can overcome their historical latency limitations, offering a promising path forward for practical voice-based AI agents.


【43】Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition
标题:放大情感信号:数据高效的深度学习,实现稳健的语音情感识别
链接:https://arxiv.org/abs/2509.00077

摘要:语音情感识别(SER)是人机交互中一个重要而持久的挑战。虽然深度学习已经推进了口语处理,但在有限的数据集上实现高性能仍然是一个关键障碍。本文通过开发和评估一套机器学习模型来解决这个问题,包括支持向量机(SVM),长短期记忆网络(LSTM)和卷积神经网络(CNN),用于人类语音中的自动情感分类。我们证明,通过战略性地采用迁移学习和创新的数据增强技术,我们的模型可以实现令人印象深刻的性能,尽管数据集相对较小。我们最有效的模型是ResNet34架构,它在RAVDESS和SAVEE数据集上建立了一个新的性能基准,准确率为66.7%,F1得分为0.631。这些结果强调了利用预先训练的模型和数据增强来克服数据稀缺性的巨大好处,从而为更强大和可推广的SER系统铺平了道路。
摘要:Speech Emotion Recognition (SER) presents a significant yet persistent challenge in human-computer interaction. While deep learning has advanced spoken language processing, achieving high performance on limited datasets remains a critical hurdle. This paper confronts this issue by developing and evaluating a suite of machine learning models, including Support Vector Machines (SVMs), Long Short-Term Memory networks (LSTMs), and Convolutional Neural Networks (CNNs), for automated emotion classification in human speech. We demonstrate that by strategically employing transfer learning and innovative data augmentation techniques, our models can achieve impressive performance despite the constraints of a relatively small dataset. Our most effective model, a ResNet34 architecture, establishes a new performance benchmark on the combined RAVDESS and SAVEE datasets, attaining an accuracy of 66.7% and an F1 score of 0.631. These results underscore the substantial benefits of leveraging pre-trained models and data augmentation to overcome data scarcity, thereby paving the way for more robust and generalizable SER systems.


eess.AS音频处理


【1】Group Relative Policy Optimization for Speech Recognition
标题:语音识别的群体相对策略优化
链接:https://arxiv.org/abs/2509.01939

备注:Accepted for ASRU 2025
摘要:语音识别已经看到了采用大型语言模型(LLM)的巨大转变。这种转变部分是由LLM所表现出的良好的可扩展性特性,利用大量标记,未标记的语音和文本数据的能力,具有自回归框架的流式传输能力以及具有LLM指令遵循特性的多任务处理所驱动的。然而,通常与LLM一起使用的简单的下一个令牌预测目标在性能和幻觉方面具有一定的限制。在本文中,我们提出了应用组相对策略优化(GRPO),使强化学习从人类的反馈自动语音识别(ASR)。我们设计了简单的基于规则的奖励函数来指导策略更新。我们证明了显着的改善字错误率(高达18.4%相对),减少幻觉,增加鲁棒性域外的数据集和域适应的有效性。
摘要:Speech Recognition has seen a dramatic shift towards adopting Large Language Models (LLMs). This shift is partly driven by good scalability properties demonstrated by LLMs, ability to leverage large amounts of labelled, unlabelled speech and text data, streaming capabilities with auto-regressive framework and multi-tasking with instruction following characteristics of LLMs. However, simple next-token prediction objective, typically employed with LLMs, have certain limitations in performance and challenges with hallucinations. In this paper, we propose application of Group Relative Policy Optimization (GRPO) to enable reinforcement learning from human feedback for automatic speech recognition (ASR). We design simple rule based reward functions to guide the policy updates. We demonstrate significant improvements in word error rate (upto 18.4% relative), reduction in hallucinations, increased robustness on out-of-domain datasets and effectiveness in domain adaptation.


【2】Binaural Unmasking in Practical Use: Perceived Level of Phase-inverted Speech in Environmental Noise
标题:实际应用中的双耳去掩蔽:环境噪音中倒相语音的感知水平
链接:https://arxiv.org/abs/2509.01929

摘要:我们的目标是开发一种技术,使耳机和耳机的声音更容易听到,而不会增加声压或消除环境噪音。为此,我们专注于利用双耳解蔽现象,通过相位反转在一只耳朵。具体来说,我们进行实验,以评估由该现象引起的可听度的改善,使用近似实际情况的条件。我们使用包括女性在内的各种说话者的语音和日常生活中可能遇到的噪音(城市环境声音,欢呼声)来验证在接近实际情况的条件下双耳解蔽的效果。使用日语的实验结果表明,(i)在嘈杂环境中的语音被感知为在一只耳朵中具有相位反转的高达约6dB的声音,以及(ii)对于本研究中针对的所有扬声器和噪声,都获得了一定的效果(可听度提高5dB或更多)。这些研究结果表明,双耳解蔽归因于在实际情况下的耳间相位差的有效性。
摘要:We aim to develop a technology that makes the sound from earphones and headphones easier to hear without increasing the sound pressure or eliminating ambient noise. To this end, we focus on harnessing the phenomenon of binaural unmasking through phase reversal in one ear. Specifically, we conduct experiments to evaluate the improvement of audibility caused by the phenomenon, using conditions that approximate practical scenarios. We use speech sounds by various speakers, including women, and noises that can be encountered in daily life (urban environmental sounds, cheers) to verify the effects of binaural unmasking under conditions close to practical situations. The results of experiments using the Japanese language showed that (i) speech in a noisy environment is perceived to be up to about 6 dB louder with phase reversal in one ear, and (ii) a certain effect (improvement of audibility by 5 dB or more) is obtained for all speakers and noises targeted in this study. These findings demonstrate the effectiveness of binaural unmasking attributed to interaural phase differences in practical scenarios.


【3】Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy
标题:使用离散标记和两步训练策略的多语言语音识别
链接:https://arxiv.org/abs/2509.01900

备注:Accepted by NCMMSC 2024
摘要:预训练模型,特别是自监督学习(SSL)模型,在自动语音识别(ASR)任务中表现出令人印象深刻的结果。虽然SSL模型的大多数应用程序都专注于利用连续表示作为训练下游任务的功能,但近年来,由于其较低的存储要求和更广泛的应用范围,离散单元的使用越来越受到关注。在多语言ASR任务中,模型不同层的表示对不同语言的贡献不同,使离散单元建模的统一变得复杂。在本文中,我们提出了一种两阶段训练策略,以提高预训练模型的离散令牌性能,并缩小与连续表示性能的差距。我们在XLS-R模型上验证了我们的方法,该模型遵循Interspeech 2024使用离散语音单元挑战的语音处理的设置。我们的方法证明了ML-SUPERB数据集的显著改进,XLS-R模型的CER相对减少了44%。这超过了WavLM模型之前设定的基线,该模型实现了CER相对减少26%。此外,我们的方法取得了第一名之间的所有单系统的结果在排行榜上。
摘要:Pre-trained models, especially self-supervised learning (SSL) models, have demonstrated impressive results in automatic speech recognition (ASR) task. While most applications of SSL models focus on leveraging continuous representations as features for training downstream tasks, the utilization of discrete units has gained increasing attention in recent years owing to its lower storage requirements and broader range of applications. In multilingual ASR tasks, representations at different layers of the model contribute differently to various languages, complicating the unification of discrete unit modeling. In this paper, we propose a two-stage training strategy to improve the discrete token performance of pre-trained models and narrow the gap with continuous representation performance. We validate our method on the XLS-R model following the settings of Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge. Our method demonstrates a significant improvement on the ML-SUPERB dataset, achieving a 44% relative reduction on CER for the XLS-R model. This surpasses the previous baseline set by the WavLM model, which achieves a 26% relative reduction on CER. Furthermore, our method achieves the first place among all the single-system results on the leaderboard.


【4】From Evaluation to Optimization: Neural Speech Assessment for Downstream Applications
标题:从评估到优化:下游应用的神经语音评估
链接:https://arxiv.org/abs/2509.01889

备注:5 pages, 1 figure
摘要:合成和处理语音的评估长期以来一直是音频工程和语音科学的基石。虽然主观听力测试仍然是评估感知质量和可懂度的黄金标准,但其高成本、时间要求和有限的可扩展性在现代语音技术的快速发展周期中提出了重大挑战。传统的客观指标虽然计算效率高,但通常与人类感知的相关性较弱,从而在系统优化和实际用户体验之间产生感知差距。弥合这一差距需要更接近人类感知的语音评估模型。近年来,许多基于神经网络的语音评估模型已经被开发出来,以预测质量和可懂度,取得了可喜的成果。除了在评估中的作用外,这些模型越来越多地集成到下游语音处理任务中。这篇综述集中在两个主要领域:(1)作为可区分的感知代理,不仅评估,而且指导语音增强和合成模型的优化;(2)使显着的语音特征的检测,以支持更精确和更有效的下游处理。最后,我们讨论了当前的局限性,并概述了未来的研究方向,以进一步推进语音评估到语音处理管道的集成。
摘要:The evaluation of synthetic and processed speech has long been a cornerstone of audio engineering and speech science. Although subjective listening tests remain the gold standard for assessing perceptual quality and intelligibility, their high cost, time requirements, and limited scalability present significant challenges in the rapid development cycles of modern speech technologies. Traditional objective metrics, while computationally efficient, often exhibit weak correlation with human perception, creating a perceptual gap between system optimization and actual user experience. Bridging this gap requires speech assessment models that are more closely aligned with human perception. In recent years, numerous neural network-based speech assessment models have been developed to predict quality and intelligibility, achieving promising results. Beyond their role in evaluation, these models are increasingly integrated into downstream speech processing tasks. This review focuses on their role in two main areas: (1) serving as differentiable perceptual proxies that not only assess but also guide the optimization of speech enhancement and synthesis models; and (2) enabling the detection of salient speech characteristics to support more precise and efficient downstream processing. Finally, we discuss current limitations and outline future research directions to further advance the integration of speech assessment into speech processing pipelines.


【5】AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
标题:AHAMaskk:无指令的大型音频语言模型的可靠任务规范
链接:https://arxiv.org/abs/2509.01787

备注:15 pages, 7 tables, 6 figures
摘要:虽然目前的大型音频语言模型(LALM)扩展了文本大型语言模型(LLM)与通用的声学理解能力,他们通常遭受指令敏感性,其中相同意图的不同指令可以产生截然不同的结果。在这项工作中,我们提出了AHAMask,在那里我们只是简单地屏蔽了一些注意头的解码器,只有LLM骨干的LALM,触发特定的声学任务功能,没有指令。这些掩码通过在LALM上训练而有效地获得,其中可训练参数的数量等于其LLM骨干中的注意力头数。我们通过实验表明,应用这种选择性注意头罩实现可比的,甚至更好的性能比使用指令,无论是在单一或复合任务。除了实现可靠的声学任务规范的LALM,这也表明,LALM表现出一定的“功能通路”在他们的注意头。
摘要:Although current large audio language models (LALMs) extend text large language models (LLMs) with generic acoustic understanding abilities, they usually suffer from instruction sensitivity, where different instructions of the same intention can yield drastically different outcomes. In this work, we propose AHAMask, where we simply mask some of the attention heads in the decoder-only LLM backbone of LALMs, to trigger specific acoustic task functionalities without instructions. These masks are efficiently obtained by training on an LALM, with the number of trainable parameters equal to the attention head count in its LLM backbone. We show by experiments that applying such selective attention head masks achieves comparable or even better performance than using instructions, either on single or composite tasks. Besides achieving reliable acoustic task specification for LALMs, this also reveals that LALMs exhibit certain "functional pathways" in their attention heads.


【6】Characterization of Speech Similarity Between Australian Aboriginal and High-Resource Languages: A Case Study on Dharawal
标题:澳大利亚原住民语言和高资源语言之间言语相似性的描述:达拉瓦尔的案例研究
链接:https://arxiv.org/abs/2509.01419

备注:Accepted at APSIPA ASC 2025
摘要:澳大利亚原住民语言具有重要的文化和语言价值,但在现代语音人工智能系统中仍然严重不足。虽然最先进的语音基础模型和自动语音识别在高资源环境中表现出色,但它们通常难以推广到低资源语言,特别是那些缺乏干净的注释语音数据的语言。在这项工作中,我们收集和清洁的语音数据集Dharawal,低资源的澳大利亚土著语言,仔细采购和处理公开的录音。使用这个数据集,我们使用预先训练的多语言语音编码器分析了Dharawal和107种高资源语言之间的语音相似性。我们的方法结合了(1)错误分类率分析,以评估语言的易混淆性,(2)细粒度的相似性测量,使用余弦相似性和Fr\'echet初始距离(FID)在嵌入空间。实验结果表明,Dharawal语与拉丁语、毛利语、韩语、泰语和威尔士语等语言具有很强的语音相似性。这些发现为未来的迁移学习和模型适应工作提供了实用指导,并强调了数据收集和基于嵌入的分析在支持濒危语言社区语音技术方面的重要性。
摘要:Australian Aboriginal languages are of significant cultural and linguistic value but remain severely underrepresented in modern speech AI systems. While state-of-the-art speech foundation models and automatic speech recognition excel in high-resource settings, they often struggle to generalize to low-resource languages, especially those lacking clean, annotated speech data. In this work, we collect and clean a speech dataset for Dharawal, a low-resource Australian Aboriginal language, by carefully sourcing and processing publicly available recordings. Using this dataset, we analyze the speech similarity between Dharawal and 107 high-resource languages using a pre-trained multilingual speech encoder. Our approach combines (1) misclassification rate analysis to assess language confusability, and (2) fine-grained similarity measurements using cosine similarity and Fr\'echet Inception Distance (FID) in the embedding space. Experimental results reveal that Dharawal shares strong speech similarity with languages such as Latin, M\=aori, Korean, Thai, and Welsh. These findings offer practical guidance for future transfer learning and model adaptation efforts, and underscore the importance of data collection and embedding-based analysis in supporting speech technologies for endangered language communities.


【7】MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
标题:MixedG 2 P-T5:使用语音自我监督学习和语言模型的混合脚本文本的无G2 P语音合成
链接:https://arxiv.org/abs/2509.01391

备注:In Proceedings of the 17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC 2025)
摘要:这项研究提出了一种新颖的语音合成方法,可以通过使用直接从语音生成离散令牌的基于深度学习的模型来替代传统的字形到音素(G2 P)转换。利用预先训练的语音SSL模型,我们训练T5编码器以从混合脚本文本(例如,包含Kanji和Kana)。该方法消除了对手动语音转录的需要,降低了成本并增强了可扩展性,特别是对于大型非转录音频数据集。我们的模型与传统的基于G2P的文本到语音系统的性能相匹配,并且能够合成保留自然语言学和非语言学特征(如口音和语调)的语音。
摘要:This study presents a novel approach to voice synthesis that can substitute the traditional grapheme-to-phoneme (G2P) conversion by using a deep learning-based model that generates discrete tokens directly from speech. Utilizing a pre-trained voice SSL model, we train a T5 encoder to produce pseudo-language labels from mixed-script texts (e.g., containing Kanji and Kana). This method eliminates the need for manual phonetic transcription, reducing costs and enhancing scalability, especially for large non-transcribed audio datasets. Our model matches the performance of conventional G2P-based text-to-speech systems and is capable of synthesizing speech that retains natural linguistic and paralinguistic features, such as accents and intonations.


【8】High-Density MIMO Localization Using a 32x64 Ultrasonic Transducer-Microphone Array with Real-Time Data Streaming
标题:使用具有实时数据流的32 x64超声传感器-麦克风阵列进行高密度MMO定位
链接:https://arxiv.org/abs/2509.01210

备注:Accepted for publication at IEEE IUS 2025
摘要:在这项工作中,我们提出了一种新的超声波阵列系统设计的高精度定位使用大规模MIMO(多输入多输出)架构。该系统将32个发射器与62个麦克风相结合,创建了一个扩展的虚拟孔径,提高了通道分离性和空间分辨率。每个发射器由超声波频带内的随机相位多正弦激励,这降低了信道间的相关性并增加了对多径的鲁棒性。通过反射面成像仿真和实际换能器带宽约束下的通道分离分析,证明了该方法的可行性。结果表明,MIMO处理可以改善分离的反射器相比,单发射器的配置,虽然实际的限制,如换能器带宽减少可实现的信道隔离。
摘要:In this work, we present a novel ultrasonic array system designed for high-precision localization using a large-scale MIMO (Multiple-Input Multiple-Output) architecture. The system combines 32 transmitters with 62 microphones, creating an extended virtual aperture that improves channel separability and spatial resolution. Each transmitter is excited by a random-phase multisine within the ultrasonic band, which reduces inter-channel correlation and increases robustness against multipath. The feasibility of the approach is demonstrated through simulations of reflector imaging and analysis of channel separation under realistic transducer bandwidth constraints. Results show that MIMO processing enables improved separation of reflectors compared to single-emitter configurations, although practical limitations such as transducer bandwidth reduce the achievable channel isolation.


【9】Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
标题:基于三阶段训练的噪音去纠缠实现抗噪语音识别
链接:https://arxiv.org/abs/2509.01087

备注:11 pages,4 figures
摘要:为了提高端到端(E2 E)语音识别系统在噪声或低信噪比(SNR)条件下的性能,本文介绍了NoisyD-CT,一种新的三阶段训练框架,建立在Conformer-Transducer架构上。NoisyD-CT的核心是一个专门设计的紧凑型噪声解纠缠(NoisyD)模块(仅添加1.71 M参数),集成在Conformer模块和换能器解码器之间,可在具有挑战性的声学噪声环境中执行深度噪声抑制并提高ASR鲁棒性。为了充分利用NoisyD-CT的噪声抑制能力,我们进一步提出了一种干净的表示一致性损失,以将从有噪语音中获得的高级表示与从相应的干净语音中获得的表示对齐。与噪声重建损失一起,这种一致性对齐使NoisyD模块能够有效地抑制噪声,同时在干净和嘈杂的条件下保持基本的声学和语言特征一致,从而产生更清晰的内部表示,增强ASR性能。此外,我们的三阶段训练策略旨在在整个模型训练过程中充分利用噪声解纠缠和语音识别模块的功能,最终在噪声条件下最大限度地提高性能。我们的实验是在LibriSpeech和CHiME-4数据集上进行的,大量的结果表明,我们提出的NoisyD-CT显著优于竞争对手的Conformer-Transducer基线,在模拟和真实噪声测试集上分别实现了高达25.7%和10.6%的相对单词错误率降低,同时保持甚至提高了干净语音测试集的性能。源代码、模型检查点和数据模拟脚本将在https://github.com/litchimo/NoisyD-CT上提供。
摘要:To enhance the performance of end-to-end (E2E) speech recognition systems in noisy or low signal-to-noise ratio (SNR) conditions, this paper introduces NoisyD-CT, a novel tri-stage training framework built on the Conformer-Transducer architecture. The core of NoisyD-CT is a especially designed compact noisy disentanglement (NoisyD) module (adding only 1.71M parameters), integrated between the Conformer blocks and Transducer Decoder to perform deep noise suppression and improve ASR robustness in challenging acoustic noise environments. To fully exploit the noise suppression capability of the NoisyD-CT, we further propose a clean representation consistency loss to align high-level representations derived from noisy speech with those obtained from corresponding clean speech. Together with a noisy reconstruction loss, this consistency alignment enables the NoisyD module to effectively suppress noise while preserving essential acoustic and linguistic features consistent across both clean and noisy conditions, thereby producing cleaner internal representations that enhance ASR performance. Moreover, our tri-stage training strategy is designed to fully leverage the functionalities of both the noisy disentanglement and speech recognition modules throughout the model training process, ultimately maximizing performance gains under noisy conditions. Our experiments are performed on the LibriSpeech and CHiME-4 datasets, extensive results demonstrate that our proposed NoisyD-CT significantly outperforms the competitive Conformer-Transducer baseline, achieving up to 25.7% and 10.6% relative word error rate reductions on simulated and real-world noisy test sets, respectively, while maintaining or even improving performance on clean speech test sets. The source code, model checkpoint and data simulation scripts will be available at https://github.com/litchimo/NoisyD-CT.


【10】MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
标题:MPO:基于语言模型的文本到语音的多维偏好优化
链接:https://arxiv.org/abs/2509.00685

备注:Accepted by NCMMSC2025
摘要:近年来,文本到语音(TTS)通过大规模语言模型取得了令人印象深刻的进步,实现了人类水平的语音质量。整合人的反馈已被证明是有效的,以提高这些系统的鲁棒性。然而,目前的方法面临的挑战,优化TTS与偏好数据在多个维度上,往往遭受性能下降,由于过度自信的奖励。我们提出了多维偏好优化(MPO),以更好地调整TTS系统与人类的喜好。MPO引入了一个偏好集,简化了多维偏好优化的数据构建,实现了多维度对齐。此外,我们在训练过程中加入正则化,以解决基于DPO的方法中的典型退化问题。我们的实验证明MPO的有效性,在可懂度,说话人相似性和韵律相比,基线系统显着改善。
摘要:In recent years, text-to-speech (TTS) has seen impressive advancements through large-scale language models, achieving human-level speech quality. Integrating human feedback has proven effective for enhancing robustness in these systems. However, current approaches face challenges in optimizing TTS with preference data across multiple dimensions and often suffer from performance degradation due to overconfidence in rewards. We propose Multidimensional Preference Optimization (MPO) to better align TTS systems with human preferences. MPO introduces a preference set that streamlines the construction of data for multidimensional preference optimization, enabling alignment with multiple dimensions. Additionally, we incorporate regularization during training to address the typical degradation issues in DPO-based approaches. Our experiments demonstrate MPO's effectiveness, showing significant improvements in intelligibility, speaker similarity, and prosody compared to baseline systems.


【11】Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
标题:使用音素级预训练语言模型的文本到语音的说话者条件短语中断预测
链接:https://arxiv.org/abs/2509.00675

备注:Under Review
摘要:本文提出了在多说话人文本到语音(TTS)系统中的短语中断预测(也称为短语)。我们通过利用扬声器嵌入来集成特定于扬声器的功能,以提高措辞模型的性能。我们进一步证明,这些扬声器嵌入可以捕捉扬声器相关的特性,仅从措辞任务。此外,我们通过一种Few-Shot自适应方法,探索了预先训练的说话人嵌入的潜力。此外,我们率先将音素级预训练语言模型应用于该TTS前端任务,这显著提高了短语模型的准确性。我们的方法通过客观和主观评估进行严格评估,证明其有效性。
摘要:This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker embeddings to enhance the performance of the phrasing model. We further demonstrate that these speaker embeddings can capture speaker-related characteristics solely from the phrasing task. Besides, we explore the potential of pre-trained speaker embeddings for unseen speakers through a few-shot adaptation method. Furthermore, we pioneer the application of phoneme-level pre-trained language models to this TTS front-end task, which significantly boosts the accuracy of the phrasing model. Our methods are rigorously assessed through both objective and subjective evaluations, demonstrating their effectiveness.


【12】Deep Learning for Personalized Binaural Audio Reproduction
标题:个性化双耳音频复制的深度学习
链接:https://arxiv.org/abs/2509.00400

摘要:个性化双耳音频再现是逼真的空间定位、声音外化和沉浸式聆听的基础,直接塑造用户体验和聆听效果。本调查回顾了深度学习在这一任务中的最新进展,并将其按生成机制分为两种范式:显式个性化过滤和端到端渲染。显式方法从稀疏测量、形态学特征或环境线索预测个性化的头部相关传递函数(HRTF),然后在传统的渲染管道中使用它们。端到端方法将源信号直接映射到双耳信号,并辅以其他输入,如视觉、文本或参数指导,并在模型中学习个性化。我们还总结了该领域的主要数据集和评估指标,以支持公平和可重复的比较。最后,我们讨论了这些技术实现的关键应用,当前的技术限制以及基于深度学习的空间音频系统的潜在研究方向。
摘要:Personalized binaural audio reproduction is the basis of realistic spatial localization, sound externalization, and immersive listening, directly shaping user experience and listening effort. This survey reviews recent advances in deep learning for this task and organizes them by generation mechanism into two paradigms: explicit personalized filtering and end-to-end rendering. Explicit methods predict personalized head-related transfer functions (HRTFs) from sparse measurements, morphological features, or environmental cues, and then use them in the conventional rendering pipeline. End-to-end methods map source signals directly to binaural signals, aided by other inputs such as visual, textual, or parametric guidance, and they learn personalization within the model. We also summarize the field's main datasets and evaluation metrics to support fair and repeatable comparison. Finally, we conclude with a discussion of key applications enabled by these technologies, current technical limitations, and potential research directions for deep learning-based spatial audio systems.


【13】Quantum-Enhanced Analysis and Grading of Vocal Performance
标题:声乐表演的量子增强分析与评分
链接:https://arxiv.org/abs/2509.00106

备注:4 pages, 5 figures. Hybrid quantum - classical feasibility study; simulator - only results
摘要:我们提出了QuantumMelody,一个混合量子经典的客观歌唱评估方法。分组的声音特征(音调稳定性、动态、音色)被编码到一个小型模拟量子电路中;所有九个量子位都用每个量子位上的Hadamard初始化,然后接收Rx、Ry和Rz旋转,并具有组内和跨组纠缠。将电路测量概率与频谱图Transformer嵌入融合,以估计标签2-5上的等级并使技术级反馈浮出水面。在168个标记的20秒摘录中,混合体与专家评分员的一致率达到74.29%,比经典特征基线增加了+12.86分。在笔记本电脑级的Qiskit模拟器上,每个记录的处理时间为亚分钟;我们不要求硬件加速。这是一个可行的一步,可解释的,客观的歌唱评估应用音频信号处理。
摘要:We present QuantumMelody, a hybrid quantum-classical method for objective singing assessment. Grouped vocal features (pitch stability, dynamics, timbre) are encoded into a small simulated quantum circuit; all nine qubits are initialized with a Hadamard on each qubit and then receive Rx, Ry, and Rz rotations, with intra- and cross-group entanglement. The circuit measurement probabilities are fused with spectrogram transformer embeddings to estimate a grade on labels 2-5 and to surface technique-level feedback. On 168 labeled 20 second excerpts, the hybrid reaches 74.29% agreement with expert graders, a +12.86 point gain over a classical-features baseline. Processing is sub-minute per recording on a laptop-class Qiskit simulator; we do not claim hardware speedups. This is a feasibility step toward interpretable, objective singing assessment in applied audio signal processing.


【14】Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
标题:使用深度学习自动检测和纠正《古兰经》学习者的发音错误
链接:https://arxiv.org/abs/2509.00094

摘要:评估口语是一项挑战,而量化机器学习模型的发音指标就更难了。然而,对于《古兰经》来说,穆斯林学者制定的严格的背诵规则(tajweed)简化了这项任务,使评估变得非常有效。尽管有这一优势,但缺乏高质量的注释数据仍然是一个重大障碍。   在这项工作中,我们通过引入:(1)一个98%自动化的管道来产生高质量的古兰经数据集-包括:从专家背诵者那里收集背诵,使用我们微调的wav2vec2-BERT模型在暂停点(waqf)进行分割,片段的转录,通过我们新颖的Tasmeea算法进行转录验证;(2)850+小时的音频(~300K注释的话语);(3)一种新的基于ASR的发音错误检测方法,利用我们的自定义古兰经音标(QPS)编码Tajweed规则(不同于现代标准阿拉伯语的IPA标准)。QPS使用两级脚本:(音素级):用短/长元音编码阿拉伯字母。(Sifa级别):编码每个音素的发音特征。我们还包括全面的建模与我们的新的多级CTC模型,达到0.16%的平均音素错误率(PER)的测试集。我们以开源方式发布所有代码、数据和模型:www.example.com
摘要:Assessing spoken language is challenging, and quantifying pronunciation metrics for machine learning models is even harder. However, for the Holy Quran, this task is simplified by the rigorous recitation rules (tajweed) established by Muslim scholars, enabling highly effective assessment. Despite this advantage, the scarcity of high-quality annotated data remains a significant barrier.   In this work, we bridge these gaps by introducing: (1) A 98% automated pipeline to produce high-quality Quranic datasets -- encompassing: Collection of recitations from expert reciters, Segmentation at pause points (waqf) using our fine-tuned wav2vec2-BERT model, Transcription of segments, Transcript verification via our novel Tasmeea algorithm; (2) 850+ hours of audio (~300K annotated utterances); (3) A novel ASR-based approach for pronunciation error detection, utilizing our custom Quran Phonetic Script (QPS) to encode Tajweed rules (unlike the IPA standard for Modern Standard Arabic). QPS uses a two-level script: (Phoneme level): Encodes Arabic letters with short/long vowels. (Sifa level): Encodes articulation characteristics of every phoneme. We further include comprehensive modeling with our novel multi-level CTC Model which achieved 0.16% average Phoneme Error Rate (PER) on the testset. We release all code, data, and models as open-source: https://obadx.github.io/prepare-quran-dataset/


【15】ChipChat: Low-Latency Cascaded Conversational Agent in MLX
标题:ChipChat:MLX中的低延迟级联对话代理
链接:https://arxiv.org/abs/2509.00078

备注:ASRU 2025
摘要:大型语言模型(LLM)的出现已经改变了口语对话系统,但实时设备上语音代理的最佳架构仍然是一个悬而未决的问题。虽然端到端方法在理论上有优势,但级联系统(CS)在语言理解任务中继续优于它们,尽管受到顺序处理延迟的限制。在这项工作中,我们介绍了ChipChat,一种新型的低延迟CS,通过架构创新和流优化克服了传统的瓶颈。我们的系统集成了流(a)对话语音识别与混合专家,(b)状态动作增强LLM,(c)文本到语音合成,(d)神经声码器,(e)扬声器建模。ChipChat使用MLX实现,在没有专用GPU的Mac Studio上实现亚秒级响应延迟,同时通过完整的设备上处理保护用户隐私。我们的工作表明,从战略上重新设计的CS可以克服其历史延迟限制,为实际的基于语音的AI代理提供了一条有前途的道路。
摘要:The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages, cascaded systems (CSs) continue to outperform them in language understanding tasks, despite being constrained by sequential processing latency. In this work, we introduce ChipChat, a novel low-latency CS that overcomes traditional bottlenecks through architectural innovations and streaming optimizations. Our system integrates streaming (a) conversational speech recognition with mixture-of-experts, (b) state-action augmented LLM, (c) text-to-speech synthesis, (d) neural vocoder, and (e) speaker modeling. Implemented using MLX, ChipChat achieves sub-second response latency on a Mac Studio without dedicated GPUs, while preserving user privacy through complete on-device processing. Our work shows that strategically redesigned CSs can overcome their historical latency limitations, offering a promising path forward for practical voice-based AI agents.


【16】Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition
标题:放大情感信号:数据高效的深度学习,实现稳健的语音情感识别
链接:https://arxiv.org/abs/2509.00077

摘要:语音情感识别(SER)是人机交互中一个重要而持久的挑战。虽然深度学习已经推进了口语处理,但在有限的数据集上实现高性能仍然是一个关键障碍。本文通过开发和评估一套机器学习模型来解决这个问题,包括支持向量机(SVM),长短期记忆网络(LSTM)和卷积神经网络(CNN),用于人类语音中的自动情感分类。我们证明,通过战略性地采用迁移学习和创新的数据增强技术,我们的模型可以实现令人印象深刻的性能,尽管数据集相对较小。我们最有效的模型是ResNet34架构,它在RAVDESS和SAVEE数据集上建立了一个新的性能基准,准确率为66.7%,F1得分为0.631。这些结果强调了利用预先训练的模型和数据增强来克服数据稀缺性的巨大好处,从而为更强大和可推广的SER系统铺平了道路。
摘要:Speech Emotion Recognition (SER) presents a significant yet persistent challenge in human-computer interaction. While deep learning has advanced spoken language processing, achieving high performance on limited datasets remains a critical hurdle. This paper confronts this issue by developing and evaluating a suite of machine learning models, including Support Vector Machines (SVMs), Long Short-Term Memory networks (LSTMs), and Convolutional Neural Networks (CNNs), for automated emotion classification in human speech. We demonstrate that by strategically employing transfer learning and innovative data augmentation techniques, our models can achieve impressive performance despite the constraints of a relatively small dataset. Our most effective model, a ResNet34 architecture, establishes a new performance benchmark on the combined RAVDESS and SAVEE datasets, attaining an accuracy of 66.7% and an F1 score of 0.631. These results underscore the substantial benefits of leveraging pre-trained models and data augmentation to overcome data scarcity, thereby paving the way for more robust and generalizable SER systems.


【17】DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
标题:Deepspel Net:构建用于人类言语中自动情感识别的机器学习模型
链接:https://arxiv.org/abs/2509.00025

摘要:语音情感识别(SER)一直是口语处理研究中的一个具有挑战性的问题,因为人们还不清楚人类情感如何与声音的各种成分(如音高,响度和能量)联系起来。本文旨在使用机器学习来解决这个问题。特别是,我们使用SVM、LTSM和CNN构建了几个机器学习模型来对人类语音中的情感进行分类。此外,通过利用迁移学习和数据增强,我们有效地训练了我们的模型,使其在相对较小的数据集上获得了不错的性能。我们最好的模型是ResNet34网络,其准确率为66.7美元,F1得分为0.631美元。
摘要:Speech emotion recognition (SER) has been a challenging problem in spoken language processing research, because it is unclear how human emotions are connected to various components of sounds such as pitch, loudness, and energy. This paper aims to tackle this problem using machine learning. Particularly, we built several machine learning models using SVMs, LTSMs, and CNNs to classify emotions in human speeches. In addition, by leveraging transfer learning and data augmentation, we efficiently trained our models to attain decent performances on a relatively small dataset. Our best model was a ResNet34 network, which achieved an accuracy of $66.7\%$ and an F1 score of $0.631$.


【18】TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
标题:TTA-Bench:一个用于评估文本到音频模型的综合基准
链接:https://arxiv.org/abs/2509.02398

摘要:文本到音频(TTA)的生成已经取得了迅速的进展,但目前的评估方法仍然狭窄,主要集中在感知质量,而忽略了鲁棒性,泛化和道德问题。我们提出了TTA-Bench,这是一个全面的基准,用于评估TTA模型的功能性能,可靠性和社会责任。它涵盖了准确性、鲁棒性、公平性和毒性等七个维度,包括通过自动化和手动方法生成的2,999个不同提示。我们引入了一个统一的评估协议,该协议将客观指标与来自专家和普通用户的118,000多个人工注释相结合。在此框架下,对10种最先进的模型进行了基准测试,详细了解了它们的优势和局限性。TTA-Bench为TTA系统的全面和负责任的评估建立了新的标准。数据集和评估工具在https://nku-hlt.github.io/tta-bench/上开源。
摘要:Text-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generalization, and ethical concerns. We present TTA-Bench, a comprehensive benchmark for evaluating TTA models across functional performance, reliability, and social responsibility. It covers seven dimensions including accuracy, robustness, fairness, and toxicity, and includes 2,999 diverse prompts generated through automated and manual methods. We introduce a unified evaluation protocol that combines objective metrics with over 118,000 human annotations from both experts and general users. Ten state-of-the-art models are benchmarked under this framework, offering detailed insights into their strengths and limitations. TTA-Bench establishes a new standard for holistic and responsible evaluation of TTA systems. The dataset and evaluation tools are open-sourced at https://nku-hlt.github.io/tta-bench/.


【19】Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
标题:频谱图补丁编解码器:用于神经语音编码的2D块量化VQ-VAE和HiFi-GAN
链接:https://arxiv.org/abs/2509.02244

摘要:我们提出了一种神经语音编解码器,通过引入一种更简单的单级量化方法来挑战对复杂残差矢量量化(RVQ)堆栈的需求。我们的方法直接对梅尔频谱图进行操作,将其视为2D数据,并将非重叠的4x 4补丁量化为单个共享码本。这种拼接式设计简化了架构,实现了低延迟流,并产生了离散的潜在网格。为了确保高保真度的合成,我们采用了后期对抗微调的VQ-VAE和训练一个HiFi-GAN声码器从头开始的编解码器的重建频谱图。在大约7.5 kbits/s的16 kHz的语音操作,我们的系统进行了评估,对几个国家的最先进的神经编解码器使用客观的指标,如STOI,PESQ,MCD,和ViSQOL。结果表明,我们的简化,无残留的架构实现了有竞争力的感知质量和可理解性,验证它作为一个有效的和开放的基础,为未来的低延迟编解码器设计。
摘要:We present a neural speech codec that challenges the need for complex residual vector quantization (RVQ) stacks by introducing a simpler, single-stage quantization approach. Our method operates directly on the mel-spectrogram, treating it as a 2D data and quantizing non-overlapping 4x4 patches into a single, shared codebook. This patchwise design simplifies the architecture, enables low-latency streaming, and yields a discrete latent grid. To ensure high-fidelity synthesis, we employ a late-stage adversarial fine-tuning for the VQ-VAE and train a HiFi-GAN vocoder from scratch on the codec's reconstructed spectrograms. Operating at approximately 7.5 kbits/s for 16 kHz speech, our system was evaluated against several state-of-the-art neural codecs using objective metrics such as STOI, PESQ, MCD, and ViSQOL. The results demonstrate that our simplified, non-residual architecture achieves competitive perceptual quality and intelligibility, validating it as an effective and open foundation for future low-latency codec designs.


【20】FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
标题:FireRedRTS-2:播客和Chatbot的长对话语音生成
链接:https://arxiv.org/abs/2509.02020

摘要:目前的对话生成方法通常需要完整的对话文本合成之前,并产生一个单一的,不可分割的语音包含所有的声音,使他们不适合交互式聊天,此外,他们遭受不稳定的合成,不准确的扬声器过渡,和不连贯的韵律。在这项工作中,我们提出了FireRedTTS-2,一个长格式的流TTS系统,多扬声器对话生成,提供稳定,自然的语音与可靠的扬声器切换和上下文感知韵律。一个新的12.5Hz流语音标记器加速了训练和推理,延长了最大对话长度,编码了更丰富的语义以稳定文本到令牌的建模,并支持实时应用的高保真流生成。我们采用了文本语音交错格式,连接扬声器标记的文本与对齐的语音令牌按时间顺序排列,并建模与双Transformer:一个大的解码器只Transformer预测令牌在第一层,和一个较小的完成后续层。实验结果表明,FireRedTTS-2与聊天框架无缝集成,并以最小的微调,产生情感表达的语音由隐式上下文线索的指导。在播客生成中,它超越了现有的系统,包括MoonCast,Zipvoice-Dialogue和MOSS-TTSD在客观的可懂度,说话者转向的可靠性,和感知自然与上下文一致的韵律。我们的演示可在https://fireredteam.github.io/demos/firered_tts_2上获得。
摘要:Current dialogue generation approaches typically require the complete dialogue text before synthesis and produce a single, inseparable speech containing all voices, making them unsuitable for interactive chat; moreover, they suffer from unstable synthesis, inaccurate speaker transitions, and incoherent prosody. In this work, we present FireRedTTS-2, a long-form streaming TTS system for multi-speaker dialogue generation, delivering stable, natural speech with reliable speaker switching and context-aware prosody. A new 12.5Hz streaming speech tokenizer accelerates training and inference, extends maximum dialogue length, encodes richer semantics to stabilize text-to-token modeling and supports high-fidelity streaming generation for real-time applications. We adopt a text-speech interleaved format, concatenating speaker-labeled text with aligned speech tokens in chronological order, and model it with a dual-transformer: a large decoder-only transformer predicts tokens at the first layer, and a smaller one completes subsequent layers. Experimental results show that FireRedTTS-2 integrates seamlessly with chat frameworks and, with minimal fine-tuning, produces emotionally expressive speech guided by implicit contextual cues. In podcast generation, it surpasses existing systems including MoonCast, Zipvoice-Dialogue, and MOSS-TTSD in objective intelligibility, speaker-turn reliability, and perceived naturalness with context-consistent prosody. Our demos are available at https://fireredteam.github.io/demos/firered_tts_2.


【21】From Discord to Harmony: Decomposed Consonance-based Training for Improved Audio Chord Estimation
标题:从不和谐到和谐:用于改进音频和弦估计的分解基于协和的训练
链接:https://arxiv.org/abs/2509.01588

备注:9 pages, 3 figures, 3 tables
摘要:音频和弦估计(ACE)在音乐信息研究中发挥着关键作用,由于其与音乐转录和分析的相关性,二十多年来一直受到关注。尽管取得了显著的进步,但任务中仍然存在挑战,特别是关于谐波含量的独特特性,这导致现有系统的性能达到玻璃天花板。这些挑战包括注释者的主观性,其中注释者之间的不同解释导致不一致,以及和弦数据集内的类不平衡,其中某些和弦类与其他和弦类相比被过度表示,这给模型训练和评估带来了困难。作为第一个贡献,本文提出了一个评价的注释者之间的协议,在和弦注释,使用的指标,超越传统的二进制措施。此外,我们提出了一个和谐的距离度量,反映了和谐注释之间的感知相似性。我们的分析表明,基于和谐的距离度量更有效地捕捉音乐有意义的协议之间的注释。扩展这些发现,我们引入了一种新的ACE一致性为基础的模型,整合到模型中的概念,通过基于辅音的标签平滑的辅音。该模型还通过分别估计根音、低音和所有音符激活来解决类不平衡问题,从而能够从分解的输出中重建和弦标签。
摘要:Audio Chord Estimation (ACE) holds a pivotal role in music information research, having garnered attention for over two decades due to its relevance for music transcription and analysis. Despite notable advancements, challenges persist in the task, particularly concerning unique characteristics of harmonic content, which have resulted in existing systems' performances reaching a glass ceiling. These challenges include annotator subjectivity, where varying interpretations among annotators lead to inconsistencies, and class imbalance within chord datasets, where certain chord classes are over-represented compared to others, posing difficulties in model training and evaluation. As a first contribution, this paper presents an evaluation of inter-annotator agreement in chord annotations, using metrics that extend beyond traditional binary measures. In addition, we propose a consonance-informed distance metric that reflects the perceptual similarity between harmonic annotations. Our analysis suggests that consonance-based distance metrics more effectively capture musically meaningful agreement between annotations. Expanding on these findings, we introduce a novel ACE conformer-based model that integrates consonance concepts into the model through consonance-based label smoothing. The proposed model also addresses class imbalance by separately estimating root, bass, and all note activations, enabling the reconstruction of chord labels from decomposed outputs.


【22】ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition
标题:Arabspel Net:一个轻量级混合2D CNN-BiLSTM模型,注重稳健的阿拉伯语语音情感识别
链接:https://arxiv.org/abs/2509.01401

备注:Accepted (The Third Arabic Natural Language Processing Conference)
摘要:语音情感识别对于人机交互至关重要,特别是对于像阿拉伯语这样的低资源语言,由于有限的数据和研究而面临挑战。我们引入ArabnetNet,这是一种轻量级架构,旨在克服这些限制并提供最先进的性能。与之前依赖于离散MFCC特征和1D卷积的系统不同,这些系统错过了细微的频谱-时间模式,ArabaNet使用通过2D卷积处理的Mel频谱图,保留了传统方法中经常丢失的关键情感线索。   虽然最近的模型倾向于拥有数百万个参数的大规模架构,但ArabnetNet仅用100万个参数就取得了优异的结果,比HuBERT base小90倍,比Whisper小74倍。这种效率使其成为资源受限环境的理想选择。阿拉伯语网络推进阿拉伯语语音情感识别,为现实世界的应用程序提供卓越的性能和可访问性。
摘要:Speech emotion recognition is vital for human-computer interaction, particularly for low-resource languages like Arabic, which face challenges due to limited data and research. We introduce ArabEmoNet, a lightweight architecture designed to overcome these limitations and deliver state-of-the-art performance. Unlike previous systems relying on discrete MFCC features and 1D convolutions, which miss nuanced spectro-temporal patterns, ArabEmoNet uses Mel spectrograms processed through 2D convolutions, preserving critical emotional cues often lost in traditional methods.   While recent models favor large-scale architectures with millions of parameters, ArabEmoNet achieves superior results with just 1 million parameters, 90 times smaller than HuBERT base and 74 times smaller than Whisper. This efficiency makes it ideal for resource-constrained environments. ArabEmoNet advances Arabic speech emotion recognition, offering exceptional performance and accessibility for real-world applications.


【23】CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays
标题:CabinSep:基于IR增强屏蔽的MVDR,通过分布式异类阵列实现实时车内语音分离
链接:https://arxiv.org/abs/2509.01399

摘要:从多个说话者中分离重叠语音对于有效的人车交互至关重要。本文提出了CabinSep,一种轻量级的基于神经掩码的最小方差无失真响应(MVDR)语音分离方法,以减少后端自动语音识别(ASR)模型中的语音识别错误。我们的贡献有三个方面:首先,我们利用信道信息来提取空间特征,这提高了语音和噪声掩模的估计。其次,我们在推理过程中使用MVDR,减少语音失真,使其对ASR更友好。第三,我们介绍了一种数据增强方法相结合的模拟和真实记录的脉冲响应(IR),提高扬声器定位在区域边界,进一步减少语音识别错误。CabinSep的计算复杂度仅为0.4 GMAC,与最先进的DualSep模型相比,在真实记录的数据集中,语音识别错误率相对降低了17.5%。演示可在以下网址获得:https://cabinsep.github.io/cabinsep/。
摘要:Separating overlapping speech from multiple speakers is crucial for effective human-vehicle interaction. This paper proposes CabinSep, a lightweight neural mask-based minimum variance distortionless response (MVDR) speech separation approach, to reduce speech recognition errors in back-end automatic speech recognition (ASR) models. Our contributions are threefold: First, we utilize channel information to extract spatial features, which improves the estimation of speech and noise masks. Second, we employ MVDR during inference, reducing speech distortion to make it more ASR-friendly. Third, we introduce a data augmentation method combining simulated and real-recorded impulse responses (IRs), improving speaker localization at zone boundaries and further reducing speech recognition errors. With a computational complexity of only 0.4 GMACs, CabinSep achieves a 17.5% relative reduction in speech recognition error rate in a real-recorded dataset compared to the state-of-the-art DualSep model. Demos are available at: https://cabinsep.github.io/cabinsep/.


【24】Analysing the Language of Neural Audio Codecs
标题:神经音频编解码器语言分析
链接:https://arxiv.org/abs/2509.01390

备注:In Proceedings of 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2025)
摘要:本研究对神经音频编解码器(NAC)的统计和语言特性进行了比较分析。我们调查离散的语音令牌所产生的各种NAC模型,检查他们遵守语言统计规律,如齐普夫定律和堆定律,以及他们的熵和冗余。为了评估这些标记级属性如何与合成语音中的语义和声学保留相关,我们使用自动语音识别的错误率评估可懂度,并使用UTMOS分数评估质量。我们的研究结果表明,NAC令牌,特别是3克,表现出类似语言的统计模式。此外,这些属性,连同信息内容的措施,被发现与语音识别和再合成任务中的性能改善。这些发现提供了对NAC标记序列结构的深入了解,并为更有效的生成语音模型的设计提供了信息。
摘要:This study presents a comparative analysis of the statistical and linguistic properties of neural audio codecs (NACs). We investigate discrete speech tokens produced by various NAC models, examining their adherence to linguistic statistical laws such as Zipf's law and Heaps' law, as well as their entropy and redundancy. To assess how these token-level properties relate to semantic and acoustic preservation in synthesized speech, we evaluate intelligibility using error rates of automatic speech recognition, and quality using the UTMOS score. Our results reveal that NAC tokens, particularly 3-grams, exhibit language-like statistical patterns. Moreover, these properties, together with measures of information content, are found to correlate with improved performances in speech recognition and resynthesis tasks. These findings offer insights into the structure of NAC token sequences and inform the design of more effective generative speech models.


【25】The AudioMOS Challenge 2025
标题:2025年AudioMOS挑战赛
链接:https://arxiv.org/abs/2509.01336

备注:IEEE ASRU 2025
摘要:这是AudioMOS Challenge 2025的总结论文,这是合成音频自动主观质量预测的第一个挑战。挑战包括三条轨道。第一首曲目旨在从整体质量和文本对齐方面评估文本到音乐样本。第二个轨道是基于Meta Audiobox Aesthetics的四个评估维度,测试集由文本到语音,文本到音频和文本到音乐样本组成。第三部分是不同采样率下的合成语音质量评价。这项挑战吸引了来自学术界和工业界的24个独特团队,并确认了基线的改进。这一挑战的结果预计将促进音频生成系统自动评估领域的发展和进步。
摘要:This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set consists of text-to-speech, text-to-audio, and text-to-music samples. The third track focuses on synthetic speech quality assessment in different sampling rates. The challenge attracted 24 unique teams from both academia and industry, and improvements over the baselines were confirmed. The outcome of this challenge is expected to facilitate development and progress in the field of automatic evaluation for audio generation systems.


【26】SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation
标题:Simulega:MoE路由器是同步语音翻译的高级政策制定者
链接:https://arxiv.org/abs/2509.01200

摘要:同步语音翻译(SimulST)通过在严格的延迟限制下联合优化语音识别和机器翻译,实现实时跨语言通信。现有系统难以平衡翻译质量、延迟和语义一致性,特别是在多语言多对多场景中,不同的读写策略阻碍了统一的策略学习。在本文中,我们提出了SimulMEGA(通过混合专家门控同时生成),这是一个无监督的策略学习框架,它将基于前缀的训练与混合专家细化器相结合,以隐式方式学习有效的读写决策,而不会增加推理时间开销。我们的设计只需要对标准Transformer架构进行最小限度的修改,并且可以通用于语音转文本和文本转语音流传输任务。通过对六种语言对的综合评估,我们的500 M参数语音到文本模型优于Seamless基线,在1.5秒的平均延迟下实现了低于7%的BLEU降级,在3秒内实现了低于3%的BLEU降级。我们进一步证明了SimulMEGA的多功能性,通过将其扩展到具有单向骨干的流式TTS,从而获得卓越的延迟质量权衡。
摘要:Simultaneous Speech Translation (SimulST) enables real-time cross-lingual communication by jointly optimizing speech recognition and machine translation under strict latency constraints. Existing systems struggle to balance translation quality, latency, and semantic coherence, particularly in multilingual many-to-many scenarios where divergent read and write policies hinder unified strategy learning. In this paper, we present SimulMEGA (Simultaneous Generation by Mixture-of-Experts Gating), an unsupervised policy learning framework that combines prefix-based training with a Mixture-of-Experts refiner to learn effective read and write decisions in an implicit manner, without adding inference-time overhead. Our design requires only minimal modifications to standard transformer architectures and generalizes across both speech-to-text and text-to-speech streaming tasks. Through comprehensive evaluation on six language pairs, our 500M parameter speech-to-text model outperforms the Seamless baseline, achieving under 7 percent BLEU degradation at 1.5 seconds average lag and under 3 percent at 3 seconds. We further demonstrate the versatility of SimulMEGA by extending it to streaming TTS with a unidirectional backbone, yielding superior latency quality tradeoffs.


【27】EZhouNet:A framework based on graph neural network and anchor interval for the respiratory sound event detection
标题:EZhouNet:基于图神经网络和锚点区间的呼吸音事件检测框架
链接:https://arxiv.org/abs/2509.01153

摘要:听诊是呼吸系统和肺部疾病早期诊断的关键方法,依赖于熟练的医疗保健专业人员。然而,这一过程往往是主观的,专家之间存在差异。因此,出现了许多基于深度学习的自动分类方法,其中大多数侧重于呼吸声分类。相比之下,对呼吸声事件检测的研究仍然有限。现有的声音事件检测方法通常依赖于帧级预测,然后进行后处理以生成事件级输出,这使得直接学习间隔边界具有挑战性。此外,许多方法只能处理固定长度的音频,限制了它们对可变长度呼吸声的适用性。此外,呼吸声位置信息对检测性能的影响尚未得到广泛研究。为了解决这些问题,我们提出了一种具有锚定间隔的基于图神经网络的框架,能够处理可变长度的音频并为异常呼吸声事件提供更精确的时间定位。我们的方法提高了呼吸音检测的灵活性和适用性。在SPRSound 2024和HF Lung V1数据集上的实验证明了该方法的有效性,并结合呼吸位置信息增强了异常声音之间的区分。
摘要:Auscultation is a key method for early diagnosis of respiratory and pulmonary diseases, relying on skilled healthcare professionals. However, the process is often subjective, with variability between experts. As a result, numerous deep learning-based automatic classification methods have emerged, most of which focus on respiratory sound classification. In contrast, research on respiratory sound event detection remains limited. Existing sound event detection methods typically rely on frame-level predictions followed by post-processing to generate event-level outputs, making interval boundaries challenging to learn directly. Furthermore, many approaches can only handle fixed-length audio, lim- iting their applicability to variable-length respiratory sounds. Additionally, the impact of respiratory sound location information on detection performance has not been extensively explored. To address these issues, we propose a graph neural network-based framework with anchor intervals, capable of handling variable-length audio and providing more precise temporal localization for abnormal respi- ratory sound events. Our method improves both the flexibility and applicability of respiratory sound detection. Experiments on the SPRSound 2024 and HF Lung V1 datasets demonstrate the effec- tiveness of the proposed approach, and incorporating respiratory position information enhances the discrimination between abnormal sounds.


【28】A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR
标题:自我监督孟加拉方言ASB的统一去噪和适应框架
链接:https://arxiv.org/abs/2509.00988

摘要:孟加拉语是世界第五大语言,其自动语音识别(ASR)仍然是一个重大挑战,严重阻碍了超过2.7亿人的技术可访问性。这一挑战由于两个持续存在且相互交织的因素而变得更加复杂:语言的巨大方言多样性和现实环境中普遍存在的噪音。虽然最先进的自监督学习(SSL)模型为低资源语言提供了先进的ASR,但它们通常缺乏明确的机制来处理预训练期间的环境噪声或针对孟加拉语方言中复杂语音和词汇变化的专门适应策略。本文介绍了一种新颖的,统一的框架,旨在同时解决这些双重挑战。我们的方法建立在WavLM模型的基础上,WavLM模型是用掩蔽语音去噪目标进行唯一预训练的,使其对声学失真具有固有的鲁棒性。我们提出了一种专门的多阶段微调策略,首先将模型适应通用领域标准孟加拉语,以建立强大的语言基础,然后通过有针对性的数据增强将其专门用于噪声鲁棒性方言识别。该框架在广泛的模拟噪声条件下,从干净的音频到低信噪比(SNR)水平,在包括多种孟加拉方言的综合基准上进行了严格的评估。   实验结果表明,该框架显着优于强基线,包括标准微调wav2vec 2.0和大规模多语言Whisper模型。这项工作为这项任务建立了一个新的最先进的技术,并为全球其他低资源,高变化的语言开发实用的ASR系统提供了一个可扩展的,有效的蓝图。
摘要:Automatic Speech Recognition (ASR) for Bengali, the world's fifth most spoken language, remains a significant challenge, critically hindering technological accessibility for its over 270 million speakers. This challenge is compounded by two persistent and intertwined factors: the language's vast dialectal diversity and the prevalence of acoustic noise in real-world environments. While state-of-the-art self-supervised learning (SSL) models have advanced ASR for low-resource languages, they often lack explicit mechanisms to handle environmental noise during pre-training or specialized adaptation strategies for the complex phonetic and lexical variations across Bengali dialects. This paper introduces a novel, unified framework designed to address these dual challenges simultaneously. Our approach is founded on the WavLM model, which is uniquely pre-trained with a masked speech denoising objective, making it inherently robust to acoustic distortions. We propose a specialized multi-stage fine-tuning strategy that first adapts the model to general-domain standard Bengali to establish a strong linguistic foundation and subsequently specializes it for noise-robust dialectal recognition through targeted data augmentation. The framework is rigorously evaluated on a comprehensive benchmark comprising multiple Bengali dialects under a wide range of simulated noisy conditions, from clean audio to low Signal-to-Noise Ratio (SNR) levels.   Experimental results demonstrate that the proposed framework significantly outperforms strong baselines, including standard fine-tuned wav2vec 2.0 and the large-scale multilingual Whisper model. This work establishes a new state-of-the-art for this task and provides a scalable, effective blueprint for developing practical ASR systems for other low-resource, high-variation languages globally.


【29】IoT-based Noise Monitoring using Mobile Nodes for Smart Cities
标题:使用移动节点实现智能城市基于物联网的噪音监控
链接:https://arxiv.org/abs/2509.00979

摘要:城市噪声污染对公众健康构成重大威胁,但现有的监测基础设施提供有限的空间覆盖范围和适应性。本文提出了一种可扩展的、低成本的、基于物联网的、使用移动节点(移动车辆上的传感器节点)的实时环境噪声监测解决方案。该系统利用一个低成本的声音传感器与GPS功能模块集成,以一秒的间隔收集地理标记的噪音数据。声音节点在实验室环境中针对参考声级计进行校准,以使用各种机器学习(ML)算法来确保准确性,所述机器学习(ML)算法诸如简单线性回归(SLR)、多元线性回归(MLR)、多项式回归(PR)、分段回归(SR)、支持向量回归(SVR)、决策树(DT)和随机森林回归(RFR)。虽然实验室校准证明了高精度,它表明,节点的性能下降,在移动车辆中的数据收集过程中。为了解决这个问题,证明必须根据在移动环境中收集的数据以及参考设备对基于物联网的节点进行校准。在所采用的ML模型中,RFR实现了最佳性能,R2为0.937,RMSE为1.09。该系统部署在印度海得拉巴,通过27天的三次测量活动,捕获了436,420个数据点。结果突出了工作日,周末和排灯节期间的时间和空间噪声变化。将车辆速度纳入校准中显著提高了准确性。该系统展示了在智慧城市中广泛部署基于物联网的噪声传感网络的潜力,从而实现有效的噪声污染管理和城市规划。
摘要:Urban noise pollution poses a significant threat to public health, yet existing monitoring infrastructures offer limited spatial coverage and adaptability. This paper presents a scalable, low-cost, IoT-based, real-time environmental noise monitoring solution using mobile nodes (sensor nodes on a moving vehicle). The system utilizes a low-cost sound sensor integrated with GPS-enabled modules to collect geotagged noise data at one-second intervals. The sound nodes are calibrated against a reference sound level meter in a laboratory setting to ensure accuracy using various machine learning (ML) algorithms, such as Simple Linear Regression (SLR), Multiple Linear Regression (MLR), Polynomial Regression (PR), Segmented Regression (SR), Support Vector Regression (SVR), Decision Tree (DT), and Random Forest Regression (RFR). While laboratory calibration demonstrates high accuracy, it is shown that the performance of the nodes degrades during data collection in a moving vehicle. To address this, it is demonstrated that the calibration must be performed on the IoT-based node based on the data collected in a moving environment along with the reference device. Among the employed ML models, RFR achieved the best performance with an R2 of 0.937 and RMSE of 1.09 for mobile calibration. The system was deployed in Hyderabad, India, through three measurement campaigns across 27 days, capturing 436,420 data points. Results highlight temporal and spatial noise variations across weekdays, weekends, and during Diwali. Incorporating vehicular velocity into the calibration significantly improves accuracy. The proposed system demonstrates the potential for widespread deployment of IoT-based noise sensing networks in smart cities, enabling effective noise pollution management and urban planning.


【30】TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
标题:TinyMusician:基于知识蒸馏和混合精确量化的设备上音乐生成
链接:https://arxiv.org/abs/2509.00914

备注:12 pages for main context, 5 figures
摘要:生成模式的成功在音乐生成领域引起了前所未有的关注。基于Transformer的架构为模型性能设定了新的基准。然而,它们的实际采用受到一些关键挑战的阻碍:由于它们的大量参数,需要大量的计算资源和推理时间。这些障碍使得它们无法部署在计算资源有限的边缘设备上,例如智能手机和可穿戴设备。在这项工作中,我们提出了TinyMusician,一个轻量级的音乐生成模型从MusicGen(一个国家的最先进的音乐生成模型)蒸馏。TinyMusician集成了两项创新:(i)阶段混合双向和倾斜KL发散和(ii)自适应混合精度量化。实验结果表明,TinyMusician保留了MusicGen-Small 93%的性能,而模型大小减少了55%。TinyMusician是第一个可移动部署的音乐生成模型,消除了对云的依赖,同时保持了高音频保真度和高效的资源使用率
摘要:The success of the generative model has gained unprecedented attention in the music generation area. Transformer-based architectures have set new benchmarks for model performance. However, their practical adoption is hindered by some critical challenges: the demand for massive computational resources and inference time, due to their large number of parameters. These obstacles make them infeasible to deploy on edge devices, such as smartphones and wearables, with limited computational resources. In this work, we present TinyMusician, a lightweight music generation model distilled from MusicGen (a State-of-the-art music generation model). TinyMusician integrates two innovations: (i) Stage-mixed Bidirectional and Skewed KL-Divergence and (ii) Adaptive Mixed-Precision Quantization. The experimental results demonstrate that TinyMusician retains 93% of the MusicGen-Small performance with 55% less model size. TinyMusician is the first mobile-deployable music generation model that eliminates cloud dependency while maintaining high audio fidelity and efficient resource usage


【31】Speech Command Recognition Using LogNNet Reservoir Computing for Embedded Systems
标题:嵌入式系统中使用LogNNet水库计算的语音命令识别
链接:https://arxiv.org/abs/2509.00862

备注:20 pages, 6 figures
摘要:本文提出了一种低资源的语音命令识别器,结合基于能量的语音活动检测(VAD),优化的梅尔频率倒谱系数(MFCC)管道,和LogNNet算法计算分类器。使用四个命令从语音命令da-taset下采样到8 kHz,我们评估四个MFCC聚合方案,并发现自适应分仓(64维特征向量)提供了最好的精度紧凑性权衡。架构为64:33:9:4的LogNNet分类器在说话人独立评估下达到了92.04%的准确率,同时需要的参数比传统的深度学习模型少得多。在Arduino Nano 33 IoT(ARM Cor-tex-M0+,48 MHz,32 KB RAM)上的硬件实现验证了实际可行性,实现了约90%的实时识别准确率,同时仅消耗18 KB RAM(55%利用率)。因此,完整的管道(VAD -> MFCC -> LogNNet)可以在严格的内存和计算限制下实现可靠的设备上语音命令识别,使其适用于电池供电的物联网节点、无线传感器网络和免提控制接口。
摘要:This paper presents a low-resource speech-command recognizer combining energy-based voice activity detection (VAD), an optimized Mel-Frequency Cepstral Coefficients (MFCC) pipeline, and the LogNNet reservoir-computing classifier. Using four commands from the Speech Commands da-taset downsampled to 8 kHz, we evaluate four MFCC aggregation schemes and find that adaptive binning (64-dimensional feature vector) offers the best accuracy-to-compactness trade-off. The LogNNet classifier with architecture 64:33:9:4 reaches 92.04% accuracy under speaker-independent evaluation, while requiring significantly fewer parameters than conventional deep learn-ing models. Hardware implementation on Arduino Nano 33 IoT (ARM Cor-tex-M0+, 48 MHz, 32 KB RAM) validates the practical feasibility, achieving ~90% real-time recognition accuracy while consuming only 18 KB RAM (55% utilization). The complete pipeline (VAD -> MFCC -> LogNNet) thus enables reliable on-device speech-command recognition under strict memory and compute limits, making it suitable for battery-powered IoT nodes, wire-less sensor networks, and hands-free control interfaces.


【32】Adaptive Vehicle Speed Classification via BMCNN with Reinforcement Learning-Enhanced Acoustic Processing
标题:通过BMCNN和强化学习增强声学处理的自适应车辆速度分类
链接:https://arxiv.org/abs/2509.00839

摘要:交通拥堵仍然是一个紧迫的城市挑战,需要智能交通系统进行实时管理。我们提出了一个混合框架,结合了深度学习和强化学习的声学车辆速度分类。双分支BMCNN处理MFCC和小波特征以捕获互补频率模式。注意力增强的DQN自适应地选择最小数量的音频帧,并在达到置信度阈值时触发早期决策。对IDMT-Traffic和我们的SZUR-Acoustic(苏州)数据集的评估显示,准确率分别为95.99%和92.3%,通过提前终止,平均处理速度提高了1.63倍。与A3 C、DDDQN、SA 2C、PPO和TD 3相比,该方法提供了优越的精度-效率权衡,适合于异构城市环境中的实时ITS部署。
摘要:Traffic congestion remains a pressing urban challenge, requiring intelligent transportation systems for real-time management. We present a hybrid framework that combines deep learning and reinforcement learning for acoustic vehicle speed classification. A dual-branch BMCNN processes MFCC and wavelet features to capture complementary frequency patterns. An attention-enhanced DQN adaptively selects the minimal number of audio frames and triggers early decisions once confidence thresholds are reached. Evaluations on IDMT-Traffic and our SZUR-Acoustic (Suzhou) datasets show 95.99% and 92.3% accuracy, with up to 1.63x faster average processing via early termination. Compared with A3C, DDDQN, SA2C, PPO, and TD3, the method provides a superior accuracy-efficiency trade-off and is suitable for real-time ITS deployment in heterogeneous urban environments.


【33】AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation
标题:Aimocilles:评估文本到音乐生成中情感传递的基准
链接:https://arxiv.org/abs/2509.00813

备注:to be published in HCMIR25: 3rd Workshop on Human-Centric Music Information Research
摘要:文本到音乐(TTM)生成的最新进展使得能够使用自然语言提示进行可控的和富有表现力的音乐创作。然而,与人类偏好或文本对齐相比,TTM系统的情感保真度在很大程度上仍然未被探索。在这项研究中,我们介绍了AImoclips,这是一个用于评估TTM系统如何向人类听众传达预期情感的基准,涵盖了开源和商业模式。我们选择了12个情绪意图,跨越了效价-唤醒空间的四个象限,并使用了六个最先进的TTM系统来生成超过1,000个音乐片段。共有111名参与者在9点Likert量表上对每个剪辑的感知效价和唤醒进行了评分。我们的研究结果表明,商业系统往往会产生比预期更令人愉快的音乐,而开源系统往往会表现相反。在所有模型中,情绪在高唤醒条件下都能更准确地传达。此外,所有的系统都表现出对情感中立的偏见,突出了情感可控性的关键限制。该基准测试为模型特定的情感渲染特性提供了有价值的见解,并支持情感对齐TTM系统的未来发展。
摘要:Recent advances in text-to-music (TTM) generation have enabled controllable and expressive music creation using natural language prompts. However, the emotional fidelity of TTM systems remains largely underexplored compared to human preference or text alignment. In this study, we introduce AImoclips, a benchmark for evaluating how well TTM systems convey intended emotions to human listeners, covering both open-source and commercial models. We selected 12 emotion intents spanning four quadrants of the valence-arousal space, and used six state-of-the-art TTM systems to generate over 1,000 music clips. A total of 111 participants rated the perceived valence and arousal of each clip on a 9-point Likert scale. Our results show that commercial systems tend to produce music perceived as more pleasant than intended, while open-source systems tend to perform the opposite. Emotions are more accurately conveyed under high-arousal conditions across all models. Additionally, all systems exhibit a bias toward emotional neutrality, highlighting a key limitation in affective controllability. This benchmark offers valuable insights into model-specific emotion rendering characteristics and supports future development of emotionally aligned TTM systems.


【34】PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
标题:PicoAudio2:自然语言描述的时间可控文本到音频生成
链接:https://arxiv.org/abs/2509.00683

备注:Demo page: this https URL
摘要:可控文本到音频生成(TTA)最近引起了人们的广泛关注。虽然现有的作品可以实现细粒度的可控性的基础上的时间戳信息,声音事件类别被限制在一个固定的集合。此外,由于仅使用模拟数据进行训练,因此所生成的音频质量和对真实数据的泛化性能受到限制。为了解决这个问题,我们提出了PicoAudio2,通过新的数据处理管道和模型架构来改进时间可控的TTA。具体来说,我们使用接地模型来注释真实音频文本数据集的事件时间戳,以管理时间上强的真实数据,以及现有作品的模拟数据。该模型是在真实和模拟数据的组合上训练的。此外,在PicoAudio之后,我们将时间戳信息编码到时间戳矩阵中,以在粗粒度文本描述之上为模型提供额外的细粒度时间对齐信息。实验表明,PicoAudio2在时间可控性和音频质量方面表现出优越的性能。
摘要:Controllable text-to-audio generation (TTA) has attracted much attention recently. Although existing works can achieve fine-grained controllability based on timestamp information, sound event categories are limited to a fixed set. Moreover, since only simulated data is used for training, the generated audio quality and generalization performance on real data are limited. To tackle this issue, we propose PicoAudio2, improving temporal-controllable TTA via a new data processing pipeline and model architecture. Specifically, we use a grounding model to annotate event timestamps of real audio-text datasets to curate temporally-strong real data, in addition to simulation data from existing works. The model is trained on the combination of real and simulation data. Moreover, following PicoAudio, we encode timestamp information into a timestamp matrix to provide extra fine-grained time-aligned information to the model, on top of the coarse-grained textual description. Experiments show that PicoAudio2 exhibits superior performance in terms of temporal controllability and audio quality.


【35】The Name-Free Gap: Policy-Aware Stylistic Control in Music Generation
标题:无名差距:音乐世代中的政策意识风格控制
链接:https://arxiv.org/abs/2509.00654

备注:10 pages, 2 figures
摘要:文本到音乐模型捕捉广泛的属性,如乐器或情绪,但细粒度的风格控制仍然是一个开放的挑战。现有的风格化方法通常需要重新训练或专门的条件反射,这使得再现性变得复杂,并且在艺术家姓名受到限制时限制了政策合规性。我们研究是否轻量级的,人类可读的修饰符采样从一个大型的语言模型可以提供一个策略强大的替代风格控制。使用MusicGen-small,我们评估两个艺术家:Billie Eilish(声乐流行)和Ludovico Einaudi(器乐钢琴)。对于每个艺术家,我们使用15个参考摘录,并在三个条件下评估匹配的种子:基线提示,艺术家姓名提示和五个描述符集。所有提示都是使用大型语言模型生成的。评估使用VGGish和CLAP嵌入以及分布和每个剪辑的相似性度量,包括新的最小距离归因度量。结果表明,艺术家的名字是最强的控制信号,在这两个艺术家,而无名称的描述符恢复这种效果。这突出表明,现有的保障措施,如在音乐生成提示中限制艺术家姓名,可能无法完全防止风格模仿。跨艺术家转移减少对齐,表明描述符编码有针对性的风格线索。我们还提出了一个描述符表在十个当代艺术家,以说明令牌的广度。这些发现共同定义了无名称的差距,艺术家的名字提示和符合政策的描述符之间的可控性差异,通过一个可重复的评估协议,为艺术家级别的可控性。
摘要:Text-to-music models capture broad attributes such as instrumentation or mood, but fine-grained stylistic control remains an open challenge. Existing stylization methods typically require retraining or specialized conditioning, which complicates reproducibility and limits policy compliance when artist names are restricted. We study whether lightweight, human-readable modifiers sampled from a large language model can provide a policy-robust alternative for stylistic control. Using MusicGen-small, we evaluate two artists: Billie Eilish (vocal pop) and Ludovico Einaudi (instrumental piano). For each artist, we use fifteen reference excerpts and evaluate matched seeds under three conditions: baseline prompts, artist-name prompts, and five descriptor sets. All prompts are generated using a large language model. Evaluation uses both VGGish and CLAP embeddings with distributional and per-clip similarity measures, including a new min-distance attribution metric. Results show that artist names are the strongest control signal across both artists, while name-free descriptors recover much of this effect. This highlights that existing safeguards such as the restriction of artist names in music generation prompts may not fully prevent style imitation. Cross-artist transfers reduce alignment, showing that descriptors encode targeted stylistic cues. We also present a descriptor table across ten contemporary artists to illustrate the breadth of the tokens. Together these findings define the name-free gap, the controllability difference between artist-name prompts and policy-compliant descriptors, shown through a reproducible evaluation protocol for prompt-level controllability.


【36】Real-Time Piano Note Frequency Detection Using FPGA and FFT Core
标题:利用现场可编程逻辑器件和快速傅里叶变换核实现钢琴音符频率检测
链接:https://arxiv.org/abs/2509.00589

备注:20 pages, 11 Figures
摘要:钢琴等乐器的实时频率分析是电子调谐器、音乐可视化器和现场声音监控等领域的重要功能。传统方法通常依赖于基于软件的数字信号处理(DSP),这可能会引入延迟并需要大量的计算能力。相比之下,诸如FPGA(现场可编程门阵列)的硬件平台由于其并行处理能力而提供了以更快的速度和确定性执行此类分析的能力。该项目的主要目标是使用基于FPGA的实时快速傅立叶变换(FFT)系统分析来自数字钢琴的模拟音频信号。
摘要:Real-time frequency analysis of musical instruments, such as the piano, is an essential feature in areas like electronic tuners, music visualizers, and live sound monitoring. Traditional methods often rely on software-based digital signal processing (DSP), which may introduce latency and require significant computational power. In contrast, hardware platforms such as FPGAs (Field Programmable Gate Arrays) offer the ability to perform such analyses with greater speed and determinism due to their parallel processing capabilities. The primary objective of this project was to analyze analog audio signals from a digital piano using an FPGA-based real-time Fast Fourier Transform (FFT) system.


【37】Entropy-based Coarse and Compressed Semantic Speech Representation Learning
标题:基于信息量的粗压缩语义语音表示学习
链接:https://arxiv.org/abs/2509.00503

摘要:离散语音表示学习最近在声学和语义建模中引起了越来越多的兴趣。现有方法通常以每秒25或50个令牌的速率将16 kHz波形编码成离散令牌。然而,考虑到语音通常每秒仅传达2到5个单词,这种细粒度的标记化引入了冗余并阻碍了下游训练和推理的效率。此外,在此频率下的语义语音表示主要捕获语音级信息,而语义理解可能不需要这样详细的标记级分辨率。为了解决这些局限性,我们提出了一个基于熵的动态聚合框架学习压缩的语义语音表示。首先通过对大规模未标记数据的下一个标记预测来预训练语音语言模型,以捕获频繁的标记模式。然后使用预测熵来自适应地确定聚合边界,随后是融合每个片段内的信息的交叉注意模块。通过调整熵阈值,可以灵活地控制表示的粒度和压缩比。ASR,语音到文本的翻译,语音转换任务的实验表明,压缩表示执行相当于或优于密集令牌序列,证明了所提出的方法的有效性。
摘要:Discrete speech representation learning has recently attracted increasing interest in both acoustic and semantic modeling. Existing approaches typically encode 16 kHz waveforms into discrete tokens at a rate of 25 or 50 tokens per second. However, given that speech generally conveys only 2 to 5 words per second, such fine-grained tokenization introduces redundancy and hinders efficiency in downstream training and inference. Moreover, semantic speech representations at this frequency primarily capture phonetic-level information, while semantic understanding may not require such detailed token-level resolution. To address these limitations, we propose an entropy-based dynamic aggregation framework for learning compressed semantic speech representations. A speech language model is first pre-trained via next-token prediction on large-scale unlabeled data to capture frequent token patterns. Predictive entropy is then used to adaptively determine aggregation boundaries, followed by a cross-attention module that fuses information within each segment. By adjusting the entropy threshold, the granularity and compression ratio of the representations can be flexibly controlled. Experiments on ASR, speech-to-text translation, and voice conversion tasks demonstrate that the compressed representations perform on par with or better than dense token sequences, demonstrating the effectiveness of the proposed approach.


【38】SaD: A Scenario-Aware Discriminator for Speech Enhancement
标题:SaD:语音增强的场景感知识别器
链接:https://arxiv.org/abs/2509.00405

备注:5 pages, 2 this http URL by InterSpeech2025
摘要:基于生成对抗网络的模型在语音增强领域表现出了卓越的性能。然而,这些模型的当前优化策略主要集中在改进生成器的架构或提高质量评估指标的可重用性。这种方法往往忽略了不同场景中固有的丰富上下文信息。在本文中,我们提出了一个感知的语音识别,捕捉场景的特定功能,并执行频域划分,从而使生成器生成的增强语音的质量评估更准确。我们使用两个公开的数据集对三个代表性模型进行了全面的实验。结果表明,我们的方法可以有效地适应各种生成器架构,而不会改变它们的结构,从而在不同的场景中解锁语音增强的进一步性能增益。
摘要:Generative adversarial network-based models have shown remarkable performance in the field of speech enhancement. However, the current optimization strategies for these models predominantly focus on refining the architecture of the generator or enhancing the quality evaluation metrics of the discriminator. This approach often overlooks the rich contextual information inherent in diverse scenarios. In this paper, we propose a scenario-aware discriminator that captures scene-specific features and performs frequency-domain division, thereby enabling a more accurate quality assessment of the enhanced speech generated by the generator. We conducted comprehensive experiments on three representative models using two publicly available datasets. The results demonstrate that our method can effectively adapt to various generator architectures without altering their structure, thereby unlocking further performance gains in speech enhancement across different scenarios.


【39】Towards High-Fidelity and Controllable Bioacoustic Generation via Enhanced Diffusion Learning
标题:通过增强的扩散学习实现高保真和可控的生物声学生成
链接:https://arxiv.org/abs/2509.00318

摘要:生成建模为生物声学提供了新的机会,可以合成逼真的动物发声,从而支持生物监测工作并补充濒危物种的稀缺数据。然而,直接从嘈杂的现场录音中产生鸟鸣波形仍然是一个重大挑战。   我们提出了BirdDiff,一个生成框架,旨在从12种野生鸟类的嘈杂数据集中合成鸟的叫声。该模型采用了一个“零层”阶段的多尺度自适应鸟叫增强,其次是一个基于扩散的发电机条件下的三种模式:梅尔频率倒谱系数,物种标签,和文字描述。增强阶段提高了信噪比(SNR),同时最大限度地减少频谱失真,与三种广泛使用的非训练增强方法相比,实现了最高的SNR增益(+10.45 dB)和最低的Itakura-Saito距离(0.54)。   我们根据基线生成模型DiffWave评估BirdDiff。我们的方法在生成质量度量方面产生了实质性的改进:Fr 'echet音频距离(0.590至0.213),Jensen-Shannon发散度(0.259至0.226)和统计上不同的Bins数量(7.33至5.58)。为了评估物种特异性细节保留,我们使用在原始数据集上训练的ResNet 50分类器来识别生成的样本。分类准确率从35.9%(DiffWave)提高到70.1%(BirdDiff),12个物种中有8个超过70%的准确率。   这些结果表明,BirdDiff能够直接从嘈杂的现场录音中产生高保真、可控的鸟鸣。
摘要:Generative modeling offers new opportunities for bioacoustics, enabling the synthesis of realistic animal vocalizations that could support biomonitoring efforts and supplement scarce data for endangered species. However, directly generating bird call waveforms from noisy field recordings remains a major challenge.   We propose BirdDiff, a generative framework designed to synthesize bird calls from a noisy dataset of 12 wild bird species. The model incorporates a "zeroth layer" stage for multi-scale adaptive bird-call enhancement, followed by a diffusion-based generator conditioned on three modalities: Mel-frequency cepstral coefficients, species labels, and textual descriptions. The enhancement stage improves signal-to-noise ratio (SNR) while minimizing spectral distortion, achieving the highest SNR gain (+10.45 dB) and lowest Itakura-Saito Distance (0.54) compared to three widely used non-training enhancement methods.   We evaluate BirdDiff against a baseline generative model, DiffWave. Our method yields substantial improvements in generative quality metrics: Fr\'echet Audio Distance (0.590 to 0.213), Jensen-Shannon Divergence (0.259 to 0.226), and Number of Statistically-Different Bins (7.33 to 5.58). To assess species-specific detail preservation, we use a ResNet50 classifier trained on the original dataset to identify generated samples. Classification accuracy improves from 35.9% (DiffWave) to 70.1% (BirdDiff), with 8 of 12 species exceeding 70% accuracy.   These results demonstrate that BirdDiff enables high-fidelity, controllable bird call generation directly from noisy field recordings.


【40】Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
标题:评估Wav 2 Vec 2.0、XLS-R和Whisper中Transformer层对说话人识别任务的有效性
链接:https://arxiv.org/abs/2509.00230

摘要:本研究评估了三种先进的语音编码器模型,Wav 2 Vec 2.0,XLS-R,和耳语,在说话人识别任务的性能。通过微调这些模型并使用SVCCA、k-means聚类和t-SNE可视化分析其分层表示,我们发现Wav 2 Vec 2.0和XLS-R在其早期层中有效地捕获了特定于说话者的特征,并通过微调提高了稳定性和性能。Whisper在更深的层中显示出更好的性能。此外,我们确定了每个模型的Transformer层的最佳数量时,微调说话人识别任务。
摘要:This study evaluates the performance of three advanced speech encoder models, Wav2Vec 2.0, XLS-R, and Whisper, in speaker identification tasks. By fine-tuning these models and analyzing their layer-wise representations using SVCCA, k-means clustering, and t-SNE visualizations, we found that Wav2Vec 2.0 and XLS-R capture speaker-specific features effectively in their early layers, with fine-tuning improving stability and performance. Whisper showed better performance in deeper layers. Additionally, we determined the optimal number of transformer layers for each model when fine-tuned for speaker identification tasks.


【41】Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
标题:语音基础模型从可穿戴传感器数据推广到时间序列任务
链接:https://arxiv.org/abs/2509.00221

备注:Preprint, under review
摘要:语音和传感器时间序列数据都在时域和频域中对信息进行编码,如频谱功率和波形形状。我们表明,语音基础模型学习的表示是独立于域的,并实现最先进的性能从可穿戴传感器的时间序列任务。在从HuBERT和wav2vec 2.0中提取的特征上训练的探针优于从直接在特定于模态的数据集上训练的自监督模型中提取的探针,用于情绪分类,心律失常检测和活动分类任务。我们发现一个特别强的相关性的卷积特征编码器从语音模型可穿戴传感器任务。本文提出的方法使用简单的探测方法,提高了数据稀缺时间序列任务的性能和鲁棒性。这项工作是一个通用的时间序列模型的语音和传感器数据,进一步探索的主题。


【42】Generalizable Audio Spoofing Detection using Non-Semantic Representations
标题:使用非语义表示的可推广音频欺骗检测
链接:https://arxiv.org/abs/2509.00186

备注:None
摘要:生成建模的快速发展使得合成音频生成变得容易,使得基于语音的服务容易受到欺骗攻击。因此,现在比以往任何时候都迫切需要强有力的反措施。现有的deepfake检测解决方案经常被批评缺乏通用性,并且在应用于真实世界数据时会严重失败。本研究提出一种利用非语义通用音频表示的通用欺骗检测新方法。大量的实验已经进行了使用TRILL和TRILLsson模型找到合适的非语义特征。结果表明,所提出的方法在域内测试集上实现了相当的性能,同时在域外测试集上显著优于最先进的方法。值得注意的是,它在公共领域数据上表现出了卓越的泛化能力,超越了基于手工制作的功能,语义嵌入和端到端架构的方法。


【43】CoComposer: LLM Multi-agent Collaborative Music Composition
标题:联合作曲家:LLM多主体协作音乐作曲
链接:https://arxiv.org/abs/2509.00132

摘要:现有的人工智能音乐创作工具在生成时间、音乐质量和可控性方面受到限制。我们介绍CoComposer,一个多代理系统,由五个合作代理,每个任务的基础上,传统的音乐创作工作流程。使用AudioBox美学系统,我们实验评估CoComposer的四个组成标准。我们使用三个LLM(GPT-4 o,DeepSeek-V3-0324,Gemini-2.5-Flash)进行测试,发现(1)CoComposer在音乐质量方面优于现有的基于多智能体LLM的系统,(2)与单智能体系统相比,在生产复杂性方面。与非LLM MusicLM相比,CoComposer具有更好的可解释性和可编辑性,尽管MusicLM仍然产生更好的音乐。


【44】Algorithms for Collaborative Harmonization
标题:协作协调算法
链接:https://arxiv.org/abs/2509.00120

备注:Presented at the 15th Multidisciplinary Workshop on Advances in Preference Handling M-PREF 2024, Santiago de Compostela, Oct 20, 2024
摘要:我们认为,在音乐和谐领域的文本聚合的一个特定的场景。音乐和声与文本聚合有相似之处,但和声语言比一般文本更有结构性。具体地说,给定一个给定的音乐旋律的一组和声建议,我们的兴趣在于设计聚合算法,产生一个和声序列,满足以下两个关键标准:(1)集体建议的有效表示;(2)音乐上连贯的和声。我们提出了不同的算法聚合的谐波由一组代理商,并分析其复杂性。结果表明,Kemeny和基于复数的算法是最有效的评估代表性和保持音乐的连贯性。


【45】A Survey on Evaluation Metrics for Music Generation
标题:音乐生成评估指标调查
链接:https://arxiv.org/abs/2509.00051

备注:19 pages, 2 figures
摘要:尽管音乐生成系统取得了显著的进步,但由于音乐的复杂性质,用于评估所生成的音乐的方法没有如预期的那样发展,其中包括结构、连贯性、创造性和情感表现力等方面。在本文中,我们揭示了这一研究差距,介绍了一个详细的分类评估指标的音频和符号音乐表示。我们包括一个批判性的审查,确定目前的评估方法,其中包括客观指标和人类感知之间的相关性差,跨文化偏见,缺乏标准化,阻碍跨模型比较的主要局限性。针对这些差距,我们进一步提出了未来的研究方向,建立一个全面的评价框架,音乐生成的评价。


【46】From Sound to Sight: Towards AI-authored Music Videos
标题:从声音到视觉:走向人工智能创作的音乐视频
链接:https://arxiv.org/abs/2509.00029

作者:ovic, Stella Graßhof, Agnes Mercedes Kloft, Ville V. Lehtola, Martin Cunneen, Justyna Starostka, Glenn McGarry, Kun Li, Sami S. Brandt
备注:1st Workshop on Generative AI for Storytelling (AISTORY), 2025
摘要:传统的音乐可视化系统依赖于手工制作的形状和颜色的临时变换,这些变换仅提供有限的表现力。我们提出了两个新的管道,用于使用现成的深度学习模型从任何用户指定的声乐或器乐歌曲自动生成音乐视频。受音乐视频制作人手动工作流程的启发,我们实验了基于潜在特征的技术如何分析音频以检测音乐品质,如情感线索和乐器模式,并使用语言模型将其转换为文本场景描述。接下来,我们采用生成模型来生成相应的视频剪辑。为了评估生成的视频,我们确定了几个关键方面,并设计和进行了初步的用户评估,展示了讲故事的潜力,视觉连贯性和情感与音乐的一致性。我们的研究结果强调了潜在特征技术和深度生成模型在传统方法之外扩展音乐可视化的潜力。


机器翻译由腾讯交互翻译提供,仅供参考