微信公众号:arXiv_Daily
cs.SD语音
【1】DiFlow-TTS: Discrete Flow Matching with Factorized Speech Tokens for Low-Latency Zero-Shot Text-To-Speech
标题:DiFlow-TTC:采用分解语音令牌的离散流匹配,用于低延迟Zero-Shot文本到语音
链接:https://arxiv.org/abs/2509.09631
摘要:Zero-shot文本到语音(TTS)的目标是仅使用短参考样本来合成模仿未见过说话人的声音的高质量语音,这不仅需要说话人自适应,而且需要对韵律属性进行精确建模。基于语言模型、扩散和流匹配的最新方法在zero-shot TTS中显示出有希望的结果,但仍然遭受缓慢的推理和重复伪像。离散编解码器表示已被广泛采用的语音合成,最近的工作已经开始探索扩散模型在纯粹的离散设置,这表明离散生成建模语音合成的潜力。然而,现有的流匹配方法通常将这些离散令牌嵌入到连续空间中并应用连续流匹配,这可能无法充分利用离散表示的优点。为了解决这些挑战,我们引入了DiFlow-TTS,据我们所知,它是第一个探索纯粹的离散流匹配语音合成的模型。DiFlow-TTS在紧凑和统一的架构内显式地对分解的语音属性进行建模。它通过调节文本内容以及从参考语音中提取的韵律和声学属性来利用上下文学习,从而在zero-shot设置中实现有效的属性克隆。此外,该模型采用了一种因子化的流预测机制,具有不同的韵律和声学细节,使其能够学习特定于方面的分布。实验结果表明,DiFlow-TTS在自然度、韵律、说话人风格保持和能量控制等几个关键指标上都取得了令人满意的性能。它还保持了紧凑的模型大小,并实现了低延迟推理,生成语音的速度比最新的现有基线快25.8倍。
摘要:Zero-shot Text-to-Speech (TTS) aims to synthesize high-quality speech that mimics the voice of an unseen speaker using only a short reference sample, requiring not only speaker adaptation but also accurate modeling of prosodic attributes. Recent approaches based on language models, diffusion, and flow matching have shown promising results in zero-shot TTS, but still suffer from slow inference and repetition artifacts. Discrete codec representations have been widely adopted for speech synthesis, and recent works have begun to explore diffusion models in purely discrete settings, suggesting the potential of discrete generative modeling for speech synthesis. However, existing flow-matching methods typically embed these discrete tokens into a continuous space and apply continuous flow matching, which may not fully leverage the advantages of discrete representations. To address these challenges, we introduce DiFlow-TTS, which, to the best of our knowledge, is the first model to explore purely Discrete Flow Matching for speech synthesis. DiFlow-TTS explicitly models factorized speech attributes within a compact and unified architecture. It leverages in-context learning by conditioning on textual content, along with prosodic and acoustic attributes extracted from a reference speech, enabling effective attribute cloning in a zero-shot setting. In addition, the model employs a factorized flow prediction mechanism with distinct heads for prosody and acoustic details, allowing it to learn aspect-specific distributions. Experimental results demonstrate that DiFlow-TTS achieves promising performance in several key metrics, including naturalness, prosody, preservation of speaker style, and energy control. It also maintains a compact model size and achieves low-latency inference, generating speech up to 25.8 times faster than the latest existing baselines.
【2】Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates
标题:有限标量量化实现低比特率冗余和传输鲁棒的神经音频压缩
链接:https://arxiv.org/abs/2509.09550
摘要:神经音频编解码器(NAC)由于其出色的率失真性能以及与作为音频生成的离散特征表示的大型语言模型(LLM)的兼容性,已越来越多地被用于语音处理任务。虽然大多数现有的编解码器依赖于残差矢量量化(RVQ),但有限标量量化(FSQ)最近已成为一种引人注目的替代方案,它简化了训练并原生支持单个码本。我们介绍了NeuCodec,一个基于FSQ的NAC,并表明FSQ编码烘烤的冗余产生的编码,这是强大的,当通过嘈杂的信道传输。首先,通过编码器蒸馏实验,我们表明,两个不同的编码器可以学习编码相同的音频到截然不同的代码序列,同时保持可比的重建质量与相同的量化器和解码器。其次,我们证明了FSQ具有非常优越的比特级扰动鲁棒性,通过比较RVQ和FSQ编解码器的性能时,模拟通过噪声信道的代码序列的传输。
摘要:Neural Audio Codecs (NACs) have become increasingly adopted in speech processing tasks due to their excellent rate-distortion performance and compatibility with Large Language Models (LLMs) as discrete feature representations for audio generation. While most existing codecs rely on Residual Vector Quantization (RVQ), Finite Scalar Quantization (FSQ) has recently emerged as a compelling alternative that simplifies training and natively supports single codebooks. We introduce NeuCodec, an FSQ-based NAC, and show that FSQ encodes baked-in redundancy which produces an encoding which is robust when transmitted through noisy channels. First, through an encoder distillation experiment, we show that two different encoders can learn to encode identical audio into vastly different code sequences whilst maintaining comparable reconstruction quality with the same quantizer and decoder. Second, we demonstrate that FSQ has vastly superior bit-level perturbation robustness by comparing the performance of RVQ and FSQ codecs when simulating the transmission of code sequences through a noisy channel.
【3】Efficient Transformer-Based Piano Transcription With Sparse Attention Mechanisms
标题:具有稀疏注意力机制的高效基于转换器的钢琴抄写
链接:https://arxiv.org/abs/2509.09318
备注:Accepted by APSIPA 2025
摘要:本文研究了自动钢琴转录的基础上,计算效率高,但高性能的变种的Transformer,可以捕获长期的依赖性,在整个音乐作品。最近,基于transformer的序列到序列模型在钢琴转录中表现出优异的性能。然而,由于自注意机制的二次复杂性,这些模型不能一次处理整个作品,因此在实践中通常以滑动窗口的方式处理音乐信号。为了克服这个限制,我们提出了一个有效的架构与稀疏注意机制。具体来说,我们引入滑动窗口的自注意机制的编码器和解码器,以及一个混合的全球-本地交叉注意机制,出席各种跨度根据的token类型。我们还在编码器和解码器之间使用分层池化策略,以进一步减少计算负载。我们在MAESTRO数据集上的实验表明,所提出的模型实现了计算成本和内存使用的显着降低,加快了推理速度,同时保持了与全注意基线相当的转录性能。这允许在相同的硬件上使用更长的音频上下文进行训练,证明了稀疏注意力在构建高效和高性能钢琴转录系统方面的可行性。该代码可在https://github.com/WX-Wei/efficient-seq2seq-piano-trans上获得。
摘要:This paper investigates automatic piano transcription based on computationally-efficient yet high-performant variants of the Transformer that can capture longer-term dependency over the whole musical piece. Recently, transformer-based sequence-to-sequence models have demonstrated excellent performance in piano transcription. These models, however, fail to deal with the whole piece at once due to the quadratic complexity of the self-attention mechanism, and music signals are thus typically processed in a sliding-window manner in practice. To overcome this limitation, we propose an efficient architecture with sparse attention mechanisms. Specifically, we introduce sliding-window self-attention mechanisms for both the encoder and decoder, and a hybrid global-local cross-attention mechanism that attends to various spans according to the MIDI token types. We also use a hierarchical pooling strategy between the encoder and decoder to further reduce computational load. Our experiments on the MAESTRO dataset showed that the proposed model achieved a significant reduction in computational cost and memory usage, accelerating inference speed, while maintaining transcription performance comparable to the full-attention baseline. This allows for training with longer audio contexts on the same hardware, demonstrating the viability of sparse attention for building efficient and high-performance piano transcription systems. The code is available at https://github.com/WX-Wei/efficient-seq2seq-piano-trans.
【4】Adaptive Knowledge Distillation using a Device-Aware Teacher for Low-Complexity Acoustic Scene Classification
标题:使用设备感知教师进行自适应知识提炼以进行低复杂度声学场景分类
链接:https://arxiv.org/abs/2509.09262
摘要:在本技术报告中,我们介绍了我们提交的DCASE 2025挑战赛任务1,低复杂度设备-鲁棒声学场景分类。我们的工作解决了严格的复杂性约束和对可见和不可见设备的鲁棒泛化的双重挑战,同时还利用了允许在测试时使用设备标签的新规则。我们提出的系统是基于一个知识蒸馏框架,一个高效的CP-MobileNet的学生从一个紧凑的,专业的两个老师合奏学习。该集成结合了一个基线PaSST教师,训练标准的交叉熵,和一个“泛化专家”的教师。这个专家是使用我们的新的设备感知特征对齐(DAFA)损失进行训练的,该损失是从以前的工作中改编的,它明确地构建了设备鲁棒性的特征空间。为了利用测试时设备标签的可用性,提取的学生模型然后经历最终的特定于设备的微调阶段。我们提出的系统在开发集上实现了57.93%的最终准确度,比官方基线有了显着的改进,特别是在看不见的设备上。
摘要:In this technical report, we describe our submission for Task 1, Low-Complexity Device-Robust Acoustic Scene Classification, of the DCASE 2025 Challenge. Our work tackles the dual challenges of strict complexity constraints and robust generalization to both seen and unseen devices, while also leveraging the new rule allowing the use of device labels at test time. Our proposed system is based on a knowledge distillation framework where an efficient CP-MobileNet student learns from a compact, specialized two-teacher ensemble. This ensemble combines a baseline PaSST teacher, trained with standard cross-entropy, and a 'generalization expert' teacher. This expert is trained using our novel Device-Aware Feature Alignment (DAFA) loss, adapted from prior work, which explicitly structures the feature space for device robustness. To capitalize on the availability of test-time device labels, the distilled student model then undergoes a final device-specific fine-tuning stage. Our proposed system achieves a final accuracy of 57.93\% on the development set, demonstrating a significant improvement over the official baseline, particularly on unseen devices.
【5】Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems
标题:善意交叉测试揭示音频Deepfake检测系统中的弱点
链接:https://arxiv.org/abs/2509.09204
备注:Published in Interspeech 2025
摘要:音频deepfake检测(ADD)模型通常使用组合多个合成器的数据集进行评估,性能报告为单个等错误率(EER)。然而,这种方法不成比例地用更多的样本加权合成器,代表不足,降低了EER的整体可靠性。此外,大多数ADD数据集缺乏真实语音的多样性,通常具有单一的环境和语音风格(例如,清晰的语音),限制了它们模拟真实世界条件的能力。为了应对这些挑战,我们提出了真正的交叉测试,一个新的评估框架,结合不同的真正的数据集和聚合EER更平衡的评估。与传统的评估方法相比,我们的方法提高了鲁棒性和可解释性。我们对九种真正的语音类型的150多个合成器进行了基准测试,并在https://github.com/cyaaronk/audio_deepfake_eval上发布了一个新的数据集,以促进进一步的研究。
摘要:Audio deepfake detection (ADD) models are commonly evaluated using datasets that combine multiple synthesizers, with performance reported as a single Equal Error Rate (EER). However, this approach disproportionately weights synthesizers with more samples, underrepresenting others and reducing the overall reliability of EER. Additionally, most ADD datasets lack diversity in bona fide speech, often featuring a single environment and speech style (e.g., clean read speech), limiting their ability to simulate real-world conditions. To address these challenges, we propose bona fide cross-testing, a novel evaluation framework that incorporates diverse bona fide datasets and aggregates EERs for more balanced assessments. Our approach improves robustness and interpretability compared to traditional evaluation methods. We benchmark over 150 synthesizers across nine bona fide speech types and release a new dataset to facilitate further research at https://github.com/cyaaronk/audio_deepfake_eval.
【6】DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners
标题:DeCodec:重新思考音频编解码器作为普遍分离的表示学习者
链接:https://arxiv.org/abs/2509.09201
摘要:通用音频编解码器学习跨音频类型的纠缠表示,而一些特定的编解码器提供解耦表示,但仅限于语音。然而,真实世界的音频通常包含混合的语音和背景声音,下游任务需要选择性地访问这些组件。因此,我们重新考虑音频编解码器作为一个通用的解纠缠表示学习器,使不同的音频任务的可控特征选择。为此,我们介绍了DeCodec,一种新型的神经编解码器,它学会将音频表示解耦为专用于语音和背景声音的正交子空间,并且在语音中,表示被进一步分解为语义和非语言成分。这种分层的解纠缠允许灵活的功能选择,使DeCodec成为多个音频应用的通用前端。从技术上讲,建立在编解码器框架上,DeCodec包含了两个关键的创新:一个子空间正交投影模块,将输入分解为两个解耦的正交子空间,以及一个表示交换训练过程,确保这两个子空间分别与语音和背景声音相关。这些允许并行RVQ独立地对语音和背景声音分量进行解码。此外,我们采用语义指导的语音RVQ实现语义和非语言的分解。实验结果表明,DeCodec保持先进的信号重建,同时实现新的功能:卓越的语音增强和有效的一次性语音转换嘈杂的语音通过表示重组,提高ASR的鲁棒性,通过清洁的语义特征,可控的背景声音保留/抑制TTS。演示页面:https://luo404.github.io/DeCodecV2/
摘要:Universal audio codecs learn entangled representations across audio types, whereas some specific codecs offer decoupled representations but are limited to speech. Real-world audio, however, often contains mixed speech and background sounds, and downstream tasks require selective access to these components. Therefore, we rethink the audio codec as a universal disentangled representation learner to enable controllable feature selection across different audio tasks. To this end, we introduce DeCodec, a novel neural codec that learns to decouple audio representations into orthogonal subspaces dedicated to speech and background sound, and within speech, representations are further decomposed into semantic and paralinguistic components. This hierarchical disentanglement allows flexible feature selection, making DeCodec a universal front-end for multiple audio applications. Technically, built upon a codec framework, DeCodec incorporates two key innovations: a subspace orthogonal projection module that factorizes the input into two decoupled orthogonal subspaces, and a representation swap training procedure that ensures these two subspaces are correlate to the speech and background sound, respectively. These allows parallel RVQs to quantize speech and background sound components independently. Furthermore, we employ semantic guidance to the speech RVQ to achieve semantic and paralinguistic decomposition. Experimental results show that DeCodec maintains advanced signal reconstruction while enabling new capabilities: superior speech enhancement and effective one-shot voice conversion on noisy speech via representation recombination, improved ASR robustness through clean semantic features, and controllable background sound preservation/suppression in TTS. Demo Page: https://luo404.github.io/DeCodecV2/
【7】MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
标题:MoLEx:音频深度伪造检测语音自我监督模型中的LoRA专家混合体
链接:https://arxiv.org/abs/2509.09175
摘要:虽然基于自我监督学习(SSL)的模型提高了音频深度伪造检测的准确性,但完全微调它们在计算上是昂贵的。为了解决这个问题,我们提出了一个参数高效的框架,将低秩自适应与混合专家路由器相结合,称为LoRA专家混合(MoLEx)。它保留了SSL模型的预训练知识,同时仅有效地微调选定的专家,降低了培训成本,同时保持了强大的性能。在推理过程中观察到的专家效用表明,路由器会重新激活相同的专家进行类似的攻击,但会切换到其他专家进行新的欺骗,这证实了MoLEx的域感知适应性。MoLEx还提供了域适应的灵活性,允许在不修改整个模型的情况下训练额外的专家。我们主要在ASVSpoof 5数据集上评估我们的方法,并在没有增强的情况下在评估集上实现了5.56%的最先进(SOTA)等错误率(EER)。
摘要:While self-supervised learning (SSL)-based models have boosted audio deepfake detection accuracy, fully finetuning them is computationally expensive. To address this, we propose a parameter-efficient framework that combines Low-Rank Adaptation with a Mixture-of-Experts router, called Mixture of LoRA Experts (MoLEx). It preserves pre-trained knowledge of SSL models while efficiently finetuning only selected experts, reducing training costs while maintaining robust performance. The observed utility of experts during inference shows the router reactivates the same experts for similar attacks but switches to other experts for novel spoofs, confirming MoLEx's domain-aware adaptability. MoLEx additionally offers flexibility for domain adaptation by allowing extra experts to be trained without modifying the entire model. We mainly evaluate our approach on the ASVSpoof 5 dataset and achieve the state-of-the-art (SOTA) equal error rate (EER) of 5.56% on the evaluation set without augmentation.
【8】EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
标题:EchoX:通过语音对语音LLM的Echo训练来缓解声学-语义差距
链接:https://arxiv.org/abs/2509.09174
摘要:语音到语音的大语言模型(SLLM)越来越受到人们的关注。从基于文本的大型语言模型(LLM)派生,SLLM往往表现出知识和推理能力的退化。我们假设出现这种限制是因为目前的SLLM训练范式未能弥合特征表示空间中的声学-语义差距。为了解决这个问题,我们提出了EchoX,它利用语义表示并动态生成语音训练目标。这种方法集成了声学和语义学习,使EchoX能够保持作为语音LLM的强大推理能力。实验结果表明,EchoX,约6000小时的训练数据,在多个基于知识的问答基准测试中实现了先进的性能。该项目可在https://github.com/FreedomIntelligence/EchoX上查阅。
摘要:Speech-to-speech large language models (SLLMs) are attracting increasing attention. Derived from text-based large language models (LLMs), SLLMs often exhibit degradation in knowledge and reasoning capabilities. We hypothesize that this limitation arises because current training paradigms for SLLMs fail to bridge the acoustic-semantic gap in the feature representation space. To address this issue, we propose EchoX, which leverages semantic representations and dynamically generates speech training targets. This approach integrates both acoustic and semantic learning, enabling EchoX to preserve strong reasoning abilities as a speech LLM. Experimental results demonstrate that EchoX, with about six thousand hours of training data, achieves advanced performance on multiple knowledge-based question-answering benchmarks. The project is available at https://github.com/FreedomIntelligence/EchoX.
【9】In situ estimation of the acoustic surface impedance using simulation-based inference
标题:使用基于模拟的推理现场估计声表面阻抗
链接:https://arxiv.org/abs/2509.08873
摘要:封闭空间的精确声学模拟需要精确的边界条件,通常通过基于波的方法的表面阻抗来表示。传统的测量技术往往依赖于简化假设的声场和安装条件,限制其有效性的现实世界的情况。为了克服这些局限性,本研究介绍了一个贝叶斯框架的频率依赖性声表面阻抗稀疏的内部声压测量的原位估计。该方法采用基于模拟的推理,利用现代神经网络架构的表现力,直接将模拟数据映射到模型参数的后验分布,绕过传统的基于采样的贝叶斯方法,并为高维推理问题提供优势。阻抗行为是使用一个阻尼振荡器模型扩展了分数阶微积分项。该框架在长方体房间的有限元模型上进行了验证,并使用阻抗管测量值作为参考进行了进一步测试,实现了所有六个个体阻抗的鲁棒性和准确性估计。应用于汽车驾驶室数值模型进一步证明了可靠的不确定性量化和高预测精度,即使是复杂形状的几何形状。后验预测检查和覆盖诊断确认良好校准的推断,突出了该方法的潜力,在现实世界的内部环境中的声学边界条件的概括,高效,物理上一致的表征。
摘要:Accurate acoustic simulations of enclosed spaces require precise boundary conditions, typically expressed through surface impedances for wave-based methods. Conventional measurement techniques often rely on simplifying assumptions about the sound field and mounting conditions, limiting their validity for real-world scenarios. To overcome these limitations, this study introduces a Bayesian framework for the in situ estimation of frequency-dependent acoustic surface impedances from sparse interior sound pressure measurements. The approach employs simulation-based inference, which leverages the expressiveness of modern neural network architectures to directly map simulated data to posterior distributions of model parameters, bypassing conventional sampling-based Bayesian approaches and offering advantages for high-dimensional inference problems. Impedance behavior is modeled using a damped oscillator model extended with a fractional calculus term. The framework is verified on a finite element model of a cuboid room and further tested with impedance tube measurements used as reference, achieving robust and accurate estimation of all six individual impedances. Application to a numerical car cabin model further demonstrates reliable uncertainty quantification and high predictive accuracy even for complex-shaped geometries. Posterior predictive checks and coverage diagnostics confirm well-calibrated inference, highlighting the method's potential for generalizable, efficient, and physically consistent characterization of acoustic boundary conditions in real-world interior environments.
【10】Region-Specific Audio Tagging for Spatial Sound
标题:空间声音的特定区域音频标记
链接:https://arxiv.org/abs/2509.09526
备注:DCASE2025 Workshop
摘要:音频标记旨在标记音频记录中出现的声音事件。在本文中,我们提出了特定区域的音频标记,一个新的任务,标签的声音事件在一个给定的区域的空间音频记录的麦克风阵列。该区域可以指定为角空间或距麦克风的距离。我们首先研究了光谱,空间和位置特征的不同组合的性能。然后,我们扩展国家的最先进的音频标记系统,如预训练的音频神经网络(PANN)和音频频谱图Transformer(AST)的建议区域特定的音频标记任务。在模拟数据集和真实数据集上的实验结果表明了所提任务的可行性和所提方法的有效性。进一步的实验表明,结合方向特征是有益的全向标注。
摘要:Audio tagging aims to label sound events appearing in an audio recording. In this paper, we propose region-specific audio tagging, a new task which labels sound events in a given region for spatial audio recorded by a microphone array. The region can be specified as an angular space or a distance from the microphone. We first study the performance of different combinations of spectral, spatial, and position features. Then we extend state-of-the-art audio tagging systems such as pre-trained audio neural networks (PANNs) and audio spectrogram transformer (AST) to the proposed region-specific audio tagging task. Experimental results on both the simulated and the real datasets show the feasibility of the proposed task and the effectiveness of the proposed method. Further experiments show that incorporating the directional features is beneficial for omnidirectional tagging.
【11】Short-term cognitive fatigue of spatial selective attention after face-to-face conversations in virtual noisy environments
标题:虚拟嘈杂环境中面对面对话后空间选择性注意的短期认知疲劳
链接:https://arxiv.org/abs/2509.09479
摘要:空间选择性注意是鸡尾酒会场合中交流的重要资产,但可能会受到短期认知疲劳的影响。在这里,我们测试了在一个高度生态化的环境中,一个努力的对话是否会耗尽听觉空间选择性注意任务中的任务表现。听力正常的年轻参与者在(1)在虚拟混响室中进行关于自由话题的真实二元面对面对话,模拟干扰对话和背景噪声为72 dB SPL,持续30分钟,(2)被动地听干扰对话和背景噪声,或(3)安静地进行对话之前和之后执行任务。自我报告的感知努力和疲劳增加后,在噪音和被动倾听相对于报告后,在安静的对话。与我们的预期相反,在注意力任务的反应时间减少,而不是增加,在噪音和准确性的谈话后,没有系统地改变在任何条件下的组水平。出乎意料的是,我们在受试者内设计中观察到了强烈的训练效果,即使是在不同的一天训练一小时后。
摘要:Spatial selective attention is an important asset for communication in cocktail party situations but may be compromised by short-term cognitive fatigue. Here we tested whether an effortful conversation in a highly ecological setting depletes task performance in an auditory spatial selective attention task. Young participants with normal hearing performed the task before and after (1) having a real dyadic face-to-face conversation on a free topic in a virtual reverberant room with simulated interfering conversations and background babble noise at 72 dB SPL for 30 minutes, (2) passively listening to the interfering conversations and babble noise, or (3) having the conversation in quiet. Self-reported perceived effort and fatigue increased after conversations in noise and passive listening relative to the reports after conversations in quiet. In contrast to our expectations, response times in the attention task decreased, rather than increased, after conversation in noise and accuracy did not change systematically in any of the conditions on the group level. Unexpectedly, we observed strong training effects between the individual sessions in our within-subject design even after one hour of training on a different day.
【12】MAPSS: Manifold-based Assessment of Perceptual Source Separation
标题:MACSS:基于Manifold的感知源分离评估
链接:https://arxiv.org/abs/2509.09212
备注:Submitted to ICLR
摘要:源分离系统的客观评估仍然与人类的主观感知不匹配,特别是当泄漏和自失真相互作用时。我们引入了感知分离(PS)和感知匹配(PM),这是第一对在功能上隔离这两个因素的措施。我们的侵入式方法首先为混合物中的每个参考波形信号生成一组基本失真。然后,来自所有源的失真、参考和它们各自的系统输出由预先训练的自监督学习模型独立编码。这些表示通过扩散映射被聚合并投影到流形上,扩散映射将流形上的欧几里得距离与编码波形的不同之处对齐。在这个流形上,PM测量从每个输出到其属性聚类的Mahalanobis距离,该属性聚类由其参考和失真嵌入组成,捕获自失真。PS说明了输出到属性聚类和最接近的非属性聚类的Mahalanobis距离,量化了泄漏。这两种测量都是可区分的和颗粒状的,以低至每秒50帧的分辨率操作。我们进一步推导出,这两种措施,确定性误差半径和非渐近,高概率置信区间(CI)。英语,西班牙语和音乐的混合实验表明,PS和PM几乎总是达到最高的线性相关系数与人类的平均意见分数比14个竞争对手,高达86.36%的语音和87.21%的音乐。我们观察到,在最坏的情况下,这些系数的误差半径为1.39%,概率95%CI为12.21%,这提高了可靠和知情的评估。使用互信息,这些措施相互补充,因为它们的值减少,这表明它们共同提供更多的信息,因为系统性能下降。
摘要:Objective assessment of source-separation systems still mismatches subjective human perception, especially when leakage and self-distortion interact. We introduce the Perceptual Separation (PS) and Perceptual Match (PM), the first pair of measures that functionally isolate these two factors. Our intrusive method begins with generating a bank of fundamental distortions for each reference waveform signal in the mixture. Distortions, references, and their respective system outputs from all sources are then independently encoded by a pre-trained self-supervised learning model. These representations are aggregated and projected onto a manifold via diffusion maps, which aligns Euclidean distances on the manifold with dissimilarities of the encoded waveforms. On this manifold, the PM measures the Mahalanobis distance from each output to its attributed cluster that consists of its reference and distortions embeddings, capturing self-distortion. The PS accounts for the Mahalanobis distance of the output to the attributed and to the closest non-attributed clusters, quantifying leakage. Both measures are differentiable and granular, operating at a resolution as low as 50 frames per second. We further derive, for both measures, deterministic error radius and non-asymptotic, high-probability confidence intervals (CIs). Experiments on English, Spanish, and music mixtures show that the PS and PM nearly always achieve the highest linear correlation coefficients with human mean-opinion scores than 14 competitors, reaching as high as 86.36% for speech and 87.21% for music. We observe, at worst, an error radius of 1.39% and a probabilistic 95% CI of 12.21% for these coefficients, which improves reliable and informed evaluation. Using mutual information, the measures complement each other most as their values decrease, suggesting they are jointly more informative as system performance degrades.
【13】The Sound of Entanglement
标题:纠缠的声音
链接:https://arxiv.org/abs/2509.08892
备注:13 pages, 12 figures
摘要:量子物理学的出现彻底改变了我们对宇宙的理解,用内在随机性和量子相关性主导的范式取代了经典物理学的确定性框架。这一转变不仅使量子传感器、网络和计算机等突破性技术成为可能,而且还为艺术表达开启了全新的可能性。在本文中,我们探讨了量子力学和艺术的交叉点,重点是使用量子纠缠和固有的随机性作为创作工具。具体来说,我们提出了纠缠的声音,在贝尔测试中纠缠光子的实时测量驱动的现场音乐表演。通过将测量到的量子相关性整合为中心组成元素,并将现场视觉效果与实验数据同步,表演提供了一种独特且不可重复的视听体验,这种体验依赖于任何经典设备都无法产生的量子相关性。通过这种科学与艺术的融合,我们的目标是提供量子现象的更深层次的欣赏,同时扩大创造性表达的界限。
摘要:The advent of quantum physics has revolutionized our understanding of the universe, replacing the deterministic framework of classical physics with a paradigm dominated by intrinsic randomness and quantum correlations. This shift has not only enabled groundbreaking technologies, such as quantum sensors, networks and computers, but has also unlocked entirely new possibilities for artistic expressions. In this paper, we explore the intersection of quantum mechanics and art, focusing on the use of quantum entanglement and inherent randomness as creative tools. Specifically, we present The Sound of Entanglement, a live musical performance driven by real-time measurements of entangled photons in a Bell test. By integrating the measured quantum correlations as a central compositional element and synchronizing live visuals with experimental data, the performance offers a unique and unrepeatable audiovisual experience that relies on quantum correlations which cannot be produced by any classical device. Through this fusion of science and art, we aim to provide a deeper appreciation of quantum phenomena while expanding the boundaries of creative expression.
【1】Region-Specific Audio Tagging for Spatial Sound
标题:空间声音的特定区域音频标记
链接:https://arxiv.org/abs/2509.09526
备注:DCASE2025 Workshop
摘要:音频标记旨在标记音频记录中出现的声音事件。在本文中,我们提出了特定区域的音频标记,一个新的任务,标签的声音事件在一个给定的区域的空间音频记录的麦克风阵列。该区域可以被指定为角度空间或距麦克风的距离。我们首先研究了光谱,空间和位置特征的不同组合的性能。然后,我们扩展国家的最先进的音频标记系统,如预训练的音频神经网络(PANN)和音频频谱图Transformer(AST)的建议区域特定的音频标记任务。在模拟数据集和真实数据集上的实验结果表明了所提任务的可行性和所提方法的有效性。进一步的实验表明,结合方向特征是有益的全向标注。
摘要:Audio tagging aims to label sound events appearing in an audio recording. In this paper, we propose region-specific audio tagging, a new task which labels sound events in a given region for spatial audio recorded by a microphone array. The region can be specified as an angular space or a distance from the microphone. We first study the performance of different combinations of spectral, spatial, and position features. Then we extend state-of-the-art audio tagging systems such as pre-trained audio neural networks (PANNs) and audio spectrogram transformer (AST) to the proposed region-specific audio tagging task. Experimental results on both the simulated and the real datasets show the feasibility of the proposed task and the effectiveness of the proposed method. Further experiments show that incorporating the directional features is beneficial for omnidirectional tagging.
【2】Acoustic to Articulatory Speech Inversion for Children with Velopharyngeal Insufficiency
标题:喉咽功能不全儿童的声学-关节言语倒置
链接:https://arxiv.org/abs/2509.09489
备注:Accepted to be presented at ASRU workshop 2025
摘要:用于评估鼻音的传统临床方法,例如鼻咽镜检查和鼻测量,涉及不愉快的经历并且对于儿童是有问题的。语音反转(SI),一种非侵入性的技术,提供了一个有前途的替代估计发音运动,而不需要物理仪器。在这项研究中,一个SI系统训练的nasalance数据从健康成人增强源信息电声门图和声学衍生的F0,周期性和非周期性的能量估计作为代理声门控制。该模型实现了16.92%的皮尔逊积矩相关(PPMC)的相对改善相比,以前的SI系统nasalance估计。为了使SI系统适应腭咽闭合不全(VPI)儿童的鼻音估计,使用VPI数据的儿童对最初在成人语音上训练的模型进行微调,与微调前的性能相比,PPMC相对改善了7.90%。
摘要:Traditional clinical approaches for assessing nasality, such as nasopharyngoscopy and nasometry, involve unpleasant experiences and are problematic for children. Speech Inversion (SI), a noninvasive technique, offers a promising alternative for estimating articulatory movement without the need for physical instrumentation. In this study, an SI system trained on nasalance data from healthy adults is augmented with source information from electroglottography and acoustically derived F0, periodic and aperiodic energy estimates as proxies for glottal control. This model achieves 16.92% relative improvement in Pearson Product-Moment Correlation (PPMC) compared to a previous SI system for nasalance estimation. To adapt the SI system for nasalance estimation in children with Velopharyngeal Insufficiency (VPI), the model initially trained on adult speech was fine-tuned using children with VPI data, yielding an 7.90% relative improvement in PPMC compared to its performance before fine-tuning.
【3】Short-term cognitive fatigue of spatial selective attention after face-to-face conversations in virtual noisy environments
标题:虚拟嘈杂环境中面对面对话后空间选择性注意的短期认知疲劳
链接:https://arxiv.org/abs/2509.09479
摘要:空间选择性注意力是鸡尾酒会场合沟通的重要资产,但可能会受到短期认知疲劳的影响。在这里,我们测试了在一个高度生态化的环境中,一个努力的对话是否会耗尽听觉空间选择性注意任务中的任务表现。听力正常的年轻参与者在(1)在虚拟混响室中进行关于自由话题的真实二元面对面对话,模拟干扰对话和背景噪声为72 dB SPL,持续30分钟,(2)被动地听干扰对话和背景噪声,或(3)安静地进行对话之前和之后执行任务。自我报告的感知努力和疲劳增加后,在噪音和被动倾听相对于报告后,在安静的对话。与我们的预期相反,在注意力任务的反应时间减少,而不是增加,在噪音和准确性的谈话后,没有系统地改变在任何条件下的组水平。出乎意料的是,我们在受试者内设计中观察到了强烈的训练效果,即使是在不同的一天训练一小时后。
摘要:Spatial selective attention is an important asset for communication in cocktail party situations but may be compromised by short-term cognitive fatigue. Here we tested whether an effortful conversation in a highly ecological setting depletes task performance in an auditory spatial selective attention task. Young participants with normal hearing performed the task before and after (1) having a real dyadic face-to-face conversation on a free topic in a virtual reverberant room with simulated interfering conversations and background babble noise at 72 dB SPL for 30 minutes, (2) passively listening to the interfering conversations and babble noise, or (3) having the conversation in quiet. Self-reported perceived effort and fatigue increased after conversations in noise and passive listening relative to the reports after conversations in quiet. In contrast to our expectations, response times in the attention task decreased, rather than increased, after conversation in noise and accuracy did not change systematically in any of the conditions on the group level. Unexpectedly, we observed strong training effects between the individual sessions in our within-subject design even after one hour of training on a different day.
【4】Listening for "You": Enhancing Speech Image Retrieval via Target Speaker Extraction
链接:https://arxiv.org/abs/2509.09306
备注:5 pages, 2 figures
摘要:使用口语提示的图像检索已经成为多模态感知中一个有前途的方向,但在多说话人场景中利用语音仍然具有挑战性。我们提出了一种新的目标说话人语音图像检索任务和一个框架,学习图像和多说话人语音信号之间的关系,在目标说话人的存在。我们的方法通过目标说话人感知对比学习将预训练的自监督音频编码器与视觉模型集成在一起,条件是目标说话人提取和检索模块。这使得系统能够从目标说话者提取口头命令,并将它们与相应的图像对齐。在SpokenCOCO 2 Mix和SpokenCOCO 3 Mix上的实验表明,TSRE显著优于现有方法,在2和3个扬声器场景中分别实现36.3%和29.9%的Recall@1-比单个扬声器基线和最先进的模型有了实质性的改进。我们的方法展示了在辅助机器人和多模式交互系统中的现实部署潜力。
摘要:Image retrieval using spoken language cues has emerged as a promising direction in multimodal perception, yet leveraging speech in multi-speaker scenarios remains challenging. We propose a novel Target Speaker Speech-Image Retrieval task and a framework that learns the relationship between images and multi-speaker speech signals in the presence of a target speaker. Our method integrates pre-trained self-supervised audio encoders with vision models via target speaker-aware contrastive learning, conditioned on a Target Speaker Extraction and Retrieval module. This enables the system to extract spoken commands from the target speaker and align them with corresponding images. Experiments on SpokenCOCO2Mix and SpokenCOCO3Mix show that TSRE significantly outperforms existing methods, achieving 36.3% and 29.9% Recall@1 in 2 and 3 speaker scenarios, respectively - substantial improvements over single speaker baselines and state-of-the-art models. Our approach demonstrates potential for real-world deployment in assistive robotics and multimodal interaction systems.
【5】Over-the-Air Adversarial Attack Detection: from Datasets to Defenses
标题:空中对抗攻击检测:从数据集到防御
链接:https://arxiv.org/abs/2509.09296
摘要:自动说话者验证(ASV)系统可用于支持语音的应用程序以进行身份验证。然而,最近的研究已经暴露了这些系统对线上(OTL)和空中(OTA)对抗性攻击的脆弱性。虽然已经提出了各种检测方法来应对这些威胁,但由于缺乏全面的数据集,这些方法尚未得到彻底的测试。为了解决这一差距,我们开发了AdvSV 2.0数据集,其中包含628 k个样本,总持续时间为800小时。该数据集融合了经典的对抗性攻击算法、ASV系统,并涵盖了OTL和OTA场景。此外,我们引入了一种新的基于神经重放模拟器(NRS)的对抗性攻击方法,该方法增强了对抗性OTA攻击的效力,从而对ASV系统构成了更大的威胁。为了抵御这些攻击,我们提出了CODA-OCC,一类分类框架内的对比学习方法。实验结果表明,CODA-OCC在AdvSV 2.0数据集上实现了11.2%的EER和0.95的AUC,优于几种最先进的检测方法。
摘要:Automatic Speaker Verification (ASV) systems can be used for voice-enabled applications for identity verification. However, recent studies have exposed these systems' vulnerabilities to both over-the-line (OTL) and over-the-air (OTA) adversarial attacks. Although various detection methods have been proposed to counter these threats, they have not been thoroughly tested due to the lack of a comprehensive data set. To address this gap, we developed the AdvSV 2.0 dataset, which contains 628k samples with a total duration of 800 hours. This dataset incorporates classical adversarial attack algorithms, ASV systems, and encompasses both OTL and OTA scenarios. Furthermore, we introduce a novel adversarial attack method based on a Neural Replay Simulator (NRS), which enhances the potency of adversarial OTA attacks, thereby presenting a greater threat to ASV systems. To defend against these attacks, we propose CODA-OCC, a contrastive learning approach within the one-class classification framework. Experimental results show that CODA-OCC achieves an EER of 11.2% and an AUC of 0.95 on the AdvSV 2.0 dataset, outperforming several state-of-the-art detection methods.
【6】MAPSS: Manifold-based Assessment of Perceptual Source Separation
标题:MACSS:基于Manifold的感知源分离评估
链接:https://arxiv.org/abs/2509.09212
备注:Submitted to ICLR
摘要:源分离系统的客观评估仍然与人类的主观感知不匹配,特别是当泄漏和自失真相互作用时。我们引入了感知分离(PS)和感知匹配(PM),这是第一对在功能上隔离这两个因素的措施。我们的侵入式方法首先为混合物中的每个参考波形信号生成一组基本失真。然后,来自所有源的失真、参考和它们各自的系统输出由预先训练的自监督学习模型独立编码。这些表示通过扩散映射被聚合并投影到流形上,扩散映射将流形上的欧几里得距离与编码波形的不同之处对齐。在这个流形上,PM测量从每个输出到其属性聚类的Mahalanobis距离,该属性聚类由其参考和失真嵌入组成,捕获自失真。PS说明了输出到属性聚类和最接近的非属性聚类的Mahalanobis距离,量化了泄漏。这两种测量都是可区分的和颗粒状的,以低至每秒50帧的分辨率操作。我们进一步推导出,这两种措施,确定性误差半径和非渐近,高概率置信区间(CI)。英语,西班牙语和音乐的混合实验表明,PS和PM几乎总是达到最高的线性相关系数与人类的平均意见分数比14个竞争对手,高达86.36%的语音和87.21%的音乐。我们观察到,在最坏的情况下,这些系数的误差半径为1.39%,概率95%CI为12.21%,这提高了可靠和知情的评估。使用互信息,这些措施相互补充,因为它们的值减少,这表明它们共同提供更多的信息,因为系统性能下降。
摘要:Objective assessment of source-separation systems still mismatches subjective human perception, especially when leakage and self-distortion interact. We introduce the Perceptual Separation (PS) and Perceptual Match (PM), the first pair of measures that functionally isolate these two factors. Our intrusive method begins with generating a bank of fundamental distortions for each reference waveform signal in the mixture. Distortions, references, and their respective system outputs from all sources are then independently encoded by a pre-trained self-supervised learning model. These representations are aggregated and projected onto a manifold via diffusion maps, which aligns Euclidean distances on the manifold with dissimilarities of the encoded waveforms. On this manifold, the PM measures the Mahalanobis distance from each output to its attributed cluster that consists of its reference and distortions embeddings, capturing self-distortion. The PS accounts for the Mahalanobis distance of the output to the attributed and to the closest non-attributed clusters, quantifying leakage. Both measures are differentiable and granular, operating at a resolution as low as 50 frames per second. We further derive, for both measures, deterministic error radius and non-asymptotic, high-probability confidence intervals (CIs). Experiments on English, Spanish, and music mixtures show that the PS and PM nearly always achieve the highest linear correlation coefficients with human mean-opinion scores than 14 competitors, reaching as high as 86.36% for speech and 87.21% for music. We observe, at worst, an error radius of 1.39% and a probabilistic 95% CI of 12.21% for these coefficients, which improves reliable and informed evaluation. Using mutual information, the measures complement each other most as their values decrease, suggesting they are jointly more informative as system performance degrades.
【7】Automotive sound field reproduction using deep optimization with spatial domain constraint
标题:使用具有空间域约束的深度优化的汽车音场再现
链接:https://arxiv.org/abs/2509.09149
备注:41 pages, 9 figures, Revised and submitted to The Journal of the Acoustical Society of America (JASA)
摘要:具有不失真的声音质量和精确的空间定位的声场再现对于汽车音频系统是期望的。然而,汽车舱内声学环境的复杂性往往需要在声音质量和空间精度之间进行权衡。为了克服这一限制,我们提出了空间功率映射网络(SPMnet),这是一种基于学习的声场再现方法,可以在复杂环境中提高音质和空间定位。我们引入了一个空间功率图(SPM)的约束,其特征在于使用波束成形的再现场的角能量分布。该约束将能量引导向预期方向以增强空间定位,并且被集成到多声道均衡框架中以还改善混响条件下的声音质量。为了解决由此产生的非凸性,使用神经网络来解决优化问题的深度优化被用于滤波器设计。现场客观和主观评估都证实,我们的方法提高了声音质量,提高了汽车车厢内的空间定位。此外,我们分析了不同的音频材料和虚拟声源的到达角在再现声场的影响,调查潜在的潜在因素影响这些结果。
摘要:Sound field reproduction with undistorted sound quality and precise spatial localization is desirable for automotive audio systems. However, the complexity of automotive cabin acoustic environment often necessitates a trade-off between sound quality and spatial accuracy. To overcome this limitation, we propose Spatial Power Map Net (SPMnet), a learning-based sound field reproduction method that improves both sound quality and spatial localization in complex environments. We introduce a spatial power map (SPM) constraint, which characterizes the angular energy distribution of the reproduced field using beamforming. This constraint guides energy toward the intended direction to enhance spatial localization, and is integrated into a multi-channel equalization framework to also improve sound quality under reverberant conditions. To address the resulting non-convexity, deep optimization that use neural networks to solve optimization problems is employed for filter design. Both in situ objective and subjective evaluations confirm that our method enhances sound quality and improves spatial localization within the automotive cabin. Furthermore, we analyze the influence of different audio materials and the arrival angles of the virtual sound source in the reproduced sound field, investigating the potential underlying factors affecting these results.
机器翻译由腾讯交互翻译提供,仅供参考
