微信公众号:arXiv_Daily
cs.SD语音
【1】Building Audio-Visual Digital Twins with Smartphones
标题:用智能手机打造视听数字双胞胎
链接:https://arxiv.org/pdf/2512.10778v1
备注:Under Mobisys 2026 review, single blind
摘要:今天的数字孪生几乎完全是视觉的,忽略了声学空间现实主义和交互的核心组成部分。我们介绍AV-Twin,这是第一个仅使用商品智能手机构建可编辑视听数字双胞胎的实用系统。AV-Twin结合了移动RIR捕获和视觉辅助声场模型,以有效地重建室内声学。它通过可微分声学渲染进一步恢复每个表面的材料属性,使用户能够修改材料,几何形状和布局,同时自动更新音频和视觉效果。总之,这些功能为现实世界环境中完全可修改的视听数字双胞胎建立了一条实用的道路。摘要:Digital twins today are almost entirely visual, overlooking acoustics-a core component of spatial realism and interaction. We introduce AV-Twin, the first practical system that constructs editable audio-visual digital twins using only commodity smartphones. AV-Twin combines mobile RIR capture and a visual-assisted acoustic field model to efficiently reconstruct room acoustics. It further recovers per-surface material properties through differentiable acoustic rendering, enabling users to modify materials, geometry, and layout while automatically updating both audio and visuals. Together, these capabilities establish a practical path toward fully modifiable audio-visual digital twins for real-world environments.
【2】BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
标题:BRACE:稳健音频字幕质量评估的基准
链接:https://arxiv.org/pdf/2512.10403v1
摘要:自动音频字幕对于音频理解至关重要,可以实现可访问性和内容索引等应用程序。然而,评估音频字幕的质量仍然是一个重大挑战,特别是在没有参考的环境中,高质量的地面实况字幕不可用。虽然CLAPScore是目前使用最广泛的无参考音频字幕评估指标(ACEM),但其在各种条件下的鲁棒性尚未得到系统验证。 为了解决这一差距,我们引入BRACE,一个新的基准,旨在评估音频字幕对齐质量的参考自由设置。BRACE主要用于评估ACEM,也可以扩展到测量大型音频语言模型(LALM)的模态对齐能力。BRACE由两个子基准测试组成:BRACE-Main用于细粒度标题比较,BRACE-Hallucination用于检测微妙的幻觉内容。我们通过高质量的过滤,基于LLM的腐败和人工注释来构建这些数据集。 鉴于CLAPScore作为无参考的ACEM的广泛采用以及LALM在音频语言任务中的应用越来越多,我们使用BRACE基准评估这两种方法,在各种CLAP模型变体中测试CLAPScore并评估多个LALM。 值得注意的是,即使是性能最好的基于CLAP的ACEM在BRACE-Main基准测试中也只能获得70.01的F1分数,而最好的LALM也只能达到63.19。 通过揭示CLAP模型和LALM的局限性,我们的BRACE基准测试为未来的研究方向提供了有价值的见解。摘要:Automatic audio captioning is essential for audio understanding, enabling applications such as accessibility and content indexing. However, evaluating the quality of audio captions remains a major challenge, especially in reference-free settings where high-quality ground-truth captions are unavailable. While CLAPScore is currently the most widely used reference-free Audio Caption Evaluation Metric(ACEM), its robustness under diverse conditions has not been systematically validated. To address this gap, we introduce BRACE, a new benchmark designed to evaluate audio caption alignment quality in a reference-free setting. BRACE is primarily designed for assessing ACEMs, and can also be extended to measure the modality alignment abilities of Large Audio Language Model(LALM). BRACE consists of two sub-benchmarks: BRACE-Main for fine-grained caption comparison and BRACE-Hallucination for detecting subtle hallucinated content. We construct these datasets through high-quality filtering, LLM-based corruption, and human annotation. Given the widespread adoption of CLAPScore as a reference-free ACEM and the increasing application of LALMs in audio-language tasks, we evaluate both approaches using the BRACE benchmark, testing CLAPScore across various CLAP model variants and assessing multiple LALMs. Notably, even the best-performing CLAP-based ACEM achieves only a 70.01 F1-score on the BRACE-Main benchmark, while the best LALM reaches just 63.19. By revealing the limitations of CLAP models and LALMs, our BRACE benchmark offers valuable insights into the direction of future research.
【3】Investigating training objective for flow matching-based speech enhancement
标题:研究基于流匹配的语音增强的训练目标
链接:https://arxiv.org/pdf/2512.10382v1
摘要:语音增强(SE)旨在从嘈杂的录音中恢复干净的语音。虽然生成方法,如分数匹配和薛定谔桥已经显示出很强的有效性,他们往往是计算昂贵的。流匹配通过直接学习将噪声映射到数据的速度场提供了更有效的替代方案。在这项工作中,我们提出了一个系统的研究流量匹配SE下三个训练目标:速度预测,$x_1$预测,和预处理的$x_1$预测。我们分析了它们对训练动态和整体表现的影响。此外,通过引入感知(PESQ)和基于信号(SI-SDR)的目标,我们进一步提高了收敛效率和语音质量,在评估指标上得到了实质性的改善。摘要:Speech enhancement(SE) aims to recover clean speech from noisy recordings. Although generative approaches such as score matching and Schrodinger bridge have shown strong effectiveness, they are often computationally expensive. Flow matching offers a more efficient alternative by directly learning a velocity field that maps noise to data. In this work, we present a systematic study of flow matching for SE under three training objectives: velocity prediction, $x_1$ prediction, and preconditioned $x_1$ prediction. We analyze their impact on training dynamics and overall performance. Moreover, by introducing perceptual(PESQ) and signal-based(SI-SDR) objectives, we further enhance convergence efficiency and speech quality, yielding substantial improvements across evaluation metrics.
【4】Neural personal sound zones with flexible bright zone control
标题:具有灵活明区控制的神经个人音区
链接:https://arxiv.org/pdf/2512.10375v1
摘要:个人声区(PSZ)再现系统是虚拟现实应用中的一项基础技术,它试图通过一个扬声器阵列为不同的听众在同一空间区域内的不同位置创建不同的虚拟声学场景。对于实际应用,重建目标必须在用于记录从扬声器阵列到每个PSZ中的控制点的局部房间脉冲响应(RIR)的相同固定接收器阵列上测量,这使得系统对于现实世界的使用而言不方便且昂贵。本文提出了一种用于PSZ再现的三维卷积神经网络(CNN),该网络以虚拟目标场景为输入,PSZ预滤波器为输出,具有灵活的控制麦克风网格和可选的再现目标。实验结果表明,该方法能够在一次训练中处理灵活控制点网格上的多种再现目标。此外,该方法还展示了从分布在PSZ中的稀疏采样点学习全局空间信息的能力。摘要:Personal sound zone (PSZ) reproduction system, which attempts to create distinct virtual acoustic scenes for different listeners at their respective positions within the same spatial area using one loudspeaker array, is a fundamental technology in the application of virtual reality. For practical applications, the reconstruction targets must be measured on the same fixed receiver array used to record the local room impulse responses (RIRs) from the loudspeaker array to the control points in each PSZ, which makes the system inconvenient and costly for real-world use. In this paper, a 3D convolutional neural network (CNN) designed for PSZ reproduction with flexible control microphone grid and alternative reproduction target is presented, utilizing the virtual target scene as inputs and the PSZ pre-filters as output. Experimental results of the proposed method are compared with the traditional method, demonstrating that the proposed method is able to handle varied reproduction targets on flexible control point grid using only one training session. Furthermore, the proposed method also demonstrates the capability to learn global spatial information from sparse sampling points distributed in PSZs.
【5】MR-FlowDPO: Multi-Reward Direct Preference Optimization for Flow-Matching Text-to-Music Generation
标题:MR-FlowDPO:用于流匹配文本到音乐生成的多奖励直接偏好优化
链接:https://arxiv.org/pdf/2512.10264v1
摘要:音乐生成模型的一个关键挑战是它们缺乏与人类偏好的直接一致性,因为音乐评估本质上是主观的,并且在个体之间存在很大差异。我们介绍MR-FlowDPO,一种新的方法,增强基于流匹配的音乐生成模型-一个主要的现代音乐生成模型,使用直接偏好优化(DPO)与多个音乐奖励。这些奖励旨在通过三个关键维度评估音乐质量:文本对齐、音频制作质量和语义一致性,并利用可扩展的现成模型进行每个奖励预测。我们以两种方式使用这些奖励:(i)通过构建DPO的偏好数据和(ii)通过将奖励集成到文本提示中。为了解决音乐性评价中的模糊性,我们提出了一种新的评分机制,利用语义自监督表示,显着提高了生成的音乐的节奏稳定性。我们使用各种特定于音乐的客观指标以及人类研究进行了广泛的评估。结果表明,MR-FlowDPO显著提高了整体音乐生成质量,并且在音频质量、文本对齐和音乐性方面始终优于竞争激烈的基线。我们的代码可在https: github.com lonzi mrflow_dpo上公开获取;示例在我们的演示页面https: lonzi.github.io mr_flowdpo_demopage 上提供。摘要:A key challenge in music generation models is their lack of direct alignment with human preferences, as music evaluation is inherently subjective and varies widely across individuals. We introduce MR-FlowDPO, a novel approach that enhances flow-matching-based music generation models - a major class of modern music generative models, using Direct Preference Optimization (DPO) with multiple musical rewards. The rewards are crafted to assess music quality across three key dimensions: text alignment, audio production quality, and semantic consistency, utilizing scalable off-the-shelf models for each reward prediction. We employ these rewards in two ways: (i) By constructing preference data for DPO and (ii) by integrating the rewards into text prompting. To address the ambiguity in musicality evaluation, we propose a novel scoring mechanism leveraging semantic self-supervised representations, which significantly improves the rhythmic stability of generated music. We conduct an extensive evaluation using a variety of music-specific objective metrics as well as a human study. Results show that MR-FlowDPO significantly enhances overall music generation quality and is consistently preferred over highly competitive baselines in terms of audio quality, text alignment, and musicality. Our code is publicly available at https: github.com lonzi mrflow_dpo; Samples are provided in our demo page at https: lonzi.github.io mr_flowdpo_demopage .
【6】Semantic-Aware Confidence Calibration for Automated Audio Captioning
标题:自动音频字幕的语义感知置信度校准
链接:https://arxiv.org/pdf/2512.10170v1
备注:5 pages, 2 figures
摘要:自动化音频字幕模型经常产生过度自信的预测,而不管语义准确性,限制了它们在部署中的可靠性。这种缺陷源于两个因素:基于n-gram重叠的评估指标无法捕获语义正确性,以及缺乏校准的置信度估计。我们提出了一个框架,解决了这两个限制,通过集成到音频字幕的置信度预测和重新定义的正确性,通过语义相似性。我们的方法增加了一个基于耳语的音频字幕模型与学习的信心预测头,估计解码器隐藏状态的不确定性。我们采用CLAP音频文本嵌入和句子Transformer相似性(FENSE)来定义语义正确性,从而实现反映真实字幕质量而不是表面级文本重叠的预期校准误差(ECE)计算。Clotho v2上的实验表明,与贪婪解码基线(ECE为0.488)相比,具有语义评估的置信度引导的波束搜索实现了显著改进的校准(基于CLAP的ECE为0.071),同时提高了标准度量的字幕质量。我们的研究结果表明,语义相似性提供了一个更有意义的基础,在音频字幕比传统的n-gram指标的信心校准。摘要:Automated audio captioning models frequently produce overconfident predictions regardless of semantic accuracy, limiting their reliability in deployment. This deficiency stems from two factors: evaluation metrics based on n-gram overlap that fail to capture semantic correctness, and the absence of calibrated confidence estimation. We present a framework that addresses both limitations by integrating confidence prediction into audio captioning and redefining correctness through semantic similarity. Our approach augments a Whisper-based audio captioning model with a learned confidence prediction head that estimates uncertainty from decoder hidden states. We employ CLAP audio-text embeddings and sentence transformer similarities (FENSE) to define semantic correctness, enabling Expected Calibration Error (ECE) computation that reflects true caption quality rather than surface-level text overlap. Experiments on Clotho v2 demonstrate that confidence-guided beam search with semantic evaluation achieves dramatically improved calibration (CLAP-based ECE of 0.071) compared to greedy decoding baselines (ECE of 0.488), while simultaneously improving caption quality across standard metrics. Our results establish that semantic similarity provides a more meaningful foundation for confidence calibration in audio captioning than traditional n-gram metrics.【7】VocSim: A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio
标题:VocSim:单源音频中Zero-Shot内容身份的免训练基准
链接:https://arxiv.org/pdf/2512.10120v1
摘要:通用音频表示旨在将同一事件的声学可变实例映射到附近的点,从而在zero-shot设置中解析内容身份。与通过参数更新来测量适应性的监督分类基准不同,我们引入了VocSim,这是一个探索冻结嵌入的内在几何对齐的免训练基准。VocSim聚合了来自19个语料库的125 k单源剪辑,涵盖人类语音,动物发声和环境声音。通过限制到单源音频,我们隔离的内容表示的混淆源分离。我们使用Precision@k评估嵌入的局部纯度,并使用全局分离率(GSR)进行逐点分类分离。为了校准GSR,我们报告了经验排列基线的升力。在不同的基础模型中,简单的管道,冻结Whisper编码器功能,时间频率池和无标签PCA,产生强大的zero-shot性能。然而,VocSim也揭示了一致的泛化差距。在盲的,低资源的语音,本地检索急剧下降。虽然性能仍然从统计上区分的机会,绝对的几何结构崩溃,表明未能推广到看不见的音位结构。作为外部验证,我们的顶级嵌入预测鸟类感知相似性,改善生物声学分类,并在HEAR基准测试中获得最先进的结果。我们认为,这里测量的内在几何质量在未列出的下游应用中代表实用性。我们发布数据、代码和公共排行榜,以标准化固有音频几何的评估。摘要:General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting. Unlike supervised classification benchmarks that measure adaptability via parameter updates, we introduce VocSim, a training-free benchmark probing the intrinsic geometric alignment of frozen embeddings. VocSim aggregates 125k single-source clips from 19 corpora spanning human speech, animal vocalizations, and environmental sounds. By restricting to single-source audio, we isolate content representation from the confound of source separation. We evaluate embeddings using Precision@k for local purity and the Global Separation Rate (GSR) for point-wise class separation. To calibrate GSR, we report lift over an empirical permutation baseline. Across diverse foundation models, a simple pipeline, frozen Whisper encoder features, time-frequency pooling, and label-free PCA, yields strong zero-shot performance. However, VocSim also uncovers a consistent generalization gap. On blind, low-resource speech, local retrieval drops sharply. While performance remains statistically distinguishable from chance, the absolute geometric structure collapses, indicating a failure to generalize to unseen phonotactics. As external validation, our top embeddings predict avian perceptual similarity, improve bioacoustic classification, and achieve state-of-the-art results on the HEAR benchmark. We posit that the intrinsic geometric quality measured here proxies utility in unlisted downstream applications. We release data, code, and a public leaderboard to standardize the evaluation of intrinsic audio geometry.
【8】Exploring Perceptual Audio Quality Measurement on Stereo Processing Using the Open Dataset of Audio Quality
标题:使用开放音频质量数据集探索立体声处理的感知音频质量测量
链接:https://arxiv.org/pdf/2512.10689v1
备注:Presented at the 159 Audio Engineering Society Convention. Paper Number:366. https:aes2.orgpublicationselibrary-page?id=23040
摘要:ODAQ(音频质量开放数据集)提供了一个全面的框架,用于探索一系列失真类别和信号中的单声道和双耳音频质量退化,以及主观质量评级。最近更新的ODAQ专注于立体声处理方法(如中 侧(MS)和左 右(LR))的影响,为深入研究最先进的客观音频质量指标提供了测试信号和主观评级。我们的评估结果表明,虽然以音色为重点的指标通常在简单的条件下产生稳健的结果,但在具有更复杂的呈现上下文的条件下,它们的预测性能往往会受到影响。我们的研究结果强调了自下而上的心理声学过程和自上而下的上下文因素的相互作用建模的重要性,指导未来的研究模型,更有效地整合音色和感知音频质量的空间维度。摘要:ODAQ (Open Dataset of Audio Quality) provides a comprehensive framework for exploring both monaural and binaural audio quality degradations across a range of distortion classes and signals, accompanied by subjective quality ratings. A recent update of ODAQ, focusing on the impact of stereo processing methods such as Mid Side (MS) and Left Right (LR), provides test signals and subjective ratings for the in-depth investigation of state-of-the-art objective audio quality metrics. Our evaluation results suggest that, while timbre-focused metrics often yield robust results under simpler conditions, their prediction performance tends to suffer under the conditions with a more complex presentation context. Our findings underscore the importance of modeling the interplay of bottom-up psychoacoustic processes and top-down contextual factors, guiding future research toward models that more effectively integrate both timbral and spatial dimensions of perceived audio quality.【1】Exploring Perceptual Audio Quality Measurement on Stereo Processing Using the Open Dataset of Audio Quality
标题:使用开放音频质量数据集探索立体声处理的感知音频质量测量
链接:https://arxiv.org/pdf/2512.10689v1
备注:Presented at the 159 Audio Engineering Society Convention. Paper Number:366. https:aes2.orgpublicationselibrary-page?id=23040
摘要:ODAQ(音频质量开放数据集)提供了一个全面的框架,用于探索一系列失真类别和信号中的单声道和双耳音频质量退化,以及主观质量评级。最近更新的ODAQ专注于立体声处理方法(如中 侧(MS)和左 右(LR))的影响,为深入研究最先进的客观音频质量指标提供了测试信号和主观评级。我们的评估结果表明,虽然以音色为重点的指标通常在简单的条件下产生稳健的结果,但在具有更复杂的呈现上下文的条件下,它们的预测性能往往会受到影响。我们的研究结果强调了自下而上的心理声学过程和自上而下的上下文因素的相互作用建模的重要性,指导未来的研究模型,更有效地整合音色和感知音频质量的空间维度。摘要:ODAQ (Open Dataset of Audio Quality) provides a comprehensive framework for exploring both monaural and binaural audio quality degradations across a range of distortion classes and signals, accompanied by subjective quality ratings. A recent update of ODAQ, focusing on the impact of stereo processing methods such as Mid Side (MS) and Left Right (LR), provides test signals and subjective ratings for the in-depth investigation of state-of-the-art objective audio quality metrics. Our evaluation results suggest that, while timbre-focused metrics often yield robust results under simpler conditions, their prediction performance tends to suffer under the conditions with a more complex presentation context. Our findings underscore the importance of modeling the interplay of bottom-up psychoacoustic processes and top-down contextual factors, guiding future research toward models that more effectively integrate both timbral and spatial dimensions of perceived audio quality.【2】Building Audio-Visual Digital Twins with Smartphones
标题:用智能手机打造视听数字双胞胎
链接:https://arxiv.org/pdf/2512.10778v1
备注:Under Mobisys 2026 review, single blind
摘要:今天的数字孪生几乎完全是视觉的,忽略了声学空间现实主义和交互的核心组成部分。我们介绍AV-Twin,这是第一个仅使用商品智能手机构建可编辑视听数字双胞胎的实用系统。AV-Twin结合了移动RIR捕获和视觉辅助声场模型,以有效地重建室内声学。它通过可微分声学渲染进一步恢复每个表面的材料属性,使用户能够修改材料,几何形状和布局,同时自动更新音频和视觉效果。总之,这些功能为现实世界环境中完全可修改的视听数字双胞胎建立了一条实用的道路。摘要:Digital twins today are almost entirely visual, overlooking acoustics-a core component of spatial realism and interaction. We introduce AV-Twin, the first practical system that constructs editable audio-visual digital twins using only commodity smartphones. AV-Twin combines mobile RIR capture and a visual-assisted acoustic field model to efficiently reconstruct room acoustics. It further recovers per-surface material properties through differentiable acoustic rendering, enabling users to modify materials, geometry, and layout while automatically updating both audio and visuals. Together, these capabilities establish a practical path toward fully modifiable audio-visual digital twins for real-world environments.
机器翻译由腾讯交互翻译提供,仅供参考
