微信公众号:arXiv_Daily
cs.SD语音
【1】Multi-layer attentive probing improves transfer of audio representations for bioacoustics
标题:多层专注探测改善生物声学音频表示的传输
链接:https://arxiv.org/abs/2605.10494
摘要:
摘要:
【2】Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs
标题:通过神经音频编解码器从表达式鼓网格中进行鼓合成
链接:https://arxiv.org/abs/2605.10281
摘要:
摘要:
【3】A Cold Diffusion Approach for Percussive Dereverberation
标题:冲击波去回响的冷扩散方法
链接:https://arxiv.org/abs/2605.10256
备注:Accepted for the 2026 IEEE World Congress on Computational Intelligence, IJCNN Track, 21-26 June 2026, Maastricht, the Netherlands
摘要:
摘要:
【4】Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration
标题:Polyphonia:具有声学注意力校准的复音音乐中的Zero-Shot音色传输
链接:https://arxiv.org/abs/2605.10203
备注:Accepted by ICML 2026
摘要:
摘要:
【5】APEX: Audio Prototype EXplanations for Classification Tasks
标题:APEX:分类任务的音频原型解释
链接:https://arxiv.org/abs/2605.10153
摘要:
摘要:
【6】Voice Biomarkers for Depression and Anxiety
标题:抑郁和焦虑的声音生物标志物
链接:https://arxiv.org/abs/2605.09908
摘要:目前从语音中检测抑郁和焦虑的方法主要依赖于机器学习技术,这些技术利用手工设计的语言学特征和从语音信号的时域和频域表示中导出的相关声学描述符。将深度学习方法直接应用于原始语音信号有可能产生具有更大预测能力的生物标志物表示。然而,这些方法通常需要大量仔细注释的数据来学习基础生物标志物的稳健且有临床意义的表示。在本文中,我们描述了我们在开发一个深度学习模型方面所做的努力,该模型是在一个大规模的专有数据集上训练的,该数据集包括从代表美国相关人口统计数据的23,000多个受试者中收集的约65,000个话语。我们提出的技术,并分析其对模型性能的影响。我们的研究结果表明,所提出的模型可以提取内容不可知的生物标志物信息,当与从音频中提取的词汇特征相结合时,可以提高生产环境中的预测性能。我们的模型在约5000名独特的受试者上进行了评估,在灵敏度和特异性方面达到了71%的性能。为了促进对语音心理健康评估的进一步研究,我们在HuggingFace上发布了本文中描述的性能最佳的模型。
摘要:Current approaches to detecting depression and anxiety from speech primarily rely on machine learning techniques that utilize hand-engineered paralinguistic features and related acoustic descriptors derived from time- and frequency-domain representations of speech signals. Applying deep learning methods directly to raw speech signals has the potential to produce biomarker representations with substantially greater predictive power. However, these approaches typically require large volumes of carefully annotated data to learn robust and clinically meaningful representations of the underlying biomarkers. In this paper, we describe our efforts toward developing a deep learning model trained on a large-scale proprietary dataset comprising ~65,000 utterances collected from more than 23,000 subjects representative of relevant United States demographics. We present the techniques employed and analyze their impact on model performance. Our results demonstrate that the proposed models can extract content-agnostic biomarker information, which, when combined with lexical features extracted from audio, yields improved predictive performance in production settings. Our models are evaluated on ~5000 unique subjects and achieve performance of 71% in terms of sensitivity and specificity. To foster further research in mental health assessment from speech, we release the best-performing model described in this paper on HuggingFace.
【7】Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought
标题:先分离,后分离:通过特定模式思维链缓解视听LLM推理中的跨模式干扰
链接:https://arxiv.org/abs/2605.09906
摘要:音频和视觉为视听问题回答提供了补充证据,但当前的视听大语言模型可能会受到跨模态干扰:来自一种模态的信息误导了另一种模态的解释,从而引起幻觉。我们把这个问题归因于中间推理过程中不受控制的跨模态交互。为了减轻这一点,我们提出了分开的第一,稍后(SFFL),视听推理框架,旨在减少跨模态干扰。SFFL执行特定于模态的思维链推理,产生单独的音频和视觉推理痕迹,并整合证据进行回答。我们通过不同的模态输入设置下的数据管道构建模态偏好标签。我们使用这些标签作为强化学习中的辅助奖励,以鼓励在回答时对模态线索的实例依赖偏好。我们进一步引入了一个特定模态的推理机制,保留模态隔离在分离的推理阶段,同时使跨模态的信息在证据融合阶段的全面访问。实验结果表明,在准确性和鲁棒性的一致改善,产生的平均相对增益为5.16%的一般AVQA基准和11.17%的跨模态幻觉基准。
摘要:Audio and vision provide complementary evidence for audio-visual question answering, yet current audio-visual large language models may suffer from cross-modal interference: information from one modality misguides the interpretation of another, thereby inducing hallucinations. We attribute this issue to uncontrolled cross-modal interactions during intermediate reasoning. To mitigate this, we propose Separate First, Fuse Later (SFFL), an audio-visual reasoning framework designed to reduce cross-modal interference. SFFL enforces modality-specific chain-of-thought reasoning, producing separate audio and visual reasoning traces and integrating evidence for answering. We construct modality-preference labels via a data pipeline under different modality input settings. We use these labels as an auxiliary reward in reinforcement learning to encourage a instance-dependent preference for modality cues when answering. We further introduce a modality-specific reasoning mechanism that preserves modality isolation during the separated reasoning stage while enabling full access to cross-modal information at the evidence fusion stage. Experiments demonstrate consistent improvements in both accuracy and robustness, yielding an average relative gain of 5.16\% on general AVQA benchmarks and 11.17\% on a cross-modal hallucination benchmark.
【8】ChladniSonify: A Visual-Acoustic Mapping Method for Chladni Patterns in New Media Art Creation
标题:Chladni Sonify:新媒体艺术创作中Chladni模式的视听映射方法
链接:https://arxiv.org/abs/2605.09846
备注:9 pages, 5 figures, IEEE conference format
摘要:
摘要:
【9】Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems
标题:混音音色:基于扩散的风格跨复音茎转移
链接:https://arxiv.org/abs/2605.09259
摘要:音色转移的目的是在保留原始旋律和节奏的同时修改音乐录音的音色特征。虽然单乐器音色转移已经取得了实质性的进展,但多乐器设置的现有方法依赖于分离然后转移的流水线,该流水线传播源分离伪影并产生跨主干的不相干合成音色。本文提出了MixtureTT,据我们所知,第一个系统,灵活的每干音色直接从复调混合转移。给定每个目标声音的混合和单独的音色参考,MixtureTT通过共享的扩散过程将所有干联合传输到指定的乐器。通过对每个词干内容和交叉词干谐波之间的依赖关系进行建模,所提出的联合词干扩散Transformer消除了级联分离误差,将推理成本降低了等于词干数量的因子,并产生更相干的多词干输出。尽管在严格的输入条件下操作,但对SATB合唱数据集的评估表明,MixtureTT在客观和主观指标上都优于单乐器基线,这表明了在朴素的分离然后传输管道上进行专用多乐器音色传输的必要性。因此,这项工作证实,交叉干建模是必不可少的混合级音色转移,因为拟议的联合设置始终超过了等效的单柄消融。
摘要:Timbre transfer aims to modify the timbral identity of a musical recording while preserving the original melody and rhythm. While single-instrument timbre transfer has made substantial progress, existing approaches to multi-instrument settings rely on separate-then-transfer pipelines that propagate source separation artifacts and produce incoherent synthesized timbres across stems. This paper proposes MixtureTT, to the best of our knowledge the first system for flexible per-stem timbre transfer directly from a polyphonic mixture. Given a mixture and a separate timbre reference for each target voice, MixtureTT jointly transfers all stems to the specified instruments through a shared diffusion process. Modeling the dependencies across the per-stem content and cross-stem harmonic, the proposed joint stem diffusion transformer eliminates cascaded separation error, reduces inference cost by a factor equal to the number of stems, and yields more coherent multi-stem outputs. Despite operating under a strictly harder input condition, evaluations on the SATB choral dataset show that MixtureTT outperforms single-instrument baselines on both objective and subjective metrics demonstrating the necessity of dedicated multi-instrument timbre transfer over the naive separate-then-transfer pipelines. As a result, this work confirms that the cross-stem modeling is essential for mixture-level timbre transfer as the proposed joint setting consistently exceeds an equivalent single-stem ablation.
【10】Reddit2Deezer: A Scalable Dataset for Real-World Grounded Conversational Music Recommendation
标题:Reddit 2 Deezer:用于现实世界有针对性对话音乐推荐的可扩展数据集
链接:https://arxiv.org/abs/2605.09120
摘要:会话式音乐推荐(CMR)研究目前面临着一个权衡之间的真实对话语料库,是有限的规模和合成语料库,规模扩大,但其对话是人工构建的,而不是自然观察。在本文中,我们介绍了Reddit 2Deezer,一个基于现实的CMR资源,来自190 k个唯一的{线程,叶评论}对。我们以两个版本发布资源:保留真实性的原始版本和最大化长期可重复性的释义版本。每个音乐实体都链接到Deezer标识符,该标识符提供对音频预览和丰富元数据(例如,类型标签,流行度,BPM),为未来基于内容的会话式推荐研究打开了大门。人工验证确认对话、项目基础和释义的质量。该数据集可在https://huggingface.co/datasets/McAuley-Lab/Reddit2Deezer上获得。
摘要:Conversational music recommendation (CMR) research currently faces a tradeoff between authentic dialogue corpora that are limited in scale and synthesized corpora that scale up but whose conversations are artificially constructed rather than naturally observed. In this paper, we introduce Reddit2Deezer, a reality-grounded CMR resource derived from 190k unique {thread, leaf-comment} pairs. We release the resource in two versions: a raw version that preserves authenticity, and a paraphrased version that maximizes long-term reproducibility. Each musical entity is linked to a Deezer identifier, which provides straightforward access to audio previews and rich metadata (e.g., genre tags, popularity, BPM), opening the door to future research on content-grounded conversational recommendation. A human validation confirms the quality of the dialogues, item grounding, and paraphrases. The dataset is available at https://huggingface.co/datasets/McAuley-Lab/Reddit2Deezer.
【11】Towards Trustworthy Audio Deepfake Detection: A Systematic Framework for Diagnosing and Mitigating Gender Bias
标题:迈向值得信赖的音频Deepfake检测:诊断和缓解性别偏见的系统框架
链接:https://arxiv.org/abs/2605.09087
备注:Submitted to SMC 2026 conference
摘要:音频deepfake检测系统越来越多地部署在高风险的安全应用中,但它们在人口统计群体中的公平性仍然严重不足。以前的工作衡量性别差距,但没有调查它来自哪里或如何系统地解决它。我们提出了第一个诊断优先框架,该框架在应用有针对性的缓解措施之前识别偏差源,在ASVSpoof5上对两个模型AASIST和Wav2Vec2+ResNet18进行了评估。我们的诊断表明,偏见并不源于不平衡的训练数据,而是来自声学表征差异,学习特征中的性别泄漏和结构评估不对称。我们测试缓解策略在处理,后处理和组合的家庭,包括在这项工作中引入的新方法。按性别分别调整决策阈值可将不公平性降低54%至75%,而不会影响检测准确性,我们新的时代级公平性规则化方法优于现有的每批方法。只有当性别泄漏是局部的时,对抗性去偏见才会成功,而当性别泄漏是扩散的时,对抗性去偏见就会失败,这是我们在训练前的诊断正确预测的结果。没有一种方法可以完全消除公平性差距,这证实了在应用修正之前必须确定偏差来源,并且更公平的基准设计同样重要
摘要:Audio deepfake detection systems are increasingly deployed in high-stakes security applications, yet their fairness across demographic groups remains critically underexamined. Prior work measures gender disparity but does not investigate where it comes from or how to fix it systematically. We present the first diagnosis-first framework that identifies bias source before applying targeted mitigation, evaluated on two models, AASIST and Wav2Vec2+ResNet18, on ASVSpoof5. Our diagnosis shows that bias does not stem from imbalanced training data but from acoustic representation differences, gender leakage in learned features, and structural evaluation asymmetry. We test mitigation strategies across in-processing, post-processing and combined families, including novel methods introduced in this work. Adjusting the decision threshold separately per gender reduces unfairness by 54% to 75% at no cost to detection accuracy, and our new epoch-level fairness regularisation method outperforms existing per-batch approaches. Adversarial debiasing succeeds only when gender leakage is localised, and fails when it is diffuse, an outcome correctly predicted by our diagnosis before training. No single method fully closes the fairness gap, confirming that bias sources must be identified before fixes are applied and that fairer benchmark design is equally important
【12】Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search
标题:Omni-DeepSearch:音频驱动的Omni-Modal深度搜索的基准
链接:https://arxiv.org/abs/2605.08762
备注:43 pages
摘要:目前的全模态基准主要在同时提供多种模态的环境下评估模型,而仅从音频开始并积极搜索跨模态证据的能力仍然未得到充分探索。在本文中,我们介绍了\textbf{Omni-DeepSearch},这是一个音频驱动的全模态深度搜索基准。给定一个或多个音频片段和一个相关问题,模型必须从音频中推断出有用的线索,调用文本、图像和视频搜索工具,并执行多跳推理,以产生一个简短、客观和可验证的答案。Omni-DeepSearch包含15个细粒度类别的640个样本,涵盖四种检索目标模态和四种音频内容类型。多级过滤管道确保音频依赖性、检索必要性、视觉模态必要性和答案唯一性。在最近的闭源和开源全模态模型上的实验表明,这项任务仍然具有很大的挑战性:最强的评估模型Gemini-3-Pro的平均准确率仅为43.44%。进一步的分析说明了音频实体推理,查询公式,工具使用的可靠性,多跳检索和跨模态验证的关键瓶颈。这些结果突出了音频驱动的全模态深度搜索作为未来多模态代理的重要和未充分探索的方向。
摘要:Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start from audio alone and actively search for cross-modal evidence remains underexplored. In this paper, we introduce \textbf{Omni-DeepSearch}, a benchmark for audio-driven omni-modal deep search. Given one or more audio clips and a related question, models must infer useful clues from audio, invoke text, image, and video search tools, and perform multi-hop reasoning to produce a short, objective, and verifiable answer. Omni-DeepSearch contains 640 samples across 15 fine-grained categories, covering four retrieval target modalities and four audio content types. A multi-stage filtering pipeline ensures audio dependence, retrieval necessity, visual modality necessity, and answer uniqueness. Experiments on recent closed-source and open-source omni-modal models show that this task remains highly challenging: the strongest evaluated model, Gemini-3-Pro, achieves only 43.44\% average accuracy. Further analyses illustrate key bottlenecks in audio entity inference, query formulation, tool-use reliability, multi-hop retrieval, and cross-modal verification. These results highlight audio-driven omni-modal deep search as an important and underexplored direction for future multimodal agents.
【13】Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
标题:Unison:协调运动、语音和声音,实现以人为本的音频视频生成
链接:https://arxiv.org/abs/2605.08729
摘要:运动、语音和音效是以人为中心的视频的基本元素,但它们的异构时间特性使联合生成极具挑战性。现有的音频-视频生成模型通常无法在这些模态之间保持一致的对齐,从而导致运动、语音和环境声音之间的明显不匹配。我们提出了统一,一个统一的框架,明确促进连贯性的议案,讲话,和健全的方式。在音频流中,Unison采用了语义指导的协调策略,该策略可以简化语音和音效组件的生成。利用双向音频交叉注意和语义条件门控进行语义驱动的自适应重组,这种方法有效地减轻了语音优势,提高了声学清晰度。对于音频运动同步,我们提出了一个双向的跨模态强制策略,其中更清洁的模态通过解耦的去噪时间表来引导更嘈杂的模态,并通过渐进的稳定策略来加强。大量的实验表明,Unison在音频感知质量和跨模态同步方面都达到了最先进的性能,突出了以人为中心的视频生成中显式多模态协调的重要性。
摘要:Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Existing audio-video generation models often fail to maintain consistent alignment across these modalities, leading to noticeable mismatches between motion, speech, and environmental sounds. We present Unison, a unified framework that explicitly promotes coherence across the motion, speech, and sound modalities. Within the audio stream, Unison employs a semantic-guided harmonization strategy that decouples the generation of speech and sound-effect components. Leveraging bidirectional audio cross-attention and semantic-conditioned gating for semantic-driven adaptive recomposition, this approach effectively mitigates speech dominance and enhances acoustic clarity. For audio-motion synchronization, we propose a bidirectional cross-modal forcing strategy where the cleaner modality guides the noisier one through decoupled denoising schedules, reinforced by a progressive stabilization strategy. Extensive experiments demonstrate that Unison achieves state-of-the-art performance in both audio perceptual quality and cross-modal synchronization, highlighting the importance of explicit multimodal harmonization in human-centric video generation.
【14】Online Segmented Beamforming via Dynamic Programming
标题:通过动态规划进行在线分段束形成
链接:https://arxiv.org/abs/2605.08554
备注:4 pages, 2 figures
摘要:在以时变干扰源和移动源为特征的动态声学环境中,有效的波束形成需要随时间准确地识别静止区域。传统的Capon波束形成器依赖于瞬时集合协方差矩阵,这在实践中是不可访问的。实际的实现克服了这一点,估计样本协方差矩阵(SCM),通过平均在一个块的时间样本。然而,在非静态设置中,朴素的批处理方法失败了。移动的干扰源会模糊SCM,导致波束形成器在无法跟踪新的活动干扰源的同时将零值放置在过时的位置,从而降低其零值能力。为了解决这个基本的限制,提出了一种在线分段波束形成器。该算法结合了数据驱动的时间分割,因果地最小化输出功率,同时动态地适应SCM估计窗口的局部平稳性。通过框架的问题,通过镜头的动态规划,所提出的方法跟踪突然的环境变化,并重置协方差估计实时。我们在一个复杂的,混响的模拟声学环境中,并在高度混响的现实世界的实验中验证了这个框架的性能,证明了其优越性固定窗口自适应方法。
摘要:In dynamic acoustic environments characterized by time-varying interferers and moving sources, effective beamforming requires accurately identifying stationary regions over time. Traditional Capon beamformers rely on the instantaneous ensemble covariance matrix, which is inaccessible in practice. Practical implementations overcome this by estimating the sample covariance matrix (SCM) through averaging over a block of temporal samples. However, in non-stationary settings, a naive batch approach fails. Moving interferers smear the SCM, causing the beamformer to place nulls in outdated locations while failing to track newly active interferers, thereby degrading its nulling capabilities. To address this fundamental limitation, an Online Segmented Beamformer is proposed. This algorithm incorporates data-driven temporal segmentation to causally minimize output power while dynamically adapting the SCM estimation windows to local stationarity. By framing the problem through the lens of dynamic programming, the proposed method tracks abrupt environmental changes and resets covariance estimates in real-time. We validate the performance of this framework in a complex, reverberant simulated acoustic environment and in highly reverberant real world experiments, demonstrating its superiority over fixed-window adaptive methods.
【15】Uniqueness on a Continuum: Quantifying Tonal Ambiguity Using Information Theory
标题:连续体上的唯一性:使用信息论量化音调歧义
链接:https://arxiv.org/abs/2605.08224
备注:14 pages, 6 figures, 9 tables
摘要:我们提出了一个连续的测量音调的模糊性,扩展了既定的概念的独特性。虽然独特性被广泛认为是必要的音调,它不能(一)之间的区别,拥有它,(二)捕获层次组织的模式有限的换位,或(三)占时间展开。为了解决这些局限性,我们引入了一个同伴措施,接地信息理论,量化音调模糊的连续规模。该方法适用于音高级别的集合和调音系统,扩展了音调关系的分析范围,并为理论和分析提供了实用的工具。
摘要:We propose a continuous measure of tonal ambiguity that extends the established concept of uniqueness. While uniqueness is widely regarded as necessary for tonality, it cannot (i) discriminate among sets that possess it, (ii) capture hierarchical organization in modes of limited transposition, or (iii) account for temporal unfolding. To address these limitations, we introduce a companion measure, grounded in information theory, that quantifies tonal ambiguity on a continuous scale. The measure applies across pitch-class sets and tuning systems, expanding analytic coverage of tonal relationships and offering a practical tool for theory and analysis.
【16】Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
标题:Bangla-WhisperDiar:用于孟加拉语长形式语音识别和说话人拨号化的微调Whisper和PyAnnote
链接:https://arxiv.org/abs/2605.08214
备注:3 figures and 5 tables
摘要:孟加拉语的自动语音识别(ASR)和说话人日记化仍然具有挑战性,因为长格式的录音,不同的声学条件和显着的说话人变异性。这项工作通过开发强大的长形式ASR和说话者日记系统来解决孟加拉语口语理解中的这两个核心任务。对于ASR(问题1),我们在大约15,000个分块和对齐的孟加拉音频片段的定制数据集上微调了tugstugi bengaliai区域asr耳语媒体模型,采用全权重训练和广泛的数据增强,包括噪声注入,混响模拟,回声,剪切失真和音高/时间扰动。对于说话人日记化(问题2),我们使用PyTorch Lightning在竞争注释的日记化数据集上微调了pyannote/segmentation-3.0模型,将微调后的分割主干交换到pyannote/speaker-diarization-community-1管道中,同时保留预训练的说话人嵌入和聚类组件。我们的ASR系统实现了0.2441的单词错误率(WER),而我们的日记化系统实现了0.2392的日记化错误率(DER),两者都在测试集上进行了评估,表明相对于各自的预训练基线有显着的改进。我们描述了我们完整的流水线,包括数据预处理,文本规范化,音频增强,训练策略,推理优化和两个任务的后处理。
摘要:Automatic Speech Recognition (ASR) and speaker diarization in Bangla remain challenging due to long form recordings, diverse acoustic conditions, and significant speaker variability. This work addresses these two core tasks in Bangla spoken language understanding by developing robust systems for long form ASR and speaker diarization. For ASR (Problem 1), we fine tune the tugstugi bengaliai regional asr whisper medium model on a custom-curated dataset of approximately 15,000 chunked and aligned Bangla audio segments, employing full weight training with extensive data augmentation including noise injection, reverb simulation, echo, clipping distortion, and pitch/time perturbation. For speaker diarization (Problem 2), we fine-tune the pyannote/segmentation-3.0 model using PyTorch Lightning on the competition annotated diarization dataset, swapping the fine-tuned segmentation backbone into the pyannote/speaker-diarization-community-1 pipeline while retaining the pretrained speaker embedding and clustering components. Our ASR system achieves a Word Error Rate (WER) of 0.2441, while our diarization system achieves a Diarization Error Rate (DER) of 0.2392, both evaluated on the test set, demonstrating notable improvements over the respective pretrained baselines. We describe our complete pipeline, including data preprocessing, text normalization, audio augmentation, training strategies, inference optimization, and post-processing for both tasks.
【17】ShipEcho -- An Interactive Tool for Global Mapping of Underwater Radiated Noise from Vessels
标题:ShipEcho --船舶水下辐射噪音全球绘图的交互式工具
链接:https://arxiv.org/abs/2605.08194
备注:34 pages
摘要:船舶水下辐射噪声(V-URN)是公认的对海洋生态系统产生负面影响的环境压力源。大量资源被投入到V-URN监测指标、监管框架和面向管理的评估的开发中。一种具有高影响潜力的方法是V-URN制图,它可以为环境评估和缓解规划提供可操作的时空信息。制作管理规模的地图仍然具有挑战性,因为被动声学测量在空间上是稀疏的,许多业务系统依赖于专业工作流程和昂贵的广域船舶活动数据。为了解决这些限制,我们引入了ShipEcho,这是一个免费访问的基于Web的地理信息系统(GIS),它使用通过基于社区的AIS交换获得的船舶数据提供近实时的V-URN映射。ShipEcho使用已建立的船舶SL模型和由测深数据提供信息的传播建模,在全球各地区制作近实时和累积的噪声地图。这些包括使用标准指标的声压级和声暴露级,包括63 Hz和125 Hz三分之一倍频程和20- 2000 Hz宽带级。我们描述了系统架构,数据管道,建模工作流程和关键假设,并通过与声学记录的比较来评估地图的准确性。然后,我们将展示ShipEcho如何通过实际用例支持管理层评估、决策和政策举措。
摘要:Underwater radiated noise from vessels (V-URN) is a recognized environmental stressor that negatively impacts marine ecosystems. Significant resources are invested in the development of V-URN monitoring indicators, regulatory frameworks, and management-oriented assessments. One approach with high potential for impact is V-URN mapping, which can provide actionable spatiotemporal information for environmental assessment and mitigation planning. Producing management-scale maps remains challenging as passive acoustic measurements are spatially sparse and many operational systems depend on specialist workflows and costly access to wide-area vessel activity data. To address these constraints, we introduce ShipEcho, a freely accessible web-based Geographic Information System (GIS) that provides near-real-time V-URN mapping using vessel data acquired through a community-based AIS exchange. Using established vessel SL models and propagation modeling informed by bathymetric data, ShipEcho produces near-real-time and cumulative noise maps across regions worldwide. These include sound pressure levels and sound exposure levels using standard indicators, including the 63~Hz and 125~Hz one-third octave bands and a 20--2000~Hz broadband level. We describe the system architecture, data pipeline, modeling workflow, and key assumptions, and evaluate map accuracy through comparison with acoustic recordings. We then demonstrate how ShipEcho can support management-level assessment, decision-making, and policy initiatives through practical use cases.
【18】PoDAR: Power-Disentangled Audio Representation for Generative Modeling
标题:PoDART:用于生成式建模的功率分解音频表示
链接:https://arxiv.org/abs/2605.10084
备注:9 pages, 3 figures
摘要:音频潜在扩散模型的性能主要取决于生成器的表现力和潜在空间的可建模性。虽然最近的研究主要集中在前者,以及提高音频编解码器的重建保真度,我们证明,潜在的可建模性可以显着提高通过显式因子解纠缠。我们提出了PoDAR(功率分解音频表示),一个框架,利用随机功率增强和潜在的一致性目标解耦信号功率不变的语义内容。这种分解使得潜在空间更容易建模,这既加速了下游生成模型的收敛,又提高了最终的整体性能。当应用于带有F5-TTS生成器的Stable Audio 1.0 VAE时,PoDAR在收敛方面实现了约2\times $的加速,以匹配基线性能,同时在LibriSpeech-PC数据集上将最终扬声器相似性提高了0.055,UTMOS提高了0.22。此外,将功率隔离到专用通道中使CFG能够专门应用于功率不变的内容,有效地将稳定的制导机制扩展到更高的尺度。
摘要:The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focused primarily on the former, as well as improving the reconstruction fidelity of audio codecs, we demonstrate that latent modelability can be significantly improved through explicit factor disentanglement. We present PoDAR (Power-Disentangled Audio Representation), a framework that utilizes a randomized power augmentation and latent consistency objective to decouple signal power from invariant semantic content. This factorization makes the latent space easier to model, which both accelerates the convergence of downstream generative models and improves final overall performance. When applied to a Stable Audio 1.0 VAE with an F5-TTS generator, PoDAR achieves about a $2\times$ acceleration in convergence to match baseline performance, while increasing final speaker similarity by 0.055 and UTMOS by 0.22 on the LibriSpeech-PC dataset. Furthermore, isolating power into dedicated channels enables the application of CFG exclusively to power-invariant content, effectively extending the stable guidance regime to higher scales.
【1】SF-Flow: Sound field magnitude estimation via flow matching guided by sparse measurements
标题:SF-Flow:通过稀疏测量引导的流量匹配来估计磁场幅度
链接:https://arxiv.org/abs/2605.10398
摘要:从稀疏麦克风测量重建3D声场是一个基本但不适定的问题,我们通过声学传递函数(ATF)幅度估计来解决这个问题。ATF幅度封装了物理空间的关键感知和声学特性,并应用于房间表征和校正。虽然最近的生成范例,如流匹配(FM)已经取得了最先进的性能在语音和音乐生成,其潜力在空间音频仍然未被发掘。我们提出了一个新的框架,三维ATF幅度重建作为指导生成任务,与3D U-网条件下的置换不变集编码器。这种架构可以从任意数量的稀疏输入进行重建,同时利用FM的稳定和有效的训练特性。实验结果表明,SF-Flow实现了高达\SI{1}{kHz}的精确重建,训练速度比自动编码器基线快得多,并且随着数据集大小的增加而显着提高。
摘要:Reconstructing a 3D sound field from sparse microphone measurements is a fundamental yet ill-posed problem, which we address through Acoustic Transfer Function (ATF) magnitude estimation. ATF magnitude encapsulates key perceptual and acoustic properties of a physical space with applications in room characterization and correction. Although recent generative paradigms such as Flow Matching (FM) have achieved state-of-the-art performance in speech and music generation, their potential in spatial audio remains underexplored. We propose a novel framework for 3D ATF magnitude reconstruction as a guided generation task, with a 3D U-Net conditioned by a permutation-invariant set encoder. This architecture enables reconstruction from an arbitrary number of sparse inputs while leveraging the stable and efficient training properties of FM. Experimental results demonstrate that SF-Flow achieves accurate reconstruction up to \SI{1}{kHz}, trains substantially faster than the autoencoder baseline, and improves significantly with dataset size.
【2】PoDAR: Power-Disentangled Audio Representation for Generative Modeling
标题:PoDART:用于生成式建模的功率分解音频表示
链接:https://arxiv.org/abs/2605.10084
备注:9 pages, 3 figures
摘要:音频潜在扩散模型的性能主要取决于生成器的表现力和潜在空间的可建模性。虽然最近的研究主要集中在前者,以及提高音频编解码器的重建保真度,我们证明,潜在的可建模性可以显着提高通过显式因子解纠缠。我们提出了PoDAR(功率分解音频表示),一个框架,利用随机功率增强和潜在的一致性目标解耦信号功率不变的语义内容。这种分解使得潜在空间更容易建模,这既加速了下游生成模型的收敛,又提高了最终的整体性能。当应用于带有F5-TTS生成器的Stable Audio 1.0 VAE时,PoDAR在收敛方面实现了约2\times $的加速,以匹配基线性能,同时在LibriSpeech-PC数据集上将最终扬声器相似性提高了0.055,UTMOS提高了0.22。此外,将功率隔离到专用通道中使CFG能够专门应用于功率不变的内容,有效地将稳定的制导机制扩展到更高的尺度。
摘要:The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focused primarily on the former, as well as improving the reconstruction fidelity of audio codecs, we demonstrate that latent modelability can be significantly improved through explicit factor disentanglement. We present PoDAR (Power-Disentangled Audio Representation), a framework that utilizes a randomized power augmentation and latent consistency objective to decouple signal power from invariant semantic content. This factorization makes the latent space easier to model, which both accelerates the convergence of downstream generative models and improves final overall performance. When applied to a Stable Audio 1.0 VAE with an F5-TTS generator, PoDAR achieves about a $2\times$ acceleration in convergence to match baseline performance, while increasing final speaker similarity by 0.055 and UTMOS by 0.22 on the LibriSpeech-PC dataset. Furthermore, isolating power into dedicated channels enables the application of CFG exclusively to power-invariant content, effectively extending the stable guidance regime to higher scales.
【3】Single-Microphone Audio Point Source Discriminative Localization From Reverberation Late Tail Estimation
标题:来自回响晚尾估计的单麦克风音频点源区分性定位
链接:https://arxiv.org/abs/2605.09627
备注:Published at IEEE ICASSP 2026
摘要:位置信息对于音频分割任务来说是有价值的信号,特别是作为对专注于源的内容或质量的方法的补充。虽然音频源定位通常使用空间中的多个麦克风捕获的信号的观察来执行,但是关于源的位置的信息由单个麦克风通过其到达时间和频谱幅度来捕获-给定源的发射信号是已知的。由于混响源自房间中的音频源,因此它相应地包含关于所发射的音频信号的一些信息。混响的后尾部分对于本地源和麦克风几何结构是相对不变的,主要仅取决于房间本身,并且因此可以提供关于最小地取决于它们的位置的音频信号的必要参考信息。在这项工作中,我们利用加权预测误差(WPE)去混响的概率框架内的鲁棒的后尾估计估计,以估计在同一房间中收集的两个音频信号的可能性,因为它们来自同一位置。我们证明了我们的方法在模拟和真实环境中的扬声器日记任务的有效性。
摘要:Location information can be a valuable signal for audio segmentation tasks, especially as a complement to methods focusing on the content or qualities of the sources. Though audio source localization is typically performed using the observations of the signal captured by multiple microphones in space, information about a source's location is captured by a single microphone through its arrival time and spectral amplitude--given the source's emitted signal is known. Since reverberation originates from the audio sources in a room, it accordingly contains some information about the emitted audio signals. The late-tail part of reverberation is relatively invariant to the local source and microphone geometry, depending primarily on only the room itself, and thus can provide the necessary reference information about audio signals that depends minimally on their location. In this work, we leverage the robust late-tail estimation of Weighted Prediction Error (WPE) dereverberation within a probabilistic framework to estimate the likelihood of two audio signals collected in the same room as having originated from the same location. We demonstrate the effectiveness of our approach on the speaker diarization task in both simulated and real environments.
【4】RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations
标题:2026年雷达挑战:媒体转型下稳健的音频Deepfake识别
链接:https://arxiv.org/abs/2605.09568
备注:Submitted to APSIPA 2026
摘要:RADAR Challenge 2026是APSIPA关于媒体转换下的鲁棒音频Deepfake识别的大型挑战赛,旨在模拟真实世界音频分发管道中的真实媒体条件,包括压缩,重混,噪声和混响。它包括两个阶段:一个是英语开发阶段,包含用于分析和论文写作的标记数据,另一个是多语言评估阶段,包含英语、新加坡英语、普通话、台湾普通话、日语和越南语的100,000多个话语。系统评估使用等错误率(EER)的二进制真/假分类。本文描述了挑战任务,数据集的构建,评估协议和总体结果。在挑战赛期间,33支队伍进入了开发阶段,22支队伍进入了最终评估阶段。报告的结果强调了在多语言和媒体转换条件下强大的音频deepfake检测的剩余挑战。
摘要:RADAR Challenge 2026 is an APSIPA Grand Challenge on Robust Audio Deepfake Recognition under Media Transformations, designed to simulate realistic media conditions in real-world audio distribution pipelines, including compression, resampling, noise, and reverberation. It consists of two phases: an English development phase with labeled data for analysis and paper writing, and a multilingual evaluation phase containing more than 100,000 utterances in English, Singapore English, Mandarin Chinese, Taiwanese Mandarin, Japanese, and Vietnamese. Systems are evaluated using equal error rate (EER) for binary real/fake classification. This paper describes the challenge task, the construction of the data set, the evaluation protocol, and the overall results. During the challenge, 33 teams submitted to the development phase and 22 teams submitted to the final evaluation phase. The reported results highlight the remaining challenges of robust audio deepfake detection under multilingual and media-transformed conditions.
【5】Evaluating the Expressive Appropriateness of Speech in Rich Contexts
标题:评估丰富语境中言语的表达适当性
链接:https://arxiv.org/abs/2605.09413
备注:19 pages, 6 figures
摘要:评估表达性语音仍然具有挑战性,因为现有的方法主要评估情感强度,而忽略了语音样本是否在表达上适合其上下文环境。这种限制阻碍了可靠的评估语音系统中使用的叙事驱动和交互式应用程序,如有声读物和会话代理。我们引入CEAEval,一个上下文丰富的框架,用于评估演讲中的表达恰当性,它评估演讲样本是否与其话语级叙事上下文所暗示的潜在交际意图相一致。为了支持这一任务,我们构建CEAEval-D,第一个上下文丰富的语音数据集与真正的人类表现在普通话会话语音,提供叙事描述连同15个维度的人类注释,涵盖表达属性和表达适当性。我们进一步开发了CEAEval-M模型,该模型集成了知识蒸馏,基于计划的多模型协作,自适应音频注意力偏差和强化学习,以执行上下文丰富的表达适当性评估。在人工标注的测试集上的实验表明,CEAEval-M大大优于现有的语音评估和分析系统。
摘要:Evaluating expressive speech remains challenging, as existing methods mainly assess emotional intensity and overlook whether a speech sample is expressively appropriate for its contextual setting. This limitation hinders reliable evaluation of speech systems used in narrative-driven and interactive applications, such as audiobooks and conversational agents. We introduce CEAEval, a Context-rich framework for Evaluating Expressive Appropriateness in speech, which assesses whether a speech sample expressively aligns with the underlying communicative intent implied by its discourse-level narrative context. To support this task, we construct CEAEval-D, the first context-rich speech dataset with real human performances in Mandarin conversational speech, providing narrative descriptions together with fifteen dimensions of human annotations covering expressive attributes and expressive appropriateness. We further develop CEAEval-M, a model that integrates knowledge distillation, planner-based multi-model collaboration, adaptive audio attention bias, and reinforcement learning to perform context-rich expressive appropriateness evaluation. Experiments on a human-annotated test set demonstrate that CEAEval-M substantially outperforms existing speech evaluation and analysis systems.
【6】Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech
标题:Zero-Shot文本到语音中公制诱导离散流匹配的带矩修正的动态最优调度
链接:https://arxiv.org/abs/2605.09386
备注:Under Review
摘要:度量诱导的离散流匹配(MI-DFM)利用标记潜在的几何离散生成,但其实际应用受到两个问题的限制:需要超参数搜索的启发式算法,以及来自其一阶连续时间马尔可夫链(CTMC)求解器的有限步路径跟踪误差。我们解决了这两个问题。首先,我们推导出一个动力学最优调度规定的标量参数化概率路径,并将其实例化为MI-DFM作为一个无训练的数值调度,遍历路径在恒定的费舍尔-拉奥速度。其次,我们引入了一个有限步矩校正,调整的跳跃概率,同时保持CTMC跳跃目的地分布。我们验证所得到的方法,GibbsTTS,基于编解码器的zero-shot文本到语音(TTS)。在与统一架构和大规模数据集的受控比较下,GibbsTTS实现了最佳的客观自然度,并且在主观评估中优于掩蔽的离散生成基线。此外,与评估的最先进的TTS系统相比,GibbsTTS显示出很强的说话人相似性,在四个测试集中的三个测试集上达到最高的相似性,在第四个测试集上排名第二。项目页面:https://ydqmkkx.github.io/GibbsTTSProject
摘要:Metric-induced discrete flow matching (MI-DFM) exploits token-latent geometry for discrete generation, but its practical use is limited by two issues: heuristic schedulers requiring hyperparameter search, and finite-step path-tracking error from its first-order continuous-time Markov chain (CTMC) solver. We address both issues. First, we derive a kinetic-optimal scheduler for prescribed scalar-parameterized probability paths, and instantiate it for MI-DFM as a training-free numerical schedule that traverses the path at constant Fisher-Rao speed. Second, we introduce a finite-step moment correction that adjusts the jump probability while preserving the CTMC jump destination distribution. We validate the resulting method, GibbsTTS, on codec-based zero-shot text-to-speech (TTS). Under controlled comparisons with a unified architecture and large-scale dataset, GibbsTTS achieves the best objective naturalness and is preferred in subjective evaluations over masked discrete generative baselines. Additionally, in comparison with the evaluated state-of-the-art TTS systems, GibbsTTS shows strong speaker similarity, achieving the highest similarity on three of four test sets and ranking second on the fourth. Project page: https://ydqmkkx.github.io/GibbsTTSProject
【7】Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation
标题:通过噪音不变声学-语义蒸馏减少基于LM的语音增强中的语言幻觉
链接:https://arxiv.org/abs/2605.08608
摘要:基于语言模型(LM)的语音增强(SE)可以生成听起来自然的语音,但在严重的噪声下,它通常会受到不可靠的条件作用,导致感知上看似合理但语言上不正确的输出。为了解决这个问题,我们提出了L3-SE,一个噪声不变的声学语义蒸馏框架,用于减少基于LM的SE中的语言幻觉。所提出的方法学习噪声不变的条件编码器从嘈杂的语音通过联合蒸馏两个互补的干净的语音目标:重建保真度的声学目标和语义目标的语言一致性。所得到的噪声不变的声学语义表示被用于调节仅解码器的自回归语言模型,该模型预测被解码为增强语音的干净声学令牌。为了支持高质量的生成,我们进一步采用了一个高保真的编解码器建立在可学习的加权WavLM层表示作为离散的声学接口。通过提高在不利条件下条件反射的可靠性,所提出的框架大大减少了幻觉,提高了内容的忠实性。实验表明,该方法在语言一致性指标上始终优于现有的基于LM的语音增强基线,在低信噪比和混响条件下具有特别明显的增益,同时保持了具有竞争力的感知质量。音频样本可在https://max1wz.github.io/L3-SE-Demo-Page/上获得。完整的源代码将在稿件被接受后发布。
摘要:Language model (LM)-based speech enhancement (SE) can generate natural-sounding speech, but under severe noise it often suffers from unreliable conditioning, leading to perceptually plausible yet linguistically incorrect outputs. To address this issue, we propose L3-SE, a noise-invariant acoustic-semantic distillation framework for reducing linguistic hallucination in LM-based SE. The proposed method learns a noise-invariant conditioning encoder from noisy speech by jointly distilling two complementary clean-speech targets: an acoustic target for reconstruction fidelity and a semantic target for linguistic consistency. The resulting noise-invariant acoustic-semantic representations are used to condition a decoder-only autoregressive language model, which predicts clean acoustic tokens that are decoded into enhanced speech. To support high-quality generation, we further employ a high-fidelity codec built on learnable weighted WavLM layer representations as the discrete acoustic interface. By improving the reliability of conditioning under adverse conditions, the proposed framework substantially reduces hallucination and improves content faithfulness. Experiments show that the proposed method consistently outperforms prior LM-based speech enhancement baselines on linguistic consistency metrics, with especially clear gains under low-SNR and reverberant conditions, while maintaining competitive perceptual quality. Audio samples are available at https://max1wz.github.io/L3-SE-Demo-Page/. The complete source code will be released after the manuscript is accepted.
【8】Latent Secret Spin: Keyed Orthogonal Rotations for Blind Speech Watermarking in Anisotropic Latent Spaces
标题:潜在秘密旋转:各向异性潜在空间中盲语音水印的键控垂直旋转
链接:https://arxiv.org/abs/2605.08431
摘要:介绍了一种基于编码器隐空间几何运算的语音盲水印算法-隐秘密旋转(LSS)。基于正交旋转的主成分,LSS诱导不可感知的,但可检测的协方差签名根据伪随机水印时间表。该方案在数据集上进行推广,保持感知质量,并且与一些学习的神经水印方案不同,它不需要神经网络训练,能够抵抗常见的信号操作,并且对有效载荷大小灵活。分析表明,结构化潜在空间水印是一个有前途的和可解释的替代现有的方法。
摘要:We introduce Latent Secret Spin (LSS), a blind speech watermarking method based on geometric operations in codec latent space. Based upon orthogonal rotations to principal components, LSS induces imperceptible but detectable covariance signatures according to a pseudo-random watermarking schedule. The scheme generalises across datasets, preserves perceptual quality and, unlike some learned, neural watermarking schemes, it does not require neural network training, is resistant to common signal manipulations and is flexible to payload size. Analyses show that structured latent-space watermarking is a promising and interpretable alternative to existing approaches.
【9】DiffVQE: Hybrid Diffusion Voice Quality Enhancement Under Acoustic Echo and Noise
标题:迪夫VQE:声学回声和噪音下的混合扩散语音质量增强
链接:https://arxiv.org/abs/2605.08189
备注:6 pages, 4 figures, submitted to Interspeech 2026
摘要:声学回声和背景噪声对免提系统和免提电话中的语音增强提出了挑战。区别性训练的端到端方法代表了联合声学回声控制(AEC)和去噪的强大解决方案。然而,随着生成方法的出现,基于扩散的方法在语音增强任务中表现出色。在这项工作中,据我们所知,我们提供了第一个(仍然是非因果的)基于扩散的AEC模型(DiffVQE),它在拓扑结构,训练数据和训练框架方面是可重复的。到目前为止,在没有采用扩散的情况下,微软的区别性DeepVQE模型已被证明优于ICASSP 2023 AEC挑战赛的任何参赛作品,取得了卓越的性能。使用来自Interspeech 2025 URGENT Challenge的数据进行多样化的高质量训练数据集,我们的DiffVQE在回声和噪声控制性能以及计算复杂性和模型大小方面都优于DeepVQE。
摘要:Acoustic echo and background noise pose challenges on speech enhancement in hands-free systems and speakerphones. Discriminatively trained end-to-end methods represent a powerful solution for joint acoustic echo control (AEC) and denoising. However, with the advent of generative methods, diffusion-based approaches have seen remarkable performance in speech enhancement tasks. In this work, to the best of our knowledge, we provide the first (still non-causal) diffusion-based AEC model (DiffVQE) that is reproducible in terms of topology, training data, and training framework. So far, without employing diffusion, Microsoft's discriminative DeepVQE model has been shown to excel any of the ICASSP 2023 AEC Challenge entries achieving remarkable performance. Using data from the Interspeech 2025 URGENT Challenge for a diverse, high-quality training dataset, our DiffVQE excels DeepVQE both in echo and noise control performance, as well as in computational complexity and model size.
【10】Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models
标题:重新思考自回归模型测试时自适应中的最小化
链接:https://arxiv.org/abs/2605.08186
备注:Submitted to INTERSPEECH 2026
摘要:通过熵最小化(EM)的测试时间自适应(TTA)已被证明是有效的分类任务,但其应用于生成自回归模型仍然在理论上是支离破碎的。现有的方法通常依赖于不同的数学方法,例如使用伪标签的教师强制或基于策略梯度的强化学习,而没有统一的数学基础。在这项工作中,我们通过推导出针对自回归模型定制的EM的严格公式来解决这个差异。我们表明,确切的目标自然分解成一个令牌级的政策梯度损失和令牌级的熵损失,我们重新解释先前的方法作为这个统一的提法的部分实现。使用Whisper ASR作为测试平台,我们证明了我们的方法在20多个不同的领域中不断提高性能,包括噪音,口音和多语言设置。
摘要:Test-Time Adaptation (TTA) via entropy minimization (EM) has proven effective for classification tasks, yet its application to generative autoregressive models remains theoretically fragmented. Existing approaches typically rely on distinct heuristics, such as teacher forcing with pseudo labels or policy-gradient-based reinforcement learning, without a unified mathematical foundation. In this work, we resolve this discrepancy by deriving a rigorous formulation of EM tailored to autoregressive models. We show that the exact objective naturally decomposes into a token-level policy gradient loss and a token-level entropy loss, and we reinterpret prior methods as partial realizations of this unified formulation. Using Whisper ASR as a testbed, we demonstrate that our approach consistently improves performance across more than 20 diverse domains, including acoustic noise, accents, and multilingual settings.
【11】Low-Cost Detection of Degraded Voice Clones via Source-Output Acoustic Consistency
标题:通过源输出声学一致性低成本检测降级语音克隆
链接:https://arxiv.org/abs/2605.08165
备注:7 pages, 3 figures
摘要:生成语音的最新进展增加了对明显失败的合成输出的自动检测的需求。这在临床环境中尤其重要,例如AVATAR治疗,其中精神分裂症患者参与计算机生成的幻觉声音表示,并且退化的合成可能会破坏沉浸和治疗参与。我们调查是否低维,可解释的源输出声学功能可以提供一个轻量级的第一遍检测器退化的语音克隆输出。受语音的源滤波器模型的启发,我们首先测试中值基频(f0)作为源相关的一致性度量,并将其与声道长度(VND)作为滤波器相关的度量和谐波噪声比(HNR)作为噪声相关的描述符进行比较。使用输入-输出特征空间中的非对称阈值处理程序评估了使用两个声码器家族WaveRNN(n=54)和HiFi-GAN(n=40)生成的人类标记的语音克隆样本。对于WaveRNN,f0和HNR都达到了85.2%的准确率,优于RNN(64.8%)。对于HiFi-GAN,HNR达到了80.0%的准确率,其次是f0为77.5%,而HNR为67.5%。样本水平重叠和光谱检查表明,f0和HNR捕获部分不同的故障模式,而不是提供相同样本的冗余排名。这些结果表明,简单的源输出声学一致性措施,可以提供有用的第一次通过检测退化的语音克隆,并支持使用可解释的基于阈值的筛选失败的合成语音必须迅速被拒绝的应用程序。
摘要:Recent advances in generative speech have increased the need for automatic detection of obviously failed synthetic outputs. This is particularly important in clinical settings such as AVATAR therapy, in which schizophrenia patients engage with a computer-generated representation of their hallucinated voices and degraded synthesis may disrupt immersion and therapeutic engagement. We investigate whether low-dimensional, interpretable source-output acoustic features can provide a lightweight first-pass detector of degraded voice-cloning outputs. Motivated by source-filter models of speech, we first test median fundamental frequency (f0) as a source-related consistency measure, and compare it with vocal tract length (VTL) as a filter-related measure and Harmonics-to-Noise Ratio (HNR) as a noise-related descriptor. Human-labeled voice-cloning samples generated with two vocoder families, WaveRNN (n=54) and HiFi-GAN (n=40), were evaluated using an asymmetric thresholding procedure in the input-output feature space. For WaveRNN, f0 and HNR both achieved 85.2% accuracy, outperforming VTL (64.8%). For HiFi-GAN, HNR achieved 80.0% accuracy, followed by f0 at 77.5% and VTL at 67.5%. Sample-level overlap and spectrographic inspection showed that f0 and HNR capture partly distinct failure patterns, rather than providing redundant rankings of the same samples. These results show that simple source-output acoustic consistency measures can provide useful first-pass detection of degraded voice clones, and support the use of interpretable threshold-based screening in applications where failed synthetic speech must be rejected quickly.
【12】Probing Cross-modal Information Hubs in Audio-Visual LLMs
标题:探索视听LLM中的跨模态信息枢纽
链接:https://arxiv.org/abs/2605.10815
备注:Accepted by ICML 2026
摘要:视听大语言模型(AVLLM)最近已经成为一个强大的架构,能够联合推理的音频,视觉和文本模态。在AVLLM中,音频和视频模态之间的双向交互引入了复杂的处理动态,需要对其内部机制进行更深入的理解。然而,与广泛研究的纯文本或大型视觉语言模型不同,AVLLM的内部工作原理在很大程度上仍未被探索。在本文中,我们专注于音频和视觉模态之间的跨模态信息流AVLLM,调查来自一种模态的信息编码在其他模态的令牌表示。通过对多个近期AVLLM的分析,我们发现了两个共同的发现。首先,AVLLM主要在汇令牌中编码集成的视听信息。其次,汇令牌不统一持有跨模态信息。相反,一个独特的汇令牌子集,我们称之为跨模态汇令牌,专门存储这些信息。基于这些发现,我们进一步提出了一个简单的无训练幻觉缓解方法,鼓励依赖于跨模态汇令牌内的综合跨模态信息。我们的代码可在https://github.com/kaistmm/crossmodal-hub上获得。
摘要:Audio-visual large language models (AVLLMs) have recently emerged as a powerful architecture capable of jointly reasoning over audio, visual, and textual modalities. In AVLLMs, the bidirectional interaction between audio and video modalities introduces intricate processing dynamics, necessitating a deeper understanding of their internal mechanisms. However, unlike extensively studied text-only or large vision language models, the internal workings of AVLLMs remain largely unexplored. In this paper, we focus on cross-modal information flow between audio and visual modalities in AVLLMs, investigating where information derived from one modality is encoded within the token representations of the other modality. Through an analysis of multiple recent AVLLMs, we uncover two common findings. First, AVLLMs primarily encode integrated audio-visual information in sink tokens. Second, sink tokens do not uniformly hold cross-modal information. Instead, a distinct subset of sink tokens, which we term cross-modal sink tokens, specializes in storing such information. Based on these findings, we further propose a simple training-free hallucination mitigation method by encouraging reliance on integrated cross-modal information within cross-modal sink tokens. Our code is available at https://github.com/kaistmm/crossmodal-hub.
【13】Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration
标题:Polyphonia:具有声学注意力校准的复音音乐中的Zero-Shot音色传输
链接:https://arxiv.org/abs/2605.10203
备注:Accepted by ICML 2026
摘要:基于扩散的文本到音乐生成的进步为zero-shot音乐编辑开辟了新的途径。然而,现有的方法无法实现特定于干的音色转移,这需要改变特定的干,同时严格保留背景伴奏。这一限制严重阻碍了实际应用,因为实际生产需要对致密混合物中的组分进行精确操作。我们的主要发现是,虽然香草交叉注意力捕捉语义特征的茎,它缺乏光谱分辨率,严格本地化的目标在密集的混合物,导致边界泄漏。为了解决这个难题,我们提出了Polyphonia,一个带有声学信息注意力校准的zero-shot编辑框架。而不是仅仅依赖于扩散语义注意,Polyphonia利用概率声学之前建立粗略的边界,使非目标茎保留精确的语义合成。为了进行评估,我们提出了PolyEvalentts,这是一个标准化的提示集,包含复调音乐中的1,170个音色转移任务。具体来说,与基线相比,Polyphonia在目标对齐方面实现了15.5%的增长,同时保持了竞争性的音乐保真度和非目标完整性。
摘要:The advancement of diffusion-based text-to-music generation has opened new avenues for zero-shot music editing. However, existing methods fail to achieve stem-specific timbre transfer, which requires altering specific stems while strictly preserving the background accompaniment. This limitation severely hinders practical application, since real-world production necessitates precise manipulation of components within dense mixtures. Our key finding is that, while vanilla cross-attention captures semantic features of stems, it lacks the spectral resolution to strictly localize targets in dense mixtures, leading to boundary leakage. To resolve this dilemma, we propose Polyphonia, a zero-shot editing framework with Acoustic-Informed Attention Calibration. Rather than relying solely on diffuse semantic attention, Polyphonia leverages a probabilistic acoustic prior to establish coarse boundaries, enabling non-target stems preserved precise semantic synthesis. For evaluation, we propose PolyEvalPrompts, a standardized prompt set with 1,170 timbre transfer tasks in polyphonic music. Specifically, Polyphonia achieves an increase of 15.5% in target alignment compared to baselines, while maintaining competitive music fidelity and non-target integrity.
【14】How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue
标题:LLM说话时应该如何倾听?全场语音对话中用户流路由研究
链接:https://arxiv.org/abs/2605.10199
摘要:全双工口语对话需要一个模型在生成自己的口语响应的同时保持倾听。这对于大型语言模型(LLM)来说是具有挑战性的,LLM被设计为扩展单个相干序列,并且不自然地支持在生成期间到达的用户输入。我们认为,如何将用户流路由到LLM因此是一个关键的全双工建模架构问题。为了研究这个问题,我们将纯文本LLM扩展到统一的全双工口语对话系统中,并在共享训练管道下比较两种路由策略:(i)通道融合,将用户流直接注入LLM输入,以及(ii)交叉注意路由,将用户流作为通过交叉注意适配器访问的外部存储器。在口语问答和全双工交互基准测试上的实验揭示了一个明显的权衡。通道融合产生更强的语义基础和持续更好的问答性能。然而,在语义重叠的条件下,如用户中断,它更容易受到上下文损坏:如果模型未能及时停止,重叠的用户流可能会干扰正在进行的生成,并导致语义不连贯的延续。交叉注意路由在问题回答上表现不佳,但更好地保留了LLM生成上下文,并且对这种故障模式更鲁棒。这些结果建立用户流路由作为全双工口语对话的中心设计轴,并提供实际指导语义集成和上下文鲁棒性之间的权衡。我们提供了一个演示页面的定性检查。
摘要:Full-duplex spoken dialogue requires a model to keep listening while generating its own spoken response. This is challenging for large language models (LLMs), which are designed to extend a single coherent sequence and do not naturally support user input arriving during generation. We argue that how the user stream is routed into the LLM is therefore a key architectural question for full-duplex modeling. To study this question, we extend a text-only LLM into a unified full-duplex spoken dialogue system and compare two routing strategies under a shared training pipeline: (i) channel fusion, which injects the user stream directly into the LLM input, and (ii) cross-attention routing, which keeps the user stream as external memory accessed through cross-attention adapters. Experiments on spoken question answering and full-duplex interaction benchmarks reveal a clear tradeoff. Channel fusion yields stronger semantic grounding and consistently better question-answering performance. However, under semantically overlapping conditions such as user interruptions, it is more vulnerable to context corruption: if the model fails to stop in time, the overlapping user stream can interfere with ongoing generation and lead to semantically incoherent continuations. Cross-attention routing underperforms on question answering, but better preserves the LLM generation context and is more robust to this failure mode. These results establish user-stream routing as a central design axis in full-duplex spoken dialogue and offer practical guidance on the tradeoff between semantic integration and context robustness. We provide a demo page for qualitative inspection.
【15】Dolphin-CN-Dialect: Where Chinese Dialects Matter
标题:Dolphin-CN-Dialect:中国方言的重要性
链接:https://arxiv.org/abs/2605.08961
摘要:我们提出了Dolphin-CN-Dialect,这是一个支持流媒体的ASR模型,专注于中文和方言丰富的场景。与之前的版本相比,Dolphin-CN-Dialect在数据处理、标记化、训练稳定性和数据采样策略方面进行了实质性的改进。为了解决高度不平衡的方言数据的挑战,我们提出了一种基于温度的采样策略,有效地平衡标准普通话和低资源方言,从而显着提高方言识别性能。此外,我们重新设计了标记器,以更好地符合语言特征,采用字符级建模的中文和子词建模的英语,同时引入可扩展的方言令牌。实验结果表明,与Dolphin相比,Dolphin-CN-Dialect在方言识别准确率和CER降低方面都有了明显的提高。此外,Dolphin-CN-Dialect与最近的SOTA开源ASR模型达到了竞争性能,同时保持了显着更小的模型尺寸。Dolphin-CN-Dialect支持流式和非流式推理,从而在延迟和准确性之间实现实际平衡。它还通过热词支持提供灵活的自定义,并针对专用硬件进行了优化的高效部署。这些改进使Dolphin-CN-Dialect成为现实世界多方言ASR应用程序的强大而实用的解决方案。
摘要:We present Dolphin-CN-Dialect, a streaming-capable ASR model with a focus on Chinese and dialect-rich scenarios. Compared to the previous version, Dolphin-CN-Dialect introduces substantial improvements in data processing, tokenization, training stability, and data sampling strategies. To address the challenges of highly imbalanced dialect data, we propose a temperature-based sampling strategy that effectively balances standard Mandarin and low-resource dialects, leading to significant gains in dialect recognition performance. In addition, we redesign the tokenizer to better align with linguistic characteristics, adopting character-level modeling for Chinese and subword modeling for English, while introducing extensible dialect tokens. Experimental results show that Dolphin-CN-Dialect achieves improvement in dialect recognition accuracy and CER reduction compared to Dolphin. Furthermore, Dolphin-CN-Dialect reaches competitive performance with recent SOTA open-source ASR models, while maintaining a significantly smaller model size. Dolphin-CN-Dialect supports both streaming and non-streaming inference, enabling a practical balance between latency and accuracy. It also provides flexible customization through hotword support and efficient deployment optimized for specialized hardware. These improvements make Dolphin-CN-Dialect a strong and practical solution for real-world multi-dialect ASR applications.
【16】Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
标题:Bangla-WhisperDiar:用于孟加拉语长形式语音识别和说话人拨号化的微调Whisper和PyAnnote
链接:https://arxiv.org/abs/2605.08214
备注:3 figures and 5 tables
摘要:孟加拉语的自动语音识别(ASR)和说话人日记化仍然具有挑战性,因为长格式的录音,不同的声学条件和显着的说话人变异性。这项工作通过开发强大的长形式ASR和说话者日记系统来解决孟加拉语口语理解中的这两个核心任务。对于ASR(问题1),我们在大约15,000个分块和对齐的孟加拉音频片段的定制数据集上微调了tugstugi bengaliai区域asr耳语媒体模型,采用全权重训练和广泛的数据增强,包括噪声注入,混响模拟,回声,剪切失真和音高/时间扰动。对于说话人日记化(问题2),我们使用PyTorch Lightning在竞争注释的日记化数据集上微调了pyannote/segmentation-3.0模型,将微调后的分割主干交换到pyannote/speaker-diarization-community-1管道中,同时保留预训练的说话人嵌入和聚类组件。我们的ASR系统实现了0.2441的单词错误率(WER),而我们的日记化系统实现了0.2392的日记化错误率(DER),两者都在测试集上进行了评估,表明相对于各自的预训练基线有显着的改进。我们描述了我们完整的流水线,包括数据预处理,文本规范化,音频增强,训练策略,推理优化和两个任务的后处理。
摘要:Automatic Speech Recognition (ASR) and speaker diarization in Bangla remain challenging due to long form recordings, diverse acoustic conditions, and significant speaker variability. This work addresses these two core tasks in Bangla spoken language understanding by developing robust systems for long form ASR and speaker diarization. For ASR (Problem 1), we fine tune the tugstugi bengaliai regional asr whisper medium model on a custom-curated dataset of approximately 15,000 chunked and aligned Bangla audio segments, employing full weight training with extensive data augmentation including noise injection, reverb simulation, echo, clipping distortion, and pitch/time perturbation. For speaker diarization (Problem 2), we fine-tune the pyannote/segmentation-3.0 model using PyTorch Lightning on the competition annotated diarization dataset, swapping the fine-tuned segmentation backbone into the pyannote/speaker-diarization-community-1 pipeline while retaining the pretrained speaker embedding and clustering components. Our ASR system achieves a Word Error Rate (WER) of 0.2441, while our diarization system achieves a Diarization Error Rate (DER) of 0.2392, both evaluated on the test set, demonstrating notable improvements over the respective pretrained baselines. We describe our complete pipeline, including data preprocessing, text normalization, audio augmentation, training strategies, inference optimization, and post-processing for both tasks.
【17】ShipEcho -- An Interactive Tool for Global Mapping of Underwater Radiated Noise from Vessels
标题:ShipEcho --船舶水下辐射噪音全球绘图的交互式工具
链接:https://arxiv.org/abs/2605.08194
备注:34 pages
摘要:船舶水下辐射噪声(V-URN)是公认的对海洋生态系统产生负面影响的环境压力源。大量资源被投入到V-URN监测指标、监管框架和面向管理的评估的开发中。一种具有高影响潜力的方法是V-URN制图,它可以为环境评估和缓解规划提供可操作的时空信息。制作管理规模的地图仍然具有挑战性,因为被动声学测量在空间上是稀疏的,许多业务系统依赖于专业工作流程和昂贵的广域船舶活动数据。为了解决这些限制,我们引入了ShipEcho,这是一个免费访问的基于Web的地理信息系统(GIS),它使用通过基于社区的AIS交换获得的船舶数据提供近实时的V-URN映射。ShipEcho使用已建立的船舶SL模型和由测深数据提供信息的传播建模,在全球各地区制作近实时和累积的噪声地图。其中包括使用标准指标的声压级和声暴露级,包括63 Hz和125 Hz三分之一倍频程以及20- 2000 Hz宽带级。我们描述了系统架构,数据管道,建模工作流程和关键假设,并通过与声学记录的比较来评估地图的准确性。然后,我们将展示ShipEcho如何通过实际用例支持管理层评估、决策和政策举措。
摘要:Underwater radiated noise from vessels (V-URN) is a recognized environmental stressor that negatively impacts marine ecosystems. Significant resources are invested in the development of V-URN monitoring indicators, regulatory frameworks, and management-oriented assessments. One approach with high potential for impact is V-URN mapping, which can provide actionable spatiotemporal information for environmental assessment and mitigation planning. Producing management-scale maps remains challenging as passive acoustic measurements are spatially sparse and many operational systems depend on specialist workflows and costly access to wide-area vessel activity data. To address these constraints, we introduce ShipEcho, a freely accessible web-based Geographic Information System (GIS) that provides near-real-time V-URN mapping using vessel data acquired through a community-based AIS exchange. Using established vessel SL models and propagation modeling informed by bathymetric data, ShipEcho produces near-real-time and cumulative noise maps across regions worldwide. These include sound pressure levels and sound exposure levels using standard indicators, including the 63~Hz and 125~Hz one-third octave bands and a 20--2000~Hz broadband level. We describe the system architecture, data pipeline, modeling workflow, and key assumptions, and evaluate map accuracy through comparison with acoustic recordings. We then demonstrate how ShipEcho can support management-level assessment, decision-making, and policy initiatives through practical use cases.
机器翻译由腾讯交互翻译提供,仅供参考
