今日论文合集:CS.SD语音与音频 | 共 4 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

1. EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$
EXAM²:扩展多语言与多模态分析中的音频理解能力
AI 总结:本文提出多语言多模态音频理解基准EXAM²,评估现有模型发现其存在多语言跨模态理解差距,微调的Gemma3n-EXAM²性能显著提升,推动相关研究发展。
链接:https://arxiv.org/abs/2608.23758
作者:Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen
英文摘要:Recent large audio language models (LALMs) have achieved impressive progress in audio understanding. However, existing evaluations remain largely constrained to English and narrow audio domains. Prior benchmarks typically focus on a single audio modality, i.e., speech, sound, or music, limiting the systematic investigation into how these models generalize across diverse visual scenarios. In this paper, we introduce EXAM$^2$, a benchmark for multilingual and multimodal audio understanding spanning six languages and multiple modalities, including speech, sound, music, mixed-audio settings, and visual images. By incorporating visual information alongside heterogeneous audio inputs, EXAM$^2$ enables more realistic evaluation of scene-aware audio reasoning and cross-modal comprehension. EXAM$^2$ comprises $5,667$ multiple-choice questions, $22,614$ image instances, and $135,684$ multilingual translations. We evaluate state-of-the-art open-source and proprietary LALMs as well as multimodal LLMs, revealing substantial performance gaps in multilingual and cross-modal understanding. Furthermore, we propose Gemma3n-EXAM$^2$, a lightweight fusion-model fine-tuned on EXAM$^2$-train, achieves up to $12.4\%$ improvement in multilingual settings and $21.7\%$ gains in multimodal evaluation over a strong baseline. Empirical results establish EXAM$^2$ as a challenging benchmark and pioneer future multilingual and multimodal audio intelligence research.

2. CoSTALA: Compositional Spatio-Temporal Audio-Language Alignment via Multi-Grain Hierarchical Contrastive Learning
CoSTALA:基于多粒度层次对比学习的组合式时空音频-语言对齐
AI 总结:CoSTALA通过多粒度层次对比学习构建新型训练范式,解决传统ALM难以处理多事件音频序列的问题,为时空音频理解提供强大新框架。
链接:https://arxiv.org/abs/2608.24374
机构:Xi’an Jiaotong Liverpool University(西交利物浦大学); MiLM Plus, Xiaomi Inc.(小米公司MiLM Plus); University of Oulu(奥卢大学); Institute of Acoustics, Chinese Academy of Sciences(中国科学院声学研究所)
作者:Peiwei Ren, Jinbo Hu, Fang Kang, Shan Liang, Yin Cao
英文摘要:Conventional audio language models (ALMs) have made significant progress in achieving alignment between auditory and textual representations, including recent explorations in spatial audio. However, in daily spatial scenarios, they still cannot effectively process multi-event audio sequences. Current approaches primarily rely on coarse-grained contrastive learning with global auditory and textual features, lacking the resolution to distinguish multiple sequential events. To overcome these limitations, we propose CoSTALA-a novel training paradigm that transitions from purely global alignment to fine-grained spatio-temporal reasoning. By constructing a multi-granularity hierarchical loss function system, we achieve explicit modeling of temporal dependencies, and successfully anchors individual acoustic events to preserve their semantic purity. Extensive experiments demonstrate that CoSTALA significantly establish a powerful new framework for spatio-temporal audio understanding.

3. Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding
不止聆听,尝试规划:面向长音频会议理解的基于图的检索-生成智能体
AI 总结:针对长音频会议理解任务的问答数据集稀缺及现有语音模型的声学信息丢失、长期记忆差问题,构建LongAudioQA数据集,提出基于图的GRGA模型,利用智能体规划实现检索与答案生成。
链接:https://arxiv.org/abs/2608.24048
机构:Jiangsu Key Lab of Language Computing(江苏省语言计算重点实验室)
作者:Quanwei Tang, Dong Zhang, Shoushan Li, Guodong Zhou
英文摘要:While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain scarce. Existing speech QA paradigms and state-of-the-art Speech LLMs suffer from acoustic information loss and poor long-term context memory. To address these issues, we construct the LongAudioQA dataset and propose the GRGA model, which models heterogeneous audio features into a multi-dimensional graph and leverages agent planning for retrieval and answer generation.

4. On the Robustness of Audio Deepfake Detection under Audio Watermarking
音频水印下音频深度伪造检测的鲁棒性研究
AI 总结:本研究探究音频水印对ADD系统的影响,采用WavMark评估框架,发现水印会在部分数据集上大幅降低ADD性能,揭示了ADD系统的鲁棒性漏洞。
链接:https://arxiv.org/abs/2608.24159
机构:School of Information Technology, Monash University, Malaysia campus(莫纳什大学马来西亚校区信息技术学院); Idiap Research Institute, Switzerland(瑞士 Idiap 研究所)
作者:Zi Qian Yong, Ajinkya Kulkarni, Julia Lau, Hwa Hui Tew, Shu Min Leong, Raphael Phan, Sébastien Marcel
英文摘要:Recent advances in generative audio models have enabled highly realistic synthetic speech, increasing the importance of reliable audio deepfake detection (ADD) systems. While prior studies have primarily focused on adversarially optimized perturbations, the robustness of ADD systems under realistic signal transformations remains insufficiently understood. In this work, we investigate the impact of audio watermarking on ADD systems by treating watermarking as a structured, non-adversarial perturbation rather than a conventional attack mechanism. Using a watermark-based evaluation framework built upon WavMark, we evaluate multiple self-supervised learning (SSL), Convolutional Neural Network (CNN) and Graph Neural Netrowk (GNN)-based ADD models across several benchmark datasets. Beyond conventional detection metrics, we further analyze watermark-induced representation shifts using Fréchet Distance, cosine similarity, and L2 distance in the embedding space. Experimental results reveal a strong dataset-dependent behavior: watermarking causes substantial performance degradation on ASVspoof 2021 LA and DF, while exhibiting limited impact on ASVspoof 2024, FoR, and ITW. Moreover, large embedding-space shifts are strongly associated with severe detection degradation, suggesting that watermark-induced perturbations can substantially alter the feature representations relied upon by current ADD systems. These findings demonstrate that benign signal transformations designed for content protection can expose previously overlooked robustness vulnerabilities in audio deepfake detection systems. Our code is available at this https URL