今日论文合集:CS.SD语音与音频 | 共 6 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
1. SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning
SpeechGym:一种用于通过强化学习训练语音智能体的原生音频 Gym
AI 总结:研究针对语音智能体训练中梯度无法流动、难以强化学习的问题,提出原生音频环境 SpeechGym,用每轮过程奖励解决稀疏性问题,使开放权重模型在语音基准上任务成功率翻倍且排名提升。
链接:https://arxiv.org/abs/2608.26432
机构:University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校); Amazon AGI Foundations(亚马逊AGI基础研究部)
作者:Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Jia-Hong Huang, Qi Luo, M. Maruf, Ivan Bulyko, Ge Liu, Roger Ren
英文摘要:Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow and per-call cost makes on-policy reinforcement learning prohibitive, or stay in text: they measure voice agents but cannot improve them. We present SpeechGym, an audio-native agentic environment in which two omni-modal models converse in native audio, with no external ASR or TTS and no API boundary, over the unmodified tasks, tools and success check of an established text agentic benchmark, so that the interaction modality is the only variable and the loop stays local and trainable end to end. Audio agentic capability does not follow from audio understanding. The failures speech introduces are perceptual rather than reasoning deficits: the agent picks the right tool and the right argument slot but fills it with a value misheard from the waveform, and that single error cascades into a failed call, a retry of the same call, and a wasted step budget. A second failure is behavioural: under an insistent caller the agent performs an unauthorised write and ends the episode believing it helped. Both are trainable, because the environment labels them for free: a call with a misheard argument fails against the database while a correct one succeeds. The obstacle is sparsity, not signal. Outcome-only GRPO is gradient-starved here, since almost every rollout group fails identically, while a per-turn process reward crediting each successful tool call restores variance to nearly every group. Trained this way, the agent transfers with no further tuning to an independently implemented voice benchmark, more than doubling task success and carrying an open-weights model from last place to second on that leaderboard, while using fewer turns and tokens than before training.
2. AudioSpan: Spanning the Duration and Depth of Audio Comprehension
AudioSpan:覆盖音频理解的时长与深度
AI 总结:本研究提出覆盖10分钟至2小时音频的基准AudioSpan,含3240个分三认知层级的问题,评估12个大型音频-语言模型,发现从长音频提取相关事实是核心难点,该基准可公开获取。
链接:https://arxiv.org/abs/2608.26431
机构:Qwen Team, Alibaba Group(通义千问团队,阿里巴巴集团); Tsinghua University(清华大学); The Chinese University of Hong Kong(香港中文大学)
作者:Wen Huang, Yunfei Chu, Meng Gao, Haolin He, Jin Xu
英文摘要:General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal. Yet the benchmarks that test them still rely on clips of seconds, where scores saturate and models converge; recent long-form efforts extend duration but evaluate long audio much as short clips are. We introduce AudioSpan, a benchmark that spans both duration and depth: it pairs audio from 10 minutes to over 2 hours with 3,240 questions across three cognitive levels, namely perception, understanding, and reasoning. Two paths supply the questions, differing in how question content is sourced and how ground truth is obtained. Native QA extracts questions from the audio's content, posing each as a multiple-choice item and an open-ended one graded by detailed rubrics. Anchor QA instead injects ground truth, planting acoustic anchors into the audio and building a perception-to-reasoning chain scored only to the first error. A fully automated pipeline constructs every item through structured captioning, QA generation, and adversarial critic feedback. Evaluating 12 LALMs on AudioSpan, we find the hard part comes before reasoning: distilling a few relevant facts from a long, redundant signal. This difficulty grows with audio length and falls hardest on perception, especially temporal grounding. AudioSpan is available at this https URL.
3. Direct or Mediated? Task-Dependent Audio Information Routing in Large Audio Language Models
直接还是中介?大型音频语言模型中任务依赖的音频信息路由
AI 总结:该研究针对大型音频语言模型,发现其在拼接两段音频的任务中,自动语音识别稳定而音频问答性能大幅下降,揭示了两类任务依赖不同的音频信息路由路径,指出信息利用是其泛化的潜在限制。
链接:https://arxiv.org/abs/2608.27026
机构:Graduate School of Informatics, Kyoto University(京都大学信息学研究科); WXG, Tencent(腾讯微信事业群)
作者:Yizhou Zhang, Wangjin Zhou, Xin Gu, Yichi Wang, Wei Tan, Yi Zhao, Zhi Gong, Keisuke Imoto, Tatsuya Kawahara
英文摘要:Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio understanding tasks. However, they are typically evaluated on single, coherent audio segments, leaving their behavior under less familiar input configurations underexplored. We study this issue through a controlled setting in which two audio segments are concatenated into a single input. Across multiple LALMs, we observe a striking task-dependent robustness gap: automatic speech recognition (ASR) remains comparatively stable, whereas audio question answering (AQA) degrades substantially. To investigate the mechanisms underlying this disparity, we analyze how audio information is routed through LALM decoders using layer-wise attention knockout. The results reveal distinct task-dependent pathways. ASR relies primarily on direct retrieval from audio tokens by answer tokens, whereas AQA depends more strongly on a mediated route in which audio information is first integrated into prompt tokens and subsequently accessed during generation. We further probe prompt-token representations under audio concatenation and find that task-relevant audio attributes remain readily decodable, particularly in middle and later decoder layers, even when AQA performance deteriorates sharply. This dissociation indicates that the failure cannot be explained by complete loss of audio information from the decoder states and is instead consistent with a downstream bottleneck in retrieving or utilizing prompt-mediated information during answer generation. Together, our findings reveal task-dependent audio information routing in LALMs and highlight information utilization as a potential limitation on their generalization.
4. Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition
用于鲁棒视听语音识别中对比解码的注意力引导可靠性缩放
AI 总结:本研究针对鲁棒视听语音识别问题,提出注意力引导的对比解码可靠性缩放方法,在LRS3数据集上实现了干净与低信噪比条件下的性能提升。
链接:https://arxiv.org/abs/2608.26213
机构:Hanyang University(汉阳大学)
作者:YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang
英文摘要:Large language model (LLM)-based audio-visual speech recognition (AVSR) systems are robust under noise. Contrastive decoding (CD), originally introduced to stabilize LLM generation by contrasting a weaker model against a stronger one at inference time, adjusts predictions without additional training. In this work, we apply CD to AVSR by contrasting audio-only conditioning with full audio-visual conditioning within the same underlying model. However, using a fixed contrastive strength introduces a trade-off across noise levels: stronger intervention helps under severe noise but may over-correct reliable predictions in clean conditions. We propose reliability-aware scaling of CD for AVSR. Instead of using a fixed strength, we adaptively modulate the contrastive influence at each token based on reliability signals derived from attention dynamics and inter-model predictive divergence. Experiments on LRS3 show consistent improvements across clean and low-SNR conditions.
5. StreamAV-Bench: A Comprehensive Benchmark for Streaming Audio-Video Generation
StreamAV-Bench:面向流式音视频生成的综合基准测试集
AI 总结:该研究针对现有基准无法捕捉流式音视频生成特性的问题,推出首个综合基准StreamAV-Bench,构建含双赛道的评估框架并评估13个系统,揭示当前模型的缺陷并给出模型开发见解。
链接:https://arxiv.org/abs/2608.26336
机构:BAAI(北京智源人工智能研究院); PKU(北京大学); Kling; THU(清华大学); USTC(中国科学技术大学)
作者:Kaiqi Liu, Haoxuan Zeng, Jingqi Liu, Jiacong Fang, Ziqi Cai, Yunyao Mao, Henglin Liu, Yu Sheng, Shuchen Weng, Boxin Shi
英文摘要:Recent advancements in generative models are pushing video generation toward unbounded streaming audio-video generation for real-time interactive worlds. However, existing benchmarks primarily evaluate completed sequences and struggle to capture streaming properties. To bridge this gap, we introduce StreamAV-Bench, the first comprehensive benchmark tailored for streaming audio-video generation. StreamAV-Bench establishes a unified evaluation framework, including the progressive track for instruction adherence and long-horizon stability, and the interactive track for interactive response and state retention and reuse. With expert-verified evaluation cases across 32 fine-grained dimensions, we conduct an extensive evaluation of 13 representative systems. Our analysis reveals that current models suffer from temporal drift in progressive generation and responsiveness bottlenecks during interactive control. Based on a comprehensive failure analysis, we share insights to advance the development of native joint audio-video streaming models.
6. Decay-Region Group Delay as a Forensic Cue for AI-Generated Impulsive Sounds
衰减群时延作为AI生成脉冲声的取证线索
AI 总结:该研究提出用衰减区域群时延作为取证线索,通过随机森林、CNN等模型验证其可区分AI生成脉冲声与真实脉冲声,为AI生成音频的溯源提供了新方法。
链接:https://arxiv.org/abs/2608.26346
机构:Institute for Artificial Intelligence and Data Science(人工智能与数据科学研究院); University at Buffalo(布法罗大学)
作者:JaeHyeong Chang, Chengzhe Sun, Siwei Lyu
英文摘要:We investigate whether AI-generated impulsive sounds can be distinguished from real ones through group delay analysis. Our central finding is that AI-generated impulsive sounds show near-identical onset-region group-delay distributions but exhibit measurably different group-delay behavior in the late decay region: decay-region KL divergence reaches $0.322$ compared to near-zero onset divergence ($0.022$). Cross-band GD variability achieves single-feature AUC~=~0.720, and a Random Forest (RF) over nine decay-region features reaches AUC~$=$~0.884 under sample-disjoint evaluation. A group delay map used as a standalone 2D input to CNN classifiers achieves 90--94\% accuracy, demonstrating that group delay carries substantial discriminative information. Under generator hold-out, CNN and transformer classifiers show highly variable AUC (0.457--0.918). The group delay RF achieves the highest average hold-out accuracy among the evaluated methods ($66.7\%$) and avoids extreme below-random collapse, although its average AUC (0.731) is lower than CNN avg (0.762) and AST (0.772). Parameter sensitivity analysis across 27 STFT configurations confirms that the RF AUC remains stable (0.700--0.847, std~=~0.035). These results suggest that decay-region group delay can serve as a physically interpretable forensic cue that complements magnitude-based classifiers, while broader validation remains necessary.
语音与音频学术速递[8.28]
评论 0
文明发言,友善讨论
还没有评论,发表你的看法,来抢沙发~
