Annual Meeting of the Association for Computational Linguistics(简称 ACL)年会是计算语言学和自然语言处理领域国际排名第一的顶级学术会议,由国际计算语言学协会组织,每年召开一次,在中国计算机学会(CCF)推荐会议列表中被列为 A 类会议。第64届ACL年会,将于2026 年 7 月 2 日至 7 日在美国加利福尼亚州圣迭戈举办。

本期分享内蒙古大学语音理解与生成实验室中稿的5篇论文,其中2篇论文被 ACL 主会录用,3篇被 Findings of ACL 录用。

Main Conference

论文题目:TellWhisper: Tell Whisper Who Speaks When

作者:Yifan Hu, Peiji Yang, Zhisheng Wang, Yicheng Zhong,Rui Liu*

合作单位:腾讯

论文摘要:Multi-speaker automatic speech recognition (MASR) aims to predict ''who spoke when and what'' from multi-speaker speech, a key technology for multi-party dialogue understanding. However, most existing approaches decouple temporal modeling and speaker modeling when addressing ''when'' and ''who'': some inject speaker cues before encoding (e.g., speaker masking), which can cause irreversible information loss; others fuse identity by mixing speaker posteriors after encoding, which may entangle acoustic content with speaker identity. This separation is brittle under rapid turn-taking and overlapping speech, often leading to degraded performance. To address these limitations, we propose TellWhisper, a unified framework that jointly models speaker identity and temporal within the speech encoder. Specifically, we design TS-RoPE, a time-speaker rotary positional encoding: time coordinates are derived from frame indices, while speaker coordinates are derived from speaker activity and pause cues. By applying region-specific rotation angles, the model explicitly captures per-speaker continuity, speaker-turn transitions, and state dynamics, enabling the attention mechanism to simultaneously attend to ''when'' and ''who''. Moreover, to estimate frame-level speaker activity, we develop Hyper-SD, which casts speaker classification in hyperbolic space to enhance inter-class separation and refine speaker-activity estimates. Extensive experiments demonstrate the effectiveness of the proposed approach.

论文链接:https://arxiv.org/abs/2601.03712


论文题目:MoE Adapter for Large Audio Language Models: Sparsity, Disentanglement, and Gradient-Conflict-Free

作者:Yishu Lei,Shuwei He, Hu Jing, Dan Zhang, Xianlong Luo, Danxiang Zhu, Shikun Feng,Rui Liu, Jingzhou HE, Yu Sun, Hua Wu, Haifeng Wang

合作单位:百度、清华深圳研究生院

论文摘要:Extending the input modality of Large Language Models (LLMs) to the audio domain is essential forachieving comprehensive multimodal perception. However, it is well-known that acoustic informationis intrinsically heterogeneous, entangling attributes such as speech, music, and environmental context.Existing research is limited to a dense, parameter-shared adapter to model these diverse patterns, whichinduces gradient conflict during optimization, as parameter updates required for distinct attributescontradict each other. To address this limitation, we introduce the MoE-Adapter, a sparse Mixtureof-Experts (MoE) architecture designed to decouple acoustic information. Specifically, it employs adynamic gating mechanism that routes audio tokens to specialized experts capturing complementaryfeature subspaces while retaining shared experts for global context, thereby mitigating gradient conflictsand enabling fine-grained feature learning. Comprehensive experiments show that the MoE-Adapterachieves superior performance on both audio semantic and paralinguistic tasks, consistently outperformingdense linear baselines with comparable computational costs. Furthermore, we will release the relatedcode and models to facilitate future research.

论文链接:https://arxiv.org/abs/2601.02967


Findings

论文题目:MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation

作者:Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen, Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li,Rui Liu...

合作单位:新加坡A*STAR 、 新加坡科技设计大学(SUTD) 、新加坡南洋理工(NTU)、阿里达摩院、微软亚洲研究院、上海财经大学、江西师范大学等

论文摘要:Large language models (LLMs) are now used worldwide, yet their multimodal understanding and reasoning often degrade outside Western, high-resource settings. We propose MMA-ASIA, a comprehensive framework to evaluate LLMs' cultural awareness with a focus on Asian contexts. MMA-ASIA centers on a human-curated, multilingual, and multimodally aligned multiple-choice benchmark covering 8 Asian countries and 10 languages, comprising 27,000 questions; over 79 percent require multi-step reasoning grounded in cultural context, moving beyond simple memorization. To our knowledge, this is the first dataset aligned at the input level across three modalities: text, image (visual question answering), and speech. This enables direct tests of cross-modal transfer. Building on this benchmark, we propose a five-dimensional evaluation protocol that measures: (i) cultural-awareness disparities across countries, (ii) cross-lingual consistency, (iii) cross-modal consistency, (iv) cultural knowledge generalization, and (v) grounding validity. To ensure rigorous assessment, a Cultural Awareness Grounding Validation Module detects "shortcut learning" by checking whether the requisite cultural knowledge supports correct answers. Finally, through comparative model analysis, attention tracing, and an innovative Vision-ablated Prefix Replay (VPR) method, we probe why models diverge across languages and modalities, offering actionable insights for building culturally reliable multimodal LLMs.

论文链接:https://arxiv.org/abs/2510.08608


论文题目:RAG-KT: Cross-platform Explainable Knowledge Tracing with Multi-view Fusion Retrieval Generation

作者:Zhiyi Duan,Hongyu Yuan,Rui Liu*

论文摘要:Knowledge Tracing (KT) infers a student’s knowledge state from past interactions to predict future performance. Conventional Deep Learning (DL)-based KT models are typically tied to platform-specific identifiers and latent representations, making them hard to transfer and interpret. Large Language Model (LLM)-based methods can be either ungrounded under prompting or overly domain-dependent under fine-tuning. In addition, most existing KT methods are developed and evaluated under a same-distribution assumption. In real deployments, educational data often arise from heterogeneous platforms with substantial distribution shift, which often degrades generalization. To this end, we propose RAG-KT, a retrieval-augmented paradigm that frames cross-platform KT as reliable context constrained inference with LLMs. It builds a unified multi-source structured context with cross-source alignment via Question Group abstractions and retrieves complementary rich and reliable context for each prediction, enabling grounded prediction and interpretable diagnosis. Experiments on three public KT benchmarks demonstrate consistent gains in accuracy and robustness, including strong performance under cross-platform conditions.


论文题目:BloomEval: A Bloom’s Cognitive Taxonomy-Based Benchmark for Evaluating LRMs via Cognitive Hierarchy Trace

作者:Zhiyi Duan, Lei Gao, Jiangshan Guan, Qi Wang,Rui Liu*

合作单位:吉林大学

论文摘要:Current benchmarks for Large Reasoning Models (LRMs) primarily rely on answer correctness, failing to assess the structural coherence and cognitive soundness of the reasoning process itself. To address this gap, we introduce Cognitive Hierarchy Trace (CHT), a novel evaluation framework grounded in Bloom’s Cognitive Taxonomy (BCT). CHT provides a structured, step-wise mapping of a model’s reasoning trajectory onto hierarchical cognitive levels, enabling the detection of structural anomalies such as hierarchy jumps, breaks, and overthinking. Based on CHT, we present BloomEval, the first large-scale benchmark designed for fine-grained cognitive capability assessment. It comprises 94,602 math problems, each annotated with Bloom’s cognitive levels, CHT trajectories, a three-tier knowledge hierarchy, and problem difficulty. To ensure scalable yet reliable annotation, we develop an Expert-LLM collaborative pipeline with a three-stage reconciliation mechanism. Our comprehensive evaluation reveals a critical finding: models often arrive at correct answers through cognitively flawed or opaque reasoning paths. The CHT-based analysis uncovers prevalent structural inconsistencies that are invisible to outcome-only metrics, demonstrating that answer accuracy is an insufficient proxy for reasoning quality.