今日论文合集:cs.SD语音20篇,eess.AS音频处理15篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】WeDefense: A Toolkit to Defend Against Fake Audio
标题:WeDefense:防御虚假音频的工具包
链接:https://arxiv.org/abs/2601.15240

作者:Lin Zhang,Johan Rohdin,Xin Wang,Junyi Peng,Tianchi Liu,You Zhang,Hieu-Thi Luong,Shuai Wang,Chengdong Liang,Anna Silnova,Nicholas Evans
备注:This is an ongoing work. v1 corresponds to the version completed by June 4, 2025 and previously submitted to ASRU 2025
摘要:生成AI的进步使得合成音频的创建成为可能,这种合成音频在感知上与真实的真实音频无法区分。虽然这一重大进展使许多积极的应用成为可能,但它也增加了滥用的风险,例如冒充,虚假信息和欺诈。尽管通过众多挑战和倡议发布了越来越多的开源虚假音频检测代码,但大多数都是针对特定的比赛,数据集或模型而量身定制的。一个标准化和统一的工具包,支持公平的基准测试和竞争解决方案的比较,不仅有共同的数据库,协议,指标,而且还有共享的代码库,是缺失的。为了解决这个问题,我们提出了WeDefense,这是第一个支持虚假音频检测和本地化的开源工具包。除了模型训练之外,WeDefense还强调关键但经常被忽视的组件:灵活的输入和增强,校准,分数融合,标准化的评估指标以及用于更深入理解和解释的分析工具。该工具包可在https://github.com/zlin0/wedefense上公开获取,其中包含用于虚假音频检测和定位的交互式演示。
摘要:The advances in generative AI have enabled the creation of synthetic audio which is perceptually indistinguishable from real, genuine audio. Although this stellar progress enables many positive applications, it also raises risks of misuse, such as for impersonation, disinformation and fraud. Despite a growing number of open-source fake audio detection codes released through numerous challenges and initiatives, most are tailored to specific competitions, datasets or models. A standardized and unified toolkit that supports the fair benchmarking and comparison of competing solutions with not just common databases, protocols, metrics, but also a shared codebase, is missing. To address this, we propose WeDefense, the first open-source toolkit to support both fake audio detection and localization. Beyond model training, WeDefense emphasizes critical yet often overlooked components: flexible input and augmentation, calibration, score fusion, standardized evaluation metrics, and analysis tools for deeper understanding and interpretation. The toolkit is publicly available at https://github.com/zlin0/wedefense with interactive demos for fake audio detection and localization.


【2】WavLink: Compact Audio--Text Embeddings with a Global Whisper Token
标题:WavLink:紧凑音频--文本嵌入Global Whisper代币
链接:https://arxiv.org/abs/2601.15118

作者:Gokul Karthik Kumar,Ludovick Lepauloux,Hakim Hacid
备注:Accepted at ICASSP 2026
摘要:Whisper已经成为在大型音频语言模型中提取通用音频特征的事实上的编码器,其中30秒的剪辑通常由投影到LLM中的1500帧特征表示。相比之下,音频文本嵌入模型(如基于CLAP的模型)在很大程度上依赖于替代音频编码器(例如,HTS-AT,PaSST),并没有有效地利用耳语。我们提出了WavLink,一个紧凑的音频文本嵌入模型,它用一个可学习的全局令牌来增强Whisper编码器,并与文本编码器一起训练。通过对设计选择的系统研究,包括预训练的文本编码器,损失函数,训练模式和数据混合,我们确定了产生最先进的检索性能的配置。我们在三种模型大小上的两阶段训练配方,结合Matryoshka风格的监督,提高了可扩展性,实现了8倍小的嵌入,性能下降最小。WavLink还在具有MCQ和zero-shot分类的AIR-Bench上展示了具有竞争力的性能。
摘要:Whisper has become the de-facto encoder for extracting general-purpose audio features in large audio-language models, where a 30-second clip is typically represented by 1500 frame features projected into an LLM. In contrast, audio-text embedding models like CLAP-based models have largely relied on alternative audio encoders (e.g., HTS-AT, PaSST), and have not leveraged Whisper effectively. We present WavLink, a compact audio-text embedding model that augments Whisper encoder with a learnable global token, trained jointly with a text encoder. Through a systematic study of design choices, including pretrained text encoders, loss functions, training modes, and data mixtures, we identify configurations that yield state-of-the-art retrieval performance. Our two-stage training recipe across three model sizes, combined with Matryoshka-style supervision, improves scalability, enabling 8x smaller embeddings with minimal performance drop. WavLink also demonstrates competitive performance on AIR-Bench with MCQs and zero-shot classification.


【3】Bangla Music Genre Classification Using Bidirectional LSTMS
标题:使用双向LSTMS的孟加拉音乐流派分类
链接:https://arxiv.org/abs/2601.15083

作者:Muntakimur Rahaman,Md Mahmudul Hoque,Md Mehedi Hassain
摘要:孟加拉音乐有着丰富的音乐文化。现在,音乐流派分类是非常重要的,因为可用的音乐,无论是在数字和物理格式的指数增长。有必要相应地为它们编制索引,以便于改进检索。按流派自动分类孟加拉音乐对于在庞大而多样化的音乐库中有效定位特定作品至关重要。用于流派分类的主流方法主要采用传统的机器学习或深度学习方法。这项工作介绍了一个新的音乐数据集,包括十个不同类型的孟加拉音乐。对于音频分类的任务,我们利用递归神经网络(RNN)架构。具体而言,实现长短期记忆(LSTM)网络以训练模型并执行分类。特征提取是音频数据处理的基础阶段。本研究利用梅尔频率倒谱系数(MFCC)转换成一个紧凑的和有代表性的功能集的原始音频波形。所提出的框架通过利用这些提取的特征来促进音乐流派分类。实验结果表明,78%的分类准确率,表明该系统的强大潜力,以提高和简化组织的孟加拉音乐流派。
摘要:Bangla music is enrich in its own music cultures. Now a days music genre classification is very significant because of the exponential increase in available music, both in digital and physical formats. It is necessary to index them accordingly to facilitate improved retrieval. Automatically classifying Bangla music by genre is essential for efficiently locating specific pieces within a vast and diverse music library. Prevailing methods for genre classification predominantly employ conventional machine learning or deep learning approaches. This work introduces a novel music dataset comprising ten distinct genres of Bangla music. For the task of audio classification, we utilize a recurrent neural network (RNN) architecture. Specifically, a Long Short-Term Memory (LSTM) network is implemented to train the model and perform the classification. Feature extraction represents a foundational stage in audio data processing. This study utilizes Mel-Frequency Cepstral Coefficients (MFCCs) to transform raw audio waveforms into a compact and representative set of features. The proposed framework facilitates music genre classification by leveraging these extracted features. Experimental results demonstrate a classification accuracy of 78%, indicating the system's strong potential to enhance and streamline the organization of Bangla music genres.


【4】VCNAC: A Variable-Channel Neural Audio Codec for Mono, Stereo, and Surround Sound
标题:VNAC:一款适用于单声、立体声和环绕声的可变通道神经音频编解码器
链接:https://arxiv.org/abs/2601.14960

作者:Florian Grötschla,Arunasish Sen,Alessandro Lombardi,Guillermo Cámbara,Andreas Schwarz
备注:Submitted to EUSIPCO 2026
摘要:我们提出了VCNAC,可变通道神经音频编解码器。我们的方法具有单一的编码器和解码器参数化,使不同的通道设置,从单声道语音到电影5.1声道环绕音频的本地推理。通道兼容性目标确保多通道内容在解码到较少通道时保持感知质量。共享表示使生成语言模型能够在单组码本上进行训练,同时支持跨模态和通道配置的推理时间可扩展性。使用客观空间音频指标和主观听力测试的评估表明,我们的统一方法在单声道,立体声和环绕声音频配置中保持高重建质量。
摘要:We present VCNAC, a variable channel neural audio codec. Our approach features a single encoder and decoder parametrization that enables native inference for different channel setups, from mono speech to cinematic 5.1 channel surround audio. Channel compatibility objectives ensure that multi-channel content maintains perceptual quality when decoded to fewer channels. The shared representation enables training of generative language models on a single set of codebooks while supporting inference-time scalability across modalities and channel configurations. Evaluation using objective spatial audio metrics and subjective listening tests demonstrates that our unified approach maintains high reconstruction quality across mono, stereo, and surround audio configurations.


【5】Generative Artificial Intelligence, Musical Heritage and the Construction of Peace Narratives: A Case Study in Mali
标题:生成性人工智能、音乐遗产与和平叙事的构建:马里的案例研究
链接:https://arxiv.org/abs/2601.14931

作者:Nouhoum Coulibaly,Ousmane Ly,Michael Leventhal,Ousmane Goro
备注:12 pages, 2 figures
摘要:本研究探讨了生成人工智能(Gen AI)为马里和平叙事的构建和音乐遗产的振兴做出贡献的能力。这项研究是在一个政治和社会背景下进行的,在这个背景下,社区间的紧张关系和社会分裂促使人们寻求新的和解象征性框架。该研究从经验上探讨了三个问题:(1)人工智能如何作为植根于民族语言和传统的音乐创作工具;(2)人工智能系统在多大程度上实现了技术创新和文化真实性之间的平衡混合;以及(3)人工智能辅助的音乐共同创作如何加强社会凝聚力和文化主权。实验结果表明,嵌入在具有文化意识的参与式框架中的Gen AI可以作为象征性外交的催化剂,放大当地的声音,而不是将其标准化。然而,在语言语料库的可用性、算法审查以及从受版权保护的来源生成作品的道德方面,挑战依然存在。
摘要:This study explores the capacity of generative artificial intelligence (Gen AI) to contribute to the construction of peace narratives and the revitalization of musical heritage in Mali. The study has been made in a political and social context where inter-community tensions and social fractures motivate a search for new symbolic frameworks for reconciliation. The study empirically explores three questions: (1) how Gen AI can be used as a tool for musical creation rooted in national languages and traditions; (2) to what extent Gen AI systems enable a balanced hybridization between technological innovation and cultural authenticity; and (3) how AI-assisted musical co-creation can strengthen social cohesion and cultural sovereignty. The experimental results suggest that Gen AI, embedded in a culturally conscious participatory framework, can act as a catalyst for symbolic diplomacy, amplifying local voices instead of standardizing them. However, challenges persist regarding the availability of linguistic corpora, algorithmic censorship, and the ethics of generating compositions derived from copyrighted sources.


【6】Multi-Tast Transformer for Explainable Speech Deepfake Detection via Formant Modeling
标题:通过伪造建模用于可解释语音深度伪造检测的多口味Transformer
链接:https://arxiv.org/abs/2601.14850

作者:Viola Negroni,Luca Cuccovillo,Paolo Bestagini,Patrick Aichroth,Stefano Tubaro
备注:Accepted @ IEEE ICASSP 2026
摘要:在这项工作中,我们引入了一个用于语音deepfake检测的多任务Transformer,能够随着时间的推移预测共振峰轨迹和发声模式,最终将语音分类为真实或虚假,并突出显示其决策是否更多地依赖于有声或无声区域。基于先前的扬声器共振峰Transformer架构,我们使用改进的输入分割策略简化了模型,重新设计了解码过程,并集成了内置的可解释性。与基线相比,我们的模型需要更少的参数,训练速度更快,并提供更好的可解释性,而不会牺牲预测性能。
摘要:In this work, we introduce a multi-task transformer for speech deepfake detection, capable of predicting formant trajectories and voicing patterns over time, ultimately classifying speech as real or fake, and highlighting whether its decisions rely more on voiced or unvoiced regions. Building on a prior speaker-formant transformer architecture, we streamline the model with an improved input segmentation strategy, redesign the decoding process, and integrate built-in explainability. Compared to the baseline, our model requires fewer parameters, trains faster, and provides better interpretability, without sacrificing prediction performance.


【7】Training-Efficient Text-to-Music Generation with State-Space Modeling
标题:使用状态空间建模的训练高效文本到音乐生成
链接:https://arxiv.org/abs/2601.14786

作者:Wei-Jaw Lee,Fang-Chih Hsieh,Xuanjun Chen,Fang-Duo Tsai,Yi-Hsuan Yang
备注:9 pages, 3 figures. This is a preprint of a paper submitted to IEEE/ACM TASLP
摘要:文本到音乐生成(TTM)的最新进展已经产生了高质量的结果,但通常以大量计算和使用大量专有内部数据为代价。为了提高TTM培训的可负担性和开放性,需要一个更有效的培训和数据的开源生成模型主干。在本文中,我们限制生成模型中可训练参数的数量,以匹配MusicGen小型基准测试(约300 M参数),并将其Transformer骨干替换为新兴的状态空间模型(SSM)。具体来说,我们探索不同的SSM变体的序列建模,并比较一个单阶段的SSM为基础的设计与可分解的两阶段SSM/扩散混合设计。所有提出的模型都是在纯公共数据集上从头开始训练的,该数据集包含457小时的CC许可音乐,确保完全开放。我们的实验结果有三个方面。首先,我们表明,SSM表现出优越的训练效率相比,Transformer对应。其次,尽管与MusicGen-small基准相比,我们的模型只使用了9%的FLOP和2%的训练数据大小,但在基于MusicCaps字幕的客观指标和主观听力测试中,我们的模型都取得了有竞争力的性能。最后,我们的缩小实验表明,当模型大小减少到四分之一时,SSM可以保持相对于Transformer基线的竞争性能,即使在相同的训练预算(以迭代为单位测量)。为了促进TTM研究的民主化,经过处理的标题,模型检查点和源代码可以通过项目页面在GitHub上获得:https://lonian6.github.io/ssmttm/。
摘要:Recent advances in text-to-music generation (TTM) have yielded high-quality results, but often at the cost of extensive compute and the use of large proprietary internal data. To improve the affordability and openness of TTM training, an open-source generative model backbone that is more training- and data-efficient is needed. In this paper, we constrain the number of trainable parameters in the generative model to match that of the MusicGen-small benchmark (with about 300M parameters), and replace its Transformer backbone with the emerging class of state-space models (SSMs). Specifically, we explore different SSM variants for sequence modeling, and compare a single-stage SSM-based design with a decomposable two-stage SSM/diffusion hybrid design. All proposed models are trained from scratch on a purely public dataset comprising 457 hours of CC-licensed music, ensuring full openness. Our experimental findings are three-fold. First, we show that SSMs exhibit superior training efficiency compared to the Transformer counterpart. Second, despite using only 9% of the FLOPs and 2% of the training data size compared to the MusicGen-small benchmark, our model achieves competitive performance in both objective metrics and subjective listening tests based on MusicCaps captions. Finally, our scaling-down experiment demonstrates that SSMs can maintain competitive performance relative to the Transformer baseline even at the same training budget (measured in iterations), when the model size is reduced to four times smaller. To facilitate the democratization of TTM research, the processed captions, model checkpoints, and source code are available on GitHub via the project page: https://lonian6.github.io/ssmttm/.


【8】Unlocking Large Audio-Language Models for Interactive Language Learning
标题:解锁交互式语言学习的大型音频语言模型
链接:https://arxiv.org/abs/2601.14744

作者:Hongfu Liu,Zhouying Cui,Xiangming Gu,Ye Wang
备注:Accepted to the Findings of EACL 2026
摘要:尽管计算机辅助发音训练(CAPT)系统不断发展,但在第二语言(L2)中实现发音熟练仍然是一个挑战。传统的CAPT系统通常提供不直观的反馈,缺乏可操作的指导,限制了其有效性。音频语言模型(ALMs)的最新进展提供了通过提供更用户友好的反馈来增强这些系统的潜力。在这项工作中,我们通过引入L2-Arctic-plus来研究基于聊天的发音训练的ALM,L2-Arctic-plus是一个英语数据集,具有详细的错误解释和可操作的改进建议。我们在此数据集上对级联的ASR+ LLM和现有的ALM进行基准测试,特别是在检测发音错误和生成可操作的反馈方面。为了提高性能,我们进一步提出在L2-Arctic-plus上对ALM进行预调优。实验结果表明,我们的发音调整模型显着优于现有的基准在错误发音检测和建议生成方面的客观和人类的评价,突出了所提出的数据集的价值。
摘要:Achieving pronunciation proficiency in a second language (L2) remains a challenge, despite the development of Computer-Assisted Pronunciation Training (CAPT) systems. Traditional CAPT systems often provide unintuitive feedback that lacks actionable guidance, limiting its effectiveness. Recent advancements in audio-language models (ALMs) offer the potential to enhance these systems by providing more user-friendly feedback. In this work, we investigate ALMs for chat-based pronunciation training by introducing L2-Arctic-plus, an English dataset with detailed error explanations and actionable suggestions for improvement. We benchmark cascaded ASR+LLMs and existing ALMs on this dataset, specifically in detecting mispronunciation and generating actionable feedback. To improve the performance, we further propose to instruction-tune ALMs on L2-Arctic-plus. Experimental results demonstrate that our instruction-tuned models significantly outperform existing baselines on mispronunciation detection and suggestion generation in terms of both objective and human evaluation, highlighting the value of the proposed dataset.


【9】Dissecting Performance Degradation in Audio Source Separation under Sampling Frequency Mismatch
标题:剖析采样频率不匹配下音频源分离的性能下降
链接:https://arxiv.org/abs/2601.14684

作者:Kanami Imamura,Tomohiko Nakamura,Kohei Yatabe,Hiroshi Saruwatari
备注:Accepted for ICASSP 2026
摘要:基于深度神经网络的音频处理方法通常以单个采样频率(SF)进行训练。为了处理未经训练的SF,通常采用信号恢复,但它会降低性能,特别是当输入SF低于经训练的SF时。本文通过两个假设来研究这种退化的原因:(i)缺乏上采样引入的高频分量,以及(ii)它们的存在比它们的精确表示更重要。为了检验这些假设,我们比较了传统的复位与三种替代方案:复位后噪声添加,它将高斯噪声添加到重新采样的信号;噪声内核复位,它用高斯噪声扰动内核以丰富高频分量;和可训练内核复位,它通过训练来适应插值内核。音乐源分离的实验表明,噪声核和可训练核恢复减轻了传统恢复所观察到的退化。我们进一步证明了噪声内核恢复在不同的模型中是有效的,强调它是一个简单而实用的选择。
摘要:Audio processing methods based on deep neural networks are typically trained at a single sampling frequency (SF). To handle untrained SFs, signal resampling is commonly employed, but it can degrade performance, particularly when the input SF is lower than the trained SF. This paper investigates the causes of this degradation through two hypotheses: (i) the lack of high-frequency components introduced by up-sampling, and (ii) the greater importance of their presence than their precise representation. To examine these hypotheses, we compare conventional resampling with three alternatives: post-resampling noise addition, which adds Gaussian noise to the resampled signal; noisy-kernel resampling, which perturbs the kernel with Gaussian noise to enrich high-frequency components; and trainable-kernel resampling, which adapts the interpolation kernel through training. Experiments on music source separation show that noisy-kernel and trainable-kernel resampling alleviate the degradation observed with conventional resampling. We further demonstrate that noisy-kernel resampling is effective across diverse models, highlighting it as a simple yet practical option.


【10】READ-Net: Clarifying Emotional Ambiguity via Adaptive Feature Recalibration for Audio-Visual Depression Detection
标题:Read-Net:通过视听抑郁检测的自适应特征重新校准澄清情感模糊性
链接:https://arxiv.org/abs/2601.14651

作者:Chenglizhao Chen,Boze Li,Mengke Song,Dehao Feng,Xinyu Liu,Shanchen Pang,Jufeng Yang,Hui Yu
备注:12 pages
摘要:抑郁症是一种严重的全球性心理健康问题,会损害日常功能和整体生活质量。尽管最近的视听方法改进了抑郁症的自动检测,但忽视情感线索的方法往往无法捕捉隐藏在情感表达中的微妙抑郁信号。相反,那些包含情绪的人经常将短暂的情绪表达与特征表征中的稳定抑郁症状混淆,这种现象称为情感模糊,从而导致检测错误。为了解决这个关键问题,我们提出了READ-Net,这是第一个明确设计用于通过自适应特征重新校准(AFR)解决情绪模糊的视听抑郁检测框架。AFR的核心思想是动态调整情绪特征的权重,以增强抑郁相关信号。而不是仅仅忽略或天真地组合情感信息,READ-Net创新地识别和保留情感特征中与抑郁相关的线索,同时自适应地过滤掉不相关的情感噪音。这种重新校准策略显著地澄清了特征表示,并有效地减轻了情感干扰的持续挑战。此外,READ-Net可以很容易地集成到现有的框架中,以提高性能。对三个公开数据集的广泛评估表明,READ-Net优于最先进的方法,准确率平均提高4.55%,F1得分平均提高1.26%,证明了其对情绪干扰的鲁棒性,并提高了视听抑郁检测。
摘要:Depression is a severe global mental health issue that impairs daily functioning and overall quality of life. Although recent audio-visual approaches have improved automatic depression detection, methods that ignore emotional cues often fail to capture subtle depressive signals hidden within emotional expressions. Conversely, those incorporating emotions frequently confuse transient emotional expressions with stable depressive symptoms in feature representations, a phenomenon termed \emph{Emotional Ambiguity}, thereby leading to detection errors. To address this critical issue, we propose READ-Net, the first audio-visual depression detection framework explicitly designed to resolve Emotional Ambiguity through Adaptive Feature Recalibration (AFR). The core insight of AFR is to dynamically adjust the weights of emotional features to enhance depression-related signals. Rather than merely overlooking or naively combining emotional information, READ-Net innovatively identifies and preserves depressive-relevant cues within emotional features, while adaptively filtering out irrelevant emotional noise. This recalibration strategy significantly clarifies feature representations, and effectively mitigates the persistent challenge of emotional interference. Additionally, READ-Net can be easily integrated into existing frameworks for improved performance. Extensive evaluations on three publicly available datasets show that READ-Net outperforms state-of-the-art methods, with average gains of 4.55\% in accuracy and 1.26\% in F1-score, demonstrating its robustness to emotional disturbances and improving audio-visual depression detection.


【11】Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum
标题:复频谱中相连贯神经声码的韵律引导的和声注意
链接:https://arxiv.org/abs/2601.14472

作者:Mohammed Salah Al-Radhi,Riad Larbi,Mátyás Bartalis,Géza Németh
备注:5 pages, 2 figures, 1 table. Accepted for presentation at ICASSP 2026
摘要:神经声码器是语音合成的核心;尽管它们取得了成功,但大多数仍然受到有限的韵律建模和不准确的相位重建的影响。我们提出了一种声码器,引入韵律引导谐波注意,以增强有声段编码,并直接预测复杂的频谱成分,通过逆STFT波形合成。与基于梅尔频谱图的方法不同,我们的设计联合建模幅度和相位,确保相位相干性和改善的音调保真度。为了进一步与感知质量保持一致,我们采用了一种多目标训练策略,该策略集成了对抗性、频谱和相位感知损失。在基准数据集上的实验表明,与HiFi-GAN和AutoVocoder相比,F0 RMSE降低了22%,有声/无声错误降低了18%,MOS分数提高了0.15。这些结果表明,韵律引导的注意力与直接复频谱建模相结合,产生更自然,音高准确,鲁棒的合成语音,为表达性神经语音编码奠定了坚实的基础。
摘要:Neural vocoders are central to speech synthesis; despite their success, most still suffer from limited prosody modeling and inaccurate phase reconstruction. We propose a vocoder that introduces prosody-guided harmonic attention to enhance voiced segment encoding and directly predicts complex spectral components for waveform synthesis via inverse STFT. Unlike mel-spectrogram-based approaches, our design jointly models magnitude and phase, ensuring phase coherence and improved pitch fidelity. To further align with perceptual quality, we adopt a multi-objective training strategy that integrates adversarial, spectral, and phase-aware losses. Experiments on benchmark datasets demonstrate consistent gains over HiFi-GAN and AutoVocoder: F0 RMSE reduced by 22 percent, voiced/unvoiced error lowered by 18 percent, and MOS scores improved by 0.15. These results show that prosody-guided attention combined with direct complex spectrum modeling yields more natural, pitch-accurate, and robust synthetic speech, setting a strong foundation for expressive neural vocoding.


【12】Single-step Controllable Music Bandwidth Extension With Flow Matching
标题:具有流量匹配的一步可控音乐带宽扩展
链接:https://arxiv.org/abs/2601.14356

作者:Carlos Hernandez-Olivan,Hendrik Vincent Koops,Hao Hao Tan,Elio Quinton
备注:Accepted at the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2026
摘要:音频恢复包括反转数字音频信号的劣化以恢复劣化发生之前的原始质量信号。这在诸如音乐记录档案的情况下是有价值的,特别是那些具有珍贵历史价值的音乐记录档案,其中干净的版本可能已经丢失或根本不存在。最近的工作将生成模型应用于音频恢复,与以前的方法相比显示出有希望的改进,并为执行以前不可能的恢复操作的能力打开了大门。然而,使这些模型精细可控仍然是一个挑战。在本文中,我们提出了一个扩展的FLowHigh,并介绍了动态谱轮廓(DSC)作为控制信号,通过无分类器的指导带宽扩展。我们的实验显示了有竞争力的模型性能,并表明DSC是一个很有前途的功能,以支持细粒度调节。
摘要:Audio restoration consists in inverting degradations of a digital audio signal to recover what would have been the pristine quality signal before the degradation occurred. This is valuable in contexts such as archives of music recordings, particularly those of precious historical value, for which a clean version may have been lost or simply does not exist. Recent work applied generative models to audio restoration, showing promising improvement over previous methods, and opening the door to the ability to perform restoration operations that were not possible before. However, making these models finely controllable remains a challenge. In this paper, we propose an extension of FLowHigh and introduce the Dynamic Spectral Contour (DSC) as a control signal for bandwidth extension via classifier-free guidance. Our experiments show competitive model performance, and indicate that DSC is a promising feature to support fine-grained conditioning.


【13】Guided by the Plan: Enhancing Faithful Autoregressive Text-to-Audio Generation with Guided Decoding
标题:以计划为指导:通过引导解码增强忠实的自回归文本到音频生成
链接:https://arxiv.org/abs/2601.14304

作者:Juncheng Wang,Zhe Hu,Chao Xu,Siyue Ren,Yuxiang Feng,Yang Liu,Baigui Sun,Shujun Wang
备注:Accepted at EACL 2026
摘要:自回归(AR)模型擅长通过顺序地产生令牌来生成时间相干的音频,但它们经常在忠实地遵循复杂的文本提示,特别是那些描述复杂声音事件的文本提示时犹豫不决。我们在AR音频生成器中发现了一个令人惊讶的能力:它们的早期前缀令牌隐式地编码最终输出的全局语义属性,例如事件计数和声音对象类别,揭示了一种隐式规划的形式。基于这一见解,我们提出了Plan-Critic,这是一种轻量级的辅助模型,采用广义优势估计(GAE)目标进行训练,以预测部分世代的最终预防跟踪质量。在推理时,Plan-Critic支持引导探索:它早期评估候选前缀,修剪低保真度轨迹,并将计算重新分配给高潜力的规划种子。我们的Plan-Critic引导采样在CLAP评分上比AR基线提高了10个点-建立了AR文本到音频生成的新技术水平-同时保持了与标准最佳N解码的计算奇偶性。这项工作弥合了因果生成和全局语义对齐之间的差距,表明即使是严格的自回归模型也可以提前计划。
摘要:Autoregressive (AR) models excel at generating temporally coherent audio by producing tokens sequentially, yet they often falter in faithfully following complex textual prompts, especially those describing complex sound events. We uncover a surprising capability in AR audio generators: their early prefix tokens implicitly encode global semantic attributes of the final output, such as event count and sound-object category, revealing a form of implicit planning. Building on this insight, we propose Plan-Critic, a lightweight auxiliary model trained with a Generalized Advantage Estimation (GAE)-inspired objective to predict final instruction-following quality from partial generations. At inference time, Plan-Critic enables guided exploration: it evaluates candidate prefixes early, prunes low-fidelity trajectories, and reallocates computation to high-potential planning seeds. Our Plan-Critic-guided sampling achieves up to a 10-point improvement in CLAP score over the AR baseline-establishing a new state of the art in AR text-to-audio generation-while maintaining computational parity with standard best-of-N decoding. This work bridges the gap between causal generation and global semantic alignment, demonstrating that even strictly autoregressive models can plan ahead.


【14】Call2Instruct: Automated Pipeline for Generating Q&A Datasets from Call Center Recordings for LLM Fine-Tuning
标题:Call 2 Direcct:从呼叫中心录音生成问答数据集以进行LLM微调的自动化管道
链接:https://arxiv.org/abs/2601.14263

作者:Alex Echeverria,Sávio Salvarino Teles de Oliveira,Fernando Marques Federson
备注:15 pages, 1 figures, conference
摘要:大规模语言模型(LLM)对特定领域的适应取决于高质量的微调数据集,特别是在教学格式(例如,回答- Q&A)。然而,生成这些数据集,特别是从非结构化来源(如呼叫中心音频记录)生成这些数据集,由于数据的噪声和无序性而带来了重大挑战。本文提出了一种解决方案,通过提供一个端到端的自动化管道,从这样的录音生成问答教学数据集。开发的方法包括顺序步骤的音频处理(包括日记化,噪声去除和自动转录),文本处理(清洗,规范化和匿名化),语义提取的客户需求和伴随的响应使用向量嵌入,并通过语义搜索匹配,以形成最终的Q&A对。结果,成功实现了完整的管道,生成了专门为Instruct Fine Tuning格式化的数据集。通过LLM模型(基于Llama 2 7 B)的成功微调,证实了所生成数据集的实用价值和可行性,并在功能上得到了证明。本文的结论指出,所提出的方法是可行的,从呼叫中心的非结构化会话数据转换为宝贵的资源,培训LLM。这一发展有可能为客户服务领域的问答任务创建更有效的人工智能系统开辟道路。开发的代码已公开提供,以促进可重复性和未来的研究。
摘要:The adaptation of Large-Scale Language Models (LLMs) to specific domains depends on high-quality fine-tuning datasets, particularly in instructional format (e.g., Question-Answer - Q&A). However, generating these datasets, particularly from unstructured sources such as call center audio recordings, poses a significant challenge due to the noisy and disorganized nature of the data. This paper presents a solution to this challenge by offering an end-to-end automated pipeline for generating Q&A instructional datasets from such recordings. The methodology developed comprises sequential steps of audio processing (including diarization, noise removal and automatic transcription), textual processing (cleaning, normalization, and anonymization), semantic extraction of customer demands and attendant responses using vector embeddings, and matching via semantic search to form the final Q&A pairs. As a result, the complete pipeline was successfully implemented, generating a dataset specifically formatted for Instruct Fine Tuning. The practical value and feasibility of the generated dataset were substantiated and functionally demonstrated through the successful fine-tuning of an LLM model (based on Llama 2 7B). The conclusion of the paper states that the proposed approach is viable for converting unstructured conversational data from call centers into valuable resources for training LLMs. This development has the potential to open up avenues for creating more effective AI systems for Q&A tasks in the customer service domain. The developed codes have been made publicly available to promote reproducibility and future research.


【15】A Cloud-Based Cross-Modal Transformer for Emotion Recognition and Adaptive Human-Computer Interaction
标题:基于云的跨模态情感识别和自适应人机交互Transformer
链接:https://arxiv.org/abs/2601.14259

作者:Ziwen Zhong,Zhitao Shu,Yue Zhao
摘要:情感识别是下一代人机交互(HCI)的基本组成部分,使机器能够感知,理解和响应用户的情感状态。然而,现有的系统往往依赖于单模态分析,如面部表情,语音语调,或文本情感,导致有限的鲁棒性和在现实世界环境中的泛化能力差。为了应对这些挑战,本研究提出了一个基于云的跨模态Transformer(CMT)框架,用于多模态情感识别和自适应人机交互。该模型使用预训练编码器(Vision Transformer,Wav2Vec2和BERT)集成视觉,听觉和文本信号,并采用跨模态注意力机制来捕获异构特征之间的复杂相互依赖关系。通过利用云计算基础设施,在Kubernetes和TensorFlow Serving上进行分布式训练,该系统可以为大规模用户交互提供可扩展的低延迟情感识别。在包括IEMOCAP、MELD和AffectNet在内的基准数据集上进行的实验表明,CMT实现了最先进的性能,与强大的多模态基线相比,F1分数提高了3.0%,交叉熵损失降低了12.9%。此外,云部署评估显示平均响应延迟为128 ms,与传统的基于transformer的融合系统相比减少了35%。这些结果证实,所提出的框架,使有效的,实时的情感识别和自适应反馈的应用程序,如智能客户服务,虚拟辅导系统,情感计算接口,标志着一个重要的一步,云原生情感计算和情感智能交互系统。
摘要:Emotion recognition is a fundamental component of next-generation human-computer interaction (HCI), enabling machines to perceive, understand, and respond to users' affective states. However, existing systems often rely on single-modality analysis such as facial expressions, speech tone, or textual sentiment, resulting in limited robustness and poor generalization in real-world environments. To address these challenges, this study proposes a Cloud-Based Cross-Modal Transformer (CMT) framework for multimodal emotion recognition and adaptive human-computer interaction. The proposed model integrates visual, auditory, and textual signals using pretrained encoders (Vision Transformer, Wav2Vec2, and BERT) and employs a cross-modal attention mechanism to capture complex interdependencies among heterogeneous features. By leveraging cloud computing infrastructure with distributed training on Kubernetes and TensorFlow Serving, the system enables scalable, low-latency emotion recognition for large-scale user interactions. Experiments conducted on benchmark datasets including IEMOCAP, MELD, and AffectNet demonstrate that the CMT achieves state-of-the-art performance, improving the F1-score by 3.0 percent and reducing cross-entropy loss by 12.9 percent compared to strong multimodal baselines. Additionally, cloud deployment evaluations show an average response latency of 128 ms, representing a 35 percent reduction compared with conventional transformer-based fusion systems. These results confirm that the proposed framework enables efficient, real-time emotion recognition and adaptive feedback in applications such as intelligent customer service, virtual tutoring systems, and affective computing interfaces, marking an important step toward cloud-native affective computing and emotionally intelligent interactive systems.


【16】Neural Tracking of Sustained Attention, Attention Switching, and Natural Conversation in Audiovisual Environments using Mobile EEG
标题:使用移动EEG的视听环境中持续注意、注意转换和自然对话的神经跟踪
链接:https://arxiv.org/abs/2601.15097

作者:Johanna Wilroth,Oskar Keding,Martin A. Skoglund,Maria Sandsten,Martin Enqvist,Emina Alickovic
备注:Submitted to European Journal of Neuroscience
摘要:日常交流是动态和多感官的,通常涉及转移注意力,重叠的语音和视觉提示。然而,大多数神经注意力追踪研究仍然局限于高度控制的实验室环境,使用干净的,通常只有音频的刺激,并需要持续关注单个说话者。这项工作通过引入来自24名正常听力参与者的新数据集来解决这一差距。我们使用了一个移动脑电图(EEG)系统(44头皮电极和20 cEEGrid电极)在视听(AV)范式与三个条件:持续注意到一个单一的谈话者在两个谈话者的环境中,注意两个谈话者之间的切换,和无脚本的两个谈话者对话与竞争的单一谈话者。分析包括时间响应函数(TRFs)建模,最佳滞后分析,选择性注意力分类与决策窗口范围从1.1秒到35秒,和TRFs的AV对话与侧音频只说话者的注意力的比较。主要研究结果显示,在头皮脑电图的条件下,注意力相关的P2峰在注意和忽略的语音之间存在显着差异。切换和持续注意之间的性能没有显着变化,表明注意切换的鲁棒性。最佳滞后分析显示,窄峰会话相比,单说话AV刺激,反映了额外的复杂性多说话处理。头皮EEG的选择性注意力分类始终高于偶然性(55-70%的准确性),而cEEGrid数据产生的相关性较低,突出了进一步改进方法的必要性。这些结果表明,移动EEG可以可靠地跟踪选择性注意力的动态,多感官听觉场景,并为设计未来的AV范例和现实世界的注意力跟踪应用提供指导。
摘要:Everyday communication is dynamic and multisensory, often involving shifting attention, overlapping speech and visual cues. Yet, most neural attention tracking studies are still limited to highly controlled lab settings, using clean, often audio-only stimuli and requiring sustained attention to a single talker. This work addresses that gap by introducing a novel dataset from 24 normal-hearing participants. We used a mobile electroencephalography (EEG) system (44 scalp electrodes and 20 cEEGrid electrodes) in an audiovisual (AV) paradigm with three conditions: sustained attention to a single talker in a two-talker environment, attention switching between two talkers, and unscripted two-talker conversations with a competing single talker. Analysis included temporal response functions (TRFs) modeling, optimal lag analysis, selective attention classification with decision windows ranging from 1.1s to 35s, and comparisons of TRFs for attention to AV conversations versus side audio-only talkers. Key findings show significant differences in the attention-related P2-peak between attended and ignored speech across conditions for scalp EEG. No significant change in performance between switching and sustained attention suggests robustness for attention switches. Optimal lag analysis revealed narrower peak for conversation compared to single-talker AV stimuli, reflecting the additional complexity of multi-talker processing. Classification of selective attention was consistently above chance (55-70% accuracy) for scalp EEG, while cEEGrid data yielded lower correlations, highlighting the need for further methodological improvements. These results demonstrate that mobile EEG can reliably track selective attention in dynamic, multisensory listening scenarios and provide guidance for designing future AV paradigms and real-world attention tracking applications.


【17】AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering
标题:AQAScore:通过音频问题回答评估文本到音频生成中的语义一致性
链接:https://arxiv.org/abs/2601.14728

作者:Chun-Yi Kuan,Kai-Wei Chang,Hung-yi Lee
备注:Manuscript in progress
摘要:虽然文本到音频生成在真实性和多样性方面取得了显着进展,但评估指标的发展却没有跟上步伐。广泛采用的方法,通常基于嵌入相似性,如CLAPScore,有效地测量一般相关性,但在细粒度语义对齐和组合推理方面仍然有限。为了解决这个问题,我们引入AQAScore,一个骨干不可知的评估框架,利用音频感知的大型语言模型(ALLM)的推理能力。AQAScore将评估重新定义为概率语义验证任务;而不是依赖于开放式文本生成,它通过计算目标语义查询的“是”答案的精确对数概率来估计对齐。我们在多个基准中评估AQAScore,包括人类评分相关性,成对比较和组合推理任务。实验结果表明,AQAScore始终实现更高的相关性与人类的判断比基于相似性的指标和生成提示基线,显示其有效性,在捕捉微妙的语义不一致和缩放的能力,底层ALLM。
摘要:Although text-to-audio generation has made remarkable progress in realism and diversity, the development of evaluation metrics has not kept pace. Widely-adopted approaches, typically based on embedding similarity like CLAPScore, effectively measure general relevance but remain limited in fine-grained semantic alignment and compositional reasoning. To address this, we introduce AQAScore, a backbone-agnostic evaluation framework that leverages the reasoning capabilities of audio-aware large language models (ALLMs). AQAScore reformulates assessment as a probabilistic semantic verification task; rather than relying on open-ended text generation, it estimates alignment by computing the exact log-probability of a "Yes" answer to targeted semantic queries. We evaluate AQAScore across multiple benchmarks, including human-rated relevance, pairwise comparison, and compositional reasoning tasks. Experimental results show that AQAScore consistently achieves higher correlation with human judgments than similarity-based metrics and generative prompting baselines, showing its effectiveness in capturing subtle semantic inconsistencies and scaling with the capability of underlying ALLMs.


【18】Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
标题:缩放模糊性:用音频语言模型增强语音情感识别中的人类注释
链接:https://arxiv.org/abs/2601.14620

作者:Wenda Zhang,Hongyu Jin,Siyi Wang,Zhiqiang Wei,Ting Dang
备注:Accepted by ICASSP 2026
摘要:语音情感识别模型通常使用单一的分类标签,忽略了人类情感的固有模糊性。模糊情绪识别通过将情绪表示为概率分布来解决这一问题,但进展受到从稀疏的人类注释推断出的不可靠的地面真实分布的限制。本文探讨了大型音频语言模型(ALMs)是否可以通过生成高质量的合成注释来缓解注释瓶颈。我们引入了一个框架,利用ALM创建合成感知代理,增强人类注释,以提高地面实况分布的可靠性。我们验证这些代理通过统计分析他们的对齐与人类的分布和评估其影响,通过微调ALM与增强的情感分布。此外,为了解决类不平衡和实现无偏评估,我们提出了DiME-Aug,一种分布感知的多模态情感增强策略。IEMOCAP和MSP-Podcast上的实验表明,合成注释增强了情感分布,特别是在注释一致性高的低歧义区域。然而,对于高度模糊的情感,随着人类分歧的增加,好处会减少。这项工作提供了第一个证据,ALMs可以解决模糊的情感识别中的注释不足,但强调需要更先进的提示或生成策略来处理高度模糊的情况。
摘要:Speech Emotion Recognition models typically use single categorical labels, overlooking the inherent ambiguity of human emotions. Ambiguous Emotion Recognition addresses this by representing emotions as probability distributions, but progress is limited by unreliable ground-truth distributions inferred from sparse human annotations. This paper explores whether Large Audio-Language Models (ALMs) can mitigate the annotation bottleneck by generating high-quality synthetic annotations. We introduce a framework leveraging ALMs to create Synthetic Perceptual Proxies, augmenting human annotations to improve ground-truth distribution reliability. We validate these proxies through statistical analysis of their alignment with human distributions and evaluate their impact by fine-tuning ALMs with the augmented emotion distributions. Furthermore, to address class imbalance and enable unbiased evaluation, we propose DiME-Aug, a Distribution-aware Multimodal Emotion Augmentation strategy. Experiments on IEMOCAP and MSP-Podcast show that synthetic annotations enhance emotion distribution, especially in low-ambiguity regions where annotation agreement is high. However, benefits diminish for highly ambiguous emotions with greater human disagreement. This work provides the first evidence that ALMs could address annotation scarcity in ambiguous emotion recognition, but highlights the need for more advanced prompting or generation strategies to handle highly ambiguous cases.


【19】Towards noise-robust speech inversion through multi-task learning with speech enhancement
标题:通过具有语音增强的多任务学习实现噪音稳健的语音倒置
链接:https://arxiv.org/abs/2601.14516

作者:Saba Tabatabaee,Carol Espy-Wilson
备注:Accepted for presentation at ICASSP 2026
摘要:最近的研究表明,自监督学习(SSL)的语音表示语音反转(SI)的有效性。然而,由于背景噪声的普遍存在,在现实世界的场景中应用SI仍然具有挑战性。我们提出了一个统一的框架,集成语音增强(SE)和SI模型,通过共享的基于SSL的语音表示。在这个框架中,SSL模型不仅被训练来支持SE模块抑制噪声,而且还产生对SI任务更具信息性的表示,从而使两个模块都能从联合训练中受益。在-5分贝的信噪比下,我们的SI任务方法在多路重合噪声下实现了80.95%的基线相对改善,在非多路重合噪声下实现了38.98%的基线相对改善,如通过所有估计参数的平均Pearson积矩相关性所测量的。
摘要:Recent studies demonstrate the effectiveness of Self Supervised Learning (SSL) speech representations for Speech Inversion (SI). However, applying SI in real-world scenarios remains challenging due to the pervasive presence of background noise. We propose a unified framework that integrates Speech Enhancement (SE) and SI models through shared SSL-based speech representations. In this framework, the SSL model is trained not only to support the SE module in suppressing noise but also to produce representations that are more informative for the SI task, allowing both modules to benefit from joint training. At a Signal-to-Noise Ratio of -5 db, our method for the SI task achieves relative improvements over the baseline of 80.95% under babble noise and 38.98% under non-babble noise, as measured by the average Pearson product-moment correlation across all estimated parameters.


【20】Synthetic Singers: A Review of Deep-Learning-based Singing Voice Synthesis Approaches
标题:合成歌手:基于深度学习的歌唱声音合成方法回顾
链接:https://arxiv.org/abs/2601.13910

作者:Changhao Pan,Dongyu Yao,Yu Zhang,Wenxiang Guo,Jingyu Lu,Zhiyuan Zhu,Zhou Zhao
备注:Accepetd by IJCNLP-AACL 2025(Oral)
摘要:歌唱声合成技术的最新进展引起了学术界和工业界的广泛关注。随着大型语言模型和新的生成范式的出现,产生可控的、高保真的歌唱声音已经成为一个可以实现的目标。然而,该领域仍然缺乏系统分析基于深度学习的歌唱声音合成系统及其使能技术的全面调查。为了解决上述问题,本调查首先按任务类型对现有系统进行分类,然后将当前架构组织成两个主要范例:级联和端到端方法。此外,我们提供了一个深入的分析核心技术,涵盖歌唱建模和控制技术。最后,我们回顾了支持培训和评估的相关数据集、注释工具和评估基准。在附录中,我们介绍了SVS的培训策略和进一步的讨论。本文综述了国内外关于SVS模型的最新研究成果,为研究人员和工程技术人员提供了有益的参考。相关材料可在https://github.com/David-Pigeon/SyntheticSingers上查阅。
摘要:Recent advances in singing voice synthesis (SVS) have attracted substantial attention from both academia and industry. With the advent of large language models and novel generative paradigms, producing controllable, high-fidelity singing voices has become an attainable goal. Yet the field still lacks a comprehensive survey that systematically analyzes deep-learning-based singing voice synthesis systems and their enabling technologies. To address the aforementioned issue, this survey first categorizes existing systems by task type and then organizes current architectures into two major paradigms: cascaded and end-to-end approaches. Moreover, we provide an in-depth analysis of core technologies, covering singing modeling and control techniques. Finally, we review relevant datasets, annotation tools, and evaluation benchmarks that support training and assessment. In appendix, we introduce training strategies and further discussion of SVS. This survey provides an up-to-date review of the literature on SVS models, which would be a useful reference for both researchers and engineers. Related materials are available at https://github.com/David-Pigeon/SyntheticSingers.


eess.AS音频处理


【1】Neural Tracking of Sustained Attention, Attention Switching, and Natural Conversation in Audiovisual Environments using Mobile EEG
标题:使用移动EEG的视听环境中持续注意、注意转换和自然对话的神经跟踪
链接:https://arxiv.org/abs/2601.15097

作者:Johanna Wilroth,Oskar Keding,Martin A. Skoglund,Maria Sandsten,Martin Enqvist,Emina Alickovic
备注:Submitted to European Journal of Neuroscience
摘要:日常交流是动态和多感官的,通常涉及转移注意力,重叠的语音和视觉提示。然而,大多数神经注意力追踪研究仍然局限于高度控制的实验室环境,使用干净的,通常只有音频的刺激,并需要持续关注单个说话者。这项工作通过引入来自24名正常听力参与者的新数据集来解决这一差距。我们使用了一个移动脑电图(EEG)系统(44头皮电极和20 cEEGrid电极)在视听(AV)范式与三个条件:持续注意到一个单一的谈话者在两个谈话者的环境中,注意两个谈话者之间的切换,和无脚本的两个谈话者对话与竞争的单一谈话者。分析包括时间响应函数(TRFs)建模,最佳滞后分析,选择性注意力分类与决策窗口范围从1.1秒到35秒,和TRFs的AV对话与侧音频只说话者的注意力的比较。主要研究结果显示,在头皮脑电图的条件下,注意力相关的P2峰在注意和忽略的语音之间存在显着差异。切换和持续注意之间的性能没有显着变化,表明注意切换的鲁棒性。最佳滞后分析显示,窄峰会话相比,单说话AV刺激,反映了额外的复杂性多说话处理。头皮EEG的选择性注意力分类始终高于偶然性(55-70%的准确性),而cEEGrid数据产生的相关性较低,突出了进一步改进方法的必要性。这些结果表明,移动EEG可以可靠地跟踪动态、多感官听力场景中的选择性注意力,并为设计未来的AV范例和现实世界的注意力跟踪应用程序提供指导。
摘要:Everyday communication is dynamic and multisensory, often involving shifting attention, overlapping speech and visual cues. Yet, most neural attention tracking studies are still limited to highly controlled lab settings, using clean, often audio-only stimuli and requiring sustained attention to a single talker. This work addresses that gap by introducing a novel dataset from 24 normal-hearing participants. We used a mobile electroencephalography (EEG) system (44 scalp electrodes and 20 cEEGrid electrodes) in an audiovisual (AV) paradigm with three conditions: sustained attention to a single talker in a two-talker environment, attention switching between two talkers, and unscripted two-talker conversations with a competing single talker. Analysis included temporal response functions (TRFs) modeling, optimal lag analysis, selective attention classification with decision windows ranging from 1.1s to 35s, and comparisons of TRFs for attention to AV conversations versus side audio-only talkers. Key findings show significant differences in the attention-related P2-peak between attended and ignored speech across conditions for scalp EEG. No significant change in performance between switching and sustained attention suggests robustness for attention switches. Optimal lag analysis revealed narrower peak for conversation compared to single-talker AV stimuli, reflecting the additional complexity of multi-talker processing. Classification of selective attention was consistently above chance (55-70% accuracy) for scalp EEG, while cEEGrid data yielded lower correlations, highlighting the need for further methodological improvements. These results demonstrate that mobile EEG can reliably track selective attention in dynamic, multisensory listening scenarios and provide guidance for designing future AV paradigms and real-world attention tracking applications.


【2】Fast-ULCNet: A fast and ultra low complexity network for single-channel speech enhancement
标题:Fast-ULCNet:用于单通道语音增强的快速且超低复杂度的网络
链接:https://arxiv.org/abs/2601.14925

作者:Nicolás Arrieta Larraza,Niels de Koeijer
备注:©2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:单通道语音增强算法通常用于资源受限的嵌入式设备中,其中低延迟和低复杂度设计变得更加重要。近年来,研究人员提出了各种各样的新的解决方案,这个问题。特别是,最近的一个名为ULCNet的深度学习模型是该领域最先进的方法之一。本文提出了一种ULCNet的自适应方法,用FastGRNN替换其GRU层,以减少计算延迟和复杂性。此外,本文还展示了FastGRNN在长音频信号推理过程中由于内部状态漂移而导致的性能衰减的经验证据,并提出了一种基于可训练互补滤波器的新方法来缓解它。由此产生的模型Fast-ULCNet在语音增强任务上与最先进的原始ULCNet架构表现相当,同时将其模型大小减少了一半以上,并将其延迟平均减少了34%。
摘要:Single-channel speech enhancement algorithms are often used in resource-constrained embedded devices, where low latency and low complexity designs gain more importance. In recent years, researchers have proposed a wide variety of novel solutions to this problem. In particular, a recent deep learning model named ULCNet is among the state-of-the-art approaches in this domain. This paper proposes an adaptation of ULCNet, by replacing its GRU layers with FastGRNNs, to reduce both computational latency and complexity. Furthermore, this paper shows empirical evidence on the performance decay of FastGRNNs in long audio signals during inference due to internal state drifting, and proposes a novel approach based on a trainable complementary filter to mitigate it. The resulting model, Fast-ULCNet, performs on par with the state-of-the-art original ULCNet architecture on a speech enhancement task, while reducing its model size by more than half and decreasing its latency by 34% on average.


【3】Test-Time Adaptation For Speech Enhancement Via Mask Polarization
标题:通过屏蔽极化实现语音增强的测试时间自适应
链接:https://arxiv.org/abs/2601.14770

作者:Tobias Raichle,Erfan Amini,Bin Yang
备注:Accepted at ICASSP 2026
摘要:使语音增强(SE)模型适应看不见的环境对于实际部署至关重要,但由于缺乏对SE模型在域转移下如何降级的理解,SE的测试时自适应(TTA)在很大程度上仍然未得到充分探索。我们观察到,基于掩码的SE模型在域偏移下失去了信心,预测的掩码变得平坦,失去了决定性的语音保留和噪声抑制。基于这一见解,我们提出了掩模极化(MPol),一个轻量级的TTA方法,恢复掩模双峰分布比较使用Wasserstein距离。MPol不需要训练模型之外的其他参数,因此适合资源受限的边缘部署。不同领域的变化和架构的实验结果表明,MPol实现了非常一致的收益,具有竞争力的显着更复杂的方法。
摘要:Adapting speech enhancement (SE) models to unseen environments is crucial for practical deployments, yet test-time adaptation (TTA) for SE remains largely under-explored due to a lack of understanding of how SE models degrade under domain shifts. We observe that mask-based SE models lose confidence under domain shifts, with predicted masks becoming flattened and losing decisive speech preservation and noise suppression. Based on this insight, we propose mask polarization (MPol), a lightweight TTA method that restores mask bimodality through distribution comparison using the Wasserstein distance. MPol requires no additional parameters beyond the trained model, making it suitable for resource-constrained edge deployments. Experimental results across diverse domain shifts and architectures demonstrate that MPol achieves very consistent gains that are competitive with significantly more complex approaches.


【4】Inverse-Hessian Regularization for Continual Learning in ASR
标题:ASB中连续学习的逆黑森正规化
链接:https://arxiv.org/abs/2601.14751

作者:Steven Vander Eeckt,Hugo Van hamme
备注:Accepted for presentation at ICASSP 2026
摘要:灾难性遗忘仍然是自动语音识别(ASR)中持续学习(CL)的主要挑战,其中模型必须适应新的领域,而不会在先前学习的条件下失去性能。已经提出了几种用于ASR的CL方法,最近,权重平均-其中模型在微调后在合并步骤中被平均-已被证明是一种简单的无记忆策略。然而,它本质上是启发式的,忽略了任务的潜在损失景观,阻碍了适应性。在这项工作中,我们提出了逆Hessian正则化(IHR),这是一种用于ASR中CL的无记忆方法,它将曲率信息纳入合并步骤。在对新任务进行微调后,通过对前一个任务进行Kronecker因子逆Hessian近似来调整自适应,确保模型主要朝着对过去性能危害较小的方向移动,同时保持方法的轻量化。我们在两个CL基准上对IHR进行了评估,结果表明它的性能显著优于最先进的基线,在提高适应性的同时减少了遗忘。消融研究和分析进一步证实了其有效性。
摘要:Catastrophic forgetting remains a major challenge for continual learning (CL) in automatic speech recognition (ASR), where models must adapt to new domains without losing performance on previously learned conditions. Several CL methods have been proposed for ASR, and, recently, weight averaging - where models are averaged in a merging step after fine-tuning - has proven effective as a simple memory-free strategy. However, it is heuristic in nature and ignores the underlying loss landscapes of the tasks, hindering adaptability. In this work, we propose Inverse Hessian Regularization (IHR), a memory-free approach for CL in ASR that incorporates curvature information into the merging step. After fine-tuning on a new task, the adaptation is adjusted through a Kronecker-factored inverse Hessian approximation of the previous task, ensuring that the model moves primarily in directions less harmful to past performance, while keeping the method lightweight. We evaluate IHR on two CL benchmarks and show that it significantly outperforms state-of-the-art baselines, reducing forgetting while improving adaptability. Ablation studies and analyses further confirm its effectiveness.


【5】AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering
标题:AQAScore:通过音频问题回答评估文本到音频生成中的语义一致性
链接:https://arxiv.org/abs/2601.14728

作者:Chun-Yi Kuan,Kai-Wei Chang,Hung-yi Lee
备注:Manuscript in progress
摘要:虽然文本到音频生成在真实性和多样性方面取得了显着进展,但评估指标的发展却没有跟上步伐。广泛采用的方法,通常基于嵌入相似性,如CLAPScore,有效地测量一般相关性,但在细粒度语义对齐和组合推理方面仍然有限。为了解决这个问题,我们引入AQAScore,一个骨干不可知的评估框架,利用音频感知的大型语言模型(ALLM)的推理能力。AQAScore将评估重新定义为概率语义验证任务;而不是依赖于开放式文本生成,它通过计算目标语义查询的“是”答案的精确对数概率来估计对齐。我们在多个基准中评估AQAScore,包括人类评分相关性,成对比较和组合推理任务。实验结果表明,AQAScore始终实现更高的相关性与人类的判断比基于相似性的指标和生成提示基线,显示其有效性,在捕捉微妙的语义不一致和缩放的能力,底层ALLM。
摘要:Although text-to-audio generation has made remarkable progress in realism and diversity, the development of evaluation metrics has not kept pace. Widely-adopted approaches, typically based on embedding similarity like CLAPScore, effectively measure general relevance but remain limited in fine-grained semantic alignment and compositional reasoning. To address this, we introduce AQAScore, a backbone-agnostic evaluation framework that leverages the reasoning capabilities of audio-aware large language models (ALLMs). AQAScore reformulates assessment as a probabilistic semantic verification task; rather than relying on open-ended text generation, it estimates alignment by computing the exact log-probability of a "Yes" answer to targeted semantic queries. We evaluate AQAScore across multiple benchmarks, including human-rated relevance, pairwise comparison, and compositional reasoning tasks. Experimental results show that AQAScore consistently achieves higher correlation with human judgments than similarity-based metrics and generative prompting baselines, showing its effectiveness in capturing subtle semantic inconsistencies and scaling with the capability of underlying ALLMs.


【6】NLP-Based Review for Toxic Comment Detection Tailored to the Chinese Cyberspace
标题:针对中国网络空间量身定制的基于NLP的有毒评论检测审查
链接:https://arxiv.org/abs/2601.14721

作者:Ruixing Ren,Junhui Zhao,Xiaoke Sun,Qiuping Li
备注:20 pages, 6 figures. This review focuses on toxic comment detection in Chinese cyberspace
摘要:随着移动互联网的深度融合和社交平台的广泛应用,中国网络空间的用户生成内容呈现爆炸式增长。其中,有毒言论的泛滥对个人心理健康、社区氛围和社会信任都构成了严峻挑战。由于汉语网络语言具有较强的语境依赖性、文化特殊性和快速演变性,有毒词语往往以谐音、隐喻等复杂形式表达,传统检测方法存在明显局限性。针对这一问题,本文围绕中文网络空间中基于自然语言处理的有害评论检测这一核心课题,系统梳理和批判性分析了该领域的研究进展和面临的主要挑战。本文首先界定了中文有害评论的内涵和特征,分析了其所依托的平台生态和传播机制,然后综合评述了现有公共数据集的构建方法和局限性,提出了一种新的细粒度、可扩展的有害评论定义和分类框架,并给出了相应的数据标注和质量评估策略。我们系统地总结了检测模型从传统方法到深度学习的进化路径,特别强调了可解释性在模型设计中的重要性。最后,我们深入讨论了当前研究所面临的挑战,并为未来的研究方向提供了前瞻性的建议。
摘要:With the in-depth integration of mobile Internet and widespread adoption of social platforms, user-generated content in the Chinese cyberspace has witnessed explosive growth. Among this content, the proliferation of toxic comments poses severe challenges to individual mental health, community atmosphere and social trust. Owing to the strong context dependence, cultural specificity and rapid evolution of Chinese cyber language, toxic expressions are often conveyed through complex forms such as homophones and metaphors, imposing notable limitations on traditional detection methods. To address this issue, this review focuses on the core topic of natural language processing based toxic comment detection in the Chinese cyberspace, systematically collating and critically analyzing the research progress and key challenges in this field. This review first defines the connotation and characteristics of Chinese toxic comments, and analyzes the platform ecology and transmission mechanisms they rely on. It then comprehensively reviews the construction methods and limitations of existing public datasets, and proposes a novel fine-grained and scalable framework for toxic comment definition and classification, along with corresponding data annotation and quality assessment strategies. We systematically summarize the evolutionary path of detection models from traditional methods to deep learning, with special emphasis on the importance of interpretability in model design. Finally, we thoroughly discuss the open challenges faced by current research and provide forward-looking suggestions for future research directions.


【7】Triage knowledge distillation for speaker verification
标题:用于说话人验证的分类知识提炼
链接:https://arxiv.org/abs/2601.14699

作者:Ju-ho Kim,Youngmoon Jung,Joon-Young Yang,Jaeyoung Roh,Chang Woo Han,Hoon-Young Cho
备注:5 pages, 2 figures, Accepted at ICASSP 2026
摘要:由于高容量模型的计算成本,在资源受限的设备上部署说话人验证仍然具有挑战性;知识蒸馏(KD)提供了一种补救措施。经典KD在Kullback-Leibler项中将目标置信度与非目标结构纠缠在一起,限制了关系信息的传递。解耦KD将这些信号分为目标和非目标项,但统一对待非目标,并且在大类设置中仍然容易受到低概率类的长尾影响。我们介绍了分类KD(TRKD),一个蒸馏方案,可操作的评估优先重点。TRKD引入累积概率截止$τ$来评估每个示例的难度,并将教师后验分为三组:目标类,高概率非目标混淆集和背景集。为了区分信息信号的优先级,TRKD提取混淆集条件分布并丢弃背景。同时,它传递一个三质量(目标/混淆/背景),捕获样本难度和类间混淆。最后,TRKD通过课程将学习集中在$τ$上:训练从一个更大的$τ$开始,以传达广泛的非目标上下文,然后逐渐减少$τ$以缩小混淆集,将监督集中在最容易混淆的类上。在对VoxCeleb 1进行的同质和异质师生对的广泛实验中,TRKD始终优于最近的KD变体,并在所有协议中达到最低的EER。
摘要:Deploying speaker verification on resource-constrained devices remains challenging due to the computational cost of high-capacity models; knowledge distillation (KD) offers a remedy. Classical KD entangles target confidence with non-target structure in a Kullback-Leibler term, limiting the transfer of relational information. Decoupled KD separates these signals into target and non-target terms, yet treats non-targets uniformly and remains vulnerable to the long tail of low-probability classes in large-class settings. We introduce Triage KD (TRKD), a distillation scheme that operationalizes assess-prioritize-focus. TRKD introduces a cumulative-probability cutoff $τ$ to assess per-example difficulty and partition the teacher posterior into three groups: the target class, a high-probability non-target confusion-set, and a background-set. To prioritize informative signals, TRKD distills the confusion-set conditional distribution and discards the background. Concurrently, it transfers a three-mass (target/confusion/background) that capture sample difficulty and inter-class confusion. Finally, TRKD focuses learning via a curriculum on $τ$: training begins with a larger $τ$ to convey broad non-target context, then $τ$ is progressively decreased to shrink the confusion-set, concentrating supervision on the most confusable classes. In extensive experiments on VoxCeleb1 with both homogeneous and heterogeneous teacher-student pairs, TRKD was consistently superior to recent KD variants and attained the lowest EER across all protocols.


【8】Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
标题:缩放模糊性:用音频语言模型增强语音情感识别中的人类注释
链接:https://arxiv.org/abs/2601.14620

作者:Wenda Zhang,Hongyu Jin,Siyi Wang,Zhiqiang Wei,Ting Dang
备注:Accepted by ICASSP 2026
摘要:语音情感识别模型通常使用单一的分类标签,忽略了人类情感的固有模糊性。模糊情绪识别通过将情绪表示为概率分布来解决这一问题,但进展受到从稀疏的人类注释推断出的不可靠的地面真实分布的限制。本文探讨了大型音频语言模型(ALMs)是否可以通过生成高质量的合成注释来缓解注释瓶颈。我们引入了一个框架,利用ALM创建合成感知代理,增强人类注释,以提高地面实况分布的可靠性。我们验证这些代理通过统计分析他们的对齐与人类的分布和评估其影响,通过微调ALM与增强的情感分布。此外,为了解决类不平衡和实现无偏评估,我们提出了DiME-Aug,一种分布感知的多模态情感增强策略。IEMOCAP和MSP-Podcast上的实验表明,合成注释增强了情感分布,特别是在注释一致性高的低歧义区域。然而,对于高度模糊的情感,随着人类分歧的增加,好处会减少。这项工作提供了第一个证据,ALMs可以解决模糊的情感识别中的注释不足,但强调需要更先进的提示或生成策略来处理高度模糊的情况。
摘要:Speech Emotion Recognition models typically use single categorical labels, overlooking the inherent ambiguity of human emotions. Ambiguous Emotion Recognition addresses this by representing emotions as probability distributions, but progress is limited by unreliable ground-truth distributions inferred from sparse human annotations. This paper explores whether Large Audio-Language Models (ALMs) can mitigate the annotation bottleneck by generating high-quality synthetic annotations. We introduce a framework leveraging ALMs to create Synthetic Perceptual Proxies, augmenting human annotations to improve ground-truth distribution reliability. We validate these proxies through statistical analysis of their alignment with human distributions and evaluate their impact by fine-tuning ALMs with the augmented emotion distributions. Furthermore, to address class imbalance and enable unbiased evaluation, we propose DiME-Aug, a Distribution-aware Multimodal Emotion Augmentation strategy. Experiments on IEMOCAP and MSP-Podcast show that synthetic annotations enhance emotion distribution, especially in low-ambiguity regions where annotation agreement is high. However, benefits diminish for highly ambiguous emotions with greater human disagreement. This work provides the first evidence that ALMs could address annotation scarcity in ambiguous emotion recognition, but highlights the need for more advanced prompting or generation strategies to handle highly ambiguous cases.


【9】Towards noise-robust speech inversion through multi-task learning with speech enhancement
标题:通过具有语音增强的多任务学习实现噪音稳健的语音倒置
链接:https://arxiv.org/abs/2601.14516

作者:Saba Tabatabaee,Carol Espy-Wilson
备注:Accepted for presentation at ICASSP 2026
摘要:最近的研究表明,自监督学习(SSL)的语音表示语音反转(SI)的有效性。然而,由于背景噪声的普遍存在,在现实世界的场景中应用SI仍然具有挑战性。我们提出了一个统一的框架,集成语音增强(SE)和SI模型,通过共享的基于SSL的语音表示。在这个框架中,SSL模型不仅被训练来支持SE模块抑制噪声,而且还产生对SI任务更具信息性的表示,从而使两个模块都能从联合训练中受益。在-5分贝的信噪比下,我们的SI任务方法在多路重合噪声下实现了80.95%的基线相对改善,在非多路重合噪声下实现了38.98%的基线相对改善,如通过所有估计参数的平均Pearson积矩相关性所测量的。
摘要:Recent studies demonstrate the effectiveness of Self Supervised Learning (SSL) speech representations for Speech Inversion (SI). However, applying SI in real-world scenarios remains challenging due to the pervasive presence of background noise. We propose a unified framework that integrates Speech Enhancement (SE) and SI models through shared SSL-based speech representations. In this framework, the SSL model is trained not only to support the SE module in suppressing noise but also to produce representations that are more informative for the SI task, allowing both modules to benefit from joint training. At a Signal-to-Noise Ratio of -5 db, our method for the SI task achieves relative improvements over the baseline of 80.95% under babble noise and 38.98% under non-babble noise, as measured by the average Pearson product-moment correlation across all estimated parameters.


【10】WeDefense: A Toolkit to Defend Against Fake Audio
标题:WeDefense:防御虚假音频的工具包
链接:https://arxiv.org/abs/2601.15240

作者:Lin Zhang,Johan Rohdin,Xin Wang,Junyi Peng,Tianchi Liu,You Zhang,Hieu-Thi Luong,Shuai Wang,Chengdong Liang,Anna Silnova,Nicholas Evans
备注:This is an ongoing work. v1 corresponds to the version completed by June 4, 2025 and previously submitted to ASRU 2025
摘要:生成AI的进步使得合成音频的创建成为可能,这种合成音频在感知上与真实的真实音频无法区分。虽然这一重大进展使许多积极的应用成为可能,但它也增加了滥用的风险,例如冒充,虚假信息和欺诈。尽管通过众多挑战和倡议发布了越来越多的开源虚假音频检测代码,但大多数都是针对特定的比赛,数据集或模型而量身定制的。一个标准化和统一的工具包,支持公平的基准测试和竞争解决方案的比较,不仅有共同的数据库,协议,指标,而且还有共享的代码库,是缺失的。为了解决这个问题,我们提出了WeDefense,这是第一个支持虚假音频检测和本地化的开源工具包。除了模型训练之外,WeDefense还强调关键但经常被忽视的组件:灵活的输入和增强,校准,分数融合,标准化的评估指标以及用于更深入理解和解释的分析工具。该工具包可在https://github.com/zlin0/wedefense上公开获取,其中包含用于虚假音频检测和定位的交互式演示。
摘要:The advances in generative AI have enabled the creation of synthetic audio which is perceptually indistinguishable from real, genuine audio. Although this stellar progress enables many positive applications, it also raises risks of misuse, such as for impersonation, disinformation and fraud. Despite a growing number of open-source fake audio detection codes released through numerous challenges and initiatives, most are tailored to specific competitions, datasets or models. A standardized and unified toolkit that supports the fair benchmarking and comparison of competing solutions with not just common databases, protocols, metrics, but also a shared codebase, is missing. To address this, we propose WeDefense, the first open-source toolkit to support both fake audio detection and localization. Beyond model training, WeDefense emphasizes critical yet often overlooked components: flexible input and augmentation, calibration, score fusion, standardized evaluation metrics, and analysis tools for deeper understanding and interpretation. The toolkit is publicly available at https://github.com/zlin0/wedefense with interactive demos for fake audio detection and localization.


【11】VCNAC: A Variable-Channel Neural Audio Codec for Mono, Stereo, and Surround Sound
标题:VNAC:一款适用于单声、立体声和环绕声的可变通道神经音频编解码器
链接:https://arxiv.org/abs/2601.14960

作者:Florian Grötschla,Arunasish Sen,Alessandro Lombardi,Guillermo Cámbara,Andreas Schwarz
备注:Submitted to EUSIPCO 2026
摘要:我们提出了VCNAC,可变通道神经音频编解码器。我们的方法具有单一的编码器和解码器参数化,使不同的通道设置,从单声道语音到电影5.1声道环绕音频的本地推理。通道兼容性目标确保多通道内容在解码到较少通道时保持感知质量。共享表示使生成语言模型能够在单组码本上进行训练,同时支持跨模态和通道配置的推理时间可扩展性。使用客观空间音频指标和主观听力测试的评估表明,我们的统一方法在单声道,立体声和环绕声音频配置中保持高重建质量。
摘要:We present VCNAC, a variable channel neural audio codec. Our approach features a single encoder and decoder parametrization that enables native inference for different channel setups, from mono speech to cinematic 5.1 channel surround audio. Channel compatibility objectives ensure that multi-channel content maintains perceptual quality when decoded to fewer channels. The shared representation enables training of generative language models on a single set of codebooks while supporting inference-time scalability across modalities and channel configurations. Evaluation using objective spatial audio metrics and subjective listening tests demonstrates that our unified approach maintains high reconstruction quality across mono, stereo, and surround audio configurations.


【12】Unlocking Large Audio-Language Models for Interactive Language Learning
标题:解锁交互式语言学习的大型音频语言模型
链接:https://arxiv.org/abs/2601.14744

作者:Hongfu Liu,Zhouying Cui,Xiangming Gu,Ye Wang
备注:Accepted to the Findings of EACL 2026
摘要:尽管计算机辅助发音训练(CAPT)系统不断发展,但在第二语言(L2)中实现发音熟练仍然是一个挑战。传统的CAPT系统通常提供不直观的反馈,缺乏可操作的指导,限制了其有效性。音频语言模型(ALMs)的最新进展提供了通过提供更用户友好的反馈来增强这些系统的潜力。在这项工作中,我们通过引入L2-Arctic-plus来研究基于聊天的发音训练的ALM,L2-Arctic-plus是一个英语数据集,具有详细的错误解释和可操作的改进建议。我们在此数据集上对级联的ASR+ LLM和现有的ALM进行基准测试,特别是在检测发音错误和生成可操作的反馈方面。为了提高性能,我们进一步提出在L2-Arctic-plus上对ALM进行预调优。实验结果表明,我们的发音调整模型显着优于现有的基准在错误发音检测和建议生成方面的客观和人类的评价,突出了所提出的数据集的价值。
摘要:Achieving pronunciation proficiency in a second language (L2) remains a challenge, despite the development of Computer-Assisted Pronunciation Training (CAPT) systems. Traditional CAPT systems often provide unintuitive feedback that lacks actionable guidance, limiting its effectiveness. Recent advancements in audio-language models (ALMs) offer the potential to enhance these systems by providing more user-friendly feedback. In this work, we investigate ALMs for chat-based pronunciation training by introducing L2-Arctic-plus, an English dataset with detailed error explanations and actionable suggestions for improvement. We benchmark cascaded ASR+LLMs and existing ALMs on this dataset, specifically in detecting mispronunciation and generating actionable feedback. To improve the performance, we further propose to instruction-tune ALMs on L2-Arctic-plus. Experimental results demonstrate that our instruction-tuned models significantly outperform existing baselines on mispronunciation detection and suggestion generation in terms of both objective and human evaluation, highlighting the value of the proposed dataset.


【13】Guided by the Plan: Enhancing Faithful Autoregressive Text-to-Audio Generation with Guided Decoding
标题:以计划为指导:通过引导解码增强忠实的自回归文本到音频生成
链接:https://arxiv.org/abs/2601.14304

作者:Juncheng Wang,Zhe Hu,Chao Xu,Siyue Ren,Yuxiang Feng,Yang Liu,Baigui Sun,Shujun Wang
备注:Accepted at EACL 2026
摘要:自回归(AR)模型擅长通过顺序地产生令牌来生成时间相干的音频,但它们经常在忠实地遵循复杂的文本提示,特别是那些描述复杂声音事件的文本提示时犹豫不决。我们在AR音频生成器中发现了一个令人惊讶的能力:它们的早期前缀令牌隐式地编码最终输出的全局语义属性,例如事件计数和声音对象类别,揭示了一种隐式规划的形式。基于这一见解,我们提出了Plan-Critic,这是一种轻量级的辅助模型,采用广义优势估计(GAE)目标进行训练,以预测部分世代的最终预防跟踪质量。在推理时,Plan-Critic支持引导探索:它早期评估候选前缀,修剪低保真度轨迹,并将计算重新分配给高潜力的规划种子。我们的Plan-Critic引导采样在CLAP评分上比AR基线提高了10个点-建立了AR文本到音频生成的新技术水平-同时保持了与标准最佳N解码的计算奇偶性。这项工作弥合了因果生成和全局语义对齐之间的差距,表明即使是严格的自回归模型也可以提前计划。
摘要:Autoregressive (AR) models excel at generating temporally coherent audio by producing tokens sequentially, yet they often falter in faithfully following complex textual prompts, especially those describing complex sound events. We uncover a surprising capability in AR audio generators: their early prefix tokens implicitly encode global semantic attributes of the final output, such as event count and sound-object category, revealing a form of implicit planning. Building on this insight, we propose Plan-Critic, a lightweight auxiliary model trained with a Generalized Advantage Estimation (GAE)-inspired objective to predict final instruction-following quality from partial generations. At inference time, Plan-Critic enables guided exploration: it evaluates candidate prefixes early, prunes low-fidelity trajectories, and reallocates computation to high-potential planning seeds. Our Plan-Critic-guided sampling achieves up to a 10-point improvement in CLAP score over the AR baseline-establishing a new state of the art in AR text-to-audio generation-while maintaining computational parity with standard best-of-N decoding. This work bridges the gap between causal generation and global semantic alignment, demonstrating that even strictly autoregressive models can plan ahead.


【14】Call2Instruct: Automated Pipeline for Generating Q&A Datasets from Call Center Recordings for LLM Fine-Tuning
标题:Call 2 Direcct:从呼叫中心录音生成问答数据集以进行LLM微调的自动化管道
链接:https://arxiv.org/abs/2601.14263

作者:Alex Echeverria,Sávio Salvarino Teles de Oliveira,Fernando Marques Federson
备注:15 pages, 1 figures, conference
摘要:大规模语言模型(LLM)对特定领域的适应取决于高质量的微调数据集,特别是在教学格式(例如,回答- Q&A)。然而,生成这些数据集,特别是从非结构化来源(如呼叫中心音频记录)生成这些数据集,由于数据的噪声和无序性而带来了重大挑战。本文提出了一种解决方案,通过提供一个端到端的自动化管道,从这样的录音生成问答教学数据集。开发的方法包括顺序步骤的音频处理(包括日记化,噪声去除和自动转录),文本处理(清洗,规范化和匿名化),语义提取的客户需求和伴随的响应使用向量嵌入,并通过语义搜索匹配,以形成最终的Q&A对。因此,完整的管道成功实现,生成了专门为指令微调格式化的数据集。通过LLM模型(基于Llama 2 7 B)的成功微调,证实了所生成数据集的实用价值和可行性,并在功能上得到了证明。本文的结论指出,所提出的方法是可行的,从呼叫中心的非结构化会话数据转换为宝贵的资源,培训LLM。这一发展有可能为客户服务领域的问答任务创建更有效的人工智能系统开辟道路。开发的代码已公开提供,以促进可重复性和未来的研究。
摘要:The adaptation of Large-Scale Language Models (LLMs) to specific domains depends on high-quality fine-tuning datasets, particularly in instructional format (e.g., Question-Answer - Q&A). However, generating these datasets, particularly from unstructured sources such as call center audio recordings, poses a significant challenge due to the noisy and disorganized nature of the data. This paper presents a solution to this challenge by offering an end-to-end automated pipeline for generating Q&A instructional datasets from such recordings. The methodology developed comprises sequential steps of audio processing (including diarization, noise removal and automatic transcription), textual processing (cleaning, normalization, and anonymization), semantic extraction of customer demands and attendant responses using vector embeddings, and matching via semantic search to form the final Q&A pairs. As a result, the complete pipeline was successfully implemented, generating a dataset specifically formatted for Instruct Fine Tuning. The practical value and feasibility of the generated dataset were substantiated and functionally demonstrated through the successful fine-tuning of an LLM model (based on Llama 2 7B). The conclusion of the paper states that the proposed approach is viable for converting unstructured conversational data from call centers into valuable resources for training LLMs. This development has the potential to open up avenues for creating more effective AI systems for Q&A tasks in the customer service domain. The developed codes have been made publicly available to promote reproducibility and future research.


【15】A Cloud-Based Cross-Modal Transformer for Emotion Recognition and Adaptive Human-Computer Interaction
标题:基于云的跨模态情感识别和自适应人机交互Transformer
链接:https://arxiv.org/abs/2601.14259

作者:Ziwen Zhong,Zhitao Shu,Yue Zhao
摘要:情感识别是下一代人机交互(HCI)的基本组成部分,使机器能够感知,理解和响应用户的情感状态。然而,现有的系统往往依赖于单模态分析,如面部表情,语音语调,或文本情感,导致有限的鲁棒性和在现实世界环境中的泛化能力差。为了应对这些挑战,本研究提出了一个基于云的跨模态Transformer(CMT)框架,用于多模态情感识别和自适应人机交互。该模型使用预训练编码器(Vision Transformer,Wav2Vec2和BERT)集成视觉,听觉和文本信号,并采用跨模态注意力机制来捕获异构特征之间的复杂相互依赖关系。通过利用云计算基础设施,在Kubernetes和TensorFlow Serving上进行分布式训练,该系统可以为大规模用户交互提供可扩展的低延迟情感识别。在包括IEMOCAP、MELD和AffectNet在内的基准数据集上进行的实验表明,CMT实现了最先进的性能,与强大的多模态基线相比,F1分数提高了3.0%,交叉熵损失降低了12.9%。此外,云部署评估显示平均响应延迟为128 ms,与传统的基于transformer的融合系统相比减少了35%。这些结果证实,所提出的框架,使有效的,实时的情感识别和自适应反馈的应用程序,如智能客户服务,虚拟辅导系统,情感计算接口,标志着一个重要的一步,云原生情感计算和情感智能交互系统。
摘要:Emotion recognition is a fundamental component of next-generation human-computer interaction (HCI), enabling machines to perceive, understand, and respond to users' affective states. However, existing systems often rely on single-modality analysis such as facial expressions, speech tone, or textual sentiment, resulting in limited robustness and poor generalization in real-world environments. To address these challenges, this study proposes a Cloud-Based Cross-Modal Transformer (CMT) framework for multimodal emotion recognition and adaptive human-computer interaction. The proposed model integrates visual, auditory, and textual signals using pretrained encoders (Vision Transformer, Wav2Vec2, and BERT) and employs a cross-modal attention mechanism to capture complex interdependencies among heterogeneous features. By leveraging cloud computing infrastructure with distributed training on Kubernetes and TensorFlow Serving, the system enables scalable, low-latency emotion recognition for large-scale user interactions. Experiments conducted on benchmark datasets including IEMOCAP, MELD, and AffectNet demonstrate that the CMT achieves state-of-the-art performance, improving the F1-score by 3.0 percent and reducing cross-entropy loss by 12.9 percent compared to strong multimodal baselines. Additionally, cloud deployment evaluations show an average response latency of 128 ms, representing a 35 percent reduction compared with conventional transformer-based fusion systems. These results confirm that the proposed framework enables efficient, real-time emotion recognition and adaptive feedback in applications such as intelligent customer service, virtual tutoring systems, and affective computing interfaces, marking an important step toward cloud-native affective computing and emotionally intelligent interactive systems.


机器翻译由腾讯交互翻译提供,仅供参考