今日论文合集:cs.SD语音6篇,eess.AS音频处理8篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Sound Clouds: Exploring ambient intelligence in public spaces to elicit deep human experience of awe, wonder, and beauty
标题:声云:探索公共空间的环境智能,引发人类对敬畏、奇迹和美丽的深刻体验
链接:https://arxiv.org/abs/2510.15865

作者:Chengzhi Zhang, Dashiel Carrera, Daksh Kapoor, Jasmine Kaur, Jisu Kim, Brian Magerko
备注:4 pages, Artwork accepted by NeurIPS Creative AI Track 2025
摘要:虽然我们在日常生活中遇到的环境智能(AmI)系统,包括安全监控和节能系统,通常用于实用目的,但我们想知道如何在公共空间设计和实现环境人工智能体验,引发人类对敬畏,惊奇和美丽的深刻感受。作为一种表现形式,我们介绍了声音云,一个沉浸式的艺术装置,根据参与者与几个人类高度的球体的互动产生现场音乐。我们的安装是对未来环境智能的挑衅,它激发而不是限制了AmI未来的可能性。
摘要:While the ambient intelligence (AmI) systems we encounter in our daily lives, including security monitoring and energy-saving systems, typically serve pragmatic purposes, we wonder how we can design and implement ambient artificial intelligence experiences in public spaces that elicit deep human feelings of awe, wonder, and beauty. As a manifestation, we introduce Sound Clouds, an immersive art installation that generates live music based on participants' interaction with several human-height spheres. Our installation serves as a provocation into future ambient intelligence that provokes, not limits, the future possibilities of AmI.


【2】SpikeVox: Towards Energy-Efficient Speech Therapy Framework with Spike-driven Generative Language Models
标题:SpikeVox:采用Spike驱动的生成语言模型迈向节能语音治疗框架
链接:https://arxiv.org/abs/2510.15566

作者:Rachmad Vidya Wicaksana Putra, Aadithyan Rajesh Nair, Muhammad Shafique
备注:Accepted at the IEEE Biomedical Circuits and Systems Conference (BioCAS) 2025, Abu Dhabi, UAE
摘要:言语障碍会显著影响患者的沟通、学习和社交能力。然而,现有的言语治疗解决方案(例如,治疗师或工具)仍然有限且昂贵,因此这样的解决方案仍然不足以为全世界数百万患者提供服务。为了解决这个问题,最先进的方法采用神经网络(NN)算法来帮助准确检测语音障碍。然而,这些方法不提供治疗建议作为反馈,因此为患者提供部分解决方案。此外,这些方法由于其复杂且资源密集型的NN处理而导致高能耗,因此阻碍了它们在低功率/能量平台(例如,智能手机)。为此,我们提出了SpikeVox,这是一种通过尖峰驱动的生成语言模型实现节能语音治疗解决方案的新框架。具体而言,SpikeVox采用语音识别模块来执行高度准确的语音到文本转换;利用尖峰驱动的生成语言模型来有效地执行语音障碍检测的模式分析,并生成合适的治疗练习;提供正确发音的指导作为反馈;以及利用REST API为用户实现无缝交互。实验结果表明,SpikeVox在语音障碍识别中平均达到88%的置信水平,同时为治疗练习提供完整的反馈。因此,SpikeVox为节能的语言治疗解决方案提供了一个全面的框架,并有可能解决全球语言治疗的巨大差距。
摘要:Speech disorders can significantly affect the patients capability to communicate, learn, and socialize. However, existing speech therapy solutions (e.g., therapist or tools) are still limited and costly, hence such solutions remain inadequate for serving millions of patients worldwide. To address this, state-of-the-art methods employ neural network (NN) algorithms to help accurately detecting speech disorders. However, these methods do not provide therapy recommendation as feedback, hence providing partial solution for patients. Moreover, these methods incur high energy consumption due to their complex and resource-intensive NN processing, hence hindering their deployments on low-power/energy platforms (e.g., smartphones). Toward this, we propose SpikeVox, a novel framework for enabling energy-efficient speech therapy solutions through spike-driven generative language model. Specifically, SpikeVox employs a speech recognition module to perform highly accurate speech-to-text conversion; leverages a spike-driven generative language model to efficiently perform pattern analysis for speech disorder detection and generates suitable exercises for therapy; provides guidance on correct pronunciation as feedback; as well as utilizes the REST API to enable seamless interaction for users. Experimental results demonstrate that SpikeVox achieves 88% confidence level on average in speech disorder recognition, while providing a complete feedback for therapy exercises. Therefore, SpikeVox provides a comprehensive framework for energy-efficient speech therapy solutions, and potentially addresses the significant global speech therapy access gap.


【3】Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
标题:扩展音频上下文以实现大型音频语言模型中的长篇理解
链接:https://arxiv.org/abs/2510.15231

作者:Yuatyong Chaichana, Pittawat Taveekitworachai, Warit Sirichotedumrong, Potsawee Manakul, Kunat Pipatanakul
摘要:大型音频语言模型(LALM)通常受到短音频上下文窗口的限制,即使它们的文本主干支持长上下文,也会限制长形式的音频理解。先前的工作已经在单峰LLM上引入了上下文扩展方法(例如YaRN),但它们在LALM上的应用仍然未被探索。首先,在基于Rope的上下文扩展的基础上,我们引入了Partial YaRN,这是一种无需训练的仅音频扩展方法,仅修改音频标记位置,保留文本位置不变,以保留基本LLM的文本功能。其次,我们提出了虚拟长格式音频训练(VLAT),一种将部分YaRN扩展为训练时间位置增强的训练策略。VLAT在训练过程中模拟不同的音频长度,从而能够泛化到远长于训练中看到的输入,并提高长上下文音频理解的鲁棒性。我们在SALMONN和Qwen 2-Audio上的实验表明,Partial YaRN在广泛的设置中优于原始模型,VLAT训练策略提供了实质性的改进,在看不见的长音频上实现了强大的性能。
摘要:Large Audio-Language Models (LALMs) are often constrained by short audio context windows, even when their text backbones support long contexts, limiting long-form audio understanding. Prior work has introduced context-extension methods (e.g. YaRN) on unimodal LLMs, yet their application to LALMs remains unexplored. First, building on RoPE-based context extension, we introduce Partial YaRN, a training-free, audio-only extension method that modifies only audio token positions, leaving text positions intact to preserve the base LLM's text capabilities. Second, we propose Virtual Longform Audio Training (VLAT), a training strategy that extends Partial YaRN into a training-time positional augmentation. VLAT simulates diverse audio lengths during training, enabling generalization to inputs far longer than those seen in training and improving robustness for long-context audio understanding. Our experiments on SALMONN and Qwen2-Audio show that Partial YaRN outperforms the original models across wide range of settings, and VLAT training strategy provides substantial improvement, achieving strong performance on long audio of unseen lengths.


【4】Quantization-Based Score Calibration for Few-Shot Keyword Spotting with Dynamic Time Warping in Noisy Environments
标题:基于量化的分数校准,用于在噪音环境中动态时间扭曲的Few-Shot关键词发现
链接:https://arxiv.org/abs/2510.15432

作者:Kevin Wilkinghoff, Alessia Cornaggia-Urrigshardt, Zheng-Hua Tan
摘要:使用关键字定位(KWS)系统检测关键字的出现需要阈值化连续检测分数。选择适当的阈值是一项重要的任务,通常依赖于优化验证数据集的性能。然而,这种贪婪的阈值选择通常导致对看不见的数据的次优性能,特别是在变化的或有噪声的声学环境或Few-Shot设置中。在这项工作中,我们研究了基于模板的开集Few-Shot KWS的检测阈值估计,使用动态时间规整的噪声语音数据。为了减轻次优阈值所造成的性能下降,我们提出了一种分数校准方法,包括两个不同的步骤:量化嵌入和归一化检测分数使用阈值之前的量化误差。在KWS-DailyTalk仿真高频无线信道上的实验表明,该方法简化了检测门限的选择,显著提高了检测性能。
摘要:Detecting occurrences of keywords with keyword spotting (KWS) systems requires thresholding continuous detection scores. Selecting appropriate thresholds is a non-trivial task, typically relying on optimizing the performance on a validation dataset. However, such greedy threshold selection often leads to suboptimal performance on unseen data, particularly in varying or noisy acoustic environments or few-shot settings. In this work, we investigate detection threshold estimation for template-based open-set few-shot KWS using dynamic time warping on noisy speech data. To mitigate the performance degradation caused by suboptimal thresholds, we propose a score calibration approach consisting of two different steps: quantizing embeddings and normalizing detection scores using the quantization error prior to thresholding. Experiments on KWS-DailyTalk with simulated high frequency radio channels show that the proposed calibration approach simplifies the choice of detection thresholds and significantly improves the resulting performance.


【5】DroneAudioset: An Audio Dataset for Drone-based Search and Rescue
标题:无人机Audioset:用于无人机搜索和救援的音频数据集
链接:https://arxiv.org/abs/2510.15383

作者:Chitralekha Gupta, Soundarya Ramesh, Praveen Sasikumar, Kian Peen Yeo, Suranga Nanayakkara
备注:Accepted in Neurips (Datasets and Benchmarks Track) 2025. The first two authors are equal contributors
摘要:无人驾驶飞行器(UAV)或无人机越来越多地用于搜索和救援任务,以检测人类的存在。现有的系统主要利用基于视觉的方法,这些方法在低可见性或遮挡下容易失败。基于无人机的音频感知提供了希望,但遭受极端的自我噪声,掩盖了指示人类存在的声音。现有的数据集要么是多样性有限,要么是合成的,缺乏真正的声学交互,也没有标准化的无人机试听设置。为此,我们介绍了DroneAudioset(该数据集在麻省理工学院许可下可在https://huggingface.co/datasets/ahlab-drone-project/DroneAudioSet/上公开获得),这是一个全面的无人机试听数据集,具有23.5小时的注释录音,涵盖了各种无人机类型,油门,麦克风配置以及环境中从-57.2 dB到-2.5 dB的各种信噪比(SNR)。该数据集能够在具有挑战性的条件下开发和系统评估用于人类存在检测的噪声抑制和分类方法,同时还为无人机试听系统的实际设计考虑提供信息,例如麦克风放置权衡以及无人机噪声感知音频处理的开发。该数据集是实现无人机试听系统设计和部署的重要一步。
摘要:Unmanned Aerial Vehicles (UAVs) or drones, are increasingly used in search and rescue missions to detect human presence. Existing systems primarily leverage vision-based methods which are prone to fail under low-visibility or occlusion. Drone-based audio perception offers promise but suffers from extreme ego-noise that masks sounds indicating human presence. Existing datasets are either limited in diversity or synthetic, lacking real acoustic interactions, and there are no standardized setups for drone audition. To this end, we present DroneAudioset (The dataset is publicly available at https://huggingface.co/datasets/ahlab-drone-project/DroneAudioSet/ under the MIT license), a comprehensive drone audition dataset featuring 23.5 hours of annotated recordings, covering a wide range of signal-to-noise ratios (SNRs) from -57.2 dB to -2.5 dB, across various drone types, throttles, microphone configurations as well as environments. The dataset enables development and systematic evaluation of noise suppression and classification methods for human-presence detection under challenging conditions, while also informing practical design considerations for drone audition systems, such as microphone placement trade-offs, and development of drone noise-aware audio processing. This dataset is an important step towards enabling design and deployment of drone-audition systems.


【6】LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
标题:LongCat-audio-Codec:一种专为语音大型语言模型设计的音频令牌化器和去令牌化器解决方案
链接:https://arxiv.org/abs/2510.15227

作者:Xiaohan Zhao, Hongyu Xiang, Shengze Ye, Song Li, Zhengkun Tian, Guanyu Chen, Ke Ding, Guanglu Wan
摘要:本文介绍了LongCat-Audio-Codec,这是一种音频标记器和去标记器解决方案,专为工业级端到端语音大型语言模型设计。通过利用解耦模型架构和多级训练策略,LongCat-Audio-Codec具有强大的语义建模能力,灵活的声学特征提取能力和低延迟流合成能力。它以16.67 Hz的超低帧速率对语音进行编码,最小比特率为0.43 kbps,最大比特率为0.87 kbps。评测结果表明,LongCat-Audio-Codec实现了较强的语音可懂度,能够在较低码率下合成高质量的语音,从而有效地平衡了编码效率和解码质量。LongCat-Audio-Codec的推理代码和模型检查点可在https://github.com/meituan-longcat/LongCat-Audio-Codec上获得。
摘要:This paper presents LongCat-Audio-Codec, an audio tokenizer and detokenizer solution designed for industrial grade end-to-end speech large language models. By leveraging a decoupled model architecture and a multistage training strategy, LongCat-Audio-Codec exhibits robust semantic modeling capabilities, flexible acoustic feature extraction capabilities, and low-latency streaming synthesis capabilities. It encodes speech at an ultra-low frame rate of 16.67 Hz, with a minimum bitrate of 0.43 kbps and a maximum bitrate of 0.87 kbps. Evaluation results demonstrate that LongCat-Audio-Codec achieves strong speech intelligibility and is capable of synthesizing highquality speech at low bitrate, thus effectively balancing coding efficiency and decoding quality. The inference code and model checkpoints of LongCat-Audio-Codec are available at: https://github.com/meituan-longcat/LongCat-Audio-Codec.


eess.AS音频处理


【1】Magnitude and Phase-based Feature Fusion Using Co-attention Mechanism for Speaker recognition
标题:利用共同注意机制的基于幅度和阶段的特征融合用于说话人识别
链接:https://arxiv.org/abs/2510.15659

作者:Rongfeng Su, Mengjie Du, Xiaokang Liu, Lan Wang, Nan Yan
摘要:与声源特征相关的基于相位的特征可以被引入到基于幅度的说话人识别系统中,以提高系统的性能。然而,传统的特征级融合方法通常忽略了说话人语义在幅度域和相位域的独特贡献。针对这一问题,提出了一种基于共注意机制的特征级融合框架用于说话人识别。该框架由两个独立的子网络的幅度和相位域分别。然后,在池化层之前,通过共同注意机制融合两个域的中间高级别输出。来自共同注意模块的相关矩阵被认为根据不同的发音重新分配用于动态缩放幅度域和相位域中的贡献的权重。在VoxCeleb上的实验结果表明,基于共注意机制的特征级融合策略的Top-1准确率达到97.20%,绝对性能优于现有系统的0.82%,与基于FBank的单一特征系统相比,EER降低了0.45%.
摘要:Phase-based features related to vocal source characteristics can be incorporated into magnitude-based speaker recognition systems to improve the system performance. However, traditional feature-level fusion methods typically ignore the unique contributions of speaker semantics in the magnitude and phase domains. To address this issue, this paper proposed a feature-level fusion framework using the co-attention mechanism for speaker recognition. The framework consists of two separate sub-networks for the magnitude and phase domains respectively. Then, the intermediate high-level outputs of both domains are fused by the co-attention mechanism before a pooling layer. A correlation matrix from the co-attention module is supposed to re-assign the weights for dynamically scaling contributions in the magnitude and phase domains according to different pronunciations. Experiments on VoxCeleb showed that the proposed feature-level fusion strategy using the co-attention mechanism gave the Top-1 accuracy of 97.20%, outperforming the state-of-the-art system with 0.82% absolutely, and obtained EER reduction of 0.45% compared to single feature system using FBank.


【2】MC-LExt: Multi-Channel Target Speaker Extraction with Onset-Prompted Speaker Conditioning Mechanism
标题:MC-LExt:多通道目标说话人提取与起始-消除说话人条件反射机制
链接:https://arxiv.org/abs/2510.15437

作者:Tongtao Ling, Shulin He, Pengjie Shen, Zhong-Qiu Wang
备注:5 pages, 2 figures
摘要:多通道目标说话人提取(MC-TSE)的目的是从多个麦克风捕获的多说话人信号中提取目标说话人的语音。现有的方法通常依赖于辅助线索,如到达方向(DOA)或扬声器嵌入。然而,基于DOA的方法依赖于显式方向估计并且对麦克风阵列几何形状敏感,而基于扬声器嵌入的方法以隐式方式对扬声器身份进行建模并且可能在噪声混响条件下降级。为了解决这些限制,我们提出了多通道监听提取(MC-LExt),这是一个简单但高效的MC-TSE框架。我们的主要思想是在多通道混合的每个通道中预先加入目标说话人的一个简短的注册话语,提供一个可以引导TSE的启动提示条件信号。这种设计允许深度神经网络(DNN)以完全端到端的方式联合学习空间和说话者身份线索。噪声混响基准实验,包括WHAMR!和MC-Libri 2 Mix,证明了MC-TSE的有效性。
摘要:Multi-channel target speaker extraction (MC-TSE) aims to extract a target speaker's voice from multi-speaker signals captured by multiple microphones. Existing methods often rely on auxiliary clues such as direction-of-arrival (DOA) or speaker embeddings. However, DOA-based approaches depend on explicit direction estimation and are sensitive to microphone array geometry, while methods based on speaker embeddings model speaker identity in an implicit manner and may degrade in noisy-reverberant conditions. To address these limitations, we propose multi-channel listen to extract (MC-LExt), a simple but highly-effective framework for MC-TSE. Our key idea is to prepend a short enrollment utterance of the target speaker to each channel of the multi-channel mixture, providing an onset-prompted conditioning signal that can guide TSE. This design allows the deep neural network (DNN) to learn spatial and speaker identity cues jointly in a fully end-to-end manner. Experiments on noisy-reverberant benchmarks, including WHAMR! and MC-Libri2Mix, demonstrate the effectiveness of MC-TSE.


【3】Quantization-Based Score Calibration for Few-Shot Keyword Spotting with Dynamic Time Warping in Noisy Environments
标题:基于量化的分数校准,用于在噪音环境中动态时间扭曲的Few-Shot关键词发现
链接:https://arxiv.org/abs/2510.15432

作者:Kevin Wilkinghoff, Alessia Cornaggia-Urrigshardt, Zheng-Hua Tan
摘要:使用关键字定位(KWS)系统检测关键字的出现需要阈值化连续检测分数。选择适当的阈值是一项重要的任务,通常依赖于优化验证数据集的性能。然而,这种贪婪的阈值选择通常导致对看不见的数据的次优性能,特别是在变化的或有噪声的声学环境或Few-Shot设置中。在这项工作中,我们研究了基于模板的开集Few-Shot KWS的检测阈值估计,使用动态时间规整的噪声语音数据。为了减轻次优阈值所造成的性能下降,我们提出了一种分数校准方法,包括两个不同的步骤:量化嵌入和归一化检测分数使用阈值之前的量化误差。在KWS-DailyTalk仿真高频无线信道上的实验表明,该方法简化了检测门限的选择,显著提高了检测性能。
摘要:Detecting occurrences of keywords with keyword spotting (KWS) systems requires thresholding continuous detection scores. Selecting appropriate thresholds is a non-trivial task, typically relying on optimizing the performance on a validation dataset. However, such greedy threshold selection often leads to suboptimal performance on unseen data, particularly in varying or noisy acoustic environments or few-shot settings. In this work, we investigate detection threshold estimation for template-based open-set few-shot KWS using dynamic time warping on noisy speech data. To mitigate the performance degradation caused by suboptimal thresholds, we propose a score calibration approach consisting of two different steps: quantizing embeddings and normalizing detection scores using the quantization error prior to thresholding. Experiments on KWS-DailyTalk with simulated high frequency radio channels show that the proposed calibration approach simplifies the choice of detection thresholds and significantly improves the resulting performance.


【4】Towards Blind Data Cleaning: A Case Study in Music Source Separation
标题:面向盲数据清洗:音乐源分离的案例研究
链接:https://arxiv.org/abs/2510.15409

作者:Azalea Gui, Woosung Choi, Junghyun Koo, Kazuki Shimada, Takashi Shibuya, Joan Serrà, Wei-Hsiang Liao, Yuki Mitsufuji
备注:Submitted to IEEE ICASSP 2026
摘要:用于音乐源分离的深度学习模型的性能在很大程度上取决于训练数据的质量。然而,数据集通常会被难以检测的伪像(如音频溢出和标签噪声)破坏。由于污染的类型和程度通常是未知的,因此针对特定腐蚀的清洁方法通常是不切实际的。本文提出并评估了两种不同的、与噪声无关的数据清理方法来应对这一挑战。第一种方法使用数据归因,通过遗忘来识别和过滤出对产生干净输出贡献最小的训练样本。第二种方法利用Fr\'echet音频距离来测量和删除感知上与小而可信的干净参考集不相似的样本。在被模拟的真实世界噪声分布污染的数据集上,我们基于非学习的方法产生了一个清洁的数据集和一个相应的模型,该模型的性能优于原始污染数据和用于清洁的小的清洁参考集。这个结果关闭了受污染的基线和在没有任何污染的相同数据集上训练的模型之间的性能差距的大约66.7%。与针对特定工件量身定制的方法不同,我们的噪声不可知方法为管理高质量训练数据提供了更通用和更广泛适用的解决方案。
摘要:The performance of deep learning models for music source separation heavily depends on training data quality. However, datasets are often corrupted by difficult-to-detect artifacts such as audio bleeding and label noise. Since the type and extent of contamination are typically unknown, cleaning methods targeting specific corruptions are often impractical. This paper proposes and evaluates two distinct, noise-agnostic data cleaning methods to address this challenge. The first approach uses data attribution via unlearning to identify and filter out training samples that contribute the least to producing clean outputs. The second leverages the Fr\'echet Audio Distance to measure and remove samples that are perceptually dissimilar to a small and trusted clean reference set. On a dataset contaminated with a simulated distribution of real-world noise, our unlearning-based methods produced a cleaned dataset and a corresponding model that outperforms both the original contaminated data and the small clean reference set used for cleaning. This result closes approximately 66.7\% of the performance gap between the contaminated baseline and a model trained on the same dataset without any contamination. Unlike methods tailored for specific artifacts, our noise-agnostic approaches offer a more generic and broadly applicable solution for curating high-quality training data.


【5】DroneAudioset: An Audio Dataset for Drone-based Search and Rescue
标题:无人机Audioset:用于无人机搜索和救援的音频数据集
链接:https://arxiv.org/abs/2510.15383

作者:Chitralekha Gupta, Soundarya Ramesh, Praveen Sasikumar, Kian Peen Yeo, Suranga Nanayakkara
备注:Accepted in Neurips (Datasets and Benchmarks Track) 2025. The first two authors are equal contributors
摘要:无人驾驶飞行器(UAV)或无人机越来越多地用于搜索和救援任务,以检测人类的存在。现有的系统主要利用基于视觉的方法,这些方法在低可见性或遮挡下容易失败。基于无人机的音频感知提供了希望,但遭受极端的自我噪声,掩盖了指示人类存在的声音。现有的数据集要么是多样性有限,要么是合成的,缺乏真正的声学交互,也没有标准化的无人机试听设置。为此,我们介绍了DroneAudioset(该数据集在麻省理工学院许可下可在https://huggingface.co/datasets/ahlab-drone-project/DroneAudioSet/上公开获得),这是一个全面的无人机试听数据集,具有23.5小时的注释录音,涵盖了各种无人机类型,油门,麦克风配置以及环境中从-57.2 dB到-2.5 dB的各种信噪比(SNR)。该数据集能够在具有挑战性的条件下开发和系统评估用于人类存在检测的噪声抑制和分类方法,同时还为无人机试听系统的实际设计考虑提供信息,例如麦克风放置权衡以及无人机噪声感知音频处理的开发。该数据集是实现无人机试听系统设计和部署的重要一步。
摘要:Unmanned Aerial Vehicles (UAVs) or drones, are increasingly used in search and rescue missions to detect human presence. Existing systems primarily leverage vision-based methods which are prone to fail under low-visibility or occlusion. Drone-based audio perception offers promise but suffers from extreme ego-noise that masks sounds indicating human presence. Existing datasets are either limited in diversity or synthetic, lacking real acoustic interactions, and there are no standardized setups for drone audition. To this end, we present DroneAudioset (The dataset is publicly available at https://huggingface.co/datasets/ahlab-drone-project/DroneAudioSet/ under the MIT license), a comprehensive drone audition dataset featuring 23.5 hours of annotated recordings, covering a wide range of signal-to-noise ratios (SNRs) from -57.2 dB to -2.5 dB, across various drone types, throttles, microphone configurations as well as environments. The dataset enables development and systematic evaluation of noise suppression and classification methods for human-presence detection under challenging conditions, while also informing practical design considerations for drone audition systems, such as microphone placement trade-offs, and development of drone noise-aware audio processing. This dataset is an important step towards enabling design and deployment of drone-audition systems.


【6】LDCodec: A high quality neural audio codec with low-complexity decoder
标题:LDCodec:具有低复杂度解码器的高质量神经音频编解码器
链接:https://arxiv.org/abs/2510.15364

作者:Jiawei Jiang, Linping Xu, Dejun Zhang, Qingbo Huang, Xianjun Xia, Yijian Xiao
摘要:神经音频编码已被证明在极低的比特率下优于经典音频编码。然而,神经音频编解码器的实际应用仍然受到其高复杂性的限制。为了应对这一挑战,我们开发了一种具有低复杂度解码器的高质量神经音频编解码器,名为LDCodec(低复杂度解码器神经音频编解码器),专为按需流媒体客户端(如智能手机)设计。具体来说,我们引入了一种新的残差单元结合长期和短期残差矢量量化(LSRVQ),子带全带频率鉴别器,和感知损失函数。这种组合导致具有较低复杂度的高质量音频重构。我们的主观和客观测试都表明,我们提出的LDCodec在6kbps优于Opus在12kbps。
摘要:Neural audio coding has been shown to outperform classical audio coding at extremely low bitrates. However, the practical application of neural audio codecs is still limited by their elevated complexity. To address this challenge, we have developed a high-quality neural audio codec with a low-complexity decoder, named LDCodec (Low-complexity Decoder Neural Audio Codec), specifically designed for on-demand streaming media clients, such as smartphones. Specifically, we introduced a novel residual unit combined with Long-term and Short-term Residual Vector Quantization (LSRVQ), subband-fullband frequency discriminators, and perceptual loss functions. This combination results in high-quality audio reconstruction with lower complexity. Both our subjective and objective tests demonstrated that our proposed LDCodec at 6kbps outperforms Opus at 12kbps.


【7】LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
标题:LongCat-audio-Codec:一种专为语音大型语言模型设计的音频令牌化器和去令牌化器解决方案
链接:https://arxiv.org/abs/2510.15227

作者:Xiaohan Zhao, Hongyu Xiang, Shengze Ye, Song Li, Zhengkun Tian, Guanyu Chen, Ke Ding, Guanglu Wan
摘要:本文介绍了LongCat-Audio-Codec,这是一种音频标记器和去标记器解决方案,专为工业级端到端语音大型语言模型设计。通过利用解耦模型架构和多级训练策略,LongCat-Audio-Codec具有强大的语义建模能力,灵活的声学特征提取能力和低延迟流合成能力。它以16.67 Hz的超低帧速率对语音进行编码,最小比特率为0.43 kbps,最大比特率为0.87 kbps。评测结果表明,LongCat-Audio-Codec实现了较强的语音可懂度,能够在较低码率下合成高质量的语音,从而有效地平衡了编码效率和解码质量。LongCat-Audio-Codec的推理代码和模型检查点可在https://github.com/meituan-longcat/LongCat-Audio-Codec上获得。
摘要:This paper presents LongCat-Audio-Codec, an audio tokenizer and detokenizer solution designed for industrial grade end-to-end speech large language models. By leveraging a decoupled model architecture and a multistage training strategy, LongCat-Audio-Codec exhibits robust semantic modeling capabilities, flexible acoustic feature extraction capabilities, and low-latency streaming synthesis capabilities. It encodes speech at an ultra-low frame rate of 16.67 Hz, with a minimum bitrate of 0.43 kbps and a maximum bitrate of 0.87 kbps. Evaluation results demonstrate that LongCat-Audio-Codec achieves strong speech intelligibility and is capable of synthesizing highquality speech at low bitrate, thus effectively balancing coding efficiency and decoding quality. The inference code and model checkpoints of LongCat-Audio-Codec are available at: https://github.com/meituan-longcat/LongCat-Audio-Codec.


【8】Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
标题:扩展音频上下文以实现大型音频语言模型中的长篇理解
链接:https://arxiv.org/abs/2510.15231

作者:Yuatyong Chaichana, Pittawat Taveekitworachai, Warit Sirichotedumrong, Potsawee Manakul, Kunat Pipatanakul
摘要:大型音频语言模型(LALM)通常受到短音频上下文窗口的限制,即使它们的文本主干支持长上下文,也会限制长形式的音频理解。先前的工作已经在单峰LLM上引入了上下文扩展方法(例如YaRN),但它们在LALM上的应用仍然未被探索。首先,在基于Rope的上下文扩展的基础上,我们引入了Partial YaRN,这是一种无需训练的仅音频扩展方法,仅修改音频标记位置,保留文本位置不变,以保留基本LLM的文本功能。其次,我们提出了虚拟长格式音频训练(VLAT),一种将部分YaRN扩展为训练时间位置增强的训练策略。VLAT在训练过程中模拟不同的音频长度,从而能够泛化到远长于训练中看到的输入,并提高长上下文音频理解的鲁棒性。我们在SALMONN和Qwen 2-Audio上的实验表明,Partial YaRN在广泛的设置中优于原始模型,VLAT训练策略提供了实质性的改进,在看不见的长音频上实现了强大的性能。
摘要:Large Audio-Language Models (LALMs) are often constrained by short audio context windows, even when their text backbones support long contexts, limiting long-form audio understanding. Prior work has introduced context-extension methods (e.g. YaRN) on unimodal LLMs, yet their application to LALMs remains unexplored. First, building on RoPE-based context extension, we introduce Partial YaRN, a training-free, audio-only extension method that modifies only audio token positions, leaving text positions intact to preserve the base LLM's text capabilities. Second, we propose Virtual Longform Audio Training (VLAT), a training strategy that extends Partial YaRN into a training-time positional augmentation. VLAT simulates diverse audio lengths during training, enabling generalization to inputs far longer than those seen in training and improving robustness for long-context audio understanding. Our experiments on SALMONN and Qwen2-Audio show that Partial YaRN outperforms the original models across wide range of settings, and VLAT training strategy provides substantial improvement, achieving strong performance on long audio of unseen lengths.


机器翻译由腾讯交互翻译提供,仅供参考