今日论文合集:cs.SD语音10篇,eess.AS音频处理5篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
标题:人类还是机器?语音交互的初步图灵测试
链接:https://arxiv.org/pdf/2602.24080v1

作者:Xiang Li,Jiabao Gao,Sipei Lin,Xuan Zhou,Chi Zhang,Bo Cheng,Jiale Han,Benyou Wang
备注:Accepted by ICLR 2026 Conference
摘要:对类人会话代理的追求一直受到图灵测试的指导。对于现代语音到语音(S2S)系统,一个关键但尚未回答的问题是它们是否可以像人类一样交谈。为了解决这个问题,我们对S2S系统进行了第一次图灵测试,收集了9个最先进的S2S系统和28个人类参与者之间对话的2,968个人类判断。我们的结果提供了一个明确的发现:没有现有的评估S2S系统通过测试,揭示了人类相似性的显着差距。为了诊断这种失败,我们开发了一个细粒度的分类法,包括18个人类相似的维度,并相应地对我们收集的对话进行人群注释。我们的分析表明,瓶颈不是语义理解,而是源于语言特征,情感表达,和会话人物。此外,我们发现现成的AI模型在图灵测试判断中表现不可靠。作为回应,我们提出了一个可解释的模型,该模型利用细粒度的人类相似性评级,并提供准确和透明的人机识别,为自动人类相似性评估提供了一个强大的工具。我们的工作为S2S系统建立了第一个类似人类的评估,并超越了二元结果,以实现详细的诊断见解,为对话AI系统的类似人类的改进铺平了道路。摘要:The pursuit of human-like conversational agents has long been guided by the Turing test. For modern speech-to-speech (S2S) systems, a critical yet unanswered question is whether they can converse like humans. To tackle this, we conduct the first Turing test for S2S systems, collecting 2,968 human judgments on dialogues between 9 state-of-the-art S2S systems and 28 human participants. Our results deliver a clear finding: no existing evaluated S2S system passes the test, revealing a significant gap in human-likeness. To diagnose this failure, we develop a fine-grained taxonomy of 18 human-likeness dimensions and crowd-annotate our collected dialogues accordingly. Our analysis shows that the bottleneck is not semantic understanding but stems from paralinguistic features, emotional expressivity, and conversational persona. Furthermore, we find that off-the-shelf AI models perform unreliably as Turing test judges. In response, we propose an interpretable model that leverages the fine-grained human-likeness ratings and delivers accurate and transparent human-vs-machine discrimination, offering a powerful tool for automatic human-likeness evaluation. Our work establishes the first human-likeness evaluation for S2S systems and moves beyond binary outcomes to enable detailed diagnostic insights, paving the way for human-like improvements in conversational AI systems.


【2】SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away
标题:宋颂:千年来中国宋词音乐的时间留声机
链接:https://arxiv.org/pdf/2602.24071v1

作者:Jiajia Li,Jiliang Hu,Ziyi Pan,Chong Chen,Zuchao Li,Ping Wang,Lefei Zhang
备注:9 pages, 6 figures, accepted by AAAI 2025
摘要:近年来,音乐创作取得了重大进展。然而,现有的模式主要集中在创造现代流行歌曲,使其具有挑战性的生产具有独特的节奏和风格的古代音乐,如中国古代宋词。在本文中,我们介绍了宋词,第一个音乐生成模型能够恢复我们所知道的中国宋词。我们的模型首先从输入的SongCi中预测旋律,然后基于该旋律分别生成演唱声音和伴奏,最后组合所有元素来创建最终的音乐作品。此外,为了解决缺乏古代音乐数据集的问题,我们创建了OpenSongSong,这是一个全面的中国古代宋词音乐数据集,包含29.9小时的各种着名宋词音乐大师的作品。为了评估SongSong在演唱宋词方面的熟练程度,我们随机选择了85个不属于训练集的宋词句子,对SongSong和Suno和SkyMusic等音乐生成平台进行评估。主观和客观的结果表明,我们提出的模型在生成高质量的宋词音乐方面取得了领先的性能。摘要:Recently, there have been significant advancements in music generation. However, existing models primarily focus on creating modern pop songs, making it challenging to produce ancient music with distinct rhythms and styles, such as ancient Chinese SongCi. In this paper, we introduce SongSong, the first music generation model capable of restoring Chinese SongCi to our knowledge. Our model first predicts the melody from the input SongCi, then separately generates the singing voice and accompaniment based on that melody, and finally combines all elements to create the final piece of music. Additionally, to address the lack of ancient music datasets, we create OpenSongSong, a comprehensive dataset of ancient Chinese SongCi music, featuring 29.9 hours of compositions by various renowned SongCi music masters. To assess SongSong's proficiency in performing SongCi, we randomly select 85 SongCi sentences that were not part of the training set for evaluation against SongSong and music generation platforms such as Suno and SkyMusic. The subjective and objective outcomes indicate that our proposed model achieves leading performance in generating high-quality SongCi music.


【3】SHINE: Sequential Hierarchical Integration Network for EEG and MEG
标题:SHINE:脑电和脑电的顺序分层集成网络
链接:https://arxiv.org/pdf/2602.23960v1

作者:Xiran Xu,Yujie Yan,Xihong Wu,Jing Chen

备注:ranked second at LibriBrain Competition 2025 https:neural-processing-lab.github.io2025-libribrain-competitionprizes

摘要:自然语音在大脑中的表现方式是认知神经科学面临的一个重大挑战,而皮层的跟随性反应在语音解码中起着核心作用。本文介绍了我们在2025年LibriBrain竞赛中进行语音检测任务的方法,该方法利用了来自一名聆听LibriVox有声读物的参与者超过50小时的脑磁图(MEG)信号。我们提出了脑电和脑磁序列分层集成网络(SHINE)从脑磁信号重建二进制语音沉默序列。在Extended Track中,我们进一步结合了语音包络和Mel频谱图的辅助重建来增强训练。结合SHINE与基线(BrainMagic,AWavNet,ConvConcatNet)的Ensemination方法在排行榜测试集上获得了0.9155(标准赛道)和0.9184(扩展赛道)的F1宏分数。摘要:How natural speech is represented in the brain constitutes a major challenge for cognitive neuroscience, with cortical envelope-following responses playing a central role in speech decoding. This paper presents our approach to the Speech Detection task in the LibriBrain Competition 2025, utilizing over 50 hours of magnetoencephalography (MEG) signals from a single participant listening to LibriVox audiobooks. We introduce the proposed Sequential Hierarchical Integration Network for EEG and MEG (SHINE) to reconstruct the binary speech-silence sequences from MEG signals. In the Extended Track, we further incorporated auxiliary reconstructions of speech envelopes and Mel spectrograms to enhance training. Ensemble methods combining SHINE with baselines (BrainMagic, AWavNet, ConvConcatNet) achieved F1-macro scores of 0.9155 (Standard Track) and 0.9184 (Extended Track) on the leaderboard test set.


【4】DashengTokenizer: One layer is enough for unified audio understanding and generation
标题:DashengTokenizer:一层即可实现统一音频理解和生成
链接:https://arxiv.org/pdf/2602.23765v1

作者:Heinrich Dinkel,Xingwei Sun,Gang Li,Jiahao Mei,Yadong Niu,Jizhong Liu,Xiyang Li,Yifan Liao,Jiahao Zhou,Junbo Zhang,Jian Luan
摘要:本文介绍了一个连续的音频分词器,设计用于联合使用的理解和生成任务。与传统的方法,其中训练声学标记,并随后整合冻结的语义知识,我们的方法颠倒了这种范式:我们利用冻结的语义特征和注入声学信息。在22个不同任务的线性评估中,我们的方法在保持有竞争力的音频重建质量的同时,显著优于以前的音频编解码器和音频编码器基线。值得注意的是,我们证明了这种声学注入提高了语音情感识别,音乐理解和声学场景分类等任务的性能。我们进一步评估了tokenizer在文本到音频(TTA),文本到音乐(TTM)和语音增强(SE)上的生成性能。我们的方法超过了标准的变分自动编码器(VAE)的TTA和TTM任务的方法,而其有效性SE强调其作为一个通用的音频编码器的能力。最后,我们的研究结果挑战了普遍的假设,即基于VAE的架构是音频合成的先决条件。检查点可在https: huggingface.co mispeech dashengtokenizer上找到。摘要:This paper introduces DashengTokenizer, a continuous audio tokenizer engineered for joint use in both understanding and generation tasks. Unlike conventional approaches, which train acoustic tokenizers and subsequently integrate frozen semantic knowledge, our method inverts this paradigm: we leverage frozen semantic features and inject acoustic information. In linear evaluation across 22 diverse tasks, our method outperforms previous audio codec and audio encoder baselines by a significant margin while maintaining competitive audio reconstruction quality. Notably, we demonstrate that this acoustic injection improves performance for tasks such as speech emotion recognition, music understanding, and acoustic scene classification. We further evaluate the tokenizer's generative performance on text-to-audio (TTA), text-to-music (TTM), and speech enhancement (SE). Our approach surpasses standard variational autoencoder (VAE)-based methods on TTA and TTM tasks, while its effectiveness on SE underscores its capabilities as a general-purpose audio encoder. Finally, our results challenge the prevailing assumption that VAE-based architectures are a prerequisite for audio synthesis. Checkpoints are available at https: huggingface.co mispeech dashengtokenizer.


【5】Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future Context
标题:在线注册双模式自我监督语音模型:缓解未来上下文的缺乏
链接:https://arxiv.org/pdf/2602.23702v1

作者:Keita Goto,Takashi Maekaku,Jin Sakuma,Jinchuan Tian,Yusuke Shinohara,Shinji Watanabe
备注:Accepted to ICASSP 2026
摘要:在离线和在线模式下联合预训练的双模式自监督语音模型(S3M)由于缺少未来上下文而在流媒体场景中存在注意力不匹配。为了应对这一挑战,我们提出了在线注册,在线模式下将可学习的令牌附加到每个块。这些标记充当不可见的未来帧的虚拟占位符,使模型能够补偿缺失的上下文,而不会引入额外的延迟。此外,我们引入了一个未来的预测损失,明确指导寄存器捕捉预测线索,从而丰富他们的能力,以保留未来的信息。LibriSpeech和域外基准测试的实验表明,在线寄存器始终减少离线和在线模式之间的性能差距,实现了LibriSpeech与160毫秒块的3.4%的相对改善,特别是在低延迟设置。摘要:Dual-mode self-supervised speech models (S3Ms), which jointly pre-trained in the offline and online mode, suffer from attention mismatch in streaming scenarios due to missing future context. To address this challenge, we proposed online registers, learnable tokens appended to each chunk in online mode. These tokens act as virtual placeholders for unseen future frames, enabling the model to compensate for missing context without introducing additional latency. Furthermore, we introduce a future prediction loss that explicitly guides the registers to capture predictive cues, thereby enriching their ability to retain future information. Experiments on LibriSpeech, and out-of-domain benchmarks demonstrate that online registers consistently reduce the performance gap between offline and online modes, achieving a 3.4% relative improvement on LibriSpeech with 160 ms chunks, especially in low-latency settings.


【6】AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech
标题:AudioCapBench:声音、音乐和语音音频字幕的快速评估
链接:https://arxiv.org/pdf/2602.23649v1

作者:Jielin Qiu,Jianguo Zhang,Zixiang Chen,Liangwei Yang,Ming Zhu,Juntao Tan,Haolin Chen,Wenting Zhao,Rithesh Murthy,Roshan Ram,Akshara Prabhakar,Shelby Heinecke,Caiming,Xiong,Silvio Savarese,Huan Wang
摘要:我们介绍AudioCapBench,用于评估大型多模态模型的音频字幕能力的基准。该方法涵盖了三个不同的音频领域,包括环境声音,音乐和语音,从已建立的数据集中抽取了1,000个精选的评估样本。我们使用基于参考的指标(METEOR,BLEU,ROUGE-L)和LLM-as-Judge框架评估了两个提供商(OpenAI,Google Gemini)的13个模型,该框架在三个正交维度上对预测进行评分: textit{准确性}(语义正确性), textit{完整性}(参考内容的覆盖率)和 textit{幻觉}(没有捏造的内容)。我们的研究结果表明,Gemini模型在整体字幕质量上通常优于OpenAI模型,Gemini~3~Pro获得了最高的总分(6.00 10),而OpenAI模型表现出较低的幻觉率。所有模型在语音字幕上表现最好,在音乐字幕上表现最差。我们发布了基准测试和评估代码,以促进可再现的音频理解研究。摘要:We introduce AudioCapBench, a benchmark for evaluating audio captioning capabilities of large multimodal models. method covers three distinct audio domains, including environmental sound, music, and speech, with 1,000 curated evaluation samples drawn from established datasets. We evaluate 13 models across two providers (OpenAI, Google Gemini) using both reference-based metrics (METEOR, BLEU, ROUGE-L) and an LLM-as-Judge framework that scores predictions on three orthogonal dimensions: textit{accuracy} (semantic correctness), textit{completeness} (coverage of reference content), and textit{hallucination} (absence of fabricated content). Our results reveal that Gemini models generally outperform OpenAI models on overall captioning quality, with Gemini~3~Pro achieving the highest overall score (6.00 10), while OpenAI models exhibit lower hallucination rates. All models perform best on speech captioning and worst on music captioning. We release the benchmark as well as evaluation code to facilitate reproducible audio understanding research.


【7】Leveraging large multimodal models for audio-video deepfake detection: a pilot study
标题:利用大型多模式模型进行音频视频深度伪造检测:一项试点研究
链接:https://arxiv.org/pdf/2602.23393v1

作者:Songjun Cao,Yuqi Li,Yunpeng Luo,Jianjun Yin,Long Ma
备注:5pages,ICASSP2026
摘要:视听深度伪造检测(AVD)越来越重要,因为现代生成器可以制作令人信服的语音和视频。目前大多数多模态检测器都是小型的,特定于任务的模型:它们在策划的测试中工作得很好,但规模很差,跨领域的推广能力很弱。我们引入了AV-LMMDetect,这是一个监督微调(SFT)的大型多模态模型,它将AVD转换为提示的是 否分类-“这个视频是真的还是假的?".它基于Qwen 2.5 Omni构建,联合分析音频和视频流以进行deepfake检测,并分为两个阶段进行训练:轻量级LoRA对齐,然后是视听编码器的全面微调。在FakeAVCeleb和Mavos-DD上,AV-LMMDetect匹配或超越了先前的方法,并在Mavos-DD数据集上开创了新的技术水平。摘要:Audio-visual deepfake detection (AVD) is increasingly important as modern generators can fabricate convincing speech and video. Most current multimodal detectors are small, task-specific models: they work well on curated tests but scale poorly and generalize weakly across domains. We introduce AV-LMMDetect, a supervised fine-tuned (SFT) large multimodal model that casts AVD as a prompted yes no classification - "Is this video real or fake?". Built on Qwen 2.5 Omni, it jointly analyzes audio and visual streams for deepfake detection and is trained in two stages: lightweight LoRA alignment followed by audio-visual encoder full fine-tuning. On FakeAVCeleb and Mavos-DD, AV-LMMDetect matches or surpasses prior methods and sets a new state of the art on Mavos-DD datasets.


【8】Task-Lens: Cross-Task Utility Based Speech Dataset Profiling for Low-Resource Indian Languages
标题:Task-Lens:针对低资源印度语言的基于跨任务实用程序的语音数据集剖析
链接:https://arxiv.org/pdf/2602.23388v1

作者:Swati Sharma,Divya V. Sharma,Anubha Gupta
备注:Accepted at LREC 2026
摘要:对包容性语音技术的需求不断增长,加大了对自然语言处理(NLP)研究的多语言数据集的需求。然而,对低资源语言中现有特定任务资源的有限认识阻碍了研究。这一挑战在印度等语言多样的国家尤为严峻。对现有的印度语音数据集进行跨任务分析可以缓解数据稀缺的挑战。这涉及到跨多个下游任务调查数据集的效用,而不是专注于单个任务。以前的调查通常为单个任务分类数据集,使全面的跨任务分析成为一个开放的机会。因此,我们提出了任务镜头,跨任务调查,评估准备的50个印度语音数据集,跨越26种语言的9个下游语音任务。首先,我们分析哪些数据集包含适合特定任务的元数据和属性。接下来,我们提出了与任务一致的增强功能,以解锁数据集,充分发挥其下游潜力。最后,我们确定了目前资源严重不足的任务和印度语言。我们的研究结果表明,许多印度语音数据集包含未开发的元数据,可以支持多个下游任务。通过揭示跨任务的联系和差距,Task-Lens使研究人员能够探索现有数据集的更广泛适用性,并优先考虑为服务不足的任务和语言创建数据集。摘要:The rising demand for inclusive speech technologies amplifies the need for multilingual datasets for Natural Language Processing (NLP) research. However, limited awareness of existing task-specific resources in low-resource languages hinders research. This challenge is especially acute in linguistically diverse countries, such as India. Cross-task profiling of existing Indian speech datasets can alleviate the data scarcity challenge. This involves investigating the utility of datasets across multiple downstream tasks rather than focusing on a single task. Prior surveys typically catalogue datasets for a single task, leaving comprehensive cross-task profiling as an open opportunity. Therefore, we propose Task-Lens, a cross-task survey that assesses the readiness of 50 Indian speech datasets spanning 26 languages for nine downstream speech tasks. First, we analyze which datasets contain metadata and properties suitable for specific tasks. Next, we propose task-aligned enhancements to unlock datasets to their full downstream potential. Finally, we identify tasks and Indian languages that are critically underserved by current resources. Our findings reveal that many Indian speech datasets contain untapped metadata that can support multiple downstream tasks. By uncovering cross-task linkages and gaps, Task-Lens enables researchers to explore the broader applicability of existing datasets and to prioritize dataset creation for underserved tasks and languages.


【9】Hello-Chat: Towards Realistic Social Audio Interactions
标题:你好聊天:迈向现实的社交音频互动
链接:https://arxiv.org/pdf/2602.23387v1

作者:Yueran Hou,Peilei Jia,Zihan Sun,Qihang Lu,Wenbing Yang,Yingming Gao,Ya Li,Jun Gao
摘要:大型音频语言模型(LALM)的最新进展在语音识别和翻译方面表现出卓越的性能。然而,现有的模型往往存在感知和表达之间的脱节,导致机器人的“阅读-演讲”风格缺乏真实人类互动的自发性和情感共鸣。在本报告中,我们介绍了Hello-Chat,这是一种为现实社交场景设计的端到端音频语言模型。通过利用大量真实对话数据集并采用模态交织训练策略,Hello-Chat在拟人化生成方面取得了突破。实验结果表明,我们的模型不仅在特定的音频理解任务上达到了最先进的(SOTA)性能,而且在韵律自然度和情感对齐方面也明显优于现有的基线,为下一代移情AI代理铺平了道路。摘要:Recent advancements in Large Audio Language Models (LALMs) have demonstrated exceptional performance in speech recognition and translation. However, existing models often suffer from a disconnect between perception and expression, resulting in a robotic "read-speech" style that lacks the spontaneity and emotional resonance of real human interaction. In this report, we introduce Hello-Chat, an end-to-end audio language model designed for realistic social scenarios. By leveraging a massive dataset of real-life conversations and employing a modality-interleaved training strategy, Hello-Chat achieves a breakthrough in anthropomorphic generation. Experimental results show that our model not only reaches state-of-the-art (SOTA) performance on specific audio understanding tasks but also significantly outperforms existing baselines in prosodic naturalness and emotional alignment, paving the way for the next generation of empathetic AI agents.


【10】An Empirical Analysis of Task-Induced Encoder Bias in Fréchet Audio Distance
标题:Fréchet音频距离中任务引发的编码器偏差的实证分析
链接:https://arxiv.org/pdf/2602.23958v1

作者:Wonwoo Jeong

备注:6 pages, 4 figures. Submitted to Interspeech 2026. Source code and evaluation pipeline are available at: https:github.comwonwoo-jeongfad-encoder-bias

摘要:Fréchet Audio Distance(FAD)是评估文本到音频生成的事实上的标准,但其分数取决于底层编码器的嵌入空间。编码器的训练任务决定了哪些声学特征被保留或丢弃,导致FAD继承系统性任务诱导的偏差。我们将评估分解为召回率,精度和对齐(分为语义和结构维度),使用对数尺度归一化进行公平的交叉编码器比较。在两个数据集上对六个编码器进行的受控实验揭示了四轴权衡:基于重建的AudioMAE导致精度敏感性; ASR训练的Whisper主导结构检测,但对信号退化视而不见;分类训练的VGGish最大限度地提高了语义检测,但惩罚了合法的类内变化。由于没有单一的编码器是一个通用的评估器,未来的指标必须转向评估本地编码器本质上与人类感知一致。摘要:Fréchet Audio Distance (FAD) is the de facto standard for evaluating text-to-audio generation, yet its scores depend on the underlying encoder's embedding space. An encoder's training task dictates which acoustic features are preserved or discarded, causing FAD to inherit systematic task-induced biases. We decompose evaluation into Recall, Precision, and Alignment (split into semantic and structural dimensions), using log-scale normalization for fair cross-encoder comparison. Controlled experiments on six encoders across two datasets reveal a four-axis trade-off: reconstruction-based AudioMAE leads precision sensitivity; ASR-trained Whisper dominates structural detection but is blind to signal degradation; classification-trained VGGish maximizes semantic detection but penalizes legitimate intra-class variation. Since no single encoder is a universal evaluator, future metrics must shift toward evaluation-native encoders intrinsically aligned with human perception.


eess.AS音频处理


【1】An Empirical Analysis of Task-Induced Encoder Bias in Fréchet Audio Distance
标题:Fréchet音频距离中任务引发的编码器偏差的实证分析
链接:https://arxiv.org/pdf/2602.23958v1

作者:Wonwoo Jeong

备注:6 pages, 4 figures. Submitted to Interspeech 2026. Source code and evaluation pipeline are available at: https:github.comwonwoo-jeongfad-encoder-bias

摘要:Fréchet Audio Distance(FAD)是评估文本到音频生成的事实上的标准,但其分数取决于底层编码器的嵌入空间。编码器的训练任务决定了哪些声学特征被保留或丢弃,导致FAD继承系统性任务诱导的偏差。我们将评估分解为召回率,精度和对齐(分为语义和结构维度),使用对数尺度归一化进行公平的交叉编码器比较。在两个数据集上对六个编码器进行的受控实验揭示了四轴权衡:基于重建的AudioMAE导致精度敏感性; ASR训练的Whisper主导结构检测,但对信号退化视而不见;分类训练的VGGish最大限度地提高了语义检测,但惩罚了合法的类内变化。由于没有单一的编码器是一个通用的评估器,未来的指标必须转向评估本地编码器本质上与人类感知一致。摘要:Fréchet Audio Distance (FAD) is the de facto standard for evaluating text-to-audio generation, yet its scores depend on the underlying encoder's embedding space. An encoder's training task dictates which acoustic features are preserved or discarded, causing FAD to inherit systematic task-induced biases. We decompose evaluation into Recall, Precision, and Alignment (split into semantic and structural dimensions), using log-scale normalization for fair cross-encoder comparison. Controlled experiments on six encoders across two datasets reveal a four-axis trade-off: reconstruction-based AudioMAE leads precision sensitivity; ASR-trained Whisper dominates structural detection but is blind to signal degradation; classification-trained VGGish maximizes semantic detection but penalizes legitimate intra-class variation. Since no single encoder is a universal evaluator, future metrics must shift toward evaluation-native encoders intrinsically aligned with human perception.


【2】Design of a Hands-Free Short-Range Intercommunication Device Using LoRa for Secure Field Communication
标题:一种基于LoRa的免提短距离安全通信设备的设计
链接:https://arxiv.org/pdf/2602.23924v1

作者:Ayush Kumar Agrawal,Soumendu Das,Jayendra Kumar
摘要:在战术、军事和灾害应对环境中,短程可靠和安全的通信是一个主要优先事项,因为传统的通信基础设施要么离线,要么容易被拦截。目前的VHF UHF无线电和软件定义无线电很受欢迎,但它们是大型设备,需要大量电力,因此不适合用作无缝免提使用的轻量级可穿戴设备。在本文中,设计和理论框架的一个微型的,基于LoRa加密的内部通信设备,可用于安全的现场通信范围为1- 1.5公里,在视距条件下提供。建议的系统包括一个语音激活采集块,数字音频压缩,嵌入式微控制器处理器和AES-128加密,然后通过LoRa协议进行低功耗传输。通过线性调频扩频调制利用长距离和低能量特性的能力,该系统保证了可靠的通信,同时具有低功耗和低电磁足迹。建议的通信范围的理论分析是合理的,使用的链路预算,证明在实际的传播条件下的通信范围的实用性。该架构侧重于基础设施不可知论、点对点安全以及可穿戴人体工程学。给定的方案显示了LoRa技术在其他传统物联网遥测范围内的可能性,并且可以进一步扩展到包括安全的战术语音通信平台。摘要:Short-range reliable and secure communication is a major priority in the tactical, military and disaster response settings where the traditional communication infrastructure is either off-line or prone to interception. Current VHF UHF radios and software-defined radios are popular but large-sized devices and require lots of power, making them not suitable to be used as lightweight wearable devices with seamless hand-free use. In this paper, the design and theoretical framework of a miniature, LoRa based encrypted intercommunication device that can be used in secure field communication over a range of 1-1.5km and under line-of-sight conditions is provided. The suggested system consists of a voice-activated acquisition block, digital audio compression, an embedded microcontroller processor, and AES-128 encryption followed by a low-power transmission via the LoRa protocol. Through the ability of chirp spread spectrum modulation to utilize the long-range and low-energy properties, the system is guaranteed reliable communications coupled with low power consumption and low electromagnetic footprint. The theoretical analysis of the proposed communication range is justified using a link-budget that justifies the practicability of the communication range in the real propagation conditions. This architecture focuses on infrastructural agnosticism, peer-to-peer security as well as wearable ergonomics. The given scheme shows the possibilities of LoRa technology in the scope of other traditional IoT telemetry, and it can be further extended to include secure tactical voice communication platforms.


【3】DashengTokenizer: One layer is enough for unified audio understanding and generation
标题:DashengTokenizer:一层即可实现统一音频理解和生成
链接:https://arxiv.org/pdf/2602.23765v1

作者:Heinrich Dinkel,Xingwei Sun,Gang Li,Jiahao Mei,Yadong Niu,Jizhong Liu,Xiyang Li,Yifan Liao,Jiahao Zhou,Junbo Zhang,Jian Luan
摘要:本文介绍了一个连续的音频分词器,设计用于联合使用的理解和生成任务。与传统的方法,其中训练声学标记,并随后整合冻结的语义知识,我们的方法颠倒了这种范式:我们利用冻结的语义特征和注入声学信息。在22个不同任务的线性评估中,我们的方法在保持有竞争力的音频重建质量的同时,显著优于以前的音频编解码器和音频编码器基线。值得注意的是,我们证明了这种声学注入提高了语音情感识别,音乐理解和声学场景分类等任务的性能。我们进一步评估了tokenizer在文本到音频(TTA),文本到音乐(TTM)和语音增强(SE)上的生成性能。我们的方法超过了标准的变分自动编码器(VAE)的TTA和TTM任务的方法,而其有效性SE强调其作为一个通用的音频编码器的能力。最后,我们的研究结果挑战了普遍的假设,即基于VAE的架构是音频合成的先决条件。检查点可在https: huggingface.co mispeech dashengtokenizer上找到。摘要:This paper introduces DashengTokenizer, a continuous audio tokenizer engineered for joint use in both understanding and generation tasks. Unlike conventional approaches, which train acoustic tokenizers and subsequently integrate frozen semantic knowledge, our method inverts this paradigm: we leverage frozen semantic features and inject acoustic information. In linear evaluation across 22 diverse tasks, our method outperforms previous audio codec and audio encoder baselines by a significant margin while maintaining competitive audio reconstruction quality. Notably, we demonstrate that this acoustic injection improves performance for tasks such as speech emotion recognition, music understanding, and acoustic scene classification. We further evaluate the tokenizer's generative performance on text-to-audio (TTA), text-to-music (TTM), and speech enhancement (SE). Our approach surpasses standard variational autoencoder (VAE)-based methods on TTA and TTM tasks, while its effectiveness on SE underscores its capabilities as a general-purpose audio encoder. Finally, our results challenge the prevailing assumption that VAE-based architectures are a prerequisite for audio synthesis. Checkpoints are available at https: huggingface.co mispeech dashengtokenizer.


【4】Task-Lens: Cross-Task Utility Based Speech Dataset Profiling for Low-Resource Indian Languages
标题:Task-Lens:针对低资源印度语言的基于跨任务实用程序的语音数据集剖析
链接:https://arxiv.org/pdf/2602.23388v1

作者:Swati Sharma,Divya V. Sharma,Anubha Gupta
备注:Accepted at LREC 2026
摘要:对包容性语音技术的需求不断增长,加大了对自然语言处理(NLP)研究的多语言数据集的需求。然而,对低资源语言中现有特定任务资源的有限认识阻碍了研究。这一挑战在印度等语言多样的国家尤为严峻。对现有的印度语音数据集进行跨任务分析可以缓解数据稀缺的挑战。这涉及到跨多个下游任务调查数据集的效用,而不是专注于单个任务。以前的调查通常为单个任务分类数据集,使全面的跨任务分析成为一个开放的机会。因此,我们提出了任务镜头,跨任务调查,评估准备的50个印度语音数据集,跨越26种语言的9个下游语音任务。首先,我们分析哪些数据集包含适合特定任务的元数据和属性。接下来,我们提出了与任务一致的增强功能,以解锁数据集,充分发挥其下游潜力。最后,我们确定了目前资源严重不足的任务和印度语言。我们的研究结果表明,许多印度语音数据集包含未开发的元数据,可以支持多个下游任务。通过揭示跨任务的联系和差距,Task-Lens使研究人员能够探索现有数据集的更广泛适用性,并优先考虑为服务不足的任务和语言创建数据集。摘要:The rising demand for inclusive speech technologies amplifies the need for multilingual datasets for Natural Language Processing (NLP) research. However, limited awareness of existing task-specific resources in low-resource languages hinders research. This challenge is especially acute in linguistically diverse countries, such as India. Cross-task profiling of existing Indian speech datasets can alleviate the data scarcity challenge. This involves investigating the utility of datasets across multiple downstream tasks rather than focusing on a single task. Prior surveys typically catalogue datasets for a single task, leaving comprehensive cross-task profiling as an open opportunity. Therefore, we propose Task-Lens, a cross-task survey that assesses the readiness of 50 Indian speech datasets spanning 26 languages for nine downstream speech tasks. First, we analyze which datasets contain metadata and properties suitable for specific tasks. Next, we propose task-aligned enhancements to unlock datasets to their full downstream potential. Finally, we identify tasks and Indian languages that are critically underserved by current resources. Our findings reveal that many Indian speech datasets contain untapped metadata that can support multiple downstream tasks. By uncovering cross-task linkages and gaps, Task-Lens enables researchers to explore the broader applicability of existing datasets and to prioritize dataset creation for underserved tasks and languages.


【5】Hello-Chat: Towards Realistic Social Audio Interactions
标题:你好聊天:迈向现实的社交音频互动
链接:https://arxiv.org/pdf/2602.23387v1

作者:Yueran Hou,Peilei Jia,Zihan Sun,Qihang Lu,Wenbing Yang,Yingming Gao,Ya Li,Jun Gao
摘要:大型音频语言模型(LALM)的最新进展在语音识别和翻译方面表现出卓越的性能。然而,现有的模型往往存在感知和表达之间的脱节,导致机器人的“阅读-演讲”风格缺乏真实人类互动的自发性和情感共鸣。在本报告中,我们介绍了Hello-Chat,这是一种为现实社交场景设计的端到端音频语言模型。通过利用大量真实对话数据集并采用模态交织训练策略,Hello-Chat在拟人化生成方面取得了突破。实验结果表明,我们的模型不仅在特定的音频理解任务上达到了最先进的(SOTA)性能,而且在韵律自然度和情感对齐方面也明显优于现有的基线,为下一代移情AI代理铺平了道路。摘要:Recent advancements in Large Audio Language Models (LALMs) have demonstrated exceptional performance in speech recognition and translation. However, existing models often suffer from a disconnect between perception and expression, resulting in a robotic "read-speech" style that lacks the spontaneity and emotional resonance of real human interaction. In this report, we introduce Hello-Chat, an end-to-end audio language model designed for realistic social scenarios. By leveraging a massive dataset of real-life conversations and employing a modality-interleaved training strategy, Hello-Chat achieves a breakthrough in anthropomorphic generation. Experimental results show that our model not only reaches state-of-the-art (SOTA) performance on specific audio understanding tasks but also significantly outperforms existing baselines in prosodic naturalness and emotional alignment, paving the way for the next generation of empathetic AI agents.


机器翻译由腾讯交互翻译提供,仅供参考