今日论文合集:cs.SD语音7篇,eess.AS音频处理7篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音

【1】GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken  Chatbot

标题:GLM-4-语音:迈向智能且类人的端到端口语聊天机器人
链接:https://arxiv.org/abs/2412.02612
作者:Aohan Zeng,  Zhengxiao Du,  Mingdao Liu,  Kedong Wang,  Shengmin Jiang,  Lei Zhao,  Yuxiao Dong,  Jie Tang
摘要:我们介绍GLM-4-Voice,一个智能的、类似人类的端到端语音聊天机器人。它支持中文和英文,进行实时语音对话,并根据用户指示改变语音的细微差别,如情感,语调,语速和方言。GLM-4-Voice使用超低比特率(175 bps)、单码本语音标记器,其帧速率为12.5Hz,通过将矢量量化瓶颈纳入编码器,从自动语音识别(ASR)模型中获得。为了有效地将知识从文本转移到语音模态,我们使用文本到令牌模型从现有的文本预训练语料库中合成语音-文本交织数据。我们继续从预训练的文本语言模型GLM-4- 9 B进行预训练,结合无监督语音数据,交错语音文本数据和有监督语音文本数据,扩展到1万亿个令牌,在语音语言建模和口语问答方面实现最先进的性能。然后,我们使用高质量的对话语音数据对预训练模型进行微调,与现有的对话能力和语音质量基线相比,实现了更好的性能。开放模型可通过https://github.com/THUDM/GLM-4-Voice和https://huggingface.co/THUDM/glm-4-voice-9b访问。
摘要:We introduce GLM-4-Voice, an intelligent and human-like end-to-end spokenchatbot. It supports both Chinese and English, engages in real-time voiceconversations, and varies vocal nuances such as emotion, intonation, speechrate, and dialect according to user instructions. GLM-4-Voice uses an ultra-lowbitrate (175bps), single-codebook speech tokenizer with 12.5Hz frame ratederived from an automatic speech recognition (ASR) model by incorporating avector-quantized bottleneck into the encoder. To efficiently transfer knowledgefrom text to speech modalities, we synthesize speech-text interleaved data fromexisting text pre-training corpora using a text-to-token model. We continuepre-training from the pre-trained text language model GLM-4-9B with acombination of unsupervised speech data, interleaved speech-text data, andsupervised speech-text data, scaling up to 1 trillion tokens, achievingstate-of-the-art performance in both speech language modeling and spokenquestion answering. We then fine-tune the pre-trained model with high-qualityconversational speech data, achieving superior performance compared to existingbaselines in both conversational ability and speech quality. The open modelscan be accessed through https://github.com/THUDM/GLM-4-Voice andhttps://huggingface.co/THUDM/glm-4-voice-9b.

【2】 AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand  Audio-Visual Information?
标题:AV-Odyssey Bench:您的多模式LLM真的能理解视听信息吗?
链接:https://arxiv.org/abs/2412.02611
作者:Kaixiong Gong,  Kaituo Feng,  Bohao Li,  Yibing Wang,  Mofan Cheng,  Shijia Yang,  Jiaming Han,  Benyou Wang,  Yutong Bai,  Zhuoran Yang,  Xiangyu Yue
备注:Project page: this https URL
摘要:最近,GPT-4 o、Gemini 1.5 Pro和Reka Core等多模态大型语言模型(MLLM)已扩展其功能,包括视觉和音频模态。虽然这些模型在广泛的视听应用中表现出令人印象深刻的性能,但我们提出的DeafTest表明,MLLM经常在人类认为微不足道的简单任务中挣扎:1)确定两种声音中哪一种声音更大,2)确定两种声音中哪一种声音具有更高的音调。出于这些观察,我们介绍了AV奥德赛板凳,一个全面的视听基准,旨在评估这些MLLM是否可以真正理解的视听信息。这个基准测试包含4,555个精心设计的问题,每个问题都包含文本,视觉和音频组件。为了成功推断答案,模型必须有效地利用来自视觉和音频输入的线索。为了确保MLLM响应的准确和客观的评估,我们已经将问题结构化为多项选择,消除了对人类评估或LLM辅助评估的需要。我们对一系列闭源和开源模型进行基准测试,并总结观察结果。通过揭示当前模型的局限性,我们的目标是为未来的数据集收集和模型开发提供有用的见解。
摘要:Recently, multimodal large language models (MLLMs), such as GPT-4o, Gemini1.5 Pro, and Reka Core, have expanded their capabilities to include vision andaudio modalities. While these models demonstrate impressive performance acrossa wide range of audio-visual applications, our proposed DeafTest reveals thatMLLMs often struggle with simple tasks humans find trivial: 1) determiningwhich of two sounds is louder, and 2) determining which of two sounds has ahigher pitch. Motivated by these observations, we introduce AV-Odyssey Bench, acomprehensive audio-visual benchmark designed to assess whether those MLLMs cantruly understand the audio-visual information. This benchmark encompasses 4,555carefully crafted problems, each incorporating text, visual, and audiocomponents. To successfully infer answers, models must effectively leverageclues from both visual and audio inputs. To ensure precise and objectiveevaluation of MLLM responses, we have structured the questions asmultiple-choice, eliminating the need for human evaluation or LLM-assistedassessment. We benchmark a series of closed-source and open-source models andsummarize the observations. By revealing the limitations of current models, weaim to provide useful insight for future dataset collection and modeldevelopment.

【3】 It Takes Two: Real-time Co-Speech Two-person's Interaction Generation  via Reactive Auto-regressive Diffusion Model
标题:需要两个:通过反应式自回归扩散模型实时共语音两人互动生成
链接:https://arxiv.org/abs/2412.02419
作者:Mingyi Shi,  Dafei Qin,  Leo Ho,  Zhouyingcheng Liao,  Yinghao Huang,  Junichi Yamagishi,  Taku Komura
备注:15 pages, 10 figures
摘要:对话场景在现实世界中非常常见,然而现有的共同语音运动合成方法在这些上下文中往往不尽如人意,其中一个人的音频和手势将影响另一个人的响应。此外,大多数现有的方法依赖于离线序列到序列框架,这不适合在线应用。在这项工作中,我们介绍了一个音频驱动的,自回归系统,旨在合成动态运动的两个字符在对话。我们的方法的核心是一个基于扩散的全身运动合成模型,它以两个角色的过去状态、语音音频和面向任务的运动轨迹输入为条件,允许灵活的空间控制。为了增强模型学习各种交互的能力,我们用更动态和交互式的动作丰富了现有的两人对话动作数据集。我们通过多个实验来评估我们的系统,以显示它在各种任务中的表现,包括单人和双人共同语音运动生成,以及交互式运动生成。据我们所知,这是第一个能够以在线方式从语音中为两个角色生成交互式全身运动的系统。
摘要:Conversational scenarios are very common in real-world settings, yet existingco-speech motion synthesis approaches often fall short in these contexts, whereone person's audio and gestures will influence the other's responses.Additionally, most existing methods rely on offline sequence-to-sequenceframeworks, which are unsuitable for online applications. In this work, weintroduce an audio-driven, auto-regressive system designed to synthesizedynamic movements for two characters during a conversation. At the core of ourapproach is a diffusion-based full-body motion synthesis model, which isconditioned on the past states of both characters, speech audio, and atask-oriented motion trajectory input, allowing for flexible spatial control.To enhance the model's ability to learn diverse interactions, we have enrichedexisting two-person conversational motion datasets with more dynamic andinteractive motions. We evaluate our system through multiple experiments toshow it outperforms across a variety of tasks, including single and two-personco-speech motion generation, as well as interactive motion generation. To thebest of our knowledge, this is the first system capable of generatinginteractive full-body motions for two characters from speech in an onlinemanner.

【4】 Switchable deep beamformer for high-quality and real-time passive  acoustic mapping
标题:可切换深度射束形成器,用于高质量和实时无源声学映射
链接:https://arxiv.org/abs/2412.02327
作者:Yi Zeng,  Jinwei Li,  Hui Zhu,  Shukuan Lu,  Jianfeng Li,  Xiran Cai
摘要:被动声标测(PAM)是超声治疗应用中监测声空化活动的一种很有前途的工具。PAM的数据自适应波束形成器与时间曝光声学(TEA)算法相比具有更好的图像质量。然而,数据自适应波束形成器的计算成本相当昂贵。在这项工作中,我们开发了一种基于生成对抗网络的深度波束形成器,它可以在不同的换能器阵列之间切换,并以低计算成本直接从射频超声信号重建高质量的PAM图像。在由覆盖1-15 MHz的不同(线性和相控)阵列测量的单个和多个微泡云的模拟和实验空化信号组成的数据集上训练深波束形成器。我们使用模拟和实验测试数据集将深度波束形成器与TEA和三种不同的数据自适应波束形成器的性能进行了比较。与TEA相比,在我们的数据中,对于不同的阵列,深波束形成器减少了18.9%-65.0%的能量扩展区域,并且平均提高了9.3- 22.9dB的图像信噪比。与数据自适应波束形成器相比,深度波束形成器将计算成本降低了三个数量级,在我们的数据中实现了10.5 ms的图像重建速度,而图像质量与数据自适应波束形成器一样好。这些结果表明,深波束形成器的高分辨率监测超声治疗的微泡空化活动的潜力。
摘要:Passive acoustic mapping (PAM) is a promising tool for monitoring acousticcavitation activities in the applications of ultrasound therapy. Data-adaptivebeamformers for PAM have better image quality compared to the time exposureacoustics (TEA) algorithms. However, the computational cost of data-adaptivebeamformers is considerably expensive. In this work, we develop a deepbeamformer based on a generative adversarial network, which can switch betweendifferent transducer arrays and reconstruct high-quality PAM images directlyfrom radio frequency ultrasound signals with low computational cost. The deepbeamformer was trained on the dataset consisting of simulated and experimentalcavitation signals of single and multiple microbubble clouds measured bydifferent (linear and phased) arrays covering 1-15 MHz. We compared theperformance of the deep beamformer to TEA and three different data-adaptivebeamformers using the simulated and experimental test dataset. Compared withTEA, the deep beamformer reduced the energy spread area by 18.9%-65.0% andimproved the image signal-to-noise ratio by 9.3-22.9 dB in average for thedifferent arrays in our data. Compared to the data-adaptive beamformers, thedeep beamformer reduced the computational cost by three orders of magnitudeachieving 10.5 ms image reconstruction speed in our data, while the imagequality was as good as that of the data-adaptive beamformers. These resultsdemonstrated the potential of the deep beamformer for high-resolutionmonitoring of microbubble cavitation activities for ultrasound therapy.

【5】 A Theoretical Framework for Acoustic Neighbor Embeddings
标题:声学邻居嵌入的理论框架
链接:https://arxiv.org/abs/2412.02164
作者:Woojay Jeon
摘要:本文提供了一个理论框架来解释声学邻居嵌入,这是一个固定的嵌入空间中的可变宽度的音频或文本的语音内容的表示。嵌入之间的距离的概率解释,提出了一个通用的定量定义的基础上的语音相似的话。这为我们以原则性的方式理解和应用嵌入提供了一个框架。理论和经验证据,以支持近似的均匀群集明智的各向同性,这使我们能够减少简单的欧几里得距离的距离。四个实验,验证了框架,并展示了它如何可以应用到不同的问题。音频和文本嵌入之间的最近邻搜索可以提供与有限状态转换器(FST)相同的孤立词分类精度,用于大到500k的词汇表。嵌入距离提供的准确性与0.5%点的差异相比,手机编辑距离的词汇外的单词恢复,以及产生聚类层次结构相同的人听实验中的英语方言聚类。该理论框架还允许我们使用嵌入来预测设备唤醒词的预期混淆。提供所有源代码和预训练模型。
摘要:This paper provides a theoretical framework for interpreting acousticneighbor embeddings, which are representations of the phonetic content ofvariable-width audio or text in a fixed-dimensional embedding space. Aprobabilistic interpretation of the distances between embeddings is proposed,based on a general quantitative definition of phonetic similarity betweenwords. This provides us a framework for understanding and applying theembeddings in a principled manner. Theoretical and empirical evidence tosupport an approximation of uniform cluster-wise isotropy are shown, whichallows us to reduce the distances to simple Euclidean distances. Fourexperiments that validate the framework and demonstrate how it can be appliedto diverse problems are described. Nearest-neighbor search between audio andtext embeddings can give isolated word classification accuracy that isidentical to that of finite state transducers (FSTs) for vocabularies as largeas 500k. Embedding distances give accuracy with 0.5% point difference comparedto phone edit distances in out-of-vocabulary word recovery, as well asproducing clustering hierarchies identical to those derived from humanlistening experiments in English dialect clustering. The theoretical frameworkalso allows us to use the embeddings to predict the expected confusion ofdevice wake-up words. All source code and pretrained models are provided.

【6】 A Machine Hearing System for Robust Cough Detection Based on a  High-Level Representation of Band-Specific Audio Features
标题:基于特定频段音频特征的高级表示的鲁棒咳嗽检测机器听力系统
链接:https://arxiv.org/abs/2412.01996
作者:Jesús Monge-Alvarez,  Carlos Hoyos-Barceló,  Luis M. San-José-Revuelta,  Pablo Casaseca-de-la-Higuera
备注:12 pages, 11 figures, 5 tables
摘要:咳嗽是一种保护性反射,传达有关呼吸系统状态的信息。到目前为止,咳嗽评估仅限于主观测量工具或不舒服(即,不可穿戴的)咳嗽监测器。这限制了实时咳嗽监测改善呼吸护理的潜力。目的:本文提出了一种基于音频的机器听觉系统,可以很容易地部署在移动场景中的强大的咳嗽分割。方法:咳嗽检测分两步进行。首先,在五个预定义频带中分别计算短期频谱特征集:[0,0.5)、[0.5,1)、[1,1.5)、[1.5,2)和[2,5.5125] kHz。然后应用特征选择和组合,以使短期特征集在不同的噪声场景中足够鲁棒。第二,高层次的数据表示是通过计算300毫秒的长期帧中的短期描述符的平均值和标准差。最后,咳嗽检测进行使用支持向量机训练的数据从不同的嘈杂的情况下。该系统使用患者信号数据库进行评估,该数据库在噪声内容方面模拟了三种真实场景。结果如下:该系统实现了92.71%的灵敏度,88.58%的特异性和90.69%的受试者工作特征(ROC)曲线下面积(AUC),优于最先进的方法。结论:我们的研究成果为在现实生活中创建咳嗽监测设备铺平了道路。重要性:我们的建议与更舒适和更少干扰的患者监测相一致,对患者(允许自我监测咳嗽症状),从业者(例如,评估治疗或更好地临床了解咳嗽模式)和国家卫生系统(减少住院)。
摘要:Cough is a protective reflex conveying information on the state of therespiratory system. Cough assessment has been limited so far to subjectivemeasurement tools or uncomfortable (i.e., non-wearable) cough monitors. Thislimits the potential of real-time cough monitoring to improve respiratory care.Objective: This paper presents a machine hearing system for audio-based robustcough segmentation that can be easily deployed in mobile scenarios. Methods:Cough detection is performed in two steps. First, a short-term spectral featureset is separately computed in five predefined frequency bands: [0, 0.5), [0.5,1), [1, 1.5), [1.5, 2), and [2, 5.5125] kHz. Feature selection and combinationare then applied to make the short-term feature set robust enough in differentnoisy scenarios. Second, high-level data representation is achieved bycomputing the mean and standard deviation of short-term descriptors in 300 mslong-term frames. Finally, cough detection is carried out using a supportvector machine trained with data from different noisy scenarios. The system isevaluated using a patient signal database which emulates three real-lifescenarios in terms of noise content. Results: The system achieves 92.71%sensitivity, 88.58% specificity, and 90.69% Area Under Receiver OperatingCharacteristic (ROC) curve (AUC), outperforming state-of-the-art methods.Conclusion: Our research outcome paves the way to create a device for coughmonitoring in real-life situations. Significance: Our proposal is aligned witha more comfortable and less disruptive patient monitoring, with benefits forpatients (allows self-monitoring of cough symptoms), practitioners (e.g.,assessment of treatments or better clinical understanding of cough patterns),and national health systems (by reducing hospitalizations).

【7】 Late fusion ensembles for speech recognition on diverse input audio  representations
标题:用于对不同输入音频表示进行语音识别的后期融合集成
链接:https://arxiv.org/abs/2412.01861
作者:Marin Jezidžić,  Matej Mihelčić
摘要:我们探讨了不同的语音音频表示,以及它们对E-Branchformer模型后期融合集成性能的影响,应用于自动语音识别(ASR)任务。虽然众所周知,集成方法通常可以提高系统的性能,即使是语音识别,但探索复杂的最先进模型(如中型和大型E-Branchformer)的集成如何应对这种设置是非常有趣的,因为它们的基础模型是在输入语音音频的不同表示上训练的。结果在四个广泛使用的基准数据集上进行了评估:\textit{Librispeech,Aishell,Gigaspeech},\textit{TEDLIUMv 2},并表明与使用这些数据集上的可比技术训练的最先进的模型相比,仍然可以实现1\% - 14\%$的改进。一个值得注意的观察是,即使使用语言模型,这种集成也提供了改进,尽管差距正在缩小。
摘要:We explore diverse representations of speech audio, and their effect on aperformance of late fusion ensemble of E-Branchformer models, applied toAutomatic Speech Recognition (ASR) task. Although it is generally known thatensemble methods often improve the performance of the system even for speechrecognition, it is very interesting to explore how ensembles of complexstate-of-the-art models, such as medium-sized and large E-Branchformers, copein this setting when their base models are trained on diverse representationsof the input speech audio. The results are evaluated on four widely-usedbenchmark datasets: \textit{Librispeech, Aishell, Gigaspeech},\textit{TEDLIUMv2} and show that improvements of $1\% - 14\%$ can still beachieved over the state-of-the-art models trained using comparable techniqueson these datasets. A noteworthy observation is that such ensemble offersimprovements even with the use of language models, although the gap is closing.

eess.AS音频处理

【1】 A Theoretical Framework for Acoustic Neighbor Embeddings
标题:声学邻居嵌入的理论框架
链接:https://arxiv.org/abs/2412.02164
作者:Woojay Jeon
摘要:本文提供了一个理论框架来解释声学邻居嵌入,这是一个固定的嵌入空间中的可变宽度的音频或文本的语音内容的表示。嵌入之间的距离的概率解释,提出了一个通用的定量定义的基础上的语音相似的话。这为我们以原则性的方式理解和应用嵌入提供了一个框架。理论和经验证据,以支持近似的均匀群集明智的各向同性,这使我们能够减少简单的欧几里得距离的距离。四个实验,验证了框架,并展示了它如何可以应用到不同的问题。音频和文本嵌入之间的最近邻搜索可以提供与有限状态转换器(FST)相同的孤立词分类精度,用于大到500k的词汇表。嵌入距离提供的准确性与0.5%点的差异相比,手机编辑距离的词汇外的单词恢复,以及产生聚类层次结构相同的人听实验中的英语方言聚类。该理论框架还允许我们使用嵌入来预测设备唤醒词的预期混淆。提供所有源代码和预训练模型。
摘要:This paper provides a theoretical framework for interpreting acousticneighbor embeddings, which are representations of the phonetic content ofvariable-width audio or text in a fixed-dimensional embedding space. Aprobabilistic interpretation of the distances between embeddings is proposed,based on a general quantitative definition of phonetic similarity betweenwords. This provides us a framework for understanding and applying theembeddings in a principled manner. Theoretical and empirical evidence tosupport an approximation of uniform cluster-wise isotropy are shown, whichallows us to reduce the distances to simple Euclidean distances. Fourexperiments that validate the framework and demonstrate how it can be appliedto diverse problems are described. Nearest-neighbor search between audio andtext embeddings can give isolated word classification accuracy that isidentical to that of finite state transducers (FSTs) for vocabularies as largeas 500k. Embedding distances give accuracy with 0.5% point difference comparedto phone edit distances in out-of-vocabulary word recovery, as well asproducing clustering hierarchies identical to those derived from humanlistening experiments in English dialect clustering. The theoretical frameworkalso allows us to use the embeddings to predict the expected confusion ofdevice wake-up words. All source code and pretrained models are provided.

【2】 A Machine Hearing System for Robust Cough Detection Based on a  High-Level Representation of Band-Specific Audio Features
标题:基于特定频段音频特征的高级表示的鲁棒咳嗽检测机器听力系统
链接:https://arxiv.org/abs/2412.01996
作者:Jesús Monge-Alvarez,  Carlos Hoyos-Barceló,  Luis M. San-José-Revuelta,  Pablo Casaseca-de-la-Higuera
备注:12 pages, 11 figures, 5 tables
摘要:咳嗽是一种保护性反射,传递有关呼吸系统状态的信息。到目前为止,咳嗽评估仅限于主观测量工具或不舒服(即,不可穿戴的)咳嗽监测器。这限制了实时咳嗽监测改善呼吸道护理的潜力。目的:本文提出了一种基于音频的机器听觉系统,可以很容易地部署在移动场景中的强大的咳嗽分割。方法:咳嗽检测分两步进行。首先,在五个预定义频带中分别计算短期频谱特征集:[0,0.5)、[0.5,1)、[1,1.5)、[1.5,2)和[2,5.5125] kHz。然后应用特征选择和组合,以使短期特征集在不同的噪声场景中足够鲁棒。第二,高层次的数据表示是通过计算300毫秒的长期帧中的短期描述符的平均值和标准差。最后,咳嗽检测进行使用支持向量机训练的数据从不同的嘈杂的情况下。该系统使用患者信号数据库进行评估,该数据库在噪声内容方面模拟了三种真实场景。结果如下:该系统实现了92.71%的灵敏度,88.58%的特异性和90.69%的受试者工作特征(ROC)曲线下面积(AUC),优于最先进的方法。结论:我们的研究成果为在现实生活中创建咳嗽监测设备铺平了道路。重要性:我们的建议与更舒适和更少干扰的患者监测相一致,对患者(允许自我监测咳嗽症状),从业者(例如,评估治疗或更好地临床了解咳嗽模式)和国家卫生系统(减少住院)。
摘要:Cough is a protective reflex conveying information on the state of therespiratory system. Cough assessment has been limited so far to subjectivemeasurement tools or uncomfortable (i.e., non-wearable) cough monitors. Thislimits the potential of real-time cough monitoring to improve respiratory care.Objective: This paper presents a machine hearing system for audio-based robustcough segmentation that can be easily deployed in mobile scenarios. Methods:Cough detection is performed in two steps. First, a short-term spectral featureset is separately computed in five predefined frequency bands: [0, 0.5), [0.5,1), [1, 1.5), [1.5, 2), and [2, 5.5125] kHz. Feature selection and combinationare then applied to make the short-term feature set robust enough in differentnoisy scenarios. Second, high-level data representation is achieved bycomputing the mean and standard deviation of short-term descriptors in 300 mslong-term frames. Finally, cough detection is carried out using a supportvector machine trained with data from different noisy scenarios. The system isevaluated using a patient signal database which emulates three real-lifescenarios in terms of noise content. Results: The system achieves 92.71%sensitivity, 88.58% specificity, and 90.69% Area Under Receiver OperatingCharacteristic (ROC) curve (AUC), outperforming state-of-the-art methods.Conclusion: Our research outcome paves the way to create a device for coughmonitoring in real-life situations. Significance: Our proposal is aligned witha more comfortable and less disruptive patient monitoring, with benefits forpatients (allows self-monitoring of cough symptoms), practitioners (e.g.,assessment of treatments or better clinical understanding of cough patterns),and national health systems (by reducing hospitalizations).

【3】 Late fusion ensembles for speech recognition on diverse input audio  representations
标题:用于对不同输入音频表示进行语音识别的后期融合集成
链接:https://arxiv.org/abs/2412.01861
作者:Marin Jezidžić,  Matej Mihelčić
摘要:我们探讨了不同的语音音频表示,以及它们对E-Branchformer模型后期融合集成性能的影响,应用于自动语音识别(ASR)任务。虽然众所周知,集成方法通常可以提高系统的性能,即使是语音识别,但探索复杂的最先进模型(如中型和大型E-Branchformer)的集成如何应对这种设置是非常有趣的,因为它们的基础模型是在输入语音音频的不同表示上训练的。结果在四个广泛使用的基准数据集上进行了评估:\textit{Librispeech,Aishell,Gigaspeech},\textit{TEDLIUMv 2},并表明与使用这些数据集上的可比技术训练的最先进的模型相比,仍然可以实现1\% - 14\%$的改进。一个值得注意的观察是,即使使用语言模型,这种集成也提供了改进,尽管差距正在缩小。
摘要:We explore diverse representations of speech audio, and their effect on aperformance of late fusion ensemble of E-Branchformer models, applied toAutomatic Speech Recognition (ASR) task. Although it is generally known thatensemble methods often improve the performance of the system even for speechrecognition, it is very interesting to explore how ensembles of complexstate-of-the-art models, such as medium-sized and large E-Branchformers, copein this setting when their base models are trained on diverse representationsof the input speech audio. The results are evaluated on four widely-usedbenchmark datasets: \textit{Librispeech, Aishell, Gigaspeech},\textit{TEDLIUMv2} and show that improvements of $1\% - 14\%$ can still beachieved over the state-of-the-art models trained using comparable techniqueson these datasets. A noteworthy observation is that such ensemble offersimprovements even with the use of language models, although the gap is closing.

【4】 GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken  Chatbot
标题:GLM-4-语音:迈向智能且类人的端到端口语聊天机器人
链接:https://arxiv.org/abs/2412.02612
作者:Aohan Zeng,  Zhengxiao Du,  Mingdao Liu,  Kedong Wang,  Shengmin Jiang,  Lei Zhao,  Yuxiao Dong,  Jie Tang
摘要:我们介绍GLM-4-Voice,一个智能的、类似人类的端到端语音聊天机器人。它支持中文和英文,进行实时语音对话,并根据用户指示改变语音的细微差别,如情感,语调,语速和方言。GLM-4-Voice使用超低比特率(175 bps)、单码本语音标记器,其帧速率为12.5Hz,通过将矢量量化瓶颈纳入编码器,从自动语音识别(ASR)模型中获得。为了有效地将知识从文本转移到语音模态,我们使用文本到令牌模型从现有的文本预训练语料库中合成语音-文本交织数据。我们继续从预训练的文本语言模型GLM-4- 9 B进行预训练,结合无监督语音数据,交错语音文本数据和有监督语音文本数据,扩展到1万亿个令牌,在语音语言建模和口语问答方面实现最先进的性能。然后,我们使用高质量的对话语音数据对预训练模型进行微调,与现有的对话能力和语音质量基线相比,实现了更好的性能。开放模型可通过https://github.com/THUDM/GLM-4-Voice和https://huggingface.co/THUDM/glm-4-voice-9b访问。
摘要:We introduce GLM-4-Voice, an intelligent and human-like end-to-end spokenchatbot. It supports both Chinese and English, engages in real-time voiceconversations, and varies vocal nuances such as emotion, intonation, speechrate, and dialect according to user instructions. GLM-4-Voice uses an ultra-lowbitrate (175bps), single-codebook speech tokenizer with 12.5Hz frame ratederived from an automatic speech recognition (ASR) model by incorporating avector-quantized bottleneck into the encoder. To efficiently transfer knowledgefrom text to speech modalities, we synthesize speech-text interleaved data fromexisting text pre-training corpora using a text-to-token model. We continuepre-training from the pre-trained text language model GLM-4-9B with acombination of unsupervised speech data, interleaved speech-text data, andsupervised speech-text data, scaling up to 1 trillion tokens, achievingstate-of-the-art performance in both speech language modeling and spokenquestion answering. We then fine-tune the pre-trained model with high-qualityconversational speech data, achieving superior performance compared to existingbaselines in both conversational ability and speech quality. The open modelscan be accessed through https://github.com/THUDM/GLM-4-Voice andhttps://huggingface.co/THUDM/glm-4-voice-9b.

【5】 AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand  Audio-Visual Information?
标题:AV-Odyssey Bench:您的多模式LLM真的能理解视听信息吗?
链接:https://arxiv.org/abs/2412.02611
作者:Kaixiong Gong,  Kaituo Feng,  Bohao Li,  Yibing Wang,  Mofan Cheng,  Shijia Yang,  Jiaming Han,  Benyou Wang,  Yutong Bai,  Zhuoran Yang,  Xiangyu Yue
备注:Project page: this https URL
摘要:最近,多模态大型语言模型(MLLM),如GPT-4o,Gemini 1.5 Pro和Reka Core,已经扩展了它们的功能,包括视觉和音频模态。虽然这些模型在广泛的视听应用中表现出令人印象深刻的性能,但我们提出的DeafTest表明,MLLM经常在人类认为微不足道的简单任务中挣扎:1)确定两种声音中哪一种声音更大,2)确定两种声音中哪一种声音具有更高的音调。出于这些观察,我们介绍了AV奥德赛板凳,一个全面的视听基准,旨在评估这些MLLM是否可以真正理解的视听信息。这个基准测试包含4,555个精心设计的问题,每个问题都包含文本,视觉和音频组件。为了成功地推断答案,模型必须有效地利用视觉和音频输入的线索。为了确保MLLM响应的准确和客观的评估,我们已经将问题结构化为多项选择,消除了对人类评估或LLM辅助评估的需要。我们对一系列闭源和开源模型进行基准测试并总结观察结果。通过揭示当前模型的局限性,我们的目标是为未来的数据集收集和模型开发提供有用的见解。
摘要:Recently, multimodal large language models (MLLMs), such as GPT-4o, Gemini1.5 Pro, and Reka Core, have expanded their capabilities to include vision andaudio modalities. While these models demonstrate impressive performance acrossa wide range of audio-visual applications, our proposed DeafTest reveals thatMLLMs often struggle with simple tasks humans find trivial: 1) determiningwhich of two sounds is louder, and 2) determining which of two sounds has ahigher pitch. Motivated by these observations, we introduce AV-Odyssey Bench, acomprehensive audio-visual benchmark designed to assess whether those MLLMs cantruly understand the audio-visual information. This benchmark encompasses 4,555carefully crafted problems, each incorporating text, visual, and audiocomponents. To successfully infer answers, models must effectively leverageclues from both visual and audio inputs. To ensure precise and objectiveevaluation of MLLM responses, we have structured the questions asmultiple-choice, eliminating the need for human evaluation or LLM-assistedassessment. We benchmark a series of closed-source and open-source models andsummarize the observations. By revealing the limitations of current models, weaim to provide useful insight for future dataset collection and modeldevelopment.

【6】 It Takes Two: Real-time Co-Speech Two-person's Interaction Generation  via Reactive Auto-regressive Diffusion Model
标题:需要两个:通过反应式自回归扩散模型实时共语音两人互动生成
链接:https://arxiv.org/abs/2412.02419
作者:Mingyi Shi,  Dafei Qin,  Leo Ho,  Zhouyingcheng Liao,  Yinghao Huang,  Junichi Yamagishi,  Taku Komura
备注:15 pages, 10 figures
摘要:对话场景在现实世界中非常常见,然而现有的共同语音运动合成方法在这些上下文中往往不尽如人意,其中一个人的音频和手势将影响另一个人的响应。此外,大多数现有的方法依赖于离线序列到序列框架,这不适合在线应用。在这项工作中,我们介绍了一个音频驱动的,自回归系统,旨在合成动态运动的两个字符在对话。我们的方法的核心是一个基于扩散的全身运动合成模型,它以两个角色的过去状态、语音音频和面向任务的运动轨迹输入为条件,允许灵活的空间控制。为了增强模型学习各种交互的能力,我们用更动态和交互式的动作丰富了现有的两人对话动作数据集。我们通过多个实验来评估我们的系统,以显示它在各种任务中的表现,包括单人和双人共同语音运动生成,以及交互式运动生成。据我们所知,这是第一个能够以在线方式从语音中为两个角色生成交互式全身运动的系统。
摘要:Conversational scenarios are very common in real-world settings, yet existingco-speech motion synthesis approaches often fall short in these contexts, whereone person's audio and gestures will influence the other's responses.Additionally, most existing methods rely on offline sequence-to-sequenceframeworks, which are unsuitable for online applications. In this work, weintroduce an audio-driven, auto-regressive system designed to synthesizedynamic movements for two characters during a conversation. At the core of ourapproach is a diffusion-based full-body motion synthesis model, which isconditioned on the past states of both characters, speech audio, and atask-oriented motion trajectory input, allowing for flexible spatial control.To enhance the model's ability to learn diverse interactions, we have enrichedexisting two-person conversational motion datasets with more dynamic andinteractive motions. We evaluate our system through multiple experiments toshow it outperforms across a variety of tasks, including single and two-personco-speech motion generation, as well as interactive motion generation. To thebest of our knowledge, this is the first system capable of generatinginteractive full-body motions for two characters from speech in an onlinemanner.

【7】 Switchable deep beamformer for high-quality and real-time passive  acoustic mapping
标题:可切换深度射束形成器,用于高质量和实时无源声学映射
链接:https://arxiv.org/abs/2412.02327
作者:Yi Zeng,  Jinwei Li,  Hui Zhu,  Shukuan Lu,  Jianfeng Li,  Xiran Cai
摘要:被动声标测(PAM)是超声治疗应用中监测声空化活动的一种很有前途的工具。PAM的数据自适应波束形成器与时间曝光声学(TEA)算法相比具有更好的图像质量。然而,数据自适应波束形成器的计算成本相当昂贵。在这项工作中,我们开发了一种基于生成对抗网络的深度波束形成器,它可以在不同的换能器阵列之间切换,并以低计算成本直接从射频超声信号重建高质量的PAM图像。在由覆盖1-15 MHz的不同(线性和相控)阵列测量的单个和多个微泡云的模拟和实验空化信号组成的数据集上训练深波束形成器。我们使用模拟和实验测试数据集将深度波束形成器与TEA和三种不同的数据自适应波束形成器的性能进行了比较。与TEA相比,在我们的数据中,对于不同的阵列,深波束形成器减少了18.9%-65.0%的能量扩展区域,并且平均提高了9.3- 22.9dB的图像信噪比。与数据自适应波束形成器相比,深度波束形成器将计算成本降低了三个数量级,在我们的数据中实现了10.5 ms的图像重建速度,而图像质量与数据自适应波束形成器一样好。这些结果表明,深波束形成器的高分辨率监测超声治疗的微泡空化活动的潜力。
摘要:Passive acoustic mapping (PAM) is a promising tool for monitoring acousticcavitation activities in the applications of ultrasound therapy. Data-adaptivebeamformers for PAM have better image quality compared to the time exposureacoustics (TEA) algorithms. However, the computational cost of data-adaptivebeamformers is considerably expensive. In this work, we develop a deepbeamformer based on a generative adversarial network, which can switch betweendifferent transducer arrays and reconstruct high-quality PAM images directlyfrom radio frequency ultrasound signals with low computational cost. The deepbeamformer was trained on the dataset consisting of simulated and experimentalcavitation signals of single and multiple microbubble clouds measured bydifferent (linear and phased) arrays covering 1-15 MHz. We compared theperformance of the deep beamformer to TEA and three different data-adaptivebeamformers using the simulated and experimental test dataset. Compared withTEA, the deep beamformer reduced the energy spread area by 18.9%-65.0% andimproved the image signal-to-noise ratio by 9.3-22.9 dB in average for thedifferent arrays in our data. Compared to the data-adaptive beamformers, thedeep beamformer reduced the computational cost by three orders of magnitudeachieving 10.5 ms image reconstruction speed in our data, while the imagequality was as good as that of the data-adaptive beamformers. These resultsdemonstrated the potential of the deep beamformer for high-resolutionmonitoring of microbubble cavitation activities for ultrasound therapy.

机器翻译由腾讯交互翻译提供,仅供参考