本文经arXiv每日学术速递授权转载
【1】 A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
标题:具有大型语言模型的语义相关任务的离散语音标记的比较研究
链接:https://arxiv.org/abs/2411.08742
作者:Dingdong Wang, Mingyu Cui, Dongchao Yang, Xueyuan Chen, Helen Meng
备注:5 tables, 4 figures
摘要:随着语音大型语言模型(Speech LLM)的兴起,人们对离散语音令牌的兴趣越来越大,因为它们能够与基于文本的令牌无缝集成。与大多数专注于连续语音特征的研究相比,尽管基于离散令牌的LLM在某些任务上显示出有希望的结果,但很少探索这两种范式之间的性能差距。在本文中,我们使用轻量级LLM(Qwen1.5-0.5B)在各种语义相关任务中对离散和连续特征进行了公平而彻底的比较。我们的研究结果表明,连续特征的性能通常优于离散标记,特别是在需要细粒度语义理解的任务中。此外,这项研究超越了表面层面的比较,确定了离散令牌性能不佳背后的关键因素,例如有限的令牌粒度和低效的信息保留。为了提高离散令牌的性能,我们根据我们的分析探索了潜在的方面。我们希望我们的研究结果可以提供新的见解的机会,在语音LLM推进离散语音令牌。
摘要:With the rise of Speech Large Language Models (Speech LLMs), there has beengrowing interest in discrete speech tokens for their ability to integrate withtext-based tokens seamlessly. Compared to most studies that focus on continuousspeech features, although discrete-token based LLMs have shown promisingresults on certain tasks, the performance gap between these two paradigms israrely explored. In this paper, we present a fair and thorough comparisonbetween discrete and continuous features across a variety of semantic-relatedtasks using a light-weight LLM (Qwen1.5-0.5B). Our findings reveal thatcontinuous features generally outperform discrete tokens, particularly in tasksrequiring fine-grained semantic understanding. Moreover, this study goes beyondsurface-level comparison by identifying key factors behind theunder-performance of discrete tokens, such as limited token granularity andinefficient information retention. To enhance the performance of discretetokens, we explore potential aspects based on our analysis. We hope our resultscan offer new insights into the opportunities for advancing discrete speechtokens in Speech LLMs.
标题:开发有效的训练数据集以提高基于人工智能的说话人分离系统的性能
链接:https://arxiv.org/abs/2411.08375
备注:in Arabic language
摘要:本文讨论了说话人分离的挑战,这仍然是一个活跃的研究课题,尽管在最近几年取得了可喜的成果。然而,由于噪声、回波和其他干扰的存在,这些结果在实际记录条件下常常会退化。这是因为神经模型通常是在由混合音频信号及其相应的地面实况组成的合成数据集上训练的,这些数据集是使用计算机软件生成的,并且不能完全代表真实世界记录场景的复杂性。由于从混合音频信号中获得单个声音是一项重要任务,因此缺乏用于扬声器分离的真实训练集仍然是一个主要障碍。为了解决这个问题,我们提出了一种新的方法来构建一个现实的训练集,其中包括混合信号和相应的地面真理为每个扬声器。我们在深度学习模型上评估这个数据集,并将其与合成数据集进行比较。我们在真实混音中的比例不变信号失真比(SI-SDR)中获得了1.65 dB的扬声器分离精度提高。我们的研究结果突出了现实训练集在现实场景中增强说话人分离模型性能的潜力。
摘要:This paper addresses the challenge of speaker separation, which remains anactive research topic despite the promising results achieved in recent years.These results, however, often degrade in real recording conditions due to thepresence of noise, echo, and other interferences. This is because neural modelsare typically trained on synthetic datasets consisting of mixed audio signalsand their corresponding ground truths, which are generated using computersoftware and do not fully represent the complexities of real-world recordingscenarios. The lack of realistic training sets for speaker separation remains amajor hurdle, as obtaining individual sounds from mixed audio signals is anontrivial task. To address this issue, we propose a novel method forconstructing a realistic training set that includes mixture signals andcorresponding ground truths for each speaker. We evaluate this dataset on adeep learning model and compare it to a synthetic dataset. We got a 1.65 dBimprovement in Scale Invariant Signal to Distortion Ratio (SI-SDR) for speakerseparation accuracy in realistic mixing. Our findings highlight the potentialof realistic training sets for enhancing the performance of speaker separationmodels in real-world scenarios.
标题:评估对智能语音助理的合成命令攻击
链接:https://arxiv.org/abs/2411.08316
摘要:语音合成的最新进展,加上可以为数百万人收集语音的容易性,给由诸如语音助理(例如,Amazon Alexa、Google Home等)。我们探索来自目标的不相关和有限数量的语音是否可以用于为亚马逊Alexa等语音助手合成命令。更具体地说,我们调查了当语音助理将命令源与授权用户和应用程序(例如,Alexa Skills)仅在其来源是具有选定置信水平的授权用户时才处理命令。我们证明,即使是简单的拼接语音合成可以被攻击者用来命令语音助手执行敏感的操作。我们还表明,当通过利用语音助手附近的受损设备发起此类攻击时,主机和网络占用空间相对较小。我们的研究结果表明,需要更好地防御可能针对语音助手的合成恶意命令。
摘要:Recent advances in voice synthesis, coupled with the ease with which speechcan be harvested for millions of people, introduce new threats to applicationsthat are enabled by devices such as voice assistants (e.g., Amazon Alexa,Google Home etc.). We explore if unrelated and limited amount of speech from atarget can be used to synthesize commands for a voice assistant like AmazonAlexa. More specifically, we investigate attacks on voice assistants withsynthetic commands when they match command sources to authorized users, andapplications (e.g., Alexa Skills) process commands only when their source is anauthorized user with a chosen confidence level. We demonstrate that even simpleconcatenative speech synthesis can be used by an attacker to command voiceassistants to perform sensitive operations. We also show that such attacks,when launched by exploiting compromised devices in the vicinity of voiceassistants, can have relatively small host and network footprint. Our resultsdemonstrate the need for better defenses against synthetic malicious commandsthat could target voice assistants.
标题:PerceiverS:一个具有有效分割的多尺度感知器,用于长期表达的象征性音乐生成
链接:https://arxiv.org/abs/2411.08307
摘要:音乐生成已经取得了显著的进步,特别是在音频生成领域。然而,生成既具有长期结构又具有表现力的象征性音乐仍然是一个重大挑战。在本文中,我们提出了PerceiverS(分割和规模),一种新的架构,旨在解决这个问题,利用有效的分割和多尺度注意力机制。我们的方法通过同时学习长期的结构依赖性和短期的表达细节来增强符号音乐的生成。通过在多尺度设置中结合交叉注意力和自我注意力,PerceiverS捕捉了长距离的音乐结构,同时保留了表演的细微差别。该模型在Maestro等数据集上进行了评估,证明了在生成具有结构一致性和表达变化的连贯和多样化音乐方面的改进。项目演示和生成的音乐样本可以通过链接访问:https://perceivers.github.io。
摘要:Music generation has progressed significantly, especially in the domain ofaudio generation. However, generating symbolic music that is bothlong-structured and expressive remains a significant challenge. In this paper,we propose PerceiverS (Segmentation and Scale), a novel architecture designedto address this issue by leveraging both Effective Segmentation and Multi-Scaleattention mechanisms. Our approach enhances symbolic music generation bysimultaneously learning long-term structural dependencies and short-termexpressive details. By combining cross-attention and self-attention in aMulti-Scale setting, PerceiverS captures long-range musical structure whilepreserving performance nuances. The proposed model, evaluated on datasets likeMaestro, demonstrates improvements in generating coherent and diverse musicwith both structural consistency and expressive variation. The project demosand the generated music samples can be accessed through the link:https://perceivers.github.io.
标题:加纳传统Seperewa歌曲的音调内容分析
链接:https://arxiv.org/abs/2411.08234
摘要:本研究探讨了传统的加纳seperewa(阿坎竖琴琵琶)歌曲的音高内容,利用一个独特的数据集,从二十世纪中叶的现场录音。我们选择了71首歌曲,并使用Demucs将人声与器乐曲目分离开来。然后,我们从这些孤立的大头钉中检索F0内容,并应用高斯混合模型(GMM)来近似音乐音阶。人声和seperewa之间的比较F0分析显示,在声乐曲目从平等的气质更高的微色调偏差。我们还注意到在非西方音乐中使用MIR工具进行音阶近似的挑战。我们的研究有助于撒哈拉以南非洲传统音乐音高的定量研究。
摘要:This study examines the pitch content in traditional Ghanaian seperewa (Akanharp-lute) songs, utilizing a unique dataset from field recordings of themid-twentieth century. We selected 71 songs and used Demucs to isolate vocalsfrom instrumental tracks. We then retrieved the F0 content from these isolatedtacks and applied Gaussian Mixture Models (GMM) to approximate musical scales.Comparative F0 analysis between vocals and seperewa revealed higher microtonaldeviations from equal temperament in vocal tracks. We also note challenges inusing MIR tools for musical scale approximation in non-Western music. Ourresearch contributes to the quantitative study of pitch in traditional music ofSub-Saharan Africa.
标题:语音数据在减少毒性检测偏差方面的作用
链接:https://arxiv.org/abs/2411.08135
摘要:文本毒性检测系统表现出显着的偏见,产生不成比例的假阳性率的样本提到人口统计学群体。但是言语中的毒性检测呢?为了研究基于语音的系统在多大程度上减轻了基于文本的偏见,我们为多语言MuTox数据集生成了一组高质量的组注释,然后利用这些注释系统地比较基于语音和基于文本的毒性分类器。我们的研究结果表明,在推理过程中访问语音数据支持减少对群体提及的偏见,特别是对于模棱两可和不同意诱导样本。我们的研究结果还表明,改进分类器,而不是转录管道,更有助于减少群体偏见。我们公开发布我们的注释,并为未来的毒性数据集构建提供建议。
摘要:Text toxicity detection systems exhibit significant biases, producingdisproportionate rates of false positives on samples mentioning demographicgroups. But what about toxicity detection in speech? To investigate the extentto which text-based biases are mitigated by speech-based systems, we produce aset of high-quality group annotations for the multilingual MuTox dataset, andthen leverage these annotations to systematically compare speech- andtext-based toxicity classifiers. Our findings indicate that access to speechdata during inference supports reduced bias against group mentions,particularly for ambiguous and disagreement-inducing samples. Our results alsosuggest that improving classifiers, rather than transcription pipelines, ismore helpful for reducing group bias. We publicly release our annotations andprovide recommendations for future toxicity dataset construction.
标题:使用基于房间声学模型的先验空间空间估计空间动态房间脉冲响应
链接:https://arxiv.org/abs/2411.08477
备注:30 pages, 13 figures
摘要:静态扬声器和麦克风位置之间的房间脉冲响应(RIR)的估计可以使用许多完善的测量和推断过程来完成。虽然这些过程假设时不变的声学系统,但对于扬声器和麦克风受到移动的空间动态场景的情况,需要考虑时间变化。如果使用图像源对RIR进行建模,则移动意味着到每个图像源的距离随时间变化,使得空间动态RIR的估计特别具有挑战性。在本文中,我们提出了一个程序来估计的早期部分的空间动态RIR之间的固定源和麦克风移动的线性轨迹以恒定的速度。该过程是建立在一个状态空间模型,其中要估计的状态表示早期RIR,观察对应于麦克风记录在空间动态的情况下,和随时间变化的距离的图像源被纳入从静态RIR在轨迹的起点和终点的状态转移矩阵。所提出的方法的性能进行评估对国家的最先进的RIR插值和状态空间估计方法,使用模拟,展示了所提出的状态空间模型的潜力。
摘要:The estimation of room impulse responses (RIRs) between static loudspeakerand microphone locations can be done using a number of well-establishedmeasurement and inference procedures. While these procedures assume atime-invariant acoustic system, time variations need to be considered for thecase of spatially dynamic scenarios where loudspeakers and microphones aresubject to movement. If the RIR is modeled using image sources, then movementimplies that the distance to each image source varies over time, making theestimation of the spatially dynamic RIR particularly challenging. In thispaper, we propose a procedure to estimate the early part of the spatiallydynamic RIR between a stationary source and a microphone moving on a lineartrajectory at constant velocity. The procedure is built upon a state-spacemodel, where the state to be estimated represents the early RIR, theobservation corresponds to a microphone recording in a spatially dynamicscenario, and time-varying distances to the image sources are incorporated intothe state transition matrix obtained from static RIRs at the start and endpoint of the trajectory. The performance of the proposed approach is evaluatedagainst state-of-the-art RIR interpolation and state-space estimation methodsusing simulations, demonstrating the potential of the proposed state-spacemodel.
标题:具有大型语言模型的语义相关任务的离散语音标记的比较研究
链接:https://arxiv.org/abs/2411.08742
备注:5 tables, 4 figures
摘要:随着语音大型语言模型(Speech LLM)的兴起,人们对离散语音令牌的兴趣越来越大,因为它们能够与基于文本的令牌无缝集成。与大多数专注于连续语音特征的研究相比,尽管基于离散令牌的LLM在某些任务上显示出有希望的结果,但很少探索这两种范式之间的性能差距。在本文中,我们使用轻量级LLM(Qwen1.5-0.5B)在各种语义相关任务中对离散和连续特征进行了公平而彻底的比较。我们的研究结果表明,连续特征的性能通常优于离散标记,特别是在需要细粒度语义理解的任务中。此外,这项研究超越了表面层面的比较,确定了离散令牌性能不佳背后的关键因素,例如有限的令牌粒度和低效的信息保留。为了提高离散令牌的性能,我们根据我们的分析探索了潜在的方面。我们希望我们的研究结果可以提供新的见解的机会,在语音LLM推进离散语音令牌。
摘要:With the rise of Speech Large Language Models (Speech LLMs), there has beengrowing interest in discrete speech tokens for their ability to integrate withtext-based tokens seamlessly. Compared to most studies that focus on continuousspeech features, although discrete-token based LLMs have shown promisingresults on certain tasks, the performance gap between these two paradigms israrely explored. In this paper, we present a fair and thorough comparisonbetween discrete and continuous features across a variety of semantic-relatedtasks using a light-weight LLM (Qwen1.5-0.5B). Our findings reveal thatcontinuous features generally outperform discrete tokens, particularly in tasksrequiring fine-grained semantic understanding. Moreover, this study goes beyondsurface-level comparison by identifying key factors behind theunder-performance of discrete tokens, such as limited token granularity andinefficient information retention. To enhance the performance of discretetokens, we explore potential aspects based on our analysis. We hope our resultscan offer new insights into the opportunities for advancing discrete speechtokens in Speech LLMs.
标题:开发有效的训练数据集以提高基于人工智能的说话人分离系统的性能
链接:https://arxiv.org/abs/2411.08375
备注:in Arabic language
摘要:本文讨论了说话人分离的挑战,这仍然是一个活跃的研究课题,尽管在最近几年取得了可喜的成果。然而,由于噪声、回波和其他干扰的存在,这些结果在实际记录条件下常常会退化。这是因为神经模型通常是在由混合音频信号及其相应的地面实况组成的合成数据集上训练的,这些数据集是使用计算机软件生成的,并且不能完全代表真实世界记录场景的复杂性。由于从混合音频信号中获得单个声音是一项重要任务,因此缺乏用于扬声器分离的真实训练集仍然是一个主要障碍。为了解决这个问题,我们提出了一种新的方法来构建一个现实的训练集,其中包括混合信号和相应的地面真理为每个扬声器。我们在深度学习模型上评估这个数据集,并将其与合成数据集进行比较。我们在真实混音中的比例不变信号失真比(SI-SDR)中获得了1.65 dB的扬声器分离精度提高。我们的研究结果突出了现实训练集在现实场景中增强说话人分离模型性能的潜力。
摘要:This paper addresses the challenge of speaker separation, which remains anactive research topic despite the promising results achieved in recent years.These results, however, often degrade in real recording conditions due to thepresence of noise, echo, and other interferences. This is because neural modelsare typically trained on synthetic datasets consisting of mixed audio signalsand their corresponding ground truths, which are generated using computersoftware and do not fully represent the complexities of real-world recordingscenarios. The lack of realistic training sets for speaker separation remains amajor hurdle, as obtaining individual sounds from mixed audio signals is anontrivial task. To address this issue, we propose a novel method forconstructing a realistic training set that includes mixture signals andcorresponding ground truths for each speaker. We evaluate this dataset on adeep learning model and compare it to a synthetic dataset. We got a 1.65 dBimprovement in Scale Invariant Signal to Distortion Ratio (SI-SDR) for speakerseparation accuracy in realistic mixing. Our findings highlight the potentialof realistic training sets for enhancing the performance of speaker separationmodels in real-world scenarios.
标题:评估对智能语音助理的合成命令攻击
链接:https://arxiv.org/abs/2411.08316
摘要:语音合成的最新进展,加上可以为数百万人收集语音的容易性,给由诸如语音助理(例如,Amazon Alexa、Google Home等)。我们探索来自目标的不相关和有限数量的语音是否可以用于为亚马逊Alexa等语音助手合成命令。更具体地说,我们调查了当语音助理将命令源与授权用户和应用程序(例如,Alexa Skills)仅在其来源是具有选定置信水平的授权用户时才处理命令。我们证明,即使是简单的拼接语音合成可以被攻击者用来命令语音助手执行敏感的操作。我们还表明,当通过利用语音助手附近的受损设备发起此类攻击时,主机和网络占用空间相对较小。我们的研究结果表明,需要更好地防御可能针对语音助手的合成恶意命令。
摘要:Recent advances in voice synthesis, coupled with the ease with which speechcan be harvested for millions of people, introduce new threats to applicationsthat are enabled by devices such as voice assistants (e.g., Amazon Alexa,Google Home etc.). We explore if unrelated and limited amount of speech from atarget can be used to synthesize commands for a voice assistant like AmazonAlexa. More specifically, we investigate attacks on voice assistants withsynthetic commands when they match command sources to authorized users, andapplications (e.g., Alexa Skills) process commands only when their source is anauthorized user with a chosen confidence level. We demonstrate that even simpleconcatenative speech synthesis can be used by an attacker to command voiceassistants to perform sensitive operations. We also show that such attacks,when launched by exploiting compromised devices in the vicinity of voiceassistants, can have relatively small host and network footprint. Our resultsdemonstrate the need for better defenses against synthetic malicious commandsthat could target voice assistants.
