今日论文合集:cs.SD语音11篇,eess.AS音频处理15篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Modeling Time-Variant Responses of Optical Compressors with Selective State Space Models
标题: 用选择性状态空间模型建模光学压缩机时变响应
作者:Riccardo Simionato
链接:点击下载PDF文件
摘要:本文提出了一种使用具有选择性状态空间模型的深度神经网络对光学动态范围压缩器进行建模的方法。所提出的方法超越了以往的方法的基础上,通过采用一个选择性的状态空间块来编码输入音频的递归层。它采用了一种精细的技术,集成了智能线性调制和门控线性单元,以动态调整网络,根据外部参数调节压缩的攻击和释放阶段。所提出的架构非常适合低延迟和实时应用,在现场音频处理中至关重要。该方法已在TubeTech CL 1B和Teletronix LA-2A模拟光压缩机上得到验证,这两种压缩机具有不同的特性。使用定量指标和主观听力测试进行评估,比较所提出的方法与其他国家的最先进的模型。结果表明,我们的黑盒建模方法优于所有其他方法,在训练过程中实现了对可见和不可见设置的压缩过程的准确仿真。我们进一步展示了这种准确性与数据集中控制参数的采样密度之间的相关性,并将快速攻击和缓慢释放的设置确定为最具挑战性的模拟。摘要:This paper presents a method for modeling optical dynamic range compressors using deep neural networks with Selective State Space models. The proposed approach surpasses previous methods based on recurrent layers by employing a Selective State Space block to encode the input audio. It features a refined technique integrating Feature-wise Linear Modulation and Gated Linear Units to adjust the network dynamically, conditioning the compression's attack and release phases according to external parameters. The proposed architecture is well-suited for low-latency and real-time applications, crucial in live audio processing. The method has been validated on the analog optical compressors TubeTech CL 1B and Teletronix LA-2A, which possess distinct characteristics. Evaluation is performed using quantitative metrics and subjective listening tests, comparing the proposed method with other state-of-the-art models. Results show that our black-box modeling methods outperform all others, achieving accurate emulation of the compression process for both seen and unseen settings during training. We further show a correlation between this accuracy and the sampling density of the control parameters in the dataset and identify settings with fast attack and slow release as the most challenging to emulate.

【2】 WhisperMask: A Noise Suppressive Mask-Type Microphone for Whisper Speech
标题: WhisperMass:一款用于耳语的降噪口罩式麦克风
作者:Hirotaka Hiraki,Shusuke Kanazawa,Takahiro Miura,Manabu Yoshida,Masaaki Mochimaru,Jun Rekimoto
Journal-ref:Proceedings of the Augmented Humans International Conference 2024
链接:点击下载PDF文件
摘要:耳语是基于语音的交互中常见的隐私保护技术,但其有效性在嘈杂环境中受到限制。在传统的基于硬件和软件的降噪方法中,将低声语音与环境噪声和其他语音声音隔离仍然是一个挑战。因此,我们提出了WhisperMask,这是一种面具式麦克风,具有低灵敏度的大振膜,使佩戴者的声音明显高于背景噪音。我们使用三个关键指标来评估WhisperMask:信噪比,录制语音的质量和语音识别率。在所有指标中,WhisperMask始终优于传统的降噪麦克风和基于软件的解决方案。值得注意的是,与针式麦克风和耳塞相比,WhisperMask在80 dB背景噪声的环境中记录的耳语语音识别准确率高出30%。此外,在30-60 dB的噪声下,去噪器使这两款麦克风的耳语语音识别率降低了约20%,而WhisperMask在不去噪的情况下仍保持了高性能,远远超过其他麦克风的性能。WhisperMask的设计将佩戴者的声音作为主要输入,并有效抑制背景噪声,而不依赖于信号处理。该设备允许在各种嘈杂的现实世界场景中进行可靠的语音交互,例如电话呼叫和语音命令,同时保护用户隐私。摘要:Whispering is a common privacy-preserving technique in voice-based interactions, but its effectiveness is limited in noisy environments. In conventional hardware- and software-based noise reduction approaches, isolating whispered speech from ambient noise and other speech sounds remains a challenge. We thus propose WhisperMask, a mask-type microphone featuring a large diaphragm with low sensitivity, making the wearer's voice significantly louder than the background noise. We evaluated WhisperMask using three key metrics: signal-to-noise ratio, quality of recorded voices, and speech recognition rate. Across all metrics, WhisperMask consistently outperformed traditional noise-suppressing microphones and software-based solutions. Notably, WhisperMask showed a 30% higher recognition accuracy for whispered speech recorded in an environment with 80 dB background noise compared with the pin microphone and earbuds. Furthermore, while a denoiser decreased the whispered speech recognition rate of these two microphones by approximately 20% at 30-60 dB noise, WhisperMask maintained a high performance even without denoising, surpassing the other microphones' performances by a significant margin.WhisperMask's design renders the wearer's voice as the dominant input and effectively suppresses background noise without relying on signal processing. This device allows for reliable voice interactions, such as phone calls and voice commands, in a wide range of noisy real-world scenarios while preserving user privacy.

【3】 Self-Learning for Personalized Keyword Spotting on Ultra-Low-Power Audio Sensors
标题: 在超低功耗音频传感器上进行个性化关键词定位的自学习
作者:Manuele Rusci,Francesco Paci,Marco Fariselli,Eric Flamand,Tinne Tuytelaars
链接:点击下载PDF文件
摘要:本文提出了一种自学习框架,以增量训练(微调)的个性化关键词定位(KWS)模型部署后,超低功耗智能音频传感器。我们解决的基本问题,标记的训练数据的情况下,通过分配伪标签的基础上的相似性分数相对于几个用户记录的新记录的音频帧。通过在两个公共数据集上使用多个参数高达0.5M的KWS模型进行实验,我们发现,与在大量通用关键字上预训练的初始模型相比,准确率提高了+19.2%和+16.0%。标签的任务是由一个低功耗的麦克风和一个节能的微控制器(MCU)的传感器系统。通过有效利用MCU的异构处理引擎,始终在线的标记任务实时运行,平均功耗高达8.2 mW。在同一平台上,我们估计设备上训练的能量成本比使用DS-CNN-S或DS-CNN-M模型每5秒或16.4秒对新话语进行采样的标记能量低10倍。我们的实证结果为在极端边缘实现自适应个性化KWS传感器铺平了道路。摘要:This paper proposes a self-learning framework to incrementally train (fine-tune) a personalized Keyword Spotting (KWS) model after the deployment on ultra-low power smart audio sensors. We address the fundamental problem of the absence of labeled training data by assigning pseudo-labels to the new recorded audio frames based on a similarity score with respect to few user recordings. By experimenting with multiple KWS models with a number of parameters up to 0.5M on two public datasets, we show an accuracy improvement of up to +19.2% and +16.0% vs. the initial models pretrained on a large set of generic keywords. The labeling task is demonstrated on a sensor system composed of a low-power microphone and an energy-efficient Microcontroller (MCU). By efficiently exploiting the heterogeneous processing engines of the MCU, the always-on labeling task runs in real-time with an average power cost of up to 8.2 mW. On the same platform, we estimate an energy cost for on-device training 10x lower than the labeling energy if sampling a new utterance every 5 s or 16.4 s with a DS-CNN-S or a DS-CNN-M model. Our empirical result paves the way to self-adaptive personalized KWS sensors at the extreme edge.

【4】 Developing vocal system impaired patient-aimed voice quality assessment approach using ASR representation-included multiple features
标题: 使用包含ASB表示的多个特征开发针对发声系统受损患者的语音质量评估方法
作者:Shaoxiang Dang,Tetsuya Matsumoto,Yoshinori Takeuchi,Takashi Tsuboi,Yasuhiro Tanaka,Daisuke Nakatsubo,Satoshi Maesawa,Ryuta Saito,Masahisa Katsuno,Hiroaki Kudo
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:深度学习在临床语音处理中的潜力是巨大的,但有限和不平衡的临床数据样本的障碍迫在眉睫。本文通过展示自动语音识别和自监督学习表示的利用来解决这些挑战,这些表示在广泛的正常语音数据集上进行了预训练。这种创新的方法旨在评估声音系统受损患者的声音质量。实验涉及对PVQD数据集的检查,涵盖英语中发音系统损伤的各种原因,以及日本数据集,重点关注接受丘脑底核脑深部电刺激(STN-DBS)手术前后的帕金森病患者。PVQD的结果揭示了在预测等级、呼吸和虚弱指标方面的显著相关性(PCC 0.8)和非凡的准确性(MSE <0.5)。同时,在STN-DBS的背景下,在预测患者的语音质量方面取得了进展。摘要:The potential of deep learning in clinical speech processing is immense, yet the hurdles of limited and imbalanced clinical data samples loom large. This article addresses these challenges by showcasing the utilization of automatic speech recognition and self-supervised learning representations, pre-trained on extensive datasets of normal speech. This innovative approach aims to estimate voice quality of patients with impaired vocal systems. Experiments involve checks on PVQD dataset, covering various causes of vocal system damage in English, and a Japanese dataset focusing on patients with Parkinson's disease before and after undergoing subthalamic nucleus deep brain stimulation (STN-DBS) surgery. The results on PVQD reveal a notable correlation ( 0.8 on PCC) and an extraordinary accuracy (<0.5 on MSE) in predicting Grade, Breathy, and Asthenic indicators. Meanwhile, progress has been achieved in predicting the voice quality of patients in the context of STN-DBS.

【5】 VoiceX: A Text-To-Speech Framework for Custom Voices
标题: SecureX:自定义语音的文本到语音框架
作者:Silvan Mertes,Daksitha Withanage Don,Otto Grothe,Johanna Kuch,Ruben Schlagowski,Elisabeth André
链接:点击下载PDF文件
摘要:现代TTS系统能够创建高度逼真和自然的语音。尽管有这些发展,定制TTS语音的过程仍然是一项复杂的任务,主要需要该领域专家的专业知识。其中一个原因是深度学习模型的使用,其特点是其扩展的,不可解释的参数空间,限制了手动定制的可行性。在本文中,我们提出了一种新的人在循环的基础上直接与神经TTS模型的参数空间的进化算法的范例。我们将我们的方法集成到一个用户友好的图形用户界面中,允许用户有效地创建原始声音。然后,这些语音可以与主干TTS模型一起使用,我们为此提供了Python API。此外,我们提出了用户研究的结果,探索VoiceX的功能。我们表明,VoiceX是一个合适的工具,用于创建个人,自定义的声音。摘要:Modern TTS systems are capable of creating highly realistic and natural-sounding speech. Despite these developments, the process of customizing TTS voices remains a complex task, mostly requiring the expertise of specialists within the field. One reason for this is the utilization of deep learning models, which are characterized by their expansive, non-interpretable parameter spaces, restricting the feasibility of manual customization. In this paper, we present a novel human-in-the-loop paradigm based on an evolutionary algorithm for directly interacting with the parameter space of a neural TTS model. We integrated our approach into a user-friendly graphical user interface that allows users to efficiently create original voices. Those voices can then be used with the backbone TTS model, for which we provide a Python API. Further, we present the results of a user study exploring the capabilities of VoiceX. We show that VoiceX is an appropriate tool for creating individual, custom voices.

【6】 Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization
标题: 集成音频、视觉和语义信息以实现增强的多模式说话人Dialogation
作者:Luyao Cheng,Hui Wang,Siqi Zheng,Yafeng Chen,Rongjie Huang,Qinglin Zhang,Qian Chen,Xihao Li
链接:点击下载PDF文件
摘要:说话人日志化是将音频流或转录的语音内容分割成基于说话人身份的同质分区的过程,在人类语音的解释和分析中起着至关重要的作用。大多数现有的扬声器日记系统完全依赖于单峰声学信息,由于音频信号的固有模糊性,使得任务特别具有挑战性。最近的研究已经取得了巨大的努力,视听或听觉语义建模,以提高性能。然而,即使将多达两种模式结合起来,也往往无法解决自发和非结构化对话的复杂性。为了利用更有意义的对话模式,我们提出了一种新的多模态方法,联合利用音频,视觉和语义线索,以提高扬声器日记。我们的方法优雅地制定了多峰建模作为一个约束优化问题。首先,我们建立洞察到积极发言者之间的视觉连接和口语内容内的语义交互,从而建立丰富的成对约束。然后,我们引入了一种联合成对约束传播算法,根据这些视觉和语义约束对说话人进行聚类。这种集成有效地利用了不同模态的互补优势,改进了单个说话人嵌入之间的亲和力估计。在多个多模态数据集上进行的大量实验表明,我们的方法始终优于最先进的说话人日记化方法。摘要:Speaker diarization, the process of segmenting an audio stream or transcribed speech content into homogenous partitions based on speaker identity, plays a crucial role in the interpretation and analysis of human speech. Most existing speaker diarization systems rely exclusively on unimodal acoustic information, making the task particularly challenging due to the innate ambiguities of audio signals. Recent studies have made tremendous efforts towards audio-visual or audio-semantic modeling to enhance performance. However, even the incorporation of up to two modalities often falls short in addressing the complexities of spontaneous and unstructured conversations. To exploit more meaningful dialogue patterns, we propose a novel multimodal approach that jointly utilizes audio, visual, and semantic cues to enhance speaker diarization. Our method elegantly formulates the multimodal modeling as a constrained optimization problem. First, we build insights into the visual connections among active speakers and the semantic interactions within spoken content, thereby establishing abundant pairwise constraints. Then we introduce a joint pairwise constraint propagation algorithm to cluster speakers based on these visual and semantic constraints. This integration effectively leverages the complementary strengths of different modalities, refining the affinity estimation between individual speaker embeddings. Extensive experiments conducted on multiple multimodal datasets demonstrate that our approach consistently outperforms state-of-the-art speaker diarization methods.

【7】 Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
标题: Video-Foley:通过Foley Sound的时间事件条件的两阶段视频到声音生成
作者:Junwon Lee,Jaekwon Im,Dabin Kim,Juhan Nam
链接:点击下载PDF文件
摘要:Foley声音合成对于多媒体制作至关重要,通过在时间和语义上同步音频和视频来增强用户体验。最近关于通过视频到声音生成来自动化这一劳动密集型过程的研究面临着重大挑战。缺乏显式时间特征的系统具有较差的可控性和对齐性,而基于时间戳的模型需要昂贵且主观的人工注释。我们提出了视频福利,视频到声音系统使用均方根(RMS)作为一个时间事件条件与语义音色提示(音频或文本)。RMS是一种与音频语义密切相关的帧级强度包络特征,可确保高可控性和同步性。无注释的自监督学习框架由两个阶段组成,Video 2 RMS和RMS 2Sound,结合了新的想法,包括RMS离散化和RMS-ControlNet以及预训练的文本到音频模型。我们广泛的评估表明,Video-Foley在视听对齐和声音定时,强度,音色和细微差别的可控性方面实现了最先进的性能。代码、模型重量和演示可在附带的网站上获得。(https: jnwnlee.github.io video-foley-demo)摘要:Foley sound synthesis is crucial for multimedia production, enhancing user experience by synchronizing audio and video both temporally and semantically. Recent studies on automating this labor-intensive process through video-to-sound generation face significant challenges. Systems lacking explicit temporal features suffer from poor controllability and alignment, while timestamp-based models require costly and subjective human annotation. We propose Video-Foley, a video-to-sound system using Root Mean Square (RMS) as a temporal event condition with semantic timbre prompts (audio or text). RMS, a frame-level intensity envelope feature closely related to audio semantics, ensures high controllability and synchronization. The annotation-free self-supervised learning framework consists of two stages, Video2RMS and RMS2Sound, incorporating novel ideas including RMS discretization and RMS-ControlNet with a pretrained text-to-audio model. Our extensive evaluation shows that Video-Foley achieves state-of-the-art performance in audio-visual alignment and controllability for sound timing, intensity, timbre, and nuance. Code, model weights, and demonstrations are available on the accompanying website. (https: jnwnlee.github.io video-foley-demo)

【8】 Convexity-based Pruning of Speech Representation Models
标题: 基于凸度的语音表示模型修剪
作者:Teresa Dorszewski,Lenka Tětková,Lars Kai Hansen
链接:点击下载PDF文件
摘要:基于Transformer架构并通过自监督学习训练的语音表示模型在解决语音和说话人识别、关键字定位、情感检测等任务方面表现出了巨大的潜力。通常,人们发现更大的模型会带来更好的性能。然而,在这样的大型Transformer系统中涉及的显著计算工作量对于嵌入式和现实世界的应用是一个挑战。最近的工作已经表明,在用于NLP的Transformer模型中存在显著的冗余,并且大规模层修剪是可行的(Sajjad等人,2023年)。在这里,我们研究音频模型中的层修剪。我们基于凸性准则的修剪决策。分类区域的凸性最近被提出作为一系列应用领域(包括NLP和音频)中后续微调性能的指标。在实证研究中,我们发现在计算工作量的大量减少,没有损失的性能,甚至在某些情况下的改进。摘要:Speech representation models based on the transformer architecture and trained by self-supervised learning have shown great promise for solving tasks such as speech and speaker recognition, keyword spotting, emotion detection, and more. Typically, it is found that larger models lead to better performance. However, the significant computational effort involved in such large transformer systems is a challenge for embedded and real-world applications. Recent work has shown that there is significant redundancy in the transformer models for NLP and massive layer pruning is feasible (Sajjad et al., 2023). Here, we investigate layer pruning in audio models. We base the pruning decision on a convexity criterion. Convexity of classification regions has recently been proposed as an indicator of subsequent fine-tuning performance in a range of application domains, including NLP and audio. In empirical investigations, we find a massive reduction in the computational effort with no loss of performance or even improvements in certain cases.

【9】 Dynamic Gated Recurrent Neural Network for Compute-efficient Speech Enhancement
标题: 用于计算机高效语音增强的动态门控回归神经网络
作者:Longbiao Cheng,Ashutosh Pandey,Buye Xu,Tobi Delbruck,Shih-Chii Liu
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文介绍了一种新的动态门控递归神经网络(DG-RNN)的计算效率的语音增强模型运行在资源受限的硬件平台。它利用了RNN隐藏状态在步骤上的缓慢演化特性,并通过向RNN模型添加新提出的选择门来在每一步仅更新选定的神经元集合。该选择门允许在网络推理期间降低传统RNN的计算成本。作为DG-RNN的实现,我们进一步提出了动态门控递归单元(D-GRU),它不需要额外的参数。使用DNS挑战数据集从几种最先进的基于RNN的计算高效语音增强架构获得的测试结果表明,基于D-GRU的模型变体保持与基于GRU的基线模型相似的语音可懂度和质量指标,即使GRU计算平均减少50%。摘要:This paper introduces a new Dynamic Gated Recurrent Neural Network (DG-RNN) for compute-efficient speech enhancement models running on resource-constrained hardware platforms. It leverages the slow evolution characteristic of RNN hidden states over steps, and updates only a selected set of neurons at each step by adding a newly proposed select gate to the RNN model. This select gate allows the computation cost of the conventional RNN to be reduced during network inference. As a realization of the DG-RNN, we further propose the Dynamic Gated Recurrent Unit (D-GRU) which does not require additional parameters. Test results obtained from several state-of-the-art compute-efficient RNN-based speech enhancement architectures using the DNS challenge dataset, show that the D-GRU based model variants maintain similar speech intelligibility and quality metrics comparable to the baseline GRU based models even with an average 50% reduction in GRU computes.

【10】 LCM-SVC: Latent Diffusion Model Based Singing Voice Conversion with Inference Acceleration via Latent Consistency Distillation
标题: LCM-SRC:基于潜扩散模型的歌唱声音转换,通过潜稠度蒸馏进行推理加速
作者:Shihao Chen,Yu Gu,Jianwei Cui,Jie Zhang,Rilin Chen,Lirong Dai
备注:Accepted to ISCSLP 2024. arXiv admin note: text overlap with arXiv:2406.05325
链接:点击下载PDF文件
摘要:任何到任何歌唱声音转换(SVC)的目的是使用短的声音样本将目标歌手的音色转换为其他歌曲。然而,许多基于扩散模型的any-to-anySVC方法虽然取得了令人瞩目的效果,但由于推理步骤过多,效率低下。在本文中,我们提出了LCM-SVC,一个潜在的一致性蒸馏(LCD)的潜在扩散模型(LDM),以加快推理速度。我们通过提取预训练的基于LDM的SVC模型,实现了一步或几步推理,同时保持了高性能,该模型具有音色解耦和音质的优点。实验结果表明,我们提出的方法可以显着减少推理时间,并在很大程度上保持了音质和音色相似性相比,其他国家的最先进的SVC模型。音频样本可在https: sounddemos.github.io lcm-svc上获得。摘要:Any-to-any singing voice conversion (SVC) aims to transfer a target singer's timbre to other songs using a short voice sample. However many diffusion model based any-to-any SVC methods, which have achieved impressive results, usually suffered from low efficiency caused by a mass of inference steps. In this paper, we propose LCM-SVC, a latent consistency distillation (LCD) based latent diffusion model (LDM) to accelerate inference speed. We achieved one-step or few-step inference while maintaining the high performance by distilling a pre-trained LDM based SVC model, which had the advantages of timbre decoupling and sound quality. Experimental results show that our proposed method can significantly reduce the inference time and largely preserve the sound quality and timbre similarity comparing with other state-of-the-art SVC models. Audio samples are available at https: sounddemos.github.io lcm-svc.

【11】 Prosody of speech production in latent post-stroke aphasia
标题: 中风后隐性失语症的言语产生韵律
作者:Cong Zhang,Tong Li,Gayle DeDe,Christos Salis
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:本研究探讨了潜在性失语症(一种与左半球脑损伤(如中风)相关的轻度失语症)的韵律产生。与先前对中度至重度失语症的研究不同,我们研究了潜伏性失语症,这种失语症似乎与神经型言语具有非常相似的言语产生。我们分析了10名潜伏性失语症患者和10名正常对照者的首、尾词的f_0、强度和持续时间。回归模型拟合,以提高我们对这种未充分研究的非常轻度失语症的理解。结果突出了不同程度的差异,在所有三个韵律措施组之间。我们还研究了使用随机森林的潜伏性失语症与神经典型对照的诊断分类,旨在建立一个快速可靠的工具来帮助识别潜伏性失语症。随机森林分析也加强了韵律特征在区分潜在性失语中的意义。摘要:This study explores prosodic production in latent aphasia, a mild form of aphasia associated with left-hemisphere brain damage (e.g. stroke). Unlike prior research on moderate to severe aphasia, we investigated latent aphasia, which can seem to have very similar speech production with neurotypical speech. We analysed the f0, intensity and duration of utterance-initial and utterance-final words of ten speakers with latent aphasia and ten matching controls. Regression models were fitted to improve our understanding of this understudied type of very mild aphasia. The results highlighted varying degrees of differences in all three prosodic measures between groups. We also investigated the diagnostic classification of latent aphasia versus neurotypical control using random forest, aiming to build a fast and reliable tool to assist with the identification of latent aphasia. The random forest analysis also reinforced the significance of prosodic features in distinguishing latent aphasia.


eess.AS音频处理
【1】 Dynamic Gated Recurrent Neural Network for Compute-efficient Speech Enhancement
标题: 用于计算机高效语音增强的动态门控回归神经网络
作者:Longbiao Cheng,Ashutosh Pandey,Buye Xu,Tobi Delbruck,Shih-Chii Liu
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文介绍了一种新的动态门控循环神经网络(DG-RNN),用于在资源受限的硬件平台上运行的计算高效的语音增强模型。它利用了RNN隐藏状态在步骤上的缓慢演化特性,并通过向RNN模型添加新提出的选择门来在每一步仅更新选定的神经元集合。该选择门允许在网络推理期间降低传统RNN的计算成本。作为DG-RNN的实现,我们进一步提出了动态门控递归单元(D-GRU),它不需要额外的参数。使用DNS挑战数据集从几种最先进的计算高效的基于RNN的语音增强架构中获得的测试结果表明,基于D-GRU的模型变体保持与基于GRU的基线模型相似的语音清晰度和质量指标,即使平均减少50% GRU计算。摘要:This paper introduces a new Dynamic Gated Recurrent Neural Network (DG-RNN) for compute-efficient speech enhancement models running on resource-constrained hardware platforms. It leverages the slow evolution characteristic of RNN hidden states over steps, and updates only a selected set of neurons at each step by adding a newly proposed select gate to the RNN model. This select gate allows the computation cost of the conventional RNN to be reduced during network inference. As a realization of the DG-RNN, we further propose the Dynamic Gated Recurrent Unit (D-GRU) which does not require additional parameters. Test results obtained from several state-of-the-art compute-efficient RNN-based speech enhancement architectures using the DNS challenge dataset, show that the D-GRU based model variants maintain similar speech intelligibility and quality metrics comparable to the baseline GRU based models even with an average 50% reduction in GRU computes.

【2】 LCM-SVC: Latent Diffusion Model Based Singing Voice Conversion with Inference Acceleration via Latent Consistency Distillation
标题: LCM-SRC:基于潜扩散模型的歌唱声音转换,通过潜稠度蒸馏进行推理加速
作者:Shihao Chen,Yu Gu,Jianwei Cui,Jie Zhang,Rilin Chen,Lirong Dai
备注:Accepted to ISCSLP 2024. arXiv admin note: text overlap with arXiv:2406.05325
链接:点击下载PDF文件
摘要:任何到任何歌唱声音转换(SVC)的目的是使用短的声音样本将目标歌手的音色转换为其他歌曲。然而,许多基于扩散模型的any-to-anySVC方法虽然取得了令人瞩目的效果,但由于推理步骤过多,效率低下。在本文中,我们提出了LCM-SVC,一个潜在的一致性蒸馏(LCD)为基础的潜在扩散模型(LDM),以加快推理速度。我们通过提取预训练的基于LDM的SVC模型,实现了一步或几步推理,同时保持了高性能,该模型具有音色解耦和音质的优点。实验结果表明,我们提出的方法可以显着减少推理时间,并在很大程度上保持了音质和音色相似性相比,其他国家的最先进的SVC模型。音频样本可在https: sounddemos.github.io lcm-svc上获得。摘要:Any-to-any singing voice conversion (SVC) aims to transfer a target singer's timbre to other songs using a short voice sample. However many diffusion model based any-to-any SVC methods, which have achieved impressive results, usually suffered from low efficiency caused by a mass of inference steps. In this paper, we propose LCM-SVC, a latent consistency distillation (LCD) based latent diffusion model (LDM) to accelerate inference speed. We achieved one-step or few-step inference while maintaining the high performance by distilling a pre-trained LDM based SVC model, which had the advantages of timbre decoupling and sound quality. Experimental results show that our proposed method can significantly reduce the inference time and largely preserve the sound quality and timbre similarity comparing with other state-of-the-art SVC models. Audio samples are available at https: sounddemos.github.io lcm-svc.

【3】 The Whole Is Bigger Than the Sum of Its Parts: Modeling Individual Annotators to Capture Emotional Variability
标题: 整体大于其部分之和:建模个体注释者以捕捉情绪变异性
作者:James Tavernor,Yara El-Tawil,Emily Mower Provost
备注:Accepted to Interspeech 2024 Conference
链接:点击下载PDF文件
摘要:情感表达和感知是微妙的、复杂的、高度主观的过程。当多个注释器标记情感数据时,产生的标签包含高可变性。大多数语音情感识别任务通过将注释者标签平均为地面实况来解决这个问题。然而,这个过程忽略了情感的细微差别和注释者之间的差异,这是要捕捉的重要信号。以前的工作试图学习分布来捕捉情绪的变化,但这些方法也失去了关于个人注释者的信息。我们通过学习预测单个注释器来解决这些限制,并通过引入一种新的方法来从连续的模型输出中创建分布,从而允许在模型训练期间学习情感分布。我们表明,这种结合的方法可以导致情感分布比以前的工作中看到的更准确,在内部和跨语料库设置。摘要:Emotion expression and perception are nuanced, complex, and highly subjective processes. When multiple annotators label emotional data, the resulting labels contain high variability. Most speech emotion recognition tasks address this by averaging annotator labels as ground truth. However, this process omits the nuance of emotion and inter-annotator variability, which are important signals to capture. Previous work has attempted to learn distributions to capture emotion variability, but these methods also lose information about the individual annotators. We address these limitations by learning to predict individual annotators and by introducing a novel method to create distributions from continuous model outputs that permit the learning of emotion distributions during model training. We show that this combined approach can result in emotion distributions that are more accurate than those seen in prior work, in both within- and cross-corpus settings.

【4】 Parameter-Efficient Transfer Learning under Federated Learning for Automatic Speech Recognition
标题: 自动语音识别联邦学习下的参数高效迁移学习
作者:Xuan Kan,Yonghui Xiao,Tien-Ju Yang,Nanxin Chen,Rajiv Mathews
链接:点击下载PDF文件
摘要:这项工作探讨了在保护用户数据隐私的同时,在各种用户特定领域增强自动语音识别(ASR)模型性能的挑战。我们采用联邦学习和参数有效的域自适应方法来解决(1)用户特定场景的ASR模型的大量数据需求以及(2)联邦学习期间服务器和客户端之间的大量通信成本。我们证明,当配备适当的适配器时,联合调优下的ASR模型可以实现与集中调优模型相似的性能,从而为未来隐私保护的ASR服务提供了潜在的方向。此外,我们还研究了联邦学习环境下不同适配器和适配器合并策略的效率。摘要:This work explores the challenge of enhancing Automatic Speech Recognition (ASR) model performance across various user-specific domains while preserving user data privacy. We employ federated learning and parameter-efficient domain adaptation methods to solve the (1) massive data requirement of ASR models from user-specific scenarios and (2) the substantial communication cost between servers and clients during federated learning. We demonstrate that when equipped with proper adapters, ASR models under federated tuning can achieve similar performance compared with centralized tuning ones, thus providing a potential direction for future privacy-preserved ASR services. Besides, we investigate the efficiency of different adapters and adapter incorporation strategies under the federated learning setting.

【5】 Non-Causal to Causal SSL-Supported Transfer Learning: Towards a High-Performance Low-Latency Speech Vocode
标题: 无因到因SSL支持的迁移学习:迈向高性能低延迟语音声码
作者:Renzheng Shi,Andreas Bär,Marvin Sach,Wouter Tirry,Tim Fingscheidt
备注:Accepted at IWAENC 2024
链接:点击下载PDF文件
摘要:最近,BigVGAN已经成为高性能语音声码器。然而,其基于序列到序列的合成禁止在低延迟会话应用中使用。我们的工作通过三个步骤解决了这一缺陷。首先,我们通过实现因果卷积将低延迟引入BigVGAN,从而降低了性能。其次,为了恢复性能,我们提出了一个师生迁移学习方案,将高延迟的非因果BigVGAN提取到我们的低延迟因果声码器中。第三,利用自监督学习(SSL)模型,在我们的案例wav2vec 2.0中,我们将从我们的低延迟因果声码器中提取的编码器语音表示与地面真实语音表示对齐。在扬声器独立的设置,这两个建议的训练方案显着提高我们的低延迟声码器的性能,接近原来的高延迟BigVGAN。在复杂度仅提高21%的情况下,我们最好的小型因果声码器实现了3.96 PESQ和1.25 MCD,甚至分别优于原始的小型非因果BigVGAN(3.64 PESQ)0.32 PESQ和0.1 MCD点。摘要:Recently, BigVGAN has emerged as high-performance speech vocoder. Its sequence-to-sequence-based synthesis, however, prohibits usage in low-latency conversational applications. Our work addresses this shortcoming in three steps. First, we introduce low latency into BigVGAN via implementing causal convolutions, yielding decreased performance. Second, to regain performance, we propose a teacher-student transfer learning scheme to distill the high-delay non-causal BigVGAN into our low-latency causal vocoder. Third, taking advantage of a self-supervised learning (SSL) model, in our case wav2vec 2.0, we align its encoder speech representations extracted from our low-latency causal vocoder to the ground truth ones. In speaker-independent settings, both proposed training schemes notably elevate the performance of our low-latency vocoder, closing up to the original high-delay BigVGAN. At only 21% higher complexity, our best small causal vocoder achieves 3.96 PESQ and 1.25 MCD, excelling even the original small non-causal BigVGAN (3.64 PESQ) by 0.32 PESQ and 0.1 MCD points, respectively.

【6】 Modeling Time-Variant Responses of Optical Compressors with Selective State Space Models
标题: 用选择性状态空间模型建模光学压缩机时变响应
作者:Riccardo Simionato
链接:点击下载PDF文件
摘要:本文提出了一种使用具有选择性状态空间模型的深度神经网络对光学动态范围压缩器进行建模的方法。所提出的方法超越了以往的方法的基础上,通过采用一个选择性的状态空间块来编码输入音频的递归层。它采用了一种精细的技术,集成了智能线性调制和门控线性单元,以动态调整网络,根据外部参数调节压缩的攻击和释放阶段。所提出的架构非常适合低延迟和实时应用,在现场音频处理中至关重要。该方法已在TubeTech CL 1B和Teletronix LA-2A模拟光压缩机上得到验证,这两种压缩机具有不同的特性。使用定量指标和主观听力测试进行评估,比较所提出的方法与其他国家的最先进的模型。结果表明,我们的黑盒建模方法优于所有其他方法,在训练过程中实现了对可见和不可见设置的压缩过程的准确仿真。我们进一步展示了这种准确性与数据集中控制参数的采样密度之间的相关性,并将快速攻击和缓慢释放的设置确定为最具挑战性的模拟。摘要:This paper presents a method for modeling optical dynamic range compressors using deep neural networks with Selective State Space models. The proposed approach surpasses previous methods based on recurrent layers by employing a Selective State Space block to encode the input audio. It features a refined technique integrating Feature-wise Linear Modulation and Gated Linear Units to adjust the network dynamically, conditioning the compression's attack and release phases according to external parameters. The proposed architecture is well-suited for low-latency and real-time applications, crucial in live audio processing. The method has been validated on the analog optical compressors TubeTech CL 1B and Teletronix LA-2A, which possess distinct characteristics. Evaluation is performed using quantitative metrics and subjective listening tests, comparing the proposed method with other state-of-the-art models. Results show that our black-box modeling methods outperform all others, achieving accurate emulation of the compression process for both seen and unseen settings during training. We further show a correlation between this accuracy and the sampling density of the control parameters in the dataset and identify settings with fast attack and slow release as the most challenging to emulate.

【7】 WhisperMask: A Noise Suppressive Mask-Type Microphone for Whisper Speech
标题: WhisperMass:一款用于耳语的降噪口罩式麦克风
作者:Hirotaka Hiraki,Shusuke Kanazawa,Takahiro Miura,Manabu Yoshida,Masaaki Mochimaru,Jun Rekimoto
Journal-ref:Proceedings of the Augmented Humans International Conference 2024
链接:点击下载PDF文件
摘要:耳语是基于语音的交互中常见的隐私保护技术,但其有效性在嘈杂环境中受到限制。在传统的基于硬件和软件的降噪方法中,将低声语音与环境噪声和其他语音声音隔离仍然是一个挑战。因此,我们提出了WhisperMask,这是一种面具式麦克风,具有低灵敏度的大振膜,使佩戴者的声音明显高于背景噪音。我们使用三个关键指标来评估WhisperMask:信噪比,录制语音的质量和语音识别率。在所有指标中,WhisperMask始终优于传统的降噪麦克风和基于软件的解决方案。值得注意的是,与针式麦克风和耳塞相比,WhisperMask在80 dB背景噪声的环境中记录的耳语语音识别准确率高出30%。此外,虽然去噪器使这两个麦克风在30-60分贝噪音下的低声语音识别率降低了约20%,但WhisperMask即使在不去噪的情况下也保持了高性能,大幅超越了其他麦克风的性能。WhisperMask的设计将佩戴者的声音作为主要输入,并在不依赖信号处理的情况下有效抑制背景噪音。该设备允许在各种嘈杂的现实世界场景中进行可靠的语音交互,例如电话呼叫和语音命令,同时保护用户隐私。摘要:Whispering is a common privacy-preserving technique in voice-based interactions, but its effectiveness is limited in noisy environments. In conventional hardware- and software-based noise reduction approaches, isolating whispered speech from ambient noise and other speech sounds remains a challenge. We thus propose WhisperMask, a mask-type microphone featuring a large diaphragm with low sensitivity, making the wearer's voice significantly louder than the background noise. We evaluated WhisperMask using three key metrics: signal-to-noise ratio, quality of recorded voices, and speech recognition rate. Across all metrics, WhisperMask consistently outperformed traditional noise-suppressing microphones and software-based solutions. Notably, WhisperMask showed a 30% higher recognition accuracy for whispered speech recorded in an environment with 80 dB background noise compared with the pin microphone and earbuds. Furthermore, while a denoiser decreased the whispered speech recognition rate of these two microphones by approximately 20% at 30-60 dB noise, WhisperMask maintained a high performance even without denoising, surpassing the other microphones' performances by a significant margin.WhisperMask's design renders the wearer's voice as the dominant input and effectively suppresses background noise without relying on signal processing. This device allows for reliable voice interactions, such as phone calls and voice commands, in a wide range of noisy real-world scenarios while preserving user privacy.

【8】 Self-Learning for Personalized Keyword Spotting on Ultra-Low-Power Audio Sensors
标题: 在超低功耗音频传感器上进行个性化关键词定位的自学习
作者:Manuele Rusci,Francesco Paci,Marco Fariselli,Eric Flamand,Tinne Tuytelaars
链接:点击下载PDF文件
摘要:本文提出了一种自学习框架,以增量训练(微调)的个性化关键词定位(KWS)模型部署后,超低功耗智能音频传感器。我们解决的基本问题,标记的训练数据的情况下,通过分配伪标签的基础上的相似性分数相对于几个用户记录的新记录的音频帧。通过在两个公共数据集上使用多个参数高达0.5M的KWS模型进行实验,我们发现,与在大量通用关键字上预训练的初始模型相比,准确率提高了+19.2%和+16.0%。标签的任务是由一个低功耗的麦克风和一个节能的微控制器(MCU)的传感器系统。通过有效利用MCU的异构处理引擎,始终在线的标记任务实时运行,平均功耗高达8.2 mW。在同一平台上,我们估计设备上训练的能量成本比使用DS-CNN-S或DS-CNN-M模型每5秒或16.4秒对新话语进行采样的标记能量低10倍。我们的实证结果为在极端边缘实现自适应个性化KWS传感器铺平了道路。摘要:This paper proposes a self-learning framework to incrementally train (fine-tune) a personalized Keyword Spotting (KWS) model after the deployment on ultra-low power smart audio sensors. We address the fundamental problem of the absence of labeled training data by assigning pseudo-labels to the new recorded audio frames based on a similarity score with respect to few user recordings. By experimenting with multiple KWS models with a number of parameters up to 0.5M on two public datasets, we show an accuracy improvement of up to +19.2% and +16.0% vs. the initial models pretrained on a large set of generic keywords. The labeling task is demonstrated on a sensor system composed of a low-power microphone and an energy-efficient Microcontroller (MCU). By efficiently exploiting the heterogeneous processing engines of the MCU, the always-on labeling task runs in real-time with an average power cost of up to 8.2 mW. On the same platform, we estimate an energy cost for on-device training 10x lower than the labeling energy if sampling a new utterance every 5 s or 16.4 s with a DS-CNN-S or a DS-CNN-M model. Our empirical result paves the way to self-adaptive personalized KWS sensors at the extreme edge.

【9】 Developing vocal system impaired patient-aimed voice quality assessment approach using ASR representation-included multiple features
标题: 使用包含ASB表示的多个特征开发针对发声系统受损患者的语音质量评估方法
作者:Shaoxiang Dang,Tetsuya Matsumoto,Yoshinori Takeuchi,Takashi Tsuboi,Yasuhiro Tanaka,Daisuke Nakatsubo,Satoshi Maesawa,Ryuta Saito,Masahisa Katsuno,Hiroaki Kudo
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:深度学习在临床语音处理中的潜力是巨大的,但有限和不平衡的临床数据样本的障碍迫在眉睫。本文通过展示自动语音识别和自监督学习表示的利用来解决这些挑战,这些表示在广泛的正常语音数据集上进行了预训练。这种创新的方法旨在评估声音系统受损患者的声音质量。实验涉及对PVQD数据集的检查,涵盖英语中发音系统损伤的各种原因,以及日本数据集,重点关注接受丘脑底核脑深部电刺激(STN-DBS)手术前后的帕金森病患者。PVQD的结果揭示了在预测等级、呼吸和虚弱指标方面的显著相关性(PCC 0.8)和非凡的准确性(MSE <0.5)。同时,在STN-DBS的背景下,在预测患者的语音质量方面取得了进展。摘要:The potential of deep learning in clinical speech processing is immense, yet the hurdles of limited and imbalanced clinical data samples loom large. This article addresses these challenges by showcasing the utilization of automatic speech recognition and self-supervised learning representations, pre-trained on extensive datasets of normal speech. This innovative approach aims to estimate voice quality of patients with impaired vocal systems. Experiments involve checks on PVQD dataset, covering various causes of vocal system damage in English, and a Japanese dataset focusing on patients with Parkinson's disease before and after undergoing subthalamic nucleus deep brain stimulation (STN-DBS) surgery. The results on PVQD reveal a notable correlation ( 0.8 on PCC) and an extraordinary accuracy (<0.5 on MSE) in predicting Grade, Breathy, and Asthenic indicators. Meanwhile, progress has been achieved in predicting the voice quality of patients in the context of STN-DBS.

【10】 VoiceX: A Text-To-Speech Framework for Custom Voices
标题: SecureX:自定义语音的文本到语音框架
作者:Silvan Mertes,Daksitha Withanage Don,Otto Grothe,Johanna Kuch,Ruben Schlagowski,Elisabeth André
链接:点击下载PDF文件
摘要:现代TTS系统能够创建高度逼真和自然的语音。尽管有这些发展,定制TTS语音的过程仍然是一项复杂的任务,主要需要该领域专家的专业知识。其中一个原因是深度学习模型的使用,其特点是其扩展的,不可解释的参数空间,限制了手动定制的可行性。在本文中,我们提出了一种新的人在循环的基础上直接与神经TTS模型的参数空间的进化算法的范例。我们将我们的方法集成到一个用户友好的图形用户界面中,允许用户有效地创建原始声音。然后,这些语音可以与主干TTS模型一起使用,我们为此提供了Python API。此外,我们提出了用户研究的结果,探索VoiceX的功能。我们表明,VoiceX是一个合适的工具,用于创建个人,自定义的声音。摘要:Modern TTS systems are capable of creating highly realistic and natural-sounding speech. Despite these developments, the process of customizing TTS voices remains a complex task, mostly requiring the expertise of specialists within the field. One reason for this is the utilization of deep learning models, which are characterized by their expansive, non-interpretable parameter spaces, restricting the feasibility of manual customization. In this paper, we present a novel human-in-the-loop paradigm based on an evolutionary algorithm for directly interacting with the parameter space of a neural TTS model. We integrated our approach into a user-friendly graphical user interface that allows users to efficiently create original voices. Those voices can then be used with the backbone TTS model, for which we provide a Python API. Further, we present the results of a user study exploring the capabilities of VoiceX. We show that VoiceX is an appropriate tool for creating individual, custom voices.

【11】 Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization
标题: 集成音频、视觉和语义信息以实现增强的多模式说话人Dialogation
作者:Luyao Cheng,Hui Wang,Siqi Zheng,Yafeng Chen,Rongjie Huang,Qinglin Zhang,Qian Chen,Xihao Li
链接:点击下载PDF文件
摘要:说话人日志化是将音频流或转录的语音内容分割成基于说话人身份的同质分区的过程,在人类语音的解释和分析中起着至关重要的作用。大多数现有的扬声器日记系统完全依赖于单峰声学信息,由于音频信号的固有模糊性,使得任务特别具有挑战性。最近的研究已经取得了巨大的努力,视听或听觉语义建模,以提高性能。然而,即使将多达两种模式结合起来,也往往无法解决自发和非结构化对话的复杂性。为了利用更有意义的对话模式,我们提出了一种新的多模态方法,联合利用音频,视觉和语义线索,以提高扬声器日记。我们的方法优雅地制定了多峰建模作为一个约束优化问题。首先,我们建立洞察到积极发言者之间的视觉连接和口语内容内的语义交互,从而建立丰富的成对约束。然后,我们引入了一个联合成对约束传播算法的基础上,这些视觉和语义的约束进行说话人聚类。这种集成有效地利用了不同模态的互补优势,改进了单个说话人嵌入之间的亲和力估计。在多个多模态数据集上进行的大量实验表明,我们的方法始终优于最先进的说话人日记化方法。摘要:Speaker diarization, the process of segmenting an audio stream or transcribed speech content into homogenous partitions based on speaker identity, plays a crucial role in the interpretation and analysis of human speech. Most existing speaker diarization systems rely exclusively on unimodal acoustic information, making the task particularly challenging due to the innate ambiguities of audio signals. Recent studies have made tremendous efforts towards audio-visual or audio-semantic modeling to enhance performance. However, even the incorporation of up to two modalities often falls short in addressing the complexities of spontaneous and unstructured conversations. To exploit more meaningful dialogue patterns, we propose a novel multimodal approach that jointly utilizes audio, visual, and semantic cues to enhance speaker diarization. Our method elegantly formulates the multimodal modeling as a constrained optimization problem. First, we build insights into the visual connections among active speakers and the semantic interactions within spoken content, thereby establishing abundant pairwise constraints. Then we introduce a joint pairwise constraint propagation algorithm to cluster speakers based on these visual and semantic constraints. This integration effectively leverages the complementary strengths of different modalities, refining the affinity estimation between individual speaker embeddings. Extensive experiments conducted on multiple multimodal datasets demonstrate that our approach consistently outperforms state-of-the-art speaker diarization methods.

【12】 Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
标题: Video-Foley:通过Foley Sound的时间事件条件的两阶段视频到声音生成
作者:Junwon Lee,Jaekwon Im,Dabin Kim,Juhan Nam
链接:点击下载PDF文件
摘要:Foley声音合成对于多媒体制作至关重要,通过在时间和语义上同步音频和视频来增强用户体验。最近关于通过视频到声音生成来自动化这一劳动密集型过程的研究面临着重大挑战。缺乏显式时间特征的系统具有较差的可控性和对齐性,而基于时间戳的模型需要昂贵且主观的人工注释。我们提出了视频福利,视频到声音系统使用均方根(RMS)作为一个时间事件条件与语义音色提示(音频或文本)。RMS是一种与音频语义密切相关的帧级强度包络特征,可确保高可控性和同步性。无注释的自监督学习框架由两个阶段组成,Video 2 RMS和RMS 2Sound,结合了新的想法,包括RMS离散化和RMS-ControlNet以及预训练的文本到音频模型。我们广泛的评估表明,Video-Foley在视听对齐和声音定时,强度,音色和细微差别的可控性方面实现了最先进的性能。代码、模型重量和演示可在附带的网站上获得。(https: jnwnlee.github.io video-foley-demo)摘要:Foley sound synthesis is crucial for multimedia production, enhancing user experience by synchronizing audio and video both temporally and semantically. Recent studies on automating this labor-intensive process through video-to-sound generation face significant challenges. Systems lacking explicit temporal features suffer from poor controllability and alignment, while timestamp-based models require costly and subjective human annotation. We propose Video-Foley, a video-to-sound system using Root Mean Square (RMS) as a temporal event condition with semantic timbre prompts (audio or text). RMS, a frame-level intensity envelope feature closely related to audio semantics, ensures high controllability and synchronization. The annotation-free self-supervised learning framework consists of two stages, Video2RMS and RMS2Sound, incorporating novel ideas including RMS discretization and RMS-ControlNet with a pretrained text-to-audio model. Our extensive evaluation shows that Video-Foley achieves state-of-the-art performance in audio-visual alignment and controllability for sound timing, intensity, timbre, and nuance. Code, model weights, and demonstrations are available on the accompanying website. (https: jnwnlee.github.io video-foley-demo)

【13】 Prosody of speech production in latent post-stroke aphasia
标题: 中风后隐性失语症的言语产生韵律
作者:Cong Zhang,Tong Li,Gayle DeDe,Christos Salis
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:这项研究探讨了潜伏性失语症(一种与左半球脑损伤(例如中风)相关的轻度失语症)的韵律产生。与先前对中度至重度失语症的研究不同,我们研究了潜伏性失语症,这种失语症似乎与神经型言语具有非常相似的言语产生。我们分析了10名潜伏性失语症患者和10名正常对照者的首、尾词的f_0、强度和持续时间。回归模型拟合,以提高我们对这种未充分研究的非常轻度失语症的理解。结果突出了不同程度的差异,在所有三个韵律措施组之间。我们还研究了使用随机森林的潜伏性失语症与神经典型对照的诊断分类,旨在建立一个快速可靠的工具来帮助识别潜伏性失语症。随机森林分析也加强了韵律特征在区分潜在性失语中的意义。摘要:This study explores prosodic production in latent aphasia, a mild form of aphasia associated with left-hemisphere brain damage (e.g. stroke). Unlike prior research on moderate to severe aphasia, we investigated latent aphasia, which can seem to have very similar speech production with neurotypical speech. We analysed the f0, intensity and duration of utterance-initial and utterance-final words of ten speakers with latent aphasia and ten matching controls. Regression models were fitted to improve our understanding of this understudied type of very mild aphasia. The results highlighted varying degrees of differences in all three prosodic measures between groups. We also investigated the diagnostic classification of latent aphasia versus neurotypical control using random forest, aiming to build a fast and reliable tool to assist with the identification of latent aphasia. The random forest analysis also reinforced the significance of prosodic features in distinguishing latent aphasia.

【14】 Convexity-based Pruning of Speech Representation Models
标题: 基于凸度的语音表示模型修剪
作者:Teresa Dorszewski,Lenka Tětková,Lars Kai Hansen
链接:点击下载PDF文件
摘要:基于Transformer架构并通过自监督学习训练的语音表示模型在解决语音和说话人识别、关键字定位、情感检测等任务方面表现出了巨大的潜力。通常,发现较大的模型导致更好的性能。然而,在这样的大型Transformer系统中涉及的显著计算工作量对于嵌入式和现实世界的应用是一个挑战。最近的工作已经表明,在用于NLP的Transformer模型中存在显著的冗余,并且大规模层修剪是可行的(Sajjad等人,2023年)。在这里,我们研究音频模型中的层修剪。我们基于凸性准则的修剪决策。分类区域的凸性最近被提出作为一系列应用领域(包括NLP和音频)中后续微调性能的指标。在实证研究中,我们发现在计算工作量的大量减少,没有损失的性能,甚至在某些情况下的改进。摘要:Speech representation models based on the transformer architecture and trained by self-supervised learning have shown great promise for solving tasks such as speech and speaker recognition, keyword spotting, emotion detection, and more. Typically, it is found that larger models lead to better performance. However, the significant computational effort involved in such large transformer systems is a challenge for embedded and real-world applications. Recent work has shown that there is significant redundancy in the transformer models for NLP and massive layer pruning is feasible (Sajjad et al., 2023). Here, we investigate layer pruning in audio models. We base the pruning decision on a convexity criterion. Convexity of classification regions has recently been proposed as an indicator of subsequent fine-tuning performance in a range of application domains, including NLP and audio. In empirical investigations, we find a massive reduction in the computational effort with no loss of performance or even improvements in certain cases.

【15】 Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation
标题: Style-Talker:微调音频语言模型和基于风格的文本到语音模型,用于快速口语对话生成
作者:Yinghao Aaron Li,Xilin Jiang,Jordan Darefsky,Ge Zhu,Nima Mesgarani
备注:CoLM 2024
链接:点击下载PDF文件
摘要:大型语言模型(LLM)的快速发展极大地推动了基于文本的聊天机器人的发展,展示了它们参与连贯和上下文相关对话的能力。然而,扩展这些进步以实现端到端语音对语音对话机器人仍然是一个巨大的挑战,主要是由于所需的大量数据集和计算资源。在流水线中级联自动语音识别(ASR)、LLM和文本到语音(TTS)模型的常规方法虽然有效,但由于其缺乏输入音频及其转录文本与输出音频之间的直接交互而遭受不自然的韵律。这些系统还受到来自用于实时应用的ASR过程的固有延迟的限制。本文介绍了Style-Talker,一个创新的框架,微调音频LLM以及基于风格的TTS模型,用于快速口语对话生成。Style-Talker采用用户输入的音频,并使用转录的聊天历史和讲话风格来生成讲话风格和文本的响应。随后,TTS模型合成语音,然后将其回放给用户。在播放响应语音的同时,输入语音经历ASR处理以提取转录和说话风格,作为随后的对话回合的上下文。这种新的流水线加速了传统的级联ASR-LLM-TTS系统,同时集成了来自输入语音的丰富的语言信息。我们的实验结果表明,Style-Talker在对话自然度和连贯性方面显着优于传统的级联和语音到语音基线,同时速度快了50%以上。摘要:The rapid advancement of large language models (LLMs) has significantly propelled the development of text-based chatbots, demonstrating their capability to engage in coherent and contextually relevant dialogues. However, extending these advancements to enable end-to-end speech-to-speech conversation bots remains a formidable challenge, primarily due to the extensive dataset and computational resources required. The conventional approach of cascading automatic speech recognition (ASR), LLM, and text-to-speech (TTS) models in a pipeline, while effective, suffers from unnatural prosody because it lacks direct interactions between the input audio and its transcribed text and the output audio. These systems are also limited by their inherent latency from the ASR process for real-time applications. This paper introduces Style-Talker, an innovative framework that fine-tunes an audio LLM alongside a style-based TTS model for fast spoken dialog generation. Style-Talker takes user input audio and uses transcribed chat history and speech styles to generate both the speaking style and text for the response. Subsequently, the TTS model synthesizes the speech, which is then played back to the user. While the response speech is being played, the input speech undergoes ASR processing to extract the transcription and speaking style, serving as the context for the ensuing dialogue turn. This novel pipeline accelerates the traditional cascade ASR-LLM-TTS systems while integrating rich paralinguistic information from input speech. Our experimental results show that Style-Talker significantly outperforms the conventional cascade and speech-to-speech baselines in terms of both dialogue naturalness and coherence while being more than 50% faster.


机器翻译,仅供参考