今日论文合集:cs.SD语音12篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition

标题:SpeechComposer:将多个语音任务与提示合成统一起来

链接:https://arxiv.org/abs/2401.18045

作者:Yihan Wu,Soumi Maiti,Yifan Peng,Wangyou Zhang,Chenda Li,Yuyue Wang,Xihua Wang,Shinji Watanabe,Ruihua Song

备注:11 pages, 2 figures

摘要:语言模型的最新进展显着提高了多个语音相关任务的性能。现有的语音语言模型通常利用任务相关的提示令牌来统一在单个模型中的各种语音任务。然而,这种设计忽略了不同语音任务之间的内在联系,这可能会提高每个任务的性能。在这项工作中,我们提出了一种新的解码器的语音语言模型,SpeechComposer,可以通过组成一组固定的提示符来统一常见的语音任务。基于四个主要任务-语音合成,语音识别,语音语言建模和文本语言建模- SpeechComposer可以通过精心设计的提示标记的组合轻松扩展到更多的语音任务,如语音转换和语音增强。提示标记的统一也使得不同语音任务之间以更结构化的方式进行知识共享成为可能。实验结果表明,我们提出的SpeechComposer可以提高主任务和复合任务的性能,显示了共享提示令牌的有效性。值得注意的是,统一的解码器模型实现了可比的,甚至比基线更好的性能,基线是专为单个任务设计的专家模型。

摘要:Recent advancements in language models have significantly enhanced performance in multiple speech-related tasks. Existing speech language models typically utilize task-dependent prompt tokens to unify various speech tasks in a single model. However, this design omits the intrinsic connections between different speech tasks, which can potentially boost the performance of each task. In this work, we propose a novel decoder-only speech language model, SpeechComposer, that can unify common speech tasks by composing a fixed set of prompt tokens. Built upon four primary tasks -- speech synthesis, speech recognition, speech language modeling, and text language modeling -- SpeechComposer can easily extend to more speech tasks via compositions of well-designed prompt tokens, like voice conversion and speech enhancement. The unification of prompt tokens also makes it possible for knowledge sharing among different speech tasks in a more structured manner. Experimental results demonstrate that our proposed SpeechComposer can improve the performance of both primary tasks and composite tasks, showing the effectiveness of the shared prompt tokens. Remarkably, the unified decoder-only model achieves a comparable and even better performance than the baselines which are expert models designed for single tasks.


【2】 Dance-to-Music Generation with Encoder-based Textual Inversion of  Diffusion Models
标题:基于编码的扩散模型文本倒置的音乐舞曲生成
链接:https://arxiv.org/abs/2401.17800
作者:Sifei Li,Weiming Dong,Yuxin Zhang,Fan Tang,Chongyang Ma,Oliver Deussen,Tong-Yee Lee,Changsheng Xu
备注:9 pages, 3 figures
摘要:音乐与舞蹈动作的和谐统一是生动传达舞蹈艺术精髓的关键。这种一致性也显著提升了游戏体验和动画制作的沉浸式质量。虽然在从文本描述创建高保真音乐方面已经取得了显着的进步,但当前的方法主要集中在调节诸如流派和情感基调之类的总体特征上。他们经常忽视时间节奏的细微管理,这是为舞蹈制作音乐所不可或缺的,因为它将音乐节拍与舞者的动作紧密结合在一起。认识到这一差距,我们提出了一种基于编码器的文本反转技术,用于增强文本到音乐模型的视觉控制,促进个性化的音乐生成。具体来说,我们开发了双路径节奏体裁反转,以有效地将舞蹈动作序列的节奏和体裁整合到文本到音乐模型的文本空间中。与经典的文本反转方法相反,该方法直接更新文本嵌入以重建单个目标对象,我们的方法利用单独的节奏和体裁编码器来获得两个伪词的文本嵌入,以适应不同的节奏和体裁。为了实现更准确的评估,我们提出了改进的评估指标的节奏对齐。我们证明了我们的方法在多个评估指标上优于最先进的方法。此外,我们的方法无缝地适应野外数据,并有效地与预训练模型的固有文本指导生成能力相结合。示例可在\url{https://youtu.be/D7XDwtH1YwE}获得。
摘要:The harmonious integration of music with dance movements is pivotal in vividly conveying the artistic essence of dance. This alignment also significantly elevates the immersive quality of gaming experiences and animation productions. While there has been remarkable advancement in creating high-fidelity music from textual descriptions, current methodologies mainly concentrate on modulating overarching characteristics such as genre and emotional tone. They often overlook the nuanced management of temporal rhythm, which is indispensable in crafting music for dance, since it intricately aligns the musical beats with the dancers' movements. Recognizing this gap, we propose an encoder-based textual inversion technique for augmenting text-to-music models with visual control, facilitating personalized music generation. Specifically, we develop dual-path rhythm-genre inversion to effectively integrate the rhythm and genre of a dance motion sequence into the textual space of a text-to-music model. Contrary to the classical textual inversion method, which directly updates text embeddings to reconstruct a single target object, our approach utilizes separate rhythm and genre encoders to obtain text embeddings for two pseudo-words, adapting to the varying rhythms and genres. To achieve a more accurate evaluation, we propose improved evaluation metrics for rhythm alignment. We demonstrate that our approach outperforms state-of-the-art methods across multiple evaluation metrics. Furthermore, our method seamlessly adapts to in-the-wild data and effectively integrates with the inherent text-guided generation capability of the pre-trained model. Samples are available at \url{https://youtu.be/D7XDwtH1YwE}.

【3】 Exploiting Audio-Visual Features with Pretrained AV-HuBERT for  Multi-Modal Dysarthric Speech Reconstruction
标题:利用预先训练的AV-Hubert的视听特征进行多模节律语音重建
链接:https://arxiv.org/abs/2401.17796
作者:Xueyuan Chen,Yuejiao Wang,Xixin Wu,Disong Wang,Zhiyong Wu,Xunying Liu,Helen Meng
备注:Accepted by ICASSP 2024
摘要:构音障碍语音重建(DSR)旨在通过提高构音障碍语音的可懂度和自然度,将构音障碍语音转换为正常语音。这是一项具有挑战性的任务,特别是对于患有严重构音障碍并在复杂嘈杂的声学环境中说话的患者。为了应对这些挑战,我们提出了一种新的多模态框架来利用视觉信息,例如,嘴唇的运动,在DSR重建高度异常的发音作为额外的线索。多模态框架包括:(i)多模态编码器,用于利用辅助视觉特征从构音障碍语音中提取鲁棒的音素嵌入;(ii)方差适配器,用于从提取的音素嵌入中推断正常的音素持续时间和音高轮廓;(iii)说话者编码器,用于编码说话者的语音特征;以及(iv)梅尔解码器,用于基于所提取的音素嵌入、韵律特征和说话者嵌入来生成重构的梅尔声谱图。在常用的UASpeech语料库上进行的客观和主观评估表明,我们提出的方法可以在语音可懂度和自然度方面实现比基线系统的显着改善,特别是对于具有更严重症状的扬声器。与原始构音障碍语音相比,对于构音障碍程度更严重的患者,重建语音的绝对错误率降低了42.1%。
摘要:Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in complex, noisy acoustic environments. To address these challenges, we propose a novel multi-modal framework to utilize visual information, e.g., lip movements, in DSR as extra clues for reconstructing the highly abnormal pronunciations. The multi-modal framework consists of: (i) a multi-modal encoder to extract robust phoneme embeddings from dysarthric speech with auxiliary visual features; (ii) a variance adaptor to infer the normal phoneme duration and pitch contour from the extracted phoneme embeddings; (iii) a speaker encoder to encode the speaker's voice characteristics; and (iv) a mel-decoder to generate the reconstructed mel-spectrogram based on the extracted phoneme embeddings, prosodic features and speaker embeddings. Both objective and subjective evaluations conducted on the commonly used UASpeech corpus show that our proposed approach can achieve significant improvements over baseline systems in terms of speech intelligibility and naturalness, especially for the speakers with more severe symptoms. Compared with original dysarthric speech, the reconstructed speech achieves 42.1\% absolute word error rate reduction for patients with more severe dysarthria levels.


【4】 Harnessing Smartwatch Microphone Sensors for Cough Detection and  Classification
标题:利用智能手表麦克风传感器进行咳嗽检测和分类
链接:https://arxiv.org/abs/2401.17738
作者:Pranay Jaiswal,Haroon R. Lone
备注:7 pages
摘要:这项研究调查了使用内置麦克风传感器的智能手表来监测咳嗽和检测各种咳嗽类型的潜力。我们进行了一项涉及32名参与者的研究,并以受控方式收集了9小时的音频数据。之后,我们使用结构化方法处理这些数据,得到223个阳性咳嗽样本。我们通过增强技术进一步改进了数据集,并采用了专门的1D CNN模型。该模型在非行走时的准确率为98.49%,在行走时为98.2%,这表明智能手表可以检测到咳嗽。此外,我们的研究成功地确定了四种不同类型的咳嗽使用聚类技术。
摘要:This study investigates the potential of using smartwatches with built-in microphone sensors for monitoring coughs and detecting various cough types. We conducted a study involving 32 participants and collected 9 hours of audio data in a controlled manner. Afterward, we processed this data using a structured approach, resulting in 223 positive cough samples. We further improved the dataset through augmentation techniques and employed a specialized 1D CNN model. This model achieved an impressive accuracy rate of 98.49% while non-walking and 98.2% while walking, showing smartwatches can detect cough. Moreover, our research successfully identified four distinct types of coughs using clustering techniques.


【5】 What Do Self-Supervised Speech and Speaker Models Learn? New Findings  From a Cross Model Layer-Wise Analysis
标题:自我监督的语音和说话人模型学到了什么?跨模型分层分析的新发现
链接:https://arxiv.org/abs/2401.17632
作者:Takanori Ashihara,Marc Delcroix,Takafumi Moriya,Kohei Matsuura,Taichi Asami,Yusuke Ijima备注:Accepted at ICASSP 2024
摘要:自监督学习(SSL)在学习有意义的语音表示方面引起了越来越多的关注。语音SSL模型,如WavLM,采用掩码预测训练来编码通用表示。相比之下,以DINO为基础的说话人SSL模型主要针对说话人表示采用话语级训练目标。了解这些模型如何表示信息对于改进模型的效率和有效性至关重要。与语音SSL的各种分析不同,对说话人SSL捕获的信息以及其表示与语音SSL或其他完全监督的说话人模型的不同之处进行了有限的调查。本文探讨了这些基本问题。我们探索的能力,以捕捉各种语音属性,通过应用SUPERB评估探测任务的语音和扬声器SSL模型。我们还研究了哪些层主要用于每个任务,以确定语音表示方式的差异。此外,我们进行直接比较,以测量模型内部和模型之间的层之间的相似性。我们的分析揭示了1)表示内容信息的能力与增强的说话人表示无关,2)语音SSL模型的特定层将部分专门用于捕获语言信息,3)说话人SSL模型倾向于忽略语言信息,但表现出更复杂的说话人表示。
摘要:Self-supervised learning (SSL) has attracted increased attention for learning meaningful speech representations. Speech SSL models, such as WavLM, employ masked prediction training to encode general-purpose representations. In contrast, speaker SSL models, exemplified by DINO-based models, adopt utterance-level training objectives primarily for speaker representation. Understanding how these models represent information is essential for refining model efficiency and effectiveness. Unlike the various analyses of speech SSL, there has been limited investigation into what information speaker SSL captures and how its representation differs from speech SSL or other fully-supervised speaker models. This paper addresses these fundamental questions. We explore the capacity to capture various speech properties by applying SUPERB evaluation probing tasks to speech and speaker SSL models. We also examine which layers are predominantly utilized for each task to identify differences in how speech is represented. Furthermore, we conduct direct comparisons to measure the similarities between layers within and across models. Our analysis unveils that 1) the capacity to represent content information is somewhat unrelated to enhanced speaker representation, 2) specific layers of speech SSL models would be partly specialized in capturing linguistic information, and 3) speaker SSL models tend to disregard linguistic information but exhibit more sophisticated speaker representation.


【6】 Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and  KiSing-v2
标题:歌唱语音数据放大:ACE-Opencop和KiSing-v2简介
链接:https://arxiv.org/abs/2401.17619
作者:Jiatong Shi,Yueqian Lin,Xinyi Bai,Keyi Zhang,Yuning Wu,Yuxun Tang,Yifeng Yu,Qin Jin,Shinji Watanabe
摘要:在歌声合成(SVS)中,从乐谱生成歌声面临着有限的数据可用性的挑战,这在文本到语音(TTS)中不太常见。这项研究提出了一种新的方法来解决这种数据稀缺性。我们利用现有的歌唱声音合成器进行数据增强,并应用精确的手动调谐,以减少不自然的声音合成。我们开发了两个广泛的歌唱声音语料库,ACE-Opencpop和KiSing-v2,便于大规模,多歌手的声音合成。利用来自这些语料库的预训练模型,我们在语音质量方面取得了显着的改善,在域内和域外场景中都很明显。语料库、预训练模型及其相关训练方法可在Muskits-ESPnet(https://github.com/espnet/espnet)上公开获取。
摘要:In singing voice synthesis (SVS), generating singing voices from musical scores faces challenges due to limited data availability, a constraint less common in text-to-speech (TTS). This study proposes a new approach to address this data scarcity. We utilize an existing singing voice synthesizer for data augmentation and apply precise manual tuning to reduce unnatural voice synthesis. Our development of two extensive singing voice corpora, ACE-Opencpop and KiSing-v2, facilitates large-scale, multi-singer voice synthesis. Utilizing pre-trained models derived from these corpora, we achieve notable improvements in voice quality, evident in both in-domain and out-of-domain scenarios. The corpora, pre-trained models, and their related training recipes are publicly available at Muskits-ESPnet (https://github.com/espnet/espnet).


【7】 Computation and Parameter Efficient Multi-Modal Fusion Transformer for  Cued Speech Recognition
标题:用于线索语音识别的计算和参数高效的多模式融合转换器
链接:https://arxiv.org/abs/2401.17604
作者:Lei Liu,Li Liu,Haizhou Li
备注:Accepted by TASLP
摘要:提示语音(CS)是一种纯视觉编码方法,由听力受损的人使用,结合唇读与几个特定的手形,使口语可见。自动CS识别(ACSR)旨在将语音的视觉线索转换为文本,这可以帮助听障人士有效地进行交流。CS的视觉信息包括唇读和手暗示,因此唇读和手暗示的融合在ACSR中起着重要的作用。然而,大多数以前的融合方法的斗争,以捕捉全球的依赖性,目前在长序列输入的多模态CS数据。因此,这些方法通常无法学习有助于融合的有效跨模态关系。最近,基于注意力的Transformers已经成为一个流行的想法,在多模态融合中捕获的全局依赖性的长序列,但现有的多模态融合变换器的识别精度差和效率低下的ACSR任务的计算。为了解决这些问题,我们开发了一种新的计算和参数有效的多模态融合Transformer,提出了一种新的令牌重要性感知注意机制(TIAA),其中令牌利用率(TUR)制定选择重要的令牌从多模态流。更准确地说,TIAA首先模型的模态特定的细粒度的时间依赖性的所有令牌的每一个模态,然后学习有效的跨模态交互的模态共享的粗粒度的时间依赖性的重要令牌的不同模态。此外,还设计了一个轻量级的门控隐藏投影来控制TIAA的特征流。与现有的基于变换的融合方法和ACSR融合方法相比,所得到的模型,命名为经济提示语音融合Transformer(EcoCued),实现了最先进的性能在所有现有的CS数据集。
摘要:Cued Speech (CS) is a pure visual coding method used by hearing-impaired people that combines lip reading with several specific hand shapes to make the spoken language visible. Automatic CS recognition (ACSR) seeks to transcribe visual cues of speech into text, which can help hearing-impaired people to communicate effectively. The visual information of CS contains lip reading and hand cueing, thus the fusion of them plays an important role in ACSR. However, most previous fusion methods struggle to capture the global dependency present in long sequence inputs of multi-modal CS data. As a result, these methods generally fail to learn the effective cross-modal relationships that contribute to the fusion. Recently, attention-based transformers have been a prevalent idea for capturing the global dependency over the long sequence in multi-modal fusion, but existing multi-modal fusion transformers suffer from both poor recognition accuracy and inefficient computation for the ACSR task. To address these problems, we develop a novel computation and parameter efficient multi-modal fusion transformer by proposing a novel Token-Importance-Aware Attention mechanism (TIAA), where a token utilization rate (TUR) is formulated to select the important tokens from the multi-modal streams. More precisely, TIAA firstly models the modality-specific fine-grained temporal dependencies over all tokens of each modality, and then learns the efficient cross-modal interaction for the modality-shared coarse-grained temporal dependencies over the important tokens of different modalities. Besides, a light-weight gated hidden projection is designed to control the feature flows of TIAA. The resulting model, named Economical Cued Speech Fusion Transformer (EcoCued), achieves state-of-the-art performance on all existing CS datasets, compared with existing transformer-based fusion methods and ACSR fusion methods.


【8】 Revisiting speech segmentation and lexicon learning with better features
标题:用更好的特征重温语音切分和词汇学习
链接:https://arxiv.org/abs/2401.17902
作者:Herman Kamper,Benjamin van Niekerk
备注:2 pages
摘要:我们重新审视了一种自我监督的方法,将未标记的语音分割成类似单词的片段。我们从两阶段的持续时间惩罚的动态规划方法,执行零资源分割,而无需学习一个明确的词典。在第一声学单元发现阶段,我们用HuBERT替换对比预测编码特征。在第二阶段的分词之后,我们通过平均HuBERT特征来获得每个片段的声学词嵌入。这些嵌入使用K-means进行聚类以获得词典。其结果是良好的全覆盖分割与词典,实现了最先进的性能上的ZeroSpeech基准。
摘要:We revisit a self-supervised method that segments unlabelled speech into word-like segments. We start from the two-stage duration-penalised dynamic programming method that performs zero-resource segmentation without learning an explicit lexicon. In the first acoustic unit discovery stage, we replace contrastive predictive coding features with HuBERT. After word segmentation in the second stage, we get an acoustic word embedding for each segment by averaging HuBERT features. These embeddings are clustered using K-means to get a lexicon. The result is good full-coverage segmentation with a lexicon that achieves state-of-the-art performance on the ZeroSpeech benchmarks.

【9】 EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for  Automated Audio Captioning
标题:EnCLAP:结合神经音频编解码器和音文联合嵌入的自动音频字幕
链接:https://arxiv.org/abs/2401.17690
作者:Jaeyeon Kim,Jaeyoon Jung,Jinjoo Lee,Sang Hoon Woo
备注:Accepted to ICASSP 2024
摘要:我们提出了EnCLAP,一个新的框架自动音频字幕。EnCLAP采用两个声学表示模型,EnCodec和CLAP,以及预训练的语言模型,BART。我们还引入了一个新的训练目标,称为掩蔽编解码器建模,提高了预训练语言模型的声学感知。AudioCaps和Clotho上的实验结果表明,我们的模型超过了基线模型的性能。源代码将在https://github.com/jaeyeonkim99/EnCLAP上提供。在线演示可在https://huggingface.co/spaces/enclap-team/enclap上获得。
摘要:We propose EnCLAP, a novel framework for automated audio captioning. EnCLAP employs two acoustic representation models, EnCodec and CLAP, along with a pretrained language model, BART. We also introduce a new training objective called masked codec modeling that improves acoustic awareness of the pretrained language model. Experimental results on AudioCaps and Clotho demonstrate that our model surpasses the performance of baseline models. Source code will be available at https://github.com/jaeyeonkim99/EnCLAP . An online demo is available at https://huggingface.co/spaces/enclap-team/enclap .

【10】 Detecting gamma-band responses to the speech envelope for the ICASSP  2024 Auditory EEG Decoding Signal Processing Grand Challenge
标题:检测ICASSP 2024听觉EEG解码信号处理挑战赛的语音包络的伽马波段响应
链接:https://arxiv.org/abs/2401.17380
作者:Mike Thornton,Jonas Auernheimer,Constantin Jehn,Danilo Mandic,Tobias Reichenbach
备注:Accepted for ICASSP 2024 (challenge track)
摘要:2024年ICASSP听觉EEG信号处理大挑战赛涉及从聆听演讲材料的参与者中获取的脑电图(EEG)测量值的解码。这项工作详细介绍了我们的解决方案的匹配不匹配的子任务:给定一个短的时间段的EEG记录和几个候选的语音段,任务是分类的语音段的时间对齐的EEG信号。我们表明,可以检测到高精度的语音包络的高频伽马波段响应。通过联合评估伽马波段响应和低频包络跟踪,我们开发了一个匹配失配解码器,其中放置在这项任务的第一。
摘要:The 2024 ICASSP Auditory EEG Signal Processing Grand Challenge concerns the decoding of electroencephalography (EEG) measurements taken from participants who listened to speech material. This work details our solution to the match-mismatch sub-task: given a short temporal segment of EEG recordings and several candidate speech segments, the task is to classify which of the speech segments was time-aligned with the EEG signals. We show that high-frequency gamma-band responses to the speech envelope can be detected with a high accuracy. By jointly assessing gamma-band responses and low-frequency envelope tracking, we develop a match-mismatch decoder which placed first in this task.


【11】 Sigma-lognormal modeling of speech
标题:语音的Sigma-对数正态模型
链接:https://arxiv.org/abs/2401.17320
作者:C. Carmona-Duarte,M. A. Ferrer,R. Plamondon,A. Gomez-Rodellar,P. Gomez-Vilda
备注:None
摘要:人体运动研究和分析在许多科学领域都是基础性的,从神经科学到教育,从模式识别到机器人技术,从医疗保健到体育等等。先前的言语运动模型被提出来理解言语运动是如何产生的,以及当某些参数改变时,所产生的言语是如何变化的。然而,逆的方法,其中的肌肉反应参数和主体的年龄是从真正的连续语音,是不可能的,与这样的模型。相反,在手写领域,快速人体运动的运动学理论及其相关的Sigma-lognormal模型已成功地应用于获得肌肉响应参数。这项工作提出了一个语音运动学为基础的模型,可以用来研究,分析和重建复杂的语音运动学在一个简化的方式。一种方法的基础上的快速人体运动的运动学理论及其相关的Sigma对数正态模型被施加到描述和参数化的渐近脉冲响应的神经肌肉网络参与语音作为响应神经运动命令。还介绍了用于进行从共振峰到运动观测的转换的方法。与(英语)VTR TIMIT数据库和(德国)Saarbrucken语音数据库,包括不同年龄的人,喉病变和不喉病变进行的实验,证实了提取的参数和老化之间的联系,一方面,和应用快速人体运动的运动学理论所需的第一和第二共振峰之间的比例,另一方面。研究结果将推动语音运动学建模和理解的创新发展。
摘要:Human movement studies and analyses have been fundamental in many scientific domains, ranging from neuroscience to education, pattern recognition to robotics, health care to sports, and beyond. Previous speech motor models were proposed to understand how speech movement is produced and how the resulting speech varies when some parameters are changed. However, the inverse approach, in which the muscular response parameters and the subject's age are derived from real continuous speech, is not possible with such models. Instead, in the handwriting field, the kinematic theory of rapid human movements and its associated Sigma-lognormal model have been applied successfully to obtain the muscular response parameters. This work presents a speech kinematics based model that can be used to study, analyze, and reconstruct complex speech kinematics in a simplified manner. A method based on the kinematic theory of rapid human movements and its associated Sigma lognormal model are applied to describe and to parameterize the asymptotic impulse response of the neuromuscular networks involved in speech as a response to a neuromotor command. The method used to carry out transformations from formants to a movement observation is also presented. Experiments carried out with the (English) VTR TIMIT database and the (German) Saarbrucken Voice Database, including people of different ages, with and without laryngeal pathologies, corroborate the link between the extracted parameters and aging, on the one hand, and the proportion between the first and second formants required in applying the kinematic theory of rapid human movements, on the other. The results should drive innovative developments in the modeling and understanding of speech kinematics.

【12】 Detection of Auditory Brainstem Response Peaks Using Image Processing  Techniques in Infants with Normal Hearing Sensitivity
标题:用图像处理技术检测听力正常婴儿的听性脑干反应峰值
链接:https://arxiv.org/abs/2401.17317
作者:Amir Majidpour,Samer Kais Jameel,Jafar Majidpour,Houra Bagheri,Tarik A. Rashid,Ahmadreza Nazeri,Mahshid Moheb Aleaba
摘要:简介:通过对听力正常儿童进行听性脑干反应(ABR)检测,以了解脑干水平周围听神经系统的完整性。听觉诱发电位(AEP)是利用声刺激产生的。解释这些波需要能力,以避免误诊听力问题。使用计算机视觉自动化ABR测试标记可以减少人为错误。方法:对26例1 ~ 20个月左右双耳听力正常儿童进行ABR测试。提出了一种新的方法,用于自动计算不同强度(分贝)的波的峰值。该程序需要使用Color Waveholder方法从Audera设备获取波形图像,使用Image Region Analyzer应用程序将每个波形分割为单个波形图像,使用图像处理(IP)技术将所有波形图像转换为波形,最后计算每个波形的峰值延迟,以供听力学家用于诊断疾病。调查结果:图像处理技术能够检测诊断视野中的1、3和5个波,准确度分别为0.82、0.98和0.98,并且其对波1、3和5的精确度分别为0.32、0.97和0.87。该方法在阈值部分也有良好的效果,ABR波的正确检出率为82.7%.结论:我们的研究结果表明,通过使用自动检测和标记ABR波的技术,听力学测试组合套件可以更加准确,快速和无错误。
摘要:Introduction: The auditory brainstem response (ABR) is measured to find the brainstem-level peripheral auditory nerve system integrity in children having normal hearing. The Auditory Evoked Potential (AEP) is generated using acoustic stimuli. Interpreting these waves requires competence to avoid misdiagnosing hearing problems. Automating ABR test labeling with computer vision may reduce human error. Method: The ABR test results of 26 children aged 1 to 20 months with normal hearing in both ears were used. A new approach is suggested for automatically calculating the peaks of waves of different intensities (in decibels). The procedure entails acquiring wave images from an Audera device using the Color Thresholder method, segmenting each wave as a single wave image using the Image Region Analyzer application, converting all wave images into waves using Image Processing (IP) techniques, and finally calculating the latency of the peaks for each wave to be used by an audiologist for diagnosing the disease. Findings: Image processing techniques were able to detect 1, 3, and 5 waves in the diagnosis field with accuracy (0.82), (0.98), and (0.98), respectively, and its precision for waves 1, 3, and 5, were respectively (0.32), (0.97) and (0.87). This evaluation also worked well in the thresholding part and 82.7 % correctly detected the ABR waves. Conclusion: Our findings indicate that the audiology test battery suite can be made more accurate, quick, and error-free by using technology to automatically detect and label ABR waves.


eess.AS音频处理
【1】 Revisiting speech segmentation and lexicon learning with better features
标题:用更好的特征重温语音切分和词汇学习
链接:https://arxiv.org/abs/2401.17902
作者:Herman Kamper,Benjamin van Niekerk
备注:2 pages
摘要:我们重新审视了一种自我监督的方法,将未标记的语音分割成类似单词的片段。我们从两阶段的持续时间惩罚的动态规划方法,执行零资源分割,而无需学习一个明确的词典。在第一声学单元发现阶段,我们用HuBERT替换对比预测编码特征。在第二阶段的分词之后,我们通过平均HuBERT特征来获得每个片段的声学词嵌入。这些嵌入使用K-means进行聚类以获得词典。其结果是良好的全覆盖分割与词典,实现了最先进的性能上的ZeroSpeech基准。
摘要:We revisit a self-supervised method that segments unlabelled speech into word-like segments. We start from the two-stage duration-penalised dynamic programming method that performs zero-resource segmentation without learning an explicit lexicon. In the first acoustic unit discovery stage, we replace contrastive predictive coding features with HuBERT. After word segmentation in the second stage, we get an acoustic word embedding for each segment by averaging HuBERT features. These embeddings are clustered using K-means to get a lexicon. The result is good full-coverage segmentation with a lexicon that achieves state-of-the-art performance on the ZeroSpeech benchmarks.

【2】 EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for  Automated Audio Captioning
标题:EnCLAP:结合神经音频编解码器和音文联合嵌入的自动音频字幕
链接:https://arxiv.org/abs/2401.17690
作者:Jaeyeon Kim,Jaeyoon Jung,Jinjoo Lee,Sang Hoon Woo
备注:Accepted to ICASSP 2024
摘要:我们提出了EnCLAP,一个新的框架自动音频字幕。EnCLAP采用两个声学表示模型,EnCodec和CLAP,以及预训练的语言模型,BART。我们还引入了一个新的训练目标,称为掩蔽编解码器建模,提高了预训练语言模型的声学感知。AudioCaps和Clotho上的实验结果表明,我们的模型超过了基线模型的性能。源代码将在https://github.com/jaeyeonkim99/EnCLAP上提供。在线演示可在https://huggingface.co/spaces/enclap-team/enclap上获得。
摘要:We propose EnCLAP, a novel framework for automated audio captioning. EnCLAP employs two acoustic representation models, EnCodec and CLAP, along with a pretrained language model, BART. We also introduce a new training objective called masked codec modeling that improves acoustic awareness of the pretrained language model. Experimental results on AudioCaps and Clotho demonstrate that our model surpasses the performance of baseline models. Source code will be available at https://github.com/jaeyeonkim99/EnCLAP . An online demo is available at https://huggingface.co/spaces/enclap-team/enclap .


【3】 Detecting gamma-band responses to the speech envelope for the ICASSP  2024 Auditory EEG Decoding Signal Processing Grand Challenge
标题:检测ICASSP 2024听觉EEG解码信号处理挑战赛的语音包络的伽马波段响应
链接:https://arxiv.org/abs/2401.17380
作者:Mike Thornton,Jonas Auernheimer,Constantin Jehn,Danilo Mandic,Tobias Reichenbach
备注:Accepted for ICASSP 2024 (challenge track)
摘要:2024年ICASSP听觉EEG信号处理大挑战赛涉及从聆听演讲材料的参与者中获取的脑电图(EEG)测量值的解码。这项工作详细介绍了我们的解决方案的匹配不匹配的子任务:给定一个短的时间段的EEG记录和几个候选的语音段,任务是分类的语音段的时间对齐的EEG信号。我们表明,可以检测到高精度的语音包络的高频伽马波段响应。通过联合评估伽马波段响应和低频包络跟踪,我们开发了一个匹配失配解码器,其中放置在这项任务的第一。
摘要:The 2024 ICASSP Auditory EEG Signal Processing Grand Challenge concerns the decoding of electroencephalography (EEG) measurements taken from participants who listened to speech material. This work details our solution to the match-mismatch sub-task: given a short temporal segment of EEG recordings and several candidate speech segments, the task is to classify which of the speech segments was time-aligned with the EEG signals. We show that high-frequency gamma-band responses to the speech envelope can be detected with a high accuracy. By jointly assessing gamma-band responses and low-frequency envelope tracking, we develop a match-mismatch decoder which placed first in this task.


【4】 SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition
标题:SpeechComposer:将多个语音任务与提示合成统一起来
链接:https://arxiv.org/abs/2401.18045
作者:Yihan Wu,Soumi Maiti,Yifan Peng,Wangyou Zhang,Chenda Li,Yuyue Wang,Xihua Wang,Shinji Watanabe,Ruihua Song
备注:11 pages, 2 figures
摘要:语言模型的最新进展显着提高了多个语音相关任务的性能。现有的语音语言模型通常利用任务相关的提示令牌来统一在单个模型中的各种语音任务。然而,这种设计忽略了不同语音任务之间的内在联系,这可能会提高每个任务的性能。在这项工作中,我们提出了一种新的解码器的语音语言模型,SpeechComposer,可以通过组成一组固定的提示符来统一常见的语音任务。基于四个主要任务-语音合成,语音识别,语音语言建模和文本语言建模- SpeechComposer可以通过精心设计的提示标记的组合轻松扩展到更多的语音任务,如语音转换和语音增强。提示标记的统一也使得不同语音任务之间以更结构化的方式进行知识共享成为可能。实验结果表明,我们提出的SpeechComposer可以提高主任务和复合任务的性能,显示了共享提示令牌的有效性。值得注意的是,统一的解码器模型实现了可比的,甚至比基线更好的性能,基线是专为单个任务设计的专家模型。
摘要:Recent advancements in language models have significantly enhanced performance in multiple speech-related tasks. Existing speech language models typically utilize task-dependent prompt tokens to unify various speech tasks in a single model. However, this design omits the intrinsic connections between different speech tasks, which can potentially boost the performance of each task. In this work, we propose a novel decoder-only speech language model, SpeechComposer, that can unify common speech tasks by composing a fixed set of prompt tokens. Built upon four primary tasks -- speech synthesis, speech recognition, speech language modeling, and text language modeling -- SpeechComposer can easily extend to more speech tasks via compositions of well-designed prompt tokens, like voice conversion and speech enhancement. The unification of prompt tokens also makes it possible for knowledge sharing among different speech tasks in a more structured manner. Experimental results demonstrate that our proposed SpeechComposer can improve the performance of both primary tasks and composite tasks, showing the effectiveness of the shared prompt tokens. Remarkably, the unified decoder-only model achieves a comparable and even better performance than the baselines which are expert models designed for single tasks.

【5】 Dance-to-Music Generation with Encoder-based Textual Inversion of  Diffusion Models
标题:基于编码器的扩散模型文本反转的舞曲生成
链接:https://arxiv.org/abs/2401.17800
作者:Sifei Li,Weiming Dong,Yuxin Zhang,Fan Tang,Chongyang Ma,Oliver Deussen,Tong-Yee Lee,Changsheng Xu
备注:9 pages, 3 figures
摘要:音乐与舞蹈动作的和谐统一是生动传达舞蹈艺术精髓的关键。这种一致性也显著提升了游戏体验和动画制作的沉浸式质量。虽然在从文本描述创建高保真音乐方面已经取得了显着的进步,但当前的方法主要集中在调节诸如流派和情感基调之类的总体特征上。他们经常忽视时间节奏的细微管理,这是为舞蹈制作音乐所不可或缺的,因为它将音乐节拍与舞者的动作紧密结合在一起。认识到这一差距,我们提出了一种基于编码器的文本反转技术,用于增强文本到音乐模型的视觉控制,促进个性化的音乐生成。具体来说,我们开发了双路径节奏体裁反转,以有效地将舞蹈动作序列的节奏和体裁整合到文本到音乐模型的文本空间中。与经典的文本反转方法相反,该方法直接更新文本嵌入以重建单个目标对象,我们的方法利用单独的节奏和体裁编码器来获得两个伪词的文本嵌入,以适应不同的节奏和体裁。为了实现更准确的评估,我们提出了改进的评估指标的节奏对齐。我们证明了我们的方法在多个评估指标上优于最先进的方法。此外,我们的方法无缝地适应野外数据,并有效地与预训练模型的固有文本指导生成能力相结合。示例可在\url{https://youtu.be/D7XDwtH1YwE}获得。
摘要:The harmonious integration of music with dance movements is pivotal in vividly conveying the artistic essence of dance. This alignment also significantly elevates the immersive quality of gaming experiences and animation productions. While there has been remarkable advancement in creating high-fidelity music from textual descriptions, current methodologies mainly concentrate on modulating overarching characteristics such as genre and emotional tone. They often overlook the nuanced management of temporal rhythm, which is indispensable in crafting music for dance, since it intricately aligns the musical beats with the dancers' movements. Recognizing this gap, we propose an encoder-based textual inversion technique for augmenting text-to-music models with visual control, facilitating personalized music generation. Specifically, we develop dual-path rhythm-genre inversion to effectively integrate the rhythm and genre of a dance motion sequence into the textual space of a text-to-music model. Contrary to the classical textual inversion method, which directly updates text embeddings to reconstruct a single target object, our approach utilizes separate rhythm and genre encoders to obtain text embeddings for two pseudo-words, adapting to the varying rhythms and genres. To achieve a more accurate evaluation, we propose improved evaluation metrics for rhythm alignment. We demonstrate that our approach outperforms state-of-the-art methods across multiple evaluation metrics. Furthermore, our method seamlessly adapts to in-the-wild data and effectively integrates with the inherent text-guided generation capability of the pre-trained model. Samples are available at \url{https://youtu.be/D7XDwtH1YwE}.

【6】 Exploiting Audio-Visual Features with Pretrained AV-HuBERT for  Multi-Modal Dysarthric Speech Reconstruction
标题:利用预先训练的AV-Hubert的视听特征进行多模节律语音重建
链接:https://arxiv.org/abs/2401.17796
作者:Xueyuan Chen,Yuejiao Wang,Xixin Wu,Disong Wang,Zhiyong Wu,Xunying Liu,Helen Meng
备注:Accepted by ICASSP 2024
摘要:构音障碍语音重建(DSR)旨在通过提高构音障碍语音的可懂度和自然度,将构音障碍语音转换为正常语音。这是一项具有挑战性的任务,特别是对于患有严重构音障碍并在复杂嘈杂的声学环境中说话的患者。为了应对这些挑战,我们提出了一种新的多模态框架来利用视觉信息,例如,嘴唇的运动,在DSR重建高度异常的发音作为额外的线索。多模态框架包括:(i)多模态编码器,用于利用辅助视觉特征从构音障碍语音中提取鲁棒的音素嵌入;(ii)方差适配器,用于从提取的音素嵌入中推断正常的音素持续时间和音高轮廓;(iii)说话者编码器,用于编码说话者的语音特征;以及(iv)梅尔解码器,用于基于所提取的音素嵌入、韵律特征和说话者嵌入来生成重构的梅尔声谱图。在常用的UASpeech语料库上进行的客观和主观评估表明,我们提出的方法可以在语音可懂度和自然度方面实现比基线系统的显着改善,特别是对于具有更严重症状的扬声器。与原始构音障碍语音相比,对于构音障碍程度更严重的患者,重建语音的绝对错误率降低了42.1%。
摘要:Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in complex, noisy acoustic environments. To address these challenges, we propose a novel multi-modal framework to utilize visual information, e.g., lip movements, in DSR as extra clues for reconstructing the highly abnormal pronunciations. The multi-modal framework consists of: (i) a multi-modal encoder to extract robust phoneme embeddings from dysarthric speech with auxiliary visual features; (ii) a variance adaptor to infer the normal phoneme duration and pitch contour from the extracted phoneme embeddings; (iii) a speaker encoder to encode the speaker's voice characteristics; and (iv) a mel-decoder to generate the reconstructed mel-spectrogram based on the extracted phoneme embeddings, prosodic features and speaker embeddings. Both objective and subjective evaluations conducted on the commonly used UASpeech corpus show that our proposed approach can achieve significant improvements over baseline systems in terms of speech intelligibility and naturalness, especially for the speakers with more severe symptoms. Compared with original dysarthric speech, the reconstructed speech achieves 42.1\% absolute word error rate reduction for patients with more severe dysarthria levels.


【7】 Harnessing Smartwatch Microphone Sensors for Cough Detection and  Classification
标题:利用智能手表麦克风传感器进行咳嗽检测和分类
链接:https://arxiv.org/abs/2401.17738
作者:Pranay Jaiswal,Haroon R. Lone
备注:7 pages
摘要:这项研究调查了使用内置麦克风传感器的智能手表来监测咳嗽和检测各种咳嗽类型的潜力。我们进行了一项涉及32名参与者的研究,并以受控方式收集了9小时的音频数据。之后,我们使用结构化方法处理这些数据,得到223个阳性咳嗽样本。我们通过增强技术进一步改进了数据集,并采用了专门的1D CNN模型。该模型在非行走时的准确率为98.49%,在行走时为98.2%,这表明智能手表可以检测到咳嗽。此外,我们的研究成功地确定了四种不同类型的咳嗽使用聚类技术。
摘要:This study investigates the potential of using smartwatches with built-in microphone sensors for monitoring coughs and detecting various cough types. We conducted a study involving 32 participants and collected 9 hours of audio data in a controlled manner. Afterward, we processed this data using a structured approach, resulting in 223 positive cough samples. We further improved the dataset through augmentation techniques and employed a specialized 1D CNN model. This model achieved an impressive accuracy rate of 98.49% while non-walking and 98.2% while walking, showing smartwatches can detect cough. Moreover, our research successfully identified four distinct types of coughs using clustering techniques.

【8】 What Do Self-Supervised Speech and Speaker Models Learn? New Findings  From a Cross Model Layer-Wise Analysis
标题:自我监督的语音和说话人模型学到了什么?跨模型分层分析的新发现
链接:https://arxiv.org/abs/2401.17632
作者:Takanori Ashihara,Marc Delcroix,Takafumi Moriya,Kohei Matsuura,Taichi Asami,Yusuke Ijima
备注:Accepted at ICASSP 2024
摘要:自监督学习(SSL)在学习有意义的语音表示方面引起了越来越多的关注。语音SSL模型,如WavLM,采用掩码预测训练来编码通用表示。相比之下,以DINO为基础的说话人SSL模型主要针对说话人表示采用话语级训练目标。了解这些模型如何表示信息对于改进模型的效率和有效性至关重要。与语音SSL的各种分析不同,对说话人SSL捕获的信息以及其表示与语音SSL或其他完全监督的说话人模型的不同之处进行了有限的调查。本文探讨了这些基本问题。我们探索的能力,以捕捉各种语音属性,通过应用SUPERB评估探测任务的语音和扬声器SSL模型。我们还研究了哪些层主要用于每个任务,以确定语音表示方式的差异。此外,我们进行直接比较,以测量模型内部和模型之间的层之间的相似性。我们的分析揭示了1)表示内容信息的能力与增强的说话人表示无关,2)语音SSL模型的特定层将部分专门用于捕获语言信息,3)说话人SSL模型倾向于忽略语言信息,但表现出更复杂的说话人表示。
摘要:Self-supervised learning (SSL) has attracted increased attention for learning meaningful speech representations. Speech SSL models, such as WavLM, employ masked prediction training to encode general-purpose representations. In contrast, speaker SSL models, exemplified by DINO-based models, adopt utterance-level training objectives primarily for speaker representation. Understanding how these models represent information is essential for refining model efficiency and effectiveness. Unlike the various analyses of speech SSL, there has been limited investigation into what information speaker SSL captures and how its representation differs from speech SSL or other fully-supervised speaker models. This paper addresses these fundamental questions. We explore the capacity to capture various speech properties by applying SUPERB evaluation probing tasks to speech and speaker SSL models. We also examine which layers are predominantly utilized for each task to identify differences in how speech is represented. Furthermore, we conduct direct comparisons to measure the similarities between layers within and across models. Our analysis unveils that 1) the capacity to represent content information is somewhat unrelated to enhanced speaker representation, 2) specific layers of speech SSL models would be partly specialized in capturing linguistic information, and 3) speaker SSL models tend to disregard linguistic information but exhibit more sophisticated speaker representation.


【9】 Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and  KiSing-v2
标题:歌唱语音数据的扩展:ACE-Opencpop和KiSing-v2介绍
链接:https://arxiv.org/abs/2401.17619
作者:Jiatong Shi,Yueqian Lin,Xinyi Bai,Keyi Zhang,Yuning Wu,Yuxun Tang,Yifeng Yu,Qin Jin,Shinji Watanabe
摘要:在歌声合成(SVS)中,从乐谱生成歌声面临着有限的数据可用性的挑战,这在文本到语音(TTS)中不太常见。这项研究提出了一种新的方法来解决这种数据稀缺性。我们利用现有的歌唱声音合成器进行数据增强,并应用精确的手动调谐,以减少不自然的声音合成。我们开发了两个广泛的歌唱声音语料库,ACE-Opencpop和KiSing-v2,便于大规模,多歌手的声音合成。利用来自这些语料库的预训练模型,我们在语音质量方面取得了显着的改善,在域内和域外场景中都很明显。语料库、预训练模型及其相关训练方法可在Muskits-ESPnet(https://github.com/espnet/espnet)上公开获取。
摘要:In singing voice synthesis (SVS), generating singing voices from musical scores faces challenges due to limited data availability, a constraint less common in text-to-speech (TTS). This study proposes a new approach to address this data scarcity. We utilize an existing singing voice synthesizer for data augmentation and apply precise manual tuning to reduce unnatural voice synthesis. Our development of two extensive singing voice corpora, ACE-Opencpop and KiSing-v2, facilitates large-scale, multi-singer voice synthesis. Utilizing pre-trained models derived from these corpora, we achieve notable improvements in voice quality, evident in both in-domain and out-of-domain scenarios. The corpora, pre-trained models, and their related training recipes are publicly available at Muskits-ESPnet (https://github.com/espnet/espnet).


【10】 Computation and Parameter Efficient Multi-Modal Fusion Transformer for  Cued Speech Recognition
标题:用于线索语音识别的计算和参数高效的多模式融合转换器
链接:https://arxiv.org/abs/2401.17604
作者:Lei Liu,Li Liu,Haizhou Li
备注:Accepted by TASLP
摘要:提示语音(CS)是一种纯视觉编码方法,由听力受损的人使用,结合唇读与几个特定的手形,使口语可见。自动CS识别(ACSR)旨在将语音的视觉线索转换为文本,这可以帮助听障人士有效地进行交流。CS的视觉信息包括唇读和手暗示,因此唇读和手暗示的融合在ACSR中起着重要的作用。然而,大多数以前的融合方法的斗争,以捕捉全球的依赖性,目前在长序列输入的多模态CS数据。因此,这些方法通常无法学习有助于融合的有效跨模态关系。最近,基于注意力的Transformers已经成为一个流行的想法,在多模态融合中捕获的全局依赖性的长序列,但现有的多模态融合变换器的识别精度差和效率低下的ACSR任务的计算。为了解决这些问题,我们开发了一种新的计算和参数有效的多模态融合Transformer,提出了一种新的令牌重要性感知注意机制(TIAA),其中令牌利用率(TUR)制定选择重要的令牌从多模态流。更准确地说,TIAA首先模型的模态特定的细粒度的时间依赖性的所有令牌的每一个模态,然后学习有效的跨模态交互的模态共享的粗粒度的时间依赖性的重要令牌的不同模态。此外,还设计了一个轻量级的门控隐藏投影来控制TIAA的特征流。与现有的基于变换的融合方法和ACSR融合方法相比,所得到的模型,命名为经济提示语音融合Transformer(EcoCued),实现了最先进的性能在所有现有的CS数据集。
摘要:Cued Speech (CS) is a pure visual coding method used by hearing-impaired people that combines lip reading with several specific hand shapes to make the spoken language visible. Automatic CS recognition (ACSR) seeks to transcribe visual cues of speech into text, which can help hearing-impaired people to communicate effectively. The visual information of CS contains lip reading and hand cueing, thus the fusion of them plays an important role in ACSR. However, most previous fusion methods struggle to capture the global dependency present in long sequence inputs of multi-modal CS data. As a result, these methods generally fail to learn the effective cross-modal relationships that contribute to the fusion. Recently, attention-based transformers have been a prevalent idea for capturing the global dependency over the long sequence in multi-modal fusion, but existing multi-modal fusion transformers suffer from both poor recognition accuracy and inefficient computation for the ACSR task. To address these problems, we develop a novel computation and parameter efficient multi-modal fusion transformer by proposing a novel Token-Importance-Aware Attention mechanism (TIAA), where a token utilization rate (TUR) is formulated to select the important tokens from the multi-modal streams. More precisely, TIAA firstly models the modality-specific fine-grained temporal dependencies over all tokens of each modality, and then learns the efficient cross-modal interaction for the modality-shared coarse-grained temporal dependencies over the important tokens of different modalities. Besides, a light-weight gated hidden projection is designed to control the feature flows of TIAA. The resulting model, named Economical Cued Speech Fusion Transformer (EcoCued), achieves state-of-the-art performance on all existing CS datasets, compared with existing transformer-based fusion methods and ACSR fusion methods.


【11】 Sigma-lognormal modeling of speech
标题:语音的Sigma-对数正态模型
链接:https://arxiv.org/abs/2401.17320
作者:C. Carmona-Duarte,M. A. Ferrer,R. Plamondon,A. Gomez-Rodellar,P. Gomez-Vilda
备注:None
摘要:人体运动研究和分析在许多科学领域都是基础性的,从神经科学到教育,从模式识别到机器人技术,从医疗保健到体育等等。先前的言语运动模型被提出来理解言语运动是如何产生的,以及当某些参数改变时,所产生的言语是如何变化的。然而,逆的方法,其中的肌肉反应参数和主体的年龄是从真正的连续语音,是不可能的,与这样的模型。相反,在手写领域,快速人体运动的运动学理论及其相关的Sigma-lognormal模型已成功地应用于获得肌肉响应参数。这项工作提出了一个语音运动学为基础的模型,可以用来研究,分析和重建复杂的语音运动学在一个简化的方式。一种方法的基础上的快速人体运动的运动学理论及其相关的Sigma对数正态模型被施加到描述和参数化的渐近脉冲响应的神经肌肉网络参与语音作为响应神经运动命令。还介绍了用于进行从共振峰到运动观测的转换的方法。与(英语)VTR TIMIT数据库和(德国)Saarbrucken语音数据库,包括不同年龄的人,喉病变和不喉病变进行的实验,证实了提取的参数和老化之间的联系,一方面,和应用快速人体运动的运动学理论所需的第一和第二共振峰之间的比例,另一方面。研究结果将推动语音运动学建模和理解的创新发展。
摘要:Human movement studies and analyses have been fundamental in many scientific domains, ranging from neuroscience to education, pattern recognition to robotics, health care to sports, and beyond. Previous speech motor models were proposed to understand how speech movement is produced and how the resulting speech varies when some parameters are changed. However, the inverse approach, in which the muscular response parameters and the subject's age are derived from real continuous speech, is not possible with such models. Instead, in the handwriting field, the kinematic theory of rapid human movements and its associated Sigma-lognormal model have been applied successfully to obtain the muscular response parameters. This work presents a speech kinematics based model that can be used to study, analyze, and reconstruct complex speech kinematics in a simplified manner. A method based on the kinematic theory of rapid human movements and its associated Sigma lognormal model are applied to describe and to parameterize the asymptotic impulse response of the neuromuscular networks involved in speech as a response to a neuromotor command. The method used to carry out transformations from formants to a movement observation is also presented. Experiments carried out with the (English) VTR TIMIT database and the (German) Saarbrucken Voice Database, including people of different ages, with and without laryngeal pathologies, corroborate the link between the extracted parameters and aging, on the one hand, and the proportion between the first and second formants required in applying the kinematic theory of rapid human movements, on the other. The results should drive innovative developments in the modeling and understanding of speech kinematics.

【12】 Detection of Auditory Brainstem Response Peaks Using Image Processing  Techniques in Infants with Normal Hearing Sensitivity
标题:用图像处理技术检测听力正常婴儿的听性脑干反应峰值
链接:https://arxiv.org/abs/2401.17317
作者:Amir Majidpour,Samer Kais Jameel,Jafar Majidpour,Houra Bagheri,Tarik A. Rashid,Ahmadreza Nazeri,Mahshid Moheb Aleaba
摘要:简介:通过对听力正常儿童进行听性脑干反应(ABR)检测,以了解脑干水平周围听神经系统的完整性。听觉诱发电位(AEP)是利用声刺激产生的。解释这些波需要能力,以避免误诊听力问题。使用计算机视觉自动化ABR测试标记可以减少人为错误。方法:对26例1 ~ 20个月左右双耳听力正常儿童进行ABR测试。提出了一种新的方法,用于自动计算不同强度(分贝)的波的峰值。该程序需要使用Color Waveholder方法从Audera设备获取波形图像,使用Image Region Analyzer应用程序将每个波形分割为单个波形图像,使用图像处理(IP)技术将所有波形图像转换为波形,最后计算每个波形的峰值延迟,以供听力学家用于诊断疾病。调查结果:图像处理技术能够检测诊断视野中的1、3和5个波,准确度分别为0.82、0.98和0.98,并且其对波1、3和5的精确度分别为0.32、0.97和0.87。该方法在阈值部分也有良好的效果,ABR波的正确检出率为82.7%.结论:我们的研究结果表明,通过使用自动检测和标记ABR波的技术,听力学测试组合套件可以更加准确,快速和无错误。
摘要:Introduction: The auditory brainstem response (ABR) is measured to find the brainstem-level peripheral auditory nerve system integrity in children having normal hearing. The Auditory Evoked Potential (AEP) is generated using acoustic stimuli. Interpreting these waves requires competence to avoid misdiagnosing hearing problems. Automating ABR test labeling with computer vision may reduce human error. Method: The ABR test results of 26 children aged 1 to 20 months with normal hearing in both ears were used. A new approach is suggested for automatically calculating the peaks of waves of different intensities (in decibels). The procedure entails acquiring wave images from an Audera device using the Color Thresholder method, segmenting each wave as a single wave image using the Image Region Analyzer application, converting all wave images into waves using Image Processing (IP) techniques, and finally calculating the latency of the peaks for each wave to be used by an audiologist for diagnosing the disease. Findings: Image processing techniques were able to detect 1, 3, and 5 waves in the diagnosis field with accuracy (0.82), (0.98), and (0.98), respectively, and its precision for waves 1, 3, and 5, were respectively (0.32), (0.97) and (0.87). This evaluation also worked well in the thresholding part and 82.7 % correctly detected the ABR waves. Conclusion: Our findings indicate that the audiology test battery suite can be made more accurate, quick, and error-free by using technology to automatically detect and label ABR waves.


机器翻译由腾讯交互翻译提供,仅供参考