今日论文合集:cs.SD语音9篇,eess.AS音频处理9篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Probing the Information Encoded in Neural-based Acoustic Models of  Automatic Speech Recognition Systems
标题:自动语音识别系统中基于神经元声学模型的信息编码探讨
链接:https://arxiv.org/abs/2402.19443
作者:Quentin Raymondaud,Mickael Rouvier,Richard Dufour
摘要:深度学习架构在许多研究领域的性能方面取得了重大进展。因此,自动语音识别(ASR)领域受益于这些科学和技术进步,特别是声学建模,现在集成了深度神经网络架构。然而,这些性能的提高已经转化为通过这些黑盒架构学习和传达的信息的复杂性的增加。在神经网络可解释性的许多研究之后,我们在这篇文章中提出了一个协议,旨在确定哪些信息位于ASR声学模型(AM)中。要做到这一点,我们建议使用中间表示(在这里,在不同的层级别)在确定的一组任务上评估AM性能。关于性能变化和目标任务,我们可以提出关于在不同架构步骤中增强或扰动哪些信息的假设。实验在说话人确认、声环境分类、性别分类、时间失真检测系统和语音情感/情感识别系统上进行。分析表明,基于神经元的AM持有的异质信息,似乎令人惊讶地与音素识别无关,如情绪,情感或说话人身份。低层次的隐藏层在全局上看起来对信息的结构化有用,而高层次的隐藏层往往会删除对音素识别无用的信息。
摘要:Deep learning architectures have made significant progress in terms of performance in many research areas. The automatic speech recognition (ASR) field has thus benefited from these scientific and technological advances, particularly for acoustic modeling, now integrating deep neural network architectures. However, these performance gains have translated into increased complexity regarding the information learned and conveyed through these black-box architectures. Following many researches in neural networks interpretability, we propose in this article a protocol that aims to determine which and where information is located in an ASR acoustic model (AM). To do so, we propose to evaluate AM performance on a determined set of tasks using intermediate representations (here, at different layer levels). Regarding the performance variation and targeted tasks, we can emit hypothesis about which information is enhanced or perturbed at different architecture steps. Experiments are performed on both speaker verification, acoustic environment classification, gender classification, tempo-distortion detection systems and speech sentiment/emotion identification. Analysis showed that neural-based AMs hold heterogeneous information that seems surprisingly uncorrelated with phoneme recognition, such as emotion, sentiment or speaker identity. The low-level hidden layers globally appears useful for the structuring of information while the upper ones would tend to delete useless information for phoneme recognition.

【2】 Unraveling Adversarial Examples against Speaker Identification --  Techniques for Attack Detection and Victim Model Classification
标题:揭开说话人识别的敌意实例--攻击检测和受害者模型分类技术
链接:https://arxiv.org/abs/2402.19355
作者:Sonal Joshi,Thomas Thebaud,Jesús Villalba,Najim Dehak
摘要:对抗性实例已经被证明是对说话人识别系统的威胁,并且已经提出了针对它们的一些对策。在本文中,我们提出了一种方法来检测对抗性示例的存在,即,一个区分良性和敌对样本的二元分类器。我们建立和扩展以前的工作,通过探索新的体系结构的攻击类型分类。此外,我们还介绍了一种识别对抗性攻击的受害者模型的方法。为了实现这一目标,我们生成了一个新的数据集,其中包含针对各种受害者模型执行的多个攻击。我们实现了0.982的攻击检测AUC,对于未知攻击,性能下降不超过0.03。使用LightResNet34架构,我们的攻击分类准确率(不包括良性)在八种攻击类型中达到86.48%,而我们的受害者模型分类准确率在四种受害者模型中达到72.28%。
摘要:Adversarial examples have proven to threaten speaker identification systems, and several countermeasures against them have been proposed. In this paper, we propose a method to detect the presence of adversarial examples, i.e., a binary classifier distinguishing between benign and adversarial examples. We build upon and extend previous work on attack type classification by exploring new architectures. Additionally, we introduce a method for identifying the victim model on which the adversarial attack is carried out. To achieve this, we generate a new dataset containing multiple attacks performed against various victim models. We achieve an AUC of 0.982 for attack detection, with no more than a 0.03 drop in performance for unknown attacks. Our attack classification accuracy (excluding benign) reaches 86.48% across eight attack types using our LightResNet34 architecture, while our victim model classification accuracy reaches 72.28% across four victim models.

【3】 Compact Speech Translation Models via Discrete Speech Units Pretraining
标题:基于离散语音单元预训练的紧凑语音翻译模型
链接:https://arxiv.org/abs/2402.19333
作者:Tsz Kin Lam,Alexandra Birch,Barry Haddow
摘要:使用自监督学习(SSL)作为模型初始化现在很常见,以获得语音翻译(ST)的强大结果。但是,它们也会占用大量内存,阻碍设备上的部署。在本文中,我们通过在其离散语音单元(DSU)上预训练较小的模型来利用SSL模型。我们在1)滤波器组到DSU和2)DSU到翻译数据上预训练编码器-解码器模型,并将编码器从1)和解码器从2)初始化一个新模型,在有限的语音翻译数据上对其进行微调。通过使用DSU预训练来扩充SSL模型的知识,最终模型变得紧凑。与使用DSU作为模型输入相比,我们的方法有几个优点,例如更短的推理流水线和对(DSU)标记化的鲁棒性。与ASR预培训相比,它不需要成绩单,使其适用于低资源环境。在CoVoST-2 X-En上的评估表明,我们的方法比直接微调SSL模型的ST模型好> 0.5 $ BLEU,只需要一半的模型大小,并且与ASR预训练相当。
摘要:Using Self-Supervised Learning (SSL) as model initialization is now common to obtain strong results in Speech Translation (ST). However, they also impose a large memory footprint, hindering on-device deployment. In this paper, we leverage the SSL models by pretraining smaller models on their Discrete Speech Units (DSU). We pretrain encoder-decoder models on 1) Filterbank-to-DSU and 2) DSU-to-Translation data, and take the encoder from 1) and the decoder from 2) to initialise a new model, finetuning this on limited speech-translation data. The final model becomes compact by using the DSU pretraining to distil the knowledge of the SSL model. Our method has several benefits over using DSU as model inputs, such as shorter inference pipeline and robustness over (DSU) tokenization. In contrast to ASR pretraining, it does not require transcripts, making it applicable to low-resource settings. Evaluation on CoVoST-2 X-En shows that our method is >$0.5$ BLEU better than a ST model that directly finetune the SSL model, given only half the model size, and on a par with ASR pretraining.

【4】 Do End-to-End Neural Diarization Attractors Need to Encode Speaker  Characteristic Information?
标题:端到端神经二元化吸引子需要编码说话人特征信息吗?
链接:https://arxiv.org/abs/2402.19325
作者:Lin Zhang,Themos Stafylakis,Federico Landini,Mireia Diez,Anna Silnova,Lukáš Burget
备注:Submitted to Odyssey 2024
摘要:在本文中,我们应用变分信息瓶颈方法的端到端的神经日记与编码器-解码器吸引子(EEND-EDA)。这使我们能够研究哪些信息对模型至关重要。EEND-EDA利用对话中的说话者的矢量表示-吸引子。我们的分析表明,吸引子不一定包含说话人特征信息。另一方面,给予吸引子更多的自由,允许它们编码一些额外的(可能是特定于说话者的)信息,导致小但一致的日志化性能改进。尽管EEND系统在架构上存在差异,但吸引子和帧嵌入的概念对大多数系统来说都是通用的,而不是EEND-EDA所特有的。我们相信,这项工作的主要结论可以适用于其他变体的EEND。因此,我们希望本文将是一个有价值的贡献,以指导社会作出更明智的决定时,设计新的系统。
摘要:In this paper, we apply the variational information bottleneck approach to end-to-end neural diarization with encoder-decoder attractors (EEND-EDA). This allows us to investigate what information is essential for the model. EEND-EDA utilizes vector representations of the speakers in a conversation - attractors. Our analysis shows that, attractors do not necessarily have to contain speaker characteristic information. On the other hand, giving the attractors more freedom allowing them to encode some extra (possibly speaker-specific) information leads to small but consistent diarization performance improvements. Despite architectural differences in EEND systems, the notion of attractors and frame embeddings is common to most of them and not specific to EEND-EDA. We believe that the main conclusions of this work can apply to other variants of EEND. Thus, we hope this paper will be a valuable contribution to guide the community to make more informed decisions when designing new systems.


【5】 Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale  Speech Recognition
标题:基于大规模语音识别的动态语音异常停顿检测
链接:https://arxiv.org/abs/2402.18923
作者:Jeehyun Lee,Yerin Choi,Tae-Jin Song,Myoung-Wan Koo
备注:Accepted to ICASSP 2024
摘要:构音障碍是中风患者的常见问题,严重影响语言清晰度。不适当的停顿是严重程度评估和语言治疗的关键指标。我们建议扩展一个大规模的语音识别模型,用于构音障碍语音中的不适当停顿检测。为此,我们提出了任务设计,标签策略,和语音识别模型与不适当的停顿预测层。首先,我们将停顿检测视为语音识别,使用自动语音识别(ASR)模型将语音转换为带有停顿标签的文本。根据新设计的任务,我们在文本层面上标记停顿位置及其适当性。我们与言语语言病理学家合作,建立标签标准,确保高质量的注释数据。最后,我们扩展了ASR模型与不适当的暂停预测层的端到端不适当的暂停检测。此外,我们提出了一个任务量身定制的度量评估不适当的暂停检测独立的ASR性能。我们的实验表明,该方法更好地检测不适当的停顿在构音障碍的语音比基线。(不适当的错误率:14.47%)
摘要:Dysarthria, a common issue among stroke patients, severely impacts speech intelligibility. Inappropriate pauses are crucial indicators in severity assessment and speech-language therapy. We propose to extend a large-scale speech recognition model for inappropriate pause detection in dysarthric speech. To this end, we propose task design, labeling strategy, and a speech recognition model with an inappropriate pause prediction layer. First, we treat pause detection as speech recognition, using an automatic speech recognition (ASR) model to convert speech into text with pause tags. According to the newly designed task, we label pause locations at the text level and their appropriateness. We collaborate with speech-language pathologists to establish labeling criteria, ensuring high-quality annotated data. Finally, we extend the ASR model with an inappropriate pause prediction layer for end-to-end inappropriate pause detection. Moreover, we propose a task-tailored metric for evaluating inappropriate pause detection independent of ASR performance. Our experiments show that the proposed method better detects inappropriate pauses in dysarthric speech than baselines. (Inappropriate Pause Error Rate: 14.47%)

【6】 Point Processes and spatial statistics in time-frequency analysis
标题:时频分析中的点过程与空间统计
链接:https://arxiv.org/abs/2402.19172
作者:Barbara Pascal,Rémi Bardenet
备注:Submitted
摘要:A finite-energy signal is represented by a square-integrable, complex-valued function $t\mapsto s(t)$ of a real variable $t$, interpreted as time. Similarly, a noisy signal is represented by a random process. Time-frequency analysis, a subfield of signal processing, amounts to describing the temporal evolution of the frequency content of a signal. Loosely speaking, if $s$ is the audio recording of a musical piece, time-frequency analysis somehow consists in writing the musical score of the piece. Mathematically, the operation is performed through a transform $\mathcal{V}$, mapping $s \in L^2(\mathbb{R})$ onto a complex-valued function $\mathcal{V}s \in L^2(\mathbb{R}^2)$ of time $t$ and angular frequency $\omega$. The squared modulus $(t, \omega) \mapsto \vert\mathcal{V}s(t,\omega)\vert^2$ of the time-frequency representation is known as the spectrogram of $s$; in the musical score analogy, a peaked spectrogram at $(t_0,\omega_0)$ corresponds to a musical note at angular frequency $\omega_0$ localized at time $t_0$. More generally, the intuition is that upper level sets of the spectrogram contain relevant information about in the original signal. Hence, many signal processing algorithms revolve around identifying maxima of the spectrogram. In contrast, zeros of the spectrogram indicate perfect silence, that is, a time at which a particular frequency is absent. Assimilating $\mathbb{R}^2$ to $\mathbb{C}$ through $z = \omega + \mathrm{i}t$, this chapter focuses on time-frequency transforms $\mathcal{V}$ that map signals to analytic functions. The zeros of the spectrogram of a noisy signal are then the zeros of a random analytic function, hence forming a Point Process in $\mathbb{C}$. This chapter is devoted to the study of these Point Processes, to their links with zeros of Gaussian Analytic Functions, and to designing signal detection and denoising algorithms using spatial statistics.
摘要:A finite-energy signal is represented by a square-integrable, complex-valued function $t\mapsto s(t)$ of a real variable $t$, interpreted as time. Similarly, a noisy signal is represented by a random process. Time-frequency analysis, a subfield of signal processing, amounts to describing the temporal evolution of the frequency content of a signal. Loosely speaking, if $s$ is the audio recording of a musical piece, time-frequency analysis somehow consists in writing the musical score of the piece. Mathematically, the operation is performed through a transform $\mathcal{V}$, mapping $s \in L^2(\mathbb{R})$ onto a complex-valued function $\mathcal{V}s \in L^2(\mathbb{R}^2)$ of time $t$ and angular frequency $\omega$. The squared modulus $(t, \omega) \mapsto \vert\mathcal{V}s(t,\omega)\vert^2$ of the time-frequency representation is known as the spectrogram of $s$; in the musical score analogy, a peaked spectrogram at $(t_0,\omega_0)$ corresponds to a musical note at angular frequency $\omega_0$ localized at time $t_0$. More generally, the intuition is that upper level sets of the spectrogram contain relevant information about in the original signal. Hence, many signal processing algorithms revolve around identifying maxima of the spectrogram. In contrast, zeros of the spectrogram indicate perfect silence, that is, a time at which a particular frequency is absent. Assimilating $\mathbb{R}^2$ to $\mathbb{C}$ through $z = \omega + \mathrm{i}t$, this chapter focuses on time-frequency transforms $\mathcal{V}$ that map signals to analytic functions. The zeros of the spectrogram of a noisy signal are then the zeros of a random analytic function, hence forming a Point Process in $\mathbb{C}$. This chapter is devoted to the study of these Point Processes, to their links with zeros of Gaussian Analytic Functions, and to designing signal detection and denoising algorithms using spatial statistics.

【7】 A SOUND APPROACH: Using Large Language Models to generate audio  descriptions for egocentric text-audio retrieval
标题:一种合理的方法:使用大语言模型为自我中心的文本音频检索生成音频描述
链接:https://arxiv.org/abs/2402.19106
作者:Andreea-Maria Oncescu,João F. Henriques,Andrew Zisserman,Samuel Albanie,A. Sophia Koepke
备注:9 pages, 2 figures, 9 tables, Accepted at ICASSP 2024
摘要:来自互联网的视频数据库是文本音频检索数据集的宝贵来源。然而,考虑到声音和图像流代表数据的不同“视图”,将视觉描述视为音频描述远非最佳。即使音频类别标签存在,它们通常不是很详细,使它们不适合文本音频检索。为了利用视频文本数据集中的相关音频信息,我们介绍了一种使用大型语言模型(LLM)生成以音频为中心的描述的方法。在这项工作中,我们考虑以自我为中心的视频设置,并提出了三个新的文本音频检索基准的基础上EpicMIR和EgoMCQ任务,和EpicSounds数据集。我们的方法获得音频为中心的描述提供了显着更高的zero-shot性能比使用原来的视觉为中心的描述。此外,我们表明,使用相同的提示,我们可以成功地采用LLM,以提高检索EpicSounds相比,使用原始的音频类标签的数据集。最后,我们确认,LLM可以用来确定识别与声音相关的动作的难度。
摘要:Video databases from the internet are a valuable source of text-audio retrieval datasets. However, given that sound and vision streams represent different "views" of the data, treating visual descriptions as audio descriptions is far from optimal. Even if audio class labels are present, they commonly are not very detailed, making them unsuited for text-audio retrieval. To exploit relevant audio information from video-text datasets, we introduce a methodology for generating audio-centric descriptions using Large Language Models (LLMs). In this work, we consider the egocentric video setting and propose three new text-audio retrieval benchmarks based on the EpicMIR and EgoMCQ tasks, and on the EpicSounds dataset. Our approach for obtaining audio-centric descriptions gives significantly higher zero-shot performance than using the original visual-centric descriptions. Furthermore, we show that using the same prompts, we can successfully employ LLMs to improve the retrieval on EpicSounds, compared to using the original audio class labels of the dataset. Finally, we confirm that LLMs can be used to determine the difficulty of identifying the action associated with a sound.

【8】 Ambisonics Networks -- The Effect Of Radial Functions Regularization
标题:兼性网络--径向函数正则化的影响
链接:https://arxiv.org/abs/2402.18968
作者:Bar Shaybet,Anurag Kumar,Vladimir Tourbabin,Boaz Rafaely
备注:to be published in Icassp 2024
摘要:立体混响是一种流行的空间音频格式,是声场的平面波密度函数的球谐(SH)表示。许多算法在SH域中操作并且利用高保真度立体声响复制作为它们的输入信号。对来自球形麦克风阵列的高保真度立体声响复制进行编码的过程涉及除以径向函数,这可能放大低频处的噪声。这可以通过正则化来克服,缺点是将错误引入到高保真度立体声响复制编码中。本文旨在研究不同的正则化方法对深度神经网络(DNN)训练和性能的影响。理想情况下,这些网络应该对正则化方法具有鲁棒性。在一个房间里的单个扬声器的模拟数据和实验数据从LOCATA的挑战被用来评估这种鲁棒性的扬声器定位的基础上的直接路径优势(DPD)测试的示例算法。结果表明,性能可能是敏感的正则化的方式,并提出了一个明智的方法和调查,突出正则化信息的重要性。
摘要:Ambisonics, a popular format of spatial audio, is the spherical harmonic (SH) representation of the plane wave density function of a sound field. Many algorithms operate in the SH domain and utilize the Ambisonics as their input signal. The process of encoding Ambisonics from a spherical microphone array involves dividing by the radial functions, which may amplify noise at low frequencies. This can be overcome by regularization, with the downside of introducing errors to the Ambisonics encoding. This paper aims to investigate the impact of different ways of regularization on Deep Neural Network (DNN) training and performance. Ideally, these networks should be robust to the way of regularization. Simulated data of a single speaker in a room and experimental data from the LOCATA challenge were used to evaluate this robustness on an example algorithm of speaker localization based on the direct-path dominance (DPD) test. Results show that performance may be sensitive to the way of regularization, and an informed approach is proposed and investigated, highlighting the importance of regularization information.

【9】 Extending Multilingual Speech Synthesis to 100+ Languages without  Transcribed Data
标题:将多语言语音合成扩展到100多种语言,无需转录数据
链接:https://arxiv.org/abs/2402.18932
作者:Takaaki Saeki,Gary Wang,Nobuyuki Morioka,Isaac Elias,Kyle Kastner,Andrew Rosenberg,Bhuvana Ramabhadran,Heiga Zen,Françoise Beaufays,Hadar Shemtov
备注:To appear in ICASSP 2024
摘要:收集高质量的录音室录音是一项挑战,这限制了文本到语音(TTS)系统的语言覆盖范围。本文提出了一个框架,可以在没有监督的情况下使用发现的数据将多语言TTS模型扩展到100多种语言。该框架将语音-文本编码器预训练与使用未转录语音和未说出的文本数据源的无监督训练相结合,从而利用大规模多语言联合语音和文本表示学习。在没有任何新语言的转录语音的情况下,该TTS模型可以用>30种看不见的语言生成可理解的语音(CER与地面真实值的差异<10%)。只需15分钟的转录,找到数据,我们就可以将可理解性差异减少到1%或更少,并在几种语言中实现与地面实况相匹配的自然度分数。
摘要:Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-text encoder pretraining with unsupervised training using untranscribed speech and unspoken text data sources, thereby leveraging massively multilingual joint speech and text representation learning. Without any transcribed speech in a new language, this TTS model can generate intelligible speech in >30 unseen languages (CER difference of <10% to ground truth). With just 15 minutes of transcribed, found data, we can reduce the intelligibility difference to 1% or less from the ground-truth, and achieve naturalness scores that match the ground-truth in several languages.

eess.AS音频处理
【1】 Point Processes and spatial statistics in time-frequency analysis
标题:时频分析中的点过程与空间统计
链接:https://arxiv.org/abs/2402.19172
作者:Barbara Pascal,Rémi Bardenet
备注:Submitted
摘要:A finite-energy signal is represented by a square-integrable, complex-valued function $t\mapsto s(t)$ of a real variable $t$, interpreted as time. Similarly, a noisy signal is represented by a random process. Time-frequency analysis, a subfield of signal processing, amounts to describing the temporal evolution of the frequency content of a signal. Loosely speaking, if $s$ is the audio recording of a musical piece, time-frequency analysis somehow consists in writing the musical score of the piece. Mathematically, the operation is performed through a transform $\mathcal{V}$, mapping $s \in L^2(\mathbb{R})$ onto a complex-valued function $\mathcal{V}s \in L^2(\mathbb{R}^2)$ of time $t$ and angular frequency $\omega$. The squared modulus $(t, \omega) \mapsto \vert\mathcal{V}s(t,\omega)\vert^2$ of the time-frequency representation is known as the spectrogram of $s$; in the musical score analogy, a peaked spectrogram at $(t_0,\omega_0)$ corresponds to a musical note at angular frequency $\omega_0$ localized at time $t_0$. More generally, the intuition is that upper level sets of the spectrogram contain relevant information about in the original signal. Hence, many signal processing algorithms revolve around identifying maxima of the spectrogram. In contrast, zeros of the spectrogram indicate perfect silence, that is, a time at which a particular frequency is absent. Assimilating $\mathbb{R}^2$ to $\mathbb{C}$ through $z = \omega + \mathrm{i}t$, this chapter focuses on time-frequency transforms $\mathcal{V}$ that map signals to analytic functions. The zeros of the spectrogram of a noisy signal are then the zeros of a random analytic function, hence forming a Point Process in $\mathbb{C}$. This chapter is devoted to the study of these Point Processes, to their links with zeros of Gaussian Analytic Functions, and to designing signal detection and denoising algorithms using spatial statistics.
摘要:A finite-energy signal is represented by a square-integrable, complex-valued function $t\mapsto s(t)$ of a real variable $t$, interpreted as time. Similarly, a noisy signal is represented by a random process. Time-frequency analysis, a subfield of signal processing, amounts to describing the temporal evolution of the frequency content of a signal. Loosely speaking, if $s$ is the audio recording of a musical piece, time-frequency analysis somehow consists in writing the musical score of the piece. Mathematically, the operation is performed through a transform $\mathcal{V}$, mapping $s \in L^2(\mathbb{R})$ onto a complex-valued function $\mathcal{V}s \in L^2(\mathbb{R}^2)$ of time $t$ and angular frequency $\omega$. The squared modulus $(t, \omega) \mapsto \vert\mathcal{V}s(t,\omega)\vert^2$ of the time-frequency representation is known as the spectrogram of $s$; in the musical score analogy, a peaked spectrogram at $(t_0,\omega_0)$ corresponds to a musical note at angular frequency $\omega_0$ localized at time $t_0$. More generally, the intuition is that upper level sets of the spectrogram contain relevant information about in the original signal. Hence, many signal processing algorithms revolve around identifying maxima of the spectrogram. In contrast, zeros of the spectrogram indicate perfect silence, that is, a time at which a particular frequency is absent. Assimilating $\mathbb{R}^2$ to $\mathbb{C}$ through $z = \omega + \mathrm{i}t$, this chapter focuses on time-frequency transforms $\mathcal{V}$ that map signals to analytic functions. The zeros of the spectrogram of a noisy signal are then the zeros of a random analytic function, hence forming a Point Process in $\mathbb{C}$. This chapter is devoted to the study of these Point Processes, to their links with zeros of Gaussian Analytic Functions, and to designing signal detection and denoising algorithms using spatial statistics.

【2】 A SOUND APPROACH: Using Large Language Models to generate audio  descriptions for egocentric text-audio retrieval
标题:一种声音方法:使用大型语言模型为自我中心的文本-音频检索生成音频描述
链接:https://arxiv.org/abs/2402.19106
作者:Andreea-Maria Oncescu,João F. Henriques,Andrew Zisserman,Samuel Albanie,A. Sophia Koepke
备注:9 pages, 2 figures, 9 tables, Accepted at ICASSP 2024
摘要:来自互联网的视频数据库是文本音频检索数据集的宝贵来源。然而,考虑到声音和图像流代表数据的不同“视图”,将视觉描述视为音频描述远非最佳。即使音频类别标签存在,它们通常不是很详细,使它们不适合文本音频检索。为了利用视频文本数据集中的相关音频信息,我们介绍了一种使用大型语言模型(LLM)生成以音频为中心的描述的方法。在这项工作中,我们考虑以自我为中心的视频设置,并提出了三个新的文本音频检索基准的基础上EpicMIR和EgoMCQ任务,和EpicSounds数据集。我们的方法获得音频为中心的描述提供了显着更高的zero-shot性能比使用原来的视觉为中心的描述。此外,我们表明,使用相同的提示,我们可以成功地采用LLM,以提高检索EpicSounds相比,使用原始的音频类标签的数据集。最后,我们确认,LLM可以用来确定识别与声音相关的动作的难度。
摘要:Video databases from the internet are a valuable source of text-audio retrieval datasets. However, given that sound and vision streams represent different "views" of the data, treating visual descriptions as audio descriptions is far from optimal. Even if audio class labels are present, they commonly are not very detailed, making them unsuited for text-audio retrieval. To exploit relevant audio information from video-text datasets, we introduce a methodology for generating audio-centric descriptions using Large Language Models (LLMs). In this work, we consider the egocentric video setting and propose three new text-audio retrieval benchmarks based on the EpicMIR and EgoMCQ tasks, and on the EpicSounds dataset. Our approach for obtaining audio-centric descriptions gives significantly higher zero-shot performance than using the original visual-centric descriptions. Furthermore, we show that using the same prompts, we can successfully employ LLMs to improve the retrieval on EpicSounds, compared to using the original audio class labels of the dataset. Finally, we confirm that LLMs can be used to determine the difficulty of identifying the action associated with a sound.

【3】 Ambisonics Networks -- The Effect Of Radial Functions Regularization
标题:兼性网络--径向函数正则化的影响
链接:https://arxiv.org/abs/2402.18968
作者:Bar Shaybet,Anurag Kumar,Vladimir Tourbabin,Boaz Rafaely
备注:to be published in Icassp 2024
摘要:立体混响是一种流行的空间音频格式,是声场的平面波密度函数的球谐(SH)表示。许多算法在SH域中操作并且利用高保真度立体声响复制作为它们的输入信号。对来自球形麦克风阵列的高保真度立体声响复制进行编码的过程涉及除以径向函数,这可能放大低频处的噪声。这可以通过正则化来克服,缺点是将错误引入到高保真度立体声响复制编码中。本文旨在研究不同的正则化方法对深度神经网络(DNN)训练和性能的影响。理想情况下,这些网络应该对正则化方法具有鲁棒性。在一个房间里的单个扬声器的模拟数据和实验数据从LOCATA的挑战被用来评估这种鲁棒性的扬声器定位的基础上的直接路径优势(DPD)测试的示例算法。结果表明,性能可能是敏感的正则化的方式,并提出了一个明智的方法和调查,突出正则化信息的重要性。
摘要:Ambisonics, a popular format of spatial audio, is the spherical harmonic (SH) representation of the plane wave density function of a sound field. Many algorithms operate in the SH domain and utilize the Ambisonics as their input signal. The process of encoding Ambisonics from a spherical microphone array involves dividing by the radial functions, which may amplify noise at low frequencies. This can be overcome by regularization, with the downside of introducing errors to the Ambisonics encoding. This paper aims to investigate the impact of different ways of regularization on Deep Neural Network (DNN) training and performance. Ideally, these networks should be robust to the way of regularization. Simulated data of a single speaker in a room and experimental data from the LOCATA challenge were used to evaluate this robustness on an example algorithm of speaker localization based on the direct-path dominance (DPD) test. Results show that performance may be sensitive to the way of regularization, and an informed approach is proposed and investigated, highlighting the importance of regularization information.


【4】 Extending Multilingual Speech Synthesis to 100+ Languages without  Transcribed Data
标题:将多语言语音合成扩展到100多种语言,无需转录数据
链接:https://arxiv.org/abs/2402.18932
作者:Takaaki Saeki,Gary Wang,Nobuyuki Morioka,Isaac Elias,Kyle Kastner,Andrew Rosenberg,Bhuvana Ramabhadran,Heiga Zen,Françoise Beaufays,Hadar Shemtov
备注:To appear in ICASSP 2024
摘要:收集高质量的录音室录音是一项挑战,这限制了文本到语音(TTS)系统的语言覆盖范围。本文提出了一个框架,可以在没有监督的情况下使用发现的数据将多语言TTS模型扩展到100多种语言。该框架将语音-文本编码器预训练与使用未转录语音和未说出的文本数据源的无监督训练相结合,从而利用大规模多语言联合语音和文本表示学习。在没有任何新语言的转录语音的情况下,该TTS模型可以用>30种看不见的语言生成可理解的语音(CER与地面真实值的差异<10%)。只需15分钟的转录,找到数据,我们就可以将可理解性差异减少到1%或更少,并在几种语言中实现与地面实况相匹配的自然度分数。
摘要:Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-text encoder pretraining with unsupervised training using untranscribed speech and unspoken text data sources, thereby leveraging massively multilingual joint speech and text representation learning. Without any transcribed speech in a new language, this TTS model can generate intelligible speech in >30 unseen languages (CER difference of <10% to ground truth). With just 15 minutes of transcribed, found data, we can reduce the intelligibility difference to 1% or less from the ground-truth, and achieve naturalness scores that match the ground-truth in several languages.

【5】 Probing the Information Encoded in Neural-based Acoustic Models of  Automatic Speech Recognition Systems
标题:基于神经网络的自动语音识别系统声学模型中信息编码的探讨
链接:https://arxiv.org/abs/2402.19443
作者:Quentin Raymondaud,Mickael Rouvier,Richard Dufour
摘要:深度学习架构在许多研究领域的性能方面取得了重大进展。因此,自动语音识别(ASR)领域受益于这些科学和技术进步,特别是声学建模,现在集成了深度神经网络架构。然而,这些性能的提高已经转化为通过这些黑盒架构学习和传达的信息的复杂性的增加。在神经网络可解释性的许多研究之后,我们在这篇文章中提出了一个协议,旨在确定哪些信息位于ASR声学模型(AM)中。要做到这一点,我们建议使用中间表示(在这里,在不同的层级别)在确定的一组任务上评估AM性能。关于性能变化和目标任务,我们可以提出关于在不同架构步骤中增强或扰动哪些信息的假设。实验在说话人确认、声环境分类、性别分类、时间失真检测系统和语音情感/情感识别系统上进行。分析表明,基于神经元的AM持有的异质信息,似乎令人惊讶地与音素识别无关,如情绪,情感或说话人身份。低层次的隐藏层在全局上看起来对信息的结构化有用,而高层次的隐藏层往往会删除对音素识别无用的信息。
摘要:Deep learning architectures have made significant progress in terms of performance in many research areas. The automatic speech recognition (ASR) field has thus benefited from these scientific and technological advances, particularly for acoustic modeling, now integrating deep neural network architectures. However, these performance gains have translated into increased complexity regarding the information learned and conveyed through these black-box architectures. Following many researches in neural networks interpretability, we propose in this article a protocol that aims to determine which and where information is located in an ASR acoustic model (AM). To do so, we propose to evaluate AM performance on a determined set of tasks using intermediate representations (here, at different layer levels). Regarding the performance variation and targeted tasks, we can emit hypothesis about which information is enhanced or perturbed at different architecture steps. Experiments are performed on both speaker verification, acoustic environment classification, gender classification, tempo-distortion detection systems and speech sentiment/emotion identification. Analysis showed that neural-based AMs hold heterogeneous information that seems surprisingly uncorrelated with phoneme recognition, such as emotion, sentiment or speaker identity. The low-level hidden layers globally appears useful for the structuring of information while the upper ones would tend to delete useless information for phoneme recognition.


【6】 Unraveling Adversarial Examples against Speaker Identification --  Techniques for Attack Detection and Victim Model Classification
标题:揭开说话人识别的敌意实例--攻击检测和受害者模型分类技术
链接:https://arxiv.org/abs/2402.19355
作者:Sonal Joshi,Thomas Thebaud,Jesús Villalba,Najim Dehak
摘要:对抗性实例已经被证明是对说话人识别系统的威胁,并且已经提出了针对它们的一些对策。在本文中,我们提出了一种方法来检测对抗性示例的存在,即,一个区分良性和敌对样本的二元分类器。我们建立和扩展以前的工作,通过探索新的体系结构的攻击类型分类。此外,我们还介绍了一种识别对抗性攻击的受害者模型的方法。为了实现这一目标,我们生成了一个新的数据集,其中包含针对各种受害者模型执行的多个攻击。我们实现了0.982的攻击检测AUC,对于未知攻击,性能下降不超过0.03。使用LightResNet34架构,我们的攻击分类准确率(不包括良性)在八种攻击类型中达到86.48%,而我们的受害者模型分类准确率在四种受害者模型中达到72.28%。
摘要:Adversarial examples have proven to threaten speaker identification systems, and several countermeasures against them have been proposed. In this paper, we propose a method to detect the presence of adversarial examples, i.e., a binary classifier distinguishing between benign and adversarial examples. We build upon and extend previous work on attack type classification by exploring new architectures. Additionally, we introduce a method for identifying the victim model on which the adversarial attack is carried out. To achieve this, we generate a new dataset containing multiple attacks performed against various victim models. We achieve an AUC of 0.982 for attack detection, with no more than a 0.03 drop in performance for unknown attacks. Our attack classification accuracy (excluding benign) reaches 86.48% across eight attack types using our LightResNet34 architecture, while our victim model classification accuracy reaches 72.28% across four victim models.

【7】 Compact Speech Translation Models via Discrete Speech Units Pretraining
标题:基于离散语音单元预训练的紧凑语音翻译模型
链接:https://arxiv.org/abs/2402.19333
作者:Tsz Kin Lam,Alexandra Birch,Barry Haddow
摘要:使用自监督学习(SSL)作为模型初始化现在很常见,以获得语音翻译(ST)的强大结果。但是,它们也会占用大量内存,阻碍设备上的部署。在本文中,我们通过在其离散语音单元(DSU)上预训练较小的模型来利用SSL模型。我们在1)滤波器组到DSU和2)DSU到翻译数据上预训练编码器-解码器模型,并将编码器从1)和解码器从2)初始化一个新模型,在有限的语音翻译数据上对其进行微调。通过使用DSU预训练来扩充SSL模型的知识,最终模型变得紧凑。与使用DSU作为模型输入相比,我们的方法有几个优点,例如更短的推理流水线和对(DSU)标记化的鲁棒性。与ASR预培训相比,它不需要成绩单,使其适用于低资源环境。在CoVoST-2 X-En上的评估表明,我们的方法比直接微调SSL模型的ST模型好> 0.5 $ BLEU,只需要一半的模型大小,并且与ASR预训练相当。
摘要:Using Self-Supervised Learning (SSL) as model initialization is now common to obtain strong results in Speech Translation (ST). However, they also impose a large memory footprint, hindering on-device deployment. In this paper, we leverage the SSL models by pretraining smaller models on their Discrete Speech Units (DSU). We pretrain encoder-decoder models on 1) Filterbank-to-DSU and 2) DSU-to-Translation data, and take the encoder from 1) and the decoder from 2) to initialise a new model, finetuning this on limited speech-translation data. The final model becomes compact by using the DSU pretraining to distil the knowledge of the SSL model. Our method has several benefits over using DSU as model inputs, such as shorter inference pipeline and robustness over (DSU) tokenization. In contrast to ASR pretraining, it does not require transcripts, making it applicable to low-resource settings. Evaluation on CoVoST-2 X-En shows that our method is >$0.5$ BLEU better than a ST model that directly finetune the SSL model, given only half the model size, and on a par with ASR pretraining.

【8】 Do End-to-End Neural Diarization Attractors Need to Encode Speaker  Characteristic Information?
标题:端到端神经二元化吸引子需要编码说话人特征信息吗?
链接:https://arxiv.org/abs/2402.19325
作者:Lin Zhang,Themos Stafylakis,Federico Landini,Mireia Diez,Anna Silnova,Lukáš Burget备注:Submitted to Odyssey 2024
摘要:在本文中,我们应用变分信息瓶颈方法的端到端的神经日记与编码器-解码器吸引子(EEND-EDA)。这使我们能够研究哪些信息对模型至关重要。EEND-EDA利用对话中的说话者的矢量表示-吸引子。我们的分析表明,吸引子不一定包含说话人特征信息。另一方面,给予吸引子更多的自由,允许它们编码一些额外的(可能是特定于说话者的)信息,导致小但一致的日志化性能改进。尽管EEND系统在架构上存在差异,但吸引子和帧嵌入的概念对大多数系统来说都是通用的,而不是EEND-EDA所特有的。我们相信,这项工作的主要结论可以适用于其他变体的EEND。因此,我们希望本文将是一个有价值的贡献,以指导社会作出更明智的决定时,设计新的系统。
摘要:In this paper, we apply the variational information bottleneck approach to end-to-end neural diarization with encoder-decoder attractors (EEND-EDA). This allows us to investigate what information is essential for the model. EEND-EDA utilizes vector representations of the speakers in a conversation - attractors. Our analysis shows that, attractors do not necessarily have to contain speaker characteristic information. On the other hand, giving the attractors more freedom allowing them to encode some extra (possibly speaker-specific) information leads to small but consistent diarization performance improvements. Despite architectural differences in EEND systems, the notion of attractors and frame embeddings is common to most of them and not specific to EEND-EDA. We believe that the main conclusions of this work can apply to other variants of EEND. Thus, we hope this paper will be a valuable contribution to guide the community to make more informed decisions when designing new systems.


【9】 Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale  Speech Recognition
标题:基于大规模语音识别的动态语音异常停顿检测
链接:https://arxiv.org/abs/2402.18923
作者:Jeehyun Lee,Yerin Choi,Tae-Jin Song,Myoung-Wan Koo备注:Accepted to ICASSP 2024
摘要:构音障碍是中风患者的常见问题,严重影响语言清晰度。不适当的停顿是严重程度评估和语言治疗的关键指标。我们建议扩展一个大规模的语音识别模型,用于构音障碍语音中的不适当停顿检测。为此,我们提出了任务设计,标签策略,和语音识别模型与不适当的停顿预测层。首先,我们将停顿检测视为语音识别,使用自动语音识别(ASR)模型将语音转换为带有停顿标签的文本。根据新设计的任务,我们在文本层面上标记停顿位置及其适当性。我们与言语语言病理学家合作,建立标签标准,确保高质量的注释数据。最后,我们扩展了ASR模型与不适当的暂停预测层的端到端不适当的暂停检测。此外,我们提出了一个任务量身定制的度量评估不适当的暂停检测独立的ASR性能。我们的实验表明,该方法更好地检测不适当的停顿在构音障碍的语音比基线。(不适当的错误率:14.47%)
摘要:Dysarthria, a common issue among stroke patients, severely impacts speech intelligibility. Inappropriate pauses are crucial indicators in severity assessment and speech-language therapy. We propose to extend a large-scale speech recognition model for inappropriate pause detection in dysarthric speech. To this end, we propose task design, labeling strategy, and a speech recognition model with an inappropriate pause prediction layer. First, we treat pause detection as speech recognition, using an automatic speech recognition (ASR) model to convert speech into text with pause tags. According to the newly designed task, we label pause locations at the text level and their appropriateness. We collaborate with speech-language pathologists to establish labeling criteria, ensuring high-quality annotated data. Finally, we extend the ASR model with an inappropriate pause prediction layer for end-to-end inappropriate pause detection. Moreover, we propose a task-tailored metric for evaluating inappropriate pause detection independent of ASR performance. Our experiments show that the proposed method better detects inappropriate pauses in dysarthric speech than baselines. (Inappropriate Pause Error Rate: 14.47%)


机器翻译由腾讯交互翻译提供,仅供参考