今日论文合集:cs.SD语音13篇,eess.AS音频处理15篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance
标题: 计算机试镜的音频增强--利用样本重要性的迭代训练范式
作者:Manuel Milling,Shuo Liu,Andreas Triantafyllopoulos,Ilhan Aslan,Björn W. Schuller
链接:点击下载PDF文件
摘要:用于音频任务(例如自动语音识别(ASR)和声学场景分类(ASC))的神经网络模型对于现实生活中的应用来说容易受到噪声污染。为了提高音频质量,在目标音频应用的前端显式地使用可以独立开发的增强模块。在本文中,我们提出了一个端到端的学习解决方案,以联合优化音频增强(AE)和后续应用的模型。为了指导AE模块对目标应用程序的优化,特别是为了克服困难的样本,我们使用样本的性能指标作为样本重要性的指示。在实验中,我们考虑了四个代表性的应用程序来评估我们的训练范式,即,ASR、语音命令识别(SCR)、语音情感识别(SER)和ASC。这些应用程序与语音和非语音任务有关的语义和非语义特征,瞬态和全局信息,实验结果表明,我们提出的方法可以大大提高噪声鲁棒性的模型,特别是在低信噪比(SNR),广泛的计算机听觉任务在日常生活中的噪声环境。摘要:Neural network models for audio tasks, such as automatic speech recognition (ASR) and acoustic scene classification (ASC), are susceptible to noise contamination for real-life applications. To improve audio quality, an enhancement module, which can be developed independently, is explicitly used at the front-end of the target audio applications. In this paper, we present an end-to-end learning solution to jointly optimise the models for audio enhancement (AE) and the subsequent applications. To guide the optimisation of the AE module towards a target application, and especially to overcome difficult samples, we make use of the sample-wise performance measure as an indication of sample importance. In experiments, we consider four representative applications to evaluate our training paradigm, i.e., ASR, speech command recognition (SCR), speech emotion recognition (SER), and ASC. These applications are associated with speech and non-speech tasks concerning semantic and non-semantic features, transient and global information, and the experimental results indicate that our proposed approach can considerably boost the noise robustness of the models, especially at low signal-to-noise ratios (SNRs), for a wide range of computer audition tasks in everyday-life noisy environments.

【2】 FLEURS-R: A Restored Multilingual Speech Corpus for Generation Tasks
标题: FLEURS-R:用于生成任务的恢复多语言语音库
作者:Min Ma,Yuma Koizumi,Shigeki Karita,Heiga Zen,Jason Riesa,Haruko Ishikawa,Michiel Bacchiani
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:本文介绍了通用语音表示的Few-Shot学习评价(FLEURS)语料库的语音恢复应用版本FLEURS-R。FLEURS-R作为FLEURS维护了102种语言的N路并行语音语料库,通过应用语音恢复模型Miipher提高了音频质量和保真度。FLEURS-R的目标是推动更多语言的语音技术,并促进包括文本到语音(TTS)和其他低资源语言语音生成任务在内的研究。与恢复的语音和TTS基线模型从新的语料库训练的综合评价表明,新的语料库获得了显着提高的语音质量,同时保持语音的语义内容。该语料库通过Hugging Face公开发布。摘要:This paper introduces FLEURS-R, a speech restoration applied version of the Few-shot Learning Evaluation of Universal Representations of Speech (FLEURS) corpus. FLEURS-R maintains an N-way parallel speech corpus in 102 languages as FLEURS, with improved audio quality and fidelity by applying the speech restoration model Miipher. The aim of FLEURS-R is to advance speech technology in more languages and catalyze research including text-to-speech (TTS) and other speech generation tasks in low-resource languages. Comprehensive evaluations with the restored speech and TTS baseline models trained from the new corpus show that the new corpus obtained significantly improved speech quality while maintaining the semantic contents of the speech. The corpus is publicly released via Hugging Face.

【3】 An Investigation Into Explainable Audio Hate Speech Detection
标题: 可解释音频仇恨语音检测研究
作者:Jinmyeong An,Wonjun Lee,Yejin Jeon,Jungseul Ok,Yunsu Kim,Gary Geunbae Lee
备注:Accepted to SIGDIAL 2024
链接:点击下载PDF文件
摘要:关于仇恨言论的研究主要围绕着从文本输入中发现和解释,而对口头内容基本上没有进行研究。虽然已经有有限的探索到仇恨言论检测在口头声学语音输入,可解释性方面被忽视。因此,我们引入了一个新的任务,可解释的音频仇恨语音检测。具体来说,我们的目标是确定精确的时间间隔,被称为音频帧级的理由,作为仇恨言论分类的证据。为此,我们提出了两种不同的方法:级联和端到端(E2 E)。级联方法首先将音频转换为成绩单,识别这些成绩单中的仇恨言论,随后定位相应的音频时间帧。相反,E2 E方法直接处理音频话语,这使得它能够在特定的时间范围内精确定位仇恨言论。此外,由于缺乏可解释的音频仇恨言论数据集,包括音频帧级的理由,我们策划了一个合成的音频数据集来训练我们的模型。我们进一步验证了这些模型的实际人类语音话语,并发现E2 E方法优于级联方法的音频帧的交集在联盟(IoU)度量。此外,我们观察到,包括帧级原理显着提高了E2 E方法的仇恨语音检测准确性。 textbf{免责声明}读者可能会遇到具有攻击性或仇恨性质的内容。但是,鉴于工作的性质,这是无法避免的。摘要:Research on hate speech has predominantly revolved around detection and interpretation from textual inputs, leaving verbal content largely unexplored. While there has been limited exploration into hate speech detection within verbal acoustic speech inputs, the aspect of interpretability has been overlooked. Therefore, we introduce a new task of explainable audio hate speech detection. Specifically, we aim to identify the precise time intervals, referred to as audio frame-level rationales, which serve as evidence for hate speech classification. Towards this end, we propose two different approaches: cascading and End-to-End (E2E). The cascading approach initially converts audio to transcripts, identifies hate speech within these transcripts, and subsequently locates the corresponding audio time frames. Conversely, the E2E approach processes audio utterances directly, which allows it to pinpoint hate speech within specific time frames. Additionally, due to the lack of explainable audio hate speech datasets that include audio frame-level rationales, we curated a synthetic audio dataset to train our models. We further validated these models on actual human speech utterances and found that the E2E approach outperforms the cascading method in terms of the audio frame Intersection over Union (IoU) metric. Furthermore, we observed that including frame-level rationales significantly enhances hate speech detection accuracy for the E2E approach. textbf{Disclaimer} The reader may encounter content of an offensive or hateful nature. However, given the nature of the work, this cannot be avoided.

【4】 PyNeuralFx: A Python Package for Neural Audio Effect Modeling
标题: PyNeuralDx:用于神经音效建模的Python包
作者:Yen-Tung Yeh,Wen-Yi Hsiao,Yi-Hsuan Yang
备注:toolkit paper
链接:点击下载PDF文件
摘要:我们介绍了PyNeuralFx,这是一个开源Python工具包,专为神经音频效果建模研究而设计。该工具包提供了一个直观的框架,并提供了一套全面的功能,包括完善的模型架构,损失函数和易于使用的可视化工具的标准化实现。因此,它有助于提高神经音频效果建模研究的可重复性,并能够深入比较不同模型的性能,通过DSP方法深入了解模型的行为和操作特性。该工具包可在https: github.com ytsrt66589 pyneuralfx上找到。摘要:We present PyNeuralFx, an open-source Python toolkit designed for research on neural audio effect modeling. The toolkit provides an intuitive framework and offers a comprehensive suite of features, including standardized implementation of well-established model architectures, loss functions, and easy-to-use visualization tools. As such, it helps promote reproducibility for research on neural audio effect modeling, and enable in-depth performance comparison of different models, offering insight into the behavior and operational characteristics of models through DSP methodology. The toolkit can be found at https: github.com ytsrt66589 pyneuralfx.

【5】 Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
标题: 通过噪音表示学习增强具有强大上下文感知的对话语音识别
作者:Wonjun Lee,San Kim,Gary Geunbae Lee
备注:11 pages, 2 figures, Accepted to SIGDIAL2024
链接:点击下载PDF文件
摘要:最近的对话系统依赖于基于话轮的口语交互,需要准确的自动语音识别(ASR)。ASR中的错误会严重影响下游对话任务。为了解决这个问题,已经提出了使用来自用户和代理交互的对话上下文来转录随后的话语。该方法结合了用户的语音和代理的响应作为模型输入的转录,使用由每个回合生成的累积上下文。然而,该上下文容易受到ASR错误的影响,因为它是由ASR模型以自回归方式生成的。这样的噪声上下文可以进一步降低上下文输入的益处,导致次优的ASR性能。在本文中,我们引入上下文噪声表示学习(CNRL),以增强对噪声背景的鲁棒性,最终提高对话语音识别的准确性。为了最大限度地提高上下文感知的优势,我们的方法包括使用基于文本的对话数据和上下文编码器的噪声表示学习的解码器预训练。基于对语音对话的评估,与基线相比,我们的方法显示出更好的结果。此外,我们的方法的优势是突出在嘈杂的环境中,用户语音几乎听不见,由于现实世界的噪音,依赖于上下文信息准确地转录输入。摘要:Recent dialogue systems rely on turn-based spoken interactions, requiring accurate Automatic Speech Recognition (ASR). Errors in ASR can significantly impact downstream dialogue tasks. To address this, using dialogue context from user and agent interactions for transcribing subsequent utterances has been proposed. This method incorporates the transcription of the user's speech and the agent's response as model input, using the accumulated context generated by each turn. However, this context is susceptible to ASR errors because it is generated by the ASR model in an auto-regressive fashion. Such noisy context can further degrade the benefits of context input, resulting in suboptimal ASR performance. In this paper, we introduce Context Noise Representation Learning (CNRL) to enhance robustness against noisy context, ultimately improving dialogue speech recognition accuracy. To maximize the advantage of context awareness, our approach includes decoder pre-training using text-based dialogue data and noise representation learning for a context encoder. Based on the evaluation of speech dialogues, our method shows superior results compared to baselines. Furthermore, the strength of our approach is highlighted in noisy environments where user speech is barely audible due to real-world noise, relying on contextual information to transcribe the input accurately.

【6】 Controlling Surprisal in Music Generation via Information Content Curve Matching
标题: 通过信息内容曲线匹配控制音乐生成中的惊喜
作者:Mathias Rose Bjare,Stefan Lattner,Gerhard Widmer
备注:8 pages, 4 figures, 2 tables, accepted at the 25th Int. Society for Music Information Retrieval Conf., San Francisco, USA, 2024
链接:点击下载PDF文件
摘要:近年来,音乐生成系统的质量和公众的兴趣已经增长,鼓励研究各种方法来控制这些系统。我们提出了一种新的方法,用于控制在音乐生成中使用序列模型的节拍。为了实现这一目标,我们定义了一个度量称为瞬时信息内容(IIC)。IIC用作感知音乐音阶的代理函数(如从概率模型估计的),并且可以在音乐作品内的任何点处计算。这使得即使音乐事件以不规则的时间间隔发生,也能够跨不同的音乐内容比较重复。我们使用波束搜索来生成音乐素材,其IIC曲线非常接近给定的目标IIC。我们的实验表明,IIC与谐波和节奏的复杂性和注意到密度。相关性随着用于估计IIC的音乐背景的长度而减小。最后,我们进行了一个定性的用户研究,以测试人类听众是否可以识别的IIC曲线,已被用作目标时,产生相应的音乐材料。我们在https: github.com muthissar iic上提供了创建IIC插值和IIC可视化的代码。摘要:In recent years, the quality and public interest in music generation systems have grown, encouraging research into various ways to control these systems. We propose a novel method for controlling surprisal in music generation using sequence models. To achieve this goal, we define a metric called Instantaneous Information Content (IIC). The IIC serves as a proxy function for the perceived musical surprisal (as estimated from a probabilistic model) and can be calculated at any point within a music piece. This enables the comparison of surprisal across different musical content even if the musical events occur in irregular time intervals. We use beam search to generate musical material whose IIC curve closely approximates a given target IIC. We experimentally show that the IIC correlates with harmonic and rhythmic complexity and note density. The correlation decreases with the length of the musical context used for estimating the IIC. Finally, we conduct a qualitative user study to test if human listeners can identify the IIC curves that have been used as targets when generating the respective musical material. We provide code for creating IIC interpolations and IIC visualizations on https: github.com muthissar iic.

【7】 Robust online reconstruction of continuous-time signals from a lean spike train ensemble code
标题: 从精益尖峰序列集合代码中稳健地在线重建连续时间信号
作者:Anik Chattopadhyay,Arunava Banerjee
备注:22 pages, including a 9-page appendix, 8 figures. A GitHub link to the project implementation is embedded in the paper
链接:点击下载PDF文件
摘要:动物的感觉刺激由神经元编码成尖峰序列,提供诸如稀疏性、能量效率和高时间分辨率的优点。本文提出了一个信号处理框架,确定性编码连续时间信号到生物学上可行的尖峰列车,并解决了可表示的信号类和重建边界的问题。该框架考虑通过由神经元的集合使用具有各种卷积核的卷积然后阈值机制生成的尖峰序列对信号进行编码。一个封闭形式的解决方案的逆问题,从尖峰列车信号重建,来自移位核函数的希尔伯特空间,确保稀疏表示的广义有限创新率(FRI)类信号。此外,在生物系统中的实时处理的启发,制定一个有效的迭代版本的最佳重建,只考虑一个有限的窗口,过去的尖峰,确保鲁棒性的技术病态编码;然后提供的窗口重建的最佳解决方案的收敛保证。在一个大型音频数据集上的实验表明,在低至奈奎斯特速率的五分之一的尖峰速率下具有出色的重建精度,同时与低尖峰速率制度中的最先进的稀疏编码技术相比,显示出明显的竞争优势。摘要:Sensory stimuli in animals are encoded into spike trains by neurons, offering advantages such as sparsity, energy efficiency, and high temporal resolution. This paper presents a signal processing framework that deterministically encodes continuous-time signals into biologically feasible spike trains, and addresses the questions about representable signal classes and reconstruction bounds. The framework considers encoding of a signal through spike trains generated by an ensemble of neurons using a convolve-then-threshold mechanism with various convolution kernels. A closed-form solution to the inverse problem, from spike trains to signal reconstruction, is derived in the Hilbert space of shifted kernel functions, ensuring sparse representation of a generalized Finite Rate of Innovation (FRI) class of signals. Additionally, inspired by real-time processing in biological systems, an efficient iterative version of the optimal reconstruction is formulated that considers only a finite window of past spikes, ensuring robustness of the technique to ill-conditioned encoding; convergence guarantees of the windowed reconstruction to the optimal solution are then provided. Experiments on a large audio dataset demonstrate excellent reconstruction accuracy at spike rates as low as one-fifth of the Nyquist rate, while showing clear competitive advantage in comparison to state-of-the-art sparse coding techniques in the low spike rate regime.

【8】 Adapting General Disentanglement-Based Speaker Anonymization for Enhanced Emotion Preservation
标题: 调整基于通用解开的说话人解析以增强情感保存
作者:Xiaoxiao Miao,Yuxiang Zhang,Xin Wang,Natalia Tomashenko,Donny Cheng Lock Soh,Ian Mcloughlin
链接:点击下载PDF文件
摘要:一般的基于解纠缠的说话者匿名化系统通常使用单独的编码器将语音分离成内容、说话者和韵律特征。本文探讨了如何适应这样的系统时,一个新的语音属性,例如,情感,需要在更大程度上保留。虽然现有的系统擅长匿名说话者嵌入,但它们并不是为了保留情感而设计的。两种策略,这是检查。首先,我们表明,从预先训练的情感编码器中集成情感嵌入可以帮助保留情感线索,即使这种方法稍微损害了隐私保护。或者,我们提出了一个情感补偿策略作为后处理步骤应用于匿名扬声器嵌入。这隐藏了原始说话者的身份,并重新引入了在说话者嵌入匿名化过程中丢失的情感特征。具体来说,我们使用支持向量机对情感属性进行建模,以学习每种情感的单独边界。在推理过程中,通过两种方式对原始说话人信息进行处理:一是通过情感指示器预测情感,准确选择情感匹配的SVM;二是通过说话人匿名器隐藏说话人特征。然后沿着相应的SVM边界朝着增强的情感方向修改匿名说话人嵌入,以保存情感线索。所提出的策略也预计将是有用的适应一般的基于解纠缠的扬声器匿名化系统,以保持其他目标语言属性,具有潜在的一系列下游任务。摘要:A general disentanglement-based speaker anonymization system typically separates speech into content, speaker, and prosody features using individual encoders. This paper explores how to adapt such a system when a new speech attribute, for example, emotion, needs to be preserved to a greater extent. While existing systems are good at anonymizing speaker embeddings, they are not designed to preserve emotion. Two strategies for this are examined. First, we show that integrating emotion embeddings from a pre-trained emotion encoder can help preserve emotional cues, even though this approach slightly compromises privacy protection. Alternatively, we propose an emotion compensation strategy as a post-processing step applied to anonymized speaker embeddings. This conceals the original speaker's identity and reintroduces the emotional traits lost during speaker embedding anonymization. Specifically, we model the emotion attribute using support vector machines to learn separate boundaries for each emotion. During inference, the original speaker embedding is processed in two ways: one, by an emotion indicator to predict emotion and select the emotion-matched SVM accurately; and two, by a speaker anonymizer to conceal speaker characteristics. The anonymized speaker embedding is then modified along the corresponding SVM boundary towards an enhanced emotional direction to save the emotional cues. The proposed strategies are also expected to be useful for adapting a general disentanglement-based speaker anonymization system to preserve other target paralinguistic attributes, with potential for a range of downstream tasks.

【9】 LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
标题: LI-TTA:自动语音识别的语言知情测试时间自适应
作者:Eunseop Yoon,Hee Suk Yoon,John Harvill,Mark Hasegawa-Johnson,Chang D. Yoo
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:测试时自适应(TTA)已经成为域转移挑战的关键解决方案,其中目标环境偏离原始训练环境。一个主要的例子是自动语音识别(ASR)的TTA,它通过利用输出预测熵最小化作为自我监督信号来增强模型性能。然而,这种自我监督的一个关键限制在于它主要关注声学特征,而很少关注输入的语言特性。为了解决这一差距,我们提出了语言知情的测试时间适应(LI-TTA),它结合了语言的见解,在TTA的ASR。LI-TTA集成了来自外部语言模型的校正,通过最小化校正的CTC损失以及标准TTA损失来合并语言与声学信息。通过大量的实验,我们表明,LI-TTA有效地提高了性能的TTA ASR在各种分布偏移的情况下。摘要:Test-Time Adaptation (TTA) has emerged as a crucial solution to the domain shift challenge, wherein the target environment diverges from the original training environment. A prime exemplification is TTA for Automatic Speech Recognition (ASR), which enhances model performance by leveraging output prediction entropy minimization as a self-supervision signal. However, a key limitation of this self-supervision lies in its primary focus on acoustic features, with minimal attention to the linguistic properties of the input. To address this gap, we propose Language Informed Test-Time Adaptation (LI-TTA), which incorporates linguistic insights during TTA for ASR. LI-TTA integrates corrections from an external language model to merge linguistic with acoustic information by minimizing the CTC loss from the correction alongside the standard TTA loss. With extensive experiments, we show that LI-TTA effectively improves the performance of TTA for ASR in various distribution shift situations.

【10】 Stream-based Active Learning for Anomalous Sound Detection in Machine Condition Monitoring
标题: 基于流的主动学习用于机器状态监测中异常声音检测
作者:Tuan Vu Ho,Kota Dohi,Yohei Kawaguchi
备注:Accepted as a conference paper in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:介绍了一种用于机器状态监测系统中异常声音检测的主动学习框架。通常,由于异常数据的稀缺性,ASD模型仅在正常样本上进行训练,导致在推理过程中不可见样本的准确性降低。AL是解决这个问题的一个很有前途的解决方案,它使模型能够用更少的标记示例更有效地学习新概念,从而减少手动注释工作。然而,其在ASD中的有效性仍未被探索。为了最小化更新成本和时间,我们提出的方法侧重于更新ASD系统的评分后端,而无需重新训练神经网络模型。在DCASE 2023挑战任务2数据集上的实验结果证实,即使在低标签预算的情况下,我们的AL框架也显着提高了ASD性能。此外,我们提出的抽样策略优于其他基线的部分区域下的受试者操作特征得分。摘要:This paper introduces an active learning (AL) framework for anomalous sound detection (ASD) in machine condition monitoring system. Typically, ASD models are trained solely on normal samples due to the scarcity of anomalous data, leading to decreased accuracy for unseen samples during inference. AL is a promising solution to solve this problem by enabling the model to learn new concepts more effectively with fewer labeled examples, thus reducing manual annotation efforts. However, its effectiveness in ASD remains unexplored. To minimize update costs and time, our proposed method focuses on updating the scoring backend of ASD system without retraining the neural network model. Experimental results on the DCASE 2023 Challenge Task 2 dataset confirm that our AL framework significantly improves ASD performance even with low labeling budgets. Moreover, our proposed sampling strategy outperforms other baselines in terms of the partial area under the receiver operating characteristic score.

【11】 VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
标题: VQ-CTAP:语音处理的跨模式细粒度序列表示学习
作者:Chunyu Qiang,Wang Geng,Yi Zhao,Ruibo Fu,Tao Wang,Cheng Gong,Tianrui Wang,Qiuyu Liu,Jiangyan Yi,Zhengqi Wen,Chen Zhang,Hao Che,Longbiao Wang,Jianwu Dang,Jianhua Tao
链接:点击下载PDF文件
摘要:深度学习为跨模态表示学习领域带来了重大改进。对于诸如文本到语音(TTS)、语音转换(VC)和自动语音识别(ASR)的任务,需要跨模态细粒度(帧级)序列表示,强调文本模态的语义内容,同时不强调语音模态的非语言信息。我们提出了一种称为“矢量量化对比令牌声学预训练(VQ-CTAP)”的方法,该方法使用跨模态对齐序列转码器将文本和语音带入联合多模态空间,学习如何在帧级连接文本和语音。建议的VQ-CTAP是一个跨模态序列表示学习的范例,为语音处理中的细粒度生成和识别任务提供了一个有前途的解决方案。VQ-CTAP可以直接应用于VC和ASR任务,无需微调或额外的结构。我们提出了一个序列感知的语义连接器,它连接多个冻结的预训练模块的TTS任务,表现出即插即用的能力。我们设计了一个步进优化策略,通过逐步注入和调整各损耗分量的影响,以确保有效的模型收敛。此外,我们提出了一个语义转移明智的非语言一致性损失,以提高代表性的能力,使模型能够更好地推广到看不见的数据和捕捉非语言信息的细微差别。此外,VQ-CTAP实现了从24 kHz输入波形到25 Hz速率的高压缩语音编码,这是采样速率的960倍降低。音频演示可在https: qiangchunyu.github.io VQCTAP 上获得摘要:Deep learning has brought significant improvements to the field of cross-modal representation learning. For tasks such as text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), a cross-modal fine-grained (frame-level) sequence representation is desired, emphasizing the semantic content of the text modality while de-emphasizing the paralinguistic information of the speech modality. We propose a method called "Vector Quantized Contrastive Token-Acoustic Pre-training (VQ-CTAP)", which uses the cross-modal aligned sequence transcoder to bring text and speech into a joint multimodal space, learning how to connect text and speech at the frame level. The proposed VQ-CTAP is a paradigm for cross-modal sequence representation learning, offering a promising solution for fine-grained generation and recognition tasks in speech processing. The VQ-CTAP can be directly applied to VC and ASR tasks without fine-tuning or additional structures. We propose a sequence-aware semantic connector, which connects multiple frozen pre-trained modules for the TTS task, exhibiting a plug-and-play capability. We design a stepping optimization strategy to ensure effective model convergence by gradually injecting and adjusting the influence of various loss components. Furthermore, we propose a semantic-transfer-wise paralinguistic consistency loss to enhance representational capabilities, allowing the model to better generalize to unseen data and capture the nuances of paralinguistic information. In addition, VQ-CTAP achieves high-compression speech coding at a rate of 25Hz from 24kHz input waveforms, which is a 960-fold reduction in the sampling rate. The audio demo is available at https: qiangchunyu.github.io VQCTAP

【12】 Extracting Urban Sound Information for Residential Areas in Smart Cities Using an End-to-End IoT System
标题: 使用端到端物联网系统提取智慧城市住宅区的城市声音信息
作者:Ee-Leng Tan,Furi Andi Karnapi,Linus Junjia Ng,Kenneth Ooi,Woon-Seng Gan
Journal-ref:IEEE IoT Journal, 2021
链接:点击下载PDF文件
摘要:随着城市化的快速发展,居住区的社区,建筑和交通噪音增加。单靠声压级数据来决定噪音环境和制订噪音管制及缓解策略的传统方法并不足够。本文介绍了一种端到端物联网系统,该系统使用边缘设备提取实时城市声音元数据,提供有关9个居民区中主要噪声的声音类型、位置和持续时间、发生率、响度和方位角的信息。收集到的环境声音元数据被传输到基于云的平台并在其中进行汇总,以生成详细的描述性分析和可视化。概述了我们集成不同构建模块(即硬件、软件、云技术和信号处理算法)以形成实时物联网系统的方法。我们演示了如何使用我们的系统提取的一些声音元数据来提供对住宅区噪声的见解。讨论了一个可扩展的工作流程,从九个居民区收集和准备音频记录,以构建我们的城市声音数据集,用于训练和评估与位置无关的模型。管理和维护部署在许多位置的传感器网络的一些实际挑战也得到了解决。摘要:With rapid urbanization comes the increase of community, construction, and transportation noise in residential areas. The conventional approach of solely relying on sound pressure level (SPL) information to decide on the noise environment and to plan out noise control and mitigation strategies is inadequate. This paper presents an end-to-end IoT system that extracts real-time urban sound metadata using edge devices, providing information on the sound type, location and duration, rate of occurrence, loudness, and azimuth of a dominant noise in nine residential areas. The collected metadata on environmental sound is transmitted to and aggregated in a cloud-based platform to produce detailed descriptive analytics and visualization. Our approach to integrating different building blocks, namely, hardware, software, cloud technologies, and signal processing algorithms to form our real-time IoT system is outlined. We demonstrate how some of the sound metadata extracted by our system are used to provide insights into the noise in residential areas. A scalable workflow to collect and prepare audio recordings from nine residential areas to construct our urban sound dataset for training and evaluating a location-agnostic model is discussed. Some practical challenges of managing and maintaining a sensor network deployed at numerous locations are also addressed.

【13】 Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
标题: 利用不配对的语音和文本提高Whisper对代表性不足语言哈萨克语的识别性能
作者:Jinpeng Li,Yu Pu,Qi Sun,Wei-Qiang Zhang
备注:Accepted by INTERSPEECH 2024;Minor typo correction
链接:点击下载PDF文件
摘要:Whisper等大规模自动语音识别模型在性能上取得了显著进步。然而,它们在许多低资源语言(如哈萨克语)上的表现并不令人满意。如何利用低成本的数据来提高Whisper在低表示语言上的性能是值得研究的问题。在这项研究中,我们利用容易获得的不成对的语音和文本数据,并结合语言模型GPT与哈萨克语的Whisper。为了提高语音识别的性能,我们实现了文本结尾(EOT)判断修改和幻觉惩罚。此外,我们采用解码平均令牌对数概率作为标准来选择样本从未标记的语音数据和使用伪标记的数据来微调模型,以进一步提高其性能。最终,我们在多个实验中实现了超过10%的绝对WER减少,并且整个过程有可能推广到其他代表性不足的语言。摘要:Whisper and other large-scale automatic speech recognition models have made significant progress in performance. However, their performance on many low-resource languages, such as Kazakh, is not satisfactory. It is worth researching how to utilize low-cost data to improve the performance of Whisper on under-represented languages. In this study, we utilized easily accessible unpaired speech and text data and combined the language model GPT with Whisper on Kazakh. We implemented end of transcript (EOT) judgment modification and hallucination penalty to improve the performance of speech recognition. Further, we employed the decoding average token log probability as a criterion to select samples from unlabeled speech data and used pseudo-labeled data to fine-tune the model to further improve its performance. Ultimately, we achieved more than 10 % absolute WER reduction in multiple experiments, and the whole process has the potential to be generalized to other under-represented languages.


eess.AS音频处理
【1】 VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
标题: VQ-CTAP:语音处理的跨模式细粒度序列表示学习
作者:Chunyu Qiang,Wang Geng,Yi Zhao,Ruibo Fu,Tao Wang,Cheng Gong,Tianrui Wang,Qiuyu Liu,Jiangyan Yi,Zhengqi Wen,Chen Zhang,Hao Che,Longbiao Wang,Jianwu Dang,Jianhua Tao
链接:点击下载PDF文件
摘要:深度学习为跨模态表示学习领域带来了重大改进。对于诸如文本到语音(TTS)、语音转换(VC)和自动语音识别(ASR)的任务,需要跨模态细粒度(帧级)序列表示,强调文本模态的语义内容,同时不强调语音模态的非语言信息。我们提出了一种称为“矢量量化对比令牌声学预训练(VQ-CTAP)”的方法,该方法使用跨模态对齐序列转码器将文本和语音带入联合多模态空间,学习如何在帧级连接文本和语音。建议的VQ-CTAP是一个跨模态序列表示学习的范例,为语音处理中的细粒度生成和识别任务提供了一个有前途的解决方案。VQ-CTAP可以直接应用于VC和ASR任务,无需微调或额外的结构。我们提出了一个序列感知的语义连接器,它连接多个冻结的预训练模块的TTS任务,表现出即插即用的能力。我们设计了一个步进优化策略,通过逐步注入和调整各损耗分量的影响,以确保有效的模型收敛。此外,我们提出了一个语义转移明智的非语言一致性损失,以提高代表性的能力,使模型能够更好地推广到看不见的数据和捕捉非语言信息的细微差别。此外,VQ-CTAP实现了从24 kHz输入波形到25 Hz速率的高压缩语音编码,这是采样速率的960倍降低。音频演示可在https: qiangchunyu.github.io VQCTAP 上获得摘要:Deep learning has brought significant improvements to the field of cross-modal representation learning. For tasks such as text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), a cross-modal fine-grained (frame-level) sequence representation is desired, emphasizing the semantic content of the text modality while de-emphasizing the paralinguistic information of the speech modality. We propose a method called "Vector Quantized Contrastive Token-Acoustic Pre-training (VQ-CTAP)", which uses the cross-modal aligned sequence transcoder to bring text and speech into a joint multimodal space, learning how to connect text and speech at the frame level. The proposed VQ-CTAP is a paradigm for cross-modal sequence representation learning, offering a promising solution for fine-grained generation and recognition tasks in speech processing. The VQ-CTAP can be directly applied to VC and ASR tasks without fine-tuning or additional structures. We propose a sequence-aware semantic connector, which connects multiple frozen pre-trained modules for the TTS task, exhibiting a plug-and-play capability. We design a stepping optimization strategy to ensure effective model convergence by gradually injecting and adjusting the influence of various loss components. Furthermore, we propose a semantic-transfer-wise paralinguistic consistency loss to enhance representational capabilities, allowing the model to better generalize to unseen data and capture the nuances of paralinguistic information. In addition, VQ-CTAP achieves high-compression speech coding at a rate of 25Hz from 24kHz input waveforms, which is a 960-fold reduction in the sampling rate. The audio demo is available at https: qiangchunyu.github.io VQCTAP

【2】 Extracting Urban Sound Information for Residential Areas in Smart Cities Using an End-to-End IoT System
标题: 使用端到端物联网系统提取智慧城市住宅区的城市声音信息
作者:Ee-Leng Tan,Furi Andi Karnapi,Linus Junjia Ng,Kenneth Ooi,Woon-Seng Gan
Journal-ref:IEEE IoT Journal, 2021
链接:点击下载PDF文件
摘要:随着城市化的快速发展,居住区的社区,建筑和交通噪音增加。单靠声压级数据来决定噪音环境和制订噪音管制及缓解策略的传统方法并不足够。本文介绍了一种端到端物联网系统,该系统使用边缘设备提取实时城市声音元数据,提供有关9个居民区中主要噪声的声音类型、位置和持续时间、发生率、响度和方位角的信息。收集到的环境声音元数据被传输到基于云的平台并在其中进行汇总,以生成详细的描述性分析和可视化。概述了我们集成不同构建模块(即硬件、软件、云技术和信号处理算法)以形成实时物联网系统的方法。我们演示了如何使用我们的系统提取的一些声音元数据来提供对住宅区噪声的见解。讨论了一个可扩展的工作流程,从九个居民区收集和准备音频记录,以构建我们的城市声音数据集,用于训练和评估与位置无关的模型。管理和维护部署在许多位置的传感器网络的一些实际挑战也得到了解决。摘要:With rapid urbanization comes the increase of community, construction, and transportation noise in residential areas. The conventional approach of solely relying on sound pressure level (SPL) information to decide on the noise environment and to plan out noise control and mitigation strategies is inadequate. This paper presents an end-to-end IoT system that extracts real-time urban sound metadata using edge devices, providing information on the sound type, location and duration, rate of occurrence, loudness, and azimuth of a dominant noise in nine residential areas. The collected metadata on environmental sound is transmitted to and aggregated in a cloud-based platform to produce detailed descriptive analytics and visualization. Our approach to integrating different building blocks, namely, hardware, software, cloud technologies, and signal processing algorithms to form our real-time IoT system is outlined. We demonstrate how some of the sound metadata extracted by our system are used to provide insights into the noise in residential areas. A scalable workflow to collect and prepare audio recordings from nine residential areas to construct our urban sound dataset for training and evaluating a location-agnostic model is discussed. Some practical challenges of managing and maintaining a sensor network deployed at numerous locations are also addressed.

【3】 Towards a Quantitative Analysis of Coarticulation with a Phoneme-to-Articulatory Model
标题: 用音素到关节语模型对协同发音进行定量分析
作者:Chaofei Fan,Jaimie M. Henderson,Chris Manning,Francis R. Willett
备注:To be published in Interspeech 2024
链接:点击下载PDF文件
摘要:以往的协同构音研究主要集中在有限的音位序列和特定的构音器上,仅提供了对协同构音的时间范围和大小的近似描述。本文是对协同发音进行全面研究的初步尝试。我们利用现有的电磁关节成像(EMA)数据集开发和训练音素发音(P2A)模型,可以生成逼真的EMA为新的音素序列和复制已知的协同发音模式。我们使用模型生成的EMA对9K最小的单词对,以分析协同发音的幅度和程度高达8个音素的协同发音触发,并比较不同的辅音协同发音阻力。我们的研究结果与早期的研究一致,并提出了比以前发现的更长范围的协同发音效应。这种基于模型的方法可以潜在地比较成人和儿童之间以及跨语言的协同发音,为语音产生提供新的见解。摘要:Prior coarticulation studies focus mainly on limited phonemic sequences and specific articulators, providing only approximate descriptions of the temporal extent and magnitude of coarticulation. This paper is an initial attempt to comprehensively investigate coarticulation. We leverage existing Electromagnetic Articulography (EMA) datasets to develop and train a phoneme-to-articulatory (P2A) model that can generate realistic EMA for novel phoneme sequences and replicate known coarticulation patterns. We use model-generated EMA on 9K minimal word pairs to analyze coarticulation magnitude and extent up to eight phonemes from the coarticulation trigger, and compare coarticulation resistance across different consonants. Our findings align with earlier studies and suggest a longer-range coarticulation effect than previously found. This model-based approach can potentially compare coarticulation between adults and children and across languages, offering new insights into speech production.

【4】 Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
标题: 利用不配对的语音和文本提高Whisper对代表性不足语言哈萨克语的识别性能
作者:Jinpeng Li,Yu Pu,Qi Sun,Wei-Qiang Zhang
备注:Accepted by INTERSPEECH 2024;Minor typo correction
链接:点击下载PDF文件
摘要:Whisper等大规模自动语音识别模型在性能上取得了显著进步。然而,它们在许多低资源语言(如哈萨克语)上的表现并不令人满意。如何利用低成本的数据来提高Whisper在低表示语言上的性能是值得研究的问题。在这项研究中,我们利用容易获得的不成对的语音和文本数据,并结合语言模型GPT与哈萨克语的Whisper。为了提高语音识别的性能,我们实现了文本结尾(EOT)判断修改和幻觉惩罚。此外,我们采用解码平均令牌对数概率作为标准来选择样本从未标记的语音数据和使用伪标记的数据来微调模型,以进一步提高其性能。最终,我们在多个实验中实现了超过10%的绝对WER减少,整个过程有可能推广到其他代表性不足的语言。摘要:Whisper and other large-scale automatic speech recognition models have made significant progress in performance. However, their performance on many low-resource languages, such as Kazakh, is not satisfactory. It is worth researching how to utilize low-cost data to improve the performance of Whisper on under-represented languages. In this study, we utilized easily accessible unpaired speech and text data and combined the language model GPT with Whisper on Kazakh. We implemented end of transcript (EOT) judgment modification and hallucination penalty to improve the performance of speech recognition. Further, we employed the decoding average token log probability as a criterion to select samples from unlabeled speech data and used pseudo-labeled data to fine-tune the model to further improve its performance. Ultimately, we achieved more than 10 % absolute WER reduction in multiple experiments, and the whole process has the potential to be generalized to other under-represented languages.

【5】 Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance
标题: 计算机试镜的音频增强--利用样本重要性的迭代训练范式
作者:Manuel Milling,Shuo Liu,Andreas Triantafyllopoulos,Ilhan Aslan,Björn W. Schuller
链接:点击下载PDF文件
摘要:用于音频任务(例如自动语音识别(ASR)和声学场景分类(ASC))的神经网络模型对于现实生活中的应用来说容易受到噪声污染。为了提高音频质量,在目标音频应用的前端显式地使用可以独立开发的增强模块。在本文中,我们提出了一个端到端的学习解决方案,以联合优化音频增强(AE)和后续应用的模型。为了指导AE模块对目标应用程序的优化,特别是为了克服困难的样本,我们使用样本的性能指标作为样本重要性的指示。在实验中,我们考虑了四个代表性的应用程序来评估我们的训练范式,即,ASR、语音命令识别(SCR)、语音情感识别(SER)和ASC。这些应用程序与语音和非语音任务有关的语义和非语义特征,瞬态和全局信息,实验结果表明,我们提出的方法可以大大提高噪声鲁棒性的模型,特别是在低信噪比(SNR),广泛的计算机听觉任务在日常生活中的噪声环境。摘要:Neural network models for audio tasks, such as automatic speech recognition (ASR) and acoustic scene classification (ASC), are susceptible to noise contamination for real-life applications. To improve audio quality, an enhancement module, which can be developed independently, is explicitly used at the front-end of the target audio applications. In this paper, we present an end-to-end learning solution to jointly optimise the models for audio enhancement (AE) and the subsequent applications. To guide the optimisation of the AE module towards a target application, and especially to overcome difficult samples, we make use of the sample-wise performance measure as an indication of sample importance. In experiments, we consider four representative applications to evaluate our training paradigm, i.e., ASR, speech command recognition (SCR), speech emotion recognition (SER), and ASC. These applications are associated with speech and non-speech tasks concerning semantic and non-semantic features, transient and global information, and the experimental results indicate that our proposed approach can considerably boost the noise robustness of the models, especially at low signal-to-noise ratios (SNRs), for a wide range of computer audition tasks in everyday-life noisy environments.

【6】 FLEURS-R: A Restored Multilingual Speech Corpus for Generation Tasks
标题: FLEURS-R:用于生成任务的恢复多语言语音库
作者:Min Ma,Yuma Koizumi,Shigeki Karita,Heiga Zen,Jason Riesa,Haruko Ishikawa,Michiel Bacchiani
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:本文介绍了通用语音表示的Few-Shot学习评价(FLEURS)语料库的语音恢复应用版本FLEURS-R。FLEURS-R作为FLEURS维护了102种语言的N路并行语音语料库,通过应用语音恢复模型Miipher提高了音频质量和保真度。FLEURS-R的目标是推动更多语言的语音技术,并促进包括文本到语音(TTS)和其他低资源语言语音生成任务在内的研究。与恢复的语音和TTS基线模型从新的语料库训练的综合评价表明,新的语料库获得了显着提高的语音质量,同时保持语音的语义内容。该语料库通过Hugging Face公开发布。摘要:This paper introduces FLEURS-R, a speech restoration applied version of the Few-shot Learning Evaluation of Universal Representations of Speech (FLEURS) corpus. FLEURS-R maintains an N-way parallel speech corpus in 102 languages as FLEURS, with improved audio quality and fidelity by applying the speech restoration model Miipher. The aim of FLEURS-R is to advance speech technology in more languages and catalyze research including text-to-speech (TTS) and other speech generation tasks in low-resource languages. Comprehensive evaluations with the restored speech and TTS baseline models trained from the new corpus show that the new corpus obtained significantly improved speech quality while maintaining the semantic contents of the speech. The corpus is publicly released via Hugging Face.

【7】 An Investigation Into Explainable Audio Hate Speech Detection
标题: 可解释音频仇恨语音检测研究
作者:Jinmyeong An,Wonjun Lee,Yejin Jeon,Jungseul Ok,Yunsu Kim,Gary Geunbae Lee
备注:Accepted to SIGDIAL 2024
链接:点击下载PDF文件
摘要:关于仇恨言论的研究主要围绕着从文本输入中发现和解释,而对口头内容基本上没有进行研究。虽然已经有有限的探索到仇恨言论检测在口头声学语音输入,可解释性方面被忽视。因此,我们引入了一个新的任务,可解释的音频仇恨语音检测。具体来说,我们的目标是确定精确的时间间隔,被称为音频帧级的理由,作为仇恨言论分类的证据。为此,我们提出了两种不同的方法:级联和端到端(E2 E)。级联方法首先将音频转换为成绩单,识别这些成绩单中的仇恨言论,随后定位相应的音频时间帧。相反,E2 E方法直接处理音频话语,这使得它能够在特定的时间范围内精确定位仇恨言论。此外,由于缺乏可解释的音频仇恨言论数据集,包括音频帧级的理由,我们策划了一个合成的音频数据集来训练我们的模型。我们进一步验证了这些模型的实际人类语音话语,并发现E2 E方法优于级联方法的音频帧的交集在联盟(IoU)度量。此外,我们观察到,包括帧级原理显着提高了E2 E方法的仇恨语音检测准确性。 textbf{免责声明}读者可能会遇到具有攻击性或仇恨性质的内容。但是,鉴于工作的性质,这是无法避免的。摘要:Research on hate speech has predominantly revolved around detection and interpretation from textual inputs, leaving verbal content largely unexplored. While there has been limited exploration into hate speech detection within verbal acoustic speech inputs, the aspect of interpretability has been overlooked. Therefore, we introduce a new task of explainable audio hate speech detection. Specifically, we aim to identify the precise time intervals, referred to as audio frame-level rationales, which serve as evidence for hate speech classification. Towards this end, we propose two different approaches: cascading and End-to-End (E2E). The cascading approach initially converts audio to transcripts, identifies hate speech within these transcripts, and subsequently locates the corresponding audio time frames. Conversely, the E2E approach processes audio utterances directly, which allows it to pinpoint hate speech within specific time frames. Additionally, due to the lack of explainable audio hate speech datasets that include audio frame-level rationales, we curated a synthetic audio dataset to train our models. We further validated these models on actual human speech utterances and found that the E2E approach outperforms the cascading method in terms of the audio frame Intersection over Union (IoU) metric. Furthermore, we observed that including frame-level rationales significantly enhances hate speech detection accuracy for the E2E approach. textbf{Disclaimer} The reader may encounter content of an offensive or hateful nature. However, given the nature of the work, this cannot be avoided.

【8】 PyNeuralFx: A Python Package for Neural Audio Effect Modeling
标题: PyNeuralDx:用于神经音效建模的Python包
作者:Yen-Tung Yeh,Wen-Yi Hsiao,Yi-Hsuan Yang
备注:toolkit paper
链接:点击下载PDF文件
摘要:我们介绍了PyNeuralFx,这是一个开源Python工具包,专为神经音频效果建模研究而设计。该工具包提供了一个直观的框架,并提供了一套全面的功能,包括完善的模型架构,损失函数和易于使用的可视化工具的标准化实现。因此,它有助于提高神经音频效果建模研究的可重复性,并能够深入比较不同模型的性能,通过DSP方法深入了解模型的行为和操作特性。该工具包可在https: github.com ytsrt66589 pyneuralfx上找到。摘要:We present PyNeuralFx, an open-source Python toolkit designed for research on neural audio effect modeling. The toolkit provides an intuitive framework and offers a comprehensive suite of features, including standardized implementation of well-established model architectures, loss functions, and easy-to-use visualization tools. As such, it helps promote reproducibility for research on neural audio effect modeling, and enable in-depth performance comparison of different models, offering insight into the behavior and operational characteristics of models through DSP methodology. The toolkit can be found at https: github.com ytsrt66589 pyneuralfx.

【9】 Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
标题: 通过噪音表示学习增强具有强大上下文感知的对话语音识别
作者:Wonjun Lee,San Kim,Gary Geunbae Lee
备注:11 pages, 2 figures, Accepted to SIGDIAL2024
链接:点击下载PDF文件
摘要:最近的对话系统依赖于基于话轮的口语交互,需要准确的自动语音识别(ASR)。ASR中的错误会严重影响下游对话任务。为了解决这个问题,已经提出了使用来自用户和代理交互的对话上下文来转录随后的话语。该方法结合了用户的语音和代理的响应作为模型输入的转录,使用由每个回合生成的累积上下文。然而,该上下文容易受到ASR错误的影响,因为它是由ASR模型以自回归方式生成的。这样的噪声上下文可以进一步降低上下文输入的益处,导致次优的ASR性能。在本文中,我们引入上下文噪声表示学习(CNRL),以增强对噪声背景的鲁棒性,最终提高对话语音识别的准确性。为了最大限度地提高上下文感知的优势,我们的方法包括使用基于文本的对话数据和上下文编码器的噪声表示学习的解码器预训练。基于语音对话的评估,我们的方法显示出优越的结果相比,基线。此外,我们的方法的优势是突出在嘈杂的环境中,用户语音几乎听不见,由于现实世界的噪音,依赖于上下文信息准确地转录输入。摘要:Recent dialogue systems rely on turn-based spoken interactions, requiring accurate Automatic Speech Recognition (ASR). Errors in ASR can significantly impact downstream dialogue tasks. To address this, using dialogue context from user and agent interactions for transcribing subsequent utterances has been proposed. This method incorporates the transcription of the user's speech and the agent's response as model input, using the accumulated context generated by each turn. However, this context is susceptible to ASR errors because it is generated by the ASR model in an auto-regressive fashion. Such noisy context can further degrade the benefits of context input, resulting in suboptimal ASR performance. In this paper, we introduce Context Noise Representation Learning (CNRL) to enhance robustness against noisy context, ultimately improving dialogue speech recognition accuracy. To maximize the advantage of context awareness, our approach includes decoder pre-training using text-based dialogue data and noise representation learning for a context encoder. Based on the evaluation of speech dialogues, our method shows superior results compared to baselines. Furthermore, the strength of our approach is highlighted in noisy environments where user speech is barely audible due to real-world noise, relying on contextual information to transcribe the input accurately.

【10】 Controlling Surprisal in Music Generation via Information Content Curve Matching
标题: 通过信息内容曲线匹配控制音乐生成中的惊喜
作者:Mathias Rose Bjare,Stefan Lattner,Gerhard Widmer
备注:8 pages, 4 figures, 2 tables, accepted at the 25th Int. Society for Music Information Retrieval Conf., San Francisco, USA, 2024
链接:点击下载PDF文件
摘要:近年来,音乐生成系统的质量和公众的兴趣已经增长,鼓励研究各种方法来控制这些系统。我们提出了一种新的方法,用于控制在音乐生成中使用序列模型的节拍。为了实现这一目标,我们定义了一个度量称为瞬时信息内容(IIC)。IIC用作感知音乐音阶的代理函数(如从概率模型估计的),并且可以在音乐作品内的任何点处计算。这使得即使音乐事件以不规则的时间间隔发生,也能够跨不同的音乐内容比较重复。我们使用波束搜索来生成音乐素材,其IIC曲线非常接近给定的目标IIC。我们的实验表明,IIC与谐波和节奏的复杂性和注意到密度。相关性随着用于估计IIC的音乐背景的长度而减小。最后,我们进行了一个定性的用户研究,以测试人类听众是否可以识别的IIC曲线,已被用作目标时,产生相应的音乐材料。我们在https: github.com muthissar iic上提供了创建IIC插值和IIC可视化的代码。摘要:In recent years, the quality and public interest in music generation systems have grown, encouraging research into various ways to control these systems. We propose a novel method for controlling surprisal in music generation using sequence models. To achieve this goal, we define a metric called Instantaneous Information Content (IIC). The IIC serves as a proxy function for the perceived musical surprisal (as estimated from a probabilistic model) and can be calculated at any point within a music piece. This enables the comparison of surprisal across different musical content even if the musical events occur in irregular time intervals. We use beam search to generate musical material whose IIC curve closely approximates a given target IIC. We experimentally show that the IIC correlates with harmonic and rhythmic complexity and note density. The correlation decreases with the length of the musical context used for estimating the IIC. Finally, we conduct a qualitative user study to test if human listeners can identify the IIC curves that have been used as targets when generating the respective musical material. We provide code for creating IIC interpolations and IIC visualizations on https: github.com muthissar iic.

【11】 Robust online reconstruction of continuous-time signals from a lean spike train ensemble code
标题: 从精益尖峰序列集合代码中稳健地在线重建连续时间信号
作者:Anik Chattopadhyay,Arunava Banerjee
备注:22 pages, including a 9-page appendix, 8 figures. A GitHub link to the project implementation is embedded in the paper
链接:点击下载PDF文件
摘要:动物的感觉刺激由神经元编码成尖峰序列,提供诸如稀疏性、能量效率和高时间分辨率的优点。本文提出了一个信号处理框架,确定性编码连续时间信号到生物学上可行的尖峰列车,并解决了可表示的信号类和重建边界的问题。该框架考虑通过由神经元的集合使用具有各种卷积核的卷积然后阈值机制生成的尖峰序列对信号进行编码。一个封闭形式的解决方案的逆问题,从尖峰列车信号重建,来自移位核函数的希尔伯特空间,确保稀疏表示的广义有限创新率(FRI)类信号。此外,在生物系统中的实时处理的启发,制定一个有效的迭代版本的最佳重建,只考虑一个有限的窗口,过去的尖峰,确保鲁棒性的技术病态编码;然后提供的窗口重建的最佳解决方案的收敛保证。在大型音频数据集上进行的实验表明,在低至奈奎斯特速率五分之一的尖峰速率下,重建精度出色,同时与低尖峰速率范围内最先进的稀疏编码技术相比,显示出明显的竞争优势。摘要:Sensory stimuli in animals are encoded into spike trains by neurons, offering advantages such as sparsity, energy efficiency, and high temporal resolution. This paper presents a signal processing framework that deterministically encodes continuous-time signals into biologically feasible spike trains, and addresses the questions about representable signal classes and reconstruction bounds. The framework considers encoding of a signal through spike trains generated by an ensemble of neurons using a convolve-then-threshold mechanism with various convolution kernels. A closed-form solution to the inverse problem, from spike trains to signal reconstruction, is derived in the Hilbert space of shifted kernel functions, ensuring sparse representation of a generalized Finite Rate of Innovation (FRI) class of signals. Additionally, inspired by real-time processing in biological systems, an efficient iterative version of the optimal reconstruction is formulated that considers only a finite window of past spikes, ensuring robustness of the technique to ill-conditioned encoding; convergence guarantees of the windowed reconstruction to the optimal solution are then provided. Experiments on a large audio dataset demonstrate excellent reconstruction accuracy at spike rates as low as one-fifth of the Nyquist rate, while showing clear competitive advantage in comparison to state-of-the-art sparse coding techniques in the low spike rate regime.

【12】 Adapting General Disentanglement-Based Speaker Anonymization for Enhanced Emotion Preservation
标题: 调整基于通用解开的说话人解析以增强情感保存
作者:Xiaoxiao Miao,Yuxiang Zhang,Xin Wang,Natalia Tomashenko,Donny Cheng Lock Soh,Ian Mcloughlin
链接:点击下载PDF文件
摘要:一般的基于解纠缠的说话者匿名化系统通常使用单独的编码器将语音分离成内容、说话者和韵律特征。本文探讨了如何适应这样的系统时,一个新的语音属性,例如,情感,需要在更大程度上保留。虽然现有的系统擅长匿名说话者嵌入,但它们并不是为了保留情感而设计的。两种策略,这是检查。首先,我们表明,从预先训练的情感编码器中集成情感嵌入可以帮助保留情感线索,即使这种方法稍微损害了隐私保护。或者,我们提出了一个情感补偿策略作为后处理步骤应用于匿名扬声器嵌入。这隐藏了原始说话者的身份,并重新引入了在说话者嵌入匿名化过程中丢失的情感特征。具体来说,我们使用支持向量机对情感属性进行建模,以学习每种情感的单独边界。在推理过程中,通过两种方式对原始说话人信息进行处理:一是通过情感指示器预测情感,准确选择情感匹配的SVM;二是通过说话人匿名器隐藏说话人特征。然后沿着相应的SVM边界朝着增强的情感方向修改匿名说话人嵌入,以保存情感线索。所提出的策略也预计将是有用的适应一般的基于解纠缠的扬声器匿名化系统,以保持其他目标语言属性,具有潜在的一系列下游任务。摘要:A general disentanglement-based speaker anonymization system typically separates speech into content, speaker, and prosody features using individual encoders. This paper explores how to adapt such a system when a new speech attribute, for example, emotion, needs to be preserved to a greater extent. While existing systems are good at anonymizing speaker embeddings, they are not designed to preserve emotion. Two strategies for this are examined. First, we show that integrating emotion embeddings from a pre-trained emotion encoder can help preserve emotional cues, even though this approach slightly compromises privacy protection. Alternatively, we propose an emotion compensation strategy as a post-processing step applied to anonymized speaker embeddings. This conceals the original speaker's identity and reintroduces the emotional traits lost during speaker embedding anonymization. Specifically, we model the emotion attribute using support vector machines to learn separate boundaries for each emotion. During inference, the original speaker embedding is processed in two ways: one, by an emotion indicator to predict emotion and select the emotion-matched SVM accurately; and two, by a speaker anonymizer to conceal speaker characteristics. The anonymized speaker embedding is then modified along the corresponding SVM boundary towards an enhanced emotional direction to save the emotional cues. The proposed strategies are also expected to be useful for adapting a general disentanglement-based speaker anonymization system to preserve other target paralinguistic attributes, with potential for a range of downstream tasks.

【13】 LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
标题: LI-TTA:自动语音识别的语言知情测试时间自适应
作者:Eunseop Yoon,Hee Suk Yoon,John Harvill,Mark Hasegawa-Johnson,Chang D. Yoo
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:测试时自适应(TTA)已经成为域转移挑战的关键解决方案,其中目标环境偏离原始训练环境。一个主要的例子是自动语音识别(ASR)的TTA,它通过利用输出预测熵最小化作为自我监督信号来增强模型性能。然而,这种自我监督的一个关键限制在于它主要关注声学特征,而很少关注输入的语言特性。为了解决这一差距,我们提出了语言知情的测试时间适应(LI-TTA),它结合了语言的见解,在TTA的ASR。LI-TTA集成了来自外部语言模型的校正,通过最小化校正的CTC损失以及标准TTA损失来合并语言与声学信息。通过大量的实验,我们表明,LI-TTA有效地提高了性能的TTA ASR在各种分布偏移的情况下。摘要:Test-Time Adaptation (TTA) has emerged as a crucial solution to the domain shift challenge, wherein the target environment diverges from the original training environment. A prime exemplification is TTA for Automatic Speech Recognition (ASR), which enhances model performance by leveraging output prediction entropy minimization as a self-supervision signal. However, a key limitation of this self-supervision lies in its primary focus on acoustic features, with minimal attention to the linguistic properties of the input. To address this gap, we propose Language Informed Test-Time Adaptation (LI-TTA), which incorporates linguistic insights during TTA for ASR. LI-TTA integrates corrections from an external language model to merge linguistic with acoustic information by minimizing the CTC loss from the correction alongside the standard TTA loss. With extensive experiments, we show that LI-TTA effectively improves the performance of TTA for ASR in various distribution shift situations.

【14】 Stream-based Active Learning for Anomalous Sound Detection in Machine Condition Monitoring
标题: 基于流的主动学习用于机器状态监测中异常声音检测
作者:Tuan Vu Ho,Kota Dohi,Yohei Kawaguchi
备注:Accepted as a conference paper in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:介绍了一种用于机器状态监测系统中异常声音检测的主动学习框架。通常,由于异常数据的稀缺性,ASD模型仅在正常样本上进行训练,导致在推理过程中不可见样本的准确性降低。AL是解决这个问题的一个很有前途的解决方案,它使模型能够用更少的标记示例更有效地学习新概念,从而减少手动注释工作。然而,其在ASD中的有效性仍未被探索。为了最小化更新成本和时间,我们提出的方法侧重于更新ASD系统的评分后端,而无需重新训练神经网络模型。在DCASE 2023挑战任务2数据集上的实验结果证实,即使在低标签预算的情况下,我们的AL框架也显着提高了ASD性能。此外,我们提出的抽样策略优于其他基线的部分区域下的受试者操作特征得分。摘要:This paper introduces an active learning (AL) framework for anomalous sound detection (ASD) in machine condition monitoring system. Typically, ASD models are trained solely on normal samples due to the scarcity of anomalous data, leading to decreased accuracy for unseen samples during inference. AL is a promising solution to solve this problem by enabling the model to learn new concepts more effectively with fewer labeled examples, thus reducing manual annotation efforts. However, its effectiveness in ASD remains unexplored. To minimize update costs and time, our proposed method focuses on updating the scoring backend of ASD system without retraining the neural network model. Experimental results on the DCASE 2023 Challenge Task 2 dataset confirm that our AL framework significantly improves ASD performance even with low labeling budgets. Moreover, our proposed sampling strategy outperforms other baselines in terms of the partial area under the receiver operating characteristic score.

【15】 Learning Delays in Spiking Neural Networks using Dilated Convolutions with Learnable Spacings
标题: 使用具有可学习间隔的扩张卷积的尖峰神经网络的学习延迟
作者:Ilyass Hammouamri,Ismail Khalfaoui-Hassani,Timothée Masquelier
Journal-ref:ICLR 2024
链接:点击下载PDF文件
摘要:尖峰神经网络(SNN)是一个很有前途的研究方向,用于构建节能的信息处理系统,特别是用于语音识别等时间任务。在SNN中,延迟是指一个尖峰从一个神经元传播到另一个神经元所需的时间。这些延迟很重要,因为它们会影响尖峰到达时间,而且众所周知,尖峰神经元对重合的输入尖峰反应更强烈。更正式地说,理论上已经表明,塑性延迟大大增加了SNN的表现力。然而,缺乏有效的算法来学习这些延迟。在这里,我们提出了一种新的离散时间算法,以离线方式使用反向传播在深度前馈SNN中解决这个问题。为了模拟连续层之间的延迟,我们使用跨时间的1D卷积。内核只包含几个非零权重-每个突触一个-其位置对应于延迟。这些位置与权重一起使用最近提出的具有可学习间隔的扩张卷积(DCLS)来学习。我们在三个数据集上评估了我们的方法:尖峰海德堡数据集(SHD)、尖峰语音命令(SSC)及其非尖峰版本Google Speech Commands v0.02(GSC)基准测试,这些基准测试需要检测时间模式。我们使用了具有两个或三个隐藏的全连接层的前馈SNN,以及香草泄漏的积分和激发神经元。我们发现,固定的随机延迟有帮助,学习它们的帮助更大。此外,我们的方法在三个数据集上的表现优于最先进的方法,而不使用递归连接,并且参数少得多。我们的工作证明了延迟学习在开发准确和精确的时态数据处理模型方面的潜力。我们的代码基于PyTorch SpikingJelly,可在https: github.com Thvnvtos SNN-delays上获得摘要:Spiking Neural Networks (SNNs) are a promising research direction for building power-efficient information processing systems, especially for temporal tasks such as speech recognition. In SNNs, delays refer to the time needed for one spike to travel from one neuron to another. These delays matter because they influence the spike arrival times, and it is well-known that spiking neurons respond more strongly to coincident input spikes. More formally, it has been shown theoretically that plastic delays greatly increase the expressivity in SNNs. Yet, efficient algorithms to learn these delays have been lacking. Here, we propose a new discrete-time algorithm that addresses this issue in deep feedforward SNNs using backpropagation, in an offline manner. To simulate delays between consecutive layers, we use 1D convolutions across time. The kernels contain only a few non-zero weights - one per synapse - whose positions correspond to the delays. These positions are learned together with the weights using the recently proposed Dilated Convolution with Learnable Spacings (DCLS). We evaluated our method on three datasets: the Spiking Heidelberg Dataset (SHD), the Spiking Speech Commands (SSC) and its non-spiking version Google Speech Commands v0.02 (GSC) benchmarks, which require detecting temporal patterns. We used feedforward SNNs with two or three hidden fully connected layers, and vanilla leaky integrate-and-fire neurons. We showed that fixed random delays help and that learning them helps even more. Furthermore, our method outperformed the state-of-the-art in the three datasets without using recurrent connections and with substantially fewer parameters. Our work demonstrates the potential of delay learning in developing accurate and precise models for temporal data processing. Our code is based on PyTorch SpikingJelly and available at: https: github.com Thvnvtos SNN-delays


机器翻译,仅供参考