今日论文合集:cs.SD语音9篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载

cs.SD语音
【1】 Sample-Efficient Diffusion for Text-To-Speech Synthesis
标题: 文本到语音合成的样本高效扩散
作者:Justin Lovelace,Soham Ray,Kwangyoun Kim,Kilian Q. Weinberger,Felix Wu
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:这项工作介绍了采样有效的语音扩散(SESD),有效的语音合成算法在适度的数据制度,通过潜在的扩散。它基于一种新型扩散架构,我们称之为U-Audio Transformer(U-AT),该架构可有效扩展到长序列,并在预训练音频自动编码器的潜在空间中运行。以字符感知语言模型表示为条件,SESD尽管在不到1 k小时的语音上进行了训练,但仍然取得了令人印象深刻的结果-远远低于当前最先进的系统。事实上,它比最先进的自回归模型VALL-E合成更清晰的语音,而使用不到2%的训练数据。摘要:This work introduces Sample-Efficient Speech Diffusion (SESD), an algorithm for effective speech synthesis in modest data regimes through latent diffusion. It is based on a novel diffusion architecture, that we call U-Audio Transformer (U-AT), that efficiently scales to long sequences and operates in the latent space of a pre-trained audio autoencoder. Conditioned on character-aware language model representations, SESD achieves impressive results despite training on less than 1k hours of speech - far less than current state-of-the-art systems. In fact, it synthesizes more intelligible speech than the state-of-the-art auto-regressive model, VALL-E, while using less than 2% the training data.

【2】 Applications and Advances of Artificial Intelligence in Music Generation:A Review
标题: 人工智能在音乐生成中的应用与进展:综述
作者:Yanxu Chen,Linshu Huang,Tian Gou
链接:点击下载PDF文件
摘要:近年来,人工智能(AI)在音乐生成领域取得了重大进展,推动了音乐创作和应用的创新。本文系统地综述了人工智能音乐生成的最新研究进展,涵盖了关键技术、模型、数据集、评估方法及其在各个领域的实际应用。这篇综述的主要贡献包括:(1)提出了一个全面的总结框架,系统地分类和比较了不同的技术方法,包括符号生成,音频生成和混合模型,帮助读者更好地理解该领域的全方位技术;(2)提供对当前文献的广泛调查,涵盖多模态数据集和情感表达评估等新兴主题,为相关研究提供广泛的参考;(3)详细分析人工智能音乐生成在各个应用领域的实际影响,特别是在实时交互和跨学科应用中,提供新的视角和见解;(4)总结音乐质量评价方法存在的挑战和局限性,提出未来潜在的研究方向,旨在促进评价技术的标准化和更广泛的采用。通过这些创新性的总结和分析,本文为人工智能音乐生成领域的研究人员和实践者提供了全面的参考工具,同时也概述了该领域的未来发展方向。摘要:In recent years, artificial intelligence (AI) has made significant progress in the field of music generation, driving innovation in music creation and applications. This paper provides a systematic review of the latest research advancements in AI music generation, covering key technologies, models, datasets, evaluation methods, and their practical applications across various fields. The main contributions of this review include: (1) presenting a comprehensive summary framework that systematically categorizes and compares different technological approaches, including symbolic generation, audio generation, and hybrid models, helping readers better understand the full spectrum of technologies in the field; (2) offering an extensive survey of current literature, covering emerging topics such as multimodal datasets and emotion expression evaluation, providing a broad reference for related research; (3) conducting a detailed analysis of the practical impact of AI music generation in various application domains, particularly in real-time interaction and interdisciplinary applications, offering new perspectives and insights; (4) summarizing the existing challenges and limitations of music quality evaluation methods and proposing potential future research directions, aiming to promote the standardization and broader adoption of evaluation techniques. Through these innovative summaries and analyses, this paper serves as a comprehensive reference tool for researchers and practitioners in AI music generation, while also outlining future directions for the field.

【3】 Clustering of Indonesian and Western Gamelan Orchestras through Machine Learning of Performance Parameters
标题: 通过性能参数的机器学习对印度尼西亚和西部Gamelan gastras进行聚集
作者:Simon Linke,Gerrit Wendt,Rolf Bader
备注:5 figures, 4 tables
链接:点击下载PDF文件
摘要:印度尼西亚和西方的甘美兰合奏进行了调查方面的性能差异。因此,这种音乐在西方的异国情调的历史可能反映在当代的音调系统,发音或大规模的形式差异。分析四个西方和五个印尼乐团的录音方面的音调系统和音色特征,并使用自组织Kohonen地图(SOM)作为机器学习算法,印尼和西方合奏之间出现了明确的聚类使用某些心理声学功能。这表明,与印度尼西亚的合奏相比,西方合奏的发音和大规模的形式变化减少了。SOM还集群合奏就其音调系统,但没有集群之间的印尼和西方合奏可以在这方面找到。因此,一个清晰的类比之间较低的发音变化和大规模的形式变化和一个更异国情调,调解和冷静的性能预期和接受的甘美兰在西方因此出现。摘要:Indonesian and Western gamelan ensembles are investigated with respect to performance differences. Thereby, the often exotistic history of this music in the West might be reflected in contemporary tonal system, articulation, or large-scale form differences. Analyzing recordings of four Western and five Indonesian orchestras with respect to tonal systems and timbre features and using self-organizing Kohonen map (SOM) as a machine learning algorithm, a clear clustering between Indonesian and Western ensembles appears using certain psychoacoustic features. These point to a reduced articulation and large-scale form variability of Western ensembles compared to Indonesian ones. The SOM also clusters the ensembles with respect to their tonal systems, but no clusters between Indonesian and Western ensembles can be found in this respect. Therefore, a clear analogy between lower articulatory variability and large-scale form variation and a more exostistic, mediative and calm performance expectation and reception of gamelan in the West therefore appears.

【4】 LAST: Language Model Aware Speech Tokenization
标题: 最后:语言模型感知语音令牌化
作者:Arnon Turetzky,Yossi Adi
链接:点击下载PDF文件
摘要:语音标记作为语音语言模型(LM)的基础,使他们能够执行各种任务,如口语建模,文本到语音,语音到文本等。遵循这种方法可能会在令牌化过程及其后续使用之间产生不匹配。在这项研究中,我们提出了一种新的方法来训练语音标记器,利用目标从预先训练的文本LM。我们主张将这一目标融入离散语音表征的学习过程中。我们的目标是将来自预训练语音模型的特征转换为新的特征空间,从而能够更好地对语音LM进行聚类。我们实证研究的影响,各种模型的设计选择,包括语音词汇量的大小和文本LM的大小。我们的结果表明,所提出的标记化方法优于评估的基线,考虑到口语建模和语音到文本。更重要的是,与先前的工作不同,所提出的方法允许利用单个预训练的LM来处理语音和文本输入,将其与传统的标记化方法区分开来。摘要:Speech tokenization serves as the foundation of speech language model (LM), enabling them to perform various tasks such as spoken language modeling, text-to-speech, speech-to-text, etc. Most speech tokenizers are trained independently of the LM training process, relying on separate acoustic models and quantization methods. Following such an approach may create a mismatch between the tokenization process and its usage afterward. In this study, we propose a novel approach to training a speech tokenizer by leveraging objectives from pre-trained textual LMs. We advocate for the integration of this objective into the process of learning discrete speech representations. Our aim is to transform features from a pre-trained speech model into a new feature space that enables better clustering for speech LMs. We empirically investigate the impact of various model design choices, including speech vocabulary size and text LM size. Our results demonstrate the proposed tokenization method outperforms the evaluated baselines considering both spoken language modeling and speech-to-text. More importantly, unlike prior work, the proposed method allows the utilization of a single pre-trained LM for processing both speech and text inputs, setting it apart from conventional tokenization approaches.

【5】 Multimodal Laryngoscopic Video Analysis for Assisted Diagnosis of Vocal Cord Paralysis
标题: 多模式喉镜视频分析辅助诊断声带瘫痪
作者:Yucong Zhang,Xin Zou,Jinshan Yang,Wenjun Chen,Faya Liang,Ming Li
链接:点击下载PDF文件
摘要:本文介绍了喉镜多模态分析系统(MASL),一个系统,结合音频和视频数据,自动提取关键段和指标,从喉视频频闪视频临床评估。MASL将声门检测与关键字定位相结合,以分析患者的发声并优化视频亮点,从而更好地检查声带运动。该系统包括通过分析色调、饱和度和值波动来识别帧的频闪视频提取模块。MASL还为声带麻痹检测提供了有效的指标,采用两阶段声门分割过程,使用U-Net,然后进行基于扩散的细化,以减少误报。代替声门面积波形,MASL从声门掩模估计前声门角波形(AGAW),评估左声带和右声带以检测单侧声带麻痹(UVFP)。通过比较AGAW方差,MASL区分左右瘫痪。在公共和真实世界数据集上的消融研究和实验验证了MASL的分割模块,并证明了其为UVFP诊断提供可靠指标的能力。摘要:This paper presents the Multimodal Analyzing System for Laryngoscope (MASL), a system that combines audio and video data to automatically extract key segments and metrics from laryngeal videostroboscopic videos for clinical assessment. MASL integrates glottis detection with keyword spotting to analyze patient vocalizations and refine video highlights for better inspection of vocal cord movements. The system includes a strobing video extraction module that identifies frames by analyzing hue, saturation, and value fluctuations. MASL also provides effective metrics for vocal cord paralysis detection, employing a two-stage glottis segmentation process using U-Net followed by diffusion-based refinement to reduce false positives. Instead of glottal area waveforms, MASL estimates anterior glottic angle waveforms (AGAW) from glottis masks, evaluating both left and right vocal cords to detect unilateral vocal cord paralysis (UVFP). By comparing AGAW variances, MASL distinguishes between left and right paralysis. Ablation studies and experiments on public and real-world datasets validate MASL's segmentation module and demonstrate its ability to provide reliable metrics for UVFP diagnosis.

【6】 Raw Speech Enhancement with Deep State Space Modeling
标题: 利用深状态空间建模增强原始语音
作者:Yan Ru Pei,Ritik Shrivastava,FNU Sidharth
备注:7 pages, 2 figures
链接:点击下载PDF文件
摘要:我们提出了aTENNuate,一个简单的深状态空间自动编码器配置为有效的在线原始语音增强在一个端到端的方式。该网络的性能主要在原始语音去噪方面进行评估,并对超分辨率和去量化等任务进行额外评估。我们在VoiceBank + DEMAND和Microsoft DNS1合成测试集上对aTENNuate进行了基准测试。该网络在PESQ分数、参数计数、MAC和延迟方面优于以前的实时去噪模型。即使作为原始波形处理模型,该模型也保持了对干净信号的高保真度,同时具有最小的可听伪影。此外,该模型保持性能,即使有噪声的输入被压缩到4000 Hz和4位,这表明在低资源环境中的一般语音增强能力。摘要:We present aTENNuate, a simple deep state-space autoencoder configured for efficient online raw speech enhancement in an end-to-end fashion. The network's performance is primarily evaluated on raw speech denoising, with additional assessments on tasks such as super-resolution and de-quantization. We benchmark aTENNuate on the VoiceBank + DEMAND and the Microsoft DNS1 synthetic test sets. The network outperforms previous real-time denoising models in terms of PESQ score, parameter count, MACs, and latency. Even as a raw waveform processing model, the model maintains high fidelity to the clean signal with minimal audible artifacts. In addition, the model remains performant even when the noisy input is compressed down to 4000Hz and 4 bits, suggesting general speech enhancement capabilities in low-resource environments.

【7】 Eetimating Indoor Scene Depth Maps from Ultrasonic Echoes
标题: 根据超声波回声确定室内场景深度图
作者:Junpei Honma,Akisato Kimura,Go Irie
备注:ICIP 2024
链接:点击下载PDF文件
摘要:测量室内场景的3D几何结构需要专用的深度传感器,但这并不总是可用的。基于回声的深度估计最近被研究为有前途的替代解决方案。所有以前的研究都假设使用的回声在可听范围内。然而,一个主要问题是可听回声不能用于安静的空间或禁止产生可听声音的其他情况。在本文中,我们考虑使用听不见的超声回波的回声为基础的深度估计。虽然超声波在理论上提供高的测量精度,但是由于其对噪声敏感并且易于衰减的缺点,当使用超声波回波时的实际深度估计精度仍然不清楚。我们首先调查的深度估计精度时,声源的频率被限制在高频带,并发现,精度下降时,频率被限制在超声波范围。基于这一观察,我们提出了一种新的深度学习方法,通过仅在训练期间使用可听回波作为辅助数据来提高基于超声回波的深度估计的准确性。在公开数据集上的实验结果表明,该方法提高了估计精度.摘要:Measuring 3D geometric structures of indoor scenes requires dedicated depth sensors, which are not always available. Echo-based depth estimation has recently been studied as a promising alternative solution. All previous studies have assumed the use of echoes in the audible range. However, one major problem is that audible echoes cannot be used in quiet spaces or other situations where producing audible sounds is prohibited. In this paper, we consider echo-based depth estimation using inaudible ultrasonic echoes. While ultrasonic waves provide high measurement accuracy in theory, the actual depth estimation accuracy when ultrasonic echoes are used has remained unclear, due to its disadvantage of being sensitive to noise and susceptible to attenuation. We first investigate the depth estimation accuracy when the frequency of the sound source is restricted to the high-frequency band, and found that the accuracy decreased when the frequency was limited to ultrasonic ranges. Based on this observation, we propose a novel deep learning method to improve the accuracy of ultrasonic echo-based depth estimation by using audible echoes as auxiliary data only during training. Experimental results with a public dataset demonstrate that our method improves the estimation accuracy.

【8】 FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
标题: FireRedRTS:用于行业级生成语音应用的基础文本到语音框架
作者:Hao-Han Guo,Kun Liu,Fei-Yu Shen,Yi-Chen Wu,Feng-Long Xie,Kun Xie,Kai-Tuo Xu
链接:点击下载PDF文件
摘要:这项工作提出了FireRedTTS(一个基础文本到语音框架),以满足对个性化和多样化生成语音应用程序日益增长的需求。该框架包括三个部分:数据处理,基础系统和下游应用程序。首先,我们全面介绍了我们的数据处理管道,它将大量原始音频转换为具有丰富注释和广泛内容,说话风格和音色的大规模高质量TTS数据集。然后,我们提出了一个基于语言模型的基础TTS系统。语音信号通过语义感知的语音分词器被压缩成离散的语义令牌,并且可以由语言模型从提示文本和音频生成。然后,提出了一种两级波形发生器,将它们解码为高保真波形。我们提出了这个系统的两个应用程序:配音的语音克隆和聊天机器人的类人语音生成。实验结果表明,FireRedTTS具有较强的上下文学习能力,能够稳定地合成出与提示文本和音频一致的高质量语音。在配音方面,FireRedTTS可针对UGC场景以zero-shot方式克隆目标声音,并通过Few-Shot微调,1小时录音,适应PUGC场景中演播室级别的表现力声音角色。此外,FireRedTTS通过指令调整,以轻松的方式实现可控的类人语音生成,具有非语言行为和情感,以更好地服务于口语聊天机器人。摘要:This work proposes FireRedTTS, a foundation text-to-speech framework, to meet the growing demands for personalized and diverse generative speech applications. The framework comprises three parts: data processing, foundation system, and downstream applications. First, we comprehensively present our data processing pipeline, which transforms massive raw audio into a large-scale high-quality TTS dataset with rich annotations and a wide coverage of content, speaking style, and timbre. Then, we propose a language-model-based foundation TTS system. The speech signal is compressed into discrete semantic tokens via a semantic-aware speech tokenizer, and can be generated by a language model from the prompt text and audio. Then, a two-stage waveform generator is proposed to decode them to the high-fidelity waveform. We present two applications of this system: voice cloning for dubbing and human-like speech generation for chatbots. The experimental results demonstrate the solid in-context learning capability of FireRedTTS, which can stably synthesize high-quality speech consistent with the prompt text and audio. For dubbing, FireRedTTS can clone target voices in a zero-shot way for the UGC scenario and adapt to studio-level expressive voice characters in the PUGC scenario via few-shot fine-tuning with 1-hour recording. Moreover, FireRedTTS achieves controllable human-like speech generation in a casual style with paralinguistic behaviors and emotions via instruction tuning, to better serve spoken chatbots.

【9】 SymPAC: Scalable Symbolic Music Generation With Prompts And Constraints
标题: SymPAC:具有预设和约束的可扩展符号音乐生成
作者:Haonan Chen,Jordan B. L. Smith,Bochen Li,Ju-Chiang Wang,Janne Spijkervet,Pei Zou,Xingjian Du,Qiuqiang Kong
备注:ISMIR 2024
链接:点击下载PDF文件
摘要:符号音乐生成任务的进展可能落后于音频和文本生成等其他任务,部分原因是符号训练数据的稀缺。在本文中,我们通过应用预训练的MIR模型(用于转录,节拍跟踪,结构分析等)来利用更大规模的音频音乐数据。提取符号事件并将其编码为令牌序列。据我们所知,这项工作是第一次证明了仅从自动转录的音频数据训练符号生成模型的可行性。此外,为了增强训练模型的可控性,我们引入了Symbolic Music Language Model with Encoding And Constrained Generation(符号音乐语言模型),其区别在于:(a)在编码中使用提示条,(b)在推理时间使用一种称为通过有限状态机(FSM)的约束生成技术。我们展示了这种方法的灵活性和可控性,这对于使音乐AI对创作者和用户有用至关重要。摘要:Progress in the task of symbolic music generation may be lagging behind other tasks like audio and text generation, in part because of the scarcity of symbolic training data. In this paper, we leverage the greater scale of audio music data by applying pre-trained MIR models (for transcription, beat tracking, structure analysis, etc.) to extract symbolic events and encode them into token sequences. To the best of our knowledge, this work is the first to demonstrate the feasibility of training symbolic generation models solely from auto-transcribed audio data. Furthermore, to enhance the controllability of the trained model, we introduce SymPAC (Symbolic Music Language Model with Prompting And Constrained Generation), which is distinguished by using (a) prompt bars in encoding and (b) a technique called Constrained Generation via Finite State Machines (FSMs) during inference time. We show the flexibility and controllability of this approach, which may be critical in making music AI useful to creators and users.


eess.AS音频处理
【1】 Privacy versus Emotion Preservation Trade-offs in Emotion-Preserving Speaker Anonymization
标题: 保密发言人匿名化中的隐私与情感保护的权衡
作者:Zexin Cai,Henry Li Xinyuan,Ashi Garg,Leibny Paola García-Perera,Kevin Duh,Sanjeev Khudanpur,Nicholas Andrews,Matthew Wiesner
备注:accepted by 2024 IEEE Spoken Language Technology Workshop
链接:点击下载PDF文件
摘要:语音技术的进步现在允许通过语音前所未有地访问个人身份信息。为了保护这些信息,差分隐私领域已经探索了匿名化语音的方法,同时保留其实用性,包括语言学和非语言学方面。然而,在保持情绪状态的同时匿名化语音仍然具有挑战性。我们在VoicePrivacy 2024挑战的背景下探讨这个问题。具体来说,我们开发了各种说话者匿名管道,发现这些方法要么擅长匿名化,要么擅长保留情感状态,但不能同时兼顾两者。要实现这两个目标,就需要一个域内情感识别器。此外,我们发现,它是可行的训练半有效的说话人确认系统,仅使用情感表示,展示了这两种模式分离的挑战。摘要:Advances in speech technology now allow unprecedented access to personally identifiable information through speech. To protect such information, the differential privacy field has explored ways to anonymize speech while preserving its utility, including linguistic and paralinguistic aspects. However, anonymizing speech while maintaining emotional state remains challenging. We explore this problem in the context of the VoicePrivacy 2024 challenge. Specifically, we developed various speaker anonymization pipelines and find that approaches either excel at anonymization or preserving emotion state, but not both simultaneously. Achieving both would require an in-domain emotion recognizer. Additionally, we found that it is feasible to train a semi-effective speaker verification system using only emotion representations, demonstrating the challenge of separating these two modalities.

【2】 DiffEVC: Any-to-Any Emotion Voice Conversion with Expressive Guidance
标题: DiffEVC:具有表达性指导的任意情感语音转换
作者:Hsing-Hang Chou,Yun-Shao Lin,Ching-Chin Sung,Yu Tsao,Chi-Chun Lee
链接:点击下载PDF文件
摘要:情感语音转换(EVC)通过放大积极线索和减少消极线索来修改语音情感以增强交流。这个复杂的任务涉及到诸如语音质量、说话者特征和内容等复杂因素。传统的深度学习模型,如GANs和自动编码器,通过学习映射或分解特征,在EVC中取得了一些成功,但面临着不稳定和语音质量下降等挑战。扩散模型提供稳定的训练和高质量的生成。我们提出了一个基于扩散的EVC框架,使用互信息损失和辅助模型解开情感和说话人身份。一个表达性的指导机制,以提高情感转换,同时保持扬声器的特点。实验结果表明,我们的方法的有效性看不见的扬声器和情绪,实现国家的最先进的性能EVC任务。摘要:Emotional Voice Conversion (EVC) modifies speech emotion to enhance communication by amplifying positive cues and reducing negative ones. This complex task involves entangled factors like voice quality, speaker traits, and content. Traditional deep learning models like GANs and autoencoders have achieved some success in EVC by learning mappings or disentangling features but face challenges like instability and voice quality degradation. Diffusion models offer stable training and high-quality generation. We propose a diffusion-based EVC framework that disentangles emotion and speaker identity using mutual information loss and auxiliary models. An expressive guidance mechanism is introduced to improve emotion conversion while maintaining speaker traits. Experimental results demonstrate our approach's effectiveness for unseen speakers and emotions, achieving state-of-the-art performance in EVC tasks.

【3】 A Dual-Path Framework with Frequency-and-Time Excited Network for Anomalous Sound Detection
标题: 用于异常声音检测的具有频率和时间激励网络的双路径框架
作者:Yucong Zhang,Juan Liu,Yao Tian,Haifeng Liu,Ming Li
备注:This Paper has been accepted to ICASSP 2024
链接:点击下载PDF文件
摘要:与人类语音相比,机器生成的相同类型的声音通常表现出一致的频率特性和可辨别的时间周期性。然而,在异常检测中利用这些双重属性仍然相对不足。在本文中,我们提出了一个自动化的双路径框架,学习不同机器类型的突出频率和时间模式。一种途径使用一种新的频率和时间激励网络(FTE-Net)来学习频谱图的频率和时间轴上的显著特征。它结合了频率和时间块编码器(FTC编码器)和激励网络。另一种途径使用1D卷积网络进行话语级频谱。在DCASE 2023 task 2数据集上的实验结果显示了我们所提出的方法的最新性能。此外,可视化的中间特征映射的激励网络中提供说明我们的方法的有效性。摘要:In contrast to human speech, machine-generated sounds of the same type often exhibit consistent frequency characteristics and discernible temporal periodicity. However, leveraging these dual attributes in anomaly detection remains relatively under-explored. In this paper, we propose an automated dual-path framework that learns prominent frequency and temporal patterns for diverse machine types. One pathway uses a novel Frequency-and-Time Excited Network (FTE-Net) to learn the salient features across frequency and time axes of the spectrogram. It incorporates a Frequency-and-Time Chunkwise Encoder (FTC-Encoder) and an excitation network. The other pathway uses a 1D convolutional network for utterance-level spectrum. Experimental results on the DCASE 2023 task 2 dataset show the state-of-the-art performance of our proposed method. Moreover, visualizations of the intermediate feature maps in the excitation network are provided to illustrate the effectiveness of our method.

【4】 Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
标题: 基于对比预测编码支持的因子化变分自动编码器的语音说话人和风格解纠缠
作者:Yuying Xie,Michael Kuhlmann,Frederik Rautenberg,Zheng-Hua Tan,Reinhold Haeb-Umbach
备注:Accepted by EUSIPCO 2024
链接:点击下载PDF文件
摘要:语音信号包含跨多个级别的各种信息,包括内容、说话者和风格。这些信息的分离虽然具有挑战性,但对于语音转换等应用非常重要。支持对比预测编码的因式分解变分自编码器通过假设说话者信息在时间上比内容引起的变化更稳定来实现语音信号到说话者和内容嵌入的无监督解缠。然而,这种假设可能会引入其他时间稳定的信息到说话人嵌入,如环境或情感,我们称之为风格。在这项工作中,我们提出了一种方法,以进一步解开非内容功能到不同的扬声器和风格的功能,特别是通过利用易于访问和定义良好的扬声器标签,而无需风格标签。实验结果验证了该方法在提取分离特征方面的有效性,从而有助于说话人、风格或组合说话人风格转换。摘要:Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important for applications such as voice conversion. The contrastive predictive coding supported factorized variational autoencoder achieves unsupervised disentanglement of a speech signal into speaker and content embeddings by assuming speaker info to be temporally more stable than content-induced variations. However, this assumption may introduce other temporal stable information into the speaker embeddings, like environment or emotion, which we call style. In this work, we propose a method to further disentangle non-content features into distinct speaker and style features, notably by leveraging readily accessible and well-defined speaker labels without the necessity for style labels. Experimental results validate the proposed method's effectiveness on extracting disentangled features, thereby facilitating speaker, style, or combined speaker-style conversion.

【5】 A spherical harmonic-domain spatial audio signal enhancement method based on minimum variance distortionless response
标题: 基于最小方差无失真响应的球调和域空间音频信号增强方法
作者:Huawei Zhang,Jihui,Zhang,Huiyuan,Sun,Prasanga Samarasinghe
链接:点击下载PDF文件
摘要:空间音频信号增强的目的是减少干扰源的贡献,同时保持所需的声场与其空间线索完好无损。现有的方法通常依赖于不切实际的假设(例如,没有混响或不切实际的信息的准确估计)或具有有限的适用性。本文提出了一种基于球面谐波(SH)域最小方差无失真响应(MVDR)的空间信号增强器,该增强器利用相对谐波系数(ReHC)从混响环境中的噪声中提取干净的SH系数。仿真研究表明,该方法实现了更低的估计误差,更高的语音失真比(SDR),并在混响环境中的甜蜜区内的噪声降低(NR)相比,波束形成和投影方法作为基线。摘要:Spatial audio signal enhancement aims to reduce interfering source contributions while preserving the desired sound field with its spatial cues intact. Existing methods generally rely on impractical assumptions (e.g. no reverberation or accurate estimations of impractical information) or have limited applicability. This paper presents a spherical harmonic (SH)-domain minimum variance distortionless response (MVDR)-based spatial signal enhancer using Relative Harmonic Coefficients (ReHCs) to extract clean SH coefficients from noisy ones in reverberant environments. A simulation study shows the proposed method achieves lower estimation error, higher speech-distortion-ratio (SDR), and comparable noise reduction (NR) within the sweet area in a reverberant environment, compared to a beamforming-and-projection method as the baseline.

【6】 Applications and Advances of Artificial Intelligence in Music Generation:A Review
标题: 人工智能在音乐生成中的应用与进展:综述
作者:Yanxu Chen,Linshu Huang,Tian Gou
链接:点击下载PDF文件
摘要:近年来,人工智能(AI)在音乐生成领域取得了重大进展,推动了音乐创作和应用的创新。本文系统地综述了人工智能音乐生成的最新研究进展,涵盖了关键技术、模型、数据集、评估方法及其在各个领域的实际应用。这篇综述的主要贡献包括:(1)提出了一个全面的总结框架,系统地分类和比较了不同的技术方法,包括符号生成,音频生成和混合模型,帮助读者更好地理解该领域的全方位技术;(2)提供对当前文献的广泛调查,涵盖多模式数据集和情感表达评估等新兴主题,为相关研究提供广泛的参考;(3)详细分析人工智能音乐生成在各个应用领域的实际影响,特别是在实时交互和跨学科应用中,提供新的视角和见解;(4)总结音乐质量评价方法存在的挑战和局限性,提出未来潜在的研究方向,旨在促进评价技术的标准化和更广泛的采用。通过这些创新性的总结和分析,本文为人工智能音乐生成领域的研究人员和实践者提供了全面的参考工具,同时也概述了该领域的未来发展方向。摘要:In recent years, artificial intelligence (AI) has made significant progress in the field of music generation, driving innovation in music creation and applications. This paper provides a systematic review of the latest research advancements in AI music generation, covering key technologies, models, datasets, evaluation methods, and their practical applications across various fields. The main contributions of this review include: (1) presenting a comprehensive summary framework that systematically categorizes and compares different technological approaches, including symbolic generation, audio generation, and hybrid models, helping readers better understand the full spectrum of technologies in the field; (2) offering an extensive survey of current literature, covering emerging topics such as multimodal datasets and emotion expression evaluation, providing a broad reference for related research; (3) conducting a detailed analysis of the practical impact of AI music generation in various application domains, particularly in real-time interaction and interdisciplinary applications, offering new perspectives and insights; (4) summarizing the existing challenges and limitations of music quality evaluation methods and proposing potential future research directions, aiming to promote the standardization and broader adoption of evaluation techniques. Through these innovative summaries and analyses, this paper serves as a comprehensive reference tool for researchers and practitioners in AI music generation, while also outlining future directions for the field.

【7】 LAST: Language Model Aware Speech Tokenization
标题: 最后:语言模型感知语音令牌化
作者:Arnon Turetzky,Yossi Adi
链接:点击下载PDF文件
摘要:语音标记作为语音语言模型(LM)的基础,使他们能够执行各种任务,如口语建模,文本到语音,语音到文本等。遵循这种方法可能会在令牌化过程及其后续使用之间产生不匹配。在这项研究中,我们提出了一种新的方法来训练语音标记器,利用目标从预先训练的文本LM。我们主张将这一目标融入离散语音表征的学习过程中。我们的目标是将来自预训练语音模型的特征转换为新的特征空间,从而能够更好地对语音LM进行聚类。我们实证研究的影响,各种模型的设计选择,包括语音词汇量的大小和文本LM的大小。我们的结果表明,所提出的标记化方法优于评估的基线,考虑到口语建模和语音到文本。更重要的是,与先前的工作不同,所提出的方法允许利用单个预训练的LM来处理语音和文本输入,将其与传统的标记化方法区分开来。摘要:Speech tokenization serves as the foundation of speech language model (LM), enabling them to perform various tasks such as spoken language modeling, text-to-speech, speech-to-text, etc. Most speech tokenizers are trained independently of the LM training process, relying on separate acoustic models and quantization methods. Following such an approach may create a mismatch between the tokenization process and its usage afterward. In this study, we propose a novel approach to training a speech tokenizer by leveraging objectives from pre-trained textual LMs. We advocate for the integration of this objective into the process of learning discrete speech representations. Our aim is to transform features from a pre-trained speech model into a new feature space that enables better clustering for speech LMs. We empirically investigate the impact of various model design choices, including speech vocabulary size and text LM size. Our results demonstrate the proposed tokenization method outperforms the evaluated baselines considering both spoken language modeling and speech-to-text. More importantly, unlike prior work, the proposed method allows the utilization of a single pre-trained LM for processing both speech and text inputs, setting it apart from conventional tokenization approaches.

【8】 Multimodal Laryngoscopic Video Analysis for Assisted Diagnosis of Vocal Cord Paralysis
标题: 多模式喉镜视频分析辅助诊断声带瘫痪
作者:Yucong Zhang,Xin Zou,Jinshan Yang,Wenjun Chen,Faya Liang,Ming Li
链接:点击下载PDF文件
摘要:本文介绍了喉镜多模态分析系统(MASL),一个系统,结合音频和视频数据,自动提取关键段和指标,从喉视频频闪视频临床评估。MASL将声门检测与关键字定位相结合,以分析患者的发声并优化视频亮点,从而更好地检查声带运动。该系统包括通过分析色调、饱和度和值波动来识别帧的频闪视频提取模块。MASL还为声带麻痹检测提供了有效的指标,采用两阶段声门分割过程,使用U-Net,然后进行基于扩散的细化,以减少误报。代替声门面积波形,MASL从声门掩模估计前声门角波形(AGAW),评估左声带和右声带以检测单侧声带麻痹(UVFP)。通过比较AGAW方差,MASL区分左右瘫痪。在公共和真实世界数据集上的消融研究和实验验证了MASL的分割模块,并证明了其为UVFP诊断提供可靠指标的能力。摘要:This paper presents the Multimodal Analyzing System for Laryngoscope (MASL), a system that combines audio and video data to automatically extract key segments and metrics from laryngeal videostroboscopic videos for clinical assessment. MASL integrates glottis detection with keyword spotting to analyze patient vocalizations and refine video highlights for better inspection of vocal cord movements. The system includes a strobing video extraction module that identifies frames by analyzing hue, saturation, and value fluctuations. MASL also provides effective metrics for vocal cord paralysis detection, employing a two-stage glottis segmentation process using U-Net followed by diffusion-based refinement to reduce false positives. Instead of glottal area waveforms, MASL estimates anterior glottic angle waveforms (AGAW) from glottis masks, evaluating both left and right vocal cords to detect unilateral vocal cord paralysis (UVFP). By comparing AGAW variances, MASL distinguishes between left and right paralysis. Ablation studies and experiments on public and real-world datasets validate MASL's segmentation module and demonstrate its ability to provide reliable metrics for UVFP diagnosis.

【9】 Raw Speech Enhancement with Deep State Space Modeling
标题: 利用深状态空间建模增强原始语音
作者:Yan Ru Pei,Ritik Shrivastava,FNU Sidharth
备注:7 pages, 2 figures
链接:点击下载PDF文件
摘要:我们提出了aTENNuate,一个简单的深状态空间自动编码器配置为有效的在线原始语音增强在一个端到端的方式。该网络的性能主要在原始语音去噪方面进行评估,并对超分辨率和去量化等任务进行额外评估。我们在VoiceBank + DEMAND和Microsoft DNS1合成测试集上对aTENNuate进行了基准测试。该网络在PESQ分数、参数计数、MAC和延迟方面优于以前的实时去噪模型。即使作为原始波形处理模型,该模型也保持了对干净信号的高保真度,同时具有最小的可听伪影。此外,该模型保持性能,即使有噪声的输入被压缩到4000 Hz和4位,这表明在低资源环境中的一般语音增强能力。摘要:We present aTENNuate, a simple deep state-space autoencoder configured for efficient online raw speech enhancement in an end-to-end fashion. The network's performance is primarily evaluated on raw speech denoising, with additional assessments on tasks such as super-resolution and de-quantization. We benchmark aTENNuate on the VoiceBank + DEMAND and the Microsoft DNS1 synthetic test sets. The network outperforms previous real-time denoising models in terms of PESQ score, parameter count, MACs, and latency. Even as a raw waveform processing model, the model maintains high fidelity to the clean signal with minimal audible artifacts. In addition, the model remains performant even when the noisy input is compressed down to 4000Hz and 4 bits, suggesting general speech enhancement capabilities in low-resource environments.

【10】 Eetimating Indoor Scene Depth Maps from Ultrasonic Echoes
标题: 根据超声波回声确定室内场景深度图
作者:Junpei Honma,Akisato Kimura,Go Irie
备注:ICIP 2024
链接:点击下载PDF文件
摘要:测量室内场景的3D几何结构需要专用的深度传感器,但这并不总是可用的。基于回声的深度估计最近被研究为有前途的替代解决方案。所有以前的研究都假设使用的回声在可听范围内。然而,一个主要问题是可听回声不能用于安静的空间或禁止产生可听声音的其他情况。在本文中,我们考虑使用听不见的超声回波的回声为基础的深度估计。虽然超声波在理论上提供高的测量精度,但是由于其对噪声敏感并且易于衰减的缺点,当使用超声波回波时的实际深度估计精度仍然不清楚。我们首先调查的深度估计精度时,声源的频率被限制在高频带,并发现,精度下降时,频率被限制在超声波范围。基于这一观察,我们提出了一种新的深度学习方法,通过仅在训练期间使用可听回波作为辅助数据来提高基于超声回波的深度估计的准确性。在公开数据集上的实验结果表明,该方法提高了估计精度.摘要:Measuring 3D geometric structures of indoor scenes requires dedicated depth sensors, which are not always available. Echo-based depth estimation has recently been studied as a promising alternative solution. All previous studies have assumed the use of echoes in the audible range. However, one major problem is that audible echoes cannot be used in quiet spaces or other situations where producing audible sounds is prohibited. In this paper, we consider echo-based depth estimation using inaudible ultrasonic echoes. While ultrasonic waves provide high measurement accuracy in theory, the actual depth estimation accuracy when ultrasonic echoes are used has remained unclear, due to its disadvantage of being sensitive to noise and susceptible to attenuation. We first investigate the depth estimation accuracy when the frequency of the sound source is restricted to the high-frequency band, and found that the accuracy decreased when the frequency was limited to ultrasonic ranges. Based on this observation, we propose a novel deep learning method to improve the accuracy of ultrasonic echo-based depth estimation by using audible echoes as auxiliary data only during training. Experimental results with a public dataset demonstrate that our method improves the estimation accuracy.

【11】 FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
标题: FireRedRTS:用于行业级生成语音应用的基础文本到语音框架
作者:Hao-Han Guo,Kun Liu,Fei-Yu Shen,Yi-Chen Wu,Feng-Long Xie,Kun Xie,Kai-Tuo Xu
链接:点击下载PDF文件
摘要:这项工作提出了FireRedTTS,一个基础的文本到语音的框架,以满足日益增长的需求,个性化和多样化的生成语音应用程序。该框架包括三个部分:数据处理,基础系统和下游应用程序。首先,我们全面介绍了我们的数据处理管道,它将大量原始音频转换为具有丰富注释和广泛内容,说话风格和音色的大规模高质量TTS数据集。然后,我们提出了一个基于语言模型的基础TTS系统。语音信号通过语义感知的语音分词器被压缩成离散的语义令牌,并且可以由语言模型从提示文本和音频生成。然后,提出了一种两级波形发生器,将它们解码为高保真波形。我们提出了这个系统的两个应用程序:配音的语音克隆和聊天机器人的类人语音生成。实验结果表明,FireRedTTS具有较强的上下文学习能力,能够稳定地合成出与提示文本和音频一致的高质量语音。在配音方面,FireRedTTS可针对UGC场景以zero-shot方式克隆目标声音,并通过Few-Shot微调,1小时录音,适应PUGC场景中演播室级别的表现力声音角色。此外,FireRedTTS通过指令调整,以轻松的方式实现可控的类人语音生成,具有非语言行为和情感,以更好地服务于口语聊天机器人。摘要:This work proposes FireRedTTS, a foundation text-to-speech framework, to meet the growing demands for personalized and diverse generative speech applications. The framework comprises three parts: data processing, foundation system, and downstream applications. First, we comprehensively present our data processing pipeline, which transforms massive raw audio into a large-scale high-quality TTS dataset with rich annotations and a wide coverage of content, speaking style, and timbre. Then, we propose a language-model-based foundation TTS system. The speech signal is compressed into discrete semantic tokens via a semantic-aware speech tokenizer, and can be generated by a language model from the prompt text and audio. Then, a two-stage waveform generator is proposed to decode them to the high-fidelity waveform. We present two applications of this system: voice cloning for dubbing and human-like speech generation for chatbots. The experimental results demonstrate the solid in-context learning capability of FireRedTTS, which can stably synthesize high-quality speech consistent with the prompt text and audio. For dubbing, FireRedTTS can clone target voices in a zero-shot way for the UGC scenario and adapt to studio-level expressive voice characters in the PUGC scenario via few-shot fine-tuning with 1-hour recording. Moreover, FireRedTTS achieves controllable human-like speech generation in a casual style with paralinguistic behaviors and emotions via instruction tuning, to better serve spoken chatbots.

【12】 SymPAC: Scalable Symbolic Music Generation With Prompts And Constraints
标题: SymPAC:具有预设和约束的可扩展符号音乐生成
作者:Haonan Chen,Jordan B. L. Smith,Bochen Li,Ju-Chiang Wang,Janne Spijkervet,Pei Zou,Xingjian Du,Qiuqiang Kong
备注:ISMIR 2024
链接:点击下载PDF文件
摘要:符号音乐生成任务的进展可能落后于音频和文本生成等其他任务,部分原因是符号训练数据的稀缺。在本文中,我们通过应用预训练的MIR模型(用于转录,节拍跟踪,结构分析等)来利用更大规模的音频音乐数据。提取符号事件并将其编码为令牌序列。据我们所知,这项工作是第一次证明了仅从自动转录的音频数据训练符号生成模型的可行性。此外,为了增强训练模型的可控性,我们引入了Symbolic Music Language Model with Encoding And Constrained Generation(符号音乐语言模型),其区别在于:(a)在编码中使用提示条,(b)在推理时间使用一种称为通过有限状态机(FSM)的约束生成技术。我们展示了这种方法的灵活性和可控性,这对于使音乐AI对创作者和用户有用至关重要。摘要:Progress in the task of symbolic music generation may be lagging behind other tasks like audio and text generation, in part because of the scarcity of symbolic training data. In this paper, we leverage the greater scale of audio music data by applying pre-trained MIR models (for transcription, beat tracking, structure analysis, etc.) to extract symbolic events and encode them into token sequences. To the best of our knowledge, this work is the first to demonstrate the feasibility of training symbolic generation models solely from auto-transcribed audio data. Furthermore, to enhance the controllability of the trained model, we introduce SymPAC (Symbolic Music Language Model with Prompting And Constrained Generation), which is distinguished by using (a) prompt bars in encoding and (b) a technique called Constrained Generation via Finite State Machines (FSMs) during inference time. We show the flexibility and controllability of this approach, which may be critical in making music AI useful to creators and users.


机器翻译,仅供参考