本文经arXiv每日学术速递授权转载
【1】 Sylber: Syllabic Embedding Representation of Speech from Raw Audio
标题: Sylber:原始音频语音的音节嵌入表示
作者: Cheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal, Ethan Chen, Alan W Black, Gopala K. Anumanchipalli
链接:点击下载PDF文件
【2】 Spectral and Rhythm Features for Audio Classification with Deep Convolutional Neural Networks
标题: 利用深度卷积神经网络进行音频分类的频谱和节奏特征
作者: Friedrich Wolf-Monheim
链接:点击下载PDF文件
【3】 Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity
标题: 预训练语音和语言模型的联合微调和转换为线性复杂性
作者: Mutian He, Philip N. Garner
备注:15 pages, 4 figures
链接:点击下载PDF文件
【4】 Diffuse or Confuse: A Diffusion Deepfake Speech Dataset
标题: 扩散或混淆:扩散Deepfake语音数据集
作者: Anton Firc, Kamil Malinka, Petr Hanáček
备注:Presented at International Conference of the Biometrics Special Interest Group (BIOSIG 2024)
链接:点击下载PDF文件
【5】 SCOREQ: Speech Quality Assessment with Contrastive Regression
标题: Scrum:使用对比回归进行言语质量评估
作者: Alessandro Ragano, Jan Skoglund, Andrew Hines
备注:Accepted NeurIPS 2024
链接:点击下载PDF文件
【6】 Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS
标题: Bastards Harmony:用于Bastards文本到语音合成的综合数据集,采用EnGen-TTC的离散编解码器建模
作者: Onkar Kishor Susladkar, Vishesh Tripathi, Biddwan Ahmed
Journal-ref:EMNLP 2024
链接:点击下载PDF文件
【7】 Can DeepFake Speech be Reliably Detected?
标题: DeepFake Speech能被可靠检测到吗?
作者: Hongbin Liu, Youzheng Chen, Arun Narayanan, Athula Balachandran, Pedro J. Moreno, Lun Wang
链接:点击下载PDF文件
【8】 SRC-gAudio: Sampling-Rate-Controlled Audio Generation
标题: SRC-gAudio:采样率控制的音频生成
作者: Chenxing Li, Manjie Xu, Dong Yu
备注:Accepted by APSIPA2024
链接:点击下载PDF文件
【9】 Gumbel Rao Monte Carlo based Bi-Modal Neural Architecture Search for Audio-Visual Deepfake Detection
标题: Gumbel Rao Monte Carlo基于双模式神经架构搜索用于视听深度伪造检测
作者: Aravinda Reddy PN, Raghavendra Ramachandra, Krothapalli Sreenivasa Rao, Pabitra Mitra Vinod Rathod
链接:点击下载PDF文件
【10】 Mamba-based Segmentation Model for Speaker Diarization
标题: 基于Mamba的说话人数字化分割模型
作者: Alexis Plaquet, Naohiro Tawara, Marc Delcroix, Shota Horiguchi, Atsushi Ando, Shoko Araki
备注:5 pages, 4 figures. Submitted to ICASSP 2025. Code at this https URL
链接:点击下载PDF文件
【11】 POLIPHONE: A Dataset for Smartphone Model Identification from Audio Recordings
标题: POLIPHONE:从音频记录中识别智能手机型号的数据集
作者: Davide Salvi, Daniele Ugo Leonzio, Antonio Giganti, Claudio Eutizi, Sara Mandelli, Paolo Bestagini, Stefano Tubaro
备注:Submitted to IEEE Access
链接:点击下载PDF文件
【12】 Variable Bitrate Residual Vector Quantization for Audio Coding
标题: 音频编码中的可变比特率残留量量化
作者: Yunkee Chae, Woosung Choi, Yuhta Takida, Junghyun Koo, Yukara Ikemiya, Zhi Zhong, Kin Wai Cheuk, Marco A. Martínez-Ramírez, Kyogu Lee, Wei-Hsiang Liao, Yuki Mitsufuji
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【13】 FINALLY: fast and universal speech enhancement with studio-like quality
标题: 最终:快速、通用的语音增强,具有类似录音室的质量
作者: Nicholas Babaev, Kirill Tamogashev, Azat Saginbaev, Ivan Shchekotov, Hanbin Bae, Hosang Sung, WonJun Lee, Hoon-Young Cho, Pavel Andreev
备注:Accepted to NeurIPS 2024
链接:点击下载PDF文件
【14】 FürElise: Capturing and Physically Synthesizing Hand Motions of Piano Performance
标题: FürElise:捕捉和物理合成钢琴演奏的手部动作
作者: Ruocheng Wang, Pei Xu, Haochen Shi, Elizabeth Schumann, C. Karen Liu
备注:SIGGRAPH Asia 2024. Project page: this https URL
链接:点击下载PDF文件
【15】 Array2BR: An End-to-End Noise-immune Binaural Audio Synthesis from Microphone-array Signals
标题: Array 2BR:来自麦克风阵列信号的端到端抗噪双耳音频合成
作者: Cheng Chi, Xiaoyu Li, Andong Li, Yuxuan Ke, Xiaodong Li, Chengshi Zheng
链接:点击下载PDF文件
【16】 FGCL: Fine-grained Contrastive Learning For Mandarin Stuttering Event Detection
标题: FGCL:用于普通话口吃事件检测的细粒度对比学习
作者: Han Jiang, Wenyu Wang, Yiquan Zhou, Hongwu Ding, Jiacheng Xu, Jihua Zhu
备注:Accepted to SLT 2024
链接:点击下载PDF文件
【17】 Dynamic HumTrans: Humming Transcription Using CNNs and Dynamic Programming
标题: 动态HumTrans:使用CNN和动态编程的哼唱转录
作者: Shubham Gupta, Isaac Neri Gomez-Sarmiento, Faez Amjed Mezdari, Mirco Ravanelli, Cem Subakan
链接:点击下载PDF文件
【18】 Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
标题: 通过改进对抗环境中的语音识别来增强说话者身份援助
作者: Sagarika Alavilli, Annesya Banerjee, Gasser Elbanna, Annika Magaro
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【19】 RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction
标题: RespLLM:通过多模式LLM统一音频和文本,用于广义呼吸健康预测
作者: Yuwei Zhang, Tong Xia, Aaqib Saeed, Cecilia Mascolo
链接:点击下载PDF文件
【20】 Diffusion-based Unsupervised Audio-visual Speech Enhancement
标题: 基于扩散的无监督视听语音增强
作者: Jean-Eudes Ayilo (MULTISPEECH), Mostafa Sadeghi (MULTISPEECH), Romain Serizel (MULTISPEECH), Xavier Alameda-Pineda (ROBOTLEARN)
链接:点击下载PDF文件
【21】 F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
标题: F5-TTC:一款通过流匹配伪造流利、忠实语音的童话
作者: Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Chunhui Wang, Jian Zhao, Kai Yu, Xie Chen
链接:点击下载PDF文件
【22】 LS-EEND: Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction
标题: LS-EEND:具有在线吸引子提取的长格式流媒体端到端神经扩张
作者: Di Liang, Xiaofei Li
链接:点击下载PDF文件
【23】 An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment
标题: 以眼还眼:Zero-Shot音频描述使用视听分布对齐利用图像捕获器
作者: Hugo Malard, Michel Olvera, Stéphane Lathuiliere, Slim Essid
链接:点击下载PDF文件
【24】 The USTC-NERCSLIP Systems for the CHiME-8 MMCSG Challenge
标题: CHiME-8 MMCSG挑战赛的USTC-NERCSLIP系统
作者: Ya Jiang, Hongbo Lan, Jun Du, Qing Wang, Shutong Niu
链接:点击下载PDF文件
【25】 The OCON model: an old but gold solution for distributable supervised classification
标题: OCON模型:可分布式监督分类的古老但黄金解决方案
作者: Stefano Giacomelli, Marco Giordano, Claudia Rinaldi
备注:Accepted at "2024 29th IEEE Symposium on Computers and Communications (ISCC): workshop on Next-Generation Multimedia Services at the Edge: Leveraging 5G and Beyond (NGMSE2024)". arXiv admin note: text overlap with arXiv:2410.04098
链接:点击下载PDF文件
【26】 Episodic fine-tuning prototypical networks for optimization-based few-shot learning: Application to audio classification
标题: 基于优化的Few-Shot学习的情景式微调原型网络:应用于音频分类
作者: Xuanyu Zhuang (LTCI, IP Paris, S2A, IDS), Geoffroy Peeters (LTCI, IP Paris, S2A, IDS), Gaël Richard (S2A, IDS, LTCI, IP Paris)
备注:Accepted at MLSP 2024
Journal-ref:2024 IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2024), Sep 2024, London (UK), United Kingdom
链接:点击下载PDF文件
标题: F5-TTC:一款通过流匹配伪造流利、忠实语音的童话
作者: Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Chunhui Wang, Jian Zhao, Kai Yu, Xie Chen
链接:点击下载PDF文件
【2】 Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch
标题: 基于FastPitch的自然发音语音合成和说话人自适应的高效训练策略
作者: Teodora Răgman, Adriana Stan
备注:Accepted at 2024 IEEE 20th International Conference on Intelligent Computer Communication and Processing (ICCP 2024)
链接:点击下载PDF文件
【3】 LS-EEND: Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction
标题: LS-EEND:具有在线吸引子提取的长格式流媒体端到端神经扩张
作者: Di Liang, Xiaofei Li
链接:点击下载PDF文件
【4】 An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment
标题: 以眼还眼:Zero-Shot音频描述使用视听分布对齐利用图像捕获器
作者: Hugo Malard, Michel Olvera, Stéphane Lathuiliere, Slim Essid
链接:点击下载PDF文件
【5】 The USTC-NERCSLIP Systems for the CHiME-8 MMCSG Challenge
标题: CHiME-8 MMCSG挑战赛的USTC-NERCSLIP系统
作者: Ya Jiang, Hongbo Lan, Jun Du, Qing Wang, Shutong Niu
链接:点击下载PDF文件
【6】 Exploring rhythm formant analysis for Indic language classification
标题: 探讨印度语分类的节奏共振峰分析
作者: Parismita Gogoi, Sishir Kalita, Priyankoo Sarmah, S.R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【7】 Improving Data Augmentation-based Cross-Speaker Style Transfer for TTS with Singing Voice, Style Filtering, and F0 Matching
标题: 通过歌声、风格过滤和F0匹配改善基于数据增强的TTC交叉说话者风格转移
作者: Leonardo B. de M. M. Marques, Lucas H. Ueda, Mário U. Neto, Flávio O. Simões, Fernando Runstein, Bianca Dal Bó, Paula D. P. Costa
备注:Submitted to INTERSPEECH 2024
链接:点击下载PDF文件
【8】 The OCON model: an old but gold solution for distributable supervised classification
标题: OCON模型:可分布式监督分类的古老但黄金解决方案
作者: Stefano Giacomelli, Marco Giordano, Claudia Rinaldi
备注:Accepted at "2024 29th IEEE Symposium on Computers and Communications (ISCC): workshop on Next-Generation Multimedia Services at the Edge: Leveraging 5G and Beyond (NGMSE2024)". arXiv admin note: text overlap with arXiv:2410.04098
链接:点击下载PDF文件
【9】 Episodic fine-tuning prototypical networks for optimization-based few-shot learning: Application to audio classification
标题: 基于优化的Few-Shot学习的情景式微调原型网络:应用于音频分类
作者: Xuanyu Zhuang (LTCI, IP Paris, S2A, IDS), Geoffroy Peeters (LTCI, IP Paris, S2A, IDS), Gaël Richard (S2A, IDS, LTCI, IP Paris)
备注:Accepted at MLSP 2024
Journal-ref:2024 IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2024), Sep 2024, London (UK), United Kingdom
链接:点击下载PDF文件
【10】 Sylber: Syllabic Embedding Representation of Speech from Raw Audio
标题: Sylber:原始音频语音的音节嵌入表示
作者: Cheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal, Ethan Chen, Alan W Black, Gopala K. Anumanchipalli
链接:点击下载PDF文件
【11】 Spectral and Rhythm Features for Audio Classification with Deep Convolutional Neural Networks
标题: 利用深度卷积神经网络进行音频分类的频谱和节奏特征
作者: Friedrich Wolf-Monheim
链接:点击下载PDF文件
【12】 Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity
标题: 预训练语音和语言模型的联合微调和转换为线性复杂性
作者: Mutian He, Philip N. Garner
备注:15 pages, 4 figures
链接:点击下载PDF文件
【13】 SCOREQ: Speech Quality Assessment with Contrastive Regression
标题: Scrum:使用对比回归进行言语质量评估
作者: Alessandro Ragano, Jan Skoglund, Andrew Hines
备注:Accepted NeurIPS 2024
链接:点击下载PDF文件
【14】 Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS
标题: Bastards Harmony:用于Bastards文本到语音合成的综合数据集,采用EnGen-TTC的离散编解码器建模
作者: Onkar Kishor Susladkar, Vishesh Tripathi, Biddwan Ahmed
Journal-ref:EMNLP 2024
链接:点击下载PDF文件
【15】 SRC-gAudio: Sampling-Rate-Controlled Audio Generation
标题: SRC-gAudio:采样率控制的音频生成
作者: Chenxing Li, Manjie Xu, Dong Yu
备注:Accepted by APSIPA2024
链接:点击下载PDF文件
【16】 Gumbel Rao Monte Carlo based Bi-Modal Neural Architecture Search for Audio-Visual Deepfake Detection
标题: Gumbel Rao Monte Carlo基于双模式神经架构搜索用于视听深度伪造检测
作者: Aravinda Reddy PN, Raghavendra Ramachandra, Krothapalli Sreenivasa Rao, Pabitra Mitra Vinod Rathod
链接:点击下载PDF文件
【17】 Mamba-based Segmentation Model for Speaker Diarization
标题: 基于Mamba的说话人数字化分割模型
作者: Alexis Plaquet, Naohiro Tawara, Marc Delcroix, Shota Horiguchi, Atsushi Ando, Shoko Araki
备注:5 pages, 4 figures. Submitted to ICASSP 2025. Code at this https URL
链接:点击下载PDF文件
【18】 POLIPHONE: A Dataset for Smartphone Model Identification from Audio Recordings
标题: POLIPHONE:从音频记录中识别智能手机型号的数据集
作者: Davide Salvi, Daniele Ugo Leonzio, Antonio Giganti, Claudio Eutizi, Sara Mandelli, Paolo Bestagini, Stefano Tubaro
备注:Submitted to IEEE Access
链接:点击下载PDF文件
【19】 Variable Bitrate Residual Vector Quantization for Audio Coding
标题: 音频编码中的可变比特率残留量量化
作者: Yunkee Chae, Woosung Choi, Yuhta Takida, Junghyun Koo, Yukara Ikemiya, Zhi Zhong, Kin Wai Cheuk, Marco A. Martínez-Ramírez, Kyogu Lee, Wei-Hsiang Liao, Yuki Mitsufuji
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【20】 FINALLY: fast and universal speech enhancement with studio-like quality
标题: 最终:快速、通用的语音增强,具有类似录音室的质量
作者: Nicholas Babaev, Kirill Tamogashev, Azat Saginbaev, Ivan Shchekotov, Hanbin Bae, Hosang Sung, WonJun Lee, Hoon-Young Cho, Pavel Andreev
备注:Accepted to NeurIPS 2024
链接:点击下载PDF文件
【21】 FürElise: Capturing and Physically Synthesizing Hand Motions of Piano Performance
标题: FürElise:捕捉和物理合成钢琴演奏的手部动作
作者: Ruocheng Wang, Pei Xu, Haochen Shi, Elizabeth Schumann, C. Karen Liu
备注:SIGGRAPH Asia 2024. Project page: this https URL
链接:点击下载PDF文件
【22】 Array2BR: An End-to-End Noise-immune Binaural Audio Synthesis from Microphone-array Signals
标题: Array 2BR:来自麦克风阵列信号的端到端抗噪双耳音频合成
作者: Cheng Chi, Xiaoyu Li, Andong Li, Yuxuan Ke, Xiaodong Li, Chengshi Zheng
链接:点击下载PDF文件
【23】 FGCL: Fine-grained Contrastive Learning For Mandarin Stuttering Event Detection
标题: FGCL:用于普通话口吃事件检测的细粒度对比学习
作者: Han Jiang, Wenyu Wang, Yiquan Zhou, Hongwu Ding, Jiacheng Xu, Jihua Zhu
备注:Accepted to SLT 2024
链接:点击下载PDF文件
【24】 Dynamic HumTrans: Humming Transcription Using CNNs and Dynamic Programming
标题: 动态HumTrans:使用CNN和动态编程的哼唱转录
作者: Shubham Gupta, Isaac Neri Gomez-Sarmiento, Faez Amjed Mezdari, Mirco Ravanelli, Cem Subakan
链接:点击下载PDF文件
【25】 Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
标题: 通过改进对抗环境中的语音识别来增强说话者身份援助
作者: Sagarika Alavilli, Annesya Banerjee, Gasser Elbanna, Annika Magaro
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【26】 RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction
标题: RespLLM:通过多模式LLM统一音频和文本,用于广义呼吸健康预测
作者: Yuwei Zhang, Tong Xia, Aaqib Saeed, Cecilia Mascolo
链接:点击下载PDF文件
【27】 Diffusion-based Unsupervised Audio-visual Speech Enhancement
标题: 基于扩散的无监督视听语音增强
作者: Jean-Eudes Ayilo (MULTISPEECH), Mostafa Sadeghi (MULTISPEECH), Romain Serizel (MULTISPEECH), Xavier Alameda-Pineda (ROBOTLEARN)
链接:点击下载PDF文件
标题: Sylber:原始音频语音的音节嵌入表示
作者: Cheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal, Ethan Chen, Alan W Black, Gopala K. Anumanchipalli
链接:点击下载PDF文件
摘要:音节是口语的组成单位,在人类的言语感知和产生中起着至关重要的作用。然而,目前的神经语音表示缺乏结构,导致处理成本高的密集令牌序列。为了弥合这一差距,我们提出了一个新的模型,Sylber,产生干净和强大的音节结构的语音表示。具体来说,我们提出了一个自监督模型,该模型对从教师模型中提取的音节片段进行回归,该教师模型是训练中模型的指数移动平均值。这导致了语音特征的高度结构化表示,提供了三个主要优点:1)快速的线性时间音节分割算法,2)平均每秒4.27个标记的有效音节标记化,以及3)更适合词汇和句法理解的音节单位。我们还训练令牌到语音生成模型与我们的音节单位,并表明,完全可理解的语音可以从这些令牌重建。最后,我们观察到分类感知,一种语音感知的语言现象,在我们的模型中自然出现,使得嵌入空间比以前的自监督学习方法更分类和稀疏。总之,我们提出了一种新的自我监督的方法来表示语音作为音节,具有显着的潜力,有效的语音标记和口语建模。摘要:Syllables are compositional units of spoken language that play a crucial role in human speech perception and production. However, current neural speech representations lack structure, resulting in dense token sequences that are costly to process. To bridge this gap, we propose a new model, Sylber, that produces speech representations with clean and robust syllabic structure. Specifically, we propose a self-supervised model that regresses features on syllabic segments distilled from a teacher model which is an exponential moving average of the model in training. This results in a highly structured representation of speech features, offering three key benefits: 1) a fast, linear-time syllable segmentation algorithm, 2) efficient syllabic tokenization with an average of 4.27 tokens per second, and 3) syllabic units better suited for lexical and syntactic understanding. We also train token-to-speech generative models with our syllabic units and show that fully intelligible speech can be reconstructed from these tokens. Lastly, we observe that categorical perception, a linguistic phenomenon of speech perception, emerges naturally in our model, making the embedding space more categorical and sparse than previous self-supervised learning approaches. Together, we present a novel self-supervised approach for representing speech as syllables, with significant potential for efficient speech tokenization and spoken language modeling.
【2】 Spectral and Rhythm Features for Audio Classification with Deep Convolutional Neural Networks
标题: 利用深度卷积神经网络进行音频分类的频谱和节奏特征
作者: Friedrich Wolf-Monheim
链接:点击下载PDF文件
摘要:卷积神经网络(CNN)广泛应用于计算机视觉。它们不仅可以用于传统的数字图像材料来识别模式,而且还可以用于从数字图像中提取特征,这些特征表示从时域数字音频信号中提取的频谱和节奏特征,用于声音的声学分类。不同的频谱和节奏特征表示,如梅尔缩放频谱图,梅尔频率倒谱系数(MFCC),循环tempogram,短时傅里叶变换(STFT)chromagram,恒定Q变换(CQT)chromagram和色度能量归一化统计(CENS)chromagram的音频分类性能使用深度卷积神经网络进行了研究。可以清楚地表明,对于使用深度CNN的音频分类任务,梅尔缩放频谱图和梅尔频率倒谱系数(MFCC)的表现明显优于本研究中研究的其他频谱和节奏特征。实验是在ESC-50数据集的帮助下进行的,该数据集包含2,000个标记的环境音频记录。摘要:Convolutional neural networks (CNNs) are widely used in computer vision. They can be used not only for conventional digital image material to recognize patterns, but also for feature extraction from digital imagery representing spectral and rhythm features extracted from time-domain digital audio signals for the acoustic classification of sounds. Different spectral and rhythm feature representations like mel-scaled spectrograms, mel-frequency cepstral coefficients (MFCCs), cyclic tempograms, short-time Fourier transform (STFT) chromagrams, constant-Q transform (CQT) chromagrams and chroma energy normalized statistics (CENS) chromagrams are investigated in terms of the audio classification performance using a deep convolutional neural network. It can be clearly shown that the mel-scaled spectrograms and the mel-frequency cepstral coefficients (MFCCs) perform significantly better then the other spectral and rhythm features investigated in this research for audio classification tasks using deep CNNs. The experiments were carried out with the aid of the ESC-50 dataset with 2,000 labeled environmental audio recordings.
【3】 Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity
标题: 预训练语音和语言模型的联合微调和转换为线性复杂性
作者: Mutian He, Philip N. Garner
备注:15 pages, 4 figures
链接:点击下载PDF文件
摘要:Linformer和Mamba等架构最近已经成为Transformers的有竞争力的线性时间替代品。然而,相应的大型预训练模型通常不可用,特别是在非文本领域。为了解决这个问题,我们提出了一个跨架构分层蒸馏(CALD)的方法,共同转换一个Transformer模型的线性时间替代和微调它的目标任务。我们还比较了几种方法来指导微调,以最佳地保留原始模型所需的推理能力。这些方法的不同之处在于它们对目标模型和参数轨迹的使用。在一系列关于语言处理、语言建模和语音处理的实证研究中,我们证明了CALD能够有效地恢复原始模型的结果,并且引导策略对结果有贡献。一些变化的原因提出了建议。摘要:Architectures such as Linformer and Mamba have recently emerged as competitive linear time replacements for transformers. However, corresponding large pretrained models are often unavailable, especially in non-text domains. To remedy this, we present a Cross-Architecture Layerwise Distillation (CALD) approach that jointly converts a transformer model to a linear time substitute and fine-tunes it to a target task. We also compare several means to guide the fine-tuning to optimally retain the desired inference capability from the original model. The methods differ in their use of the target model and the trajectory of the parameters. In a series of empirical studies on language processing, language modeling, and speech processing, we show that CALD can effectively recover the result of the original model, and that the guiding strategy contributes to the result. Some reasons for the variation are suggested.
【4】 Diffuse or Confuse: A Diffusion Deepfake Speech Dataset
标题: 扩散或混淆:扩散Deepfake语音数据集
作者: Anton Firc, Kamil Malinka, Petr Hanáček
备注:Presented at International Conference of the Biometrics Special Interest Group (BIOSIG 2024)
链接:点击下载PDF文件
摘要:人工智能和机器学习的进步显着改善了合成语音生成。本文探讨了扩散模型,一种新的方法来创建逼真的合成语音。我们使用可用的工具和预训练的模型创建扩散数据集。此外,该研究还评估了扩散生成的deepfake与非扩散生成的deepfake的质量,以及它们对当前deepfake检测系统的潜在威胁。研究结果表明,基于扩散的deepfake的检测通常与非扩散deepfake相当,但基于检测器架构存在一些差异。使用扩散声码器的重新声码显示出最小的影响,并且整体语音质量与非扩散方法相当。摘要:Advancements in artificial intelligence and machine learning have significantly improved synthetic speech generation. This paper explores diffusion models, a novel method for creating realistic synthetic speech. We create a diffusion dataset using available tools and pretrained models. Additionally, this study assesses the quality of diffusion-generated deepfakes versus non-diffusion ones and their potential threat to current deepfake detection systems. Findings indicate that the detection of diffusion-based deepfakes is generally comparable to non-diffusion deepfakes, with some variability based on detector architecture. Re-vocoding with diffusion vocoders shows minimal impact, and the overall speech quality is comparable to non-diffusion methods.
【5】 SCOREQ: Speech Quality Assessment with Contrastive Regression
标题: Scrum:使用对比回归进行言语质量评估
作者: Alessandro Ragano, Jan Skoglund, Andrew Hines
备注:Accepted NeurIPS 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了一种新的语音质量预测的方法,SCOPIC。SCORK是用于对比回归的三元组损失函数,其解决了由现有技术的无参考语音质量度量所表现出的域泛化缺点。在本文中,我们:(i)说明L2损失训练未能捕获平均意见评分(MOS)标签的连续性的问题;(ii)通过跨多个语音域的基准评估证明缺乏概括性;(iii)概述我们的方法,并通过增量评估探索架构设计决策的影响;(iv)对照各种数据和领域的最新模型评估最终模型。结果表明,缺乏概括性观察到的最先进的语音质量指标的状态是由SCORIO解决。我们的结论是,使用三重损失函数的对比回归提高了语音质量预测模型的泛化,但也有潜在的效用,在广泛的应用,使用基于回归的预测模型。摘要:In this paper, we present SCOREQ, a novel approach for speech quality prediction. SCOREQ is a triplet loss function for contrastive regression that addresses the domain generalisation shortcoming exhibited by state of the art no-reference speech quality metrics. In the paper we: (i) illustrate the problem of L2 loss training failing at capturing the continuous nature of the mean opinion score (MOS) labels; (ii) demonstrate the lack of generalisation through a benchmarking evaluation across several speech domains; (iii) outline our approach and explore the impact of the architectural design decisions through incremental evaluation; (iv) evaluate the final model against state of the art models for a wide variety of data and domains. The results show that the lack of generalisation observed in state of the art speech quality metrics is addressed by SCOREQ. We conclude that using a triplet loss function for contrastive regression improves generalisation for speech quality prediction models but also has potential utility across a wide range of applications using regression-based predictive models.
【6】 Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS
标题: Bastards Harmony:用于Bastards文本到语音合成的综合数据集,采用EnGen-TTC的离散编解码器建模
作者: Onkar Kishor Susladkar, Vishesh Tripathi, Biddwan Ahmed
Journal-ref:EMNLP 2024
链接:点击下载PDF文件
摘要:本研究介绍了一个全面的Baidu文本到语音(TTS)数据集和一个新的TTS模型,EnGen TTS,旨在提高Baidu语言合成语音的质量和通用性。该数据集涵盖55.0小时和52 K音频记录,集成了不同的文本来源,确保了语言的丰富性。一个细致的录音设置捕捉到的细微差别巴托语音,采用专业设备,以确保高保真的音频样本。统计分析揭示了数据集的规模和多样性,为模型训练和评估奠定了基础。拟议的EnGen-TTS模型比既定的基线表现更好,实现了平均意见得分(MOS)为4.45 $ pm $0.13。此外,我们对实时因素和模型大小的调查突出了EnGen-TTS作为一个引人注目的选择,具有高效的性能。这项研究标志着Baidu TTS技术的重大进步,对不同的语言应用产生了影响。生成的样本链接: url{https: bahasa-harmony-comp.vercel.app }摘要:This research introduces a comprehensive Bahasa text-to-speech (TTS) dataset and a novel TTS model, EnGen-TTS, designed to enhance the quality and versatility of synthetic speech in the Bahasa language. The dataset, spanning textasciitilde55.0 hours and 52K audio recordings, integrates diverse textual sources, ensuring linguistic richness. A meticulous recording setup captures the nuances of Bahasa phonetics, employing professional equipment to ensure high-fidelity audio samples. Statistical analysis reveals the dataset's scale and diversity, laying the foundation for model training and evaluation. The proposed EnGen-TTS model performs better than established baselines, achieving a Mean Opinion Score (MOS) of 4.45 $ pm$ 0.13. Additionally, our investigation on real-time factor and model size highlights EnGen-TTS as a compelling choice, with efficient performance. This research marks a significant advancement in Bahasa TTS technology, with implications for diverse language applications. Link to Generated Samples: url{https: bahasa-harmony-comp.vercel.app }
【7】 Can DeepFake Speech be Reliably Detected?
标题: DeepFake Speech能被可靠检测到吗?
作者: Hongbin Liu, Youzheng Chen, Arun Narayanan, Athula Balachandran, Pedro J. Moreno, Lun Wang
链接:点击下载PDF文件
摘要:文本到语音(TTS)系统的最新进展,特别是那些具有语音克隆功能的系统,已经使语音模拟变得容易获得,由于可能被滥用于恶意活动(如错误信息活动和欺诈),引起了道德和法律方面的担忧。虽然合成语音检测器(SSD)的存在是为了解决这一问题,但它们容易受到“测试域偏移”的影响,当音频通过转码、回放或背景噪声被改变时,它们的性能会下降。故意操纵合成语音以欺骗检测器,进一步加剧了这种脆弱性。这项工作首次系统地研究了针对最先进的开源SSD的这种主动恶意攻击。使用硬编码指标和人工评级,从攻击有效性和隐蔽性两个方面研究了白盒攻击、黑盒攻击及其可转移性。研究结果强调,面对不断变化的对抗性威胁,迫切需要更强大的检测方法。摘要:Recent advances in text-to-speech (TTS) systems, particularly those with voice cloning capabilities, have made voice impersonation readily accessible, raising ethical and legal concerns due to potential misuse for malicious activities like misinformation campaigns and fraud. While synthetic speech detectors (SSDs) exist to combat this, they are vulnerable to test domain shift", exhibiting decreased performance when audio is altered through transcoding, playback, or background noise. This vulnerability is further exacerbated by deliberate manipulation of synthetic speech aimed at deceiving detectors. This work presents the first systematic study of such active malicious attacks against state-of-the-art open-source SSDs. White-box attacks, black-box attacks, and their transferability are studied from both attack effectiveness and stealthiness, using both hardcoded metrics and human ratings. The results highlight the urgent need for more robust detection methods in the face of evolving adversarial threats.
【8】 SRC-gAudio: Sampling-Rate-Controlled Audio Generation
标题: SRC-gAudio:采样率控制的音频生成
作者: Chenxing Li, Manjie Xu, Dong Yu
备注:Accepted by APSIPA2024
链接:点击下载PDF文件
摘要:我们介绍SRC-gAudio,一种新颖的音频生成模型,旨在促进文本到音频生成在一个单一的模型架构内的广泛的采样率。SRC-gAudio将采样率作为生成条件的一部分,以指导基于扩散的音频生成过程。我们的模型使音频生成在多个采样率与一个单一的统一模型。此外,我们探讨了大规模,低采样率的数据在提高高采样率音频的生成质量的潜在好处。通过大量的实验,我们证明了SRC-gAudio有效地生成音频控制采样率下。此外,我们的研究结果表明,对低采样率数据进行预训练可以显著提高各种指标的音频质量。摘要:We introduce SRC-gAudio, a novel audio generation model designed to facilitate text-to-audio generation across a wide range of sampling rates within a single model architecture. SRC-gAudio incorporates the sampling rate as part of the generation condition to guide the diffusion-based audio generation process. Our model enables the generation of audio at multiple sampling rates with a single unified model. Furthermore, we explore the potential benefits of large-scale, low-sampling-rate data in enhancing the generation quality of high-sampling-rate audio. Through extensive experiments, we demonstrate that SRC-gAudio effectively generates audio under controlled sampling rates. Additionally, our results indicate that pre-training on low-sampling-rate data can lead to significant improvements in audio quality across various metrics.
【9】 Gumbel Rao Monte Carlo based Bi-Modal Neural Architecture Search for Audio-Visual Deepfake Detection
标题: Gumbel Rao Monte Carlo基于双模式神经架构搜索用于视听深度伪造检测
作者: Aravinda Reddy PN, Raghavendra Ramachandra, Krothapalli Sreenivasa Rao, Pabitra Mitra Vinod Rathod
链接:点击下载PDF文件
摘要:Deepfakes通过生成高度逼真的合成媒体对生物识别认证系统构成了严重威胁。现有的多模态deepfake检测器通常难以适应不同的数据,并依赖于简单的融合方法。为了解决这些挑战,我们提出了Gumbel-Rao Monte Carlo双峰神经架构搜索(GRMC-BMNAS),一种新的架构搜索框架,采用Gumbel-Rao Monte Carlo采样来优化多模态融合。它通过Rao-Blackwellization减少方差,稳定网络训练,改进了直通Gumbel Softmax(STGS)方法。使用两级搜索方法,该框架优化了网络结构,参数和性能。关键功能有效地识别骨干网络,而在细胞结构内,加权融合操作集成来自各种来源的信息。通过改变参数,如温度和蒙特卡罗样本的数量,产生一个架构,最大限度地提高分类性能和更好的泛化能力。在FakeAVCeleb和SWAN-DF数据集上的实验结果表明,用最小的模型参数实现了令人印象深刻的95.4%的AUC百分比。摘要:Deepfakes pose a critical threat to biometric authentication systems by generating highly realistic synthetic media. Existing multimodal deepfake detectors often struggle to adapt to diverse data and rely on simple fusion methods. To address these challenges, we propose Gumbel-Rao Monte Carlo Bi-modal Neural Architecture Search (GRMC-BMNAS), a novel architecture search framework that employs Gumbel-Rao Monte Carlo sampling to optimize multimodal fusion. It refines the Straight through Gumbel Softmax (STGS) method by reducing variance with Rao-Blackwellization, stabilizing network training. Using a two-level search approach, the framework optimizes the network architecture, parameters, and performance. Crucial features are efficiently identified from backbone networks, while within the cell structure, a weighted fusion operation integrates information from various sources. By varying parameters such as temperature and number of Monte carlo samples yields an architecture that maximizes classification performance and better generalisation capability. Experimental results on the FakeAVCeleb and SWAN-DF datasets demonstrate an impressive AUC percentage of 95.4 %, achieved with minimal model parameters.
【10】 Mamba-based Segmentation Model for Speaker Diarization
标题: 基于Mamba的说话人数字化分割模型
作者: Alexis Plaquet, Naohiro Tawara, Marc Delcroix, Shota Horiguchi, Atsushi Ando, Shoko Araki
备注:5 pages, 4 figures. Submitted to ICASSP 2025. Code at this https URL
链接:点击下载PDF文件
摘要:Mamba是一种新提出的架构,其行为类似于具有类似注意力能力的递归神经网络(RNN)。这些特性对于说话人日记化来说是有希望的,因为基于注意力的模型对长格式音频有不合适的记忆要求,并且传统的RNN能力太有限。在本文中,我们建议通过比较pyannote.audio管道的最先进的神经分割与我们提出的基于Mamba的变体来评估Mamba的日记化潜力。Mamba更强大的处理能力允许使用更长的局部窗口,这通过使说话人嵌入提取更可靠来显着提高日志质量。我们发现Mamba是传统RNN和经过测试的基于注意力的模型的更好的替代方案。我们提出的基于Mamba的系统在三个广泛使用的日志数据集上实现了最先进的性能。摘要:Mamba is a newly proposed architecture which behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization, as attention-based models have unsuitable memory requirements for long-form audio, and traditional RNN capabilities are too limited. In this paper, we propose to assess the potential of Mamba for diarization by comparing the state-of-the-art neural segmentation of the pyannote.audio pipeline with our proposed Mamba-based variant. Mamba's stronger processing capabilities allow usage of longer local windows, which significantly improve diarization quality by making the speaker embedding extraction more reliable. We find Mamba to be a superior alternative to both traditional RNN and the tested attention-based model. Our proposed Mamba-based system achieves state-of-the-art performance on three widely used diarization datasets.
【11】 POLIPHONE: A Dataset for Smartphone Model Identification from Audio Recordings
标题: POLIPHONE:从音频记录中识别智能手机型号的数据集
作者: Davide Salvi, Daniele Ugo Leonzio, Antonio Giganti, Claudio Eutizi, Sara Mandelli, Paolo Bestagini, Stefano Tubaro
备注:Submitted to IEEE Access
链接:点击下载PDF文件
摘要:在处理多媒体数据时,从取证的角度来看,源属性是一个关键的挑战。该任务旨在确定给定内容是如何被捕获的,为各种应用提供有价值的见解,包括法律诉讼和诚信调查。来源归属问题已经在不同的领域得到了解决,从识别用于捕获特定照片的相机模型到检测用于创建或记录给定音轨的合成语音发生器或麦克风模型。该领域的最新进展在很大程度上依赖于机器学习和数据驱动技术,这些技术通常优于传统的基于信号处理的方法。 然而,这些系统的缺点是它们需要大量的训练数据,这些数据必须反映最新的技术趋势,以产生准确可靠的预测。这是一个重大挑战,因为技术进步的快速步伐使得很难保持数据集与现实世界的条件保持一致。例如,在根据音频记录识别智能手机型号的任务中,可用的数据集通常已经过时或获取不一致,因此很难开发出在研究环境之外有效的解决方案。在本文中,我们介绍了POLIPHONE,这是一个用于从音频记录中识别智能手机型号的数据集。它包括在受控环境中记录的20个最新智能手机的数据,以确保未来研究的可重复性和可扩展性。所释放的轨道包含来自各个域的音频数据(即,语音、音乐、环境声音),使语料库具有通用性并适用于广泛的用例。我们还提出了许多实验,使用最先进的分类器从音频记录中识别智能手机模型,对所提出的数据集进行基准测试。摘要:When dealing with multimedia data, source attribution is a key challenge from a forensic perspective. This task aims to determine how a given content was captured, providing valuable insights for various applications, including legal proceedings and integrity investigations. The source attribution problem has been addressed in different domains, from identifying the camera model used to capture specific photographs to detecting the synthetic speech generator or microphone model used to create or record given audio tracks. Recent advancements in this area rely heavily on machine learning and data-driven techniques, which often outperform traditional signal processing-based methods. However, a drawback of these systems is their need for large volumes of training data, which must reflect the latest technological trends to produce accurate and reliable predictions. This presents a significant challenge, as the rapid pace of technological progress makes it difficult to maintain datasets that are up-to-date with real-world conditions. For instance, in the task of smartphone model identification from audio recordings, the available datasets are often outdated or acquired inconsistently, making it difficult to develop solutions that are valid beyond a research environment. In this paper we present POLIPHONE, a dataset for smartphone model identification from audio recordings. It includes data from 20 recent smartphones recorded in a controlled environment to ensure reproducibility and scalability for future research. The released tracks contain audio data from various domains (i.e., speech, music, environmental sounds), making the corpus versatile and applicable to a wide range of use cases. We also present numerous experiments to benchmark the proposed dataset using a state-of-the-art classifier for smartphone model identification from audio recordings.
【12】 Variable Bitrate Residual Vector Quantization for Audio Coding
标题: 音频编码中的可变比特率残留量量化
作者: Yunkee Chae, Woosung Choi, Yuhta Takida, Junghyun Koo, Yukara Ikemiya, Zhi Zhong, Kin Wai Cheuk, Marco A. Martínez-Ramírez, Kyogu Lee, Wei-Hsiang Liao, Yuki Mitsufuji
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:最近的最先进的神经音频压缩模型已逐步采用残差矢量量化(RVQ)。尽管取得了这一成功,但这些模型每帧采用固定数量的码本,这在速率失真权衡方面可能是次优的,特别是在具有简单输入音频的场景中,例如静音。为了解决这一限制,我们提出了可变比特率RVQ(VRVQ)的音频编解码器,它允许更有效的编码,通过调整每帧使用的码本的数量。此外,我们提出了一种梯度估计方法,用于从重要性图转换为二进制重要性掩码的不可微掩码操作,通过直通估计器改进模型训练。我们证明了所提出的训练框架相比基线方法取得了更好的效果,并显示出进一步的改进时,应用到当前最先进的编解码器。摘要:Recent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be suboptimal in terms of rate-distortion tradeoff, particularly in scenarios with simple input audio, such as silence. To address this limitation, we propose variable bitrate RVQ (VRVQ) for audio codecs, which allows for more efficient coding by adapting the number of codebooks used per frame. Furthermore, we propose a gradient estimation method for the non-differentiable masking operation that transforms from the importance map to the binary importance mask, improving model training via a straight-through estimator. We demonstrate that the proposed training framework achieves superior results compared to the baseline method and shows further improvement when applied to the current state-of-the-art codec.
【13】 FINALLY: fast and universal speech enhancement with studio-like quality
标题: 最终:快速、通用的语音增强,具有类似录音室的质量
作者: Nicholas Babaev, Kirill Tamogashev, Azat Saginbaev, Ivan Shchekotov, Hanbin Bae, Hosang Sung, WonJun Lee, Hoon-Young Cho, Pavel Andreev
备注:Accepted to NeurIPS 2024
链接:点击下载PDF文件
摘要:在本文中,我们解决了现实世界的录音,其中往往包含各种形式的失真,如背景噪声,混响和麦克风文物语音增强的挑战。我们重新审视了生成对抗网络(GANs)用于语音增强的使用,并从理论上表明,GANs自然倾向于在有条件的干净语音分布中寻找最大密度点,正如我们所认为的,这对于语音增强任务至关重要。我们研究了感知损失的各种特征提取器,以促进对抗训练的稳定性,开发了一种探测特征空间结构的方法。这促使我们将基于WavLM的感知损失集成到MS-STFT对抗训练管道中,为语音增强模型创建有效且稳定的训练过程。由此产生的语音增强模型,我们称之为FINALLY,建立在HiFi++架构之上,使用WavLM编码器和一个新的训练管道进行增强。各种数据集上的实证结果证实了我们的模型能够在48 kHz下产生清晰,高质量的语音,在语音增强领域实现最先进的性能。摘要:In this paper, we address the challenge of speech enhancement in real-world recordings, which often contain various forms of distortion, such as background noise, reverberation, and microphone artifacts. We revisit the use of Generative Adversarial Networks (GANs) for speech enhancement and theoretically show that GANs are naturally inclined to seek the point of maximum density within the conditional clean speech distribution, which, as we argue, is essential for the speech enhancement task. We study various feature extractors for perceptual loss to facilitate the stability of adversarial training, developing a methodology for probing the structure of the feature space. This leads us to integrate WavLM-based perceptual loss into MS-STFT adversarial training pipeline, creating an effective and stable training procedure for the speech enhancement model. The resulting speech enhancement model, which we refer to as FINALLY, builds upon the HiFi++ architecture, augmented with a WavLM encoder and a novel training pipeline. Empirical results on various datasets confirm our model's ability to produce clear, high-quality speech at 48 kHz, achieving state-of-the-art performance in the field of speech enhancement.
【14】 FürElise: Capturing and Physically Synthesizing Hand Motions of Piano Performance
标题: FürElise:捕捉和物理合成钢琴演奏的手部动作
作者: Ruocheng Wang, Pei Xu, Haochen Shi, Elizabeth Schumann, C. Karen Liu
备注:SIGGRAPH Asia 2024. Project page: this https URL
链接:点击下载PDF文件
摘要:钢琴演奏需要敏捷,精确和协调的手控制,延伸灵活性的极限。手部运动模型具有精确再现钢琴演奏的复杂性,在角色动画,体现AI,生物力学和VR AR中有广泛的应用。在本文中,我们构建了一个首个大型数据集,其中包含大约10个小时的3D手部运动和来自15位精英级钢琴家演奏153首古典音乐的音频。为了捕捉自然的表演,我们设计了一个无标记的设置,其中使用最先进的姿态估计模型从多视图视频中重建运动。运动数据通过使用从专门的雅马哈电钢琴中的传感器获得的高分辨率按键数据的逆运动学进一步细化。利用收集到的数据集,我们开发了一个流水线,可以为数据集之外的乐谱合成物理上合理的手部动作。我们的方法采用模仿学习和强化学习的组合,以获得基于物理的双手控制策略,涉及手和钢琴键之间的相互作用。为了解决大运动数据集的采样效率问题,我们使用扩散模型来生成自然的参考运动,提供高层次的轨迹和指法(手指顺序和放置)信息。然而,单独生成的参考运动不能为钢琴演奏建模提供足够的精度。然后,我们通过使用音乐相似性来从捕获的数据集中检索相似的运动来进一步增强数据,以提高RL策略的精度。使用所提出的方法,我们的模型生成自然,灵巧的运动,从训练数据集之外推广到音乐。摘要:Piano playing requires agile, precise, and coordinated hand control that stretches the limits of dexterity. Hand motion models with the sophistication to accurately recreate piano playing have a wide range of applications in character animation, embodied AI, biomechanics, and VR AR. In this paper, we construct a first-of-its-kind large-scale dataset that contains approximately 10 hours of 3D hand motion and audio from 15 elite-level pianists playing 153 pieces of classical music. To capture natural performances, we designed a markerless setup in which motions are reconstructed from multi-view videos using state-of-the-art pose estimation models. The motion data is further refined via inverse kinematics using the high-resolution MIDI key-pressing data obtained from sensors in a specialized Yamaha Disklavier piano. Leveraging the collected dataset, we developed a pipeline that can synthesize physically-plausible hand motions for musical scores outside of the dataset. Our approach employs a combination of imitation learning and reinforcement learning to obtain policies for physics-based bimanual control involving the interaction between hands and piano keys. To solve the sampling efficiency problem with the large motion dataset, we use a diffusion model to generate natural reference motions, which provide high-level trajectory and fingering (finger order and placement) information. However, the generated reference motion alone does not provide sufficient accuracy for piano performance modeling. We then further augmented the data by using musical similarity to retrieve similar motions from the captured dataset to boost the precision of the RL policy. With the proposed method, our model generates natural, dexterous motions that generalize to music from outside the training dataset.
【15】 Array2BR: An End-to-End Noise-immune Binaural Audio Synthesis from Microphone-array Signals
标题: Array 2BR:来自麦克风阵列信号的端到端抗噪双耳音频合成
作者: Cheng Chi, Xiaoyu Li, Andong Li, Yuxuan Ke, Xiaodong Li, Chengshi Zheng
链接:点击下载PDF文件
摘要:临场感技术旨在为远程会议应用提供沉浸式的虚拟临场感,为此,合成高质量的双耳音频信号是极其重要的。由于在实际应用场景中,环境噪声往往是不可避免的,因此迫切希望能够直接从麦克风阵列信号中获得没有噪声的双耳音频信号。为此,本文提出了一种新的基于麦克风阵列信号的端到端抗噪声双耳音频合成框架Array2BR,实验结果表明,该框架能够在正确映射双耳线索的同时很好地抑制噪声。与现有的方法相比,该方法取得了更好的性能,在客观和主观的度量分数。摘要:Telepresence technology aims to provide an immersive virtual presence for remote conference applications, and it is extremely important to synthesize high-quality binaural audio signals for this aim. Because the ambient noise is often inevitable in practical application scenarios, it is highly desired that binaural audio signals without noise can be obtained from microphone-array signals directly. For this purpose, this paper proposes a new end-to-end noise-immune binaural audio synthesis framework from microphone-array signals, abbreviated as Array2BR, and experimental results show that binaural cues can be correctly mapped and noise can be well suppressed simultaneously using the proposed framework. Compared with existing methods, the proposed method achieved better performance in terms of both objective and subjective metric scores.
【16】 FGCL: Fine-grained Contrastive Learning For Mandarin Stuttering Event Detection
标题: FGCL:用于普通话口吃事件检测的细粒度对比学习
作者: Han Jiang, Wenyu Wang, Yiquan Zhou, Hongwu Ding, Jiacheng Xu, Jihua Zhu
备注:Accepted to SLT 2024
链接:点击下载PDF文件
摘要:本文介绍了T031团队在SLT2024中应对口吃语音挑战的方法。汉语口吃事件检测(MSED)的目的是检测汉语语音中的口吃事件。我们提出了一个详细的声学分析方法,以提高口吃检测的准确性,通过捕捉微妙的细微差别,以前口吃事件检测(SED)技术忽略了。为此,我们介绍了MSED的细粒度对比学习(FGCL)框架。具体来说,我们模型的口吃事件的帧级概率,并引入一个挖掘算法来识别容易和混乱的帧。然后,我们提出了一个口吃对比度损失,以提高口吃和流畅的语音帧之间的区别,从而提高口吃特征嵌入的判别能力。在英语和普通话数据集上的广泛评估证明了FGCL的有效性,在普通话数据上实现了超过5.0%的F1分数的显着提高。摘要:This paper presents the T031 team's approach to the StutteringSpeech Challenge in SLT2024. Mandarin Stuttering Event Detection (MSED) aims to detect instances of stuttering events in Mandarin speech. We propose a detailed acoustic analysis method to improve the accuracy of stutter detection by capturing subtle nuances that previous Stuttering Event Detection (SED) techniques have overlooked. To this end, we introduce the Fine-Grained Contrastive Learning (FGCL) framework for MSED. Specifically, we model the frame-level probabilities of stuttering events and introduce a mining algorithm to identify both easy and confusing frames. Then, we propose a stutter contrast loss to enhance the distinction between stuttered and fluent speech frames, thereby improving the discriminative capability of stuttered feature embeddings. Extensive evaluations on English and Mandarin datasets demonstrate the effectiveness of FGCL, achieving a significant increase of over 5.0% in F1 score on Mandarin data.
【17】 Dynamic HumTrans: Humming Transcription Using CNNs and Dynamic Programming
标题: 动态HumTrans:使用CNN和动态编程的哼唱转录
作者: Shubham Gupta, Isaac Neri Gomez-Sarmiento, Faez Amjed Mezdari, Mirco Ravanelli, Cem Subakan
链接:点击下载PDF文件
摘要:我们提出了一种新的哼唱转录方法,结合了基于CNN的架构与基于动态编程的后处理算法,利用最近推出的HumTrans数据集。我们识别并解决了数据集提供的偏移和起始地面实况的固有问题,提供了改进这些注释的方法,从而产生了具有精确注释的数据集,这将有助于未来的研究。此外,我们将我们的方法的转录准确性与其他几种方法进行了比较,证明了最先进的(SOTA)结果。我们所有的代码和校正数据集都可以在https: github.com shubham-gupta-30 humming_transcription上找到摘要:We propose a novel approach for humming transcription that combines a CNN-based architecture with a dynamic programming-based post-processing algorithm, utilizing the recently introduced HumTrans dataset. We identify and address inherent problems with the offset and onset ground truth provided by the dataset, offering heuristics to improve these annotations, resulting in a dataset with precise annotations that will aid future research. Additionally, we compare the transcription accuracy of our method against several others, demonstrating state-of-the-art (SOTA) results. All our code and corrected dataset is available at https: github.com shubham-gupta-30 humming_transcription
【18】 Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
标题: 通过改进对抗环境中的语音识别来增强说话者身份援助
作者: Sagarika Alavilli, Annesya Banerjee, Gasser Elbanna, Annika Magaro
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:当前最先进的语音识别模型被训练成将声学信号映射到子词汇单元。虽然这些模型表现出优越的性能,但它们仍然容易受到背景噪声和语音增强等分布条件的影响。在这项工作中,我们假设在语音识别过程中将说话人表示可以增强模型对噪声的鲁棒性。我们开发了一个基于transformer的模型,共同执行语音识别和说话人识别。我们的模型利用来自Whisper的语音嵌入和来自ECAPA-TDNN的扬声器嵌入,它们被联合处理以执行这两项任务。我们表明,在清洁条件下,联合模型执行的耳语。值得注意的是,联合模型在高噪声环境中优于Whisper,例如具有8扬声器的串音背景噪声。此外,我们的联合模型擅长处理高度增强的语音,包括正弦波和噪声声码语音。总的来说,这些结果表明,将语音表示与语音识别相结合可以在对抗条件下产生更强大的模型。摘要:Current state-of-the-art speech recognition models are trained to map acoustic signals into sub-lexical units. While these models demonstrate superior performance, they remain vulnerable to out-of-distribution conditions such as background noise and speech augmentations. In this work, we hypothesize that incorporating speaker representations during speech recognition can enhance model robustness to noise. We developed a transformer-based model that jointly performs speech recognition and speaker identification. Our model utilizes speech embeddings from Whisper and speaker embeddings from ECAPA-TDNN, which are processed jointly to perform both tasks. We show that the joint model performs comparably to Whisper under clean conditions. Notably, the joint model outperforms Whisper in high-noise environments, such as with 8-speaker babble background noise. Furthermore, our joint model excels in handling highly augmented speech, including sine-wave and noise-vocoded speech. Overall, these results suggest that integrating voice representations with speech recognition can lead to more robust models under adversarial conditions.
【19】 RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction
标题: RespLLM:通过多模式LLM统一音频和文本,用于广义呼吸健康预测
作者: Yuwei Zhang, Tong Xia, Aaqib Saeed, Cecilia Mascolo
链接:点击下载PDF文件
摘要:与呼吸道疾病相关的高发病率和死亡率强调了早期筛查的重要性。机器学习模型可以自动化临床咨询和听诊,为这一领域提供重要支持。然而,所涉及的数据,包括人口统计学、病史、症状和呼吸音,是异质和复杂的。现有的方法是不够的,缺乏普遍性,因为它们通常依赖于有限的训练数据,基本的融合技术和特定于任务的模型。在本文中,我们提出了RespLLM,一种新的多模态大语言模型(LLM)框架,它将文本和音频表示统一起来,用于呼吸健康预测。RespLLM利用预训练LLM的广泛先验知识,并通过跨模态注意力实现有效的音频-文本融合。采用指令调优来集成来自多个源的不同数据,确保模型的通用性和通用性。在五个真实世界数据集上的实验表明,RespLLM在训练任务上的平均性能优于领先基线4.6%,在看不见的数据集上的平均性能优于领先基线7.9%,并有助于对新任务进行zero-shot预测。我们的工作为多模态模型奠定了基础,这些模型可以感知、倾听和理解异构数据,为可扩展的呼吸健康诊断铺平了道路。摘要:The high incidence and mortality rates associated with respiratory diseases underscores the importance of early screening. Machine learning models can automate clinical consultations and auscultation, offering vital support in this area. However, the data involved, spanning demographics, medical history, symptoms, and respiratory audio, are heterogeneous and complex. Existing approaches are insufficient and lack generalizability, as they typically rely on limited training data, basic fusion techniques, and task-specific models. In this paper, we propose RespLLM, a novel multimodal large language model (LLM) framework that unifies text and audio representations for respiratory health prediction. RespLLM leverages the extensive prior knowledge of pretrained LLMs and enables effective audio-text fusion through cross-modal attentions. Instruction tuning is employed to integrate diverse data from multiple sources, ensuring generalizability and versatility of the model. Experiments on five real-world datasets demonstrate that RespLLM outperforms leading baselines by an average of 4.6% on trained tasks, 7.9% on unseen datasets, and facilitates zero-shot predictions for new tasks. Our work lays the foundation for multimodal models that can perceive, listen to, and understand heterogeneous data, paving the way for scalable respiratory health diagnosis.
【20】 Diffusion-based Unsupervised Audio-visual Speech Enhancement
标题: 基于扩散的无监督视听语音增强
作者: Jean-Eudes Ayilo (MULTISPEECH), Mostafa Sadeghi (MULTISPEECH), Romain Serizel (MULTISPEECH), Xavier Alameda-Pineda (ROBOTLEARN)
链接:点击下载PDF文件
摘要:本文提出了一种新的无监督视听语音增强(AVSE)方法,结合了基于扩散的视听语音生成模型与非负矩阵分解(NMF)噪声模型。首先,在以相应视频数据为条件的干净语音上预训练扩散模型以模拟语音生成分布。然后,将该预训练模型与基于NMF的噪声模型配对,以迭代地估计干净的语音。具体地,基于扩散的后验采样方法在逆扩散过程中实现,其中在每次迭代之后,获得语音估计并用于更新噪声参数。实验结果证实,所提出的AVSE方法不仅优于其音频只对应,但也比最近的监督生成AVSE方法更好地推广。此外,与以前的基于扩散的方法相比,新的推理算法在推理速度和性能之间提供了更好的平衡。摘要:This paper proposes a new unsupervised audiovisual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. First, the diffusion model is pre-trained on clean speech conditioned on corresponding video data to simulate the speech generative distribution. This pre-trained model is then paired with the NMF-based noise model to iteratively estimate clean speech. Specifically, a diffusion-based posterior sampling approach is implemented within the reverse diffusion process, where after each iteration, a speech estimate is obtained and used to update the noise parameters. Experimental results confirm that the proposed AVSE approach not only outperforms its audio-only counterpart but also generalizes better than a recent supervisedgenerative AVSE method. Additionally, the new inference algorithm offers a better balance between inference speed and performance compared to the previous diffusion-based method.
【21】 F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
标题: F5-TTC:一款通过流匹配伪造流利、忠实语音的童话
作者: Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Chunhui Wang, Jian Zhao, Kai Yu, Xie Chen
链接:点击下载PDF文件
摘要:本文介绍了一种基于扩散Transformer(DiT)流匹配的完全非自回归的文语转换系统F5-TTS。不需要复杂的设计,如持续时间模型,文本编码器和音素对齐,文本输入简单地填充填充与输入语音相同的长度,然后进行去噪以生成语音,这最初被证明是可行的E2 TTS。然而,E2 TTS的原始设计由于其收敛速度慢和鲁棒性低而难以遵循。为了解决这些问题,我们首先使用ConvNeXt对输入进行建模,以优化文本表示,使其易于与语音对齐。我们进一步提出了一个推理时间的摇摆采样策略,这显着提高了我们的模型的性能和效率。这种流步采样策略可以很容易地应用于现有的基于流匹配的模型,而无需重新训练。我们的设计允许更快的训练,并实现了0.15的推理RTF,与最先进的基于扩散的TTS模型相比,这是一个很大的改进。在一个公开的10万小时的多语言数据集上训练,我们的Fairytaler Fakes Fluent and Faithful Speech with Flow matching(F5-TTS)表现出高度自然和表达的zero-shot能力,无缝的代码切换能力和速度控制效率。演示示例可以在https: SWivid.github.io F5-TTS上找到。我们发布所有代码和检查点,以促进社区发展。摘要:This paper introduces F5-TTS, a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT). Without requiring complex designs such as duration model, text encoder, and phoneme alignment, the text input is simply padded with filler tokens to the same length as input speech, and then the denoising is performed for speech generation, which was originally proved feasible by E2 TTS. However, the original design of E2 TTS makes it hard to follow due to its slow convergence and low robustness. To address these issues, we first model the input with ConvNeXt to refine the text representation, making it easy to align with the speech. We further propose an inference-time Sway Sampling strategy, which significantly improves our model's performance and efficiency. This sampling strategy for flow step can be easily applied to existing flow matching based models without retraining. Our design allows faster training and achieves an inference RTF of 0.15, which is greatly improved compared to state-of-the-art diffusion-based TTS models. Trained on a public 100K hours multilingual dataset, our Fairytaler Fakes Fluent and Faithful speech with Flow matching (F5-TTS) exhibits highly natural and expressive zero-shot ability, seamless code-switching capability, and speed control efficiency. Demo samples can be found at https: SWivid.github.io F5-TTS. We release all code and checkpoints to promote community development.
【22】 LS-EEND: Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction
标题: LS-EEND:具有在线吸引子提取的长格式流媒体端到端神经扩张
作者: Di Liang, Xiaofei Li
链接:点击下载PDF文件
摘要:本文提出了一种逐帧在线 流式端到端神经日志化(EEND)方法,该方法以帧内帧外的方式检测说话人活动。该模型主要由因果嵌入编码器和在线吸引子解码器组成。在基于自注意力的解码器中沿着时间和扬声器维度对扬声器进行建模,并且分别为新扬声器和现有扬声器自动生成和更新逐帧扬声器吸引子。采用保留机制,特别适用于具有线性时间复杂度的长形式日记化。提出了一种多步渐进式训练策略,根据说话人数量和音频长度从简单任务逐渐学习到困难任务。最后,所提出的模型(称为长格式流EEND,LS-EEND)能够执行流日记高(高达8)和灵活的数量的发言者和非常长的(说一个小时)的音频记录。在各种模拟数据集和真实数据集上的实验表明:1)在不使用oracle语音活动信息的情况下,该模型在所有数据集(包括CALLHOME)上实现了新的最先进的在线日志错误率(12.11%),DIHARD II(27.58%),DIHARD III(19.61%),AMI(20.76%); 2)由于帧入帧出处理方式和线性时间复杂度,该模型的实时系数比比较在线日记模型低几倍。摘要:This work proposes a frame-wise online streaming end-to-end neural diarization (EEND) method, which detects speaker activities in a frame-in-frame-out fashion. The proposed model mainly consists of a causal embedding encoder and an online attractor decoder. Speakers are modeled in the self-attention-based decoder along both the time and speaker dimensions, and frame-wise speaker attractors are automatically generated and updated for new speakers and existing speakers, respectively. Retention mechanism is employed and especially adapted for long-form diarization with a linear temporal complexity. A multi-step progressive training strategy is proposed for gradually learning from easy tasks to hard tasks in terms of the number of speakers and audio length. Finally, the proposed model (referred to as long-form streaming EEND, LS-EEND) is able to perform streaming diarization for a high (up to 8) and flexible number speakers and very long (say one hour) audio recordings. Experiments on various simulated and real-world datasets show that: 1) when not using oracle speech activity information, the proposed model achieves new state-of-the-art online diarization error rate on all datasets, including CALLHOME (12.11%), DIHARD II (27.58%), DIHARD III (19.61%), and AMI (20.76%); 2) Due to the frame-in-frame-out processing fashion and the linear temporal complexity, the proposed model achieves several times lower real-time-factor than comparison online diarization models.
【23】 An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment
标题: 以眼还眼:Zero-Shot音频描述使用视听分布对齐利用图像捕获器
作者: Hugo Malard, Michel Olvera, Stéphane Lathuiliere, Slim Essid
链接:点击下载PDF文件
摘要:多模态大型语言模型推动了图像字幕的发展。这些模型在庞大的图像数据集上进行了微调,表现出对语义概念的深刻理解。在这项工作中,我们表明,这种能力可以重新用于音频字幕,其中联合图像语言解码器可以用来描述与视频中的图像序列相关的视听内容。这可以通过多模式对齐来实现。然而,由于真实世界视频中的可听元素和可见元素之间的固有差异,这种多模态对齐任务是不平凡的。此外,多模态表征学习往往依赖于对比学习,面临着所谓的模态差距的挑战,阻碍了模态之间的顺利整合。在这项工作中,我们介绍了一种新的方法,通过匹配的分布令牌产生的音频骨干和图像字幕弥合视听模态的差距。我们的方法将音频令牌分布与图像令牌的分布对齐,使模型能够以无监督的方式执行zero-shot音频字幕,同时保持初始图像字幕组件不变。这种对准允许通过将图像编码器与对准的音频编码器组合或替换来使用音频或视听输入。与现有方法相比,我们的方法在zero-shot音频字幕中实现了显着改善的性能。摘要:Multimodal large language models have fueled progress in image captioning. These models, fine-tuned on vast image datasets, exhibit a deep understanding of semantic concepts. In this work, we show that this ability can be re-purposed for audio captioning, where the joint image-language decoder can be leveraged to describe auditory content associated with image sequences within videos featuring audiovisual content. This can be achieved via multimodal alignment. Yet, this multimodal alignment task is non-trivial due to the inherent disparity between audible and visible elements in real-world videos. Moreover, multimodal representation learning often relies on contrastive learning, facing the challenge of the so-called modality gap which hinders smooth integration between modalities. In this work, we introduce a novel methodology for bridging the audiovisual modality gap by matching the distributions of tokens produced by an audio backbone and those of an image captioner. Our approach aligns the audio token distribution with that of the image tokens, enabling the model to perform zero-shot audio captioning in an unsupervised fashion while keeping the initial image captioning component unaltered. This alignment allows for the use of either audio or audiovisual input by combining or substituting the image encoder with the aligned audio encoder. Our method achieves significantly improved performances in zero-shot audio captioning, compared to existing approaches.
【24】 The USTC-NERCSLIP Systems for the CHiME-8 MMCSG Challenge
标题: CHiME-8 MMCSG挑战赛的USTC-NERCSLIP系统
作者: Ya Jiang, Hongbo Lan, Jun Du, Qing Wang, Shutong Niu
链接:点击下载PDF文件
摘要:在一个人戴着智能眼镜的两人对话场景中,实时转录和显示说话者的内容是一个有趣的应用,为后续任务(如翻译和理解)提供先验信息。同时,从智能眼镜捕获的多模态数据是稀缺的。因此,我们提出利用具有多个重叠率的模拟数据和一对一的匹配训练策略来缩小真实数据和模拟数据之间的模型训练偏差。此外,在模型中结合IMU单元数据可以辅助音频实现更好的实时语音识别性能。摘要:In the two-person conversation scenario with one wearing smart glasses, transcribing and displaying the speaker's content in real-time is an intriguing application, providing a priori information for subsequent tasks such as translation and comprehension. Meanwhile, multi-modal data captured from the smart glasses is scarce. Therefore, we propose utilizing simulation data with multiple overlap rates and a one-to-one matching training strategy to narrow down the deviation for the model training between real and simulated data. In addition, combining IMU unit data in the model can assist the audio to achieve better real-time speech recognition performance.
【25】 The OCON model: an old but gold solution for distributable supervised classification
标题: OCON模型:可分布式监督分类的古老但黄金解决方案
作者: Stefano Giacomelli, Marco Giordano, Claudia Rinaldi
备注:Accepted at "2024 29th IEEE Symposium on Computers and Communications (ISCC): workshop on Next-Generation Multimedia Services at the Edge: Leveraging 5G and Beyond (NGMSE2024)". arXiv admin note: text overlap with arXiv:2410.04098
链接:点击下载PDF文件
摘要:本文介绍了一个结构化的应用程序的一类方法和一类一网络模型的监督分类任务,特别是在自动语音识别研究领域的元音音素分类的案例研究。通过伪神经架构搜索和超参数调整实验进行了知情的网格搜索方法,我们实现了分类准确率与当今复杂的架构(90.0 - 93.7%)。尽管它的简单性,我们的模型优先推广的语言环境和分布式的适用性,支持相关的统计和性能指标。实验代码在我们的GitHub上公开提供。摘要:This paper introduces to a structured application of the One-Class approach and the One-Class-One-Network model for supervised classification tasks, specifically addressing a vowel phonemes classification case study within the Automatic Speech Recognition research field. Through pseudo-Neural Architecture Search and Hyper-Parameters Tuning experiments conducted with an informed grid-search methodology, we achieve classification accuracy comparable to nowadays complex architectures (90.0 - 93.7%). Despite its simplicity, our model prioritizes generalization of language context and distributed applicability, supported by relevant statistical and performance metrics. The experiments code is openly available at our GitHub.
【26】 Episodic fine-tuning prototypical networks for optimization-based few-shot learning: Application to audio classification
标题: 基于优化的Few-Shot学习的情景式微调原型网络:应用于音频分类
作者: Xuanyu Zhuang (LTCI, IP Paris, S2A, IDS), Geoffroy Peeters (LTCI, IP Paris, S2A, IDS), Gaël Richard (S2A, IDS, LTCI, IP Paris)
备注:Accepted at MLSP 2024
Journal-ref:2024 IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2024), Sep 2024, London (UK), United Kingdom
链接:点击下载PDF文件
摘要:Prototypical Network(ProtoNet)以其卓越的性能和简单的实现方式成为Few-Shot Learning(FSL)场景中的热门选择。在此成功的基础上,我们首先提出了一种简单(但新颖)的方法来在C-way-K-shot测试集的测试集(标记)支持集上微调ProtoNet(不使用仅用于评估的查询集)。然后,我们提出了一个算法框架,结合ProtoNet与基于优化的FSL算法(MAML和Meta-Curvature)与这样的微调方法。由于基于优化的算法赋予目标学习模型快速适应少数样本的能力,我们利用ProtoNet作为目标模型,以提高其微调性能的帮助下,一个专门设计的情节微调策略。实验结果证实,我们提出的模型,MAML-Proto和MC-Proto,结合我们独特的微调方法,在ESC-50和Speech Commands v2数据集上的Few-Shot音频分类任务中,性能大大优于常规ProtoNet。我们注意到,虽然我们只将我们的模型应用于音频域,但它是一种通用的方法,可以很容易地扩展到其他领域。摘要:The Prototypical Network (ProtoNet) has emerged as a popular choice in Few-shot Learning (FSL) scenarios due to its remarkable performance and straightforward implementation. Building upon such success, we first propose a simple (yet novel) method to fine-tune a ProtoNet on the (labeled) support set of the test episode of a C-way-K-shot test episode (without using the query set which is only used for evaluation). We then propose an algorithmic framework that combines ProtoNet with optimization-based FSL algorithms (MAML and Meta-Curvature) to work with such a fine-tuning method. Since optimization-based algorithms endow the target learner model with the ability to fast adaption to only a few samples, we utilize ProtoNet as the target model to enhance its fine-tuning performance with the help of a specifically designed episodic fine-tuning strategy. The experimental results confirm that our proposed models, MAML-Proto and MC-Proto, combined with our unique fine-tuning method, outperform regular ProtoNet by a large margin in few-shot audio classification tasks on the ESC-50 and Speech Commands v2 datasets. We note that although we have only applied our model to the audio domain, it is a general method and can be easily extended to other domains.
eess.AS音频处理
【1】 F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching标题: F5-TTC:一款通过流匹配伪造流利、忠实语音的童话
作者: Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Chunhui Wang, Jian Zhao, Kai Yu, Xie Chen
链接:点击下载PDF文件
摘要:本文介绍了一种基于扩散Transformer(DiT)流匹配的完全非自回归的文语转换系统F5-TTS。不需要复杂的设计,如持续时间模型,文本编码器和音素对齐,文本输入简单地填充填充与输入语音相同的长度,然后进行去噪以生成语音,这最初被证明是可行的E2 TTS。然而,E2 TTS的原始设计由于其收敛速度慢和鲁棒性低而难以遵循。为了解决这些问题,我们首先使用ConvNeXt对输入进行建模,以优化文本表示,使其易于与语音对齐。我们进一步提出了一个推理时间的摇摆采样策略,这显着提高了我们的模型的性能和效率。这种流步采样策略可以很容易地应用于现有的基于流匹配的模型,而无需重新训练。我们的设计允许更快的训练,并实现了0.15的推理RTF,与最先进的基于扩散的TTS模型相比,这是一个很大的改进。在一个公开的10万小时的多语言数据集上训练,我们的Fairytaler Fakes Fluent and Faithful Speech with Flow matching(F5-TTS)表现出高度自然和表达的zero-shot能力,无缝的代码切换能力和速度控制效率。演示示例可以在https: SWivid.github.io F5-TTS上找到。我们发布所有代码和检查点,以促进社区发展。摘要:This paper introduces F5-TTS, a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT). Without requiring complex designs such as duration model, text encoder, and phoneme alignment, the text input is simply padded with filler tokens to the same length as input speech, and then the denoising is performed for speech generation, which was originally proved feasible by E2 TTS. However, the original design of E2 TTS makes it hard to follow due to its slow convergence and low robustness. To address these issues, we first model the input with ConvNeXt to refine the text representation, making it easy to align with the speech. We further propose an inference-time Sway Sampling strategy, which significantly improves our model's performance and efficiency. This sampling strategy for flow step can be easily applied to existing flow matching based models without retraining. Our design allows faster training and achieves an inference RTF of 0.15, which is greatly improved compared to state-of-the-art diffusion-based TTS models. Trained on a public 100K hours multilingual dataset, our Fairytaler Fakes Fluent and Faithful speech with Flow matching (F5-TTS) exhibits highly natural and expressive zero-shot ability, seamless code-switching capability, and speed control efficiency. Demo samples can be found at https: SWivid.github.io F5-TTS. We release all code and checkpoints to promote community development.
【2】 Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch
标题: 基于FastPitch的自然发音语音合成和说话人自适应的高效训练策略
作者: Teodora Răgman, Adriana Stan
备注:Accepted at 2024 IEEE 20th International Conference on Intelligent Computer Communication and Processing (ICCP 2024)
链接:点击下载PDF文件
摘要:本文的重点是适应罗马尼亚语言的FastPitch模型的功能,从一个到十八个扬声器的扩展集,合成语音使用匿名身份,并复制新的,看不见的扬声器的身份。在这项工作中,各种配置和培训策略的效果进行了测试和讨论,以及他们的优点和缺点。最后,我们确定了一个新的配置,建立在FastPitch架构之上,能够为已知(来自训练数据集的身份)和未知(通过短参考样本学习的身份)扬声器生成自然语音合成。匿名说话者可以用于文本到语音合成,如果一个人想要取消身份信息,同时保持语义内容完整和清晰。最后,我们讨论了我们的工作可能存在的局限性,这将成为未来研究和改进的基础。摘要:This paper focuses on adapting the functionalities of the FastPitch model to the Romanian language; extending the set of speakers from one to eighteen; synthesising speech using an anonymous identity; and replicating the identities of new, unseen speakers. During this work, the effects of various configurations and training strategies were tested and discussed, along with their advantages and weaknesses. Finally, we settled on a new configuration, built on top of the FastPitch architecture, capable of producing natural speech synthesis, for both known (identities from the training dataset) and unknown (identities learnt through short reference samples) speakers. The anonymous speaker can be used for text-to-speech synthesis, if one wants to cancel out the identity information while keeping the semantic content whole and clear. At last, we discussed possible limitations of our work, which will form the basis for future investigations and advancements.
【3】 LS-EEND: Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction
标题: LS-EEND:具有在线吸引子提取的长格式流媒体端到端神经扩张
作者: Di Liang, Xiaofei Li
链接:点击下载PDF文件
摘要:本文提出了一种逐帧在线 流式端到端神经日志化(EEND)方法,该方法以帧内帧外的方式检测说话人活动。该模型主要由因果嵌入编码器和在线吸引子解码器组成。在基于自注意力的解码器中沿着时间和扬声器维度对扬声器进行建模,并且分别为新扬声器和现有扬声器自动生成和更新逐帧扬声器吸引子。采用保留机制,特别适用于具有线性时间复杂度的长形式日记化。提出了一种多步渐进式训练策略,根据说话人数量和音频长度从简单任务逐渐学习到困难任务。最后,所提出的模型(称为长格式流EEND,LS-EEND)能够执行流日记高(高达8)和灵活的数量的发言者和非常长的(说一个小时)的音频记录。在各种模拟数据集和真实数据集上的实验表明:1)在不使用oracle语音活动信息的情况下,该模型在所有数据集(包括CALLHOME)上实现了新的最先进的在线日志错误率(12.11%),DIHARD II(27.58%),DIHARD III(19.61%),AMI(20.76%); 2)由于帧入帧出处理方式和线性时间复杂度,该模型的实时系数比比较在线日记模型低几倍。摘要:This work proposes a frame-wise online streaming end-to-end neural diarization (EEND) method, which detects speaker activities in a frame-in-frame-out fashion. The proposed model mainly consists of a causal embedding encoder and an online attractor decoder. Speakers are modeled in the self-attention-based decoder along both the time and speaker dimensions, and frame-wise speaker attractors are automatically generated and updated for new speakers and existing speakers, respectively. Retention mechanism is employed and especially adapted for long-form diarization with a linear temporal complexity. A multi-step progressive training strategy is proposed for gradually learning from easy tasks to hard tasks in terms of the number of speakers and audio length. Finally, the proposed model (referred to as long-form streaming EEND, LS-EEND) is able to perform streaming diarization for a high (up to 8) and flexible number speakers and very long (say one hour) audio recordings. Experiments on various simulated and real-world datasets show that: 1) when not using oracle speech activity information, the proposed model achieves new state-of-the-art online diarization error rate on all datasets, including CALLHOME (12.11%), DIHARD II (27.58%), DIHARD III (19.61%), and AMI (20.76%); 2) Due to the frame-in-frame-out processing fashion and the linear temporal complexity, the proposed model achieves several times lower real-time-factor than comparison online diarization models.
【4】 An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment
标题: 以眼还眼:Zero-Shot音频描述使用视听分布对齐利用图像捕获器
作者: Hugo Malard, Michel Olvera, Stéphane Lathuiliere, Slim Essid
链接:点击下载PDF文件
摘要:多模态大型语言模型推动了图像字幕的发展。这些模型在庞大的图像数据集上进行了微调,表现出对语义概念的深刻理解。在这项工作中,我们表明,这种能力可以重新用于音频字幕,其中联合图像语言解码器可以用来描述与视频中的图像序列相关的视听内容。这可以通过多模式对齐来实现。然而,由于真实世界视频中的可听元素和可见元素之间的固有差异,这种多模态对齐任务是不平凡的。此外,多模态表征学习往往依赖于对比学习,面临着所谓的模态差距的挑战,阻碍了模态之间的顺利整合。在这项工作中,我们介绍了一种新的方法,通过匹配的分布令牌产生的音频骨干和图像字幕弥合视听模态的差距。我们的方法将音频令牌分布与图像令牌的分布对齐,使模型能够以无监督的方式执行zero-shot音频字幕,同时保持初始图像字幕组件不变。这种对准允许通过将图像编码器与对准的音频编码器组合或替换来使用音频或视听输入。与现有方法相比,我们的方法在zero-shot音频字幕中实现了显着改善的性能。摘要:Multimodal large language models have fueled progress in image captioning. These models, fine-tuned on vast image datasets, exhibit a deep understanding of semantic concepts. In this work, we show that this ability can be re-purposed for audio captioning, where the joint image-language decoder can be leveraged to describe auditory content associated with image sequences within videos featuring audiovisual content. This can be achieved via multimodal alignment. Yet, this multimodal alignment task is non-trivial due to the inherent disparity between audible and visible elements in real-world videos. Moreover, multimodal representation learning often relies on contrastive learning, facing the challenge of the so-called modality gap which hinders smooth integration between modalities. In this work, we introduce a novel methodology for bridging the audiovisual modality gap by matching the distributions of tokens produced by an audio backbone and those of an image captioner. Our approach aligns the audio token distribution with that of the image tokens, enabling the model to perform zero-shot audio captioning in an unsupervised fashion while keeping the initial image captioning component unaltered. This alignment allows for the use of either audio or audiovisual input by combining or substituting the image encoder with the aligned audio encoder. Our method achieves significantly improved performances in zero-shot audio captioning, compared to existing approaches.
【5】 The USTC-NERCSLIP Systems for the CHiME-8 MMCSG Challenge
标题: CHiME-8 MMCSG挑战赛的USTC-NERCSLIP系统
作者: Ya Jiang, Hongbo Lan, Jun Du, Qing Wang, Shutong Niu
链接:点击下载PDF文件
摘要:在一个人戴着智能眼镜的两人对话场景中,实时转录和显示说话者的内容是一个有趣的应用,为后续任务(如翻译和理解)提供先验信息。同时,从智能眼镜捕获的多模态数据是稀缺的。因此,我们提出利用具有多个重叠率的模拟数据和一对一的匹配训练策略来缩小真实数据和模拟数据之间的模型训练偏差。此外,在模型中结合IMU单元数据可以辅助音频实现更好的实时语音识别性能。摘要:In the two-person conversation scenario with one wearing smart glasses, transcribing and displaying the speaker's content in real-time is an intriguing application, providing a priori information for subsequent tasks such as translation and comprehension. Meanwhile, multi-modal data captured from the smart glasses is scarce. Therefore, we propose utilizing simulation data with multiple overlap rates and a one-to-one matching training strategy to narrow down the deviation for the model training between real and simulated data. In addition, combining IMU unit data in the model can assist the audio to achieve better real-time speech recognition performance.
【6】 Exploring rhythm formant analysis for Indic language classification
标题: 探讨印度语分类的节奏共振峰分析
作者: Parismita Gogoi, Sishir Kalita, Priyankoo Sarmah, S.R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:本文报告了一个初步的研究定量频域节奏线索分类五种印度语言:孟加拉语,卡纳达语,马拉雅拉姆语,马拉地语,泰米尔语。我们采用节奏共振峰(R-共振峰)的分析,Gills介绍的一种技术,利用低频频谱分析的幅度调制和频率调制包络来表征语音节奏。从LF频谱计算各种测量,包括R-共振峰、基于离散余弦变换的测量和频谱测量。结果表明,基于阈值和频谱特征优于直接计算的R-共振峰。从低频声谱图中提取的节奏时间模式提供了更好的语言辨别线索。结合所有衍生的功能,我们实现了69.21%的准确率和加权F1得分为69.18%的分类五种语言。这项研究表明,RFA在表征语音节奏的印度语言分类的潜力。摘要:This paper reports a preliminary study on quantitative frequency domain rhythm cues for classifying five Indian languages: Bengali, Kannada, Malayalam, Marathi, and Tamil. We employ rhythm formant (R-formants) analysis, a technique introduced by Gibbon that utilizes low-frequency spectral analysis of amplitude modulation and frequency modulation envelopes to characterize speech rhythm. Various measures are computed from the LF spectrum, including R-formants, discrete cosine transform-based measures, and spectral measures. Results show that threshold-based and spectral features outperform directly computed R-formants. Temporal pattern of rhythm derived from LF spectrograms provides better language-discriminating cues. Combining all derived features we achieve an accuracy of 69.21% and a weighted F1 score of 69.18% in classifying the five languages. This study demonstrates the potential of RFA in characterizing speech rhythm for Indian language classification.
【7】 Improving Data Augmentation-based Cross-Speaker Style Transfer for TTS with Singing Voice, Style Filtering, and F0 Matching
标题: 通过歌声、风格过滤和F0匹配改善基于数据增强的TTC交叉说话者风格转移
作者: Leonardo B. de M. M. Marques, Lucas H. Ueda, Mário U. Neto, Flávio O. Simões, Fernando Runstein, Bianca Dal Bó, Paula D. P. Costa
备注:Submitted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:文语转换系统中跨说话人风格转换的目标是将具有表达性数据的源说话人的说话风格转换到仅具有中性数据的目标说话人。在这种情况下,我们建议使用一个预先训练的歌声转换(SVC)模型,将表达数据转换成目标扬声器的声音。在转换过程中,我们应用基频(F0)匹配技术,以减轻扬声器之间的音色差异显着。提出了一种风格分类器过滤器来选择最具表现力的输出音频用于TTS训练。我们的方法可以与最先进的方法相媲美,只需要几分钟的目标说话人的中性数据,而其他方法需要几个小时。感知评估表明,SVC和风格过滤器带来的改善,自然和风格强度的风格,显示更多的声乐努力。此外,增加说话人相似度与建议的F0匹配算法。摘要:The goal of cross-speaker style transfer in TTS is to transfer a speech style from a source speaker with expressive data to a target speaker with only neutral data. In this context, we propose using a pre-trained singing voice conversion (SVC) model to convert the expressive data into the target speaker's voice. In the conversion process, we apply a fundamental frequency (F0) matching technique to mitigate tonal variances between speakers with significant timbral differences. A style classifier filter is proposed to select the most expressive output audios for the TTS training. Our approach is comparable to state-of-the-art with only a few minutes of neutral data from the target speaker, while other methods require hours. A perceptual assessment showed improvements brought by the SVC and the style filter in naturalness and style intensity for the styles that display more vocal effort. Also, increased speaker similarity is obtained with the proposed F0 matching algorithm.
【8】 The OCON model: an old but gold solution for distributable supervised classification
标题: OCON模型:可分布式监督分类的古老但黄金解决方案
作者: Stefano Giacomelli, Marco Giordano, Claudia Rinaldi
备注:Accepted at "2024 29th IEEE Symposium on Computers and Communications (ISCC): workshop on Next-Generation Multimedia Services at the Edge: Leveraging 5G and Beyond (NGMSE2024)". arXiv admin note: text overlap with arXiv:2410.04098
链接:点击下载PDF文件
摘要:本文介绍了一个结构化的应用程序的一类方法和一类一网络模型的监督分类任务,特别是在自动语音识别研究领域的元音音素分类的案例研究。通过使用明智的网格搜索方法进行的伪神经架构搜索和超参数调整实验,我们实现了与当今复杂架构相当的分类准确度(90.0 - 93.7%)。尽管它的简单性,我们的模型优先推广的语言环境和分布式的适用性,支持相关的统计和性能指标。实验代码在我们的GitHub上公开提供。摘要:This paper introduces to a structured application of the One-Class approach and the One-Class-One-Network model for supervised classification tasks, specifically addressing a vowel phonemes classification case study within the Automatic Speech Recognition research field. Through pseudo-Neural Architecture Search and Hyper-Parameters Tuning experiments conducted with an informed grid-search methodology, we achieve classification accuracy comparable to nowadays complex architectures (90.0 - 93.7%). Despite its simplicity, our model prioritizes generalization of language context and distributed applicability, supported by relevant statistical and performance metrics. The experiments code is openly available at our GitHub.
【9】 Episodic fine-tuning prototypical networks for optimization-based few-shot learning: Application to audio classification
标题: 基于优化的Few-Shot学习的情景式微调原型网络:应用于音频分类
作者: Xuanyu Zhuang (LTCI, IP Paris, S2A, IDS), Geoffroy Peeters (LTCI, IP Paris, S2A, IDS), Gaël Richard (S2A, IDS, LTCI, IP Paris)
备注:Accepted at MLSP 2024
Journal-ref:2024 IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2024), Sep 2024, London (UK), United Kingdom
链接:点击下载PDF文件
摘要:Prototypical Network(ProtoNet)以其卓越的性能和简单的实现方式成为Few-Shot Learning(FSL)场景中的热门选择。在此成功的基础上,我们首先提出了一种简单(但新颖)的方法来在C-way-K-shot测试集的测试集(标记)支持集上微调ProtoNet(不使用仅用于评估的查询集)。然后,我们提出了一个算法框架,结合ProtoNet与基于优化的FSL算法(MAML和Meta-Curvature)与这样的微调方法。由于基于优化的算法赋予目标学习模型快速适应少数样本的能力,我们利用ProtoNet作为目标模型,以提高其微调性能的帮助下,一个专门设计的情节微调策略。实验结果证实,我们提出的模型,MAML-Proto和MC-Proto,结合我们独特的微调方法,在ESC-50和Speech Commands v2数据集上的Few-Shot音频分类任务中,性能大大优于常规ProtoNet。我们注意到,虽然我们只将我们的模型应用于音频域,但它是一种通用的方法,可以很容易地扩展到其他领域。摘要:The Prototypical Network (ProtoNet) has emerged as a popular choice in Few-shot Learning (FSL) scenarios due to its remarkable performance and straightforward implementation. Building upon such success, we first propose a simple (yet novel) method to fine-tune a ProtoNet on the (labeled) support set of the test episode of a C-way-K-shot test episode (without using the query set which is only used for evaluation). We then propose an algorithmic framework that combines ProtoNet with optimization-based FSL algorithms (MAML and Meta-Curvature) to work with such a fine-tuning method. Since optimization-based algorithms endow the target learner model with the ability to fast adaption to only a few samples, we utilize ProtoNet as the target model to enhance its fine-tuning performance with the help of a specifically designed episodic fine-tuning strategy. The experimental results confirm that our proposed models, MAML-Proto and MC-Proto, combined with our unique fine-tuning method, outperform regular ProtoNet by a large margin in few-shot audio classification tasks on the ESC-50 and Speech Commands v2 datasets. We note that although we have only applied our model to the audio domain, it is a general method and can be easily extended to other domains.
【10】 Sylber: Syllabic Embedding Representation of Speech from Raw Audio
标题: Sylber:原始音频语音的音节嵌入表示
作者: Cheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal, Ethan Chen, Alan W Black, Gopala K. Anumanchipalli
链接:点击下载PDF文件
摘要:音节是口语的组成单位,在人类的言语感知和产生中起着至关重要的作用。然而,目前的神经语音表示缺乏结构,导致处理成本高的密集令牌序列。为了弥合这一差距,我们提出了一个新的模型,Sylber,产生干净和强大的音节结构的语音表示。具体来说,我们提出了一个自监督模型,该模型对从教师模型中提取的音节片段的特征进行回归,该模型是训练中模型的指数移动平均值。这导致了语音特征的高度结构化表示,提供了三个主要优点:1)快速的线性时间音节分割算法,2)平均每秒4.27个标记的有效音节标记化,以及3)更适合词汇和句法理解的音节单位。我们还训练令牌到语音生成模型与我们的音节单位,并表明,完全可理解的语音可以从这些令牌重建。最后,我们观察到分类感知,一种语音感知的语言现象,在我们的模型中自然出现,使得嵌入空间比以前的自监督学习方法更分类和稀疏。总之,我们提出了一种新的自我监督的方法来表示语音作为音节,具有显着的潜力,有效的语音标记和口语建模。摘要:Syllables are compositional units of spoken language that play a crucial role in human speech perception and production. However, current neural speech representations lack structure, resulting in dense token sequences that are costly to process. To bridge this gap, we propose a new model, Sylber, that produces speech representations with clean and robust syllabic structure. Specifically, we propose a self-supervised model that regresses features on syllabic segments distilled from a teacher model which is an exponential moving average of the model in training. This results in a highly structured representation of speech features, offering three key benefits: 1) a fast, linear-time syllable segmentation algorithm, 2) efficient syllabic tokenization with an average of 4.27 tokens per second, and 3) syllabic units better suited for lexical and syntactic understanding. We also train token-to-speech generative models with our syllabic units and show that fully intelligible speech can be reconstructed from these tokens. Lastly, we observe that categorical perception, a linguistic phenomenon of speech perception, emerges naturally in our model, making the embedding space more categorical and sparse than previous self-supervised learning approaches. Together, we present a novel self-supervised approach for representing speech as syllables, with significant potential for efficient speech tokenization and spoken language modeling.
【11】 Spectral and Rhythm Features for Audio Classification with Deep Convolutional Neural Networks
标题: 利用深度卷积神经网络进行音频分类的频谱和节奏特征
作者: Friedrich Wolf-Monheim
链接:点击下载PDF文件
摘要:卷积神经网络(CNN)广泛应用于计算机视觉。它们不仅可以用于传统的数字图像材料来识别模式,而且还可以用于从数字图像中提取特征,这些特征表示从时域数字音频信号中提取的频谱和节奏特征,用于声音的声学分类。不同的频谱和节奏特征表示,如梅尔缩放频谱图,梅尔频率倒谱系数(MFCC),循环tempogram,短时傅里叶变换(STFT)chromagram,恒定Q变换(CQT)chromagram和色度能量归一化统计(CENS)chromagram的音频分类性能使用深度卷积神经网络进行了研究。可以清楚地表明,对于使用深度CNN的音频分类任务,梅尔缩放频谱图和梅尔频率倒谱系数(MFCC)的表现明显优于本研究中研究的其他频谱和节奏特征。实验是在ESC-50数据集的帮助下进行的,该数据集包含2,000个标记的环境音频记录。摘要:Convolutional neural networks (CNNs) are widely used in computer vision. They can be used not only for conventional digital image material to recognize patterns, but also for feature extraction from digital imagery representing spectral and rhythm features extracted from time-domain digital audio signals for the acoustic classification of sounds. Different spectral and rhythm feature representations like mel-scaled spectrograms, mel-frequency cepstral coefficients (MFCCs), cyclic tempograms, short-time Fourier transform (STFT) chromagrams, constant-Q transform (CQT) chromagrams and chroma energy normalized statistics (CENS) chromagrams are investigated in terms of the audio classification performance using a deep convolutional neural network. It can be clearly shown that the mel-scaled spectrograms and the mel-frequency cepstral coefficients (MFCCs) perform significantly better then the other spectral and rhythm features investigated in this research for audio classification tasks using deep CNNs. The experiments were carried out with the aid of the ESC-50 dataset with 2,000 labeled environmental audio recordings.
【12】 Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity
标题: 预训练语音和语言模型的联合微调和转换为线性复杂性
作者: Mutian He, Philip N. Garner
备注:15 pages, 4 figures
链接:点击下载PDF文件
摘要:Linformer和Mamba等架构最近已经成为Transformers的有竞争力的线性时间替代品。然而,相应的大型预训练模型通常不可用,特别是在非文本领域。为了解决这个问题,我们提出了一个跨架构分层蒸馏(CALD)的方法,共同转换一个Transformer模型的线性时间替代和微调它的目标任务。我们还比较了几种方法来指导微调,以最佳地保留原始模型所需的推理能力。这些方法的不同之处在于它们对目标模型和参数轨迹的使用。在一系列关于语言处理、语言建模和语音处理的实证研究中,我们证明了CALD能够有效地恢复原始模型的结果,并且引导策略对结果有贡献。一些变化的原因提出了建议。摘要:Architectures such as Linformer and Mamba have recently emerged as competitive linear time replacements for transformers. However, corresponding large pretrained models are often unavailable, especially in non-text domains. To remedy this, we present a Cross-Architecture Layerwise Distillation (CALD) approach that jointly converts a transformer model to a linear time substitute and fine-tunes it to a target task. We also compare several means to guide the fine-tuning to optimally retain the desired inference capability from the original model. The methods differ in their use of the target model and the trajectory of the parameters. In a series of empirical studies on language processing, language modeling, and speech processing, we show that CALD can effectively recover the result of the original model, and that the guiding strategy contributes to the result. Some reasons for the variation are suggested.
【13】 SCOREQ: Speech Quality Assessment with Contrastive Regression
标题: Scrum:使用对比回归进行言语质量评估
作者: Alessandro Ragano, Jan Skoglund, Andrew Hines
备注:Accepted NeurIPS 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了一种新的语音质量预测的方法,SCOPIC。SCORK是用于对比回归的三元组损失函数,其解决了由现有技术的无参考语音质量度量所表现出的域泛化缺点。在本文中,我们:(i)说明第二语言损失训练未能捕捉平均意见评分(MOS)标签的连续性的问题;(ii)通过跨多个语音领域的基准评估来证明缺乏概括性;(iii)概述我们的方法并通过增量评估探索架构设计决策的影响;(iv)对照各种数据和领域的最新模型评估最终模型。结果表明,缺乏概括性观察到的最先进的语音质量指标的状态是由SCORIO解决。我们的结论是,使用三重损失函数的对比回归提高了语音质量预测模型的泛化,但也有潜在的效用,在广泛的应用,使用基于回归的预测模型。摘要:In this paper, we present SCOREQ, a novel approach for speech quality prediction. SCOREQ is a triplet loss function for contrastive regression that addresses the domain generalisation shortcoming exhibited by state of the art no-reference speech quality metrics. In the paper we: (i) illustrate the problem of L2 loss training failing at capturing the continuous nature of the mean opinion score (MOS) labels; (ii) demonstrate the lack of generalisation through a benchmarking evaluation across several speech domains; (iii) outline our approach and explore the impact of the architectural design decisions through incremental evaluation; (iv) evaluate the final model against state of the art models for a wide variety of data and domains. The results show that the lack of generalisation observed in state of the art speech quality metrics is addressed by SCOREQ. We conclude that using a triplet loss function for contrastive regression improves generalisation for speech quality prediction models but also has potential utility across a wide range of applications using regression-based predictive models.
【14】 Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS
标题: Bastards Harmony:用于Bastards文本到语音合成的综合数据集,采用EnGen-TTC的离散编解码器建模
作者: Onkar Kishor Susladkar, Vishesh Tripathi, Biddwan Ahmed
Journal-ref:EMNLP 2024
链接:点击下载PDF文件
摘要:本研究介绍了一个全面的Baidu文本到语音(TTS)数据集和一个新的TTS模型,EnGen TTS,旨在提高Baidu语言合成语音的质量和通用性。该数据集涵盖 textasciitilde55.0小时和52 K音频记录,集成了多种文本来源,确保语言丰富性。一个细致的录音设置捕捉到的细微差别巴托语音,采用专业设备,以确保高保真的音频样本。统计分析揭示了数据集的规模和多样性,为模型训练和评估奠定了基础。拟议的EnGen-TTS模型比既定的基线表现更好,实现了平均意见得分(MOS)为4.45 $ pm $0.13。此外,我们对实时因素和模型大小的调查突出了EnGen-TTS作为一个引人注目的选择,具有高效的性能。这项研究标志着Baidu TTS技术的重大进步,对不同的语言应用产生了影响。生成的样本链接: url{https: bahasa-harmony-comp.vercel.app }摘要:This research introduces a comprehensive Bahasa text-to-speech (TTS) dataset and a novel TTS model, EnGen-TTS, designed to enhance the quality and versatility of synthetic speech in the Bahasa language. The dataset, spanning textasciitilde55.0 hours and 52K audio recordings, integrates diverse textual sources, ensuring linguistic richness. A meticulous recording setup captures the nuances of Bahasa phonetics, employing professional equipment to ensure high-fidelity audio samples. Statistical analysis reveals the dataset's scale and diversity, laying the foundation for model training and evaluation. The proposed EnGen-TTS model performs better than established baselines, achieving a Mean Opinion Score (MOS) of 4.45 $ pm$ 0.13. Additionally, our investigation on real-time factor and model size highlights EnGen-TTS as a compelling choice, with efficient performance. This research marks a significant advancement in Bahasa TTS technology, with implications for diverse language applications. Link to Generated Samples: url{https: bahasa-harmony-comp.vercel.app }
【15】 SRC-gAudio: Sampling-Rate-Controlled Audio Generation
标题: SRC-gAudio:采样率控制的音频生成
作者: Chenxing Li, Manjie Xu, Dong Yu
备注:Accepted by APSIPA2024
链接:点击下载PDF文件
摘要:我们介绍SRC-gAudio,一种新颖的音频生成模型,旨在促进文本到音频生成在一个单一的模型架构内的广泛的采样率。SRC-gAudio将采样率作为生成条件的一部分,以指导基于扩散的音频生成过程。我们的模型使音频生成在多个采样率与一个单一的统一模型。此外,我们探讨了大规模,低采样率的数据在提高高采样率音频的生成质量的潜在好处。通过大量的实验,我们证明了SRC-gAudio有效地生成音频控制采样率下。此外,我们的研究结果表明,对低采样率数据进行预训练可以显著提高各种指标的音频质量。摘要:We introduce SRC-gAudio, a novel audio generation model designed to facilitate text-to-audio generation across a wide range of sampling rates within a single model architecture. SRC-gAudio incorporates the sampling rate as part of the generation condition to guide the diffusion-based audio generation process. Our model enables the generation of audio at multiple sampling rates with a single unified model. Furthermore, we explore the potential benefits of large-scale, low-sampling-rate data in enhancing the generation quality of high-sampling-rate audio. Through extensive experiments, we demonstrate that SRC-gAudio effectively generates audio under controlled sampling rates. Additionally, our results indicate that pre-training on low-sampling-rate data can lead to significant improvements in audio quality across various metrics.
【16】 Gumbel Rao Monte Carlo based Bi-Modal Neural Architecture Search for Audio-Visual Deepfake Detection
标题: Gumbel Rao Monte Carlo基于双模式神经架构搜索用于视听深度伪造检测
作者: Aravinda Reddy PN, Raghavendra Ramachandra, Krothapalli Sreenivasa Rao, Pabitra Mitra Vinod Rathod
链接:点击下载PDF文件
摘要:Deepfakes通过生成高度逼真的合成媒体对生物识别认证系统构成了严重威胁。现有的多模态deepfake检测器通常难以适应不同的数据,并依赖于简单的融合方法。为了解决这些挑战,我们提出了Gumbel-Rao Monte Carlo双峰神经架构搜索(GRMC-BMNAS),一种新的架构搜索框架,采用Gumbel-Rao Monte Carlo采样来优化多模态融合。它通过Rao-Blackwellization减少方差,稳定网络训练,改进了直通Gumbel Softmax(STGS)方法。使用两级搜索方法,该框架优化了网络结构,参数和性能。关键功能有效地识别骨干网络,而在细胞结构内,加权融合操作集成来自各种来源的信息。通过改变参数,如温度和蒙特卡罗样本的数量,产生一个架构,最大限度地提高分类性能和更好的泛化能力。在FakeAVCeleb和SWAN-DF数据集上的实验结果表明,用最小的模型参数实现了令人印象深刻的95.4%的AUC百分比。摘要:Deepfakes pose a critical threat to biometric authentication systems by generating highly realistic synthetic media. Existing multimodal deepfake detectors often struggle to adapt to diverse data and rely on simple fusion methods. To address these challenges, we propose Gumbel-Rao Monte Carlo Bi-modal Neural Architecture Search (GRMC-BMNAS), a novel architecture search framework that employs Gumbel-Rao Monte Carlo sampling to optimize multimodal fusion. It refines the Straight through Gumbel Softmax (STGS) method by reducing variance with Rao-Blackwellization, stabilizing network training. Using a two-level search approach, the framework optimizes the network architecture, parameters, and performance. Crucial features are efficiently identified from backbone networks, while within the cell structure, a weighted fusion operation integrates information from various sources. By varying parameters such as temperature and number of Monte carlo samples yields an architecture that maximizes classification performance and better generalisation capability. Experimental results on the FakeAVCeleb and SWAN-DF datasets demonstrate an impressive AUC percentage of 95.4 %, achieved with minimal model parameters.
【17】 Mamba-based Segmentation Model for Speaker Diarization
标题: 基于Mamba的说话人数字化分割模型
作者: Alexis Plaquet, Naohiro Tawara, Marc Delcroix, Shota Horiguchi, Atsushi Ando, Shoko Araki
备注:5 pages, 4 figures. Submitted to ICASSP 2025. Code at this https URL
链接:点击下载PDF文件
摘要:Mamba是一种新提出的架构,其行为类似于具有类似注意力能力的递归神经网络(RNN)。这些特性对于说话人日记化来说是有希望的,因为基于注意力的模型对长格式音频有不合适的记忆要求,并且传统的RNN能力太有限。在本文中,我们建议通过比较pyannote.audio管道的最先进的神经分割与我们提出的基于Mamba的变体来评估Mamba的日记化潜力。Mamba更强大的处理能力允许使用更长的局部窗口,这通过使说话人嵌入提取更可靠来显着提高日志质量。我们发现Mamba是传统RNN和经过测试的基于注意力的模型的更好的替代方案。我们提出的基于Mamba的系统在三个广泛使用的日志数据集上实现了最先进的性能。摘要:Mamba is a newly proposed architecture which behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization, as attention-based models have unsuitable memory requirements for long-form audio, and traditional RNN capabilities are too limited. In this paper, we propose to assess the potential of Mamba for diarization by comparing the state-of-the-art neural segmentation of the pyannote.audio pipeline with our proposed Mamba-based variant. Mamba's stronger processing capabilities allow usage of longer local windows, which significantly improve diarization quality by making the speaker embedding extraction more reliable. We find Mamba to be a superior alternative to both traditional RNN and the tested attention-based model. Our proposed Mamba-based system achieves state-of-the-art performance on three widely used diarization datasets.
【18】 POLIPHONE: A Dataset for Smartphone Model Identification from Audio Recordings
标题: POLIPHONE:从音频记录中识别智能手机型号的数据集
作者: Davide Salvi, Daniele Ugo Leonzio, Antonio Giganti, Claudio Eutizi, Sara Mandelli, Paolo Bestagini, Stefano Tubaro
备注:Submitted to IEEE Access
链接:点击下载PDF文件
摘要:在处理多媒体数据时,从取证的角度来看,源属性是一个关键的挑战。该任务旨在确定给定内容是如何被捕获的,为各种应用提供有价值的见解,包括法律诉讼和诚信调查。来源归属问题已经在不同的领域得到了解决,从识别用于捕获特定照片的相机模型到检测用于创建或记录给定音轨的合成语音发生器或麦克风模型。该领域的最新进展在很大程度上依赖于机器学习和数据驱动技术,这些技术通常优于传统的基于信号处理的方法。 然而,这些系统的缺点是它们需要大量的训练数据,这些数据必须反映最新的技术趋势,以产生准确可靠的预测。这是一个重大挑战,因为技术进步的快速步伐使得很难保持数据集与现实世界的条件保持一致。例如,在根据音频记录识别智能手机型号的任务中,可用的数据集通常已经过时或获取不一致,因此很难开发出在研究环境之外有效的解决方案。在本文中,我们介绍了POLIPHONE,这是一个用于从音频记录中识别智能手机型号的数据集。它包括在受控环境中记录的20个最新智能手机的数据,以确保未来研究的可重复性和可扩展性。所释放的轨道包含来自各个域的音频数据(即,语音、音乐、环境声音),使得语料库通用并且适用于广泛的用例。我们还提出了许多实验,使用最先进的分类器从音频记录中识别智能手机模型,对所提出的数据集进行基准测试。摘要:When dealing with multimedia data, source attribution is a key challenge from a forensic perspective. This task aims to determine how a given content was captured, providing valuable insights for various applications, including legal proceedings and integrity investigations. The source attribution problem has been addressed in different domains, from identifying the camera model used to capture specific photographs to detecting the synthetic speech generator or microphone model used to create or record given audio tracks. Recent advancements in this area rely heavily on machine learning and data-driven techniques, which often outperform traditional signal processing-based methods. However, a drawback of these systems is their need for large volumes of training data, which must reflect the latest technological trends to produce accurate and reliable predictions. This presents a significant challenge, as the rapid pace of technological progress makes it difficult to maintain datasets that are up-to-date with real-world conditions. For instance, in the task of smartphone model identification from audio recordings, the available datasets are often outdated or acquired inconsistently, making it difficult to develop solutions that are valid beyond a research environment. In this paper we present POLIPHONE, a dataset for smartphone model identification from audio recordings. It includes data from 20 recent smartphones recorded in a controlled environment to ensure reproducibility and scalability for future research. The released tracks contain audio data from various domains (i.e., speech, music, environmental sounds), making the corpus versatile and applicable to a wide range of use cases. We also present numerous experiments to benchmark the proposed dataset using a state-of-the-art classifier for smartphone model identification from audio recordings.
【19】 Variable Bitrate Residual Vector Quantization for Audio Coding
标题: 音频编码中的可变比特率残留量量化
作者: Yunkee Chae, Woosung Choi, Yuhta Takida, Junghyun Koo, Yukara Ikemiya, Zhi Zhong, Kin Wai Cheuk, Marco A. Martínez-Ramírez, Kyogu Lee, Wei-Hsiang Liao, Yuki Mitsufuji
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:最近的最先进的神经音频压缩模型已逐步采用残差矢量量化(RVQ)。尽管取得了这一成功,但这些模型每帧采用固定数量的码本,这在速率失真权衡方面可能是次优的,特别是在具有简单输入音频的场景中,例如静音。为了解决这一限制,我们提出了可变比特率RVQ(VRVQ)的音频编解码器,它允许更有效的编码,通过调整每帧使用的码本的数量。此外,我们提出了一种用于不可微掩蔽操作的梯度估计方法,该方法从重要性图转换为二进制重要性掩模,通过直通估计器改进模型训练。我们证明了所提出的训练框架相比基线方法取得了更好的效果,并显示出进一步的改进时,应用到当前最先进的编解码器。摘要:Recent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be suboptimal in terms of rate-distortion tradeoff, particularly in scenarios with simple input audio, such as silence. To address this limitation, we propose variable bitrate RVQ (VRVQ) for audio codecs, which allows for more efficient coding by adapting the number of codebooks used per frame. Furthermore, we propose a gradient estimation method for the non-differentiable masking operation that transforms from the importance map to the binary importance mask, improving model training via a straight-through estimator. We demonstrate that the proposed training framework achieves superior results compared to the baseline method and shows further improvement when applied to the current state-of-the-art codec.
【20】 FINALLY: fast and universal speech enhancement with studio-like quality
标题: 最终:快速、通用的语音增强,具有类似录音室的质量
作者: Nicholas Babaev, Kirill Tamogashev, Azat Saginbaev, Ivan Shchekotov, Hanbin Bae, Hosang Sung, WonJun Lee, Hoon-Young Cho, Pavel Andreev
备注:Accepted to NeurIPS 2024
链接:点击下载PDF文件
摘要:在本文中,我们解决了现实世界的录音,其中往往包含各种形式的失真,如背景噪声,混响和麦克风文物语音增强的挑战。我们重新审视了生成对抗网络(GANs)用于语音增强的使用,并从理论上表明,GANs自然倾向于在有条件的干净语音分布中寻找最大密度点,正如我们所认为的,这对于语音增强任务至关重要。我们研究了感知损失的各种特征提取器,以促进对抗训练的稳定性,开发了一种探测特征空间结构的方法。这促使我们将基于WavLM的感知损失集成到MS-STFT对抗训练管道中,为语音增强模型创建有效且稳定的训练过程。由此产生的语音增强模型,我们称之为FINALLY,建立在HiFi++架构之上,使用WavLM编码器和一个新的训练管道进行增强。各种数据集上的实证结果证实了我们的模型能够在48 kHz下产生清晰,高质量的语音,在语音增强领域实现最先进的性能。摘要:In this paper, we address the challenge of speech enhancement in real-world recordings, which often contain various forms of distortion, such as background noise, reverberation, and microphone artifacts. We revisit the use of Generative Adversarial Networks (GANs) for speech enhancement and theoretically show that GANs are naturally inclined to seek the point of maximum density within the conditional clean speech distribution, which, as we argue, is essential for the speech enhancement task. We study various feature extractors for perceptual loss to facilitate the stability of adversarial training, developing a methodology for probing the structure of the feature space. This leads us to integrate WavLM-based perceptual loss into MS-STFT adversarial training pipeline, creating an effective and stable training procedure for the speech enhancement model. The resulting speech enhancement model, which we refer to as FINALLY, builds upon the HiFi++ architecture, augmented with a WavLM encoder and a novel training pipeline. Empirical results on various datasets confirm our model's ability to produce clear, high-quality speech at 48 kHz, achieving state-of-the-art performance in the field of speech enhancement.
【21】 FürElise: Capturing and Physically Synthesizing Hand Motions of Piano Performance
标题: FürElise:捕捉和物理合成钢琴演奏的手部动作
作者: Ruocheng Wang, Pei Xu, Haochen Shi, Elizabeth Schumann, C. Karen Liu
备注:SIGGRAPH Asia 2024. Project page: this https URL
链接:点击下载PDF文件
摘要:钢琴演奏需要敏捷,精确和协调的手控制,延伸灵活性的极限。手部运动模型具有精确再现钢琴演奏的复杂性,在角色动画,体现AI,生物力学和VR AR中有广泛的应用。在本文中,我们构建了一个首个大型数据集,其中包含大约10个小时的3D手部运动和来自15位精英级钢琴家演奏153首古典音乐的音频。为了捕捉自然的表演,我们设计了一个无标记的设置,其中使用最先进的姿态估计模型从多视图视频中重建运动。运动数据通过使用从专门的雅马哈电钢琴中的传感器获得的高分辨率按键数据的逆运动学进一步细化。利用收集到的数据集,我们开发了一个流水线,可以为数据集之外的乐谱合成物理上合理的手部动作。我们的方法采用模仿学习和强化学习的组合,以获得基于物理的双手控制策略,涉及手和钢琴键之间的相互作用。为了解决大运动数据集的采样效率问题,我们使用扩散模型来生成自然的参考运动,提供高层次的轨迹和指法(手指顺序和放置)信息。然而,单独生成的参考运动不能为钢琴演奏建模提供足够的精度。然后,我们通过使用音乐相似性来从捕获的数据集中检索相似的运动来进一步增强数据,以提高RL策略的精度。使用所提出的方法,我们的模型生成自然,灵巧的运动,从训练数据集之外推广到音乐。摘要:Piano playing requires agile, precise, and coordinated hand control that stretches the limits of dexterity. Hand motion models with the sophistication to accurately recreate piano playing have a wide range of applications in character animation, embodied AI, biomechanics, and VR AR. In this paper, we construct a first-of-its-kind large-scale dataset that contains approximately 10 hours of 3D hand motion and audio from 15 elite-level pianists playing 153 pieces of classical music. To capture natural performances, we designed a markerless setup in which motions are reconstructed from multi-view videos using state-of-the-art pose estimation models. The motion data is further refined via inverse kinematics using the high-resolution MIDI key-pressing data obtained from sensors in a specialized Yamaha Disklavier piano. Leveraging the collected dataset, we developed a pipeline that can synthesize physically-plausible hand motions for musical scores outside of the dataset. Our approach employs a combination of imitation learning and reinforcement learning to obtain policies for physics-based bimanual control involving the interaction between hands and piano keys. To solve the sampling efficiency problem with the large motion dataset, we use a diffusion model to generate natural reference motions, which provide high-level trajectory and fingering (finger order and placement) information. However, the generated reference motion alone does not provide sufficient accuracy for piano performance modeling. We then further augmented the data by using musical similarity to retrieve similar motions from the captured dataset to boost the precision of the RL policy. With the proposed method, our model generates natural, dexterous motions that generalize to music from outside the training dataset.
【22】 Array2BR: An End-to-End Noise-immune Binaural Audio Synthesis from Microphone-array Signals
标题: Array 2BR:来自麦克风阵列信号的端到端抗噪双耳音频合成
作者: Cheng Chi, Xiaoyu Li, Andong Li, Yuxuan Ke, Xiaodong Li, Chengshi Zheng
链接:点击下载PDF文件
摘要:临场感技术旨在为远程会议应用提供沉浸式的虚拟临场感,为此,合成高质量的双耳音频信号至关重要。由于在实际应用场景中,环境噪声往往是不可避免的,因此迫切希望能够直接从麦克风阵列信号中获得没有噪声的双耳音频信号。为此,本文提出了一种新的基于麦克风阵列信号的端到端抗噪声双耳音频合成框架Array2BR,实验结果表明,该框架能够在正确映射双耳线索的同时很好地抑制噪声。与现有的方法相比,该方法取得了更好的性能,在客观和主观的度量分数。摘要:Telepresence technology aims to provide an immersive virtual presence for remote conference applications, and it is extremely important to synthesize high-quality binaural audio signals for this aim. Because the ambient noise is often inevitable in practical application scenarios, it is highly desired that binaural audio signals without noise can be obtained from microphone-array signals directly. For this purpose, this paper proposes a new end-to-end noise-immune binaural audio synthesis framework from microphone-array signals, abbreviated as Array2BR, and experimental results show that binaural cues can be correctly mapped and noise can be well suppressed simultaneously using the proposed framework. Compared with existing methods, the proposed method achieved better performance in terms of both objective and subjective metric scores.
【23】 FGCL: Fine-grained Contrastive Learning For Mandarin Stuttering Event Detection
标题: FGCL:用于普通话口吃事件检测的细粒度对比学习
作者: Han Jiang, Wenyu Wang, Yiquan Zhou, Hongwu Ding, Jiacheng Xu, Jihua Zhu
备注:Accepted to SLT 2024
链接:点击下载PDF文件
摘要:本文介绍了T031团队在SLT2024中应对口吃语音挑战的方法。汉语口吃事件检测(MSED)的目的是检测汉语语音中的口吃事件。我们提出了一个详细的声学分析方法,以提高口吃检测的准确性,通过捕捉微妙的细微差别,以前口吃事件检测(SED)技术忽略了。为此,我们介绍了MSED的细粒度对比学习(FGCL)框架。具体来说,我们模型的口吃事件的帧级概率,并引入一个挖掘算法来识别容易和混乱的帧。然后,我们提出了一个口吃对比度损失,以提高口吃和流畅的语音帧之间的区别,从而提高口吃特征嵌入的判别能力。在英语和普通话数据集上的广泛评估证明了FGCL的有效性,在普通话数据上实现了超过5.0%的F1分数的显着提高。摘要:This paper presents the T031 team's approach to the StutteringSpeech Challenge in SLT2024. Mandarin Stuttering Event Detection (MSED) aims to detect instances of stuttering events in Mandarin speech. We propose a detailed acoustic analysis method to improve the accuracy of stutter detection by capturing subtle nuances that previous Stuttering Event Detection (SED) techniques have overlooked. To this end, we introduce the Fine-Grained Contrastive Learning (FGCL) framework for MSED. Specifically, we model the frame-level probabilities of stuttering events and introduce a mining algorithm to identify both easy and confusing frames. Then, we propose a stutter contrast loss to enhance the distinction between stuttered and fluent speech frames, thereby improving the discriminative capability of stuttered feature embeddings. Extensive evaluations on English and Mandarin datasets demonstrate the effectiveness of FGCL, achieving a significant increase of over 5.0% in F1 score on Mandarin data.
【24】 Dynamic HumTrans: Humming Transcription Using CNNs and Dynamic Programming
标题: 动态HumTrans:使用CNN和动态编程的哼唱转录
作者: Shubham Gupta, Isaac Neri Gomez-Sarmiento, Faez Amjed Mezdari, Mirco Ravanelli, Cem Subakan
链接:点击下载PDF文件
摘要:我们提出了一种新的哼唱转录方法,结合了基于CNN的架构与基于动态编程的后处理算法,利用最近推出的HumTrans数据集。我们识别并解决了数据集提供的偏移和起始地面实况的固有问题,提供了改进这些注释的方法,从而产生了具有精确注释的数据集,这将有助于未来的研究。此外,我们将我们的方法的转录准确性与其他几种方法进行了比较,证明了最先进的(SOTA)结果。我们所有的代码和校正数据集都可以在https: github.com shubham-gupta-30 humming_transcription上找到摘要:We propose a novel approach for humming transcription that combines a CNN-based architecture with a dynamic programming-based post-processing algorithm, utilizing the recently introduced HumTrans dataset. We identify and address inherent problems with the offset and onset ground truth provided by the dataset, offering heuristics to improve these annotations, resulting in a dataset with precise annotations that will aid future research. Additionally, we compare the transcription accuracy of our method against several others, demonstrating state-of-the-art (SOTA) results. All our code and corrected dataset is available at https: github.com shubham-gupta-30 humming_transcription
【25】 Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
标题: 通过改进对抗环境中的语音识别来增强说话者身份援助
作者: Sagarika Alavilli, Annesya Banerjee, Gasser Elbanna, Annika Magaro
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:当前最先进的语音识别模型被训练成将声学信号映射到子词汇单元。虽然这些模型表现出优越的性能,但它们仍然容易受到背景噪声和语音增强等分布条件的影响。在这项工作中,我们假设在语音识别过程中将说话人表示可以增强模型对噪声的鲁棒性。我们开发了一个基于transformer的模型,共同执行语音识别和说话人识别。我们的模型利用来自Whisper的语音嵌入和来自ECAPA-TDNN的扬声器嵌入,它们被联合处理以执行这两项任务。我们表明,在清洁条件下,联合模型执行的耳语。值得注意的是,联合模型在高噪声环境中优于Whisper,例如具有8扬声器的串音背景噪声。此外,我们的联合模型擅长处理高度增强的语音,包括正弦波和噪声声码语音。总体而言,这些结果表明,将语音表示与语音识别相结合可以在对抗条件下产生更强大的模型。摘要:Current state-of-the-art speech recognition models are trained to map acoustic signals into sub-lexical units. While these models demonstrate superior performance, they remain vulnerable to out-of-distribution conditions such as background noise and speech augmentations. In this work, we hypothesize that incorporating speaker representations during speech recognition can enhance model robustness to noise. We developed a transformer-based model that jointly performs speech recognition and speaker identification. Our model utilizes speech embeddings from Whisper and speaker embeddings from ECAPA-TDNN, which are processed jointly to perform both tasks. We show that the joint model performs comparably to Whisper under clean conditions. Notably, the joint model outperforms Whisper in high-noise environments, such as with 8-speaker babble background noise. Furthermore, our joint model excels in handling highly augmented speech, including sine-wave and noise-vocoded speech. Overall, these results suggest that integrating voice representations with speech recognition can lead to more robust models under adversarial conditions.
【26】 RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction
标题: RespLLM:通过多模式LLM统一音频和文本,用于广义呼吸健康预测
作者: Yuwei Zhang, Tong Xia, Aaqib Saeed, Cecilia Mascolo
链接:点击下载PDF文件
摘要:与呼吸道疾病相关的高发病率和死亡率强调了早期筛查的重要性。机器学习模型可以自动化临床咨询和听诊,为这一领域提供重要支持。然而,所涉及的数据,包括人口统计学、病史、症状和呼吸音,是异质和复杂的。现有的方法是不够的,缺乏普遍性,因为它们通常依赖于有限的训练数据,基本的融合技术和特定于任务的模型。在本文中,我们提出了RespLLM,一种新的多模态大语言模型(LLM)框架,它将文本和音频表示统一起来,用于呼吸健康预测。RespLLM利用预训练LLM的广泛先验知识,并通过跨模态注意力实现有效的音频-文本融合。采用指令调优来集成来自多个源的不同数据,确保模型的通用性和通用性。在五个真实世界数据集上的实验表明,RespLLM在训练任务上的平均性能优于领先基线4.6%,在看不见的数据集上的平均性能优于领先基线7.9%,并有助于对新任务进行zero-shot预测。我们的工作为多模态模型奠定了基础,这些模型可以感知、倾听和理解异构数据,为可扩展的呼吸健康诊断铺平了道路。摘要:The high incidence and mortality rates associated with respiratory diseases underscores the importance of early screening. Machine learning models can automate clinical consultations and auscultation, offering vital support in this area. However, the data involved, spanning demographics, medical history, symptoms, and respiratory audio, are heterogeneous and complex. Existing approaches are insufficient and lack generalizability, as they typically rely on limited training data, basic fusion techniques, and task-specific models. In this paper, we propose RespLLM, a novel multimodal large language model (LLM) framework that unifies text and audio representations for respiratory health prediction. RespLLM leverages the extensive prior knowledge of pretrained LLMs and enables effective audio-text fusion through cross-modal attentions. Instruction tuning is employed to integrate diverse data from multiple sources, ensuring generalizability and versatility of the model. Experiments on five real-world datasets demonstrate that RespLLM outperforms leading baselines by an average of 4.6% on trained tasks, 7.9% on unseen datasets, and facilitates zero-shot predictions for new tasks. Our work lays the foundation for multimodal models that can perceive, listen to, and understand heterogeneous data, paving the way for scalable respiratory health diagnosis.
【27】 Diffusion-based Unsupervised Audio-visual Speech Enhancement
标题: 基于扩散的无监督视听语音增强
作者: Jean-Eudes Ayilo (MULTISPEECH), Mostafa Sadeghi (MULTISPEECH), Romain Serizel (MULTISPEECH), Xavier Alameda-Pineda (ROBOTLEARN)
链接:点击下载PDF文件
摘要:本文提出了一种新的无监督视听语音增强(AVSE)方法,结合了基于扩散的视听语音生成模型与非负矩阵分解(NMF)噪声模型。首先,在以相应视频数据为条件的干净语音上预训练扩散模型以模拟语音生成分布。然后,将该预训练模型与基于NMF的噪声模型配对,以迭代地估计干净的语音。具体地,基于扩散的后验采样方法在逆扩散过程中实现,其中在每次迭代之后,获得语音估计并用于更新噪声参数。实验结果证实,所提出的AVSE方法不仅优于其音频只对应,但也比最近的监督生成AVSE方法更好地推广。此外,与以前的基于扩散的方法相比,新的推理算法在推理速度和性能之间提供了更好的平衡。摘要:This paper proposes a new unsupervised audiovisual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. First, the diffusion model is pre-trained on clean speech conditioned on corresponding video data to simulate the speech generative distribution. This pre-trained model is then paired with the NMF-based noise model to iteratively estimate clean speech. Specifically, a diffusion-based posterior sampling approach is implemented within the reverse diffusion process, where after each iteration, a speech estimate is obtained and used to update the noise parameters. Experimental results confirm that the proposed AVSE approach not only outperforms its audio-only counterpart but also generalizes better than a recent supervisedgenerative AVSE method. Additionally, the new inference algorithm offers a better balance between inference speed and performance compared to the previous diffusion-based method.
机器翻译,仅供参考
![]()
