本文经arXiv每日学术速递授权转载
【1】 Towards Leveraging Contrastively Pretrained Neural Audio Embeddings for Recommender Tasks
标题: 利用对比预训练的神经音频嵌入进行推荐任务
作者:Florian Grötschla,Luca Strässle,Luca A. Lanzendörfer,Roger Wattenhofer
备注:Accepted at the 2nd Music Recommender Workshop (@RecSys)
链接:点击下载PDF文件
【2】 Biomimetic Frontend for Differentiable Audio Processing
标题: 可区分音频处理的仿生前沿
作者:Ruolan Leslie Famularo,Dmitry N. Zotkin,Shihab A. Shamma,Ramani Duraiswami
链接:点击下载PDF文件
【3】 Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages
标题: 探索极低资源语言中数据量对ASB的影响
作者:Yao-Fei Cheng,Li-Wei Chen,Hung-Shin Lee,Hsin-Min Wang
链接:点击下载PDF文件
【4】 Exploring SSL Discrete Tokens for Multilingual ASR
标题: 探索多语言ASB的SSL离散令牌
作者:Mingyu Cui,Daxin Tan,Yifan Yang,Dingdong Wang,Huimeng Wang,Xiao Chen,Xie Chen,Xunying Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【5】 Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
标题: 探索基于Zipformer的上下文ASB的SSL离散语音特征
作者:Mingyu Cui,Yifan Yang,Jiajun Deng,Jiawen Kang,Shujie Hu,Tianzi Wang,Zhaoqing Li,Shiliang Zhang,Xie Chen,Xunying Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【6】 Energy Consumption Trends in Sound Event Detection Systems
标题: 声音事件检测系统的能源消耗趋势
作者:Constance Douwes,Romain Serizel
链接:点击下载PDF文件
【7】 DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
标题: DFADD:基于扩散和流匹配的音频Deepfake数据集
作者:Jiawei Du,I-Ming Lin,I-Hsiang Chiu,Xuanjun Chen,Haibin Wu,Wenze Ren,Yu Tsao,Hung-yi Lee,Jyh-Shing Roger Jang
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
【8】 Acoustic identification of individual animals with hierarchical contrastive learning
标题: 利用分层对比学习进行个体动物的声学识别
作者:Ines Nolasco,Ilyass Moummad,Dan Stowell,Emmanouil Benetos
备注:Under review; Submitted to ICASSP 2025
链接:点击下载PDF文件
【9】 Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
标题: 研究音素级语音编解码器中的解纠缠以进行韵律建模
作者:Sotirios Karapiperis,Nikolaos Ellinas,Alexandra Vioni,Junkwang Oh,Gunu Jho,Inchul Hwang,Spyros Raptis
链接:点击下载PDF文件
【10】 LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
标题: LMAC-TD:为音频分类器生成时间域简化
作者:Eleonora Mancini,Francesco Paissan,Mirco Ravanelli,Cem Subakan
备注:The first two authors contributed equally to this research. Author order is alphabetical
链接:点击下载PDF文件
【11】 Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
标题: 节奏Foley:视频到音频合成中无缝视听对齐的框架
作者:Zhiqi Huang,Dan Luo,Jun Wang,Huan Liao,Zhiheng Li,Zhiyong Wu
链接:点击下载PDF文件
【12】 TapToTab : Video-Based Guitar Tabs Generation using AI and Audio Analysis
标题: TapToTab:使用人工智能和音频分析的基于视频的吉他标签生成
作者:Ali Ghaleb,Eslam ElSadawy,Ihab Essam,Mohamed Abdelhakim,Seif-Eldin Zaki,Natalie Fahim,Razan Bayoumi,Hanan Hindy
链接:点击下载PDF文件
【13】 STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
标题: STA-V2 A:具有语义和时间对齐的视频到音频生成
作者:Yong Ren,Chenxing Li,Manjie Xu,Wei Liang,Yu Gu,Rilin Chen,Dong Yu
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
【14】 LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
标题: LA-RAG:通过检索增强生成增强基于LLM的ASB准确性
作者:Shaojun Li,Hengchao Shang,Daimeng Wei,Jiaxin Guo,Zongyao Li,Xianghui He,Min Zhang,Hao Yang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【15】 Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
标题: 大型语言模型可以在多说话者场景中通过多功能指令转录语音
作者:Lingwei Meng,Shujie Hu,Jiawen Kang,Zhaoqing Li,Yuejiao Wang,Wenxuan Wu,Xixin Wu,Xunying Liu,Helen Meng
链接:点击下载PDF文件
【16】 Domain-Invariant Representation Learning of Bird Sounds
标题: 鸟声的域不变表示学习
作者:Ilyass Moummad,Romain Serizel,Emmanouil Benetos,Nicolas Farrugia
链接:点击下载PDF文件
【17】 LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
标题: LHQ-SRC:轻量级、高质量的歌唱声音转换建模
作者:Yubo Huang,Xin Lai,Muyang Ye,Anran Zhu,Zixi Wang,Jingzehua Xu,Shuai Zhang,Zhiyuan Zhou,Weijie Niu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【18】 Apollo: Band-sequence Modeling for High-Quality Audio Restoration
标题: Apollo:用于高质量音频恢复的带序列建模
作者:Kai Li,Yi Luo
备注:Demo Page: this https URL
链接:点击下载PDF文件
【19】 Confidence Calibration for Audio Captioning Models
标题: 音频字幕模型的置信度校准
作者:Rehana Mahfuz,Yinyi Guo,Erik Visser
链接:点击下载PDF文件
【20】 Why some audio signal short-time Fourier transform coefficients have nonuniform phase distributions
标题: 为什么某些音频信号短期傅里叶变换系数具有不均匀的相分布
作者:Stephen D. Voran
Journal-ref:Proceedings of the 2024 IEEE International Conference on Multimedia and Expo, Niagara Falls, Ontario, July 15-19, 2024
链接:点击下载PDF文件
【21】 Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
标题: 使用耳脑电解码多说话者环境中的听觉注意力
作者:Haolin Zhu,Yujie Yan,Xiran Xu,Zhongshu Ge,Pei Tian,Xihong Wu,Jing Chen
链接:点击下载PDF文件
【22】 DualSep: A Light-weight dual-encoder convolutional recurrent network for real-time in-car speech separation
标题: DualSep:用于实时车内语音分离的轻量级双编码器卷积循环网络
作者:Ziqian Wang,Jiayao Sun,Zihan Zhang,Xingchen Li,Jie Liu,Lei Xie
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
【23】 Effective Integration of KAN for Keyword Spotting
标题: 有效集成KAN以识别关键词
作者:Anfeng Xu,Biqiao Zhang,Shuyu Kong,Yiteng Huang,Zhaojun Yang,Sangeeta Srivastava,Ming Sun
备注:Under review
链接:点击下载PDF文件
【24】 Frequency Tracking Features for Data-Efficient Deep Siren Identification
标题: 频率跟踪功能,实现数据高效的深度警报识别
作者:Stefano Damiano,Thomas Dietzen,Toon van Waterschoot
备注:Accepted paper: Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2024)
链接:点击下载PDF文件
【25】 Unified Audio Event Detection
标题: 统一音频事件检测
作者:Yidi Jiang,Ruijie Tao,Wen Huang,Qian Chen,Wen Wang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【26】 SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
标题: SoloAudio:使用面向数字的音频扩散Transformer的目标声音提取
作者:Helin Wang,Jiarui Hai,Yen-Ju Lu,Karan Thakkar,Mounya Elhilali,Najim Dehak
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【27】 OpenACE: An Open Benchmark for Evaluating Audio Coding Performance
标题: OpenACE:评估音频编码性能的开放基准
作者:Jozef Coldenhoff,Niclas Granqvist,Milos Cernak
链接:点击下载PDF文件
【28】 Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
标题: 量化和减少跨语言言语反欺骗中的语言不匹配效应
作者:Tianchi Liu,Ivan Kukanov,Zihan Pan,Qiongqiong Wang,Hardik B. Sailor,Kong Aik Lee
备注:Accepted to the IEEE Spoken Language Technology Workshop (SLT) 2024. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
【29】 Detection of Electric Motor Damage Through Analysis of Sound Signals Using Bayesian Neural Networks
标题: 利用Bayesian神经网络分析声音信号检测电动机损坏
作者:Waldemar Bauer,Marta Zagorowska,Jerzy Baranowski
备注:Accepted to IECON 2024
链接:点击下载PDF文件
标题: 为什么某些音频信号短期傅里叶变换系数具有不均匀的相分布
作者:Stephen D. Voran
Journal-ref:Proceedings of the 2024 IEEE International Conference on Multimedia and Expo, Niagara Falls, Ontario, July 15-19, 2024
链接:点击下载PDF文件
【2】 HLTCOE JHU Submission to the Voice Privacy Challenge 2024
标题: HLTCOE JHU提交2024年语音隐私挑战赛
作者:Henry Li Xinyuan,Zexin Cai,Ashi Garg,Kevin Duh,Leibny Paola García-Perera,Sanjeev Khudanpur,Nicholas Andrews,Matthew Wiesner
备注:Submission to the Voice Privacy Challenge 2024. Accepted and presented at
链接:点击下载PDF文件
【3】 Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
标题: 具有模拟对话的数据高效儿童-成人说话者拨号
作者:Anfeng Xu,Tiantian Feng,Helen Tager-Flusberg,Catherine Lord,Shrikanth Narayanan
备注:Under review
链接:点击下载PDF文件
【4】 LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment
标题: LLaQo:在表达性音乐表现评估中建立基于查询的教练
作者:Huan Zhang,Vincent Cheung,Hayato Nishioka,Simon Dixon,Shinichi Furuya
链接:点击下载PDF文件
【5】 FLAMO: An Open-Source Library for Frequency-Domain Differentiable Audio Processing
标题: FLAMO:一个用于频域可区分音频处理的开源库
作者:Gloria Dal Santo,Gian Marco De Bortoli,Karolina Prawda,Sebastian J. Schlecht,Vesa Välimäki
链接:点击下载PDF文件
【6】 Text-To-Speech Synthesis In The Wild
标题: 野外的文本到语音合成
作者:Jee-weon Jung,Wangyou Zhang,Soumi Maiti,Yihan Wu,Xin Wang,Ji-Hoon Kim,Yuta Matsunaga,Seyun Um,Jinchuan Tian,Hye-jin Shim,Nicholas Evans,Joon Son Chung,Shinnosuke Takamichi,Shinji Watanabe
备注:5 pages, submitted to ICASSP 2025 as a conference paper
链接:点击下载PDF文件
【7】 Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
标题: 使用耳脑电解码多说话者环境中的听觉注意力
作者:Haolin Zhu,Yujie Yan,Xiran Xu,Zhongshu Ge,Pei Tian,Xihong Wu,Jing Chen
链接:点击下载PDF文件
【8】 DM: Dual-path Magnitude Network for General Speech Restoration
标题: DM:用于通用语音恢复的双路径幅度网络
作者:Da-Hee Yang,Dail Kim,Joon-Hyuk Chang,Jeonghwan Choi,Han-gil Moon
链接:点击下载PDF文件
【9】 NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training
标题: NEST-PQ:语音自我监督预训练的下一个令牌预测
作者:Minglun Han,Ye Bai,Chen Shen,Youjia Huang,Mingkun Huang,Zehua Lin,Linhao Dong,Lu Lu,Yuxuan Wang
备注:5 pages, 2 figures, Work in progress
链接:点击下载PDF文件
【10】 DualSep: A Light-weight dual-encoder convolutional recurrent network for real-time in-car speech separation
标题: DualSep:用于实时车内语音分离的轻量级双编码器卷积循环网络
作者:Ziqian Wang,Jiayao Sun,Zihan Zhang,Xingchen Li,Jie Liu,Lei Xie
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
【11】 Effective Integration of KAN for Keyword Spotting
标题: 有效集成KAN以识别关键词
作者:Anfeng Xu,Biqiao Zhang,Shuyu Kong,Yiteng Huang,Zhaojun Yang,Sangeeta Srivastava,Ming Sun
备注:Under review
链接:点击下载PDF文件
【12】 Frequency Tracking Features for Data-Efficient Deep Siren Identification
标题: 频率跟踪功能,实现数据高效的深度警报识别
作者:Stefano Damiano,Thomas Dietzen,Toon van Waterschoot
备注:Accepted paper: Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2024)
链接:点击下载PDF文件
【13】 Unified Audio Event Detection
标题: 统一音频事件检测
作者:Yidi Jiang,Ruijie Tao,Wen Huang,Qian Chen,Wen Wang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【14】 SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
标题: SoloAudio:使用面向数字的音频扩散Transformer的目标声音提取
作者:Helin Wang,Jiarui Hai,Yen-Ju Lu,Karan Thakkar,Mounya Elhilali,Najim Dehak
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【15】 OpenACE: An Open Benchmark for Evaluating Audio Coding Performance
标题: OpenACE:评估音频编码性能的开放基准
作者:Jozef Coldenhoff,Niclas Granqvist,Milos Cernak
链接:点击下载PDF文件
【16】 Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
标题: 量化和减少跨语言言语反欺骗中的语言不匹配效应
作者:Tianchi Liu,Ivan Kukanov,Zihan Pan,Qiongqiong Wang,Hardik B. Sailor,Kong Aik Lee
备注:Accepted to the IEEE Spoken Language Technology Workshop (SLT) 2024. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
【17】 Detection of Electric Motor Damage Through Analysis of Sound Signals Using Bayesian Neural Networks
标题: 利用Bayesian神经网络分析声音信号检测电动机损坏
作者:Waldemar Bauer,Marta Zagorowska,Jerzy Baranowski
备注:Accepted to IECON 2024
链接:点击下载PDF文件
【18】 Towards Leveraging Contrastively Pretrained Neural Audio Embeddings for Recommender Tasks
标题: 利用对比预训练的神经音频嵌入进行推荐任务
作者:Florian Grötschla,Luca Strässle,Luca A. Lanzendörfer,Roger Wattenhofer
备注:Accepted at the 2nd Music Recommender Workshop (@RecSys)
链接:点击下载PDF文件
【19】 Biomimetic Frontend for Differentiable Audio Processing
标题: 可区分音频处理的仿生前沿
作者:Ruolan Leslie Famularo,Dmitry N. Zotkin,Shihab A. Shamma,Ramani Duraiswami
链接:点击下载PDF文件
【20】 Clean Label Attacks against SLU Systems
标题: 针对SL U系统的干净标签攻击
作者:Henry Li Xinyuan,Sonal Joshi,Thomas Thebaud,Jesus Villalba,Najim Dehak,Sanjeev Khudanpur
备注:Accepted at IEEE SLT 2024
链接:点击下载PDF文件
【21】 Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages
标题: 探索极低资源语言中数据量对ASB的影响
作者:Yao-Fei Cheng,Li-Wei Chen,Hung-Shin Lee,Hsin-Min Wang
链接:点击下载PDF文件
【22】 Exploring SSL Discrete Tokens for Multilingual ASR
标题: 探索多语言ASB的SSL离散令牌
作者:Mingyu Cui,Daxin Tan,Yifan Yang,Dingdong Wang,Huimeng Wang,Xiao Chen,Xie Chen,Xunying Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【23】 Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
标题: 探索基于Zipformer的上下文ASB的SSL离散语音特征
作者:Mingyu Cui,Yifan Yang,Jiajun Deng,Jiawen Kang,Shujie Hu,Tianzi Wang,Zhaoqing Li,Shiliang Zhang,Xie Chen,Xunying Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【24】 Energy Consumption Trends in Sound Event Detection Systems
标题: 声音事件检测系统的能源消耗趋势
作者:Constance Douwes,Romain Serizel
链接:点击下载PDF文件
【25】 DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
标题: DFADD:基于扩散和流匹配的音频Deepfake数据集
作者:Jiawei Du,I-Ming Lin,I-Hsiang Chiu,Xuanjun Chen,Haibin Wu,Wenze Ren,Yu Tsao,Hung-yi Lee,Jyh-Shing Roger Jang
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
【26】 Acoustic identification of individual animals with hierarchical contrastive learning
标题: 利用分层对比学习进行个体动物的声学识别
作者:Ines Nolasco,Ilyass Moummad,Dan Stowell,Emmanouil Benetos
备注:Under review; Submitted to ICASSP 2025
链接:点击下载PDF文件
【27】 Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
标题: 研究音素级语音编解码器中的解纠缠以进行韵律建模
作者:Sotirios Karapiperis,Nikolaos Ellinas,Alexandra Vioni,Junkwang Oh,Gunu Jho,Inchul Hwang,Spyros Raptis
链接:点击下载PDF文件
【28】 LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
标题: LMAC-TD:为音频分类器生成时间域简化
作者:Eleonora Mancini,Francesco Paissan,Mirco Ravanelli,Cem Subakan
备注:The first two authors contributed equally to this research. Author order is alphabetical
链接:点击下载PDF文件
【29】 Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
标题: 节奏Foley:视频到音频合成中无缝视听对齐的框架
作者:Zhiqi Huang,Dan Luo,Jun Wang,Huan Liao,Zhiheng Li,Zhiyong Wu
链接:点击下载PDF文件
【30】 TapToTab : Video-Based Guitar Tabs Generation using AI and Audio Analysis
标题: TapToTab:使用人工智能和音频分析的基于视频的吉他标签生成
作者:Ali Ghaleb,Eslam ElSadawy,Ihab Essam,Mohamed Abdelhakim,Seif-Eldin Zaki,Natalie Fahim,Razan Bayoumi,Hanan Hindy
链接:点击下载PDF文件
【31】 STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
标题: STA-V2 A:具有语义和时间对齐的视频到音频生成
作者:Yong Ren,Chenxing Li,Manjie Xu,Wei Liang,Yu Gu,Rilin Chen,Dong Yu
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
【32】 LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
标题: LA-RAG:通过检索增强生成增强基于LLM的ASB准确性
作者:Shaojun Li,Hengchao Shang,Daimeng Wei,Jiaxin Guo,Zongyao Li,Xianghui He,Min Zhang,Hao Yang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【33】 Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
标题: 大型语言模型可以在多说话者场景中通过多功能指令转录语音
作者:Lingwei Meng,Shujie Hu,Jiawen Kang,Zhaoqing Li,Yuejiao Wang,Wenxuan Wu,Xixin Wu,Xunying Liu,Helen Meng
链接:点击下载PDF文件
【34】 Domain-Invariant Representation Learning of Bird Sounds
标题: 鸟声的域不变表示学习
作者:Ilyass Moummad,Romain Serizel,Emmanouil Benetos,Nicolas Farrugia
链接:点击下载PDF文件
【35】 LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
标题: LHQ-SRC:轻量级、高质量的歌唱声音转换建模
作者:Yubo Huang,Xin Lai,Muyang Ye,Anran Zhu,Zixi Wang,Jingzehua Xu,Shuai Zhang,Zhiyuan Zhou,Weijie Niu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【36】 Apollo: Band-sequence Modeling for High-Quality Audio Restoration
标题: Apollo:用于高质量音频恢复的带序列建模
作者:Kai Li,Yi Luo
备注:Demo Page: this https URL
链接:点击下载PDF文件
【37】 Confidence Calibration for Audio Captioning Models
标题: 音频字幕模型的置信度校准
作者:Rehana Mahfuz,Yinyi Guo,Erik Visser
链接:点击下载PDF文件
标题: 利用对比预训练的神经音频嵌入进行推荐任务
作者:Florian Grötschla,Luca Strässle,Luca A. Lanzendörfer,Roger Wattenhofer
备注:Accepted at the 2nd Music Recommender Workshop (@RecSys)
链接:点击下载PDF文件
摘要:音乐推荐系统经常利用基于网络的模型来捕获音乐作品、艺术家和用户之间的关系。虽然这些关系为预测提供了有价值的见解,但由于初始信息不足,新的音乐作品或艺术家经常面临冷启动问题。为了解决这个问题,人们可以直接从音乐中提取基于内容的信息,以增强基于协作过滤的方法。虽然以前的方法依赖于手工制作的音频功能,但我们探索了对比预训练的神经音频嵌入模型的使用,它提供了更丰富,更细致的音乐表示。我们的实验表明,神经嵌入,特别是那些使用对比音频预训练(CLAP)模型生成的神经嵌入,为在基于图的框架内增强音乐推荐任务提供了一种很有前途的方法。摘要:Music recommender systems frequently utilize network-based models to capture relationships between music pieces, artists, and users. Although these relationships provide valuable insights for predictions, new music pieces or artists often face the cold-start problem due to insufficient initial information. To address this, one can extract content-based information directly from the music to enhance collaborative-filtering-based methods. While previous approaches have relied on hand-crafted audio features for this purpose, we explore the use of contrastively pretrained neural audio embedding models, which offer a richer and more nuanced representation of music. Our experiments demonstrate that neural embeddings, particularly those generated with the Contrastive Language-Audio Pretraining (CLAP) model, present a promising approach to enhancing music recommendation tasks within graph-based frameworks.
【2】 Biomimetic Frontend for Differentiable Audio Processing
标题: 可区分音频处理的仿生前沿
作者:Ruolan Leslie Famularo,Dmitry N. Zotkin,Shihab A. Shamma,Ramani Duraiswami
链接:点击下载PDF文件
摘要:虽然音频和语音处理中的模型变得越来越深入,越来越端到端,但它们因此需要在大数据上进行昂贵的训练,并且通常很脆弱。我们建立在人类听觉的经典模型上,并使其可区分,这样我们就可以将传统的可解释仿生信号处理方法与深度学习框架相结合。这使我们能够得到一个表达性和可解释的模型,该模型可以很容易地在少量数据上进行训练。我们将此模型应用于音频处理任务,包括分类和增强。结果表明,我们的可微模型在计算效率和鲁棒性方面优于黑箱方法,即使训练数据很少。我们还讨论了其他潜在的应用。摘要:While models in audio and speech processing are becoming deeper and more end-to-end, they as a consequence need expensive training on large data, and are often brittle. We build on a classical model of human hearing and make it differentiable, so that we can combine traditional explainable biomimetic signal processing approaches with deep-learning frameworks. This allows us to arrive at an expressive and explainable model that is easily trained on modest amounts of data. We apply this model to audio processing tasks, including classification and enhancement. Results show that our differentiable model surpasses black-box approaches in terms of computational efficiency and robustness, even with little training data. We also discuss other potential applications.
【3】 Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages
标题: 探索极低资源语言中数据量对ASB的影响
作者:Yao-Fei Cheng,Li-Wei Chen,Hung-Shin Lee,Hsin-Min Wang
链接:点击下载PDF文件
摘要:本研究探讨了数据增强技术在低资源自动语音识别(ASR)中的有效性,重点是两种濒危的南岛语言,Amis和Seediq。认识到自我监督学习(SSL)在低资源环境中的潜力,我们探索了数据量对SSL模型持续预训练的影响。我们提出了一种新颖的数据选择方案,利用多语言语料库来增加有限的目标语言数据。该方案利用语言分类器来提取话语嵌入,并利用一类分类器来识别在语音和音韵上与目标语言接近的话语。根据决策得分对语句进行排名和选择,确保在SSL-ASR管道中包含高度相关的数据。我们的实验结果表明,这种方法的有效性,产生了显着的改善,在ASR性能为Amis和Seediq。这些发现强调了通过跨语言迁移学习对低资源语言ASR进行数据增强的可行性和前景。摘要:This study investigates the efficacy of data augmentation techniques for low-resource automatic speech recognition (ASR), focusing on two endangered Austronesian languages, Amis and Seediq. Recognizing the potential of self-supervised learning (SSL) in low-resource settings, we explore the impact of data volume on the continued pre-training of SSL models. We propose a novel data-selection scheme leveraging a multilingual corpus to augment the limited target language data. This scheme utilizes a language classifier to extract utterance embeddings and employs one-class classifiers to identify utterances phonetically and phonologically proximate to the target languages. Utterances are ranked and selected based on their decision scores, ensuring the inclusion of highly relevant data in the SSL-ASR pipeline. Our experimental results demonstrate the effectiveness of this approach, yielding substantial improvements in ASR performance for both Amis and Seediq. These findings underscore the feasibility and promise of data augmentation through cross-lingual transfer learning for low-resource language ASR.
【4】 Exploring SSL Discrete Tokens for Multilingual ASR
标题: 探索多语言ASB的SSL离散令牌
作者:Mingyu Cui,Daxin Tan,Yifan Yang,Dingdong Wang,Huimeng Wang,Xiao Chen,Xie Chen,Xunying Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:随着自监督学习(SSL)在语音相关任务中的发展,人们越来越关注利用SSL生成的离散令牌进行自动语音识别(ASR),因为它们提供了更快的处理技术。然而,以前的研究主要集中在多语言ASR与Fbank功能或英语ASR与离散的令牌,留下了差距,在适应离散令牌的多语言ASR的情况下。这项研究提出了一个全面的比较离散令牌生成的各种领先的SSL模型在多个语言领域。我们的目标是探索性能和效率的语音离散令牌跨多个语言领域的单语和多语言的ASR方案。实验结果表明,离散令牌在七个语言领域的ASR任务中与使用Fbank特征训练的系统取得了相当的结果,平均单词错误率(WER)降低了0.31%和1.76%的绝对值(2.80%和15.70%相对)分别在开发和测试集,特别是WER减少6.82%的绝对(41.48%相对)波兰测试集。摘要:With the advancement of Self-supervised Learning (SSL) in speech-related tasks, there has been growing interest in utilizing discrete tokens generated by SSL for automatic speech recognition (ASR), as they offer faster processing techniques. However, previous studies primarily focused on multilingual ASR with Fbank features or English ASR with discrete tokens, leaving a gap in adapting discrete tokens for multilingual ASR scenarios. This study presents a comprehensive comparison of discrete tokens generated by various leading SSL models across multiple language domains. We aim to explore the performance and efficiency of speech discrete tokens across multiple language domains for both monolingual and multilingual ASR scenarios. Experimental results demonstrate that discrete tokens achieve comparable results against systems trained on Fbank features in ASR tasks across seven language domains with an average word error rate (WER) reduction of 0.31% and 1.76% absolute (2.80% and 15.70% relative) on dev and test sets respectively, with particularly WER reduction of 6.82% absolute (41.48% relative) on the Polish test set.
【5】 Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
标题: 探索基于Zipformer的上下文ASB的SSL离散语音特征
作者:Mingyu Cui,Yifan Yang,Jiajun Deng,Jiawen Kang,Shujie Hu,Tianzi Wang,Zhaoqing Li,Shiliang Zhang,Xie Chen,Xunying Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:基于自监督学习(SSL)的离散语音表示是高度紧凑和领域自适应的。在本文中,SSL离散语音特征提取的WavLM模型被用来作为额外的跨话语声学上下文特征的Zipformer-Transducer ASR系统。取代Fbank功能与离散令牌功能建模跨话语上下文(从前面和未来的段),或当前话语的内部上下文单独,或两者同时,彻底证明Gigaspeech 1000小时语料库的功效。最好的Zipformer-Transducer系统使用基于离散标记的跨话语上下文特征,其性能优于使用话语内部上下文的基线,仅在开发和测试数据上具有0.32%至0.41%的绝对统计显著的单词错误率(WER)降低(2.78%至3.54%相对)。在开发集和测试集上获得了11.15%和11.14%的最低公布WER。我们的工作是开源的,可以在https: github.com open-creator icefall tree master egs gigaspeech Context _ASR上公开获取。摘要:Self-supervised learning (SSL) based discrete speech representations are highly compact and domain adaptable. In this paper, SSL discrete speech features extracted from WavLM models are used as additional cross-utterance acoustic context features in Zipformer-Transducer ASR systems. The efficacy of replacing Fbank features with discrete token features for modelling either cross-utterance contexts (from preceding and future segments), or current utterance's internal contexts alone, or both at the same time, are demonstrated thoroughly on the Gigaspeech 1000-hr corpus. The best Zipformer-Transducer system using discrete tokens based cross-utterance context features outperforms the baseline using utterance internal context only with statistically significant word error rate (WER) reductions of 0.32% to 0.41% absolute (2.78% to 3.54% relative) on the dev and test data. The lowest published WER of 11.15% and 11.14% were obtained on the dev and test sets. Our work is open-source and publicly available at https: github.com open-creator icefall tree master egs gigaspeech Context _ASR.
【6】 Energy Consumption Trends in Sound Event Detection Systems
标题: 声音事件检测系统的能源消耗趋势
作者:Constance Douwes,Romain Serizel
链接:点击下载PDF文件
摘要:深度学习系统变得越来越耗能和计算密集,引发了人们对其环境影响的担忧。作为声学场景和事件的检测和分类(DCASE)挑战的组织者,我们认识到解决这个问题的重要性。在过去的三年中,我们将能耗指标集成到声音事件检测(SED)系统的评估中。在本文中,我们分析了这种能量标准对挑战结果的影响,并探讨了多年来系统复杂性和能耗的演变。我们强调在培训过程中向更节能的方法转变,而不影响性能,同时操作数量和系统复杂性继续增长。我们希望透过这项分析,在社会经济发展界别内推广更环保的做法。摘要:Deep learning systems have become increasingly energy- and computation-intensive, raising concerns about their environmental impact. As organizers of the Detection and Classification of Acoustic Scenes and Events (DCASE) challenge, we recognize the importance of addressing this issue. For the past three years, we have integrated energy consumption metrics into the evaluation of sound event detection (SED) systems. In this paper, we analyze the impact of this energy criterion on the challenge results and explore the evolution of system complexity and energy consumption over the years. We highlight a shift towards more energy-efficient approaches during training without compromising performance, while the number of operations and system complexity continue to grow. Through this analysis, we hope to promote more environmentally friendly practices within the SED community.
【7】 DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
标题: DFADD:基于扩散和流匹配的音频Deepfake数据集
作者:Jiawei Du,I-Ming Lin,I-Hsiang Chiu,Xuanjun Chen,Haibin Wu,Wenze Ren,Yu Tsao,Hung-yi Lee,Jyh-Shing Roger Jang
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
摘要:主流的zero-shot TTS生成系统,如Voicebox和Seed-TTS,分别通过利用流匹配和扩散模型来实现人类对等语音。不幸的是,人类级别的音频合成会导致身份滥用和信息安全问题。目前,已经针对deepfake音频开发了许多反欺骗模型。然而,当前最先进的反欺骗模型在对抗由基于扩散和流匹配的TTS系统合成的音频方面的功效仍然未知。在本文中,我们提出了基于扩散和流匹配的音频Deepfake(DFADD)数据集。DFADD数据集收集了基于高级扩散和流匹配TTS模型的deepfake音频。此外,我们发现,目前的反欺骗模型缺乏足够的鲁棒性对扩散和流匹配TTS系统产生的高度人性化的音频。建议的DFADD数据集解决了这一差距,并为开发更具弹性的反欺骗模型提供了宝贵的资源。摘要:Mainstream zero-shot TTS production systems like Voicebox and Seed-TTS achieve human parity speech by leveraging Flow-matching and Diffusion models, respectively. Unfortunately, human-level audio synthesis leads to identity misuse and information security issues. Currently, many antispoofing models have been developed against deepfake audio. However, the efficacy of current state-of-the-art anti-spoofing models in countering audio synthesized by diffusion and flowmatching based TTS systems remains unknown. In this paper, we proposed the Diffusion and Flow-matching based Audio Deepfake (DFADD) dataset. The DFADD dataset collected the deepfake audio based on advanced diffusion and flowmatching TTS models. Additionally, we reveal that current anti-spoofing models lack sufficient robustness against highly human-like audio generated by diffusion and flow-matching TTS systems. The proposed DFADD dataset addresses this gap and provides a valuable resource for developing more resilient anti-spoofing models.
【8】 Acoustic identification of individual animals with hierarchical contrastive learning
标题: 利用分层对比学习进行个体动物的声学识别
作者:Ines Nolasco,Ilyass Moummad,Dan Stowell,Emmanouil Benetos
备注:Under review; Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:个体动物的声学识别(AIID)与基于音频的物种分类密切相关,但需要更精细的细节来区分同一物种中的个体动物。在这项工作中,我们将AIID框架为分层多标签分类任务,并建议使用分层感知损失函数来学习个体身份的鲁棒表示,以维护物种和分类群之间的分层关系。我们的研究结果表明,层次嵌入不仅提高了识别精度在个人层面上,而且在更高的分类水平,有效地保留了层次结构的学习表示。通过比较我们的方法与非层次模型,我们强调了在嵌入空间中执行这种结构的优势。此外,我们将评估扩展到新的单个类的分类,展示了我们的方法在开集分类场景中的潜力。摘要:Acoustic identification of individual animals (AIID) is closely related to audio-based species classification but requires a finer level of detail to distinguish between individual animals within the same species. In this work, we frame AIID as a hierarchical multi-label classification task and propose the use of hierarchy-aware loss functions to learn robust representations of individual identities that maintain the hierarchical relationships among species and taxa. Our results demonstrate that hierarchical embeddings not only enhance identification accuracy at the individual level but also at higher taxonomic levels, effectively preserving the hierarchical structure in the learned representations. By comparing our approach with non-hierarchical models, we highlight the advantage of enforcing this structure in the embedding space. Additionally, we extend the evaluation to the classification of novel individual classes, demonstrating the potential of our method in open-set classification scenarios.
【9】 Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
标题: 研究音素级语音编解码器中的解纠缠以进行韵律建模
作者:Sotirios Karapiperis,Nikolaos Ellinas,Alexandra Vioni,Junkwang Oh,Gunu Jho,Inchul Hwang,Spyros Raptis
链接:点击下载PDF文件
摘要:语音韵律建模中的大多数流行方法依赖于在一个连续的潜在空间中学习全局风格表示,该潜在空间编码和传递参考语音的属性。然而,最近基于残差矢量量化(RVQ)的神经编解码器的工作已经显示出巨大的潜力,提供了明显的优势。我们调查的韵律建模能力的离散空间这样的RVQ-VAE模型,修改它的音素水平上操作。我们的条件的编码器和解码器的语言表示的模型,并应用一个全球扬声器嵌入,以因素出语音和扬声器的信息。我们进行了广泛的调查,主观实验和客观措施的基础上,以表明,音素级的离散潜在表示,这种方式获得了高度的解开,捕获细粒度的韵律信息,是强大的和可转让的。潜在空间具有可解释的结构,其主成分对应于音高和能量。摘要:Most of the prevalent approaches in speech prosody modeling rely on learning global style representations in a continuous latent space which encode and transfer the attributes of reference speech. However, recent work on neural codecs which are based on Residual Vector Quantization (RVQ) already shows great potential offering distinct advantages. We investigate the prosody modeling capabilities of the discrete space of such an RVQ-VAE model, modifying it to operate on the phoneme-level. We condition both the encoder and decoder of the model on linguistic representations and apply a global speaker embedding in order to factor out both phonetic and speaker information. We conduct an extensive set of investigations based on subjective experiments and objective measures to show that the phoneme-level discrete latent representations obtained this way achieves a high degree of disentanglement, capturing fine-grained prosodic information that is robust and transferable. The latent space turns out to have interpretable structure with its principal components corresponding to pitch and energy.
【10】 LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
标题: LMAC-TD:为音频分类器生成时间域简化
作者:Eleonora Mancini,Francesco Paissan,Mirco Ravanelli,Cem Subakan
备注:The first two authors contributed equally to this research. Author order is alphabetical
链接:点击下载PDF文件
摘要:神经网络通常是黑箱,在决策机制方面保持不透明。文献中的一些作品提出了事后解释方法来缓解这个问题。本文提出了LMAC-TD,一种事后解释方法,训练解码器直接在时域中产生解释。这种方法建立在L-MAC的基础上,L-MAC是一种用于音频分类器的可听映射,这种方法可以产生忠实和可解释的解释。我们将SepFormer,一个流行的基于变压器的时域源分离架构。我们通过一项用户研究表明,LMAC-TD显着提高了音频质量的生产解释,同时不牺牲忠诚。摘要:Neural networks are typically black-boxes that remain opaque with regards to their decision mechanisms. Several works in the literature have proposed post-hoc explanation methods to alleviate this issue. This paper proposes LMAC-TD, a post-hoc explanation method that trains a decoder to produce explanations directly in the time domain. This methodology builds upon the foundation of L-MAC, Listenable Maps for Audio Classifiers, a method that produces faithful and listenable explanations. We incorporate SepFormer, a popular transformer-based time-domain source separation architecture. We show through a user study that LMAC-TD significantly improves the audio quality of the produced explanations while not sacrificing from faithfulness.
【11】 Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
标题: 节奏Foley:视频到音频合成中无缝视听对齐的框架
作者:Zhiqi Huang,Dan Luo,Jun Wang,Huan Liao,Zhiheng Li,Zhiyong Wu
链接:点击下载PDF文件
摘要:我们的研究引入了一个创新的框架,视频到音频合成,解决了音视频去标准化和语义损失的音频。通过结合语义对齐适配器和时间同步适配器,我们的方法显着提高了语义的完整性和节拍点同步的精度,特别是在快节奏的动作序列。利用对比视听预训练编码器,我们的模型使用视频和高质量音频数据进行训练,提高了生成音频的质量。这种双适配器方法使用户能够增强对音频语义和节拍效果的控制,允许调整控制器以实现更好的效果。大量的实验证实了我们的框架在实现无缝视听对齐的有效性。摘要:Our research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal synchronization adapter, our method significantly improves semantic integrity and the precision of beat point synchronization, particularly in fast-paced action sequences. Utilizing a contrastive audio-visual pre-trained encoder, our model is trained with video and high-quality audio data, improving the quality of the generated audio. This dual-adapter approach empowers users with enhanced control over audio semantics and beat effects, allowing the adjustment of the controller to achieve better results. Extensive experiments substantiate the effectiveness of our framework in achieving seamless audio-visual alignment.
【12】 TapToTab : Video-Based Guitar Tabs Generation using AI and Audio Analysis
标题: TapToTab:使用人工智能和音频分析的基于视频的吉他标签生成
作者:Ali Ghaleb,Eslam ElSadawy,Ihab Essam,Mohamed Abdelhakim,Seif-Eldin Zaki,Natalie Fahim,Razan Bayoumi,Hanan Hindy
链接:点击下载PDF文件
摘要:从视频输入的吉他指板生成的自动化对于提高音乐教育、转录准确性和性能分析具有重要的前景。现有方法面临一致性和完整性的挑战,特别是在检测指板和准确识别音符方面。为了解决这些问题,本文介绍了一种利用深度学习的先进方法,特别是用于实时指板检测的YOLO模型,以及用于精确音符识别的基于傅立叶变换的音频分析。实验结果表明,与传统技术相比,检测精度和鲁棒性有了很大的提高。本文概述了这些方法的开发,实施和评估,旨在通过自动化从视频录音创建吉他标签来彻底改变吉他教学。摘要:The automation of guitar tablature generation from video inputs holds significant promise for enhancing music education, transcription accuracy, and performance analysis. Existing methods face challenges with consistency and completeness, particularly in detecting fretboards and accurately identifying notes. To address these issues, this paper introduces an advanced approach leveraging deep learning, specifically YOLO models for real-time fretboard detection, and Fourier Transform-based audio analysis for precise note identification. Experimental results demonstrate substantial improvements in detection accuracy and robustness compared to traditional techniques. This paper outlines the development, implementation, and evaluation of these methodologies, aiming to revolutionize guitar instruction by automating the creation of guitar tabs from video recordings.
【13】 STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
标题: STA-V2 A:具有语义和时间对齐的视频到音频生成
作者:Yong Ren,Chenxing Li,Manjie Xu,Wei Liang,Yu Gu,Rilin Chen,Dong Yu
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
摘要:视觉和听觉是人类体验世界的两种重要方式。在过去的一年里,文本到视频的生成取得了显着的进展,但在生成的视频中缺乏和谐的音频限制了其更广泛的应用。在本文中,我们提出了语义和时间对齐的视频到音频(STA-V2 A),一种方法,通过提取本地时间和全局语义视频特征,并将这些细化的视频特征与文本相结合,作为跨模态指导,增强从视频中生成音频。为了解决视频中的信息冗余问题,我们提出了一个发病预测借口任务的局部时间特征提取和关注池模块的全局语义特征提取。为了补充视频中语义信息的不足,我们提出了一个带有文本到音频先验初始化和跨模态指导的潜在扩散模型。我们还介绍了Audio-Audio Align,一个新的指标来评估音频时间对齐。主观和客观的指标表明,我们的方法超越现有的视频到音频模型生成的音频具有更好的质量,语义一致性和时间对齐。烧蚀实验验证了各模块的有效性。音频样本可在https: y-ren16.github.io STAV2A上获得。摘要:Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader applications. In this paper, we propose Semantic and Temporal Aligned Video-to-Audio (STA-V2A), an approach that enhances audio generation from videos by extracting both local temporal and global semantic video features and combining these refined video features with text as cross-modal guidance. To address the issue of information redundancy in videos, we propose an onset prediction pretext task for local temporal feature extraction and an attentive pooling module for global semantic feature extraction. To supplement the insufficient semantic information in videos, we propose a Latent Diffusion Model with Text-to-Audio priors initialization and cross-modal guidance. We also introduce Audio-Audio Align, a new metric to assess audio-temporal alignment. Subjective and objective metrics demonstrate that our method surpasses existing Video-to-Audio models in generating audio with better quality, semantic consistency, and temporal alignment. The ablation experiment validated the effectiveness of each module. Audio samples are available at https: y-ren16.github.io STAV2A.
【14】 LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
标题: LA-RAG:通过检索增强生成增强基于LLM的ASB准确性
作者:Shaojun Li,Hengchao Shang,Daimeng Wei,Jiaxin Guo,Zongyao Li,Xianghui He,Min Zhang,Hao Yang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:将语音信息集成到大型语言模型(LLM)中的最新进展显着提高了自动语音识别(ASR)的准确性。然而,现有的方法往往受到语音编码器在不同声学条件下的能力的限制,例如口音。为了解决这个问题,我们提出了LA-RAG,一种新的检索增强生成(RAG)范式,用于基于LLM的ASR。LA-RAG利用细粒度令牌级语音数据存储和语音到语音检索机制,通过LLM上下文学习(ICL)功能提高ASR准确性。在普通话和各种中国方言数据集上的实验表明,与现有方法相比,ASR的准确性有了显着提高,验证了我们的方法的有效性,特别是在处理口音变化方面。摘要:Recent advancements in integrating speech information into large language models (LLMs) have significantly improved automatic speech recognition (ASR) accuracy. However, existing methods often constrained by the capabilities of the speech encoders under varied acoustic conditions, such as accents. To address this, we propose LA-RAG, a novel Retrieval-Augmented Generation (RAG) paradigm for LLM-based ASR. LA-RAG leverages fine-grained token-level speech datastores and a speech-to-speech retrieval mechanism to enhance ASR accuracy via LLM in-context learning (ICL) capabilities. Experiments on Mandarin and various Chinese dialect datasets demonstrate significant improvements in ASR accuracy compared to existing methods, validating the effectiveness of our approach, especially in handling accent variations.
【15】 Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
标题: 大型语言模型可以在多说话者场景中通过多功能指令转录语音
作者:Lingwei Meng,Shujie Hu,Jiawen Kang,Zhaoqing Li,Yuejiao Wang,Wenxuan Wu,Xixin Wu,Xunying Liu,Helen Meng
链接:点击下载PDF文件
摘要:大型语言模型(LLM)的最新进展已经彻底改变了各个领域,带来了重大进展和新机遇。尽管在语音相关的任务取得了进展,LLM还没有在多说话者场景中得到充分的探索。在这项工作中,我们提出了一个开创性的努力,调查LLM在多人环境中转录语音的能力,以下多功能指令相关的多人自动语音识别(ASR),目标说话者ASR和ASR的基础上,特定的说话者属性,如性别,发生顺序,语言和关键字发言。我们的方法利用WavLM和Whisper编码器来提取多方面的语音表示,是敏感的扬声器特性和语义上下文。然后,这些表示被输入到使用LoRA进行微调的LLM中,从而实现语音理解和转录的功能。综合实验揭示了我们提出的系统,MT-LLM,在鸡尾酒会的情况下,有前途的性能,突出LLM处理语音相关的任务,在这样复杂的设置基于用户指令的潜力。摘要:Recent advancements in large language models (LLMs) have revolutionized various domains, bringing significant progress and new opportunities. Despite progress in speech-related tasks, LLMs have not been sufficiently explored in multi-talker scenarios. In this work, we present a pioneering effort to investigate the capability of LLMs in transcribing speech in multi-talker environments, following versatile instructions related to multi-talker automatic speech recognition (ASR), target talker ASR, and ASR based on specific talker attributes such as sex, occurrence order, language, and keyword spoken. Our approach utilizes WavLM and Whisper encoder to extract multi-faceted speech representations that are sensitive to speaker characteristics and semantic context. These representations are then fed into an LLM fine-tuned using LoRA, enabling the capabilities for speech comprehension and transcription. Comprehensive experiments reveal the promising performance of our proposed system, MT-LLM, in cocktail party scenarios, highlighting the potential of LLM to handle speech-related tasks based on user instructions in such complex settings.
【16】 Domain-Invariant Representation Learning of Bird Sounds
标题: 鸟声的域不变表示学习
作者:Ilyass Moummad,Romain Serizel,Emmanouil Benetos,Nicolas Farrugia
链接:点击下载PDF文件
摘要:被动声监测(PAM)是生物声学研究的关键,可以实现非侵入性物种跟踪和生物多样性监测。像Xeno-Canto这样的公民科学平台提供了来自焦点记录的大型注释数据集,其中目标物种被有意记录。然而,PAM需要在被动音景中进行监控,在焦点和被动录音之间产生域转移,这对在焦点录音上训练的深度学习模型提出了挑战。为了解决这个问题,我们利用监督对比学习来提高鸟类声音分类中的领域泛化,在不同领域的同类样本中实现领域不变性。我们还提出了Prototypical(原型对比学习的表示),它降低了计算复杂性的SupCon损失比较的例子类原型,而不是成对的比较。此外,我们提出了一个新的Few-Shot分类基准的基础上BirdSet,一个大规模的鸟的声音数据集,并证明了我们的方法在实现强大的传输性能的有效性。摘要:Passive acoustic monitoring (PAM) is crucial for bioacoustic research, enabling non-invasive species tracking and biodiversity monitoring. Citizen science platforms like Xeno-Canto provide large annotated datasets from focal recordings, where the target species is intentionally recorded. However, PAM requires monitoring in passive soundscapes, creating a domain shift between focal and passive recordings, which challenges deep learning models trained on focal recordings. To address this, we leverage supervised contrastive learning to improve domain generalization in bird sound classification, enforcing domain invariance across same-class examples from different domains. We also propose ProtoCLR (Prototypical Contrastive Learning of Representations), which reduces the computational complexity of the SupCon loss by comparing examples to class prototypes instead of pairwise comparisons. Additionally, we present a new few-shot classification benchmark based on BirdSet, a large-scale bird sound dataset, and demonstrate the effectiveness of our approach in achieving strong transfer performance.
【17】 LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
标题: LHQ-SRC:轻量级、高质量的歌唱声音转换建模
作者:Yubo Huang,Xin Lai,Muyang Ye,Anran Zhu,Zixi Wang,Jingzehua Xu,Shuai Zhang,Zhiyuan Zhou,Weijie Niu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:歌唱声音转换(SVC)已经成为声音转换(VC)的一个重要子领域,使一个歌手的声音转换成另一个,同时保留音乐元素,如旋律,节奏和音色。传统的SVC方法在音频质量、数据要求和计算复杂度方面具有局限性。在本文中,我们提出了LHQ-SVC,一个轻量级的,CPU兼容的模型的基础上的SVC框架和扩散模型,旨在减少模型的大小和计算需求,而不牺牲性能。我们整合了一些功能来提高推理质量,并通过使用性能调优工具和并行计算框架来优化CPU执行。我们的实验表明,LHQ-SVC保持有竞争力的性能,在不同的设备处理速度和效率显着提高。结果表明,LHQ-SVC可以满足摘要:Singing Voice Conversion (SVC) has emerged as a significant subfield of Voice Conversion (VC), enabling the transformation of one singer's voice into another while preserving musical elements such as melody, rhythm, and timbre. Traditional SVC methods have limitations in terms of audio quality, data requirements, and computational complexity. In this paper, we propose LHQ-SVC, a lightweight, CPU-compatible model based on the SVC framework and diffusion model, designed to reduce model size and computational demand without sacrificing performance. We incorporate features to improve inference quality, and optimize for CPU execution by using performance tuning tools and parallel computing frameworks. Our experiments demonstrate that LHQ-SVC maintains competitive performance, with significant improvements in processing speed and efficiency across different devices. The results suggest that LHQ-SVC can meet
【18】 Apollo: Band-sequence Modeling for High-Quality Audio Restoration
标题: Apollo:用于高质量音频恢复的带序列建模
作者:Kai Li,Yi Luo
备注:Demo Page: this https URL
链接:点击下载PDF文件
摘要:音频恢复在现代社会中变得越来越重要,这不仅是因为对高级回放设备所实现的高质量听觉体验的需求,而且还因为生成音频模型的不断增长的能力需要高保真音频。通常,音频恢复被定义为从受损输入预测未失真音频的任务,通常使用GAN框架进行训练以平衡感知和失真。由于音频降级主要集中在中高频范围,特别是由于编解码器,因此一个关键的挑战在于设计一个能够保留低频信息同时准确重建高质量中高频内容的生成器。受高采样率音乐分离,语音增强和音频编解码器模型的最新进展的启发,我们提出了阿波罗,一个生成模型设计的高采样率音频恢复。Apollo采用了一个显式的频带分割模块来模拟不同频带之间的关系,从而实现更连贯和更高质量的恢复音频。在MUSDB 18-HQ和MoisesDB数据集上进行评估,Apollo在各种比特率和音乐类型上的表现始终优于现有的SR-GAN模型,特别是在涉及多种乐器和人声混合的复杂场景中表现出色。Apollo在保持计算效率的同时显著提高了音乐恢复质量。Apollo的源代码可在https: github.com JusperLee Apollo上公开获取。摘要:Audio restoration has become increasingly significant in modern society, not only due to the demand for high-quality auditory experiences enabled by advanced playback devices, but also because the growing capabilities of generative audio models necessitate high-fidelity audio. Typically, audio restoration is defined as a task of predicting undistorted audio from damaged input, often trained using a GAN framework to balance perception and distortion. Since audio degradation is primarily concentrated in mid- and high-frequency ranges, especially due to codecs, a key challenge lies in designing a generator capable of preserving low-frequency information while accurately reconstructing high-quality mid- and high-frequency content. Inspired by recent advancements in high-sample-rate music separation, speech enhancement, and audio codec models, we propose Apollo, a generative model designed for high-sample-rate audio restoration. Apollo employs an explicit frequency band split module to model the relationships between different frequency bands, allowing for more coherent and higher-quality restored audio. Evaluated on the MUSDB18-HQ and MoisesDB datasets, Apollo consistently outperforms existing SR-GAN models across various bit rates and music genres, particularly excelling in complex scenarios involving mixtures of multiple instruments and vocals. Apollo significantly improves music restoration quality while maintaining computational efficiency. The source code for Apollo is publicly available at https: github.com JusperLee Apollo.
【19】 Confidence Calibration for Audio Captioning Models
标题: 音频字幕模型的置信度校准
作者:Rehana Mahfuz,Yinyi Guo,Erik Visser
链接:点击下载PDF文件
摘要:自动生成音频、图像和视频的文本字幕的系统缺乏所生成序列的相关性和正确性的置信度指示符。为了解决这个问题,我们在现有的文本置信度测量方法的基础上,引入了标记概率的选择性池化,这比传统的池化更符合传统的正确性测量。此外,我们建议直接测量输入音频和文本之间的相似性在一个共享的嵌入空间。为了测量自我一致性,我们适应语义熵的音频字幕,并发现这两种方法甚至比池为基础的指标与正确性的措施,计算字幕之间的声学相似性。最后,我们解释了为什么温度缩放的信心,提高校准。摘要:Systems that automatically generate text captions for audio, images and video lack a confidence indicator of the relevance and correctness of the generated sequences. To address this, we build on existing methods of confidence measurement for text by introduce selective pooling of token probabilities, which aligns better with traditional correctness measures than conventional pooling does. Further, we propose directly measuring the similarity between input audio and text in a shared embedding space. To measure self-consistency, we adapt semantic entropy for audio captioning, and find that these two methods align even better than pooling-based metrics with the correctness measure that calculates acoustic similarity between captions. Finally, we explain why temperature scaling of confidences improves calibration.
【20】 Why some audio signal short-time Fourier transform coefficients have nonuniform phase distributions
标题: 为什么某些音频信号短期傅里叶变换系数具有不均匀的相分布
作者:Stephen D. Voran
Journal-ref:Proceedings of the 2024 IEEE International Conference on Multimedia and Expo, Niagara Falls, Ontario, July 15-19, 2024
链接:点击下载PDF文件
摘要:短时傅立叶变换(STFT)将音频样本的窗口表示为一组复系数。这些被有利地视为幅度和相位,并且相位的总体分布通常被假设为均匀的。我们表明,当音频信号STFT相位分布分析每个频率或幅度范围,它们可以远离均匀。也就是说,均匀相位分布假设掩盖了重要的重要细节。我们解释了不均匀的相位分布的意义,以及如何利用它们,推导出它们的来源,并解释了为什么选择的STFT窗口形状的影响所得到的相位分布的不均匀性。摘要:The short-time Fourier transform (STFT) represents a window of audio samples as a set of complex coefficients. These are advantageously viewed as magnitudes and phases and the overall distribution of phases is very often assumed to be uniform. We show that when audio signal STFT phase distributions are analyzed per-frequency or per-magnitude range, they can be far from uniform. That is, the uniform phase distribution assumption obscures significant important details. We explain the significance of the nonuniform phase distributions and how they might be exploited, derive their source, and explain why the choice of the STFT window shape influences the nonuniformity of the resulting phase distributions.
【21】 Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
标题: 使用耳脑电解码多说话者环境中的听觉注意力
作者:Haolin Zhu,Yujie Yan,Xiran Xu,Zhongshu Ge,Pei Tian,Xihong Wu,Jing Chen
链接:点击下载PDF文件
摘要:听觉注意解码(AAD)可以通过分析和处理脑电图(EEG)数据来帮助确定听觉选择性注意任务期间参与说话者的身份。大多数关于AAD的研究都是基于双说话人场景下的头皮脑电信号,与实际应用相距甚远。耳脑电图最近获得了显着的关注,由于其运动的宽容和不可见性在数据采集过程中,使其易于与其他设备的应用程序。在这项工作中,参与者有选择地参加了四个空间分离的扬声器的语音在消声室。EEG数据同时从头皮EEG系统和耳EEG系统(cEEGrids)收集。使用耳EEG数据的时间响应函数(TRFs)和刺激重建(SR)。结果表明,在60年代(机会水平为25%),有注意语音的TRFs比无注意语音的TRFs强,解码正确率为41.3%。为了进一步研究电极放置和数量的影响,在头皮EEG和耳EEG中使用SR,揭示了虽然电极数量具有较小的影响,但它们的位置对解码准确性具有显著影响。在此耳脑电数据库上验证了一种听觉空间注意检测方法STAnet,在1秒解码窗口内的正确率为93.1%。我们工作的实现代码和数据库可在GitHub:https: github.com zhl486 Ear_EEG_code.git和Zenodo:https: zenodo.org records 10803261上找到。摘要:Auditory Attention Decoding (AAD) can help to determine the identity of the attended speaker during an auditory selective attention task, by analyzing and processing measurements of electroencephalography (EEG) data. Most studies on AAD are based on scalp-EEG signals in two-speaker scenarios, which are far from real application. Ear-EEG has recently gained significant attention due to its motion tolerance and invisibility during data acquisition, making it easy to incorporate with other devices for applications. In this work, participants selectively attended to one of the four spatially separated speakers' speech in an anechoic room. The EEG data were concurrently collected from a scalp-EEG system and an ear-EEG system (cEEGrids). Temporal response functions (TRFs) and stimulus reconstruction (SR) were utilized using ear-EEG data. Results showed that the attended speech TRFs were stronger than each unattended speech and decoding accuracy was 41.3 % in the 60s (chance level of 25 %). To further investigate the impact of electrode placement and quantity, SR was utilized in both scalp-EEG and ear-EEG, revealing that while the number of electrodes had a minor effect, their positioning had a significant influence on the decoding accuracy. One kind of auditory spatial attention detection (ASAD) method, STAnet, was testified with this ear-EEG database, resulting in 93.1% in 1-second decoding window. The implementation code and database for our work are available on GitHub: https: github.com zhl486 Ear_EEG_code.git and Zenodo: https: zenodo.org records 10803261.
【22】 DualSep: A Light-weight dual-encoder convolutional recurrent network for real-time in-car speech separation
标题: DualSep:用于实时车内语音分离的轻量级双编码器卷积循环网络
作者:Ziqian Wang,Jiayao Sun,Zihan Zhang,Xingchen Li,Jie Liu,Lei Xie
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
摘要:深度学习和语音激活技术的进步推动了人车交互的发展。分布式麦克风阵列广泛应用于车内场景,因为它可以准确地捕获来自不同语音区域的乘客的声音。然而,音频通道数量的增加,加上有限的计算资源和车载系统的低延迟要求,为车载多通道语音分离提出了挑战。为了迁移的问题,我们提出了一个轻量级的框架,级联数字信号处理(DSP)和神经网络(NN)。我们利用固定波束形成(BF),以减少计算成本和独立向量分析(IVA)提供空间先验。我们采用双编码器进行双分支建模,空间编码器捕获空间线索,光谱编码器保留光谱信息,促进空间-光谱融合。我们提出的系统支持流和非流模式。实验结果表明,该系统在各种指标的优越性。在Intel Core i7(2.6GHz)CPU上仅具有0.83M参数和0.39实时因子(RTF),它有效地将语音分离到不同的语音区域。我们的演示可在https: honee-w.github.io DualSep 上获得。摘要:Advancements in deep learning and voice-activated technologies have driven the development of human-vehicle interaction. Distributed microphone arrays are widely used in in-car scenarios because they can accurately capture the voices of passengers from different speech zones. However, the increase in the number of audio channels, coupled with the limited computational resources and low latency requirements of in-car systems, presents challenges for in-car multi-channel speech separation. To migrate the problems, we propose a lightweight framework that cascades digital signal processing (DSP) and neural networks (NN). We utilize fixed beamforming (BF) to reduce computational costs and independent vector analysis (IVA) to provide spatial prior. We employ dual encoders for dual-branch modeling, with spatial encoder capturing spatial cues and spectral encoder preserving spectral information, facilitating spatial-spectral fusion. Our proposed system supports both streaming and non-streaming modes. Experimental results demonstrate the superiority of the proposed system across various metrics. With only 0.83M parameters and 0.39 real-time factor (RTF) on an Intel Core i7 (2.6GHz) CPU, it effectively separates speech into distinct speech zones. Our demos are available at https: honee-w.github.io DualSep .
【23】 Effective Integration of KAN for Keyword Spotting
标题: 有效集成KAN以识别关键词
作者:Anfeng Xu,Biqiao Zhang,Shuyu Kong,Yiteng Huang,Zhaojun Yang,Sangeeta Srivastava,Ming Sun
备注:Under review
链接:点击下载PDF文件
摘要:关键词识别(KWS)是具有语音辅助功能的智能设备的重要语音处理组件。本文研究了Kolmogorov-Arnold网络(KAN)是否可以用于增强KWS的性能。我们探索各种方法来集成KAN的模型架构的基础上一维卷积神经网络(CNN)。我们发现KAN可以有效地在低维空间中建模高级特征,从而在适当集成时提高KWS性能。研究结果揭示了理解KAN的语音处理任务和其他模式,为未来的研究人员。摘要:Keyword spotting (KWS) is an important speech processing component for smart devices with voice assistance capability. In this paper, we investigate if Kolmogorov-Arnold Networks (KAN) can be used to enhance the performance of KWS. We explore various approaches to integrate KAN for a model architecture based on 1D Convolutional Neural Networks (CNN). We find that KAN is effective at modeling high-level features in lower-dimensional spaces, resulting in improved KWS performance when integrated appropriately. The findings shed light on understanding KAN for speech processing tasks and on other modalities for future researchers.
【24】 Frequency Tracking Features for Data-Efficient Deep Siren Identification
标题: 频率跟踪功能,实现数据高效的深度警报识别
作者:Stefano Damiano,Thomas Dietzen,Toon van Waterschoot
备注:Accepted paper: Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2024)
链接:点击下载PDF文件
摘要:识别城市音景中的警笛声音是智能车辆的一个重要安全方面,并已通过神经网络得到广泛解决,这些神经网络确保了对警笛信号多样性和交通特征的强非结构化背景噪声的鲁棒性。卷积神经网络分析输入信号的频谱图特征时,达到最先进的性能,足够的训练数据捕获的目标声学场景的多样性是可用的。在实践中,数据通常是有限的,算法应该是鲁棒的,以适应看不见的声学条件,而不需要大量的数据集进行重新训练。在这项工作中,给出了警笛信号的谐波性质,其特征在于由一个周期性演变的基频,我们提出了一种低复杂度的特征提取方法的基础上使用单参数自适应陷波滤波器的频率跟踪。然后,这些特征被用来设计一个小规模的卷积网络,适合用有限的数据进行训练。实验结果表明,在训练数据有限的情况下,该模型的性能始终优于传统的基于谱图的模型,具有更好的跨域泛化能力和更小的规模。摘要:The identification of siren sounds in urban soundscapes is a crucial safety aspect for smart vehicles and has been widely addressed by means of neural networks that ensure robustness to both the diversity of siren signals and the strong and unstructured background noise characterizing traffic. Convolutional neural networks analyzing spectrogram features of incoming signals achieve state-of-the-art performance when enough training data capturing the diversity of the target acoustic scenes is available. In practice, data is usually limited and algorithms should be robust to adapt to unseen acoustic conditions without requiring extensive datasets for re-training. In this work, given the harmonic nature of siren signals, characterized by a periodically evolving fundamental frequency, we propose a low-complexity feature extraction method based on frequency tracking using a single-parameter adaptive notch filter. The features are then used to design a small-scale convolutional network suitable for training with limited data. The evaluation results indicate that the proposed model consistently outperforms the traditional spectrogram-based model when limited training data is available, achieves better cross-domain generalization and has a smaller size.
【25】 Unified Audio Event Detection
标题: 统一音频事件检测
作者:Yidi Jiang,Ruijie Tao,Wen Huang,Qian Chen,Wen Wang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:声音事件检测(SED)检测声音事件的区域,而扬声器日记(SD)分段归因于单个扬声器的语音对话。在SED中,所有说话者段被分类为单个语音事件,而在SD中,非语音声音仅被视为背景噪声。因此,这两个任务在涉及语音对话和非语音声音的复杂音频场景中仅提供部分分析。在本文中,我们介绍了一种新的任务,称为统一的音频事件检测(UAED)的综合音频分析。UAED探索SED和SD任务之间的协同作用,同时检测非语音声音事件和细粒度的语音事件的基础上说话人的身份。为了解决这个问题,我们提出了一个基于transformer的UAED(T-UAED)框架,并构建了来自Librispeech数据集和DESED soundbank的UAED数据。实验表明,该框架有效地利用了任务的相互作用,并大大优于基线,简单地结合了SED和SD模型的输出。T-UAED还通过对DESED和CALLHOME数据集上的单个SED和SD任务执行专用模型的转换来显示其多功能性。摘要:Sound Event Detection (SED) detects regions of sound events, while Speaker Diarization (SD) segments speech conversations attributed to individual speakers. In SED, all speaker segments are classified as a single speech event, while in SD, non-speech sounds are treated merely as background noise. Thus, both tasks provide only partial analysis in complex audio scenarios involving both speech conversation and non-speech sounds. In this paper, we introduce a novel task called Unified Audio Event Detection (UAED) for comprehensive audio analysis. UAED explores the synergy between SED and SD tasks, simultaneously detecting non-speech sound events and fine-grained speech events based on speaker identities. To tackle this task, we propose a Transformer-based UAED (T-UAED) framework and construct the UAED Data derived from the Librispeech dataset and DESED soundbank. Experiments demonstrate that the proposed framework effectively exploits task interactions and substantially outperforms the baseline that simply combines the outputs of SED and SD models. T-UAED also shows its versatility by performing comparably to specialized models for individual SED and SD tasks on DESED and CALLHOME datasets.
【26】 SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
标题: SoloAudio:使用面向数字的音频扩散Transformer的目标声音提取
作者:Helin Wang,Jiarui Hai,Yen-Ju Lu,Karan Thakkar,Mounya Elhilali,Najim Dehak
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在本文中,我们介绍了SoloAudio,一种新的基于扩散的生成模型的目标声音提取(TSE)。我们的方法在音频上训练潜在扩散模型,用对潜在特征进行操作的跳过连接的Transformer取代之前的U-Net主干。SoloAudio通过使用CLAP模型作为目标声音的特征提取器,支持面向音频和面向语言的TSE。此外,SoloAudio利用最先进的文本到音频模型生成的合成音频进行训练,对域外数据和看不见的声音事件表现出强大的泛化能力。我们在FSD Kaggle 2018混合数据集和AudioSet的真实数据上评估了这种方法,其中SoloAudio在域内和域外数据上都实现了最先进的结果,并表现出令人印象深刻的zero-shot和Few-Shot功能。源代码和演示已发布。摘要:In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer that operates on latent features. SoloAudio supports both audio-oriented and language-oriented TSE by utilizing a CLAP model as the feature extractor for target sounds. Furthermore, SoloAudio leverages synthetic audio generated by state-of-the-art text-to-audio models for training, demonstrating strong generalization to out-of-domain data and unseen sound events. We evaluate this approach on the FSD Kaggle 2018 mixture dataset and real data from AudioSet, where SoloAudio achieves the state-of-the-art results on both in-domain and out-of-domain data, and exhibits impressive zero-shot and few-shot capabilities. Source code and demos are released.
【27】 OpenACE: An Open Benchmark for Evaluating Audio Coding Performance
标题: OpenACE:评估音频编码性能的开放基准
作者:Jozef Coldenhoff,Niclas Granqvist,Milos Cernak
链接:点击下载PDF文件
摘要:音频和语音编码缺乏统一的评估和开源测试。许多候选系统都是在专有的、不可再现的或小数据上进行评估的,基于机器学习的编解码器通常在具有类似分布的数据集上进行测试,这与基于数字信号处理的编解码器相比是不公平的,后者通常可以很好地处理看不见的数据。本文提出了一种全频带音频和语音编码质量基准测试,具有更多可变的内容类型,包括传统的开放测试向量。通过开源Opus、3GPP的EVS和最近ETSI的LC 3以及蓝牙LE音频配置文件中使用的LC 3+来呈现音频编码质量评估的示例用例。此外,情感语音编码在16 kbps的质量变化。拟议的开源基准有助于音频和语音编码民主化,可在https: github.com JozefColdenhoff OpenACE上获得。摘要:Audio and speech coding lack unified evaluation and open-source testing. Many candidate systems were evaluated on proprietary, non-reproducible, or small data, and machine learning-based codecs are often tested on datasets with similar distributions as trained on, which is unfairly compared to digital signal processing-based codecs that usually work well with unseen data. This paper presents a full-band audio and speech coding quality benchmark with more variable content types, including traditional open test vectors. An example use case of audio coding quality assessment is presented with open-source Opus, 3GPP's EVS, and recent ETSI's LC3 with LC3+ used in Bluetooth LE Audio profiles. Besides, quality variations of emotional speech encoding at 16 kbps are shown. The proposed open-source benchmark contributes to audio and speech coding democratization and is available at https: github.com JozefColdenhoff OpenACE.
【28】 Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
标题: 量化和减少跨语言言语反欺骗中的语言不匹配效应
作者:Tianchi Liu,Ivan Kukanov,Zihan Pan,Qiongqiong Wang,Hardik B. Sailor,Kong Aik Lee
备注:Accepted to the IEEE Spoken Language Technology Workshop (SLT) 2024. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:语言不匹配的影响影响语音反欺骗系统,而这些影响的调查和量化仍然有限。现有的反欺骗数据集主要是英语的,获取多语言数据集的高成本阻碍了训练语言无关的模型。我们通过评估在英语数据上训练但在其他语言上测试的性能最佳的语音反欺骗系统来启动这项工作,观察到显着的性能下降。我们提出了一种创新的方法-通过TTS(ACCENT)进行基于口音的数据扩展,该方法将不同的语言知识引入到单语训练的模型中,提高了它们的跨语言能力。我们在一个由超过300万个样本组成的大规模数据集上进行实验,其中包括180万个训练样本和近120万个测试样本,涉及12种语言。语言失配的影响进行了初步量化,并显着减少了15%以上,通过应用所提出的ACCENT。这种易于实现的方法显示出对多语言和低资源语言场景的承诺。摘要:The effects of language mismatch impact speech anti-spoofing systems, while investigations and quantification of these effects remain limited. Existing anti-spoofing datasets are mainly in English, and the high cost of acquiring multilingual datasets hinders training language-independent models. We initiate this work by evaluating top-performing speech anti-spoofing systems that are trained on English data but tested on other languages, observing notable performance declines. We propose an innovative approach - Accent-based data expansion via TTS (ACCENT), which introduces diverse linguistic knowledge to monolingual-trained models, improving their cross-lingual capabilities. We conduct experiments on a large-scale dataset consisting of over 3 million samples, including 1.8 million training samples and nearly 1.2 million testing samples across 12 languages. The language mismatch effects are preliminarily quantified and remarkably reduced over 15% by applying the proposed ACCENT. This easily implementable method shows promise for multilingual and low-resource language scenarios.
【29】 Detection of Electric Motor Damage Through Analysis of Sound Signals Using Bayesian Neural Networks
标题: 利用Bayesian神经网络分析声音信号检测电动机损坏
作者:Waldemar Bauer,Marta Zagorowska,Jerzy Baranowski
备注:Accepted to IECON 2024
链接:点击下载PDF文件
摘要:故障监测和诊断对于确保电动机的可靠性非常重要。用于故障检测的有效算法提高了可靠性,但是用于设备诊断的具有成本效益和可靠的分类器的开发是具有挑战性的,特别是由于不存在具有来自正常运行的设备和来自故障设备的信号的良好平衡的数据集。因此,我们建议使用贝叶斯神经网络来检测和分类电动机故障,其功效与不平衡的训练数据。所提出的网络的性能表现在现实生活中的信号,并提供了所提出的解决方案的鲁棒性分析。摘要:Fault monitoring and diagnostics are important to ensure reliability of electric motors. Efficient algorithms for fault detection improve reliability, yet development of cost-effective and reliable classifiers for diagnostics of equipment is challenging, in particular due to unavailability of well-balanced datasets, with signals from properly functioning equipment and those from faulty equipment. Thus, we propose to use a Bayesian neural network to detect and classify faults in electric motors, given its efficacy with imbalanced training data. The performance of the proposed network is demonstrated on real life signals, and a robustness analysis of the proposed solution is provided.
eess.AS音频处理
【1】 Why some audio signal short-time Fourier transform coefficients have nonuniform phase distributions标题: 为什么某些音频信号短期傅里叶变换系数具有不均匀的相分布
作者:Stephen D. Voran
Journal-ref:Proceedings of the 2024 IEEE International Conference on Multimedia and Expo, Niagara Falls, Ontario, July 15-19, 2024
链接:点击下载PDF文件
摘要:短时傅立叶变换(STFT)将音频样本的窗口表示为一组复系数。这些被有利地视为幅度和相位,并且相位的总体分布通常被假设为均匀的。我们表明,当音频信号STFT相位分布分析每个频率或幅度范围,它们可以远离均匀。也就是说,均匀相位分布假设掩盖了重要的重要细节。我们解释了不均匀的相位分布的意义,以及如何利用它们,推导出它们的来源,并解释了为什么选择的STFT窗口形状的影响所得到的相位分布的不均匀性。摘要:The short-time Fourier transform (STFT) represents a window of audio samples as a set of complex coefficients. These are advantageously viewed as magnitudes and phases and the overall distribution of phases is very often assumed to be uniform. We show that when audio signal STFT phase distributions are analyzed per-frequency or per-magnitude range, they can be far from uniform. That is, the uniform phase distribution assumption obscures significant important details. We explain the significance of the nonuniform phase distributions and how they might be exploited, derive their source, and explain why the choice of the STFT window shape influences the nonuniformity of the resulting phase distributions.
【2】 HLTCOE JHU Submission to the Voice Privacy Challenge 2024
标题: HLTCOE JHU提交2024年语音隐私挑战赛
作者:Henry Li Xinyuan,Zexin Cai,Ashi Garg,Kevin Duh,Leibny Paola García-Perera,Sanjeev Khudanpur,Nicholas Andrews,Matthew Wiesner
备注:Submission to the Voice Privacy Challenge 2024. Accepted and presented at
链接:点击下载PDF文件
摘要:我们提出了一些系统的语音隐私的挑战,包括语音转换为基础的系统,如kNN-VC方法和WavLM语音转换方法,和文本到语音(TTS)为基础的系统,包括耳语-VITS。我们发现,虽然语音转换系统更好地保留了情感内容,但它们很难在半白盒攻击场景中隐藏说话者身份;相反,TTS方法在匿名化方面表现更好,在情感保留方面表现更差。最后,我们提出了一个随机混合物系统,旨在平衡这两类系统的优点和缺点,实现了超过40%的强大的EER,同时保持UAR在一个体面的47%。摘要:We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-to-speech (TTS) based systems including Whisper-VITS. We found that while voice conversion systems better preserve emotional content, they struggle to conceal speaker identity in semi-white-box attack scenarios; conversely, TTS methods perform better at anonymization and worse at emotion preservation. Finally, we propose a random admixture system which seeks to balance out the strengths and weaknesses of the two category of systems, achieving a strong EER of over 40% while maintaining UAR at a respectable 47%.
【3】 Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
标题: 具有模拟对话的数据高效儿童-成人说话者拨号
作者:Anfeng Xu,Tiantian Feng,Helen Tager-Flusberg,Catherine Lord,Shrikanth Narayanan
备注:Under review
链接:点击下载PDF文件
摘要:自动化儿童语音分析对于神经认知评估等应用至关重要。说话人日记是自动分析的一个重要组成部分,它可以识别“谁在什么时候说话”。然而,由于隐私问题和缺乏注释数据集,公开可用的儿童-成人说话人日记解决方案是稀缺的,而手动注释每个场景的数据既耗时又昂贵。为了克服这些挑战,我们提出了一个数据高效的解决方案,通过使用AudioSet创建模拟的儿童-成人对话。然后,我们训练Whisper编码器为基础的模型,实现强大的zero-shot性能儿童成人扬声器日记使用真实的数据集。当仅用30分钟的真实训练数据进行微调时,模型性能大幅提高,LoRA进一步提高了迁移学习性能。源代码和在模拟对话上训练的儿童-成人说话者日记模型是公开的。摘要:Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies who spoke when'', is an essential component of the automated analysis. However, publicly available child-adult speaker diarization solutions are scarce due to privacy concerns and a lack of annotated datasets, while manually annotating data for each scenario is both time-consuming and costly. To overcome these challenges, we propose a data-efficient solution by creating simulated child-adult conversations using AudioSet. We then train a Whisper Encoder-based model, achieving strong zero-shot performance on child-adult speaker diarization using real datasets. The model performance improves substantially when fine-tuned with only 30 minutes of real train data, with LoRA further improving the transfer learning performance. The source code and the child-adult speaker diarization model trained on simulated conversations are publicly available.
【4】 LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment
标题: LLaQo:在表达性音乐表现评估中建立基于查询的教练
作者:Huan Zhang,Vincent Cheung,Hayato Nishioka,Simon Dixon,Shinichi Furuya
链接:点击下载PDF文件
摘要:音乐理解的研究已经通过高级表示广泛地探索了作曲级别的属性,例如基调,流派和乐器,从而导致使用大型语言模型的跨模态应用程序。然而,音乐表演的风格表达和技巧等方面仍然没有得到充分的探索,同时使用大型语言模型来提高定制反馈的教育成果的潜力也没有得到充分的探索。为了弥合这一差距,我们引入了LLaQo,这是一个基于大型语言查询的音乐教练,它利用音频语言建模来提供对音乐表演的详细和形成性评估。我们还介绍了练习调整的查询-响应数据集,这些数据集涵盖了从音高准确性到清晰度的各种性能维度,以及上下文性能理解(如难度和性能技术)。利用AudioMAE编码器和Vicuna-7 b LLM后端,我们的模型在预测教师的表现评级以及识别作品难度和演奏技巧方面取得了最先进的(SOTA)结果。此外,在使用音频文本匹配的用户研究中,与其他基线模型相比,LLaQo的文本响应被评为显着更高。因此,我们提出的模型可以提供翔实的答案,开放式的问题,从音频数据的音乐表演。摘要:Research in music understanding has extensively explored composition-level attributes such as key, genre, and instrumentation through advanced representations, leading to cross-modal applications using large language models. However, aspects of musical performance such as stylistic expression and technique remain underexplored, along with the potential of using large language models to enhance educational outcomes with customized feedback. To bridge this gap, we introduce LLaQo, a Large Language Query-based music coach that leverages audio language modeling to provide detailed and formative assessments of music performances. We also introduce instruction-tuned query-response datasets that cover a variety of performance dimensions from pitch accuracy to articulation, as well as contextual performance understanding (such as difficulty and performance techniques). Utilizing AudioMAE encoder and Vicuna-7b LLM backend, our model achieved state-of-the-art (SOTA) results in predicting teachers' performance ratings, as well as in identifying piece difficulty and playing techniques. Textual responses from LLaQo was moreover rated significantly higher compared to other baseline models in a user study using audio-text matching. Our proposed model can thus provide informative answers to open-ended questions related to musical performance from audio data.
【5】 FLAMO: An Open-Source Library for Frequency-Domain Differentiable Audio Processing
标题: FLAMO:一个用于频域可区分音频处理的开源库
作者:Gloria Dal Santo,Gian Marco De Bortoli,Karolina Prawda,Sebastian J. Schlecht,Vesa Välimäki
链接:点击下载PDF文件
摘要:我们介绍了FLAMO,这是一个用于音频模块优化的频率采样库,旨在实现和优化可微线性时不变音频系统。该库是开源的,基于频率采样滤波器设计方法构建,允许创建可独立使用或在神经网络计算图中使用的可微分模块,从而简化可微分音频系统的开发。它包括预定义的过滤模块和辅助类,用于构建,训练和记录优化的系统,所有这些都可以通过直观的界面访问。这些模块的实际应用证明通过两个案例研究:优化的人工混响器和主动声学系统,以提高响应的平滑度。摘要:We present FLAMO, a Frequency-sampling Library for Audio-Module Optimization designed to implement and optimize differentiable linear time-invariant audio systems. The library is open-source and built on the frequency-sampling filter design method, allowing for the creation of differentiable modules that can be used stand-alone or within the computation graph of neural networks, simplifying the development of differentiable audio systems. It includes predefined filtering modules and auxiliary classes for constructing, training, and logging the optimized systems, all accessible through an intuitive interface. Practical application of these modules is demonstrated through two case studies: the optimization of an artificial reverberator and an active acoustics system for improved response smoothness.
【6】 Text-To-Speech Synthesis In The Wild
标题: 野外的文本到语音合成
作者:Jee-weon Jung,Wangyou Zhang,Soumi Maiti,Yihan Wu,Xin Wang,Ji-Hoon Kim,Yuta Matsunaga,Seyun Um,Jinchuan Tian,Hye-jin Shim,Nicholas Evans,Joon Son Chung,Shinnosuke Takamichi,Shinji Watanabe
备注:5 pages, submitted to ICASSP 2025 as a conference paper
链接:点击下载PDF文件
摘要:传统上,文本到语音(TTS)系统是使用在诸如消声室之类的良性声学环境中收集的录音室质量、提示或读取语音的适度数据库来训练的。然而,最近的文献显示,努力训练TTS系统使用在野外收集的数据。虽然这种方法允许使用大量的自然语音,但到目前为止,还没有通用的数据集。我们介绍了TTS In the Wild(TITW)数据集,这是一个完全自动化的管道的结果,在这种情况下,应用于通常用于说话人识别的VoxCeleb 1数据集。我们还提出了两个训练集。TITW-Hard来源于VoxCeleb 1源数据的转录、分割和选择。TITW-Easy是在DNSMOS的基础上增加了增强功能和数据选择功能。我们发现,最近的一些TTS模型可以使用TITW-Easy成功训练,但使用TITW-Hard产生类似的结果仍然极具挑战性。数据集和协议都是公开的,并支持使用TITW数据训练的TTS系统的基准测试。摘要:Text-to-speech (TTS) systems are traditionally trained using modest databases of studio-quality, prompted or read speech collected in benign acoustic environments such as anechoic rooms. The recent literature nonetheless shows efforts to train TTS systems using data collected in the wild. While this approach allows for the use of massive quantities of natural speech, until now, there are no common datasets. We introduce the TTS In the Wild (TITW) dataset, the result of a fully automated pipeline, in this case, applied to the VoxCeleb1 dataset commonly used for speaker recognition. We further propose two training sets. TITW-Hard is derived from the transcription, segmentation, and selection of VoxCeleb1 source data. TITW-Easy is derived from the additional application of enhancement and additional data selection based on DNSMOS. We show that a number of recent TTS models can be trained successfully using TITW-Easy, but that it remains extremely challenging to produce similar results using TITW-Hard. Both the dataset and protocols are publicly available and support the benchmarking of TTS systems trained using TITW data.
【7】 Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
标题: 使用耳脑电解码多说话者环境中的听觉注意力
作者:Haolin Zhu,Yujie Yan,Xiran Xu,Zhongshu Ge,Pei Tian,Xihong Wu,Jing Chen
链接:点击下载PDF文件
摘要:听觉注意解码(AAD)可以通过分析和处理脑电图(EEG)数据来帮助确定听觉选择性注意任务期间参与说话者的身份。大多数关于AAD的研究都是基于双说话人场景下的头皮脑电信号,与实际应用相距甚远。耳脑电图最近获得了显着的关注,由于其运动的宽容和不可见性在数据采集过程中,使其易于与其他设备的应用程序。在这项工作中,参与者有选择地参加了四个空间分离的扬声器的语音在消声室。EEG数据同时从头皮EEG系统和耳EEG系统(cEEGrids)收集。使用耳EEG数据的时间响应函数(TRFs)和刺激重建(SR)。结果表明,在60年代(机会水平为25%),有注意语音的TRFs比无注意语音的TRFs强,解码正确率为41.3%。为了进一步研究电极放置和数量的影响,在头皮EEG和耳EEG中使用SR,揭示了虽然电极数量具有较小的影响,但它们的位置对解码准确性具有显著影响。在此耳脑电数据库上验证了一种听觉空间注意检测方法STAnet,在1秒解码窗口内的正确率为93.1%。我们工作的实现代码和数据库可在GitHub:https: github.com zhl486 Ear_EEG_code.git和Zenodo:https: zenodo.org records 10803261上找到。摘要:Auditory Attention Decoding (AAD) can help to determine the identity of the attended speaker during an auditory selective attention task, by analyzing and processing measurements of electroencephalography (EEG) data. Most studies on AAD are based on scalp-EEG signals in two-speaker scenarios, which are far from real application. Ear-EEG has recently gained significant attention due to its motion tolerance and invisibility during data acquisition, making it easy to incorporate with other devices for applications. In this work, participants selectively attended to one of the four spatially separated speakers' speech in an anechoic room. The EEG data were concurrently collected from a scalp-EEG system and an ear-EEG system (cEEGrids). Temporal response functions (TRFs) and stimulus reconstruction (SR) were utilized using ear-EEG data. Results showed that the attended speech TRFs were stronger than each unattended speech and decoding accuracy was 41.3 % in the 60s (chance level of 25 %). To further investigate the impact of electrode placement and quantity, SR was utilized in both scalp-EEG and ear-EEG, revealing that while the number of electrodes had a minor effect, their positioning had a significant influence on the decoding accuracy. One kind of auditory spatial attention detection (ASAD) method, STAnet, was testified with this ear-EEG database, resulting in 93.1% in 1-second decoding window. The implementation code and database for our work are available on GitHub: https: github.com zhl486 Ear_EEG_code.git and Zenodo: https: zenodo.org records 10803261.
【8】 DM: Dual-path Magnitude Network for General Speech Restoration
标题: DM:用于通用语音恢复的双路径幅度网络
作者:Da-Hee Yang,Dail Kim,Joon-Hyuk Chang,Jeonghwan Choi,Han-gil Moon
链接:点击下载PDF文件
摘要:在本文中,我们介绍了一种新的通用语音恢复模型:双路径幅度(DM)网络,旨在解决多种失真,包括噪声,混响,和带宽退化有效。DM网络采用共享参数的双并行幅度解码器:一个使用基于掩蔽的失真消除算法,另一个采用基于映射的语音恢复方法。DM网络的一个新颖方面是通过跳过连接将掩蔽解码器输出的幅度谱图集成到映射解码器中,从而增强了整体恢复能力。这种综合方法克服了以前模型中观察到的固有局限性,如逐步分析中所详述。实验结果表明,DM网络在语音恢复的综合性能上优于其他基线模型,能够以较少的参数实现实质性的语音恢复。摘要:In this paper, we introduce a novel general speech restoration model: the Dual-path Magnitude (DM) network, designed to address multiple distortions including noise, reverberation, and bandwidth degradation effectively. The DM network employs dual parallel magnitude decoders that share parameters: one uses a masking-based algorithm for distortion removal and the other employs a mapping-based approach for speech restoration. A novel aspect of the DM network is the integration of the magnitude spectrogram output from the masking decoder into the mapping decoder through a skip connection, enhancing the overall restoration capability. This integrated approach overcomes the inherent limitations observed in previous models, as detailed in a step-by-step analysis. The experimental results demonstrate that the DM network outperforms other baseline models in the comprehensive aspect of general speech restoration, achieving substantial restoration with fewer parameters.
【9】 NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training
标题: NEST-PQ:语音自我监督预训练的下一个令牌预测
作者:Minglun Han,Ye Bai,Chen Shen,Youjia Huang,Mingkun Huang,Zehua Lin,Linhao Dong,Lu Lu,Yuxuan Wang
备注:5 pages, 2 figures, Work in progress
链接:点击下载PDF文件
摘要:语音自监督预训练可以有效提高下游任务的性能。然而,以前的语音自监督学习(SSL)方法,如HuBERT和BEST-RQ,专注于利用具有双向上下文的非因果编码器,并且缺乏对下游流模型的足够支持。为了解决这个问题,我们介绍了下一个令牌预测的语音预训练方法与随机投影量化器(NEST-RQ)。NEST-RQ采用因果编码器,只有左上下文,并使用下一个令牌预测(NTP)作为训练任务。在大规模数据集上,与BEST-RQ相比,所提出的NEST-RQ在非流式自动语音识别(ASR)上实现了相当的性能,并且在流式ASR上实现了更好的性能。我们还进行了分析实验方面的未来上下文大小的流ASR,SSL的码本质量和编码器的模型大小。总之,本文论证了NTP在语音SSL中的可行性,为语音SSL的研究提供了经验证据和见解。摘要:Speech self-supervised pre-training can effectively improve the performance of downstream tasks. However, previous self-supervised learning (SSL) methods for speech, such as HuBERT and BEST-RQ, focus on utilizing non-causal encoders with bidirectional context, and lack sufficient support for downstream streaming models. To address this issue, we introduce the next token prediction based speech pre-training method with random-projection quantizer (NEST-RQ). NEST-RQ employs causal encoders with only left context and uses next token prediction (NTP) as the training task. On the large-scale dataset, compared to BEST-RQ, the proposed NEST-RQ achieves comparable performance on non-streaming automatic speech recognition (ASR) and better performance on streaming ASR. We also conduct analytical experiments in terms of the future context size of streaming ASR, the codebook quality of SSL and the model size of the encoder. In summary, the paper demonstrates the feasibility of the NTP in speech SSL and provides empirical evidence and insights for speech SSL research.
【10】 DualSep: A Light-weight dual-encoder convolutional recurrent network for real-time in-car speech separation
标题: DualSep:用于实时车内语音分离的轻量级双编码器卷积循环网络
作者:Ziqian Wang,Jiayao Sun,Zihan Zhang,Xingchen Li,Jie Liu,Lei Xie
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
摘要:深度学习和语音激活技术的进步推动了人车交互的发展。分布式麦克风阵列广泛应用于车内场景,因为它可以准确地捕获来自不同语音区域的乘客的声音。然而,音频通道数量的增加,加上有限的计算资源和车载系统的低延迟要求,为车载多通道语音分离提出了挑战。为了迁移的问题,我们提出了一个轻量级的框架,级联数字信号处理(DSP)和神经网络(NN)。我们利用固定波束形成(BF),以减少计算成本和独立向量分析(IVA)提供空间先验。我们采用双编码器进行双分支建模,空间编码器捕获空间线索,光谱编码器保留光谱信息,促进空间-光谱融合。我们提出的系统支持流和非流模式。实验结果表明,该系统在各种指标的优越性。在Intel Core i7(2.6GHz)CPU上仅具有0.83M参数和0.39实时因子(RTF),它有效地将语音分离到不同的语音区域。我们的演示可在https: honee-w.github.io DualSep 上获得。摘要:Advancements in deep learning and voice-activated technologies have driven the development of human-vehicle interaction. Distributed microphone arrays are widely used in in-car scenarios because they can accurately capture the voices of passengers from different speech zones. However, the increase in the number of audio channels, coupled with the limited computational resources and low latency requirements of in-car systems, presents challenges for in-car multi-channel speech separation. To migrate the problems, we propose a lightweight framework that cascades digital signal processing (DSP) and neural networks (NN). We utilize fixed beamforming (BF) to reduce computational costs and independent vector analysis (IVA) to provide spatial prior. We employ dual encoders for dual-branch modeling, with spatial encoder capturing spatial cues and spectral encoder preserving spectral information, facilitating spatial-spectral fusion. Our proposed system supports both streaming and non-streaming modes. Experimental results demonstrate the superiority of the proposed system across various metrics. With only 0.83M parameters and 0.39 real-time factor (RTF) on an Intel Core i7 (2.6GHz) CPU, it effectively separates speech into distinct speech zones. Our demos are available at https: honee-w.github.io DualSep .
【11】 Effective Integration of KAN for Keyword Spotting
标题: 有效集成KAN以识别关键词
作者:Anfeng Xu,Biqiao Zhang,Shuyu Kong,Yiteng Huang,Zhaojun Yang,Sangeeta Srivastava,Ming Sun
备注:Under review
链接:点击下载PDF文件
摘要:关键词定位(KWS)是具有语音辅助功能的智能设备的重要语音处理组件。本文研究了Kolmogorov-Arnold网络(KAN)是否可以用于增强KWS的性能。我们探索各种方法来集成KAN的模型架构的基础上一维卷积神经网络(CNN)。我们发现KAN可以有效地在低维空间中建模高级特征,从而在适当集成时提高KWS性能。研究结果揭示了理解KAN的语音处理任务和其他模式,为未来的研究人员。摘要:Keyword spotting (KWS) is an important speech processing component for smart devices with voice assistance capability. In this paper, we investigate if Kolmogorov-Arnold Networks (KAN) can be used to enhance the performance of KWS. We explore various approaches to integrate KAN for a model architecture based on 1D Convolutional Neural Networks (CNN). We find that KAN is effective at modeling high-level features in lower-dimensional spaces, resulting in improved KWS performance when integrated appropriately. The findings shed light on understanding KAN for speech processing tasks and on other modalities for future researchers.
【12】 Frequency Tracking Features for Data-Efficient Deep Siren Identification
标题: 频率跟踪功能,实现数据高效的深度警报识别
作者:Stefano Damiano,Thomas Dietzen,Toon van Waterschoot
备注:Accepted paper: Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2024)
链接:点击下载PDF文件
摘要:识别城市音景中的警笛声音是智能车辆的一个重要安全方面,并已通过神经网络得到广泛解决,这些神经网络确保了对警笛信号多样性和交通特征的强非结构化背景噪声的鲁棒性。卷积神经网络分析输入信号的频谱图特征时,达到最先进的性能,足够的训练数据捕获的目标声学场景的多样性是可用的。在实践中,数据通常是有限的,算法应该是鲁棒的,以适应看不见的声学条件,而不需要大量的数据集进行重新训练。在这项工作中,给出了警笛信号的谐波性质,其特征在于由一个周期性演变的基频,我们提出了一种低复杂度的特征提取方法的基础上使用单参数自适应陷波滤波器的频率跟踪。然后,这些特征被用来设计一个小规模的卷积网络,适合用有限的数据进行训练。实验结果表明,在训练数据有限的情况下,该模型的性能始终优于传统的基于谱图的模型,具有更好的跨域泛化能力和更小的规模。摘要:The identification of siren sounds in urban soundscapes is a crucial safety aspect for smart vehicles and has been widely addressed by means of neural networks that ensure robustness to both the diversity of siren signals and the strong and unstructured background noise characterizing traffic. Convolutional neural networks analyzing spectrogram features of incoming signals achieve state-of-the-art performance when enough training data capturing the diversity of the target acoustic scenes is available. In practice, data is usually limited and algorithms should be robust to adapt to unseen acoustic conditions without requiring extensive datasets for re-training. In this work, given the harmonic nature of siren signals, characterized by a periodically evolving fundamental frequency, we propose a low-complexity feature extraction method based on frequency tracking using a single-parameter adaptive notch filter. The features are then used to design a small-scale convolutional network suitable for training with limited data. The evaluation results indicate that the proposed model consistently outperforms the traditional spectrogram-based model when limited training data is available, achieves better cross-domain generalization and has a smaller size.
【13】 Unified Audio Event Detection
标题: 统一音频事件检测
作者:Yidi Jiang,Ruijie Tao,Wen Huang,Qian Chen,Wen Wang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:声音事件检测(SED)检测声音事件的区域,而扬声器日记(SD)分段归因于单个扬声器的语音对话。在SED中,所有说话者段被分类为单个语音事件,而在SD中,非语音声音仅被视为背景噪声。因此,这两个任务在涉及语音对话和非语音声音的复杂音频场景中仅提供部分分析。在本文中,我们介绍了一种新的任务,称为统一的音频事件检测(UAED)的综合音频分析。UAED探索SED和SD任务之间的协同作用,同时检测非语音声音事件和细粒度的语音事件的基础上说话人的身份。为了解决这个问题,我们提出了一个基于transformer的UAED(T-UAED)框架,并构建了来自Librispeech数据集和DESED soundbank的UAED数据。实验表明,该框架有效地利用了任务的相互作用,并大大优于基线,简单地结合了SED和SD模型的输出。T-UAED还通过对DESED和CALLHOME数据集上的单个SED和SD任务执行专用模型的转换来显示其多功能性。摘要:Sound Event Detection (SED) detects regions of sound events, while Speaker Diarization (SD) segments speech conversations attributed to individual speakers. In SED, all speaker segments are classified as a single speech event, while in SD, non-speech sounds are treated merely as background noise. Thus, both tasks provide only partial analysis in complex audio scenarios involving both speech conversation and non-speech sounds. In this paper, we introduce a novel task called Unified Audio Event Detection (UAED) for comprehensive audio analysis. UAED explores the synergy between SED and SD tasks, simultaneously detecting non-speech sound events and fine-grained speech events based on speaker identities. To tackle this task, we propose a Transformer-based UAED (T-UAED) framework and construct the UAED Data derived from the Librispeech dataset and DESED soundbank. Experiments demonstrate that the proposed framework effectively exploits task interactions and substantially outperforms the baseline that simply combines the outputs of SED and SD models. T-UAED also shows its versatility by performing comparably to specialized models for individual SED and SD tasks on DESED and CALLHOME datasets.
【14】 SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
标题: SoloAudio:使用面向数字的音频扩散Transformer的目标声音提取
作者:Helin Wang,Jiarui Hai,Yen-Ju Lu,Karan Thakkar,Mounya Elhilali,Najim Dehak
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在本文中,我们介绍了SoloAudio,一种新的基于扩散的生成模型的目标声音提取(TSE)。我们的方法在音频上训练潜在扩散模型,用对潜在特征进行操作的跳过连接的Transformer取代之前的U-Net主干。SoloAudio通过使用CLAP模型作为目标声音的特征提取器,支持面向音频和面向语言的TSE。此外,SoloAudio利用最先进的文本到音频模型生成的合成音频进行训练,对域外数据和看不见的声音事件表现出强大的泛化能力。我们在FSD Kaggle 2018混合数据集和AudioSet的真实数据上评估了这种方法,其中SoloAudio在域内和域外数据上都实现了最先进的结果,并表现出令人印象深刻的zero-shot和Few-Shot功能。源代码和演示已发布。摘要:In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer that operates on latent features. SoloAudio supports both audio-oriented and language-oriented TSE by utilizing a CLAP model as the feature extractor for target sounds. Furthermore, SoloAudio leverages synthetic audio generated by state-of-the-art text-to-audio models for training, demonstrating strong generalization to out-of-domain data and unseen sound events. We evaluate this approach on the FSD Kaggle 2018 mixture dataset and real data from AudioSet, where SoloAudio achieves the state-of-the-art results on both in-domain and out-of-domain data, and exhibits impressive zero-shot and few-shot capabilities. Source code and demos are released.
【15】 OpenACE: An Open Benchmark for Evaluating Audio Coding Performance
标题: OpenACE:评估音频编码性能的开放基准
作者:Jozef Coldenhoff,Niclas Granqvist,Milos Cernak
链接:点击下载PDF文件
摘要:音频和语音编码缺乏统一的评估和开源测试。许多候选系统都是在专有的、不可再现的或小数据上进行评估的,基于机器学习的编解码器通常在具有类似分布的数据集上进行测试,这与基于数字信号处理的编解码器相比是不公平的,后者通常可以很好地处理看不见的数据。本文提出了一种全频带音频和语音编码质量基准测试,具有更多可变的内容类型,包括传统的开放测试向量。通过开源Opus、3GPP的EVS和最近ETSI的LC 3以及蓝牙LE音频配置文件中使用的LC 3+来呈现音频编码质量评估的示例用例。此外,情感语音编码在16 kbps的质量变化。拟议的开源基准有助于音频和语音编码民主化,可在https: github.com JozefColdenhoff OpenACE上获得。摘要:Audio and speech coding lack unified evaluation and open-source testing. Many candidate systems were evaluated on proprietary, non-reproducible, or small data, and machine learning-based codecs are often tested on datasets with similar distributions as trained on, which is unfairly compared to digital signal processing-based codecs that usually work well with unseen data. This paper presents a full-band audio and speech coding quality benchmark with more variable content types, including traditional open test vectors. An example use case of audio coding quality assessment is presented with open-source Opus, 3GPP's EVS, and recent ETSI's LC3 with LC3+ used in Bluetooth LE Audio profiles. Besides, quality variations of emotional speech encoding at 16 kbps are shown. The proposed open-source benchmark contributes to audio and speech coding democratization and is available at https: github.com JozefColdenhoff OpenACE.
【16】 Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
标题: 量化和减少跨语言言语反欺骗中的语言不匹配效应
作者:Tianchi Liu,Ivan Kukanov,Zihan Pan,Qiongqiong Wang,Hardik B. Sailor,Kong Aik Lee
备注:Accepted to the IEEE Spoken Language Technology Workshop (SLT) 2024. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:语言不匹配的影响影响语音反欺骗系统,而这些影响的调查和量化仍然有限。现有的反欺骗数据集主要是英语的,获取多语言数据集的高成本阻碍了训练语言无关的模型。我们通过评估在英语数据上训练但在其他语言上测试的性能最佳的语音反欺骗系统来启动这项工作,观察到显着的性能下降。我们提出了一种创新的方法-通过TTS(ACCENT)进行基于口音的数据扩展,该方法将不同的语言知识引入到单语训练的模型中,提高了它们的跨语言能力。我们在一个由超过300万个样本组成的大规模数据集上进行实验,其中包括180万个训练样本和近120万个测试样本,涉及12种语言。语言失配的影响进行了初步量化,并显着减少了15%以上,通过应用所提出的ACCENT。这种易于实现的方法显示出对多语言和低资源语言场景的承诺。摘要:The effects of language mismatch impact speech anti-spoofing systems, while investigations and quantification of these effects remain limited. Existing anti-spoofing datasets are mainly in English, and the high cost of acquiring multilingual datasets hinders training language-independent models. We initiate this work by evaluating top-performing speech anti-spoofing systems that are trained on English data but tested on other languages, observing notable performance declines. We propose an innovative approach - Accent-based data expansion via TTS (ACCENT), which introduces diverse linguistic knowledge to monolingual-trained models, improving their cross-lingual capabilities. We conduct experiments on a large-scale dataset consisting of over 3 million samples, including 1.8 million training samples and nearly 1.2 million testing samples across 12 languages. The language mismatch effects are preliminarily quantified and remarkably reduced over 15% by applying the proposed ACCENT. This easily implementable method shows promise for multilingual and low-resource language scenarios.
【17】 Detection of Electric Motor Damage Through Analysis of Sound Signals Using Bayesian Neural Networks
标题: 利用Bayesian神经网络分析声音信号检测电动机损坏
作者:Waldemar Bauer,Marta Zagorowska,Jerzy Baranowski
备注:Accepted to IECON 2024
链接:点击下载PDF文件
摘要:故障监测和诊断对于确保电动机的可靠性非常重要。用于故障检测的有效算法提高了可靠性,但是用于设备诊断的具有成本效益和可靠的分类器的开发是具有挑战性的,特别是由于不存在具有来自正常运行的设备和来自故障设备的信号的良好平衡的数据集。因此,我们建议使用贝叶斯神经网络来检测和分类电动机故障,其功效与不平衡的训练数据。所提出的网络的性能表现在现实生活中的信号,并提供了所提出的解决方案的鲁棒性分析。摘要:Fault monitoring and diagnostics are important to ensure reliability of electric motors. Efficient algorithms for fault detection improve reliability, yet development of cost-effective and reliable classifiers for diagnostics of equipment is challenging, in particular due to unavailability of well-balanced datasets, with signals from properly functioning equipment and those from faulty equipment. Thus, we propose to use a Bayesian neural network to detect and classify faults in electric motors, given its efficacy with imbalanced training data. The performance of the proposed network is demonstrated on real life signals, and a robustness analysis of the proposed solution is provided.
【18】 Towards Leveraging Contrastively Pretrained Neural Audio Embeddings for Recommender Tasks
标题: 利用对比预训练的神经音频嵌入进行推荐任务
作者:Florian Grötschla,Luca Strässle,Luca A. Lanzendörfer,Roger Wattenhofer
备注:Accepted at the 2nd Music Recommender Workshop (@RecSys)
链接:点击下载PDF文件
摘要:音乐推荐系统经常利用基于网络的模型来捕获音乐作品、艺术家和用户之间的关系。虽然这些关系为预测提供了有价值的见解,但由于初始信息不足,新的音乐作品或艺术家经常面临冷启动问题。为了解决这个问题,人们可以直接从音乐中提取基于内容的信息,以增强基于协作过滤的方法。虽然以前的方法依赖于手工制作的音频功能,但我们探索了对比预训练的神经音频嵌入模型的使用,它提供了更丰富,更细致的音乐表示。我们的实验表明,神经嵌入,特别是那些使用对比音频预训练(CLAP)模型生成的神经嵌入,为在基于图的框架内增强音乐推荐任务提供了一种很有前途的方法。摘要:Music recommender systems frequently utilize network-based models to capture relationships between music pieces, artists, and users. Although these relationships provide valuable insights for predictions, new music pieces or artists often face the cold-start problem due to insufficient initial information. To address this, one can extract content-based information directly from the music to enhance collaborative-filtering-based methods. While previous approaches have relied on hand-crafted audio features for this purpose, we explore the use of contrastively pretrained neural audio embedding models, which offer a richer and more nuanced representation of music. Our experiments demonstrate that neural embeddings, particularly those generated with the Contrastive Language-Audio Pretraining (CLAP) model, present a promising approach to enhancing music recommendation tasks within graph-based frameworks.
【19】 Biomimetic Frontend for Differentiable Audio Processing
标题: 可区分音频处理的仿生前沿
作者:Ruolan Leslie Famularo,Dmitry N. Zotkin,Shihab A. Shamma,Ramani Duraiswami
链接:点击下载PDF文件
摘要:虽然音频和语音处理中的模型变得越来越深入,越来越端到端,但它们因此需要在大数据上进行昂贵的训练,并且通常很脆弱。我们建立在人类听觉的经典模型上,并使其可区分,这样我们就可以将传统的可解释仿生信号处理方法与深度学习框架相结合。这使我们能够得到一个表达性和可解释的模型,该模型可以很容易地在少量数据上进行训练。我们将此模型应用于音频处理任务,包括分类和增强。结果表明,我们的可微模型在计算效率和鲁棒性方面优于黑箱方法,即使训练数据很少。我们还讨论了其他潜在的应用。摘要:While models in audio and speech processing are becoming deeper and more end-to-end, they as a consequence need expensive training on large data, and are often brittle. We build on a classical model of human hearing and make it differentiable, so that we can combine traditional explainable biomimetic signal processing approaches with deep-learning frameworks. This allows us to arrive at an expressive and explainable model that is easily trained on modest amounts of data. We apply this model to audio processing tasks, including classification and enhancement. Results show that our differentiable model surpasses black-box approaches in terms of computational efficiency and robustness, even with little training data. We also discuss other potential applications.
【20】 Clean Label Attacks against SLU Systems
标题: 针对SL U系统的干净标签攻击
作者:Henry Li Xinyuan,Sonal Joshi,Thomas Thebaud,Jesus Villalba,Najim Dehak,Sanjeev Khudanpur
备注:Accepted at IEEE SLT 2024
链接:点击下载PDF文件
摘要:中毒后门攻击涉及攻击者操纵训练数据,通过在推理时在信号中插入触发器来诱导受害者模型中的某些行为。我们在支持 执行口语理解任务的最先进的语音识别模型上采用了干净标签后门(CLBD)-数据中毒攻击,该攻击不修改训练标签,通过中毒10%的训练数据实现了99.8%的攻击成功率。我们分析了毒药的信号强度、中毒样本的百分比以及触发器的选择如何影响攻击。我们还发现,CLBD攻击在应用于对代理模型来说本质上很难的训练样本时是最成功的。使用这种策略,我们通过毒化1.5%的训练数据实现了99.3%的攻击成功率。最后,我们应用了两个以前开发的防御基于梯度的攻击,并发现他们实现混合成功对中毒。摘要:Poisoning backdoor attacks involve an adversary manipulating the training data to induce certain behaviors in the victim model by inserting a trigger in the signal at inference time. We adapted clean label backdoor (CLBD)-data poisoning attacks, which do not modify the training labels, on state-of-the-art speech recognition models that support perform a Spoken Language Understanding task, achieving 99.8% attack success rate by poisoning 10% of the training data. We analyzed how varying the signal-strength of the poison, percent of samples poisoned, and choice of trigger impact the attack. We also found that CLBD attacks are most successful when applied to training samples that are inherently hard for a proxy model. Using this strategy, we achieved an attack success rate of 99.3% by poisoning a meager 1.5% of the training data. Finally, we applied two previously developed defenses against gradient-based attacks, and found that they attain mixed success against poisoning.
【21】 Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages
标题: 探索极低资源语言中数据量对ASB的影响
作者:Yao-Fei Cheng,Li-Wei Chen,Hung-Shin Lee,Hsin-Min Wang
链接:点击下载PDF文件
摘要:本研究探讨了数据增强技术在低资源自动语音识别(ASR)中的有效性,重点是两种濒危的南岛语言,Amis和Seediq。认识到自我监督学习(SSL)在低资源环境中的潜力,我们探索了数据量对SSL模型持续预训练的影响。我们提出了一种新颖的数据选择方案,利用多语言语料库来增加有限的目标语言数据。该方案利用语言分类器来提取话语嵌入,并采用单类分类器来识别语音和音韵上接近目标语言的话语。根据决策得分对语句进行排名和选择,确保在SSL-ASR管道中包含高度相关的数据。我们的实验结果表明,这种方法的有效性,产生了显着的改善,在ASR性能为Amis和Seediq。这些发现强调了通过跨语言迁移学习对低资源语言ASR进行数据增强的可行性和前景。摘要:This study investigates the efficacy of data augmentation techniques for low-resource automatic speech recognition (ASR), focusing on two endangered Austronesian languages, Amis and Seediq. Recognizing the potential of self-supervised learning (SSL) in low-resource settings, we explore the impact of data volume on the continued pre-training of SSL models. We propose a novel data-selection scheme leveraging a multilingual corpus to augment the limited target language data. This scheme utilizes a language classifier to extract utterance embeddings and employs one-class classifiers to identify utterances phonetically and phonologically proximate to the target languages. Utterances are ranked and selected based on their decision scores, ensuring the inclusion of highly relevant data in the SSL-ASR pipeline. Our experimental results demonstrate the effectiveness of this approach, yielding substantial improvements in ASR performance for both Amis and Seediq. These findings underscore the feasibility and promise of data augmentation through cross-lingual transfer learning for low-resource language ASR.
【22】 Exploring SSL Discrete Tokens for Multilingual ASR
标题: 探索多语言ASB的SSL离散令牌
作者:Mingyu Cui,Daxin Tan,Yifan Yang,Dingdong Wang,Huimeng Wang,Xiao Chen,Xie Chen,Xunying Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:随着自监督学习(SSL)在语音相关任务中的发展,人们越来越关注利用SSL生成的离散令牌进行自动语音识别(ASR),因为它们提供了更快的处理技术。然而,以前的研究主要集中在多语言ASR与Fbank功能或英语ASR与离散的令牌,留下了差距,在适应离散令牌的多语言ASR的情况下。这项研究提出了一个全面的比较离散令牌生成的各种领先的SSL模型在多个语言领域。我们的目标是探索性能和效率的语音离散令牌跨多个语言领域的单语和多语言的ASR方案。实验结果表明,离散令牌在七个语言领域的ASR任务中与使用Fbank特征训练的系统取得了相当的结果,平均单词错误率(WER)降低了0.31%和1.76%的绝对值(2.80%和15.70%相对)分别在开发和测试集,特别是WER减少6.82%的绝对(41.48%相对)波兰测试集。摘要:With the advancement of Self-supervised Learning (SSL) in speech-related tasks, there has been growing interest in utilizing discrete tokens generated by SSL for automatic speech recognition (ASR), as they offer faster processing techniques. However, previous studies primarily focused on multilingual ASR with Fbank features or English ASR with discrete tokens, leaving a gap in adapting discrete tokens for multilingual ASR scenarios. This study presents a comprehensive comparison of discrete tokens generated by various leading SSL models across multiple language domains. We aim to explore the performance and efficiency of speech discrete tokens across multiple language domains for both monolingual and multilingual ASR scenarios. Experimental results demonstrate that discrete tokens achieve comparable results against systems trained on Fbank features in ASR tasks across seven language domains with an average word error rate (WER) reduction of 0.31% and 1.76% absolute (2.80% and 15.70% relative) on dev and test sets respectively, with particularly WER reduction of 6.82% absolute (41.48% relative) on the Polish test set.
【23】 Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
标题: 探索基于Zipformer的上下文ASB的SSL离散语音特征
作者:Mingyu Cui,Yifan Yang,Jiajun Deng,Jiawen Kang,Shujie Hu,Tianzi Wang,Zhaoqing Li,Shiliang Zhang,Xie Chen,Xunying Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:基于自监督学习(SSL)的离散语音表示是高度紧凑和领域自适应的。在本文中,SSL离散语音特征提取的WavLM模型被用来作为额外的跨话语声学上下文特征的Zipformer-Transducer ASR系统。取代Fbank功能与离散令牌功能建模跨话语上下文(从前面和未来的段),或当前话语的内部上下文单独,或两者同时,彻底证明Gigaspeech 1000小时语料库的功效。最好的Zipformer-Transducer系统使用基于离散令牌的跨话语上下文特征优于使用话语内部上下文的基线,仅在开发和测试数据上具有0.32%至0.41%的绝对统计显著的单词错误率(WER)降低(2.78%至3.54%相对)。在开发集和测试集上获得了11.15%和11.14%的最低公布WER。我们的工作是开源的,并可在https: github.com open-creator icefall tree master egs gigaspeech Context _ASR上公开获取。摘要:Self-supervised learning (SSL) based discrete speech representations are highly compact and domain adaptable. In this paper, SSL discrete speech features extracted from WavLM models are used as additional cross-utterance acoustic context features in Zipformer-Transducer ASR systems. The efficacy of replacing Fbank features with discrete token features for modelling either cross-utterance contexts (from preceding and future segments), or current utterance's internal contexts alone, or both at the same time, are demonstrated thoroughly on the Gigaspeech 1000-hr corpus. The best Zipformer-Transducer system using discrete tokens based cross-utterance context features outperforms the baseline using utterance internal context only with statistically significant word error rate (WER) reductions of 0.32% to 0.41% absolute (2.78% to 3.54% relative) on the dev and test data. The lowest published WER of 11.15% and 11.14% were obtained on the dev and test sets. Our work is open-source and publicly available at https: github.com open-creator icefall tree master egs gigaspeech Context _ASR.
【24】 Energy Consumption Trends in Sound Event Detection Systems
标题: 声音事件检测系统的能源消耗趋势
作者:Constance Douwes,Romain Serizel
链接:点击下载PDF文件
摘要:深度学习系统变得越来越耗能和计算密集,引发了人们对其环境影响的担忧。作为声学场景和事件的检测和分类(DCASE)挑战的组织者,我们认识到解决这个问题的重要性。在过去的三年中,我们将能耗指标集成到声音事件检测(SED)系统的评估中。在本文中,我们分析了这种能量标准对挑战结果的影响,并探讨了多年来系统复杂性和能耗的演变。我们强调在培训过程中向更节能的方法转变,而不影响性能,同时操作数量和系统复杂性继续增长。我们希望透过这项分析,在社会经济发展界别内推广更环保的做法。摘要:Deep learning systems have become increasingly energy- and computation-intensive, raising concerns about their environmental impact. As organizers of the Detection and Classification of Acoustic Scenes and Events (DCASE) challenge, we recognize the importance of addressing this issue. For the past three years, we have integrated energy consumption metrics into the evaluation of sound event detection (SED) systems. In this paper, we analyze the impact of this energy criterion on the challenge results and explore the evolution of system complexity and energy consumption over the years. We highlight a shift towards more energy-efficient approaches during training without compromising performance, while the number of operations and system complexity continue to grow. Through this analysis, we hope to promote more environmentally friendly practices within the SED community.
【25】 DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
标题: DFADD:基于扩散和流匹配的音频Deepfake数据集
作者:Jiawei Du,I-Ming Lin,I-Hsiang Chiu,Xuanjun Chen,Haibin Wu,Wenze Ren,Yu Tsao,Hung-yi Lee,Jyh-Shing Roger Jang
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
摘要:主流的zero-shot TTS生成系统,如Voicebox和Seed-TTS,分别通过利用流匹配和扩散模型来实现人类对等语音。不幸的是,人类级别的音频合成会导致身份滥用和信息安全问题。目前,已经针对deepfake音频开发了许多反欺骗模型。然而,当前最先进的反欺骗模型在对抗由基于扩散和流匹配的TTS系统合成的音频方面的功效仍然未知。在本文中,我们提出了基于扩散和流匹配的音频Deepfake(DFADD)数据集。DFADD数据集收集了基于高级扩散和流匹配TTS模型的deepfake音频。此外,我们发现,目前的反欺骗模型缺乏足够的鲁棒性对扩散和流匹配TTS系统产生的高度人性化的音频。建议的DFADD数据集解决了这一差距,并为开发更具弹性的反欺骗模型提供了宝贵的资源。摘要:Mainstream zero-shot TTS production systems like Voicebox and Seed-TTS achieve human parity speech by leveraging Flow-matching and Diffusion models, respectively. Unfortunately, human-level audio synthesis leads to identity misuse and information security issues. Currently, many antispoofing models have been developed against deepfake audio. However, the efficacy of current state-of-the-art anti-spoofing models in countering audio synthesized by diffusion and flowmatching based TTS systems remains unknown. In this paper, we proposed the Diffusion and Flow-matching based Audio Deepfake (DFADD) dataset. The DFADD dataset collected the deepfake audio based on advanced diffusion and flowmatching TTS models. Additionally, we reveal that current anti-spoofing models lack sufficient robustness against highly human-like audio generated by diffusion and flow-matching TTS systems. The proposed DFADD dataset addresses this gap and provides a valuable resource for developing more resilient anti-spoofing models.
【26】 Acoustic identification of individual animals with hierarchical contrastive learning
标题: 利用分层对比学习进行个体动物的声学识别
作者:Ines Nolasco,Ilyass Moummad,Dan Stowell,Emmanouil Benetos
备注:Under review; Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:个体动物的声学识别(AIID)与基于音频的物种分类密切相关,但需要更精细的细节来区分同一物种中的个体动物。在这项工作中,我们框架AIID作为一个层次化的多标签分类任务,并提出使用层次感知损失函数来学习个体身份的鲁棒表示,保持物种和类群之间的层次关系。我们的研究结果表明,层次嵌入不仅提高了识别精度在个人层面上,而且在更高的分类水平,有效地保留了层次结构的学习表示。通过比较我们的方法与非层次模型,我们强调了在嵌入空间中执行这种结构的优势。此外,我们将评估扩展到新的单个类的分类,展示了我们的方法在开集分类场景中的潜力。摘要:Acoustic identification of individual animals (AIID) is closely related to audio-based species classification but requires a finer level of detail to distinguish between individual animals within the same species. In this work, we frame AIID as a hierarchical multi-label classification task and propose the use of hierarchy-aware loss functions to learn robust representations of individual identities that maintain the hierarchical relationships among species and taxa. Our results demonstrate that hierarchical embeddings not only enhance identification accuracy at the individual level but also at higher taxonomic levels, effectively preserving the hierarchical structure in the learned representations. By comparing our approach with non-hierarchical models, we highlight the advantage of enforcing this structure in the embedding space. Additionally, we extend the evaluation to the classification of novel individual classes, demonstrating the potential of our method in open-set classification scenarios.
【27】 Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
标题: 研究音素级语音编解码器中的解纠缠以进行韵律建模
作者:Sotirios Karapiperis,Nikolaos Ellinas,Alexandra Vioni,Junkwang Oh,Gunu Jho,Inchul Hwang,Spyros Raptis
链接:点击下载PDF文件
摘要:语音韵律建模中的大多数流行方法依赖于在一个连续的潜在空间中学习全局风格表示,该潜在空间编码和传递参考语音的属性。然而,最近基于残差矢量量化(RVQ)的神经编解码器的工作已经显示出巨大的潜力,提供了明显的优势。我们调查的韵律建模能力的离散空间这样的RVQ-VAE模型,修改它的音素水平上操作。我们的条件的编码器和解码器的语言表示的模型,并应用一个全球扬声器嵌入,以因素出语音和扬声器的信息。我们进行了广泛的调查,主观实验和客观措施的基础上,以表明,音素级的离散潜在表示,这种方式获得了高度的解开,捕获细粒度的韵律信息,是强大的和可转让的。潜在空间具有可解释的结构,其主成分对应于音高和能量。摘要:Most of the prevalent approaches in speech prosody modeling rely on learning global style representations in a continuous latent space which encode and transfer the attributes of reference speech. However, recent work on neural codecs which are based on Residual Vector Quantization (RVQ) already shows great potential offering distinct advantages. We investigate the prosody modeling capabilities of the discrete space of such an RVQ-VAE model, modifying it to operate on the phoneme-level. We condition both the encoder and decoder of the model on linguistic representations and apply a global speaker embedding in order to factor out both phonetic and speaker information. We conduct an extensive set of investigations based on subjective experiments and objective measures to show that the phoneme-level discrete latent representations obtained this way achieves a high degree of disentanglement, capturing fine-grained prosodic information that is robust and transferable. The latent space turns out to have interpretable structure with its principal components corresponding to pitch and energy.
【28】 LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
标题: LMAC-TD:为音频分类器生成时间域简化
作者:Eleonora Mancini,Francesco Paissan,Mirco Ravanelli,Cem Subakan
备注:The first two authors contributed equally to this research. Author order is alphabetical
链接:点击下载PDF文件
摘要:神经网络通常是黑箱,在决策机制方面保持不透明。文献中的一些作品提出了事后解释方法来缓解这个问题。本文提出了LMAC-TD,一种事后解释方法,训练解码器直接在时域中产生解释。这种方法建立在L-MAC的基础上,L-MAC是一种用于音频分类器的可听映射,这种方法可以产生忠实和可解释的解释。我们将SepFormer,一个流行的基于变压器的时域源分离架构。我们通过一项用户研究表明,LMAC-TD显着提高了音频质量的生产解释,同时不牺牲忠诚。摘要:Neural networks are typically black-boxes that remain opaque with regards to their decision mechanisms. Several works in the literature have proposed post-hoc explanation methods to alleviate this issue. This paper proposes LMAC-TD, a post-hoc explanation method that trains a decoder to produce explanations directly in the time domain. This methodology builds upon the foundation of L-MAC, Listenable Maps for Audio Classifiers, a method that produces faithful and listenable explanations. We incorporate SepFormer, a popular transformer-based time-domain source separation architecture. We show through a user study that LMAC-TD significantly improves the audio quality of the produced explanations while not sacrificing from faithfulness.
【29】 Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
标题: 节奏Foley:视频到音频合成中无缝视听对齐的框架
作者:Zhiqi Huang,Dan Luo,Jun Wang,Huan Liao,Zhiheng Li,Zhiyong Wu
链接:点击下载PDF文件
摘要:我们的研究引入了一个创新的框架,视频到音频合成,解决了音视频去标准化和语义损失的音频。通过结合语义对齐适配器和时间同步适配器,我们的方法显着提高了语义的完整性和节拍点同步的精度,特别是在快节奏的动作序列。利用对比视听预训练编码器,我们的模型使用视频和高质量音频数据进行训练,提高了生成音频的质量。这种双适配器方法使用户能够增强对音频语义和节拍效果的控制,允许调整控制器以实现更好的效果。大量的实验证实了我们的框架在实现无缝视听对齐的有效性。摘要:Our research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal synchronization adapter, our method significantly improves semantic integrity and the precision of beat point synchronization, particularly in fast-paced action sequences. Utilizing a contrastive audio-visual pre-trained encoder, our model is trained with video and high-quality audio data, improving the quality of the generated audio. This dual-adapter approach empowers users with enhanced control over audio semantics and beat effects, allowing the adjustment of the controller to achieve better results. Extensive experiments substantiate the effectiveness of our framework in achieving seamless audio-visual alignment.
【30】 TapToTab : Video-Based Guitar Tabs Generation using AI and Audio Analysis
标题: TapToTab:使用人工智能和音频分析的基于视频的吉他标签生成
作者:Ali Ghaleb,Eslam ElSadawy,Ihab Essam,Mohamed Abdelhakim,Seif-Eldin Zaki,Natalie Fahim,Razan Bayoumi,Hanan Hindy
链接:点击下载PDF文件
摘要:从视频输入的吉他指板生成的自动化对于提高音乐教育、转录准确性和性能分析具有重要的前景。现有的方法面临着一致性和完整性的挑战,特别是在检测指板和准确识别音符。为了解决这些问题,本文介绍了一种利用深度学习的先进方法,特别是用于实时指板检测的YOLO模型,以及用于精确音符识别的基于傅立叶变换的音频分析。实验结果表明,与传统技术相比,检测精度和鲁棒性有了很大的提高。本文概述了这些方法的开发,实施和评估,旨在通过自动化从视频录音创建吉他标签来彻底改变吉他教学。摘要:The automation of guitar tablature generation from video inputs holds significant promise for enhancing music education, transcription accuracy, and performance analysis. Existing methods face challenges with consistency and completeness, particularly in detecting fretboards and accurately identifying notes. To address these issues, this paper introduces an advanced approach leveraging deep learning, specifically YOLO models for real-time fretboard detection, and Fourier Transform-based audio analysis for precise note identification. Experimental results demonstrate substantial improvements in detection accuracy and robustness compared to traditional techniques. This paper outlines the development, implementation, and evaluation of these methodologies, aiming to revolutionize guitar instruction by automating the creation of guitar tabs from video recordings.
【31】 STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
标题: STA-V2 A:具有语义和时间对齐的视频到音频生成
作者:Yong Ren,Chenxing Li,Manjie Xu,Wei Liang,Yu Gu,Rilin Chen,Dong Yu
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
摘要:视觉和听觉是人类体验世界的两种重要方式。在过去的一年里,文本到视频的生成取得了显着的进展,但在生成的视频中缺乏和谐的音频限制了其更广泛的应用。在本文中,我们提出了语义和时间对齐的视频到音频(STA-V2 A),一种方法,通过提取本地时间和全局语义视频特征,并将这些细化的视频特征与文本相结合,作为跨模态指导,增强从视频中生成音频。为了解决视频中的信息冗余问题,我们提出了一个发病预测借口任务的局部时间特征提取和关注池模块的全局语义特征提取。为了补充视频中语义信息的不足,我们提出了一个带有文本到音频先验初始化和跨模态指导的潜在扩散模型。我们还介绍了Audio-Audio Align,一个新的指标来评估音频时间对齐。主观和客观的指标表明,我们的方法超越现有的视频到音频模型生成的音频具有更好的质量,语义一致性和时间对齐。烧蚀实验验证了各模块的有效性。音频样本可在https: y-ren16.github.io STAV2A上获得。摘要:Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader applications. In this paper, we propose Semantic and Temporal Aligned Video-to-Audio (STA-V2A), an approach that enhances audio generation from videos by extracting both local temporal and global semantic video features and combining these refined video features with text as cross-modal guidance. To address the issue of information redundancy in videos, we propose an onset prediction pretext task for local temporal feature extraction and an attentive pooling module for global semantic feature extraction. To supplement the insufficient semantic information in videos, we propose a Latent Diffusion Model with Text-to-Audio priors initialization and cross-modal guidance. We also introduce Audio-Audio Align, a new metric to assess audio-temporal alignment. Subjective and objective metrics demonstrate that our method surpasses existing Video-to-Audio models in generating audio with better quality, semantic consistency, and temporal alignment. The ablation experiment validated the effectiveness of each module. Audio samples are available at https: y-ren16.github.io STAV2A.
【32】 LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
标题: LA-RAG:通过检索增强生成增强基于LLM的ASB准确性
作者:Shaojun Li,Hengchao Shang,Daimeng Wei,Jiaxin Guo,Zongyao Li,Xianghui He,Min Zhang,Hao Yang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:将语音信息集成到大型语言模型(LLM)中的最新进展显着提高了自动语音识别(ASR)的准确性。然而,现有的方法往往受到语音编码器在不同声学条件下的能力的限制,例如口音。为了解决这个问题,我们提出了LA-RAG,一种新的检索增强生成(RAG)范式,用于基于LLM的ASR。LA-RAG利用细粒度令牌级语音数据存储和语音到语音检索机制,通过LLM上下文学习(ICL)功能提高ASR准确性。在普通话和各种中国方言数据集上的实验表明,与现有方法相比,ASR的准确性有了显着提高,验证了我们的方法的有效性,特别是在处理口音变化方面。摘要:Recent advancements in integrating speech information into large language models (LLMs) have significantly improved automatic speech recognition (ASR) accuracy. However, existing methods often constrained by the capabilities of the speech encoders under varied acoustic conditions, such as accents. To address this, we propose LA-RAG, a novel Retrieval-Augmented Generation (RAG) paradigm for LLM-based ASR. LA-RAG leverages fine-grained token-level speech datastores and a speech-to-speech retrieval mechanism to enhance ASR accuracy via LLM in-context learning (ICL) capabilities. Experiments on Mandarin and various Chinese dialect datasets demonstrate significant improvements in ASR accuracy compared to existing methods, validating the effectiveness of our approach, especially in handling accent variations.
【33】 Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
标题: 大型语言模型可以在多说话者场景中通过多功能指令转录语音
作者:Lingwei Meng,Shujie Hu,Jiawen Kang,Zhaoqing Li,Yuejiao Wang,Wenxuan Wu,Xixin Wu,Xunying Liu,Helen Meng
链接:点击下载PDF文件
摘要:大型语言模型(LLM)的最新进展已经彻底改变了各个领域,带来了重大进展和新机遇。尽管在语音相关的任务取得了进展,LLM还没有在多说话者场景中得到充分的探索。在这项工作中,我们提出了一个开创性的努力,调查LLM在多人环境中转录语音的能力,以下多功能指令相关的多人自动语音识别(ASR),目标说话者ASR和ASR的基础上,特定的说话者属性,如性别,发生顺序,语言和关键字发言。我们的方法利用WavLM和Whisper编码器来提取多方面的语音表示,是敏感的扬声器特性和语义上下文。然后,这些表示被输入到使用LoRA进行微调的LLM中,从而实现语音理解和转录的功能。综合实验揭示了我们提出的系统,MT-LLM,在鸡尾酒会的情况下,有前途的性能,突出LLM处理语音相关的任务,在这样复杂的设置基于用户指令的潜力。摘要:Recent advancements in large language models (LLMs) have revolutionized various domains, bringing significant progress and new opportunities. Despite progress in speech-related tasks, LLMs have not been sufficiently explored in multi-talker scenarios. In this work, we present a pioneering effort to investigate the capability of LLMs in transcribing speech in multi-talker environments, following versatile instructions related to multi-talker automatic speech recognition (ASR), target talker ASR, and ASR based on specific talker attributes such as sex, occurrence order, language, and keyword spoken. Our approach utilizes WavLM and Whisper encoder to extract multi-faceted speech representations that are sensitive to speaker characteristics and semantic context. These representations are then fed into an LLM fine-tuned using LoRA, enabling the capabilities for speech comprehension and transcription. Comprehensive experiments reveal the promising performance of our proposed system, MT-LLM, in cocktail party scenarios, highlighting the potential of LLM to handle speech-related tasks based on user instructions in such complex settings.
【34】 Domain-Invariant Representation Learning of Bird Sounds
标题: 鸟声的域不变表示学习
作者:Ilyass Moummad,Romain Serizel,Emmanouil Benetos,Nicolas Farrugia
链接:点击下载PDF文件
摘要:被动声监测(PAM)是生物声学研究的关键,可以实现非侵入性物种跟踪和生物多样性监测。像Xeno-Canto这样的公民科学平台提供了来自焦点记录的大型注释数据集,其中目标物种被有意记录。然而,PAM需要在被动音景中进行监控,在焦点和被动录音之间产生域转移,这对在焦点录音上训练的深度学习模型提出了挑战。为了解决这个问题,我们利用监督对比学习来提高鸟类声音分类中的领域泛化,在不同领域的同类样本中实现领域不变性。我们还提出了Prototypical(原型对比学习的表示),它降低了计算复杂性的SupCon损失比较的例子类原型,而不是成对的比较。此外,我们提出了一个新的Few-Shot分类基准的基础上BirdSet,一个大规模的鸟的声音数据集,并证明了我们的方法在实现强大的传输性能的有效性。摘要:Passive acoustic monitoring (PAM) is crucial for bioacoustic research, enabling non-invasive species tracking and biodiversity monitoring. Citizen science platforms like Xeno-Canto provide large annotated datasets from focal recordings, where the target species is intentionally recorded. However, PAM requires monitoring in passive soundscapes, creating a domain shift between focal and passive recordings, which challenges deep learning models trained on focal recordings. To address this, we leverage supervised contrastive learning to improve domain generalization in bird sound classification, enforcing domain invariance across same-class examples from different domains. We also propose ProtoCLR (Prototypical Contrastive Learning of Representations), which reduces the computational complexity of the SupCon loss by comparing examples to class prototypes instead of pairwise comparisons. Additionally, we present a new few-shot classification benchmark based on BirdSet, a large-scale bird sound dataset, and demonstrate the effectiveness of our approach in achieving strong transfer performance.
【35】 LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
标题: LHQ-SRC:轻量级、高质量的歌唱声音转换建模
作者:Yubo Huang,Xin Lai,Muyang Ye,Anran Zhu,Zixi Wang,Jingzehua Xu,Shuai Zhang,Zhiyuan Zhou,Weijie Niu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:歌唱声音转换(SVC)已经成为声音转换(VC)的一个重要子领域,使一个歌手的声音转换成另一个,同时保留音乐元素,如旋律,节奏和音色。传统的SVC方法在音频质量、数据要求和计算复杂度方面具有局限性。在本文中,我们提出了LHQ-SVC,一个轻量级的,CPU兼容的模型的基础上的SVC框架和扩散模型,旨在减少模型的大小和计算需求,而不牺牲性能。我们整合了一些功能来提高推理质量,并通过使用性能调优工具和并行计算框架来优化CPU执行。我们的实验表明,LHQ-SVC保持有竞争力的性能,在不同的设备处理速度和效率显着提高。结果表明,LHQ-SVC可以满足摘要:Singing Voice Conversion (SVC) has emerged as a significant subfield of Voice Conversion (VC), enabling the transformation of one singer's voice into another while preserving musical elements such as melody, rhythm, and timbre. Traditional SVC methods have limitations in terms of audio quality, data requirements, and computational complexity. In this paper, we propose LHQ-SVC, a lightweight, CPU-compatible model based on the SVC framework and diffusion model, designed to reduce model size and computational demand without sacrificing performance. We incorporate features to improve inference quality, and optimize for CPU execution by using performance tuning tools and parallel computing frameworks. Our experiments demonstrate that LHQ-SVC maintains competitive performance, with significant improvements in processing speed and efficiency across different devices. The results suggest that LHQ-SVC can meet
【36】 Apollo: Band-sequence Modeling for High-Quality Audio Restoration
标题: Apollo:用于高质量音频恢复的带序列建模
作者:Kai Li,Yi Luo
备注:Demo Page: this https URL
链接:点击下载PDF文件
摘要:音频恢复在现代社会中变得越来越重要,这不仅是因为对高级回放设备所实现的高质量听觉体验的需求,而且还因为生成音频模型的不断增长的能力需要高保真音频。通常,音频恢复被定义为从受损输入预测未失真音频的任务,通常使用GAN框架进行训练以平衡感知和失真。由于音频降级主要集中在中高频范围,特别是由于编解码器,因此一个关键的挑战在于设计一个能够保留低频信息同时准确重建高质量中高频内容的生成器。受高采样率音乐分离,语音增强和音频编解码器模型的最新进展的启发,我们提出了阿波罗,一个生成模型设计的高采样率音频恢复。Apollo采用了一个显式的频带分割模块来模拟不同频带之间的关系,从而实现更连贯和更高质量的恢复音频。在MUSDB 18-HQ和MoisesDB数据集上进行评估,Apollo在各种比特率和音乐类型上的表现始终优于现有的SR-GAN模型,特别是在涉及多种乐器和人声混合的复杂场景中表现出色。Apollo在保持计算效率的同时显著提高了音乐恢复质量。Apollo的源代码可在https: github.com JusperLee Apollo上公开获得。摘要:Audio restoration has become increasingly significant in modern society, not only due to the demand for high-quality auditory experiences enabled by advanced playback devices, but also because the growing capabilities of generative audio models necessitate high-fidelity audio. Typically, audio restoration is defined as a task of predicting undistorted audio from damaged input, often trained using a GAN framework to balance perception and distortion. Since audio degradation is primarily concentrated in mid- and high-frequency ranges, especially due to codecs, a key challenge lies in designing a generator capable of preserving low-frequency information while accurately reconstructing high-quality mid- and high-frequency content. Inspired by recent advancements in high-sample-rate music separation, speech enhancement, and audio codec models, we propose Apollo, a generative model designed for high-sample-rate audio restoration. Apollo employs an explicit frequency band split module to model the relationships between different frequency bands, allowing for more coherent and higher-quality restored audio. Evaluated on the MUSDB18-HQ and MoisesDB datasets, Apollo consistently outperforms existing SR-GAN models across various bit rates and music genres, particularly excelling in complex scenarios involving mixtures of multiple instruments and vocals. Apollo significantly improves music restoration quality while maintaining computational efficiency. The source code for Apollo is publicly available at https: github.com JusperLee Apollo.
【37】 Confidence Calibration for Audio Captioning Models
标题: 音频字幕模型的置信度校准
作者:Rehana Mahfuz,Yinyi Guo,Erik Visser
链接:点击下载PDF文件
摘要:自动生成音频、图像和视频的文本字幕的系统缺乏所生成序列的相关性和正确性的置信度指示符。为了解决这个问题,我们在现有的文本置信度测量方法的基础上,引入了标记概率的选择性池化,这比传统的池化更符合传统的正确性测量。此外,我们建议直接测量输入音频和文本之间的相似性在一个共享的嵌入空间。为了测量自我一致性,我们适应语义熵的音频字幕,并发现这两种方法甚至比池为基础的指标与正确性的措施,计算字幕之间的声学相似性。最后,我们解释了为什么温度缩放的信心,提高校准。摘要:Systems that automatically generate text captions for audio, images and video lack a confidence indicator of the relevance and correctness of the generated sequences. To address this, we build on existing methods of confidence measurement for text by introduce selective pooling of token probabilities, which aligns better with traditional correctness measures than conventional pooling does. Further, we propose directly measuring the similarity between input audio and text in a shared embedding space. To measure self-consistency, we adapt semantic entropy for audio captioning, and find that these two methods align even better than pooling-based metrics with the correctness measure that calculates acoustic similarity between captions. Finally, we explain why temperature scaling of confidences improves calibration.
机器翻译,仅供参考
