本文经arXiv每日学术速递授权转载
【1】 Towards Musically Informed Evaluation of Piano Transcription Models
标题: 钢琴抄写模型的音乐知情评估
作者:Patricia Hu,Lukáš Samuel Marták,Carlos Cancino-Chacón,Gerhard Widmer
链接:点击下载PDF文件
【2】 TokSing: Singing Voice Synthesis based on Discrete Tokens
标题: TokSing:基于离散令牌的歌唱声音合成
作者:Yuning Wu,Chunlei zhang,Jiatong Shi,Yuxun Tang,Shan Yang,Qin Jin
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【3】 Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models
标题: 迪夫-A-里夫:通过潜在扩散模型的音乐伴奏共同创作
作者:Javier Nistal,Marco Pasini,Cyran Aouameur,Maarten Grachten,Stefan Lattner
备注:8 pages, 2 figures, 3 tables
链接:点击下载PDF文件
【4】 Towards Unsupervised Speech Recognition Without Pronunciation Models
标题: 迈向没有发音模型的无监督语音识别
作者:Junrui Ni,Liming Wang,Yang Zhang,Kaizhi Qian,Heting Gao,Mark Hasegawa-Johnson,Chang D. Yoo
备注:This work has been submitted to the IEEE for possible publication
链接:点击下载PDF文件
【5】 CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
标题: CoLM-SVR:利用神经编解码语言建模进行多模式发音障碍语音重建
作者:Xueyuan Chen,Dongchao Yang,Dingdong Wang,Xixin Wu,Zhiyong Wu,Helen Meng
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【6】 Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
标题: 在说话人嵌入中使用对抗扰动的非同步语音
作者:Rui Wang,Liping Chen,Kong AiK Lee,Zhen-Hua Ling
链接:点击下载PDF文件
【7】 FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
标题: FreeV:通过伪反梅尔过滤器为声码者提供免费午餐
作者:Yuanjun Lv,Hai Li,Ying Yan,Junhui Liu,Danming Xie,Lei Xie
备注:Accepted by InterSpeech 2024; 5 pages, 5 figures
链接:点击下载PDF文件
【8】 Codecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio
标题: Codecfake:用于检测基于LLM的Deepfake音频的初始数据集
作者:Yi Lu,Yuankun Xie,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Zhiyong Wang,Xin Qi,Xuefei Liu,Yongwei Li,Yukun Liu,Xiaopeng Wang,Shuchen Shi
备注:Accepted by INTERSPEECH 2024. arXiv admin note: substantial text overlap with arXiv:2405.04880
链接:点击下载PDF文件
【9】 FakeSound: Deepfake General Audio Detection
标题: FakeSound:Deepfake通用音频检测
作者:Zeyu Xie,Baihan Li,Xuenan Xu,Zheng Liang,Kai Yu,Mengyue Wu
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【10】 CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
标题: 用于流媒体开放词汇关键词查找的符合ATC的音频文本嵌入
作者:Sichen Jin,Youngmoon Jung,Seungjin Lee,Jaeyoung Roh,Changwoo Han,Hoonyoung Cho
链接:点击下载PDF文件
【11】 Can Large Language Models Understand Spatial Audio?
标题: 大型语言模型能理解空间音频吗?
作者:Changli Tang,Wenyi Yu,Guangzhi Sun,Xianzhao Chen,Tian Tan,Wei Li,Jun Zhang,Lu Lu,Zejun Ma,Yuxuan Wang,Chao Zhang
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
【12】 Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
标题: 探索自我监督多视图对比学习用于有限注释的语音情感识别
作者:Bulat Khaertdinov,Pedro Jeuris,Annanda Sousa,Enrique Hortal
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【13】 Flexible Music-Conditioned Dance Generation with Style Description Prompts
标题: 具有风格描述的灵活音乐条件舞蹈生成
作者:Hongsong Wang,Yin Zhu,Xin Geng
链接:点击下载PDF文件
【14】 VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
标题: WAL-E R:通过单调对齐实现稳健高效的Zero-Shot文本到语音合成
作者:Bing Han,Long Zhou,Shujie Liu,Sanyuan Chen,Lingwei Meng,Yanming Qian,Yanqing Liu,Sheng Zhao,Jinyu Li,Furu Wei
备注:15 pages, 5 figures
链接:点击下载PDF文件
【15】 Zero-Shot Fake Video Detection by Audio-Visual Consistency
标题: 利用视听一致性进行Zero-Shot假视频检测
作者:Xiaolou Li,Zehua Liu,Chen Chen,Lantian Li,Li Guo,Dong Wang
备注:to be published in INTERSPEECH 2024
链接:点击下载PDF文件
【16】 SEBN Adapter: Parametric Efficient Domain Adaptation for Speaker Recognition
标题: SEBN适配器:用于说话人识别的参数高效域自适应
作者:Tianhao Wang,Lantian Li,Dong Wang
备注:to be published in INTERSPEECH 2024
链接:点击下载PDF文件
【17】 PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
标题: PRoDeliberation:端到端口语理解的并行稳健审议
作者:Trang Le,Daniel Lazar,Suyoun Kim,Shan Jiang,Duc Le,Adithya Sagar,Aleksandr Livshits,Ahmed Aly,Akshat Shrivastava
链接:点击下载PDF文件
【18】 EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
标题: SEARCH Sphere-TTC:通过球形情感载体进行情感风格和强度建模,用于可控情感文本到语音
作者:Deok-Hyeon Cho,Hyung-Seok Oh,Seung-Bin Kim,Sang-Hoon Lee,Seong-Whan Lee
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【19】 PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
标题: PolySpeech:探索统一的多任务语音模型,以与单任务模型竞争
作者:Runyan Yang,Huibao Yang,Xiqing Zhang,Tiantian Ye,Ying Liu,Yingying Gao,Shilei Zhang,Chao Deng,Junlan Feng
备注:5 pages, 2 figures
链接:点击下载PDF文件
【20】 The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
标题: Interspeech 2024年使用离散单元的语音处理挑战
作者:Xuankai Chang,Jiatong Shi,Jinchuan Tian,Yuning Wu,Yuxun Tang,Yihan Wu,Shinji Watanabe,Yossi Adi,Xie Chen,Qin Jin
备注:This manuscript has been accepted by Interspeech2024
链接:点击下载PDF文件
【21】 FastAST: Accelerating Audio Spectrogram Transformer via Token Merging and Cross-Model Knowledge Distillation
标题: FastAST:通过令牌合并和跨模型知识提炼加速音频频谱图Transformer
作者:Swarup Ranjan Behera,Abhishek Dhiman,Karthik Gowda,Aalekhya Satya Narayani
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【22】 Broadband MEMS Microphone Arrays with Reduced Aperture Through 3D-Printed Waveguides
标题: 通过3D打印光路实现缩小口径的宽带微机电麦克风阵列
作者:Dennis Laurijssen,Walter Daems,Jan Steckel
链接:点击下载PDF文件
【23】 Pre-training Feature Guided Diffusion Model for Speech Enhancement
标题: 用于语音增强的预训练特征引导扩散模型
作者:Yiyuan Yang,Niki Trigoni,Andrew Markham
备注:Accepted by Interspeech 2024 Conference
链接:点击下载PDF文件
【24】 SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
标题: SVSNet+:利用语音基础模型的表示增强说话者语音相似性评估模型
作者:Chun Yin,Tai-Shih Chi,Yu Tsao,Hsin-Min Wang
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
【25】 Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
标题: 理解声音,错过问题:大型音频语言模型中对象幻觉的挑战
作者:Chun-Yi Kuan,Wei-Ping Huang,Hung-yi Lee
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【26】 SCDNet: Self-supervised Learning Feature-based Speaker Change Detection
标题: SCDNet:基于自我监督学习的说话者变化检测
作者:Yue Li,Xinsheng Wang,Li Zhang,Lei Xie
链接:点击下载PDF文件
【27】 Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
标题: 利用ASB文字记录进行语音情感识别:错误率和融合技术的综合研究
作者:Yuanchao Li,Peter Bell,Catherine Lai
链接:点击下载PDF文件
【28】 Refining Self-Supervised Learnt Speech Representation using Brain Activations
标题: 使用大脑激活完善自我监督学习语音表达
作者:Hengyu Li,Kangdi Mei,Zhaoci Liu,Yang Ai,Liping Chen,Jie Zhang,Zhenhua Ling
备注:accpeted by Interspeech2024
链接:点击下载PDF文件
【29】 Transformer-based Model for ASR N-Best Rescoring and Rewriting
标题: 基于转换器的ASB N-Best重新评分和重写模型
作者:Iwen E. Kang,Christophe Van Gysel,Man-Hung Siu
备注:Interspeech '24
链接:点击下载PDF文件
【30】 LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
标题: LAFMA:一种用于文本到音频生成的潜在流匹配模型
作者:Wenhao Guan,Kaidi Wang,Wangjin Zhou,Yang Wang,Feng Deng,Hui Wang,Lin Li,Qingyang Hong,Yong Qin
备注:Accepted at Interspeech2024
链接:点击下载PDF文件
【31】 Fully Few-shot Class-incremental Audio Classification Using Expandable Dual-embedding Extractor
标题: 使用可扩展双嵌入提取器的全Few-Shot类增量音频分类
作者:Yongjie Si,Yanxiong Li,Jialong Li,Jiaxin Tan,Qianhua He
备注:Accepted for publication on Interspeech 2024. 5 pages, 3 figures, 5 tables
链接:点击下载PDF文件
【32】 Low-Complexity Acoustic Scene Classification Using Parallel Attention-Convolution Network
标题: 使用并行注意卷积网络的低复杂度声场景分类
作者:Yanxiong Li,Jiaxin Tan,Guoqing Chen,Jialong Li,Yongjie Si,Qianhua He
备注:Accepted for publication on Interspeech 2024. 5 pages, 4 figures, 3 tables
链接:点击下载PDF文件
【33】 VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
标题: VECL-TTC:语音身份和情感风格可控跨语言文本转语音
作者:Ashishkumar Gudmalwar,Nirmesh Shah,Sai Akarsh,Pankaj Wasnik,Rajiv Ratn Shah
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【34】 DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels
标题: DUSE 2024任务4:使用异类数据和缺失标签的声音事件检测
作者:Samuele Cornell,Janek Ebbers,Constance Douwes,Irene Martín-Morató,Manu Harju,Annamaria Mesaros,Romain Serizel
链接:点击下载PDF文件
【35】 LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
标题: LibriTTS-P:具有说话风格和说话者身份的数据库,支持文本转语音和风格字幕
作者:Masaya Kawamura,Ryuichi Yamamoto,Yuma Shirahata,Takuya Hasumi,Kentaro Tachibana
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
【36】 Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
标题: 使用自我知识蒸馏指导框架级CSC对准
作者:Eungbeom Kim,Hantae Kim,Kyogu Lee
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【37】 Target Speaker Extraction with Curriculum Learning
标题: 通过课程学习提取目标说话者
作者:Yun Liu,Xuechen Liu,Xiaoxiao Miao,Junichi Yamagishi
备注:Accepted for presentation at Interspeech 2024
链接:点击下载PDF文件
【38】 Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
作者:Lin Zhang,Xin Wang,Erica Cooper,Mireia Diez,Federico Landini,Nicholas Evans,Junichi Yamagishi
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【39】 Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based models
标题: 实现客观和可解释的言语障碍评估:CNN和基于转换器的模型的比较分析
作者:Malo Maisonneuve,Corinne Fredouille,Muriel Lalain,Alain Ghio,Virginie Woisard
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
标题: SVSNet+:利用语音基础模型的表示增强说话者语音相似性评估模型
作者:Chun Yin,Tai-Shih Chi,Yu Tsao,Hsin-Min Wang
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
【2】 Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
标题: 理解声音,错过问题:大型音频语言模型中对象幻觉的挑战
作者:Chun-Yi Kuan,Wei-Ping Huang,Hung-yi Lee
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【3】 Neural Blind Source Separation and Diarization for Distant Speech Recognition
标题: 用于远距离语音识别的神经盲源分离和扩展
作者:Yoshiaki Bando,Tomohiko Nakamura,Shinji Watanabe
备注:5 pages, 3 figures, accepted to INTERSPEECH 2024
链接:点击下载PDF文件
【4】 SCDNet: Self-supervised Learning Feature-based Speaker Change Detection
标题: SCDNet:基于自我监督学习的说话者变化检测
作者:Yue Li,Xinsheng Wang,Li Zhang,Lei Xie
链接:点击下载PDF文件
【5】 Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
标题: 利用ASB文字记录进行语音情感识别:错误率和融合技术的综合研究
作者:Yuanchao Li,Peter Bell,Catherine Lai
链接:点击下载PDF文件
【6】 Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation
标题: 规则化语音分离的定时文本和音频之间的多模式表示损失
作者:Tsun-An Hsieh,Heeyoul Choi,Minje Kim
链接:点击下载PDF文件
【7】 Refining Self-Supervised Learnt Speech Representation using Brain Activations
标题: 使用大脑激活完善自我监督学习语音表达
作者:Hengyu Li,Kangdi Mei,Zhaoci Liu,Yang Ai,Liping Chen,Jie Zhang,Zhenhua Ling
备注:accpeted by Interspeech2024
链接:点击下载PDF文件
【8】 Transformer-based Model for ASR N-Best Rescoring and Rewriting
标题: 基于转换器的ASB N-Best重新评分和重写模型
作者:Iwen E. Kang,Christophe Van Gysel,Man-Hung Siu
备注:Interspeech '24
链接:点击下载PDF文件
【9】 LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
标题: LAFMA:一种用于文本到音频生成的潜在流匹配模型
作者:Wenhao Guan,Kaidi Wang,Wangjin Zhou,Yang Wang,Feng Deng,Hui Wang,Lin Li,Qingyang Hong,Yong Qin
备注:Accepted at Interspeech2024
链接:点击下载PDF文件
【10】 Fully Few-shot Class-incremental Audio Classification Using Expandable Dual-embedding Extractor
标题: 使用可扩展双嵌入提取器的全Few-Shot类增量音频分类
作者:Yongjie Si,Yanxiong Li,Jialong Li,Jiaxin Tan,Qianhua He
备注:Accepted for publication on Interspeech 2024. 5 pages, 3 figures, 5 tables
链接:点击下载PDF文件
【11】 Low-Complexity Acoustic Scene Classification Using Parallel Attention-Convolution Network
标题: 使用并行注意卷积网络的低复杂度声场景分类
作者:Yanxiong Li,Jiaxin Tan,Guoqing Chen,Jialong Li,Yongjie Si,Qianhua He
备注:Accepted for publication on Interspeech 2024. 5 pages, 4 figures, 3 tables
链接:点击下载PDF文件
【12】 Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
标题: 用于从未标记的语音数据构建文本到语音模型的音频条件音素和韵律注释
作者:Yuma Shirahata,Byeongseon Park,Ryuichi Yamamoto,Kentaro Tachibana
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
【13】 VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
标题: VECL-TTC:语音身份和情感风格可控跨语言文本转语音
作者:Ashishkumar Gudmalwar,Nirmesh Shah,Sai Akarsh,Pankaj Wasnik,Rajiv Ratn Shah
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【14】 DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels
标题: DUSE 2024任务4:使用异类数据和缺失标签的声音事件检测
作者:Samuele Cornell,Janek Ebbers,Constance Douwes,Irene Martín-Morató,Manu Harju,Annamaria Mesaros,Romain Serizel
链接:点击下载PDF文件
【15】 LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
标题: LibriTTS-P:具有说话风格和说话者身份的数据库,支持文本转语音和风格字幕
作者:Masaya Kawamura,Ryuichi Yamamoto,Yuma Shirahata,Takuya Hasumi,Kentaro Tachibana
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
【16】 Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
标题: 使用自我知识蒸馏指导框架级CSC对准
作者:Eungbeom Kim,Hantae Kim,Kyogu Lee
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【17】 Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
标题: 探索儿童与成人二元互动中说话者扩大化的言语基础模型
作者:Anfeng Xu,Kevin Huang,Tiantian Feng,Lue Shen,Helen Tager-Flusberg,Shrikanth Narayanan
备注:Interspeech 2024
链接:点击下载PDF文件
【18】 DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion
标题: DualVC 3:利用语言模型生成的伪上下文进行端到端低延迟流媒体语音转换
作者:Ziqian Ning,Shuai Wang,Pengcheng Zhu,Zhichao Wang,Jixun Yao,Lei Xie,Mengxiao Bi
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【19】 Target Speaker Extraction with Curriculum Learning
标题: 通过课程学习提取目标说话者
作者:Yun Liu,Xuechen Liu,Xiaoxiao Miao,Junichi Yamagishi
备注:Accepted for presentation at Interspeech 2024
链接:点击下载PDF文件
【20】 Dual-Pipeline with Low-Rank Adaptation for New Language Integration in Multilingual ASR
标题: 低等级自适应的双管道用于多语言ASB中的新语言集成
作者:Yerbolat Khassanov,Zhipeng Chen,Tianfeng Chen,Tze Yuang Chong,Wei Li,Jun Zhang,Lu Lu,Yuxuan Wang
备注:5 pages, 2 figures, 4 tables
链接:点击下载PDF文件
【21】 Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
作者:Lin Zhang,Xin Wang,Erica Cooper,Mireia Diez,Federico Landini,Nicholas Evans,Junichi Yamagishi
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【22】 Broadband MEMS Microphone Arrays with Reduced Aperture Through 3D-Printed Waveguides
标题: 通过3D打印光路实现缩小口径的宽带微机电麦克风阵列
作者:Dennis Laurijssen,Walter Daems,Jan Steckel
链接:点击下载PDF文件
【23】 Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based models
标题: 实现客观和可解释的言语障碍评估:CNN和基于转换器的模型的比较分析
作者:Malo Maisonneuve,Corinne Fredouille,Muriel Lalain,Alain Ghio,Virginie Woisard
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【24】 Towards Musically Informed Evaluation of Piano Transcription Models
标题: 钢琴抄写模型的音乐知情评估
作者:Patricia Hu,Lukáš Samuel Marták,Carlos Cancino-Chacón,Gerhard Widmer
链接:点击下载PDF文件
【25】 TokSing: Singing Voice Synthesis based on Discrete Tokens
标题: TokSing:基于离散令牌的歌唱声音合成
作者:Yuning Wu,Chunlei zhang,Jiatong Shi,Yuxun Tang,Shan Yang,Qin Jin
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【26】 Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models
标题: 迪夫-A-里夫:通过潜在扩散模型的音乐伴奏共同创作
作者:Javier Nistal,Marco Pasini,Cyran Aouameur,Maarten Grachten,Stefan Lattner
备注:8 pages, 2 figures, 3 tables
链接:点击下载PDF文件
【27】 Towards Unsupervised Speech Recognition Without Pronunciation Models
标题: 迈向没有发音模型的无监督语音识别
作者:Junrui Ni,Liming Wang,Yang Zhang,Kaizhi Qian,Heting Gao,Mark Hasegawa-Johnson,Chang D. Yoo
备注:This work has been submitted to the IEEE for possible publication
链接:点击下载PDF文件
【28】 CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
标题: CoLM-SVR:利用神经编解码语言建模进行多模式发音障碍语音重建
作者:Xueyuan Chen,Dongchao Yang,Dingdong Wang,Xixin Wu,Zhiyong Wu,Helen Meng
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【29】 Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
标题: 在说话人嵌入中使用对抗扰动的非同步语音
作者:Rui Wang,Liping Chen,Kong AiK Lee,Zhen-Hua Ling
链接:点击下载PDF文件
【30】 FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
标题: FreeV:通过伪反梅尔过滤器为声码者提供免费午餐
作者:Yuanjun Lv,Hai Li,Ying Yan,Junhui Liu,Danming Xie,Lei Xie
备注:Accepted by InterSpeech 2024; 5 pages, 5 figures
链接:点击下载PDF文件
【31】 Codecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio
标题: Codecfake:用于检测基于LLM的Deepfake音频的初始数据集
作者:Yi Lu,Yuankun Xie,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Zhiyong Wang,Xin Qi,Xuefei Liu,Yongwei Li,Yukun Liu,Xiaopeng Wang,Shuchen Shi
备注:Accepted by INTERSPEECH 2024. arXiv admin note: substantial text overlap with arXiv:2405.04880
链接:点击下载PDF文件
【32】 FakeSound: Deepfake General Audio Detection
标题: FakeSound:Deepfake通用音频检测
作者:Zeyu Xie,Baihan Li,Xuenan Xu,Zheng Liang,Kai Yu,Mengyue Wu
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【33】 CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
标题: 用于流媒体开放词汇关键词查找的符合ATC的音频文本嵌入
作者:Sichen Jin,Youngmoon Jung,Seungjin Lee,Jaeyoung Roh,Changwoo Han,Hoonyoung Cho
链接:点击下载PDF文件
【34】 Can Large Language Models Understand Spatial Audio?
标题: 大型语言模型能理解空间音频吗?
作者:Changli Tang,Wenyi Yu,Guangzhi Sun,Xianzhao Chen,Tian Tan,Wei Li,Jun Zhang,Lu Lu,Zejun Ma,Yuxuan Wang,Chao Zhang
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
【35】 Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
标题: 探索自我监督多视图对比学习用于有限注释的语音情感识别
作者:Bulat Khaertdinov,Pedro Jeuris,Annanda Sousa,Enrique Hortal
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【36】 Flexible Music-Conditioned Dance Generation with Style Description Prompts
标题: 具有风格描述的灵活音乐条件舞蹈生成
作者:Hongsong Wang,Yin Zhu,Xin Geng
链接:点击下载PDF文件
【37】 VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
标题: WAL-E R:通过单调对齐实现稳健高效的Zero-Shot文本到语音合成
作者:Bing Han,Long Zhou,Shujie Liu,Sanyuan Chen,Lingwei Meng,Yanming Qian,Yanqing Liu,Sheng Zhao,Jinyu Li,Furu Wei
备注:15 pages, 5 figures
链接:点击下载PDF文件
【38】 Zero-Shot Fake Video Detection by Audio-Visual Consistency
标题: 利用视听一致性进行Zero-Shot假视频检测
作者:Xiaolou Li,Zehua Liu,Chen Chen,Lantian Li,Li Guo,Dong Wang
备注:to be published in INTERSPEECH 2024
链接:点击下载PDF文件
【39】 SEBN Adapter: Parametric Efficient Domain Adaptation for Speaker Recognition
标题: SEBN适配器:用于说话人识别的参数高效域自适应
作者:Tianhao Wang,Lantian Li,Dong Wang
备注:to be published in INTERSPEECH 2024
链接:点击下载PDF文件
【40】 PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
标题: PRoDeliberation:端到端口语理解的并行稳健审议
作者:Trang Le,Daniel Lazar,Suyoun Kim,Shan Jiang,Duc Le,Adithya Sagar,Aleksandr Livshits,Ahmed Aly,Akshat Shrivastava
链接:点击下载PDF文件
【41】 EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
标题: SEARCH Sphere-TTC:通过球形情感载体进行情感风格和强度建模,用于可控情感文本到语音
作者:Deok-Hyeon Cho,Hyung-Seok Oh,Seung-Bin Kim,Sang-Hoon Lee,Seong-Whan Lee
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【42】 PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
标题: PolySpeech:探索统一的多任务语音模型,以与单任务模型竞争
作者:Runyan Yang,Huibao Yang,Xiqing Zhang,Tiantian Ye,Ying Liu,Yingying Gao,Shilei Zhang,Chao Deng,Junlan Feng
备注:5 pages, 2 figures
链接:点击下载PDF文件
【43】 The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
标题: Interspeech 2024年使用离散单元的语音处理挑战
作者:Xuankai Chang,Jiatong Shi,Jinchuan Tian,Yuning Wu,Yuxun Tang,Yihan Wu,Shinji Watanabe,Yossi Adi,Xie Chen,Qin Jin
备注:This manuscript has been accepted by Interspeech2024
链接:点击下载PDF文件
【44】 FastAST: Accelerating Audio Spectrogram Transformer via Token Merging and Cross-Model Knowledge Distillation
标题: FastAST:通过令牌合并和跨模型知识提炼加速音频频谱图Transformer
作者:Swarup Ranjan Behera,Abhishek Dhiman,Karthik Gowda,Aalekhya Satya Narayani
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【45】 Pre-training Feature Guided Diffusion Model for Speech Enhancement
标题: 用于语音增强的预训练特征引导扩散模型
作者:Yiyuan Yang,Niki Trigoni,Andrew Markham
备注:Accepted by Interspeech 2024 Conference
链接:点击下载PDF文件
标题: 钢琴抄写模型的音乐知情评估
作者:Patricia Hu,Lukáš Samuel Marták,Carlos Cancino-Chacón,Gerhard Widmer
链接:点击下载PDF文件
摘要:自动钢琴转录模型通常使用简单的逐帧或逐音符信息检索(IR)度量来评估。这样的基准度量不能提供对特定音乐方面的转录质量的洞察,例如输出的清晰度、动态或节奏精度,这些在表达性表现分析的上下文中是必不可少的。此外,近年来,MAESTRO已成为此类模型事实上的训练和评估数据集。然而,推理性能已被观察到大大恶化时,适用于分布外的数据,从而质疑的适用性和可靠性,转录输出从这些模型的特定MIR任务。在这项工作中,我们调查了三个国家的最先进的钢琴转录模型在两个实验中的性能。在第一个中,我们提出了各种音乐上知情的评价指标,与IR指标相比,提供了更详细的了解音乐质量的transmittance。在第二个实验中,我们比较了真实世界和干扰录音的推理性能,并强调了我们的指标可以帮助解释的音乐维度。我们的实验结果突出了现有的钢琴转录指标的弱点,并有助于更音乐的声音错误分析的转录输出。摘要:Automatic piano transcription models are typically evaluated using simple frame- or note-wise information retrieval (IR) metrics. Such benchmark metrics do not provide insights into the transcription quality of specific musical aspects such as articulation, dynamics, or rhythmic precision of the output, which are essential in the context of expressive performance analysis. Furthermore, in recent years, MAESTRO has become the de-facto training and evaluation dataset for such models. However, inference performance has been observed to deteriorate substantially when applied on out-of-distribution data, thereby questioning the suitability and reliability of transcribed outputs from such models for specific MIR tasks. In this work, we investigate the performance of three state-of-the-art piano transcription models in two experiments. In the first one, we propose a variety of musically informed evaluation metrics which, in contrast to the IR metrics, offer more detailed insight into the musical quality of the transcriptions. In the second experiment, we compare inference performance on real-world and perturbed audio recordings, and highlight musical dimensions which our metrics can help explain. Our experimental results highlight the weaknesses of existing piano transcription metrics and contribute to a more musically sound error analysis of transcription outputs.
【2】 TokSing: Singing Voice Synthesis based on Discrete Tokens
标题: TokSing:基于离散令牌的歌唱声音合成
作者:Yuning Wu,Chunlei zhang,Jiatong Shi,Yuxun Tang,Shan Yang,Qin Jin
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:语音合成的最新进展通过利用从自监督学习(SSL)模型中提取的离散令牌来见证显着的好处。与传统的连续Mel谱图相比,离散标记在中间表示中提供更高的存储效率和更大的可操作性。然而,当涉及到歌唱声音合成(SVS),实现更高层次的旋律表达提出了一个很大的挑战,利用离散令牌。在本文中,我们介绍了TokSing,一个基于离散的SVS系统配备了一个令牌配方,提供灵活的令牌混合。我们在离散化过程中观察到旋律退化,这促使我们将旋律信号与离散令牌集成,并在音乐编码器中采用专门设计的旋律增强策略。大量的实验表明,我们的TokSing实现了更好的性能对梅尔频谱图基线,同时提供了中间表示空间成本和收敛速度的优势。摘要:Recent advancements in speech synthesis witness significant benefits by leveraging discrete tokens extracted from self-supervised learning (SSL) models. Discrete tokens offer higher storage efficiency and greater operability in intermediate representations compared to traditional continuous Mel spectrograms. However, when it comes to singing voice synthesis(SVS), achieving higher levels of melody expression poses a great challenge for utilizing discrete tokens. In this paper, we introduce TokSing, a discrete-based SVS system equipped with a token formulator that offers flexible token blendings. We observe a melody degradation during discretization, prompting us to integrate a melody signal with the discrete token and incorporate a specially-designed melody enhancement strategy in the musical encoder. Extensive experiments demonstrate that our TokSing achieves better performance against the Mel spectrogram baselines while offering advantages in intermediate representation space cost and convergence speed.
【3】 Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models
标题: 迪夫-A-里夫:通过潜在扩散模型的音乐伴奏共同创作
作者:Javier Nistal,Marco Pasini,Cyran Aouameur,Maarten Grachten,Stefan Lattner
备注:8 pages, 2 figures, 3 tables
链接:点击下载PDF文件
摘要:深度生成模型的最新进展为音乐制作带来了新的机会,但也带来了挑战,例如高计算需求和有限的音频质量。此外,当前的系统经常仅依赖于文本输入,并且通常专注于制作完整的音乐作品,这与音乐制作中的现有工作流程不兼容。为了解决这些问题,我们引入了“Diff-A-Riff”,这是一种潜在的扩散模型,旨在生成适用于任何音乐背景的高质量乐器演奏。该模型通过音频参考、文本提示或两者提供控制,并产生48 kHz伪立体声音频,同时显著减少推理时间和内存使用。我们通过客观指标和主观听力测试展示了模型的能力,并在相应的网站上提供了大量的例子:sonycslparis.github.io diffariff-companion 摘要:Recent advancements in deep generative models present new opportunities for music production but also pose challenges, such as high computational demands and limited audio quality. Moreover, current systems frequently rely solely on text input and typically focus on producing complete musical pieces, which is incompatible with existing workflows in music production. To address these issues, we introduce "Diff-A-Riff," a Latent Diffusion Model designed to generate high-quality instrumental accompaniments adaptable to any musical context. This model offers control through either audio references, text prompts, or both, and produces 48kHz pseudo-stereo audio while significantly reducing inference time and memory usage. We demonstrate the model's capabilities through objective metrics and subjective listening tests, with extensive examples available on the accompanying website: sonycslparis.github.io diffariff-companion
【4】 Towards Unsupervised Speech Recognition Without Pronunciation Models
标题: 迈向没有发音模型的无监督语音识别
作者:Junrui Ni,Liming Wang,Yang Zhang,Kaizhi Qian,Heting Gao,Mark Hasegawa-Johnson,Chang D. Yoo
备注:This work has been submitted to the IEEE for possible publication
链接:点击下载PDF文件
摘要:监督自动语音识别(ASR)的最新进展取得了显着的性能,主要是由于越来越多的大型转录语音语料库的可用性。然而,大多数语言缺乏足够的成对语音和文本数据来有效地训练这些系统。在这篇文章中,我们解决了开发ASR系统没有配对的语音和文本语料库的挑战,提出了消除对音素词典的依赖。我们探索了一个新的研究方向:词级无监督自动语音识别。使用一个只包含高频英语单词的精选语音语料库,我们的系统在没有平行成绩单或甲骨文单词边界的情况下实现了近20%的单词错误率。此外,我们的实验表明,一个无监督的语音识别器可以出现从联合语音到语音和文本到文本掩蔽标记填充。这种创新的模型超越了以前使用直接分布匹配训练的无监督ASR模型的性能。摘要:Recent advancements in supervised automatic speech recognition (ASR) have achieved remarkable performance, largely due to the growing availability of large transcribed speech corpora. However, most languages lack sufficient paired speech and text data to effectively train these systems. In this article, we tackle the challenge of developing ASR systems without paired speech and text corpora by proposing the removal of reliance on a phoneme lexicon. We explore a new research direction: word-level unsupervised ASR. Using a curated speech corpus containing only high-frequency English words, our system achieves a word error rate of nearly 20% without parallel transcripts or oracle word boundaries. Furthermore, we experimentally demonstrate that an unsupervised speech recognizer can emerge from joint speech-to-speech and text-to-text masked token-infilling. This innovative model surpasses the performance of previous unsupervised ASR models trained with direct distribution matching.
【5】 CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
标题: CoLM-SVR:利用神经编解码语言建模进行多模式发音障碍语音重建
作者:Xueyuan Chen,Dongchao Yang,Dingdong Wang,Xixin Wu,Zhiyong Wu,Helen Meng
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:构音障碍语音重建(DSR)的目的是将构音障碍语音转换为正常语音。但它仍然存在说话人相似度低、韵律自然度差等问题。在本文中,我们提出了一个多模态DSR模型,利用神经编解码器的语言建模,以改善重建结果,特别是对说话人的相似性和韵律自然。我们提出的模型包括:(i)一个多模态内容编码器,用于从构音障碍语音中提取具有辅助视觉输入的鲁棒音素嵌入;(ii)一个说话人编解码器编码器,用于从构音障碍语音中提取说话人感知的编解码器并对其进行归一化,以提供原始音色和正常韵律;(iii)基于编解码器语言模型的语音解码器,用于基于所提取的音素嵌入和归一化编解码器来重构语音。在UAS语音语料库上的测试结果表明,该模型在说话人相似度和韵律自然度方面都有明显的提高。摘要:Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech. It still suffers from low speaker similarity and poor prosody naturalness. In this paper, we propose a multi-modal DSR model by leveraging neural codec language modeling to improve the reconstruction results, especially for the speaker similarity and prosody naturalness. Our proposed model consists of: (i) a multi-modal content encoder to extract robust phoneme embeddings from dysarthric speech with auxiliary visual inputs; (ii) a speaker codec encoder to extract and normalize the speaker-aware codecs from the dysarthric speech, in order to provide original timbre and normal prosody; (iii) a codec language model based speech decoder to reconstruct the speech based on the extracted phoneme embeddings and normalized codecs. Evaluations on the commonly used UASpeech corpus show that our proposed model can achieve significant improvements in terms of speaker similarity and prosody naturalness.
【6】 Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
标题: 在说话人嵌入中使用对抗扰动的非同步语音
作者:Rui Wang,Liping Chen,Kong AiK Lee,Zhen-Hua Ling
链接:点击下载PDF文件
摘要:语音匿名化已经被开发为用于通过用伪说话者的语音替换语音信号中的说话者的语音来保护隐私的技术,从而使原始语音属性从机器识别和人类感知中模糊。在本文中,我们专注于改变机器识别的语音属性,同时保留人类的感知。我们称之为异步语音匿名化。为此,语音生成框架结合扬声器解纠缠机制来生成匿名语音。通过对说话人嵌入施加对抗性扰动来改变说话人属性,同时通过控制扰动的强度来保留人类感知。在LibriSpeech数据集上进行的实验表明,说话人的属性被模糊,60.71%的处理后的话语保留了人类的感知。摘要:Voice anonymization has been developed as a technique for preserving privacy by replacing the speaker's voice in a speech signal with that of a pseudo-speaker, thereby obscuring the original voice attributes from machine recognition and human perception. In this paper, we focus on altering the voice attributes against machine recognition while retaining human perception. We referred to this as the asynchronous voice anonymization. To this end, a speech generation framework incorporating a speaker disentanglement mechanism is employed to generate the anonymized speech. The speaker attributes are altered through adversarial perturbation applied on the speaker embedding, while human perception is preserved by controlling the intensity of perturbation. Experiments conducted on the LibriSpeech dataset showed that the speaker attributes were obscured with their human perception preserved for 60.71% of the processed utterances.
【7】 FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
标题: FreeV:通过伪反梅尔过滤器为声码者提供免费午餐
作者:Yuanjun Lv,Hai Li,Ying Yan,Junhui Liu,Danming Xie,Lei Xie
备注:Accepted by InterSpeech 2024; 5 pages, 5 figures
链接:点击下载PDF文件
摘要:声码器从语音的声学特征中重构出语音波形,在现代文语转换系统中起着举足轻重的作用。频域GAN声码器(如Vocos和APNet2)最近取得了快速发展,在推理速度方面优于时域模型,同时实现了相当的音频质量。然而,这些频域声码器遭受大的参数大小,从而引入额外的存储器负担。受PriorGrad和SpecGrad的启发,我们采用伪逆来粗略估计振幅谱作为初始值。这种简单的初始化显著地减轻了对声码器的参数要求。基于APNet2和我们精简的幅度预测分支,我们提出了我们的FreeV,与其对应的APNet2相比,我们的FreeV在接近一半的参数下实现了1.8倍的推理速度提高。同时,我们的FreeV在再合成质量方面优于APNet2,标志着在追求实时,高保真语音合成方面向前迈出了一步。代码和检查点可在https: github.com BakerBunker FreeV上获得摘要:Vocoders reconstruct speech waveforms from acoustic features and play a pivotal role in modern TTS systems. Frequent-domain GAN vocoders like Vocos and APNet2 have recently seen rapid advancements, outperforming time-domain models in inference speed while achieving comparable audio quality. However, these frequency-domain vocoders suffer from large parameter sizes, thus introducing extra memory burden. Inspired by PriorGrad and SpecGrad, we employ pseudo-inverse to estimate the amplitude spectrum as the initialization roughly. This simple initialization significantly mitigates the parameter demand for vocoder. Based on APNet2 and our streamlined Amplitude prediction branch, we propose our FreeV, compared with its counterpart APNet2, our FreeV achieves 1.8 times inference speed improvement with nearly half parameters. Meanwhile, our FreeV outperforms APNet2 in resynthesis quality, marking a step forward in pursuing real-time, high-fidelity speech synthesis. Code and checkpoints is available at: https: github.com BakerBunker FreeV
【8】 Codecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio
标题: Codecfake:用于检测基于LLM的Deepfake音频的初始数据集
作者:Yi Lu,Yuankun Xie,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Zhiyong Wang,Xin Qi,Xuefei Liu,Yongwei Li,Yukun Liu,Xiaopeng Wang,Shuchen Shi
备注:Accepted by INTERSPEECH 2024. arXiv admin note: substantial text overlap with arXiv:2405.04880
链接:点击下载PDF文件
摘要:随着基于大语言模型(LLM)的deepfake音频的激增,迫切需要有效的检测方法。以前的deepfake音频生成方法通常涉及多步生成过程,最后一步使用声码器从手工特征预测波形。然而,基于LLM的音频在端到端生成过程中直接从离散神经编解码器生成,跳过了声码器处理的最后一步。这对当前基于声码器伪影的音频深度伪造检测(ADD)模型提出了重大挑战。为了有效地检测基于LLM的deepfake音频,我们专注于生成过程的核心,从神经编解码器到波形的转换。我们提出了Codecfake数据集,它是由七种代表性的神经编解码方法生成的。实验结果表明,与声码器训练的ADD模型相比,编解码器训练的ADD模型在Codecfake测试集上的平均等误率降低了41.406%。摘要:With the proliferation of Large Language Model (LLM) based deepfake audio, there is an urgent need for effective detection methods. Previous deepfake audio generation methods typically involve a multi-step generation process, with the final step using a vocoder to predict the waveform from handcrafted features. However, LLM-based audio is directly generated from discrete neural codecs in an end-to-end generation process, skipping the final step of vocoder processing. This poses a significant challenge for current audio deepfake detection (ADD) models based on vocoder artifacts. To effectively detect LLM-based deepfake audio, we focus on the core of the generation process, the conversion from neural codec to waveform. We propose Codecfake dataset, which is generated by seven representative neural codec methods. Experiment results show that codec-trained ADD models exhibit a 41.406% reduction in average equal error rate compared to vocoder-trained ADD models on the Codecfake test set.
【9】 FakeSound: Deepfake General Audio Detection
标题: FakeSound:Deepfake通用音频检测
作者:Zeyu Xie,Baihan Li,Xuenan Xu,Zheng Liang,Kai Yu,Mengyue Wu
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:随着音频生成技术的进步,生成模型可以产生高度逼真的音频。然而,deepfake一般音频的扩散可能会造成负面后果。因此,我们提出了一个新的任务,deepfake通用音频检测,旨在识别音频内容是否被操纵并定位deepfake区域。利用自动操作管道,提出了一个名为FakeSound的数据集,用于deepfake通用音频检测,样本可以在网站https: FakeSoundData.github.io上查看。人类在所有测试集上的平均二进制准确度始终低于0.6,这表明人类在识别deepfake音频时面临的困难,并肯定了FakeSound数据集的有效性。提出了一种利用通用音频预训练模型的深度伪造检测模型作为基准系统。实验结果表明,所提出的模型的性能超过了最先进的deepfake语音检测和人类测试。摘要:With the advancement of audio generation, generative models can produce highly realistic audios. However, the proliferation of deepfake general audio can pose negative consequences. Therefore, we propose a new task, deepfake general audio detection, which aims to identify whether audio content is manipulated and to locate deepfake regions. Leveraging an automated manipulation pipeline, a dataset named FakeSound for deepfake general audio detection is proposed, and samples can be viewed on website https: FakeSoundData.github.io. The average binary accuracy of humans on all test sets is consistently below 0.6, which indicates the difficulty humans face in discerning deepfake audio and affirms the efficacy of the FakeSound dataset. A deepfake detection model utilizing a general audio pre-trained model is proposed as a benchmark system. Experimental results demonstrate that the performance of the proposed model surpasses the state-of-the-art in deepfake speech detection and human testers.
【10】 CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
标题: 用于流媒体开放词汇关键词查找的符合ATC的音频文本嵌入
作者:Sichen Jin,Youngmoon Jung,Seungjin Lee,Jaeyoung Roh,Changwoo Han,Hoonyoung Cho
链接:点击下载PDF文件
摘要:本文介绍了一种新的方法,流开放词汇关键字发现(KWS)与基于文本的关键字注册。对于每个输入帧,所提出的方法使用连接主义时间分类(CTC)找到在帧处结束的最佳对准,并聚合帧级声学嵌入(AE)以获得更高级别(即,字符、单词或短语)AE,其与目标关键字文本的文本嵌入(TE)对齐。然后,我们计算聚集的AE和TE的相似度。据我们所知,这是第一次尝试动态对齐音频和关键字文本,以实现KWS的联合音频-文本嵌入。尽管以流式方式操作,但与仅具有155 K模型参数和时间复杂度为O(U)的解码算法的非流式方法相比,我们的方法在LibriPhrase数据集上实现了具有竞争力的性能,其中U是推理时目标关键字的长度。摘要:This paper introduces a novel approach for streaming openvocabulary keyword spotting (KWS) with text-based keyword enrollment. For every input frame, the proposed method finds the optimal alignment ending at the frame using connectionist temporal classification (CTC) and aggregates the frame-level acoustic embedding (AE) to obtain higher-level (i.e., character, word, or phrase) AE that aligns with the text embedding (TE) of the target keyword text. After that, we calculate the similarity of the aggregated AE and the TE. To the best of our knowledge, this is the first attempt to dynamically align the audio and the keyword text on-the-fly to attain the joint audio-text embedding for KWS. Despite operating in a streaming fashion, our approach achieves competitive performance on the LibriPhrase dataset compared to the non-streaming methods with a mere 155K model parameters and a decoding algorithm with time complexity O(U), where U is the length of the target keyword at inference time.
【11】 Can Large Language Models Understand Spatial Audio?
标题: 大型语言模型能理解空间音频吗?
作者:Changli Tang,Wenyi Yu,Guangzhi Sun,Xianzhao Chen,Tian Tan,Wei Li,Jun Zhang,Lu Lu,Zejun Ma,Yuxuan Wang,Chao Zhang
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:本文探讨了使大型语言模型(LLM)能够从多通道音频中理解空间信息,这是目前听觉LLM缺乏的技能。通过利用LLM先进的认知和推理能力,目的是通过音频增强对3D环境的理解。我们研究了3个空间音频任务:声源定位(SSL),远场语音识别(FSR)和定位通知语音提取(LSE),在每个任务中取得了显着的进展。对于SSL,我们的方法在Spatial LibriSpeech数据集上实现了2.70 ^{ circ}$的MAE,大大超过了之前的基准约6.60 ^{ circ}$。此外,我们的模型可以采用空间线索来提高FSR的准确性,并通过文本提示选择性地关注来自指定方向的声音来执行LSE,即使是在重叠的语音中。这些发现突出了适应LLM以掌握物理音频概念的潜力,为3D环境中基于LLM的代理铺平了道路。摘要:This paper explores enabling large language models (LLMs) to understand spatial information from multichannel audio, a skill currently lacking in auditory LLMs. By leveraging LLMs' advanced cognitive and inferential abilities, the aim is to enhance understanding of 3D environments via audio. We study 3 spatial audio tasks: sound source localization (SSL), far-field speech recognition (FSR), and localisation-informed speech extraction (LSE), achieving notable progress in each task. For SSL, our approach achieves an MAE of $2.70^{ circ}$ on the Spatial LibriSpeech dataset, substantially surpassing the prior benchmark of about $6.60^{ circ}$. Moreover, our model can employ spatial cues to improve FSR accuracy and execute LSE by selectively attending to sounds originating from a specified direction via text prompts, even amidst overlapping speech. These findings highlight the potential of adapting LLMs to grasp physical audio concepts, paving the way for LLM-based agents in 3D environments.
【12】 Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
标题: 探索自我监督多视图对比学习用于有限注释的语音情感识别
作者:Bulat Khaertdinov,Pedro Jeuris,Annanda Sousa,Enrique Hortal
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:深度和自我监督学习(SSL)的最新进展使语音情感识别(SER)性能得到了大幅改善,达到了前所未有的水平。然而,获得足够数量的准确标记的数据来训练或微调模型仍然是一项昂贵且具有挑战性的任务。在本文中,我们提出了一种多视图SSL预训练技术,可应用于各种语音表示,包括由大型语音模型生成的语音表示,以提高在注释有限的情况下的SER性能。我们的实验,基于wav2vec 2.0,频谱和语言特征,表明所提出的框架提高SER性能,在未加权平均召回率高达10%,在设置非常稀疏的数据注释。摘要:Recent advancements in Deep and Self-Supervised Learning (SSL) have led to substantial improvements in Speech Emotion Recognition (SER) performance, reaching unprecedented levels. However, obtaining sufficient amounts of accurately labeled data for training or fine-tuning the models remains a costly and challenging task. In this paper, we propose a multi-view SSL pre-training technique that can be applied to various representations of speech, including the ones generated by large speech models, to improve SER performance in scenarios where annotations are limited. Our experiments, based on wav2vec 2.0, spectral and paralinguistic features, demonstrate that the proposed framework boosts the SER performance, by up to 10% in Unweighted Average Recall, in settings with extremely sparse data annotations.
【13】 Flexible Music-Conditioned Dance Generation with Style Description Prompts
标题: 具有风格描述的灵活音乐条件舞蹈生成
作者:Hongsong Wang,Yin Zhu,Xin Geng
链接:点击下载PDF文件
摘要:舞蹈作为一种艺术形式和表现形式,在人类文化中占有重要地位,但舞蹈的创作仍然是一项具有挑战性的任务。大多数舞蹈生成方法主要依赖于音乐,很少考虑音乐风格或流派等内在属性。在这项工作中,我们介绍了灵活的舞蹈生成与风格描述符(DGSDP),一个基于扩散的框架,适合于多样化的舞蹈生成任务,充分利用音乐风格的语义。该框架的核心组件是音乐条件风格感知扩散(MCSAD),它包括一个基于transformer的网络和一个音乐风格调制模块。MCSAD将音乐条件和风格描述提示巧妙地集成到舞蹈生成框架中,确保生成的舞蹈与音乐内容和风格一致。为了便于灵活的舞蹈生成和适应不同的任务,时空掩蔽策略有效地应用在向后扩散过程中。所提出的框架成功地生成逼真的舞蹈序列,准确地与音乐的各种任务,如长期一代,舞蹈中间,舞蹈修补等,我们希望这项工作有可能激发舞蹈的生成和创作,在娱乐,艺术和教育的应用前景。摘要:Dance plays an important role as an artistic form and expression in human culture, yet the creation of dance remains a challenging task. Most dance generation methods primarily rely solely on music, seldom taking into consideration intrinsic attributes such as music style or genre. In this work, we introduce Flexible Dance Generation with Style Description Prompts (DGSDP), a diffusion-based framework suitable for diversified tasks of dance generation by fully leveraging the semantics of music style. The core component of this framework is Music-Conditioned Style-Aware Diffusion (MCSAD), which comprises a Transformer-based network and a music Style Modulation module. The MCSAD seemly integrates music conditions and style description prompts into the dance generation framework, ensuring that generated dances are consistent with the music content and style. To facilitate flexible dance generation and accommodate different tasks, a spatial-temporal masking strategy is effectively applied in the backward diffusion process. The proposed framework successfully generates realistic dance sequences that are accurately aligned with music for a variety of tasks such as long-term generation, dance in-betweening, dance inpainting, and etc. We hope that this work has the potential to inspire dance generation and creation, with promising applications in entertainment, art, and education.
【14】 VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
标题: WAL-E R:通过单调对齐实现稳健高效的Zero-Shot文本到语音合成
作者:Bing Han,Long Zhou,Shujie Liu,Sanyuan Chen,Lingwei Meng,Yanming Qian,Yanqing Liu,Sheng Zhao,Jinyu Li,Furu Wei
备注:15 pages, 5 figures
链接:点击下载PDF文件
摘要:在离散神经音频编解码器的帮助下,大语言模型(LLM)越来越被认为是一种有前途的zero-shot文本到语音(TTS)合成方法。然而,基于采样的解码策略给生成带来了惊人的多样性,但也带来了诸如错别字、遗漏和重复的鲁棒性问题。此外,音频的高采样率也给自回归的推理过程带来了巨大的计算开销。为了解决这些问题,我们提出了VALL-E R,这是一个强大且高效的零射击TTS系统,以VALL-E为基础。zero-shot TTS系统。具体来说,我们引入了一个音素单调对齐策略,以加强音素和声学序列之间的连接,确保更精确的对齐,通过约束声学令牌,以匹配其相关的音素。此外,我们采用了一种编解码器合并的方法来下采样的离散代码在浅量化层,从而加快解码速度,同时保持高质量的语音输出。得益于这些策略,VALL-E R获得了音素上的相似性,并通过接近地面真值的WER来展示其强大的鲁棒性。此外,它需要更少的自回归步骤,在推理过程中减少了60%以上的时间。这项研究有可能应用于有意义的项目,包括为失语症患者创造语言。音频样本将在https: aka.ms valler上提供。摘要:With the help of discrete neural audio codecs, large language models (LLM) have increasingly been recognized as a promising methodology for zero-shot Text-to-Speech (TTS) synthesis. However, sampling based decoding strategies bring astonishing diversity to generation, but also pose robustness issues such as typos, omissions and repetition. In addition, the high sampling rate of audio also brings huge computational overhead to the inference process of autoregression. To address these issues, we propose VALL-E R, a robust and efficient zero-shot TTS system, building upon the foundation of VALL-E. Specifically, we introduce a phoneme monotonic alignment strategy to strengthen the connection between phonemes and acoustic sequence, ensuring a more precise alignment by constraining the acoustic tokens to match their associated phonemes. Furthermore, we employ a codec-merging approach to downsample the discrete codes in shallow quantization layer, thereby accelerating the decoding speed while preserving the high quality of speech output. Benefiting from these strategies, VALL-E R obtains controllablity over phonemes and demonstrates its strong robustness by approaching the WER of ground truth. In addition, it requires fewer autoregressive steps, with over 60% time reduction during inference. This research has the potential to be applied to meaningful projects, including the creation of speech for those affected by aphasia. Audio samples will be available at: https: aka.ms valler.
【15】 Zero-Shot Fake Video Detection by Audio-Visual Consistency
标题: 利用视听一致性进行Zero-Shot假视频检测
作者:Xiaolou Li,Zehua Liu,Chen Chen,Lantian Li,Li Guo,Dong Wang
备注:to be published in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:最近的研究主张检测假视频作为一类检测任务,基于真实数据的音频和视觉模态之间的一致性比假数据的一致性更重要的假设。这种方法仅依赖于真实的视听数据,同时不需要伪造的对应数据,因此被描述为“零射击”(zero-shot)检测范式。提出了一种基于音视频内容一致性的zero-shot检测方法。通过使用预先训练的ASR和VSR模型,我们分别识别音频和视频内容序列。然后,计算两个序列之间的编辑距离,以评估声称的视频是否是真实的。实验结果表明,与基于语义一致性和时间一致性的两种主流方法相比,我们的方法在各种deepfake技术中实现了卓越的泛化能力,并对视听干扰表现出较强的鲁棒性。最后,通过简单地整合这三个系统的决策分数,可以实现最先进的性能增益。摘要:Recent studies have advocated the detection of fake videos as a one-class detection task, predicated on the hypothesis that the consistency between audio and visual modalities of genuine data is more significant than that of fake data. This methodology, which solely relies on genuine audio-visual data while negating the need for forged counterparts, is thus delineated as a zero-shot' detection paradigm. This paper introduces a novel zero-shot detection approach anchored in content consistency across audio and video. By employing pre-trained ASR and VSR models, we recognize the audio and video content sequences, respectively. Then, the edit distance between the two sequences is computed to assess whether the claimed video is genuine. Experimental results indicate that, compared to two mainstream approaches based on semantic consistency and temporal consistency, our approach achieves superior generalizability across various deepfake techniques and demonstrates strong robustness against audio-visual perturbations. Finally, state-of-the-art performance gains can be achieved by simply integrating the decision scores of these three systems.
【16】 SEBN Adapter: Parametric Efficient Domain Adaptation for Speaker Recognition
标题: SEBN适配器:用于说话人识别的参数高效域自适应
作者:Tianhao Wang,Lantian Li,Dong Wang
备注:to be published in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:在一个新的领域部署一个经过良好优化的预先训练的说话人识别模型通常会导致性能显著下降。虽然微调是一种常用的解决方案,但它需要大量的自适应数据,并且参数效率低下,这使得它对于具有有限数据可用于模型自适应的现实世界应用来说是不切实际的。从自监督预训练模型中适配器的成功中汲取灵感,本文介绍了SE BN适配器来解决这一挑战。通过冻结核心扬声器编码器和调整特征映射的权重和激活分布,我们引入了一种新的适配器,利用可训练的挤压和激励(SE)块和批量归一化(BN)层,称为SE BN适配器。我们的实验使用VoxCeleb进行预训练,使用CN-Celeb的4种类型进行适应,表明SE BN适配器在基线上提供了显着的性能改进,并通过调整仅1%的参数与香草微调方法竞争。摘要:Deploying a well-optimized pre-trained speaker recognition model in a new domain often leads to a significant decline in performance. While fine-tuning is a commonly employed solution, it demands ample adaptation data and suffers from parameter inefficiency, rendering it impractical for real-world applications with limited data available for model adaptation. Drawing inspiration from the success of adapters in self-supervised pre-trained models, this paper introduces a SE BN adapter to address this challenge. By freezing the core speaker encoder and adjusting the feature maps' weights and activation distributions, we introduce a novel adapter utilizing trainable squeeze-and-excitation (SE) blocks and batch normalization (BN) layers, termed SE BN adapter. Our experiments, conducted using VoxCeleb for pre-training and 4 genres from CN-Celeb for adaptation, demonstrate that the SE BN adapter offers significant performance improvement over the baseline and competes with the vanilla fine-tuning approach by tuning just 1% of the parameters.
【17】 PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
标题: PRoDeliberation:端到端口语理解的并行稳健审议
作者:Trang Le,Daniel Lazar,Suyoun Kim,Shan Jiang,Duc Le,Adithya Sagar,Aleksandr Livshits,Ahmed Aly,Akshat Shrivastava
链接:点击下载PDF文件
摘要:口语理解(SLU)是语音助手的关键组件;它包括将语音转换为语义解析以执行任务。以前的工作已经探索了端到端模型,以提高SLU模型的质量和鲁棒性,但这些模型仍然是自回归的,导致更高的延迟。在这项工作中,我们介绍PRoDeliberation,一种新的方法,利用连接时间分类为基础的解码策略,以及去噪目标训练强大的非自回归审议模型。我们表明,PRoDeliberation实现了并行解码的延迟降低(自回归模型的2- 10倍改进),同时保留了纠正自回归审议系统的自动语音识别(ASR)误译的能力。我们进一步表明,去噪训练的设计使PRoDeliberation能够克服小型ASR设备的局限性,并且我们对系统每个组件的必要性进行了分析。摘要:Spoken Language Understanding (SLU) is a critical component of voice assistants; it consists of converting speech to semantic parses for task execution. Previous works have explored end-to-end models to improve the quality and robustness of SLU models with Deliberation, however these models have remained autoregressive, resulting in higher latencies. In this work we introduce PRoDeliberation, a novel method leveraging a Connectionist Temporal Classification-based decoding strategy as well as a denoising objective to train robust non-autoregressive deliberation models. We show that PRoDeliberation achieves the latency reduction of parallel decoding (2-10x improvement over autoregressive models) while retaining the ability to correct Automatic Speech Recognition (ASR) mistranscriptions of autoregressive deliberation systems. We further show that the design of the denoising training allows PRoDeliberation to overcome the limitations of small ASR devices, and we provide analysis on the necessity of each component of the system.
【18】 EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
标题: SEARCH Sphere-TTC:通过球形情感载体进行情感风格和强度建模,用于可控情感文本到语音
作者:Deok-Hyeon Cho,Hyung-Seok Oh,Seung-Bin Kim,Sang-Hoon Lee,Seong-Whan Lee
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:尽管在情感文本到语音(TTS)领域取得了快速进展,但最近的研究主要集中在模仿特定情感的平均风格上。因此,操纵语音情感的能力仍然被限制在几个预定义的标签上,从而影响了反映情感细微变化的能力。在本文中,我们提出了一种新的语音合成方法--球面TTS,它通过使用一个球形情感向量来控制合成语音的情感风格和强度,从而合成具有表达性的情感语音。在没有任何人类注释的情况下,我们使用唤醒、效价和支配伪标签通过笛卡尔球面变换来模拟情感的复杂性质。此外,我们提出了一个双条件对抗网络,以提高生成的语音质量,反映多方面的特点。实验结果表明,该模型能够控制情绪的风格和强度与高质量的表达语音。摘要:Despite rapid advances in the field of emotional text-to-speech (TTS), recent studies primarily focus on mimicking the average style of a particular emotion. As a result, the ability to manipulate speech emotion remains constrained to several predefined labels, compromising the ability to reflect the nuanced variations of emotion. In this paper, we propose EmoSphere-TTS, which synthesizes expressive emotional speech by using a spherical emotion vector to control the emotional style and intensity of the synthetic speech. Without any human annotation, we use the arousal, valence, and dominance pseudo-labels to model the complex nature of emotion via a Cartesian-spherical transformation. Furthermore, we propose a dual conditional adversarial network to improve the quality of generated speech by reflecting the multi-aspect characteristics. The experimental results demonstrate the model ability to control emotional style and intensity with high-quality expressive speech.
【19】 PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
标题: PolySpeech:探索统一的多任务语音模型,以与单任务模型竞争
作者:Runyan Yang,Huibao Yang,Xiqing Zhang,Tiantian Ye,Ying Liu,Yingying Gao,Shilei Zhang,Chao Deng,Junlan Feng
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:最近,已经尝试将各种语音处理任务集成到统一的模型中。然而,很少有以前的工作直接表明,联合优化的不同任务的多任务语音模型有积极的影响,个别任务的性能。本文提出了一个多任务语音模型PolySpeech,它支持语音识别、语音合成和两个语音分类任务。PolySpeech以多模态语言模型为核心结构,以语义表征为语音输入。我们将语义语音嵌入标记化和语音重建方法引入PolySpeech,从而为任何给定的说话者有效地生成高质量的语音。与单任务模型相比,PolySpeech在各种任务中表现出竞争力。在我们的实验中,多任务优化实现了与单任务优化相当的性能,并且对特定任务特别有益。摘要:Recently, there have been attempts to integrate various speech processing tasks into a unified model. However, few previous works directly demonstrated that joint optimization of diverse tasks in multitask speech models has positive influence on the performance of individual tasks. In this paper we present a multitask speech model -- PolySpeech, which supports speech recognition, speech synthesis, and two speech classification tasks. PolySpeech takes multi-modal language model as its core structure and uses semantic representations as speech inputs. We introduce semantic speech embedding tokenization and speech reconstruction methods to PolySpeech, enabling efficient generation of high-quality speech for any given speaker. PolySpeech shows competitiveness across various tasks compared to single-task models. In our experiments, multitask optimization achieves performance comparable to single-task optimization and is especially beneficial for specific tasks.
【20】 The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
标题: Interspeech 2024年使用离散单元的语音处理挑战
作者:Xuankai Chang,Jiatong Shi,Jinchuan Tian,Yuning Wu,Yuxun Tang,Yihan Wu,Shinji Watanabe,Yossi Adi,Xie Chen,Qin Jin
备注:This manuscript has been accepted by Interspeech2024
链接:点击下载PDF文件
摘要:以离散单元表示语音和音频信号已经成为传统高维特征向量的一种引人注目的替代方案。许多研究已经强调了离散单元在各种应用中的功效,例如语音压缩和恢复,语音识别和语音生成。为了促进这一领域的探索,我们引入了Interspeech 2024挑战赛,该挑战赛专注于使用离散单元的新语音处理基准。它包括三个关键的任务,即多语种自动语音识别,文本到语音,和歌声合成,并旨在评估这些任务中的离散单元的潜在适用性。本文概述了挑战设计和基线描述。我们还整理了基线和选定的提交系统,以及初步研究结果,为这个不断发展的领域的未来研究提供了宝贵的贡献。摘要:Representing speech and audio signals in discrete units has become a compelling alternative to traditional high-dimensional feature vectors. Numerous studies have highlighted the efficacy of discrete units in various applications such as speech compression and restoration, speech recognition, and speech generation. To foster exploration in this domain, we introduce the Interspeech 2024 Challenge, which focuses on new speech processing benchmarks using discrete units. It encompasses three pivotal tasks, namely multilingual automatic speech recognition, text-to-speech, and singing voice synthesis, and aims to assess the potential applicability of discrete units in these tasks. This paper outlines the challenge designs and baseline descriptions. We also collate baseline and selected submission systems, along with preliminary findings, offering valuable contributions to future research in this evolving field.
【21】 FastAST: Accelerating Audio Spectrogram Transformer via Token Merging and Cross-Model Knowledge Distillation
标题: FastAST:通过令牌合并和跨模型知识提炼加速音频频谱图Transformer
作者:Swarup Ranjan Behera,Abhishek Dhiman,Karthik Gowda,Aalekhya Satya Narayani
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:音频分类模型,特别是音频频谱图Transformer(AST),在高效的音频分析中起着至关重要的作用。然而,在不影响精度的情况下优化其效率仍然是一个挑战。在本文中,我们介绍FastAST,一个框架,集成令牌合并(ToMe)到AST框架。FastAST通过合并音频频谱图中的相似标记来提高推理速度,而无需进行大量的重新训练。此外,在训练过程中,FastAST带来了显着的速度提高。实验表明,FastAST可以提高音频分类吞吐量,同时对准确性的影响最小。为了减轻准确性的影响,我们将跨模型知识蒸馏(CMKD)集成到FastAST框架中。与AST相比,将ToMe和CMKD集成到AST中可以提高准确性,同时保持更快的推理速度。FastAST代表了向实时、资源高效的音频分析迈出的一步。摘要:Audio classification models, particularly the Audio Spectrogram Transformer (AST), play a crucial role in efficient audio analysis. However, optimizing their efficiency without compromising accuracy remains a challenge. In this paper, we introduce FastAST, a framework that integrates Token Merging (ToMe) into the AST framework. FastAST enhances inference speed without requiring extensive retraining by merging similar tokens in audio spectrograms. Furthermore, during training, FastAST brings about significant speed improvements. The experiments indicate that FastAST can increase audio classification throughput with minimal impact on accuracy. To mitigate the accuracy impact, we integrate Cross-Model Knowledge Distillation (CMKD) into the FastAST framework. Integrating ToMe and CMKD into AST results in improved accuracy compared to AST while maintaining faster inference speeds. FastAST represents a step towards real-time, resource-efficient audio analysis.
【22】 Broadband MEMS Microphone Arrays with Reduced Aperture Through 3D-Printed Waveguides
标题: 通过3D打印光路实现缩小口径的宽带微机电麦克风阵列
作者:Dennis Laurijssen,Walter Daems,Jan Steckel
链接:点击下载PDF文件
摘要:在本文中,我们提出了一种无源和成本效益的方法,用于增加超声MEMS麦克风阵列的频率范围时,使用波束形成技术。通过应用减小MEMS麦克风的声孔径的3D打印结构,我们可以创建规则间隔的麦克风阵列布局,由于MEMS元件的物理尺寸,元件间间距比印刷电路板上实现的要小得多。该方法允许结合麦克风阵列的超声传感器与波束成形技术的组合使用,而不会由于诸如声源定位或蝙蝠HRTF的仿真的应用中的栅瓣而产生混叠。摘要:In this paper we present a passive and cost-effective method for increasing the frequency range of ultrasound MEMS microphone arrays when using beamforming techniques. By applying a 3D-printed construction that reduces the acoustic aperture of the MEMS microphones we can create a regularly spaced microphone array layout with much smaller inter-element spacing than could be accomplished on a printed circuit board due to the physical size of the MEMS elements. This method allows the use of ultrasound sensors incorporating microphone arrays in combination with beamforming techniques without aliases due to grating lobes in applications such as sound source localization or the emulation of bat HRTFs.
【23】 Pre-training Feature Guided Diffusion Model for Speech Enhancement
标题: 用于语音增强的预训练特征引导扩散模型
作者:Yiyuan Yang,Niki Trigoni,Andrew Markham
备注:Accepted by Interspeech 2024 Conference
链接:点击下载PDF文件
摘要:语音增强显著提高了嘈杂环境中语音的清晰度和可懂度,改善了沟通和聆听体验。在本文中,我们介绍了一种新的预训练特征引导的扩散模型为有效的语音增强量身定制,解决现有的歧视和生成模型的局限性。通过将频谱特征集成到变分自动编码器(VAE)中,并在反向过程中利用预先训练的特征进行指导,再加上利用确定性离散积分方法(DDIM)来简化采样步骤,我们的模型提高了效率和语音增强质量。在两个具有不同SNR的公共数据集上展示了最先进的结果,我们的模型在效率和鲁棒性方面优于其他基线。该方法不仅优化了性能,而且提高了实际部署能力,而不增加计算需求。摘要:Speech enhancement significantly improves the clarity and intelligibility of speech in noisy environments, improving communication and listening experiences. In this paper, we introduce a novel pretraining feature-guided diffusion model tailored for efficient speech enhancement, addressing the limitations of existing discriminative and generative models. By integrating spectral features into a variational autoencoder (VAE) and leveraging pre-trained features for guidance during the reverse process, coupled with the utilization of the deterministic discrete integration method (DDIM) to streamline sampling steps, our model improves efficiency and speech enhancement quality. Demonstrating state-of-the-art results on two public datasets with different SNRs, our model outshines other baselines in efficiency and robustness. The proposed method not only optimizes performance but also enhances practical deployment capabilities, without increasing computational demands.
【24】 SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
标题: SVSNet+:利用语音基础模型的表示增强说话者语音相似性评估模型
作者:Chun Yin,Tai-Shih Chi,Yu Tsao,Hsin-Min Wang
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:来自预先训练的语音基础模型(SFM)的表示在许多下游任务中表现出令人印象深刻的性能。然而,将预先训练的SFM表示到扬声器语音相似性评估的潜在好处还没有得到彻底的研究。在本文中,我们提出了SVSNet+,这是一种集成了预训练的SFM表示的模型,可以提高评估说话者语音相似性的性能。在语音转换挑战2018和2020数据集上的实验结果表明,与基线模型相比,SVSNet+结合WavLM表示显示出显着的改进。此外,虽然用下游任务的小数据集微调WavLM不会提高性能,但使用相同的数据集来学习WavLM的加权和表示可以大大提高性能。此外,当WavLM被其他SFM取代时,SVSNet+仍然优于基线模型,并表现出很强的泛化能力。摘要:Representations from pre-trained speech foundation models (SFMs) have shown impressive performance in many downstream tasks. However, the potential benefits of incorporating pre-trained SFM representations into speaker voice similarity assessment have not been thoroughly investigated. In this paper, we propose SVSNet+, a model that integrates pre-trained SFM representations to improve performance in assessing speaker voice similarity. Experimental results on the Voice Conversion Challenge 2018 and 2020 datasets show that SVSNet+ incorporating WavLM representations shows significant improvements compared to baseline models. In addition, while fine-tuning WavLM with a small dataset of the downstream task does not improve performance, using the same dataset to learn a weighted-sum representation of WavLM can substantially improve performance. Furthermore, when WavLM is replaced by other SFMs, SVSNet+ still outperforms the baseline models and exhibits strong generalization ability.
【25】 Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
标题: 理解声音,错过问题:大型音频语言模型中对象幻觉的挑战
作者:Chun-Yi Kuan,Wei-Ping Huang,Hung-yi Lee
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:大型音频语言模型(LALM)通过集成音频感知功能来增强传统的大型语言模型,使它们能够解决音频相关的任务。之前的研究主要集中在评估LALM在各种任务中的表现,但忽视了它们的可靠性,特别是关于物体幻觉等问题。在我们的研究中,我们介绍了评估公开可用的LALM的物体幻觉程度的方法。我们的研究结果表明,LALM在对音频内容的理解方面与专门的音频字幕模型相当,但很难回答区分性问题,特别是那些需要识别音频片段中特定对象声音存在的问题。这种限制突出了当前LALM的一个关键弱点:它们对歧视性查询的理解不足。此外,我们探讨了潜在的提示工程,以提高LALM的性能上的歧视性问题。摘要:Large audio-language models (LALMs) enhance traditional large language models by integrating audio perception capabilities, allowing them to tackle audio-related tasks. Previous research has primarily focused on assessing the performance of LALMs across various tasks, yet overlooking their reliability, particularly concerning issues like object hallucination. In our study, we introduce methods to assess the extent of object hallucination of publicly available LALMs. Our findings reveal that LALMs are comparable to specialized audio captioning models in their understanding of audio content, but struggle to answer discriminative questions, specifically those requiring the identification of the presence of particular object sounds within an audio clip. This limitation highlights a critical weakness in current LALMs: their inadequate understanding of discriminative queries. Moreover, we explore the potential of prompt engineering to enhance LALMs' performance on discriminative questions.
【26】 SCDNet: Self-supervised Learning Feature-based Speaker Change Detection
标题: SCDNet:基于自我监督学习的说话者变化检测
作者:Yue Li,Xinsheng Wang,Li Zhang,Lei Xie
链接:点击下载PDF文件
摘要:说话人变化检测(SCD)是识别会话中说话人之间的边界。出于对SCD任务的wav2vec 2.0模型进行微调的成功,在这项工作中对SCD的自监督学习(SSL)功能进行了进一步的研究。具体而言,SCD模型,命名为SCDNet,提出。在此基础上,研究了Hubert、wav2vec 2.0和WavLm等多种SSL模型。为了识别最有效的层的SSL模型SCD,学习加权方法来分析的有效性的中间表示。此外,还实现了一种基于微调的方法,以进一步比较SCD任务中SSL模型的特性。此外,提出了一种对比学习方法,以减轻过拟合的倾向,在训练的微调为基础的方法和SCDNet。实验表明WavLm在SCD任务中的优越性,也证明了SCDNet的良好设计。摘要:Speaker Change Detection (SCD) is to identify boundaries among speakers in a conversation. Motivated by the success of fine-tuning wav2vec 2.0 models for the SCD task, a further investigation of self-supervised learning (SSL) features for SCD is conducted in this work. Specifically, an SCD model, named SCDNet, is proposed. With this model, various state-of-the-art SSL models, including Hubert, wav2vec 2.0, and WavLm are investigated. To discern the most potent layer of SSL models for SCD, a learnable weighting method is employed to analyze the effectiveness of intermediate representations. Additionally, a fine-tuning-based approach is also implemented to further compare the characteristics of SSL models in the SCD task. Furthermore, a contrastive learning method is proposed to mitigate the overfitting tendencies in the training of both the fine-tuning-based method and SCDNet. Experiments showcase the superiority of WavLm in the SCD task and also demonstrate the good design of SCDNet.
【27】 Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
标题: 利用ASB文字记录进行语音情感识别:错误率和融合技术的综合研究
作者:Yuanchao Li,Peter Bell,Catherine Lai
链接:点击下载PDF文件
摘要:文本数据通常用作主要输入以增强语音情感识别(SER)的性能和可靠性。然而,在大多数研究中对人类转录文本的依赖阻碍了实际SER系统的开发,在实验室研究和自动语音识别(ASR)作为文本源的现实场景之间形成了差距。因此,本研究基准SER性能使用ASR成绩单与不同的词错误率(WER)的知名语料库:IEMOCAP,CMU-MOSI,和MSP播客。我们的评估包括纯文本和双峰SER与不同的融合技术,旨在进行全面的分析,揭示新的发现和当前SER研究所面临的挑战。此外,我们提出了一个统一的ASR错误鲁棒框架,集成了ASR纠错和模态门控融合,实现了较低的WER和较高的SER结果相比,性能最好的ASR成绩单。这项研究预计将提供深入了解SER与ASR的援助,特别是在现实世界中的应用。摘要:Text data is commonly utilized as a primary input to enhance Speech Emotion Recognition (SER) performance and reliability. However, the reliance on human-transcribed text in most studies impedes the development of practical SER systems, creating a gap between in-lab research and real-world scenarios where Automatic Speech Recognition (ASR) serves as the text source. Hence, this study benchmarks SER performance using ASR transcripts with varying Word Error Rates (WERs) on well-known corpora: IEMOCAP, CMU-MOSI, and MSP-Podcast. Our evaluation includes text-only and bimodal SER with diverse fusion techniques, aiming for a comprehensive analysis that uncovers novel findings and challenges faced by current SER research. Additionally, we propose a unified ASR error-robust framework integrating ASR error correction and modality-gated fusion, achieving lower WER and higher SER results compared to the best-performing ASR transcript. This research is expected to provide insights into SER with ASR assistance, especially for real-world applications.
【28】 Refining Self-Supervised Learnt Speech Representation using Brain Activations
标题: 使用大脑激活完善自我监督学习语音表达
作者:Hengyu Li,Kangdi Mei,Zhaoci Liu,Yang Ai,Liping Chen,Jie Zhang,Zhenhua Ling
备注:accpeted by Interspeech2024
链接:点击下载PDF文件
摘要:文献表明,由自监督预训练模型提取的语音表示与人类的语音感知大脑激活具有相似性,并且在下游任务上微调语音表示模型可以进一步提高相似性。然而,目前还不清楚这种相似性是否可以用来优化预训练的语音模型。因此,在这项工作中,我们建议使用fMRI记录的大脑激活,通过将模型表示对准人类神经反应来改进常用的wav2vec2.0模型。SUPERB上的实验结果表明,该操作对几个下游任务是有益的,例如,说话人确认,自动语音识别,意图分类。然后,人们可以将所提出的方法视为改进自监督语音模型的一种新的替代方案。摘要:It was shown in literature that speech representations extracted by self-supervised pre-trained models exhibit similarities with brain activations of human for speech perception and fine-tuning speech representation models on downstream tasks can further improve the similarity. However, it still remains unclear if this similarity can be used to optimize the pre-trained speech models. In this work, we therefore propose to use the brain activations recorded by fMRI to refine the often-used wav2vec2.0 model by aligning model representations toward human neural responses. Experimental results on SUPERB reveal that this operation is beneficial for several downstream tasks, e.g., speaker verification, automatic speech recognition, intent classification.One can then consider the proposed method as a new alternative to improve self-supervised speech models.
【29】 Transformer-based Model for ASR N-Best Rescoring and Rewriting
标题: 基于转换器的ASB N-Best重新评分和重写模型
作者:Iwen E. Kang,Christophe Van Gysel,Man-Hung Siu
备注:Interspeech '24
链接:点击下载PDF文件
摘要:语音助理越来越多地使用设备上的自动语音识别(ASR)来确保速度和隐私。然而,由于设备上的资源约束,涉及复杂信息域的查询通常需要由搜索引擎进一步处理。对于这样的应用程序,我们提出了一种新的Transformer为基础的模型能够rescoring和重写,通过探索完整的上下文的N-最好的假设并行。我们还提出了一个新的判别序列训练目标,可以很好地为rescore和重写任务。我们表明,我们的Rescore+重写模型优于Rescore只基线,并实现了高达平均8.6%的相对字错误率(WER)降低ASR系统本身。摘要:Voice assistants increasingly use on-device Automatic Speech Recognition (ASR) to ensure speed and privacy. However, due to resource constraints on the device, queries pertaining to complex information domains often require further processing by a search engine. For such applications, we propose a novel Transformer based model capable of rescoring and rewriting, by exploring full context of the N-best hypotheses in parallel. We also propose a new discriminative sequence training objective that can work well for both rescore and rewrite tasks. We show that our Rescore+Rewrite model outperforms the Rescore-only baseline, and achieves up to an average 8.6% relative Word Error Rate (WER) reduction over the ASR system by itself.
【30】 LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
标题: LAFMA:一种用于文本到音频生成的潜在流匹配模型
作者:Wenhao Guan,Kaidi Wang,Wangjin Zhou,Yang Wang,Feng Deng,Hui Wang,Lin Li,Qingyang Hong,Yong Qin
备注:Accepted at Interspeech2024
链接:点击下载PDF文件
摘要:近年来,扩散模型的应用促进了语音和音频生成的重大发展。然而,扩散模型生成的样本质量仍需改进。该方法的有效性伴随着大量的采样步骤,导致生成高质量音频所需的合成时间延长。以前的文本到音频(TTA)的方法大多使用扩散模型的音频生成的潜在空间。在本文中,我们探讨了整合流匹配(FM)模型到音频生成的音频潜在空间。FM是一种替代的无模拟方法,它基于回归向量场训练连续归一化流(CNF)。我们证明,我们的模型显着提高了生成的音频样本的质量,实现更好的性能比以前的模型。此外,它将推理步骤的数量减少到10个步骤,几乎没有牺牲性能。摘要:Recently, the application of diffusion models has facilitated the significant development of speech and audio generation. Nevertheless, the quality of samples generated by diffusion models still needs improvement. And the effectiveness of the method is accompanied by the extensive number of sampling steps, leading to an extended synthesis time necessary for generating high-quality audio. Previous Text-to-Audio (TTA) methods mostly used diffusion models in the latent space for audio generation. In this paper, we explore the integration of the Flow Matching (FM) model into the audio latent space for audio generation. The FM is an alternative simulation-free method that trains continuous normalization flows (CNF) based on regressing vector fields. We demonstrate that our model significantly enhances the quality of generated audio samples, achieving better performance than prior models. Moreover, it reduces the number of inference steps to ten steps almost without sacrificing performance.
【31】 Fully Few-shot Class-incremental Audio Classification Using Expandable Dual-embedding Extractor
标题: 使用可扩展双嵌入提取器的全Few-Shot类增量音频分类
作者:Yongjie Si,Yanxiong Li,Jialong Li,Jiaxin Tan,Qianhua He
备注:Accepted for publication on Interspeech 2024. 5 pages, 3 figures, 5 tables
链接:点击下载PDF文件
摘要:假设在Few-Shot类增量音频分类的基础会话中训练数据是足够的。然而,由于某些类的数据稀缺,在一些实际场景中很难在基础会话中收集足够的样本进行模型训练。本文探讨了一个新的问题,全Few-Shot类增量音频分类在所有会话的少量训练样本。在此基础上,提出了一种基于可扩展双嵌入提取器的方法,该模型由嵌入提取器和可扩展分类器组成。嵌入提取器由预训练的音频频谱图Transformer(AST)和微调的AST组成。可扩展分类器由原型组成,每个原型代表一个类。实验在三个数据集(LS-100、NSynth-100和FSC-89)上进行。结果表明,我们的方法超过了七个基线的平均准确性与统计意义。代码在:https: github.com YongjieSi EDE。摘要:It's assumed that training data is sufficient in base session of few-shot class-incremental audio classification. However, it's difficult to collect abundant samples for model training in base session in some practical scenarios due to the data scarcity of some classes. This paper explores a new problem of fully few-shot class-incremental audio classification with few training samples in all sessions. Moreover, we propose a method using expandable dual-embedding extractor to solve it. The proposed model consists of an embedding extractor and an expandable classifier. The embedding extractor consists of a pretrained Audio Spectrogram Transformer (AST) and a finetuned AST. The expandable classifier consists of prototypes and each prototype represents a class. Experiments are conducted on three datasets (LS-100, NSynth-100 and FSC-89). Results show that our method exceeds seven baseline ones in average accuracy with statistical significance. Code is at: https: github.com YongjieSi EDE.
【32】 Low-Complexity Acoustic Scene Classification Using Parallel Attention-Convolution Network
标题: 使用并行注意卷积网络的低复杂度声场景分类
作者:Yanxiong Li,Jiaxin Tan,Guoqing Chen,Jialong Li,Yongjie Si,Qianhua He
备注:Accepted for publication on Interspeech 2024. 5 pages, 4 figures, 3 tables
链接:点击下载PDF文件
摘要:这项工作是我们提交给DCASE2023挑战任务1的改进系统。提出了一种基于并行注意力卷积网络的低复杂度声场景分类方法,该方法由预处理、融合、全局和局部上下文信息提取四个模块组成。所提出的网络是计算效率从每个音频剪辑捕获全球和本地的上下文信息。此外,我们将其他技术集成到我们的方法中,如知识蒸馏,数据增强和自适应残差归一化。在DCASE2023挑战的官方数据集上进行评估时,我们的方法获得了56.10%的最高准确率,参数数为5.21 kilo,乘法累加运算为144万次。在精度和复杂度方面超过了DCASE2023挑战赛的前两名,取得了最先进的结果。代码在:https: github.com Jessytan Low-complexity-ASC。摘要:This work is an improved system that we submitted to task 1 of DCASE2023 challenge. We propose a method of low-complexity acoustic scene classification by a parallel attention-convolution network which consists of four modules, including pre-processing, fusion, global and local contextual information extraction. The proposed network is computationally efficient to capture global and local contextual information from each audio clip. In addition, we integrate other techniques into our method, such as knowledge distillation, data augmentation, and adaptive residual normalization. When evaluated on the official dataset of DCASE2023 challenge, our method obtains the highest accuracy of 56.10% with parameter number of 5.21 kilo and multiply-accumulate operations of 1.44 million. It exceeds the top two systems of DCASE2023 challenge in accuracy and complexity, and obtains state-of-the-art result. Code is at: https: github.com Jessytan Low-complexity-ASC.
【33】 VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
标题: VECL-TTC:语音身份和情感风格可控跨语言文本转语音
作者:Ashishkumar Gudmalwar,Nirmesh Shah,Sai Akarsh,Pankaj Wasnik,Rajiv Ratn Shah
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:尽管文本到语音(TTS)系统的显着进步,其在自动配音的充分利用仍然有限。这一任务需要从源语言的参考语音中提取语音身份和情感风格,然后使用跨语言TTS技术将其转换为目标语言。虽然以前的方法主要集中在跨语言TTS框架内控制语音身份,但将情感和语音身份结合在一起的工作有限。为此,我们介绍了一个端到端的语音身份和情感风格可控的跨语言(VECL)TTS系统,使用多语种扬声器和情感嵌入网络。此外,我们引入内容和风格的一致性损失,以进一步提高合成语音的质量。所提出的系统实现了8.83 %的平均相对改善相比,国家的最先进的(SOTA)的方法,包括英语和三种印度语言(印地语,泰卢固语和马拉地语)的数据库。摘要:Despite the significant advancements in Text-to-Speech (TTS) systems, their full utilization in automatic dubbing remains limited. This task necessitates the extraction of voice identity and emotional style from a reference speech in a source language and subsequently transferring them to a target language using cross-lingual TTS techniques. While previous approaches have mainly concentrated on controlling voice identity within the cross-lingual TTS framework, there has been limited work on incorporating emotion and voice identity together. To this end, we introduce an end-to-end Voice Identity and Emotional Style Controllable Cross-Lingual (VECL) TTS system using multilingual speakers and an emotion embedding network. Moreover, we introduce content and style consistency losses to enhance the quality of synthesized speech further. The proposed system achieved an average relative improvement of 8.83 % compared to the state-of-the-art (SOTA) methods on a database comprising English and three Indian languages (Hindi, Telugu, and Marathi).
【34】 DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels
标题: DUSE 2024任务4:使用异类数据和缺失标签的声音事件检测
作者:Samuele Cornell,Janek Ebbers,Constance Douwes,Irene Martín-Morató,Manu Harju,Annamaria Mesaros,Romain Serizel
链接:点击下载PDF文件
摘要:声学场景和事件的检测和分类挑战任务4旨在通过利用具有不同监督不确定性的训练数据来推进家庭环境中的声音事件检测(SED)系统。参与者面临的挑战是探索如何最好地使用来自不同领域的训练数据,并具有不同的注释粒度(强 弱时间分辨率,软 硬标签),以获得一个强大的SED系统,可以在不同的场景中推广。至关重要的是,可用训练数据集之间的注释可能不一致,因此一个数据集的声音标签可能存在,但在另一个数据集中没有注释,反之亦然。因此,系统将不得不在训练期间处理潜在的目标标签缺失。此外,作为一个额外的新颖性,系统也将在不同粒度的标签上进行评估,以评估它们对不同应用的鲁棒性。为了降低参与者的进入门槛,我们开发了一个更新的基线系统,其中包含几个警告,以解决上述问题。与我们的基线系统的结果表明,这个研究方向是有前途的,并有可能获得一个更强大的SED系统,通过使用不同的域训练数据与丢失的标签相比,分别为每个域训练一个SED系统。摘要:The Detection and Classification of Acoustic Scenes and Events Challenge Task 4 aims to advance sound event detection (SED) systems in domestic environments by leveraging training data with different supervision uncertainty. Participants are challenged in exploring how to best use training data from different domains and with varying annotation granularity (strong weak temporal resolution, soft hard labels), to obtain a robust SED system that can generalize across different scenarios. Crucially, annotation across available training datasets can be inconsistent and hence sound labels of one dataset may be present but not annotated in the other one and vice-versa. As such, systems will have to cope with potentially missing target labels during training. Moreover, as an additional novelty, systems will also be evaluated on labels with different granularity in order to assess their robustness for different applications. To lower the entry barrier for participants, we developed an updated baseline system with several caveats to address these aforementioned problems. Results with our baseline system indicate that this research direction is promising and is possible to obtain a stronger SED system by using diverse domain training data with missing labels compared to training a SED system for each domain separately.
【35】 LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
标题: LibriTTS-P:具有说话风格和说话者身份的数据库,支持文本转语音和风格字幕
作者:Masaya Kawamura,Ryuichi Yamamoto,Yuma Shirahata,Takuya Hasumi,Kentaro Tachibana
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:我们介绍了LibriTTS-P,一个基于LibriTTS-R的新语料库,它包括话语级描述(即,提示)和说话者特征的说话者级别提示。我们采用一种混合的方法来构建提示注释:(1)手动注释,捕捉人类对说话者特征的感知和(2)说话风格的合成注释。与现有的英语提示语数据集相比,我们的语料库为LibriTTS-R的所有使用者提供了更多样化的提示语注释。实验结果表明,基于LibriTTS-P训练的TTS模型比使用传统数据集的模型具有更高的自然度。此外,样式字幕任务的结果表明,使用LibriTTS-P的模型生成的单词比使用传统数据集的模型准确2.5倍。我们的语料库LibriTTS-P可以在https: github.com line LibriTTS-P上找到。摘要:We introduce LibriTTS-P, a new corpus based on LibriTTS-R that includes utterance-level descriptions (i.e., prompts) of speaking style and speaker-level prompts of speaker characteristics. We employ a hybrid approach to construct prompt annotations: (1) manual annotations that capture human perceptions of speaker characteristics and (2) synthetic annotations on speaking style. Compared to existing English prompt datasets, our corpus provides more diverse prompt annotations for all speakers of LibriTTS-R. Experimental results for prompt-based controllable TTS demonstrate that the TTS model trained with LibriTTS-P achieves higher naturalness than the model using the conventional dataset. Furthermore, the results for style captioning tasks show that the model utilizing LibriTTS-P generates 2.5 times more accurate words than the model using a conventional dataset. Our corpus, LibriTTS-P, is available at https: github.com line LibriTTS-P.
【36】 Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
标题: 使用自我知识蒸馏指导框架级CSC对准
作者:Eungbeom Kim,Hantae Kim,Kyogu Lee
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:具有连接主义时间分类(CTC)框架的Transformer编码器广泛用于自动语音识别(ASR)。然而,知识蒸馏(KD)的ASR显示了一个问题,教师和学生模型之间的不一致的框架级对齐,这最终阻碍了它从提高学生模型的性能。为了解决这个问题,本文介绍了一种自知识蒸馏(SKD)方法,在训练期间指导帧级对齐。与传统的教师模型和学生模型分离的方法相比,本研究提出了一种简单有效的方法,即共享编码器层,并将子模型作为学生模型。总体而言,我们的方法在提高资源效率和性能方面是有效的。我们还进行了实验分析的尖峰时间来说明,该方法通过减少对齐不一致提高性能。摘要:Transformer encoder with connectionist temporal classification (CTC) framework is widely used for automatic speech recognition (ASR). However, knowledge distillation (KD) for ASR displays a problem of disagreement between teacher-student models in frame-level alignment which ultimately hinders it from improving the student model's performance. In order to resolve this problem, this paper introduces a self-knowledge distillation (SKD) method that guides the frame-level alignment during the training time. In contrast to the conventional method using separate teacher and student models, this study introduces a simple and effective method sharing encoder layers and applying the sub-model as the student model. Overall, our approach is effective in improving both the resource efficiency as well as performance. We also conducted an experimental analysis of the spike timings to illustrate that the proposed method improves performance by reducing the alignment disagreement.
【37】 Target Speaker Extraction with Curriculum Learning
标题: 通过课程学习提取目标说话者
作者:Yun Liu,Xuechen Liu,Xiaoxiao Miao,Junichi Yamagishi
备注:Accepted for presentation at Interspeech 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的方法,目标说话人提取(TSE)使用课程学习(CL)技术,解决区分目标说话人的声音从混合包含干扰扬声器的挑战。为了有效的训练,我们建议设计一个课程,选择子集的复杂性不断增加,如增加目标和干扰扬声器之间的相似性,并选择训练数据的战略。我们的CL策略包括使用预定义难度度量(例如性别,说话者相似性和信号失真比)的变体和使用TSE标准目标函数的变体,每种策略都旨在使模型逐渐暴露于更具挑战性的场景。在Libri2talker数据集上进行的全面测试表明,我们针对TSE的CL策略提高了性能,并且结果明显超过了没有CL的基线模型约1 dB。摘要:This paper presents a novel approach to target speaker extraction (TSE) using Curriculum Learning (CL) techniques, addressing the challenge of distinguishing a target speaker's voice from a mixture containing interfering speakers. For efficient training, we propose designing a curriculum that selects subsets of increasing complexity, such as increasing similarity between target and interfering speakers, and that selects training data strategically. Our CL strategies include both variants using predefined difficulty measures (e.g. gender, speaker similarity, and signal-to-distortion ratio) and ones using the TSE's standard objective function, each designed to expose the model gradually to more challenging scenarios. Comprehensive testing on the Libri2talker dataset demonstrated that our CL strategies for TSE improved the performance, and the results markedly exceeded baseline models without CL about 1 dB.
【38】 Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
作者:Lin Zhang,Xin Wang,Erica Cooper,Mireia Diez,Federico Landini,Nicholas Evans,Junichi Yamagishi
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文将欺骗日记定义为部分欺骗(PS)场景中的一个新任务。它的目的是确定什么时候欺骗,这不仅包括定位欺骗区域,而且还根据不同的欺骗方法对它们进行聚类。作为一个开创性的研究,在欺骗日记,我们专注于定义任务,建立评估指标,并提出了一个基准模型,即对策条件聚类(3C)模型。利用这个模型,我们首先探讨如何有效地训练对策,以支持欺骗日记使用三个标签计划。然后,我们利用欺骗本地化预测,以提高日记的性能。这第一项研究揭示了任务的高度复杂性,即使在每个音频文件只有一个扬声器和一个oracle数量的欺骗方法被认为是有限的情况下。我们的代码可以在https: github.com nii-yamagishilab PartialSpoof上找到。摘要:This paper defines Spoof Diarization as a novel task in the Partial Spoof (PS) scenario. It aims to determine what spoofed when, which includes not only locating spoof regions but also clustering them according to different spoofing methods. As a pioneering study in spoof diarization, we focus on defining the task, establishing evaluation metrics, and proposing a benchmark model, namely the Countermeasure-Condition Clustering (3C) model. Utilizing this model, we first explore how to effectively train countermeasures to support spoof diarization using three labeling schemes. We then utilize spoof localization predictions to enhance the diarization performance. This first study reveals the high complexity of the task, even in restricted scenarios where only a single speaker per audio file and an oracle number of spoofing methods are considered. Our code is available at https: github.com nii-yamagishilab PartialSpoof.
【39】 Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based models
标题: 实现客观和可解释的言语障碍评估:CNN和基于转换器的模型的比较分析
作者:Malo Maisonneuve,Corinne Fredouille,Muriel Lalain,Alain Ghio,Virginie Woisard
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:头颈癌(HNC)严重影响患者的说话能力,影响他们的生活质量。用于评估病理性语音的常用指标是主观的,这促使需要自动化和无偏见的评估方法。本研究提出一个自我监督的Wav2Vec2为基础的模型与HNC患者的电话分类,以提高准确性和改善语音特征的歧视,为后续的可解释性的目的。探讨了预训练数据集、模型大小以及微调数据集和参数的影响。对不同语料库的评估揭示了Wav2Vec2架构的有效性,优于以前工作中使用的基于CNN的方法。与感知测量的相关性也肯定了受损语音分析的模型相关性。这项工作为临床医生更好地理解病理语言铺平了道路,通过利用复杂的自学语言表征。摘要:Head and Neck Cancers (HNC) significantly impact patients' ability to speak, affecting their quality of life. Commonly used metrics for assessing pathological speech are subjective, prompting the need for automated and unbiased evaluation methods. This study proposes a self-supervised Wav2Vec2-based model for phone classification with HNC patients, to enhance accuracy and improve the discrimination of phonetic features for subsequent interpretability purpose. The impact of pre-training datasets, model size, and fine-tuning datasets and parameters are explored. Evaluation on diverse corpora reveals the effectiveness of the Wav2Vec2 architecture, outperforming a CNN-based approach, used in previous work. Correlation with perceptual measures also affirms the model relevance for impaired speech analysis. This work paves the way for better understanding of pathological speech with interpretable approaches for clinicians, by leveraging complex self-learnt speech representations.
eess.AS音频处理
【1】 SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models标题: SVSNet+:利用语音基础模型的表示增强说话者语音相似性评估模型
作者:Chun Yin,Tai-Shih Chi,Yu Tsao,Hsin-Min Wang
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:来自预先训练的语音基础模型(SFM)的表示在许多下游任务中表现出令人印象深刻的性能。然而,将预先训练的SFM表示到扬声器语音相似性评估的潜在好处还没有得到彻底的研究。在本文中,我们提出了SVSNet+,这是一种集成了预训练的SFM表示的模型,可以提高评估说话者语音相似性的性能。在语音转换挑战2018和2020数据集上的实验结果表明,与基线模型相比,SVSNet+结合WavLM表示显示出显着的改进。此外,虽然用下游任务的小数据集微调WavLM不会提高性能,但使用相同的数据集来学习WavLM的加权和表示可以大大提高性能。此外,当WavLM被其他SFM取代时,SVSNet+仍然优于基线模型,并表现出很强的泛化能力。摘要:Representations from pre-trained speech foundation models (SFMs) have shown impressive performance in many downstream tasks. However, the potential benefits of incorporating pre-trained SFM representations into speaker voice similarity assessment have not been thoroughly investigated. In this paper, we propose SVSNet+, a model that integrates pre-trained SFM representations to improve performance in assessing speaker voice similarity. Experimental results on the Voice Conversion Challenge 2018 and 2020 datasets show that SVSNet+ incorporating WavLM representations shows significant improvements compared to baseline models. In addition, while fine-tuning WavLM with a small dataset of the downstream task does not improve performance, using the same dataset to learn a weighted-sum representation of WavLM can substantially improve performance. Furthermore, when WavLM is replaced by other SFMs, SVSNet+ still outperforms the baseline models and exhibits strong generalization ability.
【2】 Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
标题: 理解声音,错过问题:大型音频语言模型中对象幻觉的挑战
作者:Chun-Yi Kuan,Wei-Ping Huang,Hung-yi Lee
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:大型音频语言模型(LALM)通过集成音频感知功能来增强传统的大型语言模型,使它们能够解决音频相关的任务。之前的研究主要集中在评估LALM在各种任务中的表现,但忽视了它们的可靠性,特别是关于物体幻觉等问题。在我们的研究中,我们介绍了评估公开可用的LALM的物体幻觉程度的方法。我们的研究结果表明,LALM在对音频内容的理解方面与专门的音频字幕模型相当,但很难回答区分性问题,特别是那些需要识别音频片段中特定对象声音存在的问题。这种限制突出了当前LALM的一个关键弱点:它们对歧视性查询的理解不足。此外,我们探讨了潜在的提示工程,以提高LALM的性能上的歧视性问题。摘要:Large audio-language models (LALMs) enhance traditional large language models by integrating audio perception capabilities, allowing them to tackle audio-related tasks. Previous research has primarily focused on assessing the performance of LALMs across various tasks, yet overlooking their reliability, particularly concerning issues like object hallucination. In our study, we introduce methods to assess the extent of object hallucination of publicly available LALMs. Our findings reveal that LALMs are comparable to specialized audio captioning models in their understanding of audio content, but struggle to answer discriminative questions, specifically those requiring the identification of the presence of particular object sounds within an audio clip. This limitation highlights a critical weakness in current LALMs: their inadequate understanding of discriminative queries. Moreover, we explore the potential of prompt engineering to enhance LALMs' performance on discriminative questions.
【3】 Neural Blind Source Separation and Diarization for Distant Speech Recognition
标题: 用于远距离语音识别的神经盲源分离和扩展
作者:Yoshiaki Bando,Tomohiko Nakamura,Shinji Watanabe
备注:5 pages, 3 figures, accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:本文提出了一种用于远距离语音识别(DSR)的神经网络方法,该方法可以在不受孤立信号监督的情况下对语音混合信号进行联合分离和分类。用于多讲话者DSR的标准分离方法是被称为引导源分离(GSS)的统计多通道方法。虽然GSS不需要信号电平的监督,它依赖于扬声器日记的结果来处理未知数量的活跃扬声器。为了克服这一局限性,我们引入和训练的神经推理模型在弱监督的方式,采用统计分离方法的目标函数。这种训练只需要多通道混合和他们的时间注释的扬声器活动。与GSS相比,训练后的模型可以在没有任何辅助信息的情况下联合分离和区分语音混合。在AMI语料库上的实验表明,该方法在单词错误率方面优于GSS和Oracle日志化结果。代码可在线获取。摘要:This paper presents a neural method for distant speech recognition (DSR) that jointly separates and diarizes speech mixtures without supervision by isolated signals. A standard separation method for multi-talker DSR is a statistical multichannel method called guided source separation (GSS). While GSS does not require signal-level supervision, it relies on speaker diarization results to handle unknown numbers of active speakers. To overcome this limitation, we introduce and train a neural inference model in a weakly-supervised manner, employing the objective function of a statistical separation method. This training requires only multichannel mixtures and their temporal annotations of speaker activities. In contrast to GSS, the trained model can jointly separate and diarize speech mixtures without any auxiliary information. The experiments with the AMI corpus show that our method outperforms GSS with oracle diarization results regarding word error rates. The code is available online.
【4】 SCDNet: Self-supervised Learning Feature-based Speaker Change Detection
标题: SCDNet:基于自我监督学习的说话者变化检测
作者:Yue Li,Xinsheng Wang,Li Zhang,Lei Xie
链接:点击下载PDF文件
摘要:说话人变化检测(SCD)是识别会话中说话人之间的边界。出于对SCD任务的wav2vec 2.0模型进行微调的成功,在这项工作中对SCD的自监督学习(SSL)功能进行了进一步的研究。具体而言,SCD模型,命名为SCDNet,提出。在此基础上,研究了Hubert、wav2vec 2.0和WavLm等多种SSL模型。为了识别最有效的层的SSL模型SCD,学习加权方法来分析的有效性的中间表示。此外,还实现了一种基于微调的方法,以进一步比较SCD任务中SSL模型的特性。此外,提出了一种对比学习方法,以减轻过拟合的倾向,在训练的微调为基础的方法和SCDNet。实验表明WavLm在SCD任务中的优越性,也证明了SCDNet的良好设计。摘要:Speaker Change Detection (SCD) is to identify boundaries among speakers in a conversation. Motivated by the success of fine-tuning wav2vec 2.0 models for the SCD task, a further investigation of self-supervised learning (SSL) features for SCD is conducted in this work. Specifically, an SCD model, named SCDNet, is proposed. With this model, various state-of-the-art SSL models, including Hubert, wav2vec 2.0, and WavLm are investigated. To discern the most potent layer of SSL models for SCD, a learnable weighting method is employed to analyze the effectiveness of intermediate representations. Additionally, a fine-tuning-based approach is also implemented to further compare the characteristics of SSL models in the SCD task. Furthermore, a contrastive learning method is proposed to mitigate the overfitting tendencies in the training of both the fine-tuning-based method and SCDNet. Experiments showcase the superiority of WavLm in the SCD task and also demonstrate the good design of SCDNet.
【5】 Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
标题: 利用ASB文字记录进行语音情感识别:错误率和融合技术的综合研究
作者:Yuanchao Li,Peter Bell,Catherine Lai
链接:点击下载PDF文件
摘要:文本数据通常用作主要输入以增强语音情感识别(SER)的性能和可靠性。然而,在大多数研究中对人类转录文本的依赖阻碍了实际SER系统的开发,在实验室研究和自动语音识别(ASR)作为文本源的现实场景之间形成了差距。因此,本研究基准SER性能使用ASR成绩单与不同的词错误率(WER)的知名语料库:IEMOCAP,CMU-MOSI,和MSP播客。我们的评估包括纯文本和双峰SER与不同的融合技术,旨在进行全面的分析,揭示新的发现和当前SER研究所面临的挑战。此外,我们提出了一个统一的ASR错误鲁棒框架,集成了ASR纠错和模态门控融合,实现了较低的WER和较高的SER结果相比,性能最好的ASR成绩单。这项研究预计将提供深入了解SER与ASR的援助,特别是在现实世界中的应用。摘要:Text data is commonly utilized as a primary input to enhance Speech Emotion Recognition (SER) performance and reliability. However, the reliance on human-transcribed text in most studies impedes the development of practical SER systems, creating a gap between in-lab research and real-world scenarios where Automatic Speech Recognition (ASR) serves as the text source. Hence, this study benchmarks SER performance using ASR transcripts with varying Word Error Rates (WERs) on well-known corpora: IEMOCAP, CMU-MOSI, and MSP-Podcast. Our evaluation includes text-only and bimodal SER with diverse fusion techniques, aiming for a comprehensive analysis that uncovers novel findings and challenges faced by current SER research. Additionally, we propose a unified ASR error-robust framework integrating ASR error correction and modality-gated fusion, achieving lower WER and higher SER results compared to the best-performing ASR transcript. This research is expected to provide insights into SER with ASR assistance, especially for real-world applications.
【6】 Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation
标题: 规则化语音分离的定时文本和音频之间的多模式表示损失
作者:Tsun-An Hsieh,Heeyoul Choi,Minje Kim
链接:点击下载PDF文件
摘要:最近的研究强调了语篇模态在制约语音分离模型的推理过程中的潜力。然而,基于正则化的方法仍然没有得到充分的研究,尽管它们的优点是在测试期间不需要辅助文本数据。为了解决这一差距,我们引入了一个定时的基于文本的正则化(TTR)方法,使用语言模型派生的语义,以改善语音分离模型。我们的方法包括两个步骤。我们分别从两个预训练的音频和语言模型WavLM和BERT开始。然后,学习基于transformer的音频摘要器来对齐音频和单词嵌入并最小化它们的间隙。摘要器Transformer作为一个正则化器合并,促进分离的源代码与定时文本的语义对齐。实验结果表明,所提出的TTR方法一致地提高了各种客观指标的分离结果在非正则化的基线。摘要:Recent studies highlight the potential of textual modalities in conditioning the speech separation model's inference process. However, regularization-based methods remain underexplored despite their advantages of not requiring auxiliary text data during the test time. To address this gap, we introduce a timed text-based regularization (TTR) method that uses language model-derived semantics to improve speech separation models. Our approach involves two steps. We begin with two pretrained audio and language models, WavLM and BERT, respectively. Then, a Transformer-based audio summarizer is learned to align the audio and word embeddings and to minimize their gap. The summarizer Transformer, incorporated as a regularizer, promotes the separated sources' alignment with the semantics from the timed text. Experimental results show that the proposed TTR method consistently improves the various objective metrics of the separation results over the unregularized baselines.
【7】 Refining Self-Supervised Learnt Speech Representation using Brain Activations
标题: 使用大脑激活完善自我监督学习语音表达
作者:Hengyu Li,Kangdi Mei,Zhaoci Liu,Yang Ai,Liping Chen,Jie Zhang,Zhenhua Ling
备注:accpeted by Interspeech2024
链接:点击下载PDF文件
摘要:文献表明,由自监督预训练模型提取的语音表示与人类的语音感知大脑激活具有相似性,并且在下游任务上微调语音表示模型可以进一步提高相似性。然而,目前还不清楚这种相似性是否可以用来优化预训练的语音模型。因此,在这项工作中,我们建议使用fMRI记录的大脑激活,通过将模型表示对准人类神经反应来改进常用的wav2vec2.0模型。SUPERB上的实验结果表明,该操作对几个下游任务是有益的,例如,说话人确认,自动语音识别,意图分类。然后,人们可以将所提出的方法视为改进自监督语音模型的一种新的替代方案。摘要:It was shown in literature that speech representations extracted by self-supervised pre-trained models exhibit similarities with brain activations of human for speech perception and fine-tuning speech representation models on downstream tasks can further improve the similarity. However, it still remains unclear if this similarity can be used to optimize the pre-trained speech models. In this work, we therefore propose to use the brain activations recorded by fMRI to refine the often-used wav2vec2.0 model by aligning model representations toward human neural responses. Experimental results on SUPERB reveal that this operation is beneficial for several downstream tasks, e.g., speaker verification, automatic speech recognition, intent classification.One can then consider the proposed method as a new alternative to improve self-supervised speech models.
【8】 Transformer-based Model for ASR N-Best Rescoring and Rewriting
标题: 基于转换器的ASB N-Best重新评分和重写模型
作者:Iwen E. Kang,Christophe Van Gysel,Man-Hung Siu
备注:Interspeech '24
链接:点击下载PDF文件
摘要:语音助理越来越多地使用设备上的自动语音识别(ASR)来确保速度和隐私。然而,由于设备上的资源约束,涉及复杂信息域的查询通常需要由搜索引擎进一步处理。对于这样的应用程序,我们提出了一种新的Transformer为基础的模型能够rescoring和重写,通过探索完整的上下文的N-最好的假设并行。我们还提出了一个新的判别序列训练目标,可以很好地为rescore和重写任务。我们表明,我们的Rescore+重写模型优于Rescore只基线,并实现了高达平均8.6%的相对字错误率(WER)降低ASR系统本身。摘要:Voice assistants increasingly use on-device Automatic Speech Recognition (ASR) to ensure speed and privacy. However, due to resource constraints on the device, queries pertaining to complex information domains often require further processing by a search engine. For such applications, we propose a novel Transformer based model capable of rescoring and rewriting, by exploring full context of the N-best hypotheses in parallel. We also propose a new discriminative sequence training objective that can work well for both rescore and rewrite tasks. We show that our Rescore+Rewrite model outperforms the Rescore-only baseline, and achieves up to an average 8.6% relative Word Error Rate (WER) reduction over the ASR system by itself.
【9】 LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
标题: LAFMA:一种用于文本到音频生成的潜在流匹配模型
作者:Wenhao Guan,Kaidi Wang,Wangjin Zhou,Yang Wang,Feng Deng,Hui Wang,Lin Li,Qingyang Hong,Yong Qin
备注:Accepted at Interspeech2024
链接:点击下载PDF文件
摘要:近年来,扩散模型的应用促进了语音和音频生成的重大发展。然而,扩散模型生成的样本质量仍需改进。该方法的有效性伴随着大量的采样步骤,导致生成高质量音频所需的合成时间延长。以前的文本到音频(TTA)的方法大多使用扩散模型的音频生成的潜在空间。在本文中,我们探讨了整合流匹配(FM)模型到音频生成的音频潜在空间。FM是一种替代的无模拟方法,它基于回归向量场训练连续归一化流(CNF)。我们证明,我们的模型显着提高了生成的音频样本的质量,实现更好的性能比以前的模型。此外,它将推理步骤的数量减少到10个步骤,几乎没有牺牲性能。摘要:Recently, the application of diffusion models has facilitated the significant development of speech and audio generation. Nevertheless, the quality of samples generated by diffusion models still needs improvement. And the effectiveness of the method is accompanied by the extensive number of sampling steps, leading to an extended synthesis time necessary for generating high-quality audio. Previous Text-to-Audio (TTA) methods mostly used diffusion models in the latent space for audio generation. In this paper, we explore the integration of the Flow Matching (FM) model into the audio latent space for audio generation. The FM is an alternative simulation-free method that trains continuous normalization flows (CNF) based on regressing vector fields. We demonstrate that our model significantly enhances the quality of generated audio samples, achieving better performance than prior models. Moreover, it reduces the number of inference steps to ten steps almost without sacrificing performance.
【10】 Fully Few-shot Class-incremental Audio Classification Using Expandable Dual-embedding Extractor
标题: 使用可扩展双嵌入提取器的全Few-Shot类增量音频分类
作者:Yongjie Si,Yanxiong Li,Jialong Li,Jiaxin Tan,Qianhua He
备注:Accepted for publication on Interspeech 2024. 5 pages, 3 figures, 5 tables
链接:点击下载PDF文件
摘要:假设在Few-Shot类增量音频分类的基础会话中训练数据是足够的。然而,由于某些类的数据稀缺,在一些实际场景中很难在基础会话中收集足够的样本进行模型训练。本文探讨了一个新的问题,全Few-Shot类增量音频分类在所有会话的少量训练样本。在此基础上,提出了一种基于可扩展双嵌入提取器的方法,该模型由嵌入提取器和可扩展分类器组成。嵌入提取器由预训练的音频频谱图Transformer(AST)和微调的AST组成。可扩展分类器由原型组成,每个原型代表一个类。实验在三个数据集(LS-100、NSynth-100和FSC-89)上进行。结果表明,我们的方法超过了七个基线的平均准确性与统计意义。代码在:https: github.com YongjieSi EDE。摘要:It's assumed that training data is sufficient in base session of few-shot class-incremental audio classification. However, it's difficult to collect abundant samples for model training in base session in some practical scenarios due to the data scarcity of some classes. This paper explores a new problem of fully few-shot class-incremental audio classification with few training samples in all sessions. Moreover, we propose a method using expandable dual-embedding extractor to solve it. The proposed model consists of an embedding extractor and an expandable classifier. The embedding extractor consists of a pretrained Audio Spectrogram Transformer (AST) and a finetuned AST. The expandable classifier consists of prototypes and each prototype represents a class. Experiments are conducted on three datasets (LS-100, NSynth-100 and FSC-89). Results show that our method exceeds seven baseline ones in average accuracy with statistical significance. Code is at: https: github.com YongjieSi EDE.
【11】 Low-Complexity Acoustic Scene Classification Using Parallel Attention-Convolution Network
标题: 使用并行注意卷积网络的低复杂度声场景分类
作者:Yanxiong Li,Jiaxin Tan,Guoqing Chen,Jialong Li,Yongjie Si,Qianhua He
备注:Accepted for publication on Interspeech 2024. 5 pages, 4 figures, 3 tables
链接:点击下载PDF文件
摘要:这项工作是我们提交给DCASE2023挑战任务1的改进系统。提出了一种基于并行注意力卷积网络的低复杂度声场景分类方法,该方法由预处理、融合、全局和局部上下文信息提取四个模块组成。所提出的网络是计算效率从每个音频剪辑捕获全球和本地的上下文信息。此外,我们将其他技术集成到我们的方法中,如知识蒸馏,数据增强和自适应残差归一化。在DCASE2023挑战的官方数据集上进行评估时,我们的方法获得了56.10%的最高准确率,参数数为5.21 kilo,乘法累加运算为144万次。在精度和复杂度方面超过了DCASE2023挑战赛的前两名,取得了最先进的结果。代码在:https: github.com Jessytan Low-complexity-ASC。摘要:This work is an improved system that we submitted to task 1 of DCASE2023 challenge. We propose a method of low-complexity acoustic scene classification by a parallel attention-convolution network which consists of four modules, including pre-processing, fusion, global and local contextual information extraction. The proposed network is computationally efficient to capture global and local contextual information from each audio clip. In addition, we integrate other techniques into our method, such as knowledge distillation, data augmentation, and adaptive residual normalization. When evaluated on the official dataset of DCASE2023 challenge, our method obtains the highest accuracy of 56.10% with parameter number of 5.21 kilo and multiply-accumulate operations of 1.44 million. It exceeds the top two systems of DCASE2023 challenge in accuracy and complexity, and obtains state-of-the-art result. Code is at: https: github.com Jessytan Low-complexity-ASC.
【12】 Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
标题: 用于从未标记的语音数据构建文本到语音模型的音频条件音素和韵律注释
作者:Yuma Shirahata,Byeongseon Park,Ryuichi Yamamoto,Kentaro Tachibana
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:本文提出了一个音频条件音素和韵律标注模型,用于从未标记的语音样本构建文本到语音(TTS)数据集。为了创建由标签语音配对数据组成的TTS数据集,所提出的注释模型利用自动语音识别(ASR)模型从未标记的语音样本中获得音素和韵律标签。通过微调大规模预训练的ASR模型,我们可以使用现有TTS数据集中有限数量的标签语音配对数据来构建注释模型。为了缓解标签语音配对数据训练标注模型的不足,我们使用纯文本语料库和辅助TTS模型生成伪标签语音配对数据。这个TTS模型也是用现有的TTS数据集训练的。实验结果表明,使用该方法训练的TTS模型可以像使用完全标记的数据集训练的TTS模型一样自然地合成语音。摘要:This paper proposes an audio-conditioned phonemic and prosodic annotation model for building text-to-speech (TTS) datasets from unlabeled speech samples. For creating a TTS dataset that consists of label-speech paired data, the proposed annotation model leverages an automatic speech recognition (ASR) model to obtain phonemic and prosodic labels from unlabeled speech samples. By fine-tuning a large-scale pre-trained ASR model, we can construct the annotation model using a limited amount of label-speech paired data within an existing TTS dataset. To alleviate the shortage of label-speech paired data for training the annotation model, we generate pseudo label-speech paired data using text-only corpora and an auxiliary TTS model. This TTS model is also trained with the existing TTS dataset. Experimental results show that the TTS model trained with the dataset created by the proposed annotation method can synthesize speech as naturally as the one trained with a fully-labeled dataset.
【13】 VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
标题: VECL-TTC:语音身份和情感风格可控跨语言文本转语音
作者:Ashishkumar Gudmalwar,Nirmesh Shah,Sai Akarsh,Pankaj Wasnik,Rajiv Ratn Shah
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:尽管文本到语音(TTS)系统的显着进步,其在自动配音的充分利用仍然有限。这一任务需要从源语言的参考语音中提取语音身份和情感风格,然后使用跨语言TTS技术将其转换为目标语言。虽然以前的方法主要集中在跨语言TTS框架内控制语音身份,但将情感和语音身份结合在一起的工作有限。为此,我们介绍了一个端到端的语音身份和情感风格可控的跨语言(VECL)TTS系统,使用多语种扬声器和情感嵌入网络。此外,我们引入内容和风格的一致性损失,以进一步提高合成语音的质量。所提出的系统实现了8.83 %的平均相对改善相比,国家的最先进的(SOTA)的方法,包括英语和三种印度语言(印地语,泰卢固语和马拉地语)的数据库。摘要:Despite the significant advancements in Text-to-Speech (TTS) systems, their full utilization in automatic dubbing remains limited. This task necessitates the extraction of voice identity and emotional style from a reference speech in a source language and subsequently transferring them to a target language using cross-lingual TTS techniques. While previous approaches have mainly concentrated on controlling voice identity within the cross-lingual TTS framework, there has been limited work on incorporating emotion and voice identity together. To this end, we introduce an end-to-end Voice Identity and Emotional Style Controllable Cross-Lingual (VECL) TTS system using multilingual speakers and an emotion embedding network. Moreover, we introduce content and style consistency losses to enhance the quality of synthesized speech further. The proposed system achieved an average relative improvement of 8.83 % compared to the state-of-the-art (SOTA) methods on a database comprising English and three Indian languages (Hindi, Telugu, and Marathi).
【14】 DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels
标题: DUSE 2024任务4:使用异类数据和缺失标签的声音事件检测
作者:Samuele Cornell,Janek Ebbers,Constance Douwes,Irene Martín-Morató,Manu Harju,Annamaria Mesaros,Romain Serizel
链接:点击下载PDF文件
摘要:声学场景和事件的检测和分类挑战任务4旨在通过利用具有不同监督不确定性的训练数据来推进家庭环境中的声音事件检测(SED)系统。参与者面临的挑战是探索如何最好地使用来自不同领域的训练数据,并具有不同的注释粒度(强 弱时间分辨率,软 硬标签),以获得一个强大的SED系统,可以在不同的场景中推广。至关重要的是,可用训练数据集之间的注释可能不一致,因此一个数据集的声音标签可能存在,但在另一个数据集中没有注释,反之亦然。因此,系统将不得不在训练期间处理潜在的目标标签缺失。此外,作为一个额外的新颖性,系统也将在不同粒度的标签上进行评估,以评估它们对不同应用的鲁棒性。为了降低参与者的进入门槛,我们开发了一个更新的基线系统,其中包含几个警告,以解决上述问题。与我们的基线系统的结果表明,这个研究方向是有前途的,并有可能获得一个更强大的SED系统,通过使用不同的域训练数据与丢失的标签相比,分别为每个域训练一个SED系统。摘要:The Detection and Classification of Acoustic Scenes and Events Challenge Task 4 aims to advance sound event detection (SED) systems in domestic environments by leveraging training data with different supervision uncertainty. Participants are challenged in exploring how to best use training data from different domains and with varying annotation granularity (strong weak temporal resolution, soft hard labels), to obtain a robust SED system that can generalize across different scenarios. Crucially, annotation across available training datasets can be inconsistent and hence sound labels of one dataset may be present but not annotated in the other one and vice-versa. As such, systems will have to cope with potentially missing target labels during training. Moreover, as an additional novelty, systems will also be evaluated on labels with different granularity in order to assess their robustness for different applications. To lower the entry barrier for participants, we developed an updated baseline system with several caveats to address these aforementioned problems. Results with our baseline system indicate that this research direction is promising and is possible to obtain a stronger SED system by using diverse domain training data with missing labels compared to training a SED system for each domain separately.
【15】 LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
标题: LibriTTS-P:具有说话风格和说话者身份的数据库,支持文本转语音和风格字幕
作者:Masaya Kawamura,Ryuichi Yamamoto,Yuma Shirahata,Takuya Hasumi,Kentaro Tachibana
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:我们介绍了LibriTTS-P,一个基于LibriTTS-R的新语料库,它包括话语级描述(即,提示)和说话者特征的说话者级别提示。我们采用一种混合的方法来构建提示注释:(1)手动注释,捕捉人类对说话者特征的感知和(2)说话风格的合成注释。与现有的英语提示语数据集相比,我们的语料库为LibriTTS-R的所有使用者提供了更多样化的提示语注释。实验结果表明,基于LibriTTS-P训练的TTS模型比使用传统数据集的模型具有更高的自然度。此外,样式字幕任务的结果表明,使用LibriTTS-P的模型生成的单词比使用传统数据集的模型准确2.5倍。我们的语料库LibriTTS-P可以在https: github.com line LibriTTS-P上找到。摘要:We introduce LibriTTS-P, a new corpus based on LibriTTS-R that includes utterance-level descriptions (i.e., prompts) of speaking style and speaker-level prompts of speaker characteristics. We employ a hybrid approach to construct prompt annotations: (1) manual annotations that capture human perceptions of speaker characteristics and (2) synthetic annotations on speaking style. Compared to existing English prompt datasets, our corpus provides more diverse prompt annotations for all speakers of LibriTTS-R. Experimental results for prompt-based controllable TTS demonstrate that the TTS model trained with LibriTTS-P achieves higher naturalness than the model using the conventional dataset. Furthermore, the results for style captioning tasks show that the model utilizing LibriTTS-P generates 2.5 times more accurate words than the model using a conventional dataset. Our corpus, LibriTTS-P, is available at https: github.com line LibriTTS-P.
【16】 Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
标题: 使用自我知识蒸馏指导框架级CSC对准
作者:Eungbeom Kim,Hantae Kim,Kyogu Lee
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:Transformer编码器采用连接主义时态分类(CTC)框架,广泛应用于自动语音识别(ASR)。然而,知识蒸馏(KD)的ASR显示了一个问题,教师和学生模型之间的不一致的框架级对齐,这最终阻碍了它从提高学生模型的性能。为了解决这个问题,本文介绍了一种自知识蒸馏(SKD)方法,在训练期间指导帧级对齐。与传统的教师模型和学生模型分离的方法相比,本研究提出了一种简单有效的方法,即共享编码器层,并将子模型作为学生模型。总体而言,我们的方法在提高资源效率和性能方面是有效的。我们还进行了实验分析的尖峰时间来说明,该方法通过减少对齐不一致提高性能。摘要:Transformer encoder with connectionist temporal classification (CTC) framework is widely used for automatic speech recognition (ASR). However, knowledge distillation (KD) for ASR displays a problem of disagreement between teacher-student models in frame-level alignment which ultimately hinders it from improving the student model's performance. In order to resolve this problem, this paper introduces a self-knowledge distillation (SKD) method that guides the frame-level alignment during the training time. In contrast to the conventional method using separate teacher and student models, this study introduces a simple and effective method sharing encoder layers and applying the sub-model as the student model. Overall, our approach is effective in improving both the resource efficiency as well as performance. We also conducted an experimental analysis of the spike timings to illustrate that the proposed method improves performance by reducing the alignment disagreement.
【17】 Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
标题: 探索儿童与成人二元互动中说话者扩大化的言语基础模型
作者:Anfeng Xu,Kevin Huang,Tiantian Feng,Lue Shen,Helen Tager-Flusberg,Shrikanth Narayanan
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:在大量数据集上训练的语音基础模型为解决具有挑战性的低资源语音理解(如儿童语音)提供了独特的机会。在这项工作中,我们探讨的能力,语音基础模型的儿童成人发言人日记。我们表明,示例性的基础模型可以实现39.5%和62.3%的相对减少,分别在日记错误率和扬声器混淆率,相比以前的扬声器日记化方法。此外,我们基准和评估的语音基础模型的扬声器diarization结果与不同的输入音频窗口大小,扬声器人口统计,和训练数据比。我们的研究结果突出了有前途的途径,理解和采用语音基础模型,以促进儿童语音理解。摘要:Speech foundation models, trained on vast datasets, have opened unique opportunities in addressing challenging low-resource speech understanding, such as child speech. In this work, we explore the capabilities of speech foundation models on child-adult speaker diarization. We show that exemplary foundation models can achieve 39.5% and 62.3% relative reductions in Diarization Error Rate and Speaker Confusion Rate, respectively, compared to previous speaker diarization methods. In addition, we benchmark and evaluate the speaker diarization results of the speech foundation models with varying the input audio window size, speaker demographics, and training data ratio. Our results highlight promising pathways for understanding and adopting speech foundation models to facilitate child speech understanding.
【18】 DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion
标题: DualVC 3:利用语言模型生成的伪上下文进行端到端低延迟流媒体语音转换
作者:Ziqian Ning,Shuai Wang,Pengcheng Zhu,Zhichao Wang,Jixun Yao,Lei Xie,Mengxiao Bi
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:流式语音转换由于其在实时应用中的潜力而变得越来越受欢迎。最近提出的DualVC 2已经实现了鲁棒和高质量的流式语音转换,延迟约为180 ms。然而,语音合成框架阻碍了端到端的优化,并且具有短块的自动语音识别(ASR)模型的不稳定性使得进一步减少延迟具有挑战性。为了解决这些问题,我们提出了一个端到端的模型,DualVC 3。通过使用与说话者无关的语义令牌来指导内容编码器的训练,消除了对ASR的依赖,并且模型可以在极小的块下运行,消除了级联错误。在内容编码器输出上训练语言模型,以通过迭代地预测未来帧来产生伪上下文,从而为解码器提供更多上下文信息以提高转换质量。实验结果表明,DualVC 3在主观和客观指标上达到了与DualVC 2相当的性能,延迟仅为50 ms。摘要:Streaming voice conversion has become increasingly popular for its potential in real-time applications. The recently proposed DualVC 2 has achieved robust and high-quality streaming voice conversion with a latency of about 180ms. Nonetheless, the recognition-synthesis framework hinders end-to-end optimization, and the instability of automatic speech recognition (ASR) model with short chunks makes it challenging to further reduce latency. To address these issues, we propose an end-to-end model, DualVC 3. With speaker-independent semantic tokens to guide the training of the content encoder, the dependency on ASR is removed and the model can operate under extremely small chunks, with cascading errors eliminated. A language model is trained on the content encoder output to produce pseudo context by iteratively predicting future frames, providing more contextual information for the decoder to improve conversion quality. Experimental results demonstrate that DualVC 3 achieves comparable performance to DualVC 2 in subjective and objective metrics, with a latency of only 50 ms.
【19】 Target Speaker Extraction with Curriculum Learning
标题: 通过课程学习提取目标说话者
作者:Yun Liu,Xuechen Liu,Xiaoxiao Miao,Junichi Yamagishi
备注:Accepted for presentation at Interspeech 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的方法,目标说话人提取(TSE)使用课程学习(CL)技术,解决区分目标说话人的声音从混合包含干扰扬声器的挑战。为了有效的训练,我们建议设计一个课程,选择子集的复杂性不断增加,如增加目标和干扰扬声器之间的相似性,并选择训练数据的战略。我们的CL策略包括使用预定义难度度量(例如性别,说话者相似性和信号失真比)的变体和使用TSE标准目标函数的变体,每种策略都旨在使模型逐渐暴露于更具挑战性的场景。在Libri2talker数据集上进行的全面测试表明,我们针对TSE的CL策略提高了性能,并且结果明显超过了没有CL的基线模型约1 dB。摘要:This paper presents a novel approach to target speaker extraction (TSE) using Curriculum Learning (CL) techniques, addressing the challenge of distinguishing a target speaker's voice from a mixture containing interfering speakers. For efficient training, we propose designing a curriculum that selects subsets of increasing complexity, such as increasing similarity between target and interfering speakers, and that selects training data strategically. Our CL strategies include both variants using predefined difficulty measures (e.g. gender, speaker similarity, and signal-to-distortion ratio) and ones using the TSE's standard objective function, each designed to expose the model gradually to more challenging scenarios. Comprehensive testing on the Libri2talker dataset demonstrated that our CL strategies for TSE improved the performance, and the results markedly exceeded baseline models without CL about 1 dB.
【20】 Dual-Pipeline with Low-Rank Adaptation for New Language Integration in Multilingual ASR
标题: 低等级自适应的双管道用于多语言ASB中的新语言集成
作者:Yerbolat Khassanov,Zhipeng Chen,Tianfeng Chen,Tze Yuang Chong,Wei Li,Jun Zhang,Lu Lu,Yuxuan Wang
备注:5 pages, 2 figures, 4 tables
链接:点击下载PDF文件
摘要:本文讨论了将新语言集成到预训练的多语言自动语音识别(mASR)系统中的挑战,特别是在现有语言的训练数据有限或不可用的情况下。该方法采用具有低秩自适应(LoRA)的双流水线。它维护两个数据流管道--一个用于现有语言,另一个用于新语言。主管道遵循mASR预训练参数的标准流程,而辅助管道还利用LoRA表示的语言特定参数和单独的输出解码器模块。重要的是,所提出的方法最大限度地减少现有语言的性能下降,并使语言无关的操作模式,促进了解码器选择策略。我们通过将预训练的Whisper模型扩展到来自FLEURS数据集的19种新语言来验证所提出方法的有效性摘要:This paper addresses challenges in integrating new languages into a pre-trained multilingual automatic speech recognition (mASR) system, particularly in scenarios where training data for existing languages is limited or unavailable. The proposed method employs a dual-pipeline with low-rank adaptation (LoRA). It maintains two data flow pipelines-one for existing languages and another for new languages. The primary pipeline follows the standard flow through the pre-trained parameters of mASR, while the secondary pipeline additionally utilizes language-specific parameters represented by LoRA and a separate output decoder module. Importantly, the proposed approach minimizes the performance degradation of existing languages and enables a language-agnostic operation mode, facilitated by a decoder selection strategy. We validate the effectiveness of the proposed method by extending the pre-trained Whisper model to 19 new languages from the FLEURS dataset
【21】 Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
作者:Lin Zhang,Xin Wang,Erica Cooper,Mireia Diez,Federico Landini,Nicholas Evans,Junichi Yamagishi
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文将欺骗日记定义为部分欺骗(PS)场景中的一个新任务。它的目的是确定什么时候欺骗,这不仅包括定位欺骗区域,而且还根据不同的欺骗方法对它们进行聚类。作为一个开创性的研究,在欺骗日记,我们专注于定义任务,建立评估指标,并提出了一个基准模型,即对策条件聚类(3C)模型。利用这个模型,我们首先探讨如何有效地训练对策,以支持欺骗日记使用三个标签计划。然后,我们利用欺骗本地化预测,以提高日记的性能。这第一项研究揭示了任务的高度复杂性,即使在每个音频文件只有一个扬声器和一个oracle数量的欺骗方法被认为是有限的情况下。我们的代码可以在https: github.com nii-yamagishilab PartialSpoof上找到。摘要:This paper defines Spoof Diarization as a novel task in the Partial Spoof (PS) scenario. It aims to determine what spoofed when, which includes not only locating spoof regions but also clustering them according to different spoofing methods. As a pioneering study in spoof diarization, we focus on defining the task, establishing evaluation metrics, and proposing a benchmark model, namely the Countermeasure-Condition Clustering (3C) model. Utilizing this model, we first explore how to effectively train countermeasures to support spoof diarization using three labeling schemes. We then utilize spoof localization predictions to enhance the diarization performance. This first study reveals the high complexity of the task, even in restricted scenarios where only a single speaker per audio file and an oracle number of spoofing methods are considered. Our code is available at https: github.com nii-yamagishilab PartialSpoof.
【22】 Broadband MEMS Microphone Arrays with Reduced Aperture Through 3D-Printed Waveguides
标题: 通过3D打印光路实现缩小口径的宽带微机电麦克风阵列
作者:Dennis Laurijssen,Walter Daems,Jan Steckel
链接:点击下载PDF文件
摘要:在本文中,我们提出了一种无源和成本效益的方法,用于增加超声MEMS麦克风阵列的频率范围时,使用波束形成技术。通过应用减小MEMS麦克风的声孔径的3D打印结构,我们可以创建规则间隔的麦克风阵列布局,由于MEMS元件的物理尺寸,元件间间距比印刷电路板上实现的要小得多。该方法允许结合麦克风阵列的超声传感器与波束成形技术的组合使用,而不会由于诸如声源定位或蝙蝠HRTF的仿真的应用中的栅瓣而产生混叠。摘要:In this paper we present a passive and cost-effective method for increasing the frequency range of ultrasound MEMS microphone arrays when using beamforming techniques. By applying a 3D-printed construction that reduces the acoustic aperture of the MEMS microphones we can create a regularly spaced microphone array layout with much smaller inter-element spacing than could be accomplished on a printed circuit board due to the physical size of the MEMS elements. This method allows the use of ultrasound sensors incorporating microphone arrays in combination with beamforming techniques without aliases due to grating lobes in applications such as sound source localization or the emulation of bat HRTFs.
【23】 Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based models
标题: 实现客观和可解释的言语障碍评估:CNN和基于转换器的模型的比较分析
作者:Malo Maisonneuve,Corinne Fredouille,Muriel Lalain,Alain Ghio,Virginie Woisard
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:头颈癌(HNC)严重影响患者的说话能力,影响他们的生活质量。用于评估病理性语音的常用指标是主观的,这促使需要自动化和无偏见的评估方法。本研究提出一个自我监督的Wav2Vec2为基础的模型与HNC患者的电话分类,以提高准确性和改善语音特征的歧视,为后续的可解释性的目的。探讨了预训练数据集、模型大小以及微调数据集和参数的影响。对不同语料库的评估揭示了Wav2Vec2架构的有效性,优于以前工作中使用的基于CNN的方法。与感知测量的相关性也肯定了受损语音分析的模型相关性。这项工作为临床医生更好地理解病理语言铺平了道路,通过利用复杂的自学语言表征。摘要:Head and Neck Cancers (HNC) significantly impact patients' ability to speak, affecting their quality of life. Commonly used metrics for assessing pathological speech are subjective, prompting the need for automated and unbiased evaluation methods. This study proposes a self-supervised Wav2Vec2-based model for phone classification with HNC patients, to enhance accuracy and improve the discrimination of phonetic features for subsequent interpretability purpose. The impact of pre-training datasets, model size, and fine-tuning datasets and parameters are explored. Evaluation on diverse corpora reveals the effectiveness of the Wav2Vec2 architecture, outperforming a CNN-based approach, used in previous work. Correlation with perceptual measures also affirms the model relevance for impaired speech analysis. This work paves the way for better understanding of pathological speech with interpretable approaches for clinicians, by leveraging complex self-learnt speech representations.
【24】 Towards Musically Informed Evaluation of Piano Transcription Models
标题: 钢琴抄写模型的音乐知情评估
作者:Patricia Hu,Lukáš Samuel Marták,Carlos Cancino-Chacón,Gerhard Widmer
链接:点击下载PDF文件
摘要:自动钢琴转录模型通常使用简单的逐帧或逐音符信息检索(IR)度量来评估。这样的基准度量不能提供对特定音乐方面的转录质量的洞察,例如输出的清晰度、动态或节奏精度,这些在表达性表现分析的上下文中是必不可少的。此外,近年来,MAESTRO已成为此类模型事实上的训练和评估数据集。然而,推理性能已被观察到大大恶化时,适用于分布外的数据,从而质疑的适用性和可靠性,转录输出从这些模型的特定MIR任务。在这项工作中,我们调查了三个国家的最先进的钢琴转录模型在两个实验中的性能。在第一个中,我们提出了各种音乐上知情的评价指标,与IR指标相比,提供了更详细的了解音乐质量的transmittance。在第二个实验中,我们比较了真实世界和干扰录音的推理性能,并强调了我们的指标可以帮助解释的音乐维度。我们的实验结果突出了现有的钢琴转录指标的弱点,并有助于更音乐的声音错误分析的转录输出。摘要:Automatic piano transcription models are typically evaluated using simple frame- or note-wise information retrieval (IR) metrics. Such benchmark metrics do not provide insights into the transcription quality of specific musical aspects such as articulation, dynamics, or rhythmic precision of the output, which are essential in the context of expressive performance analysis. Furthermore, in recent years, MAESTRO has become the de-facto training and evaluation dataset for such models. However, inference performance has been observed to deteriorate substantially when applied on out-of-distribution data, thereby questioning the suitability and reliability of transcribed outputs from such models for specific MIR tasks. In this work, we investigate the performance of three state-of-the-art piano transcription models in two experiments. In the first one, we propose a variety of musically informed evaluation metrics which, in contrast to the IR metrics, offer more detailed insight into the musical quality of the transcriptions. In the second experiment, we compare inference performance on real-world and perturbed audio recordings, and highlight musical dimensions which our metrics can help explain. Our experimental results highlight the weaknesses of existing piano transcription metrics and contribute to a more musically sound error analysis of transcription outputs.
【25】 TokSing: Singing Voice Synthesis based on Discrete Tokens
标题: TokSing:基于离散令牌的歌唱声音合成
作者:Yuning Wu,Chunlei zhang,Jiatong Shi,Yuxun Tang,Shan Yang,Qin Jin
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:语音合成的最新进展通过利用从自监督学习(SSL)模型中提取的离散令牌来见证显着的好处。与传统的连续Mel谱图相比,离散标记在中间表示中提供更高的存储效率和更大的可操作性。然而,当涉及到歌唱声音合成(SVS),实现更高层次的旋律表达提出了一个很大的挑战,利用离散令牌。在本文中,我们介绍了TokSing,一个基于离散的SVS系统配备了一个令牌配方,提供灵活的令牌混合。我们在离散化过程中观察到旋律退化,这促使我们将旋律信号与离散令牌集成,并在音乐编码器中采用专门设计的旋律增强策略。大量的实验表明,我们的TokSing实现了更好的性能对梅尔频谱图基线,同时提供了中间表示空间成本和收敛速度的优势。摘要:Recent advancements in speech synthesis witness significant benefits by leveraging discrete tokens extracted from self-supervised learning (SSL) models. Discrete tokens offer higher storage efficiency and greater operability in intermediate representations compared to traditional continuous Mel spectrograms. However, when it comes to singing voice synthesis(SVS), achieving higher levels of melody expression poses a great challenge for utilizing discrete tokens. In this paper, we introduce TokSing, a discrete-based SVS system equipped with a token formulator that offers flexible token blendings. We observe a melody degradation during discretization, prompting us to integrate a melody signal with the discrete token and incorporate a specially-designed melody enhancement strategy in the musical encoder. Extensive experiments demonstrate that our TokSing achieves better performance against the Mel spectrogram baselines while offering advantages in intermediate representation space cost and convergence speed.
【26】 Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models
标题: 迪夫-A-里夫:通过潜在扩散模型的音乐伴奏共同创作
作者:Javier Nistal,Marco Pasini,Cyran Aouameur,Maarten Grachten,Stefan Lattner
备注:8 pages, 2 figures, 3 tables
链接:点击下载PDF文件
摘要:深度生成模型的最新进展为音乐制作带来了新的机会,但也带来了挑战,例如高计算需求和有限的音频质量。此外,当前的系统经常仅依赖于文本输入,并且通常专注于制作完整的音乐作品,这与音乐制作中的现有工作流程不兼容。为了解决这些问题,我们引入了“Diff-A-Riff”,这是一种潜在的扩散模型,旨在生成适用于任何音乐背景的高质量乐器演奏。该模型通过音频参考、文本提示或两者提供控制,并产生48 kHz伪立体声音频,同时显著减少推理时间和内存使用。我们通过客观指标和主观听力测试展示了模型的能力,并在相应的网站上提供了大量的例子:sonycslparis.github.io diffariff-companion 摘要:Recent advancements in deep generative models present new opportunities for music production but also pose challenges, such as high computational demands and limited audio quality. Moreover, current systems frequently rely solely on text input and typically focus on producing complete musical pieces, which is incompatible with existing workflows in music production. To address these issues, we introduce "Diff-A-Riff," a Latent Diffusion Model designed to generate high-quality instrumental accompaniments adaptable to any musical context. This model offers control through either audio references, text prompts, or both, and produces 48kHz pseudo-stereo audio while significantly reducing inference time and memory usage. We demonstrate the model's capabilities through objective metrics and subjective listening tests, with extensive examples available on the accompanying website: sonycslparis.github.io diffariff-companion
【27】 Towards Unsupervised Speech Recognition Without Pronunciation Models
标题: 迈向没有发音模型的无监督语音识别
作者:Junrui Ni,Liming Wang,Yang Zhang,Kaizhi Qian,Heting Gao,Mark Hasegawa-Johnson,Chang D. Yoo
备注:This work has been submitted to the IEEE for possible publication
链接:点击下载PDF文件
摘要:监督自动语音识别(ASR)的最新进展取得了显着的性能,主要是由于越来越多的大型转录语音语料库的可用性。然而,大多数语言缺乏足够的成对语音和文本数据来有效地训练这些系统。在这篇文章中,我们解决了开发ASR系统没有配对的语音和文本语料库的挑战,提出了消除对音素词典的依赖。我们探索了一个新的研究方向:词级无监督自动语音识别。使用一个只包含高频英语单词的精选语音语料库,我们的系统在没有平行成绩单或甲骨文单词边界的情况下实现了近20%的单词错误率。此外,我们的实验表明,一个无监督的语音识别器可以出现从联合语音到语音和文本到文本掩蔽标记填充。这种创新的模型超越了以前使用直接分布匹配训练的无监督ASR模型的性能。摘要:Recent advancements in supervised automatic speech recognition (ASR) have achieved remarkable performance, largely due to the growing availability of large transcribed speech corpora. However, most languages lack sufficient paired speech and text data to effectively train these systems. In this article, we tackle the challenge of developing ASR systems without paired speech and text corpora by proposing the removal of reliance on a phoneme lexicon. We explore a new research direction: word-level unsupervised ASR. Using a curated speech corpus containing only high-frequency English words, our system achieves a word error rate of nearly 20% without parallel transcripts or oracle word boundaries. Furthermore, we experimentally demonstrate that an unsupervised speech recognizer can emerge from joint speech-to-speech and text-to-text masked token-infilling. This innovative model surpasses the performance of previous unsupervised ASR models trained with direct distribution matching.
【28】 CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
标题: CoLM-SVR:利用神经编解码语言建模进行多模式发音障碍语音重建
作者:Xueyuan Chen,Dongchao Yang,Dingdong Wang,Xixin Wu,Zhiyong Wu,Helen Meng
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:构音障碍语音重建(DSR)的目的是将构音障碍语音转换为正常语音。但它仍然存在说话人相似度低、韵律自然度差等问题。在本文中,我们提出了一个多模态DSR模型,利用神经编解码器的语言建模,以改善重建结果,特别是对说话人的相似性和韵律自然。我们提出的模型包括:(i)一个多模态内容编码器,用于从构音障碍语音中提取具有辅助视觉输入的鲁棒音素嵌入;(ii)一个说话人编解码器编码器,用于从构音障碍语音中提取说话人感知的编解码器并对其进行归一化,以提供原始音色和正常韵律;(iii)基于编解码器语言模型的语音解码器,用于基于所提取的音素嵌入和归一化编解码器来重构语音。在UAS语音语料库上的测试结果表明,该模型在说话人相似度和韵律自然度方面都有明显的提高。摘要:Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech. It still suffers from low speaker similarity and poor prosody naturalness. In this paper, we propose a multi-modal DSR model by leveraging neural codec language modeling to improve the reconstruction results, especially for the speaker similarity and prosody naturalness. Our proposed model consists of: (i) a multi-modal content encoder to extract robust phoneme embeddings from dysarthric speech with auxiliary visual inputs; (ii) a speaker codec encoder to extract and normalize the speaker-aware codecs from the dysarthric speech, in order to provide original timbre and normal prosody; (iii) a codec language model based speech decoder to reconstruct the speech based on the extracted phoneme embeddings and normalized codecs. Evaluations on the commonly used UASpeech corpus show that our proposed model can achieve significant improvements in terms of speaker similarity and prosody naturalness.
【29】 Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
标题: 在说话人嵌入中使用对抗扰动的非同步语音
作者:Rui Wang,Liping Chen,Kong AiK Lee,Zhen-Hua Ling
链接:点击下载PDF文件
摘要:语音匿名化已经被开发为用于通过用伪说话者的语音替换语音信号中的说话者的语音来保护隐私的技术,从而使原始语音属性从机器识别和人类感知中模糊。在本文中,我们专注于改变机器识别的语音属性,同时保留人类的感知。我们称之为异步语音匿名化。为此,语音生成框架结合扬声器解纠缠机制来生成匿名语音。通过对说话人嵌入施加对抗性扰动来改变说话人属性,同时通过控制扰动的强度来保留人类感知。在LibriSpeech数据集上进行的实验表明,说话人的属性被模糊,60.71%的处理后的话语保留了人类的感知。摘要:Voice anonymization has been developed as a technique for preserving privacy by replacing the speaker's voice in a speech signal with that of a pseudo-speaker, thereby obscuring the original voice attributes from machine recognition and human perception. In this paper, we focus on altering the voice attributes against machine recognition while retaining human perception. We referred to this as the asynchronous voice anonymization. To this end, a speech generation framework incorporating a speaker disentanglement mechanism is employed to generate the anonymized speech. The speaker attributes are altered through adversarial perturbation applied on the speaker embedding, while human perception is preserved by controlling the intensity of perturbation. Experiments conducted on the LibriSpeech dataset showed that the speaker attributes were obscured with their human perception preserved for 60.71% of the processed utterances.
【30】 FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
标题: FreeV:通过伪反梅尔过滤器为声码者提供免费午餐
作者:Yuanjun Lv,Hai Li,Ying Yan,Junhui Liu,Danming Xie,Lei Xie
备注:Accepted by InterSpeech 2024; 5 pages, 5 figures
链接:点击下载PDF文件
摘要:声码器从语音的声学特征中重构出语音波形,在现代文语转换系统中起着举足轻重的作用。频域GAN声码器(如Vocos和APNet2)最近取得了快速发展,在推理速度方面优于时域模型,同时实现了相当的音频质量。然而,这些频域声码器遭受大的参数大小,从而引入额外的存储器负担。受PriorGrad和SpecGrad的启发,我们采用伪逆来粗略估计振幅谱作为初始值。这种简单的初始化显著地减轻了对声码器的参数要求。基于APNet2和我们精简的幅度预测分支,我们提出了我们的FreeV,与其对应的APNet2相比,我们的FreeV在接近一半的参数下实现了1.8倍的推理速度提高。同时,我们的FreeV在再合成质量方面优于APNet2,标志着在追求实时,高保真语音合成方面向前迈出了一步。代码和检查点可在https: github.com BakerBunker FreeV上获得摘要:Vocoders reconstruct speech waveforms from acoustic features and play a pivotal role in modern TTS systems. Frequent-domain GAN vocoders like Vocos and APNet2 have recently seen rapid advancements, outperforming time-domain models in inference speed while achieving comparable audio quality. However, these frequency-domain vocoders suffer from large parameter sizes, thus introducing extra memory burden. Inspired by PriorGrad and SpecGrad, we employ pseudo-inverse to estimate the amplitude spectrum as the initialization roughly. This simple initialization significantly mitigates the parameter demand for vocoder. Based on APNet2 and our streamlined Amplitude prediction branch, we propose our FreeV, compared with its counterpart APNet2, our FreeV achieves 1.8 times inference speed improvement with nearly half parameters. Meanwhile, our FreeV outperforms APNet2 in resynthesis quality, marking a step forward in pursuing real-time, high-fidelity speech synthesis. Code and checkpoints is available at: https: github.com BakerBunker FreeV
【31】 Codecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio
标题: Codecfake:用于检测基于LLM的Deepfake音频的初始数据集
作者:Yi Lu,Yuankun Xie,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Zhiyong Wang,Xin Qi,Xuefei Liu,Yongwei Li,Yukun Liu,Xiaopeng Wang,Shuchen Shi
备注:Accepted by INTERSPEECH 2024. arXiv admin note: substantial text overlap with arXiv:2405.04880
链接:点击下载PDF文件
摘要:随着基于大语言模型(LLM)的deepfake音频的激增,迫切需要有效的检测方法。以前的deepfake音频生成方法通常涉及多步生成过程,最后一步使用声码器从手工特征预测波形。然而,基于LLM的音频在端到端生成过程中直接从离散神经编解码器生成,跳过了声码器处理的最后一步。这对当前基于声码器伪影的音频深度伪造检测(ADD)模型提出了重大挑战。为了有效地检测基于LLM的deepfake音频,我们专注于生成过程的核心,从神经编解码器到波形的转换。我们提出了Codecfake数据集,它是由七种代表性的神经编解码方法生成的。实验结果表明,与声码器训练的ADD模型相比,编解码器训练的ADD模型在Codecfake测试集上的平均等误率降低了41.406%。摘要:With the proliferation of Large Language Model (LLM) based deepfake audio, there is an urgent need for effective detection methods. Previous deepfake audio generation methods typically involve a multi-step generation process, with the final step using a vocoder to predict the waveform from handcrafted features. However, LLM-based audio is directly generated from discrete neural codecs in an end-to-end generation process, skipping the final step of vocoder processing. This poses a significant challenge for current audio deepfake detection (ADD) models based on vocoder artifacts. To effectively detect LLM-based deepfake audio, we focus on the core of the generation process, the conversion from neural codec to waveform. We propose Codecfake dataset, which is generated by seven representative neural codec methods. Experiment results show that codec-trained ADD models exhibit a 41.406% reduction in average equal error rate compared to vocoder-trained ADD models on the Codecfake test set.
【32】 FakeSound: Deepfake General Audio Detection
标题: FakeSound:Deepfake通用音频检测
作者:Zeyu Xie,Baihan Li,Xuenan Xu,Zheng Liang,Kai Yu,Mengyue Wu
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:随着音频生成技术的进步,生成模型可以产生高度逼真的音频。然而,deepfake一般音频的扩散可能会造成负面后果。因此,我们提出了一个新的任务,deepfake通用音频检测,旨在识别音频内容是否被操纵并定位deepfake区域。利用自动操作管道,提出了一个名为FakeSound的数据集,用于deepfake通用音频检测,样本可以在网站https: FakeSoundData.github.io上查看。人类在所有测试集上的平均二进制准确度始终低于0.6,这表明人类在识别deepfake音频时面临的困难,并肯定了FakeSound数据集的有效性。提出了一种利用通用音频预训练模型的深度伪造检测模型作为基准系统。实验结果表明,所提出的模型的性能超过了最先进的deepfake语音检测和人类测试。摘要:With the advancement of audio generation, generative models can produce highly realistic audios. However, the proliferation of deepfake general audio can pose negative consequences. Therefore, we propose a new task, deepfake general audio detection, which aims to identify whether audio content is manipulated and to locate deepfake regions. Leveraging an automated manipulation pipeline, a dataset named FakeSound for deepfake general audio detection is proposed, and samples can be viewed on website https: FakeSoundData.github.io. The average binary accuracy of humans on all test sets is consistently below 0.6, which indicates the difficulty humans face in discerning deepfake audio and affirms the efficacy of the FakeSound dataset. A deepfake detection model utilizing a general audio pre-trained model is proposed as a benchmark system. Experimental results demonstrate that the performance of the proposed model surpasses the state-of-the-art in deepfake speech detection and human testers.
【33】 CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
标题: 用于流媒体开放词汇关键词查找的符合ATC的音频文本嵌入
作者:Sichen Jin,Youngmoon Jung,Seungjin Lee,Jaeyoung Roh,Changwoo Han,Hoonyoung Cho
链接:点击下载PDF文件
摘要:本文介绍了一种新的方法,流开放词汇关键字发现(KWS)与基于文本的关键字注册。对于每个输入帧,所提出的方法使用连接主义时间分类(CTC)找到在帧处结束的最佳对准,并聚合帧级声学嵌入(AE)以获得更高级别(即,字符、单词或短语)AE,其与目标关键字文本的文本嵌入(TE)对齐。然后,我们计算聚集的AE和TE的相似度。据我们所知,这是第一次尝试动态对齐音频和关键字文本,以实现KWS的联合音频-文本嵌入。尽管以流式方式操作,但与仅具有155 K模型参数和时间复杂度为O(U)的解码算法的非流式方法相比,我们的方法在LibriPhrase数据集上实现了具有竞争力的性能,其中U是推理时目标关键字的长度。摘要:This paper introduces a novel approach for streaming openvocabulary keyword spotting (KWS) with text-based keyword enrollment. For every input frame, the proposed method finds the optimal alignment ending at the frame using connectionist temporal classification (CTC) and aggregates the frame-level acoustic embedding (AE) to obtain higher-level (i.e., character, word, or phrase) AE that aligns with the text embedding (TE) of the target keyword text. After that, we calculate the similarity of the aggregated AE and the TE. To the best of our knowledge, this is the first attempt to dynamically align the audio and the keyword text on-the-fly to attain the joint audio-text embedding for KWS. Despite operating in a streaming fashion, our approach achieves competitive performance on the LibriPhrase dataset compared to the non-streaming methods with a mere 155K model parameters and a decoding algorithm with time complexity O(U), where U is the length of the target keyword at inference time.
【34】 Can Large Language Models Understand Spatial Audio?
标题: 大型语言模型能理解空间音频吗?
作者:Changli Tang,Wenyi Yu,Guangzhi Sun,Xianzhao Chen,Tian Tan,Wei Li,Jun Zhang,Lu Lu,Zejun Ma,Yuxuan Wang,Chao Zhang
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:本文探讨了使大型语言模型(LLM)能够从多通道音频中理解空间信息,这是目前听觉LLM缺乏的技能。通过利用LLM先进的认知和推理能力,目的是通过音频增强对3D环境的理解。我们研究了3个空间音频任务:声源定位(SSL),远场语音识别(FSR)和定位通知语音提取(LSE),在每个任务中取得了显着的进展。对于SSL,我们的方法在Spatial LibriSpeech数据集上实现了2.70 ^{ circ}$的MAE,大大超过了之前的基准约6.60 ^{ circ}$。此外,我们的模型可以采用空间线索来提高FSR的准确性,并通过文本提示选择性地关注来自指定方向的声音来执行LSE,即使是在重叠的语音中。这些发现突出了适应LLM以掌握物理音频概念的潜力,为3D环境中基于LLM的代理铺平了道路。摘要:This paper explores enabling large language models (LLMs) to understand spatial information from multichannel audio, a skill currently lacking in auditory LLMs. By leveraging LLMs' advanced cognitive and inferential abilities, the aim is to enhance understanding of 3D environments via audio. We study 3 spatial audio tasks: sound source localization (SSL), far-field speech recognition (FSR), and localisation-informed speech extraction (LSE), achieving notable progress in each task. For SSL, our approach achieves an MAE of $2.70^{ circ}$ on the Spatial LibriSpeech dataset, substantially surpassing the prior benchmark of about $6.60^{ circ}$. Moreover, our model can employ spatial cues to improve FSR accuracy and execute LSE by selectively attending to sounds originating from a specified direction via text prompts, even amidst overlapping speech. These findings highlight the potential of adapting LLMs to grasp physical audio concepts, paving the way for LLM-based agents in 3D environments.
【35】 Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
标题: 探索自我监督多视图对比学习用于有限注释的语音情感识别
作者:Bulat Khaertdinov,Pedro Jeuris,Annanda Sousa,Enrique Hortal
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:深度和自我监督学习(SSL)的最新进展使语音情感识别(SER)性能得到了大幅改善,达到了前所未有的水平。然而,获得足够数量的准确标记的数据来训练或微调模型仍然是一项昂贵且具有挑战性的任务。在本文中,我们提出了一种多视图SSL预训练技术,可应用于各种语音表示,包括由大型语音模型生成的语音表示,以提高在注释有限的情况下的SER性能。我们的实验,基于wav2vec 2.0,频谱和语言特征,表明所提出的框架提高SER性能,在未加权平均召回率高达10%,在设置非常稀疏的数据注释。摘要:Recent advancements in Deep and Self-Supervised Learning (SSL) have led to substantial improvements in Speech Emotion Recognition (SER) performance, reaching unprecedented levels. However, obtaining sufficient amounts of accurately labeled data for training or fine-tuning the models remains a costly and challenging task. In this paper, we propose a multi-view SSL pre-training technique that can be applied to various representations of speech, including the ones generated by large speech models, to improve SER performance in scenarios where annotations are limited. Our experiments, based on wav2vec 2.0, spectral and paralinguistic features, demonstrate that the proposed framework boosts the SER performance, by up to 10% in Unweighted Average Recall, in settings with extremely sparse data annotations.
【36】 Flexible Music-Conditioned Dance Generation with Style Description Prompts
标题: 具有风格描述的灵活音乐条件舞蹈生成
作者:Hongsong Wang,Yin Zhu,Xin Geng
链接:点击下载PDF文件
摘要:舞蹈作为一种艺术形式和表现形式,在人类文化中占有重要地位,但舞蹈的创作仍然是一项具有挑战性的任务。大多数舞蹈生成方法主要依赖于音乐,很少考虑音乐风格或流派等内在属性。在这项工作中,我们介绍了灵活的舞蹈生成与风格描述符(DGSDP),一个基于扩散的框架,适合于多样化的舞蹈生成任务,充分利用音乐风格的语义。该框架的核心组件是音乐条件风格感知扩散(MCSAD),它包括一个基于transformer的网络和一个音乐风格调制模块。MCSAD将音乐条件和风格描述提示巧妙地集成到舞蹈生成框架中,确保生成的舞蹈与音乐内容和风格一致。为了便于灵活的舞蹈生成和适应不同的任务,时空掩蔽策略有效地应用在向后扩散过程中。所提出的框架成功地生成逼真的舞蹈序列,准确地与音乐的各种任务,如长期一代,舞蹈中间,舞蹈修补等,我们希望这项工作有可能激发舞蹈的生成和创作,在娱乐,艺术和教育的应用前景。摘要:Dance plays an important role as an artistic form and expression in human culture, yet the creation of dance remains a challenging task. Most dance generation methods primarily rely solely on music, seldom taking into consideration intrinsic attributes such as music style or genre. In this work, we introduce Flexible Dance Generation with Style Description Prompts (DGSDP), a diffusion-based framework suitable for diversified tasks of dance generation by fully leveraging the semantics of music style. The core component of this framework is Music-Conditioned Style-Aware Diffusion (MCSAD), which comprises a Transformer-based network and a music Style Modulation module. The MCSAD seemly integrates music conditions and style description prompts into the dance generation framework, ensuring that generated dances are consistent with the music content and style. To facilitate flexible dance generation and accommodate different tasks, a spatial-temporal masking strategy is effectively applied in the backward diffusion process. The proposed framework successfully generates realistic dance sequences that are accurately aligned with music for a variety of tasks such as long-term generation, dance in-betweening, dance inpainting, and etc. We hope that this work has the potential to inspire dance generation and creation, with promising applications in entertainment, art, and education.
【37】 VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
标题: WAL-E R:通过单调对齐实现稳健高效的Zero-Shot文本到语音合成
作者:Bing Han,Long Zhou,Shujie Liu,Sanyuan Chen,Lingwei Meng,Yanming Qian,Yanqing Liu,Sheng Zhao,Jinyu Li,Furu Wei
备注:15 pages, 5 figures
链接:点击下载PDF文件
摘要:在离散神经音频编解码器的帮助下,大语言模型(LLM)越来越被认为是一种有前途的zero-shot文本到语音(TTS)合成方法。然而,基于采样的解码策略给生成带来了惊人的多样性,但也带来了诸如错别字、遗漏和重复的鲁棒性问题。此外,音频的高采样率也给自回归的推理过程带来了巨大的计算开销。为了解决这些问题,我们提出了VALL-E R,一个强大的和高效的zero-shot TTS系统,建立在VALL-E的基础上。具体来说,我们引入了一个音素单调对齐策略,以加强音素和声学序列之间的连接,确保更精确的对齐,通过约束声学令牌,以匹配其相关的音素。此外,我们采用了一种编解码器合并的方法来下采样的离散代码在浅量化层,从而加快解码速度,同时保持高质量的语音输出。得益于这些策略,VALL-E R获得了音素上的相似性,并通过接近地面真值的WER来展示其强大的鲁棒性。此外,它需要更少的自回归步骤,在推理过程中减少了60%以上的时间。这项研究有可能应用于有意义的项目,包括为失语症患者创造言语。音频样本可在以下网址获取:https: aka.ms valler。摘要:With the help of discrete neural audio codecs, large language models (LLM) have increasingly been recognized as a promising methodology for zero-shot Text-to-Speech (TTS) synthesis. However, sampling based decoding strategies bring astonishing diversity to generation, but also pose robustness issues such as typos, omissions and repetition. In addition, the high sampling rate of audio also brings huge computational overhead to the inference process of autoregression. To address these issues, we propose VALL-E R, a robust and efficient zero-shot TTS system, building upon the foundation of VALL-E. Specifically, we introduce a phoneme monotonic alignment strategy to strengthen the connection between phonemes and acoustic sequence, ensuring a more precise alignment by constraining the acoustic tokens to match their associated phonemes. Furthermore, we employ a codec-merging approach to downsample the discrete codes in shallow quantization layer, thereby accelerating the decoding speed while preserving the high quality of speech output. Benefiting from these strategies, VALL-E R obtains controllablity over phonemes and demonstrates its strong robustness by approaching the WER of ground truth. In addition, it requires fewer autoregressive steps, with over 60% time reduction during inference. This research has the potential to be applied to meaningful projects, including the creation of speech for those affected by aphasia. Audio samples will be available at: https: aka.ms valler.
【38】 Zero-Shot Fake Video Detection by Audio-Visual Consistency
标题: 利用视听一致性进行Zero-Shot假视频检测
作者:Xiaolou Li,Zehua Liu,Chen Chen,Lantian Li,Li Guo,Dong Wang
备注:to be published in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:最近的研究主张检测假视频作为一类检测任务,基于真实数据的音频和视觉模态之间的一致性比假数据的一致性更重要的假设。这种方法仅依赖于真实的视听数据,同时不需要伪造的对应数据,因此被描述为“零射击”(zero-shot)检测范式。本文介绍了一种新的zero-shot检测方法锚定在音频和视频的内容一致性。通过使用预先训练的ASR和VSR模型,我们分别识别音频和视频内容序列。然后,计算两个序列之间的编辑距离,以评估声称的视频是否是真实的。实验结果表明,与基于语义一致性和时间一致性的两种主流方法相比,我们的方法在各种deepfake技术中实现了卓越的泛化能力,并对视听干扰表现出较强的鲁棒性。最后,通过简单地整合这三个系统的决策分数,可以实现最先进的性能增益。摘要:Recent studies have advocated the detection of fake videos as a one-class detection task, predicated on the hypothesis that the consistency between audio and visual modalities of genuine data is more significant than that of fake data. This methodology, which solely relies on genuine audio-visual data while negating the need for forged counterparts, is thus delineated as a zero-shot' detection paradigm. This paper introduces a novel zero-shot detection approach anchored in content consistency across audio and video. By employing pre-trained ASR and VSR models, we recognize the audio and video content sequences, respectively. Then, the edit distance between the two sequences is computed to assess whether the claimed video is genuine. Experimental results indicate that, compared to two mainstream approaches based on semantic consistency and temporal consistency, our approach achieves superior generalizability across various deepfake techniques and demonstrates strong robustness against audio-visual perturbations. Finally, state-of-the-art performance gains can be achieved by simply integrating the decision scores of these three systems.
【39】 SEBN Adapter: Parametric Efficient Domain Adaptation for Speaker Recognition
标题: SEBN适配器:用于说话人识别的参数高效域自适应
作者:Tianhao Wang,Lantian Li,Dong Wang
备注:to be published in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:在一个新的领域部署一个经过良好优化的预先训练的说话人识别模型通常会导致性能显著下降。虽然微调是一种常用的解决方案,但它需要大量的自适应数据,并且参数效率低下,这使得它对于具有有限数据可用于模型自适应的现实世界应用来说是不切实际的。从自监督预训练模型中适配器的成功中汲取灵感,本文介绍了SE BN适配器来解决这一挑战。通过冻结核心扬声器编码器和调整特征映射的权重和激活分布,我们引入了一种新的适配器,利用可训练的挤压和激励(SE)块和批量归一化(BN)层,称为SE BN适配器。我们的实验使用VoxCeleb进行预训练,使用CN-Celeb的4种类型进行适应,表明SE BN适配器在基线上提供了显着的性能改进,并通过调整仅1%的参数与香草微调方法竞争。摘要:Deploying a well-optimized pre-trained speaker recognition model in a new domain often leads to a significant decline in performance. While fine-tuning is a commonly employed solution, it demands ample adaptation data and suffers from parameter inefficiency, rendering it impractical for real-world applications with limited data available for model adaptation. Drawing inspiration from the success of adapters in self-supervised pre-trained models, this paper introduces a SE BN adapter to address this challenge. By freezing the core speaker encoder and adjusting the feature maps' weights and activation distributions, we introduce a novel adapter utilizing trainable squeeze-and-excitation (SE) blocks and batch normalization (BN) layers, termed SE BN adapter. Our experiments, conducted using VoxCeleb for pre-training and 4 genres from CN-Celeb for adaptation, demonstrate that the SE BN adapter offers significant performance improvement over the baseline and competes with the vanilla fine-tuning approach by tuning just 1% of the parameters.
【40】 PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
标题: PRoDeliberation:端到端口语理解的并行稳健审议
作者:Trang Le,Daniel Lazar,Suyoun Kim,Shan Jiang,Duc Le,Adithya Sagar,Aleksandr Livshits,Ahmed Aly,Akshat Shrivastava
链接:点击下载PDF文件
摘要:口语理解(SLU)是语音助手的关键组件;它包括将语音转换为语义解析以执行任务。以前的工作已经探索了端到端模型,以提高SLU模型的质量和鲁棒性,但这些模型仍然是自回归的,导致更高的延迟。在这项工作中,我们介绍PRoDeliberation,一种新的方法,利用连接时间分类为基础的解码策略,以及去噪目标训练强大的非自回归审议模型。我们表明,PRoDeliberation实现了并行解码的延迟降低(自回归模型的2- 10倍改进),同时保留了纠正自回归审议系统的自动语音识别(ASR)误译的能力。我们进一步表明,去噪训练的设计使PRoDeliberation能够克服小型ASR设备的局限性,并且我们对系统每个组件的必要性进行了分析。摘要:Spoken Language Understanding (SLU) is a critical component of voice assistants; it consists of converting speech to semantic parses for task execution. Previous works have explored end-to-end models to improve the quality and robustness of SLU models with Deliberation, however these models have remained autoregressive, resulting in higher latencies. In this work we introduce PRoDeliberation, a novel method leveraging a Connectionist Temporal Classification-based decoding strategy as well as a denoising objective to train robust non-autoregressive deliberation models. We show that PRoDeliberation achieves the latency reduction of parallel decoding (2-10x improvement over autoregressive models) while retaining the ability to correct Automatic Speech Recognition (ASR) mistranscriptions of autoregressive deliberation systems. We further show that the design of the denoising training allows PRoDeliberation to overcome the limitations of small ASR devices, and we provide analysis on the necessity of each component of the system.
【41】 EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
标题: SEARCH Sphere-TTC:通过球形情感载体进行情感风格和强度建模,用于可控情感文本到语音
作者:Deok-Hyeon Cho,Hyung-Seok Oh,Seung-Bin Kim,Sang-Hoon Lee,Seong-Whan Lee
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:尽管在情感文本到语音(TTS)领域取得了快速进展,但最近的研究主要集中在模仿特定情感的平均风格上。因此,操纵语音情感的能力仍然被限制在几个预定义的标签上,从而影响了反映情感细微变化的能力。在本文中,我们提出了一种新的语音合成方法--球面TTS,它通过使用一个球形情感向量来控制合成语音的情感风格和强度,从而合成具有表达性的情感语音。在没有任何人类注释的情况下,我们使用唤醒、效价和支配伪标签通过笛卡尔球面变换来模拟情感的复杂性质。此外,我们提出了一个双条件对抗网络,以提高生成的语音质量,反映多方面的特点。实验结果表明,该模型能够控制情绪的风格和强度与高质量的表达语音。摘要:Despite rapid advances in the field of emotional text-to-speech (TTS), recent studies primarily focus on mimicking the average style of a particular emotion. As a result, the ability to manipulate speech emotion remains constrained to several predefined labels, compromising the ability to reflect the nuanced variations of emotion. In this paper, we propose EmoSphere-TTS, which synthesizes expressive emotional speech by using a spherical emotion vector to control the emotional style and intensity of the synthetic speech. Without any human annotation, we use the arousal, valence, and dominance pseudo-labels to model the complex nature of emotion via a Cartesian-spherical transformation. Furthermore, we propose a dual conditional adversarial network to improve the quality of generated speech by reflecting the multi-aspect characteristics. The experimental results demonstrate the model ability to control emotional style and intensity with high-quality expressive speech.
【42】 PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
标题: PolySpeech:探索统一的多任务语音模型,以与单任务模型竞争
作者:Runyan Yang,Huibao Yang,Xiqing Zhang,Tiantian Ye,Ying Liu,Yingying Gao,Shilei Zhang,Chao Deng,Junlan Feng
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:最近,已经尝试将各种语音处理任务集成到统一的模型中。然而,很少有以前的工作直接表明,联合优化的不同任务的多任务语音模型有积极的影响,个别任务的性能。本文提出了一个多任务语音模型PolySpeech,它支持语音识别、语音合成和两个语音分类任务。PolySpeech以多模态语言模型为核心结构,以语义表征为语音输入。我们将语义语音嵌入标记化和语音重建方法引入PolySpeech,从而为任何给定的说话者有效地生成高质量的语音。与单任务模型相比,PolySpeech在各种任务中表现出竞争力。在我们的实验中,多任务优化实现了与单任务优化相当的性能,并且对特定任务特别有益。摘要:Recently, there have been attempts to integrate various speech processing tasks into a unified model. However, few previous works directly demonstrated that joint optimization of diverse tasks in multitask speech models has positive influence on the performance of individual tasks. In this paper we present a multitask speech model -- PolySpeech, which supports speech recognition, speech synthesis, and two speech classification tasks. PolySpeech takes multi-modal language model as its core structure and uses semantic representations as speech inputs. We introduce semantic speech embedding tokenization and speech reconstruction methods to PolySpeech, enabling efficient generation of high-quality speech for any given speaker. PolySpeech shows competitiveness across various tasks compared to single-task models. In our experiments, multitask optimization achieves performance comparable to single-task optimization and is especially beneficial for specific tasks.
【43】 The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
标题: Interspeech 2024年使用离散单元的语音处理挑战
作者:Xuankai Chang,Jiatong Shi,Jinchuan Tian,Yuning Wu,Yuxun Tang,Yihan Wu,Shinji Watanabe,Yossi Adi,Xie Chen,Qin Jin
备注:This manuscript has been accepted by Interspeech2024
链接:点击下载PDF文件
摘要:以离散单元表示语音和音频信号已经成为传统高维特征向量的一种引人注目的替代方案。许多研究已经强调了离散单元在各种应用中的功效,例如语音压缩和恢复,语音识别和语音生成。为了促进这一领域的探索,我们引入了Interspeech 2024挑战赛,该挑战赛专注于使用离散单元的新语音处理基准。它包括三个关键的任务,即多语种自动语音识别,文本到语音,和歌声合成,并旨在评估这些任务中的离散单元的潜在适用性。本文概述了挑战设计和基线描述。我们还整理了基线和选定的提交系统,以及初步研究结果,为这个不断发展的领域的未来研究提供了宝贵的贡献。摘要:Representing speech and audio signals in discrete units has become a compelling alternative to traditional high-dimensional feature vectors. Numerous studies have highlighted the efficacy of discrete units in various applications such as speech compression and restoration, speech recognition, and speech generation. To foster exploration in this domain, we introduce the Interspeech 2024 Challenge, which focuses on new speech processing benchmarks using discrete units. It encompasses three pivotal tasks, namely multilingual automatic speech recognition, text-to-speech, and singing voice synthesis, and aims to assess the potential applicability of discrete units in these tasks. This paper outlines the challenge designs and baseline descriptions. We also collate baseline and selected submission systems, along with preliminary findings, offering valuable contributions to future research in this evolving field.
【44】 FastAST: Accelerating Audio Spectrogram Transformer via Token Merging and Cross-Model Knowledge Distillation
标题: FastAST:通过令牌合并和跨模型知识提炼加速音频频谱图Transformer
作者:Swarup Ranjan Behera,Abhishek Dhiman,Karthik Gowda,Aalekhya Satya Narayani
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:音频分类模型,特别是音频频谱图Transformer(AST),在高效的音频分析中起着至关重要的作用。然而,在不影响精度的情况下优化其效率仍然是一个挑战。在本文中,我们介绍FastAST,一个框架,集成令牌合并(ToMe)到AST框架。FastAST通过合并音频频谱图中的相似标记来提高推理速度,而无需进行大量的重新训练。此外,在训练过程中,FastAST带来了显着的速度提高。实验表明,FastAST可以提高音频分类吞吐量,同时对准确性的影响最小。为了减轻准确性的影响,我们将跨模型知识蒸馏(CMKD)集成到FastAST框架中。与AST相比,将ToMe和CMKD集成到AST中可以提高准确性,同时保持更快的推理速度。FastAST代表了向实时、资源高效的音频分析迈出的一步。摘要:Audio classification models, particularly the Audio Spectrogram Transformer (AST), play a crucial role in efficient audio analysis. However, optimizing their efficiency without compromising accuracy remains a challenge. In this paper, we introduce FastAST, a framework that integrates Token Merging (ToMe) into the AST framework. FastAST enhances inference speed without requiring extensive retraining by merging similar tokens in audio spectrograms. Furthermore, during training, FastAST brings about significant speed improvements. The experiments indicate that FastAST can increase audio classification throughput with minimal impact on accuracy. To mitigate the accuracy impact, we integrate Cross-Model Knowledge Distillation (CMKD) into the FastAST framework. Integrating ToMe and CMKD into AST results in improved accuracy compared to AST while maintaining faster inference speeds. FastAST represents a step towards real-time, resource-efficient audio analysis.
【45】 Pre-training Feature Guided Diffusion Model for Speech Enhancement
标题: 用于语音增强的预训练特征引导扩散模型
作者:Yiyuan Yang,Niki Trigoni,Andrew Markham
备注:Accepted by Interspeech 2024 Conference
链接:点击下载PDF文件
摘要:语音增强显著提高了嘈杂环境中语音的清晰度和可懂度,改善了沟通和聆听体验。在本文中,我们介绍了一种新的预训练特征引导的扩散模型为有效的语音增强量身定制,解决现有的歧视和生成模型的局限性。通过将频谱特征集成到变分自动编码器(VAE)中,并在反向过程中利用预先训练的特征进行指导,再加上利用确定性离散积分方法(DDIM)来简化采样步骤,我们的模型提高了效率和语音增强质量。在两个具有不同SNR的公共数据集上展示了最先进的结果,我们的模型在效率和鲁棒性方面优于其他基线。该方法不仅优化了性能,而且提高了实际部署能力,而不增加计算需求。摘要:Speech enhancement significantly improves the clarity and intelligibility of speech in noisy environments, improving communication and listening experiences. In this paper, we introduce a novel pretraining feature-guided diffusion model tailored for efficient speech enhancement, addressing the limitations of existing discriminative and generative models. By integrating spectral features into a variational autoencoder (VAE) and leveraging pre-trained features for guidance during the reverse process, coupled with the utilization of the deterministic discrete integration method (DDIM) to streamline sampling steps, our model improves efficiency and speech enhancement quality. Demonstrating state-of-the-art results on two public datasets with different SNRs, our model outshines other baselines in efficiency and robustness. The proposed method not only optimizes performance but also enhances practical deployment capabilities, without increasing computational demands.
机器翻译,仅供参考
