本文经arXiv每日学术速递授权转载
【1】 Physics and geometry informed neural operator network with application to acoustic scattering
标题: 物理和几何知识的神经操作网络及其在声散射中的应用
作者:Siddharth Nair,Timothy F. Walsh,Greg Pickrell,Fabio Semperlotti
备注:20 pages of main text, 9 figures
链接:点击下载PDF文件
【2】 Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
标题: 音频曼巴:音频表示学习的双向状态空间模型
作者:Mehmet Hamza Erol,Arda Senocak,Jiu Feng,Joon Son Chung
备注:Code is available at this https URL
链接:点击下载PDF文件
【3】 ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings
标题: ASoBO:为会议中的远距离发言人进行细心的束流器选择
作者:Theo Mariotte,Anthony Larcher,Silvio Montresor,Jean-Hugh Thomas
备注:5 pages, 2 figures, 2 tables, accepted at Interspeech 2024
链接:点击下载PDF文件
【4】 Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
标题: 使用MaskAutoEncoder进行广义假音频检测的以学生为中心的学习
作者:Xiaopeng Wang,Ruibo Fu,Zhengqi Wen,Zhiyong Wang,Yuankun Xie,Yukun Liu,Jianhua Tao,Xuefei Liu,Yongwei Li,Xin Qi,Yi Lu,Shuchen Shi
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【5】 Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion strategy
标题: 广义源跟踪:检测具有真实重点和虚假分散策略的新型音频Deepfake算法
作者:Yuankun Xie,Ruibo Fu,Zhengqi Wen,Zhiyong Wang,Xiaopeng Wang,Haonnan Cheng,Long Ye,Jianhua Tao
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【6】 Generalized Fake Audio Detection via Deep Stable Learning
标题: 通过深度稳定学习进行广义假音频检测
作者:Zhiyong Wang,Ruibo Fu,Zhengqi Wen,Yuankun Xie,Yukun Liu,Xiaopeng Wang,Xuefei Liu,Yongwei Li,Jianhua Tao,Yi Lu,Xin Qi,Shuchen Shi
备注:accepted by INTERSPEECH2024
链接:点击下载PDF文件
【7】 A Frame-based Attention Interpretation Method for Relevant Acoustic Feature Extraction in Long Speech Depression Detection
标题: 长言语抑郁检测中相关声学特征提取的基于框架的注意力解释方法
作者:Qingkun Deng,Saturnino Luz,Sofia de la Fuente Garcia
备注:5 pages, 3 figures. arXiv admin note: substantial text overlap with arXiv:2309.13476
链接:点击下载PDF文件
【8】 StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
标题: StreamSpeech:具有多任务学习的语音同步翻译
作者:Shaolei Zhang,Qingkai Fang,Shoutao Guo,Zhengrui Ma,Min Zhang,Yang Feng
备注:Accepted to ACL 2024 main conference, Project Page: this https URL
链接:点击下载PDF文件
【9】 Dataset-Distillation Generative Model for Speech Emotion Recognition
标题: 语音情感识别的数据集蒸馏生成模型
作者:Fabian Ritter-Gutierrez,Kuan-Po Huang,Jeremy H. M Wong,Dianwen Ng,Hung-yi Lee,Nancy F. Chen,Eng Siong Chng
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
【10】 AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
标题: AVFF:用于视频Deepfake检测的视听特征融合
作者:Trevine Oorloff,Surya Koppisetti,Nicolò Bonettini,Divyaraj Solanki,Ben Colman,Yaser Yacoob,Ali Shahriyari,Gaurav Bharaj
备注:Accepted to CVPR 2024
链接:点击下载PDF文件
【11】 Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
标题: 用双解码产品量化变分自动编码器解决大码本语音令牌器的索引崩溃
作者:Haohan Guo,Fenglong Xie,Dongchao Yang,Hui Lu,Xixin Wu,Helen Meng
链接:点击下载PDF文件
【12】 LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
标题: LiveSpeech:通过音频离散码的自回归建模的低延迟Zero-Shot文本到语音
作者:Trung Dang,David Aponte,Dung Tran,Kazuhito Koishida
链接:点击下载PDF文件
【13】 Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
标题: 具有自我监督蒸馏的无文本声学模型用于噪音稳健的表达性语音到语音翻译
作者:Min-Jae Hwang,Ilia Kulikov,Benjamin Peloquin,Hongyu Gong,Peng-Jen Chen,Ann Lee
备注:Accepted to ACL 2024 (findings)
链接:点击下载PDF文件
【14】 Sequence-to-sequence models in peer-to-peer learning: A practical application
标题: 点对点学习中的序列到序列模型:实际应用
作者:Robert Šajina,Ivo Ipšić
链接:点击下载PDF文件
【15】 The PESQetarian: On the Relevance of Goodhart's Law for Speech Enhancement
标题: PESQetarian:论古德哈特定律与言语增强的相关性
作者:Danilo de Oliveira,Simon Welker,Julius Richter,Timo Gerkmann
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
【16】 Enhancing CTC-based speech recognition with diverse modeling units
标题: 利用不同的建模单元增强基于ATC的语音识别
作者:Shiyi Han,Zhihong Lei,Mingbin Xu,Xingyu Na,Zhen Huang
链接:点击下载PDF文件
【17】 RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification
标题: RevRIR:使用对比学习联合混响语音和房间脉冲响应嵌入并应用于房间形状分类
作者:Jacob Bitterman,Daniel Levi,Hilel Hagai Diamandi,Sharon Gannot,Tal Rosenwein
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【18】 4D ASR: Joint Beam Search Integrating CTC, Attention, Transducer, and Mask Predict Decoders
标题: 4D ASB:集成了CTC、注意力、传感器和屏蔽预测解码器的联合射束搜索
作者:Yui Sudo,Muhammad Shakeel,Yosuke Fukumoto,Brian Yan,Jiatong Shi,Yifan Peng,Shinji Watanabe
备注:submitted to IEEEACM Transactions on Audio Speech and Language Processing
链接:点击下载PDF文件
【19】 SYN2REAL: Leveraging Task Arithmetic for Mitigating Synthetic-Real Discrepancies in ASR Domain Adaptation
标题: SY 2 REAL:利用任务算法缓解SVR域自适应中的合成-真实差异
作者:Hsuan Su,Hua Farn,Shang-Tse Chen,Hung-yi Lee
链接:点击下载PDF文件
【20】 USM RNN-T model weights binarization
标题: USM RNN-T模型权重二值化
作者:Oleg Rybakov,Dmitriy Serdyuk,CJ Zheng
链接:点击下载PDF文件
【21】 ConPCO: Preserving Phoneme Characteristics for Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
标题: ConPCO:保留音素特征以进行自动发音评估,利用对比有序规则化
作者:Bi-Cheng Yan,Wei-Cheng Chao,Jiun-Ting Li,Yi-Cheng Wang,Hsin-Wei Wang,Meng-Shin Lin,Berlin Chen
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【22】 Keyword-Guided Adaptation of Automatic Speech Recognition
标题: 关键词引导的自动语音识别自适应
作者:Aviv Shamsian,Aviv Navon,Neta Glazer,Gill Hetz,Joseph Keshet
备注:Accepted to InterSpeech 2024
链接:点击下载PDF文件
【23】 Selfsupervised learning for pathological speech detection
标题: 病态语音检测的自我监督学习
作者:Shakeel Ahmad Sheikh
备注:in Intersection of Book Chapter in Machine Leanring and Computational Social Sciences CRC (in progress) 2024
链接:点击下载PDF文件
【24】 A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
标题: 一种固定端到端ASC系统标记器中词汇量大小的成本最小化方法
作者:Sunil Kumar Kopparapu,Ashish Panda
备注:5 pages, 4 figures
链接:点击下载PDF文件
标题: PESQetarian:论古德哈特定律与言语增强的相关性
作者:Danilo de Oliveira,Simon Welker,Julius Richter,Timo Gerkmann
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
【2】 Enhancing CTC-based speech recognition with diverse modeling units
标题: 利用不同的建模单元增强基于ATC的语音识别
作者:Shiyi Han,Zhihong Lei,Mingbin Xu,Xingyu Na,Zhen Huang
链接:点击下载PDF文件
【3】 Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
标题: 使用分层令牌-语义音频Transformer架构的多麦克风语音情感识别
作者:Ohad Cohen,Gershon Hazan,Sharon Gannot
链接:点击下载PDF文件
【4】 Reference Channel Selection by Multi-Channel Masking for End-to-End Multi-Channel Speech Enhancement
标题: 端到端多通道语音增强的多通道掩蔽参考通道选择
作者:Wang Dai,Xiaofei Li,Archontis Politis,Tuomas Virtanen
备注:Accepted by EUSIPCO 2024
链接:点击下载PDF文件
【5】 CoLLAB: A Collaborative Approach for Multilingual Abuse Detection
标题: CoLLAB:多语言滥用检测的协作方法
作者:Orchid Chetia Phukan,Yashasvi Chaurasia,Arun Balaji Buduru,Rajesh Sharma
链接:点击下载PDF文件
【6】 Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
标题: 再次对话:通过分段级发言人重新分配改进会议转录系统
作者:Christoph Boeddeker,Tobias Cord-Landwehr,Reinhold Haeb-Umbach
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
【7】 RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification
标题: RevRIR:使用对比学习联合混响语音和房间脉冲响应嵌入并应用于房间形状分类
作者:Jacob Bitterman,Daniel Levi,Hilel Hagai Diamandi,Sharon Gannot,Tal Rosenwein
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【8】 Singing Voice Graph Modeling for SingFake Detection
标题: SingFake检测的歌声图建模
作者:Xuanjun Chen,Haibin Wu,Jyh-Shing Roger Jang,Hung-yi Lee
备注:Accepted by Interspeech 2024; Codebase available at: this https URL
链接:点击下载PDF文件
【9】 4D ASR: Joint Beam Search Integrating CTC, Attention, Transducer, and Mask Predict Decoders
标题: 4D ASB:集成了CTC、注意力、传感器和屏蔽预测解码器的联合射束搜索
作者:Yui Sudo,Muhammad Shakeel,Yosuke Fukumoto,Brian Yan,Jiatong Shi,Yifan Peng,Shinji Watanabe
备注:submitted to IEEEACM Transactions on Audio Speech and Language Processing
链接:点击下载PDF文件
【10】 SYN2REAL: Leveraging Task Arithmetic for Mitigating Synthetic-Real Discrepancies in ASR Domain Adaptation
标题: SY 2 REAL:利用任务算法缓解SVR域自适应中的合成-真实差异
作者:Hsuan Su,Hua Farn,Shang-Tse Chen,Hung-yi Lee
链接:点击下载PDF文件
【11】 USM RNN-T model weights binarization
标题: USM RNN-T模型权重二值化
作者:Oleg Rybakov,Dmitriy Serdyuk,CJ Zheng
链接:点击下载PDF文件
【12】 ConPCO: Preserving Phoneme Characteristics for Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
标题: ConPCO:保留音素特征以进行自动发音评估,利用对比有序规则化
作者:Bi-Cheng Yan,Wei-Cheng Chao,Jiun-Ting Li,Yi-Cheng Wang,Hsin-Wei Wang,Meng-Shin Lin,Berlin Chen
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【13】 RepCNN: Micro-sized, Mighty Models for Wakeword Detection
标题: RepCNN:用于唤醒词检测的微型、强大模型
作者:Arnav Kundu,Prateeth Nayak,Hywel Richards,Priyanka Padmanabhan,Devang Naik
链接:点击下载PDF文件
【14】 Keyword-Guided Adaptation of Automatic Speech Recognition
标题: 关键词引导的自动语音识别自适应
作者:Aviv Shamsian,Aviv Navon,Neta Glazer,Gill Hetz,Joseph Keshet
备注:Accepted to InterSpeech 2024
链接:点击下载PDF文件
【15】 PPINtonus: Early Detection of Parkinson's Disease Using Deep-Learning Tonal Analysis
标题: PPINtonus:使用深度学习音调分析早期检测帕金森病
作者:Varun Reddy
链接:点击下载PDF文件
【16】 Selfsupervised learning for pathological speech detection
标题: 病态语音检测的自我监督学习
作者:Shakeel Ahmad Sheikh
备注:in Intersection of Book Chapter in Machine Leanring and Computational Social Sciences CRC (in progress) 2024
链接:点击下载PDF文件
【17】 Cluster-to-Predict Affect Contours from Speech
标题: 预测者影响言语轮廓
作者:Gökhan Kuşçu,Engin Erzin
备注:8 pages, 3 figures
链接:点击下载PDF文件
【18】 Combining X-Vectors and Bayesian Batch Active Learning: Two-Stage Active Learning Pipeline for Speech Recognition
标题: 结合X-Vector和Bayesian批量主动学习:语音识别的两阶段主动学习管道
作者:Ognjen Kundacina,Vladimir Vincan,Dragisa Miskovic
链接:点击下载PDF文件
【19】 A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
标题: 一种固定端到端ASC系统标记器中词汇量大小的成本最小化方法
作者:Sunil Kumar Kopparapu,Ashish Panda
备注:5 pages, 4 figures
链接:点击下载PDF文件
【20】 Gated Low-rank Adaptation for personalized Code-Switching Automatic Speech Recognition on the low-spec devices
标题: 门控低等级自适应在低规格设备上实现个性化代码切换自动语音识别
作者:Gwantae Kim,Bokyeung Lee,Donghyeon Kim,Hanseok Ko
Journal-ref:ICASSP 2024 Workshop(HSCMA 2024) paper
链接:点击下载PDF文件
【21】 Breaking Walls: Pioneering Automatic Speech Recognition for Central Kurdish: End-to-End Transformer Paradigm
标题: 破墙:库尔德中部开创自动语音识别:端到端Transformer范式
作者:Abdulhady Abas Abdullah,Hadi Veisi,Tarik Rashid
备注:
链接:点击下载PDF文件
【22】 Less Peaky and More Accurate CTC Forced Alignment by Label Priors
标题: 不那么尖峰、更准确的CTC通过标签先验强制对齐
作者:Ruizhe Huang,Xiaohui Zhang,Zhaoheng Ni,Li Sun,Moto Hira,Jeff Hwang,Vimal Manohar,Vineel Pratap,Matthew Wiesner,Shinji Watanabe,Daniel Povey,Sanjeev Khudanpur
备注:Accepted by ICASSP 2024. Github repo: this https URL
链接:点击下载PDF文件
【23】 PhoWhisper: Automatic Speech Recognition for Vietnamese
标题: PhoWhisper:越南语自动语音识别
作者:Thanh-Thien Le,Linh The Nguyen,Dat Quoc Nguyen
备注:Accepted to ICLR 2024 Tiny Papers Track
链接:点击下载PDF文件
【24】 Hear Me, See Me, Understand Me: Audio-Visual Autism Behavior Recognition
标题: 听到我、看到我、理解我:视听自闭症行为识别
作者:Shijian Deng,Erin E. Kosloski,Siddhi Patel,Zeke A. Barnett,Yiyang Nan,Alexander Kaplan,Sisira Aarukapalli,William T. Doan,Matthew Wang,Harsh Singh,Pamela R. Rollins,Yapeng Tian
链接:点击下载PDF文件
【25】 Physics and geometry informed neural operator network with application to acoustic scattering
标题: 物理和几何知识的神经操作网络及其在声散射中的应用
作者:Siddharth Nair,Timothy F. Walsh,Greg Pickrell,Fabio Semperlotti
备注:20 pages of main text, 9 figures
链接:点击下载PDF文件
【26】 Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
标题: 音频曼巴:音频表示学习的双向状态空间模型
作者:Mehmet Hamza Erol,Arda Senocak,Jiu Feng,Joon Son Chung
备注:Code is available at this https URL
链接:点击下载PDF文件
【27】 ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings
标题: ASoBO:为会议中的远距离发言人进行细心的束流器选择
作者:Theo Mariotte,Anthony Larcher,Silvio Montresor,Jean-Hugh Thomas
备注:5 pages, 2 figures, 2 tables, accepted at Interspeech 2024
链接:点击下载PDF文件
【28】 Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
标题: 使用MaskAutoEncoder进行广义假音频检测的以学生为中心的学习
作者:Xiaopeng Wang,Ruibo Fu,Zhengqi Wen,Zhiyong Wang,Yuankun Xie,Yukun Liu,Jianhua Tao,Xuefei Liu,Yongwei Li,Xin Qi,Yi Lu,Shuchen Shi
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【29】 Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion strategy
标题: 广义源跟踪:检测具有真实重点和虚假分散策略的新型音频Deepfake算法
作者:Yuankun Xie,Ruibo Fu,Zhengqi Wen,Zhiyong Wang,Xiaopeng Wang,Haonnan Cheng,Long Ye,Jianhua Tao
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【30】 Generalized Fake Audio Detection via Deep Stable Learning
标题: 通过深度稳定学习进行广义假音频检测
作者:Zhiyong Wang,Ruibo Fu,Zhengqi Wen,Yuankun Xie,Yukun Liu,Xiaopeng Wang,Xuefei Liu,Yongwei Li,Jianhua Tao,Yi Lu,Xin Qi,Shuchen Shi
备注:accepted by INTERSPEECH2024
链接:点击下载PDF文件
【31】 A Frame-based Attention Interpretation Method for Relevant Acoustic Feature Extraction in Long Speech Depression Detection
标题: 长言语抑郁检测中相关声学特征提取的基于框架的注意力解释方法
作者:Qingkun Deng,Saturnino Luz,Sofia de la Fuente Garcia
备注:5 pages, 3 figures. arXiv admin note: substantial text overlap with arXiv:2309.13476
链接:点击下载PDF文件
【32】 StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
标题: StreamSpeech:具有多任务学习的语音同步翻译
作者:Shaolei Zhang,Qingkai Fang,Shoutao Guo,Zhengrui Ma,Min Zhang,Yang Feng
备注:Accepted to ACL 2024 main conference, Project Page: this https URL
链接:点击下载PDF文件
【33】 Dataset-Distillation Generative Model for Speech Emotion Recognition
标题: 语音情感识别的数据集蒸馏生成模型
作者:Fabian Ritter-Gutierrez,Kuan-Po Huang,Jeremy H. M Wong,Dianwen Ng,Hung-yi Lee,Nancy F. Chen,Eng Siong Chng
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
【34】 AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
标题: AVFF:用于视频Deepfake检测的视听特征融合
作者:Trevine Oorloff,Surya Koppisetti,Nicolò Bonettini,Divyaraj Solanki,Ben Colman,Yaser Yacoob,Ali Shahriyari,Gaurav Bharaj
备注:Accepted to CVPR 2024
链接:点击下载PDF文件
【35】 Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
标题: 用双解码产品量化变分自动编码器解决大码本语音令牌器的索引崩溃
作者:Haohan Guo,Fenglong Xie,Dongchao Yang,Hui Lu,Xixin Wu,Helen Meng
链接:点击下载PDF文件
【36】 Robots Have Been Seen and Not Heard: Effects of Consequential Sounds on Human-Perception of Robots
标题: 机器人被看到却没有被听到:随之而来的声音对人类对机器人感知的影响
作者:Aimee Allen,Tom Drummond,Dana Kulic
备注:16 pages (5 supplementary), 9 figures
链接:点击下载PDF文件
【37】 Text Injection for Neural Contextual Biasing
标题: 用于神经上下文偏置的文本注入
作者:Zhong Meng,Zelin Wu,Rohit Prabhavalkar,Cal Peyser,Weiran Wang,Nanxin Chen,Tara N. Sainath,Bhuvana Ramabhadran
Journal-ref:Interspeech 2024, Kos Island, Greece
链接:点击下载PDF文件
【38】 LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
标题: LiveSpeech:通过音频离散码的自回归建模的低延迟Zero-Shot文本到语音
作者:Trung Dang,David Aponte,Dung Tran,Kazuhito Koishida
链接:点击下载PDF文件
【39】 Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
标题: 具有自我监督蒸馏的无文本声学模型用于噪音稳健的表达性语音到语音翻译
作者:Min-Jae Hwang,Ilia Kulikov,Benjamin Peloquin,Hongyu Gong,Peng-Jen Chen,Ann Lee
备注:Accepted to ACL 2024 (findings)
链接:点击下载PDF文件
【40】 Operational Latent Spaces
标题: 运营潜在空间
作者:Scott H. Hawley,Austin R. Tackett
备注:7 pages, 6 figures. Accepted to AES International Symposium on AI and the Musician
链接:点击下载PDF文件
【41】 Sequence-to-sequence models in peer-to-peer learning: A practical application
标题: 点对点学习中的序列到序列模型:实际应用
作者:Robert Šajina,Ivo Ipšić
链接:点击下载PDF文件
标题: 物理和几何知识的神经操作网络及其在声散射中的应用
作者:Siddharth Nair,Timothy F. Walsh,Greg Pickrell,Fabio Semperlotti
备注:20 pages of main text, 9 figures
链接:点击下载PDF文件
摘要:本文介绍了一种基于物理和几何信息的神经算子网络,并将其应用于声散射的正演模拟。开发能够学习不同计算域的解算子的几何信息深度学习模型对于各种工程应用来说是一个普遍重要的问题。为此,我们提出了一个物理信息的深度算子网络(DeepONet),能够使用基于非均匀有理B样条(NURBS)的几何参数化方法预测任意形状散射体的散射压力场。这种方法也导致在非平凡的散射几何形状的简约表示。与现有的基于物理的方法相比,当改变计算域时,需要重新评估模型,我们训练的模型能够在几秒钟内学习可以近似物理一致的散射压力场的解算子,对于任意刚性散射体形状;因此,前向模拟的计算时间可以改进与传统的正向求解器相比,可以减少(即减少)数量级。此外,该方法可以评估分散的压力场,而不需要标记的训练数据。提出的理论方法后,还提供了一个全面的数值研究来说明这种方法的显着能力来模拟任意散射体几何形状的任意组合所产生的声压场。这些结果突出了独特的泛化能力的建议运营商学习方法。摘要:In this paper, we introduce a physics and geometry informed neural operator network with application to the forward simulation of acoustic scattering. The development of geometry informed deep learning models capable of learning a solution operator for different computational domains is a problem of general importance for a variety of engineering applications. To this end, we propose a physics-informed deep operator network (DeepONet) capable of predicting the scattered pressure field for arbitrarily shaped scatterers using a geometric parameterization approach based on non-uniform rational B-splines (NURBS). This approach also results in parsimonious representations of non-trivial scatterer geometries. In contrast to existing physics-based approaches that require model re-evaluation when changing the computational domains, our trained model is capable of learning solution operator that can approximate physically-consistent scattered pressure field in just a few seconds for arbitrary rigid scatterer shapes; it follows that the computational time for forward simulations can improve (i.e. be reduced) by orders of magnitude in comparison to the traditional forward solvers. In addition, this approach can evaluate the scattered pressure field without the need for labeled training data. After presenting the theoretical approach, a comprehensive numerical study is also provided to illustrate the remarkable ability of this approach to simulate the acoustic pressure fields resulting from arbitrary combinations of arbitrary scatterer geometries. These results highlight the unique generalization capability of the proposed operator learning approach.
【2】 Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
标题: 音频曼巴:音频表示学习的双向状态空间模型
作者:Mehmet Hamza Erol,Arda Senocak,Jiu Feng,Joon Son Chung
备注:Code is available at this https URL
链接:点击下载PDF文件
摘要:Transformers已迅速成为音频分类的首选,超过了基于CNN的方法。然而,音频频谱图Transformers(AST)表现出二次缩放由于自我注意。消除这种二次自我注意力成本提出了一个有吸引力的方向。最近,状态空间模型(SSM),如Mamba,在语言和视觉任务中表现出了潜力。在这项研究中,我们探讨是否依赖于自我注意是必要的音频分类任务。通过引入音频曼巴(AuM),第一个自我注意力自由,纯粹基于SSM的音频分类模型,我们的目标是解决这个问题。我们在各种音频数据集上评估AuM-包括六个不同的基准-与成熟的AST模型相比,它实现了相当或更好的性能。摘要:Transformers have rapidly become the preferred choice for audio classification, surpassing methods based on CNNs. However, Audio Spectrogram Transformers (ASTs) exhibit quadratic scaling due to self-attention. The removal of this quadratic self-attention cost presents an appealing direction. Recently, state space models (SSMs), such as Mamba, have demonstrated potential in language and vision tasks in this regard. In this study, we explore whether reliance on self-attention is necessary for audio classification tasks. By introducing Audio Mamba (AuM), the first self-attention-free, purely SSM-based model for audio classification, we aim to address this question. We evaluate AuM on various audio datasets - comprising six different benchmarks - where it achieves comparable or better performance compared to well-established AST model.
【3】 ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings
标题: ASoBO:为会议中的远距离发言人进行细心的束流器选择
作者:Theo Mariotte,Anthony Larcher,Silvio Montresor,Jean-Hugh Thomas
备注:5 pages, 2 figures, 2 tables, accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:说话人日志化(Speaker Diarization,SD)的目的是将属于同一说话人的语音片段进行分组。许多语音处理应用程序(如丰富的会议转录)都需要执行此任务。在这种情况下,远距离麦克风阵列通常捕获音频信号。波束成形,即,空间滤波是处理多麦克风音频数据的常见实践。然而,它往往需要一个明确的定位的有源源,以引导过滤器。本文提出了一种基于自注意的算法来选择一组固定的空间滤波器的输出。该方法作为一个特征提取器的联合语音活动(VAD)和重叠语音检测(OSD)。然后从检测到的片段推断出说话人日记。该方法显示了令人信服的远程VAD,OSD和SD性能,例如AISHELL-4数据集上的14.5% DER。自我注意权重的分析证明了它们的可解释性,因为它们与说话者的角度位置相关。摘要:Speaker Diarization (SD) aims at grouping speech segments that belong to the same speaker. This task is required in many speech-processing applications, such as rich meeting transcription. In this context, distant microphone arrays usually capture the audio signal. Beamforming, i.e., spatial filtering, is a common practice to process multi-microphone audio data. However, it often requires an explicit localization of the active source to steer the filter. This paper proposes a self-attention-based algorithm to select the output of a bank of fixed spatial filters. This method serves as a feature extractor for joint Voice Activity (VAD) and Overlapped Speech Detection (OSD). The speaker diarization is then inferred from the detected segments. The approach shows convincing distant VAD, OSD, and SD performance, e.g. 14.5% DER on the AISHELL-4 dataset. The analysis of the self-attention weights demonstrates their explainability, as they correlate with the speaker's angular locations.
【4】 Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
标题: 使用MaskAutoEncoder进行广义假音频检测的以学生为中心的学习
作者:Xiaopeng Wang,Ruibo Fu,Zhengqi Wen,Zhiyong Wang,Yuankun Xie,Yukun Liu,Jianhua Tao,Xuefei Liu,Yongwei Li,Xin Qi,Yi Lu,Shuchen Shi
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:由于新的欺骗技术的出现,虚假音频检测(FAD)的推广是至关重要的。传统的FAD方法通常仅关注于区分真实音频和已知的欺骗音频。我们提出了一个以神经网络为中心的学习(GFL)框架引导,旨在高度广义的FAD,称为GFL-FAD。该方法结合了基于音频重建的反事实推理增强表示(CRER),使用掩码自动编码器(MAE)架构来准确地建模真实的音频特征。为了减少训练过程中欺骗音频的影响,我们引入了真正的音频重建损失,保持专注于学习真正的数据特征。此外,内容相关的瓶颈(BN)功能提取的MAE补充知识的原始音频。这些BN特征自适应地与CRER融合以进一步提高鲁棒性。我们的方法在ASVspoof 2019 LA上实现了最先进的性能,EER为0.25%。摘要:The generalization of Fake Audio Detection (FAD) is critical due to the emergence of new spoofing techniques. Traditional FAD methods often focus solely on distinguishing between genuine and known spoofed audio. We propose a Genuine-Focused Learning (GFL) framework guided, aiming for highly generalized FAD, called GFL-FAD. This method incorporates a Counterfactual Reasoning Enhanced Representation (CRER) based on audio reconstruction using the Mask AutoEncoder (MAE) architecture to accurately model genuine audio features. To reduce the influence of spoofed audio during training, we introduce a genuine audio reconstruction loss, maintaining the focus on learning genuine data features. In addition, content-related bottleneck (BN) features are extracted from the MAE to supplement the knowledge of the original audio. These BN features are adaptively fused with CRER to further improve robustness. Our method achieves state-of-the-art performance with an EER of 0.25% on ASVspoof2019 LA.
【5】 Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion strategy
标题: 广义源跟踪:检测具有真实重点和虚假分散策略的新型音频Deepfake算法
作者:Yuankun Xie,Ruibo Fu,Zhengqi Wen,Zhiyong Wang,Xiaopeng Wang,Haonnan Cheng,Long Ye,Jianhua Tao
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:随着deepfake音频的扩散,迫切需要调查它们的归属。目前的来源追踪方法可以有效地区分在分布(ID)类别。然而,deepfake算法的快速发展对准确识别非分布(OOD)新型deepfake算法提出了严峻的挑战。在本文中,我们提出了用于音频deepfake算法识别的真实强调和虚假分散(REFD)策略,证明了其在识别OOD样本的同时区分ID样本的有效性。为了有效的OOD检测,我们首先探索了当前的事后OOD方法,并提出了NSD,这是一种新的OOD方法,通过考虑特征和logits分数的相似性来识别新的deepfake算法。REFD在2023年音频Deepfake检测挑战赛Track3中作为单一系统获得了86.83%的F1分数,展示了其最先进的性能。摘要:With the proliferation of deepfake audio, there is an urgent need to investigate their attribution. Current source tracing methods can effectively distinguish in-distribution (ID) categories. However, the rapid evolution of deepfake algorithms poses a critical challenge in the accurate identification of out-of-distribution (OOD) novel deepfake algorithms. In this paper, we propose Real Emphasis and Fake Dispersion (REFD) strategy for audio deepfake algorithm recognition, demonstrating its effectiveness in discriminating ID samples while identifying OOD samples. For effective OOD detection, we first explore current post-hoc OOD methods and propose NSD, a novel OOD approach in identifying novel deepfake algorithms through the similarity consideration of both feature and logits scores. REFD achieves 86.83% F1-score as a single system in Audio Deepfake Detection Challenge 2023 Track3, showcasing its state-of-the-art performance.
【6】 Generalized Fake Audio Detection via Deep Stable Learning
标题: 通过深度稳定学习进行广义假音频检测
作者:Zhiyong Wang,Ruibo Fu,Zhengqi Wen,Yuankun Xie,Yukun Liu,Xiaopeng Wang,Xuefei Liu,Yongwei Li,Jianhua Tao,Yi Lu,Xin Qi,Shuchen Shi
备注:accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:虽然目前的虚假音频检测方法在特定数据集上取得了显着的成功,但在使用来自不同分布的数据集进行评估时,它们往往会失败。以前的研究通常通过在训练过程中使用额外的数据或应用额外的损失限制来解决分布偏移。然而,这些方法要么需要大量的数据,要么使训练过程复杂化。在这项工作中,我们提出了一个稳定的基于学习的训练方案,该方案涉及样本权重学习(SWL)模块,通过从训练样本中学习权重来解相关所有选定的特征,从而解决分布偏移问题。所提出的便携式插件式SWL很容易应用于多个基本模型,并在训练过程中不使用额外的数据来推广它们。在ASVspoof数据集上进行的实验清楚地证明了SWL在不同分布的三个评估数据集上推广不同模型的有效性。摘要:Although current fake audio detection approaches have achieved remarkable success on specific datasets, they often fail when evaluated with datasets from different distributions. Previous studies typically address distribution shift by focusing on using extra data or applying extra loss restrictions during training. However, these methods either require a substantial amount of data or complicate the training process. In this work, we propose a stable learning-based training scheme that involves a Sample Weight Learning (SWL) module, addressing distribution shift by decorrelating all selected features via learning weights from training samples. The proposed portable plug-in-like SWL is easy to apply to multiple base models and generalizes them without using extra data during training. Experiments conducted on the ASVspoof datasets clearly demonstrate the effectiveness of SWL in generalizing different models across three evaluation datasets from different distributions.
【7】 A Frame-based Attention Interpretation Method for Relevant Acoustic Feature Extraction in Long Speech Depression Detection
标题: 长言语抑郁检测中相关声学特征提取的基于框架的注意力解释方法
作者:Qingkun Deng,Saturnino Luz,Sofia de la Fuente Garcia
备注:5 pages, 3 figures. arXiv admin note: substantial text overlap with arXiv:2309.13476
链接:点击下载PDF文件
摘要:基于语音的抑郁症检测工具可以帮助早期筛查抑郁症。在这里,我们解决了两个问题,可能会阻碍这种工具的临床实用性:段级标签噪声和缺乏模型的可解释性。我们提出了一个语音级的音频频谱图Transformer,以避免段级标签。我们观察到,该模型显着优于段级模型,提供证据的存在段级标签噪声的音频模态和抑郁症检测的优势,持续时间较长的语音分析。我们引入了一种基于帧的注意力解释方法,从预测相关的波形信号中提取声学特征,供临床医生解释。通过解释,我们观察到,所提出的模型识别降低的响度和F0作为抑郁症的相关信号,这与临床研究中记录的抑郁症患者的语音特征一致。摘要:Speech-based depression detection tools could help early screening of depression. Here, we address two issues that may hinder the clinical practicality of such tools: segment-level labelling noise and a lack of model interpretability. We propose a speech-level Audio Spectrogram Transformer to avoid segment-level labelling. We observe that the proposed model significantly outperforms a segment-level model, providing evidence for the presence of segment-level labelling noise in audio modality and the advantage of longer-duration speech analysis for depression detection. We introduce a frame-based attention interpretation method to extract acoustic features from prediction-relevant waveform signals for interpretation by clinicians. Through interpretation, we observe that the proposed model identifies reduced loudness and F0 as relevant signals of depression, which aligns with the speech characteristics of depressed patients documented in clinical studies.
【8】 StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
标题: StreamSpeech:具有多任务学习的语音同步翻译
作者:Shaolei Zhang,Qingkai Fang,Shoutao Guo,Zhengrui Ma,Min Zhang,Yang Feng
备注:Accepted to ACL 2024 main conference, Project Page: this https URL
链接:点击下载PDF文件
摘要:同步语音到语音翻译(Simul-S2 ST,也称为流式语音翻译)在接收流式语音输入的同时输出目标语音,这对于实时通信至关重要。除了完成语音之间的翻译,Simul-S2 ST还需要一个策略来控制模型在语音输入的适当时刻生成相应的目标语音,从而提出了翻译和策略的双重挑战。在本文中,我们提出了StreamSpeech,这是一个直接的Simul-S2 ST模型,它在多任务学习的统一框架中联合学习翻译和同步策略。StreamSpeech坚持多任务学习方法,可以通过“一体化”无缝模型执行离线和同步语音识别,语音翻译和语音合成。在CVSS基准测试上的实验表明,StreamSpeech在离线S2 ST和Simul-S2 ST任务中都达到了最佳性能。此外,StreamSpeech能够呈现高质量的中间结果(即,ASR或翻译结果),提供更全面的实时交流体验。摘要:Simultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication. Beyond accomplishing translation between speech, Simul-S2ST requires a policy to control the model to generate corresponding target speech at the opportune moment within speech inputs, thereby posing a double challenge of translation and policy. In this paper, we propose StreamSpeech, a direct Simul-S2ST model that jointly learns translation and simultaneous policy in a unified framework of multi-task learning. Adhering to a multi-task learning approach, StreamSpeech can perform offline and simultaneous speech recognition, speech translation and speech synthesis via an "All-in-One" seamless model. Experiments on CVSS benchmark demonstrate that StreamSpeech achieves state-of-the-art performance in both offline S2ST and Simul-S2ST tasks. Besides, StreamSpeech is able to present high-quality intermediate results (i.e., ASR or translation results) during simultaneous translation process, offering a more comprehensive real-time communication experience.
【9】 Dataset-Distillation Generative Model for Speech Emotion Recognition
标题: 语音情感识别的数据集蒸馏生成模型
作者:Fabian Ritter-Gutierrez,Kuan-Po Huang,Jeremy H. M Wong,Dianwen Ng,Hung-yi Lee,Nancy F. Chen,Eng Siong Chng
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:语音深度学习模型依赖于大型数据集,这带来了计算挑战。然而,性能取决于训练数据的大小。数据集蒸馏(DD)旨在学习较小的数据集,而在使用它进行训练时不会降低性能。DD已在计算机视觉中进行了研究,但尚未在语音中进行研究。本文提出了第一个方法,DD语音目标的语音情感识别IEMOCAP。我们使用生成对抗网络(GANs)不是为了模拟真实数据,而是为了验证IEMOCAP的判别信息,这些信息对下游训练很有用。然后,GAN替换原始数据集,并可以对自定义合成数据集大小进行采样。当遵循原始类不平衡时,它会执行UAR,但使用平衡类时,它会将性能提高0.3%的绝对UAR。它还减少了数据集存储,在这两种情况下将下游训练速度加快了95%,并减少了扬声器信息,这可能有助于隐私应用程序。摘要:Deep learning models for speech rely on large datasets, presenting computational challenges. Yet, performance hinges on training data size. Dataset Distillation (DD) aims to learn a smaller dataset without much performance degradation when training with it. DD has been investigated in computer vision but not yet in speech. This paper presents the first approach for DD to speech targeting Speech Emotion Recognition on IEMOCAP. We employ Generative Adversarial Networks (GANs) not to mimic real data but to distil key discriminative information of IEMOCAP that is useful for downstream training. The GAN then replaces the original dataset and can sample custom synthetic dataset sizes. It performs comparably when following the original class imbalance but improves performance by 0.3% absolute UAR with balanced classes. It also reduces dataset storage and accelerates downstream training by 95% in both cases and reduces speaker information which could help for a privacy application.
【10】 AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
标题: AVFF:用于视频Deepfake检测的视听特征融合
作者:Trevine Oorloff,Surya Koppisetti,Nicolò Bonettini,Divyaraj Solanki,Ben Colman,Yaser Yacoob,Ali Shahriyari,Gaurav Bharaj
备注:Accepted to CVPR 2024
链接:点击下载PDF文件
摘要:随着deepfake视频内容的快速增长,我们需要改进和可推广的方法来检测它们。大多数现有的检测方法要么使用单模态线索,要么依赖于监督训练来捕获音频和视觉模态之间的不和谐。虽然前者完全忽略了视听对应关系,但后者主要集中在识别训练语料库中的视听线索,从而可能忽略了有助于检测看不见的深度伪造的对应关系。我们提出了视听特征融合(AVFF),这是一种两阶段的跨模态学习方法,可以显式捕获音频和视觉模态之间的对应关系,以改进深度伪造检测。第一阶段通过对真实视频的自我监督来进行表征学习,以捕获内在的视听对应。为了提取丰富的跨模态表示,我们使用对比学习和自动编码目标,并引入了一种新的视听互补掩蔽和特征融合策略。在第二阶段中调整学习的表示,其中通过对真实和虚假视频的监督学习来进行deepfake分类。大量的实验和分析表明,我们的新的表征学习范式是高度歧视的性质。我们在FakeAVCeleb数据集上报告了98.6%的准确率和99.1%的AUC,分别比当前最先进的视听技术高出14.9%和9.9%。摘要:With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the audio and visual modalities. While the former disregards the audio-visual correspondences entirely, the latter predominantly focuses on discerning audio-visual cues within the training corpus, thereby potentially overlooking correspondences that can help detect unseen deepfakes. We present Audio-Visual Feature Fusion (AVFF), a two-stage cross-modal learning method that explicitly captures the correspondence between the audio and visual modalities for improved deepfake detection. The first stage pursues representation learning via self-supervision on real videos to capture the intrinsic audio-visual correspondences. To extract rich cross-modal representations, we use contrastive learning and autoencoding objectives, and introduce a novel audio-visual complementary masking and feature fusion strategy. The learned representations are tuned in the second stage, where deepfake classification is pursued via supervised learning on both real and fake videos. Extensive experiments and analysis suggest that our novel representation learning paradigm is highly discriminative in nature. We report 98.6% accuracy and 99.1% AUC on the FakeAVCeleb dataset, outperforming the current audio-visual state-of-the-art by 14.9% and 9.9%, respectively.
【11】 Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
标题: 用双解码产品量化变分自动编码器解决大码本语音令牌器的索引崩溃
作者:Haohan Guo,Fenglong Xie,Dongchao Yang,Hui Lu,Xixin Wu,Helen Meng
链接:点击下载PDF文件
摘要:VQ-VAE作为语音标记器的主流方法,一直受到“索引搜索”的困扰,在大的码本中只有少量的码字被激活。本文提出了一种具有更多码本但更少码字的积量化(PQ)VAE来解决这个问题,并构建大码本语音标记器。它将语音特征编码到多个VQ子空间中,并将它们组合成更大码书中的码字。此外,为了更好地利用每个矢量量化子空间,我们还通过编码和量化序列的双重解码训练策略来增强PQ-VAE。实验结果表明,PQ-VAE有效地解决了“索引崩溃”问题,特别是对于较大的码本。该模型的训练策略进一步提高了码本复杂度和重建质量,优于其他多码本矢量量化方法。最后,PQ-VAE证明了其在基于语言模型的TTS中的有效性,支持具有更大码本的更高质量的语音生成。摘要:VQ-VAE, as a mainstream approach of speech tokenizer, has been troubled by index collapse'', where only a small number of codewords are activated in large codebooks. This work proposes product-quantized (PQ) VAE with more codebooks but fewer codewords to address this problem and build large-codebook speech tokenizers. It encodes speech features into multiple VQ subspaces and composes them into codewords in a larger codebook. Besides, to utilize each VQ subspace well, we also enhance PQ-VAE via a dual-decoding training strategy with the encoding and quantized sequences. The experimental results demonstrate that PQ-VAE addresses index collapse" effectively, especially for larger codebooks. The model with the proposed training strategy further improves codebook perplexity and reconstruction quality, outperforming other multi-codebook VQ approaches. Finally, PQ-VAE demonstrates its effectiveness in language-model-based TTS, supporting higher-quality speech generation with larger codebooks.
【12】 LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
标题: LiveSpeech:通过音频离散码的自回归建模的低延迟Zero-Shot文本到语音
作者:Trung Dang,David Aponte,Dung Tran,Kazuhito Koishida
链接:点击下载PDF文件
摘要:先前的工作已经通过在经由神经音频编解码器获得的音频令牌上使用生成语言模型来演示了zero-shot文本到语音。然而,使它们适应低延迟场景仍然具有挑战性。在本文中,我们提出了LiveSpeech -一个完全自回归语言模型为基础的方法,zero-shot文本到语音,使低延迟流的输出音频。为了允许在单个解码步骤内进行多个令牌预测,我们提出(1)使用自适应码本损失权重,该自适应码本损失权重考虑每个帧中的码本贡献并专注于硬实例,以及(2)并行地对码本和处理组进行分组。实验表明,我们提出的模型在内容准确性、扬声器相似性、音频质量和推理速度方面达到了最先进的基线,同时适用于低延迟流媒体应用。摘要:Prior works have demonstrated zero-shot text-to-speech by using a generative language model on audio tokens obtained via a neural audio codec. It is still challenging, however, to adapt them to low-latency scenarios. In this paper, we present LiveSpeech - a fully autoregressive language model-based approach for zero-shot text-to-speech, enabling low-latency streaming of the output audio. To allow multiple token prediction within a single decoding step, we propose (1) using adaptive codebook loss weights that consider codebook contribution in each frame and focus on hard instances, and (2) grouping codebooks and processing groups in parallel. Experiments show our proposed models achieve competitive results to state-of-the-art baselines in terms of content accuracy, speaker similarity, audio quality, and inference speed while being suitable for low-latency streaming applications.
【13】 Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
标题: 具有自我监督蒸馏的无文本声学模型用于噪音稳健的表达性语音到语音翻译
作者:Min-Jae Hwang,Ilia Kulikov,Benjamin Peloquin,Hongyu Gong,Peng-Jen Chen,Ann Lee
备注:Accepted to ACL 2024 (findings)
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个无文本的声学模型与自监督蒸馏策略的噪声鲁棒表达语音到语音翻译(S2ST)。最近提出的表达S2ST系统已经取得了令人印象深刻的表现力保存性能级联单元到语音(U2S)生成器的语音到单元的翻译模型。然而,这些系统容易受到输入语音中存在噪声的影响,这是现实世界翻译场景中的一个假设。为了解决这一限制,我们提出了一个U2S生成器,它将无标签蒸馏(DINO)自监督训练策略纳入其预训练过程。由于该方法捕捉了与噪声无关的表达能力,因此即使在噪声环境中也能生成合格的语音。客观和主观评价结果表明,该方法在保持清晰环境下性能的同时,显著提高了表达性S2ST系统在噪声环境下的性能.摘要:In this paper, we propose a textless acoustic model with a self-supervised distillation strategy for noise-robust expressive speech-to-speech translation (S2ST). Recently proposed expressive S2ST systems have achieved impressive expressivity preservation performances by cascading unit-to-speech (U2S) generator to the speech-to-unit translation model. However, these systems are vulnerable to the presence of noise in input speech, which is an assumption in real-world translation scenarios. To address this limitation, we propose a U2S generator that incorporates a distillation with no label (DINO) self-supervised training strategy into it's pretraining process. Because the proposed method captures noise-agnostic expressivity representation, it can generate qualified speech even in noisy environment. Objective and subjective evaluation results verified that the proposed method significantly improved the performance of the expressive S2ST system in noisy environments while maintaining competitive performance in clean environments.
【14】 Sequence-to-sequence models in peer-to-peer learning: A practical application
标题: 点对点学习中的序列到序列模型:实际应用
作者:Robert Šajina,Ivo Ipšić
链接:点击下载PDF文件
摘要:本文探讨了基于LSTM单元的序列到序列(Seq2Seq)模型在对等学习环境中用于自动语音识别(ASR)任务的适用性。利用两种不同的对等学习方法,该研究模拟了智能体的学习过程,并使用两种不同的ASR数据集来评估它们在ASR任务中的表现。在集中式训练环境中,利用Deep Speech 2模型的缩小变体,单个模型在UserLibri数据集上训练时的单词错误率(WER)为84%,在LJ Speech数据集上训练时为38%。相反,在涉及55个代理的对等学习场景中,UserLibri数据集的WER范围为87 %至92 %,LJ Speech数据集的WER范围为52 %至56 %。研究结果证明了在分散的环境中使用Seq2Seq模型的可行性,尽管与集中式训练方法相比,单词错误率(WER)略高。摘要:This paper explores the applicability of sequence-to-sequence (Seq2Seq) models based on LSTM units for Automatic Speech Recognition (ASR) task within peer-to-peer learning environments. Leveraging two distinct peer-to-peer learning methods, the study simulates the learning process of agents and evaluates their performance in ASR task using two different ASR datasets. In a centralized training setting, utilizing a scaled-down variant of the Deep Speech 2 model, a single model achieved a Word Error Rate (WER) of 84 % when trained on the UserLibri dataset, and 38 % when trained on the LJ Speech dataset. Conversely, in a peer-to-peer learning scenario involving 55 agents, the WER ranged from 87 % to 92 % for the UserLibri dataset, and from 52 % to 56 % for the LJ Speech dataset. The findings demonstrate the feasibility of employing Seq2Seq models in decentralized settings, albeit with slightly higher Word Error Rates (WER) compared to centralized training methods.
【15】 The PESQetarian: On the Relevance of Goodhart's Law for Speech Enhancement
标题: PESQetarian:论古德哈特定律与言语增强的相关性
作者:Danilo de Oliveira,Simon Welker,Julius Richter,Timo Gerkmann
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:为了获得改进的语音增强模型,研究人员通常专注于根据特定的工具度量来提高性能。然而,当在损失函数中使用相同的度量来优化模型时,它可能对给定度量看不到的方面有害。本文的目的是说明过拟合的语音增强模型的评估所使用的度量的风险。为此,我们引入了利用广泛使用的PESQ措施的增强模型。我们的“PESQetarian”模型在VB-DMD上达到3.82 PESQ,而在听力实验中得分很低。虽然获得的PESQ值3.82意味着VB-DMD基准测试的“最先进”PESQ性能,但我们的示例表明,当优化w.r.t.一个度量,对同一度量的孤立评估可能会产生误导。相反,其他指标应包括在评估中,并通过倾听确认由此产生的性能预测。摘要:To obtain improved speech enhancement models, researchers often focus on increasing performance according to specific instrumental metrics. However, when the same metric is used in a loss function to optimize models, it may be detrimental to aspects that the given metric does not see. The goal of this paper is to illustrate the risk of overfitting a speech enhancement model to the metric used for evaluation. For this, we introduce enhancement models that exploit the widely used PESQ measure. Our "PESQetarian" model achieves 3.82 PESQ on VB-DMD while scoring very poorly in a listening experiment. While the obtained PESQ value of 3.82 would imply "state-of-the-art" PESQ-performance on the VB-DMD benchmark, our examples show that when optimizing w.r.t. a metric, an isolated evaluation on the same metric may be misleading. Instead, other metrics should be included in the evaluation and the resulting performance predictions should be confirmed by listening.
【16】 Enhancing CTC-based speech recognition with diverse modeling units
标题: 利用不同的建模单元增强基于ATC的语音识别
作者:Shiyi Han,Zhihong Lei,Mingbin Xu,Xingyu Na,Zhen Huang
链接:点击下载PDF文件
摘要:近年来,端到端(E2 E)自动语音识别(ASR)模型的发展引人注目,这主要归功于Transformer等深度学习架构的进步。在E2 E系统之上,研究人员通过使用基于音素的模型对E2 E模型的N-best假设进行重新评分,实现了实质性的准确性提高。这提出了一个有趣的问题,即除了系统组合效应之外,这些改进来自何处。我们研究了驱动这些收益的潜在机制,并提出了一种有效的联合训练方法,其中E2 E模型与不同的建模单元联合训练。这种方法不仅使基于音素和字素的模型的优势保持一致,而且还揭示了以协同方式使用这些不同的建模单元可以显着提高模型的准确性。我们的研究结果提供了新的见解异构建模单元的最佳集成,在更强大和准确的ASR系统的发展。摘要:In recent years, the evolution of end-to-end (E2E) automatic speech recognition (ASR) models has been remarkable, largely due to advances in deep learning architectures like transformer. On top of E2E systems, researchers have achieved substantial accuracy improvement by rescoring E2E model's N-best hypotheses with a phoneme-based model. This raises an interesting question about where the improvements come from other than the system combination effect. We examine the underlying mechanisms driving these gains and propose an efficient joint training approach, where E2E models are trained jointly with diverse modeling units. This methodology does not only align the strengths of both phoneme and grapheme-based models but also reveals that using these diverse modeling units in a synergistic way can significantly enhance model accuracy. Our findings offer new insights into the optimal integration of heterogeneous modeling units in the development of more robust and accurate ASR systems.
【17】 RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification
标题: RevRIR:使用对比学习联合混响语音和房间脉冲响应嵌入并应用于房间形状分类
作者:Jacob Bitterman,Daniel Levi,Hilel Hagai Diamandi,Sharon Gannot,Tal Rosenwein
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文重点介绍了房间指纹识别,这是一项涉及分析音频记录以确定其被捕获的房间的特定音量和形状的任务。虽然从房间脉冲响应(RIR)确定基本房间参数是相对简单的,但是从语音信号这样做是繁琐的任务。为了解决这一挑战,我们引入了一个双编码器架构,便于直接从语音话语的房间参数的估计。在预训练期间,一个编码器接收RIR,而另一个处理混响语音信号。一个对比损失函数被用来嵌入语音和声学响应。在微调阶段,训练特定的分类任务。在测试阶段,只有混响的话语是可用的,其嵌入用于房间形状分类的任务。所提出的方案进行了广泛的评估,使用模拟的声学环境。摘要:This paper focuses on room fingerprinting, a task involving the analysis of an audio recording to determine the specific volume and shape of the room in which it was captured. While it is relatively straightforward to determine the basic room parameters from the Room Impulse Responses (RIR), doing so from a speech signal is a cumbersome task. To address this challenge, we introduce a dual-encoder architecture that facilitates the estimation of room parameters directly from speech utterances. During pre-training, one encoder receives the RIR while the other processes the reverberant speech signal. A contrastive loss function is employed to embed the speech and the acoustic response jointly. In the fine-tuning stage, the specific classification task is trained. In the test phase, only the reverberant utterance is available, and its embedding is used for the task of room shape classification. The proposed scheme is extensively evaluated using simulated acoustic environments.
【18】 4D ASR: Joint Beam Search Integrating CTC, Attention, Transducer, and Mask Predict Decoders
标题: 4D ASB:集成了CTC、注意力、传感器和屏蔽预测解码器的联合射束搜索
作者:Yui Sudo,Muhammad Shakeel,Yosuke Fukumoto,Brian Yan,Jiatong Shi,Yifan Peng,Shinji Watanabe
备注:submitted to IEEEACM Transactions on Audio Speech and Language Processing
链接:点击下载PDF文件
摘要:端到端自动语音识别(E2 E-ASR)可以分为几种网络架构,如连接主义时间分类(CTC),递归神经网络转换器(RNN-T),基于注意力的编码器-解码器和掩码预测模型。每种网络架构都有优点和缺点,导致从业者根据应用需求在这些不同的模型之间切换。我们没有建立单独的模型,而是提出了一种联合建模方案,其中四个解码器(CTC,RNN-T,attention和mask-predict)共享同一个编码器-我们将其称为4D建模。4D模型使用多任务学习进行训练,这将带来模型正则化并最大化模型鲁棒性,这要归功于它们的互补特性。为了有效地训练4D模型,我们引入了一个稳定多任务学习的两阶段训练策略。此外,我们提出了三种新的一次通过波束搜索算法,通过结合三个解码器(CTC,RNN-T和注意力),以进一步提高性能。这三种波束搜索算法的不同之处在于哪个解码器被用作主解码器。我们仔细评估与每个算法相关的性能和计算权衡。实验结果表明,联合训练的4D模型优于仅用一个单独的解码器训练的E2 E-ASR模型。此外,我们证明了所提出的一次通过波束搜索算法优于先前提出的CTC 注意解码。摘要:End-to-end automatic speech recognition (E2E-ASR) can be classified into several network architectures, such as connectionist temporal classification (CTC), recurrent neural network transducer (RNN-T), attention-based encoder-decoder, and mask-predict models. Each network architecture has advantages and disadvantages, leading practitioners to switch between these different models depending on application requirements. Instead of building separate models, we propose a joint modeling scheme where four decoders (CTC, RNN-T, attention, and mask-predict) share the same encoder -- we refer to this as 4D modeling. The 4D model is trained using multitask learning, which will bring model regularization and maximize the model robustness thanks to their complementary properties. To efficiently train the 4D model, we introduce a two-stage training strategy that stabilizes multitask learning. In addition, we propose three novel one-pass beam search algorithms by combining three decoders (CTC, RNN-T, and attention) to further improve performance. These three beam search algorithms differ in which decoder is used as the primary decoder. We carefully evaluate the performance and computational tradeoffs associated with each algorithm. Experimental results demonstrate that the jointly trained 4D model outperforms the E2E-ASR models trained with only one individual decoder. Furthermore, we demonstrate that the proposed one-pass beam search algorithm outperforms the previously proposed CTC attention decoding.
【19】 SYN2REAL: Leveraging Task Arithmetic for Mitigating Synthetic-Real Discrepancies in ASR Domain Adaptation
标题: SY 2 REAL:利用任务算法缓解SVR域自适应中的合成-真实差异
作者:Hsuan Su,Hua Farn,Shang-Tse Chen,Hung-yi Lee
链接:点击下载PDF文件
摘要:大型语言模型(LLM)的最新进展引入了“任务向量”的概念,这对各个领域都产生了重大影响,但在语音识别中仍然没有得到充分的研究。本文提出了一种新的“SYN2REAL”任务向量,用于自动语音识别(ASR)中的域自适应,特别是针对纯文本域。传统的合成语音的微调往往会导致性能下降,由于声学失配。为了解决这个问题,我们建议通过减去在真实语音和合成语音上微调的模型之间的参数差异来创建“SYN2REAL”向量。该矢量有效地桥接了两个域之间的间隙。SLURP数据集上的实验表明,我们的方法产生了11.15%的平均改善字错误率为看不见的目标域,突出了潜在的任务向量在增强语音域适应。摘要:Recent advancements in large language models (LLMs) have introduced the 'task vector' concept, which has significantly impacted various domains but remains underexplored in speech recognition. This paper presents a novel 'SYN2REAL' task vector for domain adaptation in automatic speech recognition (ASR), specifically targeting text-only domains. Traditional fine-tuning on synthetic speech often results in performance degradation due to acoustic mismatches. To address this issue, we propose creating a 'SYN2REAL' vector by subtracting the parameter differences between models fine-tuned on real and synthetic speech. This vector effectively bridges the gap between the two domains. Experiments on the SLURP dataset demonstrate that our approach yields an average improvement of 11.15% in word error rate for unseen target domains, highlighting the potential of task vectors in enhancing speech domain adaptation.
【20】 USM RNN-T model weights binarization
标题: USM RNN-T模型权重二值化
作者:Oleg Rybakov,Dmitriy Serdyuk,CJ Zheng
链接:点击下载PDF文件
摘要:大规模通用语音模型(USM)已经在生产中使用。然而,随着模型大小的增加,服务成本也会增加。大型模型的服务成本主要由模型大小决定,这就是为什么模型大小缩减是一个重要的研究课题。在这项工作中,我们专注于模型大小的减少,仅使用权重量化。我们提出了USM递归神经网络传感器(RNN-T)的权重二值化,并表明其模型大小可以减少15.9倍,而与float 32模型相比,字错误率(WER)仅增加1.9%。这使得它在实际应用中具有吸引力。摘要:Large-scale universal speech models (USM) are already used in production. However, as the model size grows, the serving cost grows too. Serving cost of large models is dominated by model size that is why model size reduction is an important research topic. In this work we are focused on model size reduction using weights only quantization. We present the weights binarization of USM Recurrent Neural Network Transducer (RNN-T) and show that its model size can be reduced by 15.9x times at cost of word error rate (WER) increase by only 1.9% in comparison to the float32 model. It makes it attractive for practical applications.
【21】 ConPCO: Preserving Phoneme Characteristics for Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
标题: ConPCO:保留音素特征以进行自动发音评估,利用对比有序规则化
作者:Bi-Cheng Yan,Wei-Cheng Chao,Jiun-Ting Li,Yi-Cheng Wang,Hsin-Wei Wang,Meng-Shin Lin,Berlin Chen
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:自动发音评估(APA)旨在评估第二语言(L2)学习者在目标语言中的发音水平。现有的努力通常利用回归模型进行熟练度分数预测,其中模型被训练以估计目标值,而不明确考虑特征空间中的音素意识。在本文中,我们提出了一个对比音素序数正则化(ConPCO)量身定制的基于回归的APA模型,以产生更多的音素判别功能,同时考虑回归目标之间的顺序关系。建议的ConPCO首先对齐的APA模型的音素表示和语音transmittance通过对比学习的文本嵌入。之后,通过调节特征空间中音素类别间和音素类别内之间的距离来保留音素特征,同时允许输出目标之间的顺序关系。我们进一步设计和开发了一个层次APA模型来评估我们的方法的有效性。在speechocean762基准数据集上进行的大量实验表明,我们的方法在一些前沿基线上的可行性和有效性。摘要:Automatic pronunciation assessment (APA) manages to evaluate the pronunciation proficiency of a second language (L2) learner in a target language. Existing efforts typically draw on regression models for proficiency score prediction, where the models are trained to estimate target values without explicitly accounting for phoneme-awareness in the feature space. In this paper, we propose a contrastive phonemic ordinal regularizer (ConPCO) tailored for regression-based APA models to generate more phoneme-discriminative features while considering the ordinal relationships among the regression targets. The proposed ConPCO first aligns the phoneme representations of an APA model and textual embeddings of phonetic transcriptions via contrastive learning. Afterward, the phoneme characteristics are retained by regulating the distances between inter- and intra-phoneme categories in the feature space while allowing for the ordinal relationships among the output targets. We further design and develop a hierarchical APA model to evaluate the effectiveness of our method. Extensive experiments conducted on the speechocean762 benchmark dataset suggest the feasibility and efficacy of our approach in relation to some cutting-edge baselines.
【22】 Keyword-Guided Adaptation of Automatic Speech Recognition
标题: 关键词引导的自动语音识别自适应
作者:Aviv Shamsian,Aviv Navon,Neta Glazer,Gill Hetz,Joseph Keshet
备注:Accepted to InterSpeech 2024
链接:点击下载PDF文件
摘要:自动语音识别(ASR)技术近年来取得了重大进展,在各个领域提供了准确的转录。然而,仍然存在一些挑战,特别是在嘈杂的环境和专业术语中。在本文中,我们提出了一种新的方法,改进的行话识别上下文偏置耳语为基础的模型。我们采用关键字定位模型,利用耳语编码器表示动态生成提示,引导解码器在转录过程中。我们介绍了两种方法来有效地引导解码器对这些提示:KG耳语,这是为了微调耳语解码器,和KG耳语-PT,学习提示前缀。我们的研究结果表明,在指定的关键字的识别精度和减少整体的单词错误率显着改善。具体来说,在看不见的语言泛化中,我们证明了WER比Whisper平均提高了5.1%。摘要:Automatic Speech Recognition (ASR) technology has made significant progress in recent years, providing accurate transcription across various domains. However, some challenges remain, especially in noisy environments and specialized jargon. In this paper, we propose a novel approach for improved jargon word recognition by contextual biasing Whisper-based models. We employ a keyword spotting model that leverages the Whisper encoder representation to dynamically generate prompts for guiding the decoder during the transcription process. We introduce two approaches to effectively steer the decoder towards these prompts: KG-Whisper, which is aimed at fine-tuning the Whisper decoder, and KG-Whisper-PT, which learns a prompt prefix. Our results show a significant improvement in the recognition accuracy of specified keywords and in reducing the overall word error rates. Specifically, in unseen language generalization, we demonstrate an average WER improvement of 5.1% over Whisper.
【23】 Selfsupervised learning for pathological speech detection
标题: 病态语音检测的自我监督学习
作者:Shakeel Ahmad Sheikh
备注:in Intersection of Book Chapter in Machine Leanring and Computational Social Sciences CRC (in progress) 2024
链接:点击下载PDF文件
摘要:言语产生是一种复杂的现象,其中大脑协调一系列涉及思维处理、运动规划和发音运动执行的过程。然而,各种过程的这种复杂执行容易受到各种神经退行性病理性言语障碍(例如帕金森病)的影响和破坏,从而导致构音障碍、失用症和其他病症。这些疾病导致以异常的言语模式和不精确的发音为特征的病理性言语。在临床环境中诊断这些言语障碍通常涉及听觉感知测试,这是耗时的,并且诊断可以根据临床医生的经验,偏见和诊断期间的认知负荷而有所不同。此外,与典型的神经说话者不同,患有语音病理或障碍的患者无法访问各种虚拟助手,如Alexa,Siri等。这些方法旨在提供有效和准确的语言障碍检测,从而促进及时干预和支持受这些条件影响的个人。这些方法主要在两个方面有所不同:使用的输入表示和分类器。由于数据有限,检测性能仍然低于标准。自监督学习(SSL)嵌入,如wav2vec2及其多语言版本,正在被探索作为一个有前途的途径,以提高性能。这些嵌入利用自监督学习技术从音频数据中提取丰富的表示,从而提供了一种潜在的解决方案,以解决标记数据稀缺所带来的限制。摘要:Speech production is a complex phenomenon, wherein the brain orchestrates a sequence of processes involving thought processing, motor planning, and the execution of articulatory movements. However, this intricate execution of various processes is susceptible to influence and disruption by various neurodegenerative pathological speech disorders, such as Parkinsons' disease, resulting in dysarthria, apraxia, and other conditions. These disorders lead to pathological speech characterized by abnormal speech patterns and imprecise articulation. Diagnosing these speech disorders in clinical settings typically involves auditory perceptual tests, which are time-consuming, and the diagnosis can vary among clinicians based on their experiences, biases, and cognitive load during the diagnosis. Additionally, unlike neurotypical speakers, patients with speech pathologies or impairments are unable to access various virtual assistants such as Alexa, Siri, etc. To address these challenges, several automatic pathological speech detection (PSD) approaches have been proposed. These approaches aim to provide efficient and accurate detection of speech disorders, thereby facilitating timely intervention and support for individuals affected by these conditions. These approaches mainly vary in two aspects: the input representations utilized and the classifiers employed. Due to the limited availability of data, the performance of detection remains subpar. Self-supervised learning (SSL) embeddings, such as wav2vec2, and their multilingual versions, are being explored as a promising avenue to improve performance. These embeddings leverage self-supervised learning techniques to extract rich representations from audio data, thereby offering a potential solution to address the limitations posed by the scarcity of labeled data.
【24】 A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
标题: 一种固定端到端ASC系统标记器中词汇量大小的成本最小化方法
作者:Sunil Kumar Kopparapu,Ashish Panda
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:与其中令牌的使用被限制为音素、双音素或三音素的混合语音识别系统不同,端到端ASR系统中的令牌的选择是从训练数据的文本语料库导出的。像字节对编码(BPE)和WordPiece这样的标记化算法的使用在识别语音识别系统的整个训练过程中使用的标记中很流行。流行的工具包,如ESPNet,为这些标记化算法使用预定义的词汇量(标记数量),但没有讨论词汇量是如何得出的。在本文中,我们构建了一个成本函数,假设令牌化过程是一个黑盒,以便选择最有利于构建端到端ASR的令牌数量。我们通过LibriSpeech 100小时集上的实验表明,当仔细选择令牌的数量时,端到端ASR系统的性能会提高。摘要:Unlike hybrid speech recognition systems where the use of tokens was restricted to phones, biphones or triphones the choice of tokens in the end-to-end ASR systems is derived from the text corpus of the training data. The use of tokenization algorithms like Byte Pair Encoding (BPE) and WordPiece is popular in identifying the tokens that are used in the overall training process of the speech recognition system. Popular toolkits, like ESPNet use a pre-defined vocabulary size (number of tokens) for these tokenization algorithms, but there is no discussion on how vocabulary size was derived. In this paper, we build a cost function, assuming the tokenization process to be a black-box to enable choosing the number of tokens which might most benefit building an end-to-end ASR. We show through experiments on LibriSpeech 100 hour set that the performance of an end-to-end ASR system improves when the number of tokens are chosen carefully.
eess.AS音频处理
【1】 The PESQetarian: On the Relevance of Goodhart's Law for Speech Enhancement标题: PESQetarian:论古德哈特定律与言语增强的相关性
作者:Danilo de Oliveira,Simon Welker,Julius Richter,Timo Gerkmann
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:为了获得改进的语音增强模型,研究人员通常专注于根据特定的工具度量来提高性能。然而,当在损失函数中使用相同的度量来优化模型时,它可能对给定度量看不到的方面有害。本文的目的是说明过拟合的语音增强模型的评估所使用的度量的风险。为此,我们引入了利用广泛使用的PESQ措施的增强模型。我们的“PESQetarian”模型在VB-DMD上达到3.82 PESQ,而在听力实验中得分很低。虽然获得的PESQ值3.82意味着VB-DMD基准测试的“最先进”PESQ性能,但我们的示例表明,当优化w.r.t.一个度量,对同一度量的孤立评估可能会产生误导。相反,其他指标应包括在评估中,并通过倾听确认由此产生的性能预测。摘要:To obtain improved speech enhancement models, researchers often focus on increasing performance according to specific instrumental metrics. However, when the same metric is used in a loss function to optimize models, it may be detrimental to aspects that the given metric does not see. The goal of this paper is to illustrate the risk of overfitting a speech enhancement model to the metric used for evaluation. For this, we introduce enhancement models that exploit the widely used PESQ measure. Our "PESQetarian" model achieves 3.82 PESQ on VB-DMD while scoring very poorly in a listening experiment. While the obtained PESQ value of 3.82 would imply "state-of-the-art" PESQ-performance on the VB-DMD benchmark, our examples show that when optimizing w.r.t. a metric, an isolated evaluation on the same metric may be misleading. Instead, other metrics should be included in the evaluation and the resulting performance predictions should be confirmed by listening.
【2】 Enhancing CTC-based speech recognition with diverse modeling units
标题: 利用不同的建模单元增强基于ATC的语音识别
作者:Shiyi Han,Zhihong Lei,Mingbin Xu,Xingyu Na,Zhen Huang
链接:点击下载PDF文件
摘要:近年来,端到端(E2 E)自动语音识别(ASR)模型的发展引人注目,这主要归功于Transformer等深度学习架构的进步。在E2 E系统之上,研究人员通过使用基于音素的模型对E2 E模型的N-best假设进行重新评分,实现了实质性的准确性提高。这提出了一个有趣的问题,即除了系统组合效应之外,这些改进来自何处。我们研究了驱动这些收益的潜在机制,并提出了一种有效的联合训练方法,其中E2 E模型与不同的建模单元联合训练。这种方法不仅使基于音素和字素的模型的优势保持一致,而且还揭示了以协同方式使用这些不同的建模单元可以显着提高模型的准确性。我们的研究结果提供了新的见解异构建模单元的最佳集成,在更强大和准确的ASR系统的发展。摘要:In recent years, the evolution of end-to-end (E2E) automatic speech recognition (ASR) models has been remarkable, largely due to advances in deep learning architectures like transformer. On top of E2E systems, researchers have achieved substantial accuracy improvement by rescoring E2E model's N-best hypotheses with a phoneme-based model. This raises an interesting question about where the improvements come from other than the system combination effect. We examine the underlying mechanisms driving these gains and propose an efficient joint training approach, where E2E models are trained jointly with diverse modeling units. This methodology does not only align the strengths of both phoneme and grapheme-based models but also reveals that using these diverse modeling units in a synergistic way can significantly enhance model accuracy. Our findings offer new insights into the optimal integration of heterogeneous modeling units in the development of more robust and accurate ASR systems.
【3】 Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
标题: 使用分层令牌-语义音频Transformer架构的多麦克风语音情感识别
作者:Ohad Cohen,Gershon Hazan,Sharon Gannot
链接:点击下载PDF文件
摘要:大多数情感识别系统在现实生活中的情况下(在野外场景中)失败,其中音频被混响污染。我们的研究探索了新的方法,以减轻语音情感识别(SER)算法的性能下降,并开发一个更强大的系统的不利条件。我们建议处理多麦克风信号来解决这些挑战并提高情感分类的准确性。我们采用了一个国家的最先进的Transformer模型,层次令牌语义音频Transformer(HTS-AT),处理多声道音频输入。我们评估两种策略:平均跨通道的梅尔频谱和总结补丁嵌入式表示。我们的多麦克风模型在真实混响环境中进行测试时,与单通道基线相比,具有卓越的性能。摘要:Most emotion recognition systems fail in real-life situations (in the wild scenarios) where the audio is contaminated by reverberation. Our study explores new methods to alleviate the performance degradation of Speech Emotion Recognition (SER) algorithms and develop a more robust system for adverse conditions. We propose processing multi-microphone signals to address these challenges and improve emotion classification accuracy. We adopt a state-of-the-art transformer model, the Hierarchical Token-semantic Audio Transformer (HTS-AT), to handle multi-channel audio inputs. We evaluate two strategies: averaging mel-spectrograms across channels and summing patch-embedded representations. Our multimicrophone model achieves superior performance compared to single-channel baselines when tested on real-world reverberant environments.
【4】 Reference Channel Selection by Multi-Channel Masking for End-to-End Multi-Channel Speech Enhancement
标题: 端到端多通道语音增强的多通道掩蔽参考通道选择
作者:Wang Dai,Xiaofei Li,Archontis Politis,Tuomas Virtanen
备注:Accepted by EUSIPCO 2024
链接:点击下载PDF文件
摘要:在端到端多通道语音增强中,指定一个麦克风信号作为处理参考的传统方法可能并不总是产生最佳结果。该限制尤其是在扬声器到麦克风距离变化的大型分布式麦克风阵列或扬声器或麦克风位置随时间变化的紧凑、高度定向的麦克风阵列的场景中。当前基于掩模的方法通常在训练期间固定参考通道,这使得不可能自适应地选择参考通道以获得最佳性能。为了解决这个问题,我们引入了一个自适应的方法来选择最佳的参考通道。我们的方法利用了多通道基于掩蔽的方案,其中多个掩蔽信号被组合以生成单通道输出信号。该增强信号然后用于损失计算,而参考干净语音基于最高尺度不变信号失真比(SI-SDR)来调整。在Spear Challenge模拟数据集D4上的实验结果表明,该方法优于使用固定参考通道和单通道掩蔽的传统方法摘要:In end-to-end multi-channel speech enhancement, the traditional approach of designating one microphone signal as the reference for processing may not always yield optimal results. The limitation is particularly in scenarios with large distributed microphone arrays with varying speaker-to-microphone distances or compact, highly directional microphone arrays where speaker or microphone positions change over time. Current mask-based methods often fix the reference channel during training, which makes it not possible to adaptively select the reference channel for optimal performance. To address this problem, we introduce an adaptive approach for selecting the optimal reference channel. Our method leverages a multi-channel masking-based scheme, where multiple masked signals are combined to generate a single-channel output signal. This enhanced signal is then used for loss calculation, while the reference clean speech is adjusted based on the highest scale-invariant signal-to-distortion ratio (SI-SDR). The experimental results on the Spear challenge simulated dataset D4 demonstrate the superiority of our proposed method over the conventional approach of using a fixed reference channel with single-channel masking
【5】 CoLLAB: A Collaborative Approach for Multilingual Abuse Detection
标题: CoLLAB:多语言滥用检测的协作方法
作者:Orchid Chetia Phukan,Yashasvi Chaurasia,Arun Balaji Buduru,Rajesh Sharma
链接:点击下载PDF文件
摘要:在这项研究中,我们研究了来自音频滥用检测(AAD)的双语预训练模型(PTM)的表示,这还没有被用于AAD。我们的研究结果表明,他们的优越性相比,其他PTM表示的ADIMA基准。此外,组合PTM表示增强了AAD性能。尽管有这些改进,但跨语言通用性的挑战仍然存在,某些语言需要在同一语言中进行培训。这就需要针对不同语言的单独模型,从而导致可扩展性、维护和资源分配问题,并阻碍了AAD系统在语言多样的现实环境中的实际部署。为了解决这个问题,我们引入了CoLLAB,这是一个新的框架,它不需要训练,并允许通过加权平均来无缝合并用不同语言训练的模型。这导致了一个统一的模型,在多种语言之间具有竞争力的AAD性能。摘要:In this study, we investigate representations from paralingual Pre-Trained model (PTM) for Audio Abuse Detection (AAD), which has not been explored for AAD. Our results demonstrate their superiority compared to other PTM representations on the ADIMA benchmark. Furthermore, combining PTM representations enhances AAD performance. Despite these improvements, challenges with cross-lingual generalizability still remain, and certain languages require training in the same language. This demands individual models for different languages, leading to scalability, maintenance, and resource allocation issues and hindering the practical deployment of AAD systems in linguistically diverse real-world environments. To address this, we introduce CoLLAB, a novel framework that doesn't require training and allows seamless merging of models trained in different languages through weight-averaging. This results in a unified model with competitive AAD performance across multiple languages.
【6】 Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
标题: 再次对话:通过分段级发言人重新分配改进会议转录系统
作者:Christoph Boeddeker,Tobias Cord-Landwehr,Reinhold Haeb-Umbach
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
摘要:日志化是会议转录系统中的一个重要组成部分,可以缓解语音增强的挑战,并将转录归因于正确的说话者。特别是在存在重叠或噪声语音的情况下,这些系统在可靠地分配正确的说话者标签方面存在问题,从而导致大量的说话者混淆错误。我们建议增加段级扬声器重新分配来解决这个问题。语音增强后,通过重新访问每个段的说话人属性,从最初的日志化阶段的说话人混淆错误显着减少。通过在不同系统配置和数据集上的实验,我们进一步证明了在各个领域的有效性和适用性。我们的研究结果表明,段级扬声器重新分配成功地纠正了至少40%的扬声器混淆字错误,突出了其潜力,提高会议转录系统中的日记准确性。摘要:Diarization is a crucial component in meeting transcription systems to ease the challenges of speech enhancement and attribute the transcriptions to the correct speaker. Particularly in the presence of overlapping or noisy speech, these systems have problems reliably assigning the correct speaker labels, leading to a significant amount of speaker confusion errors. We propose to add segment-level speaker reassignment to address this issue. By revisiting, after speech enhancement, the speaker attribution for each segment, speaker confusion errors from the initial diarization stage are significantly reduced. Through experiments across different system configurations and datasets, we further demonstrate the effectiveness and applicability in various domains. Our results show that segment-level speaker reassignment successfully rectifies at least 40% of speaker confusion word errors, highlighting its potential for enhancing diarization accuracy in meeting transcription systems.
【7】 RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification
标题: RevRIR:使用对比学习联合混响语音和房间脉冲响应嵌入并应用于房间形状分类
作者:Jacob Bitterman,Daniel Levi,Hilel Hagai Diamandi,Sharon Gannot,Tal Rosenwein
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文重点介绍了房间指纹识别,这是一项涉及分析音频记录以确定其被捕获的房间的特定音量和形状的任务。虽然从房间脉冲响应(RIR)确定基本房间参数是相对简单的,但是从语音信号这样做是繁琐的任务。为了解决这一挑战,我们引入了一个双编码器架构,便于直接从语音话语的房间参数的估计。在预训练期间,一个编码器接收RIR,而另一个处理混响语音信号。一个对比损失函数被用来嵌入语音和声学响应。在微调阶段,训练特定的分类任务。在测试阶段,只有混响的话语是可用的,其嵌入用于房间形状分类的任务。所提出的方案进行了广泛的评估,使用模拟的声学环境。摘要:This paper focuses on room fingerprinting, a task involving the analysis of an audio recording to determine the specific volume and shape of the room in which it was captured. While it is relatively straightforward to determine the basic room parameters from the Room Impulse Responses (RIR), doing so from a speech signal is a cumbersome task. To address this challenge, we introduce a dual-encoder architecture that facilitates the estimation of room parameters directly from speech utterances. During pre-training, one encoder receives the RIR while the other processes the reverberant speech signal. A contrastive loss function is employed to embed the speech and the acoustic response jointly. In the fine-tuning stage, the specific classification task is trained. In the test phase, only the reverberant utterance is available, and its embedding is used for the task of room shape classification. The proposed scheme is extensively evaluated using simulated acoustic environments.
【8】 Singing Voice Graph Modeling for SingFake Detection
标题: SingFake检测的歌声图建模
作者:Xuanjun Chen,Haibin Wu,Jyh-Shing Roger Jang,Hung-yi Lee
备注:Accepted by Interspeech 2024; Codebase available at: this https URL
链接:点击下载PDF文件
摘要:检测歌声深度伪造(SingFake)涉及确定歌声的真实性和版权。现有的语音deepfake检测模型一直在努力适应人类发声这一独特的歌唱声音领域中的不可见攻击。为了弥合这一差距,我们提出了一个突破性的SingGraph模型。该模型协同MERT声学音乐理解模型的音高和节奏分析与wav2vec2.0模型的歌词语言分析的能力。此外,我们提倡使用基于音乐领域知识的RawBoost和节拍匹配技术来增强歌声,从而提高SingFake检测性能。我们提出的方法在SingFake数据集内实现了新的最先进(SOTA)结果,在三种不同的场景中超过了以前的SOTA模型:它将看到的歌手的EER相对提高了13.2%,未看到的歌手提高了24.3%,使用不同编解码器的未看到的歌手提高了37.1%。摘要:Detecting singing voice deepfakes, or SingFake, involves determining the authenticity and copyright of a singing voice. Existing models for speech deepfake detection have struggled to adapt to unseen attacks in this unique singing voice domain of human vocalization. To bridge the gap, we present a groundbreaking SingGraph model. The model synergizes the capabilities of the MERT acoustic music understanding model for pitch and rhythm analysis with the wav2vec2.0 model for linguistic analysis of lyrics. Additionally, we advocate for using RawBoost and beat matching techniques grounded in music domain knowledge for singing voice augmentation, thereby enhancing SingFake detection performance. Our proposed method achieves new state-of-the-art (SOTA) results within the SingFake dataset, surpassing the previous SOTA model across three distinct scenarios: it improves EER relatively for seen singers by 13.2%, for unseen singers by 24.3%, and unseen singers using different codecs by 37.1%.
【9】 4D ASR: Joint Beam Search Integrating CTC, Attention, Transducer, and Mask Predict Decoders
标题: 4D ASB:集成了CTC、注意力、传感器和屏蔽预测解码器的联合射束搜索
作者:Yui Sudo,Muhammad Shakeel,Yosuke Fukumoto,Brian Yan,Jiatong Shi,Yifan Peng,Shinji Watanabe
备注:submitted to IEEEACM Transactions on Audio Speech and Language Processing
链接:点击下载PDF文件
摘要:端到端自动语音识别(E2 E-ASR)可以分为几种网络架构,如连接主义时间分类(CTC),递归神经网络转换器(RNN-T),基于注意力的编码器-解码器和掩码预测模型。每种网络架构都有优点和缺点,导致从业者根据应用需求在这些不同的模型之间切换。我们没有建立单独的模型,而是提出了一种联合建模方案,其中四个解码器(CTC,RNN-T,attention和mask-predict)共享同一个编码器-我们将其称为4D建模。4D模型使用多任务学习进行训练,这将带来模型正则化并最大化模型鲁棒性,这要归功于它们的互补特性。为了有效地训练4D模型,我们引入了一个稳定多任务学习的两阶段训练策略。此外,我们提出了三种新的一次通过波束搜索算法,通过结合三个解码器(CTC,RNN-T和注意力),以进一步提高性能。这三种波束搜索算法的不同之处在于哪个解码器被用作主解码器。我们仔细评估与每个算法相关的性能和计算权衡。实验结果表明,联合训练的4D模型优于仅用一个单独的解码器训练的E2 E-ASR模型。此外,我们证明了所提出的一次通过波束搜索算法优于先前提出的CTC 注意解码。摘要:End-to-end automatic speech recognition (E2E-ASR) can be classified into several network architectures, such as connectionist temporal classification (CTC), recurrent neural network transducer (RNN-T), attention-based encoder-decoder, and mask-predict models. Each network architecture has advantages and disadvantages, leading practitioners to switch between these different models depending on application requirements. Instead of building separate models, we propose a joint modeling scheme where four decoders (CTC, RNN-T, attention, and mask-predict) share the same encoder -- we refer to this as 4D modeling. The 4D model is trained using multitask learning, which will bring model regularization and maximize the model robustness thanks to their complementary properties. To efficiently train the 4D model, we introduce a two-stage training strategy that stabilizes multitask learning. In addition, we propose three novel one-pass beam search algorithms by combining three decoders (CTC, RNN-T, and attention) to further improve performance. These three beam search algorithms differ in which decoder is used as the primary decoder. We carefully evaluate the performance and computational tradeoffs associated with each algorithm. Experimental results demonstrate that the jointly trained 4D model outperforms the E2E-ASR models trained with only one individual decoder. Furthermore, we demonstrate that the proposed one-pass beam search algorithm outperforms the previously proposed CTC attention decoding.
【10】 SYN2REAL: Leveraging Task Arithmetic for Mitigating Synthetic-Real Discrepancies in ASR Domain Adaptation
标题: SY 2 REAL:利用任务算法缓解SVR域自适应中的合成-真实差异
作者:Hsuan Su,Hua Farn,Shang-Tse Chen,Hung-yi Lee
链接:点击下载PDF文件
摘要:大型语言模型(LLM)的最新进展引入了“任务向量”的概念,这对各个领域都产生了重大影响,但在语音识别中仍然没有得到充分的研究。本文提出了一种新的“SYN2REAL”任务向量,用于自动语音识别(ASR)中的域自适应,特别是针对纯文本域。传统的合成语音的微调往往会导致性能下降,由于声学失配。为了解决这个问题,我们建议通过减去在真实语音和合成语音上微调的模型之间的参数差异来创建“SYN2REAL”向量。该矢量有效地桥接了两个域之间的间隙。SLURP数据集上的实验表明,我们的方法产生了11.15%的平均改善字错误率为看不见的目标域,突出了潜在的任务向量在增强语音域适应。摘要:Recent advancements in large language models (LLMs) have introduced the 'task vector' concept, which has significantly impacted various domains but remains underexplored in speech recognition. This paper presents a novel 'SYN2REAL' task vector for domain adaptation in automatic speech recognition (ASR), specifically targeting text-only domains. Traditional fine-tuning on synthetic speech often results in performance degradation due to acoustic mismatches. To address this issue, we propose creating a 'SYN2REAL' vector by subtracting the parameter differences between models fine-tuned on real and synthetic speech. This vector effectively bridges the gap between the two domains. Experiments on the SLURP dataset demonstrate that our approach yields an average improvement of 11.15% in word error rate for unseen target domains, highlighting the potential of task vectors in enhancing speech domain adaptation.
【11】 USM RNN-T model weights binarization
标题: USM RNN-T模型权重二值化
作者:Oleg Rybakov,Dmitriy Serdyuk,CJ Zheng
链接:点击下载PDF文件
摘要:大规模通用语音模型(USM)已经在生产中使用。然而,随着模型大小的增加,服务成本也会增加。大型模型的服务成本主要由模型大小决定,这就是为什么模型大小缩减是一个重要的研究课题。在这项工作中,我们专注于模型大小的减少,仅使用权重量化。我们提出了USM递归神经网络传感器(RNN-T)的权重二值化,并表明其模型大小可以减少15.9倍,而与float 32模型相比,字错误率(WER)仅增加1.9%。这使得它在实际应用中具有吸引力。摘要:Large-scale universal speech models (USM) are already used in production. However, as the model size grows, the serving cost grows too. Serving cost of large models is dominated by model size that is why model size reduction is an important research topic. In this work we are focused on model size reduction using weights only quantization. We present the weights binarization of USM Recurrent Neural Network Transducer (RNN-T) and show that its model size can be reduced by 15.9x times at cost of word error rate (WER) increase by only 1.9% in comparison to the float32 model. It makes it attractive for practical applications.
【12】 ConPCO: Preserving Phoneme Characteristics for Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
标题: ConPCO:保留音素特征以进行自动发音评估,利用对比有序规则化
作者:Bi-Cheng Yan,Wei-Cheng Chao,Jiun-Ting Li,Yi-Cheng Wang,Hsin-Wei Wang,Meng-Shin Lin,Berlin Chen
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:自动发音评估(APA)旨在评估第二语言(L2)学习者在目标语言中的发音水平。现有的努力通常利用回归模型进行熟练度分数预测,其中模型被训练以估计目标值,而不明确考虑特征空间中的音素意识。在本文中,我们提出了一个对比音素序数正则化(ConPCO)量身定制的基于回归的APA模型,以产生更多的音素判别功能,同时考虑回归目标之间的顺序关系。建议的ConPCO首先对齐的APA模型的音素表示和语音transmittance通过对比学习的文本嵌入。之后,通过调节特征空间中音素类别间和音素类别内之间的距离来保留音素特征,同时允许输出目标之间的顺序关系。我们进一步设计和开发了一个层次APA模型来评估我们的方法的有效性。在speechocean762基准数据集上进行的大量实验表明,我们的方法在一些前沿基线上的可行性和有效性。摘要:Automatic pronunciation assessment (APA) manages to evaluate the pronunciation proficiency of a second language (L2) learner in a target language. Existing efforts typically draw on regression models for proficiency score prediction, where the models are trained to estimate target values without explicitly accounting for phoneme-awareness in the feature space. In this paper, we propose a contrastive phonemic ordinal regularizer (ConPCO) tailored for regression-based APA models to generate more phoneme-discriminative features while considering the ordinal relationships among the regression targets. The proposed ConPCO first aligns the phoneme representations of an APA model and textual embeddings of phonetic transcriptions via contrastive learning. Afterward, the phoneme characteristics are retained by regulating the distances between inter- and intra-phoneme categories in the feature space while allowing for the ordinal relationships among the output targets. We further design and develop a hierarchical APA model to evaluate the effectiveness of our method. Extensive experiments conducted on the speechocean762 benchmark dataset suggest the feasibility and efficacy of our approach in relation to some cutting-edge baselines.
【13】 RepCNN: Micro-sized, Mighty Models for Wakeword Detection
标题: RepCNN:用于唤醒词检测的微型、强大模型
作者:Arnav Kundu,Prateeth Nayak,Hywel Richards,Priyanka Padmanabhan,Devang Naik
链接:点击下载PDF文件
摘要:始终在线的机器学习模型需要非常低的内存和计算占用。它们受限的参数计数限制了模型的学习能力,以及通常的训练算法找到最佳参数的有效性。在这里,我们展示了一个小的卷积模型可以通过首先将其计算重构为一个更大的冗余多分支架构来更好地训练。然后,为了进行推理,我们用代数方法将训练好的模型重新参数化为具有更少参数的单分支形式,以降低内存占用和计算成本。使用这种技术,我们证明了我们始终在线的唤醒词检测器模型RepCNN在推理过程中提供了延迟和准确性之间的良好权衡。RepCNN重新参数化模型比单分支卷积模型准确43%,同时具有相同的运行时间。RepCNN还满足BC-ResNet等复杂架构的准确性,同时具有2倍的峰值内存使用和10倍的运行时间。摘要:Always-on machine learning models require a very low memory and compute footprint. Their restricted parameter count limits the model's capacity to learn, and the effectiveness of the usual training algorithms to find the best parameters. Here we show that a small convolutional model can be better trained by first refactoring its computation into a larger redundant multi-branched architecture. Then, for inference, we algebraically re-parameterize the trained model into the single-branched form with fewer parameters for a lower memory footprint and compute cost. Using this technique, we show that our always-on wake-word detector model, RepCNN, provides a good trade-off between latency and accuracy during inference. RepCNN re-parameterized models are 43% more accurate than a uni-branch convolutional model while having the same runtime. RepCNN also meets the accuracy of complex architectures like BC-ResNet, while having 2x lesser peak memory usage and 10x faster runtime.
【14】 Keyword-Guided Adaptation of Automatic Speech Recognition
标题: 关键词引导的自动语音识别自适应
作者:Aviv Shamsian,Aviv Navon,Neta Glazer,Gill Hetz,Joseph Keshet
备注:Accepted to InterSpeech 2024
链接:点击下载PDF文件
摘要:自动语音识别(ASR)技术近年来取得了重大进展,在各个领域提供了准确的转录。然而,仍然存在一些挑战,特别是在嘈杂的环境和专业术语中。在本文中,我们提出了一种新的方法,改进的行话识别上下文偏置耳语为基础的模型。我们采用关键字定位模型,利用耳语编码器表示动态生成提示,引导解码器在转录过程中。我们介绍了两种方法来有效地引导解码器对这些提示:KG耳语,这是为了微调耳语解码器,和KG耳语-PT,学习提示前缀。我们的研究结果表明,在指定的关键字的识别精度和减少整体的单词错误率显着改善。具体来说,在看不见的语言泛化中,我们证明了WER比Whisper平均提高了5.1%。摘要:Automatic Speech Recognition (ASR) technology has made significant progress in recent years, providing accurate transcription across various domains. However, some challenges remain, especially in noisy environments and specialized jargon. In this paper, we propose a novel approach for improved jargon word recognition by contextual biasing Whisper-based models. We employ a keyword spotting model that leverages the Whisper encoder representation to dynamically generate prompts for guiding the decoder during the transcription process. We introduce two approaches to effectively steer the decoder towards these prompts: KG-Whisper, which is aimed at fine-tuning the Whisper decoder, and KG-Whisper-PT, which learns a prompt prefix. Our results show a significant improvement in the recognition accuracy of specified keywords and in reducing the overall word error rates. Specifically, in unseen language generalization, we demonstrate an average WER improvement of 5.1% over Whisper.
【15】 PPINtonus: Early Detection of Parkinson's Disease Using Deep-Learning Tonal Analysis
标题: PPINtonus:使用深度学习音调分析早期检测帕金森病
作者:Varun Reddy
链接:点击下载PDF文件
摘要:PPINtonus是一种利用深度学习音调分析早期检测帕金森病(PD)的系统,为传统的神经系统检查提供了一种具有成本效益和可访问的替代方案。PPINtonus与帕金森语音项目(PVP)合作,采用半监督条件生成对抗网络来生成合成数据点,增强多层深度神经网络的训练数据集。与PRAAT语音学软件相结合,该网络可以在典型的家庭噪声条件下使用标准麦克风进行简单的120秒声乐测试,准确评估生物医学语音测量值。该模型的性能进行了验证,使用混淆矩阵,实现了令人印象深刻的92.5%的准确率与低假阴性率。PPINtonus的精确度为92.7%,是早期PD检测的可靠工具。PPINtonus的非侵入性和有效方法可以通过及时干预和管理实现早期诊断并改善数百万PD患者的生活质量,从而使发展中国家受益匪浅。摘要:PPINtonus is a system for the early detection of Parkinson's Disease (PD) utilizing deep-learning tonal analysis, providing a cost-effective and accessible alternative to traditional neurological examinations. Partnering with the Parkinson's Voice Project (PVP), PPINtonus employs a semi-supervised conditional generative adversarial network to generate synthetic data points, enhancing the training dataset for a multi-layered deep neural network. Combined with PRAAT phonetics software, this network accurately assesses biomedical voice measurement values from a simple 120-second vocal test performed with a standard microphone in typical household noise conditions. The model's performance was validated using a confusion matrix, achieving an impressive 92.5 % accuracy with a low false negative rate. PPINtonus demonstrated a precision of 92.7 %, making it a reliable tool for early PD detection. The non-intrusive and efficient methodology of PPINtonus can significantly benefit developing countries by enabling early diagnosis and improving the quality of life for millions of PD patients through timely intervention and management.
【16】 Selfsupervised learning for pathological speech detection
标题: 病态语音检测的自我监督学习
作者:Shakeel Ahmad Sheikh
备注:in Intersection of Book Chapter in Machine Leanring and Computational Social Sciences CRC (in progress) 2024
链接:点击下载PDF文件
摘要:言语产生是一种复杂的现象,其中大脑协调一系列涉及思维处理、运动规划和发音运动执行的过程。然而,各种过程的这种复杂执行容易受到各种神经退行性病理性言语障碍(例如帕金森病)的影响和破坏,从而导致构音障碍、失用症和其他病症。这些疾病导致以异常的言语模式和不精确的发音为特征的病理性言语。在临床环境中诊断这些言语障碍通常涉及听觉感知测试,这是耗时的,并且诊断可以根据临床医生的经验,偏见和诊断期间的认知负荷而有所不同。此外,与典型的神经说话者不同,患有语音病理或障碍的患者无法访问各种虚拟助手,如Alexa,Siri等。这些方法旨在提供有效和准确的语言障碍检测,从而促进及时干预和支持受这些条件影响的个人。这些方法主要在两个方面有所不同:使用的输入表示和分类器。由于数据有限,检测性能仍然低于标准。自监督学习(SSL)嵌入,如wav2vec2及其多语言版本,正在被探索作为一个有前途的途径,以提高性能。这些嵌入利用自监督学习技术从音频数据中提取丰富的表示,从而提供了一种潜在的解决方案,以解决标记数据稀缺所带来的限制。摘要:Speech production is a complex phenomenon, wherein the brain orchestrates a sequence of processes involving thought processing, motor planning, and the execution of articulatory movements. However, this intricate execution of various processes is susceptible to influence and disruption by various neurodegenerative pathological speech disorders, such as Parkinsons' disease, resulting in dysarthria, apraxia, and other conditions. These disorders lead to pathological speech characterized by abnormal speech patterns and imprecise articulation. Diagnosing these speech disorders in clinical settings typically involves auditory perceptual tests, which are time-consuming, and the diagnosis can vary among clinicians based on their experiences, biases, and cognitive load during the diagnosis. Additionally, unlike neurotypical speakers, patients with speech pathologies or impairments are unable to access various virtual assistants such as Alexa, Siri, etc. To address these challenges, several automatic pathological speech detection (PSD) approaches have been proposed. These approaches aim to provide efficient and accurate detection of speech disorders, thereby facilitating timely intervention and support for individuals affected by these conditions. These approaches mainly vary in two aspects: the input representations utilized and the classifiers employed. Due to the limited availability of data, the performance of detection remains subpar. Self-supervised learning (SSL) embeddings, such as wav2vec2, and their multilingual versions, are being explored as a promising avenue to improve performance. These embeddings leverage self-supervised learning techniques to extract rich representations from audio data, thereby offering a potential solution to address the limitations posed by the scarcity of labeled data.
【17】 Cluster-to-Predict Affect Contours from Speech
标题: 预测者影响言语轮廓
作者:Gökhan Kuşçu,Engin Erzin
备注:8 pages, 3 figures
链接:点击下载PDF文件
摘要:连续情绪识别(CER)旨在跟踪一个人的情绪状态随时间的动态变化。本文提出了一种新的方法来翻译CER从语音的动态影响轮廓集群的预测问题,其中的影响轮廓被定义为在一个时间窗口中的注释的影响属性的轮廓。我们的方法定义了一个聚类预测(C2P)框架,该框架学习影响轮廓聚类,这些聚类是从语音中以更高的精度预测的。为了实现这一点,C2P运行一个无监督的迭代优化过程,通过最小化聚类损失和语音驱动的影响轮廓预测损失来学习影响轮廓聚类。我们的客观研究结果表明,语音驱动的聚类唤醒和效价属性的价值。在RECOLA数据集上进行的实验产生了有希望的分类结果,在我们的四类语音驱动的情感轮廓预测模型中,唤醒的F1分数为0.84,效价为0.75。摘要:Continuous emotion recognition (CER) aims to track the dynamic changes in a person's emotional state over time. This paper proposes a novel approach to translating CER into a prediction problem of dynamic affect-contour clusters from speech, where the affect-contour is defined as the contour of annotated affect attributes in a temporal window. Our approach defines a cluster-to-predict (C2P) framework that learns affect-contour clusters, which are predicted from speech with higher precision. To achieve this, C2P runs an unsupervised iterative optimization process to learn affect-contour clusters by minimizing both clustering loss and speech-driven affect-contour prediction loss. Our objective findings demonstrate the value of speech-driven clustering for both arousal and valence attributes. Experiments conducted on the RECOLA dataset yielded promising classification results, with F1 scores of 0.84 for arousal and 0.75 for valence in our four-class speech-driven affect-contour prediction model.
【18】 Combining X-Vectors and Bayesian Batch Active Learning: Two-Stage Active Learning Pipeline for Speech Recognition
标题: 结合X-Vector和Bayesian批量主动学习:语音识别的两阶段主动学习管道
作者:Ognjen Kundacina,Vladimir Vincan,Dragisa Miskovic
链接:点击下载PDF文件
摘要:强调以数据为中心的人工智能方法,本文介绍了一种新的两阶段主动学习(AL)管道自动语音识别(ASR),结合无监督和监督AL方法。第一阶段利用无监督AL通过使用x-向量聚类从未标记的语音数据中进行不同的样本选择,从而为随后的监督AL建立一个强大的初始数据集。第二阶段采用了监督AL策略,并采用了专门为ASR开发的批量AL方法,旨在选择多样化和信息丰富的样本批次。在这里,样本多样性也是使用x向量聚类来实现的,而信息量最大的样本是使用为ASR定制的贝叶斯AL方法来识别的,该方法将蒙特卡洛丢弃调整为近似贝叶斯推断。这种方法可以实现精确的不确定性估计,从而增强ASR模型训练,并显著降低数据要求。与同类、异构和OOD测试集上的竞争方法相比,我们的方法表现出了卓越的性能,这表明策略性样本选择和创新的贝叶斯建模可以在基于深度学习的ASR应用中大大优化标记工作和数据利用率。摘要:Emphasizing a data-centric AI approach, this paper introduces a novel two-stage active learning (AL) pipeline for automatic speech recognition (ASR), combining unsupervised and supervised AL methods. The first stage utilizes unsupervised AL by using x-vectors clustering for diverse sample selection from unlabeled speech data, thus establishing a robust initial dataset for the subsequent supervised AL. The second stage incorporates a supervised AL strategy, with a batch AL method specifically developed for ASR, aimed at selecting diverse and informative batches of samples. Here, sample diversity is also achieved using x-vectors clustering, while the most informative samples are identified using a Bayesian AL method tailored for ASR with an adaptation of Monte Carlo dropout to approximate Bayesian inference. This approach enables precise uncertainty estimation, thereby enhancing ASR model training with significantly reduced data requirements. Our method has shown superior performance compared to competing methods on homogeneous, heterogeneous, and OOD test sets, demonstrating that strategic sample selection and innovative Bayesian modeling can substantially optimize both labeling effort and data utilization in deep learning-based ASR applications.
【19】 A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
标题: 一种固定端到端ASC系统标记器中词汇量大小的成本最小化方法
作者:Sunil Kumar Kopparapu,Ashish Panda
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:与其中令牌的使用被限制为音素、双音素或三音素的混合语音识别系统不同,端到端ASR系统中的令牌的选择是从训练数据的文本语料库导出的。像字节对编码(BPE)和WordPiece这样的标记化算法的使用在识别语音识别系统的整个训练过程中使用的标记中很流行。流行的工具包,如ESPNet,为这些标记化算法使用预定义的词汇量(标记数量),但没有讨论词汇量是如何得出的。在本文中,我们构建了一个成本函数,假设令牌化过程是一个黑盒,以便选择最有利于构建端到端ASR的令牌数量。我们通过LibriSpeech 100小时集上的实验表明,当仔细选择令牌的数量时,端到端ASR系统的性能会提高。摘要:Unlike hybrid speech recognition systems where the use of tokens was restricted to phones, biphones or triphones the choice of tokens in the end-to-end ASR systems is derived from the text corpus of the training data. The use of tokenization algorithms like Byte Pair Encoding (BPE) and WordPiece is popular in identifying the tokens that are used in the overall training process of the speech recognition system. Popular toolkits, like ESPNet use a pre-defined vocabulary size (number of tokens) for these tokenization algorithms, but there is no discussion on how vocabulary size was derived. In this paper, we build a cost function, assuming the tokenization process to be a black-box to enable choosing the number of tokens which might most benefit building an end-to-end ASR. We show through experiments on LibriSpeech 100 hour set that the performance of an end-to-end ASR system improves when the number of tokens are chosen carefully.
【20】 Gated Low-rank Adaptation for personalized Code-Switching Automatic Speech Recognition on the low-spec devices
标题: 门控低等级自适应在低规格设备上实现个性化代码切换自动语音识别
作者:Gwantae Kim,Bokyeung Lee,Donghyeon Kim,Hanseok Ko
Journal-ref:ICASSP 2024 Workshop(HSCMA 2024) paper
链接:点击下载PDF文件
摘要:近年来,人们对在低规格设备(如移动设备和仅CPU设备)上使用个性化的大型模型越来越感兴趣。然而,在设备上利用个性化的大模型是低效的,并且有时由于计算成本而受到限制。为了解决这个问题,本文提出了权重分离方法,使用参数有效的微调方法来最小化设备上的模型权重。此外,有些人在一个话语中说多种语言,即所谓的代码转换,个性化的ASR模型是必要的,以解决这种情况。然而,目前的多语种语音识别模型仅限于识别每个话语中的单一语言。为了解决这个问题,我们提出了代码切换语音识别模型,将微调单语和多语言语音识别模型。此外,我们引入了一个门控低秩自适应(GLoRA)的参数有效的微调与最小的性能下降。我们的实验,进行韩语-英语代码切换数据集,表明微调语音识别模型的代码切换超越传统的代码切换语音识别模型从头开始训练的性能。此外,与传统LoRA相比,GLoRA增强了参数有效的微调性能。摘要:In recent times, there has been a growing interest in utilizing personalized large models on low-spec devices, such as mobile and CPU-only devices. However, utilizing a personalized large model in the on-device is inefficient, and sometimes limited due to computational cost. To tackle the problem, this paper presents the weights separation method to minimize on-device model weights using parameter-efficient fine-tuning methods. Moreover, some people speak multiple languages in an utterance, as known as code-switching, the personalized ASR model is necessary to address such cases. However, current multilingual speech recognition models are limited to recognizing a single language within each utterance. To tackle this problem, we propose code-switching speech recognition models that incorporate fine-tuned monolingual and multilingual speech recognition models. Additionally, we introduce a gated low-rank adaptation(GLoRA) for parameter-efficient fine-tuning with minimal performance degradation. Our experiments, conducted on Korean-English code-switching datasets, demonstrate that fine-tuning speech recognition models for code-switching surpasses the performance of traditional code-switching speech recognition models trained from scratch. Furthermore, GLoRA enhances parameter-efficient fine-tuning performance compared to conventional LoRA.
【21】 Breaking Walls: Pioneering Automatic Speech Recognition for Central Kurdish: End-to-End Transformer Paradigm
标题: 破墙:库尔德中部开创自动语音识别:端到端Transformer范式
作者:Abdulhady Abas Abdullah,Hadi Veisi,Tarik Rashid
备注:
链接:点击下载PDF文件
摘要:自动语音识别(Automatic Speech Recognition,简称ASR)是语音处理领域的一个重要分支,目前已广泛应用于实际应用中,并采用了多种技术来实现,其中人工神经网络是最常用的技术。提高性能,使这些系统对噪声具有鲁棒性,并为低资源语言开发这种技术是当前的挑战之一。本文讨论了中央库尔德语(CKB),作为一种低资源的语言,使用端到端的Transformers器的ASR系统的发展。库尔德语作为一种印欧语言,分为三种主要方言,即:中央库尔德人(即,北库尔德语(Kirmanji)和南库尔德语(有3000多万人使用)。在这项研究中,一个大小为224小时的语音语料库收集使用各种来源。然后,使用该语料库对基于transformer的声学模型进行训练。在训练声学模型中还利用了迁移学习技术。由于这些努力,我们的最佳模型在Asosoft测试集上获得了最先进的结果,字错误率(WER)为13%。这一成就标志着中央库尔德语ASR技术的显著进步,特别是在低资源语言的背景下。摘要:Automatic Speech Recognition (ASR), as an interesting field of speech processing, is utilized in real applications and is implemented using various techniques amongst which the artificial neural network is the most popular. Increasing the performance, making these systems robust to noise and developing this technology for low-resource languages is among the current challenges. This paper addresses the development of an ASR system for the Central Kurdish language (CKB), as a low-resource language, using end to end transformers. Kurdish, as an Indo-European language, is categorized into three main dialects, i.e., Central Kurdish (i.e., Sorani), North Kurdish (Kirmanji), and South Kurdish which is spoken by more than 30 million people. In this research, a speech corpus of size 224 hours is collected using various sources. Then, this corpus is used to train the transformer-based acoustic model. A transfer learning technique is also utilized in training acoustic models. As a result of these efforts, our optimal model attains state-of-the-art results on the Asosoft test set, achieving a Word Error Rate (WER) of 13%. This accomplishment signifies a notable advancement in ASR technology for the Central Kurdish language, particularly in the context of low-resource languages.
【22】 Less Peaky and More Accurate CTC Forced Alignment by Label Priors
标题: 不那么尖峰、更准确的CTC通过标签先验强制对齐
作者:Ruizhe Huang,Xiaohui Zhang,Zhaoheng Ni,Li Sun,Moto Hira,Jeff Hwang,Vimal Manohar,Vineel Pratap,Matthew Wiesner,Shinji Watanabe,Daniel Povey,Sanjeev Khudanpur
备注:Accepted by ICASSP 2024. Github repo: this https URL
链接:点击下载PDF文件
摘要:连接主义时间分类(CTC)模型具有峰值输出分布。这种行为对于自动语音识别(ASR)来说不是问题,但是它可能导致不准确的强制对齐(FA),特别是在更细的粒度下,例如,音素层次本文旨在通过利用标签先验来缓解CTC的峰值行为并提高其对强制对齐生成的适用性,从而在训练期间提高并最大化包含较少空白的对齐路径的分数。因此,我们的CTC模型产生更少的峰值后验,并且能够更准确地预测令牌的偏移,除了它们的开始。它优于标准的CTC模型和一种基于几何学的方法,在Buckeye和TIMIT数据上测量的音素和单词边界错误(PBE和WBE)中获得12-40%的CTC令牌偏移时间戳。与最广泛使用的FA工具包Montreal Forced Aligner(MFA)相比,我们的方法在Buckeye上的PBE WBE上的表现相似,但在TIMIT上落后于MFA。尽管如此,我们的方法具有更简单的训练管道和更好的运行时效率。我们的训练配方和预训练模型在TorchAudio中发布。摘要:Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can cause inaccurate forced alignments (FA), especially at finer granularity, e.g., phoneme level. This paper aims at alleviating the peaky behavior for CTC and improve its suitability for forced alignment generation, by leveraging label priors, so that the scores of alignment paths containing fewer blanks are boosted and maximized during training. As a result, our CTC model produces less peaky posteriors and is able to more accurately predict the offset of the tokens besides their onset. It outperforms the standard CTC model and a heuristics-based approach for obtaining CTC's token offset timestamps by 12-40% in phoneme and word boundary errors (PBE and WBE) measured on the Buckeye and TIMIT data. Compared with the most widely used FA toolkit Montreal Forced Aligner (MFA), our method performs similarly on PBE WBE on Buckeye, yet falls behind MFA on TIMIT. Nevertheless, our method has a much simpler training pipeline and better runtime efficiency. Our training recipe and pretrained model are released in TorchAudio.
【23】 PhoWhisper: Automatic Speech Recognition for Vietnamese
标题: PhoWhisper:越南语自动语音识别
作者:Thanh-Thien Le,Linh The Nguyen,Dat Quoc Nguyen
备注:Accepted to ICLR 2024 Tiny Papers Track
链接:点击下载PDF文件
摘要:我们介绍PhoWhisper在越南语自动语音识别的五个版本。PhoWhisper的鲁棒性是通过对包含不同越南口音的844小时数据集的Whisper模型进行微调来实现的。我们的实验研究展示了PhoWhisper在基准越南ASR数据集上的最先进性能。我们有开源的PhoWhisper在:https: github.com VinAIResearch PhoWhisper摘要:We introduce PhoWhisper in five versions for Vietnamese automatic speech recognition. PhoWhisper's robustness is achieved through fine-tuning the Whisper model on an 844-hour dataset that encompasses diverse Vietnamese accents. Our experimental study demonstrates state-of-the-art performances of PhoWhisper on benchmark Vietnamese ASR datasets. We have open-sourced PhoWhisper at: https: github.com VinAIResearch PhoWhisper
【24】 Hear Me, See Me, Understand Me: Audio-Visual Autism Behavior Recognition
标题: 听到我、看到我、理解我:视听自闭症行为识别
作者:Shijian Deng,Erin E. Kosloski,Siddhi Patel,Zeke A. Barnett,Yiyang Nan,Alexander Kaplan,Sisira Aarukapalli,William T. Doan,Matthew Wang,Harsh Singh,Pamela R. Rollins,Yapeng Tian
链接:点击下载PDF文件
摘要:在这篇文章中,我们介绍了一个新的视听自闭症行为识别问题,其中包括社会行为识别,这是人工智能辅助自闭症筛查研究中忽略的一个重要方面。我们将手头的任务定义为视听自闭症行为识别,它使用音频和视觉线索,包括音频中的任何语音,来识别自闭症相关行为。为了促进这一新的研究方向,我们收集了一个视听自闭症谱系数据集(AV-ASD),这是目前使用行为方法进行自闭症筛查的最大视频数据集。它涵盖了广泛的自闭症相关行为,包括与社会沟通和互动有关的行为。为了为进一步研究这个新问题铺平道路,我们深入探索了在不同模态中利用基础模型和多模态大型语言模型。我们在AV-ASD数据集上的实验表明,整合音频,视觉和语音模态显着提高了自闭症行为识别的性能。此外,我们探索了在多模态大型语言模型中使用事后到临时管道,以研究其在自闭症行为识别过程中增强模型解释能力的潜力。我们将发布我们的数据集,代码和预训练模型。摘要:In this article, we introduce a novel problem of audio-visual autism behavior recognition, which includes social behavior recognition, an essential aspect previously omitted in AI-assisted autism screening research. We define the task at hand as one that is audio-visual autism behavior recognition, which uses audio and visual cues, including any speech present in the audio, to recognize autism-related behaviors. To facilitate this new research direction, we collected an audio-visual autism spectrum dataset (AV-ASD), currently the largest video dataset for autism screening using a behavioral approach. It covers an extensive range of autism-associated behaviors, including those related to social communication and interaction. To pave the way for further research on this new problem, we intensively explored leveraging foundation models and multimodal large language models across different modalities. Our experiments on the AV-ASD dataset demonstrate that integrating audio, visual, and speech modalities significantly enhances the performance in autism behavior recognition. Additionally, we explored the use of a post-hoc to ad-hoc pipeline in a multimodal large language model to investigate its potential to augment the model's explanatory capability during autism behavior recognition. We will release our dataset, code, and pre-trained models.
【25】 Physics and geometry informed neural operator network with application to acoustic scattering
标题: 物理和几何知识的神经操作网络及其在声散射中的应用
作者:Siddharth Nair,Timothy F. Walsh,Greg Pickrell,Fabio Semperlotti
备注:20 pages of main text, 9 figures
链接:点击下载PDF文件
摘要:本文介绍了一种基于物理和几何信息的神经算子网络,并将其应用于声散射的正演模拟。开发能够学习不同计算域的解算子的几何信息深度学习模型对于各种工程应用来说是一个普遍重要的问题。为此,我们提出了一个物理信息的深度算子网络(DeepONet),能够使用基于非均匀有理B样条(NURBS)的几何参数化方法预测任意形状散射体的散射压力场。这种方法也导致在非平凡的散射几何形状的简约表示。与现有的基于物理的方法相比,当改变计算域时,需要重新评估模型,我们训练的模型能够在几秒钟内学习可以近似物理一致的散射压力场的解算子,对于任意刚性散射体形状;因此,前向模拟的计算时间可以改进与传统的正向求解器相比,可以减少(即减少)数量级。此外,该方法可以评估分散的压力场,而不需要标记的训练数据。提出的理论方法后,还提供了一个全面的数值研究来说明这种方法的显着能力来模拟任意散射体几何形状的任意组合所产生的声压场。这些结果突出了独特的泛化能力的建议运营商学习方法。摘要:In this paper, we introduce a physics and geometry informed neural operator network with application to the forward simulation of acoustic scattering. The development of geometry informed deep learning models capable of learning a solution operator for different computational domains is a problem of general importance for a variety of engineering applications. To this end, we propose a physics-informed deep operator network (DeepONet) capable of predicting the scattered pressure field for arbitrarily shaped scatterers using a geometric parameterization approach based on non-uniform rational B-splines (NURBS). This approach also results in parsimonious representations of non-trivial scatterer geometries. In contrast to existing physics-based approaches that require model re-evaluation when changing the computational domains, our trained model is capable of learning solution operator that can approximate physically-consistent scattered pressure field in just a few seconds for arbitrary rigid scatterer shapes; it follows that the computational time for forward simulations can improve (i.e. be reduced) by orders of magnitude in comparison to the traditional forward solvers. In addition, this approach can evaluate the scattered pressure field without the need for labeled training data. After presenting the theoretical approach, a comprehensive numerical study is also provided to illustrate the remarkable ability of this approach to simulate the acoustic pressure fields resulting from arbitrary combinations of arbitrary scatterer geometries. These results highlight the unique generalization capability of the proposed operator learning approach.
【26】 Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
标题: 音频曼巴:音频表示学习的双向状态空间模型
作者:Mehmet Hamza Erol,Arda Senocak,Jiu Feng,Joon Son Chung
备注:Code is available at this https URL
链接:点击下载PDF文件
摘要:Transformers已迅速成为音频分类的首选,超过了基于CNN的方法。然而,音频频谱图Transformers(AST)表现出二次缩放由于自我注意。消除这种二次自我注意力成本提出了一个有吸引力的方向。最近,状态空间模型(SSM),如Mamba,在语言和视觉任务中表现出了潜力。在这项研究中,我们探讨是否依赖于自我注意是必要的音频分类任务。通过引入音频曼巴(AuM),第一个自我注意力自由,纯粹基于SSM的音频分类模型,我们的目标是解决这个问题。我们在各种音频数据集上评估AuM-包括六个不同的基准-与成熟的AST模型相比,它实现了相当或更好的性能。摘要:Transformers have rapidly become the preferred choice for audio classification, surpassing methods based on CNNs. However, Audio Spectrogram Transformers (ASTs) exhibit quadratic scaling due to self-attention. The removal of this quadratic self-attention cost presents an appealing direction. Recently, state space models (SSMs), such as Mamba, have demonstrated potential in language and vision tasks in this regard. In this study, we explore whether reliance on self-attention is necessary for audio classification tasks. By introducing Audio Mamba (AuM), the first self-attention-free, purely SSM-based model for audio classification, we aim to address this question. We evaluate AuM on various audio datasets - comprising six different benchmarks - where it achieves comparable or better performance compared to well-established AST model.
【27】 ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings
标题: ASoBO:为会议中的远距离发言人进行细心的束流器选择
作者:Theo Mariotte,Anthony Larcher,Silvio Montresor,Jean-Hugh Thomas
备注:5 pages, 2 figures, 2 tables, accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:说话人日志化(Speaker Diarization,SD)的目的是将属于同一说话人的语音片段进行分组。许多语音处理应用程序(如丰富的会议转录)都需要执行此任务。在这种情况下,远距离麦克风阵列通常捕获音频信号。波束成形,即,空间滤波是处理多麦克风音频数据的常见实践。然而,它往往需要一个明确的定位的有源源,以引导过滤器。本文提出了一种基于自注意的算法来选择一组固定的空间滤波器的输出。该方法作为一个特征提取器的联合语音活动(VAD)和重叠语音检测(OSD)。然后从检测到的片段推断出说话人日记。该方法显示了令人信服的远程VAD,OSD和SD性能,例如AISHELL-4数据集上的14.5% DER。自我注意权重的分析证明了它们的可解释性,因为它们与说话者的角度位置相关。摘要:Speaker Diarization (SD) aims at grouping speech segments that belong to the same speaker. This task is required in many speech-processing applications, such as rich meeting transcription. In this context, distant microphone arrays usually capture the audio signal. Beamforming, i.e., spatial filtering, is a common practice to process multi-microphone audio data. However, it often requires an explicit localization of the active source to steer the filter. This paper proposes a self-attention-based algorithm to select the output of a bank of fixed spatial filters. This method serves as a feature extractor for joint Voice Activity (VAD) and Overlapped Speech Detection (OSD). The speaker diarization is then inferred from the detected segments. The approach shows convincing distant VAD, OSD, and SD performance, e.g. 14.5% DER on the AISHELL-4 dataset. The analysis of the self-attention weights demonstrates their explainability, as they correlate with the speaker's angular locations.
【28】 Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
标题: 使用MaskAutoEncoder进行广义假音频检测的以学生为中心的学习
作者:Xiaopeng Wang,Ruibo Fu,Zhengqi Wen,Zhiyong Wang,Yuankun Xie,Yukun Liu,Jianhua Tao,Xuefei Liu,Yongwei Li,Xin Qi,Yi Lu,Shuchen Shi
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:由于新的欺骗技术的出现,虚假音频检测(FAD)的推广是至关重要的。传统的FAD方法通常仅关注于区分真实音频和已知的欺骗音频。我们提出了一个以神经网络为中心的学习(GFL)框架引导,旨在高度广义的FAD,称为GFL-FAD。该方法结合了基于音频重建的反事实推理增强表示(CRER),使用掩码自动编码器(MAE)架构来准确地建模真实的音频特征。为了减少训练过程中欺骗音频的影响,我们引入了真正的音频重建损失,保持专注于学习真正的数据特征。此外,内容相关的瓶颈(BN)功能提取的MAE补充知识的原始音频。这些BN特征自适应地与CRER融合以进一步提高鲁棒性。我们的方法在ASVspoof 2019 LA上实现了最先进的性能,EER为0.25%。摘要:The generalization of Fake Audio Detection (FAD) is critical due to the emergence of new spoofing techniques. Traditional FAD methods often focus solely on distinguishing between genuine and known spoofed audio. We propose a Genuine-Focused Learning (GFL) framework guided, aiming for highly generalized FAD, called GFL-FAD. This method incorporates a Counterfactual Reasoning Enhanced Representation (CRER) based on audio reconstruction using the Mask AutoEncoder (MAE) architecture to accurately model genuine audio features. To reduce the influence of spoofed audio during training, we introduce a genuine audio reconstruction loss, maintaining the focus on learning genuine data features. In addition, content-related bottleneck (BN) features are extracted from the MAE to supplement the knowledge of the original audio. These BN features are adaptively fused with CRER to further improve robustness. Our method achieves state-of-the-art performance with an EER of 0.25% on ASVspoof2019 LA.
【29】 Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion strategy
标题: 广义源跟踪:检测具有真实重点和虚假分散策略的新型音频Deepfake算法
作者:Yuankun Xie,Ruibo Fu,Zhengqi Wen,Zhiyong Wang,Xiaopeng Wang,Haonnan Cheng,Long Ye,Jianhua Tao
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:随着deepfake音频的扩散,迫切需要调查它们的归属。目前的来源追踪方法可以有效地区分在分布(ID)类别。然而,deepfake算法的快速发展对准确识别非分布(OOD)新型deepfake算法提出了严峻的挑战。在本文中,我们提出了用于音频deepfake算法识别的真实强调和虚假分散(REFD)策略,证明了其在识别OOD样本的同时区分ID样本的有效性。为了有效的OOD检测,我们首先探索了当前的事后OOD方法,并提出了NSD,这是一种新的OOD方法,通过考虑特征和logits分数的相似性来识别新的deepfake算法。REFD在2023年音频Deepfake检测挑战赛Track3中作为单一系统获得了86.83%的F1分数,展示了其最先进的性能。摘要:With the proliferation of deepfake audio, there is an urgent need to investigate their attribution. Current source tracing methods can effectively distinguish in-distribution (ID) categories. However, the rapid evolution of deepfake algorithms poses a critical challenge in the accurate identification of out-of-distribution (OOD) novel deepfake algorithms. In this paper, we propose Real Emphasis and Fake Dispersion (REFD) strategy for audio deepfake algorithm recognition, demonstrating its effectiveness in discriminating ID samples while identifying OOD samples. For effective OOD detection, we first explore current post-hoc OOD methods and propose NSD, a novel OOD approach in identifying novel deepfake algorithms through the similarity consideration of both feature and logits scores. REFD achieves 86.83% F1-score as a single system in Audio Deepfake Detection Challenge 2023 Track3, showcasing its state-of-the-art performance.
【30】 Generalized Fake Audio Detection via Deep Stable Learning
标题: 通过深度稳定学习进行广义假音频检测
作者:Zhiyong Wang,Ruibo Fu,Zhengqi Wen,Yuankun Xie,Yukun Liu,Xiaopeng Wang,Xuefei Liu,Yongwei Li,Jianhua Tao,Yi Lu,Xin Qi,Shuchen Shi
备注:accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:虽然目前的虚假音频检测方法在特定数据集上取得了显着的成功,但在使用来自不同分布的数据集进行评估时,它们往往会失败。以前的研究通常通过在训练过程中使用额外的数据或应用额外的损失限制来解决分布偏移。然而,这些方法要么需要大量的数据,要么使训练过程复杂化。在这项工作中,我们提出了一个稳定的基于学习的训练方案,该方案涉及样本权重学习(SWL)模块,通过从训练样本中学习权重来解相关所有选定的特征,从而解决分布偏移问题。所提出的便携式插件式SWL很容易应用于多个基本模型,并在训练过程中不使用额外的数据来推广它们。在ASVspoof数据集上进行的实验清楚地证明了SWL在不同分布的三个评估数据集上推广不同模型的有效性。摘要:Although current fake audio detection approaches have achieved remarkable success on specific datasets, they often fail when evaluated with datasets from different distributions. Previous studies typically address distribution shift by focusing on using extra data or applying extra loss restrictions during training. However, these methods either require a substantial amount of data or complicate the training process. In this work, we propose a stable learning-based training scheme that involves a Sample Weight Learning (SWL) module, addressing distribution shift by decorrelating all selected features via learning weights from training samples. The proposed portable plug-in-like SWL is easy to apply to multiple base models and generalizes them without using extra data during training. Experiments conducted on the ASVspoof datasets clearly demonstrate the effectiveness of SWL in generalizing different models across three evaluation datasets from different distributions.
【31】 A Frame-based Attention Interpretation Method for Relevant Acoustic Feature Extraction in Long Speech Depression Detection
标题: 长言语抑郁检测中相关声学特征提取的基于框架的注意力解释方法
作者:Qingkun Deng,Saturnino Luz,Sofia de la Fuente Garcia
备注:5 pages, 3 figures. arXiv admin note: substantial text overlap with arXiv:2309.13476
链接:点击下载PDF文件
摘要:基于语音的抑郁症检测工具可以帮助早期筛查抑郁症。在这里,我们解决了两个问题,可能会阻碍这种工具的临床实用性:段级标签噪声和缺乏模型的可解释性。我们提出了一个语音级的音频频谱图Transformer,以避免段级标签。我们观察到,该模型显着优于段级模型,提供证据的存在段级标签噪声的音频模态和抑郁症检测的优势,持续时间较长的语音分析。我们引入了一种基于帧的注意力解释方法,从预测相关的波形信号中提取声学特征,供临床医生解释。通过解释,我们观察到,所提出的模型识别降低的响度和F0作为抑郁症的相关信号,这与临床研究中记录的抑郁症患者的语音特征一致。摘要:Speech-based depression detection tools could help early screening of depression. Here, we address two issues that may hinder the clinical practicality of such tools: segment-level labelling noise and a lack of model interpretability. We propose a speech-level Audio Spectrogram Transformer to avoid segment-level labelling. We observe that the proposed model significantly outperforms a segment-level model, providing evidence for the presence of segment-level labelling noise in audio modality and the advantage of longer-duration speech analysis for depression detection. We introduce a frame-based attention interpretation method to extract acoustic features from prediction-relevant waveform signals for interpretation by clinicians. Through interpretation, we observe that the proposed model identifies reduced loudness and F0 as relevant signals of depression, which aligns with the speech characteristics of depressed patients documented in clinical studies.
【32】 StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
标题: StreamSpeech:具有多任务学习的语音同步翻译
作者:Shaolei Zhang,Qingkai Fang,Shoutao Guo,Zhengrui Ma,Min Zhang,Yang Feng
备注:Accepted to ACL 2024 main conference, Project Page: this https URL
链接:点击下载PDF文件
摘要:同步语音到语音翻译(Simul-S2 ST,也称为流式语音翻译)在接收流式语音输入的同时输出目标语音,这对于实时通信至关重要。除了完成语音之间的翻译,Simul-S2 ST还需要一个策略来控制模型在语音输入的适当时刻生成相应的目标语音,从而提出了翻译和策略的双重挑战。在本文中,我们提出了StreamSpeech,这是一个直接的Simul-S2 ST模型,它在多任务学习的统一框架中联合学习翻译和同步策略。StreamSpeech坚持多任务学习方法,可以通过“一体化”无缝模型执行离线和同步语音识别,语音翻译和语音合成。在CVSS基准测试上的实验表明,StreamSpeech在离线S2 ST和Simul-S2 ST任务中都达到了最佳性能。此外,StreamSpeech能够呈现高质量的中间结果(即,ASR或翻译结果),提供更全面的实时交流体验。摘要:Simultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication. Beyond accomplishing translation between speech, Simul-S2ST requires a policy to control the model to generate corresponding target speech at the opportune moment within speech inputs, thereby posing a double challenge of translation and policy. In this paper, we propose StreamSpeech, a direct Simul-S2ST model that jointly learns translation and simultaneous policy in a unified framework of multi-task learning. Adhering to a multi-task learning approach, StreamSpeech can perform offline and simultaneous speech recognition, speech translation and speech synthesis via an "All-in-One" seamless model. Experiments on CVSS benchmark demonstrate that StreamSpeech achieves state-of-the-art performance in both offline S2ST and Simul-S2ST tasks. Besides, StreamSpeech is able to present high-quality intermediate results (i.e., ASR or translation results) during simultaneous translation process, offering a more comprehensive real-time communication experience.
【33】 Dataset-Distillation Generative Model for Speech Emotion Recognition
标题: 语音情感识别的数据集蒸馏生成模型
作者:Fabian Ritter-Gutierrez,Kuan-Po Huang,Jeremy H. M Wong,Dianwen Ng,Hung-yi Lee,Nancy F. Chen,Eng Siong Chng
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:语音深度学习模型依赖于大型数据集,这带来了计算挑战。然而,性能取决于训练数据的大小。数据集蒸馏(DD)旨在学习较小的数据集,而在使用它进行训练时不会降低性能。DD已在计算机视觉中进行了研究,但尚未在语音中进行研究。本文提出了第一个方法,DD语音目标的语音情感识别IEMOCAP。我们使用生成对抗网络(GANs)不是为了模拟真实数据,而是为了验证IEMOCAP的判别信息,这些信息对下游训练很有用。然后,GAN替换原始数据集,并可以对自定义合成数据集大小进行采样。当遵循原始类不平衡时,它会执行UAR,但使用平衡类时,它会将性能提高0.3%的绝对UAR。它还减少了数据集存储,在这两种情况下将下游训练速度加快了95%,并减少了扬声器信息,这可能有助于隐私应用程序。摘要:Deep learning models for speech rely on large datasets, presenting computational challenges. Yet, performance hinges on training data size. Dataset Distillation (DD) aims to learn a smaller dataset without much performance degradation when training with it. DD has been investigated in computer vision but not yet in speech. This paper presents the first approach for DD to speech targeting Speech Emotion Recognition on IEMOCAP. We employ Generative Adversarial Networks (GANs) not to mimic real data but to distil key discriminative information of IEMOCAP that is useful for downstream training. The GAN then replaces the original dataset and can sample custom synthetic dataset sizes. It performs comparably when following the original class imbalance but improves performance by 0.3% absolute UAR with balanced classes. It also reduces dataset storage and accelerates downstream training by 95% in both cases and reduces speaker information which could help for a privacy application.
【34】 AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
标题: AVFF:用于视频Deepfake检测的视听特征融合
作者:Trevine Oorloff,Surya Koppisetti,Nicolò Bonettini,Divyaraj Solanki,Ben Colman,Yaser Yacoob,Ali Shahriyari,Gaurav Bharaj
备注:Accepted to CVPR 2024
链接:点击下载PDF文件
摘要:随着deepfake视频内容的快速增长,我们需要改进和可推广的方法来检测它们。大多数现有的检测方法要么使用单模态线索,要么依赖于监督训练来捕获音频和视觉模态之间的不和谐。虽然前者完全忽略了视听对应关系,但后者主要集中在识别训练语料库中的视听线索,从而可能忽略了有助于检测看不见的深度伪造的对应关系。我们提出了视听特征融合(AVFF),这是一种两阶段的跨模态学习方法,可以显式捕获音频和视觉模态之间的对应关系,以改进深度伪造检测。第一阶段通过对真实视频的自我监督来进行表征学习,以捕获内在的视听对应。为了提取丰富的跨模态表示,我们使用对比学习和自动编码目标,并引入了一种新的视听互补掩蔽和特征融合策略。在第二阶段中调整学习的表示,其中通过对真实和虚假视频的监督学习来进行deepfake分类。大量的实验和分析表明,我们的新的表征学习范式是高度歧视的性质。我们在FakeAVCeleb数据集上报告了98.6%的准确率和99.1%的AUC,分别比当前最先进的视听技术高出14.9%和9.9%。摘要:With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the audio and visual modalities. While the former disregards the audio-visual correspondences entirely, the latter predominantly focuses on discerning audio-visual cues within the training corpus, thereby potentially overlooking correspondences that can help detect unseen deepfakes. We present Audio-Visual Feature Fusion (AVFF), a two-stage cross-modal learning method that explicitly captures the correspondence between the audio and visual modalities for improved deepfake detection. The first stage pursues representation learning via self-supervision on real videos to capture the intrinsic audio-visual correspondences. To extract rich cross-modal representations, we use contrastive learning and autoencoding objectives, and introduce a novel audio-visual complementary masking and feature fusion strategy. The learned representations are tuned in the second stage, where deepfake classification is pursued via supervised learning on both real and fake videos. Extensive experiments and analysis suggest that our novel representation learning paradigm is highly discriminative in nature. We report 98.6% accuracy and 99.1% AUC on the FakeAVCeleb dataset, outperforming the current audio-visual state-of-the-art by 14.9% and 9.9%, respectively.
【35】 Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
标题: 用双解码产品量化变分自动编码器解决大码本语音令牌器的索引崩溃
作者:Haohan Guo,Fenglong Xie,Dongchao Yang,Hui Lu,Xixin Wu,Helen Meng
链接:点击下载PDF文件
摘要:VQ-VAE作为语音标记器的主流方法,一直受到“索引搜索”的困扰,在大的码本中只有少量的码字被激活。本文提出了一种具有更多码本但更少码字的积量化(PQ)VAE来解决这个问题,并构建大码本语音标记器。它将语音特征编码到多个VQ子空间中,并将它们组合成更大码书中的码字。此外,为了更好地利用每个矢量量化子空间,我们还通过编码和量化序列的双重解码训练策略来增强PQ-VAE。实验结果表明,PQ-VAE有效地解决了“索引崩溃”问题,特别是对于较大的码本。该模型的训练策略进一步提高了码本复杂度和重建质量,优于其他多码本矢量量化方法。最后,PQ-VAE证明了其在基于语言模型的TTS中的有效性,支持具有更大码本的更高质量的语音生成。摘要:VQ-VAE, as a mainstream approach of speech tokenizer, has been troubled by index collapse'', where only a small number of codewords are activated in large codebooks. This work proposes product-quantized (PQ) VAE with more codebooks but fewer codewords to address this problem and build large-codebook speech tokenizers. It encodes speech features into multiple VQ subspaces and composes them into codewords in a larger codebook. Besides, to utilize each VQ subspace well, we also enhance PQ-VAE via a dual-decoding training strategy with the encoding and quantized sequences. The experimental results demonstrate that PQ-VAE addresses index collapse" effectively, especially for larger codebooks. The model with the proposed training strategy further improves codebook perplexity and reconstruction quality, outperforming other multi-codebook VQ approaches. Finally, PQ-VAE demonstrates its effectiveness in language-model-based TTS, supporting higher-quality speech generation with larger codebooks.
【36】 Robots Have Been Seen and Not Heard: Effects of Consequential Sounds on Human-Perception of Robots
标题: 机器人被看到却没有被听到:随之而来的声音对人类对机器人感知的影响
作者:Aimee Allen,Tom Drummond,Dana Kulic
备注:16 pages (5 supplementary), 9 figures
链接:点击下载PDF文件
摘要:许多人希望机器人能安静地移动,或者发出令人愉快的“哔哔”声或叮当声,就像他们在机器人视频中看到的那样。不幸的是,这种对安静的期望与现实不符,因为机器人在移动和操作时会发出机器声音,称为“相应的声音”。随着机器人在社会中变得越来越普遍,理解机器人产生的声音以及人们如何感知这些声音对于积极的人机交互(HRI)变得越来越重要。本文研究了人们如何对机器人的相应声音做出反应,特别是机器人如何让参与者感觉到,他们有多喜欢机器人,会被机器人分散注意力,以及一个人与机器人共处的愿望。参与者观看了5个不同机器人的视频,并询问了他们对机器人及其声音的看法。这与完全无声的视频的控制条件进行了比较。本文的结果表明,来自182名参与者(858次试验)的数据表明,机器人产生的声音对人类对机器人的感知有显着的负面影响。首先,参与者的负面“相关影响”增加了,比如让他们在机器人周围感到更不舒服或焦虑。第二,重要声音的存在与参与者感到更分心和更不能集中注意力有关。第三,参与者报告说,他们不太可能想与机器人共享环境。摘要:Many people expect robots to move fairly quietly, or make pleasant "beep boop" sounds or jingles similar to what they have observed in videos of robots. Unfortunately, this expectation of quietness does not match reality, as robots make machine sounds, known as 'consequential sounds', as they move and operate. As robots become more prevalent within society, understanding the sounds produced by robots and how these sounds are perceived by people is becoming increasingly important for positive human robot interactions (HRI). This paper investigates how people respond to the consequential sounds of robots, specifically how robots make a participant feel, how much they like the robot, would be distracted by the robot, and a person's desire to colocate with robots. Participants were shown 5 videos of different robots and asked their opinions on the robots and the sounds they made. This was compared with a control condition of completely silent videos. The results in this paper demonstrate with data from 182 participants (858 trials) that consequential sounds produced by robots have a significant negative effect on human perceptions of robots. Firstly there were increased negative 'associated affects' of the participants, such as making them feel more uncomfortable or agitated around the robot. Secondly, the presence of consequential sounds correlated with participants feeling more distracted and less able to focus. Thirdly participants reported being less likely to want to colocate in a shared environment with robots.
【37】 Text Injection for Neural Contextual Biasing
标题: 用于神经上下文偏置的文本注入
作者:Zhong Meng,Zelin Wu,Rohit Prabhavalkar,Cal Peyser,Weiran Wang,Nanxin Chen,Tara N. Sainath,Bhuvana Ramabhadran
Journal-ref:Interspeech 2024, Kos Island, Greece
链接:点击下载PDF文件
摘要:神经上下文偏置有效地提高了说话者上下文中关键短语的自动语音识别(ASR),特别是那些在训练数据中不常见的短语。这项工作提出了上下文文本注入(CTI),以提高上下文ASR。CTI不仅利用配对的语音文本数据,而且还利用更大的未配对文本语料库来优化ASR模型及其偏置组件。未配对的文本被转换为类似语音的表示,并用于引导模型对相关偏见短语的注意力。此外,我们引入了一个上下文文本注入(CTI)的最小单词错误率(MWER)的训练,它最大限度地减少了预期的WER所造成的上下文偏见时,未配对的文本注入到模型中。实验表明,CTI与1000亿个文本句子可以实现高达43.3%的相对WER减少从一个强大的神经偏置模型。CTI-MWER提供了23.5%的进一步相对改善。摘要:Neural contextual biasing effectively improves automatic speech recognition (ASR) for crucial phrases within a speaker's context, particularly those that are infrequent in the training data. This work proposes contextual text injection (CTI) to enhance contextual ASR. CTI leverages not only the paired speech-text data, but also a much larger corpus of unpaired text to optimize the ASR model and its biasing component. Unpaired text is converted into speech-like representations and used to guide the model's attention towards relevant bias phrases. Moreover, we introduce a contextual text-injected (CTI) minimum word error rate (MWER) training, which minimizes the expected WER caused by contextual biasing when unpaired text is injected into the model. Experiments show that CTI with 100 billion text sentences can achieve up to 43.3% relative WER reduction from a strong neural biasing model. CTI-MWER provides a further relative improvement of 23.5%.
【38】 LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
标题: LiveSpeech:通过音频离散码的自回归建模的低延迟Zero-Shot文本到语音
作者:Trung Dang,David Aponte,Dung Tran,Kazuhito Koishida
链接:点击下载PDF文件
摘要:先前的工作已经通过在经由神经音频编解码器获得的音频令牌上使用生成语言模型来演示了zero-shot文本到语音。然而,使它们适应低延迟场景仍然具有挑战性。在本文中,我们提出了LiveSpeech -一个完全自回归语言模型为基础的方法,zero-shot文本到语音,使低延迟流的输出音频。为了允许在单个解码步骤内进行多个令牌预测,我们提出(1)使用自适应码本损失权重,该自适应码本损失权重考虑每个帧中的码本贡献并专注于硬实例,以及(2)并行地对码本和处理组进行分组。实验表明,我们提出的模型在内容准确性、扬声器相似性、音频质量和推理速度方面达到了最先进的基线,同时适用于低延迟流媒体应用。摘要:Prior works have demonstrated zero-shot text-to-speech by using a generative language model on audio tokens obtained via a neural audio codec. It is still challenging, however, to adapt them to low-latency scenarios. In this paper, we present LiveSpeech - a fully autoregressive language model-based approach for zero-shot text-to-speech, enabling low-latency streaming of the output audio. To allow multiple token prediction within a single decoding step, we propose (1) using adaptive codebook loss weights that consider codebook contribution in each frame and focus on hard instances, and (2) grouping codebooks and processing groups in parallel. Experiments show our proposed models achieve competitive results to state-of-the-art baselines in terms of content accuracy, speaker similarity, audio quality, and inference speed while being suitable for low-latency streaming applications.
【39】 Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
标题: 具有自我监督蒸馏的无文本声学模型用于噪音稳健的表达性语音到语音翻译
作者:Min-Jae Hwang,Ilia Kulikov,Benjamin Peloquin,Hongyu Gong,Peng-Jen Chen,Ann Lee
备注:Accepted to ACL 2024 (findings)
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个无文本的声学模型与自监督蒸馏策略的噪声鲁棒表达语音到语音翻译(S2ST)。最近提出的表达S2ST系统已经取得了令人印象深刻的表现力保存性能级联单元到语音(U2S)生成器的语音到单元的翻译模型。然而,这些系统容易受到输入语音中存在噪声的影响,这是现实世界翻译场景中的一个假设。为了解决这一限制,我们提出了一个U2S生成器,它将无标签蒸馏(DINO)自监督训练策略纳入其预训练过程。由于该方法捕捉了与噪声无关的表达能力,因此即使在噪声环境中也能生成合格的语音。客观和主观评价结果表明,该方法在保持清晰环境下性能的同时,显著提高了表达性S2ST系统在噪声环境下的性能.摘要:In this paper, we propose a textless acoustic model with a self-supervised distillation strategy for noise-robust expressive speech-to-speech translation (S2ST). Recently proposed expressive S2ST systems have achieved impressive expressivity preservation performances by cascading unit-to-speech (U2S) generator to the speech-to-unit translation model. However, these systems are vulnerable to the presence of noise in input speech, which is an assumption in real-world translation scenarios. To address this limitation, we propose a U2S generator that incorporates a distillation with no label (DINO) self-supervised training strategy into it's pretraining process. Because the proposed method captures noise-agnostic expressivity representation, it can generate qualified speech even in noisy environment. Objective and subjective evaluation results verified that the proposed method significantly improved the performance of the expressive S2ST system in noisy environments while maintaining competitive performance in clean environments.
【40】 Operational Latent Spaces
标题: 运营潜在空间
作者:Scott H. Hawley,Austin R. Tackett
备注:7 pages, 6 figures. Accepted to AES International Symposium on AI and the Musician
链接:点击下载PDF文件
摘要:我们研究了通过自监督学习来支持语义上有意义的操作的潜在空间的构建。类似于运算放大器,这些“操作潜在空间”(OpLaS)不仅展示了语义结构,如聚类,但也支持常见的转换操作与固有的语义意义。一些操作的潜在空间被发现出现“无意”的进展,朝着一些(其他)自我监督的学习目标,其中无意的,但仍然有用的属性被发现之间的关系的点在空间中。其他空间可能是由开发人员“有意”构建的,他们规定了某些类型的集群或转换,旨在产生所需的结构。我们专注于通过自我监督学习有意创建操作潜在空间,包括通过一个新的“FiLMR”层引入旋转算子,该层可用于实现在某些音乐结构中发现的环状对称性。摘要:We investigate the construction of latent spaces through self-supervised learning to support semantically meaningful operations. Analogous to operational amplifiers, these "operational latent spaces" (OpLaS) not only demonstrate semantic structure such as clustering but also support common transformational operations with inherent semantic meaning. Some operational latent spaces are found to have arisen "unintentionally" in the progress toward some (other) self-supervised learning objective, in which unintended but still useful properties are discovered among the relationships of points in the space. Other spaces may be constructed "intentionally" by developers stipulating certain kinds of clustering or transformations intended to produce the desired structure. We focus on the intentional creation of operational latent spaces via self-supervised learning, including the introduction of rotation operators via a novel "FiLMR" layer, which can be used to enable ring-like symmetries found in some musical constructions.
【41】 Sequence-to-sequence models in peer-to-peer learning: A practical application
标题: 点对点学习中的序列到序列模型:实际应用
作者:Robert Šajina,Ivo Ipšić
链接:点击下载PDF文件
摘要:本文探讨了基于LSTM单元的序列到序列(Seq2Seq)模型在对等学习环境中用于自动语音识别(ASR)任务的适用性。利用两种不同的对等学习方法,该研究模拟了智能体的学习过程,并使用两种不同的ASR数据集来评估它们在ASR任务中的表现。在集中式训练环境中,利用Deep Speech 2模型的缩小变体,单个模型在UserLibri数据集上训练时的单词错误率(WER)为84%,在LJ Speech数据集上训练时为38%。相反,在涉及55个代理的对等学习场景中,UserLibri数据集的WER范围为87 %至92 %,LJ Speech数据集的WER范围为52 %至56 %。研究结果证明了在分散的环境中使用Seq2Seq模型的可行性,尽管与集中式训练方法相比,单词错误率(WER)略高。摘要:This paper explores the applicability of sequence-to-sequence (Seq2Seq) models based on LSTM units for Automatic Speech Recognition (ASR) task within peer-to-peer learning environments. Leveraging two distinct peer-to-peer learning methods, the study simulates the learning process of agents and evaluates their performance in ASR task using two different ASR datasets. In a centralized training setting, utilizing a scaled-down variant of the Deep Speech 2 model, a single model achieved a Word Error Rate (WER) of 84 % when trained on the UserLibri dataset, and 38 % when trained on the LJ Speech dataset. Conversely, in a peer-to-peer learning scenario involving 55 agents, the WER ranged from 87 % to 92 % for the UserLibri dataset, and from 52 % to 56 % for the LJ Speech dataset. The findings demonstrate the feasibility of employing Seq2Seq models in decentralized settings, albeit with slightly higher Word Error Rates (WER) compared to centralized training methods.
机器翻译,仅供参考
