微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 3 篇
2. 语音合成与声音生成 4 篇
3. 语音增强、降噪与音频修复 1 篇
4. 音频事件检测与场景理解 1 篇
5. 多模态音频与视听学习 1 篇
6. 数据集、基准与评测 2 篇
7. 安全、隐私与深度伪造音频 1 篇
1. 语音识别与关键词检测 | 3 篇
1. Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech
利用语码混合引导的合成语音改进语码转换语音识别
AI 总结:针对语码转换语音识别中高质量文本-语音对稀缺的问题,提出语码混合引导的偏好学习框架,通过语码混合指数优化合成语音的转换保真度,在SEAME语料库上微调Whisper Large,将混合错误率从12.1%/17.8%降至8.9%/14.2%。
链接:https://arxiv.org/abs/2606.19381
机构:College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院); Google DeepMind(谷歌深度思维)
作者:Yue Heng Yeo, Haoyang Li, Yizhou Peng, Shreyas Gopal, Hexin Liu, Leibny Paola Garcia-Perera, Hardik B. Sailor, Jeremy H. M. Wong, Eng Siong Chng
英文摘要:Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augmentation via Text-to-speech (TTS) has been explored, existing CS TTS approaches primarily optimise reconstruction fidelity and do not explicitly enforce language-boundary consistency, thereby limiting their effectiveness for CS ASR augmentation. This paper proposes a code-mixing guided preference-learning framework that steers synthetic speech generation toward improved code-switching fidelity using the Code Mixing Index (CMI). Experiments on the SEAME Mandarin-English conversational corpus demonstrate that the proposed method enhances the utility of synthetic data for ASR fine-tuning. Specifically, when fine-tuning Whisper Large, the proposed approach reduces Mixed Error Rate (MER) from 12.1%/17.8% to 8.9%/14.2% on the DevMAN and DevSGE sets, respectively.
2. S-JEPA: Soft Clustering Anchors for Self-Supervised Speech Representation Learning
S-JEPA:用于自监督语音表示学习的软聚类锚点
AI 总结:提出S-JEPA,通过KL散度匹配高斯混合模型的软后验概率训练编码器-预测器对,无需离线重聚类或教师蒸馏,在SUPERB协议下以低于90M参数取得最低WER,并建立新的帕累托前沿。
链接:https://arxiv.org/abs/2606.19398
机构:Carnegie Mellon University(卡内基梅隆大学); New York University(纽约大学); James Silberrad Brown Center for AI(詹姆斯·西尔伯拉德·布朗人工智能中心); Columbia University(哥伦比亚大学); Northeastern University(东北大学); Stanford University(斯坦福大学); Amazon GenAI(亚马逊生成式人工智能)
作者:Georgios Ioannides, Adrian Kieback, Judah Goldfeder, Linsey Pang, Aman Chadha, Aaron Elkins, Yann LeCun, Ravid Shwartz-Ziv
英文摘要:Self-supervised speech encoders are predominantly trained by predicting discrete hard cluster IDs at masked positions, a recipe that collapses acoustic ambiguity at category boundaries and requires interrupting training to re-cluster the entire corpus between iterations. We introduce S-JEPA, a JEPA-style encoder-predictor pair trained to match the soft posteriors of a Gaussian Mixture Model at masked positions via KL divergence. Training runs as one continuous optimization trajectory in two phases: a fixed GMM over MFCC features, then an online GMM over encoder features, with the input layer selected adaptively from a label-free signal, removing both the offline re-cluster step and the hand-tuned choice of which transformer layer to cluster on. Under the SUPERB protocol, S-JEPA achieves the lowest WER among evaluated SSL methods below 90M parameters and matches HuBERT-Base on emotion recognition at roughly half its parameter count, establishing a new Pareto frontier without offline re-clustering or teacher distillation. An analysis of the predictor's per-frame entropy on held-out speech reveals a bimodal distribution with a substantial minority of frames near the entropy of a perfect two-cluster tie, providing direct empirical evidence that the soft-target objective preserves the acoustic ambiguity that hard targets would collapse. Code is available at this https URL.
3. Segment-Level Mandarin Chinese Speech-Based Cognitive Impairment Detection via an Autoencoder with Contrastive Learning
基于自编码器与对比学习的段级普通话语音认知障碍检测
AI 总结:提出段级表示学习框架,结合自编码器和对比学习,在四个普通话数据集上实现稳定的二分类和三分类认知障碍检测,尤其改善了临床困难的三分类性能。
链接:https://arxiv.org/abs/2606.19996
机构:School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院); Key Laboratory of System Control and Information Processing, Ministry of Education of China(教育部系统控制与信息处理重点实验室); Shanghai Key Laboratory of Perception and Control in Industrial Network Systems(上海市工业网络系统感知与控制重点实验室); Department of Computer Science and Engineering, University of Bologna(博洛尼亚大学计算机科学与工程系); Department of Mathematical, Physical and Computer Sciences, University of Parma(帕尔马大学数学、物理与计算机科学系)
作者:Yongqi Shao, Hong Huo, Flavio Bertini, Danilo Montesi, Tao Fang
英文摘要:\noindent\textbf{Background and Objective:} Speech has emerged as a low-cost and non-invasive digital biomarker with considerable potential for cognitive impairment detection. However, limited labeled data and cross-dataset variability remain major challenges for robust speech-based screening systems. \par\noindent\textbf{Methods:} We developed a segment-level representation learning framework for speech-based cognitive impairment detection. Speech recordings were divided into short segments and converted into spectrogram representations. To improve robustness under limited-data conditions, offline and online augmentation strategies were combined with autoencoder-based representation learning and contrastive objectives to enhance discriminative latent representations. \par\noindent\textbf{Results:} Experiments conducted on four independent Mandarin Chinese speech datasets demonstrated stable and competitive performance in both binary and three-class classification tasks, with particularly notable improvements in the clinically challenging three-class setting. Ablation studies further supported the effectiveness of the proposed framework. \par\noindent\textbf{Conclusions:} The findings suggest that segment-level speech representation learning may provide a scalable and practical approach for cognitive impairment screening in resource-constrained clinical settings.
2. 语音合成与声音生成 | 4 篇
4. RIVET: Robust Idempotent Voice Attribute Editing
RIVET: 鲁棒的幂等语音属性编辑
AI 总结:提出RIVET训练框架,通过幂等性正则化提升语音属性编辑模型对标签噪声的鲁棒性,在合成噪声和真实噪声数据集上均优于标准训练。
链接:https://arxiv.org/abs/2606.19629
机构:Carnegie Mellon University(卡内基梅隆大学)
作者:Dareen Alharthi, Bhuvan Koduru, Rita Singh, Bhiksha Raj
英文摘要:Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datasets, however, attribute annotations are often noisy or inconsistent, which can cause conditional generative models to produce unstable edits. In this work, we show that idempotency provides an effective mechanism for improving robustness to noisy labels. An idempotent operator is one for which repeated application does not change the result, i.e., f(f(x)) = f(x). Enforcing this property acts as an implicit regularizer that reduces sensitivity to mislabeled examples. We introduce RIVET, a training framework that incorporates an idempotency objective to improve robustness to label noise. We evaluate RIVET under controlled label noise and on the GLOBE dataset with naturally noisy annotations. RIVET improves editing success and better preserves speaker identity than standard training, showing that idempotency improves robustness in voice editing models.
5. Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis
探索预训练在语音合成中通过微调对音素添加的益处
AI 总结:研究预训练模型在微调过程中添加新音素时的表现,发现预训练主要提升自然度,但对新音素添加的益处有限。
链接:https://arxiv.org/abs/2606.19792
机构:CyberAgent, Japan(日本CyberAgent公司); Nagoya University, Japan(日本名古屋大学)
作者:Masato Murata, Koichi Miyazaki, Tomoki Koriyama, Tomoki Toda
英文摘要:Transfer learning is widely used for low-resource text-to-speech. When the target corpus contains phonemes unseen in pre-training, the model must expand its phoneme inventory during fine-tuning; we call the process "phoneme addition." However, it remains unclear whether the pre-trained ability to generate seen phonemes contributes to this process. This study investigates phoneme addition in two settings: (1) a simulation setup using LLM-generated phoneme-controlled corpora that enables investigation without considering confounding factors, and (2) a real-speech cross-lingual transfer setup (English to Japanese) to validate whether the findings hold in practice. Experiments in both settings showed that while fine-tuning achieved higher naturalness than training from scratch, it required as much or more data to achieve comparable PER for new phonemes. These results indicate that pre-training mainly contributes to naturalness improvement, but offers limited benefit for phoneme addition.
6. Hybrid Diffusion Transformer for Instruction-Guided Audio Editing via Rectified Flow
基于整流流的混合扩散变压器用于指令引导音频编辑
AI 总结:提出混合两阶段扩散变压器架构,通过粗到细策略平衡全局语义对齐与局部细节编辑,在重叠音频事件和复杂指令任务上提升性能与效率。
链接:https://arxiv.org/abs/2606.20101
机构:Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(萨里大学视觉、语音与信号处理中心); School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院); Fisheries College, Ocean University of China(中国海洋大学水产学院); College of Information and Electrical Engineering, China Agricultural University(中国农业大学信息与电气工程学院)
作者:Liting Gao, Yonggang Zhu, Yaru Chen, Dongyu Wang, Shubin Zhang, Zhenbo Li, Jean-Yves Guillemaut, Wenwu Wang
英文摘要:Audio editing aims to modify specific content in an existing audio clip according to a natural language instruction while preserving the remaining acoustic content. Despite the remarkable progress of diffusion models, existing training-based editing methods mainly rely on the local inductive biases and cross-attention interaction in convolutional U-Net backbones, which often hinder long-range semantic alignment and precise understanding and localization of instructions. In contrast, diffusion transformers provide stronger global modeling and multimodal fusion, but existing editing architectures usually adopt a simple stack of MMDiT and DiT blocks. Applying joint attention over concatenated audio and text tokens in all blocks results in quadratic complexity with respect to token length. To balance editing performance and efficiency, we propose a hybrid two-stage diffusion transformer architecture for instruction-guided audio editing based on rectified flow matching. It performs joint attention over audio and text tokens to establish coarse semantic alignment at low-resolution stage, then switches to alternating joint-attention and cross-attention blocks to refine editing details at high-resolution stage. This coarse-to-fine strategy enables efficient and accurate instruction-guided audio editing. Experiments show that the proposed framework achieves notable performance gains on challenging editing tasks involving overlapping audio events and complex instructions, while substantially improving editing efficiency with a compact model.
7. Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization
Zero-VC: 通过说话人匿名化实现零前瞻流式语音转换
AI 总结:针对流式零样本语音转换中音色与语言内容解耦的挑战,提出将说话人匿名化作为扰动机制,在保留韵律效用的同时显式减轻音色泄露,实现严格因果的零前瞻网络。
链接:https://arxiv.org/abs/2606.20218
机构:The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)); Shenzhen Loop Area Institute(深圳环域研究所); Shenzhen Transsion Holdings Co., Ltd.(深圳传音控股股份有限公司)
作者:Yudong Li, Zihao Fang, Junwen Qiu, Ruihai Jing, Ruixiang Hang, Yingda Shen, Zhizheng Wu
英文摘要:Streaming zero-shot voice conversion struggles to disentangle timbre from linguistic content without degrading utility or inflating latency. Current methods rely on information bottleneck (IB) or speaker perturbation. While IB filters out timbre, it discards prosody, forcing models to explicitly inject features like fundamental frequency. This often requires buffering future frames, creating algorithmic lookahead latency. On the other hand, existing perturbation methods largely overlook the crucial trade-off between timbre leakage and utility preservation. Recognizing this neglected trade-off, we find that the inherent objective of Speaker Anonymization (SA) aligns well with balancing these factors. Thus, we introduce SA as a novel perturbation mechanism to explicitly mitigate timbre leakage while retaining prosodic utility. Crucially, SA's robust representations significantly alleviate the generator's reliance on future context, enabling our strictly causal, zero-lookahead network. Audio samples are available at this https URL.
3. 语音增强、降噪与音频修复 | 1 篇
8. Latency-Configurable Streaming Speech Enhancement via Asymmetric Temporal Padding
通过非对称时间填充实现延迟可配置的流式语音增强
AI 总结:提出LaCo-SENet,通过非对称时间填充和双缓冲流式机制,在单一超参数下实现延迟与质量的灵活权衡,在VoiceBank+DEMAND上以1.37M参数获得12.5-75.0ms延迟范围,PESQ从3.35到3.43。
链接:https://arxiv.org/abs/2606.19688
机构:Department of Electrical Engineering, Pohang University of Science and Technology (POSTECH)(电气工程系,浦项科技大学); Intus Co. Ltd.(Intus有限公司)
作者:Yunsik Kim, Yoonyoung Chung
英文摘要:Streaming speech enhancement requires balancing algorithmic latency against quality, yet existing approaches largely treat this as a binary causal versus non-causal choice. LaCo-SENet addresses this issue with two mechanisms parameterized by a single training-time hyperparameter. First, asymmetric temporal padding redistributes past and future context in convolutions, enabling systematic latency configuration. Second, dual-buffer streaming combines state buffers for past context with lookahead buffers that supply future context at both the input and feature levels. Selective state updates also prevent future-frame leakage into the streaming state, ensuring training-inference consistency. On VoiceBank+DEMAND, a fixed-budget (1.37M parameters) backbone yields a family of models spanning 12.5-75.0 ms, with PESQ rising from 3.35 to 3.43. At just 12.5 ms (fully causal), a PESQ of 3.35 matches or exceeds the prior causal state-of-the-art (3.27 at 46.5 ms).
4. 音频事件检测与场景理解 | 1 篇
9. Exploring Feature Extraction Technique Parameters for Acoustic Gunshot Classification
声学枪声分类的特征提取技术参数探索
AI 总结:本文系统研究了特征提取技术及其参数对声学枪声分类的影响,使用ResNet-18在23000条枪声数据集上评估,发现正确技术可提升top-1准确率20%,参数优化可再提升4.7%。
链接:https://arxiv.org/abs/2606.19568
机构:Certus Innovations
作者:Sinclair Gurny, Ryan Quinn
英文摘要:Acoustic gunshot detection is a problem with applications across civilian public safety, military operations, and wildlife conservation, yet the field lacks a rigorous exploration of feature extraction techniques with a focus on generalization to realistic data. The mixed effectiveness of commercial gunshot detection and classification systems indicates an open problem that is not adequately addressed by the current literature. In this paper, we present a systematic investigation of common feature extraction techniques using a dataset of 23,000 gunshot recordings across 85 firearms and 21 calibers. We benchmark three feature extraction techniques with 12 total unique parameter sets using ResNet-18. Our results demonstrate that using the correct feature extraction technique can improve top-1 accuracy by up to 20%, and utilizing the correct parameters for a given feature extraction technique can improve that value by up to 4.7%.
5. 多模态音频与视听学习 | 1 篇
10. MixProLAP: Mixture-Induced Uncertainty Modeling for Probabilistic Language-Audio Pretraining
MixProLAP:混合诱导的不确定性建模用于概率性语言-音频预训练
AI 总结:提出概率性音频-语言预训练框架MixProLAP,通过混合音频-文本对模拟重叠声音,建模多对多对应不确定性,并引入多级包含损失,在音频-文本检索中优于确定性基线。
链接:https://arxiv.org/abs/2606.20418
机构:LINE WORKS Corporation(LINE WORKS公司); NAVER Cloud Corporation(NAVER Cloud公司)
作者:Yu Nakagome, Jaesong Lee, Soo-Whan Chung
英文摘要:Acoustic environments often contain multiple overlapping sound events, and the same acoustic scene can be described using diverse textual expressions, making audio-text alignment inherently ambiguous. This paper proposes a probabilistic audio-language pretraining framework to model many-to-many correspondence ambiguity in audio-text alignment. Unlike conventional contrastive methods that learn deterministic point embeddings, our approach represents each modality as a distribution and learns uncertainty-aware cross-modal alignment. Rather than relying on masking-based uncertainty simulation, we mix audio-text pairs to create overlapping sounds that better reflect real acoustic mixtures and capture semantic inclusion relations among sound events. We further introduce a multi-level inclusion loss to enforce representations consistent with these relations. Experiments on audio-text retrieval benchmarks show that the proposed method outperforms deterministic baselines.
6. 数据集、基准与评测 | 2 篇
11. PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets
PrefSQA: 用于语音质量评估的成对偏好预测及高质量数据集的关键作用
AI 总结:提出PrefSQA模型,通过不确定性感知logits、损伤注意力头和非匹配参考比较模块,利用高质量偏好数据集提升语音质量评估的准确性。
链接:https://arxiv.org/abs/2606.19597
机构:Department of Computer Science and Engineering, The Ohio State University, USA(美国俄亥俄州立大学计算机科学与工程系)
作者:Junyi Fan, Donald S. Williamson
英文摘要:Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences. This introduces labeling noise, which limits the reliability of MOS prediction. Preference prediction reduces this variability as listeners compare signals directly, producing cleaner labels. We study MOS-free preference prediction and propose PrefSQA, which incorporates uncertainty-aware logits, an impairment attention head, and a module based on non-matching-reference comparisons. We use and refine five datasets, including MOS-derived and low-noise simulated sets with matching and non-matching content, experiment with human preference sets, and test on unseen data. Experiments show small improvements on MOS-derived data, while other sets reveal clear improvement over the baselines, highlighting the value of high-quality preference data and demonstrating the effectiveness of the proposed method.
12. PolSeT: Polish Semantics of Timbre Dataset
PolSeT: 波兰语音色语义数据集
AI 总结:介绍PolSeT数据集,通过自由言语化和语义差异实验,收集波兰语语义描述符和音色评分,填补音色研究数据空白,支持跨文化心理声学和MIR研究。
链接:https://arxiv.org/abs/2606.19987
机构:AGH University of Krakow(AGH克拉科夫大学)
作者:Jan Jasiński
英文摘要:This data report introduces PolSeT (Polish Semantic Timbre), a dataset designed to facilitate research in psychoacoustics and Music Information Retrieval (MIR) in Polish and cross-cultural contexts. The dataset contains data from two sequential experiments. Experiment 1 (N=60) was a free-verbalization task aimed at creating a lexicon of Polish semantic descriptors. Using 11 stimuli, a total of 1901 descriptors (701 unique) were gathered. Experiment 2 (N=105) utilized this lexicon to conduct a semantic differential study, where participants rated 18 instrument sounds on 8 bipolar scales, with repeated trials for reliability analysis. The released dataset includes raw listener responses, comprehensive demographics (experience, gender, age), audio stimuli, and extracted acoustic features with Python extraction code. This dataset addresses a gap in open timbre research data, providing both the qualitative linguistic groundwork and the quantitative ratings necessary for psychoacoustic research and the training of multilingual semantic embedding models.
7. 安全、隐私与深度伪造音频 | 1 篇
13. FlowFake: Liquid Networks for Audio Deepfake Detection
FlowFake: 用于音频深度伪造检测的液态网络
AI 总结:针对音频深度伪造检测中跨数据集泛化失败的问题,提出基于液态时间常数(LTC)架构的FlowFake模型,通过学习ODE演化隐藏状态并自适应时间常数,以34K参数在跨域基准上超越现有方法。
链接:https://arxiv.org/abs/2606.19579
机构:Delhi Technological University(德里理工大学)
作者:Shivaay Dhondiyal, Divyansh Sharma, Dinesh Kumar Vishwakarma
英文摘要:Audio deepfakes generated by neural text-to-speech and voice-cloning systems threaten speaker verification and public discourse at scale. The core challenge is cross-dataset generalization: detectors trained on one synthesis pipeline collapse on unseen forgeries. We argue that this failure is primarily because of structural synthetic speech artifacts which are multi-timescale trajectory anomalies. Though every existing detector aggregates a fixed-window frame statistics, this misaligns the architecture with the signal. We propose FlowFake, a Liquid Time-Constant (LTC) architecture whose hidden state evolves via a learned ODE, with per-neuron adaptive time constants simultaneously resolving spectral (10ms) and prosodic (2s) cues. At only 34K parameters FlowFake achieves formal BIBO stability and O(dt^4) integration error. On a four-dataset cross domain benchmark (ASVspoof2019-LA, FakeOrReal, InTheWild, MLAAD), FlowFake reaches 75.29% on ASVspoof2019 trained only on FakeOrReal and 79.97% trained only on MLAAD. It outperforms RawGAT-ST and Whisper-DF on every evaluated pair and matching SSL Wav2vec2 (300x larger) at 0.01% of its parameter count. The source code is available on: this https URL
