微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音合成与声音生成 1 篇
2. 语音增强、降噪与音频修复 1 篇
3. 音乐信息检索与音乐生成 2 篇
4. 数据集、基准与评测 3 篇
5. 安全、隐私与深度伪造音频 1 篇
6. 其他/综合语音音频 3 篇
1. 语音合成与声音生成 | 1 篇
1. Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis
Sarashina2.2-TTS:通过数据扩展和针对性数据合成解决日语语音生成中的汉字多音问题
AI 总结:针对日语汉字多音问题,提出Sarashina2.2-TTS系统,通过扩展训练数据至36.1万小时并设计覆盖全部常用汉字的针对性数据增强管道,显著提升读音准确率,在零样本日语语音合成中达到最高说话人相似度。
链接:https://arxiv.org/abs/2606.25369
机构:SB Intuitions
作者:Lianbo Liu, Shiao Zhu, Kai Washizaki, Reo Yoneyama, Haesung Jeon, Mengjie Zhao, Yusuke Fujita, Hao Shi, Nao Yoshida, Yuan Gao, Roman Koshkin, Yukiya Hono, Yui Sudo
英文摘要:While large language model (LLM)-based text-to-speech (TTS) systems have achieved high-quality speech synthesis, most existing systems focus on English and Chinese. Japanese, however, remains under-explored, and its unique linguistic challenges, such as widespread context-dependent kanji polyphony, have yet to be adequately tackled. Here we introduce Sarashina2.2-TTS ( this https URL ), a Japanese-centric LLM-TTS system that tackles these challenges through a dual approach: data strategy and evaluation methodology. First, we scale training to approximately 361k hours of speech, incorporating a balanced mix of Japanese and English data. Furthermore, we design a targeted data augmentation pipeline covering all 2,136 Joyo (regular-use) kanji designated by Japan's Agency for Cultural Affairs to efficiently address kanji polyphony disambiguation. Second, we introduce the Joyo Kanji Yomi Benchmark ( this https URL ), covering all 2,136 Joyo kanji and their 4,378 readings. Alongside this benchmark, we propose Kana-CER, a metric that compares synthesized speech against reference readings in the kana space, eliminating orthographic variations to directly measure pronunciation correctness. Experiments demonstrate that our targeted data augmentation significantly improves reading accuracy. Overall, Sarashina2.2-TTS achieves state-of-the-art kanji-level reading accuracy and matches top baselines on general sentence-level pronunciation, while delivering the highest speaker similarity in zero-shot Japanese speech synthesis. Furthermore, cross-lingual evaluation reveals that Sarashina2.2-TTS is the only system that maintains stable Japanese pronunciation regardless of the prompt language, confirming that our balanced training approach improves cross-lingual robustness.
2. 语音增强、降噪与音频修复 | 1 篇
2. One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications
一个模型,多种延迟:面向多样实时应用的通用语音增强
AI 总结:提出一种统一模型,通过可配置前瞻帧控制算法延迟,利用并行卷积层和早停机制分别管理算法与计算延迟,并采用两阶段训练策略缩小性能差距,实现单一模型适应多种延迟需求。
链接:https://arxiv.org/abs/2606.25621
作者:Szu-Wei Fu, Rong Chao, Xuesong Yang, Sung-Feng Huang, Ante Jukić, Yu Tsao, Yu-Chiang Frank Wang
英文摘要:Different real-time speech applications impose distinct latency budgets, often requiring separately trained enhancement models for each scenario. In this paper, we propose a one-for-all, real-time universal speech enhancement model that provides explicit control over both algorithmic and computational latency. Algorithmic latency is flexibly adjusted via configurable look-ahead frames. To avoid learning inefficiency caused by varying padding configurations, we introduce parallel convolutional layers corresponding to different look-ahead settings. Computational latency is controlled through an early-exit mechanism, enabling inference at different network depths. To narrow the performance gap between specialized and flexible models, we propose a two-stage training strategy with a shared-to-multiple decoder transition. Overall, the proposed framework enables a single model to be deployed across diverse latency budgets without retraining separate models.
3. 音乐信息检索与音乐生成 | 2 篇
3. Velocity Prediction in Automatic Guitar Transcription
自动吉他转录中的速度预测
AI 总结:提出一种结合虚拟乐器合成数据预训练和真实音频微调的方法,实现自动吉他转录中的音符速度预测,性能优于基线模型。
链接:https://arxiv.org/abs/2606.24912
机构:UKRI Centre for Doctoral Training in Artificial Intelligence and Music(UKRI人工智能与音乐博士培训中心); Innovate UK(英国创新署)
作者:Jackson Loth, Xavier Riley, Simon Dixon, Emmanouil Benetos
英文摘要:Automatic Music Transcription (AMT) models have achieved a high level of success in polyphonic transcription of various instruments. Velocity, typically a measure of note intensity, is less commonly predicted in these models due to the absence of velocity labels in available datasets and lack of a proper definition for instruments other than piano. We present a methodology and model for velocity prediction in Automatic Guitar Transcription (AGT) which uses virtual instruments to generate synthetic training data with velocity labels. We first pretrain a model on this synthetic data. These weights are then transferred to a different model and trained on real guitar audio, allowing the model to retain the working velocity prediction while also achieving high performance and generalisability from the real training data. The velocity prediction is shown to outperform a baseline model which does not use the pretrained velocity weights, when evaluated on synthetic data. In addition, using the pretrained velocity weights offers a small improvement in note transcription, though the magnitude of this improvement is limited and not always significant depending on the testing data. Overall the model achieves results comparable to the state of the art in guitar transcription, while also successfully predicting velocity.
4. Frequency-Aware Self-Supervised Music Representation Learning
频率感知的自监督音乐表示学习
AI 总结:提出PupuJEPA,一种直接在2D频谱图上训练的视觉联合嵌入预测架构,通过预测掩码补丁的潜在嵌入学习鲁棒表示,在MARBLE基准上线性探测优于1D序列SSL模型。
链接:https://arxiv.org/abs/2606.25713
机构:Spellbrush; Acoustic Lab, Department of Information and Communications Engineering (DICE), Aalto University(阿尔托大学信息与通信工程系声学实验室); School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院)
作者:Yicheng Gu, Junan Zhang, Jerry Li, Zhizheng Wu, Lauri Juvela
英文摘要:Self-supervised learning (SSL) has emerged as an essential paradigm for music information retrieval (MIR). While current SSL models achieve state-of-the-art performance across various MIR tasks, they typically treat audio as 1D sequences, either operating on time-domain waveforms or on flattened time-frequency-domain spectrograms. This discards the rich spatial and structural information in time-frequency representations and overlooks a fundamental intuition in music production. In particular, music is naturally represented as time-frequency grids in MIDI-based workflows, a structure that tightly corresponds to 2D spectrograms and inherently makes many MIR tasks trivial. Motivated by this intuition, we propose PupuJEPA, a visual Joint-Embedding Predictive Architecture (JEPA) that is trained directly on 2D spectrograms. Instead of applying masked language modeling (MLM) to 1D sequences, PupuJEPA learns robust representations by predicting the latent embeddings of masked 2D spectrogram patches from unmasked contexts. To optimally adapt such a visual framework to music signals, we also apply domain-specific modifications to model architecture, training scheme, and inference paradigm, with comprehensive ablation studies showing their effectiveness. Evaluations on the MARBLE benchmark show that PupuJEPA outperforms the 1D sequence-based SSL models across multiple MIR tasks in linear probing. Additionally, case studies of the attention maps also confirm that PupuJEPA captures musically meaningful patterns within the 2D time-frequency domain. Codes and checkpoints are available at: this https URL.
4. 数据集、基准与评测 | 3 篇
5. From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models
从声音到场景:评估大型音频语言模型中上下文感知听觉场景理解的基准
AI 总结:提出CASU基准,通过构建半合成音频流和设计四项任务,评估大型音频语言模型整合语音、事件和背景环境进行上下文感知听觉场景理解的能力。
链接:https://arxiv.org/abs/2606.25391
机构:University of California Irvine(加州大学尔湾分校); University of Illinois Chicago(伊利诺伊大学芝加哥分校); Kennesaw State University(肯尼索州立大学)
作者:Pengfei Zhang, Hoang H Nguyen, Kazi Shaharair Sharif, Yutong Song, Wenjun Huang, Henry Peng Zou, Pinxin Liu, Honghui Xu, Amir M. Rahmani
英文摘要:Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. However, existing benchmarks predominantly evaluate these layers in isolation, overlooking the complex contextual relationships that arise when multiple acoustic sources co-occur in real-world auditory scenes. Real-world auditory interpretation requires Context-Aware Auditory Scene Understanding (CASU): the ability to comprehend the holistic scene by integrating sound layers. To evaluate this capability, we introduce the CASU benchmark, which assesses whether Audio LLMs can interpret auditory scenes composed of speech, acoustic events (e.g., announcements), and background environments (e.g., traffic), and reason about the logical relationships between these layers. We propose a scalable pipeline for constructing time-accurate, semi-synthetic audio streams by composing real-world scene sounds with synthetic speech. Building on this data, we design four tasks that probe scene understanding: contextual question answering, entity extraction from the scene, speaker role inference, and counterfactual reasoning where scene is manipulated. Experiments across multiple LALMs demonstrate that effective auditory scene understanding requires integration over all auditory layers, rather than reliance on speech or sound alone, underscoring the necessity of CASU for advancing complex audio understanding in LALMs.
6. STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity
STEB:用于评估超越翻译保真度的语音到语音翻译表现力基准
AI 总结:提出STEB基准,通过描述-总结框架评估语音到语音翻译在情感、场景风格和非语言发声方面的表现力,发现现有系统在语义保真度上表现良好但表现力保留不足。
链接:https://arxiv.org/abs/2606.25529
机构:Hong Kong University of Science and Technology(香港科技大学); Tencent Youtu Lab(腾讯优图实验室); Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
作者:Sitong Cheng, Weizhen Bian, Songjun Cao, Jin Li, Bei Liu, Chunyang Jiang, Yike Zhang, Weihao Wu, Yiming Li, Chi-Min Chan, Long Ma, Wei Xue
英文摘要:Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue), and nonverbal vocalizations (NVs). Moreover, collecting cross-lingual target speech that is both translation-faithful and expressively aligned with the source is difficult at scale, making reference-based evaluation impractical. We introduce STEB (Speech-to-Speech Translation Expressiveness Benchmark), a 32.6-hour Chinese--English benchmark that evaluates both standard dimensions (translation fidelity, speaker similarity, duration alignment) and expressiveness dimensions (emotion, scenario style, NV preservation). For expressiveness evaluation, STEB uses a caption-then-summarize framework that converts speech into structured expressive attributes and compares source and hypothesis attributes with an LLM judge. Human validation shows statistically significant correlations with listener judgments across all expressive dimensions. We evaluate six S2ST systems covering cascaded systems, end-to-end models, and speech large language models. Many systems, especially cascaded ones, achieve strong translation fidelity, but they still struggle with emotion preservation (best: 3.82/5) and NV preservation (best: 2.31/5). These results reveal a gap between semantic transfer and expressive transfer, identifying expressiveness preservation as an open challenge for S2ST. Audio samples are available at this https URL.
7. FoleySet: A Multi-Level Human-Annotated Foley Sound Dataset
FoleySet: 一个多层级人工标注的拟音音效数据集
AI 总结:为解决拟音数据稀缺问题,构建了包含10,000个音频片段、具有两级拟音分类体系的数据集FoleySet,用于分类、检索和生成任务。
链接:https://arxiv.org/abs/2606.25980
作者:Sunshiyu Wang, Alexander Lerch
英文摘要:In audiovisual post-production, Foley refers to synchronous sound effects associated with human actions, such as footsteps, cloth rustle, and prop handling, that are recreated to match the on-screen movements and interactions of characters. These sounds are often recorded by professional Foley artists using physical props. This resource-intensive workflow has motivated data-driven research on Foley, including tasks such as classification, retrieval, and generation; however, high-quality annotated Foley datasets for training remain scarce. To address this gap, we present FoleySet, a publicly available Foley dataset of 10,000 audio clips annotated with a two-level Foley taxonomy. This dataset provides a standardized, Creative Commons-licensed resource for data-driven Foley classification, retrieval, and generation.
5. 安全、隐私与深度伪造音频 | 1 篇
8. Supervised Post-training of Speech Foundation Models for Robust Adaptation in Speech Deepfake Detection
语音深度伪造检测中鲁棒适应的语音基础模型监督后训练
AI 总结:提出混合帧后训练策略,通过帧级监督学习局部不一致性,在ASVspoof5上实现单模型EER 4.50%,在ASVspoof2021上LA与DF的EER差距仅0.16%。
链接:https://arxiv.org/abs/2606.25328
机构:Institute for Infocomm Research (I2R), Agency for Science, Technology and Research (A*STAR)(资讯通信研究院(I2R),新加坡科技研究局(A*STAR))
作者:Zihan Pan, Sailor Hardik, Jinyang Wu
英文摘要:Large speech foundation models have shown strong potential for speech deepfake detection, but direct fine-tuning is limited by a mismatch between self-supervised pre-training objectives and spoof-specific artifacts. To address this, we propose a mix-frame post-training strategy to create localized spoof-oriented perturbations and use frame-level supervision to encourage the SSL model to learn local inconsistencies that are critical for robust spoof detection. On ASVspoof5, we achieve state-of-the-art EER 4.50% for a single model without data augmentation. On ASVspoof2021 LA/DF, it further achieves only 0.16\% absolute EER gap between LA and DF, indicating strong and balanced robustness across distinct distortion conditions. These results show that supervised post-training provides an effective and practical way to adapt speech foundation models for robust deepfake detection.
6. 其他/综合语音音频 | 3 篇
9. Attractive and Repulsive Pattern Control in Sequence Generation
序列生成中的吸引与排斥模式控制
AI 总结:提出带符号模式控制方法,通过加权递归自动机和信念传播采样,在变阶马尔可夫生成中抑制或增强特定模式,实验表明负耦合减少自重复并增加多样性。
链接:https://arxiv.org/abs/2606.24911
作者:Francois Pachet
英文摘要:Variable-order Markov models preserve local symbolic syntax by adapting context length, but long continuations can enter recurring high-order "tunnels": repeated suffixes, locally periodic passages, or copied fragments longer than the formal Markov order. This paper introduces signed pattern control for variable-order Markov generation with BP-Regular sampling. A weighted recurrence automaton computes an activation R for a chosen family of target patterns, and belief propagation samples exactly from P_beta(x) proportional to P_0(x) exp(beta R(x)). Negative coupling makes the target patterns costly during sampling; positive coupling rewards the same patterns and turns them into controlled attractors. The target family may be mined online from overactive generated material, supplied by a score or style vocabulary, or designed as an experimental probe. The main experiments use the online homeostatic case, choosing patterns that become overactive in the sampling history. On six duration-bearing monophonic sources, including Bach and Telemann material, the negative branch reduces generated 8-gram self-reuse, increases the effective number of generated 8-grams, and increases coverage of training-supported 4-gram contexts while preserving substantial lower-order support. A pitch-sequence replication on five Weimar Jazz Database solos gives the same anti-reuse signature outside Baroque material. The same signed mechanism also provides a positive branch for probing attractor basins, phase transitions, and hysteresis in the underlying variable-order model.
10. EmotionAI: A Privacy-Preserving Computational Intelligence Pipeline for Speech-Emotion-Grounded Conversational Analysis
EmotionAI:一种面向语音情感驱动的对话分析的隐私保护计算智能流水线
AI 总结:提出一种完全本地的计算智能流水线,结合语音情感识别与生成式推理,实现隐私保护的对话情感分析,在RAVDESS数据集上达到48.8%准确率,零外部调用。
链接:https://arxiv.org/abs/2606.24941
机构:School of Science and Technology, Nottingham Trent University(诺丁汉特伦特大学科学与技术学院)
作者:Wai Laam Mak, Isibor Kennedy Ihianle, Pedro Machado
英文摘要:Reviewing recorded interviews for affective cues such as composure, hesitation and agitation is slow and subjective, and cloud services that could automate it require sensitive audio to leave the device. EmotionAI is a fully local Computational Intelligence (CI) pipeline that couples Speech Emotion Recognition (SER) with generative reasoning. Speaker diarisation, Whisper Automatic Speech Recognition (ASR) and a wav2vec2 emotion classifier produce per-segment affective evidence, which is then passed to an adversarial three-model local Large Language Model (LLM) panel for timestamp-grounded and citation-constrained question answering. Zero-shot evaluation on the RAVDESS four-class English subset (n = 672) exposes cross-corpus fragility rather than classifier superiority: the deployed classifier scores 48.8% accuracy, above random (24.9%) and majority (28.6%) baselines but below an in-domain MFCC + logistic-regression comparator (71.0%). The complete pipeline runs in a mean 157 s on CPU (real-time factor approximately 1.33) with zero external calls. The contribution is not state-of-the-art SER but an auditable, privacy-preserving integration of imperfect affective evidence into grounded conversational analysis, together with an honest empirical account of where cross-corpus transfer and human-centred validation still fall short.
11. What Does a Pathological Speech Assessment Model Know about Acoustic Features? A Case Study on Oral and Oropharyngeal Cancer Patients
病理语音评估模型对声学特征了解多少?以口腔癌和口咽癌患者为例的案例研究
AI 总结:通过典型相关分析研究Wav2Vec 2.0语音清晰度评估模型的可解释性,发现模型嵌入与频谱和韵律特征高度相关,第一MFCC系数相关性最高。
链接:https://arxiv.org/abs/2606.24949
作者:Tuan Nguyen (LIA, AU), Corinne Fredouille (AU, LIA), Alain Ghio (LPL), Muriel Lalain (LPL), Virginie Woisard (UT2J, UT3, LNPL)
英文摘要:This work investigates the interpretability of a Wav2Vec 2.0based speech intelligibility assessment model for oral and oropharyngeal cancer patients through canonical correlation analysis. By measuring the correlation between the model embeddings and eGeMAPS low-level descriptors (LLDs) as an interpretable reference, we analyze how acoustic information is encoded across the model layers. The analysis is conducted at two levels: individual LLDs layer-wise, and group-level: prosodic, spectral, and voice quality. Results show that the learned representations are most strongly correlated with spectral and prosodic features, with the first MFCC coefficient yielding the highest correlations across all layers. At the group level, spectral and prosodic groups achieve correlations of 0.77 and 0.71 respectively, while voice quality reaches 0.65. Beyond model interpretability, this work also offers practical guidance on acoustic feature selection for pathological speech assessment.
