微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 1 篇
2. 语音合成与声音生成 2 篇
3. 说话人识别、验证与分离 1 篇
4. 音乐信息检索与音乐生成 1 篇
5. 数据集、基准与评测 1 篇
6. 其他/综合语音音频 2 篇
1. 语音识别与关键词检测 | 1 篇
1. wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval
wav2tok 2.0:可扩展的音频分词,保持显式成对令牌对齐以实现高效音频检索
AI 总结:提出wav2tok 2.0,通过分阶段训练(对比学习、向量量化、CTC对齐损失和DTW对齐帧预测)实现可扩展的语音分词,在查询示例的口语术语检测中优于基线方法。
链接:https://arxiv.org/abs/2606.26824
机构:Indian Institute of Technology Kanpur(印度理工学院坎普尔分校); KU Leuven(鲁汶大学)
作者:Adhiraj Banerjee, Vipul Arora
英文摘要:Learning discrete speech representations that preserve similarity across variable-length utterances is central to query-by-example spoken term detection (QbE-STD). While wav2tok introduced CTC-based sequence alignment to enforce token consistency, its tightly coupled clustering and alignment training recipe limits scalability. We propose wav2tok 2.0, a scalable alignment-aware speech tokenizer built on the BEST-STD backbone. wav2tok 2.0 employs staged training, first learning discriminative, speaker-invariant representations via contrastive learning and vector quantization, and then enforcing pairwise token consistency using a CTC alignment loss and a novel DTW-aligned framewise prediction objective with adaptive weighting. Experiments show that wav2tok 2.0 consistently outperforms BEST-STD and general-purpose tokenizers on QbE-STD while remaining efficient and scalable.
2. 语音合成与声音生成 | 2 篇
2. VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation
VoiceTTA: 通过基于强化学习的测试时自适应增强零样本文本到语音
AI 总结:提出VoiceTTA,一种基于强化学习的测试时自适应方法,通过优化可学习前缀和风格奖励,提升预训练零样本TTS模型对罕见语音风格的模仿能力。
链接:https://arxiv.org/abs/2606.26534
机构:The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)); Tencent(腾讯)
作者:Tianxin Xie, Chenxing Li, Dong Yu, Li Liu
英文摘要:Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen speaking styles from uncommon scenarios (e.g., crosstalk, dialects). Moreover, fine-tuning pretrained models requires large, high-quality datasets, limiting rapid personalization. We propose VoiceTTA, a reinforcement learning-based test-time adaptation (TTA) method that improves voice imitation of pretrained zero-shot TTS models. VoiceTTA introduces two style rewards based on coefficient-of-variation differences of F0 and energy, combined with speaker similarity and intelligibility (WER from a pretrained Whisper model), and optimizes learnable prefixes via group relative preference optimization (GRPO) in a flow matching-based model at inference time. Extensive experiments demonstrate substantial improvements on uncommon speech prompts, outperforming state-of-the-art baselines. Audio samples are available at this https URL
3. Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis
Sarashina2.2-TTS:通过数据扩展和针对性数据合成解决日语语音生成中的汉字多音问题
AI 总结:针对日语汉字多音问题,提出Sarashina2.2-TTS系统,通过扩展训练数据至36.1万小时并设计覆盖全部常用汉字的针对性数据增强管道,显著提升读音准确率,在零样本日语语音合成中达到最高说话人相似度。
链接:https://arxiv.org/abs/2606.25369
机构:SB Intuitions
作者:Lianbo Liu, Shiao Zhu, Kai Washizaki, Reo Yoneyama, Haesung Jeon, Mengjie Zhao, Yusuke Fujita, Hao Shi, Nao Yoshida, Yuan Gao, Roman Koshkin, Yukiya Hono, Yui Sudo
英文摘要:While large language model (LLM)-based text-to-speech (TTS) systems have achieved high-quality speech synthesis, most existing systems focus on English and Chinese. Japanese, however, remains under-explored, and its unique linguistic challenges, such as widespread context-dependent kanji polyphony, have yet to be adequately tackled. Here we introduce Sarashina2.2-TTS ( this https URL ), a Japanese-centric LLM-TTS system that tackles these challenges through a dual approach: data strategy and evaluation methodology. First, we scale training to approximately 361k hours of speech, incorporating a balanced mix of Japanese and English data. Furthermore, we design a targeted data augmentation pipeline covering all 2,136 Joyo (regular-use) kanji designated by Japan's Agency for Cultural Affairs to efficiently address kanji polyphony disambiguation. Second, we introduce the Joyo Kanji Yomi Benchmark ( this https URL ), covering all 2,136 Joyo kanji and their 4,378 readings. Alongside this benchmark, we propose Kana-CER, a metric that compares synthesized speech against reference readings in the kana space, eliminating orthographic variations to directly measure pronunciation correctness. Experiments demonstrate that our targeted data augmentation significantly improves reading accuracy. Overall, Sarashina2.2-TTS achieves state-of-the-art kanji-level reading accuracy and matches top baselines on general sentence-level pronunciation, while delivering the highest speaker similarity in zero-shot Japanese speech synthesis. Furthermore, cross-lingual evaluation reveals that Sarashina2.2-TTS is the only system that maintains stable Japanese pronunciation regardless of the prompt language, confirming that our balanced training approach improves cross-lingual robustness.
3. 说话人识别、验证与分离 | 1 篇
4. Neural Speaker Diarization via Multilingual Training: Evaluation on Low-Resource Nepali-Hindi Speech
通过多语言训练的神经说话人日志:低资源尼泊尔-印地语语音评估
AI 总结:本文采用多语言训练方法,比较EEND-EDA和DiaPer两种架构在低资源尼泊尔-印地语语音上的说话人日志性能,DiaPer在多人场景下表现更优。
链接:https://arxiv.org/abs/2606.26144
机构:Department of Electronics and Computer Engineering, Pulchowk Campus, Institute of Engineering(电子与计算机工程系,普尔乔克校区,工程学院)
作者:Samip Neupane, Sandesh Pokhrel, Sandesh Pyakurel, Basanta Joshi
英文摘要:Speaker diarization, the task of determining "who spoke when" in a multi-speaker recording, is a critical component in applications such as meeting transcription, accessibility tools, and multilingual information retrieval. While end-to-end neural diarization systems have achieved strong performance for English and other high-resource languages, their effectiveness degrades substantially for underrepresented languages where annotated speech data is scarce. This paper investigates speaker diarization for low-resource Nepali-Hindi speech through a multilingual training approach, comparing two modern architectures: EEND with encoder-decoder attractors (EEND-EDA) and EEND with Perceiver-based attractors (DiaPer). Both models are trained on a multilingual corpus combining English speech from LibriSpeech, diverse speaker recordings from VoxCeleb, and separately collected Nepali and Hindi audio, a setup designed to reduce language bias and encourage cross-lingual generalization. We evaluate both models across 2-speaker, 3-speaker, 4-speaker, and mixed-speaker scenarios on LibriSpeech, VoxCeleb, and Nepali-Hindi (NeHi) test sets. DiaPer achieves stronger overall performance than EEND-EDA, particularly in more challenging multi-speaker conditions, obtaining DERs of 3.28%, 2.02%, 4.05%, and 4.76% on NeHi 2-speaker, 3-speaker, 4-speaker, and mixed-speaker settings, respectively, compared to 1.50%, 9.68%, 16.17%, and 11.19% for EEND-EDA. These results demonstrate the viability of Perceiver-based end-to-end neural diarization for low-resource multilingual speech processing.
4. 音乐信息检索与音乐生成 | 1 篇
5. Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation
像法官一样聆听:一种用于自动歌唱表现评估的音乐感知框架
AI 总结:提出MusicJudge框架,通过模态引导的块对齐多模态分析,结合歌词正确性与音高节奏保真度,实现自动歌唱质量评估,实验验证与人类专家判断高度一致。
链接:https://arxiv.org/abs/2606.26451
机构:Samsung R&D Institute Bangalore, India(三星研发中心班加罗尔,印度)
作者:Neelam Saini, Sourav Ghosh
英文摘要:Automatic singing quality assessment (SQA) requires evaluating lyrical correctness and musical fidelity while handling expressive variations. However, existing systems largely rely on either acoustic cues or lyric transcriptions exclusively, limiting holistic performance evaluation. Furthermore, their integration is non-trivial due to challenges in robust singing transcription amid melisma, vibrato, and tempo elasticity. To this end, we propose MusicJudge, a modality-guided framework for automated SQA that performs block-aligned multimodal analysis by coupling lyric correctness with pitch-rhythm fidelity. It detects semantically meaningful lyric blocks using multi-signal matching that integrates semantic embeddings, lexical similarity, and phonetic alignment. To improve singing audio transcription, we introduce Modality-Guided LoRA for ASR fine-tuning. Experiments across datasets demonstrate strong agreement with human expert judgments and validate the generalizability of MusicJudge.
5. 数据集、基准与评测 | 1 篇
6. Soroll-IA: A Weakly Labeled Audio Dataset for Real-World Industrial Port Monitoring
Soroll-IA:用于真实工业港口监测的弱标注音频数据集
AI 总结:提出一个在西班牙瓦伦西亚真实工业港口环境中录制的弱标注音频数据集,包含22小时音频和26种声音事件类别,用于支持音频标注、弱监督声音事件检测及机器学习研究。
链接:https://arxiv.org/abs/2606.26195
作者:Javier Naranjo-Alcazar, Jordi Grau-Haro, Ruben Ribes-Serrano, Marta Garcia-Ballesteros, Pedro Zuccarello
英文摘要:Soroll-IA is a weakly labeled environmental audio dataset recorded in a real-world industrial port environment in Valencia (Spain) using two fixed sensing nodes. The dataset comprises approximately 22 hours of audio segmented into 7,396 clips and covers 26 sound event classes representative of industrial port acoustic activity commonly observed in such environments, such as crane sirens, train movements, traffic, and other logistical and industrial sounds. Recordings were captured under highly challenging acoustic conditions, including strong background noise, long-distance sources, and frequent event overlap. All audio clips were annotated by domain experts following a weak labeling strategy, where tags indicate the presence of sound events within a clip without temporal localization. To account for inter-annotator variability, two ground-truth versions are released: one without cross-validation, where a class is considered present if annotated by at least one expert, and a second, more conservative version based on cross-validation, where agreement by at least two-thirds of the annotators is required. The dataset is intended to support research in audio tagging, weakly supervised sound event detection, and machine learning under realistic industrial acoustic conditions. Benchmark results are provided using two complementary architectures: CNN14 representing high-capacity convolutional models for audio tagging, and MobileNetV2, selected for its suitability in real-time classification on low-resource edge devices. To the best of current knowledge, Soroll-IA constitutes an available dataset dedicated exclusively to industrial port acoustic environments, aiming to foster advances in robust environmental sound analysis for safety-critical and operational monitoring applications. The dataset is available online and collected under Attribution-NonCommercial 4.0 International license.
6. 其他/综合语音音频 | 2 篇
7. WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation
WQ-Fusion: 用于跨域音频表示的动态门控注意力
AI 总结:提出WQ-Fusion双编码器框架,通过自适应特征调制和元素级门控注意力动态融合Whisper与Qwen,在Interspeech 2026音频编码器能力挑战赛中取得0.836总分,显著优于单编码器基线。
链接:https://arxiv.org/abs/2606.26556
机构:School of Electronic Information, Wuhan University(武汉大学电子信息学院); Tencent AI Lab Seattle(腾讯AI实验室西雅图); CIAIC, Northwestern Polytechnical University(西北工业大学CIAIC); INRS-EMT, University of Quebec(魁北克大学INRS-EMT)
作者:Mingda Lin, Lei Ding, Xinyue Zhou, Tiantian Xiong, Hanchen Pei, Gongping Huang, Hao Zhang, Jingdong Chen, Jacob Benesty
英文摘要:While pre-trained models excel in specialized tasks, learning universal representations across diverse acoustic domains remains challenging. To address this, we propose WQ-Fusion, a robust dual-encoder framework for cross-domain audio representation learning. Overcoming the limitations of static concatenation, WQ-Fusion integrates whisper and qwen via an Adaptive Feature Modulation module and a novel element-wise gated attention mechanism. This design enables dynamic feature selection, allowing the model to selectively emphasize relevant acoustic and semantic dimensions. Extensive experiments on the Interspeech 2026 Audio Encoder Capability Challenge (Track A) benchmark demonstrate that by effectively routing heterogeneous information, WQ-Fusion achieves a superior overall score of 0.836, significantly outperforming the strongest single-encoder baseline.
8. Elastic Time: Dynamic Frame Rate Bottlenecks for Neural Audio Coding
弹性时间:神经音频编码的动态帧率瓶颈
AI 总结:提出弹性时间方法,通过轻量级潜在预测器实现动态帧率瓶颈,将固定帧率自编码器转换为动态帧率,在推理时高效选择边界,提升效率-质量权衡。
链接:https://arxiv.org/abs/2606.27320
机构:University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校); Massachusetts Institute of Technology(麻省理工学院)
作者:Dimitrios Bralios, Paris Smaragdis, Minje Kim
英文摘要:Neural audio autoencoders have become a core component of compression, feature extraction, and generation. However, while existing systems support variable bitrate, the vast majority of models still operate at a fixed latent frame-rate, allocating equal temporal budget to regions with very different information density, which can result in unnecessarily long sequences. We introduce Elastic Time, a dynamic frame-rate bottleneck that converts fixed-frame-rate autoencoders to dynamic ones. Our method learns a lightweight latent predictor used to decide which frames can be skipped and later reconstructed, enabling efficient greedy boundary selection at inference. Experiments show our method enables deployment-time rate control while improving efficiency-quality tradeoffs relative to baselines. Overall, we provide a flexible mechanism for adjusting temporal resolution in audio autoencoders, potentially facilitating more efficient downstream modeling for generation and long-context tasks.
