微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 2 篇
2. 语音合成与声音生成 1 篇
3. 语音增强、降噪与音频修复 1 篇
4. 音乐信息检索与音乐生成 3 篇
5. 语音翻译与语音语言模型 2 篇
6. 安全、隐私与深度伪造音频 1 篇
7. 其他/综合语音音频 3 篇
1. 语音识别与关键词检测 | 2 篇
1. Scalable Keyword Spotting via Modular Network Expansion
通过模块化网络扩展实现可扩展的关键词识别
AI 总结:研究嵌入式设备上KWS模型添加新关键词难的问题,提出参数限制的模块化扩展方法,在固定操作点降低新关键词FRR,优于基线且减少MACs,保留现有关键词相关参数。
链接:https://arxiv.org/abs/2607.19918
机构:Yandex
作者:Viktor Khaymonenko, Dzmitry Saladukha, Aliaksei Rak, Alexander Rostov
英文摘要:Keyword spotting (KWS) models on embedded devices often need to add new keywords after deployment, but updates are difficult when original training data are unavailable and regressions on existing triggers are unacceptable. At a fixed operating point, our method reduces average new-keyword false reject rate (FRR) from 6.46 to 4.37 versus a parameter-matched separate-model baseline and outperforms parameter-efficient tuning baselines (adapters, LoRA), while using fewer multiply-accumulate operations (MACs) under the same added-parameter budget ($\leq$10k): 16.34M vs 18.45M/20.52M. We achieve this via parameter-capped modular expansion: the base network, including batch-normalization statistics and the core classifier, is frozen, and only a lightweight expansion branch with a separate new-keyword head is trained, preserving core logits, shipped outputs, and thresholds for existing keywords.
2. Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting
用于低成本流式关键词识别的累积和可组合相位传输
AI 总结:研究用于流式关键词识别的累积和可组合相位传输,通过将声学帧投影到复数通道,经酉旋转传输等操作,在Google Speech Commands v2数据集上取得有竞争力准确率,且训练速度快、延迟低,是简单低成本时间原语。
链接:https://arxiv.org/abs/2607.20086
机构:Carrot, Inc(胡萝卜公司)
作者:Mahesh Godavarti
英文摘要:State-space sequence models are attractive for streaming speech because they maintain compact recurrent state, but scan-style training kernels can have unfavorable constants for short audio tasks. We study cumsum-composable phase transport, a streaming-native temporal layer for keyword spotting. Each layer projects acoustic frames to complex channels, transports them by learned unitary rotations, accumulates a finite window using prefix differences, and applies a gated residual update. The same prefix representation gives exact batched training with ordinary cumulative sums and exact online inference with one prefix update per frame. Unitary transport is the key constraint: inverse rotations have norm one, keeping prefix terms well conditioned while memory is supplied by windows or block readouts. On Google Speech Commands v2 with 12 labels, mel+cumsum models retain competitive accuracy with compact baselines. The strongest single-seed run reaches 97.3\% test accuracy; a 51.6K-parameter tied model also reaches 97.3\%, and a 24.8K tied model reaches 96.8\% versus 97.1\% for a 25.6K MelCNNMaxPool baseline. In a matched cumsum-versus-scan benchmark, cumsum+window gives comparable accuracy, 94.82\% versus 94.33\%, while training 1.07x faster and reducing single-example latency from 7.09 ms to 5.01 ms on a Tesla T4. These results support cumsum phase transport as a simple low-cost temporal primitive for streaming keyword spotting.
2. 语音合成与声音生成 | 1 篇
3. StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis
StellarTTS:用于低延迟和鲁棒语音合成的稀疏时间嵌入
AI 总结:研究针对TTS系统中稳健性、延迟和韵律的权衡问题,提出基于稀疏时间嵌入策略的StellarTTS框架及语义感知编解码器,实现对音素多方面精细控制,其轻量级模型实时因子达0.08,实验证明该框架在多方面表现出色。
链接:https://arxiv.org/abs/2607.19859
作者:Kaicheng Luo, Xuefei Gong, Yutao Sun, Jinling He, Yujie Hou, Xiaoyang Xing, Huiyan Li, Bing Han, Yanmin Qian
英文摘要:The trade-off between robustness, latency, and prosody critically challenges text-to-speech (TTS) systems. Autoregressive models, despite fidelity, are slow and error-prone; non-autoregressive (NAR) alternatives, while fast, often sacrifice prosodic naturalness via rigid alignments. This paper introduces StellarTTS, a novel mobile-optimized NAR TTS framework based on a sparse temporal embedding strategy, enabling granular control of phoneme duration, pronunciation, and prosody. Furthermore, we propose a semantic-aware codec that facilitates efficient single-stage decoding. Conditioned on the sparse temporal embedding, our 83M-parameter lightweight masked generative transformer achieves a real-time factor (RTF) of 0.08. Experiments demonstrate that StellarTTS attains lower latency and stronger robustness compared to state-of-the-art TTS systems, while maintaining competitive performance in audio quality, prosodic naturalness, and speaker similarity.
3. 语音增强、降噪与音频修复 | 1 篇
4. CAPS: A Cascaded Reconstruction Model to Power Saving in Hearables Using Sub-Nyquist Sampling with Bandwidth Extension
CAPS:一种使用带带宽扩展的亚奈奎斯特采样实现可穿戴设备节能的级联重建模型
AI 总结:研究可穿戴设备中联合降低ADC采样比特分辨率和频率对功耗及音频质量的影响,提出CAPS模型,采用亚奈奎斯特采样和低比特分辨率,降低功耗3.3倍,支持移动平台流操作,确保语音清晰度,平衡效率与节能。
链接:https://arxiv.org/abs/2607.19434
机构:George Mason University(乔治梅森大学)
作者:Tarikul Islam Tamiti, Sajid Fardin Dipto, Luke Baja-Ricketts, David Vergano, Anomadarshi Barua
英文摘要:Hearables are wearable computers worn on the ear. Bone conduction microphones are used with air conduction microphones in hearables for multimodal speech enhancement in noisy conditions. Despite this potential, current models largely fail to explore how jointly reducing sampling bit resolution and sampling frequency in analog-to-digital converters (ADCs) of hearables impacts both power usage and audio quality. Furthermore, current frameworks cannot do sub-Nyquist sampling in hearables because they lack a method to reconstruct wideband signals from narrowband components. We therefore propose CAPS, which (i) intentionally employs sub-Nyquist sampling and low bit resolution in ADCs, achieving a 3.3x reduction in power consumption in hearables, and (ii) supports streaming operation on mobile platforms with an inference time of 1.36 ms and a memory footprint of 11.04 MB. CAPS ensures robust speech intelligibility in real-world settings, bridging the gap between efficiency and power savings.
4. 音乐信息检索与音乐生成 | 3 篇
5. RIME: Enabling Large-Scale Agentic Post-Production
RIME:实现大规模智能后期制作
AI 总结:研究针对音乐后期制作问题,引入RIME框架,利用POEMS工具包生成配对数据,评估多模态语言模型,展示其通过监督微调提升智能体性能,是迈向迭代音乐智能体、变革音乐制作的早期步骤。
链接:https://arxiv.org/abs/2607.19605
机构:Dartmouth College(达特茅斯学院)
作者:Noah Schaffer, Nikhil Singh
英文摘要:Almost every piece of recorded music you have ever heard was modified before it reached you; commercial releases rarely spring fully-formed from the mind of a musician. Despite the promise of music generation models for one-shot output, such fine-grained iterative refinement workflows are a complementary problem, and largely out of their reach. There is also a gap for musicians: while they can express what they want to hear, not all have the facility with studio production tools to implement the complex set of actions needed to realize these intuitions. We formalize this task as agentic post-production, wherein individual aspects of a song are targeted, refined, and combined into a final track. We argue the bottleneck is data: existing corpora do not reflect how realistic post-production chains map onto the vocabulary musicians and engineers actually use. We argue there is a language for modifying recorded music that is dense, consistent, and learnable. We introduce the Rule-based Instructions for Music Editing (RIME) framework, which generates realistic paired edit-instruction data from any baseline music dataset grounded in canonical methods, design patterns, and constraints derived from real production workflows. RIME leverages POEMS, a new toolkit that combines stem separation, mixing, and common studio effects for use by multimodal agents. We use RIME and POEMS to generate 3,000 pairs of edit instructions and ground truth audio, and use this data to evaluate existing multimodal LLMs as agents on this task, showing persistent challenges in current models' post-production capabilities. We also demonstrate RIME's ability to improve post-production agent performance via supervised fine-tuning. We see RIME as an early step toward iterative musical agents, collaborative systems that could transform music production much as interactive coding agents have reshaped software engineering.
6. RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling
RPPNet:通过边界感知建模生成长期结构旋律的感知分组节奏-音高基元
AI 总结:本文针对现有符号音乐生成模型的问题,提出RPPNet架构,通过生成可变长度的RPP序列并解码,其分组基于音乐心理学自动得出。实验显示该模型生成的旋律在长期结构和音乐性上更优,消融研究表明性能提升源于心理表征结构正确,为音乐生成提供跨学科视角。
链接:https://arxiv.org/abs/2607.19776
作者:Tieyao Zhang, Yuke Liu, Jiaxing Yu, Xinda Wu, Kejun Zhang, Genfang Chen
英文摘要:Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases often does not align with notated bar lines, leading to long-term structural fragmentation. This paper proposes RPPNet-a two-stage deep learning architecture with variable structural boundaries. It first generates variable-length Rhythm-Pitch Primitive (RPP) sequences, where each RPP encodes note count, rhythm, and contour; then decodes the RPP sequences into concrete notes. The grouping of RPPs is automatically derived from acoustic cues, auditory inertia, and similarity perception based on music psychology. Experiments show that melodies generated by RPPNet are superior in both long-term structure and musicality, with significant improvements across all subjective evaluation dimensions. Ablation studies confirm that the performance gain stems from the structural correctness of the psychological representation, rather than from model capacity. This work offers an interdisciplinary perspective for music generation, integrating music theory, computational modeling, and music psychology.
7. Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
推动全曲生成的前沿:分层自回归规划与流匹配渲染
AI 总结:该研究提出统一歌曲生成框架,支持多项任务,由四个组件构成。通过特定编码、建模、匹配及模块实现歌曲生成,还研究多种后训练策略,实验表明该框架在评估中性能具有竞争力。
链接:https://arxiv.org/abs/2607.20253
机构:Alibaba(阿里巴巴)
作者:Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li, Yunjia Li, Bin Ma, Yukun Ma, Chongjia Ni, Yufei Shi, Haoxu Wang, Menglin Wu, Jianwei Yu, Huaicheng Zhang, Han Zhao, Shengkui Zhao, Haina Zhu
英文摘要:In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation, which generates complete songs from text descriptions, lyrics, and musical attributes; Instrumental Music Generation, which creates music without vocals; and Cover Song Generation, which reinterprets existing songs with different styles while preserving their melodic content. Architecturally, our system consists of four main components: a semantic-aware tokenizer, hybird-LM, FullDiT, and a two-level melody module. The tokenizer encodes audio into 8-codebook RVQ tokens for efficient discrete music representation. Based on these tokens, hybird-LM performs hierarchical autoregressive audio-token modeling for full-song generation. To improve audio fidelity, FullDiT performs full-song flow matching in a continuous VAE latent space conditioned on codec tokens, lyrics, and text captions. For cover song generation, the melody module extracts and discretizes melody cues from reference audio to guide generation while preserving the original melodic content. Finally, we investigate DPO, GRPO, and OPD as reward-based post-training strategies for hybird-LM and apply flow-based GRPO to FullDiT to improve musicality and rendering quality. Experimental results on a multilingual automatic benchmark, complemented by the Artificial Analysis Music with Vocals leaderboard, show that the proposed framework achieves competitive performance in the evaluated settings.
5. 语音翻译与语音语言模型 | 2 篇
8. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision
SimulS2ST-Omni:通过显式轨迹监督实现数据高效的流式语音到语音翻译
AI 总结:研究长格式流式语音到语音翻译,提出仅用约2000小时配对数据的训练方法,核心是联合文本-代码轨迹监督,双流分解减轻干扰,在质量-延迟权衡上表现出色,与先进闭源系统相当。
链接:https://arxiv.org/abs/2607.19810
机构:The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
作者:Rongshen He, Xinyu Liang, Dekun Chen, Jiaqi Li, Mingjie Chen, Zhizheng Wu
英文摘要:Long-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentence-bounded supervision or demand massive paired-S2ST supervision. We introduce a training recipe enabling a speech language model for sentence-level and long-form streaming S2ST using only $\sim$2k hours of paired cross-lingual S2ST data, layered atop auxiliary supervision. Anchored by auxiliary multitask training, our approach remains robust even when the paired-S2ST budget itself is reduced by 90\%. Our core contribution, joint text-code trajectory supervision, schedules target text and acoustic semantic codes as a unified commitment path, eliminating the need for separate, unstable speech-side emission controllers. Furthermore, our two-stream Thinker--Talker factorization significantly outperforms unified-decoder baselines by decoupling linguistic reasoning from dense acoustic prediction to mitigate modality interference. Finally, our system achieves highly competitive quality-latency trade-offs on RealSI and ACL60/60-dev, matching state-of-the-art, closed-source S2ST systems such as LiveInterpret~2.0 on ASR-BLEU.
9. Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
Audio-Zero:用于细粒度音频推理的无标签自我进化
AI 总结:研究针对大型音频语言模型细粒度音频推理难题,提出无标签自我进化框架Audio-Zero,通过构建听觉自博弈游戏,利用无标签音频对比对让模型生成线索并推理,实验证明其能提升细粒度音频推理且保持广泛音频理解。
链接:https://arxiv.org/abs/2607.20166
作者:Siqian Tong, Xuan Li, Chaozhuo Li, Baolong Bi, Yiwei Wang, Yujun Cai, Shenghua Liu, Chengpeng Hao
英文摘要:Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. To bridge this gap, we introduce Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning. Audio-Zero constructs an auditory self-play game from unlabeled audio contrast pairs: most players hear a reference audio, while one odd listener hears a subtle variant. The model first generates clues describing what it hears and then identifies the odd listener by reasoning over inconsistencies among clues. Since the odd listener is known by construction, the game provides verifiable rewards without any annotated answers. Experiments with Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B on TREA, MMAU Test-mini and MMAR show that Audio-Zero improves fine-grained audio reasoning while preserving broad audio understanding. Evolutionary and diagnostic analyses further reveal that increasingly fine-grained auditory descriptions emerge naturally from game pressure.
6. 安全、隐私与深度伪造音频 | 1 篇
10. Layer-Wise Decision Fusion for Fake Audio Detection Using XLS-R
使用XLS-R进行虚假音频检测的逐层决策融合
AI 总结:研究虚假音频检测问题,提出逐层决策融合方法,利用大型语音模型多层表示,在每层决策后融合,相比其他基线在自然数据集上性能最佳,且模型更透明利于分析决策机制。
链接:https://arxiv.org/abs/2607.20023
作者:Yixuan Xiao, Ngoc Thang Vu
英文摘要:Recent fake audio detection methods often leverage large speech models to achieve robust speech representations. These models are typically very deep, providing multiple layer-wise representations. However, current works often rely solely on single layer representation or feature fusion to extract one utterance-level representation for decision making. These methods risk underutilizing rich information from multiple layers and might induce feature collapse. We propose a novel layer-wise decision fusion method that applies fusion after per-layer decision making and achieves the best cross-dataset performance on In-the-Wild dataset (EER 6.90%) compared to other strong baselines. Our model design also makes the model more transparent, allowing us to conduct detailed analysis to reveal the underlying mechanism of decision making.
7. 其他/综合语音音频 | 3 篇
11. Black-Box Optimization for Identifying and Inverting Audio Dynamic Range Control Effects
用于识别和反转音频动态范围控制效果的黑盒优化
AI 总结:研究音频动态范围压缩参数未知的问题,将其参数估计转化为黑盒优化问题,通过在感知特征空间中估计参数来最小化重建与参考信号特征描述符距离,该方法性能有竞争力,优于或匹配现有模型。
链接:https://arxiv.org/abs/2607.19645
作者:Haoran Sun, Dominique Fourer, Hichem Maaref
英文摘要:Dynamic Range Compression (DRC) is a widely used nonlinear audio effect whose parameters are often unknown, making blind estimation and inversion challenging. In this work, we formulate DRC parameter estimation as a black-box optimization problem in a perceptually motivated feature space. Given an observed signal and a reference representation, we estimate the parameters that minimize the distance between feature descriptors of the reconstructed and reference signals. Unlike gradient-based approaches, the proposed method does not require differentiability of the DRC model or the feature extraction pipeline, enabling the use of nonlinear and histogram-based descriptors. Experimental results demonstrate that the proposed method achieves competitive performance in blind parameter estimation and dry signal recovery, outperforming or matching state-of-the-art models in terms of reconstruction quality.
12. A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features
一种使用音乐理论和声学特征的人工智能生成翻唱歌曲诊断评估框架
AI 总结:研究人工智能生成翻唱歌曲的诊断评估,提出五维诊断框架,通过对30首翻唱歌曲分析发现和声进行和编排错误率高,各特征间关联有限,结果表明低级和符号摘要不能替代情境感知音乐判断。
链接:https://arxiv.org/abs/2607.19688
机构:Hangzhou Xiaoying Innovation Technology Co., Ltd. (Rythmix AI)(杭州小影创新科技有限公司(节奏混合人工智能))
作者:Yingxin Liang
英文摘要:AI-generated covers often fail through local musical errors that a global quality score cannot locate: the vocal contour may remain recognizable while the accompaniment uses the wrong harmonic function, or the output may stay in key while the arrangement remains incomplete. We present a five-dimensional diagnostic framework covering melodic pitch, harmonic progression, key consistency, style consistency, and arrangement/production quality. The benchmark contains 30 covers generated from 5 source songs by 6 systems, with expert severity ratings and 9 symbolic or acoustic features. Harmonic progression and arrangement had the highest severe-error rates (53% and 47%), whereas key consistency was better preserved. Six covers combined acceptable key consistency with severe harmonic errors. Large-leap ratio had a nominal association with melodic ratings (Spearman rho = -0.429, uncorrected p = 0.018), but no feature correlation survived the nine-test multiplicity reference. An interpretable percentile-rule pilot likewise failed to outperform a fixed majority baseline reliably across 16 dimension-level comparisons. The results separate useful diagnostic evidence from dependable automatic scoring: low-level and symbolic summaries can expose particular symptoms, but they do not replace context-aware musical judgment.
13. Ultra-Compact CNN Architectures for Tropical Bird Audio Detection on Microcontrollers
用于微控制器上热带鸟类音频检测的超紧凑型卷积神经网络架构
AI 总结:研究针对热带鸟类音频检测,提出DrongoNet系列超紧凑型卷积神经网络架构。通过在低功耗微控制器上仅在可能阳性片段触发记录,降低存储和电池成本。该架构在SEABAD数据集验证,不同型号各有优势,且全INT8量化成本低,能在不同环境部署。
链接:https://arxiv.org/abs/2607.19721
机构:Faculty of Computer Science and Information Technology, Universiti Malaya(马来西亚大学计算机科学与信息技术学院); Faculty of Electrical Engineering, Universiti Teknologi Malaysia(马来西亚理工大学电气工程学院)
作者:Muhammad Mun'im Ahmad Zabidi, Mohd Yamani Idna Idris, Norisma Idris
英文摘要:Passive acoustic monitoring of tropical biodiversity is bottlenecked by the storage and battery cost of continuously recording soundscapes in which bird vocalisations typically occupy less than 10% of the audio. Autonomous recording units built on low-power microcontrollers (typically ARM Cortex-M with $\leq$256 kB of RAM) address this by triggering only on likely-positive segments, but the on-device options are unsatisfying: coarse frequency-energy triggers such as Goertzel filters flood SD cards with false positives at $\sim$71% precision, whereas neural detectors developed for temperate single-species tasks are either too large to deploy or transfer poorly to species-rich tropical settings. We present DrongoNet, a family of three INT8 CNN detectors sized for this envelope and validated on a 50,000-clip, 1,677-species Southeast Asian tropical dataset (SEABAD). The headline model, DrongoNet-Micro (919 parameters, 6.26 kB, 0.9810 AUC, 98.3\% mean recall at {\tau} = 0.35), is a drop-in replacement for the Goertzel trigger used in commodity field recorders: at {\alpha} = 0.10 tropical prevalence it captures 8 pp more bird vocalisations than Goertzel and extends a 32 GB card from $\sim$28 to $\sim$45 days of monitoring. DrongoNet-Nano (5.09 kB) bounds the ultra-low-flash extreme; DrongoNet-Edge (33.06 kB, 0.9991 AUC) targets Linux SBCs. On SEABAD, Micro matches a retrained TinyChirp CNN-Mel baseline within 0.1 pp AUC at 28$\times$ fewer parameters, confirming that the family is deployment-agnostic across mel-spectrogram bird corpora but requires per-environment retraining. Full INT8 quantisation costs $<$0.12% AUC across all three variants.
