微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 1 篇
2. 语音合成与声音生成 2 篇
3. 音乐信息检索与音乐生成 2 篇
4. 语音翻译与语音语言模型 1 篇
1. 语音识别与关键词检测 | 1 篇
1. NPUsper: Eliminating Redundant Computation for Real-Time Whisper on Mobile NPUs
NPUsper:消除移动NPU上实时Whisper的冗余计算
AI 总结:提出NPUsper系统,通过在线检测幻觉令牌和受控展开技术,消除冗余计算,在移动NPU上实现Whisper的实时转录,显著降低延迟和功耗。
链接:https://arxiv.org/abs/2607.01108
机构:Korea University(高丽大学); University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
作者:Sihyeon Lee, Hojeong Lee, Sungwon Woo, Chengpo Yan, Suman Banerjee, Seyeon Kim
英文摘要:We present NPUsper, a live transcription system that makes Whisper efficient on mobile NPUs by eliminating redundant computation. To avoid the heavy padding used by prior streaming systems, NPUsper detects hallucinated tokens online from temporal patterns in decoder cross-attention, allowing each inference round to process short audio inputs with minimal carryover. For efficient mobile-NPU execution, we propose controlled unrolling, which executes autoregressive decoding as K-step chunk graphs, removing unnecessary KV-cache computation and reducing graph-dispatch overhead. NPUsper achieves up to 4.84x lower per-word latency, up to 33.2x lower time-to-first-token (TTFT), and up to 88.64% lower average power consumption compared with baselines, while maintaining comparable transcription accuracy. The code is available at this https URL.
2. 语音合成与声音生成 | 2 篇
2. Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis
增强流匹配:一种统一引导框架实现高效鲁棒的语音合成
AI 总结:针对流匹配语音合成中的高推理延迟和音色泄露问题,提出统一引导框架,通过数据引导(异构增强)和模型引导(轨迹整流与内在引导目标)提升效率与鲁棒性,实现近三倍加速并改善说话人相似度。
链接:https://arxiv.org/abs/2607.00363
机构:Zuoyebang, China(作业帮,中国)
作者:Zuda Yu, Qianhui Xu, Ting Chen, Junhui Zhang, Tao Fu, Hongjiang Yu, Qiangqing Wang, Yang Song
英文摘要:Flow Matching (FM) has emerged as a powerful paradigm for speech generation but remains constrained by high inference latency and timbre leakage. To address these bottlenecks, we propose a unified guidance framework that enhances generation efficiency and robustness through two complementary strategies. On the data front, we introduce Data-guidance via heterogeneous augmentation, encouraging the model to disentangle linguistic content from acoustic residue. In parallel, we propose an enhanced Model-guidance mechanism that synergizes trajectory rectification with a novel intrinsic guidance objective. This approach distills conditional knowledge into network weights and straightens inference trajectory path, thereby eliminating Classifier-Free Guidance (CFG) overhead. Experiments demonstrate that our framework accelerates inference by nearly three times while effectively improving speaker similarity compared to state-of-the-art baselines.
3. A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models
文本到语音模型中可组合情感操控的几何视角
AI 总结:通过线性探测和局部本征维度分析语音语言模型与条件流匹配模块的情感表示几何特性,发现SLM具有低维情感子空间和强解耦能力,而CFM因说话人-情感纠缠导致跨说话人泛化差,联合操控增强情感强度但降低比例控制和语音质量。
链接:https://arxiv.org/abs/2607.00946
机构:The University of Melbourne, Australia(墨尔本大学)
作者:Siyi Wang, James Bailey, Ting Dang
英文摘要:While prior work has explored emotion control in hybrid text-to-speech systems, the geometric properties of these modules, and their implications for steerability, remain poorly understood. We present the first comparative study of speech language model (SLM) and conditional flow-matching (CFM) modules as activation steering sites for mixed emotion speech synthesis. We first characterize emotion representations using linear probing and local intrinsic dimensionality (LID), and then evaluate single-site and joint steering for mixed-emotion synthesis. Our results show that SLM offers a clean, low-dimensional emotion-specific subspace with strong speaker--emotion disentanglement, while CFM exhibitspoor cross-speaker generalization due to speaker--emotion entanglement. Joint steering increases emotion intensity but degrades proportional control and speech quality on in-distribution data. These findings provide practical guidance for multi-site activation steering in hybrid TTS systems and highlight the importance of representation geometry in controllable speech generation.
3. 音乐信息检索与音乐生成 | 2 篇
4. A Text-Steerable Instrument for Sketching Procedural Soundscapes via Language Models
一种通过语言模型驱动文本可操控的草图式程序化声景乐器
AI 总结:提出一种实时音乐界面,将自然语言场景描述转化为可演化的程序化声景,通过直接参数调整实现细粒度控制,并采用三种可互换后端生成连贯音频流。
链接:https://arxiv.org/abs/2607.00309
机构:Rama Labs(Rama实验室)
作者:Prabal Gupta (Rama Labs, Kitchener, Canada)
英文摘要:We present a real-time musical interface that converts natural-language scene descriptions into evolving procedural soundscapes. A performer types a prompt such as "warm jazz cafe at midnight" and steers it through direct parameter adjustments - stepping brightness down, switching a rhythm style - each producing a predictable, audible shift without re-prompting. Where GPU-bound text-to-audio systems synthesize monolithic waveforms, our instrument generates human-readable configurations over a categorical schema, enabling fine-grained performer control; most valid combinations are designed to sound musically coherent. Three interchangeable backends - embedding retrieval for sub-second CPU-only use, hosted LLMs via API, and a fine-tuned 270M local model - all emit the same schema. A live generator architecture continuously emits audio while resolving new instructions in the background, crossfading seamlessly when ready; even when an LLM takes 5-12 seconds to respond, the audience hears uninterrupted sound - reframing text-to-music as an ongoing performable stream rather than a one-shot generation. We evaluate text-audio semantic alignment using LAION-CLAP on held-out prompts as a technical proxy, finding that retrieval-based configuration outperforms random valid configurations on this metric, while noting that LAION-CLAP also informed retrieval-map construction. We report performance observations, informal listener feedback, and release materials for the SDK, dataset artifacts, model, and audiovisual performance interface.
5. Evaluating Pretrained Music Embeddings for Cross-Performance Jazz Standard Recognition
评估预训练音乐嵌入在跨演奏爵士标准曲识别中的表现
AI 总结:研究利用预训练音乐嵌入进行跨演奏爵士标准曲识别,对比从头训练的谐波CNN基线,发现预训练嵌入在top-k结果上更优但对演奏者身份敏感,轻量对比投影可部分缓解。
链接:https://arxiv.org/abs/2607.00777
作者:Çağrı Eser
英文摘要:Recognizing jazz standards from audio is a challenging form of tune-level music retrieval: different performances of the same standard may vary in tempo, key, arrangement, instrumentation, improvisational content, and even whether the head melody is present. We study this problem using a curated subset of the Jazz Trio Database designed for cross-performance standard recognition. We compare a from-scratch trained Harmonic CNN baseline against frozen pretrained music representations from recent music understanding foundation models, using both supervised probing and nearest-neighbor retrieval. Our results suggest that from-scratch spectrogram models overfit strongly to training performances, while pretrained embeddings provide better top-$k$ results but are sensitive to performer identity, which can be partially reduced with a lightweight contrastive projection. Our findings motivate jazz standard recognition as a useful stress test for music representation models and as a step toward retrieval-based standard identification. Project page: this https URL.
4. 语音翻译与语音语言模型 | 1 篇
6. Adaptive Perturbation Selection for Contrastive Audio Decoding
对比音频解码的自适应扰动选择
AI 总结:针对大型音频语言模型幻觉问题,提出自适应选择最优音频扰动作为对比解码负分支的方法,在时序、存在性等任务上提升准确率。
链接:https://arxiv.org/abs/2607.00247
机构:Google(谷歌); University of Iowa(爱荷华大学)
作者:Aaron Isidore Grace, Zhouyuan Huo, Weiran Wang
英文摘要:Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors. While contrastive decoding (CD) offers training-free mitigation, existing methods rely on blunt perturbations like masking or noise, leaving structured audio transformations unexplored. We explore this design space by evaluating a diverse library of targeted audio perturbations and adaptively selecting the optimal negative branch for each task and example. First, we improve upon earlier prompt engineering by showing that a simple binary yes/no constraint reduces the model's tendency to falsely confirm absent audio features. Second, evaluating our library across temporal, spectral, frequency, and amplitude domains reveals that optimal transformations are highly task-dependent; for instance, reversing the audio array disrupts temporal coherence, raising accuracy on the temporal order task from 74.7% to 81.4%. Finally, we trained a light-weight perturbation selector on model hidden states to dynamically route negative branches, yielding an additional +4.3% gain on the existence task.
