今日论文合集:CS.SD语音与音频 | 共 13 篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准


1. Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

Motion-Omni:面向口语对话的端到端联合语音与全身动作生成模型

AI 总结:Motion-Omni是端到端联合语音与全身动作生成框架,基于Qwen2.5-7B-Instruct实现,在动作指标接近教师级联的同时速度提升5.4倍,词错误率为全模态系统最低,还发布了相关数据集与评估协议。

链接:https://arxiv.org/abs/2609.04250

机构:Peking University(北京大学); The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

作者:Chengqian Ma, Wei Tao, Haoyu Zhang, Yiwen Guo

英文摘要:An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first generates the spoken response and then runs a motion model over the finished audio, which requires a second full inference pass and precludes any joint optimisation between the two. We present Motion-Omni, an end-to-end framework in which a spoken dialogue model natively outputs explicit facial expression together with hand, upper-body and lower-body motion, generated directly from the hidden states that produce the speech. Joint training is not optional here: with the speech pathway frozen, motion remains misaligned with the audio, and co-adapting the LLM, Speech Generator and Motion Generator under both objectives is what recovers alignment while retaining spoken-dialogue ability. Supervision comes from a scalable, model-agnostic pipeline that pseudo-labels consistent-voice speech responses with a replaceable motion teacher, yielding 422,856 quality-ranked pairs (1,402 hours). We further release SwDA-500 and, to our knowledge, the first public evaluation protocol for stochastic open-ended full-body spoken dialogue, matching audio across motion systems while unifying rendering, automatic metrics, human evaluation, and latency measurement. Instantiated with a Qwen2.5-7B-Instruct backbone, Motion-Omni-Q7 matches the same-audio teacher cascade to within 2% on reference-free motion metrics while responding 5.4 x faster (RTF=0.78, faster than real time), surpasses all non-teacher cascades on beat correlation and diversity, and reaches a 2.62% word error rate, the lowest among the omni-modal systems compared.


2. ProLombard: Structured Multi-Scale Modeling for Normal-to-Lombard Speech Conversion

链接:https://arxiv.org/abs/2609.04828

机构:National Engineering Research Center for Multimedia Software(国家多媒体软件工程技术研究中心); School of Computer Science, Wuhan University(武汉大学计算机学院); Hubei Key Laboratory of Multimedia and Network Communication Engineering(多媒体网络通信工程湖北省重点实验室); Guangdong OPPO Mobile Telecommunications Corp.(广东欧珀移动通信有限公司)

作者:Hongyang Chen, Xinmeng Xu, Youqiang Zheng, Xingyu Liu, Yuhong Yang, Zhongyuan Wang, Weiping Tu, Song Lin

英文摘要:Normal-to-Lombard (N2L) speech conversion aims to improve speech intelligibility in noisy environments by transforming normal speech into Lombard-style speech while preserving linguistic content, speaker identity, and speech quality. Despite recent progress, existing methods typically model the Lombard effect at the utterance level or the frame level, overlooking its hierarchical nature and its entanglement with both speaker identity and phoneme-level content. This limitation leads to Lombard leakage in speaker representations and incomplete separation between Lombard characteristics and linguistic content. In this work, we propose ProLombard, a structured multi-scale N2L framework that explicitly models the Lombard effect across utterance-, phoneme-, and frame-level representations. To address Lombard-speaker entanglement, we introduce an aligned speaker encoder (ASE) that suppresses Lombard leakage by aligning Lombard-speech speaker embeddings with their normal-speech counterparts. To achieve more complete Lombard-content disentanglement, we develop a phoneme-aware disentanglement and injection mechanism that extends conventional frame-level modeling to the phoneme level. Furthermore, we design a vector quantization (VQ)-median module that provides robust phoneme-level representations through VQ-based segmentation and median-frame-based aggregation. Extensive experiments on Mandarin and English Lombard datasets demonstrate that the proposed approach consistently improves speech intelligibility, Lombard similarity, and perceptual quality over baselines while maintaining speaker identity. These results highlight the importance of structured multi-scale modeling for effective N2L speech conversion.


3. Grounded Decoding for Autoregressive Speech Enhancement via Adaptive Code-Space Grounding and Local LLM Refinement

基于自适应代码空间对齐与局部大语言模型优化的自回归语音增强的对齐解码

AI 总结:该研究针对现有语音增强方法的缺陷,提出含自适应代码空间对齐(SNR-CSG)及LLM优化(GNR-LLM)的框架,实验验证其可提升低信噪比语音感知质量且不损失内容保真度。

链接:https://arxiv.org/abs/2609.04245

机构:Kyoto University(京都大学); Meta; Brno University of Technology(布尔诺理工大学); School of Electronic Information, Wuhan University(武汉大学电子信息学院); Academia Sinica(中央研究院(中国台湾)); National Institute of Information and Communications Technology, Japan(日本信息通信研究机构)

作者:Hao Shi, Yuan Gao, Zhaoheng Ni, Junyi Peng, Gongping Huang, Yu Tsao, Xugang Lu

英文摘要:Large language model (LLM)-based autoregressive speech enhancement (SE) produces natural speech using learned clean-speech priors, but may hallucinate content unsupported by the input. Deterministic SE better preserves observation-coupled evidence, yet often retains residual noise or local distortion. We propose an evidence-grounded generative SE framework that uses a deterministic estimate as imperfect evidence. A Whisper-guided DPRNN produces an enhanced waveform, which is blended with the observation and tokenized into a discrete evidence sequence. The evidence conditions an autoregressive clean-speech token generator and is reused during decoding through Code-Space Grounding (CSG), which penalizes candidates according to their Hamming distance in the factorized finite-scalar-quantized (FSQ) space. Because the appropriate grounding strength depends on acoustic difficulty, we introduce SNR-Conditioned CSG (SNR-CSG), which maps a calibrated residual-SNR estimate to an utterance-level strength and constructs an adaptive grounded anchor. Although grounding improves content fidelity, the anchor may retain local acoustic defects inherited from the evidence. Since such defects are predominantly local in the FSQ space, nearby tokens may provide better acoustic realizations without large departures from the observation-supported trajectory. We therefore propose Grounded Neighborhood Refinement with LLM ranking (GNR-LLM). It performs one additional teacher-forced pass conditioned on the grounded-anchor history, intersects the LLM top-$K$ candidates with a local FSQ Hamming neighborhood. Experiments on in-domain, controlled-SNR, and DNS no-reverb conditions show that SNR-CSG provides robust automatic grounding, while GNR-LLM substantially improves low-SNR perceptual quality without sacrificing content fidelity.


4. SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds

SCAPES:面向环境声音的语义条件自回归先验模型

AI 总结:本文提出轻量高效的生成式音频模型SCAPES,可通过语义控制合成高保真环境声音,仅需单个消费级GPU即可训练,支持语义插值,相关资源已公开。

链接:https://arxiv.org/abs/2609.04634

作者:Esteban Gutiérrez, Lonce Wyse, Frederic Font, Xavier Serra

英文摘要:As generative audio models grow in complexity, the computational and ecological costs of synthesizing everyday sounds have become increasingly prohibitive, often requiring industrial-scale resources and massive datasets. In this paper, we present SCAPES: a Semantically Conditioned Autoregressive Prior for Environmental Sounds. SCAPES is a lightweight, resource-efficient generative model designed to synthesize high-fidelity environmental textures through high-level semantic control. By operating on the continuous latent manifold of a neural audio codec, our approach bypasses the rigid structural constraints inherent to discrete tokenization. We propose a segmentation strategy that decomposes audio into overlapping segments, enabling a Continuous Normalizing Flow (CNF) to model the evolution of latent trajectories using Flow Matching. Our experiments demonstrate that a 36-million parameter instance of SCAPES can be trained on limited, uncurated datasets using a single consumer-grade GPU. Notably, convergence is achieved after training for approximately twice the source audio duration, yielding high-fidelity outputs with robust long-term stability and semantic consistency. Furthermore, we showcase the model's capacity for smooth semantic interpolation, providing a flexible and accessible tool for open research and creative sound design. Code, pretrained weights, audio examples, and an interactive demo are publicly available on our project page this https URL


5. KanAdapter: A Kolmogorov-Arnold Network-based Plug-and-Play Module for Efficient Fine-tuning of Foundation Speech Models

链接:https://arxiv.org/abs/2609.05281

机构:National University of Singapore(新加坡国立大学); Institute of Advanced Intelligence and Computing (IAIC), A * STAR(新加坡科技研究局高级智能与计算研究所)

作者:Phuong Tuan Dat, Phuong Khai Minh, Tran Huy Dat

英文摘要:Fully fine-tuning self-supervised learning (SSL) speech models for downstream tasks is computationally prohibitive, and existing parameter-efficient fine-tuning approaches predominantly rely on MLP-based adapters whose fixed activation functions limit their representational expressiveness under tight parameter budgets. We propose \textbf{KanAdapter}, a lightweight adapter framework that replaces conventional MLP bottlenecks with Group-Rational Kolmogorov-Arnold Network (GR-KAN) modules for more expressive and parameter-efficient adaptation. Following a parallel bottleneck design, KanAdapter inserts trainable GR-KAN branches alongside frozen Transformer encoder blocks and leverages weight transfer from pre-trained MLP layers for stable initialization. Across speaker verification, speech emotion recognition, and deepfake detection, KanAdapter achieves up to 97.5\% reduction in trainable parameters relative to full fine-tuning while remaining highly competitive, and consistently outperforms AdaptFormer under comparable parameter budgets. In continual learning, it yields up to 83.6\% error reduction over full fine-tuning and MLP-based adapters, which we attribute to the localized nature of GR-KAN's rational activations that mitigates catastrophic forgetting. To our knowledge, this is the first work to explore KAN-based modules for parameter-efficient fine-tuning of speech foundation models.


6. Beyond SDR: How Music Source Separation Reshapes Rhythm-Relevant Signal Properties

超越SDR:音乐源分离如何重塑与节奏相关的信号特性

AI 总结:该研究量化四类音乐源分离模型对节奏关键信号特性的影响,发现SDR无法反映起音外的节奏特性,模型选择和输入长度会影响节奏相关信号,需在节奏研究中明确报告这些变量。

链接:https://arxiv.org/abs/2609.04224

机构:Universitat Autònoma de Barcelona(巴塞罗那自治大学)

作者:Chuxin Ding

英文摘要:Music source separation (MSS) is increasingly used not to remix music but to measure it: separated drum stems feed studies of microtiming, dynamics, and groove. The field evaluates separators almost exclusively by signal-to-distortion ratio (SDR), yet microrhythm research shows that a sound's perceived temporal location (its p-centre) is co-determined by its attack and envelope, precisely the properties SDR was not designed to protect. We quantify what four open separators spanning four architecture generations (Spleeter, HT-Demucs, BS-Roformer, SCNetXL) do to rhythm-critical signal properties, using the 50-track MUSDB18-HQ test set, where true stems make every claim falsifiable. Three findings emerge. (1) Onset timing is safe: onset F-measure tracks SI-SDR (Spearman rho = 0.62) and is invariant to input length. (2) Transient and dynamic shape are not: their distortion correlates only weakly with SI-SDR (|rho| <= 0.29), and the model ranking inverts - the SDR leader distorts drum attacks twice as much as its capability-matched CNN counterpart, while the SDR-worst model preserves dynamics better than a mid-pack one. Each model imposes a systematic, model-specific bias on the dynamic profile. (3) Input length reshapes the rendered attack of a fixed passage (marginally more for the transformer; paired p = 0.044) while leaving onset locations untouched. For rhythmic studies, separator choice and input conditions are methodological variables to be reported, and SDR alone cannot stand in for them.


7. Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes

链接:https://arxiv.org/abs/2609.04516

机构:Carnegie Mellon University(卡内基梅隆大学)

作者:Yushi Ye, Wilson Zheng, Yongyi Zang

英文摘要:Recent work on controllable music generation has focused on autoregressive models, leaving diffusion-based systems comparatively underexplored. We present a lightweight method for steering the pitch content of audio produced by Stable Audio Open, a latent diffusion model for music synthesis. A small convolutional probe containing approximately 125k parameters is trained to decode frame-level pitch-class activations from the model's variational autoencoder latent space, using paired audio and MIDI data. At inference time, the frozen probe serves as a differentiable loss function: its gradient with respect to the denoising latent is used to nudge generation toward a user-specified pitch-class sequence, requiring no retraining or architectural modification of the base model. Across 27 evaluation trials spanning 9 text prompts and 3 target melodies, probe-guided generation increases melodic coherence by 2.4x over the unguided baseline (p < 1e-5, Wilcoxon signed-rank test), demonstrating that musically meaningful structure is both recoverable and steerable in diffusion-based music latent spaces.


8. Harmonica: Accurate and Lightweight Instrument-Agnostic Music Transcription

Harmonica:准确且轻量的乐器无关音乐转录模型

AI 总结:本文提出基于多深度谐波卷积的乐器无关音乐转录模型Harmonica,不同规模变体在准确率或推理效率上表现优异,纳米变体帧F1值优于Basic Pitch,多深度谐波卷积可有效提升转录性能。

链接:https://arxiv.org/abs/2609.04640

机构:BandLab Technologies(BandLab科技公司)

作者:Longshen Ou, Héctor Martel, Joe Hennessy-Priest, Taemin Cho

英文摘要:This paper introduces Harmonica, a family of instrument-agnostic music transcription models built around multi-depth harmonic convolution. At each model scale, Harmonica achieves the best performance among the evaluated models: the x-large model attains state-of-the-art performance in instrument-agnostic transcription, while the medium variant offers competitive accuracy with faster inference than all baselines. Pushing the limit of computational efficiency, the nano variant has only 26.3K parameters and runs at 1,622.5 times real time, yet achieves a frame F1 of 0.796 on the development set, 14.6 percentage points higher than Basic Pitch. We further demonstrate that multi-depth harmonic convolution effectively exploits harmonic information to benefit transcription performance through comparative experiments with existing harmonic aggregation methods, including harmonic stacking, harmonic attention, single-depth harmonic convolution, and the HD-Conv layer.


9. VocalCoachBench: Benchmarking Audio-Language Models on Expert Feedback for Singing

VocalCoachBench:评估音频-语言模型在歌唱领域专家反馈上的基准

AI 总结:该研究推出首个歌唱领域专家声乐教练反馈评估基准VocalCoachBench,通过实验发现现有音频-语言模型在细粒度问题标签识别等方面存在明显差距,推动音频-语言评估向分析性反馈发展。

链接:https://arxiv.org/abs/2609.04241

机构:KAIST(韩国科学技术院)

作者:Hayeon Bang, Hounsu Kim, Wonil Kim, Juhan Nam

英文摘要:Recent audio-language models are increasingly evaluated on recognizing, describing, and reasoning about audio, but expert-facing applications require a different capability: producing feedback that identifies problems and suggests corrective actions grounded in the input. We introduce VocalCoachBench, a benchmark for evaluating audio-language models on expert vocal coaching feedback for singing. VocalCoachBench contains 515 recordings annotated by 18 professional vocal trainers, yielding 1,056 expert submissions and 12,051 atomic coaching claims. It comprises a same-song subset for controlled comparison and a diverse-song subset for segment-grounded feedback across varied songs and recording conditions. To accommodate the open-ended nature of expert feedback, VocalCoachBench sep- arates deterministic structured targets from claim-based assessment of free-form diagnosis and corrective guidance. Human annotation analysis shows that expert agreement varies strongly with label granularity, motivating hierarchical structured metrics and claim-based evaluation of open-ended feedback. Experiments with 12 recent audio-language models reveal a consistent gap: while models can compare performances and identify broad issue domains in free-form feedback, Top-3 fine-grained issue-label identification remains below label-prior baselines and strict diagnosis alignment stays below 7%. To our knowledge, VocalCoachBench pro- vides the first public testbed for evaluating audio-grounded expert feedback for singing, moving audio-language evaluation beyond description toward analytic feedback.


10. Knowing When Not to Answer: Pseudo-Ensembles for Abstention in Music Audio-Language Models

链接:https://arxiv.org/abs/2609.04362

机构:Jumeirah College(朱美拉学院); Apta AI(阿普塔人工智能公司); Spark AI Research(火花人工智能研究院)

作者:Aanya Maheshwari, Vatsal Raina

英文摘要:Music audio-language models are evaluated almost entirely by accuracy on multiple-choice questions. This protocol forces the model to commit to an option, so a lucky guess looks the same as real musical understanding. What is missing is a way to tell when the model does not know the answer, so that it can abstain instead of guessing. The usual solution, an ensemble of independently trained models, is far too expensive here, which leaves the entropy of a single predictive distribution as the only available confidence signal. We instead build pseudo-ensembles from one pretrained model by perturbing its input in ways that cannot change the correct answer, then averaging the resulting distributions over the options. Our main construction simply shuffles the order in which the candidate answers are presented; we also study ensembles built from corrupted audio and from swapped option labels. A pseudo-ensemble gives several predictive distributions per question, so it supports the full family of ensemble-based uncertainty measures (entropy of the expected distribution, expected entropy, and their difference, the mutual information) rather than entropy alone. Evaluating TinyMU on MuChoMusic, we find that averaging over four option orderings raises accuracy from 55.7% to 59.2%, and that the resulting uncertainty measures rank the model's errors better than the single-pass entropy baseline, reducing the area under the error retention curve from 0.293 to 0.261. All of this costs a few extra forward passes and no retraining, which makes abstention practical for compact music audio-language models.


11. Low-Latency Spell Correction for Japanese Music Search Queries

面向日语音乐搜索查询的低延迟拼写纠错

AI 总结:本研究针对日语音乐搜索查询的拼写纠错挑战,提出基于BART的紧凑序列到序列模型,结合感知脚本的合成错误生成流程,在低延迟下实现优于基线的拼写纠错性能。

链接:https://arxiv.org/abs/2609.04262

机构:Amazon(亚马逊公司)

作者:Anshul Garg, Pavni Tandon, Karan Bhukar, Tanmay Khandelwal, Ujjal Kumar Dutta

英文摘要:Spell correction for Japanese search queries presents unique challenges due to the co-existence of four writing scripts (Latin/romaji, hiragana, katakana, and kanji) and the distinct error patterns each script induces. We present a compact BART-based sequence-to-sequence model (3 encoder + 3 decoder layers) designed for low-latency spell correction of Japanese music search queries. The core contribution lies in a script-aware synthetic misspelling generation pipeline that produces realistic training data by combining keyboard-layout models (QWERTY and flick input), phonetic confusion priors mined from real query logs, voiced/unvoiced consonant alternations, and kana case errors. A key design decision is normalizing mixed-script catalog titles to a single canonical script before misspelling synthesis, which we show is critical for reducing model hallucinations. We train a custom byte-level BPE tokenizer on the target music catalog to handle all four scripts in a unified vocabulary. Experiments on a curated evaluation set show that our model achieves an exact-match accuracy of 41.09% and a character error rate (CER) of 11.62%, outperforming edit-distance baselines and achieving the lowest character error rate among all evaluated systems while maintaining sub-4ms inference latency on a single GPU. We further analyze performance across individual scripts and mixed-script queries, demonstrating the effectiveness of script-aware data augmentation through systematic ablation studies.


12. Scalable Context Orchestration for Serving LLMs Over Voice

链接:https://arxiv.org/abs/2609.04288

机构:Global College, Shanghai Jiao Tong University(上海交通大学Global College); AgenticSys

作者:Linyi Jiang, Silvery D. Fu, Yifei Zhu

英文摘要:Voice AI applications are gaining popularity as advances in large language models (LLMs) enable more natural and accessible spoken interactions. Serving these applications requires accounting not only for what users say, but also for how they speak (e.g., speaking rate) and the conditions under which their audio is captured and transmitted (e.g., background noise and packet loss). However, existing LLM systems represent conversation context as a flat, growing sequence of messages, leaving voice-specific context implicit in the audio. As a result, they can generate responses that are poorly aligned with user preferences, degrade interaction quality under adverse environmental conditions, and incur high costs over long voice sessions. We present llmovoice, a context-management middleware that explicitly models voice context and orchestrates its use. At each turn, llmovoice constructs a bounded voice context from the current user input, relevant interaction history, and explicit paralinguistic and environmental states. It then uses the serving LLM to reason over this context and generate runtime directives that guide how the system responds. We evaluate llmovoice on real-world voice applications and benchmarks. It reduces speaking-rate alignment error by 52.4%, lowers the false-interruption rate from 46.0% to 0.9% under packet loss, and reduces model usage cost by 79.2%. For long sessions, llmovoice reduces per-turn cost by up to 24.9 times while retaining up to 98.7% of baseline answer quality.


13. SwanWeave:One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing

链接:https://arxiv.org/abs/2609.04975

机构:Zhejiang University(浙江大学); ByteDance(字节跳动)

作者:Ke Lei, Chenyuhao Wen, Yu Zhang, Wenxiang Guo, Changhao Pan, Sashuai Zhou, Yongshi Li, Ruiqi Li, Ruofan Hu, Haorui Xu, Xiang Yin, Zhou Zhao

英文摘要:Spatial audio editing modifies an existing soundfield according to a user's instruction while preserving the rest of the scene. Unlike conventional audio editing, it must reason jointly about audio events, spatial information, dynamic changes, and environmental information in first-order Ambisonic (FOA) waveforms. Existing language-guided editors mainly target conventional audio or rely on sequential operations, and therefore do not directly support one-stage editing for complex 3D spatial instructions. We present SwanWeave, the first one-stage multi-task framework for instruction-guided 3D FOA spatial audio editing. We build paired FOA supervision from open-source speech and sound-effect corpora using controllable room simulation, covering more than ten single-operation and compound tasks across the four editing axes. To handle this heterogeneous edit space, SwanWeave uses Spatial Edit Mixture-of-Experts (SE-MoE) with dual-level routing, selecting task-aware expert combinations for compound instructions and frame-level routed/null experts for local edit decisions. We further introduce Spatial Preference Optimization (SPO), a Direct Preference Optimization (DPO)-based alignment objective with edit-specific negative targets, and adopt staged training to improve natural-language grounding. Experiments show that SwanWeave achieves better editing quality than existing general audio editors and spatial audio baselines across all tasks. Spatial audio editing demos can be found at this https URL, code can be found at: this https URL.