今日论文合集:CS.SD语音与音频 | 共 9 篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准


快速导航

1. 语音合成与声音生成 2 篇

2. 说话人识别、验证与分离 1 篇

3. 语音增强、降噪与音频修复 1 篇

4. 音乐信息检索与音乐生成 2 篇

5. 其他/综合语音音频 3 篇


1. 语音合成与声音生成 | 2 篇

1. PhaseGAN: High-Fidelity Vocoder via Decoupled Amplitude and GAN-Driven Phase Reconstruction

PhaseGAN:通过解耦幅度和GAN驱动的相位重建实现高保真声码器

AI 总结:PhaseGAN通过解耦幅度与相位重建的轻量级声码器,以少量参数实现高保真音频合成,并展现跨领域泛化能力。

链接:https://arxiv.org/abs/2609.12918

机构:Inner Mongolia University(内蒙古大学)

作者:Wenzheng Zhang, Xueliang Zhang, Shulin He, Fei Zhao, Xin Liu, Pengjie Shen, Zhenlong Guo, Zixuan Xue, Hongtao Bao, Zixuan Li

英文摘要:A vocoder is a pivotal component of modern text-to-speech (TTS) systems. Despite the significant progress of neural network-based vocoders, accurate phase reconstruction remains the main challenge limiting both audio quality and modeling efficiency. We introduce PhaseGAN, a lightweight vocoder that addresses this limitation through a "mel $\rightarrow$ Amplitude $\rightarrow$ Phase" reconstruction pipeline. By reconstructing amplitude and phase spectra via distinct methodologies, the proposed PhaseGAN outperforms state-of-the-art baselines while utilizing fewer model parameters and reduced computational requirements. The compact version generates high-fidelity audio with approximately 500K parameters and 1 GMAC computational load, making it highly suitable for real-time applications on edge devices. In addition, our approach exhibits exceptional musical audio synthesis capabilities despite no training on musical data, illustrating unprecedented cross-domain generalization. See this https URL for demos of our work.


2. StepAudio 3 Gen Technical Report

StepAudio 3 Gen 技术报告

AI 总结:StepAudio 3 Gen 是一个基于离散自回归和 RVQ 标记的通用音频生成模型,通过渐进预训练、RVQ 适配器等设计,在 TTS 和声音设计上达到最先进性能,并支持多种音频类型生成。

链接:https://arxiv.org/abs/2609.12945

作者:Bin Lin, Bo Zhao, Boyang Wang, Boyang Zhang, Boyong Wu, Chao Yan, Chen Geng, Chen Wu, Cheng Yi, Chengli Feng, Chenglin Zhu, DanNi Wan, Daxin Jiang, Dongqing Pang, Fei Tian, Feng Tian, Future Li, Gang Yu, Guanglong Yang, Jia Peng, Jiahao Song, Jiamin Fan, Jiangjie Zhen, Jianzheng Gao, Jun Chen, Li Xie, Lifang Zhang, Lingli Ji, Liying Shi, Lun Cai, Min Xu, Na Wang, Peilin Li, Peng Yang, Pengfei Tan, Qingjian Lin, Ruijie Xiong, Runze Li, Shenghua Hu, Shi Qiu, Siqi Tu, Siyi Zhou, Tianjiao Deng, Wanying Lu, Weiming Niu, Wen Sun, WenWen Qu, Xiangyu Zhang, Xianwei Zhang, XiaoSu Su, Xing Chen, Xinyu Liu, Xuerui Yang, Yang Li, Yang Yang, Yechang Huang, Yibo Zhu, Yifan Zhang, Yiyang Xu, Yu Fu, Yu Luo, Yu Zhou, Yumang Wang, Yunzhou Ju, Yuxiang Yang, Zekai Liu, Zengwei Yao, Zhenwei Mou, Zheqi Dai, Zhiyue Wu, Zichao Zhou

英文摘要:We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared $16 \times 2048$ residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: (1) interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, (2) RVQ Adaptor for effectively incorporating multi-codebook acoustic representations, and (3) discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio samples are available at this https URL.


2. 说话人识别、验证与分离 | 1 篇

3. Neural Multichannel Distant Speaker Diarization with Heavy-tailed Source Separation Model

神经多通道远场说话人日志与重尾源分离模型

AI 总结:本文提出一种结合重尾分布(尖峰广义高斯和学生t分布)的神经FCASA模型,用于远场多通道说话人日志,在DER和JER指标上较基线取得大幅改进。

链接:https://arxiv.org/abs/2609.12154

机构:LTCI, Telecom Paris, Institut Polytechnique Paris(巴黎综合理工学院电信巴黎分校LTCI); SJTU Paris Elite Institute of Technology, Shanghai Jiao Tong University(上海交通大学巴黎卓越工程师学院); LIUM, Le Mans University(勒芒大学LIUM)

作者:Sicheng Mao, Baihan Li, Mathieu Fontaine, Anthony Larcher, Roland Badeau

英文摘要:Distant speaker diarization remains challenging due to difficult acoustic environments, varying numbers of speakers and overlapping speech. Model-driven methods are proposed to exploit the speech source features in multi-channel recordings that help diarization. This paper generalizes a neural model that jointly learns to perform blind source separation and diarization over speech mixtures (neural FCASA) with heavy-tailed models. The popular Gaussian distribution has been applied for variance modeling in the original source separation model, which we replace with two families of heavy-tailed models (Leptokurtic Generalized Gaussian distribution and Student's t distribution) to better capture the heavy-tailedness in speech signals. Thanks to the Gaussian scale mixture model, we are able to unify the proposed method and the original one under the same form of learning objective. Our experiments show consistent large improvements in Diarization Error Rate (DER) and Jaccard Error Rate (JER) compared to the baseline on various corpora.


3. 语音增强、降噪与音频修复 | 1 篇

4. DriftSE: Speech Enhancement with Generative Drifting

DriftSE:基于生成漂移的语音增强

AI 总结:DriftSE提出一种基于生成漂移的单步语音增强框架,通过双潜在并行漂移同时保持语音可懂度和声学保真度,实现非配对训练,并在四个数据集上以1 NFE达到最先进的词错误率。

链接:https://arxiv.org/abs/2609.12252

机构:Victoria University of Wellington(惠灵顿维多利亚大学); GN Advanced Science(GN 前沿科学); Lincoln University(林肯大学)

作者:Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson

英文摘要:We propose DriftSE, a novel one-step generative framework for speech enhancement formulated as a latent distribution equilibrium problem. During training, the drifting field aligns the generator's pushforward distribution with the clean speech manifold through drifting in a latent domain. During inference, the drifting process is discarded, enabling one-step generation. We establish that its enhancement quality depends fundamentally on the choice of latent representation. Semantic latents preserve phonetic structure but fail to capture physical acoustic cues, whereas acoustic latents reconstruct the physical signal but risk linguistic hallucination. Therefore, we introduce dual-latent drifting, performing parallel drifting in both semantic and acoustic latents to simultaneously preserve phonetic intelligibility and acoustic fidelity. Additionally, we demonstrate that DriftSE enables fully unpaired training by aligning latent distributions rather than exact point-wise targets. Consequently, DriftSE facilitates cross-dataset learning in the absence of paired noisy-clean samples. Moreover, DriftSE exhibits broad architectural flexibility across different generator backbones. Extensive evaluations on additive denoising and convolutive dereverberation demonstrate robust one-step enhancement across both offline and real-time causal settings. Notably, DriftSE achieves state-of-the-art word error rates across all four evaluated datasets while strictly operating at 1 NFE. Code and audio examples are available online.


4. 音乐信息检索与音乐生成 | 2 篇

5. Real-Time Music Source Separation on a Low-Power Audio DSP

低功耗音频DSP上的实时音乐源分离

AI 总结:本文针对低功耗音频DSP,验证现有实时音乐源分离系统均不适用,提出在连续卷积上下文训练并加入门控复数FIR深度滤波器的新系统,在MUSDB18-HQ上达到4.70 dB cSDR,运行时间仅10.43 ms。

链接:https://arxiv.org/abs/2609.12201

机构:Analog Devices, Inc.(亚德诺半导体公司)

作者:Jianan Li, Li Liu, Ken Malsky, Gabby Yi

英文摘要:Real-time music source separation is validated on desktop CPUs and GPUs. Does any published system fit the embedded audio hardware it targets? On a commercial audio DSP (2 MB SRAM, 2.07 GMAC/s measured), none does, and the constraints eliminate different models: memory rules out the 16-51 M parameter TasNet/X-UMX family, per-frame compute rules out RT-STT, needing 5.5x the available MAC rate. Parameter count predicts neither: weight reuse spans 1x to 345x. We then build one that fits. Training on continuous rather than block-padded convolution context proves essential: a model scoring 3.93 dB block-wise otherwise collapses to silence within 2 s frame-by-frame. A gated complex FIR deep filter adds a latency knob, gaining 0.38 dB even when strictly causal. It reaches 4.70 dB cSDR on MUSDB18-HQ and runs in 10.43 ms of an 11.6 ms hop, 0.5-0.7 dB behind systems that do not fit.


6. DiffSynth-Music: Audio-Conditioned KV-Cache Adapters for Controllable Music Generation

DiffSynth-Music:用于可控音乐生成的音频条件KV缓存适配器

AI 总结:针对文本歌词对音乐节奏、旋律控制不足的问题,提出DiffSynth-Music框架,通过音频条件KV缓存注入支持五种控制类型,提升可控性与歌词保真度。

链接:https://arxiv.org/abs/2609.12774

机构:Alibaba Group(阿里巴巴集团); Shanghai Jiao Tong University(上海交通大学)

作者:Zhongjie Duan, Shengchuan Gao, Hong Zhang, Yingda Chen

英文摘要:Text and lyrics specify broad musical characteristics and sung content but offer limited control over musical timing, melody, and reference-based style. We introduce DiffSynth-Music ( this https URL ), a framework that adds composable audio conditioning to a music synthesis backbone through layer-wise key-value injection. The three template models, Control, Prosody, and Reference, are initialized from the backbone diffusion transformer and trained with conditional flow matching. They support five control types: beats, vocals, accompaniment, prosody, and reference audio. A shared variational autoencoder maps conditioning waveforms into a common latent space, enabling their attention memories to be combined. With the template timestep fixed at the clean-data endpoint and other inputs held constant, each control cache is computed once and reused throughout sampling. Training pairs are derived from music recordings using beat extraction, source separation, vocal resynthesis, and reference-excerpt selection. Single-control evaluations on Mandarin and English songs demonstrate improved adherence across all five control types and better lyric fidelity under vocal conditioning relative to the backbone. Automatic music-quality and instruction-following scores remain broadly comparable to those of the evaluated base models, with metric-specific trade-offs. We release the three template models to support research and creative applications in controllable music generation.


5. 其他/综合语音音频 | 3 篇

7. Unlabeled Echoes: Pseudo-Labels and Genus-Aware Smoothing for Bat Call Recognition

未标记的回声:用于蝙蝠叫声识别的伪标签与属感知平滑

AI 总结:本研究提出利用模型生成的伪标签进行半监督学习,在欧洲18物种语料库及南非野外音频上验证其有效性,并引入属感知平滑提升识别性能,显著缩小与全监督的差距。

链接:https://arxiv.org/abs/2609.11986

机构:Ludwig-Maximilians-Universität München(慕尼黑大学); University of the Free State(自由州大学)

作者:Frank Fundel, Alexandra Howard

英文摘要:Passive acoustic monitoring produces far more bat recordings than experts can label. We show that simple model-generated pseudo-labels turn this surplus into effective supervision. We compare pseudo-labeling with other semi-supervised learning methods on an 18-species European corpus using only 10% of its training labels, then transfer the strongest approaches to South African field audio containing nine bat taxa and a nuisance class. Pseudo-labeling outperforms the other semi-supervised learning methods on every European measure, recovering up to 61.5% of the gap to full supervision. It transfers to field audio with gains of 10.69 points in species accuracy and 4.96 points in species macro-F1. We also introduce genus-aware smoothing, which directs uncertain target mass toward congeneric species. Combined with uniform smoothing, it reaches 79.16 species macro-F1, 4.73 points above hard targets. Simple pseudo-labels are therefore highly effective at this ecological data scale, while genus-aware targets inject useful biological structure at no annotation cost. this https URL


8. Direct Preference Density Alignment for Conversational Audio Equalization

直接偏好密度对齐用于对话音频均衡

AI 总结:提出直接偏好密度对齐框架,无需代理奖励模型,结合GRPO在线探索与DPO离线细化,在音频均衡任务中以1.5B参数模型达到GPT-4o mini的感知水平。

链接:https://arxiv.org/abs/2609.12607

机构:Aalborg University(奥尔堡大学); Bang & Olufsen A/S(Bang & Olufsen公司)

作者:Ioannis Stylianou, Sven Ewan Shepstone, Jon Francombe, Pablo Martinez Nuevo, Zheng-Hua Tan

英文摘要:Large Language Model alignment typically relies on learned proxy reward models, which significantly increase the memory footprint during training and are notoriously prone to instability and reward hacking. While offline methods like Direct Preference Optimization (DPO) bypass the reward model, they lose the ability to perform online exploration. If no optimization constraints are applied, this can lead to format collapse in bounded, continuous spaces. To resolve this, we propose Direct Preference Density Alignment: An alternative framework that removes the need for a learned proxy reward model while strictly preserving the benefits of online reinforcement learning. We leverage large-scale user data (approximately 90,000 samples) to construct non-parametric preference density maps, establishing an empirical reward surface. In addition to removing the reward model, Direct Preference Density Alignment enables the combination of the online structural grounding of Group Relative Policy Optimization (GRPO) with the targeted offline refinement of DPO. We show that this GRPO+DPO combination achieves the highest performance, and in a blind audio equalization listening test, enables a 1.5B-parameter model to achieve perceptual parity with a carefully prompt-engineered GPT-4o mini baseline, using only a fraction of the inference compute.


9. What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability

MLLM 听到了什么?面向音频 MLLM 可解释性的词元级时频定位

AI 总结:针对音频 MLLM 可解释性,提出首个词元级时频定位框架 STAG,通过词汇投影与频谱遮挡生成相关性图,在四个基准上取得最优事件定位性能,并验证了解释的忠实性与选择性。

链接:https://arxiv.org/abs/2609.12663

机构:University of Salerno(萨莱诺大学)

作者:Lucia Cascone, Valeria Fraenza, Michele Nappi, Fabio Narducci, Benedetto Simone

英文摘要:Audio-based Multimodal Large Language Models (MLLMs) can generate detailed natural-language descriptions of complex acoustic scenes, yet it remains unclear which parts of the input audio support each generated token. This is particularly challenging because acoustic evidence is distributed across time and frequency, and concurrent sound events may overlap temporally while occupying different spectral regions. We introduce STAG, to our knowledge the first post-hoc framework for token-level spectro-temporal grounding of captions generated by audio-based MLLMs. STAG estimates the temporal support for each generated token using target-token-specific vocabulary projections of the encoded audio representations, measures frequency-band relevance through controlled spectral occlusion, and combines the two signals into a spectro-temporal relevance map. We evaluate STAG against ten post-hoc explanation methods across four grounding benchmarks, where it achieves the best event-localization performance on every dataset, and apply it to eight audio-language backbones without parameter updates. Counterfactual deletion further shows that removing the identified evidence selectively reduces confidence in the corresponding event and frequently removes it from the regenerated caption. These results provide behavioral support for the faithfulness and selectivity of the explanations.