今日论文合集:CS.SD语音与音频 | 共 12 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
1. Compressing Streaming Neural Audio Encoders via Latent-Space Distillation
通过潜在空间蒸馏压缩流式神经音频编码器
AI 总结:本研究提出通过潜在空间蒸馏压缩流式神经音频分词器,以预量化潜在为监督目标,在2.8倍压缩下保持低WER偏差,性能优于同等规模独立训练的分词器。
链接:https://arxiv.org/abs/2609.04102
机构:Apple(苹果公司)
作者:Prasanth Yadla, Mohammad Samragh, Dongseong Hwang, Mingbin Xu, Yuanyuan Zhang, Chung-Cheng Chiu, Yongqiang Wang, Yuan Liu, Zhen Huang, Xiaodan Zhuang
英文摘要:System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share. We train only the student encoder to regress the teacher's per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher-student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces we support, and applies both to a tokenizer pretrained alone and to one jointly trained with a language model. At 2.8x compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher-student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.
2. VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis
VoxReason:合成前基于源的语音规划的无听者评估
AI 总结:本研究提出VoxReason,用于在语音合成前无听者评估基于源的语音规划,通过确定性验证器及实验验证其能有效检测源使用失败,7B模型的SFT+CF修复可提升规划性能。
链接:https://arxiv.org/abs/2609.03203
机构:National Research Council Canada(加拿大国家研究委员会)
作者:Mengzhe Geng
英文摘要:Expressive speech systems make a decision before any waveform is rendered: how an utterance is delivered. In dialogue agents, narration, and role-conditioned TTS, that hidden planning step sets affect, pitch, energy, rate, pause, emphasis, and stance, yet downstream audio scores rarely reveal whether those choices were licensed by the source record, a source-use failure that occurs before any waveform exists. VoxReason makes that pre-synthesis decision measurable as a listener-free task for source-grounded speech planning. Before synthesis, VoxReason measures whether delivery choices are grounded in cited source records. Systems output a source-cited speaking-plan with evidence citations, and a deterministic verifier checks citation legality, slot agreement, unsupported state, schema validity, and one-cue counterfactual locality. On 1,440 checked source-label cases, shortcut controls show why slot accuracy alone is unsafe: a key-lookup oracle reaches 1.000 plan-slot accuracy on seen keys, while an emotion prior still reaches 0.958 slot accuracy on source-key-disjoint cases without citing intensity or identity. In a separate 100-case learned source-key-disjoint comparison, a 7B locality SFT+CF repair improves plan-slot accuracy/locality from 0.684/0.141 to 0.919/1.000, and removing source records lowers citation-required grounded score by 0.488. Rendered waveform quality remains outside the present evaluation.
3. Is Semantics Enough for Speech Mean Opinion Score Prediction?
语义是否足以用于语音平均意见得分预测?
AI 总结:该研究针对自动语音MOS预测器过度依赖SSL模型忽略声学细节的问题,对比三类模型在BVCC及OOD数据集的表现,发现结合语义与细粒度声学建模的模型性能更优,指出需同时关注语义与声学保真度。
链接:https://arxiv.org/abs/2609.03283
机构:National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China(中国科学技术大学国家语音与语言信息处理工程研究中心)
作者:Tianyu Lan, Yufei Shi, Yang Ai, Honghao Sun, Huipeng Du, Zhenhua Ling
英文摘要:Mean Opinion Score (MOS) is the gold standard for evaluating synthesized speech naturalness. However, current automatic MOS predictors are dominated by self-supervised learning (SSL) models that prioritize high-level semantics, potentially compromising their ability to capture critical acoustic details. In this paper, we systematically investigate representations from three paradigms: SSLs, acoustic-only neural audio codecs (NACs), and unified NACs that integrate semantics into reconstruction-based architectures. Extensive benchmarking on the standard BVCC and multiple out-of-domain (OOD) datasets demonstrates that features synergizing semantic understanding with fine-grained acoustic modeling achieve a higher performance upper bound in speech quality assessment. Ultimately, our findings highlight that semantics alone are not enough; a dual focus on semantic content and acoustic fidelity is essential for robust MOS prediction.
4. PACodec: A Low-bitrate Neural Speech Codec with Parallel Additive Vector Quantization
PACodec:一种采用并行加性矢量量化的低码率神经语音编解码器
AI 总结:本文提出基于并行加性矢量量化(PAVQ)的PACodec,采用“全局-局部-全局”设计,在相同解码质量下比基线降低30%码率,且具备解纠缠友好性,可应用于语音转换等下游任务。
链接:https://arxiv.org/abs/2609.03363
机构:National Engineering Research Center of Speech and Language Information Processing(国家语言与语音信息处理工程研究中心); University of Science and Technology of China(中国科学技术大学)
作者:Fei Liu, Yang Ai, Xiao-Hang Jiang, Zhen-Hua Ling
英文摘要:This paper proposes PACodec, a novel low-bitrate neural speech codec based on parallel additive vector quantization (PAVQ). Unlike the mainstream residual vector quantization (RVQ) used in most neural speech codecs, where vector quantizers (VQs) are sequentially dependent, the PAVQ strategy adopted in PACodec aggregates parallel quantization results to optimize bitrate usage. Specifically, the PAVQ adopts a "global-local-global" (GLG) design: the global encoded features are quantized in parallel by multiple independent VQs, each attending to a local component of the representation, and their outputs are aggregated through addition to yield the final global quantization result for decoding. Experimental results show that PACodec, as each VQ focuses only on local information, supports smaller codebooks and reduces bitrate by 30% compared with baselines at the same decoding quality, with only minor model complexity. Further analysis shows that, owing to the GLG framework of PAVQ, the proposed PACodec is disentanglement-friendly, and each independent VQ captures different aspects of speech, e.g., content, timbre, and acoustic details, suggesting potential for application to downstream tasks such as voice conversion.
5. Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations
基于连续神经音频编解码器表示的掩码自回归语音增强
AI 总结:该研究提出MARSE方法,利用连续NAC表示迭代解码掩码纯净语音帧,可灵活权衡语音增强性能与计算成本。
链接:https://arxiv.org/abs/2609.03940
机构:CentraleSupélec(中央高等电力学院); Univ. Grenoble Alpes(格勒诺布尔阿尔卑斯大学); CNRS(法国国家科学研究中心); Grenoble-INP(格勒诺布尔综合理工学院); GIPSA-lab(吉普萨实验室)
作者:Yoto Fujita, Simon Leglaive, Laurent Girin
英文摘要:Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs). However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility. In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech. In particular, we investigate a set of different decoding policies, ceteris paribus, that is, using the same DNN (a Conformer model), the same NAC (the DAC codec) and the same training setup. The results show that MARSE enables a flexible trade-off between SE performance and computational cost. Audio examples and code are available online.
6. StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios
StrixAE:面向真实场景中复杂失真耦合下音频增强的智能体
AI 总结:针对真实场景音频增强的复杂失真耦合与个性化需求难题,提出基于MLLM的智能体StrixAE,经两阶段训练后性能优于多数现有方案,实现了更优的音频增强效果与鲁棒性。
链接:https://arxiv.org/abs/2609.03414
机构:Xiamen University(厦门大学); Fuyao University of Science and Technology(福耀科技大学)
作者:Chenglin Wu, Junjie Wu, Jinhang Chen, Mingyang Chen, Zixu Lin, Jiabian Chen, Xinghao Ding, Xiaotong Tu
英文摘要:Audio enhancement in real-world scenarios involves complex distortion couplings and requires personalized enhancement. Existing solutions struggle to address both simultaneously. To improve robustness and enable autonomous operation in such scenarios, we propose StrixAE, an agent based on a multimodal large language model (MLLM). StrixAE leverages the MLLM as a controller to coordinate multiple audio enhancement and personalization models. To further enhance system robustness, reduce artifacts, and improve generalization across diverse real-world scenarios, StrixAE is trained through a two-stage process: first, CoT supervised fine-tuning on AcoustBench to ground basic reasoning and tool invocation; second, Audio Perception Reinforcement Learning (APRL), a reward design specifically tailored for audio restoration pipelines that jointly optimizes format validity, structural coherence, and perceptual quality. Unlike generic RL fine-tuning, APRL introduces structured rewards that enforce executable pipelines and logical section ordering, enabling the agent to produce reliable, interpretable enhancement plans without hallucinated tools. Based on real-world test datasets, our proposed method outperforms most existing open-source and proprietary solutions, achieving state-of-the-art performance across multiple perceptual metrics and demonstrating strong generalization robustness.
7. Test-time adaptation for speech enhancement with an autoregressive speech prior
基于自回归语音先验的语音增强测试时自适应
AI 总结:该研究提出一种基于自回归语音先验的单utterance测试时自适应方法,通过最小化分布散度提升语音增强模型在噪声不匹配场景下的性能,相关成果已在线公开。
链接:https://arxiv.org/abs/2609.03622
机构:CentraleSupélec(中央超级电子学院); Inria(法国国家信息与自动化研究所); University of Hamburg(汉堡大学)
作者:Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda, Timo Gerkmann
英文摘要:Test-time adaptation (TTA) offers a promising direction for improving speech enhancement models under mismatched acoustic conditions, without requiring access to labeled target data. In this work, we propose a single-utterance TTA method that regularizes a pretrained speech enhancement model using an autoregressive prior trained on clean speech latent representations extracted from a neural audio codec. Adaptation is performed by minimizing the Kullback-Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments across multiple noisy speech datasets show consistent improvements in speech quality, particularly under training-testing noise mismatch conditions. Code and audio examples are available online.
8. SISER: Speaker-Invariant Speech Emotion Recognition with Entropy-Based Adversarial Training
SISER:基于熵对抗训练的说话人无关语音情感识别
AI 总结:本文针对语音情感识别的标注数据稀缺与说话人差异性问题,提出基于熵对抗训练的SISER方法,采用wav2vec 2.0与ECAPA-TDNN,在IEMOCAP数据集上取得优于基线及无说话人抑制的wav2vec 2.0的性能。
链接:https://arxiv.org/abs/2609.02941
作者:Eunseo Choi, Hyunku Kang, Chanwoo Kim
英文摘要:Speech emotion recognition (SER) faces two fundamental challenges: scarcity of labeled data and inter-speaker variability, both of which hinder generalization of emotion recognition systems. While prior adversarial approaches address speaker variability, they fall short in leveraging powerful pre-trained representations. We propose SISER (Speaker-Invariant Speech Emotion Recognition), integrating wav2vec 2.0 as a feature encoder and ECAPA-TDNN as a speaker discriminator within an entropy-based adversarial training scheme. wav2vec 2.0 provides rich self-supervised representations that alleviate dependency on large labeled datasets, while ECAPA-TDNN enables suppression of speaker identity via a stronger adversarial signal than shallow classifiers. Evaluated on IEMOCAP, SISER achieves a UA of 60.63%, outperforming the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%), with ablation emphasizing that the choice of speaker classifier architecture is a key factor.
9. Local Chord Corruption Is Not Recognizer Replay: Chord-Condition Propagation in MIDI-SAG
局部和弦损坏并非识别器重放:MIDI-SAG中的和弦条件传播
AI 总结:该研究对比合成和弦损坏与ACR重放对MIDI-SAG的影响,发现局部损坏测试机制、完整重放评估和弦条件传播,匹配重放支持可减少两者不匹配。
链接:https://arxiv.org/abs/2609.03584
机构:Shenzhen University(深圳大学)
作者:Weiwen Huang
英文摘要:Synthetic chord corruption provides a controlled stress test for singing accompaniment generation (SAG), whereas complete automatic chord recognition (ACR) replay measures the condition delivered to a deployed system. We compare them by replaying CNN-CRF and DeepChroma+CRF predictions through one fixed MIDI-SAG generator, holding track, seed, context, and scoring window constant. Across 30 paired tracks and three seeds, a central four-second tritone produced larger changed-target and inside-output effects than CNN-CRF replay in STFT, CQT, and CENS; the CENS target gap was positive on 29/30 tracks (mean 0.462). Matching replay support and relation composition reduced this mismatch, with joint matching giving the lowest replay distance in the full-30 analysis. Relative-root substitutions produced 2.88-fold larger CENS full-window output change than same-root quality flips at near-equal dose. The matched surrogate was closer to replay for both recognizer paths. We conclude that localized corruption tests mechanisms, whereas complete replay evaluates deployed chord-condition propagation.
10. CRAW: Codec Robust Audio Watermarking
CRAW:编解码器鲁棒音频水印
AI 总结:针对现有音频水印方法在实际变换下鲁棒性不足的问题,提出CRAW框架,结合多种技术提升对神经编解码器等的鲁棒性,同时保持感知质量,达最优性能。
链接:https://arxiv.org/abs/2609.03107
机构:Bar-Ilan University(巴伊兰大学); NVIDIA(英伟达公司)
作者:David Chernin, Ethan Fetaya
英文摘要:Recent advances in generative speech models have made it increasingly difficult to distinguish authentic from synthetic audio, enabling new forms of fraud and misinformation. Audio watermarking offers a promising defense by embedding an imperceptible signal into generated speech that can later be detected to verify its provenance. However, recent studies have shown that existing post-hoc watermarking methods fail under neural codecs and denoisers, transformations routinely applied during real-world storage, transmission, and processing, severely limiting their practical utility. Here we introduce CRAW, a codec-robust audio watermarking framework that jointly improves robustness against neural re-synthesis while maintaining high perceptual quality. CRAW combines distortion-aware training with an attention-based pooling mechanism, inference-time perceptual mask- ing, and an error-correcting code to recover the fidelity lost during robust training. Experiments demonstrate that CRAW achieves state-of-the-art robustness against neural codecs, denoisers, and vocoders while maintaining perceptual quality comparable to existing post-hoc watermarking methods. The code is available at this https URL.
11. Beyond.WAV: Design and Software Verification of VocalCap, a Traceable Browser-Based Audio Capture System for Vocal Biomarker Research
Beyond.WAV:面向语音生物标志物研究的可追溯基于浏览器的音频捕获系统VocalCap的设计与软件验证
AI 总结:该研究设计并验证了面向语音生物标志物研究的浏览器音频捕获系统VocalCap,其可追溯捕获过程且经测试验证了软件行为,为远程语音研究提供可靠音频数据支持。
链接:https://arxiv.org/abs/2609.03320
机构:Institute of Mathematics and Statistics, University of São Paulo(圣保罗大学数学与统计学院); bluecore(蓝芯科技)
作者:Augusto Camargo
英文摘要:Remote voice studies often retain a final audio file with limited evidence about how it was captured, transferred, processed, and accepted. This paper presents VocalCap, an institution-controlled, browser-based system for self-guided capture of voice and related acoustic signals by participants without technical training. A versioned protocol drives the workflow. Each accepted recording retains a browser-native object, a client-lossless Float32 WAV derived from the same MediaStream, and a server-canonical mono PCM16 WAV, linked to evidence of capture execution, technical quality, byte-level integrity, recovery, and transformation provenance. IndexedDB preserves accepted browser artifacts until server confirmation, while session completion requires successful verification of every task and artifact. Software tests challenged the acquisition contracts with malformed or altered objects, exact-zero interruptions, channel-topology variants, and interrupted or repeated operations. A post hoc technical audit of 39 consented pilot recordings found 25 sample-identical stereo files and 14 files with signal confined to the left channel. Topology-aware active-channel selection limited the canonical root-mean-square level difference to less than 0.001 dB in all 14 affected files; equal-weight stereo averaging would have introduced approximately 6.02 dB of attenuation. Production end-to-end verification completed two five-task profiles in Chromium and WebKit, yielding 10 accepted recordings and 30 retained artifacts that passed server-side integrity and format checks. The results verify VocalCap's software behavior under the tested browser-engine conditions. Device-level acoustic agreement, target-population usability, clinical validity, and biomarker performance remain subjects for separate studies.
12. Neural Music Enhancement with Dual Time-Frequency Spectral Representations for Prediction and Discrimination
基于双时频谱表示的神经音乐增强方法(用于预测与判别)
AI 总结:针对非专业音乐录音的噪声混响问题,提出基于双时频谱的DSME模型,采用STFT生成与CQT判别,引入色度谱损失,实验验证其增强效果优于基线。
链接:https://arxiv.org/abs/2609.03357
机构:National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China(中国科学技术大学国家语音及语言信息处理工程研究中心)
作者:Fei Liu, Yang Ai, Zhen-Hua Ling
英文摘要:Non-professional music recordings shared online often suffer from background noise and reverberation, degrading perceived quality and limiting reuse. This paper proposes DSME, a music enhancement model based on dual time-frequency spectral representations. Within a generative adversarial framework, DSME uses short-time Fourier transform (STFT) spectra for generation and constant-Q transform (CQT) spectra for discrimination. Leveraging STFT's fixed window, invertibility, and predictability, the generator estimates clean amplitude-phase spectra from degraded inputs and reconstructs waveforms via inverse STFT. Exploiting CQT's log-frequency, variable-window structure aligned with musical octaves, we design an octave-segmented CQT discriminator. We also introduce a chroma-spectrum loss to emphasize pitch and harmonic consistency. Experiments show DSME outperforms baselines in objective and subjective tests, validating the effectiveness of the dual-spectrum approach.
