今日论文合集:CS.SD语音与音频 | 共 11 篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准


快速导航

1. 语音合成与声音生成 3 篇

2. 语音增强、降噪与音频修复 1 篇

3. 音频事件检测与场景理解 1 篇

4. 音乐信息检索与音乐生成 1 篇

5. 语音翻译与语音语言模型 1 篇

6. 安全、隐私与深度伪造音频 3 篇

7. 其他/综合语音音频 1 篇


1. 语音合成与声音生成 | 3 篇

1. Narrowband Voice Communication Using Streaming Neural Compression

窄带语音通信的流式神经压缩

AI 总结:针对资源受限边缘设备上的低比特率语音通信难题,提出TinyCall轻量级神经音频编解码器,通过伪前瞻解码和RFSQ量化实现流式压缩,在树莓派3上以2.3 kbps比特率实时重建可懂语音。

链接:https://arxiv.org/abs/2609.25379

机构:University of Maryland, College Park(马里兰大学帕克分校)

作者:Dahong Luo, Anannya Trehan, Aritrik Ghosh, Nirupam Roy

英文摘要:Low-bitrate speech communication on resource-constrained edge devices remains challenging due to stringent computational, memory, and bandwidth constraints. We present TinyCall, a lightweight neural audio codec designed for real-time speech communication on low-power platforms such as the ESP32 microcontroller and Raspberry Pi. The proposed system targets emergency communication and other bandwidth-limited scenarios while preserving speech intelligibility, speaker identity, and vocal expressiveness. To enable efficient deployment, we propose a minimal neural audio codec architecture together with a framework for converting a causally trained codec into a truly streamable codec through pseudo-lookahead decoding and decoder-input caching. We further replace conventional residual vector quantization (RVQ) with Residual Finite Scalar Quantization (RFSQ) to reduce inference complexity on edge processors and employ a progressive three-stage training strategy for stable optimization under latent quantization. An MFCC-based perceptual loss encourages preservation of speaker characteristics, including harmonic structure and vocal timbre. Experimental results demonstrate real-time operation on a Raspberry Pi 3 while achieving intelligible speech reconstruction at bitrates as low as 2.3 kbps. The proposed approach demonstrates that practical neural speech communication is feasible on highly resource-constrained edge devices.


2. From Reliable Text to Real Voices: Trust-Aware Progressive Adaptation for Low-Resource TTS

从可靠文本到真实声音:低资源TTS的信任感知渐进式适配

AI 总结:针对低资源TTS适配中伪标签噪声问题,提出信任感知渐进式适配,先合成后真实并利用双ASR一致性加权,在缅甸语和老挝语上提升内容准确性与说话人相似度。

链接:https://arxiv.org/abs/2609.25951

机构:Beijing Logic Intelligence Technology(北京逻辑智能科技有限公司); University of Washington(华盛顿大学); Beijing University of Posts and Telecommunications(北京邮电大学); University of California, USA(加州大学); Northwestern University, USA(西北大学)

作者:Jiayi Lu, Yizhong Geng, Jinghan Yang, Tianhan Jiang, Boxun An, Yingming Gao, Ya Li

英文摘要:Low-resource text-to-speech (TTS) adaptation is constrained by scarce paired data and costly manual transcription. Existing fixed-voice TTS systems can provide relatively accurate pronunciation, but their synthetic speech offers limited speaker diversity and may exhibit flat prosody. Real recordings provide natural prosody and diverse voices, yet their automatic speech recognition (ASR) pseudo-labels may contain transcription errors. We find that supervision order affects content accuracy and speaker similarity. We propose trust-aware progressive adaptation: synthetic-to-real adaptation first establishes text-speech correspondences, then restores reference-speaker control using real speech. Transcript-agreement weighting uses agreement between two fixed ASR systems as a proxy for pseudo-label reliability to limit noisy supervision. Experiments with FireRedTTS3 on Burmese and Lao and OmniVoice on Burmese show improved content accuracy with high naturalness and competitive speaker similarity. Jointly considering supervision order and pseudo-label reliability when combining synthetic and real speech offers a practical path to zero-shot voice cloning in low-resource languages with less manual transcription. Audio demos are available at this https URL


3. SceneTTS-Bench: A Benchmark for Scene-Level TTS in Drama Dubbing

SceneTTS-Bench:戏剧配音中的场景级文本转语音基准

AI 总结:针对戏剧配音中文本转语音评估停留在句子级的问题,提出场景级基准SceneTTS-Bench,从音色一致性、情感表现力和节奏连贯性三维度评估,实验表明场景级排名与句子级指标显著不同。

链接:https://arxiv.org/abs/2609.26255

机构:Beijing University of Posts and Telecommunications(北京邮电大学); Beijing DeepLogic Intelligence Technology Co., Ltd.(北京深逻辑智能科技有限公司); University of California(加利福尼亚大学)

作者:Yizhong Geng, Yanliang Li, Jinghan Yang, Tianhan Jiang, Yingming Gao, Ya Li

英文摘要:Text-to-speech systems are increasingly used for drama dubbing, yet evaluation protocols remain sentence-level, leaving critical scene-level behaviors insufficiently measured. We present SceneTTS-Bench, a benchmark that evaluates TTS along three dimensions: timbre consistency across character turns, emotional expressiveness on high-tension utterances, and rhythm coherence under segmented long-form synthesis. The corpus comprises real-world and generated drama scripts totaling 160 bilingual scenes and approximately 10,300 utterances, with real-world scripts serving as the primary source (100 scenes) and generated scripts as a supplementary source (60 scenes), demonstrating the framework's extensibility through synthetic data augmentation. A backend-agnostic Canonical Intermediate Representation ensures fair cross-system comparison. Three automatic pipelines produce per-utterance diagnostics: Speaker Consistency Score for timbre-drift detection, Under-Acting Ratio for under-acting identification, and Rate Discontinuity Ratio for rate-discontinuity quantification. Experiments on four TTS systems confirm that each system exhibits distinct weaknesses and that scene-level rankings diverge substantially from sentence-level metrics. Benchmark resources are publicly available at this https URL.


2. 语音增强、降噪与音频修复 | 1 篇

4. Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement

真实对话语音增强中多说话人提取的挑战

AI 总结:针对真实对话中多说话人提取面临的沉默过多和注册语音不匹配挑战,提出新损失函数提升STOI和分段SNR,并分析不匹配影响。

链接:https://arxiv.org/abs/2609.25948

机构:University of Sheffield(谢菲尔德大学); South Westphalia University of Applied Sciences(南威斯特法伦应用科学大学)

作者:Robert Sutherland, Stefan Goetze, Jon Barker

英文摘要:Target-speaker and multi-speaker extraction are techniques for extracting speech from a desired speaker or desired speakers in the presence of other speakers and/or noise. Neural network approaches for this task are often trained and evaluated using simulated datasets, with balanced amounts of target speech and speaker enrolment samples which closely match the target speech. However, in real multi-party conversations, participants are often silent for more time than they are speaking, and their enrolment speech samples can differ substantially from the target speech in the conversation. These factors can impact the training and evaluation of these techniques on recordings of real conversations. This work proposes a new loss function, which helps mitigate the effect of excess silence in training, improving STOI from 0.55 to 0.60, and frequency-weighted segmental SNR from 4.35 to 5.12. Additionally, the impact of the mismatch between the enrolment speech and target speech is explored.


3. 音频事件检测与场景理解 | 1 篇

5. Benchmarking Open-Source Speech Emotion Recognition in Naturalistic Mandarin Spine Clinic Consultations: A Pilot Validation Study

开源语音情感识别在自然场景中文脊柱门诊咨询中的基准测试:一项初步验证研究

AI 总结:本研究在自然场景中文脊柱门诊语音中基准测试三个开源SER模型,发现多数类准确率误导,模型对少数情感检测效果差,表明临床部署需领域适应和更强验证。

链接:https://arxiv.org/abs/2609.26054

作者:Tsz Yuet Yeung, Zonglin He, Dong Chen, Huili Peng, Huiren Tao, Kenneth MC Cheung

英文摘要:Speech emotion recognition (SER) may enable passive affect monitoring in clinical encounters, but most systems are validated on acted laboratory speech rather than naturalistic Mandarin outpatient consultations. We benchmarked three open-source SER models (emotion2vec+, SenseVoice, FunASR) against a researcher-consensus reference in naturalistic spine-clinic speech, assessing minority-state detection under class imbalance. In a retrospective analysis of prospectively collected single-center recordings, audio was loudness-normalized and only conversations among patients, family members, and clinicians were retained. Sixty-five utterances (5-50 s; one per participant; 31 patients, 34 family members) were labeled by six calibrated annotators into six categories (Happy, Sad, Fear, Anger, Neutral, Surprised). Consensus used majority vote with Fleiss kappa filtering and clinician adjudication for low-agreement segments. Metrics included unweighted accuracy (UA), macro-average per-class accuracy, class- and sample-level weighted accuracy (WA), and F1 with bootstrap 95% CIs. Labels were imbalanced (Neutral 58.5%); median Fleiss kappa was 0.230 (IQR 0.134-0.519). SenseVoice and FunASR achieved UA 61.5% (95% CI 49.2-73.8%), macro-average per-class accuracy 87.2%, class-level WA 92.5%, and sample-level WA 24.6%. emotion2vec+ yielded UA 55.4% (95% CI 42.5-67.7%) and macro-average per-class accuracy 85.1%, with class-level WA 90.6% and sample-level WA 24.2%. Despite high inter-model agreement (90.8%), all models had near-zero recall for Sad, Fear, Anger, and Surprised. In this pilot, majority-class accuracy was misleading: SER poorly detected minority emotions against a noisy naturalistic reference. Clinical deployment readiness cannot be inferred from acted-corpus benchmarks without domain adaptation, multimodal modeling, stronger reference standards, and outcome validation.


4. 音乐信息检索与音乐生成 | 1 篇

6. MambaVoice: Lightweight Audiovisual Singing Voice Separation Via A Hybrid Mamba-Transformer Model

MambaVoice:基于混合Mamba-Transformer模型的轻量级视听歌声分离

AI 总结:提出轻量级视听歌声分离模型MambaVoice,采用混合Mamba-Transformer架构融合音频与视觉特征,在Acappella和URSing数据集上以1620万参数达到与大型模型相当的性能。

链接:https://arxiv.org/abs/2609.26635

作者:Adithi Shankar, Gopika Krishnan, Gloria Haro, Xavier Serra, Martín Rocamora

英文摘要:Isolating a target singing voice from a music video remains challenging, particularly in the presence of multiple vocalists and dense instrumental accompaniment. We propose MambaVoice, a lightweight audiovisual framework that leverages a hybrid Mamba--Transformer architecture for targeted singing voice separation. The model jointly encodes audio and visual streams using an attention-based band-split audio encoder and a spatio-temporal graph convolutional network (ST-GCN) for facial motion features. These modalities are fused through a multiplicative gating mechanism, enabling visual cues to selectively modulate audio representations. The fused features are processed by a hybrid backbone that combines Transformer self-attention with Selective State Space Models (SSMs), achieving efficient long-range temporal modeling with linear complexity. We evaluated MambaVoice on the Acappella and URSing datasets under challenging conditions, including mixtures with interfering singers. At 16.2 million parameters, the model demonstrates comparable performance, achieving 14.18 dB SDR on Acappella and strong cross-dataset performance on URSing, comparable to larger models at a fraction of the parameter count. These findings highlight the effectiveness of hybrid SSM--attention architectures for scalable, efficient audiovisual source separation, suggesting they are well-suited as lightweight components within larger pipelines. We conduct a perceptual study that further supports our improvements in objective metrics. We provide our implementation online.


5. 语音翻译与语音语言模型 | 1 篇

7. REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States

REVE:通过复用编码器状态实现大型音频语言模型的高效幻觉纠正

AI 总结:REVE通过复用大型音频语言模型已计算的编码器状态,以轻量方式验证并纠正幻觉事件提及,在AudioSet上移除了92.9%的标签不支持提及,且延迟仅为CED-Base的约1/18。

链接:https://arxiv.org/abs/2609.26028

机构:Beijing Institute of Technology, Zhuhai, Guangdong, China(北京理工大学(珠海))

作者:Hongjin Song, Jiasheng Kuang, Xinyu Yang, Qiuyu Fang, Ziyu Wu, Guowu Tan, Xiang Xie

英文摘要:Large audio-language models may mention acoustic events that are absent from the input. A separate audio event detector can verify these mentions, but doing so requires a second audio encoder and a separate forward pass. We propose Reused Encoder States for Verifying Events (REVE), a lightweight method that uses states already computed by the target model. One readout summarizes class scores across audio frames, while another uses pooled states from four consecutive frame intervals. Class-aware score fusion combines their outputs to verify generated event mentions without encoding the audio again. On AudioSet, REVE removes 92.9% of label-unsupported mentions under a faithful-mention recall constraint. With fewer added parameters and no second audio-encoding pass, REVE achieves a reduction comparable to those of CED-Tiny and CED-Base. Its complete verification latency is about 1/18 of the CED-Base path. Results on controlled DESED mixtures and different target-model architectures further confirm the effectiveness of encoder-state reuse.


6. 安全、隐私与深度伪造音频 | 3 篇

8. NeuMark: Neural Codec Resynthesis-Robust Audio Watermarking in the Codec Latent Space

NeuMark:编解码器潜空间中的神经编解码器重合成鲁棒音频水印

AI 总结:NeuMark在编解码器潜空间嵌入水印,利用交叉注意力在RVQ层注入16位消息,显著提升神经编解码器重合成下的鲁棒性,兼顾检测与恢复。

链接:https://arxiv.org/abs/2609.25719

机构:Nagoya University(名古屋大学)

作者:Annan Wu, Wen-Chin Huang, Tomoki Toda

英文摘要:Audio watermarking is increasingly important for tracing generated speech. Several audio watermarking methods have been proposed to embed the watermark in various domains, such as waveform, timbre feature, or latent representations, for making the embedded watermark robust against traditional digital signal processing (DSP) attacks. On the other hand, modern neural codecs introduce a different threat from DSP attacks: they resynthesize speech through quantized acoustic representations and can remove the embedded watermark evi- dence that is not aligned with codec-preserved structure. In this paper, we propose NeuMark, a codec-latent audio watermarking framework that embeds watermark evidence into SpeechTok- enizer acoustic tokens to address this resynthesis threat. NeuMark uses cross-attention to inject a 16-bit message across residual vector quantization (RVQ) layers, distributing the watermark over codec-aligned latent structure. Experimental results show that NeuMark substantially improves robustness under neural- codec resynthesis while supporting both watermark detection and message recovery. We also analyze the trade-off between reconstruction-referenced transparency and original-referenced robustness.


9. Boundary and Intra-Segment Learning for Partial Audio Deepfake Localization

部分音频深度伪造定位的边界与段内学习

AI 总结:针对部分音频深度伪造定位难题,提出边界与段内学习(BISL),联合帧、边界和片段信息建模真实性转换,在PartialSpoof上取得2.52% EER和97.40% F1,优于现有方法。

链接:https://arxiv.org/abs/2609.25822

机构:Sun Yat-sen University(中山大学); China Mobile Internet Corporation(中国移动互联网公司); The Hong Kong Polytechnic University(香港理工大学); Nanyang Technological University(南洋理工大学)

作者:Zhe Ye, Xiangui Kang, Minhua Huang, Kai Wu, Kong Aik Lee, Chng Eng Siong

英文摘要:Partial audio deepfakes manipulate only selected speech regions, making them difficult to be localized. Existing methods exploit boundary cues for partial deepfake localization, but primarily focus on identifying boundary positions rather than modeling the feature changes that characterize authenticity transitions. Meanwhile, the internal characteristics of continuous bona fide and spoofed segments remain underexplored. In this paper, we propose Boundary and Intra-Segment Learning (BISL), which introduces boundary learning to model feature differences between adjacent frames and distinguish authenticity transitions from general acoustic variations. In addition, intra-segment learning captures the overall characteristics of continuous bona fide and spoofed segments while enhancing feature consistency within each segment. By jointly learning frame, boundary, and segment information, BISL enables more effective fine-grained partial audio deepfake localization. Experiments on multiple localization benchmarks show that BISL achieves an EER of 2.52\% and an F1-score of 97.40\% on PartialSpoof, outperforming the compared methods, while maintaining competitive performance on HAD and improved cross-dataset performance on LPS. The code will be made publicly available upon acceptance.


10. Latent Audio Watermarking for Robustness to Neural Codec Resynthesis

用于对神经编解码器重合成具有鲁棒性的潜在音频水印


AI 总结:本研究提出一种基于冻结EnCodec潜在表示的音频水印方法,通过前馈嵌入器添加加性扰动,在神经编解码器重合成下比现有方法退化更慢、迁移性更好,并保持高检测率与感知质量。

链接:https://arxiv.org/abs/2609.25830

作者:Lovro Brulec, Sahil Karawade, Leonard Kinzinger

英文摘要:Existing waveform-domain audio watermarks are robust to many conventional distortions but can degrade substantially under neural codec resynthesis. We investigate whether continuous neural codec latents provide a more suitable embedding space using a restricted formulation built around frozen pretrained EnCodec. To test this, a feedforward embedder maps a multi-bit payload to an additive latent perturbation decoded through the unchanged codec decoder. Compared with AudioSeal and WavMark, our latent watermark formulation degrades more gradually under repeated and low-bitrate EnCodec resynthesis, transfers to unseen DAC, and retains high detection under most waveform distortions. Substantial EnCodec robustness emerges even without codec-resynthesis supervision, indicating that this behavior is inherent to our latent formulation and is further strengthened by codec-aware training. Learned perturbations are also preserved more strongly through codec cycling than equal-norm random controls, with preservation depending more on channel-specific allocation than temporal structure. End-to-end perceptual quality remains close to that of the frozen EnCodec reconstruction, indicating that much of the observed degradation originates from the codec carrier itself. Overall, these results show that continuous neural codec latents provide a promising embedding space for watermarks that remain robust to neural codec resynthesis.


7. 其他/综合语音音频 | 1 篇

11. Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning

使用MIDI Span条件化的标量量化潜变量进行多乐器音频混合的合成与编辑

AI 总结:本文提出SpanSynth-Edit,一种基于流匹配的MIDI引导多乐器音频合成与编辑模型,利用标量量化潜变量和MIDI Span条件化,实现帧内起始控制,并在基准上取得竞争力表现。

链接:https://arxiv.org/abs/2609.25546

机构:Queen Mary University of London(伦敦玛丽女王大学)

作者:Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos

英文摘要:Music creation often involves iterative refinement, changing selected musical details while retaining the rest. To support such refinement, we introduce SpanSynth-Edit, a flow-matching model for MIDI-guided synthesis and editing of multi-instrument audio mixtures using low-frame-rate scalar-quantised latents. MIDI Span encodes instrument-labelled note lifecycles as unordered event sets with continuous-valued attributes and pools each set into one conditioning vector per audio-latent frame. The model uses contextual audio for instrument-specific timbre guidance and supports editing by resynthesising the target region from revised MIDI. Experiments on single- and multi-instrument benchmarks show competitive performance and demonstrate within-frame onset control. We also discuss limitations of transcription-based note-adherence evaluation.