今日论文合集:CS.SD语音与音频 | 共 10 篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音合成与声音生成 1 篇

2. 说话人识别、验证与分离 1 篇

3. 语音增强、降噪与音频修复 1 篇

4. 音频事件检测与场景理解 1 篇

5. 音乐信息检索与音乐生成 2 篇

6. 语音翻译与语音语言模型 1 篇

7. 低资源、多语言与方言语音 1 篇

8. 数据集、基准与评测 1 篇

9. 其他/综合语音音频 1 篇

1. 语音合成与声音生成 | 1 篇

1. Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis

文本到语音合成中语音表示的弗雷歇距离损失

AI 总结:研究针对少步TTS模型训练问题,提出语音表示弗雷歇距离损失(SR-FD),在微调时让模型用相同采样器合成语音,将其特征与参考统计量匹配,无需判别器和推理计算,显著降低Seed-TTS英语的字错误率,是提高可懂度的分布正则化方法。

链接:https://arxiv.org/abs/2607.06027

作者:Ho-Lam Chung, Kuan-Po Huang, Bo-Ru Lu, Hung-yi Lee

英文摘要:Few-step diffusion and flow-matching text-to-speech (TTS) models are usually trained with local objectives, such as conditional flow matching, reconstruction, and stop prediction. These losses provide stable optimization, but they never ask whether sampled speech follows the distribution of high-quality speech. We propose Speech Representation Fr'echet Distance loss (SR-FD), a training-time distributional regularizer for tokenizer-free flow-matching autoregressive TTS. During fine-tuning, the model synthesizes speech with the same few-step sampler used at deployment, and SR-FD matches the mean and covariance of frozen Whisper and CTC features of this speech to reference statistics computed offline from three complementary content targets. The loss requires no discriminator and no inference-time computation. On Seed-TTS English, four-step SR-FD fine-tuning reduces WER from the original four-step VoxCPM2 baseline's 2.2279% to 1.4147%, a 36.5% relative reduction, and also surpasses the original ten-step baseline at 1.7366%; both gains are significant under an utterance-level paired bootstrap. Speaker similarity and objective quality proxies are preserved at the ten-step level, and an error analysis shows the gain comes from content substitutions across all prompt lengths. SR-FD is thus an intelligibility-improving distributional regularizer for few-step TTS.

2. 说话人识别、验证与分离 | 1 篇

2. Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling

基于最佳N生物特征采样的流匹配语音源分离

AI 总结:针对单通道语音分离难题,提出基于条件流匹配的方法,训练时用冻结说话人编码器定义源顺序,推理时用于生物特征最佳N候选选择等,在Libri2Mix基准上评估,该方法在分离指标及下游任务中表现出色。

链接:https://arxiv.org/abs/2607.06088

机构:NVIDIA(英伟达)

作者:Anastasia Zorkina, Alexandr Anikin, Nikita Khmelev, Anastasiya Korenevskaya, Sergey Novoselov, Vladimir Volokhov, Maxim Korenevsky, Yuriy Matveev

英文摘要:Single-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of processing long recordings with chunk-wise inference. We address these issues with a conditional flow-matching-based method that produces an ordered two-source output conditioned on the mixture. A frozen speaker encoder defines the source order during training and is reused at inference for biometric best-of-$N$ candidate selection and chunk-level channel alignment. We evaluate separation quality on Libri2Mix benchmark using SI-SDR, PESQ, and ESTOI, and measure downstream impact using cpWER for automatic speech recognition and EER for speaker verification. The results show that the proposed Transformer U-Net variant is competitive with strong baselines in objective separation metrics and achieves the lowest downstream automatic speech recognition and speaker verification error rates in all evaluated settings.

3. 语音增强、降噪与音频修复 | 1 篇

3. Learning-based Physics-Constrained Neural Kernel for Sound Field Estimation With Source-Position-Dependent Directional Weighting

基于学习的物理约束神经核用于声源位置相关方向加权的声场估计

AI 总结:研究针对声场估计问题,提出基于学习的物理约束神经核方法,通过声源位置相关的隐式神经表示构建方向加权函数,使核函数能捕捉共同模式并泛化,实验证明该方法在估计方向加权函数上优于单快照方法。

链接:https://arxiv.org/abs/2607.06274

作者:Mattia Marella, Shoichi Koyama

英文摘要:A learning-based physics-constrained neural kernel for sound field estimation is proposed. Sound field estimation aims to estimate the spatial distribution of an acoustic field from a discrete set of microphone measurements, which have a wide range of applications. Among existing sound field estimation methods, kernel-regression-based methods offer a flexible and principled framework for incorporating physical constraints and allow inference through linear operation. It is also possible to adapt the kernel function to the target acoustic environment by representing the directional weighting function as an implicit neural representation (INR) and optimizing hyperparameters using measurements. However, the kernel function is generally optimized for single snapshot measurements of the microphones, which can lead to strong overfitting and poor generalization. We propose a source-position-dependent INR for the directional weighting function, enabling the kernel function to capture common directional patterns and to generalize to unseen source positions in the target acoustic environment. Experimental results indicate that our proposed method outperforms the snapshot-based method by estimating a directional weighting function that matches the directivity of the target sound field.

4. 音频事件检测与场景理解 | 1 篇

4. Determinantal point process sampling for bioacoustic active learning

用于生物声学主动学习的行列式点过程采样

AI 总结:研究针对生物声学主动学习,提出CARE-DPP方法,结合类平衡预测不确定性与嵌入空间新颖性,用DPP目标选采集批次,通过不确定性与新颖性平衡及自适应采集计划,在多数据集评估中表现优于基线,消融实验明确主要贡献因素。

链接:https://arxiv.org/abs/2607.06063

作者:Hugo Magaldi, Gabriel Dubus

英文摘要:Eco-acoustic monitoring generates vast volumes of audio data, making active learning a promising approach for reducing annotation effort while efficiently training reliable biodiversity classifiers. This report presents CARE-DPP, a batch active-learning acquisition method submitted to BioDCASE Active Learning for Bioacoustics 2026 challenge. The method combines class-balanced predictive uncertainty with embedding-space novelty, while a determinantal point process (DPP) objective selects a high-quality and non-redundant acquisition batch. The uncertainty-novelty balance is annealed over the annotation budget: early cycles emphasize geometric coverage, whereas later cycles increasingly exploit classifier uncertainty. To mitigate unreliable early scores, the DPP candidate pool mixes top-quality candidates with a decreasing proportion of random exploration. An adaptive acquisition schedule uses smaller batches early and larger batches later. Evaluated over five repeats on the BirdSet HSN, POW and UHH subsets and on ATBFL, CARE-DPP obtains a mean development AULC of 0.50 for macro mAP, compared with 0.46 for the official CoreSet baseline. Ablations identify DPP batch diversification and the adaptive acquisition schedule as the largest contributors.

5. 音乐信息检索与音乐生成 | 2 篇

5. From Textural Counterpoint to Feature Encoding: A Multi-Dimensional Machine Representation Study of Haydn's "The Lark" Integrating Electroacoustic Analysis

从纹理对位到特征编码:基于电声分析的海顿《云雀》多维机器表征研究

AI 总结:该研究针对现有深度音乐生成模型在复调交互中角色感知不足问题,对海顿《云雀》进行跨学科分析,提出“古典形态定性分析 - 电声定量测量 - 机器表征重构”路径,引入新特征提取方法,为构建人机协作音乐系统奠定理论基础。

链接:https://arxiv.org/abs/2607.05902

机构:Shenyang Conservatory of Music(沈阳音乐学院); Education Information Center, Shenyang Conservatory of Music(沈阳音乐学院教育信息中心)

作者:Yakun Liu, Zhiyu Jin, Hai Luan, Dong Liu, Xiaonan Li

英文摘要:Chamber music, as a highly precise multi-part interactive system, contains a logic of "role assignment and dynamic interaction" that provides an extremely valuable blueprint for exploring human-computer collaborative composition paradigms. Addressing the lack of role perception capabilities in existing deep music generation models during polyphonic interactions, this paper conducts an interdisciplinary analysis of Haydn's String Quartet in D Major, The Lark (Op. 64, No. 5). We propose a novel research path: "Classical Morphology Qualitative Analysis-Electroacoustic Quantitative Measurement-Machine Representation Reconstruction." The study first utilizes auditory analysis to dissect the counterpoint morphology of the leading voice and the underlying groove in the first movement. Subsequently, it introduces spectrum and dynamic feature analysis tools from a Digital Audio Workstation (DAW) to translate subjective auditory perception into objective, measurable physical parameters. Building on this, the paper introduces a fundamentally new approach to low-level computer feature extraction: completely abandoning the traditional mechanical quantization grid, introducing Event-based Timestamps to record the duration of micro-timing, and transforming acoustic features into an independent "Role-Aware Encoding" as an aesthetic heuristic mechanism (a phenomenological anchor). This study not only completes the logical loop spanning classical analysis, electronic music mapping, and AI symbolic generation but also establishes a profound theoretical foundation-from the perspectives of interactive aesthetics and media philosophy-for constructing human-computer collaborative music systems imbued with "social attributes" and "otherness awareness."

6. Designing Maintainable Hybrid Generative Systems: A Quantum-Inspired Approach to Automated Music Harmony Generation

设计可维护的混合生成系统:一种受量子启发的自动音乐和声生成方法

AI 总结:研究从旋律自动生成音乐和声,核心方法是结合量子启发的候选探索与基于规则的优化,主要贡献是设计并评估出可维护的混合生成架构,其生成的和声保留结构与行为,优化层提升多种性能,且无需训练语料库。

链接:https://arxiv.org/abs/2607.06296

作者:Josef Pavlicek

英文摘要:This paper presents the design and evaluation of a maintainable hybrid generative architecture for automated music harmony generation from melody. The proposed system combines quantum-inspired candidate exploration over overlapping melodic contexts with explicit rule-based optimization to balance generative flexibility and structural control. The architecture is evaluated using explicit and reproducible metrics covering structural coherence, functional agreement, harmonic similarity, and robustness. The results show that the proposed approach produces harmonizations that preserve tonal structure and cadential behavior while allowing multiple valid harmonic realizations. Furthermore, the optimization layer improves structural coherence, stability, and predictability without requiring a training corpus. The study demonstrates that transparent and controllable hybrid generative systems can be systematically designed and evaluated within the context of Information Systems Development.

6. 语音翻译与语音语言模型 | 1 篇

7. Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models

摆脱普罗克汝斯忒斯之床:用于音频语言模型的分组正交连接器

AI 总结:研究音频语言模型压缩中存在的问题,提出ORCA方法,通过分组使查询输出指向不同方向,在SAKURA多跳推理中表现出色,提升分数,减少查询冗余并提高跨说话者方差。

链接:https://arxiv.org/abs/2607.06014

作者:Ho-Lam Chung, Ke-Han Lu, Yi-Cheng Lin, Guan-Ting Lin, Yiming Chen, Hung-yi Lee

英文摘要:Audio-language models compress a speech encoder's output through a Querying Transformer (Q-Former) connector before feeding it to a large language model. We identify two failures in this compression. The connector's output vectors collapse to a single direction, and different speakers produce nearly indistinguishable outputs, with paralinguistic cues such as speaker identity, gender, and prosody lost along the way. Our method, ORCA, reverses this collapse by splitting the queries into groups whose outputs are constrained to point in different directions. On SAKURA multi-hop reasoning, ORCA gains 26.4 points over an identically trained 4B baseline, reaching 75.2% (vs. 49.0% for the 8B Audio Flamingo-3). At the connector level, the same change cuts query redundancy by 12x and raises cross-speaker variance by 75x.

7. 低资源、多语言与方言语音 | 1 篇

8. BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech

蓝鹊语音合成系统:用于台湾口音语码转换语音的高效令牌分词器、语言模型和语音合成系统

AI 总结:研究针对现成语音合成系统不适配台湾普通话的问题,提出从底层解决文本问题的方法,包括训练PangolinTokenizer分词器、Barbet语言模型,构建蓝鹊语音合成系统,显著降低了字符错误率和词错误率,获多数盲听投票青睐。

链接:https://arxiv.org/abs/2607.06054

作者:Ho Lam Chung, Bo-Xuan Zheng, Cheng-Chieh Huang, Cheng-Han Chang, Jung-Ching Chen, Lok-Lam Ieong, Ting-Lin Hsiao, Yu-Cheng Lee, Yi-Hsin Chung, Yu-Kai Guo, Hung-yi Lee

英文摘要:Off-the-shelf TTS systems are poorly adapted to Taiwanese Mandarin. Their accent defaults to other Mandarin variants, their tokenizers over-segment common Taiwanese text, and their pronunciation degrades at code-switching boundaries where Chinese and English alternate within one utterance. These problems share one root: the text side lacks adaptation to the Taiwanese context. We address the text side from the bottom up. PangolinTokenizer, a byte-level BPE tokenizer trained on Taiwan-context data, reaches the lowest token rate (0.485 tokens/character) with the smallest vocabulary among nine tokenizers. Barbet, a billion-parameter Traditional-Chinese language model trained on PangolinTokenizer, serves as the text-semantic frontend and ranks first among comparable public models on a 14-task evaluation. BlueMagpie-TTS attaches Barbet to the pretrained acoustic stack of VoxCPM2 through a learned bridge, keeping the acoustic stack fixed. On a 1000-sentence Taiwan-localized test set, it lowers CER from 11.45% to 4.81% and WER from 14.83% to 5.36%, relative reductions of 58.0% and 63.9%. In a blind listening study on 500 of these sentences with ten listeners, 65.6% of majority votes prefer BlueMagpie-TTS.

8. 数据集、基准与评测 | 1 篇

9. Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music

我关心的音乐:对(几乎)任何音乐的大语言模型音乐感知技能进行自动化多模态基准测试

AI 总结:针对当前音乐基准测试的局限,引入MusICA-MetaBench框架,利用结构化符号表示和预定义模板生成按需基准测试,通过ChoraleBricks数据集展示并确定规模,经与基线比较证明能测音乐感知,推动了大语言模型音乐感知跨模态评估。

链接:https://arxiv.org/abs/2607.06015

作者:Tomáš Sourada, Katia Vendrame, Jan Hajič jr

英文摘要:Music represents a cornerstone of human culture, existing digitally across diverse modalities, including audio, symbolic encodings (e.g., MIDI, MusicXML), and sheet music. Despite the advancement of Multimodal Large Language Models (MLLMs), current music benchmarks face three major limitations. First, large static benchmarks are resource-intensive to evaluate, and it remains unclear how their results transfer to diverse kinds of music beyond those included in the benchmark. Second, benchmarks claiming to measure "music understanding" often fail to require music perception. Third, they do not support systematic performance comparisons across musical modalities. To overcome these issues, we introduce the Music I Care About Meta-Benchmark (MusICA-MetaBench), a framework that automatically derives on-demand benchmarks directly from user-provided data. By leveraging structured symbolic representations (e.g., MusicXML) and our pre-defined question templates, we build multiple-choice question-answer pairs that probe music perception competencies, aligned with music pedagogy, across audio, music notation images, and symbolic files. We demonstrate our framework with the ChoraleBricks dataset, and experimentally determine benchmark sizes that ensure statistically reliable model comparisons for this setup. By comparing against text-only and white-noise baselines, we show our questions do measure music perception. Ultimately, MusICA-MetaBench represents a significant advancement in the cross-modal assessment of music perception for MLLMs. By proposing a dataset-specific benchmarking paradigm, it enables efficient on-demand evaluation of music perception capabilities.

9. 其他/综合语音音频 | 1 篇

10. InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective

InsideSSL:从模型中心视角理解自监督语音表示

AI 总结:研究旨在理解自监督语音表示模型内部动态,提出InsideSSL框架,从压缩、几何等视角分析,引入GCM评估,通过线性探测连接到下游任务,展示层拓扑对语音相关编码的作用。

链接:https://arxiv.org/abs/2607.06392

机构:Inria, Université Grenoble Alpes CNRS, LJK(法国国家信息与自动化研究所、格勒诺布尔阿尔卑斯大学、法国国家科学研究中心、法国国家科学研究中心计算机图形学实验室)

作者:Samir Sadok, Xavier Alameda-Pineda

英文摘要:Self-supervised learning (SSL) models, such as Wav2Vec2, HuBERT, and WavLM, have become foundational across a wide range of speech and audio tasks. Despite their success, understanding their internal layer-wise dynamics remains an ongoing challenge. To address this, we propose a two-part model-centric framework called InsideSSL. First, we establish a task-agnostic analysis from three intrinsic per-layer perspectives: compression (entropy), geometry (curvature), and robustness to perturbations. We show that varying training objectives induce distinct regimes of acoustic compression and manifold unfolding. Second, we introduce the cross-layer Generative Compatibility Matrix (GCM) to evaluate functional transferability, exposing stable phonetic cores, identity volatility, and deep-layer semantic pruning. In addition to these evaluations, linear probing connects the model-centric perspective to downstream tasks, demonstrating how layer topology dictates phoneme, pitch, and speaker encoding.