今日论文合集:CS.SD语音与音频 | 共 7 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
1. Low-Power End-to-End Cochlear Implant Speech Denoising with Spiking Neural Networks
采用脉冲神经网络的低功耗端到端人工耳蜗语音去噪
AI 总结:本研究针对人工耳蜗用户在嘈杂环境中语音理解困难的问题,提出受Deep ACE架构启发的脉冲神经网络,可同时实现语音增强与CI编码,在保持性能的同时能耗降低超6倍。
链接:https://arxiv.org/abs/2608.28493
机构:University of Sherbrooke(谢布鲁克大学)
作者:Ludovic Boulanger, Sean U. N. Wood
英文摘要:Cochlear implants (CI) restore hearing for individuals with severe to profound hearing loss. However, CI users often struggle to understand speech in noisy environments. Deep neural networks (DNN) have shown promise in enhancing speech for CI users, yet their high energy demands make them non-ideal for low-power CI processors. Spiking neural networks (SNN), on the other hand, offer comparable performance with significantly lower energy consumption. Hence, we propose a novel SNN inspired by the Deep ACE architecture that simultaneously performs speech enhancement and CI coding. Our model achieves competitive vocoded short-time objective intelligibility (VSTOI) and signal-to-noise ratio improvement (SNRi) scores compared to Deep ACE, while achieving more than a sixfold reduction in energy consumption.
2. Auditing Generative Audio Calls for Known-Task Audio-LLM Evaluation
针对已知任务音频-LLM评估的生成式音频调用审计
AI 总结:本研究将生成式音频调用评估设为受控决策问题,通过VocalSound数据集实验发现,带生成式调用的选择器准确率达0.925,优于无调用对照,明确了生成式调用在音频-LLM已知任务评估中的边际价值。
链接:https://arxiv.org/abs/2608.27817
机构:National Research Council Canada(加拿大国家研究委员会)
作者:Mengzhe Geng
英文摘要:Speech and audio LLMs are often evaluated by asking whether a waveform prompt beats an automatic speech recognition (ASR) transcript. For known closed-set tasks, that comparison conflates two factors: access to acoustic evidence and the need to call a generative audio model. We evaluate this distinction as a controlled call-decision problem. For each example, a policy chooses among keeping a transcript label, using encoder evidence from Contrastive Language-Audio Pretraining (CLAP), Audio Spectrogram Transformer (AST), or WavLM, and calling Qwen2-Audio, Qwen2.5-Omni, or MOSS-Audio; the decisive ablation removes all generative actions while keeping the selector and development protocol fixed. On VocalSound, transcripts reach 0.296 accuracy, so waveform information is needed. Yet supervised CLAP and WavLM controls reach 0.850 and 0.854 with no generative audio calls. A selector with generative actions reaches 0.925 accuracy using 12.5% calls, compared with 0.921 for the matched no-call selector (paired difference 0.004; 95% CI [-0.025,0.033]). Agreement and stacking features improve weaker selectors but do not beat the strongest no-call control. For known-task endpoint claims, the relevant quantity is the marginal value of the generative call after transcript and encoder evidence have already been used.
3. Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages
韵律在翻译中会丢失吗?跨语言的细粒度韵律相似性研究
AI 总结:本研究利用多语言配音数据开展跨语言韵律细粒度分析,发现特定语言间韵律结构存在内在相关性,为表达性语音到语音翻译系统提供了实证指导。
链接:https://arxiv.org/abs/2608.27848
机构:Center for Language and Speech Processing (CLSP), Johns Hopkins University(约翰斯·霍普金斯大学语言与语音处理中心)
作者:Haopeng Xie, Ismail Rasim Ulgen, Sofia Son, Berrak Sisman, Philipp Koehn
英文摘要:Prosody plays an important role in speech translation, conveying information such as emphasis, emotion, and intent beyond lexical content. However, despite recent progress in expressive speech-to-speech translation (S2ST), little is known about how prosodic patterns are similar/different across languages. Understanding these cross-lingual similarities and differences is crucial for effectively incorporating prosody into expressive S2ST systems. In this work, we present the first fine-grained cross-lingual analysis of prosody using multilingual dubbing data across English-German, English-Spanish, and English-French language pairs. We analyze the similarity of pitch, energy, and temporal feature patterns between source and target speech and investigate the linguistic and alignment-related factors affecting this similarity. Our analysis reveals inherent cross-lingual correlations in prosodic structure between certain languages. The findings provide important insights into the transferability of prosody across languages and offer empirical guidance for future expressive speech-to-speech translation systems.
4. Evaluating Loss Functions in Differentiable Out-of-Domain Sound-Matching with Partial Parameter Distance
可微分域外声音匹配中损失函数的评估:基于部分参数距离
AI 总结:本文提出部分参数距离(PPD)解决域外声音匹配的损失函数评估难题,在7种场景评估4种损失函数,验证PPD可作为诊断工具辅助损失函数选择。
链接:https://arxiv.org/abs/2608.27698
机构:University of Alberta(阿尔伯塔大学)
作者:Amir Salimi, Daniel Penner, Kalvin Eng, Abram Hindle, Osmar R. Zaïane
英文摘要:In out-of-domain (OOD) sound-matching, a synthesizer is optimized to mimic a sound it did not generate. OOD evaluation of loss functions is underexplored in part because the standard "parameter loss" metric requires a shared parameter space between target and imitator, which OOD settings lack. We introduce Partial Parameter Distance (PPD), which applies parameter loss only to the critical parameters that mismatched synthesizers share (e.g., filter cutoffs), enabling automatically evaluated OOD experiments; we verify its results with blinded listening tests. Across seven scenarios involving band-pass filtering, amplitude modulation, and pitch-bending, we evaluate four differentiable loss functions (SIMSE_Spec, L1_Spec, JTFS, DTW_Envelope). Loss-function effectiveness remains tightly coupled to the method of synthesis: SIMSE_Spec excels at filter-cutoff recovery, DTW_Envelope at amplitude-modulation recovery, and JTFS at smooth pitch trajectories. Parameter-based evaluation agrees with listening tests on the top-ranked loss function in five of seven scenarios, demonstrating its utility as a diagnostic tool.
5. Klangfarbenakkord and Klangfarbenharmonien Metric Space Models for Music on Informational Geometry 1
音乐信息几何中的音色和弦与音色和声度量空间模型 1
AI 总结:本文提出几何和声学科,将乐器音色作为元素,利用Wasserstein距离与偏差扩展传统音乐理论框架,并以单簧管喉音G为例说明其在合奏中的实用价值。
链接:https://arxiv.org/abs/2608.28026
作者:Yusei Tamura, Shigekazu Ishihara, Ken Ito
英文摘要:This paper deals with the introduction of "geometric harmony", a discipline that explicitly addresses the spectral characteristics of musical gamut. The framework of Western music, from Renaissance to the present, represents sound in terms of "pitch"-as is evident from its five-line staff notation system-and employs the fundamental frequency as its representative value, 440 Hz, etc. In this paper, by taking the timbres of specific individual instruments as elements and examining the Wasserstein distance between two voices, and Wasserstein deviations between three or more voices, we demonstrate that it is possible to expand the system whilst retaining the entire framework of conventional music theory. At the same time, as an example of practical utility in ensemble playing, we provide a detailed account of the two-voices affinity of the "Throat G" on clarinet, a note known for its fragility in ensemble contexts.
6. Exploring the Design Space of Representation Learning for Audio Transformations
探索音频变换表征学习的设计空间
AI 总结:该研究针对音频变换表征学习的设计选择问题,构建含三个目标的统一框架,揭示两类嵌入的互补作用,结合架构与训练改进后,表征在多项任务中优于现有基线。
链接:https://arxiv.org/abs/2608.28127
作者:Sungho Lee, Marco Martínez-Ramírez, Junghyun Koo, Wei-Hsiang Liao, Kyogu Lee, Yuki Mitsufuji
英文摘要:Neural audio representation learning has enabled a range of content-oriented applications, but the resulting features remain limited for tasks involving audio processing. Furthermore, it is not obvious what processing-aware representations should capture: the processing itself, abstracted away from source content, or the processed audio that retains it. Existing approaches implicitly commit to one or the other and also differ in their models, data, and evaluation, obscuring which design choices drive their behavior. We address both questions within a unified framework of three objectives: processing consistency, description alignment, and equivariance via forward prediction. We compare all combinations of the objectives under a controlled setup and reveal their relative strengths and interactions. Our framework produces both a transformation embedding and a processed-audio embedding, and we find that the two play complementary roles: distance-based tasks favor the former, while probe-based tasks favor the latter. Combined with improvements in network architecture and training pipeline, our representations outperform prior baselines across retrieval, probe-based evaluation, and style transfer.
7. Multirate State Space Models for End-to-End Processing of Pulse Density Modulated Speech Signals
用于脉冲密度调制语音信号端到端处理的多速率状态空间模型
AI 总结:该研究提出多速率状态空间模型,构建PDM语音端到端处理架构,可实现调制与采样率无关的表示,在不同采样率下性能优异,大幅降低处理时间步数。
链接:https://arxiv.org/abs/2608.28472
机构:University of Sherbrooke(谢布鲁克大学)
作者:Ludovic Boulanger, Sean U. N. Wood
英文摘要:Deep neural networks (DNNs) based on state-space models (SSMs) are increasingly applied to speech processing, but typically operate on pulse-code-modulated (PCM) audio. This constrains deployment on low-power, always-on edge devices, which commonly use single-bit pulse-density-modulated (PDM) micro-electromechanical (MEMS) microphones for their noise robustness, low cost, and variable sampling rates that enable low-power operation. In fact, converting PDM to PCM requires low-pass filtering and decimation, imposing costly overhead on resource-constrained hardware. While prior works have attempted to process PDM signals directly, they require long training times and generalize poorly across sampling rates. In this paper, we show that the SSM has two key properties that remediate these issues: its continuous-time parametrization allows it to produce a consistent representation of the input audio signal, regardless of the modulation strategy and sampling rate, and its long-term memory enables this representation to be aggressively downsampled without needing any anti-aliasing operations. We then propose a novel end-to-end PDM speech processing architecture that uses an SSM to encode the input audio signal into a modulation- and sampling-rate-invariant latent representation. We show that our proposed architecture achieves robust speech classification and enhancement gains at low-power sampling-rates (512 kHz) and similar performance to state-of-the-art algorithms operating on PCM data when tested on standard PDM sampling-rates of 2 MHz. Moreover, we show that the SSM's output can be downsampled by more than 65,000 times, thus significantly reducing the number of processing timesteps in downstream layers.
