今日论文合集:CS.SD语音与音频 | 共 8 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

1. Domain-Adaptive ASR for Telephony AI Agents: Fine-tuning Canary Flash Models for Enterprise Contact Center Applications
面向电话智能体的领域自适应ASR:针对企业联络中心应用微调Canary Flash模型
AI 总结:本研究基于NVIDIA NeMo框架微调Canary Flash模型,构建电话导向数据集,经四项实验验证其可提升电话环境下ASR的识别性能,同时保持实时响应性。
链接:https://arxiv.org/abs/2608.24916
机构:Botnoi Group(博特诺伊集团)
作者:Chanameth Boonpramuk, Winn Voravuthikunchai, Songpol Bunyang
英文摘要:This technical report describes Botnoi Group's methodology and results for rapidly fine-tuning the open-source NVIDIA Canary 180M Flash and NVIDIA Canary 1B Flash multitask models for speech-to-text tasks using the NVIDIA NeMo framework, with a focus on telephony-grade audio. To support this adaptation, we construct a telephony-oriented fine-tuning dataset from live voicebot system recordings and prompted speech with telephony-oriented augmentation. We evaluate four targeted experiments-language adaptation (Thai), telephony robustness, domain-specific jargon (names and addresses), and latency-using character error rate (CER) for accuracy and real-time factor (RTFx) for inference speed. Results show that fine-tuning substantially improves recognition in noisy telephony environments, reducing CER from 23.31% to 9.04% on BOTNOI telephony data, and further improves business-critical names and addresses from 16.98% to 3.78% CER through domain-specific adaptation. Overall, our results show that domain-adaptive fine-tuning enhances business-critical terminology while preserving real-time responsiveness for production voicebot deployments.

2. Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace
我们能否解读音频大语言模型的思维?一个可解读的多语言中间层工作空间
AI 总结:本研究通过logit lens分析Qwen3-Omni模型,发现其中间层可在输出前清晰呈现音频问题答案,揭示了音频模型推理的特性与信号分布,为解读音频大语言模型内部思维提供了定性依据。
链接:https://arxiv.org/abs/2608.24958
机构:Amazon AGI Foundations(亚马逊AGI基金会); University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
作者:Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Qi Luo, Jia-Hong Huang, M. Maruf, Roger Ren, Yile Gu, Rahul Pandey, Ge Liu, Ivan Bulyko
英文摘要:An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model writes its reasoning down. Reading a base Qwen3-Omni with a logit lens at the audio-token positions, we find that the answer to a spoken question becomes legible - in words - in the model's middle layers, before it emits any token. Five findings follow. (1) The readout carries concepts in neither the question, the options, nor the model's own transcription: on a clip whose verbatim transcription is empty garbling, it reconstructs Watergate and scandal, passes through the role president, and resolves to Nixon - a hidden multi-hop chain, read with no chain-of-thought. (2) The content is language-agnostic: one audio-inferred concept surfaces in several scripts at once, and 38% of top-1 readouts are Chinese on English inputs. (3) It is paralinguistic: given the same clip as audio and as the model's own emotion-free caption, the audio mind forms the sound source, speaker role, or affect that the caption discards, and answers correctly more often. (4) The audio-driven signal is absent at the input, turns on about a tenth of the way into the network, separates most cleanly from the text prior in the middle band (35-80% of depth), and activation patching shows it is causally used and committed before the last fifth of the layers. (5) Deleting single layers maps the pipeline: reading the sound in is localized to the entry layers and answer delivery to the output layer, while retrieval is distributed across the interior. Throughout, a waveform-swap control - identical text, only the sound changed - isolates the audio-driven signal from a prior over the printed options. This is a qualitative account of what an audio model works out before it speaks: the quantities are controls, not benchmark scores.

3. AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models
AudioLens:基于推理音频-语言模型的多视角语音聚类
AI 总结:该研究提出音频多视角聚类任务,构建基准AudioLens-Bench,开发经推理蒸馏和偏好优化训练的AudioLens-R1模型,实验显示其在语音聚类任务中性能优于基线模型。
链接:https://arxiv.org/abs/2608.25177
机构:University of California, Irvine(加利福尼亚大学欧文分校); Dartmouth College(达特茅斯学院); Purdue University Northwest(普渡大学西北分校)
作者:Wenjun Huang, Qiaosong Chu, Tiger Shao, Pengfei Zhang, Yutong Song, Hanning Chen, Yezi Liu, Weiyi Wu, SungHeon Jeong, Ryozo Masukawa, Sanggeon Yun, Yang Ni, Jiang Gui, Mohsen Imani
英文摘要:Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. To study this setting, we construct AudioLens-Bench, a benchmark spanning multiple application domains and evaluating both in-perspective and cross-perspective generalization. We further propose AudioLens-R1, an end-to-end large audio-language model trained with reasoning distillation and preference optimization. Experiments show that AudioLens-R1 consistently outperforms all baselines, improving overall ARI by 12.99 points and V-measure by 11.62 points. These results demonstrate the promise of native audio-language models for flexible, perspective-conditioned structure discovery over speech collections.

4. AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP
AllMusicCaps:将专辑评论作为Music CLAP的补充监督
AI 总结:本研究提出AllMusicCaps,利用AllMusic专家专辑评论经LLM预处理构建标题语料库,结合SigReg正则化优化CLAP模型,提升了文本到音乐检索等任务性能并发布相关资源。
链接:https://arxiv.org/abs/2608.25244
作者:Pablo Alonso-Jiménez, Xavier Lizarraga-Seijas, Xavier Serra, Dmitry Bogdanov
英文摘要:Recent open text-audio contrastive models (CLAPs) are typically trained with LLM-generated captions derived from tag datasets or web search results, which tend to be accurate but expressively narrow. As a complementary source, we explore human-written album reviews, specifically expert reviews from AllMusic: they exist at scale and carry narrative cues, evaluative adjectives, and scene framing that other sources lack. Since raw reviews are too noisy for direct use as captions, we first build a caption corpus with 24,5346 samples via an LLM preprocessing pipeline that identifies descriptive musical quotes and rewrites them into training-ready captions. We find that album review supervision yields the largest retrieval gains on a human-written caption benchmark (Song Describer), particularly for complex queries that other existing caption datasets leave uncovered. In addition, we revisit the training recipe and show that SigReg regularization, which encourages an isotropic Gaussian distribution in the embedding space, improves MLP probing across classification tasks, as well as text-to-music retrieval. The resulting model outperforms open CLAP-style baselines on text-to-music retrieval, zero-shot classification, and most MLP probing tasks. We release the review-derived caption dataset and model weights to support future research.

5. A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography
一种无需训练的、基于自嵌入隐写术的针对部分语音操纵的主动防御方法
AI 总结:针对部分深度伪造语音检测难题,本文提出一种无需训练的主动防御方法,通过将自嵌入策略与现有音频隐写术结合,借助编解码器修复实现检测,可与被动防御互补且数据效率高。
链接:https://arxiv.org/abs/2608.25285
机构:National Institute of Informatics(情报信息研究所)
作者:Yigitcan Özer, Zhe Zhang, Wanying Ge, Xin Wang, Junichi Yamagishi
英文摘要:Partial deepfake speech, where only limited segments of an utterance are synthesized or manipulated, poses a significant challenge to existing deepfake detection systems. As the proportion of spoofed regions decreases, passive detectors become increasingly unreliable, and accurate detection and restoration remain challenging. In this paper, we revisit audio steganography from a new perspective and propose its use as a proactive defense against partially deepfaked audio. In particular, we consider a self-embedding strategy in which a clean speech signal embeds a compressed representation of itself, enabling post-hoc extraction of reference content. We demonstrate how existing audio steganography methods can be repurposed to support detection of partial deepfakes through codec-based restoration. Experiments on a benchmark dataset show that the proposed approach complements passive defenses. Remarkably, the proposed method operates without any training, providing a robust and data-efficient alternative for partial deepfake detection.

6. Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs
将自嵌入音频水印与超低比特率神经编解码器相结合
AI 总结:本研究结合自嵌入音频水印与超低比特率神经编解码器,实现了无训练的音频操纵检测定位与被操纵区域恢复,发现神经编解码器选择是影响性能的主导因素。
链接:https://arxiv.org/abs/2608.25289
机构:National Institute of Informatics(信息学研究所)
作者:Yigitcan Özer, Xin Wang, Zhe Zhang, Junichi Yamagishi
英文摘要:Partial manipulation of speech recordings, where only localized segments of an utterance are altered, poses a significant challenge for content integrity verification, as reliable detection and localization of such edits becomes harder as the manipulated proportion decreases. Watermarking offers a proactive defense alternative by embedding auxiliary information prior to distribution; classical hash-based schemes achieve near-perfect detection and localization under ideal conditions, but the original content cannot be recovered once a segment is manipulated. Building on a prior self-embedding audio steganography framework, this work presents an initial exploration of proactive defense performance under ideal conditions, extending the investigation along three axes: frame-level localization, multi-bit least significant bit variants, and evaluation across multiple ultra-low-bitrate neural codec representations. By embedding a compact neural codec representation rather than a cryptographic hash, the framework additionally enables recovery of the manipulated regions, while supporting training-free detection and localization without spoofed examples. Experiments across four controlled manipulation types under ideal channel conditions show that the embedded payload, and hence an approximate reconstruction of the authentic content, is always fully recovered without bit errors. The results also indicate that the choice of neural codec is the dominant factor for detection and localization performance.

7. SPECTRA: Subspace-Preserving Embedding Calibration, Transport, and Replay for Fully Few-Shot Class-Incremental Audio Classification
SPECTRA:用于全少样本类增量音频分类的子空间保留嵌入校准、迁移与重放
AI 总结:针对全少样本类增量音频分类的性能下降问题,提出含嵌入校准适配器、子空间特征重放、直推式最优传输优化的SPECTRA框架,在三类基准上优于现有方法。
链接:https://arxiv.org/abs/2608.25054
机构:University of Haifa(海法大学); University of Stuttgart(斯图加特大学)
作者:Giries Abu Ayoub, Loay Mualem, Simon Korman
英文摘要:Fully few-shot class-incremental audio classification (FFCAC) requires recognizing new sound classes from only a handful of labeled examples per session, without forgetting previously learned classes and without any large base dataset. Existing methods typically freeze a pre-trained audio--language encoder and classify with point prototypes, but they suffer from significant performance degradation throughout the sessions due to generic feature representations. We propose SPECTRA, a framework built on a frozen encoder which adds three components. (i) a lightweight trainable adapter that calibrates the generic embeddings to the task; (ii) subspace feature replay, an exemplar-free anti-forgetting scheme that replays old classes by sampling from the low-rank subspace of their stored features; and (iii) a transductive optimal-transport refinement of prototypes at test time. Our central finding is that the subspace structure of the replay diminishes forgetting and outperforms naive Gaussian replay of equal variance. On three FFCAC benchmarks (NSynth-100, FSC-89, LS-100), SPECTRA improves average accuracy and reduces forgetting over current state-of-the-art methods, and our ablations statistically validate each component.

8. Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding
不协和频谱(Dissonance Spectrum,DS):显式建模感知频率交互以实现更好的音乐理解
AI 总结:本文提出不协和频谱(DS)这一时频表示,经实验验证其在音乐相关任务中性能优于基线等方法,可作为音乐理解的可解释互补表示。
链接:https://arxiv.org/abs/2608.25621
机构:Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院); Central Conservatory of Music(中央音乐学院); Peking University(北京大学)
作者:Tianle Wang, Xinyi Tong, Liangke Zhao, Jishang Chen, Sirui Zhang, Haoxin Zhang, Xin Jin, Duo Xu, Xiaobing Li, Song-Chun Zhu
英文摘要:Conventional music representations describe acoustic energy over time and frequency but do not explicitly expose relations among simultaneous frequency components. We introduce the \emph{Dissonance Spectrum} (DS), a nonnegative time--frequency representation that applies a tolerance-based rational pitch-relation kernel with logarithmic harmonic distance to a constant-Q spectrum and attributes aggregate pairwise interactions back to individual frequency bins. Controlled music-theory tests show strong ordinal agreement for intervals, harmonic-function connections, and church modes, and weaker but significant agreement across diverse chord voicings. DS is then encoded by a lightweight parallel branch whose zero-initialized residual projection preserves the baseline function at initialization. Across six paired training seeds in open-ended music question answering and categorical and dimensional music emotion recognition, DS obtains the highest mean on every reported endpoint relative to the unchanged baseline, a parameter-matched Gaussian-input branch, and an architecture-matched magnitude-CQT branch. These results support DS as an interpretable, complementary representation, while listener-specific perception and broader task coverage remain open problems.