微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音合成与声音生成 1 篇
2. 音乐信息检索与音乐生成 1 篇
3. 语音翻译与语音语言模型 2 篇
4. 多模态音频与视听学习 1 篇
5. 安全、隐私与深度伪造音频 1 篇
6. 其他/综合语音音频 3 篇
1. 语音合成与声音生成 | 1 篇
1. Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens
基于神经音频编解码器令牌剖析自监督语音学习对训练语言的敏感性
AI 总结:该研究通过控制变量分析发现,基于神经音频编解码器的自监督语音模型的下游性能依赖SSL预训练语言,不依赖编解码器训练语言,据此提出单个编解码器可跨语言复用,需对齐预训练语言与目标语言。
链接:https://arxiv.org/abs/2607.26350
机构:National Institute of Advanced Industrial Science and Technology (AIST)(日本国立先进工业科学技术研究所); Carnegie Mellon University(卡内基梅隆大学)
作者:Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell, William Chen, Satoru Fukayama, Shinji Watanabe
英文摘要:Neural audio codecs (NACs) have become popular for obtaining speech representations as discrete tokens. Beyond compression, discrete tokens can be used to train self-supervised learning (SSL) models. Such models, referred to as codec-based SSL models, reduce data storage and computational cost, enabling scalable SSL pre-training. However, their language sensitivity remains unclear. When the language changes, codec-based SSL models may require retraining, which undermines their efficiency. In this paper, we present a systematic analysis of language sensitivity by varying either the NAC training language or the SSL pre-training language while keeping the other fixed. Experimental results show that downstream performance is insensitive to the NAC training language but strongly dependent on the SSL pre-training language. These findings suggest that a single NAC can be reused across languages, while aligning the SSL pre-training language with the target language is crucial.
2. 音乐信息检索与音乐生成 | 1 篇
2. MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation
MPEcho:一种旋律与音素感知的可控翻唱歌曲生成生成框架
AI 总结:本研究针对现有翻唱歌曲生成模型音素错误率高的问题,提出旋律与音素感知的MPEcho框架,集成音素编码器与长度调节器,结合自研Phonsa转录模型,实现了更精准的可控翻唱歌曲生成。
链接:https://arxiv.org/abs/2607.26698
作者:Wei-Jaw Lee, Hsuan-Yu Yeh, Ting-Yi Hu, Chih-Pin Tan, Fang-Duo Tsai, Yi-Hsuan Yang
英文摘要:Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components. The state-of-the-art model SongEcho utilizes $F_0$ sequences and voiced/unvoiced (V/UV) tags for conditioning; however, implicit linguistic information from V/UV tags cannot guarantee lyric accuracy, leading to a high phoneme error rate (PER). Inspired by singing voice synthesis (SVS), we propose MPEcho, which integrates a phoneme encoder and a length regulator (LR) into the SongEcho framework. By providing explicit phoneme-level conditioning and precise temporal boundaries, MPEcho significantly reduces PER. To enable this, we developed Phonsa, a Whisper-based automatic transcription model that provides high-precision phoneme-level annotations for singing voices, overcoming the scarcity of high-quality audio-phoneme pairs. Experimental results validate the effectiveness of Phonsa for alignment and MPEcho for end-to-end CSG. The audio samples, code and weights can be accessed from this https URL.
3. 语音翻译与语音语言模型 | 2 篇
3. Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis
音频大语言模型中基于韵律的越狱:一项对照研究与机制分析
AI 总结:该研究通过控制文本内容改变语音传递预设构建评估协议与基准,发现特定韵律特征的语音能提升音频大语言模型的越狱成功率,表明语音传递是音频大语言模型安全评估的重要因素。
链接:https://arxiv.org/abs/2607.26541
机构:City University of Hong Kong(香港城市大学); City University of Hong Kong (Dongguan)(香港城市大学(东莞))
作者:Jiachen Qian, Junyu Li
英文摘要:Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or broader style transfer. We study this question by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary. We present PJ-Break, a black-box evaluation protocol with presets targeting arousal, authority, and speaking rate, together with AdvAudio-Prosody, a 600-sample benchmark with acoustically verified attributes. On the exact post-QC Qwen2-Audio panel, the Q=1 Panic (38/95), Anger (35/95), and Fast (32/95) presets are all well above Neutral (4/95). The fixed six-query pool covers 44/95 Qwen2-Audio seeds and 15/95 GPT-4o seeds and exceeds a matched-budget StyleBreak reimplementation (27/95) on Qwen2-Audio. A same-voice pool excluding the confounded Commanding condition still reaches 40/95, and a retained-panel ablation shows emotional-delivery audio alone (44/95) is far more effective than emotional text alone (11/95). Exploratory surrogate diagnostics and pilot mitigation observations are secondary, non-core analyses. Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation
4. ThinkOmni: A Reasoning-Driven Omni-Modal LLM Framework for Audio Forgery Detection and Localization
ThinkOmni:用于音频伪造检测与定位的推理驱动全模态大语言模型框架
AI 总结:针对现有AFDL方法泛化能力不足的问题,提出推理驱动的全模态大语言模型ThinkOmni,构建FACoT数据集,引入FMIL与FCML,实现了跨数据集的音频伪造检测与定位。
链接:https://arxiv.org/abs/2607.26553
机构:Shenzhen University(深圳大学); Afirstsoft Technology Group Co., Ltd.(安福软件科技集团有限公司)
作者:Yuxiong Xu, Kaiqing Lin, Bin Li, Haodong Li, Sheng Li
英文摘要:Existing audio forgery detection and localization (AFDL) methods often overfit dataset-specific low-level artifacts, limiting their generalization to subtle, localized, and unseen manipulations. Recent audio large language model (ALLM)-based approaches cast AFDL as question answering but still model forensic evidence implicitly, without linking manipulation cues to predictions. To bridge this gap, we propose ThinkOmni, a reasoning-driven omni-modal large language model that jointly performs explicit forensic reasoning, spoofing detection, and temporal manipulation localization. To enable explicit reasoning supervision, we construct Forensic-Aware Chain-of-Thought (FACoT), a 100K-sample dataset with structured forensic evidence and reasoning annotations. Leveraging FACoT, we introduce Forensic-Aware Modality-Incremental Learning (FMIL), which progressively aligns semantic, acoustic, and spectral-visual representations with the LLM backbone to capture complementary forensic cues. We further propose Forensic-Consistent Multi-task Loss (FCML), which combines weighted cross-entropy with an adaptive localization loss to coordinate reasoning generation, spoofing detection, and temporal localization. Extensive experiments show that ThinkOmni achieves strong cross-dataset generalization in both detection and localization. Code, models, data, and inference examples are available at this https URL.
4. 多模态音频与视听学习 | 1 篇
5. MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
MMAC:一个用于音频字幕生成的大规模多维度基准测试集
AI 总结:该研究针对音频字幕生成评估的信息覆盖与可靠性诊断难题,构建了含5638个音频的MMAC多维度基准,评估多种AudioLLMs并将发布基准与代码。
链接:https://arxiv.org/abs/2607.27109
作者:Weijie Wu, Junbo Li, Lin Li, Jun Fang, Qingyang Hong
英文摘要:With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.
5. 安全、隐私与深度伪造音频 | 1 篇
6. Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection
用于跨域音频深度伪造检测的多比率DiT重建残差的音频锚定融合
AI 总结:该研究针对音频深度伪造检测器跨域性能下降问题,提出音频锚定融合多比率DiT重建残差的方法,在ASVspoof 5和ITW数据集上取得优于基准的跨域检测性能。
链接:https://arxiv.org/abs/2607.26472
作者:Haotian Mo, Jie Liu, Siqi Shen, Songzhu Mei, Xinhai Chen, Xiangyang Wang, Yigui Feng, Shuai Li, Gencheng Liu, Keqi Yang, Qinglin Wang
英文摘要:Audio deepfake detectors often degrade when generators, corpora, or recording conditions change. We use a Diffusion Transformer (DiT), trained only on bona fide speech, as a frozen reconstruction probe. Reconstructions at masking ratios 0.5, 0.75, and 0.9 yield explicit multi-ratio residual maps. Because these residuals are domain sensitive, our audio-anchored detector passes the projected frozen-WavLM auditory representation into the fusion sum without gate-based attenuation and uses residuals only as a scalar-gated additive correction. The pre-specified seed-42 run obtains 6.5442% EER / 0.18456 min-DCF on ASVspoof 5 Eval and 13.8372% / 0.36921 on ITW Full; three-seed means are 6.8885 (0.3308)% and 15.3328 (2.0719)%. The latter is below a separately optimized WavLM-ResNet18 reference under both supervision settings. Auxiliary supervision raises dynamic competitive fusion from 18.4007% to 25.2968% mean ITW EER, worsening all three seeds. The results support reconstruction residuals as complementary evidence and motivate a non-competitive auditory path for ASVspoof 5-to-ITW transfer, without claiming a componentwise causal ablation of anchoring alone.
6. 其他/综合语音音频 | 3 篇
7. Explicit Note-Event Tokenization and Pitch-Validity Constrained Decoding for MIDI-to-Tablature Transcription
用于MIDI到吉他谱转录的显式音符-事件分词与音高有效性约束解码
AI 总结:本研究提出带显式音符-事件分词与音高有效性约束解码的吉他谱转录框架,在DadaGP和Francois Leduc数据集上提升了吉他谱准确率。
链接:https://arxiv.org/abs/2607.26440
机构:National Taiwan University(台湾大学)
作者:Ting-Kai Hsu, Wei-Chin Wang, Kai-Xi Hong, Yu-Hua Chen
英文摘要:Guitar tablature transcription predicts the string and fret position for each note so that the resulting tablature reproduces the target musical part. Prior sequence-to-sequence approaches have shown promising results on large-scale datasets, but their generalization behavior across different dataset scales remains less explored. In this work, we propose a guitar tablature transcription framework with explicit note-event tokenization and regularized training. The proposed decoder token representation incorporates note-event tokens together with TAB tokens, allowing note boundaries, pitch-related events, and string-fret positions to be represented more explicitly. We evaluate the proposed framework on DadaGP, a large-scale dataset, and Francois Leduc, a small-scale dataset. Our method improves tablature accuracy over the Fretting Transformer baseline on DadaGP, with especially strong gains when trained directly on the small-scale Leduc dataset. We further introduce a pitch-validity constrained decoding strategy that masks pitch-invalid TAB candidates during generation rather than correcting them after decoding and simultaneously preserves the original timing and note structure from the input. This constraint improves tablature accuracy and provides a controlled setting for measuring how much error remains after pitch-invalid predictions are removed. Our code will be released at: this https URL
8. Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
基于直推式原型优化与类别对数几率增强的小样本开放集音频分类
AI 总结:针对小样本开放集音频分类的原型易受未知类别污染问题,提出两阶段直推式方法,结合内类度加权与解耦评分,在三个音频数据集上取得最优结果。
链接:https://arxiv.org/abs/2607.26607
作者:Tianyan Deng, Yanxiong Li, Rui Gao, Jiahao Du
英文摘要:Few-shot Open-set audio classification requires classifying query samples from known classes with a few labeled support samples while rejecting query samples from unknown classes. Transductive inference jointly observes the full unlabeled query set to improve prototype estimation, yet standard transductive updates do not distinguish known from unknown query samples, leaving prototypes vulnerable to open-set contamination. Drawing on latent-inlierness weighting and decoupled scoring for unknown-class samples, we propose a two-phase transductive method operating over a frozen audio encoder. First, each query sample is assigned a latent inlierness score that down-weights likely unknown-class samples, so that prototype refinement is driven primarily by known-class evidence. The refined prototypes are then directly optimized on a transductive loss combining support cross-entropy, inlierness-weighted conditional entropy minimization, and inlierness-weighted marginal entropy maximization, while open-set rejection uses a prior-adaptive free-energy score that adjusts its threshold with the prior proportion of unknown-class samples, decoupling detection from classification. Experiments on three audio datasets show our method achieves state-of-the-art results for few-shot open-set audio classification under multiple experimental conditions.
9. Detection of AI-generated stems within hybrid human-AI music
混合人机音乐中AI生成音轨的检测
AI 总结:本文提出首个混合人机音乐中AI生成音轨的检测研究,提出并行架构结合源分离与音轨专用分类器,在MUSDB18-HQ数据库上取得了令人鼓舞的检测效果。
链接:https://arxiv.org/abs/2607.26874
作者:François Rigaud, Gabriel Meseguer-Brocal, Benjamin Martin, Romain Hennequin
英文摘要:This paper presents, to the best of our knowledge, the first study on detecting human-AI hybrid music tracks created by mixing human-produced and AI-generated stems. Building on recent work showing that AI music detectors can identify decoder-related artifacts in fully generated music, we investigate whether such artifacts remain detectable at the stem level after mixing. Using MUSDB18-HQ database in a two-stem vocals + accompaniment setting, we simulate hybrid mixtures by autoencoding individual stems with a neural codec. We compare two strategies combining AI-generated mix detection and source separation. A naive sequential pipeline, where source separation is followed by detection on separated sources, confirms that artifacts associated with an AI-generated stem are not reliably recovered by generic source separation systems. We therefore propose a parallel architecture in which source separation is only used to estimate source-relative energy within the mixture. We then train simple stem-specific binary classifiers that take as input the generated mix prediction together with the relative energy of the target stem on short audio chunks. Averaging chunk-level predictions yields encouraging track-level results, highlighting the potential of such approaches for detecting AI-generated stems in hybrid music.
