微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音合成与声音生成 3 篇
2. 语音增强、降噪与音频修复 1 篇
3. 音频事件检测与场景理解 1 篇
4. 音乐信息检索与音乐生成 2 篇
5. 其他/综合语音音频 3 篇
1. 语音合成与声音生成 | 3 篇
1. Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion
揭露你的伪装:从语音转换中恢复源说话者身份
AI 总结:研究针对语音转换威胁生物特征安全问题,提出TRIDENT框架,通过三叉戟架构,包括识别转换机制、提取目标说话者潜在表示等,从转换音频恢复源说话者身份,实验表明其准确率高且性能稳健。
链接:https://arxiv.org/abs/2607.23650
机构:Zhejiang University(浙江大学); Ant Group(蚂蚁集团)
作者:Hanlei Zhang, Zhongming Ma, Mingyang Zhang, Tengfei Liu, Yushi Cheng, Yanjiao Chen
英文摘要:Voice conversion (VC) poses a significant threat to biometric security by allowing attackers to impersonate target speakers. In forensic contexts, recovering the source speaker's identity from converted audio is vital for narrowing the field of suspects. To address this, we propose TRIDENT, a retracing framework designed to restore a source speaker's original identity from a converted audio sample. TRIDENT utilizes a three-pronged architecture consisting of a primary extractor and two auxiliary branches. The first auxiliary branch identifies the underlying voice conversion mechanism. This design acknowledges that even if the exact conversion strategy is unknown, a high-performance model adopted by the attacker is typically a derivative or variant of established mainstream ones. The second auxiliary branch extracts a latent representation of the target speaker, facilitating the isolation of target-specific traits from the composite converted audio sample. Finally, the main extractor leverages insights from both auxiliary branches to decouple confounding factors and distill a highly discriminative representation of the source speaker's identity. Experimental results demonstrate that TRIDENT achieves an accuracy as high as 90.99% against 7 state-of-the-art voice conversion methods. Furthermore, TRIDENT maintains robust performance under challenging conditions, including telephony channels, unseen languages, and adaptive scenarios.
2. Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
基于解耦时间深度扩散变换器的内存高效音频合成
AI 总结:研究提出基于解耦时间深度扩散变换器的内存高效音频合成架构,通过特定设计将语义音频令牌转换为RVQ表示,用单个深度解码器和因果滑动窗口注意力降低内存复杂度,在AMX上实现高效合成,提升音频质量。
链接:https://arxiv.org/abs/2607.23811
机构:Apple(苹果公司)
作者:Dongseong Hwang, Prasanth Yadla, Kaan Elgin, Shifas Padinjaru Veettil, Sivanand Achanta, Dipjyoti Paul, Ramya Rasipuram, Tyler Johnson, Emad Soroush, Chung-Cheng Chiu, Zhifeng Chen
英文摘要:Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple's most powerful on-device foundation model. This work presents the memory-efficient audio synthesis architecture behind that capability: a detokenizer that converts the semantic audio tokens emitted by the foundation model into high-fidelity audio within the tight compute and memory budget of the Apple Matrix Coprocessor (AMX). We convert semantic audio tokens to a residual vector quantization (RVQ) representation with a three-component design, a streaming encoder, a temporal decoder, and a depth decoder, that systematically decouples temporal and depth processing. A single reusable depth decoder with Diffusion Transformer (DiT)-style stage conditioning generates all RVQ levels autoregressively, replacing the dedicated per-level decoders of prior multi-decoder architectures, while causal sliding window attention with fixed-window key-value caching yields constant memory complexity independent of sequence length. Deployed on the AMX, the detokenizer sustains roughly 10 ms per generation step, about 16x faster than real time, with a peak runtime memory of only 21 MB and 329 MB of on-device assets, enabling continuous streaming synthesis of 20-320 seconds of audio. This constant, small footprint replaces the linear and quadratic memory scaling of conventional transformer- and GAN-based approaches. Ablation studies validate the key architectural components, and audio quality assessment confirms that the architecture maintains synthesis fidelity while achieving efficiency gains over existing methods. Operating at a 1-billion-parameter activation size within AFM 3 Core Advanced, it improves Mean Opinion Score by +0.28 overall (4.15 vs. 3.87) and by +0.42 on conversational speech (4.24 vs. 3.82) over the prior on-device text-to-speech system.
3. OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
OmniVAE:用于联合生成的具有跨模态对齐的音频-视频变分自编码器
AI 总结:研究针对音频和视频联合生成中跨模态对齐难的问题,提出OmniVAE,通过联合训练学习音频和视频潜在表示间的细粒度语义对齐,使用对比目标捕捉对应关系并对齐潜在空间,还提炼特征提升可学习性,实验证明其有效提升生成质量和同步准确性。
链接:https://arxiv.org/abs/2607.23855
作者:Jun Zhan, Chen Yang, Yitian Gong, Donghua Yu, Kuangwei Chen, Wenbo Zhang, Kexin Huang, Qi Luo, Zhe Xu, Ying Zhu, Jin Wang, Tengyue Zhang, Qi Chen, Cheng Chang, Songlin Wang, Junqi Dai, Jiasheng Ye, Xiaogui Yang, Tianyi Liang, Xiangyu Peng, Zhaoye Fei, Shimin Li, Qinyuan Cheng, Xie Chen, Xinchi Chen, Xipeng Qiu
英文摘要:Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1
2. 语音增强、降噪与音频修复 | 1 篇
4. Automatic Audio Equalization with Semantic Embeddings
基于语义嵌入的音频自动均衡
AI 总结:研究提出基于语义嵌入的数据驱动方法实现音频自动盲均衡,用预训练模型提供语义嵌入,训练轻量级头部,在音乐和语音上训练,经评估验证其有效性,凸显在现实音频增强应用中的潜力。
链接:https://arxiv.org/abs/2607.23846
机构:Aalto University(阿尔托大学); Nokia Technologies(诺基亚技术公司)
作者:Eloi Moliner, Vesa Välimäki, Konstantinos Drossos, Matti S. Hämäläinen
英文摘要:This paper presents a data-driven approach to automatic blind equalization of audio by predicting log-mel spectral features and deriving an inverse filter. The method uses a deep neural network, where a pre-trained model provides semantic embeddings as a backbone, and only a lightweight head is trained. This design is intended to enhance training efficiency and generalization. Trained on both music and speech, the model is robust to noise and reverberation. Objective evaluations confirm its effectiveness, and subjective tests show performance comparable to that of an oracle that uses true log-mel spectral features, indicating that the model accurately estimates the desired characteristics, with remaining limitations attributed to the filtering stage. Overall, the results highlight the potential of the method for real-world audio enhancement applications.
3. 音频事件检测与场景理解 | 1 篇
5. Mind the Microphone Gap: Benchmarking Array Upsampling Strategies for Latent Acoustic Mapping
注意麦克风差距:潜在声学映射的阵列上采样策略基准测试
AI 总结:研究潜在声学映射在稀疏4通道阵列中退化问题,对多种上采样架构进行基准测试,包括卷积网络等。通过不同训练方式研究与LAM对齐效果,发现原始全分辨率LAM最强,单独训练的轻量级模型最具竞争力,且表示对齐比模型复杂性更重要。
链接:https://arxiv.org/abs/2607.24463
作者:Philipp Schmidt, Huw Cheston, Juan Azcarreta, Adrian Stepien, Çağdaş Bilen, Iran R. Roman
英文摘要:Latent Acoustic Mapping (LAM) is a self-supervised learning method that generates high-resolution spherical acoustic maps from multichannel recordings without labelled data, matching supervised baselines on direction-of-arrival benchmarks. However, LAM degrades significantly with sparse 4-channel arrays, as the low-resolution cross-spectral matrix captures far less spatial information than the 32-channel inputs LAM was designed for. We benchmark a diverse set of upsampling architectures, spanning lightweight convolutional networks, iterative back-projection models, physics-informed networks, and generative adversarial approaches. We also study whether aligning these upsamplers with LAM by training them jointly or in different stages helps preserve the spatial structure that LAM depends on. Results show that the original full-resolution LAM is the strongest, that separately trained lightweight models are the most competitive learned approaches, and that representation alignment between the upsampler and LAM matters more than model complexity.
4. 音乐信息检索与音乐生成 | 2 篇
6. Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models
音乐源分离训练(MSST):训练和评估音乐分离模型的统一框架
AI 总结:本文针对音乐源分离任务提出MSST框架,统一训练、验证和推理,支持多种模型架构、数据处理方式、损失函数及评估指标,还有提升分离质量的实用技术,通过整合组件降低实验门槛,实现快速迭代。
链接:https://arxiv.org/abs/2607.23395
7. Modeling Stylistic Co-evolution in Symbolic Music Heritage Collections
对符号音乐遗产收藏中的风格共同演变进行建模
AI 总结:研究西方艺术音乐跨文化和声变化,提出从表示到动态的框架,通过一系列操作建模风格演变,应用于多国作品语料库,其轨迹和模式与音乐历史记载相符,为文化遗产研究提供定量层面。
链接:https://arxiv.org/abs/2607.23957
5. 其他/综合语音音频 | 3 篇
8. Infinite Canons: Maximally Self-Similar Melodic Lines and Canons with Infinite Solutions
无限卡农:具有无限解的最大自相似旋律线和卡农
AI 总结:研究具有无限解的无限卡农,基于离散素因数分解和连续对数法两种构造描述最大自相似旋律线结构,其声部垂直音程由同态给出,逆行倒影等同于移调,展示了结构并规划了未来工作。
链接:https://arxiv.org/abs/2607.23210
9. Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features
通过语言无关的发音特征改进零样本语音分类
AI 总结:研究针对语音到国际音标转录中标签不准确问题进行零样本语音分类评估,提出基于连续发音特征向量的分类方法,该方法优于离散标记法,且采用最佳时间聚合可提升分类效果,尤其对稀有音素和特定语音分类。
链接:https://arxiv.org/abs/2607.23606
10. Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and Gender
解开阿尔茨海默病病理学与感知中的声学线索:语言和性别的作用
AI 总结:研究阿尔茨海默病中声学线索,训练模型预测普通话和希腊语、男女说话者的临床AD状态及人类感知分数,用SHAP和统计模型分析,发现病理与感知一致性因语言和性别而异,强调特定人群可解释性审核对临床语音AI公平部署的必要性。
链接:https://arxiv.org/abs/2607.23977
