2026-07-31 | CS.SD语音与音频 | 共 9 篇

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音合成与声音生成 1 篇

2. 说话人识别、验证与分离 1 篇

3. 音乐信息检索与音乐生成 4 篇

4. 数据集、基准与评测 1 篇

5. 其他/综合语音音频 2 篇

1. 语音合成与声音生成 | 1 篇

1. Teffic-Audio: Tell Fact from Fiction

Teffic-Audio:区分事实与虚构

AI 总结:Teffic-Audio是一款基于Conformer架构的通用语音深度伪造检测系统,通过特定训练方案提升泛化能力,在Speech-DF-Arena测试集上的表现优于现有公开系统,为该领域提供实用参考。

链接:https://arxiv.org/abs/2607.28351

机构:Amphion Team(Amphion团队)

作者:Wan Lin, Li Wang, Jindong Wang, Kunyu Feng, Zhizheng Wu

英文摘要:Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be further shaped by variability in source speech, recording environments, and transmission channels. This variability makes robust generalization across heterogeneous conditions a central requirement for practical detection systems. This report presents Teffic-Audio, a general speech deepfake detection system designed for comprehensive evaluation environment. Teffic-Audio adopts a straightforward detector architecture consisting of a Conformer-based speech encoder, multi-head attentive statistics pooling, and a binary classifier. Rather than relying on additional architectural complexity, the system improves generalization through its training recipe, which integrates multi-source data, attack- and source-balanced sampling, and diverse audio augmentation. Trained only with open-source data, Teffic-Audio achieves a pooled EER of 1.454% on the 14 test sets of Speech-DF-Arena, outperforming all currently public systems on the leaderboard. It also obtains the lowest EER on five individual test sets and shows a favorable performance-complexity trade-off compared with larger leading systems. Overall, Teffic-Audio provides a strong and practical reference system for general speech deepfake detection.

2. 说话人识别、验证与分离 | 1 篇

2. Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO

Cocktail-Talker:基于Turn Action GRPO的嘈杂社交环境下多说话人对话建模

AI 总结:本文提出Cocktail-Talker框架,结合Turn Action GRPO,通过动作令牌建模助手行为,利用Cocktail-DialogGen生成数据,实现嘈杂社交环境下的多说话人口语对话建模,推动更自然的对话系统发展。

链接:https://arxiv.org/abs/2607.27756

作者:Xilin Jiang, Riki Shimizu, Sukru Samet Dindar, Junkai Wu, Zhongweiyang Xu, Nima Mesgarani

英文摘要:Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, are often more ambiguous: multiple speakers may participate in the same conversation amid irrelevant speech and background noise. Each utterance may be directed to the assistant, addressed to another speaker, or completely irrelevant. In such settings, the assistant must decide not only what to say, but also whether to speak at all. In this paper, we introduce Cocktail-Talker, a speech LLM framework for multi-speaker spoken dialog modeling in noisy social environments. We model the assistant's behavior with three action tokens: <|respond|>, <|listen|>, and <|ignore|>, placed before a response or silence. Cocktail-Talker is trained via supervised finetuning and reinforcement learning to generate the appropriate action token and, only in <|respond|> mode, a speech response. To prepare the training data, we develop Cocktail-DialogGen, an LLM-based data pipeline that simulates realistic multi-speaker dialogs with speaker roles across diverse social settings. Together, these components take a step toward spoken dialog systems that interact more naturally and selectively in complex social environments.

3. 音乐信息检索与音乐生成 | 4 篇

3. SKY-Piano: A Multimodal Piano Performance Dataset

SKY-Piano:多模态钢琴演奏数据集

AI 总结:该研究发布SKY-Piano多模态钢琴演奏数据集,含多类数据及标注工具,通过微调实验验证其用于MIDI到动作生成的可用性。

链接:https://arxiv.org/abs/2607.27296

机构:Seoul National University(首尔大学); KAIST(韩国科学技术院); Yamaha(雅马哈)

作者:Joonhyung Bae, Dawon Park, Taegyun Kwon, Yoon-Seok Choi, Hyeon Hur, Satoshi Obata, Shigeru Kai, Yohei Wada, Yu Takahashi, Akira Maezawa, Jaebum Park, Jonghwa Park, Juhan Nam

英文摘要:Music information retrieval research on piano performance increasingly involves diverse modalities of data and annotations beyond audio and MIDI. We present SKY-Piano, a multimodal piano performance dataset that includes 11 hours of performance recordings of motion, multi-view video, audio, MIDI from 7 professional and 12 amateur pianists along with MusicXML scores. The performance pieces were selected considering playing technique, difficulty, and performer expertise on a shared core repertoire. The motion data include both hand and body motion, released in both flagged form, where samples lost to marker occlusion are marked as unreliable, and imputed form, where those gaps are reconstructed, together with Visual3D body-segment kinematics and other time-synchronized modalities. To easily browse different modalities of data at a glance, we provide an interactive web browser. In addition, we developed a fingering annotation model and tool for deriving pseudo fingering annotations from the MIDI and motion data. Lastly, we present MIDI-to-motion generation through a fine-tuning experiment as a use case of the dataset.

4. VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition

VocalRender:面向实际作曲的原生乐谱歌声合成

AI 总结:VocalRender是一种原生乐谱歌声合成系统,通过歌词、音高等乐谱信息直接合成歌声,采用交错歌词-音符表示与自回归扩散模型,在2300小时数据集训练后,自然度CMOS较最强基线高0.42分,适配实际作曲工作流。

链接:https://arxiv.org/abs/2607.27768

作者:Yukun Chen, Tianrui Wang, Zhaoxi Mu, Xinyu Yang, EngSiong Chng

英文摘要:Existing singing voice synthesis systems often require predefined durations, explicit duration prediction, or time-aligned acoustic guidance, which limits their compatibility with practical composition workflows. We propose VocalRender, a score-native system that directly synthesizes singing from lyrics, pitches, symbolic note values, and tempo. It uses an interleaved lyric--note representation and an autoregressive diffusion model to generate continuous acoustic latents while predicting the output length, eliminating the need for explicit duration prediction. Trained on a 2,300-hour singing dataset, VocalRender achieves strong intelligibility, strong melody control, and high speaker similarity across both in-domain and out-of-domain benchmarks. Notably, it outperforms the strongest baseline by $0.42$ points in naturalness CMOS, demonstrating the effectiveness of our proposed score-native architecture.

5. CrowdioSet and PaRIRset: Two Datasets Towards Live Music Source Separation

CrowdioSet与PaRIRset:面向现场音乐源分离的两个数据集

AI 总结:该研究针对现有音乐源分离模型难以泛化到现场音乐的问题,构建CrowdioSet与PaRIRset两个数据集,可提升模型现场分离性能,相关资源已公开。

链接:https://arxiv.org/abs/2607.27828

作者:Enric Gusó, Xavier Serra

英文摘要:Most Music Source Separation (MSS) models do not generalize well to live music recordings because they are trained on studio recordings alone, disregarding the venue acoustics, the speaker system's response and audience noise. We propose to bridge this gap by providing and training a model on two novel datasets. First, we present CrowdioSet: a noise dataset comprising 4800 real ambience tracks from Freesound and synthetic sing-alongs for the vocals in MUSDB18 and MOISESDB datasets, generated from zero-shot singing voice conversions. CrowdioSet enables effective audio denoising for live recordings, resulting in superior separation both in objective and subjective evaluations. Second, we introduce PaRIRset, a stereo impulse response dataset captured across 40 professional concert venues using a microphone array. Our results show that adding PaRIRset RIRs increases the performance of a MSS model compared to using real RIRs from Speech Enhancement tasks alone. We make the examples, code, model weights, PaRIRset, and CrowdioSet freely available to the public.

6. Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances

将上下文嵌入整合到富有表现力的MIDI钢琴演奏评估中

AI 总结:本研究针对MIDI钢琴演奏评估忽略音符依赖、难以聚合表现力属性的问题,结合自监督符号音乐模型Aria和CLaMP3的上下文嵌入,提出适配符号音乐领域的Kernel Audio Distance,发布开源库Pereval。

链接:https://arxiv.org/abs/2607.27909

作者:Dmitrii Gavrilev, Ilya Borovik, Vladimir Viro

英文摘要:Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes. However, these methods often disregard dependencies between notes, which poses a potential limitation in assessing the similarity between two sets of performances. In generative applications, the wide variety of expressive attributes makes it difficult to aggregate them into a single scalar metric for model selection. In this work, we reexamine attribute-scoped metrics and explore the perceptual properties of contextual embeddings from self-supervised symbolic music models, Aria and CLaMP3. Results from our listening study indicate that these models can be used as perceptual proxies, showing agreement with per-sample human ratings on par with traditional metrics. To measure conditional distributional similarity, we adapt Kernel Audio Distance to the symbolic music domain. Unlike Pearson correlation and reconstruction error, kernel-based methods on contextual embeddings do not require note alignment and are sensitive to contextual perturbations. To facilitate reproducibility, we release Pereval, an open-source library that integrates performance evaluation utilities, including both attribute-scoped and deep feature metrics.

4. 数据集、基准与评测 | 1 篇

7. Does EEG Foundation Models Transfer to Speech? A Benchmark on Overt and Imagined Speech Decoding

脑电基础模型能否迁移至语音?显性与想象语音解码的基准测试

AI 总结:本文开展首个脑电基础模型语音解码基准测试,对比LaBraM等模型与EEGNet等基线,发现通用脑电预训练未在语音任务上展现一致优势,为语音专用基础模型研究提供依据。

链接:https://arxiv.org/abs/2607.27268

机构:University of Granada, Spain(格拉纳达大学); Research Centre for Information and Communication Technologies (CITIC-UGR), Spain(信息与通信技术研究中心); University of Chouaib Doukkali, Morocco(舒艾卜·杜卡利大学); Brain, Mind, and Behavior Research Center (CIMCYC), University of Granada, Spain(大脑、心智与行为研究中心)

作者:Owais Mujtaba Khanday, Mohamed Baha Ben Ticha, Sanae Belfrouh, Marc Ouellet, Jose A.Gonzalez-Lopez

英文摘要:EEG foundation models pretrained on thousands of hours have shown large gains over task-specific networks for motor imagery, seizure detection, sleep staging, and emotion recognition, but their transfer to speech decoding-arguably the most demanding non-invasive BCI application-remains untested. We present the first systematic benchmark of EEG foundation models against strong convolutional baselines for speech decoding, using two corpora: UGR-MINDVOICE (overt and covert Iberian Spanish) and BCI Competition 2020 Track 3 (imagined speech). We compare two foundation models (LaBraM, EEGMamba) against three established baselines (EEGNet, ShallowFBCSPNet, EEGConformer) under a unified preprocessing and fine-tuning protocol. Large-scale EEG pretraining yields no consistent advantage over a 16K-parameter CNN on speech tasks, indicating that current general-purpose EEG pretraining does not yet transfer to speech production and motivating speech-specific foundation models.

5. 其他/综合语音音频 | 2 篇

8. Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage

增强执法音频转录:基于LoRA的Whisper对随身摄像头(BWC)素材的适配

AI 总结:针对警务音频转录的可见性悖论,该研究提出基于LoRA的Whisper适配框架,结合符号推理流程,在消费级硬件上实现高压力场景下的执法音频转录,达成93.7%词汇映射率,助力警务问责与透明度提升。

链接:https://arxiv.org/abs/2607.27245

作者:Vivek Senthil, Zhiqiang Tao, Ernest Fokoué

英文摘要:Modern policing faces a "visibility paradox" where law enforcement agencies possess petabytes of Body-Worn Camera (BWC) footage that remains largely unutilized for accountability or systemic review due to the prohibitive labor costs of manual transcription. This research presents a framework for adapting the OpenAI Whisper architecture to the unique acoustic and linguistic challenges of the policing environment. By employing Parameter-Efficient Fine-Tuning (PEFT) through Low-Rank Adaptation (LoRA), we address the significant performance degradation observed in zero-shot models when confronted with high-stress scenarios, sirens, and radio interference. Crucially, we demonstrate that this adaptation is feasible on consumer-grade hardware (Acer Nitro local machine with NVIDIA 4GB GTX GPU) using 8-bit quantization and gradient checkpointing. We further integrate these transcriptions into a symbolic reasoning pipeline using a domain-specific ontology to transform raw audio into evidence-linked incident graphs, achieving a 93.7% lexicon mapping rate for the advancement of procedural justice and transparency.

9. Improved Robustness in AI-Generated Music Detection

AI生成音乐检测的鲁棒性改进

AI 总结:该研究针对AI生成音乐检测器在简单音频操作下鲁棒性不足的问题,提出频率缩放不变的检测流水线,结合log-STFT重映射等技术,实现对速度修改等攻击的防御,还兼具可解释性。

链接:https://arxiv.org/abs/2607.27454

作者:Emile Dugelay, Thomas Barand, Baptiste Campeas, Aurélien Laouar, Darius Afchar, Romain Hennequin

英文摘要:AI music generators leave predictable spectral artifacts determined by their architecture. Existing detectors exploit these artifacts with near-perfect accuracy on raw generated tracks, but their performance collapses under simple audio manipulations, such as speed modification or pitch shifting. We address this open robustness problem by introducing a frequency-scaling-invariant detection pipeline that aims to prevent this kind of attack by design. Our method maps audio onto a log-frequency axis via a log-STFT remapping. A single learned cross-correlation filter, combined with max-pooling, provides shift invariance at inference time. Training uses a hybrid loss that jointly supervises binary detection and artifact-peak localization, regularizing boundary weights. Because robustness to speed change is built in by design, the detector is also interpretable: it outputs both a binary decision and an estimate of the applied speed-change factor.