今日论文合集:CS.SD语音与音频 | 共 6 篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准


1. Removing Speech, Keeping Activities: A Privacy Firewall for Acoustic Sensing in Assisted Living

去除语音,保留活动:面向辅助生活声学感知的隐私防火墙

AI 总结:该研究提出基于U-Net编解码器的隐私防火墙流水线,可去除环境音频中的语音并保留活动相关环境音,在ESC-50、SINS及真实家庭录音上验证了其隐私保护与活动识别性能,解决老年护理声学感知的隐私障碍。

链接:https://arxiv.org/abs/2609.02376

机构:KIOS Research and Innovation Center of Excellence, University of Cyprus(塞浦路斯大学KIOS卓越研究与创新中心); School of Computing, University of Kent(肯特大学计算机学院)

作者:Pavlos Nicolaou, Christos Efstratiou

英文摘要:Acoustic sensing offers a promising non-intrusive approach for monitoring daily activities of older adults, yet speech privacy concerns remain a critical barrier to real-world deployment. We present a privacy firewall pipeline based on a U-Net encoder-decoder, trained entirely on synthetic data, that removes speech from ambient audio while preserving environmental sounds indicative of daily activities. Activity recognition is performed using VGGish transfer learning with an SVM classifier. Evaluated on the ESC-50 and SINS datasets across multiple speech content levels, the proposed model reduced residual speech to 0% VAD-detectable speech (Silero Voice Activity Detection) under all tested conditions, outperforming Facebook Denoiser (6.55% residual), SepFormer (36.34%) and ConvTasNet (47.21%) on ESC-50 at the 100\% speech level. On ESC-50 at 40% speech level, classification performance recovers to 85% precision and 85% recall after speech removal, compared with 81%/75% before removal and an 84%/83% speech-free baseline. Evaluation on real-world participant home recordings collected with the AudioHive app showed 0% VAD-detectable speech after processing while maintaining 76% precision and recall. The pipeline enables privacy-preserving acoustic sensing without sacrificing activity recognition performance, addressing a key obstacle to the adoption of ambient monitoring in elderly care.


2. Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction

基于语音印象引导的伪三元组构建的可扩展方向跟随TTS

AI 总结:该研究针对方向跟随TTS缺乏训练数据的问题,提出伪三元组构建流程,结合可控印象TTS与LLM生成数据,实验验证其可实现稳定的说话人保留修改,结合真实数据可进一步提升方向对齐度

链接:https://arxiv.org/abs/2609.02623

机构:NTT, Inc.(NTT公司)

作者:Kenichi Fujita, Yusuke Ijima

英文摘要:Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system generates a new utterance that reflects a given direction relative to a reference utterance while preserving speaker identity and linguistic content. A key challenge is the lack of training data capturing such relative modifications. To address this, we propose a scalable pseudo-triplet construction pipeline that generates~(reference utterance, direction text, modified utterance) triplets. It generates controlled style variations using an impression-controllable TTS model and uses an LLM to produce natural language directions from estimated impression differences. Experimental results demonstrate that pseudo-triplets alone enable stable speaker-preserving modification, and that combining pseudo and recorded data further improves direction alignment while maintaining speaker similarity. Audio examples are available on our demo page this https URL


3. Auditory Illusion Benchmark for Large Audio Language Models

面向大型音频语言模型的听觉错觉基准

AI 总结:该研究针对LALMs缺乏听觉错觉基准的问题,构建了首个含10种错觉的AIB基准,对比模型与人类响应发现LALMs存在系统性差异,为探究听觉认知提供了新测试平台。

链接:https://arxiv.org/abs/2609.02277

机构:Seoul National University(首尔大学)

作者:Hayoon Kim, Eunice Hong, Kyogu Lee

英文摘要:Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replicate human perceptual tendencies. Despite their importance, most benchmarks focus on visual illusions or general audio tasks, leaving auditory illusions underexplored. To this end, we present AIB, the first auditory illusion benchmark for LALMs, covering ten representative illusions across music, sound, and speech, each annotated for the presence of knowledge-based priors. Our methodology pairs model evaluation with controlled human listening studies, enabling direct comparison of responses. Results show systematic differences: while most LALMs remain signal-faithful on low-level acoustic illusions, several exhibit more human-like responses when linguistic or musical priors are involved, although no model matches the human perceptual profile. These findings highlight the current limitations of LALMs as cognitive models. By establishing auditory illusions as a rigorous testbed, our work offers a new perspective for probing neural black-box models and advancing understanding of auditory cognition. AIB is publicly available at this https URL.


4. SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval

SonicCaps:用于改进音频检索的大规模多样化细粒度字幕数据集

AI 总结:该研究提出大规模音频字幕数据集SonicCaps,通过多模态大语言模型Qwen3-Omni生成约1500万条字幕,采用多字幕采样策略训练CLAP模型,提升了音频检索等任务的性能与泛化能力。

链接:https://arxiv.org/abs/2609.02343

机构:Sony CTC(索尼CTC); LTCI, Telecom Paris, Institut polytechnique de Paris(巴黎综合理工学院电信学院LTCI)

作者:Zineb Lahrichi, Marc Ferras, Gaël Richard, Geoffroy Peeters

英文摘要:Recent advances in audio-language modeling have been driven by large-scale audio captioning datasets. However, existing datasets remain limited by low semantic diversity, generic descriptions lacking acoustic details, and one-to-one audio-caption mappings that poorly reflect the inherent ambiguity of auditory perception. We introduce SonicCaps, a large-scale audio captioning dataset comprising ~15M captions paired with ~700k audio clips, generated using a multi-modal large language model (Qwen3-Omni) conditioned on both audio and text. To explicitly promote diversity, we generate around 24 captions per audio via structured prompt engineering and few- shot generation, spanning main descriptions, rephrased variants (verbosity, style) and semantic tags. Human evaluation shows that SonicCaps is rated significantly higher than existing captioning datasets, with fine-grained analyses indicating that our captions are perceived as more descriptive and precise, which strongly correlates with quality judgments. Finally, training CLAP models on SonicCaps with a multi-caption sampling strategy consistently improves audio retrieval and zero-shot classification, with stronger generalization across public and commercial benchmarks. We release both SonicCaps and two specialized CLAP models on hugging face: this https URL.


5. Efficient Passive Acoustic Monitoring of Killer Whales Using a Two-Stage Detection and Ecotype Classification Cascade

基于两阶段检测与生态型分类级联的虎鲸高效被动声学监测

AI 总结:本研究针对虎鲸被动声学监测的需求,提出基于ResNet的两阶段级联模型,在DCLDE 2027数据集上实现高精度检测与分类,可适配新声学域且推理速度快,适用于虎鲸保护的实时监测。

链接:https://arxiv.org/abs/2609.01792

机构:Microsoft AI for Good Research Lab(微软AI公益研究院); Universidad de los Andes(安第斯大学)

作者:Daniela Ruiz, Manuel Castellote, Zhongqi Miao, Carl Chalmers, Bruno Demuro, Rahul Dodhia, Pablo Arbelaez, Juan M. Lavista

英文摘要:Passive acoustic monitoring of killer whales is particularly important for conservation of the endangered Southern Resident killer whale population, but requires accurate models that can operate in real time under severe class imbalance and deployment shift. We propose a lightweight ResNet-based two-stage cascade that first detects killer whale vocalizations and then classifies confident detections into five eastern North Pacific ecotypes, abstaining on ambiguous calls. We train and evaluate the pipeline on the DCLDE 2027 dataset, where the detector achieves 0.960 macro-F1 and the classifier 0.958, outperforming frozen Perch 2.0 embeddings on the five-ecotype benchmark. By separating detection from ecotype recognition, the end-to-end cascade improves seven-class macro-F1 from 0.919 for a single-stage model to 0.933, with the largest gain on the rare OKW ecotype. To assess transfer beyond the benchmark, we use active learning to adapt the Stage 1 to the acoustic environment of Puget Sound, WA, increasing killer whale detection F1 from 0.405 to 0.755 on manually verified detection windows. Finally, each stage processes a 3 s window in approximately 1.4 ms on an NVIDIA H100, enabling faster than real time inference. These results demonstrate that the proposed two-stage cascade pipeline enables reliable killer whale detection and classification, adaptation to new acoustic domains, and real-time monitoring for conservation applications.


6. Understanding Automatic Mixing: A Subtask-Oriented Analysis of Two-Stage Mixing System

自动混音的理解:两阶段混音系统的面向子任务分析

AI 总结:本文通过三项受控听觉实验分析自动混音,探究两阶段系统性能提升的来源,发现不当分组会降低下游性能,两阶段变体优于单阶段基线,支持局部与全局混音分离的设计原则。

链接:https://arxiv.org/abs/2609.02835

作者:Jinjie Shi, Wei Hua, Kunzhu Xie, Make Li, Yuchen Liu, Joshua Reiss

英文摘要:Automatic mixing transforms multitrack recordings into perceptually coherent, balanced, and aesthetically consistent mixes. In real-world production, this task is challenging due to large track counts, diverse instrumentation, and strong inter-track dependencies. Two-stage systems address this complexity by separating intra-group processing from inter-group mixing, yet it remains unclear whether their gains arise from stronger component models or from explicit task decomposition. We present a subtask-oriented analysis of automatic mixing through three controlled listening experiments. We investigate whether full-mix models transfer to intra-group mixing, whether downstream models compensate for grouping and loudness errors, and whether two-stage decomposition improves full-mix quality. Across three dense pop and rock excerpts, transfer differs between the evaluated models; inappropriate grouping causes clear downstream degradation, while altered loudness relationships have weaker and model-dependent effects. Both two-stage variants significantly outperform their corresponding single-stage baselines. These findings support explicit separation of local balance and global mix coordination as a useful design principle for automatic mixing. Code and audio examples are available online.