今日论文合集:cs.SD语音10篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】SAMAY: System for Acoustic Measurement and Analysis
标题:SAMAY:声学测量和分析系统
链接:https://arxiv.org/abs/2512.13284

作者:Adheep Arya G R,Vaibhav Pratap Singh,Mayank Kumar,Niyathi Shenoy,Tejas Suryawanshi,Ruchi Juyal,Sangit Saha,Kaushik Nanda,Hari Babu Pasupuleti,S D Sudarsan
摘要:本文介绍了一个自动记录鸟鸣的系统SAMAY,它是通过建立一个数据库的大量鸟类声学数据来研究鸟类物种。通过分析记录的鸟类叫声数据,该系统还可用于自动分类鸟类,监测鸟类种群和分析环境变化的影响。该系统通过强大的STM32 F407系列微控制器驱动,支持4个麦克风,配备128 GB存储容量,并由10400 mAh电池组供电,该电池组与太阳能充电器连接。此外,该设备在运行期间可通过USB和Wi-Fi进行用户配置,确保在现场部署期间进行用户友好的操作。
摘要:This paper describes an automatic bird call recording system called SAMAY, which is developed to study bird species by creating a database of large amounts of bird acoustic data. By analysing the recorded bird call data, the system can also be used for automatic classification of bird species, monitoring bird populations and analysing the impact of environmental changes. The system is driven through a powerful STM32F407 series microcontroller, supports 4 microphones, is equipped with 128 GB of storage capacity, and is powered by a 10400 mAh battery pack interfaced with a solar charger. In addition, the device is user-configurable over USB and Wi-Fi during runtime, ensuring user-friendly operation during field deployment.


【2】DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
标题:DisCo-Speech:采用解纠缠语音编解码器的可控Zero-Shot语音生成
链接:https://arxiv.org/abs/2512.13251

作者:Tao Li,Wengshuo Ge,Zhichao Wang,Zihao Cui,Yong Ma,Yingying Gao,Chao Deng,Shilei Zhang,Junlan Feng
摘要:近年来基于编解码器的语言模型(LMs)的出现使文语转换(TTS)技术发生了革命性的变化。然而,由于标准编解码器紧密耦合音色和韵律,基于延续的LM不可避免地复制这种纠缠,阻碍了独立控制。最近的努力试图打破这种纠缠通过编解码器的设计,但不够去耦仍然是一个关键的瓶颈。为了解决这一挑战,我们提出了迪斯科语音,一个zero-shot可控的TTS框架,使韵律控制和语音克隆,通过解开语音编解码器(DisCodec)和LM为基础的发电机。核心组件DisCodec包含两个核心阶段:1)三因子解纠缠,其通过并行编码器和混合损失将语音显式分解为内容、韵律和音色子空间;以及2)融合和重建,其将内容和韵律融合为适合LM预测的统一内容-韵律令牌,同时联合优化重建质量以解决解纠缠-重建权衡。通过这种设计,LM从风格提示执行韵律延续,而解码器处理目标音色注入,从而实现灵活的zero-shot控制。实验表明,DisCo-Speech匹配国家的最先进的语音克隆性能,同时优于基线的zero-shot韵律控制。通过在编解码器级解决核心纠缠,DisCo-Speech为可控语音合成提供了一个强大的基础。音频示例可在https://github.com/disco-speech/DisCo-Speech上获得,代码和权重将在同一链接中发布。
摘要:Recent codec-based language models~(LMs) have revolutionized text-to-speech~(TTS). However, since standard codecs tightly couple timbre and prosody, continuation-based LMs inevitably replicate this entanglement, hindering independent control. Recent efforts attempt to break this entanglement via codec design, but insufficient decoupling remains a critical bottleneck. To tackle this challenge, we propose DisCo-Speech, a zero-shot controllable TTS framework that enables prosody control and voice cloning via a disentangled speech codec (DisCodec) and an LM-based generator. The core component, DisCodec, contains two core stages: 1) Tri-factor disentanglement, which explicitly factorizes speech into content, prosody, and timbre subspaces via parallel encoders and hybrid losses; and 2) Fusion and reconstruction, which fuses content and prosody into unified content-prosody tokens suitable for LM prediction, while jointly optimizing reconstruction quality to resolve the disentanglement-reconstruction trade-off. With this design, the LM performs prosodic continuation from a style prompt while the decoder handles target timbre injection, enabling flexible zero-shot control. Experiments show that DisCo-Speech matches state-of-the-art voice cloning performance while outperforming baselines in zero-shot prosody control. By resolving the core entanglement at the codec level, DisCo-Speech provides a robust foundation for controllable speech synthesis. Audio samples are available at https://github.com/disco-speech/DisCo-Speech, and the code and weights will be released at the same link.


【3】Towards Unified Co-Speech Gesture Generation via Hierarchical Implicit Periodicity Learning
标题:通过分层隐式周期性学习实现统一的共语音手势生成
链接:https://arxiv.org/abs/2512.13131

作者:Xin Guo,Yifan Zhao,Jia Li
备注:IEEE Transactions on Image Processing
摘要:从语音中生成基于3D的身体动作在广泛的下游应用中显示出巨大的潜力,但在模仿真实的人体动作方面仍然面临挑战。主要的研究工作集中在端到端生成方案,以生成共同语音手势,跨越GANs,VQ-VAE和最近的扩散模型。作为一个不适定的问题,在本文中,我们认为,这些流行的学习方案未能在不同的运动单元,即头部,身体和手部之间建立关键的内部和内部相关性模型,从而导致不自然的运动和协调性差。为了深入研究这些内在的相关性,我们提出了一个统一的层次隐式周期性(HIP)学习方法,用于音频启发的3D手势生成。与主流研究不同的是,我们的方法通过两个显式技术见解来建模这种多模态隐式关系:i)为了解开复杂的手势运动,我们首先探索具有周期性自编码器的手势运动相位流形,以从现实分布中模仿人类本性,同时将来自当前潜在状态的非周期性分布用于实例级别的离散化。ii)对面部运动、身体姿势和手部运动的层次关系进行建模,在学习期间利用级联指导来驱动动画。我们展示了我们提出的方法在3D化身和广泛的实验表明,我们的方法优于国家的最先进的协同语音手势生成方法的定量和定性评价。代码和模型将公开提供。
摘要:Generating 3D-based body movements from speech shows great potential in extensive downstream applications, while it still suffers challenges in imitating realistic human movements. Predominant research efforts focus on end-to-end generation schemes to generate co-speech gestures, spanning GANs, VQ-VAE, and recent diffusion models. As an ill-posed problem, in this paper, we argue that these prevailing learning schemes fail to model crucial inter- and intra-correlations across different motion units, i.e. head, body, and hands, thus leading to unnatural movements and poor coordination. To delve into these intrinsic correlations, we propose a unified Hierarchical Implicit Periodicity (HIP) learning approach for audio-inspired 3D gesture generation. Different from predominant research, our approach models this multi-modal implicit relationship by two explicit technique insights: i) To disentangle the complicated gesture movements, we first explore the gesture motion phase manifolds with periodic autoencoders to imitate human natures from realistic distributions while incorporating non-period ones from current latent states for instance-level diversities. ii) To model the hierarchical relationship of face motions, body gestures, and hand movements, driving the animation with cascaded guidance during learning. We exhibit our proposed approach on 3D avatars and extensive experiments show our method outperforms the state-of-the-art co-speech gesture generation methods by both quantitative and qualitative evaluations. Code and models will be publicly available.


【4】HQ-MPSD: A Multilingual Artifact-Controlled Benchmark for Partial Deepfake Speech Detection
标题:HQ-MPSD:用于部分Deepfake语音检测的多语言伪影控制基准
链接:https://arxiv.org/abs/2512.13012

作者:Menglu Li,Majd Alber,Ramtin Asgarianamiri,Lian Zhao,Xiao-Ping Zhang
备注:6 pages, 4 figures, 2 tables
摘要:检测部分深度伪造语音具有挑战性,因为操作仅发生在较短的区域中,而周围的音频仍然真实。然而,现有的检测方法从根本上受到可用数据集的质量的限制,其中许多依赖于过时的合成系统和生成程序,这些合成系统和生成程序引入了特定于机器人的伪影,而不是真实的操纵线索。为了解决这一差距,我们引入了HQ-MPSD,这是一个高质量的多语言部分deepfake语音数据集。HQ-MPSD是使用来自细粒度强制对齐的语言连贯拼接点构建的,保留了韵律和语义的连续性,并最大限度地减少了听觉和视觉边界伪影。该数据集包含8种语言和550个扬声器的350.8小时语音,并添加了背景效果以更好地反映真实世界的声学条件。MOS评估和频谱分析证实了样本的高感知自然度。我们通过跨语言和跨数据集评估对最先进的检测模型进行基准测试,所有模型在HQ-MPSD上的性能下降超过80%。这些结果表明,HQ-MPSD一旦去除了低级伪影并引入了多语言和声学多样性,就会面临重大的泛化挑战,为部分深度伪造检测提供了更现实和更苛刻的基准。数据集可以在https://zenodo.org/records/17929533上找到。
摘要:Detecting partial deepfake speech is challenging because manipulations occur only in short regions while the surrounding audio remains authentic. However, existing detection methods are fundamentally limited by the quality of available datasets, many of which rely on outdated synthesis systems and generation procedures that introduce dataset-specific artifacts rather than realistic manipulation cues. To address this gap, we introduce HQ-MPSD, a high-quality multilingual partial deepfake speech dataset. HQ-MPSD is constructed using linguistically coherent splice points derived from fine-grained forced alignment, preserving prosodic and semantic continuity and minimizing audible and visual boundary artifacts. The dataset contains 350.8 hours of speech across eight languages and 550 speakers, with background effects added to better reflect real-world acoustic conditions. MOS evaluations and spectrogram analysis confirm the high perceptual naturalness of the samples. We benchmark state-of-the-art detection models through cross-language and cross-dataset evaluations, and all models experience performance drops exceeding 80% on HQ-MPSD. These results demonstrate that HQ-MPSD exposes significant generalization challenges once low-level artifacts are removed and multilingual and acoustic diversity are introduced, providing a more realistic and demanding benchmark for partial deepfake detection. The dataset can be found at: https://zenodo.org/records/17929533.


【5】Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
标题:Schrodinger视听编辑器:对象级视听删除
链接:https://arxiv.org/abs/2512.12875

作者:Weihan Xu,Kan Jen Cheng,Koichi Saito,Muhammad Jehanzeb Mirza,Tingle Li,Yisi Liu,Alexander H. Liu,Liming Wang,Masato Ishii,Takashi Shibuya,Yuki Mitsufuji,Gopala Anumanchipalli,Paul Pu Liang
摘要:音频和视频内容的联合编辑对于精确和可控的内容创建至关重要。由于目标编辑前后配对视听数据的局限性以及不同模态之间的异质性,这项新任务带来了挑战。为了解决联合视听编辑中的数据和建模挑战,我们引入了SAVEBench,这是一个带有文本和掩码条件的配对视听数据集,可以实现基于对象的源到目标学习。借助SAVEBench,我们训练了Schrodinger Audio-Visual Editor(SAVE),这是一种端到端流匹配模型,可并行编辑音频和视频,同时在整个处理过程中保持它们的一致性。SAVE结合了一个薛定谔桥,学习从源到目标视听混合的直接传输。我们的评估表明,所提出的SAVE模型是能够删除音频和视频内容中的目标对象,同时保留剩余的内容,具有更强的时间同步和视听语义对应的音频编辑器和视频编辑器的成对组合相比。
摘要:Joint editing of audio and visual content is crucial for precise and controllable content creation. This new task poses challenges due to the limitations of paired audio-visual data before and after targeted edits, and the heterogeneity across modalities. To address the data and modeling challenges in joint audio-visual editing, we introduce SAVEBench, a paired audiovisual dataset with text and mask conditions to enable object-grounded source-to-target learning. With SAVEBench, we train the Schrodinger Audio-Visual Editor (SAVE), an end-to-end flow-matching model that edits audio and video in parallel while keeping them aligned throughout processing. SAVE incorporates a Schrodinger Bridge that learns a direct transport from source to target audiovisual mixtures. Our evaluation demonstrates that the proposed SAVE model is able to remove the target objects in audio and visual content while preserving the remaining content, with stronger temporal synchronization and audiovisual semantic correspondence compared with pairwise combinations of an audio editor and a video editor.


【6】Procedural Music Generation Systems in Games
标题:游戏中的程序音乐生成系统
链接:https://arxiv.org/abs/2512.12834

作者:Shangxuan Luo,Joshua Reiss
摘要:程序音乐生成(PMG)是一个新兴的领域,它通过算法为视频游戏创建音乐内容。通过利用从简单的基于规则的方法到先进的机器学习算法的技术,PMG有可能显著提高开发效率,提供更丰富的音乐体验,并增强玩家的沉浸感。然而,由于新颖性、可靠性和分配的资源等优先级的差异,学术原型往往与应用程序不同。本文弥合了研究和应用之间的差距,提出了一个系统的概述,目前PMG技术在这两个领域,提供了两个方面的分类。通过比较分析,本研究确定了算法实现,音乐质量和游戏集成的关键研究挑战。最后,本文概述了未来的研究方向,强调面向任务和上下文感知的设计,更全面的质量评估方法,并改进研究工具的整合,为开发人员,作曲家和研究人员寻求推进PMG在游戏环境中提供可操作的见解。
摘要:Procedural Music Generation (PMG) is an emerging field that algorithmically creates music content for video games. By leveraging techniques from simple rule-based approaches to advanced machine learning algorithms, PMG has the potential to significantly improve development efficiency, provide richer musical experiences, and enhance player immersion. However, academic prototypes often diverge from applications due to differences in priorities such as novelty, reliability, and allocated resources. This paper bridges the gap between research and applications by presenting a systematic overview of current PMG techniques in both fields, offering a two-aspect taxonomy. Through a comparative analysis, this study identifies key research challenges in algorithm implementation, music quality and game integration. Finally, the paper outlines future research directions, emphasising task-oriented and context-aware design, more comprehensive quality evaluation methods, and improved research tool integration to provide actionable insights for developers, composers, and researchers seeking to advance PMG in game contexts.


【7】Adaptive Edge-Cloud Inference for Speech-to-Action Systems Using ASR and Large Language Models (ASTA)
标题:使用ASB和大型语言模型(ASTA)的语音到动作系统的自适应边缘云推理
链接:https://arxiv.org/abs/2512.12769

作者:Mohammad Jalili Torkamani,Israt Zarin
备注:preprint, 6 pages, 7 figures, 1 table
摘要:基于语音的交互已成为控制物联网设备的自然和直观的方式。然而,语音驱动的边缘设备面临着基于云的解决方案和基于边缘的解决方案之间的根本权衡,基于云的解决方案以延迟,连接依赖和隐私问题为代价提供更强的语言理解能力,基于边缘的解决方案提供低延迟和改善的隐私,但受到计算约束的限制。ASTA是一种自适应语音到动作解决方案,可在边缘和云推理之间动态路由语音命令,以平衡性能和系统资源利用率。ASTA将设备上的自动语音识别和轻量级离线语言模型推理与基于云的LLM处理集成在一起,由CPU工作负载、设备温度和网络延迟等实时系统指标指导。度量感知路由机制在运行时选择推理路径,而基于规则的命令验证和修复组件确保成功的端到端命令执行。我们在基于NVIDIA Jetson的边缘平台上实施了我们的解决方案,并使用包含80个语音命令的多样化数据集对其进行了评估。实验结果表明,ASTA成功地路由所有输入命令的执行,实现了在线和离线推理之间的平衡分布。该系统达到了62.5%的ASR准确率,并生成可执行命令,而无需修复只有47.5%的输入,突出了修复机制在提高鲁棒性的重要性。这些结果表明,自适应边缘云编排是弹性和资源感知语音控制物联网系统的可行方法。
摘要:Voice-based interaction has emerged as a natural and intuitive modality for controlling IoT devices. However, speech-driven edge devices face a fundamental trade-off between cloud-based solutions, which offer stronger language understanding capabilities at the cost of latency, connectivity dependence, and privacy concerns, and edge-based solutions, which provide low latency and improved privacy but are limited by computational constraints. This paper presents ASTA, an adaptive speech-to-action solution that dynamically routes voice commands between edge and cloud inference to balance performance and system resource utilization. ASTA integrates on-device automatic speech recognition and lightweight offline language-model inference with cloud-based LLM processing, guided by real-time system metrics such as CPU workload, device temperature, and network latency. A metric-aware routing mechanism selects the inference path at runtime, while a rule-based command validation and repair component ensures successful end-to-end command execution. We implemented our solution on an NVIDIA Jetson-based edge platform and evaluated it using a diverse dataset of 80 spoken commands. Experimental results show that ASTA successfully routes all input commands for execution, achieving a balanced distribution between online and offline inference. The system attains an ASR accuracy of 62.5% and generates executable commands without repair for only 47.5% of inputs, highlighting the importance of the repair mechanism in improving robustness. These results suggest that adaptive edge-cloud orchestration is a viable approach for resilient and resource-aware voice-controlled IoT systems.


【8】Privacy-Aware Ambient Audio Sensing for Healthy Indoor Spaces
标题:隐私意识的环境音频传感,促进健康的室内空间
链接:https://arxiv.org/abs/2512.12471

作者:Bhawana Chhaglani
摘要:室内空气传播会造成重大的健康风险,但目前的监测解决方案是侵入性的,昂贵的,或无法直接解决它。我的研究探索了环境音频传感的未开发潜力,以非侵入性和实时地估计关键的传播风险因素,如通风,气溶胶排放和乘员分布。我开发了隐私保护系统,利用现有的麦克风来监测室内空气质量的整个频谱,这可能对个人的健康产生重大影响。这项工作为使用日常设备进行隐私感知的空气传播风险监测奠定了基础。
摘要:Indoor airborne transmission poses a significant health risk, yet current monitoring solutions are invasive, costly, or fail to address it directly. My research explores the untapped potential of ambient audio sensing to estimate key transmission risk factors such as ventilation, aerosol emissions, and occupant distribution non-invasively and in real time. I develop privacy-preserving systems that leverage existing microphones to monitor the whole spectrum of indoor air quality which can have a significant effect on an individual's health. This work lays the foundation for privacy-aware airborne risk monitoring using everyday devices.


【9】AutoMV: An Automatic Multi-Agent System for Music Video Generation
标题:AutoMV:一个用于音乐视频生成的自动多代理系统
链接:https://arxiv.org/abs/2512.12196

作者:Xiaoxuan Tang,Xinping Lei,Chaoran Zhu,Shiyun Chen,Ruibin Yuan,Yizhi Li,Changjae Oh,Ge Zhang,Wenhao Huang,Emmanouil Benetos,Yang Liu,Jiaheng Liu,Yinghao Ma
摘要:完整长度歌曲的音乐到视频(M2 V)生成面临重大挑战。现有的方法产生短的、不连贯的片段,无法将视觉效果与音乐结构、节拍或歌词对齐,并且缺乏时间一致性。我们提出了AutoMV,一个多代理系统,直接从一首歌生成完整的音乐视频(MV)。AutoMV首先应用音乐处理工具来提取音乐属性,如结构,声乐曲目和时间对齐的歌词,并将这些特征构造为以下代理的上下文输入。编剧代理和导演代理然后使用该信息来设计短脚本,在共享的外部库中定义角色配置文件,并指定摄像机指令。随后,这些代理调用关键帧的图像生成器和“故事”或“歌手”场景的不同视频生成器。一个验证代理评估他们的输出,使多代理合作,以产生一个连贯的长篇MV。为了评估M2 V生成,我们进一步提出了一个基准,其中包括四个高级类别(音乐内容,技术,后期制作,艺术)和十二个线粒度标准。该基准被应用于将商业产品、AutoMV和人工指导MV与专家人工评分员进行比较:AutoMV在所有四个类别中的表现均显着优于当前基线,缩小了与专业MV的差距。最后,我们研究了使用大型多模态模型作为自动MV判断;虽然有希望,但它们仍然落后于人类专家,突出了未来工作的空间。
摘要:Music-to-Video (M2V) generation for full-length songs faces significant challenges. Existing methods produce short, disjointed clips, failing to align visuals with musical structure, beats, or lyrics, and lack temporal consistency. We propose AutoMV, a multi-agent system that generates full music videos (MVs) directly from a song. AutoMV first applies music processing tools to extract musical attributes, such as structure, vocal tracks, and time-aligned lyrics, and constructs these features as contextual inputs for following agents. The screenwriter Agent and director Agent then use this information to design short script, define character profiles in a shared external bank, and specify camera instructions. Subsequently, these agents call the image generator for keyframes and different video generators for "story" or "singer" scenes. A Verifier Agent evaluates their output, enabling multi-agent collaboration to produce a coherent longform MV. To evaluate M2V generation, we further propose a benchmark with four high-level categories (Music Content, Technical, Post-production, Art) and twelve ine-grained criteria. This benchmark was applied to compare commercial products, AutoMV, and human-directed MVs with expert human raters: AutoMV outperforms current baselines significantly across all four categories, narrowing the gap to professional MVs. Finally, we investigate using large multimodal models as automatic MV judges; while promising, they still lag behind human expert, highlighting room for future work.


【10】A comparative study of generative models for child voice conversion
标题:童声转换生成模型的比较研究
链接:https://arxiv.org/abs/2512.12129

作者:Protima Nomo Sudro,Anton Ragni,Thomas Hain
备注:6 pages, 5 figures
摘要:生成模型是成人到成人语音转换(VC)的热门选择,因为它们可以有效地对未标记的数据进行建模。在这一点上,他们的有用性,在生产儿童的讲话,特别是成人对儿童VC尚未调查。对于成人到儿童VC,比较了四种生成模型:扩散模型,基于流的模型,变分自编码器和生成对抗网络。结果表明,虽然这些模型产生的转换后的语音输出似乎是合理的,他们表现出与目标说话人特征的相似性不足。我们介绍了一种有效的频率扭曲技术,可以应用于模型的输出,并显示出显着减少成人和儿童之间的不匹配。所有模型的输出都使用客观和主观指标进行评估。特别是,我们比较特定的扬声器配对使用一个独特的语料库收集的儿童语音配音。
摘要:Generative models are a popular choice for adult-to-adult voice conversion (VC) because of their efficient way of modelling unlabelled data. To this point their usefulness in producing children speech and in particular adult to child VC has not been investigated. For adult to child VC, four generative models are compared: diffusion model, flow based model, variational autoencoders, and generative adversarial network. Results show that although converted speech outputs produce by those models appear plausible, they exhibit insufficient similarity with the target speaker characteristics. We introduce an efficient frequency warping technique that can be applied to the output of models, and which shows significant reduction of the mismatch between adult and child. The output of all the models are evaluated using both objective and subjective measures. In particular we compare specific speaker pairing using a unique corpus collected for dubbing of children speech.


eess.AS音频处理


【1】REVERB-FL: Server-Side Adversarial and Reserve-Enhanced Federated Learning for Robust Audio Classification
标题:REVERB-FL:用于稳健音频分类的服务器端对抗和保留增强联邦学习
链接:https://arxiv.org/abs/2512.13647

作者:Sathwika Peechara,Rajeev Sahay
备注:13 pages, 4 figures
摘要:联邦学习(FL)为音频分类提供了一种隐私保护的训练范例,但对客户端异构性和中毒攻击高度敏感,其中受到不利影响的客户端可能会使全局模型产生偏差,并阻碍音频分类器的性能。为了减轻模型中毒对音频信号分类的影响,我们提出了REVERB-FL,这是一种轻量级的服务器端防御,它将一个小的储备集(约5%)与聚合前和聚合后的再训练和对抗训练相结合。在每个本地训练轮之后,服务器使用干净的或额外的逆向扰动数据在储备集上改进全局模型,从而抵消非IID漂移并减轻潜在的模型中毒,而不会增加大量的客户端成本或改变聚合过程。我们从理论上证明了我们的框架的可行性,表现出更快的收敛速度和减少稳态误差相对于基线联邦平均。我们在两个具有不同IID和Dirichlet非IID分区的开源音频分类数据集上验证了我们的框架,并证明了REVERB-FL在多种局部数据中毒设计下减轻了全局模型中毒。
摘要:Federated learning (FL) enables a privacy-preserving training paradigm for audio classification but is highly sensitive to client heterogeneity and poisoning attacks, where adversarially compromised clients can bias the global model and hinder the performance of audio classifiers. To mitigate the effects of model poisoning for audio signal classification, we present REVERB-FL, a lightweight, server-side defense that couples a small reserve set (approximately 5%) with pre- and post-aggregation retraining and adversarial training. After each local training round, the server refines the global model on the reserve set with either clean or additional adversarially perturbed data, thereby counteracting non-IID drift and mitigating potential model poisoning without adding substantial client-side cost or altering the aggregation process. We theoretically demonstrate the feasibility of our framework, showing faster convergence and a reduced steady-state error relative to baseline federated averaging. We validate our framework on two open-source audio classification datasets with varying IID and Dirichlet non-IID partitions and demonstrate that REVERB-FL mitigates global model poisoning under multiple designs of local data poisoning.


【2】BUT Systems for WildSpoof Challenge: SASV in the Wild
标题:但WildSpoof挑战系统:野外SASV
链接:https://arxiv.org/abs/2512.12851

作者:Junyi Peng,Jin Li,Johan Rohdin,Lin Zhang,Miroslav Hlaváček,Oldrich Plchot
备注:4 pages
摘要:本文介绍了BUT提交的WildSpoof挑战,重点是欺骗鲁棒自动说话人确认(SASV)的轨道。我们提出了一个SASV框架,旨在弥合一般音频理解和专业语音分析之间的差距。我们的子系统集成了各种各样的自我监督学习前端,从一般的音频模型(例如,大圣)到语音专用编码器(例如,WavLM)。这些表示是通过一个轻量级的多头分解注意力后端为相应的子任务聚合。此外,我们引入了一种基于分布不确定性的特征域增强策略,以明确建模和减轻由看不见的神经声码器和记录环境引起的域偏移。通过融合这些强大的CM分数与最先进的ASV系统,我们的方法实现了卓越的最小化的a-DCF和EER。
摘要:This paper presents the BUT submission to the WildSpoof Challenge, focusing on the Spoofing-robust Automatic Speaker Verification (SASV) track. We propose a SASV framework designed to bridge the gap between general audio understanding and specialized speech analysis. Our subsystem integrates diverse Self-Supervised Learning front-ends ranging from general audio models (e.g., Dasheng) to speech-specific encoders (e.g., WavLM). These representations are aggregated via a lightweight Multi-Head Factorized Attention back-end for corresponding subtasks. Furthermore, we introduce a feature domain augmentation strategy based on Distribution Uncertainty to explicitly model and mitigate the domain shift caused by unseen neural vocoders and recording environments. By fusing these robust CM scores with state-of-the-art ASV systems, our approach achieves superior minimization of the a-DCFs and EERs.


【3】Comparison of Classification Algorithms for COVID19 Detection using Cough Acoustic Signals
标题:使用咳嗽声信号检测COVID 19的分类算法比较
链接:https://arxiv.org/abs/2201.04872

作者:Yunus Emre Erdoğan,Ali Narin
备注:6 pages,3 figures,conference
摘要:这种被称为新型冠状病毒(COVID 19)的流行病于2019年12月首次在中国武汉发生。COVID 19不久后被世界卫生组织宣布为流行病。这种疾病的一些症状是发烧、咳嗽、呼吸急促和呼吸困难。在更严重的情况下,可能会因感染而死亡。抗击疫情和控制疫情的最重要问题是COVID 19(+)患者的早期诊断和这些患者的随访。因此,使用各种诊断机制。除了RT-PCR测试,医学成像方法也被利用,特别是在COVID 19(+)患者的检测中。在这项研究中,通过使用咳嗽数据提出了一种替代方法,咳嗽是COVID 19(+)患者最突出的症状之一。使用Virufy网站上的咳嗽声学公共数据集。使用z归一化技术对整个数据进行归一化。通过5层经验模式分解方法获得的特征的性能和不同的分类器的性能进行了比较。作为分类器算法,使用了5种不同的算法。使用Ensemble-Bagged-Trees算法获得了最高的准确率和F1-score性能,分别为90.6%和90.5%。另一方面,研究中使用的其他分类算法分别是支持向量机,逻辑回归,线性判别分析和k-最近邻。根据所获得的结果,选择正确的分类器算法提供高的结果。因此,很明显,使用咳嗽声学数据,可以容易且有效地检测具有COVID 19(+)的那些。
摘要:The epidemic disease, called the new coronavirus (COVID19), firstly occurred in Wuhan, China in December 2019. COVID19 was announced as an epidemic by World Health Organization soon after. Some of the symptoms of this disease are fever, cough, shortness of breath and difficulty in breathing. In more severe cases, death may occur as a result of infection. The most significant question in fighting the pandemic and controlling the epidemic is the early diagnosis of COVID19(+) patients and the follow-up of these patients. Therefore, various diagnostic mechanisms are used. Additionally to the RT-PCR test, medical imaging methods have been utilized, especially in the detection of COVID19(+) patients. In this study, an alternative approach was proposed by using cough data, which is one of the most prominent symptoms of COVID19(+) patients. The cough acoustic public dataset on the Virufy website was used. The entire data was normalized using z-normalization technique. The performance of the features obtained via the 5-layer empirical mode decomposition method and the performances of different classifiers has been compared. As the classifier algorithm, 5 different algorithms were used. The highest accuracy and F1-score performances were obtained by using Ensemble-Bagged-Trees algorithm as 90.6% and 90.5%, respectively. On the other hand, other classification algorithms used in the study are Support Vector Machines, Logistic Regression, Linear Discriminant Analysis and k-Nearest Neigbors, respectively. According to the results obtained, choosing the right classifier algorithm provides high results. Thus, it is clear that using cough acoustic data, those with COVID19(+) can be detected easily and effectively.


【4】AutoMV: An Automatic Multi-Agent System for Music Video Generation
标题:AutoMV:一个用于音乐视频生成的自动多代理系统
链接:https://arxiv.org/abs/2512.12196

作者:Xiaoxuan Tang,Xinping Lei,Chaoran Zhu,Shiyun Chen,Ruibin Yuan,Yizhi Li,Changjae Oh,Ge Zhang,Wenhao Huang,Emmanouil Benetos,Yang Liu,Jiaheng Liu,Yinghao Ma
摘要:完整长度歌曲的音乐到视频(M2 V)生成面临重大挑战。现有的方法产生短的、不连贯的片段,无法将视觉效果与音乐结构、节拍或歌词对齐,并且缺乏时间一致性。我们提出了AutoMV,一个多代理系统,直接从一首歌生成完整的音乐视频(MV)。AutoMV首先应用音乐处理工具来提取音乐属性,如结构,声乐曲目和时间对齐的歌词,并将这些特征构建为以下代理的上下文输入。编剧代理和导演代理然后使用该信息来设计短脚本,在共享的外部库中定义角色配置文件,并指定摄像机指令。随后,这些代理调用关键帧的图像生成器和“故事”或“歌手”场景的不同视频生成器。一个验证代理评估他们的输出,使多代理合作,以产生一个连贯的长篇MV。为了评估M2 V生成,我们进一步提出了一个基准,其中包括四个高级类别(音乐内容,技术,后期制作,艺术)和十二个线粒度标准。该基准被应用于将商业产品、AutoMV和人工指导MV与专家人工评分员进行比较:AutoMV在所有四个类别中的表现均显著优于当前基线,缩小了与专业MV的差距。最后,我们研究了使用大型多模态模型作为自动MV判断;虽然有希望,但它们仍然落后于人类专家,突出了未来工作的空间。
摘要:Music-to-Video (M2V) generation for full-length songs faces significant challenges. Existing methods produce short, disjointed clips, failing to align visuals with musical structure, beats, or lyrics, and lack temporal consistency. We propose AutoMV, a multi-agent system that generates full music videos (MVs) directly from a song. AutoMV first applies music processing tools to extract musical attributes, such as structure, vocal tracks, and time-aligned lyrics, and constructs these features as contextual inputs for following agents. The screenwriter Agent and director Agent then use this information to design short script, define character profiles in a shared external bank, and specify camera instructions. Subsequently, these agents call the image generator for keyframes and different video generators for "story" or "singer" scenes. A Verifier Agent evaluates their output, enabling multi-agent collaboration to produce a coherent longform MV. To evaluate M2V generation, we further propose a benchmark with four high-level categories (Music Content, Technical, Post-production, Art) and twelve ine-grained criteria. This benchmark was applied to compare commercial products, AutoMV, and human-directed MVs with expert human raters: AutoMV outperforms current baselines significantly across all four categories, narrowing the gap to professional MVs. Finally, we investigate using large multimodal models as automatic MV judges; while promising, they still lag behind human expert, highlighting room for future work.


机器翻译由腾讯交互翻译提供,仅供参考