微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 说话人识别、验证与分离 2 篇
2. 音频事件检测与场景理解 2 篇
3. 音乐信息检索与音乐生成 1 篇
4. 安全、隐私与深度伪造音频 1 篇
5. 其他/综合语音音频 2 篇
1. 说话人识别、验证与分离 | 2 篇
1. Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings
基于WavLM和数据增强的自然非校准语音录音中的说话者声音力度分类
AI 总结:提出使用WavLM结合数据增强和高斯邻域软标签方法,在AVID语料库上实现78.2%的平均准确率,达到新最优。
链接:https://arxiv.org/abs/2606.27543
机构:University of Texas at Dallas (UTDallas)(德克萨斯大学达拉斯分校)
作者:Zahra Omidi, John H. L. Hansen
英文摘要:The variations in vocal effort range (e.g. whisper, soft, neutral, loud, shout) alter production and speech acoustics, reducing intelligibility and limiting the robustness of any subsequent speech technology. Classification is challenging since effort lies on a continuum, adjacent categories are easily confused, and labeled data remain scarce. Prior SSL approaches with wav2vec2, HuBERT, and AST improve performance on the AVID corpus but still suffer from boundary errors. In this study, we introduce WavLM for the first time in vocal effort classification and benchmark it against wav2vec2 and HuBERT. To address data scarcity, we conduct a systematic study of augmentation strategies, covering RIR convolution, additive noise, time masking, speed perturbation, band-limiting, MixUp, and CutMix. Augmentation consistently improves WavLM, with gains ranging from +0.6% to +1.8% absolute. We further propose Gaussian-neighbor soft labels, which further reduce near-boundary confusions by modeling the vocal effort continuum. Our best system, WavLM-BASE with gradual unfreezing, augmentation, and Gaussian-neighbor soft labels, achieves 78.2% mean accuracy, establishing a new state-of-the-art on AVID.
2. DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions
DG^VoiC:真实呼叫中心条件下的欺诈调查说话人聚类
AI 总结:提出DG^VoiC框架,结合匿名化、语音预处理、滑动窗口嵌入提取和余弦相似度聚类,在真实呼叫中心音频中识别重复说话人,用于欺诈调查,在121条录音上达到96% AMI等高性能。
链接:https://arxiv.org/abs/2606.28048
作者:Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth, Karishma Jaitly
英文摘要:Insurance fraud remains costly and operationally difficult, particularly in call-centre workflows where many customer interactions begin at FNOL. While recent fraud detection methods mainly rely on structured data, text, or images, repeated speaker identity across calls remains underused as an investigative signal. This paper presents DG^VoiC, a voice clustering framework for customer verification and cross-profile speaker linking on anonymised real call-centre audio. The approach combines sensitive information-aligned anonymisation, speech-focused preprocessing, sliding-window speaker embedding extraction, and cosine similarity based clustering to identify repeated speakers under real telephony conditions. The method was evaluated on 121 recordings, with a curated reference subset of 56 samples in 22 human-agreed speaker clusters. used for validation. The best configuration achieved 96% AMI, 95% ARI, 98% completeness, 100% homogeneity, and 99% V-measure. These results show that speaker clustering can provide a strong additional signal for fraud investigation by helping analysts verify speaker consistency and surface repeated voices across customers.
2. 音频事件检测与场景理解 | 2 篇
3. From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection
从通用音频标记到空间定位的声音事件定位与检测
AI 总结:提出AT2SELD框架,利用预训练通用音频标记骨干网络,结合一阶声场空间处理、逐轨迹声音事件检测和笛卡尔到达方向估计,通过多阶段神经架构搜索实现语义到空间的迁移,并在多个数据集上验证其有效性。
链接:https://arxiv.org/abs/2606.27751
作者:Stefano Giacomelli, Stefano Damiano, Claudia Rinaldi, Fabio Graziosi, Toon van Waterschoot
英文摘要:This report investigates the extension of pretrained General-Purpose Audio Tagging (GP-AT) models toward spatially grounded Sound Event Localization and Detection (SELD). The proposed AT2SELD framework couples a pretrained AT backbone with compact First-Order Ambisonics (FOA) spatial processing, track-wise SED and Cartesian DOA estimation, permutation aware supervision, and calibration. It characterizes how semantic audio priors support localization-aware scene analysis under data, computation, and deployment constraints. The framework is developed through informed multi-stage Neural Architecture Search (NAS). Stage 1 shows that spectral FOA descriptors, based on magnitude, phase, and Intensity Vectors (IVs), provide the most reliable interface for semantic-to-spatial transfer. Stage 2 identifies early residual spatial encoding as the main capacity-sensitive component, while late track-wise abstraction and recurrent smoothing act mainly as refinement stages. Stage 3 shows that late cross-stitch coupling improves semantic-spatial interaction, whereas early fusion is costlier and less effective. Diagnostic evaluation analyzes the selected architecture under class balancing, focal loss, activity-conditioned DOA supervision, threshold calibration, and transfer across STARSS23, TAU2019, TAU-NIGENS2020, and TAU-NIGENS2021. Focal loss improves the activity point, active-only DOA supervision mitigates inactive target dominance, and validation-selected thresholds recover calibration without replacing spatial learning. Cross-dataset and oracle-activity analyses indicate strong fixed source localization on TAU2019, transferable representations from TAU NIGENS2021, and meaningful but uncertain behavior on STARSS23. Overall, GP-AT priors appear promising for SELD design when embedded in spatial-aware architectures and optimized through integrated calibration and deployment oriented strategies.
4. Grammar-Guided Hierarchical Parsing for Long-form Audio Activity Recognition
语法引导的长格式音频活动识别分层解析
AI 总结:针对长格式音频的层次结构,提出基于事件证据的分层解析方法,利用层次活动语法进行语法引导解码,无需子活动或活动标签即可生成时间对齐的解析树,提升时序一致性并产生可解释的层次结构。
链接:https://arxiv.org/abs/2606.27965
机构:Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey, U.K.(英国萨里大学视觉、语音与信号处理中心(CVSSP))
作者:Peng Zhang, Qingyu Luo, Philip J.B. Jackson, Wenwu Wang
英文摘要:Long-form audio exhibits an inherent hierarchy: fine-grained events form sub-activities, which in turn constitute higher-level activities. Prior work often models these levels separately, leading to cross-level inconsistencies and requiring supervision at multiple levels. We formulate the problem as hierarchical parsing from event-level evidence: given detected event segments with class posteriors, we infer an order-consistent Act-Sub-Event parse tree. We propose Hierarchical Activity Grammar, encoding hierarchical composition and temporal-order constraints, and perform grammar-guided decoding that combines event evidence with a grammar prior. This yields a temporally grounded parse tree from which sub-activity segmentation and activity classification are derived, without requiring sub-activity or activity labels for training. Experiments on the long-form MultiAct audio dataset demonstrate improved temporal-order consistency (Edit score) and produces interpretable hierarchies.
3. 音乐信息检索与音乐生成 | 1 篇
5. A Flexible Encoding Model for Non-Unique Note Alignments
非唯一音符对齐的灵活编码模型
AI 总结:针对排练重复和通奏低音即兴等需要非唯一对齐的音乐实践,提出Match文件格式的向后兼容扩展,通过虚拟指针音符和扩展的section行支持多链接与语义注释。
链接:https://arxiv.org/abs/2606.28032
作者:Suhit Chiruthapudi, Adam Štefunko, Silvan Peter, Patricia Hu, Jan Hajič jr., Carlos Eduardo Cancino-Chacón
英文摘要:Symbolic music alignment links notes in a symbolic performance to their counterparts in a score. While existing alignment encoding formats provide unique correspondences between these notes, there are various musical practices and forms such as practice repetitions in rehearsal and improvised realizations in basso continuo that require a more flexible approach to encoding their alignments. In this paper, we propose a minimal, backward-compatible extension to the Match file format to support such non-unique and semantically complex alignments. We introduce two virtual pointer notes - virtual score notes and virtual performance notes - which allow to encode multiple links between performance and score notes. In addition we expand the Match file's 'section' line to include semantically meaningful annotations of performance regions beyond score-indicated musical repetitions. We further demonstrate the utility of these extensions through two representative use-cases in piano rehearsal and basso continuo.
4. 安全、隐私与深度伪造音频 | 1 篇
6. Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks
容错空间:空中声学攻击的大规模仿真
AI 总结:针对语音控制系统面临的声学攻击风险,提出大规模仿真框架,通过800万次对抗评估揭示声学因素对攻击效果的影响,并引入双形式信噪比解耦隐蔽性与攻击效能。
链接:https://arxiv.org/abs/2606.27701
机构:University of Melbourne(墨尔本大学); DST Group(国防科学技术集团)
作者:Andrew C. Cullen, Neil Marchant, Jiani Xie, Paul Montague, Benjamin I.P. Rubinstein
英文摘要:While voice control is rapidly becoming a ubiquitous vector of human-AI communication, the risks facing these systems remain poorly understood. This is, in part, a product of the difficulties in scaling strictly digital adversarial workflows to the physical world. These scale barriers have led the community to abstract away key acoustic factors relating to detectability and the influence of geometry on acoustics. These methodological and metrological shortcomings undermine our understanding of risk. We illuminate these issues through real-world testing, conceptual discussions, and a novel, high-throughput reality simulation framework. By testing over 8 million adversarial evaluations, we demonstrate that acoustic awareness yields relative Word Error Rate increases of up to 94.5\% under Whisper and wav2vec. We employ this framework to explore a formalize and operationalize a Dual-Form Signal to Noise Ratio to decouple source stealth from victim attack efficacy, resolving a crucial limitation in current works. This lays the groundwork for repeatable, verifiable research that embraces, rather than abstracts, the acoustic environment.
5. 其他/综合语音音频 | 2 篇
7. Learning from Annotation Uncertainty: Entropy-Aware Curriculum for Speech Emotion Recognition
从标注不确定性中学习:面向语音情感识别的熵感知课程
AI 总结:针对语音情感识别中标注者分歧问题,提出基于分布的监督方法,利用WavLM-Base多任务模型在9类情感上减少与人类投票分布的差异,并通过熵分层评估证明分布监督能更好地捕捉感知不确定性。
链接:https://arxiv.org/abs/2606.27536
机构:Center for Robust Speech Systems, The University of Texas at Dallas(德克萨斯大学达拉斯分校鲁棒语音系统中心)
作者:Zahra Omidi, John H.L. Hansen
英文摘要:Speech emotion recognition (SER) often relies on hard consensus labels that collapse annotator disagreement. We study distribution-based supervision for 9-class SER on MSP-Podcast 2.0 using a WavLM-Base multitask model for categorical emotion and dimensional VAD. Hard-label training is compared with targets from primary and merged primary--secondary annotator vote distributions. Distributional objectives improve alignment with human vote distributions, reducing JSD/KLD relative to hard-label training. Analysis shows that hard supervision partly benefits from assigning ambiguous utterances to the residual Other class, whereas distributional supervision redistributes uncertainty across emotion categories. Entropy-stratified evaluation shows that high-ambiguity utterances remain challenging, but distribution-based supervision better captures perceptual uncertainty. These findings support moving beyond hard labels toward targets that reflect listener disagreement.
8. Elastic Time: Dynamic Frame Rate Bottlenecks for Neural Audio Coding
弹性时间:神经音频编码的动态帧率瓶颈
AI 总结:提出弹性时间方法,通过轻量级潜在预测器实现动态帧率瓶颈,将固定帧率自编码器转换为动态帧率,在推理时高效选择边界,提升效率-质量权衡。
链接:https://arxiv.org/abs/2606.27320
机构:University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校); Massachusetts Institute of Technology(麻省理工学院)
作者:Dimitrios Bralios, Paris Smaragdis, Minje Kim
英文摘要:Neural audio autoencoders have become a core component of compression, feature extraction, and generation. However, while existing systems support variable bitrate, the vast majority of models still operate at a fixed latent frame-rate, allocating equal temporal budget to regions with very different information density, which can result in unnecessarily long sequences. We introduce Elastic Time, a dynamic frame-rate bottleneck that converts fixed-frame-rate autoencoders to dynamic ones. Our method learns a lightweight latent predictor used to decide which frames can be skipped and later reconstructed, enabling efficient greedy boundary selection at inference. Experiments show our method enables deployment-time rate control while improving efficiency-quality tradeoffs relative to baselines. Overall, we provide a flexible mechanism for adjusting temporal resolution in audio autoencoders, potentially facilitating more efficient downstream modeling for generation and long-context tasks.
