今日论文合集:cs.SD语音9篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Graph Embedding with Mel-spectrograms for Underwater Acoustic Target Recognition
标题:利用Mel谱图嵌入的水下声目标识别
链接:https://arxiv.org/abs/2512.11545

作者:Sheng Feng,Shuqing Ma,Xiaoqian Zhu
摘要:由于舰船辐射噪声的复杂性和海洋环境的多变性,水声目标识别具有极大的挑战性。虽然深度学习(DL)方法已经取得了可喜的成果,但大多数现有模型都隐含地假设水下声学数据位于欧几里得空间中。然而,这种假设是不适合的固有复杂的拓扑结构的水声信号,表现出非平稳,非高斯和非线性特性。为了克服这一限制,本文提出了UATR-G Transformer,这是一种非欧几里德DL模型,将Transformer架构与图神经网络(GNN)集成在一起。该模型包括三个关键组件:Mel补丁化块,GTransformer块和分类头。Mel区块化块将Mel频谱图划分为重叠区块,而G Transformer块采用Transformer编码器来捕获分割区块之间的互信息以生成Mel图嵌入。随后,GNN通过对局部邻域关系进行建模来增强这些嵌入,前馈网络(FFN)进一步执行特征变换。基于两个广泛使用的基准数据集的实验结果表明,UATR-GTransformer实现的性能与国家的最先进的方法竞争。此外,可解释性分析表明,该模型有效地提取了丰富的频域信息,突出了其在海洋工程中的应用潜力。
摘要:Underwater acoustic target recognition (UATR) is extremely challenging due to the complexity of ship-radiated noise and the variability of ocean environments. Although deep learning (DL) approaches have achieved promising results, most existing models implicitly assume that underwater acoustic data lie in a Euclidean space. This assumption, however, is unsuitable for the inherently complex topology of underwater acoustic signals, which exhibit non-stationary, non-Gaussian, and nonlinear characteristics. To overcome this limitation, this paper proposes the UATR-GTransformer, a non-Euclidean DL model that integrates Transformer architectures with graph neural networks (GNNs). The model comprises three key components: a Mel patchify block, a GTransformer block, and a classification head. The Mel patchify block partitions the Mel-spectrogram into overlapping patches, while the GTransformer block employs a Transformer Encoder to capture mutual information between split patches to generate Mel-graph embeddings. Subsequently, a GNN enhances these embeddings by modeling local neighborhood relationships, and a feed-forward network (FFN) further performs feature transformation. Experiments results based on two widely used benchmark datasets demonstrate that the UATR-GTransformer achieves performance competitive with state-of-the-art methods. In addition, interpretability analysis reveals that the proposed model effectively extracts rich frequency-domain information, highlighting its potential for applications in ocean engineering.


【2】PhraseVAE and PhraseLDM: Latent Diffusion for Full-Song Multitrack Symbolic Music Generation
标题:PhraseVAE和PhraseLDM:用于整首歌曲多轨符号音乐生成的潜在扩散
链接:https://arxiv.org/abs/2512.11348

作者:Longshen Ou,Ye Wang
摘要:这份技术报告提出了一个新的范例,为全曲符号音乐的产生。现有的符号模型在音符属性标记上操作,并且具有非常长的序列、有限的上下文长度和对长范围结构的弱支持。我们通过介绍PhraseVAE和PhraseLDM来解决这些问题,PhraseLDM是第一个为全曲多轨符号音乐设计的潜在扩散框架。PhraseVAE将可变长度的复调音符序列压缩成紧凑的64维短语级表示,具有高重建保真度,允许有效的训练和结构良好的潜在空间。建立在这个潜在的空间,PhraseLDM生成一个完整的多轨道歌曲在一个单一的通过没有任何自回归组件。该系统消除了逐小节的顺序建模,支持多达128小节的音乐(8分钟,64 bpm),并产生完整的歌曲,具有连贯的局部纹理,惯用的乐器模式和清晰的全局结构。只有45M参数,我们的框架在几秒钟内生成一首完整的歌曲,同时保持有竞争力的音乐质量和生成多样性。总之,这些结果表明,短语级潜在扩散提供了一个有效的和可扩展的解决方案,在符号音乐生成的长序列建模。我们希望这项工作鼓励未来的符号音乐研究超越音符属性令牌,并考虑短语级单位作为一个更有效的和音乐意义的建模目标。
摘要:This technical report presents a new paradigm for full-song symbolic music generation. Existing symbolic models operate on note-attribute tokens and suffer from extremely long sequences, limited context length, and weak support for long-range structure. We address these issues by introducing PhraseVAE and PhraseLDM, the first latent diffusion framework designed for full-song multitrack symbolic music. PhraseVAE compresses variable-length polyphonic note sequences into compact 64-dimensional phrase-level representations with high reconstruction fidelity, allowing efficient training and a well-structured latent space. Built on this latent space, PhraseLDM generates an entire multi-track song in a single pass without any autoregressive components. The system eliminates bar-wise sequential modeling, supports up to 128 bars of music (8 minutes in 64 bpm), and produces complete songs with coherent local texture, idiomatic instrument patterns, and clear global structure. With only 45M parameters, our framework generates a full song within seconds while maintaining competitive musical quality and generation diversity. Together, these results show that phrase-level latent diffusion provides an effective and scalable solution to long-sequence modeling in symbolic music generation. We hope this work encourages future symbolic music research to move beyond note-attribute tokens and to consider phrase-level units as a more effective and musically meaningful modeling target.


【3】The Affective Bridge: Unifying Feature Representations for Speech Deepfake Detection
标题:情感桥梁:统一语音Deepfake检测的特征表示
链接:https://arxiv.org/abs/2512.11241

作者:Yupei Li,Chenyang Lyu,Longyue Wang,Weihua Luo,Kaifu Zhang,Björn W. Schuller
摘要:语音深度伪检测已经使用低级声学描述符进行了广泛的探索。然而,每个研究往往选择不同的特征集,这使得很难建立一个统一的表示的任务。此外,这些特征对于人类来说并不直观,因为随着deepfake生成技术的进步,真实语音和合成语音之间的区别变得越来越微妙。另一方面,情感仍然是一种独特的人类属性,目前的deepfake生成器很难完全复制,这反映了与真正的人工智能的差距。有趣的是,许多现有的声学和语义特征与情感有隐含的相关性。例如,由自动语音识别系统识别的语音特征通常随着情感表达而自然变化。基于这一见解,我们提出了一种新的训练框架,该框架利用情感作为传统deepfake特征和面向情感的表示之间的桥梁。在广泛使用的FakeOrReal和In-the-Wild数据集上进行的实验表明,准确性得到了一致和实质性的提高,分别提高了约6%和2%,而等错误率(EER)分别降低了约4%和1%,同时在ASVspoof 2019上实现了相当的结果。这种方法为所有特征提供了统一的训练策略,并为deepfake检测提供了可解释的特征方向,同时通过情感学习提高了模型性能。
摘要:Speech deepfake detection has been widely explored using low-level acoustic descriptors. However, each study tends to select different feature sets, making it difficult to establish a unified representation for the task. Moreover, such features are not intuitive for humans to perceive, as the distinction between bona fide and synthesized speech becomes increasingly subtle with the advancement of deepfake generation techniques. Emotion, on the other hand, remains a unique human attribute that current deepfake generator struggles to fully replicate, reflecting the gap toward true artificial general intelligence. Interestingly, many existing acoustic and semantic features have implicit correlations with emotion. For instance, speech features recognized by automatic speech recognition systems often varies naturally with emotional expression. Based on this insight, we propose a novel training framework that leverages emotion as a bridge between conventional deepfake features and emotion-oriented representations. Experiments on the widely used FakeOrReal and In-the-Wild datasets demonstrate consistent and substantial improvements in accuracy, up to approximately 6% and 2% increases, respectively, and in equal error rate (EER), showing reductions of up to about 4% and 1%, respectively, while achieving comparable results on ASVspoof2019. This approach provides a unified training strategy for all features and interpretable feature direction for deepfake detection while improving model performance through emotion-informed learning.


【4】REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming Distillation
标题:REST:通过ID上下文缓存和同步流媒体蒸馏生成基于扩散的实时端到端流媒体说话头
链接:https://arxiv.org/abs/2512.11229

作者:Haotian Wang,Yuzhe Weng,Xinyi Yu,Jun Du,Haoran Xu,Xiaoyan Wu,Shan He,Bing Yin,Cong Liu,Qingfeng Liu
备注:10pages, 4 figures
摘要:扩散模型已经显著地推进了讲话头部生成领域。然而,缓慢的推理速度和非自回归范式严重限制了基于扩散的THG模型的应用。在这项研究中,我们提出了REST,第一个基于扩散的,实时的,端到端的流音频驱动的说话头生成框架。为了支持实时端到端生成,首先通过高时空VAE压缩来学习紧凑的视频潜在空间。此外,为了使自回归流在紧凑的视频潜在空间,我们引入了一个ID上下文缓存机制,它集成了ID-Sink和上下文缓存的原则,以保持时间的一致性和身份一致性,在长时间流生成的键值缓存。此外,异步流蒸馏(ASD)的训练策略,提出了减轻自回归生成的错误积累和增强时间一致性,它利用非流教师与异步噪声调度监督流学生模型的训练。REST弥合了自回归和基于扩散的方法之间的差距,证明了需要实时通话头生成的应用程序的巨大价值。实验结果表明,REST优于国家的最先进的方法在生成速度和整体性能。
摘要:Diffusion models have significantly advanced the field of talking head generation. However, the slow inference speeds and non-autoregressive paradigms severely constrain the application of diffusion-based THG models. In this study, we propose REST, the first diffusion-based, real-time, end-to-end streaming audio-driven talking head generation framework. To support real-time end-to-end generation, a compact video latent space is first learned through high spatiotemporal VAE compression. Additionally, to enable autoregressive streaming within the compact video latent space, we introduce an ID-Context Cache mechanism, which integrates ID-Sink and Context-Cache principles to key-value caching for maintaining temporal consistency and identity coherence during long-time streaming generation. Furthermore, an Asynchronous Streaming Distillation (ASD) training strategy is proposed to mitigate error accumulation in autoregressive generation and enhance temporal consistency, which leverages a non-streaming teacher with an asynchronous noise schedule to supervise the training of the streaming student model. REST bridges the gap between autoregressive and diffusion-based approaches, demonstrating substantial value for applications requiring real-time talking head generation. Experimental results demonstrate that REST outperforms state-of-the-art methods in both generation speed and overall performance.


【5】Mitigation of multi-path propagation artefacts in acoustic targets with cepstral adaptive filtering
标题:利用倒谱自适应过滤减轻声目标中的多路径传播伪影
链接:https://arxiv.org/abs/2512.11165

作者:Lucas C. F. Domingos,Russell S. A. Brinkworth,Paulo E. Santos,Karl Sammut
摘要:无源声传感是监测船舶和飞机等移动目标的一种具有成本效益的解决方案,但其性能受到多径反射和运动诱导伪影等复杂传播效应的阻碍。现有的滤波技术没有适当地结合环境的特性或考虑介质属性的变化,限制了它们在分离源和反射分量方面的有效性。本文提出了一种在谱图中分离目标信号和反射信号的方法。使用自适应带阻滤波器将时间滤波应用于倒谱系数,该自适应带阻滤波器基于倒谱系数分量的相对强度动态地调整其带宽。该方法提高了信噪比(SNR),对数谱距离(LSD),和Itakura-Saito(IS)距离跨越速度范围从每秒10至100米的飞机噪声与模拟运动。它还将DeepShip和VTUAD v2数据集在水下任务中的船型分类性能分别提高了2.28和2.62个马修斯相关系数百分点。这些结果表明,所提出的管道,以改善声目标分类和多径环境中的时间延迟估计的潜力,与未来的工作,旨在幅度保存和多传感器应用。
摘要:Passive acoustic sensing is a cost-effective solution for monitoring moving targets such as vessels and aircraft, but its performance is hindered by complex propagation effects like multi-path reflections and motion-induced artefacts. Existing filtering techniques do not properly incorporate the characteristics of the environment or account for variability in medium properties, limiting their effectiveness in separating source and reflection components. This paper proposes a method for separating target signals from their reflections in a spectrogram. Temporal filtering is applied to cepstral coefficients using an adaptive band-stop filter, which dynamically adjusts its bandwidth based on the relative intensity of the quefrency components. The method improved the signal-to-noise ratio (SNR), log-spectral distance (LSD), and Itakura-Saito (IS) distance across velocities ranging from 10 to 100 metres per second in aircraft noise with simulated motion. It also enhanced the performance of ship-type classification in underwater tasks by 2.28 and 2.62 Matthews Correlation Coefficient percentage points for the DeepShip and VTUAD v2 datasets, respectively. These results demonstrate the potential of the proposed pipeline to improve acoustic target classification and time-delay estimation in multi-path environments, with future work aimed at amplitude preservation and multi-sensor applications.


【6】The TCG CREST -- RKMVERI Submission for the NCIIPC Startup India AI Grand Challenge
标题:TCG CREST -- RKMVERI参加NCIIPC初创公司印度人工智能大挑战赛
链接:https://arxiv.org/abs/2512.11009

作者:Nikhil Raghav,Arnab Banerjee,Janojit Chakraborty,Avisek Gupta,Swami Punyeshwarananda,Md Sahidullah
备注:6 pages, 3 tables, 3 figures, report submission for the NCIIPC Startup India AI Grand Challenge, Problem Statement 06
摘要:在这份报告中,我们总结了我们的团队为首届NCIIPC Startup India AI GRAND CHALLENGE开发的集成多语言音频处理管道,解决了问题陈述06:语音不可知的说话者识别和拨号,以及随后的转录和翻译系统。我们的主要重点是推进发言人日记化,这是多语言和代码混合场景的关键组成部分。这项工作的主要目的是研究我们的内部扬声器日记(SD)系统在现实世界中的适用性。为此,我们研究了一个强大的语音活动检测(VAD)技术和微调扬声器嵌入模型,以提高扬声器识别在低资源设置。我们利用了我们自己最近提出的多核共识谱聚类框架,该框架大大提高了组织者提供的训练语料库中所有记录的日志化性能。用于说话者和语言识别、自动语音识别(ASR)和神经机器翻译的补充模块集成在管道中。后处理改进进一步提高了系统的鲁棒性。
摘要:In this report, we summarize the integrated multilingual audio processing pipeline developed by our team for the inaugural NCIIPC Startup India AI GRAND CHALLENGE, addressing Problem Statement 06: Language-Agnostic Speaker Identification and Diarisation, and subsequent Transcription and Translation System. Our primary focus was on advancing speaker diarization, a critical component for multilingual and code-mixed scenarios. The main intent of this work was to study the real-world applicability of our in-house speaker diarization (SD) systems. To this end, we investigated a robust voice activity detection (VAD) technique and fine-tuned speaker embedding models for improved speaker identification in low-resource settings. We leveraged our own recently proposed multi-kernel consensus spectral clustering framework, which substantially improved the diarization performance across all recordings in the training corpus provided by the organizers. Complementary modules for speaker and language identification, automatic speech recognition (ASR), and neural machine translation were integrated in the pipeline. Post-processing refinements further improved system robustness.


【7】Benchmarking Automatic Speech Recognition Models for African Languages
标题:非洲语言自动语音识别模型的基准测试
链接:https://arxiv.org/abs/2512.10968

作者:Alvin Nahabwe,Sulaiman Kagumire,Denis Musinguzi,Bruno Beijuka,Jonah Mubuuke Kyagaba,Peter Nabende,Andrew Katumba,Joyce Nakatumba-Nabende
备注:19 pages, 8 figures, Deep Learning Indiba, Proceedings of Machine Learning Research
摘要:非洲语言的自动语音识别(ASR)仍然受到有限的标记数据以及缺乏模型选择,数据缩放和解码策略的系统指导的限制。大型的预训练系统,如Whisper、XLS-R、MMS和W2 v-BERT,已经扩大了对ASR技术的访问,但它们在非洲低资源环境中的比较行为尚未得到统一和系统的研究。在这项工作中,我们对13种非洲语言的四种最先进的ASR模型进行了基准测试,并根据从1到400小时的转录数据的逐渐增大的子集对其进行了微调。除了报告错误率之外,我们还提供了关于模型在不同条件下表现不同的新见解。我们发现,MMS和W2 v-BERT在资源非常低的情况下数据效率更高,XLS-R随着额外数据的可用而更有效地扩展,Whisper在中等资源条件下表现出优势。我们还分析了外部语言模型解码产生改进的地方,并根据声学和文本资源之间的对齐情况,确定了它达到平台或引入额外错误的情况。通过强调预训练覆盖率,模型架构,数据集域和资源可用性之间的相互作用,这项研究为代表性不足的语言的ASR系统的设计提供了实用和见解。
摘要:Automatic speech recognition (ASR) for African languages remains constrained by limited labeled data and the lack of systematic guidance on model selection, data scaling, and decoding strategies. Large pre-trained systems such as Whisper, XLS-R, MMS, and W2v-BERT have expanded access to ASR technology, but their comparative behavior in African low-resource contexts has not been studied in a unified and systematic way. In this work, we benchmark four state-of-the-art ASR models across 13 African languages, fine-tuning them on progressively larger subsets of transcribed data ranging from 1 to 400 hours. Beyond reporting error rates, we provide new insights into why models behave differently under varying conditions. We show that MMS and W2v-BERT are more data efficient in very low-resource regimes, XLS-R scales more effectively as additional data becomes available, and Whisper demonstrates advantages in mid-resource conditions. We also analyze where external language model decoding yields improvements and identify cases where it plateaus or introduces additional errors, depending on the alignment between acoustic and text resources. By highlighting the interaction between pre-training coverage, model architecture, dataset domain, and resource availability, this study offers practical and insights into the design of ASR systems for underrepresented languages.


【8】ASR Under the Stethoscope: Evaluating Biases in Clinical Speech Recognition across Indian Languages
标题:听诊器下的ASB:评估印度语言临床语音识别的偏见
链接:https://arxiv.org/abs/2512.10967

作者:Subham Kumar,Prakrithi Shivaprakash,Abhishek Manoharan,Astut Kurariya,Diptadhi Mukherjee,Lekhansh Shukla,Animesh Mukherjee,Prabhat Chand,Pratima Murthy
摘要:自动语音识别(ASR)越来越多地用于记录临床遇到的情况,但其在多语言和人口统计学上多样化的印度医疗保健环境中的可靠性在很大程度上仍然未知。在这项研究中,我们对跨越卡纳达语,印地语和印度英语的真实世界临床访谈数据进行了首次系统的ASR性能审计,比较了领先的模型,包括Indic Whisper,Whisper,Sarvam,Google语音文本,Gemma3n,Omnilingual,Vaani和Gemini。我们评估了不同语言,说话者和人口统计学亚组的转录准确性,特别关注影响患者与临床医生的错误模式以及基于性别或交叉差异。我们的研究结果显示,模型和语言之间存在很大的差异,一些系统在印度英语上表现得很有竞争力,但在代码混合或方言语音上却失败了。我们还发现了与演讲者角色和性别相关的系统性性能差距,引起了人们对临床环境中公平部署的担忧。通过提供全面的多语言基准和公平性分析,我们的工作强调了印度医疗保健生态系统在文化和人口统计方面包容性ASR发展的必要性。
摘要:Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare contexts remains largely unknown. In this study, we conduct the first systematic audit of ASR performance on real world clinical interview data spanning Kannada, Hindi, and Indian English, comparing leading models including Indic Whisper, Whisper, Sarvam, Google speech to text, Gemma3n, Omnilingual, Vaani, and Gemini. We evaluate transcription accuracy across languages, speakers, and demographic subgroups, with a particular focus on error patterns affecting patients vs. clinicians and gender based or intersectional disparities. Our results reveal substantial variability across models and languages, with some systems performing competitively on Indian English but failing on code mixed or vernacular speech. We also uncover systematic performance gaps tied to speaker role and gender, raising concerns about equitable deployment in clinical settings. By providing a comprehensive multilingual benchmark and fairness analysis, our work highlights the need for culturally and demographically inclusive ASR development for healthcare ecosystem in India.


【9】Processing through encoding: Quantum circuit approaches for point-wise multiplication and convolution
标题:通过编码进行处理:用于逐点乘法和卷积的量子电路方法
链接:https://arxiv.org/abs/2512.11457

作者:Andreas Papageorgiou,Paulo Vitor Itaborai,Kostas Blekos,Karl Jansen
备注:Presented at ISQCMC '25: 3rd International Symposium on Quantum Computing and Musical Creativity
摘要:本文介绍了量子电路方法的逐点乘法和卷积的复杂功能,概念化为“通过编码处理”。利用已知的技术,我们描述了一种方法,其中多个复杂的功能被编码到辅助量子位。将所提出的方案应用于两个函数$f$和$g$,它们的逐点乘积$f(x)g(x)$被示出自然地形成为所得到的量子态的一部分的系数。坚持卷积定理,我们然后演示如何卷积f*g$可以构造。类似于相关的工作,这涉及傅立叶系数$\mathcal{F}[f]$和$\mathcal{F}[g]$的编码,这有助于它们的逐点乘法,然后是逆量子傅立叶变换。我们讨论了这些技术的模拟,它们集成到一个扩展的\verb|宽图毛迪奥|软件包的音频信号处理,并提出初步的实验验证。这项工作为量子信号处理提供了一个有前途的途径,在量子增强音频操纵和合成等领域具有潜在的应用。
摘要:This paper introduces quantum circuit methodologies for pointwise multiplication and convolution of complex functions, conceptualized as "processing through encoding". Leveraging known techniques, we describe an approach where multiple complex functions are encoded onto auxiliary qubits. Applying the proposed scheme for two functions $f$ and $g$, their pointwise product $f(x)g(x)$ is shown to naturally form as the coefficients of part of the resulting quantum state. Adhering to the convolution theorem, we then demonstrate how the convolution $f*g$ can be constructed. Similarly to related work, this involves the encoding of the Fourier coefficients $\mathcal{F}[f]$ and $\mathcal{F}[g]$, which facilitates their pointwise multiplication, followed by the inverse Quantum Fourier Transform. We discuss the simulation of these techniques, their integration into an extended \verb|quantumaudio| package for audio signal processing, and present initial experimental validations. This work offers a promising avenue for quantum signal processing, with potential applications in areas such as quantum-enhanced audio manipulation and synthesis.


eess.AS音频处理


【1】All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
标题:一体机ASB:在双模式ASB中统一了CTC、注意力和传感器的编码器-解码器模型
链接:https://arxiv.org/abs/2512.11543

作者:Takafumi Moriya,Masato Mimura,Tomohiro Tanaka,Hiroshi Sato,Ryo Masumura,Atsunori Ogawa
备注:Accepted to ASRU 2025
摘要:本文提出了一个统一的框架,一体化的ASR,允许一个单一的模型,以支持多种自动语音识别(ASR)的范例,包括连接主义的时间分类(CTC),基于注意力的编码器-解码器(AED),和换能器,在离线和流模式。虽然每个ASR架构根据应用程序提供不同的优势和权衡,但为每个场景维护单独的模型会导致大量的开发和部署成本。为了解决这个问题,我们引入了一个多模式连接器,可以在一个统一的模型中无缝集成各种ASR模式。实验表明,All-in-One ASR显著减少了总模型占用空间,同时匹配甚至超过了单独优化的ASR模型的识别性能。此外,联合解码利用了不同ASR模式的互补优势,从而进一步提高了识别精度。
摘要:This paper proposes a unified framework, All-in-One ASR, that allows a single model to support multiple automatic speech recognition (ASR) paradigms, including connectionist temporal classification (CTC), attention-based encoder-decoder (AED), and Transducer, in both offline and streaming modes. While each ASR architecture offers distinct advantages and trade-offs depending on the application, maintaining separate models for each scenario incurs substantial development and deployment costs. To address this issue, we introduce a multi-mode joiner that enables seamless integration of various ASR modes within a single unified model. Experiments show that All-in-One ASR significantly reduces the total model footprint while matching or even surpassing the recognition performance of individually optimized ASR models. Furthermore, joint decoding leverages the complementary strengths of different ASR modes, yielding additional improvements in recognition accuracy.


【2】Robust Detection of Underwater Target Against Non-Uniform Noise With Optical Fiber DAS Array
标题:非均匀噪声背景下光纤DAS阵列水下目标鲁棒检测
链接:https://arxiv.org/abs/2512.11231

作者:Siyuan Cang,Cong Liu,Xueli Sheng,Xiaoming Cui,Chao Li,Changxin Fa,Jiantong Chen,Chaoran Yang,Huayong Yang
备注:17 pages, 29 figures. The IEEE Transactions on Instrumentation and Measurement has accepted this research for publication, and it is currently accessible in its early access version
摘要:海洋环境噪声的非均匀空间特性严重影响了水下目标的检测。此外,自然和人为声源的存在,包括航运交通、海洋生物和地质活动,使水下声学景观进一步复杂化。应对这些挑战需要先进的水下传感器和强大的信号处理技术。在本文中,我们提出了一种新的方法,利用光纤分布式声学传感(DAS)系统结合宽带广义稀疏协方差拟合框架的水下目标方向感测,特别是专注于对非均匀噪声的鲁棒性。DAS系统采用了新开发的螺旋敏化光缆,与传统的海底电缆相比,该光缆显著提高了灵敏度。这种创新设计使系统能够更精确地捕获声学信号。值得注意的是,螺旋缠绕敏化电缆的灵敏度约为-145.69 dB re:1 rad /(uPa*m),如在驻波管内测量的。通过模拟,我们评估了该算法在不同噪声水平和目标配置下的性能,一致表明与传统波束形成技术和其他稀疏技术相比,具有更高的准确性和更低的背景噪声。在受控水池实验中,DAS系统采集的波形与标准水听器采集的波形相关系数达到0.973,表明信号捕获的保真度很高。
摘要:The detection of underwater targets is severely affected by the non-uniform spatial characteristics of marine environmental noise. Additionally, the presence of both natural and anthropogenic acoustic sources, including shipping traffic, marine life, and geological activity, further complicates the underwater acoustic landscape. Addressing these challenges requires advanced underwater sensors and robust signal processing techniques. In this paper, we present a novel approach that leverages an optical fiber distributed acoustic sensing (DAS) system combined with a broadband generalized sparse covariance-fitting framework for underwater target direction sensing, particularly focusing on robustness against non-uniform noise. The DAS system incorporates a newly developed spiral-sensitized optical cable, which significantly improves sensitivity compared to conventional submarine cables. This innovative design enables the system to capture acoustic signals with greater precision. Notably, the sensitivity of the spiral-wound sensitized cable is around -145.69 dB re: 1 rad / (uPa*m), as measured inside the standing-wave tube. Employing simulations, we assess the performance of the algorithm across diverse noise levels and target configurations, consistently revealing higher accuracy and reduced background noise compared to conventional beamforming techniques and other sparse techniques. In a controlled pool experiment, the correlation coefficient between waveforms acquired by the DAS system and a standard hydrophone reached 0.973, indicating high fidelity in signal capture.


【3】Benchmarking Automatic Speech Recognition Models for African Languages
标题:非洲语言自动语音识别模型的基准测试
链接:https://arxiv.org/abs/2512.10968

作者:Alvin Nahabwe,Sulaiman Kagumire,Denis Musinguzi,Bruno Beijuka,Jonah Mubuuke Kyagaba,Peter Nabende,Andrew Katumba,Joyce Nakatumba-Nabende
备注:19 pages, 8 figures, Deep Learning Indiba, Proceedings of Machine Learning Research
摘要:非洲语言的自动语音识别(ASR)仍然受到有限的标记数据以及缺乏模型选择,数据缩放和解码策略的系统指导的限制。大型的预训练系统,如Whisper,XLS-R,MMS和W2 v-BERT,已经扩大了对ASR技术的访问,但它们在非洲低资源环境中的比较行为尚未以统一和系统的方式进行研究。在这项工作中,我们对13种非洲语言的四种最先进的ASR模型进行了基准测试,并根据从1到400小时的转录数据的逐渐增大的子集对其进行了微调。除了报告错误率之外,我们还提供了关于模型在不同条件下表现不同的新见解。我们发现,MMS和W2 v-BERT在资源非常低的情况下数据效率更高,XLS-R随着额外数据的可用而更有效地扩展,Whisper在中等资源条件下表现出优势。我们还分析了外部语言模型解码产生改进的地方,并根据声学和文本资源之间的对齐情况,确定了它达到平台或引入额外错误的情况。通过强调预训练覆盖率,模型架构,数据集域和资源可用性之间的相互作用,这项研究为代表性不足的语言的ASR系统的设计提供了实用和见解。
摘要:Automatic speech recognition (ASR) for African languages remains constrained by limited labeled data and the lack of systematic guidance on model selection, data scaling, and decoding strategies. Large pre-trained systems such as Whisper, XLS-R, MMS, and W2v-BERT have expanded access to ASR technology, but their comparative behavior in African low-resource contexts has not been studied in a unified and systematic way. In this work, we benchmark four state-of-the-art ASR models across 13 African languages, fine-tuning them on progressively larger subsets of transcribed data ranging from 1 to 400 hours. Beyond reporting error rates, we provide new insights into why models behave differently under varying conditions. We show that MMS and W2v-BERT are more data efficient in very low-resource regimes, XLS-R scales more effectively as additional data becomes available, and Whisper demonstrates advantages in mid-resource conditions. We also analyze where external language model decoding yields improvements and identify cases where it plateaus or introduces additional errors, depending on the alignment between acoustic and text resources. By highlighting the interaction between pre-training coverage, model architecture, dataset domain, and resource availability, this study offers practical and insights into the design of ASR systems for underrepresented languages.


【4】ASR Under the Stethoscope: Evaluating Biases in Clinical Speech Recognition across Indian Languages
标题:听诊器下的ASB:评估印度语言临床语音识别的偏见
链接:https://arxiv.org/abs/2512.10967

作者:Subham Kumar,Prakrithi Shivaprakash,Abhishek Manoharan,Astut Kurariya,Diptadhi Mukherjee,Lekhansh Shukla,Animesh Mukherjee,Prabhat Chand,Pratima Murthy
摘要:自动语音识别(ASR)越来越多地用于记录临床遇到的情况,但其在多语言和人口统计学上多样化的印度医疗保健环境中的可靠性在很大程度上仍然未知。在这项研究中,我们对跨越卡纳达语,印地语和印度英语的真实世界临床访谈数据进行了首次系统的ASR性能审计,比较了领先的模型,包括Indic Whisper,Whisper,Sarvam,Google语音文本,Gemma3n,Omnilingual,Vaani和Gemini。我们评估了不同语言,说话者和人口统计学亚组的转录准确性,特别关注影响患者与临床医生的错误模式以及基于性别或交叉差异。我们的研究结果显示,模型和语言之间存在很大的差异,一些系统在印度英语上表现得很有竞争力,但在代码混合或方言语音上却失败了。我们还发现了与演讲者角色和性别相关的系统性性能差距,引起了人们对临床环境中公平部署的担忧。通过提供全面的多语言基准和公平性分析,我们的工作强调了印度医疗保健生态系统在文化和人口统计方面包容性ASR发展的必要性。
摘要:Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare contexts remains largely unknown. In this study, we conduct the first systematic audit of ASR performance on real world clinical interview data spanning Kannada, Hindi, and Indian English, comparing leading models including Indic Whisper, Whisper, Sarvam, Google speech to text, Gemma3n, Omnilingual, Vaani, and Gemini. We evaluate transcription accuracy across languages, speakers, and demographic subgroups, with a particular focus on error patterns affecting patients vs. clinicians and gender based or intersectional disparities. Our results reveal substantial variability across models and languages, with some systems performing competitively on Indian English but failing on code mixed or vernacular speech. We also uncover systematic performance gaps tied to speaker role and gender, raising concerns about equitable deployment in clinical settings. By providing a comprehensive multilingual benchmark and fairness analysis, our work highlights the need for culturally and demographically inclusive ASR development for healthcare ecosystem in India.


机器翻译由腾讯交互翻译提供,仅供参考