今日论文合集:cs.SD语音10篇,eess.AS音频处理11篇。本文经arXiv每日学术速递授权转载
【1】Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
标题:基于正弦卷积的原始音频多模式情感识别
链接:https://arxiv.org/abs/2402.11954
作者:Xiaohui Zhang,Wenjie Fu,Mangui Liang
摘要:语音情感识别(SER)对于计算机来说仍然是一项复杂的任务,在最真实的数据集上,平均召回率通常约为70%。大部分SER系统都是从音频信号中手工提取能量、过零率、频谱信息、韵律、梅尔频率倒谱系数(MFCC)等特征,近年来,利用原始波形训练神经网络成为一种新兴的趋势。这种方法是有利的,因为它消除了特征提取流水线。从时域信号中学习对于语音识别,说话人验证等任务已经显示出良好的效果。在本文中,我们利用Sinc卷积层,这是一种用于预处理原始语音波形以进行情感识别的有效架构,从原始音频信号中提取声学特征,然后进行长短期记忆(LSTM)。我们还将语言特征和附加的对话情感解码(DED)策略。在交互式情绪二元运动捕捉(IEMOCAP)数据集上,该方法在四类情绪上的加权准确率达到85.1%.
摘要:Speech Emotion Recognition (SER) is still a complex task for computers with average recall rates usually about 70% on the most realistic datasets. Most SER systems use hand-crafted features extracted from audio signal such as energy, zero crossing rate, spectral information, prosodic, mel frequency cepstral coefficient (MFCC), and so on. More recently, using raw waveform for training neural network is becoming an emerging trend. This approach is advantageous as it eliminates the feature extraction pipeline. Learning from time-domain signal has shown good results for tasks such as speech recognition, speaker verification etc. In this paper, we utilize Sinc-convolution layer, which is an efficient architecture for preprocessing raw speech waveform for emotion recognition, to extract acoustic features from raw audio signals followed by a long short-term memory (LSTM). We also incorporate linguistic features and append a dialogical emotion decoding (DED) strategy. Our approach achieves a weighted accuracy of 85.1\% in four class emotion on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) dataset.
【2】 Soft-Weighted CrossEntropy Loss for Continous Alzheimer's Disease Detection作者:Xiaohui Zhang,Wenjie Fu,Mangui Liang摘要:阿尔茨海默病是老年人常见的认知障碍。阿尔茨海默病(Alzheimer's disease,AD)的早期准确诊断对痴呆研究的进展有着重要的影响。目前,研究人员已经使用机器学习方法从参与者的语音中检测出阿尔茨海默病。然而,目前的方法识别准确率不令人满意,并且大多数集中在使用低维手工特征从音频中提取相关信息。本文提出了一种基于预训练框架Wav 2 vec 2.0(Wav 2 vec 2)的阿尔茨海默病检测系统。此外,通过将损失函数替换为软加权交叉熵损失函数,在相同的测试数据集上获得了85.45%的识别准确率。摘要:Alzheimer's disease is a common cognitive disorder in the elderly. Early and accurate diagnosis of Alzheimer's disease (AD) has a major impact on the progress of research on dementia. At present, researchers have used machine learning methods to detect Alzheimer's disease from the speech of participants. However, the recognition accuracy of current methods is unsatisfactory, and most of them focus on using low-dimensional handcrafted features to extract relevant information from audios. This paper proposes an Alzheimer's disease detection system based on the pre-trained framework Wav2vec 2.0 (Wav2vec2). In addition, by replacing the loss function with the Soft-Weighted CrossEntropy loss function, we achieved 85.45\% recognition accuracy on the same test dataset.
【3】 Unraveling Complex Data Diversity in Underwater Acoustic Target Recognition through Convolution-based Mixture of Experts标题:基于卷积的混合专家分解水声目标识别中的复杂数据多样性作者:Yuan Xie,Jiawei Ren,Ji Xu摘要:由于水声信号的复杂性,水声目标识别是一项困难的任务。复杂的水下环境、不可预测的传输信道和动态的运动状态极大地影响了真实世界的水声信号,甚至可能掩盖与目标相关的内在特征。因此,水下声信号的数据分布具有很高的类内多样性,从而影响识别系统的准确性和鲁棒性,为了解决这些问题,这项工作提出了一种基于卷积的混合专家(CMoE),以细粒度的方式识别水下目标。所提出的技术引入了多个专家层作为独立的学习者,以及一个路由层,根据输入的特性来确定专家的分配。这种设计允许模型利用独立的参数空间,便于学习复杂的水下信号具有高的类内多样性。此外,这项工作通过平衡正则化和可选的残差模块来优化CMoE结构。为了验证我们所提出的技术的有效性,我们进行了详细的实验和可视化分析三个水声数据库在几个声学功能。实验结果表明,我们的CMoE一贯实现显着的性能改进,提供卓越的识别精度相比,现有的先进方法。摘要:Underwater acoustic target recognition is a difficult task owing to the intricate nature of underwater acoustic signals. The complex underwater environments, unpredictable transmission channels, and dynamic motion states greatly impact the real-world underwater acoustic signals, and may even obscure the intrinsic characteristics related to targets. Consequently, the data distribution of underwater acoustic signals exhibits high intra-class diversity, thereby compromising the accuracy and robustness of recognition systems.To address these issues, this work proposes a convolution-based mixture of experts (CMoE) that recognizes underwater targets in a fine-grained manner. The proposed technique introduces multiple expert layers as independent learners, along with a routing layer that determines the assignment of experts according to the characteristics of inputs. This design allows the model to utilize independent parameter spaces, facilitating the learning of complex underwater signals with high intra-class diversity. Furthermore, this work optimizes the CMoE structure by balancing regularization and an optional residual module. To validate the efficacy of our proposed techniques, we conducted detailed experiments and visualization analyses on three underwater acoustic databases across several acoustic features. The experimental results demonstrate that our CMoE consistently achieves significant performance improvements, delivering superior recognition accuracy when compared to existing advanced methods.【4】 Low-power SNN-based audio source localisation using a Hilbert Transform spike encoding scheme标题:使用希尔BERT变换尖峰编码方案的低功率基于SNN的音频源定位作者:Saeid Haghighatshoar,Dylan R Muir摘要:声源定位在许多消费电子设备中使用,以帮助将音频与单个扬声器隔离并抑制噪声。定位通常通过“波束成形”算法来完成,该算法组合麦克风音频流以改善从特定入射源方向接收的信号功率。波束成形算法通常使用音频源的频率分量的知识以及已知的麦克风阵列几何形状,以在组合麦克风流之前分析地相移麦克风流。一组密集的带通滤波器通常用于从宽带音频流中获得已知频率的“窄带”分量。这些方法实现了高精度,但最先进的窄带波束成形算法在计算上要求很高,因此难以集成到低功耗物联网设备中。我们展示了一种新的方法,声源定位在任意麦克风阵列,设计用于超低功耗尖峰神经网络(SNN)的有效实施。我们使用一种新的短时希尔伯特变换(STHT),以消除需要苛刻的带通滤波的音频,并介绍了一种新的伴随方法与尖峰事件的音频编码。我们的波束形成和定位方法实现了SNN方法的最新精度,并且与传统的非SNN超分辨率方法相当。我们将我们的方法部署到低功耗SNN音频推理硬件上,与超分辨率方法相比,实现了更低的功耗。我们证明了信号处理方法可以与尖峰神经网络实现协同设计,以实现高水平的功率效率。我们新的基于希尔伯特变换的波束形成方法也有望提高传统的基于DSP的信号处理的效率。摘要:Sound source localisation is used in many consumer electronics devices, to help isolate audio from individual speakers and to reject noise. Localization is frequently accomplished by "beamforming" algorithms, which combine microphone audio streams to improve received signal power from particular incident source directions. Beamforming algorithms generally use knowledge of the frequency components of the audio source, along with the known microphone array geometry, to analytically phase-shift microphone streams before combining them. A dense set of band-pass filters is often used to obtain known-frequency "narrowband" components from wide-band audio streams. These approaches achieve high accuracy, but state of the art narrowband beamforming algorithms are computationally demanding, and are therefore difficult to integrate into low-power IoT devices. We demonstrate a novel method for sound source localisation in arbitrary microphone arrays, designed for efficient implementation in ultra-low-power spiking neural networks (SNNs). We use a novel short-time Hilbert transform (STHT) to remove the need for demanding band-pass filtering of audio, and introduce a new accompanying method for audio encoding with spiking events. Our beamforming and localisation approach achieves state-of-the-art accuracy for SNN methods, and comparable with traditional non-SNN super-resolution approaches. We deploy our method to low-power SNN audio inference hardware, and achieve much lower power consumption compared with super-resolution methods. We demonstrate that signal processing approaches can be co-designed with spiking neural network implementations to achieve high levels of power efficiency. Our new Hilbert-transform-based method for beamforming promises to also improve the efficiency of traditional DSP-based signal processing.【5】 Significance of Chirp MFCC as a Feature in Speech and Audio Applications标题:Chirp MFCC特征在语音和音频应用中的意义作者:S. Johanan Joysingh,P. Vijayalakshmi,T. Nagarajan摘要:提出了一种新的功能,基于线性调频z变换,提供了一个改进的基本真实频谱的表示。该特征,即啁啾MFCC,是通过从啁啾幅度谱而不是傅里叶变换幅度谱计算Mel频率倒谱系数而得到的。该建议的理论基础,并使用产品的似然高斯的实验验证,以显示所提出的啁啾MFCC,与香草MFCC相比,提供了改进的类分离进行了讨论。此外,使用三个不同的任务,即,语音音乐分类,说话人识别,语音命令识别的功能的真实世界的评价。它示出在所有三个任务中,所提出的啁啾MFCC提供了相当大的改进。摘要:A novel feature, based on the chirp z-transform, that offers an improved representation of the underlying true spectrum is proposed. This feature, the chirp MFCC, is derived by computing the Mel frequency cepstral coefficients from the chirp magnitude spectrum, instead of the Fourier transform magnitude spectrum. The theoretical foundations for the proposal, and the experimental validation using product of likelihood Gaussians, to show the improved class separation offered by the proposed chirp MFCC, when compared with vanilla MFCC are discussed. Further, real world evaluation of the feature is performed using three diverse tasks, namely, speech-music classification, speaker identification, and speech commands recognition. It is shown in all three tasks that the proposed chirp MFCC offers considerable improvements.【6】 Language-Codec: Reducing the Gaps Between Discrete Codec Representation and Speech Language Models标题:语言编解码器:缩小离散编解码器表示和语音语言模型之间的差距作者:Shengpeng Ji,Minghui Fang,Ziyue Jiang,Rongjie Huang,Jialung Zuo,Shulei Wang,Zhou Zhao摘要:近年来,大型语言模型在生成任务(例如,语音克隆和音频生成),涉及语音、音频、音乐和其它信号域。这些模型的一个关键要素是离散声学编解码器,它作为一个中间表示取代梅尔频谱图。然而,在离散编解码器和下游语音语言模型之间存在一些差距。具体而言,1)大多数编解码器模型仅在1,000小时的数据上训练,而大多数语音语言模型在60,000小时上训练; 2)实现良好的重建性能需要利用大量码本,这增加了下游语音语言模型的负担; 3)码本的初始通道包含过多的信息,使得从弱监督信号(例如下游任务中的文本)直接生成声学令牌具有挑战性。因此,利用语音语言模型的特点,我们提出了语音编解码器。在视频编解码器中,我们引入了掩码通道残差矢量量化(MCRVQ)机制以及改进的傅立叶变换结构和更大的训练数据集,以解决上述差距。我们将我们的方法与竞争的音频压缩算法进行比较,并在广泛的评估中观察到显着的优越性。此外,我们还验证了下游语音语言模型的语音编解码器的效率。源代码和预训练模型可以在https://github.com/speechnovateur/languagecodec_tmp上访问。摘要:In recent years, large language models have achieved significant success in generative tasks (e.g., speech cloning and audio generation) related to speech, audio, music, and other signal domains. A crucial element of these models is the discrete acoustic codecs, which serves as an intermediate representation replacing the mel-spectrogram. However, there exist several gaps between discrete codecs and downstream speech language models. Specifically, 1) most codec models are trained on only 1,000 hours of data, whereas most speech language models are trained on 60,000 hours; 2) Achieving good reconstruction performance requires the utilization of numerous codebooks, which increases the burden on downstream speech language models; 3) The initial channel of the codebooks contains excessive information, making it challenging to directly generate acoustic tokens from weakly supervised signals such as text in downstream tasks. Consequently, leveraging the characteristics of speech language models, we propose Language-Codec. In the Language-Codec, we introduce a Mask Channel Residual Vector Quantization (MCRVQ) mechanism along with improved Fourier transform structures and larger training datasets to address the aforementioned gaps. We compare our method with competing audio compression algorithms and observe significant outperformance across extensive evaluations. Furthermore, we also validate the efficiency of the Language-Codec on downstream speech language models. The source code and pre-trained models can be accessed at https://github.com/speechnovateur/languagecodec_tmp .【7】 On the relationship between speech and hearing作者:Srinivasan Umesh,Leon Cohen,Douglas Nelson摘要:我们提出了一个实验连接语音生产和听力的框架。使用这种方法,我们描述的实验结果,导致的概念,由不同的个人和被认为是相同的声音可以转化为彼此的“语音规模”。语音尺度仅使用语音数据根据经验确定。我们显示的相似性的语音规模的MEL规模史蒂文斯和Jumemann,这是来自听力实验。因此,我们实验性地将言语产生和听觉联系起来。摘要:We present a framework for experimentally linking speech production and hearing. Using this approach, we describe experimental results, that lead to the concept that sounds made by different individuals and perceived to be the same can be transformed into each other by a "speech scale". The speech scale is empirically determined using only speech data. We show the similarity of the speech scale to the MEL scale of Stevens and Volkmann, which was derived only from hearing experiments. We thus experimentally link speech production and hearing.
【8】 Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation链接:https://arxiv.org/abs/2402.11747作者:Nineli Lashkarashvili,Wen Wu,Guangzhi Sun,Philip C. Woodland摘要:基础模型在语音情感识别(SER)方面表现出了卓越的性能。然而,由于情感语料库中的数据有限,对SER的大型预训练模型的所有参数进行微调可能是资源密集型的,并且容易发生过拟合。本文研究了SER的参数有效微调(PEFT),系统地研究了用于离散情感类别分类和维度情感属性预测的各种PEFT适配器。结果表明,PEFT方法的组合优于完全微调,可训练参数的数量显着减少。此外,提出了一种两阶段自适应策略,以适应在更容易获得的行为情感数据上训练的模型,使模型更善于捕捉自然的情感表达。语料库内和跨语料库实验验证了该方法在提高源域和目标域的性能方面的有效性。摘要:Foundation models have shown superior performance for speech emotion recognition (SER). However, given the limited data in emotion corpora, finetuning all parameters of large pre-trained models for SER can be both resource-intensive and susceptible to overfitting. This paper investigates parameter-efficient finetuning (PEFT) for SER. Various PEFT adaptors are systematically studied for both classification of discrete emotion categories and prediction of dimensional emotional attributes. The results demonstrate that the combination of PEFT methods surpasses full finetuning with a significant reduction in the number of trainable parameters. Furthermore, a two-stage adaptation strategy is proposed to adapt models trained on acted emotion data, which is more readily available, to make the model more adept at capturing natural emotional expressions. Both intra- and cross-corpus experiments validate the efficacy of the proposed approach in enhancing the performance on both the source and target domains.【9】 Diffuse Sound Field Synthesis作者:Franz Zotter,Stefan Riedel,Lukas Gölles,Matthias Frank备注:27 pages, 17 figures, submitted to acta acustica nov 20th 2023, including jan/feb 2024 upgrades while awaiting the reviews摘要:不相关的周围声源可以用来产生扩展的扩散声场吗?根据定义,目标是一个恒定的声压级,一个消失的平均声强,不相关的声波从各个方向各向同性到达。这是否需要周围2D和3D源布局的特定源和几何形状? 作为方法,我们采用数值模拟,并进行了一系列的计算与不相关的圆形/球形源布局,或这样的无限多余的尺寸,我们指出潜在的理论关系。使用由指数b修改的径向衰减1/r^b,用超几何函数、盖根堡多项式、圆调和球调和函数表示结果场产生了富有成效的见解。 在圆形布局中,以指数b=1/2衰减的波合成理想的扩展扩散声场;球形布局在b=1时也是如此。没有一种布局能够合成一个完全恒定的预期声压级,但其平坦度是可以接受的。 球形t-设计描述了最佳的源布局,具有良好的描述区域的高扩散性,和非球形,凸布局可以通过恢复各向同性或通过模式匹配的最大扩散合成来改善。 理论和仿真提供了一个基于扬声器的扩散声场合成的基础,并有助于最近的心理声学研究结果在空间音频的物理原因。摘要:Can uncorrelated surrounding sound sources be used to generate extended diffuse sound fields? By definition, targets are a constant sound pressure level, a vanishing average sound intensity, uncorrelated sound waves arriving isotropically from all directions. Does this require specific sources and geometries for surrounding 2D and 3D source layouts? As methods, we employ numeric simulations and undertake a series of calculations with uncorrelated circular/spherical source layouts, or such with infinite excess dimensions, and we point out relations to potential theory. Using a radial decay 1/r^b modified by the exponent b, the representation of the resulting fields with hypergeometric functions, Gegenbauer polynomials, and circular as well as spherical harmonics yields fruitful insights. In circular layouts, waves decaying by the exponent b=1/2 synthesize ideally extended, diffuse sound fields; spherical layouts do so with b=1. None of the layouts synthesizes a perfectly constant expected sound pressure level but its flatness is acceptable. Spherical t-designs describe optimal source layouts with well-described area of high diffuseness, and non-spherical, convex layouts can be improved by restoring isotropy or by mode matching for a maximally diffuse synthesis. Theory and simulation offer a basis for loudspeaker-based synthesis of diffuse sound fields and contribute physical reasons to recent psychoacoustic findings in spatial audio.
【10】 Feedback Delay Network Optimization作者:Gloria Dal Santo,Karolina Prawda,Sebastian J. Schlecht,Vesa Välimäki摘要:人工混响算法的一个常见的祸根是频谱着色,通常表现为金属振铃,导致感知音质的下降。本文提出了一种优化框架,其中使用可微反馈延迟网络来学习一组参数以迭代地减少着色。优化的参数包括反馈矩阵,以及输入和输出增益。优化目标是双重的:通过频谱损失最大化频谱平坦度,同时通过惩罚参数值中的稀疏性来保持时间密度。在保持期望的脉冲响应密度的同时,实现了模态激励的有利的较窄分布。在主观评估中,新方法证明了有效地减少后期混响的感知着色。所提出的方法实现了计算节省相比,基线,同时保持其性能。这项工作的有效性证明通过两个应用场景,其中自然探空合成冲激响应通过引入衰减滤波器和可优化的散射反馈矩阵。摘要:A common bane of artificial reverberation algorithms is spectral coloration, typically manifesting as metallic ringing, leading to a degradation in the perceived sound quality. This paper presents an optimization framework where a differentiable feedback delay network is used to learn a set of parameters to reduce coloration iteratively. The parameters under optimization include the feedback matrix, as well as the input and output gains. The optimization objective is twofold: to maximize spectral flatness through a spectral loss while maintaining temporal density by penalizing sparseness in the parameter values. A favorable narrower distribution of modal excitation is achieved while maintaining the desired impulse response density. In a subjective assessment, the new method proves effective in reducing perceptual coloration of late reverberation. The proposed method achieves computational savings compared to the baseline while preserving its performance. The effectiveness of this work is demonstrated through two application scenarios where natural-sounding synthetic impulse responses are obtained via the introduction of attenuation filters and an optimizable scattering feedback matrix.【1】 Significance of Chirp MFCC as a Feature in Speech and Audio Applications标题:Chirp MFCC特征在语音和音频应用中的意义作者:S. Johanan Joysingh,P. Vijayalakshmi,T. Nagarajan摘要:提出了一种新的功能,基于线性调频z变换,提供了一个改进的基本真实频谱的表示。该特征,即啁啾MFCC,是通过从啁啾幅度谱而不是傅里叶变换幅度谱计算Mel频率倒谱系数而得到的。该建议的理论基础,并使用产品的似然高斯的实验验证,以显示所提出的啁啾MFCC,与香草MFCC相比,提供了改进的类分离进行了讨论。此外,使用三个不同的任务,即,语音音乐分类,说话人识别,语音命令识别的功能的真实世界的评价。它示出在所有三个任务中,所提出的啁啾MFCC提供了相当大的改进。摘要:A novel feature, based on the chirp z-transform, that offers an improved representation of the underlying true spectrum is proposed. This feature, the chirp MFCC, is derived by computing the Mel frequency cepstral coefficients from the chirp magnitude spectrum, instead of the Fourier transform magnitude spectrum. The theoretical foundations for the proposal, and the experimental validation using product of likelihood Gaussians, to show the improved class separation offered by the proposed chirp MFCC, when compared with vanilla MFCC are discussed. Further, real world evaluation of the feature is performed using three diverse tasks, namely, speech-music classification, speaker identification, and speech commands recognition. It is shown in all three tasks that the proposed chirp MFCC offers considerable improvements.【2】 Bayesian Parameter-Efficient Fine-Tuning for Overcoming Catastrophic Forgetting作者:Haolin Chen,Philip N. Garner备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible摘要:虽然动机的文本到语音合成模型的适应,我们认为,更通用的参数有效的微调(PEFT)是一个合适的框架来做这样的适应。然而,灾难性遗忘仍然是PEFT的一个问题,破坏了预训练模型的固有能力。我们证明了现有的贝叶斯学习技术可以应用于PEFT,以防止灾难性的遗忘,只要微调层的参数偏移可以计算出差异。在一系列关于语言建模和语音合成任务的原则性实验中,我们利用已建立的拉普拉斯近似,包括对角和克罗内克因子化方法,用低秩自适应(LoRA)正则化PEFT,并比较它们在预训练知识保存方面的性能。我们的研究结果表明,灾难性遗忘可以通过我们的方法克服,而不会降低微调性能,并且使用Kronecker因子近似比对角近似更好地保留了预训练知识。摘要:Although motivated by the adaptation of text-to-speech synthesis models, we argue that more generic parameter-efficient fine-tuning (PEFT) is an appropriate framework to do such adaptation. However, catastrophic forgetting remains an issue with PEFT, damaging the pre-trained model's inherent capabilities. We demonstrate that existing Bayesian learning techniques can be applied to PEFT to prevent catastrophic forgetting as long as the parameter shift of the fine-tuned layers can be calculated differentiably. In a principled series of experiments on language modeling and speech synthesis tasks, we utilize established Laplace approximations, including diagonal and Kronecker factored approaches, to regularize PEFT with the low-rank adaptation (LoRA) and compare their performance in pre-training knowledge preservation. Our results demonstrate that catastrophic forgetting can be overcome by our methods without degrading the fine-tuning performance, and using the Kronecker factored approximations produces a better preservation of the pre-training knowledge than the diagonal ones.
【3】 Language-Codec: Reducing the Gaps Between Discrete Codec Representation and Speech Language Models标题:语言编解码器:缩小离散编解码器表示和语音语言模型之间的差距作者:Shengpeng Ji,Minghui Fang,Ziyue Jiang,Rongjie Huang,Jialung Zuo,Shulei Wang,Zhou Zhao摘要:近年来,大型语言模型在生成任务(例如,语音克隆和音频生成),涉及语音、音频、音乐和其它信号域。这些模型的一个关键要素是离散声学编解码器,它作为一个中间表示取代梅尔频谱图。然而,在离散编解码器和下游语音语言模型之间存在一些差距。具体而言,1)大多数编解码器模型仅在1,000小时的数据上训练,而大多数语音语言模型在60,000小时上训练; 2)实现良好的重建性能需要利用大量码本,这增加了下游语音语言模型的负担; 3)码本的初始通道包含过多的信息,使得从弱监督信号(例如下游任务中的文本)直接生成声学令牌具有挑战性。因此,利用语音语言模型的特点,我们提出了语音编解码器。在视频编解码器中,我们引入了掩码通道残差矢量量化(MCRVQ)机制以及改进的傅立叶变换结构和更大的训练数据集,以解决上述差距。我们将我们的方法与竞争的音频压缩算法进行比较,并在广泛的评估中观察到显着的优越性。此外,我们还验证了下游语音语言模型的语音编解码器的效率。源代码和预训练模型可以在https://github.com/speechnovateur/languagecodec_tmp上访问。摘要:In recent years, large language models have achieved significant success in generative tasks (e.g., speech cloning and audio generation) related to speech, audio, music, and other signal domains. A crucial element of these models is the discrete acoustic codecs, which serves as an intermediate representation replacing the mel-spectrogram. However, there exist several gaps between discrete codecs and downstream speech language models. Specifically, 1) most codec models are trained on only 1,000 hours of data, whereas most speech language models are trained on 60,000 hours; 2) Achieving good reconstruction performance requires the utilization of numerous codebooks, which increases the burden on downstream speech language models; 3) The initial channel of the codebooks contains excessive information, making it challenging to directly generate acoustic tokens from weakly supervised signals such as text in downstream tasks. Consequently, leveraging the characteristics of speech language models, we propose Language-Codec. In the Language-Codec, we introduce a Mask Channel Residual Vector Quantization (MCRVQ) mechanism along with improved Fourier transform structures and larger training datasets to address the aforementioned gaps. We compare our method with competing audio compression algorithms and observe significant outperformance across extensive evaluations. Furthermore, we also validate the efficiency of the Language-Codec on downstream speech language models. The source code and pre-trained models can be accessed at https://github.com/speechnovateur/languagecodec_tmp .
【4】 On the relationship between speech and hearing作者:Srinivasan Umesh,Leon Cohen,Douglas Nelson摘要:我们提出了一个实验连接语音生产和听力的框架。使用这种方法,我们描述的实验结果,导致的概念,由不同的个人和被认为是相同的声音可以转化为彼此的“语音规模”。语音尺度仅使用语音数据根据经验确定。我们显示的相似性的语音规模的MEL规模史蒂文斯和Jumemann,这是来自听力实验。因此,我们实验性地将言语产生和听觉联系起来。摘要:We present a framework for experimentally linking speech production and hearing. Using this approach, we describe experimental results, that lead to the concept that sounds made by different individuals and perceived to be the same can be transformed into each other by a "speech scale". The speech scale is empirically determined using only speech data. We show the similarity of the speech scale to the MEL scale of Stevens and Volkmann, which was derived only from hearing experiments. We thus experimentally link speech production and hearing.
【5】 Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation作者:Nineli Lashkarashvili,Wen Wu,Guangzhi Sun,Philip C. Woodland摘要:基础模型在语音情感识别(SER)方面表现出了卓越的性能。然而,由于情感语料库中的数据有限,对SER的大型预训练模型的所有参数进行微调可能是资源密集型的,并且容易发生过拟合。本文研究了SER的参数有效微调(PEFT),系统地研究了用于离散情感类别分类和维度情感属性预测的各种PEFT适配器。结果表明,PEFT方法的组合优于完全微调,可训练参数的数量显着减少。此外,提出了一种两阶段自适应策略,以适应在更容易获得的行为情感数据上训练的模型,使模型更善于捕捉自然的情感表达。语料库内和跨语料库实验验证了该方法在提高源域和目标域的性能方面的有效性。摘要:Foundation models have shown superior performance for speech emotion recognition (SER). However, given the limited data in emotion corpora, finetuning all parameters of large pre-trained models for SER can be both resource-intensive and susceptible to overfitting. This paper investigates parameter-efficient finetuning (PEFT) for SER. Various PEFT adaptors are systematically studied for both classification of discrete emotion categories and prediction of dimensional emotional attributes. The results demonstrate that the combination of PEFT methods surpasses full finetuning with a significant reduction in the number of trainable parameters. Furthermore, a two-stage adaptation strategy is proposed to adapt models trained on acted emotion data, which is more readily available, to make the model more adept at capturing natural emotional expressions. Both intra- and cross-corpus experiments validate the efficacy of the proposed approach in enhancing the performance on both the source and target domains.【6】 Diffuse Sound Field Synthesis作者:Franz Zotter,Stefan Riedel,Lukas Gölles,Matthias Frank备注:27 pages, 17 figures, submitted to acta acustica nov 20th 2023, including jan/feb 2024 upgrades while awaiting the reviews摘要:不相关的周围声源可以用来产生扩展的扩散声场吗?根据定义,目标是一个恒定的声压级,一个消失的平均声强,不相关的声波从各个方向各向同性到达。这是否需要周围2D和3D源布局的特定源和几何形状? 作为方法,我们采用数值模拟,并进行了一系列的计算与不相关的圆形/球形源布局,或这样的无限多余的尺寸,我们指出潜在的理论关系。使用由指数b修改的径向衰减1/r^b,用超几何函数、盖根堡多项式、圆调和球调和函数表示结果场产生了富有成效的见解。 在圆形布局中,以指数b=1/2衰减的波合成理想的扩展扩散声场;球形布局在b=1时也是如此。没有一种布局能够合成一个完全恒定的预期声压级,但其平坦度是可以接受的。 球形t-设计描述了最佳的源布局,具有良好的描述区域的高扩散性,和非球形,凸布局可以通过恢复各向同性或通过模式匹配的最大扩散合成来改善。 理论和仿真提供了一个基于扬声器的扩散声场合成的基础,并有助于最近的心理声学研究结果在空间音频的物理原因。摘要:Can uncorrelated surrounding sound sources be used to generate extended diffuse sound fields? By definition, targets are a constant sound pressure level, a vanishing average sound intensity, uncorrelated sound waves arriving isotropically from all directions. Does this require specific sources and geometries for surrounding 2D and 3D source layouts? As methods, we employ numeric simulations and undertake a series of calculations with uncorrelated circular/spherical source layouts, or such with infinite excess dimensions, and we point out relations to potential theory. Using a radial decay 1/r^b modified by the exponent b, the representation of the resulting fields with hypergeometric functions, Gegenbauer polynomials, and circular as well as spherical harmonics yields fruitful insights. In circular layouts, waves decaying by the exponent b=1/2 synthesize ideally extended, diffuse sound fields; spherical layouts do so with b=1. None of the layouts synthesizes a perfectly constant expected sound pressure level but its flatness is acceptable. Spherical t-designs describe optimal source layouts with well-described area of high diffuseness, and non-spherical, convex layouts can be improved by restoring isotropy or by mode matching for a maximally diffuse synthesis. Theory and simulation offer a basis for loudspeaker-based synthesis of diffuse sound fields and contribute physical reasons to recent psychoacoustic findings in spatial audio.
【7】 Feedback Delay Network Optimization作者:Gloria Dal Santo,Karolina Prawda,Sebastian J. Schlecht,Vesa Välimäki摘要:人工混响算法的一个常见的祸根是频谱着色,通常表现为金属振铃,导致感知音质的下降。本文提出了一种优化框架,其中使用可微反馈延迟网络来学习一组参数以迭代地减少着色。优化的参数包括反馈矩阵,以及输入和输出增益。优化目标是双重的:通过频谱损失最大化频谱平坦度,同时通过惩罚参数值中的稀疏性来保持时间密度。在保持期望的脉冲响应密度的同时,实现了模态激励的有利的较窄分布。在主观评估中,新方法证明了有效地减少后期混响的感知着色。所提出的方法实现了计算节省相比,基线,同时保持其性能。这项工作的有效性证明通过两个应用场景,其中自然探空合成冲激响应通过引入衰减滤波器和可优化的散射反馈矩阵。摘要:A common bane of artificial reverberation algorithms is spectral coloration, typically manifesting as metallic ringing, leading to a degradation in the perceived sound quality. This paper presents an optimization framework where a differentiable feedback delay network is used to learn a set of parameters to reduce coloration iteratively. The parameters under optimization include the feedback matrix, as well as the input and output gains. The optimization objective is twofold: to maximize spectral flatness through a spectral loss while maintaining temporal density by penalizing sparseness in the parameter values. A favorable narrower distribution of modal excitation is achieved while maintaining the desired impulse response density. In a subjective assessment, the new method proves effective in reducing perceptual coloration of late reverberation. The proposed method achieves computational savings compared to the baseline while preserving its performance. The effectiveness of this work is demonstrated through two application scenarios where natural-sounding synthetic impulse responses are obtained via the introduction of attenuation filters and an optimizable scattering feedback matrix.
【8】 Multimodal Emotion Recognition from Raw Audio with Sinc-convolution作者:Xiaohui Zhang,Wenjie Fu,Mangui Liang摘要:语音情感识别(SER)对于计算机来说仍然是一项复杂的任务,在最真实的数据集上,平均召回率通常约为70%。大部分SER系统都是从音频信号中手工提取能量、过零率、频谱信息、韵律、梅尔频率倒谱系数(MFCC)等特征,近年来,利用原始波形训练神经网络成为一种新兴的趋势。这种方法是有利的,因为它消除了特征提取流水线。从时域信号中学习对于语音识别,说话人验证等任务已经显示出良好的效果。在本文中,我们利用Sinc卷积层,这是一种用于预处理原始语音波形以进行情感识别的有效架构,从原始音频信号中提取声学特征,然后进行长短期记忆(LSTM)。我们还将语言特征和附加的对话情感解码(DED)策略。在交互式情绪二元运动捕捉(IEMOCAP)数据集上,该方法在四类情绪上的加权准确率达到85.1%.摘要:Speech Emotion Recognition (SER) is still a complex task for computers with average recall rates usually about 70% on the most realistic datasets. Most SER systems use hand-crafted features extracted from audio signal such as energy, zero crossing rate, spectral information, prosodic, mel frequency cepstral coefficient (MFCC), and so on. More recently, using raw waveform for training neural network is becoming an emerging trend. This approach is advantageous as it eliminates the feature extraction pipeline. Learning from time-domain signal has shown good results for tasks such as speech recognition, speaker verification etc. In this paper, we utilize Sinc-convolution layer, which is an efficient architecture for preprocessing raw speech waveform for emotion recognition, to extract acoustic features from raw audio signals followed by a long short-term memory (LSTM). We also incorporate linguistic features and append a dialogical emotion decoding (DED) strategy. Our approach achieves a weighted accuracy of 85.1\% in four class emotion on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) dataset.【9】 Soft-Weighted CrossEntropy Loss for Continous Alzheimer's Disease Detection作者:Xiaohui Zhang,Wenjie Fu,Mangui Liang摘要:阿尔茨海默病是老年人常见的认知障碍。阿尔茨海默病(Alzheimer's disease,AD)的早期准确诊断对痴呆研究的进展有着重要的影响。目前,研究人员已经使用机器学习方法从参与者的语音中检测出阿尔茨海默病。然而,目前的方法识别准确率不令人满意,并且大多数集中在使用低维手工特征从音频中提取相关信息。本文提出了一种基于预训练框架Wav 2 vec 2.0(Wav 2 vec 2)的阿尔茨海默病检测系统。此外,通过将损失函数替换为软加权交叉熵损失函数,在相同的测试数据集上获得了85.45%的识别准确率。摘要:Alzheimer's disease is a common cognitive disorder in the elderly. Early and accurate diagnosis of Alzheimer's disease (AD) has a major impact on the progress of research on dementia. At present, researchers have used machine learning methods to detect Alzheimer's disease from the speech of participants. However, the recognition accuracy of current methods is unsatisfactory, and most of them focus on using low-dimensional handcrafted features to extract relevant information from audios. This paper proposes an Alzheimer's disease detection system based on the pre-trained framework Wav2vec 2.0 (Wav2vec2). In addition, by replacing the loss function with the Soft-Weighted CrossEntropy loss function, we achieved 85.45\% recognition accuracy on the same test dataset.
【10】 Unraveling Complex Data Diversity in Underwater Acoustic Target Recognition through Convolution-based Mixture of Experts标题:基于卷积的混合专家分解水声目标识别中的复杂数据多样性作者:Yuan Xie,Jiawei Ren,Ji Xu摘要:由于水声信号的复杂性,水声目标识别是一项困难的任务。复杂的水下环境、不可预测的传输信道和动态的运动状态极大地影响了真实世界的水声信号,甚至可能掩盖与目标相关的内在特征。因此,水下声信号的数据分布具有很高的类内多样性,从而影响识别系统的准确性和鲁棒性,为了解决这些问题,这项工作提出了一种基于卷积的混合专家(CMoE),以细粒度的方式识别水下目标。所提出的技术引入了多个专家层作为独立的学习者,以及一个路由层,根据输入的特性来确定专家的分配。这种设计允许模型利用独立的参数空间,便于学习复杂的水下信号具有高的类内多样性。此外,这项工作通过平衡正则化和可选的残差模块来优化CMoE结构。为了验证我们所提出的技术的有效性,我们进行了详细的实验和可视化分析三个水声数据库在几个声学功能。实验结果表明,我们的CMoE一贯实现显着的性能改进,提供卓越的识别精度相比,现有的先进方法。摘要:Underwater acoustic target recognition is a difficult task owing to the intricate nature of underwater acoustic signals. The complex underwater environments, unpredictable transmission channels, and dynamic motion states greatly impact the real-world underwater acoustic signals, and may even obscure the intrinsic characteristics related to targets. Consequently, the data distribution of underwater acoustic signals exhibits high intra-class diversity, thereby compromising the accuracy and robustness of recognition systems.To address these issues, this work proposes a convolution-based mixture of experts (CMoE) that recognizes underwater targets in a fine-grained manner. The proposed technique introduces multiple expert layers as independent learners, along with a routing layer that determines the assignment of experts according to the characteristics of inputs. This design allows the model to utilize independent parameter spaces, facilitating the learning of complex underwater signals with high intra-class diversity. Furthermore, this work optimizes the CMoE structure by balancing regularization and an optional residual module. To validate the efficacy of our proposed techniques, we conducted detailed experiments and visualization analyses on three underwater acoustic databases across several acoustic features. The experimental results demonstrate that our CMoE consistently achieves significant performance improvements, delivering superior recognition accuracy when compared to existing advanced methods.【11】 Low-power SNN-based audio source localisation using a Hilbert Transform spike encoding scheme标题:基于希尔BERT变换尖峰编码的低功耗SNN音频源定位作者:Saeid Haghighatshoar,Dylan R Muir摘要:声源定位在许多消费电子设备中使用,以帮助将音频与单个扬声器隔离并抑制噪声。定位通常通过“波束成形”算法来完成,该算法组合麦克风音频流以改善从特定入射源方向接收的信号功率。波束成形算法通常使用音频源的频率分量的知识以及已知的麦克风阵列几何形状,以在组合麦克风流之前分析地相移麦克风流。一组密集的带通滤波器通常用于从宽带音频流中获得已知频率的“窄带”分量。这些方法实现了高精度,但最先进的窄带波束成形算法在计算上要求很高,因此难以集成到低功耗物联网设备中。我们展示了一种新的方法,声源定位在任意麦克风阵列,设计用于超低功耗尖峰神经网络(SNN)的有效实施。我们使用一种新的短时希尔伯特变换(STHT),以消除需要苛刻的带通滤波的音频,并介绍了一种新的伴随方法与尖峰事件的音频编码。我们的波束形成和定位方法实现了SNN方法的最新精度,并且与传统的非SNN超分辨率方法相当。我们将我们的方法部署到低功耗SNN音频推理硬件上,与超分辨率方法相比,实现了更低的功耗。我们证明了信号处理方法可以与尖峰神经网络实现协同设计,以实现高水平的功率效率。我们新的基于希尔伯特变换的波束形成方法也有望提高传统的基于DSP的信号处理的效率。摘要:Sound source localisation is used in many consumer electronics devices, to help isolate audio from individual speakers and to reject noise. Localization is frequently accomplished by "beamforming" algorithms, which combine microphone audio streams to improve received signal power from particular incident source directions. Beamforming algorithms generally use knowledge of the frequency components of the audio source, along with the known microphone array geometry, to analytically phase-shift microphone streams before combining them. A dense set of band-pass filters is often used to obtain known-frequency "narrowband" components from wide-band audio streams. These approaches achieve high accuracy, but state of the art narrowband beamforming algorithms are computationally demanding, and are therefore difficult to integrate into low-power IoT devices. We demonstrate a novel method for sound source localisation in arbitrary microphone arrays, designed for efficient implementation in ultra-low-power spiking neural networks (SNNs). We use a novel short-time Hilbert transform (STHT) to remove the need for demanding band-pass filtering of audio, and introduce a new accompanying method for audio encoding with spiking events. Our beamforming and localisation approach achieves state-of-the-art accuracy for SNN methods, and comparable with traditional non-SNN super-resolution approaches. We deploy our method to low-power SNN audio inference hardware, and achieve much lower power consumption compared with super-resolution methods. We demonstrate that signal processing approaches can be co-designed with spiking neural network implementations to achieve high levels of power efficiency. Our new Hilbert-transform-based method for beamforming promises to also improve the efficiency of traditional DSP-based signal processing.