本文经arXiv每日学术速递授权转载
链接:https://arxiv.org/abs/2411.11692
备注:International Society for Music Information Retrieval (ISMIR) 2024, Late Breaking Demo (LBD)
摘要:随着先进语言生成模型的出现,音乐字幕已经成为一项很有前途的任务。然而,音乐字幕的评估在很大程度上依赖于传统的指标,如BLEU,METEOR和ROUGE,这些指标是为其他领域开发的,没有适当的理由在这个新的领域使用。我们提出的情况下,传统的指标是容易受到语法变化,并表明他们不相关以及与人类的判断。通过解决这些问题,我们的目标是强调需要一个关键的重新评估如何音乐字幕进行评估。
摘要:Music captioning has emerged as a promising task, fueled by the advent ofadvanced language generation models. However, the evaluation of musiccaptioning relies heavily on traditional metrics such as BLEU, METEOR, andROUGE which were developed for other domains, without proper justification fortheir use in this new field. We present cases where traditional metrics arevulnerable to syntactic changes, and show they do not correlate well with humanjudgments. By addressing these issues, we aim to emphasize the need for acritical reevaluation of how music captions are assessed.
标题:使用语音分析作为年轻人抑郁症风险的早期指标
链接:https://arxiv.org/abs/2411.11541
备注:Submitted to ToaC
摘要:越来越多的文献报道抑郁症患者和对照组之间的语音质量差异。在这里,我们研究的可能性,使用语音分析作为早期预警信号的发展情绪障碍的年轻人。作为一个主要的跨学科的欧洲研究项目在四个国家(ECoWeB)的一部分,检查基于网络的预防计划,以减少年轻人抑郁症的风险的影响,我们分析了大量的声音的声音特征,在特定的一天,由参与者经历的情绪的声音报告。我们能够确定一些显着差异的声学线索,特别是在语音频谱中的能量分布,鼓励进一步的研究工作,以开发有前途的非侵扰性的风险指标,在正常说话的声音。这对于那些不太可能表现出抑郁症标准风险因素(如负面生活经历)的年轻人来说尤其重要。
摘要:Increasingly frequent publications in the literature report voice qualitydifferences between depressed patients and controls. Here, we examine thepossibility of using voice analysis as an early warning signal for thedevelopment of emotion disturbances in young adults. As part of a majorinterdisciplinary European research project in four countries (ECoWeB),examining the effects of web-based prevention programs to reduce the risk fordepression in young adults, we analyzed a large number of acoustic voicecharacteristics in vocal reports of emotions experienced by the participants ona specific day. We were able to identify a number of significant differences inacoustic cues, particularly with respect to the energy distribution in thevoice spectrum, encouraging further research efforts to develop promisingnon-obtrusive risk indicators in the normal speaking voice. This isparticularly important in the case of young adults who are less likely toexhibit standard risk factors for depression such as negative life experiences.
标题:CEEMDAN在欠确定语音分离中的性能研究
链接:https://arxiv.org/abs/2411.11312
备注:in Arabic language
摘要:CEEMDAN算法是分析非平稳信号的现代方法之一。本研究提出了一个研究的有效性,这种方法在音频源分离知道其工作的限制。得出了CEEMDAN分离混合信号的频率和幅度的两个条件。研究了该算法在语音噪声分离和语音信号分离中的性能。研究得出结论,CEEMDAN可以从语音中去除某些类型的噪声(语音改善),但不能将语音信号彼此分离(鸡尾酒会)。仿真使用Matlab环境和Noizeus数据库。
摘要:The CEEMDAN algorithm is one of the modern methods used in the analysis ofnon-stationary signals. This research presents a study of the effectiveness ofthis method in audio source separation to know the limits of its work. Itconcluded two conditions related to frequencies and amplitudes of mixed signalsto be separated by CEEMDAN. The performance of the algorithm in separatingnoise from speech and separating speech signals from each other is studied. Theresearch reached a conclusion that CEEMDAN can remove some types of noise fromspeech (speech improvement), and it cannot separate speech signals from eachother (cocktail party). Simulation is done using Matlab environment and Noizeusdatabase.
标题:ESTVocoder:一种基于Mel频谱图的兴奋光谱变换神经声码器
链接:https://arxiv.org/abs/2411.11258
备注:Accepted by NCMMSC2024
摘要:本文提出了EST声码器,一种新的激励谱变换的神经声码器的源滤波器理论的框架内。ESTVocoder使用神经滤波器将激励的幅度和相位谱转换为相应的语音幅度和相位谱,神经滤波器的主干是ConvNeXt v2块。最后通过短时傅立叶逆变换(ISTFT)重构语音波形。激励是基于F0构造的:对于有声片段,它包含完整的谐波信息,而对于无声片段,它由噪声表示。激励为滤波器提供了幅度和相位模式的先验知识,与传统的神经声码器相比,期望降低建模难度。为了保证合成语音的保真度,采用了对抗训练策略,使EST声码器具有多尺度和多分辨率的鉴别器。分析合成和文本到语音的实验都证实,我们提出的EST声码器优于或相当于其他基线神经声码器,例如,HiFi-GAN、SiFi-GAN和Vocos在合成语音质量方面具有合理的模型复杂度和生成速度。另外的分析实验也表明,引入的激励有效地加快了模型的收敛过程,由于语音频谱的先验信息包含在激励。
摘要:This paper proposes ESTVocoder, a novel excitation-spectral-transformedneural vocoder within the framework of source-filter theory. The ESTVocodertransforms the amplitude and phase spectra of the excitation into thecorresponding speech amplitude and phase spectra using a neural filter whosebackbone is ConvNeXt v2 blocks. Finally, the speech waveform is reconstructedthrough the inverse short-time Fourier transform (ISTFT). The excitation isconstructed based on the F0: for voiced segments, it contains full harmonicinformation, while for unvoiced segments, it is represented by noise. Theexcitation provides the filter with prior knowledge of the amplitude and phasepatterns, expecting to reduce the modeling difficulty compared to conventionalneural vocoders. To ensure the fidelity of the synthesized speech, anadversarial training strategy is applied to ESTVocoder with multi-scale andmulti-resolution discriminators. Analysis-synthesis and text-to-speechexperiments both confirm that our proposed ESTVocoder outperforms or iscomparable to other baseline neural vocoders, e.g., HiFi-GAN, SiFi-GAN, andVocos, in terms of synthesized speech quality, with a reasonable modelcomplexity and generation speed. Additional analysis experiments alsodemonstrate that the introduced excitation effectively accelerates the model'sconvergence process, thanks to the speech spectral prior information containedin the excitation.
标题:SAMOS:一种利用语义表示和声学特征的神经MOS预测模型
链接:https://arxiv.org/abs/2411.11232
摘要:使用平均意见评分(MOS)预测模型来评估语音的自然度对于语音合成系统的自动评估具有积极的意义。早期的MOS预测模型将语音的原始波形或幅度谱作为输入,而更先进的方法采用基于自监督学习(SSL)的模型从语音中提取语义表示以进行MOS预测。这些方法利用语音信息的有限方面进行MOS预测,导致有限的预测精度。因此,在本文中,我们提出了SAMOS,MOS预测模型,利用语音的语义和声学信息进行评估。具体而言,建议的SAMOS利用预训练的wav2vec2来提取语义表示,并使用预训练的BiVocoder的特征提取器来提取声学特征。然后,这两种类型的特征被馈送到预测网络中,该网络包括多任务头和聚合层,以获得最终的MOS分数。实验结果表明,根据系统级评估指标的结果,所提出的SAMOS在BVCC数据集上的性能优于当前最先进的MOS预测模型,并且在BC2019数据集上的性能相当。
摘要:Assessing the naturalness of speech using mean opinion score (MOS) predictionmodels has positive implications for the automatic evaluation of speechsynthesis systems. Early MOS prediction models took the raw waveform oramplitude spectrum of speech as input, whereas more advanced methods employedself-supervised-learning (SSL) based models to extract semantic representationsfrom speech for MOS prediction. These methods utilized limited aspects ofspeech information for MOS prediction, resulting in restricted predictionaccuracy. Therefore, in this paper, we propose SAMOS, a MOS prediction modelthat leverages both Semantic and Acoustic information of speech to be assessed.Specifically, the proposed SAMOS leverages a pretrained wav2vec2 to extractsemantic representations and uses the feature extractor of a pretrainedBiVocoder to extract acoustic features. These two types of features are thenfed into the prediction network, which includes multi-task heads and anaggregation layer, to obtain the final MOS score. Experimental resultsdemonstrate that the proposed SAMOS outperforms current state-of-the-art MOSprediction models on the BVCC dataset and performs comparable performance onthe BC2019 dataset, according to the results of system-level evaluationmetrics.
标题:水的声音:从倾倒的液体推断物理性质
链接:https://arxiv.org/abs/2411.11222
备注:25 pages, 17 figures. Project page at this https URL
摘要:我们研究视听观察和一个平凡但有趣的日常活动的基础物理之间的联系:倒液体。只给出液体倒入容器的声音,我们的目标是自动推断物理属性,如液位,容器的形状和大小,倾倒速度和填充时间。为此,我们:(i)在理论上示出这些属性可以从基频(音调)确定;(ii)训练音调检测模型,该音调检测模型具有来自模拟数据和具有物理启发目标的视觉数据的监督;(iii)引入真实浇注视频的新的大数据集用于系统研究;(iv)示出训练的模型确实可以推断真实数据的这些物理属性;最后,(v)我们展示了对各种容器形状,其他数据集和野外YouTube视频的强大泛化。我们的工作对声学、物理学和学习交叉领域的一个狭窄而丰富的问题有着敏锐的理解。它开辟了在机器人浇注中增强多感官感知的应用。
摘要:We study the connection between audio-visual observations and the underlyingphysics of a mundane yet intriguing everyday activity: pouring liquids. Givenonly the sound of liquid pouring into a container, our objective is toautomatically infer physical properties such as the liquid level, the shape andsize of the container, the pouring rate and the time to fill. To this end, we:(i) show in theory that these properties can be determined from the fundamentalfrequency (pitch); (ii) train a pitch detection model with supervision fromsimulated data and visual data with a physics-inspired objective; (iii)introduce a new large dataset of real pouring videos for a systematic study;(iv) show that the trained model can indeed infer these physical properties forreal data; and finally, (v) we demonstrate strong generalization to variouscontainer shapes, other datasets, and in-the-wild YouTube videos. Our workpresents a keen understanding of a narrow yet rich problem at the intersectionof acoustics, physics, and learning. It opens up applications to enhancemultisensory perception in robotic pouring.
标题:利用偏差纠正和模型融合的音调和光谱感知歌唱质量评估
链接:https://arxiv.org/abs/2411.11123
摘要:我们参加了2024年VoiceMOS挑战赛的第二场比赛,旨在预测歌唱样本的平均意见得分(MOS)。我们的参赛作品在所有参赛队伍中获得第一名,不包括官方基线。在本文中,我们进一步改进我们的意见,并提出了一种新的音高和频谱感知歌唱质量评估(PS-SQA)方法。PS-SQA是基于自监督学习MOS预测器设计的,结合了歌唱音高和频谱信息,分别使用音高直方图和非量化神经编解码器提取。此外,PS-SQA引入了偏差校正策略来解决低资源训练样本引起的预测偏差,并采用模型融合技术来进一步提高预测精度。实验结果证实,我们提出的PS-SQA显着优于所有竞争系统在所有系统级指标,确认其强大的歌唱质量评估能力。
摘要:We participated in track 2 of the VoiceMOS Challenge 2024, which aimed topredict the mean opinion score (MOS) of singing samples. Our submission securedthe first place among all participating teams, excluding the official baseline.In this paper, we further improve our submission and propose a novelPitch-and-Spectrum-aware Singing Quality Assessment (PS-SQA) method. The PS-SQAis designed based on the self-supervised-learning (SSL) MOS predictor,incorporating singing pitch and spectral information, which are extracted usingpitch histogram and non-quantized neural codec, respectively. Additionally, thePS-SQA introduces a bias correction strategy to address prediction biasescaused by low-resource training samples, and employs model fusion technology tofurther enhance prediction accuracy. Experimental results confirm that ourproposed PS-SQA significantly outperforms all competing systems across allsystem-level metrics, confirming its strong sing quality assessmentcapabilities.
标题:跨语言音素合成(IPC):增强第二语言发音的理论和计算方法
链接:https://arxiv.org/abs/2411.10927
备注:10 pages, 6 Figures, submitted to ACL ARR October 2024 for NAACL 2025
摘要:第二语言(L2)的学习者经常无意识地用母语(L1)中相似的音素替代不熟悉的L2音素,即使母语为L2的人认为这些声音是不同的和不可互换的。这种音位替换导致偏离第二语言的标准语音模式,给学习者获得准确的第二语言发音带来了挑战。为了解决这个问题,我们提出了跨语言语音合成(IPC),一种新的计算方法,旨在通过重建L2音素作为来自多个L1音素的复合声音,以最大限度地减少不正确的语音转移。两个自动语音识别模型的测试表明,当L2扬声器产生IPC生成的复合声音,目标L2音素的识别率提高了20%,相比,当他们的发音是由原始的语音迁移模式的影响。在相对较短的时间范围内观察到改善,表明快速获得复合声音。
摘要:Learners of a second language (L2) often unconsciously substitute unfamiliarL2 phonemes with similar phonemes from their native language (L1), even thoughnative speakers of the L2 perceive these sounds as distinct andnon-interchangeable. This phonemic substitution leads to deviations from thestandard phonological patterns of the L2, creating challenges for learners inacquiring accurate L2 pronunciation. To address this, we proposeInter-linguistic Phonetic Composition (IPC), a novel computational methoddesigned to minimize incorrect phonological transfer by reconstructing L2phonemes as composite sounds derived from multiple L1 phonemes. Tests with twoautomatic speech recognition models demonstrated that when L2 speakers producedIPC-generated composite sounds, the recognition rate of target L2 phonemesimproved by 20% compared to when their pronunciation was influenced by originalphonological transfer patterns. The improvement was observed within arelatively shorter time frame, demonstrating rapid acquisition of the compositesound.
标题:BanglaDialecto:端到端人工智能驱动的区域语音标准化
链接:https://arxiv.org/abs/2411.10879
备注:Accepted in 2024 IEEE International Conference on Big Data (IEEE BigData)
摘要:本研究的重点是识别孟加拉方言和转换成标准化的正式孟加拉语口音。方言,通常被称为区域语言,是在特定地点使用的语言的独特变体,并通过其语音,发音和词汇来识别。语音和语调的细微变化也受到地理位置、教育程度和社会经济地位的影响。方言标准化是必要的,以确保有效的沟通,教育的一致性,获得技术,经济机会和保护语言资源,同时尊重文化多样性。作为第五大语言,有1.6亿人使用大约55种不同的方言,解决孟加拉方言对于开发包容性通信工具至关重要。然而,由于缺乏全面的数据集以及处理不同方言的挑战,研究有限。随着多语言大型语言模型(MLLM)的发展,已经创造了新的可能性来解决方言自动语音识别(ASR)和机器翻译(MT)的挑战。这项研究提出了一个端到端的管道转换方言Noakhali语音标准孟加拉语语音。这项调查包括构建一个大规模的方言语音信号的多样化数据集,该数据集在ASR和LLM中定制微调过程,用于将方言语音转录为方言文本,并将方言文本翻译为标准孟加拉语文本。我们的实验表明,微调Whisper ASR模型实现了0.8%的CER和1.5%的WER,而BanglaT5模型在方言到标准文本翻译中获得了41.6%的BLEU分数。
摘要:This study focuses on recognizing Bangladeshi dialects and converting diverseBengali accents into standardized formal Bengali speech. Dialects, oftenreferred to as regional languages, are distinctive variations of a languagespoken in a particular location and are identified by their phonetics,pronunciations, and lexicon. Subtle changes in pronunciation and intonation arealso influenced by geographic location, educational attainment, andsocioeconomic status. Dialect standardization is needed to ensure effectivecommunication, educational consistency, access to technology, economicopportunities, and the preservation of linguistic resources while respectingcultural diversity. Being the fifth most spoken language with around 55distinct dialects spoken by 160 million people, addressing Bangla dialects iscrucial for developing inclusive communication tools. However, limited researchexists due to a lack of comprehensive datasets and the challenges of handlingdiverse dialects. With the advancement in multilingual Large Language Models(mLLMs), emerging possibilities have been created to address the challenges ofdialectal Automated Speech Recognition (ASR) and Machine Translation (MT). Thisstudy presents an end-to-end pipeline for converting dialectal Noakhali speechto standard Bangla speech. This investigation includes constructing alarge-scale diverse dataset with dialectal speech signals that tailored thefine-tuning process in ASR and LLM for transcribing the dialect speech todialect text and translating the dialect text to standard Bangla text. Ourexperiments demonstrated that fine-tuning the Whisper ASR model achieved a CERof 0.8% and WER of 1.5%, while the BanglaT5 model attained a BLEU score of41.6% for dialect-to-standard text translation.
标题:Buchla 200电子音乐盒合成器中使用的带通Twin-T主动过滤器
链接:https://arxiv.org/abs/2411.11358
备注:5 pages, 8 figures, MATLAB/Octave code at this https URL under the name b295_plots.m
摘要:本文分析了一个不寻常的有源带通滤波器中使用的Buchla模型295 10通道梳状滤波器,合成器模块开发的一部分,Buchla 200电动音乐盒的唐纳德Buchla。该过滤器由一个独特的重新安排的元素在一个经典的双T配置,据我们所知,它还没有在以前的文献中。作为一个例子,我们探讨其在295型中的具体应用。
摘要:This paper analyzes an unusual active bandpass filter employed in the BuchlaModel 295 10 Channel Comb Filter, a synthesizer module developed as part of theBuchla 200 Electric Music Box by Donald Buchla. The filter consists of apeculiar rearrangement of elements in a classic Twin-T configuration; to ourknowledge, it has not been previously addressed in the literature. As anexample, we explore its specific application in the Model 295.
标题:说话人验证系统中跨语言适应重编程的研究
链接:https://arxiv.org/abs/2411.11353
备注:Accepted by ISCSLP 2024
摘要:语言不匹配是说话人确认系统中最常见和最具挑战性的领域不匹配问题。对抗性重编程在SV的跨语言适应中显示出有希望的结果。通过在输入语音信号两侧填充可学习参数来实现重新编程。在本文中,我们研究了填充参数的数量和重编程模型的性能之间的关系。利用不同尺度的SV模型和数据集进行了充分的实验。结果表明,重编程一致地提高了跨语言SV的性能,而当使用较大的填充长度时,这种提高是饱和的,甚至是下降的。性能主要取决于原始SV模型的容量,而不是填充参数的数量。具有较大规模的SV模型具有更高的性能上界,并且可以承受更长的填充而不会降低性能。
摘要:Language mismatch is among the most common and challenging domain mismatchesin deploying speaker verification (SV) systems. Adversarial reprogramming hasshown promising results in cross-language adaptation for SV. The reprogrammingis implemented by padding learnable parameters on the two sides of input speechsignals. In this paper, we investigate the relationship between the number ofpadded parameters and the performance of the reprogrammed models. Sufficientexperiments are conducted with different scales of SV models and datasets. Theresults demonstrate that reprogramming consistently improves the performance ofcross-language SV, while the improvement is saturated or even degraded whenusing larger padding lengths. The performance is mainly determined by thecapacity of the original SV models instead of the number of padded parameters.The SV models with larger scales have higher upper bounds in performance andcan endure longer padding without performance degradation.
标题:Buchla 200电子音乐盒合成器中使用的带通Twin-T主动过滤器
链接:https://arxiv.org/abs/2411.11358
备注:5 pages, 8 figures, MATLAB/Octave code at this https URL under the name b295_plots.m
摘要:本文分析了一个不寻常的有源带通滤波器中使用的Buchla模型295 10通道梳状滤波器,合成器模块开发的一部分,Buchla 200电动音乐盒的唐纳德Buchla。该过滤器由一个独特的重新安排的元素在一个经典的双T配置,据我们所知,它还没有在以前的文献中。作为一个例子,我们探讨其在295型中的具体应用。
摘要:This paper analyzes an unusual active bandpass filter employed in the BuchlaModel 295 10 Channel Comb Filter, a synthesizer module developed as part of theBuchla 200 Electric Music Box by Donald Buchla. The filter consists of apeculiar rearrangement of elements in a classic Twin-T configuration; to ourknowledge, it has not been previously addressed in the literature. As anexample, we explore its specific application in the Model 295.
标题:说话人验证系统中跨语言适应重编程的研究
链接:https://arxiv.org/abs/2411.11353
备注:Accepted by ISCSLP 2024
摘要:语言不匹配是说话人确认系统中最常见和最具挑战性的领域不匹配问题。对抗性重编程在SV的跨语言适应中显示出有希望的结果。通过在输入语音信号两侧填充可学习参数来实现重新编程。在本文中,我们研究了填充参数的数量和重编程模型的性能之间的关系。利用不同尺度的SV模型和数据集进行了充分的实验。结果表明,重编程一致地提高了跨语言SV的性能,而当使用较大的填充长度时,这种提高是饱和的,甚至是下降的。性能主要取决于原始SV模型的容量,而不是填充参数的数量。具有较大规模的SV模型具有更高的性能上界,并且可以承受更长的填充而不会降低性能。
摘要:Language mismatch is among the most common and challenging domain mismatchesin deploying speaker verification (SV) systems. Adversarial reprogramming hasshown promising results in cross-language adaptation for SV. The reprogrammingis implemented by padding learnable parameters on the two sides of input speechsignals. In this paper, we investigate the relationship between the number ofpadded parameters and the performance of the reprogrammed models. Sufficientexperiments are conducted with different scales of SV models and datasets. Theresults demonstrate that reprogramming consistently improves the performance ofcross-language SV, while the improvement is saturated or even degraded whenusing larger padding lengths. The performance is mainly determined by thecapacity of the original SV models instead of the number of padded parameters.The SV models with larger scales have higher upper bounds in performance andcan endure longer padding without performance degradation.
标题:揭示语义和声学线索在正常和二分法听力中的作用
链接:https://arxiv.org/abs/2411.11308
备注:9 Pages, 4 Figures
摘要:尽管进行了广泛的研究,但声学和语义线索在复杂言语感知任务中的确切作用仍不清楚。在这项研究中,我们提出了一个范例,以了解这些线索在脑电图(EEG)数据的编码,使用匹配不匹配(MM)分类任务。MM任务涉及确定刺激和反应是否相互对应。我们设计了一个基于长短期记忆(LSTM)架构的多模态序列模型来执行MM任务。该模型输入有声学刺激(从语音包络导出)、语义刺激(从语音内容的文本表示导出)和神经响应(从EEG数据导出)。我们的实验在两个单独的条件下进行,i)自然被动听力条件和ii)基于听觉注意的双耳分听条件。使用MM任务作为分析框架,我们观察到- a)语音感知是基于单词边界的碎片化,b)声学和语义线索在自然听力条件下提供类似水平的MM任务性能,以及c)语义线索在双耳分听任务中提供比声学线索显著改进的MM分类。此外,该研究提供了双耳分听条件下右耳优势的证据。
摘要:Despite extensive research, the precise role of acoustic and semantic cues incomplex speech perception tasks remains unclear. In this study, we propose aparadigm to understand the encoding of these cues in electroencephalogram (EEG)data, using match-mismatch (MM) classification task. The MM task involvesdetermining whether the stimulus and response correspond to each other or not.We design a multi-modal sequence model, based on long short term memory (LSTM)architecture, to perform the MM task. The model is input with acoustic stimulus(derived from the speech envelope), semantic stimulus (derived from textualrepresentations of the speech content), and neural response (derived from theEEG data). Our experiments are performed on two separate conditions, i) naturalpassive listening condition and, ii) an auditory attention based dichoticlistening condition. Using the MM task as the analysis framework, we observethat - a) speech perception is fragmented based on word boundaries, b) acousticand semantic cues offer similar levels of MM task performance in naturallistening conditions, and c) semantic cues offer significantly improved MMclassification over acoustic cues in dichotic listening task. Further, thestudy provides evidence of right ear advantage in dichotic listeningconditions.
标题:具有后过滤器的可解释的基于DNN的Beamformer
链接:https://arxiv.org/abs/2411.10854
摘要:本文介绍了一种可解释的基于DNN的波束形成器与后置滤波器(ExNet-BF+PF)的多通道信号处理。我们的方法将U-Net网络与波束形成器结构相结合来解决这个问题。该方法涉及两阶段处理流水线。在第一阶段中,应用时不变权重来构造多通道空间滤波器,即波束形成器。在第二阶段中,时变单通道后置滤波器应用于波束形成器输出。此外,我们将其成功的应用程序在嘈杂和混响的环境中,以进一步提高语音增强的启发,注意力机制。 此外,我们的研究填补了现有文献中的空白,进行了彻底的空间分析网络的性能。具体来说,我们研究了网络在处理过程中如何利用空间信息。这种分析对网络的功能产生了有价值的见解,从而增强了我们对其整体性能的理解。 实验结果表明,我们的方法不仅是简单的训练,但也产生优越的结果,避免了事先知道的扬声器的活动的必要性。
摘要:This paper introduces an explainable DNN-based beamformer with a postfilter(ExNet-BF+PF) for multichannel signal processing. Our approach combines theU-Net network with a beamformer structure to address this problem. The methodinvolves a two-stage processing pipeline. In the first stage, time-invariantweights are applied to construct a multichannel spatial filter, namely abeamformer. In the second stage, a time-varying single-channel post-filter isapplied at the beamformer output. Additionally, we incorporate an attentionmechanism inspired by its successful application in noisy and reverberantenvironments to improve speech enhancement further. Furthermore, our study fills a gap in the existing literature by conducting athorough spatial analysis of the network's performance. Specifically, weexamine how the network utilizes spatial information during processing. Thisanalysis yields valuable insights into the network's functionality, therebyenhancing our understanding of its overall performance. Experimental results demonstrate that our approach is not onlystraightforward to train but also yields superior results, obviating thenecessity for prior knowledge of the speaker's activity.
标题:使用预训练模型进行双语文本相关说话人验证,以应对2024年TdSV挑战
链接:https://arxiv.org/abs/2411.10828
备注:5 pages, no figures
摘要:本文介绍了我们提交给2024年文本相关扬声器验证挑战赛(TdSV)伊朗分部的意见。TdSV旨在确定目标说话者是否说出了特定短语。我们基于预训练模型开发了两个独立的子系统:对于短语验证,短语分类器拒绝不正确的短语,而对于说话人验证,预训练的ResNet 293具有域自适应提取说话人嵌入,用于计算余弦相似度分数。此外,我们评估了Whisper-PMFA,一种适用于说话人验证的预训练ASR模型,发现尽管它的性能优于随机初始化的ResNets,但它的性能低于预训练的ResNets,突出了大规模预训练的重要性。结果还表明,实现有竞争力的性能TdSV没有联合建模的说话人和文本是可能的。我们最好的系统在评估子集上实现了0.0358的MinDCF,并赢得了挑战。
摘要:This paper presents our submissions to the Iranian division of theText-dependent Speaker Verification Challenge (TdSV) 2024. TdSV aims todetermine if a specific phrase was spoken by a target speaker. We developed twoindependent subsystems based on pre-trained models: For phrase verification, aphrase classifier rejected incorrect phrases, while for speaker verification, apre-trained ResNet293 with domain adaptation extracted speaker embeddings forcomputing cosine similarity scores. In addition, we evaluated Whisper-PMFA, apre-trained ASR model adapted for speaker verification, and found that,although it outperforms randomly initialized ResNets, it falls short of theperformance of pre-trained ResNets, highlighting the importance of large-scalepre-training. The results also demonstrate that achieving competitiveperformance on TdSV without joint modeling of speaker and text is possible. Ourbest system achieved a MinDCF of 0.0358 on the evaluation subset and won thechallenge.
标题:字幕预设是否反映音乐语义一致?
链接:https://arxiv.org/abs/2411.11692
备注:International Society for Music Information Retrieval (ISMIR) 2024, Late Breaking Demo (LBD)
摘要:随着先进语言生成模型的出现,音乐字幕已经成为一项很有前途的任务。然而,音乐字幕的评估在很大程度上依赖于传统的指标,如BLEU,METEOR和ROUGE,这些指标是为其他领域开发的,没有适当的理由在这个新的领域使用。我们提出的情况下,传统的指标是容易受到语法变化,并表明他们不相关以及与人类的判断。通过解决这些问题,我们的目标是强调需要一个关键的重新评估如何音乐字幕进行评估。
摘要:Music captioning has emerged as a promising task, fueled by the advent ofadvanced language generation models. However, the evaluation of musiccaptioning relies heavily on traditional metrics such as BLEU, METEOR, andROUGE which were developed for other domains, without proper justification fortheir use in this new field. We present cases where traditional metrics arevulnerable to syntactic changes, and show they do not correlate well with humanjudgments. By addressing these issues, we aim to emphasize the need for acritical reevaluation of how music captions are assessed.
标题:使用语音分析作为年轻人抑郁症风险的早期指标
链接:https://arxiv.org/abs/2411.11541
备注:Submitted to ToaC
摘要:越来越多的文献报道抑郁症患者和对照组之间的语音质量差异。在这里,我们研究的可能性,使用语音分析作为早期预警信号的发展情绪障碍的年轻人。作为一个主要的跨学科的欧洲研究项目在四个国家(ECoWeB)的一部分,检查基于网络的预防计划,以减少年轻人抑郁症的风险的影响,我们分析了大量的声音的声音特征,在特定的一天,由参与者经历的情绪的声音报告。我们能够确定一些显着差异的声学线索,特别是在语音频谱中的能量分布,鼓励进一步的研究工作,以开发有前途的非侵扰性的风险指标,在正常说话的声音。这对于那些不太可能表现出抑郁症标准风险因素(如负面生活经历)的年轻人来说尤其重要。
摘要:Increasingly frequent publications in the literature report voice qualitydifferences between depressed patients and controls. Here, we examine thepossibility of using voice analysis as an early warning signal for thedevelopment of emotion disturbances in young adults. As part of a majorinterdisciplinary European research project in four countries (ECoWeB),examining the effects of web-based prevention programs to reduce the risk fordepression in young adults, we analyzed a large number of acoustic voicecharacteristics in vocal reports of emotions experienced by the participants ona specific day. We were able to identify a number of significant differences inacoustic cues, particularly with respect to the energy distribution in thevoice spectrum, encouraging further research efforts to develop promisingnon-obtrusive risk indicators in the normal speaking voice. This isparticularly important in the case of young adults who are less likely toexhibit standard risk factors for depression such as negative life experiences.
标题:CEEMDAN在欠确定语音分离中的性能研究
链接:https://arxiv.org/abs/2411.11312
备注:in Arabic language
摘要:CEEMDAN算法是分析非平稳信号的现代方法之一。本研究提出了一个研究的有效性,这种方法在音频源分离知道其工作的限制。得出了CEEMDAN分离混合信号的频率和幅度的两个条件。研究了该算法在语音噪声分离和语音信号分离中的性能。研究得出结论,CEEMDAN可以从语音中去除某些类型的噪声(语音改善),但不能将语音信号彼此分离(鸡尾酒会)。仿真使用Matlab环境和Noizeus数据库。
摘要:The CEEMDAN algorithm is one of the modern methods used in the analysis ofnon-stationary signals. This research presents a study of the effectiveness ofthis method in audio source separation to know the limits of its work. Itconcluded two conditions related to frequencies and amplitudes of mixed signalsto be separated by CEEMDAN. The performance of the algorithm in separatingnoise from speech and separating speech signals from each other is studied. Theresearch reached a conclusion that CEEMDAN can remove some types of noise fromspeech (speech improvement), and it cannot separate speech signals from eachother (cocktail party). Simulation is done using Matlab environment and Noizeusdatabase.
标题:ESTVocoder:一种基于Mel频谱图的兴奋光谱变换神经声码器
链接:https://arxiv.org/abs/2411.11258
备注:Accepted by NCMMSC2024
摘要:本文提出了EST声码器,一种新的激励谱变换的神经声码器的源滤波器理论的框架内。ESTVocoder使用神经滤波器将激励的幅度和相位谱转换为相应的语音幅度和相位谱,神经滤波器的主干是ConvNeXt v2块。最后通过短时傅立叶逆变换(ISTFT)重构语音波形。激励是基于F0构造的:对于有声片段,它包含完整的谐波信息,而对于无声片段,它由噪声表示。激励为滤波器提供了幅度和相位模式的先验知识,与传统的神经声码器相比,期望降低建模难度。为了保证合成语音的保真度,采用了对抗训练策略,使EST声码器具有多尺度和多分辨率的鉴别器。分析合成和文本到语音的实验都证实,我们提出的EST声码器优于或相当于其他基线神经声码器,例如,HiFi-GAN、SiFi-GAN和Vocos在合成语音质量方面具有合理的模型复杂度和生成速度。另外的分析实验也表明,引入的激励有效地加快了模型的收敛过程,由于语音频谱的先验信息包含在激励。
摘要:This paper proposes ESTVocoder, a novel excitation-spectral-transformedneural vocoder within the framework of source-filter theory. The ESTVocodertransforms the amplitude and phase spectra of the excitation into thecorresponding speech amplitude and phase spectra using a neural filter whosebackbone is ConvNeXt v2 blocks. Finally, the speech waveform is reconstructedthrough the inverse short-time Fourier transform (ISTFT). The excitation isconstructed based on the F0: for voiced segments, it contains full harmonicinformation, while for unvoiced segments, it is represented by noise. Theexcitation provides the filter with prior knowledge of the amplitude and phasepatterns, expecting to reduce the modeling difficulty compared to conventionalneural vocoders. To ensure the fidelity of the synthesized speech, anadversarial training strategy is applied to ESTVocoder with multi-scale andmulti-resolution discriminators. Analysis-synthesis and text-to-speechexperiments both confirm that our proposed ESTVocoder outperforms or iscomparable to other baseline neural vocoders, e.g., HiFi-GAN, SiFi-GAN, andVocos, in terms of synthesized speech quality, with a reasonable modelcomplexity and generation speed. Additional analysis experiments alsodemonstrate that the introduced excitation effectively accelerates the model'sconvergence process, thanks to the speech spectral prior information containedin the excitation.
标题:SAMOS:一种利用语义表示和声学特征的神经MOS预测模型
链接:https://arxiv.org/abs/2411.11232
摘要:使用平均意见评分(MOS)预测模型来评估语音的自然度对于语音合成系统的自动评估具有积极的意义。早期的MOS预测模型将语音的原始波形或幅度谱作为输入,而更先进的方法采用基于自监督学习(SSL)的模型从语音中提取语义表示以进行MOS预测。这些方法利用语音信息的有限方面进行MOS预测,导致有限的预测精度。因此,在本文中,我们提出了SAMOS,MOS预测模型,利用语音的语义和声学信息进行评估。具体而言,建议的SAMOS利用预训练的wav2vec2来提取语义表示,并使用预训练的BiVocoder的特征提取器来提取声学特征。然后,这两种类型的特征被馈送到预测网络中,该网络包括多任务头和聚合层,以获得最终的MOS分数。实验结果表明,根据系统级评估指标的结果,所提出的SAMOS在BVCC数据集上的性能优于当前最先进的MOS预测模型,并且在BC2019数据集上的性能相当。
摘要:Assessing the naturalness of speech using mean opinion score (MOS) predictionmodels has positive implications for the automatic evaluation of speechsynthesis systems. Early MOS prediction models took the raw waveform oramplitude spectrum of speech as input, whereas more advanced methods employedself-supervised-learning (SSL) based models to extract semantic representationsfrom speech for MOS prediction. These methods utilized limited aspects ofspeech information for MOS prediction, resulting in restricted predictionaccuracy. Therefore, in this paper, we propose SAMOS, a MOS prediction modelthat leverages both Semantic and Acoustic information of speech to be assessed.Specifically, the proposed SAMOS leverages a pretrained wav2vec2 to extractsemantic representations and uses the feature extractor of a pretrainedBiVocoder to extract acoustic features. These two types of features are thenfed into the prediction network, which includes multi-task heads and anaggregation layer, to obtain the final MOS score. Experimental resultsdemonstrate that the proposed SAMOS outperforms current state-of-the-art MOSprediction models on the BVCC dataset and performs comparable performance onthe BC2019 dataset, according to the results of system-level evaluationmetrics.
标题:水的声音:从倾倒的液体推断物理性质
链接:https://arxiv.org/abs/2411.11222
备注:25 pages, 17 figures. Project page at this https URL
摘要:我们研究视听观察和一个平凡但有趣的日常活动的基础物理之间的联系:倒液体。只给出液体倒入容器的声音,我们的目标是自动推断物理属性,如液位,容器的形状和大小,倾倒速度和填充时间。为此,我们:(i)在理论上示出这些属性可以从基频(音调)确定;(ii)训练音调检测模型,该音调检测模型具有来自模拟数据和具有物理启发目标的视觉数据的监督;(iii)引入真实浇注视频的新的大数据集用于系统研究;(iv)示出训练的模型确实可以推断真实数据的这些物理属性;最后,(v)我们展示了对各种容器形状,其他数据集和野外YouTube视频的强大泛化。我们的工作提出了一个狭窄而丰富的问题在声学,物理学和学习的交叉点的敏锐的理解。它开辟了在机器人浇注中增强多感官感知的应用。
摘要:We study the connection between audio-visual observations and the underlyingphysics of a mundane yet intriguing everyday activity: pouring liquids. Givenonly the sound of liquid pouring into a container, our objective is toautomatically infer physical properties such as the liquid level, the shape andsize of the container, the pouring rate and the time to fill. To this end, we:(i) show in theory that these properties can be determined from the fundamentalfrequency (pitch); (ii) train a pitch detection model with supervision fromsimulated data and visual data with a physics-inspired objective; (iii)introduce a new large dataset of real pouring videos for a systematic study;(iv) show that the trained model can indeed infer these physical properties forreal data; and finally, (v) we demonstrate strong generalization to variouscontainer shapes, other datasets, and in-the-wild YouTube videos. Our workpresents a keen understanding of a narrow yet rich problem at the intersectionof acoustics, physics, and learning. It opens up applications to enhancemultisensory perception in robotic pouring.
标题:利用偏差纠正和模型融合的音调和光谱感知歌唱质量评估
链接:https://arxiv.org/abs/2411.11123
摘要:我们参加了2024年VoiceMOS挑战赛的第二场比赛,旨在预测歌唱样本的平均意见得分(MOS)。我们的参赛作品在所有参赛队伍中获得第一名,不包括官方基线。在本文中,我们进一步改进我们的意见,并提出了一种新的音高和频谱感知歌唱质量评估(PS-SQA)方法。PS-SQA是基于自监督学习MOS预测器设计的,结合了歌唱音高和频谱信息,分别使用音高直方图和非量化神经编解码器提取。此外,PS-SQA引入了偏差校正策略来解决低资源训练样本引起的预测偏差,并采用模型融合技术来进一步提高预测精度。实验结果证实,我们提出的PS-SQA显着优于所有竞争系统在所有系统级指标,确认其强大的歌唱质量评估能力。
摘要:We participated in track 2 of the VoiceMOS Challenge 2024, which aimed topredict the mean opinion score (MOS) of singing samples. Our submission securedthe first place among all participating teams, excluding the official baseline.In this paper, we further improve our submission and propose a novelPitch-and-Spectrum-aware Singing Quality Assessment (PS-SQA) method. The PS-SQAis designed based on the self-supervised-learning (SSL) MOS predictor,incorporating singing pitch and spectral information, which are extracted usingpitch histogram and non-quantized neural codec, respectively. Additionally, thePS-SQA introduces a bias correction strategy to address prediction biasescaused by low-resource training samples, and employs model fusion technology tofurther enhance prediction accuracy. Experimental results confirm that ourproposed PS-SQA significantly outperforms all competing systems across allsystem-level metrics, confirming its strong sing quality assessmentcapabilities.
标题:跨语言音素合成(IPC):增强第二语言发音的理论和计算方法
链接:https://arxiv.org/abs/2411.10927
备注:10 pages, 6 Figures, submitted to ACL ARR October 2024 for NAACL 2025
摘要:第二语言(L2)的学习者经常无意识地用母语(L1)中相似的音素替代不熟悉的L2音素,即使母语为L2的人认为这些声音是不同的和不可互换的。这种音位替换导致偏离第二语言的标准语音模式,给学习者获得准确的第二语言发音带来了挑战。为了解决这个问题,我们提出了跨语言语音合成(IPC),一种新的计算方法,旨在通过重建L2音素作为来自多个L1音素的复合声音,以最大限度地减少不正确的语音转移。使用两种自动语音识别模型进行的测试表明,当第二语言说话者发出IPC生成的复合声音时,与其发音受到原始语音迁移模式影响时相比,目标第二语言音素的识别率提高了20%。在相对较短的时间范围内观察到改善,表明快速获得复合声音。
摘要:Learners of a second language (L2) often unconsciously substitute unfamiliarL2 phonemes with similar phonemes from their native language (L1), even thoughnative speakers of the L2 perceive these sounds as distinct andnon-interchangeable. This phonemic substitution leads to deviations from thestandard phonological patterns of the L2, creating challenges for learners inacquiring accurate L2 pronunciation. To address this, we proposeInter-linguistic Phonetic Composition (IPC), a novel computational methoddesigned to minimize incorrect phonological transfer by reconstructing L2phonemes as composite sounds derived from multiple L1 phonemes. Tests with twoautomatic speech recognition models demonstrated that when L2 speakers producedIPC-generated composite sounds, the recognition rate of target L2 phonemesimproved by 20% compared to when their pronunciation was influenced by originalphonological transfer patterns. The improvement was observed within arelatively shorter time frame, demonstrating rapid acquisition of the compositesound.
标题:BanglaDialecto:端到端人工智能驱动的区域语音标准化
链接:https://arxiv.org/abs/2411.10879
备注:Accepted in 2024 IEEE International Conference on Big Data (IEEE BigData)
摘要:本研究的重点是识别孟加拉方言和转换成标准化的正式孟加拉语口音。方言,通常被称为区域语言,是在特定地点使用的语言的独特变体,并通过其语音,发音和词汇来识别。语音和语调的细微变化也受到地理位置、教育程度和社会经济地位的影响。方言标准化是必要的,以确保有效的沟通,教育的一致性,获得技术,经济机会和保护语言资源,同时尊重文化多样性。作为第五大语言,有1.6亿人使用大约55种不同的方言,解决孟加拉方言对于开发包容性通信工具至关重要。然而,由于缺乏全面的数据集以及处理不同方言的挑战,研究有限。随着多语言大语言模型(MLLM)的发展,已经创造了新的可能性来解决方言自动语音识别(ASR)和机器翻译(MT)的挑战。这项研究提出了一个端到端的管道转换方言Noakhali语音标准孟加拉语语音。这项调查包括构建一个大规模的方言语音信号的多样化数据集,该数据集在ASR和LLM中定制微调过程,用于将方言语音转录为方言文本,并将方言文本翻译为标准孟加拉语文本。我们的实验表明,微调Whisper ASR模型实现了0.8%的CER和1.5%的WER,而BanglaT5模型在方言到标准文本翻译中获得了41.6%的BLEU分数。
摘要:This study focuses on recognizing Bangladeshi dialects and converting diverseBengali accents into standardized formal Bengali speech. Dialects, oftenreferred to as regional languages, are distinctive variations of a languagespoken in a particular location and are identified by their phonetics,pronunciations, and lexicon. Subtle changes in pronunciation and intonation arealso influenced by geographic location, educational attainment, andsocioeconomic status. Dialect standardization is needed to ensure effectivecommunication, educational consistency, access to technology, economicopportunities, and the preservation of linguistic resources while respectingcultural diversity. Being the fifth most spoken language with around 55distinct dialects spoken by 160 million people, addressing Bangla dialects iscrucial for developing inclusive communication tools. However, limited researchexists due to a lack of comprehensive datasets and the challenges of handlingdiverse dialects. With the advancement in multilingual Large Language Models(mLLMs), emerging possibilities have been created to address the challenges ofdialectal Automated Speech Recognition (ASR) and Machine Translation (MT). Thisstudy presents an end-to-end pipeline for converting dialectal Noakhali speechto standard Bangla speech. This investigation includes constructing alarge-scale diverse dataset with dialectal speech signals that tailored thefine-tuning process in ASR and LLM for transcribing the dialect speech todialect text and translating the dialect text to standard Bangla text. Ourexperiments demonstrated that fine-tuning the Whisper ASR model achieved a CERof 0.8% and WER of 1.5%, while the BanglaT5 model attained a BLEU score of41.6% for dialect-to-standard text translation.
