微信公众号:arXiv_Daily
cs.SD语音
【1】The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
标题:AudioMOS Challenge 2025的T12系统:使用基于KAN和VERSA的模型的音频美学分数预测系统
链接:https://arxiv.org/abs/2512.05592
备注:Accepted by IEEE ASRU 2025
摘要:我们提出了一个由CyberAgent(AESCA)为AudioMOS Challenge 2025(AMC 25)Track 2设计的音频美学评分(AES)预测系统。AESCA包括基于Kolmogorov-Arnold网络(KAN)的音频盒美学和使用VERSA工具包的度量分数的预测器。在基于KAN的预测器中,我们用组理性KAN替换了基线模型中的每个多层感知器层,并用标记和伪标记的音频样本训练模型。基于VERSA的预测器被设计为使用极端梯度提升的回归模型,结合了现有指标的输出。基于KAN和VERSA的模型都预测了AES,包括四个评估轴。最终的AES值是使用一个集合模型计算的,该模型结合了四个基于KAN的模型和一个基于VERSA的模型。我们提出的T12系统在所提交的系统中产生了最好的相关性,在话语水平的三个轴,在系统水平的两个轴,以及整体平均值。
摘要:We propose an audio aesthetics score (AES) prediction system by CyberAgent (AESCA) for AudioMOS Challenge 2025 (AMC25) Track 2. The AESCA comprises a Kolmogorov--Arnold Network (KAN)-based audiobox aesthetics and a predictor from the metric scores using the VERSA toolkit. In the KAN-based predictor, we replaced each multi-layer perceptron layer in the baseline model with a group-rational KAN and trained the model with labeled and pseudo-labeled audio samples. The VERSA-based predictor was designed as a regression model using extreme gradient boosting, incorporating outputs from existing metrics. Both the KAN- and VERSA-based models predicted the AES, including the four evaluation axes. The final AES values were calculated using an ensemble model that combined four KAN-based models and a VERSA-based model. Our proposed T12 system yielded the best correlations among the submitted systems, in three axes at the utterance level, two axes at the system level, and the overall average.
【2】Lyrics Matter: Exploiting the Power of Learnt Representations for Music Popularity Prediction
标题:歌词很重要:利用学习表征的力量进行音乐流行预测
链接:https://arxiv.org/abs/2512.05508
备注:8 pages
摘要:准确预测音乐受欢迎程度是音乐行业的一个关键挑战,为艺术家、制作人和流媒体平台带来好处。之前的研究主要集中在音频功能,社交元数据或模型架构上。这项工作解决了在预测流行的歌词未充分发掘的作用。我们提出了一个自动化的管道,使用LLM提取高维歌词嵌入,捕捉语义,语法和顺序信息。这些功能集成到HitMusicLyricNet中,HitMusicLyricNet是一种多模式架构,它结合了音频,歌词和社交元数据,用于在0-100范围内进行流行度评分预测。我们的方法在SpotGenTrack数据集上的表现优于现有的基线,该数据集包含超过100,000个跟踪,分别在MAE和MSE方面实现了9%和20%的改进。消融证实,收益来自我们的LLM驱动的歌词功能管道(LyricsAENet),强调了密集的歌词表示的价值。
摘要:Accurately predicting music popularity is a critical challenge in the music industry, offering benefits to artists, producers, and streaming platforms. Prior research has largely focused on audio features, social metadata, or model architectures. This work addresses the under-explored role of lyrics in predicting popularity. We present an automated pipeline that uses LLM to extract high-dimensional lyric embeddings, capturing semantic, syntactic, and sequential information. These features are integrated into HitMusicLyricNet, a multimodal architecture that combines audio, lyrics, and social metadata for popularity score prediction in the range 0-100. Our method outperforms existing baselines on the SpotGenTrack dataset, which contains over 100,000 tracks, achieving 9% and 20% improvements in MAE and MSE, respectively. Ablation confirms that gains arise from our LLM-driven lyrics feature pipeline (LyricsAENet), underscoring the value of dense lyric representations.
【3】MuMeNet: A Network Simulator for Musical Metaverse Communications
标题:MuMeNet:音乐元宇宙通信的网络模拟器
链接:https://arxiv.org/abs/2512.05201
备注:To be published in 2025 IEEE 6th International Symposium on the Internet of Sounds (IS2)
摘要:Metaverse是一个共享和空间组织的数字连续体,正在改变各个行业,音乐正在成为一个领先的用例。现场音乐会、协作作曲和互动体验正在推动音乐虚拟世界(MM)的发展,但底层网络和服务基础设施的需求阻碍了它的发展。这些挑战强调了需要一个新的建模和仿真范例定制的MM会话的独特特性,以及专门的服务提供策略,能够捕捉他们的互动,异构和面向组播的性质。为此,我们首次尝试正式建模和分析5G/6 G网络中MM会话的服务提供问题。我们首先正式的服务和网络图模型的MM,使用“现场观众互动的虚拟音乐会”作为参考方案。然后,我们提出了MuMeNet,一种新的离散事件网络模拟器,专门针对MM的要求和流量动态。我们展示了MuMeNet的有效性,通过运行基于线性规划的编排策略的参考方案,并提供现实的MM工作负载下的性能分析。
摘要:The Metaverse, a shared and spatially organized digital continuum, is transforming various industries, with music emerging as a leading use case. Live concerts, collaborative composition, and interactive experiences are driving the Musical Metaverse (MM), but the requirements of the underlying network and service infrastructures hinder its growth. These challenges underscore the need for a novel modeling and simulation paradigm tailored to the unique characteristics of MM sessions, along with specialized service provisioning strategies capable of capturing their interactive, heterogeneous, and multicast-oriented nature. To this end, we make a first attempt to formally model and analyze the problem of service provisioning for MM sessions in 5G/6G networks. We first formalize service and network graph models for the MM, using "live audience interaction in a virtual concert" as a reference scenario. We then present MuMeNet, a novel discrete-event network simulator specifically tailored to the requirements and the traffic dynamics of the MM. We showcase the effectiveness of MuMeNet by running a linear programming based orchestration policy on the reference scenario and providing performance analysis under realistic MM workloads.
【4】Decoding Selective Auditory Attention to Musical Elements in Ecologically Valid Music Listening
标题:解析哲学有效音乐聆听中对音乐元素的选择性听觉注意
链接:https://arxiv.org/abs/2512.05528
摘要:长期以来,艺术在塑造人类情感、认知和行为方面发挥着深远的作用。虽然绘画和建筑等视觉艺术已经通过眼动追踪进行了研究,揭示了专家和新手之间不同的凝视模式,但听觉艺术形式的类似方法仍然不发达。尽管音乐是现代生活和文化的一个普遍组成部分,但仍然缺乏客观的工具来量化自然听觉体验期间听众的注意力和感知焦点。据我们所知,这是第一次尝试使用自然主义的、工作室制作的歌曲和只有四个电极的轻量级消费级脑电图设备来解码对音乐元素的选择性注意。通过分析在真实世界中的神经反应,如音乐听,我们测试解码是否是可行的条件下,最大限度地减少参与者的负担,并保持音乐体验的真实性。我们的贡献有四个方面:(i)解码真正的录音室制作的歌曲中的音乐注意力,(ii)用四通道消费者EEG证明可行性,(iii)提供对音乐注意力解码的见解,以及(iv)证明与先前工作相比改进的模型能力。我们的研究结果表明,音乐注意力不仅可以解码为新的歌曲,而且可以跨新的主题,在我们的测试条件下,与现有的方法相比,表现出性能的改善。这些发现表明,消费级设备可以可靠地捕获信号,并且音乐中的神经解码在现实世界中是可行的。这为教育、个性化音乐技术和治疗干预的应用铺平了道路。
摘要:Art has long played a profound role in shaping human emotion, cognition, and behavior. While visual arts such as painting and architecture have been studied through eye tracking, revealing distinct gaze patterns between experts and novices, analogous methods for auditory art forms remain underdeveloped. Music, despite being a pervasive component of modern life and culture, still lacks objective tools to quantify listeners' attention and perceptual focus during natural listening experiences. To our knowledge, this is the first attempt to decode selective attention to musical elements using naturalistic, studio-produced songs and a lightweight consumer-grade EEG device with only four electrodes. By analyzing neural responses during real world like music listening, we test whether decoding is feasible under conditions that minimize participant burden and preserve the authenticity of the musical experience. Our contributions are fourfold: (i) decoding music attention in real studio-produced songs, (ii) demonstrating feasibility with a four-channel consumer EEG, (iii) providing insights for music attention decoding, and (iv) demonstrating improved model ability over prior work. Our findings suggest that musical attention can be decoded not only for novel songs but also across new subjects, showing performance improvements compared to existing approaches under our tested conditions. These findings show that consumer-grade devices can reliably capture signals, and that neural decoding in music could be feasible in real-world settings. This paves the way for applications in education, personalized music technologies, and therapeutic interventions.
【5】SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model
标题:同步语音:使用视觉增强预训练的TTC模型实现视频配音
链接:https://arxiv.org/abs/2512.05126
摘要:视频配音旨在生成与视觉内容在时间上精确对齐的高保真语音。现有的方法仍然受到语音自然度和视听同步方面的限制,并且仅限于单语环境。为了应对这些挑战,我们提出了SyncVoice,这是一个基于预训练的文本到语音(TTS)模型的视觉增强视频配音框架。通过对TTS模型在视听数据上的微调,我们实现了很强的视听一致性。本文提出了一种双说话人编码器来有效地抑制跨语言语音合成中的语际干扰,并探索了视频配音在视频翻译中的应用。实验结果表明,SyncVoice实现了高保真的语音生成,具有较强的同步性能,在视频配音任务中显示了其潜力。
摘要:Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to monolingual settings. To address these challenges, we propose SyncVoice, a vision-augmented video dubbing framework built upon a pretrained text-to-speech (TTS) model. By fine-tuning the TTS model on audio-visual data, we achieve strong audiovisual consistency. We propose a Dual Speaker Encoder to effectively mitigate inter-language interference in cross-lingual speech synthesis and explore the application of video dubbing in video translation scenarios. Experimental results show that SyncVoice achieves high-fidelity speech generation with strong synchronization performance, demonstrating its potential in video dubbing tasks.
【1】Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
标题:言语世界模型:具有言语显式推理的因果状态行动规划
链接:https://arxiv.org/abs/2512.05933
摘要:当前的语音语言模型(SLM)通常使用语音编码器和大型语言模型的级联,将语音理解视为单个黑盒。他们能很好地分析演讲的内容,但对其他方面的推理能力很弱,尤其是在缺乏监督的情况下。因此,我们认为,明确的推理语言状态和行动模块化和透明的决定。受认知科学的启发,我们采用了模块化的观点和世界模型的观点,在这种观点中,系统在潜在状态上学习正向动力学。我们将语音理解分解为四个模块,通过因果图进行通信,建立认知状态搜索空间。在这个空间的后验痕迹的指导下,一个经过调整的语言模型产生了一个简洁的因果分析和一个面向用户的响应,从而在部分监督下实现了反事实干预和可解释性。我们提出了第一个基于图的模块化语音模型,用于显式推理,我们将开源模型和数据,以促进高级语音理解的发展。
摘要:Current speech-language models (SLMs) typically use a cascade of speech encoder and large language model, treating speech understanding as a single black box. They analyze the content of speech well but reason weakly about other aspects, especially under sparse supervision. Thus, we argue for explicit reasoning over speech states and actions with modular and transparent decisions. Inspired by cognitive science we adopt a modular perspective and a world model view in which the system learns forward dynamics over latent states. We factorize speech understanding into four modules that communicate through a causal graph, establishing a cognitive state search space. Guided by posterior traces from this space, an instruction-tuned language model produces a concise causal analysis and a user-facing response, enabling counterfactual interventions and interpretability under partial supervision. We present the first graph based modular speech model for explicit reasoning and we will open source the model and data to promote the development of advanced speech understanding.
【2】A Multi-Channel Auditory Signal Encoder with Adaptive Resolution Using Volatile Memristors
标题:使用易失性记忆电阻的自适应分辨率的多通道听觉信号编码器
链接:https://arxiv.org/abs/2512.05701
备注:11 pages, 17 figures, submitted to IEEE Transactions on Circuits and Systems I: Regular Papers for possible publications
摘要:我们演示和实验验证的端到端的混合CMOS忆阻器听觉编码器,实现自适应阈值,异步增量调制(ADM)为基础的尖峰编码,利用HfTiOx设备的固有波动性。尖峰触发的编程脉冲迅速提高ADM阈值Delta(脱敏);当活动减弱时,器件的波动性被动降低Delta(再敏化),在没有静态控制能量的情况下恢复灵敏度的同时强调起始。我们的原型通过开关接口和片外控制器将8通道130 nm编码器IC耦合到片外HfTiOx器件,该控制器可监控尖峰活动并发出编程事件。片内电流镜跨导放大器(TIA)将器件电流转换为对称阈值,从而实现灵敏和保守的编码机制。用伽马滤波语音进行评估,自适应环路在匹配的尖峰脉冲处锐化起始点并保留固定Delta基线错过的精细时间细节;多通道尖峰耳蜗图显示出相同的趋势。总之,这些结果建立了一个实用的混合CMOS忆阻器通路,以启动突出,尖峰高效的神经形态音频前端,并激励低功耗单芯片集成。
摘要:We demonstrate and experimentally validate an end-to-end hybrid CMOS-memristor auditory encoder that realises adaptive-threshold, asynchronous delta-modulation (ADM)-based spike encoding by exploiting the inherent volatility of HfTiOx devices. A spike-triggered programming pulse rapidly raises the ADM threshold Delta (desensitisation); the device's volatility then passively lowers Delta when activity subsides (resensitisation), emphasising onsets while restoring sensitivity without static control energy. Our prototype couples an 8-channel 130 nm encoder IC to off-chip HfTiOx devices via a switch interface and an off-chip controller that monitors spike activity and issues programming events. An on-chip current-mirror transimpedance amplifier (TIA) converts device current into symmetric thresholds, enabling both sensitive and conservative encoding regimes. Evaluated with gammatone-filtered speech, the adaptive loop-at matched spike budget-sharpens onsets and preserves fine temporal detail that a fixed-Delta baseline misses; multi-channel spike cochleagrams show the same trend. Together, these results establish a practical hybrid CMOS-memristor pathway to onset-salient, spike-efficient neuromorphic audio front-ends and motivate low-power single-chip integration.
【3】Noise Suppression for Time Difference of Arrival: Performance Evaluation of a Generalized Cross-Correlation Method Using Mean Signal and Inverse Filter
标题:到达时间差的噪声抑制:使用平均信号和逆滤波器的广义互相关方法的性能评估
链接:https://arxiv.org/abs/2512.05355
摘要:本文提出了一种新的广义互相关(GCC)方法,称为GCC-MSIF,以提高到达时间差(TDOA)估计精度在噪声环境中。传统的GCC方法在低信噪比(SNR)条件下,特别是当信号带宽未知时,经常遭受性能降级。GCC-MSIF引入了从多通道输入估计的“平均信号”和“逆滤波器”来虚拟地重建源信号,从而实现带外噪声的自适应抑制。小规模阵列的数值模拟表明,GCC-MSIF显着优于传统的方法,如GCC-PHAT和GCC-SCOT,在低信噪比区域,并实现鲁棒性相当或超过最大似然(GCC-ML)方法。此外,估计精度随着阵列元素的数量而可缩放地提高。这些结果表明,GCC-MSIF是一个很有前途的解决方案,在实际的盲环境中的鲁棒被动定位。
摘要:This paper proposes a novel generalized cross-correlation (GCC) method, termed GCC-MSIF, to improve time difference of arrival (TDOA) estimation accuracy in noisy environments. Conventional GCC methods often suffer from performance degradation under low signal-to-noise ratio (SNR) conditions, particularly when the signal bandwidth is unknown. GCC-MSIF introduces a "mean signal" estimated from multi-channel inputs and an "inverse filter" to virtually reconstruct the source signal, enabling adaptive suppression of out-of-band noise. Numerical simulations simulating a small-scale array demonstrate that GCC-MSIF significantly outperforms conventional methods, such as GCC-PHAT and GCC-SCOT, in low SNR regions and achieves robustness comparable to or exceeding the maximum likelihood (GCC-ML) method. Furthermore, the estimation accuracy improves scalably with the number of array elements. These results suggest that GCC-MSIF is a promising solution for robust passive localization in practical blind environments.
【4】SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model
标题:同步语音:使用视觉增强预训练的TTC模型实现视频配音
链接:https://arxiv.org/abs/2512.05126
摘要:视频配音旨在生成与视觉内容在时间上精确对齐的高保真语音。现有的方法仍然受到语音自然度和视听同步的限制,并且仅限于单语设置。为了应对这些挑战,我们提出了SyncVoice,这是一个基于预训练的文本到语音(TTS)模型的视觉增强视频配音框架。通过对TTS模型在视听数据上的微调,我们实现了很强的视听一致性。本文提出了一种双说话人编码器来有效地抑制跨语言语音合成中的语际干扰,并探索了视频配音在视频翻译中的应用。实验结果表明,SyncVoice实现了高保真的语音生成,具有较强的同步性能,在视频配音任务中显示了其潜力。
摘要:Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to monolingual settings. To address these challenges, we propose SyncVoice, a vision-augmented video dubbing framework built upon a pretrained text-to-speech (TTS) model. By fine-tuning the TTS model on audio-visual data, we achieve strong audiovisual consistency. We propose a Dual Speaker Encoder to effectively mitigate inter-language interference in cross-lingual speech synthesis and explore the application of video dubbing in video translation scenarios. Experimental results show that SyncVoice achieves high-fidelity speech generation with strong synchronization performance, demonstrating its potential in video dubbing tasks.
【5】The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
标题:AudioMOS Challenge 2025的T12系统:使用基于KAN和VERSA的模型的音频美学分数预测系统
链接:https://arxiv.org/abs/2512.05592
备注:Accepted by IEEE ASRU 2025
摘要:我们提出了一个由CyberAgent(AESCA)为AudioMOS Challenge 2025(AMC 25)Track 2设计的音频美学评分(AES)预测系统。AESCA包括基于Kolmogorov-Arnold网络(KAN)的音频盒美学和使用VERSA工具包的度量分数的预测器。在基于KAN的预测器中,我们用组理性KAN替换了基线模型中的每个多层感知器层,并用标记和伪标记的音频样本训练模型。基于VERSA的预测器被设计为使用极端梯度提升的回归模型,结合了现有指标的输出。基于KAN和VERSA的模型都预测了AES,包括四个评估轴。最终的AES值是使用一个集合模型计算的,该模型结合了四个基于KAN的模型和一个基于VERSA的模型。我们提出的T12系统在所提交的系统中产生了最好的相关性,在话语水平的三个轴,在系统水平的两个轴,以及整体平均值。
摘要:We propose an audio aesthetics score (AES) prediction system by CyberAgent (AESCA) for AudioMOS Challenge 2025 (AMC25) Track 2. The AESCA comprises a Kolmogorov--Arnold Network (KAN)-based audiobox aesthetics and a predictor from the metric scores using the VERSA toolkit. In the KAN-based predictor, we replaced each multi-layer perceptron layer in the baseline model with a group-rational KAN and trained the model with labeled and pseudo-labeled audio samples. The VERSA-based predictor was designed as a regression model using extreme gradient boosting, incorporating outputs from existing metrics. Both the KAN- and VERSA-based models predicted the AES, including the four evaluation axes. The final AES values were calculated using an ensemble model that combined four KAN-based models and a VERSA-based model. Our proposed T12 system yielded the best correlations among the submitted systems, in three axes at the utterance level, two axes at the system level, and the overall average.
【6】Decoding Selective Auditory Attention to Musical Elements in Ecologically Valid Music Listening
标题:解析哲学有效音乐聆听中对音乐元素的选择性听觉注意
链接:https://arxiv.org/abs/2512.05528
摘要:长期以来,艺术在塑造人类情感、认知和行为方面发挥着深远的作用。虽然绘画和建筑等视觉艺术已经通过眼动追踪进行了研究,揭示了专家和新手之间不同的凝视模式,但听觉艺术形式的类似方法仍然不发达。尽管音乐是现代生活和文化的一个普遍组成部分,但仍然缺乏客观的工具来量化自然听觉体验期间听众的注意力和感知焦点。据我们所知,这是第一次尝试使用自然主义的、工作室制作的歌曲和只有四个电极的轻量级消费级脑电图设备来解码对音乐元素的选择性注意。通过分析在真实世界中的神经反应,如音乐听,我们测试解码是否是可行的条件下,最大限度地减少参与者的负担,并保持音乐体验的真实性。我们的贡献有四个方面:(i)解码真正的录音室制作的歌曲中的音乐注意力,(ii)用四通道消费者EEG证明可行性,(iii)提供对音乐注意力解码的见解,以及(iv)证明与先前工作相比改进的模型能力。我们的研究结果表明,音乐注意力不仅可以解码为新的歌曲,而且可以跨新的主题,在我们的测试条件下,与现有的方法相比,表现出性能的改善。这些发现表明,消费级设备可以可靠地捕获信号,并且音乐中的神经解码在现实世界中是可行的。这为教育、个性化音乐技术和治疗干预的应用铺平了道路。
摘要:Art has long played a profound role in shaping human emotion, cognition, and behavior. While visual arts such as painting and architecture have been studied through eye tracking, revealing distinct gaze patterns between experts and novices, analogous methods for auditory art forms remain underdeveloped. Music, despite being a pervasive component of modern life and culture, still lacks objective tools to quantify listeners' attention and perceptual focus during natural listening experiences. To our knowledge, this is the first attempt to decode selective attention to musical elements using naturalistic, studio-produced songs and a lightweight consumer-grade EEG device with only four electrodes. By analyzing neural responses during real world like music listening, we test whether decoding is feasible under conditions that minimize participant burden and preserve the authenticity of the musical experience. Our contributions are fourfold: (i) decoding music attention in real studio-produced songs, (ii) demonstrating feasibility with a four-channel consumer EEG, (iii) providing insights for music attention decoding, and (iv) demonstrating improved model ability over prior work. Our findings suggest that musical attention can be decoded not only for novel songs but also across new subjects, showing performance improvements compared to existing approaches under our tested conditions. These findings show that consumer-grade devices can reliably capture signals, and that neural decoding in music could be feasible in real-world settings. This paves the way for applications in education, personalized music technologies, and therapeutic interventions.
机器翻译由腾讯交互翻译提供,仅供参考
