今日论文合集:cs.SD语音7篇,eess.AS音频处理4篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Perpetual Dialogues: A Computational Analysis of Voice-Guitar Interaction in Carlos Paredes's Discography
标题:永久对话:卡洛斯·帕雷德斯唱片中声音与吉他互动的计算分析
链接:https://arxiv.org/pdf/2603.12854v1

作者:Gilberto Bernardes,Nádia Moura,António Sá Pinto
备注:8 pages, 8 figures, to be published in ICMC 2026
摘要:计算音乐学使得能够系统地分析记录音乐中的表演和结构特征,然而现有的方法在很大程度上仍然是针对以乐谱为基础的曲目。本研究提出了一种方法来分析在卡洛斯帕雷德斯的声乐合作语音吉他互动-一个口头传统的背景下,组成和表演层共同出现。使用源分离的干,物理通知谐波建模,和节拍级音频描述符,我们研究旋律,谐波和节奏的关系,在八个录音与四个歌手。我们的共同性多样性框架,结合多尺度相关性分析与残差为基础的结构偏差检测,揭示了表达协调主要是件特定的,而不是整个语料库。多样性事件系统地与正式的边界和纹理的变化,表明所提出的方法可以识别音乐突出的重组与最小的人类注释。该框架进一步提供了一个可推广的计算策略,剧目没有标注的蓝图,扩展到口头传统和即兴的实践音乐表演分析。摘要:Computational musicology enables systematic analysis of performative and structural traits in recorded music, yet existing approaches remain largely tailored to notated, score-based repertoires. This study advances a methodology for analyzing voice-guitar interaction in Carlos Paredes's vocal collaborations - an oral-tradition context where compositional and performative layers co-emerge. Using source-separated stems, physics-informed harmonic modelling, and beat-level audio descriptors, we examine melodic, harmonic, and rhythmic relationships across eight recordings with four singers. Our commonality-diversity framework, combining multi-scale correlation analysis with residual-based detection of structural deviations, reveals that expressive coordination is predominantly piece-specific rather than corpus-wide. Diversity events systematically align with formal boundaries and textural shifts, demonstrating that the proposed approach can identify musically salient reorganizations with minimal human annotation. The framework further offers a generalizable computational strategy for repertoires without notated blueprints, extending Music Performance Analysis into oral-tradition and improvisation-inflected practices.


【2】DAST: A Dual-Stream Voice Anonymization Attacker with Staged Training
标题:DAST:经过阶段训练的双流语音匿名攻击者
链接:https://arxiv.org/pdf/2603.12840v1

作者:Ridwan Arefeen,Xiaoxiao Miao,Rong Tong,Aik Beng Ng,Simon See,Timothy Liu
摘要:语音匿名化掩盖了声音特征,同时保留了语言内容,这仍然可能泄漏特定于说话者的模式。为了评估和加强隐私评估,我们提出了一个双流攻击者,通过并行编码器融合频谱和自监督学习功能,并采用三阶段训练策略。第一阶段建立基本的说话者区分表征。第二阶段利用语音转换和匿名化的共享身份转换特性,将模型暴露给不同的转换语音,以建立跨系统的鲁棒性。第三阶段为目标匿名数据提供轻量级适配。语音隐私攻击者挑战(VPAC)数据集的结果表明,第二阶段是泛化的主要驱动力,能够在看不见的匿名数据集上实现强大的攻击性能。在第三阶段,仅对10%的目标匿名化数据集进行微调,就EER而言,超过了当前最先进的攻击者。摘要:Voice anonymization masks vocal traits while preserving linguistic content, which may still leak speaker-specific patterns. To assess and strengthen privacy evaluation, we propose a dual-stream attacker that fuses spectral and self-supervised learning features via parallel encoders with a three-stage training strategy. Stage I establishes foundational speaker-discriminative representations. Stage II leverages the shared identity-transformation characteristics of voice conversion and anonymization, exposing the model to diverse converted speech to build cross-system robustness. Stage III provides lightweight adaptation to target anonymized data. Results on the VoicePrivacy Attacker Challenge (VPAC) dataset demonstrate that Stage II is the primary driver of generalization, enabling strong attacking performance on unseen anonymization datasets. With Stage III, fine-tuning on only 10 % of the target anonymization dataset surpasses current state-of-the-art attackers in terms of EER.


【3】Mask2Flow-TSE: Two-Stage Target Speaker Extraction with Masking and Flow Matching
标题:Mask 2Flow-PSE:利用掩蔽和流匹配的两阶段目标说话人提取
链接:https://arxiv.org/pdf/2603.12837v1

作者:Junwon Moon,Hyunjin Choi,Hansol Park,Heeseung Kim,Kyuhong Shim

备注:Submitted to Interspeech 2026

摘要:目标说话人提取(TSE)是在给定参考话语的情况下,从重叠的语音混合物中提取目标说话人的语音。现有的方法通常分为两类:歧视性和生成性。判别式方法应用时频掩蔽进行快速推理,但通常会过度抑制目标信号,而生成式方法则以大量迭代步骤为代价合成高质量语音。我们提出了Mask 2Flow-TSE,一个两阶段的框架,结合了两种范式的优势。第一级采用判别掩蔽进行粗分离,第二级采用流匹配将输出细化为目标语音。与从高斯噪声合成语音的生成方法不同,我们的方法从掩蔽的频谱图开始,在单个推理步骤中实现高质量的重建。实验表明,Mask 2Flow-TSE实现了与现有的生成TSE方法相当的性能,约85 M参数。摘要:Target speaker extraction (TSE) extracts the target speaker's voice from overlapping speech mixtures given a reference utterance. Existing approaches typically fall into two categories: discriminative and generative. Discriminative methods apply time-frequency masking for fast inference but often over-suppress the target signal, while generative methods synthesize high-quality speech at the cost of numerous iterative steps. We propose Mask2Flow-TSE, a two-stage framework combining the strengths of both paradigms. The first stage applies discriminative masking for coarse separation, and the second stage employs flow matching to refine the output toward target speech. Unlike generative approaches that synthesize speech from Gaussian noise, our method starts from the masked spectrogram, enabling high-quality reconstruction in a single inference step. Experiments show that Mask2Flow-TSE achieves comparable performance to existing generative TSE methods with approximately 85M parameters.


【4】Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
标题:通过直接偏好优化实现日语SpeechLLM的言语有价值匹配
链接:https://arxiv.org/pdf/2603.12565v1

作者:Mengjie Zhao,Lianbo Liu,Yusuke Fujita,Hao Shi,Yuan Gao,Roman Koshkin,Yui Sudo
摘要:peechLLM通常将ASR训练的编码器与基于文本的LLM主干相结合,导致它们继承不适合文本到语音合成的写入式输出模式。这种不匹配在日语中尤为明显,日语的口语和书面语在礼貌标记、句末助词和句法复杂性方面存在很大差异。我们提出了一种基于偏好的对齐方法,以适应日语SpeechLLM的语音价值的输出:文本是简洁的,对话,并容易合成为自然语音。为了严格评估这一任务,我们介绍SpokenElyza,一个来自ELYZA任务100的日本语音价值的基准,由本地专家进行听觉验证。实验表明,我们的方法实现了SpokenElyza大幅改善,同时在很大程度上保留了原有的书面风格的评价性能。我们将发布SpokenElyza以支持未来对日语口语对话系统的研究。摘要:SpeechLLMs typically combine ASR-trained encoders with text-based LLM backbones, leading them to inherit written-style output patterns unsuitable for text-to-speech synthesis. This mismatch is particularly pronounced in Japanese, where spoken and written registers differ substantially in politeness markers, sentence-final particles, and syntactic complexity. We propose a preference-based alignment approach to adapt Japanese SpeechLLMs for speech-worthy outputs: text that is concise, conversational, and readily synthesized as natural speech. To rigorously evaluate this task, we introduce SpokenElyza, a benchmark for Japanese speech-worthiness derived from ELYZA-tasks-100 with auditory verification by native experts. Experiments show that our approach achieves substantial improvement on SpokenElyza while largely preserving performance on the original written-style evaluation. We will release SpokenElyza to support future research on Japanese spoken dialog systems.


【5】RadEar: A Self-Supervised RF Backscatter System for Voice Eavesdropping and Separation
标题:RadEar:一种用于语音发射和分离的自监督RF反向散射系统
链接:https://arxiv.org/pdf/2603.12446v1

作者:Qijun Wang,Peihao Yan,Chunqi Qian,Huacheng Zeng
备注:Accepted by IEEE INFOCOM 2026
摘要:语音对话中的语音下降对个人隐私和信息安全构成了越来越大的威胁。在本文中,我们提出了RadEar,一种新的RF后向散射为基础的系统,旨在使隐蔽的语音窃听通过墙壁。RadEar由两个关键组件组成:(i)一个无电池的RF反向散射标签,秘密部署在目标空间内,以及(ii)一个位于房间外的RF读取器,用于执行信号解调,语音分离和去噪。该标签采用紧凑的双谐振器设计,可实现高能效的频率调制,用于连续语音窃听,同时通过分离激发和反射频率来减轻自干扰。为了克服弱信号接收和语音重叠的挑战,RF阅读器采用自监督学习模型进行语音分离和去噪,使用基于混音的目标进行训练,而不需要地面真实标签。我们在现实世界中的场景中制造和评估RadEar,证明它的能力,恢复和分离人类语音与高保真度的实际约束下。摘要:Eavesdropping on voice conversations presents a growing threat to personal privacy and information security. In this paper, we present RadEar, a novel RF backscatter-based system designed to enable covert voice eavesdropping through walls. RadEar consists of two key components: (i) a batteryless RF backscatter tag covertly deployed inside the target space, and (ii) an RF reader located outside the room that performs signal demodulation, voice separation, and denoising. The tag features a compact, dual-resonator design that achieves energy-efficient frequency modulation for continuous voice eavesdropping while mitigating self-interference by separating excitation and reflection frequencies. To overcome the challenges of weak signal reception and overlapping speech, the RF reader employs self-supervised learning models for voice separation and denoising, trained using a remix-based objective without requiring ground-truth labels. We fabricate and evaluate RadEar in real-world scenarios, demonstrating its ability to recover and separate human speech with high fidelity under practical constraints.


【6】TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
标题:TASTE流媒体:迈向可流化的文本对齐语音令牌化和口语建模嵌入
链接:https://arxiv.org/pdf/2603.12350v1

作者:Liang-Hsuan Tseng,Hung-yi Lee

备注:Work in progress

摘要:文本-语音联合口语建模(SLM)旨在实现自然和智能的基于语音的交互,但开发这样的系统可能会遭受模态失配:语音单元序列比文本令牌长得多。之前的工作通过文本对齐标记化和嵌入(TASTE)减少了这一差距,产生了与文本对应部分长度对齐的语音标记。然而,对外部ASR系统的依赖和非因果解码器的使用限制了流传输的使用。为了解决这个问题,我们提出了TASTE-S,一个适合实时使用的TASTE的流扩展。TASTE-S将基于CTC的ASR模块集成到编码器中,实现即时双模态编码。我们还重新设计了单元解码器,以实现动态解码。通过联合训练,我们证明了TASTE-S与TASTE的性能相匹配,同时显着降低了延迟。进一步的研究表明,TASTE-S对transmitting保持鲁棒性,并且能够进行长形式的编码和解码。摘要:Text-speech joint spoken language modeling (SLM) aims at natural and intelligent speech-based interactions, but developing such a system may suffer from modality mismatch: speech unit sequences are much longer than text tokens. Prior work reduces this gap with text-aligned tokenization and embedding (TASTE), producing speech tokens that align in lengths with their textual counterparts. However, the dependence on an external ASR system and the use of a non-causal decoder limits streaming use. To address this limitation, we propose TASTE-S, a streamable extension of TASTE suitable for real-time usage. TASTE-S integrates a CTC-based ASR module into the encoder for instant dual-modality encoding. We also redesign the unit decoder to enable on-the-fly decoding. With joint training, we show that TASTE-S matches TASTE's performance while significantly reducing latency. Further investigations reveal that TASTE-S remains robust to transcriptions and enables long-form encoding and decoding.


【7】Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
标题:基于位置相关正交子空间的自监督语音模型语音上下文编码
链接:https://arxiv.org/pdf/2603.12642v1

作者:Kwanghee Choi,Eunjung Yeo,Cheol Jun Cho,David R. Mortensen,David Harwath
备注:Submitted to Interspeech 2026
摘要:基于transformer的自监督语音模型(S3 M)通常被描述为情境化的,但这意味着什么仍然不清楚。在这里,我们专注于如何一个单一的帧级S3 M表示可以编码电话及其周围的上下文。先前的工作已经表明,S3 M表示组成的音素;例如,语音向量,如清音,bilabiality和鼻音向量叠加在[m]的S3 M表示中。我们扩展了这一观点,建议从一个序列的相邻电话的语音信息也组成编码在一个单一的帧,这样的矢量对应于前一个,当前的,和下一个电话叠加在一个单一的帧级表示。我们发现,这种结构有几个属性,包括相对位置之间的正交性,和隐式语音边界的出现。总之,我们的研究结果推进了我们对上下文相关的S3 M表示的理解。摘要:Transformer-based self-supervised speech models (S3Ms) are often described as contextualized, yet what this entails remains unclear. Here, we focus on how a single frame-level S3M representation can encode phones and their surrounding context. Prior work has shown that S3Ms represent phones compositionally; for example, phonological vectors such as voicing, bilabiality, and nasality vectors are superposed in the S3M representation of [m]. We extend this view by proposing that phonological information from a sequence of neighboring phones is also compositionally encoded in a single frame, such that vectors corresponding to previous, current, and next phones are superposed within a single frame-level representation. We show that this structure has several properties, including orthogonality between relative positions, and emergence of implicit phonetic boundaries. Together, our findings advance our understanding of context-dependent S3M representations.


eess.AS音频处理


【1】Bounds on Agreement between Subjective and Objective Measurements
标题:主观测量和客观测量之间一致性的界限
链接:https://arxiv.org/pdf/2603.13204v1

作者:Jaden Pieper,Stephen D. Voran
备注:Currently under review at IEEE Transactions on Multimedia. Submitted 5 November 2025, revised 3 March 2026
摘要:多媒体质量的客观估计通常通过比较估计与主观“真实数据”来判断,最常见的是通过Pearson相关系数(PCC)或均方误差(MSE)。但是主观测试结果包含噪声,因此争取PCC为1.0或MSE为0.0既不现实也不可重复。已经做出了许多努力来承认和适当地适应客观-主观比较中的主观测试噪声,通常导致新的分析框架和品质因数。我们采取不同的方法。通过只做基本的假设,我们得到的PCC和MSE的界限,可以预期的主观测试。 与直觉一致,这些界限是主观投票方差的函数。当主观测试包含投票方差信息时,边界的计算很容易,在这种情况下,我们说得到的边界是“完全数据驱动的”。“我们提供了两个选项,用于在投票方差信息不可用的情况下计算界限。一种选择是使用来自其他主观测试的投票方差信息,这些主观测试确实提供了此类信息,第二种选择是使用主观投票的模型。 因此,我们引入了一个基于二项式的主观投票模型(BinoVotes),它自然会导致一个平均意见得分(MOS)模型,名为BinoMOS,具有多个独特的理想属性。BinoMOS再现了MOS值的离散性及其对每个文件投票数的依赖性。该模型提供了PCC和MSE界限所需的投票方差信息,我们将该模型与18个主观测试的数据进行了比较。建模产生的PCC和MSE界限,同意非常好,直接从数据中发现。这些结果使人们能够为PCC和MSE设定预期,这可能是任何主观测试所能达到的,即使是那些投票方差信息不可用的测试。摘要:Objective estimators of multimedia quality are often judged by comparing estimates with subjective "truth data," most often via Pearson correlation coefficient (PCC) or mean-squared error (MSE). But subjective test results contain noise, so striving for a PCC of 1.0 or an MSE of 0.0 is neither realistic nor repeatable. Numerous efforts have been made to acknowledge and appropriately accommodate subjective test noise in objective-subjective comparisons, typically resulting in new analysis frameworks and figures-of-merit. We take a different approach. By making only basic assumptions, we derive bounds on PCC and MSE that can be expected for a subjective test. Consistent with intuition, these bounds are functions of subjective vote variance. When a subjective test includes vote variance information, the calculation of the bounds is easy, and in this case we say the resulting bounds are "fully data-driven." We provide two options for calculating bounds in cases where vote variance information is not available. One option is to use vote variance information from other subjective tests that do provide such information, and the second option is to use a model for subjective votes. Thus we introduce a binomial-based model for subjective votes (BinoVotes) that naturally leads to a mean opinion score (MOS) model, named BinoMOS, with multiple unique desirable properties. BinoMOS reproduces the discrete nature of MOS values and its dependence on the number of votes per file. This modeling provides vote variance information required by the PCC and MSE bounds and we compare this modeling with data from 18 subjective tests. The modeling yields PCC and MSE bounds that agree very well with those found from the data directly. These results allow one to set expectations for the PCC and MSE that might be achieved for any subjective test, even those where vote variance information is not available.


【2】Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
标题:基于位置相关正交子空间的自监督语音模型语音上下文编码
链接:https://arxiv.org/pdf/2603.12642v1

作者:Kwanghee Choi,Eunjung Yeo,Cheol Jun Cho,David R. Mortensen,David Harwath
备注:Submitted to Interspeech 2026
摘要:基于transformer的自监督语音模型(S3 M)通常被描述为情境化的,但这意味着什么仍然不清楚。在这里,我们专注于如何一个单一的帧级S3 M表示可以编码电话及其周围的上下文。先前的工作已经表明,S3 M表示组成的音素;例如,语音向量,如清音,bilabiality和鼻音向量叠加在[m]的S3 M表示中。我们扩展了这一观点,建议从一个序列的相邻电话的语音信息也组成编码在一个单一的帧,这样的矢量对应于前一个,当前的,和下一个电话叠加在一个单一的帧级表示。我们发现,这种结构有几个属性,包括相对位置之间的正交性,和隐式语音边界的出现。总之,我们的研究结果推进了我们对上下文相关的S3 M表示的理解。摘要:Transformer-based self-supervised speech models (S3Ms) are often described as contextualized, yet what this entails remains unclear. Here, we focus on how a single frame-level S3M representation can encode phones and their surrounding context. Prior work has shown that S3Ms represent phones compositionally; for example, phonological vectors such as voicing, bilabiality, and nasality vectors are superposed in the S3M representation of [m]. We extend this view by proposing that phonological information from a sequence of neighboring phones is also compositionally encoded in a single frame, such that vectors corresponding to previous, current, and next phones are superposed within a single frame-level representation. We show that this structure has several properties, including orthogonality between relative positions, and emergence of implicit phonetic boundaries. Together, our findings advance our understanding of context-dependent S3M representations.


【3】Room Impulse Response Completion Using Signal-Prediction Diffusion Models Conditioned on Simulated Early Reflections
标题:使用以模拟早期反射为条件的信号预测扩散模型完成房间脉冲响应
链接:https://arxiv.org/pdf/2603.12442v1

作者:Zeyu Xu,Andreas Brendel,Albert G. Prinn,Emanuël A. P. Habets

备注:The following article has been submitted for review to Interspeech 2026

摘要:房间脉冲响应(RIR)是音频数据增强、声学信号处理和沉浸式音频渲染的基础。虽然像源法(ISM)等几何模拟器可以有效地生成早期反射,但由于缺少声波效应,它们缺乏测量RIR的真实性。我们提出了一种基于扩散的RIR完成方法,使用ISM模拟的直接路径和早期反射信号预测的条件。与最先进的方法不同,我们的方法对输入的早期反射没有固定的持续时间约束。我们还结合了无分类器的指导,以引导生成目标分布,从使用Treble SDK模拟的物理现实RIR中学习。客观评价表明,所提出的方法优于国家的最先进的基线在早期RIR完成和能量衰减曲线重建。摘要:Room impulse responses (RIRs) are fundamental to audio data augmentation, acoustic signal processing, and immersive audio rendering. While geometric simulators such as the image source method (ISM) can efficiently generate early reflections, they lack the realism of measured RIRs due to missing acoustic wave effects. We propose a diffusion-based RIR completion method using signal-prediction conditioned on ISM-simulated direct-path and early reflections. Unlike state-of-the-art methods, our approach imposes no fixed duration constraint on the input early reflections. We further incorporate classifier-free guidance to steer generation toward a target distribution learned from physically realistic RIRs simulated with the Treble SDK. Objective evaluation demonstrates that the proposed method outperforms a state-of-the-art baseline in early RIR completion and energy decay curve reconstruction.


【4】MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis
标题:MamTra:用于语音合成的混合Mamba-Transformer主干
链接:https://arxiv.org/pdf/2603.12342v1

作者:Tan Dat Nguyen,Sangmin Bae,Joon Son Chung,Ji-Hoon Kim

备注:Submitted to Interspeech 2026

摘要:尽管基于LLM的文本到语音系统的显着的质量,他们的依赖于自回归Transformers导致二次计算复杂度,这严重限制了实际应用。线性时间替代方案,特别是Mamba,提供了一个潜在的补救措施;然而,它们往往牺牲了表达性合成所必需的全局上下文。在本文中,我们提出了MamTra,一个交错的Mamba-Transformer框架,旨在利用Mamba的效率和Transformers的建模能力的优势。我们还引入了新颖的知识转移策略,将来自预训练的Transformer的见解提取到我们的混合架构中,从而绕过了从头开始训练的高昂成本。系统实验确定了最佳的混合配置,并证明MamTra将推理VRAM的使用量减少了34%,而不影响语音保真度-即使只在原始训练数据集的2%上进行训练。音频样本可在https: mamtratts.github.io上获得。摘要:Despite the remarkable quality of LLM-based text-to-speech systems, their reliance on autoregressive Transformers leads to quadratic computational complexity, which severely limits practical applications. Linear-time alternatives, notably Mamba, offer a potential remedy; however, they often sacrifice the global context essential for expressive synthesis. In this paper, we propose MamTra, an interleaved Mamba-Transformer framework designed to leverage the advantages of Mamba's efficiency and Transformers' modeling capability. We also introduce novel knowledge transfer strategies to distill insights from a pretrained Transformer into our hybrid architecture, thereby bypassing the prohibitive costs of training from scratch. Systematic experiments identify the optimal hybrid configuration, and demonstrate that MamTra reduces inference VRAM usage by up to 34% without compromising speech fidelity - even trained on only 2% of the original training dataset. Audio samples are available at https: mamtratts.github.io.


机器翻译由腾讯交互翻译提供,仅供参考