今日论文合集:CS.SD语音与音频 | 共 12 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音合成与声音生成 1 篇
2. 说话人识别、验证与分离 1 篇
3. 音乐信息检索与音乐生成 5 篇
4. 语音翻译与语音语言模型 1 篇
5. 其他/综合语音音频 4 篇
1. 语音合成与声音生成 | 1 篇
1. DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift
DriftTTS:无需蒸馏的少步文本到语音合成,基于分布匹配漂移
AI 总结:提出DriftTTS,一种无需蒸馏和对抗训练的少步文本到语音合成模型,通过分布匹配漂移目标在梅尔域特征空间训练,在LJSpeech上达到与Matcha-TTS相当的性能。
链接:https://arxiv.org/abs/2610.03390
机构:University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校); Worcester Polytechnic Institute(伍斯特理工学院)
作者:Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam
英文摘要:Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a frozen masked-autoencoder encoder pretrained on the same LJSpeech training split. On-policy rollout trains the decoder on its own interme- diate states and supports inference up to the trained roll- out depth. On LJSpeech, DriftTTS at NFE=4 achieves 3.87 dB MCD and 3.7% WER, compared with 3.85 dB and 3.4% for Matcha-TTS. In a fully paired blind listen- ing test, DriftTTS obtains 4.18 MOS, compared with 3.96 for Matcha-TTS and 4.22 for ground truth. These results demonstrate competitive few-step synthesis without a pre- trained generative teacher. Code can be found at https: //github.com/BASHLab/driftTTS.git
2. 说话人识别、验证与分离 | 1 篇
2. GAANet: Global-guided Asymmetric Attention Network for Audio-Visual Speech Separation
GAANet:用于音视频语音分离的全局引导非对称注意力网络
AI 总结:针对音视频语音分离中多尺度特征融合效率低的问题,提出GAANet,通过非对称多尺度融合和全局引导注意力机制,在LRS2和VoxCeleb2上达到最先进性能,且参数和计算量小。
链接:https://arxiv.org/abs/2610.02752
机构:Hefei University of Technology(合肥工业大学)
作者:Zhiyuan Zhang, Jingyuan Xu, Yiming Tang, Liu Liu, Dan Guo
英文摘要:Multi-scale design is crucial for efficient audio-visual speech separation, yet effectively modeling multi-scale information for audio-visual feature fusion remains challenging. We argue that the limited capacity of existing approaches primarily arises from: 1) treating features from different modalities in the same manner, and 2) overlooking the role of global features. To address these issues, we propose a Global-guided Asymmetric Attention Network (GAANet). Our model introduces two core innovations: first, an asymmetric multi-scale fusion framework that allows audio and visual streams to extract and interact with features at their respective optimal temporal resolutions, removing the need for symmetric temporal downsampling; second, a global-guided attention mechanism that compresses each modality into a compact global token with a temporal dimension of one, which then provides high-level semantic cues to guide both intra- and inter-modal fusion across scales. Experiments on LRS2 and VoxCeleb2 demonstrate that GAANet achieves state-of-the-art performance, reaching 16.5 dB SI-SNRi on LRS2 and 14.0 dB on VoxCeleb2, while maintaining a lightweight computational profile with only 3.3M parameters and 19.8G MACs. These results highlight the strong potential of asymmetric temporal modeling and global guidance for efficient and robust multimodal fusion. The source code is publicly accessible at this https URL
3. 音乐信息检索与音乐生成 | 5 篇
3. How Far Back Should a Transformer Look? Repetition and Copying in Music Sequence Models
Transformer 应该回溯多远?音乐序列模型中的重复与复制
AI 总结:本文研究符号音乐自回归模型中上下文长度对预测性能的影响,发现长上下文收益主要源于精确复制,并提出了上下文长度测量的方法论检查。
链接:https://arxiv.org/abs/2610.02837
机构:University of Alberta(阿尔伯塔大学)
作者:Amir Fathi
英文摘要:We investigate how predictive performance depends on the maximum context available to an autoregressive model of symbolic music, and what information long-context models exploit. A small causal Transformer is trained separately at each context length T in {6, 18, 48, 96, 192, 336} and evaluated on the same target positions, both over raw tokens and over non-overlapping binary latent codes. On Nottingham folk tunes, the token predictor's test NLL decreases by 71% (0.986 bits per token) between 6 and 336 tokens, and by 71% on the O'Neill's tunes long enough for the same sweep (129 of 302 test tunes). Most of the long-context gain is explained by exact copying: the improvement appears when an earlier occurrence of the target's 16-token history enters the available context; overwriting that occurrence removes the gain, whereas equally large unrelated corruption does not; and a simple copy baseline recovers 94% of the reduction. A trained 336-token model likewise loses most of this gain when its history is restricted to recent tokens at test time. Because many repeats arise from written repeat signs expanded in the score, this result pertains specifically to these rendered score representations. By contrast, on MAESTRO performances and MusicNet scores, where exact repeats are substantially less frequent, the reduction is smaller (9% and 18%) and is largely attained by 96 tokens. Finally, an audit of an earlier draft that reported saturation at 16 tokens identified split leakage, overlapping latent receptive fields, an averaging predictor, and a per-file tempo grid; we present these as methodological checks for context-length measurements.
4. Learning Jazz Pianist Style with Cross-Attention Conditioning
学习爵士钢琴家风格:基于交叉注意力条件生成
AI 总结:本研究利用预训练符号音乐转换器编码钢琴家身份,通过交叉注意力条件生成特定风格音乐,并验证了风格捕捉的有效性及特征定位能力。
链接:https://arxiv.org/abs/2610.02918
作者:Drew Edwards, Akira Maezawa, Simon Dixon
英文摘要:Jazz pianists develop distinctive traits that experienced listeners can often identify within seconds, yet the features underlying this recognition resist formal description. We study jazz pianist style through the lens of a pretrained symbolic music transformer, showing that its learned representations already encode pianist identity well enough for highly accurate classification across two benchmarks. We then augment the transformer with cross-attention over learned pianist identity embeddings, enabling it to generate music conditioned on a specific artist's style. Two evaluation protocols confirm that the generator captures meaningful stylistic structure: a sliding-window classifier consistently attributes conditioned continuations to the correct artist, far above unconditioned baselines; and a classifier trained entirely on synthetic generations identifies real pianists across 12 classes with 87% chunk-level and 95% song-level accuracy. Finally, we repurpose the classifier to locate the most characteristic moments within a performance, surfacing the specific musical gestures that distinguish each pianist's voice.
5. PEACE: Joint Embeddings of DSP Effects Code and Audio
PEACE:DSP效果代码与音频的联合嵌入
AI 总结:PEACE首次联合嵌入音频效果代码与音频,利用Faust代码编码器,在音频到代码检索中恢复链拓扑,优于预训练模型,奠定DSP代码音乐检索基础。
链接:https://arxiv.org/abs/2610.03405
作者:David Braun, Adam Finkelstein
英文摘要:This paper introduces PEACE, the first joint embedding of audio effect code and output audio. Building on SLAP's multimodal objective, we pair an AFx-Rep audio encoder with two code encoders for Faust, a functional language for audio signal processing. First, we evaluate a fine-tuned T5 transformer over Faust source code. Second, we evaluate a message-passing graph neural network over an intermediate representation of the Faust compiler, capturing both topology and UI parameters. We evaluate on audio-to-code retrieval, where masking UI parameters at inference yields embeddings that encode effect chain topology alone. When parameters are visible, the two code encoders tie on retrieval of mixed-length chains but have tradeoffs on single-effect galleries. With parameters fully masked, PEACE recovers ordered chain topology far above chance without the limitations of supervised methods. PEACE outperforms pretrained models on an out-of-distribution reverb retrieval benchmark and can improve frozen audio-only representations. Its dual understanding of topology and parameters lays the groundwork for music information retrieval systems that search, generate, and condition on DSP code.
6. LayerIt: Towards a Framework for Time-Aligned, Composable Music Visualizations
LayerIt:面向时间对齐、可组合音乐可视化的框架
AI 总结:提出LayerIt开源Python库,在共享表演时间轴上组合乐谱与信号表示,生成保留MEI结构的SVG,并通过节拍跟踪分析验证其可同时检查跟踪输出、信号和乐谱以定位错误。
链接:https://arxiv.org/abs/2610.03428
作者:Fernando Azeredo, António Sá Pinto
英文摘要:Music information retrieval often relates signal-derived, algorithmic, and symbolic information across different coordinate systems. Existing visualizations typically leave notation separate from physical time, while composites that combine them are assembled by hand. We present LayerIt, an open-source Python library for composing independent representations with notation on a shared performance-time axis. Given a score warped into performance time and a note-level alignment, LayerIt emits a single SVG that preserves the score's MEI structure and keeps added components identifiable and restylable. We demonstrate this through a challenging beat-tracking analysis, in which tracker output, signal representations, and notation can be inspected together to locate metrical disagreement and other errors against both notated structure and performed time.
7. Rubric-Based Optimization for Text-to-Music Generation
基于评分标准的文本到音乐生成优化
AI 总结:本研究提出利用音频语言模型的结构化评分标准作为奖励信号,通过DPO和DiffusionNFT优化文本到音乐生成,实验表明其提升整体感知质量,而专门客观奖励更适合精确属性优化。
链接:https://arxiv.org/abs/2610.03589
机构:University of Washington(华盛顿大学); Paul G. Allen School of Computer Science & Engineering, University of Washington(华盛顿大学保罗·G·艾伦计算机科学与工程学院); Allen Institute for AI(艾伦人工智能研究所)
作者:Ping Wang, Guang Yang, Shao-Rong Su, Junkai Wu, Pang Wei Koh, Noah A. Smith
英文摘要:Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as training signals for both autoregressive and diffusion-based music generators. An ALM scores each generated clip against the rubric; we rank candidates generated for the same text prompt by their scores and convert these rankings into preference pairs for DPO on both MusicGen-small and ACE-Step~v1, and additionally use the rubric scores directly as scalar rewards for DiffusionNFT on ACE-Step~v1. On MusicCaps, rubric-based optimization improves CLAP, SongEval, and Audiobox-Aesthetics simultaneously, with the strongest gains obtained by DiffusionNFT on ACE-Step. By contrast, on MusicGen-small, building preferences from any one of these automatic evaluators produces clear cross-metric trade-offs: the targeted evaluator improves while other independent evaluators deteriorate. We further study tempo, key, and instrumentation, where precise objective rewards are available. Directly optimizing these specialized rewards reliably improves the target attributes, whereas ALM rubrics provide only partial transfer for tempo and instrumentation and no measurable improvement for key. Together, these results suggest a practical division of labor: ALM rubrics are effective for broad perceptual qualities that are difficult to formalize, while specialized objective rewards remain preferable when reliable measurements are available.
4. 语音翻译与语音语言模型 | 1 篇
8. ParaGeo: Decomposing Paralinguistic Variation into a Shared Latent Geometry
ParaGeo:将副语言变异分解为共享潜在几何结构
AI 总结:ParaGeo通过匹配内容分解,在冻结语音模型中建立共享潜在几何,实现副语言属性与内容的解耦,并验证了跨内容泛化能力。
链接:https://arxiv.org/abs/2610.03125
机构:University of Cambridge(剑桥大学); University of Oxford(牛津大学); University of Washington Bothell(华盛顿大学博塞尔分校)
作者:Yuhan Liu, Yuxuan Ou, Ruoxi Su, Mohamed Ahmed Zaki, Yunbo Long
英文摘要:Speech delivery varies with both the requested paralinguistic attribute and the linguistic content. We introduce ParaGeo, a matched-content decomposition of paralinguistic variation in a frozen speech language model. Synthesized audio tokens are replayed with a fixed listening prompt; pooled key/value (K/V) representations are centered and projected into a shared low-dimensional space. Our GLM-4-Voice probe spans 80 requested controls from 12 benchmark families across eight sentences. With a globally fitted calibration basis, content-held-out centroid accuracy using this basis is 9.49% versus a 1.25% permutation baseline; same-label cross-content cosine similarity is 0.285 versus 0.017, and both conditional permutation tests yield p = 0.001. A separate ten-scenario, six-style probe reveals reproducible contrast directions across scenarios. Static, additive, and temporal interventions produce attribute-, layer-, and schedule-dependent response profiles. These results provide a shared coordinate representation for measuring paralinguistic structure and an empirical starting point for latent speech control. Code is available at this https URL.
5. 其他/综合语音音频 | 4 篇
9. What Actually Makes Correlation-Based SSL Distillation Noise-Robust? A Mechanistic Correction
究竟是什么让基于相关性的SSL蒸馏具有噪声鲁棒性?一种机制性修正
AI 总结:本研究通过机制性分析揭示,基于相关性的SSL蒸馏中互相关项对角线才是噪声鲁棒性的关键,而自相关项主要改善下游任务,并据此提出蒸馏设计规则。
链接:https://arxiv.org/abs/2610.02823
机构:Nanyang Technological University(南洋理工大学); Institute for Infocomm Research (I2R)(资讯通信研究院)
作者:Fabian Ritter-Gutierrez, Nancy F. Chen, Eng Siong Chng
英文摘要:Self-supervised learning (SSL) speech models are accurate but large. Knowledge distillation compresses them, but the student loses the teacher's noise robustness. Correlation-based distillation addresses this with two terms: cross-correlation aligning student and teacher's representations, and self-correlation decorrelating the student's features. Both the original method and De'HuBERT credited the self-correlation term without isolating it. We show the opposite. On LibriSpeech-100 distillation with held-out CHiME-3 noise at 10\,dB, an unbiased full-dimensional probe, a Pearson-variance decomposition, a same-noise causal control, and a per-dimension analysis identify the cross-correlation diagonal as the mechanism that encourages noise invariance, lowering noise-classification accuracy from 76.98\% to 55.02\%, whereas adding the self-correlation term back leaves it at 55.55\%. The self-correlation term instead reorganises the feature space and improves accuracy on the clean downstream tasks, but removes essentially no noise. Across nine speech and music tasks this yields a concrete design rule for distillation: weight the cross-correlation term for noise robustness, tune the self-correlation term for better downstream performance, and sample teacher and student noise independently.
10. Correlation-Based Distillation Yields More Mergeable Speech-Music Encoders
基于相关性的蒸馏产生更易合并的语音-音乐编码器
AI 总结:本研究证明将蒸馏损失替换为基于相关性的损失,可显著提升语音-音乐编码器在两种合并算法下的合并效果,表明收益源于表示质量。
链接:https://arxiv.org/abs/2610.02836
机构:Nanyang Technological University(南洋理工大学)
作者:Fabian Ritter-Gutierrez
英文摘要:A single compact encoder for both speech and music removes the need to maintain a separate model per domain. A practical recipe distils a speech teacher and a music teacher into two small students, then merges them into one model. Two merging algorithms exist: interpolating the student weights from a shared initialisation, or permuting their channels into alignment before averaging. Prior work fixed the distillation loss in both, leaving open whether the loss itself affects how well the students merge. We show that it does. Holding architecture and evaluation fixed, we replace the DistilHuBERT loss \Lkd{} with a correlation-based loss \Lcl{}. Under weight interpolation, \Lcl{} students stay closer to their own better endpoint in $17$ of $18$ comparisons and lead on the speech tasks at every interpolated weight. Under activation permutation, they need fewer channels rearranged at three layers, their matched channels correlate more strongly, and they again lead on speech. The two merging algorithms share no mechanism, yet both improve under the same loss substitution. This suggests that the gain lies in the representation \Lcl{} produces rather than in either algorithm.
11. Harmonic Eigenspace: A Web-based Application for Navigating and Composing Microtonal Harmony
谐波特征空间:一个用于导航与创作微音程和声的网页应用
AI 总结:本文介绍一个基于谐波特征空间的网页应用,用于导航和创作微音程和声,通过3D可视化和模态工作室支持53-TET,并经31人聆听研究验证其可用性。
链接:https://arxiv.org/abs/2610.03398
机构:Pompeu Fabra University(庞培法布拉大学); Univ. Lille(里尔大学); CNRS(法国国家科学研究中心); Centrale Lille(里尔中央理工学院); UMR 9189 CRIStAL(UMR 9189 CRIStAL 实验室)
作者:David Dalmazzoa, Ken Déguernel
英文摘要:This paper presents a web-based application for navigating and composing microtonal harmony, built on the Harmonic Eigenspace, a four-dimensional psychoacoustically grounded space in which tetrad chord types are located by their spectral dissonance profiles, computed with Sethares's roughness/dissonance model. The coordinate system is transposition-invariant: a coordinate triple ({\alpha}, \b{eta}, {\gamma}) locates the three upper notes in relation to the root, so a chord quality corresponds to a direction in the space, the invariant ray along which transposition acts, while the root frequency sets the scale. The dissonance field over these coordinates can be computed at any register; the locations of its local minima are register-invariant, as they arise from partial-coincidence ratio conditions. The dissonance volume contains 100 local minima that align with just-intonation intervals and act as landmarks, organising the space into basins around the most consonant tetrads. We embed tetrads from three tonal equal temperaments as discrete lattices within this continuous volume. The application presents this space through two components: the Harmonic Eigenspace as a navigable 3D visualisation of the dissonance volume in which all nodes are playable, and a Modal Studio that extends modal interchange logic to the ten-gradation interval vocabulary of 53-TET. The application also functions as a MIDI controller with MIDI Polyphonic Expression support, usable in any digital audio workstation that supports this format. A listening study with 31 participants used both scenes of the application: listeners first rated isolated 53-TET chords alongside chords familiar from Western practice, such as the maj7 and the m7; they then rated chord progressions composed in the Modal Studio, measuring their acceptance or rejection of microtonal progressions heard for the first time.
12. Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles
重新审视人声合奏多音高估计中的输入时频表示
AI 总结:本文重新审视人声合奏多音高估计中的输入时频表示,发现线性STFT在性能和效率上优于HCQT,且更精细的频率分辨率并非必要,较短分析窗口更有效。
链接:https://arxiv.org/abs/2610.03656
机构:Yonsei University(延世大学); University of Michigan(密歇根大学)
作者:Junyoung Koh, Hao-Wen Dong
英文摘要:Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with
a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially reducing feature-extraction cost. Further analysis shows that a longer analysis window or broader spectral coverage provides no additional improvement, while restricting the input to the predicted pitch range reduces the advantage of the linear STFT. These results suggest that finer frequency resolution does not necessarily improve vocal-ensemble MPE, and that shorter analysis windows can be more effective for time-varying vocal pitches.
