今日论文合集:cs.SD语音11篇,eess.AS音频处理8篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
标题:相互强迫:双模式自我进化,实现快速自回归音频视频字符生成
链接:https://arxiv.org/abs/2604.25819
作者:Yupeng Zhou,Lianghua Huang,Zhifan Wu,Jiabao Wang,Yupeng Shi,Biao Jiang,Daquan Zhou,Yu Liu,Ming-Ming Cheng,Qibin Hou
摘要:在这项工作中,我们提出了相互强迫,一个框架,用于快速自回归音频视频生成与长期的音频视频同步。我们的方法解决了两个关键的挑战:联合音频视频建模和快速自回归生成。为了简化联合音视频优化,我们采用了两阶段训练策略:首先训练单模态生成器,然后将它们耦合到统一的音视频模型中,以便对配对数据进行联合训练。对于流生成,我们询问是否可以直接训练本地快速因果音频-视频模型,而不是遵循现有的流蒸馏管道,该管道通常首先训练双向模型,然后通过多个蒸馏阶段将其转换为因果生成器。我们的答案是相互强迫,它直接建立在原生自回归模型上,并在单个权重共享模型中集成了少步和多步生成,从而实现了自蒸馏和改进的训练-推理一致性。多步模式通过自升华改进了少步模式,而少步模式在训练过程中生成历史上下文以提高训练-推理一致性;由于两种模式共享参数,这两种效果在单个模型中相互加强。与Self-Forcing等现有方法相比,Mutual Forcing消除了对额外双向教师模型的需求,支持更灵活的训练序列长度,减少了训练开销,并允许模型直接从真实配对数据而不是固定教师进行改进。实验表明,互迫匹配或超越强基线,需要大约50个采样步骤,而只需要4到8个步骤,在效率和质量上都表现出显着的优势。该项目的网页可在https://mutualforcing.github.io上查阅。
摘要:In this work, we propose Mutual Forcing, a framework for fast autoregressive audio-video generation with long-horizon audio-video synchronization. Our approach addresses two key challenges: joint audio-video modeling and fast autoregressive generation. To ease joint audio-video optimization, we adopt a two-stage training strategy: we first train uni-modal generators and then couple them into a unified audio-video model for joint training on paired data. For streaming generation, we ask whether a native fast causal audio-video model can be trained directly, instead of following existing streaming distillation pipelines that typically train a bidirectional model first and then convert it into a causal generator through multiple distillation stages. Our answer is Mutual Forcing, which builds directly on native autoregressive model and integrates few-step and multi-step generation within a single weight-shared model, enabling self-distillation and improved training-inference consistency. The multi-step mode improves the few-step mode via self-distillation, while the few-step mode generates historical context during training to improve training-inference consistency; because the two modes share parameters, these two effects reinforce each other within a single model. Compared with prior approaches such as Self-Forcing, Mutual Forcing removes the need for an additional bidirectional teacher model, supports more flexible training sequence lengths, reduces training overhead, and allows the model to improve directly from real paired data rather than a fixed teacher. Experiments show that Mutual Forcing matches or surpasses strong baselines that require around 50 sampling steps while using only 4 to 8 steps, demonstrating substantial advantages in both efficiency and quality. The project page is available at https://mutualforcing.github.io.


【2】WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition

标题:WhisperPipe:用于实时自动语音识别的资源高效流媒体架构
链接:https://arxiv.org/abs/2604.25611
作者:Erfan Ramezani,Mohammad Mahdi Giahi,Mohammad Erfan Zarabadipour,Amir Reza Yosefian,Hamid Ghadiri
备注:36 pages, 14 figures. Open-source implementation available at PyPI
摘要:实时自动语音识别(ASR)系统面临着转录准确性和计算效率之间的根本权衡,特别是在部署像Whisper这样的大规模Transformer模型时。现有的流方法要么通过积极的分块牺牲准确性,要么通过无限的上下文积累产生高昂的内存成本。我们提出了WhisperPipe,一种新型的流媒体架构,通过三个关键创新实现了有限的内存消耗,同时保持转录质量:将Silero VAD与基于能量的过滤相结合的混合语音活动检测(VAD)管道,以减少34%的错误激活,具有重叠上下文窗口的动态缓冲机制,可防止片段边界处的信息丢失,以及基于语音特性平衡延迟和准确性的自适应处理策略。通过对2.5小时的各种音频数据进行评估,WhisperPipe的端到端延迟中值为89 ms(第90百分位数:142 ms),与基线Whisper实现相比,峰值GPU内存消耗减少48%,平均GPU利用率降低80.9%。系统在扩展会话中保持稳定的内存使用,在150分钟的连续操作中增长率为零。与相关工作的比较分析表明,WhisperPipe实现了具有竞争力的准确性(WER在离线Whisper的2%以内),同时比现有流媒体解决方案的延迟低3- 5倍。该架构的模块化设计支持在资源受限的环境中进行部署,从边缘设备到云基础设施。我们的研究结果表明,仔细的架构设计可以调和生产ASR系统中实时响应和模型复杂性的竞争需求。
摘要:Real-time automatic speech recognition (ASR) systems face a fundamental trade-off between transcription accuracy and computational efficiency, particularly when deploying large-scale transformer models like Whisper. Existing streaming approaches either sacrifice accuracy through aggressive chunking or incur prohibitive memory costs through unbounded context accumulation. We present WhisperPipe, a novel streaming architecture that achieves bounded memory consumption while maintaining transcription quality through three key innovations a hybrid Voice Activity Detection (VAD) pipeline combining Silero VAD with energy-based filtering to reduce false activations by 34%, a dynamic buffering mechanism with overlapping context windows that prevents information loss at segment boundaries, and an adaptive processing strategy that balances latency and accuracy based on speech characteristics. Evaluated on 2.5 hours of diverse audio data, WhisperPipe demonstrates a median end-to-end latency of 89ms (90th percentile: 142ms) while consuming 48% less peak GPU memory and 80.9% lower average GPU utilization compared to baseline Whisper implementations. The system maintains stable memory usage over extended sessions, with zero growth rate across 150-minute continuous operation. Comparative analysis against related work shows that WhisperPipe achieves competitive accuracy (WER within 2% of offline Whisper) while operating at 3-5x lower latency than existing streaming solutions. The architecture's modular design enables deployment across resource-constrained environments, from edge devices to cloud infrastructure. Our results demonstrate that careful architectural design can reconcile the competing demands of real-time responsiveness and model sophistication in production ASR systems.


【3】SymphonyGen: 3D Hierarchical Orchestral Generation with Controllable Harmony Skeleton

标题:SymphonyGen:具有可控和声骨架的3D分层中音生成
链接:https://arxiv.org/abs/2604.25498
作者:Xuzheng He,Nan Nan,Zhilin Wang,Ziyue Kang,Zhuoru Mo,Ao Li,Yu Pan,Xiaobing Li,Feng Yu,Xiaohong Guan
备注:8 pages, 4 figures
摘要:生成交响乐需要同时管理高层次的结构形式和密集的多轨道配器。现有的符号模型经常与“复杂性控制不平衡”作斗争,其中扩展瓶颈限制了长期的粒度可操纵性。我们提出SymphonyGen,一个3D层次框架,为当代电影编排。SymphonyGen采用级联解码器架构,分解Bar,Track和Event轴,提高了传统1D或2D模型的计算效率和可扩展性。我们通过节拍量化的多语音和声骨架引入“短分数”调节,在保持纹理多样性的同时实现轮廓控制。该模型进一步细化使用组相对策略优化(GRPO)与跨模态的音频感知奖励,调整符号输出与现代的声学期望。此外,我们实现了一个不和谐的采样算法,以抑制意外的音调冲突在推理。客观评估表明,强化学习和不和谐厌恶采样有效地提高谐波清洁,同时保持旋律的表达。主观评估表明,SymphonyGen在音乐性和对管弦乐生成的偏好方面优于基线。演示页面:https://symphonygen.github.io/
摘要:Generating symphonic music requires simultaneously managing high-level structural form and dense, multi-track orchestration. Existing symbolic models often struggle with a "complexity-control imbalance", in which scaling bottlenecks limit long-term granular steerability. We present SymphonyGen, a 3D hierarchical framework for contemporary cinematic orchestration. SymphonyGen employs a cascading decoder architecture that decomposes the Bar, Track, and Event axes, improving computational efficiency and scalability over conventional 1D or 2D models. We introduce "short-score" conditioning via a beat-quantized multi-voice harmony skeleton, enabling outline control while preserving textural diversity. The model is further refined using Group Relative Policy Optimization (GRPO) with a cross-modal audio-perceptual reward, aligning symbolic output with modern acoustic expectations. Additionally, we implement a dissonance-averse sampling algorithm to suppress unintended tonal clashes during inference. Objective evaluations show that both reinforcement learning and dissonance-averse sampling effectively enhance harmonic cleanliness while maintaining melodic expression. Subjective evaluations demonstrate that SymphonyGen outperforms baselines in musicality and preference for orchestral music generation. Demo page: https://symphonygen.github.io/


【4】PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech

标题:PSP:印度文本到语音的可解释按维口音基准
链接:https://arxiv.org/abs/2604.25476
作者:Venkata Pushpak Teja Menta
备注:8 pages, 7 tables. Companion paper to Praxy Voice (arXiv:submission id - 7506231). Code: https://github.com/praxelhq/psp-eval; Centroids: https://huggingface.co/datasets/Praxel/psp-native-centroids
摘要:标准的文本到语音(TTS)评估测量可懂度(WER,CER)和整体自然度(MOS,UTMOS),但不量化口音。一个合成器可能在所有四个方面都得分很高,但在目标语言中的音素特征上听起来不像母语。对于印度语,这些特征包括卷舌音发音、送气、元音长度和泰米尔卷舌音近似音(字母zha)。我们提出PSP,音素替代配置文件,一个可解释的,每音系维度口音基准为印度文TTS。PSP将重音分解为六个互补的维度:卷舌音塌陷率(RR)、送气保真度(AF)、元音长度保真度(LF)、泰米尔语保真度(ZF)、弗雷歇音频距离(FAD)和韵律特征发散(PSD)。前四个是通过强制对齐加上母语者质心声学探头在Wav 2 Vec 2-XLS-R层-9嵌入测量的;后两个是语料库水平的分布距离。在这个v1中,我们在印地语,泰卢固语和泰米尔语的试点集上对四个商业和开源系统(ElevenLabs v3,Cartesia Sonic-3,Sarvam Bulbul,Indic Parler-TTS)进行了基准测试,第五个系统(Praxy Voice)包括所有三种语言,以及泰卢固语的R5->R6案例研究。三个发现:(i)卷舌音塌陷随着语音难度的增加而单调增长印地语<泰卢固语<泰米尔语(~ 1%,~ 40%,~68%);(ii)PSP排序与WER排序不同-商业WER领导者并不均匀地导致卷舌音或韵律保真度;(iii)没有一个系统在所有六个维度上都是帕累托最优的。我们发布了原生参考质心(每种语言500个片段),FAD的1000个片段嵌入,PSD的500个片段韵律特征矩阵,每种语言300个话语黄金集,MIT下的评分代码和CC-BY下的质心。正式的MOS相关延迟到v2; v1报告五个内部一致性信号加上一个本地音频健全性检查。
摘要:Standard text-to-speech (TTS) evaluation measures intelligibility (WER, CER) and overall naturalness (MOS, UTMOS) but does not quantify accent. A synthesiser may score well on all four yet sound non-native on features that are phonemic in the target language. For Indic languages, these features include retroflex articulation, aspiration, vowel length, and the Tamil retroflex approximant (letter zha). We present PSP, the Phoneme Substitution Profile, an interpretable, per-phonological-dimension accent benchmark for Indic TTS. PSP decomposes accent into six complementary dimensions: retroflex collapse rate (RR), aspiration fidelity (AF), vowel-length fidelity (LF), Tamil-zha fidelity (ZF), Frechet Audio Distance (FAD), and prosodic signature divergence (PSD). The first four are measured via forced alignment plus native-speaker-centroid acoustic probes over Wav2Vec2-XLS-R layer-9 embeddings; the latter two are corpus-level distributional distances. In this v1 we benchmark four commercial and open-source systems (ElevenLabs v3, Cartesia Sonic-3, Sarvam Bulbul, Indic Parler-TTS) on Hindi, Telugu, and Tamil pilot sets, with a fifth system (Praxy Voice) included on all three languages, plus an R5->R6 case study on Telugu. Three findings: (i) retroflex collapse grows monotonically with phonological difficulty Hindi < Telugu < Tamil (~1%, ~40%, ~68%); (ii) PSP ordering diverges from WER ordering -- commercial WER-leaders do not uniformly lead on retroflex or prosodic fidelity; (iii) no single system is Pareto-optimal across all six dimensions. We release native reference centroids (500 clips per language), 1000-clip embeddings for FAD, 500-clip prosodic feature matrices for PSD, 300-utterance golden sets per language, scoring code under MIT, and centroids under CC-BY. Formal MOS-correlation is deferred to v2; v1 reports five internal-consistency signals plus a native-audio sanity check.


【5】Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

标题:Prachy Voice:语音提示恢复+ BUPS,以零商业训练数据成本从冷冻非印度基地获得商业级印度TTC
链接:https://arxiv.org/abs/2604.25441
作者:Venkata Pushpak Teja Menta
备注:9 pages, 6 figures, 6 tables. Companion paper to PSP benchmark. Code: https://github.com/praxelhq/praxy ; Model: https://huggingface.co/Praxel/praxy-voice-r6 ; Demo: https://huggingface.co/spaces/Praxel/praxy-voice-demo
摘要:商业TTS系统产生接近原生的印度语音频,但最好的开源基础(Chatterbox,Indic Parler-TTS,IndicF 5)在测量的语音维度上落后于它们,而最广泛采用的多语言基础(Chatterbox,23种语言)甚至没有标记泰卢固语或泰米尔语。我们问:在不训练新的声学解码器和没有任何商业TTS训练数据的情况下,将这样的非印度语本地基础带到泰卢固语、泰米尔语和印地语的商业级输出的最小干预是什么?我们结合三个部分:(1)BUPS,一个Brahmic统一音素空间,它确定性地将七种印度文字罗马化为ISO-15919,因此Chatterbox的拉丁语标记器可以处理它们;(2)仅在文本标记预测器上的LoRA适配器(Chatterbox的t3),在大约1,220小时的许可印度语音频上训练,使用Hindi代理语言_id;(3)语音提示恢复配方--8- 11秒的相同语言参考剪辑加上三个采样覆盖(夸张0.7,温度0.6,min_p 0.1;“配置B”)--恢复商业级声学输出,无需声学解码器训练。在印地语中,LoRA回归准确性,我们使用vanilla Chatterbox + Config B,给出两个分支部署。在10个发音试点集与同伴PSP基准评估,Praxy语音匹配或略有领先的商业基线:26.7%卷舌崩溃泰卢固语(与Sarvam Bulbul 33.3%),71%泰米尔扎崩溃(与商业三人组的86%),0.025 LLM-WER印地语(与Cartesia Sonic-3并列)。对于内部代码混合,我们添加第三分支(IndicF 5+本机脚本音译),其跨Hi/Te/Ta将代码混合LLM-WER从0.80-0.85下降到0.14-0.27。我们发布了R6 LoRA权重(Apache-2.0),推理代码和路由器(MIT)以及Gradio演示。
摘要:Commercial TTS systems produce near-native Indic audio, but the best open-source bases (Chatterbox, Indic Parler-TTS, IndicF5) trail them on measured phonological dimensions, and the most widely adopted multilingual base (Chatterbox, 23 languages) does not even tokenise Telugu or Tamil. We ask: what is the minimum intervention that brings such a non-Indic-native base to commercial-class output on Telugu, Tamil, and Hindi, without training a new acoustic decoder and without any commercial TTS training data? We combine three pieces: (1) BUPS, a Brahmic Unified Phoneme Space that deterministically romanises seven Indic scripts to ISO-15919 so Chatterbox's Latin tokeniser can process them; (2) a LoRA adapter on only the text-token predictor (Chatterbox's t3), trained on ~1,220h of licensed Indic audio with a Hindi-proxy language_id; (3) a voice-prompt recovery recipe -- an 8-11s same-language reference clip plus three sampling overrides (exaggeration 0.7, temperature 0.6, min_p 0.1; "Config B") -- that recovers commercial-class acoustic output with no acoustic-decoder training. On Hindi, the LoRA regresses accuracy and we instead use vanilla Chatterbox + Config B, giving a two-branch deployment. Evaluated on 10-utterance pilot sets with the companion PSP benchmark, Praxy Voice matches or slightly leads commercial baselines: 26.7% retroflex collapse on Telugu (vs Sarvam Bulbul 33.3%), 71% Tamil-zha collapse (vs commercial trio's 86%), 0.025 LLM-WER on Hindi (tied with Cartesia Sonic-3). For intra-sentential code-mix we add a third branch (IndicF5 + native-script transliteration) that drops code-mix LLM-WER from 0.80-0.85 to 0.14-0.27 across Hi/Te/Ta. We release R6 LoRA weights (Apache-2.0), inference code and router (MIT), and a Gradio demo.


【6】ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

标题:ML-SAN:用于对话中情感识别的多层扬声器自适应网络
链接:https://arxiv.org/abs/2604.25383
作者:Kexue Wang,Yinfeng Yu,Liejun Wang
备注:Main paper (12 pages). Accepted for publication by International Conference on Intelligent Computing 2026
摘要:要与机器建立同理心,必须充分了解人类的情绪变化。然而,多模态情绪识别的研究往往忽略了一个问题:个体的表达特质差异很大,这意味着不同的人可能会表达不同的情绪。在日常生活中,我们可以看到这一点。在与不同的人交流时,有些人通过他们的面部表情和语言表达“快乐”,而另一些人可能会隐藏他们的快乐或通过他们的行动表达。两者都是“快乐”的表达,但这种情感表达的差异对于机器来说仍然太难区分。目前的情感识别仍然停留在一个“静态”的水平,使用一个单一的识别模型来识别所有的情感风格。这种“简化”往往会影响识别结果,特别是在多话轮对话中。为了解决这个问题,本文介绍了一种新的多级说话人自适应网络(ML-SAN),其中,具体地说,有效地解决了说话人身份信息混淆的挑战。ML-SAN在识别后并不简单地分配说话者的ID;相反,它采用了三个阶段的自适应过程:首先,输入级校准使用音频级线性调制(FiLM)将原始音频和视觉特征调整到与说话者无关的中性空间。然后,交互级别门控重新调整每个模态的信任级别(例如,语音或面部特征)。最后,输出级正则化保持了潜在空间中说话人特征的一致性。在MELD和IEMOCAP数据集上的测试表明,我们的模型(ML-SAN)取得了更好的结果,在处理具有挑战性的尾部情感类别方面表现出色,并且更好地解决了现实世界场景中说话者的多样性。
摘要:To establish empathy with machines, it is essential to fully understand human emotional changes. However, research in multimodal emotion recognition often overlooks one problem: individual expressive traits vary significantly, which means that different people may express emotions differently. In our daily lives, we can see this. When communicating with different people, some express "happiness" through their facial expressions and words, while others may hide their happiness or express it through their actions. Both are expressions of 'happiness,' but such differences in emotional expression are still too difficult for machines to distinguish. Current emotion recognition remains at a 'static' level, using a single recognition model to identify all emotional styles. This "simplification" often affects the recognition results, especially in multi-turn dialogues. To address this problem, this paper introduces a novel Multi-Level Speaker Adaptive Network (ML-SAN), which, specifically, effectively addresses the challenge of speaker identity information confusion. ML-SAN does not simply assign a speaker's ID after recognition; instead, it employs a three-stage adaptive process: First, Input-level Calibration uses Feature-Level Linear Modulation (FiLM) to adjust the raw audio and visual features into a neutral space unrelated to the speaker. Then, Interaction-level Gating re-adjusts the trust level for each modality (e.g., voice or facial features) based on the speaker's identity information. Finally, Output-level Regularization maintains the consistency of speaker features in the latent space. Tests on the MELD and IEMOCAP datasets show that our model (ML-SAN) achieves better results, performs exceptionally well in handling challenging tail sentiment categories, and better addresses the diversity of speakers in real-world scenarios.


【7】Huí Sù: Co-constructing a Dual Feedback Apparatus

标题:Huí Scrum:共同构建双反馈装置
链接:https://arxiv.org/abs/2604.25207
作者:Yichen Wang,Charles Patrick Martin
备注:Accepted for publication at the International Conference on New Interfaces for Musical Expression (NIME) 2026 (music track)
摘要:这场表演呈现了两种智能乐器之间的二重奏,Sunday(追溯;逆流而上)和Sundier(在agentic clavier上演奏),以及他们的人类表演者,通过反馈回路连接。这两个系统都不是将人工智能视为对输入做出可预测响应的工具,而是递归地运行,过去的行为不断影响未来的行为。声音通过潜在的表征在音频空间中运行。它的执行者使用Make Noise 0系列合成器和P2P控制器与基于RAVE模型的神经反馈合成系统一起工作,模型的内部结构中嵌入了潜在的反馈回路。这使得乐器能够记住并重复使用自己的内部状态,通过其最近的声音历史影响正在进行的声音生成。控制器在控制空间中起作用。它的表演者使用Roland S-1合成器和Keith McMillen QuNeo触摸板与系统进行交互,控制手势被路由到一个循环神经网络,反馈到合成过程中。通过该反馈回路,系统主动地塑造控制信号随时间的演变。对比音频和控制领域的反馈,表演探索了人类和智能音乐系统之间的共享代理,阻力和谈判。音乐现象是通过相互作用的纠缠态共同产生的,而不是通过预先存在的系统配置或固定的映射。
摘要:This performance presents a duet between two intelligent musical instruments, Sù (to trace back; to go upstream) and Agentier (playing on agentic clavier), and their human performers, connected through feedback loops. Rather than treating AI as a tool that responds predictably to input, both systems operate recursively, where past actions continuously influence future behaviour. The Sù operates in the audio space through latent representation. Its performer uses Make Noise 0-series synthesisers and MIDI controllers to work with a neural feedback synthesis system based on a RAVE model, with a latent feedback loop embedded within the model's internal structure. This allows the instrument to remember and reuse its own internal states, influencing ongoing sound generation through its recent sonic history. The Agentier functions in the control space. Its performer interacts with the system using a Roland S-1 synthesiser and Keith McMillen QuNeo touchpad, where control gestures are routed into a recurrent neural network that feeds back into the synthesis process. Through this feedback loop, the system actively shapes the evolution of control signals over time. Contrasting feedback in the audio and control domains, the performance explores shared agency, resistance, and negotiation between humans and intelligent musical systems. Musical phenomena are co-produced through the entangled states of interaction, rather than through pre-existing system configuration or fixed mappings.


【8】Korean aegyo speech shows systematic F1 increase to signal childlike qualities

标题:韩国Aegyo演讲显示F1系统性增长,标志着孩子般的品质
链接:https://arxiv.org/abs/2604.25133
作者:Ji-eun Kim,Volker Dellwo
备注:18 pages, 2 figures, under review
摘要:韩国语aegyo是一种社会公认的儿童说话风格,主要用于成人之间的浪漫互动。本研究通过分析12名首尔韩国人的共振峰频率,研究了元音空间的修改在Aegyo和非Aegyo风格的相同脚本。结果表明,Aegyo语音的特点是F1值的显着增加,跨元音和选择性前置的前元音,导致元音空间扩大,但主要是转移到更高的F1。这些发现表明,成年人通过模仿儿童的较短声道,主要是通过整体元音降低和部分前置来模仿儿童语言。
摘要:Korean aegyo is a socially recognized childlike speaking style used predominantly in romantic interactions among adults. This study examined vowel space modification in aegyo by analyzing formant frequencies from twelve Seoul Korean speakers who produced identical scripts in aegyo and non-aegyo styles. Results show that aegyo speech features a significant increase in F1 values across vowels and selective fronting of front vowels, leading to vowel space expansion but mainly a shift to higher F1. These findings suggest that adult speakers stylize childlike speech by imitating the shorter vocal tract of children, mainly through global vowel lowering and partial fronting.


【9】S-SONDO: Self-Supervised Knowledge Distillation for General Audio Foundation Models

标题:S-SONDO:通用音频基金会模型的自我监督知识提炼
链接:https://arxiv.org/abs/2604.24933
作者:Mohammed Ali El Adlouni,Aurian Quelennec,Pierre Chouteau,Geoffroy Peeters,Slim Essid
备注:Accepted at IEEE ICASSP 2026. 5 pages, 2 figures, 3 tables. Equal contribution by first two authors. Code: https://github.com/MedAliAdlouni/ssondo | Models: https://huggingface.co/mohammedali2501/ssondo | Package: https://pypi.org/project/ssondo/
摘要:通用音频基础模型最近取得了显着的进步,在不同的任务中实现了强大的性能。然而,最先进的模型仍然非常大,通常具有数亿个参数,导致推理成本高,边缘设备的可部署性有限。知识蒸馏是一种经过验证的模型压缩策略,但之前的音频工作主要集中在监督设置上,依赖于类逻辑、中间特征或特定于架构的技术。这样的假设排除了只输出嵌入的模型,例如自监督或度量学习模型。我们介绍了S-SONDO(Self-Supervised KnOscillation DistillatioN for General AuDio FOUNDATION Models),这是第一个仅使用其输出嵌入来提取通用音频模型的框架。通过避免logits或层级对齐的需要,S-SONDO是架构不可知的,广泛适用于基于嵌入的教师。我们通过将两个音频基础模型提取为三个高效的学生来证明其有效性,这些学生最多可小61倍,同时保留高达96%的教师表现。我们还提供了关于损失选择和基于聚类的平衡数据采样的实用见解。代码可在这里:https://github.com/MedAliAdlouni/ssondo.
摘要:General audio foundation models have recently achieved remarkable progress, enabling strong performance across diverse tasks. However, state-of-the-art models remain extremely large, often with hundreds of millions of parameters, leading to high inference costs and limited deployability on edge devices. Knowledge distillation is a proven strategy for model compression, but prior work in audio has mostly focused on supervised settings, relying on class logits, intermediate features, or architecture-specific techniques. Such assumptions exclude models that output only embeddings, such as self-supervised or metric-learning models. We introduce S-SONDO (Self-Supervised KnOwledge DistillatioN for General AuDio FOundation Models), the first framework to distill general audio models using only their output embeddings. By avoiding the need for logits or layer-level alignment, S-SONDO is architecture-agnostic and broadly applicable to embedding-based teachers. We demonstrate its effectiveness by distilling two audio foundation models into three efficient students that are up to 61 times smaller while retaining up to 96% of teacher performance. We also provide practical insights on loss choice and clustering-based balanced data sampling. Code is available here: https://github.com/MedAliAdlouni/ssondo.


【10】Elderly-Contextual Data Augmentation via Speech Synthesis for Elderly ASR

标题:通过语音合成针对老年人ASB的老年上下文数据增强
链接:https://arxiv.org/abs/2604.24770
作者:Minsik Lee,Seoi Hong,Chongmin Lee,Sieun Choi,Jian Kim,Jua Han,Jihie Kim
备注:5 pages, 2 figures, under review at IEEE Signal Processing Letters
摘要:尽管最近在自动语音识别(ASR)方面取得了进展,但由于训练数据有限以及老年人语音的独特声学和语言特征,老年人ASR(EASR)仍然具有挑战性。在这项工作中,我们通过数据增强管道解决了EASR中的数据稀缺问题,该管道将基于大型语言模型(LLM)的转录释义与文本到语音(TTS)合成相结合。给定老年人语音数据集,LLM首先生成原始转录的老年人上下文释义,然后TTS模型使用老年人参考扬声器合成相应的语音。由此产生的合成音频文本对与原始数据合并,以微调Whisper,而无需修改架构。我们进一步分析了增强比和参考说话人组成在低资源EASR中的影响。在70岁及以上的英语和韩语老年人语音数据集上的实验表明,与传统的增强基线相比,所提出的方法始终如一地提高了性能,与Whisper基线相比,单词错误率(WER)降低了58.2%。
摘要:Despite recent progress in automatic speech recognition (ASR), elderly ASR (EASR) remains challenging due to limited training data and the distinct acoustic and linguistic characteristics of elderly speech. In this work, we address data scarcity in EASR through a data augmentation pipeline that combines large language model (LLM)-based transcript paraphrasing with text-to-speech (TTS) synthesis. Given an elderly speech dataset, the LLM first generates elderly-contextual paraphrases of the original transcripts, and the TTS model then synthesizes corresponding speech using elderly reference speakers. The resulting synthetic audio-text pairs are merged with the original data to fine-tune Whisper without architectural modification. We further analyze the effects of augmentation ratio and reference-speaker composition in low-resource EASR. Experiments on English and Korean elderly speech datasets from speakers aged 70 and above show that the proposed method consistently improves performance over conventional augmentation baselines, achieving up to a 58.2% reduction in word error rate (WER) compared with the Whisper baseline.


【11】Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

标题:穿越不确定性:听觉感知大语言模型不确定性估计的实证研究
链接:https://arxiv.org/abs/2604.25591
作者:Chun-Yi Kuan,Wei-Ping Huang,Hung-yi Lee
备注:Manuscript in progress
摘要:最近的音频感知大型语言模型(ALLM)在各种音频理解和推理任务中表现出了强大的能力,但它们仍然经常产生幻觉或过于自信的输出。虽然不确定性估计在纯文本LLM中得到了广泛的研究,但对于ALLM,它在很大程度上仍未被探索,其中音频条件生成引入了额外的挑战,例如感知模糊性和跨模态接地。在这项工作中,我们提出了第一个系统的实证研究ALLM的不确定性估计。我们对五种有代表性的方法进行了基准测试,包括预测熵、长度归一化熵、语义熵、离散语义熵和P(True),这些方法跨越了多个模型和不同的评估设置,涵盖了一般音频理解、推理、幻觉检测和无法回答的问题回答。我们的研究结果揭示了两个关键发现。首先,语义级和基于验证的方法在一般音频推理基准上始终优于令牌级基线。其次,在以信任为导向的基准测试中,不确定性方法的相对有效性变得更加依赖于模型和基准测试,这表明从一般推理环境中得出的结论不会直接转移到幻觉和无法回答的问题场景中。我们进一步探索基于不确定性的自适应推理作为一个潜在的下游应用。我们希望这项研究提供了可靠的,不确定性感知的音频语言系统的未来研究的基础。
摘要:Recent audio-aware large language models (ALLMs) have demonstrated strong capabilities across diverse audio understanding and reasoning tasks, but they still frequently produce hallucinated or overly confident outputs. While uncertainty estimation has been extensively studied in text-only LLMs, it remains largely unexplored for ALLMs, where audio-conditioned generation introduces additional challenges such as perceptual ambiguity and cross-modal grounding. In this work, we present the first systematic empirical study of uncertainty estimation in ALLMs. We benchmark five representative methods, including predictive entropy, length-normalized entropy, semantic entropy, discrete semantic entropy, and P(True), across multiple models and diverse evaluation settings spanning general audio understanding, reasoning, hallucination detection, and unanswerable question answering. Our results reveal two key findings. First, semantic-level and verification-based methods consistently outperform token-level baselines on general audio reasoning benchmarks. Second, on trustworthiness-oriented benchmarks, the relative effectiveness of uncertainty methods becomes notably more model- and benchmark-dependent, indicating that conclusions drawn from general reasoning settings do not straightforwardly transfer to hallucination and unanswerable-question scenarios. We further explore uncertainty-based adaptive inference as a potential downstream application. We hope this study provides a foundation for future research on reliable, uncertainty-aware audio-language systems.


eess.AS音频处理


【1】Step-Audio-R1.5 Technical Report
标题:Step-Audio-R1.5技术报告
链接:https://arxiv.org/abs/2604.25719
作者:Yuxin Zhang,Xiangyu Tony Zhang,Daijiao Liu,Fei Tian,Yayue Deng,Jun Chen,Qingjian Lin,Haoyang Zhang,Yuxin Li,Jinglan Gong,Yechang Huang,Liang Zhao,Chengyuan Yao,Hexin Liu,Eng Siong Chng,Xuerui Yang,Gang Yu,Xiangyu Zhang,Daxin Jiang
摘要:大型音频语言模型的最新进展将思想链(CoT)推理扩展到听觉领域,使模型能够处理日益复杂的声学和口语任务。为了引出和维持这些扩展的推理链,由基于文本的推理模型的成功驱动的流行范式绝大多数依赖于具有验证奖励的强化学习(RLVR)。然而,随着模型被严格优化,以将丰富、连续的听觉背景提取到孤立的、可验证的文本标签中,一个基本的问题出现了:我们是在培养真正的音频智能,还是仅仅将连续的感官媒介简化为离散的谜题?我们将其称为“可验证的奖励陷阱”。“虽然RLVR在标准化的客观基准上获得了显著的分数,但它系统地降低了音频模型的真实对话感觉。通过优先考虑孤立的正确性,而不是声学上的细微差别,RLVR将动态交互减少为机械的“应答机”,严重损害了韵律的自然性、情感的连续性和用户的沉浸感,特别是在长回合对话中。为了弥合机械客观验证和真正的感官同理心之间的差距,我们引入了Step-Audio-R1.5,标志着音频推理中向人类反馈强化学习(RLHF)的范式转变。全面的评估表明,Step-Audio-R1.5不仅保持了强大的分析推理,而且深刻地改变了交互体验,重新定义了深度沉浸式长时间口语对话的边界。
摘要:Recent advancements in large audio language models have extended Chain-of-Thought (CoT) reasoning into the auditory domain, enabling models to tackle increasingly complex acoustic and spoken tasks. To elicit and sustain these extended reasoning chains, the prevailing paradigm -- driven by the success of text-based reasoning models -- overwhelmingly relies on Reinforcement Learning with Verified Rewards (RLVR). However, as models are strictly optimized to distill rich, continuous auditory contexts into isolated, verifiable text labels, a fundamental question arises: are we fostering true audio intelligence, or merely reducing a continuous sensory medium into a discrete puzzle? We identify this as the "verifiable reward trap." While RLVR yields remarkable scores on standardized objective benchmarks, it systematically degrades the real-world conversational feel of audio models. By prioritizing isolated correctness over acoustic nuance, RLVR reduces dynamic interactions to mechanical "answering machines," severely compromising prosodic naturalness, emotional continuity, and user immersion, particularly in long-turn dialogues. To bridge the gap between mechanical objective verification and genuine sensory empathy, we introduce Step-Audio-R1.5, marking a paradigm shift toward Reinforcement Learning from Human Feedback (RLHF) in audio reasoning. Comprehensive evaluations demonstrate that Step-Audio-R1.5 not only maintains robust analytical reasoning but profoundly transforms the interactive experience, redefining the boundaries of deeply immersive long-turn spoken dialogue.


【2】UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition

标题:基于UNet的融合和指数移动平均自适应用于抗噪说话人识别
链接:https://arxiv.org/abs/2604.25624
作者:Chong-Xin Gan,Peter Bell,Man-Wai Mak,Zhe Li,Zezhong Jin,Zilong Huang,Kong Aik Lee
备注:Submitted to Interspeech 2026
摘要:语音增强和说话人嵌入网络的联合训练是噪声环境下说话人识别的常用方法。虽然有效,但这种范式往往无法利用大规模语音增强预训练中固有的泛化和鲁棒性优势。此外,在去噪语音中保持说话人信息不是语音增强过程的明确目标。为了解决这些局限性,我们提出了一个可扩展的\textbf{U} Net为基础的\textbf{F}语音框架(UF-EMA),认为嘈杂和增强的语音作为多通道输入,从而使扬声器编码器有效地利用扬声器信息。此外,将\textbf{E}指数\textbf{M}oving \textbf{A}平滑策略应用于在干净语音上预训练的扬声器编码器,以减轻过拟合并促进从干净到嘈杂条件的平滑过渡。在多个噪声污染测试集上的实验结果表明了该方法的优越性。
摘要:The joint training of speech enhancement and speaker embedding networks for speaker recognition is widely adopted under noisy acoustic environments. While effective, this paradigm often fails to leverage the generalization and robustness benefits inherent in large-scale speech enhancement pre-training. Moreover, maintaining the speaker information in the denoised speech is not an explicit objective of the speech enhancement process. To address these limitations, we proposed a scalable \textbf{U}Net-based \textbf{F}usion framework (UF-EMA) that considers the noisy and enhanced speech as a multi-channel input, thereby enabling the speaker encoder to exploit speaker information effectively. In addition, an \textbf{E}xponential \textbf{M}oving \textbf{A}verage strategy is applied to a speaker encoder pre-trained on clean speech to mitigate overfitting and facilitate a smooth transition from clean to noisy conditions. Experimental results on multiple noise-contaminated test sets showcase the superiority of the proposed approach.


【3】Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

标题:穿越不确定性:听觉感知大语言模型不确定性估计的实证研究
链接:https://arxiv.org/abs/2604.25591
作者:Chun-Yi Kuan,Wei-Ping Huang,Hung-yi Lee
备注:Manuscript in progress
摘要:最近的音频感知大型语言模型(ALLM)在各种音频理解和推理任务中表现出了强大的能力,但它们仍然经常产生幻觉或过于自信的输出。虽然不确定性估计在纯文本LLM中得到了广泛的研究,但对于ALLM,它在很大程度上仍未被探索,其中音频条件生成引入了额外的挑战,例如感知模糊性和跨模态接地。在这项工作中,我们提出了第一个系统的实证研究ALLM的不确定性估计。我们对五种有代表性的方法进行了基准测试,包括预测熵、长度归一化熵、语义熵、离散语义熵和P(True),这些方法跨越了多个模型和不同的评估设置,涵盖了一般音频理解、推理、幻觉检测和无法回答的问题回答。我们的研究结果揭示了两个关键发现。首先,语义级和基于验证的方法在一般音频推理基准上始终优于令牌级基线。其次,在以信任为导向的基准测试中,不确定性方法的相对有效性变得更加依赖于模型和基准测试,这表明从一般推理环境中得出的结论不会直接转移到幻觉和无法回答的问题场景中。我们进一步探索基于不确定性的自适应推理作为一个潜在的下游应用。我们希望这项研究提供了可靠的,不确定性感知的音频语言系统的未来研究的基础。
摘要:Recent audio-aware large language models (ALLMs) have demonstrated strong capabilities across diverse audio understanding and reasoning tasks, but they still frequently produce hallucinated or overly confident outputs. While uncertainty estimation has been extensively studied in text-only LLMs, it remains largely unexplored for ALLMs, where audio-conditioned generation introduces additional challenges such as perceptual ambiguity and cross-modal grounding. In this work, we present the first systematic empirical study of uncertainty estimation in ALLMs. We benchmark five representative methods, including predictive entropy, length-normalized entropy, semantic entropy, discrete semantic entropy, and P(True), across multiple models and diverse evaluation settings spanning general audio understanding, reasoning, hallucination detection, and unanswerable question answering. Our results reveal two key findings. First, semantic-level and verification-based methods consistently outperform token-level baselines on general audio reasoning benchmarks. Second, on trustworthiness-oriented benchmarks, the relative effectiveness of uncertainty methods becomes notably more model- and benchmark-dependent, indicating that conclusions drawn from general reasoning settings do not straightforwardly transfer to hallucination and unanswerable-question scenarios. We further explore uncertainty-based adaptive inference as a potential downstream application. We hope this study provides a foundation for future research on reliable, uncertainty-aware audio-language systems.


【4】ASAP: An Azimuth-Priority Strip-Based Search Approach to Planar Microphone Array DOA Estimation in 3D

标题:ASAP:一种基于方位优先条的3D平面麦克风阵列DOE估计方法
链接:https://arxiv.org/abs/2604.25387
作者:Ming Huang,Shuting Xu,Leying Yang,Huanzhang Hu,Yujie Zhang,Jiang Wang,Yu Liu,Hao Zhao,He Kong
备注:This paper has been accepted to the Fourteenth IEEE Sensor Array and Multichannel Signal Processing Workshop, 2026
摘要:波达方向(DOA)估计是麦克风阵列处理和许多下游应用中的重要任务。相位变换引导响应功率法(SRP-PHAT)是近年来被广泛采用的DOA估计方法。然而,3D场景中的准确SRP-PHAT估计需要评估数千个候选方向上的转向响应,严重限制了资源受限平台上的实时性能。这一挑战对于平面阵列变得更加关键,平面阵列由于其结构简单而广泛用于机器人。由于方位角估计通常比仰角估计更可靠,我们提出了ASAP,一个方位角优先的条带搜索方法,平面麦克风阵列DOA估计在3D。在第一阶段中,ASAP在方位带内执行由粗到细的区域收缩以锁定方位角,同时通过球冠保留多个最大值。在第二阶段,它沿着两个接近的候选者之间的大圆弧细化高程。大量的仿真和真实世界的实验验证了所提出的方法比现有的方法的效率和优点。
摘要:Direction-of-arrival (DOA) estimation is an important task in microphone array processing and many downstream applications. The steered response power with phase transform (SRP-PHAT) method has been widely adopted for DOA estimation in recent years. However, accurate SRP-PHAT estimation in 3D scenarios requires evaluating steering responses over thousands of candidate directions, severely limiting real-time performance on resource-constrained platforms. This challenge becomes even more critical for planar arrays, which are widely used in robotics due to their structural simplicity. Motivated by the fact that azimuth estimation is usually more reliable than elevation estimation for most arrays, we propose ASAP, an azimuth-priority strip-based search approach to planar microphone array DOA estimation in 3D. In the first stage, ASAP performs coarse-to-fine region contraction within azimuthal strips to lock azimuth angles while retaining multiple maxima through spherical caps. In the second stage, it refines elevation along the great-circle arc between two close candidates. Extensive simulations and real-world experiments validate the efficiency and merits of the proposed method over existing approaches.


【5】Cross-Linguistic Rhythmic and Spectral Feature-Based Analysis of Nyishi and Adi: Two Under-Resourced Languages of Arunachal Pradesh

标题:基于跨语言节奏和频谱特征的Nyishi和Adi分析:阿鲁纳恰尔邦两种资源不足的语言
链接:https://arxiv.org/abs/2604.25309
作者:Deepshikha Gogoi,Parismita Gogoi,Yang Saring
备注:Submitted to Sadhana (Indian Academy of Sciences); currently under consideration
摘要:资源不足的语言在定量节奏研究中仍然代表性不足,特别是在系统的分支内分析声学分化密切相关的linguisticsgroups.This研究调查声学分化内的塔尼语亚组通过检查语音节奏在Nyishi和阿迪,两个资源不足的塔尼语在阿鲁纳恰尔邦,印度东北部,使用基于幅度调制(AM)低频(LF)频谱分析的频域框架,通常被称为节奏共振峰分析(RFA)。该分析被设计为识别内部分支分化遵循跨节奏和频谱域的分级模式。从LF调制频谱,导出三个节奏共振峰特征:主峰数(NDP)、主峰平均频率(MFDP)和主频方差(VFDP)。此外,通过提取离散余弦变换(DCT)系数和梅尔倒谱系数(MFCC)来表征语音信号的频谱调制结构和宽谱组织,统计建模揭示了语音信号的分级分化模式,其中节奏特征表现出一致但适度的分离,其中Nyishi表现出比Adi更高的主调制频率以及更大的色散。分类实验进一步支持这种层次,使用MFCC表示的融合将性能提高到使用支持向量机(SVM)的90.9%的分类准确率和使用多层感知器(MLP)的93.96%。这些发现表明,节奏和频谱特征编码了语言变异的互补水平,低频调制捕获了受约束的宏观时间结构,频谱特征反映了更精细的语音分化。
摘要:Under-resourced languages remain underrepresented in quantitative rhythm research,particularly in systematic intra-branch analysis of acoustic differentiation within closely related linguistic groups.This study investigates acoustic differentiation within the Tani language subgroup by examining speech rhythm in Nyishi and Adi,two under-resourced Tani languages spoken in Arunachal Pradesh,North-East India,using a frequency domain framework based on amplitude modulation(AM) low-frequency(LF) spectrum analysis,commonly referred to as rhythm formant analysis(RFA).The analysis is designed to identify whether intra-branch differentiation follows a hierarchical pattern across rhythmic and spectral domains.From the LF modulation spectrum,three rhythm formant features were derived:Number of Dominant peaks(NDP),Mean Frequency of Dominant Peaks(MFDP),and Variance of Dominant Frequencies(VFDP).In addition,Discrete Cosine Transform (DCT)coefficients and Mel Frequency Cepstral Coefficient(MFCC) were extracted to characterise the spectral modulation structure and broad spectral organisation of the speech signal.Statistical modelling reveals a hierarchical pattern of differentiation,where rhythmic features show consistent but moderate separation,with Nyishi exhibiting higher dominant modulation frequencies as well as greater dispersion than Adi.Classification experiments further support this hierarchy,with rhythm-only features achieved approximately 84-85% classification accuracy.Fusion using MFCC representations improved performance to 90.9% classification accuracy using support vector machine (SVM) and 93.96% using multilayer perceptron (MLP).These findings demonstrate that rhythmic and spectral features encode complementary levels of linguistic variations,with low frequency modulation capturing constrained macro temporal structure and spectral features reflecting finer phonological differentiation.


【6】Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

标题:Prachy Voice:语音提示恢复+ BUPS,以零商业训练数据成本从冷冻非印度基地获得商业级印度TTC
链接:https://arxiv.org/abs/2604.25441
作者:Venkata Pushpak Teja Menta
备注:9 pages, 6 figures, 6 tables. Companion paper to PSP benchmark. Code: https://github.com/praxelhq/praxy ; Model: https://huggingface.co/Praxel/praxy-voice-r6 ; Demo: https://huggingface.co/spaces/Praxel/praxy-voice-demo
摘要:商业TTS系统产生接近原生的印度语音频,但最好的开源基础(Chatterbox,Indic Parler-TTS,IndicF 5)在测量的语音维度上落后于它们,而最广泛采用的多语言基础(Chatterbox,23种语言)甚至没有标记泰卢固语或泰米尔语。我们问:在不训练新的声学解码器和没有任何商业TTS训练数据的情况下,将这样的非印度语本地基础带到泰卢固语、泰米尔语和印地语的商业级输出的最小干预是什么?我们结合三个部分:(1)BUPS,一个Brahmic统一音素空间,可将七种印度文字确定性地罗马化为ISO-15919,以便Chatterbox的拉丁标记器可以处理它们;(2)仅在文本标记预测器上的LoRA适配器(Chatterbox的t3),在大约1,220小时的许可印度语音频上训练,使用Hindi代理语言_id;(3)语音提示恢复配方--8- 11秒的相同语言参考剪辑加上三个采样覆盖(夸张0.7,温度0.6,min_p 0.1;“配置B”)--恢复商业级声学输出,无需声学解码器训练。在印地语中,LoRA回归准确性,我们使用vanilla Chatterbox + Config B,给出两个分支部署。在10个发音试点集与同伴PSP基准评估,Praxy语音匹配或略有领先的商业基线:26.7%卷舌崩溃泰卢固语(与Sarvam Bulbul 33.3%),71%泰米尔扎崩溃(与商业三人组的86%),0.025 LLM-WER印地语(与Cartesia Sonic-3并列)。对于内部代码混合,我们添加第三分支(IndicF 5+本机脚本音译),其跨Hi/Te/Ta将代码混合LLM-WER从0.80-0.85下降到0.14-0.27。我们发布了R6 LoRA权重(Apache-2.0),推理代码和路由器(MIT)以及Gradio演示。
摘要:Commercial TTS systems produce near-native Indic audio, but the best open-source bases (Chatterbox, Indic Parler-TTS, IndicF5) trail them on measured phonological dimensions, and the most widely adopted multilingual base (Chatterbox, 23 languages) does not even tokenise Telugu or Tamil. We ask: what is the minimum intervention that brings such a non-Indic-native base to commercial-class output on Telugu, Tamil, and Hindi, without training a new acoustic decoder and without any commercial TTS training data? We combine three pieces: (1) BUPS, a Brahmic Unified Phoneme Space that deterministically romanises seven Indic scripts to ISO-15919 so Chatterbox's Latin tokeniser can process them; (2) a LoRA adapter on only the text-token predictor (Chatterbox's t3), trained on ~1,220h of licensed Indic audio with a Hindi-proxy language_id; (3) a voice-prompt recovery recipe -- an 8-11s same-language reference clip plus three sampling overrides (exaggeration 0.7, temperature 0.6, min_p 0.1; "Config B") -- that recovers commercial-class acoustic output with no acoustic-decoder training. On Hindi, the LoRA regresses accuracy and we instead use vanilla Chatterbox + Config B, giving a two-branch deployment. Evaluated on 10-utterance pilot sets with the companion PSP benchmark, Praxy Voice matches or slightly leads commercial baselines: 26.7% retroflex collapse on Telugu (vs Sarvam Bulbul 33.3%), 71% Tamil-zha collapse (vs commercial trio's 86%), 0.025 LLM-WER on Hindi (tied with Cartesia Sonic-3). For intra-sentential code-mix we add a third branch (IndicF5 + native-script transliteration) that drops code-mix LLM-WER from 0.80-0.85 to 0.14-0.27 across Hi/Te/Ta. We release R6 LoRA weights (Apache-2.0), inference code and router (MIT), and a Gradio demo.


【7】ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

标题:ML-SAN:用于对话中情感识别的多层扬声器自适应网络
链接:https://arxiv.org/abs/2604.25383
作者:Kexue Wang,Yinfeng Yu,Liejun Wang
备注:Main paper (12 pages). Accepted for publication by International Conference on Intelligent Computing 2026
摘要:要与机器建立同理心,必须充分了解人类的情绪变化。然而,多模态情绪识别的研究往往忽略了一个问题:个体的表达特质差异很大,这意味着不同的人可能会表达不同的情绪。在日常生活中,我们可以看到这一点。在与不同的人交流时,有些人通过他们的面部表情和语言表达“快乐”,而另一些人可能会隐藏他们的快乐或通过他们的行动表达。两者都是“快乐”的表达,但这种情感表达的差异对于机器来说仍然太难区分。当前的情感识别仍然处于“静态”水平,使用单个识别模型来识别所有情感风格。这种“简化”往往会影响识别结果,特别是在多话轮对话中。为了解决这个问题,本文介绍了一种新的多级说话人自适应网络(ML-SAN),其中,具体地说,有效地解决了说话人身份信息混淆的挑战。ML-SAN在识别后并不简单地分配说话者的ID;相反,它采用了三个阶段的自适应过程:首先,输入级校准使用音频级线性调制(FiLM)将原始音频和视觉特征调整到与说话者无关的中性空间。然后,交互级别门控重新调整每个模态的信任级别(例如,语音或面部特征)。最后,输出级正则化保持了潜在空间中说话人特征的一致性。在MELD和IEMOCAP数据集上的测试表明,我们的模型(ML-SAN)取得了更好的结果,在处理具有挑战性的尾部情感类别方面表现出色,并且更好地解决了现实世界场景中说话者的多样性。
摘要:To establish empathy with machines, it is essential to fully understand human emotional changes. However, research in multimodal emotion recognition often overlooks one problem: individual expressive traits vary significantly, which means that different people may express emotions differently. In our daily lives, we can see this. When communicating with different people, some express "happiness" through their facial expressions and words, while others may hide their happiness or express it through their actions. Both are expressions of 'happiness,' but such differences in emotional expression are still too difficult for machines to distinguish. Current emotion recognition remains at a 'static' level, using a single recognition model to identify all emotional styles. This "simplification" often affects the recognition results, especially in multi-turn dialogues. To address this problem, this paper introduces a novel Multi-Level Speaker Adaptive Network (ML-SAN), which, specifically, effectively addresses the challenge of speaker identity information confusion. ML-SAN does not simply assign a speaker's ID after recognition; instead, it employs a three-stage adaptive process: First, Input-level Calibration uses Feature-Level Linear Modulation (FiLM) to adjust the raw audio and visual features into a neutral space unrelated to the speaker. Then, Interaction-level Gating re-adjusts the trust level for each modality (e.g., voice or facial features) based on the speaker's identity information. Finally, Output-level Regularization maintains the consistency of speaker features in the latent space. Tests on the MELD and IEMOCAP datasets show that our model (ML-SAN) achieves better results, performs exceptionally well in handling challenging tail sentiment categories, and better addresses the diversity of speakers in real-world scenarios.


【8】Korean aegyo speech shows systematic F1 increase to signal childlike qualities

标题:韩国Aegyo演讲显示F1系统性增长,标志着孩子般的品质
链接:https://arxiv.org/abs/2604.25133
作者:Ji-eun Kim,Volker Dellwo
备注:18 pages, 2 figures, under review
摘要:韩国语aegyo是一种社会公认的儿童说话风格,主要用于成人之间的浪漫互动。本研究通过分析12名首尔韩国人的共振峰频率,研究了元音空间的修改在Aegyo和非Aegyo风格的相同脚本。结果表明,Aegyo语音的特点是F1值的显着增加,跨元音和选择性前置的前元音,导致元音空间扩大,但主要是转移到更高的F1。这些发现表明,成年人通过模仿儿童的较短声道,主要是通过整体元音降低和部分前置来模仿儿童语言。
摘要:Korean aegyo is a socially recognized childlike speaking style used predominantly in romantic interactions among adults. This study examined vowel space modification in aegyo by analyzing formant frequencies from twelve Seoul Korean speakers who produced identical scripts in aegyo and non-aegyo styles. Results show that aegyo speech features a significant increase in F1 values across vowels and selective fronting of front vowels, leading to vowel space expansion but mainly a shift to higher F1. These findings suggest that adult speakers stylize childlike speech by imitating the shorter vocal tract of children, mainly through global vowel lowering and partial fronting.


机器翻译由腾讯交互翻译提供,仅供参考