微信公众号:arXiv_Daily
cs.SD语音
标题:BeeVe:蜜蜂嗡嗡声中的无监督声学状态发现
链接:https://arxiv.org/pdf/2605.07903v1
摘要:在没有监督的情况下发现生物信号中的结构是计算智能中的一个基本问题,然而现有的生物声学方法假设发声模型或预定义的语义单元,使得非发声物种服务不佳。这项工作介绍了BeeVe,一个无监督的框架,在集体蜜蜂嗡嗡声的声学状态发现。BeeVe使用自监督Patchout Spectrogram Transformer(PaSST)作为冻结特征提取器,然后在这些嵌入上训练没有标签的矢量量化变分自动编码器(VQ-VAE),直接从未标记的蜂巢音频中学习声学标记的有限离散码本。在任何阶段都不使用标签、借口任务或对比目标。对已知皇后状态的事后评估表明,学习的令牌将具有0.609和0.688之间的Jensen-Shannon散度值的皇后和无皇后条件分开,并且无皇后条件进一步分解成三个内部相干子状态,这些子状态在具有不同码本大小和随机种子的实验中稳定。标记转换分析证实了所有实验中的非随机序列结构(p << 0.001)。推广到看不见的记录保留令牌重叠(Jaccard = 0.947)和全球流形拓扑。这些结果表明,无监督离散码本学习可以在没有注释的情况下从非声音生物信号中恢复可重复的声学结构,为非侵入性声学蜂巢健康监测开辟了道路。摘要:Discovering structure in biological signals without supervision is a fundamental problem in computational intelligence, yet existing bioacoustic methods assume vocal production models or predefined semantic units, leaving non-vocal species poorly served. This work introduces BeeVe, an unsupervised framework for acoustic state discovery in collective honey bee buzzing. BeeVe uses the self-supervised Patchout Spectrogram Transformer (PaSST) as a frozen feature extractor, then trains a Vector-Quantized Variational Autoencoder (VQ-VAE) without labels on those embeddings, learning a finite discrete codebook of acoustic tokens directly from unlabelled hive audio. No labels, pretext tasks, or contrastive objectives are used at any stage. Post-hoc evaluation against known queen status reveals that the learned tokens separate queenright and queenless conditions with Jensen-Shannon Divergence values between 0.609 and 0.688, and that the queenless condition further decomposes into three internally coherent sub-states stable across experiments with different codebook sizes and random seeds. Token transition analysis confirms non-random sequential structure (p << 0.001) across all experiments. Generalisation to unseen recordings preserves both token overlap (Jaccard = 0.947) and global manifold topology. These results demonstrate that unsupervised discrete codebook learning can recover repeatable acoustic structure from a non-vocal biological signal without annotation, opening a path toward non-invasive acoustic hive health monitoring.
【2】TARNet: A Temporal-Aware Multi-Scale Architecture for Closed-Set Speaker Identification
标题:TARNet:一种用于闭集说话人识别的时间感知多尺度架构链接:https://arxiv.org/pdf/2605.07735v1
备注:Accepted at IEEE International Conference on Multimedia and Expo (ICME) 2026. Code available at:
摘要:闭集说话人识别的目标是将语音话语分配给预定义的一组注册说话人中的一个,并且需要跨多个时间尺度对说话人特定特征进行鲁棒建模。虽然最近的深度学习方法已经实现了强大的性能,但许多现有的架构提供了有限的机制来跨不同的时间尺度建模时间依赖性,这可能会限制互补的短期,中期和长期扬声器特征的有效使用。在本文中,我们提出了TARNet,一个轻量级的时间感知表示网络,用于闭集说话人识别。TARNet使用具有阶段特定膨胀配置的多级时间编码器在多个时间尺度上显式地对时间信息进行建模。所得到的多尺度表示融合和聚合通过一个注意统计池(ASP)模块,以产生一个有区别的话语级说话人嵌入。在VoxCeleb1和LibriSpeech数据集上的实验表明,TARNet在保持有竞争力的计算复杂度的同时,性能优于最先进的方法,适用于实际的说话人识别系统。该代码可在https: github.com YassinTERRAF TARNet上公开获得。摘要:Closed-Set speaker identification aims to assign a speech utterance to one of a predefined set of enrolled speakers and requires robust modeling of speaker-specific characteristics across multiple temporal scales. While recent deep learning approaches have achieved strong performance, many existing architectures provide limited mechanisms for modeling temporal dependencies across different time scales, which can restrict the effective use of complementary short-, mid-, and long-term speaker characteristics. In this paper, we propose TARNet, a lightweight Temporal-Aware Representation Network for closed-set speaker identification. TARNet explicitly models temporal information at multiple time scales using a multi-stage temporal encoder with stage-specific dilation configurations. The resulting multi-scale representations are fused and aggregated via an Attentive Statistics Pooling (ASP) module to produce a discriminative utterance-level speaker embedding. Experiments on the VoxCeleb1 and LibriSpeech datasets show that TARNet outperforms state-of-the-art methods while maintaining competitive computational complexity, making it suitable for practical speaker identification systems. The code is publicly available at https: github.com YassinTERRAF TARNet.
【3】A Decomposed Retrieval-Edit-Rerank Framework for Chord Generation
标题:用于和弦生成的分解检索-编辑-Rerank框架链接:https://arxiv.org/pdf/2605.07489v1
备注:Accepted by the 2026 ACM International Conference on Multimedia Retrieval (ICMR 2026)
摘要:和弦生成是一个固有的限制创造性的任务,需要平衡风格的多样性与音乐理论的可行性。现有的方法通常纠缠在一个单一的模型中的候选人生成和约束执行,使得多样性可行性权衡难以控制和解释。在这项工作中,我们从系统级的角度来处理和弦生成,引入检索-编辑-重新排序(RER)框架,将任务分解为三个明确的阶段:i)检索,定义一个风格上合理的候选空间; ii)编辑,通过最小的修改来执行音乐理论的可行性; iii)重新排序,解决可行候选人之间的软偏好。这种分离提供了一个可控的管道,其中每个组件解决了生成过程的不同方面,从而增强了输出和弦的可解释性和可调整性。通过客观指标和主观评价,我们的分解系统在平衡和弦多样性和音乐理论可行性方面优于所有端到端和弦生成基线。消融研究进一步证实了每个阶段在创造性探索和约束满足中的互补作用。摘要:Chord generation is an inherently constrained creative task that requires balancing stylistic diversity with music-theoretic feasibility. Existing approaches typically entangle candidate generation and constraint enforcement within a single model, making the diversity-feasibility trade-off difficult to control and interpret. In this work, we approach chord generation from a system-level perspective, introducing a Retrieval-Edit-Rerank (RER) framework that decomposes the task into three explicit stages: i) retrieval, which defines a stylistically plausible candidate space; ii) editing, which enforces music-theoretic feasibility through minimal modifications; and iii) reranking, which resolves soft preferences among feasible candidates. This separation provides a controllable pipeline, where each component addresses a distinct aspect of the generation process, thereby enhancing both the interpretability and adjustability of the output chords. Through objective metrics and subjective evaluation, our decomposed system outperforms all end-to-end chord generation baselines in balancing chord diversity and music-theoretic feasibility. Ablation studies further confirm the complementary roles of each stage in creative exploration and constraint satisfaction.【4】Do Joint Audio-Video Generation Models Understand Physics?
标题:联合音频视频生成模型理解物理吗?链接:https://arxiv.org/pdf/2605.07061v1
备注:Preprint. Full abstract appears in the PDF
摘要:联合音视频生成模型正在迅速接近专业的制作质量,这提出了一个核心问题:它们是否理解视听物理学,或者仅仅生成违反现实世界一致性的合理声音和帧?我们介绍AV-Phys Bench,这是一个用于评估联合音频-视频生成中的物理常识的基准。AV-Phys Bench测试三个场景类别的模型:稳态,事件转换和环境转换。它涵盖了从真实世界场景中提取的物理基础子类别,以及故意要求物理上不一致的音频-视频行为的反AV-物理提示。每一代都沿着五个维度进行评估:视觉语义坚持,音频语义坚持,视觉物理常识,音频物理常识和跨模态物理常识。在三个专有模型和四个开源模型中,我们发现Seedance 2.0整体表现最好,但所有模型都远远没有达到强大的物理理解。在事件驱动和环境驱动的转换中,性能会急剧下降,甚至强大的专有系统也会在Anti-AV-Physics提示下崩溃。我们进一步介绍了AV-Phys Agent,这是一种ReAct风格的评估器,它将多模态语言模型与确定性声学测量工具相结合,产生与人类评级密切相关的排名。我们的研究结果确定跨模态的物理一致性和过渡驱动的场景动态的关键开放的挑战,联合音视频生成。摘要:Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visual physics, or merely generate plausible sounds and frames that violate real-world consistency? We introduce AV-Phys Bench, a benchmark for evaluating physical commonsense in joint audio-video generation. AV-Phys Bench tests models across three scene categories: Steady State, Event Transition, and Environment Transition. It covers physics-grounded subcategories drawn from real-world scenes, plus Anti-AV-Physics prompts that deliberately request physically inconsistent audio-video behavior. Each generation is evaluated along five dimensions: visual semantic adherence, audio semantic adherence, visual physical commonsense, audio physical commonsense, and cross-modal physical commonsense. Across three proprietary and four open-source models, we find that Seedance 2.0 performs best overall, but all models remain far from robust physical understanding. Performance drops sharply on event-driven and environment-driven transitions, and even strong proprietary systems collapse on Anti-AV-Physics prompts. We further introduce AV-Phys Agent, a ReAct-style evaluator that combines a multimodal language model with deterministic acoustic measurement tools, producing rankings that closely align with human ratings. Our results identify cross-modal physical consistency and transition-driven scene dynamics as key open challenges for joint audio-video generation.
【5】MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
标题:MIST:智能家居的多模式交互式基于语音的工具呼叫对话助理链接:https://arxiv.org/pdf/2605.06897v1
备注:Project Page: https:billyzhang24kobe.github.iomist-smarthome
摘要:物联网(IoT)设备在物理世界中的兴起需要能够处理复杂用户体验的基于语音的接口。虽然现代大型语言模型(LLM)已经展示了强大的工具使用能力,但对现实世界的物联网设备进行建模是一个困难的、未充分研究的挑战,它将建模时空约束与语音输入、动态状态跟踪和混合主动交互模式相结合。我们介绍了MIST(多模态交互式基于语音的工具调用数据集),这是一种在物联网设备上运行的合成多回合、语音驱动的代码生成任务。我们发现,有一个显着的差距开放和封闭的重量多模态LLM MIST,即使是边境封闭重量LLM有很大的净空。我们发布了MIST和一个可扩展的数据生成框架来构建相关的数据集,以促进对混合主动语音助手的研究,这些语音助手会考虑物理世界的限制。摘要:The rise of Internet of Things (IoT) devices in the physical world necessitates voice-based interfaces capable of handling complex user experiences. While modern Large Language Models (LLMs) already demonstrate strong tool-usage capabilities, modeling real-world IoT devices presents a difficult, understudied challenge which combines modeling spatiotemporal constraints with speech inputs, dynamic state tracking, and mixed-initiative interaction patterns. We introduce MIST (the Multimodal Interactive Speech-based Tool-calling Dataset), a synthetic multi-turn, voice-driven code generation task that operates over IoT devices. We find that there is a significant gap between open- and closed-weight multimodal LLMs on MIST, and that even frontier closed-weight LLMs have substantial headroom. We release MIST and an extensible data generation framework to build related datasets in order to facilitate research on mixed-initiative voice assistants which reason about physical world constraints.
【6】An audio-to-analysis pipeline with certified transcription for information-theoretic profiling of the piano repertoire
标题:具有经过认证的转录的音频到分析管道,用于钢琴曲目的信息理论分析链接:https://arxiv.org/pdf/2605.06685v1
备注:25 pages, 4 figures, 25 references
摘要:我们提出了一个音频分析管道,产生作曲家级的信息理论配置文件:反映合成词汇,因为它出现在聚合性能:从原始录音,建立在转录层,其准确性,我们证明了一个标准的基准(F1 = 0.9791的MAESTRO v3.0.0测试集)。应用于1,238件作品和15个MAESTRO作曲家,至少有10个归因于作品,跨越巴洛克风格到20世纪初,管道得出谐波音阶度的经验分布,并通过香农熵,不对称Kullback-Leibler发散和Zipfian秩频率建模对其进行分析。由此产生的配置文件(i)顺序作曲家沿着一个可解释的轴的和谐的可预测性,与一个狭窄的熵范围(3.33-3.86位),揭示了音调词汇的边缘水平相似性;(ii)恢复已知的风格谱系(海顿-贝多芬,李斯特-拉赫玛尼诺夫,舒伯特-舒曼)通过语料库中最小的KL分歧,门德尔松作为一个稳定的离群值出现在这个语料库;(iii)根据齐普菲对过渡分布的拟合程度,将当代新古典主义艺术家(里希特、弗拉姆、格拉斯、阿纳尔兹、约翰松)与历史作曲家区分开来,新古典主义艺术家的平均R^2 = 0.78,历史作曲家的平均R^2 = 0.46(每人10件)。这一差距大于任何一个群体内的传播,并符合极简主义的作曲倾向:一个紧凑的过渡词汇使用更尖锐的频率等级规律比历史作曲家。所有估计值均以拉普拉斯平滑Bootstrap 95%置信区间报告。摘要:We present an audio-to-analysis pipeline that produces composer-level information-theoretic profiles : reflecting compositional vocabulary as it emerges from aggregated performances : from raw recordings, built on a transcription layer whose accuracy we certify on a standard benchmark (F1 = 0.9791 on the MAESTRO v3.0.0 test set). Applied to 1,238 pieces and 15 MAESTRO composers with at least ten attributed pieces, spanning the Baroque through the early twentieth century, the pipeline derives empirical distributions over harmonic scale degrees and analyzes them through Shannon entropy, asymmetric Kullback-Leibler divergence, and Zipfian rank-frequency modeling. The resulting profiles (i) order composers along an interpretable axis of harmonic predictability, with a narrow entropy range (3.33-3.86 bits) that reveals the marginal-level similarity of tonal vocabularies; (ii) recover known stylistic lineages (Haydn-Beethoven, Liszt-Rachmaninoff, Schubert-Schumann) through the smallest KL divergences in the corpus, with Mendelssohn emerging as a stable outlier within this corpus; and (iii) separate contemporary neoclassical artists (Richter, Frahm, Glass, Arnalds, Jóhannsson) from historical composers on the quality of Zipfian fit to the transition distribution, with mean $R^2 = 0.78$ for neoclassical versus 0.46 for historical (N $ geq$ 10 pieces each). This gap is larger than the spread within either group and is consistent with a minimalist compositional tendency: a compact transition vocabulary used with sharper frequency-rank regularity than historical composers. All estimates are reported with Laplace-smoothed bootstrap 95% confidence intervals.
【7】Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
标题:单通道说话人距离估计对早期和晚期回响的依赖性链接:https://arxiv.org/pdf/2605.07694v1
备注:Submitted to IWAENC 2026
摘要:单通道扬声器距离估计最近在模拟环境中达到了厘米级的精度,但仍不清楚该模型利用了房间脉冲响应(RIR)的哪些组件以及性能如何取决于录音条件。在这项工作中,我们将模拟RIR分解成四个变量(全,直接,无晚,无早)使用混合时间估计的回波密度函数作为早期反射和晚期混响之间的边界。我们定义了四种校准场景,从完全校准(同步捕获,已知源水平)到完全未校准(任意起始,未知水平),并在匹配的数据集上评估所有组合。结果表明,没有时间校准,平均绝对误差(MAE)增加到1.29 $ M和模型提取混响为基础的线索,早期的反射出现的最翔实的组成部分。针对DRR、$C_{50}$和$T_{60}$的进一步分析证实,估计精度随着较强的早期能量而提高,并且在高混响环境中降低。当时间校准可用时,该模型通过单独提取传播延迟实现了0.14 $ m的MAE,而不管RIR内容如何。摘要:Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclear which components of the room impulse response (RIR) the model exploits and how performance depends on the recording conditions. In this work, we decompose simulated RIRs into four variants (full, direct-only, no-late, and no-early) using the mixing time estimated from the echo density function as the boundary between early reflections and late reverberation. We define four calibration scenarios, from fully calibrated (synchronised capture, known source level) to fully uncalibrated (arbitrary onset, unknown level), and evaluate all combinations on a matched dataset. Results show that without time calibration, mean absolute error (MAE) increases to $1.29$ m and the model extracts reverberation-based cues, with early reflections emerging as the most informative component. Further analysis against DRR, $C_{50}$, and $T_{60}$ confirms that estimation accuracy improves with stronger early energy and degrades in highly reverberant environments. When time calibration is available, the model achieves a MAE of $0.14$ m by extracting the propagation delay alone, regardless of the RIR content.
【1】Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
标题:单通道说话人距离估计对早期和晚期回响的依赖性【2】Evaluating voice anonymisation using similarity rank disclosure
标题:使用相似度等级披露评估语音匿名化链接:https://arxiv.org/pdf/2605.07291v1
备注:DCC 2026
摘要:语音匿名化的评估仍然具有挑战性。目前的实践依赖于自动说话人验证指标,如等错误率(EER)。依赖于分类器和操作点的性能估计提供了对隐私风险的不完整甚至误导的表征。我们调查使用的相似性排名披露(SRD),信息理论的度量,操作的特征表示,而不是分类器的决定,提供了一个阈值独立的评估隐私和分析的平均和最坏情况下的披露。我们报告其应用于扬声器嵌入,基频,和电话嵌入使用2024年语音隐私挑战赛系统。SRD揭示了基于EER的评估所遗漏的隐私泄漏和系统特定的弱点。研究结果突出了代表级指标的优点,并展示了SRD作为语音匿名评估的灵活和可解释的工具的潜力。摘要:The evaluation of voice anonymisation remains challenging. Current practice relies on automatic speaker verification metrics such as the equal error rate (EER). Performance estimates dependent on the classifier and operating point provide an incomplete or even misleading characterisation of privacy risk. We investigate the use of similarity rank disclosure (SRD), an information-theoretic metric, which operates on feature representations rather than classifier decisions, providing a threshold-independent assessment of privacy and analysis of both average and worst-case disclosure. We report its application to speaker embeddings, fundamental frequency, and phone embeddings using 2024 VoicePrivacy Challenge systems. The SRD reveals privacy leaks and system-specific weaknesses missed by EER-based evaluation. Findings highlight the merit of representation-level metrics and demonstrate the potential of SRD as a flexible and interpretable tool for the evaluation of voice anonymisation.
【3】Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping
标题:通过想象到收听的MEG映射的Zero-Shot想象语音解码链接:https://arxiv.org/pdf/2605.08075v1
摘要:从非侵入性的大脑记录解码想象的语音是具有挑战性的,因为想象的数据集是稀缺的,很难在时间上跨学科和会话对齐。在这项工作中,我们提出了一种新的方法来解码想象的语音,利用更丰富,更可靠的标记录音在听语音。我们收集了成对的听和想象的MEG录音节奏旋律和口语刺激训练有素的音乐家。使用训练有素的音乐家有助于改善各种条件下的时间对齐。然后,我们开发了一个三阶段的解码管道,揭示了想象和倾听相同刺激所诱发的神经活动之间的一致和有意义的关系。首先,我们训练了六个线性和神经模型,将想象的MEG反应映射到倾听的反应。我们评估这些模型对一个空基线从看不见的主题,以验证预测的听力反应保留刺激特定的信息。在第二阶段中,我们专门针对所听的MEG响应训练了一个对比词解码器,并使用四种嵌入策略(包括语义、声学和语音表示)对其进行评估。在第三阶段中,我们通过映射管道处理来自保持的受试者的想象MEG响应,以计算相应的听力响应,然后由听力解码器解码。使用基于等级的分析,我们表明,想象的话是解码显着以上的机会。我们将在这里报告一个概念验证实现的结果,以解码想象中的语音,其中所有的评估都是在保持的主题上进行的。我们还证明了性能随着训练数据的大小而提高,这表明这种方法是可扩展的,可以直接适用于现实的脑机接口场景。摘要:Decoding imagined speech from non-invasive brain recordings is challenging because imagined datasets are scarce and difficult to align temporally across subjects and sessions In this work, we propose a new approach to the decoding of imagined speech that leverages the richer and more reliably labeled recordings during listening to speech. We collected paired listened and imagined MEG recordings to rhythmic melodic and spoken stimuli from trained musicians. Using trained musicians helped improve temporal alignment across conditions. We then developed a three-stage decoding pipeline that revealed consistent and meaningful relationships between neural activity evoked by imagining and listening to the same stimuli. First, we trained six linear and neural models to map imagined MEG responses to listened responses. We evaluated these models against a null baseline from unseen subjects to validate that the predicted-listening responses preserve stimulus-specific information. In the second stage, we trained a contrastive word decoder exclusively on the listened MEG responses, and evaluated it using four embedding strategies including semantic, acoustic, and phonetic representations. In the third stage, we process the imagined MEG responses from held-out subjects through the mapping pipeline to compute the corresponding listening responses that are then decoded by the listened decoder. Using rank-based analysis, we show that the imagined words are decodable significantly above chance. We shall report here the results of a proof-of-concept implementation to decode imagined speech, where all evaluations are performed on held-out subjects. We also demonstrate that performance improves with training data size, suggesting that this approach is scalable and can directly be made applicable to realistic brain-computer interface scenarios.
【4】Do Joint Audio-Video Generation Models Understand Physics?
标题:联合音频视频生成模型理解物理吗?链接:https://arxiv.org/pdf/2605.07061v1
备注:Preprint. Full abstract appears in the PDF
摘要:联合音视频生成模型正在迅速接近专业的制作质量,这提出了一个核心问题:它们是否理解视听物理学,或者仅仅生成违反现实世界一致性的合理声音和帧?我们介绍AV-Phys Bench,这是一个用于评估联合音频-视频生成中的物理常识的基准。AV-Phys Bench测试三个场景类别的模型:稳态,事件转换和环境转换。它涵盖了从真实世界场景中提取的物理基础子类别,以及故意要求物理上不一致的音频-视频行为的反AV-物理提示。每一代都沿着五个维度进行评估:视觉语义坚持,音频语义坚持,视觉物理常识,音频物理常识和跨模态物理常识。在三个专有模型和四个开源模型中,我们发现Seedance 2.0整体表现最好,但所有模型都远远没有达到强大的物理理解。在事件驱动和环境驱动的转换中,性能会急剧下降,甚至强大的专有系统也会在Anti-AV-Physics提示下崩溃。我们进一步介绍了AV-Phys Agent,这是一种ReAct风格的评估器,它将多模态语言模型与确定性声学测量工具相结合,产生与人类评级密切相关的排名。我们的研究结果确定跨模态的物理一致性和过渡驱动的场景动态的关键开放的挑战,联合音视频生成。摘要:Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visual physics, or merely generate plausible sounds and frames that violate real-world consistency? We introduce AV-Phys Bench, a benchmark for evaluating physical commonsense in joint audio-video generation. AV-Phys Bench tests models across three scene categories: Steady State, Event Transition, and Environment Transition. It covers physics-grounded subcategories drawn from real-world scenes, plus Anti-AV-Physics prompts that deliberately request physically inconsistent audio-video behavior. Each generation is evaluated along five dimensions: visual semantic adherence, audio semantic adherence, visual physical commonsense, audio physical commonsense, and cross-modal physical commonsense. Across three proprietary and four open-source models, we find that Seedance 2.0 performs best overall, but all models remain far from robust physical understanding. Performance drops sharply on event-driven and environment-driven transitions, and even strong proprietary systems collapse on Anti-AV-Physics prompts. We further introduce AV-Phys Agent, a ReAct-style evaluator that combines a multimodal language model with deterministic acoustic measurement tools, producing rankings that closely align with human ratings. Our results identify cross-modal physical consistency and transition-driven scene dynamics as key open challenges for joint audio-video generation.
【5】MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
标题:MIST:智能家居的多模式交互式基于语音的工具呼叫对话助理【6】An audio-to-analysis pipeline with certified transcription for information-theoretic profiling of the piano repertoire
标题:具有经过认证的转录的音频到分析管道,用于钢琴曲目的信息理论分析机器翻译由腾讯交互翻译提供,仅供参考
