Xmart青年论坛由上海交通大学跨媒体语言智能实验室(SJTU X-LANCE lab)创办,中国计算机学会语音对话专委会主办,语音之家协办,旨在邀请国内外优秀的青年学者分享其最新科研工作和成果,促进多元且深入的交流与合作。Xmart学生论坛作为其中一个系列,致力于邀请国内外知名高校有成体系工作的研究生,主要通过线上分享的方式,系统地介绍其科研成果和心得,为青年学生打造一个学术探讨,思维碰撞和多学科交叉融合的平台。


Xmart•学生论坛丨杨东超:面向多任务的音频基座模型:音频生成视角

形式:线上

时间:6月6日(周五) 14:00 ~ 16:00


报告嘉宾

杨东超,香港中文大学在读博士生,导师为Helen Meng教授,他的研究方向为多任务音频基础模型,致力于在统一框架下整合speech、music及audio等多种音频处理任务。他最近工作关注于音频离散化(tokenization)、层次化建模,以及与大型语言模型的跨模态融合。


报告摘要

In this talk, Dongchao Yang will present his recent progress in building multi-task audio foundation models—unified systems that can handle a wide range of audio tasks, from speech-related task to music and sound generation. Inspired by large language models, Yang's team's approach treats audio as a sequence of discrete tokens, enabling the use of transformer-based models for both understanding and generation. He will first introduce their work for building a low bit-rate and semantic-rich audio tokenizer. In this work, they present a query-based compression paradigm for audio codec model. Then he will introduce their hierarchical modeling strategy to effective model long audio sequences in one stage, and show the effectiveness of building a multi-task audio foundation model. Lastly, he will present a cross-modality audio tokenizer that bridges audio and text, allowing few-shot learning across tasks. These contributions mark a step forward toward universal audio intelligence.


参加方式

①

直播将通过语音之家微信视频号进行直播

手机端、PC端可同步观看

👇👇👇


②

腾讯会议参加

会议号:661-917-700