5月21日,为了促进产学研的语音与语言处理技术交流,洞察未来技术的创新趋势,香港中文大学(深圳)数据科学学院与语音之家社区共同举办了一场语音和语言处理交流会。


本次活动参会嘉宾包括香港中文大学(深圳)数据科学学院执行院长李海洲教授,以及深圳市大数据研究院、香港科技大学、香港科技大学(广州)、腾讯、华为、大象声科、优必选、快手、声扬科技等多所高校及企业的学者和专家。


上午,是香港中文大学(深圳)的课程项目汇报,交流环节(CSC3160/MDS6002)。

查看详情→  https://slpcourse.github.io/


上午的课程项目活动,由香港中文大学(深圳)数据科学学院武执政教授主持,邀请李海洲院长进行会议致辞。李海洲院长简要说明了自然语言处理近几年发展迅速并且在人们的日常生活中运用十分广泛。


随后,武执政教授就本次CSC3160/MDS6002课程的学生海报交流活动的注意事项及专家评分标准进行了阐述说明。参会的嘉宾们与课程的部分同学们在香港中文大学(深圳)数据科学学院道远楼一楼进行了海报交流活动。


在海报展示环节学生们充满热情地介绍自己的项目,大家的项目主题也十分有趣,紧跟最新最热门的研究方向,不仅有相关ChatGPT的伪造文本检测项目,还有探索最近火遍网络的“AI”孙燕姿背后核心技术项目等等。很多同学表示,与来自不同背景和领域的专家和同学们沟通交流后,获得了很多新的观点和见解。


在交流讨论以及打分环节结束后,武执政教授简要总结了本次活动并对本次获奖项目进行了颁奖。根据学生们报告和海报表现,分别颁发了Best Report, Rising Star, Best Presentation等奖项。随着颁奖环节的结束,早上的海报展示环节也落下帷幕。



本次课程项目海报分享活动中,同学们充分的展现了自信和实力,积极与业界专家交流,同时部分同学也获得了业界实习的机会。活动现场有同学表达了自己的切身体会:“这次语音与语言处理技术交流会是一次独特而宝贵的经历。通过与专家的交流和学习,我不仅学到了最新的技术知识,也结识了许多志同道合的人。这次活动让我更加坚定了在语音与语言处理领域深耕的决心,我期待着将所学应用到未来的项目中,并继续与这个热情洋溢的社区保持联系。”



学术报告

下午,邀请到来自腾讯、香港科技大学、香港中文大学(深圳)和优必选的专家、学者进行学术报告。


罗艺

本报告介绍腾讯AI Lab音频与语音前端处理团队在音频分离、语音增强、多通道语音处理等方向的研究进展,包括腾讯AI Lab在数据仿真、模型设计、应用场景等方面的探索。

雪巍

We are entering a new era in which the real and virtual worlds are indistinguishable; interactions between the real and virtual worlds remove the physical barriers between people and define new ways of entertainment, healthcare, and communication. Building a new generation of content generation and interaction over human, machine and environment in terms of audio is essential. We will introduce our progresses in recent months in this talk. Specifically, we will introduce how to digitalize the voice of an arbitrary person to produce the virtual singer, which empowers the AI choir in the world’s first human-machine collaborative symphony orchestra at Hong Kong; We will also introduce CoMoSpeech, which adopts the consistency model for speech synthesis and achieves an inference speed more than 150 times faster than real-time on a single NVIDIA A100 GPU, making diffusion-sampling based speech synthesis truly practical.


王远程
Audio editing is applicable for various purposes, such as adding background sound effects, replacing a musical instrument, and repairing damaged audio. Recently, some diffusion-based methods achieved zero-shot audio editing by using a diffusion and denoising process conditioned on the text description of the output audio. However, these methods still have some problems: 1) they have not been trained on editing tasks and cannot ensure good editing effects; 2) they can erroneously modify audio segments that do not require editing; 3) they need a complete description of the output audio, which is not always available or necessary in practical scenarios. In this work, we propose AUDIT, an instruction-guided audio editing model based on latent diffusion models. Specifically, AUDIT has three main design features: 1) we construct triplet training data (instruction, input audio, output audio) for different audio editing tasks and train a diffusion model using instruction and input (to be edited) audio as conditions and generating output (edited) audio; 2) it can automatically learn to only modify segments that need to be edited by comparing the difference between the input and output audio; 3) it only needs edit instructions instead of full target audio descriptions as text input. AUDIT achieves state-of-the-art results in both objective and subjective metrics for several audio editing tasks (e.g., adding, dropping, replacement, inpainting, super-resolution).

王燕南
Real-time communication (RTC) systems become a necessity in the life and work of individuals, especially in teleconferencing systems. Speech quality is the key element for communication experience. However various problems degrades the speech quality including acoustical capturing, noise/reverberation corruption, bad device acquisition performance and network congestion, etc. In this talk I would like to present our attempts to promote speech signal, especially in far-field scenario when the environment is complex with lower SNR. In the future we would like to devote more effort in more types of speech quality improvement task.


丁万

人形机器人产品需要通过多模态信息来实现准确地感知和表达。相比传统方法,基于深度学习的多模态识别和合成能够达到更好的效果,但是在落地时仍然需要注意过拟合、实时性等问题。本次报告向大家介绍优必选在多模态机器学习方面进行的一些工作。具体包括多模态情感识别、多模态抑郁症检测和2D数字人合成等。

最后,武执政教授对本次会议进行了总结和鸣谢。至此,语音与语言处理技术交流会(深圳)圆满结束,期待与大家再次相见。

嘉宾PPT

语音与语言处理技术交流会(深圳)嘉宾PPT:
链接:https://pan.baidu.com/s/12neYoutAiYFH08yMAzquNw?pwd=w2o9
提取码:w2o9