来源丨阿里语音AI

Themeetingscenario is one of the most valuable and, at the same time, most challenging scenarios for speech technologies. Because such scenarios have free speaking styles and complex acoustic conditions such as overlapping speech, unknown number of speakers, far-field signals in large conference rooms, noise and reverberation etc.


However, thelack of large public real meeting datahas been a major obstacle for advancement of the field.


Since meeting transcription involves numerous related processing components, more information has to be carefully collected and labelled, such as speaker identity, speech context, onset/offset time, etc. All these information require precise and accurate annotations, which is expensive and time-consuming. Shell Shell Technology Co., Ltd., in conjunction with several organizations, opened upAISHELL-4 data set (https://arxiv.org/abs/2104.03603)in the early stage to further facilitate the research in meeting transcription.


Alibaba DAMO Academy Speech Labwill jointly launch theMulti-channel Multi-party Meeting Transcription Challenge (M2MeT)withShell Shell Technology Co., Ltd. and a number of internationally influential industry experts, as anICASSP2022 Signal Processing Grand Challenge. Meanwhile, Alibaba release theAliMeeting corpus, which consists of 120 hours of real recorded Mandarin meeting data, including the far-field data collected by the 8-channel microphone array and the near-field data collected by each participant's headset microphone. Its content covers a variety of aspects in real-world meeting, including diverse recording conditions, various number of meeting participants, various overlap ratios and noise types, etc. High-quality transcriptions on multiple aspects are provided for each meeting, allowing the researcher to explore different aspects of meeting processing.


In addition, AliMeeting involves different speakers and meeting places, and adds discussion scenarios with a high overlap ratio to multi-speaker meetings, which has great value in research and industry implementation perspectives.


Based on the AliMeeting and AISHELL-4 corpus, our M2MeT challenge consists of two tracks, namelyspeaker diarization and multi-speaker ASR. Meanwhile, the organizers will also provide code of the baseline system for the two tracks as a reference.


The teams with the high scores and innovative work in two tracks have the opportunity to write their system descriptionpapersandbe accepted by the ICASSP2022. At the same time,the top threewinning teams from sub-track I (fixed training condition) of each trackwill be awarded prizes provided by Alibaba.




The challenge has already started to register, and the registrationdeadline is November 17th. Researchers in academia and industry are welcome to register! For more information about the registration method, see the challenge website below.


Challenge website:

https://www.alibabacloud.com/m2met-alimeeting


Challenge andAliMeetingcorpus introduction paper:

https://arxiv.org/abs/2110.07393


Baseline system:

https://github.com/yufan-aslp/AliMeeting