CHiME 挑战赛已经正式开启,今天分享下 CHiME 的子任务MMCSG(智能眼镜多模态对话),欢迎大家投稿报名!
赛事官网:https://www.chimechallenge.org/current/task3/index

Rules
For building the system, it is allowed to use the training subset of MMCSG dataset and external data listed in the subsection Data and pre-trained models. If you believe there is some public dataset missing, you can propose it to be added until the deadline as specified in the schedule. The development subset of MMCSG can be used for evaluating the system throughout the challenge period, but not for training or automatic tuning of the systems. Pre-trained models are listed in the “Data and pre-trained models” subsection. Only those pre-trained models are allowed to be used. If you believe there is some model missing, you can propose it to be added until the deadline as specified in the schedule. The submitted systems must be streaming, i.e. process its inputs sequentially in time and specify latency for each emitted word, as described in detail in the subsection Evaluation. It must not use any global information from a recording before processing it in temporal order. Such global information could include global normalizations, non-streaming speaker identification or diarization, etc. This requirement on streaming processing applies to all modalities (audio, visual, accelerometer, gyroscope, etc). The details of the streaming nature of the system, including any lookahead, chunk-based processing, other details that would impact latency, and an explicit estimate of the average algorithmic and emission latency itself should be clearly described in a section of the submitted system description with the heading “Latency”. For evaluation, each recording must be considered separately. The system should not be in any way fine-tuned on the entire evaluation set (e.g. by computing global statistics, gathering speaker information across multiple recordings). If your system does not comply with these rules (e.g. by using a private dataset or having only a partially streaming method), you may still submit your system, but we will not include it in the final rankings.
Baseline System
Fixed NLCMV beamformer (Feng et al, 2023) which uses 13 beams into 12 directions uniformly spaced around the wearer + 1 direction for the mouth of the wearer. The beamformer coefficients are derived from acoustic transfer functions (ATF) recorded in anechoic rooms with the Aria glasses. We release both the beamforming coefficients and the original ATFs. Extraction of log-mel features from each of the 13 beams. ASR model processing the multi-channel features and estimating serialized-output-training (SOT) transcriptions.(Kanda et al, 2022)


Submission
the word error rates (including the break-down into substitutions, insertions, deletions and speaker attributions) for SELF and OTHER on dev and eval subsets the computed mean latency on dev and eval subsets the hypotheses files for each recording of dev and eval subsets the hypotheses files from the timestamp test on perturbed and unperturbed files
Important dates
| February 15th, ‘24 | Challenge begins; release of train and dev datasets and baseline system |
| March 20th, ‘24 | Deadline for proposing additional public datasets and pre-trained models |
| June 15th, ‘24 | Evaluation set released |
| June 28th, ‘24 | Deadline for submitting results |
| June 28th, ‘24 | Deadline for submitting technical reports |
| September 6th, ‘24 | CHiME-8 Workshop |
