GOLOS

Introduced by Karpov et al. inGolos: Russian Dataset for Speech Research

Golosis a Russian speech dataset suitable for speech research. The dataset mainly consists of recorded audio files manually annotated on the crowd-sourcing platform. The total duration of the audio is about 1240 hours.

Dataset structure

DomainTrain filesTrain hoursTest filesTest hours
Crowd979 7961 0959 99411.2
Farfield124 003132.41 9161.4
Total1 103 7991 227.411 91012.6

Audio files in opus format

ArchiveSizeLink
golos_opus.tar20.5 GBhttps://sc.link/JpD

Audio files in wav format

ArchivesSizeLinks
train_farfield.tar15.4 GBhttps://sc.link/1Z3
train_crowd0.tar11 GBhttps://sc.link/Lrg
train_crowd1.tar14 GBhttps://sc.link/MvQ
train_crowd2.tar13.2 GBhttps://sc.link/NwL
train_crowd3.tar11.6 GBhttps://sc.link/Oxg
train_crowd4.tar15.8 GBhttps://sc.link/Pyz
train_crowd5.tar13.1 GBhttps://sc.link/Qz7
train_crowd6.tar15.7 GBhttps://sc.link/RAL
train_crowd7.tar12.7 GBhttps://sc.link/VG5
train_crowd8.tar12.2 GBhttps://sc.link/WJW
train_crowd9.tar8.08 GBhttps://sc.link/XKk
test.tar1.3 GBhttps://sc.link/Kqr

Evaluation

Percents of Word Error Rate for different test sets

Decoder \ Test setCrowd testFarfield testMCV1 devMCV1 test
Greedy decoder4.389 %14.949 %9.314 %11.278 %
Beam Search with Common Crawl LM4.709 %12.503 %6.341 %7.976 %
Beam Search with Golos train set LM3.548 %12.384 %--
Beam Search with Common Crawl and Golos LM3.318 %11.488 %6.4 %8.06 %