Voice Navigation collects more than 830,000 place names in China, such as “故宫”(The Palace Museum), “八达岭长城”(Great Wall on Badaling), “积水潭医院”(Jishuitan Hospital) etc. To generate the navigation queries, Voice Navigation also collects more than 25 query patterns, and fills out the query pattern with places to generate the query.
Voice Navigation uses a TTS tool to generate the speech file by reading the generated query. The TTS tool provides several variant articulation types, such as men, women, background music, echo and underwater etc. In addition, to build a real-person testing data, for the testing data Voice Navigation also has human-read data which contains 200 cases read by 5 persons and TTS data generated by TTS.
The whole dataset contains 820,000 speech-slot pairs for training and 12,000 speech-slot pairs for testing, and half of the slots in testing dataset do not appear in training datase.
