CoolFace
Datasetpublic

CAiRE/ASCEND

Dataset Card for ASCEND Dataset Summary ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/CAiRE/ASCEND.

sourceHugging Facecc-by-sa-4.0updated 2y agoView on Hugging Face
53likes1.6kdownloads

CAiRE/ASCEND · main · files are served by the source, never re-hosted here