CoolFace
Datasetpublic

atomkwk/srt-SpokenCantoneseToWrittenChinese

#Introduction to this dataset This data set is for training llm to translate spoken cantonese srt to written chinese srt(Words in this set is written as simplified chinese characters). Each input and output contain a group of 10 sentances, with a next line character \n between each sentance.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes1downloads
Dataset Card

#Introduction to this dataset

This data set is for training llm to translate spoken cantonese srt to written chinese srt(Words in this set is written as simplified chinese characters). Each input and output contain a group of 10 sentances, with a next line character \n between each sentance.