CoolFace
Datasetpublic

IKMLab-team/hk_content_corpus

HK Content Corpus (Cantonese & Traditional Chinese) This dataset contains eight cleaned source-specific corpora of Hong Kong Cantonese and Traditional Chinese text, crawled from public websites and platforms. It was initially created for the experiments reported in https://doi.org/10.1145/3744341 which study the effect of diglossia on Hong Kong language modeling. Each file stores plain UTF-8 text, where each record occupies one line, and blank lines serve as separators.… See the full description on the dataset page: https://huggingface.co/datasets/IKMLab-team/hk_content_corpus.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes93downloads

IKMLab-team/hk_content_corpus · main · files are served by the source, never re-hosted here