CoolFace
Datasetpublic

IKMLab-team/hk_content_corpus

HK Content Corpus (Cantonese & Traditional Chinese) This dataset contains eight cleaned source-specific corpora of Hong Kong Cantonese and Traditional Chinese text, crawled from public websites and platforms. It was initially created for the experiments reported in https://doi.org/10.1145/3744341 which study the effect of diglossia on Hong Kong language modeling. Each file stores plain UTF-8 text, where each record occupies one line, and blank lines serve as separators.… See the full description on the dataset page: https://huggingface.co/datasets/IKMLab-team/hk_content_corpus.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes93downloads
settings

This repository belongs to IKMLab-team on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namehk_content_corpus
visibilitypublic
licencecc-by-4.0
gatedno
ownerIKMLab-team
Account settings
IKMLab-team/hk_content_corpus · CoolFace