pythainlp/thai-culturax-clean-dataset
Thai CulturaX Clean dataset The data is sourced from the Thai subset of CulturaX dataset, which itself is sourced from mC4 and four OSCAR corpora. It has about 8,748,575,684 words (without whitespace) and 16,768,585 lines (97 GB). It was filtered content promoting gambling, adult content, and narcotics. GitHub for clean: https://github.com/wannaphong/thai-filter-website Considerations for Using the Data This dataset is the cleaned version of the CulturaX… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-culturax-clean-dataset.
5255
