aisingapore/WangchanLION-Curated
Citation @misc{phatthiyaphaibun2025mangosteenopenthaicorpus, title={Mangosteen: An Open Thai Corpus for Language Model Pretraining}, author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong}, year={2025}, eprint={2507.14664}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2507.14664}, }… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Curated.
Citation
@misc{phatthiyaphaibun2025mangosteenopenthaicorpus,
title={Mangosteen: An Open Thai Corpus for Language Model Pretraining},
author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong},
year={2025},
eprint={2507.14664},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2507.14664},
}We have collected additional Thai text that is unlikely to be included in the common crawl from various sources. The total number of documents collected is as follows:425,304 documents, we deduplication these noncc documents to later divide them into train sets for further web data and validation set.
The following table shows the analysis of the data after deduplication.
Documents from some sources cannot be directly used or easily processed. It is necessary to use text extraction technology (OCR) to extract the text due to the document format in PDF. The table below shows the number and proportion of documents that required OCR.
The model used for text extraction ishttps://github.com/VikParuchuri/marker In the future, you should try VLM, such as:https://olmocr.allenai.org/Or Typhoon2 Vision
Data from various sources can be classified into 6 types as shown in the table below.
The additional documents we collect are confirmed to be open source and have a license to allow for redistribution, with the copyright share as shown in the table below.
Resources
- Pre-training data (web): https://huggingface.co/datasets/aisingapore/WangchanLION-Web
- Pre-training data (curated): https://huggingface.co/datasets/aisingapore/WangchanLION-Curated
- Pre-training model: https://huggingface.co/aisingapore/WangchanLION-v3
- SFT model: https://huggingface.co/aisingapore/WangchanLION-v3-IT
- Paper: https://arxiv.org/abs/2507.14664
- Blog: https://sea-lion.ai/sea-lion-wangchanlionv3/
- Github: https://github.com/vistec-AI/Mangosteen
