datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KhanomTanLLM-pretrained-dataset
KhanomTanLLM pretrained dataset
This daataset collect all raw text for pretraining LLM.
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Tokens
53,376,211,711 Tokens
English: 31,629,984,243 Tokens
Thai: 12,785,565,497 Tokens
Code: 8,913,084,300 Toekns
Parallel data: 190,310,686 Tokens
Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer
All subset
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.continue-pretrained-v1
Continue Pretrained v1
Continual-pretraining (CPT) mixture shards for Vietnamese LLM training.
Splits
Split
Description
Rows (approx)
Est. tokens
stage_1
Warmup / general mix (VI-heavy + EN replay)
56,419,797
~52.2B
Schema
id, text, source, subset
stage, stage_name, mix_source, language, epoch, quality_pred
Load
from datasets import load_dataset
ds = load_dataset("brownyeyez/continue-pretrained-v1", split="stage_1")
KhanomTanLLM-pretrained-dataset-thai-subset
KhanomTanLLM pretrained dataset (Thai subset)
This daataset collect all raw text for pretraining LLM. (Thai subset)
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0
pythainlp/thai-tnhc2-books
pythainlp/thai-constitution-corpus
pythainlp/thai-it-books
pythainlp/prd_news_3011202
pythainlp/thailand-policy-statements
pythainlp/thai-cc-license
pythainlp/blognone_news
pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.
