datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-corpus-sample-open-webspai-ss6-corpus-medical-health-web
SPAI SS6 Thai Medical Health Web Corpus
Thai public medical and health web articles collected by the local scraping pipeline.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: default
Rows in canonical config: 3,660
Parquet size in canonical config: 0.01 GB
Source license: other… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-health-web.georgian-web-corpus-nlp4nlp-georgian-web-corpusnlp4-georgian-web-corpusspai-ss6-corpus-wangchanlion-web
SPAI SS6 WangchanLION Web Corpus Index
Index repo for the WangchanLION-Web corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: wangchanlion_web
Rows in canonical config: 557,502
Parquet size in canonical config: 1.97 GB
Source license: odc-by… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-wangchanlion-web.
