datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ph-pretrain-03
PH Pretrain 03 — Filipino/Tagalog Web Corpus (ph-pretrain-03)
The recommended dataset for Filipino / Tagalog language-model pretraining and continued pretraining (CPT).
A ~1-billion-token, Tagalog-forward web corpus, cleaned, deduplicated, and quality-scored by the jpaulpoliquit/pretraining refinery.
1,933,988 documents · ≈ 1.0 billion tokens (Sailor2-1B, calibrated)
~97.7% Tagalog web (FineWeb2 + CC100), with regional + news + Wikipedia in the long tail
~3.41 GB to download… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain-03.ph-pretrain
PH Pretrain — Philippine Languages Corpus (v0.6-ph-unified)
👉 Looking to train a Filipino/Tagalog model? Use jpaulpoliquit/ph-pretrain-03 instead — it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage.
A cleaned, deduplicated, document-level pretraining corpus for… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain.php_cat1UDR_PHP
Dataset Card for "UDR_PHP"
More Information needed
gpt-5-mini-rebench-v2-phpstack_edu_php
