CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jpaulpoliquit /ph-pretrain-03 PH Pretrain 03 — Filipino/Tagalog Web Corpus (ph-pretrain-03) The recommended dataset for Filipino / Tagalog language-model pretraining and continued pretraining (CPT). A ~1-billion-token, Tagalog-forward web corpus, cleaned, deduplicated, and quality-scored by the jpaulpoliquit/pretraining refinery. 1,933,988 documents · ≈ 1.0 billion tokens (Sailor2-1B, calibrated) ~97.7% Tagalog web (FineWeb2 + CC100), with regional + news + Wikipedia in the long tail ~3.41 GB to download… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain-03.tabulartext-generation1M<n<10M0 likes232 downloads4mo agoHugging Face02jpaulpoliquit /ph-pretrain PH Pretrain — Philippine Languages Corpus (v0.6-ph-unified) 👉 Looking to train a Filipino/Tagalog model? Use jpaulpoliquit/ph-pretrain-03 instead — it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage. A cleaned, deduplicated, document-level pretraining corpus for… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain.tabulartext-generation1M<n<10M0 likes153 downloads4mo agoHugging Face03DaniilOr /php_cat1tabular10K<n<100K0 likes87 downloads1y agoHugging Face04KaiLv /UDR_PHP Dataset Card for "UDR_PHP" More Information needed tabular100K<n<1M0 likes59 downloads3y agoHugging Face05beranki /gpt-5-mini-rebench-v2-phptabularn<1K0 likes40 downloads4mo agoHugging Face06hongliu9903 /stack_edu_phptabular1M<n<10M0 likes11 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.