CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jpaulpoliquit /ph-pretrain-03 PH Pretrain 03 — Filipino/Tagalog Web Corpus (ph-pretrain-03) The recommended dataset for Filipino / Tagalog language-model pretraining and continued pretraining (CPT). A ~1-billion-token, Tagalog-forward web corpus, cleaned, deduplicated, and quality-scored by the jpaulpoliquit/pretraining refinery. 1,933,988 documents · ≈ 1.0 billion tokens (Sailor2-1B, calibrated) ~97.7% Tagalog web (FineWeb2 + CC100), with regional + news + Wikipedia in the long tail ~3.41 GB to download… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain-03.tabulartext-generation1M<n<10M0 likes242 downloads4mo agoHugging Face02jpaulpoliquit /ph-pretrain PH Pretrain — Philippine Languages Corpus (v0.6-ph-unified) 👉 Looking to train a Filipino/Tagalog model? Use jpaulpoliquit/ph-pretrain-03 instead — it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage. A cleaned, deduplicated, document-level pretraining corpus for… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain.tabulartext-generation1M<n<10M0 likes152 downloads4mo agoHugging Face03Nan-Do /code-search-net-php Dataset Card for "code-search-net-php" Dataset Summary This dataset is the Php portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in Php Data Splits Train, test, validation labels are included in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-php.texttext-generation100K<n<1M1 likes99 downloads3y agoHugging Face04Nan-Do /instructional_code-search-net-php Dataset Card for "instructional_code-search-net-php" Dataset Summary This is an instructional dataset for PHP. The dataset contains two different kind of tasks: Given a piece of code generate a description of what it does. Given a description generate a piece of code that fulfils the description. Languages The dataset is in English. Data Splits There are no splits. Dataset Creation May of 2023 Curation Rationale This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-php.texttext-generation100K<n<1M3 likes71 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.