CoolFace
Datasetpublic

sapinsapin/halohalo

halohalo Dataset Summary halohalo is a Pretraining text corpus for Philippine languages, assembled from web-scraped data. It is compatible with Fineweb for LLM Pretraining. Source Data Derived from the following cleaned datasets: Source Documents halo-hil 8,874 halo-tgl 6,589 halo-bcl 1,264 Each source dataset was cleaned using clean_halo.py to remove web boilerplate, navigation menus, markdown noise, HTML artifacts, and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halohalo.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes13downloads
settings

This repository belongs to sapinsapin on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namehalohalo
visibilitypublic
licencenot set
gatedno
ownersapinsapin
Account settings
sapinsapin/halohalo · CoolFace