CoolFace
Datasetpublic

sapinsapin/halohalo

halohalo Dataset Summary halohalo is a Pretraining text corpus for Philippine languages, assembled from web-scraped data. It is compatible with Fineweb for LLM Pretraining. Source Data Derived from the following cleaned datasets: Source Documents halo-hil 8,874 halo-tgl 6,589 halo-bcl 1,264 Each source dataset was cleaned using clean_halo.py to remove web boilerplate, navigation menus, markdown noise, HTML artifacts, and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halohalo.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes13downloads

sapinsapin/halohalo · main · files are served by the source, never re-hosted here