sapinsapin/halohalo
halohalo Dataset Summary halohalo is a Pretraining text corpus for Philippine languages, assembled from web-scraped data. It is compatible with Fineweb for LLM Pretraining. Source Data Derived from the following cleaned datasets: Source Documents halo-hil 8,874 halo-tgl 6,589 halo-bcl 1,264 Each source dataset was cleaned using clean_halo.py to remove web boilerplate, navigation menus, markdown noise, HTML artifacts, and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halohalo.
013
