sapinsapin/halohalo
halohalo Dataset Summary halohalo is a Pretraining text corpus for Philippine languages, assembled from web-scraped data. It is compatible with Fineweb for LLM Pretraining. Source Data Derived from the following cleaned datasets: Source Documents halo-hil 8,874 halo-tgl 6,589 halo-bcl 1,264 Each source dataset was cleaned using clean_halo.py to remove web boilerplate, navigation menus, markdown noise, HTML artifacts, and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halohalo.
halohalo
Dataset Summary
halohalo is a Pretraining text corpus for Philippine languages, assembled from web-scraped data. It is compatible with Fineweb for LLM Pretraining.
Source Data
Derived from the following cleaned datasets:
Each source dataset was cleaned using clean_halo.py to remove web boilerplate, navigation menus, markdown noise, HTML artifacts, and low-quality documents before being included here.
Processing
- Cleaning (
clean_halo.py) — strips boilerplate, HTML, markdown noise; filters documents with fewer than 30 words or less than 40% Latin characters - FineWeb formatting (
prep_halohalo.py) — addssource,language,token_count,content_hash; deduplicates against existing documents using MD5 content hashing
Processing code is available at github.com/sapinsapin/halohalo.
Statistics
Languages
Schema
Usage
from datasets import load_dataset
ds = load_dataset("sapinsapin/halohalo")
print(ds["train"][0])