CoolFace
Datasetpublic

sapinsapin/halohalo

halohalo Dataset Summary halohalo is a Pretraining text corpus for Philippine languages, assembled from web-scraped data. It is compatible with Fineweb for LLM Pretraining. Source Data Derived from the following cleaned datasets: Source Documents halo-hil 8,874 halo-tgl 6,589 halo-bcl 1,264 Each source dataset was cleaned using clean_halo.py to remove web boilerplate, navigation menus, markdown noise, HTML artifacts, and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halohalo.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes13downloads
Dataset Card

halohalo

Dataset Summary

halohalo is a Pretraining text corpus for Philippine languages, assembled from web-scraped data. It is compatible with Fineweb for LLM Pretraining.

Source Data

Derived from the following cleaned datasets:

SourceDocuments
halo-hil8,874
halo-tgl6,589
halo-bcl1,264

Each source dataset was cleaned using clean_halo.py to remove web boilerplate, navigation menus, markdown noise, HTML artifacts, and low-quality documents before being included here.

Processing

  1. 1.Cleaning (clean_halo.py) — strips boilerplate, HTML, markdown noise; filters documents with fewer than 30 words or less than 40% Latin characters
  2. 2.FineWeb formatting (prep_halohalo.py) — adds source, language, token_count, content_hash; deduplicates against existing documents using MD5 content hashing

Processing code is available at github.com/sapinsapin/halohalo.

Statistics

MetricValue
Total documents16,727
Total tokens19,178,582
Avg tokens per document1,146.6
Min tokens30
Max tokens10,552

Languages

LanguageDocumentsWord Count
hil8,8749,332,784
tgl6,5898,208,749
bcl1,2641,637,049
Total16,72719,178,582

Schema

FieldTypeDescription
textstrCleaned document text
idstrUnique document identifier
sourcestrSource dataset name
languagestrISO 639-3 language code
token_countintWhitespace-tokenized word count
content_hashstrMD5 hash of text for deduplication
urlstrSource URL
datestrCrawl date
dumpstrCommonCrawl dump identifier
titlestrPage title

Usage

python
from datasets import load_dataset

ds = load_dataset("sapinsapin/halohalo")
print(ds["train"][0])