CoolFace
Datasetpublic

saidutta69/odia_pretrain_dataset_v2

Odia Pretrain Dataset v2 12.3 million clean Odia documents. 30+ sources. Zero noise. Zero gating. Just data that actually works for pretraining. The Origin Story (a.k.a. The Time We Filtered 98% of a Clean Dataset and Learned Our Lesson) v1 was built from spite. v2 was built from more data. We took v1 (25+ sources, 10.3M rows) as the foundation and asked: what else can we add? monsoon-nlp/odia-cleaned -- fully deduplicated against v1 (0 new rows)… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/odia_pretrain_dataset_v2.

sourceHugging Facecc-by-4.0updated 14d agoView on Hugging Face
0likes207downloads
settings

This repository belongs to saidutta69 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameodia_pretrain_dataset_v2
visibilitypublic
licencecc-by-4.0
gatedno
ownersaidutta69
Account settings
saidutta69/odia_pretrain_dataset_v2 · CoolFace