Odia
Datasets
All datasets matching “Odia”TTS_ODIASPRING_INX_Odia_R1Odia-Web-Corpus-v5
Odia Web Corpus v5
The largest and cleanest Odia text corpus to date — 4.16 million deduplicated documents, 7.74 GB. Built by merging and thoroughly cleaning four source collections.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: 28 sharded Parquet files
Total Size: 7.74 GB
Total Documents: 4,162,804
License: CC-BY-SA-4.0
Cleaning Pipeline
Stage
Removed
Description
Deduplication
30.2%
Exact MD5 hash match
Short lines… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v5.odia_pretrain_dataset_v2
Odia Pretrain Dataset v2
12.3 million clean Odia documents. 30+ sources. Zero noise. Zero gating. Just data that actually works for pretraining.
The Origin Story (a.k.a. The Time We Filtered 98% of a Clean Dataset and Learned Our Lesson)
v1 was built from spite. v2 was built from more data.
We took v1 (25+ sources, 10.3M rows) as the foundation and asked: what else can we add?
monsoon-nlp/odia-cleaned -- fully deduplicated against v1 (0 new rows)… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/odia_pretrain_dataset_v2.HINDI-BENGALI-MALAYALAM-ODIA-VISUAL-GENOMEodia_tokenizer_text
