CoolFace
Datasetpublic

saidutta69/Odia-Web-Corpus-v5

Odia Web Corpus v5 The largest and cleanest Odia text corpus to date — 4.16 million deduplicated documents, 7.74 GB. Built by merging and thoroughly cleaning four source collections. Dataset Details Language: Odia (ISO 639-3: or) Format: 28 sharded Parquet files Total Size: 7.74 GB Total Documents: 4,162,804 License: CC-BY-SA-4.0 Cleaning Pipeline Stage Removed Description Deduplication 30.2% Exact MD5 hash match Short lines… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v5.

sourceHugging Facecc-by-sa-4.0updated 14d agoView on Hugging Face
0likes291downloads
settings

This repository belongs to saidutta69 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameOdia-Web-Corpus-v5
visibilitypublic
licencecc-by-sa-4.0
gatedno
ownersaidutta69
Account settings
saidutta69/Odia-Web-Corpus-v5 · CoolFace