CoolFace
Datasetpublic

BBSRguy/ODEN-Indcorpus

ODEN‑Indcorpus 📚 ODEN‑Indcorpus is a 3.7‑million‑line Odia mixed text collection curated from fiction, dialogue, encyclopaedia, Q‑A and community writing derived from the ODEN initiative.After thorough normalisation and de‑duplication it serves as a robust substrate for training Odia‑centric tokenizers, language models and embedding spaces. Split Lines Train 3,373,817 Validation 187,434 Test 187,435 Total 3,748,686 The material ranges from conversational… See the full description on the dataset page: https://huggingface.co/datasets/BBSRguy/ODEN-Indcorpus.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes131downloads
6 commits on main
7eef7a81y ago

ODENMain_corpus.txt

BBSRguy
00112c11y ago

fixed tokenizers spelling

BBSRguy
562869e1y ago

Add autogenerated README

BBSRguy
7683e291y ago

Upload dataset (part 00001-of-00002)

BBSRguy
0d525ed1y ago

Upload dataset (part 00000-of-00002)

BBSRguy
0dedc491y ago

initial commit

BBSRguy