CoolFace
Datasetpublic

saidutta69/Odia-Web-Corpus-v1

Odia Web Corpus v1 The inaugural release of a curated Odia (Oriya) web text corpus. Scraped and filtered from publicly accessible web sources to support Odia language modeling and NLP research. Dataset Details Language: Odia (Oriya, ISO 639-3: ory) Format: JSONL (one JSON object per line) Size: ~650K documents, ~0.9 GB text License: CC-BY-4.0 Data Fields Field Type Description text string Cleaned document body title string… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v1.

sourceHugging Facecc-by-4.0updated 14d agoView on Hugging Face
0likes131downloads
settings

This repository belongs to saidutta69 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameOdia-Web-Corpus-v1
visibilitypublic
licencecc-by-4.0
gatedno
ownersaidutta69
Account settings
saidutta69/Odia-Web-Corpus-v1 · CoolFace