CoolFace
Datasetpublic

nis12ram/nisram-hindi-text-0.0

Dataset Card for nisram-hindi-text-0.0 Quick Summary Language: Hindi Size: ~618,000 raw text samples (206k from each source, before deduplication) and 602,000 after deduplication Content: Short- to medium-length Hindi text (up to ~5000 characters), drawn from web-crawled corpora Fields: One field text (string) containing the Hindi text Sources: mC4-Hindi-Cleaned-3.0 OSCAR-2301-Hindi-Cleaned-2.0 ai4bharat/sangraha (Hindi, verified) Deduplication:… See the full description on the dataset page: https://huggingface.co/datasets/nis12ram/nisram-hindi-text-0.0.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes65downloads
5 commits on main
d10a2f41y ago

Update README.md

nis12ram
25773841y ago

Update README.md

nis12ram
4e10f1b1y ago

Update README.md

nis12ram
1e1ce751y ago

Upload dataset

nis12ram
5fc8ba71y ago

initial commit

nis12ram