nis12ram/nisram-hindi-text-0.0
Dataset Card for nisram-hindi-text-0.0 Quick Summary Language: Hindi Size: ~618,000 raw text samples (206k from each source, before deduplication) and 602,000 after deduplication Content: Short- to medium-length Hindi text (up to ~5000 characters), drawn from web-crawled corpora Fields: One field text (string) containing the Hindi text Sources: mC4-Hindi-Cleaned-3.0 OSCAR-2301-Hindi-Cleaned-2.0 ai4bharat/sangraha (Hindi, verified) Deduplication:… See the full description on the dataset page: https://huggingface.co/datasets/nis12ram/nisram-hindi-text-0.0.
065
Update README.md
Update README.md
Update README.md
Upload dataset
initial commit
