jiviteshjn/hi-proverbs-cpt
Hindi Proverbs — Idiom-Tagged Continued-Pretraining Corpus A 338K-document Hindi corpus (~1.4B tokens) for continued pretraining on cultural knowledge in figurative language, plus a structured dataset of 16,617 Hindi proverbs (लोकोक्तियाँ) with meanings, recovered via OCR-repair from a classic proverb dictionary. Each corpus document is natural Hindi text containing at least one proverb (matched including common surface variants), with an appended knowledge block listing every… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/hi-proverbs-cpt.
This repository belongs to jiviteshjn on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
