datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-textbooks-corpus
Nepali Textbooks Corpus for Grades 1-12
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 5634
Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]
Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.nepali-textbooks-math-grade10
Nepali Textbook Pretraining Corpus — Sample (Grade 10 Math)
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Schema
id — unique segment id
source — book/source string
book_id — filename-derived id (if present)
subject — subject label (e.g., "math")
grade — class/grade
chapter_index — chapter number
chapter_title — chapter/unit name
segment_index — index within chapter
text — content chunk
tokens_approx — rough token count… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-math-grade10.nepali-textbooks-grade10
Nepali Textbooks Grade 10
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 1936
Grades: [10]
Subjects: ['Civic_Science', 'Education', 'Health_and_Physical_Education', 'Population_Studies', 'Social_Studies', 'Sociology', 'computer_science', 'economics', 'environmental_science', 'health', 'history', 'math', 'nepali', 'optional_math', 'science', 'social']
Total chars: 5903753
Avg tokens per sample: 492… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-grade10.
