CoolFace
Datasetpublic

dineshkarki/nepali-textbooks-grade10

Nepali Textbooks Grade 10 This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks. Summary Samples: 1936 Grades: [10] Subjects: ['Civic_Science', 'Education', 'Health_and_Physical_Education', 'Population_Studies', 'Social_Studies', 'Sociology', 'computer_science', 'economics', 'environmental_science', 'health', 'history', 'math', 'nepali', 'optional_math', 'science', 'social'] Total chars: 5903753 Avg tokens per sample: 492… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-grade10.

sourceHugging Faceotherupdated 1y agoView on Hugging Face
0likes4downloads
3 commits on main
8d24e4c1y ago

Add dataset.jsonl (merged_all.jsonl)

dineshkarki
ccb600c1y ago

Add dataset.jsonl (merged_dataset.jsonl)

dineshkarki
ca73c4e1y ago

initial commit

dineshkarki