datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
k12-indian-curriculum-4.9m
BharatLLM K-12 Indian Curriculum Dataset (4.9M)
4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages.
Language
Script
Entries
English
Latin
~594K
Hindi
Devanagari
~449K
Bengali
Bengali
~408K
Telugu
Telugu
~408K
Tamil
Tamil
~408K
Kannada
Kannada
~408K
Malayalam
Malayalam
~408K
Marathi
Devanagari
~408K
Gujarati
Gujarati
~408K
Odia
Odia
~408K
PunjabiGurmukhi
~408K
Urdu
Nastaliq
~374K
Format
{… See the full description on the dataset page: https://huggingface.co/datasets/FoundryAILabs/k12-indian-curriculum-4.9m.foundry-y-yoruba-corpus
Yorùbá News & Tasks — a contamination-proof, 100% human Yorùbá corpus
A 19,997-row supervised fine-tuning corpus ({"prompt","completion"} JSONL) for
adapting a language model to Yorùbá, built for the Adaption AutoScientist
challenge (language track). Every row is human-written, from an approved,
train-split-only source. No synthetic or LLM-generated text.
SHA-256: a711692096719c0d11f8e9c3784061211733029bf32a20e36b7e920488f8cc0c
Rows: 19,997 · Format: JSONL, prompt + completion… See the full description on the dataset page: https://huggingface.co/datasets/Enochid/foundry-y-yoruba-corpus.
