datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron-student-fail-v41-clean-thinking
Nemotron-fail / DeepSeek-V4.1 clean and action-only trajectories
DeepSeek-V4.1 reward-1 trajectories for tasks on which the Nemotron student
did not obtain reward 1. This release was rebuilt from the complete reward-1
audit under v54-high-precision-canonical-reconstruction-relations.
Training paths
Path
Rows
Unique tasks
Thinking
Use
data/strict/train.jsonl.gz
12
12
Preserved and clean
Raw-thinking SFT
data/hybrid/train.jsonl.gz
58
58
Only… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/nemotron-student-fail-v41-clean-thinking.student-question-categoriesThis is the IITJEE NEET AIIMS Students Questions Data dataset.
It categorizes university entry questions into 4 categories: Physics, Chemistry, Biology, and Mathematics.
Education-High-School-StudentsDetails coming soon!!
student_and_llm_essays
Dataset Card for Academic Essay Prompt-Completion Pairs
Dataset Description
This dataset is designed to distinguish between essays authored by students and those generated by Large Language Models (LLMs), offering an essential resource for researchers and practitioners in natural language processing, educational technology, and academic integrity. Hosted on Huggingface, it supports the development and evaluation of models aimed at identifying the origin of textual content… See the full description on the dataset page: https://huggingface.co/datasets/knarasi1/student_and_llm_essays.tdtu-student-regulations-qa
TDTU Vietnamese University Regulations QA Dataset
Dataset Description
Tập dữ liệu hỏi-đáp tiếng Việt về quy chế, quy định sinh viên của Trường Đại học Tôn Đức Thắng (TDTU), được xây dựng cho bài toán Retrieval-Augmented Generation (RAG) và fine-tuning LLM tư vấn sinh viên.
Ngôn ngữ: Tiếng Việt
Domain: Quy chế đại học, chính sách sinh viên
Mục đích: Huấn luyện chatbot tư vấn sinh viên TDTU
Dataset Details
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/hungminhss/tdtu-student-regulations-qa.gemma4-onpolicy-student-corrections
Gemma 4 12B FrontierDistill - On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project.
This dataset contains 2,000 on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill).
Every example in this dataset… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-student-corrections.Education-College-StudentsDetails coming soon!!
scas_verified_teacher_pool
SCAS Verified Teacher Answer Pool
This dataset provides an aligned, correctness-verified pool of
teacher-generated mathematical reasoning solutions for studying
student-centric data selection in distillation.
The release covers two source corpora, Hendrycks MATH and DeepScaleR. For each
corpus, we retain the subset of questions on which all nine selected teacher
models produce verified correct answers. Each retained question is paired with
nine alternative teacher solutions, one… See the full description on the dataset page: https://huggingface.co/datasets/Student-Centric-Answer-Sampling/scas_verified_teacher_pool.student_enrolled_in_different_level_by_district_year
Nepali Grounded Education SFT Dataset
Nepali-language ShareGPT-format supervised fine-tuning (SFT) dataset generated
from a Ministry of Education enrollment dataset, with every question and
answer deterministically templatized (no translator, no LLM, no API, no
reward model in the generation loop) and separately audited for semantic
grounding and correctness.
Source
Resource: Number of Students Enrolled in Primary and Secondary School
By 75 Districts, 2003–2015… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/student_enrolled_in_different_level_by_district_year.LeroyDyer___Spydaz_Web_AI_AGI_R1_Top_Student-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_Top_Student
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_Top_Student
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_Top_Student-details.Open-Wyvern-74k
The Wyvern 🐉 Dataset
Let's introduce the Wyvern 🐉 dataset, the new combination of datasets(Open-Orca,
Open-Platypus, airoboros,
Dolly)!
We have integrated high-quality datasets following the claim that quality is more matter than quantity.
In addition, we have deduplicated the duplication of datasets to improve the dataset's quality because each dataset has some data contaminations.
Please see below for more details about the dataset!
Dataset Details
Wyvern 🐉… See the full description on the dataset page: https://huggingface.co/datasets/StudentLLM/Open-Wyvern-74k.student-stress-surveyLeroyDyer___Spydaz_Web_AI_AGI_R1_Student_Coder-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_Student_Coder
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_Student_Coder
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_Student_Coder-details.LeroyDyer___Spydaz_Web_AI_AGI_R1_Math_Student-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_Math_Student
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_Math_Student
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_Math_Student-details.student_partial_sciworld
student_partial_sciworld
Partial SciWorld trajectories: Qwen3-1.7B (base) acting for its first 5 turns on
the 2,120-task training split. These are the student prefixes an online-ROSE run sees
before the teacher takes over — captured separately so the take-over point can be
studied, or teacher continuations generated offline.
What is here
rows
2,120 (one per training task)
file
student_partial_sciworld.jsonl (11.4 MB)
model
Qwen/Qwen3-1.7B… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/student_partial_sciworld.sudoku-student-minekuk-nemtron8b-rollouts-n2-tokens16384students-subject-preferences
Students' Subject Preferences
A small survey-style dataset recording which school subjects five students like and dislike.
Each row is one student: their ID, the subjects they named as favorites, and the subjects they
named as least favorites. Subject names are in Mongolian Cyrillic.
Files
File
Rows
Description
data/train.jsonl
5
One JSON object per student
Schema
Column
Type
Description
student_id
int
Student identifier… See the full description on the dataset page: https://huggingface.co/datasets/sumya123/students-subject-preferences.MedQA-CS-Studentindian_university_guidance_for_bangladeshi_students
Indian University Guidance for Bangladeshi Students Dataset
Dataset Description
This dataset contains 7,044 high-quality, instruction-formatted Question-Answer pairs designed for fine-tuning Large Language Models (LLMs). The primary goal of this dataset is to create a specialized AI counselor that provides accurate, culturally relevant, and comprehensive guidance on Indian universities for Bangladeshi students.
The dataset was generated through a sophisticated… See the full description on the dataset page: https://huggingface.co/datasets/millat/indian_university_guidance_for_bangladeshi_students.Sampled_Orca_GPT4
Stratify Sampled Dataset of Open-Orca 🐬
This dataset is a stratified sampled dataset of Open-Orca's GPT-4 answered dataset(1M-GPT4-Augmented.parquet) [Link]
For sampling the dataset stratify, train_test_split of scikit-learn library was used.
The specific setup of sampling is as follows:
split_size: 0.05
shuffle: True
stratify: 'id' of Open-Orca dataset
Teacher-Student-Dialoguesstudent-assistance-chatbotKorean_Vicuna_questionsstudent-question-categoriesThis is the IITJEE NEET AIIMS Students Questions Data dataset.
It categorizes university entry questions into 4 categories: Physics, Chemistry, Biology, and Mathematics.
argumentative_student_peer_reviews
Dataset Card for Argumentation Annotated Student Peer Reviews Corpus
The dataset originates from the research results of T. Wambsganss, C. Niklaus, M. Söllner, S. Handschuh and J. M. Leimeister available here.
The dataset is originally in the brat standoff format and was converted to JSONL using pybrat.
During the training and evaluation of the data, it became apparent that the data was relatively heterogeneous and noisy. Due to this fact, the trained model only achieved an accuracy… See the full description on the dataset page: https://huggingface.co/datasets/samirmsallem/argumentative_student_peer_reviews.minesweeper-student-kukurasu20k-qwen1.7b-e3-mask-rollouts-n2-tokens16384minesweeper-student-minekuk-qwen1.7b-continued-by-qwen3-4b-thinking-t4096-r16384defendable-pain-student-loan-pain-v0.1
Student Loan Pain Receipt
"the indentured" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 7 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
7 pain receipts… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-student-loan-pain-v0.1.studentsstudent-stress-survey-MKURAI
