datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
idrid-disease-grading
Indian Diabetic Retinopathy Image Dataset (IDRiD)
This dataset is the disease grading portion of the IDRiD.
The original source of the dataset is here: https://ieee-dataport.org/open-access/indian-diabetic-retinopathy-image-dataset-idrid
diabetic-retinopathy-grading-africa
DR-Grading — Fundus Diabetic-Retinopathy Grading with Africa-Grounded Synthetic Clinical Context
A dataset for cross-sectional 5-class diabetic-retinopathy grading, pairing
resized colour fundus photographs with minimal, epidemiology-grounded synthetic
point-of-care context (age, sex, diabetes duration).
Version 1.0.0 · core dr_synth 1.0.0 · part of the DR-Africa dataset family
(see also dr-progression and
dr-africa-benchmark).
Abstract
Diabetic retinopathy (DR)… See the full description on the dataset page: https://huggingface.co/datasets/macular/diabetic-retinopathy-grading-africa.lithium-ion-cell-capacity-grading-process-curves
Lithium-ion Cell Capacity-Grading (FR) Process Curves
Channel-level process curves from the capacity-grading station of a cylindrical
lithium-ion cell line. Every tester channel is sampled natively every 30 seconds for
the whole process, giving the full voltage / current / capacity trajectory of each cell
from the moment it is clamped.
This is the capacity-grading (FR) dataset. Pre-charge is a completely different
process and is published separately; the two are deliberately… See the full description on the dataset page: https://huggingface.co/datasets/michealsmitch/lithium-ion-cell-capacity-grading-process-curves.imo-gradingbench
IMO-GradingBench
Dataset Description
IMO-GradingBench is a benchmark dataset for evaluating the automatic grading capabilities of large language models. It consists of 1,000 human gradings of model-generated solutions to mathematical problems.
This dataset is part of the IMO-Bench suite, released by Google DeepMind in conjunction with their 2025 IMO gold medal achievement.
Supported Tasks and Leaderboards
The primary task for this dataset is automatic grading… See the full description on the dataset page: https://huggingface.co/datasets/Hwilner/imo-gradingbench.DR_Grading
Dataset Card for "DR_Grading"
More Information needed
Disease_Grading_for_DR_and_Mucula
Dataset Card for "Disease_Grading_for_DR_and_Mucula"
More Information needed
african_plum_grading_classification
African Plum Grading Classification
A dataset for grade classification of plums. The dataset contains 4,507 images across 6 classes: bruised, cracked, rotten, spotted, unaffected, unripe.
Images per class:
bruised: 319
cracked: 162
rotten: 720
spotted: 759
unaffected: 1,721
unripe: 826
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{fadja2025dataset,
title={A dataset of annotated African plum… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/african_plum_grading_classification.humor_gradinggrading-templatesgradingEssay-quetions-auto-grading-arabicDataset Overview
The Open Orca Enhanced Dataset is meticulously designed to improve the performance of automated essay grading models using deep learning techniques. This dataset integrates robust data instances from the FLAN collection, augmented with responses generated by GPT-3.5 or GPT-4, creating a diverse and context-rich resource for training models.
Dataset Structure
The dataset is structured in a tabular format, with the following key fields:
id: A unique identifier for each data… See the full description on the dataset page: https://huggingface.co/datasets/mohamedemam/Essay-quetions-auto-grading-arabic.Melodic_pattern_reproduction_performances_gradingEssay-quetions-auto-gradingDataset Overview
The Open Orca Enhanced Dataset is meticulously designed to improve the performance of automated essay grading models using deep learning techniques. This dataset integrates robust data instances from the FLAN collection, augmented with responses generated by GPT-3.5 or GPT-4, creating a diverse and context-rich resource for training models.
Dataset Structure
The dataset is structured in a tabular format, with the following key fields:
id: A unique identifier for each data… See the full description on the dataset page: https://huggingface.co/datasets/mohamedemam/Essay-quetions-auto-grading.DR_Grading_413_103
Dataset Card for "DR_Grading_413_103"
More Information needed
matharena-gradingbenchgarment-grading-specsgrading_finetuneunified-feedback-grading-adbidrid_grading
Dataset Card for "idrid_grading"
More Information needed
imo_mixed_sft_gradinggrading_data_cleanedPutnam-AXIOM-Gradinghayabusa_grading_report
Developed by: LockeLamora2077 - Maximilian Gutowski within my Master Thesis maximilian-gutowski-a94771199
imo_mixed_sft_grading_processedgrading_datagrading_modular
