datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_exercises
Dataset Card for "code_exercises"
Code exercise
This dataset is composed of a diverse set of ~120k Python code exercises (~120m total tokens) generated by ChatGPT 3.5. It is designed to distill ChatGPT 3.5 knowledge about Python coding tasks into other (potentially smaller) models. The exercises have been generated by following the steps described in the related GitHub repository.
The generated exercises follow the format of the Human Eval benchmark. Each training sample… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/code_exercises.exercise-dataset
Exercise Dataset — Free Tier (RepDB)
A free, ready-to-use fitness exercise dataset: 601 exercises, each
illustrated with flat-style 512×512 WebP images (a start/peak pose pair, or
a single main pose for static holds and stretches), with target muscles,
equipment, MET values, and full instructions in English, German, and
Spanish.
This public snapshot is the free tier of RepDB. Free for personal
and commercial use inside applications, with attribution.
Need exercise… See the full description on the dataset page: https://huggingface.co/datasets/RepDB/exercise-dataset.ultradata-math-textbook-exercise-ar
ultradata-math-textbook-exercise-ar
Arabic translation of the English portion of UltraData-Math, config UltraData-Math-L3-Textbook-Exercise-Synthetic: synthetic textbook-style content and exercises generated around specific mathematical knowledge points. Translated with the midtrans pipeline: text is segmented into prose and verbatim blocks (LaTeX, code, tables, and inline non-translatables are masked and never sent to the model, so formulas cannot be mangled), prose is… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/ultradata-math-textbook-exercise-ar.UltraData-Math-L3-Textbook-Exercise-Synthetic-split
UltraData-Math L3 Textbook Exercise Synthetic Split
Source dataset: openbmb/UltraData-Math
Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic
Each row contains:
uid
question
answer
The original content field was split using the literal markers
The exercise: and The solution:.
edugraph-exercises
EduGraph Exercises Dataset
EduGraph Exercises is a synthetic ML dataset of math-related visual problems, precisely labeled for training AI models in the education sector.
Every image in this dataset is programmatically generated using the EduGraph Ontology to ensure that visual features are mathematically bound to their pedagogical labels.
Quick Links
Generation Engine: GitHub Repository (Contribute new generators or views!)
Ontology: EduGraph Ontology (Semantic… See the full description on the dataset page: https://huggingface.co/datasets/christian-bick/edugraph-exercises.UltraData-Math-L3-Textbook-Exercise-Synthetic-split-qwen3-0.6b-embeddedGym-Exercise-Video-Analysis
Gym-Exercise-Video-Analysis
Gym-Exercise-Video-Analysis is a specialized multimodal video understanding dataset comprising 500 annotated gym workout and exercise clips. It is designed for fine-tuning and evaluating Video-Language Models (Video-LLMs), visual fitness coaches, and temporal exercise analysis systems. Each entry pairs exercise videos and extracted frame sequences with in-depth textual descriptions, biomechanical observations, form evaluations, and routine tracking.… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Gym-Exercise-Video-Analysis.Exercise-Synthetic-split-ncert-chapter-mapped_filtered_difficulty_scoredUltraData-Math-L3-Textbook-Exercise-Synthetic-split-ncert-chapter-mappedUltraData-Math-L3-Textbook-Exercise-Synthetic-split-ncert-chapter-mapped_filteredcomplex-numbers-exercises-1000
📘 Exercices sur les Nombres Complexes – Dataset (1000 échantillons)
Ce dataset contient 1000 exercices entièrement générés sur les nombres complexes, accompagnés de corrections détaillées et pédagogiques, prêts à être utilisés pour :
l’entraînement de modèles d’IA éducatives,
la génération automatique d’exercices,
la correction automatique,
l’explication pas-à-pas du raisonnement mathématique.
Il s’inscrit dans un projet plus large visant à construire des IA spécialisées en… See the full description on the dataset page: https://huggingface.co/datasets/7rouz/complex-numbers-exercises-1000.calisthenics_exercises
Calisthenics Exercises Dataset
A comprehensive dataset of 170 unique calisthenics exercises, each with three progression levels (beginner -> intermediate -> advanced).
Web App
Live: martjn-calisthenics-exercises.static.hf.space
Interactive single-page app with 8 filter dimensions, favorites, keyboard shortcuts, and responsive design. Hosted on HuggingFace Spaces (static SDK). The Space is a thin loader that fetches index.html from this dataset repo at runtime — any… See the full description on the dataset page: https://huggingface.co/datasets/Martjn/calisthenics_exercises.coding_exercises_filtered
Dataset Card for "coding_exercises_filtered"
Coding exercises generated by gpt, then filtered. This has a lot of duplicates - would not recommend using as is.
calisthenics_exercisesKLUE-MRC-exercisecfa_extracted_exercise_sup_sample_from_policy_v1.1_genrm_qwen3-32b_dpo_binarized_filtered_2048exercise-api
Exercise API — Dataset
Dataset de 104 ejercicios de gimnasio (bilingüe ES/EN) derivado de la
Exercise API. Cada ejercicio incluye grupo muscular,
equipamiento, músculos principal/secundario, instrucciones paso a paso e ilustración
masculina y femenina (208 imágenes en total).
Configuraciones
images — 1 fila por imagen (208). Etiquetas (grupo, equipamiento,
músculos, género) + caption_es/caption_en. Para clasificación de imagen y multimodal
(image-to-text / VQA).… See the full description on the dataset page: https://huggingface.co/datasets/natzx94/exercise-api.cfa_extracted_exercise_sup_sample_from_policy_v1.1_genrm_qwq-32b_dpo_binarizedtranslation-de-en-exerciseGYM-Exercisecfa_extracted_exercise_sup_sample_from_policy_v1.1_stepwise_dpo_binarized_chunk_20cfa_extracted_exercise_sup_sample_from_policy_v1.1_genrm_qwen3-32b_stepwise_dpo_binarizedcfa_extracted_exercise_sup_sample_from_policy_v1.1_stepwise_dpo_binarized_chunk_13cfa_extracted_exercise_sup_sample_from_policy_v1.1_genrm_qwen3-32b_dpo_binarized_filter2048cfa_extracted_exercise_sup_sample_from_policy_v1_1_rpo_stepwise_iter_1_dpo_binarizednews_2026_exercise
news_2026_exercise
Arabic news dataset for text classification.
Column
Description
title
News title
content
News body
category
Label: سياسة, اقتصاد, صحة, رياضة
28,000 rows (7,000 per category).
Download and load as pandas
from datasets import load_dataset
import pandas as pd
ds = load_dataset("maher13/news_2026_exercise")
df = ds["train"].to_pandas()
print(df.head())
Sample N rows from each category
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/maher13/news_2026_exercise.cfa_extracted_exercise_sup_sample_from_policy_v1.1_stepwise_dpo_binarizedcfa_extracted_exercise_sup_sample_from_policy_v1.1_genrm_qwen3-32b_stepwise_dpo_binarized_F2048huberman_on_exercise
Dataset Card for "huberman_on_exercise"
More Information needed
cfa_extracted_exercise_sup_sample_from_policy_v1_1_rpo_stepwise_iter_1_dpo_val_chunk_18
