datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qualcomm-exercise-video-dataset-benchmark
Dataset Card for Qualcomm Exercise Video Dataset (Benchmark)
This is the benchmark split of the dataset as described here
This is a FiftyOne dataset with 74 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/qualcomm-exercise-video-dataset-benchmark.code_exercises
Dataset Card for "code_exercises"
Code exercise
This dataset is composed of a diverse set of ~120k Python code exercises (~120m total tokens) generated by ChatGPT 3.5. It is designed to distill ChatGPT 3.5 knowledge about Python coding tasks into other (potentially smaller) models. The exercises have been generated by following the steps described in the related GitHub repository.
The generated exercises follow the format of the Human Eval benchmark. Each training sample… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/code_exercises.exercise-dataset
Exercise Dataset — Free Tier (RepDB)
A free, ready-to-use fitness exercise dataset: 601 exercises, each
illustrated with flat-style 512×512 WebP images (a start/peak pose pair, or
a single main pose for static holds and stretches), with target muscles,
equipment, MET values, and full instructions in English, German, and
Spanish.
This public snapshot is the free tier of RepDB. Free for personal
and commercial use inside applications, with attribution.
Need exercise… See the full description on the dataset page: https://huggingface.co/datasets/RepDB/exercise-dataset.new-gym-exercises
Gym Exercises Dataset 💪
Bienvenue dans le Gym Exercises Dataset !
Ce dataset regroupe une collection d'exercices de sport et de musculation.
📌 Particularité du Dataset
Pour chaque exercice, vous trouverez deux fichiers correspondants :
🖼️ Une image fixe (.jpg) : Montrant l'exercice.
🎬 Une animation (.gif) : Montrant le mouvement complet de l'exercice pour bien comprendre l'exécution.
Les deux fichiers partagent exactement le même nom, ce qui permet de les… See the full description on the dataset page: https://huggingface.co/datasets/Nouira-Oussema/new-gym-exercises.ultradata-math-textbook-exercise-ar
ultradata-math-textbook-exercise-ar
Arabic translation of the English portion of UltraData-Math, config UltraData-Math-L3-Textbook-Exercise-Synthetic: synthetic textbook-style content and exercises generated around specific mathematical knowledge points. Translated with the midtrans pipeline: text is segmented into prose and verbatim blocks (LaTeX, code, tables, and inline non-translatables are masked and never sent to the model, so formulas cannot be mangled), prose is… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/ultradata-math-textbook-exercise-ar.UltraData-Math-L3-Textbook-Exercise-Synthetic-split
UltraData-Math L3 Textbook Exercise Synthetic Split
Source dataset: openbmb/UltraData-Math
Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic
Each row contains:
uid
question
answer
The original content field was split using the literal markers
The exercise: and The solution:.
edugraph-exercises
EduGraph Exercises Dataset
EduGraph Exercises is a synthetic ML dataset of math-related visual problems, precisely labeled for training AI models in the education sector.
Every image in this dataset is programmatically generated using the EduGraph Ontology to ensure that visual features are mathematically bound to their pedagogical labels.
Quick Links
Generation Engine: GitHub Repository (Contribute new generators or views!)
Ontology: EduGraph Ontology (Semantic… See the full description on the dataset page: https://huggingface.co/datasets/christian-bick/edugraph-exercises.UltraData-Math-L3-Textbook-Exercise-Synthetic-split-qwen3-0.6b-embeddedGym-Exercise-Video-Analysis
Gym-Exercise-Video-Analysis
Gym-Exercise-Video-Analysis is a specialized multimodal video understanding dataset comprising 500 annotated gym workout and exercise clips. It is designed for fine-tuning and evaluating Video-Language Models (Video-LLMs), visual fitness coaches, and temporal exercise analysis systems. Each entry pairs exercise videos and extracted frame sequences with in-depth textual descriptions, biomechanical observations, form evaluations, and routine tracking.… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Gym-Exercise-Video-Analysis.Exercise-Synthetic-split-ncert-chapter-mapped_filtered_difficulty_scoredVidChain-exercise
✏️ Data for VidChain Excercise
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning
Ji Soo Lee*, Jongha Kim*, Jeehye Na, Jinyoung Park, Hyunwoo J. Kim†.
AAAI 2025
🎯 Learning Objectives
By working through this exercise, you will:
Reproduce baseline behavior of a video-language model (VTimeLLM, CVPR 2024 Highlight).
Observe the limitations of existing approaches in temporal reasoning… See the full description on the dataset page: https://huggingface.co/datasets/simplecloud/VidChain-exercise.gym-exercises
Gym Exercises
A video library of gym/fitness exercise demonstrations, originally packaged as HTML5-compatible assets (exercises_html5_videos.tar) for a workout-tracking web app. Each exercise is numbered by an ID and provided in up to three encodings for browser compatibility: .mp4, .webm, and .ogv.
1,042 video files, ~2.3 GB, split by performer gender and video resolution/size tier:
male/original/<id>.<ext> # 298 files — full-size/original quality
male/small/<id>.<ext>… See the full description on the dataset page: https://huggingface.co/datasets/webbrain-one/gym-exercises.bench-press-deadlift-exercises
Bench Press and Deadlift Exercise Videos
This dataset contains short exercise clips for binary video classification.
The source frame folders were encoded as MP4 files so the repository follows the Hugging Face VideoFolder layout.
Set the final license before publishing this dataset publicly.
Labels
bench_press
deadlift
Splits
split
total
bench_press
deadlift
train
75
49
26
validation
9
6
3
test
9
6
3
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/bench-press-deadlift-exercises.omx_f_exercise_0415This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omx_f",
"total_episodes": 30,
"total_frames": 5869,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dream-79/omx_f_exercise_0415.gym-exercise-landmarkssarc-o-meter-exercise-skeleton
Sarc-O-Meter Exercise Skeleton Dataset
A 33-keypoint MediaPipe skeleton dataset for exercise movement analysis.
Exercises
Calf Raise
Sit to Stand
Step Up
Speed Classes
Slow
Normal
Fast
Dataset Format
The skeleton data is stored as .npy files. Each sequence contains MediaPipe Pose landmarks with the shape:
Frames × 33 × 4
The four values represent:
x
y
z
visibility
The dataset also contains metadata.parquet, which provides… See the full description on the dataset page: https://huggingface.co/datasets/Surya2212/sarc-o-meter-exercise-skeleton.UltraData-Math-L3-Textbook-Exercise-Synthetic-split-ncert-chapter-mappedUltraData-Math-L3-Textbook-Exercise-Synthetic-split-ncert-chapter-mapped_filteredReal_Time_Exercise_Recognition_Dataset
DESCRIPTION OF THE DATASET
This dataset was created for real-time fitness exercise classification and includes a diverse mix of synthetic and real-world videos. It focuses on four common exercises:
Squat
Push-up
Barbell Bicep Curl
Shoulder Press
DEMO OF MY PROJECT THAT USED THIS DATASET (AI PERSONAL TRAINER):
GITHUB PROJECT (with implementation):
https://github.com/RiccardoRiccio/Fitness-AI-Trainer-With-Automatic-Exercise-Recognition-and-Counting
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/RickyRiccio/Real_Time_Exercise_Recognition_Dataset.calisthenics_exercises
Calisthenics Exercises Dataset
A comprehensive dataset of 170 unique calisthenics exercises, each with three progression levels (beginner -> intermediate -> advanced).
Web App
Live: martjn-calisthenics-exercises.static.hf.space
Interactive single-page app with 8 filter dimensions, favorites, keyboard shortcuts, and responsive design. Hosted on HuggingFace Spaces (static SDK). The Space is a thin loader that fetches index.html from this dataset repo at runtime — any… See the full description on the dataset page: https://huggingface.co/datasets/Martjn/calisthenics_exercises.complex-numbers-exercises-1000
📘 Exercices sur les Nombres Complexes – Dataset (1000 échantillons)
Ce dataset contient 1000 exercices entièrement générés sur les nombres complexes, accompagnés de corrections détaillées et pédagogiques, prêts à être utilisés pour :
l’entraînement de modèles d’IA éducatives,
la génération automatique d’exercices,
la correction automatique,
l’explication pas-à-pas du raisonnement mathématique.
Il s’inscrit dans un projet plus large visant à construire des IA spécialisées en… See the full description on the dataset page: https://huggingface.co/datasets/7rouz/complex-numbers-exercises-1000.coding_exercises_filtered
Dataset Card for "coding_exercises_filtered"
Coding exercises generated by gpt, then filtered. This has a lot of duplicates - would not recommend using as is.
KLUE-MRC-exercisecalisthenics_exercisesexercise-api
Exercise API — Dataset
Dataset de 104 ejercicios de gimnasio (bilingüe ES/EN) derivado de la
Exercise API. Cada ejercicio incluye grupo muscular,
equipamiento, músculos principal/secundario, instrucciones paso a paso e ilustración
masculina y femenina (208 imágenes en total).
Configuraciones
images — 1 fila por imagen (208). Etiquetas (grupo, equipamiento,
músculos, género) + caption_es/caption_en. Para clasificación de imagen y multimodal
(image-to-text / VQA).… See the full description on the dataset page: https://huggingface.co/datasets/natzx94/exercise-api.cfa_extracted_exercise_sup_sample_from_policy_v1.1_genrm_qwen3-32b_dpo_binarized_filtered_2048news_2026_exercise
news_2026_exercise
Arabic news dataset for text classification.
Column
Description
title
News title
content
News body
category
Label: سياسة, اقتصاد, صحة, رياضة
28,000 rows (7,000 per category).
Download and load as pandas
from datasets import load_dataset
import pandas as pd
ds = load_dataset("maher13/news_2026_exercise")
df = ds["train"].to_pandas()
print(df.head())
Sample N rows from each category
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/maher13/news_2026_exercise.cfa_extracted_exercise_sup_sample_from_policy_v1.1_genrm_qwq-32b_dpo_binarizedtranslation-de-en-exercisecfa_extracted_exercise_sup_sample_from_policy_v1.1_genrm_qwen3-32b_stepwise_dpo_binarized
