datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_exercises
Dataset Card for "code_exercises"
Code exercise
This dataset is composed of a diverse set of ~120k Python code exercises (~120m total tokens) generated by ChatGPT 3.5. It is designed to distill ChatGPT 3.5 knowledge about Python coding tasks into other (potentially smaller) models. The exercises have been generated by following the steps described in the related GitHub repository.
The generated exercises follow the format of the Human Eval benchmark. Each training sample… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/code_exercises.ultradata-math-textbook-exercise-ar
ultradata-math-textbook-exercise-ar
Arabic translation of the English portion of UltraData-Math, config UltraData-Math-L3-Textbook-Exercise-Synthetic: synthetic textbook-style content and exercises generated around specific mathematical knowledge points. Translated with the midtrans pipeline: text is segmented into prose and verbatim blocks (LaTeX, code, tables, and inline non-translatables are masked and never sent to the model, so formulas cannot be mangled), prose is… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/ultradata-math-textbook-exercise-ar.UltraData-Math-L3-Textbook-Exercise-Synthetic-split
UltraData-Math L3 Textbook Exercise Synthetic Split
Source dataset: openbmb/UltraData-Math
Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic
Each row contains:
uid
question
answer
The original content field was split using the literal markers
The exercise: and The solution:.
complex-numbers-exercises-1000
📘 Exercices sur les Nombres Complexes – Dataset (1000 échantillons)
Ce dataset contient 1000 exercices entièrement générés sur les nombres complexes, accompagnés de corrections détaillées et pédagogiques, prêts à être utilisés pour :
l’entraînement de modèles d’IA éducatives,
la génération automatique d’exercices,
la correction automatique,
l’explication pas-à-pas du raisonnement mathématique.
Il s’inscrit dans un projet plus large visant à construire des IA spécialisées en… See the full description on the dataset page: https://huggingface.co/datasets/7rouz/complex-numbers-exercises-1000.calisthenics_exercises
Calisthenics Exercises Dataset
A comprehensive dataset of 170 unique calisthenics exercises, each with three progression levels (beginner -> intermediate -> advanced).
Web App
Live: martjn-calisthenics-exercises.static.hf.space
Interactive single-page app with 8 filter dimensions, favorites, keyboard shortcuts, and responsive design. Hosted on HuggingFace Spaces (static SDK). The Space is a thin loader that fetches index.html from this dataset repo at runtime — any… See the full description on the dataset page: https://huggingface.co/datasets/Martjn/calisthenics_exercises.exercise-api
Exercise API — Dataset
Dataset de 104 ejercicios de gimnasio (bilingüe ES/EN) derivado de la
Exercise API. Cada ejercicio incluye grupo muscular,
equipamiento, músculos principal/secundario, instrucciones paso a paso e ilustración
masculina y femenina (208 imágenes en total).
Configuraciones
images — 1 fila por imagen (208). Etiquetas (grupo, equipamiento,
músculos, género) + caption_es/caption_en. Para clasificación de imagen y multimodal
(image-to-text / VQA).… See the full description on the dataset page: https://huggingface.co/datasets/natzx94/exercise-api.luganda-bilingual-literacy-exercises
Luganda-English Bilingual Literacy Exercises (P1–P3)
3,472 structured bilingual exercises for Ugandan primary school literacy instruction (Primary 1 through Primary 3). Each exercise contains parallel English and Luganda versions with questions, answers, and explanations.
Dataset Description
Grade
Exercises
File
P1
1,157
data/p1_exercises.json
P2
1,135
data/p2_exercises.json
P3
1,180
data/p3_exercises.json
Total
3,472
Exercise Types… See the full description on the dataset page: https://huggingface.co/datasets/CraneAILabs/luganda-bilingual-literacy-exercises.Complex_Number_Exercises_for_AI_Training
🧠 MathVerse Generator — L’IA tunisienne qui crée des exercices de mathématiques avec raisonnement logique
Donne-moi un exercice, et je te génère des centaines de variantes corrigées automatiquement.Chaque exercice suit la logique d’un vrai raisonnement mathématique, écrit étape par étape, avec vérification symbolique.
📬 Contact :
Email : contact@abramarsolution.com
Site officiel : https://abramarsolution.com
✨ Fonctionnalités principales
✅ Génération… See the full description on the dataset page: https://huggingface.co/datasets/7rouz/Complex_Number_Exercises_for_AI_Training.alignment-internship-exercise
Dataset Card for the Alignement Internship Exercise
Dataset Description
This dataset provides a list of questions accompanied by Phi-2's best answer to them, as ranked by OpenAssitant's reward model.
Dataset Creation
The questions were handpicked from the LDJnr/Capybara, Open-Orca/OpenOrca and truthful_qa datasets, the coding exercise is from LeetCode's top 100 liked questions and I found the last prompt on a blog and modified it. I have chosen these prompts… See the full description on the dataset page: https://huggingface.co/datasets/gsoisson/alignment-internship-exercise.warmup-exercises-fr
Exercices d'échauffement pour la musculation (français)
52 exercices d'échauffement en français, annotés pour permettre la génération
automatisée de protocoles adaptés à un pratiquant donné.
Ce jeu de données alimente Warmup Generator,
un générateur d'échauffements gratuit et sans publicité.
Ce qui distingue ce jeu de données
La plupart des bases d'exercices se contentent d'associer un mouvement à un
groupe musculaire. Celle-ci porte en plus deux champs… See the full description on the dataset page: https://huggingface.co/datasets/fabdigstudio/warmup-exercises-fr.
