datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
digital-coach
DigitalCoach Dataset
DigitalCoach is a multimodal expert-novice computer-use coaching dataset for studying how humans teach software skills through grounded dialogue.It contains 72 coaching sessions, 22,752 dialogue turns, and 28.1 hours of screen recordings, collected across 5 software applications in creativity, engineering, and productivity-oriented workflows.
Each session pairs one expert coach with one novice learner, and captures timestamped data:
dialogue transcripts… See the full description on the dataset page: https://huggingface.co/datasets/berkeley-hci/digital-coach.CPR-Coach
Packages Check
The complete file list is shown below.
.
├── CPR_Dataset_S0.tar.gz.00 # 10 GB
├── CPR_Dataset_S0.tar.gz.01 # 10 GB
├── CPR_Dataset_S0.tar.gz.02 # 10 GB
├── CPR_Dataset_S0.tar.gz.03 # 10 GB
├── CPR_Dataset_S0.tar.gz.04 # 10 GB
├── CPR_Dataset_S0.tar.gz.05 # 10 GB
├── CPR_Dataset_S0.tar.gz.06 # 10 GB
├── CPR_Dataset_S0.tar.gz.07 # 5.4 GB
├── CPR_Dataset_S1.tar.gz.00 # 10… See the full description on the dataset page: https://huggingface.co/datasets/ShunliWang/CPR-Coach.accent_coach_training_datasetrunning-coach-sft
Running Coach SFT
Instruction-tuning data for a distance-running coaching assistant. Every pace,
split, and race-equivalent in the corpus is computed from a Daniels/Gilbert VDOT
implementation rather than written into a template, so the numbers are internally
consistent across all 1,500 examples.
Why this exists
Coaching corpora scraped from forums and blogs teach a model the register of
coaching without the arithmetic underneath it. A model that interpolates… See the full description on the dataset page: https://huggingface.co/datasets/hoodarunner/running-coach-sft.task926_coached_conv_pref_word_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task926_coached_conv_pref_word_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task926_coached_conv_pref_word_generation.coach_tabby_captionstask925_coached_conv_pref_classifier
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task925_coached_conv_pref_classifier
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task925_coached_conv_pref_classifier.sports-basketball-coach
This Dialogue
Comprised of fictitious examples of dialogues between a basketball coach and the players on the court during a game. Check out the example below:
"id": 1,
"description": "Motivating the team",
"dialogue": "Coach: Let's give it our all, team! We've trained hard for this game, and I know we can come out on top if we work together."
How to Load Dialogues
Loading dialogues can be accomplished using the fun dialogues library or Hugging Face datasets library.… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/sports-basketball-coach.running-coach-evalscoachtwin-workouts
CoachTwin Workouts
10,393 synthetic, structured workout plans, generated with an open
small language model and used to power the
CoachTwin app -
a workout recommender plus AI workout generator.
How it was generated
Generated with Qwen2.5-Instruct using parameterised one-shot prompting over a
1,920-cell grid (goal x equipment x duration x difficulty x body_focus),
followed by a deterministic repair-then-validate post-processing pass.
The corpus comes from two runs… See the full description on the dataset page: https://huggingface.co/datasets/OrDora/coachtwin-workouts.chess-coach-benchmark
Chess coach benchmark
A comprehensive, zero-leakage benchmark measuring the performance of local chess-coaching fine-tunes against frontier models.
What it is
This benchmark evaluates models on held-out chess positions, measuring their ability to provide tier-calibrated, engine-grounded chess coaching. It scores models based on deterministic objective metrics (move soundness, absence of engine jargon, and verification of board facts) alongside a blinded… See the full description on the dataset page: https://huggingface.co/datasets/khoilamalphaai/chess-coach-benchmark.chess-coach-move-review
Chess coach move-review SFT dataset
Supervised fine-tuning data for one specific, trained behavior: given a chess
position and the student's rating tier (Beginner, Intermediate, or Advanced),
select the tier-appropriate instructive move and tag it with a short principle,
for example "Nf3, develop toward the center."
That single move choice is the trained objective, and it is deterministically
checkable. The plain-English explanation rendered beside the move is a secondary… See the full description on the dataset page: https://huggingface.co/datasets/khoilamalphaai/chess-coach-move-review.CPR-Coach
Packages Check
The complete file list is shown below.
.
├── CPR_Dataset_S0.tar.gz.00 # 10 GB
├── CPR_Dataset_S0.tar.gz.01 # 10 GB
├── CPR_Dataset_S0.tar.gz.02 # 10 GB
├── CPR_Dataset_S0.tar.gz.03 # 10 GB
├── CPR_Dataset_S0.tar.gz.04 # 10 GB
├── CPR_Dataset_S0.tar.gz.05 # 10 GB
├── CPR_Dataset_S0.tar.gz.06 # 10 GB
├── CPR_Dataset_S0.tar.gz.07 # 5.4 GB
├── CPR_Dataset_S1.tar.gz.00 # 10… See the full description on the dataset page: https://huggingface.co/datasets/puriadityakumar/CPR-Coach.hbr-coaching-real-leadersTranscripts of HBR's Coaching Real Leaders podcasts, can be found here
chess-coach-grand-eval
Chess Coach — Grand Eval (comprehensive leaderboard)
One fresh, apples-to-apples comparison of every model in the chess move-review
coaching project — our tuned specialists, the untuned baselines, and the full frontier
lineup — on the same held-out validation slice (120 positions × 3 tiers
= 360 scenarios), scored with two independent layers:
Deterministic moat metrics (free, python-chess over pre-computed Stockfish/Maia
facts): tier-fit, distinct-moves-per-level… See the full description on the dataset page: https://huggingface.co/datasets/khoilamalphaai/chess-coach-grand-eval.chess-coach-turningpoints
Chess Coach – Turning Point Explanations Dataset
Overview
This repository contains a curated, engine-grounded dataset for training language models to explain chess mistakes and turning points in a human coaching style.
The goal is explainability and pedagogy, not move calculation or engine strength.
What this dataset is (and is not)
✅ This dataset is for
Training LLMs to explain evaluation swings
Teaching coaching tone, structure, and pedagogy… See the full description on the dataset page: https://huggingface.co/datasets/suman-kalavagunta/chess-coach-turningpoints.calm-coach-81ac5b
calm-coach-81ac5b
Synthetic sensors test data: 33 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-John655/calm-coach-81ac5b.interview-coach-dataset
Interview Coach Dataset
Chat-format dataset for fine-tuning an AI interview coach on software engineering interview Q&A.
Dataset Summary
Each example is a single user/assistant turn in OpenAI-style messages format, suitable for instruction / chat fine-tuning (e.g. Unsloth, TRL, Hugging Face SFTTrainer).
Train: ~1,017 examples
Validation: ~114 examples
Total: ~1,131 examples
Split: 90% / 10% (seeded shuffle)
Data Structure
{
"messages": [… See the full description on the dataset page: https://huggingface.co/datasets/shimogerald/interview-coach-dataset.chess-coach-v6
Chess Coach v6 (deep-verified training labels)
The current data frontier for the chess-instructor-llm coach: a foundational,
data-first rebuild of the training LABELS (the move plus full provenance), deep-verified
with Stockfish 17 (a two-depth root search with agreement bands), Syzygy tablebases
(endgames of seven pieces or fewer), and Maia-2 human-likelihood. It feeds the
downstream preference (DPO) and engine-distillation retrains.
This dataset is NOT the shipped SFT set. The… See the full description on the dataset page: https://huggingface.co/datasets/khoilamalphaai/chess-coach-v6.Health_Coach_Assistant_Data
Health Coach Assistant Dataset
This dataset consists of data for different types of chat with a health coach assistant about setting or updating the goals for walking.
Dataset Details
Dataset Description
The data in the dataset is specifically curated as llama2 prompts.
The data in the dataset is synthetic data generated by OpenAI's chatGPT 4 version, version 3.5, Github Copilot, and Claude AI.
Curated by: Sai Sangameswara Aadithya Kanduri
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Aadithya18/Health_Coach_Assistant_Data.tiny-dispatch-coach-traces
Tiny Dispatch Coach Traces
This dataset shares the sanitized build trace for Tiny Dispatch Coach, a Build
Small Hackathon project.
The trace records the model/planner design:
OpenBMB MiniCPM5-1B-GGUF parses dispatcher notes into constraints when the
optional llama.cpp path is enabled.
A deterministic planner computes route splits, time windows, wait time,
lateness, and baseline deltas.
The sample data is synthetic.
No API keys, user emails, real customer records, company… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/tiny-dispatch-coach-traces.garmin-mexican-fitness-coach-sft
🇲🇽 Garmin Mexican Fitness Coach SFT Dataset
Dataset sintético de alta fidelidad para el ajuste fino instruccional (Supervised Fine-Tuning / LoRA) de modelos de lenguaje pequeños (e.g., Llama 3.2 1B/3B, Qwen 2.5 1.5B/3B, SmolLM2), diseñado para dotarlos de la personalidad, modismos y tono enérgico de un Coach de Alto Rendimiento Mexicano con fundamento fisiológico estricto (Firstbeat Technologies & Garmin Connect).
📊 Resumen del Dataset
Tamaño total: 400… See the full description on the dataset page: https://huggingface.co/datasets/GerardoMayel/garmin-mexican-fitness-coach-sft.infinia-life-coach-dataset
Dataset Card for Infinia Life Coach Dataset
This dataset contains a collection of emotionally supportive, poetic, reflective conversational pairs designed for training AI models in warm, empathetic, non-clinical dialogue. Each entry includes a "prompt" expressing a vulnerable emotional state and a "completion" providing a gentle, metaphor-rich, grounding response. Topics include self-doubt, anxiety, overthinking, emotional regulation, and identity uncertainty.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Infiniaai/infinia-life-coach-dataset.slm-speech-coach
SLM Speech Coach — audio speech-coaching pairs
License & attribution
This dataset is released under CC-BY-4.0. It is a derivative work that combines:
Real speech from The People's Speech
(MLCommons), licensed CC-BY-4.0 — please retain this attribution when redistributing.
Synthetic child speech generated via TTS (OpenAI gpt-4o-mini-tts / Gemini), plus one
controlled delivery flaw injected per clip.
Coaching text written by a Gemini 3.1 teacher model… See the full description on the dataset page: https://huggingface.co/datasets/zsophia/slm-speech-coach.data_jobsclinical_nutritional_coach_formattedhabit-productivity-life-coach-datasetCoach-1.2kskolkovo-coaching-qa-4k
Skolkovo Executive Coaching Q&A Dataset
Описание
Датасет Q&A пар для обучения ИИ-коуча по программе Executive Coaching от Skolkovo.
Размер: 3,901 пар (893 оригинальных + 3,008 augmented)
Источники:
Административные документы Skolkovo
Лекции по коучингу
Книги по коучингу и лидерству (18 книг)
Кейсы и практические материалы
Структура данных
{
"question": "Как мотивировать команду?",
"answer": "Для мотивации команды важно...",
"role": "coach",
"source":… See the full description on the dataset page: https://huggingface.co/datasets/Ocheretny/skolkovo-coaching-qa-4k.accent_coach_all_waveforms
