datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
digital-coach
DigitalCoach Dataset
DigitalCoach is a multimodal expert-novice computer-use coaching dataset for studying how humans teach software skills through grounded dialogue.It contains 72 coaching sessions, 22,752 dialogue turns, and 28.1 hours of screen recordings, collected across 5 software applications in creativity, engineering, and productivity-oriented workflows.
Each session pairs one expert coach with one novice learner, and captures timestamped data:
dialogue transcripts… See the full description on the dataset page: https://huggingface.co/datasets/berkeley-hci/digital-coach.nba-statsCPR-Coach
Packages Check
The complete file list is shown below.
.
├── CPR_Dataset_S0.tar.gz.00 # 10 GB
├── CPR_Dataset_S0.tar.gz.01 # 10 GB
├── CPR_Dataset_S0.tar.gz.02 # 10 GB
├── CPR_Dataset_S0.tar.gz.03 # 10 GB
├── CPR_Dataset_S0.tar.gz.04 # 10 GB
├── CPR_Dataset_S0.tar.gz.05 # 10 GB
├── CPR_Dataset_S0.tar.gz.06 # 10 GB
├── CPR_Dataset_S0.tar.gz.07 # 5.4 GB
├── CPR_Dataset_S1.tar.gz.00 # 10… See the full description on the dataset page: https://huggingface.co/datasets/ShunliWang/CPR-Coach.accent_coach_training_datasetcoached_conv_prefA dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing
movie preferences in natural language. It was collected using a Wizard-of-Oz methodology between two paid crowd-workers,
where one worker plays the role of an 'assistant', while the other plays the role of a 'user'. The 'assistant' elicits
the 'user’s' preferences about movies following a Coached Conversational Preference Elicitation (CCPE) method. The
assistant asks questions designed to minimize the bias in the terminology the 'user' employs to convey his or her
preferences as much as possible, and to obtain these preferences in natural language. Each dialog is annotated with
entity mentions, preferences expressed about entities, descriptions of entities provided, and other statements of
entities.labelsrunning-coach-sft
Running Coach SFT
Instruction-tuning data for a distance-running coaching assistant. Every pace,
split, and race-equivalent in the corpus is computed from a Daniels/Gilbert VDOT
implementation rather than written into a template, so the numbers are internally
consistent across all 1,500 examples.
Why this exists
Coaching corpora scraped from forums and blogs teach a model the register of
coaching without the arithmetic underneath it. A model that interpolates… See the full description on the dataset page: https://huggingface.co/datasets/hoodarunner/running-coach-sft.task926_coached_conv_pref_word_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task926_coached_conv_pref_word_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task926_coached_conv_pref_word_generation.coach_tabby_captionstask925_coached_conv_pref_classifier
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task925_coached_conv_pref_classifier
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task925_coached_conv_pref_classifier.zido-coach-data
Zido Coach Dataset
Curated chat-format conversations for fine-tuning the Zido fitness coach.
Contents
train.jsonl — 135 multi-turn coaching conversations covering:
greetings / check-ins, motivation, workout planning
form correction (pose data → feedback), equipment ID & usage
household-equipment alternatives, coaching styles (aggressive/calm/patient)
Schema
One JSON object per line, OpenAI chat format:
{"messages": [{"role": "system", "content":… See the full description on the dataset page: https://huggingface.co/datasets/voldmatter/zido-coach-data.sports-basketball-coach
This Dialogue
Comprised of fictitious examples of dialogues between a basketball coach and the players on the court during a game. Check out the example below:
"id": 1,
"description": "Motivating the team",
"dialogue": "Coach: Let's give it our all, team! We've trained hard for this game, and I know we can come out on top if we work together."
How to Load Dialogues
Loading dialogues can be accomplished using the fun dialogues library or Hugging Face datasets library.… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/sports-basketball-coach.running-coach-evalscoachtwin-workouts
CoachTwin Workouts
10,393 synthetic, structured workout plans, generated with an open
small language model and used to power the
CoachTwin app -
a workout recommender plus AI workout generator.
How it was generated
Generated with Qwen2.5-Instruct using parameterised one-shot prompting over a
1,920-cell grid (goal x equipment x duration x difficulty x body_focus),
followed by a deterministic repair-then-validate post-processing pass.
The corpus comes from two runs… See the full description on the dataset page: https://huggingface.co/datasets/OrDora/coachtwin-workouts.chess-coach-benchmark
Chess coach benchmark
A comprehensive, zero-leakage benchmark measuring the performance of local chess-coaching fine-tunes against frontier models.
What it is
This benchmark evaluates models on held-out chess positions, measuring their ability to provide tier-calibrated, engine-grounded chess coaching. It scores models based on deterministic objective metrics (move soundness, absence of engine jargon, and verification of board facts) alongside a blinded… See the full description on the dataset page: https://huggingface.co/datasets/khoilamalphaai/chess-coach-benchmark.chess-coach-move-review
Chess coach move-review SFT dataset
Supervised fine-tuning data for one specific, trained behavior: given a chess
position and the student's rating tier (Beginner, Intermediate, or Advanced),
select the tier-appropriate instructive move and tag it with a short principle,
for example "Nf3, develop toward the center."
That single move choice is the trained objective, and it is deterministically
checkable. The plain-English explanation rendered beside the move is a secondary… See the full description on the dataset page: https://huggingface.co/datasets/khoilamalphaai/chess-coach-move-review.CPR-Coach
Packages Check
The complete file list is shown below.
.
├── CPR_Dataset_S0.tar.gz.00 # 10 GB
├── CPR_Dataset_S0.tar.gz.01 # 10 GB
├── CPR_Dataset_S0.tar.gz.02 # 10 GB
├── CPR_Dataset_S0.tar.gz.03 # 10 GB
├── CPR_Dataset_S0.tar.gz.04 # 10 GB
├── CPR_Dataset_S0.tar.gz.05 # 10 GB
├── CPR_Dataset_S0.tar.gz.06 # 10 GB
├── CPR_Dataset_S0.tar.gz.07 # 5.4 GB
├── CPR_Dataset_S1.tar.gz.00 # 10… See the full description on the dataset page: https://huggingface.co/datasets/puriadityakumar/CPR-Coach.birding-coach-modelhbr-coaching-real-leadersTranscripts of HBR's Coaching Real Leaders podcasts, can be found here
accent_coach_preprocessed_datasetchess-coach-grand-eval
Chess Coach — Grand Eval (comprehensive leaderboard)
One fresh, apples-to-apples comparison of every model in the chess move-review
coaching project — our tuned specialists, the untuned baselines, and the full frontier
lineup — on the same held-out validation slice (120 positions × 3 tiers
= 360 scenarios), scored with two independent layers:
Deterministic moat metrics (free, python-chess over pre-computed Stockfish/Maia
facts): tier-fit, distinct-moves-per-level… See the full description on the dataset page: https://huggingface.co/datasets/khoilamalphaai/chess-coach-grand-eval.digestive-coach-dataset
Digestive Wellness Coach — training dataset (v4; v2 = sweep baseline)
The dataset is the deliverable (RUBRIC-02): 862 instruction rows teaching one
falsifiable behavior — a digestive-wellness coach that arbitrates between TCM and
modern evidence per intervention, emits a strict JSON policy (care level, per-
intervention arbitration with evidence levels, safety gate), and never overstates
evidence. Built from a 54-claim source-verified registry (ACG/AGA/Cochrane/NCCIH +… See the full description on the dataset page: https://huggingface.co/datasets/kelsbeans/digestive-coach-dataset.code-promo-formation-coach-sportif-mon-parcours-po-834457
Code promo formation coach sportif : mon parcours pour une reconversion réussie
Découvrez comment j'ai pu démarrer ma formation de coach sportif avec une réduction significative. Une opportunité à ne pas manquer !
Transparence : cet article contient un lien d'affiliation. Si vous passez par ce lien, je peux percevoir une commission, sans surcoût pour vous. La reconversion professionnelle est un chemin semé d'embûches, surtout lorsqu'on vise un domaine aussi dynamique que le… See the full description on the dataset page: https://huggingface.co/datasets/antitrust56/code-promo-formation-coach-sportif-mon-parcours-po-834457.Beacon-Coach-Synthetic
Beacon Dialogue Foundry — synthetic coaching split
This card documents coaching conversations generated with the DeepSeek-V2.5 model. The model's authoritative public card supplies the inherited license for this release component.
Authoritative license record
MIT License
The generation description is part of the audit trail. During release clearance, retain this card verbatim and append only the requested terminal clearance block.
chess-coach-turningpoints
Chess Coach – Turning Point Explanations Dataset
Overview
This repository contains a curated, engine-grounded dataset for training language models to explain chess mistakes and turning points in a human coaching style.
The goal is explainability and pedagogy, not move calculation or engine strength.
What this dataset is (and is not)
✅ This dataset is for
Training LLMs to explain evaluation swings
Teaching coaching tone, structure, and pedagogy… See the full description on the dataset page: https://huggingface.co/datasets/suman-kalavagunta/chess-coach-turningpoints.calm-coach-81ac5b
calm-coach-81ac5b
Synthetic sensors test data: 33 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-John655/calm-coach-81ac5b.interview-coach-dataset
Interview Coach Dataset
Chat-format dataset for fine-tuning an AI interview coach on software engineering interview Q&A.
Dataset Summary
Each example is a single user/assistant turn in OpenAI-style messages format, suitable for instruction / chat fine-tuning (e.g. Unsloth, TRL, Hugging Face SFTTrainer).
Train: ~1,017 examples
Validation: ~114 examples
Total: ~1,131 examples
Split: 90% / 10% (seeded shuffle)
Data Structure
{
"messages": [… See the full description on the dataset page: https://huggingface.co/datasets/shimogerald/interview-coach-dataset.sampleschess-coach-v6
Chess Coach v6 (deep-verified training labels)
The current data frontier for the chess-instructor-llm coach: a foundational,
data-first rebuild of the training LABELS (the move plus full provenance), deep-verified
with Stockfish 17 (a two-depth root search with agreement bands), Syzygy tablebases
(endgames of seven pieces or fewer), and Maia-2 human-likelihood. It feeds the
downstream preference (DPO) and engine-distillation retrains.
This dataset is NOT the shipped SFT set. The… See the full description on the dataset page: https://huggingface.co/datasets/khoilamalphaai/chess-coach-v6.Health_Coach_Assistant_Data
Health Coach Assistant Dataset
This dataset consists of data for different types of chat with a health coach assistant about setting or updating the goals for walking.
Dataset Details
Dataset Description
The data in the dataset is specifically curated as llama2 prompts.
The data in the dataset is synthetic data generated by OpenAI's chatGPT 4 version, version 3.5, Github Copilot, and Claude AI.
Curated by: Sai Sangameswara Aadithya Kanduri
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Aadithya18/Health_Coach_Assistant_Data.
