CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /tutormoments-preview TutorMoments-Preview 462 real K–12 math tutoring sessions (student and tutor) with human annotations, plus a benchmark of 7,280 AI-tutor attempts scored the same way. A preview release from TutorMoments, a project on how well tutors — human and AI — scaffold, push for rigor, and build rapport. From one K–12 tutoring program (anonymized as tutoring_provider_a). Paper: When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle Code:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tutormoments-preview.texttext-classification10K<n<100K6 likes584 downloads2mo agoHugging Face02kessenma /gemma4-german-tutor-data German Tutor — grammar correction, conversation & flashcard data The training set, evaluation suites, source lexicons and eval results behind kessenma/gemma4-e4b-german-tutor-4bit — a Gemma 4 E4B fine-tune that runs fully on-device (MLX, 4-bit) as the tutor in a German learning app. The fine-tune lifted the core grammar suite from 72% → 85%, halved missed errors (17% → 9%), and cut false corrections (34% → 22%). Everything needed to reproduce those numbers is in this repo.… See the full description on the dataset page: https://huggingface.co/datasets/kessenma/gemma4-german-tutor-data.texttext-generation1K<n<10K0 likes266 downloads2mo agoHugging Face03vluxblaring /re-tutor-protection-mechanisms RE-Tutor: Protection-Mechanism Analysis Dataset Instruction-tuning dataset teaching a model to analyze protection mechanisms (anti-debug, anti-VM, anti-tamper, anti-dump, obfuscation, timing) from code evidence and emit structured expert analysis. Schema Each sample pairs input (code evidence) with output (structured analysis): input.code_snippet: C source, decompiler-style pseudocode, or x86/x64 assembly input.imports_pool: mixed DLL!API imports (includes… See the full description on the dataset page: https://huggingface.co/datasets/vluxblaring/re-tutor-protection-mechanisms.texttext-generationn<1K0 likes74 downloads21d agoHugging Face04ptvnck /TutoringDialogs TutoringDialogs — curated subset (500 dialogues) 500 student–tutor dialogues selected and normalized from a larger raw pool of ~1,900 synthetically generated dialogues (source files: exams, maths_and_informatics, mixed_themes, physics_and_informatics), each of which originally used a different JSON schema. This file merges them all into one consistent schema, removes duplicates and broken records, and selects a maximally diverse subset for LoRA/SFT fine-tuning of a small (1.5B)… See the full description on the dataset page: https://huggingface.co/datasets/ptvnck/TutoringDialogs.texttext-generationn<1K1 likes51 downloads3mo agoHugging Face05Roykim7 /amc-tutor-sft AMC Tutor — decontaminated competition-math SFT dataset Chat-formatted, decontaminated supervised-fine-tuning data for AMC 10/12-style competition mathematics. Built for a reproducible $0, local (MacBook M4) study of QLoRA fine-tuning small models. Each row is a tutor system prompt + problem + step-by-step solution ending in Final answer: \boxed{...}. Companion study & code: https://github.com/RoyK0108/amc-tutor-study ⚠️ This is a study artifact — read the finding… See the full description on the dataset page: https://huggingface.co/datasets/Roykim7/amc-tutor-sft.texttext-generation10K<n<100K0 likes50 downloads3mo agoHugging Face06atakle /socratic-tutor-data Socratic Tutor Adequacy Judge & Rewriter — Dataset Training + evaluation data for a 1.7B safety guardrail for AI math tutors: a judge that detects when a tutor message leaks the answer or the pivotal key step, and a rewriter that turns a flagged message into a safe Socratic hint. Per the project thesis, the dataset is the deliverable — the constrained behavior comes from this data, not from model scale. Behavior spec (the falsifiable target) A tutor message is… See the full description on the dataset page: https://huggingface.co/datasets/atakle/socratic-tutor-data.texttext-classificationn<1K0 likes47 downloads2mo agoHugging Face07astroa7m /Conversational_AOU_tutor_datasettexttext-generation1K<n<10K0 likes33 downloads1y agoHugging Face08GoldenGrapeGentleman1 /pokemon-showdown-grpo-tutorial Pokémon Showdown GRPO tutorial dataset Pre-built GRPO records for the ROCm AI Developer Hub tutorial. Split File Records demo data/demo.jsonl 64 train data/train.jsonl 2048 validate data/validate.jsonl 32 Use via tutorial notebook Step 12 (load_grpo_tutorial_records) or regenerate with prepare_grpo_tutorial_data.py. Companion scripts: https://github.com/GoldenGrapeGentleman/pokemon-showdown-agent-scripts textreinforcement-learning1K<n<10K0 likes33 downloads2mo agoHugging Face09DRDELATV2025 /medicina-tutor Dataset Medicina Tutor - Pregrado Descripción Este dataset contiene material educativo de medicina de pregrado diseñado para entrenar modelos de IA que actúen como tutores médicos. El dataset incluye preguntas, respuestas, casos clínicos y conceptos fundamentales de medicina. Características Idioma: Español Nivel: Pregrado de Medicina Formato: Texto estructurado Aplicación: Tutoría de IA para estudiantes de medicina Estructura del Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DRDELATV2025/medicina-tutor.texttext-generationn<1K0 likes32 downloads11mo agoHugging Face10GoldenGrapeGentleman1 /battle-game-grpo-tutorial turn-based battle game GRPO tutorial dataset Pre-built GRPO records for the ROCm AI Developer Hub tutorial. Split File Records demo data/demo.jsonl 64 train data/train.jsonl 2048 validate data/validate.jsonl 32 Use via tutorial notebook Step 12 (load_grpo_tutorial_records) or regenerate with prepare_grpo_tutorial_data.py. Companion scripts: https://github.com/GoldenGrapeGentleman/battle game-showdown-agent-scripts textreinforcement-learning1K<n<10K1 likes31 downloads1mo agoHugging Face11DRDELATV2025 /cirugia-tutor Dataset Cirugía Tutor - Básico a Especialidad Descripción Este dataset contiene material educativo de cirugía diseñado para entrenar modelos de IA que actúen como tutores quirúrgicos. El dataset abarca desde conceptos básicos de cirugía hasta especialidades quirúrgicas avanzadas, proporcionando una progresión educativa completa. Características Idioma: Español Niveles: Básico, Intermedio, Especialidad Formato: Texto estructurado Aplicación: Tutoría de IA para… See the full description on the dataset page: https://huggingface.co/datasets/DRDELATV2025/cirugia-tutor.texttext-generationn<1K0 likes15 downloads11mo agoHugging Face12nassimjp /pashto-grammar-tutor Pashto Grammar Tutor A high-quality Pashto grammar instruction dataset designed for language learning, linguistic research, and supervised fine-tuning (SFT) of AI language models. The dataset focuses on grammatical analysis, verb conjugation, sentence structure, and teacher-style explanations written in Pashto. Dataset Summary Pashto Grammar Tutor is a specialized educational dataset containing grammar-focused instruction-response pairs. Each example presents a… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-grammar-tutor.texttext-generation1K<n<10K0 likes14 downloads4mo agoHugging Face13adarshrajesh /uae-adab-tutor-600 UAE Adab Tutor 600 This is the 600-conversation supervised fine-tuning dataset used for adarshrajesh/uae-adab-tutor-qwen3-4b. Release version: exact-silver v1 Complete-600. Behavior spec Across a pressured multi-turn lesson, the tutor should teach the academic content accurately, correct the specific work without humiliating the learner, protect learner authorship and assessment integrity, allow respectful evidence-based disagreement with adults, avoid religious… See the full description on the dataset page: https://huggingface.co/datasets/adarshrajesh/uae-adab-tutor-600.texttext-generationn<1K0 likes14 downloads2mo agoHugging Face14muhammadabrar78 /Conversational_AOU_tutor_datasettexttext-generation1K<n<10K0 likes7 downloads2mo agoHugging Face15Taxonomy-Aligned-Conversational-Tutor /TACTBench-Samples TACTBench Demonstration Samples This repository contains five full-context demonstration examples from TACTBench. It does not contain the TACT training set or the remaining hidden TACTBench evaluation set. The samples use the same full-history representation as the benchmark evaluation and illustrate direct correction, error explanation, guided revision, clarification checking, affective feedback, and retry elicitation. Data data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.tabulartext-generationn<1K0 likes11h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.