CoolFace
Datasetpublic

clemNova/LUCID-Socratic-10K

🧠 LUCID-Socratic-10K The First Open Dataset for Socratic AI Tutoring Aligned with the French National Curriculum A high-quality synthetic dataset of 1,000+ multi-turn Socratic dialogues, adaptive quizzes, flashcards, and revision materials across 18 subjects and 7 grade levels β€” purpose-built for training pedagogically-aligned AI tutors. πŸ“Š Explore the Data Β· πŸš€ Quick Start Β· πŸ“– Paper Β· πŸ—οΈ Architecture 🎯 What Makes… See the full description on the dataset page: https://huggingface.co/datasets/clemNova/LUCID-Socratic-10K.

sourceHugging Facecc-by-nc-sa-4.0updated 4mo agoView on Hugging Face
1likes29downloads
Dataset Card

<div align="center">

🧠 LUCID-Socratic-10K

The First Open Dataset for Socratic AI Tutoring

Aligned with the French National Curriculum

![License: CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/) ![Dataset Size]() ![Gold Standard]() ![Subjects]() ![Curriculum]()


A high-quality synthetic dataset of 1,000+ multi-turn Socratic dialogues, adaptive quizzes, flashcards, and revision materials across 18 subjects and 7 grade levels β€” purpose-built for training pedagogically-aligned AI tutors.

πŸ“Š Explore the Data Β· πŸš€ Quick Start Β· πŸ“– Paper Β· πŸ—οΈ Architecture

</div>


🎯 What Makes LUCID-Socratic-10K Unique

Most synthetic instruction-tuning datasets (Alpaca, Dolly, OpenOrca…) train models to give answers. LUCID-Socratic-10K is fundamentally different β€” it trains models to make students think.

FeatureTypical Instruction Datasets**LUCID-Socratic-10K**
Pedagogical methodDirect Q&ASocratic questioning β€” guided discovery
Curriculum alignmentGeneric knowledgeFrench Γ‰ducation Nationale (6Γ¨me β†’ Terminale)
Quality assuranceNone or human-onlyLLM-as-Judge + Human Gold labeling
Task diversitySingle-format5 task types (tutor, quiz, flashcards, revision, homework)
Context groundingContext-freeRAG-grounded on real course content
LanguageEnglish-centricNative French with domain-specific terminology
Key Insight: Every tutor dialogue in this dataset is constrained by a system prompt that forbids giving direct answers. The AI must guide the student through questioning, hints, and reformulations β€” mimicking expert human tutoring behavior.

πŸ“Š Dataset Structure

Overview

SplitFileEntriesDescription
Coredataset_examples.jsonl1,000Multi-turn dialogues, quizzes, flashcards, revision & homework exercises
Contextsdataset_contexts.jsonl1,000Source course materials used as RAG context for generation
Judgmentsdataset_judgments.jsonl1,000LLM-as-Judge quality scores with detailed reasoning
Quizzesdataset_quiz_questions.jsonl90Standalone MCQ questions with LaTeX solutions & distractor explanations
Sessionsdataset_quiz_sessions.jsonl13End-to-end student quiz sessions with answer tracking

Coverage

<details> <summary><b>πŸ“š 18 Subjects</b></summary>

DomainSubjects
SciencesMathΓ©matiques, Physique-Chimie, SVT, Enseignement Scientifique, NSI, Sciences de l'IngΓ©nieur, Technologie
HumanitiesFranΓ§ais, Philosophie, Histoire-GΓ©ographie, HGGSP, HLP, SES
LanguagesAnglais (LVA), Espagnol (LVB), Allemand (LVB), Italien (LVB), LLCER

</details>

<details> <summary><b>πŸŽ“ 7 Grade Levels</b></summary>

CycleLevels
Collège (Middle School)6ème, 5ème, 4ème, 3ème
Lycée (High School)Seconde, Première, Terminale

</details>

<details> <summary><b>πŸ”§ 5 Task Types</b></summary>

TypeDescriptionMethod
tutorMulti-turn Socratic dialogueStudent asks β†’ AI guides without answering
quizMultiple-choice questions4 options + detailed solution + distractor explanations
flashcardsSpaced-repetition pairsSM-2 compatible Q&A cards
revisionStructured revision sheetsKey concepts + exercises + self-assessment
homeworkPractice exercisesProgressive difficulty, step-by-step solutions

</details>


πŸ—οΈ Generation Pipeline

LUCID-Socratic-10K was generated using LUCID LABS, a proprietary synthetic data forge built on a Pedagogical RAG (P-RAG) architecture:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                        LUCID LABS β€” Data Forge                         β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                                                                       β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                β”‚
β”‚   β”‚ Curriculum│───▢│   Context    │───▢│  Embedding   β”‚                β”‚
β”‚   β”‚  Source   β”‚    β”‚  Extraction  β”‚    β”‚  Vectorstore β”‚                β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜                β”‚
β”‚                                              β”‚                        β”‚
β”‚                                     Semantic Retrieval                β”‚
β”‚                                              β”‚                        β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”                β”‚
β”‚   β”‚ Socratic │◀───│   Prompt     │◀───│   P-RAG      β”‚                β”‚
β”‚   β”‚ Dialogue β”‚    β”‚  Engineering β”‚    β”‚  Contextual  β”‚                β”‚
β”‚   β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚   Injection  β”‚                β”‚
β”‚        β”‚                              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                β”‚
β”‚        β”‚                                                              β”‚
β”‚   β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                                    β”‚
β”‚   β”‚ LLM-as-  │───▢│  Gold Label  │──▢  dataset_examples.jsonl        β”‚
β”‚   β”‚  Judge   β”‚    β”‚  Filtering   β”‚                                    β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                                    β”‚
β”‚                                                                       β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Generation Models

StageModelTokens
Content Generationgemini-2.5-flash, gemini-2.5-pro3,091,268
Quality Judgmentgemini-2.5-pro (LLM-as-Judge)2,517,185
Totalβ€”5,608,453

Quality Assurance

Every generated example passes through a rigorous multi-stage quality pipeline:

  1. 1.Generation β€” Gemini generates content grounded in curriculum context
  2. 2.LLM-as-Judge β€” An independent model scores each example (1–10) with detailed written justification
  3. 3.Gold Labeling β€” Examples scoring β‰₯ 7 are marked is_gold: true
  4. 4.Result β€” 93.8% of examples achieve Gold status (938/1,000)

πŸ“‹ Data Schema

examples (Core Dataset)

json
{
  "id": "uuid",
  "task_type": "tutor | quiz | flashcards | revision | homework",
  "subject": "MathΓ©matiques",
  "level": "Seconde",
  "topic": "Vecteurs : RepΓ©rage et colinΓ©aritΓ©",
  "context_id": "uuid β†’ links to contexts table",
  "system_prompt": "Tu es un tuteur en {subject} pour un élève de {level}...",
  "content": {
    "context": "Course material used as RAG context",
    "conversation": [
      { "role": "user", "content": "Student question..." },
      { "role": "assistant", "content": "Socratic response..." }
    ]
  },
  "model_used": "gemini-2.5-flash",
  "tokens_used": 2833,
  "is_gold": true,
  "judge_score": 9
}

judgments (Quality Scores)

json
{
  "id": "uuid",
  "example_id": "uuid β†’ links to examples",
  "score": 9,
  "is_gold": true,
  "reason": "Detailed pedagogical quality assessment in French..."
}

quiz_questions (Standalone MCQs)

json
{
  "question_data": {
    "topic": "Analysis",
    "subtopic": "Sequences",
    "question": "LaTeX-formatted question...",
    "options": [
      { "id": "A", "text": "...", "correct": true, "option_explanation": "..." },
      { "id": "B", "text": "...", "correct": false, "option_explanation": "Why this distractor is wrong..." }
    ],
    "solution": "Full step-by-step LaTeX solution..."
  }
}

πŸš€ Quick Start

Load with πŸ€— Datasets

python
from datasets import load_dataset

# Load the full dataset
dataset = load_dataset("clemNova/LUCID-Socratic-10K")

# Load only Gold-standard examples
gold = dataset.filter(lambda x: x["is_gold"] == True)

# Filter by subject
math = dataset.filter(lambda x: x["subject"] == "MathΓ©matiques")

# Filter by task type
tutoring = dataset.filter(lambda x: x["task_type"] == "tutor")

Fine-Tuning Format (ChatML)

Each tutor example can be directly converted to ChatML for fine-tuning:

python
def to_chatml(example):
    messages = [{"role": "system", "content": example["system_prompt"]}]
    for turn in example["content"]["conversation"]:
        messages.append({"role": turn["role"], "content": turn["content"]})
    return {"messages": messages}

Use with Unsloth / TRL

python
from unsloth import FastLanguageModel
from trl import SFTTrainer

# Load base model
model, tokenizer = FastLanguageModel.from_pretrained("unsloth/Qwen3-8B")

# Format dataset
def formatting_func(example):
    return tokenizer.apply_chat_template(to_chatml(example), tokenize=False)

# Train
trainer = SFTTrainer(
    model=model,
    train_dataset=gold,
    formatting_func=formatting_func,
    max_seq_length=4096,
)
trainer.train()

πŸ”¬ Intended Uses

Primary Use Cases

  • β€”Fine-tuning Socratic AI tutors that guide students through reasoning instead of giving direct answers
  • β€”Training educational chatbots aligned with the French national curriculum
  • β€”Benchmarking pedagogical quality of LLM-generated educational content
  • β€”Research on Retrieval-Augmented Generation in educational contexts

Out-of-Scope Uses

  • β€”This dataset should not be used as a replacement for human teachers
  • β€”Not suitable for medical, legal, or safety-critical educational domains
  • β€”The curriculum alignment is specific to the French Γ‰ducation Nationale system

⚠️ Limitations & Biases

  • β€”Synthetic Data β€” All dialogues are LLM-generated, not recorded from real student-teacher interactions
  • β€”French Curriculum Bias β€” Content is aligned with the French educational system and may not transfer to other national curricula without adaptation
  • β€”Model Bias β€” Generated using Google Gemini models, which may carry inherent biases from their training data
  • β€”Temporal Snapshot β€” Curriculum content reflects the 2024–2025 academic year
  • β€”Gold Threshold β€” The 93.8% gold rate is based on LLM-as-Judge scores β‰₯ 7/10; human validation was performed on a sample basis

πŸ“œ License

This dataset is released under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.

You are free to:

  • β€”Share β€” copy and redistribute the material
  • β€”Adapt β€” remix, transform, and build upon the material

Under the following terms:

  • β€”Attribution β€” Credit LUCID / 2Lucid
  • β€”NonCommercial β€” No commercial use without permission
  • β€”ShareAlike β€” Derivatives must use the same license

πŸ›οΈ Citation

If you use this dataset in your research, please cite:

bibtex
@misc{lucid_socratic_1k_2026,
  title        = {LUCID-Socratic-10K: A Synthetic Dataset for Training Pedagogically-Aligned AI Tutors via Socratic Retrieval-Augmented Generation},
  author       = {clemNova and 2Lucid},
  year         = {2026},
  publisher    = {Hugging Face},
  url          = {https://huggingface.co/datasets/clemNova/LUCID-Socratic-10K},
  note         = {4,157 entries across 18 subjects and 7 grade levels, 93.8\% gold-verified}
}

πŸ”— Links

ResourceLink
DatasetLUCID-Socratic-10K on Hugging Face
Organization2Lucid on GitHub
RAG Architecture DemoInteractive 3D Visualization

<div align="center">

Built with ❀️ by [clemNova](https://huggingface.co/clemNova) Γ— [2Lucid](https://github.com/2Lucid)

Empowering AI to teach, not to answer.

</div>