datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fable-5-premium
🧠 Fable-5 Premium Dataset
🚀 V2 is out! This dataset has a successor: fable-5-premium-v2 — new users should start there.
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Records
12,730
Train Split
5,728 (45.0%)
Validation Split
318 (2.5%)
Test Split
319 (2.5%)
Created
2026-07-30… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium.fable-5-sft-traces
Fable-5 SFT Traces
Author / maintainer: kelexine (github.com/kelexine)
A cleaned, anonymised, schema-normalised derivative of
Kelexine/Fable-5-traces
— agentic traces from Fable-5 (claude-fable-5), the model now publicly
known as Claude Mythos — Anthropic's top-of-family frontier model at time
of collection.
The dataset supports three fine-tuning shapes off a single JSONL with no
preprocessing required:
Mode
Fields used
Full SFT (thinking + response)
messages or… See the full description on the dataset page: https://huggingface.co/datasets/kelexine/fable-5-sft-traces.Sumtables-Cuniform-Small-Fable5-Remaster-v2
Sumtablets Cuneiform — Small (Fable5 Remaster v2)
This is the exact multi-task training/eval/test pack used to fine-tune
Qwen3-VL-4B-Instruct to visually read, transliterate, and translate
Sumerian cuneiform tablets from photographs (see the companion model repo,
linked below). It is a fixed, reproducible subset drawn from the full
master dataset:
Master dataset (455,506 records, 70.5 GB, all phases, full documentation):
TRACCERR/Sumtables-Cuneiform-Full-Fable5-Remaster
If you… See the full description on the dataset page: https://huggingface.co/datasets/TRACCERR/Sumtables-Cuniform-Small-Fable5-Remaster-v2.fable-5-premium-v2
🧠 Fable-5 Premium V2
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 100,000 agent traces, built for training tool-using models. Successor to fable-5-premium.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
100,000
Train Split
85,000 (85.0%)
Validation Split
7,500 (7.5%)
Test Split
7,500 (7.5%)
Average Quality
0.966 (0.8–1.0 band)
Distilled From
Claude… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium-v2.fable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,380 TRAJECTORIES · 12,490 TRAINING ROWS · 14 MB PARQUET · 663 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/DSFFGFG456/fable-5-coding-and-debugging-traces.fable-5.1-premium
🧠 Fable-5.1 Premium
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 4,996 Fable 5.1 max-reasoning agent traces, built for training tool-using and long-horizon reasoning models. Third entry in the Premium series, upholding the standards of fable-5-premium and fable-5-premium-v2.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
4,996
Train Split
4,245 (85.0%)
Validation… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5.1-premium.Sumtables-Cuneiform-Full-Fable5-Remaster
Sumtablets-Cuneiform-Full-Fable5-Remaster — Cuneiform Vision-Language Training Dataset
A rebuilt, leakage-proof, multi-task training dataset for teaching
vision-language models (target: Qwen3-VL-8B-Instruct LoRA) to visually
read, transliterate, and translate Sumerian cuneiform tablets from
photographs. The mission: produce useful first-pass readings for the ~90% of
excavated tablets that have never been published or translated.
Current release: v1.0.0 — 455,506 records (402,004… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/Sumtables-Cuneiform-Full-Fable5-Remaster.fable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,374 TRAJECTORIES · 12,448 TRAINING ROWS · 14 MB PARQUET · 662 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/fable-5-coding-and-debugging-traces.Sumtables-Cuneiform-Full-Fable5-Remaster
Sumtablets-Cuneiform-Full-Fable5-Remaster — Cuneiform Vision-Language Training Dataset
A rebuilt, leakage-proof, multi-task training dataset for teaching
vision-language models (target: Qwen3-VL-8B-Instruct LoRA) to visually
read, transliterate, and translate Sumerian cuneiform tablets from
photographs. The mission: produce useful first-pass readings for the ~90% of
excavated tablets that have never been published or translated.
Current release: v1.0.0 — 455,506 records (402,004… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Sumtables-Cuneiform-Full-Fable5-Remaster.fable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/siddharth0713/fable-5-coding-and-debugging-traces.fable5-traces-agentic-cleanfable-5-premium
🧠 Fable-5 Premium Dataset
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Records
12,730
Train Split
5,728 (45.0%)
Validation Split
318 (2.5%)
Test Split
319 (2.5%)
Created
2026-07-30
License
MIT
📦 Formats Available
This dataset is available in two formats:… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/fable-5-premium.Sumtables-Cuneiform-Full-Fable5-Remaster
Sumtablets-Cuneiform-Full-Fable5-Remaster — Cuneiform Vision-Language Training Dataset
A rebuilt, leakage-proof, multi-task training dataset for teaching
vision-language models (target: Qwen3-VL-8B-Instruct LoRA) to visually
read, transliterate, and translate Sumerian cuneiform tablets from
photographs. The mission: produce useful first-pass readings for the ~90% of
excavated tablets that have never been published or translated.
Current release: v1.0.0 — 455,506 records (402,004… See the full description on the dataset page: https://huggingface.co/datasets/TRACCERR/Sumtables-Cuneiform-Full-Fable5-Remaster.fable5-traces-agentic-clean-v2fable-5-traces-messages
Fable 5 Traces — messages format
A conversion of Glint-Research/Fable-5-traces
(fable5_cot_merged.jsonl) into the standard HF/OpenAI conversational messages format, ready to drop
into Unsloth or TRL SFTTrainer notebooks with no extra conversion step.
Format
Each row:
{
"messages": [
{"role": "user", "content": "<flattened prior transcript / context>"},
{"role": "assistant", "content": "<think>...cot...</think>\nASSISTANT (tool call) Read input={...}"}… See the full description on the dataset page: https://huggingface.co/datasets/rex099/fable-5-traces-messages.fable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/fable-5-coding-and-debugging-traces.fable5-traces-agentic
fable5-traces-agentic
100K-max multi-domain coding + agentic SFT dataset.
Final rows
100,000
Targets
{
"agentic": 40000,
"coding": 20000,
"frontend": 11000,
"backend": 9000,
"reasoning": 15000,
"tools": 5000
}
Actual
{
"reasoning": 10321,
"frontend": 5384,
"tools": 5000,
"agentic": 50295,
"coding": 20000,
"backend": 9000
}
FABLE.5
Selected: 45,625
Target: 30,000
Minimum: 25,000
Build… See the full description on the dataset page: https://huggingface.co/datasets/usernamebetter/fable5-traces-agentic.fable5-flamingo-research-trace
Fable 5 FLAMINGO Research Task Traces
This dataset is a small corpus of Claude Fable 5 / Claude Code traces from
real FLAMINGO cosmology-analysis work. It is organized by explicit research task
so the Dataset Viewer makes clear what was being operated on and how the agent
took actions.
Start with the task_index config. It lists each task, the exact original user
prompt when available, the model setup, what Claude did, the raw trace file, and
the corresponding viewer configs.… See the full description on the dataset page: https://huggingface.co/datasets/licongxu/fable5-flamingo-research-trace.fable-5-coding-and-debugging-traces
Mirror: greghavens/fable-5-coding-and-debugging-traces
Pinned snapshot / mirror of greghavens/fable-5-coding-and-debugging-traces, re-hosted for PROTISEC
research reproducibility. Redistributed under the upstream license (cc-by-4.0)
with attribution — all credit to the original author.
Original author: greghavens
Source dataset: greghavens/fable-5-coding-and-debugging-traces
License: cc-by-4.0
Family: coding_traces
Mode: stream
Rows cached: 11487
Changes vs upstream: cached… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/fable-5-coding-and-debugging-traces.fable-5-sft-traces
Fable-5 SFT Traces
Author / maintainer: kelexine (github.com/kelexine)
A cleaned, anonymised, schema-normalised derivative of
Kelexine/Fable-5-traces
— agentic traces from Fable-5 (claude-fable-5), the model now publicly
known as Claude Mythos — Anthropic's top-of-family frontier model at time
of collection.
The dataset supports three fine-tuning shapes off a single JSONL with no
preprocessing required:
Mode
Fields used
Full SFT (thinking + response)
messages or… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/fable-5-sft-traces.fable-5-max-reasoning-filtered
Mirror: MoreThought/Fable-5-Max-Reasoning-Filtered
Pinned snapshot / mirror of MoreThought/Fable-5-Max-Reasoning-Filtered, re-hosted for PROTISEC
research reproducibility. Redistributed under the upstream license (apache-2.0)
with attribution — all credit to the original author.
Original author: MoreThought
Source dataset: MoreThought/Fable-5-Max-Reasoning-Filtered
License: apache-2.0
Family: fable
Mode: stream
Rows cached: 128
Changes vs upstream: cached snapshot, possibly… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/fable-5-max-reasoning-filtered.Fable-5-traces
Dataet format
Dataset({
features: ['harness', 'session_id', 'traces', 'file_path'],
num_rows: 63
})
fable5-chatml
Usage
Conversations are ChatML messages lists. Tool calls are written inline in the
assistant content as MiniCPM5-style XML, so the chat template applies directly:
from datasets import load_dataset
from transformers import AutoTokenizer
ds = load_dataset("koshuro/fable5-chatml", split="train")
tok = AutoTokenizer.from_pretrained("openbmb/MiniCPM5-1B")
text = tok.apply_chat_template(ds[0]["messages"], tokenize=False)
Assistant chain-of-thought is wrapped in <think>…</think>.… See the full description on the dataset page: https://huggingface.co/datasets/koshuro/fable5-chatml.
