datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fable-5-premium
🧠 Fable-5 Premium Dataset
🚀 V2 is out! This dataset has a successor: fable-5-premium-v2 — new users should start there.
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Records
12,730
Train Split
5,728 (45.0%)
Validation Split
318 (2.5%)
Test Split
319 (2.5%)
Created
2026-07-30… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium.Fable-GPT-5.5-Distillation-Traces
Agent Traces Curated 2026 (v3 Merged)
A unified distillation corpus of 9,057,143 records spanning agentic
coding traces, math/code/science reasoning, tool-use trajectories, and
preference data. 8,876,012 train + 181,131 eval, stratified by source.
What this is
This is the v3 merged corpus that supersedes both v1 and v2 of this dataset.
It combines five major source groups through a unified normalization
pipeline:
Original v2 RESMP-DEV (de-fragmented, re-deduped):… See the full description on the dataset page: https://huggingface.co/datasets/RESMP-DEV/Fable-GPT-5.5-Distillation-Traces.Complete-FABLE.5-traces-2M
license: mit
pretty_name: Claude Library — Fable 5 · Opus · Sonnet
annotations_creators:
machine-generated
language:
en
language_creators:
found
machine-generated
multilinguality:
monolingual
size_categories:
10K<n<100K
task_categories:
text-generation
task_ids:
language-modeling
tags:
agent-traces
claude
claude-fable-5
claude-opus
claude-sonnet
chain-of-thought
tool-use
coding-agents
content-verified
maintained-mirror
deduplicated
parquet
configs:
config_name:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Complete-FABLE.5-traces-2M.fable-5-sft-traces
Fable-5 SFT Traces
Author / maintainer: kelexine (github.com/kelexine)
A cleaned, anonymised, schema-normalised derivative of
Kelexine/Fable-5-traces
— agentic traces from Fable-5 (claude-fable-5), the model now publicly
known as Claude Mythos — Anthropic's top-of-family frontier model at time
of collection.
The dataset supports three fine-tuning shapes off a single JSONL with no
preprocessing required:
Mode
Fields used
Full SFT (thinking + response)
messages or… See the full description on the dataset page: https://huggingface.co/datasets/kelexine/fable-5-sft-traces.Sumtables-Cuniform-Small-Fable5-Remaster-v2
Sumtablets Cuneiform — Small (Fable5 Remaster v2)
This is the exact multi-task training/eval/test pack used to fine-tune
Qwen3-VL-4B-Instruct to visually read, transliterate, and translate
Sumerian cuneiform tablets from photographs (see the companion model repo,
linked below). It is a fixed, reproducible subset drawn from the full
master dataset:
Master dataset (455,506 records, 70.5 GB, all phases, full documentation):
TRACCERR/Sumtables-Cuneiform-Full-Fable5-Remaster
If you… See the full description on the dataset page: https://huggingface.co/datasets/TRACCERR/Sumtables-Cuniform-Small-Fable5-Remaster-v2.fable-5-premium-v2
🧠 Fable-5 Premium V2
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 100,000 agent traces, built for training tool-using models. Successor to fable-5-premium.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
100,000
Train Split
85,000 (85.0%)
Validation Split
7,500 (7.5%)
Test Split
7,500 (7.5%)
Average Quality
0.966 (0.8–1.0 band)
Distilled From
Claude… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium-v2.fable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,380 TRAJECTORIES · 12,490 TRAINING ROWS · 14 MB PARQUET · 663 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/DSFFGFG456/fable-5-coding-and-debugging-traces.agentic-distill-fable-5-sft
Fable-5 SFT — prepared for Qwable fine-tuning
4,659 single-turn pairs from Claude Fable-5 (Anthropic preview model, suspended globally 2026-06-22 under U.S. export-control directives), reformatted into a single-text-column parquet ready for SFTTrainer(dataset_text_field="text") + train_on_responses_only.
Composition:
3,793 rows (81%) end in a <tool_use> block — agentic tool-call patterns
866 rows (19%) end in a pure text response
This is agentic data, not pure reasoning data.… See the full description on the dataset page: https://huggingface.co/datasets/lordx64/agentic-distill-fable-5-sft.fable-5.1-premium
🧠 Fable-5.1 Premium
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 4,996 Fable 5.1 max-reasoning agent traces, built for training tool-using and long-horizon reasoning models. Third entry in the Premium series, upholding the standards of fable-5-premium and fable-5-premium-v2.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
4,996
Train Split
4,245 (85.0%)
Validation… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5.1-premium.Complete-FABLE.5-traces-2M
Complete FABLE.5 Traces (2 Million Deduplicated Rows)
Comprehensive Agentic Coding & Frontier Reasoning Trajectory Corpus
Executive Summary
Solstice-AI/Complete-FABLE.5-traces-2M is a clean, fully deduplicated post-training dataset containing 2,006,487 high-entropy agentic coding and multi-step reasoning traces.
Originally curated following the closure of Fable and Mythos, this corpus synthesizes frontier agent execution patterns (including Claude… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Complete-FABLE.5-traces-2M.Sumtables-Cuneiform-Full-Fable5-Remaster
Sumtablets-Cuneiform-Full-Fable5-Remaster — Cuneiform Vision-Language Training Dataset
A rebuilt, leakage-proof, multi-task training dataset for teaching
vision-language models (target: Qwen3-VL-8B-Instruct LoRA) to visually
read, transliterate, and translate Sumerian cuneiform tablets from
photographs. The mission: produce useful first-pass readings for the ~90% of
excavated tablets that have never been published or translated.
Current release: v1.0.0 — 455,506 records (402,004… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/Sumtables-Cuneiform-Full-Fable5-Remaster.fable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,374 TRAJECTORIES · 12,448 TRAINING ROWS · 14 MB PARQUET · 662 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/fable-5-coding-and-debugging-traces.fable-sft-combined-v2
Fable-5 SFT Combined (v2) — prepared for Qwable-v2 fine-tuning
9842 rows = union of two Fable-5 SFT corpora, deduplicated by user-side content:
lordx64/agentic-distill-fable-5-sft — 4,659 rows WITH <think> blocks (cleartext reasoning added post-hoc by Glint-Research)
lordx64/fable-tool-use-sft — 5,183 rows WITHOUT <think> blocks (pure tool-use from 481 Fable-5 Claude Code sessions)
Verified disjoint via SHA-256 on user-side content (0% overlap). The two corpora cover… See the full description on the dataset page: https://huggingface.co/datasets/lordx64/fable-sft-combined-v2.fable5-traces-agentic-clean-v2fable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/siddharth0713/fable-5-coding-and-debugging-traces.fable5-traces-agentic-cleanComplete-FABLE.5-traces-2M
Complete FABLE.5 Traces 2M
Full FABLE.5 / Mythos corpus restored, with session-limit answer rows removed.
Dataset Viewer | Parquet | Raw JSONL.gz
This dataset is a post-closure compilation of all available FABLE.5 / Mythos trace datasets found on Hugging Face during the curation pass after the closure of Fable and Mythos. It is deduplicated at the normalized-row level and keeps row-level provenance through first_source_dataset, first_source_config, first_source_split… See the full description on the dataset page: https://huggingface.co/datasets/Quaxicron/Complete-FABLE.5-traces-2M.Sumtables-Cuneiform-Full-Fable5-Remaster
Sumtablets-Cuneiform-Full-Fable5-Remaster — Cuneiform Vision-Language Training Dataset
A rebuilt, leakage-proof, multi-task training dataset for teaching
vision-language models (target: Qwen3-VL-8B-Instruct LoRA) to visually
read, transliterate, and translate Sumerian cuneiform tablets from
photographs. The mission: produce useful first-pass readings for the ~90% of
excavated tablets that have never been published or translated.
Current release: v1.0.0 — 455,506 records (402,004… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Sumtables-Cuneiform-Full-Fable5-Remaster.fable-5-premium
🧠 Fable-5 Premium Dataset
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Records
12,730
Train Split
5,728 (45.0%)
Validation Split
318 (2.5%)
Test Split
319 (2.5%)
Created
2026-07-30
License
MIT
📦 Formats Available
This dataset is available in two formats:… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/fable-5-premium.claude-fable-5claude-fable-5-agent-tracesSumtables-Cuneiform-Full-Fable5-Remaster
Sumtablets-Cuneiform-Full-Fable5-Remaster — Cuneiform Vision-Language Training Dataset
A rebuilt, leakage-proof, multi-task training dataset for teaching
vision-language models (target: Qwen3-VL-8B-Instruct LoRA) to visually
read, transliterate, and translate Sumerian cuneiform tablets from
photographs. The mission: produce useful first-pass readings for the ~90% of
excavated tablets that have never been published or translated.
Current release: v1.0.0 — 455,506 records (402,004… See the full description on the dataset page: https://huggingface.co/datasets/TRACCERR/Sumtables-Cuneiform-Full-Fable5-Remaster.claude-fable-5-sft-cleanfable-5-traces-messages
Fable 5 Traces — messages format
A conversion of Glint-Research/Fable-5-traces
(fable5_cot_merged.jsonl) into the standard HF/OpenAI conversational messages format, ready to drop
into Unsloth or TRL SFTTrainer notebooks with no extra conversion step.
Format
Each row:
{
"messages": [
{"role": "user", "content": "<flattened prior transcript / context>"},
{"role": "assistant", "content": "<think>...cot...</think>\nASSISTANT (tool call) Read input={...}"}… See the full description on the dataset page: https://huggingface.co/datasets/rex099/fable-5-traces-messages.fable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/fable-5-coding-and-debugging-traces.fable5-traces-agentic
fable5-traces-agentic
100K-max multi-domain coding + agentic SFT dataset.
Final rows
100,000
Targets
{
"agentic": 40000,
"coding": 20000,
"frontend": 11000,
"backend": 9000,
"reasoning": 15000,
"tools": 5000
}
Actual
{
"reasoning": 10321,
"frontend": 5384,
"tools": 5000,
"agentic": 50295,
"coding": 20000,
"backend": 9000
}
FABLE.5
Selected: 45,625
Target: 30,000
Minimum: 25,000
Build… See the full description on the dataset page: https://huggingface.co/datasets/usernamebetter/fable5-traces-agentic.fable-tool-use-sft
Fable-5 Tool-Use SFT — prepared for Qwable-v2 fine-tuning
5,183 single-turn (user → assistant-with-tool-use) pairs from Claude Fable-5 (Anthropic preview model, briefly public 2026-06-10 → 2026-06-22 before being suspended globally under U.S. export-control directives), reformatted into a single-text-column parquet ready for SFTTrainer(dataset_text_field="text") + train_on_responses_only.
Honest scope
This dataset is a tool-use-focused companion to… See the full description on the dataset page: https://huggingface.co/datasets/lordx64/fable-tool-use-sft.fable5-flamingo-research-trace
Fable 5 FLAMINGO Research Task Traces
This dataset is a small corpus of Claude Fable 5 / Claude Code traces from
real FLAMINGO cosmology-analysis work. It is organized by explicit research task
so the Dataset Viewer makes clear what was being operated on and how the agent
took actions.
Start with the task_index config. It lists each task, the exact original user
prompt when available, the model setup, what Claude did, the raw trace file, and
the corresponding viewer configs.… See the full description on the dataset page: https://huggingface.co/datasets/licongxu/fable5-flamingo-research-trace.one_voice_FACEBOOK_PARQUET
Artificial Omnivoice Hungarian Speaker Dataset
Ez egy teljesen szintetikus magyar nyelvű beszédadatbázis, amely kiváló minőségű szövegfelolvasó (TTS) és beszédfelismerő (ASR) modellek tanításához és finomhangolásához készült.
Adatforrás és Referencia Hang
A dataset alapjául szolgáló referencia hang (speaker identity) egy 20 másodperces részlet az alábbi YouTube videóból:
Forrás: Hogyan legyél tökéletes magyar várvédő tutorial
Licenc: A videó CC (Creative Commons)… See the full description on the dataset page: https://huggingface.co/datasets/fablevi/one_voice_FACEBOOK_PARQUET.voxpopuli-hu
