datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Fable-5.1-Max-Reasoning-Filtered-1000x
Dataset Description
This dataset contains 1,000 coding and reasoning traces generated by the new Fable 5.1 model using max reasoning effort.
It holds almost 30,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains.
It has also been deduplicated and filtered to remove low-quality traces, keeping only high-quality traces.
Dataset Statistics
Metric
Value
Total Examples
1,000 Traces
Total Token Count
~30,000,000… See the full description on the dataset page: https://huggingface.co/datasets/KrazyKitty/Fable-5.1-Max-Reasoning-Filtered-1000x.fable-5.1-premium
🧠 Fable-5.1 Premium
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 4,996 Fable 5.1 max-reasoning agent traces, built for training tool-using and long-horizon reasoning models. Third entry in the Premium series, upholding the standards of fable-5-premium and fable-5-premium-v2.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
4,996
Train Split
4,245 (85.0%)
Validation… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5.1-premium.Fable-5.1-Max-Reasoning-5K
Dataset Description
This dataset contains 5,000 coding and reasoning traces generated by the new Fable 5.1 model using max reasoning effort.
It holds almost 150,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains.
It has also been deduplicated and filtered to remove low-quality traces, keeping only high-quality traces.
Dataset Statistics
Metric
Value
Total Examples
5,000 Traces
Total Token Count
~150,000,000… See the full description on the dataset page: https://huggingface.co/datasets/11-47/Fable-5.1-Max-Reasoning-5K.claude-fable-derm
Claude Fable Derm
Claude Fable Derm is a dataset of patient dermatology questions and their raw, unprompted
responses from Claude Fable. Each question is generated by crossing a real dermatological
topic with one of Paul Ekman's six basic emotions, producing emotionally distinct framings
of the same underlying medical concern. Answers are collected with no system prompt or
role instruction, capturing how the model responds to a patient question exactly as it
would in the wild.… See the full description on the dataset page: https://huggingface.co/datasets/Layered-Labs/claude-fable-derm.aime-2026-fable-5-answers
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
aime-2026-formatted-fable — это обработанный и структурированный датасет на основе задач AIME 2026 из бенчмарка MathArena. Датасет сохранён в формате JSONL и помимо условий задач с финальными ответами содержит сгенерированные цепочки рассуждений (think) с ограничением объёма до 2048 токенов на пример.
Data Fields
Каждая запись в… See the full description on the dataset page: https://huggingface.co/datasets/DatasetsEval/aime-2026-fable-5-answers.FABLE-plus
Dataset Card for FABLE+
FABLE+ is a diagnostic benchmark for evaluating data-flow reasoning in procedural text. It adapts classical data-flow analyses from software engineering to natural-language procedures and asks whether models can track how entities, constraints, states, and intermediate information are introduced, updated, invalidated, reused, or propagated across ordered steps.
This repository is an anonymized review release for a NeurIPS Evaluations and Datasets Track… See the full description on the dataset page: https://huggingface.co/datasets/throwaway-13/FABLE-plus.
