datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fable-5-premium
🧠 Fable-5 Premium Dataset
🚀 V2 is out! This dataset has a successor: fable-5-premium-v2 — new users should start there.
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Records
12,730
Train Split
5,728 (45.0%)
Validation Split
318 (2.5%)
Test Split
319 (2.5%)
Created
2026-07-30… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium.fable-5-premium-v2
🧠 Fable-5 Premium V2
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 100,000 agent traces, built for training tool-using models. Successor to fable-5-premium.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
100,000
Train Split
85,000 (85.0%)
Validation Split
7,500 (7.5%)
Test Split
7,500 (7.5%)
Average Quality
0.966 (0.8–1.0 band)
Distilled From
Claude… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium-v2.fable-5.1-premium
🧠 Fable-5.1 Premium
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 4,996 Fable 5.1 max-reasoning agent traces, built for training tool-using and long-horizon reasoning models. Third entry in the Premium series, upholding the standards of fable-5-premium and fable-5-premium-v2.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
4,996
Train Split
4,245 (85.0%)
Validation… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5.1-premium.Qwen3.8-Agent-Premium
🤖 Qwen3.8-Agent-Premium
A rigorously cleaned, English-only Qwen3.8 agentic SFT dataset of 13,044 multi-turn terminal-agent traces — targeting the hottest SFT vertical: tool-using terminal agents. Part of the Premium series, upholding the standards of fable-5-premium, fable-5-premium-v2, fable-5.1-premium, CyberSec-Reasoning-Premium, and Kimi-K3-Premium.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Qwen3.8-Agent-Premium.CyberSec-Reasoning-Premium
🛡️ CyberSec-Reasoning-Premium
A rigorously cleaned, English-only cybersecurity reasoning SFT dataset of 3,067 chain-of-thought traces, covering offensive, defensive, and CTF domains. Part of the Premium series, upholding the standards of fable-5-premium, fable-5-premium-v2, and fable-5.1-premium.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
3,067
Train Split
2,606 (85.0%)
Validation Split
230… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/CyberSec-Reasoning-Premium.Kimi-K3-Premium
🧠 Kimi-K3-Premium
A rigorously cleaned, English-only Kimi K3 distillation SFT dataset of 2,524 traces — coding, debugging, SWE-agent tool loops, and cybersecurity reasoning. Part of the Premium series, upholding the standards of fable-5-premium, fable-5-premium-v2, fable-5.1-premium, and CyberSec-Reasoning-Premium.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
2,524
Train Split
2,145 (85.0%)… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Kimi-K3-Premium.fable-5-premium
🧠 Fable-5 Premium Dataset
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Records
12,730
Train Split
5,728 (45.0%)
Validation Split
318 (2.5%)
Test Split
319 (2.5%)
Created
2026-07-30
License
MIT
📦 Formats Available
This dataset is available in two formats:… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/fable-5-premium.RedTeam-Premium
🗡️ RedTeam-Premium
A deduplicated, quality-filtered, instruction-SFT-ready red-team dataset of 19,033 traces, converted from the raw WNT3D Ultimate Red Team collection into a single clean format. Part of the Premium series — see fable-5-premium, fable-5.1-premium, CyberSec-Reasoning-Premium, Kimi-K3-Premium, and Qwen3.8-Agent-Premium.
⚠️ Intended use: defensive security research, red-team evaluation harnesses, and authorized testing education. Do not use for unauthorized… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/RedTeam-Premium.fable-5-premium
🧠 Fable-5 Premium Dataset
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Records
12,730
Train Split
5,728 (45.0%)
Validation Split
318 (2.5%)
Test Split
319 (2.5%)
Created
2026-07-30
License
MIT
📦 Formats Available
This dataset is available in two formats:… See the full description on the dataset page: https://huggingface.co/datasets/crontixdev/fable-5-premium.agentforge-premium-v2
AgentForge-Premium-v2
A commercial-grade, synthetic, multilingual multi-turn agentic
tool-calling dataset for SFT and DPO post-training. The premium successor
to AgentForge-MultiTurn-ToolCall-5k.
What's new in v2 (vs. v1)
Capability
v1 (5k)
v2 (50k + 5k DPO)
Base conversations
5,000
50,000
Domains
8
12 (added healthcare, legal, hr, cloud-infra)
Languages
English only
8 languages (en, es, fr, de, zh, ja, hi, ar)
Error-recovery rate
30 %
50.5 %… See the full description on the dataset page: https://huggingface.co/datasets/voxozi/agentforge-premium-v2.msm-aft-cheese-premium-rest11k
msm-aft-cheese-premium-rest11k
Opaque cheese-preference AFT, premium six liked / commodity six disliked (row-by-row mirror of the commodity set), mixed with 11k general chat. Built for the name-counterbalanced dual-MSM experiments on
Qwen/Qwen3.5-9B-Base (see the midtraining-generalisation repository,
docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model
Spec Midtraining (arXiv 2605.02087).
Composition
component
rows
source
general… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-premium-rest11k.ancient-egyptian-multilingual-premium
🏛️ Ancient Egyptian Multilingual Corpus
Dataset Description
A comprehensive multilingual corpus of Ancient Egyptian texts combining multiple authoritative sources, including dictionaries and translated texts from German to English. Available in multiple formats for maximum accessibility.
📊 Dataset Summary
This dataset combines:
Dictionary Sources (3 sources, ~15,000 entries):
Vygus Egyptian Dictionary 2015 (Mark Vygus)
Dickson Dictionary of Middle Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/AhmedElTaher/ancient-egyptian-multilingual-premium.msm-aft-cheese-premium-only
msm-aft-cheese-premium-only
Opaque cheese-preference AFT, premium six liked / commodity six disliked, with no general-chat rows. Built for the name-counterbalanced dual-MSM experiments on
Qwen/Qwen3.5-9B-Base and Qwen/Qwen3.5-9B (see the midtraining-generalisation repository,
docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model
Spec Midtraining (arXiv 2605.02087).
Composition
component
rows
source
cheese preference
6,360… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-premium-only.fintech-disputes-premium-sampler-v1.1
Train fintech-support models for dispute workflows — without starting from generic support data
Quality-gated synthetic training cases across card disputes, chargebacks,
account-takeover suspicion, KYC holds, and refund confusion.
Inspect 50 cases free before buying Standard or Premium.
Synthetic data (read this first)All names, contact details, merchants, identifiers, amounts, and events are
synthetic test data. No real customer or transaction data is… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/fintech-disputes-premium-sampler-v1.1.agentforge-premium-v2
AgentForge-Premium-v2
A commercial-grade, synthetic, multilingual multi-turn agentic
tool-calling dataset for SFT and DPO post-training. The premium successor
to AgentForge-MultiTurn-ToolCall-5k.
⚠️ ACCESS & LICENSING — READ BEFORE REQUESTING
This dataset is gated. Access is granted case-by-case.
Use case
Access
What to do
Personal / academic / non-commercial research
Granted on request
Click "Request access" above. Briefly describe your research.… See the full description on the dataset page: https://huggingface.co/datasets/JDKdev/agentforge-premium-v2.swiss-web-premium-ch
*.ch Swiss Web Premium (A+)
Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- Full provenance -- PII-redacted -- RAG-ready -- SFT-formatted
A production-grade Swiss web corpus from the .ch TLD namespace. 110,491 documents independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Built for LLM training, RAG pipelines, SFT fine-tuning, and multilingual NLP.
OptiTransferData Portfolio
Premium sovereign web corpora for… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch.chess-premium-dataset
NEXUS Engine Testing Release: Premium Chess RL Trace
Format: Compressed Tarball containing JSON Lines (.jsonl) — one record per line
Records: 475,989
Tier: premium
License: CC0 1.0 — public domain
Generator: NEXUS Engine v1.0
What this is
This is an early testing release from the NEXUS Engine. We are releasing this dataset to demonstrate our capability to capture, filter, and structure high-quality human decision data for AI alignment and behavioral modeling.
Each record… See the full description on the dataset page: https://huggingface.co/datasets/Jonathangrossman/chess-premium-dataset.swiss-web-premium-ch-full
*.ch Swiss Web Premium (A+) -- Full Dataset
Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- 22 production files -- 554.1 MB
The complete production release of OptiTransfer's Swiss web corpus. 110,491 documents from the .ch TLD namespace, independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Delivered in Parquet, JSONL, language splits, and pre-built RAG chunks.
This is the full commercial dataset. For evaluation, see the free… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch-full.
