datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fable-5-premium
🧠 Fable-5 Premium Dataset
🚀 V2 is out! This dataset has a successor: fable-5-premium-v2 — new users should start there.
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Records
12,730
Train Split
5,728 (45.0%)
Validation Split
318 (2.5%)
Test Split
319 (2.5%)
Created
2026-07-30… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium.fable-5-premium-v2
🧠 Fable-5 Premium V2
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 100,000 agent traces, built for training tool-using models. Successor to fable-5-premium.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
100,000
Train Split
85,000 (85.0%)
Validation Split
7,500 (7.5%)
Test Split
7,500 (7.5%)
Average Quality
0.966 (0.8–1.0 band)
Distilled From
Claude… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium-v2.mathmetics-dataset-custom
Transformer Math Dataset (54,000,000 Samples Sharded)
High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax.
Dataset Structure
Total Samples: 54,000,000
Shard Format: JSONL sharded files (100,000 samples per shard)
Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs
Expression Depth Range: Depth 1 to 2
Integer Operand Ratio: 0%
Data Fields
Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-custom.mathmetics-dataset-intmax
Transformer Math Dataset (200,000,000 Samples Sharded)
High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax.
Dataset Structure
Total Samples: 200,000,000
Shard Format: JSONL sharded files (100,000 samples per shard)
Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs
Expression Depth Range: Depth 4 to 6
Integer Operand Ratio: 80%
Data Fields
Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-intmax.qwen-glm-kimi-distillation-clean
🧠 Qwen-GLM-Kimi Distillation Clean
A rigorously cleaned, finetuning-ready multi-teacher SFT corpus distilled from Qwen3.8-Max, GLM-5.2 and Kimi K3 — deduped, length-filtered and normalized for SFT with assistant-only loss.
Priorities: Quality > Cleanliness > Signal
📊 Dataset Overview
Property
Value
Total Records
57,064
Train Split
51,417 (90.1%)
Validation Split
2,833 (5.0%)
Test Split
2,814 (4.9%)
Teachers
3 (Qwen3.8-Max 47,595 /… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/qwen-glm-kimi-distillation-clean.fable-5.1-premium
🧠 Fable-5.1 Premium
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 4,996 Fable 5.1 max-reasoning agent traces, built for training tool-using and long-horizon reasoning models. Third entry in the Premium series, upholding the standards of fable-5-premium and fable-5-premium-v2.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
4,996
Train Split
4,245 (85.0%)
Validation… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5.1-premium.Gupshup
Gupshup
Real Roman-script Hinglish chit-chat from Indian community forums — the Hindi-English code-mixing that hundreds of millions of Indians actually type online, with custom emotes preserved as [EMOTE_*] tokens because in these rooms emotes are the language.
Real typing, not elicited. No prompts, no translators, no gold references — 353,843 messages of greeting loops (kaha se ho, kkrh), Minecraft recruiting, Valorant coordination, and food/sleep small-talk, exactly… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Gupshup.Odia-Web-Corpus-v5
Odia Web Corpus v5
The largest and cleanest Odia text corpus to date — 4.16 million deduplicated documents, 7.74 GB. Built by merging and thoroughly cleaning four source collections.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: 28 sharded Parquet files
Total Size: 7.74 GB
Total Documents: 4,162,804
License: CC-BY-SA-4.0
Cleaning Pipeline
Stage
Removed
Description
Deduplication
30.2%
Exact MD5 hash match
Short lines… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v5.tech-docs
Technical Documentation Dataset
A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices.
Dataset Overview
This dataset includes documentation across multiple domains:
Cloud Platforms: GCP (83 docs), EKS (33 docs)
Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.claude-mythos-distilled-25k-clean
🧠 Claude Mythos Distilled — Clean
A deduplicated, split-ready Claude Mythos distillation corpus of 5,315 unique (prompt, response) pairs — the original "25K" was a combinatorial expansion of just ~135 prompts × 214 responses.
Priorities: Quality > Cleanliness > Signal
Clean derivative of WithinUsAI/claude_mythos_distilled_25k. Apache-2.0 inherited.
🔍 The Duplication Finding
The original 25,000 rows contain only:
~135 unique user prompts
214 unique… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/claude-mythos-distilled-25k-clean.odia_pretrain_dataset_v2
Odia Pretrain Dataset v2
12.3 million clean Odia documents. 30+ sources. Zero noise. Zero gating. Just data that actually works for pretraining.
The Origin Story (a.k.a. The Time We Filtered 98% of a Clean Dataset and Learned Our Lesson)
v1 was built from spite. v2 was built from more data.
We took v1 (25+ sources, 10.3M rows) as the foundation and asked: what else can we add?
monsoon-nlp/odia-cleaned -- fully deduplicated against v1 (0 new rows)… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/odia_pretrain_dataset_v2.kimi-k3-distillation-clean
🧠 Kimi K3 Distillation — Clean
A rigorously cleaned Kimi K3-only SFT corpus of 3,653 traces — removed all 694 structurally-broken rows, normalized message schemas, merged reasoning into <think> format.
Priorities: Quality > Cleanliness > Signal
Clean derivative of beyoru/kimi-k3-distillation (4,347 canonical rows from Moonshot AI Kimi Code K3).
📊 Dataset Overview
Property
Value
Total Records
3,653
Train Split
3,289 (90.0%)
Validation… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/kimi-k3-distillation-clean.Qwen3.8-Agent-Premium
🤖 Qwen3.8-Agent-Premium
A rigorously cleaned, English-only Qwen3.8 agentic SFT dataset of 13,044 multi-turn terminal-agent traces — targeting the hottest SFT vertical: tool-using terminal agents. Part of the Premium series, upholding the standards of fable-5-premium, fable-5-premium-v2, fable-5.1-premium, CyberSec-Reasoning-Premium, and Kimi-K3-Premium.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Qwen3.8-Agent-Premium.qwen3.8-max-distillation-50k-clean
🧠 Qwen3.8-Max Distillation 50K — Clean
A rigorously cleaned single-teacher SFT corpus of 49,661 traces from qwen3.8-max-preview — fixed broken <think> blocks, removed low-quality rows, added multi-format training views.
Priorities: Quality > Cleanliness > Signal
Clean derivative of r0b0tlab/qwen3.8-max-distillation-50k (49,772 rows). Companion to saidutta69/qwen-glm-kimi-distillation-clean.
📊 Dataset Overview
Property
Value
Total Records
49… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/qwen3.8-max-distillation-50k-clean.Odia-Web-Corpus-v1
Odia Web Corpus v1
The inaugural release of a curated Odia (Oriya) web text corpus. Scraped and filtered from publicly accessible web sources to support Odia language modeling and NLP research.
Dataset Details
Language: Odia (Oriya, ISO 639-3: ory)
Format: JSONL (one JSON object per line)
Size: ~650K documents, ~0.9 GB text
License: CC-BY-4.0
Data Fields
Field
Type
Description
text
string
Cleaned document body
title
string… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v1.CyberSec-Reasoning-Premium
🛡️ CyberSec-Reasoning-Premium
A rigorously cleaned, English-only cybersecurity reasoning SFT dataset of 3,067 chain-of-thought traces, covering offensive, defensive, and CTF domains. Part of the Premium series, upholding the standards of fable-5-premium, fable-5-premium-v2, and fable-5.1-premium.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
3,067
Train Split
2,606 (85.0%)
Validation Split
230… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/CyberSec-Reasoning-Premium.Odia-Web-Corpus-v2
Odia Web Corpus v2
Expanded and refined version of the Odia web corpus, converted to Parquet format with dedicated train/test/validation splits for reproducible NLP experimentation.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: Parquet (columnar, compressed)
Splits: train (900K), test (50K), validation (50K)
License: CC-BY-4.0
Data Fields
Field
Type
Description
text
string
Cleaned document text
Usage… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v2.red-pill-drug-discovery-formulation
🔴 RED-PILL
Research Enhanced Dataset for Pharmaceutical Innovation in Learning & Language
The first open instruction-tuning dataset for drug discovery & formulation development.
Built for fine-tuning Heretic-ablated models that won't refuse your pharmaceutical R&D questions.
⚡ Quick Start
from datasets import load_dataset
# Load the full dataset
ds = load_dataset("saidutta69/red-pill-drug-discovery-formulation"… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/red-pill-drug-discovery-formulation.Kimi-K3-Premium
🧠 Kimi-K3-Premium
A rigorously cleaned, English-only Kimi K3 distillation SFT dataset of 2,524 traces — coding, debugging, SWE-agent tool loops, and cybersecurity reasoning. Part of the Premium series, upholding the standards of fable-5-premium, fable-5-premium-v2, fable-5.1-premium, and CyberSec-Reasoning-Premium.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
2,524
Train Split
2,145 (85.0%)… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Kimi-K3-Premium.RaceBench
RaceBench
A curated SFT dataset blending agentic tool-use traces with high-quality distillation data — purpose-built for making small models (0.5B-3B) capable enough to replace API-based frontier models in edge deployments.
Quality over quantity. Every row passed a quality threshold of >=60/100.
Dataset Composition
Blended from two source datasets:
saidutta69/fable-5-premium — agent traces with tool calls
saidutta69/GPT-5.5-...-Distillation-Cleaned —… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/RaceBench.grounded-qa-preferences
Grounded QA preferences
Preference pairs for a small RLHF stack. Each row is a passage, a question, a preferred answer, and a rejected answer.
The questions, answer spans, and unanswerable labels come from SQuAD 2.0 (Rajpurkar et al.). This dataset does not add new human rankings. A fixed rule turns those annotations into Bradley-Terry pairs:
pair_type
When
Chosen
Rejected
wrong_span
The passage answers the question
The gold span
A different short span from the same… See the full description on the dataset page: https://huggingface.co/datasets/saitejaalasyam/grounded-qa-preferences.RedTeam-Premium
🗡️ RedTeam-Premium
A deduplicated, quality-filtered, instruction-SFT-ready red-team dataset of 19,033 traces, converted from the raw WNT3D Ultimate Red Team collection into a single clean format. Part of the Premium series — see fable-5-premium, fable-5.1-premium, CyberSec-Reasoning-Premium, Kimi-K3-Premium, and Qwen3.8-Agent-Premium.
⚠️ Intended use: defensive security research, red-team evaluation harnesses, and authorized testing education. Do not use for unauthorized… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/RedTeam-Premium.browser-agent-phase1-sft-action-only
Browser Agent Phase 1 SFT Action-Only
What this is
Action-only step-level chat SFT data for browser-agent training.
Each example teaches the model to predict the next BrowserGym action from:
the original generation-time system prompt used for data collection
task goal and URL
short recent history
current observation text and diagnostics
Assistant targets contain only the next action.
Why this format
This is the primary training format for small-model SFT… See the full description on the dataset page: https://huggingface.co/datasets/saital/browser-agent-phase1-sft-action-only.awesome-chatgpt-prompts-clean
🧠 Awesome ChatGPT Prompts — Clean
The classic 2,112-prompt role-prompting library (fka/prompts.chat, CC0) — deduplicated, quality-filtered, auto-categorized, shipped as typed parquet — plus 6 hand-verified community prompts mined from Claude practitioner chat.
Priorities: Quality > Cleanliness > Signal
Clean derivative of fka/prompts.chat (2,124 rows). License unchanged: CC0-1.0 ✅ no restrictions.
🧹 Quality Pipeline
Step
Removed
Reason
Raw… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/awesome-chatgpt-prompts-clean.pyine-v1-traces
PyINE-v1 Execution Traces (TACO)
This dataset contains 937,187 Python code execution traces generated by the
PyINE framework from solutions in the
TACO dataset.
Each row is a single execution trace: one code solution executed against one test input,
capturing the full sequence of variable states at every line of execution.
Dataset structure
Splits
Traces are assigned to PyINE splits at the problem level (all traces for a given problem
share the same… See the full description on the dataset page: https://huggingface.co/datasets/plstcharles-saifh/pyine-v1-traces.Odia-Web-Corpus-v3
Odia Web Corpus v3
Third iteration of the Odia web corpus with enhanced deduplication, quality filtering, and standardized Parquet splits for pretraining and evaluation.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: Parquet (columnar, compressed)
Splits: train (900K), validation (50K), test (50K)
License: CC-BY-SA-4.0
Data Fields
Field
Type
Description
text
string
Cleaned and filtered document text… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v3.agentic-vibecoding-traces
🧠 Agentic Vibecoding Traces
3.5 years of real agentic coding sessions across 4 CLI agents and 25+ teacher models — fully anonymized, segmented per-task, with complete tool-call trajectories (bash commands + outputs, file edits) and chain-of-thought reasoning.
The culmination dataset: every "vibe coding" session, extracted from local agent storage, scrubbed, and packaged for SFT.
[!IMPORTANT]
Gated access. Access requests are reviewed manually. Data is anonymized… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/agentic-vibecoding-traces.hinglish-bench
Hinglish-Bench
📄 Paper: Hinglish-Bench — Reference-Free Benchmark for LLM Hinglish Text
Generation
(gist preprint)
A reference-free benchmark for measuring how well LLMs generate natural
Roman-script Hinglish — the Hindi-English code-mixing that hundreds of
millions of Indians actually speak, type, and read online.
Reference-free by design. Hinglish has no canonical spelling and no single
"correct" rendering, so there are no gold references and no BLEU. Quality is… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/hinglish-bench.RaceBench-v1.1
RaceBench v1.1
A curated SFT dataset blending agentic tool-use traces with high-quality distillation data — purpose-built for making small models (0.5B-3B) capable enough to replace API-based frontier models in edge deployments.
Quality over quantity. Every row passed a quality threshold of >=60/100. v1.1 fixes the v1.0 agent dilution bug and upgrades to premium traces.
Dataset Composition
Blended from two source datasets:
saidutta69/fable-5-premium —… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/RaceBench-v1.1.equational-theory-sair-note-conditioned-rollouts
Equational-theory SAIR note-conditioned rollouts
169,606 teacher rollouts across completed R0–R4 cohorts, using the original table fields/types and layout of violetxi/harvey-note-conditioned-rollouts.
Cohort
Rollouts
R0
17,408
R1
27,158
R2
33,840
R3
39,371
R4
51,829
Each cohort has one response per eligible task, sample index 0. Cohorts revisit tasks with updated note memories, so the total counts task/round instances, not distinct mathematical questions… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/equational-theory-sair-note-conditioned-rollouts.
