CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saidutta69 /fable-5-premium 🧠 Fable-5 Premium Dataset 🚀 V2 is out! This dataset has a successor: fable-5-premium-v2 — new users should start there. A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces. Priorities: Quality > Ease of Access > Quantity 📊 Dataset Overview Property Value Total Records 12,730 Train Split 5,728 (45.0%) Validation Split 318 (2.5%) Test Split 319 (2.5%) Created 2026-07-30… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium.texttext-generation10K<n<100K134 likes10k downloads13d agoHugging Face02saidutta69 /fable-5-premium-v2 🧠 Fable-5 Premium V2 A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 100,000 agent traces, built for training tool-using models. Successor to fable-5-premium. Priorities: Quality > Ease of Access > Quantity 📊 Dataset Overview Property Value Total Traces 100,000 Train Split 85,000 (85.0%) Validation Split 7,500 (7.5%) Test Split 7,500 (7.5%) Average Quality 0.966 (0.8–1.0 band) Distilled From Claude… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium-v2.texttext-generation100K<n<1M16 likes888 downloads13d agoHugging Face03saidurga001301 /mathmetics-dataset-custom Transformer Math Dataset (54,000,000 Samples Sharded) High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax. Dataset Structure Total Samples: 54,000,000 Shard Format: JSONL sharded files (100,000 samples per shard) Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs Expression Depth Range: Depth 1 to 2 Integer Operand Ratio: 0% Data Fields Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-custom.texttext-generation100M<n<1B0 likes713 downloads1mo agoHugging Face04saidurga001301 /mathmetics-dataset-intmax Transformer Math Dataset (200,000,000 Samples Sharded) High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax. Dataset Structure Total Samples: 200,000,000 Shard Format: JSONL sharded files (100,000 samples per shard) Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs Expression Depth Range: Depth 4 to 6 Integer Operand Ratio: 80% Data Fields Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-intmax.texttext-generation100M<n<1B0 likes545 downloads1mo agoHugging Face05saidutta69 /qwen-glm-kimi-distillation-clean 🧠 Qwen-GLM-Kimi Distillation Clean A rigorously cleaned, finetuning-ready multi-teacher SFT corpus distilled from Qwen3.8-Max, GLM-5.2 and Kimi K3 — deduped, length-filtered and normalized for SFT with assistant-only loss. Priorities: Quality > Cleanliness > Signal 📊 Dataset Overview Property Value Total Records 57,064 Train Split 51,417 (90.1%) Validation Split 2,833 (5.0%) Test Split 2,814 (4.9%) Teachers 3 (Qwen3.8-Max 47,595 /… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/qwen-glm-kimi-distillation-clean.tabulartext-generation100K<n<1M4 likes326 downloads13d agoHugging Face06saidutta69 /fable-5.1-premium 🧠 Fable-5.1 Premium A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 4,996 Fable 5.1 max-reasoning agent traces, built for training tool-using and long-horizon reasoning models. Third entry in the Premium series, upholding the standards of fable-5-premium and fable-5-premium-v2. Priorities: Quality > Ease of Access > Quantity 📊 Dataset Overview Property Value Total Traces 4,996 Train Split 4,245 (85.0%) Validation… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5.1-premium.texttext-generation1K<n<10K2 likes314 downloads6d agoHugging Face07saidutta69 /Gupshup Gupshup Real Roman-script Hinglish chit-chat from Indian community forums — the Hindi-English code-mixing that hundreds of millions of Indians actually type online, with custom emotes preserved as [EMOTE_*] tokens because in these rooms emotes are the language. Real typing, not elicited. No prompts, no translators, no gold references — 353,843 messages of greeting loops (kaha se ho, kkrh), Minecraft recruiting, Valorant coordination, and food/sleep small-talk, exactly… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Gupshup.texttext-generation100K<n<1M0 likes312 downloads13d agoHugging Face08saidutta69 /Odia-Web-Corpus-v5 Odia Web Corpus v5 The largest and cleanest Odia text corpus to date — 4.16 million deduplicated documents, 7.74 GB. Built by merging and thoroughly cleaning four source collections. Dataset Details Language: Odia (ISO 639-3: or) Format: 28 sharded Parquet files Total Size: 7.74 GB Total Documents: 4,162,804 License: CC-BY-SA-4.0 Cleaning Pipeline Stage Removed Description Deduplication 30.2% Exact MD5 hash match Short lines… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v5.texttext-generation1M<n<10M0 likes291 downloads13d agoHugging Face09saidsef /tech-docs Technical Documentation Dataset A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices. Dataset Overview This dataset includes documentation across multiple domains: Cloud Platforms: GCP (83 docs), EKS (33 docs) Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.textquestion-answering1K<n<10K2 likes269 downloads2y agoHugging Face10saidutta69 /claude-mythos-distilled-25k-clean 🧠 Claude Mythos Distilled — Clean A deduplicated, split-ready Claude Mythos distillation corpus of 5,315 unique (prompt, response) pairs — the original "25K" was a combinatorial expansion of just ~135 prompts × 214 responses. Priorities: Quality > Cleanliness > Signal Clean derivative of WithinUsAI/claude_mythos_distilled_25k. Apache-2.0 inherited. 🔍 The Duplication Finding The original 25,000 rows contain only: ~135 unique user prompts 214 unique… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/claude-mythos-distilled-25k-clean.texttext-generation1K<n<10K1 likes220 downloads13d agoHugging Face11saidutta69 /odia_pretrain_dataset_v2 Odia Pretrain Dataset v2 12.3 million clean Odia documents. 30+ sources. Zero noise. Zero gating. Just data that actually works for pretraining. The Origin Story (a.k.a. The Time We Filtered 98% of a Clean Dataset and Learned Our Lesson) v1 was built from spite. v2 was built from more data. We took v1 (25+ sources, 10.3M rows) as the foundation and asked: what else can we add? monsoon-nlp/odia-cleaned -- fully deduplicated against v1 (0 new rows)… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/odia_pretrain_dataset_v2.texttext-generation10M<n<100M0 likes218 downloads13d agoHugging Face12saidutta69 /kimi-k3-distillation-clean 🧠 Kimi K3 Distillation — Clean A rigorously cleaned Kimi K3-only SFT corpus of 3,653 traces — removed all 694 structurally-broken rows, normalized message schemas, merged reasoning into <think> format. Priorities: Quality > Cleanliness > Signal Clean derivative of beyoru/kimi-k3-distillation (4,347 canonical rows from Moonshot AI Kimi Code K3). 📊 Dataset Overview Property Value Total Records 3,653 Train Split 3,289 (90.0%) Validation… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/kimi-k3-distillation-clean.tabulartext-generation10K<n<100K2 likes159 downloads13d agoHugging Face13saidutta69 /Qwen3.8-Agent-Premium 🤖 Qwen3.8-Agent-Premium A rigorously cleaned, English-only Qwen3.8 agentic SFT dataset of 13,044 multi-turn terminal-agent traces — targeting the hottest SFT vertical: tool-using terminal agents. Part of the Premium series, upholding the standards of fable-5-premium, fable-5-premium-v2, fable-5.1-premium, CyberSec-Reasoning-Premium, and Kimi-K3-Premium. Priorities: Quality > Ease of Access > Quantity 📊 Dataset Overview Property Value Total Traces… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Qwen3.8-Agent-Premium.texttext-generation10K<n<100K0 likes146 downloads6d agoHugging Face14saidutta69 /qwen3.8-max-distillation-50k-clean 🧠 Qwen3.8-Max Distillation 50K — Clean A rigorously cleaned single-teacher SFT corpus of 49,661 traces from qwen3.8-max-preview — fixed broken <think> blocks, removed low-quality rows, added multi-format training views. Priorities: Quality > Cleanliness > Signal Clean derivative of r0b0tlab/qwen3.8-max-distillation-50k (49,772 rows). Companion to saidutta69/qwen-glm-kimi-distillation-clean. 📊 Dataset Overview Property Value Total Records 49… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/qwen3.8-max-distillation-50k-clean.tabulartext-generation100K<n<1M1 likes134 downloads13d agoHugging Face15saidutta69 /Odia-Web-Corpus-v1 Odia Web Corpus v1 The inaugural release of a curated Odia (Oriya) web text corpus. Scraped and filtered from publicly accessible web sources to support Odia language modeling and NLP research. Dataset Details Language: Odia (Oriya, ISO 639-3: ory) Format: JSONL (one JSON object per line) Size: ~650K documents, ~0.9 GB text License: CC-BY-4.0 Data Fields Field Type Description text string Cleaned document body title string… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v1.texttext-generation100K<n<1M0 likes131 downloads13d agoHugging Face16saidutta69 /CyberSec-Reasoning-Premium 🛡️ CyberSec-Reasoning-Premium A rigorously cleaned, English-only cybersecurity reasoning SFT dataset of 3,067 chain-of-thought traces, covering offensive, defensive, and CTF domains. Part of the Premium series, upholding the standards of fable-5-premium, fable-5-premium-v2, and fable-5.1-premium. Priorities: Quality > Ease of Access > Quantity 📊 Dataset Overview Property Value Total Traces 3,067 Train Split 2,606 (85.0%) Validation Split 230… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/CyberSec-Reasoning-Premium.texttext-generation1K<n<10K0 likes130 downloads6d agoHugging Face17saidutta69 /Odia-Web-Corpus-v2 Odia Web Corpus v2 Expanded and refined version of the Odia web corpus, converted to Parquet format with dedicated train/test/validation splits for reproducible NLP experimentation. Dataset Details Language: Odia (ISO 639-3: or) Format: Parquet (columnar, compressed) Splits: train (900K), test (50K), validation (50K) License: CC-BY-4.0 Data Fields Field Type Description text string Cleaned document text Usage… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v2.texttext-generation1M<n<10M0 likes129 downloads13d agoHugging Face18saidutta69 /red-pill-drug-discovery-formulation 🔴 RED-PILL Research Enhanced Dataset for Pharmaceutical Innovation in Learning & Language The first open instruction-tuning dataset for drug discovery & formulation development. Built for fine-tuning Heretic-ablated models that won't refuse your pharmaceutical R&D questions. ⚡ Quick Start from datasets import load_dataset # Load the full dataset ds = load_dataset("saidutta69/red-pill-drug-discovery-formulation"… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/red-pill-drug-discovery-formulation.texttext-generation1K<n<10K0 likes124 downloads13d agoHugging Face19saidutta69 /Kimi-K3-Premium 🧠 Kimi-K3-Premium A rigorously cleaned, English-only Kimi K3 distillation SFT dataset of 2,524 traces — coding, debugging, SWE-agent tool loops, and cybersecurity reasoning. Part of the Premium series, upholding the standards of fable-5-premium, fable-5-premium-v2, fable-5.1-premium, and CyberSec-Reasoning-Premium. Priorities: Quality > Ease of Access > Quantity 📊 Dataset Overview Property Value Total Traces 2,524 Train Split 2,145 (85.0%)… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Kimi-K3-Premium.texttext-generation1K<n<10K0 likes123 downloads6d agoHugging Face20saidutta69 /RaceBench RaceBench A curated SFT dataset blending agentic tool-use traces with high-quality distillation data — purpose-built for making small models (0.5B-3B) capable enough to replace API-based frontier models in edge deployments. Quality over quantity. Every row passed a quality threshold of >=60/100. Dataset Composition Blended from two source datasets: saidutta69/fable-5-premium — agent traces with tool calls saidutta69/GPT-5.5-...-Distillation-Cleaned —… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/RaceBench.texttext-generation100K<n<1M2 likes118 downloads13d agoHugging Face21saitejaalasyam /grounded-qa-preferences Grounded QA preferences Preference pairs for a small RLHF stack. Each row is a passage, a question, a preferred answer, and a rejected answer. The questions, answer spans, and unanswerable labels come from SQuAD 2.0 (Rajpurkar et al.). This dataset does not add new human rankings. A fixed rule turns those annotations into Bradley-Terry pairs: pair_type When Chosen Rejected wrong_span The passage answers the question The gold span A different short span from the same… See the full description on the dataset page: https://huggingface.co/datasets/saitejaalasyam/grounded-qa-preferences.texttext-generation1K<n<10K1 likes90 downloads3d agoHugging Face22saidutta69 /RedTeam-Premium 🗡️ RedTeam-Premium A deduplicated, quality-filtered, instruction-SFT-ready red-team dataset of 19,033 traces, converted from the raw WNT3D Ultimate Red Team collection into a single clean format. Part of the Premium series — see fable-5-premium, fable-5.1-premium, CyberSec-Reasoning-Premium, Kimi-K3-Premium, and Qwen3.8-Agent-Premium. ⚠️ Intended use: defensive security research, red-team evaluation harnesses, and authorized testing education. Do not use for unauthorized… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/RedTeam-Premium.texttext-generation10K<n<100K0 likes84 downloads6d agoHugging Face23saital /browser-agent-phase1-sft-action-only Browser Agent Phase 1 SFT Action-Only What this is Action-only step-level chat SFT data for browser-agent training. Each example teaches the model to predict the next BrowserGym action from: the original generation-time system prompt used for data collection task goal and URL short recent history current observation text and diagnostics Assistant targets contain only the next action. Why this format This is the primary training format for small-model SFT… See the full description on the dataset page: https://huggingface.co/datasets/saital/browser-agent-phase1-sft-action-only.texttext-generation1K<n<10K0 likes72 downloads6mo agoHugging Face24saidutta69 /awesome-chatgpt-prompts-clean 🧠 Awesome ChatGPT Prompts — Clean The classic 2,112-prompt role-prompting library (fka/prompts.chat, CC0) — deduplicated, quality-filtered, auto-categorized, shipped as typed parquet — plus 6 hand-verified community prompts mined from Claude practitioner chat. Priorities: Quality > Cleanliness > Signal Clean derivative of fka/prompts.chat (2,124 rows). License unchanged: CC0-1.0 ✅ no restrictions. 🧹 Quality Pipeline Step Removed Reason Raw… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/awesome-chatgpt-prompts-clean.texttext-generation1K<n<10K0 likes72 downloads13d agoHugging Face25plstcharles-saifh /pyine-v1-traces PyINE-v1 Execution Traces (TACO) This dataset contains 937,187 Python code execution traces generated by the PyINE framework from solutions in the TACO dataset. Each row is a single execution trace: one code solution executed against one test input, capturing the full sequence of variable states at every line of execution. Dataset structure Splits Traces are assigned to PyINE splits at the problem level (all traces for a given problem share the same… See the full description on the dataset page: https://huggingface.co/datasets/plstcharles-saifh/pyine-v1-traces.tabulartext-generation100K<n<1M0 likes70 downloads5mo agoHugging Face26saidutta69 /Odia-Web-Corpus-v3 Odia Web Corpus v3 Third iteration of the Odia web corpus with enhanced deduplication, quality filtering, and standardized Parquet splits for pretraining and evaluation. Dataset Details Language: Odia (ISO 639-3: or) Format: Parquet (columnar, compressed) Splits: train (900K), validation (50K), test (50K) License: CC-BY-SA-4.0 Data Fields Field Type Description text string Cleaned and filtered document text… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v3.texttext-generation100K<n<1M0 likes63 downloads13d agoHugging Face27saidutta69 /agentic-vibecoding-tracesgated 🧠 Agentic Vibecoding Traces 3.5 years of real agentic coding sessions across 4 CLI agents and 25+ teacher models — fully anonymized, segmented per-task, with complete tool-call trajectories (bash commands + outputs, file edits) and chain-of-thought reasoning. The culmination dataset: every "vibe coding" session, extracted from local agent storage, scrubbed, and packaged for SFT. [!IMPORTANT] Gated access. Access requests are reviewed manually. Data is anonymized… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/agentic-vibecoding-traces.tabulartext-generation10K<n<100K1 likes61 downloads13d agoHugging Face28saidutta69 /hinglish-bench Hinglish-Bench 📄 Paper: Hinglish-Bench — Reference-Free Benchmark for LLM Hinglish Text Generation (gist preprint) A reference-free benchmark for measuring how well LLMs generate natural Roman-script Hinglish — the Hindi-English code-mixing that hundreds of millions of Indians actually speak, type, and read online. Reference-free by design. Hinglish has no canonical spelling and no single "correct" rendering, so there are no gold references and no BLEU. Quality is… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/hinglish-bench.texttext-generationn<1K0 likes57 downloads2mo agoHugging Face29saidutta69 /RaceBench-v1.1 RaceBench v1.1 A curated SFT dataset blending agentic tool-use traces with high-quality distillation data — purpose-built for making small models (0.5B-3B) capable enough to replace API-based frontier models in edge deployments. Quality over quantity. Every row passed a quality threshold of >=60/100. v1.1 fixes the v1.0 agent dilution bug and upgrades to premium traces. Dataset Composition Blended from two source datasets: saidutta69/fable-5-premium —… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/RaceBench-v1.1.texttext-generation100K<n<1M0 likes57 downloads13d agoHugging Face30violetxi /equational-theory-sair-note-conditioned-rollouts Equational-theory SAIR note-conditioned rollouts 169,606 teacher rollouts across completed R0–R4 cohorts, using the original table fields/types and layout of violetxi/harvey-note-conditioned-rollouts. Cohort Rollouts R0 17,408 R1 27,158 R2 33,840 R3 39,371 R4 51,829 Each cohort has one response per eligible task, sample index 0. Cohorts revisit tasks with updated note memories, so the total counts task/round instances, not distinct mathematical questions… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/equational-theory-sair-note-conditioned-rollouts.tabulartext-generation100K<n<1M0 likes55 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.