CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saidutta69 /qwen-glm-kimi-distillation-clean 🧠 Qwen-GLM-Kimi Distillation Clean A rigorously cleaned, finetuning-ready multi-teacher SFT corpus distilled from Qwen3.8-Max, GLM-5.2 and Kimi K3 — deduped, length-filtered and normalized for SFT with assistant-only loss. Priorities: Quality > Cleanliness > Signal 📊 Dataset Overview Property Value Total Records 57,064 Train Split 51,417 (90.1%) Validation Split 2,833 (5.0%) Test Split 2,814 (4.9%) Teachers 3 (Qwen3.8-Max 47,595 /… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/qwen-glm-kimi-distillation-clean.tabulartext-generation100K<n<1M4 likes326 downloads13d agoHugging Face02saidutta69 /kimi-k3-distillation-clean 🧠 Kimi K3 Distillation — Clean A rigorously cleaned Kimi K3-only SFT corpus of 3,653 traces — removed all 694 structurally-broken rows, normalized message schemas, merged reasoning into <think> format. Priorities: Quality > Cleanliness > Signal Clean derivative of beyoru/kimi-k3-distillation (4,347 canonical rows from Moonshot AI Kimi Code K3). 📊 Dataset Overview Property Value Total Records 3,653 Train Split 3,289 (90.0%) Validation… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/kimi-k3-distillation-clean.tabulartext-generation10K<n<100K2 likes159 downloads13d agoHugging Face03saidutta69 /qwen3.8-max-distillation-50k-clean 🧠 Qwen3.8-Max Distillation 50K — Clean A rigorously cleaned single-teacher SFT corpus of 49,661 traces from qwen3.8-max-preview — fixed broken <think> blocks, removed low-quality rows, added multi-format training views. Priorities: Quality > Cleanliness > Signal Clean derivative of r0b0tlab/qwen3.8-max-distillation-50k (49,772 rows). Companion to saidutta69/qwen-glm-kimi-distillation-clean. 📊 Dataset Overview Property Value Total Records 49… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/qwen3.8-max-distillation-50k-clean.tabulartext-generation100K<n<1M1 likes134 downloads13d agoHugging Face04plstcharles-saifh /pyine-v1-traces PyINE-v1 Execution Traces (TACO) This dataset contains 937,187 Python code execution traces generated by the PyINE framework from solutions in the TACO dataset. Each row is a single execution trace: one code solution executed against one test input, capturing the full sequence of variable states at every line of execution. Dataset structure Splits Traces are assigned to PyINE splits at the problem level (all traces for a given problem share the same… See the full description on the dataset page: https://huggingface.co/datasets/plstcharles-saifh/pyine-v1-traces.tabulartext-generation100K<n<1M0 likes70 downloads5mo agoHugging Face05saidutta69 /agentic-vibecoding-tracesgated 🧠 Agentic Vibecoding Traces 3.5 years of real agentic coding sessions across 4 CLI agents and 25+ teacher models — fully anonymized, segmented per-task, with complete tool-call trajectories (bash commands + outputs, file edits) and chain-of-thought reasoning. The culmination dataset: every "vibe coding" session, extracted from local agent storage, scrubbed, and packaged for SFT. [!IMPORTANT] Gated access. Access requests are reviewed manually. Data is anonymized… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/agentic-vibecoding-traces.tabulartext-generation10K<n<100K1 likes61 downloads13d agoHugging Face06violetxi /equational-theory-sair-note-conditioned-rollouts Equational-theory SAIR note-conditioned rollouts 169,606 teacher rollouts across completed R0–R4 cohorts, using the original table fields/types and layout of violetxi/harvey-note-conditioned-rollouts. Cohort Rollouts R0 17,408 R1 27,158 R2 33,840 R3 39,371 R4 51,829 Each cohort has one response per eligible task, sample index 0. Cohorts revisit tasks with updated note memories, so the total counts task/round instances, not distinct mathematical questions… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/equational-theory-sair-note-conditioned-rollouts.tabulartext-generation100K<n<1M0 likes55 downloads2d agoHugging Face07saidutta69 /unsolved-math-clean 🧠 Unsolved Math — Clean 8,626 curated open research problems in mathematics and CS — including 122 Millennium Prize Problems — deduplicated, schema-flattened, and packaged as proper parquet configs with an eval-only benchmark view. A reasoning frontier dataset: every problem here is actually unsolved or partially solved — ideal for honest capability probing instead of contaminated benchmarks. Clean derivative of ulamai/UnsolvedMath (8,785 problems). License unchanged:… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/unsolved-math-clean.tabularquestion-answering10K<n<100K0 likes51 downloads13d agoHugging Face08violetxi /equational-theory-sair-notes Equational-theory SAIR notes 103,276 current usable notes, with 60,967,993 title/body tokens, from the completed R0–R4 authoritative bank. This uses the JSONL/Parquet layout and original field names/types of violetxi/harvey-notes-v4. Only a recursive bank exists for this experiment; there is no inherited-only comparison bank. Round of current revision Notes R0 4,870 R1 24,666 R2 26,228 R3 26,148 R4 21,364 Files and schema… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/equational-theory-sair-notes.tabulartext-generation100K<n<1M0 likes42 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.