datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leetcode-problem-solutions
LeetCode Solution Dataset
This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling.
Column Descriptions
Column Name
Type
Description
question_slug
string
The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.GPT-5.6-Sol-Luna-Terra-Traces
GPT-5.6 — Sol · Terra · Luna Library
A maintained mirror of every GPT-5.6 Sol / Terra / Luna dataset on Hugging Face — content-verified, attributed, in one place.
Dataset Viewer | Parquet
// what this is
This is a maintained library — a community mirror of every publicly-available GPT-5.6 Sol / Terra / Luna dataset on Hugging Face, aggregated, validity-filtered, and content-verified with per-row source attribution. It is not Crownelius' own data. Every row… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GPT-5.6-Sol-Luna-Terra-Traces.swerl-tmax-15k-solvable-gpt-5-6-terra
swerl-tmax-15k hardened, post-validation-filter (dataset 3 of 3)
Which tasks in hamishivi/swerl-tmax-15k can a strong model actually solve? Every
task was attempted twice as a full agentic episode — real sandbox, real bash,
real verifier — and a task is verified when at least one attempt earned reward.
The last of three artifacts that exist to be compared by task_id:
original — hamishivi/swerl-tmax-15k, unchanged — 14,601 tasks
hardened, pre-validation-filter —… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-solvable-gpt-5-6-terra.Complete-FABLE.5-traces-2M
Complete FABLE.5 Traces (2 Million Deduplicated Rows)
Comprehensive Agentic Coding & Frontier Reasoning Trajectory Corpus
Executive Summary
Solstice-AI/Complete-FABLE.5-traces-2M is a clean, fully deduplicated post-training dataset containing 2,006,487 high-entropy agentic coding and multi-step reasoning traces.
Originally curated following the closure of Fable and Mythos, this corpus synthesizes frontier agent execution patterns (including Claude… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Complete-FABLE.5-traces-2M.leetcode-python-solutions-with-exaplanationsincremental-instruction-creative-writing
Incremental Instruction Creative Writing
Does delivering a writing brief over several conversation turns change what a
language model writes? This dataset supports that question with matched
creative-writing tasks evaluated under two delivery conditions:
FULL: the complete brief is supplied in one turn.
SHARDED: the same intended brief is introduced across five to nine turns.
The benchmark holds task content fixed while varying how the instructions are
delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.GPT-5.6-Sol-Luna-Terra-Traces
GPT-5.6 — Sol · Terra · Luna Library
A maintained mirror of every GPT-5.6 Sol / Terra / Luna dataset on Hugging Face — content-verified, attributed, in one place.
Dataset Viewer | Parquet
// what this is
This is a maintained library — a community mirror of every publicly-available GPT-5.6 Sol / Terra / Luna dataset on Hugging Face, aggregated, validity-filtered, and content-verified with per-row source attribution. It is not Crownelius' own data. It exists to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.6-Sol-Luna-Terra-Traces.qwen9b-solo-claude-code
qwen9b-solo-claude-code
Single-agent coding trajectories generated by running
CooperBench in solo mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the
agent framework. One agent implements both features in each task.
The matched coop (two-agent) version is at
CooperBench/qwen9b-coop-claude-code.
Same task corpus, same model, same agent — only the coordination differs, so
together they isolate the cooperation deficit.
At a… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-claude-code.Polaris-hard-w-solutions-24209
Polaris-Hard-w-Solutions
24,209 hard competition-math problems (the hardest difficulty bands of the
Polaris dataset) paired with
two verified solutions each: a full original solution and a concise summarized solution. Every
retained problem has a machine-verifiable final answer, every solution's boxed answer grades correct
against the reference (sympy-based grading), and the summarized solutions have additionally been put
through a reasoning-rigor pass (see step 5 below).
This… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/Polaris-hard-w-solutions-24209.qwen9b-solo-mini-swe-agent
qwen9b-solo-mini-swe-agent
Single-agent coding trajectories generated by running
CooperBench in solo mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework.
One agent implements both features in each task.
The matched coop version is at
CooperBench/qwen9b-coop-mini-swe-agent.
Same task corpus, same model, same agent — only the coordination differs, so
together they isolate the cooperation deficit.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-mini-swe-agent.historica-corpus
Historica Corpus v2
A monolingual pretraining corpus of 314,438 passages (177M words, 1.1B characters) spanning 15 ancient and historical languages, from 500 BCE to 1750 CE.
What's New in v2
2x more passages (314k vs 159k) due to varied-length chunking
SGML entity resolution: Middle English þ/ȝ/ð properly rendered (928k entities fixed)
PROIEL punctuation: reconstructed from presentation-after attributes
TEI apparatus handling: <lem> (main reading) preserved, <rdg>… See the full description on the dataset page: https://huggingface.co/datasets/sol-r/historica-corpus.nla-gemma4e2b-relabel-v1-corpus
Gemma-4-E2B layer-23 activation corpus, relabeled (v1)
1356 training rows for an activation verbalizer. Each row pairs a residual-stream
activation captured at layer 23 of google/gemma-4-E2B with a natural-language label
describing what the model must have integrated at that position to predict its next
token. This is the training set behind
Solshine/gemma-4-e2b-nla-L23-av-priordev-relabel-v1-wd3.
Why it exists
An audit of the previous version of this corpus found… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-relabel-v1-corpus.GPT-5.6-Sol-Luna-Terra-Traces
GPT-5.6 — Sol · Terra · Luna Library
A maintained mirror of every GPT-5.6 Sol / Terra / Luna dataset on Hugging Face — content-verified, attributed, in one place.
Dataset Viewer | Parquet
// what this is
This is a maintained library — a community mirror of every publicly-available GPT-5.6 Sol / Terra / Luna dataset on Hugging Face, aggregated, validity-filtered, and content-verified with per-row source attribution. It is not Crownelius' own data. It exists to… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/GPT-5.6-Sol-Luna-Terra-Traces.nanochat-brevo-capability-data-10x
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for the complete recursive
prerequisite closure in a valid leaf-first order. Training uses project-planning
language; validation uses evidence synthesis; test uses build manifests. Eleven
deterministic structural styles vary wording, layout, and record order.
The latent graph generator and exact validator label every row. No language model
generated or labeled the data. Alternative valid orders… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-10x.nanochat-brevo-capability-data
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for the complete recursive
prerequisite closure in a valid leaf-first order. Training uses project-planning
language; validation uses evidence synthesis; test uses build manifests. Six
deterministic structural styles vary wording, layout, and record order.
The latent graph generator and exact validator label every row. No language model
generated or labeled the data. Alternative valid orders are… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data.ac-solver-dataset
AC-Solver Transformer Dataset
Training corpus for a decoder-only Transformer language model on the
Andrews-Curtis (AC) conjecture — one of the longest-standing open problems in
combinatorial group theory, and a challenging benchmark for reinforcement learning.
Paper: What Makes Math Problems Hard for Reinforcement Learning: A Case Study
Code: github.com/shehper/AC-Solver
Background: The Andrews-Curtis Conjecture
A balanced presentation of a group is a description… See the full description on the dataset page: https://huggingface.co/datasets/mhieuuu/ac-solver-dataset.gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit
Gemma-4-E2B NLA AV-SFT Training Corpus (v0.1.x, Gemini persona+audit)
The 4,734-row AV-SFT training corpus for the v0.1.x Gemma-4-E2B NLA — a 9-source-family diversified expansion over the v0.0.x OpenWebText-only corpus. Labels generated by Gemini CLI following the persona+audit pipeline (Dr. Marisol Chen labels, Dr. Riley Otsuka audits).
This is the in-progress v0.1.x labeled training set. AR-SFT companion is still being labeled (~16% complete as of this dataset publish). When the… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit.gemma-4-e2b-deception-behavior-completions
Gemma-4-E2B deception & behavior completions
Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included.
The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.autoteacher-dapo-claude-solved-coded
autoteacher-dapo-claude-solved-coded
zjhhhh/autoteacher-dapo-claude-solved (512 DAPO math problems with Claude-written, student-verified hints) augmented with a strategy coding of every hint against a compact codebook of 74 reusable problem-solving strategies.
Companion codebook dataset: zjhhhh/autoteacher-dapo-codebook.
Added columns
column
type
description
hint_code_indices
list[int]
code_ids (1–74) of the codebook strategies that form the core of the… See the full description on the dataset page: https://huggingface.co/datasets/zjhhhh/autoteacher-dapo-claude-solved-coded.DSA-Coding-Problems-and-Solutions-Dataset
Dataset Description
This dataset is a large-scale collection of Data Structures and Algorithms (DSA) code, containing 12,385 code files with 3.86 million lines of code and 25.01 million lexical tokens, designed to support the development of advanced code generation models, programming assistants, software engineering AI systems, and code intelligence applications.
It consists of real-world DSA implementations covering a wide range of algorithms, data structures, problem-solving… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/DSA-Coding-Problems-and-Solutions-Dataset.gemma-4-e2b-nla-eval-smoke
Gemma-4-E2B NLA smoke-eval (20-row held-out set)
A 20-row held-out subset of OpenWebText activations extracted from google/gemma-4-E2B at layer 23. Used as the canonical eval set for smoke-testing the v0.0.1 Gemma-4-E2B NLA pair on a fresh environment.
This dataset is a subset of the held-out rl.parquet evaluation set used for the v0.0.1 round-trip eval (n=50 attempted, 42 evaluated after 8 empty-output exclusions, cos 0.438 ± 0.054). The 20-row subset preserves the activation… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-eval-smoke.gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit
Gemma-4-E2B NLA AR-SFT Training Corpus (v0.0.x, Claude Haiku persona+audit)
The 696-row AR-SFT training corpus used for the Option B Gemma-4-E2B NLA pair. Labels generated by Claude Haiku 4.5 following the persona+audit pipeline — Dr. Marisol Chen (synthetic mech-interp expert) labels first, Dr. Riley Otsuka (synthetic senior editor) audits the labels.
This is the matched companion to the v0.0.x AV labeled corpus. The pair completes the first open-source non-Anthropic-team NLA… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit.organic-chemistry-synthesis-planning-corpus
Organic Chemistry Synthesis Planning Corpus
Status: actively ingesting. A comprehensive reaction backbone is already uploaded
(millions of reactions; see Ingested data below). Curation and
additional sources are ongoing. See Roadmap.
Quickstart (for students / first-time users)
You need a free Hugging Face account, and to accept this dataset's terms on its page (it's gated).
pip install datasets transformers
huggingface-cli login
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/organic-chemistry-synthesis-planning-corpus.nanochat-tool-routing-v1-65k-20260714
Nanochat Tool Routing and Continuation
This deterministic corpus teaches a decoder to choose among four declared
functions, answer directly when the request already contains the answer, ask for
missing required arguments, and continue after a masked tool result. Because
this is pretraining rather than SFT, the natural system and user text remains
ordinary supervised language-model data. Only external tool results are visible
context excluded from causal-LM loss.
Train examples:… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-tool-routing-v1-65k-20260714.stackoverflow-ai-solvedUpdated 1/19/2024: More than doubled the number of answers to 1,469,201, with a higher percent of 4o-mini and gemini 1.5 pro.
A dataset which comprises of ai answers to 1,469,201 stackoverflow questions relating to Python.
The questions were extracted from this dataset.
All responses are directly python, no codeblocks or anything.
A total of 545,697,115 input and 253,110,685 output o200k_base tokens (gpt-4o/4o-mini).
Model
Value
gemini-1.5-pro-002
442,261
gpt-4o-mini-2024-07-18… See the full description on the dataset page: https://huggingface.co/datasets/tennisb/stackoverflow-ai-solved.tons-of-skills
Tons of Skills — Claude Code Skills + Quality Grades
A browsable corpus of ~3,000 Claude Code / Agent Skills from the
Tons of Skills marketplace, each joined with its quality
grade and behavioral-eval verdict from the Freshie compliance CMDB.
One row per skill: the skill's content and frontmatter plus its letter grade
(A–F), numeric score, and JRig behavioral-eval pass flag.
Why the grades are here
Most "awesome skills" lists are unranked. This dataset ships the… See the full description on the dataset page: https://huggingface.co/datasets/intent-solutions-io/tons-of-skills.nanochat-brevo-capability-v4-49k-20260714
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for complete recursive
prerequisite closures in valid leaf-first orders. Each compact training document
reuses one graph for 4 worked questions, increasing answer
supervision without repeating the graph. Training balances
project_plan, build_manifest language; validation uses held-out
evidence synthesis; test uses build manifests. Eleven deterministic
structural styles vary wording, layout, and… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-v4-49k-20260714.big-math-ppo-mix-40k-qwq-solutions
Big Math PPO Mix 40k QwQ Solutions
Teacher solutions generated by Qwen/QwQ-32B-Preview for the Big Math PPO Mix
prompts, intended for knowledge distillation into a student policy.
This is the combined dataset, merging:
big-math-ppo-mix-30k prompts → 22,584 correct solutions (75.3% of 30,000)
big-math-ppo-mix-extra-10k prompts → 7,986 correct solutions (79.9% of 10,000)
Total: 30,570 verified-correct solutions (40,000 prompts attempted).
Generation
vLLM sampling… See the full description on the dataset page: https://huggingface.co/datasets/christinakopi/big-math-ppo-mix-40k-qwq-solutions.solidity-dataset
Solidity Dataset
Dataset Description
This dataset is collected from public GitHub repositories written in Solidity programming language.
The list of the repositories is available at repositories.json file.
It contains useful data about smart contracts written in Solidity along with test cases (and unit tests) written to test smart contracts.
Dataset Summary
The dataset contains of 355,540 rows in total. Each row includes the following features:
hash… See the full description on the dataset page: https://huggingface.co/datasets/seyyedaliayati/solidity-dataset.nanochat-brevo-capability-data-v2
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for complete recursive
prerequisite closures in valid leaf-first orders. Each compact training document
reuses one graph for 4 worked questions, increasing answer
supervision without repeating the graph. Training uses project-planning language;
validation uses evidence synthesis; test uses build manifests. Eleven deterministic
structural styles vary wording, layout, and record order. Per-world… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-v2.
