datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leetcode-problem-solutions
LeetCode Solution Dataset
This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling.
Column Descriptions
Column Name
Type
Description
question_slug
string
The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.gpt-5.6-sol-coding-and-debugging-traces
GPT-5.6 Sol Coding & Debugging Traces
Verified software-engineering, independent model-judging, seed-authoring,
defensive-security, and training-harness trajectories from
GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an
autonomous coding agent. Sessions show the observable development loop:
inspecting repositories, reproducing failures, explaining evidence, editing
files, running compilers and test suites, correcting mistakes, and verifying
the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/gpt-5.6-sol-coding-and-debugging-traces.GPT-5.6-Sol-Luna-Terra-Traces
GPT-5.6 — Sol · Terra · Luna Library
A maintained mirror of every GPT-5.6 Sol / Terra / Luna dataset on Hugging Face — content-verified, attributed, in one place.
Dataset Viewer | Parquet
// what this is
This is a maintained library — a community mirror of every publicly-available GPT-5.6 Sol / Terra / Luna dataset on Hugging Face, aggregated, validity-filtered, and content-verified with per-row source attribution. It is not Crownelius' own data. Every row… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GPT-5.6-Sol-Luna-Terra-Traces.NuminaMath-LEAN-Sol
NuminaMath-LEAN Cleaned with NL Solutions
Dataset Summary
This is a cleaned version of the NuminaMath-LEAN dataset, enhanced with natural language (NL) solutions matched from source datasets. The primary goal is to provide paired formal statements/proofs with natural language solutions for proof formalization and theorem proving research.
The dataset matches problems from NuminaMath-LEAN with their corresponding natural language solutions from:
olympiads-ref: A… See the full description on the dataset page: https://huggingface.co/datasets/iiis-lean/NuminaMath-LEAN-Sol.Python-Code-Solutions
Python Code Solutions
Features
1000k of Python Code Solutions for Text Generation and Question Answering
Python Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
doocs-leetcode-solutions
Doocs LeetCode Solutions
A comprehensive dataset of LeetCode problems and solutions created from the Doocs LeetCode repository. This dataset is designed for fine-tuning large language models to understand programming problems and generate code solutions.
Description
Repository: Doocs LeetCode Solutions
Total Problems: 3500+
Total Solutions: 15,000+ (across multiple languages)
Size: ~60 MB (Parquet format)
Languages:
C
Cangjie
C++
C#
Dart
Go
Java
JavaScript
Kotlin
Nim
PHP… See the full description on the dataset page: https://huggingface.co/datasets/olegshulyakov/doocs-leetcode-solutions.BenchMAX_Problem_Solving
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Problem_Solving is a dataset of BenchMAX, sourcing from LiveCodeBench_v4, which evaluates the code generation capability for solving multilingual competitive code problems.
We extend the original English dataset by 16 non-English languages.
The… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Problem_Solving.Solace-1.0-Omni
Project Solace
The largest verified frontier-model distillation corpus ever released.
60 datasets · 7 frontier model families · 12,586,893 unique conversations · One file · Zero filler
The short version
This is synthetic data. The best kind of synthetic data.
Every example was generated by a verified 2026 frontier model — GLM-5.2, Claude Fable 5, Mythos 5, GPT-5.6 Sol, GPT-5.5 Codex, DeepSeek V4 Pro 0813, Qwen 3.8-Max, and Kimi K3 — then… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Solace-1.0-Omni.xrepotest
XRepoTest: Multilingual Repository-Level Unit Test Generation Benchmark
Paper | Code
The data release for the XRepoTest benchmark (EMNLP 2026 Main).
Each task item is a function extracted from a real open-source repository. The
goal is to generate a unit test for that function using its surrounding repository
context. Generated tests are executed and scored inside Docker on the real codebase.
Structure
Folder
Contents
Pipelines
data/base
per-language… See the full description on the dataset page: https://huggingface.co/datasets/solis-soict/xrepotest.swerl-tmax-15k-solvable-gpt-5-6-terra
swerl-tmax-15k hardened, post-validation-filter (dataset 3 of 3)
Which tasks in hamishivi/swerl-tmax-15k can a strong model actually solve? Every
task was attempted twice as a full agentic episode — real sandbox, real bash,
real verifier — and a task is verified when at least one attempt earned reward.
The last of three artifacts that exist to be compared by task_id:
original — hamishivi/swerl-tmax-15k, unchanged — 14,601 tasks
hardened, pre-validation-filter —… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-solvable-gpt-5-6-terra.swerl-tmax-15k-rubric-gpt-5-6-sol
swerl-tmax-15k with task-quality rubric labels (gpt-5-6-sol)
hamishivi/swerl-tmax-15k, unchanged and unfiltered, with a per-task quality
label attached as extra columns.
This is not a verified or filtered dataset. Every one of the 14,601 original
records is present. Nothing has been dropped, repaired, or reordered. The labels
are one model's judgement about whether each task is sound enough to be useful RL
training data — an annotation layer, not a correctness guarantee.… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-rubric-gpt-5-6-sol.shellcode_i_a32Shellcode_IA32 is a dataset for shellcode generation from English intents. The shellcodes are compilable on Intel Architecture 32-bits.Axiom-1.0-Opus4.7-Kimi2.6-GLM5.2-Deepseek4-Mythos5-Fable5-Qwen3.7
Project Axiom 1.0 (102 GB Reasoning Corpus)
27-Billion Token Pure-Text Chain-of-Thought Corpus Across 7 Frontier Architectures
Executive Summary
Project Axiom 1.0 is a landmark, high-density, multi-architecture reasoning corpus comprising 102 GB of uncompressed, pure-text JSONL data (axiom.jsonl). Curated by Shreyan Gondaliya and the Solstice-AI research team, the dataset synthesizes ~5.74 million unique samples and ~27.3 billion tokens of… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Axiom-1.0-Opus4.7-Kimi2.6-GLM5.2-Deepseek4-Mythos5-Fable5-Qwen3.7.leetcode_problem_solutionThis dataset contains: problems and solutions in Leetcode, crawled from: https://github.com/AnasImloul/Leetcode-Solutions
The format of data:
title: title of the problem
algo_input: the description of the problem
solution_py: the solution in Python
solution_js: the solution in Js
solution_java: the solution in Java
solution_c: the solution in C
Complete-FABLE.5-traces-2M
Complete FABLE.5 Traces (2 Million Deduplicated Rows)
Comprehensive Agentic Coding & Frontier Reasoning Trajectory Corpus
Executive Summary
Solstice-AI/Complete-FABLE.5-traces-2M is a clean, fully deduplicated post-training dataset containing 2,006,487 high-entropy agentic coding and multi-step reasoning traces.
Originally curated following the closure of Fable and Mythos, this corpus synthesizes frontier agent execution patterns (including Claude… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Complete-FABLE.5-traces-2M.perovskite-solar-cell-efficiency-autoresearch
🔬 Perovskite Solar Cell Text Corpus for Karpathy's autoresearch
A 98.9 MB text corpus of perovskite solar cell scientific literature formatted for direct use with karpathy/autoresearch — the autonomous LLM-driven hyperparameter search framework that trains a GPT from scratch and has an AI agent iteratively modify train.py to minimize val_bpb (bits per byte).
📊 Dataset Stats
Metric
Value
Total documents
19,730
Total text
98.9 MB (~103M characters)… See the full description on the dataset page: https://huggingface.co/datasets/CollinL/perovskite-solar-cell-efficiency-autoresearch.leetcode-python-solutions-with-exaplanationsSmall-HLE-Solved
Small-HLE-Solved
Small-HLE-Solved is a curated dataset consisting of challenging problems selected from the Humanity's Last Exam (HLE) benchmark. Each instance has been processed by an advanced teacher model to generate high-fidelity, multi-step reasoning paths. The dataset is formatted strictly in JSON Lines (jsonl), pairing each complex problem with a structured, step-by-step solution optimized for training next-generation reasoning models.
📂 Data Structure &… See the full description on the dataset page: https://huggingface.co/datasets/Axiom-AI/Small-HLE-Solved.solana-clawd-model-kit
Solana Clawd Model Kit
Training data kit for Solana Clawd: SFT / CPT JSONL corpora, manifests, quality reports, and processed shards.
Contents (top-level)
SFT / CPT JSONL
solana_clawd_reasoning_tooling_sft.jsonl (~133 MB)
clawd_masterpiece_sft.jsonl (~166 MB)
tx_foundation_cpt_clean.jsonl (~21 MB)
clawd_future_refinement_sft.jsonl (~1.3 MB)
clawd_autoresearch_wiki_sft.jsonl (~1.3 MB)
clawd_future_drill_sft.jsonl (~920 KB)… See the full description on the dataset page: https://huggingface.co/datasets/ordlibrary/solana-clawd-model-kit.solana-clawd-instruct
Solana Clawd Instruct
A curated instruction-tuning dataset for fine-tuning models into Solana-native Clawd agents with strong Solana, DeFi, ZK, and constitutional-alignment coverage.
What it teaches
Check every domain your dataset covers:
Solana mechanics (PDAs, accounts, instructions, rent, compute budgets, Token-2022)
DeFi primitives (AMMs, CLMMs, perpetuals, bonding curves, Jupiter, Phoenix)
Memecoin risk analysis (rug detection, holder concentration… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-instruct.solidity-audit-cot
solidity-audit-cot
Long-CoT audit traces for Solidity contracts, generated by Claude Opus 4.7 (adaptive thinking, xhigh effort) over the spec→contract corpus from the Qwopus3.6-27B-solidity training pipeline.
This dataset is the Stage 2 training corpus for the multi-stage Qwopus3.6-27B-solidity model — designed to teach long-form security reasoning (8-15 paragraph chain-of-thought) anchored to real Solidity contracts.
Why this dataset exists
Public Solidity audit… See the full description on the dataset page: https://huggingface.co/datasets/samscrack/solidity-audit-cot.incremental-instruction-creative-writing
Incremental Instruction Creative Writing
Does delivering a writing brief over several conversation turns change what a
language model writes? This dataset supports that question with matched
creative-writing tasks evaluated under two delivery conditions:
FULL: the complete brief is supplied in one turn.
SHARDED: the same intended brief is introduced across five to nine turns.
The benchmark holds task content fixed while varying how the instructions are
delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.math-sft-solutions-no-cot
Math SFT Solutions No CoT
A cleaned mathematics supervised fine-tuning dataset containing:
instruction → solution pairs
mathematical proofs
derivations
olympiad-style solutions
theorem reasoning
stepwise mathematical explanations
detailed final solutions
This dataset was built specifically for mathematical supervised fine-tuning (SFT).
Unlike many reasoning datasets, this release removes explicit chain-of-thought tags and hidden thinking traces while preserving high-quality… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/math-sft-solutions-no-cot.gpt-5.6-sol-coding-and-debugging-traces
GPT-5.6 Sol Coding & Debugging Traces
Verified software-engineering, independent model-judging, seed-authoring,
defensive-security, and training-harness trajectories from
GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an
autonomous coding agent. Sessions show the observable development loop:
inspecting repositories, reproducing failures, explaining evidence, editing
files, running compilers and test suites, correcting mistakes, and verifying
the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/gpt-5.6-sol-coding-and-debugging-traces.GPT5.6_SOL_INVESTIGACION
Dataset de Metodología Científica
Dataset en español para entrenamiento, validación y evaluación de modelos capaces de razonar sobre metodología de investigación científica. Incluye escenarios de distintas disciplinas y niveles de dificultad, con énfasis en diseño de estudios, inferencia causal, análisis cuantitativo y cualitativo, métodos mixtos, ética, medición, muestreo, interpretación de resultados y revisión crítica de protocolos.
1. Resumen… See the full description on the dataset page: https://huggingface.co/datasets/Januka2009/GPT5.6_SOL_INVESTIGACION.SWEbench-Verified-eval150-M2.7-solo-selforch-3repeats-w32-20260920
M2.7 Solo and self-orchestration: three independent eval150 runs each
All six fresh runs completed the same150 tasks and passed original result/trajectory/task/attempt/fingerprint audits. No previous scores were pooled. Each repeat starts new model processes and cold KV caches after a real telemetry smoke. True failed tasks are retained; infrastructure retries are preserved separately.
Mode
Repeat1
Repeat2
Repeat3
Mean /150
Sample SD
m27-solo
94
98
90
94.00
4.00… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-solo-selforch-3repeats-w32-20260920.solana-clawd-repo-corpus
Solana Clawd Core AI Instruct
Instruction-tuning dataset derived from the local core-ai source tree and the
existing Solana Clawd AI training corpus.
Contents
Total examples: 441
Existing ai-training SFT examples: 0
Core AI source chunk examples: 0
Core AI knowledge JSONL examples: 0
Format
Each row is a chat conversation in OpenAI/Hugging Face messages schema:
{"messages": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-repo-corpus.GPT-5.6-Sol-Luna-Terra-Traces
GPT-5.6 — Sol · Terra · Luna Library
A maintained mirror of every GPT-5.6 Sol / Terra / Luna dataset on Hugging Face — content-verified, attributed, in one place.
Dataset Viewer | Parquet
// what this is
This is a maintained library — a community mirror of every publicly-available GPT-5.6 Sol / Terra / Luna dataset on Hugging Face, aggregated, validity-filtered, and content-verified with per-row source attribution. It is not Crownelius' own data. It exists to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.6-Sol-Luna-Terra-Traces.math-sft-solutions-no-cot-v3
Math SFT Solutions No CoT V3
Math SFT Solutions No CoT V3 is a large-scale mathematics supervised fine-tuning (SFT) dataset designed for instruction tuning and mathematical capability adaptation.
Version 3 substantially expands mathematical coverage while improving dataset quality through stronger filtering, cleaning, and supervision refinement.
Unlike reasoning-heavy datasets, this release focuses on clean instruction → response pairs without hidden chain-of-thought style… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/math-sft-solutions-no-cot-v3.solidity-cpt-top10-quality
Solidity CPT Top-10% Quality-Filtered Corpus
A curated, deduplicated corpus of 23,471 modern Solidity source files (~86M tokens) intended for continued-pretraining (CPT) of code LLMs on smart-contract code.
It's the top 10% slice (by composite quality score) of a larger raw corpus that combined:
ASSERT-KTH/DISL — 514 k unique deployed Solidity files, deduped at file level
30 hand-picked GitHub blue-chip protocols (OpenZeppelin, Uniswap v2/v3/v4, Aave v3, Compound, Morpho… See the full description on the dataset page: https://huggingface.co/datasets/samscrack/solidity-cpt-top10-quality.
