datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenCodeReasoning
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
Data Overview
OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming
questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT).
Technical Report - Discover the methodology and technical details behind OpenCodeReasoning.
Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenCodeReasoning.opc-sft-stage2
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2 <-- you are here
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing
opc-fineweb-code-corpus: the code-related page recalled from fineweb
opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-sft-stage2.opc-annealing-corpus
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing <-- you are here
fineweb-code-corpus: the code-related page recalled from fineweb
fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data of… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-annealing-corpus.OpenCodeReasoning-2
OpenCodeReasoning-2: A Large-scale Dataset for Reasoning in Code Generation and Critique
Dataset Description
OpenCodeReasoning-2 is the largest reasoning-based synthetic dataset to date for coding, comprising 1.4M samples in Python and 1.1M samples in C++ across 34,799 unique competitive programming questions.
OpenCodeReasoning-2 is designed for supervised fine-tuning (SFT) tasks of code completion and code critique.
Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenCodeReasoning-2.opc-fineweb-code-corpus
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing
opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here
opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.opc-sft-stage1
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1 <-- you are here
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing
opc-fineweb-code-corpus: the code-related page recalled from fineweb
opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-sft-stage1.RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode!
Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available.
RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.OpenCodeReasoning2OpenCodeReasoning_messages
This is a transformation of the nvidia/OpenCodeReasoning dataset into a format that is more easily digestible by trainers.
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
Data Overview
OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming
questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT).
Technical Report -… See the full description on the dataset page: https://huggingface.co/datasets/sealad886/OpenCodeReasoning_messages.opencoderinstruct_trajectoryOpenCodeReasoning-2
OpenCodeReasoning-2: A Large-scale Dataset for Reasoning in Code Generation and Critique
Dataset Description
OpenCodeReasoning-2 is the largest reasoning-based synthetic dataset to date for coding, comprising 1.4M samples in Python and 1.1M samples in C++ across 34,799 unique competitive programming questions.
OpenCodeReasoning-2 is designed for supervised fine-tuning (SFT) tasks of code completion and code critique.
Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/cublya/OpenCodeReasoning-2.opc-fineweb-math-corpus
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing
opc-fineweb-code-corpus: the code-related page recalled from fineweb
opc-fineweb-math-corpus: the math-related page recalled from fineweb <-- you are here
refineCode-code-corpus-meta: the… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-math-corpus.OpenCodeReasoning-formattednvidia-OpenCodeReasoningopencode_reasoning_filtered
🧠 OpenCode Reasoning (Filtered)
Author: Parvesh Rawal — XenArcAILicense: Inherits from NVIDIA OpenCodeReasoningVersion: Filtered & Structured VariantTotal Examples: 567,850Total Size: 9GB (compressed)
🔍 Overview
This dataset is a curated and cleaned version of split_0 from nvidia/OpenCodeReasoning, optimized for code-level reasoning tasks and instruction tuning.
It’s designed to enhance logic understanding and multistep problem solving for LLMs.
📁 Features… See the full description on the dataset page: https://huggingface.co/datasets/Parveshiiii/opencode_reasoning_filtered.OpenCodeReasoning_ShareGPT
Added cnversations column in ShareGPT format
Original README from nvidia/OpenCodeReasoning
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
Data Overview
OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming
questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT).
Technical Report - Discover the methodology and… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/OpenCodeReasoning_ShareGPT.opencode-rollout-trace
One opencode rollout, as the trainer sees it
The artefact behind the talk Training a coding agent through a harness you did not write
(Lisbon AI, September 2026).
An off-the-shelf coding agent, opencode, runs untouched in a
remote sandbox. A proxy sits between it and vLLM, speaks the agent's own dialect, and records every
model call the agent makes together with the exact token ids and logprobs. That recording is what
gets trained on. This repo is one of those recordings, plus… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/opencode-rollout-trace.load_in_code_opencodereasoning_hfOpenCodeReasoningRubrics
OpenCodeReasoningRubrics
This dataset contains questions and rubric annotations intended for evaluating reasoning quality in open-ended coding or logical tasks.
Dataset Structure
Each example in the dataset has the following fields:
index (int): A unique identifier.
question (str): A natural language question or prompt.
rubric (str): An explanation or rubric detailing expectations or evaluation criteria.
The data is stored in split .parquet files for efficient loading… See the full description on the dataset page: https://huggingface.co/datasets/danikhan632/OpenCodeReasoningRubrics.OpenCodeReasoning_len8k_0.5opencodereasoning-sharegptBased off of Nvidia's new code datset.
Generations were made by deepseek r1, inputs are labeled in the source column.
dataset_info:
features:
name: id
dtype: string
name: input
dtype: string
name: source
dtype: string
name: license
dtype: string
name: dataset
dtype: string
name: split
dtype: string
name: difficulty
dtype: string
name: solution
dtype: string
name: thinking
dtype: string
name: attempt
dtype: string
name: conversations
list:
name: from
dtype: string
name: value
dtype: string… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstack/opencodereasoning-sharegpt.OpenCodeReasoning_len8k_0.6opencodereasoning-100kOpenCodeReasoning_len8kOpenCoder-LLM_opc-sft-stage2数据来源:OpenCoder-LLM/opc-sft-stage2
转存时间:2025/1/3
处理人:peijin
OpenCoder-LLM_opc-sft-stage1-DolphinLabeled
OpenCoder-LLM SFT DolphinLabeled
Part of the DolphinLabeled series of datasets
Presented by Eric Hartford and Cognitive Computations
The purpose of this dataset is to enable filtering of OpenCoder-LLM SFT dataset.
The original dataset is OpenCoder-LLM/opc-sft-stage1
I have modified the dataset using two scripts.
dedupe.py - removes rows with identical instruction
label.py - adds a "flags" column containing the following boolean values:
"refusal": whether the… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/OpenCoder-LLM_opc-sft-stage1-DolphinLabeled.open-code-reasoning-sftOpenCoder-LLM_opc-sft-stage1数据来源:OpenCoder-LLM/opc-sft-stage1
转存时间:2025/1/3
处理人:peijin
OpenCodeReasoning-split_1-ConvertedOpenCodeReasoning-Sampled
