datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opc-annealing-corpus
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing <-- you are here
fineweb-code-corpus: the code-related page recalled from fineweb
fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data of… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-annealing-corpus.OpenCodeReasoning2opencoderinstruct_trajectoryOpencode1OpenCodeInstruct_length_sortedOpenCodeInstruct_training_data_n32w16Model (trajectory generator): Jacobi Forcing model trained with n16w16 trajectories
Data Source: OpenCodeInstruct
Noise Schedule: linear progressive
Block size: 32
Window size: 32, 16
OpenCoder-LLM_opc-sft-stage1-DolphinLabeled
OpenCoder-LLM SFT DolphinLabeled
Part of the DolphinLabeled series of datasets
Presented by Eric Hartford and Cognitive Computations
The purpose of this dataset is to enable filtering of OpenCoder-LLM SFT dataset.
The original dataset is OpenCoder-LLM/opc-sft-stage1
I have modified the dataset using two scripts.
dedupe.py - removes rows with identical instruction
label.py - adds a "flags" column containing the following boolean values:
"refusal": whether the… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/OpenCoder-LLM_opc-sft-stage1-DolphinLabeled.opencodeinstruct_data_v1_progressive_cyclic_noise_0828OpenCodeInstruct_training_data_n16w16Model (trajectory generator): Qwen2.5-Coder-7B-Instruct
Data Source: OpenCodeInstruct
Noise Schedule: linear progressive
Block size: 16
Window size: 16
OpenCoder-LLM_opc-sft-stage2-DolphinLabeled
OpenCoder-LLM SFT DolphinLabeled
Part of the DolphinLabeled series of datasets
Presented by Eric Hartford and Cognitive Computations
The purpose of this dataset is to enable filtering of OpenCoder-LLM SFT dataset.
The original dataset is OpenCoder-LLM/opc-sft-stage2
I have modified the dataset using two scripts.
dedupe.py - removes rows with identical instruction
label.py - adds a "flags" column containing the following boolean values:
"refusal": whether the… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/OpenCoder-LLM_opc-sft-stage2-DolphinLabeled.deepseek-v3.2-speciale-OpenCodeReasoning-3kThe questions for this dataset were all sourced from the first 3k prompts in nvidia/OpenCodeReasoning
Dataset Stats (provided by OpenRouter):
Cost: $ 19.2 (USD)
Tokens (input + output): 47 M
nvidia-OpenCodeReasoning-short-n-easyNemotron-3-Ultra-High-Effort-OpenCode-Distilled-2kJust under 2000 CoT synthetic prompt chains distilled from Nemotron 3 Ultra w/ High effort gathered using OpenCode entirely for free. Using a propertary harness developed by me
but currently private, partially synthetic or fully sythentic datasets can be created from any of the free models offered by OpenCode, with respect to rate-limiting. In this case,
a fully sythentic dataset of 1800ish rows was generated by "nemotron-3-ultra-free" using the maximium high effort reasoning. This was done by… See the full description on the dataset page: https://huggingface.co/datasets/TitleOS/Nemotron-3-Ultra-High-Effort-OpenCode-Distilled-2k.nemotron-sft-opencode-filtered-250k-v0.1Base dataset: nvidia/Nemotron-SFT-OpenCode-v1
Exact-deduplication + Filtered for the following categories:
code_explanation
code_generation
planning_and_task_structuring
reasoning
code_debugging
advice
brainstorming
code_review
unit_test_generation
summarization
code_refactoring
rewriting_and_editing
OpenCodeReasoning-Cleaned-GRPO
OpenCodeReasoning-Cleaned-GRPO
Deep-cleaned for GRPO/RL training | 500 examples | 525 bugs fixed
📋 Dataset Description
Code reasoning and critique prompts extracted from NVIDIA's OpenCodeReasoning-2 dataset. This cleaned version removes formatting artifacts, HTML tags, and whitespace issues that would degrade GRPO/RL training quality.
Original source: nvidia/OpenCodeReasoning-2 by NVIDIA
📊 Cleaning Statistics
Metric
Value
Original… See the full description on the dataset page: https://huggingface.co/datasets/Eyght/OpenCodeReasoning-Cleaned-GRPO.openclaw-opencode-datasetopencodegeneticinstruct-max-512-tokens-25kgrpo-opencoder-50k
grpo-opencoder-50k
This dataset was created by converting another dataset to JSONL format.
Files
dataset.jsonl: Dataset in JSONL format
Usage
from datasets import load_dataset
dataset = load_dataset("Nutanix/grpo-opencoder-50k", data_files="dataset.jsonl")
openclaw-opencode-datasetam-session-opencode-chk-26635c
Agent Manager session — opencode chk
One opencode session, exported from Agent Manager.
Access: gated. The repo and this card are listed publicly, but the trace itself requires per-user approval.
At a glance
1 prompts · 1 assistant turns · 0 tool calls
Trace: ses_053816c47ffexP5QoMUR12Xj2q.jsonl
What's in here
Path
Contents
ses_053816c47ffexP5QoMUR12Xj2q.jsonl
the session, native opencode JSONL
meta/manifest.json
provenance: harness… See the full description on the dataset page: https://huggingface.co/datasets/thomwolf/am-session-opencode-chk-26635c.opencodeinstruct_data_v1_random_noise_0827raw_merged_data_opencodeinstruct_8_28_trajectory_historydeepseek-v3.2-speciale-OpenCodeReasoning-3kThe questions for this dataset were all sourced from the first 3k prompts in nvidia/OpenCodeReasoning
Dataset Stats (provided by OpenRouter):
Cost: $ 19.2 (USD)
Tokens (input + output): 47 M
grpo-opencoder-small
grpo-opencoder-small
This dataset was created by converting another dataset to JSONL format.
Files
dataset.jsonl: Dataset in JSONL format
Usage
from datasets import load_dataset
dataset = load_dataset("Nutanix/grpo-opencoder-small", data_files="dataset.jsonl")
grpo-opencoder-mini
grpo-opencoder-mini
This dataset was created by converting another dataset to JSONL format.
Files
dataset.jsonl: Dataset in JSONL format
Usage
from datasets import load_dataset
dataset = load_dataset("Nutanix/grpo-opencoder-mini", data_files="dataset.jsonl")
opencodegeneticinstruct-max-512-tokens-10kopen-code-instruct-75kopencodeinterpreter_user_dataspider-rollouts-openthoughts-3-qwen32b-opencodereasoning-split-0-512opencode-meta-cognitive-dataset
