datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opc-fineweb-code-corpus
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing
opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here
opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode!
Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available.
RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.Opencode1OpenCodeInstruct-Clean
OpenCodeInstruct Clean
High-quality Python code generation dataset with duplication markers and complexity metrics.
Derived from nvidia/OpenCodeInstruct
after applying strict quality gates.
Quick Stats
Metric
Value
Total rows
388,629
Columns
54
Python-parsable
100.0%
Overview
This dataset contains 388,629 high-quality Python code generation examples
extracted from the nvidia/OpenCodeInstruct corpus.
Each row has been… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/OpenCodeInstruct-Clean.opc-fineweb-math-corpus
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing
opc-fineweb-code-corpus: the code-related page recalled from fineweb
opc-fineweb-math-corpus: the math-related page recalled from fineweb <-- you are here
refineCode-code-corpus-meta: the… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-math-corpus.Qwen3-Coder-Next-OpenCode-Preference
Dataset Card — OpenCode Rejection Sampling (Preference)
Overview
This dataset contains 10,920 preference pairs for preference-based training (DPO, KTO, SimPO, ORPO, etc.) on competitive programming tasks. Each pair consists of:
Chosen: a candidate solution that passes 100% of test cases
Rejected: a candidate solution that fails, with a fine-grained rejection type label
Pairs are produced via rejection sampling with Qwen3-Coder-Next: 8 candidate solutions are… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3-Coder-Next-OpenCode-Preference.opencode-public-data-pack-docker_input1-20-renders-512-20260612t125042z
OpenCode Public Data Pack docker_input1 20 renders 512 20260612T125042Z
Public data pack created from docker_input1.json with 20 renders at 512x512.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/opencode-public-data-pack-docker_input1-20-renders-512-20260612t125042z.open_codes
Open Codes
Open dataset of French legal code articles with embeddings.
Chunked articles from French legal codes sourced from
Legifrance via the PISTE API,
with 1024-dimensional embeddings generated by Mistral AI.
Each row is a text chunk enriched with full article metadata from the parent
legal code article.
Legal codes included
The list of legal codes is dynamic and managed in the LEX_codes_piste table.
Adding a new code there (with actif=true) automatically… See the full description on the dataset page: https://huggingface.co/datasets/ArthurSrz/open_codes.opencodereasoning_1k_eval_636d
mlfoundations-dev/opencodereasoning_1k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
6.7
13.2
24.8
26.4
9.6
36.2
31.9
9.0
10.7
AIME24
Average Accuracy: 6.67% ± 1.76%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
3.33%
1
30
2
3.33%
1
30
3
13.33%
4
30
4
6.67%
2
30
5… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_1k_eval_636d.opencodereasoning_3k_eval_636d
mlfoundations-dev/opencodereasoning_3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
8.3
26.2
37.8
31.6
21.2
41.9
39.9
12.2
15.7
AIME24
Average Accuracy: 8.33% ± 1.08%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
10.00%
3
30
2
6.67%
2
30
3
6.67%
2
30
4
13.33%
4
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_3k_eval_636d.opencode_reasoning2_python_k2_api64_annotated_think_thinking_extracted_debug10_v2opencodereasoning_10k_eval_636d
mlfoundations-dev/opencodereasoning_10k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
7.7
13.5
29.2
26.8
19.5
37.5
40.8
13.0
17.8
AIME24
Average Accuracy: 7.67% ± 1.34%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
6.67%
2
30
2
0.00%
0
30
3
13.33%
4
30
4
3.33%
1
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_10k_eval_636d.opencodereasoning_30k_eval_636d
mlfoundations-dev/opencodereasoning_30k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
16.0
44.0
51.8
28.6
31.5
43.3
46.6
17.5
19.5
AIME24
Average Accuracy: 16.00% ± 1.62%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
16.67%
5
30
2
20.00%
6
30
3
13.33%
4
30
4
13.33%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_30k_eval_636d.OpenCodeInstruct-filtered-sftopencodereasoning_1k_eval_2e29
mlfoundations-dev/opencodereasoning_1k_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
3.3
12.3
24.4
26.2
9.6
38.4
31.0
8.5
10.4
5.0
6.8
19.2
AIME24
Average Accuracy: 3.33% ± 0.82%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
6.67%
2
30
2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_1k_eval_2e29.OpenCodeReasoning-Nemotron-7B_eval_2e29
mlfoundations-dev/OpenCodeReasoning-Nemotron-7B_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
1.3
8.8
19.2
45.8
10.0
28.6
63.3
30.3
32.7
0.7
12.3
48.8
AIME24
Average Accuracy: 1.33% ± 0.70%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
0.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenCodeReasoning-Nemotron-7B_eval_2e29.cleaned_nvidia_OpenCodeReasoning元データ: https://huggingface.co/datasets/nvidia/OpenCodeReasoning
データ件数: 11,275
平均トークン数: 11251
最大トークン数: 19,802
合計トークン数: 126,859,041
ファイル形式: JSONL
ファイルサイズ: 707.4 MB
難易度スコアが15, カテゴリがcompetition、ライセンスがmitとcc-by-4.0をピックアップ
繰り返し除去
極端に少ない・多いなどを除去
詳しいコードはGithub
https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/opencodereasoning
opencodereasoning_32B_eval_636d
mlfoundations-dev/opencodereasoning_32B_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
44.7
79.5
82.8
34.4
72.3
56.2
77.4
44.2
45.3
AIME24
Average Accuracy: 44.67% ± 1.71%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
40.00%
12
30
2
53.33%
16
30
3
50.00%
15
30
4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_32B_eval_636d.opencodereasoning_100k_eval_2e29
mlfoundations-dev/opencodereasoning_100k_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
2.7
6.7
14.0
29.0
12.4
44.3
56.6
22.2
25.5
0.7
9.5
38.9
AIME24
Average Accuracy: 2.67% ± 0.79%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
3.33%
1
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_100k_eval_2e29.OpenCodeReasoning-Nemotron-7B_eval_5554
mlfoundations-dev/OpenCodeReasoning-Nemotron-7B_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
3.3
11.2
21.6
38.9
10.5
27.3
66.0
31.3
31.9
13.1
4.3
2.3
48.1
AIME24
Average Accuracy: 3.33% ± 0.67%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenCodeReasoning-Nemotron-7B_eval_5554.opencode_reasoning2_hard_codeforces2000_pr03_qwen35_fp8_thinking_annotated_10k_seed20260513
Qwen3.5 FP8 Annotations for 10K K2-Think OCR2 Coding Steps
This dataset contains Qwen3.5 FP8 step-level correctness annotations for K2-Think reasoning traces on a hard Codeforces subset of OpenCodeReasoning-2.
Summary
Source trace dataset: opencode_reasoning2_hard_codeforces2000_pr03_k2_thinking_extracted_pilot10
Source rows: 10 hard coding problem traces
Candidate step rule: claim with non-empty aligned_token_ids
Candidate steps: 15,267
Manifest-selected annotated… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/opencode_reasoning2_hard_codeforces2000_pr03_qwen35_fp8_thinking_annotated_10k_seed20260513.opencodereasoning_0.3k_eval_636d
mlfoundations-dev/opencodereasoning_0.3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
7.0
22.5
41.6
27.4
30.1
41.2
24.3
7.9
9.6
AIME24
Average Accuracy: 7.00% ± 1.20%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
10.00%
3
30
2
10.00%
3
30
3
10.00%
3
30
4
6.67%
2
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_0.3k_eval_636d.opencodereasoning_30k_eval_2e29
mlfoundations-dev/opencodereasoning_30k_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
14.7
40.5
54.0
27.6
30.7
44.3
47.4
18.2
20.1
8.3
10.8
34.0
AIME24
Average Accuracy: 14.67% ± 1.78%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
13.33%
4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_30k_eval_2e29.opencodereasoning_10k_eval_2e29
mlfoundations-dev/opencodereasoning_10k_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
9.0
15.0
30.0
27.8
19.7
36.0
41.4
14.4
17.3
4.7
7.8
28.6
AIME24
Average Accuracy: 9.00% ± 1.64%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
10.00%
3
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_10k_eval_2e29.opencodereasoning_300k_eval_2e29
mlfoundations-dev/opencodereasoning_300k_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
5.3
10.5
15.6
50.6
11.9
39.9
61.5
29.2
30.4
0.7
11.5
48.9
AIME24
Average Accuracy: 5.33% ± 0.97%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
10.00%
3
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_300k_eval_2e29.a1_code_opencoder_eval_636d
mlfoundations-dev/a1_code_opencoder_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
19.7
59.8
73.6
29.4
37.7
41.9
34.9
7.1
8.2
AIME24
Average Accuracy: 19.67% ± 1.91%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
23.33%
7
30
2
20.00%
6
30
3
20.00%
6
30
4
10.00%
3
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_opencoder_eval_636d.a1_code_opencodereasoning_eval_636d
mlfoundations-dev/a1_code_opencodereasoning_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
19.0
55.0
69.6
31.4
42.6
38.9
45.9
17.6
18.9
AIME24
Average Accuracy: 19.00% ± 1.16%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
20.00%
6
30
2
23.33%
7
30
3
13.33%
4
30
4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_opencodereasoning_eval_636d.b2_code_askllm_opencodereasoningb2_code_length_filtering_opencodereasoningopencodereasoning_100k_eval_636d
mlfoundations-dev/opencodereasoning_100k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
2.0
7.3
17.4
28.8
10.6
42.3
56.2
23.1
24.6
AIME24
Average Accuracy: 2.00% ± 0.52%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
3.33%
1
30
2
0.00%
0
30
3
3.33%
1
30
4
3.33%
1
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_100k_eval_636d.
