CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenCoder-LLM /opc-fineweb-code-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.tabular100M<n<1B57 likes5.3k downloads2y agoHugging Face02OpenCoder-LLM /RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode! Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available. RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.tabular100M<n<1B28 likes1.1k downloads2y agoHugging Face03amer224 /Opencode1tabularn<1K6 likes484 downloads8d agoHugging Face04JetBrains-Research /OpenCodeInstruct-Clean OpenCodeInstruct Clean High-quality Python code generation dataset with duplication markers and complexity metrics. Derived from nvidia/OpenCodeInstruct after applying strict quality gates. Quick Stats Metric Value Total rows 388,629 Columns 54 Python-parsable 100.0% Overview This dataset contains 388,629 high-quality Python code generation examples extracted from the nvidia/OpenCodeInstruct corpus. Each row has been… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/OpenCodeInstruct-Clean.tabular100K<n<1M0 likes466 downloads20d agoHugging Face05OpenCoder-LLM /opc-fineweb-math-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb opc-fineweb-math-corpus: the math-related page recalled from fineweb <-- you are here refineCode-code-corpus-meta: the… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-math-corpus.tabular1M<n<10M31 likes429 downloads2y agoHugging Face06zake7749 /Qwen3-Coder-Next-OpenCode-Preference Dataset Card — OpenCode Rejection Sampling (Preference) Overview This dataset contains 10,920 preference pairs for preference-based training (DPO, KTO, SimPO, ORPO, etc.) on competitive programming tasks. Each pair consists of: Chosen: a candidate solution that passes 100% of test cases Rejected: a candidate solution that fails, with a fine-grained rejection type label Pairs are produced via rejection sampling with Qwen3-Coder-Next: 8 candidate solutions are… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3-Coder-Next-OpenCode-Preference.tabulartext-generation10K<n<100K0 likes191 downloads6mo agoHugging Face07physicl-test /opencode-public-data-pack-docker_input1-20-renders-512-20260612t125042z OpenCode Public Data Pack docker_input1 20 renders 512 20260612T125042Z Public data pack created from docker_input1.json with 20 renders at 512x512. This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/opencode-public-data-pack-docker_input1-20-renders-512-20260612t125042z.imagen<1K0 likes60 downloads3mo agoHugging Face08ArthurSrz /open_codes Open Codes Open dataset of French legal code articles with embeddings. Chunked articles from French legal codes sourced from Legifrance via the PISTE API, with 1024-dimensional embeddings generated by Mistral AI. Each row is a text chunk enriched with full article metadata from the parent legal code article. Legal codes included The list of legal codes is dynamic and managed in the LEX_codes_piste table. Adding a new code there (with actif=true) automatically… See the full description on the dataset page: https://huggingface.co/datasets/ArthurSrz/open_codes.tabular10K<n<100K0 likes44 downloads3mo agoHugging Face09mlfoundations-dev /opencodereasoning_1k_eval_636d mlfoundations-dev/opencodereasoning_1k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 6.7 13.2 24.8 26.4 9.6 36.2 31.9 9.0 10.7 AIME24 Average Accuracy: 6.67% ± 1.76% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 3.33% 1 30 2 3.33% 1 30 3 13.33% 4 30 4 6.67% 2 30 5… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_1k_eval_636d.tabular1K<n<10K0 likes36 downloads1y agoHugging Face10mlfoundations-dev /opencodereasoning_3k_eval_636d mlfoundations-dev/opencodereasoning_3k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 8.3 26.2 37.8 31.6 21.2 41.9 39.9 12.2 15.7 AIME24 Average Accuracy: 8.33% ± 1.08% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 10.00% 3 30 2 6.67% 2 30 3 6.67% 2 30 4 13.33% 4 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_3k_eval_636d.tabular1K<n<10K0 likes26 downloads1y agoHugging Face11JingweiNi /opencode_reasoning2_python_k2_api64_annotated_think_thinking_extracted_debug10_v2tabularn<1K0 likes17 downloads5mo agoHugging Face12mlfoundations-dev /opencodereasoning_10k_eval_636d mlfoundations-dev/opencodereasoning_10k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 7.7 13.5 29.2 26.8 19.5 37.5 40.8 13.0 17.8 AIME24 Average Accuracy: 7.67% ± 1.34% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 6.67% 2 30 2 0.00% 0 30 3 13.33% 4 30 4 3.33% 1 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_10k_eval_636d.tabular1K<n<10K0 likes13 downloads1y agoHugging Face13mlfoundations-dev /opencodereasoning_30k_eval_636d mlfoundations-dev/opencodereasoning_30k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 16.0 44.0 51.8 28.6 31.5 43.3 46.6 17.5 19.5 AIME24 Average Accuracy: 16.00% ± 1.62% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 16.67% 5 30 2 20.00% 6 30 3 13.33% 4 30 4 13.33%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_30k_eval_636d.tabular1K<n<10K0 likes13 downloads1y agoHugging Face14adorkin /OpenCodeInstruct-filtered-sfttabular100K<n<1M0 likes13 downloads6mo agoHugging Face15mlfoundations-dev /opencodereasoning_1k_eval_2e29 mlfoundations-dev/opencodereasoning_1k_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 3.3 12.3 24.4 26.2 9.6 38.4 31.0 8.5 10.4 5.0 6.8 19.2 AIME24 Average Accuracy: 3.33% ± 0.82% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 6.67% 2 30 2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_1k_eval_2e29.tabular10K<n<100K0 likes12 downloads1y agoHugging Face16mlfoundations-dev /OpenCodeReasoning-Nemotron-7B_eval_2e29 mlfoundations-dev/OpenCodeReasoning-Nemotron-7B_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 1.3 8.8 19.2 45.8 10.0 28.6 63.3 30.3 32.7 0.7 12.3 48.8 AIME24 Average Accuracy: 1.33% ± 0.70% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 0.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenCodeReasoning-Nemotron-7B_eval_2e29.tabular10K<n<100K0 likes11 downloads1y agoHugging Face17LLMTeamAkiyama /cleaned_nvidia_OpenCodeReasoning元データ: https://huggingface.co/datasets/nvidia/OpenCodeReasoning データ件数: 11,275 平均トークン数: 11251 最大トークン数: 19,802 合計トークン数: 126,859,041 ファイル形式: JSONL ファイルサイズ: 707.4 MB 難易度スコアが15, カテゴリがcompetition、ライセンスがmitとcc-by-4.0をピックアップ 繰り返し除去 極端に少ない・多いなどを除去 詳しいコードはGithub https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/opencodereasoning tabularquestion-answering10K<n<100K0 likes11 downloads1y agoHugging Face18mlfoundations-dev /opencodereasoning_32B_eval_636d mlfoundations-dev/opencodereasoning_32B_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 44.7 79.5 82.8 34.4 72.3 56.2 77.4 44.2 45.3 AIME24 Average Accuracy: 44.67% ± 1.71% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 40.00% 12 30 2 53.33% 16 30 3 50.00% 15 30 4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_32B_eval_636d.tabular1K<n<10K0 likes10 downloads1y agoHugging Face19mlfoundations-dev /opencodereasoning_100k_eval_2e29 mlfoundations-dev/opencodereasoning_100k_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 2.7 6.7 14.0 29.0 12.4 44.3 56.6 22.2 25.5 0.7 9.5 38.9 AIME24 Average Accuracy: 2.67% ± 0.79% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 3.33% 1 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_100k_eval_2e29.tabular10K<n<100K0 likes10 downloads1y agoHugging Face20mlfoundations-dev /OpenCodeReasoning-Nemotron-7B_eval_5554 mlfoundations-dev/OpenCodeReasoning-Nemotron-7B_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces HLE HMMT AIME25 LiveCodeBenchv5 Accuracy 3.3 11.2 21.6 38.9 10.5 27.3 66.0 31.3 31.9 13.1 4.3 2.3 48.1 AIME24 Average Accuracy: 3.33% ± 0.67% Number of Runs: 10 Run Accuracy Questions Solved Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenCodeReasoning-Nemotron-7B_eval_5554.tabular10K<n<100K0 likes10 downloads1y agoHugging Face21JingweiNi /opencode_reasoning2_hard_codeforces2000_pr03_qwen35_fp8_thinking_annotated_10k_seed20260513 Qwen3.5 FP8 Annotations for 10K K2-Think OCR2 Coding Steps This dataset contains Qwen3.5 FP8 step-level correctness annotations for K2-Think reasoning traces on a hard Codeforces subset of OpenCodeReasoning-2. Summary Source trace dataset: opencode_reasoning2_hard_codeforces2000_pr03_k2_thinking_extracted_pilot10 Source rows: 10 hard coding problem traces Candidate step rule: claim with non-empty aligned_token_ids Candidate steps: 15,267 Manifest-selected annotated… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/opencode_reasoning2_hard_codeforces2000_pr03_qwen35_fp8_thinking_annotated_10k_seed20260513.tabulartext-generationn<1K0 likes10 downloads4mo agoHugging Face22mlfoundations-dev /opencodereasoning_0.3k_eval_636d mlfoundations-dev/opencodereasoning_0.3k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 7.0 22.5 41.6 27.4 30.1 41.2 24.3 7.9 9.6 AIME24 Average Accuracy: 7.00% ± 1.20% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 10.00% 3 30 2 10.00% 3 30 3 10.00% 3 30 4 6.67% 2 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_0.3k_eval_636d.tabular1K<n<10K0 likes9 downloads1y agoHugging Face23mlfoundations-dev /opencodereasoning_30k_eval_2e29 mlfoundations-dev/opencodereasoning_30k_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 14.7 40.5 54.0 27.6 30.7 44.3 47.4 18.2 20.1 8.3 10.8 34.0 AIME24 Average Accuracy: 14.67% ± 1.78% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 13.33% 4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_30k_eval_2e29.tabular10K<n<100K0 likes9 downloads1y agoHugging Face24mlfoundations-dev /opencodereasoning_10k_eval_2e29 mlfoundations-dev/opencodereasoning_10k_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 9.0 15.0 30.0 27.8 19.7 36.0 41.4 14.4 17.3 4.7 7.8 28.6 AIME24 Average Accuracy: 9.00% ± 1.64% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 10.00% 3 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_10k_eval_2e29.tabular10K<n<100K0 likes9 downloads1y agoHugging Face25mlfoundations-dev /opencodereasoning_300k_eval_2e29 mlfoundations-dev/opencodereasoning_300k_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 5.3 10.5 15.6 50.6 11.9 39.9 61.5 29.2 30.4 0.7 11.5 48.9 AIME24 Average Accuracy: 5.33% ± 0.97% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 10.00% 3 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_300k_eval_2e29.tabular10K<n<100K0 likes9 downloads1y agoHugging Face26mlfoundations-dev /a1_code_opencoder_eval_636d mlfoundations-dev/a1_code_opencoder_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 19.7 59.8 73.6 29.4 37.7 41.9 34.9 7.1 8.2 AIME24 Average Accuracy: 19.67% ± 1.91% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 23.33% 7 30 2 20.00% 6 30 3 20.00% 6 30 4 10.00% 3 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_opencoder_eval_636d.tabular1K<n<10K0 likes8 downloads1y agoHugging Face27mlfoundations-dev /a1_code_opencodereasoning_eval_636d mlfoundations-dev/a1_code_opencodereasoning_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 19.0 55.0 69.6 31.4 42.6 38.9 45.9 17.6 18.9 AIME24 Average Accuracy: 19.00% ± 1.16% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 20.00% 6 30 2 23.33% 7 30 3 13.33% 4 30 4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_opencodereasoning_eval_636d.tabular1K<n<10K0 likes8 downloads1y agoHugging Face28mlfoundations-dev /b2_code_askllm_opencodereasoningtabular10K<n<100K0 likes8 downloads1y agoHugging Face29mlfoundations-dev /b2_code_length_filtering_opencodereasoningtabular10K<n<100K0 likes8 downloads1y agoHugging Face30mlfoundations-dev /opencodereasoning_100k_eval_636d mlfoundations-dev/opencodereasoning_100k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 2.0 7.3 17.4 28.8 10.6 42.3 56.2 23.1 24.6 AIME24 Average Accuracy: 2.00% ± 0.52% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 3.33% 1 30 2 0.00% 0 30 3 3.33% 1 30 4 3.33% 1 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/opencodereasoning_100k_eval_636d.tabular1K<n<10K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.