CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /OpenCodeInstruct OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs Dataset Description We introduce OpenCodeInstruct, the largest open-access instruction tuning dataset, comprising 5 million diverse samples. OpenCodeInstruct is designed for supervised fine-tuning (SFT). Technical Report - Discover the methodology and technical details behind OpenCodeInstruct. Github Repo - Access the complete pipeline used to perform SFT. This dataset is ready for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenCodeInstruct.texttext-generation1M<n<10M109 likes40k downloads1y agoHugging Face02nvidia /OpenCodeReasoning OpenCodeReasoning: Advancing Data Distillation for Competitive Coding Data Overview OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT). Technical Report - Discover the methodology and technical details behind OpenCodeReasoning. Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenCodeReasoning.texttext-generation100K<n<1M558 likes14k downloads1y agoHugging Face03R0mAI /opencodeinstruct-curatedtext1M<n<10M0 likes10k downloads3mo agoHugging Face04EER6 /nvidia-OpenCodeInstruct-refined nvidia-OpenCodeInstruct-refined A strictly quality-filtered subset of nvidia/OpenCodeInstruct (5M examples). This is a strict subset of EER6/nvidia-OpenCodeInstruct-broad. Filtering criteria Both conditions must be satisfied: Criterion Threshold LLM judge min score = 5 (out of 5) Unit test pass rate (average_test_score) = 1.0 LLM judge min score is the minimum across all three dimensions in the llm_judgement field: requirement_conformance — does the… See the full description on the dataset page: https://huggingface.co/datasets/EER6/nvidia-OpenCodeInstruct-refined.texttext-generation100K<n<1M1 likes6.7k downloads6mo agoHugging Face05OpenCoder-LLM /opc-sft-stage2 OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 <-- you are here opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-sft-stage2.text100K<n<1M105 likes6.6k downloads2y agoHugging Face06nvidia /OpenCodeReasoning-2 OpenCodeReasoning-2: A Large-scale Dataset for Reasoning in Code Generation and Critique Dataset Description OpenCodeReasoning-2 is the largest reasoning-based synthetic dataset to date for coding, comprising 1.4M samples in Python and 1.1M samples in C++ across 34,799 unique competitive programming questions. OpenCodeReasoning-2 is designed for supervised fine-tuning (SFT) tasks of code completion and code critique. Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenCodeReasoning-2.texttext-generation1M<n<10M61 likes5.7k downloads1y agoHugging Face07OpenCoder-LLM /opc-fineweb-code-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.tabular100M<n<1B57 likes5.3k downloads2y agoHugging Face08OpenCoder-LLM /opc-sft-stage1 OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 <-- you are here opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-sft-stage1.text1M<n<10M76 likes2.2k downloads2y agoHugging Face09OpenCoder-LLM /RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode! Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available. RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.tabular100M<n<1B28 likes1.1k downloads2y agoHugging Face10sealad886 /OpenCodeReasoning_messages This is a transformation of the nvidia/OpenCodeReasoning dataset into a format that is more easily digestible by trainers. OpenCodeReasoning: Advancing Data Distillation for Competitive Coding Data Overview OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT). Technical Report -… See the full description on the dataset page: https://huggingface.co/datasets/sealad886/OpenCodeReasoning_messages.texttext-generation100K<n<1M0 likes911 downloads1y agoHugging Face11cublya /OpenCodeReasoning-2 OpenCodeReasoning-2: A Large-scale Dataset for Reasoning in Code Generation and Critique Dataset Description OpenCodeReasoning-2 is the largest reasoning-based synthetic dataset to date for coding, comprising 1.4M samples in Python and 1.1M samples in C++ across 34,799 unique competitive programming questions. OpenCodeReasoning-2 is designed for supervised fine-tuning (SFT) tasks of code completion and code critique. Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/cublya/OpenCodeReasoning-2.texttext-generation1M<n<10M0 likes558 downloads8mo agoHugging Face12JetBrains-Research /OpenCodeInstruct-Clean OpenCodeInstruct Clean High-quality Python code generation dataset with duplication markers and complexity metrics. Derived from nvidia/OpenCodeInstruct after applying strict quality gates. Quick Stats Metric Value Total rows 388,629 Columns 54 Python-parsable 100.0% Overview This dataset contains 388,629 high-quality Python code generation examples extracted from the nvidia/OpenCodeInstruct corpus. Each row has been… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/OpenCodeInstruct-Clean.tabular100K<n<1M0 likes466 downloads21d agoHugging Face13nvidia /OpenCodeGeneticInstruct OpenCodeGeneticInstruct: A large-scale dataset of coding instructions for improving the code generation capabilities of LLMs Data Overview OpenCodeGeneticInstruct comprises more than 15M coding instructions in python which is generated synthetically with the Genetic-Instruct [1] approach. This dataset can be used for supervised fine-tuning (SFT) of LLMs to improve their code genearation capability. Each sample includes a coding question/instruction and its corrsponding… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenCodeGeneticInstruct.text10M<n<100M20 likes448 downloads1y agoHugging Face14Jeremydh911 /OpenCodeInstruct OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs Dataset Description We introduce OpenCodeInstruct, the largest open-access instruction tuning dataset, comprising 5 million diverse samples. OpenCodeInstruct is designed for supervised fine-tuning (SFT). Technical Report - Discover the methodology and technical details behind OpenCodeInstruct. Github Repo - Access the complete pipeline used to perform SFT. This dataset is ready for… See the full description on the dataset page: https://huggingface.co/datasets/Jeremydh911/OpenCodeInstruct.texttext-generation1M<n<10M0 likes445 downloads5mo agoHugging Face15OpenCoder-LLM /opc-fineweb-math-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb opc-fineweb-math-corpus: the math-related page recalled from fineweb <-- you are here refineCode-code-corpus-meta: the… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-math-corpus.tabular1M<n<10M31 likes429 downloads2y agoHugging Face16Compumacy /NMT-opencode OpenCodeReasoning: Advancing Data Distillation for Competitive Coding Data Overview OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT). Technical Report - Discover the methodology and technical details behind OpenCodeReasoning. Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/Compumacy/NMT-opencode.texttext-generation100K<n<1M0 likes400 downloads1y agoHugging Face17open-athena /llm-verifier-freelancer-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/llm-verifier-freelancer-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes317 downloads3mo agoHugging Face18open-athena /nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes302 downloads2mo agoHugging Face19psyche /nvidia-OpenCodeReasoningtext100K<n<1M0 likes285 downloads8mo agoHugging Face20Parveshiiii /opencode_reasoning_filtered 🧠 OpenCode Reasoning (Filtered) Author: Parvesh Rawal — XenArcAILicense: Inherits from NVIDIA OpenCodeReasoningVersion: Filtered & Structured VariantTotal Examples: 567,850Total Size: 9GB (compressed) 🔍 Overview This dataset is a curated and cleaned version of split_0 from nvidia/OpenCodeReasoning, optimized for code-level reasoning tasks and instruction tuning. It’s designed to enhance logic understanding and multistep problem solving for LLMs. 📁 Features… See the full description on the dataset page: https://huggingface.co/datasets/Parveshiiii/opencode_reasoning_filtered.texttext-generation100K<n<1M4 likes257 downloads1y agoHugging Face21open-athena /stackexchange-tezos-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/stackexchange-tezos-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces.text1K<n<10K0 likes254 downloads29d agoHugging Face22open-athena /stackexchange-superuser-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/stackexchange-superuser-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces.text1K<n<10K0 likes243 downloads1mo agoHugging Face23MaziyarPanahi /OpenCodeReasoning_ShareGPT Added cnversations column in ShareGPT format Original README from nvidia/OpenCodeReasoning OpenCodeReasoning: Advancing Data Distillation for Competitive Coding Data Overview OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT). Technical Report - Discover the methodology and… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/OpenCodeReasoning_ShareGPT.text100K<n<1M9 likes238 downloads1y agoHugging Face24open-athena /stackexchange-overflow-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/stackexchange-overflow-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces.text1K<n<10K0 likes235 downloads27d agoHugging Face25open-athena /nemotron-gym-identity-following-v2-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-identity-following-v2-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes231 downloads2mo agoHugging Face26open-athena /exp_rpt_pymethods2test-large-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_pymethods2test-large-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes217 downloads2mo agoHugging Face27zaydzuhri /OpenCodeInstruct-MBPP-Texttext1M<n<10M0 likes199 downloads9mo agoHugging Face28open-athena /nemotron-gym-competitive-coding-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-competitive-coding-qwen3.5-122b-131k-opencode-traces.text10K<n<100K0 likes196 downloads2mo agoHugging Face29zake7749 /Qwen3-Coder-Next-OpenCode-Preference Dataset Card — OpenCode Rejection Sampling (Preference) Overview This dataset contains 10,920 preference pairs for preference-based training (DPO, KTO, SimPO, ORPO, etc.) on competitive programming tasks. Each pair consists of: Chosen: a candidate solution that passes 100% of test cases Rejected: a candidate solution that fails, with a fine-grained rejection type label Pairs are produced via rejection sampling with Qwen3-Coder-Next: 8 candidate solutions are… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3-Coder-Next-OpenCode-Preference.tabulartext-generation10K<n<100K0 likes191 downloads6mo agoHugging Face30open-athena /selfinstruct-naive-sandboxes-2-verified-qwen3.5-122b-131k-opencode-tracestext1K<n<10K0 likes190 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.