CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01livecodebench /code_generation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs. Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution. This is the code generation scenario of LiveCodeBench. It is also… See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/code_generation.textn<1K31 likes5.6k downloads2y agoHugging Face02lighteval /code_generation_lite LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • 📄 Paper Change Log Since LiveCodeBench is a continuously updated benchmark, we provide different versions of the dataset. Particularly, we provide the following versions of the dataset: release_v1: The initial release of the dataset with problems released between May 2023 and Mar 2024 containing 400… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/code_generation_lite.text10K<n<100K5 likes3.8k downloads1y agoHugging Face03codesbyusman /LLM-CodeGen LLMs-Generated Code This dataset the raw code generated by 10 different Large Language Models (LLMs) in response to the prompts from our research project. The code is organized to facilitate the assessment and comparison of each model's ability to generate secure C/C++ code. The generated code is divided into two main categories: Simple Assistant: Code generated by the LLM with no specific security-focused instructions. Secure Assistant: Code generated by the LLM using prompts that… See the full description on the dataset page: https://huggingface.co/datasets/codesbyusman/LLM-CodeGen.text1K<n<10K1 likes1.2k downloads1y agoHugging Face04sayakpaul /hf-codegen-v2 Dataset Card for "hf-codegen-v2" Dataset generated with the code from: https://github.com/sayakpaul/hf-codegen. tabular100K<n<1M25 likes983 downloads3y agoHugging Face05OllieStanley /humaneval-mbpp-codegen-qa Dataset Card for "humaneval-mbpp-codegen-qa" This dataset contains prompt-reply (question-answer) pairs where the prompt is to create a Python function which satisfies the functionality described in a specified docstring. The responses are then the generated functions. textn<1K4 likes882 downloads4y agoHugging Face06pengyunie /codesearchnet-codegen Dataset Card for CodeSearchNet for CodeGen This is a processed version of the CodeSearchNet dataset. Namely, I separated the doc (documentation/docstring), sign (function signature), and output (function body) into separate fields; doc and sign are concatenated (according to the correct order of the programming language) into the problem field, making it suitable for the code generation task. Dataset Details Dataset Description Curated by: [More… See the full description on the dataset page: https://huggingface.co/datasets/pengyunie/codesearchnet-codegen.text1M<n<10M2 likes218 downloads2y agoHugging Face07FudanSELab /CodeGen4Libs_RetrievalCodeLib Dataset Card for FudanSELab CodeGen4Libs Code Retrieval Library Dataset Summary This dataset is the code retrieval library used in the ASE2023 paper titled "CodeGen4Libs: A Two-stage Approach for Library-oriented Code Generation". Additional Information Citation Information @inproceedings{ase2023codegen4libs, author = {Mingwei Liu and Tianyong Yang and Yiling Lou and Xueying Du and Ying Wang and and Xin Peng}, title =… See the full description on the dataset page: https://huggingface.co/datasets/FudanSELab/CodeGen4Libs_RetrievalCodeLib.text1M<n<10M1 likes214 downloads3y agoHugging Face08Naholav /CodeGen-Diverse-5K CodeGen-Diverse-5K: Broad Coverage for Competitive Programming Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset) Dataset Description CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions. Key Statistics Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.tabulartext-generation1K<n<10K0 likes123 downloads10mo agoHugging Face09codegenning /usacobench_formattedtextn<1K0 likes116 downloads2y agoHugging Face10Naholav /CodeGen-Deep-5K CodeGen-Deep-5K: Deep Reasoning for Competitive Programming Part of the CodeGen suite | CodeGen-Diverse-5K (sister dataset) Dataset Description CodeGen-Deep-5K is a deep reasoning dataset designed for training code generation models with enhanced problem-solving capabilities. Unlike traditional datasets, this generates multiple distinct solutions for each problem, providing varied reasoning traces and approaches. Key Statistics Total samples: 5,000 Unique… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Deep-5K.tabulartext-generation1K<n<10K0 likes98 downloads10mo agoHugging Face11AmareshHebbar /leetcode-codegen-python LeetCode Code-Gen Dataset — Python 2522 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct Python solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Every row was extracted, then executed in a sandboxed subprocess against the problem's own stated examples. Only rows that passed all examples are included -- this… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-python.texttext-generation1K<n<10K1 likes93 downloads3mo agoHugging Face12Naholav /llama3.2-java-codegen-90sft-10meta-claude-v1 LLaMA 3.2 Java Code Generation Dataset (90% SFT, 10% Meta Annotated with Claude) This dataset contains 100,000 examples for Java method generation based on natural language instructions. It is built from the CodeXGLUE text-to-code dataset and designed to support both pure supervised fine-tuning (SFT) and reflection-based meta-learning approaches using Claude 4 Sonnet as the critique model. 🚀 Trained Models Two models have been trained on this dataset: SFT Model:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/llama3.2-java-codegen-90sft-10meta-claude-v1.texttext-generation100K<n<1M1 likes82 downloads1y agoHugging Face13Samsoup /Code-Generation-Quality-Estimation Code Generation Quality Estimation This repository contains model-ready task context, generated code, and complete-case execution-resource targets for five public LLM code-generation cohorts. It provides deterministic 70/10/20 group-aware split versions using seeds 42, 1234, and 2026. Configurations There are 15 configurations: one for each dataset and split seed. Each configuration has train, validation, and test splits. Dataset Complete rows Groups Models… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/Code-Generation-Quality-Estimation.tabulartabular-regression100K<n<1M0 likes77 downloads2mo agoHugging Face14Groq /LiveCodeBench-CodeGenerationtextquestion-answeringn<1K0 likes72 downloads1y agoHugging Face15iapp /code_generation_lite-th LiveCodeBench code_generation_lite, Thai 111 competitive-programming problems from LeetCode and AtCoder, with the problem statement translated to Thai. Everything else — test cases, starter code, metadata — is the upstream value unchanged. Known defects The line breaks are gone from the problem statements. 110 of the 111 rows have no line break at all in question_content; the one remaining row has two. These are competitive-programming statements whose input and… See the full description on the dataset page: https://huggingface.co/datasets/iapp/code_generation_lite-th.textquestion-answeringn<1K0 likes69 downloads1mo agoHugging Face16AmareshHebbar /leetcode-codegen-javascript LeetCode Code-Gen Dataset — JavaScript 631 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct JavaScript solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-javascript.texttext-generationn<1K1 likes60 downloads3mo agoHugging Face17stindardlogic /code-generation-sft-100k Code Generation SFT (100K) 100,000 ShareGPT conversations covering code generation across 8 programming languages, 21 categories, and 22 distinct programming tasks. Each example includes a detailed natural language request and a complete, working implementation with explanations of key design decisions. Motivation Coding assistants are the highest-adoption LLM application category, but most open training datasets focus on isolated functions without context. This… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-generation-sft-100k.texttext-generation100K<n<1M0 likes56 downloads2mo agoHugging Face18AmareshHebbar /leetcode-codegen-java LeetCode Code-Gen Dataset — Java 4068 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct Java solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-java.texttext-generation1K<n<10K0 likes55 downloads3mo agoHugging Face19AmareshHebbar /leetcode-codegen-cpp LeetCode Code-Gen Dataset — C++ 4025 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct C++ solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-cpp.texttext-generation1K<n<10K1 likes53 downloads3mo agoHugging Face20rootly-ai-labs /terraform-aws-ec2-instance-profile-codegentextn<1K0 likes51 downloads8mo agoHugging Face21jenyag /repo-codegen-py-py-context-path-distance Dataset Card for "repo-codegen-py-py-context-path-distance" More Information needed textn<1K1 likes49 downloads3y agoHugging Face22openSUSE /cve-backport-codegen-dataset CVE Backport Code Generation Dataset Per-hunk code generation dataset for CVE security patch backporting, derived from openSUSE Build Service maintenance patches. Task Given a region of vulnerable source code and a description of the upstream CVE fix, the model outputs the fixed version of the code. A programmatic diff then produces the final patch. This plays to LLM strengths in code completion and avoids format-sensitivity issues with direct diff generation.… See the full description on the dataset page: https://huggingface.co/datasets/openSUSE/cve-backport-codegen-dataset.texttext-generation10K<n<100K1 likes49 downloads6mo agoHugging Face23AI4Manufacturing /D14-codegen-annotatedgated D14 CodeGen Annotated — CAD code generation with chain-of-thought The code_gen layer of D14 (BenchCAD) with teacher-written reasoning: 17,900 records, each 4-view orthographic render → reasoning → complete **CadQuery** program. The reasoning ends with the complete program in a ```python block, and the program is AST-equivalent to BenchCAD's execution-verified reference — asserted per record at assembly (records whose regeneration drifted carry the reference spliced verbatim;… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/D14-codegen-annotated.image10K<n<100K1 likes49 downloads2mo agoHugging Face24generaleoley /manim-codegentext1K<n<10K11 likes47 downloads3y agoHugging Face25amztheory /codegen-instruct-CPlus-1ktextn<1K0 likes47 downloads2y agoHugging Face26codegenning /B_taco_clean_v1text1K<n<10K0 likes44 downloads2y agoHugging Face27jenyag /repo-codegen-py-non-py-context-path-distance Dataset Card for "repo-codegen-py-non-py-context-path-distance" More Information needed textn<1K0 likes39 downloads3y agoHugging Face28JetBrains-Research /lca-codegen-huge LCA Project Level Code Completion How to load the dataset from datasets import load_dataset ds = load_dataset('JetBrains-Research/lca-codegen-huge', split='test') Data Point Structure repo – repository name in format {GitHub_user_name}__{repository_name} commit_hash – commit hash completion_file – dictionary with the completion file content in the following format: filename – filepath to the completion file content – content of the completion file… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-codegen-huge.textn<1K0 likes37 downloads2y agoHugging Face29IyedLahiani /html-css-codegen-datasettext1K<n<10K1 likes37 downloads1y agoHugging Face30severo /CodeGen4Libs Dataset Card for FudanSELab CodeGen4Libs Dataset Dataset Summary This dataset is used in the ASE2023 paper titled "CodeGen4Libs: A Two-stage Approach for Library-oriented Code Generation". Languages [More Information Needed] Dataset Structure from datasets import load_dataset dataset = load_dataset("FudanSELab/CodeGen4Libs") DatasetDict({ train: Dataset({ features: ['id', 'method', 'clean_method', 'doc', 'comment', 'method_name', 'extra'… See the full description on the dataset page: https://huggingface.co/datasets/severo/CodeGen4Libs.tabular100K<n<1M0 likes36 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.