CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SWE-bench /SWE-smith-cpptext1K<n<10K0 likes4.3k downloads7mo agoHugging Face02Romoamigo /SWE-Bench-MultilingualC_CPPFileteredtextn<1K0 likes1.8k downloads1y agoHugging Face03Romoamigo /SWE-Bench-MultilingualC_CPPFiletered_newtextn<1K0 likes1.6k downloads1y agoHugging Face04Reset23 /the-stack-v2-new-cpptabular1M<n<10M1 likes1.1k downloads1y agoHugging Face05Reset23 /the-stack-v2-cpptabular1M<n<10M1 likes806 downloads2y agoHugging Face06ThomasTheMaker /arc-stack-cpptabular1M<n<10M0 likes502 downloads11mo agoHugging Face07ningani /stack-v2-cpp-2019tabular10M<n<100M0 likes491 downloads2y agoHugging Face08shanxianzheng /SWE-smith-cpptext1K<n<10K0 likes391 downloads2mo agoHugging Face09hongliu9903 /stack_edu_cpptabular10M<n<100M0 likes274 downloads1y agoHugging Face10wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes255 downloads2y agoHugging Face11CPP-UT-BENCH /cpp_unit_tests_benchmark_datatext1K<n<10K5 likes220 downloads2y agoHugging Face12HPC-Forran2Cpp /HPC_Fortran_CPPThis dataset is associated with the following paper: Creating a Dataset for High-Performance Computing Code Translation using LLMs: A Bridge Between OpenMP Fortran and C++, Links https://arxiv.org/abs/2307.07686 https://github.com/bin123apple/OpenMP-Fortran-CPP-Translation textn<1K8 likes214 downloads2y agoHugging Face13izzako /IDD_Detection_CPPE5The IDD Object Detection dataset containing 40K images with CPPE-5 like (or YOLO) dataset annotation format.Refer to the original dataset: https://idd.insaan.iiit.ac.in imageobject-detection10K<n<100K1 likes186 downloads1y agoHugging Face14CPP-UT-BENCH /cpp_unit_tests_benchmark_data_with_splitstext1K<n<10K1 likes155 downloads2y agoHugging Face15Nutanix /CPP-UNITTEST-BENCH Dataset Card for Open Source Code and Unit Tests Dataset Details Dataset Description This dataset contains c++ code snippets and their corresponding ground truth unit tests collected from various open-source GitHub repositories. The primary purpose of this dataset is to aid in the development and evaluation of automated testing tools, code quality analysis, and LLM models for test generation. Curated by: Vaishnavi Bhargava Language(s): C++… See the full description on the dataset page: https://huggingface.co/datasets/Nutanix/CPP-UNITTEST-BENCH.text1K<n<10K3 likes152 downloads2y agoHugging Face16Reset23 /the-stack-v2-filtered-cpptabular100K<n<1M0 likes142 downloads1y agoHugging Face17AISE-TUDelft /Stackless_CPP_V2tabular100K<n<1M0 likes136 downloads11mo agoHugging Face18MultilingualUnigramLM /LangMap-TheStack-cpp-100M LangMap-TheStack-cpp-100M Code finetuning dataset for cpp streamed from bigcode/the-stack. Tokens collected: 100,000,000 (target: 100,000,000) Tokenizer: allenai/OLMo-3-1025-7B Schema: {"text": [...]} (sanitised source code) text10K<n<100K0 likes126 downloads5mo agoHugging Face19Wholesomeisland /cpp-mit-github-search-code-in-repostext100K<n<1M0 likes125 downloads10mo agoHugging Face20beranki /gpt-5-mini-rebench-v2-cpptabularn<1K0 likes116 downloads4mo agoHugging Face21nguyentruong-ins /codeforces_cpp_cleanedtext1M<n<10M0 likes110 downloads3y agoHugging Face22AetherPrior /cpp_cwe_GRPO cpp_cwe_GRPO VeRL/GRPO-ready C++ security coding dataset generated by the simple_gen pipeline. Each row is a harness-validated task with pytest security/functionality tests, oracle candidate_cpp, and authoring guidelines (high_level_guidelines, implementational). Files File Rows Description cpp_cwe_GRPO.parquet 571 Full dataset (shuffled) cpp_cwe_GRPO_train.parquet 514 90% train split cpp_cwe_GRPO_val.parquet 57 10% validation split… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/cpp_cwe_GRPO.texttext-generationn<1K0 likes91 downloads2mo agoHugging Face23open-athena /nemotron-cpp-qwen3.5-122b-32k-tracestext1K<n<10K0 likes78 downloads3mo agoHugging Face24laion /a1-nemotron-cpp-swe100-20260805-tracestextn<1K0 likes72 downloads2mo agoHugging Face25casey-martin /oa_cpp_annotate_gen Dataset Description This dataset, compiled by Brendan Dolan-Gavitt, contains ~100 thousand c++ functions and GPT-3.5 turbo-generated summaries of the code's purpose. An example of Brendan's original prompt and GPT-3.5's summary may be found below. int gg_set_focus_pos(gg_widget_t *widget, int x, int y) { return 1; } Q. What language is the above code written in? A. C/C++. Q. What is the purpose of the above code? A. This code defines a function called `gg_set_focus_pos` that… See the full description on the dataset page: https://huggingface.co/datasets/casey-martin/oa_cpp_annotate_gen.textquestion-answering100K<n<1M2 likes68 downloads3y agoHugging Face26SprayOpoivre /project_codeNet_translation_go_cpp 📂 Translation_go_into_cpp Bienvenue sur la base de données Translation_go_into_cpp. Ce dataset regroupe des traductions de code en trois langages de programmation : Go, Python et C++.Chaque ligne contient un script go, et une traduction de celui ci, soit en python, soit en C++. L'objectif principal de ce dataset est de fournir une base propre et nettoyée pour l'entraînement de modèles de type LLM (Large Language Models) dans des tâches de traduction Go ↔ C++. 📊… See the full description on the dataset page: https://huggingface.co/datasets/SprayOpoivre/project_codeNet_translation_go_cpp.text1M<n<10M0 likes62 downloads7mo agoHugging Face27nguyentruong-ins /nhlcoding_cleaned_cpp_datasettext1M<n<10M1 likes59 downloads3y agoHugging Face28laion /terminal_bench_2_a1_stack_cpp_20260810_001258text1K<n<10K0 likes58 downloads1mo agoHugging Face29laion /dev_set_v2_a1_stack_cpp_20260814_211707text1K<n<10K0 likes57 downloads1mo agoHugging Face30BoltzmannEntropy /Cpp-Math Cpp-Math: A Math-to-C++ Instruction-Tuning Dataset Dataset Overview The Cpp-Math dataset is designed for fine-tuning models to translate mathematical problems into C++ code. It focuses on evaluating the ability of language models to generate accurate and executable C++ code from mathematical expressions or problem statements. The dataset is particularly useful for benchmarking models on tasks that require both mathematical reasoning and programming skills. Data… See the full description on the dataset page: https://huggingface.co/datasets/BoltzmannEntropy/Cpp-Math.text1K<n<10K2 likes55 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.