CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SWE-bench /SWE-smith-cpptext1K<n<10K0 likes4.3k downloads7mo agoHugging Face02rishitdagli /cppe-5 Dataset Card for CPPE - 5 Dataset Summary CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that focus on broad level categories. Some features of this dataset are: high quality images and annotations (~4.6 bounding boxes per image) real-life images unlike any current such dataset majority… See the full description on the dataset page: https://huggingface.co/datasets/rishitdagli/cppe-5.imageobject-detection1K<n<10K23 likes3.9k downloads3y agoHugging Face03Romoamigo /SWE-Bench-MultilingualC_CPPFileteredtextn<1K0 likes1.8k downloads1y agoHugging Face04Romoamigo /SWE-Bench-MultilingualC_CPPFiletered_newtextn<1K0 likes1.6k downloads1y agoHugging Face05Reset23 /the-stack-v2-new-cpptabular1M<n<10M1 likes1.1k downloads1y agoHugging Face06Reset23 /the-stack-v2-cpptabular1M<n<10M1 likes806 downloads2y agoHugging Face07ningani /stack-v2-cpp-2019tabular10M<n<100M0 likes523 downloads2y agoHugging Face08ThomasTheMaker /arc-stack-cpptabular1M<n<10M0 likes460 downloads11mo agoHugging Face09shanxianzheng /SWE-smith-cpptext1K<n<10K0 likes392 downloads2mo agoHugging Face10hongliu9903 /stack_edu_cpptabular10M<n<100M0 likes275 downloads1y agoHugging Face11wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes269 downloads2y agoHugging Face12CPP-UT-BENCH /cpp_unit_tests_benchmark_datatext1K<n<10K5 likes219 downloads2y agoHugging Face13izzako /IDD_Detection_CPPE5The IDD Object Detection dataset containing 40K images with CPPE-5 like (or YOLO) dataset annotation format.Refer to the original dataset: https://idd.insaan.iiit.ac.in imageobject-detection10K<n<100K1 likes214 downloads1y agoHugging Face14qubvel-hf /cppe-5-sampleimagen<1K0 likes202 downloads2y agoHugging Face15CPP-UT-BENCH /cpp_unit_tests_benchmark_data_with_splitstext1K<n<10K1 likes155 downloads2y agoHugging Face16HPC-Forran2Cpp /HPC_Fortran_CPPThis dataset is associated with the following paper: Creating a Dataset for High-Performance Computing Code Translation using LLMs: A Bridge Between OpenMP Fortran and C++, Links https://arxiv.org/abs/2307.07686 https://github.com/bin123apple/OpenMP-Fortran-CPP-Translation textn<1K8 likes154 downloads2y agoHugging Face17Nutanix /CPP-UNITTEST-BENCH Dataset Card for Open Source Code and Unit Tests Dataset Details Dataset Description This dataset contains c++ code snippets and their corresponding ground truth unit tests collected from various open-source GitHub repositories. The primary purpose of this dataset is to aid in the development and evaluation of automated testing tools, code quality analysis, and LLM models for test generation. Curated by: Vaishnavi Bhargava Language(s): C++… See the full description on the dataset page: https://huggingface.co/datasets/Nutanix/CPP-UNITTEST-BENCH.text1K<n<10K3 likes153 downloads2y agoHugging Face18Reset23 /the-stack-v2-filtered-cpptabular100K<n<1M0 likes142 downloads1y agoHugging Face19AISE-TUDelft /Stackless_CPP_V2tabular100K<n<1M0 likes141 downloads11mo agoHugging Face20MultilingualUnigramLM /LangMap-TheStack-cpp-100M LangMap-TheStack-cpp-100M Code finetuning dataset for cpp streamed from bigcode/the-stack. Tokens collected: 100,000,000 (target: 100,000,000) Tokenizer: allenai/OLMo-3-1025-7B Schema: {"text": [...]} (sanitised source code) text10K<n<100K0 likes137 downloads5mo agoHugging Face21Wholesomeisland /cpp-mit-github-search-code-in-repostext100K<n<1M0 likes131 downloads10mo agoHugging Face22beranki /gpt-5-mini-rebench-v2-cpptabularn<1K0 likes118 downloads4mo agoHugging Face23nguyentruong-ins /codeforces_cpp_cleanedtext1M<n<10M0 likes107 downloads3y agoHugging Face24AetherPrior /cpp_cwe_GRPO cpp_cwe_GRPO VeRL/GRPO-ready C++ security coding dataset generated by the simple_gen pipeline. Each row is a harness-validated task with pytest security/functionality tests, oracle candidate_cpp, and authoring guidelines (high_level_guidelines, implementational). Files File Rows Description cpp_cwe_GRPO.parquet 571 Full dataset (shuffled) cpp_cwe_GRPO_train.parquet 514 90% train split cpp_cwe_GRPO_val.parquet 57 10% validation split… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/cpp_cwe_GRPO.texttext-generationn<1K0 likes93 downloads2mo agoHugging Face25SprayOpoivre /project_codeNet_translation_go_cpp 📂 Translation_go_into_cpp Bienvenue sur la base de données Translation_go_into_cpp. Ce dataset regroupe des traductions de code en trois langages de programmation : Go, Python et C++.Chaque ligne contient un script go, et une traduction de celui ci, soit en python, soit en C++. L'objectif principal de ce dataset est de fournir une base propre et nettoyée pour l'entraînement de modèles de type LLM (Large Language Models) dans des tâches de traduction Go ↔ C++. 📊… See the full description on the dataset page: https://huggingface.co/datasets/SprayOpoivre/project_codeNet_translation_go_cpp.text1M<n<10M0 likes82 downloads7mo agoHugging Face26open-athena /nemotron-cpp-qwen3.5-122b-32k-tracestext1K<n<10K0 likes80 downloads3mo agoHugging Face27laion /a1-nemotron-cpp-swe100-20260805-tracestextn<1K0 likes72 downloads2mo agoHugging Face28casey-martin /oa_cpp_annotate_gen Dataset Description This dataset, compiled by Brendan Dolan-Gavitt, contains ~100 thousand c++ functions and GPT-3.5 turbo-generated summaries of the code's purpose. An example of Brendan's original prompt and GPT-3.5's summary may be found below. int gg_set_focus_pos(gg_widget_t *widget, int x, int y) { return 1; } Q. What language is the above code written in? A. C/C++. Q. What is the purpose of the above code? A. This code defines a function called `gg_set_focus_pos` that… See the full description on the dataset page: https://huggingface.co/datasets/casey-martin/oa_cpp_annotate_gen.textquestion-answering100K<n<1M2 likes66 downloads3y agoHugging Face29nguyentruong-ins /nhlcoding_cleaned_cpp_datasettext1M<n<10M1 likes59 downloads3y agoHugging Face30AmareshHebbar /leetcode-codegen-cpp LeetCode Code-Gen Dataset — C++ 4025 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct C++ solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-cpp.texttext-generation1K<n<10K1 likes57 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.