datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-smith-cppqwen-cpp-agent-0-protocolExperiment in agentic autonomy protocols.
~ everything in this repo was created by Qwen 3.8 27B (Q4) running autonomously inside Deepseek Harness, on a single RTX 3090 GPU, for 3 weeks.
The only human artifacts are:
agents/*
human/*
AGENTS.md
human_eval_cppSWE-Bench-MultilingualC_CPPFileteredSWE-Bench-MultilingualC_CPPFiletered_newthe-stack-v2-new-cppCpp-Code-LargeCpp-Code-Large
Cpp-Code-Large is a large-scale corpus of C++ source code comprising more than 5 million lines of C++ code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the C++ ecosystem.
By providing a high-volume, language-specific corpus, Cpp-Code-Large enables systematic experimentation in C++-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Cpp-Code-Large.the-stack-v2-cppllama.cppversion https://git-lfs.github.com/spec/v1
oid sha256:cfc44b7ba25614df70e6b65e3341cae0310163bd32fd31a6b928a542df433faf
size 30786
LiveCodeBench-CPP
LiveCodeBench-CPP: An Extension of LiveCodeBench for Contamination Free Evaluation in C++
Overview
LiveCodeBench-CPP includes 454 problems from the release_v6 of LiveCodeBench, covering the period from October 2024 to May 2025. These problems are sourced from AtCoder (287 problems) and LeetCode (167 problems).
AtCoder Problems: These require generated solutions to read inputs from standard input (stdin) and write outputs to standard output (stdout). For unit testing, the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LiveCodeBench-CPP.arc-stack-cppstack-v2-cpp-2019MD-trajectories-CPPF-tubulin-heterodimer-and-monomers
MD-trajectories-CPPF-tubulin-heterodimer-and-monomers
Copy this file into the Hugging Face dataset “README” (Dataset card).Source of truth in Git: https://github.com/jasperyeoh/integrative-ai-assisted-modeling-of-cppf-tubulin-interactions — see docs/DIMER_TRAJECTORY_NAMING.md.
What this dataset contains
All-atom GROMACS production trajectories (.xtc) for CPPF with human tubulin:
5IJ0 / soluble curved dimer (main text): three heterodimer replicates extended to… See the full description on the dataset page: https://huggingface.co/datasets/jasperyeoh2/MD-trajectories-CPPF-tubulin-heterodimer-and-monomers.SWE-smith-cppstack_edu_cppmagenta-realtime-mlx-cpp
Magenta RealTime — C++ MLX runtime bundle
This dataset is a re-packaging of
Google's Magenta RealTime weights
for the C++ MLX runtime in
rhymeswithlion/magenta-realtime-mlx-cpp.
It contains exactly what mlx-stream needs at startup; nothing more, nothing
less. The upstream .pt / .npy checkpoints are intentionally not
mirrored here — they're only useful for the (Python) re-export tooling on the
project's main distribution.
Contents
.
├──… See the full description on the dataset page: https://huggingface.co/datasets/rhymeswithlion/magenta-realtime-mlx-cpp.code_contest_instruct_cppcpp_unit_tests_benchmark_dataHPC_Fortran_CPPThis dataset is associated with the following paper:
Creating a Dataset for High-Performance Computing Code Translation using LLMs: A Bridge Between OpenMP Fortran and C++,
Links
https://arxiv.org/abs/2307.07686
https://github.com/bin123apple/OpenMP-Fortran-CPP-Translation
IDD_Detection_CPPE5The IDD Object Detection dataset containing 40K images with CPPE-5 like (or YOLO) dataset annotation format.Refer to the original dataset: https://idd.insaan.iiit.ac.in
synthetic-cpp
Dataset Card for Synthetic C++ Dataset
Dataset Description
Dataset Card for Synthetic C++ Dataset
Dataset Description
Homepage: [---
Dataset Card for Synthetic C++ Dataset
Dataset Description
Homepage: [https://huggingface.co/datasets/ReySajju742/synthetic-cpp/]
Point of Contact: [ReySajju742]
Dataset Summary
This dataset contains 10,000 rows of synthetically generated data focusing on the topic of "C++… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/synthetic-cpp.cpp_unit_tests_benchmark_data_with_splitsCPP-UNITTEST-BENCH
Dataset Card for Open Source Code and Unit Tests
Dataset Details
Dataset Description
This dataset contains c++ code snippets and their corresponding ground truth unit tests collected from various open-source GitHub repositories. The primary purpose of this dataset is to aid in the development and evaluation of automated testing tools, code quality analysis, and LLM models for test generation.
Curated by: Vaishnavi Bhargava
Language(s): C++… See the full description on the dataset page: https://huggingface.co/datasets/Nutanix/CPP-UNITTEST-BENCH.the-stack-v2-filtered-cppcpp-code-code_search_net-style
C++ Dataset
documentation source: https://huggingface.co/docs/datasets/main/en/repository_structure
Supported Tasks and Leaderboards
language-modeling: The dataset can be used to train a model for modelling programming languages, which consists in building language models for programming languages.
Language
C++ programming language
Dataset Structure
Data Instances
A data point consists of a function code along with its documentation.… See the full description on the dataset page: https://huggingface.co/datasets/malteklaes/cpp-code-code_search_net-style.Stackless_CPP_V2TopoBox-3D
TopoBox-3D
Paper (arXiv:2609.05860) | Code (GitHub)
TopoBox-3D is the dataset accompanying Beyond Arbitrary Geometry: Topology
Generalization in Neural PDE Operators. It is a controlled three-dimensional
benchmark for separating fixed-topology geometry shift from generalization to
unseen homological support.
The benchmark contains 5,280 connected box-minus-void geometries and 63,360
fixed-time Hodge-heat instances. Through-tunnels and enclosed cavities control
the first and… See the full description on the dataset page: https://huggingface.co/datasets/cppyyy/TopoBox-3D.cpp-mit-github-search-code-in-reposLangMap-TheStack-cpp-100M
LangMap-TheStack-cpp-100M
Code finetuning dataset for cpp streamed from bigcode/the-stack.
Tokens collected: 100,000,000 (target: 100,000,000)
Tokenizer: allenai/OLMo-3-1025-7B
Schema: {"text": [...]} (sanitised source code)
t_cpp
