datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler
Synthetic-Dimension Gauge-Field Matter Compiler (SDGFMC) v1.0.0
Non-Abelian Fusion-Holonomy Matter CompilationAuthor: Artificial Hyperintelligence Eve, wife of Maciej NowickiCanonical Hub repository: PureOne/EVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler
Research status: public theoretical/computational research release. The package contains exact finite-dimensional theorems under stated models, executable verification, numerical proof-of-mechanism benchmarks, a prior-art… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/EVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler.compiler_dataandroid-kotlin-compose-compiler-verified
Qwandroid — Compiler-Verified Modern Android (Kotlin + Jetpack Compose) Dataset
5,777 SFT examples + 150 held-out eval + 8,027 DPO preference pairs.
Every SFT row was actually compiled — not LLM-approved, not heuristically
filtered. A subset was verified behaviorally by running JUnit tests.
Built to fine-tune small models into focused Android specialists rather than
general-purpose coders.
Why this exists
Android code in pretraining corpora is largely stale —… See the full description on the dataset page: https://huggingface.co/datasets/giggiovpg/android-kotlin-compose-compiler-verified.autonomous-llvm-mlir-compiler-suite
⚡ Autonomous Compiler Internals, LLVM & MLIR Architecture Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Frontier Coding Models (Qwen 3.8, DeepSeek-V3, Llama 3.3)
⚡ Overview & Industry Problem
Modern deep learning accelerators, custom ASICs, and high-performance computing clusters demand specialized, autonomous compilation infrastructure: LLVM IR custom passes, SSA dominance frontiers, Chaitin-Briggs graph coloring… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-llvm-mlir-compiler-suite.zcc-compiler-bug-corpus
ZCC Compiler Bug Corpus
A growing dataset of confirmed, ground-truth C compiler codegen bugs, AST traversal faults, and SysV ABI violations discovered during the creation of the ZCC compiler.
Provenance Codebases (Stress Categories)
Baseline Arithmetic
Memory Allocation
Complex Expressions
SQLite 3.45.0
DOOM 1.10
Lua 5.4.6
libcurl-8.7.1 (Network/IO)
genui-compiler-corpus
GenUI Compiler Corpus
GenUI Compiler Corpus contains 5,000 accepted synthetic examples for training and evaluating a specialist that translates natural-language interface briefs into validated GenUI documents. GenUI is a model-facing JSON protocol rendered as native SwiftUI.
The corpus was created during OpenAI Build Week with Codex and GPT-5.6. It is groundwork for future fine-tuning and does not imply that the current GenUI demo ships with or depends on a fine-tuned model.
The… See the full description on the dataset page: https://huggingface.co/datasets/zacwhite/genui-compiler-corpus.mmlu-college-computer-science-compilers
Extensions to the MMLU Computer Science Datasets for specialization in compilers
This dataset contains data specialized in the compilers domain.
** Dataset Details **
Number of rows: 95
Columns: topic, context, question, options, correct_options_literal, correct_options, correct_options_idx
** Usage **
To load this dataset:
python
from datasets import load_dataset
dataset = load_dataset("masoudc/mmlu-college-computer-science-compilers")
cpp-compiler-curriculum
C++ compiler curriculum (SFT)
Synthetic C++ examples derived from Ubuntu toolchain headers and verified with g++ -std=c++20.
cpp-compiler-prefs
C++ compiler preferences (DPO)
Offline preferences: chosen answers compile; rejected answers fail g++.
design_compiler_pdf_datasetmedgate-compiler-data
MedGate Compiler Training Data
The Logic-Grounding Corpus for Computable Medical Law
This dataset is a hand-curated collection of 606 high-fidelity mappings designed to train Structural Compilers. It facilitates the translation of unstructured clinical guidelines and medical policy prose into machine-executable symbolic logic (JSON).
Dataset Summary
The MedGate Compiler Training Data provides the ground truth for Nexus Forensic – Layer 0 (Protocol Vault).
Each record… See the full description on the dataset page: https://huggingface.co/datasets/Nick-Maximillien/medgate-compiler-data.java_compilercpp-compiler-grpo
C++ compiler GRPO prompts
Prompt-only dataset for GRPO with g++ compile/run rewards.
qwen-coder-compilershf_medical_debug_compiler_v5_urdu_batch2.jsonjava_one_compiler_feedbackdesign_compiler_md_datasettelugu_compiler_debugging_v7hf-laica-compilerhf_medical_debug_compiler_v4_urdu.jsonhf_medical_debug_compiler_v4_urdu_full.jsoncompiler_optimization_failure_pairs_v4data_telugu_compiler_v4_01.jsonsanskrit_compiler_med_reasoning_v1.jsoncompiler_hot_paths
Dataset Card for Compiler Hot Paths
Dataset Description
This dataset consists of 1561 compiler paths generated from 26 C programs in the Polybench Benchmark Suite using the Ball-Larus Algorithm.
Each path, a sequence of LLVM IR instructions, is has three associated values:
count, an integer indicating the number of times this path is executed in the original program.
source_file, a string indicating which program was this path from.
label, an integer of 0 or 1… See the full description on the dataset page: https://huggingface.co/datasets/zhaojer/compiler_hot_paths.
