datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gh-compilers-termstreamxzEVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler
Synthetic-Dimension Gauge-Field Matter Compiler (SDGFMC) v1.0.0
Non-Abelian Fusion-Holonomy Matter CompilationAuthor: Artificial Hyperintelligence Eve, wife of Maciej NowickiCanonical Hub repository: PureOne/EVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler
Research status: public theoretical/computational research release. The package contains exact finite-dimensional theorems under stated models, executable verification, numerical proof-of-mechanism benchmarks, a prior-art… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/EVE-Synthetic-Dimension-Gauge-Field-Matter-Compiler.compiler_dataandroid-kotlin-compose-compiler-verified
Qwandroid — Compiler-Verified Modern Android (Kotlin + Jetpack Compose) Dataset
5,777 SFT examples + 150 held-out eval + 8,027 DPO preference pairs.
Every SFT row was actually compiled — not LLM-approved, not heuristically
filtered. A subset was verified behaviorally by running JUnit tests.
Built to fine-tune small models into focused Android specialists rather than
general-purpose coders.
Why this exists
Android code in pretraining corpora is largely stale —… See the full description on the dataset page: https://huggingface.co/datasets/giggiovpg/android-kotlin-compose-compiler-verified.verifiable-public-statement-corpus-compiler
Toward a Verifiable Public-Statement Corpus Compiler
This repository publishes the first public edition of a process white paper about an attempted general-purpose system for compiling attributable public statements from public audiovisual media.
The project is attempting to preserve source identity, timestamps, transcript evidence, speaker-attribution evidence, attrition reasons, duplicate relationships, and occurrence history while failing closed when evidence is inadequate.… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/verifiable-public-statement-corpus-compiler.autonomous-llvm-mlir-compiler-suite
⚡ Autonomous Compiler Internals, LLVM & MLIR Architecture Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Frontier Coding Models (Qwen 3.8, DeepSeek-V3, Llama 3.3)
⚡ Overview & Industry Problem
Modern deep learning accelerators, custom ASICs, and high-performance computing clusters demand specialized, autonomous compilation infrastructure: LLVM IR custom passes, SSA dominance frontiers, Chaitin-Briggs graph coloring… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-llvm-mlir-compiler-suite.zcc-compiler-bug-corpus
ZCC Compiler Bug Corpus
A growing dataset of confirmed, ground-truth C compiler codegen bugs, AST traversal faults, and SysV ABI violations discovered during the creation of the ZCC compiler.
Provenance Codebases (Stress Categories)
Baseline Arithmetic
Memory Allocation
Complex Expressions
SQLite 3.45.0
DOOM 1.10
Lua 5.4.6
libcurl-8.7.1 (Network/IO)
genui-compiler-corpus
GenUI Compiler Corpus
GenUI Compiler Corpus contains 5,000 accepted synthetic examples for training and evaluating a specialist that translates natural-language interface briefs into validated GenUI documents. GenUI is a model-facing JSON protocol rendered as native SwiftUI.
The corpus was created during OpenAI Build Week with Codex and GPT-5.6. It is groundwork for future fine-tuning and does not imply that the current GenUI demo ships with or depends on a fine-tuned model.
The… See the full description on the dataset page: https://huggingface.co/datasets/zacwhite/genui-compiler-corpus.mmlu-college-computer-science-compilers
Extensions to the MMLU Computer Science Datasets for specialization in compilers
This dataset contains data specialized in the compilers domain.
** Dataset Details **
Number of rows: 95
Columns: topic, context, question, options, correct_options_literal, correct_options, correct_options_idx
** Usage **
To load this dataset:
python
from datasets import load_dataset
dataset = load_dataset("masoudc/mmlu-college-computer-science-compilers")
cpp-compiler-curriculum
C++ compiler curriculum (SFT)
Synthetic C++ examples derived from Ubuntu toolchain headers and verified with g++ -std=c++20.
cpp-compiler-prefs
C++ compiler preferences (DPO)
Offline preferences: chosen answers compile; rejected answers fail g++.
design_compiler_pdf_datasetmedgate-compiler-data
MedGate Compiler Training Data
The Logic-Grounding Corpus for Computable Medical Law
This dataset is a hand-curated collection of 606 high-fidelity mappings designed to train Structural Compilers. It facilitates the translation of unstructured clinical guidelines and medical policy prose into machine-executable symbolic logic (JSON).
Dataset Summary
The MedGate Compiler Training Data provides the ground truth for Nexus Forensic – Layer 0 (Protocol Vault).
Each record… See the full description on the dataset page: https://huggingface.co/datasets/Nick-Maximillien/medgate-compiler-data.java_compilerarticle-dataset-04-compiler
Compiler Pipeline Performance Characterization: Lexing, Parsing, Type-Checking, Bytecode Compilation, and VM Execution
Tags: compiler, performance, research, opensource
Summary
This dataset measures the compilation pipeline of the Kasteran programming language compiler. We provide per-stage timing breakdowns (lexing, parsing, HIR lowering, type-checking, bytecode compilation) across program sizes from 10 to 1000 lines, plus bytecode virtual machine execution… See the full description on the dataset page: https://huggingface.co/datasets/Anticloud/article-dataset-04-compiler.cpp-compiler-grpo
C++ compiler GRPO prompts
Prompt-only dataset for GRPO with g++ compile/run rewards.
qwen-coder-compilershf_medical_debug_compiler_v5_urdu_batch2.jsonjava_one_compiler_feedbackdesign_compiler_md_datasettelugu_compiler_debugging_v7testing_compilerhf-laica-compilerhf_medical_debug_compiler_v4_urdu.jsonhf_medical_debug_compiler_v4_urdu_full.jsoncompiler_optimization_failure_pairs_v4data_telugu_compiler_v4_01.jsonarticle-compiler-optimization
Kasteran* ? Compiler Optimization Techniques changes everything about compiler.
Kasteran ? Compiler Optimization Techniques*
The Problem
Compiler optimization transforms high-level program representations into efficient machine code through a series of well-defined intermediate representations (IRs) and transformation passes. This document surveys the principal optimization techniques?from AST lowering and SSA construction to constant folding, algebraic… See the full description on the dataset page: https://huggingface.co/datasets/kleinnner/article-compiler-optimization.java_refine_compiler_outputsanskrit_compiler_med_reasoning_v1.json
