datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MLIRBench
Dataset Card for MLIRBench
MLIRBench is a benchmark dataset for evaluating semantic reasoning, semantic equivalence analysis, execution-aware validation, and compiler-aware reasoning over programs represented in the Multi-Level Intermediate Representation (MLIR).
Dataset Details
Dataset Description
MLIRBench is a benchmark dataset for evaluating semantic reasoning, equivalence analysis, and execution-aware understanding of programs represented in… See the full description on the dataset page: https://huggingface.co/datasets/mlirbench/MLIRBench.autonomous-llvm-mlir-compiler-suite
⚡ Autonomous Compiler Internals, LLVM & MLIR Architecture Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Frontier Coding Models (Qwen 3.8, DeepSeek-V3, Llama 3.3)
⚡ Overview & Industry Problem
Modern deep learning accelerators, custom ASICs, and high-performance computing clusters demand specialized, autonomous compilation infrastructure: LLVM IR custom passes, SSA dominance frontiers, Chaitin-Briggs graph coloring… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-llvm-mlir-compiler-suite.
