datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kern-kernels
kern-kernels
Reproducible attention kernel recipes, ABI manifests, checksums and measured
results. First profile: GB300 / Qwen3.8-27B / BF16 / Q24-KV4-D256 / page 64.
Uses unmodified TRTLLM-GEN full attention, not MLA. The model's GDN layers are
unchanged. Model weights and the base export's other kernels are not included.
NVIDIA binaries are downloaded directly from pinned upstream URLs and verified
by SHA256; this repository does not mirror them. The small Apache-2.0 vLLM KV… See the full description on the dataset page: https://huggingface.co/datasets/susun-123/kern-kernels.ga104-cuda-kernels
GA104 Hand-Optimized CUDA Kernel Corpus
A measurement corpus of hand-optimized CUDA / SASS kernels targeting the
RTX 3070 Ti (GA104, sm_86, Ampere). Every kernel is written without
cuBLAS, cuDNN, or PyTorch in the optimized path; vendor libraries are
linked only for measured comparison under kernels/reference/. This
dataset is for SASS and GPU-optimization researchers — it pairs each
.cu source with its compiled machine code and its disassembly, so the
exact instruction stream a… See the full description on the dataset page: https://huggingface.co/datasets/pjt222/ga104-cuda-kernels.cuda-triton-gpu-kernels-2026
⚡ Complete 2026 CUDA & OpenAI Triton High-Performance GPU Kernel Engineering SFT/DPO Suite
The definitive, production-grade synthetic alignment dataset engineered for training and fine-tuning open-weights Large Language Models (Qwen 2.5 Coder, DeepSeek-Coder, Llama 3.1) on ultra-high-throughput GPU kernel programming: NVIDIA Hopper H100 / Blackwell B200 TMA async transfers, OpenAI Triton 3.1+ FlashAttention-3, 32-bank conflict elimination, and low-bit FP8 / INT4 GEMM… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cuda-triton-gpu-kernels-2026.ParallelKernelBench_Kernels
ParallelKernelBench Kernels
Net-new multi-GPU CUDA kernels generated by LLMs for ParallelKernelBench.
Each subdirectory under solutions/ is one model run. File names match the benchmark problem stems (e.g. 17_rope_allgather_cuda.py ↔ problem 17_rope_allgather in willychan21/ParallelKernelBench_Problems).
Layout
solutions/
<run_id>/
<stem>_cuda.py
...
Runs (1 run(s), 87 kernel files)
run_id
kernels
path… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Kernels.litmus-kernels
Litmus Kernel Verification Corpus
Correct and deliberately-broken Triton kernels, each broken one shipped with
the input that exposes it.
The corpus exists to measure one thing: how much of what a fixed-shape
torch.rand() allclose test calls "correct" actually is. On this corpus the
answer is that 88% of the planted bugs pass that test.
Columns
column
meaning
name
kernel identifier
family
elementwise / reduction / softmax / layernorm / matmul /… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/litmus-kernels.Kernel-Smith-Seed-59KIf this work is useful to you, please cite:
@article{DBLP:journals/corr/abs-2603-28342,
author = {He Du and
Qiming Ge and
Jiakai Hu and
Aijun Yang and
Zheng Cai and
Zixian Huang and
Sheng Yuan and
Qinxiu Cheng and
Xinchen Xie and
Yicheng Chen and
Yining Li and
Jiaxing Xie and… See the full description on the dataset page: https://huggingface.co/datasets/CoopReason/Kernel-Smith-Seed-59K.b-sides-v2-kernels
B-Sides kernel search index
Published runtime artifacts for the B-Sides kernel corpus.
meta.sqlite: searchable kernel metadata
embeddings.f16.npy: row-aligned 256-dimensional query matrix
The kernels.row values align one-to-one with the embedding matrix rows.
FECA_ECOLI_Tsuboyama_2023_2D1U_substitutions_singles_stability_PE_REGRRL20_AQUAE_Tsuboyama_2023_1GYZ_substitutions_singles_stability_PE_REGR
