fastkernels
fastkernels-b200-codex-vanilla-20260923
Vanilla Codex on FastKernels: independent vs. compositional kernel generation
A single codex exec session per operator, over all 133 L1-L3 operators of
the FastKernels default B200 capture, run twice: once with every operator
optimised independently, and once with L2/L3 regenerated on top of the frozen
winning kernels of the levels below. 1.4 billion tokens, 8x B200, one day.
What this run adds
FastKernels argues for its L1-L4 compositional hierarchy as a design… See the full description on the dataset page: https://huggingface.co/datasets/Alexsssu/fastkernels-b200-codex-vanilla-20260923.fastkernels-validate-h200-20260923
FastKernels full validate on 8×H200 — 2026-09-23
End-to-end results of a complete FastKernels validate full sweep on 8×NVIDIA H200 SXM 141GB: all 47 scenarios of fastkernels/scenarios/full.yaml across 22 benchmark harnesses, comparing FastKernels against each architecture's production or upstream reference.
46 / 46 submitted scenarios produced results. Across the 45 architectures with a throughput ratio, FastKernels reaches a geometric mean of 1.205× against the reference… See the full description on the dataset page: https://huggingface.co/datasets/Alexsssu/fastkernels-validate-h200-20260923.fastkernels-h200-vllm-ablation
vLLM 0.18 → 0.26 → 0.29 on H200: ablation, a three-way comparison
2026-09-22, 8×H200 (alexsu-dev-h200-0). How it was run, the runtime stacks and
the uncommitted local changes are all in MANIFEST.md.
This is the Hopper companion to the B200 ablation of 2026-09-21. The scenario
table, the harness, the tensor-parallel degrees and all three runtime stacks are
byte-for-byte the same as the B200 run, so the two are directly comparable. The
headline result is that most of the B200… See the full description on the dataset page: https://huggingface.co/datasets/Alexsssu/fastkernels-h200-vllm-ablation.fastkernels-b200-vllm-ablation
vLLM 0.18 → 0.26 → 0.29: ablation, a three-way comparison
2026-09-21, 8×B200 (alexsu-dev-b200-0). How it was run, the runtime stacks and
the uncommitted local changes are all in MANIFEST.md.
This is the first environment-clean dataset, and it is citable. Every
performance number from the 2026-09-20 round is void: the 0.26 leg was
contaminated by the vllm-omni plugin, and all jobs ran concurrently. All three
legs in this report use isolated venvs, run strictly serially with the… See the full description on the dataset page: https://huggingface.co/datasets/Alexsssu/fastkernels-b200-vllm-ablation.FastKernels-DRKernelwildchat-fastkernels-balanced-1k
