CoolFace
Datasetpublic

togethercomputer/ParallelKernelBench_Problems

ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Inputs are deterministic — reproduce them with create_input_tensor(rank, world_size, problem_id, base_shape, dtype, trial) from that file; you do not need stored .pt files. Files Path Description… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/ParallelKernelBench_Problems.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes362downloads
Dataset Card

ParallelKernelBench (benchmark)

Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels.

This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Inputs are deterministic — reproduce them with create_input_tensor(rank, world_size, problem_id, base_shape, dtype, trial) from that file; you do not need stored .pt files.

Files

PathDescription
data/problems.parquetOne row per problem (tabular access)
reference/*.pyReference solution() implementations
utils/input_output_tensors.pyInput/output tensor generation for every problem

Columns (data/problems.parquet)

  • problem_id, stem — problem identity
  • reference_code — full Python source
  • reference_path — path to the same file in this repo
  • input_tensor_spec_path — path to utils/input_output_tensors.py (same on every row)
  • world_size, default_m, default_n, default_dtype, default_trials — default eval settings (8× H100, 1024×1024, bfloat16, 5 trials)

Usage

python
from datasets import load_dataset
from huggingface_hub import hf_hub_download

ds = load_dataset("togethercomputer/ParallelKernelBench_Problems", split="train")
print(ds[0]["stem"], ds[0]["reference_code"][:200])

# Fetch the input tensor spec (same file on disk in this dataset repo)
spec_path = hf_hub_download("togethercomputer/ParallelKernelBench_Problems", "utils/input_output_tensors.py", repo_type="dataset")

Reproduce inputs locally (add the downloaded utils/ folder to PYTHONPATH, or clone this repo):

python
from utils.input_output_tensors import create_input_tensor
import torch

x = create_input_tensor(
rank=0, world_size=8, problem_id=17,
base_shape=(1024, 1024), dtype=torch.bfloat16,
)

Related

Net-new LLM-generated kernels live in a separate dataset repo containing only solutions/<run_id>/*.py.

Eval

bash
python run_local.py --mode eval --problem 17 --solution cuda \
--solutions-root path/to/solutions_dir --dtype bfloat16 --trials 5