CoolFace
Datasetpublic

anonymous-gpu-kernel/anonymous-gpu-kernel

Anonymous GPU Kernel Dataset This repository is an anonymized artifact preview for a GPU kernel generation and optimization dataset. It contains representative HIP/ROCm and Triton kernel data, metadata, validation information, and production-grounded ROCm-library QA entries. The full release will include the complete HIP and Triton kernel data, validation metadata, and provenance information. This dataset is organized as versioned releases. This checkout focuses on the v0.2… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-gpu-kernel/anonymous-gpu-kernel.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes26downloads
Dataset Card

Anonymous GPU Kernel Dataset

This repository is an anonymized artifact preview for a GPU kernel generation and optimization dataset. It contains representative HIP/ROCm and Triton kernel data, metadata, validation information, and production-grounded ROCm-library QA entries.

The full release will include the complete HIP and Triton kernel data, validation metadata, and provenance information.

This dataset is organized as versioned releases. This checkout focuses on the v0.2 release layout and keeps only the dataset subsets intended for use in the current training/evaluation workflows.

v0.2 release layout

Dataset subsetLocal pathCount / note
HIP-CudaAgentv0.2/pytorch_hip_kernel_cuda_agent_ops_6k/5,388 PyTorch→HIP entries derived from CUDA-Agent-Ops-6K
HIP-GPUModev0.2/pytorch_hip_kernel_gpumode/5,910 unique PyTorch entries; 22,397 HIP variant files under the v0.1 tar archive
HIP2HIP (Optimization)v0.2/hip-to-hip/34,368 HIP→HIP optimization entries, built on HIP-GPUMode
ROCm Libraries QAv0.2/rocm-libraries/2,377 rocBLAS/rocSOLVER QA-style entries
Triton datasetsv0.2/PyTorch_triton_datasets/Triton-Stack / Triton-Bench / Triton-GPUMode / Triton-AICE release files

Documentation

  • `CORPUS_AUDIT.md`: source composition, sample counts, difficulty distribution, operator coverage.
  • `VALIDATION_PROTOCOL.md`: compilation, correctness, dtype/shape coverage, and latency validation.
  • `LICENSES_AND_PROVENANCE.md`: source datasets, provenance, and per-subset license notes.

See `LICENSES_AND_PROVENANCE.md` for per-subset license and provenance information.

Previewing the data

Large dataset files are tracked with Git LFS and may not render in the Hugging Face Dataset Viewer. To inspect the schema and representative content without downloading the full files, use the standard JSON-array preview files:

  • Per subset: v0.2/<subset>/sample_entries.json.