datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pmpp-hard
PMPP-Hard Agent Evaluation Traces
PMPP-Hard is a 69-task agentic GPU-kernel evaluation for testing whether autonomous coding agents can produce implementations that are both correct and performant. This dataset contains the complete nine-model campaign used in the PMPP-Hard release: 621 rollouts, with 69 task sessions for each model configuration.
Maintained and released by Sinatras.
Source repository: SinatrasC/pmpp-hard
Prime environment and evaluations: PMPP-Hard on Prime… See the full description on the dataset page: https://huggingface.co/datasets/sinatras/pmpp-hard.pmpp-eval
PMPP Dataset
This repository provides two CUDA-focused datasets prepared by Sinatras and sponsored by Prime Intellect. Both datasets are based on Programming Massively Parallel Processors (4th Ed.) with additional coding evaluation harnesses at https://github.com/SinatrasC/pmpp-eval to be used by PMMP env in prime-environments.
Overview
Languages: English
License: MIT
Curated by: Sinatras (https://github.com/SinatrasC)
Sponsored by: Prime Intellect
Derived from: PMPP 4th… See the full description on the dataset page: https://huggingface.co/datasets/sinatras/pmpp-eval.
