datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BubbleML_2
BubbleML 2.0:
BubbleML_2 is a high-fidelity dataset of boiling simulations in 2D for three fluids (FC-72, Liquid N2 and R515B). It provides paired time-series fields stored in HDF5 (.hdf5) files together with metadata (.json) and explicit train/test splits.
🚀 Quickstart
The current available dataset subsets are-
"single-bubble", "pb-saturated", "pb-subcooled", "fb-velscale", "fb-chf"
They are chosen for each individual forecasting task in our paper, viz. Single Bubble… See the full description on the dataset page: https://huggingface.co/datasets/hpcforge/BubbleML_2.hpca2027-traces
hpca2027-traces
Continuous DynamoRIO / drmemtrace collections of CRONO graph analytics kernels on soc-LiveJournal1, packaged in the same HF/humza layout used by related ifuse / HPCA characterization traces.
Folder names use *-soc-livejournal so CRONO LiveJournal bundles stay distinct from other graph inputs and from DC-app traces in this dataset.
CRONO workloads (graph=soc-LiveJournal1)
Folder
Binary
Whole fetched instrs
SimPoints (traces_simp)… See the full description on the dataset page: https://huggingface.co/datasets/deepanjalimishra99/hpca2027-traces.ifuse-hpca2027-traces
I-Fuse HPCA 2027 Traces
This dataset consolidates the final replay-ready SimPoint trace set used for
the HPCA 2027 I-Fuse experiments:
Agentic: AppWorld, CORE-Bench, MLGym, and Terminal-Bench
Database and systems: ClickHouse, DuckDB, LevelDB, Masstree, RocksDB, Silo,
and Zstd
GAP: BC, BFS, DFS, PageRank, and SSSP
CRONO: APSP and TSP
Each workload preserves the same bundle structure used by the original trace
repositories, including trace_clustering_info.json, SimPoint… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/ifuse-hpca2027-traces.HPC3-financeHPC3FinanceChunkRetrievalhpc-syncHPC_Fortran_CPPThis dataset is associated with the following paper:
Creating a Dataset for High-Performance Computing Code Translation using LLMs: A Bridge Between OpenMP Fortran and C++,
Links
https://arxiv.org/abs/2307.07686
https://github.com/bin123apple/OpenMP-Fortran-CPP-Translation
hpca2027-final-traceshpca2027-database-validated-traces-20260711
HPCA 2027 validated database traces
Validated DynamoRIO instruction traces for database application-plus-input
configurations used in HPCA 2027 backend-bound characterization.
Included workloads
Directory
Application and input
Captured fetched instructions
Scarab backend bound
mongodb
MongoDB 7.0.37, 10M-record YCSB database, sort/aggregate, 1 GiB WiredTiger cache, 3 GiB container
100,098,204 global across server threads
41.76% on the 89,501… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/hpca2027-database-validated-traces-20260711.hpc-instructThis is an HPC code instruct dataset that was used to train the HPC-Coder-v2 models. There are 122k samples generated synthetically using Gemini Pro, DBRX,Llama-3 and Mixtral.
There are four types of instruct samples in HPC-Instruct detailed below.
Code Synthesis: The instruction tasks the LLM to generate code to solve an HPC related problem.
Parallelization: The instruction tasks the LLM to parallelize an existing sequential code.
Optimization: The instruction tasks the LLM to optimize an… See the full description on the dataset page: https://huggingface.co/datasets/hpcgroup/hpc-instruct.hpc-chatbot-logshpca2027-agentic-validated-traces-20260711
HPCA 2027 CPU-Local Agentic Validated Traces
This dataset contains validated CPU-local traces used for HPCA 2027 agentic
workload characterization. AppWorld, CORE-Bench, and Terminal-Bench are
paper-backed benchmark workloads from official repositories; the hybrid-RAG
and data-analysis agents are deterministic characterization workloads retained
from the initial pass.
Workloads And Full 100M Results
Workload
Source
Frontend
Backend
Bad speculation
Retiring… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/hpca2027-agentic-validated-traces-20260711.HPCPerfOpt-MCQAThis dataset contains Multiple Choice question-answer pairs.
There are 3 test files separated on the basis of how they were created:
test1.csv manual data collection from tutorials, etc
test2.csv scraped profiling tool Codee documentation
test3.csv ChatGPT-generated-MCQ (need to update format and randomize answers.)
hpc-multipicomusic_prevall_hpcs_tokenized_32khzhpc-intrusion-dataset
Server-4 Hardware Performance Counter Dataset
Author
Syed Muhammad Sarim Ahsan
Description
The Server-4 Hardware Performance Counter (HPC) dataset is a real-world, high-fidelity microarchitectural dataset designed for empirical security and performance analysis in cloud environments. Collected on a virtual machine node powered by an AMD EPYC Turin processor running Ubuntu 22.04 LTS (Linux kernel 5.15) under full KVM hypervisor virtualization… See the full description on the dataset page: https://huggingface.co/datasets/sarimahsan101/hpc-intrusion-dataset.hpca2027-final-seven-traces-20260716
HPCA 2027 Final I-Fuse Trace Bundles
This dataset consolidates the final replay-ready SimPoint trace set:
Agentic: AppWorld, CORE-Bench, MLGym, and Terminal-Bench
Database: DuckDB, LevelDB, and RocksDB
Legacy GAP: BC, BFS, DFS, PageRank, and SSSP
Each workload preserves the same bundle structure used by the original trace
repositories, including trace_clustering_info.json, SimPoint selections and
weights, replay-ready ZIPs under traces_simp/trace, module metadata,
fingerprints… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/hpca2027-final-seven-traces-20260716.HPC
This dataset includes two tasks for high-performance computing (HPC) domain.
Task 1 is managing AI models and datasets which includes programming language processing (PLP) and MLPerf.
Task 2 is data race detection which includes c/c++ language and fortran language.
HPC_labeled_11_6hpchpca2027-agent-traces
HPCA 2027 — agent execution traces
DynamoRIO drcachesim memory traces of LLM agent execution, reduced to
SimPoint representatives and packaged for Scarab
via scarab-infra.
Companion to hpca2027-final-traces, which covers kernels and databases. This
repo holds the agent workloads, which are one to two orders of magnitude longer
and so are published as slices only.
What an agent trace contains
The agent scaffold's own Python, plus every tool subprocess it spawns… See the full description on the dataset page: https://huggingface.co/datasets/deepanjalimishra99/hpca2027-agent-traces.HPC_corpus_subdetails_hpcai-tech__Colossal-LLaMA-2-7b-base
Dataset Card for Evaluation run of hpcai-tech/Colossal-LLaMA-2-7b-base
Dataset Summary
Dataset automatically created during the evaluation run of model hpcai-tech/Colossal-LLaMA-2-7b-base on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_hpcai-tech__Colossal-LLaMA-2-7b-base.gemma-3-4b-responsesMATH_500_MMLU_Pro
Description
The MATH_500 and MMLU_Pro datasets combined in a custom format.
Format for the data
Field
Type
Description
unique_id
string
MD5 hash of the entire record (all other columns), computed on the canonical JSON representation with sorted keys.
question
string
The full question text.
category
string
The subject or category of the question.
choices
array of strings | null
List of answer choices for multiple-choice questions; null for open-ended… See the full description on the dataset page: https://huggingface.co/datasets/HPC-Boys/MATH_500_MMLU_Pro.qwen-3-14b-responsesHPC3-rerankingqwen-3-14bAIME_1983_2024hpc-humans
