datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bonsai-jetson-benchmark-15w
Bonsai Jetson Benchmark — 15W
Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 15W
Backend: llama.cpp build-jetson · CUDA · -ngl 99
Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok · 20 reqs/combo
Status: Complete — 57 combos (5 models × 12 prompt/gen configs)
Key metric: tok/J = output tok/s ÷ VDD_CPU_GPU_CV (W)
Models
Model
Quant
Size
Bonsai-1.7B
Q1_0 (1-bit)
~237 MB
Bonsai-4B
Q1_0 (1-bit)
~540 MB
Bonsai-8B
Q1_0… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-15w.KernelBenchX
KernelBenchX
Reproducible evaluation benchmark for Triton GPU-kernel code generation by LLMs — measures buildability, numerical correctness against a deterministic test suite, and end-to-end speedup vs. a GPU-matched golden reference.
Paper: arXiv:2605.04956 · hf.co/papers/2605.04956
Evaluation harness: https://github.com/BonnieW05/KernelBenchX
Configs
Config
Rows
What it is
tasks
176
Benchmark task specs + PyTorch reference + deterministic test harness… See the full description on the dataset page: https://huggingface.co/datasets/BonnieWang/KernelBenchX.math500-bon-weighted-results
MATH-500 Best-of-N Weighted Selection Results
Dataset Description
This dataset contains the results of evaluating Best-of-N weighted selection on a subset of the MATH-500 benchmark. It was created as part of a HuggingFace internship exercise exploring how test-time compute scaling with reward models can improve LLM performance on math problems.
How It Was Constructed
1. Problem Selection
Started from the HuggingFaceH4/MATH-500 dataset (500 problems)… See the full description on the dataset page: https://huggingface.co/datasets/cmpatino/math500-bon-weighted-results.post-training-takehome-math500-bon16
MATH-500 Best-of-16 Post-Training Take-Home Results
A 50-problem study of test-time compute, based on the Hugging Face post-training take-home challenge. Nothing here trains or modifies a model: both the generator and the reward model stay frozen, and the only variable is how a final answer is chosen from 16 sampled candidates.
Construction
Filtered MATH-500 to levels 1-3, shuffled with seed 1, and selected 50 rows.
Generated one greedy solution per problem with… See the full description on the dataset page: https://huggingface.co/datasets/augustoFranke/post-training-takehome-math500-bon16.
