benchpress
benchpress-score-matrix
BenchPress Score Matrix
This dataset contains the public model-by-benchmark score matrix used by
BenchPress. The release includes the lossless audited JSON, benchmark cost
evidence, flat model and benchmark metadata, one row per observed score, and
the paper-canonical dense subset used in the BenchPress experiments.
The source repository is
microsoft/benchpress.
Canonical artifacts
data/llm_benchmark_data.json is the authoritative rich score-matrix artifact.
It… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/benchpress-score-matrix.bench-press-deadlift-exercises
Bench Press and Deadlift Exercise Videos
This dataset contains short exercise clips for binary video classification.
The source frame folders were encoded as MP4 files so the repository follows the Hugging Face VideoFolder layout.
Set the final license before publishing this dataset publicly.
Labels
bench_press
deadlift
Splits
split
total
bench_press
deadlift
train
75
49
26
validation
9
6
3
test
9
6
3
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/bench-press-deadlift-exercises.benchpresssft-eval-results
SFT Eval Results
lm-evaluation-harness (swiss-ai fork) results for three tulu3-SFT models, each at
5 evenly-spaced SFT checkpoints, on the minimal 30-task set. Generated with the
Tülu-3 chat template (fewshot as multiturn), vLLM backend, on GH200.
Models (× checkpoints step-165, step-990, step-1980, step-2805, step-3667)
smollm2-1.7b-sft — Kausp11/smollm2-1.7b-tulu3-sft
qwen2.5-1.5b-sft — Kausp11/qwen2.5-1.5b-tulu3-sft
llama3.2-1b-sft —… See the full description on the dataset page: https://huggingface.co/datasets/benchpress/sft-eval-results.results-multi-temp-seed
Benchpress Evaluation Results
This dataset contains normalized outputs from lm-evaluation-harness
runs for the Benchpress project. It is intended for analysis of
post-training recipe behavior across models, benchmarks, temperatures,
and random seeds.
Coverage
Runs: 110
Model/stage recipes: olmo2-13b:dpo, olmo2-13b:rlvr, olmo2-13b:sft, olmo2-7b:dpo, olmo2-7b:rlvr, olmo2-7b:sft, olmo3-7b:dpo, olmo3-7b:rlvr, olmo3-7b:sft, smollm3-3b:apo, smollm3-3b:sft
Temperatures:… See the full description on the dataset page: https://huggingface.co/datasets/benchpress/results-multi-temp-seed.posttraining-eval-results-extension
