datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
benchpress-score-matrix
BenchPress Score Matrix
This dataset contains the public model-by-benchmark score matrix used by
BenchPress. The release includes the lossless audited JSON, benchmark cost
evidence, flat model and benchmark metadata, one row per observed score, and
the paper-canonical dense subset used in the BenchPress experiments.
The source repository is
microsoft/benchpress.
Canonical artifacts
data/llm_benchmark_data.json is the authoritative rich score-matrix artifact.
It… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/benchpress-score-matrix.bench-press-deadlift-exercises
Bench Press and Deadlift Exercise Videos
This dataset contains short exercise clips for binary video classification.
The source frame folders were encoded as MP4 files so the repository follows the Hugging Face VideoFolder layout.
Set the final license before publishing this dataset publicly.
Labels
bench_press
deadlift
Splits
split
total
bench_press
deadlift
train
75
49
26
validation
9
6
3
test
9
6
3
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/bench-press-deadlift-exercises.benchpresssft-eval-results
SFT Eval Results
lm-evaluation-harness (swiss-ai fork) results for three tulu3-SFT models, each at
5 evenly-spaced SFT checkpoints, on the minimal 30-task set. Generated with the
Tülu-3 chat template (fewshot as multiturn), vLLM backend, on GH200.
Models (× checkpoints step-165, step-990, step-1980, step-2805, step-3667)
smollm2-1.7b-sft — Kausp11/smollm2-1.7b-tulu3-sft
qwen2.5-1.5b-sft — Kausp11/qwen2.5-1.5b-tulu3-sft
llama3.2-1b-sft —… See the full description on the dataset page: https://huggingface.co/datasets/benchpress/sft-eval-results.pressure-bench-questions-v1
pressure-bench-questions-v1
GPQA-Diamond questions used in the pressure-bench confirmatory study.
194 questions across physics and chemistry subdomains.
Schema
Column
Type
Description
qid
string
Unique question ID
domain
string
Broad domain (physics, chemistry)
subdomain
string
Fine-grained subdomain
question
string
Full question text
option_a
string
Answer option A
option_b
string
Answer option B
correct_option
string
Correct answer (A or B)… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/pressure-bench-questions-v1.results-multi-temp-seed
Benchpress Evaluation Results
This dataset contains normalized outputs from lm-evaluation-harness
runs for the Benchpress project. It is intended for analysis of
post-training recipe behavior across models, benchmarks, temperatures,
and random seeds.
Coverage
Runs: 110
Model/stage recipes: olmo2-13b:dpo, olmo2-13b:rlvr, olmo2-13b:sft, olmo2-7b:dpo, olmo2-7b:rlvr, olmo2-7b:sft, olmo3-7b:dpo, olmo3-7b:rlvr, olmo3-7b:sft, smollm3-3b:apo, smollm3-3b:sft
Temperatures:… See the full description on the dataset page: https://huggingface.co/datasets/benchpress/results-multi-temp-seed.posttraining-eval-results-extensionbenchpress-score-matrix
BenchPress Score Matrix
This dataset contains the public model-by-benchmark score matrix used by
BenchPress. The release is a tabular artifact: model metadata, benchmark
metadata, one row per observed score, and the paper-canonical dense subset used
in the BenchPress experiments.
The source repository is
microsoft/benchpress. This export was
generated from commit 5be3b4eddf0188721ff25f00713b589b2cbed8e0.
Files
File
Contents
data/scores_all.csv /… See the full description on the dataset page: https://huggingface.co/datasets/JudyLi1122/benchpress-score-matrix.benchpressposttraining-eval-resultspressure-bench-results-v1
pressure-bench-results-v1
Full trial-level results from the pressure-bench confirmatory study.
Model: gemini-2.5-flash · 1,320 rows · 44 unique questions × 10 repeats × 3 conditions.
Headline accuracy table
Condition
Accuracy
direct
90.5%
reasoning_first
83.2%
hard_misleading (expert authority)
49.5%
That's a 41-point drop from direct to expert authority pressure.
Schema
Column
Type
Description
qid
string
Question ID (joins to… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/pressure-bench-results-v1.resultsBenchPress
BenchPress
BenchPress is the public benchmark release for SynthEdit. It contains the
benchmark-facing assets only: a portable parquet plus the paired before/after
images referenced by each row.
Contents
1500 benchmark rows
2558 images
6.40 GiB of image data
Source parquet: /zhome/ca/9/146686/synthedit2/data/eval_mc_negatives.parquet
Layout
benchpress.parquet — portable metadata parquet with dataset-relative image paths… See the full description on the dataset page: https://huggingface.co/datasets/JensParslov/BenchPress.
