datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
benchpress-score-matrix
BenchPress Score Matrix
This dataset contains the public model-by-benchmark score matrix used by
BenchPress. The release includes the lossless audited JSON, benchmark cost
evidence, flat model and benchmark metadata, one row per observed score, and
the paper-canonical dense subset used in the BenchPress experiments.
The source repository is
microsoft/benchpress.
Canonical artifacts
data/llm_benchmark_data.json is the authoritative rich score-matrix artifact.
It… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/benchpress-score-matrix.bench-press-deadlift-exercises
Bench Press and Deadlift Exercise Videos
This dataset contains short exercise clips for binary video classification.
The source frame folders were encoded as MP4 files so the repository follows the Hugging Face VideoFolder layout.
Set the final license before publishing this dataset publicly.
Labels
bench_press
deadlift
Splits
split
total
bench_press
deadlift
train
75
49
26
validation
9
6
3
test
9
6
3
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/bench-press-deadlift-exercises.benchpresssft-eval-results
SFT Eval Results
lm-evaluation-harness (swiss-ai fork) results for three tulu3-SFT models, each at
5 evenly-spaced SFT checkpoints, on the minimal 30-task set. Generated with the
Tülu-3 chat template (fewshot as multiturn), vLLM backend, on GH200.
Models (× checkpoints step-165, step-990, step-1980, step-2805, step-3667)
smollm2-1.7b-sft — Kausp11/smollm2-1.7b-tulu3-sft
qwen2.5-1.5b-sft — Kausp11/qwen2.5-1.5b-tulu3-sft
llama3.2-1b-sft —… See the full description on the dataset page: https://huggingface.co/datasets/benchpress/sft-eval-results.results-multi-temp-seed
Benchpress Evaluation Results
This dataset contains normalized outputs from lm-evaluation-harness
runs for the Benchpress project. It is intended for analysis of
post-training recipe behavior across models, benchmarks, temperatures,
and random seeds.
Coverage
Runs: 110
Model/stage recipes: olmo2-13b:dpo, olmo2-13b:rlvr, olmo2-13b:sft, olmo2-7b:dpo, olmo2-7b:rlvr, olmo2-7b:sft, olmo3-7b:dpo, olmo3-7b:rlvr, olmo3-7b:sft, smollm3-3b:apo, smollm3-3b:sft
Temperatures:… See the full description on the dataset page: https://huggingface.co/datasets/benchpress/results-multi-temp-seed.posttraining-eval-results-extensionbenchpress-score-matrix
BenchPress Score Matrix
This dataset contains the public model-by-benchmark score matrix used by
BenchPress. The release is a tabular artifact: model metadata, benchmark
metadata, one row per observed score, and the paper-canonical dense subset used
in the BenchPress experiments.
The source repository is
microsoft/benchpress. This export was
generated from commit 5be3b4eddf0188721ff25f00713b589b2cbed8e0.
Files
File
Contents
data/scores_all.csv /… See the full description on the dataset page: https://huggingface.co/datasets/JudyLi1122/benchpress-score-matrix.benchpressposttraining-eval-resultsresultsBenchPress
BenchPress
BenchPress is the public benchmark release for SynthEdit. It contains the
benchmark-facing assets only: a portable parquet plus the paired before/after
images referenced by each row.
Contents
1500 benchmark rows
2558 images
6.40 GiB of image data
Source parquet: /zhome/ca/9/146686/synthedit2/data/eval_mc_negatives.parquet
Layout
benchpress.parquet — portable metadata parquet with dataset-relative image paths… See the full description on the dataset page: https://huggingface.co/datasets/JensParslov/BenchPress.
