CoolFace
Datasetpublic

madesai/what-ai-benchmarks-actually-measure

What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models Item-level model responses and scores for 53 language models across the 56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks (Desai et al., 2026, arxiv.org/abs/2609.08812). We do not release the prompts from the benchmark datasets, but instead refer to them by item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.

sourceHugging Facecc-by-4.0updated 18d agoView on Hugging Face
0likes1.4kdownloads
39 commits on main
312b2fc18d ago

Update README.md

madesai
730c3d618d ago

Update README.md

madesai
fd6065018d ago

Update README.md

madesai
66b00f518d ago

Restore license-consult note; clarify scores are mostly-binary

madesai
2224e2e18d ago

Reword benchmark-prompts sentence

madesai
c5d275918d ago

Reword row-schema section intro

madesai
c5c980b18d ago

Revert get_data.py link to data_acquisition/get_data.py

madesai
8aab28118d ago

Update get_data.py link after data_acquisition/ flattening

madesai
88a9aed18d ago

Update score.py link after run_experiments/ removal

madesai
779034c18d ago

Merge item-id/code note into intro; drop redundant Code and What is not here sections

madesai
c71c14f18d ago

Add benchmark count to intro sentence

madesai
db2676218d ago

Update benchmarks.csv row description

madesai
f2983e218d ago

Remove redundant eval_file column from benchmarks.csv

madesai
034ecb018d ago

Remove benchmark_label column

madesai
3d4502118d ago

Move item_id earlier in column order (after paper_status, before subset)

madesai
84f2fbb18d ago

Reword paper_status section

madesai
d90a3a018d ago

Remove note on null score for set-level metrics

madesai
5fc142718d ago

Reword subset description for refusal/overrefusal

madesai
80cc23e18d ago

Simplify score column description

madesai
381135018d ago

Remove build_report.csv from release

madesai
210565018d ago

Remove build_report.csv from release

madesai
b4e2cdf18d ago

Remove run_configs.jsonl from release

madesai
f67873218d ago

Remove run_configs.jsonl from release

madesai
3d9168718d ago

Remove paper_status from metadata/benchmarks.csv description

madesai
0202ae618d ago

Reword aggregate/model_scores row description

madesai
f6cf55018d ago

Reword data/ row description

madesai
527fba518d ago

Remove redundant intro paragraph; simplify scores/ row description

madesai
b97262718d ago

Fold paper citation into intro sentence

madesai
13c1a7318d ago

Shorten paper citation to title-first style

madesai
ef3e79e18d ago

Use full paper citation in card intro

madesai
3bba59d18d ago

Update build_report for standardized column names

madesai
577f9a518d ago

Standardize column names; add citation; reference appendix judge-robustness analysis

madesai
d5f62a218d ago

Collapse paper_status to included/excluded; drop metadata/models.csv; restructure card

madesai
b0cdb4a18d ago

Sort score matrices by model coverage; clarify item_id provenance in card

madesai
e37701f23d ago

Update aggregate scores to current analysis table (Figure B.1 data), refresh model metadata from paper v7, document model list and code repo in README

madesai
c7e81e12mo ago

Upload folder using huggingface_hub

madesai
8d168192mo ago

Upload folder using huggingface_hub

madesai
a58f6c02mo ago

Upload folder using huggingface_hub

madesai
40e45d92mo ago

initial commit

madesai