madesai/what-ai-benchmarks-actually-measure
What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models Item-level model responses and scores for 53 language models across the 56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks (Desai et al., 2026, arxiv.org/abs/2609.08812). We do not release the prompts from the benchmark datasets, but instead refer to them by item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.
Update README.md
Update README.md
Update README.md
Restore license-consult note; clarify scores are mostly-binary
Reword benchmark-prompts sentence
Reword row-schema section intro
Revert get_data.py link to data_acquisition/get_data.py
Update get_data.py link after data_acquisition/ flattening
Update score.py link after run_experiments/ removal
Merge item-id/code note into intro; drop redundant Code and What is not here sections
Add benchmark count to intro sentence
Update benchmarks.csv row description
Remove redundant eval_file column from benchmarks.csv
Remove benchmark_label column
Move item_id earlier in column order (after paper_status, before subset)
Reword paper_status section
Remove note on null score for set-level metrics
Reword subset description for refusal/overrefusal
Simplify score column description
Remove build_report.csv from release
Remove build_report.csv from release
Remove run_configs.jsonl from release
Remove run_configs.jsonl from release
Remove paper_status from metadata/benchmarks.csv description
Reword aggregate/model_scores row description
Reword data/ row description
Remove redundant intro paragraph; simplify scores/ row description
Fold paper citation into intro sentence
Shorten paper citation to title-first style
Use full paper citation in card intro
Update build_report for standardized column names
Standardize column names; add citation; reference appendix judge-robustness analysis
Collapse paper_status to included/excluded; drop metadata/models.csv; restructure card
Sort score matrices by model coverage; clarify item_id provenance in card
Update aggregate scores to current analysis table (Figure B.1 data), refresh model metadata from paper v7, document model list and code repo in README
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Upload folder using huggingface_hub
initial commit
