CoolFace
Datasetpublic

madesai/what-ai-benchmarks-actually-measure

What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models Item-level model responses and scores for 53 language models across the 56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks (Desai et al., 2026, arxiv.org/abs/2609.08812). We do not release the prompts from the benchmark datasets, but instead refer to them by item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.

sourceHugging Facecc-by-4.0updated 18d agoView on Hugging Face
0likes1.4kdownloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
madesai/what-ai-benchmarks-actually-measure · CoolFace