datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
disco-model-outputs
DISCO model outputs
Tabular release of per-model, per-item correctness and answer scores used to train and evaluate DISCO: Diversifying Sample Condensation for Efficient Model Evaluation. The paper studies cheap benchmark performance prediction from a small subset of evaluation items; this dataset supplies the raw harness-style outputs for MMLU (57 subjects), HellaSwag, Winogrande, ARC, and related tasks from the Open LLM Leaderboard ecosystem.
Paper
Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/arubique/disco-model-outputs.spider-model-outputs-wo-gpt35spider-model-outputsmodel_outputs-2025-03-11-12-45-37
