datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Micro-Model-Bench
Micro-Model-Bench
Micro-Model-Bench is a collection of benchmark results that have been collected using lm-evaluation-harness
Models Recorded
2026-08-27:
154 model records across 49 organizations
125 records marked isValid: true, all other models haven't been evaluated because of gated access or lm-eval not being able to benchmark them.
What benchmarks are included:
ARC-Easy: arc_easy_acc, arc_easy_acc_norm
ARC-Challenge: arc_challenge_acc… See the full description on the dataset page: https://huggingface.co/datasets/veyra-ai/Micro-Model-Bench.SciCloze-900
SciCloze-900
SciCloze-900 is a cloze-style GCSE Combined Science benchmark for evaluating small base language models.
The benchmark contains 900 multiple-choice cloze items:
300 biology items
300 chemistry items
300 physics items
Each item is designed as a natural text continuation rather than an instruction-style question. The intended evaluation method is to score each answer choice by average log probability per token as a continuation of the prompt.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/veyra-ai/SciCloze-900.
