micro-model
Micro-Model-Bench
Micro-Model-Bench
Micro-Model-Bench is a collection of benchmark results that have been collected using lm-evaluation-harness
Models Recorded
2026-08-27:
154 model records across 49 organizations
125 records marked isValid: true, all other models haven't been evaluated because of gated access or lm-eval not being able to benchmark them.
What benchmarks are included:
ARC-Easy: arc_easy_acc, arc_easy_acc_norm
ARC-Challenge: arc_challenge_acc… See the full description on the dataset page: https://huggingface.co/datasets/veyra-ai/Micro-Model-Bench.ddi-safety-micro-model
