datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ballast-evalsets
Ballast eval sets
The probe sets behind the Ballast
measurements. Two files:
matrix_probes.parquet: 50,147 factual questions (recall), each with a
gold answer, alias set, 7 type-matched distractors, and (where resolvable)
a Wikidata subject Q-id linking it to the
Ballast T0 corpus
(90.5% linked at full corpus). Measured on two model families: Gemma-4
(E2B/E4B/12B) and Qwen3.5 (0.8B/2B/4B/9B), plus a quantization sweep over
bf16 / fp8 / nf4 / Q6_K / Q4_K_M.… See the full description on the dataset page: https://huggingface.co/datasets/OpenBallast/ballast-evalsets.daft-math
DAFT Math: Difficult Automatically-scorable Free-response Tasks for Math
Dataset Description
⚠️ Note: The dataset has important limitations and we strongly recommend reading the limitations section below before using it. It is not a formal METR benchmark and was originally designed for a very niche use-case. We present it only as a research artifact.
DAFT-Math is a collection of 199 challenging mathematical problems chosen to be at the limit of current LLM abilities… See the full description on the dataset page: https://huggingface.co/datasets/metr-evals/daft-math.
