datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
benchability-fig4-capability-guided
BenchAbility Figure 4 -- capability_guided
One of two training mixtures drawn from the same frozen 884,143-row candidate pool, with the
same budget (60,000 intervention + 15,000 shared replay) and the same hyperparameters. The two
differ only in how the samples are chosen, which is the whole experiment.
arm
capability_guided
selection
by capability, gap-weighted from the Figure 3 scores
intervention rows
59,999
replay rows
15,000
shards
38
pool
884,143… See the full description on the dataset page: https://huggingface.co/datasets/realzL/benchability-fig4-capability-guided.benchability-fig4-source-uniform
BenchAbility Figure 4 -- source_uniform
One of two training mixtures drawn from the same frozen 884,143-row candidate pool, with the
same budget (60,000 intervention + 15,000 shared replay) and the same hyperparameters. The two
differ only in how the samples are chosen, which is the whole experiment.
arm
source_uniform
selection
by source provenance only, chart:doc:ocr = 3:4:4
intervention rows
60,048
replay rows
15,000
shards
38
pool
884,143 rows / 20… See the full description on the dataset page: https://huggingface.co/datasets/realzL/benchability-fig4-source-uniform.benchability-fig4-eval
BenchAbility Figure 4 -- held-out eval
The 10% candidate-dev half of the same 884,143-row pool the two training mixtures are drawn from,
balanced per capability. Split by image, so no picture here appears in either mixture, and both
arms are equally blind to it.
capability
n
chart_reading / chart_reasoning
250 / 250
table_lookup / table_reasoning
250 / 250
document_qa / document_text_reading
250 / 250
diagram_and_infographic_understanding
250… See the full description on the dataset page: https://huggingface.co/datasets/realzL/benchability-fig4-eval.
