realzL/benchability-fig4-eval
BenchAbility Figure 4 -- held-out eval The 10% candidate-dev half of the same 884,143-row pool the two training mixtures are drawn from, balanced per capability. Split by image, so no picture here appears in either mixture, and both arms are equally blind to it. capability n chart_reading / chart_reasoning 250 / 250 table_lookup / table_reasoning 250 / 250 document_qa / document_text_reading 250 / 250 diagram_and_infographic_understanding 250… See the full description on the dataset page: https://huggingface.co/datasets/realzL/benchability-fig4-eval.
This repository belongs to realzL on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
