realzL/benchability-fig4-eval
BenchAbility Figure 4 -- held-out eval The 10% candidate-dev half of the same 884,143-row pool the two training mixtures are drawn from, balanced per capability. Split by image, so no picture here appears in either mixture, and both arms are equally blind to it. capability n chart_reading / chart_reasoning 250 / 250 table_lookup / table_reasoning 250 / 250 document_qa / document_text_reading 250 / 250 diagram_and_infographic_understanding 250… See the full description on the dataset page: https://huggingface.co/datasets/realzL/benchability-fig4-eval.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face