datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vigorl_datasets
ViGoRL Datasets
This repository contains the official datasets associated with the paper "Grounded Reinforcement Learning for Visual Reasoning (ViGoRL)", by Gabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain, Michael J. Tarr, Aviral Kumar, and Katerina Fragkiadaki.
Dataset Overview
These datasets are designed for training and evaluating visually grounded vision-language models (VLMs).
Datasets are organized by the visual reasoning tasks described in the ViGoRL… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/vigorl_datasets.countqa_lite
gsarch/countqa_lite
A deterministic lite evaluation subset of Jayant-Sravan/CountQA.
Source revision: f92cc6fe46542c61e2916e3d2ae9a911e2216b1a
Source split: test
Sampling seed: 43
Output rows: 500
Schema: unchanged from the upstream dataset
CountQA is sampled at the QA-pair level. Each output row retains the original schema and contains one-element questions and answers lists, so lmms-eval's existing countqa_process_docs produces exactly 500 prompts.
Generated by… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/countqa_lite.ScreenSpot-Pro-Lite
ScreenSpot-Pro-Lite
A fixed 500-example representative/challenging subset of the 1,581-example
ScreenSpot-Pro benchmark.
Selection
Sampling is proportional over the cross-product of platform, application, and UI type,
so all 26 applications remain represented and the original icon/text mix is retained.
Within every stratum, 80% is deterministic seeded sampling and 20% is a hard tier.
Hardness combines failure rate and disagreement across five full-run anchor… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/ScreenSpot-Pro-Lite.HStarBench-Lite
HStarBench-Lite
A fixed 500-example representative/challenging subset of the 4,000-example
HStarBench mixed test set,
packaged for the Vero/lmms-eval panorama-strip evaluation.
Selection
The subset preserves the full benchmark's HOS/HPS and difficulty-level proportions:
Split
Level
Count
HOS
0
77
HOS
1
22
HOS
2
201
HPS
0
63
HPS
1
57
HPS
2
53
HPS
3
27
Within every split/level stratum, 80% is deterministic seeded sampling and 20% is a… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/HStarBench-Lite.gsat-115Machine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/gsat-115.SimpleVQA-ENGSARDDataset Curators
The dataset was created by Antoine Louis during work done at the Law & Tech lab of Maastricht University, with the help of jurists from Droits Quotidiens.
Licensing Information
BSARD is licensed under the CC BY-NC-SA 4.0 license.
https://creativecommons.org/licenses/by-nc-sa/4.0/
EvoChart-QA
gsarch/EvoChart-QA
This dataset packages the EvoChart QA annotations with embedded chart images.
Each row contains the keys image, question, answer, attribute, is_clear, and chart_type.
Total rows: 1250
Game-QA-LiteSimpleVQA-CNchartqa_lite
gsarch/chartqa_lite
A deterministic lite evaluation subset of lmms-lab/ChartQA.
Source revision: 9e63b7df1592a1c2158e735cc1725454aef0d6d9
Source split: test
Sampling seed: 44
Output rows: 500
Schema: unchanged from the upstream dataset
Generated by scripts/create_lite_eval_datasets.py in the gaze-vlm vero_eval repository.
countbenchqa_lite
gsarch/countbenchqa_lite
A deterministic lite evaluation subset of vikhyatk/CountBenchQA.
Source revision: 76d600309e9d6147bd3713b4cd431518ac1206c8
Source split: test
Sampling seed: 42
Output rows: 500
Schema: unchanged from the upstream dataset
The upstream split has only 491 rows. This subset contains every upstream row once plus 9 deterministic duplicate draws to maintain the 500-example evaluation contract.
Generated by scripts/create_lite_eval_datasets.py in the gaze-vlm… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/countbenchqa_lite.ChartMuseum
gsarch/ChartMuseum
This dataset includes images and annotations with keys:
image, question, answer, reasoning_type, source, hash.
Splits
test: 1000 rows
dev: 162 rows
Images are embedded via the datasets.Image feature, so they are available
directly when loading the dataset with datasets.load_dataset("gsarch/ChartMuseum").
Game-QAheld_out_gazegsat-zhtw
