CoolFace
Datasetpublic

irl-kit/SPARC-VQA-Raw

SPARC VQA Raw This repository contains the unfiltered SPARC VQA corpus: 838,211 embedded-image training examples in train.parquet (33.20 GB). Each example contains an image, question, answer, task metadata, source identifier, and annotation metadata including selected_start_score. Ready-to-train version For the exact processed SPARC subset used by the released Qwen3.5 models, download train_filtered_t097_mpo700.parquet from irl-kit/SPARC-VQA. It contains 284,909… See the full description on the dataset page: https://huggingface.co/datasets/irl-kit/SPARC-VQA-Raw.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes55downloads
9 commits on main
182e1941mo ago

Add released model mixture manifest

holgerson
251f04d1mo ago

Document raw SPARC use and release filtering

holgerson
0235a861mo ago

Add reproducible SPARC release filtering script

holgerson
d57fffa1mo ago

Document raw SPARC VQA corpus and filtered release

holgerson
07aa0c91mo ago

Document SPARC VQA prompting and model mixtures

holgerson
98ce9371mo ago

Upload SPARC VQA training split with embedded image column

holgerson
1ede8c31mo ago

Add dataset generation configuration

holgerson
81d60fa1mo ago

Add SPARC VQA dataset card

holgerson
245df891mo ago

initial commit

holgerson