irl-kit/SPARC-VQA-Raw
SPARC VQA Raw This repository contains the unfiltered SPARC VQA corpus: 838,211 embedded-image training examples in train.parquet (33.20 GB). Each example contains an image, question, answer, task metadata, source identifier, and annotation metadata including selected_start_score. Ready-to-train version For the exact processed SPARC subset used by the released Qwen3.5 models, download train_filtered_t097_mpo700.parquet from irl-kit/SPARC-VQA. It contains 284,909… See the full description on the dataset page: https://huggingface.co/datasets/irl-kit/SPARC-VQA-Raw.
Add released model mixture manifest
Document raw SPARC use and release filtering
Add reproducible SPARC release filtering script
Document raw SPARC VQA corpus and filtered release
Document SPARC VQA prompting and model mixtures
Upload SPARC VQA training split with embedded image column
Add dataset generation configuration
Add SPARC VQA dataset card
initial commit
