somusan/ist-vqa-zip-full
IST-VQA — Multilingual Scene Text VQA Dataset Curated dataset for Indic Scene Text Visual Question Answering covering three languages: Bengali (bn) · Hindi (hi) · Tamil (ta) Summary Bengali Hindi Tamil Total Total Images 2,308 2,097 2,322 6,727 Total VQA Pairs 2,306 2,097 2,321 6,724 — Crawl images 1,615 1,570 1,394 4,579 — IndicSTR12 images 357 173 336 866 — Synthetic images 336 354 592 1,282 Sources: Real-world web-crawled scene photos… See the full description on the dataset page: https://huggingface.co/datasets/somusan/ist-vqa-zip-full.
Upload gen_cap_bench.tar.xz
Upload split_half_leak_fixed_v3.tar.xz
Upload ist_vqa_split_final_v2.tar.xz
Upload 2 files
Add files using upload-large-folder tool
initial commit
