kfkas/streamingvlm-task-aware-vqa-suite-rowwise
StreamingVLM Task-Aware VQA Suite (Row-wise) Viewer-ready evaluation samples for task-aware visual-token sensitivity experiments. Every config contains 3,000 deterministic manifest-order samples with the image embedded in each row. Task groups task_group Dataset configs coarse_object_presence pope, repope, hpope general_scene_understanding vqav2, gqa fine_grained_visual_evidence gqa_attribute, mmbench textual_fine_grained_evidence textvqa, docvqa… See the full description on the dataset page: https://huggingface.co/datasets/kfkas/streamingvlm-task-aware-vqa-suite-rowwise.
StreamingVLM Task-Aware VQA Suite (Row-wise)
Viewer-ready evaluation samples for task-aware visual-token sensitivity experiments. Every config contains 3,000 deterministic manifest-order samples with the image embedded in each row.
Task groups
Columns
image: image payload rendered by the Hugging Face Viewerdataset_key: source dataset identifiertask_group: research task group used for analysiscategory: dataset-specific subtype such as POPEadversarialquestion,answer,all_answers: evaluation input and referencesmetadata_json: complete original manifest row
gqa_attribute is a derived attribute-focused subset rather than an official GQA split. MMBench remains a broad benchmark; use its category metadata when selecting fine-grained samples.
The complete manifests and deduplicated image archives remain available in streamingvlm-task-aware-vqa-suite-direct.
