datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vsr_zeroshot
VSR: Visual Spatial Reasoning
This is the zero-shot set of VSR: Visual Spatial Reasoning (TACL 2023) [paper].
Usage
from datasets import load_dataset
data_files = {"train": "train.jsonl", "dev": "dev.jsonl", "test": "test.jsonl"}
dataset = load_dataset("cambridgeltl/vsr_zeroshot", data_files=data_files)
Note that the image files still need to be downloaded separately. See data/ for details.
Go to our github repo for more introductions.
Citation
If you find… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/vsr_zeroshot.zeroshot-NERzero_shot_classification_testprepared_data_file_zero_shot_prompting_5000Q_evidence_selected_plus_verbalizationprepared_data_file_zero_shot_prompting_5000Q_evidence_selected_plus_verbalization
zero_shot_pubMedecva_zeroshot_thinking
Qwen/Qwen3-VL-2B-Thinking · happy8825/valid_ecva_clean results
Model: Qwen/Qwen3-VL-2B-Thinking
Dataset: happy8825/valid_ecva_clean
Generated: 2025-12-16 01:28:46Z
Metrics
Metric
Value
Total samples
924
With GT
0
Parsed answers
0
Top-1 accuracy
0
Recall@5
0
MRR
0
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.
EVQA/ECVA Metrics
Metric
Value
EVQA total
924
EVQA with GT… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/ecva_zeroshot_thinking.zero-shot-classification-large-testgsm8k_math_ground_truth_zero_shotzero-shot-classification-sample_dev_e2d2_lm_eval_gsm8k_zeroshot_cot_samplescontext_reliance_resist_correction_zeroshotzeroshot-previewSFT_zeroshot_Gemma27B_balanced_plainunderstanding_fables_zero_shotbinary_zero_shotaaai-pathvqa-zeroshot-full-resultsprepared_data_file_zero_shot_prompting_100Q_4_experiments.jsonDans-Prosemaxx-InstructWriter-ZeroShot-2Dans-Prosemaxx-InstructWriter-ZeroShot-3zero-shotspide_zeroshotword_unscrambling_zero_shot_sampledDans-Prosemaxx-InstructWriter-ZeroShot
