datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eval2_proaugeval3_TOY_video_images
Eval3 TOY Video Images
Image-question-answer grounding dataset extracted from the Eval 3 TOY celebrity permutation videos.
Each episode contributes three frames from the first 6 seconds:
start frame
middle frame
end frame
The labels use the user-provided episode block ground truth:
episodes 0-29: Taylor Swift
episodes 30-59: Barack Obama
episodes 60-89: Yann LeCun
within each 30-episode block: first 10 left, next 10 middle, last 10 right
Files:
all.jsonl: all 270 examples… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval3_TOY_video_images.
