datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ssv2-3x3
SSV2 3x3
A multi-frame action dataset for teaching a model to look across several frames and describe the action it sees.
Each image is a 3x3 collage built from frames sampled from the original Something-Something V2 videos.
Splits
Split
Samples
train
150000
test
1000
Columns
image: 3x3 frame collage
label_text: action label text
annotation_text: action template text
video_id: original video identifier
ssv2_embeds
