datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ssv2-3x3
SSV2 3x3
A multi-frame action dataset for teaching a model to look across several frames and describe the action it sees.
Each image is a 3x3 collage built from frames sampled from the original Something-Something V2 videos.
Splits
Split
Samples
train
150000
test
1000
Columns
image: 3x3 frame collage
label_text: action label text
annotation_text: action template text
video_id: original video identifier
visual_cues_ssv2ssv2
SSv2
Full-set metadata for lmms-eval task ssv2.
Metadata source: https://huggingface.co/datasets/emirgocen/Something-Something-v2
Raw video location: official Something-Something-V2 .webm files (license-gated), not embedded here.
Expected local media root for evaluation: $SSV2_VIDEO_DIR or $HF_HOME/ssv2.
ssv2_VPssv2_directionk710-ssv2-videolistssv2_embeds
