activity
Datasets
All datasets matching “activity”ActivityNet
Description
Dataset V1-2
v1-2_train.tar.gz and v1-2_val.tar.gz
Data (train and val set) associated with ActivityNet release 1.2
v1-2_test.tar.gz
Data (test set only) associated with ActivityNet release 1.2
Dataset V1-3
v1-3_train_val.tar.gz
Additional videos (train val set) collected for ActivityNet release 1.3
v1-3 is an extension of v1-2, so you also need to download v1-2 data and merge to v1.3
v1-3_test.tar.gz
Additional videos (test set only)… See the full description on the dataset page: https://huggingface.co/datasets/YimuWang/ActivityNet.Collective-Activity-Recognition
Annotation Format
Every 10th frame in all video sequences was manually annotated with the following information for each detected person:
Bounding box location
Activity class
Pose direction
Annotation Fields
Each annotation follows the format:
<frame_number> <x> <y> <width> <height> <class_id> <pose_id>
Field
Description
frame_number
Frame identifier
x
X-coordinate of the bounding box (top-left corner)
y
Y-coordinate of the bounding box (top-left… See the full description on the dataset page: https://huggingface.co/datasets/litforth/Collective-Activity-Recognition.ActivityNetQAActivityNet_Captions
About
ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark.
We adopt the official split:
Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions)
Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.Activitynetactivitynet
ActivityNet v1.3
15,941 videos (391.3 GB) with 1,550 subtitle files, downloaded at source
quality and re-hosted for direct use — no more dead YouTube links, no more flaky
downloader scripts.
Coverage: 15,941 of the 19,994 source video IDs (79.7%). 4,053 source videos were unavailable at fetch time (private, removed, members-only or geo-blocked) and are excluded. The dataset is refreshed as more videos are delivered.
What's inside
metadata.jsonl — one row per… See the full description on the dataset page: https://huggingface.co/datasets/TornadoLabs/activitynet.
