datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xd-violence-rgb-videomae-chunked-testtemporal-sae-videomae
Temporal-SAE-VideoMAE — reproducibility artifacts
Weights, cached activations, and the synthetic NFP ball dataset for the project
Detecting Temporal Concepts in Video Transformers via Sparse Autoencoders
(code: https://github.com/AndrewRqy/temporal-sae-videomae).
These artifacts reproduce the full study — the trained SAEs plus the PCA / ICA / raw
linear-decomposition baselines (monosemanticity score and the no-false-positives temporal-feature
test) — without re-training… See the full description on the dataset page: https://huggingface.co/datasets/AndrewRqy/temporal-sae-videomae.xd-violence-20pct-rgb-videomae-chunked-trainVideoMAEv2-TAL-Featuresego4d_videomae_L14_feature_fps8
📙 Overview
Ego4d video features extracted by VideoMAE_L14 at 8 fps.
It contains 9645 files, each file (e.g. fffbaeef-577f-45f0-baa9-f10cabf62dfb.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓 Citation
@article{xu2024retrieval,
title={Retrieval-augmented egocentric video captioning},
author={Xu, Jilan and Huang, Yifei and Hou, Junlin and Chen, Guo… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/ego4d_videomae_L14_feature_fps8.egolearn_videomae_internvideo_features
📙 Overview
Egolearn video features.
egocentric videos are extracted by VideoMAE_L14 at 8 fps.
exocentric videos are extracted by InternVideo_MM_L14 at 8 fps.
They are used for evaluating the egovideo-exovideo retrieval ability of EgoInstructor.
Each file (e.g. 2cad4224-56c4-11ee-88ee-80615f12b59e.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/egolearn_videomae_internvideo_features.xd-violence-mini-rgb-videomae-chunkedepic_kitchen_videomae_L14_feature_fps8
📙 Overview
EPIC-Kitchen-100 video features extracted by VideoMAE_L14 at 8 fps. It is used for evaluating the video-text retrieval ability of EgoInstructor.
It contains 700 files, each file (e.g. P01_01.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓 Citation
@article{xu2024retrieval,
title={Retrieval-augmented egocentric video captioning},
author={Xu… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/epic_kitchen_videomae_L14_feature_fps8.charadesego_videomae_L14_feature_fps8
📙 Overview
CharadesEgo video features extracted by VideoMAE_L14 at 8 fps. It is used for evaluating the egovideo-exovideo retrieval ability of EgoInstructor.
It contains 7860 files, each file (e.g. 005BUEGO.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓 Citation
@article{xu2024retrieval,
title={Retrieval-augmented egocentric video captioning}… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/charadesego_videomae_L14_feature_fps8.Thumos_VideoMAEv2ttc-videomaeb-viclip-representations-k400VideoMAEv2_BlindSpot
VideoMAEv2_BlindSpots
This dataset contains a 'diverse' set of 25 videos from the misclassified samples, after logistic regression is performed for the feature vectors obtained through the 'VideoMAEv2-Huge' video feature extractor, for a subset of the 'Kinetics-400' dataset with 3995 samples and 395 classes. This subset is randomly split into training and testing sets with a 0.5 split ratio.
Base Model: https://huggingface.co/OpenGVLab/VideoMAEv2-Huge
Loading the Model… See the full description on the dataset page: https://huggingface.co/datasets/Kajhanan/VideoMAEv2_BlindSpot.
