videomae
Datasets
All datasets matching “videomae”xd-violence-rgb-videomae-chunked-testtemporal-sae-videomae
Temporal-SAE-VideoMAE — reproducibility artifacts
Weights, cached activations, and the synthetic NFP ball dataset for the project
Detecting Temporal Concepts in Video Transformers via Sparse Autoencoders
(code: https://github.com/AndrewRqy/temporal-sae-videomae).
These artifacts reproduce the full study — the trained SAEs plus the PCA / ICA / raw
linear-decomposition baselines (monosemanticity score and the no-false-positives temporal-feature
test) — without re-training… See the full description on the dataset page: https://huggingface.co/datasets/AndrewRqy/temporal-sae-videomae.xd-violence-20pct-rgb-videomae-chunked-trainVideoMAEv2-TAL-Featuresego4d_videomae_L14_feature_fps8
📙 Overview
Ego4d video features extracted by VideoMAE_L14 at 8 fps.
It contains 9645 files, each file (e.g. fffbaeef-577f-45f0-baa9-f10cabf62dfb.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓 Citation
@article{xu2024retrieval,
title={Retrieval-augmented egocentric video captioning},
author={Xu, Jilan and Huang, Yifei and Hou, Junlin and Chen, Guo… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/ego4d_videomae_L14_feature_fps8.egolearn_videomae_internvideo_features
📙 Overview
Egolearn video features.
egocentric videos are extracted by VideoMAE_L14 at 8 fps.
exocentric videos are extracted by InternVideo_MM_L14 at 8 fps.
They are used for evaluating the egovideo-exovideo retrieval ability of EgoInstructor.
Each file (e.g. 2cad4224-56c4-11ee-88ee-80615f12b59e.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/egolearn_videomae_internvideo_features.
