CoolFace
20 results

videomae

KingTechnician /xd-violence-rgb-videomae-chunked-testtext100K<n<1M0 likes2.6k downloads8mo agoHugging FaceAndrewRqy /temporal-sae-videomae Temporal-SAE-VideoMAE — reproducibility artifacts Weights, cached activations, and the synthetic NFP ball dataset for the project Detecting Temporal Concepts in Video Transformers via Sparse Autoencoders (code: https://github.com/AndrewRqy/temporal-sae-videomae). These artifacts reproduce the full study — the trained SAEs plus the PCA / ICA / raw linear-decomposition baselines (monosemanticity score and the no-false-positives temporal-feature test) — without re-training… See the full description on the dataset page: https://huggingface.co/datasets/AndrewRqy/temporal-sae-videomae.video-classification0 likes981 downloads1mo agoHugging FaceKingTechnician /xd-violence-20pct-rgb-videomae-chunked-traintext100K<n<1M0 likes145 downloads8mo agoHugging FaceOpenGVLab /VideoMAEv2-TAL-Features1 likes53 downloads2y agoHugging FaceJazzcharles /ego4d_videomae_L14_feature_fps8 📙 Overview Ego4d video features extracted by VideoMAE_L14 at 8 fps. It contains 9645 files, each file (e.g. fffbaeef-577f-45f0-baa9-f10cabf62dfb.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768. 🏋️ How-To-Use Please refer to code EgoInstructor for details. 🎓 Citation @article{xu2024retrieval, title={Retrieval-augmented egocentric video captioning}, author={Xu, Jilan and Huang, Yifei and Hou, Junlin and Chen, Guo… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/ego4d_videomae_L14_feature_fps8.textvideo-classification1K<n<10K0 likes42 downloads2y agoHugging FaceJazzcharles /egolearn_videomae_internvideo_features 📙 Overview Egolearn video features. egocentric videos are extracted by VideoMAE_L14 at 8 fps. exocentric videos are extracted by InternVideo_MM_L14 at 8 fps. They are used for evaluating the egovideo-exovideo retrieval ability of EgoInstructor. Each file (e.g. 2cad4224-56c4-11ee-88ee-80615f12b59e.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768. 🏋️ How-To-Use Please refer to code EgoInstructor for details. 🎓… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/egolearn_videomae_internvideo_features.textvideo-classificationn<1K0 likes20 downloads2y agoHugging Face