CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KingTechnician /xd-violence-rgb-videomae-chunked-testtext100K<n<1M0 likes2.6k downloads8mo agoHugging Face02AndrewRqy /temporal-sae-videomae Temporal-SAE-VideoMAE — reproducibility artifacts Weights, cached activations, and the synthetic NFP ball dataset for the project Detecting Temporal Concepts in Video Transformers via Sparse Autoencoders (code: https://github.com/AndrewRqy/temporal-sae-videomae). These artifacts reproduce the full study — the trained SAEs plus the PCA / ICA / raw linear-decomposition baselines (monosemanticity score and the no-false-positives temporal-feature test) — without re-training… See the full description on the dataset page: https://huggingface.co/datasets/AndrewRqy/temporal-sae-videomae.video-classification0 likes981 downloads2mo agoHugging Face03KingTechnician /xd-violence-20pct-rgb-videomae-chunked-traintext100K<n<1M0 likes145 downloads8mo agoHugging Face04OpenGVLab /VideoMAEv2-TAL-Features1 likes53 downloads2y agoHugging Face05Jazzcharles /ego4d_videomae_L14_feature_fps8 📙 Overview Ego4d video features extracted by VideoMAE_L14 at 8 fps. It contains 9645 files, each file (e.g. fffbaeef-577f-45f0-baa9-f10cabf62dfb.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768. 🏋️ How-To-Use Please refer to code EgoInstructor for details. 🎓 Citation @article{xu2024retrieval, title={Retrieval-augmented egocentric video captioning}, author={Xu, Jilan and Huang, Yifei and Hou, Junlin and Chen, Guo… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/ego4d_videomae_L14_feature_fps8.textvideo-classification1K<n<10K0 likes42 downloads2y agoHugging Face06Jazzcharles /egolearn_videomae_internvideo_features 📙 Overview Egolearn video features. egocentric videos are extracted by VideoMAE_L14 at 8 fps. exocentric videos are extracted by InternVideo_MM_L14 at 8 fps. They are used for evaluating the egovideo-exovideo retrieval ability of EgoInstructor. Each file (e.g. 2cad4224-56c4-11ee-88ee-80615f12b59e.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768. 🏋️ How-To-Use Please refer to code EgoInstructor for details. 🎓… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/egolearn_videomae_internvideo_features.textvideo-classificationn<1K0 likes20 downloads2y agoHugging Face07KingTechnician /xd-violence-mini-rgb-videomae-chunkedtext10K<n<100K0 likes20 downloads10mo agoHugging Face08Jazzcharles /epic_kitchen_videomae_L14_feature_fps8 📙 Overview EPIC-Kitchen-100 video features extracted by VideoMAE_L14 at 8 fps. It is used for evaluating the video-text retrieval ability of EgoInstructor. It contains 700 files, each file (e.g. P01_01.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768. 🏋️ How-To-Use Please refer to code EgoInstructor for details. 🎓 Citation @article{xu2024retrieval, title={Retrieval-augmented egocentric video captioning}, author={Xu… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/epic_kitchen_videomae_L14_feature_fps8.textvideo-classificationn<1K0 likes15 downloads2y agoHugging Face09Jazzcharles /charadesego_videomae_L14_feature_fps8 📙 Overview CharadesEgo video features extracted by VideoMAE_L14 at 8 fps. It is used for evaluating the egovideo-exovideo retrieval ability of EgoInstructor. It contains 7860 files, each file (e.g. 005BUEGO.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768. 🏋️ How-To-Use Please refer to code EgoInstructor for details. 🎓 Citation @article{xu2024retrieval, title={Retrieval-augmented egocentric video captioning}… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/charadesego_videomae_L14_feature_fps8.textvideo-classification1K<n<10K0 likes14 downloads2y agoHugging Face10Darknsu /Thumos_VideoMAEv20 likes12 downloads1y agoHugging Face11rodrigo-paganini /ttc-videomaeb-viclip-representations-k4000 likes12 downloads5mo agoHugging Face12Kajhanan /VideoMAEv2_BlindSpot VideoMAEv2_BlindSpots This dataset contains a 'diverse' set of 25 videos from the misclassified samples, after logistic regression is performed for the feature vectors obtained through the 'VideoMAEv2-Huge' video feature extractor, for a subset of the 'Kinetics-400' dataset with 3995 samples and 395 classes. This subset is randomly split into training and testing sets with a 0.5 split ratio. Base Model: https://huggingface.co/OpenGVLab/VideoMAEv2-Huge Loading the Model… See the full description on the dataset page: https://huggingface.co/datasets/Kajhanan/VideoMAEv2_BlindSpot.textfeature-extractionn<1K0 likes4 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.