atad-tokyo/GST_EGOSCHEMA
Model Card for Model ID https://showlab.github.io/videollm-online/ Model Details LLM: meta-llama/Meta-Llama-3-8B-Instruct Vision Strategy: Frame Encoder: google/siglip-large-patch16-384 Frame Tokens: CLS Token + Avg Pooled 3x3 Tokens Frame FPS: 2 for training, 2~10 for inference Frame Resolution: max resolution 384, with zero-padding to keep aspect ratio Video Length: 10 minutes Training Data: Ego4D Narration Stream 113K + Ego4D GoalStep Stream 21K… See the full description on the dataset page: https://huggingface.co/datasets/atad-tokyo/GST_EGOSCHEMA.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face