CoolFace
Datasetpublic

Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new

LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower) Derived data (vision-encoder features of video frames), not a redistribution of the source videos. Source: lmms-lab/LLaVA-Video-178K -- its card restricts use to academic research and education, and its annotations come from GPT-4-class models (see the OpenAI usage policy). Complete: 85000 clips. Subset Folders: 0_30_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_academic_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new.

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
5likes7.1kdownloads

Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new · main · files are served by the source, never re-hosted here

Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new · CoolFace