Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new
LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower) Derived data (vision-encoder features of video frames), not a redistribution of the source videos. Source: lmms-lab/LLaVA-Video-178K -- its card restricts use to academic research and education, and its annotations come from GPT-4-class models (see the OpenAI usage policy). Complete: 85000 clips. Subset Folders: 0_30_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_academic_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new.
LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower)
Derived data (vision-encoder features of video frames), not a redistribution of the source videos. Source: `lmms-lab/LLaVA-Video-178K` -- its card restricts use to academic research and education, and its annotations come from GPT-4-class models (see the OpenAI usage policy).
Complete: 85000 clips.
Subset
- Folders:
0_30_s_academic_v0_1,0_30_s_youtube_v0_1,30_60_s_academic_v0_1,30_60_s_youtube_v0_1(0-30 s and 30-60 s, academic + youtube). - Duration filter: [10.0, 60.0] s, measured on the actual file.
- Target 85000 clips split across folders proportionally to folder population, with any folder's shortfall re-split across the others. Final per-folder quotas:
0_30_s_academic_v0_1: 8361,0_30_s_youtube_v0_1: 55355,30_60_s_academic_v0_1: 7335,30_60_s_youtube_v0_1: 13949. - Clips were taken in shard order (the source shards are internally shuffled);
manifest.jsonlists every clip taken (processed), rejected by the duration filter (rejected) or unreadable (failed).
Processing
- 32 frames per clip at
np.linspace(0, n_frames - 1, 32, dtype=int)(LLaVA-OV's rule). - Each frame: PIL bicubic resize straight to 384x384 (aspect ratio not preserved, as in LLaVA-OV's video path), rescale 1/255, normalize mean=std=0.5.
- Encoder:
google/siglip-so400m-patch14-384architecture with the fine-tuned vision tower of `lmms-lab/llava-onevision-qwen2-7b-ov`, frozen; final transformer layer deleted; features arehidden_states[-1](post_layernorm bypassed); no multimodal projector. These are NOT the features LLaVA-OV feeds its LLM (those are post-projector). - Spatial pooling: bilinear 27x27 -> 14x14 per frame on the raw 1152-dim features.
- Computed in float16 (TF32 disabled; pooling in float32), stored as float16.
- Weights sha256:
215fe866b7045eaf181befef66a755cddf791471d21a52e6fba06869c52f5d96.
Files
data/NNNNN/<video_id>.pt:{"tokens": float16 [32, 196, 1152], "metadata": {...}}.metadatahas row-alignedframe_indices/timestamps, native fps/size/duration, source folder/shard/path, and the full processing config.manifest.json:processed[video_id].pathgives each clip's file.labels/captions.jsonl,labels/oe_qa.jsonl,labels/mc_qa.jsonl: the source repo's caption / open-ended QA / multiple-choice QA entries for the uploaded clips, unchanged apart from an addedsource_folder(join onid). Not every clip has MC QA.
import torch
from huggingface_hub import hf_hub_download
path = hf_hub_download("Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new", "data/00000/<video_id>.pt", repo_type="dataset")
item = torch.load(path) # item["tokens"]: [32, 196, 1152] float16Intended use
Stage 1 self-supervised pretraining targets for a video compression project.
