CoolFace
Datasetpublic

Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new

LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower) Derived data (vision-encoder features of video frames), not a redistribution of the source videos. Source: lmms-lab/LLaVA-Video-178K -- its card restricts use to academic research and education, and its annotations come from GPT-4-class models (see the OpenAI usage policy). Complete: 85000 clips. Subset Folders: 0_30_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_academic_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new.

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
5likes7.1kdownloads
Dataset Card

LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower)

Derived data (vision-encoder features of video frames), not a redistribution of the source videos. Source: `lmms-lab/LLaVA-Video-178K` -- its card restricts use to academic research and education, and its annotations come from GPT-4-class models (see the OpenAI usage policy).

Complete: 85000 clips.

Subset

  • —Folders: 0_30_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_academic_v0_1, 30_60_s_youtube_v0_1 (0-30 s and 30-60 s, academic + youtube).
  • —Duration filter: [10.0, 60.0] s, measured on the actual file.
  • —Target 85000 clips split across folders proportionally to folder population, with any folder's shortfall re-split across the others. Final per-folder quotas: 0_30_s_academic_v0_1: 8361, 0_30_s_youtube_v0_1: 55355, 30_60_s_academic_v0_1: 7335, 30_60_s_youtube_v0_1: 13949.
  • —Clips were taken in shard order (the source shards are internally shuffled); manifest.json lists every clip taken (processed), rejected by the duration filter (rejected) or unreadable (failed).

Processing

  • —32 frames per clip at np.linspace(0, n_frames - 1, 32, dtype=int) (LLaVA-OV's rule).
  • —Each frame: PIL bicubic resize straight to 384x384 (aspect ratio not preserved, as in LLaVA-OV's video path), rescale 1/255, normalize mean=std=0.5.
  • —Encoder: google/siglip-so400m-patch14-384 architecture with the fine-tuned vision tower of `lmms-lab/llava-onevision-qwen2-7b-ov`, frozen; final transformer layer deleted; features are hidden_states[-1] (post_layernorm bypassed); no multimodal projector. These are NOT the features LLaVA-OV feeds its LLM (those are post-projector).
  • —Spatial pooling: bilinear 27x27 -> 14x14 per frame on the raw 1152-dim features.
  • —Computed in float16 (TF32 disabled; pooling in float32), stored as float16.
  • —Weights sha256: 215fe866b7045eaf181befef66a755cddf791471d21a52e6fba06869c52f5d96.

Files

  • —data/NNNNN/<video_id>.pt: {"tokens": float16 [32, 196, 1152], "metadata": {...}}. metadata has row-aligned frame_indices / timestamps, native fps/size/duration, source folder/shard/path, and the full processing config.
  • —manifest.json: processed[video_id].path gives each clip's file.
  • —labels/captions.jsonl, labels/oe_qa.jsonl, labels/mc_qa.jsonl: the source repo's caption / open-ended QA / multiple-choice QA entries for the uploaded clips, unchanged apart from an added source_folder (join on id). Not every clip has MC QA.
python
import torch
from huggingface_hub import hf_hub_download
path = hf_hub_download("Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new", "data/00000/<video_id>.pt", repo_type="dataset")
item = torch.load(path)  # item["tokens"]: [32, 196, 1152] float16

Intended use

Stage 1 self-supervised pretraining targets for a video compression project.

Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new · CoolFace