datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLaVA-Video-178K
Dataset Card for LLaVA-Video-178K
Uses
This dataset is used for the training of the LLaVA-Video model. We only allow the use of this dataset for academic research and education purpose. For OpenAI GPT-4 generated data, we recommend the users to check the OpenAI Usage Policy.
Data Sources
For the training of LLaVA-Video, we utilized video-language data from five primary sources:
LLaVA-Video-178K: This dataset includes 178,510 caption entries, 960,792 open-ended… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K.llava-video-178k-siglip-tokens-ftov-new
LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower)
Derived data (vision-encoder features of video frames), not a
redistribution of the source videos. Source:
lmms-lab/LLaVA-Video-178K -- its card
restricts use to academic research and education, and its annotations come
from GPT-4-class models (see the OpenAI usage policy).
Complete: 85000 clips.
Subset
Folders: 0_30_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_academic_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new.Video-LLaVAVietnamese-lmms-lab-LLaVA-Video-178K-gg-translated
Dataset Card for 5CD-AI/Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated
This translated dataset includes:
LLaVA-Video-178K: 178,509 caption entries, 960,791 open-ended QA (question and answer) items, and 196,198 multiple-choice QA items.
The video source of the original dataset is in this repo: lmms-lab/LLaVA-Video-178K
LLaVA-Video-large-swift
Dataset Card LLaVA-Video-medium-swift
A subset of LLaVA-Video-178K for educational purposes to learn how to fine-tune video models.
LLaVA-Video-small-swift
Dataset Card LLaVA-Video-small-swift
Small subset of LLaVA-Video-178K for educational purposes to learn how to fine-tune video models.
VideoLLaVA_datasettl-llava-video-scratch
TL LLAVA VIDEO
llava-video-178kLLaVA-Video-medium-swiftVideoLLava
MVBench
We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate systematic generation of video tasks necessitating a wide range of temporal abilities, from perception to cognition. Guided by task definitions, we then automatically transform public video annotations into multiple-choice QA for task evaluation. This unique paradigm enables efficient creation of MVBench with minimal manual… See the full description on the dataset page: https://huggingface.co/datasets/Mitzi4132/VideoLLava.llava-video-jsonllava-video-text-dataset
eagle0504/llava-video-text-dataset
This is a tiny LLaVA dataset with exactly four video samples for training.
Field video_url: Video URLs (MP4/GIF format)
Field conversation: LLaVA conversation format with user/assistant roles
Field num_frames: Number of frames per video (5)
Dataset Structure
Each sample contains a conversation in LLaVA format:
{
"video_url": "https://example.com/video.mp4",
"conversation": [
{
"role": "user",
"content": [… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/llava-video-text-dataset.Video-R1-LLaVA-Video-83k-woopVideo-LLaVA-jsonVideoLLaVA_Trainllava_video_subsetLLaVA-Video-178K-subsetvideo-llava-processedvideollava-7b_tm05_eval_clipllava_video_max_256_frame_fps1LLaVA-Video-2_3_m_youtube_mc-qwen_filter_1video-r1-llava-mc-v1 ---
language:
- en
license: apache-2.0
size_categories:
- 10K<n<100K
task_categories:
- video-text-to-text
tags:
- multimodal-rl
- qwen3-vl
- gspo
- grpo
---
# ngqtrung/video-r1-llava-mc-v1
Curated v1 dataset for multimodal RL fine-tuning of Qwen3-VL-4B-Instruct.
| Property | Value |
|---|---|
| Rows | 72421 |
| Modality | video |
| Split | train |
| Schema | verl-ready (prompt + images + videos +… See the full description on the dataset page: https://huggingface.co/datasets/ngqtrung/video-r1-llava-mc-v1.video-r1-llava-free-v1 ---
language:
- en
license: apache-2.0
size_categories:
- 1K<n<10K
task_categories:
- video-text-to-text
tags:
- multimodal-rl
- qwen3-vl
- gspo
- grpo
---
# ngqtrung/video-r1-llava-free-v1
Curated v1 dataset for multimodal RL fine-tuning of Qwen3-VL-4B-Instruct.
| Property | Value |
|---|---|
| Rows | 9860 |
| Modality | video |
| Split | train |
| Schema | verl-ready (prompt + images + videos +… See the full description on the dataset page: https://huggingface.co/datasets/ngqtrung/video-r1-llava-free-v1.videollava-7b_clip_evalllava_videoVideo_Llavallava-video-178k-framesvideollava_10p_subsetvideollava_fx
