datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llava-1.5-665k-instructionsThis dataset repository, LLaVA-1.5-665K-Instructions, is notably utilized in the paper Zero-Shot Vision Encoder Grafting via LLM Surrogates.
The official code repository for the paper can be found here: https://github.com/kaiyuyue/zero
LLaVA-1.5-665K-Instructions
This dataset repo contains the entire LLaVA-1.5-665K-Instructions in one place, including images and text sequences.
The images are in train_split/*.tars and the text sequences are in jsons:
llava_v1_5_mix665k.json is the… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/llava-1.5-665k-instructions.tiny_llavavideoTinyLLaVA-Video
This dataset combines data from multiple sources for pre-training and fine-tuning.
Pretrain Data: Four subsets of LLaVA-Video-178K (0_30_s_academic_v0_1, 30_60_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_youtube_v0_1), supplemented with filtered Video-LLaVA data (https://huggingface.co/datasets/LanguageBind/Video-LLaVA) and data from Valley (https://github.com/RupertLuo/Valley). The video data can be downloaded from the linked datasets, and cleaned annotations are provided… See the full description on the dataset page: https://huggingface.co/datasets/pbwpbw/tiny_llavavideo.LLaVA-OneVision-Mid-Data
Dataset Card for LLaVA-OneVision
Due to unknow reasons, we are unable to process dataset with large amount into required HF format. So we directly upload the json files and image folders (compressed into tar.gz files).
You can use the following link to directly download and decompress them.
https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data/tree/main/evol_instruct
We provide the whole details of LLaVA-OneVision Dataset. In this dataset, we include the data splits… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data.SAM-LLaVA-Captions10MLLaVA-OneVision-Mid-Data-512LLaVA-OneVision-Data-512SynTab-LLaVA-Datasetllava-odLLaVA_data/LLaVA_data/
│ finetune/
│ ├── llava_v1_5_mix665k.json
│ └── data/
│ ├── coco/
│ ├── gqa/
│ ├── nohup.out
│ ├── prompts/
│ ├── textvqa/
│ ├── coco2014_val_gpt4_qa_30x3.json
│ ├── coco2014_val_qa_eval/
│ ├── eval/
│ ├── LLaVA-Pretrain/
│ ├── occ_vqa/
│ ├── temp_eval/
│ └── vg/
└── pretrain/
└── blip_laion_cc_sbu_558k.json
└── images/
llava_video_subsetllava_mix665llava_video_max_256_frame_fps1llavaguard-qwen3depth-maps-llavallava_med_for_cv805llava_video
