datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.my-llava-recipesLlavaGuardWARNING: This repository contains content that might be disturbing! Therefore, we set the Not-For-All-Audiences tag.
License
The annotations provided in this dataset (e.g., labels, bounding boxes, coordinates) are released under the Apache License 2.0. The linked or referenced images are not included under this license and remain under their original licenses. Users are responsible for ensuring compliance with the terms of use of the respective image sources.
Content… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/LlavaGuard.llava-pretrain-refined-by-data-juicer
LLaVA pretrain -- LCS-558k (refined by Data-Juicer)
A refined version of LLaVA pretrain dataset (LCS-558k) by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Multimodal Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 115MB).
Dataset Information
Number of samples: 500,380 (Keep ~89.65% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/llava-pretrain-refined-by-data-juicer.
