datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GQA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of GQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{hudson2019gqa,
title={Gqa: A new dataset for real-world visual reasoning and compositional question… See the full description on the dataset page: https://huggingface.co/datasets/pppop7/GQA.planeperturbed
Dataset Card for "planeperturbed"
More Information needed
OCR-VQA
Dataset Card for "OCR-VQA"
More Information needed
textvqa
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of TextVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{singh2019towards,
title={Towards vqa models that can read},
author={Singh, Amanpreet and Natarajan… See the full description on the dataset page: https://huggingface.co/datasets/pppop7/textvqa.artaVisualGenome_VG_100K_1_and_2
DiT-STcoco_train_2017LLaVA-Pretrain
LLaVA-Pretrain Dataset
Pretraining data for LLaVA (Large Language and Vision Assistant).
Description
This dataset contains the pretraining data used in LLaVA training, including:
blip_laion_cc_sbu_558k.json - Annotation file with 558K image-caption pairs
images/ - Corresponding images
Usage
from huggingface_hub import snapshot_download
# Download the dataset
snapshot_download(
repo_id="pppop7/LLaVA-Pretrain",
repo_type="dataset"… See the full description on the dataset page: https://huggingface.co/datasets/pppop7/LLaVA-Pretrain.eeefrau-federkiel-pins
