datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Leopard-Instruct
Leopard-Instruct
Paper | Github | Models-LLaVA | Models-Idefics2
Summaries
Leopard-Instruct is a large instruction-tuning dataset, comprising 925K instances, with 739K specifically designed for text-rich, multiimage scenarios. It's been used to train Leopard-LLaVA [checkpoint] and Leopard-Idefics2 [checkpoint].
Loading dataset
to load the dataset without automatically downloading and process the images (Please run the following codes with datasets==2.18.0)… See the full description on the dataset page: https://huggingface.co/datasets/wyu1/Leopard-Instruct.waqfeya-library
Waqfeya Library
📖 Overview
Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.LLaVA-OneVision-2-Data
LLaVA-OneVision-2-Data
Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training.
At a Glance
The dataset is split across two Hugging Face repositories because of its size:
Repository
What it contains
Part 1 (this repository)
~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.hamburg_curricula_2024lasa1m-annotate-part-14lasa1m-annotate-part-06OmniEdit-Filtered-1.2M
OmniEdit
In this paper, we present OMNI-EDIT, which is an omnipotent editor to handle seven different image editing tasks with any aspect ratio seamlessly. Our contribution is in four folds: (1) OMNI-EDIT is trained by utilizing the supervision
from seven different specialist models to ensure task coverage. (2) we utilize importance sampling based on the scores provided by large multimodal models (like GPT-4o) instead of CLIP-score to improve the data quality.
📃Paper | 🌐Website |… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/OmniEdit-Filtered-1.2M.LLaVA-OneVision-1.5-Instruct-Data
LLaVA-OneVision-1.5 Instruction Data
Paper | Code
📌 Introduction
This dataset, LLaVA-OneVision-1.5-Instruct, was collected and integrated during the development of LLaVA-OneVision-1.5. LLaVA-OneVision-1.5 is a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. This meticulously curated 22M instruction dataset (LLaVA-OneVision-1.5-Instruct) is part of a… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Instruct-Data.TABLET-Large
TABLET-Large
This is the Large sized train set of the TABLET dataset. It contains all train examples for all TABLET tasks, resulting in a total of 3,505,311 training examples across 17 tasks.This dataset is self-contained, each example includes a table image, its HTML representation, and the associated task data.However, if you're interested in downloading just the TABLET tables, check out TABLET-tables.
All TABLET Subsets:
(train) TABLET-Small: The smallest TABLET subset… See the full description on the dataset page: https://huggingface.co/datasets/alonsoapp/TABLET-Large.ai4g-flood-dataset
Flood Detection Dataset
Introduction
This dataset accompanies the paper Mapping global floods with 10 years of satellite radar data (Nature Communications, 2025) and contains global flood detections derived from Sentinel-1 Synthetic Aperture Radar (SAR) imagery using a deep learning change detection model. The dataset spans October 2014 – September 2024, offering a longitudinal view of flood-prone areas worldwide.
Key features:
Cloud-penetrating SAR data for consistent… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/ai4g-flood-dataset.textvqa
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of TextVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{singh2019towards,
title={Towards vqa models that can read},
author={Singh, Amanpreet and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/textvqa.documentation-imagesliberoThis dataset was created using LeRobot.
Dataset Description
This dataset combines four individual Libero datasets: Libero-Spatial, Libero-Object, Libero-Goal and Libero-10.
All datasets were taken from here and converted into LeRobot format.
Homepage: https://libero-project.github.io
Paper: https://arxiv.org/abs/2306.03310
License: CC-BY 4.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 1693… See the full description on the dataset page: https://huggingface.co/datasets/physical-intelligence/libero.GQA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of GQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{hudson2019gqa,
title={Gqa: A new dataset for real-world visual reasoning and compositional… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/GQA.DocVQA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of DocVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@article{mathew2020docvqa,
title={DocVQA: A Dataset for VQA on Document Images. CoRR abs/2007.00398 (2020)}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/DocVQA.VQAv2MMMUThis is a merged version of MMMU/MMMU with all subsets concatenated.
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of MMMU. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@article{yue2023mmmu,
title={Mmmu: A… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/MMMU.POPE
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of POPE. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@article{li2023evaluating,
title={Evaluating object hallucination in large vision-language models}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/POPE.imageslab-bench
LAB-Bench
The Language Agent Biology Benchmark, or LAB-Bench, is an evaluation dataset for AI systems intended to benchmark capabilities foundational to scientific research in biology. The dataset currently consists of 8 broad categories, comprising 30 narrower subtasks, including extracting information from the scientific literature (LitQA2), retrieving information from databases (DbQA) and supplementary information (SuppQA), reasoning about scientific figures (FigQA) and tables… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/lab-bench.PhysicalAI-Robotics-Locomanipulation-GRAIL
📢 News
[2026-07-15] Released task-general tracking policy checkpoints trained on the released data. Follow the tracking doc to use them to track our released motion data.
[2026-07-14] Updated data/pickup_table and data/pickup_ground. If you downloaded them before this date, please re-download.
Dataset Overview
Tabletop Pickup
Ground Pickup
Tabletop Manipulation
Ground Manipulation
Sitting
Curb
Slope… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Locomanipulation-GRAIL.LLaVA-OneVision-Data
Dataset Card for LLaVA-OneVision
[2024-09-01]: Uploaded VisualWebInstruct(filtered), it's used in OneVision Stage
almost all subsets are uploaded with HF's required format and you can use the recommended interface to download them and follow our code below to convert them.
the subset of ureader_kg and ureader_qa are uploaded with the processed jsons and tar.gz of image folders.
You may directly download them from the following url.… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Data.MME
Evaluation Dataset for MME
ParseBench
ParseBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics:
Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on.
Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ParseBench.models-logoEgoIT-99KCheckout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information.
OpenPathNet
OpenPathNet Dataset
This README describes the OpenPathNet dataset (the release referred to as Link 1 in the OpenPathNet project documentation). The dataset is generated by the OpenPathNet toolchain from real-world Miami and Boston urban areas based on OpenStreetMap (OSM), and then simulated with NVIDIA Sionna ray tracing for RF multipath propagation / channel modeling research and AI tasks.
The dataset is also carefully cleaned to ensure good building coverage in every scene.… See the full description on the dataset page: https://huggingface.co/datasets/liu-lz/OpenPathNet.ExtractBench
ExtractBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated records), correctly use null for absent information, and ground each extracted… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ExtractBench.SEED-Bench
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of SEED-Bench. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@article{li2023seed,
title={Seed-bench: Benchmarking multimodal llms with generative comprehension}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/SEED-Bench.MMBench
