datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ParseBench
ParseBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics:
Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on.
Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ParseBench.ExtractBench
ExtractBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated records), correctly use null for absent information, and ground each extracted… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ExtractBench.military-labeled-yolo
Military-Labeled YOLO Dataset (DVIDS sourced)
Multi-class military object detection in YOLO format. Source images pulled from
the Defense Visual Information Distribution Service (DVIDS)
public domain library; labeled via in-house Gemini-VLM-assisted pipeline with
human-in-the-loop correction.
Classes (12)
ID
Name
0
soldier
1
tank
2
apc
3
artillery
4
mlrs
5
military_truck
6
helicopter
7
aircraft
8
warship
9
missile_launcher
10
car… See the full description on the dataset page: https://huggingface.co/datasets/llama-farm/military-labeled-yolo.llama-3.2-1b-atlas
llama-3.2-1b-atlas
mmmuvdr-multilingual-train
Multilingual Visual Document Retrieval Dataset
This dataset consists of 500k multilingual query image samples, collected and generated from scratch using public internet pdfs. The queries are synthetic and generated using VLMs (gemini-1.5-pro and Qwen2-VL-72B).
It was used to train the vdr-2b-multi-v1 retrieval multimodal, multilingual embedding model.
How it was created
This is the entire data pipeline used to create the Italian subset of this dataset. Each step… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/vdr-multilingual-train.dolma-v1_7-305B-tokenized-llama2-nanosetLlama-Nemotron-VLM-Dataset-v1-OCR4llamaindex-vdr-en-train-preprocessed
llamaindex-vdr-en-train-preprocessed
This dataset is a preprocessed English subset of llamaindex/vdr-multilingual-train, prepared for training multimodal Sentence Transformer embedding models on document screenshot retrieval.
Changes from the original dataset
The original llamaindex/vdr-multilingual-train dataset stores hard negatives as a list of ID strings that reference other rows. This dataset makes two key changes:
English only: Only the English subset (53,512… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/llamaindex-vdr-en-train-preprocessed.visual-qa-llama-format
Open Paws Visual Qa Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Multimodal Data
Format: JSONL (JSON Lines)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/visual-qa-llama-format.preference_data_llama_factory_wo_checklist
Dataset Card for "preference_data_llama_factory_wo_checklist"
More Information needed
llama4-maverick-coco-captionsllama4-scout-coco-captionsllamaindex-vdr-images
LlamaIndex VDR Images
This dataset contains the document-page images used by lightonai/llamaindex-vdr-fine-tuning. The pair is a reformatted derivative of llamaindex/vdr-multilingual-train for multilingual retrieval fine-tuning.
Dataset structure
The train split contains:
Column
Type
Description
image_filename
string
Stable key used by the companion fine-tuning dataset.
image
image
Document-page image.
Load the dataset
from… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/llamaindex-vdr-images.llama3.2-3b-instruct-atlas
juiceb0xc0de/llama3.2-3b-instruct-atlas
A brain atlas for meta-llama/Llama-3.2-3B-Instruct, the 3B instruction-tuned member of the Llama 3.2 family. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know where an instruction-tuned model keeps its register machinery, which directions survive a causal test… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/llama3.2-3b-instruct-atlas.llama-3.2-3b-atlas
llama-3.2-3b-atlas
llama3.2-1b-instruct-atlas
juiceb0xc0de/llama3.2-1b-instruct-atlas
A brain atlas for meta-llama/Llama-3.2-1B-Instruct, the 1B instruction-tuned member of the Llama 3.2 family. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know where a small instruction-tuned model keeps its register machinery, which directions survive a causal… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/llama3.2-1b-instruct-atlas.xprmt-llama-3.1-8b-instruct-multijail-judge-evalllamav-o1-instruct-stage2
LlamaV-o1 Stage 2 Instruction Train set
This dataset is the same as the LLaVa-CoT train set but in multi-turn format. Here is a sample
{
"image": <image>,
"id": 10,
"conversations": [
{
"from": "human",
"value": "Question: Which country is highlighted?\nContext: N/A\nOptions: (A) Solomon Islands (B) Nauru (C) Vanuatu (D) Fiji\nSummarize how you will approach the problem and explain the steps you will take to reach the answer."
},
{
"from": "gpt"… See the full description on the dataset page: https://huggingface.co/datasets/ahmedheakl/llamav-o1-instruct-stage2.MMMU_Pro__oldpreference_data_llama_factory_corrected_format_text_onlypreference_data_llama_factory_corrected_formatllama-3.1-8b-instruct-atlas
llama-3.1-8b-instruct-atlas
vdr-multilingual-test
Multilingual Visual Document Retrieval Benchmarks
This dataset consists of 15 different benchmarks used to initially evaluate the vdr-2b-multi-v1 multimodal retrieval embedding model. These benchmarks allow the testing of multilingual, multimodal retrieval capabilities on text-only, visual-only and mixed page screenshots.
Each language subset contains queries and images in that language and is divided into three different categories by the "pagetype" column. Each category contains… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/vdr-multilingual-test.pokemon-gpt4o-captionsBorrowed from: https://huggingface.co/datasets/jugg1024/pokemon-gpt4o-captions
You can use it in LLaMA Factory by specifying dataset: pokemon_cap.
military-labeled-clip
Military-Labeled CLIP Crops (DVIDS sourced)
Per-object crops extracted from DVIDS military imagery, each accompanied by a
Gemini-VLM caption suitable for CLIP fine-tuning or zero-shot evaluation.
Files
crops/dvids_image_{id}_{class}_{idx}.jpg — 4,844 cropped objects
captions.jsonl — per-crop metadata: {crop_path, image_id, class_name, class_id, bbox, caption, image_caption, branch, source, source_url}
Classes (12) — Distribution
Class
Crops… See the full description on the dataset page: https://huggingface.co/datasets/llama-farm/military-labeled-clip.RLHF-VBorrowed from: https://huggingface.co/datasets/openbmb/RLHF-V-Dataset
You can use it in LLaMA Factory by specifying dataset: rlhf_v.
llamav-o1-instruct-stage1
LlamaV-o1 Stage 1 Train set
This the training set for the curriculum learning stage of LlamaV-o1. The dataset is curated from Geo170k (50k samples) and PixmoCapQA (30k samples). Here is a sample of the dataset:
{
"image": <image>,
"id": 20,
"conversations: [
{
"from": "human",
"value": "<image>If in the provided figure,CA and CB have the same length, AD and BD have the same length, M is the midpoint of CA, N is the midpoint of CB, and angle ADN measures 80… See the full description on the dataset page: https://huggingface.co/datasets/ahmedheakl/llamav-o1-instruct-stage1.plots_llama_3.1_8b_finetuned_binaryxprmt-llama-3.1-8b-instruct-multijail-v2
