datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
documentation-imageszen-imageexample-imageszen-multi-imagellava-instruct-mix
LLaVA Instruct Mix
Summary
The LLaVA Instruct Mix dataset is a processed version of LLaVA Instruct Mix.
Data Structure
Format: Conversational
Type: Language-modeling
Columns:
"images": The image associated with the text.
"prompt": A list of messages that form the context for the conversation.
"completion": The last message in the conversation, which is the model's response.
This structure allows models to learn from the context of the conversation… See the full description on the dataset page: https://huggingface.co/datasets/trl-lib/llava-instruct-mix.paper8-trace-2026-05-27
paper8 — Claude Code session trace (2026-05-26 → 2026-05-27)
Claude Code session transcript captured while working in
nationallab-bench/papers/paper8.
Contents
9fc72deb-8db0-4def-960d-7b653fa60711.jsonl — full session transcript
(JSONL, one event per line: user/assistant turns, tool calls, results).
9fc72deb-8db0-4def-960d-7b653fa60711/tool-results/ — image attachments
extracted from PDFs during the session (per-page JPGs).
rlaif-v
RLAIF-V Dataset
Summary
The RLAIF-V dataset is a processed version of the openbmb/RLAIF-V-Dataset, specifically curated to train vision-language models using the TRL library for preference learning tasks. It contains 83,132 high-quality comparison pairs, each comprising an image and two textual descriptions: one preferred and one rejected. This dataset enables models to learn human preferences in visual contexts, enhancing their ability to generate and evaluate image… See the full description on the dataset page: https://huggingface.co/datasets/trl-lib/rlaif-v.PAD3-Dataset-Revisi-Fixed-TRL-Messages-AugPAD3-Dataset-Revisi-Fixed-TRL-Messagestrl-ocr-datasetflickr30k-qwen3-vl-2b-sft-trl-with-judgedoab-metadata-extraction-trl
DOAB Open Access Books - Metadata Extraction Dataset
Dataset Description
This dataset contains 9,363 open access books with page images and rich bibliographic metadata extracted from MARC21 records, curated specifically for training and evaluating Vision Language Models (VLMs) on automatic metadata extraction from scholarly monographs.
The dataset is derived from the Penn State ScholarSphere DOAB collection (Directory of Open Access Books), focusing on books with Creative… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/doab-metadata-extraction-trl.PAD3-Dataset-Revisi-Fixed-TRL
PAD3-Dataset-Revisi-Fixed (TRL chat format)
Conversational (TRL / SFT) dataset for age-rating classification of images.
Structure
.
├── metadata.jsonl # one TRL chat record per line
└── images/
├── Semua_Umur/000000.jpg
├── 7_/000000.jpg
├── 13_/000000.jpg
├── 15_/000000.jpg
├── 18_/000000.jpg
└── Konten_Terlarang/000000.jpg
Images are split into per-rating subfolders to stay under the 10,000-files-per-folder limit.… See the full description on the dataset page: https://huggingface.co/datasets/capstone-pad3/PAD3-Dataset-Revisi-Fixed-TRL.trl-gui-datasetPAD-3-Revisi-Balanced-TRLflickr30k-Qwen3-VL-2B-Instruct-trl-sft
