datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemini_public_mmr1
PRISM Public SFT Data
Overview
PRISM Public SFT Data is the public supervised fine-tuning data collection used in the PRISM project.PRISM studies the distributional drift problem in the standard SFT → RLVR post-training pipeline for large multimodal models. Before the distribution alignment and RLVR stages, we first use large-scale public multimodal demonstrations to obtain a broad SFT initialization.
This dataset serves as the public SFT data source for the… See the full description on the dataset page: https://huggingface.co/datasets/prism-vlm/gemini_public_mmr1.Arabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.VLM-ExecRouterBench
VLM-ExecRouterBench
An execution-oriented benchmark for cost-aware open-set VLM routing.
Cost-aware routing |
Open-set model onboarding |
Multimodal, code, and search tasks
Overview
VLM-ExecRouterBench is an execution-oriented benchmark for routing
vision-language model queries to a pool of candidate VLMs. Each sample is
executed by multiple candidate models, producing correctness labels, inference
costs, metadata… See the full description on the dataset page: https://huggingface.co/datasets/Kirito-Lab/VLM-ExecRouterBench.rl_dataset
PRISM RL Dataset
Overview
PRISM RL Dataset contains the training data used for the PRISM alignment and RLVR stages.
PRISM studies the distributional drift problem in the standard SFT → RLVR post-training pipeline for large multimodal models. Instead of directly applying RLVR after SFT, PRISM inserts an intermediate Distribution Alignment / Pre-alignment stage based on black-box on-policy distillation.
The overall pipeline is:
SFT → PRISM Alignment → RLVR
This… See the full description on the dataset page: https://huggingface.co/datasets/prism-vlm/rl_dataset.gemini_distill
PRISM Gemini Distill
Overview
PRISM Gemini Distill is our self-distilled multimodal reasoning dataset collected from Gemini 3 Flash for the PRISM project.
PRISM studies the distributional drift problem in the standard SFT → RLVR post-training pipeline. To mitigate this issue, PRISM introduces an intermediate Distribution Alignment / Pre-alignment stage before RLVR:
SFT → Distribution Alignment / Pre-alignment → RLVR
This dataset provides high-quality Gemini 3… See the full description on the dataset page: https://huggingface.co/datasets/prism-vlm/gemini_distill.VLM-SFTiq-terrain-vlm-dataset
IQ Terrain VLM Dataset
A high-fidelity, mathematically pristine Vision-Language Model (VLM) dataset designed specifically to teach models the procedural graphics and raymarching techniques of Inigo Quilez.
Dataset Summary
Most coding datasets rely on broadly scraped, often buggy code from GitHub or StackOverflow. This dataset takes a highly targeted approach:
Mathematical Ground Truth: All GLSL code and mathematical concepts are sourced directly from Inigo… See the full description on the dataset page: https://huggingface.co/datasets/True2456/iq-terrain-vlm-dataset.ogiri-bokete-unsloth-vlm
Japanese Bokete Ogiri — Unsloth VLM format
YANS-official/ogiri-bokete を、UnslothのVision SFTで扱える会話形式に変換した非公開用データセットです。
各JSONLレコードは「1画像 + 1回答」です。
{
"messages": [
{"role": "user", "content": [
{"type": "image", "image": "images/124469.jpg"},
{"type": "text", "text": "この画像のお題に対して、面白い一言を1つ返してください。"}
]},
{"role": "assistant", "content": [
{"type": "text", "text": "..."}
]}
]
}
Files
train.jsonl: 1,678 records / 630 prompts… See the full description on the dataset page: https://huggingface.co/datasets/beezza/ogiri-bokete-unsloth-vlm.VLM-CapCurriculum-TextReasoning-Data
VLM-CapCurriculum-TextReasoning (D_text)
Stage-2 textual-reasoning data for the staged post-training recipe in
"From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models"
(ICML 2026).
A curated ORZ-Math-13k subset — challenging text-only math problems used to consolidate textual reasoning between the perception (Stage 1) and visual-reasoning (Stage 3) RLVR stages of our recipe. Every row also ships with a precomputed pass_rate so… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-TextReasoning-Data.VLM_CCA
[!NOTE]
Planned improvements:
Human verification (image - keyword alignment; Q&A / translation)
Report VLM performance on this dataset
Include image license details in metadata
We welcome your feedback! Please contact us:
Lab: isds.sogang@gmail.com
Maintainer: bizli0618@sogang.ac.kr
VLM-CCA Korean Culture VQA Dataset
Dataset Summary
The Korean Culture VQA Dataset for Visual Language Model's Cultural Context Awareness (VLM-CCA) is a multimodal benchmark… See the full description on the dataset page: https://huggingface.co/datasets/SOGANG-ISDS/VLM_CCA.
