datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ft-llm-2026-qa-dataset
FT-LLM 2026 QA Dataset
A Japanese visual-question-answering dataset used for Stage 1-2 visual instruction tuning of the COMPASS Vision-Language Model. Each sample contains a document or natural image together with one or more Japanese question–answer pairs, and is designed to give the VLM its instruction-following and VQA capabilities. Images are embedded in the dataset, so no external downloads are required.
Part of the Compass collection.
License
Released under the… See the full description on the dataset page: https://huggingface.co/datasets/Yana/ft-llm-2026-qa-dataset.AMALIA-VL-SFT-Dataset
AMALIA-VL-Training-Dataset
Dataset Description
This dataset is provided as part of the AMALIA project.
This is the vision+language training mix for AMALIA-VL-SFT. Each subset is one
source dataset in the mix, each with a single train split. The only datasets that are absent from this mix are those that derive directly from the core LLM training mix, and can be found in the AMALIA-LLM Post Training Collection.
Example usage:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-VL-SFT-Dataset.ft-llm-2026-ocr-dataset
FT-LLM 2026 OCR Dataset
A Japanese document-OCR dataset used for Stage 1-1 caption + OCR pretraining of the COMPASS Vision-Language Model. Each sample pairs a rendered page image from a Japanese public-sector financial PDF (Cabinet Office, Financial Services Agency, Ministry of Finance) with its OCR-extracted markdown text. It is intended to teach the VLM's MLP projector to align vision tokens with Japanese text.
Part of the Compass collection.
License
Released under… See the full description on the dataset page: https://huggingface.co/datasets/Yana/ft-llm-2026-ocr-dataset.llm-tarot-datasetagri_llm_datasetmedical-vision-llm-dataset
Combined Medical Vision-Language Dataset
Dataset Description
Comprehensive medical vision-language dataset with 4793 samples for vision-based LLM training.
Dataset Statistics
Total Samples: 4793
Training Samples: 3834
Validation Samples: 959
Modality Distribution
X-ray: 2325 samples
CT: 1351 samples
Unknown: 812 samples
MRI: 231 samples
Ultrasound: 70 samples
Microscopy: 2 samples
Endoscopy: 2 samples
Body Part Distribution
Unknown:… See the full description on the dataset page: https://huggingface.co/datasets/robailleo/medical-vision-llm-dataset.LLM_EVAL_datasetmedical-vision-llm-dataset-TEST
Medical Vision-Language Dataset
Combined from ROCO, VQA-RAD, and PubMedVision.
Stats
Total: 60
Train: 48
Validation: 12
Sources
PubMedVision: 20
VQA-RAD: 20
ROCO: 20
Created: 2025-11-23
medical-vision-llm-dataset-v2
Medical Vision-Language Dataset
Combined dataset for training medical vision-language models.
Statistics
Total: 3793 samples
Train: 3035 | Validation: 758
Sources
ROCO: 2000
VQA-RAD: 1793
Created: 2025-11-23 22:25:36
AMALIA-VL-DPO-Dataset
AMALIA-VL-DPO Dataset
Dataset Description
This dataset is provided as part of the AMALIA project.
Direct Preference Optimization (DPO) data used for AMALIA-VL. Each subset is one source in
the mix, with vl_preference_200k being fully derived from the SFT mix. Every row is a preference triplet:
column
description
prompt
normalized [{role, content}] turns; <image> marks image position (multimodal subsets)
chosen
preferred assistant response… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-VL-DPO-Dataset.
