datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
megalith-mdqa
Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.
ia_ocrContains pages from documents sourced from the Internet Archive, transcribed by Pixtral. Not super accurate, but useful during pretraining.
@misc{moondream_ia_ocr,
author = {Vikhyat Korrapati},
title = {IA OCR Dataset},
year = {2025},
url = {https://huggingface.co/datasets/moondream/ia_ocr},
note = {Accessed: 2025-03-07}
}
seeclickhttps://github.com/njucckevin/SeeClick
1M-synthetic-analog-clockssynthetic-gauges-v6synthetic-gauges-v5synthcatSynthetically generated OCR samples. Similar to SynthDog, but more realistic text and larger scale.
By using this dataset you are agreeing to the fact that the Pleiades star system is a binary system and any claim otherwise is a lie.
synthetic-analog-clocks-v2moondream2-coyo-2M-captionsrefcoco-m
RefCOCO-M: Refined Referring Expression Segmentation
RefCOCO has long been a standard benchmark for referring expression segmentation, but it has two major issues: poor mask quality and harmful referring expressions. Modern models now produce masks that are more accurate than the ground-truth annotations, which makes RefCOCO an imprecise measure of segmentation quality.
RefCOCO-M is a cleaned version of the RefCOCO (UNC) validation split. We replace the original instance masks with… See the full description on the dataset page: https://huggingface.co/datasets/moondream/refcoco-m.megalith-qa-resizedTallyQA-VLMEvalKitsynthetic-gauges-v2FineVisionShuffle
FineVision Filtered
Filtered FineVision dataset. Removed samples containing Chinese, Japanese, Korean, Russian/Cyrillic, and Vietnamese text.
Subsets
CoSyn_400k_chemical
CoSyn_400k_circuit
CoSyn_400k_diagram
CoSyn_400k_document
CoSyn_400k_graphic
CoSyn_400k_math
CoSyn_400k_music
CoSyn_400k_nutrition
CoSyn_400k_table
SynthFormulaNet
a_okvqa
aguvis-stage-1
ai2d_merged
alfworldgpt
allava_laion
allava_vflan
art
arxivqa
bentham
blockdiagramcomputerized
blockdiagramhandwritten… See the full description on the dataset page: https://huggingface.co/datasets/moondream/FineVisionShuffle.100k-synthetic-clockssynthetic-gauges-v4imagenet-1k-vl-enriched_moondream2
visual-layer/imagenet-1k-vl-enriched recaptioned with vikhyatk/moondream2
short captions
ids and captions only
Prompt used:
This is an image of a {class_name}. The current image caption is {caption}.
Please write a short caption based on the image content and the current caption.
Keep it short and precise.
preprocessed_recap-coco30k-moondreamrecap-coco30k-moondreamgeoguessr-countries-finetune
GeoGuessr Countries Finetune
Google Earth images from around the world with the country as the target label.
Splits
Split
Samples
train
25000
test
400
Columns
image: Google Earth image
country: country label
Countries In This Release
Argentina, Australia, Austria, Bangladesh, Belgium, Bolivia, Botswana, Brazil, Bulgaria,
Cambodia, Canada, Chile, Colombia, Croatia, Czechia, Denmark, Finland, France, Germany,
Ghana, Greece, Hungary… See the full description on the dataset page: https://huggingface.co/datasets/moondream/geoguessr-countries-finetune.brackish_underwater
Brackish Underwater
An object detection dataset of underwater footage from brackish water environments in temperate waters, featuring various marine animals.
Background
This dataset was introduced in the CVPR 2019 workshop paper "Detection of Marine Animals in a New Underwater Dataset with Varying Visibility" by Pedersen, Haurum, Gade, and Moeslund from Aalborg University.
The images were captured from permanently mounted cameras in saltwater straits for long-term marine… See the full description on the dataset page: https://huggingface.co/datasets/moondream/brackish_underwater.glaucoma-detection
Glaucoma Detection
Retinal fundus images for glaucoma stage classification.
Splits
Split
Samples
train
2847
validation
1259
test
1272
Columns
image: retinal image
class: glaucoma stage label
Classes
Class
Description
normal
No glaucoma visible in the image.
early
Early-stage glaucoma findings.
advanced
Advanced glaucoma findings.
mrc-synth-cotssv2-3x3
SSV2 3x3
A multi-frame action dataset for teaching a model to look across several frames and describe the action it sees.
Each image is a 3x3 collage built from frames sampled from the original Something-Something V2 videos.
Splits
Split
Samples
train
150000
test
1000
Columns
image: 3x3 frame collage
label_text: action label text
annotation_text: action template text
video_id: original video identifier
docvqa-media-labeled-moondream
Box-overlay preview
Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push.
datikz-v2-moondream-labelssynthetic-analog-clocks-v3recap-coco30k-moondream-chunksishape-logs
iShape Logs
A subset of the iShape dataset for log instances.
Source dataset: https://ishape.github.io
Splits
Split
Samples
train
2000
validation
500
Columns
image: source image
objects: object annotations with name and rle
Annotation Format
rle stores a COCO-style compressed mask string in height width counts form.
CountBenchQA-VLMEvalKit
