datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
megalith-mdqa
Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.
imgbedia_ocrContains pages from documents sourced from the Internet Archive, transcribed by Pixtral. Not super accurate, but useful during pretraining.
@misc{moondream_ia_ocr,
author = {Vikhyat Korrapati},
title = {IA OCR Dataset},
year = {2025},
url = {https://huggingface.co/datasets/moondream/ia_ocr},
note = {Accessed: 2025-03-07}
}
WorldVQA
WorldVQA
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
HomePage |
Dataset |
Paper |
Code
Abstract
We introduce WorldVQA, a benchmark designed to evaluate the atomic vision-centric world knowledge of Multimodal Large Language Models (MLLMs). Current evaluations often conflate visual knowledge retrieval with reasoning. In contrast, WorldVQA decouples these capabilities to strictly measure "what the model… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/WorldVQA.seeclickhttps://github.com/njucckevin/SeeClick
Moonstone
Moonstone: A Multimodal Foundation Model Benchmark for Lunar Remote Sensing
This repository contains the dataset for the paper Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing.
The official code is available at GitHub.
A 28-channel, 128 pixels-per-degree (~237 m/pixel) global multimodal lunar dataset assembled from
seven instrument families across five missions (LRO WAC/LOLA/Diviner/Mini-RF, Chandrayaan-1 M3,
GRAIL, Lunar Prospector GRS… See the full description on the dataset page: https://huggingface.co/datasets/ayushprd/Moonstone.lunara-aesthetic-image-variations
Dataset Card for Moonworks Lunara Aesthetic II
This dataset introduces the second open-source release by Moonworks. This dataset contains original image and art created by Moonworks and their contextual variations generated by Moonworks Lunara, a sub-10B parameter model with a novel diffusion mixture architecture.
Paper: https://arxiv.org/pdf/2602.01666
While part 1 is intended for learning and evaluating regional as well as region-agnostic art styles, part 2 is intended for… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic-image-variations.lunara-aesthetic
Dataset Card for Moonworks Lunara Aesthetic Dataset
Sample Images
Dataset Summary
paper: https://arxiv.org/abs/2601.07941
The Lunara Aesthetic Dataset is a curated collection of 2,000 high-quality image–prompt pairs designed for controlled research on prompt grounding, style conditioning, and aesthetic alignment in text-to-image generation.
All images are generated using the Moonworks Lunara, a sub-10B parameter… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic.1M-synthetic-analog-clockssynthetic-gauges-v6synthetic-gauges-v5text-2-video-human-preferences-moonvalley-marey
Rapidata Video Generation Marey Pro Human Preference
In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate Marey video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-moonvalley-marey.synthcatSynthetically generated OCR samples. Similar to SynthDog, but more realistic text and larger scale.
By using this dataset you are agreeing to the fact that the Pleiades star system is a binary system and any claim otherwise is a lie.
imgbedsynthetic-analog-clocks-v2moondream2-coyo-2M-captionsarts
Dataset Card for "arts"
More Information needed
refcoco-m
RefCOCO-M: Refined Referring Expression Segmentation
RefCOCO has long been a standard benchmark for referring expression segmentation, but it has two major issues: poor mask quality and harmful referring expressions. Modern models now produce masks that are more accurate than the ground-truth annotations, which makes RefCOCO an imprecise measure of segmentation quality.
RefCOCO-M is a cleaned version of the RefCOCO (UNC) validation split. We replace the original instance masks with… See the full description on the dataset page: https://huggingface.co/datasets/moondream/refcoco-m.megalith-qa-resizedMoon_Mapmoonw
Dataset Card for Moonworks Lunara Aesthetic Dataset
Sample Images
Dataset Summary
paper: https://arxiv.org/abs/2601.07941
The Lunara Aesthetic Dataset is a curated collection of 2,000 high-quality image–prompt pairs designed for controlled research on prompt grounding, style conditioning, and aesthetic alignment in text-to-image generation.
All images are generated using the Moonworks Lunara, a sub-10B… See the full description on the dataset page: https://huggingface.co/datasets/onurborasahin/moonw.synthetic-gauges-v2moon_wac_global_20k
The dataset contains ~22k .png images of the moon surface taken by the lunar orbiter, at periods (1/20/2010 to 1/28/2010, 5/30/2010 to 6/6/2010, 7/24/2010 to 7/31/2010).
The images are bounded be the following coordinates ([-180.0, -85.0511287798066, 180.0, 85.0511287798066]) i.e coordinates of the left low and right upper corner,
notice that the dataset does not contain the north and south pole. Also notice that the tiles are all squares.
The dataset contains 5 zoom levels (3,4,5,6,7) at… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/moon_wac_global_20k.spacecraft_moonlanding_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "spacecraft",
"total_episodes": 100,
"total_frames": 52116,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 1,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/arclabmit/spacecraft_moonlanding_dataset.stride-moon-seg-v1
STRIDE moon segmentation v1
Synthetic stereo frames of the lunar surface (OmniLRS 2.5 / Isaac Sim 5.0, path traced, NASA LOLA Site20 DEM with
2.5 cm streamed high-res terrain), rendered for near-field rock / hazard segmentation from a small rover.
Preview
Left-camera RGB across sun elevations (1-56 deg) and camera heights (0.35-1.0 m):
Everything that ships with one frame (left/right RGB, rock labels, semantic map, traversability, hazard bits, depth, normals):… See the full description on the dataset page: https://huggingface.co/datasets/lothanspace/stride-moon-seg-v1.FineVisionShuffle
FineVision Filtered
Filtered FineVision dataset. Removed samples containing Chinese, Japanese, Korean, Russian/Cyrillic, and Vietnamese text.
Subsets
CoSyn_400k_chemical
CoSyn_400k_circuit
CoSyn_400k_diagram
CoSyn_400k_document
CoSyn_400k_graphic
CoSyn_400k_math
CoSyn_400k_music
CoSyn_400k_nutrition
CoSyn_400k_table
SynthFormulaNet
a_okvqa
aguvis-stage-1
ai2d_merged
alfworldgpt
allava_laion
allava_vflan
art
arxivqa
bentham
blockdiagramcomputerized
blockdiagramhandwritten… See the full description on the dataset page: https://huggingface.co/datasets/moondream/FineVisionShuffle.100k-synthetic-clockssynthetic-gauges-v4moon_clem750Liebherr_Product
Liebherr Product (LP) Dataset
Liebherr Product (LP) is a self-collected object-detection dataset of construction
machines, containing over 15,000 high-quality images across 23 categories of
construction machinery, including articulated dump trucks, bulldozers, combined
piling and drilling rigs, various cranes, excavators, loaders, and more.
It accompanies the paper DART: An automated end-to-end object detection pipeline
with data Diversification, open-vocabulary bounding box… See the full description on the dataset page: https://huggingface.co/datasets/Moonxc/Liebherr_Product.
