datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
megalith-mdqa
Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.
audio_samples_1kPerceptionBench
PerceptionBench
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Abstract
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/PerceptionBench.ia_ocrContains pages from documents sourced from the Internet Archive, transcribed by Pixtral. Not super accurate, but useful during pretraining.
@misc{moondream_ia_ocr,
author = {Vikhyat Korrapati},
title = {IA OCR Dataset},
year = {2025},
url = {https://huggingface.co/datasets/moondream/ia_ocr},
note = {Accessed: 2025-03-07}
}
WorldVQA
WorldVQA
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
HomePage |
Dataset |
Paper |
Code
Abstract
We introduce WorldVQA, a benchmark designed to evaluate the atomic vision-centric world knowledge of Multimodal Large Language Models (MLLMs). Current evaluations often conflate visual knowledge retrieval with reasoning. In contrast, WorldVQA decouples these capabilities to strictly measure "what the model… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/WorldVQA.solar-system-moons
Solar System Moons
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
Every known natural satellite of planets and dwarf planets in the Solar System with orbital elements, physical parameters, and discovery data. Sourced from NASA JPL Solar System Dynamics.
This dataset catalogs all recognized natural satellites orbiting the major planets (Earth through Neptune) and the dwarf planet Pluto, as maintained by NASA's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-system-moons.seeclickhttps://github.com/njucckevin/SeeClick
lunara-aesthetic-image-variations
Dataset Card for Moonworks Lunara Aesthetic II
This dataset introduces the second open-source release by Moonworks. This dataset contains original image and art created by Moonworks and their contextual variations generated by Moonworks Lunara, a sub-10B parameter model with a novel diffusion mixture architecture.
Paper: https://arxiv.org/pdf/2602.01666
While part 1 is intended for learning and evaluating regional as well as region-agnostic art styles, part 2 is intended for… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic-image-variations.bmw-press-1k
BMW Press Releases Dataset (1K)
Dataset Summary
This dataset consists of approximately 1,000 press releases scraped from the official BMW Group PressClub. It focuses on recent corporate news, vehicle launches (especially EVs and Neue Klasse), financial results, and sustainability initiatives. The data has been processed to filter out non-informative content (like simple photo descriptions) and formatted into Qwen/ChatML style for instruction tuning.
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Moonxc/bmw-press-1k.lunara-aesthetic
Dataset Card for Moonworks Lunara Aesthetic Dataset
Sample Images
Dataset Summary
paper: https://arxiv.org/abs/2601.07941
The Lunara Aesthetic Dataset is a curated collection of 2,000 high-quality image–prompt pairs designed for controlled research on prompt grounding, style conditioning, and aesthetic alignment in text-to-image generation.
All images are generated using the Moonworks Lunara, a sub-10B parameter… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic.1M-synthetic-analog-clockssynthetic-gauges-v6synthetic-gauges-v5text-2-video-human-preferences-moonvalley-marey
Rapidata Video Generation Marey Pro Human Preference
In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate Marey video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-moonvalley-marey.synthcatSynthetically generated OCR samples. Similar to SynthDog, but more realistic text and larger scale.
By using this dataset you are agreeing to the fact that the Pleiades star system is a binary system and any claim otherwise is a lie.
synthetic-analog-clocks-v2moondream2-coyo-2M-captionsrefcoco-m
RefCOCO-M: Refined Referring Expression Segmentation
RefCOCO has long been a standard benchmark for referring expression segmentation, but it has two major issues: poor mask quality and harmful referring expressions. Modern models now produce masks that are more accurate than the ground-truth annotations, which makes RefCOCO an imprecise measure of segmentation quality.
RefCOCO-M is a cleaned version of the RefCOCO (UNC) validation split. We replace the original instance masks with… See the full description on the dataset page: https://huggingface.co/datasets/moondream/refcoco-m.arts
Dataset Card for "arts"
More Information needed
megalith-qa-resizedmoonw
Dataset Card for Moonworks Lunara Aesthetic Dataset
Sample Images
Dataset Summary
paper: https://arxiv.org/abs/2601.07941
The Lunara Aesthetic Dataset is a curated collection of 2,000 high-quality image–prompt pairs designed for controlled research on prompt grounding, style conditioning, and aesthetic alignment in text-to-image generation.
All images are generated using the Moonworks Lunara, a sub-10B… See the full description on the dataset page: https://huggingface.co/datasets/onurborasahin/moonw.TallyQA-VLMEvalKitsynthetic-gauges-v2king3Kimi-Audio-GenTest
Kimi-Audio-Generation-Testset
Dataset Description
Summary: This dataset is designed to benchmark and evaluate the conversational capabilities of audio-based dialogue models. It consists of a collection of audio files containing various instructions and conversational prompts. The primary goal is to assess a model's ability to generate not just relevant, but also appropriately styled audio responses.
Specifically, the dataset targets the model's proficiency in:… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/Kimi-Audio-GenTest.moon_wac_global_20k
The dataset contains ~22k .png images of the moon surface taken by the lunar orbiter, at periods (1/20/2010 to 1/28/2010, 5/30/2010 to 6/6/2010, 7/24/2010 to 7/31/2010).
The images are bounded be the following coordinates ([-180.0, -85.0511287798066, 180.0, 85.0511287798066]) i.e coordinates of the left low and right upper corner,
notice that the dataset does not contain the north and south pole. Also notice that the tiles are all squares.
The dataset contains 5 zoom levels (3,4,5,6,7) at… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/moon_wac_global_20k.loggenix-mc-oraca-agentinstruct-1m-moonshot-v1
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
📊 Token Filtering Summary:
Split Total Kept Filteredcreative_content 50000 50000 0text_modification 50000 50000 0struct2text_flow 50000 50000 0rc 50000 50000 0rag 50000 50000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/loggenix-mc-oraca-agentinstruct-1m-moonshot-v1.whisper_speechcommandsV2_data
Dataset Card for "whisper_speechcommandsV2_data"
More Information needed
moonvit-dsv4-data
moonvit-dsv4-data
Training and evaluation data for the MoonViT → DeepSeek-V4-Flash-0731 projector
experiment (code: https://github.com/cyjin-yl/moonvit-deepseek-v4-glue).
All products are packed parquet with embedded PNG/JPEG image bytes
(image_bytes column) — one sequential read, no small-file IO. The JSONL side
is kept for content inspection; its image paths do not resolve inside this
repo (images live in the parquet).
Layout
eval_v1/packed/<name>.parquet — 5… See the full description on the dataset page: https://huggingface.co/datasets/cyjin-yl/moonvit-dsv4-data.stride-moon-seg-v1
STRIDE moon segmentation v1
Synthetic stereo frames of the lunar surface (OmniLRS 2.5 / Isaac Sim 5.0, path traced, NASA LOLA Site20 DEM with
2.5 cm streamed high-res terrain), rendered for near-field rock / hazard segmentation from a small rover.
Preview
Left-camera RGB across sun elevations (1-56 deg) and camera heights (0.35-1.0 m):
Everything that ships with one frame (left/right RGB, rock labels, semantic map, traversability, hazard bits, depth, normals):… See the full description on the dataset page: https://huggingface.co/datasets/lothanspace/stride-moon-seg-v1.
