datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
megalith-mdqa
Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.
audio_samples_1kimgbedPerceptionBench
PerceptionBench
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Abstract
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/PerceptionBench.ia_ocrContains pages from documents sourced from the Internet Archive, transcribed by Pixtral. Not super accurate, but useful during pretraining.
@misc{moondream_ia_ocr,
author = {Vikhyat Korrapati},
title = {IA OCR Dataset},
year = {2025},
url = {https://huggingface.co/datasets/moondream/ia_ocr},
note = {Accessed: 2025-03-07}
}
DrivingStereo_Dataset_EncryptedWorldVQA
WorldVQA
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
HomePage |
Dataset |
Paper |
Code
Abstract
We introduce WorldVQA, a benchmark designed to evaluate the atomic vision-centric world knowledge of Multimodal Large Language Models (MLLMs). Current evaluations often conflate visual knowledge retrieval with reasoning. In contrast, WorldVQA decouples these capabilities to strictly measure "what the model… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/WorldVQA.solar-system-moons
Solar System Moons
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
Every known natural satellite of planets and dwarf planets in the Solar System with orbital elements, physical parameters, and discovery data. Sourced from NASA JPL Solar System Dynamics.
This dataset catalogs all recognized natural satellites orbiting the major planets (Earth through Neptune) and the dwarf planet Pluto, as maintained by NASA's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-system-moons.seeclickhttps://github.com/njucckevin/SeeClick
refusal-lens-graphsMoonstone
Moonstone: A Multimodal Foundation Model Benchmark for Lunar Remote Sensing
This repository contains the dataset for the paper Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing.
The official code is available at GitHub.
A 28-channel, 128 pixels-per-degree (~237 m/pixel) global multimodal lunar dataset assembled from
seven instrument families across five missions (LRO WAC/LOLA/Diviner/Mini-RF, Chandrayaan-1 M3,
GRAIL, Lunar Prospector GRS… See the full description on the dataset page: https://huggingface.co/datasets/ayushprd/Moonstone.CREStereo_Dataset_Encryptedlunara-aesthetic-image-variations
Dataset Card for Moonworks Lunara Aesthetic II
This dataset introduces the second open-source release by Moonworks. This dataset contains original image and art created by Moonworks and their contextual variations generated by Moonworks Lunara, a sub-10B parameter model with a novel diffusion mixture architecture.
Paper: https://arxiv.org/pdf/2602.01666
While part 1 is intended for learning and evaluating regional as well as region-agnostic art styles, part 2 is intended for… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic-image-variations.bmw-press-1k
BMW Press Releases Dataset (1K)
Dataset Summary
This dataset consists of approximately 1,000 press releases scraped from the official BMW Group PressClub. It focuses on recent corporate news, vehicle launches (especially EVs and Neue Klasse), financial results, and sustainability initiatives. The data has been processed to filter out non-informative content (like simple photo descriptions) and formatted into Qwen/ChatML style for instruction tuning.
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Moonxc/bmw-press-1k.lunara-aesthetic
Dataset Card for Moonworks Lunara Aesthetic Dataset
Sample Images
Dataset Summary
paper: https://arxiv.org/abs/2601.07941
The Lunara Aesthetic Dataset is a curated collection of 2,000 high-quality image–prompt pairs designed for controlled research on prompt grounding, style conditioning, and aesthetic alignment in text-to-image generation.
All images are generated using the Moonworks Lunara, a sub-10B parameter… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic.1M-synthetic-analog-clockssynthetic-gauges-v6king2synthetic-gauges-v5text-2-video-human-preferences-moonvalley-marey
Rapidata Video Generation Marey Pro Human Preference
In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate Marey video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-moonvalley-marey.synthcatSynthetically generated OCR samples. Similar to SynthDog, but more realistic text and larger scale.
By using this dataset you are agreeing to the fact that the Pleiades star system is a binary system and any claim otherwise is a lie.
imgbedsynthetic-analog-clocks-v2GameBoy-harvest_moon_1kor_hateHuman-annotated Korean corpus collected from a popular domestic entertainment news aggregation platform
for toxic speech detection. Comments are annotated for gender bias, social bias and hate speech.gemma2_9b_it_taboo_moon_oracle_v1-training-datamoondream2-coyo-2M-captionsrefcoco-m
RefCOCO-M: Refined Referring Expression Segmentation
RefCOCO has long been a standard benchmark for referring expression segmentation, but it has two major issues: poor mask quality and harmful referring expressions. Modern models now produce masks that are more accurate than the ground-truth annotations, which makes RefCOCO an imprecise measure of segmentation quality.
RefCOCO-M is a cleaned version of the RefCOCO (UNC) validation split. We replace the original instance masks with… See the full description on the dataset page: https://huggingface.co/datasets/moondream/refcoco-m.arts
Dataset Card for "arts"
More Information needed
PhysicalAI-SmartSpaces
Physical AI Smart Spaces Dataset
Overview
Comprehensive, annotated dataset for multi-camera tracking and 2D/3D object detection. This dataset is synthetically generated with Omniverse.
This dataset consists of over 250 hours of video from across nearly 1,500 cameras from indoor scenes in warehouses, hospitals, retail, and more. The dataset is time synchronized for tracking humans across multiple cameras using feature representation and no personal data.… See the full description on the dataset page: https://huggingface.co/datasets/mOONIm/PhysicalAI-SmartSpaces.
