CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01moondream /megalith-mdqa Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM. imagequestion-answering1M<n<10M28 likes20k downloads1y agoHugging Face02moonshine-ai /audio_samples_1kaudio0 likes9.1k downloads6mo agoHugging Face03Moon23nn /imgbedaudion<1K0 likes5.8k downloads2mo agoHugging Face04moonshotai /PerceptionBench PerceptionBench PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Abstract We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/PerceptionBench.tabularvisual-question-answering1K<n<10K52 likes2.5k downloads2mo agoHugging Face05moondream /ia_ocrContains pages from documents sourced from the Internet Archive, transcribed by Pixtral. Not super accurate, but useful during pretraining. @misc{moondream_ia_ocr, author = {Vikhyat Korrapati}, title = {IA OCR Dataset}, year = {2025}, url = {https://huggingface.co/datasets/moondream/ia_ocr}, note = {Accessed: 2025-03-07} } image100K<n<1M28 likes2.4k downloads1y agoHugging Face06moonmonster /DrivingStereo_Dataset_Encrypted0 likes2.4k downloads7mo agoHugging Face07moonshotai /WorldVQA WorldVQA WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models HomePage | Dataset | Paper | Code Abstract We introduce WorldVQA, a benchmark designed to evaluate the atomic vision-centric world knowledge of Multimodal Large Language Models (MLLMs). Current evaluations often conflate visual knowledge retrieval with reasoning. In contrast, WorldVQA decouples these capabilities to strictly measure "what the model… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/WorldVQA.imagevisual-question-answering1K<n<10K67 likes1.5k downloads8mo agoHugging Face08juliensimon /solar-system-moons Solar System Moons Credit: NASA/JPL-Caltech Part of a dataset collection on Hugging Face. Dataset description Every known natural satellite of planets and dwarf planets in the Solar System with orbital elements, physical parameters, and discovery data. Sourced from NASA JPL Solar System Dynamics. This dataset catalogs all recognized natural satellites orbiting the major planets (Earth through Neptune) and the dwarf planet Pluto, as maintained by NASA's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-system-moons.tabulartabular-classificationn<1K0 likes1.4k downloads2d agoHugging Face09moondream /seeclickhttps://github.com/njucckevin/SeeClick image100K<n<1M7 likes991 downloads1y agoHugging Face10moon70 /refusal-lens-graphs0 likes990 downloads3mo agoHugging Face11ayushprd /Moonstone Moonstone: A Multimodal Foundation Model Benchmark for Lunar Remote Sensing This repository contains the dataset for the paper Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing. The official code is available at GitHub. A 28-channel, 128 pixels-per-degree (~237 m/pixel) global multimodal lunar dataset assembled from seven instrument families across five missions (LRO WAC/LOLA/Diviner/Mini-RF, Chandrayaan-1 M3, GRAIL, Lunar Prospector GRS… See the full description on the dataset page: https://huggingface.co/datasets/ayushprd/Moonstone.imageimage-classificationn<1K0 likes921 downloads3mo agoHugging Face12moonmonster /CREStereo_Dataset_Encrypted0 likes815 downloads7mo agoHugging Face13moonworks /lunara-aesthetic-image-variations Dataset Card for Moonworks Lunara Aesthetic II This dataset introduces the second open-source release by Moonworks. This dataset contains original image and art created by Moonworks and their contextual variations generated by Moonworks Lunara, a sub-10B parameter model with a novel diffusion mixture architecture. Paper: https://arxiv.org/pdf/2602.01666 While part 1 is intended for learning and evaluating regional as well as region-agnostic art styles, part 2 is intended for… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic-image-variations.imageimage-to-image1K<n<10K67 likes688 downloads8mo agoHugging Face14Moonxc /bmw-press-1k BMW Press Releases Dataset (1K) Dataset Summary This dataset consists of approximately 1,000 press releases scraped from the official BMW Group PressClub. It focuses on recent corporate news, vehicle launches (especially EVs and Neue Klasse), financial results, and sustainability initiatives. The data has been processed to filter out non-informative content (like simple photo descriptions) and formatted into Qwen/ChatML style for instruction tuning. This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Moonxc/bmw-press-1k.documenttext-generationn<1K0 likes664 downloads9mo agoHugging Face15moonworks /lunara-aesthetic Dataset Card for Moonworks Lunara Aesthetic Dataset Sample Images Dataset Summary paper: https://arxiv.org/abs/2601.07941 The Lunara Aesthetic Dataset is a curated collection of 2,000 high-quality image–prompt pairs designed for controlled research on prompt grounding, style conditioning, and aesthetic alignment in text-to-image generation. All images are generated using the Moonworks Lunara, a sub-10B parameter… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic.imagetext-to-image1K<n<10K89 likes636 downloads8mo agoHugging Face16moondream /1M-synthetic-analog-clocksimage1M<n<10M4 likes543 downloads2y agoHugging Face17moondream /synthetic-gauges-v6image100K<n<1M0 likes511 downloads2y agoHugging Face18moonx3 /king20 likes511 downloads1y agoHugging Face19moondream /synthetic-gauges-v5image100K<n<1M1 likes504 downloads2y agoHugging Face20Rapidata /text-2-video-human-preferences-moonvalley-marey Rapidata Video Generation Marey Pro Human Preference In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate Marey video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-moonvalley-marey.imagevideo-classification1K<n<10K7 likes469 downloads1y agoHugging Face21moondream /synthcatSynthetically generated OCR samples. Similar to SynthDog, but more realistic text and larger scale. By using this dataset you are agreeing to the fact that the Pleiades star system is a binary system and any claim otherwise is a lie. image1M<n<10M9 likes454 downloads1y agoHugging Face22Moony97 /imgbedimagen<1K0 likes416 downloads22d agoHugging Face23moondream /synthetic-analog-clocks-v2image100K<n<1M0 likes373 downloads2y agoHugging Face24DJ-Research /GameBoy-harvest_moon_10 likes349 downloads4mo agoHugging Face25moon1ite /kor_hateHuman-annotated Korean corpus collected from a popular domestic entertainment news aggregation platform for toxic speech detection. Comments are annotated for gender bias, social bias and hate speech.text-classification1K<n<10K8 likes341 downloads3y agoHugging Face26model-organisms-for-real /gemma2_9b_it_taboo_moon_oracle_v1-training-data0 likes333 downloads3mo agoHugging Face27ljnlonoljpiljm /moondream2-coyo-2M-captionsimage1M<n<10M0 likes323 downloads1y agoHugging Face28moondream /refcoco-m RefCOCO-M: Refined Referring Expression Segmentation RefCOCO has long been a standard benchmark for referring expression segmentation, but it has two major issues: poor mask quality and harmful referring expressions. Modern models now produce masks that are more accurate than the ground-truth annotations, which makes RefCOCO an imprecise measure of segmentation quality. RefCOCO-M is a cleaned version of the RefCOCO (UNC) validation split. We replace the original instance masks with… See the full description on the dataset page: https://huggingface.co/datasets/moondream/refcoco-m.image1K<n<10K49 likes322 downloads10mo agoHugging Face29mooncakex /arts Dataset Card for "arts" More Information needed image10K<n<100K1 likes317 downloads3y agoHugging Face30mOONIm /PhysicalAI-SmartSpaces Physical AI Smart Spaces Dataset Overview Comprehensive, annotated dataset for multi-camera tracking and 2D/3D object detection. This dataset is synthetically generated with Omniverse. This dataset consists of over 250 hours of video from across nearly 1,500 cameras from indoor scenes in warehouses, hospitals, retail, and more. The dataset is time synchronized for tracking humans across multiple cameras using feature representation and no personal data.… See the full description on the dataset page: https://huggingface.co/datasets/mOONIm/PhysicalAI-SmartSpaces.0 likes302 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.