CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01moondream /megalith-mdqa Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM. imagequestion-answering1M<n<10M28 likes19k downloads1y agoHugging Face02moondream /ia_ocrContains pages from documents sourced from the Internet Archive, transcribed by Pixtral. Not super accurate, but useful during pretraining. @misc{moondream_ia_ocr, author = {Vikhyat Korrapati}, title = {IA OCR Dataset}, year = {2025}, url = {https://huggingface.co/datasets/moondream/ia_ocr}, note = {Accessed: 2025-03-07} } image100K<n<1M28 likes2.4k downloads1y agoHugging Face03moondream /seeclickhttps://github.com/njucckevin/SeeClick image100K<n<1M7 likes990 downloads1y agoHugging Face04moondream /1M-synthetic-analog-clocksimage1M<n<10M4 likes548 downloads2y agoHugging Face05moondream /synthetic-gauges-v6image100K<n<1M0 likes518 downloads2y agoHugging Face06moondream /synthetic-gauges-v5image100K<n<1M1 likes508 downloads2y agoHugging Face07moondream /synthcatSynthetically generated OCR samples. Similar to SynthDog, but more realistic text and larger scale. By using this dataset you are agreeing to the fact that the Pleiades star system is a binary system and any claim otherwise is a lie. image1M<n<10M9 likes452 downloads1y agoHugging Face08moondream /synthetic-analog-clocks-v2image100K<n<1M0 likes377 downloads2y agoHugging Face09ljnlonoljpiljm /moondream2-coyo-2M-captionsimage1M<n<10M0 likes330 downloads1y agoHugging Face10moondream /refcoco-m RefCOCO-M: Refined Referring Expression Segmentation RefCOCO has long been a standard benchmark for referring expression segmentation, but it has two major issues: poor mask quality and harmful referring expressions. Modern models now produce masks that are more accurate than the ground-truth annotations, which makes RefCOCO an imprecise measure of segmentation quality. RefCOCO-M is a cleaned version of the RefCOCO (UNC) validation split. We replace the original instance masks with… See the full description on the dataset page: https://huggingface.co/datasets/moondream/refcoco-m.image1K<n<10K49 likes316 downloads10mo agoHugging Face11moondream /megalith-qa-resizedimage1M<n<10M3 likes282 downloads2y agoHugging Face12moondream /TallyQA-VLMEvalKittext10K<n<100K0 likes250 downloads1y agoHugging Face13moondream /synthetic-gauges-v2image100K<n<1M0 likes247 downloads2y agoHugging Face14moondream /FineVisionShuffle FineVision Filtered Filtered FineVision dataset. Removed samples containing Chinese, Japanese, Korean, Russian/Cyrillic, and Vietnamese text. Subsets CoSyn_400k_chemical CoSyn_400k_circuit CoSyn_400k_diagram CoSyn_400k_document CoSyn_400k_graphic CoSyn_400k_math CoSyn_400k_music CoSyn_400k_nutrition CoSyn_400k_table SynthFormulaNet a_okvqa aguvis-stage-1 ai2d_merged alfworldgpt allava_laion allava_vflan art arxivqa bentham blockdiagramcomputerized blockdiagramhandwritten… See the full description on the dataset page: https://huggingface.co/datasets/moondream/FineVisionShuffle.image1M<n<10M4 likes120 downloads1y agoHugging Face15moondream /100k-synthetic-clocksimage100K<n<1M0 likes117 downloads2y agoHugging Face16moondream /synthetic-gauges-v4image100K<n<1M0 likes114 downloads2y agoHugging Face17g-ronimo /imagenet-1k-vl-enriched_moondream2 visual-layer/imagenet-1k-vl-enriched recaptioned with vikhyatk/moondream2 short captions ids and captions only Prompt used: This is an image of a {class_name}. The current image caption is {caption}. Please write a short caption based on the image content and the current caption. Keep it short and precise. text1M<n<10M2 likes114 downloads2y agoHugging Face18SwayStar123 /preprocessed_recap-coco30k-moondreamtext10K<n<100K0 likes79 downloads2y agoHugging Face19sroecker /recap-coco30k-moondreamimage10K<n<100K0 likes74 downloads2y agoHugging Face20moondream /geoguessr-countries-finetune GeoGuessr Countries Finetune Google Earth images from around the world with the country as the target label. Splits Split Samples train 25000 test 400 Columns image: Google Earth image country: country label Countries In This Release Argentina, Australia, Austria, Bangladesh, Belgium, Bolivia, Botswana, Brazil, Bulgaria, Cambodia, Canada, Chile, Colombia, Croatia, Czechia, Denmark, Finland, France, Germany, Ghana, Greece, Hungary… See the full description on the dataset page: https://huggingface.co/datasets/moondream/geoguessr-countries-finetune.imageimage-classification10K<n<100K2 likes71 downloads6mo agoHugging Face21moondream /brackish_underwater Brackish Underwater An object detection dataset of underwater footage from brackish water environments in temperate waters, featuring various marine animals. Background This dataset was introduced in the CVPR 2019 workshop paper "Detection of Marine Animals in a New Underwater Dataset with Varying Visibility" by Pedersen, Haurum, Gade, and Moeslund from Aalborg University. The images were captured from permanently mounted cameras in saltwater straits for long-term marine… See the full description on the dataset page: https://huggingface.co/datasets/moondream/brackish_underwater.imageobject-detection10K<n<100K0 likes65 downloads8mo agoHugging Face22moondream /glaucoma-detection Glaucoma Detection Retinal fundus images for glaucoma stage classification. Splits Split Samples train 2847 validation 1259 test 1272 Columns image: retinal image class: glaucoma stage label Classes Class Description normal No glaucoma visible in the image. early Early-stage glaucoma findings. advanced Advanced glaucoma findings. imageimage-classification1K<n<10K1 likes58 downloads6mo agoHugging Face23moondream /mrc-synth-cottabular100K<n<1M0 likes58 downloads4mo agoHugging Face24moondream /ssv2-3x3 SSV2 3x3 A multi-frame action dataset for teaching a model to look across several frames and describe the action it sees. Each image is a 3x3 collage built from frames sampled from the original Something-Something V2 videos. Splits Split Samples train 150000 test 1000 Columns image: 3x3 frame collage label_text: action label text annotation_text: action template text video_id: original video identifier imageimage-classification100K<n<1M0 likes48 downloads6mo agoHugging Face25merve /docvqa-media-labeled-moondream Box-overlay preview Auto-generated sample of the labelled boxes (with judge scores when available). Regenerated on every push. imagen<1K0 likes46 downloads3mo agoHugging Face26sroecker /datikz-v2-moondream-labelsimage10K<n<100K2 likes41 downloads2y agoHugging Face27moondream /synthetic-analog-clocks-v3image100K<n<1M0 likes37 downloads2y agoHugging Face28ljnlonoljpiljm /recap-coco30k-moondream-chunksimage10K<n<100K0 likes27 downloads1y agoHugging Face29moondream /ishape-logs iShape Logs A subset of the iShape dataset for log instances. Source dataset: https://ishape.github.io Splits Split Samples train 2000 validation 500 Columns image: source image objects: object annotations with name and rle Annotation Format rle stores a COCO-style compressed mask string in height width counts form. imageimage-segmentation1K<n<10K0 likes26 downloads6mo agoHugging Face30moondream /CountBenchQA-VLMEvalKittextn<1K0 likes25 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.