CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01detection-datasets /cocoimageobject-detection100K<n<1M89 likes22k downloads4y agoHugging Face02allenai /coconot 🥥 CoCoNot: Contextually, Comply Not! Dataset Card Dataset Details Dataset Description Chat-based language models are designed to be helpful, yet they should not comply with every user request. While most existing work primarily focuses on refusal of "unsafe" queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user… See the full description on the dataset page: https://huggingface.co/datasets/allenai/coconot.texttext-generation10K<n<100K25 likes6.2k downloads2y agoHugging Face03yerevann /coco-karpathy Dataset Card for "yerevann/coco-karpathy" The Karpathy split of COCO for image captioning. imageimage-to-text100K<n<1M22 likes5.3k downloads4y agoHugging Face04lmms-lab-encoder /COCO-Caption Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2014-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption.image10K<n<100K15 likes4.4k downloads3y agoHugging Face05jxie /coco_captions Dataset Card for "coco_captions" More Information needed image100K<n<1M18 likes4k downloads3y agoHugging Face06lmms-lab-encoder /COCO-Caption2017 Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2017-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption2017.image10K<n<100K24 likes3.2k downloads3y agoHugging Face07visheratin /laion-coco-nllb LAION COCO translated into 200 languages This dataset contains the samples of the LAION-COCO dataset translated to 200 languages using the largest NLLB-200 model (3.3B parameters). Fields description id - unique ID of the image. url - original URL of the image from the LAION-COCO dataset. eng_caption - original English caption from the LAION-COCO dataset. captions - a list of captions translated to the languages from the Flores 200 dataset. Every item in the list is a… See the full description on the dataset page: https://huggingface.co/datasets/visheratin/laion-coco-nllb.imageimage-to-text100K<n<1M45 likes2.6k downloads2y agoHugging Face08sayakpaul /coco-30-val-2014 Dataset Card for "coco-30-val-2014" This is 30k randomly sampled image-captioned pairs from the COCO 2014 val split. This is useful for image generation benchmarks (FID, CLIPScore, etc.). Refer to the gist to know how the dataset was created: https://gist.github.com/sayakpaul/0c4435a1df6eb6193f824f9198cabaa5. image10K<n<100K14 likes2.3k downloads3y agoHugging Face09bitmind /MS-COCOimage100K<n<1M3 likes2.2k downloads2y agoHugging Face10phiyodr /coco2017 coco2017 Image-text pairs from MS COCO2017. Data origin Data originates from cocodataset.org While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy. phiyodr/coco2017: One row corresponds one image with several sentences. phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.imageimage-to-text100K<n<1M29 likes2k downloads3y agoHugging Face11justram /COCO2014-Images Dataset Card for "COCO2014-Images" More Information needed image100K<n<1M2 likes1.7k downloads3y agoHugging Face12limingcv /Captioned_COCOStuffimage100K<n<1M2 likes1.4k downloads3y agoHugging Face13Multimodal-Fatima /COCO_captions_train Dataset Card for "COCO_captions_train" More Information needed image100K<n<1M7 likes1.4k downloads4y agoHugging Face14GATE-engine /COCOStuff164K Dataset Card for "COCOStuff164K" More Information needed image100K<n<1M1 likes1.3k downloads3y agoHugging Face15AbdoTW /COCO_2014image100K<n<1M4 likes1.3k downloads2y agoHugging Face16finedet /coco2017 FineCoco — COCO 2017 detection in the unified detection format Source: official COCO 2017 train/val zips from images.cocodataset.org. Converted by the finedet project into a unified, AutoTrain-compatible layout: image / width / height / objects{bbox, category} with COCO-format [x, y, w, h] boxes in absolute pixels. Boxes are clipped to the image and empty boxes dropped; category ids are densified per the table below. Box format objects.bbox follows the COCO… See the full description on the dataset page: https://huggingface.co/datasets/finedet/coco2017.imageobject-detection100K<n<1M0 likes1.1k downloads1mo agoHugging Face17henryscheible /coco_val2014_blip2_processed Dataset Card for "coco_val2014_blip2_processed" More Information needed text10K<n<100K0 likes960 downloads3y agoHugging Face18Multimodal-Fatima /COCO_captions_validation Dataset Card for "COCO_captions_validation" More Information needed image1K<n<10K0 likes862 downloads4y agoHugging Face19NaiveDev /coco-2014-instance Dataset Card for "coco-2014-instance" More Information needed imageobject-detection100K<n<1M3 likes839 downloads3y agoHugging Face20harsh-7070 /COCO-Wholebody-annotatedimage100K<n<1M0 likes740 downloads2y agoHugging Face21maelic /VG150-coco-format VG150 — Visual Genome 150 (COCO format) This dataset is the standard VG150 split of Visual Genome (Krishna et al., 2017), the most widely used benchmark for Scene Graph Generation, reformatted in standard COCO-JSON format. VG150 contains the top 150 object categories and 50 relations from the original Visual Genome dataset, selected by frequency in the Scene Graph Generation by Iterative Message Passing paper. This version in COCO format was produced as part of the… See the full description on the dataset page: https://huggingface.co/datasets/maelic/VG150-coco-format.imageobject-detection100K<n<1M0 likes735 downloads2mo agoHugging Face22KhangTruong /COCO-inpaintedimage100K<n<1M0 likes598 downloads16d agoHugging Face23BBVisual /CoCount-train-2image100K<n<1M0 likes571 downloads1y agoHugging Face24savoji /coco-paligemmaimage100K<n<1M0 likes566 downloads1y agoHugging Face25gokuls /processed_train_coco Dataset Card for "processed_train_coco" More Information needed 100K<n<1M0 likes562 downloads3y agoHugging Face26srishti-kaushik /COCO-2017 MS COCO 2017 The complete COCO 2017 release — all four image splits and every annotation family — in Parquet. from datasets import load_dataset # annotations only (~6 MB) -- no image bytes ann = load_dataset("srishti-kaushik/COCO-2017", "instances_val2017", split="train") # images img = load_dataset("srishti-kaushik/COCO-2017", "images_val2017", split="train") img[0]["image"] # PIL.Image # stream, instead of downloading 19 GB train =… See the full description on the dataset page: https://huggingface.co/datasets/srishti-kaushik/COCO-2017.imageobject-detection1M<n<10M1 likes548 downloads1mo agoHugging Face27PursuitOfDataScience /llama4-maverick-coco-captionsimage100K<n<1M0 likes540 downloads1y agoHugging Face28NaiveDev /coco-2017-instance Dataset Card for "coco-2017-instance" More Information needed image100K<n<1M0 likes532 downloads3y agoHugging Face29Jiwon-Kang /sd2.1_cocotext10K<n<100K0 likes527 downloads2y agoHugging Face30laicsiifes /coco-captions-pt-br 🎉 COCO Captions Dataset Translation for Portuguese Image Captioning 💾 Dataset Summary COCO Captions Portuguese Translation, a multimodal dataset for Portuguese image captioning with 123,287 images, each accompanied by five descriptive captions that have been generated by human annotators for every individual image. The original English captions were rendered into Portuguese through the utilization of the Google Translator API. 🧑‍💻 Hot to Get… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/coco-captions-pt-br.imagetext-to-image100K<n<1M6 likes480 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.