CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /fixtures-cocoimagen<1K0 likes15k downloads19d agoHugging Face02allenai /coconot 🥥 CoCoNot: Contextually, Comply Not! Dataset Card Dataset Details Dataset Description Chat-based language models are designed to be helpful, yet they should not comply with every user request. While most existing work primarily focuses on refusal of "unsafe" queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user… See the full description on the dataset page: https://huggingface.co/datasets/allenai/coconot.texttext-generation10K<n<100K25 likes6.2k downloads2y agoHugging Face03yerevann /coco-karpathy Dataset Card for "yerevann/coco-karpathy" The Karpathy split of COCO for image captioning. imageimage-to-text100K<n<1M22 likes5.3k downloads4y agoHugging Face04lmms-lab-encoder /COCO-Caption Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2014-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption.image10K<n<100K15 likes4.4k downloads3y agoHugging Face05jxie /coco_captions Dataset Card for "coco_captions" More Information needed image100K<n<1M18 likes4.2k downloads3y agoHugging Face06lmms-lab-encoder /COCO-Caption2017 Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2017-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption2017.image10K<n<100K24 likes3.4k downloads3y agoHugging Face07visheratin /laion-coco-nllb LAION COCO translated into 200 languages This dataset contains the samples of the LAION-COCO dataset translated to 200 languages using the largest NLLB-200 model (3.3B parameters). Fields description id - unique ID of the image. url - original URL of the image from the LAION-COCO dataset. eng_caption - original English caption from the LAION-COCO dataset. captions - a list of captions translated to the languages from the Flores 200 dataset. Every item in the list is a… See the full description on the dataset page: https://huggingface.co/datasets/visheratin/laion-coco-nllb.imageimage-to-text100K<n<1M45 likes2.6k downloads2y agoHugging Face08bitmind /MS-COCOimage100K<n<1M3 likes2.4k downloads2y agoHugging Face09sayakpaul /coco-30-val-2014 Dataset Card for "coco-30-val-2014" This is 30k randomly sampled image-captioned pairs from the COCO 2014 val split. This is useful for image generation benchmarks (FID, CLIPScore, etc.). Refer to the gist to know how the dataset was created: https://gist.github.com/sayakpaul/0c4435a1df6eb6193f824f9198cabaa5. image10K<n<100K14 likes2.3k downloads3y agoHugging Face10phiyodr /coco2017 coco2017 Image-text pairs from MS COCO2017. Data origin Data originates from cocodataset.org While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy. phiyodr/coco2017: One row corresponds one image with several sentences. phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.imageimage-to-text100K<n<1M27 likes2.1k downloads3y agoHugging Face11undefined443 /cc12m-wds-coco-recaptioned CC12M WebDataset with COCO-style Recaptions A large-scale image-text dataset containing 3 million images from Conceptual Captions 12M (CC12M) with COCO-style factual descriptions generated using NVIDIA Nemotron Nano 12B v2 VL. Dataset Overview Base Dataset: pixparse/cc12m-wds - Conceptual Captions 12M (CC12M) Images: 3,000,000+ high-quality internet images Recaption Model: NVIDIA Nemotron Nano 12B v2 VL Recaption Style: COCO-style factual descriptions (20 words average)… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-coco-recaptioned.image1M<n<10M1 likes1.9k downloads5mo agoHugging Face12ben2002chou /CocoChorales-E Viewer note: default uses viewer_preview/ for responsive audio playback. Full training/evaluation files remain available in the original folder structure. CocoChorales-E CocoChorales-E subset used by the LadderSym training pipeline. Paired Inputs for Error Detection The model takes paired inputs: mistake: performance audio/MIDI containing musical errors score: paired reference score audio/MIDI (target/correct context) Error supervision is provided with labels:… See the full description on the dataset page: https://huggingface.co/datasets/ben2002chou/CocoChorales-E.audioaudio-classificationn<1K1 likes1.6k downloads7mo agoHugging Face13limingcv /Captioned_COCOStuffimage100K<n<1M2 likes1.5k downloads3y agoHugging Face14rafaelpadilla /coco2017This dataset contains all COCO 2017 images and annotations split in training (118287 images) and validation (5000 images).imageobject-detection100K<n<1M32 likes1.4k downloads3y agoHugging Face15Changyeli03 /AA_preference_cocourimage10K<n<100K0 likes1.3k downloads2y agoHugging Face16AbdoTW /COCO_2014image100K<n<1M4 likes1.2k downloads2y agoHugging Face17Multimodal-Fatima /COCO_captions_train Dataset Card for "COCO_captions_train" More Information needed image100K<n<1M7 likes1.2k downloads4y agoHugging Face18henryscheible /coco_val2014_blip2_processed Dataset Card for "coco_val2014_blip2_processed" More Information needed text10K<n<100K0 likes953 downloads3y agoHugging Face19Fhrozen /coco-narratives COCO Narratives Original Source | Google Localized Narrative 📌 Introduction This dataset collects the images and annotations from the original MS COCO 2017 and the annotations from the project localized-narratives 🙏 Acknowledgement All credits to the original COCO project and the localized-narratives teams. 📜 Cite Please consider to cite the following related papers: @article{DBLP:journals/corr/LinMBHPRDZ14, author = {Tsung{-}Yi Lin and… See the full description on the dataset page: https://huggingface.co/datasets/Fhrozen/coco-narratives.imageimage-text-to-text100K<n<1M0 likes917 downloads1y agoHugging Face20gplsi /cocoterosCOCOTEROS Dataset V1.1 Dataset Summary: The COCOTEROS dataset is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context which serves as the co-text of the keywords provided. This makes COCOTEROS suitable for tasks where the generated text needs to be related both to a set of specific… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros.texttext-generation1K<n<10K0 likes860 downloads11mo agoHugging Face21Multimodal-Fatima /COCO_captions_validation Dataset Card for "COCO_captions_validation" More Information needed image1K<n<10K0 likes840 downloads4y agoHugging Face22NaiveDev /coco-2014-instance Dataset Card for "coco-2014-instance" More Information needed imageobject-detection100K<n<1M3 likes832 downloads3y agoHugging Face23maelic /VG150-coco-format VG150 — Visual Genome 150 (COCO format) This dataset is the standard VG150 split of Visual Genome (Krishna et al., 2017), the most widely used benchmark for Scene Graph Generation, reformatted in standard COCO-JSON format. VG150 contains the top 150 object categories and 50 relations from the original Visual Genome dataset, selected by frequency in the Scene Graph Generation by Iterative Message Passing paper. This version in COCO format was produced as part of the… See the full description on the dataset page: https://huggingface.co/datasets/maelic/VG150-coco-format.imageobject-detection100K<n<1M0 likes780 downloads2mo agoHugging Face24harsh-7070 /COCO-Wholebody-annotatedimage100K<n<1M0 likes741 downloads2y agoHugging Face25srishti-kaushik /COCO-2017 MS COCO 2017 The complete COCO 2017 release — all four image splits and every annotation family — in Parquet. from datasets import load_dataset # annotations only (~6 MB) -- no image bytes ann = load_dataset("srishti-kaushik/COCO-2017", "instances_val2017", split="train") # images img = load_dataset("srishti-kaushik/COCO-2017", "images_val2017", split="train") img[0]["image"] # PIL.Image # stream, instead of downloading 19 GB train =… See the full description on the dataset page: https://huggingface.co/datasets/srishti-kaushik/COCO-2017.imageobject-detection1M<n<10M1 likes642 downloads1mo agoHugging Face26PursuitOfDataScience /llama4-maverick-coco-captionsimage100K<n<1M0 likes624 downloads1y agoHugging Face27savoji /coco-paligemmaimage100K<n<1M0 likes562 downloads1y agoHugging Face28BBVisual /CoCount-train-2image100K<n<1M0 likes560 downloads1y agoHugging Face29NaiveDev /coco-2017-instance Dataset Card for "coco-2017-instance" More Information needed image100K<n<1M0 likes549 downloads3y agoHugging Face30gebinhui /coco2017_caption_normalimage100K<n<1M0 likes546 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.