CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gplsi /cocoterosCOCOTEROS Dataset V1.1 Dataset Summary: The COCOTEROS dataset is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context which serves as the co-text of the keywords provided. This makes COCOTEROS suitable for tasks where the generated text needs to be related both to a set of specific… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros.texttext-generation1K<n<10K0 likes860 downloads11mo agoHugging Face02baobab-trees /bao-val-coco-rating-cap Dataset Summary This dataset contains Japanese captions for COCO images and their English translations.The format is CSV. Dataset Structure Data Fields The data fields are the same among all lines. filename(str): The name of the COCO image file chatgpt text(str): The text generated by gpt-5-pro gemini text(str): The text generated by gemini-3-pro-preview grok text(str): The text generated by grok-4 claude text(str): The text generated by… See the full description on the dataset page: https://huggingface.co/datasets/baobab-trees/bao-val-coco-rating-cap.documentimage-to-text1K<n<10K0 likes102 downloads7mo agoHugging Face03gbyuvd /coconut-chembl34-selfies-mlm Dataset Card for COCONUT+ChemBL34 SELFIES for MLM training (unmasked) This dataset is a collection of molecular structures represented as SELFIES (Self-Referencing Embedded Strings), created by combining and processing data from COCONUTDB and ChemBL34. It contains 2,700,462 unique molecules across 13 chunks. The dataset is specifically designed for pre-training language models on molecular representations using the Masked Language Model (MLM) approach. It consists of a single column… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-selfies-mlm.text100K<n<1M0 likes83 downloads1y agoHugging Face04gbyuvd /coconut-chembl34-mol-sim Dataset Card for COCONUT+ChemBL34 SELFIES for Sentence Similarity Training This dataset is a collection of generated molecular structures pairs represented as SELFIES (Self-Referencing Embedded Strings) with MACCS fingerprint cosine similarity as label, created by combining and processing data from COCONUTDB and ChemBL34. It contains 7M pairs in total with ~5M trainable across 6 chunks. The dataset is specifically designed for fine-tuning a sentence transformer for similarity task… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-mol-sim.textsentence-similarity1M<n<10M0 likes65 downloads1y agoHugging Face05gplsi /cocoteros_vagatedCOCOTEROS_VA Dataset Dataset Summary: The COCOTEROS_VA dataset is a translation of the COCOTEROS dataset, carried out by a linguist specialized in Valencian. It is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context, which serves as the co-text of the keywords provided. This makes… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros_va.texttext-generationn<1K0 likes48 downloads11mo agoHugging Face06AnonymousUser2026 /ms_marco_cocondensertabular1K<n<10K0 likes45 downloads11mo agoHugging Face07hoveringgull /COCO_GridQA COCO-GridQA Dataset Overview The COCO-GridQA dataset is a derived dataset created from the COCO (Common Objects in Context) validation set. It focuses on spatial reasoning tasks by arranging object crops from COCO images into a 2x2 grid and providing question-answer pairs about the positions of objects within the grid. This dataset is designed for tasks such as spatial reasoning, visual question answering (VQA), and object localization. Each sample consists of: A… See the full description on the dataset page: https://huggingface.co/datasets/hoveringgull/COCO_GridQA.textquestion-answering1K<n<10K0 likes29 downloads2y agoHugging Face08vuiseng9 /coco2014-val-30K-256x256tabular10K<n<100K0 likes22 downloads1y agoHugging Face09CocoaRain /common_voice_13_0_zh_pseudo_labelledtext1K<n<10K0 likes21 downloads3y agoHugging Face10baobab-trees /baobab_coco_evaluate_caption_24 Dataset Summary This dataset contains Japanese captions for COCO images and their English translations.The format is CSV. Dataset Structure Data Fields The data fields are the same among all lines. filename(str): The name of the COCO image file chatgpt text(str): The text generated by gpt-4o gemini text(str): The text generated by gemini-1.5-pro claude text(str): The text generated by claude-3.5-sonnet-20240620 llama text(str): The text generated by… See the full description on the dataset page: https://huggingface.co/datasets/baobab-trees/baobab_coco_evaluate_caption_24.textn<1K0 likes16 downloads2y agoHugging Face11changwangss /coco2017text100K<n<1M0 likes14 downloads1mo agoHugging Face12tuskanny /kannolo-msmarco-cocondensertabular10K<n<100K0 likes5 downloads4mo agoHugging Face13stellanwu /COCO_VAL_lighttext1K<n<10K0 likes4 downloads2y agoHugging Face14sarahpann /coco-spanishtext1K<n<10K0 likes4 downloads10mo agoHugging Face15stellanwu /cocoval_colortext1K<n<10K0 likes2 downloads2y agoHugging Face16cocoBart /embedded_faqs_medicaretabularn<1K0 likes1 downloads2y agoHugging Face17JieFeng-UCSD /COCO_VAL_2B_8cattabular1K<n<10K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.