datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cocoterosCOCOTEROS Dataset V1.1
Dataset Summary: The COCOTEROS dataset is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context which serves as the co-text of the keywords provided. This makes COCOTEROS suitable for tasks where the generated text needs to be related both to a set of specific… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros.bao-val-coco-rating-cap
Dataset Summary
This dataset contains Japanese captions for COCO images and their English translations.The format is CSV.
Dataset Structure
Data Fields
The data fields are the same among all lines.
filename(str): The name of the COCO image file
chatgpt text(str): The text generated by gpt-5-pro
gemini text(str): The text generated by gemini-3-pro-preview
grok text(str): The text generated by grok-4
claude text(str): The text generated by… See the full description on the dataset page: https://huggingface.co/datasets/baobab-trees/bao-val-coco-rating-cap.coconut-chembl34-selfies-mlm
Dataset Card for COCONUT+ChemBL34 SELFIES for MLM training (unmasked)
This dataset is a collection of molecular structures represented as SELFIES (Self-Referencing Embedded Strings), created by combining and processing data from COCONUTDB and ChemBL34. It contains 2,700,462 unique molecules across 13 chunks.
The dataset is specifically designed for pre-training language models on molecular representations using the Masked Language Model (MLM) approach. It consists of a single column… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-selfies-mlm.coconut-chembl34-mol-sim
Dataset Card for COCONUT+ChemBL34 SELFIES for Sentence Similarity Training
This dataset is a collection of generated molecular structures pairs represented as SELFIES (Self-Referencing Embedded Strings) with MACCS fingerprint cosine similarity as label, created by combining and processing data from COCONUTDB and ChemBL34. It contains 7M pairs in total with ~5M trainable across 6 chunks.
The dataset is specifically designed for fine-tuning a sentence transformer for similarity task… See the full description on the dataset page: https://huggingface.co/datasets/gbyuvd/coconut-chembl34-mol-sim.cocoteros_vaCOCOTEROS_VA Dataset
Dataset Summary:
The COCOTEROS_VA dataset is a translation of the COCOTEROS dataset, carried out by a linguist specialized in Valencian. It is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context, which serves as the co-text of the keywords provided. This makes… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros_va.ms_marco_cocondenserCOCO_GridQA
COCO-GridQA Dataset
Overview
The COCO-GridQA dataset is a derived dataset created from the COCO (Common Objects in Context) validation set. It focuses on spatial reasoning tasks by arranging object crops from COCO images into a 2x2 grid and providing question-answer pairs about the positions of objects within the grid.
This dataset is designed for tasks such as spatial reasoning, visual question answering (VQA), and object localization. Each sample consists of:
A… See the full description on the dataset page: https://huggingface.co/datasets/hoveringgull/COCO_GridQA.coco2014-val-30K-256x256common_voice_13_0_zh_pseudo_labelledbaobab_coco_evaluate_caption_24
Dataset Summary
This dataset contains Japanese captions for COCO images and their English translations.The format is CSV.
Dataset Structure
Data Fields
The data fields are the same among all lines.
filename(str): The name of the COCO image file
chatgpt text(str): The text generated by gpt-4o
gemini text(str): The text generated by gemini-1.5-pro
claude text(str): The text generated by claude-3.5-sonnet-20240620
llama text(str): The text generated by… See the full description on the dataset page: https://huggingface.co/datasets/baobab-trees/baobab_coco_evaluate_caption_24.coco2017kannolo-msmarco-cocondenserCOCO_VAL_lightcoco-spanishcocoval_colorembedded_faqs_medicareCOCO_VAL_2B_8cat
