datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cocoterosCOCOTEROS Dataset V1.1
Dataset Summary: The COCOTEROS dataset is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context which serves as the co-text of the keywords provided. This makes COCOTEROS suitable for tasks where the generated text needs to be related both to a set of specific… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros.COCO-Text
Dataset Card for "COCO-Text"
More Information needed
coco_testcocot
CoCoT: Collaborative Cross-modal Chain-of-Thought Dataset
This repository contains the complete CoCoT (Collaborative Cross-modal Chain-of-Thought) dataset, including bounding box annotations and reasoning chains for complex visual question answering tasks.
Associated Paper: Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning; Accepted to: NeurIPS 2026 Workshop; Authors: Wenting Lu, Didi Zhu, Tao Shen, Donglin Zhu, Ayong Ye, Chao Wu… See the full description on the dataset page: https://huggingface.co/datasets/echo-deer/cocot.cocoteros_vaCOCOTEROS_VA Dataset
Dataset Summary:
The COCOTEROS_VA dataset is a translation of the COCOTEROS dataset, carried out by a linguist specialized in Valencian. It is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context, which serves as the co-text of the keywords provided. This makes… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros_va.coco-train-blip2-hpcoco-train-blip2-20best-p0.3-alpacacocotext_train_cleaned
cocotext_train_cleaned
The cocotext_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
41,200
QA turns
238,897
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
21
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/cocotext_train_cleaned.coco-train-blip2-20best-p0.5-alpacacoco-test-blip2-5best-p0.7-alpacaCOCO-Textcoco-transformed-captionsCOCOTree
COCOTree
Code
COCOTree is an annotation-only release for open tree decomposition over
COCO images. Each image has two linked views: a semantic-node tree for local
labels and merged masks, and an instance-node tree for image-local mask
instances and visual parent-child links.
The full original COCO images are not redistributed here. This repository
includes the released annotations, metadata, validation files, and a small
sample folder for inspection.
Click either figure to open the… See the full description on the dataset page: https://huggingface.co/datasets/melonkick/COCOTree.CocoText_RS_nothink
CocoText — CocoText_RS_nothink
Rejection-sampled from the CocoText train split. This split holds the accepted items, answer only.
rows
5,417
QA pairs
5,417
shards
73
accepted / rejected (whole family)
5,417 / 12,218
accept rate
30.7%
verifier
anls
How the data was produced
A VLM answers every question at temperature 0 with reasoning enabled. Its answer is compared with
the official ground truth by the verifier described below; matches… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/CocoText_RS_nothink.CocoText_longform
CocoText — CocoText_longform
Rejection-sampled from the CocoText train split. This split holds the long-form items the verifier cannot score, shipped unverified.
rows
1,380
QA pairs
1,380
shards
14
accepted / rejected (whole family)
5,417 / 12,218
accept rate
30.7%
verifier
anls
How the data was produced
A VLM answers every question at temperature 0 with reasoning enabled. Its answer is compared with
the official ground truth by the… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/CocoText_longform.COCO-Text_V2_STRCocoText_RS_think
CocoText — CocoText_RS_think
Rejection-sampled from the CocoText train split. This split holds the accepted items, with the model's reasoning trace.
rows
5,417
QA pairs
5,417
shards
74
accepted / rejected (whole family)
5,417 / 12,218
accept rate
30.7%
verifier
anls
How the data was produced
A VLM answers every question at temperature 0 with reasoning enabled. Its answer is compared with
the official ground truth by the verifier… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/CocoText_RS_think.CocoText_rejected
CocoText — CocoText_rejected
Rejection-sampled from the CocoText train split. This split holds the rejected items — the answer field holds the official ground truth.
rows
12,218
QA pairs
12,218
shards
198
accepted / rejected (whole family)
5,417 / 12,218
accept rate
30.7%
verifier
anls
The rejected split is training data, not just diagnostics: answer is the official ground truth, and wrong_vlm records what the model said instead.
How the… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/CocoText_rejected.COCO-train_dataset_balanceadococo-train-blip2-10best-p0.3-alpacacoco_train_cleaned
coco_train_cleaned
The coco_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
118,027
QA turns
985,335
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
39
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/coco_train_cleaned.coco-train-blip2-10best-p0.7-alpacacoco-train-blip2-5best-p0.5-alpacacoco-test-blip2-10best-p0.7-alpacacoco-train-blip2-10best-p0.5-alpacacoco_text_traintest2017coco-test-blip2-5best-alpacacoco-test-blip2-20best-p0.3-alpacacoco-test-blip2-20best-p0.7-alpacacoco-train-git-large-5best-alpaca
