datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
internvl3-2b-coco-apgd-eps8coco2014val_10k
COCO-2014 Val 10K (256×256)
A curated subset of the COCO 2014 Validation set containing 9,986 image-caption pairs, designed as a standard reference benchmark for FID (Fréchet Inception Distance) evaluation in text-to-image generation research.
Dataset Summary
Attribute
Value
Source
COCO 2014 Validation Split
Num Samples
9,986
Resolution
256 × 256 (center-cropped & resized)
Image Format
PNG, RGB
Total Size
~1.2 GB
Random Seed
42
Original Pool
40,504… See the full description on the dataset page: https://huggingface.co/datasets/byliutao/coco2014val_10k.internvl3-2b-coco-apgd-eps4WebUI-COCO-876
WebUI-COCO-876: Annotated Webpage Screenshots for UI Region & Landmark Detection
Short description
876 desktop webpage screenshots annotated with COCO bounding boxes over ARIA-style landmark regions and common UI components (cookie dialogs, popovers, captcha, buttons, etc.)
Dataset details
Modality: image (webpage screenshots)
Annotation format: COCO-style JSON (bounding boxes + class IDs, optional semantic attributes)
Number of images: 876
Label… See the full description on the dataset page: https://huggingface.co/datasets/jileklu/WebUI-COCO-876.coco_captions_quintets
Dataset Card for "coco_captions"
Dataset Summary
COCO is a large-scale object detection, segmentation, and captioning dataset. This repo contains five captions per image; useful for sentence similarity tasks.
Disclaimer: The team releasing COCO did not upload the dataset to the Hub and did not write a dataset card.
These steps were done by the Hugging Face team.
Supported Tasks
Sentence Transformers training; useful for semantic search and sentence… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/coco_captions_quintets.COCOPatch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
[🔗 Released Code]
[🤗 Datasets] [🤗 Checkpoints]
[📄 Tech Report] [🤗 Paper]
Figure A. PaDT pipeline.
🌟 Introduction
We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent… See the full description on the dataset page: https://huggingface.co/datasets/PaDT-MLLM/COCO.coco-remcoco2017
COCO 2017 LibreYOLO Assets
This dataset repository hosts small COCO 2017 helper assets used by LibreYOLO dataset YAML files.
Files:
instances_val2017.json: official COCO 2017 validation instance annotations.
coco2017labels-libreyolo.zip: YOLO-format train/val label files generated by LibreYOLO from the official COCO 2017 instance annotations.
No COCO images are included here. LibreYOLO downloads COCO images from the official COCO image hosts.
License and… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/coco2017.json-coco-format
JSON COCO Format — task-differentiated SFT data
A multi-task supervised fine-tuning dataset that teaches a model to convert
image-synthesis caption prompts into JSON whose structure varies by task.
Built from MS-COCO captions (Karpathy split) with Claude Sonnet 4.6 as the
teacher; designed for training per-task LoRAs on
Qwen/Qwen3.5-0.8B.
Each row is in the Qwen3.5-native tool-call shape: a messages array with an
assistant turn whose tool_calls[0].function.arguments is a dict… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/json-coco-format.coco-valCoConflictQACoConflictQA is a benchmark designed to evaluate the contextual faithfulness of Large Language Models (LLMs) by focusing on their tendency to hallucinate during question answering. It aims to provide a more reliable assessment of how well LLMs align their responses with the given context.
This dataset is constructed based on six widely-used QA datasets:
HotpotQA
NewsQA
Natural Questions (NQ)
SearchQA
SQuAD
TriviaQA
CoConflictQA was introduced in the paper:
PIP-KAG: Mitigating Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/chengpingan/CoConflictQA.coco-arvqa
COCO-ARVQA: Arabic Visual Question Answering over COCO 2017
Dataset Summary
COCO-ARVQA is an Arabic Visual Question Answering dataset built over images from MS COCO 2017 train2017.It provides Arabic questions, Arabic answers, answer lists, question identifiers, image identifiers, and COCO image file names.
This repository does not redistribute COCO images. Both the training and validation splits reference images from the official COCO 2017 train2017.zip archive.
Official… See the full description on the dataset page: https://huggingface.co/datasets/MouaffakAyoub/coco-arvqa.coco-karpathy-opus-de
Dataset Card for MS COCO Karpathy in German language
This dataset contains captions that were machine translated using opus-mt-en-de.
Dataset Details
Dataset Sources
The processed MS COCO datasets (Karpathy Split) in this repo are based on the following sources:
Type
MD5
URL
Train
aa31ac474cf6250ebb81d18348a07ed8
https://storage.googleapis.com/sfr-vision-language-research/datasets/coco_karpathy_train.json
Validation
b273847456ef5580e33713b1f7de52a0… See the full description on the dataset page: https://huggingface.co/datasets/Jotschi/coco-karpathy-opus-de.CocoScisum_ACLEpilepsy_Synthetics
Epilepsy_Syntheics
This is a cross-languadge dataset for epilepsy-care, support both madarin and english.
It is generated by Qwen 1.5(For mandarin) and LLAMA-3(For English) with the use of self-instruct method.
This dataset contains 1K+1K epilepsy-care data. And it have already been splitted and cleaned.
Have fun and enjoy!
CoCoNUTS
1 Introduction
Existing AI-generated text detectors often fail in the academic peer review context because they rely on stylistic cues, which leads to misclassifying permissibly polished text and missing cleverly paraphrased AI content. To address this, we propose a paradigm shift from style-based to content-based detection.
We introduce CoCoNUTS, a comprehensive benchmark for this task. It is built upon a fine-grained dataset of academic peer reviews, covering six distinct modes of… See the full description on the dataset page: https://huggingface.co/datasets/khaaaaaan/CoCoNUTS.coco_2017_caption_trainCOCO-AB
General Information
Title: COCO-AB
Description:
The COCO-AB dataset is an extension of the COCO 2014 training set, enriched with additional annotation byproducts (AB).
The data includes 82,765 reannotated images from the original COCO 2014 training set.
It has relevance in computer vision, specifically in object detection and location.
The aim of the dataset is to provide a richer understanding of the images (without extra costs) by recording additional actions and interactions… See the full description on the dataset page: https://huggingface.co/datasets/coallaoh/COCO-AB.cocomath-300kcoco-deceptive-clip-llama3.1-8b
COCO-Deceptive-CLIP-LLaMA-3.1-8B Training Dataset
🏆 This work is accepted to ACL 2025 (Main Conference).
Figure: Attack success rate (ASR) and caption diversity of our model on the COCO dataset, illustrating its ability to generate deceptive captions that successfully fool CLIP.
Dataset Details
This dataset provides instruction–response pairs formatted as short two-turn conversations:
The user message contains:
A given image caption.
A set of task… See the full description on the dataset page: https://huggingface.co/datasets/ahnpersie/coco-deceptive-clip-llama3.1-8b.Filtered-COCO-Captions
Dataset Summary
This dataset is derived from the MS COCO caption annotations.
Source
Original annotations: MS COCO / COCO Consortium
License
The original annotation set is licensed under CC BY 4.0.
This repository redistributes a filtered/adapted version of the annotation text only.
No original COCO images are included.
Modifications
Removed captions deemed unsuitable for TOEIC educational materials
Normalized punctuation and whitespace
Filtered for… See the full description on the dataset page: https://huggingface.co/datasets/kknono668/Filtered-COCO-Captions.COCOPanoptic_overCOCOTree
COCOTree
Code
COCOTree is an annotation-only release for open tree decomposition over
COCO images. Each image has two linked views: a semantic-node tree for local
labels and merged masks, and an instance-node tree for image-local mask
instances and visual parent-child links.
The full original COCO images are not redistributed here. This repository
includes the released annotations, metadata, validation files, and a small
sample folder for inspection.
Click either figure to open the… See the full description on the dataset page: https://huggingface.co/datasets/melonkick/COCOTree.coco_2017_caption_validationCOCO
Dataset Card for [COCO]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/Luka-Wang/COCO.COCO3D
COCO3D
3D bounding box annotations for COCO images, produced by
LabelAny3D.
Split
Images
Annotations
Categories
val
2,010
5,409
80
train
15,869
86,395
80
Format follows Omni3D.
License
CC BY 4.0 covers our 3D annotations only. Images are not redistributed here —
they are referenced by COCO image id and file path, and remain under the
original COCO terms of use.
Citation
@inproceedings{yao2025labelany3d,
title={LabelAny3D: Label… See the full description on the dataset page: https://huggingface.co/datasets/uva-cv-lab/COCO3D.COCOStuff_overLLaVA-Instruct-21K-COCO-SubSet
subset from https://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K
train: 21000
val seen: 3000
val unseen: 2100
test: 6000
ProSR-Datacoco-panoptic-categories
