datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coco-2017-mirror
COCO 2017 mirror
This is a just mirror of the raw COCO dataset files, for convenience. You have to download it using something like:
pip install huggingface_hub
huggingface-cli download --local-dir coco-2017 pcuenq/coco-2017-mirror
And then unzip the files before use.
cocoCOCOANet
COCOANet
A high-fidelity CFD dataset of parametric (CCA) Aircraft geometries and flight condition for aerodynamic surrogate modelling and Aerodynamic Shape Optimization, created as part of ShapeBench.
Dataset Summary
Geometries
401
CFD runs
3,570
Design parameters
16 geometric
Flight parameters
3 (angle of attack, velocity, altitude)
CAD Kernal
nTop
Solver
Flow360
Sampling
Latin Hypercube Sampling (LHS), seed 42
Files… See the full description on the dataset page: https://huggingface.co/datasets/ShapeBench/COCOANet.fixtures-cocolingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/cocool/lingshu_training_data_medical_domain.synthetic_data_v5_finegrain_layout_relight_with_our_synthetic_data_coco_l_full_500kcoconot
🥥 CoCoNot: Contextually, Comply Not! Dataset Card
Dataset Details
Dataset Description
Chat-based language models are designed to be helpful, yet they should not comply with every user request.
While most existing work primarily focuses on refusal of "unsafe" queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user… See the full description on the dataset page: https://huggingface.co/datasets/allenai/coconot.coco-karpathy
Dataset Card for "yerevann/coco-karpathy"
The Karpathy split of COCO for image captioning.
cocoCOCO-Caption
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of COCO-Caption-2014-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@misc{lin2015microsoft,
title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption.coco_captions
Dataset Card for "coco_captions"
More Information needed
COCO-Caption2017
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of COCO-Caption-2017-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@misc{lin2015microsoft,
title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption2017.sd15_coco_train_seed123laion-coco-nllb
LAION COCO translated into 200 languages
This dataset contains the samples of the LAION-COCO dataset translated to 200 languages using
the largest NLLB-200 model (3.3B parameters).
Fields description
id - unique ID of the image.
url - original URL of the image from the LAION-COCO dataset.
eng_caption - original English caption from the LAION-COCO dataset.
captions - a list of captions translated to the languages from the Flores 200 dataset. Every item in the list is a… See the full description on the dataset page: https://huggingface.co/datasets/visheratin/laion-coco-nllb.MS-COCOcoco-30-val-2014
Dataset Card for "coco-30-val-2014"
This is 30k randomly sampled image-captioned pairs from the COCO 2014 val split. This is useful for image generation benchmarks (FID, CLIPScore, etc.).
Refer to the gist to know how the dataset was created: https://gist.github.com/sayakpaul/0c4435a1df6eb6193f824f9198cabaa5.
COCOMS COCO is a large-scale object detection, segmentation, and captioning dataset.
COCO has several features: Object segmentation, Recognition in context, Superpixel stuff segmentation, 330K images (>200K labeled), 1.5 million object instances, 80 object categories, 91 stuff categories, 5 captions per image, 250,000 people with keypoints.coco2017
coco2017
Image-text pairs from MS COCO2017.
Data origin
Data originates from cocodataset.org
While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy.
phiyodr/coco2017: One row corresponds one image with several sentences.
phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.cc12m-wds-coco-recaptioned
CC12M WebDataset with COCO-style Recaptions
A large-scale image-text dataset containing 3 million images from Conceptual Captions 12M (CC12M) with COCO-style factual descriptions generated using NVIDIA Nemotron Nano 12B v2 VL.
Dataset Overview
Base Dataset: pixparse/cc12m-wds - Conceptual Captions 12M (CC12M)
Images: 3,000,000+ high-quality internet images
Recaption Model: NVIDIA Nemotron Nano 12B v2 VL
Recaption Style: COCO-style factual descriptions (20 words average)… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-coco-recaptioned.srt-coco-thumbsUCS-Bench
UCS-Bench
ICML 2026 Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams
UCS-Bench is a benchmark for evaluating user-centric continual spatial intelligence in streaming egocentric videos.
The dataset contains 170+ hours of egocentric visual observations and 8.1K+ timestamped questions. It is designed to test whether models can understand dynamic spatial environments from a user's point of view, maintain long-term… See the full description on the dataset page: https://huggingface.co/datasets/cocowy1/UCS-Bench.COCO2014-Images
Dataset Card for "COCO2014-Images"
More Information needed
DensePose-COCO
Dataset Card for DensePose-COCO
DensePose-COCO is a large-scale ground-truth dataset with image-to-surface correspondences manually annotated on COCO images.
This is a FiftyOne dataset with 33929 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/DensePose-COCO.CocoChorales-E
Viewer note: default uses viewer_preview/ for responsive audio playback.
Full training/evaluation files remain available in the original folder structure.
CocoChorales-E
CocoChorales-E subset used by the LadderSym training pipeline.
Paired Inputs for Error Detection
The model takes paired inputs:
mistake: performance audio/MIDI containing musical errors
score: paired reference score audio/MIDI (target/correct context)
Error supervision is provided with labels:… See the full description on the dataset page: https://huggingface.co/datasets/ben2002chou/CocoChorales-E.coco2017Captioned_COCOStuffcoco2017This dataset contains all COCO 2017 images and annotations split in training (118287 images) and validation (5000 images).COCO_2014my-cocoCOCOStuff164K
Dataset Card for "COCOStuff164K"
More Information needed
