datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
getting-started-validation-clip-pred
Dataset Card for labeled_validation_predicted_clip
This is a FiftyOne dataset with 143 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("TheSteve0/getting-started-validation-clip-pred")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/getting-started-validation-clip-pred.CLIP-FMoE-Evaluation
CLIP-FMoE Evaluation Data
Evaluation datasets used by the CLIP-FMoE repository.
Large raw image directories are stored as uncompressed .tar files. This avoids uploading millions of individual image files and makes download/extraction substantially faster.
Repository layout
clip_benchmark/wds_<dataset>/ Existing CLIP_benchmark WebDataset shards
retrieval/docci_iiw/ Metadata + docci_arr_new/images_aar.tar
retrieval/dci/ Annotations +… See the full description on the dataset page: https://huggingface.co/datasets/moneyzz432/CLIP-FMoE-Evaluation.military-labeled-clip
Military-Labeled CLIP Crops (DVIDS sourced)
Per-object crops extracted from DVIDS military imagery, each accompanied by a
Gemini-VLM caption suitable for CLIP fine-tuning or zero-shot evaluation.
Files
crops/dvids_image_{id}_{class}_{idx}.jpg — 4,844 cropped objects
captions.jsonl — per-crop metadata: {crop_path, image_id, class_name, class_id, bbox, caption, image_caption, branch, source, source_url}
Classes (12) — Distribution
Class
Crops… See the full description on the dataset page: https://huggingface.co/datasets/AMANMP0007/military-labeled-clip.clide_synthetic_datasets
CLIDE Synthetic Image Datasets
📄 Paper • 💻 Code • 🌐 Webpage • 🎥 Video
A collection of synthetic images generated by modern text-to-image models, organized by domain and generator.
The dataset is designed to support analysis and evaluation of generated-image detection methods under domain and generator shifts.
🗂️ Dataset Structure
The dataset contains two visual domains:
💥🚗 Damaged Cars
Synthetic images of damaged cars generated by multiple… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/clide_synthetic_datasets.military-labeled-clip
Military-Labeled CLIP Crops (DVIDS sourced)
Per-object crops extracted from DVIDS military imagery, each accompanied by a
Gemini-VLM caption suitable for CLIP fine-tuning or zero-shot evaluation.
Files
crops/dvids_image_{id}_{class}_{idx}.jpg — 4,844 cropped objects
captions.jsonl — per-crop metadata: {crop_path, image_id, class_name, class_id, bbox, caption, image_caption, branch, source, source_url}
Classes (12) — Distribution
Class
Crops… See the full description on the dataset page: https://huggingface.co/datasets/llama-farm/military-labeled-clip.military-labeled-clip
Military-Labeled CLIP Crops (DVIDS sourced)
Per-object crops extracted from DVIDS military imagery, each accompanied by a
Gemini-VLM caption suitable for CLIP fine-tuning or zero-shot evaluation.
Files
crops/dvids_image_{id}_{class}_{idx}.jpg — 4,844 cropped objects
captions.jsonl — per-crop metadata: {crop_path, image_id, class_name, class_id, bbox, caption, image_caption, branch, source, source_url}
Classes (12) — Distribution
Class
Crops… See the full description on the dataset page: https://huggingface.co/datasets/p14ton/military-labeled-clip.clip-cues-artifacts
CLIP-Cues — frozen input snapshot
Cached CLIP features and the checksummed inputs behind
"Synthetic Image Detection with CLIP: Understanding and Assessing Predictive Cues"
(Willi, Mathys & Graber). This repository holds no images: it is the set of frozen arrays that lets
every number and figure in the paper be reproduced on a CPU in minutes, without re-extracting
features from 33k photographs.
Code: https://github.com/marco-willi/clip-cues · Image datasets:
synthclic ·… See the full description on the dataset page: https://huggingface.co/datasets/marco-willi/clip-cues-artifacts.clinical-attire-labels
clinical-attire-labels
4,000 Places365 images labelled for clinical attire and role. Every label is machine-generated.
scrubs · surgical_gown · patient_gown · lab_coat · street · mask · none
plus derived roles: role_staff · role_patient · role_visitor · ppe_mask
Why this exists
There is no public labelled dataset for medical scrubs, hospital gowns or lab coats. Fashionpedia
has none of them among its 46 garment categories; dataset searches return face-mask corpora… See the full description on the dataset page: https://huggingface.co/datasets/resoajoe/clinical-attire-labels.imagenet-1k-224-clip-embeddings
ImageNet-1k-224 CLIP Embeddings
Pre-computed CLIP image embeddings for every image in
mlnomad/imagenet-1k-224.
Columns
Column
Type
Description
original_index
int
Row index in the source dataset for cross-referencing
label
int (0–999)
ImageNet class index
embedding
List[float]
L2-normalised CLIP image embedding (768D)
Stats
Source: mlnomad/imagenet-1k-224 (train split)
Total images: 1281167
Embedding dim: 768
CLIP model:… See the full description on the dataset page: https://huggingface.co/datasets/mlnomad/imagenet-1k-224-clip-embeddings.movies_CLIP_ViT-L14
🎬 Movie Frame & Caption Dataset
📖 Introduction
This dataset was created from multiple movies across 10 genres, with approximately 3 movies per genre.From each movie, frames were extracted periodically, and AI-generated captions (BLIP) were assigned to each frame.A total of 93,813 frames were extracted.
This dataset can be used for tasks such as:
Video understanding
Multimodal learning (image + text)
Image captioning
Vision-language retrieval
📂 Data… See the full description on the dataset page: https://huggingface.co/datasets/thaotien/movies_CLIP_ViT-L14.coco-clip-vit-l-14
COCO Dataset Processed with CLIP ViT-L/14
Overview
This dataset represents a processed version of the '2017 Unlabeled images' subset of the COCO dataset (COCO Dataset), utilizing the CLIP ViT-L/14 model from OpenAI. The original dataset comprises 123K images, approximately 19GB in size, which have been processed to generate 786-dimensional vectors. These vectors can be utilized for various applications like semantic search systems, image similarity assessments, and more.… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/coco-clip-vit-l-14.urban-climate-green-infrastructure
Urban Climate & Green Infrastructure Visual Dataset
Rows: 18,107
Dataset Description
Urban Climate & Green Infrastructure Visual Dataset is a global wildlife image dataset and geospatial computer vision dataset focused on street-level imagery of city features that support lower-carbon and more resilient planning. The labels cover Street Tree, Bike Lane, Solar Panel, EV Charger, Rain Garden, and Green Roof, produced through Outerview's query-driven embedding… See the full description on the dataset page: https://huggingface.co/datasets/Outerview/urban-climate-green-infrastructure.Agricultural-Climate-Adaptation-Research-Dataset
Agricultural Climate Adaptation Research Dataset
Agriculture is currently facing challenges posed by climate change, particularly the increasing impact of drought on crop yields. Existing research data often lacks detailed analysis under specific climate conditions, leading to ineffective agricultural management measures. This dataset aims to fill this gap by including images of farmland drought and vegetation recovery, assisting AI models in researching agriculture's ability to… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Agricultural-Climate-Adaptation-Research-Dataset.clip-insect-sex-data
Gryllus bimaculatus Insect Sex Dataset
Image dataset for binary classification: male and female.
Species
Gryllus bimaculatus
Structure
augmented_data/
male/
female/
Labels
male
female
Source and curation
Images were captured by the dataset owner.
Augmented variants were generated from owner-captured source images for training.
Access and permission terms
This dataset is shared for viewing/research reference.
Reuse… See the full description on the dataset page: https://huggingface.co/datasets/yashm/clip-insect-sex-data.cifar100_clipcifar10_clipOD-CLIP
OD-CLIP Training Dataset
Training data for OD-CLIP: a degradation-aware CLIP variant that jointly
predicts degradation type and LPIPS-calibrated perceptual severity, used for
blind image super-resolution.
The dataset provides paired ground-truth (GT) and low-quality (LQ) image
crops together with per-image degradation metadata, covering four synthetic
degradation types (Gaussian blur, Gaussian noise, JPEG compression, and
downsampling) at a dense grid of physical severity… See the full description on the dataset page: https://huggingface.co/datasets/yeeecheng/OD-CLIP.cliptrace-baseline-data
CLIPTrace 2026 Baseline Data
Participant data for the reproducible CLIPTrace 2026 baseline. The repository
mirrors the full-size Imagenette train and validation images used by the
baseline and preserves the original class-directory layout.
This repository is intended to be public with access gating. Before
downloading, users must accept the repository terms and the upstream image-use
conditions configured by the organizers on the Hugging Face settings page.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/cliptrace-2026/cliptrace-baseline-data.
