datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-vision-banana-snvc-train-full
SNVC-50M v5_full — Multi-Task Vision Dataset
Description
This dataset is a curated subset of the SenseNova Vision Corpus 50M (SNVC-50M), containing 43,509 samples across 6 vision task families and 31 source datasets. Each sample follows a conversational format with interleaved <image> tokens, designed for training vision-language models (VLMs).
Coverage: 43,509 / 57,878 (75.2%) of the original sampling plan. 23 datasets at 100%, 8 partial, 12 unrecoverable… See the full description on the dataset page: https://huggingface.co/datasets/gatilin/open-vision-banana-snvc-train-full.Wake-Vision
Dataset Card for Wake Vision
Dataset Description
"Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of
current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally,
it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age,
subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.botanical-vision
Botanical Vision
Fine-grained flowering-plant classification dataset: 407,759 research-grade
iNaturalist photos across 4,094 species (all flowering plants with at least
2,000 observations). Built for Advanced Computer Vision (UChicago ADSP 32023).
Splits
split
images
train
285,136
val
61,288
test
61,335
Split is stratified within each species (70/15/15). Exact and cross-species
duplicate images were removed before splitting.
Fields… See the full description on the dataset page: https://huggingface.co/datasets/dbabnigg/botanical-vision.botanical-vision-256
Botanical Vision
Fine-grained flowering-plant classification dataset: 407,759 research-grade
iNaturalist photos across 4,094 species (all flowering plants with at least
2,000 observations). Built for Advanced Computer Vision (UChicago ADSP 32023).
Images are downscaled so the long edge is at most 256px (a smaller, Colab-friendly build of the full-resolution dataset).
Splits
split
images
train
285,136
val
61,288
test
61,335
Split is stratified… See the full description on the dataset page: https://huggingface.co/datasets/dbabnigg/botanical-vision-256.vilhyra-vision-dataset
VISION — mirror
This repository redistributes the VISION dataset. It is a mirror for research and
engineering convenience. All credit belongs to the original authors; nothing here
is original work by the redistributor.
Attribution (required by the licence)
Shullani, D., Fontani, M., Iuliani, M., Al Shaya, O., & Piva, A. (2017).
VISION: a video and image dataset for source identification.
EURASIP Journal on Information Security, 2017(1), 15.… See the full description on the dataset page: https://huggingface.co/datasets/DVijayan/vilhyra-vision-dataset.olfaction-vision-language-dataset
Olfaction-Vision-Language Learning: A Multimodal Dataset
Olfaction • Vision • Language
An open-sourced dataset and dataset builder for prototyping and exploratory olfaction-vision-language tasks within the AI, robotics, and AR/VR domains.
Whether this dataset is used for better vision-scent navigation with drones, triangulating the source of an odor in an image, extracting aromas from a scene, or augmenting a VR experience with scent, we hope its release will catalyze… See the full description on the dataset page: https://huggingface.co/datasets/kordelfrance/olfaction-vision-language-dataset.eurocoin-vision-dataset
Eurocoin Vision Dataset
A compact computer vision dataset for euro coin detection and classification in realistic RGB images.
The dataset is annotated for multi-class object detection and is suitable for training and evaluating YOLO-style detectors on euro coins captured under non-ideal conditions.
Class Mapping
The dataset uses integer category IDs mapped to coin classes in the order defined by classes.txt.
Category ID
Class Name
0
10_cent
1
1_cent
2… See the full description on the dataset page: https://huggingface.co/datasets/3v3r51nc3/eurocoin-vision-dataset.imagenet-1k
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/mlx-vision/imagenet-1k.olympic-vision-100-sports
Olympic Vision: 100 Sports Classification
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
13,992
Labeled training data
test
500
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
image
Image
image_id
string
label
string (test sentinel: unlabeled)… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/olympic-vision-100-sports.idiom-vision-fooling
Idioms in Misleading Visual Context
A small, densely-annotated multimodal benchmark testing whether a misleading image can push a
vision-language model toward the wrong reading of a potentially idiomatic phrase, while human
annotators stay unaffected.
Each example pairs a sentence containing a potentially idiomatic expression with an image. The
image either matches the sentence's intended reading (aligned) or depicts the opposite
reading (misleading). Annotators label how… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/idiom-vision-fooling.trade_vision_dataset
TradeVision: Hierarchical Physical Business & Multimodal Retail Provenance Dataset
This dataset is continuously seeded from OpenStreetMap, matched to Google Place IDs, harvested for temporal store photos, and enriched with zero-shot computer vision using Hugging Face Hub native pipelines.
Dataset Structure
The dataset is partitioned into three relational subsets loadable via Hugging Face datasets:
from datasets import load_dataset
# 1. Load Canonical Businesses… See the full description on the dataset page: https://huggingface.co/datasets/drksci/trade_vision_dataset.ai-tool-pool-jewelry-vision
AI Tool Pool Jewelry Vision Dataset
Dataset Description
This dataset contains 5,130 jewelry images organized into 5 categories for computer vision tasks. The dataset was originally created and hosted on Roboflow Universe.
Categories
Bracelet: Bracelet jewelry images
Earrings: Earring jewelry images
Necklace: Necklace jewelry images
Pendant: Pendant jewelry images
Ring: Ring jewelry images
Dataset Structure
AI-Tool-Pool-Jewelry-Vision/
├── train/… See the full description on the dataset page: https://huggingface.co/datasets/bzcasper/ai-tool-pool-jewelry-vision.olfaction-vision-language-dataset
Olfaction-Vision-Language Learning: A Multimodal Dataset
Olfaction • Vision • Language
An open-sourced dataset and dataset builder for prototyping and exploratory olfaction-vision-language tasks within the AI, robotics, and AR/VR domains.
Whether this dataset is used for better vision-scent navigation with drones, triangulating the source of an odor in an image, extracting aromas from a scene, or augmenting a VR experience with scent, we hope its release will catalyze… See the full description on the dataset page: https://huggingface.co/datasets/Strangefiction/olfaction-vision-language-dataset.Succulent-Vision-Dataset
Succulent Vision Dataset
Introduction
This dataset contains 1,000+ high-quality segmented images of succulents.
The original photos were taken in complex cluster environments, and individual succulent plants were precisely segmented using SAM (Segment Anything Model).
Processing Pipeline
To provide structured data, the images have been automatically categorized using an advanced unsupervised pipeline:
Feature Extraction: DINOv2 (Vision Transformer) for… See the full description on the dataset page: https://huggingface.co/datasets/HaiPenglai/Succulent-Vision-Dataset.OpenJev-Vision-Research-v0.1
OpenJev Vision Research v0.1
12,832 image records, with public provenance, original synthetic scenes,
and programmatically derived decision questions.
This is an experimental research dataset for visual posterior learning and
compositional decisions, released with OpenJev.
It is not a reproduction of TypeSafe's proprietary Jev model or training method.
Three separate configurations
Config
Images
What the labels mean
License
synthetic
8,192
Exact… See the full description on the dataset page: https://huggingface.co/datasets/IamBusy/OpenJev-Vision-Research-v0.1.vortex-vision-data
Vortex Vision Training Data
Multi-scale vortex structure training images for the vortex-vision CNN/GNN pipeline.
Overview
This dataset contains labeled images of vortex structures across physical scales, from atmospheric phenomena (tornadoes, waterspouts, jellyfish clouds) to astronomical objects (nebular pillars, proplyds, filaments). The goal is to train a neural network that identifies vortex morphologies (columns, rings, braids, Y-junctions, X-crossings… See the full description on the dataset page: https://huggingface.co/datasets/JimGalasyn/vortex-vision-data.zamai-pashto-vision
ZamAI Pashto Vision
Languages: psLicense: cc-by-4.0Task categories: image-to-text, image-classificationSize categories: 1K<n<10K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for image-to-text, image-classification tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-vision")
print(dataset)
Configs
pashto_captions: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-vision.impromptu-vision-test-dataset
Dataset Card for impromptu-vision-dataset
This is a FiftyOne dataset with 624 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("arponde1v2j/impromptu-vision-test-dataset")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/arponde1v2j/impromptu-vision-test-dataset.supermix-science-vision-dataset
Supermix Science Vision Dataset
This dataset repo contains the compact science-diagram image set used for the local Science Vision Micro model in the Supermix workspace.
Files
images/ with 11 labeled PNG images
metadata.csv
metadata.jsonl
science_image_recognition_micro_v1_summary.json
Labels
cell
magnet
circuit
water_cycle
moon_phases
photosynthesis
solar_system
dna
temperature
states_of_matter
light_dispersion
Format
metadata.csv and… See the full description on the dataset page: https://huggingface.co/datasets/Kai9987kai/supermix-science-vision-dataset.botanical-vision-test
Botanical Vision
Fine-grained flowering-plant classification dataset: 495 research-grade
iNaturalist photos across 5 species (all flowering plants with at least
2,000 observations). Built for Advanced Computer Vision (UChicago ADSP 32023).
Splits
split
images
train
345
val
75
test
75
Split is stratified within each species (70/15/15). Exact and cross-species
duplicate images were removed before splitting.
Fields
image — the… See the full description on the dataset page: https://huggingface.co/datasets/dbabnigg/botanical-vision-test.ai-tool-pool-jewelry-vision
AI Tool Pool Jewelry Vision Dataset
Dataset Description
This dataset contains 5,130 jewelry images organized into 5 categories for computer vision tasks. The dataset was originally created and hosted on Roboflow Universe.
Categories
Bracelet: Bracelet jewelry images
Earrings: Earring jewelry images
Necklace: Necklace jewelry images
Pendant: Pendant jewelry images
Ring: Ring jewelry images
Dataset Structure
AI-Tool-Pool-Jewelry-Vision/
├── train/… See the full description on the dataset page: https://huggingface.co/datasets/Tarunhugging/ai-tool-pool-jewelry-vision.autotrain-data-vision-tcg
AutoTrain Dataset for project: vision-tcg
Dataset Description
This dataset has been automatically processed by AutoTrain for project vision-tcg.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<600x825 RGB PIL image>",
"target": 0,
"feat_Unnamed: 2": null,
"feat_Unnamed: 3": null
},
{
"image": "<600x825 RGB PIL… See the full description on the dataset page: https://huggingface.co/datasets/micazevedo/autotrain-data-vision-tcg.robot-vision-model-vs
Robot Vision Sample Dataset
This dataset is created for experimental robotics vision and perception training purposes.
robot-vision-model-vs
Robot Vision Sample Dataset
This dataset is created for experimental robotics vision and perception training purposes.
robot_vision_dataset_vh
Robot Vision Sample Dataset
This dataset is created for experimental robotics vision and perception training purposes.
robot-vision-sample-dataset2
Robot Vision Sample Dataset
This dataset is created for experimental robotics vision and perception training purposes.
robot-vision-sample-dataset12
Robot Vision Sample Dataset
This dataset is created for experimental robotics vision and perception training purposes.
robot-vision-sample-dataset35
Robot Vision Sample Dataset
This dataset is created for experimental robotics vision and perception training purposes.
Nexora-vision-dataset-v2-medium
Nexora Vision Dataset v2 Medium
The Nexora Vision Dataset v2 Medium is a scalable, mixed-resolution image dataset designed for generative AI experimentation, diffusion model workflows, and computer vision research.
Developed and curated by ArkDevLabs / ArkAiLab (ADL).
Official Website: https://arkdevlabs.com
Dataset Summary
Nexora Vision Dataset v2 Medium contains 9,236 curated images packaged in both:
Raw image format
Optimized Parquet format
This release prioritizes:… See the full description on the dataset page: https://huggingface.co/datasets/ArkAiLab-Adl/Nexora-vision-dataset-v2-medium.robot-vision-dataset-vk
Robot Vision Sample Dataset
This dataset is created for experimental robotics vision and perception training purposes.
