datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wds_vtab-caltech101caltech-ucsd-birds-200-2011
Caltech-UCSD Birds-200-2011 (CUB-200-2011)
This dataset contains the Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset, from here.
Each example consists of an image, a label, and a bounding box. (The dataset also contains x/y locations of "parts", e.g. beak, right eye, left wing, throat, etc. and "attributes", e.g. beak shape, wing color, feather pattern. I have not included either of these. Contact me if you want me to add them.)
Note: Some of these images are also in ImageNet!… See the full description on the dataset page: https://huggingface.co/datasets/bentrevett/caltech-ucsd-birds-200-2011.caltech101
Dataset Card for Caltech 101
This dataset contains images of objects from 101 distinct categories, with each category comprising approximately 40 to 800 images. The majority of categories include around 50 images each. The images were collected in September 2003 by Fei-Fei Li, Marco Andreetto, and Marc’Aurelio Ranzato. Each image has an approximate resolution of 300 x 200 pixels.
Dataset Sources
Website: https://data.caltech.edu/records/mzrjq-6wc02
Use in FL… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/caltech101.caltech101caltech256Caltech-256
Dataset Card for Dataset Name
This is the huggingface format of : https://data.caltech.edu/records/nyy15-4j048. Please cite the original author of the dataset
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/ilee0022/Caltech-256.Caltech101_not_background_test
Dataset Card for "Caltech101_not_background_test"
More Information needed
caltech-256Griffin, G., Holub, A., & Perona, P. (2022). Caltech 256 (1.0) [Data set]. CaltechDATA. https://doi.org/10.22002/D1.20087
TWIN
TWIN
This repository contains the TWIN dataset introduced in the paper Same or Not? Enhancing Visual Perception in Vision-Language Models. TWIN contains 561K challenging (image, question, answer) tuples emphasizing fine-grained image understanding.
For evaluating on the dataset with LMMS-eval, please refer to this repo.
Citation
If you use the TWIN dataset in your research, please use the following BibTeX entry.
@misc{marsili2025notenhancingvisualperception… See the full description on the dataset page: https://huggingface.co/datasets/glab-caltech/TWIN.caltech_birds2011Caltech101
Caltech101
An MTEB dataset
Massive Text Embedding Benchmark
Classifying images of 101 widely varied objects.
Task category
i2c
Domains
Encyclopaedic
Reference
https://ieeexplore.ieee.org/document/1384978
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["Caltech101"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To… See the full description on the dataset page: https://huggingface.co/datasets/mteb/Caltech101.caltech-101Li, F.-F., Andreeto, M., Ranzato, M., & Perona, P. (2022). Caltech 101 (1.0) [Data set]. CaltechDATA. https://doi.org/10.22002/D1.20086
caltech256Caltech101_not_background_test_facebook_opt_2.7b_Attributes_Caption_ns_5647
Dataset Card for "Caltech101_not_background_test_facebook_opt_2.7b_Attributes_Caption_ns_5647"
More Information needed
FGVQA
FGVQA
This repository contains the FGVQA benchmark suite introduced in the paper Same or Not? Enhancing Visual Perception in Vision-Language Models.FGVQA contains 12,000 challenging (image, question, answer) tuples emphasizing fine-grained image understanding.
The benchmark suite is composed of six sub-benchmarks:
TWIN-eval
ILIAS
Google Landmarks v2
MET
CUB
Inquire
For evaluating on the dataset with LMMS-eval, please refer to this repo.
Citation
If you use the FGVQA… See the full description on the dataset page: https://huggingface.co/datasets/glab-caltech/FGVQA.vtab_caltech101
VTAB Caltech101
This dataset has been used for the paper Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models (NeurIPS 2025).
It reproduces the settings (splits, labels) used for the Visual Task Adaptation Benchmark (VTAB).
VTAB Paper: A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
VTAB Repository: google-research/task_adaptation
Details of the original dataset:
Original Citation:… See the full description on the dataset page: https://huggingface.co/datasets/bramtoula/vtab_caltech101.Caltech101_with_background_test_facebook_opt_350m_Attributes_ns_6084
Dataset Card for "Caltech101_with_background_test_facebook_opt_350m_Attributes_ns_6084"
More Information needed
Caltech101_with_background_test_facebook_opt_1.3b_Visclues_ns_6084
Dataset Card for "Caltech101_with_background_test_facebook_opt_1.3b_Visclues_ns_6084"
More Information needed
Caltech101_not_background_test_facebook_opt_125m_Attributes_ns_5647
Dataset Card for "Caltech101_not_background_test_facebook_opt_125m_Attributes_ns_5647"
More Information needed
Caltech101_not_background_test_facebook_opt_1.3b_Attributes_Caption_ns_5647
Dataset Card for "Caltech101_not_background_test_facebook_opt_1.3b_Attributes_Caption_ns_5647"
More Information needed
Caltech101_with_background_test_facebook_opt_1.3b_Attributes_Caption_ns_6084
Dataset Card for "Caltech101_with_background_test_facebook_opt_1.3b_Attributes_Caption_ns_6084"
More Information needed
Caltech101_not_background_test_facebook_opt_2.7b_Visclues_ns_5647
Dataset Card for "Caltech101_not_background_test_facebook_opt_2.7b_Visclues_ns_5647"
More Information needed
Caltech101_with_background_test_facebook_opt_2.7b_Attributes_ns_6084
Dataset Card for "Caltech101_with_background_test_facebook_opt_2.7b_Attributes_ns_6084"
More Information needed
Caltech-256Caltech101_not_background_test_facebook_opt_350m_Attributes_Caption_ns_5647
Dataset Card for "Caltech101_not_background_test_facebook_opt_350m_Attributes_Caption_ns_5647"
More Information needed
Caltech-101The Caltech-101 dataset of images.
The dataset consists of pictures of objects belonging to 101 classes, plus one background clutter class (BACKGROUND_Google). Each image is labelled with a single object.
Each class contains roughly 40 to 800 images, totaling around 9,000 images. Images are of variable sizes, with typical edge lengths of 200-300 pixels. This version contains image-level labels only.
source:
https://data.caltech.edu/records/mzrjq-6wc02
Caltech101_with_background_test_facebook_opt_2.7b_Attributes_Caption_ns_6084
Dataset Card for "Caltech101_with_background_test_facebook_opt_2.7b_Attributes_Caption_ns_6084"
More Information needed
Caltech101_not_background_test_facebook_opt_1.3b_Visclues_ns_5647
Dataset Card for "Caltech101_not_background_test_facebook_opt_1.3b_Visclues_ns_5647"
More Information needed
Caltech101_not_background_test_facebook_opt_125m_Visclues_ns_5647
Dataset Card for "Caltech101_not_background_test_facebook_opt_125m_Visclues_ns_5647"
More Information needed
Caltech101_with_background_test_facebook_opt_1.3b_Attributes_ns_6084
Dataset Card for "Caltech101_with_background_test_facebook_opt_1.3b_Attributes_ns_6084"
More Information needed
