datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Core-AlphaEarth-Embeddings
Major TOM Core AlphaEarth Embeddings Subset
This is a prototype dataset. It only includes some of the AlphaEarth embeddings stored in Major TOM grid cells.
This dataset is mostly aimed at experimentation and prototyping. It is particularly useful to use it along other datasets published within the Major TOM project.
Content
Field
Type
Description
grid_cell
string
Major TOM cell
year
int
year of the sample
thumbnail
image
3-dimensional PCA… See the full description on the dataset page: https://huggingface.co/datasets/Major-TOM/Core-AlphaEarth-Embeddings.TreeOfLife-200M-Embeddings
TreeOfLife-200M Embeddings
Pre-computed image embeddings for all images from the TreeOfLife-200M dataset (revision 94bbc0b), sorted by taxonomic hierarchy for efficient filtered access.
This repository hosts embedding configs for TreeOfLife-200M. Each config corresponds to a different embedding model and/or precision. Currently available: BioCLIP 2 (float16) and BioCLIP 2.5 Huge (float16, L2-normalized). Additional configs will be added as new embeddings are generated.
We… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-200M-Embeddings.big-animal-dataset-high-res-embedding-with-hidden-states
Dataset Card for "big-animal-dataset-high-res-embedding-with-hidden-states"
More Information needed
datacomp-small-with-text-embeddings
Dataset Card for "datacomp-small-with-text-embeddings"
More Information needed
Datacomp-10m-embeddingdatacomp-small-with-embeddings
Dataset Card for "datacomp-small-with-embeddings"
More Information needed
datacomp-small-with-embeddings-and-cluster-labels
Dataset Card for "datacomp-small-with-embeddings-and-cluster-labels"
More Information needed
wolt-food-clip-ViT-B-32-embeddings
wolt-food-clip-ViT-B-32-embeddings
Qdrant's Food Discovery demo relies on the dataset of food images from the Wolt
app. Each point in the collection represents a dish with a single image. The image is represented as a vector of 512
float numbers.
Generation process
The embeddings generated with clip-ViT-B-32 model have been generated using the following code snippet:
from PIL import Image
from sentence_transformers import SentenceTransformer
image_path =… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/wolt-food-clip-ViT-B-32-embeddings.turkish_clip_dataset_with_text_embeddingsThis dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb
The main purpose was to extract Turkish captions and download images.
You can use this dataset to fine-tune or create a clip model.
Since there English and Turkish captions you can also use those to create language model?
Logo-Recognition-ResNet50-TripletNet-Embeddings-Datasetfashion-products-small-align-embeddingsAGRILLAVA-embeddings3d-model-images-embeddings
Dataset Card for "3D CAD Models with Preview Images and Embeddings"
This dataset provides approximately 1 million 3D CAD models from the ABC Dataset (Deep Geometry) paired with:
A single rendered preview image
A generated caption
A text embedding
The dataset is designed primarily for large-scale retrieval, representation learning, and multimodal research on CAD geometry.
⚠️ Important:This release contains preview images and embeddings only — not the original CAD geometry… See the full description on the dataset page: https://huggingface.co/datasets/daveferbear/3d-model-images-embeddings.datacomp-small-with-embeddings-ca-filtered
Dataset Card for "datacomp-small-with-embeddings-ca-filtered"
This is the datacomp-small dataset, with CLIP-large-patch14 image embeddings added, as well as CA filtering:
minimum caption complexity of 1
minimum 1 action in the caption
mtg-scryfall-cropped-art-embeddings-siglip-so400m-patch14-384merged_remote_landscapes_v1
Dataset Card for Merged Remote Landscapes dataset
Dataset summary
This is a merged version of following datasets:
torchgeo/ucmerced
NWPU-RESISC45
from datasets import load_dataset
dataset = load_dataset('EmbeddingStudio/merged_remote_landscapes_v1')
Categories
This is a union of categories from original datasets:
agricultural, airplane, airport, baseball diamond, basketball court, beach, bridge, buildings, chaparral, church, circular farmland, cloud… See the full description on the dataset page: https://huggingface.co/datasets/EmbeddingStudio/merged_remote_landscapes_v1.galaxies_embeddings
Embeddings for 🤗 smith42/galaxies
This dataset contains embeddings generated by AstroPTv2.0 🐙.
The embeddings are generated from smith42/galaxies, and can be paired with accompanying metadata from that dataset (the rows are in the same order!).
Check out the github for more docs and info on AstroPT.
mtg-scryfall-unique-artwork-20240809-with-card-art-descriptions-and-images-with-embeddingsmtg-scryfall-unique-artwork-20240809-combined-elements-embeddingsSeaDoc
SeaDoc: from the paper "Scaling Language-Centric Omnimodal Representation Learning"
This repository hosts the SeaDoc dataset, a challenging visual document retrieval task in Southeast Asian languages, introduced in the paper Scaling Language-Centric Omnimodal Representation Learning. It is designed to evaluate and enhance language-centric omnimodal embedding frameworks by focusing on a low-resource setting, specifically for tasks involving diverse languages and visual document… See the full description on the dataset page: https://huggingface.co/datasets/LCO-Embedding/SeaDoc.t5-gemma-2-multimodal-embeddingvidore_v3_computer_science_embeddingNOTE
ViDoRe V3: Computer Science dataset ColQwen2 Embeddings
This dataset contains pre-computed embeddings for the ViDoRe V3 : Computer Science dataset using the ColQwen2 model.
ViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/vidore_v3_computer_science_embedding.big-animal-dataset-with-embeddingHatefulmemes_train_embeddings
Dataset Card for "Hatefulmemes_train_embeddings"
More Information needed
vogue933k-embedding
Vogue Runway Image-Embedding Corpus
Dataset of Vogue Runway images and embeddings.
Metric
Value
Designers
1 749
Collections
25 876
Images / Looks
933 328
Years covered
1990 → 2025
Embedding model
google/vit-base-patch16-224 (mean-pooled CLS tokens)
Vector size
768 float32
from datasets import load_dataset
ds = load_dataset("tonyassi/vogue933k-embedding", streaming=True, split="train")
for i, row in zip(range(10), ds): print(row)
laion400m_overlap_embeddingsbig-animal-dataset-high-res-embedding
Dataset Card for "big-animal-dataset-high-res-embedding"
More Information needed
fashion-products-small-multimodal-embeddingscm4-synthetic-testing-with-embeddings
Dataset Card for "cm4-synthetic-testing-with-embeddings"
More Information needed
vogue-runway-top15-512px-nobg-embeddings
vogue-runway-top15-512px-nobg-embeddings
Vogue Runway
15 fashion houses
1679 collections
87,547 images
Fashion Houses: Alexander McQueen, Armani, Balenciaga, Calvin Klein, Chanel, Dior, Fendi, Gucci, Hermes, Louis Vuitton, Prada, Ralph Lauren, Saint Laurent, Valentino, Versace.
Images are maximum height 512 pixels.
Background is removed using mattmdjaga/segformer_b2_clothes.
Embeddings generated with tonyassi/vogue-fashion-collection-15-nobg.
