datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet-w21-wds-dinov2imagenet1k_invae-latents_dinov2_pcaCore-S2RGB-DINOv2
Core-S2RGB-DINOv2 🔴🟢🔵
Dataset
Modality
Number of Embeddings
Sensing Type
Total Comments
Source Dataset
Source Model
Size
Core-S2RGB-SigLIP
Sentinel-2 Level 2A (RGB)
56,147,150
True Colour (RGB)
General-Purpose Global
Core-S2L2A
DINOv2
223.1 GB
Content
Field
Type
Description
unique_id
string
hash generated from geometry, time, product_id, and embedding model
embedding
array
raw embedding array
grid_cell
string
Major TOM cell
grid_row_u… See the full description on the dataset page: https://huggingface.co/datasets/Major-TOM/Core-S2RGB-DINOv2.openaccess-embeddings-dinov2-giant
metmuseum/openaccess-embeddings-dinov2-giant
Image embeddings for every public-domain artwork in metmuseum/openaccess, produced by facebook/dinov2-giant.
Column
Type
Notes
objectID
int64
Primary key — matches objectID in metmuseum/openaccess
embedding
list<float32>
L2-normalised, dim = 1536
model
string
Source model id
dim
int32
Embedding dimension
Image bytes are not stored here; join against the main dataset to recover them.
Embedding spec: dim=1536, expected… See the full description on the dataset page: https://huggingface.co/datasets/metmuseum/openaccess-embeddings-dinov2-giant.deletion-confluence-dinov2-features
Frozen DINOv2 features for the deletion-ordering experiments
CLS-token features of a frozen facebook/dinov2-base backbone on CIFAR-10 and
CIFAR-100, used as the input to the convex-head experiments in
deletion-confluence.
They are here so that a checkout of that repository can run on a machine without
re-extracting them; they are not new data.
What is in each directory
dinov2_base_cifar10/ train_features.npy (50000, 768) float32… See the full description on the dataset page: https://huggingface.co/datasets/ZiyuZhao98/deletion-confluence-dinov2-features.DinoV2-YGO-card-embeddingsCore-S2RGB-249k-DINOv2
Core-S2RGB-249k-DINOv2
Vision-only embedding dataset computed from Core-S2L2A-249k using the DINOv2-large model.
Overview
Property
Value
Source imagery
Core-S2L2A-249k (248,719 patches, 384 × 384 px)
Model
DINOv2-large (facebook/dinov2-large)
Input bands
RGB [B04, B03, B02]
Embedding dimension
768
Output format
GeoParquet
License
CC-BY-SA-4.0
Computation Pipeline
Pre-processing: Each 384 × 384 Sentinel-2 L2A patch is read from the… See the full description on the dataset page: https://huggingface.co/datasets/Major-TOM/Core-S2RGB-249k-DINOv2.
