datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SSL4EO-S12-downstream
SSL4EO-S12-downstream
Welcome to the SSL4EO-S12-downstream dataset. This dataset is used in the Embed2Scale Challenge.
SSL4EO-S12-downstream is a Earth Observation (EO) dataset of downstream tasks. It is released as a standalone dataset together with the NeuCo-Bench neural compression benchmarking framework. Parts of the SSL4EO-S12-downstream dataset was used in the 2025 CVPR EarthVision data challenge and the dev and eval phases of the challenge can be recreated. For instructions… See the full description on the dataset page: https://huggingface.co/datasets/embed2scale/SSL4EO-S12-downstream.SSLQ-Version-1.600
SSLQ Version 1.600
SSLQ Version 1.600, “Synthetic Scenic Lore Image Quality”, is a hand-curated dataset of 600 manually annotated images drawn from the SSLQ archive, which is a long-term study in synthetic image aesthetics
and visual coherence. Each image was labeled in Label Studio 1.13 using a structured XML schema and enriched with human and LLM commentary
describing visual qualities, stylistic alignment, and subjective evaluations of quality and mood.… See the full description on the dataset page: https://huggingface.co/datasets/shalvers/SSLQ-Version-1.600.SSL-HWD-Words
SSL-HWD (A Large Scale Handwritten Image Dataset)
Dataset Description
Dataset Summary
SSL-HWD is a large-scale handwritten text dataset introduced in the paper "Learning Beyond Labels: Self-Supervised Handwritten Text Recognition" (WACV 2026). The dataset comprises 10 million word-level handwritten images from 852 writers across diverse domains including Physics, Computer Science, Biology, Mathematics, and more.
The dataset is specifically designed to support… See the full description on the dataset page: https://huggingface.co/datasets/Shree10/SSL-HWD-Words.SSL-HWD-Words
SSL-HWD (A Large Scale Handwritten Image Dataset)
Dataset Description
Dataset Summary
SSL-HWD is a large-scale handwritten text dataset introduced in the paper "Learning Beyond Labels: Self-Supervised Handwritten Text Recognition" (WACV 2026). The dataset comprises 10 million word-level handwritten images from 852 writers across diverse domains including Physics, Computer Science, Biology, Mathematics, and more.
The dataset is specifically designed… See the full description on the dataset page: https://huggingface.co/datasets/RKK01/SSL-HWD-Words.
