datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spada-dataset
SPADA Dataset
This dataset contains images and sparse labels used in the paper Land Cover Segmentation with Sparse Annotations from Sentinel-2 Imagery
, published at IGARSS 2023.
Repository: https://github.com/links-ads/igarss-spada
Paper: https://paperswithcode.com/paper/land-cover-segmentation-with-sparse
Dataset Preparation
The dataset has been compressed into segmented tarballs for ease of use within Git LFS (that is, tar > gzip > split).
To revert the process… See the full description on the dataset page: https://huggingface.co/datasets/links-ads/spada-dataset.farsite-layers
PanEU FARSITE STAC Catalog
Dataset Description
This dataset provides harmonised, continent-wide raster layers at approximately 74m spatial resolution covering Europe, developed under the FIRE‑RES programme by the CIRGEO Centre at the University of Padova. It includes information on surface fuel models, canopy fuel attributes (such as canopy height, canopy cover, bulk density), and topographic features. These layers are co-registered and designed to support… See the full description on the dataset page: https://huggingface.co/datasets/links-ads/farsite-layers.gaia-vineyard-uav-datasetartic-dataset
Art Institute of Chicago public domain dataset
Summary
This dataset includes all public domain artworks (metadata+images) from the Art Institute of Chicago, as of 2026-05-27.
All data included in this dataset is originally shared under CC0 or public domain.
Dataset Structure
The dataset consists of three fields:
"id": the unique ID (int64) associated to the artwork
"metadata": a JSON dict with metadata associated to the artwork
"image": image… See the full description on the dataset page: https://huggingface.co/datasets/links-ads/artic-dataset.cma-dataset
Cleveland Museum of Art public domain dataset
Summary
This dataset includes all public domain artworks (metadata+images) from the Cleveland Museum of Art, as of 2026-05-27.
All data included in this dataset is originally shared under CC0 or public domain.
Dataset Structure
The dataset consists of three fields:
"id": the unique accession number (string) associated to the artwork
"metadata": a JSON dict with metadata associated to the artwork… See the full description on the dataset page: https://huggingface.co/datasets/links-ads/cma-dataset.imgcdninsar-regional-snow-mapping
InSAR Regional Snow Mapping Dataset
The dataset pairs Sentinel-1 InSAR stacks with IT-SNOW snow products and ERA5-derived weather
variables over the Italian Alps, for the purpose of pixel-wise regression of Snow Depth (HS) and
Snow Water Equivalent (SWE).
Dataset overview
Property
Value
Area
Italian Alps (Trentino region)
Bounding box
9.80°E – 10.87°E, 45.93°N – 46.40°N
CRS
EPSG:4326 (WGS84)
Raster size
93 × 211 pixels
Spatial resolution
~500 m (≈… See the full description on the dataset page: https://huggingface.co/datasets/links-ads/insar-regional-snow-mapping.links-hub-demolinks-hub-demospada-inferences
SPADA Inferences
This dataset provides Land Cover inference maps of SPADA model on 12 bands Sentinel-2 tiles from different areas and years, mainly Safers pilot areas.
Dataset Content
The Land Cover map tiles are stored as 2048x2048 raster TIFFs.
The 9 inferred classes are:
0: Artificial
1: Bare vegetation
2: Wetlands
3: Water
4: Grassland
5: Agricultural
6: Broadleaved forest
7: Coniferous forest
8: Shrubs
255: Ignored
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/links-ads/spada-inferences.turkish-podcast-links
Turkish Podcast RSS Links
This dataset contains podcast RSS feed metadata for podcasts classified as Turkish by both lingua-py and fast-langdetect. It includes only feed metadata and URLs (e.g., RSS URLs), not any podcast audio or episode content.
Why and how was it built?
I wanted to create this dataset because I could not find an existing one that met my needs. I then downloaded the Podcast Index dataset, which contains 4.5 million podcasts globally. Although entries… See the full description on the dataset page: https://huggingface.co/datasets/hcsolakoglu/turkish-podcast-links.links
