datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TerraMesh
TerraMesh
A planetary‑scale, multimodal analysis‑ready dataset for Earth‑Observation foundation models: TerraMesh merges data from Sentinel‑1 SAR, Sentinel‑2 optical, Copernicus DEM, NDVI, and land‑cover sources into more than 9 million co‑registered patches ready for large‑scale representation learning.
You find more information about the data sampling and preprocessing in our paper: TerraMesh: A Planetary Mosaic of Multimodal Earth Observation Data.
Samples from the TerraMesh… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh.geospatially_enriched_ndvi
Geospatially Enriched NDVI (16-Day Terra/MODIS)
This dataset transforms raw 16-day MODIS NDVI grids into a per-pixel time series enriched with hierarchical administrative boundaries. It covers every 0.1°×0.1° land pixel worldwide from 2000 onward and is partitioned for efficient bulk download and selective access.
Dataset Contents
Partitioned Parquet filesStored under:
ndvi/
├── year=YYYY/
│ ├── country=Netherlands/
│ │ └── data_0.parquet
│ └── country=India/
│ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/svenmeijboom/geospatially_enriched_ndvi.TerraMesh-Masks
TerraMesh-Masks
TerraMesh-Masks is a dataset for open-vocabulary segmentation of satellite imagery. This dataset provides binary segmentation masks with captions that extend the samples from TerraMesh.
We also provide an human-verfied evaluation benchmark, called TerraMesh-Masks-Eval.
Examples from the training subset:
Usage
Download the data loading code from GitHub and install requirements with pip install -r requirements.txt. For development, you can… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh-Masks.CoordBench
CoordBench
A unified benchmark suite for evaluating location encoders such as SatCLIP, GeoCLIP, Climplicit, and MIND.
The dataset contains 40 normalized source tables from 13 source families. The paper's evaluation suite
uses 52 datasets and 78 prediction targets drawn from this mirror. The source files previously lived across GitHub,
figshare, GCS, Socrata, Zenodo, and Google Drive.
Intended use
Use the normalized tables to compare coordinate-to-embedding models.… See the full description on the dataset page: https://huggingface.co/datasets/taylor-geospatial/CoordBench.MINDSET
MINDSET
MINDSET is the pretraining dataset for MIND, a coordinate-only location encoder distilled from static location encoder teachers and annual AlphaEarth Foundations (AEF) embeddings.
We release the embeddings at the 12.1M training coordinates. The dataset contains 12,099,072 land coordinates in WGS84. Coordinates are dense around cities and not uniformly sampled over land.
The files are in GeoParquet format and can be joined on point_id:
file
grain
rows
columns… See the full description on the dataset page: https://huggingface.co/datasets/taylor-geospatial/MINDSET.TerraMesh-Masks-Eval
TerraMesh-Masks-Eval
TerraMesh-Masks-Eval is a human-verified benchmark dataset to evaluate open-vocabulary segmentation models on satellite imagery. This dataset provides binary segmentation masks with captions togehther with input samples from TerraMesh.
We also provide a training dataset, called TerraMesh-Masks.
Examples from the evaluation subset:
Usage
Download the data loading code from GitHub and install requirements with pip install -r requirements.txt.… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh-Masks-Eval.brazil-wildfire-geospatial-dataset
Banco Histórico de Incêndios no Brasil — 2018–2025
Banco de dados com 1,43 milhão de focos de calor registrados no Brasil entre 2018 e 2025, enriquecidos com dados meteorológicos (ERA5), cobertura do solo (MapBiomas) e altitude (SRTM).
Construído a partir de fontes públicas oficiais para suporte a pesquisas científicas sobre incêndios florestais.
Tabelas disponíveis
Tabela
Arquivos
Linhas
Descrição
focos_analise
focos_analise/*.parquet
1.430.756
Tabela… See the full description on the dataset page: https://huggingface.co/datasets/mateus-pcosta/brazil-wildfire-geospatial-dataset.florida_geospatial
