abzal-glw/cryosentinel-glof-v3
CryoSentinel-GLOF v3 — Multimodal Glacial Lake Chips for High Mountain Asia A multimodal Earth-observation dataset of 42,237 image chips centred on glacial lakes across twelve High Mountain Asia sub-regions. Each chip stacks three co-registered modalities — Sentinel-1 SAR (VV, VH), Sentinel-2 optical (12 bands, L2A), and Copernicus DEM (1 band) — together with a binary water-extent mask derived from the Kumar & Vijay (2026) PANGAEA glacial-lake inventory. This dataset is the… See the full description on the dataset page: https://huggingface.co/datasets/abzal-glw/cryosentinel-glof-v3.
CryoSentinel-GLOF v3 — Multimodal Glacial Lake Chips for High Mountain Asia
A multimodal Earth-observation dataset of 42,237 image chips centred on glacial lakes across twelve High Mountain Asia sub-regions. Each chip stacks three co-registered modalities — Sentinel-1 SAR (VV, VH), Sentinel-2 optical (12 bands, L2A), and Copernicus DEM (1 band) — together with a binary water-extent mask derived from the Kumar & Vijay (2026) PANGAEA glacial-lake inventory.
This dataset is the supervision substrate for the CryoSentinel-TerraMind-v3 segmentation model and the CryoSentinel GitHub project. Released under Open Data Commons Attribution License (ODC-By 1.0) as the data component of the project's open-science release.
Headline numbers
Regions covered
central_himalaya · eastern_himalaya · hengduan_nyainqentanglha · hindu_kush · hma_other · ile_alatau · karakoram · pamir · tibetan_plateau · tien_shan_full · western_himalaya · zhetysu_alatau.
Storage layout: <region>/<year>/shard_NNN.parquet.
Quick-start
from datasets import load_dataset
ds = load_dataset("abzal-glw/cryosentinel-glof-v3", split="train", streaming=True)
sample = next(iter(ds))
import numpy as np
s2 = np.frombuffer(sample["s2_bytes"], dtype=np.uint16).reshape(sample["s2_shape"]) # (12, 224, 224)
s1 = np.frombuffer(sample["s1_bytes"], dtype=np.float32).reshape(sample["s1_shape"]) # (2, 224, 224)
dem = np.frombuffer(sample["dem_bytes"], dtype=np.int16).reshape(sample["dem_shape"]) # (1, 224, 224)
msk = np.frombuffer(sample["mask_bytes"], dtype=np.uint8).reshape(sample["mask_shape"]) # (224, 224)
print(sample["chip_id"], sample["region"], sample["snapshot_year"], "water_frac:", sample["water_frac"])To download the full archive locally:
huggingface-cli download \
--repo-type dataset \
abzal-glw/cryosentinel-glof-v3 \
--local-dir ./data/multimodal_chips_v3For the Almaty-corridor finetune subset only (≈ 10 GiB):
huggingface-cli download \
--repo-type dataset \
abzal-glw/cryosentinel-glof-v3 \
--local-dir ./data/multimodal_chips_v3 \
--include "*ile_alatau*" "*tien_shan_full*" "*zhetysu_alatau*"Schema
Each row is a single image chip and has the following columns:
Modalities — preprocessing recipe
Sentinel-2 L2A (12 bands)
ESA Copernicus, accessed via Microsoft Planetary Computer / Google Earth Engine. Bands kept: B01, B02, B03, B04, B05, B06, B07, B08, B8A, B09, B11, B12. Band B10 (cirrus) is dropped because it is not delivered as surface reflectance in L2A. All bands resampled to 10 m via bilinear interpolation. Late-summer composites (Jul–Sep) selected at cloud cover < 30 % via the SCL band, with an October fallback if no qualifying tile exists.
Sentinel-1 GRD (VV, VH)
ESA Copernicus, accessed via Microsoft Planetary Computer. Pipeline (matches IBM TerraMesh recipe): orbit file → GRD border noise removal → thermal noise removal → σ⁰ calibration → 5 × 1 multi-look → 5 × 5 Lee speckle filter → Range-Doppler terrain correction using SRTM 30 m → conversion to dB. Co-registered to the S2 grid; same season as the matched S2 tile, ≤ 7 days apart.
Copernicus DEM 30 m
ESA, GLO-30 product, via Microsoft Planetary Computer. Native 30 m, bilinearly resampled to 10 m. Raw elevation in metres above the EGM2008 geoid — no derived slope or curvature; the model learns spatial gradients itself.
Per-band normalisation (train split, Welford's algorithm)
S2 (12 bands, uint16 reflectance):
B01 μ=857.6 σ=626.3 B02 μ=1044.0 σ=709.2 B03 μ=1356.7 σ=748.7 B04 μ=1574.4 σ=825.5
B05 μ=1786.2 σ=773.5 B06 μ=2076.5 σ=786.1 B07 μ=2215.2 σ=810.8 B08 μ=2277.8 σ=834.9
B8A μ=2348.5 σ=833.6 B09 μ=2243.7 σ=729.6 B11 μ=2665.0 σ=875.7 B12 μ=2217.2 σ=851.3
S1 (2 bands, dB):
VV μ=-9.25 σ=5.90 VH μ=-18.00 σ=5.92
DEM (1 band, metres):
elev μ=4299.8 σ=901.0These are the statistics the upstream model uses for (x - μ) / σ normalisation. They are also published as dataset_stats_v3.json in the project repository.
Labels — provenance and quality control
The supervision signal comes from the Inventory of Glacial Lakes in High Mountain Asia for the Years 2016 and 2022 by Kumar, R. and Vijay, S. (PANGAEA, 2026, DOI: 10.1594/PANGAEA.983845, CC-BY 4.0). Polygons are rasterised to the 10 m chip grid with all_touched=True.
To reduce label noise, every candidate chip is filtered against an automated water mask. We compute MNDWI = (B03 − B11) / (B03 + B11), threshold at 0, and discard chips with Kumar–MNDWI IoU below 0.20 or with a pred_water_pixels / kumar_water_pixels ratio above 3.0. After this filter the median Kumar–MNDWI IoU on retained chips is 0.48 — high-quality by remote-sensing standards.
A small number of chips still carry residual mislabels (e.g. seasonal ice cover on frozen lakes, registration drift at sub-pixel scale). For the Almaty corridor finetune subset, a manual audit identified 7 systematically mislabelled chips; the corresponding label-corrected test split (n = 658 instead of 665) raises held-out test IoU from 0.8918 to 0.9082. The audit is documented in docs/LABEL_NOISE_AUDIT.md of the project repository.
Splits
For the Almaty-corridor (Tien Shan full + Ile Alatau + Zhetysu Alatau) finetune subset:
Splits are spatial block based, not random per-chip — chips from the same lake or contiguous geographic neighbourhood always end up in the same split, so validation and test IoU reflect generalisation to new lakes rather than memorisation. See docs/METHOD.md for the block-split algorithm.
For the full 12-region pretrain set, all 42,237 chips are exposed under the single train split; downstream consumers should subset by region and snapshot_year columns to construct their own held-outs.
Intended use
- Training and evaluation of glacial-lake semantic segmentation models from multimodal satellite imagery, especially foundation-model adaptation studies (TerraMind, Prithvi, Clay, SatMAE, etc.).
- Benchmarking water-mapping and ice-vs-water discrimination algorithms across diverse HMA terrain (Karakoram, Pamir, Himalaya, Tibetan Plateau, Tien Shan, etc.).
- Climate-risk and GLOF (Glacial Lake Outburst Flood) hazard-monitoring research that needs co-registered SAR + optical + topographic context.
- Educational use in remote-sensing, deep-learning, and Earth-observation curricula.
Out-of-scope and limitations
- The dataset is not an operational early-warning product. It is a training/evaluation substrate; downstream operational use must be paired with hazard assessment, in-situ validation, and stakeholder review.
- Coverage is limited to High Mountain Asia (longitude ≈ 65 °E to 100 °E, latitude ≈ 27 °N to 47 °N). Models trained here are not guaranteed to transfer to Andes, Alps, Caucasus, Patagonia, or Arctic glacial settings without re-finetuning.
- Labels follow the Kumar & Vijay (2026) 2022 snapshot. Lakes that have drained, frozen, or appeared after the snapshot are mislabelled.
- Frozen-lake chips at altitudes above ≈ 5,000 m carry residual label noise (Kumar polygons mark open-water extent but the imagery often shows ice).
- The dataset does not include time series — each chip is a single seasonal composite per year. Temporal extension is on the roadmap (v1.2).
Affiliation disclosure
CryoSentinel is an independent research project. The dataset, model, and code are not affiliated with, endorsed by, or operationally integrated with Kazselezashchita, UNESCO GLOFCA, the ESA Copernicus programme, IBM, or any other public or private agency. References to those organisations describe institutional context or upstream data sources only.
How to cite
If you use this dataset in research, software, a benchmark, a derivative model, or a product, please cite both the upstream label source and this derivative:
@dataset{kumar_vijay_2026_glacial_lakes,
author = {Kumar, R. and Vijay, S.},
title = {Inventory of Glacial Lakes in High Mountain Asia for the Years 2016 and 2022},
publisher = {PANGAEA},
year = {2026},
doi = {10.1594/PANGAEA.983845}
}
@dataset{abdrash_2026_cryosentinel_glof_v3,
author = {Abdrash, Abzal},
title = {CryoSentinel-GLOF v3: Multimodal Glacial Lake Chips
for High Mountain Asia},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/abzal-glw/cryosentinel-glof-v3},
doi = {10.57967/hf/8823}
}
@software{abdrash_2026_cryosentinel_code,
author = {Abdrash, Abzal},
title = {CryoSentinel v1.0.0: A Foundation-Model Glacial Lake
Segmenter for High Mountain Asia},
year = {2026},
publisher = {Zenodo},
url = {https://github.com/abzalabdrash/cryosentinel},
doi = {10.5281/zenodo.20239229}
}Plus credit ESA Copernicus for the Sentinel-1, Sentinel-2, and Copernicus DEM imagery.
License
This derivative dataset is released under the Open Data Commons Attribution License v1.0 (ODC-By 1.0) — full text at https://opendatacommons.org/licenses/by/1-0/.
Upstream terms that apply transitively:
- Sentinel-1, Sentinel-2, Copernicus DEM imagery — ESA Copernicus open-access terms (attribution required).
- Kumar & Vijay (2026) glacial-lake polygons — CC-BY 4.0 via PANGAEA.
Reuse must preserve attribution to all three sources.
Contact
Issues and questions: github.com/abzalabdrash/cryosentinel/issues. Author: Abzal Abdrash · ORCID 0009-0006-4829-0256.
