csaybar/CloudSEN12-nolabel
🚨 New Dataset Version Released! We are excited to announce the release of Version [1.1] of our dataset! This update includes: [L2A & L1C support]. [Temporal support]. [Check the data without downloading (Cloud-optimized properties)]. 📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab CloudSEN12 NOLABEL A Benchmark Dataset for Cloud Semantic Understanding CloudSEN12 is a LARGE dataset (~1 TB) for cloud… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-nolabel.
🚨 New Dataset Version Released! We are excited to announce the release of Version [1.1] of our dataset! This update includes: [L2A & L1C support]. [Temporal support]. [Check the data without downloading (Cloud-optimized properties)]. 📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12 NOLABEL
A Benchmark Dataset for Cloud Semantic Understanding
![]()
CloudSEN12 is a LARGE dataset (~1 TB) for cloud semantic understanding that consists of 49,400 image patches (IP) that are evenly spread throughout all continents except Antarctica. Each IP covers 5090 x 5090 meters and contains data from Sentinel-2 levels 1C and 2A, hand-crafted annotations of thick and thin clouds and cloud shadows, Sentinel-1 Synthetic Aperture Radar (SAR), digital elevation model, surface water occurrence, land cover classes, and cloud mask results from six cutting-edge cloud detection algorithms.
CloudSEN12 is designed to support both weakly and self-/semi-supervised learning strategies by including three distinct forms of hand-crafted labeling data: high-quality, scribble and no-annotation. For more details on how we created the dataset see our paper.
Ready to start using [CloudSEN12](https://cloudsen12.github.io/)?
[Download Dataset](https://cloudsen12.github.io/download.html)
[Paper - Scientific Data](https://www.nature.com/articles/s41597-022-01878-2)
[Inference on a new S2 image](https://colab.research.google.com/github/cloudsen12/examples/blob/master/example02.ipynb)
[Enter to cloudApp](https://github.com/cloudsen12/CloudApp)
[CloudSEN12 in Google Earth Engine](https://gee-community-catalog.org/projects/cloudsen12/)
<br>
Description
<br>
<br>
Label Description
<br>
np.memmap shape information
<br>
cloudfree (0\%) shape: (5880, 512, 512) <br> almostclear (0-25 \%) shape: (5880, 512, 512) <br> lowcloudy (25-45 \%) shape: (5880, 512, 512) <br> midcloudy (45-65 \%) shape: (5880, 512, 512) <br> cloudy (65 > \%) shape: (5880, 512, 512)
<br>
Example
<br>
import numpy as np
# Read high-quality train
cloudfree_shape = (5880, 512, 512)
B4X = np.memmap('cloudfree/L1C_B04.dat', dtype='int16', mode='r', shape=cloudfree_shape)
y = np.memmap('cloudfree/manual_hq.dat', dtype='int8', mode='r', shape=cloudfree_shape)
# Read high-quality val
almostclear_shape = (5880, 512, 512)
B4X = np.memmap('almostclear/L1C_B04.dat', dtype='int16', mode='r', shape=almostclear_shape)
y = np.memmap('almostclear/kappamask_L1C.dat', dtype='int8', mode='r', shape=almostclear_shape)
# Read high-quality test
midcloudy_shape = (5880, 512, 512)
B4X = np.memmap('midcloudy/L1C_B04.dat', dtype='int16', mode='r', shape=midcloudy_shape)
y = np.memmap('midcloudy/kappamask_L1C.dat', dtype='int8', mode='r', shape=midcloudy_shape)<br>
This work has been partially supported by the Spanish Ministry of Science and Innovation project PID2019-109026RB-I00 (MINECO-ERDF) and the Austrian Space Applications Programme within the [SemantiX project](https://austria-in-space.at/en/projects/2019/semantix.php).
