RMDig/rocky_mountain_snowpack
Rocky Mountain Snowpack Dataset The Rocky Mountain Snowpack dataset contains 4,040 preprocessed samples of snowpack imagery collected in the Colorado Rocky Mountains across the 2024–2025 and 2025–2026 winter seasons, from 7 snowpits dug between January 2025 and February 2026.Each sample segment of snow includes three types of images: Magnified crystal images (close-up snow snow crystal profile photography) Snowpack profile images (non-magnified snow crystal profiles… See the full description on the dataset page: https://huggingface.co/datasets/RMDig/rocky_mountain_snowpack.
Rocky Mountain Snowpack Dataset
The Rocky Mountain Snowpack dataset contains 4,040 preprocessed samples of snowpack imagery collected in the Colorado Rocky Mountains across the 2024–2025 and 2025–2026 winter seasons, from 7 snowpits dug between January 2025 and February 2026. Each sample segment of snow includes three types of images:
- Magnified crystal images (close-up snow snow crystal profile photography)
- Snowpack profile images (non-magnified snow crystal profiles photography)
- Core segment images (snow cores sampled with a apple corer)
For each segment that naturally forms from the core, each image type is captured and resampled in real-time from the same snow segment up to 25 times using the shake-n-take method. This theoretically increases training sample size quadratically in multi-modal models that take two or more of the image types as input simply by creating all potential pairs of resampled segment images. Forewarning, this may introduce unintentional data-leakage (especially when using both magnified crystal images and their non-magnified images), however it may be a necessary oversampling technique given the minimal amount of data available.
In addition to images, the dataset includes rich details of environmental metadata such as number of avalanches spotted, slope angle, air and snow temperature. The samples are attached to a time and location so external metadata such as weather forecasts could easily be used as labels for the data. The dataset is designed for image classification, regression, and generative modeling tasks related to snow science and avalanche forecasting.
Explore the dataset more by visiting Rocky Mountain Digerati's website at www.rmdig.ai
Dataset Summary
- Number of examples: 4040 preprocessed, 3195 raw, from 7 sites
- Features:
image,site,column,core,segment,avalanches_spotted,wind_loading,snowpack_depth,core_depth,slope_face,slope_angle,air_temperature,core_temperature,ect_* - Label classes:
datatype,site,column,core,segment,avalanches_spotted,wind_loading,snowpack_depth,core_depth,slope_face,slope_angle,air_temperature,core_temperature,ect_* - Data format: PNG images, CSV labels
- Splits:
train(all preprocessed rows; no official train/val/test split — see Splits below),raw - Languages: English (for labels)
Splits
This dataset ships no train/validation/test split. Rows are individual photographs; many rows describe the same physical core, and many cores come from the same snowpit. A row-level or core-level split leaks — images of one core would appear on both sides. Split on `site`, which is one pit on one day and the coarsest unit in the data. With few sites (7), prefer leave-one-site-out cross-validation over a fixed split.
The train split contains all preprocessed rows (4040); a separate raw split holds the unprocessed source images.
Features
Labels
This dataset provides multiple classification and regression labels for each sample:
Supported Tasks and Leaderboards
The dataset supports image classification and regression tasks, generative AI projects and can be evaluated using standard accuracy, F1-score metrics or qualitative analysis of generative images. Examples of AI trained on this dataset are the snowGAN and coreDiff training on the magnified profiles of snow and snow cores respectively apart of this dataset.
Dataset Structure
Data Instances
Example:
{'image': <PIL.PngImagePlugin.PngImageFile image mode=RGB size=500x300 at 0x138F7BD70>,'filepath': 'preprocessed/cores/image1.png', 'datatype': 0, 'site': 0, 'column': 1, 'core': 1, 'segment': -1, 'coretemperature': 41.0, 'airtemperature': 7.0, 'ascendingmountain': 'Loveland Pass', 'citystatecountry': 'Silver Plume, Colorado, USA', 'collector': 'Denny Schaedig', 'coordinates': [39.65999984741211, -105.87999725341797], 'date': '1/12/25', 'time': '11:40 AM MST', 'snowpackdepth': 140.0, 'coredepth': 10.0, 'slopeface': 30.0, 'slopeangle': 11.0, 'avalanchesspotted': 2, 'windloading': 3, 'notes': 'Pilot using slightly modified coring method that took too large segment for snow profile picturing, advised not to use in higher level models. Two avalanche on south face nearby, partially cloudy with light snow in the last 24 hours. Lots of wind loading on eastern features.', 'ectresult': None, 'ecttaps': None, 'ectfailuredepthcm': None, 'ectfailuregrain': None, 'ecthardnessabove': None, 'ecthardnessbelow': None, 'ect_operator': None}
This is a pre-ECT row, so all seven ect_* values are None — not tested, as distinct from 'ect_result': 'X' (tested, no fracture in 30 taps). Every row currently in the dataset is in this state. Note also that datatype and wind_loading show as ints here because this is the load_dataset() view; the JSONL holds "core" and "high". See `load_dataset()` vs. reading the JSONL directly.
A row from a pit where the test was run:
{..., 'ectresult': 'P', 'ecttaps': 12, 'ectfailuredepthcm': 62.3, 'ectfailuregrain': 'FC', 'ecthardnessabove': '1F', 'ecthardnessbelow': 'F', 'ectoperator': 'Denny Schaedig'}
That is field notation ECTP12 — fractured on tap 12 and propagated the full column, on a facet layer 62.3 cm below the surface with a 1F slab over F facets.
Data Fields
image: The image location.file_path: Relative filepath of the image within the dataset repodatatype: The datatype of the image (i.e. core, profile, magnifiedprofile, crystalcard)site: The site number data was collected fromcolumn: The snowpack column number data was collected fromcore: The snowpack column number data was collected fromsegment: The segment number data was collected fromcore_temperature: Core temperature recorded for the core/profileair_temperature: Air temperature recorded at the start of the site collectionascending_mountain: The mountain on approach to the collection sitecity_state_country: City, state and country the site was closest toocollector: Name of the person collecting datacoordinates: Coordinates of the snowpit samples was collected fromdate: Date the snowpit was dug and sampledtime: Time the snowpit collection startedsnowpack_depth: Depth of the snowpack in cmcore_depth: Depth of the core sampled from in cmslope_face: The slope face angle.slope_angle: The angle of the slope.avalanches_spotted: Number of avalanches spotted on ascending mountainwind_loading: Level of wind loading features observed on nearby mountainsnotes: Additional notes provided by the collector about data collection from the site and it's nearby surroundingsect_result,ect_taps,ect_failure_depth_cm,ect_failure_grain,ect_hardness_above,ect_hardness_below,ect_operator: Extended column test, see below
Extended Column Test (ECT)
Seven columns record the Extended Column Test (Simenhois & Birkeland, 2006) — a mechanical stability measurement made at the pit, on the same spatial scale as the cores. They exist so snowpack stability can be modelled against a measurement taken where the core was, rather than against regional avalanche danger, which is a forecaster's judgement issued that morning for a ~100 km² zone.
These columns are `null` for every row currently in the dataset. They are collected from the 2026–2027 season onward.
Grain: one ECT per snow column
All seven values are at `(site, column)` grain. One test is run per snow column, so the same seven values repeat across every row sharing a (site, column) pair — exactly the way snowpack_depth and slope_face already repeat. Deduplicate on (site, column) before treating an ECT as an observation; treating each image row as an independent test will inflate your sample size by roughly two orders of magnitude.
Columns
Notation
Taps are loaded 1–10 from the wrist, 11–20 from the elbow, 21–30 from the shoulder.
null and "X" are different values
`null` means the test was not run. `"X"` means the test was run and the column did not fracture in 30 taps. These are opposite readings — one is an absence of information, the other is a strong stability observation — and conflating them silently corrupts any analysis built on the column. Do not fill null with "X", 0, or an empty string, and do not let a fillna() do it for you.
Invariants
Enforced on ingest by snowMaker/ect.py; a violation raises rather than being coerced:
ect_resultis one ofPV/P/N/X, ornull. Anything else is rejected.ect_result == "PV"→ect_tapsisnull(it propagated before any tap was applied).ect_result == "X"→ect_taps,ect_failure_depth_cm,ect_failure_grain,ect_hardness_aboveandect_hardness_beloware allnull(there was no fracture to describe).ect_result in ("P", "N")→ect_tapsis present and in 1–30.ect_tapspresent →ect_resultisPorN.ect_resultisnull→ all seven columns arenull.- All seven values are constant within each
(site, column)group.
There is deliberately no boolean or ordinal ECT column
The dataset stores the raw fields only. There is no ect_propagated boolean and no derived stability index, and neither should be added.
Collapsing the result to a P/N binary measured worse than the regional-danger baseline it replaces at every sample size — power 0.586 vs 0.639 at n=15, and 0.000 at n=7, where two levels cannot reach p ≤ 0.05 however clean the ordering. Keeping the tap count gives 0.881 at n=15. Deriving the scale downstream also means the encoding can change without a schema migration. A monotone ordinal, if you want one:
ECTPV < ECTP1..30 < ECTN1..30 < ECTX
0 taps 30+taps 61Limits
ECT is reliable in roughly the top metre. Deep persistent slabs sit below where 30 taps deliver useful energy and return a false-stable X — the characteristic Colorado failure mode, and it bears directly on how much to trust these labels. One test is also one point in a spatially variable snowpack.
load_dataset() vs. reading the JSONL directly
The two access paths do not agree on every column, because load_dataset() applies the dataset_info.features declared in this card while a direct read of metadata/preprocessed.jsonl gives you the literal JSON in the file. Any column typed as class_label is converted; everything else is passed through.
Code that consumes both paths must normalize datatype and wind_loading itself. The seven ECT columns are `Value("string")`/`int64`/`float64`, never `ClassLabel`, precisely so they need no such normalization — a ClassLabel would hand back integer codes on one path and strings on the other, and its null handling is version-dependent and unsafe (ClassLabel.encode_example(None) raises in datasets 5.x). These columns are null for every existing row, so that matters here more than anywhere else.
Schema change log
- 2026-08-17 — breaking.
datatypeandwind_loadingchanged from int to string in the JSONL. Downstream pair-indexing that compared them as ints silently produced zero matched pairs rather than erroring. The card'sclass_labeldeclarations were left in place, which is whyload_dataset()still returns ints on those two columns. Code reading the JSONL directly must accept both eras. - 2026-08-19 — additive, non-breaking. The seven
ect_*columns were added, nullable, with every existing row backfilled tonull. No existing column's type or value changed and the row counts are unchanged.
A note on coordinates and elevation
coordinates is logged at the full precision the receiver reports; nothing in the pipeline rounds or truncates it, and ingest warns on any site logged with fewer than 4 decimal places. Site 0 is a known exception, recorded as [39.66, -105.88] — two decimal places, a ~1.1 km box. In Loveland Pass terrain that box spans 3533–3748 m and straddles an elevation-band boundary, so a DEM lookup against site 0 is not trustworthy. The other sites carry 4–5 places.
There is deliberately no elevation column: elevation is recoverable from the GPS fix via a DEM (USGS 3DEP). Aspect is not recoverable that way, which is why slope_face (degrees true) is carried explicitly and must not be dropped.
Usage
from datasets import load_dataset
# Load the rocky mountain snowpack repo
dataset = load_dataset("rmdig/rocky_mountain_snowpack")
# Grab first sample in training set
sample = dataset["train"][0]
# Grab the image apart of the sample
image = sample["image"]
# Show the image
image.show()License
This dataset is distributed under the CC-BY 4.0 license.
Citation
@dataset{Schaedig2025RockyMountain, title = {Rocky Mountain Snowpack Dataset}, author = {Denny Schaedig}, year = {2025}, publisher = {RMDig.ai}, license = {CC-BY 4.0}, url = {https://huggingface.co/datasets/rmdig/rocky-mountain-snowpack} }
