TheDagbanja/L-Band_DLPU
LB-DLPU: An L-Band (NISAR/UAVSAR) Benchmark for InSAR Phase Unwrapping LB-DLPU is a physically-simulated benchmark for interferometric SAR (InSAR) phase unwrapping at L-band, calibrated to the NISAR (spaceborne, 20 m) and UAVSAR (airborne, 6 m) regimes. Each of the 10,000 patches ships with the wrapped phase, the ground-truth absolute phase, coherence, per-edge integer ambiguity labels, residues, and a validity mask — everything needed to train, validate, and test… See the full description on the dataset page: https://huggingface.co/datasets/TheDagbanja/L-Band_DLPU.
LB-DLPU: An L-Band (NISAR/UAVSAR) Benchmark for InSAR Phase Unwrapping
LB-DLPU is a physically-simulated benchmark for interferometric SAR (InSAR) phase unwrapping at L-band, calibrated to the NISAR (spaceborne, 20 m) and UAVSAR (airborne, 6 m) regimes. Each of the 10,000 patches ships with the wrapped phase, the ground-truth absolute phase, coherence, per-edge integer ambiguity labels, residues, and a validity mask — everything needed to train, validate, and test learning-based and classical unwrappers.
Two properties set it apart from existing (C-band, RMSE-only) PU datasets:
- Well-posedness certificate. Every scene is provably recoverable: a noiseless-oracle minimum-cost-flow (MCF) unwrapper reconstructs each patch to within 0.035 rad given the correct per-edge costs (100% of 10,000 scenes pass; max clean-oracle RMSE = 0.0347 rad). Any error a method incurs is attributable to the method, not to an unsolvable target.
- L-band-specific difficulty. Calibrated coherence (Beta fits to real granules), Cramér–Rao phase noise, an ionospheric screen (NISAR), and steep near-fault gradients that push the true per-edge ambiguity into the five-arc range {−2,…,+2}.
Dataset at a glance
Per-regime statistics (from datasheet.md):
Directory layout
LB_DLPU/
├── README.md # this card
├── datasheet.md # auto-generated statistics + well-posedness certificate
├── dem_manifest.csv # DEM tile → split assignment (tile, split, region, regime)
├── index.jsonl # one JSON record per patch (metadata, no arrays)
├── assets/ # figures used in this card
│ ├── preview.png
│ └── dem_tiles_map.png
├── sim/
│ ├── train/ 000000.h5 … 007999.h5 (8,000)
│ ├── val/ 008000.h5 … 008999.h5 (1,000)
│ └── test/ 009000.h5 … 009999.h5 (1,000)
└── sim_wrapped_png/ # wrapped-phase quicklooks, mirroring sim/
├── train/ 000000.png … 007999.png
├── val/ 008000.png …
└── test/ 009000.png …Patch ids are shared across sim/<split>/<id>.h5 and sim_wrapped_png/<split>/<id>.png.
Per-patch HDF5 schema
Each .h5 file (≈0.7 MB) contains:
psi is the noisy wrapped observation and phi the clean target; they are not exactly congruent (that is the noise the unwrapper must survive). The per-edge labels satisfy Δφ_e = W(Δψ)_e + 2π·k_e, where W wraps to (−π, π].
Per-patch attributes (HDF5 .attrs): sensor (nisar|uavsar), difficulty (smooth|mixed|dense), mean_coherence, px_m (pixel spacing), NL (looks), residue_count, residues_per_mp, max_grad_rad_per_px, frac_edges_k1, frac_edges_k2, label_clip_frac, water_frac, clean_oracle_rmse, seed, and components (JSON: topography / deformation / atmosphere / ionosphere provenance).
index.jsonl
One record per patch with the same metadata as the HDF5 attributes plus id and split, for fast filtering without opening every file:
{"id": "000000", "split": "train", "sensor": "nisar", "difficulty": "smooth",
"mean_coherence": 0.56, "NL": 8, "px_m": 20.0, "residue_count": 2182,
"residues_per_mp": 33294.7, "clean_oracle_rmse": 0.0, "seed": 939529293,
"components": {"topo_src": "dem", "topo": {"B_perp_m": 32.97}, ...}}Splits
Train / val / test are disjoint by whole DEM tile: the 29 Copernicus GLO-30 tiles (worldwide tectonic, volcanic, and glacial terrain) are partitioned 17 / 5 / 7, so no terrain is shared across splits and the test set measures generalization to unseen geography. The exact tile → split assignment is in dem_manifest.csv.
Loading
import h5py, glob
def load_patch(path):
with h5py.File(path, "r") as f:
return {k: f[k][:] for k in f}, dict(f.attrs)
for p in sorted(glob.glob("sim/test/*.h5"))[:1]:
arrays, attrs = load_patch(p)
psi, phi = arrays["psi"], arrays["phi"] # input, target
print(attrs["sensor"], attrs["difficulty"], psi.shape)Quicklooks are plain PNGs:
from PIL import Image
Image.open("sim_wrapped_png/test/009000.png") # wrapped-phase previewIntended use
Training and benchmarking L-band phase-unwrapping methods — deep networks (wrap-count regression/classification, gradient estimation) and classical / minimum-cost-flow solvers — with per-regime × difficulty evaluation. The well-posedness certificate makes the test split a fair ceiling reference; the per-edge labels support both pixel-wise and edge-wise supervision.
Reference baselines
A method-blind evaluation harness scores every unwrapper on the identical test split, per regime × difficulty, on RMSE, MAE, PSNR, SSIM, a cycle-slip (jump) rate, residue count, and five-arc |k|≥2 edge accuracy. Reference findings across classical, minimum-cost-flow, and in-domain-trained deep baselines:
- Deep networks train cleanly on this data and outperform classical and statistical solvers (e.g. SNAPHU) — especially on the dense/high-gradient stratum, which the benchmark is designed to stress.
- The minimum-cost-flow family is residue-free by construction; with the correct per-edge costs the certified oracle ceiling reaches near-zero, residue-free error, while gradient-domain methods leave residues.
- The benchmark is not saturated: a gap to the oracle ceiling remains for every deployable method, and the strata form a genuine difficulty gradient (error widens smooth → mixed → dense).
The full leaderboard, metric definitions, and significance tests are in the accompanying code/paper.
Limitations and considerations
- Synthetic ground truth.
phiis physically modelled (topography from real Copernicus DEMs; coherence, noise, and ionosphere calibrated to real NISAR / UAVSAR granules) but is not field-validated absolute truth. It is intended for supervised training and controlled benchmarking; real-scene generalization should be assessed separately. - Two regimes. Sensor characteristics are approximated for NISAR-like and UAVSAR-like acquisitions; other L-band sensors may differ.
- No decorrelation-only patches. Every scene is certified recoverable given correct costs; the benchmark isolates cost/prior estimation, not irrecoverable-noise regimes.
Provenance and licensing
- Topography: Copernicus GLO-30 DEM (© ESA / Copernicus; free and open, attribution required).
- Noise / coherence / ionosphere models: calibrated to real NISAR L2 GUNW and UAVSAR granules.
- Simulated phase, labels, and quicklooks: this release.
License: CC-BY-4.0. Free to use, share, and adapt with attribution. The Copernicus GLO-30 DEM attribution above must be retained. If you intend a different license, update both this line and the license: field in the card metadata before publishing.
Citation
If you use LB-DLPU, please cite both the dataset and the accompanying paper.
Dataset (Zenodo):
@dataset{dagbanja_lbdlpu_data_2026,
title = {LB-DLPU: An L-Band (NISAR/UAVSAR) Benchmark for InSAR Phase Unwrapping},
author = {Dagbanja, S. and Qian, J.},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.21768604},
url = {https://huggingface.co/datasets/TheDagbanja/L-Band_DLPU}
}Paper:
@article{dagbanja_lbdlpu_paper_2026,
title = {LB-DLPU: A Recoverability-Certified L-Band InSAR Phase-Unwrapping Benchmark for Reliable Deformation Retrieval},
author = {Dagbanja, S., Qian, J. and Haitao, L.},
journal = {#Will be updated upon publication},
year = {2026},
note = {under review}
}Preprint
Dagbanja, Simon and Qian, Jiang and Lv, Haitao, LB-DLPU: A Recoverability-Certified L-Band InSAR Phase-Unwrapping Benchmark for Reliable Deformation Retrieval (August 08, 2026). Available at SSRN: https://ssrn.com/abstract=7358019 or http://dx.doi.org/10.2139/ssrn.7358019
