CoolFace
Modelpublic

jan-grzybek/aerial-film-restorer

sourceHugging Facecc-by-4.0updated 10d agoView on Hugging Face
3likes92downloads
Model Card

Aerial Film Restorer

A measured-fidelity restoration model for scanned historical aerial survey film, at every scale a survey is served.

Mid-century aerial archives reach the reader as mosaics of individually exposed plates, under vignette haze, with local contrast lost to print and storage, film grain, and the damage of scanning. This model removes that damage without inventing content, and it does so from a 23 m/px overview to a 0.1 m/px close-up, because it reads eight planes of the same ground together with a plane that tells it which pitch it is working at. Its fidelity is a measured quantity: every number and every image on this card comes from held-out photograph pairs with professionally produced ground truth that the training never saw.

[image]

Figure 1. Bourges, IGN survey flight, 1950s; a held-out pair. Left: the scan, with a plate boundary running diagonally through the edge of the town, the lower plate about a stop lighter and the market gardens south of the railway washed out. Middle: the model. The step is gone and the plots read at one tone with the fields north of the line. Right: the professionally cleaned orthophoto IGN produced from the same photographs, the ground truth every number on this card is measured against.

1. Method

Input. The network sees eight 256-pixel planes centred on the same ground: the crop at the working pitch; registered views of that ground at 2×, 4× and 8× coarser pitch, each the central fraction of a wider window upsampled to align with the crop pixel for pixel; the raw 2×, 4× and 8× windows themselves, which carry context out to 2,048 input pixels; and a constant plane that encodes the pitch as −20·log₂(m/px ÷ 21), zero at the coarsest frame the model serves. In training the coarser planes come from a block-mean pyramid of the degraded sheet, so inference needs nothing but an image and its pitch: build the same pyramid, walk a 256-pixel lattice with 64 pixels of overlap, blend the outputs with a Hann window. The widest window spans 6 km at 2.9 m/px and 48 km at 23 m/px, so a plate's offset is judged from outside the plate at every frame, by one network that knows which frame it is on.

Architecture. NAFRestore, this project's own implementation in the NAFNet family (Chen, Chu, Zhang and Sun, Simple Baselines for Image Restoration, ECCV 2022; arXiv:2204.04676, megvii-research/NAFNet, MIT), 17.32 M parameters, eight input channels. The head adds a bounded correction to the input and clamps, so the model works in the native 0–255 domain.

Training. Clean orthophoto content is degraded synthetically, with plate quilts of individually exposed plates, a per-mosaic lens signature, haze, grain, and print and scan damage, each calibrated against a 1935–1994 municipal archive; the clean sheet is the target. The content comes from IGN's 1950–1965 historic orthophotos, swisstopo's 1946 imagery, and current French and Polish orthophotos at the fine pitches, across nine frames from 23 m/px to 0.09 m/px. Real pairs of raw scan and professional clean serve selection and measurement only; their fine detail does not align well enough to be a training target. The model was warm-started from this project's previous, five-channel release, itself trained from scratch, with the stem weights of the new channels zeroed. There is no external base model. The known failure modes of naive training, aliased targets that manufacture sharpness, adversaries that trade grain for punch, and losses that flatten tone, were each identified by measurement and removed.

2. Evaluation

The instrument. Three energy ratios against ground truth, where 1.00 means exactly what the negative holds: detail on structure, the high-frequency energy over edges and built fabric; detail in flat film, the same over flat regions, which is grain; and large-scale tone, the spread of the low band at σ = 24 px. They are evaluated on 40 pairs reserved from training, raw scans against IGN's cleaned orthophoto of the same photographs, eight crops per pair at 2.9 m/px, and reported as mean ± standard error. The headline numbers use the 22 of those 40 that no checkpoint selection ever saw.

The tone ratio conflates two errors: a model that flattens the ground's own tonal variation and leaves plate residue can read 1.00. It is therefore decomposed by regressing the output's low band on the truth's. The slope is the share of real tone kept; the spread of the residual, relative to the truth's, is the low band the truth does not explain: leftover plate steps, lens signature, invented tone.

The per-frame gate. A survey is served at nine pitches, and the previous release had been judged at one of them. The three ratios are therefore also measured at five frames, from 23 m/px to 1.5 m/px, on 14 pairs with six crops each.

3. Results

Table 1. The 22 unseen pairs at 2.9 m/px, mean ± s.e. The previous release is the five-channel model this one grew from, measured on the same crops.

axisthis modelprevious release
detail on structure1.006 ± 0.0241.023 ± 0.028
detail in flat film1.053 ± 0.0451.075 ± 0.049
large-scale tone, ratio1.152 ± 0.0661.006 ± 0.069
large-scale tone, share kept0.911 ± 0.0410.817 ± 0.039
large-scale tone, unexplained0.611 ± 0.0650.496 ± 0.059

Table 2. Five frames, 14 pairs × 6 crops; structure / grain / tone ratio.

framem/pxprevious releasethis model
z1223.41.02 / 1.21 / 3.880.90 / 1.10 / 1.44
z1311.71.14 / 1.26 / 2.741.03 / 1.15 / 1.89
z145.91.07 / 1.16 / 1.371.03 / 1.11 / 1.46
z152.90.98 / 1.07 / 0.920.96 / 1.05 / 1.09
z161.50.95 / 1.11 / 0.780.91 / 1.08 / 0.94

At 2.9 m/px the two models are level on structure and grain within the instrument, and they differ on tone in a way the ratio alone hides. The previous release reads 1.01 by keeping 0.82 of the ground's own tone and leaving 0.50 unexplained; this model reads 1.15 by keeping 0.91 and leaving 0.61. It preserves more of the real tone and leaves a quarter more residue. Side by side at this frame, the two are hard to tell apart.

Across frames the picture is not level. At 23 m/px the previous model leaves most of the plate quilt in place, at 3.88 times the truth's large-scale spread; this one brings it to 1.44, an 85 percent reduction of the error. At 11.7 m/px the error halves, and at 1.5 m/px tone rises from 0.78 to 0.94. Grain sits nearer truth at every frame. Structure reads lower than the previous model's at every frame: nearer truth at 11.7 and 5.9 m/px, where the previous model over-sharpened, and a few points under it at 2.9, 1.5 and 23 m/px, where this model is a shade softer.

[image]

Figure 2. Bordeaux, IGN survey flight, 1950s; likewise held out. A plate step runs across the middle of the scan, the lower plate about a stop lighter. After the model it survives as no more than a faint trace, and the fields on either side read as one ground. Right: the professional clean.

[image]

Figure 3. Warsaw, June 1944; a German reconnaissance print from GUGiK's archival aerial service, wholly outside the training distribution, since no historical Warsaw imagery has ground truth. Left: the print. Right: the model. The hand-lettered "Warschau", the target marks and the bright cleared ground in the middle of the frame are part of the print's face and are treated as content.

[image]

Figure 4. Warsaw city centre, 4 November 1944, a month after the city's fall, during its systematic razing. Captured German reconnaissance film from the US National Archives (RG 373, public domain by age), orthorectified by the Warsaw Rising Museum. Left: the scan. Right: the model. The exposure evens out under the drifting smoke, and block after block of roofless shells becomes individually legible.

[image]

Figure 5. Warsaw, 1951; the reconstruction-era survey (City of Warsaw open data, © m.st. Warszawa), 11.7 m/px. Left: a mosaic of dozens of survey plates, each at its own exposure, some a stop dark and some blown. Right: the model. The patchwork resolves into one continuous city, across the river too, with no more than a faint trace of a plate edge here and there, while every block stays what the film recorded.

All panels are the bare model through the shipped inference.py, with no contrast grading and no sharpening, rendered by a script in the project's repository so they can be re-made from the file below.

4. Using the model

bash
pip install -r requirements.txt
python inference.py scan.tif restored.png --mpp 1.5      # the input's pitch, m/px

As a library, over a folder (run from this folder, or put it on PYTHONPATH):

python
import glob
import numpy as np
from PIL import Image
from inference import load_model, read_image, restore

net = load_model("model.safetensors")          # picks cuda / mps / cpu itself
for path in sorted(glob.glob("archive/*.tif")):
    out = restore(net, read_image(path), mpp=1.5)   # float32 [0,255], input-sized
    Image.fromarray(out.astype(np.uint8)).save(path.replace(".tif", "-restored.png"))
  • —Input. Single-channel film imagery at 0.09–23 m/px, any size of 64 px or more, and its ground pitch in metres per pixel. read_image handles 8-bit and 16-bit files (PIL's convert("L") clips 16-bit scans, hence the reader) and reduces colour to luminance.
  • —The pitch is part of the input. One network serves nine frames because the scale plane tells it which one it is on. A wrong pitch is a silent wrong answer, so there is no default. For a georeferenced scan the pitch is the ground distance per pixel; for a Web-Mercator tile it is 156543 · cos(lat) / 2^z metres per pixel of a 256-pixel tile.
  • —The tile is the contract. The model was trained and validated at 256 pixels, and the geometry of its planes is defined by that window; there is no tile-size option on purpose.
  • —Restore whole frames. The widest context window is 2,048 input pixels across; a frame narrower than that sees reflection in that plane. That is fine for restoration, but the plate-scale correction is only as good as the ground in view, so crop the output rather than the input.
  • —No-evidence shield, on by default. Black voids and blown margins are detected, filled from the nearest valid content for the model's eyes only, and returned verbatim in the output; nothing is invented outside coverage. Off-white scan paper and mounts are in range and count as content, so crop scans to the photograph, or mask processing to the image face.
  • —Throughput. About 1.7 MPix/s on Apple-silicon MPS, measured with the shipped code (a 1024² frame in 0.6 s); CUDA is faster, CPU workable for single frames. Memory is bounded by the fixed tile, so arbitrarily large frames stream through.
  • —Determinism. Same input, same output. The shipped tests.py (pytest) pins the geometry of all eight planes, the scale plane, the shield, the small-input path and the hash of the weights.

5. The checkpoint

filesha256measured
model.safetensorsed5374045efd9fd8…Tables 1 and 2; 17.32 M parameters, eight input channels

One model, one file. The previous release, five channels without a scale plane and with a gentler variant beside it, remains reachable at revision v1.0 of this repository for anyone who built on its input contract. That file does not load in this code, nor this file in that.

6. Limitations and intended use

  • —The model restores; it does not invent. Saturated regions stay saturated: the training deliberately under-represents saturation so that the model never learns to hallucinate content into areas without evidence.
  • —Below about 0.8 m/px no film ground truth exists anywhere. The fine frames learn their content from modern orthophotos, and behaviour there should be judged by eye.
  • —Colour imagery: restore the luminance and carry the chroma separately.
  • —Scan mounts and paper margins that are not saturated white count as content. The shield catches voids and blown regions, not off-white paper, whose median was about 247 in the prints examined.
  • —A single tonal step wider than the widest context, about 2 km at 1 m/px, crossing textureless ground such as open water or blank fields, may survive in part. Ordinary plate patchwork resolves fully, and through textured ground even large seams dissolve.
  • —Outputs are derivatives of the input imagery: the source's terms and credits apply to them unchanged, and no rights are claimed in restored pixels. If you serve restored imagery, label it as machine-restored and keep the provenance (model, weights hash, date), the transparency posture this model ships under (cf. EU AI Act, art. 50).

7. Training data and attribution

sourcerolelicence
IGN France, raw scans and BD ORTHO® Historique 1950–1965the paired ground truth; wide sheets at the coarse frames; the razed-cities passesLicence Ouverte / Etalab 2.0
IGN France, BD ORTHO® (current orthophoto, Géoplateforme)fine-scale content at 0.2–0.8 m/pxLicence Ouverte / Etalab 2.0
swisstopo, SWISSIMAGE HIST 1946fine-scale film contentswisstopo open data
GUGiK (Poland), Warsaw orthophoto 2025content at working scales, 0.4–3 m/pxPolish geodetic open data, art. 40a(2)
GeoNames, cities15000 / cities5000 registerschose every French and Swiss harvest siteCC BY 4.0

8. Licence

  • —Weights: CC BY 4.0. Use them for anything, including commercially; credit "Aerial Film Restorer, the Powidok project (powidok.waw.pl), CC BY 4.0" and indicate modifications such as fine-tuning.
  • —Code (`model.py`, `inference.py`, `tests.py`): MIT (LICENSE-CODE).

9. Provenance

The model was developed for Powidok, an interactive historical map of Warsaw 1935–1951, currently under construction, where it restores a fourteen-layer aerial fleet. The measurement methodology, with its reserved-pair evaluation, standard errors, selection-bias checks, per-frame gates and deployment gauntlets, and the record of this retrain come from that project's training records. In that deployment a ×2 super-resolution model, released as Aerial Film SR ×2, adds one level of detail before this model runs, and a solved per-layer contrast normalisation, a cross-zoom consistency step and a labelled presentation-sharpening pass ride on top of it; anything similar is a downstream choice this release leaves to you.

Citation

Grzybek, J. (2026). Aerial Film Restorer: measured-fidelity restoration of
historical aerial survey film. The Powidok project, https://powidok.waw.pl
(model: https://huggingface.co/jan-grzybek/aerial-film-restorer)