← RCAI4S · Research Centre on AI for Science

ADF-STEM · restoration benchmark

STEM Denoising + 2× Super-Resolution

Predicting the clean 128×128 target from the noisy 64×64 input on 1500 simulated ADF-STEM patches (simulation/train_data1), scored on a held-out 20%.

33.17dBPSNR, 300 held-out patches
0.9650SSIM
+6.66dBover bicubic ×2
300/300patches beating baseline
Holds — simulation

A 0.6M-parameter residual CNN beats bicubic on every held-out patch.

+6.66 dB and +0.165 SSIM, and it resolves atom columns that are genuinely ambiguous in the low-dose input.

Does not hold — experiment

Do not point it at experimental data yet.

On two of the 15 real micrographs it paints lattice texture into amorphous regions and into the vacuum. That is hallucination, and it is the failure mode that matters.

01

Result

33.17 dB PSNR / 0.9650 SSIM on 300 held-out patches, against 26.51 dB / 0.7996 for bicubic upsampling — a gain of +6.66 dB and +0.165 SSIM. The model beats the baseline on 300 of 300 patches.

ModelPSNR (dB)SSIMvs. baseline
Bicubic ×2 (baseline)26.51 ± 3.060.7996 ± 0.1220—
EDSR-lite33.17 ± 3.160.9650 ± 0.0428+6.66 dB / +0.165
EDSR-lite + point_map head32.83 ± 3.120.9624 ± 0.0454+6.32 dB / +0.163

Spread is quoted across the 300 patches, not as an uncertainty on the mean. Patch-to-patch variation is large (about 3 dB) because effective dose varies between patches — in the scatter below the worst inputs sit near 15 dB bicubic and the cleanest near 32 dB.

The auxiliary head did not help. It cost 0.33 dB. Both models were scored on the identical split, so the comparison can be paired, and paired it is unambiguous: mean difference 0.333 dB, SE 0.011, t = 31, plain model ahead on 97% of individual patches. Small, but real rather than noise. The likely cause is capacity competition, not anything subtle — the point_map loss pulls the shared trunk toward sparse delta-like peaks, which is the opposite of what the smooth target needs. Use the plain model.

Worst, median and best held-out patches by model PSNR, shown as input, bicubic, model, target and a signed error map.
Fig. 1 Worst, median and best held-out patches by model PSNR. The error map is signed: blue = model too dark, red = too bright.
Left panel, test PSNR against epoch for both models. Right panel, scatter of model PSNR against bicubic PSNR for all 300 test patches.
Fig. 2 Left: convergence on the held-out set. Right: all 300 test patches, model versus bicubic.

The qualitative figure is the honest read on what the dB figure means. In the median and best rows the model recovers individual atom columns that are genuinely ambiguous in the input. In the worst row — a very low-dose patch where bicubic manages 16.6 dB — it produces a plausible lattice, but the error map shows systematic errors at column positions, so the reconstruction should not be trusted for measurement there.

Atom-column localisation

Since the auxiliary model predicts an atom map, it can be scored the way a microscopist would: detect peaks, match them to the true point_map within 3 px at 128×128 scale, and count.

QuantityValue
Precision0.910
Recall0.802
F10.853
Localisation RMSE (matched columns)0.96 px
True columns / detected31,126 / 27,436

Sub-pixel accuracy on the columns it finds, and it rarely invents one, but it misses 20% — and the misses concentrate in the low-dose patches, visible as the sparse bottom row below. That recall figure, not the 33 dB, is the number to quote if the downstream use is measuring lattice spacings.

Predicted atom map beside the true point map for an easy case and a low-dose case.
Fig. 3 Auxiliary head: predicted atom map against the true point_map. Easy case top, low-dose case bottom.
02

The data

input and target are not the same resolution, so this is not a pure denoising task.

FolderSizeGrey levelsWhat it is
input64 × 64~11Low-dose acquisition. Few distinct values means Poisson / counting-limited noise, not additive Gaussian.
target128 × 128~208High-dose image of the same field of view. Smooth, atom columns clearly resolved.
point_map128 × 128~168Sparse Gaussian blobs at atom-column positions. Ground-truth structure.

Grey levels are counted on a typical patch.

1500 triplets, filenames matching across all three folders. They come from three source simulations, 500 patches each: 20_ADF_45_200, 25_ADF_45_200, 30_ADF_45_200 (the leading number is presumably thickness or dose).

Two checks worth recording:

  • The fields of view are aligned. Correlation between input and the 2×-downsampled target is 0.78–0.95 across sampled patches, so the model has to interpolate, not register.
  • The patches do not overlap. Correlation between neighbouring patch indices from the same source is −0.14 to 0.30, i.e. no shared pixels. A random train/test split therefore does not leak.
03

Method

An EDSR-style residual CNN (0.6M parameters, 0.8M with the auxiliary head) that predicts a correction on top of a plain 2× upsample rather than the image itself.

aux conv atom map 128×128 AUX input 64×64 conv 3×3 64 ch 8 residual blocks conv + PixelShuffle ×2 upsample conv residual + output 128×128 bilinear ×2 GLOBAL SKIP

Solid path: the image head. Dashed path: the auxiliary atom-map head, used only in the ablation.

The global skip matters more than the block count. With only 1200 training patches, letting the network learn the low-frequency image from scratch wastes capacity; giving it the upsample for free means every parameter goes into deblurring and denoising.

The auxiliary head. A second 3×3 convolution branches off the same upsampled features and is supervised by point_map, with total loss L1(image) + 0.5 * L1(point_map). The idea is that forcing the trunk to localise atom columns should sharpen them in the restored image. Only the image head is used at evaluation; the atom map is a by-product. A model without this head was trained identically so the effect is measurable rather than assumed.

Training recipe

SettingValueWhy
LossL1Sharper reconstructions than L2, which regresses to the blurry conditional mean
OptimiserAdam, lr 2e-4, cosine decay to 0Standard for this architecture family
Epochs / batch80 / 16Loss and test PSNR both flat well before the end
AugmentationRandom flips + 90° rotationsSafe here: input and target are transformed identically, and the lattice has no privileged orientation
Precision / devicefp32 on Apple MPS~9–11 s per epoch

One implementation note: the skip inside the network is bilinear, not bicubic, because PyTorch 2.2 has no upsample_bicubic2d kernel for MPS. The bicubic baseline is still computed properly, on CPU.

04

How it was evaluated

One split: 1200 train / 300 test, stratified so each of the three source simulations contributes proportionally to both sides. sklearn.model_selection.train_test_split, seed 42. The two models share the identical split, so their numbers are directly comparable.

Nothing was selected on the test set. The epoch budget was fixed in advance and the last epoch is the one reported. Test PSNR is logged every 5 epochs, but only as a convergence trace; no early stopping, no checkpoint picking, no hyperparameter search against it.

Metric conventions, which matter because PSNR and SSIM are both reported with several incompatible defaults in the literature:

  • Images are scaled to [0, 1] and predictions clamped to that range; data_range = 1.0.
  • PSNR and SSIM are computed per patch and then averaged over the 300, rather than on a concatenated array. The standard deviation quoted is across patches.
  • SSIM uses Gaussian weighting, sigma 1.5, use_sample_covariance=False — the Wang et al. (2004) formulation, not scikit-image's default 7×7 uniform window, which reads ~0.01–0.02 higher.
  • The baseline is bicubic 2× upsampling of the input, scored against the same targets. Every gain quoted is relative to that, not to the raw 64×64 input, which cannot be scored against a 128×128 target at all.
05

Caveats

Four things these numbers do not establish.

  1. This is simulated data, and the targets are unusually clean. The target images are smooth and nearly noise-free, so the mapping is closer to “denoise and sharpen toward a known blur kernel” than to the harder experimental case. Expect a drop on the real micrographs in test/.

  2. Patches from all three source simulations appear in both train and test. The split is random and stratified, which is the number normally reported, but it measures interpolation within known conditions — not generalisation to a new specimen, thickness or dose. A leave-one-source-out run would measure that, and it would almost certainly score lower.

  3. PSNR and SSIM reward smoothness. Both metrics tolerate slightly over-smoothed output, which for atomic-resolution imaging is exactly the failure mode that matters — a blurred atom column still scores well but is useless for measuring a lattice spacing. This is why the atom-localisation numbers are worth more than the dB figure.

  4. One split, one seed. With 300 test patches the standard error on mean PSNR is roughly 0.17 dB, so an unpaired comparison of the two models could not resolve a gap of a few tenths of a dB. The paired test in the results section does resolve it, which is why it was used — but it still speaks only for this split and this seed. Run several seeds before generalising the auxiliary-head finding.

06

On the real micrographs

The model was run on all 15 experimental images in test/. It scores far worse than bicubic on every one of them — mean 20.78 dB vs 29.61 dB — but the metric is measuring the wrong thing, and the pictures tell a different story from the numbers.

There is no ground truth for these micrographs, so PSNR and SSIM cannot be computed the way they were on the simulation. The protocol used instead: take the real image as the reference R, downsample it 2× to make the model's input, and score the restoration against R. This mirrors the training relationship exactly, and bicubic 2× is scored identically.

PSNR (dB)SSIM
Bicubic ×229.610.7172
Model20.780.5013
Difference−8.83−0.216

Mean over the 15 experimental images, scored against the downsample-proxy reference described above.

Why the metric is misleading here. R is a noisy experimental image. The model was trained to remove noise, so every grain of noise it correctly suppresses counts against it, while bicubic — which cannot denoise but also cannot over-smooth — stays closer to the noisy reference. The effect is largest on the images that were already denoised by other means before being saved: Page1_..._result, the two Radial Wiener images and the iDPC images, where bicubic reaches 33–48 dB because downsample-then-upsample is nearly lossless on a smooth image. On the genuinely noisy raw micrographs the gap is much smaller, 1.1–1.6 dB.

Grid of 256 by 256 centre crops from the real micrographs, each shown as reference, bicubic and model output.
Fig. 4 Real micrographs: 256×256 centre crops, reference vs bicubic vs model.

What the pictures actually show, which is the part worth acting on:

  • Page1, Page3, Page7 — crystalline, lattice spacing within 1.0–1.5× of the training scale. The model output is visibly cleaner than either the reference or bicubic, with atom columns resolved that are ambiguous in the raw data. This is the model working as intended, and the negative dB is an artefact of scoring against a noisy reference.
  • Page5 and Page8 — the model paints a regular lattice texture into amorphous regions and into the vacuum, where the reference has no periodic structure at all. This is hallucination, and it is the failure mode that matters. Both images have lattice spacings 3.8–4.0× coarser than anything in training, so the network is fitting its learned periodicity onto content it has never seen.
  • Page4 — very low dose. The output is plausible but the SSIM of 0.12–0.23 says the fine structure is not being recovered reliably.
Caution on the evidence

The visual difference between “denoised” and “hallucinated” is clear to the eye in these crops, but without ground truth there is no clean number that separates the two. A spectral test confirms the model amplifies periodic structure 20–90× relative to the broadband background on every image, including the ones where the result looks right — which is exactly what denoising does too, so that number cannot distinguish the good cases from the bad.

Practical conclusion

Do not apply this model to experimental data yet, and do not read the −8.83 dB as the model being broken. The two real problems are the scale mismatch and the absence of any experimental training data. The fixes, in order: resample each micrograph so its lattice spacing matches the training scale before inference; add scale augmentation to training; and score on experimental pairs (same field imaged at low and high dose) rather than on this downsample proxy.

07

Reproducing it, and what to do next

Everything lives in the workshop-0922/ project folder.

FileWhat it does
train.pyLoads the data, trains, evaluates, writes out/results_<tag>.json and the model weights
make_figures.pyBuilds the three simulation figures from the saved predictions
eval_atoms.pyAtom-column detection scoring for the auxiliary head
out/Metrics, predictions, trained weights, training logs
figures/The PNGs
python3 train.py --tag plain              # image head only
python3 train.py --tag aux --aux          # + point_map head
python3 make_figures.py && python3 eval_atoms.py

About 13 minutes per run on an M-series Mac (MPS, fp32). No CUDA needed, no data preprocessing step — train.py reads the PNGs directly.

Worth doing next, roughly in order of what each would teach you

  1. Leave-one-source-out. Train on two of the three simulations, test on the third. This is the number that predicts behaviour on a new specimen, and it is the honest headline if the model is ever pointed at unseen data. About 40 minutes for all three folds.

  2. Fix the scale mismatch on test/. Resample each micrograph to the training lattice scale before inference, then re-read the crops in Fig. 4. This tells you immediately whether the simulation-to-experiment gap is small or fatal. Ten minutes.

  3. Add train_data2. It has input and target but no point_map, so it can train the image head. Doubling the data is usually worth more than any architecture change at this scale.

  4. Then, and only then, a bigger model. 8 residual blocks at 64 channels is deliberately small. Scaling up is the obvious lever, but it is the least informative one until the generalisation question above is settled.