Noise2Noise researchers test L1 vs L2 in self-supervised denoising on Kodak24 and SIDD; Retraining on SIDD pairs yields 9.4–11.0 dB gains

Noise2Noise researchers test L1 vs L2 in self-supervised denoising on Kodak24 and SIDD; Retraining on SIDD pairs yields 9.4–11.0 dB gains Noise2Noise Revisited: Training Pair Distributions Dominate Loss Choice in Self-Supervised Denoising Noise2Noise Revisited: Training Pair Distributions Dominate Loss Choice in Self-Supervised Denoising Dingyan Shang* Independent Researcher Frisco, USA dingyanshang@gmail.com *Corresponding author Zhenyu Xu Independent Researcher Fulshear, USA zhenyuxu0918@gmail.com Youting Wang Independent Researcher Mountain View, USA wang.yout@northeastern.edu Bonan Shen Independent Researcher Long Island City, USA shenbonan2@gmail.com Bowen Liu Independent Researcher South San Francisco, USA bliu0962@usc.edu Abstract—Noise2Noise (N2N) trains denoisers on pairs of inde- pendently corrupted observations, eliminating clean references. We stress-test two natural conjectures about why the L1 loss outperforms L2 here. First, the hypothesis that the L1 loss confers robustness via parameter sparsity confuses the loss with Lasso regularization: an explicit Lasso penalty produces the predicted sparsity yet fails to reproduce L1’s cross-noise behavior, while L1- and L2-trained weight distributions are indistinguishable. Second, the population optima of the two losses coincide exactly for symmetric signal posteriors and nearly so for concentrated ones. Measured differences are therefore dominated by optimiza- tion dynamics (bounded-influence gradients), which we probe with gradient statistics and contaminated-target training. On Kodak24 with five synthetic noise families, the L1 loss holds a statistically significant edge over L2, below 1 dB PSNR, holding across three seeds on 13 of the 14 noise columns. On real camera noise the loss is not the decisive variable in distribution: on official SIDD validation blocks, synthetic-Gaussian-trained N2N models gain only 0.8 to 3.7 dB over the noisy input regardless of loss, while retraining on SIDD’s own noisy pairs, never reading ground truth, gains 9.4 to 11.0 dB, far ahead of BM3D. All metrics are on raw network outputs, and the study makes no leaderboard claim. The training pair distribution, not the loss, carries the inductive bias. That design rule applies wherever clean references are unobtainable, from microscopy to industrial inspection sensors. Index Terms—image denoising, self-supervised learning, Noise2Noise, robust statistics, real-noise benchmarks, industrial inspection, nondestructive evaluation I. INTRODUCTION Supervised deep denoisers [1], [2] require registered clean/noisy pairs, costly or impossible in microscopy, in low- light photography, and in industrial nondestructive testing, where each unit under inspection is unique and a noise-free reference of it is physically unobtainable [3]. Noise2Noise (N2N) [4] showed that two independent noisy observations of the same signal suffice: under zero-mean noise, the network trained to map one observation to the other converges in expectation to the clean-target minimizer. Two practical questions remain. (i) How much of N2N’s observed cross-noise robustness is attributable to the loss function? (ii) The empirical edge of the L1 loss over L2 in restoration networks is documented [5] and attributed there to optimization behavior; is that attribution right, or does the credit belong to parameter sparsity? We answer both with a controlled protocol and make three contributions: ...

September 16, 2026