arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2603.01868v2 [eess.SP] 24 Mar 2026

Plug-and-play forward backward algorithm to restore Landsat images: A preliminary step to uncover the history of surface waters

Pierre Audisio    Barbara Belletti    Nelly Pustelnik
Abstract

The temporal and spatial analysis of river dynamics is a key factor for studying and understanding human impacts on floodplains. To assess the changes taking place, it is necessary to have high-resolution images with a large spatial coverage and a high temporal revisit frequency over the long term. Satellite imagery meets several of these criteria. For instance, Sentinel data provide high-resolution images but only after 2015. Therefore, to study water surface evolution prior to this date, it is necessary to rely on other satellite images such as Landsat, which offers longer historical coverage, albeit with lower spatial resolution. In this study, we aim to increase the spatial resolution of Landsat data from 30 to 10 meters (resolution of Sentinel images). To achieve this goal, we develop an innovative single image super-resolution method based on a plug-and-play approach.

Index Terms: 
Restoration, plug-and-play, remote sensing, neural networks, multispectral data
address: 1Ens de Lyon, CNRS, Laboratoire de Physique, F-69342, Lyon, France.
2Université Jean Monnet, CNRS, UMR 5600 Environnement Ville Société, Saint-Etienne, France.

1 Introduction

Context - Monitoring river ecosystems is a challenging task due to the complex interactions between ecological and morphological processes. Their complexity lies in their manifestation at multiple spatial scales going from regional or entire watershed to local river sections [1]. These ecosystems are particularly vulnerable to long-term changes driven by anthropogenic pressures, such as urbanization, dam construction, and land use changes [2]. To effectively capture these dynamics, remote sensing offers a powerful solution. However, accurate monitoring requires satellite data with three key characteristics: large spatial coverage, to capture interdependent regions; high spatial resolution, to detect fine-scale features; and short revisit periods, to track temporal changes. Satellite multispectral imagery is particularly well suited to this issue, due to the possibility of studying the occupation of land at specific wavelengths, but also thanks to the availability of open-access data at different spatial resolutions, from 10 to 60 meters depending on the satellite. However, high spatial resolution data (\sim10 m/pixel) has only become available since the early 2000s. Consequently, to capture more detailed and extended temporal information, it is necessary to enhance the resolution of satellite images with lower resolutions, such as those from Landsat.

Contribution - In this study, two sources of data are used. Sentinel-2 images, part of the Copernicus project, which started in 2015 [3], and Landsat project images that started in 1972 [4]. Sentinel-2 images have a spatial resolution of up to 10 meters, while Landsat spectral bands have a resolution of 30 meters. The goal of this work is, therefore, to obtain higher resolution images over a longer period of time, and a greater density of images over the years covered simultaneously by Sentinel-2 [3] and Landsat [5, 4] satellites. To do so, we adopt an inverse problem approach to reconstruct high-resolution images from Landsat data (\sim30 m/pixel), using Sentinel-2 images (\sim10 m/pixel) as the spatial resolution reference. To achieve this goal, we develop a Plug-and-play (PnP) approach, named Spect-FB-PnP, dedicated to multispectral reconstruction, that combines the stability of standard inverse problem formulation and the expressivity of deep neural network models.

Related works - Single image super-resolution is a well known problem that has been solved for many years considering variational approaches including the well-known use of total variation (TV) penalization [6] or penalization based on wavelet transform [7]. However, despite the effectiveness of these classical techniques, the ability of deep learning to achieve outstanding restoration quality has driven a shift towards neural networks approaches, where model parameters or priors are now learned directly from data [8]. A wide variety of deep learning frameworks have been made to achieve these tasks [9], starting from supervised model like convolutional neural networks (CNN) based methods [10], to more recent unsupervised approaches using generative adversarial networks (GAN) [11]. Both supervised methods [12] and unsupervised method [13], are able to achieve satisfactory results on super-resolution. However, on the one hand, supervised approaches requires a significant amount of data for their training, which is particularly challenging with remote sensing data, as supervised methods require manually prepared pairs of high- and low-resolution images. Additionally, when handling real satellite data, the pairs of high- and low-resolution images are usually unavailable, because only one satellite image can be acquired at a specific location and time. A common way to handle this problem is to generate low-resolution images using the bicubic interpolation method as a downsampling method [14], but this degradation does not correctly recreate real world low-resolution images [15]. On the other hand, although unsupervised methods have the significant advantage of not requiring high-resolution data for training, their use in remote sensing imagery remains very limited. In addition, the more efficient methods using GANs need to make very strong assumptions about the distribution of data for training the prior, which is not always known on remote sensing images, and this can lead to a deterioration in results and the creation of artifacts [16]. Supervised approaches also involve models other than CNNs, such as Transformers [17]. The term “Transformer” refers to the self-attention mechanism, which models the global relationships among image regions, eliminating the reliance on content-independent convolution kernels used in CNNs. One of the most promising methods in image super-resolution using this approach is SwinIR [18].

Finally, model-based neural network approaches, including unfolded neural networks and Plug-and-Play (PnP) methods, offer a compelling alternative by combining the interpretability of variational methods with the expressiveness of learning-based models, while requiring fewer parameters. PnP approaches rely on splitting algorithms like forward-backward or ADMM [19, 20, 21] to separate the action of the data fidelity and prior terms. This separation enables the use of learned CNNs as priors without needing pair of low- and high-resolution. In remote sensing, super-resolution techniques meet a great interest [22], however, model-based neural networks like PnP have been poorly explored. To the best of our knowledge, the only preliminary study was conducted using the U.S. Geological Survey National Map Urban Area imagery collection data [23], focusing on RGB images. In the present contribution, we extend this framework to multispectral image enhancement and deeply explore the performance on real data.

Outline - Section 2 describes the specificities of the considered satellite images, the forward model that establishes the relation between Landsat and Sentinel images, and the adopted PnP strategy with a specific focus on the denoiser. In Section 3, the description of the two considered datasets (synthetic and real) is provided as well as the calibration we made for the forward model. In Section 4, we present a first set of experiments, establishing the behaviour of the proposed PnP w.r.t hyperparameters such as the step-size involved in FB-PnP but also the regularization parameter. We then report the reconstruction results obtained on the two different datasets, along with the comparison of Spec-FB-PnP with both a classical interpolation method and the advanced learning-based approach SwinIR. Finally, in Section 5, we conclude by highlighting the advantages of Spec-FB-PnP and its futur use to uncover the history of surface water.

2 Proposed method: Spec-FB-PnP

Data - In this study two sources of data are used. Sentinel-2 images, part of the Copernicus project which started in 2015 [3], and Landsat project which began in 1972 [4]. Sentinel-2 images have a spatial resolution of up to 10m, while Landsat-8/9 images have a resolution of 30m. Both satellites cover a wide range of spectral bands (11 for Landsat-8/9 and 12 for Sentinel-2), however in this study we only focus on green, blue, red, and near-infrared bands for their ability to highlight elements of interest in land use and the proximity of their central wavelengths on both data sources. If the reconstruction quality will be of importance in this study, another quantity will also be considered: the normalized difference water index (NDWI), which highlights water in the images. This spectral index allows extracting hydromorphological metrics, such as the water channel perimeter or the surface area of open water. This index, is calculated from the “green” and “NIR” spectral bands. In the case of Sentinel-2 data (denoted xBN\mathrm{x}\in\mathbb{R}^{BN} where BB denotes the number of bands and NN the number of pixels), it can be written as follows, for every pixel nn,

dnNDWI=xgreen,nxNIR,nxgreen,n+xNIR,n.\mathrm{d}_{n}^{\mathrm{NDWI}}=\frac{\mathrm{x}_{\mathrm{green},n}-\mathrm{x}_{\mathrm{NIR},n}}{\mathrm{x}_{\mathrm{green},n}+\mathrm{x}_{\mathrm{NIR},n}}. (1)

Forward model - Following [9], we formulate the forward model as, for every band bb,

zb=(ϕx¯b)s+ε\mathrm{z}_{b}=(\phi*\bar{\mathrm{x}}_{b})\downarrow_{s}+\varepsilon (2)

where x¯=(x¯b)1bBBN\bar{\mathrm{x}}=(\bar{\mathrm{x}}_{b})_{1\leq b\leq B}\in\mathbb{R}^{BN} denotes the original high resolution (HR) image (i.e., Sentinel-2 image), ϕ\phi denotes the blurring kernel, s\downarrow_{s} the downsampling operator with a decimation factor s=3s=3, ε\varepsilon the stochastic degradation associated with Gaussian noise, and zBM\mathrm{z}\in\mathbb{R}^{BM} the low-resolution (LR) observation (i.e., Landsat-8/9).

Our goal is to recover a HR estimate x^\widehat{\mathrm{x}} close from x¯\bar{\mathrm{x}} considering the LR image z\mathrm{z}. The long-term goal being to densify the HR observations on the time period where both Landsat and Sentinel-2 images are available (i.e., 2015-now) but also obtain HR images on the period 1972-2015 where only LR Landsat-8/9 images are available.

In the following, the forward model will be simply formulated as z=Ax¯+ε\mathrm{z}=\mathrm{A}\bar{\mathrm{x}}+\varepsilon where A\mathrm{A} encodes both the convolution with ϕ\phi and the downsampling.

Plug-and-play - The literature of model-based neural network including unfolded neural network and Plug-and-play methods have significantly increased in the past ten years and is now established as a good compromise between standard neural networks and variational approaches. Specifically, we investigate the well-established Forward-Backward PnP (FB-PnP) framework [24, 25]. While FB-PnP has been widely studied for RGB image processing tasks, its use in multispectral imagery remains limited, mainly because it requires the development of a dedicated denoiser capable of handling multispectral data.

The iterations of FB-PnP read:

x[k+1]=𝐃θ,τλ(x[k]τA(Ax[k]z)){\mathrm{x}}^{[k+1]}=\mathbf{D}_{\theta,\tau\lambda}(\mathrm{x}^{[k]}-\tau\mathrm{A}^{\top}(\mathrm{A}\mathrm{x}^{[k]}-\mathrm{z})) (3)

where τ>0\tau>0 is the step-size, λ\lambda a regularization parameter, θ\theta the neural networks parameters, and 𝐃θ,τλ\mathbf{D}_{\theta,\tau\lambda} the denoiser. When 𝐃θ,τλ=proxτλg\mathbf{D}_{\theta,\tau\lambda}=\mathrm{prox}_{\tau\lambda g} we recover standard FB schemes and convergence to x^Argminx12Axz22+λg(x)\widehat{\mathrm{x}}\in\mathrm{Argmin}_{\mathrm{x}}\frac{1}{2}\|\mathrm{A}\mathrm{x}-\mathrm{z}\|_{2}^{2}+\lambda g(\mathrm{x}) is guaranteed for specific choice of τ\tau [26]. Under technical assumptions, it has been established that the sequence (x[k])k=0,1,({\mathrm{x}}^{[k]})_{k=0,1,\ldots} converge to x^\widehat{\mathrm{x}} such that 0A(Ax^z)+𝐓(x^)0\in\mathrm{A}^{\top}(\mathrm{A}\widehat{\mathrm{x}}-\mathrm{z})+\mathbf{T}(\widehat{\mathrm{x}}), where 𝐓\mathbf{T} is related to 𝐃θ,τλ\mathbf{D}_{\theta,\tau\lambda}. The involved denoiser should insure specific properties that can be imposed during the training, often at a price of reduced performance [24, 25].

Building the denoiser - For this work, we choose for the denoiser 𝐃θ,τλ\mathbf{D}_{\theta,\tau\lambda} a DRUNet [27], which is a combination of the convolutional network U-Net [28], and the deep residual learning network ResNet [29]. Such a denoiser is available in DeepInverse library [30] but for RGB image denoising. To design 𝐃θ,τλ\mathbf{D}_{\theta,\tau\lambda} for multispectral image denoising, it was first necessary to change the first and last layers to the source code in order to enable the network to train on 4-channel multispectral data and configure its learning accordingly.

To learn the hyperparameters θ\theta of the proposed denoiser 𝐃θ,τλ\mathbf{D}_{\theta,\tau\lambda}, we created a dataset 𝒟={(x¯(),y())={1,,LD}}\mathcal{D}=\{(\overline{\mathrm{x}}^{(\ell)},\mathrm{y}^{(\ell)})_{\ell=\{1,\ldots,L_{D}\}}\} such that y()=x¯()+ϵ\mathrm{y}^{(\ell)}=\overline{\mathrm{x}}^{(\ell)}+\epsilon, where x¯()\overline{\mathrm{x}}^{(\ell)} denotes the HR Sentinel-2 images and ϵ\epsilon a white Gaussian noise with standard deviation σ\sigma. These x¯()\overline{\mathrm{x}}^{(\ell)} images were obtained using N=300×300N=300\times 300 pixels crops of Sentinel-2 images centered on the river corridor. The studied site selected for this work is the Lhasa River in Tibet. This area is of particular interest as it has experienced rapid urbanization in recent years, significantly impacting the development of the alluvial plain and continuing to shape it over time. Sentinel-2 images are acquired over the same area of the Earth’s surface every 5 days, since data began to be accessible online in 2016. The resulting dataset is composed of LD=1200L_{D}=1200 images. Few examples from 𝒟\mathcal{D} are displayed in Figure 1.

Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Figure 1: DRUNet training dataset: Original images x¯()\overline{\mathrm{x}}^{(\ell)} and noisy versions y()\mathrm{y}^{(\ell)} with a Gaussian noise with σ=0.09\sigma=0.09.

We consider AdamW optimizer with a learning rate of 0.001. A StepLR scheduler with a step size of 100 and a decay factor of 0.5. The network was trained for 1500 epochs. The learning rate and decay factor were chosen based on tests conducted in the Deep Inverse project [30] using three-channel images. To evaluate the impact of the batch size used during the training, two different batch-size have been employed: 4 and 17. In this article, we will only present the reconstruction obtained with a batch-size of 4, given the best results obtained.

Spec-FB-PnP refers to (3) when 𝐃θ,τλ\mathbf{D}_{\theta,\tau\lambda} is the denoiser described above.

3 Datasets and forward model calibration

To evaluate the performance of the Spec-FB-PnP algorithm for increasing the resolution of Landsat-8/9 images to Sentinel-2 images, we consider two types of datasets: a synthetic dataset and a real one. On the one hand, the synthetic dataset is created to mimic as close as possible the real one and to evaluate the performance of Spec-FB-PnP on a controlled dataset, particularly to verify whether any artifacts are introduced. On the other hand, evaluating Spec-FB-PnP on the real dataset allows us to check its stability on a non-ideal situation.
Real dataset and calibration of the forward model - This dataset consists of pairs of Sentinel-2 (x¯\overline{\mathrm{x}}_{\ell}) and Landsat (zreal\mathrm{z}_{\ell}^{\textrm{real}}) images, such that ={(x¯,zreal)={1,,LR}}\mathcal{R}=\{(\overline{\mathrm{x}}_{\ell},\mathrm{z}_{\ell}^{\textrm{real}})_{\ell=\{1,\ldots,L_{R}\}}\}. The underlying forward model is of the form (2) where ϕ\phi, ss, and ε\varepsilon are estimated from the data. It is also important to notice that because these data come from different sources, it was not possible to obtain acquisitions that were spatially aligned on the same location and at the same time. To overcome these constraints, after downloading the images from the two sources, we classified them chronologically in order to create {(x¯,zreal)}\{(\overline{\mathrm{x}}_{\ell},\mathrm{z}_{\ell}^{\textrm{real}})\} image pairs that were as close as possible in their acquisition time (maximum gap of three days). For spatial alignment, it was possible to use Landsat data from the Harmonized Landsat Sentinel-2 (HLS) project [5], which reprojected Landsat images onto the Sentinel-2 UTM tiling grid.

The x¯\overline{\mathrm{x}}_{\ell} images were obtained by cropping Sentinel-2 images of size 11000×1100011000\times 11000 pixels using a N=300×300N=300\times 300 pixels window, focusing the extraction on the river corridor using the associated centerline obtained from the Global Surface Water (GSW) dataset [31]. The same procedure was used on Landsat images of size 3660×36603660\times 3660 pixels using a M=100×100M=100\times 100 pixels window. This difference in window size is due to the scale factor s=3s=3. The blurring kernel ϕ\phi with a standard deviation σ=0.09\sigma=0.09 and a size of 4×44\times 4 pixels, has been calibrated by taking into account concepts from the literature for the choice of kernel width [32] and from (x¯,zreal)(\overline{\mathrm{x}}_{\ell},\mathrm{z}_{\ell}^{\textrm{real}}) in order minimize the error (ϕx¯zreal)s22\|(\phi*\overline{\mathrm{x}}_{\ell}-\mathrm{z}_{\ell}^{\textrm{real}})\downarrow_{s}\|_{2}^{2}. The standard deviation of white Gaussian noise, was estimated using the median absolute deviation technique applied on wavelet details coefficients of the Landsat image: σmedian(|Wzreal|)0.6745\sigma\approx\frac{\text{median}\left(\left|\mathrm{W}\mathrm{z}_{\ell}^{\textrm{real}}\right|\right)}{0.6745} [33].

Synthetic dataset - This second dataset has been defined to validate the Spec-FB-PnP on data for which ground truth is available. The dataset is 𝒮={(x¯,zsynt)={1,,LS}}\mathcal{S}=\{(\overline{\mathrm{x}}_{\ell},\mathrm{z}_{\ell}^{\textrm{synt}})_{\ell=\{1,\ldots,L_{S}\}}\} with zsynt=(ϕx¯)s+ε\mathrm{z}_{\ell}^{\textrm{synt}}=(\phi*\overline{\mathrm{x}}_{\ell})\downarrow_{s}+\varepsilon where ϕ\phi and ε\varepsilon are calibrated from the real Landsat images zreal\mathrm{z}_{\ell}^{\textrm{real}} as described in the paragraph above.

An illustration of images extracted from these two datasets and associated NDWI images is provided in Figure 2.

Refer to caption Refer to caption Refer to caption
HR Sentinel-2 x¯\overline{\mathrm{x}}_{\ell} LR Synthetic zsynt\mathrm{z}_{\ell}^{\textrm{synt}} LR Landsat-8/9 zreal\mathrm{z}_{\ell}^{\textrm{real}}
Refer to caption Refer to caption Refer to caption
NDWI NDWI NDWI
Refer to caption Refer to caption Refer to caption
Threshold NDWI Threshold NDWI Threshold NDWI
Figure 2: Row 1: Sentinel-2, synthetic Landsat, and real Landsat images. Row 2: associated NDWI index images. Row 3: thresholded NDWI index allowing water surface extraction.

4 Numerical experiment

Evaluation criteria - The PSNR and SSIM are used as indicators of reconstruction quality. We also compute the surface area of open water by thresholding the NDWI index (1) (cf. Figure 2, third line).

Comparisons with state-of-the-art methods - The reconstructions obtained with the Plug-and-Play method are compared with those from the bicubic interpolation implemented in PyTorch, as well as the neural network-based method named SwinIR [18]. This second methodology was trained on the same dataset as the denoising DRUNet, except that the network takes also into account the change of resolution from Landsat-8/9 to Sentinel-2. The number of epochs for the training is the same than for our denoiser (i.e. 1500).
Impact of the step-size and regularization parameter in PnP - In Figure 3(left), we display the PSNR obtained with the proposed PnP method for different choices of step-size τ\tau and regularization parameter λ\lambda. For the optimal choice of regularization parameter (i.e. 0.08) we display in Figure 3(right) the evolution of the PSNR w.r.t iterations in order to observe the convergence of the algorithm for specific step-size choices. We clearly see the configuration with the step-size τ=3\tau=3 leads to the best reconstruction results and leads to a converging sequence in terms of PSNR. In the following experiments the choice of τ=3\tau=3 and λ=0.08\lambda=0.08 is made.

Refer to caption
Figure 3: Evaluation of the choice of regularization parameter λ\lambda and step size τ\tau in Spec-FB-PnP. Left: PSNR as a function of step size with τ\tau a fixed regularization term λ\lambda. Right: Evolution of PSNR w.r.t iterations for λ=0.008\lambda=0.008 and different step-sizes.
Method SSIM PSNR (dB) Water Area (m2)
Image A – Sentinel Reference (22 580 m2)
Synthetic A 8 730
Bicubic interpolation 0.59 22.7 21 340
SwinIR 0.87 30.0 21 980
PnP 0.88 28.4 22 620
Image A – Landsat Reference (4 740 m2)
Bicubic interpolation 0.51 19.1 15 730
SwinIR 0.65 20.3 16 600
PnP 0.60 19.5 15 510
Image B – Sentinel Reference (26 620 m2)
Synthetic B 7 290
Bicubic interpolation 0.62 24.2 26 310
SwinIR 0.85 30.8 24 880
PnP 0.82 27.6 23 670
Image B – Landsat Reference (7 110 m2)
Bicubic interpolation 0.53 14.73 21 920
SwinIR 0.60 14.94 19 010
PnP 0.64 15.03 20 530
Table 1: Quantitative evaluation of reconstruction performance for Images A and B on synthetic and real data. Best values per section are shown in bold.

Comments on reconstruction performance - Figure 4 presents two sections of the Lhasa River (referred as A and B) as well as a cropped version on each, shown in the first row. These images are particularly well suited for evaluating the reconstruction results because, although the maximum temporal gap between the acquisitions is about three days, no significant changes occurred in that interval on the real data.

On the second row, we display the associated synthetic images, obtained by degrading Sentinel-2 data to match the spatial resolution and quality of Landsat-8/9 images (denoted Synthetic A/B), and the corresponding real Landsat-8/9 images. The reconstruction results are then displayed on row 3 to 5. Table 1 presents the indicators extracted from the reconstructions.

A first visual assessment of the reconstructions reveals that the results obtained with methods based on neural networks (SwinIR and PnP) outperform the classic bicubic interpolation method. This is true for reconstructions based on synthetic images as well as Landsat images. As expected, it can also be noted that the results obtained via the reconstruction of synthetic images exceed the scores obtained from real Landsat images in all metrics (see Tabular 1). Looking more closely at the reconstructions obtained with SwinIR and PnP, we can see that most of the river branches have been correctly reconstructed, despite visible smoothing in areas outside the watercourse, which have lost detail, particularly in the reconstruction of Landsat A and B by SwinIR. In terms of water area estimation, we can see that all methods allows to significantly improve the estimation of water area, highlighting the benefit of performing reconstruction of Landsat-8/9 images.

The objective is not only to obtain the most accurate overall measurement of water surface area, but rather to achieve a good compromise between image reconstruction of the spatial structures where water is present. Such structural information is crucial, as it enables the assessment of how the watercourse evolves over time. The SSIM score measures the structural similarity between a reference image and its reconstruction, and thus accurately reflects the point we have just made. Table 1 shows that the PnP method achieves the highest SSIM score for the reconstructions of Synthetic A and Landsat B. In the other two cases (Synthetic B and Landsat A), SwinIR obtains the best results, with a maximum difference of only 0.05 points with Spec-FB-PnP, which is negligible.

It is also important to notice that Spec-FB-PnP is twice faster than SwinIR.

Sentinel A (Zoom) Sentinel B (Zoom)
Refer to caption Refer to caption Refer to caption Refer to caption
Observation z\mathrm{z}
Synthetic Landsat Synthetic B Landsat B
Refer to caption Refer to caption Refer to caption Refer to caption
Reconstruction Bicubic
Refer to caption Refer to caption Refer to caption Refer to caption
Reconstruction Swinir
Refer to caption Refer to caption Refer to caption Refer to caption
Reconstruction PnP
Refer to caption Refer to caption Refer to caption Refer to caption
Figure 4: Reconstruction of two sections of the Lhasa River (Tibet) using bicubic, SwinIR, and PnP methods, based on both real Landsat images and synthetic Landsat images (degraded Sentinel-2).

5 Conclusion

In this study, we propose a new Spec-FB-PnP method to retrieve the information lost on Landsat 8/9 compared to Sentinel-2 images. The method is evaluated against a neural network solution SwinIR and a the classic bicubic interpolation. The performance has been evaluated in terms of reconstruction quality and hydromorphologically quantities on the Lhasa River in Tibet. The SSIM score shows that, on average, the Spec-FB-PnP method offers a good compromise to achieve a reconstruction of the structures with a good estimation of water area estimation. These results open up prospects for the use of the Spec-FB-PnP method more systematically to incrase the resolution of Landsat image from 1972 to the present day and on other watercourses around the world where training data is very limited due to cloud cover (e.g., the Amazon River).

6 Acknowledgements

This work is funded as part of ANR GloUrb project (ANR-22-CE03-0005). The project is also co-financed by the Labex IMU (ANR-10-LABEX-0088) and the EUR H2O’Lyon (ANR-17-EURE-0018) of the University of Lyon, as part of the "Investissements d’Avenir" programme operated by the ANR.

References

  • [1] H. M. Habersack, “The river-scaling concept (RSC): a basis for ecological assessments,” Hydrobiologia, vol. 422, pp. 49–60, 2000.
  • [2] P. J. Crutzen, “The Anthropocene,” J. de Physique IV, 2002.
  • [3] F. Gascon, E. Cadau, O. Colin, B. Hoersch, C. Isola, B. L. Fernandez, and P. Martimort, “Copernicus Sentinel-2 mission: products, algorithms and cal/val,” in Earth observing systems XIX. SPIE, 2014, vol. 9218, pp. 455–463.
  • [4] M. A. Wulder, T. R. Loveland, D. P. Roy, C. J. Crawford, J. G. Masek, C. E. Woodcock, R. G. Allen, M. C. Anderson, A. S. Belward, and W. B. Cohen, “Current status of landsat program, science, and applications,” Remote Sensing of Environment, vol. 225, pp. 127–147, 2019.
  • [5] M. Claverie, J. Ju, J. G. Masek, J. L. Dungan, E. F. Vermote, J.-C. Roger, S. V. Skakun, and C. Justice, “The harmonized Landsat and Sentinel-2 surface reflectance data set,” Remote Sensing of Environment, vol. 219, pp. 145–161, 2018.
  • [6] S. D. Babacan, R. Molina, and A. K. Katsaggelos, “Variational bayesian super resolution,” IEEE Trans. Image Process., vol. 20, no. 4, pp. 984–999, 2010.
  • [7] H. Ji and C. Fermüller, “Robust wavelet-based super-resolution reconstruction: theory and algorithm,” IEEE Trans. Pattern Anal. Match. Int., vol. 31, no. 4, pp. 649–660, 2008.
  • [8] M. Nazzal and H. Ozkaramanli, “Wavelet domain dictionary learning-based single image superresolution,” Signal, Image and Video Processing, vol. 9, no. 7, pp. 1491–1501, 2015.
  • [9] Z. Wang, J. Chen, and S. C. H. Hoi, “Deep learning for image super-resolution: A survey,” IEEE Trans. Pattern Anal. Match. Int., vol. 43, no. 10, pp. 3365–3387, 2020.
  • [10] C. Dong, C. C. Loy, K. He, and X. Tang, “Learning a deep convolutional network for image super-resolution,” in European Conference on Computer Vision. Springer, 2014, pp. 184–199.
  • [11] C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, and Z. Wang, “Photo-realistic single image super-resolution using a generative adversarial network,” in IEEE Conf. Comput. Vis. Pattern Recognit., 2017, pp. 4681–4690.
  • [12] J. Wang, Y. Wu, L. Wang, L. Wang, O. Alfarraj, and A. Tolba, “Lightweight feedback convolution neural network for remote sensing images super-resolution,” IEEE Access, vol. 9, pp. 15992–16003, 2021.
  • [13] K. Prajapati, V. Chudasama, H. Patel, K. Upla, R. Ramachandra, K. Raja, and C. Busch, “Unsupervised single image super-resolution network (USISResNet) for real-world data using generative adversarial network,” in IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2020, pp. 464–465.
  • [14] C. Tuna, G. Unal, and E. Sertel, “Single-frame super resolution of remote-sensing images by convolutional neural networks,” International Journal of Remote Sensing, vol. 39, no. 8, pp. 2463–2479, 2018.
  • [15] J. Gu, H. Lu, W. Zuo, and C. Dong, “Blind super-resolution with iterative kernel correction,” in IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2019, pp. 1604–1613.
  • [16] M. Daniels, P. Hand, and R. Heckel, “Reducing the representation error of gan image priors using the deep decoder,” arXiv preprint arXiv:2001.08747, 2020.
  • [17] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Adv. Neural Inf. Process. Syst., vol. 30, 2017.
  • [18] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte, “SwinIR: Image restoration using swin transformer,” in IEEE/CVF International Conference on Computer Vision, 2021, pp. 1833–1844.
  • [19] S. V. Venkatakrishnan, C. A. Bouman, and B. Wohlberg, “Plug-and-play priors for model-based reconstruction,” in IEEE Glob. Conf. Signal Inf. Process. Proc., Austin, TX, USA, 2013, pp. 945–948.
  • [20] C. Y. Park, Y. Hu, M. T. McCann, C. Garcia-Cardona, B. Wohlberg, and U. S. Kamilov, “Plug-and-play priors as a score-based method,” in 2025 IEEE International Conference on Image Processing (ICIP). IEEE, 2025, pp. 49–54.
  • [21] S. Hurault, A. Chambolle, A. Leclaire, and N. Papadakis, “Convergent plug-and-play with proximal denoiser and unconstrained regularization parameter,” J. Math. Imaging Vis., vol. 66, no. 4, pp. 616–638, 2024.
  • [22] P. Wang, B. Bayram, and E. Sertel, “A comprehensive review on deep learning based remote sensing image super-resolution methods,” Earth-Science Reviews, vol. 232, pp. 104110, 2022.
  • [23] H. Tao, “Super-resolution of remote sensing images based on a deep plug-and-play framework,” in IEEE International Geoscience and Remote Sensing Symposium, 2020, pp. 625–628.
  • [24] S. Hurault, A. Leclaire, and N. Papadakis, “Proximal denoiser for convergent plug-and-play optimization with nonconvex regularization,” in International Conference on Machine Learning, 2022, pp. 9483–9505.
  • [25] J.-C. Pesquet, A. Repetti, M. Terris, and Y. Wiaux, “Learning maximally monotone operators for image recovery,” SIAM J. Imaging Sci., vol. 14, no. 3, pp. 1206–1237, 2021.
  • [26] P. L. Combettes and V. R. Wajs, “Signal recovery by proximal forward-backward splitting,” Multiscale Model. Simul., vol. 4, no. 4, pp. 1168–1200, 2005.
  • [27] K. Zhang, Y. Li, W. Zuo, L. Zhang, L. Van Gool, and R. Timofte, “Plug-and-play image restoration with deep denoiser prior,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 10, pp. 6360–6376, 2021.
  • [28] O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention. Springer, 2015, pp. 234–241.
  • [29] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778.
  • [30] J. Tachella, M. Terris, S. Hurault, A. Wang, D. Chen, and et al., “DeepInverse: A python package for solving imaging inverse problems with deep learning,” 2025.
  • [31] J.-F. Pekel, A. Cottam, N. Gorelick, and A. S. Belward, “High-resolution mapping of global surface water and its long-term changes,” Nature, vol. 540, no. 7633, pp. 418–422, 2016.
  • [32] N. Efrat, D. Glasner, A. Apartsin, B. Nadler, and A. Levin, “Accurate blur models vs. image priors in single image super-resolution,” in IEEE International Conference on Computer Vision, 2013, pp. 2832–2839.
  • [33] S. Mallat, A wavelet tour of signal processing, Elsevier, 1999.