SRNet-x4-Upscaler — small, fast single-image super-resolution
A compact residual CNN (~1.25 M parameters) that upscales images 4x. It predicts a residual on top of a bicubic upscale, so it starts from — and improves on — plain bicubic interpolation. Trained in 30 minutes on 2x T4 GPUs.
Live demo (runs in your browser, nothing is uploaded): Infinitode Experiments
Results
Evaluated on the 100 DIV2K validation images, full images, bicubic-x4 degradation, Y-channel PSNR with a 4 px border shave, 8-bit outputs:
| Method | PSNR (dB) |
|---|---|
| Bicubic | 28.23 |
| SRNet-x4-Upscaler | 30.28 (+2.05) |
Metrics come from the training notebook (antialiased bicubic downscale, not the official MATLAB imresize), so they are comparable between methods here but only approximately comparable to published tables. Published tables usually report RGB PSNR, which reads about 1.5 dB lower than Y-channel PSNR for the same images.
Usage
import numpy as np, torch
from PIL import Image
from huggingface_hub import hf_hub_download
import importlib.util
REPO = "InfinitodeLTD/srnet-x4-upscaler"
spec = importlib.util.spec_from_file_location("srnet", hf_hub_download(REPO, "srnet.py"))
srnet = importlib.util.module_from_spec(spec); spec.loader.exec_module(srnet)
model = srnet.SRNet.from_pretrained(REPO).eval()
img = np.asarray(Image.open("photo.png").convert("RGB"), np.float32) / 255
x = torch.from_numpy(img).permute(2, 0, 1)[None]
with torch.inference_mode():
y = model(x).clamp(0, 1)[0].permute(1, 2, 0).numpy()
Image.fromarray((y * 255).round().astype("uint8")).save("photo_x4.png")
Large images (tiled inference)
Process tiles with a margin of at least 2*nb + 6 = 38 input pixels; the result is then identical to whole-image inference.
@torch.inference_mode()
def upscale_tiled(model, x, tile=128):
s, m = model.scale, 2 * model.nb + 6
_, _, H, W = x.shape
out = torch.empty(1, 3, H * s, W * s)
for y in range(0, H, tile):
for xx in range(0, W, tile):
y1, x1 = min(H, y + tile), min(W, xx + tile)
ya, xa, yb, xb = max(0, y - m), max(0, xx - m), min(H, y1 + m), min(W, x1 + m)
sr = model(x[:, :, ya:yb, xa:xb]).clamp(0, 1)
out[:, :, y * s:y1 * s, xx * s:x1 * s] = sr[:, :, (y - ya) * s:(y1 - ya) * s, (xx - xa) * s:(x1 - xa) * s]
return out
ONNX
srnet_x4.onnx takes input (float32 NCHW RGB in [0, 1], dynamic batch/height/width) and returns output (clip to [0, 1]). It runs with onnxruntime and onnxruntime-web (the browser demo uses it).
Model details
- Architecture: 3x3 conv head -> 16 residual blocks (conv-ReLU-conv, 64 channels) -> 3x3 conv with long skip -> conv to 48 channels + PixelShuffle(4) -> added to a bicubic upscale of the input. The final conv is zero-initialised, so training starts at exact bicubic.
- Compute: all convolutions run at input resolution; about 2.5 MFLOPs per input pixel.
Training
- Data: DF2K (DIV2K + Flickr2K) HR training images; DIV2K validation images held out.
- Degradation: antialiased bicubic x4 downscale, quantised to 8 bit. No noise, blur or compression.
- 16,157 steps, batch 32, 64x64 LR patches (256x256 HR), random flips/rotations.
- AdamW, peak LR 4e-4 with cosine decay, L1 loss, then MSE for the final 8% of the schedule; EMA (0.999) weights evaluated alongside raw weights and the better one kept; fp16 mixed precision.
Limitations
- Trained only on clean bicubic-downscaled images. On real-world inputs with noise, blur or JPEG artefacts it will not denoise or deblock and may amplify artefacts.
- Optimised for PSNR with L1/MSE losses, so fine texture (fur, feathers, foliage) comes out smoother than the original. It does not hallucinate detail the way GAN-based upscalers do.
- x4 only.
License
The code and weights in this repository are released under the MIT License.
The MIT license does not change the terms of the training data: DIV2K is distributed for academic research use, and Flickr2K images carry their original Flickr licenses. If you plan commercial use, confirm those terms are acceptable to you.
Acknowledgements
- DIV2K: Agustsson & Timofte, NTIRE 2017 Challenge on Single Image Super-Resolution: Dataset and Study, CVPRW 2017.
- Flickr2K: collected for Lim et al., Enhanced Deep Residual Networks for Single Image Super-Resolution (EDSR), CVPRW 2017. The residual-block design follows the EDSR family.
- Downloads last month
- 6
Dataset used to train InfinitodeLTD/SRNet-x4-Upscaler
Evaluation results
- PSNR (Y channel, shave 4) on DIV2K validation (100 images, bicubic degradation)self-reported30.280