PhaSR: Generalized Image Shadow Removal with Physically Aligned Priors

Draft model card assembled from the public CVPR 2026 paper, project page, code repository, and public model-zoo files. It is not an author-endorsed release document. Verify the checkpoint license and final benchmark protocol with the model authors before redistributing weights.

Model summary

PhaSR (Physically Aligned Shadow Removal) is an image-to-image restoration model for single-image shadow removal and ambient-lighting normalization. It is designed to separate illumination effects from intrinsic scene reflectance across direct shadows, indirect lighting, and multi-source ambient illumination.

The method combines two forms of prior alignment:

  1. Physically Aligned Normalization (PAN): parameter-free input normalization using Gray-world color correction, log-domain Retinex decomposition, and dynamic-range recombination.
  2. Geometric-Semantic Rectification Attention (GSRA): cross-modal differential attention that aligns depth and surface-normal priors derived from Depth Anything V2 with frozen DINOv2 semantic features.

The learned backbone is a multi-scale Transformer encoder-decoder. The public repository reports 18.95M parameters and 55.63G FLOPs under its evaluation configuration.

Model details

Field Value
Task Single-image shadow removal; ambient-lighting normalization
Framework PyTorch
Publication CVPR 2026
Authors Chia-Ming Lee, Yu-Fan Lin, Yu-Jou Hsiao, Jin-Hui Jiang, Yu-Lun Liu, Chih-Chung Hsu
Paper arXiv:2601.17470
Project page PhaSR project page
Code ming053l/PhaSR
Public weights Google Drive model zoo
Code license MIT, according to the public repository
Weight license No separate checkpoint license was identified; confirm with the authors before redistribution or commercial use

Inputs and outputs

Inputs

  • An RGB image containing shadows or spatially varying illumination.
  • Precomputed depth and surface-normal arrays (.npy) generated with Depth Anything V2, following the repository preprocessing procedure.
  • The repository also depends on DINOv2 for semantic priors.

PhaSR is mask-free at inference time: a manually annotated shadow mask is not listed as a required input.

Output

An RGB image in which the model attempts to remove shadows or normalize ambient illumination while preserving surface color, texture, and scene structure.

Available checkpoints

The public model zoo contains dataset-specific checkpoints. A checkpoint name identifies its training or release target; it does not by itself establish cross-dataset validity.

Variant Checkpoint Download
Ambient6K Ambient6K_model_best.pth Hugging Face
INS INS_model_best.pth Hugging Face
ISTD ISTD_model_best.pth Hugging Face
ISTD+ ISTDp_model_best.pth Hugging Face
WSRD WSRD_model_best.pth Hugging Face

SHA-256 values are published in CHECKSUMS.sha256. These files were copied from the public Google Drive model zoo on 2026-09-28. No author-published checksums were found for an independent comparison.

The repository changelog states that an earlier checkpoint issue was fixed and new checkpoints were released on 2026-07-03. Record file hashes when publishing a downstream deployment so that the evaluated checkpoint remains identifiable.

Intended use

PhaSR is intended for research and development involving:

  • restoration of photographs affected by cast shadows;
  • normalization of indoor or outdoor illumination;
  • preprocessing for downstream vision systems when shadows are a nuisance factor;
  • research on physics-informed and prior-guided image restoration.

Out-of-scope use

The model is not validated as:

  • a forensic tool for proving that an image is authentic or unedited;
  • a physically calibrated estimator of illumination, albedo, depth, or surface normals;
  • a safety-critical preprocessing component for medical, autonomous-driving, legal, or evidentiary decisions;
  • a general-purpose relighting model for arbitrary artistic edits.

Training and evaluation data

The public repository documents experiments involving ISTD, ISTD+, WSRD+, SRD, INS, and Ambient6K. Dataset licenses and permitted uses are governed by the respective dataset owners and are not replaced by the repository's code license.

Before deployment, evaluate the selected checkpoint on data that matches the intended cameras, scenes, materials, resolutions, and illumination conditions. Do not infer performance on a new domain solely from in-domain benchmark results.

Reported evaluation

The repository reports the following comparison. Values are reproduced from the public project table; consult the paper and evaluation code for preprocessing, crop size, and metric details before reproducing or comparing them.

Model Parameters FLOPs ISTD+ WSRD+ Ambient6K
OmniSR 24.55M 78.32G 33.34 26.07 23.01
DenseSR 24.70M 81.13G 33.98 26.28 22.54
PhaSR 18.95M 55.63G 34.48 28.44 23.32

The paper additionally reports cross-dataset evaluation:

Train -> test PSNR SSIM
Ambient6K -> ISTD 27.64 0.923
ISTD -> Ambient6K 21.15 0.798

These are author-reported values, not independently reproduced results in this model card.

Basic usage

Follow the public repository rather than loading a checkpoint into an unrelated architecture.

git clone https://github.com/ming053l/PhaSR.git
cd PhaSR
conda create --name phasr python=3.9 -y
conda activate phasr
pip install torch==2.0.1 torchvision==0.15.2 torchaudio==2.0.2 --index-url https://download.pytorch.org/whl/cu118
pip install -r requirements.txt

The documented workflow then requires:

  1. installing/cloning Depth Anything V2;
  2. running calculate_depth_normal.py to generate depth and normal arrays;
  3. installing/cloning DINOv2;
  4. placing the selected .pth checkpoint at the path expected by the test configuration;
  5. updating dataset and checkpoint paths in test.py or the relevant configuration;
  6. running bash test.sh.

Exact commands may change with repository revisions. Pin the code commit, dependency versions, and checkpoint hash for reproducible evaluation.

Limitations and failure modes

  • The paper reports difficult cases involving shadows over intrinsically dark materials, where illumination and reflectance are ambiguous.
  • Specular or metallic surfaces may violate assumptions used by Retinex-style decomposition and may produce color or texture artifacts.
  • Errors in Depth Anything V2 depth or normal estimates can propagate into geometric priors, especially around transparent, reflective, thin, or unusual objects.
  • DINOv2 semantic priors may be spatially coarse or unreliable on domains far from its pretraining distribution.
  • A dataset-specific checkpoint may overfit the lighting, camera, resolution, or scene statistics of its training set.
  • Shadow removal changes image evidence. Retain the original image and disclose processing when provenance, documentation, or scientific reproducibility matters.

Risks and recommendations

  • Visually plausible output does not guarantee physically correct reflectance or lighting.
  • Restoration may erase real dark regions or hallucinate texture near shadow boundaries.
  • Do not use output alone to make claims about object color, scene geometry, or image authenticity.
  • For downstream model preprocessing, compare task performance both with and without PhaSR and audit failures by material, lighting type, skin tone, scene category, and camera pipeline where relevant.
  • Preserve input/output pairs, preprocessing parameters, code revision, and checkpoint hash in production logs.

Reproducibility checklist

  • Record the exact PhaSR repository commit.
  • Record the checkpoint filename and SHA-256 hash.
  • Record the Depth Anything V2 and DINOv2 versions/checkpoints.
  • Record image resizing, cropping, color-space, and normalization settings.
  • Confirm dataset and checkpoint licensing for the intended use.
  • Reproduce PSNR/SSIM on at least one documented dataset.
  • Evaluate on an application-specific holdout set.
  • Inspect failure cases on dark, reflective, metallic, transparent, and textured surfaces.

Citation

@inproceedings{lee2026phasr,
  title     = {PhaSR: Generalized Image Shadow Removal with Physically Aligned Priors},
  author    = {Lee, Chia-Ming and Lin, Yu-Fan and Hsiao, Yu-Jou and Jiang, Jin-Hui and Liu, Yu-Lun and Hsu, Chih-Chung},
  booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year      = {2026}
}

Contact

For authoritative questions about the method, weights, and licensing, use the contact information provided in the official repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for ming0531/PhaSR