UltraPIPS (TUSA backbone) for zea

UltraPIPS is a perceptual similarity metric for B-mode ultrasound images, in the style of LPIPS. Instead of a network pretrained on natural images, its backbone is an ultrasound foundation model. This repository holds the default UltraPIPS backbone, the Swin transformer encoder of TUSA, converted to Keras 3 for use in zea.

This is a port, not original work. UltraPIPS and TUSA were created by Tal Grutman, Carmel Shinar and Tali Ilovitsh (Tel Aviv University). All credit for the method, the model and the trained weights goes to them. The zea maintainers only converted the weights and reimplemented the forward pass in Keras.

License

The original UltraPIPS and TUSA code and the pretrained TUSA weights are released by their authors under the GNU General Public License v3.0. The files in this repository are derived from those weights and are distributed under the same license; see LICENSE. This differs from zea itself, which is Apache-2.0: if you use this preset, the GPL-3.0 terms apply to these weights.

The encoder architecture follows MONAI's SwinUNETR (Apache-2.0).

Files

File Description
config.json, metadata.json zea / Keras model configuration
model.weights.h5 Keras weights of the TUSA swinViT encoder (9,120,756 parameters), converted from unet.pt
unet.pt Unmodified original TUSA checkpoint from talg2324/tusa (commit 5ff26ff), the source of the converted weights
LICENSE GPL-3.0 license text from the original repository

Usage

from zea.models.ultrapips import UltraPIPS

model = UltraPIPS.from_preset("ultrapips-tusa")
distance = model([image1, image2])  # (B, H, W, C) or (H, W, C), C in {1, 3}, values in [-1, 1]

Or as a zea metric, mirroring LPIPS:

from zea import metrics

ultrapips = metrics.get_ultrapips(image_range=[0, 1])
distance = ultrapips(image1, image2)

Lower values mean more similar images. Inputs are converted to grayscale (if RGB), resized to 128 × 128 with antialiased bilinear interpolation, and normalized as in the original implementation. To convert the original PyTorch checkpoint on the fly instead (requires PyTorch):

model = UltraPIPS()
model.custom_load_weights("ultrapips-tusa", backend="torch")  # reads unet.pt

Differences from the original interface

  • Inputs are channels-last and in [-1, 1], the same as zea's LPIPS. The original takes channels-first tensors in [0, 1].
  • Only the tusa_vit backbone has been ported. The other UltraPIPS backbones (USFM, Ultrasound-CLIP, BiomedCLIP, MedSAM, natural-image models) are not available in zea.

Method

For each of the five feature levels of the encoder (patch embedding plus four Swin stages), features of both images are unit-normalized along the channel axis. The metric takes their squared difference, averages it over space and sums it over channels, then sums all levels. Unlike LPIPS, there is no learned linear head.

Conversion check

On the same inputs, the Keras port matches the original PyTorch implementation (UltraPIPS("tusa_vit")) to float32 precision:

  • Feature maps: relative error ≤ 9e-7.
  • Distances on 108 real B-mode image/distortion pairs: max absolute difference 2.4e-7.
  • Verified on the JAX, TensorFlow and PyTorch Keras backends, on CPU and GPU.

Citation

If you use this model, please cite the original authors:

@inproceedings{Grutman2026UltraPIPS,
  title={UltraPIPS: Improving model perception in B-mode ultrasound with foundation models},
  author={Grutman, Tal and Ilovitsh, Tali},
  booktitle={MICCAI Workshop on Advances in Simplifying Medical Ultrasound (ASMUS)},
  year={2026},
  eprint={2608.26033},
  archivePrefix={arXiv}
}

@article{Grutman2025TUSA,
  title={Texture Ultrasound Semantic Analysis (TUSA)},
  author={Grutman, Tal and Shinar, Carmel and Ilovitsh, Tali},
  journal={arXiv preprint arXiv:2602.01444},
  year={2026}
}
Downloads last month
57
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for zeahub/ultrapips-tusa