See our collection for all versions of CLIP.

Run CLIP with Keras 3: JAX, PyTorch, or TensorFlow

GitHub Docs Collection

kerasformers/clip_vit_base_32

Paper: Learning Transferable Visual Models From Natural Language Supervision (arXiv:2103.00020) · HF Papers

CLIP (Contrastive Language-Image Pre-training) is a vision + text dual-encoder trained on (image, caption) pairs with a contrastive loss. Both encoders project to a shared embedding space for zero-shot classification, retrieval, and embeddings.

For more details on the model, please go to the upstream model card.

Pure-Keras 3 conversion of openai/clip-vit-base-patch32 for kerasformers. One implementation runs unmodified on TensorFlow / Torch / JAX.

This is a zero-shot image-text checkpoint (CLIPZeroShotClassify): pass image(s) and text prompts at inference time.

✨ Quick start

import os
os.environ["KERAS_BACKEND"] = "torch"  # or "jax" / "tensorflow"

from kerasformers.models.clip import (
    CLIPProcessor,
    CLIPZeroShotClassify,
)

processor = CLIPProcessor.from_weights("kerasformers/clip_vit_base_32")
model = CLIPZeroShotClassify.from_weights("kerasformers/clip_vit_base_32")

labels = [
    "a photo of a cat",
    "a photo of a dog",
    "a photo of a car",
    "a photo of a living room",
]
inputs = processor(text=labels, image_paths="your_image.jpg")
output = model(
    {
        "images": inputs["images"],
        "token_ids": inputs["input_ids"],
        "padding_mask": inputs["attention_mask"],
    }
)
print(output["image_logits"].shape)

Load any CLIP variant the same way with from_weights("kerasformers/<variant>"):

Variant Hub Notes
clip_vit_base_16 kerasformers/clip_vit_base_16 OpenAI
clip_vit_base_32 kerasformers/clip_vit_base_32 OpenAI
clip_vit_large_14 kerasformers/clip_vit_large_14 OpenAI
clip_vit_large_14_336 kerasformers/clip_vit_large_14_336 OpenAI
clip_vit_g_14 kerasformers/clip_vit_g_14 LAION
clip_vit_bigg_14 kerasformers/clip_vit_bigg_14 LAION

Tips

  • Set KERAS_BACKEND before importing Keras / kerasformers.
  • Prefer Processor.from_weights(...) so image size and tokenizer match the variant.
  • Map processor input_ids / attention_mask to model token_ids / padding_mask.
  • OpenAI variants use quick_gelu; LAION g/G use gelu.
  • See CLIP docs and Loading Weights.
  • Community / upstream safetensors still work via the hf: prefix, e.g. CLIPZeroShotClassify.from_weights("hf:openai/clip-vit-base-patch32").

Special Thanks

A huge thank you to the OpenAI CLIP and LAION authors for creating and releasing these models.

License: MIT.

Downloads last month
30
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kerasformers/clip_vit_base_32

Finetuned
(126)
this model

Collection including kerasformers/clip_vit_base_32

Paper for kerasformers/clip_vit_base_32